@tangle-network/agent-eval 0.180.0 → 0.181.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +60 -0
- package/README.md +119 -159
- package/dist/adapters/http.d.ts +2 -2
- package/dist/{agent-profile-B7yErX0q.d.ts → agent-profile-CivaSsSy.d.ts} +4 -4
- package/dist/{agent-profile-B7yErX0q.d.ts.map → agent-profile-CivaSsSy.d.ts.map} +1 -1
- package/dist/{agent-profile-cell-0gSi5ffD.js → agent-profile-cell-Cv6UA-W_.js} +20 -57
- package/dist/agent-profile-cell-Cv6UA-W_.js.map +1 -0
- package/dist/{agent-profile-cell-CTOZJUuE.d.ts → agent-profile-cell-s__adRnK.d.ts} +3 -3
- package/dist/agent-profile-cell-s__adRnK.d.ts.map +1 -0
- package/dist/analyst/index.d.ts +10 -10
- package/dist/analyst/index.js +3 -3
- package/dist/ast-CP9ae9B0.js +557 -0
- package/dist/ast-CP9ae9B0.js.map +1 -0
- package/dist/ast-hI-vjW6J.d.ts +457 -0
- package/dist/ast-hI-vjW6J.d.ts.map +1 -0
- package/dist/{benchmark-command-D4vpnAdO.js → benchmark-command-B57n9vjz.js} +7 -6
- package/dist/{benchmark-command-D4vpnAdO.js.map → benchmark-command-B57n9vjz.js.map} +1 -1
- package/dist/benchmarks/index.d.ts +4 -4
- package/dist/benchmarks/index.js +3 -3
- package/dist/campaign/index.d.ts +6 -6
- package/dist/campaign/index.js +8 -8
- package/dist/{campaign-BGEurASO.js → campaign-4_ppJW5X.js} +12 -12
- package/dist/{campaign-BGEurASO.js.map → campaign-4_ppJW5X.js.map} +1 -1
- package/dist/{campaign-evidence-D8DBLqLI.js → campaign-evidence-B8oF9xQ6.js} +515 -471
- package/dist/campaign-evidence-B8oF9xQ6.js.map +1 -0
- package/dist/cli.js +4 -7
- package/dist/cli.js.map +1 -1
- package/dist/{client-BlLY6o2w.js → client-CXE-U1SA.js} +3 -1
- package/dist/client-CXE-U1SA.js.map +1 -0
- package/dist/{client-CuQgX33c.d.ts → client-kh2jOjTK.d.ts} +4 -4
- package/dist/{client-CuQgX33c.d.ts.map → client-kh2jOjTK.d.ts.map} +1 -1
- package/dist/contract/index.d.ts +13 -13
- package/dist/contract/index.js +10 -9
- package/dist/contract/index.js.map +1 -1
- package/dist/{default-registry-IGDE9XIC.d.ts → default-registry-BwDSWVzg.d.ts} +6 -6
- package/dist/{default-registry-IGDE9XIC.d.ts.map → default-registry-BwDSWVzg.d.ts.map} +1 -1
- package/dist/{define-agent-eval-Cx4Ls9ta.d.ts → define-agent-eval-CwOWWQt_.d.ts} +33 -12
- package/dist/define-agent-eval-CwOWWQt_.d.ts.map +1 -0
- package/dist/{define-agent-eval-Dzidv34q.js → define-agent-eval-Ddu33JH9.js} +134 -67
- package/dist/define-agent-eval-Ddu33JH9.js.map +1 -0
- package/dist/{dspy-rlm-engine-xKiWmj_G.js → dspy-rlm-engine-S53V0HhE.js} +2 -2
- package/dist/{dspy-rlm-engine-xKiWmj_G.js.map → dspy-rlm-engine-S53V0HhE.js.map} +1 -1
- package/dist/{engine-CX8ReXkn.d.ts → engine-DS1cysJy.d.ts} +10 -7
- package/dist/engine-DS1cysJy.d.ts.map +1 -0
- package/dist/{eval-campaign-Cs-7MiCs.js → eval-campaign-aYdtjtJR.js} +4 -4
- package/dist/{eval-campaign-Cs-7MiCs.js.map → eval-campaign-aYdtjtJR.js.map} +1 -1
- package/dist/{exact-types-B7LC1EyX.d.ts → exact-types-BZDe0W2D.d.ts} +2 -2
- package/dist/{exact-types-B7LC1EyX.d.ts.map → exact-types-BZDe0W2D.d.ts.map} +1 -1
- package/dist/experiment/index.d.ts +27 -477
- package/dist/experiment/index.d.ts.map +1 -1
- package/dist/experiment/index.js +95 -559
- package/dist/experiment/index.js.map +1 -1
- package/dist/{experiment-tracker-B3TiF5-u.d.ts → experiment-tracker-C7PfnF4b.d.ts} +2 -2
- package/dist/{experiment-tracker-B3TiF5-u.d.ts.map → experiment-tracker-C7PfnF4b.d.ts.map} +1 -1
- package/dist/{external-optimizer-process-Dlz8YxrT.js → external-optimizer-process-QDRURJAM.js} +3 -3
- package/dist/{external-optimizer-process-Dlz8YxrT.js.map → external-optimizer-process-QDRURJAM.js.map} +1 -1
- package/dist/{external-optimizer-subprocess-q3VzlGAO.js → external-optimizer-subprocess-D4dzUBZI.js} +3 -2
- package/dist/{external-optimizer-subprocess-q3VzlGAO.js.map → external-optimizer-subprocess-D4dzUBZI.js.map} +1 -1
- package/dist/{feedback-trajectory-eHWNv5Aj.d.ts → feedback-trajectory-CXmtITBo.d.ts} +3 -3
- package/dist/{feedback-trajectory-eHWNv5Aj.d.ts.map → feedback-trajectory-CXmtITBo.d.ts.map} +1 -1
- package/dist/hosted/index.d.ts +2 -2
- package/dist/hosted/index.d.ts.map +1 -1
- package/dist/hosted/index.js +1 -1
- package/dist/{index-BxWvILU8.d.ts → index-Bp_6sj3x.d.ts} +109 -56
- package/dist/index-Bp_6sj3x.d.ts.map +1 -0
- package/dist/{index-e7LXeRVa.d.ts → index-CJ3LhKIX.d.ts} +2 -2
- package/dist/{index-e7LXeRVa.d.ts.map → index-CJ3LhKIX.d.ts.map} +1 -1
- package/dist/{index-CiUjjEIa.d.ts → index-DNntP4ch.d.ts} +7 -7
- package/dist/{index-CiUjjEIa.d.ts.map → index-DNntP4ch.d.ts.map} +1 -1
- package/dist/{index-DxNYmx4a.d.ts → index-DoykkxW0.d.ts} +11 -11
- package/dist/{index-DxNYmx4a.d.ts.map → index-DoykkxW0.d.ts.map} +1 -1
- package/dist/index.d.ts +28 -28
- package/dist/index.js +24 -15
- package/dist/index.js.map +1 -1
- package/dist/{insight-report-DETqPc_A.d.ts → insight-report-D1qa0HWs.d.ts} +9 -5
- package/dist/{insight-report-DETqPc_A.d.ts.map → insight-report-D1qa0HWs.d.ts.map} +1 -1
- package/dist/{integrity-BKTcA-HP.d.ts → integrity-rGOfSUle.d.ts} +2 -2
- package/dist/{integrity-BKTcA-HP.d.ts.map → integrity-rGOfSUle.d.ts.map} +1 -1
- package/dist/{ledger-core-Cs9f7385.js → journal-Cs9f7385.js} +1 -1
- package/dist/journal-Cs9f7385.js.map +1 -0
- package/dist/{judge-calibration-C5CbMYce.d.ts → judge-calibration-DFtEMlde.d.ts} +31 -2
- package/dist/judge-calibration-DFtEMlde.d.ts.map +1 -0
- package/dist/{judge-calibration-BnpVKtnb.js → judge-calibration-DYmaBtJr.js} +48 -2
- package/dist/{judge-calibration-BnpVKtnb.js.map → judge-calibration-DYmaBtJr.js.map} +1 -1
- package/dist/ledger-core/index.d.ts +1 -1
- package/dist/ledger-core/index.js +1 -1
- package/dist/{llm-judge-v80Kmu9g.js → llm-judge-DEFZeSiu.js} +645 -456
- package/dist/llm-judge-DEFZeSiu.js.map +1 -0
- package/dist/{matrix-DeMmnWrP.d.ts → matrix-CyhW-vgJ.d.ts} +2 -2
- package/dist/{matrix-DeMmnWrP.d.ts.map → matrix-CyhW-vgJ.d.ts.map} +1 -1
- package/dist/meta-eval/index.d.ts +138 -7
- package/dist/meta-eval/index.d.ts.map +1 -1
- package/dist/meta-eval/index.js +245 -97
- package/dist/meta-eval/index.js.map +1 -1
- package/dist/{mint-Cc1_zwRQ.js → mint-ySIIkKlV.js} +2 -2
- package/dist/{mint-Cc1_zwRQ.js.map → mint-ySIIkKlV.js.map} +1 -1
- package/dist/multishot/golden/index.d.ts +1 -1
- package/dist/multishot/index.d.ts +2 -2
- package/dist/openapi.json +1 -1
- package/dist/outcome-store-BXlkwMPR.js +131 -0
- package/dist/outcome-store-BXlkwMPR.js.map +1 -0
- package/dist/{outcome-store-BYHIuO0e.d.ts → outcome-store-CNt4iZ67.d.ts} +18 -25
- package/dist/outcome-store-CNt4iZ67.d.ts.map +1 -0
- package/dist/{paired-promotion-decision-CGzg0cI_.d.ts → paired-promotion-decision-DPsMQm-0.d.ts} +13 -7
- package/dist/{paired-promotion-decision-CGzg0cI_.d.ts.map → paired-promotion-decision-DPsMQm-0.d.ts.map} +1 -1
- package/dist/pipelines/index.js +1 -1
- package/dist/{produced-state-Cv0kJJuP.js → produced-state-BHboMaab.js} +3 -3
- package/dist/{produced-state-Cv0kJJuP.js.map → produced-state-BHboMaab.js.map} +1 -1
- package/dist/profile-cell.d.ts +1 -1
- package/dist/profile-cell.js +1 -1
- package/dist/{promotion-policy-DWOm70gx.js → promotion-policy-CDMMxzb6.js} +28 -40
- package/dist/promotion-policy-CDMMxzb6.js.map +1 -0
- package/dist/{registry-ByVld1-5.d.ts → registry-BRbB6Y0v.d.ts} +4 -4
- package/dist/{registry-ByVld1-5.d.ts.map → registry-BRbB6Y0v.d.ts.map} +1 -1
- package/dist/{release-confidence-BAcNYOf1.d.ts → release-confidence-BcqeQTHW.d.ts} +3 -3
- package/dist/{release-confidence-BAcNYOf1.d.ts.map → release-confidence-BcqeQTHW.d.ts.map} +1 -1
- package/dist/{release-confidence-BcGCclTB.js → release-confidence-DMg8n18l.js} +2 -2
- package/dist/{release-confidence-BcGCclTB.js.map → release-confidence-DMg8n18l.js.map} +1 -1
- package/dist/reporting.d.ts +4 -4
- package/dist/reporting.js +3 -3
- package/dist/{researcher-jsW1X94L.d.ts → researcher-64T49THL.d.ts} +6 -6
- package/dist/{researcher-jsW1X94L.d.ts.map → researcher-64T49THL.d.ts.map} +1 -1
- package/dist/{reward-hacking-ZXEi9VCq.d.ts → reward-hacking-uzO_ihep.d.ts} +2 -2
- package/dist/{reward-hacking-ZXEi9VCq.d.ts.map → reward-hacking-uzO_ihep.d.ts.map} +1 -1
- package/dist/rl.d.ts +53 -99
- package/dist/rl.d.ts.map +1 -1
- package/dist/rl.js +182 -169
- package/dist/rl.js.map +1 -1
- package/dist/rollout/index.d.ts +1 -1
- package/dist/rollout/index.js +2 -2
- package/dist/{rollout-DmoJVqrF.js → rollout-B-UF5R6w.js} +2 -2
- package/dist/{rollout-DmoJVqrF.js.map → rollout-B-UF5R6w.js.map} +1 -1
- package/dist/rubric-predictive-validity-Bmj2_cll.d.ts +79 -0
- package/dist/rubric-predictive-validity-Bmj2_cll.d.ts.map +1 -0
- package/dist/rubric-predictive-validity-CCK-1B7w.js +178 -0
- package/dist/rubric-predictive-validity-CCK-1B7w.js.map +1 -0
- package/dist/{run-record-DTv1MdjK.d.ts → run-record-BiTWauyO.d.ts} +2 -2
- package/dist/{run-record-DTv1MdjK.d.ts.map → run-record-BiTWauyO.d.ts.map} +1 -1
- package/dist/run-record-Br-Yzt_k.js +464 -0
- package/dist/run-record-Br-Yzt_k.js.map +1 -0
- package/dist/{run-record-DQpSf7t-.js → run-record-DualPTn2.js} +2 -2
- package/dist/{run-record-DQpSf7t-.js.map → run-record-DualPTn2.js.map} +1 -1
- package/dist/{semantic-concept-judge-Bi6_iGqg.js → semantic-concept-judge-Bm5JDEKO.js} +3 -3
- package/dist/{semantic-concept-judge-Bi6_iGqg.js.map → semantic-concept-judge-Bm5JDEKO.js.map} +1 -1
- package/dist/{sequential-B5gXgcyp.js → sequential-DAsyV2T9.js} +42 -25
- package/dist/sequential-DAsyV2T9.js.map +1 -0
- package/dist/{series-convergence-DeG33RpC.d.ts → series-convergence-BnMs_uAr.d.ts} +3 -3
- package/dist/{series-convergence-DeG33RpC.d.ts.map → series-convergence-BnMs_uAr.d.ts.map} +1 -1
- package/dist/{skillopt-optimization-method-C3oYul8v.js → skillopt-optimization-method-CL_0aArC.js} +5 -5
- package/dist/{skillopt-optimization-method-C3oYul8v.js.map → skillopt-optimization-method-CL_0aArC.js.map} +1 -1
- package/dist/{statistical-heldout-0La5ZTlv.d.ts → statistical-heldout-CpVd6FmY.d.ts} +207 -144
- package/dist/statistical-heldout-CpVd6FmY.d.ts.map +1 -0
- package/dist/{store-tool-spans-4J1EDElP.d.ts → store-tool-spans-Dt-YdAuE.d.ts} +6 -6
- package/dist/{store-tool-spans-4J1EDElP.d.ts.map → store-tool-spans-Dt-YdAuE.d.ts.map} +1 -1
- package/dist/{summary-report-gMrbYawB.d.ts → summary-report-D1h4dlrK.d.ts} +3 -3
- package/dist/{summary-report-gMrbYawB.d.ts.map → summary-report-D1h4dlrK.d.ts.map} +1 -1
- package/dist/{summary-report-B16xy9Kd.js → summary-report-e-MaOAHV.js} +2 -2
- package/dist/{summary-report-B16xy9Kd.js.map → summary-report-e-MaOAHV.js.map} +1 -1
- package/dist/{tool-groups-2QA0S7dK.d.ts → tool-groups-B2bSNaJB.d.ts} +3 -3
- package/dist/tool-groups-B2bSNaJB.d.ts.map +1 -0
- package/dist/{tool-waste-B9tdWV6g.js → tool-waste-C7MU9u1e.js} +2 -2
- package/dist/{tool-waste-B9tdWV6g.js.map → tool-waste-C7MU9u1e.js.map} +1 -1
- package/dist/trace-repair/index.d.ts +2 -2
- package/dist/traces.d.ts +6 -6
- package/dist/traces.js +1 -1
- package/dist/{types-BmlkCrg0.d.ts → types-BvZoPTGa.d.ts} +3 -3
- package/dist/{types-BmlkCrg0.d.ts.map → types-BvZoPTGa.d.ts.map} +1 -1
- package/dist/{types-gvRsyJLh.d.ts → types-CBbLtr2J.d.ts} +38 -3
- package/dist/{types-gvRsyJLh.d.ts.map → types-CBbLtr2J.d.ts.map} +1 -1
- package/dist/{types-C34V4Vto.d.ts → types-CS0qk_Yp.d.ts} +4 -4
- package/dist/{types-C34V4Vto.d.ts.map → types-CS0qk_Yp.d.ts.map} +1 -1
- package/dist/{types-DzuaM493.d.ts → types-D7gEdPoQ.d.ts} +3 -3
- package/dist/{types-DzuaM493.d.ts.map → types-D7gEdPoQ.d.ts.map} +1 -1
- package/dist/wire/index.d.ts +2 -2
- package/docs/adapters-observability.md +14 -0
- package/docs/campaign-proposers.md +86 -128
- package/docs/charter.md +108 -112
- package/docs/concepts.md +157 -69
- package/docs/design/mlbenchmarks-book-review.md +440 -0
- package/docs/design/mlbenchmarks-review/observations.json +713 -0
- package/docs/design/mlbenchmarks-review/probes.mts +476 -0
- package/docs/design/mlbenchmarks-review/sources.json +200 -0
- package/docs/design/self-improvement-evidence-audit.md +263 -0
- package/docs/design.md +2 -1
- package/docs/eval-surface-map.md +95 -42
- package/docs/evaluation-integrity.md +220 -0
- package/docs/experiment.md +111 -55
- package/docs/feature-guide.md +5 -6
- package/docs/hosted-ingest-spec.md +4 -11
- package/docs/insight-report.md +187 -455
- package/docs/outcome-validity.md +182 -0
- package/docs/product-eval-adoption.md +1 -2
- package/docs/research-report-methodology.md +7 -7
- package/docs/search-history-receipts.md +8 -0
- package/docs/statistical-evidence.md +129 -0
- package/docs/verdicts.md +76 -49
- package/package.json +1 -1
- package/dist/agent-profile-cell-0gSi5ffD.js.map +0 -1
- package/dist/agent-profile-cell-CTOZJUuE.d.ts.map +0 -1
- package/dist/campaign-evidence-D8DBLqLI.js.map +0 -1
- package/dist/client-BlLY6o2w.js.map +0 -1
- package/dist/define-agent-eval-Cx4Ls9ta.d.ts.map +0 -1
- package/dist/define-agent-eval-Dzidv34q.js.map +0 -1
- package/dist/engine-CX8ReXkn.d.ts.map +0 -1
- package/dist/index-BxWvILU8.d.ts.map +0 -1
- package/dist/judge-calibration-C5CbMYce.d.ts.map +0 -1
- package/dist/ledger-core-Cs9f7385.js.map +0 -1
- package/dist/llm-judge-v80Kmu9g.js.map +0 -1
- package/dist/outcome-store-BYHIuO0e.d.ts.map +0 -1
- package/dist/outcome-store-ChBKlTd_.js +0 -75
- package/dist/outcome-store-ChBKlTd_.js.map +0 -1
- package/dist/promotion-policy-DWOm70gx.js.map +0 -1
- package/dist/rubric-predictive-validity-2D5Gw9z9.js +0 -131
- package/dist/rubric-predictive-validity-2D5Gw9z9.js.map +0 -1
- package/dist/rubric-predictive-validity-Dl1dvKCv.d.ts +0 -75
- package/dist/rubric-predictive-validity-Dl1dvKCv.d.ts.map +0 -1
- package/dist/run-record-CR63CpHK.js +0 -216
- package/dist/run-record-CR63CpHK.js.map +0 -1
- package/dist/sequential-B5gXgcyp.js.map +0 -1
- package/dist/statistical-heldout-0La5ZTlv.d.ts.map +0 -1
- package/dist/tool-groups-2QA0S7dK.d.ts.map +0 -1
package/docs/concepts.md
CHANGED
|
@@ -2,52 +2,50 @@
|
|
|
2
2
|
|
|
3
3
|
`agent-eval` records agent runs, scores their outputs, compares variants, and applies caller-defined release rules.
|
|
4
4
|
|
|
5
|
-
|
|
5
|
+
An agent can claim success while a build, browser flow, or integration fails.
|
|
6
|
+
Required source evidence can also be missing.
|
|
6
7
|
This package lets code, model judges, and human feedback check those outcomes through the same run format.
|
|
7
8
|
|
|
8
9
|
## The top-level functions
|
|
9
10
|
|
|
10
|
-
Start with
|
|
11
|
-
|
|
11
|
+
Start with `defineAgentEval()` from `/contract` for one agent, judge, case set, and baseline surface.
|
|
12
|
+
Its `evaluate()` method returns campaign measurements.
|
|
13
|
+
Its `improve()` method searches and returns a final comparison with a release decision.
|
|
14
|
+
Use `selfImprove()` directly when you do not need shared configuration.
|
|
12
15
|
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
| **`selfImprove()`** | You want candidate generation, scoring, and a release decision in one call. | scenarios, agent, judge, baseline surface | report, winner surface, and a `gateDecision` (see below) |
|
|
17
|
-
| **`loadEvalFixtureScenarios()`** | You want agents to add evals as folders with `PROMPT.md`, checks, and starter files. | `evals/<name>/PROMPT.md + EVAL.ts + package.json` | `Scenario[]` that runs through `runCampaign`; pair with `planEvalFixtureRun()` before spending tokens |
|
|
18
|
-
| **`analyzeRuns()`** | You have existing runs and do not need to invoke an agent. | `RunRecord[]` and options | `InsightReport` |
|
|
19
|
-
| **Intake adapters** (`fromFeedbackTable`, `fromOtelSpans`) | Your data isn't already in `RunRecord` shape: it's in Obsidian, Sheets, an OTel collector, etc. | source-specific input | `RunRecord[]` ready to pipe into `analyzeRuns()` |
|
|
20
|
-
| **`sealExperiment()` / `openSealedExperiment()`** | The result must convince a reader who does not trust you, so the rules must be fixed before the data arrives. | arms, admission funnel, estimand, interval, decision table | a hashed rule tree plus executors that can run no other rule ([`experiment.md`](./experiment.md)) |
|
|
21
|
-
| **`runEquivalenceCheck()`** | The work has no held-out test suite, so no answer key exists to grade against. | a claim, two blind arms, an injected checker | a certification naming who vouched and how it can fail ([`verification-strategies.md`](./verification-strategies.md)) |
|
|
22
|
-
| **`AnalystRegistry.runExact()`** | A batch of runs failed and you need cited findings, with the caller owning every execution choice. | recorded evidence, a declared analyst list | findings with evidence references, an execution plan, and a receipt ([`trace-analysis.md`](./trace-analysis.md)) |
|
|
16
|
+
Use `analyzeRuns()` from `/contract` for existing `RunRecord[]` evidence.
|
|
17
|
+
[The README workflows](../README.md#choose-a-workflow) link to runnable examples.
|
|
18
|
+
[The surface map](./eval-surface-map.md) lists lower-level execution and analysis APIs.
|
|
23
19
|
|
|
24
|
-
|
|
25
|
-
The
|
|
20
|
+
Root `Scenario`, `JudgeScore`, and `GateDecision` use the same definitions as `/contract` and `/campaign`.
|
|
21
|
+
The product-judging shapes have explicit root names: `ProductScenario` and `DimensionJudgeScore`.
|
|
22
|
+
The separate `HeldOutGate` class returns `HeldOutGateDecision` over `RunRecord` comparisons.
|
|
26
23
|
|
|
27
24
|
### The five release decisions
|
|
28
25
|
|
|
29
|
-
`selfImprove()`
|
|
30
|
-
|
|
26
|
+
`selfImprove()` returns a `gateDecision` from the campaign `GateDecision` union.
|
|
27
|
+
Keep its five values distinct because they require different actions.
|
|
31
28
|
|
|
32
29
|
| Decision | What it means | What to do next |
|
|
33
30
|
|---|---|---|
|
|
34
|
-
| `ship` |
|
|
35
|
-
| `hold` |
|
|
36
|
-
| `need_more_work` |
|
|
31
|
+
| `ship` | All required configured checks support release. | Review the evidence and release the candidate. |
|
|
32
|
+
| `hold` | The gate does not justify release. A required check can fail or lack sufficient evidence. | Inspect the contributions to distinguish regression from an unresolved comparison. |
|
|
33
|
+
| `need_more_work` | The gate reports that more work or evidence is required. | Address the reported gap before another decision. |
|
|
37
34
|
| `model_ceiling` | Reserved for a caller-supplied gate that attributes the limit to the model. | Handle it; no gate in this package emits it. |
|
|
38
35
|
| `arch_ceiling` | Reserved for a caller-supplied gate that attributes the limit to the architecture. | Handle it; no gate in this package emits it. |
|
|
39
36
|
|
|
40
|
-
|
|
41
|
-
Handle all five anyway: a caller's own gate may return either, and the type will not let you ignore them.
|
|
37
|
+
Handle all five values when accepting caller-supplied gates.
|
|
42
38
|
|
|
43
|
-
|
|
44
|
-
|
|
39
|
+
Read the contributing checks before interpreting a refused release.
|
|
40
|
+
An unresolved comparison does not establish that the candidate is worse.
|
|
41
|
+
Absent optional checks remain `not_evaluated`, including when the required checks support `ship`.
|
|
45
42
|
|
|
46
43
|
When gates are composed, `ship` requires every gate to ship.
|
|
47
44
|
Otherwise the strongest hold wins, in this order: `arch_ceiling`, `model_ceiling`, `hold`, `need_more_work`.
|
|
48
45
|
|
|
49
|
-
`analyzeRuns()`
|
|
50
|
-
|
|
46
|
+
`analyzeRuns()` returns an `InsightReport`; `selfImprove()` includes one in its result.
|
|
47
|
+
The report includes score distributions, cost, and recommendations.
|
|
48
|
+
Paired lift, failure clusters, contamination checks, and outcome associations require their corresponding inputs.
|
|
51
49
|
[`insight-report.md`](./insight-report.md) defines every field.
|
|
52
50
|
|
|
53
51
|
## Package Boundary
|
|
@@ -68,7 +66,8 @@ Use the profile improvement functions from `/contract` when a host owns immutabl
|
|
|
68
66
|
|
|
69
67
|
This API never activates a candidate or runs an agent itself.
|
|
70
68
|
The host owns authorization, billing, task isolation, profile materialization, execution, and durable evidence.
|
|
71
|
-
The
|
|
69
|
+
The portable profile contract accepts prompt and skill changes.
|
|
70
|
+
A host needs an adapter for exact state before measuring tools, MCP servers, hooks, subagents, or external knowledge.
|
|
72
71
|
|
|
73
72
|
## Main Objects
|
|
74
73
|
|
|
@@ -85,7 +84,7 @@ Traces, datasets, optimization, statistics, and reports build on these objects.
|
|
|
85
84
|
|
|
86
85
|
Every entry in `GateResult.contributingGates` has a `status` of `pass`, `fail`, or `not_evaluated`.
|
|
87
86
|
`pass` and `fail` mean the check ran with sufficient input.
|
|
88
|
-
`not_evaluated` means the check lacked
|
|
87
|
+
`not_evaluated` means the check was unconfigured or lacked required input or evidence.
|
|
89
88
|
`defaultProductionGate` always requires held-out significance.
|
|
90
89
|
Its other checks are optional until their input is configured or their name is included in `requiredChecks`.
|
|
91
90
|
A required check with missing or insufficient evidence remains `not_evaluated` and holds the release decision.
|
|
@@ -93,6 +92,29 @@ An absent optional check records `not_evaluated` and never appears as a successf
|
|
|
93
92
|
Run history is shared input only.
|
|
94
93
|
Enable reward-hacking and canary monitoring independently with `rewardHacking` and `canary`.
|
|
95
94
|
|
|
95
|
+
`selfImprove()` uses `defaultProductionGate()` unless you pass a custom `gate`.
|
|
96
|
+
The optional `paretoSignificanceGate()` applies each objective's regression floor to its deciding confidence interval.
|
|
97
|
+
Its confidence level applies per objective; the gate does not adjust for multiple objectives or repeated comparisons.
|
|
98
|
+
It pairs execution cells directly; custom gates must apply any grouping required by a claim.
|
|
99
|
+
For detected binary outcomes, that interval accounts for uncertainty even when every observed pair agrees.
|
|
100
|
+
At 95% confidence, 20 matching all-positive binary pairs leave approximately 16 percentage points of uncertainty in either direction.
|
|
101
|
+
A declared five-point regression tolerance therefore holds that candidate; 100 matching pairs narrow the interval enough to clear that floor.
|
|
102
|
+
Another objective must still show a significant gain before promotion.
|
|
103
|
+
An axis labeled `regressed` has not cleared its configured floor.
|
|
104
|
+
Its interval can permit a regression without demonstrating one.
|
|
105
|
+
An interval with zero width or non-finite bounds produces an `indeterminate` axis verdict and a `not_evaluated` check.
|
|
106
|
+
If no other axis breaches its regression floor, the gate returns `need_more_work`.
|
|
107
|
+
This includes undeclared all-zero outcomes, whose binary scale cannot be inferred, and continuous observations with constant paired differences.
|
|
108
|
+
Additional identical observations do not resolve an unknown outcome scale or a collapsed bootstrap interval.
|
|
109
|
+
Consumers with custom policies must handle `indeterminate` as unresolved evidence.
|
|
110
|
+
|
|
111
|
+
Set an objective's `binaryScale: 1` for known `{0, 1}` observations, including all-zero error indicators.
|
|
112
|
+
Use `binaryScale: 100` for `{0, 100}` observations; its default regression tolerance is 5.
|
|
113
|
+
The shared `decidePairedPromotion()` function accepts the same declaration.
|
|
114
|
+
The scale must be positive and finite, and every paired cell score must be zero or that scale.
|
|
115
|
+
Declared binary outcomes use the risk-difference mean and reject `statistic: 'median'`.
|
|
116
|
+
The declaration identifies the outcome support; it does not reduce sample requirements or supply missing observations.
|
|
117
|
+
|
|
96
118
|
When the thing being evaluated is an agent that should keep working, use
|
|
97
119
|
[`runAgentControlLoop`](./control-runtime.md). It turns validators into a
|
|
98
120
|
runtime loop: observe typed state, validate it, decide the next action, act,
|
|
@@ -115,19 +137,19 @@ that can seed memory, replay scenarios, and optimization.
|
|
|
115
137
|
| **Layer** | One stage of a verifier pipeline (install, typecheck, build, semantic, …). |
|
|
116
138
|
| **Finding** | A specific issue a judge found: file, line, severity, message. |
|
|
117
139
|
| **Trace store** | The append-only log of every span/event during a run. Replay = read this back. |
|
|
118
|
-
| **Composite score** |
|
|
119
|
-
| **Rubric version** | A stable hash of the rubric.
|
|
140
|
+
| **Composite score** | An aggregate on the judge's declared scale. Gates must use thresholds on that scale. |
|
|
141
|
+
| **Rubric version** | A stable hash of the rubric. Comparing revisions requires calibration against shared independent labels. |
|
|
120
142
|
|
|
121
143
|
### Running an evaluation
|
|
122
144
|
|
|
123
145
|
| Term | Plain English |
|
|
124
146
|
|---|---|
|
|
125
|
-
| **Case** (`Scenario`) | One task the agent must do.
|
|
147
|
+
| **Case** (`Scenario`) | One task the agent must do. Variants can share an independent source unit. |
|
|
126
148
|
| **Surface** | The value being changed: a prompt, a skill, or a serialized configuration. |
|
|
127
149
|
| **Dispatch** | The function that runs your agent on one case and returns the artifact. |
|
|
128
150
|
| **Campaign** | One complete pass of every case, executed, scored, and cached under a run directory. |
|
|
129
|
-
| **Cell** | One (case × replicate) of a campaign.
|
|
130
|
-
| **Receipt** |
|
|
151
|
+
| **Cell** | One (case × replicate) of a campaign. With caching enabled, matching completed cells can be reused. |
|
|
152
|
+
| **Receipt** | A settled call record with cost, token usage, and flags for unknown measurements. |
|
|
131
153
|
| **Cost ledger** | The spend account receipts are written to. A capped ledger refuses a call that would exceed the cap. |
|
|
132
154
|
| **Provenance** | Where a number came from: the package version, the source revision, the run identity, the exact attempt. |
|
|
133
155
|
| **`RunRecord`** | The analysis-time projection of one run: who ran, on what, with which seed, at what cost, and what it scored. |
|
|
@@ -145,20 +167,52 @@ that can seed memory, replay scenarios, and optimization.
|
|
|
145
167
|
| **Selection cases** | Evidence the optimizer reads to choose among its candidates. |
|
|
146
168
|
| **Final cases** | Held back from the optimizer entirely. They produce the reported lift. |
|
|
147
169
|
|
|
148
|
-
|
|
149
|
-
|
|
170
|
+
Keep scenario identifiers disjoint across the three partitions.
|
|
171
|
+
For new-unit claims or fresh final evidence, also keep source units separate between development and final cases.
|
|
172
|
+
Fixed-roster development can share sources while retaining independent-unit counts in its reports.
|
|
173
|
+
Renamed variants from one incident can leak information across splits.
|
|
174
|
+
A final comparison supports only the declared population and measured conditions.
|
|
150
175
|
|
|
151
176
|
### Proving a result
|
|
152
177
|
|
|
153
|
-
| Term |
|
|
178
|
+
| Term | Meaning |
|
|
154
179
|
|---|---|
|
|
155
|
-
| **
|
|
156
|
-
| **
|
|
157
|
-
| **
|
|
158
|
-
| **
|
|
159
|
-
| **
|
|
160
|
-
| **
|
|
161
|
-
| **
|
|
180
|
+
| **Claim** | The intended use, population, sampling frame, independent unit, generalization target, and optional minimum useful effect. |
|
|
181
|
+
| **Independent unit** | The source task, incident, or family that contributes one independent observation to an inference. |
|
|
182
|
+
| **Experiment** | Arms, admission, estimand, interval, and decision rules declared before results are inspected. |
|
|
183
|
+
| **Seal** | A digest binding the experiment's rules and claim to the executed specification. |
|
|
184
|
+
| **Estimand** | The quantity being estimated, such as the mean difference across independent task families. |
|
|
185
|
+
| **Funnel** | Counts of input rows, exclusions at each stage, and retained evidence. |
|
|
186
|
+
| **Final-evidence reservation** | A durable claim on source units before search; exposure records the evaluated candidates before dispatch. |
|
|
187
|
+
| **Verification strategy** | A method of checking a result, with documented assumptions and failure modes. |
|
|
188
|
+
| **Certification** | The checker identity, strategy, unverified assumptions, and evidence associated with a verdict. |
|
|
189
|
+
| **Analyst** | A function that reads recorded evidence and returns cited findings. |
|
|
190
|
+
|
|
191
|
+
Use `defineEvaluationClaim()` from `/experiment` to declare what a result can describe.
|
|
192
|
+
Pass it as the top-level `claim` when improving a surface or comparing optimization methods.
|
|
193
|
+
`fixed-roster` concerns the listed units; `new-units` attempts to generalize to further units from the declared population.
|
|
194
|
+
A declared sampling frame does not itself establish representative sampling.
|
|
195
|
+
|
|
196
|
+
Count repetitions separately from independent units.
|
|
197
|
+
For example, 100 retries of one incident produce 100 observations and one independent incident.
|
|
198
|
+
Campaign aggregate `n` describes its observed scores.
|
|
199
|
+
The default improvement gate and method comparisons retain their independent-unit and paired-cell counts.
|
|
200
|
+
|
|
201
|
+
Set `minimumEffect` when the decision concerns a practically useful change.
|
|
202
|
+
Development and absolute-rate claims can omit it.
|
|
203
|
+
Sealed power checks assess the declared effect.
|
|
204
|
+
A design that detects only much larger effects cannot pass that adequacy check.
|
|
205
|
+
Inference also needs the interval and clustering rule to match the claim.
|
|
206
|
+
|
|
207
|
+
Unit-aware comparison does not require a final-evidence ledger.
|
|
208
|
+
For fresh confirmation, add `finalEvidence: { ledger, requestId, evaluatorDigest }` alongside the top-level `claim`.
|
|
209
|
+
This reserves final units before search.
|
|
210
|
+
Use one durable ledger across related campaigns and stable source identities across renamed variants.
|
|
211
|
+
Exposure remains recorded if measurement fails or the process stops.
|
|
212
|
+
The host enforces access isolation; the ledger cannot inspect reads outside this workflow.
|
|
213
|
+
|
|
214
|
+
See [evaluation integrity](./evaluation-integrity.md) for claims, final evidence, and evaluator admission.
|
|
215
|
+
[Registered experiments](./experiment.md) describes seals, decision rules, and refusal artifacts.
|
|
162
216
|
|
|
163
217
|
## The feedback trajectory loop
|
|
164
218
|
|
|
@@ -178,7 +232,8 @@ rows, optimizer rows, and held-out examples for overfit checks.
|
|
|
178
232
|
|
|
179
233
|
## Code Generator Eval
|
|
180
234
|
|
|
181
|
-
|
|
235
|
+
Generated-code evaluations can score the agent session, the build, and the running application.
|
|
236
|
+
Each layer detects different failures:
|
|
182
237
|
|
|
183
238
|
```
|
|
184
239
|
L0 builder Did the agent's session itself work?
|
|
@@ -190,21 +245,22 @@ L1 app-build Does the artifact build / typecheck / test?
|
|
|
190
245
|
│
|
|
191
246
|
▼
|
|
192
247
|
L2 app-runtime Does the artifact actually run end-to-end?
|
|
193
|
-
(Dynamic signal:
|
|
248
|
+
(Dynamic signal: requires a runnable application.)
|
|
194
249
|
```
|
|
195
250
|
|
|
196
|
-
`BuilderSession`
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
|
|
200
|
-
- L1 misses: files exist but typecheck fails. LLM judges can't reliably catch this.
|
|
201
|
-
- L2 misses: code compiles but does the wrong thing at runtime.
|
|
251
|
+
`BuilderSession` coordinates these checks.
|
|
252
|
+
It opens at `startChat`, runs the build at `ship`, and runs the application check at `runAppScenario`.
|
|
253
|
+
Each layer emits a trace span.
|
|
254
|
+
`scoreProject` reports each layer's score and whether the required measurements are complete.
|
|
202
255
|
|
|
203
|
-
|
|
256
|
+
- L0: The agent crashed during generation and left an incomplete artifact.
|
|
257
|
+
- L1: Files exist but do not typecheck or build.
|
|
258
|
+
- L2: Code compiles but behaves incorrectly when executed.
|
|
204
259
|
|
|
205
260
|
## How rubrics work
|
|
206
261
|
|
|
207
262
|
A rubric describes:
|
|
263
|
+
|
|
208
264
|
1. **Dimensions**: the axes you score on (e.g. `buyer_quality`, `voice`, `signal`).
|
|
209
265
|
2. **Weights**: how to combine dimensions into a composite (`0.5 * buyer_quality + 0.3 * voice + 0.2 * signal`).
|
|
210
266
|
3. **Failure modes**: named patterns the judge looks for ("ai-cadence", "vague-claim").
|
|
@@ -214,7 +270,10 @@ A rubric describes:
|
|
|
214
270
|
Built-in rubrics ship in `src/wire/rubrics.ts`, including `anti-slop` for technical-buyer voice.
|
|
215
271
|
You can also pass the same rubric shape inline at the call site.
|
|
216
272
|
|
|
217
|
-
A rubric is plain data.
|
|
273
|
+
A rubric is plain data.
|
|
274
|
+
Its digest and encoding scheme identify the `rubricVersion`.
|
|
275
|
+
Changing the rubric starts a new comparison series.
|
|
276
|
+
Evaluate rubric revisions against independent labels before combining their scores.
|
|
218
277
|
|
|
219
278
|
## How verifiers work
|
|
220
279
|
|
|
@@ -229,7 +288,7 @@ const verifier = new MultiLayerVerifier([
|
|
|
229
288
|
])
|
|
230
289
|
|
|
231
290
|
const report = await verifier.run({ env })
|
|
232
|
-
report.allPass //
|
|
291
|
+
report.allPass // every layer passed and the task measurement is complete
|
|
233
292
|
report.taskScore // complete task score, or undefined
|
|
234
293
|
report.blendedScore // diagnostic weighted aggregate, possibly partial
|
|
235
294
|
report.layers // per-layer status, findings, duration
|
|
@@ -238,21 +297,27 @@ report.layers // per-layer status, findings, duration
|
|
|
238
297
|
`env` carries the sandbox driver, the working directory, and the harness commands each layer runs.
|
|
239
298
|
|
|
240
299
|
Use `taskScore` when creating task labels or training data.
|
|
241
|
-
|
|
300
|
+
A complete scoring panel needs at least one valid score from a positive-weight layer.
|
|
301
|
+
Positive-weight layers that error, time out, or skip leave `taskScore` undefined.
|
|
302
|
+
Passing layers can omit a numeric score.
|
|
303
|
+
A failed layer contributes only when `failContributesToScore` is enabled and it supplies a valid score.
|
|
304
|
+
Zero-weight layers can still fail `allPass` without removing `taskScore`.
|
|
242
305
|
Use `blendedScore` only to inspect the measurements that did complete.
|
|
243
306
|
|
|
244
307
|
Two rules that will save you bugs:
|
|
245
308
|
|
|
246
|
-
1.
|
|
247
|
-
|
|
248
|
-
2.
|
|
309
|
+
1. Run build checks and structural assertions.
|
|
310
|
+
They detect different failures.
|
|
311
|
+
2. Preserve a failed build as a deterministic release failure.
|
|
312
|
+
A semantic score cannot override it.
|
|
249
313
|
|
|
250
314
|
## Judge calibration
|
|
251
315
|
|
|
252
|
-
|
|
316
|
+
Compare a judge with independent labels and other judges before using its scores for decisions:
|
|
253
317
|
|
|
254
|
-
1. **Does it agree with humans?** `calibrateJudge(golden, candidate)` reports Pearson, MAE, integer-rounded κ, and
|
|
255
|
-
2. **Does it agree with
|
|
318
|
+
1. **Does it agree with humans?** `calibrateJudge(golden, candidate)` reports Pearson, MAE, integer-rounded κ, and the five largest errors on matched item IDs.
|
|
319
|
+
2. **Does it agree with other judges?**
|
|
320
|
+
`continuousAgreement()` and `calibrateJudgeContinuous()` report agreement and bootstrap intervals on continuous scores.
|
|
256
321
|
|
|
257
322
|
Each statistic answers a different question:
|
|
258
323
|
|
|
@@ -262,7 +327,7 @@ Each statistic answers a different question:
|
|
|
262
327
|
| Spearman | Do they rank the same way? | The size of any gap |
|
|
263
328
|
| MAE (mean absolute error) | How far apart are they, on average? | Whether the gap is systematic |
|
|
264
329
|
| κ (Cohen's kappa) | Do they agree more than chance? | Everything below the rounding step |
|
|
265
|
-
| ICC(2,1) | Do
|
|
330
|
+
| ICC(2,1) | Do raters agree in absolute score under its variance model? | Shared errors against the intended outcome |
|
|
266
331
|
|
|
267
332
|
Use two flavours of κ for one reason.
|
|
268
333
|
`calibrateJudge` rounds each score to an integer first.
|
|
@@ -273,15 +338,36 @@ ICC(2,1) catches a bias Pearson cannot see.
|
|
|
273
338
|
If judge B always scores twice judge A, the two move together perfectly and Pearson stays near 1, while ICC drops.
|
|
274
339
|
That drop is the signal.
|
|
275
340
|
|
|
276
|
-
|
|
341
|
+
ICC and continuous weighted κ have percentile bootstrap intervals, with `ciLevel: 0.95` by default.
|
|
342
|
+
Pearson, Spearman, and MAE are point estimates in these reports.
|
|
343
|
+
|
|
344
|
+
Import calibration and bias functions from `/meta-eval`.
|
|
345
|
+
|
|
346
|
+
| Probe | Input | Observation |
|
|
347
|
+
|---|---|---|
|
|
348
|
+
| `positionalBias()` | The same items judged with their presentation order swapped. | Mean paired score difference by position. |
|
|
349
|
+
| `verbosityBias()` | Output lengths and judge scores. | Correlation between length and score. |
|
|
350
|
+
| `selfPreference()` | Scores grouped by whether judge and output share a model family. | Difference between the group means. |
|
|
351
|
+
|
|
352
|
+
These probes are descriptive diagnostics.
|
|
353
|
+
Length and family groups can also differ in task quality; an observed association alone does not isolate bias.
|
|
354
|
+
Inspect sample counts before interpreting a diagnostic, especially `n: 0`.
|
|
355
|
+
|
|
356
|
+
Use `auditEvaluator()` for admission against predeclared false-acceptance and false-rejection limits.
|
|
357
|
+
Its observation records distinguish fresh controls, development exposure, and unknown judgments.
|
|
358
|
+
It counts source families rather than repeated variants and reports simultaneous exact bounds for both error rates.
|
|
359
|
+
The host must enforce independent authorship and control access.
|
|
277
360
|
|
|
278
|
-
`
|
|
279
|
-
|
|
280
|
-
|
|
361
|
+
Use `rubricPredictiveValidity()` to compare rubric scores with declared deployment outcomes.
|
|
362
|
+
Specify whether each outcome should increase or decrease.
|
|
363
|
+
The report preserves signed associations, direction-aligned associations, and exclusions.
|
|
364
|
+
An `inverse` association is a reason to investigate; it does not prove that reversing a rubric will improve behavior.
|
|
365
|
+
See [outcome validity](./outcome-validity.md).
|
|
281
366
|
|
|
282
367
|
## Trace Model
|
|
283
368
|
|
|
284
|
-
|
|
369
|
+
Instrumented execution writes structured spans into a `TraceStore`.
|
|
370
|
+
A builder run can have this tree:
|
|
285
371
|
|
|
286
372
|
```
|
|
287
373
|
builder-session [span]
|
|
@@ -294,7 +380,9 @@ builder-session [span]
|
|
|
294
380
|
└── scenario.run [span]
|
|
295
381
|
```
|
|
296
382
|
|
|
297
|
-
|
|
383
|
+
Recorded spans preserve their identifiers and relationships.
|
|
384
|
+
Trace inspection reads this evidence; executable replay separately reruns recorded operations.
|
|
385
|
+
OTLP export sends spans to distributed tracing systems.
|
|
298
386
|
|
|
299
387
|
You usually should not build this tree by hand. Product runtimes,
|
|
300
388
|
`runAgentControlLoop`, harnesses, and verifiers should emit it while they run.
|
|
@@ -317,4 +405,4 @@ release decision.
|
|
|
317
405
|
- **Certifying a result with no answer key?** Read [verification-strategies.md](./verification-strategies.md) for the ten-member family and the blind two-arm protocol.
|
|
318
406
|
- **Reading a verdict someone else produced?** Read [verdicts.md](./verdicts.md) for what `certification` carries and what an absent one means.
|
|
319
407
|
- **Grading a finding by executing its repair?** Read [trace-repair-grader.md](./trace-repair-grader.md), and [trajectory-replay.md](./trajectory-replay.md) for re-executing a recorded failure.
|
|
320
|
-
- **
|
|
408
|
+
- **Checking package ownership?** Read [charter.md](./charter.md) for the implemented foundations and host responsibilities.
|