@tangle-network/agent-eval 0.124.0 → 0.126.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +60 -35
- package/README.md +270 -189
- package/dist/analyst/index.d.ts +15 -145
- package/dist/analyst/index.js +33 -47
- package/dist/analyst/index.js.map +1 -1
- package/dist/benchmarks/index.d.ts +45 -162
- package/dist/benchmarks/index.js +8 -9
- package/dist/campaign/index.d.ts +3655 -5365
- package/dist/campaign/index.js +21 -95
- package/dist/{chunk-R226UZOI.js → chunk-474LBSOX.js} +2 -2
- package/dist/{chunk-W5B3ZGP3.js → chunk-4B7ZZHPX.js} +8 -6
- package/dist/{chunk-W5B3ZGP3.js.map → chunk-4B7ZZHPX.js.map} +1 -1
- package/dist/{chunk-DT7OXY3C.js → chunk-CM4OILD2.js} +535 -846
- package/dist/chunk-CM4OILD2.js.map +1 -0
- package/dist/{chunk-HM6V7F3M.js → chunk-FO7HEH76.js} +3 -3
- package/dist/chunk-IILEIWGW.js +635 -0
- package/dist/chunk-IILEIWGW.js.map +1 -0
- package/dist/{chunk-EQUK3RFS.js → chunk-J5SQWP6Y.js} +8 -5
- package/dist/chunk-J5SQWP6Y.js.map +1 -0
- package/dist/chunk-KO2PZOGP.js +4637 -0
- package/dist/chunk-KO2PZOGP.js.map +1 -0
- package/dist/{chunk-4Y7AAATF.js → chunk-LKKT3IVV.js} +574 -81
- package/dist/chunk-LKKT3IVV.js.map +1 -0
- package/dist/chunk-M7AH34KV.js +155 -0
- package/dist/chunk-M7AH34KV.js.map +1 -0
- package/dist/chunk-NTOV7RU5.js +7152 -0
- package/dist/chunk-NTOV7RU5.js.map +1 -0
- package/dist/{chunk-QFQZ3U3X.js → chunk-OCFJACJU.js} +2 -2
- package/dist/{chunk-GID26AN4.js → chunk-P22LJ3Y2.js} +4 -6
- package/dist/{chunk-GID26AN4.js.map → chunk-P22LJ3Y2.js.map} +1 -1
- package/dist/{chunk-SJT4OBVL.js → chunk-SDPM6554.js} +3 -3
- package/dist/{chunk-D5JZ7UDZ.js → chunk-UCLVDLCH.js} +136 -50
- package/dist/chunk-UCLVDLCH.js.map +1 -0
- package/dist/chunk-UI4YMIN2.js +105 -0
- package/dist/chunk-UI4YMIN2.js.map +1 -0
- package/dist/chunk-VBQ3CRKH.js +291 -0
- package/dist/chunk-VBQ3CRKH.js.map +1 -0
- package/dist/{chunk-JKDNAOF5.js → chunk-W4L6C2XT.js} +2 -2
- package/dist/{chunk-GRCDRKII.js → chunk-WS3NZZQQ.js} +58 -20
- package/dist/chunk-WS3NZZQQ.js.map +1 -0
- package/dist/cli.js +3 -3
- package/dist/contract/index.d.ts +3221 -3094
- package/dist/contract/index.js +173 -42
- package/dist/contract/index.js.map +1 -1
- package/dist/control.js +2 -3
- package/dist/fuzz.d.ts +14 -1
- package/dist/fuzz.js +1 -1
- package/dist/hosted/index.d.ts +8 -1
- package/dist/index.d.ts +208 -690
- package/dist/index.js +185 -500
- package/dist/index.js.map +1 -1
- package/dist/openapi.json +1 -1
- package/dist/rl.d.ts +5 -100
- package/dist/rl.js +4 -5
- package/dist/rl.js.map +1 -1
- package/dist/rollout/index.d.ts +9 -1
- package/dist/rollout/index.js +6 -6
- package/dist/{run-campaign-I3JXKVAK.js → run-campaign-LVFKZCEU.js} +3 -3
- package/dist/supervisor-run/index.d.ts +156 -4
- package/dist/supervisor-run/index.js +14 -2
- package/dist/traces.js +2 -3
- package/dist/wire/index.d.ts +14 -1
- package/dist/wire/index.js +3 -3
- package/docs/campaign-proposers.md +363 -168
- package/docs/design/loop-taxonomy.md +142 -190
- package/docs/design.md +1 -1
- package/docs/distributed-driver.md +8 -11
- package/docs/feature-guide.md +20 -19
- package/docs/knowledge-readiness.md +2 -5
- package/docs/multi-shot-optimization.md +35 -27
- package/docs/rollout.md +5 -5
- package/package.json +4 -4
- package/dist/chunk-4Y7AAATF.js.map +0 -1
- package/dist/chunk-5PVZVCZB.js +0 -9190
- package/dist/chunk-5PVZVCZB.js.map +0 -1
- package/dist/chunk-A6GT67HT.js +0 -550
- package/dist/chunk-A6GT67HT.js.map +0 -1
- package/dist/chunk-D5JZ7UDZ.js.map +0 -1
- package/dist/chunk-DT7OXY3C.js.map +0 -1
- package/dist/chunk-EQUK3RFS.js.map +0 -1
- package/dist/chunk-GC4ATIKK.js +0 -317
- package/dist/chunk-GC4ATIKK.js.map +0 -1
- package/dist/chunk-GRCDRKII.js.map +0 -1
- package/dist/chunk-LOW3U7JZ.js +0 -328
- package/dist/chunk-LOW3U7JZ.js.map +0 -1
- package/dist/chunk-MGGFVCJ7.js +0 -288
- package/dist/chunk-MGGFVCJ7.js.map +0 -1
- package/dist/chunk-PMITBABE.js +0 -3841
- package/dist/chunk-PMITBABE.js.map +0 -1
- package/dist/chunk-R7ZRE2KV.js +0 -138
- package/dist/chunk-R7ZRE2KV.js.map +0 -1
- /package/dist/{chunk-R226UZOI.js.map → chunk-474LBSOX.js.map} +0 -0
- /package/dist/{chunk-HM6V7F3M.js.map → chunk-FO7HEH76.js.map} +0 -0
- /package/dist/{chunk-QFQZ3U3X.js.map → chunk-OCFJACJU.js.map} +0 -0
- /package/dist/{chunk-SJT4OBVL.js.map → chunk-SDPM6554.js.map} +0 -0
- /package/dist/{chunk-JKDNAOF5.js.map → chunk-W4L6C2XT.js.map} +0 -0
- /package/dist/{run-campaign-I3JXKVAK.js.map → run-campaign-LVFKZCEU.js.map} +0 -0
package/README.md
CHANGED
|
@@ -1,284 +1,365 @@
|
|
|
1
1
|
# `@tangle-network/agent-eval`
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
Measure agent behavior, compare changes on the same cases, and improve prompts or skills without exposing final test cases to the optimizer.
|
|
4
4
|
|
|
5
5
|
[](https://www.npmjs.com/package/@tangle-network/agent-eval)
|
|
6
6
|
[](https://pypi.org/project/agent-eval-rpc/)
|
|
7
7
|
[](https://github.com/tangle-network/agent-eval/actions/workflows/ci.yml)
|
|
8
8
|
[](./LICENSE)
|
|
9
9
|
|
|
10
|
-
|
|
11
|
-
It gives you numbers you can act on: how much the new prompt changed outcomes, how uncertain that estimate is, what failed and why, and whether the change meets your release rule.
|
|
10
|
+
Use this package to:
|
|
12
11
|
|
|
13
|
-
|
|
12
|
+
- run an agent over representative cases and score every result,
|
|
13
|
+
- compare a candidate with a baseline using paired statistics,
|
|
14
|
+
- analyze existing runs, traces, or human feedback,
|
|
15
|
+
- optimize a prompt or skill with official GEPA or SkillOpt,
|
|
16
|
+
- supply your own candidate generator for product-specific changes.
|
|
14
17
|
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
- run an automated improve-and-verify loop over a prompt, held to a promotion rule you choose,
|
|
18
|
-
- explain failures by cluster, cost, and judge disagreement.
|
|
19
|
-
|
|
20
|
-
The deterministic evaluator runs in your process and makes no network calls.
|
|
21
|
-
Features that use a model send their inputs to the model client you pass.
|
|
22
|
-
Trace exporters and hosted ingestion are also opt-in.
|
|
23
|
-
Python can drive the same engine over HTTP via [`agent-eval-rpc`](./clients/python/README.md).
|
|
24
|
-
|
|
25
|
-
---
|
|
18
|
+
The evaluation path runs in your TypeScript process.
|
|
19
|
+
Model calls occur only through the clients and agents you configure.
|
|
26
20
|
|
|
27
21
|
## Install
|
|
28
22
|
|
|
29
23
|
```sh
|
|
30
|
-
pnpm add @tangle-network/agent-eval
|
|
24
|
+
pnpm add @tangle-network/agent-eval
|
|
31
25
|
```
|
|
32
26
|
|
|
27
|
+
The official optimizers use the Python bridge.
|
|
28
|
+
Install only the optimizer you plan to run:
|
|
29
|
+
|
|
33
30
|
```sh
|
|
34
|
-
|
|
31
|
+
# Microsoft SkillOpt at the tested source revision
|
|
32
|
+
python -m pip install agent-eval-rpc
|
|
33
|
+
python -m pip install \
|
|
34
|
+
"skillopt @ git+https://github.com/microsoft/SkillOpt.git@61735e3922efc2b90c6d6cab561e62e98452ca90"
|
|
35
|
+
|
|
36
|
+
# GEPA Optimize Anything at the source revision tested by this release
|
|
37
|
+
python -m pip install agent-eval-rpc
|
|
38
|
+
python -m pip install "gepa[full] @ git+https://github.com/gepa-ai/gepa.git@f919db0a622e2e9f9204779b81fe00cc1b2d808f"
|
|
39
|
+
|
|
40
|
+
# DSPy 3.2.1 with Agent Eval metrics
|
|
41
|
+
python -m pip install "agent-eval-rpc[dspy]"
|
|
35
42
|
```
|
|
36
43
|
|
|
37
|
-
|
|
44
|
+
The published `gepa==0.1.4` wheel does not contain the Optimize Anything API used here.
|
|
45
|
+
The Git revision is intentional and should move only after compatibility tests pass.
|
|
46
|
+
The published `skillopt==0.2.0` wheel omits the prompt files required by `ReflACTTrainer`, so the tested SkillOpt source revision is also intentional.
|
|
47
|
+
DSPy 3.2.1 requires GEPA 0.0.27, while the general bridge requires GEPA 0.1.4.
|
|
48
|
+
Install the DSPy adapter and the general GEPA bridge in separate Python environments.
|
|
38
49
|
|
|
39
|
-
##
|
|
50
|
+
## Evaluate An Agent
|
|
40
51
|
|
|
41
|
-
|
|
42
|
-
|
|
52
|
+
This example is offline.
|
|
53
|
+
Replace the agent and judge functions with your product code.
|
|
43
54
|
|
|
44
55
|
```ts
|
|
45
56
|
import { defineAgentEval } from '@tangle-network/agent-eval/contract'
|
|
46
57
|
|
|
47
|
-
interface
|
|
58
|
+
interface SupportCase {
|
|
48
59
|
id: string
|
|
49
60
|
kind: 'support'
|
|
50
61
|
}
|
|
51
62
|
|
|
52
|
-
|
|
53
|
-
|
|
63
|
+
const evalKit = defineAgentEval<SupportCase, string>({
|
|
64
|
+
scenarios: [
|
|
54
65
|
{ id: 'refund', kind: 'support' },
|
|
55
66
|
{ id: 'shipping', kind: 'support' },
|
|
56
67
|
{ id: 'cancel', kind: 'support' },
|
|
57
|
-
]
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
name: 'cites-ticket',
|
|
67
|
-
dimensions: [{ key: 'ticket_id', description: 'The answer includes the ticket id' }],
|
|
68
|
-
score: ({ artifact, scenario }) => {
|
|
69
|
-
const ticketId = artifact.includes(scenario.id) ? 1 : 0
|
|
70
|
-
return { dimensions: { ticket_id: ticketId }, composite: ticketId, notes: '' }
|
|
71
|
-
},
|
|
68
|
+
],
|
|
69
|
+
agent: async (prompt, scenario) =>
|
|
70
|
+
String(prompt).includes('ticket') ? `Ticket ${scenario.id}: on it.` : 'On it.',
|
|
71
|
+
judge: {
|
|
72
|
+
name: 'ticket-id',
|
|
73
|
+
dimensions: [{ key: 'present', description: 'The answer includes the ticket id' }],
|
|
74
|
+
score: ({ artifact, scenario }) => {
|
|
75
|
+
const present = artifact.includes(scenario.id) ? 1 : 0
|
|
76
|
+
return { dimensions: { present }, composite: present, notes: '' }
|
|
72
77
|
},
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
78
|
+
},
|
|
79
|
+
baselineSurface: 'Answer politely.',
|
|
80
|
+
expectUsage: 'off',
|
|
81
|
+
})
|
|
82
|
+
|
|
83
|
+
console.log((await evalKit.evaluate()).aggregates.byJudge)
|
|
84
|
+
console.log(
|
|
85
|
+
(await evalKit.evaluate({ surface: 'Answer politely and cite the ticket id.' })).aggregates
|
|
86
|
+
.byJudge,
|
|
87
|
+
)
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
Each call runs every case, records the artifact, applies the same judge, and returns score distributions.
|
|
91
|
+
The surface is the value being changed, such as a prompt, skill, or serialized configuration.
|
|
92
|
+
|
|
93
|
+
## Adapt Another Text Optimizer
|
|
94
|
+
|
|
95
|
+
Use `externalTextOptimizationMethod()` when an existing package owns search and selection for a text prompt or named text components.
|
|
96
|
+
Its `run` callback receives the starting candidate plus serialized train and selection cases, but it never receives final test cases.
|
|
97
|
+
The optimizer must score candidates through `context.evaluate()` so Agent Eval can enforce the evaluation limit and use the configured execution and judges.
|
|
98
|
+
Every optimizer-owned paid call must use `context.cost.runPaidCall()`.
|
|
99
|
+
Set `source` to the package version and revision, and set `evaluationId` to a commit, content hash, or other stable identity for the execution and scoring behavior.
|
|
100
|
+
Agent Eval derives the run identity from those values, the optimizer settings, the starting surface, the described data, and the seed.
|
|
101
|
+
The callback returns the selected candidate, whether compatible state was restored, and how optimizer spend was recorded.
|
|
102
|
+
|
|
103
|
+
See [Adapt A Third-Party Text Optimizer](./docs/campaign-proposers.md#adapt-a-third-party-text-optimizer) for a complete minimal adapter.
|
|
104
|
+
|
|
105
|
+
## Optimize With Official GEPA
|
|
106
|
+
|
|
107
|
+
`gepaOptimizationMethod()` delegates candidate search and recipe composition to the installed GEPA package.
|
|
108
|
+
Agent Eval supplies the train and selection cases, executes candidates, records cost, and evaluates the selected result on final cases after GEPA exits.
|
|
76
109
|
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
110
|
+
```ts
|
|
111
|
+
import { gepaOptimizationMethod } from '@tangle-network/agent-eval/campaign'
|
|
112
|
+
|
|
113
|
+
const optimizerPricing = {
|
|
114
|
+
inputUsdPerMillion: Number(process.env.OPTIMIZER_INPUT_USD_PER_MILLION),
|
|
115
|
+
outputUsdPerMillion: Number(process.env.OPTIMIZER_OUTPUT_USD_PER_MILLION),
|
|
80
116
|
}
|
|
81
117
|
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
118
|
+
const gepa = gepaOptimizationMethod<MyCase, MyArtifact>({
|
|
119
|
+
objective: 'Improve the instructions so the agent returns valid, complete JSON.',
|
|
120
|
+
evaluationId: 'json-agent',
|
|
121
|
+
recipe: {
|
|
122
|
+
kind: 'engine',
|
|
123
|
+
run: {
|
|
124
|
+
engine: 'gepa',
|
|
125
|
+
maxEvaluations: 40,
|
|
126
|
+
maxProposerCostUsd: 5,
|
|
127
|
+
},
|
|
128
|
+
},
|
|
129
|
+
optimizer: {
|
|
130
|
+
model: 'gpt-4.1-mini',
|
|
131
|
+
baseUrl: 'https://api.openai.com/v1',
|
|
132
|
+
apiKey: process.env.OPENAI_API_KEY!,
|
|
133
|
+
budget: {
|
|
134
|
+
maxCostUsd: 5,
|
|
135
|
+
maxRequests: 100,
|
|
136
|
+
maxRequestBytes: 2_000_000,
|
|
137
|
+
maxResponseBytes: 2_000_000,
|
|
138
|
+
maxOutputTokensPerRequest: 32_768,
|
|
139
|
+
pricing: optimizerPricing,
|
|
140
|
+
},
|
|
141
|
+
},
|
|
142
|
+
describeScenario: (scenario) => ({ input: scenario.input }),
|
|
143
|
+
describeArtifact: (artifact) => ({ output: artifact.output }),
|
|
85
144
|
})
|
|
86
145
|
```
|
|
87
146
|
|
|
88
|
-
|
|
147
|
+
GEPA also supports official sequential, adaptive, best-of, vote, and Omni recipes through the same factory.
|
|
148
|
+
When every recipe stage uses the standard GEPA engine, Agent Eval keeps the provider key outside Python and records exact reflection usage through a local proxy.
|
|
149
|
+
Other official GEPA engines can still run through `engineConfig`, but their model spend remains incomplete unless the engine reports it.
|
|
150
|
+
Custom engines can register through GEPA's official registry by listing their Python modules in `engineModules`.
|
|
89
151
|
|
|
90
|
-
|
|
91
|
-
baseline: { 'cites-ticket': { mean: 0, stdev: 0, ci95: [ 0, 0 ], n: 3 } }
|
|
92
|
-
candidate: { 'cites-ticket': { mean: 1, stdev: 0, ci95: [ 1, 1 ], n: 3 } }
|
|
93
|
-
```
|
|
152
|
+
Run the repository example:
|
|
94
153
|
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
154
|
+
```sh
|
|
155
|
+
OPTIMIZERS=gepa \
|
|
156
|
+
LLM_API_KEY="$OPENAI_API_KEY" \
|
|
157
|
+
GEPA_PRICE_IN_PER_M=0.4 \
|
|
158
|
+
GEPA_PRICE_OUT_PER_M=1.6 \
|
|
159
|
+
pnpm tsx examples/compare-optimization-methods/index.ts
|
|
160
|
+
```
|
|
99
161
|
|
|
100
|
-
|
|
162
|
+
Replace the example rates with the exact endpoint rates.
|
|
101
163
|
|
|
102
|
-
|
|
164
|
+
## Optimize With Official SkillOpt
|
|
103
165
|
|
|
104
|
-
`
|
|
166
|
+
`skillOptOptimizationMethod()` runs Microsoft's `ReflACTTrainer` against the same TypeScript execution and scoring path.
|
|
167
|
+
SkillOpt receives train and selection cases but never receives final cases.
|
|
105
168
|
|
|
106
169
|
```ts
|
|
107
|
-
import {
|
|
170
|
+
import { skillOptOptimizationMethod } from '@tangle-network/agent-eval/campaign'
|
|
108
171
|
|
|
109
|
-
const
|
|
110
|
-
|
|
172
|
+
const optimizerPricing = {
|
|
173
|
+
inputUsdPerMillion: Number(process.env.OPTIMIZER_INPUT_USD_PER_MILLION),
|
|
174
|
+
outputUsdPerMillion: Number(process.env.OPTIMIZER_OUTPUT_USD_PER_MILLION),
|
|
175
|
+
}
|
|
111
176
|
|
|
112
|
-
const
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
177
|
+
const skillopt = skillOptOptimizationMethod<MyCase, MyArtifact>({
|
|
178
|
+
objective: 'Improve the skill so the agent returns valid, complete JSON.',
|
|
179
|
+
evaluationId: 'json-agent',
|
|
180
|
+
trainer: {
|
|
181
|
+
epochs: 2,
|
|
182
|
+
batchSize: 4,
|
|
183
|
+
},
|
|
184
|
+
optimizer: {
|
|
185
|
+
model: 'gpt-4.1-mini',
|
|
186
|
+
baseUrl: 'https://api.openai.com/v1',
|
|
187
|
+
apiKey: process.env.OPENAI_API_KEY!,
|
|
188
|
+
budget: {
|
|
189
|
+
maxCostUsd: 5,
|
|
190
|
+
maxRequests: 100,
|
|
191
|
+
maxRequestBytes: 2_000_000,
|
|
192
|
+
maxResponseBytes: 2_000_000,
|
|
193
|
+
maxOutputTokensPerRequest: 32_768,
|
|
194
|
+
pricing: optimizerPricing,
|
|
195
|
+
},
|
|
196
|
+
},
|
|
197
|
+
maxEvaluations: 80,
|
|
198
|
+
describeScenario: (scenario) => ({ input: scenario.input }),
|
|
199
|
+
describeArtifact: (artifact) => ({ output: artifact.output }),
|
|
117
200
|
})
|
|
201
|
+
```
|
|
118
202
|
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
203
|
+
Replace the example rates with the exact rates for your endpoint.
|
|
204
|
+
Agent Eval places a local OpenAI-compatible proxy between each standard optimizer and the provider.
|
|
205
|
+
The proxy enforces request, byte, token, and dollar limits before forwarding calls, records exact usage from provider responses, and does not pass the provider key to the optimizer process.
|
|
206
|
+
|
|
207
|
+
## Optimize A DSPy Program
|
|
208
|
+
|
|
209
|
+
DSPy owns its program optimizers.
|
|
210
|
+
`DspyJudgeMetric` lets them use the same Agent Eval rubric as TypeScript agents.
|
|
211
|
+
|
|
212
|
+
```python
|
|
213
|
+
import dspy
|
|
214
|
+
|
|
215
|
+
from agent_eval_rpc import DspyJudgeMetric
|
|
216
|
+
|
|
217
|
+
dspy.configure(lm=dspy.LM("openai/gpt-4.1-mini"))
|
|
218
|
+
metric = DspyJudgeMetric(rubric_name="answer-quality")
|
|
219
|
+
|
|
220
|
+
# GEPA needs both the numeric score and diagnostic feedback.
|
|
221
|
+
optimizer = dspy.GEPA(
|
|
222
|
+
metric=metric.feedback,
|
|
223
|
+
reflection_lm=dspy.LM("openai/gpt-4.1-mini"),
|
|
224
|
+
max_metric_calls=100,
|
|
129
225
|
)
|
|
226
|
+
optimized_program = optimizer.compile(program, trainset=train, valset=selection)
|
|
227
|
+
|
|
228
|
+
# MIPROv2, SIMBA, and few-shot optimizers use the numeric metric directly.
|
|
229
|
+
mipro = dspy.MIPROv2(metric=metric, auto="light")
|
|
130
230
|
```
|
|
131
231
|
|
|
132
|
-
|
|
133
|
-
|
|
232
|
+
Use official DSPy directly for DSPy programs.
|
|
233
|
+
Use `gepaOptimizationMethod()` for text or named component surfaces in non-DSPy agents.
|
|
234
|
+
Use agent-runtime's worktree path for executable code changes.
|
|
134
235
|
|
|
135
|
-
|
|
236
|
+
## Compare Complete Methods
|
|
136
237
|
|
|
137
|
-
|
|
138
|
-
|
|
238
|
+
`compareOptimizationMethods()` gives each method the same starting surface, execution function, judges, train cases, and selection cases.
|
|
239
|
+
It waits for optimization to finish before it evaluates any selected surface on the final cases.
|
|
139
240
|
|
|
140
241
|
```ts
|
|
141
|
-
import {
|
|
142
|
-
type BuiltinOptimizationMethodConfig,
|
|
143
|
-
compareOptimizationMethods,
|
|
144
|
-
gepaParetoMethod,
|
|
145
|
-
gepaReflectionMethod,
|
|
146
|
-
skillOptMethod,
|
|
147
|
-
} from '@tangle-network/agent-eval/campaign'
|
|
148
|
-
|
|
149
|
-
const methodConfig: BuiltinOptimizationMethodConfig<MyScenario, MyArtifact> = {
|
|
150
|
-
llm,
|
|
151
|
-
model,
|
|
152
|
-
target: 'the complete prompt being improved',
|
|
153
|
-
}
|
|
242
|
+
import { compareOptimizationMethods } from '@tangle-network/agent-eval/campaign'
|
|
154
243
|
|
|
155
|
-
const result = await compareOptimizationMethods
|
|
156
|
-
methods: [
|
|
157
|
-
gepaReflectionMethod(methodConfig),
|
|
158
|
-
gepaParetoMethod(methodConfig),
|
|
159
|
-
skillOptMethod(methodConfig),
|
|
160
|
-
],
|
|
244
|
+
const result = await compareOptimizationMethods({
|
|
245
|
+
methods: [gepa, skillopt],
|
|
161
246
|
baselineSurface,
|
|
162
247
|
trainScenarios,
|
|
163
248
|
selectionScenarios,
|
|
164
249
|
testScenarios,
|
|
165
250
|
dispatchWithSurface,
|
|
166
251
|
judges,
|
|
167
|
-
runDir,
|
|
252
|
+
runDir: '.agent-eval/optimizer-comparison',
|
|
168
253
|
})
|
|
169
|
-
```
|
|
170
254
|
|
|
171
|
-
|
|
172
|
-
|
|
173
|
-
Ranks follow estimated lift; use the intervals and pairwise results to determine whether the observed difference excludes zero.
|
|
174
|
-
See the [method-comparison guide](./docs/campaign-proposers.md) and [runnable example](./examples/compare-optimization-methods/).
|
|
175
|
-
|
|
176
|
-
---
|
|
177
|
-
|
|
178
|
-
## Core APIs
|
|
255
|
+
console.table(result.scores)
|
|
256
|
+
```
|
|
179
257
|
|
|
180
|
-
|
|
181
|
-
|
|
258
|
+
Read `scores` for final-case lift and intervals.
|
|
259
|
+
Read `pairwise` before claiming one method beat another.
|
|
260
|
+
Read `totalCost.accountingComplete` before using the reported dollars as a complete total.
|
|
261
|
+
Each official method score records the optimizer and bridge package versions, source revisions and source-tree hashes, Python runtime, custom engine module hashes, compatible run ID, exact attempt ID, resume status, evaluation count, artifact directory, and available optimizer token usage in `provenance`.
|
|
182
262
|
|
|
183
|
-
|
|
184
|
-
|
|
185
|
-
| **Evaluation** (`runEval`, `runCampaign`) | Run agent × scenarios × repetitions, score every run, and record the result. |
|
|
186
|
-
| **Scoring** (`JudgeConfig`, `llmJudge`, calibration) | Score one output on weighted dimensions with code or a model, then compare model scores against human ratings. |
|
|
187
|
-
| **Release rules** (`heldOutGate`, `paretoSignificanceGate`, `composeGate`, …) | Decide whether a candidate ships, such as requiring an improvement on scenarios that candidate generation never saw. |
|
|
188
|
-
| **Candidate generation** (`gepaProposer`, `evolutionaryProposer`, …) | Generate candidate prompts or configs from prior failures. |
|
|
189
|
-
| **Method comparison** (`compareOptimizationMethods`) | Run complete optimization methods on shared train and selection data, then rank them on separate final test data. |
|
|
190
|
-
| **Run analysis** (`analyzeRuns`, `diffRuns`) | Turn any set of `RunRecord`s into a report: score distributions, baseline-vs-candidate lift with confidence intervals, failure clusters, cost breakdown, recommendations. |
|
|
191
|
-
| **Intake adapters** (`fromFeedbackTable`, `fromOtelSpans`) | Convert data you already have, such as human ratings tables and OpenTelemetry spans, into `RunRecord`s. |
|
|
192
|
-
| **Cost tracking** | Attribute every model call's tokens and dollars to the run, phase, and judge that spent them, including interrupted calls. |
|
|
193
|
-
| **Human feedback storage** | Persist runs with approved, rejected, or edited labels so review activity becomes training and eval data. |
|
|
194
|
-
| **Statistics** (`pairedBootstrap`, `benjaminiHochberg`, sequential tests) | The release-decision math, usable standalone. |
|
|
195
|
-
| **Trace tools** (`/traces`, `/analyst`) | Store and replay structured run traces; cluster failures with an LLM analyst panel. |
|
|
196
|
-
| **HTTP and RPC** (`/wire`) | Expose judging and ingestion to non-TypeScript stacks, including the Python client. |
|
|
197
|
-
|
|
198
|
-
Our own experiments with these primitives live in [`examples/`](./examples/README.md); they are demonstrations, not part of the API.
|
|
199
|
-
|
|
200
|
-
| Runnable example | Shows |
|
|
201
|
-
|---|---|
|
|
202
|
-
| [`examples/selfimprove-quickstart/`](./examples/selfimprove-quickstart/) | The closed improve-and-verify loop, fully offline |
|
|
203
|
-
| [`examples/customer-feedback-loop/`](./examples/customer-feedback-loop/) | Multi-rater human feedback (CSV/Sheets/Obsidian) → per-rater judges → report |
|
|
204
|
-
| [`examples/customer-otel-traces/`](./examples/customer-otel-traces/) | Production OpenTelemetry traces → report, no closed loop required |
|
|
205
|
-
| [`examples/compare-optimization-methods/`](./examples/compare-optimization-methods/) | Compare complete optimization methods with separate train, selection, and test data |
|
|
263
|
+
The [optimizer guide](./docs/campaign-proposers.md) covers recipes, budgets, resuming, and data separation.
|
|
264
|
+
The [runnable comparison](./examples/compare-optimization-methods/) can run GEPA, SkillOpt, or both.
|
|
206
265
|
|
|
207
|
-
|
|
266
|
+
## Supply Your Own Candidate Generator
|
|
208
267
|
|
|
209
|
-
|
|
268
|
+
Use `SurfaceProposer` when candidate creation belongs to your product or an agent runtime.
|
|
269
|
+
The campaign still owns execution, scoring, history, stopping, and release decisions.
|
|
210
270
|
|
|
211
|
-
|
|
271
|
+
```ts
|
|
272
|
+
import {
|
|
273
|
+
defineAgentEval,
|
|
274
|
+
type SurfaceProposer,
|
|
275
|
+
} from '@tangle-network/agent-eval/contract'
|
|
276
|
+
|
|
277
|
+
const proposer: SurfaceProposer = {
|
|
278
|
+
kind: 'product-rules',
|
|
279
|
+
async propose({ currentSurface, populationSize }) {
|
|
280
|
+
const prompt = String(currentSurface)
|
|
281
|
+
return [
|
|
282
|
+
{
|
|
283
|
+
surface: `${prompt}\nReturn JSON only.`,
|
|
284
|
+
label: 'json-only',
|
|
285
|
+
rationale: 'Training failures contained prose around the JSON object.',
|
|
286
|
+
},
|
|
287
|
+
].slice(0, populationSize)
|
|
288
|
+
},
|
|
289
|
+
}
|
|
212
290
|
|
|
213
|
-
|
|
291
|
+
const result = await defineAgentEval({
|
|
292
|
+
scenarios,
|
|
293
|
+
agent,
|
|
294
|
+
judge,
|
|
295
|
+
baselineSurface,
|
|
296
|
+
proposer,
|
|
297
|
+
budget: { generations: 1, populationSize: 1, holdoutFraction: 0.3 },
|
|
298
|
+
}).improve()
|
|
299
|
+
```
|
|
214
300
|
|
|
215
|
-
|
|
216
|
-
|---|---|
|
|
217
|
-
| `/contract` | **Start here.** Stable APIs for defining an eval, running it, improving a prompt, judging outputs, analyzing existing runs, and storing results. |
|
|
218
|
-
| `/campaign` | Lower-level control over candidate generation, release rules, storage, and comparisons. |
|
|
219
|
-
| `/reporting` | Statistical comparisons and report renderers. |
|
|
220
|
-
| `/analyst` | Model-based failure clustering and stored findings. |
|
|
221
|
-
| `/traces` | Trace stores, emitters, deterministic replay, trace analysis. |
|
|
222
|
-
| `/rl` | Export eval artifacts as training signal: rewards, preferences, trainer-format datasets. |
|
|
223
|
-
| `/benchmarks` | Benchmark adapter contract + retrieval metrics + a bundled reference benchmark. |
|
|
224
|
-
| `/wire` | The HTTP/RPC server and Zod schemas (what the Python client speaks). |
|
|
225
|
-
| `/hosted` | Client for shipping eval-run events to a remote orchestrator (see below). |
|
|
226
|
-
| `/control` | A generic observe → validate → decide → act agent loop with eval-backed stopping rules. |
|
|
227
|
-
| `/matrix`, `/multishot` | N-axis configuration sweeps; multi-turn persona × turn-count runners. |
|
|
228
|
-
| `/meta-eval`, `/belief-state`, `/builder-eval`, `/pipelines`, `/storyboard`, `/authenticity`, `/fuzz`, `/trace-attributes` | Specialized surfaces: judge calibration, decision-point extraction, code-generator grading, trace diagnostics, run replay rendering, anti-gaming output checks, input fuzzing, trace attribute vocabulary. |
|
|
301
|
+
Run the complete offline example:
|
|
229
302
|
|
|
230
|
-
|
|
303
|
+
```sh
|
|
304
|
+
pnpm tsx examples/selfimprove-quickstart/index.ts
|
|
305
|
+
```
|
|
231
306
|
|
|
232
|
-
|
|
307
|
+
## Start From Existing Runs
|
|
233
308
|
|
|
234
|
-
|
|
309
|
+
You do not need a runnable agent to analyze data you already captured.
|
|
310
|
+
Use `analyzeRuns()` for `RunRecord[]`, or use the feedback and OpenTelemetry adapters to normalize existing data first.
|
|
235
311
|
|
|
236
|
-
|
|
237
|
-
- [`docs/customer-journeys.md`](./docs/customer-journeys.md): three complete adoption paths with code
|
|
238
|
-
- [`docs/insight-report.md`](./docs/insight-report.md): annotated walkthrough of every section of the `analyzeRuns()` report
|
|
239
|
-
- [`docs/campaign-proposers.md`](./docs/campaign-proposers.md): candidate generation and fair comparison of complete optimization methods
|
|
240
|
-
- [`docs/adapters-observability.md`](./docs/adapters-observability.md): composing with LangSmith, Langfuse, Phoenix, and OpenLLMetry
|
|
241
|
-
- [`docs/wire-protocol.md`](./docs/wire-protocol.md): the HTTP/RPC contract for other languages
|
|
242
|
-
- [`docs/design.md`](./docs/design.md): how this package relates to the rest of the Tangle agent stack, and the dependency rules that keep it reusable
|
|
243
|
-
- [`CHANGELOG.md`](./CHANGELOG.md): every release, with additive and breaking changes identified
|
|
312
|
+
See [concepts](./docs/concepts.md), [customer paths](./docs/customer-journeys.md), and [trace analysis](./docs/trace-analysis.md).
|
|
244
313
|
|
|
245
|
-
|
|
314
|
+
## Entry Points
|
|
246
315
|
|
|
247
|
-
|
|
316
|
+
| Import | Use |
|
|
317
|
+
|---|---|
|
|
318
|
+
| `@tangle-network/agent-eval/contract` | Define an evaluation, run it, improve with a custom candidate generator, and analyze runs. |
|
|
319
|
+
| `@tangle-network/agent-eval/campaign` | Control campaigns, official optimization methods, comparisons, storage, and release rules. |
|
|
320
|
+
| `@tangle-network/agent-eval/reporting` | Statistical comparisons and report rendering. |
|
|
321
|
+
| `@tangle-network/agent-eval/analyst` | Model-assisted failure analysis. |
|
|
322
|
+
| `@tangle-network/agent-eval/traces` | Store, replay, and inspect structured traces. |
|
|
323
|
+
| `@tangle-network/agent-eval/benchmarks` | Benchmark adapters and retrieval metrics. |
|
|
324
|
+
| `@tangle-network/agent-eval/rl` | Export rewards, preferences, and training rows. |
|
|
325
|
+
| `@tangle-network/agent-eval/wire` | HTTP and RPC schemas for other languages. |
|
|
248
326
|
|
|
249
|
-
|
|
250
|
-
|
|
327
|
+
Prefer these subpaths for new code.
|
|
328
|
+
The root export remains broad for compatibility.
|
|
251
329
|
|
|
252
|
-
|
|
253
|
-
await evalKit.improve({
|
|
254
|
-
hostedTenant: {
|
|
255
|
-
endpoint: 'https://intelligence.tangle.tools',
|
|
256
|
-
apiKey: process.env.TANGLE_API_KEY!,
|
|
257
|
-
tenantId: 'your-tenant',
|
|
258
|
-
},
|
|
259
|
-
})
|
|
260
|
-
```
|
|
330
|
+
## Examples
|
|
261
331
|
|
|
262
|
-
|
|
263
|
-
|
|
264
|
-
|
|
265
|
-
|
|
332
|
+
| Goal | Example |
|
|
333
|
+
|---|---|
|
|
334
|
+
| Evaluate and improve with a custom candidate generator | [`selfimprove-quickstart`](./examples/selfimprove-quickstart/) |
|
|
335
|
+
| Run official GEPA or SkillOpt | [`compare-optimization-methods`](./examples/compare-optimization-methods/) |
|
|
336
|
+
| Analyze human feedback | [`customer-feedback-loop`](./examples/customer-feedback-loop/) |
|
|
337
|
+
| Analyze OpenTelemetry traces | [`customer-otel-traces`](./examples/customer-otel-traces/) |
|
|
338
|
+
| Run public benchmark adapters | [`benchmarks`](./examples/benchmarks/) |
|
|
266
339
|
|
|
267
|
-
|
|
340
|
+
See the [example index](./examples/README.md) for the full list.
|
|
268
341
|
|
|
269
342
|
## Development
|
|
270
343
|
|
|
271
344
|
```sh
|
|
272
345
|
pnpm install
|
|
273
|
-
pnpm build
|
|
274
|
-
pnpm test # vitest, ~3300 tests
|
|
275
346
|
pnpm typecheck
|
|
347
|
+
pnpm typecheck:examples
|
|
348
|
+
pnpm test
|
|
349
|
+
pnpm build
|
|
276
350
|
```
|
|
277
351
|
|
|
278
|
-
|
|
352
|
+
Python compatibility tests use the locked dependencies:
|
|
353
|
+
|
|
354
|
+
```sh
|
|
355
|
+
cd clients/python
|
|
356
|
+
uv sync --frozen --extra dev --group skillopt-source --group gepa-source
|
|
357
|
+
uv run --frozen pytest
|
|
279
358
|
|
|
280
|
-
|
|
359
|
+
uv sync --frozen --extra dev --extra dspy
|
|
360
|
+
uv run --frozen pytest tests/test_dspy_metric.py
|
|
361
|
+
```
|
|
281
362
|
|
|
282
363
|
## License
|
|
283
364
|
|
|
284
|
-
MIT.
|
|
365
|
+
MIT.
|