@tangle-network/agent-eval 0.94.0 → 0.95.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +32 -0
- package/README.md +44 -30
- package/dist/adapters/http.d.ts +8 -7
- package/dist/adapters/http.js.map +1 -1
- package/dist/adapters/langchain.d.ts +3 -2
- package/dist/adapters/otel.d.ts +5 -4
- package/dist/analyst/index.d.ts +11 -31
- package/dist/analyst/index.js +5 -65
- package/dist/analyst/index.js.map +1 -1
- package/dist/{analyze-runs-B6Ljo_dI.d.ts → analyze-runs-DtT6F_6T.d.ts} +3 -3
- package/dist/belief-state/index.d.ts +4 -3
- package/dist/benchmarks/index.d.ts +3 -2
- package/dist/campaign/index.d.ts +727 -616
- package/dist/campaign/index.js +1863 -1316
- package/dist/campaign/index.js.map +1 -1
- package/dist/{chunk-2K6UUZ7P.js → chunk-2T4EZACH.js} +1 -1
- package/dist/chunk-2T4EZACH.js.map +1 -0
- package/dist/{chunk-CTBHKLEU.js → chunk-77T4STFI.js} +59 -86
- package/dist/chunk-77T4STFI.js.map +1 -0
- package/dist/{chunk-EGPMSBEZ.js → chunk-7QTQKIDD.js} +178 -177
- package/dist/chunk-7QTQKIDD.js.map +1 -0
- package/dist/{chunk-MIFZUPEK.js → chunk-AQ5WQAIV.js} +21 -6
- package/dist/chunk-AQ5WQAIV.js.map +1 -0
- package/dist/chunk-DJWX3GVS.js +81 -0
- package/dist/chunk-DJWX3GVS.js.map +1 -0
- package/dist/{chunk-S6OZEZQK.js → chunk-HMA63UEO.js} +37 -9
- package/dist/{chunk-S6OZEZQK.js.map → chunk-HMA63UEO.js.map} +1 -1
- package/dist/{chunk-TBDR6PAI.js → chunk-IZCEK2HR.js} +2 -2
- package/dist/{chunk-SD2YFWQQ.js → chunk-KKWJD5E6.js} +20 -20
- package/dist/chunk-KKWJD5E6.js.map +1 -0
- package/dist/{chunk-KWRRMR3J.js → chunk-LO6IOIJ2.js} +10 -10
- package/dist/chunk-LO6IOIJ2.js.map +1 -0
- package/dist/{chunk-E4GH6USR.js → chunk-NZEQVRH5.js} +2 -2
- package/dist/chunk-NZEQVRH5.js.map +1 -0
- package/dist/{chunk-MPQWFX6Y.js → chunk-PSWWQXHF.js} +13 -88
- package/dist/chunk-PSWWQXHF.js.map +1 -0
- package/dist/{chunk-Q5LIB7BC.js → chunk-S4SYLDFX.js} +2 -2
- package/dist/chunk-S4SYLDFX.js.map +1 -0
- package/dist/{chunk-KW53MSA5.js → chunk-X74V6ESX.js} +2 -2
- package/dist/{chunk-QMUEXQJS.js → chunk-YBIGNSCZ.js} +81 -4
- package/dist/chunk-YBIGNSCZ.js.map +1 -0
- package/dist/{chunk-2KNZHH3P.js → chunk-Z6L6YSU6.js} +2 -2
- package/dist/{code-agent-session-BO8nCnv3.d.ts → code-agent-session-CPHRCb4-.d.ts} +1 -1
- package/dist/contract/index.d.ts +91 -43
- package/dist/contract/index.js +127 -17
- package/dist/contract/index.js.map +1 -1
- package/dist/{control-D6qwHXIR.d.ts → control-Doncu-B_.d.ts} +2 -2
- package/dist/control.d.ts +3 -2
- package/dist/control.js +2 -2
- package/dist/{corpus-B8A4BDR3.d.ts → corpus-D4YW9UoJ.d.ts} +1 -1
- package/dist/{default-registry-6dhErQbs.d.ts → default-registry-GyE8X5SP.d.ts} +3 -3
- package/dist/diagnose.d.ts +4 -3
- package/dist/diagnose.js +1 -1
- package/dist/{run-improvement-loop-DBahB8Ax.d.ts → gepa-C1NCIZ9o.d.ts} +117 -130
- package/dist/hosted/index.d.ts +5 -4
- package/dist/{index-Bx3gZ8xl.d.ts → index-_Y4oNOOb.d.ts} +1 -1
- package/dist/index.d.ts +76 -81
- package/dist/index.js +66 -31
- package/dist/index.js.map +1 -1
- package/dist/{insight-report-DWl3z9tl.d.ts → insight-report-BnRjTibG.d.ts} +1 -1
- package/dist/{kind-factory-0BhLSI27.d.ts → kind-factory-X3eDYbKn.d.ts} +2 -3
- package/dist/matrix/index.d.ts +1 -1
- package/dist/meta-eval/index.d.ts +3 -2
- package/dist/multishot/index.d.ts +4 -4
- package/dist/multishot/index.js.map +1 -1
- package/dist/openapi.json +1 -1
- package/dist/{pre-registration-mAnCugl9.d.ts → pre-registration-nfUdc9EQ.d.ts} +2 -42
- package/dist/{provenance-P-bCL2Fo.d.ts → provenance-CncDq9qE.d.ts} +26 -41
- package/dist/{release-report-BEbWmVYj.d.ts → release-report-pidWUMZ2.d.ts} +2 -2
- package/dist/reporting.d.ts +5 -4
- package/dist/{researcher-B0C2_fVO.d.ts → researcher-Jr8ME1dZ.d.ts} +2 -2
- package/dist/rl.d.ts +516 -515
- package/dist/rl.js +612 -612
- package/dist/rl.js.map +1 -1
- package/dist/{rubric-predictive-validity-Cy_W-hWZ.d.ts → rubric-predictive-validity-C2hDKM8Z.d.ts} +1 -1
- package/dist/{run-campaign-7WNXMDSN.js → run-campaign-WXY7KI67.js} +2 -2
- package/dist/{run-record-e7vj1uZQ.d.ts → run-record-CP2ObebC.d.ts} +14 -18
- package/dist/{runtime-trajectory-BDgfGZSr.d.ts → runtime-trajectory-BOUUjI0y.d.ts} +1 -1
- package/dist/{semantic-concept-judge-B9MgmBnM.d.ts → semantic-concept-judge-DSBB2Cfp.d.ts} +2 -2
- package/dist/{summary-report-BDOFevaT.d.ts → summary-report-CInXwsza.d.ts} +1 -1
- package/dist/testing-C21CHsq2.d.ts +20 -0
- package/dist/testing.d.ts +1 -0
- package/dist/testing.js +8 -0
- package/dist/testing.js.map +1 -0
- package/dist/traces.d.ts +26 -10
- package/dist/traces.js +41 -11
- package/dist/{types-Ce17tDlG.d.ts → types-B5x54y6n.d.ts} +1 -1
- package/dist/{types-mn5Aqk7x.d.ts → types-BUxNaJ8c.d.ts} +2 -4
- package/dist/{types-BU-7W85F.d.ts → types-DQRY8ZT-.d.ts} +60 -58
- package/dist/workflow/index.d.ts +5 -4
- package/dist/workflow/index.js +1 -1
- package/docs/campaign-proposers.md +170 -0
- package/docs/concepts.md +8 -4
- package/docs/customer-journeys.md +15 -13
- package/docs/design/loop-taxonomy.md +34 -66
- package/docs/distributed-driver.md +14 -14
- package/docs/feature-guide.md +1 -1
- package/docs/hosted-ingest-spec.md +2 -3
- package/docs/multi-shot-optimization.md +8 -8
- package/docs/product-eval-adoption.md +1 -1
- package/docs/self-improvement-map.md +33 -29
- package/package.json +8 -14
- package/dist/chunk-2K6UUZ7P.js.map +0 -1
- package/dist/chunk-CTBHKLEU.js.map +0 -1
- package/dist/chunk-E4GH6USR.js.map +0 -1
- package/dist/chunk-EGPMSBEZ.js.map +0 -1
- package/dist/chunk-KWRRMR3J.js.map +0 -1
- package/dist/chunk-MIFZUPEK.js.map +0 -1
- package/dist/chunk-MPQWFX6Y.js.map +0 -1
- package/dist/chunk-Q5LIB7BC.js.map +0 -1
- package/dist/chunk-QMUEXQJS.js.map +0 -1
- package/dist/chunk-SD2YFWQQ.js.map +0 -1
- package/docs/design/external-agent-wedge.md +0 -89
- package/docs/design/phase-d-rfc.md +0 -125
- package/docs/design/phase4-consumer-migration.md +0 -70
- package/docs/design/primitives-integration-spec.md +0 -393
- package/docs/design/product-self-improvement-loop.md +0 -146
- package/docs/design/self-improvement-engine.md +0 -140
- package/docs/design/self-improvement-protocol.md +0 -223
- package/docs/design/self-improvement-roadmap.md +0 -106
- package/docs/design/substrate-gaps.md +0 -118
- package/docs/phase-b-pairing-kit.md +0 -188
- package/docs/phase-b-runbook.md +0 -176
- package/docs/pilot/README.md +0 -62
- package/docs/pilot/customer-checklist.md +0 -90
- package/docs/pilot/integration-foreign-stack.md +0 -296
- package/docs/pilot/integration-tangle-stack.md +0 -248
- package/docs/pilot/one-pager.md +0 -161
- package/docs/pilot/sample-insight-report.json +0 -172
- package/docs/quickstart-external.md +0 -229
- package/docs/research/belief-state-agent-eval-roadmap.md +0 -593
- package/docs/research/research-roadmap.md +0 -205
- package/docs/specs/driver-honest-spec.md +0 -251
- package/docs/specs/hermes-self-improvement-audit.md +0 -93
- package/docs/specs/profile-versioning.md +0 -291
- package/docs/three-package-architecture.md +0 -168
- /package/dist/{chunk-TBDR6PAI.js.map → chunk-IZCEK2HR.js.map} +0 -0
- /package/dist/{chunk-KW53MSA5.js.map → chunk-X74V6ESX.js.map} +0 -0
- /package/dist/{chunk-2KNZHH3P.js.map → chunk-Z6L6YSU6.js.map} +0 -0
- /package/dist/{run-campaign-7WNXMDSN.js.map → run-campaign-WXY7KI67.js.map} +0 -0
|
@@ -1,229 +0,0 @@
|
|
|
1
|
-
# Quickstart — self-improvement loop for any agent (15 minutes)
|
|
2
|
-
|
|
3
|
-
The standalone walkthrough mirroring
|
|
4
|
-
`examples/foreign-agent-quickstart/`. Read this first; copy the runnable
|
|
5
|
-
example second.
|
|
6
|
-
|
|
7
|
-
## What you get
|
|
8
|
-
|
|
9
|
-
After 15 minutes you have a closed self-improvement loop running
|
|
10
|
-
against your agent — measured, gated, and reproducible — with no
|
|
11
|
-
Tangle sandbox, no Tangle account, and no hosted infrastructure.
|
|
12
|
-
|
|
13
|
-
## Install
|
|
14
|
-
|
|
15
|
-
```sh
|
|
16
|
-
npm i @tangle-network/agent-eval@^0.46.0
|
|
17
|
-
```
|
|
18
|
-
|
|
19
|
-
The package's `@tangle-network/sandbox` peer is `optional`. Foreign
|
|
20
|
-
consumers install agent-eval and run the full LAND tier without our
|
|
21
|
-
sandbox or its dependencies.
|
|
22
|
-
|
|
23
|
-
## The one-shot happy path
|
|
24
|
-
|
|
25
|
-
If you don't want to learn the substrate, the entire LAND tier reduces
|
|
26
|
-
to one function call:
|
|
27
|
-
|
|
28
|
-
```ts
|
|
29
|
-
import { selfImprove } from '@tangle-network/agent-eval/contract'
|
|
30
|
-
|
|
31
|
-
const result = await selfImprove({
|
|
32
|
-
agent: (surface, scenario, ctx) =>
|
|
33
|
-
runYourAgent({ systemPrompt: surface as string, scenario, signal: ctx.signal }),
|
|
34
|
-
scenarios,
|
|
35
|
-
judge,
|
|
36
|
-
baselineSurface: 'You are a senior copywriter…',
|
|
37
|
-
budget: { dollars: 10, generations: 3 },
|
|
38
|
-
})
|
|
39
|
-
|
|
40
|
-
console.log(`lift: ${result.lift.toFixed(3)} (${result.gateDecision})`)
|
|
41
|
-
if (result.gateDecision === 'ship') {
|
|
42
|
-
// result.winner.surface is the optimized prompt
|
|
43
|
-
}
|
|
44
|
-
```
|
|
45
|
-
|
|
46
|
-
That's the LAND happy path. Smart defaults pick: in-memory storage,
|
|
47
|
-
`gepaDriver` with copywriting-flavored mutation primitives,
|
|
48
|
-
`defaultProductionGate` with `deltaThreshold: 0.05`, 25% deterministic
|
|
49
|
-
train/holdout split.
|
|
50
|
-
|
|
51
|
-
Every escape hatch the substrate exposes is reachable from
|
|
52
|
-
`selfImprove` — custom `driver`, custom `gate`, distributed-driver
|
|
53
|
-
`cellPlacement`, `onProgress` streaming callback, `autoOnPromote: 'pr'`
|
|
54
|
-
to open a GitHub PR with the winner. See the type signatures in
|
|
55
|
-
[`src/contract/self-improve.ts`](../src/contract/self-improve.ts) for
|
|
56
|
-
the full surface.
|
|
57
|
-
|
|
58
|
-
The sections below are the lower-level path — useful when you want
|
|
59
|
-
fine-grained control over each piece. Read those next if `selfImprove`
|
|
60
|
-
isn't enough.
|
|
61
|
-
|
|
62
|
-
## Five types, four functions
|
|
63
|
-
|
|
64
|
-
```ts
|
|
65
|
-
import {
|
|
66
|
-
// Types
|
|
67
|
-
type Scenario, // what you evaluate against (id + kind + your fields)
|
|
68
|
-
type Dispatch, // your agent, wrapped as one function
|
|
69
|
-
type JudgeConfig, // pluggable dimensional scorer
|
|
70
|
-
type Mutator, // proposes a next surface
|
|
71
|
-
type Gate, // promotion guard
|
|
72
|
-
|
|
73
|
-
// Functions
|
|
74
|
-
runEval,
|
|
75
|
-
runCampaign,
|
|
76
|
-
runImprovementLoop,
|
|
77
|
-
defaultProductionGate,
|
|
78
|
-
|
|
79
|
-
// Storage
|
|
80
|
-
fsCampaignStorage,
|
|
81
|
-
inMemoryCampaignStorage,
|
|
82
|
-
} from '@tangle-network/agent-eval/contract'
|
|
83
|
-
```
|
|
84
|
-
|
|
85
|
-
Every export above is committed under semver. New minors only ADD;
|
|
86
|
-
nothing here changes shape in a 0.x minor.
|
|
87
|
-
|
|
88
|
-
## Three steps to wire your agent
|
|
89
|
-
|
|
90
|
-
### 1. Scenarios
|
|
91
|
-
|
|
92
|
-
```ts
|
|
93
|
-
interface MarketingScenario extends Scenario {
|
|
94
|
-
blurb: string
|
|
95
|
-
surface: 'landing-hero' | 'tweet' | 'email-subject'
|
|
96
|
-
audience: string
|
|
97
|
-
}
|
|
98
|
-
|
|
99
|
-
const scenarios: MarketingScenario[] = [
|
|
100
|
-
{ id: 's1', kind: 'marketing-rewrite', blurb: '...', surface: 'tweet', audience: '...' },
|
|
101
|
-
// ...
|
|
102
|
-
]
|
|
103
|
-
```
|
|
104
|
-
|
|
105
|
-
### 2. Wrap your agent as `Dispatch`
|
|
106
|
-
|
|
107
|
-
```ts
|
|
108
|
-
const dispatch: Dispatch<MarketingScenario, MarketingArtifact> = async (scenario, ctx) => {
|
|
109
|
-
const rewrite = await callYourAgent(scenario, { signal: ctx.signal })
|
|
110
|
-
return { rewrite, modelUsed: '...' }
|
|
111
|
-
}
|
|
112
|
-
```
|
|
113
|
-
|
|
114
|
-
`ctx` carries `signal` (cancellation), `trace` (write spans), `artifacts`
|
|
115
|
-
(write blobs), `cost` (token + $ meter). Use them or ignore them.
|
|
116
|
-
|
|
117
|
-
### 3. Bring a judge
|
|
118
|
-
|
|
119
|
-
```ts
|
|
120
|
-
const judge: JudgeConfig<MarketingArtifact, MarketingScenario> = {
|
|
121
|
-
name: 'marketing-quality',
|
|
122
|
-
dimensions: [
|
|
123
|
-
{ key: 'hook_strength', description: '...' },
|
|
124
|
-
{ key: 'voice_match', description: '...' },
|
|
125
|
-
{ key: 'cta_clarity', description: '...' },
|
|
126
|
-
{ key: 'factual_grounding', description: '...' },
|
|
127
|
-
],
|
|
128
|
-
async score({ artifact, scenario, signal }) {
|
|
129
|
-
// LLM call, heuristic, ensemble — anything. Return JudgeScore.
|
|
130
|
-
return { dimensions: { ... }, composite: 0.72, notes: '...' }
|
|
131
|
-
},
|
|
132
|
-
}
|
|
133
|
-
```
|
|
134
|
-
|
|
135
|
-
Throw on failure; the substrate records it as a failed cell. No silent
|
|
136
|
-
zeros.
|
|
137
|
-
|
|
138
|
-
## Baseline
|
|
139
|
-
|
|
140
|
-
```ts
|
|
141
|
-
const baseline = await runEval({
|
|
142
|
-
scenarios,
|
|
143
|
-
dispatch,
|
|
144
|
-
judges: [judge],
|
|
145
|
-
storage: inMemoryCampaignStorage(),
|
|
146
|
-
runDir: 'mem://my-baseline',
|
|
147
|
-
})
|
|
148
|
-
|
|
149
|
-
const score = Object.values(baseline.aggregates.byScenario)
|
|
150
|
-
.reduce((sum, s) => sum + s.meanComposite, 0) / scenarios.length
|
|
151
|
-
|
|
152
|
-
console.log(`Baseline composite: ${score.toFixed(3)}`)
|
|
153
|
-
```
|
|
154
|
-
|
|
155
|
-
## Self-improvement loop
|
|
156
|
-
|
|
157
|
-
```ts
|
|
158
|
-
import { gepaDriver, defaultProductionGate } from '@tangle-network/agent-eval/contract'
|
|
159
|
-
|
|
160
|
-
const result = await runImprovementLoop({
|
|
161
|
-
scenarios: trainScenarios,
|
|
162
|
-
baselineSurface,
|
|
163
|
-
dispatchWithSurface: (surface, scenario, ctx) =>
|
|
164
|
-
runYourAgent({ systemPrompt: surface as string }, scenario, ctx),
|
|
165
|
-
driver: gepaDriver({
|
|
166
|
-
llm: { apiKey: process.env.OPENAI_API_KEY, baseUrl: '...' },
|
|
167
|
-
model: 'gpt-4o-mini',
|
|
168
|
-
target: 'marketing copywriting system prompt',
|
|
169
|
-
mutationPrimitives: [
|
|
170
|
-
'Tighten the hook: lead with the concrete user outcome.',
|
|
171
|
-
'Replace generic adjectives with specific verbs.',
|
|
172
|
-
// ...
|
|
173
|
-
],
|
|
174
|
-
}),
|
|
175
|
-
judges: [judge],
|
|
176
|
-
populationSize: 2,
|
|
177
|
-
maxGenerations: 3,
|
|
178
|
-
holdoutScenarios,
|
|
179
|
-
gate: defaultProductionGate({
|
|
180
|
-
holdoutScenarios,
|
|
181
|
-
deltaThreshold: 0.05,
|
|
182
|
-
}),
|
|
183
|
-
autoOnPromote: 'none',
|
|
184
|
-
storage: inMemoryCampaignStorage(),
|
|
185
|
-
runDir: 'mem://my-improve',
|
|
186
|
-
})
|
|
187
|
-
|
|
188
|
-
if (result.gateResult.decision === 'ship') {
|
|
189
|
-
// Deploy result.winnerSurface — we don't push it for you.
|
|
190
|
-
}
|
|
191
|
-
```
|
|
192
|
-
|
|
193
|
-
The gate decision is `'ship'` | `'hold'` | `'need_more_work'` |
|
|
194
|
-
`'model_ceiling'` | `'arch_ceiling'`. You define what each means in
|
|
195
|
-
your deploy pipeline.
|
|
196
|
-
|
|
197
|
-
## What you control
|
|
198
|
-
|
|
199
|
-
- The agent (any framework, any model, any backend).
|
|
200
|
-
- The judge (LLM, heuristic, ensemble; we don't pick).
|
|
201
|
-
- The mutation strategy (`gepaDriver` for reflective LLM mutation,
|
|
202
|
-
`evolutionaryDriver({ mutator })` for population search, or
|
|
203
|
-
implement `ImprovementDriver` directly).
|
|
204
|
-
- The gate (compose `defaultProductionGate` with custom checks via
|
|
205
|
-
`composeGate`).
|
|
206
|
-
- The deploy step (`autoOnPromote: 'pr'` opens a GitHub PR with the
|
|
207
|
-
winner; `'none'` returns the surface and you ship however you ship).
|
|
208
|
-
|
|
209
|
-
## What this does NOT install
|
|
210
|
-
|
|
211
|
-
- No `@tangle-network/sandbox` — nothing runs in a Tangle sandbox.
|
|
212
|
-
- No hosted orchestrator — traces, artifacts, judge scores stay on
|
|
213
|
-
your machine (or in `inMemoryCampaignStorage` for Workers/edge).
|
|
214
|
-
- No daemons — `runEval` and `runImprovementLoop` complete in-process
|
|
215
|
-
and return.
|
|
216
|
-
|
|
217
|
-
## When you want more
|
|
218
|
-
|
|
219
|
-
The wedge doc (`docs/design/external-agent-wedge.md`) lays out three
|
|
220
|
-
graduated tiers:
|
|
221
|
-
|
|
222
|
-
| Tier | What you do | What you get |
|
|
223
|
-
|---|---|---|
|
|
224
|
-
| **LAND** (this quickstart) | `npm i @tangle-network/agent-eval`, wrap dispatch + judge, run loops | Local artifacts; full self-improvement; no Tangle infra |
|
|
225
|
-
| **EXPAND** | Point trace/eval data at our hosted orchestrator | Hosted dashboards, cross-run intelligence, billing on data routed to us |
|
|
226
|
-
| **PLATFORM** | Move execution into our sandbox | Substrate + orchestrator data pre-wired; sandbox usage billing |
|
|
227
|
-
|
|
228
|
-
Each tier is opt-in. EXPAND and PLATFORM build on the same primitives;
|
|
229
|
-
upgrading is adding configuration, not rewriting your wiring.
|