raindrop-ai 0.4.0 → 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +168 -0
- package/dist/{chunk-MZZJWWWB.mjs → chunk-AUNFORWV.mjs} +1213 -860
- package/dist/chunk-HFHHGRTW.mjs +1201 -0
- package/dist/index-CQj4bn2v.d.mts +60760 -0
- package/dist/index-CQj4bn2v.d.ts +60760 -0
- package/dist/index.d.mts +2 -1
- package/dist/index.d.ts +2 -1
- package/dist/index.js +11422 -6431
- package/dist/index.mjs +3397 -176
- package/dist/portable-runtime-GBE3SHBZ.mjs +193 -0
- package/dist/tracing/index.d.mts +2 -1
- package/dist/tracing/index.d.ts +2 -1
- package/dist/tracing/index.js +1301 -982
- package/dist/tracing/index.mjs +1 -1
- package/package.json +19 -12
- package/dist/index-CDmWtd83.d.mts +0 -1525
- package/dist/index-CDmWtd83.d.ts +0 -1525
package/README.md
CHANGED
|
@@ -65,6 +65,174 @@ On runtimes that lack a usable `AsyncLocalStorage.enterWith` (notably Cloudflare
|
|
|
65
65
|
- `withSpan`/`withTool` invoke your callback as `fn.apply(thisArg, args)` — the callback receives exactly the arguments you pass (an earlier version prepended a spurious leading `undefined`). If you had code compensating for that shifted argument, remove the workaround.
|
|
66
66
|
- With per-client interaction registries, `resumeInteraction(eventId)` returns only interactions that **this** client began; a second client resuming the same event id gets its own fresh interaction rather than the first client's.
|
|
67
67
|
|
|
68
|
+
## Eval mode
|
|
69
|
+
|
|
70
|
+
A replay calls your agent once per row and needs every span and event that call
|
|
71
|
+
emits to land on that row. `withEvalScope` binds one opaque **correlation id**
|
|
72
|
+
to the execution context, so it does, with no change to the agent:
|
|
73
|
+
|
|
74
|
+
```ts
|
|
75
|
+
import { withEvalScope } from "raindrop-ai";
|
|
76
|
+
|
|
77
|
+
for (const row of rows) {
|
|
78
|
+
await withEvalScope({ correlationId: row.correlationId }, () => supportAgent(row.input));
|
|
79
|
+
}
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
Everything started inside the callback carries the id: spans you create with
|
|
83
|
+
`withSpan`/`withTool`, auto-instrumented LLM spans from libraries you never call
|
|
84
|
+
directly, and events from `trackAi`. The agent itself is unchanged and never
|
|
85
|
+
sees a correlation id.
|
|
86
|
+
|
|
87
|
+
`currentEvalScope()` reads the scope if you need it, and returns `undefined`
|
|
88
|
+
outside one. Nothing changes for a caller who never enters a scope. Spans and
|
|
89
|
+
events carry exactly the attributes they carried before.
|
|
90
|
+
|
|
91
|
+
**Concurrency.** Rows run concurrently and each keeps its own id. The scope is
|
|
92
|
+
bound in `AsyncLocalStorage`, so two agents in flight at once never see each
|
|
93
|
+
other's row.
|
|
94
|
+
|
|
95
|
+
**Nesting replaces, it does not merge.** A correlation id names exactly one
|
|
96
|
+
row, and a span belongs to exactly one row, so a nested `withEvalScope` wins for
|
|
97
|
+
its own duration and the outer scope is restored when it returns. Merging two
|
|
98
|
+
ids would produce a third identity that is neither.
|
|
99
|
+
|
|
100
|
+
### Known limits
|
|
101
|
+
|
|
102
|
+
**Work whose async context predates the scope carries no id.** A continuation
|
|
103
|
+
attached before the row started, a background flush loop, or a pooled
|
|
104
|
+
connection's pending callback runs outside the scope even when the row triggers
|
|
105
|
+
it. Those spans land in normal traffic.
|
|
106
|
+
|
|
107
|
+
**Work detached inside the scope keeps the id and can outlive the row.** This
|
|
108
|
+
is the opposite of what it sounds like. `AsyncLocalStorage` propagates into
|
|
109
|
+
timers and unawaited promises, so a detached task started inside the callback
|
|
110
|
+
still carries the correlation id when it emits, possibly after the row has been
|
|
111
|
+
recorded. Narrow it by awaiting everything the row starts before the callback
|
|
112
|
+
returns.
|
|
113
|
+
|
|
114
|
+
Both limits are demonstrated by tests in `packages/core/src/eval-scope.test.ts`.
|
|
115
|
+
|
|
116
|
+
### Replaying a dataset
|
|
117
|
+
|
|
118
|
+
`replay` is the loop that uses the scope. It asks Raindrop for the dataset's
|
|
119
|
+
rows, calls your agent once per row inside that row's scope, says which rows
|
|
120
|
+
failed, and waits for the traces to land:
|
|
121
|
+
|
|
122
|
+
```ts
|
|
123
|
+
import Raindrop, { replay } from "raindrop-ai";
|
|
124
|
+
|
|
125
|
+
const raindrop = new Raindrop({ apiKey: process.env.RAINDROP_QUERY_API_KEY });
|
|
126
|
+
|
|
127
|
+
const result = await replay(raindrop, {
|
|
128
|
+
dataset: "support-cases",
|
|
129
|
+
run: (row) => supportAgent(row.input),
|
|
130
|
+
evaluators: ["concise", "task-completed"],
|
|
131
|
+
});
|
|
132
|
+
|
|
133
|
+
for (const row of result.rows) {
|
|
134
|
+
console.log(row.id, row.status);
|
|
135
|
+
for (const verdict of row.verdicts) {
|
|
136
|
+
console.log(verdict.evaluator, verdict);
|
|
137
|
+
}
|
|
138
|
+
}
|
|
139
|
+
```
|
|
140
|
+
|
|
141
|
+
After the traces land, `replay` starts each requested evaluator against the
|
|
142
|
+
replay and waits for its result. Each item in `result.rows` contains the
|
|
143
|
+
original dataset row, its trace status, and one verdict per evaluator in
|
|
144
|
+
`row.verdicts`. A verdict identifies its `evaluator` and contains either the
|
|
145
|
+
grade (`pass`, `score`, or `value`), an evaluator error, or the reason it was
|
|
146
|
+
not graded.
|
|
147
|
+
|
|
148
|
+
`result.evaluators` contains aggregate status, summary, error, and URL for each
|
|
149
|
+
evaluator. A failed evaluator is returned there instead of making the whole
|
|
150
|
+
replay throw.
|
|
151
|
+
|
|
152
|
+
An evaluator can run beside the agent when its verdict depends on state that is
|
|
153
|
+
not in the trace. Keep the evaluator saved in Raindrop so its name and output
|
|
154
|
+
type appear consistently, then provide its local `judge`:
|
|
155
|
+
|
|
156
|
+
```ts
|
|
157
|
+
const result = await replay(raindrop, {
|
|
158
|
+
dataset: "support-cases",
|
|
159
|
+
run: (row) => supportAgent(row.input),
|
|
160
|
+
evaluators: [
|
|
161
|
+
{
|
|
162
|
+
evaluator: "ticket-closed",
|
|
163
|
+
judge: async ({ row }) => {
|
|
164
|
+
const ticket = await testEnvironment.tickets.get(row.properties.ticketId);
|
|
165
|
+
return {
|
|
166
|
+
pass: ticket.status === "closed",
|
|
167
|
+
note: `final status: ${ticket.status}`,
|
|
168
|
+
};
|
|
169
|
+
},
|
|
170
|
+
},
|
|
171
|
+
],
|
|
172
|
+
});
|
|
173
|
+
```
|
|
174
|
+
|
|
175
|
+
The callback receives the dataset `row` and the value returned by `run`. It
|
|
176
|
+
runs immediately after the row finishes, while the local environment is still
|
|
177
|
+
available. Return `{ pass, note? }`, `{ score, note? }`, or `{ value, note? }`
|
|
178
|
+
to match the saved evaluator's output type. A thrown error is stored as that
|
|
179
|
+
row's errored verdict. Raindrop marks the resulting run as `local` in the UI.
|
|
180
|
+
|
|
181
|
+
If rows mutate one shared test environment, set `concurrency: 1`. The default
|
|
182
|
+
of four is appropriate when each row has isolated state.
|
|
183
|
+
|
|
184
|
+
The row counts are computed on the server, so they match what the web screen
|
|
185
|
+
shows. By default, evaluators are checked every two seconds for up to ten
|
|
186
|
+
minutes. Use `evalPollIntervalMs` and `evalWaitMs` to change those waits.
|
|
187
|
+
|
|
188
|
+
**A row that fails is `missing`, never dropped.** A callback that throws is
|
|
189
|
+
reported with its error, and so is a row whose agent returned but whose spans
|
|
190
|
+
never shipped, after `traceWaitMs` (30 seconds by default). A replay in which
|
|
191
|
+
every row failed still finishes, with `missing` equal to the row count. A
|
|
192
|
+
replay that quietly dropped rows would report a better number than it earned.
|
|
193
|
+
|
|
194
|
+
**Concurrency** defaults to four rows at once. Each row keeps its own
|
|
195
|
+
correlation id, so a span never lands on the wrong row.
|
|
196
|
+
|
|
197
|
+
`createReplay` and `runReplay` are the trace-only halves if you want them
|
|
198
|
+
separately. Call `evaluateReplay(raindrop, replayId, { evaluators })` after
|
|
199
|
+
`runReplay` to apply the evaluators and wait for their aggregate results.
|
|
200
|
+
|
|
201
|
+
A newly created replay also returns `ingestUrl`, the dedicated replay endpoint.
|
|
202
|
+
`runReplay` binds it around each agent call automatically. Spans created by
|
|
203
|
+
those calls use the client's Query SDK `apiKey` and carry the replay id and project in
|
|
204
|
+
the `X-Raindrop-Replay-Id` and `X-Raindrop-Project-Id` headers. Dataset and eval control
|
|
205
|
+
calls use that same organization API key. A `writeKey` is needed only if the client
|
|
206
|
+
also sends ordinary production telemetry. Existing eval callers must migrate from
|
|
207
|
+
`writeKey` to `apiKey`; write keys no longer authorize evals or datasets.
|
|
208
|
+
|
|
209
|
+
Deploy backend support for SDK-key replay ingest before upgrading eval consumers.
|
|
210
|
+
Older SDK eval clients must migrate when the backend changes. Production telemetry
|
|
211
|
+
continues to use `writeKey`; eval requests do not fall back to it.
|
|
212
|
+
|
|
213
|
+
**Running a replay somebody started in the dashboard.** A replay created there
|
|
214
|
+
is queued with no runner, because only your process can call your agent.
|
|
215
|
+
`claimReplay` picks one up and returns the same handle `createReplay` does, so
|
|
216
|
+
it goes straight to `runReplay`:
|
|
217
|
+
|
|
218
|
+
```ts
|
|
219
|
+
import Raindrop, { claimReplay, runReplay } from "raindrop-ai";
|
|
220
|
+
|
|
221
|
+
const handle = await claimReplay(raindrop, replayId);
|
|
222
|
+
const result = await runReplay(raindrop, handle, {
|
|
223
|
+
run: (row) => supportAgent(row.input),
|
|
224
|
+
});
|
|
225
|
+
```
|
|
226
|
+
|
|
227
|
+
The process that acquires the runner lease gets `ingestUrl`. A converged create
|
|
228
|
+
retry or later claim gets `ingestUrl: null`, so a second process cannot run the
|
|
229
|
+
same replay concurrently. `runReplay` refuses that handle before calling the
|
|
230
|
+
agent.
|
|
231
|
+
|
|
232
|
+
**One limit worth knowing.** In `useExternalOtel` mode you own the tracer
|
|
233
|
+
provider, so `replay`'s own flush cannot reach your span processor. Flush your
|
|
234
|
+
provider yourself, or keep its batch delay well under `traceWaitMs`.
|
|
235
|
+
|
|
68
236
|
## Detached sub-agents
|
|
69
237
|
|
|
70
238
|
A **detached** sub-agent runs in another process (a queue worker, a container,
|