raindrop-ai 0.4.0 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -65,6 +65,174 @@ On runtimes that lack a usable `AsyncLocalStorage.enterWith` (notably Cloudflare
65
65
  - `withSpan`/`withTool` invoke your callback as `fn.apply(thisArg, args)` — the callback receives exactly the arguments you pass (an earlier version prepended a spurious leading `undefined`). If you had code compensating for that shifted argument, remove the workaround.
66
66
  - With per-client interaction registries, `resumeInteraction(eventId)` returns only interactions that **this** client began; a second client resuming the same event id gets its own fresh interaction rather than the first client's.
67
67
 
68
+ ## Eval mode
69
+
70
+ A replay calls your agent once per row and needs every span and event that call
71
+ emits to land on that row. `withEvalScope` binds one opaque **correlation id**
72
+ to the execution context, so it does, with no change to the agent:
73
+
74
+ ```ts
75
+ import { withEvalScope } from "raindrop-ai";
76
+
77
+ for (const row of rows) {
78
+ await withEvalScope({ correlationId: row.correlationId }, () => supportAgent(row.input));
79
+ }
80
+ ```
81
+
82
+ Everything started inside the callback carries the id: spans you create with
83
+ `withSpan`/`withTool`, auto-instrumented LLM spans from libraries you never call
84
+ directly, and events from `trackAi`. The agent itself is unchanged and never
85
+ sees a correlation id.
86
+
87
+ `currentEvalScope()` reads the scope if you need it, and returns `undefined`
88
+ outside one. Nothing changes for a caller who never enters a scope. Spans and
89
+ events carry exactly the attributes they carried before.
90
+
91
+ **Concurrency.** Rows run concurrently and each keeps its own id. The scope is
92
+ bound in `AsyncLocalStorage`, so two agents in flight at once never see each
93
+ other's row.
94
+
95
+ **Nesting replaces, it does not merge.** A correlation id names exactly one
96
+ row, and a span belongs to exactly one row, so a nested `withEvalScope` wins for
97
+ its own duration and the outer scope is restored when it returns. Merging two
98
+ ids would produce a third identity that is neither.
99
+
100
+ ### Known limits
101
+
102
+ **Work whose async context predates the scope carries no id.** A continuation
103
+ attached before the row started, a background flush loop, or a pooled
104
+ connection's pending callback runs outside the scope even when the row triggers
105
+ it. Those spans land in normal traffic.
106
+
107
+ **Work detached inside the scope keeps the id and can outlive the row.** This
108
+ is the opposite of what it sounds like. `AsyncLocalStorage` propagates into
109
+ timers and unawaited promises, so a detached task started inside the callback
110
+ still carries the correlation id when it emits, possibly after the row has been
111
+ recorded. Narrow it by awaiting everything the row starts before the callback
112
+ returns.
113
+
114
+ Both limits are demonstrated by tests in `packages/core/src/eval-scope.test.ts`.
115
+
116
+ ### Replaying a dataset
117
+
118
+ `replay` is the loop that uses the scope. It asks Raindrop for the dataset's
119
+ rows, calls your agent once per row inside that row's scope, says which rows
120
+ failed, and waits for the traces to land:
121
+
122
+ ```ts
123
+ import Raindrop, { replay } from "raindrop-ai";
124
+
125
+ const raindrop = new Raindrop({ apiKey: process.env.RAINDROP_QUERY_API_KEY });
126
+
127
+ const result = await replay(raindrop, {
128
+ dataset: "support-cases",
129
+ run: (row) => supportAgent(row.input),
130
+ evaluators: ["concise", "task-completed"],
131
+ });
132
+
133
+ for (const row of result.rows) {
134
+ console.log(row.id, row.status);
135
+ for (const verdict of row.verdicts) {
136
+ console.log(verdict.evaluator, verdict);
137
+ }
138
+ }
139
+ ```
140
+
141
+ After the traces land, `replay` starts each requested evaluator against the
142
+ replay and waits for its result. Each item in `result.rows` contains the
143
+ original dataset row, its trace status, and one verdict per evaluator in
144
+ `row.verdicts`. A verdict identifies its `evaluator` and contains either the
145
+ grade (`pass`, `score`, or `value`), an evaluator error, or the reason it was
146
+ not graded.
147
+
148
+ `result.evaluators` contains aggregate status, summary, error, and URL for each
149
+ evaluator. A failed evaluator is returned there instead of making the whole
150
+ replay throw.
151
+
152
+ An evaluator can run beside the agent when its verdict depends on state that is
153
+ not in the trace. Keep the evaluator saved in Raindrop so its name and output
154
+ type appear consistently, then provide its local `judge`:
155
+
156
+ ```ts
157
+ const result = await replay(raindrop, {
158
+ dataset: "support-cases",
159
+ run: (row) => supportAgent(row.input),
160
+ evaluators: [
161
+ {
162
+ evaluator: "ticket-closed",
163
+ judge: async ({ row }) => {
164
+ const ticket = await testEnvironment.tickets.get(row.properties.ticketId);
165
+ return {
166
+ pass: ticket.status === "closed",
167
+ note: `final status: ${ticket.status}`,
168
+ };
169
+ },
170
+ },
171
+ ],
172
+ });
173
+ ```
174
+
175
+ The callback receives the dataset `row` and the value returned by `run`. It
176
+ runs immediately after the row finishes, while the local environment is still
177
+ available. Return `{ pass, note? }`, `{ score, note? }`, or `{ value, note? }`
178
+ to match the saved evaluator's output type. A thrown error is stored as that
179
+ row's errored verdict. Raindrop marks the resulting run as `local` in the UI.
180
+
181
+ If rows mutate one shared test environment, set `concurrency: 1`. The default
182
+ of four is appropriate when each row has isolated state.
183
+
184
+ The row counts are computed on the server, so they match what the web screen
185
+ shows. By default, evaluators are checked every two seconds for up to ten
186
+ minutes. Use `evalPollIntervalMs` and `evalWaitMs` to change those waits.
187
+
188
+ **A row that fails is `missing`, never dropped.** A callback that throws is
189
+ reported with its error, and so is a row whose agent returned but whose spans
190
+ never shipped, after `traceWaitMs` (30 seconds by default). A replay in which
191
+ every row failed still finishes, with `missing` equal to the row count. A
192
+ replay that quietly dropped rows would report a better number than it earned.
193
+
194
+ **Concurrency** defaults to four rows at once. Each row keeps its own
195
+ correlation id, so a span never lands on the wrong row.
196
+
197
+ `createReplay` and `runReplay` are the trace-only halves if you want them
198
+ separately. Call `evaluateReplay(raindrop, replayId, { evaluators })` after
199
+ `runReplay` to apply the evaluators and wait for their aggregate results.
200
+
201
+ A newly created replay also returns `ingestUrl`, the dedicated replay endpoint.
202
+ `runReplay` binds it around each agent call automatically. Spans created by
203
+ those calls use the client's Query SDK `apiKey` and carry the replay id and project in
204
+ the `X-Raindrop-Replay-Id` and `X-Raindrop-Project-Id` headers. Dataset and eval control
205
+ calls use that same organization API key. A `writeKey` is needed only if the client
206
+ also sends ordinary production telemetry. Existing eval callers must migrate from
207
+ `writeKey` to `apiKey`; write keys no longer authorize evals or datasets.
208
+
209
+ Deploy backend support for SDK-key replay ingest before upgrading eval consumers.
210
+ Older SDK eval clients must migrate when the backend changes. Production telemetry
211
+ continues to use `writeKey`; eval requests do not fall back to it.
212
+
213
+ **Running a replay somebody started in the dashboard.** A replay created there
214
+ is queued with no runner, because only your process can call your agent.
215
+ `claimReplay` picks one up and returns the same handle `createReplay` does, so
216
+ it goes straight to `runReplay`:
217
+
218
+ ```ts
219
+ import Raindrop, { claimReplay, runReplay } from "raindrop-ai";
220
+
221
+ const handle = await claimReplay(raindrop, replayId);
222
+ const result = await runReplay(raindrop, handle, {
223
+ run: (row) => supportAgent(row.input),
224
+ });
225
+ ```
226
+
227
+ The process that acquires the runner lease gets `ingestUrl`. A converged create
228
+ retry or later claim gets `ingestUrl: null`, so a second process cannot run the
229
+ same replay concurrently. `runReplay` refuses that handle before calling the
230
+ agent.
231
+
232
+ **One limit worth knowing.** In `useExternalOtel` mode you own the tracer
233
+ provider, so `replay`'s own flush cannot reach your span processor. Flush your
234
+ provider yourself, or keep its batch delay well under `traceWaitMs`.
235
+
68
236
  ## Detached sub-agents
69
237
 
70
238
  A **detached** sub-agent runs in another process (a queue worker, a container,