raindrop-ai 0.7.1 → 0.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +69 -1
- package/dist/{chunk-6WL655QL.mjs → chunk-22CNYHSA.mjs} +399 -110
- package/dist/{chunk-RMP6BZSY.mjs → chunk-EFATLUNL.mjs} +6 -4
- package/dist/{chunk-SHH4SYVY.mjs → chunk-EWHJ2YBW.mjs} +1 -1
- package/dist/evals/cli.js +207 -153
- package/dist/evals/cli.mjs +3 -3
- package/dist/{index-DuJkG47V.d.ts → index-gf1epf_P.d.mts} +186 -86
- package/dist/{index-DuJkG47V.d.mts → index-gf1epf_P.d.ts} +186 -86
- package/dist/index.d.mts +1 -1
- package/dist/index.d.ts +1 -1
- package/dist/index.js +466 -172
- package/dist/index.mjs +9 -3
- package/dist/{portable-runtime-MFVOVFPB.mjs → portable-runtime-6X2HHB4Y.mjs} +1 -1
- package/dist/tracing/index.d.mts +1 -1
- package/dist/tracing/index.d.ts +1 -1
- package/dist/tracing/index.js +1 -1
- package/dist/tracing/index.mjs +1 -1
- package/package.json +2 -2
package/README.md
CHANGED
|
@@ -127,6 +127,74 @@ The Query SDK organization `apiKey` authenticates evals, dataset publication/rea
|
|
|
127
127
|
|
|
128
128
|
Use the returned `runId` with `readEvalRun`, `evaluateEvalRun`, or `compareEvalRuns`. To execute a queued run, pass `{ runId }` to `runEvalSuite`. Use `traceTools.getOutput(trace)` and `traceTools.getToolCalls(trace, name)` in local evaluators.
|
|
129
129
|
|
|
130
|
+
### Run eval agents on your server
|
|
131
|
+
|
|
132
|
+
Protect the route with your existing authentication before calling `handleEval`. The handler does not authenticate incoming requests. Run it on a server that can await the entire execution. Do not return early from a serverless request.
|
|
133
|
+
|
|
134
|
+
The simplest agent is a slug and a `run` function:
|
|
135
|
+
|
|
136
|
+
```ts
|
|
137
|
+
import { createEvalHandler, defineAgent } from "raindrop-ai";
|
|
138
|
+
|
|
139
|
+
const support = defineAgent({
|
|
140
|
+
slug: "support",
|
|
141
|
+
run: (row) => supportAgent(row.input),
|
|
142
|
+
});
|
|
143
|
+
|
|
144
|
+
const handleEval = createEvalHandler(client, { agents: [support] });
|
|
145
|
+
|
|
146
|
+
export async function route(request: Request): Promise<Response> {
|
|
147
|
+
if (!(await isAuthorized(request))) return new Response("Unauthorized", { status: 401 });
|
|
148
|
+
return handleEval(request);
|
|
149
|
+
}
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
Declare `parameters` to vary the agent per run without redeploying. Each key is a `z.string()`, `z.number()`, `z.boolean()`, or `z.enum()` schema, optionally with `.default()`, `.optional()`, and `.describe()`. Raindrop reads the schema from `GET`, renders one form control per key, and saves each edit as a new version of a named parameter set, so runs pin an exact version. The handler validates incoming values and applies defaults before any row runs; a mismatch returns 400 with the schema error instead of claiming the run. `context.parameters` is typed from the declaration. Agents without `parameters` accept only `{}`.
|
|
153
|
+
|
|
154
|
+
```ts
|
|
155
|
+
import { z } from "zod";
|
|
156
|
+
|
|
157
|
+
const support = defineAgent({
|
|
158
|
+
slug: "support",
|
|
159
|
+
parameters: {
|
|
160
|
+
model: z.enum(["gpt-4o", "gpt-4o-mini"]).default("gpt-4o"),
|
|
161
|
+
temperature: z.number().min(0).max(2).default(0.2),
|
|
162
|
+
systemPrompt: z.string().max(4000).describe("Prepended to every conversation"),
|
|
163
|
+
},
|
|
164
|
+
run: (row, { parameters }) => supportAgent(row.input, parameters),
|
|
165
|
+
});
|
|
166
|
+
```
|
|
167
|
+
|
|
168
|
+
Add `setup` when rows need shared infrastructure. It runs once per execution with the run's parameters, every row receives its result as `environment`, and `cleanup` receives the same value once the rows settle, including after failures. `cleanup` does not run when `setup` itself throws. Give temporary infrastructure an expiry because a terminated process cannot execute cleanup. Concurrent rows share the environment, so isolate row state or use an agent that supports concurrent requests.
|
|
169
|
+
|
|
170
|
+
```ts
|
|
171
|
+
const sandboxed = defineAgent({
|
|
172
|
+
slug: "support-sandboxed",
|
|
173
|
+
name: "Support agent in a sandbox",
|
|
174
|
+
parameters: { image: z.string().default("support:latest") },
|
|
175
|
+
setup: (parameters) => createSandbox(parameters.image),
|
|
176
|
+
run: (row, { environment }) => supportAgent(row.input, environment),
|
|
177
|
+
cleanup: (environment) => environment.dispose(),
|
|
178
|
+
});
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
The same agent runs locally through a suite. Set either `agent` or `run`, not both; `runEvalSuite` and `createEvalSuiteRun` accept `parameters` and start one execution per run.
|
|
182
|
+
|
|
183
|
+
```ts
|
|
184
|
+
const suite = defineEvalSuite({
|
|
185
|
+
name: "Support regression",
|
|
186
|
+
dataset: "support-examples",
|
|
187
|
+
agent: support,
|
|
188
|
+
evaluators: [{ evaluator: "task-completed" }],
|
|
189
|
+
});
|
|
190
|
+
|
|
191
|
+
await runEvalSuite(client, suite, { parameters: { model: "gpt-4o-mini", systemPrompt: "Be brief" } });
|
|
192
|
+
```
|
|
193
|
+
|
|
194
|
+
GET returns registered agent metadata and parameter fields. POST accepts `{ runId, agent, parameters }`. Raindrop creates the queued run and pins its dataset before dispatch. The handler claims the run with your client's Query SDK credentials, executes the agent over the claimed rows, and uploads its traces. Raindrop grades them separately with the selected hosted evaluators.
|
|
195
|
+
|
|
196
|
+
The handler returns `{ runId, status, counts, traceEvidence }` after execution. Repeated requests for terminal runs return their existing status and counts without rerunning. A run claimed by another process returns 409. Execution failures return 500 with the thrown error's message in `error`, which Raindrop shows on the run. `queryUrl` and `replayIngestUrl` are server configuration, never request fields.
|
|
197
|
+
|
|
130
198
|
<details>
|
|
131
199
|
<summary>Legacy replay APIs and tracing internals</summary>
|
|
132
200
|
|
|
@@ -254,7 +322,7 @@ minutes. Use `evalPollIntervalMs` and `evalWaitMs` to change those waits.
|
|
|
254
322
|
|
|
255
323
|
**A row that fails is `missing`, never dropped.** A callback that throws is
|
|
256
324
|
reported with its error, and so is a row whose agent returned but whose spans
|
|
257
|
-
never shipped, after `traceWaitMs` (
|
|
325
|
+
never shipped, after `traceWaitMs` (60 seconds by default). A replay in which
|
|
258
326
|
every row failed still finishes, with `missing` equal to the row count. A
|
|
259
327
|
replay that quietly dropped rows would report a better number than it earned.
|
|
260
328
|
|