raindrop-ai 0.6.0 → 0.7.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +17 -8
- package/dist/chunk-4Z7UHC7X.mjs +14850 -0
- package/dist/{chunk-ZVHQBBST.mjs → chunk-OX6X4ZWQ.mjs} +5 -1
- package/dist/{chunk-CFGEE5ID.mjs → chunk-RMP6BZSY.mjs} +59 -16
- package/dist/evals/cli.d.mts +1 -0
- package/dist/evals/cli.d.ts +1 -0
- package/dist/evals/cli.js +21135 -0
- package/dist/evals/cli.mjs +125 -0
- package/dist/{index-APN1jN-i.d.mts → index-n8AIAiRR.d.mts} +18838 -18626
- package/dist/{index-APN1jN-i.d.ts → index-n8AIAiRR.d.ts} +18838 -18626
- package/dist/index.d.mts +1 -1
- package/dist/index.d.ts +1 -1
- package/dist/index.js +731 -294
- package/dist/index.mjs +57 -14425
- package/dist/{portable-runtime-QZKU2VXS.mjs → portable-runtime-MFVOVFPB.mjs} +7 -2
- package/dist/tracing/index.d.mts +1 -1
- package/dist/tracing/index.d.ts +1 -1
- package/dist/tracing/index.js +5 -1
- package/dist/tracing/index.mjs +1 -1
- package/package.json +5 -1
package/README.md
CHANGED
|
@@ -81,25 +81,32 @@ const client = new Raindrop({
|
|
|
81
81
|
|
|
82
82
|
const suite = defineEvalSuite({
|
|
83
83
|
name: "Support responses",
|
|
84
|
-
dataset: "support-
|
|
84
|
+
dataset: "support-examples",
|
|
85
85
|
run: ({ input }) => supportAgent(input ?? ""),
|
|
86
|
-
evaluators: [{ evaluator: "task-completed"
|
|
86
|
+
evaluators: [{ evaluator: "task-completed" }],
|
|
87
87
|
});
|
|
88
88
|
|
|
89
89
|
try {
|
|
90
|
-
const result = await runEvalSuite(client, suite
|
|
91
|
-
destination: { kind: "raindrop" },
|
|
92
|
-
});
|
|
90
|
+
const result = await runEvalSuite(client, suite);
|
|
93
91
|
console.log(result);
|
|
94
92
|
} finally {
|
|
95
93
|
await client.close();
|
|
96
94
|
}
|
|
97
95
|
```
|
|
98
96
|
|
|
99
|
-
The dataset string names a published, project-scoped eval dataset, not an ordinary custom-upload dataset ID.
|
|
97
|
+
The dataset string names a published, project-scoped eval dataset, not an ordinary custom-upload dataset ID. You can also pass a local `defineDataset` object with rows containing just `id` and `input`. `runEvalSuite` and `evalTests` upload local dataset rows and resolve existing hosted evaluators automatically. Create hosted evaluators in Raindrop and reference their slugs; evaluator source is not uploaded. Dataset versions and evaluator scopes are inferred; Raindrop is the default destination. Existing version pins and update-conflict protections remain in effect.
|
|
98
|
+
|
|
99
|
+
Use `publishEvalDataset` when supplying captured reference traces. Local evaluators can declare `requiresReference: true`, and hosted programs can export `const requiresReference = true`, to reject missing references before agent execution. See the [author/run/compare flow](./EVALS.md) and [Vitest guide](../vitest/README.md).
|
|
100
100
|
|
|
101
101
|
The Query SDK organization `apiKey` authenticates evals, dataset publication/reads, and replay trace uploads. Browser login and production `writeKey` credentials do not replace it. The agent's model credentials and test environment are separate. Running locally describes where your callback executes, not where results are stored.
|
|
102
102
|
|
|
103
|
+
Use the returned `runId` with `readEvalRun`, `evaluateEvalRun`, or `compareEvalRuns`. To execute a queued run, pass `{ runId }` to `runEvalSuite`. Use `traceTools.getOutput(trace)` and `traceTools.getToolCalls(trace, name)` in local evaluators.
|
|
104
|
+
|
|
105
|
+
<details>
|
|
106
|
+
<summary>Legacy replay APIs and tracing internals</summary>
|
|
107
|
+
|
|
108
|
+
New integrations should use the suite functions above. The following lower-level APIs remain available for existing integrations and custom workers.
|
|
109
|
+
|
|
103
110
|
### Trace correlation
|
|
104
111
|
|
|
105
112
|
A replay calls your agent once per row and needs every span and event that call
|
|
@@ -160,7 +167,7 @@ import Raindrop, { replay } from "raindrop-ai";
|
|
|
160
167
|
const raindrop = new Raindrop({ apiKey: process.env.RAINDROP_QUERY_API_KEY });
|
|
161
168
|
|
|
162
169
|
const result = await replay(raindrop, {
|
|
163
|
-
dataset: "support-
|
|
170
|
+
dataset: "support-examples",
|
|
164
171
|
run: (row) => supportAgent(row.input),
|
|
165
172
|
evaluators: ["concise", "task-completed"],
|
|
166
173
|
});
|
|
@@ -190,7 +197,7 @@ type appear consistently, then provide its local `judge`:
|
|
|
190
197
|
|
|
191
198
|
```ts
|
|
192
199
|
const result = await replay(raindrop, {
|
|
193
|
-
dataset: "support-
|
|
200
|
+
dataset: "support-examples",
|
|
194
201
|
run: (row) => supportAgent(row.input),
|
|
195
202
|
evaluators: [
|
|
196
203
|
{
|
|
@@ -268,6 +275,8 @@ agent.
|
|
|
268
275
|
provider, so `replay`'s own flush cannot reach your span processor. Flush your
|
|
269
276
|
provider yourself, or keep its batch delay well under `traceWaitMs`.
|
|
270
277
|
|
|
278
|
+
</details>
|
|
279
|
+
|
|
271
280
|
## Detached sub-agents
|
|
272
281
|
|
|
273
282
|
A **detached** sub-agent runs in another process (a queue worker, a container,
|