raindrop-ai 0.6.0 → 0.7.0-otelv2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -81,25 +81,32 @@ const client = new Raindrop({
81
81
 
82
82
  const suite = defineEvalSuite({
83
83
  name: "Support responses",
84
- dataset: "support-cases",
84
+ dataset: "support-examples",
85
85
  run: ({ input }) => supportAgent(input ?? ""),
86
- evaluators: [{ evaluator: "task-completed", threshold: { equals: true } }],
86
+ evaluators: [{ evaluator: "task-completed" }],
87
87
  });
88
88
 
89
89
  try {
90
- const result = await runEvalSuite(client, suite, {
91
- destination: { kind: "raindrop" },
92
- });
90
+ const result = await runEvalSuite(client, suite);
93
91
  console.log(result);
94
92
  } finally {
95
93
  await client.close();
96
94
  }
97
95
  ```
98
96
 
99
- The dataset string names a published, project-scoped eval dataset, not an ordinary custom-upload dataset ID. Publish its immutable version with `publishEvalDataset`. Cases can contain only input and scenario properties; an optional reference must be a complete captured trace. Local evaluators can declare `requiresReference: true`, and hosted programs can export `const requiresReference = true`, to reject missing references before agent execution. A local `defineDataset` object does not publish itself and is not accepted by the hosted runner. See the [Vitest guide](../vitest/README.md) for publication, evaluator requirements, and credential setup.
97
+ The dataset string names a published, project-scoped eval dataset, not an ordinary custom-upload dataset ID. You can also pass a local `defineDataset` object with rows containing just `id` and `input`. `runEvalSuite` and `evalTests` upload local dataset rows and resolve existing hosted evaluators automatically. Create hosted evaluators in Raindrop and reference their slugs; evaluator source is not uploaded. Dataset versions and evaluator scopes are inferred; Raindrop is the default destination. Existing version pins and update-conflict protections remain in effect.
98
+
99
+ Use `publishEvalDataset` when supplying captured reference traces. Local evaluators can declare `requiresReference: true`, and hosted programs can export `const requiresReference = true`, to reject missing references before agent execution. See the [author/run/compare flow](./EVALS.md) and [Vitest guide](../vitest/README.md).
100
100
 
101
101
  The Query SDK organization `apiKey` authenticates evals, dataset publication/reads, and replay trace uploads. Browser login and production `writeKey` credentials do not replace it. The agent's model credentials and test environment are separate. Running locally describes where your callback executes, not where results are stored.
102
102
 
103
+ Use the returned `runId` with `readEvalRun`, `evaluateEvalRun`, or `compareEvalRuns`. To execute a queued run, pass `{ runId }` to `runEvalSuite`. Use `traceTools.getOutput(trace)` and `traceTools.getToolCalls(trace, name)` in local evaluators.
104
+
105
+ <details>
106
+ <summary>Legacy replay APIs and tracing internals</summary>
107
+
108
+ New integrations should use the suite functions above. The following lower-level APIs remain available for existing integrations and custom workers.
109
+
103
110
  ### Trace correlation
104
111
 
105
112
  A replay calls your agent once per row and needs every span and event that call
@@ -160,7 +167,7 @@ import Raindrop, { replay } from "raindrop-ai";
160
167
  const raindrop = new Raindrop({ apiKey: process.env.RAINDROP_QUERY_API_KEY });
161
168
 
162
169
  const result = await replay(raindrop, {
163
- dataset: "support-cases",
170
+ dataset: "support-examples",
164
171
  run: (row) => supportAgent(row.input),
165
172
  evaluators: ["concise", "task-completed"],
166
173
  });
@@ -190,7 +197,7 @@ type appear consistently, then provide its local `judge`:
190
197
 
191
198
  ```ts
192
199
  const result = await replay(raindrop, {
193
- dataset: "support-cases",
200
+ dataset: "support-examples",
194
201
  run: (row) => supportAgent(row.input),
195
202
  evaluators: [
196
203
  {
@@ -268,6 +275,8 @@ agent.
268
275
  provider, so `replay`'s own flush cannot reach your span processor. Flush your
269
276
  provider yourself, or keep its batch delay well under `traceWaitMs`.
270
277
 
278
+ </details>
279
+
271
280
  ## Detached sub-agents
272
281
 
273
282
  A **detached** sub-agent runs in another process (a queue worker, a container,