@databricks/appkit 0.73.0 → 0.74.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (60) hide show
  1. package/CLAUDE.md +7 -0
  2. package/dist/appkit/package.js +1 -1
  3. package/dist/beta.d.ts +6 -6
  4. package/dist/beta.js +6 -6
  5. package/dist/cli/commands/agent/eval.js +85 -21
  6. package/dist/cli/commands/agent/eval.js.map +1 -1
  7. package/dist/cli/commands/registry/add.js +77 -6
  8. package/dist/cli/commands/registry/add.js.map +1 -1
  9. package/dist/evals/dataset.d.ts +14 -1
  10. package/dist/evals/dataset.d.ts.map +1 -1
  11. package/dist/evals/dataset.js +16 -1
  12. package/dist/evals/dataset.js.map +1 -1
  13. package/dist/evals/define-eval.d.ts +4 -2
  14. package/dist/evals/define-eval.d.ts.map +1 -1
  15. package/dist/evals/define-eval.js +5 -1
  16. package/dist/evals/define-eval.js.map +1 -1
  17. package/dist/evals/discover.d.ts +15 -1
  18. package/dist/evals/discover.d.ts.map +1 -1
  19. package/dist/evals/discover.js +40 -10
  20. package/dist/evals/discover.js.map +1 -1
  21. package/dist/evals/http-driver.d.ts.map +1 -1
  22. package/dist/evals/http-driver.js +31 -9
  23. package/dist/evals/http-driver.js.map +1 -1
  24. package/dist/evals/index.d.ts +5 -5
  25. package/dist/evals/index.js +5 -5
  26. package/dist/evals/report.d.ts +16 -1
  27. package/dist/evals/report.d.ts.map +1 -1
  28. package/dist/evals/report.js +64 -2
  29. package/dist/evals/report.js.map +1 -1
  30. package/dist/evals/run-eval.d.ts +5 -0
  31. package/dist/evals/run-eval.d.ts.map +1 -1
  32. package/dist/evals/run-eval.js +44 -3
  33. package/dist/evals/run-eval.js.map +1 -1
  34. package/dist/evals/run-evals.d.ts +32 -3
  35. package/dist/evals/run-evals.d.ts.map +1 -1
  36. package/dist/evals/run-evals.js +128 -26
  37. package/dist/evals/run-evals.js.map +1 -1
  38. package/dist/evals/types.d.ts +40 -2
  39. package/dist/evals/types.d.ts.map +1 -1
  40. package/dist/registry/manifest-loader.d.ts +1 -1
  41. package/docs/api/appkit/Function.defineEvalConfig.md +18 -0
  42. package/docs/api/appkit/Function.discoverEvalConfigs.md +18 -0
  43. package/docs/api/appkit/Function.formatResultsJUnit.md +18 -0
  44. package/docs/api/appkit/Function.formatResultsJson.md +18 -0
  45. package/docs/api/appkit/Function.runWithRetries.md +28 -0
  46. package/docs/api/appkit/Function.userTurns.md +20 -0
  47. package/docs/api/appkit/Interface.DiscoveredEvalConfig.md +25 -0
  48. package/docs/api/appkit/Interface.DriveResult.md +28 -0
  49. package/docs/api/appkit/Interface.EvalDefinition.md +22 -0
  50. package/docs/api/appkit/Interface.EvalDriver.md +10 -4
  51. package/docs/api/appkit/Interface.EvalResult.md +11 -0
  52. package/docs/api/appkit/Interface.EvalSummary.md +11 -0
  53. package/docs/api/appkit/Interface.RunEvalOptions.md +11 -0
  54. package/docs/api/appkit/Interface.RunEvalsOptions.md +23 -1
  55. package/docs/api/appkit/Interface.TestContext.md +30 -8
  56. package/docs/api/appkit.md +7 -0
  57. package/docs/plugins/agents.md +255 -4
  58. package/llms.txt +7 -0
  59. package/package.json +1 -1
  60. package/sbom.cdx.json +1 -1
@@ -52,6 +52,28 @@ optional description: string;
52
52
 
53
53
  Short human description, shown in reports.
54
54
 
55
+ ***
56
+
57
+ ### tags?[​](#tags "Direct link to tags?")
58
+
59
+ ```ts
60
+ optional tags: string[];
61
+
62
+ ```
63
+
64
+ Free-form tags for filtering (see the runner's `tags` / `--tag` option).
65
+
66
+ ***
67
+
68
+ ### timeoutMs?[​](#timeoutms "Direct link to timeoutMs?")
69
+
70
+ ```ts
71
+ optional timeoutMs: number;
72
+
73
+ ```
74
+
75
+ Per-eval timeout (ms): `runEval` races the test against it and records a non-passing result instead of hanging. Overrides the runner/CLI default.
76
+
55
77
  ## Methods[​](#methods "Direct link to Methods")
56
78
 
57
79
  ### test()[​](#test "Direct link to test()")
@@ -22,15 +22,21 @@ Drop the current conversation so the next `send` starts a fresh thread. Optional
22
22
  ### send()[​](#send "Direct link to send()")
23
23
 
24
24
  ```ts
25
- send(message: string): Promise<DriveResult>;
25
+ send(message: string, options?: {
26
+ signal?: AbortSignal;
27
+ }): Promise<DriveResult>;
26
28
 
27
29
  ```
28
30
 
31
+ Drive one turn. `options.signal`, when provided, aborts the in-flight turn: the runner passes its per-eval timeout signal so a timed-out eval cancels the request instead of leaking a live stream.
32
+
29
33
  #### Parameters[​](#parameters "Direct link to Parameters")
30
34
 
31
- | Parameter | Type |
32
- | --------- | -------- |
33
- | `message` | `string` |
35
+ | Parameter | Type |
36
+ | ----------------- | ----------------------------- |
37
+ | `message` | `string` |
38
+ | `options?` | { `signal?`: `AbortSignal`; } |
39
+ | `options.signal?` | `AbortSignal` |
34
40
 
35
41
  #### Returns[​](#returns-1 "Direct link to Returns")
36
42
 
@@ -42,6 +42,17 @@ id: string;
42
42
 
43
43
  ***
44
44
 
45
+ ### infraFailure?[​](#infrafailure "Direct link to infraFailure?")
46
+
47
+ ```ts
48
+ optional infraFailure: boolean;
49
+
50
+ ```
51
+
52
+ A turn failed at the transport/agent level (`succeeded: false`), not on an assertion — a retryable infra flake, distinct from `error`.
53
+
54
+ ***
55
+
45
56
  ### passed[​](#passed "Direct link to passed")
46
57
 
47
58
  ```ts
@@ -31,6 +31,17 @@ passed: number;
31
31
 
32
32
  ***
33
33
 
34
+ ### passRate[​](#passrate "Direct link to passRate")
35
+
36
+ ```ts
37
+ passRate: number;
38
+
39
+ ```
40
+
41
+ Fraction of scored (non-skipped) evals that passed, 0..1 (1 when none scored).
42
+
43
+ ***
44
+
34
45
  ### skipped[​](#skipped "Direct link to skipped")
35
46
 
36
47
  ```ts
@@ -43,3 +43,14 @@ optional strict: boolean;
43
43
  ```
44
44
 
45
45
  When true, soft assertion failures also fail the eval.
46
+
47
+ ***
48
+
49
+ ### timeoutMs?[​](#timeoutms "Direct link to timeoutMs?")
50
+
51
+ ```ts
52
+ optional timeoutMs: number;
53
+
54
+ ```
55
+
56
+ Runner-level default per-eval timeout (ms). `def.timeoutMs` wins over this; when both are unset the eval runs unbounded (current behavior).
@@ -160,6 +160,17 @@ Progress callback, invoked as evals are discovered, started, and finished.
160
160
 
161
161
  ***
162
162
 
163
+ ### retries?[​](#retries "Direct link to retries?")
164
+
165
+ ```ts
166
+ optional retries: number;
167
+
168
+ ```
169
+
170
+ Re-run an eval up to this many extra times when it fails on infrastructure — a thrown error/timeout (`result.error`) or a transport/agent turn failure (`result.infraFailure`). Assertion failures are never retried. Defaults to `0`.
171
+
172
+ ***
173
+
163
174
  ### rootDir?[​](#rootdir "Direct link to rootDir?")
164
175
 
165
176
  ```ts
@@ -182,6 +193,17 @@ Soft assertion failures also fail the eval.
182
193
 
183
194
  ***
184
195
 
196
+ ### tags?[​](#tags "Direct link to tags?")
197
+
198
+ ```ts
199
+ optional tags: string[];
200
+
201
+ ```
202
+
203
+ Only run evals whose `tags` intersect this list. Empty/undefined runs all. Tags live on the eval def, so filtering happens after each file is loaded.
204
+
205
+ ***
206
+
185
207
  ### timeoutMs?[​](#timeoutms "Direct link to timeoutMs?")
186
208
 
187
209
  ```ts
@@ -189,7 +211,7 @@ optional timeoutMs: number;
189
211
 
190
212
  ```
191
213
 
192
- Per-turn wall-clock timeout (ms) before a turn is failed. Defaults to 120s.
214
+ Default per-eval timeout (ms): `runEval` races the whole test against it and it also caps each driver turn. A per-eval `def.timeoutMs` overrides it, and it wins over an agent's `evals.config.ts` `timeoutMs`. Unbounded when unset.
193
215
 
194
216
  ***
195
217
 
@@ -152,6 +152,28 @@ Assert a tool was called during the run (gate by default).
152
152
 
153
153
  ***
154
154
 
155
+ ### calledToolWith()[​](#calledtoolwith "Direct link to calledToolWith()")
156
+
157
+ ```ts
158
+ calledToolWith(name: string, expected: Record<string, unknown>): AssertionHandle;
159
+
160
+ ```
161
+
162
+ Assert a tool was called with arguments that deep-contain `expected`: every key in `expected` must equal the actual argument (recursively for nested objects; arrays match element-for-element), so extra arguments are ignored. Gate by default.
163
+
164
+ #### Parameters[​](#parameters-4 "Direct link to Parameters")
165
+
166
+ | Parameter | Type |
167
+ | ---------- | ----------------------------- |
168
+ | `name` | `string` |
169
+ | `expected` | `Record`<`string`, `unknown`> |
170
+
171
+ #### Returns[​](#returns-4 "Direct link to Returns")
172
+
173
+ [`AssertionHandle`](./docs/api/appkit/Interface.AssertionHandle.md)
174
+
175
+ ***
176
+
155
177
  ### check()[​](#check "Direct link to check()")
156
178
 
157
179
  ```ts
@@ -161,14 +183,14 @@ check(value: string, matcher: Matcher): AssertionHandle;
161
183
 
162
184
  Assert a value against a matcher, e.g. `t.check(t.reply, includes("Sunny"))`.
163
185
 
164
- #### Parameters[​](#parameters-4 "Direct link to Parameters")
186
+ #### Parameters[​](#parameters-5 "Direct link to Parameters")
165
187
 
166
188
  | Parameter | Type |
167
189
  | --------- | --------------------------------------------------------- |
168
190
  | `value` | `string` |
169
191
  | `matcher` | [`Matcher`](./docs/api/appkit/TypeAlias.Matcher.md) |
170
192
 
171
- #### Returns[​](#returns-4 "Direct link to Returns")
193
+ #### Returns[​](#returns-5 "Direct link to Returns")
172
194
 
173
195
  [`AssertionHandle`](./docs/api/appkit/Interface.AssertionHandle.md)
174
196
 
@@ -183,7 +205,7 @@ reset(): void;
183
205
 
184
206
  Start a fresh conversation: the next `send` opens a new thread with no history. Use to run several independent one-shot checks in one test. Consecutive `send`s (without a `reset`) stay in one multi-turn conversation.
185
207
 
186
- #### Returns[​](#returns-5 "Direct link to Returns")
208
+ #### Returns[​](#returns-6 "Direct link to Returns")
187
209
 
188
210
  `void`
189
211
 
@@ -198,13 +220,13 @@ send(message: string): Promise<void>;
198
220
 
199
221
  Send a user message to the agent and capture its response.
200
222
 
201
- #### Parameters[​](#parameters-5 "Direct link to Parameters")
223
+ #### Parameters[​](#parameters-6 "Direct link to Parameters")
202
224
 
203
225
  | Parameter | Type |
204
226
  | --------- | -------- |
205
227
  | `message` | `string` |
206
228
 
207
- #### Returns[​](#returns-6 "Direct link to Returns")
229
+ #### Returns[​](#returns-7 "Direct link to Returns")
208
230
 
209
231
  `Promise`<`void`>
210
232
 
@@ -219,13 +241,13 @@ skip(reason?: string): never;
219
241
 
220
242
  Skip this eval with an optional reason.
221
243
 
222
- #### Parameters[​](#parameters-6 "Direct link to Parameters")
244
+ #### Parameters[​](#parameters-7 "Direct link to Parameters")
223
245
 
224
246
  | Parameter | Type |
225
247
  | --------- | -------- |
226
248
  | `reason?` | `string` |
227
249
 
228
- #### Returns[​](#returns-7 "Direct link to Returns")
250
+ #### Returns[​](#returns-8 "Direct link to Returns")
229
251
 
230
252
  `never`
231
253
 
@@ -240,6 +262,6 @@ succeeded(): AssertionHandle;
240
262
 
241
263
  Assert the last turn completed successfully (gate by default).
242
264
 
243
- #### Returns[​](#returns-8 "Direct link to Returns")
265
+ #### Returns[​](#returns-9 "Direct link to Returns")
244
266
 
245
267
  [`AssertionHandle`](./docs/api/appkit/Interface.AssertionHandle.md)
@@ -54,6 +54,7 @@ Documentation merge entry for Typedoc — combines the stable `@databricks/appki
54
54
  | [DatabricksAuth](./docs/api/appkit/Interface.DatabricksAuth.md) | Resolved Databricks host + bearer token for the eval runner's REST calls. |
55
55
  | [DatasetRow](./docs/api/appkit/Interface.DatasetRow.md) | One row of a managed evaluation dataset. `inputs` are the kwargs passed to the agent for the turn; `expectations` (when present) is the row's ground truth / guidelines. Mirrors the `{inputs, expectations}` shape of `mlflow.genai` datasets and of the Unity Catalog table backing a managed eval dataset. |
56
56
  | [DiscoveredEval](./docs/api/appkit/Interface.DiscoveredEval.md) | An eval file found under `server/agents/<agent>/evals/`. |
57
+ | [DiscoveredEvalConfig](./docs/api/appkit/Interface.DiscoveredEvalConfig.md) | A per-agent `evals.config.ts` found under `server/agents/<agent>/evals/`. |
57
58
  | [DriveResult](./docs/api/appkit/Interface.DriveResult.md) | What a driver returns for a single `t.send`. |
58
59
  | [EndpointConfig](./docs/api/appkit/Interface.EndpointConfig.md) | - |
59
60
  | [EntityMutationHooks](./docs/api/appkit/Interface.EntityMutationHooks.md) | Mutation lifecycle for one entity. A before hook may return a replacement payload, which is revalidated against the trusted schema before it is persisted. Every hook, the mutation, and any write a hook issues through `ctx.app.database` share one transaction, so a rejection anywhere rolls all of them back. Throw `DatabaseValidationError` to answer a generated route with `422`; any other failure stays an opaque server error. |
@@ -198,9 +199,11 @@ Documentation merge entry for Typedoc — combines the stable `@databricks/appki
198
199
  | [createWorkspaceClient](./docs/api/appkit/Function.createWorkspaceClient.md) | Construct an AppKit workspace client. |
199
200
  | [database](./docs/api/appkit/Function.database.md) | Create a typed database plugin registration for a finalized schema. |
200
201
  | [defineEval](./docs/api/appkit/Function.defineEval.md) | Define an agent eval. Default-export the result from a `server/agents/<id>/evals/*.eval.ts` file. |
202
+ | [defineEvalConfig](./docs/api/appkit/Function.defineEvalConfig.md) | Define per-directory eval config. Default-export from `evals.config.ts`. |
201
203
  | [defineManifest](./docs/api/appkit/Function.defineManifest.md) | Validates a raw manifest (typically a `manifest.json` import) against the canonical Zod schema and returns it as a strict [PluginManifest](./docs/api/appkit/Interface.PluginManifest.md). |
202
204
  | [defineSchema](./docs/api/appkit/Function.defineSchema.md) | Compile one declared schema. The returned type keeps the table names the builder returned, so `api.tables` and `hooks` can name only real tables. |
203
205
  | [defineTool](./docs/api/appkit/Function.defineTool.md) | Defines a single tool entry for a plugin's internal registry. |
206
+ | [discoverEvalConfigs](./docs/api/appkit/Function.discoverEvalConfigs.md) | Discover the per-agent `evals.config.ts` (from [defineEvalConfig](./docs/api/appkit/Function.defineEvalConfig.md)) at `<rootDir>/server/agents/<agent>/evals/evals.config.ts`. Config is per-agent: each agent's config applies only to that agent's evals. Agents without a config file are omitted. Returns a stable, sorted list. |
204
207
  | [discoverEvalFiles](./docs/api/appkit/Function.discoverEvalFiles.md) | Discover evals under `<rootDir>/server/agents/<agent>/evals/` — co-located with each agent's `agent.{md,ts}` (same folder-per-agent layout the agents plugin discovers). The agent id is the folder name; the eval id is the file path relative to that evals dir with `.eval.ts` stripped. Sorted + stable. |
205
208
  | [enumColumn](./docs/api/appkit/Function.enumColumn.md) | - |
206
209
  | [equals](./docs/api/appkit/Function.equals.md) | Passes when the value equals `expected` exactly. |
@@ -212,6 +215,8 @@ Documentation merge entry for Typedoc — combines the stable `@databricks/appki
212
215
  | [formatEvalDetail](./docs/api/appkit/Function.formatEvalDetail.md) | Indented detail lines for a failing eval (error + failing assertions). |
213
216
  | [formatEvalHeadline](./docs/api/appkit/Function.formatEvalHeadline.md) | The one-line header for a single eval result (no failure detail). |
214
217
  | [formatEvalResults](./docs/api/appkit/Function.formatEvalResults.md) | Render all results as a human-readable console report (non-streaming). |
218
+ | [formatResultsJson](./docs/api/appkit/Function.formatResultsJson.md) | Render results as a machine-readable JSON report (2-space indented): `{ summary: EvalSummary, results: EvalResult[] }`. Faithful to the types — every field present on a result round-trips. |
219
+ | [formatResultsJUnit](./docs/api/appkit/Function.formatResultsJUnit.md) | Render results as JUnit XML for standard CI test reporters: a single `<testsuite name="appkit-agent-evals">` with one `<testcase>` per result. Failures carry a `<failure>` (error or failing-gate summary); skips a `<skipped>`. All attribute/text values are XML-escaped. |
215
220
  | [formatSummaryLine](./docs/api/appkit/Function.formatSummaryLine.md) | The final PASS/FAIL summary line. |
216
221
  | [fromSupervisorApi](./docs/api/appkit/Function.fromSupervisorApi.md) | Creates an [AgentAdapter](./docs/api/appkit/Interface.AgentAdapter.md) backed by the Databricks AI Gateway Responses API (`/ai-gateway/mlflow/v1/responses`). |
217
222
  | [functionToolToDefinition](./docs/api/appkit/Function.functionToolToDefinition.md) | - |
@@ -247,10 +252,12 @@ Documentation merge entry for Typedoc — combines the stable `@databricks/appki
247
252
  | [runAgent](./docs/api/appkit/Function.runAgent.md) | Standalone agent execution without `createApp`. Resolves the adapter, binds inline tools, and drives the adapter's `run()` loop to completion. |
248
253
  | [runEval](./docs/api/appkit/Function.runEval.md) | Run a single eval against a driver. Never throws for assertion or agent failures — those become a non-passing [EvalResult](./docs/api/appkit/Interface.EvalResult.md). Only a malformed eval definition surfaces as `result.error`. |
249
254
  | [runEvalsInDir](./docs/api/appkit/Function.runEvalsInDir.md) | Discover, load, and run every eval under each agent's `evals/` dir, driving the agents on a running app. Never throws for an individual eval — load/run failures become non-passing [EvalResult](./docs/api/appkit/Interface.EvalResult.md)s. |
255
+ | [runWithRetries](./docs/api/appkit/Function.runWithRetries.md) | Run `attempt` up to `1 + retries` times, stopping as soon as it returns a result that is neither a thrown error / per-eval timeout (`error`) nor a transport/agent turn failure (`infraFailure`). Assertion failures set neither, so a failed-but-completed eval is returned on the first try and never retried. Returns the last result when every attempt failed on infra. |
250
256
  | [summarize](./docs/api/appkit/Function.summarize.md) | - |
251
257
  | [text](./docs/api/appkit/Function.text.md) | - |
252
258
  | [timestamp](./docs/api/appkit/Function.timestamp.md) | - |
253
259
  | [tool](./docs/api/appkit/Function.tool.md) | Factory for defining function tools with Zod schemas. |
254
260
  | [toolsFromRegistry](./docs/api/appkit/Function.toolsFromRegistry.md) | Produces the `AgentToolDefinition[]` a ToolProvider exposes to the LLM, deriving `parameters` JSON Schema from each entry's Zod schema. |
261
+ | [userTurns](./docs/api/appkit/Function.userTurns.md) | Extract every user-message content, in order, from an MLflow `{"messages":[{"role":"user","content":"..."}]}` input. A dataset row can carry a full multi-turn conversation; replaying these against one thread (one `t.send` per returned string) lets the agent see the accumulating history. |
255
262
  | [uuid](./docs/api/appkit/Function.uuid.md) | - |
256
263
  | [varchar](./docs/api/appkit/Function.varchar.md) | - |
@@ -73,7 +73,7 @@ Migrating from `config/agents/`
73
73
 
74
74
  Earlier versions kept markdown agents under `config/agents/<id>/agent.md`. That location is still read as a deprecated fallback (one-time warning on boot); move each folder to `server/agents/<id>/agent.md` so every agent — markdown and code — lives in one place.
75
75
 
76
- Requests land at `POST /invocations` (or its alias `POST /responses`) with an OpenAI Responses-compatible body. These endpoints run the agent to completion and return a single JSON response — no SSE. Streaming clients should use `POST /chat`. Every tool call runs through `asUser(req)` so SQL executes as the requesting user, file access respects Unity Catalog ACLs, and telemetry spans are created automatically.
76
+ Requests land at `POST /invocations` (or its alias `POST /responses`) with an OpenAI Responses-compatible body. These endpoints run the agent to completion and return a single JSON response — no SSE. Streaming clients should use `POST /chat`. Every tool call is traced automatically. Plugin-toolkit tool calls (the `plugin:<name>` entries / `plugins.<name>.toolkit()`) additionally run through `asUser(req)`, so their SQL executes as the requesting user and file access respects Unity Catalog ACLs. A hand-rolled `tool({ execute })` is **not** wrapped: its `execute` receives only the validated tool arguments (no `req`), so it runs with the app's service-principal identity and cannot opt into OBO. If a tool must act as the requesting user, expose it as a plugin tool rather than a hand-rolled `execute`. See [Execution context](./docs/plugins/execution-context.md).
77
77
 
78
78
  No HITL on `/invocations` and `/responses`
79
79
 
@@ -218,7 +218,7 @@ Skills are on-demand instruction packs — the same `SKILL.md` format Claude Cod
218
218
  A skill is a directory with a `SKILL.md` plus any bundled reference files:
219
219
 
220
220
  ```text
221
- config/agents/
221
+ server/agents/
222
222
  skills/ # shared pool — any agent can opt in
223
223
  pdf-forms/
224
224
  SKILL.md
@@ -248,8 +248,8 @@ To fill a PDF form:
248
248
 
249
249
  ### Visibility[​](#visibility "Direct link to Visibility")
250
250
 
251
- * **Per-agent skills** (`config/agents/<id>/skills/`) are always visible to that agent.
252
- * **Global skills** (`config/agents/skills/`, and catalog-volume skills) are **opt-in**: list them in the agent's frontmatter, `skills: [pdf-forms]`. Set `autoInheritSkills: true` (or `{ file, code }`) on the plugin to make every global skill visible without listing — off by default so each agent's always-on catalog stays lean.
251
+ * **Per-agent skills** (`server/agents/<id>/skills/`) are always visible to that agent.
252
+ * **Global skills** (`server/agents/skills/`, and catalog-volume skills) are **opt-in**: list them in the agent's frontmatter, `skills: [pdf-forms]`. Set `autoInheritSkills: true` (or `{ file, code }`) on the plugin to make every global skill visible without listing — off by default so each agent's always-on catalog stays lean.
253
253
 
254
254
  ### How the agent uses a skill[​](#how-the-agent-uses-a-skill "Direct link to How the agent uses a skill")
255
255
 
@@ -687,6 +687,257 @@ appkit.agents.getThreads(userId); // list user's threads
687
687
 
688
688
  ```
689
689
 
690
+ ## Evaluating agents[​](#evaluating-agents "Direct link to Evaluating agents")
691
+
692
+ AppKit ships an eval framework for the agents you build here. You author evals in TypeScript with `defineEval`, drive the agent by sending it messages, and assert on its reply and tool usage with deterministic matchers or LLM judges. Evals run against a **running app** over HTTP (`--url`), and — with Databricks creds and an experiment — report to MLflow as native "Evaluation runs" with per-assertion and per-judge feedback attached to each turn's trace. The eval API is part of the beta surface: import it from `@databricks/appkit/beta`.
693
+
694
+ Evals live beside each agent: `server/agents/<agent-id>/evals/*.eval.ts`. Each file default-exports one `defineEval({ test })`. The agent under test defaults to the parent `<agent-id>` directory; set `agent:` to target a different one.
695
+
696
+ ### A first eval[​](#a-first-eval "Direct link to A first eval")
697
+
698
+ ```ts
699
+ // server/agents/query/evals/smoke.eval.ts
700
+ import { defineEval } from "@databricks/appkit/beta";
701
+
702
+ export default defineEval({
703
+ description: "Query agent responds to a greeting",
704
+ async test(t) {
705
+ await t.send("Hi there!");
706
+ t.succeeded(); // gate: the turn completed without an agent/stream error
707
+ },
708
+ });
709
+
710
+ ```
711
+
712
+ Start the app, then run the evals against it:
713
+
714
+ ```bash
715
+ # the app must be running and reachable at --url
716
+ appkit agent eval --url http://localhost:3000
717
+
718
+ # scope to one agent/eval by substring, and point at a project root
719
+ appkit agent eval query --root apps/dev-playground --url http://localhost:3000
720
+
721
+ ```
722
+
723
+ The positional `[filter]` matches evals whose `<agent>/<id>` contains the substring (or an exact agent id). The command discovers every `*.eval.ts` under `server/agents/*/evals/`, drives each against the running app, and exits non-zero if any gate fails.
724
+
725
+ ### Assertions[​](#assertions "Direct link to Assertions")
726
+
727
+ Every assertion returns a chainable handle. Assertions are **gates by default** — a failure fails the eval (non-zero exit). Chain `.soft()` to demote to a tracked-only metric, `.gate()` to promote a soft assertion back, or `.atLeast(n)` to set the pass threshold on a scored assertion.
728
+
729
+ | Assertion | Passes when |
730
+ | ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ |
731
+ | `t.succeeded()` | The last turn completed without an agent/stream error. |
732
+ | `t.calledTool(name)` | The agent called `name` during the run. |
733
+ | `t.calledToolWith(name, expected)` | `name` was called with arguments that deep-contain `expected` (every key in `expected` matches recursively; extra args are ignored). |
734
+ | `t.check(value, matcher)` | `value` satisfies the matcher — `includes(substring)`, `equals(expected)`, or `matches(pattern)`. |
735
+
736
+ ```ts
737
+ import { defineEval, includes } from "@databricks/appkit/beta";
738
+
739
+ export default defineEval({
740
+ description: "Helper agent answers a math question",
741
+ agent: "helper",
742
+ async test(t) {
743
+ await t.send("What is 2 + 2?");
744
+ t.succeeded(); // gate
745
+ t.check(t.reply, includes("4")).soft(); // tracked metric, won't fail the gate
746
+ },
747
+ });
748
+
749
+ ```
750
+
751
+ ```ts
752
+ // deep-partial tool-arg check
753
+ await t.send("What's the weather in Brooklyn?");
754
+ t.calledTool("get_weather");
755
+ t.calledToolWith("get_weather", { city: "Brooklyn" });
756
+
757
+ ```
758
+
759
+ Call `t.skip("reason")` to skip an eval, and read `t.reply`, `t.toolCalls`, and `t.sessionId` to inspect the last turn.
760
+
761
+ ### LLM-as-judge[​](#llm-as-judge "Direct link to LLM-as-judge")
762
+
763
+ `t.judge.*` scores the last reply with an LLM judge (via `autoevals` pointed at a Databricks serving endpoint). Each judge returns a scored assertion (0..1) that **gates by default** — a miss fails the eval. Chain `.atLeast(n)` to set the pass threshold, or `.soft()` to track it only. Judges require a judge model: pass `--judge-model <endpoint>` (or set `APPKIT_JUDGE_MODEL`) plus Databricks auth; without one, `t.judge.*` throws with a clear message.
764
+
765
+ ```ts
766
+ async test(t) {
767
+ await t.send("What's the weather in Brooklyn?");
768
+ t.succeeded();
769
+
770
+ // closedQA needs no ground truth — it judges the reply against a question.
771
+ (await t.judge.closedQA(
772
+ "Does the response describe weather conditions for Brooklyn?",
773
+ )).atLeast(0.5);
774
+ }
775
+
776
+ ```
777
+
778
+ * `t.judge.factuality(expected)` — score the reply against an expected reference answer.
779
+ * `t.judge.closedQA(criteria)` — score whether the reply answers the question, per `criteria`.
780
+ * `t.judge.custom(spec)` — a prompt-template judge (`{ name, promptTemplate, choiceScores }`), the TS analog of MLflow's `@scorer`.
781
+
782
+ Guard judge calls with `isJudgeConfigured()` when an eval should still exercise the drive path without a judge model configured:
783
+
784
+ ```ts
785
+ import { defineEval, isJudgeConfigured } from "@databricks/appkit/beta";
786
+ // ...
787
+ if (isJudgeConfigured()) {
788
+ (await t.judge.closedQA(guideline)).atLeast(0.5);
789
+ }
790
+
791
+ ```
792
+
793
+ ### Conversations[​](#conversations "Direct link to Conversations")
794
+
795
+ Each `t.send` is one user turn. How you sequence them controls the thread:
796
+
797
+ ```ts
798
+ // One-shot: a single turn.
799
+ await t.send("Summarize Q3 revenue.");
800
+ t.succeeded();
801
+
802
+ // Multi-turn: consecutive sends share one thread, so the agent sees history.
803
+ await t.send("Show me the orders table.");
804
+ await t.send("Now filter it to last week.");
805
+ t.succeeded();
806
+
807
+ // t.reset() drops the conversation: the next send opens a fresh thread with
808
+ // no history. Use it to run several independent one-shot checks in one test.
809
+ await t.send("What's 2 + 2?");
810
+ t.check(t.reply, includes("4"));
811
+ t.reset();
812
+ await t.send("What's the capital of France?");
813
+ t.check(t.reply, includes("Paris"));
814
+
815
+ ```
816
+
817
+ ### Datasets[​](#datasets "Direct link to Datasets")
818
+
819
+ Add `dataset: { table }` to sweep a Databricks **managed evaluation dataset** — a Unity Catalog `catalog.schema.table` with `inputs`/`expectations` columns. The eval runs once per row; the runner binds each row's `inputs` to `t.input` and `expectations` to `t.expected`. Reading the dataset requires a workspace client and warehouse (`--warehouse-id` + auth).
820
+
821
+ ```ts
822
+ import { defineEval, isJudgeConfigured, userTurns } from "@databricks/appkit/beta";
823
+
824
+ export default defineEval({
825
+ description: "Query agent satisfies each dataset row's guidelines",
826
+ dataset: { table: "main.mario.appkit_eval_dataset" }, // optional `limit?`
827
+ async test(t) {
828
+ // Replay every user turn in the row against one thread, so the agent sees
829
+ // the accumulating conversation. A single-user-turn row sends once.
830
+ for (const turn of userTurns(t.input)) {
831
+ await t.send(turn);
832
+ }
833
+ t.succeeded();
834
+
835
+ if (isJudgeConfigured()) {
836
+ for (const guideline of guidelines(t.expected)) {
837
+ (await t.judge.closedQA(guideline)).atLeast(0.5);
838
+ }
839
+ }
840
+ },
841
+ });
842
+
843
+ ```
844
+
845
+ The row shapes match the MLflow managed-dataset UI:
846
+
847
+ ```text
848
+ inputs {"messages":[{"role":"user","content":"..."}]}
849
+ expectations {"guidelines":{"value":["...","..."]}} (optional)
850
+
851
+ ```
852
+
853
+ `userTurns(t.input)` extracts every `role: "user"` message content in order from the `{messages:[...]}` input — a row can carry a full multi-turn conversation, and replaying each user turn against one thread lets the agent build up history (interleaved assistant/system turns are ignored; the agent generates its own). Read `expectations.guidelines.value` yourself; the UI wraps the array as `{value: [...]}`:
854
+
855
+ ```ts
856
+ function guidelines(expected: Record<string, unknown> | undefined): string[] {
857
+ const g = (expected?.guidelines as { value?: unknown } | undefined)?.value;
858
+ return Array.isArray(g) ? g.map(String) : [];
859
+ }
860
+
861
+ ```
862
+
863
+ Run a dataset eval:
864
+
865
+ ```bash
866
+ appkit agent eval dataset --root apps/dev-playground --url http://localhost:3000 \
867
+ --profile <profile> --warehouse-id <warehouse-id> --judge-model <endpoint>
868
+
869
+ ```
870
+
871
+ ### Running evals & CI[​](#running-evals--ci "Direct link to Running evals & CI")
872
+
873
+ `appkit agent eval [filter]` — run agent evals (`server/agents/<id>/evals/*.eval.ts`) against a running app.
874
+
875
+ | Flag | Description |
876
+ | ---------------------------- | --------------------------------------------------------------------------------------------------------------- |
877
+ | `[filter]` | Only run evals whose `<agent>/<id>` contains this substring (or an exact agent id) |
878
+ | `--url <url>` | Base URL of the running app (default `http://localhost:3000`) |
879
+ | `--strict` | Fail on soft-assertion misses too |
880
+ | `--root <dir>` | Project root containing `server/agents/` (default: cwd) |
881
+ | `--header <header...>` | Extra request header as `'Key: value'` (repeatable) |
882
+ | `--tag <tag...>` | Only run evals tagged with one of these tags (repeatable) |
883
+ | `--profile <name>` | Databricks CLI profile to authenticate with via OAuth (default: `DATABRICKS_CONFIG_PROFILE`) |
884
+ | `--databricks-host <host>` | Databricks host for writing MLflow assessments (default: `DATABRICKS_HOST`) |
885
+ | `--databricks-token <token>` | Databricks token for writing MLflow assessments (default: `DATABRICKS_TOKEN`) |
886
+ | `--experiment <id>` | MLflow experiment id for the evaluation run (default: `MLFLOW_EXPERIMENT_ID`) |
887
+ | `--warehouse-id <id>` | SQL warehouse id for reading managed evaluation datasets (default: `DATABRICKS_WAREHOUSE_ID`) |
888
+ | `--judge-model <endpoint>` | Databricks serving endpoint to use as the LLM judge for `t.judge.*` (default: `APPKIT_JUDGE_MODEL`) |
889
+ | `--concurrency <n>` | Max evals/dataset rows to drive concurrently (default: 4) |
890
+ | `--timeout <ms>` | Default per-eval timeout in ms (a per-eval `timeoutMs` overrides it) |
891
+ | `--retries <n>` | Re-run an eval up to N times when it fails on an infra error (turn/timeout); assertion failures are not retried |
892
+ | `--min-pass-rate <rate>` | Gate on aggregate pass rate (0..1) instead of requiring every eval to pass; exit 1 when below |
893
+ | `--reporter <format>` | Report format: `text` (live console), `json` (dashboards), or `junit` (CI test reporters) |
894
+ | `--output <file>` | Write the json/junit report to this file instead of stdout (ignored for text) |
895
+
896
+ Notes:
897
+
898
+ * **Auth is OAuth-first.** `--profile <name>` mints an OAuth token from your Databricks CLI profile — no PAT needed. An explicit `--databricks-host`/`--databricks-token` (or the `DATABRICKS_*` env vars) wins over the profile.
899
+ * **`--retries` only absorbs infra flakiness.** A retry fires only when an eval throws or times out (`result.error` set); a wrong reply is real signal and is never retried. Each attempt gets a fresh driver.
900
+ * **Gating.** By default the run exits non-zero if any eval fails. `--min-pass-rate 0.9` switches to threshold mode: it exits non-zero only when the aggregate pass rate drops below the threshold. `--strict` additionally treats soft-assertion misses as failures.
901
+ * **CI reports.** `--reporter junit --output results.xml` writes a JUnit file for CI test reporters; `--reporter json` emits machine-readable results. In machine reporters, human-facing lines go to stderr so stdout stays clean for the report.
902
+
903
+ ```bash
904
+ # CI: gate on 90% pass rate, emit JUnit, authenticate + report to MLflow
905
+ appkit agent eval --url "$APP_URL" \
906
+ --profile ci \
907
+ --experiment "$MLFLOW_EXPERIMENT_ID" \
908
+ --concurrency 4 --retries 1 \
909
+ --min-pass-rate 0.9 \
910
+ --reporter junit --output eval-results.xml
911
+
912
+ ```
913
+
914
+ ### Per-directory config[​](#per-directory-config "Direct link to Per-directory config")
915
+
916
+ Drop an `evals.config.ts` beside an agent's evals to set defaults for that agent's runs:
917
+
918
+ ```ts
919
+ // server/agents/query/evals/evals.config.ts
920
+ import { defineEvalConfig } from "@databricks/appkit/beta";
921
+
922
+ export default defineEvalConfig({
923
+ maxConcurrency: 4, // run up to 4 evals/rows concurrently
924
+ timeoutMs: 30_000, // default per-eval timeout
925
+ });
926
+
927
+ ```
928
+
929
+ Precedence: a **CLI flag** wins over the **`evals.config.ts`** value, which wins over the **built-in default** (concurrency `4`, no timeout). A per-eval `def.timeoutMs` overrides both for that eval.
930
+
931
+ ### MLflow reporting[​](#mlflow-reporting "Direct link to MLflow reporting")
932
+
933
+ When `--experiment <id>` (or `MLFLOW_EXPERIMENT_ID`) is set together with Databricks auth, the runner creates a native MLflow **Evaluation run** up front. As each eval runs against the app, its turn trace is linked to the run, and every assertion and judge score is written back as feedback on that trace:
934
+
935
+ * Each deterministic assertion becomes a `CODE`-sourced Feedback with a boolean value.
936
+ * Each judge assertion becomes an `LLM_JUDGE`-sourced Feedback with its numeric 0..1 score and rationale.
937
+ * An overall `appkit_eval` pass/fail Feedback is attached per eval, and aggregate metrics are logged when the run finishes.
938
+
939
+ Skip `--experiment` (and `MLFLOW_EXPERIMENT_ID`) to run evals purely locally with no MLflow side effects — the CLI prints a reminder that the evaluation run was skipped.
940
+
690
941
  ## Frontmatter schema[​](#frontmatter-schema "Direct link to Frontmatter schema")
691
942
 
692
943
  | Key | Type | Notes |