@databricks/appkit-ui 0.72.0 → 0.73.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (36) hide show
  1. package/CLAUDE.md +17 -0
  2. package/NOTICE.md +1 -0
  3. package/dist/cli/commands/agent/eval.js +9 -1
  4. package/dist/cli/commands/agent/eval.js.map +1 -1
  5. package/dist/react/ui/button.d.ts +1 -1
  6. package/docs/api/appkit/Class.AppKitError.md +1 -0
  7. package/docs/api/appkit/Class.DatabaseValidationError.md +191 -0
  8. package/docs/api/appkit/Function.defineSchema.md +1 -1
  9. package/docs/api/appkit/Function.readEvalDataset.md +21 -0
  10. package/docs/api/appkit/Function.resolveWorkspaceClient.md +18 -0
  11. package/docs/api/appkit/Interface.AssertionHandle.md +1 -1
  12. package/docs/api/appkit/Interface.DatabaseValidationIssue.md +21 -0
  13. package/docs/api/appkit/Interface.DatasetRow.md +21 -0
  14. package/docs/api/appkit/Interface.EntityMutationHooks.md +173 -0
  15. package/docs/api/appkit/Interface.EvalDefinition.md +28 -0
  16. package/docs/api/appkit/Interface.EvalDriver.md +16 -1
  17. package/docs/api/appkit/Interface.HookApp.md +12 -0
  18. package/docs/api/appkit/Interface.HookContext.md +21 -0
  19. package/docs/api/appkit/Interface.ReadEvalDatasetOptions.md +34 -0
  20. package/docs/api/appkit/Interface.ReadSerializerContext.md +21 -0
  21. package/docs/api/appkit/Interface.RunEvalOptions.md +11 -0
  22. package/docs/api/appkit/Interface.RunEvalsOptions.md +22 -0
  23. package/docs/api/appkit/Interface.TestContext.md +41 -4
  24. package/docs/api/appkit/TypeAlias.DatabaseApiConfig.md +53 -0
  25. package/docs/api/appkit/TypeAlias.DatabaseApiWriteOperation.md +8 -0
  26. package/docs/api/appkit/TypeAlias.DatabaseApiWritesConfig.md +49 -0
  27. package/docs/api/appkit/TypeAlias.DatabaseExports.md +3 -3
  28. package/docs/api/appkit/TypeAlias.EntityHooks.md +25 -0
  29. package/docs/api/appkit/TypeAlias.IDatabaseConfig.md +16 -5
  30. package/docs/api/appkit/TypeAlias.ReadSerializer.md +19 -0
  31. package/docs/api/appkit/TypeAlias.TransactionClient.md +19 -0
  32. package/docs/api/appkit.md +135 -119
  33. package/docs/plugins/database.md +144 -0
  34. package/llms.txt +17 -0
  35. package/package.json +1 -1
  36. package/sbom.cdx.json +1 -1
package/CLAUDE.md CHANGED
@@ -48,6 +48,7 @@ npx @databricks/appkit docs <query>
48
48
  - [Analytics plugin](./docs/plugins/analytics.md): Enables SQL query execution against Databricks SQL Warehouses.
49
49
  - [Caching](./docs/plugins/caching.md): AppKit provides both global and plugin-level caching capabilities.
50
50
  - [Creating custom plugins](./docs/plugins/custom-plugins.md): If you need custom API routes or background logic, implement an AppKit plugin. The fastest way is to use the CLI:
51
+ - [Database plugin](./docs/plugins/database.md): This plugin is currently beta. APIs may change between minor releases. Import from @databricks/appkit/beta. See Plugin Stability Tiers.
51
52
  - [Execution context](./docs/plugins/execution-context.md): AppKit manages Databricks authentication via two contexts:
52
53
  - [Files plugin](./docs/plugins/files.md): File operations against Databricks Unity Catalog Volumes. Supports listing, reading, downloading, uploading, deleting, and previewing files with built-in caching, retry, and timeout handling via the execution interceptor pipeline.
53
54
  - [Genie plugin](./docs/plugins/genie.md): Integrates Databricks AI/BI Genie spaces into your AppKit application, enabling natural language data queries via a conversational interface.
@@ -68,6 +69,7 @@ npx @databricks/appkit docs <query>
68
69
  - [Class: AuthenticationError](./docs/api/appkit/Class.AuthenticationError.md): Error thrown when authentication fails.
69
70
  - [Class: ConfigurationError](./docs/api/appkit/Class.ConfigurationError.md): Error thrown when configuration is missing or invalid.
70
71
  - [Class: ConnectionError](./docs/api/appkit/Class.ConnectionError.md): Error thrown when a connection or network operation fails.
72
+ - [Class: DatabaseValidationError](./docs/api/appkit/Class.DatabaseValidationError.md): Deliberate validation failure raised by a database mutation hook. Generated
71
73
  - [Class: DatabricksAdapter](./docs/api/appkit/Class.DatabricksAdapter.md): Adapter that talks directly to Databricks Model Serving /invocations endpoint.
72
74
  - [Class: ExecutionError](./docs/api/appkit/Class.ExecutionError.md): Error thrown when an operation execution fails.
73
75
  - [Class: InitializationError](./docs/api/appkit/Class.InitializationError.md): Error thrown when a service or component is not properly initialized.
@@ -138,9 +140,11 @@ npx @databricks/appkit docs <query>
138
140
  - [Function: mcpServer()](./docs/api/appkit/Function.mcpServer.md): Factory for declaring a custom MCP server tool.
139
141
  - [Function: normalizeHost()](./docs/api/appkit/Function.normalizeHost.md): Ensure the host has a scheme (Databricks env often lacks https://).
140
142
  - [Function: parseTextToolCalls()](./docs/api/appkit/Function.parseTextToolCalls.md): Parses text-based tool calls from model output.
143
+ - [Function: readEvalDataset()](./docs/api/appkit/Function.readEvalDataset.md): Read a Databricks managed evaluation dataset (a Unity Catalog table with
141
144
  - [Function: reportToMlflow()](./docs/api/appkit/Function.reportToMlflow.md): Write one pass/fail assessment per eval result to the Databricks MLflow REST
142
145
  - [Function: resolveDatabricksAuth()](./docs/api/appkit/Function.resolveDatabricksAuth.md): Parameters
143
146
  - [Function: resolveHostedTools()](./docs/api/appkit/Function.resolveHostedTools.md): Parameters
147
+ - [Function: resolveWorkspaceClient()](./docs/api/appkit/Function.resolveWorkspaceClient.md): Construct a Databricks WorkspaceClient for the eval runner — the object the
144
148
  - [Function: runAgent()](./docs/api/appkit/Function.runAgent.md): Standalone agent execution without createApp. Resolves the adapter, binds
145
149
  - [Function: runEval()](./docs/api/appkit/Function.runEval.md): Run a single eval against a driver. Never throws for assertion or agent
146
150
  - [Function: runEvalsInDir()](./docs/api/appkit/Function.runEvalsInDir.md): Discover, load, and run every eval under each agent's evals/ dir, driving
@@ -166,10 +170,13 @@ npx @databricks/appkit docs <query>
166
170
  - [Interface: CustomJudgeSpec](./docs/api/appkit/Interface.CustomJudgeSpec.md): A custom LLM-judge definition: a prompt template and choice→score mapping.
167
171
  - [Interface: DatabaseCredential](./docs/api/appkit/Interface.DatabaseCredential.md): Database credentials with OAuth token for Postgres connection
168
172
  - [Interface: DatabaseRegistry](./docs/api/appkit/Interface.DatabaseRegistry.md): CANONICAL augmentation target. Empty by default; the generated database.d.ts
173
+ - [Interface: DatabaseValidationIssue](./docs/api/appkit/Interface.DatabaseValidationIssue.md): One rejected field; path names public columns, never their values.
169
174
  - [Interface: DatabricksAuth](./docs/api/appkit/Interface.DatabricksAuth.md): Resolved Databricks host + bearer token for the eval runner's REST calls.
175
+ - [Interface: DatasetRow](./docs/api/appkit/Interface.DatasetRow.md): One row of a managed evaluation dataset. inputs are the kwargs passed to the
170
176
  - [Interface: DiscoveredEval](./docs/api/appkit/Interface.DiscoveredEval.md): An eval file found under server/agents//evals/.
171
177
  - [Interface: DriveResult](./docs/api/appkit/Interface.DriveResult.md): What a driver returns for a single t.send.
172
178
  - [Interface: EndpointConfig](./docs/api/appkit/Interface.EndpointConfig.md): Properties
179
+ - [Interface: EntityMutationHooks<TTable>](./docs/api/appkit/Interface.EntityMutationHooks.md): Mutation lifecycle for one entity. A before hook may return a replacement
173
180
  - [Interface: EvalDefinition](./docs/api/appkit/Interface.EvalDefinition.md): A single eval, default-exported from a *.eval.ts file.
174
181
  - [Interface: EvalDriver](./docs/api/appkit/Interface.EvalDriver.md): Abstraction over how the agent is driven. The HTTP driver posts to a running
175
182
  - [Interface: EvalResult](./docs/api/appkit/Interface.EvalResult.md): The outcome of running one eval.
@@ -180,6 +187,8 @@ npx @databricks/appkit docs <query>
180
187
  - [Interface: FunctionTool](./docs/api/appkit/Interface.FunctionTool.md): Properties
181
188
  - [Interface: GenerateDatabaseCredentialRequest](./docs/api/appkit/Interface.GenerateDatabaseCredentialRequest.md): Request parameters for generating database OAuth credentials
182
189
  - [Interface: GenerationParams](./docs/api/appkit/Interface.GenerationParams.md): Optional generation parameters forwarded to the OpenAI-compatible serving
190
+ - [Interface: HookApp](./docs/api/appkit/Interface.HookApp.md): The only capability a hook receives: entities bound to its transaction.
191
+ - [Interface: HookContext](./docs/api/appkit/Interface.HookContext.md): Which entity is being mutated, and the surface a hook may write through.
183
192
  - [Interface: HostedSupervisorTool](./docs/api/appkit/Interface.HostedSupervisorTool.md): Tagged record returned by every supervisorTools factory. The
184
193
  - [Interface: HttpDriverOptions](./docs/api/appkit/Interface.HttpDriverOptions.md): Properties
185
194
  - [Interface: IAiSearchConfig](./docs/api/appkit/Interface.IAiSearchConfig.md): Base configuration interface for AppKit plugins
@@ -201,6 +210,8 @@ npx @databricks/appkit docs <query>
201
210
  - [Interface: PluginToolkitProvider](./docs/api/appkit/Interface.PluginToolkitProvider.md): Minimum shape every entry in the Plugins map must expose. Core
202
211
  - [Interface: PostResult](./docs/api/appkit/Interface.PostResult.md): Structured result for a best-effort POST that must not throw.
203
212
  - [Interface: PromptContext](./docs/api/appkit/Interface.PromptContext.md): Context passed to baseSystemPrompt callbacks.
213
+ - [Interface: ReadEvalDatasetOptions](./docs/api/appkit/Interface.ReadEvalDatasetOptions.md): Properties
214
+ - [Interface: ReadSerializerContext](./docs/api/appkit/Interface.ReadSerializerContext.md): Which entity and generated operation produced the row being shaped.
204
215
  - [Interface: RegisteredAgent](./docs/api/appkit/Interface.RegisteredAgent.md): Properties
205
216
  - [Interface: ReportOutcome](./docs/api/appkit/Interface.ReportOutcome.md): Properties
206
217
  - [Interface: RequestedClaims](./docs/api/appkit/Interface.RequestedClaims.md): Optional claims for fine-grained Unity Catalog table permissions
@@ -242,7 +253,11 @@ npx @databricks/appkit docs <query>
242
253
  - [Type Alias: AgentToolsFn()](./docs/api/appkit/TypeAlias.AgentToolsFn.md): Function form of AgentDefinition.tools. Receives the typed
243
254
  - [Type Alias: BaseSystemPromptOption](./docs/api/appkit/TypeAlias.BaseSystemPromptOption.md)
244
255
  - [Type Alias: ConfigSchema](./docs/api/appkit/TypeAlias.ConfigSchema.md): Configuration schema definition for plugin config.
256
+ - [Type Alias: DatabaseApiConfig<TSchema>](./docs/api/appkit/TypeAlias.DatabaseApiConfig.md): Full generated CRUD for every declared table by default. Set false to disable
257
+ - [Type Alias: DatabaseApiWriteOperation](./docs/api/appkit/TypeAlias.DatabaseApiWriteOperation.md): Generated HTTP write operations.
258
+ - [Type Alias: DatabaseApiWritesConfig<TSchema>](./docs/api/appkit/TypeAlias.DatabaseApiWritesConfig.md): All writes by default; false keeps reads only, and an object narrows writes.
245
259
  - [Type Alias: DatabaseExports](./docs/api/appkit/TypeAlias.DatabaseExports.md): Typed database API published by the plugin.
260
+ - [Type Alias: EntityHooks<TTable>](./docs/api/appkit/TypeAlias.EntityHooks.md): Response shaping and mutation lifecycle declared for one table.
246
261
  - [Type Alias: EvalProgress](./docs/api/appkit/TypeAlias.EvalProgress.md)
247
262
  - [Type Alias: ExecutionResult<T>](./docs/api/appkit/TypeAlias.ExecutionResult.md): Discriminated union for plugin execution results.
248
263
  - [Type Alias: FileAction](./docs/api/appkit/TypeAlias.FileAction.md): Every action the files plugin can perform.
@@ -254,6 +269,7 @@ npx @databricks/appkit docs <query>
254
269
  - [Type Alias: Matcher()](./docs/api/appkit/TypeAlias.Matcher.md): A deterministic matcher: inspects a string value and returns a result.
255
270
  - [Type Alias: PluginData<T, U, N>](./docs/api/appkit/TypeAlias.PluginData.md): Tuple of plugin class, config, and name. Created by toPlugin() and passed to createApp().
256
271
  - [Type Alias: Plugins](./docs/api/appkit/TypeAlias.Plugins.md): Plugin map passed to the function form of AgentDefinition.tools.
272
+ - [Type Alias: ReadSerializer()](./docs/api/appkit/TypeAlias.ReadSerializer.md): Shape one already private-safe row before it reaches the wire. A Promise
257
273
  - [Type Alias: ResolvedToolEntry](./docs/api/appkit/TypeAlias.ResolvedToolEntry.md): Internal tool-index entry after a tool record has been resolved to a dispatchable form.
258
274
  - [Type Alias: ResourceFieldEntry](./docs/api/appkit/TypeAlias.ResourceFieldEntry.md)
259
275
  - [Type Alias: ResourcePermission](./docs/api/appkit/TypeAlias.ResourcePermission.md): Union of all possible permission levels across all resource types.
@@ -263,6 +279,7 @@ npx @databricks/appkit docs <query>
263
279
  - [Type Alias: SupervisorTool](./docs/api/appkit/TypeAlias.SupervisorTool.md): Tools supported by the Databricks AI Gateway Responses API. The shapes match
264
280
  - [Type Alias: ToolRegistry](./docs/api/appkit/TypeAlias.ToolRegistry.md)
265
281
  - [Type Alias: ToPlugin()<T, U, N>](./docs/api/appkit/TypeAlias.ToPlugin.md): Factory function type returned by toPlugin(). Accepts optional config and returns a PluginData tuple.
282
+ - [Type Alias: TransactionClient](./docs/api/appkit/TypeAlias.TransactionClient.md): Entity and SQL capabilities bound to one transaction.
266
283
  - [Variable: agents](./docs/api/appkit/Variable.agents.md): Plugin factory for the agents plugin. Discovers agents from
267
284
  - [Variable: aiSearch](./docs/api/appkit/Variable.aiSearch.md)
268
285
  - [Variable: READ_ACTIONS](./docs/api/appkit/Variable.READ_ACTIONS.md): Actions that only read data.
package/NOTICE.md CHANGED
@@ -54,6 +54,7 @@ This Software contains code from the following open source projects:
54
54
  | [@tanstack/react-table](https://www.npmjs.com/package/@tanstack/react-table) | 8.21.3 | MIT | https://tanstack.com/table |
55
55
  | [@types/semver](https://www.npmjs.com/package/@types/semver) | 7.7.1 | MIT | https://github.com/DefinitelyTyped/DefinitelyTyped/tree/master/types/semver |
56
56
  | [apache-arrow](https://www.npmjs.com/package/apache-arrow) | 21.1.0 | Apache-2.0 | https://arrow.apache.org/js/ |
57
+ | [autoevals](https://www.npmjs.com/package/autoevals) | 0.3.0 | MIT | https://www.braintrust.dev/docs |
57
58
  | [class-variance-authority](https://www.npmjs.com/package/class-variance-authority) | 0.7.1 | Apache-2.0 | https://github.com/joe-bell/cva#readme |
58
59
  | [clsx](https://www.npmjs.com/package/clsx) | 2.1.1 | MIT | https://github.com/lukeed/clsx#readme |
59
60
  | [cmdk](https://www.npmjs.com/package/cmdk) | 1.1.1 | MIT | https://github.com/pacocoursey/cmdk#readme |
@@ -82,6 +82,12 @@ async function runAgentEval(filter, opts) {
82
82
  host: opts.databricksHost ?? process.env.DATABRICKS_HOST,
83
83
  token: opts.databricksToken ?? process.env.DATABRICKS_TOKEN
84
84
  }) ?? {};
85
+ const warehouseId = opts.warehouseId ?? process.env.DATABRICKS_WAREHOUSE_ID;
86
+ const workspaceClient = runner.resolveWorkspaceClient({
87
+ profile: opts.profile ?? process.env.DATABRICKS_CONFIG_PROFILE,
88
+ host: opts.databricksHost ?? process.env.DATABRICKS_HOST,
89
+ token: opts.databricksToken ?? process.env.DATABRICKS_TOKEN
90
+ });
85
91
  let summary;
86
92
  try {
87
93
  summary = await runner.runEvalsInDir({
@@ -93,6 +99,8 @@ async function runAgentEval(filter, opts) {
93
99
  concurrency: opts.concurrency,
94
100
  mlflow: resolveMlflow(opts, auth),
95
101
  judge: resolveJudge(opts, auth),
102
+ workspaceClient,
103
+ warehouseId,
96
104
  onEvent: makeProgressReporter(runner, opts.url)
97
105
  });
98
106
  } catch (err) {
@@ -105,7 +113,7 @@ async function runAgentEval(filter, opts) {
105
113
  else console.log("\nMLflow evaluation run skipped — pass --experiment (or set MLFLOW_EXPERIMENT_ID) plus --profile/--databricks-host to create one.");
106
114
  if (!runner.summarize(summary.results).allPassed) process.exitCode = 1;
107
115
  }
108
- const agentEvalCommand = new Command("eval").description("Run agent evals (server/agents/<id>/evals/*.eval.ts) against a running app").argument("[filter]", "Only run evals whose <agent>/<id> contains this substring (or an exact agent id)").option("--url <url>", "Base URL of the running app", "http://localhost:3000").option("--strict", "Fail on soft-assertion misses too", false).option("--concurrency <n>", "Max evals to run concurrently (default 4; keep at or below the app's max concurrent streams per user)", (v) => Number.parseInt(v, 10)).option("--root <dir>", "Project root containing server/agents/ (default: cwd)").option("--header <header...>", "Extra request header as 'Key: value' (repeatable)").option("--profile <name>", "Databricks CLI profile to authenticate with via OAuth (default: DATABRICKS_CONFIG_PROFILE)").option("--databricks-host <host>", "Databricks host for writing MLflow assessments (default: DATABRICKS_HOST)").option("--databricks-token <token>", "Databricks token for writing MLflow assessments (default: DATABRICKS_TOKEN)").option("--experiment <id>", "MLflow experiment id for the evaluation run (default: MLFLOW_EXPERIMENT_ID)").option("--warehouse-id <id>", "SQL warehouse id for writing assessments to UC-backed experiments (default: MLFLOW_TRACING_SQL_WAREHOUSE_ID or DATABRICKS_WAREHOUSE_ID)").option("--judge-model <endpoint>", "Databricks serving endpoint to use as the LLM judge for t.judge.* (default: APPKIT_JUDGE_MODEL)").action(runAgentEval);
116
+ const agentEvalCommand = new Command("eval").description("Run agent evals (server/agents/<id>/evals/*.eval.ts) against a running app").argument("[filter]", "Only run evals whose <agent>/<id> contains this substring (or an exact agent id)").option("--url <url>", "Base URL of the running app", "http://localhost:3000").option("--strict", "Fail on soft-assertion misses too", false).option("--concurrency <n>", "Max evals to run concurrently (default 4; keep at or below the app's max concurrent streams per user)", (v) => Number.parseInt(v, 10)).option("--root <dir>", "Project root containing server/agents/ (default: cwd)").option("--header <header...>", "Extra request header as 'Key: value' (repeatable)").option("--profile <name>", "Databricks CLI profile to authenticate with via OAuth (default: DATABRICKS_CONFIG_PROFILE)").option("--databricks-host <host>", "Databricks host for writing MLflow assessments (default: DATABRICKS_HOST)").option("--databricks-token <token>", "Databricks token for writing MLflow assessments (default: DATABRICKS_TOKEN)").option("--experiment <id>", "MLflow experiment id for the evaluation run (default: MLFLOW_EXPERIMENT_ID)").option("--warehouse-id <id>", "SQL warehouse id for reading managed eval datasets and writing assessments to UC-backed experiments (default: DATABRICKS_WAREHOUSE_ID, or MLFLOW_TRACING_SQL_WAREHOUSE_ID for assessments)").option("--judge-model <endpoint>", "Databricks serving endpoint to use as the LLM judge for t.judge.* (default: APPKIT_JUDGE_MODEL)").action(runAgentEval);
109
117
 
110
118
  //#endregion
111
119
  export { agentEvalCommand };
@@ -1 +1 @@
1
- {"version":3,"file":"eval.js","names":[],"sources":["../../../../src/cli/commands/agent/eval.ts"],"sourcesContent":["import { Command } from \"commander\";\n\ninterface EvalRunSummary {\n results: unknown[];\n mlflow?: {\n runId: string;\n report: {\n written: number;\n skipped: number;\n failures: Array<{ traceId: string; status?: number; error?: string }>;\n };\n finish: { finished: boolean; metricsError?: string; finishError?: string };\n };\n}\n\ntype EvalProgress =\n | { type: \"discovered\"; total: number }\n | { type: \"run-created\"; runId: string }\n | { type: \"start\"; id: string; index: number; total: number }\n | { type: \"result\"; result: unknown; index: number; total: number };\n\n/** Subset of `@databricks/appkit/beta`'s eval runner used by this command. */\ninterface EvalRunner {\n runEvalsInDir(opts: {\n rootDir?: string;\n baseUrl: string;\n filter?: string;\n strict?: boolean;\n headers?: Record<string, string>;\n concurrency?: number;\n mlflow?: {\n host: string;\n token: string;\n experimentId: string;\n sqlWarehouseId?: string;\n };\n judge?: { host: string; token: string; model: string };\n onEvent?: (event: EvalProgress) => void;\n }): Promise<EvalRunSummary>;\n resolveDatabricksAuth(opts: {\n profile?: string;\n host?: string;\n token?: string;\n }): Promise<{ host: string; token: string } | undefined>;\n formatEvalHeadline(result: unknown): string;\n formatEvalDetail(result: unknown): string[];\n formatSummaryLine(results: unknown[]): string;\n summarize(results: unknown[]): { allPassed: boolean };\n}\n\n/**\n * Loaded at runtime from the consuming project so this command (which ships in\n * `@databricks/shared`) doesn't take a build-time dependency on appkit. The\n * specifier is a variable so the type checker treats it as `any`.\n */\nasync function loadRunner(): Promise<EvalRunner> {\n const spec = \"@databricks/appkit/beta\";\n try {\n return (await import(spec)) as unknown as EvalRunner;\n } catch (err) {\n throw new Error(\n \"Could not load @databricks/appkit. Run `appkit agent eval` from a \" +\n \"project with @databricks/appkit installed. \" +\n `Cause: ${err instanceof Error ? err.message : String(err)}`,\n );\n }\n}\n\nfunction parseHeaders(values: string[]): Record<string, string> {\n const headers: Record<string, string> = {};\n for (const v of values) {\n const i = v.indexOf(\":\");\n if (i === -1) continue;\n headers[v.slice(0, i).trim()] = v.slice(i + 1).trim();\n }\n return headers;\n}\n\ninterface EvalOptions {\n url: string;\n strict?: boolean;\n root?: string;\n header?: string[];\n profile?: string;\n databricksHost?: string;\n databricksToken?: string;\n experiment?: string;\n judgeModel?: string;\n concurrency?: number;\n warehouseId?: string;\n}\n\n/** Resolved Databricks host + bearer (either field may be absent). */\ntype Auth = { host?: string; token?: string };\n\n/**\n * Native MLflow \"Evaluation run\" config — only when creds + an experiment are\n * all present (traces live in the app; the run + scores are driven from here).\n */\nfunction resolveMlflow(opts: EvalOptions, auth: Auth) {\n const experimentId = opts.experiment ?? process.env.MLFLOW_EXPERIMENT_ID;\n if (!(auth.host && auth.token && experimentId)) return undefined;\n // UC-backed experiments need a SQL warehouse to write assessments to their\n // V4 traces. Mirror mlflow's env var, and accept the common DATABRICKS one.\n const sqlWarehouseId =\n opts.warehouseId ??\n process.env.MLFLOW_TRACING_SQL_WAREHOUSE_ID ??\n process.env.DATABRICKS_WAREHOUSE_ID;\n return {\n host: auth.host,\n token: auth.token,\n experimentId,\n ...(sqlWarehouseId ? { sqlWarehouseId } : {}),\n };\n}\n\n/** LLM-as-judge config — reuses the Databricks creds + a judge serving endpoint. */\nfunction resolveJudge(opts: EvalOptions, auth: Auth) {\n const model = opts.judgeModel ?? process.env.APPKIT_JUDGE_MODEL;\n return model && auth.host && auth.token\n ? { host: auth.host, token: auth.token, model }\n : undefined;\n}\n\n/** Progress reporter: stream each eval as it runs instead of going silent. */\nfunction makeProgressReporter(\n runner: EvalRunner,\n url: string,\n): (event: EvalProgress) => void {\n return (event) => {\n switch (event.type) {\n case \"discovered\":\n console.log(\n `Running ${event.total} eval${event.total === 1 ? \"\" : \"s\"} against ${url}\\n`,\n );\n break;\n case \"run-created\":\n console.log(`MLflow evaluation run: ${event.runId}\\n`);\n break;\n case \"result\": {\n // One full line per completion — evals run concurrently, so a split\n // \"start … glyph\" prefix would interleave into garbage.\n console.log(\n `[${event.index + 1}/${event.total}] ${runner.formatEvalHeadline(event.result)}`,\n );\n for (const line of runner.formatEvalDetail(event.result)) {\n console.log(line);\n }\n break;\n }\n }\n };\n}\n\nfunction formatFailureLine(f: {\n traceId: string;\n status?: number;\n error?: string;\n}): string {\n return ` ✗ trace ${f.traceId}: ${f.status ?? \"\"} ${f.error ?? \"\"}`.trim();\n}\n\n/** Print the MLflow assessment/finish outcome after a run that created one. */\nfunction printMlflowOutcome(\n mlflow: NonNullable<EvalRunSummary[\"mlflow\"]>,\n): void {\n const { report, finish } = mlflow;\n console.log(\n `MLflow: ${report.written} assessment(s) written` +\n (report.skipped ? `, ${report.skipped} skipped` : \"\") +\n (report.failures.length ? `, ${report.failures.length} failed` : \"\"),\n );\n for (const f of report.failures) {\n console.error(formatFailureLine(f));\n }\n if (finish.metricsError) {\n console.error(` ⚠ metrics not logged: ${finish.metricsError}`);\n }\n if (!finish.finished) {\n console.error(\n ` ✗ run left RUNNING — failed to finish: ${finish.finishError ?? \"unknown\"}`,\n );\n }\n}\n\nasync function runAgentEval(\n filter: string | undefined,\n opts: EvalOptions,\n): Promise<void> {\n const runner = await loadRunner();\n\n // Resolve Databricks host + bearer the AppKit-native way: an explicit\n // host/token (or DATABRICKS_* env) wins; otherwise the SDK mints an OAuth\n // token from the CLI profile — so no hand-set PAT is required.\n const auth: Auth =\n (await runner.resolveDatabricksAuth({\n profile: opts.profile ?? process.env.DATABRICKS_CONFIG_PROFILE,\n host: opts.databricksHost ?? process.env.DATABRICKS_HOST,\n token: opts.databricksToken ?? process.env.DATABRICKS_TOKEN,\n })) ?? {};\n\n let summary: EvalRunSummary;\n try {\n summary = await runner.runEvalsInDir({\n rootDir: opts.root,\n baseUrl: opts.url,\n filter,\n strict: opts.strict,\n headers: opts.header ? parseHeaders(opts.header) : undefined,\n concurrency: opts.concurrency,\n mlflow: resolveMlflow(opts, auth),\n judge: resolveJudge(opts, auth),\n onEvent: makeProgressReporter(runner, opts.url),\n });\n } catch (err) {\n // Setup failures (e.g. a bad --experiment for the MLflow run) reject before\n // any eval runs; surface a clean message + non-zero exit rather than an\n // unhandled promise rejection with a raw stack.\n console.error(\n `\\nEval run failed: ${err instanceof Error ? err.message : String(err)}`,\n );\n process.exitCode = 1;\n return;\n }\n console.log(`\\n${runner.formatSummaryLine(summary.results)}`);\n\n if (summary.mlflow) {\n printMlflowOutcome(summary.mlflow);\n } else {\n console.log(\n \"\\nMLflow evaluation run skipped — pass --experiment (or set\" +\n \" MLFLOW_EXPERIMENT_ID) plus --profile/--databricks-host to create one.\",\n );\n }\n\n if (!runner.summarize(summary.results).allPassed) {\n process.exitCode = 1;\n }\n}\n\nexport const agentEvalCommand = new Command(\"eval\")\n .description(\n \"Run agent evals (server/agents/<id>/evals/*.eval.ts) against a running app\",\n )\n .argument(\n \"[filter]\",\n \"Only run evals whose <agent>/<id> contains this substring (or an exact agent id)\",\n )\n .option(\"--url <url>\", \"Base URL of the running app\", \"http://localhost:3000\")\n .option(\"--strict\", \"Fail on soft-assertion misses too\", false)\n .option(\n \"--concurrency <n>\",\n \"Max evals to run concurrently (default 4; keep at or below the app's max concurrent streams per user)\",\n (v) => Number.parseInt(v, 10),\n )\n .option(\n \"--root <dir>\",\n \"Project root containing server/agents/ (default: cwd)\",\n )\n .option(\n \"--header <header...>\",\n \"Extra request header as 'Key: value' (repeatable)\",\n )\n .option(\n \"--profile <name>\",\n \"Databricks CLI profile to authenticate with via OAuth (default: DATABRICKS_CONFIG_PROFILE)\",\n )\n .option(\n \"--databricks-host <host>\",\n \"Databricks host for writing MLflow assessments (default: DATABRICKS_HOST)\",\n )\n .option(\n \"--databricks-token <token>\",\n \"Databricks token for writing MLflow assessments (default: DATABRICKS_TOKEN)\",\n )\n .option(\n \"--experiment <id>\",\n \"MLflow experiment id for the evaluation run (default: MLFLOW_EXPERIMENT_ID)\",\n )\n .option(\n \"--warehouse-id <id>\",\n \"SQL warehouse id for writing assessments to UC-backed experiments (default: MLFLOW_TRACING_SQL_WAREHOUSE_ID or DATABRICKS_WAREHOUSE_ID)\",\n )\n .option(\n \"--judge-model <endpoint>\",\n \"Databricks serving endpoint to use as the LLM judge for t.judge.* (default: APPKIT_JUDGE_MODEL)\",\n )\n .action(runAgentEval);\n"],"mappings":";;;;;;;;AAuDA,eAAe,aAAkC;CAC/C,MAAM,OAAO;AACb,KAAI;AACF,SAAQ,MAAM,OAAO;UACd,KAAK;AACZ,QAAM,IAAI,MACR,yHAEY,eAAe,QAAQ,IAAI,UAAU,OAAO,IAAI,GAC7D;;;AAIL,SAAS,aAAa,QAA0C;CAC9D,MAAM,UAAkC,EAAE;AAC1C,MAAK,MAAM,KAAK,QAAQ;EACtB,MAAM,IAAI,EAAE,QAAQ,IAAI;AACxB,MAAI,MAAM,GAAI;AACd,UAAQ,EAAE,MAAM,GAAG,EAAE,CAAC,MAAM,IAAI,EAAE,MAAM,IAAI,EAAE,CAAC,MAAM;;AAEvD,QAAO;;;;;;AAwBT,SAAS,cAAc,MAAmB,MAAY;CACpD,MAAM,eAAe,KAAK,cAAc,QAAQ,IAAI;AACpD,KAAI,EAAE,KAAK,QAAQ,KAAK,SAAS,cAAe,QAAO;CAGvD,MAAM,iBACJ,KAAK,eACL,QAAQ,IAAI,mCACZ,QAAQ,IAAI;AACd,QAAO;EACL,MAAM,KAAK;EACX,OAAO,KAAK;EACZ;EACA,GAAI,iBAAiB,EAAE,gBAAgB,GAAG,EAAE;EAC7C;;;AAIH,SAAS,aAAa,MAAmB,MAAY;CACnD,MAAM,QAAQ,KAAK,cAAc,QAAQ,IAAI;AAC7C,QAAO,SAAS,KAAK,QAAQ,KAAK,QAC9B;EAAE,MAAM,KAAK;EAAM,OAAO,KAAK;EAAO;EAAO,GAC7C;;;AAIN,SAAS,qBACP,QACA,KAC+B;AAC/B,SAAQ,UAAU;AAChB,UAAQ,MAAM,MAAd;GACE,KAAK;AACH,YAAQ,IACN,WAAW,MAAM,MAAM,OAAO,MAAM,UAAU,IAAI,KAAK,IAAI,WAAW,IAAI,IAC3E;AACD;GACF,KAAK;AACH,YAAQ,IAAI,0BAA0B,MAAM,MAAM,IAAI;AACtD;GACF,KAAK;AAGH,YAAQ,IACN,IAAI,MAAM,QAAQ,EAAE,GAAG,MAAM,MAAM,IAAI,OAAO,mBAAmB,MAAM,OAAO,GAC/E;AACD,SAAK,MAAM,QAAQ,OAAO,iBAAiB,MAAM,OAAO,CACtD,SAAQ,IAAI,KAAK;AAEnB;;;;AAMR,SAAS,kBAAkB,GAIhB;AACT,QAAO,aAAa,EAAE,QAAQ,IAAI,EAAE,UAAU,GAAG,GAAG,EAAE,SAAS,KAAK,MAAM;;;AAI5E,SAAS,mBACP,QACM;CACN,MAAM,EAAE,QAAQ,WAAW;AAC3B,SAAQ,IACN,WAAW,OAAO,QAAQ,2BACvB,OAAO,UAAU,KAAK,OAAO,QAAQ,YAAY,OACjD,OAAO,SAAS,SAAS,KAAK,OAAO,SAAS,OAAO,WAAW,IACpE;AACD,MAAK,MAAM,KAAK,OAAO,SACrB,SAAQ,MAAM,kBAAkB,EAAE,CAAC;AAErC,KAAI,OAAO,aACT,SAAQ,MAAM,2BAA2B,OAAO,eAAe;AAEjE,KAAI,CAAC,OAAO,SACV,SAAQ,MACN,4CAA4C,OAAO,eAAe,YACnE;;AAIL,eAAe,aACb,QACA,MACe;CACf,MAAM,SAAS,MAAM,YAAY;CAKjC,MAAM,OACH,MAAM,OAAO,sBAAsB;EAClC,SAAS,KAAK,WAAW,QAAQ,IAAI;EACrC,MAAM,KAAK,kBAAkB,QAAQ,IAAI;EACzC,OAAO,KAAK,mBAAmB,QAAQ,IAAI;EAC5C,CAAC,IAAK,EAAE;CAEX,IAAI;AACJ,KAAI;AACF,YAAU,MAAM,OAAO,cAAc;GACnC,SAAS,KAAK;GACd,SAAS,KAAK;GACd;GACA,QAAQ,KAAK;GACb,SAAS,KAAK,SAAS,aAAa,KAAK,OAAO,GAAG;GACnD,aAAa,KAAK;GAClB,QAAQ,cAAc,MAAM,KAAK;GACjC,OAAO,aAAa,MAAM,KAAK;GAC/B,SAAS,qBAAqB,QAAQ,KAAK,IAAI;GAChD,CAAC;UACK,KAAK;AAIZ,UAAQ,MACN,sBAAsB,eAAe,QAAQ,IAAI,UAAU,OAAO,IAAI,GACvE;AACD,UAAQ,WAAW;AACnB;;AAEF,SAAQ,IAAI,KAAK,OAAO,kBAAkB,QAAQ,QAAQ,GAAG;AAE7D,KAAI,QAAQ,OACV,oBAAmB,QAAQ,OAAO;KAElC,SAAQ,IACN,oIAED;AAGH,KAAI,CAAC,OAAO,UAAU,QAAQ,QAAQ,CAAC,UACrC,SAAQ,WAAW;;AAIvB,MAAa,mBAAmB,IAAI,QAAQ,OAAO,CAChD,YACC,6EACD,CACA,SACC,YACA,mFACD,CACA,OAAO,eAAe,+BAA+B,wBAAwB,CAC7E,OAAO,YAAY,qCAAqC,MAAM,CAC9D,OACC,qBACA,0GACC,MAAM,OAAO,SAAS,GAAG,GAAG,CAC9B,CACA,OACC,gBACA,wDACD,CACA,OACC,wBACA,oDACD,CACA,OACC,oBACA,6FACD,CACA,OACC,4BACA,4EACD,CACA,OACC,8BACA,8EACD,CACA,OACC,qBACA,8EACD,CACA,OACC,uBACA,0IACD,CACA,OACC,4BACA,kGACD,CACA,OAAO,aAAa"}
1
+ {"version":3,"file":"eval.js","names":[],"sources":["../../../../src/cli/commands/agent/eval.ts"],"sourcesContent":["import { Command } from \"commander\";\n\ninterface EvalRunSummary {\n results: unknown[];\n mlflow?: {\n runId: string;\n report: {\n written: number;\n skipped: number;\n failures: Array<{ traceId: string; status?: number; error?: string }>;\n };\n finish: { finished: boolean; metricsError?: string; finishError?: string };\n };\n}\n\ntype EvalProgress =\n | { type: \"discovered\"; total: number }\n | { type: \"run-created\"; runId: string }\n | { type: \"start\"; id: string; index: number; total: number }\n | { type: \"result\"; result: unknown; index: number; total: number };\n\n/** Subset of `@databricks/appkit/beta`'s eval runner used by this command. */\ninterface EvalRunner {\n runEvalsInDir(opts: {\n rootDir?: string;\n baseUrl: string;\n filter?: string;\n strict?: boolean;\n headers?: Record<string, string>;\n concurrency?: number;\n mlflow?: {\n host: string;\n token: string;\n experimentId: string;\n sqlWarehouseId?: string;\n };\n judge?: { host: string; token: string; model: string };\n workspaceClient?: unknown;\n warehouseId?: string;\n onEvent?: (event: EvalProgress) => void;\n }): Promise<EvalRunSummary>;\n resolveDatabricksAuth(opts: {\n profile?: string;\n host?: string;\n token?: string;\n }): Promise<{ host: string; token: string } | undefined>;\n resolveWorkspaceClient(opts: {\n profile?: string;\n host?: string;\n token?: string;\n }): unknown;\n formatEvalHeadline(result: unknown): string;\n evalGlyph(result: unknown): string;\n formatEvalDetail(result: unknown): string[];\n formatSummaryLine(results: unknown[]): string;\n summarize(results: unknown[]): { allPassed: boolean };\n}\n\n/**\n * Loaded at runtime from the consuming project so this command (which ships in\n * `@databricks/shared`) doesn't take a build-time dependency on appkit. The\n * specifier is a variable so the type checker treats it as `any`.\n */\nasync function loadRunner(): Promise<EvalRunner> {\n const spec = \"@databricks/appkit/beta\";\n try {\n return (await import(spec)) as unknown as EvalRunner;\n } catch (err) {\n throw new Error(\n \"Could not load @databricks/appkit. Run `appkit agent eval` from a \" +\n \"project with @databricks/appkit installed. \" +\n `Cause: ${err instanceof Error ? err.message : String(err)}`,\n );\n }\n}\n\nfunction parseHeaders(values: string[]): Record<string, string> {\n const headers: Record<string, string> = {};\n for (const v of values) {\n const i = v.indexOf(\":\");\n if (i === -1) continue;\n headers[v.slice(0, i).trim()] = v.slice(i + 1).trim();\n }\n return headers;\n}\n\ninterface EvalOptions {\n url: string;\n strict?: boolean;\n root?: string;\n header?: string[];\n profile?: string;\n databricksHost?: string;\n databricksToken?: string;\n experiment?: string;\n judgeModel?: string;\n concurrency?: number;\n warehouseId?: string;\n}\n\n/** Resolved Databricks host + bearer (either field may be absent). */\ntype Auth = { host?: string; token?: string };\n\n/**\n * Native MLflow \"Evaluation run\" config — only when creds + an experiment are\n * all present (traces live in the app; the run + scores are driven from here).\n */\nfunction resolveMlflow(opts: EvalOptions, auth: Auth) {\n const experimentId = opts.experiment ?? process.env.MLFLOW_EXPERIMENT_ID;\n if (!(auth.host && auth.token && experimentId)) return undefined;\n // UC-backed experiments need a SQL warehouse to write assessments to their\n // V4 traces. Mirror mlflow's env var, and accept the common DATABRICKS one.\n const sqlWarehouseId =\n opts.warehouseId ??\n process.env.MLFLOW_TRACING_SQL_WAREHOUSE_ID ??\n process.env.DATABRICKS_WAREHOUSE_ID;\n return {\n host: auth.host,\n token: auth.token,\n experimentId,\n ...(sqlWarehouseId ? { sqlWarehouseId } : {}),\n };\n}\n\n/** LLM-as-judge config — reuses the Databricks creds + a judge serving endpoint. */\nfunction resolveJudge(opts: EvalOptions, auth: Auth) {\n const model = opts.judgeModel ?? process.env.APPKIT_JUDGE_MODEL;\n return model && auth.host && auth.token\n ? { host: auth.host, token: auth.token, model }\n : undefined;\n}\n\n/** Progress reporter: stream each eval as it runs instead of going silent. */\nfunction makeProgressReporter(\n runner: EvalRunner,\n url: string,\n): (event: EvalProgress) => void {\n return (event) => {\n switch (event.type) {\n case \"discovered\":\n console.log(\n `Running ${event.total} eval${event.total === 1 ? \"\" : \"s\"} against ${url}\\n`,\n );\n break;\n case \"run-created\":\n console.log(`MLflow evaluation run: ${event.runId}\\n`);\n break;\n case \"result\": {\n // One full line per completion — evals run concurrently, so a split\n // \"start … glyph\" prefix would interleave into garbage.\n console.log(\n `[${event.index + 1}/${event.total}] ${runner.formatEvalHeadline(event.result)}`,\n );\n for (const line of runner.formatEvalDetail(event.result)) {\n console.log(line);\n }\n break;\n }\n }\n };\n}\n\nfunction formatFailureLine(f: {\n traceId: string;\n status?: number;\n error?: string;\n}): string {\n return ` ✗ trace ${f.traceId}: ${f.status ?? \"\"} ${f.error ?? \"\"}`.trim();\n}\n\n/** Print the MLflow assessment/finish outcome after a run that created one. */\nfunction printMlflowOutcome(\n mlflow: NonNullable<EvalRunSummary[\"mlflow\"]>,\n): void {\n const { report, finish } = mlflow;\n console.log(\n `MLflow: ${report.written} assessment(s) written` +\n (report.skipped ? `, ${report.skipped} skipped` : \"\") +\n (report.failures.length ? `, ${report.failures.length} failed` : \"\"),\n );\n for (const f of report.failures) {\n console.error(formatFailureLine(f));\n }\n if (finish.metricsError) {\n console.error(` ⚠ metrics not logged: ${finish.metricsError}`);\n }\n if (!finish.finished) {\n console.error(\n ` ✗ run left RUNNING — failed to finish: ${finish.finishError ?? \"unknown\"}`,\n );\n }\n}\n\nasync function runAgentEval(\n filter: string | undefined,\n opts: EvalOptions,\n): Promise<void> {\n const runner = await loadRunner();\n\n // Resolve Databricks host + bearer the AppKit-native way: an explicit\n // host/token (or DATABRICKS_* env) wins; otherwise the SDK mints an OAuth\n // token from the CLI profile — so no hand-set PAT is required.\n const auth: Auth =\n (await runner.resolveDatabricksAuth({\n profile: opts.profile ?? process.env.DATABRICKS_CONFIG_PROFILE,\n host: opts.databricksHost ?? process.env.DATABRICKS_HOST,\n token: opts.databricksToken ?? process.env.DATABRICKS_TOKEN,\n })) ?? {};\n\n // Managed-dataset reads: a workspace client (same profile/host/token) + a SQL\n // warehouse. Only needed by evals that declare `dataset`.\n const warehouseId = opts.warehouseId ?? process.env.DATABRICKS_WAREHOUSE_ID;\n const workspaceClient = runner.resolveWorkspaceClient({\n profile: opts.profile ?? process.env.DATABRICKS_CONFIG_PROFILE,\n host: opts.databricksHost ?? process.env.DATABRICKS_HOST,\n token: opts.databricksToken ?? process.env.DATABRICKS_TOKEN,\n });\n\n let summary: EvalRunSummary;\n try {\n summary = await runner.runEvalsInDir({\n rootDir: opts.root,\n baseUrl: opts.url,\n filter,\n strict: opts.strict,\n headers: opts.header ? parseHeaders(opts.header) : undefined,\n concurrency: opts.concurrency,\n mlflow: resolveMlflow(opts, auth),\n judge: resolveJudge(opts, auth),\n workspaceClient,\n warehouseId,\n onEvent: makeProgressReporter(runner, opts.url),\n });\n } catch (err) {\n // Setup failures (e.g. a bad --experiment for the MLflow run) reject before\n // any eval runs; surface a clean message + non-zero exit rather than an\n // unhandled promise rejection with a raw stack.\n console.error(\n `\\nEval run failed: ${err instanceof Error ? err.message : String(err)}`,\n );\n process.exitCode = 1;\n return;\n }\n console.log(`\\n${runner.formatSummaryLine(summary.results)}`);\n\n if (summary.mlflow) {\n printMlflowOutcome(summary.mlflow);\n } else {\n console.log(\n \"\\nMLflow evaluation run skipped — pass --experiment (or set\" +\n \" MLFLOW_EXPERIMENT_ID) plus --profile/--databricks-host to create one.\",\n );\n }\n\n if (!runner.summarize(summary.results).allPassed) {\n process.exitCode = 1;\n }\n}\n\nexport const agentEvalCommand = new Command(\"eval\")\n .description(\n \"Run agent evals (server/agents/<id>/evals/*.eval.ts) against a running app\",\n )\n .argument(\n \"[filter]\",\n \"Only run evals whose <agent>/<id> contains this substring (or an exact agent id)\",\n )\n .option(\"--url <url>\", \"Base URL of the running app\", \"http://localhost:3000\")\n .option(\"--strict\", \"Fail on soft-assertion misses too\", false)\n .option(\n \"--concurrency <n>\",\n \"Max evals to run concurrently (default 4; keep at or below the app's max concurrent streams per user)\",\n (v) => Number.parseInt(v, 10),\n )\n .option(\n \"--root <dir>\",\n \"Project root containing server/agents/ (default: cwd)\",\n )\n .option(\n \"--header <header...>\",\n \"Extra request header as 'Key: value' (repeatable)\",\n )\n .option(\n \"--profile <name>\",\n \"Databricks CLI profile to authenticate with via OAuth (default: DATABRICKS_CONFIG_PROFILE)\",\n )\n .option(\n \"--databricks-host <host>\",\n \"Databricks host for writing MLflow assessments (default: DATABRICKS_HOST)\",\n )\n .option(\n \"--databricks-token <token>\",\n \"Databricks token for writing MLflow assessments (default: DATABRICKS_TOKEN)\",\n )\n .option(\n \"--experiment <id>\",\n \"MLflow experiment id for the evaluation run (default: MLFLOW_EXPERIMENT_ID)\",\n )\n .option(\n \"--warehouse-id <id>\",\n \"SQL warehouse id for reading managed eval datasets and writing assessments to UC-backed experiments (default: DATABRICKS_WAREHOUSE_ID, or MLFLOW_TRACING_SQL_WAREHOUSE_ID for assessments)\",\n )\n .option(\n \"--judge-model <endpoint>\",\n \"Databricks serving endpoint to use as the LLM judge for t.judge.* (default: APPKIT_JUDGE_MODEL)\",\n )\n .action(runAgentEval);\n"],"mappings":";;;;;;;;AA+DA,eAAe,aAAkC;CAC/C,MAAM,OAAO;AACb,KAAI;AACF,SAAQ,MAAM,OAAO;UACd,KAAK;AACZ,QAAM,IAAI,MACR,yHAEY,eAAe,QAAQ,IAAI,UAAU,OAAO,IAAI,GAC7D;;;AAIL,SAAS,aAAa,QAA0C;CAC9D,MAAM,UAAkC,EAAE;AAC1C,MAAK,MAAM,KAAK,QAAQ;EACtB,MAAM,IAAI,EAAE,QAAQ,IAAI;AACxB,MAAI,MAAM,GAAI;AACd,UAAQ,EAAE,MAAM,GAAG,EAAE,CAAC,MAAM,IAAI,EAAE,MAAM,IAAI,EAAE,CAAC,MAAM;;AAEvD,QAAO;;;;;;AAwBT,SAAS,cAAc,MAAmB,MAAY;CACpD,MAAM,eAAe,KAAK,cAAc,QAAQ,IAAI;AACpD,KAAI,EAAE,KAAK,QAAQ,KAAK,SAAS,cAAe,QAAO;CAGvD,MAAM,iBACJ,KAAK,eACL,QAAQ,IAAI,mCACZ,QAAQ,IAAI;AACd,QAAO;EACL,MAAM,KAAK;EACX,OAAO,KAAK;EACZ;EACA,GAAI,iBAAiB,EAAE,gBAAgB,GAAG,EAAE;EAC7C;;;AAIH,SAAS,aAAa,MAAmB,MAAY;CACnD,MAAM,QAAQ,KAAK,cAAc,QAAQ,IAAI;AAC7C,QAAO,SAAS,KAAK,QAAQ,KAAK,QAC9B;EAAE,MAAM,KAAK;EAAM,OAAO,KAAK;EAAO;EAAO,GAC7C;;;AAIN,SAAS,qBACP,QACA,KAC+B;AAC/B,SAAQ,UAAU;AAChB,UAAQ,MAAM,MAAd;GACE,KAAK;AACH,YAAQ,IACN,WAAW,MAAM,MAAM,OAAO,MAAM,UAAU,IAAI,KAAK,IAAI,WAAW,IAAI,IAC3E;AACD;GACF,KAAK;AACH,YAAQ,IAAI,0BAA0B,MAAM,MAAM,IAAI;AACtD;GACF,KAAK;AAGH,YAAQ,IACN,IAAI,MAAM,QAAQ,EAAE,GAAG,MAAM,MAAM,IAAI,OAAO,mBAAmB,MAAM,OAAO,GAC/E;AACD,SAAK,MAAM,QAAQ,OAAO,iBAAiB,MAAM,OAAO,CACtD,SAAQ,IAAI,KAAK;AAEnB;;;;AAMR,SAAS,kBAAkB,GAIhB;AACT,QAAO,aAAa,EAAE,QAAQ,IAAI,EAAE,UAAU,GAAG,GAAG,EAAE,SAAS,KAAK,MAAM;;;AAI5E,SAAS,mBACP,QACM;CACN,MAAM,EAAE,QAAQ,WAAW;AAC3B,SAAQ,IACN,WAAW,OAAO,QAAQ,2BACvB,OAAO,UAAU,KAAK,OAAO,QAAQ,YAAY,OACjD,OAAO,SAAS,SAAS,KAAK,OAAO,SAAS,OAAO,WAAW,IACpE;AACD,MAAK,MAAM,KAAK,OAAO,SACrB,SAAQ,MAAM,kBAAkB,EAAE,CAAC;AAErC,KAAI,OAAO,aACT,SAAQ,MAAM,2BAA2B,OAAO,eAAe;AAEjE,KAAI,CAAC,OAAO,SACV,SAAQ,MACN,4CAA4C,OAAO,eAAe,YACnE;;AAIL,eAAe,aACb,QACA,MACe;CACf,MAAM,SAAS,MAAM,YAAY;CAKjC,MAAM,OACH,MAAM,OAAO,sBAAsB;EAClC,SAAS,KAAK,WAAW,QAAQ,IAAI;EACrC,MAAM,KAAK,kBAAkB,QAAQ,IAAI;EACzC,OAAO,KAAK,mBAAmB,QAAQ,IAAI;EAC5C,CAAC,IAAK,EAAE;CAIX,MAAM,cAAc,KAAK,eAAe,QAAQ,IAAI;CACpD,MAAM,kBAAkB,OAAO,uBAAuB;EACpD,SAAS,KAAK,WAAW,QAAQ,IAAI;EACrC,MAAM,KAAK,kBAAkB,QAAQ,IAAI;EACzC,OAAO,KAAK,mBAAmB,QAAQ,IAAI;EAC5C,CAAC;CAEF,IAAI;AACJ,KAAI;AACF,YAAU,MAAM,OAAO,cAAc;GACnC,SAAS,KAAK;GACd,SAAS,KAAK;GACd;GACA,QAAQ,KAAK;GACb,SAAS,KAAK,SAAS,aAAa,KAAK,OAAO,GAAG;GACnD,aAAa,KAAK;GAClB,QAAQ,cAAc,MAAM,KAAK;GACjC,OAAO,aAAa,MAAM,KAAK;GAC/B;GACA;GACA,SAAS,qBAAqB,QAAQ,KAAK,IAAI;GAChD,CAAC;UACK,KAAK;AAIZ,UAAQ,MACN,sBAAsB,eAAe,QAAQ,IAAI,UAAU,OAAO,IAAI,GACvE;AACD,UAAQ,WAAW;AACnB;;AAEF,SAAQ,IAAI,KAAK,OAAO,kBAAkB,QAAQ,QAAQ,GAAG;AAE7D,KAAI,QAAQ,OACV,oBAAmB,QAAQ,OAAO;KAElC,SAAQ,IACN,oIAED;AAGH,KAAI,CAAC,OAAO,UAAU,QAAQ,QAAQ,CAAC,UACrC,SAAQ,WAAW;;AAIvB,MAAa,mBAAmB,IAAI,QAAQ,OAAO,CAChD,YACC,6EACD,CACA,SACC,YACA,mFACD,CACA,OAAO,eAAe,+BAA+B,wBAAwB,CAC7E,OAAO,YAAY,qCAAqC,MAAM,CAC9D,OACC,qBACA,0GACC,MAAM,OAAO,SAAS,GAAG,GAAG,CAC9B,CACA,OACC,gBACA,wDACD,CACA,OACC,wBACA,oDACD,CACA,OACC,oBACA,6FACD,CACA,OACC,4BACA,4EACD,CACA,OACC,8BACA,8EACD,CACA,OACC,qBACA,8EACD,CACA,OACC,uBACA,6LACD,CACA,OACC,4BACA,kGACD,CACA,OAAO,aAAa"}
@@ -5,7 +5,7 @@ import * as class_variance_authority_types0 from "class-variance-authority/types
5
5
 
6
6
  //#region src/react/ui/button.d.ts
7
7
  declare const buttonVariants: (props?: ({
8
- variant?: "default" | "destructive" | "secondary" | "outline" | "ghost" | "link" | null | undefined;
8
+ variant?: "default" | "link" | "destructive" | "secondary" | "outline" | "ghost" | null | undefined;
9
9
  size?: "default" | "sm" | "lg" | "icon" | "icon-sm" | "icon-lg" | null | undefined;
10
10
  } & class_variance_authority_types0.ClassProp) | undefined) => string;
11
11
  /** Clickable button with multiple variants and sizes */
@@ -30,6 +30,7 @@ console.error(error.toJSON()); // Safe for logging, sensitive values redacted
30
30
  * [`AuthenticationError`](./docs/api/appkit/Class.AuthenticationError.md)
31
31
  * [`ConfigurationError`](./docs/api/appkit/Class.ConfigurationError.md)
32
32
  * [`ConnectionError`](./docs/api/appkit/Class.ConnectionError.md)
33
+ * [`DatabaseValidationError`](./docs/api/appkit/Class.DatabaseValidationError.md)
33
34
  * [`ExecutionError`](./docs/api/appkit/Class.ExecutionError.md)
34
35
  * [`InitializationError`](./docs/api/appkit/Class.InitializationError.md)
35
36
  * [`ServerError`](./docs/api/appkit/Class.ServerError.md)
@@ -0,0 +1,191 @@
1
+ # Class: DatabaseValidationError
2
+
3
+ Deliberate validation failure raised by a database mutation hook. Generated routes answer `422` and echo only the issues naming a public column; every other failure raised inside a hook stays an opaque server error.
4
+
5
+ ## Extends[​](#extends "Direct link to Extends")
6
+
7
+ * [`AppKitError`](./docs/api/appkit/Class.AppKitError.md)
8
+
9
+ ## Constructors[​](#constructors "Direct link to Constructors")
10
+
11
+ ### Constructor[​](#constructor "Direct link to Constructor")
12
+
13
+ ```ts
14
+ new DatabaseValidationError(message: string, issues: readonly DatabaseValidationIssue[]): DatabaseValidationError;
15
+
16
+ ```
17
+
18
+ #### Parameters[​](#parameters "Direct link to Parameters")
19
+
20
+ | Parameter | Type | Default value |
21
+ | --------- | ----------------------------------------------------------------------------------------------------- | ------------- |
22
+ | `message` | `string` | `undefined` |
23
+ | `issues` | readonly [`DatabaseValidationIssue`](./docs/api/appkit/Interface.DatabaseValidationIssue.md)\[] | `[]` |
24
+
25
+ #### Returns[​](#returns "Direct link to Returns")
26
+
27
+ `DatabaseValidationError`
28
+
29
+ #### Overrides[​](#overrides "Direct link to Overrides")
30
+
31
+ [`AppKitError`](./docs/api/appkit/Class.AppKitError.md).[`constructor`](./docs/api/appkit/Class.AppKitError.md#constructor)
32
+
33
+ ## Properties[​](#properties "Direct link to Properties")
34
+
35
+ ### \_clientMessage?[​](#_clientmessage "Direct link to _clientMessage?")
36
+
37
+ ```ts
38
+ protected readonly optional _clientMessage: string;
39
+
40
+ ```
41
+
42
+ Client-safe error message. When set, callers serializing the error to a client (SSE, HTTP body) MUST prefer `clientMessage` over `message` — `message` may contain raw upstream / SDK text including statement fragments, internal object names, and correlation IDs.
43
+
44
+ Subclasses can set this in their constructor for a fixed sanitized string. When unset, `clientMessage` defaults to a generic per-code string (see the getter), and the raw `message` is kept server-side only.
45
+
46
+ #### Inherited from[​](#inherited-from "Direct link to Inherited from")
47
+
48
+ [`AppKitError`](./docs/api/appkit/Class.AppKitError.md).[`_clientMessage`](./docs/api/appkit/Class.AppKitError.md#_clientmessage)
49
+
50
+ ***
51
+
52
+ ### cause?[​](#cause "Direct link to cause?")
53
+
54
+ ```ts
55
+ readonly optional cause: Error;
56
+
57
+ ```
58
+
59
+ Optional cause of the error
60
+
61
+ #### Inherited from[​](#inherited-from-1 "Direct link to Inherited from")
62
+
63
+ [`AppKitError`](./docs/api/appkit/Class.AppKitError.md).[`cause`](./docs/api/appkit/Class.AppKitError.md#cause)
64
+
65
+ ***
66
+
67
+ ### code[​](#code "Direct link to code")
68
+
69
+ ```ts
70
+ readonly code: "DATABASE_VALIDATION_ERROR" = "DATABASE_VALIDATION_ERROR";
71
+
72
+ ```
73
+
74
+ Error code for programmatic error handling
75
+
76
+ #### Overrides[​](#overrides-1 "Direct link to Overrides")
77
+
78
+ [`AppKitError`](./docs/api/appkit/Class.AppKitError.md).[`code`](./docs/api/appkit/Class.AppKitError.md#code)
79
+
80
+ ***
81
+
82
+ ### context?[​](#context "Direct link to context?")
83
+
84
+ ```ts
85
+ readonly optional context: Record<string, unknown>;
86
+
87
+ ```
88
+
89
+ Additional context for the error
90
+
91
+ #### Inherited from[​](#inherited-from-2 "Direct link to Inherited from")
92
+
93
+ [`AppKitError`](./docs/api/appkit/Class.AppKitError.md).[`context`](./docs/api/appkit/Class.AppKitError.md#context)
94
+
95
+ ***
96
+
97
+ ### isRetryable[​](#isretryable "Direct link to isRetryable")
98
+
99
+ ```ts
100
+ readonly isRetryable: false = false;
101
+
102
+ ```
103
+
104
+ Whether this error type is generally safe to retry
105
+
106
+ #### Overrides[​](#overrides-2 "Direct link to Overrides")
107
+
108
+ [`AppKitError`](./docs/api/appkit/Class.AppKitError.md).[`isRetryable`](./docs/api/appkit/Class.AppKitError.md#isretryable)
109
+
110
+ ***
111
+
112
+ ### issues[​](#issues "Direct link to issues")
113
+
114
+ ```ts
115
+ readonly issues: readonly DatabaseValidationIssue[];
116
+
117
+ ```
118
+
119
+ ***
120
+
121
+ ### statusCode[​](#statuscode "Direct link to statusCode")
122
+
123
+ ```ts
124
+ readonly statusCode: 422 = 422;
125
+
126
+ ```
127
+
128
+ HTTP status code suggestion (can be overridden)
129
+
130
+ #### Overrides[​](#overrides-3 "Direct link to Overrides")
131
+
132
+ [`AppKitError`](./docs/api/appkit/Class.AppKitError.md).[`statusCode`](./docs/api/appkit/Class.AppKitError.md#statuscode)
133
+
134
+ ## Accessors[​](#accessors "Direct link to Accessors")
135
+
136
+ ### clientMessage[​](#clientmessage "Direct link to clientMessage")
137
+
138
+ #### Get Signature[​](#get-signature "Direct link to Get Signature")
139
+
140
+ ```ts
141
+ get clientMessage(): string;
142
+
143
+ ```
144
+
145
+ Sanitized message safe to forward to clients. Override in subclasses if a more specific default is appropriate.
146
+
147
+ ##### Returns[​](#returns-1 "Direct link to Returns")
148
+
149
+ `string`
150
+
151
+ #### Inherited from[​](#inherited-from-3 "Direct link to Inherited from")
152
+
153
+ [`AppKitError`](./docs/api/appkit/Class.AppKitError.md).[`clientMessage`](./docs/api/appkit/Class.AppKitError.md#clientmessage)
154
+
155
+ ## Methods[​](#methods "Direct link to Methods")
156
+
157
+ ### toJSON()[​](#tojson "Direct link to toJSON()")
158
+
159
+ ```ts
160
+ toJSON(): Record<string, unknown>;
161
+
162
+ ```
163
+
164
+ Convert error to JSON for logging/serialization. Sensitive values in context are automatically redacted.
165
+
166
+ #### Returns[​](#returns-2 "Direct link to Returns")
167
+
168
+ `Record`<`string`, `unknown`>
169
+
170
+ #### Inherited from[​](#inherited-from-4 "Direct link to Inherited from")
171
+
172
+ [`AppKitError`](./docs/api/appkit/Class.AppKitError.md).[`toJSON`](./docs/api/appkit/Class.AppKitError.md#tojson)
173
+
174
+ ***
175
+
176
+ ### toString()[​](#tostring "Direct link to toString()")
177
+
178
+ ```ts
179
+ toString(): string;
180
+
181
+ ```
182
+
183
+ Create a human-readable string representation
184
+
185
+ #### Returns[​](#returns-3 "Direct link to Returns")
186
+
187
+ `string`
188
+
189
+ #### Inherited from[​](#inherited-from-5 "Direct link to Inherited from")
190
+
191
+ [`AppKitError`](./docs/api/appkit/Class.AppKitError.md).[`toString`](./docs/api/appkit/Class.AppKitError.md#tostring)
@@ -5,7 +5,7 @@ function defineSchema<TTables>(builder: (context: SchemaBuilderContext) => TTabl
5
5
 
6
6
  ```
7
7
 
8
- Compile one declared schema. The returned type keeps the table names the builder returned, so `crudRoutes` and `hooks` can name only real tables.
8
+ Compile one declared schema. The returned type keeps the table names the builder returned, so `api.tables` and `hooks` can name only real tables.
9
9
 
10
10
  ## Type Parameters[​](#type-parameters "Direct link to Type Parameters")
11
11
 
@@ -0,0 +1,21 @@
1
+ # Function: readEvalDataset()
2
+
3
+ ```ts
4
+ function readEvalDataset(client: WorkspaceClient, options: ReadEvalDatasetOptions): Promise<DatasetRow[]>;
5
+
6
+ ```
7
+
8
+ Read a Databricks managed evaluation dataset (a Unity Catalog table with `inputs`/`expectations` columns) into rows, over the public SQL Statement Execution API. Reuses SQLWarehouseConnector for submit/poll/transform — its result transform already JSON-parses string columns into objects, so `inputs`/`expectations` come back as records whether the table stores them as JSON strings or structs.
9
+
10
+ The Python `mlflow.genai.datasets` API needs a Spark session (no TS equivalent), so we read the backing table directly.
11
+
12
+ ## Parameters[​](#parameters "Direct link to Parameters")
13
+
14
+ | Parameter | Type |
15
+ | --------- | --------------------------------------------------------------------------------------- |
16
+ | `client` | [`WorkspaceClient`](./docs/api/appkit/Interface.WorkspaceClient.md) |
17
+ | `options` | [`ReadEvalDatasetOptions`](./docs/api/appkit/Interface.ReadEvalDatasetOptions.md) |
18
+
19
+ ## Returns[​](#returns "Direct link to Returns")
20
+
21
+ `Promise`<[`DatasetRow`](./docs/api/appkit/Interface.DatasetRow.md)\[]>
@@ -0,0 +1,18 @@
1
+ # Function: resolveWorkspaceClient()
2
+
3
+ ```ts
4
+ function resolveWorkspaceClient(options: ResolveDatabricksAuthOptions): WorkspaceClient | undefined;
5
+
6
+ ```
7
+
8
+ Construct a Databricks `WorkspaceClient` for the eval runner — the object the SDK-backed connectors (e.g. `SQLWarehouseConnector`) take. An explicit host+token builds a PAT client; otherwise the profile (or ambient config) is used and the SDK resolves credentials, minting OAuth as needed. Returns `undefined` if construction throws (missing/invalid config).
9
+
10
+ ## Parameters[​](#parameters "Direct link to Parameters")
11
+
12
+ | Parameter | Type |
13
+ | --------- | --------------------------------------------------------------------------------------------------- |
14
+ | `options` | [`ResolveDatabricksAuthOptions`](./docs/api/appkit/Interface.ResolveDatabricksAuthOptions.md) |
15
+
16
+ ## Returns[​](#returns "Direct link to Returns")
17
+
18
+ [`WorkspaceClient`](./docs/api/appkit/Interface.WorkspaceClient.md) | `undefined`
@@ -11,7 +11,7 @@ atLeast(threshold: number): AssertionHandle;
11
11
 
12
12
  ```
13
13
 
14
- Soft assertion that passes only when the score is at least `threshold`.
14
+ Set the pass threshold for a scored assertion: it passes only when the score is at least `threshold`. Keeps the current severity (gate unless also chained with `.soft()`).
15
15
 
16
16
  #### Parameters[​](#parameters "Direct link to Parameters")
17
17
 
@@ -0,0 +1,21 @@
1
+ # Interface: DatabaseValidationIssue
2
+
3
+ One rejected field; `path` names public columns, never their values.
4
+
5
+ ## Properties[​](#properties "Direct link to Properties")
6
+
7
+ ### message[​](#message "Direct link to message")
8
+
9
+ ```ts
10
+ readonly message: string;
11
+
12
+ ```
13
+
14
+ ***
15
+
16
+ ### path[​](#path "Direct link to path")
17
+
18
+ ```ts
19
+ readonly path: readonly string[];
20
+
21
+ ```
@@ -0,0 +1,21 @@
1
+ # Interface: DatasetRow
2
+
3
+ One row of a managed evaluation dataset. `inputs` are the kwargs passed to the agent for the turn; `expectations` (when present) is the row's ground truth / guidelines. Mirrors the `{inputs, expectations}` shape of `mlflow.genai` datasets and of the Unity Catalog table backing a managed eval dataset.
4
+
5
+ ## Properties[​](#properties "Direct link to Properties")
6
+
7
+ ### expectations?[​](#expectations "Direct link to expectations?")
8
+
9
+ ```ts
10
+ optional expectations: Record<string, unknown>;
11
+
12
+ ```
13
+
14
+ ***
15
+
16
+ ### inputs[​](#inputs "Direct link to inputs")
17
+
18
+ ```ts
19
+ inputs: Record<string, unknown>;
20
+
21
+ ```