@databricks/appkit 0.74.0 → 0.75.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CLAUDE.md +3 -0
- package/dist/appkit/package.js +1 -1
- package/dist/beta.d.ts +4 -4
- package/dist/beta.js +3 -3
- package/dist/cli/commands/agent/eval.js +84 -9
- package/dist/cli/commands/agent/eval.js.map +1 -1
- package/dist/cli/commands/registry/add.js +78 -7
- package/dist/cli/commands/registry/add.js.map +1 -1
- package/dist/cli/commands/registry/config-writer.js +1 -1
- package/dist/cli/index.js +4 -1
- package/dist/cli/index.js.map +1 -1
- package/dist/evals/discover.d.ts +8 -1
- package/dist/evals/discover.d.ts.map +1 -1
- package/dist/evals/discover.js +33 -17
- package/dist/evals/discover.js.map +1 -1
- package/dist/evals/index.d.ts +3 -3
- package/dist/evals/index.js +2 -2
- package/dist/evals/run-evals.d.ts +9 -2
- package/dist/evals/run-evals.d.ts.map +1 -1
- package/dist/evals/run-evals.js +13 -2
- package/dist/evals/run-evals.js.map +1 -1
- package/dist/evals/types.d.ts +40 -2
- package/dist/evals/types.d.ts.map +1 -1
- package/dist/shared/src/schemas/manifest.d.ts +87 -87
- package/docs/api/appkit/Function.findRootEvalConfig.md +18 -0
- package/docs/api/appkit/Function.loadRootEvalConfig.md +18 -0
- package/docs/api/appkit/Interface.EvalWebServer.md +47 -0
- package/docs/api/appkit.md +3 -0
- package/docs/plugins/agents.md +255 -4
- package/llms.txt +3 -0
- package/package.json +1 -1
- package/sbom.cdx.json +1 -1
package/docs/api/appkit.md
CHANGED
|
@@ -63,6 +63,7 @@ Documentation merge entry for Typedoc — combines the stable `@databricks/appki
|
|
|
63
63
|
| [EvalResult](./docs/api/appkit/Interface.EvalResult.md) | The outcome of running one eval. |
|
|
64
64
|
| [EvalRunSummary](./docs/api/appkit/Interface.EvalRunSummary.md) | - |
|
|
65
65
|
| [EvalSummary](./docs/api/appkit/Interface.EvalSummary.md) | - |
|
|
66
|
+
| [EvalWebServer](./docs/api/appkit/Interface.EvalWebServer.md) | Auto-start config for the app under test, à la Playwright's `webServer`. When set in a root `evals.config.ts`, the CLI boots the app before running evals and tears it down after — so you don't have to start the server by hand. |
|
|
66
67
|
| [FilePolicyUser](./docs/api/appkit/Interface.FilePolicyUser.md) | Minimal user identity passed to the policy function. |
|
|
67
68
|
| [FileResource](./docs/api/appkit/Interface.FileResource.md) | Describes the file or directory being acted upon. |
|
|
68
69
|
| [FunctionTool](./docs/api/appkit/Interface.FunctionTool.md) | - |
|
|
@@ -210,6 +211,7 @@ Documentation merge entry for Typedoc — combines the stable `@databricks/appki
|
|
|
210
211
|
| [evalGlyph](./docs/api/appkit/Function.evalGlyph.md) | Status glyph for a single eval result. |
|
|
211
212
|
| [executeFromRegistry](./docs/api/appkit/Function.executeFromRegistry.md) | Validates tool-call arguments against the entry's schema and invokes its handler. On validation failure, returns an LLM-friendly error string (matching the behavior of `tool()`) rather than throwing, so the model can self-correct on its next turn. |
|
|
212
213
|
| [extractServingEndpoints](./docs/api/appkit/Function.extractServingEndpoints.md) | Extract serving endpoint config from a server file by AST-parsing it. Looks for `serving({ endpoints: { alias: { env: "..." }, ... } })` calls and extracts the endpoint alias names and their environment variable mappings. |
|
|
214
|
+
| [findRootEvalConfig](./docs/api/appkit/Function.findRootEvalConfig.md) | Path to the root `evals.config.ts` (from [defineEvalConfig](./docs/api/appkit/Function.defineEvalConfig.md)) at `<rootDir>/evals.config.ts`, or `undefined` when absent. The root config holds run-wide settings (`baseUrl`, `webServer`); it's distinct from the per-agent configs found by [discoverEvalConfigs](./docs/api/appkit/Function.discoverEvalConfigs.md). |
|
|
213
215
|
| [findServerFile](./docs/api/appkit/Function.findServerFile.md) | Find the server entry file by checking candidate paths in order. |
|
|
214
216
|
| [fk](./docs/api/appkit/Function.fk.md) | Declare foreign-key to another column. |
|
|
215
217
|
| [formatEvalDetail](./docs/api/appkit/Function.formatEvalDetail.md) | Indented detail lines for a failing eval (error + failing assertions). |
|
|
@@ -240,6 +242,7 @@ Documentation merge entry for Typedoc — combines the stable `@databricks/appki
|
|
|
240
242
|
| [jsonb](./docs/api/appkit/Function.jsonb.md) | - |
|
|
241
243
|
| [loadAgentFromFile](./docs/api/appkit/Function.loadAgentFromFile.md) | Loads a single markdown agent file and resolves its frontmatter against registered plugin toolkits + ambient tool library. |
|
|
242
244
|
| [loadAgentsFromDir](./docs/api/appkit/Function.loadAgentsFromDir.md) | Scans a directory for one subdirectory per agent, each containing `agent.md` (frontmatter + body). Produces an `AgentDefinition` record keyed by agent id (folder name). Throws on frontmatter errors or unresolved references. Returns an empty map if the directory does not exist. |
|
|
245
|
+
| [loadRootEvalConfig](./docs/api/appkit/Function.loadRootEvalConfig.md) | Load the root `evals.config.ts` under `rootDir` (the project root), or return `undefined` when there is none. This is the run-wide config carrying `baseUrl`/`webServer`; the CLI reads it to resolve options and manage the app-under-test lifecycle before calling [runEvalsInDir](./docs/api/appkit/Function.runEvalsInDir.md). |
|
|
243
246
|
| [matches](./docs/api/appkit/Function.matches.md) | Passes when the value matches `pattern`. |
|
|
244
247
|
| [mcpServer](./docs/api/appkit/Function.mcpServer.md) | Factory for declaring a custom MCP server tool. |
|
|
245
248
|
| [normalizeHost](./docs/api/appkit/Function.normalizeHost.md) | Ensure the host has a scheme (Databricks env often lacks `https://`). |
|
package/docs/plugins/agents.md
CHANGED
|
@@ -73,7 +73,7 @@ Migrating from `config/agents/`
|
|
|
73
73
|
|
|
74
74
|
Earlier versions kept markdown agents under `config/agents/<id>/agent.md`. That location is still read as a deprecated fallback (one-time warning on boot); move each folder to `server/agents/<id>/agent.md` so every agent — markdown and code — lives in one place.
|
|
75
75
|
|
|
76
|
-
Requests land at `POST /invocations` (or its alias `POST /responses`) with an OpenAI Responses-compatible body. These endpoints run the agent to completion and return a single JSON response — no SSE. Streaming clients should use `POST /chat`. Every tool call
|
|
76
|
+
Requests land at `POST /invocations` (or its alias `POST /responses`) with an OpenAI Responses-compatible body. These endpoints run the agent to completion and return a single JSON response — no SSE. Streaming clients should use `POST /chat`. Every tool call is traced automatically. Plugin-toolkit tool calls (the `plugin:<name>` entries / `plugins.<name>.toolkit()`) additionally run through `asUser(req)`, so their SQL executes as the requesting user and file access respects Unity Catalog ACLs. A hand-rolled `tool({ execute })` is **not** wrapped: its `execute` receives only the validated tool arguments (no `req`), so it runs with the app's service-principal identity and cannot opt into OBO. If a tool must act as the requesting user, expose it as a plugin tool rather than a hand-rolled `execute`. See [Execution context](./docs/plugins/execution-context.md).
|
|
77
77
|
|
|
78
78
|
No HITL on `/invocations` and `/responses`
|
|
79
79
|
|
|
@@ -218,7 +218,7 @@ Skills are on-demand instruction packs — the same `SKILL.md` format Claude Cod
|
|
|
218
218
|
A skill is a directory with a `SKILL.md` plus any bundled reference files:
|
|
219
219
|
|
|
220
220
|
```text
|
|
221
|
-
|
|
221
|
+
server/agents/
|
|
222
222
|
skills/ # shared pool — any agent can opt in
|
|
223
223
|
pdf-forms/
|
|
224
224
|
SKILL.md
|
|
@@ -248,8 +248,8 @@ To fill a PDF form:
|
|
|
248
248
|
|
|
249
249
|
### Visibility[](#visibility "Direct link to Visibility")
|
|
250
250
|
|
|
251
|
-
* **Per-agent skills** (`
|
|
252
|
-
* **Global skills** (`
|
|
251
|
+
* **Per-agent skills** (`server/agents/<id>/skills/`) are always visible to that agent.
|
|
252
|
+
* **Global skills** (`server/agents/skills/`, and catalog-volume skills) are **opt-in**: list them in the agent's frontmatter, `skills: [pdf-forms]`. Set `autoInheritSkills: true` (or `{ file, code }`) on the plugin to make every global skill visible without listing — off by default so each agent's always-on catalog stays lean.
|
|
253
253
|
|
|
254
254
|
### How the agent uses a skill[](#how-the-agent-uses-a-skill "Direct link to How the agent uses a skill")
|
|
255
255
|
|
|
@@ -687,6 +687,257 @@ appkit.agents.getThreads(userId); // list user's threads
|
|
|
687
687
|
|
|
688
688
|
```
|
|
689
689
|
|
|
690
|
+
## Evaluating agents[](#evaluating-agents "Direct link to Evaluating agents")
|
|
691
|
+
|
|
692
|
+
AppKit ships an eval framework for the agents you build here. You author evals in TypeScript with `defineEval`, drive the agent by sending it messages, and assert on its reply and tool usage with deterministic matchers or LLM judges. Evals run against a **running app** over HTTP (`--url`), and — with Databricks creds and an experiment — report to MLflow as native "Evaluation runs" with per-assertion and per-judge feedback attached to each turn's trace. The eval API is part of the beta surface: import it from `@databricks/appkit/beta`.
|
|
693
|
+
|
|
694
|
+
Evals live beside each agent: `server/agents/<agent-id>/evals/*.eval.ts`. Each file default-exports one `defineEval({ test })`. The agent under test defaults to the parent `<agent-id>` directory; set `agent:` to target a different one.
|
|
695
|
+
|
|
696
|
+
### A first eval[](#a-first-eval "Direct link to A first eval")
|
|
697
|
+
|
|
698
|
+
```ts
|
|
699
|
+
// server/agents/query/evals/smoke.eval.ts
|
|
700
|
+
import { defineEval } from "@databricks/appkit/beta";
|
|
701
|
+
|
|
702
|
+
export default defineEval({
|
|
703
|
+
description: "Query agent responds to a greeting",
|
|
704
|
+
async test(t) {
|
|
705
|
+
await t.send("Hi there!");
|
|
706
|
+
t.succeeded(); // gate: the turn completed without an agent/stream error
|
|
707
|
+
},
|
|
708
|
+
});
|
|
709
|
+
|
|
710
|
+
```
|
|
711
|
+
|
|
712
|
+
Start the app, then run the evals against it:
|
|
713
|
+
|
|
714
|
+
```bash
|
|
715
|
+
# the app must be running and reachable at --url
|
|
716
|
+
appkit agent eval --url http://localhost:3000
|
|
717
|
+
|
|
718
|
+
# scope to one agent/eval by substring, and point at a project root
|
|
719
|
+
appkit agent eval query --root apps/dev-playground --url http://localhost:3000
|
|
720
|
+
|
|
721
|
+
```
|
|
722
|
+
|
|
723
|
+
The positional `[filter]` matches evals whose `<agent>/<id>` contains the substring (or an exact agent id). The command discovers every `*.eval.ts` under `server/agents/*/evals/`, drives each against the running app, and exits non-zero if any gate fails.
|
|
724
|
+
|
|
725
|
+
### Assertions[](#assertions "Direct link to Assertions")
|
|
726
|
+
|
|
727
|
+
Every assertion returns a chainable handle. Assertions are **gates by default** — a failure fails the eval (non-zero exit). Chain `.soft()` to demote to a tracked-only metric, `.gate()` to promote a soft assertion back, or `.atLeast(n)` to set the pass threshold on a scored assertion.
|
|
728
|
+
|
|
729
|
+
| Assertion | Passes when |
|
|
730
|
+
| ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ |
|
|
731
|
+
| `t.succeeded()` | The last turn completed without an agent/stream error. |
|
|
732
|
+
| `t.calledTool(name)` | The agent called `name` during the run. |
|
|
733
|
+
| `t.calledToolWith(name, expected)` | `name` was called with arguments that deep-contain `expected` (every key in `expected` matches recursively; extra args are ignored). |
|
|
734
|
+
| `t.check(value, matcher)` | `value` satisfies the matcher — `includes(substring)`, `equals(expected)`, or `matches(pattern)`. |
|
|
735
|
+
|
|
736
|
+
```ts
|
|
737
|
+
import { defineEval, includes } from "@databricks/appkit/beta";
|
|
738
|
+
|
|
739
|
+
export default defineEval({
|
|
740
|
+
description: "Helper agent answers a math question",
|
|
741
|
+
agent: "helper",
|
|
742
|
+
async test(t) {
|
|
743
|
+
await t.send("What is 2 + 2?");
|
|
744
|
+
t.succeeded(); // gate
|
|
745
|
+
t.check(t.reply, includes("4")).soft(); // tracked metric, won't fail the gate
|
|
746
|
+
},
|
|
747
|
+
});
|
|
748
|
+
|
|
749
|
+
```
|
|
750
|
+
|
|
751
|
+
```ts
|
|
752
|
+
// deep-partial tool-arg check
|
|
753
|
+
await t.send("What's the weather in Brooklyn?");
|
|
754
|
+
t.calledTool("get_weather");
|
|
755
|
+
t.calledToolWith("get_weather", { city: "Brooklyn" });
|
|
756
|
+
|
|
757
|
+
```
|
|
758
|
+
|
|
759
|
+
Call `t.skip("reason")` to skip an eval, and read `t.reply`, `t.toolCalls`, and `t.sessionId` to inspect the last turn.
|
|
760
|
+
|
|
761
|
+
### LLM-as-judge[](#llm-as-judge "Direct link to LLM-as-judge")
|
|
762
|
+
|
|
763
|
+
`t.judge.*` scores the last reply with an LLM judge (via `autoevals` pointed at a Databricks serving endpoint). Each judge returns a scored assertion (0..1) that **gates by default** — a miss fails the eval. Chain `.atLeast(n)` to set the pass threshold, or `.soft()` to track it only. Judges require a judge model: pass `--judge-model <endpoint>` (or set `APPKIT_JUDGE_MODEL`) plus Databricks auth; without one, `t.judge.*` throws with a clear message.
|
|
764
|
+
|
|
765
|
+
```ts
|
|
766
|
+
async test(t) {
|
|
767
|
+
await t.send("What's the weather in Brooklyn?");
|
|
768
|
+
t.succeeded();
|
|
769
|
+
|
|
770
|
+
// closedQA needs no ground truth — it judges the reply against a question.
|
|
771
|
+
(await t.judge.closedQA(
|
|
772
|
+
"Does the response describe weather conditions for Brooklyn?",
|
|
773
|
+
)).atLeast(0.5);
|
|
774
|
+
}
|
|
775
|
+
|
|
776
|
+
```
|
|
777
|
+
|
|
778
|
+
* `t.judge.factuality(expected)` — score the reply against an expected reference answer.
|
|
779
|
+
* `t.judge.closedQA(criteria)` — score whether the reply answers the question, per `criteria`.
|
|
780
|
+
* `t.judge.custom(spec)` — a prompt-template judge (`{ name, promptTemplate, choiceScores }`), the TS analog of MLflow's `@scorer`.
|
|
781
|
+
|
|
782
|
+
Guard judge calls with `isJudgeConfigured()` when an eval should still exercise the drive path without a judge model configured:
|
|
783
|
+
|
|
784
|
+
```ts
|
|
785
|
+
import { defineEval, isJudgeConfigured } from "@databricks/appkit/beta";
|
|
786
|
+
// ...
|
|
787
|
+
if (isJudgeConfigured()) {
|
|
788
|
+
(await t.judge.closedQA(guideline)).atLeast(0.5);
|
|
789
|
+
}
|
|
790
|
+
|
|
791
|
+
```
|
|
792
|
+
|
|
793
|
+
### Conversations[](#conversations "Direct link to Conversations")
|
|
794
|
+
|
|
795
|
+
Each `t.send` is one user turn. How you sequence them controls the thread:
|
|
796
|
+
|
|
797
|
+
```ts
|
|
798
|
+
// One-shot: a single turn.
|
|
799
|
+
await t.send("Summarize Q3 revenue.");
|
|
800
|
+
t.succeeded();
|
|
801
|
+
|
|
802
|
+
// Multi-turn: consecutive sends share one thread, so the agent sees history.
|
|
803
|
+
await t.send("Show me the orders table.");
|
|
804
|
+
await t.send("Now filter it to last week.");
|
|
805
|
+
t.succeeded();
|
|
806
|
+
|
|
807
|
+
// t.reset() drops the conversation: the next send opens a fresh thread with
|
|
808
|
+
// no history. Use it to run several independent one-shot checks in one test.
|
|
809
|
+
await t.send("What's 2 + 2?");
|
|
810
|
+
t.check(t.reply, includes("4"));
|
|
811
|
+
t.reset();
|
|
812
|
+
await t.send("What's the capital of France?");
|
|
813
|
+
t.check(t.reply, includes("Paris"));
|
|
814
|
+
|
|
815
|
+
```
|
|
816
|
+
|
|
817
|
+
### Datasets[](#datasets "Direct link to Datasets")
|
|
818
|
+
|
|
819
|
+
Add `dataset: { table }` to sweep a Databricks **managed evaluation dataset** — a Unity Catalog `catalog.schema.table` with `inputs`/`expectations` columns. The eval runs once per row; the runner binds each row's `inputs` to `t.input` and `expectations` to `t.expected`. Reading the dataset requires a workspace client and warehouse (`--warehouse-id` + auth).
|
|
820
|
+
|
|
821
|
+
```ts
|
|
822
|
+
import { defineEval, isJudgeConfigured, userTurns } from "@databricks/appkit/beta";
|
|
823
|
+
|
|
824
|
+
export default defineEval({
|
|
825
|
+
description: "Query agent satisfies each dataset row's guidelines",
|
|
826
|
+
dataset: { table: "main.mario.appkit_eval_dataset" }, // optional `limit?`
|
|
827
|
+
async test(t) {
|
|
828
|
+
// Replay every user turn in the row against one thread, so the agent sees
|
|
829
|
+
// the accumulating conversation. A single-user-turn row sends once.
|
|
830
|
+
for (const turn of userTurns(t.input)) {
|
|
831
|
+
await t.send(turn);
|
|
832
|
+
}
|
|
833
|
+
t.succeeded();
|
|
834
|
+
|
|
835
|
+
if (isJudgeConfigured()) {
|
|
836
|
+
for (const guideline of guidelines(t.expected)) {
|
|
837
|
+
(await t.judge.closedQA(guideline)).atLeast(0.5);
|
|
838
|
+
}
|
|
839
|
+
}
|
|
840
|
+
},
|
|
841
|
+
});
|
|
842
|
+
|
|
843
|
+
```
|
|
844
|
+
|
|
845
|
+
The row shapes match the MLflow managed-dataset UI:
|
|
846
|
+
|
|
847
|
+
```text
|
|
848
|
+
inputs {"messages":[{"role":"user","content":"..."}]}
|
|
849
|
+
expectations {"guidelines":{"value":["...","..."]}} (optional)
|
|
850
|
+
|
|
851
|
+
```
|
|
852
|
+
|
|
853
|
+
`userTurns(t.input)` extracts every `role: "user"` message content in order from the `{messages:[...]}` input — a row can carry a full multi-turn conversation, and replaying each user turn against one thread lets the agent build up history (interleaved assistant/system turns are ignored; the agent generates its own). Read `expectations.guidelines.value` yourself; the UI wraps the array as `{value: [...]}`:
|
|
854
|
+
|
|
855
|
+
```ts
|
|
856
|
+
function guidelines(expected: Record<string, unknown> | undefined): string[] {
|
|
857
|
+
const g = (expected?.guidelines as { value?: unknown } | undefined)?.value;
|
|
858
|
+
return Array.isArray(g) ? g.map(String) : [];
|
|
859
|
+
}
|
|
860
|
+
|
|
861
|
+
```
|
|
862
|
+
|
|
863
|
+
Run a dataset eval:
|
|
864
|
+
|
|
865
|
+
```bash
|
|
866
|
+
appkit agent eval dataset --root apps/dev-playground --url http://localhost:3000 \
|
|
867
|
+
--profile <profile> --warehouse-id <warehouse-id> --judge-model <endpoint>
|
|
868
|
+
|
|
869
|
+
```
|
|
870
|
+
|
|
871
|
+
### Running evals & CI[](#running-evals--ci "Direct link to Running evals & CI")
|
|
872
|
+
|
|
873
|
+
`appkit agent eval [filter]` — run agent evals (`server/agents/<id>/evals/*.eval.ts`) against a running app.
|
|
874
|
+
|
|
875
|
+
| Flag | Description |
|
|
876
|
+
| ---------------------------- | --------------------------------------------------------------------------------------------------------------- |
|
|
877
|
+
| `[filter]` | Only run evals whose `<agent>/<id>` contains this substring (or an exact agent id) |
|
|
878
|
+
| `--url <url>` | Base URL of the running app (default `http://localhost:3000`) |
|
|
879
|
+
| `--strict` | Fail on soft-assertion misses too |
|
|
880
|
+
| `--root <dir>` | Project root containing `server/agents/` (default: cwd) |
|
|
881
|
+
| `--header <header...>` | Extra request header as `'Key: value'` (repeatable) |
|
|
882
|
+
| `--tag <tag...>` | Only run evals tagged with one of these tags (repeatable) |
|
|
883
|
+
| `--profile <name>` | Databricks CLI profile to authenticate with via OAuth (default: `DATABRICKS_CONFIG_PROFILE`) |
|
|
884
|
+
| `--databricks-host <host>` | Databricks host for writing MLflow assessments (default: `DATABRICKS_HOST`) |
|
|
885
|
+
| `--databricks-token <token>` | Databricks token for writing MLflow assessments (default: `DATABRICKS_TOKEN`) |
|
|
886
|
+
| `--experiment <id>` | MLflow experiment id for the evaluation run (default: `MLFLOW_EXPERIMENT_ID`) |
|
|
887
|
+
| `--warehouse-id <id>` | SQL warehouse id for reading managed evaluation datasets (default: `DATABRICKS_WAREHOUSE_ID`) |
|
|
888
|
+
| `--judge-model <endpoint>` | Databricks serving endpoint to use as the LLM judge for `t.judge.*` (default: `APPKIT_JUDGE_MODEL`) |
|
|
889
|
+
| `--concurrency <n>` | Max evals/dataset rows to drive concurrently (default: 4) |
|
|
890
|
+
| `--timeout <ms>` | Default per-eval timeout in ms (a per-eval `timeoutMs` overrides it) |
|
|
891
|
+
| `--retries <n>` | Re-run an eval up to N times when it fails on an infra error (turn/timeout); assertion failures are not retried |
|
|
892
|
+
| `--min-pass-rate <rate>` | Gate on aggregate pass rate (0..1) instead of requiring every eval to pass; exit 1 when below |
|
|
893
|
+
| `--reporter <format>` | Report format: `text` (live console), `json` (dashboards), or `junit` (CI test reporters) |
|
|
894
|
+
| `--output <file>` | Write the json/junit report to this file instead of stdout (ignored for text) |
|
|
895
|
+
|
|
896
|
+
Notes:
|
|
897
|
+
|
|
898
|
+
* **Auth is OAuth-first.** `--profile <name>` mints an OAuth token from your Databricks CLI profile — no PAT needed. An explicit `--databricks-host`/`--databricks-token` (or the `DATABRICKS_*` env vars) wins over the profile.
|
|
899
|
+
* **`--retries` only absorbs infra flakiness.** A retry fires only when an eval throws or times out (`result.error` set); a wrong reply is real signal and is never retried. Each attempt gets a fresh driver.
|
|
900
|
+
* **Gating.** By default the run exits non-zero if any eval fails. `--min-pass-rate 0.9` switches to threshold mode: it exits non-zero only when the aggregate pass rate drops below the threshold. `--strict` additionally treats soft-assertion misses as failures.
|
|
901
|
+
* **CI reports.** `--reporter junit --output results.xml` writes a JUnit file for CI test reporters; `--reporter json` emits machine-readable results. In machine reporters, human-facing lines go to stderr so stdout stays clean for the report.
|
|
902
|
+
|
|
903
|
+
```bash
|
|
904
|
+
# CI: gate on 90% pass rate, emit JUnit, authenticate + report to MLflow
|
|
905
|
+
appkit agent eval --url "$APP_URL" \
|
|
906
|
+
--profile ci \
|
|
907
|
+
--experiment "$MLFLOW_EXPERIMENT_ID" \
|
|
908
|
+
--concurrency 4 --retries 1 \
|
|
909
|
+
--min-pass-rate 0.9 \
|
|
910
|
+
--reporter junit --output eval-results.xml
|
|
911
|
+
|
|
912
|
+
```
|
|
913
|
+
|
|
914
|
+
### Per-directory config[](#per-directory-config "Direct link to Per-directory config")
|
|
915
|
+
|
|
916
|
+
Drop an `evals.config.ts` beside an agent's evals to set defaults for that agent's runs:
|
|
917
|
+
|
|
918
|
+
```ts
|
|
919
|
+
// server/agents/query/evals/evals.config.ts
|
|
920
|
+
import { defineEvalConfig } from "@databricks/appkit/beta";
|
|
921
|
+
|
|
922
|
+
export default defineEvalConfig({
|
|
923
|
+
maxConcurrency: 4, // run up to 4 evals/rows concurrently
|
|
924
|
+
timeoutMs: 30_000, // default per-eval timeout
|
|
925
|
+
});
|
|
926
|
+
|
|
927
|
+
```
|
|
928
|
+
|
|
929
|
+
Precedence: a **CLI flag** wins over the **`evals.config.ts`** value, which wins over the **built-in default** (concurrency `4`, no timeout). A per-eval `def.timeoutMs` overrides both for that eval.
|
|
930
|
+
|
|
931
|
+
### MLflow reporting[](#mlflow-reporting "Direct link to MLflow reporting")
|
|
932
|
+
|
|
933
|
+
When `--experiment <id>` (or `MLFLOW_EXPERIMENT_ID`) is set together with Databricks auth, the runner creates a native MLflow **Evaluation run** up front. As each eval runs against the app, its turn trace is linked to the run, and every assertion and judge score is written back as feedback on that trace:
|
|
934
|
+
|
|
935
|
+
* Each deterministic assertion becomes a `CODE`-sourced Feedback with a boolean value.
|
|
936
|
+
* Each judge assertion becomes an `LLM_JUDGE`-sourced Feedback with its numeric 0..1 score and rationale.
|
|
937
|
+
* An overall `appkit_eval` pass/fail Feedback is attached per eval, and aggregate metrics are logged when the run finishes.
|
|
938
|
+
|
|
939
|
+
Skip `--experiment` (and `MLFLOW_EXPERIMENT_ID`) to run evals purely locally with no MLflow side effects — the CLI prints a reminder that the evaluation run was skipped.
|
|
940
|
+
|
|
690
941
|
## Frontmatter schema[](#frontmatter-schema "Direct link to Frontmatter schema")
|
|
691
942
|
|
|
692
943
|
| Key | Type | Notes |
|
package/llms.txt
CHANGED
|
@@ -110,6 +110,7 @@ npx @databricks/appkit docs <query>
|
|
|
110
110
|
- [Function: evalGlyph()](./docs/api/appkit/Function.evalGlyph.md): Status glyph for a single eval result.
|
|
111
111
|
- [Function: executeFromRegistry()](./docs/api/appkit/Function.executeFromRegistry.md): Validates tool-call arguments against the entry's schema and invokes its
|
|
112
112
|
- [Function: extractServingEndpoints()](./docs/api/appkit/Function.extractServingEndpoints.md): Extract serving endpoint config from a server file by AST-parsing it.
|
|
113
|
+
- [Function: findRootEvalConfig()](./docs/api/appkit/Function.findRootEvalConfig.md): Path to the root evals.config.ts (from defineEvalConfig) at
|
|
113
114
|
- [Function: findServerFile()](./docs/api/appkit/Function.findServerFile.md): Find the server entry file by checking candidate paths in order.
|
|
114
115
|
- [Function: fk()](./docs/api/appkit/Function.fk.md): Declare foreign-key to another column.
|
|
115
116
|
- [Function: formatEvalDetail()](./docs/api/appkit/Function.formatEvalDetail.md): Indented detail lines for a failing eval (error + failing assertions).
|
|
@@ -140,6 +141,7 @@ npx @databricks/appkit docs <query>
|
|
|
140
141
|
- [Function: jsonb()](./docs/api/appkit/Function.jsonb.md): Returns
|
|
141
142
|
- [Function: loadAgentFromFile()](./docs/api/appkit/Function.loadAgentFromFile.md): Loads a single markdown agent file and resolves its frontmatter against
|
|
142
143
|
- [Function: loadAgentsFromDir()](./docs/api/appkit/Function.loadAgentsFromDir.md): Scans a directory for one subdirectory per agent, each containing
|
|
144
|
+
- [Function: loadRootEvalConfig()](./docs/api/appkit/Function.loadRootEvalConfig.md): Load the root evals.config.ts under rootDir (the project root), or return
|
|
143
145
|
- [Function: matches()](./docs/api/appkit/Function.matches.md): Passes when the value matches pattern.
|
|
144
146
|
- [Function: mcpServer()](./docs/api/appkit/Function.mcpServer.md): Factory for declaring a custom MCP server tool.
|
|
145
147
|
- [Function: normalizeHost()](./docs/api/appkit/Function.normalizeHost.md): Ensure the host has a scheme (Databricks env often lacks https://).
|
|
@@ -189,6 +191,7 @@ npx @databricks/appkit docs <query>
|
|
|
189
191
|
- [Interface: EvalResult](./docs/api/appkit/Interface.EvalResult.md): The outcome of running one eval.
|
|
190
192
|
- [Interface: EvalRunSummary](./docs/api/appkit/Interface.EvalRunSummary.md): Properties
|
|
191
193
|
- [Interface: EvalSummary](./docs/api/appkit/Interface.EvalSummary.md): Properties
|
|
194
|
+
- [Interface: EvalWebServer](./docs/api/appkit/Interface.EvalWebServer.md): Auto-start config for the app under test, à la Playwright's webServer. When
|
|
192
195
|
- [Interface: FilePolicyUser](./docs/api/appkit/Interface.FilePolicyUser.md): Minimal user identity passed to the policy function.
|
|
193
196
|
- [Interface: FileResource](./docs/api/appkit/Interface.FileResource.md): Describes the file or directory being acted upon.
|
|
194
197
|
- [Interface: FunctionTool](./docs/api/appkit/Interface.FunctionTool.md): Properties
|