@cursor/july 0.1.108 → 0.1.109
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +2 -6
- package/README.md +3 -6
- package/dist/channels/bitbucket/api.d.ts +41 -0
- package/dist/channels/bitbucket/api.d.ts.map +1 -1
- package/dist/channels/bitbucket/api.js +260 -0
- package/dist/channels/bitbucket/binding.d.ts +4 -0
- package/dist/channels/bitbucket/binding.d.ts.map +1 -1
- package/dist/channels/bitbucket/binding.js +16 -0
- package/dist/channels/bitbucket/index.d.ts +1 -1
- package/dist/channels/bitbucket/index.d.ts.map +1 -1
- package/dist/channels/bitbucket/index.js +1 -1
- package/dist/channels/gitlab/api.d.ts +27 -0
- package/dist/channels/gitlab/api.d.ts.map +1 -1
- package/dist/channels/gitlab/api.js +88 -0
- package/dist/channels/gitlab/binding.d.ts +5 -0
- package/dist/channels/gitlab/binding.d.ts.map +1 -1
- package/dist/channels/gitlab/binding.js +10 -0
- package/dist/channels/gitlab/index.d.ts +1 -1
- package/dist/channels/gitlab/index.d.ts.map +1 -1
- package/dist/channels/gitlab/index.js +1 -1
- package/dist/channels/slack/dispatch.d.ts +10 -0
- package/dist/channels/slack/dispatch.d.ts.map +1 -1
- package/dist/channels/slack/dispatch.js +20 -3
- package/dist/channels/slack/slack-channel.d.ts +12 -5
- package/dist/channels/slack/slack-channel.d.ts.map +1 -1
- package/dist/channels/slack/slack-channel.js +59 -8
- package/dist/docs/404.html +2 -2
- package/dist/docs/assets/{app.BASKM3M4.js → app.Cr-wVbnB.js} +1 -1
- package/dist/docs/assets/{building-with-agents.md.CUSWxlP_.js → building-with-agents.md.D0KbSkJn.js} +2 -2
- package/dist/docs/assets/{building-with-agents.md.CUSWxlP_.lean.js → building-with-agents.md.D0KbSkJn.lean.js} +1 -1
- package/dist/docs/assets/chunks/@localSearchIndexroot.CFVQ4S17.js +1 -0
- package/dist/docs/assets/chunks/{VPLocalSearchBox.BOKxlYGP.js → VPLocalSearchBox.CVQERt56.js} +1 -1
- package/dist/docs/assets/chunks/{theme.DbDZW-zb.js → theme.Dnsd3XOn.js} +2 -2
- package/dist/docs/assets/concepts.md.B4o63Gul.js +1 -0
- package/dist/docs/assets/evals.md.C7JLjoEP.js +211 -0
- package/dist/docs/assets/evals.md.C7JLjoEP.lean.js +1 -0
- package/dist/docs/assets/guides_bitbucket.md.DZPzmUjU.js +10 -0
- package/dist/docs/assets/guides_bitbucket.md.DZPzmUjU.lean.js +1 -0
- package/dist/docs/assets/{guides_github.md.BH33UEBJ.js → guides_github.md.TZaTZlfz.js} +13 -3
- package/dist/docs/assets/{guides_github.md.BH33UEBJ.lean.js → guides_github.md.TZaTZlfz.lean.js} +1 -1
- package/dist/docs/assets/guides_gitlab.md.BmqwQfdG.js +14 -0
- package/dist/docs/assets/guides_gitlab.md.BmqwQfdG.lean.js +1 -0
- package/dist/docs/assets/{guides_webhooks.md.DKdA43Qm.js → guides_webhooks.md.CJK484ex.js} +2 -2
- package/dist/docs/assets/{hillclimbing.md.CpTGTCle.js → hillclimbing.md.BOiVo1tf.js} +1 -1
- package/dist/docs/assets/index.md.DD9Q2XuJ.js +5 -0
- package/dist/docs/assets/{index.md.BjH1w2ZW.lean.js → index.md.DD9Q2XuJ.lean.js} +1 -1
- package/dist/docs/assets/{reference_channels.md.D-qTqwcq.js → reference_channels.md.CAo-iK4j.js} +2 -2
- package/dist/docs/assets/{reference_channels.md.D-qTqwcq.lean.js → reference_channels.md.CAo-iK4j.lean.js} +1 -1
- package/dist/docs/assets/{reference_cli.md.Ca26u0Es.js → reference_cli.md.Deg7849l.js} +1 -1
- package/dist/docs/assets/{reference_extensions.md.CPWt00ds.js → reference_extensions.md.ZAVUyuEX.js} +2 -2
- package/dist/docs/assets/{reference_hooks.md.Ddt5DdgJ.js → reference_hooks.md.BlM_bOg6.js} +3 -3
- package/dist/docs/assets/{reference_hooks.md.Ddt5DdgJ.lean.js → reference_hooks.md.BlM_bOg6.lean.js} +1 -1
- package/dist/docs/assets/{reference_http-api.md.oySXBO8o.js → reference_http-api.md.BwaCo-VO.js} +1 -1
- package/dist/docs/assets/{reference_playground.md.4myJPxrf.js → reference_playground.md.DLnoaczX.js} +1 -1
- package/dist/docs/assets/{reference_playground.md.4myJPxrf.lean.js → reference_playground.md.DLnoaczX.lean.js} +1 -1
- package/dist/docs/assets/{reference_project-layout.md.DuBu9a96.js → reference_project-layout.md.BEU8MtQV.js} +3 -3
- package/dist/docs/assets/{reference_project-layout.md.DuBu9a96.lean.js → reference_project-layout.md.BEU8MtQV.lean.js} +1 -1
- package/dist/docs/assets/reference_sessions.md.CyXV1MUw.js +1 -0
- package/dist/docs/assets/skills_deploy.md.CWqi_ZxW.js +35 -0
- package/dist/docs/assets/skills_deploy.md.CWqi_ZxW.lean.js +1 -0
- package/dist/docs/assets/{skills_evals.md.BhovOvrl.js → skills_evals.md.DFxYPErF.js} +3 -3
- package/dist/docs/assets/skills_evals.md.DFxYPErF.lean.js +1 -0
- package/dist/docs/assets/{skills_framework-map.md.D-tFZhFS.js → skills_framework-map.md.BxLSOhSY.js} +1 -1
- package/dist/docs/assets/skills_index.md.DL7EHaQ-.js +1 -0
- package/dist/docs/assets/skills_index.md.DL7EHaQ-.lean.js +1 -0
- package/dist/docs/assets/{storage.md.BOHeqk2M.js → storage.md.BUrhJ-Zz.js} +4 -4
- package/dist/docs/assets/{storage.md.BOHeqk2M.lean.js → storage.md.BUrhJ-Zz.lean.js} +1 -1
- package/dist/docs/building-with-agents.html +5 -5
- package/dist/docs/building-with-agents.md +3 -2
- package/dist/docs/concepts.html +5 -5
- package/dist/docs/concepts.md +2 -4
- package/dist/docs/deployment.html +4 -4
- package/dist/docs/evals.html +161 -35
- package/dist/docs/evals.md +612 -297
- package/dist/docs/guides/agent-to-agent.html +4 -4
- package/dist/docs/guides/bitbucket.html +36 -0
- package/dist/docs/guides/bitbucket.md +84 -0
- package/dist/docs/guides/cloud-agents.html +4 -4
- package/dist/docs/guides/convert-automation.html +4 -4
- package/dist/docs/guides/github.html +17 -7
- package/dist/docs/guides/github.md +31 -0
- package/dist/docs/guides/gitlab.html +40 -0
- package/dist/docs/guides/gitlab.md +92 -0
- package/dist/docs/guides/grokbot-agents.html +4 -4
- package/dist/docs/guides/human-in-the-loop.html +4 -4
- package/dist/docs/guides/improve.html +4 -4
- package/dist/docs/guides/mcp-oauth.html +4 -4
- package/dist/docs/guides/opentelemetry.html +4 -4
- package/dist/docs/guides/slack.html +5 -5
- package/dist/docs/guides/webhooks.html +6 -6
- package/dist/docs/guides/webhooks.md +4 -2
- package/dist/docs/hashmap.json +1 -1
- package/dist/docs/hillclimbing.html +6 -6
- package/dist/docs/hillclimbing.md +2 -2
- package/dist/docs/index.html +6 -6
- package/dist/docs/index.md +10 -6
- package/dist/docs/llms-full.txt +1113 -778
- package/dist/docs/llms.txt +6 -5
- package/dist/docs/quickstart.html +4 -4
- package/dist/docs/reference/agent-config.html +4 -4
- package/dist/docs/reference/artifacts.html +4 -4
- package/dist/docs/reference/channels.html +6 -6
- package/dist/docs/reference/channels.md +16 -3
- package/dist/docs/reference/cli.html +6 -6
- package/dist/docs/reference/cli.md +6 -4
- package/dist/docs/reference/connections.html +4 -4
- package/dist/docs/reference/extensions.html +7 -7
- package/dist/docs/reference/extensions.md +0 -3
- package/dist/docs/reference/hooks.html +7 -7
- package/dist/docs/reference/hooks.md +9 -12
- package/dist/docs/reference/http-api.html +6 -6
- package/dist/docs/reference/http-api.md +2 -3
- package/dist/docs/reference/instructions.html +4 -4
- package/dist/docs/reference/playground.html +5 -5
- package/dist/docs/reference/playground.md +0 -4
- package/dist/docs/reference/project-layout.html +7 -7
- package/dist/docs/reference/project-layout.md +1 -8
- package/dist/docs/reference/prompt.html +4 -4
- package/dist/docs/reference/result.html +4 -4
- package/dist/docs/reference/schedules.html +4 -4
- package/dist/docs/reference/sessions.html +5 -5
- package/dist/docs/reference/sessions.md +1 -2
- package/dist/docs/reference/skills.html +4 -4
- package/dist/docs/reference/subagents.html +4 -4
- package/dist/docs/reference/tools.html +4 -4
- package/dist/docs/scaffolding-agents.html +4 -4
- package/dist/docs/skills/create-agent.html +4 -4
- package/dist/docs/skills/debug.html +4 -4
- package/dist/docs/skills/deploy.html +61 -0
- package/dist/docs/skills/deploy.md +161 -0
- package/dist/docs/skills/evals.html +8 -8
- package/dist/docs/skills/evals.md +44 -6
- package/dist/docs/skills/framework-map.html +5 -5
- package/dist/docs/skills/framework-map.md +1 -2
- package/dist/docs/skills/github.html +4 -4
- package/dist/docs/skills/hillclimb.html +4 -4
- package/dist/docs/skills/index.html +6 -6
- package/dist/docs/skills/index.md +1 -1
- package/dist/docs/skills/mcp-auth.html +4 -4
- package/dist/docs/skills/otel.html +4 -4
- package/dist/docs/skills/setup-slack.html +4 -4
- package/dist/docs/storage.html +8 -8
- package/dist/docs/storage.md +15 -24
- package/dist/docs/templates/agentic-owners.html +4 -4
- package/dist/docs/templates/agents-md.html +4 -4
- package/dist/docs/templates/code-wiki.html +4 -4
- package/dist/docs/templates/demo.html +4 -4
- package/dist/docs/templates/grokbot-agents.html +4 -4
- package/dist/docs/templates/pr-autofixer.html +4 -4
- package/dist/docs/templates/security-help.html +4 -4
- package/dist/docs/templates/security-reviewer.html +4 -4
- package/dist/docs/templates/triage.html +4 -4
- package/dist/docs/troubleshooting.html +4 -4
- package/dist/extensions.d.ts +2 -3
- package/dist/extensions.d.ts.map +1 -1
- package/dist/extensions.js +2 -5
- package/dist/index.d.ts +2 -3
- package/dist/index.d.ts.map +1 -1
- package/dist/index.js +1 -2
- package/dist/internal/authored-alias-hooks.d.ts +5 -0
- package/dist/internal/authored-alias-hooks.d.ts.map +1 -1
- package/dist/internal/authored-alias-hooks.js +17 -0
- package/dist/internal/authored-loaders.d.ts +4 -0
- package/dist/internal/authored-loaders.d.ts.map +1 -1
- package/dist/internal/authored-loaders.js +21 -2
- package/dist/internal/cli-ax.d.ts +1 -2
- package/dist/internal/cli-ax.d.ts.map +1 -1
- package/dist/internal/cli-ax.js +1 -2
- package/dist/internal/continuation-channel.d.ts.map +1 -1
- package/dist/internal/continuation-channel.js +2 -2
- package/dist/internal/continuation-identity.js +8 -3
- package/dist/internal/discovery/agent.d.ts +1 -1
- package/dist/internal/discovery/agent.d.ts.map +1 -1
- package/dist/internal/discovery/agent.js +0 -5
- package/dist/internal/discovery/extension-overlay.d.ts +1 -2
- package/dist/internal/discovery/extension-overlay.d.ts.map +1 -1
- package/dist/internal/discovery/extension-overlay.js +0 -14
- package/dist/internal/discovery/extensions.d.ts +1 -2
- package/dist/internal/discovery/extensions.d.ts.map +1 -1
- package/dist/internal/discovery/extensions.js +4 -22
- package/dist/internal/discovery/info.d.ts.map +1 -1
- package/dist/internal/discovery/info.js +42 -24
- package/dist/internal/discovery/modules.js +0 -1
- package/dist/internal/discovery/project.d.ts.map +1 -1
- package/dist/internal/discovery/project.js +0 -16
- package/dist/internal/eval-runner.js +0 -1
- package/dist/internal/framework-file-storage.d.ts +4 -5
- package/dist/internal/framework-file-storage.d.ts.map +1 -1
- package/dist/internal/framework-file-storage.js +4 -5
- package/dist/internal/framework-storage-selection.d.ts +2 -2
- package/dist/internal/framework-storage-selection.js +2 -2
- package/dist/internal/init-scaffold.d.ts.map +1 -1
- package/dist/internal/init-scaffold.js +1 -2
- package/dist/internal/install-cursor-skills.d.ts +5 -2
- package/dist/internal/install-cursor-skills.d.ts.map +1 -1
- package/dist/internal/install-cursor-skills.js +25 -5
- package/dist/internal/run-client.d.ts +1 -1
- package/dist/internal/sdk-runner.js +1 -1
- package/dist/internal/server.js +0 -14
- package/dist/internal/session-engine.d.ts +5 -35
- package/dist/internal/session-engine.d.ts.map +1 -1
- package/dist/internal/session-engine.js +17 -162
- package/dist/internal/storage-coordinator.d.ts +5 -19
- package/dist/internal/storage-coordinator.d.ts.map +1 -1
- package/dist/internal/storage-coordinator.js +3 -62
- package/dist/internal/storage-roles.d.ts +4 -9
- package/dist/internal/storage-roles.d.ts.map +1 -1
- package/dist/internal/storage-roles.js +2 -2
- package/dist/playground/assets/index-Bhxzrcf6.css +1 -0
- package/dist/playground/assets/index-CqLX5uF3.js +67 -0
- package/dist/playground/index.html +2 -2
- package/dist/storage-backends/cursor-hosted-v2.d.ts +4 -5
- package/dist/storage-backends/cursor-hosted-v2.d.ts.map +1 -1
- package/dist/storage-backends/cursor-hosted-v2.js +4 -5
- package/dist/storage-backends/cursor-hosted.d.ts +1 -2
- package/dist/storage-backends/cursor-hosted.d.ts.map +1 -1
- package/dist/storage-backends/cursor-hosted.js +0 -30
- package/dist/storage-backends/file-kv.d.ts +9 -12
- package/dist/storage-backends/file-kv.d.ts.map +1 -1
- package/dist/storage-backends/file-kv.js +11 -47
- package/dist/storage-protocol.d.ts +3 -11
- package/dist/storage-protocol.d.ts.map +1 -1
- package/dist/storage-protocol.js +3 -11
- package/dist/storage.d.ts +8 -36
- package/dist/storage.d.ts.map +1 -1
- package/dist/storage.js +8 -44
- package/dist/types.d.ts +20 -57
- package/dist/types.d.ts.map +1 -1
- package/docs/README.md +10 -6
- package/docs/building-with-agents.md +3 -2
- package/docs/concepts.md +2 -4
- package/docs/evals.md +613 -298
- package/docs/guides/bitbucket.md +89 -0
- package/docs/guides/github.md +31 -0
- package/docs/guides/gitlab.md +97 -0
- package/docs/guides/webhooks.md +4 -2
- package/docs/hillclimbing.md +2 -2
- package/docs/reference/channels.md +16 -3
- package/docs/reference/cli.md +6 -4
- package/docs/reference/extensions.md +0 -3
- package/docs/reference/hooks.md +9 -12
- package/docs/reference/http-api.md +2 -3
- package/docs/reference/playground.md +0 -4
- package/docs/reference/project-layout.md +1 -8
- package/docs/reference/sessions.md +1 -2
- package/docs/skills/index.md +2 -2
- package/docs/storage.md +15 -24
- package/package.json +1 -7
- package/skills/deploy/SKILL.md +169 -0
- package/skills/evals/SKILL.md +45 -8
- package/skills/framework-map/SKILL.md +1 -2
- package/src/channels/bitbucket/api.ts +341 -0
- package/src/channels/bitbucket/binding.ts +25 -0
- package/src/channels/bitbucket/index.ts +2 -0
- package/src/channels/gitlab/api.ts +123 -0
- package/src/channels/gitlab/binding.ts +12 -0
- package/src/channels/gitlab/index.ts +1 -0
- package/src/channels/slack/dispatch.ts +30 -0
- package/src/channels/slack/slack-channel.ts +69 -7
- package/src/extensions.ts +2 -6
- package/src/index.ts +0 -3
- package/src/internal/authored-alias-hooks.ts +32 -0
- package/src/internal/authored-loaders.ts +26 -2
- package/src/internal/cli-ax.ts +1 -2
- package/src/internal/continuation-channel.ts +2 -1
- package/src/internal/continuation-identity.ts +10 -2
- package/src/internal/discovery/agent.ts +1 -7
- package/src/internal/discovery/extension-overlay.ts +0 -18
- package/src/internal/discovery/extensions.ts +2 -26
- package/src/internal/discovery/info.ts +3 -13
- package/src/internal/discovery/modules.ts +0 -1
- package/src/internal/discovery/project.ts +0 -16
- package/src/internal/eval-runner.ts +0 -1
- package/src/internal/framework-file-storage.ts +4 -5
- package/src/internal/framework-storage-selection.ts +2 -2
- package/src/internal/init-scaffold.ts +1 -2
- package/src/internal/install-cursor-skills.ts +38 -5
- package/src/internal/run-client.ts +1 -1
- package/src/internal/sdk-runner.ts +1 -1
- package/src/internal/server.ts +0 -16
- package/src/internal/session-engine.ts +14 -202
- package/src/internal/storage-coordinator.ts +5 -76
- package/src/internal/storage-roles.ts +4 -9
- package/src/storage-backends/cursor-hosted-v2.ts +4 -7
- package/src/storage-backends/cursor-hosted.ts +0 -37
- package/src/storage-backends/file-kv.ts +10 -51
- package/src/storage-protocol.ts +3 -17
- package/src/storage.ts +10 -101
- package/src/types.ts +19 -56
- package/templates/demo/README.md +10 -6
- package/templates/demo/agent/channels/github.ts +2 -0
- package/templates/demo/agent/channels/queue.ts +6 -2
- package/templates/demo/agent/lib/repos.ts +5 -0
- package/templates/demo/init.json +25 -0
- package/dist/ab.d.ts +0 -209
- package/dist/ab.d.ts.map +0 -1
- package/dist/ab.js +0 -246
- package/dist/docs/ab.html +0 -80
- package/dist/docs/ab.md +0 -332
- package/dist/docs/assets/ab.md.mlVgqvSk.js +0 -54
- package/dist/docs/assets/ab.md.mlVgqvSk.lean.js +0 -1
- package/dist/docs/assets/chunks/@localSearchIndexroot.BHZYsNVi.js +0 -1
- package/dist/docs/assets/concepts.md.DgEcZOfT.js +0 -1
- package/dist/docs/assets/evals.md.C1ekS3k2.js +0 -85
- package/dist/docs/assets/evals.md.C1ekS3k2.lean.js +0 -1
- package/dist/docs/assets/index.md.BjH1w2ZW.js +0 -5
- package/dist/docs/assets/reference_sessions.md.CueyOHSL.js +0 -1
- package/dist/docs/assets/skills_ab.md.CsFNatVx.js +0 -26
- package/dist/docs/assets/skills_ab.md.CsFNatVx.lean.js +0 -1
- package/dist/docs/assets/skills_evals.md.BhovOvrl.lean.js +0 -1
- package/dist/docs/assets/skills_index.md.DKwIxzGg.js +0 -1
- package/dist/docs/assets/skills_index.md.DKwIxzGg.lean.js +0 -1
- package/dist/docs/skills/ab.html +0 -52
- package/dist/docs/skills/ab.md +0 -50
- package/dist/internal/ab-collector.d.ts +0 -44
- package/dist/internal/ab-collector.d.ts.map +0 -1
- package/dist/internal/ab-collector.js +0 -142
- package/dist/internal/ab-fold.d.ts +0 -36
- package/dist/internal/ab-fold.d.ts.map +0 -1
- package/dist/internal/ab-fold.js +0 -175
- package/dist/internal/ab-snapshot.d.ts +0 -68
- package/dist/internal/ab-snapshot.d.ts.map +0 -1
- package/dist/internal/ab-snapshot.js +0 -208
- package/dist/internal/discovery/ab.d.ts +0 -9
- package/dist/internal/discovery/ab.d.ts.map +0 -1
- package/dist/internal/discovery/ab.js +0 -113
- package/dist/playground/assets/index-BMDqAeXu.js +0 -67
- package/dist/playground/assets/index-LUgJoWdL.css +0 -1
- package/docs/ab.md +0 -337
- package/skills/ab/SKILL.md +0 -58
- package/src/ab.ts +0 -430
- package/src/internal/ab-collector.ts +0 -200
- package/src/internal/ab-fold.ts +0 -232
- package/src/internal/ab-snapshot.ts +0 -331
- package/src/internal/discovery/ab.ts +0 -131
- /package/dist/docs/assets/{concepts.md.DgEcZOfT.lean.js → concepts.md.B4o63Gul.lean.js} +0 -0
- /package/dist/docs/assets/{guides_webhooks.md.DKdA43Qm.lean.js → guides_webhooks.md.CJK484ex.lean.js} +0 -0
- /package/dist/docs/assets/{hillclimbing.md.CpTGTCle.lean.js → hillclimbing.md.BOiVo1tf.lean.js} +0 -0
- /package/dist/docs/assets/{reference_cli.md.Ca26u0Es.lean.js → reference_cli.md.Deg7849l.lean.js} +0 -0
- /package/dist/docs/assets/{reference_extensions.md.CPWt00ds.lean.js → reference_extensions.md.ZAVUyuEX.lean.js} +0 -0
- /package/dist/docs/assets/{reference_http-api.md.oySXBO8o.lean.js → reference_http-api.md.BwaCo-VO.lean.js} +0 -0
- /package/dist/docs/assets/{reference_sessions.md.CueyOHSL.lean.js → reference_sessions.md.CyXV1MUw.lean.js} +0 -0
- /package/dist/docs/assets/{skills_framework-map.md.D-tFZhFS.lean.js → skills_framework-map.md.BxLSOhSY.lean.js} +0 -0
package/dist/docs/llms-full.txt
CHANGED
|
@@ -4,349 +4,13 @@
|
|
|
4
4
|
|
|
5
5
|
---
|
|
6
6
|
|
|
7
|
-
Source: /docs/ab.md
|
|
8
|
-
|
|
9
|
-
# Live A/B metrics
|
|
10
|
-
|
|
11
|
-
Use `defineAB` to compare variants on live agent sessions. New sessions
|
|
12
|
-
receive a sticky assignment in each enrolled experiment. The Agent SDK folds
|
|
13
|
-
their durable event streams into tool, token, failure, and wall-time
|
|
14
|
-
metrics. You can send cumulative samples to your metrics backend and
|
|
15
|
-
inspect aggregates in the playground.
|
|
16
|
-
|
|
17
|
-
`defineAB` compares live variants through sticky assignment,
|
|
18
|
-
instruction overlays, optional tool branches, and cumulative metrics.
|
|
19
|
-
Metric callbacks observe the result without approving, rejecting, or
|
|
20
|
-
failing a turn. Use [evals](/docs/evals.md) for pass/fail regression checks
|
|
21
|
-
on fixed inputs.
|
|
22
|
-
|
|
23
|
-
## Choose live A/B metrics or evals
|
|
24
|
-
|
|
25
|
-
Both features read the session event stream, but they answer different
|
|
26
|
-
questions.
|
|
27
|
-
|
|
28
|
-
| | Live A/B metrics | Evals |
|
|
29
|
-
| --- | --- | --- |
|
|
30
|
-
| Question | How do variants compare on live sessions? | Does the agent still meet a fixed contract? |
|
|
31
|
-
| Location | `agent/ab.ts` or `agent/ab/<name>.ts` | `evals/**/*.eval.ts` |
|
|
32
|
-
| Input | Dev or production traffic | Frozen prompts and fixtures |
|
|
33
|
-
| Output | Cumulative metrics by session and arm | Pass/fail assertions |
|
|
34
|
-
| How it runs | Automatically on new live sessions | `agent-sdk eval` |
|
|
35
|
-
|
|
36
|
-
There is no `agent-sdk ab` command or assertion API.
|
|
37
|
-
|
|
38
|
-
## Define an experiment
|
|
39
|
-
|
|
40
|
-
Author one experiment in `agent/ab.ts`, add more under
|
|
41
|
-
`agent/ab/<name>.ts`, or use both forms. Each file defines one
|
|
42
|
-
experiment. The experiment name comes from `name` when set. Otherwise,
|
|
43
|
-
the Agent SDK uses `ab` for `agent/ab.ts` and the file stem for files under
|
|
44
|
-
`agent/ab/`.
|
|
45
|
-
|
|
46
|
-
```ts
|
|
47
|
-
// agent/ab/concise-weather.ts
|
|
48
|
-
import {
|
|
49
|
-
defineAB,
|
|
50
|
-
splitBySessionHash,
|
|
51
|
-
} from "@cursor/july/ab";
|
|
52
|
-
|
|
53
|
-
export default defineAB({
|
|
54
|
-
name: "concise-weather",
|
|
55
|
-
variants: {
|
|
56
|
-
control: {
|
|
57
|
-
label: "Baseline",
|
|
58
|
-
},
|
|
59
|
-
treatment: {
|
|
60
|
-
label: "Short replies",
|
|
61
|
-
description: "Adds a one-paragraph response limit.",
|
|
62
|
-
instructions: "Keep weather replies to one short paragraph.",
|
|
63
|
-
},
|
|
64
|
-
},
|
|
65
|
-
split: splitBySessionHash({
|
|
66
|
-
weights: { control: 1, treatment: 1 },
|
|
67
|
-
holdout: 0.1,
|
|
68
|
-
}),
|
|
69
|
-
derive: {
|
|
70
|
-
weatherCalls: (event) =>
|
|
71
|
-
event.type === "action.result" &&
|
|
72
|
-
event.data.toolName === "get_weather"
|
|
73
|
-
? 1
|
|
74
|
-
: null,
|
|
75
|
-
},
|
|
76
|
-
onSample(sample) {
|
|
77
|
-
console.log(
|
|
78
|
-
sample.experiment,
|
|
79
|
-
sample.variant,
|
|
80
|
-
sample.metrics.toolCalls,
|
|
81
|
-
sample.metrics.wallTimeMs
|
|
82
|
-
);
|
|
83
|
-
},
|
|
84
|
-
});
|
|
85
|
-
```
|
|
86
|
-
|
|
87
|
-
Every definition needs:
|
|
88
|
-
|
|
89
|
-
- At least two variants. Variant keys cannot be empty or contain `/` or
|
|
90
|
-
`\`.
|
|
91
|
-
- A `split` function that returns a variant key or `null`.
|
|
92
|
-
- An `onSample` callback for completed or failed turns.
|
|
93
|
-
|
|
94
|
-
`label` and `description` appear with the arm in result surfaces.
|
|
95
|
-
`instructions` changes the prompt for sessions in that arm. `derive`
|
|
96
|
-
adds custom counters.
|
|
97
|
-
|
|
98
|
-
Duplicate experiment names are validation errors. Check discovery
|
|
99
|
-
before you serve:
|
|
100
|
-
|
|
101
|
-
```bash
|
|
102
|
-
agent-sdk validate --dir .
|
|
103
|
-
agent-sdk info --dir . --json
|
|
104
|
-
```
|
|
105
|
-
|
|
106
|
-
The `abs` field in `info` lists the discovered experiment names.
|
|
107
|
-
|
|
108
|
-
## Assign sticky variants
|
|
109
|
-
|
|
110
|
-
Enrollment happens once, when a live session is created and before its
|
|
111
|
-
first turn:
|
|
112
|
-
|
|
113
|
-
1. The Agent SDK records `session.started`.
|
|
114
|
-
2. Each experiment runs its `split` function.
|
|
115
|
-
3. The Agent SDK records one durable `ab.assigned` event per experiment.
|
|
116
|
-
4. The selected arms become available on `session.abs`.
|
|
117
|
-
5. Variant instruction overlays reach the first model turn.
|
|
118
|
-
|
|
119
|
-
A split can return a variant key or `null`. A null assignment is a
|
|
120
|
-
sticky skip for that experiment. It increments the experiment's
|
|
121
|
-
`skipped` total, still appears in the snapshot's `sessions` list with
|
|
122
|
-
`variant: null`, and does not collect arm metrics or call `onSample`.
|
|
123
|
-
|
|
124
|
-
Use the split helper that matches your rollout:
|
|
125
|
-
|
|
126
|
-
| Helper | Behavior |
|
|
127
|
-
| --- | --- |
|
|
128
|
-
| `splitBySessionHash({ weights?, holdout?, salt? })` | Hashes the session id into a reproducible arm; the recommended default |
|
|
129
|
-
| `splitByRandom({ weights?, holdout? })` | Draws once when the session starts, then persists the result |
|
|
130
|
-
| `splitAlways("control")` | Pins every new session to one arm |
|
|
131
|
-
| `splitNone()` | Skips every new session without deleting the experiment |
|
|
132
|
-
| `splitIf(predicate, inner)` | Runs `inner` only when the predicate passes |
|
|
133
|
-
| Custom `split(ctx)` | Returns a declared variant key or `null` |
|
|
134
|
-
|
|
135
|
-
The split context includes the agent name, channel id, session info,
|
|
136
|
-
experiment name, and declared variant keys. For example, enroll only
|
|
137
|
-
Slack sessions:
|
|
138
|
-
|
|
139
|
-
```ts
|
|
140
|
-
split: splitIf(
|
|
141
|
-
(ctx) => ctx.channel.id === "slack",
|
|
142
|
-
splitBySessionHash()
|
|
143
|
-
),
|
|
144
|
-
```
|
|
145
|
-
|
|
146
|
-
Weights default to equal. Non-positive weights leave an arm out of the
|
|
147
|
-
draw, and at least one arm must have a positive weight. `holdout` is the
|
|
148
|
-
fraction of sessions assigned `null`, from `0` through `1`. Change
|
|
149
|
-
`salt` to reshuffle future hash assignments without renaming the
|
|
150
|
-
experiment.
|
|
151
|
-
|
|
152
|
-
If a custom split throws or returns an unknown variant, the Agent SDK logs
|
|
153
|
-
the error and records `variant: null`. The failed decision becomes a
|
|
154
|
-
sticky skip instead of breaking the session.
|
|
155
|
-
|
|
156
|
-
Enrollment only applies to new sessions. Adding an experiment does not
|
|
157
|
-
assign existing conversations. Follow-ups keep the session's original
|
|
158
|
-
arms. Keep experiment names and variant keys stable while you collect
|
|
159
|
-
and compare results.
|
|
160
|
-
|
|
161
|
-
## Change behavior by variant
|
|
162
|
-
|
|
163
|
-
Variant instructions are appended to the agent's base instructions.
|
|
164
|
-
Local sessions receive the merged instructions in `AGENTS.md` before
|
|
165
|
-
every turn. Cloud sessions receive them in the first-turn preamble
|
|
166
|
-
only. For cloud follow-ups, branch through `session.abs` when the arm
|
|
167
|
-
must remain visible to deterministic behavior.
|
|
168
|
-
|
|
169
|
-
Tools can branch on the assignment through `ctx.session.abs`. Hooks can
|
|
170
|
-
read the same field for logging or export:
|
|
171
|
-
|
|
172
|
-
```ts
|
|
173
|
-
const treatment =
|
|
174
|
-
ctx.session.abs?.["concise-weather"] === "treatment";
|
|
175
|
-
|
|
176
|
-
if (treatment) {
|
|
177
|
-
return conciseWeatherResult;
|
|
178
|
-
}
|
|
179
|
-
|
|
180
|
-
return baselineWeatherResult;
|
|
181
|
-
```
|
|
182
|
-
|
|
183
|
-
This makes the assignment available to deterministic code as well as
|
|
184
|
-
the model prompt. Use both patterns together when one experiment must
|
|
185
|
-
steer the prompt and host code at once.
|
|
186
|
-
|
|
187
|
-
`defineAB` does not select a different model or runtime for each arm.
|
|
188
|
-
Keep those settings in `agent/agent.ts`, or write explicit host logic
|
|
189
|
-
when your experiment needs another behavior lever.
|
|
190
|
-
|
|
191
|
-
The split and selected arm can affect agent behavior. `derive` and
|
|
192
|
-
`onSample` only observe the resulting event stream. Errors in either
|
|
193
|
-
callback are logged and never fail the turn.
|
|
194
|
-
|
|
195
|
-
## Collect built-in and custom metrics
|
|
196
|
-
|
|
197
|
-
Metrics accumulate for each session and experiment. When one session
|
|
198
|
-
joins several experiments, every enrolled experiment folds the same
|
|
199
|
-
turn and tool events into its own counters.
|
|
200
|
-
|
|
201
|
-
| Metric | How the Agent SDK calculates it |
|
|
202
|
-
| --- | --- |
|
|
203
|
-
| `turns` | Adds one on `turn.completed` or `turn.failed` |
|
|
204
|
-
| `turnFailures` | Adds one on `turn.failed` |
|
|
205
|
-
| `toolCalls` | Adds one for each `action.result` |
|
|
206
|
-
| `toolErrors` | Adds one when `action.result.data.isError` is true |
|
|
207
|
-
| `inputTokens`, `outputTokens` | Adds usage from completed turns |
|
|
208
|
-
| `cacheReadTokens`, `cacheWriteTokens` | Adds cache usage from completed turns |
|
|
209
|
-
| `costUsd` | Sums the estimated turn cost recorded on `turn.completed` (turns whose model has no known rates contribute 0) |
|
|
210
|
-
| `wallTimeMs` | Sums the time from `turn.started` to its completed or failed event |
|
|
211
|
-
| `custom` | Sums finite numeric deltas returned by `derive` |
|
|
212
|
-
|
|
213
|
-
`onSample` fires after every `turn.completed` and `turn.failed` event
|
|
214
|
-
for an enrolled arm. The sample contains:
|
|
215
|
-
|
|
216
|
-
| Field | Value |
|
|
217
|
-
| --- | --- |
|
|
218
|
-
| `experiment` | Experiment name |
|
|
219
|
-
| `variant`, `variantLabel?` | Sticky arm and optional display label |
|
|
220
|
-
| `sessionId`, `channelId` | Source session |
|
|
221
|
-
| `metrics` | Cumulative metrics through this turn |
|
|
222
|
-
| `reason` | `turn.completed` or `turn.failed` |
|
|
223
|
-
| `at` | Terminal event timestamp |
|
|
224
|
-
|
|
225
|
-
The metrics are cumulative, not per-turn deltas. A second sample from
|
|
226
|
-
the same session includes the first turn's counts.
|
|
227
|
-
|
|
228
|
-
Each `derive` extractor runs on every session event for its enrolled
|
|
229
|
-
experiment, including streamed `message.appended` events. Keep it
|
|
230
|
-
synchronous and cheap. Return a finite number to add a delta, or
|
|
231
|
-
`null` to skip the event. Send samples to your metrics service from
|
|
232
|
-
`onSample`; do not perform network or disk work in `derive`.
|
|
233
|
-
|
|
234
|
-
Skipped sessions never call `onSample`. Errors from `derive` or
|
|
235
|
-
`onSample` are logged, then metric collection continues.
|
|
236
|
-
|
|
237
|
-
## Inspect assignments and results
|
|
238
|
-
|
|
239
|
-
Open the playground's **A/Bs** tab to see aggregate arm totals and
|
|
240
|
-
per-session assignments. The tab reads `GET /v1/abs`.
|
|
241
|
-
|
|
242
|
-
The response has two views of the same durable data:
|
|
243
|
-
|
|
244
|
-
| Field | Contents |
|
|
245
|
-
| --- | --- |
|
|
246
|
-
| `experiments` | Declared variants, skipped-session count, arm session counts, and aggregate metrics |
|
|
247
|
-
| `sessions` | Visible sessions with their assignments and cumulative metrics |
|
|
248
|
-
|
|
249
|
-
`GET /v1/abs` returns sessions visible to the current principal by
|
|
250
|
-
default. In `--dev`, loopback requests include every session. Add
|
|
251
|
-
`--allow-anonymous` to include every session from non-loopback callers
|
|
252
|
-
too.
|
|
253
|
-
|
|
254
|
-
The session event stream is the source of truth for assignment + fold.
|
|
255
|
-
`GET /v1/abs` recomputes aggregates from those logs. Any
|
|
256
|
-
`agent/storage.ts` exports samples and snapshots durably: an authored
|
|
257
|
-
`abs` table when the backend has a native shape for it, or the table
|
|
258
|
-
derived over the KV core otherwise. See
|
|
259
|
-
[Storage](/docs/storage.md#eval-and-a-b-tables).
|
|
260
|
-
|
|
261
|
-
## Configure the playground fold window
|
|
262
|
-
|
|
263
|
-
Assignments and foldable metrics already persist in each session's
|
|
264
|
-
event stream. The optional `agent/ab.config.ts` only caps how many
|
|
265
|
-
sessions the playground and `GET /v1/abs` fold:
|
|
266
|
-
|
|
267
|
-
```ts
|
|
268
|
-
import { defineABConfig } from "@cursor/july/ab";
|
|
269
|
-
|
|
270
|
-
export default defineABConfig({
|
|
271
|
-
// Optional. Defaults to 200. Only affects GET /v1/abs / A/Bs tab.
|
|
272
|
-
maxPlaygroundSessions: 500,
|
|
273
|
-
});
|
|
274
|
-
```
|
|
275
|
-
|
|
276
|
-
`maxPlaygroundSessions` keeps the newest sessions in the fold. It does
|
|
277
|
-
not prune session logs or change assignment. For export to S3, a DB, or
|
|
278
|
-
your metrics vendor, send samples from `onSample` or declare a storage
|
|
279
|
-
`abs` table.
|
|
280
|
-
|
|
281
|
-
## Keep assignments durable
|
|
282
|
-
|
|
283
|
-
The append-only event stream is the source of truth. Each
|
|
284
|
-
`ab.assigned` event persists a variant key or null skip. Built-in
|
|
285
|
-
metrics come from the turn and tool events that follow it.
|
|
286
|
-
|
|
287
|
-
After a server restart or a parked session resumes, the live collector
|
|
288
|
-
replays the stream to rebuild cumulative counters. Replay does not call
|
|
289
|
-
`onSample` (or write to the storage `abs` table) for historical turns.
|
|
290
|
-
Only a new completed or failed turn emits another sample.
|
|
291
|
-
|
|
292
|
-
The snapshot API also replays `derive` across the full stream, so
|
|
293
|
-
custom totals match the current extractor. Changing a derive function
|
|
294
|
-
can change historical snapshot totals. Treat metric definitions as
|
|
295
|
-
versioned experiment code.
|
|
296
|
-
|
|
297
|
-
## Keep eval traffic separate
|
|
298
|
-
|
|
299
|
-
Sessions created by `agent-sdk eval` and the playground Evals runner use
|
|
300
|
-
`purpose: "eval"`. They skip A/B enrollment entirely:
|
|
301
|
-
|
|
302
|
-
- No split function runs.
|
|
303
|
-
- No `ab.assigned` event is recorded.
|
|
304
|
-
- No `onSample` callback fires.
|
|
305
|
-
- The session is omitted from `GET /v1/abs`.
|
|
306
|
-
|
|
307
|
-
Ordinary chat, `agent-sdk run`, Slack, GitHub, and other channel sessions
|
|
308
|
-
use the live purpose. You do not need `splitIf` to exclude eval traffic.
|
|
309
|
-
|
|
310
|
-
## Know the boundaries
|
|
311
|
-
|
|
312
|
-
`defineAB` provides sticky assignment, variant instructions,
|
|
313
|
-
`session.abs` for tools, cumulative metrics, and local inspection. It
|
|
314
|
-
does not provide:
|
|
315
|
-
|
|
316
|
-
- A test command, assertion API, or pass/fail result
|
|
317
|
-
- Statistical significance calculations
|
|
318
|
-
- An experiment rollout or lifecycle service
|
|
319
|
-
- Per-variant model or runtime configuration
|
|
320
|
-
- A built-in analytics warehouse (bring your own via `onSample` or the
|
|
321
|
-
storage `abs` table)
|
|
322
|
-
|
|
323
|
-
Use [evals](/docs/evals.md) to protect known behavior. Use `onSample` or a
|
|
324
|
-
storage `abs` table when you need sample/snapshot exports beyond the
|
|
325
|
-
session event log.
|
|
326
|
-
|
|
327
|
-
## What's next
|
|
328
|
-
|
|
329
|
-
Continue with these pages:
|
|
330
|
-
|
|
331
|
-
- [Evals](/docs/evals.md): pass/fail regression checks on fixed inputs
|
|
332
|
-
- [Hillclimbing](/docs/hillclimbing.md): improve an agent against fixed
|
|
333
|
-
fixtures
|
|
334
|
-
- [Hooks](/docs/reference/hooks.md): other event-stream consumers
|
|
335
|
-
- [Sessions and streaming](/docs/reference/sessions.md): the
|
|
336
|
-
`ab.assigned` event and durable log
|
|
337
|
-
- [Playground](/docs/reference/playground.md): the A/Bs tab
|
|
338
|
-
- [HTTP API](/docs/reference/http-api.md): `GET /v1/abs`
|
|
339
|
-
- [Live A/B metrics skill](/docs/skills/ab.md): have a coding agent
|
|
340
|
-
wire an experiment
|
|
341
|
-
|
|
342
|
-
---
|
|
343
|
-
|
|
344
7
|
Source: /docs/building-with-agents.md
|
|
345
8
|
|
|
346
9
|
# Building agents with agents
|
|
347
10
|
|
|
348
11
|
Give a coding agent the goal. The built-in skills guide it through
|
|
349
|
-
scaffolding, channels, verification, evals, and measured
|
|
12
|
+
scaffolding, channels, verification, deployment, evals, and measured
|
|
13
|
+
improvement.
|
|
350
14
|
|
|
351
15
|
## What can a coding agent build for me?
|
|
352
16
|
|
|
@@ -393,12 +57,12 @@ The package ships task-specific guides under [`skills/`](/docs/skills/index.md):
|
|
|
393
57
|
| Understand the project layout and runtimes | [`framework-map`](/docs/skills/framework-map.md) |
|
|
394
58
|
| Create and verify a new agent | [`create-agent`](/docs/skills/create-agent.md) |
|
|
395
59
|
| Write fixtures and regression checks | [`evals`](/docs/skills/evals.md) |
|
|
396
|
-
| Live A/B metrics on traffic (`defineAB`) | [`ab`](/docs/skills/ab.md) |
|
|
397
60
|
| Export OpenTelemetry traces | [`otel`](/docs/skills/otel.md) |
|
|
398
61
|
| Improve an agent against fixed inputs | [`hillclimb`](/docs/skills/hillclimb.md) |
|
|
399
62
|
| Add GitHub webhooks and replay events | [`github`](/docs/skills/github.md) |
|
|
400
63
|
| Connect an agent to Slack | [`setup-slack`](/docs/skills/setup-slack.md) |
|
|
401
64
|
| Authorize host MCP OAuth | [`mcp-auth`](/docs/skills/mcp-auth.md) |
|
|
65
|
+
| Deploy with an attached service account | [`deploy`](/docs/skills/deploy.md) |
|
|
402
66
|
| Diagnose a local run | [`debug`](/docs/skills/debug.md) |
|
|
403
67
|
|
|
404
68
|
Point your coding agent at the matching `SKILL.md`. The guide contains
|
|
@@ -505,7 +169,6 @@ name. For example, `agent/tools/get_weather.ts` creates a tool named
|
|
|
505
169
|
| `agent/mcp-connections/<name>.ts` | Tools from external MCP servers |
|
|
506
170
|
| `agent/host-connections/<name>.ts` | Privileged MCP servers for host tools only |
|
|
507
171
|
| `agent/channels/*.ts` | HTTP, Slack, and GitHub entry points |
|
|
508
|
-
| `agent/ab.ts` or `agent/ab/*.ts` | Sticky variants and live performance metrics |
|
|
509
172
|
| `agent/result.ts` | Optional host `commit` on the final assistant text |
|
|
510
173
|
| `evals/**/*.eval.ts` | Repeatable checks at the project root |
|
|
511
174
|
|
|
@@ -541,8 +204,8 @@ Each session records an append-only event stream. It includes:
|
|
|
541
204
|
- Turn completion and token usage
|
|
542
205
|
|
|
543
206
|
Sessions and their event streams survive server restarts. The
|
|
544
|
-
playground renders the stream. Evals assert against it.
|
|
545
|
-
`agent-sdk trajectory` command turns a saved stream into a short
|
|
207
|
+
playground renders the stream. [Evals](/docs/evals.md) assert against it.
|
|
208
|
+
The `agent-sdk trajectory` command turns a saved stream into a short
|
|
546
209
|
summary.
|
|
547
210
|
|
|
548
211
|
When a run surprises you, inspect its event stream first. See
|
|
@@ -636,7 +299,6 @@ See [Agent-to-agent](/docs/guides/agent-to-agent.md) for a complete example.
|
|
|
636
299
|
- [Project layout](/docs/reference/project-layout.md)
|
|
637
300
|
- [Sessions and streaming](/docs/reference/sessions.md)
|
|
638
301
|
- [Channels](/docs/reference/channels.md)
|
|
639
|
-
- [Live A/B metrics](/docs/ab.md)
|
|
640
302
|
|
|
641
303
|
---
|
|
642
304
|
|
|
@@ -3240,31 +2902,64 @@ Source: /docs/evals.md
|
|
|
3240
2902
|
|
|
3241
2903
|
# Evals
|
|
3242
2904
|
|
|
3243
|
-
An eval
|
|
3244
|
-
|
|
3245
|
-
|
|
3246
|
-
tweak helped, a refactor didn't regress the agent, and last
|
|
3247
|
-
|
|
2905
|
+
An eval sends a fixed message to your agent and asserts over the
|
|
2906
|
+
trajectory it records: the turn completed, the right tool ran with the
|
|
2907
|
+
right input, the reply has the right shape. Evals are how you know a
|
|
2908
|
+
prompt tweak helped, a refactor didn't regress the agent, and last
|
|
2909
|
+
month's fix still holds.
|
|
3248
2910
|
|
|
3249
|
-
|
|
3250
|
-
|
|
3251
|
-
|
|
3252
|
-
|
|
2911
|
+
Nothing is mocked. The runner starts (or targets) a real agent server,
|
|
2912
|
+
drives sessions over the public API, and grades the events it gets
|
|
2913
|
+
back. The model runs and server tools execute, so
|
|
2914
|
+
[keep side effects out of eval sessions](#keep-side-effects-out-of-eval-sessions)
|
|
2915
|
+
before you point an eval at an agent that posts anywhere.
|
|
3253
2916
|
|
|
3254
|
-
##
|
|
2917
|
+
## Evals, hooks, or hillclimbing?
|
|
3255
2918
|
|
|
3256
|
-
|
|
3257
|
-
|
|
3258
|
-
inside it (`agent/evals/` is silently ignored). TypeScript is the normal
|
|
3259
|
-
authoring format.
|
|
2919
|
+
All three read the same session event stream. Pick by the question you
|
|
2920
|
+
are asking.
|
|
3260
2921
|
|
|
3261
|
-
|
|
3262
|
-
|
|
3263
|
-
|
|
3264
|
-
|
|
2922
|
+
| You want to | Use |
|
|
2923
|
+
| --- | --- |
|
|
2924
|
+
| Gate one fixed input's behavior, locally and in CI | Evals (this page) |
|
|
2925
|
+
| Observe every live session: metrics, audit, alerts | [Hooks](/docs/reference/hooks.md) |
|
|
2926
|
+
| Improve an agent one measured round at a time | [Hillclimbing](/docs/hillclimbing.md); each kept win lands an eval |
|
|
2927
|
+
|
|
2928
|
+
[Hooks, channel events, or evals?](/docs/reference/hooks.md#hooks-channel-events-or-evals)
|
|
2929
|
+
has the side-by-side table.
|
|
2930
|
+
|
|
2931
|
+
### When not to write an eval
|
|
2932
|
+
|
|
2933
|
+
- Test a server tool's own logic with
|
|
2934
|
+
`agent-sdk call <tool> --dir . --input '{...}'` or a unit test. No
|
|
2935
|
+
model turn, no credential.
|
|
2936
|
+
- Explore a prompt with `agent-sdk run --dir . --message "..."` and
|
|
2937
|
+
read the trajectory. Write the eval once you know which decision to
|
|
2938
|
+
gate.
|
|
2939
|
+
- Stop a bad turn while it runs with
|
|
2940
|
+
[`needsApproval`](/docs/reference/tools.md#gate-a-tool-on-human-approval)
|
|
2941
|
+
on the tool or [`defineResult`](/docs/reference/result.md). Evals grade
|
|
2942
|
+
after the fact.
|
|
2943
|
+
|
|
2944
|
+
## Write your first eval
|
|
3265
2945
|
|
|
3266
|
-
|
|
3267
|
-
|
|
2946
|
+
Evals live under the project-root `evals/` directory, a sibling of
|
|
2947
|
+
`agent/`. `agent/evals/` is silently ignored. Discovery loads every
|
|
2948
|
+
`.eval.ts` (or `.eval.js`) file under `evals/`, plus one config file.
|
|
2949
|
+
|
|
2950
|
+
```text
|
|
2951
|
+
my-agent/
|
|
2952
|
+
agent/
|
|
2953
|
+
agent.ts
|
|
2954
|
+
tools/inspect_pr.ts
|
|
2955
|
+
evals/
|
|
2956
|
+
evals.config.ts # required to run: maxConcurrency
|
|
2957
|
+
readiness.eval.ts # id: readiness
|
|
2958
|
+
prs.eval.ts # cases: prs/checkout, prs/search
|
|
2959
|
+
```
|
|
2960
|
+
|
|
2961
|
+
An eval is a single `async test(t)`. You drive the agent with `t.send`
|
|
2962
|
+
and assert on the recorded run with the same `t`:
|
|
3268
2963
|
|
|
3269
2964
|
```ts
|
|
3270
2965
|
// evals/readiness.eval.ts
|
|
@@ -3286,12 +2981,49 @@ export default defineEval({
|
|
|
3286
2981
|
});
|
|
3287
2982
|
```
|
|
3288
2983
|
|
|
3289
|
-
|
|
3290
|
-
|
|
3291
|
-
|
|
2984
|
+
```ts
|
|
2985
|
+
// evals/evals.config.ts
|
|
2986
|
+
import { defineEvalConfig } from "@cursor/july/evals";
|
|
2987
|
+
|
|
2988
|
+
export default defineEvalConfig({ maxConcurrency: 20 });
|
|
2989
|
+
```
|
|
2990
|
+
|
|
2991
|
+
Run it under Node 22.13 or newer (never Bun) with a Cursor credential
|
|
2992
|
+
in place; see [Credentials](#credentials):
|
|
2993
|
+
|
|
2994
|
+
```bash
|
|
2995
|
+
agent-sdk eval --dir . --list
|
|
2996
|
+
agent-sdk eval --dir . readiness
|
|
2997
|
+
```
|
|
2998
|
+
|
|
2999
|
+
```text
|
|
3000
|
+
PASS readiness (14.2s) — Inspects a PR without approving it.
|
|
3001
|
+
✓ succeeded
|
|
3002
|
+
✓ calledTool(inspect_pr)
|
|
3003
|
+
✓ notCalledTool(approve_pr)
|
|
3004
|
+
✓ check(includes)
|
|
3005
|
+
|
|
3006
|
+
1 passed, 0 failed, 1 total
|
|
3007
|
+
artifacts: <project state directory>/evals/2026-09-11T15-02-11-402Z
|
|
3008
|
+
```
|
|
3009
|
+
|
|
3010
|
+
Every local run writes each case's assertions, inputs, tool calls, and
|
|
3011
|
+
`t.log` lines under that artifacts directory. Open
|
|
3012
|
+
`evals/<case-id>.json` there when a case fails; see
|
|
3013
|
+
[Where results land](#where-results-land).
|
|
3014
|
+
|
|
3015
|
+
## Name cases by path
|
|
3016
|
+
|
|
3017
|
+
The file path is the eval's identity, so you don't author an id.
|
|
3018
|
+
`evals/builds/api.eval.ts` becomes `builds/api`. An `index` filename
|
|
3019
|
+
collapses to its directory: `evals/builds/index.eval.ts` becomes
|
|
3020
|
+
`builds`.
|
|
3021
|
+
|
|
3022
|
+
One file can hold several datapoints through `cases`. Provide either
|
|
3023
|
+
`test` or `cases`, not both. Each case id becomes `<fileId>/<case.id>`:
|
|
3292
3024
|
|
|
3293
3025
|
```ts
|
|
3294
|
-
// evals/prs.eval.ts
|
|
3026
|
+
// evals/prs.eval.ts: prs/checkout, prs/search
|
|
3295
3027
|
export default defineEval({
|
|
3296
3028
|
tags: ["smoke", "prs"],
|
|
3297
3029
|
cases: [
|
|
@@ -3320,93 +3052,75 @@ export default defineEval({
|
|
|
3320
3052
|
});
|
|
3321
3053
|
```
|
|
3322
3054
|
|
|
3323
|
-
Case ids
|
|
3324
|
-
|
|
3325
|
-
`
|
|
3326
|
-
datapoint
|
|
3055
|
+
Case ids are single path segments, unique within the file. A case can
|
|
3056
|
+
set its own `description`, `tags`, `timeoutMs`, `iterations`, `judge`,
|
|
3057
|
+
`reporters`, and `metadata`. A case-level value replaces the file-level
|
|
3058
|
+
one for that datapoint, except `metadata`, which merges with case keys
|
|
3059
|
+
winning, and `reporters`, which adds to the file's list. `metadata` is
|
|
3060
|
+
free-form data carried onto the result and every reporter.
|
|
3061
|
+
|
|
3062
|
+
A file may instead export an array of `defineEval` calls to fan out
|
|
3063
|
+
over a dataset. Ids are then the file id plus a zero-padded index
|
|
3064
|
+
(`sql/0000`, `sql/0001`, ...); see [Load a dataset](#load-a-dataset).
|
|
3065
|
+
Prefer `cases` when datapoints are hand-written and deserve stable
|
|
3066
|
+
names.
|
|
3327
3067
|
|
|
3328
3068
|
### Iterations
|
|
3329
3069
|
|
|
3330
|
-
`iterations` (file or case, default `1`) runs a datapoint
|
|
3331
|
-
Discovery expands `iterations: 3` on case `nyc` to runnable
|
|
3332
|
-
`weather/nyc/1`, `weather/nyc/2`, `weather/nyc/3
|
|
3333
|
-
`weather/nyc` still selects all three
|
|
3334
|
-
`t.iteration`
|
|
3070
|
+
`iterations` (file or case, default `1`, cap `100`) runs a datapoint
|
|
3071
|
+
repeatedly. Discovery expands `iterations: 3` on case `nyc` to runnable
|
|
3072
|
+
ids `weather/nyc/1`, `weather/nyc/2`, and `weather/nyc/3`. The filter
|
|
3073
|
+
`weather/nyc` still selects all three. Each expanded case exposes
|
|
3074
|
+
`t.iteration` and `t.iterations`.
|
|
3335
3075
|
|
|
3336
|
-
`maxConcurrency` counts
|
|
3337
|
-
|
|
3338
|
-
|
|
3339
|
-
`maxConcurrency: 20`
|
|
3076
|
+
`maxConcurrency` counts authored datapoints, not expanded iterations.
|
|
3077
|
+
Iterations of one datapoint share a concurrency slot and run in
|
|
3078
|
+
sequence, so a suite of 11 cases with 3 iterations each and
|
|
3079
|
+
`maxConcurrency: 20` has at most 11 cases in flight.
|
|
3340
3080
|
|
|
3341
|
-
##
|
|
3081
|
+
## Drive the agent with `t.send`
|
|
3342
3082
|
|
|
3343
|
-
|
|
3344
|
-
|
|
3345
|
-
|
|
3346
|
-
200. Existing projects use 20. Discovery with `eval --list` works
|
|
3347
|
-
without this file, but running a case does not.
|
|
3083
|
+
`t.send(message, options?)` runs one turn and waits for it to settle:
|
|
3084
|
+
complete, park on an approval request, or fail. Several sends in one
|
|
3085
|
+
case share the session, which is how you write multi-turn evals.
|
|
3348
3086
|
|
|
3349
|
-
|
|
3350
|
-
|
|
3087
|
+
Each send resolves to a turn result: `message` (the assistant text),
|
|
3088
|
+
`sessionId`, `events`, `toolCalls` (tool names in order), `ok`, and
|
|
3089
|
+
`index`. The turn carries the same assertion vocabulary as `t`, scoped
|
|
3090
|
+
to that turn, so you can grade an intermediate turn before the next
|
|
3091
|
+
send overwrites `t.reply`. `turn.expectOk()` throws when the turn
|
|
3092
|
+
failed, for later steps that depend on it.
|
|
3351
3093
|
|
|
3352
|
-
|
|
3353
|
-
|
|
3354
|
-
|
|
3355
|
-
|
|
3356
|
-
|
|
3357
|
-
|
|
3094
|
+
Read the whole case with `t.reply` (last assistant text), `t.events`
|
|
3095
|
+
(every event so far), `t.turns` (settled turns, oldest first), and
|
|
3096
|
+
`t.sessionId`. `t.signal` aborts when the case hits its timeout; pass
|
|
3097
|
+
it to your own async work.
|
|
3098
|
+
|
|
3099
|
+
Three options apply on the first send only, because they shape session
|
|
3100
|
+
creation:
|
|
3101
|
+
|
|
3102
|
+
| Option | Effect |
|
|
3103
|
+
| --- | --- |
|
|
3104
|
+
| `workspaceFiles` | `{ path: contents }` seeded into the local session workspace. Prefer this over machine-local paths |
|
|
3105
|
+
| `workspaceDir` | Absolute harness cwd for the local runtime |
|
|
3106
|
+
| `cloud` | Per-session cloud options merged over the agent's static `cloud` config. Pin a fixture repo here for cloud evals instead of on the agent's default `cloud.repos`. On the cloud runtime, seeded files reach the agent as described under [Where does a turn run?](/docs/concepts.md#where-does-a-turn-run) |
|
|
3107
|
+
|
|
3108
|
+
```ts
|
|
3109
|
+
await t.send("Review pr/diff.patch and post findings.", {
|
|
3110
|
+
workspaceFiles: {
|
|
3111
|
+
"pr/diff.patch": [
|
|
3112
|
+
"diff --git a/app/routes/search.ts b/app/routes/search.ts",
|
|
3113
|
+
"+res.send(`<h1>Results for ${req.query.q}</h1>`);",
|
|
3114
|
+
].join("\n"),
|
|
3115
|
+
},
|
|
3358
3116
|
});
|
|
3359
3117
|
```
|
|
3360
3118
|
|
|
3361
|
-
|
|
3362
|
-
project config `timeoutMs`, then the 180-second runner default.
|
|
3363
|
-
|
|
3364
|
-
The optional fields:
|
|
3119
|
+
## Assert over the trajectory
|
|
3365
3120
|
|
|
3366
|
-
|
|
3367
|
-
|
|
3368
|
-
|
|
3369
|
-
| `judge` | unset | Default judge model for `t.judge.*`; see [Judge free-form output](#judge-free-form-output) |
|
|
3370
|
-
| `reporters` | unset | Destinations that observe every case; `--skip-report` suppresses them |
|
|
3371
|
-
| `maxPlaygroundRuns` | `20` | Max batches in the playground / `/v1/dev/evals*` history (not CLI `eval`). Hard-capped at 500. |
|
|
3372
|
-
|
|
3373
|
-
Reporters come from `@cursor/july/evals/reporters`: `JUnit` writes a
|
|
3374
|
-
JUnit XML file for CI, `Artifacts` writes per-case files, and
|
|
3375
|
-
`combineReporters` merges several into one (`renderJUnitXml` renders
|
|
3376
|
-
the XML for a custom destination). A file or case can add its own
|
|
3377
|
-
`reporters` on top of the config list.
|
|
3378
|
-
|
|
3379
|
-
Playground batches survive restarts whenever `agent/storage.ts` exists
|
|
3380
|
-
with an `evals` table or a KV core providing `delete` and `list` (the
|
|
3381
|
-
table is derived over the core); see
|
|
3382
|
-
[Storage](/docs/storage.md#eval-and-a-b-tables). Without storage they live
|
|
3383
|
-
in process memory and disappear when `serve` exits. Navigating away
|
|
3384
|
-
and back still works while the process is up.
|
|
3385
|
-
|
|
3386
|
-
## Drive and assert with `t`
|
|
3387
|
-
|
|
3388
|
-
`t` is both the driver and the assertion surface. You write ordinary
|
|
3389
|
-
control flow, sending turns and asserting inline.
|
|
3390
|
-
|
|
3391
|
-
Drive the agent with `t.send(message, options?)`. It runs one turn and
|
|
3392
|
-
waits for the session to park or fail. Multiple sends in one case share
|
|
3393
|
-
the session, which is how you write multi-turn evals.
|
|
3394
|
-
|
|
3395
|
-
Each `t.send` resolves to a turn result with `message`,
|
|
3396
|
-
`sessionId`, `events`, `toolCalls`, `ok`, and `index`. The turn carries
|
|
3397
|
-
the same assertion vocabulary as `t`, scoped to that turn, so you can
|
|
3398
|
-
grade an intermediate turn before the next send overwrites `t.reply`.
|
|
3399
|
-
`turn.expectOk()` throws when the turn failed, for later
|
|
3400
|
-
steps that depend on it.
|
|
3401
|
-
|
|
3402
|
-
Read the full case state with `t.reply` (the last assistant text),
|
|
3403
|
-
`t.events` (session events captured so far), `t.turns` (settled
|
|
3404
|
-
turns, oldest first), and `t.sessionId`. `t.signal` aborts when the
|
|
3405
|
-
case hits its timeout; pass it to your own async work. A thrown
|
|
3406
|
-
[turn result](/docs/reference/result.md) `commit` fails the turn, so
|
|
3407
|
-
`t.succeeded()` fails too.
|
|
3408
|
-
|
|
3409
|
-
Assert with the gates:
|
|
3121
|
+
Assertions record; they never throw. One run reports every failure
|
|
3122
|
+
instead of dying on the first. Assertions on `t` read the whole run.
|
|
3123
|
+
Assertions on a turn read only that turn.
|
|
3410
3124
|
|
|
3411
3125
|
| Gate | Checks |
|
|
3412
3126
|
| --- | --- |
|
|
@@ -3415,158 +3129,483 @@ Assert with the gates:
|
|
|
3415
3129
|
| `t.messageIncludes(token)` | the joined assistant text matches a string or `RegExp` |
|
|
3416
3130
|
| `t.calledTool(name, matcher?)` | a matching call to `name` happened |
|
|
3417
3131
|
| `t.notCalledTool(name)` | no request for `name`, in any lifecycle state |
|
|
3418
|
-
| `t.loadedSkill(name)` | the agent opened
|
|
3419
|
-
| `t.toolOrder(names)` | tool requests appear in this relative order
|
|
3132
|
+
| `t.loadedSkill(name)` | the agent opened `skills/<name>/SKILL.md` (read, grep, or shell `cat`) |
|
|
3133
|
+
| `t.toolOrder(names)` | tool requests appear in this relative order; extra calls allowed |
|
|
3420
3134
|
| `t.usedNoTools()` | no tool calls at all |
|
|
3421
3135
|
| `t.maxToolCalls(max)` | at most `max` tool calls |
|
|
3422
3136
|
| `t.noFailedActions()` | no tool call reported an error |
|
|
3423
3137
|
| `t.calledSubagent(name, matcher?)` | a matching subagent delegation happened |
|
|
3424
3138
|
| `t.taggedArtifact(kind?, predicate?)` | at least one [artifact](/docs/reference/artifacts.md) was tagged |
|
|
3425
|
-
| `t.event(type, matcher?)` | at least one matching event of `type`
|
|
3426
|
-
| `t.notEvent(type, matcher?)` | no matching event of `type`
|
|
3139
|
+
| `t.event(type, matcher?)` | at least one matching [event](/docs/reference/sessions.md#which-events-can-i-stream) of `type` |
|
|
3140
|
+
| `t.notEvent(type, matcher?)` | no matching event of `type` |
|
|
3427
3141
|
| `t.eventOrder(matchers)` | matching event groups occur in this relative order |
|
|
3428
3142
|
| `t.eventsSatisfy(label, predicate)` | your predicate over the typed event stream |
|
|
3429
|
-
| `t.check(value, expectation)` | any value, against a builder |
|
|
3430
|
-
| `t.score(name, value)` |
|
|
3431
|
-
|
|
3432
|
-
|
|
3143
|
+
| `t.check(value, expectation)` | any value, against a [builder](#grade-values-with-expectation-builders) |
|
|
3144
|
+
| `t.score(name, value)` | a 0-1 score you computed; soft until you add a bar |
|
|
3145
|
+
|
|
3146
|
+
Three more assertions gate and return the matched fact. They stop the
|
|
3147
|
+
test body when nothing matches, without a duplicate execution error.
|
|
3148
|
+
`t.requireToolCall(name, matcher?)` returns the call so later code can
|
|
3149
|
+
read its `input` and `output`. `t.requireInputRequest(filter?)` returns
|
|
3150
|
+
the single pending approval request. `await t.require(value, expectation)`
|
|
3151
|
+
does the same for a value check.
|
|
3152
|
+
|
|
3153
|
+
A case with no assertions passes when at least one turn completed. Add
|
|
3154
|
+
`t.succeeded()` and behavior gates anyway. They make the contract
|
|
3155
|
+
visible in review.
|
|
3433
3156
|
|
|
3434
|
-
|
|
3435
|
-
|
|
3436
|
-
|
|
3157
|
+
### What good cases assert
|
|
3158
|
+
|
|
3159
|
+
Gate decisions and shape, not prose. Model wording varies run to run.
|
|
3160
|
+
Tool choice, tool avoidance, and output structure are the stable
|
|
3161
|
+
contract.
|
|
3162
|
+
|
|
3163
|
+
1. `t.succeeded()`: always, first.
|
|
3164
|
+
2. The tool decision: `calledTool` for the intended path,
|
|
3165
|
+
`notCalledTool` for the likely wrong alternative. The pair is
|
|
3166
|
+
stronger than either alone.
|
|
3167
|
+
3. Output shape: a regex for the contract (`/ready|blocked/i`, a JSON
|
|
3168
|
+
marker, a findings-block fence), never exact sentences.
|
|
3169
|
+
4. For structured output, parse `t.reply` and check fields with
|
|
3170
|
+
`matches` or `satisfies` instead of substring-matching JSON.
|
|
3171
|
+
|
|
3172
|
+
The common failure modes: asserting exact phrasing, packing more than
|
|
3173
|
+
about five gates into one case (split it), and cases that depend on
|
|
3174
|
+
live external state that drifts (pin the input).
|
|
3175
|
+
|
|
3176
|
+
### Narrow tool assertions with matchers
|
|
3437
3177
|
|
|
3438
3178
|
With no matcher, `calledTool` is request-based: a requested call counts
|
|
3439
|
-
even
|
|
3440
|
-
|
|
3441
|
-
|
|
3442
|
-
|
|
3443
|
-
|
|
3444
|
-
|
|
3445
|
-
|
|
3446
|
-
|
|
3447
|
-
|
|
3448
|
-
|
|
3449
|
-
|
|
3450
|
-
|
|
3451
|
-
|
|
3452
|
-
|
|
3453
|
-
|
|
3454
|
-
|
|
3455
|
-
|
|
3456
|
-
|
|
3457
|
-
|
|
3458
|
-
|
|
3459
|
-
|
|
3460
|
-
|
|
3461
|
-
|
|
3462
|
-
|
|
3463
|
-
|
|
3464
|
-
|
|
3465
|
-
|
|
3466
|
-
|
|
3467
|
-
|
|
3468
|
-
|
|
3469
|
-
|
|
3470
|
-
|
|
3471
|
-
|
|
3472
|
-
|
|
3179
|
+
even before its result arrives. A matcher narrows it:
|
|
3180
|
+
|
|
3181
|
+
```ts
|
|
3182
|
+
t.calledTool("inspect_pr", { status: "completed" });
|
|
3183
|
+
t.calledTool("apply_agents", { input: { verdict: "update" } });
|
|
3184
|
+
t.calledTool("bash", { input: { command: /^gh pr view/ }, count: 1 });
|
|
3185
|
+
t.calledTool("read_file", {
|
|
3186
|
+
output: (value) => String(value).includes("TODO"),
|
|
3187
|
+
});
|
|
3188
|
+
```
|
|
3189
|
+
|
|
3190
|
+
`input`, `output`, and `count` accept a literal, a `RegExp`, or a
|
|
3191
|
+
predicate. Object literals partial-deep-match, so `{ verdict: "update" }`
|
|
3192
|
+
matches arguments that also carry other keys. `status` is one of
|
|
3193
|
+
`completed`, `failed`, `pending`, or `rejected` (a human denied the
|
|
3194
|
+
approval). `calledSubagent` takes `{ output, status, count, callId }`.
|
|
3195
|
+
`event`, `notEvent`, and `eventOrder` take `{ data, count }`.
|
|
3196
|
+
|
|
3197
|
+
### Grade values with expectation builders
|
|
3198
|
+
|
|
3199
|
+
`t.check(value, expectation)` grades any value: `t.reply`, a parsed
|
|
3200
|
+
JSON field, a tool's output.
|
|
3201
|
+
|
|
3202
|
+
| Builder | Checks | Severity |
|
|
3203
|
+
| --- | --- | --- |
|
|
3204
|
+
| `includes(string \| RegExp)` | substring or match; structured values are stringified first | gate |
|
|
3205
|
+
| `equals(value)` | deep equality | gate |
|
|
3206
|
+
| `matches(schema)` | a Standard Schema (Zod, Valibot, ...) or anything with `safeParse` | gate |
|
|
3207
|
+
| `similarity(expected)` | normalized text similarity, 0-1 | soft |
|
|
3208
|
+
| `satisfies(predicate, label)` | your predicate; `label` is the failure detail | gate |
|
|
3209
|
+
|
|
3210
|
+
```ts
|
|
3211
|
+
import { matches, satisfies } from "@cursor/july/evals";
|
|
3212
|
+
import { z } from "zod";
|
|
3213
|
+
|
|
3214
|
+
const verdict = JSON.parse(t.reply ?? "{}");
|
|
3215
|
+
t.check(
|
|
3216
|
+
verdict,
|
|
3217
|
+
matches(z.object({ ready: z.boolean(), blockers: z.array(z.string()) }))
|
|
3218
|
+
);
|
|
3473
3219
|
t.check(
|
|
3474
|
-
|
|
3475
|
-
satisfies((n) => (n as number) <=
|
|
3220
|
+
verdict.blockers.length,
|
|
3221
|
+
satisfies((n) => (n as number) <= 3, "at most 3 blockers")
|
|
3476
3222
|
);
|
|
3477
3223
|
```
|
|
3478
3224
|
|
|
3479
|
-
|
|
3480
|
-
|
|
3481
|
-
|
|
3225
|
+
`normalizedSimilarity(actual, expected)` returns the same 0-1 score as
|
|
3226
|
+
`similarity`, for use with `t.score`.
|
|
3227
|
+
|
|
3228
|
+
### Record without gating
|
|
3229
|
+
|
|
3230
|
+
- `t.metric(name, value)` records a structured score or label. It shows
|
|
3231
|
+
on the CLI result, the playground case card, JUnit output, and
|
|
3232
|
+
artifacts.
|
|
3233
|
+
- `t.log(message)` records a debug line, streamed under `--verbose`.
|
|
3234
|
+
- `t.skip(reason)` ends the case as skipped. Skipped cases report
|
|
3235
|
+
separately and never change the exit code. Call it before sending
|
|
3236
|
+
messages.
|
|
3237
|
+
|
|
3238
|
+
## Gates, soft scores, and verdicts
|
|
3239
|
+
|
|
3240
|
+
Every assertion returns a handle, so severity rides on the assertion
|
|
3241
|
+
instead of a separate thresholds map:
|
|
3242
|
+
|
|
3243
|
+
```ts
|
|
3244
|
+
t.succeeded(); // gate (default)
|
|
3245
|
+
t.calledTool("get_weather").soft(); // tracked, never fails
|
|
3246
|
+
t.check(t.reply, similarity("Sunny, 72°F")).atLeast(0.8); // soft, with a bar
|
|
3247
|
+
t.judge.closedQA("cites a source").gate(0.8); // promoted to a gate
|
|
3248
|
+
```
|
|
3249
|
+
|
|
3250
|
+
- `.gate(threshold?)` is hard. A miss fails the case and `eval` exits 1.
|
|
3251
|
+
- `.soft(threshold?)` is tracked. With no threshold it never fails.
|
|
3252
|
+
- `.atLeast(threshold)` is soft with a bar. A miss marks the case
|
|
3253
|
+
`scored`.
|
|
3482
3254
|
|
|
3483
|
-
|
|
3255
|
+
Each case ends with one verdict:
|
|
3484
3256
|
|
|
3485
|
-
|
|
3486
|
-
|
|
3487
|
-
|
|
3488
|
-
|
|
3257
|
+
| Verdict | Meaning | Exit code |
|
|
3258
|
+
| --- | --- | --- |
|
|
3259
|
+
| `passed` | every gate passed and no soft bar was missed | 0 |
|
|
3260
|
+
| `failed` | a gate failed, or the test body threw | 1 |
|
|
3261
|
+
| `scored` | only soft bars were missed | 0, or 1 under `--strict` |
|
|
3262
|
+
| `skipped` | `t.skip(reason)`, or a judge with no credentials | 0 |
|
|
3263
|
+
|
|
3264
|
+
The CLI prints soft misses as `~` and gate misses as `✗`. Start a new
|
|
3265
|
+
benchmark with `t.score("recall", recall)` and `.atLeast()` so its
|
|
3266
|
+
number reports for a while without blocking merges. Add `--strict`
|
|
3267
|
+
once the bars are trustworthy.
|
|
3268
|
+
|
|
3269
|
+
## Judge free-form output
|
|
3270
|
+
|
|
3271
|
+
When wording matters and no regex captures it, `t.judge` grades with an
|
|
3272
|
+
LLM. The graders are `factuality(expected)`, `summarizes(expected)`,
|
|
3273
|
+
`closedQA(criteria)`, and `sql(expected)`. Each scores `t.reply` by
|
|
3274
|
+
default; pass `{ on }` to grade another value.
|
|
3489
3275
|
|
|
3490
3276
|
```ts
|
|
3491
|
-
t.
|
|
3277
|
+
const summary = await t.send("Why did CI fail on PR 42?");
|
|
3278
|
+
t.judge.factuality("The lint step failed on src/sidebar.ts.").atLeast(0.7);
|
|
3279
|
+
t.judge.closedQA("names the failing step", { on: summary.message }).gate(1);
|
|
3492
3280
|
```
|
|
3493
3281
|
|
|
3494
3282
|
Judge assertions are soft by default, so a judge never fails a build
|
|
3495
|
-
until you give it a bar with `.atLeast(
|
|
3496
|
-
|
|
3283
|
+
until you give it a bar with `.atLeast()` or promote it with `.gate()`.
|
|
3284
|
+
The recorded detail names the choice the judge made and its rationale.
|
|
3285
|
+
|
|
3286
|
+
The judge model comes from `defineEvalConfig({ judge })`,
|
|
3497
3287
|
`defineEval({ judge })`, a case-level `judge`, or a per-call
|
|
3498
|
-
`{ model }
|
|
3499
|
-
|
|
3500
|
-
|
|
3501
|
-
|
|
3288
|
+
`{ model }`. The nearest one wins. A judge call with no model
|
|
3289
|
+
configured fails the case. A judge that cannot reach a model (no
|
|
3290
|
+
credential) ends the case as `skipped`, unless a deterministic gate
|
|
3291
|
+
already failed.
|
|
3292
|
+
|
|
3293
|
+
For a domain-specific judge whose verdict is not a single score,
|
|
3294
|
+
`t.judge.model(prompt)` sends a raw prompt to the same model and
|
|
3295
|
+
returns the reply. Record the parsed result with `t.score` or
|
|
3296
|
+
`t.check`. Anything derived from the agent under test is untrusted
|
|
3297
|
+
input to your prompt: wrap it with `fenceUntrusted` and include
|
|
3298
|
+
`EVAL_JUDGE_INJECTION_GUARD`, as the built-in graders do.
|
|
3299
|
+
|
|
3300
|
+
```ts
|
|
3301
|
+
import {
|
|
3302
|
+
EVAL_JUDGE_INJECTION_GUARD,
|
|
3303
|
+
fenceUntrusted,
|
|
3304
|
+
} from "@cursor/july/evals";
|
|
3305
|
+
|
|
3306
|
+
const gold = ["XSS in search.ts", "open redirect in login.ts"];
|
|
3307
|
+
const reply = await t.judge.model(
|
|
3308
|
+
[
|
|
3309
|
+
"For each GOLD finding, answer whether SUBMISSION reports it.",
|
|
3310
|
+
"Reply with one line per finding: <index> YES|NO.",
|
|
3311
|
+
EVAL_JUDGE_INJECTION_GUARD,
|
|
3312
|
+
fenceUntrusted("GOLD", gold.map((g, i) => `${i + 1}. ${g}`).join("\n")),
|
|
3313
|
+
fenceUntrusted("SUBMISSION", t.reply ?? ""),
|
|
3314
|
+
].join("\n\n")
|
|
3315
|
+
);
|
|
3316
|
+
const hits = reply.match(/\bYES\b/g)?.length ?? 0;
|
|
3317
|
+
t.score("recall", hits / gold.length).atLeast(0.5);
|
|
3318
|
+
```
|
|
3319
|
+
|
|
3320
|
+
## Keep side effects out of eval sessions
|
|
3321
|
+
|
|
3322
|
+
Eval sessions run the real agent, tools included. A reviewer that
|
|
3323
|
+
comments on GitHub or posts to Slack will do so from an eval unless
|
|
3324
|
+
the tool checks the session's purpose. Eval sessions carry
|
|
3325
|
+
`purpose: "eval"`; live traffic carries `"live"`. Branch on it in the
|
|
3326
|
+
tool, hook, or result handler that actuates:
|
|
3327
|
+
|
|
3328
|
+
```ts
|
|
3329
|
+
// agent/tools/post_findings.ts
|
|
3330
|
+
async execute({ findings }, ctx) {
|
|
3331
|
+
if (ctx.session.purpose === "eval") {
|
|
3332
|
+
return { posted: false, reason: "eval", count: findings.length };
|
|
3333
|
+
}
|
|
3334
|
+
// post the review
|
|
3335
|
+
}
|
|
3336
|
+
```
|
|
3337
|
+
|
|
3338
|
+
Return a shaped result instead of throwing, so the eval can still
|
|
3339
|
+
assert `t.calledTool("post_findings", { input: ... })` on the decision.
|
|
3340
|
+
The same check belongs in [hooks](/docs/reference/hooks.md) that meter or
|
|
3341
|
+
page and in [`defineResult`](/docs/reference/result.md) commits.
|
|
3342
|
+
|
|
3343
|
+
## Worked examples
|
|
3344
|
+
|
|
3345
|
+
### Multi-turn: grade each turn
|
|
3346
|
+
|
|
3347
|
+
```ts
|
|
3348
|
+
// evals/intro.eval.ts
|
|
3349
|
+
import { defineEval, includes, satisfies } from "@cursor/july/evals";
|
|
3350
|
+
|
|
3351
|
+
export default defineEval({
|
|
3352
|
+
description: "Introduces itself once; a repeat mention gets a short ack.",
|
|
3353
|
+
async test(t) {
|
|
3354
|
+
const intro = await t.send("Meet Jenny! @Jenny introduce yourself.");
|
|
3355
|
+
intro.expectOk();
|
|
3356
|
+
t.check(intro.message, includes(/jenny/i));
|
|
3357
|
+
|
|
3358
|
+
const repeat = await t.send("Meet, @Jenny!");
|
|
3359
|
+
t.succeeded();
|
|
3360
|
+
repeat.usedNoTools();
|
|
3361
|
+
t.check(
|
|
3362
|
+
repeat.message,
|
|
3363
|
+
satisfies((r) => (r as string).trim().length <= 280, "short ack")
|
|
3364
|
+
);
|
|
3365
|
+
t.check(
|
|
3366
|
+
repeat.message,
|
|
3367
|
+
satisfies((r) => !/what i can do/i.test(r as string), "no second intro")
|
|
3368
|
+
);
|
|
3369
|
+
},
|
|
3370
|
+
});
|
|
3371
|
+
```
|
|
3372
|
+
|
|
3373
|
+
`t.succeeded()` grades the whole session. `repeat.usedNoTools()` and the
|
|
3374
|
+
checks on `repeat.message` read only the second turn, even though
|
|
3375
|
+
`t.reply` now holds its text.
|
|
3376
|
+
|
|
3377
|
+
### Approvals: assert the parked decision
|
|
3378
|
+
|
|
3379
|
+
For a tool with `needsApproval`, the turn parks instead of finishing.
|
|
3380
|
+
Gate on `t.parked()` and on the arguments the model chose:
|
|
3381
|
+
|
|
3382
|
+
```ts
|
|
3383
|
+
// evals/agents.eval.ts (one case; RULE and SLACK are fixture strings)
|
|
3384
|
+
{
|
|
3385
|
+
id: "update-rule",
|
|
3386
|
+
description: "A repeated billing rule parks the AGENTS.md write.",
|
|
3387
|
+
async test(t) {
|
|
3388
|
+
await t.send("Weekly AGENTS.md review. Read week/ and call apply_agents once.", {
|
|
3389
|
+
workspaceFiles: {
|
|
3390
|
+
"week/prs.md": RULE,
|
|
3391
|
+
"week/slack.md": SLACK,
|
|
3392
|
+
"week/tree/AGENTS.md.txt": "# API\n\nKeep handlers thin.\n",
|
|
3393
|
+
},
|
|
3394
|
+
});
|
|
3395
|
+
t.parked();
|
|
3396
|
+
t.calledTool("apply_agents", { input: { verdict: "update" } });
|
|
3397
|
+
},
|
|
3398
|
+
},
|
|
3399
|
+
```
|
|
3400
|
+
|
|
3401
|
+
`t.parked()` and `t.succeeded()` are exclusive: a parked run is a clean
|
|
3402
|
+
stop on an unanswered approval, not a completed one. Pair the parked
|
|
3403
|
+
case with a sibling that expects `verdict: "skip"` and `t.succeeded()`,
|
|
3404
|
+
so both branches stay pinned.
|
|
3405
|
+
|
|
3406
|
+
## Pin fixtures
|
|
3407
|
+
|
|
3408
|
+
A fixed input is what makes an eval repeatable. Pick the fixture by the
|
|
3409
|
+
surface under test.
|
|
3410
|
+
|
|
3411
|
+
| Agent surface | Fixture |
|
|
3412
|
+
| --- | --- |
|
|
3413
|
+
| Chat or domain assistant | One canonical prompt string, chosen once and frozen |
|
|
3414
|
+
| Tool-heavy | Run `agent-sdk call <tool>` first to pin what the tool returns, then freeze the prompt that triggers it |
|
|
3415
|
+
| GitHub webhook | `agent-sdk github replay <pr> --events '*' --dry-run --out fixtures/github` snapshots real payloads for offline replay ([GitHub guide](/docs/guides/github.md#test-with-github-replay)) |
|
|
3416
|
+
| PR reviewer with host preparation | Diff, metadata, and gold labels pinned to commit SHAs; keep any live PR matrix small |
|
|
3417
|
+
| Workspace-dependent | `workspaceFiles` on the first `t.send`, never developer-machine paths |
|
|
3418
|
+
|
|
3419
|
+
Tag the fast, reliably passing core `smoke` and run `--tag smoke` in
|
|
3420
|
+
the inner loop. Leave slow or drift-prone cases untagged for explicit
|
|
3421
|
+
runs.
|
|
3422
|
+
|
|
3423
|
+
### Materialize API-backed fixtures
|
|
3424
|
+
|
|
3425
|
+
An input that only points at external data (a pull request URL, a
|
|
3426
|
+
snapshot id, a pair of commit SHAs) is not self-contained. Fetch it once
|
|
3427
|
+
and commit the rendered fixture before you expand the suite:
|
|
3428
|
+
|
|
3429
|
+
1. Save the diff, metadata, and labels under `fixtures/` at pinned
|
|
3430
|
+
revisions.
|
|
3431
|
+
2. Seed those files with `workspaceFiles`, or read them from the
|
|
3432
|
+
fixture directory.
|
|
3433
|
+
3. Assert decisions and output shape against the saved evidence.
|
|
3434
|
+
4. Keep a small `smoke` subset for any remaining live checks.
|
|
3435
|
+
|
|
3436
|
+
`maxConcurrency` limits parallel datapoints, not the model or API
|
|
3437
|
+
fan-out inside one datapoint. Materialized fixtures keep a large suite
|
|
3438
|
+
from exhausting provider and GitHub rate limits.
|
|
3439
|
+
|
|
3440
|
+
### Load a dataset
|
|
3441
|
+
|
|
3442
|
+
Read committed fixtures with `loadJson`, `loadJsonl`, and `loadYaml`
|
|
3443
|
+
from `@cursor/july/evals/loaders`. Relative paths resolve against the
|
|
3444
|
+
project root the runner discovered, not the cwd the CLI ran from. Eval
|
|
3445
|
+
files are ES modules, so top-level `await` can load a dataset and fan
|
|
3446
|
+
one file out over it:
|
|
3447
|
+
|
|
3448
|
+
```ts
|
|
3449
|
+
// evals/sql.eval.ts: sql/0000, sql/0001, ...
|
|
3450
|
+
import { defineEval, equals } from "@cursor/july/evals";
|
|
3451
|
+
import { loadYaml } from "@cursor/july/evals/loaders";
|
|
3452
|
+
|
|
3453
|
+
const rows = await loadYaml<{ task: string; prompt: string; sql: string }[]>(
|
|
3454
|
+
"evals/data/cases.yaml"
|
|
3455
|
+
);
|
|
3456
|
+
|
|
3457
|
+
export default rows.map((row) =>
|
|
3458
|
+
defineEval({
|
|
3459
|
+
description: row.task,
|
|
3460
|
+
async test(t) {
|
|
3461
|
+
await t.send(row.prompt);
|
|
3462
|
+
t.succeeded();
|
|
3463
|
+
t.check(t.reply, equals(row.sql));
|
|
3464
|
+
},
|
|
3465
|
+
})
|
|
3466
|
+
);
|
|
3467
|
+
```
|
|
3468
|
+
|
|
3469
|
+
## Configure eval runs
|
|
3470
|
+
|
|
3471
|
+
`evals/evals.config.ts` (or `.js`) holds project-wide defaults. It must
|
|
3472
|
+
set `maxConcurrency`. Each case issues real model requests, so
|
|
3473
|
+
concurrency is hard-capped at 200; the templates use 10.
|
|
3474
|
+
`eval --list` works without the file. Running a case does not.
|
|
3475
|
+
|
|
3476
|
+
```ts
|
|
3477
|
+
import { defineEvalConfig } from "@cursor/july/evals";
|
|
3478
|
+
|
|
3479
|
+
export default defineEvalConfig({
|
|
3480
|
+
maxConcurrency: 20,
|
|
3481
|
+
timeoutMs: 180_000,
|
|
3482
|
+
judge: { model: "gpt-5.4-mini" },
|
|
3483
|
+
});
|
|
3484
|
+
```
|
|
3485
|
+
|
|
3486
|
+
| Option | Default | Meaning |
|
|
3487
|
+
| --- | --- | --- |
|
|
3488
|
+
| `maxConcurrency` | required | Datapoints in flight at once, 1-200 |
|
|
3489
|
+
| `timeoutMs` | `180_000` | Per-case timeout. Precedence: case or file `timeoutMs`, then `--timeout-ms`, then this value |
|
|
3490
|
+
| `judge` | unset | Default judge model for `t.judge.*` |
|
|
3491
|
+
| `reporters` | unset | Destinations that observe every case; `--skip-report` suppresses them |
|
|
3492
|
+
| `maxPlaygroundRuns` | `20` | Batches kept in playground history (cap 500). Counts `--prod` and `--url` batches, not the default local run |
|
|
3502
3493
|
|
|
3503
|
-
|
|
3494
|
+
Reporters ship results somewhere; the runner still does the grading.
|
|
3495
|
+
`JUnit({ filePath, suiteName? })` writes JUnit XML and
|
|
3496
|
+
`Artifacts({ dir })` writes per-case files, both from
|
|
3497
|
+
`@cursor/july/evals/reporters`. A custom reporter is an object with any
|
|
3498
|
+
of `onRunStart`, `onEvalComplete`, and `onRunComplete`. A reporter that
|
|
3499
|
+
throws is logged and never fails the run. CI usually attaches the
|
|
3500
|
+
built-in two with `--junit` and `--artifacts` instead of `reporters`,
|
|
3501
|
+
so output paths stay with the pipeline, not the eval author.
|
|
3504
3502
|
|
|
3505
|
-
|
|
3503
|
+
Playground batches survive restarts when the project has
|
|
3504
|
+
[storage](/docs/storage.md#eval-table). Otherwise they live in process
|
|
3505
|
+
memory until `serve` exits.
|
|
3506
3506
|
|
|
3507
|
-
Run the CLI
|
|
3508
|
-
breaks tool-result streams and causes eval turns to fail.
|
|
3507
|
+
## Run evals from the CLI
|
|
3509
3508
|
|
|
3510
3509
|
```bash
|
|
3511
|
-
agent-sdk eval --dir . --list
|
|
3512
|
-
agent-sdk eval --dir .
|
|
3513
|
-
agent-sdk eval --dir . builds/checkout
|
|
3514
|
-
agent-sdk eval --dir . builds search
|
|
3515
|
-
agent-sdk eval --dir . --tag smoke --tag pull-request
|
|
3516
|
-
agent-sdk eval --dir . --
|
|
3517
|
-
agent-sdk eval --dir . --
|
|
3510
|
+
agent-sdk eval --dir . --list # discover only
|
|
3511
|
+
agent-sdk eval --dir . # run all
|
|
3512
|
+
agent-sdk eval --dir . builds/checkout # one datapoint
|
|
3513
|
+
agent-sdk eval --dir . builds search # several ids or prefixes
|
|
3514
|
+
agent-sdk eval --dir . --tag smoke --tag pull-request # any matching tag
|
|
3515
|
+
agent-sdk eval --dir . --verbose # t.log lines + reply snippets
|
|
3516
|
+
agent-sdk eval --dir . --timeout-ms 300000 # override the per-case timeout
|
|
3518
3517
|
```
|
|
3519
3518
|
|
|
3520
3519
|
Id filters use OR semantics. Each filter selects an exact id and its
|
|
3521
|
-
descendants
|
|
3522
|
-
|
|
3523
|
-
|
|
3524
|
-
match both groups.
|
|
3520
|
+
descendants: `builds` selects `builds`, `builds/checkout`, and every
|
|
3521
|
+
other case below that path. Repeated tags also use OR. With both ids
|
|
3522
|
+
and tags, a case must match both groups.
|
|
3525
3523
|
|
|
3526
|
-
`eval` boots
|
|
3527
|
-
|
|
3528
|
-
|
|
3529
|
-
|
|
3524
|
+
By default `eval` boots a throwaway server with its own state root, so
|
|
3525
|
+
cases don't inherit your checkout's `AGENTS.md` and session state stays
|
|
3526
|
+
out of the project. Artifacts still land in the project state
|
|
3527
|
+
directory; see [Where results land](#where-results-land). `--slug`
|
|
3528
|
+
picks the target in a multi-agent directory.
|
|
3529
|
+
|
|
3530
|
+
`--url` runs the batch on a running server instead, the same way
|
|
3531
|
+
`--prod` does: that server discovers its own `evals/`, results land in
|
|
3532
|
+
its playground history, and the local-only flags (`--junit`,
|
|
3533
|
+
`--artifacts`, `--max-concurrency`, `--skip-report`) do not apply. See
|
|
3534
|
+
[Run evals in the playground or on a deployment](#run-evals-in-the-playground-or-on-a-deployment).
|
|
3530
3535
|
|
|
3531
3536
|
```bash
|
|
3532
|
-
agent-sdk eval --
|
|
3533
|
-
--url http://127.0.0.1:3000/weather-agent \
|
|
3537
|
+
agent-sdk eval --url http://127.0.0.1:3000/weather-agent \
|
|
3534
3538
|
--bearer-token "$AGENT_TOKEN"
|
|
3535
3539
|
```
|
|
3536
3540
|
|
|
3537
|
-
|
|
3538
|
-
|
|
3539
|
-
|
|
3540
|
-
|
|
3541
|
-
|
|
3542
|
-
|
|
3543
|
-
|
|
3544
|
-
|
|
3541
|
+
See [CLI: eval](/docs/reference/cli.md#eval) for every flag.
|
|
3542
|
+
|
|
3543
|
+
### Credentials
|
|
3544
|
+
|
|
3545
|
+
Model turns need a
|
|
3546
|
+
[Cursor credential](/docs/reference/cli.md#environment-variables):
|
|
3547
|
+
`CURSOR_API_KEY`, `CURSOR_SERVICE_ACCOUNT_KEY`, or a saved
|
|
3548
|
+
`agent-sdk login`. The judge uses the same one. `eval --list` needs
|
|
3549
|
+
none.
|
|
3545
3550
|
|
|
3546
|
-
|
|
3547
|
-
`CURSOR_API_KEY_FILE` (hosted default `/run/cursor/secrets/CURSOR_API_KEY`), then
|
|
3548
|
-
`CURSOR_SERVICE_ACCOUNT_KEY`, then `agent-sdk login`.
|
|
3551
|
+
### Where results land
|
|
3549
3552
|
|
|
3550
|
-
|
|
3553
|
+
Every local run writes artifacts to a timestamped directory under
|
|
3554
|
+
`evals/` in the project state directory, whatever `--state-root` says.
|
|
3555
|
+
`--artifacts <dir>` chooses the path and `--no-artifacts` skips them.
|
|
3556
|
+
The directory holds `summary.json`,
|
|
3557
|
+
`results.jsonl`, and `evals/<case-id>.json` with every assertion, the
|
|
3558
|
+
inputs, tool calls with arguments and output, the final text, and
|
|
3559
|
+
`t.log` lines. Start there when a case fails. `--out <file>` also
|
|
3560
|
+
writes the full results JSON to a path of your choice.
|
|
3551
3561
|
|
|
3552
|
-
|
|
3562
|
+
The artifact does not include the session's event stream. Pass
|
|
3563
|
+
`--state-root <path>` to keep the ephemeral server's
|
|
3564
|
+
[session data](/docs/reference/sessions.md#where-does-the-agent-sdk-store-session-data)
|
|
3565
|
+
on disk when you need the raw events.
|
|
3553
3566
|
|
|
3554
|
-
|
|
3555
|
-
|
|
3567
|
+
## Run evals in CI
|
|
3568
|
+
|
|
3569
|
+
Run the suite non-interactively, write JUnit for the CI annotations,
|
|
3570
|
+
and fail the job on a red gate:
|
|
3571
|
+
|
|
3572
|
+
```bash
|
|
3573
|
+
# CURSOR_API_KEY comes from the CI secret store
|
|
3574
|
+
agent-sdk eval --dir . --json --no-stream \
|
|
3575
|
+
--junit reports/evals.xml \
|
|
3576
|
+
--artifacts reports/evals \
|
|
3577
|
+
> reports/evals.json
|
|
3578
|
+
```
|
|
3579
|
+
|
|
3580
|
+
The exit code follows the
|
|
3581
|
+
[verdict table](#gates-soft-scores-and-verdicts); `2` means nothing
|
|
3582
|
+
matched the selection. `--max-concurrency` overrides the project
|
|
3583
|
+
setting, for example to run lower on a shared runner.
|
|
3584
|
+
|
|
3585
|
+
The JSON on stdout carries the totals and one result per case:
|
|
3556
3586
|
|
|
3557
3587
|
```json
|
|
3558
3588
|
{
|
|
3559
3589
|
"ok": true,
|
|
3560
3590
|
"passed": 1,
|
|
3561
3591
|
"failed": 0,
|
|
3592
|
+
"scored": 0,
|
|
3593
|
+
"skipped": 0,
|
|
3594
|
+
"strict": false,
|
|
3595
|
+
"artifactsDir": "/work/my-agent/reports/evals",
|
|
3562
3596
|
"results": [
|
|
3563
3597
|
{
|
|
3564
3598
|
"id": "readiness",
|
|
3599
|
+
"verdict": "passed",
|
|
3565
3600
|
"ok": true,
|
|
3566
|
-
"assertions": [
|
|
3601
|
+
"assertions": [
|
|
3602
|
+
{ "name": "succeeded", "passed": true },
|
|
3603
|
+
{ "name": "calledTool(inspect_pr)", "passed": true }
|
|
3604
|
+
],
|
|
3567
3605
|
"sessionId": "ses_123",
|
|
3568
|
-
"inputs": ["Is checkout
|
|
3606
|
+
"inputs": ["Is https://github.com/acme/checkout/pull/42 ready to approve?"],
|
|
3569
3607
|
"toolCalls": [{ "toolName": "inspect_pr", "isError": false }],
|
|
3608
|
+
"metrics": {},
|
|
3570
3609
|
"logs": [],
|
|
3571
3610
|
"durationMs": 12340
|
|
3572
3611
|
}
|
|
@@ -3574,126 +3613,64 @@ the totals and one result per case:
|
|
|
3574
3613
|
}
|
|
3575
3614
|
```
|
|
3576
3615
|
|
|
3577
|
-
Each
|
|
3578
|
-
`error`, and tool
|
|
3579
|
-
|
|
3616
|
+
Each result can also include `description`, `finalText`, `tools`,
|
|
3617
|
+
`error`, `skipReason`, `metadata`, `tags`, and tool `args` and `output`.
|
|
3618
|
+
A soft miss shows as `"severity": "soft"` with `score` and `threshold`
|
|
3619
|
+
on the assertion. This shape lets CI report the failed assertion
|
|
3620
|
+
without parsing terminal text.
|
|
3621
|
+
|
|
3622
|
+
Keep CI green without weakening gates:
|
|
3580
3623
|
|
|
3581
|
-
|
|
3624
|
+
- Run `--tag smoke` on every push and the full suite on a schedule.
|
|
3625
|
+
- For probabilistic behavior, use `iterations` and a soft bar instead
|
|
3626
|
+
of one hard gate.
|
|
3582
3627
|
|
|
3583
|
-
|
|
3584
|
-
|
|
3585
|
-
|
|
3628
|
+
## Run evals in the playground or on a deployment
|
|
3629
|
+
|
|
3630
|
+
Start the server, open the playground, and choose **Evals**. Run every
|
|
3631
|
+
case or one case, watch progress, and open the resulting session trace.
|
|
3632
|
+
Playground runs target the live server instead of an ephemeral one, so
|
|
3633
|
+
their sessions appear in the session list. One batch runs at a time.
|
|
3586
3634
|
|
|
3587
3635
|
```bash
|
|
3588
3636
|
agent-sdk serve --dir .
|
|
3589
3637
|
```
|
|
3590
3638
|
|
|
3591
|
-
|
|
3592
|
-
|
|
3593
|
-
|
|
3594
|
-
[Configure eval runs](#configure-eval-runs). See
|
|
3595
|
-
[Playground eval routes](/docs/reference/http-api.md#playground-eval-routes).
|
|
3596
|
-
The start request returns `202` while cases run in the background.
|
|
3597
|
-
Poll until the snapshot status becomes `completed`, `failed`, or `cancelled`.
|
|
3598
|
-
Configuration errors appear on a failed snapshot.
|
|
3599
|
-
|
|
3600
|
-
On `--prod` / `--url`, the CLI prints the Eval ID as soon as the batch is
|
|
3601
|
-
accepted (and a Playground deep link with `?view=evals&evalRunId=…`):
|
|
3639
|
+
`--prod` (or `--url`) starts the same server-side batch on the team's
|
|
3640
|
+
hosted deployment (or the server you name), so results land in that
|
|
3641
|
+
server's playground history:
|
|
3602
3642
|
|
|
3603
3643
|
```bash
|
|
3604
3644
|
agent-sdk eval --prod --slug vulnerability-scanner --tag deepsec
|
|
3605
|
-
# Eval ID:
|
|
3606
|
-
#
|
|
3607
|
-
|
|
3608
|
-
|
|
3609
|
-
agent-sdk eval cancel evalrun_… --prod --slug vulnerability-scanner
|
|
3610
|
-
agent-sdk eval status evalrun_… --prod --slug vulnerability-scanner
|
|
3645
|
+
# Eval ID: <evalId>
|
|
3646
|
+
# Playground: <deployment>/playground?view=evals&evalRunId=<evalId>
|
|
3647
|
+
agent-sdk eval status <evalId> --prod --slug vulnerability-scanner
|
|
3648
|
+
agent-sdk eval cancel <evalId> --prod --slug vulnerability-scanner
|
|
3611
3649
|
```
|
|
3612
3650
|
|
|
3613
|
-
|
|
3614
|
-
|
|
3615
|
-
|
|
3616
|
-
|
|
3617
|
-
|
|
3618
|
-
|
|
3619
|
-
1. `t.succeeded()`: always, first.
|
|
3620
|
-
2. The tool decision: `calledTool` for the intended path,
|
|
3621
|
-
`notCalledTool` for the likely wrong alternative. The pair is
|
|
3622
|
-
stronger than either alone.
|
|
3623
|
-
3. Output shape: a regex for the contract (`/ready|blocked/i`, a JSON
|
|
3624
|
-
marker, a findings-block fence), never exact sentences.
|
|
3625
|
-
4. For structured output, parse `t.reply` and check fields with
|
|
3626
|
-
`satisfies` instead of substring-matching JSON.
|
|
3627
|
-
|
|
3628
|
-
The common failure modes: asserting exact phrasing, packing more than
|
|
3629
|
-
about five gates into one case (split it), and cases that depend on live
|
|
3630
|
-
external state that drifts (pin the input; see fixtures).
|
|
3631
|
-
|
|
3632
|
-
## Pick fixtures by agent type
|
|
3633
|
-
|
|
3634
|
-
The right fixture depends on the surface under test.
|
|
3635
|
-
|
|
3636
|
-
| Agent surface | Fixture |
|
|
3637
|
-
| --- | --- |
|
|
3638
|
-
| Chat / domain assistant | A canonical prompt string, chosen once and frozen |
|
|
3639
|
-
| Tool-heavy | Run `agent-sdk call <tool>` first to pin what the tool returns, then freeze the prompt that triggers it |
|
|
3640
|
-
| GitHub webhook | `agent-sdk github replay <pr> --events '*' --dry-run --out fixtures/github` snapshots real payloads for offline replay ([GitHub guide](/docs/guides/github.md)) |
|
|
3641
|
-
| PR reviewer with host preparation | Diff, metadata, and gold labels pinned to commit SHAs; keep any live PR matrix small |
|
|
3642
|
-
| Workspace-dependent | `workspaceFiles` in `t.send` options, never developer-machine paths |
|
|
3643
|
-
|
|
3644
|
-
Tag the fast, reliably passing core `smoke` and run `--tag smoke` in the
|
|
3645
|
-
inner loop. Leave slow or flaky-prone cases untagged for explicit runs.
|
|
3646
|
-
|
|
3647
|
-
### Materialize API-backed fixtures
|
|
3648
|
-
|
|
3649
|
-
An input that only points at external data, such as a pull request URL,
|
|
3650
|
-
snapshot id, or pair of commit SHAs, is not self-contained. Fetch it
|
|
3651
|
-
once and commit the rendered fixture before you expand the suite.
|
|
3652
|
-
|
|
3653
|
-
1. Save the diff, metadata, and labels under `fixtures/` at pinned
|
|
3654
|
-
revisions.
|
|
3655
|
-
2. Seed those files with `workspaceFiles`, or read them from the fixture
|
|
3656
|
-
directory.
|
|
3657
|
-
3. Assert decisions and output shape against the saved evidence.
|
|
3658
|
-
4. Keep a small `smoke` subset for any remaining live pipeline checks.
|
|
3659
|
-
|
|
3660
|
-
Read committed fixtures with `@cursor/july/evals/loaders`: `loadJson`,
|
|
3661
|
-
`loadJsonl`, and `loadYaml` resolve relative paths against the project
|
|
3662
|
-
root the runner discovered, not the cwd the CLI was invoked from
|
|
3663
|
-
(`resolveFixturePath` and `evalFixtureRoot` expose the same
|
|
3664
|
-
resolution for other file formats).
|
|
3665
|
-
|
|
3666
|
-
`maxConcurrency` limits parallel datapoints. It does not limit model or
|
|
3667
|
-
API fan-out inside one datapoint. Materialized fixtures prevent a large
|
|
3668
|
-
suite from exhausting provider and GitHub rate limits. The
|
|
3669
|
-
[evals skill](/docs/skills/evals.md) has the full fixture workflow.
|
|
3651
|
+
The CLI prints the Eval ID as soon as the batch is accepted. Pass
|
|
3652
|
+
`--no-wait` to return right away and poll with `eval status` later; it
|
|
3653
|
+
exits `3` while the batch is still running. Hosted history follows
|
|
3654
|
+
`maxPlaygroundRuns` and the persistence rule under
|
|
3655
|
+
[Configure eval runs](#configure-eval-runs). The HTTP surface is under
|
|
3656
|
+
[Playground eval routes](/docs/reference/http-api.md#playground-eval-routes).
|
|
3670
3657
|
|
|
3671
3658
|
## Keep improvements with regression evals
|
|
3672
3659
|
|
|
3673
|
-
Every [hillclimb](/docs/hillclimbing.md) round that keeps a change must
|
|
3674
|
-
an eval that would have failed before the change. If you can't
|
|
3675
|
-
the improvement as a gate (a `calledTool` shift, a bounded
|
|
3676
|
-
`
|
|
3677
|
-
|
|
3660
|
+
Every [hillclimb](/docs/hillclimbing.md) round that keeps a change must
|
|
3661
|
+
land an eval that would have failed before the change. If you can't
|
|
3662
|
+
express the improvement as a gate (a `calledTool` shift, a bounded
|
|
3663
|
+
`maxToolCalls`, an output-shape check), the improvement is unverified,
|
|
3664
|
+
and it'll regress silently.
|
|
3678
3665
|
|
|
3679
|
-
The rule cuts the other way too: never weaken an existing gate to make
|
|
3680
|
-
round pass. That's the freeze line moving, and it turns your
|
|
3681
|
-
suite into a list of checks that no longer protect anything.
|
|
3682
|
-
|
|
3683
|
-
## Compare variants on live traffic
|
|
3684
|
-
|
|
3685
|
-
Use `defineAB` to compare variant metrics on live sessions. It is not a
|
|
3686
|
-
test runner and has no `agent-sdk ab` command. Keep `defineEval` as the
|
|
3687
|
-
regression ratchet. Eval sessions do not enroll or change live metrics.
|
|
3688
|
-
See [Live A/B metrics](/docs/ab.md) for assignment, behavior, collection,
|
|
3689
|
-
and inspection.
|
|
3666
|
+
The rule cuts the other way too: never weaken an existing gate to make
|
|
3667
|
+
a round pass. That's the freeze line moving, and it turns your
|
|
3668
|
+
regression suite into a list of checks that no longer protect anything.
|
|
3690
3669
|
|
|
3691
3670
|
## What's next
|
|
3692
3671
|
|
|
3693
3672
|
Continue with these pages:
|
|
3694
3673
|
|
|
3695
|
-
- [Live A/B metrics](/docs/ab.md): sticky variants and cumulative metrics
|
|
3696
|
-
on live sessions
|
|
3697
3674
|
- [Hillclimbing](/docs/hillclimbing.md): the loop evals make trustworthy
|
|
3698
3675
|
- [Building agents with agents](/docs/building-with-agents.md): have a
|
|
3699
3676
|
coding agent write the first suite
|
|
@@ -3822,6 +3799,95 @@ Continue with these pages:
|
|
|
3822
3799
|
|
|
3823
3800
|
---
|
|
3824
3801
|
|
|
3802
|
+
Source: /docs/guides/bitbucket.md
|
|
3803
|
+
|
|
3804
|
+
# Bitbucket agents
|
|
3805
|
+
|
|
3806
|
+
Use `bitbucketChannel()` for Bitbucket Cloud and Bitbucket Data Center. The
|
|
3807
|
+
channel detects the source from each signed payload and normalizes Data Center
|
|
3808
|
+
events to the Bitbucket Cloud event vocabulary.
|
|
3809
|
+
|
|
3810
|
+
## Define the channel
|
|
3811
|
+
|
|
3812
|
+
Author `agent/channels/bitbucket.ts`:
|
|
3813
|
+
|
|
3814
|
+
```ts
|
|
3815
|
+
import { bitbucketChannel } from "@cursor/july/channels/bitbucket";
|
|
3816
|
+
|
|
3817
|
+
export default bitbucketChannel({});
|
|
3818
|
+
```
|
|
3819
|
+
|
|
3820
|
+
Without `onPullRequest`, new pull requests start turns. Data Center updates
|
|
3821
|
+
also start turns when the source branch has new commits. Comments and pushes
|
|
3822
|
+
are opt-in through `onPullRequestComment` and `onPush`. Use `onEvent` for
|
|
3823
|
+
other event types. Set `webhookEvents` when managed event delivery must
|
|
3824
|
+
subscribe to an event that only `onEvent` handles.
|
|
3825
|
+
|
|
3826
|
+
Filter comment hooks by author before starting a turn. This prevents comments
|
|
3827
|
+
posted by the agent from triggering another turn.
|
|
3828
|
+
|
|
3829
|
+
`ctx.bitbucket.api` can read pull requests, post comments, and create build
|
|
3830
|
+
statuses. See the [Channels reference](/docs/reference/channels.md) for hook
|
|
3831
|
+
return values and session behavior.
|
|
3832
|
+
|
|
3833
|
+
## Connect Bitbucket
|
|
3834
|
+
|
|
3835
|
+
Set a repository hook secret and API token:
|
|
3836
|
+
|
|
3837
|
+
```bash
|
|
3838
|
+
BITBUCKET_WEBHOOK_SECRET=...
|
|
3839
|
+
BITBUCKET_TOKEN=...
|
|
3840
|
+
```
|
|
3841
|
+
|
|
3842
|
+
Add a repository webhook for
|
|
3843
|
+
`https://<your-host>/<slug>/v1/channels/bitbucket`. Use the same secret on
|
|
3844
|
+
both sides. The channel verifies the `X-Hub-Signature` HMAC before it parses
|
|
3845
|
+
the payload.
|
|
3846
|
+
|
|
3847
|
+
Bitbucket Cloud uses its 2.0 API by default. Data Center also needs its REST
|
|
3848
|
+
API base:
|
|
3849
|
+
|
|
3850
|
+
```ts
|
|
3851
|
+
export default bitbucketChannel({
|
|
3852
|
+
apiBaseUrl: "https://bitbucket.example.com/rest/api/1.0",
|
|
3853
|
+
repos: ["PLATFORM/api"],
|
|
3854
|
+
});
|
|
3855
|
+
```
|
|
3856
|
+
|
|
3857
|
+
You can set `BITBUCKET_API_BASE_URL` instead.
|
|
3858
|
+
|
|
3859
|
+
## Test locally
|
|
3860
|
+
|
|
3861
|
+
Start the agent, then replay a pull request you can read:
|
|
3862
|
+
|
|
3863
|
+
```bash
|
|
3864
|
+
agent-sdk dev
|
|
3865
|
+
agent-sdk bitbucket replay \
|
|
3866
|
+
https://bitbucket.example.com/projects/PLATFORM/repos/api/pull-requests/42
|
|
3867
|
+
```
|
|
3868
|
+
|
|
3869
|
+
Replay supports Bitbucket Cloud and Data Center URLs. `BITBUCKET_TOKEN` needs
|
|
3870
|
+
pull request read access. The command reads the pull request, creates a payload
|
|
3871
|
+
in the matching dialect, and sends it through the same channel route. Use
|
|
3872
|
+
`--events '*'` to replay the supported events declared by the channel. Use
|
|
3873
|
+
`--dry-run --out fixtures/bitbucket` to save fixtures.
|
|
3874
|
+
|
|
3875
|
+
```bash
|
|
3876
|
+
agent-sdk bitbucket events --dir .
|
|
3877
|
+
agent-sdk bitbucket forward --dir .
|
|
3878
|
+
```
|
|
3879
|
+
|
|
3880
|
+
`forward` prints the repository-hook and HTTPS tunnel setup for live
|
|
3881
|
+
deliveries.
|
|
3882
|
+
|
|
3883
|
+
## Related
|
|
3884
|
+
|
|
3885
|
+
- [Channels reference](/docs/reference/channels.md)
|
|
3886
|
+
- [Webhooks and custom channels](/docs/guides/webhooks.md)
|
|
3887
|
+
- [Evals](/docs/evals.md)
|
|
3888
|
+
|
|
3889
|
+
---
|
|
3890
|
+
|
|
3825
3891
|
Source: /docs/guides/cloud-agents.md
|
|
3826
3892
|
|
|
3827
3893
|
# Cursor cloud agents
|
|
@@ -4173,6 +4239,37 @@ snapshot in the wake.
|
|
|
4173
4239
|
This is the preferred production path: no public URL, no repo admin
|
|
4174
4240
|
webhook, and no inbound network for GitHub deliveries.
|
|
4175
4241
|
|
|
4242
|
+
## Connect GitHub Enterprise Server
|
|
4243
|
+
|
|
4244
|
+
GitHub Enterprise Server uses the same `githubChannel()` hooks and normalized
|
|
4245
|
+
events. Connect it through the direct webhook route. Set the REST API base and
|
|
4246
|
+
credentials for your server:
|
|
4247
|
+
|
|
4248
|
+
```ts
|
|
4249
|
+
export default githubChannel({
|
|
4250
|
+
api: {
|
|
4251
|
+
apiBaseUrl: "https://github.example.com/api/v3",
|
|
4252
|
+
},
|
|
4253
|
+
credentials: {
|
|
4254
|
+
token: () => process.env.GITHUB_TOKEN,
|
|
4255
|
+
webhookSecret: () => process.env.GITHUB_WEBHOOK_SECRET,
|
|
4256
|
+
},
|
|
4257
|
+
onPullRequest: (ctx, pr) =>
|
|
4258
|
+
pr.action === "opened" ? { auth: defaultGitHubAuth(ctx) } : null,
|
|
4259
|
+
});
|
|
4260
|
+
```
|
|
4261
|
+
|
|
4262
|
+
Add a repository webhook for
|
|
4263
|
+
`https://<your-host>/<slug>/v1/channels/github`. Use the same webhook secret
|
|
4264
|
+
on the server and in `GITHUB_WEBHOOK_SECRET`. The channel verifies
|
|
4265
|
+
`X-Hub-Signature-256` before it parses the payload.
|
|
4266
|
+
|
|
4267
|
+
Cursor account event pull, `github replay`, and `github forward` target
|
|
4268
|
+
GitHub.com. Test Enterprise Server integrations by posting saved webhook
|
|
4269
|
+
fixtures to a local `--dev` server. Leave `GITHUB_WEBHOOK_SECRET` unset for
|
|
4270
|
+
this local test so the channel admits unsigned loopback deliveries. Use a real
|
|
4271
|
+
delivery from your server so the fixture matches its version.
|
|
4272
|
+
|
|
4176
4273
|
## Define the channel
|
|
4177
4274
|
|
|
4178
4275
|
Author `agent/channels/github.ts` with `githubChannel()` from
|
|
@@ -4380,6 +4477,103 @@ key. Handlers you author replace the matching defaults (same as
|
|
|
4380
4477
|
|
|
4381
4478
|
---
|
|
4382
4479
|
|
|
4480
|
+
Source: /docs/guides/gitlab.md
|
|
4481
|
+
|
|
4482
|
+
# GitLab agents
|
|
4483
|
+
|
|
4484
|
+
Use `gitlabChannel()` for GitLab.com and self-managed GitLab. The channel
|
|
4485
|
+
verifies project hooks, normalizes their payloads, and gives each hook a
|
|
4486
|
+
project-bound `ctx.gitlab` API client.
|
|
4487
|
+
|
|
4488
|
+
## Define the channel
|
|
4489
|
+
|
|
4490
|
+
Author `agent/channels/gitlab.ts`:
|
|
4491
|
+
|
|
4492
|
+
```ts
|
|
4493
|
+
import { gitlabChannel } from "@cursor/july/channels/gitlab";
|
|
4494
|
+
|
|
4495
|
+
export default gitlabChannel({});
|
|
4496
|
+
```
|
|
4497
|
+
|
|
4498
|
+
Without `onMergeRequest`, new and reopened merge requests start turns.
|
|
4499
|
+
Updates start turns only when they include new commits. Notes, pushes, and
|
|
4500
|
+
pipelines are opt-in through `onNote`, `onPush`, and `onPipeline`.
|
|
4501
|
+
Use `onEvent` for other GitLab event types. Set `webhookEvents` when managed
|
|
4502
|
+
event delivery must subscribe to an event that only `onEvent` handles.
|
|
4503
|
+
|
|
4504
|
+
Filter note hooks by author before starting a turn. This prevents comments
|
|
4505
|
+
posted by the agent from triggering another turn.
|
|
4506
|
+
|
|
4507
|
+
`ctx.gitlab` can call the project REST API and create commit statuses. See the
|
|
4508
|
+
[Channels reference](/docs/reference/channels.md) for hook return values and
|
|
4509
|
+
session behavior.
|
|
4510
|
+
|
|
4511
|
+
## Connect GitLab
|
|
4512
|
+
|
|
4513
|
+
On Cursor-managed hosting, use the signed-in Cursor account:
|
|
4514
|
+
|
|
4515
|
+
```ts
|
|
4516
|
+
export default gitlabChannel({
|
|
4517
|
+
cursorAccount: {
|
|
4518
|
+
projects: ["acme/platform"],
|
|
4519
|
+
},
|
|
4520
|
+
});
|
|
4521
|
+
```
|
|
4522
|
+
|
|
4523
|
+
For direct webhooks, set a project hook secret and API token:
|
|
4524
|
+
|
|
4525
|
+
```bash
|
|
4526
|
+
GITLAB_WEBHOOK_SECRET=...
|
|
4527
|
+
GITLAB_TOKEN=...
|
|
4528
|
+
```
|
|
4529
|
+
|
|
4530
|
+
Add a project webhook for
|
|
4531
|
+
`https://<your-host>/<slug>/v1/channels/gitlab`. Use the same value for the
|
|
4532
|
+
GitLab secret token and `GITLAB_WEBHOOK_SECRET`. The channel checks
|
|
4533
|
+
`X-Gitlab-Token` before it parses the payload.
|
|
4534
|
+
|
|
4535
|
+
Self-managed GitLab also needs its REST API base:
|
|
4536
|
+
|
|
4537
|
+
```ts
|
|
4538
|
+
export default gitlabChannel({
|
|
4539
|
+
apiBaseUrl: "https://gitlab.example.com/api/v4",
|
|
4540
|
+
projects: ["acme/platform"],
|
|
4541
|
+
});
|
|
4542
|
+
```
|
|
4543
|
+
|
|
4544
|
+
You can set `GITLAB_API_BASE_URL` instead.
|
|
4545
|
+
|
|
4546
|
+
## Test locally
|
|
4547
|
+
|
|
4548
|
+
Start the agent, then replay a merge request you can read:
|
|
4549
|
+
|
|
4550
|
+
```bash
|
|
4551
|
+
agent-sdk dev
|
|
4552
|
+
agent-sdk gitlab replay \
|
|
4553
|
+
https://gitlab.example.com/acme/platform/-/merge_requests/42
|
|
4554
|
+
```
|
|
4555
|
+
|
|
4556
|
+
`GITLAB_TOKEN` needs API read access. Replay reads the merge request,
|
|
4557
|
+
creates GitLab-shaped payloads, and sends them through the same channel route.
|
|
4558
|
+
Use `--events '*'` to replay the supported events declared by the channel.
|
|
4559
|
+
Use `--dry-run --out fixtures/gitlab` to save fixtures.
|
|
4560
|
+
|
|
4561
|
+
```bash
|
|
4562
|
+
agent-sdk gitlab events --dir .
|
|
4563
|
+
agent-sdk gitlab forward --dir .
|
|
4564
|
+
```
|
|
4565
|
+
|
|
4566
|
+
GitLab has no local webhook relay. `forward` prints the project-hook and HTTPS
|
|
4567
|
+
tunnel setup for live deliveries.
|
|
4568
|
+
|
|
4569
|
+
## Related
|
|
4570
|
+
|
|
4571
|
+
- [Channels reference](/docs/reference/channels.md)
|
|
4572
|
+
- [Webhooks and custom channels](/docs/guides/webhooks.md)
|
|
4573
|
+
- [Evals](/docs/evals.md)
|
|
4574
|
+
|
|
4575
|
+
---
|
|
4576
|
+
|
|
4383
4577
|
Source: /docs/guides/grokbot-agents.md
|
|
4384
4578
|
|
|
4385
4579
|
# Cursor Grok Bot agents
|
|
@@ -5395,7 +5589,8 @@ A custom channel gives the agent its own HTTP surface. You get routes
|
|
|
5395
5589
|
with validated payloads, sessions keyed to something in your domain (a
|
|
5396
5590
|
thread, a ticket, a PR), and replies delivered back to the caller. The
|
|
5397
5591
|
[Slack](/docs/guides/slack.md) and [GitHub](/docs/guides/github.md) packs build on this
|
|
5398
|
-
mechanism.
|
|
5592
|
+
mechanism. The [GitLab](/docs/guides/gitlab.md) and [Bitbucket](/docs/guides/bitbucket.md) packs
|
|
5593
|
+
use it too. This page is the mechanism itself.
|
|
5399
5594
|
|
|
5400
5595
|
## What you already have
|
|
5401
5596
|
|
|
@@ -5849,7 +6044,8 @@ from any PR you can read. See the [GitHub guide](/docs/guides/github.md).
|
|
|
5849
6044
|
Continue with these pages:
|
|
5850
6045
|
|
|
5851
6046
|
- [Channels reference](/docs/reference/channels.md): the full authoring API
|
|
5852
|
-
- [GitHub](/docs/guides/github.md)
|
|
6047
|
+
- [GitHub](/docs/guides/github.md), [GitLab](/docs/guides/gitlab.md),
|
|
6048
|
+
[Bitbucket](/docs/guides/bitbucket.md), and [Slack](/docs/guides/slack.md): the packaged channels
|
|
5853
6049
|
- [Sessions and streaming](/docs/reference/sessions.md): events your
|
|
5854
6050
|
channel can subscribe to
|
|
5855
6051
|
|
|
@@ -5926,9 +6122,9 @@ Pin the input first. A moving fixture is noise. For GitHub agents, use `agent-sd
|
|
|
5926
6122
|
|
|
5927
6123
|
## How do I lock a hillclimb improvement with an eval?
|
|
5928
6124
|
|
|
5929
|
-
Every kept change needs an eval that would have failed before the change: a tool-choice gate,
|
|
6125
|
+
Every kept change needs an eval that would have failed before the change: a tool-choice gate, a `maxToolCalls` bound, or an output-shape check. Run `agent-sdk eval --dir . --json` between rounds. Never weaken an existing gate to pass the round.
|
|
5930
6126
|
|
|
5931
|
-
Details live in [
|
|
6127
|
+
Details live in [Keep improvements with regression evals](/docs/evals.md#keep-improvements-with-regression-evals). The evals skill will author the case with you.
|
|
5932
6128
|
|
|
5933
6129
|
## What habits help hillclimbing stay reliable?
|
|
5934
6130
|
|
|
@@ -5983,14 +6179,16 @@ npx @cursor/july docs
|
|
|
5983
6179
|
| Turning a Cursor Automation into a project | [Convert a Cursor Automation](/docs/guides/convert-automation.md) |
|
|
5984
6180
|
| Wiring an agent to Slack | [Slack guide](/docs/guides/slack.md) |
|
|
5985
6181
|
| Starting from a packaged template | [Demo](/docs/templates/demo.md), [Grok Bot agents](/docs/templates/grokbot-agents.md), [Code wiki](/docs/templates/code-wiki.md), [Living AGENTS.md](/docs/templates/agents-md.md), [Security reviewer](/docs/templates/security-reviewer.md), [Security help](/docs/templates/security-help.md), [Triage](/docs/templates/triage.md), or [Agentic Owners](/docs/templates/agentic-owners.md) |
|
|
5986
|
-
| Wiring an agent to GitHub
|
|
6182
|
+
| Wiring an agent to GitHub or GitHub Enterprise Server | [GitHub guide](/docs/guides/github.md) |
|
|
6183
|
+
| Wiring an agent to GitLab | [GitLab guide](/docs/guides/gitlab.md) |
|
|
6184
|
+
| Wiring an agent to Bitbucket | [Bitbucket guide](/docs/guides/bitbucket.md) |
|
|
5987
6185
|
| Driving PRs from a cloud VM | [PR autofixer template](/docs/templates/pr-autofixer.md) |
|
|
5988
6186
|
| Handing coding work to Cursor cloud agents | [Cursor cloud agents](/docs/guides/cloud-agents.md) |
|
|
5989
6187
|
| Talking to your Grok Bot agents | [Cursor Grok Bot agents](/docs/guides/grokbot-agents.md) |
|
|
5990
6188
|
| Letting an agent change its own source | [Self-improvement](/docs/guides/improve.md) |
|
|
5991
6189
|
| Driving an agent from Linear (or another tracker) | [Webhooks guide: Linear example](/docs/guides/webhooks.md#example-linear-as-the-control-plane) |
|
|
5992
6190
|
| Making an existing agent measurably better | [Evals](/docs/evals.md), then [Hillclimbing](/docs/hillclimbing.md) |
|
|
5993
|
-
|
|
|
6191
|
+
| Gating agent behavior in CI | [Run evals in CI](/docs/evals.md#run-evals-in-ci) |
|
|
5994
6192
|
| Deploying with Cursor or on your own infrastructure | [Deployment](/docs/deployment.md) |
|
|
5995
6193
|
| Debugging something that misbehaves | [Fix common agent problems](/docs/troubleshooting.md) |
|
|
5996
6194
|
|
|
@@ -6031,10 +6229,8 @@ npx @cursor/july docs
|
|
|
6031
6229
|
|
|
6032
6230
|
- [Building agents with agents](/docs/building-with-agents.md): use a coding
|
|
6033
6231
|
agent to scaffold, run, and iterate on your agent.
|
|
6034
|
-
- [Evals](/docs/evals.md): author `defineEval` cases,
|
|
6035
|
-
|
|
6036
|
-
- [Live A/B metrics](/docs/ab.md): assign sticky variants and compare
|
|
6037
|
-
cumulative metrics on live sessions.
|
|
6232
|
+
- [Evals](/docs/evals.md): author `defineEval` cases, assert over the
|
|
6233
|
+
trajectory, and run them locally and in CI.
|
|
6038
6234
|
- [Storage](/docs/storage.md): point durable storage at a backend you own
|
|
6039
6235
|
with `defineStorage`.
|
|
6040
6236
|
- [Hillclimbing](/docs/hillclimbing.md): measure and improve an agent
|
|
@@ -6046,6 +6242,10 @@ npx @cursor/july docs
|
|
|
6046
6242
|
its own HTTP surface.
|
|
6047
6243
|
- [GitHub](/docs/guides/github.md): trigger the agent from pull requests,
|
|
6048
6244
|
CI, and comments.
|
|
6245
|
+
- [GitLab](/docs/guides/gitlab.md): trigger the agent from merge requests,
|
|
6246
|
+
notes, pipelines, and pushes.
|
|
6247
|
+
- [Bitbucket](/docs/guides/bitbucket.md): trigger the agent from pull requests,
|
|
6248
|
+
comments, and pushes.
|
|
6049
6249
|
- [Slack](/docs/guides/slack.md): put the agent in Slack over Socket Mode.
|
|
6050
6250
|
- [Human-in-the-loop approvals](/docs/guides/human-in-the-loop.md): park a
|
|
6051
6251
|
tool call until a person signs off.
|
|
@@ -6862,7 +7062,8 @@ Source: /docs/reference/channels.md
|
|
|
6862
7062
|
A channel is the surface an agent lives on. The built-in HTTP session
|
|
6863
7063
|
channel is always mounted. Custom channels declare their own routes
|
|
6864
7064
|
under `/v1/channels/<id>`. The Slack and GitHub packs are prebuilt
|
|
6865
|
-
channels with platform transports.
|
|
7065
|
+
channels with platform transports. GitLab and Bitbucket packs cover their
|
|
7066
|
+
hosted and self-managed products. This page is the authoring reference;
|
|
6866
7067
|
for the walkthrough, see the [Webhooks guide](/docs/guides/webhooks.md).
|
|
6867
7068
|
|
|
6868
7069
|
## Built-in HTTP channel
|
|
@@ -6988,7 +7189,7 @@ Handlers receive the Fetch `Request` and an args object:
|
|
|
6988
7189
|
`workspaceFiles`, `workspaceDir`, `cloud` (attach cloud repos for this
|
|
6989
7190
|
session), `auth` (defaults to the request principal), `state` (starting
|
|
6990
7191
|
channel state for new sessions), `title` (session display title), and
|
|
6991
|
-
`purpose` (`"eval"`
|
|
7192
|
+
`purpose` (`"eval"` marks the session as regression traffic).
|
|
6992
7193
|
|
|
6993
7194
|
## Events
|
|
6994
7195
|
|
|
@@ -7069,7 +7270,19 @@ model turn), `{ task }` (host work), or `null`, and CLI tooling for
|
|
|
7069
7270
|
replay and live forwarding. Author `agent/channels/github.ts` with
|
|
7070
7271
|
`githubChannel()`. Opt-in `progress.commitStatus` and `progress.banner`
|
|
7071
7272
|
converge a merge-box check and sticky PR comment from default stream
|
|
7072
|
-
events.
|
|
7273
|
+
events. Supports GitHub.com and GitHub Enterprise Server. Guide:
|
|
7274
|
+
[GitHub](/docs/guides/github.md).
|
|
7275
|
+
|
|
7276
|
+
**GitLab** (`@cursor/july/channels/gitlab`): verified project hooks for
|
|
7277
|
+
merge requests, notes, pipelines, pushes, and custom event types. Supports
|
|
7278
|
+
GitLab.com and self-managed GitLab. Author `agent/channels/gitlab.ts` with
|
|
7279
|
+
`gitlabChannel()`. Guide: [GitLab](/docs/guides/gitlab.md).
|
|
7280
|
+
|
|
7281
|
+
**Bitbucket** (`@cursor/july/channels/bitbucket`): verified repository hooks
|
|
7282
|
+
for pull requests, comments, pushes, and custom event types. Supports
|
|
7283
|
+
Bitbucket Cloud and Bitbucket Data Center through one normalized hook API.
|
|
7284
|
+
Author `agent/channels/bitbucket.ts` with `bitbucketChannel()`. Guide:
|
|
7285
|
+
[Bitbucket](/docs/guides/bitbucket.md).
|
|
7073
7286
|
|
|
7074
7287
|
**Deployments** (`@cursor/july/channels/deployments`): pull deploy
|
|
7075
7288
|
events. Declare `events` and handle each one in `onEvent`. Each event
|
|
@@ -7491,8 +7704,10 @@ agent-sdk eval status <evalId> --prod --slug pr-approver
|
|
|
7491
7704
|
agent-sdk eval cancel <evalId> --prod --slug pr-approver
|
|
7492
7705
|
```
|
|
7493
7706
|
|
|
7494
|
-
`eval` runs `evals/**/*.eval.{ts,js}` on an ephemeral server
|
|
7495
|
-
`--
|
|
7707
|
+
`eval` runs `evals/**/*.eval.{ts,js}` on an ephemeral server. With
|
|
7708
|
+
`--prod` or `--url`, the target server runs its own evals as a
|
|
7709
|
+
server-side batch. Select one or more exact case IDs, file ID prefixes,
|
|
7710
|
+
or tags.
|
|
7496
7711
|
Omit selectors to run all cases. Repeated `--tag` flags use OR matching.
|
|
7497
7712
|
|
|
7498
7713
|
An eval run requires `evals/evals.config.{ts,js}` with `maxConcurrency`
|
|
@@ -7503,13 +7718,13 @@ between 1 and 200. Timeout priority is the case's `timeoutMs`, the CLI's
|
|
|
7503
7718
|
| --- | --- |
|
|
7504
7719
|
| `--list` | Print discovered cases without running. `--list --json` prints them as an array. |
|
|
7505
7720
|
| `--tag <tag>` | Run cases with this tag. Repeated flags use OR matching. |
|
|
7506
|
-
| `--json` | Print `{ ok, passed, failed, results }
|
|
7721
|
+
| `--json` | Print `{ ok, passed, failed, scored, skipped, strict, results }`; see [Run evals in CI](/docs/evals.md#run-evals-in-ci) for the result shape. |
|
|
7507
7722
|
| `--verbose` | Stream `t.log` lines and reply snippets. |
|
|
7508
7723
|
| `--no-stream` | Hide live progress on stderr. |
|
|
7509
7724
|
| `--strict` | Exit `1` when a scored case misses a soft threshold. |
|
|
7510
7725
|
| `--max-concurrency <n>` | Override `maxConcurrency` from `evals.config.ts`. |
|
|
7511
7726
|
| `--junit <path>` | Write JUnit XML for CI annotations. |
|
|
7512
|
-
| `--artifacts <dir>` | Write run artifacts here. The default is a timestamped directory under
|
|
7727
|
+
| `--artifacts <dir>` | Write run artifacts here. The default is a timestamped directory under `evals/` in the project state directory (not affected by `--state-root`). |
|
|
7513
7728
|
| `--no-artifacts` | Skip run artifacts. |
|
|
7514
7729
|
| `--skip-report` | Ignore reporters from `evals.config.ts` and eval files. |
|
|
7515
7730
|
| `--out <path>` | Also write the full results JSON to this path (also for `eval status <evalId>`). |
|
|
@@ -8473,7 +8688,6 @@ namespace:
|
|
|
8473
8688
|
| `subagents/<id>/` | subagent `<ns>__<id>` |
|
|
8474
8689
|
| `instructions.md` / `.ts` / dir | appended to the agent's system prompt |
|
|
8475
8690
|
| `sandbox/workspace/**` | seeded into each local session workspace |
|
|
8476
|
-
| `ab.ts` / `ab/<name>.ts` | A/B experiment `<ns>__<name>` |
|
|
8477
8691
|
| `artifacts.ts` | artifact kinds `<ns>__<kind>` |
|
|
8478
8692
|
|
|
8479
8693
|
The root agent still needs its own `instructions.md`. Extension
|
|
@@ -8490,7 +8704,6 @@ or is outside discovery:
|
|
|
8490
8704
|
| `agent.ts` | One `defineAgent` runtime per agent |
|
|
8491
8705
|
| `storage.ts` | One `host.kv` / `host.files` backend |
|
|
8492
8706
|
| `otel.ts` | One OTLP exporter |
|
|
8493
|
-
| `ab.config.ts` | One experiment-platform config; `ab.ts` / `ab/` still merge |
|
|
8494
8707
|
| `playground/` | Custom chips are a Vite glob of the agent tree, not a discovery walk |
|
|
8495
8708
|
| `extensions/` | Nested mounts are not loaded |
|
|
8496
8709
|
| `sandbox.ts` | Custom sandbox backends stay on the agent |
|
|
@@ -8531,7 +8744,6 @@ it alive.
|
|
|
8531
8744
|
| `schedules/<name>.ts` | `disableSchedule()` |
|
|
8532
8745
|
| `subagents/<id>.ts` | `disableSubagent()` |
|
|
8533
8746
|
| `instructions.ts` | `disableInstructions()` |
|
|
8534
|
-
| `ab.ts` / `ab/<name>.ts` | `disableAB()` |
|
|
8535
8747
|
| `artifacts.ts` | `disableArtifacts()` |
|
|
8536
8748
|
|
|
8537
8749
|
`disable()` is the same brand as the slot helpers above and works in
|
|
@@ -8775,7 +8987,7 @@ exported from `@cursor/july`.
|
|
|
8775
8987
|
|
|
8776
8988
|
| Member | What it is |
|
|
8777
8989
|
| --- | --- |
|
|
8778
|
-
| `ctx.session` | Read-only session info: `id`, `channelId`, `mode` (`chat` or `task`), `purpose` (`live` or `eval`), `auth`, plus `title
|
|
8990
|
+
| `ctx.session` | Read-only session info: `id`, `channelId`, `mode` (`chat` or `task`), `purpose` (`live` or `eval`), `auth`, plus `title` and `sdkAgentId` when set |
|
|
8779
8991
|
| `ctx.agent` | `{ name }` of the agent the event belongs to |
|
|
8780
8992
|
| `ctx.channel` | `{ id, continuationToken }`. The token is `null` when the session can't take follow-ups |
|
|
8781
8993
|
| `ctx.host.kv` | Durable JSON, shared by every session of the agent; the [storage backend](/docs/storage.md#author-kv-ctx-host-kv) decides whether it survives a hosted replace. Prefix keys with `ctx.session.id` for per-session state |
|
|
@@ -8806,17 +9018,17 @@ Each event reaches a hook at most once. A restart doesn't replay the log
|
|
|
8806
9018
|
into hooks, so a mirror needs no dedupe, and the event log rather than
|
|
8807
9019
|
the hook's copy is the source of truth.
|
|
8808
9020
|
|
|
8809
|
-
## Hooks, channel events,
|
|
9021
|
+
## Hooks, channel events, or evals?
|
|
8810
9022
|
|
|
8811
9023
|
All of them consume the same stream, for different jobs:
|
|
8812
9024
|
|
|
8813
|
-
| | Hooks | Channel `events` | Evals |
|
|
8814
|
-
| --- | --- | --- | --- |
|
|
8815
|
-
| Scope | every session of the agent | sessions the channel owns | one test turn |
|
|
8816
|
-
| Job | observe: audit, metrics, mirrors, derived state | deliver: replies back to the channel's surface | assert: gates over the trajectory |
|
|
8817
|
-
| Context | `ctx.host`, `ctx.artifacts`, session info | `channel.state`, `setContinuationToken`, `ctx.host`, session info | the `t` assertion helpers |
|
|
8818
|
-
| Can affect the run | no | yes, it owns the surface | n/a |
|
|
8819
|
-
| Authored at | `agent/hooks/*.ts` | channel config | `evals/**/*.eval.ts` |
|
|
9025
|
+
| | Hooks | Channel `events` | Evals |
|
|
9026
|
+
| --- | --- | --- | --- |
|
|
9027
|
+
| Scope | every session of the agent | sessions the channel owns | one test turn |
|
|
9028
|
+
| Job | observe: audit, metrics, mirrors, derived state | deliver: replies back to the channel's surface | assert: gates over the trajectory |
|
|
9029
|
+
| Context | `ctx.host`, `ctx.artifacts`, session info | `channel.state`, `setContinuationToken`, `ctx.host`, session info | the `t` assertion helpers |
|
|
9030
|
+
| Can affect the run | no | yes, it owns the surface | n/a |
|
|
9031
|
+
| Authored at | `agent/hooks/*.ts` | channel config | `evals/**/*.eval.ts` |
|
|
8820
9032
|
|
|
8821
9033
|
## When not to use a hook
|
|
8822
9034
|
|
|
@@ -8828,7 +9040,6 @@ All of them consume the same stream, for different jobs:
|
|
|
8828
9040
|
| Block, approve, or rewrite a tool call | [`needsApproval`](/docs/reference/tools.md#gate-a-tool-on-human-approval) on the tool |
|
|
8829
9041
|
| Act on the final assistant text, or fail a bad turn | [`defineResult`](/docs/reference/result.md) |
|
|
8830
9042
|
| Gate a change on behavior | [Evals](/docs/evals.md) |
|
|
8831
|
-
| Compare two prompts on live traffic | [`defineAB`](/docs/ab.md) |
|
|
8832
9043
|
|
|
8833
9044
|
## Patterns
|
|
8834
9045
|
|
|
@@ -8954,8 +9165,6 @@ Continue with these pages:
|
|
|
8954
9165
|
- [Deployment](/docs/deployment.md#observability): runtime logs and export
|
|
8955
9166
|
paths
|
|
8956
9167
|
- [Channels](/docs/reference/channels.md#events): the delivery-side counterpart
|
|
8957
|
-
- [Live A/B metrics](/docs/ab.md): sticky variants over the same event
|
|
8958
|
-
stream
|
|
8959
9168
|
|
|
8960
9169
|
---
|
|
8961
9170
|
|
|
@@ -9110,12 +9319,11 @@ These read-only routes describe the running agent.
|
|
|
9110
9319
|
|
|
9111
9320
|
| Route | What it does |
|
|
9112
9321
|
| ---------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
9113
|
-
| `GET /v1/info` | The discovered surface: model, tools, skills, MCP connections, subagents, channels and routes (with schemas), schedules, hooks,
|
|
9322
|
+
| `GET /v1/info` | The discovered surface: model, tools, skills, MCP connections, subagents, channels and routes (with schemas), schedules, hooks, diagnostics |
|
|
9114
9323
|
| `GET /v1/tools` | The live tool catalog: authored server tools plus advertised MCP passthroughs under model-facing names, as light `{ name, title?, source? }` entries. `session` / `continuationToken` query parameters bind the listing to a session identity (advertised inventories can be tenant-scoped); a connection whose listing fails is skipped and reported in `connectionErrors` |
|
|
9115
9324
|
| `GET /v1/tools/:name` | One catalog tool's full description: description, execution, `needsApproval`, `effect`, input and output schemas, source connection. Same session binding as the listing; unknown names get `404` with the available names |
|
|
9116
9325
|
| `GET /v1/health` | Per-agent liveness, no auth |
|
|
9117
9326
|
| `GET /v1/logs?after=N` | Recent server log lines, with a polling cursor |
|
|
9118
|
-
| `GET /v1/abs` | [Live A/B metrics](/docs/ab.md): per-session assignments and aggregate arm totals |
|
|
9119
9327
|
|
|
9120
9328
|
## Artifacts
|
|
9121
9329
|
|
|
@@ -9174,7 +9382,7 @@ the final `completed` or `failed` status. Batch errors appear on the
|
|
|
9174
9382
|
snapshot returned by the poll. Entries within `filterIds` and `tags`
|
|
9175
9383
|
use OR semantics. When both fields are present, a case must match one
|
|
9176
9384
|
entry from each field. Listed runs persist across restarts when storage is configured; see
|
|
9177
|
-
[Storage](/docs/storage.md#eval-
|
|
9385
|
+
[Storage](/docs/storage.md#eval-table). Otherwise they are
|
|
9178
9386
|
process-memory only.
|
|
9179
9387
|
|
|
9180
9388
|
## Dev-mode routes
|
|
@@ -9336,8 +9544,6 @@ Use the playground to chat, try channel routes, and inspect sessions.
|
|
|
9336
9544
|
an [extension](/docs/reference/extensions.md) cannot contribute them.
|
|
9337
9545
|
- **Raw events pane**: flip it on to inspect the event stream.
|
|
9338
9546
|
- **Logs tab**: recent server log lines, polled from `GET /v1/logs`.
|
|
9339
|
-
- **A/Bs tab**: per-session and aggregate
|
|
9340
|
-
[live A/B metrics](/docs/ab.md) from `GET /v1/abs`.
|
|
9341
9547
|
|
|
9342
9548
|
In multi-agent mode each agent has its own playground at
|
|
9343
9549
|
`/<slug>/playground`, and `/` is an index of them all.
|
|
@@ -9360,8 +9566,6 @@ Continue with these pages:
|
|
|
9360
9566
|
- [Sessions and streaming](/docs/reference/sessions.md): the streams it renders
|
|
9361
9567
|
- [Human-in-the-loop](/docs/guides/human-in-the-loop.md): the approval
|
|
9362
9568
|
buttons in context
|
|
9363
|
-
- [Live A/B metrics](/docs/ab.md): the assignments and results in the A/Bs
|
|
9364
|
-
tab
|
|
9365
9569
|
|
|
9366
9570
|
---
|
|
9367
9571
|
|
|
@@ -9375,8 +9579,7 @@ how the Agent SDK loads it.
|
|
|
9375
9579
|
|
|
9376
9580
|
## Folder structure
|
|
9377
9581
|
|
|
9378
|
-
For the capabilities below, identity
|
|
9379
|
-
experiments can override their file-derived name.
|
|
9582
|
+
For the capabilities below, identity comes from the path.
|
|
9380
9583
|
|
|
9381
9584
|
| Path | Resolves to |
|
|
9382
9585
|
| --- | --- |
|
|
@@ -9388,8 +9591,6 @@ experiments can override their file-derived name.
|
|
|
9388
9591
|
| `agent/extensions/ci.ts` | extension mount `ci`; its contributions become `ci__<name>` |
|
|
9389
9592
|
| `agent/extensions/notion.ts` | Cursor plugin mount `notion` (`cursorPlugin`); its skills, agents, and MCP servers become `notion__<name>` |
|
|
9390
9593
|
| `agent/channels/drive.ts` | channel `drive`, routes under `/v1/channels/drive` |
|
|
9391
|
-
| `agent/ab.ts` | A/B experiment `ab` unless `name` overrides it |
|
|
9392
|
-
| `agent/ab/concise.ts` | A/B experiment `concise` unless `name` overrides it |
|
|
9393
9594
|
|
|
9394
9595
|
The root agent takes its name from `package.json` `name`, falling back
|
|
9395
9596
|
to the directory name. When serving multiple agents, the slug is the
|
|
@@ -9441,8 +9642,6 @@ Each path maps to a capability and a reference page.
|
|
|
9441
9642
|
| `agent/channels/*.ts` | HTTP surfaces beyond the built-in session API; `slack.ts` and `github.ts` use the platform packs | [Channels](/docs/reference/channels.md) |
|
|
9442
9643
|
| `agent/hooks/*.ts` | Observe-only event subscribers, never fatal | [Hooks](/docs/reference/hooks.md) |
|
|
9443
9644
|
| `agent/otel.ts` | `defineOtel` OTLP export (traces, metrics, optional logs) | [OpenTelemetry](/docs/guides/opentelemetry.md) |
|
|
9444
|
-
| `agent/ab.ts`, `agent/ab/*.ts` | `defineAB` experiments with sticky variants and live metrics | [Live A/B metrics](/docs/ab.md) |
|
|
9445
|
-
| `agent/ab.config.ts` | `defineABConfig` shared A/B settings | [Live A/B metrics](/docs/ab.md) |
|
|
9446
9645
|
| `agent/storage.ts` | `defineStorage` backend for the durable `host.kv` / `host.files` APIs | [Storage](/docs/storage.md) |
|
|
9447
9646
|
| `agent/artifacts.ts` | `defineArtifacts` kinds, the `tag_artifact` opt-in, and retention | [Artifacts](/docs/reference/artifacts.md) |
|
|
9448
9647
|
| `agent/result.ts` | `defineResult` host `commit` on the final assistant text | [Turn result](/docs/reference/result.md) |
|
|
@@ -9479,8 +9678,6 @@ Continue with these pages:
|
|
|
9479
9678
|
|
|
9480
9679
|
- [Agent config](/docs/reference/agent-config.md): the runtime config at the root
|
|
9481
9680
|
- [Tools](/docs/reference/tools.md): add typed actions under `agent/tools/`
|
|
9482
|
-
- [Live A/B metrics](/docs/ab.md): compare variants from `agent/ab.ts` or
|
|
9483
|
-
`agent/ab/`
|
|
9484
9681
|
- [Concepts](/docs/concepts.md): why the filesystem is the interface
|
|
9485
9682
|
|
|
9486
9683
|
---
|
|
@@ -9924,7 +10121,7 @@ within one session. The `at` field is an ISO-8601 timestamp.
|
|
|
9924
10121
|
|
|
9925
10122
|
| Phase | Events | What they tell you |
|
|
9926
10123
|
| --- | --- | --- |
|
|
9927
|
-
| Session | `session.started`,
|
|
10124
|
+
| Session | `session.started`, `session.waiting`, `session.completed`, `session.failed` | Session creation, readiness, and task completion |
|
|
9928
10125
|
| Agent | `agent.bound` | Cloud conversation URL |
|
|
9929
10126
|
| Input | `message.received` | A user message was accepted |
|
|
9930
10127
|
| Turn | `turn.queued`, `turn.started`, `turn.completed`, `turn.failed` | Queue position under a [`maxRunningTurns` cap](/docs/reference/agent-config.md#concurrency), then turn status, final result, and token usage |
|
|
@@ -10008,7 +10205,6 @@ view.
|
|
|
10008
10205
|
|
|
10009
10206
|
- [HTTP API](/docs/reference/http-api.md)
|
|
10010
10207
|
- [Hooks](/docs/reference/hooks.md)
|
|
10011
|
-
- [Live A/B metrics](/docs/ab.md)
|
|
10012
10208
|
- [How the Agent SDK works](/docs/concepts.md)
|
|
10013
10209
|
|
|
10014
10210
|
---
|
|
@@ -10661,61 +10857,6 @@ inputs again, and adds an eval for each improvement you keep.
|
|
|
10661
10857
|
|
|
10662
10858
|
---
|
|
10663
10859
|
|
|
10664
|
-
Source: /docs/skills/ab.md
|
|
10665
|
-
|
|
10666
|
-
# Agent SDK A/B metrics (`defineAB`)
|
|
10667
|
-
|
|
10668
|
-
Live metrics plug-in. No `agent-sdk ab` CLI. No assertion API.
|
|
10669
|
-
Reference: `docs/ab.md`.
|
|
10670
|
-
|
|
10671
|
-
| | `defineEval` | `defineAB` |
|
|
10672
|
-
| --- | --- | --- |
|
|
10673
|
-
| Job | Gates on frozen fixtures | Metrics on live runs |
|
|
10674
|
-
| Location | `evals/**/*.eval.ts` | `agent/ab.ts` or `agent/ab/<name>.ts` |
|
|
10675
|
-
| How it runs | `agent-sdk eval` | Under `serve` / `run` |
|
|
10676
|
-
|
|
10677
|
-
```ts
|
|
10678
|
-
import { defineAB, splitBySessionHash } from "@cursor/july/ab";
|
|
10679
|
-
|
|
10680
|
-
export default defineAB({
|
|
10681
|
-
name: "concise-instructions",
|
|
10682
|
-
variants: {
|
|
10683
|
-
control: { label: "Baseline" },
|
|
10684
|
-
treatment: {
|
|
10685
|
-
label: "Shorter",
|
|
10686
|
-
instructions: "Keep replies to one short paragraph.",
|
|
10687
|
-
},
|
|
10688
|
-
},
|
|
10689
|
-
split: splitBySessionHash({ holdout: 0.1 }),
|
|
10690
|
-
derive: {
|
|
10691
|
-
weatherCalls: (event) =>
|
|
10692
|
-
event.type === "action.result" && event.data.toolName === "get_weather"
|
|
10693
|
-
? 1
|
|
10694
|
-
: null,
|
|
10695
|
-
},
|
|
10696
|
-
onSample(sample) {
|
|
10697
|
-
console.log(sample.variant, sample.metrics.toolCalls, sample.metrics.wallTimeMs);
|
|
10698
|
-
},
|
|
10699
|
-
});
|
|
10700
|
-
```
|
|
10701
|
-
|
|
10702
|
-
```ts
|
|
10703
|
-
async execute(input, ctx) {
|
|
10704
|
-
if (ctx.session.abs?.["concise-instructions"] === "treatment") {
|
|
10705
|
-
// treatment-specific behavior
|
|
10706
|
-
}
|
|
10707
|
-
}
|
|
10708
|
-
```
|
|
10709
|
-
|
|
10710
|
-
Enrollment is at session creation. Eval sessions skip it. Do not
|
|
10711
|
-
use `splitIf` to filter evals. Split helpers and `onSample`
|
|
10712
|
-
fields: `docs/ab.md`.
|
|
10713
|
-
|
|
10714
|
-
Pick a name, arm labels, a split, and a real `onSample` sink. Do
|
|
10715
|
-
not invent credentials.
|
|
10716
|
-
|
|
10717
|
-
---
|
|
10718
|
-
|
|
10719
10860
|
Source: /docs/skills/create-agent.md
|
|
10720
10861
|
|
|
10721
10862
|
# Create an Agent SDK agent
|
|
@@ -10922,15 +11063,182 @@ trace.
|
|
|
10922
11063
|
|
|
10923
11064
|
---
|
|
10924
11065
|
|
|
11066
|
+
Source: /docs/skills/deploy.md
|
|
11067
|
+
|
|
11068
|
+
# Deploy an Agent SDK agent
|
|
11069
|
+
|
|
11070
|
+
Use this skill only after a person asks for a deployment. It deploys a
|
|
11071
|
+
pushed Git ref to Cursor-managed hosting. Local files never upload.
|
|
11072
|
+
|
|
11073
|
+
Use the exact published `@cursor/july` release embedded in the installed
|
|
11074
|
+
skill. Do not use a floating npm tag or a workspace build of the CLI.
|
|
11075
|
+
|
|
11076
|
+
## Gather the target
|
|
11077
|
+
|
|
11078
|
+
Resolve these values from the request and checkout:
|
|
11079
|
+
|
|
11080
|
+
- Agent directory. Default to the current directory only when it contains
|
|
11081
|
+
one Agent SDK project.
|
|
11082
|
+
- Deployment slug. Default to the normalized directory name.
|
|
11083
|
+
- Git ref. Default to the exact pushed `HEAD` commit.
|
|
11084
|
+
- Team. Use the service account's team unless the request names another.
|
|
11085
|
+
- Cursor-event repositories and CLI-only egress domains.
|
|
11086
|
+
|
|
11087
|
+
Ask one focused question when the target or requested ref is ambiguous.
|
|
11088
|
+
Do not ask for values the checkout or existing deployment supplies.
|
|
11089
|
+
|
|
11090
|
+
## Use the attached service account
|
|
11091
|
+
|
|
11092
|
+
Continue only when `CURSOR_SERVICE_ACCOUNT_KEY` is present. Never print
|
|
11093
|
+
the value. Do not run `agent-sdk login`.
|
|
11094
|
+
|
|
11095
|
+
On Linux and macOS, run every Agent SDK command through this wrapper:
|
|
11096
|
+
|
|
11097
|
+
```bash
|
|
11098
|
+
CURSOR_JULY_VERSION="__CURSOR_JULY_VERSION__"
|
|
11099
|
+
case "$CURSOR_JULY_VERSION" in
|
|
11100
|
+
__CURSOR_JULY_*__)
|
|
11101
|
+
echo "The Agent SDK deploy skill is not bound to a package version." >&2
|
|
11102
|
+
exit 1
|
|
11103
|
+
;;
|
|
11104
|
+
esac
|
|
11105
|
+
|
|
11106
|
+
run_agent_sdk() {
|
|
11107
|
+
env -u NODE_OPTIONS -u CURSOR_API_KEY \
|
|
11108
|
+
CURSOR_API_KEY_FILE=/dev/null \
|
|
11109
|
+
npx --yes "@cursor/july@$CURSOR_JULY_VERSION" "$@"
|
|
11110
|
+
}
|
|
11111
|
+
```
|
|
11112
|
+
|
|
11113
|
+
The installed copy pins the release it came from. The wrapper isolates that
|
|
11114
|
+
published CLI and the attached service account. The presence check prevents a
|
|
11115
|
+
stored personal login from being used when the attachment is missing.
|
|
11116
|
+
|
|
11117
|
+
Verify the principal before any write:
|
|
11118
|
+
|
|
11119
|
+
```bash
|
|
11120
|
+
test -n "${CURSOR_SERVICE_ACCOUNT_KEY:-}" || {
|
|
11121
|
+
echo "No attached Cursor service account." >&2
|
|
11122
|
+
exit 1
|
|
11123
|
+
}
|
|
11124
|
+
run_agent_sdk whoami --json
|
|
11125
|
+
```
|
|
11126
|
+
|
|
11127
|
+
Continue only when `credentialSource` is `service-account`. Record the
|
|
11128
|
+
team and complete service-account ID for the final report. Stop on an
|
|
11129
|
+
authentication, team-access, repository-scope, or hosting-entitlement
|
|
11130
|
+
error.
|
|
11131
|
+
|
|
11132
|
+
## Pin pushed source
|
|
11133
|
+
|
|
11134
|
+
Find the Git root and commit:
|
|
11135
|
+
|
|
11136
|
+
```bash
|
|
11137
|
+
GIT_ROOT="$(git -C "$AGENT_DIR" rev-parse --show-toplevel)"
|
|
11138
|
+
SOURCE_REF="$(git -C "$GIT_ROOT" rev-parse HEAD)"
|
|
11139
|
+
git -C "$GIT_ROOT" status --short
|
|
11140
|
+
git -C "$GIT_ROOT" branch -r --contains "$SOURCE_REF"
|
|
11141
|
+
```
|
|
11142
|
+
|
|
11143
|
+
Use a ref from the request instead of `SOURCE_REF` when the person names
|
|
11144
|
+
one. Confirm the selected commit exists on the remote. If intended
|
|
11145
|
+
changes are uncommitted or unpushed, stop and explain they will not ship.
|
|
11146
|
+
Do not commit or push unless the request includes that work.
|
|
11147
|
+
|
|
11148
|
+
The CLI infers the normalized HTTPS `origin`, agent path, and slug from
|
|
11149
|
+
`--dir`. Pass `--repo`, `--path`, or `--slug` when the request overrides
|
|
11150
|
+
the inferred value. Never print a remote URL that contains credentials.
|
|
11151
|
+
|
|
11152
|
+
## Preserve deployment inputs
|
|
11153
|
+
|
|
11154
|
+
A redeploy replaces its source, Cursor-event repositories, and
|
|
11155
|
+
CLI-supplied egress domains. Preserve the existing
|
|
11156
|
+
`cursorEventRepos` and `egressAllowedDomains` unless the request or
|
|
11157
|
+
project changes them. Reapply each value with:
|
|
11158
|
+
|
|
11159
|
+
```text
|
|
11160
|
+
--cursor-events-repo <owner/repo>
|
|
11161
|
+
--allow-domain <hostname>
|
|
11162
|
+
```
|
|
11163
|
+
|
|
11164
|
+
Read existing settings through a filter that selects only those two
|
|
11165
|
+
fields:
|
|
11166
|
+
|
|
11167
|
+
```bash
|
|
11168
|
+
run_agent_sdk deployment "$SLUG" --json |
|
|
11169
|
+
node -e '
|
|
11170
|
+
let raw = "";
|
|
11171
|
+
process.stdin.setEncoding("utf8");
|
|
11172
|
+
process.stdin.on("data", (chunk) => { raw += chunk; });
|
|
11173
|
+
process.stdin.on("end", () => {
|
|
11174
|
+
const value = JSON.parse(raw);
|
|
11175
|
+
console.log(JSON.stringify({
|
|
11176
|
+
cursorEventRepos: value.cursorEventRepos ?? [],
|
|
11177
|
+
egressAllowedDomains: value.egressAllowedDomains ?? [],
|
|
11178
|
+
}, null, 2));
|
|
11179
|
+
});
|
|
11180
|
+
'
|
|
11181
|
+
```
|
|
11182
|
+
|
|
11183
|
+
Never print or save the full response because it can contain short-lived
|
|
11184
|
+
engine-access headers.
|
|
11185
|
+
|
|
11186
|
+
For a new SCM-channel deployment, ask which repositories should wake the
|
|
11187
|
+
agent when the project does not declare the answer. The source repository
|
|
11188
|
+
is not always the event repository.
|
|
11189
|
+
|
|
11190
|
+
## Validate, deploy, and verify
|
|
11191
|
+
|
|
11192
|
+
Validate with the same published package and principal:
|
|
11193
|
+
|
|
11194
|
+
```bash
|
|
11195
|
+
run_agent_sdk validate --dir "$AGENT_DIR"
|
|
11196
|
+
```
|
|
11197
|
+
|
|
11198
|
+
Stop on validation errors. Review hosting warnings about secret names,
|
|
11199
|
+
egress domains, and channel configuration before continuing.
|
|
11200
|
+
|
|
11201
|
+
Deploy the pinned source. Add the preserved or requested repeatable
|
|
11202
|
+
flags:
|
|
11203
|
+
|
|
11204
|
+
```bash
|
|
11205
|
+
run_agent_sdk deploy \
|
|
11206
|
+
--dir "$AGENT_DIR" \
|
|
11207
|
+
--ref "$SOURCE_REF"
|
|
11208
|
+
```
|
|
11209
|
+
|
|
11210
|
+
The command waits for a terminal deployment state. Do not treat
|
|
11211
|
+
`pending`, `accepted`, or `deploying` as success. If the command is
|
|
11212
|
+
interrupted, resume inspection with:
|
|
11213
|
+
|
|
11214
|
+
```bash
|
|
11215
|
+
run_agent_sdk deployment "$SLUG"
|
|
11216
|
+
```
|
|
11217
|
+
|
|
11218
|
+
A first deployment prints a one-time alias token. Never paste it into
|
|
11219
|
+
chat or expose it in an agent-captured terminal. Before a first deploy,
|
|
11220
|
+
ask for a secure destination or ask the person to run the final command
|
|
11221
|
+
in a private terminal. When writing the token, use mode `0600` and print
|
|
11222
|
+
only the path. Redeploys do not print the token. JSON deployment output
|
|
11223
|
+
can also contain credentials, so filter it before display.
|
|
11224
|
+
|
|
11225
|
+
Finish only when the deployment reports `running`. Report the slug,
|
|
11226
|
+
team, generation, pinned ref, and complete service-account ID. On
|
|
11227
|
+
failure, report the deployment status and sanitized error without
|
|
11228
|
+
retrying a different ref or principal.
|
|
11229
|
+
|
|
11230
|
+
---
|
|
11231
|
+
|
|
10925
11232
|
Source: /docs/skills/evals.md
|
|
10926
11233
|
|
|
10927
11234
|
# Agent SDK evals
|
|
10928
11235
|
|
|
10929
11236
|
Fixed input, model turn, gates on the trajectory. Files live at
|
|
10930
11237
|
project-root `evals/**/*.eval.ts`. `agent/evals/` is ignored.
|
|
11238
|
+
Guide: `docs/evals.md`.
|
|
10931
11239
|
|
|
10932
|
-
|
|
10933
|
-
|
|
11240
|
+
Observing production: hooks. Tool logic without a model:
|
|
11241
|
+
`agent-sdk call <tool> --input '{...}'`.
|
|
10934
11242
|
|
|
10935
11243
|
```bash
|
|
10936
11244
|
agent-sdk eval --dir . --list
|
|
@@ -10944,9 +11252,13 @@ agent-sdk eval --dir . --tag smoke
|
|
|
10944
11252
|
| `evals/weather.eval.ts` + `test` | `weather` |
|
|
10945
11253
|
| `evals/weather/nyc.eval.ts` + `test` | `weather/nyc` |
|
|
10946
11254
|
| `evals/weather.eval.ts` + `{ id: "nyc" }` | `weather/nyc` |
|
|
11255
|
+
| `evals/sql.eval.ts` exporting an array of `defineEval` calls | `sql/0000`, `sql/0001`, ... |
|
|
10947
11256
|
|
|
10948
11257
|
`eval` boots an ephemeral server and a temp state root. `--url`
|
|
10949
|
-
points at a running agent. Model turns need a Cursor credential
|
|
11258
|
+
points at a running agent. Model turns need a Cursor credential
|
|
11259
|
+
(`CURSOR_API_KEY`, then `CURSOR_API_KEY_FILE`, then
|
|
11260
|
+
`CURSOR_SERVICE_ACCOUNT_KEY`, then `agent-sdk login`); `--list` does
|
|
11261
|
+
not. Node 22.13+, never Bun.
|
|
10950
11262
|
|
|
10951
11263
|
## Seeding
|
|
10952
11264
|
|
|
@@ -10996,9 +11308,30 @@ export default defineEvalConfig({
|
|
|
10996
11308
|
});
|
|
10997
11309
|
```
|
|
10998
11310
|
|
|
10999
|
-
Either `test(t)` or `cases`, not both. `t.send` waits for
|
|
11000
|
-
|
|
11001
|
-
|
|
11311
|
+
Either `test(t)` or `cases`, not both. `t.send` waits for the turn to
|
|
11312
|
+
settle (complete, park, or fail) and returns it; `turn.calledTool(...)`
|
|
11313
|
+
reads only that turn, `t.*` reads the whole run. First-send options:
|
|
11314
|
+
`workspaceFiles`, `workspaceDir`, `cloud`.
|
|
11315
|
+
|
|
11316
|
+
Gates: `t.succeeded()` / `t.parked()`,
|
|
11317
|
+
`calledTool(name, { input, output, status, count })` / `notCalledTool`,
|
|
11318
|
+
`toolOrder`, `maxToolCalls`, `taggedArtifact`, `event` / `notEvent`,
|
|
11319
|
+
`t.check(value, includes | equals | matches | similarity | satisfies)`.
|
|
11320
|
+
Records: `t.metric`, `t.log`, `t.score` (soft 0-1).
|
|
11321
|
+
|
|
11322
|
+
Severity is on the handle: default gate; `.soft()` tracked;
|
|
11323
|
+
`.atLeast(0.8)` soft with a bar (verdict `scored`, exit 1 only under
|
|
11324
|
+
`--strict`); `.gate(0.8)` hard. `t.judge.factuality`, `summarizes`,
|
|
11325
|
+
`closedQA`, and `sql` are soft by default and need a judge model
|
|
11326
|
+
(`judge` in `evals.config.ts`, on the eval or case, or per-call
|
|
11327
|
+
`{ model }`).
|
|
11328
|
+
|
|
11329
|
+
## Side effects
|
|
11330
|
+
|
|
11331
|
+
Eval sessions run real tools. Guard actuation with
|
|
11332
|
+
`ctx.session.purpose === "eval"` in the tool, hook, or `defineResult`
|
|
11333
|
+
and return a shaped result so `calledTool` can still assert the
|
|
11334
|
+
decision.
|
|
11002
11335
|
|
|
11003
11336
|
## What to gate
|
|
11004
11337
|
|
|
@@ -11020,6 +11353,18 @@ drifting inputs (pin them).
|
|
|
11020
11353
|
| GitHub | `github replay … --dry-run --out fixtures/github` |
|
|
11021
11354
|
| Host-prep PR review | A team-owned PR; gate findings shape, not counts |
|
|
11022
11355
|
| Workspace | `workspaceFiles` in `t.send` |
|
|
11356
|
+
| Dataset | `loadJson` / `loadJsonl` / `loadYaml` from `@cursor/july/evals/loaders`, paths from the project root |
|
|
11357
|
+
|
|
11358
|
+
## Debug and CI
|
|
11359
|
+
|
|
11360
|
+
A failed local run leaves `evals/<case-id>.json` (assertions, inputs,
|
|
11361
|
+
tool I/O, final text, logs) under `evals/<stamp>/` in the project state
|
|
11362
|
+
directory (the CLI prints the path); read it before editing the eval.
|
|
11363
|
+
`--state-root <path>` keeps the raw session events too.
|
|
11364
|
+
|
|
11365
|
+
CI: `agent-sdk eval --dir . --json --no-stream --junit reports/evals.xml`.
|
|
11366
|
+
Exit `1` on a failed gate, `2` when nothing matched, `--strict` to fail
|
|
11367
|
+
on `scored`.
|
|
11023
11368
|
|
|
11024
11369
|
Every kept hillclimb change lands an eval that would have failed
|
|
11025
11370
|
before it. Never weaken a gate to pass a round.
|
|
@@ -11073,7 +11418,6 @@ Path is identity. Full list: README "Folder structure".
|
|
|
11073
11418
|
| `agent/hooks/*.ts` | Observe-only |
|
|
11074
11419
|
| `agent/artifacts.ts` | Durable tagged outputs (`defineArtifacts`) |
|
|
11075
11420
|
| `agent/result.ts` | Host `commit` on the final assistant text (`defineResult`) |
|
|
11076
|
-
| `agent/ab.ts` or `agent/ab/*.ts` | Live A/B (`defineAB`) |
|
|
11077
11421
|
| `agent/otel.ts` | OpenTelemetry (`defineOtel`) |
|
|
11078
11422
|
| `agent/schedules/*` | Cron. Never auto-fire under `--dev` |
|
|
11079
11423
|
| `agent/sandbox/workspace/` | Session seed files (local only) |
|
|
@@ -11121,11 +11465,11 @@ only this agent's directory.
|
|
|
11121
11465
|
| --- | --- |
|
|
11122
11466
|
| Scaffold | `skills/create-agent/SKILL.md` |
|
|
11123
11467
|
| Evals | `skills/evals/SKILL.md` |
|
|
11124
|
-
| Live A/B | `skills/ab/SKILL.md` |
|
|
11125
11468
|
| OpenTelemetry | `skills/otel/SKILL.md` |
|
|
11126
11469
|
| GitHub | `skills/github/SKILL.md` |
|
|
11127
11470
|
| Slack | `skills/setup-slack/SKILL.md` |
|
|
11128
11471
|
| Host MCP OAuth | `skills/mcp-auth/SKILL.md` |
|
|
11472
|
+
| Managed deployment | `skills/deploy/SKILL.md` |
|
|
11129
11473
|
| Local triage | `skills/debug/SKILL.md` |
|
|
11130
11474
|
| Measured improvement | `skills/hillclimb/SKILL.md` |
|
|
11131
11475
|
|
|
@@ -11305,12 +11649,12 @@ to use each one.
|
|
|
11305
11649
|
| [framework-map](/docs/skills/framework-map.md) | Learn the project layout and runtimes |
|
|
11306
11650
|
| [create-agent](/docs/skills/create-agent.md) | Scaffold and verify a new agent |
|
|
11307
11651
|
| [evals](/docs/skills/evals.md) | Write fixtures and regression checks |
|
|
11308
|
-
| [ab](/docs/skills/ab.md) | Compare variants on live traffic |
|
|
11309
11652
|
| [otel](/docs/skills/otel.md) | Export OpenTelemetry traces |
|
|
11310
11653
|
| [hillclimb](/docs/skills/hillclimb.md) | Improve an agent against fixed inputs |
|
|
11311
11654
|
| [github](/docs/skills/github.md) | Add GitHub webhooks and replay events |
|
|
11312
11655
|
| [setup-slack](/docs/skills/setup-slack.md) | Connect an agent to Slack |
|
|
11313
11656
|
| [mcp-auth](/docs/skills/mcp-auth.md) | Authorize host MCP OAuth |
|
|
11657
|
+
| [deploy](/docs/skills/deploy.md) | Deploy with an attached service account |
|
|
11314
11658
|
| [debug](/docs/skills/debug.md) | Diagnose a local run |
|
|
11315
11659
|
|
|
11316
11660
|
---
|
|
@@ -11605,8 +11949,8 @@ Source: /docs/storage.md
|
|
|
11605
11949
|
# Storage
|
|
11606
11950
|
|
|
11607
11951
|
The Agent SDK owns durable storage for sessions, continuation tokens,
|
|
11608
|
-
reminders, playground eval history
|
|
11609
|
-
|
|
11952
|
+
reminders, and playground eval history. It chooses the keys, when to
|
|
11953
|
+
read and write, and how to restore after restart.
|
|
11610
11954
|
|
|
11611
11955
|
The Agent SDK owns key encoding. Backends must accept the keys they are
|
|
11612
11956
|
given. Do not fail `put` to enforce a shorter cap.
|
|
@@ -11640,9 +11984,9 @@ export default defineStorage({
|
|
|
11640
11984
|
## Which fields to provide
|
|
11641
11985
|
|
|
11642
11986
|
Implement the small KV core (`put`/`get`/`delete`/`list` plus the `cas`
|
|
11643
|
-
group) and you get **full functionality**: eval-run
|
|
11644
|
-
|
|
11645
|
-
|
|
11987
|
+
group) and you get **full functionality**: eval-run history is derived
|
|
11988
|
+
over the core automatically. The dedicated `evals` group is a
|
|
11989
|
+
backend-native optimization, not a required-or-lose-history hook.
|
|
11646
11990
|
|
|
11647
11991
|
| Field | Required | Role |
|
|
11648
11992
|
| --- | --- | --- |
|
|
@@ -11653,36 +11997,28 @@ are backend-native optimizations, not required-or-lose-history hooks.
|
|
|
11653
11997
|
| `delete` | For cleanup | Remove a key |
|
|
11654
11998
|
| `name` | No | Label surfaced on `GET /v1/info` diagnostics |
|
|
11655
11999
|
| `policy` | No | Timing knobs; see [Policy](#policy) |
|
|
11656
|
-
| `evals` | No | Backend-native eval-runs table; derived over the core when omitted. See [Eval
|
|
11657
|
-
| `abs` | No | Backend-native A/B metrics table; derived over the core when omitted. See [Eval and A/B tables](#eval-and-a-b-tables) |
|
|
12000
|
+
| `evals` | No | Backend-native eval-runs table; derived over the core when omitted. See [Eval table](#eval-table) |
|
|
11658
12001
|
|
|
11659
12002
|
A throwing `put` is logged and dropped. It never fails a turn. When
|
|
11660
12003
|
resolving a missing continuation token, a throwing `get` fails the
|
|
11661
12004
|
follow-up so a store outage does not open a new session. Return
|
|
11662
12005
|
`undefined` only for a real miss.
|
|
11663
12006
|
|
|
11664
|
-
## Eval
|
|
12007
|
+
## Eval table
|
|
11665
12008
|
|
|
11666
|
-
|
|
11667
|
-
values.
|
|
11668
|
-
authored, `defineStorage` derives it over the KV core, so a backend
|
|
11669
|
-
implements only the core loses nothing. Author
|
|
11670
|
-
backend has a better native shape (a real database table, an
|
|
11671
|
-
pipeline). The built-in `fileKv` and `cursorHostedStorage`
|
|
12009
|
+
The dedicated `evals` table group carries structured rows instead of
|
|
12010
|
+
opaque KV values. It is an **optional optimization**: when the group is
|
|
12011
|
+
not authored, `defineStorage` derives it over the KV core, so a backend
|
|
12012
|
+
that implements only the core loses nothing. Author the group only when
|
|
12013
|
+
the backend has a better native shape (a real database table, an
|
|
12014
|
+
analytics pipeline). The built-in `fileKv` and `cursorHostedStorage`
|
|
12015
|
+
both do.
|
|
11672
12016
|
|
|
11673
12017
|
`evals` keeps playground eval batches across restarts (`put`, `delete`,
|
|
11674
12018
|
`list` over run snapshots keyed by `runId`). A core missing `delete` or
|
|
11675
12019
|
`list` leaves eval history in memory until restart.
|
|
11676
12020
|
See [Evals](/docs/evals.md#configure-eval-runs).
|
|
11677
12021
|
|
|
11678
|
-
`abs` exports live A/B metrics: `putSample` appends one cumulative
|
|
11679
|
-
metric sample per enrolled experiment on each completed or failed turn;
|
|
11680
|
-
optional `putSnapshot` / `getSnapshot` store and serve back the latest
|
|
11681
|
-
aggregate so a replacement host can still serve the A/Bs surface.
|
|
11682
|
-
`putSample` and `putSnapshot` need only core `put`; `getSnapshot` needs
|
|
11683
|
-
core `get`. Session event logs remain the assignment source of truth
|
|
11684
|
-
either way. See [Live A/B metrics](/docs/ab.md).
|
|
11685
|
-
|
|
11686
12022
|
## Policy
|
|
11687
12023
|
|
|
11688
12024
|
Two knobs change behavior:
|
|
@@ -11717,8 +12053,7 @@ With `get` and `list`, serve can rebuild local state from your store:
|
|
|
11717
12053
|
- At startup, the Agent SDK loads recent sessions up to the restore caps.
|
|
11718
12054
|
- On demand, a missing continuation token resolves through the store
|
|
11719
12055
|
and resumes that session.
|
|
11720
|
-
- Playground eval history
|
|
11721
|
-
sink.
|
|
12056
|
+
- Playground eval history can load from the same sink.
|
|
11722
12057
|
|
|
11723
12058
|
A turn in flight at crash time is not replayed. The next follow-up
|
|
11724
12059
|
resumes from the last flushed state.
|