@cursor/july 0.1.93 → 0.1.94
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +8 -20
- package/README.md +4 -26
- package/dist/channels/slack/attachments.js +2 -2
- package/dist/channels/slack/dispatch.d.ts +0 -7
- package/dist/channels/slack/dispatch.d.ts.map +1 -1
- package/dist/channels/slack/dispatch.js +4 -7
- package/dist/channels/slack/eval-directive.d.ts +5 -12
- package/dist/channels/slack/eval-directive.d.ts.map +1 -1
- package/dist/channels/slack/eval-directive.js +8 -19
- package/dist/channels/slack/index.d.ts +0 -6
- package/dist/channels/slack/index.d.ts.map +1 -1
- package/dist/channels/slack/index.js +0 -6
- package/dist/channels/slack/setup.d.ts +4 -4
- package/dist/channels/slack/setup.d.ts.map +1 -1
- package/dist/channels/slack/setup.js +8 -15
- package/dist/channels/slack/slack-channel.d.ts +6 -13
- package/dist/channels/slack/slack-channel.d.ts.map +1 -1
- package/dist/channels/slack/slack-channel.js +15 -101
- package/dist/channels/slack/types.d.ts +12 -79
- package/dist/channels/slack/types.d.ts.map +1 -1
- package/dist/channels/slack/types.js +1 -15
- package/dist/client.d.ts +14 -0
- package/dist/client.d.ts.map +1 -0
- package/dist/client.js +12 -0
- package/dist/connections.d.ts +18 -9
- package/dist/connections.d.ts.map +1 -1
- package/dist/connections.js +17 -8
- package/dist/docs/404.html +2 -2
- package/dist/docs/ab.html +4 -4
- package/dist/docs/assets/{app.CjWU-x0z.js → app.CFDEas4I.js} +1 -1
- package/dist/docs/assets/chunks/@localSearchIndexroot.DU3U2Ij2.js +1 -0
- package/dist/docs/assets/chunks/{VPLocalSearchBox.Cxy8ySFQ.js → VPLocalSearchBox.B1IIYpYS.js} +1 -1
- package/dist/docs/assets/chunks/{theme.Dvq1Bktu.js → theme.Ct4NSiLm.js} +2 -2
- package/dist/docs/assets/concepts.md.lwAgBIMI.js +1 -0
- package/dist/docs/assets/{deployment.md.DoLFAzfm.js → deployment.md.D9msOFOW.js} +3 -8
- package/dist/docs/assets/{deployment.md.DoLFAzfm.lean.js → deployment.md.D9msOFOW.lean.js} +1 -1
- package/dist/docs/assets/{guides_agent-to-agent.md.B3JIaAqz.js → guides_agent-to-agent.md.BDb0t1QV.js} +1 -1
- package/dist/docs/assets/guides_cloud-runtime.md.CkYbjnAX.js +9 -0
- package/dist/docs/assets/guides_cloud-runtime.md.CkYbjnAX.lean.js +1 -0
- package/dist/docs/assets/{guides_convert-automation.md.Bboisykk.js → guides_convert-automation.md.B4sjlodG.js} +1 -1
- package/dist/docs/assets/{guides_github.md.DqJhuaN1.js → guides_github.md.Cnh2mL4a.js} +1 -1
- package/dist/docs/assets/{guides_mcp-oauth.md.CJvrXtkN.js → guides_mcp-oauth.md.DPYmBCbV.js} +7 -9
- package/dist/docs/assets/{guides_mcp-oauth.md.CJvrXtkN.lean.js → guides_mcp-oauth.md.DPYmBCbV.lean.js} +1 -1
- package/dist/docs/assets/{guides_slack.md.mqeNKs84.js → guides_slack.md.C32HsdKk.js} +5 -11
- package/dist/docs/assets/guides_slack.md.C32HsdKk.lean.js +1 -0
- package/dist/docs/assets/index.md.DRakGHFe.js +5 -0
- package/dist/docs/assets/{index.md.B-lVR4wT.lean.js → index.md.DRakGHFe.lean.js} +1 -1
- package/dist/docs/assets/{quickstart.md.BrmfrrIr.js → quickstart.md.Nj_LjW_a.js} +1 -1
- package/dist/docs/assets/{reference_cli.md.D9KESDsD.js → reference_cli.md.Cw6_ICYG.js} +1 -1
- package/dist/docs/assets/{reference_connections.md.DB6SsN6U.js → reference_connections.md.BH8Oc0D0.js} +5 -5
- package/dist/docs/assets/{reference_connections.md.DB6SsN6U.lean.js → reference_connections.md.BH8Oc0D0.lean.js} +1 -1
- package/dist/docs/assets/{reference_hooks.md.BxN87gCw.js → reference_hooks.md.a8BJxMR5.js} +1 -1
- package/dist/docs/assets/reference_http-api.md.D89k1mdm.js +11 -0
- package/dist/docs/assets/reference_http-api.md.D89k1mdm.lean.js +1 -0
- package/dist/docs/assets/reference_project-layout.md.Bv4KOtlB.js +19 -0
- package/dist/docs/assets/{reference_skills.md.BFW9retM.js → reference_skills.md.8son6Hjm.js} +3 -3
- package/dist/docs/assets/{reference_subagents.md.Xoav0AII.js → reference_subagents.md.CfsIloPm.js} +1 -1
- package/dist/docs/assets/{reference_tools.md.DuKvkYWG.js → reference_tools.md.BHeXn2id.js} +3 -3
- package/dist/docs/assets/{reference_tools.md.DuKvkYWG.lean.js → reference_tools.md.BHeXn2id.lean.js} +1 -1
- package/dist/docs/assets/{templates_pr-autofixer.md.R4K_qytS.js → templates_pr-autofixer.md.DU7dQpor.js} +2 -2
- package/dist/docs/assets/{templates_pr-autofixer.md.R4K_qytS.lean.js → templates_pr-autofixer.md.DU7dQpor.lean.js} +1 -1
- package/dist/docs/assets/{templates_security-reviewer.md.ByFyRta2.js → templates_security-reviewer.md.CTa7u_l1.js} +2 -2
- package/dist/docs/assets/{templates_security-reviewer.md.ByFyRta2.lean.js → templates_security-reviewer.md.CTa7u_l1.lean.js} +1 -1
- package/dist/docs/assets/troubleshooting.md.Ctv3T8C2.js +1 -0
- package/dist/docs/building-with-agents.html +4 -4
- package/dist/docs/concepts.html +5 -5
- package/dist/docs/concepts.md +1 -0
- package/dist/docs/deployment.html +7 -12
- package/dist/docs/deployment.md +1 -20
- package/dist/docs/design/agsh.md +406 -0
- package/dist/docs/evals.html +4 -4
- package/dist/docs/guides/agent-to-agent.html +6 -6
- package/dist/docs/guides/agent-to-agent.md +2 -2
- package/dist/docs/guides/cloud-runtime.html +6 -6
- package/dist/docs/guides/cloud-runtime.md +1 -0
- package/dist/docs/guides/convert-automation.html +6 -6
- package/dist/docs/guides/convert-automation.md +1 -1
- package/dist/docs/guides/github.html +6 -6
- package/dist/docs/guides/github.md +4 -4
- package/dist/docs/guides/human-in-the-loop.html +4 -4
- package/dist/docs/guides/mcp-oauth.html +11 -13
- package/dist/docs/guides/mcp-oauth.md +10 -18
- package/dist/docs/guides/opentelemetry.html +5 -5
- package/dist/docs/guides/slack.html +9 -15
- package/dist/docs/guides/slack.md +9 -46
- package/dist/docs/guides/webhooks.html +4 -4
- package/dist/docs/hashmap.json +1 -1
- package/dist/docs/hillclimbing.html +4 -4
- package/dist/docs/index.html +6 -6
- package/dist/docs/index.md +0 -28
- package/dist/docs/llms-full.txt +712 -2830
- package/dist/docs/llms.txt +2 -16
- package/dist/docs/quickstart.html +6 -6
- package/dist/docs/quickstart.md +2 -3
- package/dist/docs/reference/agent-config.html +4 -4
- package/dist/docs/reference/artifacts.html +4 -4
- package/dist/docs/reference/channels.html +4 -4
- package/dist/docs/reference/cli.html +6 -6
- package/dist/docs/reference/cli.md +2 -1
- package/dist/docs/reference/connections.html +9 -9
- package/dist/docs/reference/connections.md +15 -11
- package/dist/docs/reference/hooks.html +6 -6
- package/dist/docs/reference/hooks.md +2 -3
- package/dist/docs/reference/http-api.html +6 -6
- package/dist/docs/reference/http-api.md +8 -0
- package/dist/docs/reference/instructions.html +4 -4
- package/dist/docs/reference/playground.html +4 -4
- package/dist/docs/reference/project-layout.html +8 -6
- package/dist/docs/reference/project-layout.md +5 -1
- package/dist/docs/reference/prompt.html +4 -4
- package/dist/docs/reference/schedules.html +4 -4
- package/dist/docs/reference/sessions.html +4 -4
- package/dist/docs/reference/skills.html +7 -7
- package/dist/docs/reference/subagents.html +6 -6
- package/dist/docs/reference/subagents.md +2 -2
- package/dist/docs/reference/tools.html +7 -7
- package/dist/docs/reference/tools.md +19 -3
- package/dist/docs/scaffolding-agents.html +4 -4
- package/dist/docs/storage.html +4 -4
- package/dist/docs/templates/agentic-owners.html +4 -4
- package/dist/docs/templates/demo.html +4 -4
- package/dist/docs/templates/pr-autofixer.html +6 -6
- package/dist/docs/templates/pr-autofixer.md +4 -3
- package/dist/docs/templates/security-reviewer.html +5 -5
- package/dist/docs/templates/security-reviewer.md +2 -3
- package/dist/docs/templates/triage.html +4 -4
- package/dist/docs/troubleshooting.html +5 -5
- package/dist/docs/troubleshooting.md +2 -2
- package/dist/index.d.ts +1 -1
- package/dist/index.d.ts.map +1 -1
- package/dist/index.js +1 -1
- package/dist/internal/advertise-tools.d.ts +11 -0
- package/dist/internal/advertise-tools.d.ts.map +1 -1
- package/dist/internal/advertise-tools.js +47 -9
- package/dist/internal/cli-mcp-oauth.d.ts.map +1 -1
- package/dist/internal/cli-mcp-oauth.js +7 -4
- package/dist/internal/convert-automation/convert-workflow.d.ts.map +1 -1
- package/dist/internal/convert-automation/convert-workflow.js +26 -15
- package/dist/internal/convert-automation/slug.d.ts +0 -2
- package/dist/internal/convert-automation/slug.d.ts.map +1 -1
- package/dist/internal/convert-automation/slug.js +0 -8
- package/dist/internal/cursor/account-mcp.d.ts.map +1 -1
- package/dist/internal/cursor/account-mcp.js +5 -1
- package/dist/internal/discovery.d.ts.map +1 -1
- package/dist/internal/discovery.js +88 -13
- package/dist/internal/hosted-delivery.d.ts.map +1 -1
- package/dist/internal/hosted-delivery.js +22 -9
- package/dist/internal/mcp-endpoint.js +3 -3
- package/dist/internal/mcp-host.d.ts +8 -7
- package/dist/internal/mcp-host.d.ts.map +1 -1
- package/dist/internal/mcp-host.js +8 -7
- package/dist/internal/peer-connections.d.ts.map +1 -1
- package/dist/internal/peer-connections.js +5 -1
- package/dist/internal/playground/static.d.ts +0 -3
- package/dist/internal/playground/static.d.ts.map +1 -1
- package/dist/internal/resolved-connections.d.ts.map +1 -1
- package/dist/internal/resolved-connections.js +5 -7
- package/dist/internal/server.d.ts.map +1 -1
- package/dist/internal/server.js +113 -172
- package/dist/internal/session-engine.d.ts +45 -10
- package/dist/internal/session-engine.d.ts.map +1 -1
- package/dist/internal/session-engine.js +208 -65
- package/dist/internal/tool-catalog.d.ts +31 -0
- package/dist/internal/tool-catalog.d.ts.map +1 -0
- package/dist/internal/tool-catalog.js +67 -0
- package/dist/playground/assets/{index-D9MFzhNE.js → index-B3JCyigB.js} +1 -1
- package/dist/playground/index.html +1 -1
- package/dist/types.d.ts +72 -23
- package/dist/types.d.ts.map +1 -1
- package/dist/types.js +19 -0
- package/docs/README.md +0 -28
- package/docs/concepts.md +1 -0
- package/docs/deployment.md +1 -20
- package/docs/design/agsh.md +406 -0
- package/docs/guides/agent-to-agent.md +2 -2
- package/docs/guides/cloud-runtime.md +1 -0
- package/docs/guides/convert-automation.md +1 -1
- package/docs/guides/github.md +4 -4
- package/docs/guides/mcp-oauth.md +10 -18
- package/docs/guides/slack.md +10 -47
- package/docs/quickstart.md +2 -3
- package/docs/reference/cli.md +2 -1
- package/docs/reference/connections.md +15 -11
- package/docs/reference/hooks.md +2 -3
- package/docs/reference/http-api.md +8 -0
- package/docs/reference/project-layout.md +5 -1
- package/docs/reference/subagents.md +2 -2
- package/docs/reference/tools.md +19 -3
- package/docs/templates/pr-autofixer.md +4 -3
- package/docs/templates/security-reviewer.md +2 -3
- package/docs/troubleshooting.md +2 -2
- package/package.json +9 -2
- package/skills/create-agent/SKILL.md +6 -13
- package/skills/debug/SKILL.md +2 -4
- package/skills/evals/SKILL.md +1 -1
- package/skills/framework-map/SKILL.md +3 -2
- package/skills/mcp-auth/SKILL.md +10 -13
- package/skills/setup-slack/SKILL.md +21 -137
- package/src/channels/slack/attachments.ts +2 -2
- package/src/channels/slack/dispatch.ts +2 -16
- package/src/channels/slack/eval-directive.ts +8 -27
- package/src/channels/slack/index.ts +0 -6
- package/src/channels/slack/setup.ts +8 -15
- package/src/channels/slack/slack-channel.ts +14 -125
- package/src/channels/slack/types.ts +12 -96
- package/src/client.ts +23 -0
- package/src/connections.ts +20 -7
- package/src/index.ts +2 -0
- package/src/internal/advertise-tools.ts +45 -7
- package/src/internal/cli-mcp-oauth.ts +6 -4
- package/src/internal/convert-automation/convert-workflow.ts +29 -17
- package/src/internal/convert-automation/slug.ts +0 -9
- package/src/internal/cursor/account-mcp.ts +4 -1
- package/src/internal/discovery.ts +104 -13
- package/src/internal/fixtures/units-server.ts +52 -0
- package/src/internal/hosted-delivery.ts +60 -28
- package/src/internal/mcp-endpoint.ts +3 -3
- package/src/internal/mcp-host.ts +8 -7
- package/src/internal/peer-connections.ts +4 -1
- package/src/internal/playground/static.ts +1 -3
- package/src/internal/resolved-connections.ts +8 -10
- package/src/internal/server.ts +151 -251
- package/src/internal/session-engine.ts +254 -69
- package/src/internal/tool-catalog.ts +106 -0
- package/src/types.ts +90 -23
- package/templates/pr-autofixer/agent/channels/slack.ts +8 -2
- package/templates/triage/README.md +2 -1
- package/templates/triage/overlays/jira/agent/mcp-connections/tracker.ts +0 -1
- package/templates/triage/overlays/linear/agent/mcp-connections/tracker.ts +0 -1
- package/dist/channels/slack/cursor-account.d.ts +0 -87
- package/dist/channels/slack/cursor-account.d.ts.map +0 -1
- package/dist/channels/slack/cursor-account.js +0 -100
- package/dist/docs/assets/chunks/@localSearchIndexroot.ChpIC3Zy.js +0 -1
- package/dist/docs/assets/concepts.md.F6AiPorA.js +0 -1
- package/dist/docs/assets/example-agents_approval-buddy.md.DmezILPg.js +0 -10
- package/dist/docs/assets/example-agents_approval-buddy.md.DmezILPg.lean.js +0 -1
- package/dist/docs/assets/example-agents_benny.md.B0kwY7D_.js +0 -5
- package/dist/docs/assets/example-agents_benny.md.B0kwY7D_.lean.js +0 -1
- package/dist/docs/assets/example-agents_bugbot.md.BRGMi9O2.js +0 -11
- package/dist/docs/assets/example-agents_bugbot.md.BRGMi9O2.lean.js +0 -1
- package/dist/docs/assets/example-agents_codebase-wiki.md.BBNw9Ekr.js +0 -8
- package/dist/docs/assets/example-agents_codebase-wiki.md.BBNw9Ekr.lean.js +0 -1
- package/dist/docs/assets/example-agents_codeowners-review.md.Bfta-lBU.js +0 -8
- package/dist/docs/assets/example-agents_codeowners-review.md.Bfta-lBU.lean.js +0 -1
- package/dist/docs/assets/example-agents_concierge.md.BzB2b20R.js +0 -22
- package/dist/docs/assets/example-agents_concierge.md.BzB2b20R.lean.js +0 -1
- package/dist/docs/assets/example-agents_index.md.ChBp0AX6.js +0 -2
- package/dist/docs/assets/example-agents_index.md.ChBp0AX6.lean.js +0 -1
- package/dist/docs/assets/example-agents_knowledge-base.md.CrA85ig-.js +0 -11
- package/dist/docs/assets/example-agents_knowledge-base.md.CrA85ig-.lean.js +0 -1
- package/dist/docs/assets/example-agents_oncall.md.DK4XkYTd.js +0 -10
- package/dist/docs/assets/example-agents_oncall.md.DK4XkYTd.lean.js +0 -1
- package/dist/docs/assets/example-agents_security-reviewer.md.74pPpWYj.js +0 -19
- package/dist/docs/assets/example-agents_security-reviewer.md.74pPpWYj.lean.js +0 -1
- package/dist/docs/assets/example-agents_slack-agent.md.D7Kdj5BV.js +0 -5
- package/dist/docs/assets/example-agents_slack-agent.md.D7Kdj5BV.lean.js +0 -1
- package/dist/docs/assets/example-agents_weather-agent.md.CaGpmw3Y.js +0 -25
- package/dist/docs/assets/example-agents_weather-agent.md.CaGpmw3Y.lean.js +0 -1
- package/dist/docs/assets/guides_cloud-runtime.md.BnvjPiia.js +0 -9
- package/dist/docs/assets/guides_cloud-runtime.md.BnvjPiia.lean.js +0 -1
- package/dist/docs/assets/guides_slack.md.mqeNKs84.lean.js +0 -1
- package/dist/docs/assets/index.md.B-lVR4wT.js +0 -5
- package/dist/docs/assets/reference_http-api.md.C68BERYr.js +0 -11
- package/dist/docs/assets/reference_http-api.md.C68BERYr.lean.js +0 -1
- package/dist/docs/assets/reference_project-layout.md.WN9nwJht.js +0 -17
- package/dist/docs/assets/troubleshooting.md.vCWwvqcJ.js +0 -1
- package/dist/docs/example-agents/approval-buddy.html +0 -36
- package/dist/docs/example-agents/approval-buddy.md +0 -266
- package/dist/docs/example-agents/benny.html +0 -31
- package/dist/docs/example-agents/benny.md +0 -173
- package/dist/docs/example-agents/bugbot.html +0 -37
- package/dist/docs/example-agents/bugbot.md +0 -229
- package/dist/docs/example-agents/codebase-wiki.html +0 -34
- package/dist/docs/example-agents/codebase-wiki.md +0 -167
- package/dist/docs/example-agents/codeowners-review.html +0 -34
- package/dist/docs/example-agents/codeowners-review.md +0 -192
- package/dist/docs/example-agents/concierge.html +0 -48
- package/dist/docs/example-agents/concierge.md +0 -200
- package/dist/docs/example-agents/index.html +0 -28
- package/dist/docs/example-agents/index.md +0 -99
- package/dist/docs/example-agents/knowledge-base.html +0 -37
- package/dist/docs/example-agents/knowledge-base.md +0 -168
- package/dist/docs/example-agents/oncall.html +0 -36
- package/dist/docs/example-agents/oncall.md +0 -212
- package/dist/docs/example-agents/security-reviewer.html +0 -45
- package/dist/docs/example-agents/security-reviewer.md +0 -265
- package/dist/docs/example-agents/slack-agent.html +0 -31
- package/dist/docs/example-agents/slack-agent.md +0 -142
- package/dist/docs/example-agents/weather-agent.html +0 -51
- package/dist/docs/example-agents/weather-agent.md +0 -297
- package/dist/internal/cursor-slack-relay.d.ts +0 -96
- package/dist/internal/cursor-slack-relay.d.ts.map +0 -1
- package/dist/internal/cursor-slack-relay.js +0 -176
- package/docs/example-agents/approval-buddy.md +0 -271
- package/docs/example-agents/benny.md +0 -178
- package/docs/example-agents/bugbot.md +0 -234
- package/docs/example-agents/codebase-wiki.md +0 -172
- package/docs/example-agents/codeowners-review.md +0 -197
- package/docs/example-agents/concierge.md +0 -205
- package/docs/example-agents/index.md +0 -104
- package/docs/example-agents/knowledge-base.md +0 -173
- package/docs/example-agents/oncall.md +0 -217
- package/docs/example-agents/security-reviewer.md +0 -270
- package/docs/example-agents/slack-agent.md +0 -147
- package/docs/example-agents/weather-agent.md +0 -302
- package/src/channels/slack/cursor-account.ts +0 -202
- package/src/internal/cursor-slack-relay.ts +0 -249
- /package/dist/docs/assets/{concepts.md.F6AiPorA.lean.js → concepts.md.lwAgBIMI.lean.js} +0 -0
- /package/dist/docs/assets/{guides_agent-to-agent.md.B3JIaAqz.lean.js → guides_agent-to-agent.md.BDb0t1QV.lean.js} +0 -0
- /package/dist/docs/assets/{guides_convert-automation.md.Bboisykk.lean.js → guides_convert-automation.md.B4sjlodG.lean.js} +0 -0
- /package/dist/docs/assets/{guides_github.md.DqJhuaN1.lean.js → guides_github.md.Cnh2mL4a.lean.js} +0 -0
- /package/dist/docs/assets/{quickstart.md.BrmfrrIr.lean.js → quickstart.md.Nj_LjW_a.lean.js} +0 -0
- /package/dist/docs/assets/{reference_cli.md.D9KESDsD.lean.js → reference_cli.md.Cw6_ICYG.lean.js} +0 -0
- /package/dist/docs/assets/{reference_hooks.md.BxN87gCw.lean.js → reference_hooks.md.a8BJxMR5.lean.js} +0 -0
- /package/dist/docs/assets/{reference_project-layout.md.WN9nwJht.lean.js → reference_project-layout.md.Bv4KOtlB.lean.js} +0 -0
- /package/dist/docs/assets/{reference_skills.md.BFW9retM.lean.js → reference_skills.md.8son6Hjm.lean.js} +0 -0
- /package/dist/docs/assets/{reference_subagents.md.Xoav0AII.lean.js → reference_subagents.md.CfsIloPm.lean.js} +0 -0
- /package/dist/docs/assets/{troubleshooting.md.vCWwvqcJ.lean.js → troubleshooting.md.Ctv3T8C2.lean.js} +0 -0
package/dist/docs/llms-full.txt
CHANGED
|
@@ -500,6 +500,7 @@ name. For example, `agent/tools/get_weather.ts` creates a tool named
|
|
|
500
500
|
| `agent/tools/<name>.ts` | Typed actions the model can call |
|
|
501
501
|
| `agent/skills/*` | Procedures loaded when needed |
|
|
502
502
|
| `agent/mcp-connections/<name>.ts` | Tools from external MCP servers |
|
|
503
|
+
| `agent/host-connections/<name>.ts` | Privileged MCP servers for host tools only |
|
|
503
504
|
| `agent/channels/*.ts` | HTTP, Slack, and GitHub entry points |
|
|
504
505
|
| `agent/ab.ts` or `agent/ab/*.ts` | Sticky variants and live performance metrics |
|
|
505
506
|
| `evals/**/*.eval.ts` | Repeatable checks at the project root |
|
|
@@ -863,26 +864,7 @@ See the [GitHub guide](/docs/guides/github.md).
|
|
|
863
864
|
|
|
864
865
|
### Connect Slack
|
|
865
866
|
|
|
866
|
-
Hosted Slack
|
|
867
|
-
Mode app.
|
|
868
|
-
|
|
869
|
-
Use the Cursor Slack app when mentions and direct messages are enough:
|
|
870
|
-
|
|
871
|
-
```ts
|
|
872
|
-
import { slackChannel } from "@cursor/july/channels/slack";
|
|
873
|
-
|
|
874
|
-
export default slackChannel({
|
|
875
|
-
cursorAccount: true,
|
|
876
|
-
agentName: "PrApprover",
|
|
877
|
-
});
|
|
878
|
-
```
|
|
879
|
-
|
|
880
|
-
The team must have the Cursor Slack app installed and Slack event relay
|
|
881
|
-
access enabled. This mode needs no Slack token secrets. It doesn't
|
|
882
|
-
support channel-post watches, tool approvals, or interactivity.
|
|
883
|
-
|
|
884
|
-
Use a dedicated Socket Mode app for those features or a separate bot
|
|
885
|
-
identity:
|
|
867
|
+
Hosted Slack uses a dedicated Socket Mode app:
|
|
886
868
|
|
|
887
869
|
```ts
|
|
888
870
|
import { slackChannel } from "@cursor/july/channels/slack";
|
|
@@ -1105,6 +1087,417 @@ Continue with these pages:
|
|
|
1105
1087
|
|
|
1106
1088
|
---
|
|
1107
1089
|
|
|
1090
|
+
Source: /docs/design/agsh.md
|
|
1091
|
+
|
|
1092
|
+
# agsh: a shell for deployed agents
|
|
1093
|
+
|
|
1094
|
+
## What this is
|
|
1095
|
+
|
|
1096
|
+
`agsh` (agent shell) is a standalone CLI that connects to one agent-sdk
|
|
1097
|
+
deployment and turns the agent's live tool surface into commands. Every tool
|
|
1098
|
+
the deployment can execute (authored server tools and tools provided by the
|
|
1099
|
+
agent's MCP connections) becomes a subcommand with a synopsis derived from its
|
|
1100
|
+
input schema, a man-page style `--help`, and a place in an interactive shell.
|
|
1101
|
+
|
|
1102
|
+
It is a separate binary and a separate package from `agent-sdk`. The
|
|
1103
|
+
`agent-sdk` CLI stays what it is today: the developer workflow tool for
|
|
1104
|
+
authoring, validating, deploying, and debugging agent projects. `agsh` is the
|
|
1105
|
+
operator's tool for working *inside* one deployed agent. The split also keeps
|
|
1106
|
+
heavy presentation dependencies (markdown rendering, syntax highlighting, the
|
|
1107
|
+
shell interpreter) out of `@cursor/july`, which ships to every agent project.
|
|
1108
|
+
|
|
1109
|
+
## The experience
|
|
1110
|
+
|
|
1111
|
+
```
|
|
1112
|
+
$ agsh help # list of commands, man-page style
|
|
1113
|
+
$ agsh read --help # man-page style: NAME, SYNOPSIS, DESCRIPTION, OPTIONS
|
|
1114
|
+
$ agsh read /repo/README.md
|
|
1115
|
+
$ agsh datadog_list_monitors --query "service:api"
|
|
1116
|
+
$ agsh # bare: interactive shell on a TTY, script from stdin otherwise
|
|
1117
|
+
❯ ls /repo | grep -i readme
|
|
1118
|
+
❯ read /repo/config.json | jq .version
|
|
1119
|
+
```
|
|
1120
|
+
|
|
1121
|
+
Every invocation binds to the deployment's latest session by default, with
|
|
1122
|
+
`--session` and `--continuation-token` overrides, and prints the session
|
|
1123
|
+
identifier as a final stderr line.
|
|
1124
|
+
|
|
1125
|
+
## Configuration
|
|
1126
|
+
|
|
1127
|
+
`agsh` is a client only; it never boots an agent. Every invocation needs a
|
|
1128
|
+
target deployment, given by flags or by environment variables. Flags always
|
|
1129
|
+
win over the environment.
|
|
1130
|
+
|
|
1131
|
+
Global command line options, accepted on every command and on the bare shell
|
|
1132
|
+
launch:
|
|
1133
|
+
|
|
1134
|
+
| Option | Environment default | Meaning |
|
|
1135
|
+
| --- | --- | --- |
|
|
1136
|
+
| `--target <url \| name>` | `AGENT_SHELL_TARGET` | The deployment to talk to: a URL is a local deployment (`http://127.0.0.1:39400/executor`), a name a production one (`change-monitor-executor`). |
|
|
1137
|
+
| `--team <team>` | `AGENT_SHELL_TEAM` | Team override for production resolution, when the login spans several. |
|
|
1138
|
+
| `--bearer-token <token>` | `AGENT_SHELL_BEARER_TOKEN` | Explicit bearer auth for a deployment that is not behind the Cursor login. |
|
|
1139
|
+
| `--session <id>` | | Bind to a specific session instead of the latest. |
|
|
1140
|
+
| `--continuation-token <token>` | | Bind by continuation token instead of session id. |
|
|
1141
|
+
| `--output <text\|json>` | | Result rendering: human-friendly views (default) or raw JSON. |
|
|
1142
|
+
| `-h`, `--help` | | Per-command help. |
|
|
1143
|
+
|
|
1144
|
+
One parameter carries the whole target selection, and the value's shape
|
|
1145
|
+
encodes the mode: a URL (`http://` or `https://`) targets a local
|
|
1146
|
+
deployment, anything else names a production one. Two options with a
|
|
1147
|
+
precedence rule would invite exactly the confusion a target selector must
|
|
1148
|
+
not have; with one parameter the only rule is that the flag beats the
|
|
1149
|
+
environment. A URL is self-contained down to the agent because one local
|
|
1150
|
+
agent-sdk serve process hosts every agent of the project (change-monitor's
|
|
1151
|
+
dev stack mounts `/executor` and `/planner` from a single port); a
|
|
1152
|
+
production deployment is a single agent, so its name is the complete
|
|
1153
|
+
address (`--team` narrows resolution when the login spans several).
|
|
1154
|
+
Authentication defaults to the stored Cursor login (the same engine-access
|
|
1155
|
+
credential agent-sdk uses); `--bearer-token` is the escape hatch for direct
|
|
1156
|
+
deployments. Session flags are per invocation and have no environment
|
|
1157
|
+
default: a session is state, not configuration. Color output follows the
|
|
1158
|
+
`NO_COLOR` convention and TTY detection; there is no agsh-specific color
|
|
1159
|
+
setting. No configuration file: one environment variable pins a working
|
|
1160
|
+
target for a terminal session
|
|
1161
|
+
(`AGENT_SHELL_TARGET=change-monitor-executor`, or a URL for a local stack),
|
|
1162
|
+
which is the whole persistent-configuration need.
|
|
1163
|
+
|
|
1164
|
+
With no target from flags or environment, every command fails with a message
|
|
1165
|
+
naming both ways to provide one.
|
|
1166
|
+
|
|
1167
|
+
## Architecture
|
|
1168
|
+
|
|
1169
|
+
### A new package
|
|
1170
|
+
|
|
1171
|
+
A new workspace package (working name `packages/agsh`, bin `agsh`) that
|
|
1172
|
+
depends on `@cursor/july` for target resolution, stored Cursor login, and the
|
|
1173
|
+
HTTP client plumbing. It owns the presentation stack: `marked` for terminal
|
|
1174
|
+
markdown (moved out of `@cursor/july`), with syntax highlighting (`shiki`)
|
|
1175
|
+
arriving in the phase that renders code; the shell interpreter is
|
|
1176
|
+
purpose-built (see Rationale).
|
|
1177
|
+
No new abstraction seam between the two packages; `agsh` imports what it
|
|
1178
|
+
needs until a second consumer justifies extracting a thin client.
|
|
1179
|
+
|
|
1180
|
+
### The tool catalog
|
|
1181
|
+
|
|
1182
|
+
At startup `agsh` fetches one live catalog of everything invocable on the
|
|
1183
|
+
deployment. This is the piece the current `/v1/info` cannot provide: `/v1/info`
|
|
1184
|
+
projects the authored manifest, and connection tools only exist at runtime,
|
|
1185
|
+
resolved per session under the connection's auth. A new endpoint provides the
|
|
1186
|
+
live view (see Backend changes).
|
|
1187
|
+
|
|
1188
|
+
Catalog entries carry exactly one identifier each: the tool name exactly as
|
|
1189
|
+
the agent sees it. Authored server tools keep their authored name (`read`).
|
|
1190
|
+
Connection tools appear under their model-facing advertised name (the
|
|
1191
|
+
sanitized passthrough name from `advertise-tools.ts`, e.g.
|
|
1192
|
+
`datadog_list_monitors`). The CLI never invents a different naming format:
|
|
1193
|
+
a tool name copied from a session transcript is a valid `agsh` command, and
|
|
1194
|
+
vice versa. Where a tool came from — the upstream connector name when the
|
|
1195
|
+
tool declares one, the connection name otherwise — is a field on the
|
|
1196
|
+
catalog entry, not part of the identifier.
|
|
1197
|
+
|
|
1198
|
+
### Two command tiers
|
|
1199
|
+
|
|
1200
|
+
Each catalog entry becomes a command, through one of two shapes:
|
|
1201
|
+
|
|
1202
|
+
**Curated commands for builtin tools.** The well-known tool names (`ls`,
|
|
1203
|
+
`read`, `grep`, `glob`, `diff`, ...) get hand-designed, POSIX-flavored
|
|
1204
|
+
command shapes, hardcoded in `agsh` next to their titles. These tools are
|
|
1205
|
+
what an operator types all day; their shapes should feel like the unix
|
|
1206
|
+
commands they mirror, not like generated bindings. A curated shape decides
|
|
1207
|
+
which schema fields are positional operands and which are flags, and every
|
|
1208
|
+
input has exactly one spelling: an operand is only an operand, never also a
|
|
1209
|
+
flag.
|
|
1210
|
+
|
|
1211
|
+
```
|
|
1212
|
+
$ agsh read /repo/package.json --limit 2
|
|
1213
|
+
{
|
|
1214
|
+
"name": "change-monitor",
|
|
1215
|
+
→ ses_a99d1b69c329eb75a2ec8603
|
|
1216
|
+
|
|
1217
|
+
$ agsh grep -i -A 2 toolEffect /repo/src
|
|
1218
|
+
src/tool-policy.ts:12:export type ToolEffect = "read" | "write";
|
|
1219
|
+
...
|
|
1220
|
+
→ ses_a99d1b69c329eb75a2ec8603
|
|
1221
|
+
|
|
1222
|
+
$ agsh ls /repo --ignore-globs '*.test.ts' --ignore-globs 'node_modules/**'
|
|
1223
|
+
```
|
|
1224
|
+
|
|
1225
|
+
`read` takes its path as an operand mapped to the schema's `path` field, with
|
|
1226
|
+
`--offset` and `--limit` as integer flags. `grep` follows POSIX grep:
|
|
1227
|
+
`grep [options] <pattern> [path]`, with the rg-style options (`-i`, `-A`,
|
|
1228
|
+
`-B`, `-C`, `--output-mode`, `--head-limit`) mapping onto the schema fields
|
|
1229
|
+
of the same names (kebab-cased). `ls` shows array input: an array field's flag repeats once
|
|
1230
|
+
per element. A curated shape binds to the deployment's live schema at
|
|
1231
|
+
startup; when a deployment's tool lacks the expected field, the command
|
|
1232
|
+
degrades to the generic shape below rather than guessing.
|
|
1233
|
+
|
|
1234
|
+
A curated shape may also reformat the tool's text result toward the unix
|
|
1235
|
+
command's own output conventions: the VFS ls tool returns the model-facing
|
|
1236
|
+
tree (` - name/` rows under a header), and `agsh ls` prints it as standard
|
|
1237
|
+
ls does, one name per line with the trailing slash kept on directories. The
|
|
1238
|
+
tool's result string itself stays what the model sees; when a result does
|
|
1239
|
+
not match the expected shape it prints verbatim.
|
|
1240
|
+
|
|
1241
|
+
**Generated commands for MCP tools.** Connection tools are dynamically
|
|
1242
|
+
discovered, so no special treatment is possible; they get a uniform
|
|
1243
|
+
schema-derived mapping:
|
|
1244
|
+
|
|
1245
|
+
- Every schema property is accepted as one flag, spelled as the
|
|
1246
|
+
kebab-cased property name (`org_slug` → `--org-slug`) — the unix
|
|
1247
|
+
convention; kebab collisions gain a numeric suffix. Properties already
|
|
1248
|
+
shaped like flags (grep's `-i`) stay literal. No positionals, no other
|
|
1249
|
+
aliases.
|
|
1250
|
+
- Object-typed properties flatten recursively into one flag per leaf,
|
|
1251
|
+
dash-joined (`--telemetry-context` for `telemetry.context`), so every
|
|
1252
|
+
option reads as a plain value; a free-form object with no declared
|
|
1253
|
+
properties stays one JSON-valued flag. A leaf is required only when its
|
|
1254
|
+
whole ancestor chain is.
|
|
1255
|
+
- Values are coerced by schema type: booleans are valueless flags, numbers
|
|
1256
|
+
and integers are parsed, arrays accept the flag repeated once per element,
|
|
1257
|
+
enums are validated before the call.
|
|
1258
|
+
|
|
1259
|
+
```
|
|
1260
|
+
$ agsh datadog_list_monitors --query "service:api" --limit 10
|
|
1261
|
+
```
|
|
1262
|
+
|
|
1263
|
+
In both tiers `-h`/`--help` and the global target and session flags are
|
|
1264
|
+
reserved and injected, a flag that names no schema property fails before any
|
|
1265
|
+
request (listing the tool's actual properties), and the bound session prints
|
|
1266
|
+
as a final stderr line.
|
|
1267
|
+
|
|
1268
|
+
### Result rendering
|
|
1269
|
+
|
|
1270
|
+
Raw JSON on a terminal is not an experience for people, so `--output=text`
|
|
1271
|
+
(the default) renders structured results through a small set of views,
|
|
1272
|
+
selected automatically by the shape of the value each call actually returned;
|
|
1273
|
+
tool metadata plays no part, since most tools advertise no output schema, and
|
|
1274
|
+
many return structured data as JSON text. A string result that parses as a
|
|
1275
|
+
JSON object or array counts as structured. An array of objects renders as a
|
|
1276
|
+
table (columns are the union of keys, missing cells stay blank, the table
|
|
1277
|
+
clamps to the terminal width); a single object renders as a property view
|
|
1278
|
+
(aligned keys, scalar lists as bullets, nested structures indented); an
|
|
1279
|
+
object that is nothing but an error wrapper renders as an `Error:` line;
|
|
1280
|
+
plain text prints verbatim. `--output=json` renders the structured value as
|
|
1281
|
+
raw JSON. The rendering never depends on the TTY: piped and interactive
|
|
1282
|
+
output carry the same content, only color follows TTY detection.
|
|
1283
|
+
|
|
1284
|
+
### Help rendering
|
|
1285
|
+
|
|
1286
|
+
`--help` on a tool renders a man-page layout: NAME (the tool name, with the
|
|
1287
|
+
tool's `title` beside it when the catalog carries one; titles are curated
|
|
1288
|
+
data, never derived from the description), SYNOPSIS (operands from the
|
|
1289
|
+
curated shape; options never enumerate — they summarize as `[options...]`,
|
|
1290
|
+
man-page style, so the line stays bounded), DESCRIPTION (the tool
|
|
1291
|
+
description rendered as terminal markdown), OPERANDS (positional arguments,
|
|
1292
|
+
curated commands only), and OPTIONS. Descriptions of operands and options
|
|
1293
|
+
come from the schema's property descriptions. Effect and approval metadata
|
|
1294
|
+
render as notes when declared. Everything except the curated shape derives
|
|
1295
|
+
from `GET /v1/tools/:name`; nothing else is hand-written per tool.
|
|
1296
|
+
|
|
1297
|
+
```
|
|
1298
|
+
$ agsh read --help
|
|
1299
|
+
NAME
|
|
1300
|
+
read - Read a file
|
|
1301
|
+
|
|
1302
|
+
SYNOPSIS
|
|
1303
|
+
read [options...] <path>
|
|
1304
|
+
|
|
1305
|
+
DESCRIPTION
|
|
1306
|
+
Reads a file from the local filesystem. This tool can also read image
|
|
1307
|
+
files when called with the appropriate path. Formats supported:
|
|
1308
|
+
jpeg/jpg, png, gif, webp.
|
|
1309
|
+
|
|
1310
|
+
OPERANDS
|
|
1311
|
+
<path>
|
|
1312
|
+
The absolute path of the file to read.
|
|
1313
|
+
|
|
1314
|
+
OPTIONS
|
|
1315
|
+
--offset <integer>
|
|
1316
|
+
The line number to start reading from. Positive values are 1-indexed
|
|
1317
|
+
from the start of the file. Negative values count backwards from the
|
|
1318
|
+
end. Only provide if the file is too large to read at once.
|
|
1319
|
+
|
|
1320
|
+
--limit <integer>
|
|
1321
|
+
The number of lines to read. Only provide if the file is too large
|
|
1322
|
+
to read at once.
|
|
1323
|
+
|
|
1324
|
+
NOTES
|
|
1325
|
+
Effect: read (performs no writes).
|
|
1326
|
+
```
|
|
1327
|
+
|
|
1328
|
+
`agsh help` lists the available command names grouped by source, authored
|
|
1329
|
+
tools first, then one group per upstream connector (its name is the group
|
|
1330
|
+
header — one aggregating connection can host tools from several connectors,
|
|
1331
|
+
and the connector name is what an operator recognizes). Each row is the
|
|
1332
|
+
name, with the title beside it when the tool declares one; everything else
|
|
1333
|
+
lives behind the command's `--help`:
|
|
1334
|
+
|
|
1335
|
+
The agent's description renders as a DESCRIPTION section when the deployment
|
|
1336
|
+
declares one (`/v1/info` carries both name and description).
|
|
1337
|
+
|
|
1338
|
+
```
|
|
1339
|
+
$ agsh help
|
|
1340
|
+
NAME
|
|
1341
|
+
change-monitor-executor
|
|
1342
|
+
|
|
1343
|
+
DESCRIPTION
|
|
1344
|
+
Executes monitoring plans against changed code.
|
|
1345
|
+
|
|
1346
|
+
COMMANDS
|
|
1347
|
+
diff Show workspace changes
|
|
1348
|
+
glob Find files by pattern
|
|
1349
|
+
grep Search file contents
|
|
1350
|
+
ls List a directory
|
|
1351
|
+
read Read a file
|
|
1352
|
+
report_change_issue
|
|
1353
|
+
report_change_succeeded
|
|
1354
|
+
|
|
1355
|
+
DATADOG
|
|
1356
|
+
datadog_list_monitors List monitors
|
|
1357
|
+
...
|
|
1358
|
+
|
|
1359
|
+
Run any command with --help for its synopsis and options.
|
|
1360
|
+
```
|
|
1361
|
+
|
|
1362
|
+
### Shell mode
|
|
1363
|
+
|
|
1364
|
+
Invoked bare, `agsh` starts a shell. On a TTY this is a REPL; on a pipe it
|
|
1365
|
+
reads a script from stdin, so `echo 'ls /' | agsh` and here-docs work.
|
|
1366
|
+
|
|
1367
|
+
The interpreter is purpose-built and minimal: tokenizing (quotes, escapes),
|
|
1368
|
+
pipelines, and `;` / `&&` / `||`. The command namespace is exactly the
|
|
1369
|
+
deployment's tool catalog plus a small curated set of local pipe filters
|
|
1370
|
+
(`head`, `tail`, `wc`, stdin-filtering `grep`), so a tool name can never be
|
|
1371
|
+
shadowed. There is no local filesystem, no variables, no control flow: agsh
|
|
1372
|
+
has nothing local to operate on, and every command is a single traced
|
|
1373
|
+
`POST /v1/tools/:name` call.
|
|
1374
|
+
|
|
1375
|
+
The shell binds one session identity at launch (latest by default) and keeps
|
|
1376
|
+
it for the whole run, so a sequence of tool calls observes one consistent
|
|
1377
|
+
session context.
|
|
1378
|
+
|
|
1379
|
+
## Backend changes on the agent-sdk runtime
|
|
1380
|
+
|
|
1381
|
+
Two read endpoints, mirroring the invocation path:
|
|
1382
|
+
|
|
1383
|
+
**`GET /v1/tools`: the live tool listing.** Returns the session's tool
|
|
1384
|
+
namespace exactly as a turn would assemble it: authored server tools plus the
|
|
1385
|
+
advertised passthrough tools synthesized from connections, under their
|
|
1386
|
+
model-facing names. Entries are light (name, source, and `title` when one is
|
|
1387
|
+
known); everything else lives behind the detail endpoint. Titles have two
|
|
1388
|
+
sources and no new authoring surface: connection tools inherit the upstream
|
|
1389
|
+
server's MCP title, which the host already propagates length-capped off
|
|
1390
|
+
listings; tools that do not come from MCP get theirs from a hardcoded
|
|
1391
|
+
name-to-title table in the runtime's endpoint implementation, covering the
|
|
1392
|
+
well-known tool names. A tool in neither place has no title. Accepts the same
|
|
1393
|
+
optional session binding as invocation (`session` or `continuationToken`)
|
|
1394
|
+
because advertised inventories can be tenant-scoped and resolved per session.
|
|
1395
|
+
Implementation reuses the existing plumbing: the discovered manifest for
|
|
1396
|
+
authored tools and the advertise-tools synthesis (`McpHost.listTools`, or the
|
|
1397
|
+
`oneOff` path when per-session auth substitution applies) for connection
|
|
1398
|
+
tools. This is not a duplicate of `/v1/info`: the info document stays the
|
|
1399
|
+
static authored manifest; the listing is the runtime view that only the
|
|
1400
|
+
running deployment can answer.
|
|
1401
|
+
|
|
1402
|
+
**`GET /v1/tools/:name`: one tool's full description.** Description, input
|
|
1403
|
+
schema, output schema when declared, effect when declared, approval
|
|
1404
|
+
requirement, and source connection. Same path as invocation
|
|
1405
|
+
(`POST /v1/tools/:name`), different method: GET describes what POST executes,
|
|
1406
|
+
for the same identifier.
|
|
1407
|
+
|
|
1408
|
+
Invocation needs no new naming scheme. Advertised connection tools are
|
|
1409
|
+
synthesized as ordinary server tools in the session's namespace, so
|
|
1410
|
+
`POST /v1/tools/:name` addresses them by their model-facing name like any
|
|
1411
|
+
authored tool, with the same session binding, policy checks, and per-call
|
|
1412
|
+
tracing. (The direct-call path did need the synthesis step added: it now
|
|
1413
|
+
resolves the advertised listing for the call's session identity when the
|
|
1414
|
+
authored lookup misses.)
|
|
1415
|
+
|
|
1416
|
+
Phase 1 ships the minimal runtime surface agsh calls: the `effect`
|
|
1417
|
+
projection in `/v1/info` (rendered in per-tool help), the scratch-workspace
|
|
1418
|
+
fallback on direct calls, and `continuationToken` binding on
|
|
1419
|
+
`POST /v1/tools/:toolName`. The detail endpoint in phase 2 also closes the
|
|
1420
|
+
output-schema gap; `/v1/info` stays as it is.
|
|
1421
|
+
|
|
1422
|
+
## Local development loop
|
|
1423
|
+
|
|
1424
|
+
`factory/change-monitor` is the test bed. Its `pnpm start` already serves the
|
|
1425
|
+
planner and executor locally through the agent-sdk dev runtime
|
|
1426
|
+
(`agent-sdk serve --dir . --dev`). The loop:
|
|
1427
|
+
|
|
1428
|
+
1. `cd factory/change-monitor && pnpm start` (local stack, both agents).
|
|
1429
|
+
2. `agsh --target http://127.0.0.1:<port>/<agent>` against it, via a dev shim
|
|
1430
|
+
analogous to `agent-sdk-dev` so the CLI runs from the worktree.
|
|
1431
|
+
3. Iterate end to end: VFS verbs (`ls`, `read`, `grep`, `glob`, `diff`) for the
|
|
1432
|
+
authored-tool path, and the planner's tenant connectors for the
|
|
1433
|
+
connection-tool path once `GET /v1/tools` exists.
|
|
1434
|
+
|
|
1435
|
+
## Removing the inspector surface from agent-sdk
|
|
1436
|
+
|
|
1437
|
+
The inspector CLI is still on development branches, so nothing migrates: the
|
|
1438
|
+
CLI-side code is removed from `agent-sdk` and `agsh` is built in its place.
|
|
1439
|
+
|
|
1440
|
+
- The verb commands (`ls`, `read`, `grep`, `glob`, `diff`) become the
|
|
1441
|
+
curated tier: their hand-designed shapes, schema-binding logic (including
|
|
1442
|
+
the candidate-field fallback), and session binding carry over. The
|
|
1443
|
+
schema-to-argv flag mapping seeds the generated tier for MCP tools.
|
|
1444
|
+
- The `tools` and `skills` commands disappear entirely. `agsh help` and
|
|
1445
|
+
per-tool `--help` are the discovery surface.
|
|
1446
|
+
- `marked` and `shiki` leave `@cursor/july`; agsh's help rendering takes
|
|
1447
|
+
`marked`, and `shiki` returns when agsh ships syntax highlighting.
|
|
1448
|
+
`agent-sdk` keeps its developer workflow commands unchanged.
|
|
1449
|
+
|
|
1450
|
+
## Plan
|
|
1451
|
+
|
|
1452
|
+
1. **Package and core invocation.** Create the package, port target
|
|
1453
|
+
resolution, the schema-to-argv mapping, and help rendering from the
|
|
1454
|
+
inspector code. Authored tools only, against the existing endpoints.
|
|
1455
|
+
Verified end to end on the local change-monitor stack.
|
|
1456
|
+
2. **Live catalog.** Add `GET /v1/tools` and `GET /v1/tools/:name` to
|
|
1457
|
+
the agent-sdk runtime, with the hardcoded title table for non-MCP tools, and verify
|
|
1458
|
+
direct invocation resolves advertised connection tools by their
|
|
1459
|
+
model-facing names. Connection tools appear as commands. Verified against
|
|
1460
|
+
the planner's connectors.
|
|
1461
|
+
3. **Shell mode.** The purpose-built mini-shell: REPL on TTY, script on
|
|
1462
|
+
stdin, tools as the command namespace, one session per shell run.
|
|
1463
|
+
4. **Cleanup.** Remove the inspector CLI surface and presentation
|
|
1464
|
+
dependencies from `@cursor/july`.
|
|
1465
|
+
|
|
1466
|
+
## Rationale and rejected alternatives
|
|
1467
|
+
|
|
1468
|
+
**Why not extend `agent-sdk`.** The audiences differ: `agent-sdk` is for the
|
|
1469
|
+
person building and deploying an agent; this tool is for the person operating
|
|
1470
|
+
inside one. Bundling also forces every agent project to carry markdown
|
|
1471
|
+
rendering, syntax highlighting, and a bash interpreter it never uses.
|
|
1472
|
+
|
|
1473
|
+
**Name.** `agsh` reads as "agent shell", is four characters, collides with
|
|
1474
|
+
nothing common, and works as a shell prompt name. Considered: `august`
|
|
1475
|
+
(pairs with `july` but says nothing about purpose), `toolsh` (awkward to
|
|
1476
|
+
pronounce), `cursor-shell` (too broad; this is scoped to one agent).
|
|
1477
|
+
|
|
1478
|
+
**Why a REST catalog instead of the MCP endpoint.** The deployment already
|
|
1479
|
+
speaks MCP at `/v1/mcp/tools`, including a per-connection bridge, but the
|
|
1480
|
+
bridge is bound to an active turn and speaks JSON-RPC. The CLI wants a plain
|
|
1481
|
+
authenticated GET with session binding that returns the assembled tool
|
|
1482
|
+
namespace under the names the model sees. Wrapping that in MCP framing buys
|
|
1483
|
+
nothing for a first-party client.
|
|
1484
|
+
|
|
1485
|
+
**Why a purpose-built interpreter instead of just-bash.** just-bash was the
|
|
1486
|
+
original plan (a full bash emulation with a custom-command extension point),
|
|
1487
|
+
and a prototype disproved it: custom commands replace its coreutils but can
|
|
1488
|
+
never shadow its shell builtins, and `read`, `test`, `type`, and `help` are
|
|
1489
|
+
builtins — so the flagship `read` tool is unreachable, and the precedence is
|
|
1490
|
+
not ours to control (vercel-labs owns the package). No other embeddable JS
|
|
1491
|
+
shell interpreter has a workable custom-command story (mvdan-sh's JS build
|
|
1492
|
+
does not expose one; bash-parser is a parser only). agsh also needs almost
|
|
1493
|
+
none of bash: no local filesystem, no variables, no control flow — just
|
|
1494
|
+
tokenizing, pipelines, and a command namespace it fully owns. A TypeScript
|
|
1495
|
+
REPL with tools as async functions (the shape of change-monitor's `script`
|
|
1496
|
+
tool) was considered and kept as a possible later addition; it trades away
|
|
1497
|
+
the unix muscle memory the curated commands exist for.
|
|
1498
|
+
|
|
1499
|
+
---
|
|
1500
|
+
|
|
1108
1501
|
Source: /docs/evals.md
|
|
1109
1502
|
|
|
1110
1503
|
# Evals
|
|
@@ -1294,2749 +1687,279 @@ Assert with the gates:
|
|
|
1294
1687
|
| `t.eventOrder(matchers)` | matching event groups occur in this relative order |
|
|
1295
1688
|
| `t.eventsSatisfy(label, predicate)` | your predicate over the typed event stream |
|
|
1296
1689
|
| `t.check(value, expectation)` | any value, against a builder |
|
|
1297
|
-
| `t.score(name, value)` | records a 0–1 score you computed; soft until you add a bar |
|
|
1298
|
-
| `t.requireToolCall(name, matcher?)` | gates on a matching call and returns it, so later code can read its input and output |
|
|
1299
|
-
| `t.requireInputRequest(filter?)` | gates on exactly one pending approval request and returns it |
|
|
1300
|
-
|
|
1301
|
-
Every gate returns a handle: `.soft()` demotes it to tracked-only,
|
|
1302
|
-
`.atLeast(0.7)` adds a soft score bar, and `.gate(0.8)` promotes a
|
|
1303
|
-
scored assertion into a hard gate.
|
|
1304
|
-
|
|
1305
|
-
With no matcher, `calledTool` is request-based: a requested call counts
|
|
1306
|
-
even when its result has not arrived. Pass
|
|
1307
|
-
`t.calledTool("inspect_pr", { status: "completed" })` to require the
|
|
1308
|
-
call to return. `input`, `output`, and `count` matcher fields accept a
|
|
1309
|
-
literal, a `RegExp`, or a predicate.
|
|
1310
|
-
|
|
1311
|
-
The expectation builders are `includes(string | RegExp)`,
|
|
1312
|
-
`equals(value)`, `matches(schema)`, `similarity(expected)`, and
|
|
1313
|
-
`satisfies(predicate, label)`. `includes` stringifies its input,
|
|
1314
|
-
`equals` compares values deeply, `matches` validates against a Standard
|
|
1315
|
-
Schema (or anything with `safeParse`, like Zod), `similarity` scores
|
|
1316
|
-
normalized text similarity, and `satisfies` runs your predicate. The
|
|
1317
|
-
plain function `normalizedSimilarity(actual, expected)` returns the
|
|
1318
|
-
same 0–1 score for use with `t.score`.
|
|
1319
|
-
|
|
1320
|
-
A few more context members shape a case: `t.require(value, expectation)`
|
|
1321
|
-
records a gate and stops the test body when it fails, without a
|
|
1322
|
-
duplicate execution error. `t.skip(reason)` ends the case as skipped
|
|
1323
|
-
(reported separately, never changes the exit code; call it before
|
|
1324
|
-
sending messages). `t.metric(name, value)` records a structured score
|
|
1325
|
-
for the playground case card. `t.log(message)` records a debug line for
|
|
1326
|
-
the CLI and playground result.
|
|
1327
|
-
|
|
1328
|
-
Three `t.send` options apply on session create (first `t.send` only):
|
|
1329
|
-
|
|
1330
|
-
- `workspaceFiles`: `{ path: contents }`, seeded into the local session
|
|
1331
|
-
workspace. Prefer this over machine-local paths.
|
|
1332
|
-
- `workspaceDir`: absolute harness cwd (local runtime).
|
|
1333
|
-
- `cloud`: per-session cloud options merged over the agent's static
|
|
1334
|
-
`cloud` config (repos / env / …). Use a pinned `repos` override to
|
|
1335
|
-
attach a fixture repo for cloud evals without putting it on the
|
|
1336
|
-
agent's default `cloud.repos`. Cloud ignores `workspaceFiles` seeds.
|
|
1337
|
-
|
|
1338
|
-
```ts
|
|
1339
|
-
const toolResults = t.events.filter((e) => e.type === "action.result");
|
|
1340
|
-
t.check(
|
|
1341
|
-
toolResults.length,
|
|
1342
|
-
satisfies((n) => (n as number) <= 4, "at most 4 tool calls")
|
|
1343
|
-
);
|
|
1344
|
-
```
|
|
1345
|
-
|
|
1346
|
-
A case with no explicit gates falls back to whether at least one turn
|
|
1347
|
-
completed successfully. Add `t.succeeded()` and behavior-specific gates
|
|
1348
|
-
anyway. They make the contract visible during review.
|
|
1349
|
-
|
|
1350
|
-
### Judge free-form output
|
|
1351
|
-
|
|
1352
|
-
When wording matters and no regex captures it, `t.judge` grades the
|
|
1353
|
-
reply with an LLM. The built-in graders are `factuality(expected)`,
|
|
1354
|
-
`summarizes(expected)`, `closedQA(criteria)`, and `sql(expected)`. Each
|
|
1355
|
-
scores `t.reply` by default; pass `{ on }` to grade another value.
|
|
1356
|
-
|
|
1357
|
-
```ts
|
|
1358
|
-
t.judge.factuality("It is 54°F in NYC right now.").atLeast(0.7);
|
|
1359
|
-
```
|
|
1360
|
-
|
|
1361
|
-
Judge assertions are soft by default, so a judge never fails a build
|
|
1362
|
-
until you give it a bar with `.atLeast(0.7)` or promote it with
|
|
1363
|
-
`.gate(0.8)`. The judge model comes from `defineEvalConfig({ judge })`,
|
|
1364
|
-
`defineEval({ judge })`, a case-level `judge`, or a per-call
|
|
1365
|
-
`{ model }` override; the nearest one wins. For a domain-specific judge
|
|
1366
|
-
whose verdict is not a single score, `t.judge.model(prompt)` sends a
|
|
1367
|
-
raw prompt to the same model and returns the reply. You then record the
|
|
1368
|
-
parsed result with `t.score` or `t.check`.
|
|
1369
|
-
|
|
1370
|
-
## Run evals from the CLI
|
|
1371
|
-
|
|
1372
|
-
The `eval` command discovers, filters, and runs cases.
|
|
1373
|
-
|
|
1374
|
-
Run the CLI under Node 22.13 or newer. Do not use Bun. Its HTTP/2 client
|
|
1375
|
-
breaks tool-result streams and causes eval turns to fail.
|
|
1376
|
-
|
|
1377
|
-
```bash
|
|
1378
|
-
agent-sdk eval --dir . --list # discover only
|
|
1379
|
-
agent-sdk eval --dir . # run all
|
|
1380
|
-
agent-sdk eval --dir . builds/checkout # one datapoint
|
|
1381
|
-
agent-sdk eval --dir . builds search # several ids or prefixes
|
|
1382
|
-
agent-sdk eval --dir . --tag smoke --tag pull-request # any matching tag
|
|
1383
|
-
agent-sdk eval --dir . --json --no-stream # machine-readable results
|
|
1384
|
-
agent-sdk eval --dir . --verbose # logs + reply snippets
|
|
1385
|
-
```
|
|
1386
|
-
|
|
1387
|
-
Id filters use OR semantics. Each filter selects an exact id and its
|
|
1388
|
-
descendants. For example, `builds` selects `builds`,
|
|
1389
|
-
`builds/checkout`, and every other case below that path. Repeated tags
|
|
1390
|
-
also use OR semantics. When you provide both ids and tags, a case must
|
|
1391
|
-
match both groups.
|
|
1392
|
-
|
|
1393
|
-
`eval` boots an ephemeral server on port 0 with a temp state root
|
|
1394
|
-
outside the project, so cases don't inherit ambient monorepo rules and
|
|
1395
|
-
don't write into the project state directory. Point `--url` at a running server to eval
|
|
1396
|
-
a live agent instead:
|
|
1397
|
-
|
|
1398
|
-
```bash
|
|
1399
|
-
agent-sdk eval --dir . \
|
|
1400
|
-
--url http://127.0.0.1:3000/weather-agent \
|
|
1401
|
-
--bearer-token "$AGENT_TOKEN"
|
|
1402
|
-
```
|
|
1403
|
-
|
|
1404
|
-
The eval definitions still come from `--dir`; `--url` only changes the
|
|
1405
|
-
agent that receives the turns. For a locally mounted multi-agent
|
|
1406
|
-
directory, `--slug weather-agent` chooses the target. Use
|
|
1407
|
-
`--state-root` to keep ephemeral session state at a chosen path,
|
|
1408
|
-
`--timeout-ms` to override the project timeout, and `--no-stream` to
|
|
1409
|
-
keep live progress off stderr. A TTY streams turn progress by default.
|
|
1410
|
-
`--verbose` still writes `t.log` lines to stderr and adds reply snippets
|
|
1411
|
-
to text results.
|
|
1412
|
-
|
|
1413
|
-
Model turns need a Cursor credential from `agent-sdk login` or
|
|
1414
|
-
`CURSOR_API_KEY`.
|
|
1415
|
-
|
|
1416
|
-
See [CLI: eval](/docs/reference/cli.md#eval) for flags and exit codes.
|
|
1417
|
-
|
|
1418
|
-
### JSON results
|
|
1419
|
-
|
|
1420
|
-
Use `--json --no-stream` in scripts and CI. The top-level result carries
|
|
1421
|
-
the totals and one result per case:
|
|
1422
|
-
|
|
1423
|
-
```json
|
|
1424
|
-
{
|
|
1425
|
-
"ok": true,
|
|
1426
|
-
"passed": 1,
|
|
1427
|
-
"failed": 0,
|
|
1428
|
-
"results": [
|
|
1429
|
-
{
|
|
1430
|
-
"id": "readiness",
|
|
1431
|
-
"ok": true,
|
|
1432
|
-
"assertions": [{ "name": "succeeded", "passed": true }],
|
|
1433
|
-
"sessionId": "ses_123",
|
|
1434
|
-
"inputs": ["Is checkout pull request 42 ready to approve?"],
|
|
1435
|
-
"toolCalls": [{ "toolName": "inspect_pr", "isError": false }],
|
|
1436
|
-
"logs": [],
|
|
1437
|
-
"durationMs": 12340
|
|
1438
|
-
}
|
|
1439
|
-
]
|
|
1440
|
-
}
|
|
1441
|
-
```
|
|
1442
|
-
|
|
1443
|
-
Each case result can also include `description`, `finalText`, `tools`,
|
|
1444
|
-
`error`, and tool arguments or output. This shape lets CI report the
|
|
1445
|
-
failed assertion without parsing terminal text.
|
|
1446
|
-
|
|
1447
|
-
## Run evals in the playground
|
|
1448
|
-
|
|
1449
|
-
Start the server, open the playground, and choose **Evals**. You can run
|
|
1450
|
-
every case or one case, watch progress, and open the resulting session
|
|
1451
|
-
trace. The Evals tab works on a normal `serve`.
|
|
1452
|
-
|
|
1453
|
-
```bash
|
|
1454
|
-
agent-sdk serve --dir .
|
|
1455
|
-
```
|
|
1456
|
-
|
|
1457
|
-
Playground runs target the live server instead of an ephemeral one.
|
|
1458
|
-
Their sessions appear in the session list. One eval batch can run at a
|
|
1459
|
-
time. Persistence follows the rule under
|
|
1460
|
-
[Configure eval runs](#configure-eval-runs). See
|
|
1461
|
-
[Playground eval routes](/docs/reference/http-api.md#playground-eval-routes).
|
|
1462
|
-
The start request returns `202` while cases run in the background.
|
|
1463
|
-
Poll until the snapshot status becomes `completed`, `failed`, or `cancelled`.
|
|
1464
|
-
Configuration errors appear on a failed snapshot.
|
|
1465
|
-
|
|
1466
|
-
On `--prod` / `--url`, the CLI prints the Eval ID as soon as the batch is
|
|
1467
|
-
accepted (and a Playground deep link with `?view=evals&evalRunId=…`):
|
|
1468
|
-
|
|
1469
|
-
```bash
|
|
1470
|
-
agent-sdk eval --prod --slug vulnerability-scanner --tag deepsec
|
|
1471
|
-
# Eval ID: evalrun_…
|
|
1472
|
-
# Cancel: agent-sdk eval cancel evalrun_… --prod --slug vulnerability-scanner
|
|
1473
|
-
# Playground: https://…/playground?view=evals&evalRunId=evalrun_…
|
|
1474
|
-
|
|
1475
|
-
agent-sdk eval cancel evalrun_… --prod --slug vulnerability-scanner
|
|
1476
|
-
agent-sdk eval status evalrun_… --prod --slug vulnerability-scanner
|
|
1477
|
-
```
|
|
1478
|
-
|
|
1479
|
-
## What good cases assert
|
|
1480
|
-
|
|
1481
|
-
Gate decisions and shape, not prose. Model wording varies run to run.
|
|
1482
|
-
Tool choice, tool avoidance, and output structure are the stable
|
|
1483
|
-
contract.
|
|
1484
|
-
|
|
1485
|
-
1. `t.succeeded()`: always, first.
|
|
1486
|
-
2. The tool decision: `calledTool` for the intended path,
|
|
1487
|
-
`notCalledTool` for the likely wrong alternative. The pair is
|
|
1488
|
-
stronger than either alone.
|
|
1489
|
-
3. Output shape: a regex for the contract (`/ready|blocked/i`, a JSON
|
|
1490
|
-
marker, a findings-block fence), never exact sentences.
|
|
1491
|
-
4. For structured output, parse `t.reply` and check fields with
|
|
1492
|
-
`satisfies` instead of substring-matching JSON.
|
|
1493
|
-
|
|
1494
|
-
The common failure modes: asserting exact phrasing, packing more than
|
|
1495
|
-
about five gates into one case (split it), and cases that depend on live
|
|
1496
|
-
external state that drifts (pin the input; see fixtures).
|
|
1497
|
-
|
|
1498
|
-
## Pick fixtures by agent type
|
|
1499
|
-
|
|
1500
|
-
The right fixture depends on the surface under test.
|
|
1501
|
-
|
|
1502
|
-
| Agent surface | Fixture |
|
|
1503
|
-
| --- | --- |
|
|
1504
|
-
| Chat / domain assistant | A canonical prompt string, chosen once and frozen |
|
|
1505
|
-
| Tool-heavy | Run `agent-sdk call <tool>` first to pin what the tool returns, then freeze the prompt that triggers it |
|
|
1506
|
-
| GitHub webhook | `agent-sdk github replay <pr> --events '*' --dry-run --out fixtures/github` snapshots real payloads for offline replay ([GitHub guide](/docs/guides/github.md)) |
|
|
1507
|
-
| PR reviewer with host preparation | Diff, metadata, and gold labels pinned to commit SHAs; keep any live PR matrix small |
|
|
1508
|
-
| Workspace-dependent | `workspaceFiles` in `t.send` options, never developer-machine paths |
|
|
1509
|
-
|
|
1510
|
-
Tag the fast, reliably passing core `smoke` and run `--tag smoke` in the
|
|
1511
|
-
inner loop. Leave slow or flaky-prone cases untagged for explicit runs.
|
|
1512
|
-
|
|
1513
|
-
### Materialize API-backed fixtures
|
|
1514
|
-
|
|
1515
|
-
An input that only points at external data, such as a pull request URL,
|
|
1516
|
-
snapshot id, or pair of commit SHAs, is not self-contained. Fetch it
|
|
1517
|
-
once and commit the rendered fixture before you expand the suite.
|
|
1518
|
-
|
|
1519
|
-
1. Save the diff, metadata, and labels under `fixtures/` at pinned
|
|
1520
|
-
revisions.
|
|
1521
|
-
2. Seed those files with `workspaceFiles`, or read them from the fixture
|
|
1522
|
-
directory.
|
|
1523
|
-
3. Assert decisions and output shape against the saved evidence.
|
|
1524
|
-
4. Keep a small `smoke` subset for any remaining live pipeline checks.
|
|
1525
|
-
|
|
1526
|
-
Read committed fixtures with `@cursor/july/evals/loaders`: `loadJson`,
|
|
1527
|
-
`loadJsonl`, and `loadYaml` resolve relative paths against the project
|
|
1528
|
-
root the runner discovered, not the cwd the CLI was invoked from
|
|
1529
|
-
(`resolveFixturePath` and `evalFixtureRoot` expose the same
|
|
1530
|
-
resolution for other file formats).
|
|
1531
|
-
|
|
1532
|
-
`maxConcurrency` limits parallel datapoints. It does not limit model or
|
|
1533
|
-
API fan-out inside one datapoint. Materialized fixtures prevent a large
|
|
1534
|
-
suite from exhausting provider and GitHub rate limits. The
|
|
1535
|
-
[evals skill](https://github.com/cursor/cursor/blob/main/packages/agent-serve/skills/evals/SKILL.md) has the full fixture workflow.
|
|
1536
|
-
|
|
1537
|
-
## Keep improvements with regression evals
|
|
1538
|
-
|
|
1539
|
-
Every [hillclimb](/docs/hillclimbing.md) round that keeps a change must land
|
|
1540
|
-
an eval that would have failed before the change. If you can't express
|
|
1541
|
-
the improvement as a gate (a `calledTool` shift, a bounded
|
|
1542
|
-
`action.result` count, an output-shape regex), the improvement is
|
|
1543
|
-
unverified, and it'll regress silently.
|
|
1544
|
-
|
|
1545
|
-
The rule cuts the other way too: never weaken an existing gate to make a
|
|
1546
|
-
round pass. That's the freeze line moving, and it turns your regression
|
|
1547
|
-
suite into a list of checks that no longer protect anything.
|
|
1548
|
-
|
|
1549
|
-
## Compare variants on live traffic
|
|
1550
|
-
|
|
1551
|
-
Use `defineAB` to compare variant metrics on live sessions. It is not a
|
|
1552
|
-
test runner and has no `agent-sdk ab` command. Keep `defineEval` as the
|
|
1553
|
-
regression ratchet. Eval sessions do not enroll or change live metrics.
|
|
1554
|
-
See [Live A/B metrics](/docs/ab.md) for assignment, behavior, collection,
|
|
1555
|
-
and inspection.
|
|
1556
|
-
|
|
1557
|
-
## What's next
|
|
1558
|
-
|
|
1559
|
-
Continue with these pages:
|
|
1560
|
-
|
|
1561
|
-
- [Live A/B metrics](/docs/ab.md): sticky variants and cumulative metrics
|
|
1562
|
-
on live sessions
|
|
1563
|
-
- [Hillclimbing](/docs/hillclimbing.md): the loop evals make trustworthy
|
|
1564
|
-
- [Building agents with agents](/docs/building-with-agents.md): have a
|
|
1565
|
-
coding agent write the first suite
|
|
1566
|
-
- [GitHub guide](/docs/guides/github.md): deterministic webhook fixtures
|
|
1567
|
-
with `github replay`
|
|
1568
|
-
- [Sessions and streaming](/docs/reference/sessions.md): the events
|
|
1569
|
-
`t.events` contains
|
|
1570
|
-
|
|
1571
|
-
---
|
|
1572
|
-
|
|
1573
|
-
Source: /docs/example-agents/approval-buddy.md
|
|
1574
|
-
|
|
1575
|
-
# Keep PR approval policy deterministic with Approval Buddy
|
|
1576
|
-
|
|
1577
|
-
Approval Buddy approves eligible pull requests from a fixed roster and
|
|
1578
|
-
declines every other request. GitHub still blocks self-approval when the stamp
|
|
1579
|
-
identity authored the PR. Code decides eligibility. The model prepares
|
|
1580
|
-
evidence, runs two specialist reviews, and passes their findings to the
|
|
1581
|
-
approval tool without changing the policy decision.
|
|
1582
|
-
|
|
1583
|
-
Use this example when an agent can make a judgment inside a workflow, but
|
|
1584
|
-
authorization and the final side effect must stay in deterministic code.
|
|
1585
|
-
|
|
1586
|
-
[Browse the Approval Buddy source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/approval-buddy/)
|
|
1587
|
-
|
|
1588
|
-
## Keep approval policy in code
|
|
1589
|
-
|
|
1590
|
-
Approval Buddy draws three hard boundaries:
|
|
1591
|
-
|
|
1592
|
-
- `prepare_review` and `approve_pr` re-read the live PR and apply the same
|
|
1593
|
-
eligibility rules.
|
|
1594
|
-
- Two subagents inspect prepared evidence, but their findings never grant or
|
|
1595
|
-
block approval.
|
|
1596
|
-
- Only `approve_pr` posts the GitHub review.
|
|
1597
|
-
|
|
1598
|
-
A spoofed webhook, Slack message, or model claim can't add someone to the
|
|
1599
|
-
buddy roster. The mutating tool checks the source of truth immediately before it
|
|
1600
|
-
acts.
|
|
1601
|
-
|
|
1602
|
-
## Follow the intended stamp flow
|
|
1603
|
-
|
|
1604
|
-
The root instructions ask the model to run this sequence for a qualifying PR:
|
|
1605
|
-
|
|
1606
|
-
1. A non-draft `pull_request` event arrives with action `opened`, `reopened`,
|
|
1607
|
-
or `ready_for_review`.
|
|
1608
|
-
2. The GitHub channel checks its repository allowlist and starts a session.
|
|
1609
|
-
3. `turn.started` posts a pending commit status.
|
|
1610
|
-
4. The model calls `prepare_review`.
|
|
1611
|
-
5. Host code fetches the live PR. It checks the author, open state, merged
|
|
1612
|
-
state, and draft state.
|
|
1613
|
-
6. A qualifying PR gets `pr/MANIFEST.md`, `pr/meta.json`, and
|
|
1614
|
-
`pr/diff.patch` in the session workspace. Diffs above 2,000,000
|
|
1615
|
-
characters are truncated and marked in metadata.
|
|
1616
|
-
7. The model calls both review subagents through the built-in `task` tool.
|
|
1617
|
-
8. It concatenates their contracted replies and calls `approve_pr`.
|
|
1618
|
-
9. `approve_pr` re-runs eligibility, posts an `APPROVE` review, and returns
|
|
1619
|
-
the outcome.
|
|
1620
|
-
10. The channel posts a final commit status. A self-approval block also gets
|
|
1621
|
-
a short timeline comment because no approval review can appear.
|
|
1622
|
-
|
|
1623
|
-
Ineligible PRs skip evidence and subagents. The model still calls
|
|
1624
|
-
`approve_pr` so the deterministic tool returns the formal decline reason.
|
|
1625
|
-
|
|
1626
|
-
Steps 4 through 9 are prompt-driven. The channel doesn't enforce tool order
|
|
1627
|
-
or prove both subagents ran, and `approve_pr` accepts missing findings. A
|
|
1628
|
-
failed turn clears the pending status with a green non-blocking result without
|
|
1629
|
-
approving the PR.
|
|
1630
|
-
|
|
1631
|
-
## Map the framework features
|
|
1632
|
-
|
|
1633
|
-
| Capability | Source | Role |
|
|
1634
|
-
| --- | --- | --- |
|
|
1635
|
-
| Root agent and policy prompt | [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/approval-buddy/agent/agent.ts), [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/approval-buddy/agent/instructions.md) | Configure the local agent and describe orchestration order. |
|
|
1636
|
-
| GitHub channel | [`agent/channels/github.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/approval-buddy/agent/channels/github.ts) | Filter wakes, lease GitHub access, and publish status events. |
|
|
1637
|
-
| Slack channel | [`agent/channels/slack.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/approval-buddy/agent/channels/slack.ts) | Accept approval-bot stamp and qualification requests. |
|
|
1638
|
-
| Server tools | [`agent/tools/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/approval-buddy/agent/tools/) | Prepare evidence, approve, list buddies, and search GIFs. |
|
|
1639
|
-
| Deterministic policy | [`agent/lib/approve.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/approval-buddy/agent/lib/approve.ts), [`agent/lib/buddies.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/approval-buddy/agent/lib/buddies.ts) | Own the roster and live eligibility checks. |
|
|
1640
|
-
| Review subagents | [`agent/subagents/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/approval-buddy/agent/subagents/) | Run deep audit and code-quality passes over the same evidence. |
|
|
1641
|
-
| Storage | [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/approval-buddy/agent/storage.ts) | Persist sessions and events with `cursorHostedStorage`. See [Storage](/docs/storage.md). |
|
|
1642
|
-
| Live A/B experiment | [`agent/ab.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/approval-buddy/agent/ab.ts) | Compare baseline responses with a concise, presentation-only treatment (`concise-results`). |
|
|
1643
|
-
| Evals and unit tests | [`evals/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/approval-buddy/evals/), [`agent/lib/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/approval-buddy/agent/lib/) | Protect routing, output contracts, policy, and GitHub behavior. |
|
|
1644
|
-
|
|
1645
|
-
There are no authored skills, MCP connections, schedules, reminders, hooks,
|
|
1646
|
-
sandbox seeds, or tool approvals.
|
|
1647
|
-
|
|
1648
|
-
## Prepare credentials
|
|
1649
|
-
|
|
1650
|
-
You need:
|
|
1651
|
-
|
|
1652
|
-
- Node 22.13 or newer.
|
|
1653
|
-
- An agent-runtime credential.
|
|
1654
|
-
- GitHub access to read PRs, post reviews, create commit statuses, and
|
|
1655
|
-
post the self-approval visibility comment.
|
|
1656
|
-
|
|
1657
|
-
Optional GIF selection uses:
|
|
1658
|
-
|
|
1659
|
-
- `GIPHY_API_KEY` or `APPROVAL_BUDDY_GIPHY_API_KEY`,
|
|
1660
|
-
- `APPROVAL_BUDDY_STAMP_GIF`, or
|
|
1661
|
-
- severity-specific `APPROVAL_BUDDY_STAMP_GIF_<LEVEL>` variables.
|
|
1662
|
-
|
|
1663
|
-
If you enable Giphy in a hosted copy, declare its secret and
|
|
1664
|
-
`api.giphy.com` egress.
|
|
1665
|
-
|
|
1666
|
-
## Validate without approving a PR
|
|
1667
|
-
|
|
1668
|
-
```bash
|
|
1669
|
-
agent-sdk validate --dir examples/approval-buddy
|
|
1670
|
-
agent-sdk info --dir examples/approval-buddy --json
|
|
1671
|
-
```
|
|
1672
|
-
|
|
1673
|
-
List the deterministic roster:
|
|
1674
|
-
|
|
1675
|
-
```bash
|
|
1676
|
-
agent-sdk call list_buddies \
|
|
1677
|
-
--dir examples/approval-buddy \
|
|
1678
|
-
--input '{}'
|
|
1679
|
-
```
|
|
1680
|
-
|
|
1681
|
-
Set a known merged PR, then run the read-only precheck:
|
|
1682
|
-
|
|
1683
|
-
```bash
|
|
1684
|
-
MERGED_PR_URL=https://github.com/your-org/your-repo/pull/123
|
|
1685
|
-
agent-sdk call prepare_review \
|
|
1686
|
-
--dir examples/approval-buddy \
|
|
1687
|
-
--input "{\"prUrl\":\"$MERGED_PR_URL\"}"
|
|
1688
|
-
```
|
|
1689
|
-
|
|
1690
|
-
The result should decline because the PR is no longer open. `prepare_review`
|
|
1691
|
-
never posts an approval.
|
|
1692
|
-
|
|
1693
|
-
> [!CAUTION]
|
|
1694
|
-
> Don't use `agent-sdk call approve_pr` as a smoke test. The tool has no
|
|
1695
|
-
> `needsApproval` gate and posts a real GitHub review when the PR qualifies.
|
|
1696
|
-
|
|
1697
|
-
## See why preparation is separate
|
|
1698
|
-
|
|
1699
|
-
`prepare_review` is read-only. It checks policy before fetching a large diff,
|
|
1700
|
-
so declined requests don't spend review-agent work.
|
|
1701
|
-
|
|
1702
|
-
Direct calls return the evidence file map because their scratch workspace is
|
|
1703
|
-
deleted after the call. In-session calls write the tree to
|
|
1704
|
-
`ctx.workspaceDir`, where both subagents can read it.
|
|
1705
|
-
|
|
1706
|
-
`approve_pr` repeats the live check instead of trusting preparation. A PR can
|
|
1707
|
-
close, merge, become a draft, or change author-related context between the two
|
|
1708
|
-
steps. Revalidation keeps the final write bound to current state.
|
|
1709
|
-
|
|
1710
|
-
This is a reusable two-tool pattern:
|
|
1711
|
-
|
|
1712
|
-
- a read-only tool prepares and explains the decision,
|
|
1713
|
-
- a mutating tool repeats policy at the side-effect boundary.
|
|
1714
|
-
|
|
1715
|
-
## Fan out two review contracts
|
|
1716
|
-
|
|
1717
|
-
The two discovered subagents have different contracts:
|
|
1718
|
-
|
|
1719
|
-
- The security reviewer reports bugs, breaking changes, and security findings
|
|
1720
|
-
with `High`, `Medium`, or `Low` tags.
|
|
1721
|
-
- The code-quality reviewer reports maintainability and structure concerns
|
|
1722
|
-
with `Blocker`, `Major`, or `Minor` tags.
|
|
1723
|
-
|
|
1724
|
-
The parent calls both through the harness `task` tool. They inherit the root
|
|
1725
|
-
agent's execution surface and read the same `pr/` workspace. The prompt asks
|
|
1726
|
-
the parent not to rewrite either reply. The review body trims the combined
|
|
1727
|
-
text and caps it at 16,000 characters.
|
|
1728
|
-
|
|
1729
|
-
Findings are informational. A high-severity finding doesn't veto the stamp.
|
|
1730
|
-
That policy is explicit in the root instructions and approval code.
|
|
1731
|
-
|
|
1732
|
-
## Trace GitHub channel behavior
|
|
1733
|
-
|
|
1734
|
-
The channel uses `githubChannel` with:
|
|
1735
|
-
|
|
1736
|
-
- a configured repository allowlist on the account-linked GitHub transport,
|
|
1737
|
-
- a second optional `APPROVAL_BUDDY_REPOS` wake filter,
|
|
1738
|
-
- `deliverReplies: false`,
|
|
1739
|
-
- progress reactions disabled, and
|
|
1740
|
-
- event handlers for turn start, `approve_pr` results, and failed turns.
|
|
1741
|
-
|
|
1742
|
-
The source requests `contents-write`, even though the documented workflow
|
|
1743
|
-
posts reviews, statuses, and comments. When adapting the example, start with
|
|
1744
|
-
`pr-write` and opt up only if a tool must push code.
|
|
1745
|
-
|
|
1746
|
-
Every terminal status is green by design. Declines and crashed turns are
|
|
1747
|
-
informational, not merge-blocking. This is a product decision in the example,
|
|
1748
|
-
not an Agent SDK default.
|
|
1749
|
-
|
|
1750
|
-
A successful turn that never calls `approve_pr` leaves the pending status in
|
|
1751
|
-
place. The channel clears it on `approve_pr` results and `turn.failed`, but
|
|
1752
|
-
has no `turn.completed` fallback.
|
|
1753
|
-
|
|
1754
|
-
`github replay` reaches the same channel and can post a real approval, status,
|
|
1755
|
-
or comment. Use replay only against a repository and PR created for this
|
|
1756
|
-
test.
|
|
1757
|
-
|
|
1758
|
-
## Use Slack for explicit requests
|
|
1759
|
-
|
|
1760
|
-
Start the dev server:
|
|
1761
|
-
|
|
1762
|
-
```bash
|
|
1763
|
-
agent-sdk dev examples/approval-buddy
|
|
1764
|
-
```
|
|
1765
|
-
|
|
1766
|
-
Then ask through the signed-in account-linked Slack connection:
|
|
1767
|
-
|
|
1768
|
-
> Would this PR qualify for a stamp?
|
|
1769
|
-
|
|
1770
|
-
The instructions route qualification questions to `prepare_review` only. A
|
|
1771
|
-
stamp request runs the complete flow and may approve the PR.
|
|
1772
|
-
|
|
1773
|
-
This channel uses the account-linked transport instead of a dedicated Socket
|
|
1774
|
-
Mode app.
|
|
1775
|
-
|
|
1776
|
-
## See how durable storage fits
|
|
1777
|
-
|
|
1778
|
-
`defineStorage` replaces the default local session store with a shared,
|
|
1779
|
-
durable key-value adapter. Approval Buddy chooses:
|
|
1780
|
-
|
|
1781
|
-
- a 15-second write debounce,
|
|
1782
|
-
- startup restoration for up to 200 sessions, and
|
|
1783
|
-
- a 14-day restore window.
|
|
1784
|
-
|
|
1785
|
-
That policy fits long-lived Slack threads and a small webhook fleet. The
|
|
1786
|
-
security reviewer uses the same adapter with lazy restore, which fits its
|
|
1787
|
-
shorter sessions.
|
|
1788
|
-
|
|
1789
|
-
## Run the regression suite
|
|
1790
|
-
|
|
1791
|
-
List the four eval cases:
|
|
1792
|
-
|
|
1793
|
-
```bash
|
|
1794
|
-
agent-sdk eval --dir examples/approval-buddy --list
|
|
1795
|
-
```
|
|
1796
|
-
|
|
1797
|
-
The suite covers:
|
|
1798
|
-
|
|
1799
|
-
- buddy-list routing,
|
|
1800
|
-
- declining a merged PR,
|
|
1801
|
-
- using only `prepare_review` for a qualification question, and
|
|
1802
|
-
- the combined findings headings and severity format over seeded evidence.
|
|
1803
|
-
|
|
1804
|
-
Run the safe qualification case:
|
|
1805
|
-
|
|
1806
|
-
```bash
|
|
1807
|
-
agent-sdk eval \
|
|
1808
|
-
--dir examples/approval-buddy \
|
|
1809
|
-
qualify/merged-pr-question \
|
|
1810
|
-
--json
|
|
1811
|
-
```
|
|
1812
|
-
|
|
1813
|
-
The qualification case reads a live merged PR. The seeded format case shown
|
|
1814
|
-
by `--list` uses a planted auth-bypass diff
|
|
1815
|
-
and checks for a `task` call, both headings, and severity tags. It doesn't
|
|
1816
|
-
prove both named subagents ran or whether their output reached `approve_pr`.
|
|
1817
|
-
Unit tests under `agent/lib/` cover policy, self-approval handling, evidence
|
|
1818
|
-
limits, status mapping, GIF selection, and severity parsing.
|
|
1819
|
-
|
|
1820
|
-
## Reuse the policy boundary
|
|
1821
|
-
|
|
1822
|
-
Keep these properties when you replace the buddy policy:
|
|
1823
|
-
|
|
1824
|
-
1. Put authorization in typed code.
|
|
1825
|
-
2. Fetch the source of truth inside both prepare and mutate steps.
|
|
1826
|
-
3. Give the model evidence only after the request qualifies.
|
|
1827
|
-
4. Treat specialist findings as data, not authority.
|
|
1828
|
-
5. Keep the core domain mutation in one named tool. Treat channel status and
|
|
1829
|
-
visibility writes as separate, audited effects.
|
|
1830
|
-
6. Add a human approval gate if your policy still needs operator consent.
|
|
1831
|
-
7. Test read-only routing separately from mutation.
|
|
1832
|
-
|
|
1833
|
-
## Where to go next
|
|
1834
|
-
|
|
1835
|
-
- [GitHub](/docs/guides/github.md)
|
|
1836
|
-
- [Tools](/docs/reference/tools.md)
|
|
1837
|
-
- [Subagents](/docs/reference/subagents.md)
|
|
1838
|
-
- [Storage](/docs/storage.md)
|
|
1839
|
-
- [Slack](/docs/guides/slack.md)
|
|
1840
|
-
- [Evals](/docs/evals.md)
|
|
1841
|
-
|
|
1842
|
-
---
|
|
1843
|
-
|
|
1844
|
-
Source: /docs/example-agents/benny.md
|
|
1845
|
-
|
|
1846
|
-
# Route Slack work through repository playbooks
|
|
1847
|
-
|
|
1848
|
-
This agent is a Slack teammate for a product team. Mentions and direct
|
|
1849
|
-
messages reach it through an account-linked transport. New top-level posts in
|
|
1850
|
-
an allowlisted issue channel reach it through a dedicated Slack app, even
|
|
1851
|
-
without a mention. The agent then selects a repository playbook for triage,
|
|
1852
|
-
reproduction, fixes, reviews, on-call work, or design critique.
|
|
1853
|
-
|
|
1854
|
-
Use this example when Slack is the intake surface and your durable procedures
|
|
1855
|
-
already live as repository skills.
|
|
1856
|
-
|
|
1857
|
-
[Browse the current playbook-router source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/benny/)
|
|
1858
|
-
|
|
1859
|
-
## Combine two Slack transports with repo skills
|
|
1860
|
-
|
|
1861
|
-
The playbook router uniquely combines three decisions:
|
|
1862
|
-
|
|
1863
|
-
- Two Slack transports serve different engagement modes.
|
|
1864
|
-
- `local.cwd` keeps session workspaces inside the monorepo.
|
|
1865
|
-
- Instructions route work to inherited repository playbooks
|
|
1866
|
-
instead of authored `agent/skills/`.
|
|
1867
|
-
|
|
1868
|
-
The result is a thin agent project over a mature procedure library.
|
|
1869
|
-
|
|
1870
|
-
## Follow an issue report
|
|
1871
|
-
|
|
1872
|
-
1. A teammate creates a top-level post in the allowlisted issue channel.
|
|
1873
|
-
2. The dedicated Socket Mode channel accepts the allowlisted channel.
|
|
1874
|
-
3. A 15-second debounce lets edits settle. Deleting the post during that
|
|
1875
|
-
window cancels the dispatch.
|
|
1876
|
-
4. The Agent SDK creates a thread-scoped session and sends the report to the
|
|
1877
|
-
playbook router.
|
|
1878
|
-
5. The instructions select the matching triage playbook.
|
|
1879
|
-
6. The harness finds the repository root, opens the inherited playbook, and
|
|
1880
|
-
follows its procedure.
|
|
1881
|
-
7. The agent posts only in the source thread and reports the evidence it
|
|
1882
|
-
gathered.
|
|
1883
|
-
|
|
1884
|
-
Mentions and direct messages follow the same agent instructions. They don't
|
|
1885
|
-
need the watched-channel path.
|
|
1886
|
-
|
|
1887
|
-
## Map the playbook router files
|
|
1888
|
-
|
|
1889
|
-
| File | Purpose |
|
|
1890
|
-
| --- | --- |
|
|
1891
|
-
| [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/benny/agent/agent.ts) | Names the agent, selects its model, and points the harness at a project-local cwd so inherited playbooks load. |
|
|
1892
|
-
| [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/benny/agent/instructions.md) | Defines engagement rules, evidence policy, and the playbook routing map. |
|
|
1893
|
-
| [`agent/channels/slack.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/benny/agent/channels/slack.ts) | Handles account-linked mentions and direct messages. |
|
|
1894
|
-
| [`agent/channels/slack-app.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/benny/agent/channels/slack-app.ts) | Runs the dedicated app and watches one allowlisted channel. |
|
|
1895
|
-
| [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/benny/agent/storage.ts) | Persists sessions and events with `cursorHostedStorage`. |
|
|
1896
|
-
| [`evals/evals.config.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/benny/evals/evals.config.ts) | Caps eval run concurrency. |
|
|
1897
|
-
| [`evals/smoke.eval.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/benny/evals/smoke.eval.ts) | Checks the agent identity and expected triage route. |
|
|
1898
|
-
|
|
1899
|
-
The playbook router authors no tools, MCP connections, subagents, schedules, hooks, A/B
|
|
1900
|
-
experiments, or sandbox seeds.
|
|
1901
|
-
|
|
1902
|
-
## See why `local.cwd` matters
|
|
1903
|
-
|
|
1904
|
-
The Agent SDK normally keeps an ephemeral `run` or `eval` workspace outside a
|
|
1905
|
-
large monorepo. This prevents ancestor instruction and repository-rule files
|
|
1906
|
-
from leaking into an unrelated agent.
|
|
1907
|
-
|
|
1908
|
-
The playbook router needs the opposite. Its procedures live at the repository
|
|
1909
|
-
root, so `agent.ts` points `local.cwd` at a harness directory under the
|
|
1910
|
-
project. Each harness workspace is a child of that directory. Walking up
|
|
1911
|
-
reaches the host repository and its inherited playbook directory.
|
|
1912
|
-
|
|
1913
|
-
Those playbooks are inherited context. `agent-sdk info` reports zero authored
|
|
1914
|
-
skills for the agent. Copying this project into another repository removes
|
|
1915
|
-
its main procedures unless you copy or replace the skill library too.
|
|
1916
|
-
|
|
1917
|
-
## Connect both Slack paths
|
|
1918
|
-
|
|
1919
|
-
The account-linked path needs an agent-runtime login and a connected Slack
|
|
1920
|
-
account:
|
|
1921
|
-
|
|
1922
|
-
```bash
|
|
1923
|
-
agent-sdk login
|
|
1924
|
-
agent-sdk whoami
|
|
1925
|
-
```
|
|
1926
|
-
|
|
1927
|
-
It routes explicit mentions without a dedicated Slack token on the host.
|
|
1928
|
-
|
|
1929
|
-
For the watched-channel path, configure a dedicated Socket Mode app with:
|
|
1930
|
-
|
|
1931
|
-
- subscribe to `message.channels` and `message.groups`,
|
|
1932
|
-
- have an App-Level Token with `connections:write`, and
|
|
1933
|
-
- be a member of the watched channel.
|
|
1934
|
-
|
|
1935
|
-
Run `agent-sdk slack create --dir examples/benny --channel-posts` for a
|
|
1936
|
-
dedicated Socket Mode app, then `agent-sdk slack doctor`.
|
|
1937
|
-
|
|
1938
|
-
Missing dedicated-app tokens leave that channel idle. They don't stop the
|
|
1939
|
-
account-linked channel.
|
|
1940
|
-
|
|
1941
|
-
## Validate and start the server
|
|
1942
|
-
|
|
1943
|
-
```bash
|
|
1944
|
-
agent-sdk validate --dir examples/benny
|
|
1945
|
-
agent-sdk info --dir examples/benny --json
|
|
1946
|
-
agent-sdk dev examples/benny
|
|
1947
|
-
```
|
|
1948
|
-
|
|
1949
|
-
The info output should show two Slack channels and no authored skill. That
|
|
1950
|
-
combination confirms the example is using inherited playbooks.
|
|
1951
|
-
|
|
1952
|
-
## Exercise each engagement mode
|
|
1953
|
-
|
|
1954
|
-
Test the explicit account-linked path by asking:
|
|
1955
|
-
|
|
1956
|
-
> Which playbook would you use to triage a product UI bug?
|
|
1957
|
-
|
|
1958
|
-
Test the dedicated app:
|
|
1959
|
-
|
|
1960
|
-
1. Create a top-level post in the allowlisted issue channel.
|
|
1961
|
-
2. Don't mention the bot.
|
|
1962
|
-
3. Wait for the debounce window.
|
|
1963
|
-
4. Confirm the agent replies in the post's thread.
|
|
1964
|
-
|
|
1965
|
-
Thread replies don't trigger the proactive watch. Mentions still use Slack's
|
|
1966
|
-
normal mention path. Bot-authored posts are ignored to prevent loops.
|
|
1967
|
-
|
|
1968
|
-
The channel uses the default handler after filtering. It doesn't apply a
|
|
1969
|
-
second code-level classifier, so every accepted top-level post spends a model
|
|
1970
|
-
turn and reaches the prompt.
|
|
1971
|
-
|
|
1972
|
-
## Inspect thread continuity
|
|
1973
|
-
|
|
1974
|
-
The Agent SDK keys Slack sessions by channel and thread timestamp. A follow-up in
|
|
1975
|
-
the same thread resumes the conversation and workspace. A new top-level issue
|
|
1976
|
-
gets a new session.
|
|
1977
|
-
|
|
1978
|
-
This lets a playbook gather evidence over several turns without mixing two
|
|
1979
|
-
reports. The playground shows both the account-linked and dedicated-app
|
|
1980
|
-
sessions while the dev server runs.
|
|
1981
|
-
|
|
1982
|
-
## Run the smoke eval
|
|
1983
|
-
|
|
1984
|
-
```bash
|
|
1985
|
-
agent-sdk eval --dir examples/benny --list
|
|
1986
|
-
agent-sdk eval --dir examples/benny smoke --json
|
|
1987
|
-
```
|
|
1988
|
-
|
|
1989
|
-
The case asks for the agent identity and the playbook used for issue triage.
|
|
1990
|
-
It checks the configured identity and route label.
|
|
1991
|
-
|
|
1992
|
-
This is a lexical smoke test. It doesn't prove Slack delivery, skill
|
|
1993
|
-
selection, skill loading, procedure execution, or thread-only behavior. Add
|
|
1994
|
-
fixture-backed evals around the playbooks when you reuse this design.
|
|
1995
|
-
|
|
1996
|
-
## Build a playbook-routed teammate
|
|
1997
|
-
|
|
1998
|
-
Use this structure when your organization already has tested skills:
|
|
1999
|
-
|
|
2000
|
-
1. Put the playbooks under a stable repository path.
|
|
2001
|
-
2. Set `local.cwd` so harness workspaces can inherit that path.
|
|
2002
|
-
3. Write a short routing table in `instructions.md`.
|
|
2003
|
-
4. Use account-linked Slack for explicit requests.
|
|
2004
|
-
5. Add a dedicated app only for allowlisted proactive intake.
|
|
2005
|
-
6. Keep the channel allowlist narrow and debounce edited posts.
|
|
2006
|
-
7. Add an eval for every important request-to-playbook route.
|
|
2007
|
-
|
|
2008
|
-
If the procedures should ship with the agent, put them under
|
|
2009
|
-
`agent/skills/` instead. Authored skills appear in the manifest and travel
|
|
2010
|
-
with the project.
|
|
2011
|
-
|
|
2012
|
-
## Where to go next
|
|
2013
|
-
|
|
2014
|
-
- [Slack](/docs/guides/slack.md)
|
|
2015
|
-
- [Agent config](/docs/reference/agent-config.md)
|
|
2016
|
-
- [Skills](/docs/reference/skills.md)
|
|
2017
|
-
- [Sessions and streaming](/docs/reference/sessions.md)
|
|
2018
|
-
- [Evals](/docs/evals.md)
|
|
2019
|
-
|
|
2020
|
-
---
|
|
2021
|
-
|
|
2022
|
-
Source: /docs/example-agents/bugbot.md
|
|
2023
|
-
|
|
2024
|
-
# Review prepared pull-request evidence
|
|
2025
|
-
|
|
2026
|
-
This GitHub-read-only reviewer uses host code to fetch the PR
|
|
2027
|
-
with `gh` and `git`, builds a trimmed `pr/` evidence tree, then hands that tree
|
|
2028
|
-
to the model. The model reads the diff, loads a review skill, and returns at
|
|
2029
|
-
most three high-confidence findings.
|
|
2030
|
-
|
|
2031
|
-
Use this example when the host should control evidence collection and the
|
|
2032
|
-
model shouldn't browse or mutate the source repository.
|
|
2033
|
-
|
|
2034
|
-
[Browse the current reviewer source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/bugbot/)
|
|
2035
|
-
|
|
2036
|
-
## Separate evidence preparation from review
|
|
2037
|
-
|
|
2038
|
-
The reviewer separates preparation from judgment:
|
|
2039
|
-
|
|
2040
|
-
- Host code owns GitHub and Git access.
|
|
2041
|
-
- A server tool turns untrusted PR input into bounded workspace files.
|
|
2042
|
-
- A custom channel seeds those files before the model starts.
|
|
2043
|
-
- An on-demand skill defines the review procedure and output contract.
|
|
2044
|
-
- The model returns chat text. No path posts a GitHub review.
|
|
2045
|
-
|
|
2046
|
-
This architecture gives the model a purpose-built evidence package instead of
|
|
2047
|
-
a checkout.
|
|
2048
|
-
|
|
2049
|
-
## Follow a review
|
|
2050
|
-
|
|
2051
|
-
The custom HTTP path runs this sequence:
|
|
2052
|
-
|
|
2053
|
-
1. `POST /v1/channels/review/` receives a PR reference.
|
|
2054
|
-
2. The handler calls `prepare_pr` without a model turn.
|
|
2055
|
-
3. Host code reads PR metadata and the unified diff.
|
|
2056
|
-
4. It reuses a matching checkout, force-fetching the PR ref there when the
|
|
2057
|
-
commit is missing. Without a matching checkout, it uses a temporary bare
|
|
2058
|
-
cache.
|
|
2059
|
-
5. It creates `pr/MANIFEST.md`, `pr/meta.json`, `pr/diff.patch`, and selected
|
|
2060
|
-
small files and rules.
|
|
2061
|
-
6. `send({ workspaceFiles })` creates the model session with that evidence.
|
|
2062
|
-
7. The model reads the manifest and diff, then loads `pr-review`.
|
|
2063
|
-
8. The channel returns session and playground URLs while the review streams.
|
|
2064
|
-
|
|
2065
|
-
If a normal chat starts without evidence, the model can call `prepare_pr`
|
|
2066
|
-
mid-turn. That form writes the same files into the active session workspace.
|
|
2067
|
-
|
|
2068
|
-
## Map the evidence-review files
|
|
2069
|
-
|
|
2070
|
-
| File | Purpose |
|
|
2071
|
-
| --- | --- |
|
|
2072
|
-
| [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/agent/agent.ts) | Selects the local runtime and model. |
|
|
2073
|
-
| [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/agent/instructions.md) | Requires diff-first review and confines model work to `pr/`. |
|
|
2074
|
-
| [`agent/tools/prepare_pr.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/agent/tools/prepare_pr.ts) | Exposes host preparation as a typed server tool. |
|
|
2075
|
-
| [`agent/lib/prepare-pr.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/agent/lib/prepare-pr.ts) | Parses PR references, runs `gh` and `git`, and builds the evidence map. |
|
|
2076
|
-
| [`agent/channels/review.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/agent/channels/review.ts) | Provides the loopback-only prepare-and-send HTTP route. |
|
|
2077
|
-
| [`agent/channels/slack.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/agent/channels/slack.ts) | Extracts PR references and prepares evidence for mentions and direct messages. |
|
|
2078
|
-
| [`agent/skills/pr-review.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/agent/skills/pr-review.md) | Sets finding limits, severities, and the machine-readable review format. |
|
|
2079
|
-
| [`agent/lib/log.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/agent/lib/log.ts) | Writes timing logs for the host tools to stderr. |
|
|
2080
|
-
| [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/agent/storage.ts) | Persists sessions and events with `cursorHostedStorage`. |
|
|
2081
|
-
| [`evals/evals.config.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/evals/evals.config.ts) | Caps eval run concurrency. |
|
|
2082
|
-
| [`evals/review/smoke.eval.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/evals/review/smoke.eval.ts) | Seeds fake evidence and checks the review path without GitHub. |
|
|
2083
|
-
|
|
2084
|
-
There is no authored GitHub channel, MCP connection, subagent, schedule,
|
|
2085
|
-
hook, A/B experiment, approval, or custom storage.
|
|
2086
|
-
|
|
2087
|
-
## Prepare the host
|
|
2088
|
-
|
|
2089
|
-
You need:
|
|
2090
|
-
|
|
2091
|
-
- Node 22.13 or newer.
|
|
2092
|
-
- An agent-runtime credential for model turns and account-linked Slack.
|
|
2093
|
-
- `gh` and `git` on `PATH`.
|
|
2094
|
-
- `gh` access to the target PR.
|
|
2095
|
-
- Network access to GitHub and a writable temporary directory.
|
|
2096
|
-
|
|
2097
|
-
The preparer can prefer a configured local checkout. Its `origin` must match
|
|
2098
|
-
the target repository. Otherwise the reviewer uses its bare cache. It never
|
|
2099
|
-
checks out the PR into the serve host's working tree.
|
|
2100
|
-
|
|
2101
|
-
## Validate the surface
|
|
2102
|
-
|
|
2103
|
-
```bash
|
|
2104
|
-
agent-sdk validate --dir examples/bugbot
|
|
2105
|
-
agent-sdk info --dir examples/bugbot --json
|
|
2106
|
-
```
|
|
2107
|
-
|
|
2108
|
-
The manifest should show one server tool, one skill, and two authored
|
|
2109
|
-
channels.
|
|
2110
|
-
|
|
2111
|
-
## Inspect evidence without a model turn
|
|
2112
|
-
|
|
2113
|
-
Call the preparation tool directly:
|
|
2114
|
-
|
|
2115
|
-
```bash
|
|
2116
|
-
agent-sdk call prepare_pr \
|
|
2117
|
-
--dir examples/bugbot \
|
|
2118
|
-
--input '{"pr":"https://github.com/owner/repo/pull/123"}'
|
|
2119
|
-
```
|
|
2120
|
-
|
|
2121
|
-
Direct tool calls use a scratch workspace removed after the call.
|
|
2122
|
-
`prepare_pr` detects this path and returns the complete file map in its
|
|
2123
|
-
result. In a model session, it writes the files and returns a smaller summary.
|
|
2124
|
-
|
|
2125
|
-
The evidence builder applies explicit limits:
|
|
2126
|
-
|
|
2127
|
-
| Evidence | Limit |
|
|
2128
|
-
| --- | --- |
|
|
2129
|
-
| Post-change file | 12,000 characters |
|
|
2130
|
-
| One rule file | 8,000 characters |
|
|
2131
|
-
| Combined rules | 12,000 characters |
|
|
2132
|
-
| PR body in metadata | 2,000 characters |
|
|
2133
|
-
|
|
2134
|
-
Large files remain visible in `diff.patch`. The manifest records which full
|
|
2135
|
-
files or rules were omitted.
|
|
2136
|
-
|
|
2137
|
-
The per-file limits aren't an aggregate context cap. Every changed file below
|
|
2138
|
-
12,000 characters can be included. The diff command has a 12 MiB output
|
|
2139
|
-
buffer; a larger diff fails preparation instead of being truncated.
|
|
2140
|
-
|
|
2141
|
-
## Run the HTTP review path
|
|
2142
|
-
|
|
2143
|
-
Start the server:
|
|
2144
|
-
|
|
2145
|
-
```bash
|
|
2146
|
-
agent-sdk dev examples/bugbot
|
|
2147
|
-
```
|
|
2148
|
-
|
|
2149
|
-
From another terminal:
|
|
2150
|
-
|
|
2151
|
-
```bash
|
|
2152
|
-
curl -s -X POST \
|
|
2153
|
-
http://127.0.0.1:3000/bugbot/v1/channels/review/ \
|
|
2154
|
-
-H 'content-type: application/json' \
|
|
2155
|
-
-d '{"pr":"https://github.com/owner/repo/pull/123"}'
|
|
2156
|
-
```
|
|
2157
|
-
|
|
2158
|
-
The route returns `status: "started"`, a continuation token, and session and
|
|
2159
|
-
playground URLs. Open the session URL to watch the model read the evidence and
|
|
2160
|
-
produce findings.
|
|
2161
|
-
|
|
2162
|
-
The channel declares `localDevStrict()`. Direct loopback callers can use it.
|
|
2163
|
-
Proxy-forwarding headers and non-loopback hosts are rejected.
|
|
2164
|
-
|
|
2165
|
-
Send a follow-up by passing the returned key:
|
|
2166
|
-
|
|
2167
|
-
```bash
|
|
2168
|
-
curl -s -X POST \
|
|
2169
|
-
http://127.0.0.1:3000/bugbot/v1/channels/review/ \
|
|
2170
|
-
-H 'content-type: application/json' \
|
|
2171
|
-
-d '{"pr":"owner/repo#123","key":"<continuation-token>"}'
|
|
2172
|
-
```
|
|
2173
|
-
|
|
2174
|
-
The follow-up resumes the session without fetching a new evidence tree.
|
|
2175
|
-
|
|
2176
|
-
## Run the Slack path
|
|
2177
|
-
|
|
2178
|
-
The account-linked Slack channel handles review-bot mentions and direct
|
|
2179
|
-
messages:
|
|
2180
|
-
|
|
2181
|
-
> Review https://github.com/owner/repo/pull/123
|
|
2182
|
-
|
|
2183
|
-
Slack handlers don't receive the channel `callTool` helper. This example calls
|
|
2184
|
-
the shared `preparePrReview` host function, then returns `workspaceFiles` in
|
|
2185
|
-
the Slack message preparation result. The model sees the same evidence and
|
|
2186
|
-
prompt as the HTTP path.
|
|
2187
|
-
|
|
2188
|
-
If a message contains no PR reference, the handler asks for one. Thread
|
|
2189
|
-
follow-ups keep the same session.
|
|
2190
|
-
|
|
2191
|
-
## See how the skill constrains review
|
|
2192
|
-
|
|
2193
|
-
`pr-review.md` tells the model to:
|
|
2194
|
-
|
|
2195
|
-
- read the manifest and unified diff first,
|
|
2196
|
-
- open at most one supporting file or rules file when a hunk is ambiguous,
|
|
2197
|
-
- avoid shell, network, `gh`, and `git`,
|
|
2198
|
-
- report no more than three findings,
|
|
2199
|
-
- keep each description under 120 words, and
|
|
2200
|
-
- emit the machine-readable review contract.
|
|
2201
|
-
|
|
2202
|
-
The root instructions set the evidence boundary. The skill holds the reusable
|
|
2203
|
-
review procedure. Keeping those roles separate lets another agent reuse the
|
|
2204
|
-
same skill with different intake channels.
|
|
2205
|
-
|
|
2206
|
-
## Run the fixture-backed eval
|
|
2207
|
-
|
|
2208
|
-
```bash
|
|
2209
|
-
agent-sdk eval --dir examples/bugbot --list
|
|
2210
|
-
agent-sdk eval --dir examples/bugbot review/smoke --json
|
|
2211
|
-
```
|
|
2212
|
-
|
|
2213
|
-
The eval constructs a `PreparedPrReview`, seeds its file map through
|
|
2214
|
-
`workspaceFiles`, and checks for at least one read call with no shell call. It
|
|
2215
|
-
doesn't assert which evidence file was read or whether the skill loaded. It
|
|
2216
|
-
accepts either a formatted review or a clean result.
|
|
2217
|
-
|
|
2218
|
-
This case tests review behavior without GitHub credentials or network data.
|
|
2219
|
-
Add fixtures with reachable bugs when you need stricter location and severity
|
|
2220
|
-
checks.
|
|
2221
|
-
|
|
2222
|
-
## Keep the side-effect boundary clear
|
|
2223
|
-
|
|
2224
|
-
The reviewer makes no remote GitHub writes. It doesn't author a GitHub channel and
|
|
2225
|
-
doesn't call a review API. Host preparation does write session evidence and
|
|
2226
|
-
force-update `refs/pull/<N>/head` in either its bare cache or a matching local
|
|
2227
|
-
checkout when the commit is missing. Every result ends with a note saying no
|
|
2228
|
-
GitHub review was posted.
|
|
2229
|
-
|
|
2230
|
-
If you add publishing later, keep it in a separate tool. This preserves a
|
|
2231
|
-
read-only preparation and review path safe to run in evals.
|
|
2232
|
-
|
|
2233
|
-
## Reuse the evidence handoff
|
|
2234
|
-
|
|
2235
|
-
Use host-prepared workspaces when:
|
|
2236
|
-
|
|
2237
|
-
- external APIs should stay off the model's tool surface,
|
|
2238
|
-
- context needs hard size limits,
|
|
2239
|
-
- the model should inspect a snapshot instead of a live checkout, or
|
|
2240
|
-
- several channels need the same preparation.
|
|
2241
|
-
|
|
2242
|
-
Return `workspaceFiles` from direct host preparation, write into
|
|
2243
|
-
`ctx.workspaceDir` for mid-turn recovery, and encode the reading order in both
|
|
2244
|
-
the manifest and a skill.
|
|
2245
|
-
|
|
2246
|
-
## Where to go next
|
|
2247
|
-
|
|
2248
|
-
- [Webhooks and custom channels](/docs/guides/webhooks.md)
|
|
2249
|
-
- [Tools](/docs/reference/tools.md)
|
|
2250
|
-
- [Skills](/docs/reference/skills.md)
|
|
2251
|
-
- [Slack](/docs/guides/slack.md)
|
|
2252
|
-
- [Evals](/docs/evals.md)
|
|
2253
|
-
|
|
2254
|
-
---
|
|
2255
|
-
|
|
2256
|
-
Source: /docs/example-agents/codebase-wiki.md
|
|
2257
|
-
|
|
2258
|
-
# Build a feature wiki from merged pull requests
|
|
2259
|
-
|
|
2260
|
-
Codebase wiki keeps a living, feature-organized wiki of a repository.
|
|
2261
|
-
The GitHub channel acknowledges every closed pull request instantly,
|
|
2262
|
-
fetches a compact digest on the host, and spends a model turn only on
|
|
2263
|
-
merged PRs. The turn maps the change onto feature pages; a daily
|
|
2264
|
-
schedule writes a digest of what changed and rebuilds the index. Chat
|
|
2265
|
-
sessions answer codebase questions from the wiki with page citations.
|
|
2266
|
-
|
|
2267
|
-
Use this project when documentation should accumulate from merges
|
|
2268
|
-
instead of being regenerated from scratch. Use
|
|
2269
|
-
[Knowledge base](/docs/example-agents/knowledge-base.md) when people should curate
|
|
2270
|
-
organizational context through conversation.
|
|
2271
|
-
|
|
2272
|
-
[Browse the codebase wiki source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/codebase-wiki/)
|
|
2273
|
-
|
|
2274
|
-
## Treat PRs as evidence and features as pages
|
|
2275
|
-
|
|
2276
|
-
The wiki refuses to become a merge log:
|
|
2277
|
-
|
|
2278
|
-
- The page tree is rigid: `index`, `features/<slug>`, and
|
|
2279
|
-
`digests/<yyyy-mm-dd>`. The store rejects anything else, so the wiki
|
|
2280
|
-
can't sprawl.
|
|
2281
|
-
- The `feature-mapping` skill requires a `wiki_search` before every
|
|
2282
|
-
write. A PR updates the page that owns its feature; a new page needs
|
|
2283
|
-
a genuinely new feature; chores change nothing.
|
|
2284
|
-
- Every touched page gets a dated changelog entry citing the PR
|
|
2285
|
-
number, so each fact traces back to a merge.
|
|
2286
|
-
|
|
2287
|
-
The wiki itself is markdown on the serve host, in a wiki directory by
|
|
2288
|
-
default with a `CODEBASE_WIKI_DIR` override. Sessions are disposable;
|
|
2289
|
-
the wiki is the durable state.
|
|
2290
|
-
|
|
2291
|
-
## Follow a merged PR
|
|
2292
|
-
|
|
2293
|
-
1. GitHub delivers `pull_request` with action `closed`. The channel
|
|
2294
|
-
returns a task acknowledgement immediately.
|
|
2295
|
-
2. The task fetches the digest with the host `gh` CLI: title, body,
|
|
2296
|
-
labels, changed files, and a bounded diff excerpt. No checkout.
|
|
2297
|
-
3. The webhook payload can't say whether the PR merged, so the host
|
|
2298
|
-
checks `mergedAt` and skips abandoned PRs without a model turn.
|
|
2299
|
-
4. For merged PRs, the task starts the turn with `pr/DIGEST.md` seeded
|
|
2300
|
-
through `workspaceFiles` and a `pr:<owner/repo#N>` continuation
|
|
2301
|
-
token, so redeliveries resume instead of double-ingesting.
|
|
2302
|
-
5. The model follows `feature-mapping`: search, update or create
|
|
2303
|
-
feature pages, add changelog entries, and refresh `index` when pages
|
|
2304
|
-
were added.
|
|
2305
|
-
|
|
2306
|
-
In chat, "ingest PR #123" runs the same flow through the `ingest_pr`
|
|
2307
|
-
tool, which writes the digest into the active session workspace.
|
|
2308
|
-
|
|
2309
|
-
## Map the wiki files
|
|
2310
|
-
|
|
2311
|
-
| File | Purpose |
|
|
2312
|
-
| --- | --- |
|
|
2313
|
-
| [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/agent.ts) | Selects the local runtime and model. |
|
|
2314
|
-
| [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/instructions.md) | Splits the job into merge ingestion and wiki-cited Q&A. |
|
|
2315
|
-
| [`agent/lib/wiki-store.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/lib/wiki-store.ts) | Enforces the rigid page tree and owns reads, writes, and search. |
|
|
2316
|
-
| [`agent/lib/pr-digest.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/lib/pr-digest.ts) | Fetches PR metadata and diff, and formats `pr/DIGEST.md`. |
|
|
2317
|
-
| [`agent/tools/ingest_pr.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/tools/ingest_pr.ts) | Exposes host digest preparation for chat-driven backfills. |
|
|
2318
|
-
| [`agent/tools/wiki_read.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/tools/wiki_read.ts), [`wiki_search.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/tools/wiki_search.ts), [`wiki_write.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/tools/wiki_write.ts) | Read, search, and rewrite wiki pages. |
|
|
2319
|
-
| [`agent/skills/feature-mapping.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/skills/feature-mapping.md) | Maps changes onto features and fixes the page and changelog shape. |
|
|
2320
|
-
| [`agent/schedules/daily-digest.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/schedules/daily-digest.md) | Writes `digests/<date>`, rebuilds the index, and flags stale pages. |
|
|
2321
|
-
| [`agent/channels/github.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/channels/github.ts) | Acknowledges closed PRs and starts merged-only ingest turns. |
|
|
2322
|
-
| [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/storage.ts) | Persists sessions and events with `cursorHostedStorage`. |
|
|
2323
|
-
| [`evals/evals.config.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/evals/evals.config.ts) | Caps eval run concurrency. |
|
|
2324
|
-
| [`evals/ingest.eval.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/evals/ingest.eval.ts) | Gates ingest decisions against the wiki filesystem. |
|
|
2325
|
-
|
|
2326
|
-
There is no MCP connection, subagent, hook, or A/B experiment.
|
|
2327
|
-
|
|
2328
|
-
## Prepare credentials and services
|
|
2329
|
-
|
|
2330
|
-
You need:
|
|
2331
|
-
|
|
2332
|
-
- Node 22.13 or newer.
|
|
2333
|
-
- An agent-runtime credential for model turns.
|
|
2334
|
-
- `gh` on `PATH` with read access to the PRs you ingest.
|
|
2335
|
-
|
|
2336
|
-
The channel verifies webhook signatures when `GITHUB_WEBHOOK_SECRET` is
|
|
2337
|
-
set and narrows repositories with
|
|
2338
|
-
`CODEBASE_WIKI_REPOS=owner/repo,owner/other`. The agent never writes to
|
|
2339
|
-
GitHub. Its only side effects are wiki files on the serve host.
|
|
2340
|
-
|
|
2341
|
-
## Validate the surface
|
|
2342
|
-
|
|
2343
|
-
```bash
|
|
2344
|
-
agent-sdk validate --dir examples/codebase-wiki
|
|
2345
|
-
agent-sdk info --dir examples/codebase-wiki --json
|
|
2346
|
-
```
|
|
2347
|
-
|
|
2348
|
-
The manifest should report four server tools, one skill, one schedule,
|
|
2349
|
-
and the authored GitHub channel.
|
|
2350
|
-
|
|
2351
|
-
## Ingest without webhook plumbing
|
|
2352
|
-
|
|
2353
|
-
Replay a real merged PR as a closed delivery:
|
|
2354
|
-
|
|
2355
|
-
```bash
|
|
2356
|
-
agent-sdk dev examples/codebase-wiki
|
|
2357
|
-
|
|
2358
|
-
agent-sdk github replay https://github.com/owner/repo/pull/123 \
|
|
2359
|
-
--dir examples/codebase-wiki --action closed
|
|
2360
|
-
```
|
|
2361
|
-
|
|
2362
|
-
The reply is a 202 acknowledgement; the ingest continues in the task.
|
|
2363
|
-
Watch the session in the playground, then open the wiki directory on
|
|
2364
|
-
the serve host. Feature pages land under `features/`.
|
|
2365
|
-
|
|
2366
|
-
Each ingested feature page carries an overview, a "How it works"
|
|
2367
|
-
section, and a changelog line citing the PR. Deterministic digest
|
|
2368
|
-
preparation works without a model turn:
|
|
2369
|
-
|
|
2370
|
-
```bash
|
|
2371
|
-
agent-sdk call ingest_pr \
|
|
2372
|
-
--dir examples/codebase-wiki \
|
|
2373
|
-
--input '{"pr":"https://github.com/owner/repo/pull/123"}'
|
|
2374
|
-
```
|
|
2375
|
-
|
|
2376
|
-
A PR closed without merging returns `merged: false` and a note telling
|
|
2377
|
-
the model to change nothing.
|
|
2378
|
-
|
|
2379
|
-
## Run the daily digest
|
|
2380
|
-
|
|
2381
|
-
The schedule fires at 07:00 UTC. Under `agent-sdk dev`, trigger it by
|
|
2382
|
-
hand:
|
|
2383
|
-
|
|
2384
|
-
```bash
|
|
2385
|
-
curl -s -X POST http://127.0.0.1:3000/codebase-wiki/v1/dev/schedules/daily-digest
|
|
2386
|
-
```
|
|
2387
|
-
|
|
2388
|
-
The turn reads every feature changelog, writes
|
|
2389
|
-
`digests/<today>` grouped by feature with PR citations, rebuilds
|
|
2390
|
-
`index`, and reports one line per page it wrote. Entries dated today
|
|
2391
|
-
always count; a digest only claims a quiet day when no entry qualifies.
|
|
2392
|
-
|
|
2393
|
-
## Run the evals
|
|
2394
|
-
|
|
2395
|
-
```bash
|
|
2396
|
-
agent-sdk eval --dir examples/codebase-wiki --list
|
|
2397
|
-
agent-sdk eval --dir examples/codebase-wiki ingest/update-existing
|
|
2398
|
-
```
|
|
2399
|
-
|
|
2400
|
-
The cases seed a temp wiki through `CODEBASE_WIKI_DIR` and build
|
|
2401
|
-
digests with the same formatter the channel uses, so they run without
|
|
2402
|
-
GitHub or network access. The gates check the filesystem, not prose:
|
|
2403
|
-
a new feature page lands on a new slug, a related PR updates the
|
|
2404
|
-
existing page instead of duplicating it, an unmerged PR changes
|
|
2405
|
-
nothing, and the daily pass writes a digest naming both seeded
|
|
2406
|
-
features.
|
|
2407
|
-
|
|
2408
|
-
## Reuse the merge-ingestion pattern
|
|
2409
|
-
|
|
2410
|
-
Copy this shape when events should accumulate into curated state:
|
|
2411
|
-
|
|
2412
|
-
- Acknowledge webhooks with a task and decide host-side whether a
|
|
2413
|
-
model turn is worth spending.
|
|
2414
|
-
- Seed evidence through `workspaceFiles` so the model never fetches.
|
|
2415
|
-
- Constrain the durable store's shape in code and its content in a
|
|
2416
|
-
skill.
|
|
2417
|
-
- Add a consolidation schedule so incremental writes stay coherent.
|
|
2418
|
-
|
|
2419
|
-
## Where to go next
|
|
2420
|
-
|
|
2421
|
-
- [GitHub webhooks](/docs/guides/github.md)
|
|
2422
|
-
- [Schedules](/docs/reference/schedules.md)
|
|
2423
|
-
- [Tools](/docs/reference/tools.md)
|
|
2424
|
-
- [Evals](/docs/evals.md)
|
|
2425
|
-
|
|
2426
|
-
---
|
|
2427
|
-
|
|
2428
|
-
Source: /docs/example-agents/codeowners-review.md
|
|
2429
|
-
|
|
2430
|
-
# Route PR reviews by code ownership
|
|
2431
|
-
|
|
2432
|
-
Codeowners review gives each part of a codebase its own review. A
|
|
2433
|
-
CODEOWNERS-style table maps changed paths to review areas; each area
|
|
2434
|
-
has a markdown playbook with the team's rules for that domain; and one
|
|
2435
|
-
`area-reviewer` subagent runs per routed area, in parallel. A billing
|
|
2436
|
-
change gets the billing review, a migration gets the migration review,
|
|
2437
|
-
and an author's personal style rides along as advisory notes. The lead
|
|
2438
|
-
aggregates: approve only when every area approves.
|
|
2439
|
-
|
|
2440
|
-
Use this project when review quality depends on domain-specific values
|
|
2441
|
-
instead of one generic checklist.
|
|
2442
|
-
|
|
2443
|
-
[Browse the codeowners review source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/codeowners-review/)
|
|
2444
|
-
|
|
2445
|
-
## Keep routing in code and judgment in playbooks
|
|
2446
|
-
|
|
2447
|
-
The pipeline separates three concerns:
|
|
2448
|
-
|
|
2449
|
-
- `reviews/REVIEWERS` routes. Host code matches every changed path
|
|
2450
|
-
against the table; every matching rule applies, and unmatched paths
|
|
2451
|
-
fall back to the `general` playbook. Routing is glob code with unit
|
|
2452
|
-
tests, not model judgment.
|
|
2453
|
-
- `reviews/<area>.md` judges. Each playbook is a severity-ordered rule
|
|
2454
|
-
list the team owns: billing mandates integer cents and idempotent
|
|
2455
|
-
webhooks, migrations forbid destructive DDL beside code changes,
|
|
2456
|
-
background jobs demand idempotency and dead-letter paths.
|
|
2457
|
-
- Subagents review. The lead reads nothing but the manifest and
|
|
2458
|
-
routes; each `area-reviewer` reads one playbook plus its files' diff
|
|
2459
|
-
hunks and returns a mechanical verdict: request changes on any High
|
|
2460
|
-
finding or two Mediums.
|
|
2461
|
-
|
|
2462
|
-
Personal styles extend the same mechanism. `reviews/people/<login>.md`
|
|
2463
|
-
attaches automatically, as advisory notes, whenever that person authors
|
|
2464
|
-
the PR. Adding an area or a style is a markdown file plus at most one
|
|
2465
|
-
routing line.
|
|
2466
|
-
|
|
2467
|
-
## Follow a review
|
|
2468
|
-
|
|
2469
|
-
1. A PR arrives: a GitHub `pull_request` event, a chat message, or a
|
|
2470
|
-
bundled fixture reference.
|
|
2471
|
-
2. `prepare_review` fetches metadata and the diff with the host `gh`
|
|
2472
|
-
CLI, routes every changed file, and writes the `pr/` evidence tree:
|
|
2473
|
-
`MANIFEST.md`, `ROUTES.md`, `diff.patch`, and a copy of each matched
|
|
2474
|
-
playbook.
|
|
2475
|
-
3. The lead follows the `review-process` skill and issues one
|
|
2476
|
-
`area-reviewer` delegation per routed area, plus one per personal
|
|
2477
|
-
style, all in one step so they run in parallel.
|
|
2478
|
-
4. Each reviewer reads its playbook, reviews only its files, and
|
|
2479
|
-
returns a verdict line with at most three findings.
|
|
2480
|
-
5. The lead aggregates per-area sections and the overall verdict:
|
|
2481
|
-
APPROVE only when every non-advisory area approved.
|
|
2482
|
-
|
|
2483
|
-
Nothing posts to GitHub. Verdicts live in the session; the
|
|
2484
|
-
[Approval Buddy guide](/docs/example-agents/approval-buddy.md) shows how to wire a real
|
|
2485
|
-
APPROVE and commit statuses on top of the same shape.
|
|
2486
|
-
|
|
2487
|
-
## Map the review files
|
|
2488
|
-
|
|
2489
|
-
| File | Purpose |
|
|
2490
|
-
| --- | --- |
|
|
2491
|
-
| [`reviews/REVIEWERS`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/reviews/REVIEWERS) | Routes path patterns to review areas. |
|
|
2492
|
-
| [`reviews/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/codeowners-review/reviews/) | Holds the area playbooks and `people/<login>.md` styles. |
|
|
2493
|
-
| [`agent/lib/routing.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/agent/lib/routing.ts) | Parses the table, matches globs, and unions areas per file. |
|
|
2494
|
-
| [`agent/lib/prepare-review.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/agent/lib/prepare-review.ts) | Fetches PRs or fixtures and builds the evidence tree. |
|
|
2495
|
-
| [`agent/tools/prepare_review.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/agent/tools/prepare_review.ts) | Exposes host preparation as a typed server tool. |
|
|
2496
|
-
| [`agent/tools/list_review_areas.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/agent/tools/list_review_areas.ts) | Answers routing questions deterministically. |
|
|
2497
|
-
| [`agent/skills/review-process.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/agent/skills/review-process.md) | Fixes the fan-out procedure and the verdict rule. |
|
|
2498
|
-
| [`agent/subagents/area-reviewer/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/codeowners-review/agent/subagents/area-reviewer/) | Defines the one-area, one-playbook reviewer contract. |
|
|
2499
|
-
| [`agent/channels/github.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/agent/channels/github.ts) | Reviews opened, reopened, synchronized, and undrafted PRs. |
|
|
2500
|
-
| [`fixtures/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/codeowners-review/fixtures/) | Ships two reviewable PRs with known planted findings. |
|
|
2501
|
-
| [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/agent/storage.ts) | Persists sessions and events with `cursorHostedStorage`. |
|
|
2502
|
-
| [`evals/evals.config.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/evals/evals.config.ts) | Caps eval run concurrency. |
|
|
2503
|
-
| [`evals/review.eval.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/evals/review.eval.ts) | Gates routing, fan-out, planted bugs, and verdicts. |
|
|
2504
|
-
|
|
2505
|
-
There is no MCP connection, schedule, hook, A/B experiment, or custom
|
|
2506
|
-
storage.
|
|
2507
|
-
|
|
2508
|
-
## Prepare credentials and services
|
|
2509
|
-
|
|
2510
|
-
You need:
|
|
2511
|
-
|
|
2512
|
-
- Node 22.13 or newer.
|
|
2513
|
-
- An agent-runtime credential for model turns.
|
|
2514
|
-
- `gh` on `PATH` with read access to real PRs you review. The bundled
|
|
2515
|
-
fixtures need no network at all.
|
|
2516
|
-
|
|
2517
|
-
The channel verifies webhook signatures when `GITHUB_WEBHOOK_SECRET` is
|
|
2518
|
-
set and narrows repositories with
|
|
2519
|
-
`CODEOWNERS_REVIEW_REPOS=owner/repo,owner/other`. Pushes re-review in
|
|
2520
|
-
the same session through the `pr:<label>` continuation token.
|
|
2521
|
-
|
|
2522
|
-
## Validate the surface
|
|
2523
|
-
|
|
2524
|
-
```bash
|
|
2525
|
-
agent-sdk validate --dir examples/codeowners-review
|
|
2526
|
-
agent-sdk info --dir examples/codeowners-review --json
|
|
2527
|
-
```
|
|
2528
|
-
|
|
2529
|
-
The manifest should report two server tools, one skill, one subagent,
|
|
2530
|
-
and the authored GitHub channel.
|
|
2531
|
-
|
|
2532
|
-
## Inspect routing without a model turn
|
|
2533
|
-
|
|
2534
|
-
```bash
|
|
2535
|
-
agent-sdk call list_review_areas --dir examples/codeowners-review --input '{}'
|
|
2536
|
-
|
|
2537
|
-
agent-sdk call prepare_review \
|
|
2538
|
-
--dir examples/codeowners-review \
|
|
2539
|
-
--input '{"pr":"fixture:multi-area"}'
|
|
2540
|
-
```
|
|
2541
|
-
|
|
2542
|
-
The fixture routes to `billing`, `database-migrations`, and `frontend`,
|
|
2543
|
-
attaches `people/alice` because alice authored it, and returns the full
|
|
2544
|
-
evidence map. Point the same tool at a real PR URL and the routing runs
|
|
2545
|
-
against the live file list. The example table maps a hypothetical
|
|
2546
|
-
`src/` layout, so most real repositories route to `general` until you
|
|
2547
|
-
adapt `reviews/REVIEWERS`.
|
|
2548
|
-
|
|
2549
|
-
## Review the planted fixture
|
|
2550
|
-
|
|
2551
|
-
```bash
|
|
2552
|
-
agent-sdk dev examples/codeowners-review
|
|
2553
|
-
```
|
|
2554
|
-
|
|
2555
|
-
In the playground:
|
|
2556
|
-
|
|
2557
|
-
> Review fixture:multi-area
|
|
2558
|
-
|
|
2559
|
-
The fixture plants one violation per area: float dollar math in
|
|
2560
|
-
`src/billing/invoice.ts`, a `DROP COLUMN` plus a non-concurrent index
|
|
2561
|
-
in the migration, and a clickable `div` without loading states in the
|
|
2562
|
-
UI. The trace shows `prepare_review`, the evidence reads, four parallel
|
|
2563
|
-
`area-reviewer` cards, and an aggregated CHANGES REQUESTED verdict with
|
|
2564
|
-
each planted bug filed under its own area. The second fixture,
|
|
2565
|
-
`fixture:jobs-clean`, routes to `background-jobs` alone and ends in
|
|
2566
|
-
APPROVE.
|
|
2567
|
-
|
|
2568
|
-
Review a real PR the same way:
|
|
2569
|
-
|
|
2570
|
-
> Review https://github.com/owner/repo/pull/123
|
|
2571
|
-
|
|
2572
|
-
Or replay one as a webhook delivery:
|
|
2573
|
-
|
|
2574
|
-
```bash
|
|
2575
|
-
agent-sdk github replay https://github.com/owner/repo/pull/123 \
|
|
2576
|
-
--dir examples/codeowners-review --action opened
|
|
2577
|
-
```
|
|
2578
|
-
|
|
2579
|
-
## See how the verdict stays mechanical
|
|
2580
|
-
|
|
2581
|
-
The reviewer contract computes verdicts from findings instead of
|
|
2582
|
-
letting the model pick a mood: findings first, then
|
|
2583
|
-
`request-changes` if any High exists or two Mediums do, otherwise
|
|
2584
|
-
`approve`. Pre-existing issues visible in context are scoped out, at
|
|
2585
|
-
most one advisory Low. The lead applies one rule on top: the PR is
|
|
2586
|
-
APPROVE only when every non-advisory area approved.
|
|
2587
|
-
|
|
2588
|
-
## Run the evals
|
|
2589
|
-
|
|
2590
|
-
```bash
|
|
2591
|
-
agent-sdk eval --dir examples/codeowners-review --list
|
|
2592
|
-
agent-sdk eval --dir examples/codeowners-review review/multi-area
|
|
2593
|
-
```
|
|
2594
|
-
|
|
2595
|
-
`review/multi-area` gates the whole pipeline: `prepare_review` runs, at
|
|
2596
|
-
least three subagent delegations happen, the reply carries every area
|
|
2597
|
-
section plus alice's advisory notes, the planted billing and migration
|
|
2598
|
-
bugs surface, and the verdict requests changes. `review/clean-approve`
|
|
2599
|
-
proves the approval path on the clean fixture, and
|
|
2600
|
-
`review/routing-question` gates that routing answers come from
|
|
2601
|
-
`list_review_areas`.
|
|
2602
|
-
|
|
2603
|
-
## Reuse the ownership-routing pattern
|
|
2604
|
-
|
|
2605
|
-
Copy this shape when different code deserves different judgment:
|
|
2606
|
-
|
|
2607
|
-
- Route with data and code, not prompt instructions. Tables and globs
|
|
2608
|
-
are testable.
|
|
2609
|
-
- Write one playbook per domain and keep each reviewer blind to the
|
|
2610
|
-
others.
|
|
2611
|
-
- Make verdicts mechanical so aggregation is arithmetic, not
|
|
2612
|
-
negotiation.
|
|
2613
|
-
- Ship fixtures with planted findings so the review quality itself is
|
|
2614
|
-
testable offline.
|
|
2615
|
-
|
|
2616
|
-
## Where to go next
|
|
2617
|
-
|
|
2618
|
-
- [Approval Buddy](/docs/example-agents/approval-buddy.md) for posting real approvals
|
|
2619
|
-
- [Subagents](/docs/reference/subagents.md)
|
|
2620
|
-
- [GitHub webhooks](/docs/guides/github.md)
|
|
2621
|
-
- [Evals](/docs/evals.md)
|
|
2622
|
-
|
|
2623
|
-
---
|
|
2624
|
-
|
|
2625
|
-
Source: /docs/example-agents/concierge.md
|
|
2626
|
-
|
|
2627
|
-
# Compose agents with a concierge
|
|
2628
|
-
|
|
2629
|
-
Concierge answers general questions itself and sends every weather question
|
|
2630
|
-
to the weather agent. The connection is one file. The Agent SDK turns the target
|
|
2631
|
-
agent's MCP endpoint into tools the concierge can call.
|
|
2632
|
-
|
|
2633
|
-
Use this example when two agents are useful on their own and one should
|
|
2634
|
-
delegate a narrow class of work to the other.
|
|
2635
|
-
|
|
2636
|
-
[Browse the Concierge source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/concierge/)
|
|
2637
|
-
|
|
2638
|
-
## Delegate through a peer MCP connection
|
|
2639
|
-
|
|
2640
|
-
Concierge has no domain tool of its own. Its capability comes from a peer MCP
|
|
2641
|
-
connection:
|
|
2642
|
-
|
|
2643
|
-
```ts
|
|
2644
|
-
export default defineConnection({
|
|
2645
|
-
agent: "weather-agent",
|
|
2646
|
-
description:
|
|
2647
|
-
"The weather-agent peer: delegate weather questions with ask; it runs its own tools (live Open-Meteo data) in its own context.",
|
|
2648
|
-
});
|
|
2649
|
-
```
|
|
2650
|
-
|
|
2651
|
-
The filename
|
|
2652
|
-
[`weather.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/concierge/agent/mcp-connections/weather.ts)
|
|
2653
|
-
makes the MCP server name `weather`. The `agent` field points to the sibling
|
|
2654
|
-
project's mount slug.
|
|
2655
|
-
|
|
2656
|
-
This differs from a subagent. A peer keeps its own:
|
|
2657
|
-
|
|
2658
|
-
- root instructions,
|
|
2659
|
-
- tools and MCP connections,
|
|
2660
|
-
- durable sessions,
|
|
2661
|
-
- playground, and
|
|
2662
|
-
- public MCP endpoint.
|
|
2663
|
-
|
|
2664
|
-
An SDK subagent inherits the parent's execution surface and only its parent
|
|
2665
|
-
can invoke it. See [Agent-to-agent](/docs/guides/agent-to-agent.md) for the full
|
|
2666
|
-
comparison.
|
|
2667
|
-
|
|
2668
|
-
## Follow a delegated request
|
|
2669
|
-
|
|
2670
|
-
1. A user asks Concierge what to pack for Paris.
|
|
2671
|
-
2. [`instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/concierge/agent/instructions.md)
|
|
2672
|
-
classifies packing advice as weather-related.
|
|
2673
|
-
3. The model calls `weather.ask` with the city, timeframe, units, and the
|
|
2674
|
-
complete question.
|
|
2675
|
-
4. The Agent SDK creates an MCP-channel session inside `weather-agent`.
|
|
2676
|
-
5. Weather agent calls its own Open-Meteo tools and returns a reply.
|
|
2677
|
-
6. If the turn exceeds the bounded MCP wait, `ask` returns
|
|
2678
|
-
`status: "running"`. Concierge calls `weather.check` with the returned
|
|
2679
|
-
`sessionId`.
|
|
2680
|
-
7. Concierge relays the result and may add one sentence of travel advice.
|
|
2681
|
-
|
|
2682
|
-
The weather session appears in the weather agent's playground. It doesn't
|
|
2683
|
-
share Concierge's conversation history.
|
|
2684
|
-
|
|
2685
|
-
## Map the delegation files
|
|
2686
|
-
|
|
2687
|
-
| File | Purpose |
|
|
2688
|
-
| --- | --- |
|
|
2689
|
-
| [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/concierge/agent/agent.ts) | Describes the root agent and selects the local runtime. |
|
|
2690
|
-
| [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/concierge/agent/instructions.md) | Draws a strict weather-only delegation boundary. |
|
|
2691
|
-
| [`agent/mcp-connections/weather.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/concierge/agent/mcp-connections/weather.ts) | Resolves the peer by its `weather-agent` slug. |
|
|
2692
|
-
| [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/concierge/agent/storage.ts) | Persists sessions and events with `cursorHostedStorage`. |
|
|
2693
|
-
|
|
2694
|
-
Concierge doesn't author channels, tools, skills, subagents, schedules,
|
|
2695
|
-
hooks, A/B experiments, or evals. The built-in HTTP and MCP surfaces still
|
|
2696
|
-
exist.
|
|
2697
|
-
|
|
2698
|
-
Its own MCP endpoint exposes `ask` and `check`. It doesn't expose
|
|
2699
|
-
`call_tool` because Concierge has no server tools. The target weather agent
|
|
2700
|
-
does expose `call_tool`, so that tool also appears under Concierge's
|
|
2701
|
-
`weather` connection.
|
|
2702
|
-
|
|
2703
|
-
## Mount both agents
|
|
2704
|
-
|
|
2705
|
-
A peer can only resolve within a multi-agent serve host. Validating Concierge
|
|
2706
|
-
alone checks its files, but serving it alone fails because `weather-agent`
|
|
2707
|
-
isn't mounted.
|
|
2708
|
-
|
|
2709
|
-
From this package, validate both projects:
|
|
2710
|
-
|
|
2711
|
-
```bash
|
|
2712
|
-
agent-sdk validate --dir examples/concierge
|
|
2713
|
-
agent-sdk validate --dir examples/weather-agent
|
|
2714
|
-
```
|
|
2715
|
-
|
|
2716
|
-
Don't serve the repository's whole `examples/` directory for this proof.
|
|
2717
|
-
Several advanced examples subscribe to live GitHub events. Create an ignored
|
|
2718
|
-
two-project mount instead. Copy only the authored files needed for this proof,
|
|
2719
|
-
leaving Weather's Slack channels out:
|
|
2720
|
-
|
|
2721
|
-
```bash
|
|
2722
|
-
PAIR_DIR=$(mktemp -d "${TMPDIR:-/tmp}/concierge-weather.XXXXXX")
|
|
2723
|
-
mkdir -p "$PAIR_DIR/concierge" "$PAIR_DIR/weather-agent/agent"
|
|
2724
|
-
cp -R examples/concierge/agent "$PAIR_DIR/concierge/"
|
|
2725
|
-
cp examples/concierge/package.json "$PAIR_DIR/concierge/"
|
|
2726
|
-
cp examples/weather-agent/agent/{agent.ts,instructions.md,ab.ts,ab.config.ts} \
|
|
2727
|
-
"$PAIR_DIR/weather-agent/agent/"
|
|
2728
|
-
cp -R examples/weather-agent/agent/{tools,skills,mcp-connections,subagents,schedules,hooks,lib} \
|
|
2729
|
-
"$PAIR_DIR/weather-agent/agent/"
|
|
2730
|
-
cp -R examples/weather-agent/mcp "$PAIR_DIR/weather-agent/"
|
|
2731
|
-
cp examples/weather-agent/package.json "$PAIR_DIR/weather-agent/"
|
|
2732
|
-
agent-sdk dev "$PAIR_DIR"
|
|
2733
|
-
```
|
|
2734
|
-
|
|
2735
|
-
The host resolves the peer after it knows every mount. The local peer URL is
|
|
2736
|
-
`http://127.0.0.1:3000/weather-agent/v1/mcp`. You still need an agent-runtime
|
|
2737
|
-
credential for both model turns.
|
|
2738
|
-
|
|
2739
|
-
## Exercise delegation
|
|
2740
|
-
|
|
2741
|
-
Send a weather request to the running Concierge:
|
|
2742
|
-
|
|
2743
|
-
```bash
|
|
2744
|
-
agent-sdk chat \
|
|
2745
|
-
--url http://127.0.0.1:3000/concierge \
|
|
2746
|
-
--message "What should I pack for Paris tomorrow?"
|
|
2747
|
-
```
|
|
2748
|
-
|
|
2749
|
-
Open both playgrounds:
|
|
2750
|
-
|
|
2751
|
-
- `http://127.0.0.1:3000/concierge/playground`
|
|
2752
|
-
- `http://127.0.0.1:3000/weather-agent/playground`
|
|
2753
|
-
|
|
2754
|
-
The Concierge transcript shows the MCP call. The weather playground shows a
|
|
2755
|
-
separate session on the `mcp` channel with live weather tool calls.
|
|
2756
|
-
|
|
2757
|
-
Now send a general request:
|
|
2758
|
-
|
|
2759
|
-
```bash
|
|
2760
|
-
agent-sdk chat \
|
|
2761
|
-
--url http://127.0.0.1:3000/concierge \
|
|
2762
|
-
--message "Give me three ideas for a quiet weekend."
|
|
2763
|
-
```
|
|
2764
|
-
|
|
2765
|
-
The instructions tell Concierge to answer without delegating. This contrast is
|
|
2766
|
-
the proof loop: weather goes to the peer, unrelated work stays local.
|
|
2767
|
-
|
|
2768
|
-
## Preserve peer context
|
|
2769
|
-
|
|
2770
|
-
`weather.ask` returns a peer `sessionId`. Passing it back to a later `ask`
|
|
2771
|
-
continues the same weather conversation. Concierge's instructions require
|
|
2772
|
-
this for follow-ups dependent on an earlier answer.
|
|
2773
|
-
|
|
2774
|
-
Use a fresh call when the tasks are independent. Reuse the peer session when
|
|
2775
|
-
the second question needs facts or choices from the first.
|
|
2776
|
-
|
|
2777
|
-
## Keep delegation bounded
|
|
2778
|
-
|
|
2779
|
-
The Agent SDK rejects unknown peer slugs and self-references during startup. It
|
|
2780
|
-
doesn't stop a cycle across several valid peers. If agent A delegates all work
|
|
2781
|
-
to B and B delegates all work to A, they can recurse.
|
|
2782
|
-
|
|
2783
|
-
The prompt provides the guardrail here:
|
|
2784
|
-
|
|
2785
|
-
- delegate every weather request,
|
|
2786
|
-
- include complete context, and
|
|
2787
|
-
- never delegate unrelated work.
|
|
2788
|
-
|
|
2789
|
-
Write similarly narrow routing rules for each peer. A tool description helps
|
|
2790
|
-
the model choose the connection, but the always-on instructions own the
|
|
2791
|
-
policy.
|
|
2792
|
-
|
|
2793
|
-
## Use peers from cloud turns
|
|
2794
|
-
|
|
2795
|
-
Local turns reach peers over loopback. A cloud VM can't reach the serve
|
|
2796
|
-
host's loopback address. Set a public URL when a cloud agent needs the peer:
|
|
2797
|
-
|
|
2798
|
-
```bash
|
|
2799
|
-
agent-sdk serve --dir "$PAIR_DIR" \
|
|
2800
|
-
--public-url https://agents.example.com \
|
|
2801
|
-
--bearer-token "$AGENT_TOKEN"
|
|
2802
|
-
```
|
|
2803
|
-
|
|
2804
|
-
The Agent SDK attaches the bearer token to peer calls. Without `--public-url`,
|
|
2805
|
-
cloud turns omit peer connections and the server logs a warning.
|
|
2806
|
-
|
|
2807
|
-
## Compose your own pair
|
|
2808
|
-
|
|
2809
|
-
To compose your own agents:
|
|
2810
|
-
|
|
2811
|
-
1. Give each project a stable directory slug.
|
|
2812
|
-
2. Add `agent/mcp-connections/<name>.ts` to the caller.
|
|
2813
|
-
3. Set `agent` to the target slug.
|
|
2814
|
-
4. Describe the exact work the peer owns.
|
|
2815
|
-
5. Mount both projects from their parent directory.
|
|
2816
|
-
6. Add evals for delegated and non-delegated requests.
|
|
2817
|
-
|
|
2818
|
-
Keep the peer independently useful. If the specialist only makes sense inside
|
|
2819
|
-
one parent and needs no independent sessions, use a subagent instead.
|
|
2820
|
-
|
|
2821
|
-
## Where to go next
|
|
2822
|
-
|
|
2823
|
-
- [Agent-to-agent](/docs/guides/agent-to-agent.md)
|
|
2824
|
-
- [MCP connections](/docs/reference/connections.md)
|
|
2825
|
-
- [Subagents](/docs/reference/subagents.md)
|
|
2826
|
-
- [Sessions and streaming](/docs/reference/sessions.md)
|
|
2827
|
-
|
|
2828
|
-
---
|
|
2829
|
-
|
|
2830
|
-
Source: /docs/example-agents/index.md
|
|
2831
|
-
|
|
2832
|
-
# Choose the right Agent SDK example
|
|
2833
|
-
|
|
2834
|
-
The examples progress from one-channel assistants to durable, event-driven
|
|
2835
|
-
workflows. Start with the smallest agent for your use case. Each guide
|
|
2836
|
-
explains its request flow, framework features, verification path, and reusable
|
|
2837
|
-
design.
|
|
2838
|
-
|
|
2839
|
-
The source projects live under
|
|
2840
|
-
[`examples/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/). Run the commands below from
|
|
2841
|
-
this package. See [Run the CLI](/docs/index.md#run-the-cli) if the
|
|
2842
|
-
`agent-sdk` command isn't installed.
|
|
2843
|
-
|
|
2844
|
-
## Compare the examples
|
|
2845
|
-
|
|
2846
|
-
| Agent | Runtime | Intake | Framework focus | What sets it apart |
|
|
2847
|
-
| --- | --- | --- | --- | --- |
|
|
2848
|
-
| [Weather agent](/docs/example-agents/weather-agent.md) | Cloud | HTTP and two Slack transports | Tools, stdio MCP, skill, subagent, schedule, hooks, A/B, and evals | It demonstrates the broad cloud-runtime surface in one domain. |
|
|
2849
|
-
| [Slack agent](/docs/example-agents/slack-agent.md) | Local | Account-linked Slack | Channel identity, threads, and suggested prompts | It reaches Slack without authored tools. |
|
|
2850
|
-
| [Concierge](/docs/example-agents/concierge.md) | Local | Built-in HTTP | Peer MCP and multi-agent serving | It delegates to a separate agent with its own tools, sessions, and context. |
|
|
2851
|
-
| [Playbook router](/docs/example-agents/benny.md) | Local with repo context | Two Slack transports | Channel watching, inherited skills, custom cwd, and an eval | An allowlisted Slack channel becomes an intake queue for repo playbooks. |
|
|
2852
|
-
| [Alert investigator](/docs/example-agents/oncall.md) | Local | Watched Slack alerts channel | Bot-post channel watching, per-thread debounce, reminder tools, and host Slack calls | Every alert gets a thread-pinned investigation that schedules its own re-checks. |
|
|
2853
|
-
| [PR evidence reviewer](/docs/example-agents/bugbot.md) | Local | Custom HTTP and Slack | Host tool, skill, seeded workspaces, and an eval | The model receives a prepared diff-first evidence tree instead of a checkout. |
|
|
2854
|
-
| [Approval Buddy](/docs/example-agents/approval-buddy.md) | Local | GitHub and Slack | Policy tools, two subagents, durable storage, and evals | Code decides whether a PR may be approved. Reviews stay informational. |
|
|
2855
|
-
| [Security Reviewer](/docs/example-agents/security-reviewer.md) | Local host pipeline | GitHub and chat | Staged tools, parallel SDK agents, progress UI, durable storage, A/B, and evals | Reviewers and triage overlap while the playground shows every stage. |
|
|
2856
|
-
| [Knowledge base](/docs/example-agents/knowledge-base.md) | Local | Built-in HTTP chat | Durable host-side state, a conventions skill, a schedule, unit tests, and evals | People curate shared facts in chat, and fresh sessions retrieve them from markdown. |
|
|
2857
|
-
| [Codebase wiki](/docs/example-agents/codebase-wiki.md) | Local | GitHub and chat | Task-dispatch webhooks, seeded digests, a mapping skill, a schedule, and evals | Merged PRs accumulate into per-feature wiki pages with a daily digest. |
|
|
2858
|
-
| [Codeowners review](/docs/example-agents/codeowners-review.md) | Local | GitHub, chat, and fixtures | Ownership routing in code, playbook data files, parallel subagents, and evals | Each product area reviews with its own playbook, and verdicts aggregate mechanically. |
|
|
2859
|
-
|
|
2860
|
-
## Pick a learning path
|
|
2861
|
-
|
|
2862
|
-
Use this order when you want to learn the Agent SDK one capability at a time:
|
|
2863
|
-
|
|
2864
|
-
1. Start with [Weather agent](/docs/example-agents/weather-agent.md) to explore the filesystem
|
|
2865
|
-
conventions and cloud runtime.
|
|
2866
|
-
2. Strip the project back to [Slack agent](/docs/example-agents/slack-agent.md) to see the
|
|
2867
|
-
minimum channel surface.
|
|
2868
|
-
3. Read [Playbook router](/docs/example-agents/benny.md) when Slack should route requests into repo
|
|
2869
|
-
playbooks.
|
|
2870
|
-
4. Continue to [Alert investigator](/docs/example-agents/oncall.md) when the intake is bot
|
|
2871
|
-
posts and the agent must pace its own engagement and re-checks.
|
|
2872
|
-
5. Add composition with [Concierge](/docs/example-agents/concierge.md).
|
|
2873
|
-
6. Study [PR evidence reviewer](/docs/example-agents/bugbot.md) before giving a model repository
|
|
2874
|
-
evidence.
|
|
2875
|
-
7. Move policy into code with [Approval Buddy](/docs/example-agents/approval-buddy.md).
|
|
2876
|
-
8. Study [Security Reviewer](/docs/example-agents/security-reviewer.md) for host-side PR
|
|
2877
|
-
work.
|
|
2878
|
-
9. See parallel subagent delegation carry team judgment in
|
|
2879
|
-
[Codeowners review](/docs/example-agents/codeowners-review.md).
|
|
2880
|
-
10. Curate team context through conversation with
|
|
2881
|
-
[Knowledge base](/docs/example-agents/knowledge-base.md), then let GitHub events maintain
|
|
2882
|
-
product documentation in [Codebase wiki](/docs/example-agents/codebase-wiki.md).
|
|
2883
|
-
|
|
2884
|
-
## Common prerequisites
|
|
2885
|
-
|
|
2886
|
-
All examples require:
|
|
2887
|
-
|
|
2888
|
-
- Node 22.13 or newer. Don't run the Agent SDK under Bun.
|
|
2889
|
-
- Workspace dependencies installed.
|
|
2890
|
-
- An agent-runtime credential for model turns.
|
|
2891
|
-
|
|
2892
|
-
Several examples need more:
|
|
2893
|
-
|
|
2894
|
-
- Account-linked Slack channels require a connected host account.
|
|
2895
|
-
- Alert investigator needs a dedicated Socket Mode app with channel-post
|
|
2896
|
-
events and membership in the watched alerts channel.
|
|
2897
|
-
- GitHub examples require access to the target repository. Codebase wiki and
|
|
2898
|
-
Codeowners review call the host `gh` CLI for PR data; the codeowners
|
|
2899
|
-
fixtures run without network.
|
|
2900
|
-
- Example agents use `cursorHostedStorage` in `agent/storage.ts` for hosted session storage. See [Storage](/docs/storage.md).
|
|
2901
|
-
|
|
2902
|
-
Each guide lists its own credentials, services, and side effects.
|
|
2903
|
-
|
|
2904
|
-
## Validate any example
|
|
2905
|
-
|
|
2906
|
-
Discovery commands don't start a model turn:
|
|
2907
|
-
|
|
2908
|
-
```bash
|
|
2909
|
-
agent-sdk validate --dir examples/weather-agent
|
|
2910
|
-
agent-sdk info --dir examples/weather-agent --json
|
|
2911
|
-
```
|
|
2912
|
-
|
|
2913
|
-
Start one development server with `agent-sdk dev examples/<name>`.
|
|
2914
|
-
Concierge depends on Weather agent, so its guide creates an isolated
|
|
2915
|
-
two-project mount. Don't mount the whole examples directory to test one
|
|
2916
|
-
agent; several advanced examples subscribe to live GitHub events.
|
|
2917
|
-
|
|
2918
|
-
## Read by framework feature
|
|
2919
|
-
|
|
2920
|
-
- [Concepts](/docs/concepts.md) explains filesystem discovery and runtime
|
|
2921
|
-
boundaries.
|
|
2922
|
-
- [Project layout](/docs/reference/project-layout.md) lists every authored
|
|
2923
|
-
folder.
|
|
2924
|
-
- [Tools](/docs/reference/tools.md), [channels](/docs/reference/channels.md), and
|
|
2925
|
-
[MCP connections](/docs/reference/connections.md) cover the core extension
|
|
2926
|
-
points.
|
|
2927
|
-
- [Evals](/docs/evals.md) and [live A/B metrics](/docs/ab.md) cover measured
|
|
2928
|
-
iteration.
|
|
2929
|
-
- [Deployment](/docs/deployment.md) covers credentials, auth, storage, and
|
|
2930
|
-
hosting.
|
|
2931
|
-
|
|
2932
|
-
---
|
|
2933
|
-
|
|
2934
|
-
Source: /docs/example-agents/knowledge-base.md
|
|
2935
|
-
|
|
2936
|
-
# Build a team knowledge base through conversation
|
|
2937
|
-
|
|
2938
|
-
Knowledge base turns conversations into shared team context. Teach the agent
|
|
2939
|
-
about people, systems, decisions, and standing preferences. Three server
|
|
2940
|
-
tools read, search, and write human-readable markdown pages; a conventions
|
|
2941
|
-
skill shapes each write; and a daily schedule merges duplicates and rebuilds
|
|
2942
|
-
the index. A fresh session retrieves what an earlier conversation captured.
|
|
2943
|
-
|
|
2944
|
-
Use this project when people should curate organizational knowledge through
|
|
2945
|
-
chat. Use [Codebase wiki](/docs/example-agents/codebase-wiki.md) when merged PRs should maintain
|
|
2946
|
-
feature documentation instead.
|
|
2947
|
-
|
|
2948
|
-
[Browse the knowledge base source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/knowledge-base/)
|
|
2949
|
-
|
|
2950
|
-
## Keep shared knowledge on the filesystem
|
|
2951
|
-
|
|
2952
|
-
The knowledge base lives outside any session workspace, in a wiki
|
|
2953
|
-
directory on the serve host by default. `KNOWLEDGE_BASE_DIR` overrides the location,
|
|
2954
|
-
and the tools resolve it on every call, so tests and evals can point the same
|
|
2955
|
-
code at a temp directory.
|
|
2956
|
-
|
|
2957
|
-
The store enforces its own safety:
|
|
2958
|
-
|
|
2959
|
-
- Page ids are one to three lowercase kebab-case segments, so a page id
|
|
2960
|
-
can't escape the wiki directory.
|
|
2961
|
-
- Pages cap at 64 KiB. Oversized writes fail with instructions to split
|
|
2962
|
-
the page.
|
|
2963
|
-
- `wiki_write` replaces whole pages. The instructions require reading a
|
|
2964
|
-
page before updating it, so rewrites carry existing facts forward.
|
|
2965
|
-
|
|
2966
|
-
Every page is plain markdown. You can open the wiki in an editor,
|
|
2967
|
-
review it in a PR, or grep it.
|
|
2968
|
-
|
|
2969
|
-
## Follow a fact through the agent
|
|
2970
|
-
|
|
2971
|
-
1. You tell the agent something durable: a system, an owner, a standing
|
|
2972
|
-
preference.
|
|
2973
|
-
2. The instructions require a `wiki_search` before claiming knowledge
|
|
2974
|
-
and a `wiki_write` after learning something worth keeping.
|
|
2975
|
-
3. The `wiki-conventions` skill picks the page id (`staging-database`,
|
|
2976
|
-
`people/jane-doe`), the page shape, and the dated fact format.
|
|
2977
|
-
4. The tool writes the page under the durable wiki root and returns
|
|
2978
|
-
whether it created or updated the page.
|
|
2979
|
-
5. A later session, on any channel, finds the fact with `wiki_search`
|
|
2980
|
-
and cites the knowledge-base page in its answer.
|
|
2981
|
-
|
|
2982
|
-
Ephemeral chatter stays out. The instructions tell the model to skip
|
|
2983
|
-
one-off questions and to ask before saving anything borderline.
|
|
2984
|
-
|
|
2985
|
-
## Map the knowledge-base files
|
|
2986
|
-
|
|
2987
|
-
| File | Purpose |
|
|
2988
|
-
| --- | --- |
|
|
2989
|
-
| [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/agent.ts) | Selects the local runtime and model. |
|
|
2990
|
-
| [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/instructions.md) | Sets the read-before-answer and save-after-learning policy. |
|
|
2991
|
-
| [`agent/lib/wiki-store.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/lib/wiki-store.ts) | Validates page ids, lists, reads, writes, and searches the knowledge base. |
|
|
2992
|
-
| [`agent/tools/wiki_read.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/tools/wiki_read.ts) | Reads one page or lists every page with titles and timestamps. |
|
|
2993
|
-
| [`agent/tools/wiki_search.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/tools/wiki_search.ts) | Searches titles and bodies with per-page match lines. |
|
|
2994
|
-
| [`agent/tools/wiki_write.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/tools/wiki_write.ts) | Creates or replaces a page and reports created versus updated. |
|
|
2995
|
-
| [`agent/skills/wiki-conventions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/skills/wiki-conventions.md) | Names pages, shapes them, and dates every fact. |
|
|
2996
|
-
| [`agent/schedules/gardener.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/schedules/gardener.md) | Merges duplicates, rebuilds the index, and flags stale facts daily. |
|
|
2997
|
-
| [`agent/lib/wiki-store.test.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/lib/wiki-store.test.ts) | Unit-tests slug safety and store round-trips. |
|
|
2998
|
-
| [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/storage.ts) | Persists sessions and events with `cursorHostedStorage`. |
|
|
2999
|
-
| [`evals/evals.config.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/evals/evals.config.ts) | Caps eval run concurrency. |
|
|
3000
|
-
| [`evals/knowledge.eval.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/evals/knowledge.eval.ts) | Seeds a temp knowledge base and gates recall, save, and no-write decisions. |
|
|
3001
|
-
|
|
3002
|
-
There is no authored channel, MCP connection, subagent, hook, or A/B
|
|
3003
|
-
experiment. The wiki directory is the durable knowledge store.
|
|
3004
|
-
|
|
3005
|
-
## Prepare the example
|
|
3006
|
-
|
|
3007
|
-
You need:
|
|
3008
|
-
|
|
3009
|
-
- Node 22.13 or newer.
|
|
3010
|
-
- An agent-runtime credential for model turns.
|
|
3011
|
-
|
|
3012
|
-
Nothing else. The wiki is created on first write.
|
|
3013
|
-
|
|
3014
|
-
## Validate the surface
|
|
3015
|
-
|
|
3016
|
-
```bash
|
|
3017
|
-
agent-sdk validate --dir examples/knowledge-base
|
|
3018
|
-
agent-sdk info --dir examples/knowledge-base --json
|
|
3019
|
-
```
|
|
3020
|
-
|
|
3021
|
-
The manifest should report three server tools, one skill, and one
|
|
3022
|
-
schedule.
|
|
3023
|
-
|
|
3024
|
-
## Exercise the store without a model turn
|
|
3025
|
-
|
|
3026
|
-
```bash
|
|
3027
|
-
agent-sdk call wiki_write \
|
|
3028
|
-
--dir examples/knowledge-base \
|
|
3029
|
-
--input '{"page":"staging-database","content":"# Staging database\n\n- Port: 6432 (recorded 2026-07-19)\n"}'
|
|
3030
|
-
|
|
3031
|
-
agent-sdk call wiki_search \
|
|
3032
|
-
--dir examples/knowledge-base \
|
|
3033
|
-
--input '{"query":"6432"}'
|
|
3034
|
-
|
|
3035
|
-
agent-sdk call wiki_read --dir examples/knowledge-base --input '{}'
|
|
3036
|
-
```
|
|
3037
|
-
|
|
3038
|
-
Invalid page ids fail fast. Try `{"page":"../escape"}` and the tool
|
|
3039
|
-
returns the validation error instead of touching the filesystem.
|
|
3040
|
-
|
|
3041
|
-
## Prove recall across sessions
|
|
3042
|
-
|
|
3043
|
-
```bash
|
|
3044
|
-
agent-sdk dev examples/knowledge-base
|
|
3045
|
-
```
|
|
3046
|
-
|
|
3047
|
-
Teach it something in the playground:
|
|
3048
|
-
|
|
3049
|
-
> Remember: our staging database is Postgres at
|
|
3050
|
-
> staging-db.internal.example.com, port 6432 via PgBouncer. Jane Doe
|
|
3051
|
-
> owns it.
|
|
3052
|
-
|
|
3053
|
-
The trace shows the conventions skill load, then `wiki_write` calls
|
|
3054
|
-
for `staging-database`, `people/jane-doe`, and `index`. Start a new
|
|
3055
|
-
session and ask:
|
|
3056
|
-
|
|
3057
|
-
> What port does our staging database use, and who owns it?
|
|
3058
|
-
|
|
3059
|
-
The fresh session finds the answer with `wiki_search` and `wiki_read`
|
|
3060
|
-
and cites the pages. The conversation history is empty; the wiki is the
|
|
3061
|
-
source of truth.
|
|
3062
|
-
|
|
3063
|
-
## Run the gardener
|
|
3064
|
-
|
|
3065
|
-
The `gardener` schedule fires at 06:00 UTC and rewrites the wiki for
|
|
3066
|
-
consistency: merge near-duplicate pages, rebuild `index`, and flag
|
|
3067
|
-
facts older than 90 days. Under `agent-sdk dev`, timers don't auto-fire.
|
|
3068
|
-
Trigger it by hand:
|
|
3069
|
-
|
|
3070
|
-
```bash
|
|
3071
|
-
curl -s -X POST http://127.0.0.1:3000/knowledge-base/v1/dev/schedules/gardener
|
|
3072
|
-
```
|
|
3073
|
-
|
|
3074
|
-
## Run the evals
|
|
3075
|
-
|
|
3076
|
-
```bash
|
|
3077
|
-
agent-sdk eval --dir examples/knowledge-base --list
|
|
3078
|
-
agent-sdk eval --dir examples/knowledge-base knowledge/recall
|
|
3079
|
-
```
|
|
3080
|
-
|
|
3081
|
-
The eval file seeds a temp directory through `KNOWLEDGE_BASE_DIR`
|
|
3082
|
-
inside the cases, so the durable knowledge base never sees test data.
|
|
3083
|
-
`knowledge/recall` proves the fact comes from disk, not the conversation.
|
|
3084
|
-
`knowledge/save`
|
|
3085
|
-
gates the write decision, and `knowledge/no-write-on-ephemera` proves small
|
|
3086
|
-
talk stays out of the knowledge base.
|
|
3087
|
-
|
|
3088
|
-
## Reuse the knowledge-base pattern
|
|
3089
|
-
|
|
3090
|
-
Copy this shape when an agent needs durable, inspectable team knowledge:
|
|
3091
|
-
|
|
3092
|
-
- Resolve the storage root lazily behind an environment override.
|
|
3093
|
-
- Validate identifiers in the store, not in the prompt.
|
|
3094
|
-
- Put naming and structure conventions in a skill so writes stay
|
|
3095
|
-
consistent.
|
|
3096
|
-
- Add a consolidation schedule instead of letting pages rot.
|
|
3097
|
-
|
|
3098
|
-
## Where to go next
|
|
3099
|
-
|
|
3100
|
-
- [Tools](/docs/reference/tools.md)
|
|
3101
|
-
- [Skills](/docs/reference/skills.md)
|
|
3102
|
-
- [Schedules](/docs/reference/schedules.md)
|
|
3103
|
-
- [Evals](/docs/evals.md)
|
|
3104
|
-
|
|
3105
|
-
---
|
|
3106
|
-
|
|
3107
|
-
Source: /docs/example-agents/oncall.md
|
|
3108
|
-
|
|
3109
|
-
# Investigate every alert in its own Slack thread
|
|
3110
|
-
|
|
3111
|
-
This agent is an on-call teammate. Alert feeds post into an alerts channel
|
|
3112
|
-
as bots. Each new alert dispatches an investigation session pinned to that
|
|
3113
|
-
post's thread: the agent reacts 👀 the moment it locks in, investigates
|
|
3114
|
-
immediately, and posts brief findings backed by evidence it observed.
|
|
3115
|
-
Replies in the thread reach it only after the thread has been quiet for
|
|
3116
|
-
about a minute, and reminder tools let it wake itself later to re-check a
|
|
3117
|
-
baseline or confirm an alert cleared.
|
|
3118
|
-
|
|
3119
|
-
Use this example when alerts land in Slack and you want one thread-scoped
|
|
3120
|
-
investigation per alert, with an agent that paces its own engagement
|
|
3121
|
-
instead of answering every message.
|
|
3122
|
-
|
|
3123
|
-
[Browse the current alert-investigator source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/oncall/)
|
|
3124
|
-
|
|
3125
|
-
## Follow an alert
|
|
3126
|
-
|
|
3127
|
-
1. An alert feed (Alertmanager, PagerDuty, Datadog) posts a new top-level
|
|
3128
|
-
message in the watched alerts channel.
|
|
3129
|
-
2. The channel watch accepts it. `includeBotPosts` lets bot authors
|
|
3130
|
-
through; the agent's own posts always stay dropped.
|
|
3131
|
-
3. The handler reacts 👀 on the alert post and sets "Investigating…"
|
|
3132
|
-
typing. The reaction is the lock-in signal: this alert has an owner.
|
|
3133
|
-
4. The Agent SDK creates a session keyed to the alert's thread and dispatches
|
|
3134
|
-
immediately. New alerts get no debounce.
|
|
3135
|
-
5. The agent reads the alert, gathers evidence, and posts findings to the
|
|
3136
|
-
thread once it has a hypothesis.
|
|
3137
|
-
6. People discuss in the thread. Replies buffer per thread and dispatch as
|
|
3138
|
-
one coalesced follow-up after roughly a minute of quiet.
|
|
3139
|
-
7. The agent arms reminders for anything that needs time and posts interim
|
|
3140
|
-
updates when new evidence changes the picture.
|
|
3141
|
-
|
|
3142
|
-
Mentions and DMs skip the watch entirely and behave like ordinary chat.
|
|
3143
|
-
|
|
3144
|
-
## Map the files
|
|
3145
|
-
|
|
3146
|
-
| File | Purpose |
|
|
3147
|
-
| --- | --- |
|
|
3148
|
-
| [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/agent.ts) | Names the agent and keeps harness workspaces outside any monorepo checkout. |
|
|
3149
|
-
| [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/instructions.md) | Engagement rules, the investigation loop, and the message discipline. |
|
|
3150
|
-
| [`agent/channels/slack-app.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/channels/slack-app.ts) | Dedicated Socket Mode app: watch configuration and handler wiring. |
|
|
3151
|
-
| [`agent/lib/alert-watch.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/lib/alert-watch.ts) | The engagement policy: lock in on new alerts, coalesce replies. |
|
|
3152
|
-
| [`agent/lib/thread-debounce.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/lib/thread-debounce.ts) | Per-thread quiet window. |
|
|
3153
|
-
| [`agent/lib/alerts.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/lib/alerts.ts) | Dispatch classification, prompt building, and thread addressing. |
|
|
3154
|
-
| [`agent/lib/slack-api.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/lib/slack-api.ts) | Reactions and thread posts on this agent's own token pair. |
|
|
3155
|
-
| [`agent/tools/reminders_create.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/tools/reminders_create.ts) | Self-scheduled wakes bound to the thread (plus `reminders_list` and `reminders_cancel`). |
|
|
3156
|
-
| [`agent/tools/post_thread_update.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/tools/post_thread_update.ts) | Interim updates to the thread mid-turn. |
|
|
3157
|
-
| [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/storage.ts) | Persists sessions and events with `cursorHostedStorage`. |
|
|
3158
|
-
| [`evals/evals.config.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/evals/evals.config.ts) | Caps eval run concurrency. |
|
|
3159
|
-
| [`evals/smoke.eval.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/evals/smoke.eval.ts) | Checks identity and the reminder-tool route. |
|
|
3160
|
-
|
|
3161
|
-
## Let bot posts through the watch
|
|
3162
|
-
|
|
3163
|
-
Channel watching drops bot-authored posts by default so two agents can
|
|
3164
|
-
never feed each other. Alert channels invert the assumption: the posts
|
|
3165
|
-
worth watching come from bots. `channelPosts.includeBotPosts` opts in per
|
|
3166
|
-
channel:
|
|
3167
|
-
|
|
3168
|
-
```ts
|
|
3169
|
-
engagement: {
|
|
3170
|
-
channelPosts: {
|
|
3171
|
-
allow: ["#alerts"],
|
|
3172
|
-
posts: "all",
|
|
3173
|
-
includeBotPosts: true,
|
|
3174
|
-
},
|
|
3175
|
-
},
|
|
3176
|
-
```
|
|
3177
|
-
|
|
3178
|
-
Loop safety survives the opt-in. The pack matches the watching app's own
|
|
3179
|
-
posts by the `bot_id` and bot user id from `auth.test` and drops them, so
|
|
3180
|
-
the agent's findings never re-dispatch it. Posts that mention the bot stay
|
|
3181
|
-
on the mention path.
|
|
3182
|
-
|
|
3183
|
-
`posts: "all"` also delivers thread replies. The handler, not the pack,
|
|
3184
|
-
decides their pace.
|
|
3185
|
-
|
|
3186
|
-
## Pace the engagement
|
|
3187
|
-
|
|
3188
|
-
The example runs two rhythms:
|
|
3189
|
-
|
|
3190
|
-
- A new alert dispatches immediately.
|
|
3191
|
-
- Thread replies produce one engagement per lull.
|
|
3192
|
-
|
|
3193
|
-
The pack's `debounceMs` is per message; it exists to let edits settle. This
|
|
3194
|
-
agent needs a per-thread window instead, so the handler owns it
|
|
3195
|
-
([`lib/thread-debounce.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/lib/thread-debounce.ts)).
|
|
3196
|
-
Every reply restarts a 60-second timer keyed by thread. Superseded waiters
|
|
3197
|
-
resolve `null` and the handler returns `null` for them. When the thread
|
|
3198
|
-
goes quiet, the newest waiter receives the whole batch and dispatches one
|
|
3199
|
-
follow-up that lists every message with mentionable attribution.
|
|
3200
|
-
|
|
3201
|
-
Two details make the window matter. A follow-up that arrives while a turn
|
|
3202
|
-
runs preempts that turn (latest message wins), so engaging per message
|
|
3203
|
-
would keep cancelling the investigation. And @mentions bypass the window
|
|
3204
|
-
through Slack's mention path, so a person who needs the agent now still
|
|
3205
|
-
gets it now.
|
|
3206
|
-
|
|
3207
|
-
## Schedule your own re-checks
|
|
3208
|
-
|
|
3209
|
-
Investigations rarely finish in one pass. A baseline comparison needs 20
|
|
3210
|
-
minutes of data. An alert that cleared may re-fire. The example hands the
|
|
3211
|
-
model three tools over `host.reminders`:
|
|
3212
|
-
|
|
3213
|
-
- `reminders_create` arms a one-shot (`delay: "20m"`) or recurring
|
|
3214
|
-
(`every: "30m"` with a plain-language stop condition) wake bound to the
|
|
3215
|
-
thread's conversation.
|
|
3216
|
-
- `reminders_list` shows the thread's standing watches.
|
|
3217
|
-
- `reminders_cancel` disarms one, and refuses ids that belong to another
|
|
3218
|
-
thread's conversation.
|
|
3219
|
-
|
|
3220
|
-
When a reminder fires, its prompt returns to the same session as a
|
|
3221
|
-
follow-up turn, and the reply lands in the alert thread. The instructions
|
|
3222
|
-
keep wake prompts generic (re-read live state instead of replaying stale
|
|
3223
|
-
numbers) and wake replies to one line, for example "re-checked p99 on
|
|
3224
|
-
api-gateway: 120ms, back at baseline, cancelling the watch."
|
|
3225
|
-
|
|
3226
|
-
Keep these tool filenames if you copy the design: the framework's reminder
|
|
3227
|
-
fire prompt tells the model to call `reminders_cancel` by name when a stop
|
|
3228
|
-
condition is set.
|
|
3229
|
-
|
|
3230
|
-
## Alert people mid-investigation
|
|
3231
|
-
|
|
3232
|
-
The final reply of each turn posts to the thread on its own.
|
|
3233
|
-
`post_thread_update` covers evidence that shouldn't wait for the turn to
|
|
3234
|
-
finish: it posts a one-or-two-sentence update through the agent's token,
|
|
3235
|
-
with `<@USERID>` mentions for the people who need to act. The instructions
|
|
3236
|
-
restrict it to changes in hypothesis, severity, or blast radius. Progress
|
|
3237
|
-
narration doesn't qualify.
|
|
3238
|
-
|
|
3239
|
-
## Connect the Slack app
|
|
3240
|
-
|
|
3241
|
-
Channel watching is Socket Mode only, so this example uses a dedicated
|
|
3242
|
-
app:
|
|
3243
|
-
|
|
3244
|
-
```bash
|
|
3245
|
-
agent-sdk slack create --dir examples/oncall --name "Oncall" --channel-posts
|
|
3246
|
-
agent-sdk slack doctor --prefix ONCALL
|
|
3247
|
-
```
|
|
3248
|
-
|
|
3249
|
-
`--channel-posts` prefills channel-watch events (`message.channels` /
|
|
3250
|
-
`message.groups`). Invite the bot to each watched channel after the
|
|
3251
|
-
wizard finishes.
|
|
3252
|
-
|
|
3253
|
-
`ONCALL_ALERTS_CHANNELS` sets the watch list as comma-separated ids or
|
|
3254
|
-
`#names`. It defaults to `#alerts`.
|
|
3255
|
-
|
|
3256
|
-
Wire observability MCP servers under `agent/mcp-connections/` so evidence
|
|
3257
|
-
gathering reaches your logs, metrics, and dashboards. The example ships
|
|
3258
|
-
none; without them the agent works from the alert text, its links, and the
|
|
3259
|
-
thread.
|
|
3260
|
-
|
|
3261
|
-
## Validate and start the server
|
|
3262
|
-
|
|
3263
|
-
```bash
|
|
3264
|
-
agent-sdk validate --dir examples/oncall
|
|
3265
|
-
agent-sdk info --dir examples/oncall --json
|
|
3266
|
-
agent-sdk dev examples/oncall
|
|
3267
|
-
```
|
|
3268
|
-
|
|
3269
|
-
The info output lists four server tools and the watched channel on the
|
|
3270
|
-
`slack-app` channel. Missing tokens leave that channel idle without
|
|
3271
|
-
stopping the server.
|
|
3272
|
-
|
|
3273
|
-
In dev mode, reminder timers don't auto-fire. List and fire them by hand
|
|
3274
|
-
through the dev routes described in
|
|
3275
|
-
[Schedules and reminders](/docs/reference/schedules.md#dispatch-and-dev-mode).
|
|
3276
|
-
|
|
3277
|
-
## Test the policy without Slack
|
|
3278
|
-
|
|
3279
|
-
The engagement policy is plain code with unit tests:
|
|
3280
|
-
|
|
3281
|
-
```bash
|
|
3282
|
-
pnpm exec vitest run examples/oncall
|
|
3283
|
-
```
|
|
3284
|
-
|
|
3285
|
-
The integration test drives a synthetic Events API delivery through the
|
|
3286
|
-
real parse, watch, and dispatch plumbing. It asserts a bot alert
|
|
3287
|
-
dispatches pinned to its thread after the lock-in reaction, the agent's
|
|
3288
|
-
own posts never loop, and replies coalesce behind the quiet window.
|
|
3289
|
-
|
|
3290
|
-
The smoke eval spends a model turn:
|
|
3291
|
-
|
|
3292
|
-
```bash
|
|
3293
|
-
agent-sdk eval --dir examples/oncall smoke --json
|
|
3294
|
-
```
|
|
3295
|
-
|
|
3296
|
-
It checks identity and the reminder-tool route lexically. It doesn't prove
|
|
3297
|
-
Slack delivery or reaction behavior; the unit tests cover the dispatch
|
|
3298
|
-
side, and a live check needs the dedicated app connected.
|
|
3299
|
-
|
|
3300
|
-
## Build an alert investigator
|
|
3301
|
-
|
|
3302
|
-
Use this structure when a bot feed should drive thread-scoped work:
|
|
3303
|
-
|
|
3304
|
-
1. Watch the feed channel with `includeBotPosts: true` and a narrow
|
|
3305
|
-
allowlist.
|
|
3306
|
-
2. Acknowledge on the triggering post before dispatching, so people see
|
|
3307
|
-
ownership without opening the thread.
|
|
3308
|
-
3. Dispatch new items immediately; coalesce thread chatter behind a
|
|
3309
|
-
per-thread quiet window.
|
|
3310
|
-
4. Give the agent reminder tools for anything that needs time, and make
|
|
3311
|
-
cancel discipline part of the instructions.
|
|
3312
|
-
5. Keep every posted message brief and tied to evidence the agent saw.
|
|
3313
|
-
|
|
3314
|
-
## Where to go next
|
|
3315
|
-
|
|
3316
|
-
- [Slack](/docs/guides/slack.md)
|
|
3317
|
-
- [Schedules and reminders](/docs/reference/schedules.md)
|
|
3318
|
-
- [Tools](/docs/reference/tools.md)
|
|
3319
|
-
- [Playbook router](/docs/example-agents/benny.md) for the human-post variant of channel
|
|
3320
|
-
watching
|
|
3321
|
-
|
|
3322
|
-
---
|
|
3323
|
-
|
|
3324
|
-
Source: /docs/example-agents/security-reviewer.md
|
|
3325
|
-
|
|
3326
|
-
# Run staged security reviews from GitHub events
|
|
3327
|
-
|
|
3328
|
-
Security Reviewer turns a pull request into a staged host-side review. One
|
|
3329
|
-
tool prepares the diff and selects modules. A second fans out specialized
|
|
3330
|
-
reviewers and triages candidates as they arrive. A third deduplicates the
|
|
3331
|
-
confirmed findings, writes artifacts, and may publish a GitHub review.
|
|
3332
|
-
|
|
3333
|
-
Use this example when the workflow needs several model workers, but the host
|
|
3334
|
-
must own orchestration, progress, artifacts, and the final write.
|
|
3335
|
-
|
|
3336
|
-
Source lives under [`factory/security-reviewer/`](https://github.com/cursor/cursor/tree/main/factory/security-reviewer/) (factory agent, not under `examples/`).
|
|
3337
|
-
|
|
3338
|
-
Want one model turn and one comment? Scaffold the
|
|
3339
|
-
[security-reviewer template](/docs/templates/security-reviewer.md).
|
|
3340
|
-
|
|
3341
|
-
[Browse the Security Reviewer source.](https://github.com/cursor/cursor/tree/main/factory/security-reviewer/)
|
|
3342
|
-
|
|
3343
|
-
## Run a three-stage host pipeline
|
|
3344
|
-
|
|
3345
|
-
Security Reviewer is a pipeline, not one long agent turn:
|
|
3346
|
-
|
|
3347
|
-
| Stage | Tool | Result |
|
|
3348
|
-
| --- | --- | --- |
|
|
3349
|
-
| Prepare | `prepare_review` | Fetch metadata and diff, create a `runId`, and select security modules. |
|
|
3350
|
-
| Review and triage | `run_reviewers` | Run module reviewers in parallel and start triage as each candidate arrives. |
|
|
3351
|
-
| Finalize | `finalize_review` | Apply thresholds, deduplicate findings, write artifacts, and optionally post a review. |
|
|
3352
|
-
|
|
3353
|
-
`run_triage` remains available as a compatibility stage. In the normal flow,
|
|
3354
|
-
triage has already completed inside `run_reviewers`, so it reports existing
|
|
3355
|
-
results. If candidates exist without triage output, it starts triage workers
|
|
3356
|
-
and writes their state.
|
|
3357
|
-
|
|
3358
|
-
The configured root agent chooses and sequences tools in chat. The review
|
|
3359
|
-
workers use a model selected by the host pipeline. They are
|
|
3360
|
-
created programmatically with the agent SDK, not discovered from
|
|
3361
|
-
`agent/subagents/`.
|
|
3362
|
-
|
|
3363
|
-
## Follow a GitHub review
|
|
3364
|
-
|
|
3365
|
-
1. A pull request event starts a review and opens a playground session.
|
|
3366
|
-
2. The playground shows reviewer and triage progress.
|
|
3367
|
-
3. Confirmed findings appear in the PR review.
|
|
3368
|
-
4. A GitHub Check reports completion or a processing failure.
|
|
3369
|
-
|
|
3370
|
-
## Map the framework features
|
|
3371
|
-
|
|
3372
|
-
| Capability | Source | Role |
|
|
3373
|
-
| --- | --- | --- |
|
|
3374
|
-
| Root agent | [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/factory/security-reviewer/agent/agent.ts), [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/factory/security-reviewer/agent/instructions.md) | Configure local chat and explain the three-stage contract. |
|
|
3375
|
-
| Server tools | [`agent/tools/`](https://github.com/cursor/cursor/tree/main/factory/security-reviewer/agent/tools/) | Expose each review stage to chat and host orchestration. |
|
|
3376
|
-
| GitHub channel | [`agent/channels/github.ts`](https://github.com/cursor/cursor/blob/main/factory/security-reviewer/agent/channels/github.ts) | Filter wakes, run background tasks, and publish status. |
|
|
3377
|
-
| Progress channel | [`agent/channels/asr-progress.ts`](https://github.com/cursor/cursor/blob/main/factory/security-reviewer/agent/channels/asr-progress.ts) | Serve live reviewer and triage state by `runId`. |
|
|
3378
|
-
| Playground renderer | [`agent/playground/tools/run_reviewers.tsx`](https://github.com/cursor/cursor/blob/main/factory/security-reviewer/agent/playground/tools/run_reviewers.tsx) | Replace the generic tool chip with live module rows. |
|
|
3379
|
-
| SDK review pipeline | [`review-stages.ts`](https://github.com/cursor/cursor/blob/main/factory/security-reviewer/agent/lib/review-stages.ts), [`@anysphere/security-review-lib`](https://github.com/cursor/cursor/blob/main/packages/security-review-lib/src/index.ts) | Select modules, call model workers, triage, deduplicate, and write artifacts. |
|
|
3380
|
-
| Storage | [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/factory/security-reviewer/agent/storage.ts) | Persist framework sessions with `cursorHostedStorage` (lazy restore). |
|
|
3381
|
-
| A/B | [`agent/ab.ts`](https://github.com/cursor/cursor/blob/main/factory/security-reviewer/agent/ab.ts) | Compare all-severity versus high-only GitHub comments. |
|
|
3382
|
-
| Eval | [`evals/`](https://github.com/cursor/cursor/tree/main/factory/security-reviewer/evals/) | Check stage-tool presence against a pinned sample. |
|
|
3383
|
-
|
|
3384
|
-
There is no Slack channel, authored skill, discovered subagent, MCP
|
|
3385
|
-
connection, schedule, reminder, hook, tool approval, or cloud runtime.
|
|
3386
|
-
|
|
3387
|
-
## Prepare the host
|
|
3388
|
-
|
|
3389
|
-
You need:
|
|
3390
|
-
|
|
3391
|
-
- Node 22.13 or newer.
|
|
3392
|
-
- An agent-runtime credential for the root turn and review workers.
|
|
3393
|
-
- GitHub read access for preparation.
|
|
3394
|
-
- GitHub write access for webhook-driven reviews and Checks.
|
|
3395
|
-
|
|
3396
|
-
The pipeline exposes settings for:
|
|
3397
|
-
|
|
3398
|
-
- the worker model,
|
|
3399
|
-
- reviewer and triage parallelism,
|
|
3400
|
-
- reviewer, triage, duplicate-gate, and final-dedupe timeouts, and
|
|
3401
|
-
- prior-comment loading.
|
|
3402
|
-
|
|
3403
|
-
The active names live beside the orchestration in
|
|
3404
|
-
[`review-stages.ts`](https://github.com/cursor/cursor/blob/main/factory/security-reviewer/agent/lib/review-stages.ts).
|
|
3405
|
-
|
|
3406
|
-
## Validate the discovered agent
|
|
3407
|
-
|
|
3408
|
-
```bash
|
|
3409
|
-
agent-sdk validate --dir ../../factory/security-reviewer
|
|
3410
|
-
agent-sdk info --dir ../../factory/security-reviewer --json
|
|
3411
|
-
agent-sdk eval --dir ../../factory/security-reviewer --list
|
|
3412
|
-
```
|
|
3413
|
-
|
|
3414
|
-
`validate` should pass. `info` and `eval --list` should match the capabilities
|
|
3415
|
-
mapped above.
|
|
3416
|
-
|
|
3417
|
-
## Know the chat path's write boundary
|
|
3418
|
-
|
|
3419
|
-
In chat, the root instructions ask the model to use this order:
|
|
3420
|
-
|
|
3421
|
-
```text
|
|
3422
|
-
prepare_review -> run_reviewers -> finalize_review
|
|
3423
|
-
```
|
|
3424
|
-
|
|
3425
|
-
They also ask the model to set `postComment: true` only on request. This is
|
|
3426
|
-
prompt policy, not a deterministic safety gate. The model chooses tool
|
|
3427
|
-
arguments, and `finalize_review` has no human approval. Use the direct stage
|
|
3428
|
-
calls below when a no-post proof must be enforced.
|
|
3429
|
-
|
|
3430
|
-
## Call stages directly without publishing
|
|
3431
|
-
|
|
3432
|
-
Call each stage and pass `postComment: false` yourself:
|
|
3433
|
-
|
|
3434
|
-
```bash
|
|
3435
|
-
agent-sdk call prepare_review \
|
|
3436
|
-
--dir ../../factory/security-reviewer \
|
|
3437
|
-
--input '{"prUrl":"https://github.com/owner/repo/pull/123"}'
|
|
3438
|
-
|
|
3439
|
-
agent-sdk call run_reviewers \
|
|
3440
|
-
--dir ../../factory/security-reviewer \
|
|
3441
|
-
--input '{"runId":"<run-id>"}'
|
|
3442
|
-
|
|
3443
|
-
agent-sdk call finalize_review \
|
|
3444
|
-
--dir ../../factory/security-reviewer \
|
|
3445
|
-
--input '{"runId":"<run-id>","postComment":false}'
|
|
3446
|
-
```
|
|
3447
|
-
|
|
3448
|
-
Review state lives under the project's run-artifact directory, so later
|
|
3449
|
-
stages can open the prepared `runId`.
|
|
3450
|
-
|
|
3451
|
-
> [!CAUTION]
|
|
3452
|
-
> `finalize_review` with `postComment: true` writes to GitHub. The webhook
|
|
3453
|
-
> path always requests that write. Chat instructions alone don't prevent it.
|
|
3454
|
-
|
|
3455
|
-
## Watch parallel work in the playground
|
|
3456
|
-
|
|
3457
|
-
Run the dev server:
|
|
3458
|
-
|
|
3459
|
-
```bash
|
|
3460
|
-
agent-sdk dev ../../factory/security-reviewer
|
|
3461
|
-
```
|
|
3462
|
-
|
|
3463
|
-
Open the printed playground and start a review. The custom
|
|
3464
|
-
`run_reviewers` renderer polls the progress channel's `GET /:runId` route.
|
|
3465
|
-
|
|
3466
|
-
It refreshes every 500 ms while the stage runs. Each row shows a reviewer
|
|
3467
|
-
module's state, candidates, reviewed areas, and failure. A second section
|
|
3468
|
-
shows triage jobs and confirmed or rejected counts.
|
|
3469
|
-
|
|
3470
|
-
This is an authored playground extension. The Agent SDK discovers it by the tool
|
|
3471
|
-
name, so the generic `run_reviewers` chip becomes a domain-specific view
|
|
3472
|
-
without changing the framework playground.
|
|
3473
|
-
|
|
3474
|
-
## Fan out reviewers while triage starts
|
|
3475
|
-
|
|
3476
|
-
Module selection uses repository and path rules. The current module set
|
|
3477
|
-
covers:
|
|
3478
|
-
|
|
3479
|
-
- agent tooling trust boundaries,
|
|
3480
|
-
- privileged service RPCs,
|
|
3481
|
-
- product-specific security risks,
|
|
3482
|
-
- dependency and supply-chain changes,
|
|
3483
|
-
- deployment and infrastructure code,
|
|
3484
|
-
- filesystem and workspace boundaries,
|
|
3485
|
-
- privacy, and
|
|
3486
|
-
- general security review.
|
|
3487
|
-
|
|
3488
|
-
Selected modules may run more than once. Candidates pass through a duplicate
|
|
3489
|
-
gate, then bounded triage. Reviewer or triage failures can produce partial
|
|
3490
|
-
results. A final dedupe failure stops finalization.
|
|
3491
|
-
|
|
3492
|
-
The pipeline writes JSONL journals as work completes. Final artifacts include
|
|
3493
|
-
the review bundle, patch, reviewer outputs, candidates, triage decisions,
|
|
3494
|
-
findings, accounting, and audit events.
|
|
3495
|
-
|
|
3496
|
-
## Separate session storage from review artifacts
|
|
3497
|
-
|
|
3498
|
-
`cursorHostedStorage` keeps Agent SDK session and event records on
|
|
3499
|
-
Cursor-managed hosting. Security Reviewer sets `restore: "off"` so startup
|
|
3500
|
-
doesn't load old review sessions in bulk. A continuation lookup can still
|
|
3501
|
-
fetch a needed session. See [Storage](/docs/storage.md).
|
|
3502
|
-
|
|
3503
|
-
The staged review files are separate from session storage. Session-store
|
|
3504
|
-
durability doesn't preserve those files. All stages for one `runId` must see
|
|
3505
|
-
the same filesystem.
|
|
3506
|
-
|
|
3507
|
-
This split is useful when conversation history needs shared durability but
|
|
3508
|
-
large review artifacts belong on attached storage or an object store.
|
|
3509
|
-
|
|
3510
|
-
## Compare live comment variants
|
|
3511
|
-
|
|
3512
|
-
The comment-severity experiment uses sticky session assignment with a 5%
|
|
3513
|
-
holdout:
|
|
3514
|
-
|
|
3515
|
-
- `control` posts every finding.
|
|
3516
|
-
- `treatment` posts only high and critical findings.
|
|
3517
|
-
|
|
3518
|
-
Finalization enforces the comment filter. The treatment also adds an
|
|
3519
|
-
instruction overlay asking chat and playground summaries to lead with high
|
|
3520
|
-
and critical findings. Full artifacts, `finalResponse`, and finding counts
|
|
3521
|
-
still include every finding. Stage-tool counters appear in the
|
|
3522
|
-
playground A/B view. Local sample and snapshot files persist under
|
|
3523
|
-
the project state directory.
|
|
3524
|
-
|
|
3525
|
-
When a treatment session has only low or medium findings, the filtered review
|
|
3526
|
-
body currently says no vulnerabilities were found even though artifacts and
|
|
3527
|
-
status retain findings. Account for that mismatch before using this
|
|
3528
|
-
experiment as a publishing policy.
|
|
3529
|
-
|
|
3530
|
-
Eval sessions skip A/B enrollment.
|
|
3531
|
-
|
|
3532
|
-
## Test the GitHub channel carefully
|
|
3533
|
-
|
|
3534
|
-
The channel uses the host's Cursor account repository scope. It wakes on
|
|
3535
|
-
`opened` and `synchronize`, skips drafts, and posts its own GitHub Check.
|
|
3536
|
-
|
|
3537
|
-
Inspect its event surface:
|
|
3538
|
-
|
|
3539
|
-
```bash
|
|
3540
|
-
agent-sdk github events \
|
|
3541
|
-
--dir ../../factory/security-reviewer \
|
|
3542
|
-
--json
|
|
3543
|
-
```
|
|
3544
|
-
|
|
3545
|
-
Replay reaches the full publishing path:
|
|
3546
|
-
|
|
3547
|
-
```bash
|
|
3548
|
-
TEST_PR_URL=https://github.com/your-org/allowlisted-test-repo/pull/123
|
|
3549
|
-
agent-sdk github replay \
|
|
3550
|
-
"$TEST_PR_URL" \
|
|
3551
|
-
--dir ../../factory/security-reviewer \
|
|
3552
|
-
--action opened
|
|
3553
|
-
```
|
|
3554
|
-
|
|
3555
|
-
Set `TEST_PR_URL` to a PR in the channel's configured repository allowlist.
|
|
3556
|
-
Run the command only against a PR intended for test reviews. It posts a
|
|
3557
|
-
GitHub Check and may post findings.
|
|
3558
|
-
|
|
3559
|
-
## Inspect the eval before running it
|
|
3560
|
-
|
|
3561
|
-
```bash
|
|
3562
|
-
agent-sdk eval --dir ../../factory/security-reviewer --list
|
|
3563
|
-
```
|
|
3564
|
-
|
|
3565
|
-
The case gates the review flow and prevents comment posting. It still fetches
|
|
3566
|
-
the live PR, so it needs GitHub access.
|
|
3567
|
-
|
|
3568
|
-
## Build another staged pipeline
|
|
3569
|
-
|
|
3570
|
-
Use staged host orchestration when:
|
|
3571
|
-
|
|
3572
|
-
- each phase needs its own timeout and artifact,
|
|
3573
|
-
- model workers should run in bounded parallel,
|
|
3574
|
-
- later work can start as soon as partial results arrive,
|
|
3575
|
-
- a webhook must acknowledge before the work finishes, or
|
|
3576
|
-
- operators need live progress beyond one tool spinner.
|
|
3577
|
-
|
|
3578
|
-
Keep external writes in finalization. Pass a `runId` between stages, journal
|
|
3579
|
-
progress before publishing, and make partial-worker failures visible in the
|
|
3580
|
-
result.
|
|
3581
|
-
|
|
3582
|
-
## Where to go next
|
|
3583
|
-
|
|
3584
|
-
- [GitHub](/docs/guides/github.md)
|
|
3585
|
-
- [Tools](/docs/reference/tools.md)
|
|
3586
|
-
- [Channels](/docs/reference/channels.md)
|
|
3587
|
-
- [Playground](/docs/reference/playground.md)
|
|
3588
|
-
- [Storage](/docs/storage.md)
|
|
3589
|
-
- [Live A/B metrics](/docs/ab.md)
|
|
3590
|
-
- [Evals](/docs/evals.md)
|
|
3591
|
-
|
|
3592
|
-
---
|
|
3593
|
-
|
|
3594
|
-
Source: /docs/example-agents/slack-agent.md
|
|
3595
|
-
|
|
3596
|
-
# Put a minimal agent in Slack
|
|
3597
|
-
|
|
3598
|
-
Slack agent is the smallest channel example. It has one runtime config, one
|
|
3599
|
-
instruction file, and one authored channel. A teammate mentions the agent,
|
|
3600
|
-
the local runtime harness runs a turn, and the answer
|
|
3601
|
-
returns to the same Slack thread.
|
|
3602
|
-
|
|
3603
|
-
Use it to learn the minimum needed for a Slack agent before adding tools,
|
|
3604
|
-
workflows, or a dedicated app.
|
|
3605
|
-
|
|
3606
|
-
[Browse the Slack agent source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/slack-agent/)
|
|
3607
|
-
|
|
3608
|
-
## Keep the Slack channel small
|
|
3609
|
-
|
|
3610
|
-
Slack agent delegates transport details to the host connection. The authored
|
|
3611
|
-
file selects the account-linked transport, gives the agent a single-token
|
|
3612
|
-
router name and icon, and supplies suggested prompts.
|
|
3613
|
-
|
|
3614
|
-
The complete channel lives in
|
|
3615
|
-
[`agent/channels/slack.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/slack-agent/agent/channels/slack.ts).
|
|
3616
|
-
The framework supplies message intake, thread-scoped sessions, delivery,
|
|
3617
|
-
status updates, and suggested prompts.
|
|
3618
|
-
|
|
3619
|
-
## Follow a Slack message
|
|
3620
|
-
|
|
3621
|
-
1. A user mentions the agent or sends the host app a direct message naming
|
|
3622
|
-
it.
|
|
3623
|
-
2. The Slack relay selects this channel by its single-token `agentName`.
|
|
3624
|
-
3. The Agent SDK maps the Slack channel and thread timestamp to a continuation
|
|
3625
|
-
key.
|
|
3626
|
-
4. The local harness runs with
|
|
3627
|
-
[`instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/slack-agent/agent/instructions.md).
|
|
3628
|
-
5. The response returns to the triggering thread.
|
|
3629
|
-
6. A later message in the same thread resumes the durable session.
|
|
3630
|
-
|
|
3631
|
-
The prompt asks for concise threaded replies. It doesn't define domain policy
|
|
3632
|
-
or tool routing.
|
|
3633
|
-
|
|
3634
|
-
## Map the Slack agent files
|
|
3635
|
-
|
|
3636
|
-
| File | Purpose |
|
|
3637
|
-
| --- | --- |
|
|
3638
|
-
| [`package.json`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/slack-agent/package.json) | Declares the example package and Agent SDK dependency. |
|
|
3639
|
-
| [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/slack-agent/agent/agent.ts) | Names the agent and selects the model. The omitted `runtime` defaults to local. |
|
|
3640
|
-
| [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/slack-agent/agent/instructions.md) | Sets the always-on response style. |
|
|
3641
|
-
| [`agent/channels/slack.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/slack-agent/agent/channels/slack.ts) | Connects the signed-in host account to Slack. |
|
|
3642
|
-
| [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/slack-agent/agent/storage.ts) | Persists sessions and events with `cursorHostedStorage`. |
|
|
3643
|
-
|
|
3644
|
-
There are no authored tools, skills, MCP connections, subagents, schedules,
|
|
3645
|
-
hooks, A/B experiments, or evals. This small surface is the lesson.
|
|
3646
|
-
|
|
3647
|
-
## Connect the host
|
|
3648
|
-
|
|
3649
|
-
You need:
|
|
3650
|
-
|
|
3651
|
-
- Node 22.13 or newer.
|
|
3652
|
-
- An agent-runtime credential.
|
|
3653
|
-
- Slack connected through the selected channel transport.
|
|
3654
|
-
|
|
3655
|
-
Sign in and confirm the active account:
|
|
3656
|
-
|
|
3657
|
-
```bash
|
|
3658
|
-
agent-sdk login
|
|
3659
|
-
agent-sdk whoami
|
|
3660
|
-
```
|
|
3661
|
-
|
|
3662
|
-
The selected transport owns Slack credential setup. See the
|
|
3663
|
-
[Slack guide](/docs/guides/slack.md) for account-linked and dedicated-app
|
|
3664
|
-
options.
|
|
3665
|
-
|
|
3666
|
-
## Validate and start the server
|
|
3667
|
-
|
|
3668
|
-
```bash
|
|
3669
|
-
agent-sdk validate --dir examples/slack-agent
|
|
3670
|
-
agent-sdk info --dir examples/slack-agent --json
|
|
3671
|
-
agent-sdk dev examples/slack-agent
|
|
3672
|
-
```
|
|
3673
|
-
|
|
3674
|
-
The dev command prints the playground URL. It also mounts the Slack channel
|
|
3675
|
-
and waits for relayed messages.
|
|
3676
|
-
|
|
3677
|
-
In Slack, address the configured host app and router name, then send:
|
|
3678
|
-
|
|
3679
|
-
> `<host-app mention> <router name>` Explain the Agent SDK in three bullets.
|
|
3680
|
-
|
|
3681
|
-
Reply in the generated thread:
|
|
3682
|
-
|
|
3683
|
-
> Make the second bullet simpler.
|
|
3684
|
-
|
|
3685
|
-
The second message reaches the same session. You can open that session in the
|
|
3686
|
-
playground to inspect the received message, model events, final reply, and
|
|
3687
|
-
usage.
|
|
3688
|
-
|
|
3689
|
-
## Test without Slack
|
|
3690
|
-
|
|
3691
|
-
Every project gets the built-in HTTP channel even when no HTTP file exists.
|
|
3692
|
-
Run a one-shot turn through it:
|
|
3693
|
-
|
|
3694
|
-
```bash
|
|
3695
|
-
agent-sdk run --dir examples/slack-agent \
|
|
3696
|
-
--message "Explain the Agent SDK simply."
|
|
3697
|
-
```
|
|
3698
|
-
|
|
3699
|
-
The same project also exposes an MCP endpoint. Since this agent has no server
|
|
3700
|
-
tools, its MCP surface contains `ask` and `check`, but not `call_tool`.
|
|
3701
|
-
|
|
3702
|
-
These automatic surfaces let you test the prompt from the CLI and let another
|
|
3703
|
-
agent delegate to it later. The authored Slack channel only changes how work
|
|
3704
|
-
arrives and where replies go.
|
|
3705
|
-
|
|
3706
|
-
## Know when to add a dedicated app
|
|
3707
|
-
|
|
3708
|
-
An account-linked Slack transport is a fit for mentions, direct messages, thread
|
|
3709
|
-
continuity, and agent-branded replies. Move to a dedicated Socket Mode channel
|
|
3710
|
-
when you need:
|
|
3711
|
-
|
|
3712
|
-
- top-level channel watching,
|
|
3713
|
-
- interactive approval buttons,
|
|
3714
|
-
- a separate bot identity, or
|
|
3715
|
-
- Slack app events unsupported by the account-linked relay.
|
|
3716
|
-
|
|
3717
|
-
Compare this example with [Playbook router](/docs/example-agents/benny.md), which adds allowlisted channel
|
|
3718
|
-
watching, and [Weather agent](/docs/example-agents/weather-agent.md), which adds approval buttons
|
|
3719
|
-
through a second Slack channel.
|
|
3720
|
-
|
|
3721
|
-
## Turn the channel into your own Slack agent
|
|
3722
|
-
|
|
3723
|
-
Copy the three authored files, then change:
|
|
3724
|
-
|
|
3725
|
-
- `name` in `agent.ts` for the harness identity,
|
|
3726
|
-
- `agentName` in `slack.ts` for the single-token router name,
|
|
3727
|
-
- the instructions for your domain, and
|
|
3728
|
-
- suggested prompts for the tasks teammates should try.
|
|
3729
|
-
|
|
3730
|
-
Keep `agentName` free of whitespace. Use PascalCase for multiword names.
|
|
3731
|
-
|
|
3732
|
-
## Where to go next
|
|
3733
|
-
|
|
3734
|
-
- [Slack](/docs/guides/slack.md)
|
|
3735
|
-
- [Channels](/docs/reference/channels.md)
|
|
3736
|
-
- [Sessions and streaming](/docs/reference/sessions.md)
|
|
3737
|
-
- [Playground](/docs/reference/playground.md)
|
|
3738
|
-
|
|
3739
|
-
---
|
|
3740
|
-
|
|
3741
|
-
Source: /docs/example-agents/weather-agent.md
|
|
3742
|
-
|
|
3743
|
-
# Explore the full Agent SDK surface with a weather agent
|
|
3744
|
-
|
|
3745
|
-
The weather agent is the broadest small example in the repository. It fetches
|
|
3746
|
-
live conditions and forecasts, converts units through MCP, writes notes in a
|
|
3747
|
-
session workspace, and runs from HTTP, Slack, a schedule, and the MCP endpoint.
|
|
3748
|
-
|
|
3749
|
-
Use this project when you want to see how the Agent SDK's filesystem pieces fit
|
|
3750
|
-
together before you design a larger agent.
|
|
3751
|
-
|
|
3752
|
-
[Browse the weather agent source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/weather-agent/)
|
|
3753
|
-
|
|
3754
|
-
## See the runtime features together
|
|
3755
|
-
|
|
3756
|
-
Most examples focus on one feature. Weather agent puts the major runtime
|
|
3757
|
-
features side by side:
|
|
3758
|
-
|
|
3759
|
-
| Capability | Source | Role |
|
|
3760
|
-
| --- | --- | --- |
|
|
3761
|
-
| Root config and instructions | [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/agent.ts), [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/instructions.md) | Select the cloud runtime and route each request. |
|
|
3762
|
-
| Server tools | [`agent/tools/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/weather-agent/agent/tools/) | Fetch Open-Meteo data and call MCP from the serve host. |
|
|
3763
|
-
| Agent tool | [`save_weather_note.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/tools/save_weather_note.ts) | Run a Python script inside the session workspace. |
|
|
3764
|
-
| Stdio MCP | [`units.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/mcp-connections/units.ts), [`probe.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/mcp-connections/probe.ts) | Expose conversion tools to the model, host tools, and channel handlers. Author VM-side probe tools as TypeScript `execute` functions. |
|
|
3765
|
-
| Custom HTTP | [`webhook.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/channels/webhook.ts) | Start a turn or call MCP without a model turn. |
|
|
3766
|
-
| Slack | [`slack.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/channels/slack.ts), [`slack-app.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/channels/slack-app.ts) | Compare account-linked chat with a dedicated app. |
|
|
3767
|
-
| Skill and subagent | [`forecast.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/skills/forecast.md), [`researcher/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/weather-agent/agent/subagents/researcher/) | Load a procedure on demand or delegate broad research. |
|
|
3768
|
-
| Schedule and hooks | [`heartbeat.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/schedules/heartbeat.md), [`audit.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/hooks/audit.ts), [`journal.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/hooks/journal.ts) | Start recurring tasks, log usage, and save turn summaries. |
|
|
3769
|
-
| A/B and evals | [`agent/ab.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/ab.ts), [`evals/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/weather-agent/evals/) | Compare a sticky variant and protect tool routing with regression cases. |
|
|
3770
|
-
|
|
3771
|
-
## Follow one request
|
|
3772
|
-
|
|
3773
|
-
A current-weather question takes this path:
|
|
3774
|
-
|
|
3775
|
-
1. The built-in HTTP channel, Slack, or the custom `/report` route creates a
|
|
3776
|
-
durable session.
|
|
3777
|
-
2. `instructions.md` tells the model to call `get_weather` instead of
|
|
3778
|
-
guessing.
|
|
3779
|
-
3. The server tool geocodes the city, fetches Open-Meteo, validates the
|
|
3780
|
-
response, and returns normalized fields.
|
|
3781
|
-
4. The agent writes a short answer. The Agent SDK records every event in the
|
|
3782
|
-
session stream.
|
|
3783
|
-
5. The audit hook observes `turn.completed`. If the session joined the A/B
|
|
3784
|
-
experiment, the collector updates its metrics too.
|
|
3785
|
-
|
|
3786
|
-
Forecasts route to `get_forecast`. Unit conversions route to
|
|
3787
|
-
`convert_temperature`, which calls the `units` MCP server through
|
|
3788
|
-
`ctx.host.mcp`. Climate history and broad comparisons route to the
|
|
3789
|
-
`researcher` subagent.
|
|
3790
|
-
|
|
3791
|
-
## Prepare the example
|
|
3792
|
-
|
|
3793
|
-
You need:
|
|
3794
|
-
|
|
3795
|
-
- Node 22.13 or newer.
|
|
3796
|
-
- An agent-runtime credential.
|
|
3797
|
-
- Network access to Open-Meteo.
|
|
3798
|
-
- Python 3 for `save_weather_note`.
|
|
3799
|
-
|
|
3800
|
-
The project mounts an account-linked Slack channel. The Agent SDK checks the
|
|
3801
|
-
connection at startup, so sign in even when you plan to call a deterministic
|
|
3802
|
-
tool.
|
|
3803
|
-
|
|
3804
|
-
The optional dedicated Slack app also needs a token pair.
|
|
3805
|
-
`agent-sdk slack create --dir examples/weather-agent` provisions the
|
|
3806
|
-
app and writes the tokens for you (see the
|
|
3807
|
-
[Slack guide](/docs/guides/slack.md#set-it-up)); with hand-minted tokens,
|
|
3808
|
-
export them instead:
|
|
3809
|
-
|
|
3810
|
-
```bash
|
|
3811
|
-
export WEATHER_AGENT_SLACK_BOT_TOKEN=xoxb-...
|
|
3812
|
-
export WEATHER_AGENT_SLACK_APP_TOKEN=xapp-...
|
|
3813
|
-
agent-sdk slack doctor --prefix WEATHER_AGENT
|
|
3814
|
-
```
|
|
1690
|
+
| `t.score(name, value)` | records a 0–1 score you computed; soft until you add a bar |
|
|
1691
|
+
| `t.requireToolCall(name, matcher?)` | gates on a matching call and returns it, so later code can read its input and output |
|
|
1692
|
+
| `t.requireInputRequest(filter?)` | gates on exactly one pending approval request and returns it |
|
|
3815
1693
|
|
|
3816
|
-
|
|
3817
|
-
|
|
1694
|
+
Every gate returns a handle: `.soft()` demotes it to tracked-only,
|
|
1695
|
+
`.atLeast(0.7)` adds a soft score bar, and `.gate(0.8)` promotes a
|
|
1696
|
+
scored assertion into a hard gate.
|
|
3818
1697
|
|
|
3819
|
-
|
|
1698
|
+
With no matcher, `calledTool` is request-based: a requested call counts
|
|
1699
|
+
even when its result has not arrived. Pass
|
|
1700
|
+
`t.calledTool("inspect_pr", { status: "completed" })` to require the
|
|
1701
|
+
call to return. `input`, `output`, and `count` matcher fields accept a
|
|
1702
|
+
literal, a `RegExp`, or a predicate.
|
|
3820
1703
|
|
|
3821
|
-
|
|
3822
|
-
|
|
3823
|
-
|
|
3824
|
-
|
|
3825
|
-
|
|
1704
|
+
The expectation builders are `includes(string | RegExp)`,
|
|
1705
|
+
`equals(value)`, `matches(schema)`, `similarity(expected)`, and
|
|
1706
|
+
`satisfies(predicate, label)`. `includes` stringifies its input,
|
|
1707
|
+
`equals` compares values deeply, `matches` validates against a Standard
|
|
1708
|
+
Schema (or anything with `safeParse`, like Zod), `similarity` scores
|
|
1709
|
+
normalized text similarity, and `satisfies` runs your predicate. The
|
|
1710
|
+
plain function `normalizedSimilarity(actual, expected)` returns the
|
|
1711
|
+
same 0–1 score for use with `t.score`.
|
|
3826
1712
|
|
|
3827
|
-
|
|
3828
|
-
|
|
1713
|
+
A few more context members shape a case: `t.require(value, expectation)`
|
|
1714
|
+
records a gate and stops the test body when it fails, without a
|
|
1715
|
+
duplicate execution error. `t.skip(reason)` ends the case as skipped
|
|
1716
|
+
(reported separately, never changes the exit code; call it before
|
|
1717
|
+
sending messages). `t.metric(name, value)` records a structured score
|
|
1718
|
+
for the playground case card. `t.log(message)` records a debug line for
|
|
1719
|
+
the CLI and playground result.
|
|
3829
1720
|
|
|
3830
|
-
|
|
1721
|
+
Three `t.send` options apply on session create (first `t.send` only):
|
|
3831
1722
|
|
|
3832
|
-
|
|
1723
|
+
- `workspaceFiles`: `{ path: contents }`, seeded into the local session
|
|
1724
|
+
workspace. Prefer this over machine-local paths.
|
|
1725
|
+
- `workspaceDir`: absolute harness cwd (local runtime).
|
|
1726
|
+
- `cloud`: per-session cloud options merged over the agent's static
|
|
1727
|
+
`cloud` config (repos / env / …). Use a pinned `repos` override to
|
|
1728
|
+
attach a fixture repo for cloud evals without putting it on the
|
|
1729
|
+
agent's default `cloud.repos`. Cloud ignores `workspaceFiles` seeds.
|
|
3833
1730
|
|
|
3834
|
-
```
|
|
3835
|
-
|
|
3836
|
-
|
|
3837
|
-
|
|
1731
|
+
```ts
|
|
1732
|
+
const toolResults = t.events.filter((e) => e.type === "action.result");
|
|
1733
|
+
t.check(
|
|
1734
|
+
toolResults.length,
|
|
1735
|
+
satisfies((n) => (n as number) <= 4, "at most 4 tool calls")
|
|
1736
|
+
);
|
|
3838
1737
|
```
|
|
3839
1738
|
|
|
3840
|
-
|
|
3841
|
-
|
|
3842
|
-
|
|
1739
|
+
A case with no explicit gates falls back to whether at least one turn
|
|
1740
|
+
completed successfully. Add `t.succeeded()` and behavior-specific gates
|
|
1741
|
+
anyway. They make the contract visible during review.
|
|
3843
1742
|
|
|
3844
|
-
|
|
1743
|
+
### Judge free-form output
|
|
3845
1744
|
|
|
3846
|
-
|
|
3847
|
-
|
|
3848
|
-
|
|
3849
|
-
|
|
1745
|
+
When wording matters and no regex captures it, `t.judge` grades the
|
|
1746
|
+
reply with an LLM. The built-in graders are `factuality(expected)`,
|
|
1747
|
+
`summarizes(expected)`, `closedQA(criteria)`, and `sql(expected)`. Each
|
|
1748
|
+
scores `t.reply` by default; pass `{ on }` to grade another value.
|
|
1749
|
+
|
|
1750
|
+
```ts
|
|
1751
|
+
t.judge.factuality("It is 54°F in NYC right now.").atLeast(0.7);
|
|
3850
1752
|
```
|
|
3851
1753
|
|
|
3852
|
-
|
|
3853
|
-
|
|
1754
|
+
Judge assertions are soft by default, so a judge never fails a build
|
|
1755
|
+
until you give it a bar with `.atLeast(0.7)` or promote it with
|
|
1756
|
+
`.gate(0.8)`. The judge model comes from `defineEvalConfig({ judge })`,
|
|
1757
|
+
`defineEval({ judge })`, a case-level `judge`, or a per-call
|
|
1758
|
+
`{ model }` override; the nearest one wins. For a domain-specific judge
|
|
1759
|
+
whose verdict is not a single score, `t.judge.model(prompt)` sends a
|
|
1760
|
+
raw prompt to the same model and returns the reply. You then record the
|
|
1761
|
+
parsed result with `t.score` or `t.check`.
|
|
3854
1762
|
|
|
3855
|
-
##
|
|
1763
|
+
## Run evals from the CLI
|
|
3856
1764
|
|
|
3857
|
-
|
|
3858
|
-
runs inside the serve host and can reach `ctx.host` services.
|
|
1765
|
+
The `eval` command discovers, filters, and runs cases.
|
|
3859
1766
|
|
|
3860
|
-
|
|
3861
|
-
|
|
3862
|
-
appends to `weather-notes.md` in that session's workspace:
|
|
1767
|
+
Run the CLI under Node 22.13 or newer. Do not use Bun. Its HTTP/2 client
|
|
1768
|
+
breaks tool-result streams and causes eval turns to fail.
|
|
3863
1769
|
|
|
3864
1770
|
```bash
|
|
3865
|
-
agent-sdk
|
|
3866
|
-
|
|
1771
|
+
agent-sdk eval --dir . --list # discover only
|
|
1772
|
+
agent-sdk eval --dir . # run all
|
|
1773
|
+
agent-sdk eval --dir . builds/checkout # one datapoint
|
|
1774
|
+
agent-sdk eval --dir . builds search # several ids or prefixes
|
|
1775
|
+
agent-sdk eval --dir . --tag smoke --tag pull-request # any matching tag
|
|
1776
|
+
agent-sdk eval --dir . --json --no-stream # machine-readable results
|
|
1777
|
+
agent-sdk eval --dir . --verbose # logs + reply snippets
|
|
3867
1778
|
```
|
|
3868
1779
|
|
|
3869
|
-
|
|
3870
|
-
example
|
|
3871
|
-
|
|
3872
|
-
|
|
3873
|
-
|
|
3874
|
-
|
|
3875
|
-
## Verify tool execution on the agent VM
|
|
3876
|
-
|
|
3877
|
-
Ask the agent to call `probe_cloud_tool` on the `probe` MCP server.
|
|
3878
|
-
`agent/mcp-connections/probe.ts` authors that tool as TypeScript. The Agent
|
|
3879
|
-
SDK packages it as stdio MCP so a cloud VM with no checkout of this example
|
|
3880
|
-
can still run it. The model lists the server and calls the tool; it does
|
|
3881
|
-
not write a `.sh`.
|
|
3882
|
-
|
|
3883
|
-
A real call writes `vm-tool-observations/<id>.json` in the agent cwd and
|
|
3884
|
-
returns hostname, cwd, and pid. Stream events show `probe:probe_cloud_tool`,
|
|
3885
|
-
not `shell`.
|
|
1780
|
+
Id filters use OR semantics. Each filter selects an exact id and its
|
|
1781
|
+
descendants. For example, `builds` selects `builds`,
|
|
1782
|
+
`builds/checkout`, and every other case below that path. Repeated tags
|
|
1783
|
+
also use OR semantics. When you provide both ids and tags, a case must
|
|
1784
|
+
match both groups.
|
|
3886
1785
|
|
|
3887
|
-
|
|
3888
|
-
|
|
1786
|
+
`eval` boots an ephemeral server on port 0 with a temp state root
|
|
1787
|
+
outside the project, so cases don't inherit ambient monorepo rules and
|
|
1788
|
+
don't write into the project state directory. Point `--url` at a running server to eval
|
|
1789
|
+
a live agent instead:
|
|
3889
1790
|
|
|
3890
1791
|
```bash
|
|
3891
|
-
agent-sdk
|
|
3892
|
-
--
|
|
1792
|
+
agent-sdk eval --dir . \
|
|
1793
|
+
--url http://127.0.0.1:3000/weather-agent \
|
|
1794
|
+
--bearer-token "$AGENT_TOKEN"
|
|
3893
1795
|
```
|
|
3894
1796
|
|
|
3895
|
-
|
|
1797
|
+
The eval definitions still come from `--dir`; `--url` only changes the
|
|
1798
|
+
agent that receives the turns. For a locally mounted multi-agent
|
|
1799
|
+
directory, `--slug weather-agent` chooses the target. Use
|
|
1800
|
+
`--state-root` to keep ephemeral session state at a chosen path,
|
|
1801
|
+
`--timeout-ms` to override the project timeout, and `--no-stream` to
|
|
1802
|
+
keep live progress off stderr. A TTY streams turn progress by default.
|
|
1803
|
+
`--verbose` still writes `t.log` lines to stderr and adds reply snippets
|
|
1804
|
+
to text results.
|
|
3896
1805
|
|
|
3897
|
-
|
|
3898
|
-
|
|
1806
|
+
Model turns need a Cursor credential from `agent-sdk login` or
|
|
1807
|
+
`CURSOR_API_KEY`.
|
|
3899
1808
|
|
|
3900
|
-
|
|
3901
|
-
- server tools through `ctx.host.mcp`, and
|
|
3902
|
-
- channel handlers through `host.mcp`.
|
|
1809
|
+
See [CLI: eval](/docs/reference/cli.md#eval) for flags and exit codes.
|
|
3903
1810
|
|
|
3904
|
-
|
|
3905
|
-
functions; the Agent SDK packages them so a cloud VM can spawn the server
|
|
3906
|
-
without this checkout. The model calls `probe_cloud_tool` directly; no host
|
|
3907
|
-
tool wraps it.
|
|
1811
|
+
### JSON results
|
|
3908
1812
|
|
|
3909
|
-
`
|
|
1813
|
+
Use `--json --no-stream` in scripts and CI. The top-level result carries
|
|
1814
|
+
the totals and one result per case:
|
|
3910
1815
|
|
|
3911
|
-
```
|
|
3912
|
-
|
|
3913
|
-
|
|
3914
|
-
|
|
1816
|
+
```json
|
|
1817
|
+
{
|
|
1818
|
+
"ok": true,
|
|
1819
|
+
"passed": 1,
|
|
1820
|
+
"failed": 0,
|
|
1821
|
+
"results": [
|
|
1822
|
+
{
|
|
1823
|
+
"id": "readiness",
|
|
1824
|
+
"ok": true,
|
|
1825
|
+
"assertions": [{ "name": "succeeded", "passed": true }],
|
|
1826
|
+
"sessionId": "ses_123",
|
|
1827
|
+
"inputs": ["Is checkout pull request 42 ready to approve?"],
|
|
1828
|
+
"toolCalls": [{ "toolName": "inspect_pr", "isError": false }],
|
|
1829
|
+
"logs": [],
|
|
1830
|
+
"durationMs": 12340
|
|
1831
|
+
}
|
|
1832
|
+
]
|
|
1833
|
+
}
|
|
3915
1834
|
```
|
|
3916
1835
|
|
|
3917
|
-
|
|
1836
|
+
Each case result can also include `description`, `finalText`, `tools`,
|
|
1837
|
+
`error`, and tool arguments or output. This shape lets CI report the
|
|
1838
|
+
failed assertion without parsing terminal text.
|
|
3918
1839
|
|
|
3919
|
-
|
|
3920
|
-
agent-sdk dev examples/weather-agent
|
|
3921
|
-
```
|
|
1840
|
+
## Run evals in the playground
|
|
3922
1841
|
|
|
3923
|
-
|
|
1842
|
+
Start the server, open the playground, and choose **Evals**. You can run
|
|
1843
|
+
every case or one case, watch progress, and open the resulting session
|
|
1844
|
+
trace. The Evals tab works on a normal `serve`.
|
|
3924
1845
|
|
|
3925
1846
|
```bash
|
|
3926
|
-
|
|
3927
|
-
http://127.0.0.1:3000/weather-agent/v1/channels/webhook/convert \
|
|
3928
|
-
-H 'content-type: application/json' \
|
|
3929
|
-
-d '{"value":20,"from":"C"}'
|
|
1847
|
+
agent-sdk serve --dir .
|
|
3930
1848
|
```
|
|
3931
1849
|
|
|
3932
|
-
|
|
3933
|
-
|
|
3934
|
-
|
|
3935
|
-
|
|
1850
|
+
Playground runs target the live server instead of an ephemeral one.
|
|
1851
|
+
Their sessions appear in the session list. One eval batch can run at a
|
|
1852
|
+
time. Persistence follows the rule under
|
|
1853
|
+
[Configure eval runs](#configure-eval-runs). See
|
|
1854
|
+
[Playground eval routes](/docs/reference/http-api.md#playground-eval-routes).
|
|
1855
|
+
The start request returns `202` while cases run in the background.
|
|
1856
|
+
Poll until the snapshot status becomes `completed`, `failed`, or `cancelled`.
|
|
1857
|
+
Configuration errors appear on a failed snapshot.
|
|
3936
1858
|
|
|
3937
|
-
`
|
|
1859
|
+
On `--prod` / `--url`, the CLI prints the Eval ID as soon as the batch is
|
|
1860
|
+
accepted (and a Playground deep link with `?view=evals&evalRunId=…`):
|
|
3938
1861
|
|
|
3939
1862
|
```bash
|
|
3940
|
-
|
|
3941
|
-
|
|
3942
|
-
-
|
|
3943
|
-
|
|
3944
|
-
```
|
|
3945
|
-
|
|
3946
|
-
The response includes a `key`. Send it back on the next request to continue
|
|
3947
|
-
the same session:
|
|
1863
|
+
agent-sdk eval --prod --slug vulnerability-scanner --tag deepsec
|
|
1864
|
+
# Eval ID: evalrun_…
|
|
1865
|
+
# Cancel: agent-sdk eval cancel evalrun_… --prod --slug vulnerability-scanner
|
|
1866
|
+
# Playground: https://…/playground?view=evals&evalRunId=evalrun_…
|
|
3948
1867
|
|
|
3949
|
-
|
|
3950
|
-
|
|
3951
|
-
http://127.0.0.1:3000/weather-agent/v1/channels/webhook/report \
|
|
3952
|
-
-H 'content-type: application/json' \
|
|
3953
|
-
-d '{"message":"How about tomorrow?","key":"<key>"}'
|
|
1868
|
+
agent-sdk eval cancel evalrun_… --prod --slug vulnerability-scanner
|
|
1869
|
+
agent-sdk eval status evalrun_… --prod --slug vulnerability-scanner
|
|
3954
1870
|
```
|
|
3955
1871
|
|
|
3956
|
-
|
|
3957
|
-
[webhooks and custom channels](/docs/guides/webhooks.md) for route schemas,
|
|
3958
|
-
authentication, and asynchronous handlers.
|
|
3959
|
-
|
|
3960
|
-
## Load procedures and delegate research
|
|
1872
|
+
## What good cases assert
|
|
3961
1873
|
|
|
3962
|
-
|
|
3963
|
-
|
|
3964
|
-
|
|
1874
|
+
Gate decisions and shape, not prose. Model wording varies run to run.
|
|
1875
|
+
Tool choice, tool avoidance, and output structure are the stable
|
|
1876
|
+
contract.
|
|
3965
1877
|
|
|
3966
|
-
|
|
3967
|
-
|
|
3968
|
-
|
|
1878
|
+
1. `t.succeeded()`: always, first.
|
|
1879
|
+
2. The tool decision: `calledTool` for the intended path,
|
|
1880
|
+
`notCalledTool` for the likely wrong alternative. The pair is
|
|
1881
|
+
stronger than either alone.
|
|
1882
|
+
3. Output shape: a regex for the contract (`/ready|blocked/i`, a JSON
|
|
1883
|
+
marker, a findings-block fence), never exact sentences.
|
|
1884
|
+
4. For structured output, parse `t.reply` and check fields with
|
|
1885
|
+
`satisfies` instead of substring-matching JSON.
|
|
3969
1886
|
|
|
3970
|
-
|
|
3971
|
-
|
|
3972
|
-
|
|
3973
|
-
```
|
|
1887
|
+
The common failure modes: asserting exact phrasing, packing more than
|
|
1888
|
+
about five gates into one case (split it), and cases that depend on live
|
|
1889
|
+
external state that drifts (pin the input; see fixtures).
|
|
3974
1890
|
|
|
3975
|
-
|
|
3976
|
-
parent should hand a bounded task to a specialist. The
|
|
3977
|
-
[subagents reference](/docs/reference/subagents.md) explains the current
|
|
3978
|
-
inheritance limits.
|
|
1891
|
+
## Pick fixtures by agent type
|
|
3979
1892
|
|
|
3980
|
-
|
|
1893
|
+
The right fixture depends on the surface under test.
|
|
3981
1894
|
|
|
3982
|
-
|
|
3983
|
-
|
|
1895
|
+
| Agent surface | Fixture |
|
|
1896
|
+
| --- | --- |
|
|
1897
|
+
| Chat / domain assistant | A canonical prompt string, chosen once and frozen |
|
|
1898
|
+
| Tool-heavy | Run `agent-sdk call <tool>` first to pin what the tool returns, then freeze the prompt that triggers it |
|
|
1899
|
+
| GitHub webhook | `agent-sdk github replay <pr> --events '*' --dry-run --out fixtures/github` snapshots real payloads for offline replay ([GitHub guide](/docs/guides/github.md)) |
|
|
1900
|
+
| PR reviewer with host preparation | Diff, metadata, and gold labels pinned to commit SHAs; keep any live PR matrix small |
|
|
1901
|
+
| Workspace-dependent | `workspaceFiles` in `t.send` options, never developer-machine paths |
|
|
3984
1902
|
|
|
3985
|
-
|
|
3986
|
-
|
|
3987
|
-
http://127.0.0.1:3000/weather-agent/v1/dev/schedules/heartbeat
|
|
3988
|
-
```
|
|
1903
|
+
Tag the fast, reliably passing core `smoke` and run `--tag smoke` in the
|
|
1904
|
+
inner loop. Leave slow or flaky-prone cases untagged for explicit runs.
|
|
3989
1905
|
|
|
3990
|
-
|
|
3991
|
-
The audit hook logs usage after each completed turn. Hooks observe recorded
|
|
3992
|
-
events; their failures don't fail the turn.
|
|
1906
|
+
### Materialize API-backed fixtures
|
|
3993
1907
|
|
|
3994
|
-
|
|
1908
|
+
An input that only points at external data, such as a pull request URL,
|
|
1909
|
+
snapshot id, or pair of commit SHAs, is not self-contained. Fetch it
|
|
1910
|
+
once and commit the rendered fixture before you expand the suite.
|
|
3995
1911
|
|
|
3996
|
-
|
|
3997
|
-
|
|
1912
|
+
1. Save the diff, metadata, and labels under `fixtures/` at pinned
|
|
1913
|
+
revisions.
|
|
1914
|
+
2. Seed those files with `workspaceFiles`, or read them from the fixture
|
|
1915
|
+
directory.
|
|
1916
|
+
3. Assert decisions and output shape against the saved evidence.
|
|
1917
|
+
4. Keep a small `smoke` subset for any remaining live pipeline checks.
|
|
3998
1918
|
|
|
3999
|
-
|
|
4000
|
-
|
|
4001
|
-
|
|
1919
|
+
Read committed fixtures with `@cursor/july/evals/loaders`: `loadJson`,
|
|
1920
|
+
`loadJsonl`, and `loadYaml` resolve relative paths against the project
|
|
1921
|
+
root the runner discovered, not the cwd the CLI was invoked from
|
|
1922
|
+
(`resolveFixturePath` and `evalFixtureRoot` expose the same
|
|
1923
|
+
resolution for other file formats).
|
|
4002
1924
|
|
|
4003
|
-
|
|
4004
|
-
|
|
4005
|
-
|
|
4006
|
-
|
|
4007
|
-
policy.
|
|
1925
|
+
`maxConcurrency` limits parallel datapoints. It does not limit model or
|
|
1926
|
+
API fan-out inside one datapoint. Materialized fixtures prevent a large
|
|
1927
|
+
suite from exhausting provider and GitHub rate limits. The
|
|
1928
|
+
[evals skill](https://github.com/cursor/cursor/blob/main/packages/agent-serve/skills/evals/SKILL.md) has the full fixture workflow.
|
|
4008
1929
|
|
|
4009
|
-
|
|
1930
|
+
## Keep improvements with regression evals
|
|
4010
1931
|
|
|
4011
|
-
|
|
4012
|
-
|
|
4013
|
-
|
|
4014
|
-
|
|
1932
|
+
Every [hillclimb](/docs/hillclimbing.md) round that keeps a change must land
|
|
1933
|
+
an eval that would have failed before the change. If you can't express
|
|
1934
|
+
the improvement as a gate (a `calledTool` shift, a bounded
|
|
1935
|
+
`action.result` count, an output-shape regex), the improvement is
|
|
1936
|
+
unverified, and it'll regress silently.
|
|
4015
1937
|
|
|
4016
|
-
The
|
|
4017
|
-
|
|
1938
|
+
The rule cuts the other way too: never weaken an existing gate to make a
|
|
1939
|
+
round pass. That's the freeze line moving, and it turns your regression
|
|
1940
|
+
suite into a list of checks that no longer protect anything.
|
|
4018
1941
|
|
|
4019
|
-
##
|
|
1942
|
+
## Compare variants on live traffic
|
|
4020
1943
|
|
|
4021
|
-
|
|
1944
|
+
Use `defineAB` to compare variant metrics on live sessions. It is not a
|
|
1945
|
+
test runner and has no `agent-sdk ab` command. Keep `defineEval` as the
|
|
1946
|
+
regression ratchet. Eval sessions do not enroll or change live metrics.
|
|
1947
|
+
See [Live A/B metrics](/docs/ab.md) for assignment, behavior, collection,
|
|
1948
|
+
and inspection.
|
|
4022
1949
|
|
|
4023
|
-
|
|
4024
|
-
- Keep deterministic transforms behind direct server tools or MCP.
|
|
4025
|
-
- Use an agent tool only when code must run in the agent workspace.
|
|
4026
|
-
- Put reusable procedures in skills and narrow specialist work into
|
|
4027
|
-
subagents.
|
|
4028
|
-
- Add a channel only when the external surface needs its own identity,
|
|
4029
|
-
continuation key, or delivery behavior.
|
|
1950
|
+
## What's next
|
|
4030
1951
|
|
|
4031
|
-
|
|
1952
|
+
Continue with these pages:
|
|
4032
1953
|
|
|
4033
|
-
- [
|
|
4034
|
-
|
|
4035
|
-
- [
|
|
4036
|
-
- [
|
|
4037
|
-
|
|
4038
|
-
- [
|
|
4039
|
-
|
|
1954
|
+
- [Live A/B metrics](/docs/ab.md): sticky variants and cumulative metrics
|
|
1955
|
+
on live sessions
|
|
1956
|
+
- [Hillclimbing](/docs/hillclimbing.md): the loop evals make trustworthy
|
|
1957
|
+
- [Building agents with agents](/docs/building-with-agents.md): have a
|
|
1958
|
+
coding agent write the first suite
|
|
1959
|
+
- [GitHub guide](/docs/guides/github.md): deterministic webhook fixtures
|
|
1960
|
+
with `github replay`
|
|
1961
|
+
- [Sessions and streaming](/docs/reference/sessions.md): the events
|
|
1962
|
+
`t.events` contains
|
|
4040
1963
|
|
|
4041
1964
|
---
|
|
4042
1965
|
|
|
@@ -4143,8 +2066,8 @@ Peer MCP connections are ordinary MCP connections, so deterministic host code
|
|
|
4143
2066
|
can use them too. A channel handler or server tool can call
|
|
4144
2067
|
`host.mcp.callTool("weather", "ask", { message: "…" })` without any
|
|
4145
2068
|
model turn deciding to. See
|
|
4146
|
-
[MCP connections](/docs/reference/connections.md#every-mcp-connection-is-available-in-three-places)
|
|
4147
|
-
for the three places every MCP connection is available.
|
|
2069
|
+
[MCP connections](/docs/reference/connections.md#every-model-visible-mcp-connection-is-available-in-three-places)
|
|
2070
|
+
for the three places every model-visible MCP connection is available.
|
|
4148
2071
|
|
|
4149
2072
|
## What's next
|
|
4150
2073
|
|
|
@@ -4223,6 +2146,7 @@ mapping shifts:
|
|
|
4223
2146
|
| Agent tools (`execution: "agent"`) | scripts in the session workspace | catalog + script bodies on the first prompt |
|
|
4224
2147
|
| `skills/*` | `.cursor/skills/` in the workspace | native discovery after the first turn, from the hosted store or the signed-in account |
|
|
4225
2148
|
| `mcp-connections/*.ts` | SDK `mcpServers` | SDK `mcpServers` (peers need `--public-url`) |
|
|
2149
|
+
| `host-connections/*.ts` | `ctx.host.mcp` only | `ctx.host.mcp` only |
|
|
4226
2150
|
| `sandbox/workspace/**` | seeded into the session workspace | ignored |
|
|
4227
2151
|
| Tool approvals (`needsApproval`) | supported | not supported; keep approval-gated tools on local turns |
|
|
4228
2152
|
|
|
@@ -4354,7 +2278,7 @@ or run lifecycle scripts from an existing `package.json`.
|
|
|
4354
2278
|
| Prompt model | Pins the model on `defineAgent` in `agent/agent.ts`. `git_config` and `agent_options` remain comments. |
|
|
4355
2279
|
| Cron trigger | Creates `agent/schedules/<slug>.ts` with `defineSchedule` in UTC. |
|
|
4356
2280
|
| GitHub trigger | Creates `agent/channels/github.ts`. It converts pull-request action, push branch, issue action, and user allowlist filters. |
|
|
4357
|
-
| Slack trigger | Creates `agent/channels/slack.ts
|
|
2281
|
+
| Slack trigger | Creates `agent/channels/slack.ts` as Socket Mode with `envPrefix` from the automation name (same names `slack create` writes). Watches add `engagement.channelPosts`. Run `agent-sdk slack create` for the bot. |
|
|
4358
2282
|
| Linear, PagerDuty, Sentry, Teams, or generic webhook | Creates a boilerplate `agent/channels/<slug>.ts`. |
|
|
4359
2283
|
| HTTP or SSE MCP server | Creates a name-based Cursor-account connection under `agent/mcp-connections/`. The project contains no server URL or credentials. |
|
|
4360
2284
|
| Stdio MCP server | Writes `agent/mcp-connections/<slug>.todo.md`. |
|
|
@@ -4716,10 +2640,10 @@ A comment-only first wake has no head SHA, so the check waits for a
|
|
|
4716
2640
|
PR or CI event. The banner still posts. A later turn on the same SHA
|
|
4717
2641
|
creates a new check run; GitHub cannot reopen a completed run.
|
|
4718
2642
|
|
|
4719
|
-
Override `events` when the mapping is custom.
|
|
4720
|
-
|
|
4721
|
-
|
|
4722
|
-
|
|
2643
|
+
Override `events` when the mapping is custom. A handler can post commit
|
|
2644
|
+
status from `turn.started` / `action.result` / `turn.failed` and stay
|
|
2645
|
+
never-red; that pattern still wins when you replace a default handler
|
|
2646
|
+
key. Handlers you author replace the matching defaults (same as
|
|
4723
2647
|
`progress.reactions` composition today).
|
|
4724
2648
|
|
|
4725
2649
|
## Related
|
|
@@ -4885,10 +2809,8 @@ The companion skill is
|
|
|
4885
2809
|
to that connection's resource URL
|
|
4886
2810
|
- Upsert deployment secrets with `--store` so hosted engines seed the
|
|
4887
2811
|
same tokens from env
|
|
4888
|
-
-
|
|
4889
|
-
tools still call
|
|
4890
|
-
connectors the playground or local chat should call. Use
|
|
4891
|
-
`advertiseTools: true` for those.
|
|
2812
|
+
- Use `advertiseTools: true` when local turns should call the server by
|
|
2813
|
+
name. Host tools can still call it through `ctx.host.mcp`.
|
|
4892
2814
|
|
|
4893
2815
|
Prefer a Cursor account MCP connection when the connector already lives
|
|
4894
2816
|
in the signed-in account dashboard:
|
|
@@ -4903,7 +2825,8 @@ and the host must hold tokens.
|
|
|
4903
2825
|
|
|
4904
2826
|
## How do I declare a host-OAuth connection?
|
|
4905
2827
|
|
|
4906
|
-
Add one file under `agent/mcp-connections
|
|
2828
|
+
Add one file under `agent/mcp-connections/` (model + host) or
|
|
2829
|
+
`agent/host-connections/` (host + `mcp oauth` only). The filename is the
|
|
4907
2830
|
connection name you pass to the CLI and to `host.mcp`.
|
|
4908
2831
|
|
|
4909
2832
|
```ts
|
|
@@ -4913,17 +2836,14 @@ import { defineConnection } from "@cursor/july/connections";
|
|
|
4913
2836
|
export default defineConnection({
|
|
4914
2837
|
url: "https://mcp.example.com/inventory",
|
|
4915
2838
|
oauth: true,
|
|
4916
|
-
|
|
4917
|
-
description:
|
|
4918
|
-
"Inventory MCP (privileged). Call only from host tools, not the model.",
|
|
2839
|
+
description: "Inventory MCP.",
|
|
4919
2840
|
});
|
|
4920
2841
|
```
|
|
4921
2842
|
|
|
4922
2843
|
Rules of the road:
|
|
4923
2844
|
|
|
4924
2845
|
- `oauth: true` is required for `agent-sdk mcp oauth`
|
|
4925
|
-
-
|
|
4926
|
-
channel handlers still see it
|
|
2846
|
+
- A file under `mcp-connections/` is visible to the model and to `ctx.host.mcp`. A file under `host-connections/` stays on the host.
|
|
4927
2847
|
- Declare expected secret names on the agent when you plan to `--store`:
|
|
4928
2848
|
|
|
4929
2849
|
```ts
|
|
@@ -4956,8 +2876,7 @@ agent-sdk mcp oauth inventory
|
|
|
4956
2876
|
|
|
4957
2877
|
What happens:
|
|
4958
2878
|
|
|
4959
|
-
1. The Agent SDK loads
|
|
4960
|
-
`oauth: true`
|
|
2879
|
+
1. The Agent SDK loads the connection file and checks `oauth: true`
|
|
4961
2880
|
2. It opens the authorization URL in your browser
|
|
4962
2881
|
3. The callback lands on `http://localhost:8787/callback`
|
|
4963
2882
|
4. Tokens land in `mcp-auth.json` under the CLI config directory
|
|
@@ -4996,8 +2915,6 @@ on the pod.
|
|
|
4996
2915
|
|
|
4997
2916
|
## How do host tools call the server?
|
|
4998
2917
|
|
|
4999
|
-
Keep privileged calls on the host:
|
|
5000
|
-
|
|
5001
2918
|
```ts
|
|
5002
2919
|
const result = await ctx.host.mcp.callTool(
|
|
5003
2920
|
"inventory",
|
|
@@ -5006,9 +2923,8 @@ const result = await ctx.host.mcp.callTool(
|
|
|
5006
2923
|
);
|
|
5007
2924
|
```
|
|
5008
2925
|
|
|
5009
|
-
The model
|
|
5010
|
-
|
|
5011
|
-
tool instead.
|
|
2926
|
+
The model can call the same server. Use a host tool when the write needs
|
|
2927
|
+
an allowlist or other deterministic gate.
|
|
5012
2928
|
|
|
5013
2929
|
## What if authorization fails?
|
|
5014
2930
|
|
|
@@ -5021,7 +2937,7 @@ tool instead.
|
|
|
5021
2937
|
|
|
5022
2938
|
## What's next
|
|
5023
2939
|
|
|
5024
|
-
- [MCP connections](/docs/reference/connections.md): transports,
|
|
2940
|
+
- [MCP connections](/docs/reference/connections.md): transports, account MCP
|
|
5025
2941
|
- [CLI](/docs/reference/cli.md#mcp-oauth): full flag list for `mcp oauth`
|
|
5026
2942
|
- [Deployment](/docs/deployment.md): secrets, egress, and hosted engines
|
|
5027
2943
|
- [Fix common agent problems](/docs/troubleshooting.md): more symptom → fix tables
|
|
@@ -5246,10 +3162,10 @@ Source: /docs/guides/slack.md
|
|
|
5246
3162
|
|
|
5247
3163
|
# Slack agents
|
|
5248
3164
|
|
|
5249
|
-
The Slack channel puts your agent in Slack
|
|
5250
|
-
|
|
5251
|
-
|
|
5252
|
-
|
|
3165
|
+
The Slack channel puts your agent in Slack as its own Socket Mode bot.
|
|
3166
|
+
`agent-sdk slack create` opens the dashboard wizard and writes tokens
|
|
3167
|
+
to `.env.local`. To own the Slack app yourself, run
|
|
3168
|
+
`agent-sdk slack init --manual` and paste the manifests at
|
|
5253
3169
|
[api.slack.com](https://api.slack.com/apps). Socket Mode has no
|
|
5254
3170
|
public Request URL. Replies stream in threads, with tool "thinking" steps,
|
|
5255
3171
|
suggested prompts, and opt-in approval buttons.
|
|
@@ -5288,41 +3204,12 @@ Missing tokens leave the channel idle (`channel idle … missing
|
|
|
5288
3204
|
credentials`) rather than failing `serve`. That's useful when you mount
|
|
5289
3205
|
many agents and only some have Slack apps.
|
|
5290
3206
|
|
|
5291
|
-
## Use the Cursor Slack connection
|
|
5292
|
-
|
|
5293
|
-
If the Cursor Slack app is already installed in your workspace and linked
|
|
5294
|
-
to your Cursor account, skip the dedicated Slack app:
|
|
5295
|
-
|
|
5296
|
-
```ts
|
|
5297
|
-
import { slackChannel } from "@cursor/july/channels/slack";
|
|
5298
|
-
|
|
5299
|
-
export default slackChannel({
|
|
5300
|
-
cursorAccount: true,
|
|
5301
|
-
agentName: "Weatherbot", // single token — no spaces; defaults from mount slug (PascalCase)
|
|
5302
|
-
agentIcon: { emoji: ":robot_face:" },
|
|
5303
|
-
});
|
|
5304
|
-
```
|
|
5305
|
-
|
|
5306
|
-
Sign the host in (`agent-sdk login` or `CURSOR_API_KEY`), then mention the
|
|
5307
|
-
agent in Slack as `@Cursor Weatherbot …`. Thread replies and DMs keep going to
|
|
5308
|
-
the same agent. Messages appear as the Cursor app under that agent's name
|
|
5309
|
-
and icon. Slack shows its working status, then posts one final reply.
|
|
5310
|
-
|
|
5311
|
-
Use a dedicated Socket Mode Slack app when you need your own bot user,
|
|
5312
|
-
channel watching (`engagement.channelPosts`), or approval buttons. On
|
|
5313
|
-
`cursorAccount`, agents must be explicitly addressed (@mention, DM, or
|
|
5314
|
-
claimed-thread reply). Channel watching and `toolApprovals` /
|
|
5315
|
-
`interactivity` are Socket Mode only; the Cursor connection does not relay
|
|
5316
|
-
Block Kit clicks. Agent names must be unique on the host; an unmatched
|
|
5317
|
-
`@Cursor <name>` stays on Cursor's normal Slack agent.
|
|
5318
|
-
|
|
5319
3207
|
## Control who can message the agent
|
|
5320
3208
|
|
|
5321
3209
|
External senders are blocked by default. Slack Connect users, guests, and
|
|
5322
3210
|
people whose home workspace is not the install team never reach the
|
|
5323
|
-
handler.
|
|
5324
|
-
|
|
5325
|
-
your org:
|
|
3211
|
+
handler. Set `blockExternals: false` only when the agent should serve
|
|
3212
|
+
people outside your org:
|
|
5326
3213
|
|
|
5327
3214
|
```ts
|
|
5328
3215
|
export default slackChannel({
|
|
@@ -5485,8 +3372,7 @@ export default slackChannel({
|
|
|
5485
3372
|
```
|
|
5486
3373
|
|
|
5487
3374
|
Channel watching needs the `message.channels` / `message.groups` events
|
|
5488
|
-
on the Slack app
|
|
5489
|
-
`cursorAccount: true`). Pass `--channel-posts` on `slack create` or
|
|
3375
|
+
on the Slack app. Pass `--channel-posts` on `slack create` or
|
|
5490
3376
|
`slack init --manual`. The bot must also be a member of each watched
|
|
5491
3377
|
channel.
|
|
5492
3378
|
|
|
@@ -5496,9 +3382,7 @@ Set `includeBotPosts: true` when the posts worth watching come from bots:
|
|
|
5496
3382
|
alert feeds, webhook integrations, or other agents posting notes. The
|
|
5497
3383
|
watching app's own posts stay dropped either way, matched by the `bot_id`
|
|
5498
3384
|
and bot user id from `auth.test`, so an agent can never dispatch on its
|
|
5499
|
-
own replies.
|
|
5500
|
-
[alert investigator example](/docs/example-agents/oncall.md) watches a
|
|
5501
|
-
bot-fed alerts channel this way.
|
|
3385
|
+
own replies. Use this for a bot-fed alerts channel.
|
|
5502
3386
|
|
|
5503
3387
|
## Prepare work on the host
|
|
5504
3388
|
|
|
@@ -5527,11 +3411,6 @@ Approval cards need interactivity on the Slack app. Recreate with
|
|
|
5527
3411
|
`buildToolApprovalEvents({ credentials })` into `events` and set
|
|
5528
3412
|
`interactivity: true` on the channel so Socket Mode routes the clicks.
|
|
5529
3413
|
|
|
5530
|
-
Approval buttons need Socket Mode. `slackChannel({ cursorAccount: true })`
|
|
5531
|
-
rejects `toolApprovals` and `interactivity` at construction, since the
|
|
5532
|
-
Cursor Slack connection does not relay Block Kit clicks. Use a dedicated
|
|
5533
|
-
Slack app to run approvals for a cursor-account agent.
|
|
5534
|
-
|
|
5535
3414
|
Cards show redacted, truncated arguments (Block Kit size limits);
|
|
5536
3415
|
execution still uses the full validated input, so review sensitive tools
|
|
5537
3416
|
in the playground when the arguments may exceed the card. Approvals
|
|
@@ -5563,7 +3442,7 @@ Two habits matter most.
|
|
|
5563
3442
|
The `slack` subcommands cover setup end to end.
|
|
5564
3443
|
|
|
5565
3444
|
```bash
|
|
5566
|
-
agent-sdk slack setup #
|
|
3445
|
+
agent-sdk slack setup # printed setup guide
|
|
5567
3446
|
agent-sdk slack create --dir . # dashboard wizard (dev app)
|
|
5568
3447
|
agent-sdk slack create --dir . --prod # prod app
|
|
5569
3448
|
agent-sdk slack destroy --dir . # delete the provisioned app
|
|
@@ -6178,7 +4057,6 @@ npx @cursor/july docs
|
|
|
6178
4057
|
| New to the Agent SDK | [Quickstart](/docs/quickstart.md) (PR reviewer), then [Concepts](/docs/concepts.md) |
|
|
6179
4058
|
| Building a new agent with Cursor | [Scaffold an agent with Cursor](/docs/scaffolding-agents.md) |
|
|
6180
4059
|
| Turning a Cursor Automation into a project | [Convert a Cursor Automation](/docs/guides/convert-automation.md) |
|
|
6181
|
-
| Learning from working agents | [Example agents](/docs/example-agents/index.md) |
|
|
6182
4060
|
| Wiring an agent to Slack | [Slack guide](/docs/guides/slack.md) |
|
|
6183
4061
|
| Starting from a packaged template | [Demo](/docs/templates/demo.md), [Security reviewer](/docs/templates/security-reviewer.md), [Triage](/docs/templates/triage.md), or [Agentic Owners](/docs/templates/agentic-owners.md) |
|
|
6184
4062
|
| Wiring an agent to GitHub webhooks | [GitHub guide](/docs/guides/github.md) |
|
|
@@ -6246,33 +4124,6 @@ npx @cursor/july docs
|
|
|
6246
4124
|
- [OpenTelemetry](/docs/guides/opentelemetry.md): push session, turn, and
|
|
6247
4125
|
tool traces to an OTLP collector you run.
|
|
6248
4126
|
|
|
6249
|
-
**Example agents**
|
|
6250
|
-
|
|
6251
|
-
- [Choose the right example](/docs/example-agents/index.md): compare the
|
|
6252
|
-
example agents by runtime, channels, tools, and state.
|
|
6253
|
-
- [Weather agent](/docs/example-agents/weather-agent.md): explore tools, MCP,
|
|
6254
|
-
approvals, skills, subagents, schedules, hooks, A/B metrics, and evals.
|
|
6255
|
-
- [Slack agent](/docs/example-agents/slack-agent.md): put a minimal agent in
|
|
6256
|
-
Slack through an account-linked transport.
|
|
6257
|
-
- [Concierge](/docs/example-agents/concierge.md): delegate work to a peer agent
|
|
6258
|
-
with its own context and sessions.
|
|
6259
|
-
- [Playbook router](/docs/example-agents/benny.md): route Slack intake through inherited
|
|
6260
|
-
repository playbooks.
|
|
6261
|
-
- [Alert investigator](/docs/example-agents/oncall.md): watch a Slack alerts
|
|
6262
|
-
channel and pin a self-rechecking investigation to every alert thread.
|
|
6263
|
-
- [PR evidence reviewer](/docs/example-agents/bugbot.md): review a host-prepared,
|
|
6264
|
-
diff-first pull-request evidence tree.
|
|
6265
|
-
- [Approval Buddy](/docs/example-agents/approval-buddy.md): keep approval policy
|
|
6266
|
-
in code while subagents supply review findings.
|
|
6267
|
-
- [Security Reviewer](/docs/example-agents/security-reviewer.md): run a staged,
|
|
6268
|
-
parallel security pipeline with live playground progress.
|
|
6269
|
-
- [Knowledge base](/docs/example-agents/knowledge-base.md): turn conversations
|
|
6270
|
-
about people, systems, decisions, and preferences into shared markdown.
|
|
6271
|
-
- [Codebase wiki](/docs/example-agents/codebase-wiki.md): ingest merged PRs into
|
|
6272
|
-
per-feature pages with a daily digest schedule.
|
|
6273
|
-
- [Codeowners review](/docs/example-agents/codeowners-review.md): route PR
|
|
6274
|
-
reviews by ownership to per-area playbooks and aggregate verdicts.
|
|
6275
|
-
|
|
6276
4127
|
**Operating**
|
|
6277
4128
|
|
|
6278
4129
|
- [Deployment](/docs/deployment.md): Cursor-managed hosting, self-hosting,
|
|
@@ -6676,9 +4527,8 @@ See [GitHub](/docs/guides/github.md) for local event delivery and
|
|
|
6676
4527
|
|
|
6677
4528
|
## Where to go next
|
|
6678
4529
|
|
|
6679
|
-
- [
|
|
6680
|
-
|
|
6681
|
-
policy
|
|
4530
|
+
- [PR autofixer template](/docs/templates/pr-autofixer.md): drive a PR on a
|
|
4531
|
+
Cursor cloud VM
|
|
6682
4532
|
- [Evals](/docs/evals.md): freeze these two PRs as regression checks so
|
|
6683
4533
|
prompt changes can't flip a verdict
|
|
6684
4534
|
- [Tools](/docs/reference/tools.md): more on typed tools, approvals, and
|
|
@@ -8062,7 +5912,8 @@ command prints a note when you pass it anyway.
|
|
|
8062
5912
|
agent-sdk mcp oauth <connection> [--dir .] [--store] [--slug <slug>] [--team <id>]
|
|
8063
5913
|
```
|
|
8064
5914
|
|
|
8065
|
-
`<connection>` is the `agent/mcp-connections
|
|
5915
|
+
`<connection>` is the basename under `agent/mcp-connections/` or
|
|
5916
|
+
`agent/host-connections/`.
|
|
8066
5917
|
`--slug` defaults to the `--dir` basename. `--team` defaults to the
|
|
8067
5918
|
signed-in account's team. You need `agent-sdk login` (or `--api-key`)
|
|
8068
5919
|
before `--store`.
|
|
@@ -8275,6 +6126,11 @@ from `@cursor/july/connections`, and the transport comes in
|
|
|
8275
6126
|
four shapes: remote HTTP, local stdio, the signed-in Cursor account's
|
|
8276
6127
|
connectors, and peer agents on the same host.
|
|
8277
6128
|
|
|
6129
|
+
Put a server in `agent/host-connections/` when host tools should call it
|
|
6130
|
+
and the model should not. Same `defineConnection` shape. `agent-sdk mcp
|
|
6131
|
+
oauth` still works. The playground and the turn's MCP servers never see
|
|
6132
|
+
those files.
|
|
6133
|
+
|
|
8278
6134
|
## Remote MCP server
|
|
8279
6135
|
|
|
8280
6136
|
Point an MCP connection at a remote server with a URL.
|
|
@@ -8311,12 +6167,10 @@ agent-sdk mcp oauth inventory --store # also upsert deployment secrets
|
|
|
8311
6167
|
Full walkthrough: [Host MCP OAuth](/docs/guides/mcp-oauth.md). Companion
|
|
8312
6168
|
skill: [`skills/mcp-auth/SKILL.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/skills/mcp-auth/SKILL.md).
|
|
8313
6169
|
|
|
8314
|
-
|
|
8315
|
-
|
|
8316
|
-
|
|
8317
|
-
|
|
8318
|
-
connected connector. If the model should call those tools by name on
|
|
8319
|
-
local turns, set `advertiseTools: true`.
|
|
6170
|
+
Account MCP (`cursorAccount: true`) is the right choice for connectors
|
|
6171
|
+
already linked in the Cursor dashboard. Omit `servers` (or pass `"*"`)
|
|
6172
|
+
to forward every connected connector. If the model should call those
|
|
6173
|
+
tools by name on local turns, set `advertiseTools: true`.
|
|
8320
6174
|
|
|
8321
6175
|
## Per-session auth (`auth`)
|
|
8322
6176
|
|
|
@@ -8376,7 +6230,7 @@ export default defineConnection({
|
|
|
8376
6230
|
|
|
8377
6231
|
A listing failure, invalid tool name, or name collision fails the turn.
|
|
8378
6232
|
Advertised tools follow the same runtime support as server tools. They cannot
|
|
8379
|
-
be
|
|
6233
|
+
be called through the direct tool API.
|
|
8380
6234
|
|
|
8381
6235
|
In a dry-run session, MCP tools marked read-only run normally. Tools marked
|
|
8382
6236
|
as writes are stubbed. Tools without effect annotations are unavailable.
|
|
@@ -8484,13 +6338,14 @@ Unknown slugs and self-references fail `serve` at startup. Resolution
|
|
|
8484
6338
|
(loopback versus `--public-url`), loop caveats, and the delegation model
|
|
8485
6339
|
are in the [Agent-to-agent guide](/docs/guides/agent-to-agent.md).
|
|
8486
6340
|
|
|
8487
|
-
## Every MCP connection is available in three places
|
|
6341
|
+
## Every model-visible MCP connection is available in three places
|
|
8488
6342
|
|
|
8489
|
-
|
|
6343
|
+
A file under `agent/mcp-connections/` serves three consumers. Host
|
|
6344
|
+
connections skip the first one.
|
|
8490
6345
|
|
|
8491
6346
|
1. **Cursor agent:** Attached connections ride SDK `mcpServers` behind
|
|
8492
6347
|
harness MCP meta-tools. Set `advertiseTools: true` so local turns see
|
|
8493
|
-
named tools.
|
|
6348
|
+
named tools.
|
|
8494
6349
|
2. **Server tools:** Deterministic host code composes MCP calls
|
|
8495
6350
|
through `ctx.host.mcp`:
|
|
8496
6351
|
|
|
@@ -8525,7 +6380,7 @@ lazily on first use.
|
|
|
8525
6380
|
|
|
8526
6381
|
Continue with these pages:
|
|
8527
6382
|
|
|
8528
|
-
- [Host MCP OAuth](/docs/guides/mcp-oauth.md): `mcp oauth`, `--store
|
|
6383
|
+
- [Host MCP OAuth](/docs/guides/mcp-oauth.md): `mcp oauth`, `--store`
|
|
8529
6384
|
- [Agent-to-agent](/docs/guides/agent-to-agent.md): peers in depth
|
|
8530
6385
|
- [Tools](/docs/reference/tools.md): authored tools that wrap MCP connections
|
|
8531
6386
|
- [Webhooks](/docs/guides/webhooks.md): calling MCP connections from handlers
|
|
@@ -8602,9 +6457,8 @@ For GitHub merge-box checks and sticky PR banners, use
|
|
|
8602
6457
|
`githubChannel({ progress: { commitStatus, banner } })` from
|
|
8603
6458
|
`@cursor/july/channels/github`. That is the supported Autofix-style
|
|
8604
6459
|
path. See [GitHub: Show PR progress](/docs/guides/github.md#show-pr-progress).
|
|
8605
|
-
Override channel `events` only when the lifecycle is custom
|
|
8606
|
-
|
|
8607
|
-
never-red status from tool output). Do not use `defineHook` for those
|
|
6460
|
+
Override channel `events` only when the lifecycle is custom, such as
|
|
6461
|
+
never-red status from tool output. Do not use `defineHook` for those
|
|
8608
6462
|
writes.
|
|
8609
6463
|
|
|
8610
6464
|
## Patterns
|
|
@@ -8772,6 +6626,12 @@ while a turn runs). Agent-execution tools are rejected with `400`, and
|
|
|
8772
6626
|
unknown tools with `404` and the list of available names. For the
|
|
8773
6627
|
semantics, see [Tools](/docs/reference/tools.md#call-a-tool-without-a-model-turn).
|
|
8774
6628
|
|
|
6629
|
+
An optional `"continuationToken"` (`<channelId>:<key>`, as
|
|
6630
|
+
`/v1/sessions` lists it; mutually exclusive with `sessionId`) addresses
|
|
6631
|
+
the session by continuation token instead; malformed tokens are
|
|
6632
|
+
rejected with `400 invalid_continuation_token`. For the semantics, see
|
|
6633
|
+
[Tools](/docs/reference/tools.md#call-a-tool-without-a-model-turn).
|
|
6634
|
+
|
|
8775
6635
|
## Discovery
|
|
8776
6636
|
|
|
8777
6637
|
These read-only routes describe the running agent.
|
|
@@ -8779,6 +6639,8 @@ These read-only routes describe the running agent.
|
|
|
8779
6639
|
| Route | What it does |
|
|
8780
6640
|
| ---------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
8781
6641
|
| `GET /v1/info` | The discovered surface: model, tools, skills, MCP connections, subagents, channels and routes (with schemas), schedules, hooks, A/B experiments, diagnostics |
|
|
6642
|
+
| `GET /v1/tools` | The live tool catalog: authored server tools plus advertised MCP passthroughs under model-facing names, as light `{ name, title?, source? }` entries. `session` / `continuationToken` query parameters bind the listing to a session identity (advertised inventories can be tenant-scoped); a connection whose listing fails is skipped and reported in `connectionErrors` |
|
|
6643
|
+
| `GET /v1/tools/:name` | One catalog tool's full description: description, execution, `needsApproval`, `effect`, input and output schemas, source connection. Same session binding as the listing; unknown names get `404` with the available names |
|
|
8782
6644
|
| `GET /v1/health` | Per-agent liveness, no auth |
|
|
8783
6645
|
| `GET /v1/logs?after=N` | Recent server log lines, with a polling cursor |
|
|
8784
6646
|
| `GET /v1/abs` | [Live A/B metrics](/docs/ab.md): per-session assignments and aggregate arm totals |
|
|
@@ -9044,7 +6906,8 @@ experiments can override their file-derived name.
|
|
|
9044
6906
|
| Path | Resolves to |
|
|
9045
6907
|
| --- | --- |
|
|
9046
6908
|
| `agent/tools/approve_pr.ts` | tool `approve_pr` |
|
|
9047
|
-
| `agent/mcp-connections/linear.ts` | MCP connection `linear` |
|
|
6909
|
+
| `agent/mcp-connections/linear.ts` | MCP connection `linear` (model + host) |
|
|
6910
|
+
| `agent/host-connections/anytool.ts` | Host MCP connection `anytool` (host + `mcp oauth` only) |
|
|
9048
6911
|
| `agent/skills/pr-review.md` | skill `pr-review` |
|
|
9049
6912
|
| `agent/subagents/reviewer/` | subagent `reviewer` |
|
|
9050
6913
|
| `agent/channels/drive.ts` | channel `drive`, routes under `/v1/channels/drive` |
|
|
@@ -9072,6 +6935,8 @@ my-agent/
|
|
|
9072
6935
|
│ │ └── pr-review.md # on-demand procedures (SKILL.md convention)
|
|
9073
6936
|
│ ├── mcp-connections/
|
|
9074
6937
|
│ │ └── linear.ts # tools from external MCP servers
|
|
6938
|
+
│ ├── host-connections/
|
|
6939
|
+
│ │ └── anytool.ts # privileged MCP, host tools only
|
|
9075
6940
|
│ └── channels/
|
|
9076
6941
|
│ └── github.ts # messages and external events
|
|
9077
6942
|
└── evals/
|
|
@@ -9093,6 +6958,7 @@ Each path maps to a capability and a reference page.
|
|
|
9093
6958
|
| `agent/tools/<name>.ts` | One typed tool; filename = tool name. `execution: "server"` (in-process, default) or `"agent"` (a script that runs where the agent runs) | [Tools](/docs/reference/tools.md) |
|
|
9094
6959
|
| `agent/skills/*` | SKILL.md-convention procedures, loaded on demand | [Skills](/docs/reference/skills.md) |
|
|
9095
6960
|
| `agent/mcp-connections/<name>.ts` | MCP servers, available to the model, to server tools (`ctx.host.mcp`), and to channel/schedule handlers (`args.host.mcp`) | [MCP connections](/docs/reference/connections.md) |
|
|
6961
|
+
| `agent/host-connections/<name>.ts` | Privileged MCP servers for `ctx.host.mcp` and `mcp oauth`. The model never sees them. | [MCP connections](/docs/reference/connections.md) |
|
|
9096
6962
|
| `agent/subagents/<id>/` | Child agent directory; `description` required | [Subagents](/docs/reference/subagents.md) |
|
|
9097
6963
|
| `agent/channels/*.ts` | HTTP surfaces beyond the built-in session API; `slack.ts` and `github.ts` use the platform packs | [Channels](/docs/reference/channels.md) |
|
|
9098
6964
|
| `agent/hooks/*.ts` | Observe-only event subscribers, never fatal | [Hooks](/docs/reference/hooks.md) |
|
|
@@ -9698,8 +7564,8 @@ different one.
|
|
|
9698
7564
|
Subagents inherit the parent's execution surface. Every per-subagent
|
|
9699
7565
|
capability directory is reported as a warning and ignored: `tools/`,
|
|
9700
7566
|
`skills/`, `mcp-connections/` (and the legacy `connections/` alias),
|
|
9701
|
-
`channels/`, `schedules/`, `hooks/`, `sandbox/`,
|
|
9702
|
-
`subagents/`.
|
|
7567
|
+
`host-connections/`, `channels/`, `schedules/`, `hooks/`, `sandbox/`,
|
|
7568
|
+
and nested `subagents/`.
|
|
9703
7569
|
|
|
9704
7570
|
Delegation needs both halves: the description makes it possible, and the
|
|
9705
7571
|
parent's [instructions](/docs/reference/instructions.md) make it happen. "When a
|
|
@@ -9920,8 +7786,9 @@ passthrough server tools from the connection's live `listTools` on every
|
|
|
9920
7786
|
local turn. See
|
|
9921
7787
|
[MCP Connections](/docs/reference/connections.md#advertise-tools).
|
|
9922
7788
|
Advertised tools ride the same execution path as authored server tools,
|
|
9923
|
-
|
|
9924
|
-
|
|
7789
|
+
and [direct tool calls](#call-a-tool-without-a-model-turn) address them
|
|
7790
|
+
by the same model-facing names: the call's session identity resolves the
|
|
7791
|
+
advertised listing when the authored lookup misses.
|
|
9925
7792
|
|
|
9926
7793
|
## Gate a tool on human approval
|
|
9927
7794
|
|
|
@@ -9995,7 +7862,22 @@ when the call returns. Pass a `sessionId` (a body field over
|
|
|
9995
7862
|
HTTP, `--session` on the CLI, `options.sessionId` programmatically) to
|
|
9996
7863
|
run inside an existing session instead: the tool sees that session's
|
|
9997
7864
|
workspace, and the call is recorded on the session's event stream.
|
|
9998
|
-
Session-bound calls return `409 session_busy` while a turn runs.
|
|
7865
|
+
Session-bound calls return `409 session_busy` while a turn runs. When the
|
|
7866
|
+
session's harness cwd cannot be materialized, a read-effect call runs
|
|
7867
|
+
in a scratch workspace instead and the outcome carries
|
|
7868
|
+
`scratchWorkspace: true`; a write-effect call fails with
|
|
7869
|
+
`workspace_unavailable`.
|
|
7870
|
+
|
|
7871
|
+
A session can also be addressed by its continuation token: an optional
|
|
7872
|
+
`continuationToken` (`<channelId>:<key>`, as `/v1/sessions` lists it;
|
|
7873
|
+
mutually exclusive with `sessionId`). A token that maps to a live
|
|
7874
|
+
session behaves exactly like passing that session's id — same ownership
|
|
7875
|
+
check, same `409 session_busy`, same event recording. A token with no
|
|
7876
|
+
session behind it runs the call scratch-bound with the token's channel
|
|
7877
|
+
id and continuation key as the call's session identity, so a deployment
|
|
7878
|
+
whose tools resolve state from the continuation key can serve it with
|
|
7879
|
+
no live session. Malformed tokens are rejected with
|
|
7880
|
+
`400 invalid_continuation_token`.
|
|
9999
7881
|
|
|
10000
7882
|
The error semantics match the model path. Unknown tools are rejected
|
|
10001
7883
|
with the available names, agent-execution tools cannot be called on the
|
|
@@ -10630,9 +8512,10 @@ fixture replay.
|
|
|
10630
8512
|
|
|
10631
8513
|
## Slack
|
|
10632
8514
|
|
|
10633
|
-
`agent/channels/slack.ts`
|
|
10634
|
-
|
|
10635
|
-
|
|
8515
|
+
`agent/channels/slack.ts` is a dedicated Socket Mode bot
|
|
8516
|
+
(`PR_AUTOFIXER_SLACK_*`). Mint it with `agent-sdk slack create`, then
|
|
8517
|
+
mention the bot or DM it with a PR URL.
|
|
8518
|
+
Same `drive_pr` path as the playground. See [Slack](/docs/guides/slack.md).
|
|
10636
8519
|
|
|
10637
8520
|
## Evals
|
|
10638
8521
|
|
|
@@ -10669,9 +8552,8 @@ This agent reads a pull request diff and posts one review comment. It
|
|
|
10669
8552
|
reports exploitable bugs: injection, authz bypass, secret leaks, SSRF,
|
|
10670
8553
|
RCE. Style nits stay out.
|
|
10671
8554
|
|
|
10672
|
-
The factory
|
|
10673
|
-
|
|
10674
|
-
model turn.
|
|
8555
|
+
The in-repo factory Security Reviewer runs a staged pipeline with
|
|
8556
|
+
parallel workers. This template is one model turn.
|
|
10675
8557
|
|
|
10676
8558
|
## Scaffold
|
|
10677
8559
|
|
|
@@ -10884,7 +8766,7 @@ not on `PATH`, use `npx @cursor/july`.
|
|
|
10884
8766
|
| Built-in file reads and greps fail; the turn retries for a long time | Run under Node 22.13+ (or `tsx`), never Bun. Look for `NGHTTP2_FRAME_SIZE_ERROR` in logs. |
|
|
10885
8767
|
| The turn fails immediately with an API-key error | Sign in with `agent-sdk login`, or set `CURSOR_API_KEY`. Discovery, `info`, `call`, and serve bring-up work without a key; model turns need one. |
|
|
10886
8768
|
| Replies quote rules or `AGENTS.md` from outside your agent project | The session workspace inherited parent-folder config. Nested git checkouts default `local.cwd` to a per-project cache directory under `~/.cache`. Point `defineAgent({ local: { cwd } })` at a checkout only when the agent should inherit that tree, or set `--state-root` to a clean directory (for example under `/tmp`). |
|
|
10887
|
-
| Yellow box shows Datadog/Linear tools, but the model lists `GetDynamicTools` / IDE `cursor` tools and never calls them | Attached MCP sits behind harness meta-tools, or
|
|
8769
|
+
| Yellow box shows Datadog/Linear tools, but the model lists `GetDynamicTools` / IDE `cursor` tools and never calls them | Attached MCP sits behind harness meta-tools, or the harness cwd is still inside another checkout. Set `advertiseTools: true` for named tools on local turns. Check `GET /v1/info` `local.cwd` and `connections[].advertiseTools`. |
|
|
10888
8770
|
| Server tools, skills, or workspace seed files never appear | Server tools and sandbox seeds apply on the local runtime (cloud server tools need `--public-url` / `--cloud-tools-url`). Skills reach cloud through the Agent Store when hosting or a personal `CURSOR_API_KEY` is available; otherwise only skills already in the cloud repo. `validate` warns when this combination is present. |
|
|
10889
8771
|
| `validate` and `run` succeed, but typecheck fails in CI | The CLI runs TypeScript with type-stripping only. Keep tool `execute` return types as object literals or `type` aliases, not `interface` types. |
|
|
10890
8772
|
| Login works, but turns are rejected when using custom API hosts | Point login and model traffic at the same host (`CURSOR_API_BASE_URL` and `CURSOR_BACKEND_URL`). A key from one host is rejected by the other. |
|
|
@@ -10923,7 +8805,7 @@ not on `PATH`, use `npx @cursor/july`.
|
|
|
10923
8805
|
| --- | --- |
|
|
10924
8806
|
| `must be defineConnection({ url, oauth: true })` | The connection file needs `oauth: true`, or you passed the wrong connection name to `agent-sdk mcp oauth`. |
|
|
10925
8807
|
| Local auth works; hosted calls unauthorized | Run `agent-sdk mcp oauth <name> --store`, confirm names with `agent-sdk secrets list <slug>`, then redeploy. |
|
|
10926
|
-
| Model asks for `mcp_auth` or IDE MCP for a
|
|
8808
|
+
| Model asks for `mcp_auth` or IDE MCP for a connector it already has | Attached MCP is behind meta-tools. Set `advertiseTools: true` for named tools on local turns, or call it from a host tool via `ctx.host.mcp`. |
|
|
10927
8809
|
|
|
10928
8810
|
See [Host MCP OAuth](/docs/guides/mcp-oauth.md) and
|
|
10929
8811
|
[`skills/mcp-auth/SKILL.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/skills/mcp-auth/SKILL.md).
|