@cursor/july 0.1.93 → 0.1.94

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (318) hide show
  1. package/AGENTS.md +8 -20
  2. package/README.md +4 -26
  3. package/dist/channels/slack/attachments.js +2 -2
  4. package/dist/channels/slack/dispatch.d.ts +0 -7
  5. package/dist/channels/slack/dispatch.d.ts.map +1 -1
  6. package/dist/channels/slack/dispatch.js +4 -7
  7. package/dist/channels/slack/eval-directive.d.ts +5 -12
  8. package/dist/channels/slack/eval-directive.d.ts.map +1 -1
  9. package/dist/channels/slack/eval-directive.js +8 -19
  10. package/dist/channels/slack/index.d.ts +0 -6
  11. package/dist/channels/slack/index.d.ts.map +1 -1
  12. package/dist/channels/slack/index.js +0 -6
  13. package/dist/channels/slack/setup.d.ts +4 -4
  14. package/dist/channels/slack/setup.d.ts.map +1 -1
  15. package/dist/channels/slack/setup.js +8 -15
  16. package/dist/channels/slack/slack-channel.d.ts +6 -13
  17. package/dist/channels/slack/slack-channel.d.ts.map +1 -1
  18. package/dist/channels/slack/slack-channel.js +15 -101
  19. package/dist/channels/slack/types.d.ts +12 -79
  20. package/dist/channels/slack/types.d.ts.map +1 -1
  21. package/dist/channels/slack/types.js +1 -15
  22. package/dist/client.d.ts +14 -0
  23. package/dist/client.d.ts.map +1 -0
  24. package/dist/client.js +12 -0
  25. package/dist/connections.d.ts +18 -9
  26. package/dist/connections.d.ts.map +1 -1
  27. package/dist/connections.js +17 -8
  28. package/dist/docs/404.html +2 -2
  29. package/dist/docs/ab.html +4 -4
  30. package/dist/docs/assets/{app.CjWU-x0z.js → app.CFDEas4I.js} +1 -1
  31. package/dist/docs/assets/chunks/@localSearchIndexroot.DU3U2Ij2.js +1 -0
  32. package/dist/docs/assets/chunks/{VPLocalSearchBox.Cxy8ySFQ.js → VPLocalSearchBox.B1IIYpYS.js} +1 -1
  33. package/dist/docs/assets/chunks/{theme.Dvq1Bktu.js → theme.Ct4NSiLm.js} +2 -2
  34. package/dist/docs/assets/concepts.md.lwAgBIMI.js +1 -0
  35. package/dist/docs/assets/{deployment.md.DoLFAzfm.js → deployment.md.D9msOFOW.js} +3 -8
  36. package/dist/docs/assets/{deployment.md.DoLFAzfm.lean.js → deployment.md.D9msOFOW.lean.js} +1 -1
  37. package/dist/docs/assets/{guides_agent-to-agent.md.B3JIaAqz.js → guides_agent-to-agent.md.BDb0t1QV.js} +1 -1
  38. package/dist/docs/assets/guides_cloud-runtime.md.CkYbjnAX.js +9 -0
  39. package/dist/docs/assets/guides_cloud-runtime.md.CkYbjnAX.lean.js +1 -0
  40. package/dist/docs/assets/{guides_convert-automation.md.Bboisykk.js → guides_convert-automation.md.B4sjlodG.js} +1 -1
  41. package/dist/docs/assets/{guides_github.md.DqJhuaN1.js → guides_github.md.Cnh2mL4a.js} +1 -1
  42. package/dist/docs/assets/{guides_mcp-oauth.md.CJvrXtkN.js → guides_mcp-oauth.md.DPYmBCbV.js} +7 -9
  43. package/dist/docs/assets/{guides_mcp-oauth.md.CJvrXtkN.lean.js → guides_mcp-oauth.md.DPYmBCbV.lean.js} +1 -1
  44. package/dist/docs/assets/{guides_slack.md.mqeNKs84.js → guides_slack.md.C32HsdKk.js} +5 -11
  45. package/dist/docs/assets/guides_slack.md.C32HsdKk.lean.js +1 -0
  46. package/dist/docs/assets/index.md.DRakGHFe.js +5 -0
  47. package/dist/docs/assets/{index.md.B-lVR4wT.lean.js → index.md.DRakGHFe.lean.js} +1 -1
  48. package/dist/docs/assets/{quickstart.md.BrmfrrIr.js → quickstart.md.Nj_LjW_a.js} +1 -1
  49. package/dist/docs/assets/{reference_cli.md.D9KESDsD.js → reference_cli.md.Cw6_ICYG.js} +1 -1
  50. package/dist/docs/assets/{reference_connections.md.DB6SsN6U.js → reference_connections.md.BH8Oc0D0.js} +5 -5
  51. package/dist/docs/assets/{reference_connections.md.DB6SsN6U.lean.js → reference_connections.md.BH8Oc0D0.lean.js} +1 -1
  52. package/dist/docs/assets/{reference_hooks.md.BxN87gCw.js → reference_hooks.md.a8BJxMR5.js} +1 -1
  53. package/dist/docs/assets/reference_http-api.md.D89k1mdm.js +11 -0
  54. package/dist/docs/assets/reference_http-api.md.D89k1mdm.lean.js +1 -0
  55. package/dist/docs/assets/reference_project-layout.md.Bv4KOtlB.js +19 -0
  56. package/dist/docs/assets/{reference_skills.md.BFW9retM.js → reference_skills.md.8son6Hjm.js} +3 -3
  57. package/dist/docs/assets/{reference_subagents.md.Xoav0AII.js → reference_subagents.md.CfsIloPm.js} +1 -1
  58. package/dist/docs/assets/{reference_tools.md.DuKvkYWG.js → reference_tools.md.BHeXn2id.js} +3 -3
  59. package/dist/docs/assets/{reference_tools.md.DuKvkYWG.lean.js → reference_tools.md.BHeXn2id.lean.js} +1 -1
  60. package/dist/docs/assets/{templates_pr-autofixer.md.R4K_qytS.js → templates_pr-autofixer.md.DU7dQpor.js} +2 -2
  61. package/dist/docs/assets/{templates_pr-autofixer.md.R4K_qytS.lean.js → templates_pr-autofixer.md.DU7dQpor.lean.js} +1 -1
  62. package/dist/docs/assets/{templates_security-reviewer.md.ByFyRta2.js → templates_security-reviewer.md.CTa7u_l1.js} +2 -2
  63. package/dist/docs/assets/{templates_security-reviewer.md.ByFyRta2.lean.js → templates_security-reviewer.md.CTa7u_l1.lean.js} +1 -1
  64. package/dist/docs/assets/troubleshooting.md.Ctv3T8C2.js +1 -0
  65. package/dist/docs/building-with-agents.html +4 -4
  66. package/dist/docs/concepts.html +5 -5
  67. package/dist/docs/concepts.md +1 -0
  68. package/dist/docs/deployment.html +7 -12
  69. package/dist/docs/deployment.md +1 -20
  70. package/dist/docs/design/agsh.md +406 -0
  71. package/dist/docs/evals.html +4 -4
  72. package/dist/docs/guides/agent-to-agent.html +6 -6
  73. package/dist/docs/guides/agent-to-agent.md +2 -2
  74. package/dist/docs/guides/cloud-runtime.html +6 -6
  75. package/dist/docs/guides/cloud-runtime.md +1 -0
  76. package/dist/docs/guides/convert-automation.html +6 -6
  77. package/dist/docs/guides/convert-automation.md +1 -1
  78. package/dist/docs/guides/github.html +6 -6
  79. package/dist/docs/guides/github.md +4 -4
  80. package/dist/docs/guides/human-in-the-loop.html +4 -4
  81. package/dist/docs/guides/mcp-oauth.html +11 -13
  82. package/dist/docs/guides/mcp-oauth.md +10 -18
  83. package/dist/docs/guides/opentelemetry.html +5 -5
  84. package/dist/docs/guides/slack.html +9 -15
  85. package/dist/docs/guides/slack.md +9 -46
  86. package/dist/docs/guides/webhooks.html +4 -4
  87. package/dist/docs/hashmap.json +1 -1
  88. package/dist/docs/hillclimbing.html +4 -4
  89. package/dist/docs/index.html +6 -6
  90. package/dist/docs/index.md +0 -28
  91. package/dist/docs/llms-full.txt +712 -2830
  92. package/dist/docs/llms.txt +2 -16
  93. package/dist/docs/quickstart.html +6 -6
  94. package/dist/docs/quickstart.md +2 -3
  95. package/dist/docs/reference/agent-config.html +4 -4
  96. package/dist/docs/reference/artifacts.html +4 -4
  97. package/dist/docs/reference/channels.html +4 -4
  98. package/dist/docs/reference/cli.html +6 -6
  99. package/dist/docs/reference/cli.md +2 -1
  100. package/dist/docs/reference/connections.html +9 -9
  101. package/dist/docs/reference/connections.md +15 -11
  102. package/dist/docs/reference/hooks.html +6 -6
  103. package/dist/docs/reference/hooks.md +2 -3
  104. package/dist/docs/reference/http-api.html +6 -6
  105. package/dist/docs/reference/http-api.md +8 -0
  106. package/dist/docs/reference/instructions.html +4 -4
  107. package/dist/docs/reference/playground.html +4 -4
  108. package/dist/docs/reference/project-layout.html +8 -6
  109. package/dist/docs/reference/project-layout.md +5 -1
  110. package/dist/docs/reference/prompt.html +4 -4
  111. package/dist/docs/reference/schedules.html +4 -4
  112. package/dist/docs/reference/sessions.html +4 -4
  113. package/dist/docs/reference/skills.html +7 -7
  114. package/dist/docs/reference/subagents.html +6 -6
  115. package/dist/docs/reference/subagents.md +2 -2
  116. package/dist/docs/reference/tools.html +7 -7
  117. package/dist/docs/reference/tools.md +19 -3
  118. package/dist/docs/scaffolding-agents.html +4 -4
  119. package/dist/docs/storage.html +4 -4
  120. package/dist/docs/templates/agentic-owners.html +4 -4
  121. package/dist/docs/templates/demo.html +4 -4
  122. package/dist/docs/templates/pr-autofixer.html +6 -6
  123. package/dist/docs/templates/pr-autofixer.md +4 -3
  124. package/dist/docs/templates/security-reviewer.html +5 -5
  125. package/dist/docs/templates/security-reviewer.md +2 -3
  126. package/dist/docs/templates/triage.html +4 -4
  127. package/dist/docs/troubleshooting.html +5 -5
  128. package/dist/docs/troubleshooting.md +2 -2
  129. package/dist/index.d.ts +1 -1
  130. package/dist/index.d.ts.map +1 -1
  131. package/dist/index.js +1 -1
  132. package/dist/internal/advertise-tools.d.ts +11 -0
  133. package/dist/internal/advertise-tools.d.ts.map +1 -1
  134. package/dist/internal/advertise-tools.js +47 -9
  135. package/dist/internal/cli-mcp-oauth.d.ts.map +1 -1
  136. package/dist/internal/cli-mcp-oauth.js +7 -4
  137. package/dist/internal/convert-automation/convert-workflow.d.ts.map +1 -1
  138. package/dist/internal/convert-automation/convert-workflow.js +26 -15
  139. package/dist/internal/convert-automation/slug.d.ts +0 -2
  140. package/dist/internal/convert-automation/slug.d.ts.map +1 -1
  141. package/dist/internal/convert-automation/slug.js +0 -8
  142. package/dist/internal/cursor/account-mcp.d.ts.map +1 -1
  143. package/dist/internal/cursor/account-mcp.js +5 -1
  144. package/dist/internal/discovery.d.ts.map +1 -1
  145. package/dist/internal/discovery.js +88 -13
  146. package/dist/internal/hosted-delivery.d.ts.map +1 -1
  147. package/dist/internal/hosted-delivery.js +22 -9
  148. package/dist/internal/mcp-endpoint.js +3 -3
  149. package/dist/internal/mcp-host.d.ts +8 -7
  150. package/dist/internal/mcp-host.d.ts.map +1 -1
  151. package/dist/internal/mcp-host.js +8 -7
  152. package/dist/internal/peer-connections.d.ts.map +1 -1
  153. package/dist/internal/peer-connections.js +5 -1
  154. package/dist/internal/playground/static.d.ts +0 -3
  155. package/dist/internal/playground/static.d.ts.map +1 -1
  156. package/dist/internal/resolved-connections.d.ts.map +1 -1
  157. package/dist/internal/resolved-connections.js +5 -7
  158. package/dist/internal/server.d.ts.map +1 -1
  159. package/dist/internal/server.js +113 -172
  160. package/dist/internal/session-engine.d.ts +45 -10
  161. package/dist/internal/session-engine.d.ts.map +1 -1
  162. package/dist/internal/session-engine.js +208 -65
  163. package/dist/internal/tool-catalog.d.ts +31 -0
  164. package/dist/internal/tool-catalog.d.ts.map +1 -0
  165. package/dist/internal/tool-catalog.js +67 -0
  166. package/dist/playground/assets/{index-D9MFzhNE.js → index-B3JCyigB.js} +1 -1
  167. package/dist/playground/index.html +1 -1
  168. package/dist/types.d.ts +72 -23
  169. package/dist/types.d.ts.map +1 -1
  170. package/dist/types.js +19 -0
  171. package/docs/README.md +0 -28
  172. package/docs/concepts.md +1 -0
  173. package/docs/deployment.md +1 -20
  174. package/docs/design/agsh.md +406 -0
  175. package/docs/guides/agent-to-agent.md +2 -2
  176. package/docs/guides/cloud-runtime.md +1 -0
  177. package/docs/guides/convert-automation.md +1 -1
  178. package/docs/guides/github.md +4 -4
  179. package/docs/guides/mcp-oauth.md +10 -18
  180. package/docs/guides/slack.md +10 -47
  181. package/docs/quickstart.md +2 -3
  182. package/docs/reference/cli.md +2 -1
  183. package/docs/reference/connections.md +15 -11
  184. package/docs/reference/hooks.md +2 -3
  185. package/docs/reference/http-api.md +8 -0
  186. package/docs/reference/project-layout.md +5 -1
  187. package/docs/reference/subagents.md +2 -2
  188. package/docs/reference/tools.md +19 -3
  189. package/docs/templates/pr-autofixer.md +4 -3
  190. package/docs/templates/security-reviewer.md +2 -3
  191. package/docs/troubleshooting.md +2 -2
  192. package/package.json +9 -2
  193. package/skills/create-agent/SKILL.md +6 -13
  194. package/skills/debug/SKILL.md +2 -4
  195. package/skills/evals/SKILL.md +1 -1
  196. package/skills/framework-map/SKILL.md +3 -2
  197. package/skills/mcp-auth/SKILL.md +10 -13
  198. package/skills/setup-slack/SKILL.md +21 -137
  199. package/src/channels/slack/attachments.ts +2 -2
  200. package/src/channels/slack/dispatch.ts +2 -16
  201. package/src/channels/slack/eval-directive.ts +8 -27
  202. package/src/channels/slack/index.ts +0 -6
  203. package/src/channels/slack/setup.ts +8 -15
  204. package/src/channels/slack/slack-channel.ts +14 -125
  205. package/src/channels/slack/types.ts +12 -96
  206. package/src/client.ts +23 -0
  207. package/src/connections.ts +20 -7
  208. package/src/index.ts +2 -0
  209. package/src/internal/advertise-tools.ts +45 -7
  210. package/src/internal/cli-mcp-oauth.ts +6 -4
  211. package/src/internal/convert-automation/convert-workflow.ts +29 -17
  212. package/src/internal/convert-automation/slug.ts +0 -9
  213. package/src/internal/cursor/account-mcp.ts +4 -1
  214. package/src/internal/discovery.ts +104 -13
  215. package/src/internal/fixtures/units-server.ts +52 -0
  216. package/src/internal/hosted-delivery.ts +60 -28
  217. package/src/internal/mcp-endpoint.ts +3 -3
  218. package/src/internal/mcp-host.ts +8 -7
  219. package/src/internal/peer-connections.ts +4 -1
  220. package/src/internal/playground/static.ts +1 -3
  221. package/src/internal/resolved-connections.ts +8 -10
  222. package/src/internal/server.ts +151 -251
  223. package/src/internal/session-engine.ts +254 -69
  224. package/src/internal/tool-catalog.ts +106 -0
  225. package/src/types.ts +90 -23
  226. package/templates/pr-autofixer/agent/channels/slack.ts +8 -2
  227. package/templates/triage/README.md +2 -1
  228. package/templates/triage/overlays/jira/agent/mcp-connections/tracker.ts +0 -1
  229. package/templates/triage/overlays/linear/agent/mcp-connections/tracker.ts +0 -1
  230. package/dist/channels/slack/cursor-account.d.ts +0 -87
  231. package/dist/channels/slack/cursor-account.d.ts.map +0 -1
  232. package/dist/channels/slack/cursor-account.js +0 -100
  233. package/dist/docs/assets/chunks/@localSearchIndexroot.ChpIC3Zy.js +0 -1
  234. package/dist/docs/assets/concepts.md.F6AiPorA.js +0 -1
  235. package/dist/docs/assets/example-agents_approval-buddy.md.DmezILPg.js +0 -10
  236. package/dist/docs/assets/example-agents_approval-buddy.md.DmezILPg.lean.js +0 -1
  237. package/dist/docs/assets/example-agents_benny.md.B0kwY7D_.js +0 -5
  238. package/dist/docs/assets/example-agents_benny.md.B0kwY7D_.lean.js +0 -1
  239. package/dist/docs/assets/example-agents_bugbot.md.BRGMi9O2.js +0 -11
  240. package/dist/docs/assets/example-agents_bugbot.md.BRGMi9O2.lean.js +0 -1
  241. package/dist/docs/assets/example-agents_codebase-wiki.md.BBNw9Ekr.js +0 -8
  242. package/dist/docs/assets/example-agents_codebase-wiki.md.BBNw9Ekr.lean.js +0 -1
  243. package/dist/docs/assets/example-agents_codeowners-review.md.Bfta-lBU.js +0 -8
  244. package/dist/docs/assets/example-agents_codeowners-review.md.Bfta-lBU.lean.js +0 -1
  245. package/dist/docs/assets/example-agents_concierge.md.BzB2b20R.js +0 -22
  246. package/dist/docs/assets/example-agents_concierge.md.BzB2b20R.lean.js +0 -1
  247. package/dist/docs/assets/example-agents_index.md.ChBp0AX6.js +0 -2
  248. package/dist/docs/assets/example-agents_index.md.ChBp0AX6.lean.js +0 -1
  249. package/dist/docs/assets/example-agents_knowledge-base.md.CrA85ig-.js +0 -11
  250. package/dist/docs/assets/example-agents_knowledge-base.md.CrA85ig-.lean.js +0 -1
  251. package/dist/docs/assets/example-agents_oncall.md.DK4XkYTd.js +0 -10
  252. package/dist/docs/assets/example-agents_oncall.md.DK4XkYTd.lean.js +0 -1
  253. package/dist/docs/assets/example-agents_security-reviewer.md.74pPpWYj.js +0 -19
  254. package/dist/docs/assets/example-agents_security-reviewer.md.74pPpWYj.lean.js +0 -1
  255. package/dist/docs/assets/example-agents_slack-agent.md.D7Kdj5BV.js +0 -5
  256. package/dist/docs/assets/example-agents_slack-agent.md.D7Kdj5BV.lean.js +0 -1
  257. package/dist/docs/assets/example-agents_weather-agent.md.CaGpmw3Y.js +0 -25
  258. package/dist/docs/assets/example-agents_weather-agent.md.CaGpmw3Y.lean.js +0 -1
  259. package/dist/docs/assets/guides_cloud-runtime.md.BnvjPiia.js +0 -9
  260. package/dist/docs/assets/guides_cloud-runtime.md.BnvjPiia.lean.js +0 -1
  261. package/dist/docs/assets/guides_slack.md.mqeNKs84.lean.js +0 -1
  262. package/dist/docs/assets/index.md.B-lVR4wT.js +0 -5
  263. package/dist/docs/assets/reference_http-api.md.C68BERYr.js +0 -11
  264. package/dist/docs/assets/reference_http-api.md.C68BERYr.lean.js +0 -1
  265. package/dist/docs/assets/reference_project-layout.md.WN9nwJht.js +0 -17
  266. package/dist/docs/assets/troubleshooting.md.vCWwvqcJ.js +0 -1
  267. package/dist/docs/example-agents/approval-buddy.html +0 -36
  268. package/dist/docs/example-agents/approval-buddy.md +0 -266
  269. package/dist/docs/example-agents/benny.html +0 -31
  270. package/dist/docs/example-agents/benny.md +0 -173
  271. package/dist/docs/example-agents/bugbot.html +0 -37
  272. package/dist/docs/example-agents/bugbot.md +0 -229
  273. package/dist/docs/example-agents/codebase-wiki.html +0 -34
  274. package/dist/docs/example-agents/codebase-wiki.md +0 -167
  275. package/dist/docs/example-agents/codeowners-review.html +0 -34
  276. package/dist/docs/example-agents/codeowners-review.md +0 -192
  277. package/dist/docs/example-agents/concierge.html +0 -48
  278. package/dist/docs/example-agents/concierge.md +0 -200
  279. package/dist/docs/example-agents/index.html +0 -28
  280. package/dist/docs/example-agents/index.md +0 -99
  281. package/dist/docs/example-agents/knowledge-base.html +0 -37
  282. package/dist/docs/example-agents/knowledge-base.md +0 -168
  283. package/dist/docs/example-agents/oncall.html +0 -36
  284. package/dist/docs/example-agents/oncall.md +0 -212
  285. package/dist/docs/example-agents/security-reviewer.html +0 -45
  286. package/dist/docs/example-agents/security-reviewer.md +0 -265
  287. package/dist/docs/example-agents/slack-agent.html +0 -31
  288. package/dist/docs/example-agents/slack-agent.md +0 -142
  289. package/dist/docs/example-agents/weather-agent.html +0 -51
  290. package/dist/docs/example-agents/weather-agent.md +0 -297
  291. package/dist/internal/cursor-slack-relay.d.ts +0 -96
  292. package/dist/internal/cursor-slack-relay.d.ts.map +0 -1
  293. package/dist/internal/cursor-slack-relay.js +0 -176
  294. package/docs/example-agents/approval-buddy.md +0 -271
  295. package/docs/example-agents/benny.md +0 -178
  296. package/docs/example-agents/bugbot.md +0 -234
  297. package/docs/example-agents/codebase-wiki.md +0 -172
  298. package/docs/example-agents/codeowners-review.md +0 -197
  299. package/docs/example-agents/concierge.md +0 -205
  300. package/docs/example-agents/index.md +0 -104
  301. package/docs/example-agents/knowledge-base.md +0 -173
  302. package/docs/example-agents/oncall.md +0 -217
  303. package/docs/example-agents/security-reviewer.md +0 -270
  304. package/docs/example-agents/slack-agent.md +0 -147
  305. package/docs/example-agents/weather-agent.md +0 -302
  306. package/src/channels/slack/cursor-account.ts +0 -202
  307. package/src/internal/cursor-slack-relay.ts +0 -249
  308. /package/dist/docs/assets/{concepts.md.F6AiPorA.lean.js → concepts.md.lwAgBIMI.lean.js} +0 -0
  309. /package/dist/docs/assets/{guides_agent-to-agent.md.B3JIaAqz.lean.js → guides_agent-to-agent.md.BDb0t1QV.lean.js} +0 -0
  310. /package/dist/docs/assets/{guides_convert-automation.md.Bboisykk.lean.js → guides_convert-automation.md.B4sjlodG.lean.js} +0 -0
  311. /package/dist/docs/assets/{guides_github.md.DqJhuaN1.lean.js → guides_github.md.Cnh2mL4a.lean.js} +0 -0
  312. /package/dist/docs/assets/{quickstart.md.BrmfrrIr.lean.js → quickstart.md.Nj_LjW_a.lean.js} +0 -0
  313. /package/dist/docs/assets/{reference_cli.md.D9KESDsD.lean.js → reference_cli.md.Cw6_ICYG.lean.js} +0 -0
  314. /package/dist/docs/assets/{reference_hooks.md.BxN87gCw.lean.js → reference_hooks.md.a8BJxMR5.lean.js} +0 -0
  315. /package/dist/docs/assets/{reference_project-layout.md.WN9nwJht.lean.js → reference_project-layout.md.Bv4KOtlB.lean.js} +0 -0
  316. /package/dist/docs/assets/{reference_skills.md.BFW9retM.lean.js → reference_skills.md.8son6Hjm.lean.js} +0 -0
  317. /package/dist/docs/assets/{reference_subagents.md.Xoav0AII.lean.js → reference_subagents.md.CfsIloPm.lean.js} +0 -0
  318. /package/dist/docs/assets/{troubleshooting.md.vCWwvqcJ.lean.js → troubleshooting.md.Ctv3T8C2.lean.js} +0 -0
@@ -500,6 +500,7 @@ name. For example, `agent/tools/get_weather.ts` creates a tool named
500
500
  | `agent/tools/<name>.ts` | Typed actions the model can call |
501
501
  | `agent/skills/*` | Procedures loaded when needed |
502
502
  | `agent/mcp-connections/<name>.ts` | Tools from external MCP servers |
503
+ | `agent/host-connections/<name>.ts` | Privileged MCP servers for host tools only |
503
504
  | `agent/channels/*.ts` | HTTP, Slack, and GitHub entry points |
504
505
  | `agent/ab.ts` or `agent/ab/*.ts` | Sticky variants and live performance metrics |
505
506
  | `evals/**/*.eval.ts` | Repeatable checks at the project root |
@@ -863,26 +864,7 @@ See the [GitHub guide](/docs/guides/github.md).
863
864
 
864
865
  ### Connect Slack
865
866
 
866
- Hosted Slack supports the team's Cursor Slack app or a dedicated Socket
867
- Mode app.
868
-
869
- Use the Cursor Slack app when mentions and direct messages are enough:
870
-
871
- ```ts
872
- import { slackChannel } from "@cursor/july/channels/slack";
873
-
874
- export default slackChannel({
875
- cursorAccount: true,
876
- agentName: "PrApprover",
877
- });
878
- ```
879
-
880
- The team must have the Cursor Slack app installed and Slack event relay
881
- access enabled. This mode needs no Slack token secrets. It doesn't
882
- support channel-post watches, tool approvals, or interactivity.
883
-
884
- Use a dedicated Socket Mode app for those features or a separate bot
885
- identity:
867
+ Hosted Slack uses a dedicated Socket Mode app:
886
868
 
887
869
  ```ts
888
870
  import { slackChannel } from "@cursor/july/channels/slack";
@@ -1105,6 +1087,417 @@ Continue with these pages:
1105
1087
 
1106
1088
  ---
1107
1089
 
1090
+ Source: /docs/design/agsh.md
1091
+
1092
+ # agsh: a shell for deployed agents
1093
+
1094
+ ## What this is
1095
+
1096
+ `agsh` (agent shell) is a standalone CLI that connects to one agent-sdk
1097
+ deployment and turns the agent's live tool surface into commands. Every tool
1098
+ the deployment can execute (authored server tools and tools provided by the
1099
+ agent's MCP connections) becomes a subcommand with a synopsis derived from its
1100
+ input schema, a man-page style `--help`, and a place in an interactive shell.
1101
+
1102
+ It is a separate binary and a separate package from `agent-sdk`. The
1103
+ `agent-sdk` CLI stays what it is today: the developer workflow tool for
1104
+ authoring, validating, deploying, and debugging agent projects. `agsh` is the
1105
+ operator's tool for working *inside* one deployed agent. The split also keeps
1106
+ heavy presentation dependencies (markdown rendering, syntax highlighting, the
1107
+ shell interpreter) out of `@cursor/july`, which ships to every agent project.
1108
+
1109
+ ## The experience
1110
+
1111
+ ```
1112
+ $ agsh help # list of commands, man-page style
1113
+ $ agsh read --help # man-page style: NAME, SYNOPSIS, DESCRIPTION, OPTIONS
1114
+ $ agsh read /repo/README.md
1115
+ $ agsh datadog_list_monitors --query "service:api"
1116
+ $ agsh # bare: interactive shell on a TTY, script from stdin otherwise
1117
+ ❯ ls /repo | grep -i readme
1118
+ ❯ read /repo/config.json | jq .version
1119
+ ```
1120
+
1121
+ Every invocation binds to the deployment's latest session by default, with
1122
+ `--session` and `--continuation-token` overrides, and prints the session
1123
+ identifier as a final stderr line.
1124
+
1125
+ ## Configuration
1126
+
1127
+ `agsh` is a client only; it never boots an agent. Every invocation needs a
1128
+ target deployment, given by flags or by environment variables. Flags always
1129
+ win over the environment.
1130
+
1131
+ Global command line options, accepted on every command and on the bare shell
1132
+ launch:
1133
+
1134
+ | Option | Environment default | Meaning |
1135
+ | --- | --- | --- |
1136
+ | `--target <url \| name>` | `AGENT_SHELL_TARGET` | The deployment to talk to: a URL is a local deployment (`http://127.0.0.1:39400/executor`), a name a production one (`change-monitor-executor`). |
1137
+ | `--team <team>` | `AGENT_SHELL_TEAM` | Team override for production resolution, when the login spans several. |
1138
+ | `--bearer-token <token>` | `AGENT_SHELL_BEARER_TOKEN` | Explicit bearer auth for a deployment that is not behind the Cursor login. |
1139
+ | `--session <id>` | | Bind to a specific session instead of the latest. |
1140
+ | `--continuation-token <token>` | | Bind by continuation token instead of session id. |
1141
+ | `--output <text\|json>` | | Result rendering: human-friendly views (default) or raw JSON. |
1142
+ | `-h`, `--help` | | Per-command help. |
1143
+
1144
+ One parameter carries the whole target selection, and the value's shape
1145
+ encodes the mode: a URL (`http://` or `https://`) targets a local
1146
+ deployment, anything else names a production one. Two options with a
1147
+ precedence rule would invite exactly the confusion a target selector must
1148
+ not have; with one parameter the only rule is that the flag beats the
1149
+ environment. A URL is self-contained down to the agent because one local
1150
+ agent-sdk serve process hosts every agent of the project (change-monitor's
1151
+ dev stack mounts `/executor` and `/planner` from a single port); a
1152
+ production deployment is a single agent, so its name is the complete
1153
+ address (`--team` narrows resolution when the login spans several).
1154
+ Authentication defaults to the stored Cursor login (the same engine-access
1155
+ credential agent-sdk uses); `--bearer-token` is the escape hatch for direct
1156
+ deployments. Session flags are per invocation and have no environment
1157
+ default: a session is state, not configuration. Color output follows the
1158
+ `NO_COLOR` convention and TTY detection; there is no agsh-specific color
1159
+ setting. No configuration file: one environment variable pins a working
1160
+ target for a terminal session
1161
+ (`AGENT_SHELL_TARGET=change-monitor-executor`, or a URL for a local stack),
1162
+ which is the whole persistent-configuration need.
1163
+
1164
+ With no target from flags or environment, every command fails with a message
1165
+ naming both ways to provide one.
1166
+
1167
+ ## Architecture
1168
+
1169
+ ### A new package
1170
+
1171
+ A new workspace package (working name `packages/agsh`, bin `agsh`) that
1172
+ depends on `@cursor/july` for target resolution, stored Cursor login, and the
1173
+ HTTP client plumbing. It owns the presentation stack: `marked` for terminal
1174
+ markdown (moved out of `@cursor/july`), with syntax highlighting (`shiki`)
1175
+ arriving in the phase that renders code; the shell interpreter is
1176
+ purpose-built (see Rationale).
1177
+ No new abstraction seam between the two packages; `agsh` imports what it
1178
+ needs until a second consumer justifies extracting a thin client.
1179
+
1180
+ ### The tool catalog
1181
+
1182
+ At startup `agsh` fetches one live catalog of everything invocable on the
1183
+ deployment. This is the piece the current `/v1/info` cannot provide: `/v1/info`
1184
+ projects the authored manifest, and connection tools only exist at runtime,
1185
+ resolved per session under the connection's auth. A new endpoint provides the
1186
+ live view (see Backend changes).
1187
+
1188
+ Catalog entries carry exactly one identifier each: the tool name exactly as
1189
+ the agent sees it. Authored server tools keep their authored name (`read`).
1190
+ Connection tools appear under their model-facing advertised name (the
1191
+ sanitized passthrough name from `advertise-tools.ts`, e.g.
1192
+ `datadog_list_monitors`). The CLI never invents a different naming format:
1193
+ a tool name copied from a session transcript is a valid `agsh` command, and
1194
+ vice versa. Where a tool came from — the upstream connector name when the
1195
+ tool declares one, the connection name otherwise — is a field on the
1196
+ catalog entry, not part of the identifier.
1197
+
1198
+ ### Two command tiers
1199
+
1200
+ Each catalog entry becomes a command, through one of two shapes:
1201
+
1202
+ **Curated commands for builtin tools.** The well-known tool names (`ls`,
1203
+ `read`, `grep`, `glob`, `diff`, ...) get hand-designed, POSIX-flavored
1204
+ command shapes, hardcoded in `agsh` next to their titles. These tools are
1205
+ what an operator types all day; their shapes should feel like the unix
1206
+ commands they mirror, not like generated bindings. A curated shape decides
1207
+ which schema fields are positional operands and which are flags, and every
1208
+ input has exactly one spelling: an operand is only an operand, never also a
1209
+ flag.
1210
+
1211
+ ```
1212
+ $ agsh read /repo/package.json --limit 2
1213
+ {
1214
+ "name": "change-monitor",
1215
+ → ses_a99d1b69c329eb75a2ec8603
1216
+
1217
+ $ agsh grep -i -A 2 toolEffect /repo/src
1218
+ src/tool-policy.ts:12:export type ToolEffect = "read" | "write";
1219
+ ...
1220
+ → ses_a99d1b69c329eb75a2ec8603
1221
+
1222
+ $ agsh ls /repo --ignore-globs '*.test.ts' --ignore-globs 'node_modules/**'
1223
+ ```
1224
+
1225
+ `read` takes its path as an operand mapped to the schema's `path` field, with
1226
+ `--offset` and `--limit` as integer flags. `grep` follows POSIX grep:
1227
+ `grep [options] <pattern> [path]`, with the rg-style options (`-i`, `-A`,
1228
+ `-B`, `-C`, `--output-mode`, `--head-limit`) mapping onto the schema fields
1229
+ of the same names (kebab-cased). `ls` shows array input: an array field's flag repeats once
1230
+ per element. A curated shape binds to the deployment's live schema at
1231
+ startup; when a deployment's tool lacks the expected field, the command
1232
+ degrades to the generic shape below rather than guessing.
1233
+
1234
+ A curated shape may also reformat the tool's text result toward the unix
1235
+ command's own output conventions: the VFS ls tool returns the model-facing
1236
+ tree (` - name/` rows under a header), and `agsh ls` prints it as standard
1237
+ ls does, one name per line with the trailing slash kept on directories. The
1238
+ tool's result string itself stays what the model sees; when a result does
1239
+ not match the expected shape it prints verbatim.
1240
+
1241
+ **Generated commands for MCP tools.** Connection tools are dynamically
1242
+ discovered, so no special treatment is possible; they get a uniform
1243
+ schema-derived mapping:
1244
+
1245
+ - Every schema property is accepted as one flag, spelled as the
1246
+ kebab-cased property name (`org_slug` → `--org-slug`) — the unix
1247
+ convention; kebab collisions gain a numeric suffix. Properties already
1248
+ shaped like flags (grep's `-i`) stay literal. No positionals, no other
1249
+ aliases.
1250
+ - Object-typed properties flatten recursively into one flag per leaf,
1251
+ dash-joined (`--telemetry-context` for `telemetry.context`), so every
1252
+ option reads as a plain value; a free-form object with no declared
1253
+ properties stays one JSON-valued flag. A leaf is required only when its
1254
+ whole ancestor chain is.
1255
+ - Values are coerced by schema type: booleans are valueless flags, numbers
1256
+ and integers are parsed, arrays accept the flag repeated once per element,
1257
+ enums are validated before the call.
1258
+
1259
+ ```
1260
+ $ agsh datadog_list_monitors --query "service:api" --limit 10
1261
+ ```
1262
+
1263
+ In both tiers `-h`/`--help` and the global target and session flags are
1264
+ reserved and injected, a flag that names no schema property fails before any
1265
+ request (listing the tool's actual properties), and the bound session prints
1266
+ as a final stderr line.
1267
+
1268
+ ### Result rendering
1269
+
1270
+ Raw JSON on a terminal is not an experience for people, so `--output=text`
1271
+ (the default) renders structured results through a small set of views,
1272
+ selected automatically by the shape of the value each call actually returned;
1273
+ tool metadata plays no part, since most tools advertise no output schema, and
1274
+ many return structured data as JSON text. A string result that parses as a
1275
+ JSON object or array counts as structured. An array of objects renders as a
1276
+ table (columns are the union of keys, missing cells stay blank, the table
1277
+ clamps to the terminal width); a single object renders as a property view
1278
+ (aligned keys, scalar lists as bullets, nested structures indented); an
1279
+ object that is nothing but an error wrapper renders as an `Error:` line;
1280
+ plain text prints verbatim. `--output=json` renders the structured value as
1281
+ raw JSON. The rendering never depends on the TTY: piped and interactive
1282
+ output carry the same content, only color follows TTY detection.
1283
+
1284
+ ### Help rendering
1285
+
1286
+ `--help` on a tool renders a man-page layout: NAME (the tool name, with the
1287
+ tool's `title` beside it when the catalog carries one; titles are curated
1288
+ data, never derived from the description), SYNOPSIS (operands from the
1289
+ curated shape; options never enumerate — they summarize as `[options...]`,
1290
+ man-page style, so the line stays bounded), DESCRIPTION (the tool
1291
+ description rendered as terminal markdown), OPERANDS (positional arguments,
1292
+ curated commands only), and OPTIONS. Descriptions of operands and options
1293
+ come from the schema's property descriptions. Effect and approval metadata
1294
+ render as notes when declared. Everything except the curated shape derives
1295
+ from `GET /v1/tools/:name`; nothing else is hand-written per tool.
1296
+
1297
+ ```
1298
+ $ agsh read --help
1299
+ NAME
1300
+ read - Read a file
1301
+
1302
+ SYNOPSIS
1303
+ read [options...] <path>
1304
+
1305
+ DESCRIPTION
1306
+ Reads a file from the local filesystem. This tool can also read image
1307
+ files when called with the appropriate path. Formats supported:
1308
+ jpeg/jpg, png, gif, webp.
1309
+
1310
+ OPERANDS
1311
+ <path>
1312
+ The absolute path of the file to read.
1313
+
1314
+ OPTIONS
1315
+ --offset <integer>
1316
+ The line number to start reading from. Positive values are 1-indexed
1317
+ from the start of the file. Negative values count backwards from the
1318
+ end. Only provide if the file is too large to read at once.
1319
+
1320
+ --limit <integer>
1321
+ The number of lines to read. Only provide if the file is too large
1322
+ to read at once.
1323
+
1324
+ NOTES
1325
+ Effect: read (performs no writes).
1326
+ ```
1327
+
1328
+ `agsh help` lists the available command names grouped by source, authored
1329
+ tools first, then one group per upstream connector (its name is the group
1330
+ header — one aggregating connection can host tools from several connectors,
1331
+ and the connector name is what an operator recognizes). Each row is the
1332
+ name, with the title beside it when the tool declares one; everything else
1333
+ lives behind the command's `--help`:
1334
+
1335
+ The agent's description renders as a DESCRIPTION section when the deployment
1336
+ declares one (`/v1/info` carries both name and description).
1337
+
1338
+ ```
1339
+ $ agsh help
1340
+ NAME
1341
+ change-monitor-executor
1342
+
1343
+ DESCRIPTION
1344
+ Executes monitoring plans against changed code.
1345
+
1346
+ COMMANDS
1347
+ diff Show workspace changes
1348
+ glob Find files by pattern
1349
+ grep Search file contents
1350
+ ls List a directory
1351
+ read Read a file
1352
+ report_change_issue
1353
+ report_change_succeeded
1354
+
1355
+ DATADOG
1356
+ datadog_list_monitors List monitors
1357
+ ...
1358
+
1359
+ Run any command with --help for its synopsis and options.
1360
+ ```
1361
+
1362
+ ### Shell mode
1363
+
1364
+ Invoked bare, `agsh` starts a shell. On a TTY this is a REPL; on a pipe it
1365
+ reads a script from stdin, so `echo 'ls /' | agsh` and here-docs work.
1366
+
1367
+ The interpreter is purpose-built and minimal: tokenizing (quotes, escapes),
1368
+ pipelines, and `;` / `&&` / `||`. The command namespace is exactly the
1369
+ deployment's tool catalog plus a small curated set of local pipe filters
1370
+ (`head`, `tail`, `wc`, stdin-filtering `grep`), so a tool name can never be
1371
+ shadowed. There is no local filesystem, no variables, no control flow: agsh
1372
+ has nothing local to operate on, and every command is a single traced
1373
+ `POST /v1/tools/:name` call.
1374
+
1375
+ The shell binds one session identity at launch (latest by default) and keeps
1376
+ it for the whole run, so a sequence of tool calls observes one consistent
1377
+ session context.
1378
+
1379
+ ## Backend changes on the agent-sdk runtime
1380
+
1381
+ Two read endpoints, mirroring the invocation path:
1382
+
1383
+ **`GET /v1/tools`: the live tool listing.** Returns the session's tool
1384
+ namespace exactly as a turn would assemble it: authored server tools plus the
1385
+ advertised passthrough tools synthesized from connections, under their
1386
+ model-facing names. Entries are light (name, source, and `title` when one is
1387
+ known); everything else lives behind the detail endpoint. Titles have two
1388
+ sources and no new authoring surface: connection tools inherit the upstream
1389
+ server's MCP title, which the host already propagates length-capped off
1390
+ listings; tools that do not come from MCP get theirs from a hardcoded
1391
+ name-to-title table in the runtime's endpoint implementation, covering the
1392
+ well-known tool names. A tool in neither place has no title. Accepts the same
1393
+ optional session binding as invocation (`session` or `continuationToken`)
1394
+ because advertised inventories can be tenant-scoped and resolved per session.
1395
+ Implementation reuses the existing plumbing: the discovered manifest for
1396
+ authored tools and the advertise-tools synthesis (`McpHost.listTools`, or the
1397
+ `oneOff` path when per-session auth substitution applies) for connection
1398
+ tools. This is not a duplicate of `/v1/info`: the info document stays the
1399
+ static authored manifest; the listing is the runtime view that only the
1400
+ running deployment can answer.
1401
+
1402
+ **`GET /v1/tools/:name`: one tool's full description.** Description, input
1403
+ schema, output schema when declared, effect when declared, approval
1404
+ requirement, and source connection. Same path as invocation
1405
+ (`POST /v1/tools/:name`), different method: GET describes what POST executes,
1406
+ for the same identifier.
1407
+
1408
+ Invocation needs no new naming scheme. Advertised connection tools are
1409
+ synthesized as ordinary server tools in the session's namespace, so
1410
+ `POST /v1/tools/:name` addresses them by their model-facing name like any
1411
+ authored tool, with the same session binding, policy checks, and per-call
1412
+ tracing. (The direct-call path did need the synthesis step added: it now
1413
+ resolves the advertised listing for the call's session identity when the
1414
+ authored lookup misses.)
1415
+
1416
+ Phase 1 ships the minimal runtime surface agsh calls: the `effect`
1417
+ projection in `/v1/info` (rendered in per-tool help), the scratch-workspace
1418
+ fallback on direct calls, and `continuationToken` binding on
1419
+ `POST /v1/tools/:toolName`. The detail endpoint in phase 2 also closes the
1420
+ output-schema gap; `/v1/info` stays as it is.
1421
+
1422
+ ## Local development loop
1423
+
1424
+ `factory/change-monitor` is the test bed. Its `pnpm start` already serves the
1425
+ planner and executor locally through the agent-sdk dev runtime
1426
+ (`agent-sdk serve --dir . --dev`). The loop:
1427
+
1428
+ 1. `cd factory/change-monitor && pnpm start` (local stack, both agents).
1429
+ 2. `agsh --target http://127.0.0.1:<port>/<agent>` against it, via a dev shim
1430
+ analogous to `agent-sdk-dev` so the CLI runs from the worktree.
1431
+ 3. Iterate end to end: VFS verbs (`ls`, `read`, `grep`, `glob`, `diff`) for the
1432
+ authored-tool path, and the planner's tenant connectors for the
1433
+ connection-tool path once `GET /v1/tools` exists.
1434
+
1435
+ ## Removing the inspector surface from agent-sdk
1436
+
1437
+ The inspector CLI is still on development branches, so nothing migrates: the
1438
+ CLI-side code is removed from `agent-sdk` and `agsh` is built in its place.
1439
+
1440
+ - The verb commands (`ls`, `read`, `grep`, `glob`, `diff`) become the
1441
+ curated tier: their hand-designed shapes, schema-binding logic (including
1442
+ the candidate-field fallback), and session binding carry over. The
1443
+ schema-to-argv flag mapping seeds the generated tier for MCP tools.
1444
+ - The `tools` and `skills` commands disappear entirely. `agsh help` and
1445
+ per-tool `--help` are the discovery surface.
1446
+ - `marked` and `shiki` leave `@cursor/july`; agsh's help rendering takes
1447
+ `marked`, and `shiki` returns when agsh ships syntax highlighting.
1448
+ `agent-sdk` keeps its developer workflow commands unchanged.
1449
+
1450
+ ## Plan
1451
+
1452
+ 1. **Package and core invocation.** Create the package, port target
1453
+ resolution, the schema-to-argv mapping, and help rendering from the
1454
+ inspector code. Authored tools only, against the existing endpoints.
1455
+ Verified end to end on the local change-monitor stack.
1456
+ 2. **Live catalog.** Add `GET /v1/tools` and `GET /v1/tools/:name` to
1457
+ the agent-sdk runtime, with the hardcoded title table for non-MCP tools, and verify
1458
+ direct invocation resolves advertised connection tools by their
1459
+ model-facing names. Connection tools appear as commands. Verified against
1460
+ the planner's connectors.
1461
+ 3. **Shell mode.** The purpose-built mini-shell: REPL on TTY, script on
1462
+ stdin, tools as the command namespace, one session per shell run.
1463
+ 4. **Cleanup.** Remove the inspector CLI surface and presentation
1464
+ dependencies from `@cursor/july`.
1465
+
1466
+ ## Rationale and rejected alternatives
1467
+
1468
+ **Why not extend `agent-sdk`.** The audiences differ: `agent-sdk` is for the
1469
+ person building and deploying an agent; this tool is for the person operating
1470
+ inside one. Bundling also forces every agent project to carry markdown
1471
+ rendering, syntax highlighting, and a bash interpreter it never uses.
1472
+
1473
+ **Name.** `agsh` reads as "agent shell", is four characters, collides with
1474
+ nothing common, and works as a shell prompt name. Considered: `august`
1475
+ (pairs with `july` but says nothing about purpose), `toolsh` (awkward to
1476
+ pronounce), `cursor-shell` (too broad; this is scoped to one agent).
1477
+
1478
+ **Why a REST catalog instead of the MCP endpoint.** The deployment already
1479
+ speaks MCP at `/v1/mcp/tools`, including a per-connection bridge, but the
1480
+ bridge is bound to an active turn and speaks JSON-RPC. The CLI wants a plain
1481
+ authenticated GET with session binding that returns the assembled tool
1482
+ namespace under the names the model sees. Wrapping that in MCP framing buys
1483
+ nothing for a first-party client.
1484
+
1485
+ **Why a purpose-built interpreter instead of just-bash.** just-bash was the
1486
+ original plan (a full bash emulation with a custom-command extension point),
1487
+ and a prototype disproved it: custom commands replace its coreutils but can
1488
+ never shadow its shell builtins, and `read`, `test`, `type`, and `help` are
1489
+ builtins — so the flagship `read` tool is unreachable, and the precedence is
1490
+ not ours to control (vercel-labs owns the package). No other embeddable JS
1491
+ shell interpreter has a workable custom-command story (mvdan-sh's JS build
1492
+ does not expose one; bash-parser is a parser only). agsh also needs almost
1493
+ none of bash: no local filesystem, no variables, no control flow — just
1494
+ tokenizing, pipelines, and a command namespace it fully owns. A TypeScript
1495
+ REPL with tools as async functions (the shape of change-monitor's `script`
1496
+ tool) was considered and kept as a possible later addition; it trades away
1497
+ the unix muscle memory the curated commands exist for.
1498
+
1499
+ ---
1500
+
1108
1501
  Source: /docs/evals.md
1109
1502
 
1110
1503
  # Evals
@@ -1294,2749 +1687,279 @@ Assert with the gates:
1294
1687
  | `t.eventOrder(matchers)` | matching event groups occur in this relative order |
1295
1688
  | `t.eventsSatisfy(label, predicate)` | your predicate over the typed event stream |
1296
1689
  | `t.check(value, expectation)` | any value, against a builder |
1297
- | `t.score(name, value)` | records a 0–1 score you computed; soft until you add a bar |
1298
- | `t.requireToolCall(name, matcher?)` | gates on a matching call and returns it, so later code can read its input and output |
1299
- | `t.requireInputRequest(filter?)` | gates on exactly one pending approval request and returns it |
1300
-
1301
- Every gate returns a handle: `.soft()` demotes it to tracked-only,
1302
- `.atLeast(0.7)` adds a soft score bar, and `.gate(0.8)` promotes a
1303
- scored assertion into a hard gate.
1304
-
1305
- With no matcher, `calledTool` is request-based: a requested call counts
1306
- even when its result has not arrived. Pass
1307
- `t.calledTool("inspect_pr", { status: "completed" })` to require the
1308
- call to return. `input`, `output`, and `count` matcher fields accept a
1309
- literal, a `RegExp`, or a predicate.
1310
-
1311
- The expectation builders are `includes(string | RegExp)`,
1312
- `equals(value)`, `matches(schema)`, `similarity(expected)`, and
1313
- `satisfies(predicate, label)`. `includes` stringifies its input,
1314
- `equals` compares values deeply, `matches` validates against a Standard
1315
- Schema (or anything with `safeParse`, like Zod), `similarity` scores
1316
- normalized text similarity, and `satisfies` runs your predicate. The
1317
- plain function `normalizedSimilarity(actual, expected)` returns the
1318
- same 0–1 score for use with `t.score`.
1319
-
1320
- A few more context members shape a case: `t.require(value, expectation)`
1321
- records a gate and stops the test body when it fails, without a
1322
- duplicate execution error. `t.skip(reason)` ends the case as skipped
1323
- (reported separately, never changes the exit code; call it before
1324
- sending messages). `t.metric(name, value)` records a structured score
1325
- for the playground case card. `t.log(message)` records a debug line for
1326
- the CLI and playground result.
1327
-
1328
- Three `t.send` options apply on session create (first `t.send` only):
1329
-
1330
- - `workspaceFiles`: `{ path: contents }`, seeded into the local session
1331
- workspace. Prefer this over machine-local paths.
1332
- - `workspaceDir`: absolute harness cwd (local runtime).
1333
- - `cloud`: per-session cloud options merged over the agent's static
1334
- `cloud` config (repos / env / …). Use a pinned `repos` override to
1335
- attach a fixture repo for cloud evals without putting it on the
1336
- agent's default `cloud.repos`. Cloud ignores `workspaceFiles` seeds.
1337
-
1338
- ```ts
1339
- const toolResults = t.events.filter((e) => e.type === "action.result");
1340
- t.check(
1341
- toolResults.length,
1342
- satisfies((n) => (n as number) <= 4, "at most 4 tool calls")
1343
- );
1344
- ```
1345
-
1346
- A case with no explicit gates falls back to whether at least one turn
1347
- completed successfully. Add `t.succeeded()` and behavior-specific gates
1348
- anyway. They make the contract visible during review.
1349
-
1350
- ### Judge free-form output
1351
-
1352
- When wording matters and no regex captures it, `t.judge` grades the
1353
- reply with an LLM. The built-in graders are `factuality(expected)`,
1354
- `summarizes(expected)`, `closedQA(criteria)`, and `sql(expected)`. Each
1355
- scores `t.reply` by default; pass `{ on }` to grade another value.
1356
-
1357
- ```ts
1358
- t.judge.factuality("It is 54°F in NYC right now.").atLeast(0.7);
1359
- ```
1360
-
1361
- Judge assertions are soft by default, so a judge never fails a build
1362
- until you give it a bar with `.atLeast(0.7)` or promote it with
1363
- `.gate(0.8)`. The judge model comes from `defineEvalConfig({ judge })`,
1364
- `defineEval({ judge })`, a case-level `judge`, or a per-call
1365
- `{ model }` override; the nearest one wins. For a domain-specific judge
1366
- whose verdict is not a single score, `t.judge.model(prompt)` sends a
1367
- raw prompt to the same model and returns the reply. You then record the
1368
- parsed result with `t.score` or `t.check`.
1369
-
1370
- ## Run evals from the CLI
1371
-
1372
- The `eval` command discovers, filters, and runs cases.
1373
-
1374
- Run the CLI under Node 22.13 or newer. Do not use Bun. Its HTTP/2 client
1375
- breaks tool-result streams and causes eval turns to fail.
1376
-
1377
- ```bash
1378
- agent-sdk eval --dir . --list # discover only
1379
- agent-sdk eval --dir . # run all
1380
- agent-sdk eval --dir . builds/checkout # one datapoint
1381
- agent-sdk eval --dir . builds search # several ids or prefixes
1382
- agent-sdk eval --dir . --tag smoke --tag pull-request # any matching tag
1383
- agent-sdk eval --dir . --json --no-stream # machine-readable results
1384
- agent-sdk eval --dir . --verbose # logs + reply snippets
1385
- ```
1386
-
1387
- Id filters use OR semantics. Each filter selects an exact id and its
1388
- descendants. For example, `builds` selects `builds`,
1389
- `builds/checkout`, and every other case below that path. Repeated tags
1390
- also use OR semantics. When you provide both ids and tags, a case must
1391
- match both groups.
1392
-
1393
- `eval` boots an ephemeral server on port 0 with a temp state root
1394
- outside the project, so cases don't inherit ambient monorepo rules and
1395
- don't write into the project state directory. Point `--url` at a running server to eval
1396
- a live agent instead:
1397
-
1398
- ```bash
1399
- agent-sdk eval --dir . \
1400
- --url http://127.0.0.1:3000/weather-agent \
1401
- --bearer-token "$AGENT_TOKEN"
1402
- ```
1403
-
1404
- The eval definitions still come from `--dir`; `--url` only changes the
1405
- agent that receives the turns. For a locally mounted multi-agent
1406
- directory, `--slug weather-agent` chooses the target. Use
1407
- `--state-root` to keep ephemeral session state at a chosen path,
1408
- `--timeout-ms` to override the project timeout, and `--no-stream` to
1409
- keep live progress off stderr. A TTY streams turn progress by default.
1410
- `--verbose` still writes `t.log` lines to stderr and adds reply snippets
1411
- to text results.
1412
-
1413
- Model turns need a Cursor credential from `agent-sdk login` or
1414
- `CURSOR_API_KEY`.
1415
-
1416
- See [CLI: eval](/docs/reference/cli.md#eval) for flags and exit codes.
1417
-
1418
- ### JSON results
1419
-
1420
- Use `--json --no-stream` in scripts and CI. The top-level result carries
1421
- the totals and one result per case:
1422
-
1423
- ```json
1424
- {
1425
- "ok": true,
1426
- "passed": 1,
1427
- "failed": 0,
1428
- "results": [
1429
- {
1430
- "id": "readiness",
1431
- "ok": true,
1432
- "assertions": [{ "name": "succeeded", "passed": true }],
1433
- "sessionId": "ses_123",
1434
- "inputs": ["Is checkout pull request 42 ready to approve?"],
1435
- "toolCalls": [{ "toolName": "inspect_pr", "isError": false }],
1436
- "logs": [],
1437
- "durationMs": 12340
1438
- }
1439
- ]
1440
- }
1441
- ```
1442
-
1443
- Each case result can also include `description`, `finalText`, `tools`,
1444
- `error`, and tool arguments or output. This shape lets CI report the
1445
- failed assertion without parsing terminal text.
1446
-
1447
- ## Run evals in the playground
1448
-
1449
- Start the server, open the playground, and choose **Evals**. You can run
1450
- every case or one case, watch progress, and open the resulting session
1451
- trace. The Evals tab works on a normal `serve`.
1452
-
1453
- ```bash
1454
- agent-sdk serve --dir .
1455
- ```
1456
-
1457
- Playground runs target the live server instead of an ephemeral one.
1458
- Their sessions appear in the session list. One eval batch can run at a
1459
- time. Persistence follows the rule under
1460
- [Configure eval runs](#configure-eval-runs). See
1461
- [Playground eval routes](/docs/reference/http-api.md#playground-eval-routes).
1462
- The start request returns `202` while cases run in the background.
1463
- Poll until the snapshot status becomes `completed`, `failed`, or `cancelled`.
1464
- Configuration errors appear on a failed snapshot.
1465
-
1466
- On `--prod` / `--url`, the CLI prints the Eval ID as soon as the batch is
1467
- accepted (and a Playground deep link with `?view=evals&evalRunId=…`):
1468
-
1469
- ```bash
1470
- agent-sdk eval --prod --slug vulnerability-scanner --tag deepsec
1471
- # Eval ID: evalrun_…
1472
- # Cancel: agent-sdk eval cancel evalrun_… --prod --slug vulnerability-scanner
1473
- # Playground: https://…/playground?view=evals&evalRunId=evalrun_…
1474
-
1475
- agent-sdk eval cancel evalrun_… --prod --slug vulnerability-scanner
1476
- agent-sdk eval status evalrun_… --prod --slug vulnerability-scanner
1477
- ```
1478
-
1479
- ## What good cases assert
1480
-
1481
- Gate decisions and shape, not prose. Model wording varies run to run.
1482
- Tool choice, tool avoidance, and output structure are the stable
1483
- contract.
1484
-
1485
- 1. `t.succeeded()`: always, first.
1486
- 2. The tool decision: `calledTool` for the intended path,
1487
- `notCalledTool` for the likely wrong alternative. The pair is
1488
- stronger than either alone.
1489
- 3. Output shape: a regex for the contract (`/ready|blocked/i`, a JSON
1490
- marker, a findings-block fence), never exact sentences.
1491
- 4. For structured output, parse `t.reply` and check fields with
1492
- `satisfies` instead of substring-matching JSON.
1493
-
1494
- The common failure modes: asserting exact phrasing, packing more than
1495
- about five gates into one case (split it), and cases that depend on live
1496
- external state that drifts (pin the input; see fixtures).
1497
-
1498
- ## Pick fixtures by agent type
1499
-
1500
- The right fixture depends on the surface under test.
1501
-
1502
- | Agent surface | Fixture |
1503
- | --- | --- |
1504
- | Chat / domain assistant | A canonical prompt string, chosen once and frozen |
1505
- | Tool-heavy | Run `agent-sdk call <tool>` first to pin what the tool returns, then freeze the prompt that triggers it |
1506
- | GitHub webhook | `agent-sdk github replay <pr> --events '*' --dry-run --out fixtures/github` snapshots real payloads for offline replay ([GitHub guide](/docs/guides/github.md)) |
1507
- | PR reviewer with host preparation | Diff, metadata, and gold labels pinned to commit SHAs; keep any live PR matrix small |
1508
- | Workspace-dependent | `workspaceFiles` in `t.send` options, never developer-machine paths |
1509
-
1510
- Tag the fast, reliably passing core `smoke` and run `--tag smoke` in the
1511
- inner loop. Leave slow or flaky-prone cases untagged for explicit runs.
1512
-
1513
- ### Materialize API-backed fixtures
1514
-
1515
- An input that only points at external data, such as a pull request URL,
1516
- snapshot id, or pair of commit SHAs, is not self-contained. Fetch it
1517
- once and commit the rendered fixture before you expand the suite.
1518
-
1519
- 1. Save the diff, metadata, and labels under `fixtures/` at pinned
1520
- revisions.
1521
- 2. Seed those files with `workspaceFiles`, or read them from the fixture
1522
- directory.
1523
- 3. Assert decisions and output shape against the saved evidence.
1524
- 4. Keep a small `smoke` subset for any remaining live pipeline checks.
1525
-
1526
- Read committed fixtures with `@cursor/july/evals/loaders`: `loadJson`,
1527
- `loadJsonl`, and `loadYaml` resolve relative paths against the project
1528
- root the runner discovered, not the cwd the CLI was invoked from
1529
- (`resolveFixturePath` and `evalFixtureRoot` expose the same
1530
- resolution for other file formats).
1531
-
1532
- `maxConcurrency` limits parallel datapoints. It does not limit model or
1533
- API fan-out inside one datapoint. Materialized fixtures prevent a large
1534
- suite from exhausting provider and GitHub rate limits. The
1535
- [evals skill](https://github.com/cursor/cursor/blob/main/packages/agent-serve/skills/evals/SKILL.md) has the full fixture workflow.
1536
-
1537
- ## Keep improvements with regression evals
1538
-
1539
- Every [hillclimb](/docs/hillclimbing.md) round that keeps a change must land
1540
- an eval that would have failed before the change. If you can't express
1541
- the improvement as a gate (a `calledTool` shift, a bounded
1542
- `action.result` count, an output-shape regex), the improvement is
1543
- unverified, and it'll regress silently.
1544
-
1545
- The rule cuts the other way too: never weaken an existing gate to make a
1546
- round pass. That's the freeze line moving, and it turns your regression
1547
- suite into a list of checks that no longer protect anything.
1548
-
1549
- ## Compare variants on live traffic
1550
-
1551
- Use `defineAB` to compare variant metrics on live sessions. It is not a
1552
- test runner and has no `agent-sdk ab` command. Keep `defineEval` as the
1553
- regression ratchet. Eval sessions do not enroll or change live metrics.
1554
- See [Live A/B metrics](/docs/ab.md) for assignment, behavior, collection,
1555
- and inspection.
1556
-
1557
- ## What's next
1558
-
1559
- Continue with these pages:
1560
-
1561
- - [Live A/B metrics](/docs/ab.md): sticky variants and cumulative metrics
1562
- on live sessions
1563
- - [Hillclimbing](/docs/hillclimbing.md): the loop evals make trustworthy
1564
- - [Building agents with agents](/docs/building-with-agents.md): have a
1565
- coding agent write the first suite
1566
- - [GitHub guide](/docs/guides/github.md): deterministic webhook fixtures
1567
- with `github replay`
1568
- - [Sessions and streaming](/docs/reference/sessions.md): the events
1569
- `t.events` contains
1570
-
1571
- ---
1572
-
1573
- Source: /docs/example-agents/approval-buddy.md
1574
-
1575
- # Keep PR approval policy deterministic with Approval Buddy
1576
-
1577
- Approval Buddy approves eligible pull requests from a fixed roster and
1578
- declines every other request. GitHub still blocks self-approval when the stamp
1579
- identity authored the PR. Code decides eligibility. The model prepares
1580
- evidence, runs two specialist reviews, and passes their findings to the
1581
- approval tool without changing the policy decision.
1582
-
1583
- Use this example when an agent can make a judgment inside a workflow, but
1584
- authorization and the final side effect must stay in deterministic code.
1585
-
1586
- [Browse the Approval Buddy source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/approval-buddy/)
1587
-
1588
- ## Keep approval policy in code
1589
-
1590
- Approval Buddy draws three hard boundaries:
1591
-
1592
- - `prepare_review` and `approve_pr` re-read the live PR and apply the same
1593
- eligibility rules.
1594
- - Two subagents inspect prepared evidence, but their findings never grant or
1595
- block approval.
1596
- - Only `approve_pr` posts the GitHub review.
1597
-
1598
- A spoofed webhook, Slack message, or model claim can't add someone to the
1599
- buddy roster. The mutating tool checks the source of truth immediately before it
1600
- acts.
1601
-
1602
- ## Follow the intended stamp flow
1603
-
1604
- The root instructions ask the model to run this sequence for a qualifying PR:
1605
-
1606
- 1. A non-draft `pull_request` event arrives with action `opened`, `reopened`,
1607
- or `ready_for_review`.
1608
- 2. The GitHub channel checks its repository allowlist and starts a session.
1609
- 3. `turn.started` posts a pending commit status.
1610
- 4. The model calls `prepare_review`.
1611
- 5. Host code fetches the live PR. It checks the author, open state, merged
1612
- state, and draft state.
1613
- 6. A qualifying PR gets `pr/MANIFEST.md`, `pr/meta.json`, and
1614
- `pr/diff.patch` in the session workspace. Diffs above 2,000,000
1615
- characters are truncated and marked in metadata.
1616
- 7. The model calls both review subagents through the built-in `task` tool.
1617
- 8. It concatenates their contracted replies and calls `approve_pr`.
1618
- 9. `approve_pr` re-runs eligibility, posts an `APPROVE` review, and returns
1619
- the outcome.
1620
- 10. The channel posts a final commit status. A self-approval block also gets
1621
- a short timeline comment because no approval review can appear.
1622
-
1623
- Ineligible PRs skip evidence and subagents. The model still calls
1624
- `approve_pr` so the deterministic tool returns the formal decline reason.
1625
-
1626
- Steps 4 through 9 are prompt-driven. The channel doesn't enforce tool order
1627
- or prove both subagents ran, and `approve_pr` accepts missing findings. A
1628
- failed turn clears the pending status with a green non-blocking result without
1629
- approving the PR.
1630
-
1631
- ## Map the framework features
1632
-
1633
- | Capability | Source | Role |
1634
- | --- | --- | --- |
1635
- | Root agent and policy prompt | [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/approval-buddy/agent/agent.ts), [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/approval-buddy/agent/instructions.md) | Configure the local agent and describe orchestration order. |
1636
- | GitHub channel | [`agent/channels/github.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/approval-buddy/agent/channels/github.ts) | Filter wakes, lease GitHub access, and publish status events. |
1637
- | Slack channel | [`agent/channels/slack.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/approval-buddy/agent/channels/slack.ts) | Accept approval-bot stamp and qualification requests. |
1638
- | Server tools | [`agent/tools/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/approval-buddy/agent/tools/) | Prepare evidence, approve, list buddies, and search GIFs. |
1639
- | Deterministic policy | [`agent/lib/approve.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/approval-buddy/agent/lib/approve.ts), [`agent/lib/buddies.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/approval-buddy/agent/lib/buddies.ts) | Own the roster and live eligibility checks. |
1640
- | Review subagents | [`agent/subagents/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/approval-buddy/agent/subagents/) | Run deep audit and code-quality passes over the same evidence. |
1641
- | Storage | [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/approval-buddy/agent/storage.ts) | Persist sessions and events with `cursorHostedStorage`. See [Storage](/docs/storage.md). |
1642
- | Live A/B experiment | [`agent/ab.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/approval-buddy/agent/ab.ts) | Compare baseline responses with a concise, presentation-only treatment (`concise-results`). |
1643
- | Evals and unit tests | [`evals/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/approval-buddy/evals/), [`agent/lib/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/approval-buddy/agent/lib/) | Protect routing, output contracts, policy, and GitHub behavior. |
1644
-
1645
- There are no authored skills, MCP connections, schedules, reminders, hooks,
1646
- sandbox seeds, or tool approvals.
1647
-
1648
- ## Prepare credentials
1649
-
1650
- You need:
1651
-
1652
- - Node 22.13 or newer.
1653
- - An agent-runtime credential.
1654
- - GitHub access to read PRs, post reviews, create commit statuses, and
1655
- post the self-approval visibility comment.
1656
-
1657
- Optional GIF selection uses:
1658
-
1659
- - `GIPHY_API_KEY` or `APPROVAL_BUDDY_GIPHY_API_KEY`,
1660
- - `APPROVAL_BUDDY_STAMP_GIF`, or
1661
- - severity-specific `APPROVAL_BUDDY_STAMP_GIF_<LEVEL>` variables.
1662
-
1663
- If you enable Giphy in a hosted copy, declare its secret and
1664
- `api.giphy.com` egress.
1665
-
1666
- ## Validate without approving a PR
1667
-
1668
- ```bash
1669
- agent-sdk validate --dir examples/approval-buddy
1670
- agent-sdk info --dir examples/approval-buddy --json
1671
- ```
1672
-
1673
- List the deterministic roster:
1674
-
1675
- ```bash
1676
- agent-sdk call list_buddies \
1677
- --dir examples/approval-buddy \
1678
- --input '{}'
1679
- ```
1680
-
1681
- Set a known merged PR, then run the read-only precheck:
1682
-
1683
- ```bash
1684
- MERGED_PR_URL=https://github.com/your-org/your-repo/pull/123
1685
- agent-sdk call prepare_review \
1686
- --dir examples/approval-buddy \
1687
- --input "{\"prUrl\":\"$MERGED_PR_URL\"}"
1688
- ```
1689
-
1690
- The result should decline because the PR is no longer open. `prepare_review`
1691
- never posts an approval.
1692
-
1693
- > [!CAUTION]
1694
- > Don't use `agent-sdk call approve_pr` as a smoke test. The tool has no
1695
- > `needsApproval` gate and posts a real GitHub review when the PR qualifies.
1696
-
1697
- ## See why preparation is separate
1698
-
1699
- `prepare_review` is read-only. It checks policy before fetching a large diff,
1700
- so declined requests don't spend review-agent work.
1701
-
1702
- Direct calls return the evidence file map because their scratch workspace is
1703
- deleted after the call. In-session calls write the tree to
1704
- `ctx.workspaceDir`, where both subagents can read it.
1705
-
1706
- `approve_pr` repeats the live check instead of trusting preparation. A PR can
1707
- close, merge, become a draft, or change author-related context between the two
1708
- steps. Revalidation keeps the final write bound to current state.
1709
-
1710
- This is a reusable two-tool pattern:
1711
-
1712
- - a read-only tool prepares and explains the decision,
1713
- - a mutating tool repeats policy at the side-effect boundary.
1714
-
1715
- ## Fan out two review contracts
1716
-
1717
- The two discovered subagents have different contracts:
1718
-
1719
- - The security reviewer reports bugs, breaking changes, and security findings
1720
- with `High`, `Medium`, or `Low` tags.
1721
- - The code-quality reviewer reports maintainability and structure concerns
1722
- with `Blocker`, `Major`, or `Minor` tags.
1723
-
1724
- The parent calls both through the harness `task` tool. They inherit the root
1725
- agent's execution surface and read the same `pr/` workspace. The prompt asks
1726
- the parent not to rewrite either reply. The review body trims the combined
1727
- text and caps it at 16,000 characters.
1728
-
1729
- Findings are informational. A high-severity finding doesn't veto the stamp.
1730
- That policy is explicit in the root instructions and approval code.
1731
-
1732
- ## Trace GitHub channel behavior
1733
-
1734
- The channel uses `githubChannel` with:
1735
-
1736
- - a configured repository allowlist on the account-linked GitHub transport,
1737
- - a second optional `APPROVAL_BUDDY_REPOS` wake filter,
1738
- - `deliverReplies: false`,
1739
- - progress reactions disabled, and
1740
- - event handlers for turn start, `approve_pr` results, and failed turns.
1741
-
1742
- The source requests `contents-write`, even though the documented workflow
1743
- posts reviews, statuses, and comments. When adapting the example, start with
1744
- `pr-write` and opt up only if a tool must push code.
1745
-
1746
- Every terminal status is green by design. Declines and crashed turns are
1747
- informational, not merge-blocking. This is a product decision in the example,
1748
- not an Agent SDK default.
1749
-
1750
- A successful turn that never calls `approve_pr` leaves the pending status in
1751
- place. The channel clears it on `approve_pr` results and `turn.failed`, but
1752
- has no `turn.completed` fallback.
1753
-
1754
- `github replay` reaches the same channel and can post a real approval, status,
1755
- or comment. Use replay only against a repository and PR created for this
1756
- test.
1757
-
1758
- ## Use Slack for explicit requests
1759
-
1760
- Start the dev server:
1761
-
1762
- ```bash
1763
- agent-sdk dev examples/approval-buddy
1764
- ```
1765
-
1766
- Then ask through the signed-in account-linked Slack connection:
1767
-
1768
- > Would this PR qualify for a stamp?
1769
-
1770
- The instructions route qualification questions to `prepare_review` only. A
1771
- stamp request runs the complete flow and may approve the PR.
1772
-
1773
- This channel uses the account-linked transport instead of a dedicated Socket
1774
- Mode app.
1775
-
1776
- ## See how durable storage fits
1777
-
1778
- `defineStorage` replaces the default local session store with a shared,
1779
- durable key-value adapter. Approval Buddy chooses:
1780
-
1781
- - a 15-second write debounce,
1782
- - startup restoration for up to 200 sessions, and
1783
- - a 14-day restore window.
1784
-
1785
- That policy fits long-lived Slack threads and a small webhook fleet. The
1786
- security reviewer uses the same adapter with lazy restore, which fits its
1787
- shorter sessions.
1788
-
1789
- ## Run the regression suite
1790
-
1791
- List the four eval cases:
1792
-
1793
- ```bash
1794
- agent-sdk eval --dir examples/approval-buddy --list
1795
- ```
1796
-
1797
- The suite covers:
1798
-
1799
- - buddy-list routing,
1800
- - declining a merged PR,
1801
- - using only `prepare_review` for a qualification question, and
1802
- - the combined findings headings and severity format over seeded evidence.
1803
-
1804
- Run the safe qualification case:
1805
-
1806
- ```bash
1807
- agent-sdk eval \
1808
- --dir examples/approval-buddy \
1809
- qualify/merged-pr-question \
1810
- --json
1811
- ```
1812
-
1813
- The qualification case reads a live merged PR. The seeded format case shown
1814
- by `--list` uses a planted auth-bypass diff
1815
- and checks for a `task` call, both headings, and severity tags. It doesn't
1816
- prove both named subagents ran or whether their output reached `approve_pr`.
1817
- Unit tests under `agent/lib/` cover policy, self-approval handling, evidence
1818
- limits, status mapping, GIF selection, and severity parsing.
1819
-
1820
- ## Reuse the policy boundary
1821
-
1822
- Keep these properties when you replace the buddy policy:
1823
-
1824
- 1. Put authorization in typed code.
1825
- 2. Fetch the source of truth inside both prepare and mutate steps.
1826
- 3. Give the model evidence only after the request qualifies.
1827
- 4. Treat specialist findings as data, not authority.
1828
- 5. Keep the core domain mutation in one named tool. Treat channel status and
1829
- visibility writes as separate, audited effects.
1830
- 6. Add a human approval gate if your policy still needs operator consent.
1831
- 7. Test read-only routing separately from mutation.
1832
-
1833
- ## Where to go next
1834
-
1835
- - [GitHub](/docs/guides/github.md)
1836
- - [Tools](/docs/reference/tools.md)
1837
- - [Subagents](/docs/reference/subagents.md)
1838
- - [Storage](/docs/storage.md)
1839
- - [Slack](/docs/guides/slack.md)
1840
- - [Evals](/docs/evals.md)
1841
-
1842
- ---
1843
-
1844
- Source: /docs/example-agents/benny.md
1845
-
1846
- # Route Slack work through repository playbooks
1847
-
1848
- This agent is a Slack teammate for a product team. Mentions and direct
1849
- messages reach it through an account-linked transport. New top-level posts in
1850
- an allowlisted issue channel reach it through a dedicated Slack app, even
1851
- without a mention. The agent then selects a repository playbook for triage,
1852
- reproduction, fixes, reviews, on-call work, or design critique.
1853
-
1854
- Use this example when Slack is the intake surface and your durable procedures
1855
- already live as repository skills.
1856
-
1857
- [Browse the current playbook-router source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/benny/)
1858
-
1859
- ## Combine two Slack transports with repo skills
1860
-
1861
- The playbook router uniquely combines three decisions:
1862
-
1863
- - Two Slack transports serve different engagement modes.
1864
- - `local.cwd` keeps session workspaces inside the monorepo.
1865
- - Instructions route work to inherited repository playbooks
1866
- instead of authored `agent/skills/`.
1867
-
1868
- The result is a thin agent project over a mature procedure library.
1869
-
1870
- ## Follow an issue report
1871
-
1872
- 1. A teammate creates a top-level post in the allowlisted issue channel.
1873
- 2. The dedicated Socket Mode channel accepts the allowlisted channel.
1874
- 3. A 15-second debounce lets edits settle. Deleting the post during that
1875
- window cancels the dispatch.
1876
- 4. The Agent SDK creates a thread-scoped session and sends the report to the
1877
- playbook router.
1878
- 5. The instructions select the matching triage playbook.
1879
- 6. The harness finds the repository root, opens the inherited playbook, and
1880
- follows its procedure.
1881
- 7. The agent posts only in the source thread and reports the evidence it
1882
- gathered.
1883
-
1884
- Mentions and direct messages follow the same agent instructions. They don't
1885
- need the watched-channel path.
1886
-
1887
- ## Map the playbook router files
1888
-
1889
- | File | Purpose |
1890
- | --- | --- |
1891
- | [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/benny/agent/agent.ts) | Names the agent, selects its model, and points the harness at a project-local cwd so inherited playbooks load. |
1892
- | [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/benny/agent/instructions.md) | Defines engagement rules, evidence policy, and the playbook routing map. |
1893
- | [`agent/channels/slack.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/benny/agent/channels/slack.ts) | Handles account-linked mentions and direct messages. |
1894
- | [`agent/channels/slack-app.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/benny/agent/channels/slack-app.ts) | Runs the dedicated app and watches one allowlisted channel. |
1895
- | [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/benny/agent/storage.ts) | Persists sessions and events with `cursorHostedStorage`. |
1896
- | [`evals/evals.config.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/benny/evals/evals.config.ts) | Caps eval run concurrency. |
1897
- | [`evals/smoke.eval.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/benny/evals/smoke.eval.ts) | Checks the agent identity and expected triage route. |
1898
-
1899
- The playbook router authors no tools, MCP connections, subagents, schedules, hooks, A/B
1900
- experiments, or sandbox seeds.
1901
-
1902
- ## See why `local.cwd` matters
1903
-
1904
- The Agent SDK normally keeps an ephemeral `run` or `eval` workspace outside a
1905
- large monorepo. This prevents ancestor instruction and repository-rule files
1906
- from leaking into an unrelated agent.
1907
-
1908
- The playbook router needs the opposite. Its procedures live at the repository
1909
- root, so `agent.ts` points `local.cwd` at a harness directory under the
1910
- project. Each harness workspace is a child of that directory. Walking up
1911
- reaches the host repository and its inherited playbook directory.
1912
-
1913
- Those playbooks are inherited context. `agent-sdk info` reports zero authored
1914
- skills for the agent. Copying this project into another repository removes
1915
- its main procedures unless you copy or replace the skill library too.
1916
-
1917
- ## Connect both Slack paths
1918
-
1919
- The account-linked path needs an agent-runtime login and a connected Slack
1920
- account:
1921
-
1922
- ```bash
1923
- agent-sdk login
1924
- agent-sdk whoami
1925
- ```
1926
-
1927
- It routes explicit mentions without a dedicated Slack token on the host.
1928
-
1929
- For the watched-channel path, configure a dedicated Socket Mode app with:
1930
-
1931
- - subscribe to `message.channels` and `message.groups`,
1932
- - have an App-Level Token with `connections:write`, and
1933
- - be a member of the watched channel.
1934
-
1935
- Run `agent-sdk slack create --dir examples/benny --channel-posts` for a
1936
- dedicated Socket Mode app, then `agent-sdk slack doctor`.
1937
-
1938
- Missing dedicated-app tokens leave that channel idle. They don't stop the
1939
- account-linked channel.
1940
-
1941
- ## Validate and start the server
1942
-
1943
- ```bash
1944
- agent-sdk validate --dir examples/benny
1945
- agent-sdk info --dir examples/benny --json
1946
- agent-sdk dev examples/benny
1947
- ```
1948
-
1949
- The info output should show two Slack channels and no authored skill. That
1950
- combination confirms the example is using inherited playbooks.
1951
-
1952
- ## Exercise each engagement mode
1953
-
1954
- Test the explicit account-linked path by asking:
1955
-
1956
- > Which playbook would you use to triage a product UI bug?
1957
-
1958
- Test the dedicated app:
1959
-
1960
- 1. Create a top-level post in the allowlisted issue channel.
1961
- 2. Don't mention the bot.
1962
- 3. Wait for the debounce window.
1963
- 4. Confirm the agent replies in the post's thread.
1964
-
1965
- Thread replies don't trigger the proactive watch. Mentions still use Slack's
1966
- normal mention path. Bot-authored posts are ignored to prevent loops.
1967
-
1968
- The channel uses the default handler after filtering. It doesn't apply a
1969
- second code-level classifier, so every accepted top-level post spends a model
1970
- turn and reaches the prompt.
1971
-
1972
- ## Inspect thread continuity
1973
-
1974
- The Agent SDK keys Slack sessions by channel and thread timestamp. A follow-up in
1975
- the same thread resumes the conversation and workspace. A new top-level issue
1976
- gets a new session.
1977
-
1978
- This lets a playbook gather evidence over several turns without mixing two
1979
- reports. The playground shows both the account-linked and dedicated-app
1980
- sessions while the dev server runs.
1981
-
1982
- ## Run the smoke eval
1983
-
1984
- ```bash
1985
- agent-sdk eval --dir examples/benny --list
1986
- agent-sdk eval --dir examples/benny smoke --json
1987
- ```
1988
-
1989
- The case asks for the agent identity and the playbook used for issue triage.
1990
- It checks the configured identity and route label.
1991
-
1992
- This is a lexical smoke test. It doesn't prove Slack delivery, skill
1993
- selection, skill loading, procedure execution, or thread-only behavior. Add
1994
- fixture-backed evals around the playbooks when you reuse this design.
1995
-
1996
- ## Build a playbook-routed teammate
1997
-
1998
- Use this structure when your organization already has tested skills:
1999
-
2000
- 1. Put the playbooks under a stable repository path.
2001
- 2. Set `local.cwd` so harness workspaces can inherit that path.
2002
- 3. Write a short routing table in `instructions.md`.
2003
- 4. Use account-linked Slack for explicit requests.
2004
- 5. Add a dedicated app only for allowlisted proactive intake.
2005
- 6. Keep the channel allowlist narrow and debounce edited posts.
2006
- 7. Add an eval for every important request-to-playbook route.
2007
-
2008
- If the procedures should ship with the agent, put them under
2009
- `agent/skills/` instead. Authored skills appear in the manifest and travel
2010
- with the project.
2011
-
2012
- ## Where to go next
2013
-
2014
- - [Slack](/docs/guides/slack.md)
2015
- - [Agent config](/docs/reference/agent-config.md)
2016
- - [Skills](/docs/reference/skills.md)
2017
- - [Sessions and streaming](/docs/reference/sessions.md)
2018
- - [Evals](/docs/evals.md)
2019
-
2020
- ---
2021
-
2022
- Source: /docs/example-agents/bugbot.md
2023
-
2024
- # Review prepared pull-request evidence
2025
-
2026
- This GitHub-read-only reviewer uses host code to fetch the PR
2027
- with `gh` and `git`, builds a trimmed `pr/` evidence tree, then hands that tree
2028
- to the model. The model reads the diff, loads a review skill, and returns at
2029
- most three high-confidence findings.
2030
-
2031
- Use this example when the host should control evidence collection and the
2032
- model shouldn't browse or mutate the source repository.
2033
-
2034
- [Browse the current reviewer source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/bugbot/)
2035
-
2036
- ## Separate evidence preparation from review
2037
-
2038
- The reviewer separates preparation from judgment:
2039
-
2040
- - Host code owns GitHub and Git access.
2041
- - A server tool turns untrusted PR input into bounded workspace files.
2042
- - A custom channel seeds those files before the model starts.
2043
- - An on-demand skill defines the review procedure and output contract.
2044
- - The model returns chat text. No path posts a GitHub review.
2045
-
2046
- This architecture gives the model a purpose-built evidence package instead of
2047
- a checkout.
2048
-
2049
- ## Follow a review
2050
-
2051
- The custom HTTP path runs this sequence:
2052
-
2053
- 1. `POST /v1/channels/review/` receives a PR reference.
2054
- 2. The handler calls `prepare_pr` without a model turn.
2055
- 3. Host code reads PR metadata and the unified diff.
2056
- 4. It reuses a matching checkout, force-fetching the PR ref there when the
2057
- commit is missing. Without a matching checkout, it uses a temporary bare
2058
- cache.
2059
- 5. It creates `pr/MANIFEST.md`, `pr/meta.json`, `pr/diff.patch`, and selected
2060
- small files and rules.
2061
- 6. `send({ workspaceFiles })` creates the model session with that evidence.
2062
- 7. The model reads the manifest and diff, then loads `pr-review`.
2063
- 8. The channel returns session and playground URLs while the review streams.
2064
-
2065
- If a normal chat starts without evidence, the model can call `prepare_pr`
2066
- mid-turn. That form writes the same files into the active session workspace.
2067
-
2068
- ## Map the evidence-review files
2069
-
2070
- | File | Purpose |
2071
- | --- | --- |
2072
- | [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/agent/agent.ts) | Selects the local runtime and model. |
2073
- | [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/agent/instructions.md) | Requires diff-first review and confines model work to `pr/`. |
2074
- | [`agent/tools/prepare_pr.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/agent/tools/prepare_pr.ts) | Exposes host preparation as a typed server tool. |
2075
- | [`agent/lib/prepare-pr.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/agent/lib/prepare-pr.ts) | Parses PR references, runs `gh` and `git`, and builds the evidence map. |
2076
- | [`agent/channels/review.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/agent/channels/review.ts) | Provides the loopback-only prepare-and-send HTTP route. |
2077
- | [`agent/channels/slack.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/agent/channels/slack.ts) | Extracts PR references and prepares evidence for mentions and direct messages. |
2078
- | [`agent/skills/pr-review.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/agent/skills/pr-review.md) | Sets finding limits, severities, and the machine-readable review format. |
2079
- | [`agent/lib/log.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/agent/lib/log.ts) | Writes timing logs for the host tools to stderr. |
2080
- | [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/agent/storage.ts) | Persists sessions and events with `cursorHostedStorage`. |
2081
- | [`evals/evals.config.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/evals/evals.config.ts) | Caps eval run concurrency. |
2082
- | [`evals/review/smoke.eval.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/bugbot/evals/review/smoke.eval.ts) | Seeds fake evidence and checks the review path without GitHub. |
2083
-
2084
- There is no authored GitHub channel, MCP connection, subagent, schedule,
2085
- hook, A/B experiment, approval, or custom storage.
2086
-
2087
- ## Prepare the host
2088
-
2089
- You need:
2090
-
2091
- - Node 22.13 or newer.
2092
- - An agent-runtime credential for model turns and account-linked Slack.
2093
- - `gh` and `git` on `PATH`.
2094
- - `gh` access to the target PR.
2095
- - Network access to GitHub and a writable temporary directory.
2096
-
2097
- The preparer can prefer a configured local checkout. Its `origin` must match
2098
- the target repository. Otherwise the reviewer uses its bare cache. It never
2099
- checks out the PR into the serve host's working tree.
2100
-
2101
- ## Validate the surface
2102
-
2103
- ```bash
2104
- agent-sdk validate --dir examples/bugbot
2105
- agent-sdk info --dir examples/bugbot --json
2106
- ```
2107
-
2108
- The manifest should show one server tool, one skill, and two authored
2109
- channels.
2110
-
2111
- ## Inspect evidence without a model turn
2112
-
2113
- Call the preparation tool directly:
2114
-
2115
- ```bash
2116
- agent-sdk call prepare_pr \
2117
- --dir examples/bugbot \
2118
- --input '{"pr":"https://github.com/owner/repo/pull/123"}'
2119
- ```
2120
-
2121
- Direct tool calls use a scratch workspace removed after the call.
2122
- `prepare_pr` detects this path and returns the complete file map in its
2123
- result. In a model session, it writes the files and returns a smaller summary.
2124
-
2125
- The evidence builder applies explicit limits:
2126
-
2127
- | Evidence | Limit |
2128
- | --- | --- |
2129
- | Post-change file | 12,000 characters |
2130
- | One rule file | 8,000 characters |
2131
- | Combined rules | 12,000 characters |
2132
- | PR body in metadata | 2,000 characters |
2133
-
2134
- Large files remain visible in `diff.patch`. The manifest records which full
2135
- files or rules were omitted.
2136
-
2137
- The per-file limits aren't an aggregate context cap. Every changed file below
2138
- 12,000 characters can be included. The diff command has a 12 MiB output
2139
- buffer; a larger diff fails preparation instead of being truncated.
2140
-
2141
- ## Run the HTTP review path
2142
-
2143
- Start the server:
2144
-
2145
- ```bash
2146
- agent-sdk dev examples/bugbot
2147
- ```
2148
-
2149
- From another terminal:
2150
-
2151
- ```bash
2152
- curl -s -X POST \
2153
- http://127.0.0.1:3000/bugbot/v1/channels/review/ \
2154
- -H 'content-type: application/json' \
2155
- -d '{"pr":"https://github.com/owner/repo/pull/123"}'
2156
- ```
2157
-
2158
- The route returns `status: "started"`, a continuation token, and session and
2159
- playground URLs. Open the session URL to watch the model read the evidence and
2160
- produce findings.
2161
-
2162
- The channel declares `localDevStrict()`. Direct loopback callers can use it.
2163
- Proxy-forwarding headers and non-loopback hosts are rejected.
2164
-
2165
- Send a follow-up by passing the returned key:
2166
-
2167
- ```bash
2168
- curl -s -X POST \
2169
- http://127.0.0.1:3000/bugbot/v1/channels/review/ \
2170
- -H 'content-type: application/json' \
2171
- -d '{"pr":"owner/repo#123","key":"<continuation-token>"}'
2172
- ```
2173
-
2174
- The follow-up resumes the session without fetching a new evidence tree.
2175
-
2176
- ## Run the Slack path
2177
-
2178
- The account-linked Slack channel handles review-bot mentions and direct
2179
- messages:
2180
-
2181
- > Review https://github.com/owner/repo/pull/123
2182
-
2183
- Slack handlers don't receive the channel `callTool` helper. This example calls
2184
- the shared `preparePrReview` host function, then returns `workspaceFiles` in
2185
- the Slack message preparation result. The model sees the same evidence and
2186
- prompt as the HTTP path.
2187
-
2188
- If a message contains no PR reference, the handler asks for one. Thread
2189
- follow-ups keep the same session.
2190
-
2191
- ## See how the skill constrains review
2192
-
2193
- `pr-review.md` tells the model to:
2194
-
2195
- - read the manifest and unified diff first,
2196
- - open at most one supporting file or rules file when a hunk is ambiguous,
2197
- - avoid shell, network, `gh`, and `git`,
2198
- - report no more than three findings,
2199
- - keep each description under 120 words, and
2200
- - emit the machine-readable review contract.
2201
-
2202
- The root instructions set the evidence boundary. The skill holds the reusable
2203
- review procedure. Keeping those roles separate lets another agent reuse the
2204
- same skill with different intake channels.
2205
-
2206
- ## Run the fixture-backed eval
2207
-
2208
- ```bash
2209
- agent-sdk eval --dir examples/bugbot --list
2210
- agent-sdk eval --dir examples/bugbot review/smoke --json
2211
- ```
2212
-
2213
- The eval constructs a `PreparedPrReview`, seeds its file map through
2214
- `workspaceFiles`, and checks for at least one read call with no shell call. It
2215
- doesn't assert which evidence file was read or whether the skill loaded. It
2216
- accepts either a formatted review or a clean result.
2217
-
2218
- This case tests review behavior without GitHub credentials or network data.
2219
- Add fixtures with reachable bugs when you need stricter location and severity
2220
- checks.
2221
-
2222
- ## Keep the side-effect boundary clear
2223
-
2224
- The reviewer makes no remote GitHub writes. It doesn't author a GitHub channel and
2225
- doesn't call a review API. Host preparation does write session evidence and
2226
- force-update `refs/pull/<N>/head` in either its bare cache or a matching local
2227
- checkout when the commit is missing. Every result ends with a note saying no
2228
- GitHub review was posted.
2229
-
2230
- If you add publishing later, keep it in a separate tool. This preserves a
2231
- read-only preparation and review path safe to run in evals.
2232
-
2233
- ## Reuse the evidence handoff
2234
-
2235
- Use host-prepared workspaces when:
2236
-
2237
- - external APIs should stay off the model's tool surface,
2238
- - context needs hard size limits,
2239
- - the model should inspect a snapshot instead of a live checkout, or
2240
- - several channels need the same preparation.
2241
-
2242
- Return `workspaceFiles` from direct host preparation, write into
2243
- `ctx.workspaceDir` for mid-turn recovery, and encode the reading order in both
2244
- the manifest and a skill.
2245
-
2246
- ## Where to go next
2247
-
2248
- - [Webhooks and custom channels](/docs/guides/webhooks.md)
2249
- - [Tools](/docs/reference/tools.md)
2250
- - [Skills](/docs/reference/skills.md)
2251
- - [Slack](/docs/guides/slack.md)
2252
- - [Evals](/docs/evals.md)
2253
-
2254
- ---
2255
-
2256
- Source: /docs/example-agents/codebase-wiki.md
2257
-
2258
- # Build a feature wiki from merged pull requests
2259
-
2260
- Codebase wiki keeps a living, feature-organized wiki of a repository.
2261
- The GitHub channel acknowledges every closed pull request instantly,
2262
- fetches a compact digest on the host, and spends a model turn only on
2263
- merged PRs. The turn maps the change onto feature pages; a daily
2264
- schedule writes a digest of what changed and rebuilds the index. Chat
2265
- sessions answer codebase questions from the wiki with page citations.
2266
-
2267
- Use this project when documentation should accumulate from merges
2268
- instead of being regenerated from scratch. Use
2269
- [Knowledge base](/docs/example-agents/knowledge-base.md) when people should curate
2270
- organizational context through conversation.
2271
-
2272
- [Browse the codebase wiki source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/codebase-wiki/)
2273
-
2274
- ## Treat PRs as evidence and features as pages
2275
-
2276
- The wiki refuses to become a merge log:
2277
-
2278
- - The page tree is rigid: `index`, `features/<slug>`, and
2279
- `digests/<yyyy-mm-dd>`. The store rejects anything else, so the wiki
2280
- can't sprawl.
2281
- - The `feature-mapping` skill requires a `wiki_search` before every
2282
- write. A PR updates the page that owns its feature; a new page needs
2283
- a genuinely new feature; chores change nothing.
2284
- - Every touched page gets a dated changelog entry citing the PR
2285
- number, so each fact traces back to a merge.
2286
-
2287
- The wiki itself is markdown on the serve host, in a wiki directory by
2288
- default with a `CODEBASE_WIKI_DIR` override. Sessions are disposable;
2289
- the wiki is the durable state.
2290
-
2291
- ## Follow a merged PR
2292
-
2293
- 1. GitHub delivers `pull_request` with action `closed`. The channel
2294
- returns a task acknowledgement immediately.
2295
- 2. The task fetches the digest with the host `gh` CLI: title, body,
2296
- labels, changed files, and a bounded diff excerpt. No checkout.
2297
- 3. The webhook payload can't say whether the PR merged, so the host
2298
- checks `mergedAt` and skips abandoned PRs without a model turn.
2299
- 4. For merged PRs, the task starts the turn with `pr/DIGEST.md` seeded
2300
- through `workspaceFiles` and a `pr:<owner/repo#N>` continuation
2301
- token, so redeliveries resume instead of double-ingesting.
2302
- 5. The model follows `feature-mapping`: search, update or create
2303
- feature pages, add changelog entries, and refresh `index` when pages
2304
- were added.
2305
-
2306
- In chat, "ingest PR #123" runs the same flow through the `ingest_pr`
2307
- tool, which writes the digest into the active session workspace.
2308
-
2309
- ## Map the wiki files
2310
-
2311
- | File | Purpose |
2312
- | --- | --- |
2313
- | [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/agent.ts) | Selects the local runtime and model. |
2314
- | [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/instructions.md) | Splits the job into merge ingestion and wiki-cited Q&A. |
2315
- | [`agent/lib/wiki-store.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/lib/wiki-store.ts) | Enforces the rigid page tree and owns reads, writes, and search. |
2316
- | [`agent/lib/pr-digest.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/lib/pr-digest.ts) | Fetches PR metadata and diff, and formats `pr/DIGEST.md`. |
2317
- | [`agent/tools/ingest_pr.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/tools/ingest_pr.ts) | Exposes host digest preparation for chat-driven backfills. |
2318
- | [`agent/tools/wiki_read.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/tools/wiki_read.ts), [`wiki_search.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/tools/wiki_search.ts), [`wiki_write.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/tools/wiki_write.ts) | Read, search, and rewrite wiki pages. |
2319
- | [`agent/skills/feature-mapping.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/skills/feature-mapping.md) | Maps changes onto features and fixes the page and changelog shape. |
2320
- | [`agent/schedules/daily-digest.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/schedules/daily-digest.md) | Writes `digests/<date>`, rebuilds the index, and flags stale pages. |
2321
- | [`agent/channels/github.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/channels/github.ts) | Acknowledges closed PRs and starts merged-only ingest turns. |
2322
- | [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/agent/storage.ts) | Persists sessions and events with `cursorHostedStorage`. |
2323
- | [`evals/evals.config.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/evals/evals.config.ts) | Caps eval run concurrency. |
2324
- | [`evals/ingest.eval.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codebase-wiki/evals/ingest.eval.ts) | Gates ingest decisions against the wiki filesystem. |
2325
-
2326
- There is no MCP connection, subagent, hook, or A/B experiment.
2327
-
2328
- ## Prepare credentials and services
2329
-
2330
- You need:
2331
-
2332
- - Node 22.13 or newer.
2333
- - An agent-runtime credential for model turns.
2334
- - `gh` on `PATH` with read access to the PRs you ingest.
2335
-
2336
- The channel verifies webhook signatures when `GITHUB_WEBHOOK_SECRET` is
2337
- set and narrows repositories with
2338
- `CODEBASE_WIKI_REPOS=owner/repo,owner/other`. The agent never writes to
2339
- GitHub. Its only side effects are wiki files on the serve host.
2340
-
2341
- ## Validate the surface
2342
-
2343
- ```bash
2344
- agent-sdk validate --dir examples/codebase-wiki
2345
- agent-sdk info --dir examples/codebase-wiki --json
2346
- ```
2347
-
2348
- The manifest should report four server tools, one skill, one schedule,
2349
- and the authored GitHub channel.
2350
-
2351
- ## Ingest without webhook plumbing
2352
-
2353
- Replay a real merged PR as a closed delivery:
2354
-
2355
- ```bash
2356
- agent-sdk dev examples/codebase-wiki
2357
-
2358
- agent-sdk github replay https://github.com/owner/repo/pull/123 \
2359
- --dir examples/codebase-wiki --action closed
2360
- ```
2361
-
2362
- The reply is a 202 acknowledgement; the ingest continues in the task.
2363
- Watch the session in the playground, then open the wiki directory on
2364
- the serve host. Feature pages land under `features/`.
2365
-
2366
- Each ingested feature page carries an overview, a "How it works"
2367
- section, and a changelog line citing the PR. Deterministic digest
2368
- preparation works without a model turn:
2369
-
2370
- ```bash
2371
- agent-sdk call ingest_pr \
2372
- --dir examples/codebase-wiki \
2373
- --input '{"pr":"https://github.com/owner/repo/pull/123"}'
2374
- ```
2375
-
2376
- A PR closed without merging returns `merged: false` and a note telling
2377
- the model to change nothing.
2378
-
2379
- ## Run the daily digest
2380
-
2381
- The schedule fires at 07:00 UTC. Under `agent-sdk dev`, trigger it by
2382
- hand:
2383
-
2384
- ```bash
2385
- curl -s -X POST http://127.0.0.1:3000/codebase-wiki/v1/dev/schedules/daily-digest
2386
- ```
2387
-
2388
- The turn reads every feature changelog, writes
2389
- `digests/<today>` grouped by feature with PR citations, rebuilds
2390
- `index`, and reports one line per page it wrote. Entries dated today
2391
- always count; a digest only claims a quiet day when no entry qualifies.
2392
-
2393
- ## Run the evals
2394
-
2395
- ```bash
2396
- agent-sdk eval --dir examples/codebase-wiki --list
2397
- agent-sdk eval --dir examples/codebase-wiki ingest/update-existing
2398
- ```
2399
-
2400
- The cases seed a temp wiki through `CODEBASE_WIKI_DIR` and build
2401
- digests with the same formatter the channel uses, so they run without
2402
- GitHub or network access. The gates check the filesystem, not prose:
2403
- a new feature page lands on a new slug, a related PR updates the
2404
- existing page instead of duplicating it, an unmerged PR changes
2405
- nothing, and the daily pass writes a digest naming both seeded
2406
- features.
2407
-
2408
- ## Reuse the merge-ingestion pattern
2409
-
2410
- Copy this shape when events should accumulate into curated state:
2411
-
2412
- - Acknowledge webhooks with a task and decide host-side whether a
2413
- model turn is worth spending.
2414
- - Seed evidence through `workspaceFiles` so the model never fetches.
2415
- - Constrain the durable store's shape in code and its content in a
2416
- skill.
2417
- - Add a consolidation schedule so incremental writes stay coherent.
2418
-
2419
- ## Where to go next
2420
-
2421
- - [GitHub webhooks](/docs/guides/github.md)
2422
- - [Schedules](/docs/reference/schedules.md)
2423
- - [Tools](/docs/reference/tools.md)
2424
- - [Evals](/docs/evals.md)
2425
-
2426
- ---
2427
-
2428
- Source: /docs/example-agents/codeowners-review.md
2429
-
2430
- # Route PR reviews by code ownership
2431
-
2432
- Codeowners review gives each part of a codebase its own review. A
2433
- CODEOWNERS-style table maps changed paths to review areas; each area
2434
- has a markdown playbook with the team's rules for that domain; and one
2435
- `area-reviewer` subagent runs per routed area, in parallel. A billing
2436
- change gets the billing review, a migration gets the migration review,
2437
- and an author's personal style rides along as advisory notes. The lead
2438
- aggregates: approve only when every area approves.
2439
-
2440
- Use this project when review quality depends on domain-specific values
2441
- instead of one generic checklist.
2442
-
2443
- [Browse the codeowners review source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/codeowners-review/)
2444
-
2445
- ## Keep routing in code and judgment in playbooks
2446
-
2447
- The pipeline separates three concerns:
2448
-
2449
- - `reviews/REVIEWERS` routes. Host code matches every changed path
2450
- against the table; every matching rule applies, and unmatched paths
2451
- fall back to the `general` playbook. Routing is glob code with unit
2452
- tests, not model judgment.
2453
- - `reviews/<area>.md` judges. Each playbook is a severity-ordered rule
2454
- list the team owns: billing mandates integer cents and idempotent
2455
- webhooks, migrations forbid destructive DDL beside code changes,
2456
- background jobs demand idempotency and dead-letter paths.
2457
- - Subagents review. The lead reads nothing but the manifest and
2458
- routes; each `area-reviewer` reads one playbook plus its files' diff
2459
- hunks and returns a mechanical verdict: request changes on any High
2460
- finding or two Mediums.
2461
-
2462
- Personal styles extend the same mechanism. `reviews/people/<login>.md`
2463
- attaches automatically, as advisory notes, whenever that person authors
2464
- the PR. Adding an area or a style is a markdown file plus at most one
2465
- routing line.
2466
-
2467
- ## Follow a review
2468
-
2469
- 1. A PR arrives: a GitHub `pull_request` event, a chat message, or a
2470
- bundled fixture reference.
2471
- 2. `prepare_review` fetches metadata and the diff with the host `gh`
2472
- CLI, routes every changed file, and writes the `pr/` evidence tree:
2473
- `MANIFEST.md`, `ROUTES.md`, `diff.patch`, and a copy of each matched
2474
- playbook.
2475
- 3. The lead follows the `review-process` skill and issues one
2476
- `area-reviewer` delegation per routed area, plus one per personal
2477
- style, all in one step so they run in parallel.
2478
- 4. Each reviewer reads its playbook, reviews only its files, and
2479
- returns a verdict line with at most three findings.
2480
- 5. The lead aggregates per-area sections and the overall verdict:
2481
- APPROVE only when every non-advisory area approved.
2482
-
2483
- Nothing posts to GitHub. Verdicts live in the session; the
2484
- [Approval Buddy guide](/docs/example-agents/approval-buddy.md) shows how to wire a real
2485
- APPROVE and commit statuses on top of the same shape.
2486
-
2487
- ## Map the review files
2488
-
2489
- | File | Purpose |
2490
- | --- | --- |
2491
- | [`reviews/REVIEWERS`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/reviews/REVIEWERS) | Routes path patterns to review areas. |
2492
- | [`reviews/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/codeowners-review/reviews/) | Holds the area playbooks and `people/<login>.md` styles. |
2493
- | [`agent/lib/routing.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/agent/lib/routing.ts) | Parses the table, matches globs, and unions areas per file. |
2494
- | [`agent/lib/prepare-review.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/agent/lib/prepare-review.ts) | Fetches PRs or fixtures and builds the evidence tree. |
2495
- | [`agent/tools/prepare_review.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/agent/tools/prepare_review.ts) | Exposes host preparation as a typed server tool. |
2496
- | [`agent/tools/list_review_areas.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/agent/tools/list_review_areas.ts) | Answers routing questions deterministically. |
2497
- | [`agent/skills/review-process.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/agent/skills/review-process.md) | Fixes the fan-out procedure and the verdict rule. |
2498
- | [`agent/subagents/area-reviewer/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/codeowners-review/agent/subagents/area-reviewer/) | Defines the one-area, one-playbook reviewer contract. |
2499
- | [`agent/channels/github.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/agent/channels/github.ts) | Reviews opened, reopened, synchronized, and undrafted PRs. |
2500
- | [`fixtures/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/codeowners-review/fixtures/) | Ships two reviewable PRs with known planted findings. |
2501
- | [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/agent/storage.ts) | Persists sessions and events with `cursorHostedStorage`. |
2502
- | [`evals/evals.config.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/evals/evals.config.ts) | Caps eval run concurrency. |
2503
- | [`evals/review.eval.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/codeowners-review/evals/review.eval.ts) | Gates routing, fan-out, planted bugs, and verdicts. |
2504
-
2505
- There is no MCP connection, schedule, hook, A/B experiment, or custom
2506
- storage.
2507
-
2508
- ## Prepare credentials and services
2509
-
2510
- You need:
2511
-
2512
- - Node 22.13 or newer.
2513
- - An agent-runtime credential for model turns.
2514
- - `gh` on `PATH` with read access to real PRs you review. The bundled
2515
- fixtures need no network at all.
2516
-
2517
- The channel verifies webhook signatures when `GITHUB_WEBHOOK_SECRET` is
2518
- set and narrows repositories with
2519
- `CODEOWNERS_REVIEW_REPOS=owner/repo,owner/other`. Pushes re-review in
2520
- the same session through the `pr:<label>` continuation token.
2521
-
2522
- ## Validate the surface
2523
-
2524
- ```bash
2525
- agent-sdk validate --dir examples/codeowners-review
2526
- agent-sdk info --dir examples/codeowners-review --json
2527
- ```
2528
-
2529
- The manifest should report two server tools, one skill, one subagent,
2530
- and the authored GitHub channel.
2531
-
2532
- ## Inspect routing without a model turn
2533
-
2534
- ```bash
2535
- agent-sdk call list_review_areas --dir examples/codeowners-review --input '{}'
2536
-
2537
- agent-sdk call prepare_review \
2538
- --dir examples/codeowners-review \
2539
- --input '{"pr":"fixture:multi-area"}'
2540
- ```
2541
-
2542
- The fixture routes to `billing`, `database-migrations`, and `frontend`,
2543
- attaches `people/alice` because alice authored it, and returns the full
2544
- evidence map. Point the same tool at a real PR URL and the routing runs
2545
- against the live file list. The example table maps a hypothetical
2546
- `src/` layout, so most real repositories route to `general` until you
2547
- adapt `reviews/REVIEWERS`.
2548
-
2549
- ## Review the planted fixture
2550
-
2551
- ```bash
2552
- agent-sdk dev examples/codeowners-review
2553
- ```
2554
-
2555
- In the playground:
2556
-
2557
- > Review fixture:multi-area
2558
-
2559
- The fixture plants one violation per area: float dollar math in
2560
- `src/billing/invoice.ts`, a `DROP COLUMN` plus a non-concurrent index
2561
- in the migration, and a clickable `div` without loading states in the
2562
- UI. The trace shows `prepare_review`, the evidence reads, four parallel
2563
- `area-reviewer` cards, and an aggregated CHANGES REQUESTED verdict with
2564
- each planted bug filed under its own area. The second fixture,
2565
- `fixture:jobs-clean`, routes to `background-jobs` alone and ends in
2566
- APPROVE.
2567
-
2568
- Review a real PR the same way:
2569
-
2570
- > Review https://github.com/owner/repo/pull/123
2571
-
2572
- Or replay one as a webhook delivery:
2573
-
2574
- ```bash
2575
- agent-sdk github replay https://github.com/owner/repo/pull/123 \
2576
- --dir examples/codeowners-review --action opened
2577
- ```
2578
-
2579
- ## See how the verdict stays mechanical
2580
-
2581
- The reviewer contract computes verdicts from findings instead of
2582
- letting the model pick a mood: findings first, then
2583
- `request-changes` if any High exists or two Mediums do, otherwise
2584
- `approve`. Pre-existing issues visible in context are scoped out, at
2585
- most one advisory Low. The lead applies one rule on top: the PR is
2586
- APPROVE only when every non-advisory area approved.
2587
-
2588
- ## Run the evals
2589
-
2590
- ```bash
2591
- agent-sdk eval --dir examples/codeowners-review --list
2592
- agent-sdk eval --dir examples/codeowners-review review/multi-area
2593
- ```
2594
-
2595
- `review/multi-area` gates the whole pipeline: `prepare_review` runs, at
2596
- least three subagent delegations happen, the reply carries every area
2597
- section plus alice's advisory notes, the planted billing and migration
2598
- bugs surface, and the verdict requests changes. `review/clean-approve`
2599
- proves the approval path on the clean fixture, and
2600
- `review/routing-question` gates that routing answers come from
2601
- `list_review_areas`.
2602
-
2603
- ## Reuse the ownership-routing pattern
2604
-
2605
- Copy this shape when different code deserves different judgment:
2606
-
2607
- - Route with data and code, not prompt instructions. Tables and globs
2608
- are testable.
2609
- - Write one playbook per domain and keep each reviewer blind to the
2610
- others.
2611
- - Make verdicts mechanical so aggregation is arithmetic, not
2612
- negotiation.
2613
- - Ship fixtures with planted findings so the review quality itself is
2614
- testable offline.
2615
-
2616
- ## Where to go next
2617
-
2618
- - [Approval Buddy](/docs/example-agents/approval-buddy.md) for posting real approvals
2619
- - [Subagents](/docs/reference/subagents.md)
2620
- - [GitHub webhooks](/docs/guides/github.md)
2621
- - [Evals](/docs/evals.md)
2622
-
2623
- ---
2624
-
2625
- Source: /docs/example-agents/concierge.md
2626
-
2627
- # Compose agents with a concierge
2628
-
2629
- Concierge answers general questions itself and sends every weather question
2630
- to the weather agent. The connection is one file. The Agent SDK turns the target
2631
- agent's MCP endpoint into tools the concierge can call.
2632
-
2633
- Use this example when two agents are useful on their own and one should
2634
- delegate a narrow class of work to the other.
2635
-
2636
- [Browse the Concierge source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/concierge/)
2637
-
2638
- ## Delegate through a peer MCP connection
2639
-
2640
- Concierge has no domain tool of its own. Its capability comes from a peer MCP
2641
- connection:
2642
-
2643
- ```ts
2644
- export default defineConnection({
2645
- agent: "weather-agent",
2646
- description:
2647
- "The weather-agent peer: delegate weather questions with ask; it runs its own tools (live Open-Meteo data) in its own context.",
2648
- });
2649
- ```
2650
-
2651
- The filename
2652
- [`weather.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/concierge/agent/mcp-connections/weather.ts)
2653
- makes the MCP server name `weather`. The `agent` field points to the sibling
2654
- project's mount slug.
2655
-
2656
- This differs from a subagent. A peer keeps its own:
2657
-
2658
- - root instructions,
2659
- - tools and MCP connections,
2660
- - durable sessions,
2661
- - playground, and
2662
- - public MCP endpoint.
2663
-
2664
- An SDK subagent inherits the parent's execution surface and only its parent
2665
- can invoke it. See [Agent-to-agent](/docs/guides/agent-to-agent.md) for the full
2666
- comparison.
2667
-
2668
- ## Follow a delegated request
2669
-
2670
- 1. A user asks Concierge what to pack for Paris.
2671
- 2. [`instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/concierge/agent/instructions.md)
2672
- classifies packing advice as weather-related.
2673
- 3. The model calls `weather.ask` with the city, timeframe, units, and the
2674
- complete question.
2675
- 4. The Agent SDK creates an MCP-channel session inside `weather-agent`.
2676
- 5. Weather agent calls its own Open-Meteo tools and returns a reply.
2677
- 6. If the turn exceeds the bounded MCP wait, `ask` returns
2678
- `status: "running"`. Concierge calls `weather.check` with the returned
2679
- `sessionId`.
2680
- 7. Concierge relays the result and may add one sentence of travel advice.
2681
-
2682
- The weather session appears in the weather agent's playground. It doesn't
2683
- share Concierge's conversation history.
2684
-
2685
- ## Map the delegation files
2686
-
2687
- | File | Purpose |
2688
- | --- | --- |
2689
- | [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/concierge/agent/agent.ts) | Describes the root agent and selects the local runtime. |
2690
- | [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/concierge/agent/instructions.md) | Draws a strict weather-only delegation boundary. |
2691
- | [`agent/mcp-connections/weather.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/concierge/agent/mcp-connections/weather.ts) | Resolves the peer by its `weather-agent` slug. |
2692
- | [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/concierge/agent/storage.ts) | Persists sessions and events with `cursorHostedStorage`. |
2693
-
2694
- Concierge doesn't author channels, tools, skills, subagents, schedules,
2695
- hooks, A/B experiments, or evals. The built-in HTTP and MCP surfaces still
2696
- exist.
2697
-
2698
- Its own MCP endpoint exposes `ask` and `check`. It doesn't expose
2699
- `call_tool` because Concierge has no server tools. The target weather agent
2700
- does expose `call_tool`, so that tool also appears under Concierge's
2701
- `weather` connection.
2702
-
2703
- ## Mount both agents
2704
-
2705
- A peer can only resolve within a multi-agent serve host. Validating Concierge
2706
- alone checks its files, but serving it alone fails because `weather-agent`
2707
- isn't mounted.
2708
-
2709
- From this package, validate both projects:
2710
-
2711
- ```bash
2712
- agent-sdk validate --dir examples/concierge
2713
- agent-sdk validate --dir examples/weather-agent
2714
- ```
2715
-
2716
- Don't serve the repository's whole `examples/` directory for this proof.
2717
- Several advanced examples subscribe to live GitHub events. Create an ignored
2718
- two-project mount instead. Copy only the authored files needed for this proof,
2719
- leaving Weather's Slack channels out:
2720
-
2721
- ```bash
2722
- PAIR_DIR=$(mktemp -d "${TMPDIR:-/tmp}/concierge-weather.XXXXXX")
2723
- mkdir -p "$PAIR_DIR/concierge" "$PAIR_DIR/weather-agent/agent"
2724
- cp -R examples/concierge/agent "$PAIR_DIR/concierge/"
2725
- cp examples/concierge/package.json "$PAIR_DIR/concierge/"
2726
- cp examples/weather-agent/agent/{agent.ts,instructions.md,ab.ts,ab.config.ts} \
2727
- "$PAIR_DIR/weather-agent/agent/"
2728
- cp -R examples/weather-agent/agent/{tools,skills,mcp-connections,subagents,schedules,hooks,lib} \
2729
- "$PAIR_DIR/weather-agent/agent/"
2730
- cp -R examples/weather-agent/mcp "$PAIR_DIR/weather-agent/"
2731
- cp examples/weather-agent/package.json "$PAIR_DIR/weather-agent/"
2732
- agent-sdk dev "$PAIR_DIR"
2733
- ```
2734
-
2735
- The host resolves the peer after it knows every mount. The local peer URL is
2736
- `http://127.0.0.1:3000/weather-agent/v1/mcp`. You still need an agent-runtime
2737
- credential for both model turns.
2738
-
2739
- ## Exercise delegation
2740
-
2741
- Send a weather request to the running Concierge:
2742
-
2743
- ```bash
2744
- agent-sdk chat \
2745
- --url http://127.0.0.1:3000/concierge \
2746
- --message "What should I pack for Paris tomorrow?"
2747
- ```
2748
-
2749
- Open both playgrounds:
2750
-
2751
- - `http://127.0.0.1:3000/concierge/playground`
2752
- - `http://127.0.0.1:3000/weather-agent/playground`
2753
-
2754
- The Concierge transcript shows the MCP call. The weather playground shows a
2755
- separate session on the `mcp` channel with live weather tool calls.
2756
-
2757
- Now send a general request:
2758
-
2759
- ```bash
2760
- agent-sdk chat \
2761
- --url http://127.0.0.1:3000/concierge \
2762
- --message "Give me three ideas for a quiet weekend."
2763
- ```
2764
-
2765
- The instructions tell Concierge to answer without delegating. This contrast is
2766
- the proof loop: weather goes to the peer, unrelated work stays local.
2767
-
2768
- ## Preserve peer context
2769
-
2770
- `weather.ask` returns a peer `sessionId`. Passing it back to a later `ask`
2771
- continues the same weather conversation. Concierge's instructions require
2772
- this for follow-ups dependent on an earlier answer.
2773
-
2774
- Use a fresh call when the tasks are independent. Reuse the peer session when
2775
- the second question needs facts or choices from the first.
2776
-
2777
- ## Keep delegation bounded
2778
-
2779
- The Agent SDK rejects unknown peer slugs and self-references during startup. It
2780
- doesn't stop a cycle across several valid peers. If agent A delegates all work
2781
- to B and B delegates all work to A, they can recurse.
2782
-
2783
- The prompt provides the guardrail here:
2784
-
2785
- - delegate every weather request,
2786
- - include complete context, and
2787
- - never delegate unrelated work.
2788
-
2789
- Write similarly narrow routing rules for each peer. A tool description helps
2790
- the model choose the connection, but the always-on instructions own the
2791
- policy.
2792
-
2793
- ## Use peers from cloud turns
2794
-
2795
- Local turns reach peers over loopback. A cloud VM can't reach the serve
2796
- host's loopback address. Set a public URL when a cloud agent needs the peer:
2797
-
2798
- ```bash
2799
- agent-sdk serve --dir "$PAIR_DIR" \
2800
- --public-url https://agents.example.com \
2801
- --bearer-token "$AGENT_TOKEN"
2802
- ```
2803
-
2804
- The Agent SDK attaches the bearer token to peer calls. Without `--public-url`,
2805
- cloud turns omit peer connections and the server logs a warning.
2806
-
2807
- ## Compose your own pair
2808
-
2809
- To compose your own agents:
2810
-
2811
- 1. Give each project a stable directory slug.
2812
- 2. Add `agent/mcp-connections/<name>.ts` to the caller.
2813
- 3. Set `agent` to the target slug.
2814
- 4. Describe the exact work the peer owns.
2815
- 5. Mount both projects from their parent directory.
2816
- 6. Add evals for delegated and non-delegated requests.
2817
-
2818
- Keep the peer independently useful. If the specialist only makes sense inside
2819
- one parent and needs no independent sessions, use a subagent instead.
2820
-
2821
- ## Where to go next
2822
-
2823
- - [Agent-to-agent](/docs/guides/agent-to-agent.md)
2824
- - [MCP connections](/docs/reference/connections.md)
2825
- - [Subagents](/docs/reference/subagents.md)
2826
- - [Sessions and streaming](/docs/reference/sessions.md)
2827
-
2828
- ---
2829
-
2830
- Source: /docs/example-agents/index.md
2831
-
2832
- # Choose the right Agent SDK example
2833
-
2834
- The examples progress from one-channel assistants to durable, event-driven
2835
- workflows. Start with the smallest agent for your use case. Each guide
2836
- explains its request flow, framework features, verification path, and reusable
2837
- design.
2838
-
2839
- The source projects live under
2840
- [`examples/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/). Run the commands below from
2841
- this package. See [Run the CLI](/docs/index.md#run-the-cli) if the
2842
- `agent-sdk` command isn't installed.
2843
-
2844
- ## Compare the examples
2845
-
2846
- | Agent | Runtime | Intake | Framework focus | What sets it apart |
2847
- | --- | --- | --- | --- | --- |
2848
- | [Weather agent](/docs/example-agents/weather-agent.md) | Cloud | HTTP and two Slack transports | Tools, stdio MCP, skill, subagent, schedule, hooks, A/B, and evals | It demonstrates the broad cloud-runtime surface in one domain. |
2849
- | [Slack agent](/docs/example-agents/slack-agent.md) | Local | Account-linked Slack | Channel identity, threads, and suggested prompts | It reaches Slack without authored tools. |
2850
- | [Concierge](/docs/example-agents/concierge.md) | Local | Built-in HTTP | Peer MCP and multi-agent serving | It delegates to a separate agent with its own tools, sessions, and context. |
2851
- | [Playbook router](/docs/example-agents/benny.md) | Local with repo context | Two Slack transports | Channel watching, inherited skills, custom cwd, and an eval | An allowlisted Slack channel becomes an intake queue for repo playbooks. |
2852
- | [Alert investigator](/docs/example-agents/oncall.md) | Local | Watched Slack alerts channel | Bot-post channel watching, per-thread debounce, reminder tools, and host Slack calls | Every alert gets a thread-pinned investigation that schedules its own re-checks. |
2853
- | [PR evidence reviewer](/docs/example-agents/bugbot.md) | Local | Custom HTTP and Slack | Host tool, skill, seeded workspaces, and an eval | The model receives a prepared diff-first evidence tree instead of a checkout. |
2854
- | [Approval Buddy](/docs/example-agents/approval-buddy.md) | Local | GitHub and Slack | Policy tools, two subagents, durable storage, and evals | Code decides whether a PR may be approved. Reviews stay informational. |
2855
- | [Security Reviewer](/docs/example-agents/security-reviewer.md) | Local host pipeline | GitHub and chat | Staged tools, parallel SDK agents, progress UI, durable storage, A/B, and evals | Reviewers and triage overlap while the playground shows every stage. |
2856
- | [Knowledge base](/docs/example-agents/knowledge-base.md) | Local | Built-in HTTP chat | Durable host-side state, a conventions skill, a schedule, unit tests, and evals | People curate shared facts in chat, and fresh sessions retrieve them from markdown. |
2857
- | [Codebase wiki](/docs/example-agents/codebase-wiki.md) | Local | GitHub and chat | Task-dispatch webhooks, seeded digests, a mapping skill, a schedule, and evals | Merged PRs accumulate into per-feature wiki pages with a daily digest. |
2858
- | [Codeowners review](/docs/example-agents/codeowners-review.md) | Local | GitHub, chat, and fixtures | Ownership routing in code, playbook data files, parallel subagents, and evals | Each product area reviews with its own playbook, and verdicts aggregate mechanically. |
2859
-
2860
- ## Pick a learning path
2861
-
2862
- Use this order when you want to learn the Agent SDK one capability at a time:
2863
-
2864
- 1. Start with [Weather agent](/docs/example-agents/weather-agent.md) to explore the filesystem
2865
- conventions and cloud runtime.
2866
- 2. Strip the project back to [Slack agent](/docs/example-agents/slack-agent.md) to see the
2867
- minimum channel surface.
2868
- 3. Read [Playbook router](/docs/example-agents/benny.md) when Slack should route requests into repo
2869
- playbooks.
2870
- 4. Continue to [Alert investigator](/docs/example-agents/oncall.md) when the intake is bot
2871
- posts and the agent must pace its own engagement and re-checks.
2872
- 5. Add composition with [Concierge](/docs/example-agents/concierge.md).
2873
- 6. Study [PR evidence reviewer](/docs/example-agents/bugbot.md) before giving a model repository
2874
- evidence.
2875
- 7. Move policy into code with [Approval Buddy](/docs/example-agents/approval-buddy.md).
2876
- 8. Study [Security Reviewer](/docs/example-agents/security-reviewer.md) for host-side PR
2877
- work.
2878
- 9. See parallel subagent delegation carry team judgment in
2879
- [Codeowners review](/docs/example-agents/codeowners-review.md).
2880
- 10. Curate team context through conversation with
2881
- [Knowledge base](/docs/example-agents/knowledge-base.md), then let GitHub events maintain
2882
- product documentation in [Codebase wiki](/docs/example-agents/codebase-wiki.md).
2883
-
2884
- ## Common prerequisites
2885
-
2886
- All examples require:
2887
-
2888
- - Node 22.13 or newer. Don't run the Agent SDK under Bun.
2889
- - Workspace dependencies installed.
2890
- - An agent-runtime credential for model turns.
2891
-
2892
- Several examples need more:
2893
-
2894
- - Account-linked Slack channels require a connected host account.
2895
- - Alert investigator needs a dedicated Socket Mode app with channel-post
2896
- events and membership in the watched alerts channel.
2897
- - GitHub examples require access to the target repository. Codebase wiki and
2898
- Codeowners review call the host `gh` CLI for PR data; the codeowners
2899
- fixtures run without network.
2900
- - Example agents use `cursorHostedStorage` in `agent/storage.ts` for hosted session storage. See [Storage](/docs/storage.md).
2901
-
2902
- Each guide lists its own credentials, services, and side effects.
2903
-
2904
- ## Validate any example
2905
-
2906
- Discovery commands don't start a model turn:
2907
-
2908
- ```bash
2909
- agent-sdk validate --dir examples/weather-agent
2910
- agent-sdk info --dir examples/weather-agent --json
2911
- ```
2912
-
2913
- Start one development server with `agent-sdk dev examples/<name>`.
2914
- Concierge depends on Weather agent, so its guide creates an isolated
2915
- two-project mount. Don't mount the whole examples directory to test one
2916
- agent; several advanced examples subscribe to live GitHub events.
2917
-
2918
- ## Read by framework feature
2919
-
2920
- - [Concepts](/docs/concepts.md) explains filesystem discovery and runtime
2921
- boundaries.
2922
- - [Project layout](/docs/reference/project-layout.md) lists every authored
2923
- folder.
2924
- - [Tools](/docs/reference/tools.md), [channels](/docs/reference/channels.md), and
2925
- [MCP connections](/docs/reference/connections.md) cover the core extension
2926
- points.
2927
- - [Evals](/docs/evals.md) and [live A/B metrics](/docs/ab.md) cover measured
2928
- iteration.
2929
- - [Deployment](/docs/deployment.md) covers credentials, auth, storage, and
2930
- hosting.
2931
-
2932
- ---
2933
-
2934
- Source: /docs/example-agents/knowledge-base.md
2935
-
2936
- # Build a team knowledge base through conversation
2937
-
2938
- Knowledge base turns conversations into shared team context. Teach the agent
2939
- about people, systems, decisions, and standing preferences. Three server
2940
- tools read, search, and write human-readable markdown pages; a conventions
2941
- skill shapes each write; and a daily schedule merges duplicates and rebuilds
2942
- the index. A fresh session retrieves what an earlier conversation captured.
2943
-
2944
- Use this project when people should curate organizational knowledge through
2945
- chat. Use [Codebase wiki](/docs/example-agents/codebase-wiki.md) when merged PRs should maintain
2946
- feature documentation instead.
2947
-
2948
- [Browse the knowledge base source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/knowledge-base/)
2949
-
2950
- ## Keep shared knowledge on the filesystem
2951
-
2952
- The knowledge base lives outside any session workspace, in a wiki
2953
- directory on the serve host by default. `KNOWLEDGE_BASE_DIR` overrides the location,
2954
- and the tools resolve it on every call, so tests and evals can point the same
2955
- code at a temp directory.
2956
-
2957
- The store enforces its own safety:
2958
-
2959
- - Page ids are one to three lowercase kebab-case segments, so a page id
2960
- can't escape the wiki directory.
2961
- - Pages cap at 64 KiB. Oversized writes fail with instructions to split
2962
- the page.
2963
- - `wiki_write` replaces whole pages. The instructions require reading a
2964
- page before updating it, so rewrites carry existing facts forward.
2965
-
2966
- Every page is plain markdown. You can open the wiki in an editor,
2967
- review it in a PR, or grep it.
2968
-
2969
- ## Follow a fact through the agent
2970
-
2971
- 1. You tell the agent something durable: a system, an owner, a standing
2972
- preference.
2973
- 2. The instructions require a `wiki_search` before claiming knowledge
2974
- and a `wiki_write` after learning something worth keeping.
2975
- 3. The `wiki-conventions` skill picks the page id (`staging-database`,
2976
- `people/jane-doe`), the page shape, and the dated fact format.
2977
- 4. The tool writes the page under the durable wiki root and returns
2978
- whether it created or updated the page.
2979
- 5. A later session, on any channel, finds the fact with `wiki_search`
2980
- and cites the knowledge-base page in its answer.
2981
-
2982
- Ephemeral chatter stays out. The instructions tell the model to skip
2983
- one-off questions and to ask before saving anything borderline.
2984
-
2985
- ## Map the knowledge-base files
2986
-
2987
- | File | Purpose |
2988
- | --- | --- |
2989
- | [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/agent.ts) | Selects the local runtime and model. |
2990
- | [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/instructions.md) | Sets the read-before-answer and save-after-learning policy. |
2991
- | [`agent/lib/wiki-store.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/lib/wiki-store.ts) | Validates page ids, lists, reads, writes, and searches the knowledge base. |
2992
- | [`agent/tools/wiki_read.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/tools/wiki_read.ts) | Reads one page or lists every page with titles and timestamps. |
2993
- | [`agent/tools/wiki_search.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/tools/wiki_search.ts) | Searches titles and bodies with per-page match lines. |
2994
- | [`agent/tools/wiki_write.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/tools/wiki_write.ts) | Creates or replaces a page and reports created versus updated. |
2995
- | [`agent/skills/wiki-conventions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/skills/wiki-conventions.md) | Names pages, shapes them, and dates every fact. |
2996
- | [`agent/schedules/gardener.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/schedules/gardener.md) | Merges duplicates, rebuilds the index, and flags stale facts daily. |
2997
- | [`agent/lib/wiki-store.test.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/lib/wiki-store.test.ts) | Unit-tests slug safety and store round-trips. |
2998
- | [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/agent/storage.ts) | Persists sessions and events with `cursorHostedStorage`. |
2999
- | [`evals/evals.config.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/evals/evals.config.ts) | Caps eval run concurrency. |
3000
- | [`evals/knowledge.eval.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/knowledge-base/evals/knowledge.eval.ts) | Seeds a temp knowledge base and gates recall, save, and no-write decisions. |
3001
-
3002
- There is no authored channel, MCP connection, subagent, hook, or A/B
3003
- experiment. The wiki directory is the durable knowledge store.
3004
-
3005
- ## Prepare the example
3006
-
3007
- You need:
3008
-
3009
- - Node 22.13 or newer.
3010
- - An agent-runtime credential for model turns.
3011
-
3012
- Nothing else. The wiki is created on first write.
3013
-
3014
- ## Validate the surface
3015
-
3016
- ```bash
3017
- agent-sdk validate --dir examples/knowledge-base
3018
- agent-sdk info --dir examples/knowledge-base --json
3019
- ```
3020
-
3021
- The manifest should report three server tools, one skill, and one
3022
- schedule.
3023
-
3024
- ## Exercise the store without a model turn
3025
-
3026
- ```bash
3027
- agent-sdk call wiki_write \
3028
- --dir examples/knowledge-base \
3029
- --input '{"page":"staging-database","content":"# Staging database\n\n- Port: 6432 (recorded 2026-07-19)\n"}'
3030
-
3031
- agent-sdk call wiki_search \
3032
- --dir examples/knowledge-base \
3033
- --input '{"query":"6432"}'
3034
-
3035
- agent-sdk call wiki_read --dir examples/knowledge-base --input '{}'
3036
- ```
3037
-
3038
- Invalid page ids fail fast. Try `{"page":"../escape"}` and the tool
3039
- returns the validation error instead of touching the filesystem.
3040
-
3041
- ## Prove recall across sessions
3042
-
3043
- ```bash
3044
- agent-sdk dev examples/knowledge-base
3045
- ```
3046
-
3047
- Teach it something in the playground:
3048
-
3049
- > Remember: our staging database is Postgres at
3050
- > staging-db.internal.example.com, port 6432 via PgBouncer. Jane Doe
3051
- > owns it.
3052
-
3053
- The trace shows the conventions skill load, then `wiki_write` calls
3054
- for `staging-database`, `people/jane-doe`, and `index`. Start a new
3055
- session and ask:
3056
-
3057
- > What port does our staging database use, and who owns it?
3058
-
3059
- The fresh session finds the answer with `wiki_search` and `wiki_read`
3060
- and cites the pages. The conversation history is empty; the wiki is the
3061
- source of truth.
3062
-
3063
- ## Run the gardener
3064
-
3065
- The `gardener` schedule fires at 06:00 UTC and rewrites the wiki for
3066
- consistency: merge near-duplicate pages, rebuild `index`, and flag
3067
- facts older than 90 days. Under `agent-sdk dev`, timers don't auto-fire.
3068
- Trigger it by hand:
3069
-
3070
- ```bash
3071
- curl -s -X POST http://127.0.0.1:3000/knowledge-base/v1/dev/schedules/gardener
3072
- ```
3073
-
3074
- ## Run the evals
3075
-
3076
- ```bash
3077
- agent-sdk eval --dir examples/knowledge-base --list
3078
- agent-sdk eval --dir examples/knowledge-base knowledge/recall
3079
- ```
3080
-
3081
- The eval file seeds a temp directory through `KNOWLEDGE_BASE_DIR`
3082
- inside the cases, so the durable knowledge base never sees test data.
3083
- `knowledge/recall` proves the fact comes from disk, not the conversation.
3084
- `knowledge/save`
3085
- gates the write decision, and `knowledge/no-write-on-ephemera` proves small
3086
- talk stays out of the knowledge base.
3087
-
3088
- ## Reuse the knowledge-base pattern
3089
-
3090
- Copy this shape when an agent needs durable, inspectable team knowledge:
3091
-
3092
- - Resolve the storage root lazily behind an environment override.
3093
- - Validate identifiers in the store, not in the prompt.
3094
- - Put naming and structure conventions in a skill so writes stay
3095
- consistent.
3096
- - Add a consolidation schedule instead of letting pages rot.
3097
-
3098
- ## Where to go next
3099
-
3100
- - [Tools](/docs/reference/tools.md)
3101
- - [Skills](/docs/reference/skills.md)
3102
- - [Schedules](/docs/reference/schedules.md)
3103
- - [Evals](/docs/evals.md)
3104
-
3105
- ---
3106
-
3107
- Source: /docs/example-agents/oncall.md
3108
-
3109
- # Investigate every alert in its own Slack thread
3110
-
3111
- This agent is an on-call teammate. Alert feeds post into an alerts channel
3112
- as bots. Each new alert dispatches an investigation session pinned to that
3113
- post's thread: the agent reacts 👀 the moment it locks in, investigates
3114
- immediately, and posts brief findings backed by evidence it observed.
3115
- Replies in the thread reach it only after the thread has been quiet for
3116
- about a minute, and reminder tools let it wake itself later to re-check a
3117
- baseline or confirm an alert cleared.
3118
-
3119
- Use this example when alerts land in Slack and you want one thread-scoped
3120
- investigation per alert, with an agent that paces its own engagement
3121
- instead of answering every message.
3122
-
3123
- [Browse the current alert-investigator source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/oncall/)
3124
-
3125
- ## Follow an alert
3126
-
3127
- 1. An alert feed (Alertmanager, PagerDuty, Datadog) posts a new top-level
3128
- message in the watched alerts channel.
3129
- 2. The channel watch accepts it. `includeBotPosts` lets bot authors
3130
- through; the agent's own posts always stay dropped.
3131
- 3. The handler reacts 👀 on the alert post and sets "Investigating…"
3132
- typing. The reaction is the lock-in signal: this alert has an owner.
3133
- 4. The Agent SDK creates a session keyed to the alert's thread and dispatches
3134
- immediately. New alerts get no debounce.
3135
- 5. The agent reads the alert, gathers evidence, and posts findings to the
3136
- thread once it has a hypothesis.
3137
- 6. People discuss in the thread. Replies buffer per thread and dispatch as
3138
- one coalesced follow-up after roughly a minute of quiet.
3139
- 7. The agent arms reminders for anything that needs time and posts interim
3140
- updates when new evidence changes the picture.
3141
-
3142
- Mentions and DMs skip the watch entirely and behave like ordinary chat.
3143
-
3144
- ## Map the files
3145
-
3146
- | File | Purpose |
3147
- | --- | --- |
3148
- | [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/agent.ts) | Names the agent and keeps harness workspaces outside any monorepo checkout. |
3149
- | [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/instructions.md) | Engagement rules, the investigation loop, and the message discipline. |
3150
- | [`agent/channels/slack-app.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/channels/slack-app.ts) | Dedicated Socket Mode app: watch configuration and handler wiring. |
3151
- | [`agent/lib/alert-watch.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/lib/alert-watch.ts) | The engagement policy: lock in on new alerts, coalesce replies. |
3152
- | [`agent/lib/thread-debounce.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/lib/thread-debounce.ts) | Per-thread quiet window. |
3153
- | [`agent/lib/alerts.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/lib/alerts.ts) | Dispatch classification, prompt building, and thread addressing. |
3154
- | [`agent/lib/slack-api.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/lib/slack-api.ts) | Reactions and thread posts on this agent's own token pair. |
3155
- | [`agent/tools/reminders_create.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/tools/reminders_create.ts) | Self-scheduled wakes bound to the thread (plus `reminders_list` and `reminders_cancel`). |
3156
- | [`agent/tools/post_thread_update.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/tools/post_thread_update.ts) | Interim updates to the thread mid-turn. |
3157
- | [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/storage.ts) | Persists sessions and events with `cursorHostedStorage`. |
3158
- | [`evals/evals.config.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/evals/evals.config.ts) | Caps eval run concurrency. |
3159
- | [`evals/smoke.eval.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/evals/smoke.eval.ts) | Checks identity and the reminder-tool route. |
3160
-
3161
- ## Let bot posts through the watch
3162
-
3163
- Channel watching drops bot-authored posts by default so two agents can
3164
- never feed each other. Alert channels invert the assumption: the posts
3165
- worth watching come from bots. `channelPosts.includeBotPosts` opts in per
3166
- channel:
3167
-
3168
- ```ts
3169
- engagement: {
3170
- channelPosts: {
3171
- allow: ["#alerts"],
3172
- posts: "all",
3173
- includeBotPosts: true,
3174
- },
3175
- },
3176
- ```
3177
-
3178
- Loop safety survives the opt-in. The pack matches the watching app's own
3179
- posts by the `bot_id` and bot user id from `auth.test` and drops them, so
3180
- the agent's findings never re-dispatch it. Posts that mention the bot stay
3181
- on the mention path.
3182
-
3183
- `posts: "all"` also delivers thread replies. The handler, not the pack,
3184
- decides their pace.
3185
-
3186
- ## Pace the engagement
3187
-
3188
- The example runs two rhythms:
3189
-
3190
- - A new alert dispatches immediately.
3191
- - Thread replies produce one engagement per lull.
3192
-
3193
- The pack's `debounceMs` is per message; it exists to let edits settle. This
3194
- agent needs a per-thread window instead, so the handler owns it
3195
- ([`lib/thread-debounce.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/oncall/agent/lib/thread-debounce.ts)).
3196
- Every reply restarts a 60-second timer keyed by thread. Superseded waiters
3197
- resolve `null` and the handler returns `null` for them. When the thread
3198
- goes quiet, the newest waiter receives the whole batch and dispatches one
3199
- follow-up that lists every message with mentionable attribution.
3200
-
3201
- Two details make the window matter. A follow-up that arrives while a turn
3202
- runs preempts that turn (latest message wins), so engaging per message
3203
- would keep cancelling the investigation. And @mentions bypass the window
3204
- through Slack's mention path, so a person who needs the agent now still
3205
- gets it now.
3206
-
3207
- ## Schedule your own re-checks
3208
-
3209
- Investigations rarely finish in one pass. A baseline comparison needs 20
3210
- minutes of data. An alert that cleared may re-fire. The example hands the
3211
- model three tools over `host.reminders`:
3212
-
3213
- - `reminders_create` arms a one-shot (`delay: "20m"`) or recurring
3214
- (`every: "30m"` with a plain-language stop condition) wake bound to the
3215
- thread's conversation.
3216
- - `reminders_list` shows the thread's standing watches.
3217
- - `reminders_cancel` disarms one, and refuses ids that belong to another
3218
- thread's conversation.
3219
-
3220
- When a reminder fires, its prompt returns to the same session as a
3221
- follow-up turn, and the reply lands in the alert thread. The instructions
3222
- keep wake prompts generic (re-read live state instead of replaying stale
3223
- numbers) and wake replies to one line, for example "re-checked p99 on
3224
- api-gateway: 120ms, back at baseline, cancelling the watch."
3225
-
3226
- Keep these tool filenames if you copy the design: the framework's reminder
3227
- fire prompt tells the model to call `reminders_cancel` by name when a stop
3228
- condition is set.
3229
-
3230
- ## Alert people mid-investigation
3231
-
3232
- The final reply of each turn posts to the thread on its own.
3233
- `post_thread_update` covers evidence that shouldn't wait for the turn to
3234
- finish: it posts a one-or-two-sentence update through the agent's token,
3235
- with `<@USERID>` mentions for the people who need to act. The instructions
3236
- restrict it to changes in hypothesis, severity, or blast radius. Progress
3237
- narration doesn't qualify.
3238
-
3239
- ## Connect the Slack app
3240
-
3241
- Channel watching is Socket Mode only, so this example uses a dedicated
3242
- app:
3243
-
3244
- ```bash
3245
- agent-sdk slack create --dir examples/oncall --name "Oncall" --channel-posts
3246
- agent-sdk slack doctor --prefix ONCALL
3247
- ```
3248
-
3249
- `--channel-posts` prefills channel-watch events (`message.channels` /
3250
- `message.groups`). Invite the bot to each watched channel after the
3251
- wizard finishes.
3252
-
3253
- `ONCALL_ALERTS_CHANNELS` sets the watch list as comma-separated ids or
3254
- `#names`. It defaults to `#alerts`.
3255
-
3256
- Wire observability MCP servers under `agent/mcp-connections/` so evidence
3257
- gathering reaches your logs, metrics, and dashboards. The example ships
3258
- none; without them the agent works from the alert text, its links, and the
3259
- thread.
3260
-
3261
- ## Validate and start the server
3262
-
3263
- ```bash
3264
- agent-sdk validate --dir examples/oncall
3265
- agent-sdk info --dir examples/oncall --json
3266
- agent-sdk dev examples/oncall
3267
- ```
3268
-
3269
- The info output lists four server tools and the watched channel on the
3270
- `slack-app` channel. Missing tokens leave that channel idle without
3271
- stopping the server.
3272
-
3273
- In dev mode, reminder timers don't auto-fire. List and fire them by hand
3274
- through the dev routes described in
3275
- [Schedules and reminders](/docs/reference/schedules.md#dispatch-and-dev-mode).
3276
-
3277
- ## Test the policy without Slack
3278
-
3279
- The engagement policy is plain code with unit tests:
3280
-
3281
- ```bash
3282
- pnpm exec vitest run examples/oncall
3283
- ```
3284
-
3285
- The integration test drives a synthetic Events API delivery through the
3286
- real parse, watch, and dispatch plumbing. It asserts a bot alert
3287
- dispatches pinned to its thread after the lock-in reaction, the agent's
3288
- own posts never loop, and replies coalesce behind the quiet window.
3289
-
3290
- The smoke eval spends a model turn:
3291
-
3292
- ```bash
3293
- agent-sdk eval --dir examples/oncall smoke --json
3294
- ```
3295
-
3296
- It checks identity and the reminder-tool route lexically. It doesn't prove
3297
- Slack delivery or reaction behavior; the unit tests cover the dispatch
3298
- side, and a live check needs the dedicated app connected.
3299
-
3300
- ## Build an alert investigator
3301
-
3302
- Use this structure when a bot feed should drive thread-scoped work:
3303
-
3304
- 1. Watch the feed channel with `includeBotPosts: true` and a narrow
3305
- allowlist.
3306
- 2. Acknowledge on the triggering post before dispatching, so people see
3307
- ownership without opening the thread.
3308
- 3. Dispatch new items immediately; coalesce thread chatter behind a
3309
- per-thread quiet window.
3310
- 4. Give the agent reminder tools for anything that needs time, and make
3311
- cancel discipline part of the instructions.
3312
- 5. Keep every posted message brief and tied to evidence the agent saw.
3313
-
3314
- ## Where to go next
3315
-
3316
- - [Slack](/docs/guides/slack.md)
3317
- - [Schedules and reminders](/docs/reference/schedules.md)
3318
- - [Tools](/docs/reference/tools.md)
3319
- - [Playbook router](/docs/example-agents/benny.md) for the human-post variant of channel
3320
- watching
3321
-
3322
- ---
3323
-
3324
- Source: /docs/example-agents/security-reviewer.md
3325
-
3326
- # Run staged security reviews from GitHub events
3327
-
3328
- Security Reviewer turns a pull request into a staged host-side review. One
3329
- tool prepares the diff and selects modules. A second fans out specialized
3330
- reviewers and triages candidates as they arrive. A third deduplicates the
3331
- confirmed findings, writes artifacts, and may publish a GitHub review.
3332
-
3333
- Use this example when the workflow needs several model workers, but the host
3334
- must own orchestration, progress, artifacts, and the final write.
3335
-
3336
- Source lives under [`factory/security-reviewer/`](https://github.com/cursor/cursor/tree/main/factory/security-reviewer/) (factory agent, not under `examples/`).
3337
-
3338
- Want one model turn and one comment? Scaffold the
3339
- [security-reviewer template](/docs/templates/security-reviewer.md).
3340
-
3341
- [Browse the Security Reviewer source.](https://github.com/cursor/cursor/tree/main/factory/security-reviewer/)
3342
-
3343
- ## Run a three-stage host pipeline
3344
-
3345
- Security Reviewer is a pipeline, not one long agent turn:
3346
-
3347
- | Stage | Tool | Result |
3348
- | --- | --- | --- |
3349
- | Prepare | `prepare_review` | Fetch metadata and diff, create a `runId`, and select security modules. |
3350
- | Review and triage | `run_reviewers` | Run module reviewers in parallel and start triage as each candidate arrives. |
3351
- | Finalize | `finalize_review` | Apply thresholds, deduplicate findings, write artifacts, and optionally post a review. |
3352
-
3353
- `run_triage` remains available as a compatibility stage. In the normal flow,
3354
- triage has already completed inside `run_reviewers`, so it reports existing
3355
- results. If candidates exist without triage output, it starts triage workers
3356
- and writes their state.
3357
-
3358
- The configured root agent chooses and sequences tools in chat. The review
3359
- workers use a model selected by the host pipeline. They are
3360
- created programmatically with the agent SDK, not discovered from
3361
- `agent/subagents/`.
3362
-
3363
- ## Follow a GitHub review
3364
-
3365
- 1. A pull request event starts a review and opens a playground session.
3366
- 2. The playground shows reviewer and triage progress.
3367
- 3. Confirmed findings appear in the PR review.
3368
- 4. A GitHub Check reports completion or a processing failure.
3369
-
3370
- ## Map the framework features
3371
-
3372
- | Capability | Source | Role |
3373
- | --- | --- | --- |
3374
- | Root agent | [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/factory/security-reviewer/agent/agent.ts), [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/factory/security-reviewer/agent/instructions.md) | Configure local chat and explain the three-stage contract. |
3375
- | Server tools | [`agent/tools/`](https://github.com/cursor/cursor/tree/main/factory/security-reviewer/agent/tools/) | Expose each review stage to chat and host orchestration. |
3376
- | GitHub channel | [`agent/channels/github.ts`](https://github.com/cursor/cursor/blob/main/factory/security-reviewer/agent/channels/github.ts) | Filter wakes, run background tasks, and publish status. |
3377
- | Progress channel | [`agent/channels/asr-progress.ts`](https://github.com/cursor/cursor/blob/main/factory/security-reviewer/agent/channels/asr-progress.ts) | Serve live reviewer and triage state by `runId`. |
3378
- | Playground renderer | [`agent/playground/tools/run_reviewers.tsx`](https://github.com/cursor/cursor/blob/main/factory/security-reviewer/agent/playground/tools/run_reviewers.tsx) | Replace the generic tool chip with live module rows. |
3379
- | SDK review pipeline | [`review-stages.ts`](https://github.com/cursor/cursor/blob/main/factory/security-reviewer/agent/lib/review-stages.ts), [`@anysphere/security-review-lib`](https://github.com/cursor/cursor/blob/main/packages/security-review-lib/src/index.ts) | Select modules, call model workers, triage, deduplicate, and write artifacts. |
3380
- | Storage | [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/factory/security-reviewer/agent/storage.ts) | Persist framework sessions with `cursorHostedStorage` (lazy restore). |
3381
- | A/B | [`agent/ab.ts`](https://github.com/cursor/cursor/blob/main/factory/security-reviewer/agent/ab.ts) | Compare all-severity versus high-only GitHub comments. |
3382
- | Eval | [`evals/`](https://github.com/cursor/cursor/tree/main/factory/security-reviewer/evals/) | Check stage-tool presence against a pinned sample. |
3383
-
3384
- There is no Slack channel, authored skill, discovered subagent, MCP
3385
- connection, schedule, reminder, hook, tool approval, or cloud runtime.
3386
-
3387
- ## Prepare the host
3388
-
3389
- You need:
3390
-
3391
- - Node 22.13 or newer.
3392
- - An agent-runtime credential for the root turn and review workers.
3393
- - GitHub read access for preparation.
3394
- - GitHub write access for webhook-driven reviews and Checks.
3395
-
3396
- The pipeline exposes settings for:
3397
-
3398
- - the worker model,
3399
- - reviewer and triage parallelism,
3400
- - reviewer, triage, duplicate-gate, and final-dedupe timeouts, and
3401
- - prior-comment loading.
3402
-
3403
- The active names live beside the orchestration in
3404
- [`review-stages.ts`](https://github.com/cursor/cursor/blob/main/factory/security-reviewer/agent/lib/review-stages.ts).
3405
-
3406
- ## Validate the discovered agent
3407
-
3408
- ```bash
3409
- agent-sdk validate --dir ../../factory/security-reviewer
3410
- agent-sdk info --dir ../../factory/security-reviewer --json
3411
- agent-sdk eval --dir ../../factory/security-reviewer --list
3412
- ```
3413
-
3414
- `validate` should pass. `info` and `eval --list` should match the capabilities
3415
- mapped above.
3416
-
3417
- ## Know the chat path's write boundary
3418
-
3419
- In chat, the root instructions ask the model to use this order:
3420
-
3421
- ```text
3422
- prepare_review -> run_reviewers -> finalize_review
3423
- ```
3424
-
3425
- They also ask the model to set `postComment: true` only on request. This is
3426
- prompt policy, not a deterministic safety gate. The model chooses tool
3427
- arguments, and `finalize_review` has no human approval. Use the direct stage
3428
- calls below when a no-post proof must be enforced.
3429
-
3430
- ## Call stages directly without publishing
3431
-
3432
- Call each stage and pass `postComment: false` yourself:
3433
-
3434
- ```bash
3435
- agent-sdk call prepare_review \
3436
- --dir ../../factory/security-reviewer \
3437
- --input '{"prUrl":"https://github.com/owner/repo/pull/123"}'
3438
-
3439
- agent-sdk call run_reviewers \
3440
- --dir ../../factory/security-reviewer \
3441
- --input '{"runId":"<run-id>"}'
3442
-
3443
- agent-sdk call finalize_review \
3444
- --dir ../../factory/security-reviewer \
3445
- --input '{"runId":"<run-id>","postComment":false}'
3446
- ```
3447
-
3448
- Review state lives under the project's run-artifact directory, so later
3449
- stages can open the prepared `runId`.
3450
-
3451
- > [!CAUTION]
3452
- > `finalize_review` with `postComment: true` writes to GitHub. The webhook
3453
- > path always requests that write. Chat instructions alone don't prevent it.
3454
-
3455
- ## Watch parallel work in the playground
3456
-
3457
- Run the dev server:
3458
-
3459
- ```bash
3460
- agent-sdk dev ../../factory/security-reviewer
3461
- ```
3462
-
3463
- Open the printed playground and start a review. The custom
3464
- `run_reviewers` renderer polls the progress channel's `GET /:runId` route.
3465
-
3466
- It refreshes every 500 ms while the stage runs. Each row shows a reviewer
3467
- module's state, candidates, reviewed areas, and failure. A second section
3468
- shows triage jobs and confirmed or rejected counts.
3469
-
3470
- This is an authored playground extension. The Agent SDK discovers it by the tool
3471
- name, so the generic `run_reviewers` chip becomes a domain-specific view
3472
- without changing the framework playground.
3473
-
3474
- ## Fan out reviewers while triage starts
3475
-
3476
- Module selection uses repository and path rules. The current module set
3477
- covers:
3478
-
3479
- - agent tooling trust boundaries,
3480
- - privileged service RPCs,
3481
- - product-specific security risks,
3482
- - dependency and supply-chain changes,
3483
- - deployment and infrastructure code,
3484
- - filesystem and workspace boundaries,
3485
- - privacy, and
3486
- - general security review.
3487
-
3488
- Selected modules may run more than once. Candidates pass through a duplicate
3489
- gate, then bounded triage. Reviewer or triage failures can produce partial
3490
- results. A final dedupe failure stops finalization.
3491
-
3492
- The pipeline writes JSONL journals as work completes. Final artifacts include
3493
- the review bundle, patch, reviewer outputs, candidates, triage decisions,
3494
- findings, accounting, and audit events.
3495
-
3496
- ## Separate session storage from review artifacts
3497
-
3498
- `cursorHostedStorage` keeps Agent SDK session and event records on
3499
- Cursor-managed hosting. Security Reviewer sets `restore: "off"` so startup
3500
- doesn't load old review sessions in bulk. A continuation lookup can still
3501
- fetch a needed session. See [Storage](/docs/storage.md).
3502
-
3503
- The staged review files are separate from session storage. Session-store
3504
- durability doesn't preserve those files. All stages for one `runId` must see
3505
- the same filesystem.
3506
-
3507
- This split is useful when conversation history needs shared durability but
3508
- large review artifacts belong on attached storage or an object store.
3509
-
3510
- ## Compare live comment variants
3511
-
3512
- The comment-severity experiment uses sticky session assignment with a 5%
3513
- holdout:
3514
-
3515
- - `control` posts every finding.
3516
- - `treatment` posts only high and critical findings.
3517
-
3518
- Finalization enforces the comment filter. The treatment also adds an
3519
- instruction overlay asking chat and playground summaries to lead with high
3520
- and critical findings. Full artifacts, `finalResponse`, and finding counts
3521
- still include every finding. Stage-tool counters appear in the
3522
- playground A/B view. Local sample and snapshot files persist under
3523
- the project state directory.
3524
-
3525
- When a treatment session has only low or medium findings, the filtered review
3526
- body currently says no vulnerabilities were found even though artifacts and
3527
- status retain findings. Account for that mismatch before using this
3528
- experiment as a publishing policy.
3529
-
3530
- Eval sessions skip A/B enrollment.
3531
-
3532
- ## Test the GitHub channel carefully
3533
-
3534
- The channel uses the host's Cursor account repository scope. It wakes on
3535
- `opened` and `synchronize`, skips drafts, and posts its own GitHub Check.
3536
-
3537
- Inspect its event surface:
3538
-
3539
- ```bash
3540
- agent-sdk github events \
3541
- --dir ../../factory/security-reviewer \
3542
- --json
3543
- ```
3544
-
3545
- Replay reaches the full publishing path:
3546
-
3547
- ```bash
3548
- TEST_PR_URL=https://github.com/your-org/allowlisted-test-repo/pull/123
3549
- agent-sdk github replay \
3550
- "$TEST_PR_URL" \
3551
- --dir ../../factory/security-reviewer \
3552
- --action opened
3553
- ```
3554
-
3555
- Set `TEST_PR_URL` to a PR in the channel's configured repository allowlist.
3556
- Run the command only against a PR intended for test reviews. It posts a
3557
- GitHub Check and may post findings.
3558
-
3559
- ## Inspect the eval before running it
3560
-
3561
- ```bash
3562
- agent-sdk eval --dir ../../factory/security-reviewer --list
3563
- ```
3564
-
3565
- The case gates the review flow and prevents comment posting. It still fetches
3566
- the live PR, so it needs GitHub access.
3567
-
3568
- ## Build another staged pipeline
3569
-
3570
- Use staged host orchestration when:
3571
-
3572
- - each phase needs its own timeout and artifact,
3573
- - model workers should run in bounded parallel,
3574
- - later work can start as soon as partial results arrive,
3575
- - a webhook must acknowledge before the work finishes, or
3576
- - operators need live progress beyond one tool spinner.
3577
-
3578
- Keep external writes in finalization. Pass a `runId` between stages, journal
3579
- progress before publishing, and make partial-worker failures visible in the
3580
- result.
3581
-
3582
- ## Where to go next
3583
-
3584
- - [GitHub](/docs/guides/github.md)
3585
- - [Tools](/docs/reference/tools.md)
3586
- - [Channels](/docs/reference/channels.md)
3587
- - [Playground](/docs/reference/playground.md)
3588
- - [Storage](/docs/storage.md)
3589
- - [Live A/B metrics](/docs/ab.md)
3590
- - [Evals](/docs/evals.md)
3591
-
3592
- ---
3593
-
3594
- Source: /docs/example-agents/slack-agent.md
3595
-
3596
- # Put a minimal agent in Slack
3597
-
3598
- Slack agent is the smallest channel example. It has one runtime config, one
3599
- instruction file, and one authored channel. A teammate mentions the agent,
3600
- the local runtime harness runs a turn, and the answer
3601
- returns to the same Slack thread.
3602
-
3603
- Use it to learn the minimum needed for a Slack agent before adding tools,
3604
- workflows, or a dedicated app.
3605
-
3606
- [Browse the Slack agent source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/slack-agent/)
3607
-
3608
- ## Keep the Slack channel small
3609
-
3610
- Slack agent delegates transport details to the host connection. The authored
3611
- file selects the account-linked transport, gives the agent a single-token
3612
- router name and icon, and supplies suggested prompts.
3613
-
3614
- The complete channel lives in
3615
- [`agent/channels/slack.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/slack-agent/agent/channels/slack.ts).
3616
- The framework supplies message intake, thread-scoped sessions, delivery,
3617
- status updates, and suggested prompts.
3618
-
3619
- ## Follow a Slack message
3620
-
3621
- 1. A user mentions the agent or sends the host app a direct message naming
3622
- it.
3623
- 2. The Slack relay selects this channel by its single-token `agentName`.
3624
- 3. The Agent SDK maps the Slack channel and thread timestamp to a continuation
3625
- key.
3626
- 4. The local harness runs with
3627
- [`instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/slack-agent/agent/instructions.md).
3628
- 5. The response returns to the triggering thread.
3629
- 6. A later message in the same thread resumes the durable session.
3630
-
3631
- The prompt asks for concise threaded replies. It doesn't define domain policy
3632
- or tool routing.
3633
-
3634
- ## Map the Slack agent files
3635
-
3636
- | File | Purpose |
3637
- | --- | --- |
3638
- | [`package.json`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/slack-agent/package.json) | Declares the example package and Agent SDK dependency. |
3639
- | [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/slack-agent/agent/agent.ts) | Names the agent and selects the model. The omitted `runtime` defaults to local. |
3640
- | [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/slack-agent/agent/instructions.md) | Sets the always-on response style. |
3641
- | [`agent/channels/slack.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/slack-agent/agent/channels/slack.ts) | Connects the signed-in host account to Slack. |
3642
- | [`agent/storage.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/slack-agent/agent/storage.ts) | Persists sessions and events with `cursorHostedStorage`. |
3643
-
3644
- There are no authored tools, skills, MCP connections, subagents, schedules,
3645
- hooks, A/B experiments, or evals. This small surface is the lesson.
3646
-
3647
- ## Connect the host
3648
-
3649
- You need:
3650
-
3651
- - Node 22.13 or newer.
3652
- - An agent-runtime credential.
3653
- - Slack connected through the selected channel transport.
3654
-
3655
- Sign in and confirm the active account:
3656
-
3657
- ```bash
3658
- agent-sdk login
3659
- agent-sdk whoami
3660
- ```
3661
-
3662
- The selected transport owns Slack credential setup. See the
3663
- [Slack guide](/docs/guides/slack.md) for account-linked and dedicated-app
3664
- options.
3665
-
3666
- ## Validate and start the server
3667
-
3668
- ```bash
3669
- agent-sdk validate --dir examples/slack-agent
3670
- agent-sdk info --dir examples/slack-agent --json
3671
- agent-sdk dev examples/slack-agent
3672
- ```
3673
-
3674
- The dev command prints the playground URL. It also mounts the Slack channel
3675
- and waits for relayed messages.
3676
-
3677
- In Slack, address the configured host app and router name, then send:
3678
-
3679
- > `<host-app mention> <router name>` Explain the Agent SDK in three bullets.
3680
-
3681
- Reply in the generated thread:
3682
-
3683
- > Make the second bullet simpler.
3684
-
3685
- The second message reaches the same session. You can open that session in the
3686
- playground to inspect the received message, model events, final reply, and
3687
- usage.
3688
-
3689
- ## Test without Slack
3690
-
3691
- Every project gets the built-in HTTP channel even when no HTTP file exists.
3692
- Run a one-shot turn through it:
3693
-
3694
- ```bash
3695
- agent-sdk run --dir examples/slack-agent \
3696
- --message "Explain the Agent SDK simply."
3697
- ```
3698
-
3699
- The same project also exposes an MCP endpoint. Since this agent has no server
3700
- tools, its MCP surface contains `ask` and `check`, but not `call_tool`.
3701
-
3702
- These automatic surfaces let you test the prompt from the CLI and let another
3703
- agent delegate to it later. The authored Slack channel only changes how work
3704
- arrives and where replies go.
3705
-
3706
- ## Know when to add a dedicated app
3707
-
3708
- An account-linked Slack transport is a fit for mentions, direct messages, thread
3709
- continuity, and agent-branded replies. Move to a dedicated Socket Mode channel
3710
- when you need:
3711
-
3712
- - top-level channel watching,
3713
- - interactive approval buttons,
3714
- - a separate bot identity, or
3715
- - Slack app events unsupported by the account-linked relay.
3716
-
3717
- Compare this example with [Playbook router](/docs/example-agents/benny.md), which adds allowlisted channel
3718
- watching, and [Weather agent](/docs/example-agents/weather-agent.md), which adds approval buttons
3719
- through a second Slack channel.
3720
-
3721
- ## Turn the channel into your own Slack agent
3722
-
3723
- Copy the three authored files, then change:
3724
-
3725
- - `name` in `agent.ts` for the harness identity,
3726
- - `agentName` in `slack.ts` for the single-token router name,
3727
- - the instructions for your domain, and
3728
- - suggested prompts for the tasks teammates should try.
3729
-
3730
- Keep `agentName` free of whitespace. Use PascalCase for multiword names.
3731
-
3732
- ## Where to go next
3733
-
3734
- - [Slack](/docs/guides/slack.md)
3735
- - [Channels](/docs/reference/channels.md)
3736
- - [Sessions and streaming](/docs/reference/sessions.md)
3737
- - [Playground](/docs/reference/playground.md)
3738
-
3739
- ---
3740
-
3741
- Source: /docs/example-agents/weather-agent.md
3742
-
3743
- # Explore the full Agent SDK surface with a weather agent
3744
-
3745
- The weather agent is the broadest small example in the repository. It fetches
3746
- live conditions and forecasts, converts units through MCP, writes notes in a
3747
- session workspace, and runs from HTTP, Slack, a schedule, and the MCP endpoint.
3748
-
3749
- Use this project when you want to see how the Agent SDK's filesystem pieces fit
3750
- together before you design a larger agent.
3751
-
3752
- [Browse the weather agent source.](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/weather-agent/)
3753
-
3754
- ## See the runtime features together
3755
-
3756
- Most examples focus on one feature. Weather agent puts the major runtime
3757
- features side by side:
3758
-
3759
- | Capability | Source | Role |
3760
- | --- | --- | --- |
3761
- | Root config and instructions | [`agent/agent.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/agent.ts), [`agent/instructions.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/instructions.md) | Select the cloud runtime and route each request. |
3762
- | Server tools | [`agent/tools/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/weather-agent/agent/tools/) | Fetch Open-Meteo data and call MCP from the serve host. |
3763
- | Agent tool | [`save_weather_note.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/tools/save_weather_note.ts) | Run a Python script inside the session workspace. |
3764
- | Stdio MCP | [`units.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/mcp-connections/units.ts), [`probe.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/mcp-connections/probe.ts) | Expose conversion tools to the model, host tools, and channel handlers. Author VM-side probe tools as TypeScript `execute` functions. |
3765
- | Custom HTTP | [`webhook.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/channels/webhook.ts) | Start a turn or call MCP without a model turn. |
3766
- | Slack | [`slack.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/channels/slack.ts), [`slack-app.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/channels/slack-app.ts) | Compare account-linked chat with a dedicated app. |
3767
- | Skill and subagent | [`forecast.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/skills/forecast.md), [`researcher/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/weather-agent/agent/subagents/researcher/) | Load a procedure on demand or delegate broad research. |
3768
- | Schedule and hooks | [`heartbeat.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/schedules/heartbeat.md), [`audit.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/hooks/audit.ts), [`journal.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/hooks/journal.ts) | Start recurring tasks, log usage, and save turn summaries. |
3769
- | A/B and evals | [`agent/ab.ts`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/examples/weather-agent/agent/ab.ts), [`evals/`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/weather-agent/evals/) | Compare a sticky variant and protect tool routing with regression cases. |
3770
-
3771
- ## Follow one request
3772
-
3773
- A current-weather question takes this path:
3774
-
3775
- 1. The built-in HTTP channel, Slack, or the custom `/report` route creates a
3776
- durable session.
3777
- 2. `instructions.md` tells the model to call `get_weather` instead of
3778
- guessing.
3779
- 3. The server tool geocodes the city, fetches Open-Meteo, validates the
3780
- response, and returns normalized fields.
3781
- 4. The agent writes a short answer. The Agent SDK records every event in the
3782
- session stream.
3783
- 5. The audit hook observes `turn.completed`. If the session joined the A/B
3784
- experiment, the collector updates its metrics too.
3785
-
3786
- Forecasts route to `get_forecast`. Unit conversions route to
3787
- `convert_temperature`, which calls the `units` MCP server through
3788
- `ctx.host.mcp`. Climate history and broad comparisons route to the
3789
- `researcher` subagent.
3790
-
3791
- ## Prepare the example
3792
-
3793
- You need:
3794
-
3795
- - Node 22.13 or newer.
3796
- - An agent-runtime credential.
3797
- - Network access to Open-Meteo.
3798
- - Python 3 for `save_weather_note`.
3799
-
3800
- The project mounts an account-linked Slack channel. The Agent SDK checks the
3801
- connection at startup, so sign in even when you plan to call a deterministic
3802
- tool.
3803
-
3804
- The optional dedicated Slack app also needs a token pair.
3805
- `agent-sdk slack create --dir examples/weather-agent` provisions the
3806
- app and writes the tokens for you (see the
3807
- [Slack guide](/docs/guides/slack.md#set-it-up)); with hand-minted tokens,
3808
- export them instead:
3809
-
3810
- ```bash
3811
- export WEATHER_AGENT_SLACK_BOT_TOKEN=xoxb-...
3812
- export WEATHER_AGENT_SLACK_APP_TOKEN=xapp-...
3813
- agent-sdk slack doctor --prefix WEATHER_AGENT
3814
- ```
1690
+ | `t.score(name, value)` | records a 0–1 score you computed; soft until you add a bar |
1691
+ | `t.requireToolCall(name, matcher?)` | gates on a matching call and returns it, so later code can read its input and output |
1692
+ | `t.requireInputRequest(filter?)` | gates on exactly one pending approval request and returns it |
3815
1693
 
3816
- Without those two tokens, the dedicated channel stays idle. The
3817
- account-linked channel still works.
1694
+ Every gate returns a handle: `.soft()` demotes it to tracked-only,
1695
+ `.atLeast(0.7)` adds a soft score bar, and `.gate(0.8)` promotes a
1696
+ scored assertion into a hard gate.
3818
1697
 
3819
- ## Inspect before running
1698
+ With no matcher, `calledTool` is request-based: a requested call counts
1699
+ even when its result has not arrived. Pass
1700
+ `t.calledTool("inspect_pr", { status: "completed" })` to require the
1701
+ call to return. `input`, `output`, and `count` matcher fields accept a
1702
+ literal, a `RegExp`, or a predicate.
3820
1703
 
3821
- ```bash
3822
- agent-sdk validate --dir examples/weather-agent
3823
- agent-sdk info --dir examples/weather-agent --json
3824
- agent-sdk eval --dir examples/weather-agent --list
3825
- ```
1704
+ The expectation builders are `includes(string | RegExp)`,
1705
+ `equals(value)`, `matches(schema)`, `similarity(expected)`, and
1706
+ `satisfies(predicate, label)`. `includes` stringifies its input,
1707
+ `equals` compares values deeply, `matches` validates against a Standard
1708
+ Schema (or anything with `safeParse`, like Zod), `similarity` scores
1709
+ normalized text similarity, and `satisfies` runs your predicate. The
1710
+ plain function `normalizedSimilarity(actual, expected)` returns the
1711
+ same 0–1 score for use with `t.score`.
3826
1712
 
3827
- `validate` should pass. `info` and `eval --list` should match the capabilities
3828
- described above.
1713
+ A few more context members shape a case: `t.require(value, expectation)`
1714
+ records a gate and stops the test body when it fails, without a
1715
+ duplicate execution error. `t.skip(reason)` ends the case as skipped
1716
+ (reported separately, never changes the exit code; call it before
1717
+ sending messages). `t.metric(name, value)` records a structured score
1718
+ for the playground case card. `t.log(message)` records a debug line for
1719
+ the CLI and playground result.
3829
1720
 
3830
- ## Call the typed tools
1721
+ Three `t.send` options apply on session create (first `t.send` only):
3831
1722
 
3832
- Start with the current-weather server tool:
1723
+ - `workspaceFiles`: `{ path: contents }`, seeded into the local session
1724
+ workspace. Prefer this over machine-local paths.
1725
+ - `workspaceDir`: absolute harness cwd (local runtime).
1726
+ - `cloud`: per-session cloud options merged over the agent's static
1727
+ `cloud` config (repos / env / …). Use a pinned `repos` override to
1728
+ attach a fixture repo for cloud evals without putting it on the
1729
+ agent's default `cloud.repos`. Cloud ignores `workspaceFiles` seeds.
3833
1730
 
3834
- ```bash
3835
- agent-sdk call get_weather \
3836
- --dir examples/weather-agent \
3837
- --input '{"city":"New York City"}'
1731
+ ```ts
1732
+ const toolResults = t.events.filter((e) => e.type === "action.result");
1733
+ t.check(
1734
+ toolResults.length,
1735
+ satisfies((n) => (n as number) <= 4, "at most 4 tool calls")
1736
+ );
3838
1737
  ```
3839
1738
 
3840
- `defineTool` gives the input a Zod schema. The Agent SDK validates the JSON before
3841
- `execute` runs. The result includes the matched place, condition,
3842
- temperature, humidity, wind, gusts, and precipitation.
1739
+ A case with no explicit gates falls back to whether at least one turn
1740
+ completed successfully. Add `t.succeeded()` and behavior-specific gates
1741
+ anyway. They make the contract visible during review.
3843
1742
 
3844
- Try the forecast:
1743
+ ### Judge free-form output
3845
1744
 
3846
- ```bash
3847
- agent-sdk call get_forecast \
3848
- --dir examples/weather-agent \
3849
- --input '{"city":"Lisbon","days":5}'
1745
+ When wording matters and no regex captures it, `t.judge` grades the
1746
+ reply with an LLM. The built-in graders are `factuality(expected)`,
1747
+ `summarizes(expected)`, `closedQA(criteria)`, and `sql(expected)`. Each
1748
+ scores `t.reply` by default; pass `{ on }` to grade another value.
1749
+
1750
+ ```ts
1751
+ t.judge.factuality("It is 54°F in NYC right now.").atLeast(0.7);
3850
1752
  ```
3851
1753
 
3852
- The tool accepts one to seven days. Shared Open-Meteo code lives under
3853
- `agent/lib/`, so the Agent SDK imports it without discovering another tool.
1754
+ Judge assertions are soft by default, so a judge never fails a build
1755
+ until you give it a bar with `.atLeast(0.7)` or promote it with
1756
+ `.gate(0.8)`. The judge model comes from `defineEvalConfig({ judge })`,
1757
+ `defineEval({ judge })`, a case-level `judge`, or a per-call
1758
+ `{ model }` override; the nearest one wins. For a domain-specific judge
1759
+ whose verdict is not a single score, `t.judge.model(prompt)` sends a
1760
+ raw prompt to the same model and returns the reply. You then record the
1761
+ parsed result with `t.score` or `t.check`.
3854
1762
 
3855
- ## Compare server and agent execution
1763
+ ## Run evals from the CLI
3856
1764
 
3857
- Most weather tools use the default `execution: "server"`. Their TypeScript
3858
- runs inside the serve host and can reach `ctx.host` services.
1765
+ The `eval` command discovers, filters, and runs cases.
3859
1766
 
3860
- `save_weather_note` uses `execution: "agent"` instead. The Agent SDK materializes
3861
- its script into the agent environment. The script reads JSON from stdin and
3862
- appends to `weather-notes.md` in that session's workspace:
1767
+ Run the CLI under Node 22.13 or newer. Do not use Bun. Its HTTP/2 client
1768
+ breaks tool-result streams and causes eval turns to fail.
3863
1769
 
3864
1770
  ```bash
3865
- agent-sdk run --dir examples/weather-agent \
3866
- --message "Save a note that Boston is cold and windy."
1771
+ agent-sdk eval --dir . --list # discover only
1772
+ agent-sdk eval --dir . # run all
1773
+ agent-sdk eval --dir . builds/checkout # one datapoint
1774
+ agent-sdk eval --dir . builds search # several ids or prefixes
1775
+ agent-sdk eval --dir . --tag smoke --tag pull-request # any matching tag
1776
+ agent-sdk eval --dir . --json --no-stream # machine-readable results
1777
+ agent-sdk eval --dir . --verbose # logs + reply snippets
3867
1778
  ```
3868
1779
 
3869
- Each session gets its own workspace. Saving a note doesn't edit the authored
3870
- example.
3871
-
3872
- This split matters on cloud. Server tools stay on the Agent SDK host.
3873
- Agent tools run inside the cloud workspace.
3874
-
3875
- ## Verify tool execution on the agent VM
3876
-
3877
- Ask the agent to call `probe_cloud_tool` on the `probe` MCP server.
3878
- `agent/mcp-connections/probe.ts` authors that tool as TypeScript. The Agent
3879
- SDK packages it as stdio MCP so a cloud VM with no checkout of this example
3880
- can still run it. The model lists the server and calls the tool; it does
3881
- not write a `.sh`.
3882
-
3883
- A real call writes `vm-tool-observations/<id>.json` in the agent cwd and
3884
- returns hostname, cwd, and pid. Stream events show `probe:probe_cloud_tool`,
3885
- not `shell`.
1780
+ Id filters use OR semantics. Each filter selects an exact id and its
1781
+ descendants. For example, `builds` selects `builds`,
1782
+ `builds/checkout`, and every other case below that path. Repeated tags
1783
+ also use OR semantics. When you provide both ids and tags, a case must
1784
+ match both groups.
3886
1785
 
3887
- A local tool script or a marker under `probes/` means the model
3888
- invented a substitute.
1786
+ `eval` boots an ephemeral server on port 0 with a temp state root
1787
+ outside the project, so cases don't inherit ambient monorepo rules and
1788
+ don't write into the project state directory. Point `--url` at a running server to eval
1789
+ a live agent instead:
3889
1790
 
3890
1791
  ```bash
3891
- agent-sdk run --dir examples/weather-agent \
3892
- --message "Test custom tool execution from this cloud agent. Call probe_cloud_tool."
1792
+ agent-sdk eval --dir . \
1793
+ --url http://127.0.0.1:3000/weather-agent \
1794
+ --bearer-token "$AGENT_TOKEN"
3893
1795
  ```
3894
1796
 
3895
- ## Use MCP connections in three places
1797
+ The eval definitions still come from `--dir`; `--url` only changes the
1798
+ agent that receives the turns. For a locally mounted multi-agent
1799
+ directory, `--slug weather-agent` chooses the target. Use
1800
+ `--state-root` to keep ephemeral session state at a chosen path,
1801
+ `--timeout-ms` to override the project timeout, and `--no-stream` to
1802
+ keep live progress off stderr. A TTY streams turn progress by default.
1803
+ `--verbose` still writes `t.log` lines to stderr and adds reply snippets
1804
+ to text results.
3896
1805
 
3897
- `agent/mcp-connections/units.ts` starts a local stdio server. The filename
3898
- makes its server name `units`. The Agent SDK exposes it to:
1806
+ Model turns need a Cursor credential from `agent-sdk login` or
1807
+ `CURSOR_API_KEY`.
3899
1808
 
3900
- - the model as MCP tools,
3901
- - server tools through `ctx.host.mcp`, and
3902
- - channel handlers through `host.mcp`.
1809
+ See [CLI: eval](/docs/reference/cli.md#eval) for flags and exit codes.
3903
1810
 
3904
- `probe` is a second stdio connection. Its tools are TypeScript `execute`
3905
- functions; the Agent SDK packages them so a cloud VM can spawn the server
3906
- without this checkout. The model calls `probe_cloud_tool` directly; no host
3907
- tool wraps it.
1811
+ ### JSON results
3908
1812
 
3909
- `convert_temperature` demonstrates the server-tool path:
1813
+ Use `--json --no-stream` in scripts and CI. The top-level result carries
1814
+ the totals and one result per case:
3910
1815
 
3911
- ```bash
3912
- agent-sdk call convert_temperature \
3913
- --dir examples/weather-agent \
3914
- --input '{"value":72,"from":"F"}'
1816
+ ```json
1817
+ {
1818
+ "ok": true,
1819
+ "passed": 1,
1820
+ "failed": 0,
1821
+ "results": [
1822
+ {
1823
+ "id": "readiness",
1824
+ "ok": true,
1825
+ "assertions": [{ "name": "succeeded", "passed": true }],
1826
+ "sessionId": "ses_123",
1827
+ "inputs": ["Is checkout pull request 42 ready to approve?"],
1828
+ "toolCalls": [{ "toolName": "inspect_pr", "isError": false }],
1829
+ "logs": [],
1830
+ "durationMs": 12340
1831
+ }
1832
+ ]
1833
+ }
3915
1834
  ```
3916
1835
 
3917
- The custom channel demonstrates the handler path. Start the dev server:
1836
+ Each case result can also include `description`, `finalText`, `tools`,
1837
+ `error`, and tool arguments or output. This shape lets CI report the
1838
+ failed assertion without parsing terminal text.
3918
1839
 
3919
- ```bash
3920
- agent-sdk dev examples/weather-agent
3921
- ```
1840
+ ## Run evals in the playground
3922
1841
 
3923
- Then call MCP deterministically through `/convert`:
1842
+ Start the server, open the playground, and choose **Evals**. You can run
1843
+ every case or one case, watch progress, and open the resulting session
1844
+ trace. The Evals tab works on a normal `serve`.
3924
1845
 
3925
1846
  ```bash
3926
- curl -s -X POST \
3927
- http://127.0.0.1:3000/weather-agent/v1/channels/webhook/convert \
3928
- -H 'content-type: application/json' \
3929
- -d '{"value":20,"from":"C"}'
1847
+ agent-sdk serve --dir .
3930
1848
  ```
3931
1849
 
3932
- No model chooses a tool in this route. The handler calls the MCP server and
3933
- returns its result.
3934
-
3935
- ## Keep conversation state in a custom channel
1850
+ Playground runs target the live server instead of an ephemeral one.
1851
+ Their sessions appear in the session list. One eval batch can run at a
1852
+ time. Persistence follows the rule under
1853
+ [Configure eval runs](#configure-eval-runs). See
1854
+ [Playground eval routes](/docs/reference/http-api.md#playground-eval-routes).
1855
+ The start request returns `202` while cases run in the background.
1856
+ Poll until the snapshot status becomes `completed`, `failed`, or `cancelled`.
1857
+ Configuration errors appear on a failed snapshot.
3936
1858
 
3937
- `POST /report` starts a model turn and waits for it:
1859
+ On `--prod` / `--url`, the CLI prints the Eval ID as soon as the batch is
1860
+ accepted (and a Playground deep link with `?view=evals&evalRunId=…`):
3938
1861
 
3939
1862
  ```bash
3940
- curl -s -X POST \
3941
- http://127.0.0.1:3000/weather-agent/v1/channels/webhook/report \
3942
- -H 'content-type: application/json' \
3943
- -d '{"message":"What is the weather in Paris?"}'
3944
- ```
3945
-
3946
- The response includes a `key`. Send it back on the next request to continue
3947
- the same session:
1863
+ agent-sdk eval --prod --slug vulnerability-scanner --tag deepsec
1864
+ # Eval ID: evalrun_…
1865
+ # Cancel: agent-sdk eval cancel evalrun_… --prod --slug vulnerability-scanner
1866
+ # Playground: https://…/playground?view=evals&evalRunId=evalrun_…
3948
1867
 
3949
- ```bash
3950
- curl -s -X POST \
3951
- http://127.0.0.1:3000/weather-agent/v1/channels/webhook/report \
3952
- -H 'content-type: application/json' \
3953
- -d '{"message":"How about tomorrow?","key":"<key>"}'
1868
+ agent-sdk eval cancel evalrun_… --prod --slug vulnerability-scanner
1869
+ agent-sdk eval status evalrun_… --prod --slug vulnerability-scanner
3954
1870
  ```
3955
1871
 
3956
- This is the custom-channel version of a continuation token. See
3957
- [webhooks and custom channels](/docs/guides/webhooks.md) for route schemas,
3958
- authentication, and asynchronous handlers.
3959
-
3960
- ## Load procedures and delegate research
1872
+ ## What good cases assert
3961
1873
 
3962
- The forecast skill gives the root agent an on-demand procedure. The Agent SDK
3963
- advertises the skill's description, then the harness loads its content when
3964
- the request matches.
1874
+ Gate decisions and shape, not prose. Model wording varies run to run.
1875
+ Tool choice, tool avoidance, and output structure are the stable
1876
+ contract.
3965
1877
 
3966
- The `researcher` directory is an SDK subagent. Its description tells the
3967
- parent when to delegate. It inherits the parent's execution surface, but
3968
- gets its own instructions:
1878
+ 1. `t.succeeded()`: always, first.
1879
+ 2. The tool decision: `calledTool` for the intended path,
1880
+ `notCalledTool` for the likely wrong alternative. The pair is
1881
+ stronger than either alone.
1882
+ 3. Output shape: a regex for the contract (`/ready|blocked/i`, a JSON
1883
+ marker, a findings-block fence), never exact sentences.
1884
+ 4. For structured output, parse `t.reply` and check fields with
1885
+ `satisfies` instead of substring-matching JSON.
3969
1886
 
3970
- ```bash
3971
- agent-sdk run --dir examples/weather-agent \
3972
- --message "Compare record summer temperatures across Paris, London, and Rome."
3973
- ```
1887
+ The common failure modes: asserting exact phrasing, packing more than
1888
+ about five gates into one case (split it), and cases that depend on live
1889
+ external state that drifts (pin the input; see fixtures).
3974
1890
 
3975
- Use a skill when the same agent needs a procedure. Use a subagent when the
3976
- parent should hand a bounded task to a specialist. The
3977
- [subagents reference](/docs/reference/subagents.md) explains the current
3978
- inheritance limits.
1891
+ ## Pick fixtures by agent type
3979
1892
 
3980
- ## Trigger the schedule and inspect the hook
1893
+ The right fixture depends on the surface under test.
3981
1894
 
3982
- The heartbeat schedule runs at 09:00 UTC on weekdays. Automatic schedule
3983
- timers stay off under `--dev`, so dispatch it manually:
1895
+ | Agent surface | Fixture |
1896
+ | --- | --- |
1897
+ | Chat / domain assistant | A canonical prompt string, chosen once and frozen |
1898
+ | Tool-heavy | Run `agent-sdk call <tool>` first to pin what the tool returns, then freeze the prompt that triggers it |
1899
+ | GitHub webhook | `agent-sdk github replay <pr> --events '*' --dry-run --out fixtures/github` snapshots real payloads for offline replay ([GitHub guide](/docs/guides/github.md)) |
1900
+ | PR reviewer with host preparation | Diff, metadata, and gold labels pinned to commit SHAs; keep any live PR matrix small |
1901
+ | Workspace-dependent | `workspaceFiles` in `t.send` options, never developer-machine paths |
3984
1902
 
3985
- ```bash
3986
- curl -s -X POST \
3987
- http://127.0.0.1:3000/weather-agent/v1/dev/schedules/heartbeat
3988
- ```
1903
+ Tag the fast, reliably passing core `smoke` and run `--tag smoke` in the
1904
+ inner loop. Leave slow or flaky-prone cases untagged for explicit runs.
3989
1905
 
3990
- It creates a task session to check San Francisco, New York, and London.
3991
- The audit hook logs usage after each completed turn. Hooks observe recorded
3992
- events; their failures don't fail the turn.
1906
+ ### Materialize API-backed fixtures
3993
1907
 
3994
- ## Measure variants and regressions
1908
+ An input that only points at external data, such as a pull request URL,
1909
+ snapshot id, or pair of commit SHAs, is not self-contained. Fetch it
1910
+ once and commit the rendered fixture before you expand the suite.
3995
1911
 
3996
- The `weather-tool-efficiency` A/B experiment assigns sessions by a sticky
3997
- hash:
1912
+ 1. Save the diff, metadata, and labels under `fixtures/` at pinned
1913
+ revisions.
1914
+ 2. Seed those files with `workspaceFiles`, or read them from the fixture
1915
+ directory.
1916
+ 3. Assert decisions and output shape against the saved evidence.
1917
+ 4. Keep a small `smoke` subset for any remaining live pipeline checks.
3998
1918
 
3999
- - `control` returns current conditions in Fahrenheit.
4000
- - `treatment` adds a brief Celsius instruction and changes `get_weather` to
4001
- return Celsius fields.
1919
+ Read committed fixtures with `@cursor/july/evals/loaders`: `loadJson`,
1920
+ `loadJsonl`, and `loadYaml` resolve relative paths against the project
1921
+ root the runner discovered, not the cwd the CLI was invoked from
1922
+ (`resolveFixturePath` and `evalFixtureRoot` expose the same
1923
+ resolution for other file formats).
4002
1924
 
4003
- Samples and aggregate snapshots persist under the project state
4004
- directory. The treatment
4005
- only changes current conditions; `get_forecast` still returns Fahrenheit.
4006
- Treat the branch as an example of `ctx.session.abs`, not a complete unit
4007
- policy.
1925
+ `maxConcurrency` limits parallel datapoints. It does not limit model or
1926
+ API fan-out inside one datapoint. Materialized fixtures prevent a large
1927
+ suite from exhausting provider and GitHub rate limits. The
1928
+ [evals skill](https://github.com/cursor/cursor/blob/main/packages/agent-serve/skills/evals/SKILL.md) has the full fixture workflow.
4008
1929
 
4009
- List and run the evals:
1930
+ ## Keep improvements with regression evals
4010
1931
 
4011
- ```bash
4012
- agent-sdk eval --dir examples/weather-agent --list
4013
- agent-sdk eval --dir examples/weather-agent --json
4014
- ```
1932
+ Every [hillclimb](/docs/hillclimbing.md) round that keeps a change must land
1933
+ an eval that would have failed before the change. If you can't express
1934
+ the improvement as a gate (a `calledTool` shift, a bounded
1935
+ `action.result` count, an output-shape regex), the improvement is
1936
+ unverified, and it'll regress silently.
4015
1937
 
4016
- The suite covers weather and forecast routing, the local converter, workspace
4017
- notes, and the VM-side probe.
1938
+ The rule cuts the other way too: never weaken an existing gate to make a
1939
+ round pass. That's the freeze line moving, and it turns your regression
1940
+ suite into a list of checks that no longer protect anything.
4018
1941
 
4019
- ## Turn the weather tour into your own agent
1942
+ ## Compare variants on live traffic
4020
1943
 
4021
- Keep the shape and replace the domain:
1944
+ Use `defineAB` to compare variant metrics on live sessions. It is not a
1945
+ test runner and has no `agent-sdk ab` command. Keep `defineEval` as the
1946
+ regression ratchet. Eval sessions do not enroll or change live metrics.
1947
+ See [Live A/B metrics](/docs/ab.md) for assignment, behavior, collection,
1948
+ and inspection.
4022
1949
 
4023
- - Swap Open-Meteo tools for your typed service clients.
4024
- - Keep deterministic transforms behind direct server tools or MCP.
4025
- - Use an agent tool only when code must run in the agent workspace.
4026
- - Put reusable procedures in skills and narrow specialist work into
4027
- subagents.
4028
- - Add a channel only when the external surface needs its own identity,
4029
- continuation key, or delivery behavior.
1950
+ ## What's next
4030
1951
 
4031
- ## Where to go next
1952
+ Continue with these pages:
4032
1953
 
4033
- - [Tools](/docs/reference/tools.md)
4034
- - [MCP connections](/docs/reference/connections.md)
4035
- - [Human-in-the-loop approvals](/docs/guides/human-in-the-loop.md)
4036
- - [Slack](/docs/guides/slack.md)
4037
- - [Schedules and reminders](/docs/reference/schedules.md)
4038
- - [Evals](/docs/evals.md)
4039
- - [Live A/B metrics](/docs/ab.md)
1954
+ - [Live A/B metrics](/docs/ab.md): sticky variants and cumulative metrics
1955
+ on live sessions
1956
+ - [Hillclimbing](/docs/hillclimbing.md): the loop evals make trustworthy
1957
+ - [Building agents with agents](/docs/building-with-agents.md): have a
1958
+ coding agent write the first suite
1959
+ - [GitHub guide](/docs/guides/github.md): deterministic webhook fixtures
1960
+ with `github replay`
1961
+ - [Sessions and streaming](/docs/reference/sessions.md): the events
1962
+ `t.events` contains
4040
1963
 
4041
1964
  ---
4042
1965
 
@@ -4143,8 +2066,8 @@ Peer MCP connections are ordinary MCP connections, so deterministic host code
4143
2066
  can use them too. A channel handler or server tool can call
4144
2067
  `host.mcp.callTool("weather", "ask", { message: "…" })` without any
4145
2068
  model turn deciding to. See
4146
- [MCP connections](/docs/reference/connections.md#every-mcp-connection-is-available-in-three-places)
4147
- for the three places every MCP connection is available.
2069
+ [MCP connections](/docs/reference/connections.md#every-model-visible-mcp-connection-is-available-in-three-places)
2070
+ for the three places every model-visible MCP connection is available.
4148
2071
 
4149
2072
  ## What's next
4150
2073
 
@@ -4223,6 +2146,7 @@ mapping shifts:
4223
2146
  | Agent tools (`execution: "agent"`) | scripts in the session workspace | catalog + script bodies on the first prompt |
4224
2147
  | `skills/*` | `.cursor/skills/` in the workspace | native discovery after the first turn, from the hosted store or the signed-in account |
4225
2148
  | `mcp-connections/*.ts` | SDK `mcpServers` | SDK `mcpServers` (peers need `--public-url`) |
2149
+ | `host-connections/*.ts` | `ctx.host.mcp` only | `ctx.host.mcp` only |
4226
2150
  | `sandbox/workspace/**` | seeded into the session workspace | ignored |
4227
2151
  | Tool approvals (`needsApproval`) | supported | not supported; keep approval-gated tools on local turns |
4228
2152
 
@@ -4354,7 +2278,7 @@ or run lifecycle scripts from an existing `package.json`.
4354
2278
  | Prompt model | Pins the model on `defineAgent` in `agent/agent.ts`. `git_config` and `agent_options` remain comments. |
4355
2279
  | Cron trigger | Creates `agent/schedules/<slug>.ts` with `defineSchedule` in UTC. |
4356
2280
  | GitHub trigger | Creates `agent/channels/github.ts`. It converts pull-request action, push branch, issue action, and user allowlist filters. |
4357
- | Slack trigger | Creates `agent/channels/slack.ts`. Mention-only uses `cursorAccount`; watches, reactions, and channel-created triggers use Socket Mode. Watches add `engagement.channelPosts`. |
2281
+ | Slack trigger | Creates `agent/channels/slack.ts` as Socket Mode with `envPrefix` from the automation name (same names `slack create` writes). Watches add `engagement.channelPosts`. Run `agent-sdk slack create` for the bot. |
4358
2282
  | Linear, PagerDuty, Sentry, Teams, or generic webhook | Creates a boilerplate `agent/channels/<slug>.ts`. |
4359
2283
  | HTTP or SSE MCP server | Creates a name-based Cursor-account connection under `agent/mcp-connections/`. The project contains no server URL or credentials. |
4360
2284
  | Stdio MCP server | Writes `agent/mcp-connections/<slug>.todo.md`. |
@@ -4716,10 +2640,10 @@ A comment-only first wake has no head SHA, so the check waits for a
4716
2640
  PR or CI event. The banner still posts. A later turn on the same SHA
4717
2641
  creates a new check run; GitHub cannot reopen a completed run.
4718
2642
 
4719
- Override `events` when the mapping is custom. [Approval Buddy](/docs/example-agents/approval-buddy.md)
4720
- posts commit status from `turn.started` / `action.result` / `turn.failed`
4721
- and stays never-red; that pattern still wins when you replace a default
4722
- handler key. Handlers you author replace the matching defaults (same as
2643
+ Override `events` when the mapping is custom. A handler can post commit
2644
+ status from `turn.started` / `action.result` / `turn.failed` and stay
2645
+ never-red; that pattern still wins when you replace a default handler
2646
+ key. Handlers you author replace the matching defaults (same as
4723
2647
  `progress.reactions` composition today).
4724
2648
 
4725
2649
  ## Related
@@ -4885,10 +2809,8 @@ The companion skill is
4885
2809
  to that connection's resource URL
4886
2810
  - Upsert deployment secrets with `--store` so hosted engines seed the
4887
2811
  same tokens from env
4888
- - Keep privileged servers off the model with `hostOnly: true` while
4889
- tools still call them through `ctx.host.mcp`. Do not set `hostOnly` on
4890
- connectors the playground or local chat should call. Use
4891
- `advertiseTools: true` for those.
2812
+ - Use `advertiseTools: true` when local turns should call the server by
2813
+ name. Host tools can still call it through `ctx.host.mcp`.
4892
2814
 
4893
2815
  Prefer a Cursor account MCP connection when the connector already lives
4894
2816
  in the signed-in account dashboard:
@@ -4903,7 +2825,8 @@ and the host must hold tokens.
4903
2825
 
4904
2826
  ## How do I declare a host-OAuth connection?
4905
2827
 
4906
- Add one file under `agent/mcp-connections/`. The filename is the
2828
+ Add one file under `agent/mcp-connections/` (model + host) or
2829
+ `agent/host-connections/` (host + `mcp oauth` only). The filename is the
4907
2830
  connection name you pass to the CLI and to `host.mcp`.
4908
2831
 
4909
2832
  ```ts
@@ -4913,17 +2836,14 @@ import { defineConnection } from "@cursor/july/connections";
4913
2836
  export default defineConnection({
4914
2837
  url: "https://mcp.example.com/inventory",
4915
2838
  oauth: true,
4916
- hostOnly: true,
4917
- description:
4918
- "Inventory MCP (privileged). Call only from host tools, not the model.",
2839
+ description: "Inventory MCP.",
4919
2840
  });
4920
2841
  ```
4921
2842
 
4922
2843
  Rules of the road:
4923
2844
 
4924
2845
  - `oauth: true` is required for `agent-sdk mcp oauth`
4925
- - `hostOnly: true` hides the server from the model; `ctx.host.mcp` and
4926
- channel handlers still see it
2846
+ - A file under `mcp-connections/` is visible to the model and to `ctx.host.mcp`. A file under `host-connections/` stays on the host.
4927
2847
  - Declare expected secret names on the agent when you plan to `--store`:
4928
2848
 
4929
2849
  ```ts
@@ -4956,8 +2876,7 @@ agent-sdk mcp oauth inventory
4956
2876
 
4957
2877
  What happens:
4958
2878
 
4959
- 1. The Agent SDK loads `agent/mcp-connections/inventory.ts` and checks
4960
- `oauth: true`
2879
+ 1. The Agent SDK loads the connection file and checks `oauth: true`
4961
2880
  2. It opens the authorization URL in your browser
4962
2881
  3. The callback lands on `http://localhost:8787/callback`
4963
2882
  4. Tokens land in `mcp-auth.json` under the CLI config directory
@@ -4996,8 +2915,6 @@ on the pod.
4996
2915
 
4997
2916
  ## How do host tools call the server?
4998
2917
 
4999
- Keep privileged calls on the host:
5000
-
5001
2918
  ```ts
5002
2919
  const result = await ctx.host.mcp.callTool(
5003
2920
  "inventory",
@@ -5006,9 +2923,8 @@ const result = await ctx.host.mcp.callTool(
5006
2923
  );
5007
2924
  ```
5008
2925
 
5009
- The model never sees `hostOnly` tools in its MCP namespace list. If the
5010
- agent asks to "check IDE MCP" or run `mcp_auth`, point it at your host
5011
- tool instead.
2926
+ The model can call the same server. Use a host tool when the write needs
2927
+ an allowlist or other deterministic gate.
5012
2928
 
5013
2929
  ## What if authorization fails?
5014
2930
 
@@ -5021,7 +2937,7 @@ tool instead.
5021
2937
 
5022
2938
  ## What's next
5023
2939
 
5024
- - [MCP connections](/docs/reference/connections.md): transports, `hostOnly`, account MCP
2940
+ - [MCP connections](/docs/reference/connections.md): transports, account MCP
5025
2941
  - [CLI](/docs/reference/cli.md#mcp-oauth): full flag list for `mcp oauth`
5026
2942
  - [Deployment](/docs/deployment.md): secrets, egress, and hosted engines
5027
2943
  - [Fix common agent problems](/docs/troubleshooting.md): more symptom → fix tables
@@ -5246,10 +3162,10 @@ Source: /docs/guides/slack.md
5246
3162
 
5247
3163
  # Slack agents
5248
3164
 
5249
- The Slack channel puts your agent in Slack. Two products: the Cursor-hosted
5250
- connection (`cursorAccount: true`), or a dedicated Socket Mode app created
5251
- in the dashboard wizard (`agent-sdk slack create`). To own the Slack app
5252
- yourself, run `agent-sdk slack init --manual` and paste the manifests at
3165
+ The Slack channel puts your agent in Slack as its own Socket Mode bot.
3166
+ `agent-sdk slack create` opens the dashboard wizard and writes tokens
3167
+ to `.env.local`. To own the Slack app yourself, run
3168
+ `agent-sdk slack init --manual` and paste the manifests at
5253
3169
  [api.slack.com](https://api.slack.com/apps). Socket Mode has no
5254
3170
  public Request URL. Replies stream in threads, with tool "thinking" steps,
5255
3171
  suggested prompts, and opt-in approval buttons.
@@ -5288,41 +3204,12 @@ Missing tokens leave the channel idle (`channel idle … missing
5288
3204
  credentials`) rather than failing `serve`. That's useful when you mount
5289
3205
  many agents and only some have Slack apps.
5290
3206
 
5291
- ## Use the Cursor Slack connection
5292
-
5293
- If the Cursor Slack app is already installed in your workspace and linked
5294
- to your Cursor account, skip the dedicated Slack app:
5295
-
5296
- ```ts
5297
- import { slackChannel } from "@cursor/july/channels/slack";
5298
-
5299
- export default slackChannel({
5300
- cursorAccount: true,
5301
- agentName: "Weatherbot", // single token — no spaces; defaults from mount slug (PascalCase)
5302
- agentIcon: { emoji: ":robot_face:" },
5303
- });
5304
- ```
5305
-
5306
- Sign the host in (`agent-sdk login` or `CURSOR_API_KEY`), then mention the
5307
- agent in Slack as `@Cursor Weatherbot …`. Thread replies and DMs keep going to
5308
- the same agent. Messages appear as the Cursor app under that agent's name
5309
- and icon. Slack shows its working status, then posts one final reply.
5310
-
5311
- Use a dedicated Socket Mode Slack app when you need your own bot user,
5312
- channel watching (`engagement.channelPosts`), or approval buttons. On
5313
- `cursorAccount`, agents must be explicitly addressed (@mention, DM, or
5314
- claimed-thread reply). Channel watching and `toolApprovals` /
5315
- `interactivity` are Socket Mode only; the Cursor connection does not relay
5316
- Block Kit clicks. Agent names must be unique on the host; an unmatched
5317
- `@Cursor <name>` stays on Cursor's normal Slack agent.
5318
-
5319
3207
  ## Control who can message the agent
5320
3208
 
5321
3209
  External senders are blocked by default. Slack Connect users, guests, and
5322
3210
  people whose home workspace is not the install team never reach the
5323
- handler. That applies to Socket Mode and `cursorAccount: true`. Set
5324
- `blockExternals: false` only when the agent should serve people outside
5325
- your org:
3211
+ handler. Set `blockExternals: false` only when the agent should serve
3212
+ people outside your org:
5326
3213
 
5327
3214
  ```ts
5328
3215
  export default slackChannel({
@@ -5485,8 +3372,7 @@ export default slackChannel({
5485
3372
  ```
5486
3373
 
5487
3374
  Channel watching needs the `message.channels` / `message.groups` events
5488
- on the Slack app (Socket Mode only; not available with
5489
- `cursorAccount: true`). Pass `--channel-posts` on `slack create` or
3375
+ on the Slack app. Pass `--channel-posts` on `slack create` or
5490
3376
  `slack init --manual`. The bot must also be a member of each watched
5491
3377
  channel.
5492
3378
 
@@ -5496,9 +3382,7 @@ Set `includeBotPosts: true` when the posts worth watching come from bots:
5496
3382
  alert feeds, webhook integrations, or other agents posting notes. The
5497
3383
  watching app's own posts stay dropped either way, matched by the `bot_id`
5498
3384
  and bot user id from `auth.test`, so an agent can never dispatch on its
5499
- own replies. The
5500
- [alert investigator example](/docs/example-agents/oncall.md) watches a
5501
- bot-fed alerts channel this way.
3385
+ own replies. Use this for a bot-fed alerts channel.
5502
3386
 
5503
3387
  ## Prepare work on the host
5504
3388
 
@@ -5527,11 +3411,6 @@ Approval cards need interactivity on the Slack app. Recreate with
5527
3411
  `buildToolApprovalEvents({ credentials })` into `events` and set
5528
3412
  `interactivity: true` on the channel so Socket Mode routes the clicks.
5529
3413
 
5530
- Approval buttons need Socket Mode. `slackChannel({ cursorAccount: true })`
5531
- rejects `toolApprovals` and `interactivity` at construction, since the
5532
- Cursor Slack connection does not relay Block Kit clicks. Use a dedicated
5533
- Slack app to run approvals for a cursor-account agent.
5534
-
5535
3414
  Cards show redacted, truncated arguments (Block Kit size limits);
5536
3415
  execution still uses the full validated input, so review sensitive tools
5537
3416
  in the playground when the arguments may exceed the card. Approvals
@@ -5563,7 +3442,7 @@ Two habits matter most.
5563
3442
  The `slack` subcommands cover setup end to end.
5564
3443
 
5565
3444
  ```bash
5566
- agent-sdk slack setup # two-product chooser plus manual phases
3445
+ agent-sdk slack setup # printed setup guide
5567
3446
  agent-sdk slack create --dir . # dashboard wizard (dev app)
5568
3447
  agent-sdk slack create --dir . --prod # prod app
5569
3448
  agent-sdk slack destroy --dir . # delete the provisioned app
@@ -6178,7 +4057,6 @@ npx @cursor/july docs
6178
4057
  | New to the Agent SDK | [Quickstart](/docs/quickstart.md) (PR reviewer), then [Concepts](/docs/concepts.md) |
6179
4058
  | Building a new agent with Cursor | [Scaffold an agent with Cursor](/docs/scaffolding-agents.md) |
6180
4059
  | Turning a Cursor Automation into a project | [Convert a Cursor Automation](/docs/guides/convert-automation.md) |
6181
- | Learning from working agents | [Example agents](/docs/example-agents/index.md) |
6182
4060
  | Wiring an agent to Slack | [Slack guide](/docs/guides/slack.md) |
6183
4061
  | Starting from a packaged template | [Demo](/docs/templates/demo.md), [Security reviewer](/docs/templates/security-reviewer.md), [Triage](/docs/templates/triage.md), or [Agentic Owners](/docs/templates/agentic-owners.md) |
6184
4062
  | Wiring an agent to GitHub webhooks | [GitHub guide](/docs/guides/github.md) |
@@ -6246,33 +4124,6 @@ npx @cursor/july docs
6246
4124
  - [OpenTelemetry](/docs/guides/opentelemetry.md): push session, turn, and
6247
4125
  tool traces to an OTLP collector you run.
6248
4126
 
6249
- **Example agents**
6250
-
6251
- - [Choose the right example](/docs/example-agents/index.md): compare the
6252
- example agents by runtime, channels, tools, and state.
6253
- - [Weather agent](/docs/example-agents/weather-agent.md): explore tools, MCP,
6254
- approvals, skills, subagents, schedules, hooks, A/B metrics, and evals.
6255
- - [Slack agent](/docs/example-agents/slack-agent.md): put a minimal agent in
6256
- Slack through an account-linked transport.
6257
- - [Concierge](/docs/example-agents/concierge.md): delegate work to a peer agent
6258
- with its own context and sessions.
6259
- - [Playbook router](/docs/example-agents/benny.md): route Slack intake through inherited
6260
- repository playbooks.
6261
- - [Alert investigator](/docs/example-agents/oncall.md): watch a Slack alerts
6262
- channel and pin a self-rechecking investigation to every alert thread.
6263
- - [PR evidence reviewer](/docs/example-agents/bugbot.md): review a host-prepared,
6264
- diff-first pull-request evidence tree.
6265
- - [Approval Buddy](/docs/example-agents/approval-buddy.md): keep approval policy
6266
- in code while subagents supply review findings.
6267
- - [Security Reviewer](/docs/example-agents/security-reviewer.md): run a staged,
6268
- parallel security pipeline with live playground progress.
6269
- - [Knowledge base](/docs/example-agents/knowledge-base.md): turn conversations
6270
- about people, systems, decisions, and preferences into shared markdown.
6271
- - [Codebase wiki](/docs/example-agents/codebase-wiki.md): ingest merged PRs into
6272
- per-feature pages with a daily digest schedule.
6273
- - [Codeowners review](/docs/example-agents/codeowners-review.md): route PR
6274
- reviews by ownership to per-area playbooks and aggregate verdicts.
6275
-
6276
4127
  **Operating**
6277
4128
 
6278
4129
  - [Deployment](/docs/deployment.md): Cursor-managed hosting, self-hosting,
@@ -6676,9 +4527,8 @@ See [GitHub](/docs/guides/github.md) for local event delivery and
6676
4527
 
6677
4528
  ## Where to go next
6678
4529
 
6679
- - [`examples/approval-buddy`](https://github.com/cursor/cursor/tree/main/packages/agent-serve/examples/approval-buddy/): an example
6680
- with commit statuses, review subagents, and a deterministic stamp
6681
- policy
4530
+ - [PR autofixer template](/docs/templates/pr-autofixer.md): drive a PR on a
4531
+ Cursor cloud VM
6682
4532
  - [Evals](/docs/evals.md): freeze these two PRs as regression checks so
6683
4533
  prompt changes can't flip a verdict
6684
4534
  - [Tools](/docs/reference/tools.md): more on typed tools, approvals, and
@@ -8062,7 +5912,8 @@ command prints a note when you pass it anyway.
8062
5912
  agent-sdk mcp oauth <connection> [--dir .] [--store] [--slug <slug>] [--team <id>]
8063
5913
  ```
8064
5914
 
8065
- `<connection>` is the `agent/mcp-connections/<connection>.ts` basename.
5915
+ `<connection>` is the basename under `agent/mcp-connections/` or
5916
+ `agent/host-connections/`.
8066
5917
  `--slug` defaults to the `--dir` basename. `--team` defaults to the
8067
5918
  signed-in account's team. You need `agent-sdk login` (or `--api-key`)
8068
5919
  before `--store`.
@@ -8275,6 +6126,11 @@ from `@cursor/july/connections`, and the transport comes in
8275
6126
  four shapes: remote HTTP, local stdio, the signed-in Cursor account's
8276
6127
  connectors, and peer agents on the same host.
8277
6128
 
6129
+ Put a server in `agent/host-connections/` when host tools should call it
6130
+ and the model should not. Same `defineConnection` shape. `agent-sdk mcp
6131
+ oauth` still works. The playground and the turn's MCP servers never see
6132
+ those files.
6133
+
8278
6134
  ## Remote MCP server
8279
6135
 
8280
6136
  Point an MCP connection at a remote server with a URL.
@@ -8311,12 +6167,10 @@ agent-sdk mcp oauth inventory --store # also upsert deployment secrets
8311
6167
  Full walkthrough: [Host MCP OAuth](/docs/guides/mcp-oauth.md). Companion
8312
6168
  skill: [`skills/mcp-auth/SKILL.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/skills/mcp-auth/SKILL.md).
8313
6169
 
8314
- Set `hostOnly: true` only when host tools should call the server and the
8315
- model should not. Playground chat will not see those tools. Account MCP
8316
- (`cursorAccount: true`) is the right choice for connectors already linked
8317
- in the Cursor dashboard. Omit `servers` (or pass `"*"`) to forward every
8318
- connected connector. If the model should call those tools by name on
8319
- local turns, set `advertiseTools: true`.
6170
+ Account MCP (`cursorAccount: true`) is the right choice for connectors
6171
+ already linked in the Cursor dashboard. Omit `servers` (or pass `"*"`)
6172
+ to forward every connected connector. If the model should call those
6173
+ tools by name on local turns, set `advertiseTools: true`.
8320
6174
 
8321
6175
  ## Per-session auth (`auth`)
8322
6176
 
@@ -8376,7 +6230,7 @@ export default defineConnection({
8376
6230
 
8377
6231
  A listing failure, invalid tool name, or name collision fails the turn.
8378
6232
  Advertised tools follow the same runtime support as server tools. They cannot
8379
- be combined with `hostOnly` or called through the direct tool API.
6233
+ be called through the direct tool API.
8380
6234
 
8381
6235
  In a dry-run session, MCP tools marked read-only run normally. Tools marked
8382
6236
  as writes are stubbed. Tools without effect annotations are unavailable.
@@ -8484,13 +6338,14 @@ Unknown slugs and self-references fail `serve` at startup. Resolution
8484
6338
  (loopback versus `--public-url`), loop caveats, and the delegation model
8485
6339
  are in the [Agent-to-agent guide](/docs/guides/agent-to-agent.md).
8486
6340
 
8487
- ## Every MCP connection is available in three places
6341
+ ## Every model-visible MCP connection is available in three places
8488
6342
 
8489
- One authored MCP connection serves three consumers.
6343
+ A file under `agent/mcp-connections/` serves three consumers. Host
6344
+ connections skip the first one.
8490
6345
 
8491
6346
  1. **Cursor agent:** Attached connections ride SDK `mcpServers` behind
8492
6347
  harness MCP meta-tools. Set `advertiseTools: true` so local turns see
8493
- named tools. `hostOnly` keeps the connection off the model.
6348
+ named tools.
8494
6349
  2. **Server tools:** Deterministic host code composes MCP calls
8495
6350
  through `ctx.host.mcp`:
8496
6351
 
@@ -8525,7 +6380,7 @@ lazily on first use.
8525
6380
 
8526
6381
  Continue with these pages:
8527
6382
 
8528
- - [Host MCP OAuth](/docs/guides/mcp-oauth.md): `mcp oauth`, `--store`, `hostOnly`
6383
+ - [Host MCP OAuth](/docs/guides/mcp-oauth.md): `mcp oauth`, `--store`
8529
6384
  - [Agent-to-agent](/docs/guides/agent-to-agent.md): peers in depth
8530
6385
  - [Tools](/docs/reference/tools.md): authored tools that wrap MCP connections
8531
6386
  - [Webhooks](/docs/guides/webhooks.md): calling MCP connections from handlers
@@ -8602,9 +6457,8 @@ For GitHub merge-box checks and sticky PR banners, use
8602
6457
  `githubChannel({ progress: { commitStatus, banner } })` from
8603
6458
  `@cursor/july/channels/github`. That is the supported Autofix-style
8604
6459
  path. See [GitHub: Show PR progress](/docs/guides/github.md#show-pr-progress).
8605
- Override channel `events` only when the lifecycle is custom (for
8606
- example [Approval Buddy](/docs/example-agents/approval-buddy.md)'s
8607
- never-red status from tool output). Do not use `defineHook` for those
6460
+ Override channel `events` only when the lifecycle is custom, such as
6461
+ never-red status from tool output. Do not use `defineHook` for those
8608
6462
  writes.
8609
6463
 
8610
6464
  ## Patterns
@@ -8772,6 +6626,12 @@ while a turn runs). Agent-execution tools are rejected with `400`, and
8772
6626
  unknown tools with `404` and the list of available names. For the
8773
6627
  semantics, see [Tools](/docs/reference/tools.md#call-a-tool-without-a-model-turn).
8774
6628
 
6629
+ An optional `"continuationToken"` (`<channelId>:<key>`, as
6630
+ `/v1/sessions` lists it; mutually exclusive with `sessionId`) addresses
6631
+ the session by continuation token instead; malformed tokens are
6632
+ rejected with `400 invalid_continuation_token`. For the semantics, see
6633
+ [Tools](/docs/reference/tools.md#call-a-tool-without-a-model-turn).
6634
+
8775
6635
  ## Discovery
8776
6636
 
8777
6637
  These read-only routes describe the running agent.
@@ -8779,6 +6639,8 @@ These read-only routes describe the running agent.
8779
6639
  | Route | What it does |
8780
6640
  | ---------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
8781
6641
  | `GET /v1/info` | The discovered surface: model, tools, skills, MCP connections, subagents, channels and routes (with schemas), schedules, hooks, A/B experiments, diagnostics |
6642
+ | `GET /v1/tools` | The live tool catalog: authored server tools plus advertised MCP passthroughs under model-facing names, as light `{ name, title?, source? }` entries. `session` / `continuationToken` query parameters bind the listing to a session identity (advertised inventories can be tenant-scoped); a connection whose listing fails is skipped and reported in `connectionErrors` |
6643
+ | `GET /v1/tools/:name` | One catalog tool's full description: description, execution, `needsApproval`, `effect`, input and output schemas, source connection. Same session binding as the listing; unknown names get `404` with the available names |
8782
6644
  | `GET /v1/health` | Per-agent liveness, no auth |
8783
6645
  | `GET /v1/logs?after=N` | Recent server log lines, with a polling cursor |
8784
6646
  | `GET /v1/abs` | [Live A/B metrics](/docs/ab.md): per-session assignments and aggregate arm totals |
@@ -9044,7 +6906,8 @@ experiments can override their file-derived name.
9044
6906
  | Path | Resolves to |
9045
6907
  | --- | --- |
9046
6908
  | `agent/tools/approve_pr.ts` | tool `approve_pr` |
9047
- | `agent/mcp-connections/linear.ts` | MCP connection `linear` |
6909
+ | `agent/mcp-connections/linear.ts` | MCP connection `linear` (model + host) |
6910
+ | `agent/host-connections/anytool.ts` | Host MCP connection `anytool` (host + `mcp oauth` only) |
9048
6911
  | `agent/skills/pr-review.md` | skill `pr-review` |
9049
6912
  | `agent/subagents/reviewer/` | subagent `reviewer` |
9050
6913
  | `agent/channels/drive.ts` | channel `drive`, routes under `/v1/channels/drive` |
@@ -9072,6 +6935,8 @@ my-agent/
9072
6935
  │ │ └── pr-review.md # on-demand procedures (SKILL.md convention)
9073
6936
  │ ├── mcp-connections/
9074
6937
  │ │ └── linear.ts # tools from external MCP servers
6938
+ │ ├── host-connections/
6939
+ │ │ └── anytool.ts # privileged MCP, host tools only
9075
6940
  │ └── channels/
9076
6941
  │ └── github.ts # messages and external events
9077
6942
  └── evals/
@@ -9093,6 +6958,7 @@ Each path maps to a capability and a reference page.
9093
6958
  | `agent/tools/<name>.ts` | One typed tool; filename = tool name. `execution: "server"` (in-process, default) or `"agent"` (a script that runs where the agent runs) | [Tools](/docs/reference/tools.md) |
9094
6959
  | `agent/skills/*` | SKILL.md-convention procedures, loaded on demand | [Skills](/docs/reference/skills.md) |
9095
6960
  | `agent/mcp-connections/<name>.ts` | MCP servers, available to the model, to server tools (`ctx.host.mcp`), and to channel/schedule handlers (`args.host.mcp`) | [MCP connections](/docs/reference/connections.md) |
6961
+ | `agent/host-connections/<name>.ts` | Privileged MCP servers for `ctx.host.mcp` and `mcp oauth`. The model never sees them. | [MCP connections](/docs/reference/connections.md) |
9096
6962
  | `agent/subagents/<id>/` | Child agent directory; `description` required | [Subagents](/docs/reference/subagents.md) |
9097
6963
  | `agent/channels/*.ts` | HTTP surfaces beyond the built-in session API; `slack.ts` and `github.ts` use the platform packs | [Channels](/docs/reference/channels.md) |
9098
6964
  | `agent/hooks/*.ts` | Observe-only event subscribers, never fatal | [Hooks](/docs/reference/hooks.md) |
@@ -9698,8 +7564,8 @@ different one.
9698
7564
  Subagents inherit the parent's execution surface. Every per-subagent
9699
7565
  capability directory is reported as a warning and ignored: `tools/`,
9700
7566
  `skills/`, `mcp-connections/` (and the legacy `connections/` alias),
9701
- `channels/`, `schedules/`, `hooks/`, `sandbox/`, and nested
9702
- `subagents/`.
7567
+ `host-connections/`, `channels/`, `schedules/`, `hooks/`, `sandbox/`,
7568
+ and nested `subagents/`.
9703
7569
 
9704
7570
  Delegation needs both halves: the description makes it possible, and the
9705
7571
  parent's [instructions](/docs/reference/instructions.md) make it happen. "When a
@@ -9920,8 +7786,9 @@ passthrough server tools from the connection's live `listTools` on every
9920
7786
  local turn. See
9921
7787
  [MCP Connections](/docs/reference/connections.md#advertise-tools).
9922
7788
  Advertised tools ride the same execution path as authored server tools,
9923
- but cannot be invoked via
9924
- [direct tool calls](#call-a-tool-without-a-model-turn).
7789
+ and [direct tool calls](#call-a-tool-without-a-model-turn) address them
7790
+ by the same model-facing names: the call's session identity resolves the
7791
+ advertised listing when the authored lookup misses.
9925
7792
 
9926
7793
  ## Gate a tool on human approval
9927
7794
 
@@ -9995,7 +7862,22 @@ when the call returns. Pass a `sessionId` (a body field over
9995
7862
  HTTP, `--session` on the CLI, `options.sessionId` programmatically) to
9996
7863
  run inside an existing session instead: the tool sees that session's
9997
7864
  workspace, and the call is recorded on the session's event stream.
9998
- Session-bound calls return `409 session_busy` while a turn runs.
7865
+ Session-bound calls return `409 session_busy` while a turn runs. When the
7866
+ session's harness cwd cannot be materialized, a read-effect call runs
7867
+ in a scratch workspace instead and the outcome carries
7868
+ `scratchWorkspace: true`; a write-effect call fails with
7869
+ `workspace_unavailable`.
7870
+
7871
+ A session can also be addressed by its continuation token: an optional
7872
+ `continuationToken` (`<channelId>:<key>`, as `/v1/sessions` lists it;
7873
+ mutually exclusive with `sessionId`). A token that maps to a live
7874
+ session behaves exactly like passing that session's id — same ownership
7875
+ check, same `409 session_busy`, same event recording. A token with no
7876
+ session behind it runs the call scratch-bound with the token's channel
7877
+ id and continuation key as the call's session identity, so a deployment
7878
+ whose tools resolve state from the continuation key can serve it with
7879
+ no live session. Malformed tokens are rejected with
7880
+ `400 invalid_continuation_token`.
9999
7881
 
10000
7882
  The error semantics match the model path. Unknown tools are rejected
10001
7883
  with the available names, agent-execution tools cannot be called on the
@@ -10630,9 +8512,10 @@ fixture replay.
10630
8512
 
10631
8513
  ## Slack
10632
8514
 
10633
- `agent/channels/slack.ts` uses your signed-in Cursor account. Mention
10634
- the agent or DM it with a PR URL. Same `drive_pr` path as the
10635
- playground. See [Slack](/docs/guides/slack.md).
8515
+ `agent/channels/slack.ts` is a dedicated Socket Mode bot
8516
+ (`PR_AUTOFIXER_SLACK_*`). Mint it with `agent-sdk slack create`, then
8517
+ mention the bot or DM it with a PR URL.
8518
+ Same `drive_pr` path as the playground. See [Slack](/docs/guides/slack.md).
10636
8519
 
10637
8520
  ## Evals
10638
8521
 
@@ -10669,9 +8552,8 @@ This agent reads a pull request diff and posts one review comment. It
10669
8552
  reports exploitable bugs: injection, authz bypass, secret leaks, SSRF,
10670
8553
  RCE. Style nits stay out.
10671
8554
 
10672
- The factory [Security Reviewer](/docs/example-agents/security-reviewer.md)
10673
- runs a staged pipeline with parallel workers. This template is one
10674
- model turn.
8555
+ The in-repo factory Security Reviewer runs a staged pipeline with
8556
+ parallel workers. This template is one model turn.
10675
8557
 
10676
8558
  ## Scaffold
10677
8559
 
@@ -10884,7 +8766,7 @@ not on `PATH`, use `npx @cursor/july`.
10884
8766
  | Built-in file reads and greps fail; the turn retries for a long time | Run under Node 22.13+ (or `tsx`), never Bun. Look for `NGHTTP2_FRAME_SIZE_ERROR` in logs. |
10885
8767
  | The turn fails immediately with an API-key error | Sign in with `agent-sdk login`, or set `CURSOR_API_KEY`. Discovery, `info`, `call`, and serve bring-up work without a key; model turns need one. |
10886
8768
  | Replies quote rules or `AGENTS.md` from outside your agent project | The session workspace inherited parent-folder config. Nested git checkouts default `local.cwd` to a per-project cache directory under `~/.cache`. Point `defineAgent({ local: { cwd } })` at a checkout only when the agent should inherit that tree, or set `--state-root` to a clean directory (for example under `/tmp`). |
10887
- | Yellow box shows Datadog/Linear tools, but the model lists `GetDynamicTools` / IDE `cursor` tools and never calls them | Attached MCP sits behind harness meta-tools, or `hostOnly` hid the connection, or the harness cwd is still inside another checkout. Set `advertiseTools: true` for named tools on local turns. Check `GET /v1/info` `local.cwd` and `connections[].advertiseTools`. |
8769
+ | Yellow box shows Datadog/Linear tools, but the model lists `GetDynamicTools` / IDE `cursor` tools and never calls them | Attached MCP sits behind harness meta-tools, or the harness cwd is still inside another checkout. Set `advertiseTools: true` for named tools on local turns. Check `GET /v1/info` `local.cwd` and `connections[].advertiseTools`. |
10888
8770
  | Server tools, skills, or workspace seed files never appear | Server tools and sandbox seeds apply on the local runtime (cloud server tools need `--public-url` / `--cloud-tools-url`). Skills reach cloud through the Agent Store when hosting or a personal `CURSOR_API_KEY` is available; otherwise only skills already in the cloud repo. `validate` warns when this combination is present. |
10889
8771
  | `validate` and `run` succeed, but typecheck fails in CI | The CLI runs TypeScript with type-stripping only. Keep tool `execute` return types as object literals or `type` aliases, not `interface` types. |
10890
8772
  | Login works, but turns are rejected when using custom API hosts | Point login and model traffic at the same host (`CURSOR_API_BASE_URL` and `CURSOR_BACKEND_URL`). A key from one host is rejected by the other. |
@@ -10923,7 +8805,7 @@ not on `PATH`, use `npx @cursor/july`.
10923
8805
  | --- | --- |
10924
8806
  | `must be defineConnection({ url, oauth: true })` | The connection file needs `oauth: true`, or you passed the wrong connection name to `agent-sdk mcp oauth`. |
10925
8807
  | Local auth works; hosted calls unauthorized | Run `agent-sdk mcp oauth <name> --store`, confirm names with `agent-sdk secrets list <slug>`, then redeploy. |
10926
- | Model asks for `mcp_auth` or IDE MCP for a privileged server | That connection is `hostOnly`. Call it from a host tool via `ctx.host.mcp`, and update instructions. |
8808
+ | Model asks for `mcp_auth` or IDE MCP for a connector it already has | Attached MCP is behind meta-tools. Set `advertiseTools: true` for named tools on local turns, or call it from a host tool via `ctx.host.mcp`. |
10927
8809
 
10928
8810
  See [Host MCP OAuth](/docs/guides/mcp-oauth.md) and
10929
8811
  [`skills/mcp-auth/SKILL.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/skills/mcp-auth/SKILL.md).