@cursor/july 0.1.108 → 0.1.109

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (343) hide show
  1. package/AGENTS.md +2 -6
  2. package/README.md +3 -6
  3. package/dist/channels/bitbucket/api.d.ts +41 -0
  4. package/dist/channels/bitbucket/api.d.ts.map +1 -1
  5. package/dist/channels/bitbucket/api.js +260 -0
  6. package/dist/channels/bitbucket/binding.d.ts +4 -0
  7. package/dist/channels/bitbucket/binding.d.ts.map +1 -1
  8. package/dist/channels/bitbucket/binding.js +16 -0
  9. package/dist/channels/bitbucket/index.d.ts +1 -1
  10. package/dist/channels/bitbucket/index.d.ts.map +1 -1
  11. package/dist/channels/bitbucket/index.js +1 -1
  12. package/dist/channels/gitlab/api.d.ts +27 -0
  13. package/dist/channels/gitlab/api.d.ts.map +1 -1
  14. package/dist/channels/gitlab/api.js +88 -0
  15. package/dist/channels/gitlab/binding.d.ts +5 -0
  16. package/dist/channels/gitlab/binding.d.ts.map +1 -1
  17. package/dist/channels/gitlab/binding.js +10 -0
  18. package/dist/channels/gitlab/index.d.ts +1 -1
  19. package/dist/channels/gitlab/index.d.ts.map +1 -1
  20. package/dist/channels/gitlab/index.js +1 -1
  21. package/dist/channels/slack/dispatch.d.ts +10 -0
  22. package/dist/channels/slack/dispatch.d.ts.map +1 -1
  23. package/dist/channels/slack/dispatch.js +20 -3
  24. package/dist/channels/slack/slack-channel.d.ts +12 -5
  25. package/dist/channels/slack/slack-channel.d.ts.map +1 -1
  26. package/dist/channels/slack/slack-channel.js +59 -8
  27. package/dist/docs/404.html +2 -2
  28. package/dist/docs/assets/{app.BASKM3M4.js → app.Cr-wVbnB.js} +1 -1
  29. package/dist/docs/assets/{building-with-agents.md.CUSWxlP_.js → building-with-agents.md.D0KbSkJn.js} +2 -2
  30. package/dist/docs/assets/{building-with-agents.md.CUSWxlP_.lean.js → building-with-agents.md.D0KbSkJn.lean.js} +1 -1
  31. package/dist/docs/assets/chunks/@localSearchIndexroot.CFVQ4S17.js +1 -0
  32. package/dist/docs/assets/chunks/{VPLocalSearchBox.BOKxlYGP.js → VPLocalSearchBox.CVQERt56.js} +1 -1
  33. package/dist/docs/assets/chunks/{theme.DbDZW-zb.js → theme.Dnsd3XOn.js} +2 -2
  34. package/dist/docs/assets/concepts.md.B4o63Gul.js +1 -0
  35. package/dist/docs/assets/evals.md.C7JLjoEP.js +211 -0
  36. package/dist/docs/assets/evals.md.C7JLjoEP.lean.js +1 -0
  37. package/dist/docs/assets/guides_bitbucket.md.DZPzmUjU.js +10 -0
  38. package/dist/docs/assets/guides_bitbucket.md.DZPzmUjU.lean.js +1 -0
  39. package/dist/docs/assets/{guides_github.md.BH33UEBJ.js → guides_github.md.TZaTZlfz.js} +13 -3
  40. package/dist/docs/assets/{guides_github.md.BH33UEBJ.lean.js → guides_github.md.TZaTZlfz.lean.js} +1 -1
  41. package/dist/docs/assets/guides_gitlab.md.BmqwQfdG.js +14 -0
  42. package/dist/docs/assets/guides_gitlab.md.BmqwQfdG.lean.js +1 -0
  43. package/dist/docs/assets/{guides_webhooks.md.DKdA43Qm.js → guides_webhooks.md.CJK484ex.js} +2 -2
  44. package/dist/docs/assets/{hillclimbing.md.CpTGTCle.js → hillclimbing.md.BOiVo1tf.js} +1 -1
  45. package/dist/docs/assets/index.md.DD9Q2XuJ.js +5 -0
  46. package/dist/docs/assets/{index.md.BjH1w2ZW.lean.js → index.md.DD9Q2XuJ.lean.js} +1 -1
  47. package/dist/docs/assets/{reference_channels.md.D-qTqwcq.js → reference_channels.md.CAo-iK4j.js} +2 -2
  48. package/dist/docs/assets/{reference_channels.md.D-qTqwcq.lean.js → reference_channels.md.CAo-iK4j.lean.js} +1 -1
  49. package/dist/docs/assets/{reference_cli.md.Ca26u0Es.js → reference_cli.md.Deg7849l.js} +1 -1
  50. package/dist/docs/assets/{reference_extensions.md.CPWt00ds.js → reference_extensions.md.ZAVUyuEX.js} +2 -2
  51. package/dist/docs/assets/{reference_hooks.md.Ddt5DdgJ.js → reference_hooks.md.BlM_bOg6.js} +3 -3
  52. package/dist/docs/assets/{reference_hooks.md.Ddt5DdgJ.lean.js → reference_hooks.md.BlM_bOg6.lean.js} +1 -1
  53. package/dist/docs/assets/{reference_http-api.md.oySXBO8o.js → reference_http-api.md.BwaCo-VO.js} +1 -1
  54. package/dist/docs/assets/{reference_playground.md.4myJPxrf.js → reference_playground.md.DLnoaczX.js} +1 -1
  55. package/dist/docs/assets/{reference_playground.md.4myJPxrf.lean.js → reference_playground.md.DLnoaczX.lean.js} +1 -1
  56. package/dist/docs/assets/{reference_project-layout.md.DuBu9a96.js → reference_project-layout.md.BEU8MtQV.js} +3 -3
  57. package/dist/docs/assets/{reference_project-layout.md.DuBu9a96.lean.js → reference_project-layout.md.BEU8MtQV.lean.js} +1 -1
  58. package/dist/docs/assets/reference_sessions.md.CyXV1MUw.js +1 -0
  59. package/dist/docs/assets/skills_deploy.md.CWqi_ZxW.js +35 -0
  60. package/dist/docs/assets/skills_deploy.md.CWqi_ZxW.lean.js +1 -0
  61. package/dist/docs/assets/{skills_evals.md.BhovOvrl.js → skills_evals.md.DFxYPErF.js} +3 -3
  62. package/dist/docs/assets/skills_evals.md.DFxYPErF.lean.js +1 -0
  63. package/dist/docs/assets/{skills_framework-map.md.D-tFZhFS.js → skills_framework-map.md.BxLSOhSY.js} +1 -1
  64. package/dist/docs/assets/skills_index.md.DL7EHaQ-.js +1 -0
  65. package/dist/docs/assets/skills_index.md.DL7EHaQ-.lean.js +1 -0
  66. package/dist/docs/assets/{storage.md.BOHeqk2M.js → storage.md.BUrhJ-Zz.js} +4 -4
  67. package/dist/docs/assets/{storage.md.BOHeqk2M.lean.js → storage.md.BUrhJ-Zz.lean.js} +1 -1
  68. package/dist/docs/building-with-agents.html +5 -5
  69. package/dist/docs/building-with-agents.md +3 -2
  70. package/dist/docs/concepts.html +5 -5
  71. package/dist/docs/concepts.md +2 -4
  72. package/dist/docs/deployment.html +4 -4
  73. package/dist/docs/evals.html +161 -35
  74. package/dist/docs/evals.md +612 -297
  75. package/dist/docs/guides/agent-to-agent.html +4 -4
  76. package/dist/docs/guides/bitbucket.html +36 -0
  77. package/dist/docs/guides/bitbucket.md +84 -0
  78. package/dist/docs/guides/cloud-agents.html +4 -4
  79. package/dist/docs/guides/convert-automation.html +4 -4
  80. package/dist/docs/guides/github.html +17 -7
  81. package/dist/docs/guides/github.md +31 -0
  82. package/dist/docs/guides/gitlab.html +40 -0
  83. package/dist/docs/guides/gitlab.md +92 -0
  84. package/dist/docs/guides/grokbot-agents.html +4 -4
  85. package/dist/docs/guides/human-in-the-loop.html +4 -4
  86. package/dist/docs/guides/improve.html +4 -4
  87. package/dist/docs/guides/mcp-oauth.html +4 -4
  88. package/dist/docs/guides/opentelemetry.html +4 -4
  89. package/dist/docs/guides/slack.html +5 -5
  90. package/dist/docs/guides/webhooks.html +6 -6
  91. package/dist/docs/guides/webhooks.md +4 -2
  92. package/dist/docs/hashmap.json +1 -1
  93. package/dist/docs/hillclimbing.html +6 -6
  94. package/dist/docs/hillclimbing.md +2 -2
  95. package/dist/docs/index.html +6 -6
  96. package/dist/docs/index.md +10 -6
  97. package/dist/docs/llms-full.txt +1113 -778
  98. package/dist/docs/llms.txt +6 -5
  99. package/dist/docs/quickstart.html +4 -4
  100. package/dist/docs/reference/agent-config.html +4 -4
  101. package/dist/docs/reference/artifacts.html +4 -4
  102. package/dist/docs/reference/channels.html +6 -6
  103. package/dist/docs/reference/channels.md +16 -3
  104. package/dist/docs/reference/cli.html +6 -6
  105. package/dist/docs/reference/cli.md +6 -4
  106. package/dist/docs/reference/connections.html +4 -4
  107. package/dist/docs/reference/extensions.html +7 -7
  108. package/dist/docs/reference/extensions.md +0 -3
  109. package/dist/docs/reference/hooks.html +7 -7
  110. package/dist/docs/reference/hooks.md +9 -12
  111. package/dist/docs/reference/http-api.html +6 -6
  112. package/dist/docs/reference/http-api.md +2 -3
  113. package/dist/docs/reference/instructions.html +4 -4
  114. package/dist/docs/reference/playground.html +5 -5
  115. package/dist/docs/reference/playground.md +0 -4
  116. package/dist/docs/reference/project-layout.html +7 -7
  117. package/dist/docs/reference/project-layout.md +1 -8
  118. package/dist/docs/reference/prompt.html +4 -4
  119. package/dist/docs/reference/result.html +4 -4
  120. package/dist/docs/reference/schedules.html +4 -4
  121. package/dist/docs/reference/sessions.html +5 -5
  122. package/dist/docs/reference/sessions.md +1 -2
  123. package/dist/docs/reference/skills.html +4 -4
  124. package/dist/docs/reference/subagents.html +4 -4
  125. package/dist/docs/reference/tools.html +4 -4
  126. package/dist/docs/scaffolding-agents.html +4 -4
  127. package/dist/docs/skills/create-agent.html +4 -4
  128. package/dist/docs/skills/debug.html +4 -4
  129. package/dist/docs/skills/deploy.html +61 -0
  130. package/dist/docs/skills/deploy.md +161 -0
  131. package/dist/docs/skills/evals.html +8 -8
  132. package/dist/docs/skills/evals.md +44 -6
  133. package/dist/docs/skills/framework-map.html +5 -5
  134. package/dist/docs/skills/framework-map.md +1 -2
  135. package/dist/docs/skills/github.html +4 -4
  136. package/dist/docs/skills/hillclimb.html +4 -4
  137. package/dist/docs/skills/index.html +6 -6
  138. package/dist/docs/skills/index.md +1 -1
  139. package/dist/docs/skills/mcp-auth.html +4 -4
  140. package/dist/docs/skills/otel.html +4 -4
  141. package/dist/docs/skills/setup-slack.html +4 -4
  142. package/dist/docs/storage.html +8 -8
  143. package/dist/docs/storage.md +15 -24
  144. package/dist/docs/templates/agentic-owners.html +4 -4
  145. package/dist/docs/templates/agents-md.html +4 -4
  146. package/dist/docs/templates/code-wiki.html +4 -4
  147. package/dist/docs/templates/demo.html +4 -4
  148. package/dist/docs/templates/grokbot-agents.html +4 -4
  149. package/dist/docs/templates/pr-autofixer.html +4 -4
  150. package/dist/docs/templates/security-help.html +4 -4
  151. package/dist/docs/templates/security-reviewer.html +4 -4
  152. package/dist/docs/templates/triage.html +4 -4
  153. package/dist/docs/troubleshooting.html +4 -4
  154. package/dist/extensions.d.ts +2 -3
  155. package/dist/extensions.d.ts.map +1 -1
  156. package/dist/extensions.js +2 -5
  157. package/dist/index.d.ts +2 -3
  158. package/dist/index.d.ts.map +1 -1
  159. package/dist/index.js +1 -2
  160. package/dist/internal/authored-alias-hooks.d.ts +5 -0
  161. package/dist/internal/authored-alias-hooks.d.ts.map +1 -1
  162. package/dist/internal/authored-alias-hooks.js +17 -0
  163. package/dist/internal/authored-loaders.d.ts +4 -0
  164. package/dist/internal/authored-loaders.d.ts.map +1 -1
  165. package/dist/internal/authored-loaders.js +21 -2
  166. package/dist/internal/cli-ax.d.ts +1 -2
  167. package/dist/internal/cli-ax.d.ts.map +1 -1
  168. package/dist/internal/cli-ax.js +1 -2
  169. package/dist/internal/continuation-channel.d.ts.map +1 -1
  170. package/dist/internal/continuation-channel.js +2 -2
  171. package/dist/internal/continuation-identity.js +8 -3
  172. package/dist/internal/discovery/agent.d.ts +1 -1
  173. package/dist/internal/discovery/agent.d.ts.map +1 -1
  174. package/dist/internal/discovery/agent.js +0 -5
  175. package/dist/internal/discovery/extension-overlay.d.ts +1 -2
  176. package/dist/internal/discovery/extension-overlay.d.ts.map +1 -1
  177. package/dist/internal/discovery/extension-overlay.js +0 -14
  178. package/dist/internal/discovery/extensions.d.ts +1 -2
  179. package/dist/internal/discovery/extensions.d.ts.map +1 -1
  180. package/dist/internal/discovery/extensions.js +4 -22
  181. package/dist/internal/discovery/info.d.ts.map +1 -1
  182. package/dist/internal/discovery/info.js +42 -24
  183. package/dist/internal/discovery/modules.js +0 -1
  184. package/dist/internal/discovery/project.d.ts.map +1 -1
  185. package/dist/internal/discovery/project.js +0 -16
  186. package/dist/internal/eval-runner.js +0 -1
  187. package/dist/internal/framework-file-storage.d.ts +4 -5
  188. package/dist/internal/framework-file-storage.d.ts.map +1 -1
  189. package/dist/internal/framework-file-storage.js +4 -5
  190. package/dist/internal/framework-storage-selection.d.ts +2 -2
  191. package/dist/internal/framework-storage-selection.js +2 -2
  192. package/dist/internal/init-scaffold.d.ts.map +1 -1
  193. package/dist/internal/init-scaffold.js +1 -2
  194. package/dist/internal/install-cursor-skills.d.ts +5 -2
  195. package/dist/internal/install-cursor-skills.d.ts.map +1 -1
  196. package/dist/internal/install-cursor-skills.js +25 -5
  197. package/dist/internal/run-client.d.ts +1 -1
  198. package/dist/internal/sdk-runner.js +1 -1
  199. package/dist/internal/server.js +0 -14
  200. package/dist/internal/session-engine.d.ts +5 -35
  201. package/dist/internal/session-engine.d.ts.map +1 -1
  202. package/dist/internal/session-engine.js +17 -162
  203. package/dist/internal/storage-coordinator.d.ts +5 -19
  204. package/dist/internal/storage-coordinator.d.ts.map +1 -1
  205. package/dist/internal/storage-coordinator.js +3 -62
  206. package/dist/internal/storage-roles.d.ts +4 -9
  207. package/dist/internal/storage-roles.d.ts.map +1 -1
  208. package/dist/internal/storage-roles.js +2 -2
  209. package/dist/playground/assets/index-Bhxzrcf6.css +1 -0
  210. package/dist/playground/assets/index-CqLX5uF3.js +67 -0
  211. package/dist/playground/index.html +2 -2
  212. package/dist/storage-backends/cursor-hosted-v2.d.ts +4 -5
  213. package/dist/storage-backends/cursor-hosted-v2.d.ts.map +1 -1
  214. package/dist/storage-backends/cursor-hosted-v2.js +4 -5
  215. package/dist/storage-backends/cursor-hosted.d.ts +1 -2
  216. package/dist/storage-backends/cursor-hosted.d.ts.map +1 -1
  217. package/dist/storage-backends/cursor-hosted.js +0 -30
  218. package/dist/storage-backends/file-kv.d.ts +9 -12
  219. package/dist/storage-backends/file-kv.d.ts.map +1 -1
  220. package/dist/storage-backends/file-kv.js +11 -47
  221. package/dist/storage-protocol.d.ts +3 -11
  222. package/dist/storage-protocol.d.ts.map +1 -1
  223. package/dist/storage-protocol.js +3 -11
  224. package/dist/storage.d.ts +8 -36
  225. package/dist/storage.d.ts.map +1 -1
  226. package/dist/storage.js +8 -44
  227. package/dist/types.d.ts +20 -57
  228. package/dist/types.d.ts.map +1 -1
  229. package/docs/README.md +10 -6
  230. package/docs/building-with-agents.md +3 -2
  231. package/docs/concepts.md +2 -4
  232. package/docs/evals.md +613 -298
  233. package/docs/guides/bitbucket.md +89 -0
  234. package/docs/guides/github.md +31 -0
  235. package/docs/guides/gitlab.md +97 -0
  236. package/docs/guides/webhooks.md +4 -2
  237. package/docs/hillclimbing.md +2 -2
  238. package/docs/reference/channels.md +16 -3
  239. package/docs/reference/cli.md +6 -4
  240. package/docs/reference/extensions.md +0 -3
  241. package/docs/reference/hooks.md +9 -12
  242. package/docs/reference/http-api.md +2 -3
  243. package/docs/reference/playground.md +0 -4
  244. package/docs/reference/project-layout.md +1 -8
  245. package/docs/reference/sessions.md +1 -2
  246. package/docs/skills/index.md +2 -2
  247. package/docs/storage.md +15 -24
  248. package/package.json +1 -7
  249. package/skills/deploy/SKILL.md +169 -0
  250. package/skills/evals/SKILL.md +45 -8
  251. package/skills/framework-map/SKILL.md +1 -2
  252. package/src/channels/bitbucket/api.ts +341 -0
  253. package/src/channels/bitbucket/binding.ts +25 -0
  254. package/src/channels/bitbucket/index.ts +2 -0
  255. package/src/channels/gitlab/api.ts +123 -0
  256. package/src/channels/gitlab/binding.ts +12 -0
  257. package/src/channels/gitlab/index.ts +1 -0
  258. package/src/channels/slack/dispatch.ts +30 -0
  259. package/src/channels/slack/slack-channel.ts +69 -7
  260. package/src/extensions.ts +2 -6
  261. package/src/index.ts +0 -3
  262. package/src/internal/authored-alias-hooks.ts +32 -0
  263. package/src/internal/authored-loaders.ts +26 -2
  264. package/src/internal/cli-ax.ts +1 -2
  265. package/src/internal/continuation-channel.ts +2 -1
  266. package/src/internal/continuation-identity.ts +10 -2
  267. package/src/internal/discovery/agent.ts +1 -7
  268. package/src/internal/discovery/extension-overlay.ts +0 -18
  269. package/src/internal/discovery/extensions.ts +2 -26
  270. package/src/internal/discovery/info.ts +3 -13
  271. package/src/internal/discovery/modules.ts +0 -1
  272. package/src/internal/discovery/project.ts +0 -16
  273. package/src/internal/eval-runner.ts +0 -1
  274. package/src/internal/framework-file-storage.ts +4 -5
  275. package/src/internal/framework-storage-selection.ts +2 -2
  276. package/src/internal/init-scaffold.ts +1 -2
  277. package/src/internal/install-cursor-skills.ts +38 -5
  278. package/src/internal/run-client.ts +1 -1
  279. package/src/internal/sdk-runner.ts +1 -1
  280. package/src/internal/server.ts +0 -16
  281. package/src/internal/session-engine.ts +14 -202
  282. package/src/internal/storage-coordinator.ts +5 -76
  283. package/src/internal/storage-roles.ts +4 -9
  284. package/src/storage-backends/cursor-hosted-v2.ts +4 -7
  285. package/src/storage-backends/cursor-hosted.ts +0 -37
  286. package/src/storage-backends/file-kv.ts +10 -51
  287. package/src/storage-protocol.ts +3 -17
  288. package/src/storage.ts +10 -101
  289. package/src/types.ts +19 -56
  290. package/templates/demo/README.md +10 -6
  291. package/templates/demo/agent/channels/github.ts +2 -0
  292. package/templates/demo/agent/channels/queue.ts +6 -2
  293. package/templates/demo/agent/lib/repos.ts +5 -0
  294. package/templates/demo/init.json +25 -0
  295. package/dist/ab.d.ts +0 -209
  296. package/dist/ab.d.ts.map +0 -1
  297. package/dist/ab.js +0 -246
  298. package/dist/docs/ab.html +0 -80
  299. package/dist/docs/ab.md +0 -332
  300. package/dist/docs/assets/ab.md.mlVgqvSk.js +0 -54
  301. package/dist/docs/assets/ab.md.mlVgqvSk.lean.js +0 -1
  302. package/dist/docs/assets/chunks/@localSearchIndexroot.BHZYsNVi.js +0 -1
  303. package/dist/docs/assets/concepts.md.DgEcZOfT.js +0 -1
  304. package/dist/docs/assets/evals.md.C1ekS3k2.js +0 -85
  305. package/dist/docs/assets/evals.md.C1ekS3k2.lean.js +0 -1
  306. package/dist/docs/assets/index.md.BjH1w2ZW.js +0 -5
  307. package/dist/docs/assets/reference_sessions.md.CueyOHSL.js +0 -1
  308. package/dist/docs/assets/skills_ab.md.CsFNatVx.js +0 -26
  309. package/dist/docs/assets/skills_ab.md.CsFNatVx.lean.js +0 -1
  310. package/dist/docs/assets/skills_evals.md.BhovOvrl.lean.js +0 -1
  311. package/dist/docs/assets/skills_index.md.DKwIxzGg.js +0 -1
  312. package/dist/docs/assets/skills_index.md.DKwIxzGg.lean.js +0 -1
  313. package/dist/docs/skills/ab.html +0 -52
  314. package/dist/docs/skills/ab.md +0 -50
  315. package/dist/internal/ab-collector.d.ts +0 -44
  316. package/dist/internal/ab-collector.d.ts.map +0 -1
  317. package/dist/internal/ab-collector.js +0 -142
  318. package/dist/internal/ab-fold.d.ts +0 -36
  319. package/dist/internal/ab-fold.d.ts.map +0 -1
  320. package/dist/internal/ab-fold.js +0 -175
  321. package/dist/internal/ab-snapshot.d.ts +0 -68
  322. package/dist/internal/ab-snapshot.d.ts.map +0 -1
  323. package/dist/internal/ab-snapshot.js +0 -208
  324. package/dist/internal/discovery/ab.d.ts +0 -9
  325. package/dist/internal/discovery/ab.d.ts.map +0 -1
  326. package/dist/internal/discovery/ab.js +0 -113
  327. package/dist/playground/assets/index-BMDqAeXu.js +0 -67
  328. package/dist/playground/assets/index-LUgJoWdL.css +0 -1
  329. package/docs/ab.md +0 -337
  330. package/skills/ab/SKILL.md +0 -58
  331. package/src/ab.ts +0 -430
  332. package/src/internal/ab-collector.ts +0 -200
  333. package/src/internal/ab-fold.ts +0 -232
  334. package/src/internal/ab-snapshot.ts +0 -331
  335. package/src/internal/discovery/ab.ts +0 -131
  336. /package/dist/docs/assets/{concepts.md.DgEcZOfT.lean.js → concepts.md.B4o63Gul.lean.js} +0 -0
  337. /package/dist/docs/assets/{guides_webhooks.md.DKdA43Qm.lean.js → guides_webhooks.md.CJK484ex.lean.js} +0 -0
  338. /package/dist/docs/assets/{hillclimbing.md.CpTGTCle.lean.js → hillclimbing.md.BOiVo1tf.lean.js} +0 -0
  339. /package/dist/docs/assets/{reference_cli.md.Ca26u0Es.lean.js → reference_cli.md.Deg7849l.lean.js} +0 -0
  340. /package/dist/docs/assets/{reference_extensions.md.CPWt00ds.lean.js → reference_extensions.md.ZAVUyuEX.lean.js} +0 -0
  341. /package/dist/docs/assets/{reference_http-api.md.oySXBO8o.lean.js → reference_http-api.md.BwaCo-VO.lean.js} +0 -0
  342. /package/dist/docs/assets/{reference_sessions.md.CueyOHSL.lean.js → reference_sessions.md.CyXV1MUw.lean.js} +0 -0
  343. /package/dist/docs/assets/{skills_framework-map.md.D-tFZhFS.lean.js → skills_framework-map.md.BxLSOhSY.lean.js} +0 -0
@@ -4,349 +4,13 @@
4
4
 
5
5
  ---
6
6
 
7
- Source: /docs/ab.md
8
-
9
- # Live A/B metrics
10
-
11
- Use `defineAB` to compare variants on live agent sessions. New sessions
12
- receive a sticky assignment in each enrolled experiment. The Agent SDK folds
13
- their durable event streams into tool, token, failure, and wall-time
14
- metrics. You can send cumulative samples to your metrics backend and
15
- inspect aggregates in the playground.
16
-
17
- `defineAB` compares live variants through sticky assignment,
18
- instruction overlays, optional tool branches, and cumulative metrics.
19
- Metric callbacks observe the result without approving, rejecting, or
20
- failing a turn. Use [evals](/docs/evals.md) for pass/fail regression checks
21
- on fixed inputs.
22
-
23
- ## Choose live A/B metrics or evals
24
-
25
- Both features read the session event stream, but they answer different
26
- questions.
27
-
28
- | | Live A/B metrics | Evals |
29
- | --- | --- | --- |
30
- | Question | How do variants compare on live sessions? | Does the agent still meet a fixed contract? |
31
- | Location | `agent/ab.ts` or `agent/ab/<name>.ts` | `evals/**/*.eval.ts` |
32
- | Input | Dev or production traffic | Frozen prompts and fixtures |
33
- | Output | Cumulative metrics by session and arm | Pass/fail assertions |
34
- | How it runs | Automatically on new live sessions | `agent-sdk eval` |
35
-
36
- There is no `agent-sdk ab` command or assertion API.
37
-
38
- ## Define an experiment
39
-
40
- Author one experiment in `agent/ab.ts`, add more under
41
- `agent/ab/<name>.ts`, or use both forms. Each file defines one
42
- experiment. The experiment name comes from `name` when set. Otherwise,
43
- the Agent SDK uses `ab` for `agent/ab.ts` and the file stem for files under
44
- `agent/ab/`.
45
-
46
- ```ts
47
- // agent/ab/concise-weather.ts
48
- import {
49
- defineAB,
50
- splitBySessionHash,
51
- } from "@cursor/july/ab";
52
-
53
- export default defineAB({
54
- name: "concise-weather",
55
- variants: {
56
- control: {
57
- label: "Baseline",
58
- },
59
- treatment: {
60
- label: "Short replies",
61
- description: "Adds a one-paragraph response limit.",
62
- instructions: "Keep weather replies to one short paragraph.",
63
- },
64
- },
65
- split: splitBySessionHash({
66
- weights: { control: 1, treatment: 1 },
67
- holdout: 0.1,
68
- }),
69
- derive: {
70
- weatherCalls: (event) =>
71
- event.type === "action.result" &&
72
- event.data.toolName === "get_weather"
73
- ? 1
74
- : null,
75
- },
76
- onSample(sample) {
77
- console.log(
78
- sample.experiment,
79
- sample.variant,
80
- sample.metrics.toolCalls,
81
- sample.metrics.wallTimeMs
82
- );
83
- },
84
- });
85
- ```
86
-
87
- Every definition needs:
88
-
89
- - At least two variants. Variant keys cannot be empty or contain `/` or
90
- `\`.
91
- - A `split` function that returns a variant key or `null`.
92
- - An `onSample` callback for completed or failed turns.
93
-
94
- `label` and `description` appear with the arm in result surfaces.
95
- `instructions` changes the prompt for sessions in that arm. `derive`
96
- adds custom counters.
97
-
98
- Duplicate experiment names are validation errors. Check discovery
99
- before you serve:
100
-
101
- ```bash
102
- agent-sdk validate --dir .
103
- agent-sdk info --dir . --json
104
- ```
105
-
106
- The `abs` field in `info` lists the discovered experiment names.
107
-
108
- ## Assign sticky variants
109
-
110
- Enrollment happens once, when a live session is created and before its
111
- first turn:
112
-
113
- 1. The Agent SDK records `session.started`.
114
- 2. Each experiment runs its `split` function.
115
- 3. The Agent SDK records one durable `ab.assigned` event per experiment.
116
- 4. The selected arms become available on `session.abs`.
117
- 5. Variant instruction overlays reach the first model turn.
118
-
119
- A split can return a variant key or `null`. A null assignment is a
120
- sticky skip for that experiment. It increments the experiment's
121
- `skipped` total, still appears in the snapshot's `sessions` list with
122
- `variant: null`, and does not collect arm metrics or call `onSample`.
123
-
124
- Use the split helper that matches your rollout:
125
-
126
- | Helper | Behavior |
127
- | --- | --- |
128
- | `splitBySessionHash({ weights?, holdout?, salt? })` | Hashes the session id into a reproducible arm; the recommended default |
129
- | `splitByRandom({ weights?, holdout? })` | Draws once when the session starts, then persists the result |
130
- | `splitAlways("control")` | Pins every new session to one arm |
131
- | `splitNone()` | Skips every new session without deleting the experiment |
132
- | `splitIf(predicate, inner)` | Runs `inner` only when the predicate passes |
133
- | Custom `split(ctx)` | Returns a declared variant key or `null` |
134
-
135
- The split context includes the agent name, channel id, session info,
136
- experiment name, and declared variant keys. For example, enroll only
137
- Slack sessions:
138
-
139
- ```ts
140
- split: splitIf(
141
- (ctx) => ctx.channel.id === "slack",
142
- splitBySessionHash()
143
- ),
144
- ```
145
-
146
- Weights default to equal. Non-positive weights leave an arm out of the
147
- draw, and at least one arm must have a positive weight. `holdout` is the
148
- fraction of sessions assigned `null`, from `0` through `1`. Change
149
- `salt` to reshuffle future hash assignments without renaming the
150
- experiment.
151
-
152
- If a custom split throws or returns an unknown variant, the Agent SDK logs
153
- the error and records `variant: null`. The failed decision becomes a
154
- sticky skip instead of breaking the session.
155
-
156
- Enrollment only applies to new sessions. Adding an experiment does not
157
- assign existing conversations. Follow-ups keep the session's original
158
- arms. Keep experiment names and variant keys stable while you collect
159
- and compare results.
160
-
161
- ## Change behavior by variant
162
-
163
- Variant instructions are appended to the agent's base instructions.
164
- Local sessions receive the merged instructions in `AGENTS.md` before
165
- every turn. Cloud sessions receive them in the first-turn preamble
166
- only. For cloud follow-ups, branch through `session.abs` when the arm
167
- must remain visible to deterministic behavior.
168
-
169
- Tools can branch on the assignment through `ctx.session.abs`. Hooks can
170
- read the same field for logging or export:
171
-
172
- ```ts
173
- const treatment =
174
- ctx.session.abs?.["concise-weather"] === "treatment";
175
-
176
- if (treatment) {
177
- return conciseWeatherResult;
178
- }
179
-
180
- return baselineWeatherResult;
181
- ```
182
-
183
- This makes the assignment available to deterministic code as well as
184
- the model prompt. Use both patterns together when one experiment must
185
- steer the prompt and host code at once.
186
-
187
- `defineAB` does not select a different model or runtime for each arm.
188
- Keep those settings in `agent/agent.ts`, or write explicit host logic
189
- when your experiment needs another behavior lever.
190
-
191
- The split and selected arm can affect agent behavior. `derive` and
192
- `onSample` only observe the resulting event stream. Errors in either
193
- callback are logged and never fail the turn.
194
-
195
- ## Collect built-in and custom metrics
196
-
197
- Metrics accumulate for each session and experiment. When one session
198
- joins several experiments, every enrolled experiment folds the same
199
- turn and tool events into its own counters.
200
-
201
- | Metric | How the Agent SDK calculates it |
202
- | --- | --- |
203
- | `turns` | Adds one on `turn.completed` or `turn.failed` |
204
- | `turnFailures` | Adds one on `turn.failed` |
205
- | `toolCalls` | Adds one for each `action.result` |
206
- | `toolErrors` | Adds one when `action.result.data.isError` is true |
207
- | `inputTokens`, `outputTokens` | Adds usage from completed turns |
208
- | `cacheReadTokens`, `cacheWriteTokens` | Adds cache usage from completed turns |
209
- | `costUsd` | Sums the estimated turn cost recorded on `turn.completed` (turns whose model has no known rates contribute 0) |
210
- | `wallTimeMs` | Sums the time from `turn.started` to its completed or failed event |
211
- | `custom` | Sums finite numeric deltas returned by `derive` |
212
-
213
- `onSample` fires after every `turn.completed` and `turn.failed` event
214
- for an enrolled arm. The sample contains:
215
-
216
- | Field | Value |
217
- | --- | --- |
218
- | `experiment` | Experiment name |
219
- | `variant`, `variantLabel?` | Sticky arm and optional display label |
220
- | `sessionId`, `channelId` | Source session |
221
- | `metrics` | Cumulative metrics through this turn |
222
- | `reason` | `turn.completed` or `turn.failed` |
223
- | `at` | Terminal event timestamp |
224
-
225
- The metrics are cumulative, not per-turn deltas. A second sample from
226
- the same session includes the first turn's counts.
227
-
228
- Each `derive` extractor runs on every session event for its enrolled
229
- experiment, including streamed `message.appended` events. Keep it
230
- synchronous and cheap. Return a finite number to add a delta, or
231
- `null` to skip the event. Send samples to your metrics service from
232
- `onSample`; do not perform network or disk work in `derive`.
233
-
234
- Skipped sessions never call `onSample`. Errors from `derive` or
235
- `onSample` are logged, then metric collection continues.
236
-
237
- ## Inspect assignments and results
238
-
239
- Open the playground's **A/Bs** tab to see aggregate arm totals and
240
- per-session assignments. The tab reads `GET /v1/abs`.
241
-
242
- The response has two views of the same durable data:
243
-
244
- | Field | Contents |
245
- | --- | --- |
246
- | `experiments` | Declared variants, skipped-session count, arm session counts, and aggregate metrics |
247
- | `sessions` | Visible sessions with their assignments and cumulative metrics |
248
-
249
- `GET /v1/abs` returns sessions visible to the current principal by
250
- default. In `--dev`, loopback requests include every session. Add
251
- `--allow-anonymous` to include every session from non-loopback callers
252
- too.
253
-
254
- The session event stream is the source of truth for assignment + fold.
255
- `GET /v1/abs` recomputes aggregates from those logs. Any
256
- `agent/storage.ts` exports samples and snapshots durably: an authored
257
- `abs` table when the backend has a native shape for it, or the table
258
- derived over the KV core otherwise. See
259
- [Storage](/docs/storage.md#eval-and-a-b-tables).
260
-
261
- ## Configure the playground fold window
262
-
263
- Assignments and foldable metrics already persist in each session's
264
- event stream. The optional `agent/ab.config.ts` only caps how many
265
- sessions the playground and `GET /v1/abs` fold:
266
-
267
- ```ts
268
- import { defineABConfig } from "@cursor/july/ab";
269
-
270
- export default defineABConfig({
271
- // Optional. Defaults to 200. Only affects GET /v1/abs / A/Bs tab.
272
- maxPlaygroundSessions: 500,
273
- });
274
- ```
275
-
276
- `maxPlaygroundSessions` keeps the newest sessions in the fold. It does
277
- not prune session logs or change assignment. For export to S3, a DB, or
278
- your metrics vendor, send samples from `onSample` or declare a storage
279
- `abs` table.
280
-
281
- ## Keep assignments durable
282
-
283
- The append-only event stream is the source of truth. Each
284
- `ab.assigned` event persists a variant key or null skip. Built-in
285
- metrics come from the turn and tool events that follow it.
286
-
287
- After a server restart or a parked session resumes, the live collector
288
- replays the stream to rebuild cumulative counters. Replay does not call
289
- `onSample` (or write to the storage `abs` table) for historical turns.
290
- Only a new completed or failed turn emits another sample.
291
-
292
- The snapshot API also replays `derive` across the full stream, so
293
- custom totals match the current extractor. Changing a derive function
294
- can change historical snapshot totals. Treat metric definitions as
295
- versioned experiment code.
296
-
297
- ## Keep eval traffic separate
298
-
299
- Sessions created by `agent-sdk eval` and the playground Evals runner use
300
- `purpose: "eval"`. They skip A/B enrollment entirely:
301
-
302
- - No split function runs.
303
- - No `ab.assigned` event is recorded.
304
- - No `onSample` callback fires.
305
- - The session is omitted from `GET /v1/abs`.
306
-
307
- Ordinary chat, `agent-sdk run`, Slack, GitHub, and other channel sessions
308
- use the live purpose. You do not need `splitIf` to exclude eval traffic.
309
-
310
- ## Know the boundaries
311
-
312
- `defineAB` provides sticky assignment, variant instructions,
313
- `session.abs` for tools, cumulative metrics, and local inspection. It
314
- does not provide:
315
-
316
- - A test command, assertion API, or pass/fail result
317
- - Statistical significance calculations
318
- - An experiment rollout or lifecycle service
319
- - Per-variant model or runtime configuration
320
- - A built-in analytics warehouse (bring your own via `onSample` or the
321
- storage `abs` table)
322
-
323
- Use [evals](/docs/evals.md) to protect known behavior. Use `onSample` or a
324
- storage `abs` table when you need sample/snapshot exports beyond the
325
- session event log.
326
-
327
- ## What's next
328
-
329
- Continue with these pages:
330
-
331
- - [Evals](/docs/evals.md): pass/fail regression checks on fixed inputs
332
- - [Hillclimbing](/docs/hillclimbing.md): improve an agent against fixed
333
- fixtures
334
- - [Hooks](/docs/reference/hooks.md): other event-stream consumers
335
- - [Sessions and streaming](/docs/reference/sessions.md): the
336
- `ab.assigned` event and durable log
337
- - [Playground](/docs/reference/playground.md): the A/Bs tab
338
- - [HTTP API](/docs/reference/http-api.md): `GET /v1/abs`
339
- - [Live A/B metrics skill](/docs/skills/ab.md): have a coding agent
340
- wire an experiment
341
-
342
- ---
343
-
344
7
  Source: /docs/building-with-agents.md
345
8
 
346
9
  # Building agents with agents
347
10
 
348
11
  Give a coding agent the goal. The built-in skills guide it through
349
- scaffolding, channels, verification, evals, and measured improvement.
12
+ scaffolding, channels, verification, deployment, evals, and measured
13
+ improvement.
350
14
 
351
15
  ## What can a coding agent build for me?
352
16
 
@@ -393,12 +57,12 @@ The package ships task-specific guides under [`skills/`](/docs/skills/index.md):
393
57
  | Understand the project layout and runtimes | [`framework-map`](/docs/skills/framework-map.md) |
394
58
  | Create and verify a new agent | [`create-agent`](/docs/skills/create-agent.md) |
395
59
  | Write fixtures and regression checks | [`evals`](/docs/skills/evals.md) |
396
- | Live A/B metrics on traffic (`defineAB`) | [`ab`](/docs/skills/ab.md) |
397
60
  | Export OpenTelemetry traces | [`otel`](/docs/skills/otel.md) |
398
61
  | Improve an agent against fixed inputs | [`hillclimb`](/docs/skills/hillclimb.md) |
399
62
  | Add GitHub webhooks and replay events | [`github`](/docs/skills/github.md) |
400
63
  | Connect an agent to Slack | [`setup-slack`](/docs/skills/setup-slack.md) |
401
64
  | Authorize host MCP OAuth | [`mcp-auth`](/docs/skills/mcp-auth.md) |
65
+ | Deploy with an attached service account | [`deploy`](/docs/skills/deploy.md) |
402
66
  | Diagnose a local run | [`debug`](/docs/skills/debug.md) |
403
67
 
404
68
  Point your coding agent at the matching `SKILL.md`. The guide contains
@@ -505,7 +169,6 @@ name. For example, `agent/tools/get_weather.ts` creates a tool named
505
169
  | `agent/mcp-connections/<name>.ts` | Tools from external MCP servers |
506
170
  | `agent/host-connections/<name>.ts` | Privileged MCP servers for host tools only |
507
171
  | `agent/channels/*.ts` | HTTP, Slack, and GitHub entry points |
508
- | `agent/ab.ts` or `agent/ab/*.ts` | Sticky variants and live performance metrics |
509
172
  | `agent/result.ts` | Optional host `commit` on the final assistant text |
510
173
  | `evals/**/*.eval.ts` | Repeatable checks at the project root |
511
174
 
@@ -541,8 +204,8 @@ Each session records an append-only event stream. It includes:
541
204
  - Turn completion and token usage
542
205
 
543
206
  Sessions and their event streams survive server restarts. The
544
- playground renders the stream. Evals assert against it. The
545
- `agent-sdk trajectory` command turns a saved stream into a short
207
+ playground renders the stream. [Evals](/docs/evals.md) assert against it.
208
+ The `agent-sdk trajectory` command turns a saved stream into a short
546
209
  summary.
547
210
 
548
211
  When a run surprises you, inspect its event stream first. See
@@ -636,7 +299,6 @@ See [Agent-to-agent](/docs/guides/agent-to-agent.md) for a complete example.
636
299
  - [Project layout](/docs/reference/project-layout.md)
637
300
  - [Sessions and streaming](/docs/reference/sessions.md)
638
301
  - [Channels](/docs/reference/channels.md)
639
- - [Live A/B metrics](/docs/ab.md)
640
302
 
641
303
  ---
642
304
 
@@ -3240,31 +2902,64 @@ Source: /docs/evals.md
3240
2902
 
3241
2903
  # Evals
3242
2904
 
3243
- An eval is a repeatable check that runs your agent against a fixed input
3244
- and gates the recorded trajectory: the run completed, the right tool
3245
- ran, the reply has the right shape. Evals are how you know a prompt
3246
- tweak helped, a refactor didn't regress the agent, and last month's fix
3247
- is still holding.
2905
+ An eval sends a fixed message to your agent and asserts over the
2906
+ trajectory it records: the turn completed, the right tool ran with the
2907
+ right input, the reply has the right shape. Evals are how you know a
2908
+ prompt tweak helped, a refactor didn't regress the agent, and last
2909
+ month's fix still holds.
3248
2910
 
3249
- Evals exercise the same surface your users hit. The runner starts (or
3250
- targets) a real agent server, drives sessions over the public API, and
3251
- grades what comes back. A passing eval means the agent started,
3252
- accepted a message, and did what you asserted.
2911
+ Nothing is mocked. The runner starts (or targets) a real agent server,
2912
+ drives sessions over the public API, and grades the events it gets
2913
+ back. The model runs and server tools execute, so
2914
+ [keep side effects out of eval sessions](#keep-side-effects-out-of-eval-sessions)
2915
+ before you point an eval at an agent that posts anywhere.
3253
2916
 
3254
- ## Define evals with `defineEval`
2917
+ ## Evals, hooks, or hillclimbing?
3255
2918
 
3256
- The Agent SDK discovers evals under the project-root `evals/` directory,
3257
- in `.eval.ts` or `.eval.js` files. That's a sibling of `agent/`, never
3258
- inside it (`agent/evals/` is silently ignored). TypeScript is the normal
3259
- authoring format.
2919
+ All three read the same session event stream. Pick by the question you
2920
+ are asking.
3260
2921
 
3261
- The file path is the eval's identity, so you don't author an id.
3262
- Directories group related evals: `evals/builds/api.eval.ts` becomes id
3263
- `builds/api`. An `index` filename collapses to its directory, so
3264
- `evals/builds/index.eval.ts` becomes `builds`.
2922
+ | You want to | Use |
2923
+ | --- | --- |
2924
+ | Gate one fixed input's behavior, locally and in CI | Evals (this page) |
2925
+ | Observe every live session: metrics, audit, alerts | [Hooks](/docs/reference/hooks.md) |
2926
+ | Improve an agent one measured round at a time | [Hillclimbing](/docs/hillclimbing.md); each kept win lands an eval |
2927
+
2928
+ [Hooks, channel events, or evals?](/docs/reference/hooks.md#hooks-channel-events-or-evals)
2929
+ has the side-by-side table.
2930
+
2931
+ ### When not to write an eval
2932
+
2933
+ - Test a server tool's own logic with
2934
+ `agent-sdk call <tool> --dir . --input '{...}'` or a unit test. No
2935
+ model turn, no credential.
2936
+ - Explore a prompt with `agent-sdk run --dir . --message "..."` and
2937
+ read the trajectory. Write the eval once you know which decision to
2938
+ gate.
2939
+ - Stop a bad turn while it runs with
2940
+ [`needsApproval`](/docs/reference/tools.md#gate-a-tool-on-human-approval)
2941
+ on the tool or [`defineResult`](/docs/reference/result.md). Evals grade
2942
+ after the fact.
2943
+
2944
+ ## Write your first eval
3265
2945
 
3266
- An eval is a single `async test(t)`. You drive the agent with `t` and
3267
- assert on the run with the same `t`:
2946
+ Evals live under the project-root `evals/` directory, a sibling of
2947
+ `agent/`. `agent/evals/` is silently ignored. Discovery loads every
2948
+ `.eval.ts` (or `.eval.js`) file under `evals/`, plus one config file.
2949
+
2950
+ ```text
2951
+ my-agent/
2952
+ agent/
2953
+ agent.ts
2954
+ tools/inspect_pr.ts
2955
+ evals/
2956
+ evals.config.ts # required to run: maxConcurrency
2957
+ readiness.eval.ts # id: readiness
2958
+ prs.eval.ts # cases: prs/checkout, prs/search
2959
+ ```
2960
+
2961
+ An eval is a single `async test(t)`. You drive the agent with `t.send`
2962
+ and assert on the recorded run with the same `t`:
3268
2963
 
3269
2964
  ```ts
3270
2965
  // evals/readiness.eval.ts
@@ -3286,12 +2981,49 @@ export default defineEval({
3286
2981
  });
3287
2982
  ```
3288
2983
 
3289
- One file can also hold several datapoints through `cases` (provide
3290
- either `test` or `cases`, not both). Each case id becomes
3291
- `<fileId>/<case.id>`:
2984
+ ```ts
2985
+ // evals/evals.config.ts
2986
+ import { defineEvalConfig } from "@cursor/july/evals";
2987
+
2988
+ export default defineEvalConfig({ maxConcurrency: 20 });
2989
+ ```
2990
+
2991
+ Run it under Node 22.13 or newer (never Bun) with a Cursor credential
2992
+ in place; see [Credentials](#credentials):
2993
+
2994
+ ```bash
2995
+ agent-sdk eval --dir . --list
2996
+ agent-sdk eval --dir . readiness
2997
+ ```
2998
+
2999
+ ```text
3000
+ PASS readiness (14.2s) — Inspects a PR without approving it.
3001
+ ✓ succeeded
3002
+ ✓ calledTool(inspect_pr)
3003
+ ✓ notCalledTool(approve_pr)
3004
+ ✓ check(includes)
3005
+
3006
+ 1 passed, 0 failed, 1 total
3007
+ artifacts: <project state directory>/evals/2026-09-11T15-02-11-402Z
3008
+ ```
3009
+
3010
+ Every local run writes each case's assertions, inputs, tool calls, and
3011
+ `t.log` lines under that artifacts directory. Open
3012
+ `evals/<case-id>.json` there when a case fails; see
3013
+ [Where results land](#where-results-land).
3014
+
3015
+ ## Name cases by path
3016
+
3017
+ The file path is the eval's identity, so you don't author an id.
3018
+ `evals/builds/api.eval.ts` becomes `builds/api`. An `index` filename
3019
+ collapses to its directory: `evals/builds/index.eval.ts` becomes
3020
+ `builds`.
3021
+
3022
+ One file can hold several datapoints through `cases`. Provide either
3023
+ `test` or `cases`, not both. Each case id becomes `<fileId>/<case.id>`:
3292
3024
 
3293
3025
  ```ts
3294
- // evals/prs.eval.ts prs/checkout, prs/search
3026
+ // evals/prs.eval.ts: prs/checkout, prs/search
3295
3027
  export default defineEval({
3296
3028
  tags: ["smoke", "prs"],
3297
3029
  cases: [
@@ -3320,93 +3052,75 @@ export default defineEval({
3320
3052
  });
3321
3053
  ```
3322
3054
 
3323
- Case ids must be single path segments, unique within the file.
3324
- Each case can set its own `description`, `tags`, `timeoutMs`, and
3325
- `iterations`. A case-level value replaces the file-level value for that
3326
- datapoint.
3055
+ Case ids are single path segments, unique within the file. A case can
3056
+ set its own `description`, `tags`, `timeoutMs`, `iterations`, `judge`,
3057
+ `reporters`, and `metadata`. A case-level value replaces the file-level
3058
+ one for that datapoint, except `metadata`, which merges with case keys
3059
+ winning, and `reporters`, which adds to the file's list. `metadata` is
3060
+ free-form data carried onto the result and every reporter.
3061
+
3062
+ A file may instead export an array of `defineEval` calls to fan out
3063
+ over a dataset. Ids are then the file id plus a zero-padded index
3064
+ (`sql/0000`, `sql/0001`, ...); see [Load a dataset](#load-a-dataset).
3065
+ Prefer `cases` when datapoints are hand-written and deserve stable
3066
+ names.
3327
3067
 
3328
3068
  ### Iterations
3329
3069
 
3330
- `iterations` (file or case, default `1`) runs a datapoint repeatedly.
3331
- Discovery expands `iterations: 3` on case `nyc` to runnable ids
3332
- `weather/nyc/1`, `weather/nyc/2`, `weather/nyc/3` (filter prefix
3333
- `weather/nyc` still selects all three). Each expanded case exposes
3334
- `t.iteration` / `t.iterations` on the test context. Cap is 100.
3070
+ `iterations` (file or case, default `1`, cap `100`) runs a datapoint
3071
+ repeatedly. Discovery expands `iterations: 3` on case `nyc` to runnable
3072
+ ids `weather/nyc/1`, `weather/nyc/2`, and `weather/nyc/3`. The filter
3073
+ `weather/nyc` still selects all three. Each expanded case exposes
3074
+ `t.iteration` and `t.iterations`.
3335
3075
 
3336
- `maxConcurrency` counts **authored datapoints**, not expanded
3337
- iterations: siblings `…/1`…`…/n` share one concurrency slot and run
3338
- sequentially. A suite with 11 cases × 3 iterations and
3339
- `maxConcurrency: 20` therefore has at most 11 cases in flight, not 33.
3076
+ `maxConcurrency` counts authored datapoints, not expanded iterations.
3077
+ Iterations of one datapoint share a concurrency slot and run in
3078
+ sequence, so a suite of 11 cases with 3 iterations each and
3079
+ `maxConcurrency: 20` has at most 11 cases in flight.
3340
3080
 
3341
- ## Configure eval runs
3081
+ ## Drive the agent with `t.send`
3342
3082
 
3343
- Each project with evals needs `evals/evals.config.ts` or
3344
- `evals/evals.config.js`, and it must set `maxConcurrency`. Each case
3345
- issues real model-provider requests, so concurrency is capped hard at
3346
- 200. Existing projects use 20. Discovery with `eval --list` works
3347
- without this file, but running a case does not.
3083
+ `t.send(message, options?)` runs one turn and waits for it to settle:
3084
+ complete, park on an approval request, or fail. Several sends in one
3085
+ case share the session, which is how you write multi-turn evals.
3348
3086
 
3349
- ```ts
3350
- import { defineEvalConfig } from "@cursor/july/evals";
3087
+ Each send resolves to a turn result: `message` (the assistant text),
3088
+ `sessionId`, `events`, `toolCalls` (tool names in order), `ok`, and
3089
+ `index`. The turn carries the same assertion vocabulary as `t`, scoped
3090
+ to that turn, so you can grade an intermediate turn before the next
3091
+ send overwrites `t.reply`. `turn.expectOk()` throws when the turn
3092
+ failed, for later steps that depend on it.
3351
3093
 
3352
- export default defineEvalConfig({
3353
- maxConcurrency: 20, // required
3354
- // timeoutMs: 180_000, // optional project-wide default
3355
- // judge: { model: "..." }, // default judge model for t.judge.*
3356
- // reporters: [], // destinations that observe every case
3357
- // maxPlaygroundRuns: 50, // playground history only (default 20)
3094
+ Read the whole case with `t.reply` (last assistant text), `t.events`
3095
+ (every event so far), `t.turns` (settled turns, oldest first), and
3096
+ `t.sessionId`. `t.signal` aborts when the case hits its timeout; pass
3097
+ it to your own async work.
3098
+
3099
+ Three options apply on the first send only, because they shape session
3100
+ creation:
3101
+
3102
+ | Option | Effect |
3103
+ | --- | --- |
3104
+ | `workspaceFiles` | `{ path: contents }` seeded into the local session workspace. Prefer this over machine-local paths |
3105
+ | `workspaceDir` | Absolute harness cwd for the local runtime |
3106
+ | `cloud` | Per-session cloud options merged over the agent's static `cloud` config. Pin a fixture repo here for cloud evals instead of on the agent's default `cloud.repos`. On the cloud runtime, seeded files reach the agent as described under [Where does a turn run?](/docs/concepts.md#where-does-a-turn-run) |
3107
+
3108
+ ```ts
3109
+ await t.send("Review pr/diff.patch and post findings.", {
3110
+ workspaceFiles: {
3111
+ "pr/diff.patch": [
3112
+ "diff --git a/app/routes/search.ts b/app/routes/search.ts",
3113
+ "+res.send(`<h1>Results for ${req.query.q}</h1>`);",
3114
+ ].join("\n"),
3115
+ },
3358
3116
  });
3359
3117
  ```
3360
3118
 
3361
- The timeout order is case or file `timeoutMs`, CLI `--timeout-ms`,
3362
- project config `timeoutMs`, then the 180-second runner default.
3363
-
3364
- The optional fields:
3119
+ ## Assert over the trajectory
3365
3120
 
3366
- | Option | Default | Meaning |
3367
- | --- | --- | --- |
3368
- | `timeoutMs` | `180_000` | Project-wide per-case timeout |
3369
- | `judge` | unset | Default judge model for `t.judge.*`; see [Judge free-form output](#judge-free-form-output) |
3370
- | `reporters` | unset | Destinations that observe every case; `--skip-report` suppresses them |
3371
- | `maxPlaygroundRuns` | `20` | Max batches in the playground / `/v1/dev/evals*` history (not CLI `eval`). Hard-capped at 500. |
3372
-
3373
- Reporters come from `@cursor/july/evals/reporters`: `JUnit` writes a
3374
- JUnit XML file for CI, `Artifacts` writes per-case files, and
3375
- `combineReporters` merges several into one (`renderJUnitXml` renders
3376
- the XML for a custom destination). A file or case can add its own
3377
- `reporters` on top of the config list.
3378
-
3379
- Playground batches survive restarts whenever `agent/storage.ts` exists
3380
- with an `evals` table or a KV core providing `delete` and `list` (the
3381
- table is derived over the core); see
3382
- [Storage](/docs/storage.md#eval-and-a-b-tables). Without storage they live
3383
- in process memory and disappear when `serve` exits. Navigating away
3384
- and back still works while the process is up.
3385
-
3386
- ## Drive and assert with `t`
3387
-
3388
- `t` is both the driver and the assertion surface. You write ordinary
3389
- control flow, sending turns and asserting inline.
3390
-
3391
- Drive the agent with `t.send(message, options?)`. It runs one turn and
3392
- waits for the session to park or fail. Multiple sends in one case share
3393
- the session, which is how you write multi-turn evals.
3394
-
3395
- Each `t.send` resolves to a turn result with `message`,
3396
- `sessionId`, `events`, `toolCalls`, `ok`, and `index`. The turn carries
3397
- the same assertion vocabulary as `t`, scoped to that turn, so you can
3398
- grade an intermediate turn before the next send overwrites `t.reply`.
3399
- `turn.expectOk()` throws when the turn failed, for later
3400
- steps that depend on it.
3401
-
3402
- Read the full case state with `t.reply` (the last assistant text),
3403
- `t.events` (session events captured so far), `t.turns` (settled
3404
- turns, oldest first), and `t.sessionId`. `t.signal` aborts when the
3405
- case hits its timeout; pass it to your own async work. A thrown
3406
- [turn result](/docs/reference/result.md) `commit` fails the turn, so
3407
- `t.succeeded()` fails too.
3408
-
3409
- Assert with the gates:
3121
+ Assertions record; they never throw. One run reports every failure
3122
+ instead of dying on the first. Assertions on `t` read the whole run.
3123
+ Assertions on a turn read only that turn.
3410
3124
 
3411
3125
  | Gate | Checks |
3412
3126
  | --- | --- |
@@ -3415,158 +3129,483 @@ Assert with the gates:
3415
3129
  | `t.messageIncludes(token)` | the joined assistant text matches a string or `RegExp` |
3416
3130
  | `t.calledTool(name, matcher?)` | a matching call to `name` happened |
3417
3131
  | `t.notCalledTool(name)` | no request for `name`, in any lifecycle state |
3418
- | `t.loadedSkill(name)` | the agent opened the skill's `SKILL.md` (read, grep, or shell `cat`) |
3419
- | `t.toolOrder(names)` | tool requests appear in this relative order (extra calls allowed) |
3132
+ | `t.loadedSkill(name)` | the agent opened `skills/<name>/SKILL.md` (read, grep, or shell `cat`) |
3133
+ | `t.toolOrder(names)` | tool requests appear in this relative order; extra calls allowed |
3420
3134
  | `t.usedNoTools()` | no tool calls at all |
3421
3135
  | `t.maxToolCalls(max)` | at most `max` tool calls |
3422
3136
  | `t.noFailedActions()` | no tool call reported an error |
3423
3137
  | `t.calledSubagent(name, matcher?)` | a matching subagent delegation happened |
3424
3138
  | `t.taggedArtifact(kind?, predicate?)` | at least one [artifact](/docs/reference/artifacts.md) was tagged |
3425
- | `t.event(type, matcher?)` | at least one matching event of `type` occurred |
3426
- | `t.notEvent(type, matcher?)` | no matching event of `type` occurred |
3139
+ | `t.event(type, matcher?)` | at least one matching [event](/docs/reference/sessions.md#which-events-can-i-stream) of `type` |
3140
+ | `t.notEvent(type, matcher?)` | no matching event of `type` |
3427
3141
  | `t.eventOrder(matchers)` | matching event groups occur in this relative order |
3428
3142
  | `t.eventsSatisfy(label, predicate)` | your predicate over the typed event stream |
3429
- | `t.check(value, expectation)` | any value, against a builder |
3430
- | `t.score(name, value)` | records a 01 score you computed; soft until you add a bar |
3431
- | `t.requireToolCall(name, matcher?)` | gates on a matching call and returns it, so later code can read its input and output |
3432
- | `t.requireInputRequest(filter?)` | gates on exactly one pending approval request and returns it |
3143
+ | `t.check(value, expectation)` | any value, against a [builder](#grade-values-with-expectation-builders) |
3144
+ | `t.score(name, value)` | a 0-1 score you computed; soft until you add a bar |
3145
+
3146
+ Three more assertions gate and return the matched fact. They stop the
3147
+ test body when nothing matches, without a duplicate execution error.
3148
+ `t.requireToolCall(name, matcher?)` returns the call so later code can
3149
+ read its `input` and `output`. `t.requireInputRequest(filter?)` returns
3150
+ the single pending approval request. `await t.require(value, expectation)`
3151
+ does the same for a value check.
3152
+
3153
+ A case with no assertions passes when at least one turn completed. Add
3154
+ `t.succeeded()` and behavior gates anyway. They make the contract
3155
+ visible in review.
3433
3156
 
3434
- Every gate returns a handle: `.soft()` demotes it to tracked-only,
3435
- `.atLeast(0.7)` adds a soft score bar, and `.gate(0.8)` promotes a
3436
- scored assertion into a hard gate.
3157
+ ### What good cases assert
3158
+
3159
+ Gate decisions and shape, not prose. Model wording varies run to run.
3160
+ Tool choice, tool avoidance, and output structure are the stable
3161
+ contract.
3162
+
3163
+ 1. `t.succeeded()`: always, first.
3164
+ 2. The tool decision: `calledTool` for the intended path,
3165
+ `notCalledTool` for the likely wrong alternative. The pair is
3166
+ stronger than either alone.
3167
+ 3. Output shape: a regex for the contract (`/ready|blocked/i`, a JSON
3168
+ marker, a findings-block fence), never exact sentences.
3169
+ 4. For structured output, parse `t.reply` and check fields with
3170
+ `matches` or `satisfies` instead of substring-matching JSON.
3171
+
3172
+ The common failure modes: asserting exact phrasing, packing more than
3173
+ about five gates into one case (split it), and cases that depend on
3174
+ live external state that drifts (pin the input).
3175
+
3176
+ ### Narrow tool assertions with matchers
3437
3177
 
3438
3178
  With no matcher, `calledTool` is request-based: a requested call counts
3439
- even when its result has not arrived. Pass
3440
- `t.calledTool("inspect_pr", { status: "completed" })` to require the
3441
- call to return. `input`, `output`, and `count` matcher fields accept a
3442
- literal, a `RegExp`, or a predicate.
3443
-
3444
- The expectation builders are `includes(string | RegExp)`,
3445
- `equals(value)`, `matches(schema)`, `similarity(expected)`, and
3446
- `satisfies(predicate, label)`. `includes` stringifies its input,
3447
- `equals` compares values deeply, `matches` validates against a Standard
3448
- Schema (or anything with `safeParse`, like Zod), `similarity` scores
3449
- normalized text similarity, and `satisfies` runs your predicate. The
3450
- plain function `normalizedSimilarity(actual, expected)` returns the
3451
- same 0–1 score for use with `t.score`.
3452
-
3453
- A few more context members shape a case: `t.require(value, expectation)`
3454
- records a gate and stops the test body when it fails, without a
3455
- duplicate execution error. `t.skip(reason)` ends the case as skipped
3456
- (reported separately, never changes the exit code; call it before
3457
- sending messages). `t.metric(name, value)` records a structured score
3458
- for the playground case card. `t.log(message)` records a debug line for
3459
- the CLI and playground result.
3460
-
3461
- Three `t.send` options apply on session create (first `t.send` only):
3462
-
3463
- - `workspaceFiles`: `{ path: contents }`, seeded into the local session
3464
- workspace. Prefer this over machine-local paths.
3465
- - `workspaceDir`: absolute harness cwd (local runtime).
3466
- - `cloud`: per-session cloud options merged over the agent's static
3467
- `cloud` config (repos / env / …). Use a pinned `repos` override to
3468
- attach a fixture repo for cloud evals without putting it on the
3469
- agent's default `cloud.repos`. Cloud ignores `workspaceFiles` seeds.
3470
-
3471
- ```ts
3472
- const toolResults = t.events.filter((e) => e.type === "action.result");
3179
+ even before its result arrives. A matcher narrows it:
3180
+
3181
+ ```ts
3182
+ t.calledTool("inspect_pr", { status: "completed" });
3183
+ t.calledTool("apply_agents", { input: { verdict: "update" } });
3184
+ t.calledTool("bash", { input: { command: /^gh pr view/ }, count: 1 });
3185
+ t.calledTool("read_file", {
3186
+ output: (value) => String(value).includes("TODO"),
3187
+ });
3188
+ ```
3189
+
3190
+ `input`, `output`, and `count` accept a literal, a `RegExp`, or a
3191
+ predicate. Object literals partial-deep-match, so `{ verdict: "update" }`
3192
+ matches arguments that also carry other keys. `status` is one of
3193
+ `completed`, `failed`, `pending`, or `rejected` (a human denied the
3194
+ approval). `calledSubagent` takes `{ output, status, count, callId }`.
3195
+ `event`, `notEvent`, and `eventOrder` take `{ data, count }`.
3196
+
3197
+ ### Grade values with expectation builders
3198
+
3199
+ `t.check(value, expectation)` grades any value: `t.reply`, a parsed
3200
+ JSON field, a tool's output.
3201
+
3202
+ | Builder | Checks | Severity |
3203
+ | --- | --- | --- |
3204
+ | `includes(string \| RegExp)` | substring or match; structured values are stringified first | gate |
3205
+ | `equals(value)` | deep equality | gate |
3206
+ | `matches(schema)` | a Standard Schema (Zod, Valibot, ...) or anything with `safeParse` | gate |
3207
+ | `similarity(expected)` | normalized text similarity, 0-1 | soft |
3208
+ | `satisfies(predicate, label)` | your predicate; `label` is the failure detail | gate |
3209
+
3210
+ ```ts
3211
+ import { matches, satisfies } from "@cursor/july/evals";
3212
+ import { z } from "zod";
3213
+
3214
+ const verdict = JSON.parse(t.reply ?? "{}");
3215
+ t.check(
3216
+ verdict,
3217
+ matches(z.object({ ready: z.boolean(), blockers: z.array(z.string()) }))
3218
+ );
3473
3219
  t.check(
3474
- toolResults.length,
3475
- satisfies((n) => (n as number) <= 4, "at most 4 tool calls")
3220
+ verdict.blockers.length,
3221
+ satisfies((n) => (n as number) <= 3, "at most 3 blockers")
3476
3222
  );
3477
3223
  ```
3478
3224
 
3479
- A case with no explicit gates falls back to whether at least one turn
3480
- completed successfully. Add `t.succeeded()` and behavior-specific gates
3481
- anyway. They make the contract visible during review.
3225
+ `normalizedSimilarity(actual, expected)` returns the same 0-1 score as
3226
+ `similarity`, for use with `t.score`.
3227
+
3228
+ ### Record without gating
3229
+
3230
+ - `t.metric(name, value)` records a structured score or label. It shows
3231
+ on the CLI result, the playground case card, JUnit output, and
3232
+ artifacts.
3233
+ - `t.log(message)` records a debug line, streamed under `--verbose`.
3234
+ - `t.skip(reason)` ends the case as skipped. Skipped cases report
3235
+ separately and never change the exit code. Call it before sending
3236
+ messages.
3237
+
3238
+ ## Gates, soft scores, and verdicts
3239
+
3240
+ Every assertion returns a handle, so severity rides on the assertion
3241
+ instead of a separate thresholds map:
3242
+
3243
+ ```ts
3244
+ t.succeeded(); // gate (default)
3245
+ t.calledTool("get_weather").soft(); // tracked, never fails
3246
+ t.check(t.reply, similarity("Sunny, 72°F")).atLeast(0.8); // soft, with a bar
3247
+ t.judge.closedQA("cites a source").gate(0.8); // promoted to a gate
3248
+ ```
3249
+
3250
+ - `.gate(threshold?)` is hard. A miss fails the case and `eval` exits 1.
3251
+ - `.soft(threshold?)` is tracked. With no threshold it never fails.
3252
+ - `.atLeast(threshold)` is soft with a bar. A miss marks the case
3253
+ `scored`.
3482
3254
 
3483
- ### Judge free-form output
3255
+ Each case ends with one verdict:
3484
3256
 
3485
- When wording matters and no regex captures it, `t.judge` grades the
3486
- reply with an LLM. The built-in graders are `factuality(expected)`,
3487
- `summarizes(expected)`, `closedQA(criteria)`, and `sql(expected)`. Each
3488
- scores `t.reply` by default; pass `{ on }` to grade another value.
3257
+ | Verdict | Meaning | Exit code |
3258
+ | --- | --- | --- |
3259
+ | `passed` | every gate passed and no soft bar was missed | 0 |
3260
+ | `failed` | a gate failed, or the test body threw | 1 |
3261
+ | `scored` | only soft bars were missed | 0, or 1 under `--strict` |
3262
+ | `skipped` | `t.skip(reason)`, or a judge with no credentials | 0 |
3263
+
3264
+ The CLI prints soft misses as `~` and gate misses as `✗`. Start a new
3265
+ benchmark with `t.score("recall", recall)` and `.atLeast()` so its
3266
+ number reports for a while without blocking merges. Add `--strict`
3267
+ once the bars are trustworthy.
3268
+
3269
+ ## Judge free-form output
3270
+
3271
+ When wording matters and no regex captures it, `t.judge` grades with an
3272
+ LLM. The graders are `factuality(expected)`, `summarizes(expected)`,
3273
+ `closedQA(criteria)`, and `sql(expected)`. Each scores `t.reply` by
3274
+ default; pass `{ on }` to grade another value.
3489
3275
 
3490
3276
  ```ts
3491
- t.judge.factuality("It is 54°F in NYC right now.").atLeast(0.7);
3277
+ const summary = await t.send("Why did CI fail on PR 42?");
3278
+ t.judge.factuality("The lint step failed on src/sidebar.ts.").atLeast(0.7);
3279
+ t.judge.closedQA("names the failing step", { on: summary.message }).gate(1);
3492
3280
  ```
3493
3281
 
3494
3282
  Judge assertions are soft by default, so a judge never fails a build
3495
- until you give it a bar with `.atLeast(0.7)` or promote it with
3496
- `.gate(0.8)`. The judge model comes from `defineEvalConfig({ judge })`,
3283
+ until you give it a bar with `.atLeast()` or promote it with `.gate()`.
3284
+ The recorded detail names the choice the judge made and its rationale.
3285
+
3286
+ The judge model comes from `defineEvalConfig({ judge })`,
3497
3287
  `defineEval({ judge })`, a case-level `judge`, or a per-call
3498
- `{ model }` override; the nearest one wins. For a domain-specific judge
3499
- whose verdict is not a single score, `t.judge.model(prompt)` sends a
3500
- raw prompt to the same model and returns the reply. You then record the
3501
- parsed result with `t.score` or `t.check`.
3288
+ `{ model }`. The nearest one wins. A judge call with no model
3289
+ configured fails the case. A judge that cannot reach a model (no
3290
+ credential) ends the case as `skipped`, unless a deterministic gate
3291
+ already failed.
3292
+
3293
+ For a domain-specific judge whose verdict is not a single score,
3294
+ `t.judge.model(prompt)` sends a raw prompt to the same model and
3295
+ returns the reply. Record the parsed result with `t.score` or
3296
+ `t.check`. Anything derived from the agent under test is untrusted
3297
+ input to your prompt: wrap it with `fenceUntrusted` and include
3298
+ `EVAL_JUDGE_INJECTION_GUARD`, as the built-in graders do.
3299
+
3300
+ ```ts
3301
+ import {
3302
+ EVAL_JUDGE_INJECTION_GUARD,
3303
+ fenceUntrusted,
3304
+ } from "@cursor/july/evals";
3305
+
3306
+ const gold = ["XSS in search.ts", "open redirect in login.ts"];
3307
+ const reply = await t.judge.model(
3308
+ [
3309
+ "For each GOLD finding, answer whether SUBMISSION reports it.",
3310
+ "Reply with one line per finding: <index> YES|NO.",
3311
+ EVAL_JUDGE_INJECTION_GUARD,
3312
+ fenceUntrusted("GOLD", gold.map((g, i) => `${i + 1}. ${g}`).join("\n")),
3313
+ fenceUntrusted("SUBMISSION", t.reply ?? ""),
3314
+ ].join("\n\n")
3315
+ );
3316
+ const hits = reply.match(/\bYES\b/g)?.length ?? 0;
3317
+ t.score("recall", hits / gold.length).atLeast(0.5);
3318
+ ```
3319
+
3320
+ ## Keep side effects out of eval sessions
3321
+
3322
+ Eval sessions run the real agent, tools included. A reviewer that
3323
+ comments on GitHub or posts to Slack will do so from an eval unless
3324
+ the tool checks the session's purpose. Eval sessions carry
3325
+ `purpose: "eval"`; live traffic carries `"live"`. Branch on it in the
3326
+ tool, hook, or result handler that actuates:
3327
+
3328
+ ```ts
3329
+ // agent/tools/post_findings.ts
3330
+ async execute({ findings }, ctx) {
3331
+ if (ctx.session.purpose === "eval") {
3332
+ return { posted: false, reason: "eval", count: findings.length };
3333
+ }
3334
+ // post the review
3335
+ }
3336
+ ```
3337
+
3338
+ Return a shaped result instead of throwing, so the eval can still
3339
+ assert `t.calledTool("post_findings", { input: ... })` on the decision.
3340
+ The same check belongs in [hooks](/docs/reference/hooks.md) that meter or
3341
+ page and in [`defineResult`](/docs/reference/result.md) commits.
3342
+
3343
+ ## Worked examples
3344
+
3345
+ ### Multi-turn: grade each turn
3346
+
3347
+ ```ts
3348
+ // evals/intro.eval.ts
3349
+ import { defineEval, includes, satisfies } from "@cursor/july/evals";
3350
+
3351
+ export default defineEval({
3352
+ description: "Introduces itself once; a repeat mention gets a short ack.",
3353
+ async test(t) {
3354
+ const intro = await t.send("Meet Jenny! @Jenny introduce yourself.");
3355
+ intro.expectOk();
3356
+ t.check(intro.message, includes(/jenny/i));
3357
+
3358
+ const repeat = await t.send("Meet, @Jenny!");
3359
+ t.succeeded();
3360
+ repeat.usedNoTools();
3361
+ t.check(
3362
+ repeat.message,
3363
+ satisfies((r) => (r as string).trim().length <= 280, "short ack")
3364
+ );
3365
+ t.check(
3366
+ repeat.message,
3367
+ satisfies((r) => !/what i can do/i.test(r as string), "no second intro")
3368
+ );
3369
+ },
3370
+ });
3371
+ ```
3372
+
3373
+ `t.succeeded()` grades the whole session. `repeat.usedNoTools()` and the
3374
+ checks on `repeat.message` read only the second turn, even though
3375
+ `t.reply` now holds its text.
3376
+
3377
+ ### Approvals: assert the parked decision
3378
+
3379
+ For a tool with `needsApproval`, the turn parks instead of finishing.
3380
+ Gate on `t.parked()` and on the arguments the model chose:
3381
+
3382
+ ```ts
3383
+ // evals/agents.eval.ts (one case; RULE and SLACK are fixture strings)
3384
+ {
3385
+ id: "update-rule",
3386
+ description: "A repeated billing rule parks the AGENTS.md write.",
3387
+ async test(t) {
3388
+ await t.send("Weekly AGENTS.md review. Read week/ and call apply_agents once.", {
3389
+ workspaceFiles: {
3390
+ "week/prs.md": RULE,
3391
+ "week/slack.md": SLACK,
3392
+ "week/tree/AGENTS.md.txt": "# API\n\nKeep handlers thin.\n",
3393
+ },
3394
+ });
3395
+ t.parked();
3396
+ t.calledTool("apply_agents", { input: { verdict: "update" } });
3397
+ },
3398
+ },
3399
+ ```
3400
+
3401
+ `t.parked()` and `t.succeeded()` are exclusive: a parked run is a clean
3402
+ stop on an unanswered approval, not a completed one. Pair the parked
3403
+ case with a sibling that expects `verdict: "skip"` and `t.succeeded()`,
3404
+ so both branches stay pinned.
3405
+
3406
+ ## Pin fixtures
3407
+
3408
+ A fixed input is what makes an eval repeatable. Pick the fixture by the
3409
+ surface under test.
3410
+
3411
+ | Agent surface | Fixture |
3412
+ | --- | --- |
3413
+ | Chat or domain assistant | One canonical prompt string, chosen once and frozen |
3414
+ | Tool-heavy | Run `agent-sdk call <tool>` first to pin what the tool returns, then freeze the prompt that triggers it |
3415
+ | GitHub webhook | `agent-sdk github replay <pr> --events '*' --dry-run --out fixtures/github` snapshots real payloads for offline replay ([GitHub guide](/docs/guides/github.md#test-with-github-replay)) |
3416
+ | PR reviewer with host preparation | Diff, metadata, and gold labels pinned to commit SHAs; keep any live PR matrix small |
3417
+ | Workspace-dependent | `workspaceFiles` on the first `t.send`, never developer-machine paths |
3418
+
3419
+ Tag the fast, reliably passing core `smoke` and run `--tag smoke` in
3420
+ the inner loop. Leave slow or drift-prone cases untagged for explicit
3421
+ runs.
3422
+
3423
+ ### Materialize API-backed fixtures
3424
+
3425
+ An input that only points at external data (a pull request URL, a
3426
+ snapshot id, a pair of commit SHAs) is not self-contained. Fetch it once
3427
+ and commit the rendered fixture before you expand the suite:
3428
+
3429
+ 1. Save the diff, metadata, and labels under `fixtures/` at pinned
3430
+ revisions.
3431
+ 2. Seed those files with `workspaceFiles`, or read them from the
3432
+ fixture directory.
3433
+ 3. Assert decisions and output shape against the saved evidence.
3434
+ 4. Keep a small `smoke` subset for any remaining live checks.
3435
+
3436
+ `maxConcurrency` limits parallel datapoints, not the model or API
3437
+ fan-out inside one datapoint. Materialized fixtures keep a large suite
3438
+ from exhausting provider and GitHub rate limits.
3439
+
3440
+ ### Load a dataset
3441
+
3442
+ Read committed fixtures with `loadJson`, `loadJsonl`, and `loadYaml`
3443
+ from `@cursor/july/evals/loaders`. Relative paths resolve against the
3444
+ project root the runner discovered, not the cwd the CLI ran from. Eval
3445
+ files are ES modules, so top-level `await` can load a dataset and fan
3446
+ one file out over it:
3447
+
3448
+ ```ts
3449
+ // evals/sql.eval.ts: sql/0000, sql/0001, ...
3450
+ import { defineEval, equals } from "@cursor/july/evals";
3451
+ import { loadYaml } from "@cursor/july/evals/loaders";
3452
+
3453
+ const rows = await loadYaml<{ task: string; prompt: string; sql: string }[]>(
3454
+ "evals/data/cases.yaml"
3455
+ );
3456
+
3457
+ export default rows.map((row) =>
3458
+ defineEval({
3459
+ description: row.task,
3460
+ async test(t) {
3461
+ await t.send(row.prompt);
3462
+ t.succeeded();
3463
+ t.check(t.reply, equals(row.sql));
3464
+ },
3465
+ })
3466
+ );
3467
+ ```
3468
+
3469
+ ## Configure eval runs
3470
+
3471
+ `evals/evals.config.ts` (or `.js`) holds project-wide defaults. It must
3472
+ set `maxConcurrency`. Each case issues real model requests, so
3473
+ concurrency is hard-capped at 200; the templates use 10.
3474
+ `eval --list` works without the file. Running a case does not.
3475
+
3476
+ ```ts
3477
+ import { defineEvalConfig } from "@cursor/july/evals";
3478
+
3479
+ export default defineEvalConfig({
3480
+ maxConcurrency: 20,
3481
+ timeoutMs: 180_000,
3482
+ judge: { model: "gpt-5.4-mini" },
3483
+ });
3484
+ ```
3485
+
3486
+ | Option | Default | Meaning |
3487
+ | --- | --- | --- |
3488
+ | `maxConcurrency` | required | Datapoints in flight at once, 1-200 |
3489
+ | `timeoutMs` | `180_000` | Per-case timeout. Precedence: case or file `timeoutMs`, then `--timeout-ms`, then this value |
3490
+ | `judge` | unset | Default judge model for `t.judge.*` |
3491
+ | `reporters` | unset | Destinations that observe every case; `--skip-report` suppresses them |
3492
+ | `maxPlaygroundRuns` | `20` | Batches kept in playground history (cap 500). Counts `--prod` and `--url` batches, not the default local run |
3502
3493
 
3503
- ## Run evals from the CLI
3494
+ Reporters ship results somewhere; the runner still does the grading.
3495
+ `JUnit({ filePath, suiteName? })` writes JUnit XML and
3496
+ `Artifacts({ dir })` writes per-case files, both from
3497
+ `@cursor/july/evals/reporters`. A custom reporter is an object with any
3498
+ of `onRunStart`, `onEvalComplete`, and `onRunComplete`. A reporter that
3499
+ throws is logged and never fails the run. CI usually attaches the
3500
+ built-in two with `--junit` and `--artifacts` instead of `reporters`,
3501
+ so output paths stay with the pipeline, not the eval author.
3504
3502
 
3505
- The `eval` command discovers, filters, and runs cases.
3503
+ Playground batches survive restarts when the project has
3504
+ [storage](/docs/storage.md#eval-table). Otherwise they live in process
3505
+ memory until `serve` exits.
3506
3506
 
3507
- Run the CLI under Node 22.13 or newer. Do not use Bun. Its HTTP/2 client
3508
- breaks tool-result streams and causes eval turns to fail.
3507
+ ## Run evals from the CLI
3509
3508
 
3510
3509
  ```bash
3511
- agent-sdk eval --dir . --list # discover only
3512
- agent-sdk eval --dir . # run all
3513
- agent-sdk eval --dir . builds/checkout # one datapoint
3514
- agent-sdk eval --dir . builds search # several ids or prefixes
3515
- agent-sdk eval --dir . --tag smoke --tag pull-request # any matching tag
3516
- agent-sdk eval --dir . --json --no-stream # machine-readable results
3517
- agent-sdk eval --dir . --verbose # logs + reply snippets
3510
+ agent-sdk eval --dir . --list # discover only
3511
+ agent-sdk eval --dir . # run all
3512
+ agent-sdk eval --dir . builds/checkout # one datapoint
3513
+ agent-sdk eval --dir . builds search # several ids or prefixes
3514
+ agent-sdk eval --dir . --tag smoke --tag pull-request # any matching tag
3515
+ agent-sdk eval --dir . --verbose # t.log lines + reply snippets
3516
+ agent-sdk eval --dir . --timeout-ms 300000 # override the per-case timeout
3518
3517
  ```
3519
3518
 
3520
3519
  Id filters use OR semantics. Each filter selects an exact id and its
3521
- descendants. For example, `builds` selects `builds`,
3522
- `builds/checkout`, and every other case below that path. Repeated tags
3523
- also use OR semantics. When you provide both ids and tags, a case must
3524
- match both groups.
3520
+ descendants: `builds` selects `builds`, `builds/checkout`, and every
3521
+ other case below that path. Repeated tags also use OR. With both ids
3522
+ and tags, a case must match both groups.
3525
3523
 
3526
- `eval` boots an ephemeral server on port 0 with a temp state root
3527
- outside the project, so cases don't inherit ambient monorepo rules and
3528
- don't write into the project state directory. Point `--url` at a running server to eval
3529
- a live agent instead:
3524
+ By default `eval` boots a throwaway server with its own state root, so
3525
+ cases don't inherit your checkout's `AGENTS.md` and session state stays
3526
+ out of the project. Artifacts still land in the project state
3527
+ directory; see [Where results land](#where-results-land). `--slug`
3528
+ picks the target in a multi-agent directory.
3529
+
3530
+ `--url` runs the batch on a running server instead, the same way
3531
+ `--prod` does: that server discovers its own `evals/`, results land in
3532
+ its playground history, and the local-only flags (`--junit`,
3533
+ `--artifacts`, `--max-concurrency`, `--skip-report`) do not apply. See
3534
+ [Run evals in the playground or on a deployment](#run-evals-in-the-playground-or-on-a-deployment).
3530
3535
 
3531
3536
  ```bash
3532
- agent-sdk eval --dir . \
3533
- --url http://127.0.0.1:3000/weather-agent \
3537
+ agent-sdk eval --url http://127.0.0.1:3000/weather-agent \
3534
3538
  --bearer-token "$AGENT_TOKEN"
3535
3539
  ```
3536
3540
 
3537
- The eval definitions still come from `--dir`; `--url` only changes the
3538
- agent that receives the turns. For a locally mounted multi-agent
3539
- directory, `--slug weather-agent` chooses the target. Use
3540
- `--state-root` to keep ephemeral session state at a chosen path,
3541
- `--timeout-ms` to override the project timeout, and `--no-stream` to
3542
- keep live progress off stderr. A TTY streams turn progress by default.
3543
- `--verbose` still writes `t.log` lines to stderr and adds reply snippets
3544
- to text results.
3541
+ See [CLI: eval](/docs/reference/cli.md#eval) for every flag.
3542
+
3543
+ ### Credentials
3544
+
3545
+ Model turns need a
3546
+ [Cursor credential](/docs/reference/cli.md#environment-variables):
3547
+ `CURSOR_API_KEY`, `CURSOR_SERVICE_ACCOUNT_KEY`, or a saved
3548
+ `agent-sdk login`. The judge uses the same one. `eval --list` needs
3549
+ none.
3545
3550
 
3546
- Model turns need a Cursor credential: `CURSOR_API_KEY`, then
3547
- `CURSOR_API_KEY_FILE` (hosted default `/run/cursor/secrets/CURSOR_API_KEY`), then
3548
- `CURSOR_SERVICE_ACCOUNT_KEY`, then `agent-sdk login`.
3551
+ ### Where results land
3549
3552
 
3550
- See [CLI: eval](/docs/reference/cli.md#eval) for flags and exit codes.
3553
+ Every local run writes artifacts to a timestamped directory under
3554
+ `evals/` in the project state directory, whatever `--state-root` says.
3555
+ `--artifacts <dir>` chooses the path and `--no-artifacts` skips them.
3556
+ The directory holds `summary.json`,
3557
+ `results.jsonl`, and `evals/<case-id>.json` with every assertion, the
3558
+ inputs, tool calls with arguments and output, the final text, and
3559
+ `t.log` lines. Start there when a case fails. `--out <file>` also
3560
+ writes the full results JSON to a path of your choice.
3551
3561
 
3552
- ### JSON results
3562
+ The artifact does not include the session's event stream. Pass
3563
+ `--state-root <path>` to keep the ephemeral server's
3564
+ [session data](/docs/reference/sessions.md#where-does-the-agent-sdk-store-session-data)
3565
+ on disk when you need the raw events.
3553
3566
 
3554
- Use `--json --no-stream` in scripts and CI. The top-level result carries
3555
- the totals and one result per case:
3567
+ ## Run evals in CI
3568
+
3569
+ Run the suite non-interactively, write JUnit for the CI annotations,
3570
+ and fail the job on a red gate:
3571
+
3572
+ ```bash
3573
+ # CURSOR_API_KEY comes from the CI secret store
3574
+ agent-sdk eval --dir . --json --no-stream \
3575
+ --junit reports/evals.xml \
3576
+ --artifacts reports/evals \
3577
+ > reports/evals.json
3578
+ ```
3579
+
3580
+ The exit code follows the
3581
+ [verdict table](#gates-soft-scores-and-verdicts); `2` means nothing
3582
+ matched the selection. `--max-concurrency` overrides the project
3583
+ setting, for example to run lower on a shared runner.
3584
+
3585
+ The JSON on stdout carries the totals and one result per case:
3556
3586
 
3557
3587
  ```json
3558
3588
  {
3559
3589
  "ok": true,
3560
3590
  "passed": 1,
3561
3591
  "failed": 0,
3592
+ "scored": 0,
3593
+ "skipped": 0,
3594
+ "strict": false,
3595
+ "artifactsDir": "/work/my-agent/reports/evals",
3562
3596
  "results": [
3563
3597
  {
3564
3598
  "id": "readiness",
3599
+ "verdict": "passed",
3565
3600
  "ok": true,
3566
- "assertions": [{ "name": "succeeded", "passed": true }],
3601
+ "assertions": [
3602
+ { "name": "succeeded", "passed": true },
3603
+ { "name": "calledTool(inspect_pr)", "passed": true }
3604
+ ],
3567
3605
  "sessionId": "ses_123",
3568
- "inputs": ["Is checkout pull request 42 ready to approve?"],
3606
+ "inputs": ["Is https://github.com/acme/checkout/pull/42 ready to approve?"],
3569
3607
  "toolCalls": [{ "toolName": "inspect_pr", "isError": false }],
3608
+ "metrics": {},
3570
3609
  "logs": [],
3571
3610
  "durationMs": 12340
3572
3611
  }
@@ -3574,126 +3613,64 @@ the totals and one result per case:
3574
3613
  }
3575
3614
  ```
3576
3615
 
3577
- Each case result can also include `description`, `finalText`, `tools`,
3578
- `error`, and tool arguments or output. This shape lets CI report the
3579
- failed assertion without parsing terminal text.
3616
+ Each result can also include `description`, `finalText`, `tools`,
3617
+ `error`, `skipReason`, `metadata`, `tags`, and tool `args` and `output`.
3618
+ A soft miss shows as `"severity": "soft"` with `score` and `threshold`
3619
+ on the assertion. This shape lets CI report the failed assertion
3620
+ without parsing terminal text.
3621
+
3622
+ Keep CI green without weakening gates:
3580
3623
 
3581
- ## Run evals in the playground
3624
+ - Run `--tag smoke` on every push and the full suite on a schedule.
3625
+ - For probabilistic behavior, use `iterations` and a soft bar instead
3626
+ of one hard gate.
3582
3627
 
3583
- Start the server, open the playground, and choose **Evals**. You can run
3584
- every case or one case, watch progress, and open the resulting session
3585
- trace. The Evals tab works on a normal `serve`.
3628
+ ## Run evals in the playground or on a deployment
3629
+
3630
+ Start the server, open the playground, and choose **Evals**. Run every
3631
+ case or one case, watch progress, and open the resulting session trace.
3632
+ Playground runs target the live server instead of an ephemeral one, so
3633
+ their sessions appear in the session list. One batch runs at a time.
3586
3634
 
3587
3635
  ```bash
3588
3636
  agent-sdk serve --dir .
3589
3637
  ```
3590
3638
 
3591
- Playground runs target the live server instead of an ephemeral one.
3592
- Their sessions appear in the session list. One eval batch can run at a
3593
- time. Persistence follows the rule under
3594
- [Configure eval runs](#configure-eval-runs). See
3595
- [Playground eval routes](/docs/reference/http-api.md#playground-eval-routes).
3596
- The start request returns `202` while cases run in the background.
3597
- Poll until the snapshot status becomes `completed`, `failed`, or `cancelled`.
3598
- Configuration errors appear on a failed snapshot.
3599
-
3600
- On `--prod` / `--url`, the CLI prints the Eval ID as soon as the batch is
3601
- accepted (and a Playground deep link with `?view=evals&evalRunId=…`):
3639
+ `--prod` (or `--url`) starts the same server-side batch on the team's
3640
+ hosted deployment (or the server you name), so results land in that
3641
+ server's playground history:
3602
3642
 
3603
3643
  ```bash
3604
3644
  agent-sdk eval --prod --slug vulnerability-scanner --tag deepsec
3605
- # Eval ID: evalrun_…
3606
- # Cancel: agent-sdk eval cancel evalrun_… --prod --slug vulnerability-scanner
3607
- # Playground: https://…/playground?view=evals&evalRunId=evalrun_…
3608
-
3609
- agent-sdk eval cancel evalrun_… --prod --slug vulnerability-scanner
3610
- agent-sdk eval status evalrun_… --prod --slug vulnerability-scanner
3645
+ # Eval ID: <evalId>
3646
+ # Playground: <deployment>/playground?view=evals&evalRunId=<evalId>
3647
+ agent-sdk eval status <evalId> --prod --slug vulnerability-scanner
3648
+ agent-sdk eval cancel <evalId> --prod --slug vulnerability-scanner
3611
3649
  ```
3612
3650
 
3613
- ## What good cases assert
3614
-
3615
- Gate decisions and shape, not prose. Model wording varies run to run.
3616
- Tool choice, tool avoidance, and output structure are the stable
3617
- contract.
3618
-
3619
- 1. `t.succeeded()`: always, first.
3620
- 2. The tool decision: `calledTool` for the intended path,
3621
- `notCalledTool` for the likely wrong alternative. The pair is
3622
- stronger than either alone.
3623
- 3. Output shape: a regex for the contract (`/ready|blocked/i`, a JSON
3624
- marker, a findings-block fence), never exact sentences.
3625
- 4. For structured output, parse `t.reply` and check fields with
3626
- `satisfies` instead of substring-matching JSON.
3627
-
3628
- The common failure modes: asserting exact phrasing, packing more than
3629
- about five gates into one case (split it), and cases that depend on live
3630
- external state that drifts (pin the input; see fixtures).
3631
-
3632
- ## Pick fixtures by agent type
3633
-
3634
- The right fixture depends on the surface under test.
3635
-
3636
- | Agent surface | Fixture |
3637
- | --- | --- |
3638
- | Chat / domain assistant | A canonical prompt string, chosen once and frozen |
3639
- | Tool-heavy | Run `agent-sdk call <tool>` first to pin what the tool returns, then freeze the prompt that triggers it |
3640
- | GitHub webhook | `agent-sdk github replay <pr> --events '*' --dry-run --out fixtures/github` snapshots real payloads for offline replay ([GitHub guide](/docs/guides/github.md)) |
3641
- | PR reviewer with host preparation | Diff, metadata, and gold labels pinned to commit SHAs; keep any live PR matrix small |
3642
- | Workspace-dependent | `workspaceFiles` in `t.send` options, never developer-machine paths |
3643
-
3644
- Tag the fast, reliably passing core `smoke` and run `--tag smoke` in the
3645
- inner loop. Leave slow or flaky-prone cases untagged for explicit runs.
3646
-
3647
- ### Materialize API-backed fixtures
3648
-
3649
- An input that only points at external data, such as a pull request URL,
3650
- snapshot id, or pair of commit SHAs, is not self-contained. Fetch it
3651
- once and commit the rendered fixture before you expand the suite.
3652
-
3653
- 1. Save the diff, metadata, and labels under `fixtures/` at pinned
3654
- revisions.
3655
- 2. Seed those files with `workspaceFiles`, or read them from the fixture
3656
- directory.
3657
- 3. Assert decisions and output shape against the saved evidence.
3658
- 4. Keep a small `smoke` subset for any remaining live pipeline checks.
3659
-
3660
- Read committed fixtures with `@cursor/july/evals/loaders`: `loadJson`,
3661
- `loadJsonl`, and `loadYaml` resolve relative paths against the project
3662
- root the runner discovered, not the cwd the CLI was invoked from
3663
- (`resolveFixturePath` and `evalFixtureRoot` expose the same
3664
- resolution for other file formats).
3665
-
3666
- `maxConcurrency` limits parallel datapoints. It does not limit model or
3667
- API fan-out inside one datapoint. Materialized fixtures prevent a large
3668
- suite from exhausting provider and GitHub rate limits. The
3669
- [evals skill](/docs/skills/evals.md) has the full fixture workflow.
3651
+ The CLI prints the Eval ID as soon as the batch is accepted. Pass
3652
+ `--no-wait` to return right away and poll with `eval status` later; it
3653
+ exits `3` while the batch is still running. Hosted history follows
3654
+ `maxPlaygroundRuns` and the persistence rule under
3655
+ [Configure eval runs](#configure-eval-runs). The HTTP surface is under
3656
+ [Playground eval routes](/docs/reference/http-api.md#playground-eval-routes).
3670
3657
 
3671
3658
  ## Keep improvements with regression evals
3672
3659
 
3673
- Every [hillclimb](/docs/hillclimbing.md) round that keeps a change must land
3674
- an eval that would have failed before the change. If you can't express
3675
- the improvement as a gate (a `calledTool` shift, a bounded
3676
- `action.result` count, an output-shape regex), the improvement is
3677
- unverified, and it'll regress silently.
3660
+ Every [hillclimb](/docs/hillclimbing.md) round that keeps a change must
3661
+ land an eval that would have failed before the change. If you can't
3662
+ express the improvement as a gate (a `calledTool` shift, a bounded
3663
+ `maxToolCalls`, an output-shape check), the improvement is unverified,
3664
+ and it'll regress silently.
3678
3665
 
3679
- The rule cuts the other way too: never weaken an existing gate to make a
3680
- round pass. That's the freeze line moving, and it turns your regression
3681
- suite into a list of checks that no longer protect anything.
3682
-
3683
- ## Compare variants on live traffic
3684
-
3685
- Use `defineAB` to compare variant metrics on live sessions. It is not a
3686
- test runner and has no `agent-sdk ab` command. Keep `defineEval` as the
3687
- regression ratchet. Eval sessions do not enroll or change live metrics.
3688
- See [Live A/B metrics](/docs/ab.md) for assignment, behavior, collection,
3689
- and inspection.
3666
+ The rule cuts the other way too: never weaken an existing gate to make
3667
+ a round pass. That's the freeze line moving, and it turns your
3668
+ regression suite into a list of checks that no longer protect anything.
3690
3669
 
3691
3670
  ## What's next
3692
3671
 
3693
3672
  Continue with these pages:
3694
3673
 
3695
- - [Live A/B metrics](/docs/ab.md): sticky variants and cumulative metrics
3696
- on live sessions
3697
3674
  - [Hillclimbing](/docs/hillclimbing.md): the loop evals make trustworthy
3698
3675
  - [Building agents with agents](/docs/building-with-agents.md): have a
3699
3676
  coding agent write the first suite
@@ -3822,6 +3799,95 @@ Continue with these pages:
3822
3799
 
3823
3800
  ---
3824
3801
 
3802
+ Source: /docs/guides/bitbucket.md
3803
+
3804
+ # Bitbucket agents
3805
+
3806
+ Use `bitbucketChannel()` for Bitbucket Cloud and Bitbucket Data Center. The
3807
+ channel detects the source from each signed payload and normalizes Data Center
3808
+ events to the Bitbucket Cloud event vocabulary.
3809
+
3810
+ ## Define the channel
3811
+
3812
+ Author `agent/channels/bitbucket.ts`:
3813
+
3814
+ ```ts
3815
+ import { bitbucketChannel } from "@cursor/july/channels/bitbucket";
3816
+
3817
+ export default bitbucketChannel({});
3818
+ ```
3819
+
3820
+ Without `onPullRequest`, new pull requests start turns. Data Center updates
3821
+ also start turns when the source branch has new commits. Comments and pushes
3822
+ are opt-in through `onPullRequestComment` and `onPush`. Use `onEvent` for
3823
+ other event types. Set `webhookEvents` when managed event delivery must
3824
+ subscribe to an event that only `onEvent` handles.
3825
+
3826
+ Filter comment hooks by author before starting a turn. This prevents comments
3827
+ posted by the agent from triggering another turn.
3828
+
3829
+ `ctx.bitbucket.api` can read pull requests, post comments, and create build
3830
+ statuses. See the [Channels reference](/docs/reference/channels.md) for hook
3831
+ return values and session behavior.
3832
+
3833
+ ## Connect Bitbucket
3834
+
3835
+ Set a repository hook secret and API token:
3836
+
3837
+ ```bash
3838
+ BITBUCKET_WEBHOOK_SECRET=...
3839
+ BITBUCKET_TOKEN=...
3840
+ ```
3841
+
3842
+ Add a repository webhook for
3843
+ `https://<your-host>/<slug>/v1/channels/bitbucket`. Use the same secret on
3844
+ both sides. The channel verifies the `X-Hub-Signature` HMAC before it parses
3845
+ the payload.
3846
+
3847
+ Bitbucket Cloud uses its 2.0 API by default. Data Center also needs its REST
3848
+ API base:
3849
+
3850
+ ```ts
3851
+ export default bitbucketChannel({
3852
+ apiBaseUrl: "https://bitbucket.example.com/rest/api/1.0",
3853
+ repos: ["PLATFORM/api"],
3854
+ });
3855
+ ```
3856
+
3857
+ You can set `BITBUCKET_API_BASE_URL` instead.
3858
+
3859
+ ## Test locally
3860
+
3861
+ Start the agent, then replay a pull request you can read:
3862
+
3863
+ ```bash
3864
+ agent-sdk dev
3865
+ agent-sdk bitbucket replay \
3866
+ https://bitbucket.example.com/projects/PLATFORM/repos/api/pull-requests/42
3867
+ ```
3868
+
3869
+ Replay supports Bitbucket Cloud and Data Center URLs. `BITBUCKET_TOKEN` needs
3870
+ pull request read access. The command reads the pull request, creates a payload
3871
+ in the matching dialect, and sends it through the same channel route. Use
3872
+ `--events '*'` to replay the supported events declared by the channel. Use
3873
+ `--dry-run --out fixtures/bitbucket` to save fixtures.
3874
+
3875
+ ```bash
3876
+ agent-sdk bitbucket events --dir .
3877
+ agent-sdk bitbucket forward --dir .
3878
+ ```
3879
+
3880
+ `forward` prints the repository-hook and HTTPS tunnel setup for live
3881
+ deliveries.
3882
+
3883
+ ## Related
3884
+
3885
+ - [Channels reference](/docs/reference/channels.md)
3886
+ - [Webhooks and custom channels](/docs/guides/webhooks.md)
3887
+ - [Evals](/docs/evals.md)
3888
+
3889
+ ---
3890
+
3825
3891
  Source: /docs/guides/cloud-agents.md
3826
3892
 
3827
3893
  # Cursor cloud agents
@@ -4173,6 +4239,37 @@ snapshot in the wake.
4173
4239
  This is the preferred production path: no public URL, no repo admin
4174
4240
  webhook, and no inbound network for GitHub deliveries.
4175
4241
 
4242
+ ## Connect GitHub Enterprise Server
4243
+
4244
+ GitHub Enterprise Server uses the same `githubChannel()` hooks and normalized
4245
+ events. Connect it through the direct webhook route. Set the REST API base and
4246
+ credentials for your server:
4247
+
4248
+ ```ts
4249
+ export default githubChannel({
4250
+ api: {
4251
+ apiBaseUrl: "https://github.example.com/api/v3",
4252
+ },
4253
+ credentials: {
4254
+ token: () => process.env.GITHUB_TOKEN,
4255
+ webhookSecret: () => process.env.GITHUB_WEBHOOK_SECRET,
4256
+ },
4257
+ onPullRequest: (ctx, pr) =>
4258
+ pr.action === "opened" ? { auth: defaultGitHubAuth(ctx) } : null,
4259
+ });
4260
+ ```
4261
+
4262
+ Add a repository webhook for
4263
+ `https://<your-host>/<slug>/v1/channels/github`. Use the same webhook secret
4264
+ on the server and in `GITHUB_WEBHOOK_SECRET`. The channel verifies
4265
+ `X-Hub-Signature-256` before it parses the payload.
4266
+
4267
+ Cursor account event pull, `github replay`, and `github forward` target
4268
+ GitHub.com. Test Enterprise Server integrations by posting saved webhook
4269
+ fixtures to a local `--dev` server. Leave `GITHUB_WEBHOOK_SECRET` unset for
4270
+ this local test so the channel admits unsigned loopback deliveries. Use a real
4271
+ delivery from your server so the fixture matches its version.
4272
+
4176
4273
  ## Define the channel
4177
4274
 
4178
4275
  Author `agent/channels/github.ts` with `githubChannel()` from
@@ -4380,6 +4477,103 @@ key. Handlers you author replace the matching defaults (same as
4380
4477
 
4381
4478
  ---
4382
4479
 
4480
+ Source: /docs/guides/gitlab.md
4481
+
4482
+ # GitLab agents
4483
+
4484
+ Use `gitlabChannel()` for GitLab.com and self-managed GitLab. The channel
4485
+ verifies project hooks, normalizes their payloads, and gives each hook a
4486
+ project-bound `ctx.gitlab` API client.
4487
+
4488
+ ## Define the channel
4489
+
4490
+ Author `agent/channels/gitlab.ts`:
4491
+
4492
+ ```ts
4493
+ import { gitlabChannel } from "@cursor/july/channels/gitlab";
4494
+
4495
+ export default gitlabChannel({});
4496
+ ```
4497
+
4498
+ Without `onMergeRequest`, new and reopened merge requests start turns.
4499
+ Updates start turns only when they include new commits. Notes, pushes, and
4500
+ pipelines are opt-in through `onNote`, `onPush`, and `onPipeline`.
4501
+ Use `onEvent` for other GitLab event types. Set `webhookEvents` when managed
4502
+ event delivery must subscribe to an event that only `onEvent` handles.
4503
+
4504
+ Filter note hooks by author before starting a turn. This prevents comments
4505
+ posted by the agent from triggering another turn.
4506
+
4507
+ `ctx.gitlab` can call the project REST API and create commit statuses. See the
4508
+ [Channels reference](/docs/reference/channels.md) for hook return values and
4509
+ session behavior.
4510
+
4511
+ ## Connect GitLab
4512
+
4513
+ On Cursor-managed hosting, use the signed-in Cursor account:
4514
+
4515
+ ```ts
4516
+ export default gitlabChannel({
4517
+ cursorAccount: {
4518
+ projects: ["acme/platform"],
4519
+ },
4520
+ });
4521
+ ```
4522
+
4523
+ For direct webhooks, set a project hook secret and API token:
4524
+
4525
+ ```bash
4526
+ GITLAB_WEBHOOK_SECRET=...
4527
+ GITLAB_TOKEN=...
4528
+ ```
4529
+
4530
+ Add a project webhook for
4531
+ `https://<your-host>/<slug>/v1/channels/gitlab`. Use the same value for the
4532
+ GitLab secret token and `GITLAB_WEBHOOK_SECRET`. The channel checks
4533
+ `X-Gitlab-Token` before it parses the payload.
4534
+
4535
+ Self-managed GitLab also needs its REST API base:
4536
+
4537
+ ```ts
4538
+ export default gitlabChannel({
4539
+ apiBaseUrl: "https://gitlab.example.com/api/v4",
4540
+ projects: ["acme/platform"],
4541
+ });
4542
+ ```
4543
+
4544
+ You can set `GITLAB_API_BASE_URL` instead.
4545
+
4546
+ ## Test locally
4547
+
4548
+ Start the agent, then replay a merge request you can read:
4549
+
4550
+ ```bash
4551
+ agent-sdk dev
4552
+ agent-sdk gitlab replay \
4553
+ https://gitlab.example.com/acme/platform/-/merge_requests/42
4554
+ ```
4555
+
4556
+ `GITLAB_TOKEN` needs API read access. Replay reads the merge request,
4557
+ creates GitLab-shaped payloads, and sends them through the same channel route.
4558
+ Use `--events '*'` to replay the supported events declared by the channel.
4559
+ Use `--dry-run --out fixtures/gitlab` to save fixtures.
4560
+
4561
+ ```bash
4562
+ agent-sdk gitlab events --dir .
4563
+ agent-sdk gitlab forward --dir .
4564
+ ```
4565
+
4566
+ GitLab has no local webhook relay. `forward` prints the project-hook and HTTPS
4567
+ tunnel setup for live deliveries.
4568
+
4569
+ ## Related
4570
+
4571
+ - [Channels reference](/docs/reference/channels.md)
4572
+ - [Webhooks and custom channels](/docs/guides/webhooks.md)
4573
+ - [Evals](/docs/evals.md)
4574
+
4575
+ ---
4576
+
4383
4577
  Source: /docs/guides/grokbot-agents.md
4384
4578
 
4385
4579
  # Cursor Grok Bot agents
@@ -5395,7 +5589,8 @@ A custom channel gives the agent its own HTTP surface. You get routes
5395
5589
  with validated payloads, sessions keyed to something in your domain (a
5396
5590
  thread, a ticket, a PR), and replies delivered back to the caller. The
5397
5591
  [Slack](/docs/guides/slack.md) and [GitHub](/docs/guides/github.md) packs build on this
5398
- mechanism. This page is the mechanism itself.
5592
+ mechanism. The [GitLab](/docs/guides/gitlab.md) and [Bitbucket](/docs/guides/bitbucket.md) packs
5593
+ use it too. This page is the mechanism itself.
5399
5594
 
5400
5595
  ## What you already have
5401
5596
 
@@ -5849,7 +6044,8 @@ from any PR you can read. See the [GitHub guide](/docs/guides/github.md).
5849
6044
  Continue with these pages:
5850
6045
 
5851
6046
  - [Channels reference](/docs/reference/channels.md): the full authoring API
5852
- - [GitHub](/docs/guides/github.md) and [Slack](/docs/guides/slack.md): the packaged channels
6047
+ - [GitHub](/docs/guides/github.md), [GitLab](/docs/guides/gitlab.md),
6048
+ [Bitbucket](/docs/guides/bitbucket.md), and [Slack](/docs/guides/slack.md): the packaged channels
5853
6049
  - [Sessions and streaming](/docs/reference/sessions.md): events your
5854
6050
  channel can subscribe to
5855
6051
 
@@ -5926,9 +6122,9 @@ Pin the input first. A moving fixture is noise. For GitHub agents, use `agent-sd
5926
6122
 
5927
6123
  ## How do I lock a hillclimb improvement with an eval?
5928
6124
 
5929
- Every kept change needs an eval that would have failed before the change: a tool-choice gate, an `action.result` count bound, or an output-shape check. Run `agent-sdk eval --dir . --json` between rounds. Never weaken an existing gate to pass the round.
6125
+ Every kept change needs an eval that would have failed before the change: a tool-choice gate, a `maxToolCalls` bound, or an output-shape check. Run `agent-sdk eval --dir . --json` between rounds. Never weaken an existing gate to pass the round.
5930
6126
 
5931
- Details live in [Evals](/docs/evals.md). The evals skill will author the case with you.
6127
+ Details live in [Keep improvements with regression evals](/docs/evals.md#keep-improvements-with-regression-evals). The evals skill will author the case with you.
5932
6128
 
5933
6129
  ## What habits help hillclimbing stay reliable?
5934
6130
 
@@ -5983,14 +6179,16 @@ npx @cursor/july docs
5983
6179
  | Turning a Cursor Automation into a project | [Convert a Cursor Automation](/docs/guides/convert-automation.md) |
5984
6180
  | Wiring an agent to Slack | [Slack guide](/docs/guides/slack.md) |
5985
6181
  | Starting from a packaged template | [Demo](/docs/templates/demo.md), [Grok Bot agents](/docs/templates/grokbot-agents.md), [Code wiki](/docs/templates/code-wiki.md), [Living AGENTS.md](/docs/templates/agents-md.md), [Security reviewer](/docs/templates/security-reviewer.md), [Security help](/docs/templates/security-help.md), [Triage](/docs/templates/triage.md), or [Agentic Owners](/docs/templates/agentic-owners.md) |
5986
- | Wiring an agent to GitHub webhooks | [GitHub guide](/docs/guides/github.md) |
6182
+ | Wiring an agent to GitHub or GitHub Enterprise Server | [GitHub guide](/docs/guides/github.md) |
6183
+ | Wiring an agent to GitLab | [GitLab guide](/docs/guides/gitlab.md) |
6184
+ | Wiring an agent to Bitbucket | [Bitbucket guide](/docs/guides/bitbucket.md) |
5987
6185
  | Driving PRs from a cloud VM | [PR autofixer template](/docs/templates/pr-autofixer.md) |
5988
6186
  | Handing coding work to Cursor cloud agents | [Cursor cloud agents](/docs/guides/cloud-agents.md) |
5989
6187
  | Talking to your Grok Bot agents | [Cursor Grok Bot agents](/docs/guides/grokbot-agents.md) |
5990
6188
  | Letting an agent change its own source | [Self-improvement](/docs/guides/improve.md) |
5991
6189
  | Driving an agent from Linear (or another tracker) | [Webhooks guide: Linear example](/docs/guides/webhooks.md#example-linear-as-the-control-plane) |
5992
6190
  | Making an existing agent measurably better | [Evals](/docs/evals.md), then [Hillclimbing](/docs/hillclimbing.md) |
5993
- | Comparing variants on live traffic | [Live A/B metrics](/docs/ab.md) |
6191
+ | Gating agent behavior in CI | [Run evals in CI](/docs/evals.md#run-evals-in-ci) |
5994
6192
  | Deploying with Cursor or on your own infrastructure | [Deployment](/docs/deployment.md) |
5995
6193
  | Debugging something that misbehaves | [Fix common agent problems](/docs/troubleshooting.md) |
5996
6194
 
@@ -6031,10 +6229,8 @@ npx @cursor/july docs
6031
6229
 
6032
6230
  - [Building agents with agents](/docs/building-with-agents.md): use a coding
6033
6231
  agent to scaffold, run, and iterate on your agent.
6034
- - [Evals](/docs/evals.md): author `defineEval` cases, pick fixtures, and use
6035
- evals as regression checks.
6036
- - [Live A/B metrics](/docs/ab.md): assign sticky variants and compare
6037
- cumulative metrics on live sessions.
6232
+ - [Evals](/docs/evals.md): author `defineEval` cases, assert over the
6233
+ trajectory, and run them locally and in CI.
6038
6234
  - [Storage](/docs/storage.md): point durable storage at a backend you own
6039
6235
  with `defineStorage`.
6040
6236
  - [Hillclimbing](/docs/hillclimbing.md): measure and improve an agent
@@ -6046,6 +6242,10 @@ npx @cursor/july docs
6046
6242
  its own HTTP surface.
6047
6243
  - [GitHub](/docs/guides/github.md): trigger the agent from pull requests,
6048
6244
  CI, and comments.
6245
+ - [GitLab](/docs/guides/gitlab.md): trigger the agent from merge requests,
6246
+ notes, pipelines, and pushes.
6247
+ - [Bitbucket](/docs/guides/bitbucket.md): trigger the agent from pull requests,
6248
+ comments, and pushes.
6049
6249
  - [Slack](/docs/guides/slack.md): put the agent in Slack over Socket Mode.
6050
6250
  - [Human-in-the-loop approvals](/docs/guides/human-in-the-loop.md): park a
6051
6251
  tool call until a person signs off.
@@ -6862,7 +7062,8 @@ Source: /docs/reference/channels.md
6862
7062
  A channel is the surface an agent lives on. The built-in HTTP session
6863
7063
  channel is always mounted. Custom channels declare their own routes
6864
7064
  under `/v1/channels/<id>`. The Slack and GitHub packs are prebuilt
6865
- channels with platform transports. This page is the authoring reference;
7065
+ channels with platform transports. GitLab and Bitbucket packs cover their
7066
+ hosted and self-managed products. This page is the authoring reference;
6866
7067
  for the walkthrough, see the [Webhooks guide](/docs/guides/webhooks.md).
6867
7068
 
6868
7069
  ## Built-in HTTP channel
@@ -6988,7 +7189,7 @@ Handlers receive the Fetch `Request` and an args object:
6988
7189
  `workspaceFiles`, `workspaceDir`, `cloud` (attach cloud repos for this
6989
7190
  session), `auth` (defaults to the request principal), `state` (starting
6990
7191
  channel state for new sessions), `title` (session display title), and
6991
- `purpose` (`"eval"` skips sticky A/B enrollment).
7192
+ `purpose` (`"eval"` marks the session as regression traffic).
6992
7193
 
6993
7194
  ## Events
6994
7195
 
@@ -7069,7 +7270,19 @@ model turn), `{ task }` (host work), or `null`, and CLI tooling for
7069
7270
  replay and live forwarding. Author `agent/channels/github.ts` with
7070
7271
  `githubChannel()`. Opt-in `progress.commitStatus` and `progress.banner`
7071
7272
  converge a merge-box check and sticky PR comment from default stream
7072
- events. Guide: [GitHub](/docs/guides/github.md).
7273
+ events. Supports GitHub.com and GitHub Enterprise Server. Guide:
7274
+ [GitHub](/docs/guides/github.md).
7275
+
7276
+ **GitLab** (`@cursor/july/channels/gitlab`): verified project hooks for
7277
+ merge requests, notes, pipelines, pushes, and custom event types. Supports
7278
+ GitLab.com and self-managed GitLab. Author `agent/channels/gitlab.ts` with
7279
+ `gitlabChannel()`. Guide: [GitLab](/docs/guides/gitlab.md).
7280
+
7281
+ **Bitbucket** (`@cursor/july/channels/bitbucket`): verified repository hooks
7282
+ for pull requests, comments, pushes, and custom event types. Supports
7283
+ Bitbucket Cloud and Bitbucket Data Center through one normalized hook API.
7284
+ Author `agent/channels/bitbucket.ts` with `bitbucketChannel()`. Guide:
7285
+ [Bitbucket](/docs/guides/bitbucket.md).
7073
7286
 
7074
7287
  **Deployments** (`@cursor/july/channels/deployments`): pull deploy
7075
7288
  events. Declare `events` and handle each one in `onEvent`. Each event
@@ -7491,8 +7704,10 @@ agent-sdk eval status <evalId> --prod --slug pr-approver
7491
7704
  agent-sdk eval cancel <evalId> --prod --slug pr-approver
7492
7705
  ```
7493
7706
 
7494
- `eval` runs `evals/**/*.eval.{ts,js}` on an ephemeral server or against
7495
- `--url`. Select one or more exact case IDs, file ID prefixes, or tags.
7707
+ `eval` runs `evals/**/*.eval.{ts,js}` on an ephemeral server. With
7708
+ `--prod` or `--url`, the target server runs its own evals as a
7709
+ server-side batch. Select one or more exact case IDs, file ID prefixes,
7710
+ or tags.
7496
7711
  Omit selectors to run all cases. Repeated `--tag` flags use OR matching.
7497
7712
 
7498
7713
  An eval run requires `evals/evals.config.{ts,js}` with `maxConcurrency`
@@ -7503,13 +7718,13 @@ between 1 and 200. Timeout priority is the case's `timeoutMs`, the CLI's
7503
7718
  | --- | --- |
7504
7719
  | `--list` | Print discovered cases without running. `--list --json` prints them as an array. |
7505
7720
  | `--tag <tag>` | Run cases with this tag. Repeated flags use OR matching. |
7506
- | `--json` | Print `{ ok, passed, failed, results }`. |
7721
+ | `--json` | Print `{ ok, passed, failed, scored, skipped, strict, results }`; see [Run evals in CI](/docs/evals.md#run-evals-in-ci) for the result shape. |
7507
7722
  | `--verbose` | Stream `t.log` lines and reply snippets. |
7508
7723
  | `--no-stream` | Hide live progress on stderr. |
7509
7724
  | `--strict` | Exit `1` when a scored case misses a soft threshold. |
7510
7725
  | `--max-concurrency <n>` | Override `maxConcurrency` from `evals.config.ts`. |
7511
7726
  | `--junit <path>` | Write JUnit XML for CI annotations. |
7512
- | `--artifacts <dir>` | Write run artifacts here. The default is a timestamped directory under `<state-root>/evals/`. |
7727
+ | `--artifacts <dir>` | Write run artifacts here. The default is a timestamped directory under `evals/` in the project state directory (not affected by `--state-root`). |
7513
7728
  | `--no-artifacts` | Skip run artifacts. |
7514
7729
  | `--skip-report` | Ignore reporters from `evals.config.ts` and eval files. |
7515
7730
  | `--out <path>` | Also write the full results JSON to this path (also for `eval status <evalId>`). |
@@ -8473,7 +8688,6 @@ namespace:
8473
8688
  | `subagents/<id>/` | subagent `<ns>__<id>` |
8474
8689
  | `instructions.md` / `.ts` / dir | appended to the agent's system prompt |
8475
8690
  | `sandbox/workspace/**` | seeded into each local session workspace |
8476
- | `ab.ts` / `ab/<name>.ts` | A/B experiment `<ns>__<name>` |
8477
8691
  | `artifacts.ts` | artifact kinds `<ns>__<kind>` |
8478
8692
 
8479
8693
  The root agent still needs its own `instructions.md`. Extension
@@ -8490,7 +8704,6 @@ or is outside discovery:
8490
8704
  | `agent.ts` | One `defineAgent` runtime per agent |
8491
8705
  | `storage.ts` | One `host.kv` / `host.files` backend |
8492
8706
  | `otel.ts` | One OTLP exporter |
8493
- | `ab.config.ts` | One experiment-platform config; `ab.ts` / `ab/` still merge |
8494
8707
  | `playground/` | Custom chips are a Vite glob of the agent tree, not a discovery walk |
8495
8708
  | `extensions/` | Nested mounts are not loaded |
8496
8709
  | `sandbox.ts` | Custom sandbox backends stay on the agent |
@@ -8531,7 +8744,6 @@ it alive.
8531
8744
  | `schedules/<name>.ts` | `disableSchedule()` |
8532
8745
  | `subagents/<id>.ts` | `disableSubagent()` |
8533
8746
  | `instructions.ts` | `disableInstructions()` |
8534
- | `ab.ts` / `ab/<name>.ts` | `disableAB()` |
8535
8747
  | `artifacts.ts` | `disableArtifacts()` |
8536
8748
 
8537
8749
  `disable()` is the same brand as the slot helpers above and works in
@@ -8775,7 +8987,7 @@ exported from `@cursor/july`.
8775
8987
 
8776
8988
  | Member | What it is |
8777
8989
  | --- | --- |
8778
- | `ctx.session` | Read-only session info: `id`, `channelId`, `mode` (`chat` or `task`), `purpose` (`live` or `eval`), `auth`, plus `title`, `sdkAgentId`, and `abs` when set |
8990
+ | `ctx.session` | Read-only session info: `id`, `channelId`, `mode` (`chat` or `task`), `purpose` (`live` or `eval`), `auth`, plus `title` and `sdkAgentId` when set |
8779
8991
  | `ctx.agent` | `{ name }` of the agent the event belongs to |
8780
8992
  | `ctx.channel` | `{ id, continuationToken }`. The token is `null` when the session can't take follow-ups |
8781
8993
  | `ctx.host.kv` | Durable JSON, shared by every session of the agent; the [storage backend](/docs/storage.md#author-kv-ctx-host-kv) decides whether it survives a hosted replace. Prefix keys with `ctx.session.id` for per-session state |
@@ -8806,17 +9018,17 @@ Each event reaches a hook at most once. A restart doesn't replay the log
8806
9018
  into hooks, so a mirror needs no dedupe, and the event log rather than
8807
9019
  the hook's copy is the source of truth.
8808
9020
 
8809
- ## Hooks, channel events, evals, or A/B?
9021
+ ## Hooks, channel events, or evals?
8810
9022
 
8811
9023
  All of them consume the same stream, for different jobs:
8812
9024
 
8813
- | | Hooks | Channel `events` | Evals | A/B (`defineAB`) |
8814
- | --- | --- | --- | --- | --- |
8815
- | Scope | every session of the agent | sessions the channel owns | one test turn | every live session; enrollment at creation, metrics on each turn |
8816
- | Job | observe: audit, metrics, mirrors, derived state | deliver: replies back to the channel's surface | assert: gates over the trajectory | compare: sticky arms, then fold the stream into `onSample` metrics |
8817
- | Context | `ctx.host`, `ctx.artifacts`, session info | `channel.state`, `setContinuationToken`, `ctx.host`, session info | the `t` assertion helpers | per-session samples in `onSample` |
8818
- | Can affect the run | no | yes, it owns the surface | n/a | yes through arm instructions or `session.abs`; collection is observe-only |
8819
- | Authored at | `agent/hooks/*.ts` | channel config | `evals/**/*.eval.ts` | [`agent/ab.ts` or `agent/ab/*.ts`](/docs/ab.md) |
9025
+ | | Hooks | Channel `events` | Evals |
9026
+ | --- | --- | --- | --- |
9027
+ | Scope | every session of the agent | sessions the channel owns | one test turn |
9028
+ | Job | observe: audit, metrics, mirrors, derived state | deliver: replies back to the channel's surface | assert: gates over the trajectory |
9029
+ | Context | `ctx.host`, `ctx.artifacts`, session info | `channel.state`, `setContinuationToken`, `ctx.host`, session info | the `t` assertion helpers |
9030
+ | Can affect the run | no | yes, it owns the surface | n/a |
9031
+ | Authored at | `agent/hooks/*.ts` | channel config | `evals/**/*.eval.ts` |
8820
9032
 
8821
9033
  ## When not to use a hook
8822
9034
 
@@ -8828,7 +9040,6 @@ All of them consume the same stream, for different jobs:
8828
9040
  | Block, approve, or rewrite a tool call | [`needsApproval`](/docs/reference/tools.md#gate-a-tool-on-human-approval) on the tool |
8829
9041
  | Act on the final assistant text, or fail a bad turn | [`defineResult`](/docs/reference/result.md) |
8830
9042
  | Gate a change on behavior | [Evals](/docs/evals.md) |
8831
- | Compare two prompts on live traffic | [`defineAB`](/docs/ab.md) |
8832
9043
 
8833
9044
  ## Patterns
8834
9045
 
@@ -8954,8 +9165,6 @@ Continue with these pages:
8954
9165
  - [Deployment](/docs/deployment.md#observability): runtime logs and export
8955
9166
  paths
8956
9167
  - [Channels](/docs/reference/channels.md#events): the delivery-side counterpart
8957
- - [Live A/B metrics](/docs/ab.md): sticky variants over the same event
8958
- stream
8959
9168
 
8960
9169
  ---
8961
9170
 
@@ -9110,12 +9319,11 @@ These read-only routes describe the running agent.
9110
9319
 
9111
9320
  | Route | What it does |
9112
9321
  | ---------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
9113
- | `GET /v1/info` | The discovered surface: model, tools, skills, MCP connections, subagents, channels and routes (with schemas), schedules, hooks, A/B experiments, diagnostics |
9322
+ | `GET /v1/info` | The discovered surface: model, tools, skills, MCP connections, subagents, channels and routes (with schemas), schedules, hooks, diagnostics |
9114
9323
  | `GET /v1/tools` | The live tool catalog: authored server tools plus advertised MCP passthroughs under model-facing names, as light `{ name, title?, source? }` entries. `session` / `continuationToken` query parameters bind the listing to a session identity (advertised inventories can be tenant-scoped); a connection whose listing fails is skipped and reported in `connectionErrors` |
9115
9324
  | `GET /v1/tools/:name` | One catalog tool's full description: description, execution, `needsApproval`, `effect`, input and output schemas, source connection. Same session binding as the listing; unknown names get `404` with the available names |
9116
9325
  | `GET /v1/health` | Per-agent liveness, no auth |
9117
9326
  | `GET /v1/logs?after=N` | Recent server log lines, with a polling cursor |
9118
- | `GET /v1/abs` | [Live A/B metrics](/docs/ab.md): per-session assignments and aggregate arm totals |
9119
9327
 
9120
9328
  ## Artifacts
9121
9329
 
@@ -9174,7 +9382,7 @@ the final `completed` or `failed` status. Batch errors appear on the
9174
9382
  snapshot returned by the poll. Entries within `filterIds` and `tags`
9175
9383
  use OR semantics. When both fields are present, a case must match one
9176
9384
  entry from each field. Listed runs persist across restarts when storage is configured; see
9177
- [Storage](/docs/storage.md#eval-and-a-b-tables). Otherwise they are
9385
+ [Storage](/docs/storage.md#eval-table). Otherwise they are
9178
9386
  process-memory only.
9179
9387
 
9180
9388
  ## Dev-mode routes
@@ -9336,8 +9544,6 @@ Use the playground to chat, try channel routes, and inspect sessions.
9336
9544
  an [extension](/docs/reference/extensions.md) cannot contribute them.
9337
9545
  - **Raw events pane**: flip it on to inspect the event stream.
9338
9546
  - **Logs tab**: recent server log lines, polled from `GET /v1/logs`.
9339
- - **A/Bs tab**: per-session and aggregate
9340
- [live A/B metrics](/docs/ab.md) from `GET /v1/abs`.
9341
9547
 
9342
9548
  In multi-agent mode each agent has its own playground at
9343
9549
  `/<slug>/playground`, and `/` is an index of them all.
@@ -9360,8 +9566,6 @@ Continue with these pages:
9360
9566
  - [Sessions and streaming](/docs/reference/sessions.md): the streams it renders
9361
9567
  - [Human-in-the-loop](/docs/guides/human-in-the-loop.md): the approval
9362
9568
  buttons in context
9363
- - [Live A/B metrics](/docs/ab.md): the assignments and results in the A/Bs
9364
- tab
9365
9569
 
9366
9570
  ---
9367
9571
 
@@ -9375,8 +9579,7 @@ how the Agent SDK loads it.
9375
9579
 
9376
9580
  ## Folder structure
9377
9581
 
9378
- For the capabilities below, identity usually comes from the path. A/B
9379
- experiments can override their file-derived name.
9582
+ For the capabilities below, identity comes from the path.
9380
9583
 
9381
9584
  | Path | Resolves to |
9382
9585
  | --- | --- |
@@ -9388,8 +9591,6 @@ experiments can override their file-derived name.
9388
9591
  | `agent/extensions/ci.ts` | extension mount `ci`; its contributions become `ci__<name>` |
9389
9592
  | `agent/extensions/notion.ts` | Cursor plugin mount `notion` (`cursorPlugin`); its skills, agents, and MCP servers become `notion__<name>` |
9390
9593
  | `agent/channels/drive.ts` | channel `drive`, routes under `/v1/channels/drive` |
9391
- | `agent/ab.ts` | A/B experiment `ab` unless `name` overrides it |
9392
- | `agent/ab/concise.ts` | A/B experiment `concise` unless `name` overrides it |
9393
9594
 
9394
9595
  The root agent takes its name from `package.json` `name`, falling back
9395
9596
  to the directory name. When serving multiple agents, the slug is the
@@ -9441,8 +9642,6 @@ Each path maps to a capability and a reference page.
9441
9642
  | `agent/channels/*.ts` | HTTP surfaces beyond the built-in session API; `slack.ts` and `github.ts` use the platform packs | [Channels](/docs/reference/channels.md) |
9442
9643
  | `agent/hooks/*.ts` | Observe-only event subscribers, never fatal | [Hooks](/docs/reference/hooks.md) |
9443
9644
  | `agent/otel.ts` | `defineOtel` OTLP export (traces, metrics, optional logs) | [OpenTelemetry](/docs/guides/opentelemetry.md) |
9444
- | `agent/ab.ts`, `agent/ab/*.ts` | `defineAB` experiments with sticky variants and live metrics | [Live A/B metrics](/docs/ab.md) |
9445
- | `agent/ab.config.ts` | `defineABConfig` shared A/B settings | [Live A/B metrics](/docs/ab.md) |
9446
9645
  | `agent/storage.ts` | `defineStorage` backend for the durable `host.kv` / `host.files` APIs | [Storage](/docs/storage.md) |
9447
9646
  | `agent/artifacts.ts` | `defineArtifacts` kinds, the `tag_artifact` opt-in, and retention | [Artifacts](/docs/reference/artifacts.md) |
9448
9647
  | `agent/result.ts` | `defineResult` host `commit` on the final assistant text | [Turn result](/docs/reference/result.md) |
@@ -9479,8 +9678,6 @@ Continue with these pages:
9479
9678
 
9480
9679
  - [Agent config](/docs/reference/agent-config.md): the runtime config at the root
9481
9680
  - [Tools](/docs/reference/tools.md): add typed actions under `agent/tools/`
9482
- - [Live A/B metrics](/docs/ab.md): compare variants from `agent/ab.ts` or
9483
- `agent/ab/`
9484
9681
  - [Concepts](/docs/concepts.md): why the filesystem is the interface
9485
9682
 
9486
9683
  ---
@@ -9924,7 +10121,7 @@ within one session. The `at` field is an ISO-8601 timestamp.
9924
10121
 
9925
10122
  | Phase | Events | What they tell you |
9926
10123
  | --- | --- | --- |
9927
- | Session | `session.started`, [`ab.assigned`](/docs/ab.md#assign-sticky-variants), `session.waiting`, `session.completed`, `session.failed` | Session creation, A/B enrollment, readiness, and task completion |
10124
+ | Session | `session.started`, `session.waiting`, `session.completed`, `session.failed` | Session creation, readiness, and task completion |
9928
10125
  | Agent | `agent.bound` | Cloud conversation URL |
9929
10126
  | Input | `message.received` | A user message was accepted |
9930
10127
  | Turn | `turn.queued`, `turn.started`, `turn.completed`, `turn.failed` | Queue position under a [`maxRunningTurns` cap](/docs/reference/agent-config.md#concurrency), then turn status, final result, and token usage |
@@ -10008,7 +10205,6 @@ view.
10008
10205
 
10009
10206
  - [HTTP API](/docs/reference/http-api.md)
10010
10207
  - [Hooks](/docs/reference/hooks.md)
10011
- - [Live A/B metrics](/docs/ab.md)
10012
10208
  - [How the Agent SDK works](/docs/concepts.md)
10013
10209
 
10014
10210
  ---
@@ -10661,61 +10857,6 @@ inputs again, and adds an eval for each improvement you keep.
10661
10857
 
10662
10858
  ---
10663
10859
 
10664
- Source: /docs/skills/ab.md
10665
-
10666
- # Agent SDK A/B metrics (`defineAB`)
10667
-
10668
- Live metrics plug-in. No `agent-sdk ab` CLI. No assertion API.
10669
- Reference: `docs/ab.md`.
10670
-
10671
- | | `defineEval` | `defineAB` |
10672
- | --- | --- | --- |
10673
- | Job | Gates on frozen fixtures | Metrics on live runs |
10674
- | Location | `evals/**/*.eval.ts` | `agent/ab.ts` or `agent/ab/<name>.ts` |
10675
- | How it runs | `agent-sdk eval` | Under `serve` / `run` |
10676
-
10677
- ```ts
10678
- import { defineAB, splitBySessionHash } from "@cursor/july/ab";
10679
-
10680
- export default defineAB({
10681
- name: "concise-instructions",
10682
- variants: {
10683
- control: { label: "Baseline" },
10684
- treatment: {
10685
- label: "Shorter",
10686
- instructions: "Keep replies to one short paragraph.",
10687
- },
10688
- },
10689
- split: splitBySessionHash({ holdout: 0.1 }),
10690
- derive: {
10691
- weatherCalls: (event) =>
10692
- event.type === "action.result" && event.data.toolName === "get_weather"
10693
- ? 1
10694
- : null,
10695
- },
10696
- onSample(sample) {
10697
- console.log(sample.variant, sample.metrics.toolCalls, sample.metrics.wallTimeMs);
10698
- },
10699
- });
10700
- ```
10701
-
10702
- ```ts
10703
- async execute(input, ctx) {
10704
- if (ctx.session.abs?.["concise-instructions"] === "treatment") {
10705
- // treatment-specific behavior
10706
- }
10707
- }
10708
- ```
10709
-
10710
- Enrollment is at session creation. Eval sessions skip it. Do not
10711
- use `splitIf` to filter evals. Split helpers and `onSample`
10712
- fields: `docs/ab.md`.
10713
-
10714
- Pick a name, arm labels, a split, and a real `onSample` sink. Do
10715
- not invent credentials.
10716
-
10717
- ---
10718
-
10719
10860
  Source: /docs/skills/create-agent.md
10720
10861
 
10721
10862
  # Create an Agent SDK agent
@@ -10922,15 +11063,182 @@ trace.
10922
11063
 
10923
11064
  ---
10924
11065
 
11066
+ Source: /docs/skills/deploy.md
11067
+
11068
+ # Deploy an Agent SDK agent
11069
+
11070
+ Use this skill only after a person asks for a deployment. It deploys a
11071
+ pushed Git ref to Cursor-managed hosting. Local files never upload.
11072
+
11073
+ Use the exact published `@cursor/july` release embedded in the installed
11074
+ skill. Do not use a floating npm tag or a workspace build of the CLI.
11075
+
11076
+ ## Gather the target
11077
+
11078
+ Resolve these values from the request and checkout:
11079
+
11080
+ - Agent directory. Default to the current directory only when it contains
11081
+ one Agent SDK project.
11082
+ - Deployment slug. Default to the normalized directory name.
11083
+ - Git ref. Default to the exact pushed `HEAD` commit.
11084
+ - Team. Use the service account's team unless the request names another.
11085
+ - Cursor-event repositories and CLI-only egress domains.
11086
+
11087
+ Ask one focused question when the target or requested ref is ambiguous.
11088
+ Do not ask for values the checkout or existing deployment supplies.
11089
+
11090
+ ## Use the attached service account
11091
+
11092
+ Continue only when `CURSOR_SERVICE_ACCOUNT_KEY` is present. Never print
11093
+ the value. Do not run `agent-sdk login`.
11094
+
11095
+ On Linux and macOS, run every Agent SDK command through this wrapper:
11096
+
11097
+ ```bash
11098
+ CURSOR_JULY_VERSION="__CURSOR_JULY_VERSION__"
11099
+ case "$CURSOR_JULY_VERSION" in
11100
+ __CURSOR_JULY_*__)
11101
+ echo "The Agent SDK deploy skill is not bound to a package version." >&2
11102
+ exit 1
11103
+ ;;
11104
+ esac
11105
+
11106
+ run_agent_sdk() {
11107
+ env -u NODE_OPTIONS -u CURSOR_API_KEY \
11108
+ CURSOR_API_KEY_FILE=/dev/null \
11109
+ npx --yes "@cursor/july@$CURSOR_JULY_VERSION" "$@"
11110
+ }
11111
+ ```
11112
+
11113
+ The installed copy pins the release it came from. The wrapper isolates that
11114
+ published CLI and the attached service account. The presence check prevents a
11115
+ stored personal login from being used when the attachment is missing.
11116
+
11117
+ Verify the principal before any write:
11118
+
11119
+ ```bash
11120
+ test -n "${CURSOR_SERVICE_ACCOUNT_KEY:-}" || {
11121
+ echo "No attached Cursor service account." >&2
11122
+ exit 1
11123
+ }
11124
+ run_agent_sdk whoami --json
11125
+ ```
11126
+
11127
+ Continue only when `credentialSource` is `service-account`. Record the
11128
+ team and complete service-account ID for the final report. Stop on an
11129
+ authentication, team-access, repository-scope, or hosting-entitlement
11130
+ error.
11131
+
11132
+ ## Pin pushed source
11133
+
11134
+ Find the Git root and commit:
11135
+
11136
+ ```bash
11137
+ GIT_ROOT="$(git -C "$AGENT_DIR" rev-parse --show-toplevel)"
11138
+ SOURCE_REF="$(git -C "$GIT_ROOT" rev-parse HEAD)"
11139
+ git -C "$GIT_ROOT" status --short
11140
+ git -C "$GIT_ROOT" branch -r --contains "$SOURCE_REF"
11141
+ ```
11142
+
11143
+ Use a ref from the request instead of `SOURCE_REF` when the person names
11144
+ one. Confirm the selected commit exists on the remote. If intended
11145
+ changes are uncommitted or unpushed, stop and explain they will not ship.
11146
+ Do not commit or push unless the request includes that work.
11147
+
11148
+ The CLI infers the normalized HTTPS `origin`, agent path, and slug from
11149
+ `--dir`. Pass `--repo`, `--path`, or `--slug` when the request overrides
11150
+ the inferred value. Never print a remote URL that contains credentials.
11151
+
11152
+ ## Preserve deployment inputs
11153
+
11154
+ A redeploy replaces its source, Cursor-event repositories, and
11155
+ CLI-supplied egress domains. Preserve the existing
11156
+ `cursorEventRepos` and `egressAllowedDomains` unless the request or
11157
+ project changes them. Reapply each value with:
11158
+
11159
+ ```text
11160
+ --cursor-events-repo <owner/repo>
11161
+ --allow-domain <hostname>
11162
+ ```
11163
+
11164
+ Read existing settings through a filter that selects only those two
11165
+ fields:
11166
+
11167
+ ```bash
11168
+ run_agent_sdk deployment "$SLUG" --json |
11169
+ node -e '
11170
+ let raw = "";
11171
+ process.stdin.setEncoding("utf8");
11172
+ process.stdin.on("data", (chunk) => { raw += chunk; });
11173
+ process.stdin.on("end", () => {
11174
+ const value = JSON.parse(raw);
11175
+ console.log(JSON.stringify({
11176
+ cursorEventRepos: value.cursorEventRepos ?? [],
11177
+ egressAllowedDomains: value.egressAllowedDomains ?? [],
11178
+ }, null, 2));
11179
+ });
11180
+ '
11181
+ ```
11182
+
11183
+ Never print or save the full response because it can contain short-lived
11184
+ engine-access headers.
11185
+
11186
+ For a new SCM-channel deployment, ask which repositories should wake the
11187
+ agent when the project does not declare the answer. The source repository
11188
+ is not always the event repository.
11189
+
11190
+ ## Validate, deploy, and verify
11191
+
11192
+ Validate with the same published package and principal:
11193
+
11194
+ ```bash
11195
+ run_agent_sdk validate --dir "$AGENT_DIR"
11196
+ ```
11197
+
11198
+ Stop on validation errors. Review hosting warnings about secret names,
11199
+ egress domains, and channel configuration before continuing.
11200
+
11201
+ Deploy the pinned source. Add the preserved or requested repeatable
11202
+ flags:
11203
+
11204
+ ```bash
11205
+ run_agent_sdk deploy \
11206
+ --dir "$AGENT_DIR" \
11207
+ --ref "$SOURCE_REF"
11208
+ ```
11209
+
11210
+ The command waits for a terminal deployment state. Do not treat
11211
+ `pending`, `accepted`, or `deploying` as success. If the command is
11212
+ interrupted, resume inspection with:
11213
+
11214
+ ```bash
11215
+ run_agent_sdk deployment "$SLUG"
11216
+ ```
11217
+
11218
+ A first deployment prints a one-time alias token. Never paste it into
11219
+ chat or expose it in an agent-captured terminal. Before a first deploy,
11220
+ ask for a secure destination or ask the person to run the final command
11221
+ in a private terminal. When writing the token, use mode `0600` and print
11222
+ only the path. Redeploys do not print the token. JSON deployment output
11223
+ can also contain credentials, so filter it before display.
11224
+
11225
+ Finish only when the deployment reports `running`. Report the slug,
11226
+ team, generation, pinned ref, and complete service-account ID. On
11227
+ failure, report the deployment status and sanitized error without
11228
+ retrying a different ref or principal.
11229
+
11230
+ ---
11231
+
10925
11232
  Source: /docs/skills/evals.md
10926
11233
 
10927
11234
  # Agent SDK evals
10928
11235
 
10929
11236
  Fixed input, model turn, gates on the trajectory. Files live at
10930
11237
  project-root `evals/**/*.eval.ts`. `agent/evals/` is ignored.
11238
+ Guide: `docs/evals.md`.
10931
11239
 
10932
- Live traffic variants: `skills/ab/SKILL.md`. That is not a test
10933
- runner.
11240
+ Observing production: hooks. Tool logic without a model:
11241
+ `agent-sdk call <tool> --input '{...}'`.
10934
11242
 
10935
11243
  ```bash
10936
11244
  agent-sdk eval --dir . --list
@@ -10944,9 +11252,13 @@ agent-sdk eval --dir . --tag smoke
10944
11252
  | `evals/weather.eval.ts` + `test` | `weather` |
10945
11253
  | `evals/weather/nyc.eval.ts` + `test` | `weather/nyc` |
10946
11254
  | `evals/weather.eval.ts` + `{ id: "nyc" }` | `weather/nyc` |
11255
+ | `evals/sql.eval.ts` exporting an array of `defineEval` calls | `sql/0000`, `sql/0001`, ... |
10947
11256
 
10948
11257
  `eval` boots an ephemeral server and a temp state root. `--url`
10949
- points at a running agent. Model turns need a Cursor credential (`CURSOR_API_KEY`, then `CURSOR_API_KEY_FILE`, then `CURSOR_SERVICE_ACCOUNT_KEY`, then `agent-sdk login`).
11258
+ points at a running agent. Model turns need a Cursor credential
11259
+ (`CURSOR_API_KEY`, then `CURSOR_API_KEY_FILE`, then
11260
+ `CURSOR_SERVICE_ACCOUNT_KEY`, then `agent-sdk login`); `--list` does
11261
+ not. Node 22.13+, never Bun.
10950
11262
 
10951
11263
  ## Seeding
10952
11264
 
@@ -10996,9 +11308,30 @@ export default defineEvalConfig({
10996
11308
  });
10997
11309
  ```
10998
11310
 
10999
- Either `test(t)` or `cases`, not both. `t.send` waits for park/fail.
11000
- `workspaceFiles` seeds the first turn. Assert with `t.succeeded()`,
11001
- `calledTool` / `notCalledTool`, `t.check(t.reply, …)`, `t.metric`.
11311
+ Either `test(t)` or `cases`, not both. `t.send` waits for the turn to
11312
+ settle (complete, park, or fail) and returns it; `turn.calledTool(...)`
11313
+ reads only that turn, `t.*` reads the whole run. First-send options:
11314
+ `workspaceFiles`, `workspaceDir`, `cloud`.
11315
+
11316
+ Gates: `t.succeeded()` / `t.parked()`,
11317
+ `calledTool(name, { input, output, status, count })` / `notCalledTool`,
11318
+ `toolOrder`, `maxToolCalls`, `taggedArtifact`, `event` / `notEvent`,
11319
+ `t.check(value, includes | equals | matches | similarity | satisfies)`.
11320
+ Records: `t.metric`, `t.log`, `t.score` (soft 0-1).
11321
+
11322
+ Severity is on the handle: default gate; `.soft()` tracked;
11323
+ `.atLeast(0.8)` soft with a bar (verdict `scored`, exit 1 only under
11324
+ `--strict`); `.gate(0.8)` hard. `t.judge.factuality`, `summarizes`,
11325
+ `closedQA`, and `sql` are soft by default and need a judge model
11326
+ (`judge` in `evals.config.ts`, on the eval or case, or per-call
11327
+ `{ model }`).
11328
+
11329
+ ## Side effects
11330
+
11331
+ Eval sessions run real tools. Guard actuation with
11332
+ `ctx.session.purpose === "eval"` in the tool, hook, or `defineResult`
11333
+ and return a shaped result so `calledTool` can still assert the
11334
+ decision.
11002
11335
 
11003
11336
  ## What to gate
11004
11337
 
@@ -11020,6 +11353,18 @@ drifting inputs (pin them).
11020
11353
  | GitHub | `github replay … --dry-run --out fixtures/github` |
11021
11354
  | Host-prep PR review | A team-owned PR; gate findings shape, not counts |
11022
11355
  | Workspace | `workspaceFiles` in `t.send` |
11356
+ | Dataset | `loadJson` / `loadJsonl` / `loadYaml` from `@cursor/july/evals/loaders`, paths from the project root |
11357
+
11358
+ ## Debug and CI
11359
+
11360
+ A failed local run leaves `evals/<case-id>.json` (assertions, inputs,
11361
+ tool I/O, final text, logs) under `evals/<stamp>/` in the project state
11362
+ directory (the CLI prints the path); read it before editing the eval.
11363
+ `--state-root <path>` keeps the raw session events too.
11364
+
11365
+ CI: `agent-sdk eval --dir . --json --no-stream --junit reports/evals.xml`.
11366
+ Exit `1` on a failed gate, `2` when nothing matched, `--strict` to fail
11367
+ on `scored`.
11023
11368
 
11024
11369
  Every kept hillclimb change lands an eval that would have failed
11025
11370
  before it. Never weaken a gate to pass a round.
@@ -11073,7 +11418,6 @@ Path is identity. Full list: README "Folder structure".
11073
11418
  | `agent/hooks/*.ts` | Observe-only |
11074
11419
  | `agent/artifacts.ts` | Durable tagged outputs (`defineArtifacts`) |
11075
11420
  | `agent/result.ts` | Host `commit` on the final assistant text (`defineResult`) |
11076
- | `agent/ab.ts` or `agent/ab/*.ts` | Live A/B (`defineAB`) |
11077
11421
  | `agent/otel.ts` | OpenTelemetry (`defineOtel`) |
11078
11422
  | `agent/schedules/*` | Cron. Never auto-fire under `--dev` |
11079
11423
  | `agent/sandbox/workspace/` | Session seed files (local only) |
@@ -11121,11 +11465,11 @@ only this agent's directory.
11121
11465
  | --- | --- |
11122
11466
  | Scaffold | `skills/create-agent/SKILL.md` |
11123
11467
  | Evals | `skills/evals/SKILL.md` |
11124
- | Live A/B | `skills/ab/SKILL.md` |
11125
11468
  | OpenTelemetry | `skills/otel/SKILL.md` |
11126
11469
  | GitHub | `skills/github/SKILL.md` |
11127
11470
  | Slack | `skills/setup-slack/SKILL.md` |
11128
11471
  | Host MCP OAuth | `skills/mcp-auth/SKILL.md` |
11472
+ | Managed deployment | `skills/deploy/SKILL.md` |
11129
11473
  | Local triage | `skills/debug/SKILL.md` |
11130
11474
  | Measured improvement | `skills/hillclimb/SKILL.md` |
11131
11475
 
@@ -11305,12 +11649,12 @@ to use each one.
11305
11649
  | [framework-map](/docs/skills/framework-map.md) | Learn the project layout and runtimes |
11306
11650
  | [create-agent](/docs/skills/create-agent.md) | Scaffold and verify a new agent |
11307
11651
  | [evals](/docs/skills/evals.md) | Write fixtures and regression checks |
11308
- | [ab](/docs/skills/ab.md) | Compare variants on live traffic |
11309
11652
  | [otel](/docs/skills/otel.md) | Export OpenTelemetry traces |
11310
11653
  | [hillclimb](/docs/skills/hillclimb.md) | Improve an agent against fixed inputs |
11311
11654
  | [github](/docs/skills/github.md) | Add GitHub webhooks and replay events |
11312
11655
  | [setup-slack](/docs/skills/setup-slack.md) | Connect an agent to Slack |
11313
11656
  | [mcp-auth](/docs/skills/mcp-auth.md) | Authorize host MCP OAuth |
11657
+ | [deploy](/docs/skills/deploy.md) | Deploy with an attached service account |
11314
11658
  | [debug](/docs/skills/debug.md) | Diagnose a local run |
11315
11659
 
11316
11660
  ---
@@ -11605,8 +11949,8 @@ Source: /docs/storage.md
11605
11949
  # Storage
11606
11950
 
11607
11951
  The Agent SDK owns durable storage for sessions, continuation tokens,
11608
- reminders, playground eval history, and live A/B samples. It chooses the
11609
- keys, when to read and write, and how to restore after restart.
11952
+ reminders, and playground eval history. It chooses the keys, when to
11953
+ read and write, and how to restore after restart.
11610
11954
 
11611
11955
  The Agent SDK owns key encoding. Backends must accept the keys they are
11612
11956
  given. Do not fail `put` to enforce a shorter cap.
@@ -11640,9 +11984,9 @@ export default defineStorage({
11640
11984
  ## Which fields to provide
11641
11985
 
11642
11986
  Implement the small KV core (`put`/`get`/`delete`/`list` plus the `cas`
11643
- group) and you get **full functionality**: eval-run and A/B history are
11644
- derived over the core automatically. The dedicated `evals` / `abs` groups
11645
- are backend-native optimizations, not required-or-lose-history hooks.
11987
+ group) and you get **full functionality**: eval-run history is derived
11988
+ over the core automatically. The dedicated `evals` group is a
11989
+ backend-native optimization, not a required-or-lose-history hook.
11646
11990
 
11647
11991
  | Field | Required | Role |
11648
11992
  | --- | --- | --- |
@@ -11653,36 +11997,28 @@ are backend-native optimizations, not required-or-lose-history hooks.
11653
11997
  | `delete` | For cleanup | Remove a key |
11654
11998
  | `name` | No | Label surfaced on `GET /v1/info` diagnostics |
11655
11999
  | `policy` | No | Timing knobs; see [Policy](#policy) |
11656
- | `evals` | No | Backend-native eval-runs table; derived over the core when omitted. See [Eval and A/B tables](#eval-and-a-b-tables) |
11657
- | `abs` | No | Backend-native A/B metrics table; derived over the core when omitted. See [Eval and A/B tables](#eval-and-a-b-tables) |
12000
+ | `evals` | No | Backend-native eval-runs table; derived over the core when omitted. See [Eval table](#eval-table) |
11658
12001
 
11659
12002
  A throwing `put` is logged and dropped. It never fails a turn. When
11660
12003
  resolving a missing continuation token, a throwing `get` fails the
11661
12004
  follow-up so a store outage does not open a new session. Return
11662
12005
  `undefined` only for a real miss.
11663
12006
 
11664
- ## Eval and A/B tables
12007
+ ## Eval table
11665
12008
 
11666
- Two dedicated table groups carry structured rows instead of opaque KV
11667
- values. Both are **optional optimizations**: when a group is not
11668
- authored, `defineStorage` derives it over the KV core, so a backend that
11669
- implements only the core loses nothing. Author a group only when the
11670
- backend has a better native shape (a real database table, an analytics
11671
- pipeline). The built-in `fileKv` and `cursorHostedStorage` both do.
12009
+ The dedicated `evals` table group carries structured rows instead of
12010
+ opaque KV values. It is an **optional optimization**: when the group is
12011
+ not authored, `defineStorage` derives it over the KV core, so a backend
12012
+ that implements only the core loses nothing. Author the group only when
12013
+ the backend has a better native shape (a real database table, an
12014
+ analytics pipeline). The built-in `fileKv` and `cursorHostedStorage`
12015
+ both do.
11672
12016
 
11673
12017
  `evals` keeps playground eval batches across restarts (`put`, `delete`,
11674
12018
  `list` over run snapshots keyed by `runId`). A core missing `delete` or
11675
12019
  `list` leaves eval history in memory until restart.
11676
12020
  See [Evals](/docs/evals.md#configure-eval-runs).
11677
12021
 
11678
- `abs` exports live A/B metrics: `putSample` appends one cumulative
11679
- metric sample per enrolled experiment on each completed or failed turn;
11680
- optional `putSnapshot` / `getSnapshot` store and serve back the latest
11681
- aggregate so a replacement host can still serve the A/Bs surface.
11682
- `putSample` and `putSnapshot` need only core `put`; `getSnapshot` needs
11683
- core `get`. Session event logs remain the assignment source of truth
11684
- either way. See [Live A/B metrics](/docs/ab.md).
11685
-
11686
12022
  ## Policy
11687
12023
 
11688
12024
  Two knobs change behavior:
@@ -11717,8 +12053,7 @@ With `get` and `list`, serve can rebuild local state from your store:
11717
12053
  - At startup, the Agent SDK loads recent sessions up to the restore caps.
11718
12054
  - On demand, a missing continuation token resolves through the store
11719
12055
  and resumes that session.
11720
- - Playground eval history and A/B aggregates can load from the same
11721
- sink.
12056
+ - Playground eval history can load from the same sink.
11722
12057
 
11723
12058
  A turn in flight at crash time is not replayed. The next follow-up
11724
12059
  resumes from the last flushed state.