@cursor/july 0.1.91 → 0.1.93

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (351) hide show
  1. package/AGENTS.md +4 -0
  2. package/README.md +117 -162
  3. package/dist/channels/deployments/deployments-channel.d.ts +7 -0
  4. package/dist/channels/deployments/deployments-channel.d.ts.map +1 -1
  5. package/dist/channels/deployments/deployments-channel.js +26 -2
  6. package/dist/channels/deployments/types.d.ts +8 -0
  7. package/dist/channels/deployments/types.d.ts.map +1 -1
  8. package/dist/channels/github/github-channel.d.ts +3 -0
  9. package/dist/channels/github/github-channel.d.ts.map +1 -1
  10. package/dist/channels/github/github-channel.js +28 -56
  11. package/dist/continuation.d.ts +1 -1
  12. package/dist/continuation.js +1 -1
  13. package/dist/docs/404.html +4 -2
  14. package/dist/docs/ab.html +10 -8
  15. package/dist/docs/ab.md +332 -0
  16. package/dist/docs/assets/{ab.md.CVzWxLoB.js → ab.md.DJo5r4R-.js} +4 -4
  17. package/dist/docs/assets/{ab.md.CVzWxLoB.lean.js → ab.md.DJo5r4R-.lean.js} +1 -1
  18. package/dist/docs/assets/{app.Bci6CM9E.js → app.CjWU-x0z.js} +1 -1
  19. package/dist/docs/assets/building-with-agents.md.DI4mEzlt.js +13 -0
  20. package/dist/docs/assets/{building-with-agents.md.DH8A_cHA.lean.js → building-with-agents.md.DI4mEzlt.lean.js} +1 -1
  21. package/dist/docs/assets/chunks/@localSearchIndexroot.ChpIC3Zy.js +1 -0
  22. package/dist/docs/assets/chunks/{VPLocalSearchBox.BCPT6xA-.js → VPLocalSearchBox.Cxy8ySFQ.js} +1 -1
  23. package/dist/docs/assets/chunks/{theme.BEA8BF3c.js → theme.Dvq1Bktu.js} +2 -2
  24. package/dist/docs/assets/concepts.md.F6AiPorA.js +1 -0
  25. package/dist/docs/assets/{concepts.md.CRfU3bVg.lean.js → concepts.md.F6AiPorA.lean.js} +1 -1
  26. package/dist/docs/assets/{deployment.md.DX_hc3ze.js → deployment.md.DoLFAzfm.js} +6 -6
  27. package/dist/docs/assets/{evals.md.a0SMN6r9.js → evals.md.lfJoEVc8.js} +6 -6
  28. package/dist/docs/assets/{evals.md.a0SMN6r9.lean.js → evals.md.lfJoEVc8.lean.js} +1 -1
  29. package/dist/docs/assets/{example-agents_approval-buddy.md.DNL83puR.js → example-agents_approval-buddy.md.DmezILPg.js} +1 -1
  30. package/dist/docs/assets/{example-agents_benny.md.C40vHRLc.js → example-agents_benny.md.B0kwY7D_.js} +2 -4
  31. package/dist/docs/assets/{example-agents_benny.md.C40vHRLc.lean.js → example-agents_benny.md.B0kwY7D_.lean.js} +1 -1
  32. package/dist/docs/assets/{example-agents_codebase-wiki.md.Dftj_tPp.js → example-agents_codebase-wiki.md.BBNw9Ekr.js} +3 -3
  33. package/dist/docs/assets/{example-agents_codebase-wiki.md.Dftj_tPp.lean.js → example-agents_codebase-wiki.md.BBNw9Ekr.lean.js} +1 -1
  34. package/dist/docs/assets/{example-agents_concierge.md.MrKpQndp.js → example-agents_concierge.md.BzB2b20R.js} +2 -3
  35. package/dist/docs/assets/example-agents_index.md.ChBp0AX6.js +2 -0
  36. package/dist/docs/assets/example-agents_index.md.ChBp0AX6.lean.js +1 -0
  37. package/dist/docs/assets/{example-agents_knowledge-base.md.DqKqHQ9u.js → example-agents_knowledge-base.md.CrA85ig-.js} +1 -1
  38. package/dist/docs/assets/{example-agents_security-reviewer.md.Bai6D0Ee.js → example-agents_security-reviewer.md.74pPpWYj.js} +1 -1
  39. package/dist/docs/assets/{example-agents_weather-agent.md.lVEAbWFf.js → example-agents_weather-agent.md.CaGpmw3Y.js} +2 -2
  40. package/dist/docs/assets/{guides_agent-to-agent.md.BCeVdJRJ.js → guides_agent-to-agent.md.B3JIaAqz.js} +1 -1
  41. package/dist/docs/assets/{guides_cloud-runtime.md.BSMLIBHr.js → guides_cloud-runtime.md.BnvjPiia.js} +2 -2
  42. package/dist/docs/assets/{guides_cloud-runtime.md.BSMLIBHr.lean.js → guides_cloud-runtime.md.BnvjPiia.lean.js} +1 -1
  43. package/dist/docs/assets/{guides_convert-automation.md.D06eIzea.js → guides_convert-automation.md.Bboisykk.js} +1 -1
  44. package/dist/docs/assets/{guides_github.md.Cdt1s2QC.js → guides_github.md.DqJhuaN1.js} +5 -5
  45. package/dist/docs/assets/{guides_github.md.Cdt1s2QC.lean.js → guides_github.md.DqJhuaN1.lean.js} +1 -1
  46. package/dist/docs/assets/{guides_mcp-oauth.md.Du0f7pGU.js → guides_mcp-oauth.md.CJvrXtkN.js} +2 -2
  47. package/dist/docs/assets/{guides_slack.md.DiUmk_Oi.js → guides_slack.md.mqeNKs84.js} +2 -2
  48. package/dist/docs/assets/{guides_webhooks.md.BpnIdO0i.js → guides_webhooks.md.DKdA43Qm.js} +2 -2
  49. package/dist/docs/assets/{hillclimbing.md.ywF3yDAd.js → hillclimbing.md.DhESf3OO.js} +1 -1
  50. package/dist/docs/assets/{index.md.BAaMXLFd.js → index.md.B-lVR4wT.js} +3 -3
  51. package/dist/docs/assets/{index.md.BAaMXLFd.lean.js → index.md.B-lVR4wT.lean.js} +1 -1
  52. package/dist/docs/assets/{quickstart.md.DsrarzEg.js → quickstart.md.BrmfrrIr.js} +1 -1
  53. package/dist/docs/assets/{reference_agent-config.md.Bqylgw50.js → reference_agent-config.md.Cp_x38Nl.js} +3 -3
  54. package/dist/docs/assets/{reference_channels.md.DQZjCnyh.js → reference_channels.md.Cd2f2iyV.js} +2 -2
  55. package/dist/docs/assets/{reference_channels.md.DQZjCnyh.lean.js → reference_channels.md.Cd2f2iyV.lean.js} +1 -1
  56. package/dist/docs/assets/{reference_cli.md.B7GkAJRC.js → reference_cli.md.D9KESDsD.js} +10 -11
  57. package/dist/docs/assets/{reference_cli.md.B7GkAJRC.lean.js → reference_cli.md.D9KESDsD.lean.js} +1 -1
  58. package/dist/docs/assets/{reference_connections.md.DYidrb-j.js → reference_connections.md.DB6SsN6U.js} +3 -3
  59. package/dist/docs/assets/reference_hooks.md.BxN87gCw.js +14 -0
  60. package/dist/docs/assets/{reference_hooks.md.B9FSgdDe.lean.js → reference_hooks.md.BxN87gCw.lean.js} +1 -1
  61. package/dist/docs/assets/reference_http-api.md.C68BERYr.js +11 -0
  62. package/dist/docs/assets/reference_http-api.md.C68BERYr.lean.js +1 -0
  63. package/dist/docs/assets/{reference_instructions.md.DhNCOl7r.js → reference_instructions.md.CR7XSsGk.js} +3 -3
  64. package/dist/docs/assets/{reference_instructions.md.DhNCOl7r.lean.js → reference_instructions.md.CR7XSsGk.lean.js} +1 -1
  65. package/dist/docs/assets/reference_playground.md.DnX5nL-B.js +1 -0
  66. package/dist/docs/assets/reference_playground.md.DnX5nL-B.lean.js +1 -0
  67. package/dist/docs/assets/{reference_project-layout.md.CwkSbEWT.js → reference_project-layout.md.WN9nwJht.js} +2 -2
  68. package/dist/docs/assets/{reference_prompt.md.DZUMtLPD.js → reference_prompt.md.DnaD5dNK.js} +1 -1
  69. package/dist/docs/assets/{reference_schedules.md.DNipebiG.js → reference_schedules.md.DI_JrHgq.js} +1 -1
  70. package/dist/docs/assets/reference_sessions.md.D0mIh4KK.js +1 -0
  71. package/dist/docs/assets/{reference_sessions.md.tUFzz98S.lean.js → reference_sessions.md.D0mIh4KK.lean.js} +1 -1
  72. package/dist/docs/assets/{reference_skills.md.B5ZEuHfG.js → reference_skills.md.BFW9retM.js} +1 -1
  73. package/dist/docs/assets/{reference_tools.md.wpaJtHn6.js → reference_tools.md.DuKvkYWG.js} +4 -4
  74. package/dist/docs/assets/{reference_tools.md.wpaJtHn6.lean.js → reference_tools.md.DuKvkYWG.lean.js} +1 -1
  75. package/dist/docs/assets/scaffolding-agents.md.D7UUkWw0.js +1 -0
  76. package/dist/docs/assets/{scaffolding-agents.md.CRDDUtYJ.lean.js → scaffolding-agents.md.D7UUkWw0.lean.js} +1 -1
  77. package/dist/docs/assets/{storage.md.JbjlHWZ6.js → storage.md.BOHeqk2M.js} +5 -5
  78. package/dist/docs/assets/{storage.md.JbjlHWZ6.lean.js → storage.md.BOHeqk2M.lean.js} +1 -1
  79. package/dist/docs/assets/{templates_agentic-owners.md.DSJSIpWU.js → templates_agentic-owners.md.DqtPdm6f.js} +2 -2
  80. package/dist/docs/assets/{templates_pr-autofixer.md.1HAR3RXE.js → templates_pr-autofixer.md.R4K_qytS.js} +2 -2
  81. package/dist/docs/assets/{templates_pr-autofixer.md.1HAR3RXE.lean.js → templates_pr-autofixer.md.R4K_qytS.lean.js} +1 -1
  82. package/dist/docs/assets/troubleshooting.md.vCWwvqcJ.js +1 -0
  83. package/dist/docs/assets/{troubleshooting.md.DYECCZiJ.lean.js → troubleshooting.md.vCWwvqcJ.lean.js} +1 -1
  84. package/dist/docs/building-with-agents.html +9 -7
  85. package/dist/docs/building-with-agents.md +118 -0
  86. package/dist/docs/concepts.html +7 -8
  87. package/dist/docs/concepts.md +169 -0
  88. package/dist/docs/deployment.html +13 -11
  89. package/dist/docs/deployment.md +462 -0
  90. package/dist/docs/evals.html +12 -10
  91. package/dist/docs/evals.md +460 -0
  92. package/dist/docs/example-agents/approval-buddy.html +7 -5
  93. package/dist/docs/example-agents/approval-buddy.md +266 -0
  94. package/dist/docs/example-agents/benny.html +7 -7
  95. package/dist/docs/example-agents/benny.md +173 -0
  96. package/dist/docs/example-agents/bugbot.html +6 -4
  97. package/dist/docs/example-agents/bugbot.md +229 -0
  98. package/dist/docs/example-agents/codebase-wiki.html +8 -6
  99. package/dist/docs/example-agents/codebase-wiki.md +167 -0
  100. package/dist/docs/example-agents/codeowners-review.html +6 -4
  101. package/dist/docs/example-agents/codeowners-review.md +192 -0
  102. package/dist/docs/example-agents/concierge.html +9 -8
  103. package/dist/docs/example-agents/concierge.md +200 -0
  104. package/dist/docs/example-agents/index.html +8 -6
  105. package/dist/docs/example-agents/index.md +99 -0
  106. package/dist/docs/example-agents/knowledge-base.html +8 -6
  107. package/dist/docs/example-agents/knowledge-base.md +168 -0
  108. package/dist/docs/example-agents/oncall.html +6 -4
  109. package/dist/docs/example-agents/oncall.md +212 -0
  110. package/dist/docs/example-agents/security-reviewer.html +9 -7
  111. package/dist/docs/example-agents/security-reviewer.md +265 -0
  112. package/dist/docs/example-agents/slack-agent.html +6 -4
  113. package/dist/docs/example-agents/slack-agent.md +142 -0
  114. package/dist/docs/example-agents/weather-agent.html +9 -7
  115. package/dist/docs/example-agents/weather-agent.md +297 -0
  116. package/dist/docs/guides/agent-to-agent.html +7 -5
  117. package/dist/docs/guides/agent-to-agent.md +113 -0
  118. package/dist/docs/guides/cloud-runtime.html +8 -6
  119. package/dist/docs/guides/cloud-runtime.md +114 -0
  120. package/dist/docs/guides/convert-automation.html +8 -6
  121. package/dist/docs/guides/convert-automation.md +171 -0
  122. package/dist/docs/guides/github.html +11 -9
  123. package/dist/docs/guides/github.md +275 -0
  124. package/dist/docs/guides/human-in-the-loop.html +6 -4
  125. package/dist/docs/guides/human-in-the-loop.md +126 -0
  126. package/dist/docs/guides/mcp-oauth.html +8 -6
  127. package/dist/docs/guides/mcp-oauth.md +159 -0
  128. package/dist/docs/guides/opentelemetry.html +6 -4
  129. package/dist/docs/guides/opentelemetry.md +209 -0
  130. package/dist/docs/guides/slack.html +9 -7
  131. package/dist/docs/guides/slack.md +337 -0
  132. package/dist/docs/guides/webhooks.html +8 -6
  133. package/dist/docs/guides/webhooks.md +463 -0
  134. package/dist/docs/hashmap.json +1 -1
  135. package/dist/docs/hillclimbing.html +8 -6
  136. package/dist/docs/hillclimbing.md +88 -0
  137. package/dist/docs/index.html +8 -6
  138. package/dist/docs/index.md +171 -0
  139. package/dist/docs/llms-full.txt +10968 -0
  140. package/dist/docs/llms.txt +74 -0
  141. package/dist/docs/quickstart.html +7 -5
  142. package/dist/docs/quickstart.md +364 -0
  143. package/dist/docs/reference/agent-config.html +10 -8
  144. package/dist/docs/reference/agent-config.md +251 -0
  145. package/dist/docs/reference/artifacts.html +6 -4
  146. package/dist/docs/reference/artifacts.md +112 -0
  147. package/dist/docs/reference/channels.html +8 -6
  148. package/dist/docs/reference/channels.md +244 -0
  149. package/dist/docs/reference/cli.html +16 -15
  150. package/dist/docs/reference/cli.md +947 -0
  151. package/dist/docs/reference/connections.html +10 -8
  152. package/dist/docs/reference/connections.md +263 -0
  153. package/dist/docs/reference/hooks.html +8 -6
  154. package/dist/docs/reference/hooks.md +98 -0
  155. package/dist/docs/reference/http-api.html +9 -7
  156. package/dist/docs/reference/http-api.md +247 -0
  157. package/dist/docs/reference/instructions.html +8 -6
  158. package/dist/docs/reference/instructions.md +74 -0
  159. package/dist/docs/reference/playground.html +7 -5
  160. package/dist/docs/reference/playground.md +57 -0
  161. package/dist/docs/reference/project-layout.html +9 -7
  162. package/dist/docs/reference/project-layout.md +107 -0
  163. package/dist/docs/reference/prompt.html +8 -6
  164. package/dist/docs/reference/prompt.md +42 -0
  165. package/dist/docs/reference/schedules.html +8 -6
  166. package/dist/docs/reference/schedules.md +214 -0
  167. package/dist/docs/reference/sessions.html +7 -12
  168. package/dist/docs/reference/sessions.md +159 -0
  169. package/dist/docs/reference/skills.html +8 -6
  170. package/dist/docs/reference/skills.md +83 -0
  171. package/dist/docs/reference/subagents.html +6 -4
  172. package/dist/docs/reference/subagents.md +71 -0
  173. package/dist/docs/reference/tools.html +10 -8
  174. package/dist/docs/reference/tools.md +293 -0
  175. package/dist/docs/scaffolding-agents.html +7 -5
  176. package/dist/docs/scaffolding-agents.md +129 -0
  177. package/dist/docs/storage.html +11 -9
  178. package/dist/docs/storage.md +176 -0
  179. package/dist/docs/templates/agentic-owners.html +9 -7
  180. package/dist/docs/templates/agentic-owners.md +92 -0
  181. package/dist/docs/templates/demo.html +6 -4
  182. package/dist/docs/templates/demo.md +79 -0
  183. package/dist/docs/templates/pr-autofixer.html +8 -6
  184. package/dist/docs/templates/pr-autofixer.md +128 -0
  185. package/dist/docs/templates/security-reviewer.html +6 -4
  186. package/dist/docs/templates/security-reviewer.md +84 -0
  187. package/dist/docs/templates/triage.html +6 -4
  188. package/dist/docs/templates/triage.md +98 -0
  189. package/dist/docs/troubleshooting.html +7 -5
  190. package/dist/docs/troubleshooting.md +111 -0
  191. package/dist/internal/authored-alias-hooks.d.ts +14 -11
  192. package/dist/internal/authored-alias-hooks.d.ts.map +1 -1
  193. package/dist/internal/authored-alias-hooks.js +14 -11
  194. package/dist/internal/authored-loaders.d.ts +7 -6
  195. package/dist/internal/authored-loaders.d.ts.map +1 -1
  196. package/dist/internal/authored-loaders.js +14 -10
  197. package/dist/internal/cli-deploy.d.ts +1 -1
  198. package/dist/internal/cli-deploy.js +5 -5
  199. package/dist/internal/continuation-channel.d.ts +6 -3
  200. package/dist/internal/continuation-channel.d.ts.map +1 -1
  201. package/dist/internal/continuation-channel.js +44 -40
  202. package/dist/internal/continuation-identity.d.ts +17 -16
  203. package/dist/internal/continuation-identity.d.ts.map +1 -1
  204. package/dist/internal/continuation-identity.js +109 -36
  205. package/dist/internal/deploy-manifest.d.ts +2 -2
  206. package/dist/internal/deploy-manifest.d.ts.map +1 -1
  207. package/dist/internal/deploy-manifest.js +4 -9
  208. package/dist/internal/discovery.d.ts.map +1 -1
  209. package/dist/internal/discovery.js +3 -0
  210. package/dist/internal/distribution.d.ts +4 -3
  211. package/dist/internal/distribution.d.ts.map +1 -1
  212. package/dist/internal/distribution.js +4 -3
  213. package/dist/internal/hosted-delivery-protocol.d.ts +38 -0
  214. package/dist/internal/hosted-delivery-protocol.d.ts.map +1 -0
  215. package/dist/internal/hosted-delivery-protocol.js +70 -0
  216. package/dist/internal/hosted-delivery.d.ts +35 -0
  217. package/dist/internal/hosted-delivery.d.ts.map +1 -0
  218. package/dist/internal/hosted-delivery.js +226 -0
  219. package/dist/internal/http-channel.d.ts.map +1 -1
  220. package/dist/internal/http-channel.js +1 -1
  221. package/dist/internal/init-scaffold.d.ts.map +1 -1
  222. package/dist/internal/init-scaffold.js +1 -0
  223. package/dist/internal/playground/static.d.ts.map +1 -1
  224. package/dist/internal/playground/static.js +2 -0
  225. package/dist/internal/review-comments.d.ts +186 -63
  226. package/dist/internal/review-comments.d.ts.map +1 -1
  227. package/dist/internal/review-comments.js +350 -168
  228. package/dist/internal/server.d.ts.map +1 -1
  229. package/dist/internal/server.js +21 -3
  230. package/dist/internal/session-engine.d.ts +5 -0
  231. package/dist/internal/session-engine.d.ts.map +1 -1
  232. package/dist/internal/session-engine.js +16 -4
  233. package/dist/internal/shallow-clone.d.ts +8 -2
  234. package/dist/internal/shallow-clone.d.ts.map +1 -1
  235. package/dist/internal/shallow-clone.js +17 -10
  236. package/dist/playground/assets/{index-DDvyC2z6.js → index-D9MFzhNE.js} +1 -1
  237. package/dist/playground/index.html +1 -1
  238. package/dist/types.d.ts +9 -17
  239. package/dist/types.d.ts.map +1 -1
  240. package/docs/README.md +2 -10
  241. package/docs/ab.md +7 -13
  242. package/docs/building-with-agents.md +5 -11
  243. package/docs/concepts.md +12 -17
  244. package/docs/deployment.md +8 -10
  245. package/docs/evals.md +16 -37
  246. package/docs/example-agents/approval-buddy.md +1 -1
  247. package/docs/example-agents/benny.md +4 -13
  248. package/docs/example-agents/codebase-wiki.md +5 -8
  249. package/docs/example-agents/concierge.md +2 -3
  250. package/docs/example-agents/index.md +6 -9
  251. package/docs/example-agents/knowledge-base.md +2 -2
  252. package/docs/example-agents/security-reviewer.md +5 -5
  253. package/docs/example-agents/weather-agent.md +4 -3
  254. package/docs/guides/agent-to-agent.md +1 -1
  255. package/docs/guides/cloud-runtime.md +8 -25
  256. package/docs/guides/convert-automation.md +3 -3
  257. package/docs/guides/github.md +11 -23
  258. package/docs/guides/mcp-oauth.md +4 -4
  259. package/docs/guides/slack.md +4 -4
  260. package/docs/guides/webhooks.md +3 -3
  261. package/docs/hillclimbing.md +1 -1
  262. package/docs/quickstart.md +1 -1
  263. package/docs/reference/agent-config.md +10 -15
  264. package/docs/reference/channels.md +20 -31
  265. package/docs/reference/cli.md +27 -37
  266. package/docs/reference/connections.md +9 -14
  267. package/docs/reference/hooks.md +10 -14
  268. package/docs/reference/http-api.md +18 -38
  269. package/docs/reference/instructions.md +1 -1
  270. package/docs/reference/playground.md +14 -19
  271. package/docs/reference/project-layout.md +2 -2
  272. package/docs/reference/prompt.md +1 -1
  273. package/docs/reference/schedules.md +1 -2
  274. package/docs/reference/sessions.md +8 -19
  275. package/docs/reference/skills.md +3 -3
  276. package/docs/reference/tools.md +12 -17
  277. package/docs/scaffolding-agents.md +4 -5
  278. package/docs/storage.md +37 -80
  279. package/docs/templates/agentic-owners.md +2 -2
  280. package/docs/templates/pr-autofixer.md +3 -6
  281. package/docs/troubleshooting.md +6 -6
  282. package/package.json +9 -2
  283. package/skills/ab/SKILL.md +3 -0
  284. package/skills/create-agent/SKILL.md +3 -0
  285. package/skills/debug/SKILL.md +3 -0
  286. package/skills/evals/SKILL.md +3 -0
  287. package/skills/framework-map/SKILL.md +3 -0
  288. package/skills/github/SKILL.md +3 -0
  289. package/skills/hillclimb/SKILL.md +3 -0
  290. package/skills/mcp-auth/SKILL.md +3 -0
  291. package/skills/otel/SKILL.md +3 -0
  292. package/skills/setup-slack/SKILL.md +3 -0
  293. package/src/channels/deployments/deployments-channel.ts +32 -2
  294. package/src/channels/deployments/types.ts +8 -0
  295. package/src/channels/github/github-channel.ts +71 -21
  296. package/src/continuation.ts +1 -1
  297. package/src/internal/authored-alias-hooks.ts +14 -11
  298. package/src/internal/authored-loaders.ts +14 -10
  299. package/src/internal/cli-deploy.ts +5 -5
  300. package/src/internal/continuation-channel.ts +62 -45
  301. package/src/internal/continuation-identity.ts +123 -38
  302. package/src/internal/deploy-manifest.ts +5 -9
  303. package/src/internal/discovery.ts +3 -0
  304. package/src/internal/distribution.ts +4 -3
  305. package/src/internal/hosted-delivery-protocol.ts +114 -0
  306. package/src/internal/hosted-delivery.ts +327 -0
  307. package/src/internal/http-channel.ts +0 -2
  308. package/src/internal/init-scaffold.ts +1 -0
  309. package/src/internal/playground/static.ts +2 -0
  310. package/src/internal/review-comments.ts +542 -229
  311. package/src/internal/server.ts +29 -2
  312. package/src/internal/session-engine.ts +29 -7
  313. package/src/internal/shallow-clone.ts +30 -16
  314. package/src/types.ts +9 -17
  315. package/dist/docs/assets/building-with-agents.md.DH8A_cHA.js +0 -13
  316. package/dist/docs/assets/chunks/@localSearchIndexroot.Dv-Q0XtU.js +0 -1
  317. package/dist/docs/assets/concepts.md.CRfU3bVg.js +0 -4
  318. package/dist/docs/assets/example-agents_fsd.md.ZWHWWZPE.js +0 -15
  319. package/dist/docs/assets/example-agents_fsd.md.ZWHWWZPE.lean.js +0 -1
  320. package/dist/docs/assets/example-agents_index.md.QZ8mhr6n.js +0 -2
  321. package/dist/docs/assets/example-agents_index.md.QZ8mhr6n.lean.js +0 -1
  322. package/dist/docs/assets/reference_hooks.md.B9FSgdDe.js +0 -14
  323. package/dist/docs/assets/reference_http-api.md.CSHVobzG.js +0 -11
  324. package/dist/docs/assets/reference_http-api.md.CSHVobzG.lean.js +0 -1
  325. package/dist/docs/assets/reference_playground.md.Dfb92yQf.js +0 -1
  326. package/dist/docs/assets/reference_playground.md.Dfb92yQf.lean.js +0 -1
  327. package/dist/docs/assets/reference_sessions.md.tUFzz98S.js +0 -8
  328. package/dist/docs/assets/scaffolding-agents.md.CRDDUtYJ.js +0 -1
  329. package/dist/docs/assets/troubleshooting.md.DYECCZiJ.js +0 -1
  330. package/dist/docs/example-agents/fsd.html +0 -39
  331. package/docs/example-agents/fsd.md +0 -334
  332. /package/dist/docs/assets/{deployment.md.DX_hc3ze.lean.js → deployment.md.DoLFAzfm.lean.js} +0 -0
  333. /package/dist/docs/assets/{example-agents_approval-buddy.md.DNL83puR.lean.js → example-agents_approval-buddy.md.DmezILPg.lean.js} +0 -0
  334. /package/dist/docs/assets/{example-agents_concierge.md.MrKpQndp.lean.js → example-agents_concierge.md.BzB2b20R.lean.js} +0 -0
  335. /package/dist/docs/assets/{example-agents_knowledge-base.md.DqKqHQ9u.lean.js → example-agents_knowledge-base.md.CrA85ig-.lean.js} +0 -0
  336. /package/dist/docs/assets/{example-agents_security-reviewer.md.Bai6D0Ee.lean.js → example-agents_security-reviewer.md.74pPpWYj.lean.js} +0 -0
  337. /package/dist/docs/assets/{example-agents_weather-agent.md.lVEAbWFf.lean.js → example-agents_weather-agent.md.CaGpmw3Y.lean.js} +0 -0
  338. /package/dist/docs/assets/{guides_agent-to-agent.md.BCeVdJRJ.lean.js → guides_agent-to-agent.md.B3JIaAqz.lean.js} +0 -0
  339. /package/dist/docs/assets/{guides_convert-automation.md.D06eIzea.lean.js → guides_convert-automation.md.Bboisykk.lean.js} +0 -0
  340. /package/dist/docs/assets/{guides_mcp-oauth.md.Du0f7pGU.lean.js → guides_mcp-oauth.md.CJvrXtkN.lean.js} +0 -0
  341. /package/dist/docs/assets/{guides_slack.md.DiUmk_Oi.lean.js → guides_slack.md.mqeNKs84.lean.js} +0 -0
  342. /package/dist/docs/assets/{guides_webhooks.md.BpnIdO0i.lean.js → guides_webhooks.md.DKdA43Qm.lean.js} +0 -0
  343. /package/dist/docs/assets/{hillclimbing.md.ywF3yDAd.lean.js → hillclimbing.md.DhESf3OO.lean.js} +0 -0
  344. /package/dist/docs/assets/{quickstart.md.DsrarzEg.lean.js → quickstart.md.BrmfrrIr.lean.js} +0 -0
  345. /package/dist/docs/assets/{reference_agent-config.md.Bqylgw50.lean.js → reference_agent-config.md.Cp_x38Nl.lean.js} +0 -0
  346. /package/dist/docs/assets/{reference_connections.md.DYidrb-j.lean.js → reference_connections.md.DB6SsN6U.lean.js} +0 -0
  347. /package/dist/docs/assets/{reference_project-layout.md.CwkSbEWT.lean.js → reference_project-layout.md.WN9nwJht.lean.js} +0 -0
  348. /package/dist/docs/assets/{reference_prompt.md.DZUMtLPD.lean.js → reference_prompt.md.DnaD5dNK.lean.js} +0 -0
  349. /package/dist/docs/assets/{reference_schedules.md.DNipebiG.lean.js → reference_schedules.md.DI_JrHgq.lean.js} +0 -0
  350. /package/dist/docs/assets/{reference_skills.md.B5ZEuHfG.lean.js → reference_skills.md.BFW9retM.lean.js} +0 -0
  351. /package/dist/docs/assets/{templates_agentic-owners.md.DSJSIpWU.lean.js → templates_agentic-owners.md.DqtPdm6f.lean.js} +0 -0
@@ -0,0 +1,460 @@
1
+ # Evals
2
+
3
+ An eval is a repeatable check that runs your agent against a fixed input
4
+ and gates the recorded trajectory: the run completed, the right tool
5
+ ran, the reply has the right shape. Evals are how you know a prompt
6
+ tweak helped, a refactor didn't regress the agent, and last month's fix
7
+ is still holding.
8
+
9
+ Evals exercise the same surface your users hit. The runner starts (or
10
+ targets) a real agent server, drives sessions over the public API, and
11
+ grades what comes back. A passing eval means the agent started,
12
+ accepted a message, and did what you asserted.
13
+
14
+ ## Define evals with `defineEval`
15
+
16
+ The Agent SDK discovers evals under the project-root `evals/` directory,
17
+ in `.eval.ts` or `.eval.js` files. That's a sibling of `agent/`, never
18
+ inside it (`agent/evals/` is silently ignored). TypeScript is the normal
19
+ authoring format.
20
+
21
+ The file path is the eval's identity, so you don't author an id.
22
+ Directories group related evals: `evals/builds/api.eval.ts` becomes id
23
+ `builds/api`. An `index` filename collapses to its directory, so
24
+ `evals/builds/index.eval.ts` becomes `builds`.
25
+
26
+ An eval is a single `async test(t)`. You drive the agent with `t` and
27
+ assert on the run with the same `t`:
28
+
29
+ ```ts
30
+ // evals/readiness.eval.ts
31
+ import { defineEval, includes } from "@cursor/july/evals";
32
+
33
+ export default defineEval({
34
+ description: "Inspects a PR without approving it.",
35
+ tags: ["smoke"],
36
+ timeoutMs: 120_000,
37
+ async test(t) {
38
+ await t.send(
39
+ "Is https://github.com/acme/checkout/pull/42 ready to approve?"
40
+ );
41
+ t.succeeded();
42
+ t.calledTool("inspect_pr");
43
+ t.notCalledTool("approve_pr");
44
+ t.check(t.reply, includes(/ready|approve/i));
45
+ },
46
+ });
47
+ ```
48
+
49
+ One file can also hold several datapoints through `cases` (provide
50
+ either `test` or `cases`, not both). Each case id becomes
51
+ `<fileId>/<case.id>`:
52
+
53
+ ```ts
54
+ // evals/prs.eval.ts → prs/checkout, prs/search
55
+ export default defineEval({
56
+ tags: ["smoke", "prs"],
57
+ cases: [
58
+ {
59
+ id: "checkout",
60
+ description: "Checkout PR readiness.",
61
+ async test(t) {
62
+ await t.send(
63
+ "Is https://github.com/acme/checkout/pull/42 ready to approve?"
64
+ );
65
+ t.succeeded();
66
+ t.calledTool("inspect_pr");
67
+ },
68
+ },
69
+ {
70
+ id: "search",
71
+ async test(t) {
72
+ await t.send(
73
+ "Check https://github.com/acme/search/pull/7 before approval."
74
+ );
75
+ t.succeeded();
76
+ t.calledTool("inspect_pr");
77
+ },
78
+ },
79
+ ],
80
+ });
81
+ ```
82
+
83
+ Case ids must be single path segments, unique within the file.
84
+ Each case can set its own `description`, `tags`, `timeoutMs`, and
85
+ `iterations`. A case-level value replaces the file-level value for that
86
+ datapoint.
87
+
88
+ ### Iterations
89
+
90
+ `iterations` (file or case, default `1`) runs a datapoint repeatedly.
91
+ Discovery expands `iterations: 3` on case `nyc` to runnable ids
92
+ `weather/nyc/1`, `weather/nyc/2`, `weather/nyc/3` (filter prefix
93
+ `weather/nyc` still selects all three). Each expanded case exposes
94
+ `t.iteration` / `t.iterations` on the test context. Cap is 100.
95
+
96
+ `maxConcurrency` counts **authored datapoints**, not expanded
97
+ iterations: siblings `…/1`…`…/n` share one concurrency slot and run
98
+ sequentially. A suite with 11 cases × 3 iterations and
99
+ `maxConcurrency: 20` therefore has at most 11 cases in flight, not 33.
100
+
101
+ ## Configure eval runs
102
+
103
+ Each project with evals needs `evals/evals.config.ts` or
104
+ `evals/evals.config.js`, and it must set `maxConcurrency`. Each case
105
+ issues real model-provider requests, so concurrency is capped hard at
106
+ 200. Existing projects use 20. Discovery with `eval --list` works
107
+ without this file, but running a case does not.
108
+
109
+ ```ts
110
+ import { defineEvalConfig } from "@cursor/july/evals";
111
+
112
+ export default defineEvalConfig({
113
+ maxConcurrency: 20, // required
114
+ // timeoutMs: 180_000, // optional project-wide default
115
+ // judge: { model: "..." }, // default judge model for t.judge.*
116
+ // reporters: [], // destinations that observe every case
117
+ // maxPlaygroundRuns: 50, // playground history only (default 20)
118
+ });
119
+ ```
120
+
121
+ The timeout order is case or file `timeoutMs`, CLI `--timeout-ms`,
122
+ project config `timeoutMs`, then the 180-second runner default.
123
+
124
+ The optional fields:
125
+
126
+ | Option | Default | Meaning |
127
+ | --- | --- | --- |
128
+ | `timeoutMs` | `180_000` | Project-wide per-case timeout |
129
+ | `judge` | unset | Default judge model for `t.judge.*`; see [Judge free-form output](#judge-free-form-output) |
130
+ | `reporters` | unset | Destinations that observe every case; `--skip-report` suppresses them |
131
+ | `maxPlaygroundRuns` | `20` | Max batches in the playground / `/v1/dev/evals*` history (not CLI `eval`). Hard-capped at 500. |
132
+
133
+ Reporters come from `@cursor/july/evals/reporters`: `JUnit` writes a
134
+ JUnit XML file for CI, `Artifacts` writes per-case files, and
135
+ `combineReporters` merges several into one (`renderJUnitXml` renders
136
+ the XML for a custom destination). A file or case can add its own
137
+ `reporters` on top of the config list.
138
+
139
+ Playground batches survive restarts whenever `agent/storage.ts` exists
140
+ with an `evals` table or a KV core providing `delete` and `list` (the
141
+ table is derived over the core); see
142
+ [Storage](/docs/storage.md#eval-and-a-b-tables). Without storage they live
143
+ in process memory and disappear when `serve` exits. Navigating away
144
+ and back still works while the process is up.
145
+
146
+ ## Drive and assert with `t`
147
+
148
+ `t` is both the driver and the assertion surface. You write ordinary
149
+ control flow, sending turns and asserting inline.
150
+
151
+ Drive the agent with `t.send(message, options?)`. It runs one turn and
152
+ waits for the session to park or fail. Multiple sends in one case share
153
+ the session, which is how you write multi-turn evals.
154
+
155
+ Each `t.send` resolves to a turn result with `message`, `sessionId`,
156
+ `events`, `toolCalls`, `ok`, and `index`. The turn carries the same
157
+ assertion vocabulary as `t`, scoped to that turn, so you can grade an
158
+ intermediate turn before the next send overwrites `t.reply`.
159
+ `turn.expectOk()` throws when the turn failed, for later steps that
160
+ depend on it.
161
+
162
+ Read the full case state with `t.reply` (the last assistant text),
163
+ `t.events` (session events captured so far), `t.turns` (settled
164
+ turns, oldest first), and `t.sessionId`. `t.signal` aborts when the
165
+ case hits its timeout; pass it to your own async work.
166
+
167
+ Assert with the gates:
168
+
169
+ | Gate | Checks |
170
+ | --- | --- |
171
+ | `t.succeeded()` | the run did not fail and is not parked on an unanswered approval |
172
+ | `t.parked()` | the run cleanly parked on an unanswered approval request |
173
+ | `t.messageIncludes(token)` | the joined assistant text matches a string or `RegExp` |
174
+ | `t.calledTool(name, matcher?)` | a matching call to `name` happened |
175
+ | `t.notCalledTool(name)` | no request for `name`, in any lifecycle state |
176
+ | `t.loadedSkill(name)` | the agent opened the skill's `SKILL.md` (read, grep, or shell `cat`) |
177
+ | `t.toolOrder(names)` | tool requests appear in this relative order (extra calls allowed) |
178
+ | `t.usedNoTools()` | no tool calls at all |
179
+ | `t.maxToolCalls(max)` | at most `max` tool calls |
180
+ | `t.noFailedActions()` | no tool call reported an error |
181
+ | `t.calledSubagent(name, matcher?)` | a matching subagent delegation happened |
182
+ | `t.taggedArtifact(kind?, predicate?)` | at least one [artifact](/docs/reference/artifacts.md) was tagged |
183
+ | `t.event(type, matcher?)` | at least one matching event of `type` occurred |
184
+ | `t.notEvent(type, matcher?)` | no matching event of `type` occurred |
185
+ | `t.eventOrder(matchers)` | matching event groups occur in this relative order |
186
+ | `t.eventsSatisfy(label, predicate)` | your predicate over the typed event stream |
187
+ | `t.check(value, expectation)` | any value, against a builder |
188
+ | `t.score(name, value)` | records a 0–1 score you computed; soft until you add a bar |
189
+ | `t.requireToolCall(name, matcher?)` | gates on a matching call and returns it, so later code can read its input and output |
190
+ | `t.requireInputRequest(filter?)` | gates on exactly one pending approval request and returns it |
191
+
192
+ Every gate returns a handle: `.soft()` demotes it to tracked-only,
193
+ `.atLeast(0.7)` adds a soft score bar, and `.gate(0.8)` promotes a
194
+ scored assertion into a hard gate.
195
+
196
+ With no matcher, `calledTool` is request-based: a requested call counts
197
+ even when its result has not arrived. Pass
198
+ `t.calledTool("inspect_pr", { status: "completed" })` to require the
199
+ call to return. `input`, `output`, and `count` matcher fields accept a
200
+ literal, a `RegExp`, or a predicate.
201
+
202
+ The expectation builders are `includes(string | RegExp)`,
203
+ `equals(value)`, `matches(schema)`, `similarity(expected)`, and
204
+ `satisfies(predicate, label)`. `includes` stringifies its input,
205
+ `equals` compares values deeply, `matches` validates against a Standard
206
+ Schema (or anything with `safeParse`, like Zod), `similarity` scores
207
+ normalized text similarity, and `satisfies` runs your predicate. The
208
+ plain function `normalizedSimilarity(actual, expected)` returns the
209
+ same 0–1 score for use with `t.score`.
210
+
211
+ A few more context members shape a case: `t.require(value, expectation)`
212
+ records a gate and stops the test body when it fails, without a
213
+ duplicate execution error. `t.skip(reason)` ends the case as skipped
214
+ (reported separately, never changes the exit code; call it before
215
+ sending messages). `t.metric(name, value)` records a structured score
216
+ for the playground case card. `t.log(message)` records a debug line for
217
+ the CLI and playground result.
218
+
219
+ Three `t.send` options apply on session create (first `t.send` only):
220
+
221
+ - `workspaceFiles`: `{ path: contents }`, seeded into the local session
222
+ workspace. Prefer this over machine-local paths.
223
+ - `workspaceDir`: absolute harness cwd (local runtime).
224
+ - `cloud`: per-session cloud options merged over the agent's static
225
+ `cloud` config (repos / env / …). Use a pinned `repos` override to
226
+ attach a fixture repo for cloud evals without putting it on the
227
+ agent's default `cloud.repos`. Cloud ignores `workspaceFiles` seeds.
228
+
229
+ ```ts
230
+ const toolResults = t.events.filter((e) => e.type === "action.result");
231
+ t.check(
232
+ toolResults.length,
233
+ satisfies((n) => (n as number) <= 4, "at most 4 tool calls")
234
+ );
235
+ ```
236
+
237
+ A case with no explicit gates falls back to whether at least one turn
238
+ completed successfully. Add `t.succeeded()` and behavior-specific gates
239
+ anyway. They make the contract visible during review.
240
+
241
+ ### Judge free-form output
242
+
243
+ When wording matters and no regex captures it, `t.judge` grades the
244
+ reply with an LLM. The built-in graders are `factuality(expected)`,
245
+ `summarizes(expected)`, `closedQA(criteria)`, and `sql(expected)`. Each
246
+ scores `t.reply` by default; pass `{ on }` to grade another value.
247
+
248
+ ```ts
249
+ t.judge.factuality("It is 54°F in NYC right now.").atLeast(0.7);
250
+ ```
251
+
252
+ Judge assertions are soft by default, so a judge never fails a build
253
+ until you give it a bar with `.atLeast(0.7)` or promote it with
254
+ `.gate(0.8)`. The judge model comes from `defineEvalConfig({ judge })`,
255
+ `defineEval({ judge })`, a case-level `judge`, or a per-call
256
+ `{ model }` override; the nearest one wins. For a domain-specific judge
257
+ whose verdict is not a single score, `t.judge.model(prompt)` sends a
258
+ raw prompt to the same model and returns the reply. You then record the
259
+ parsed result with `t.score` or `t.check`.
260
+
261
+ ## Run evals from the CLI
262
+
263
+ The `eval` command discovers, filters, and runs cases.
264
+
265
+ Run the CLI under Node 22.13 or newer. Do not use Bun. Its HTTP/2 client
266
+ breaks tool-result streams and causes eval turns to fail.
267
+
268
+ ```bash
269
+ agent-sdk eval --dir . --list # discover only
270
+ agent-sdk eval --dir . # run all
271
+ agent-sdk eval --dir . builds/checkout # one datapoint
272
+ agent-sdk eval --dir . builds search # several ids or prefixes
273
+ agent-sdk eval --dir . --tag smoke --tag pull-request # any matching tag
274
+ agent-sdk eval --dir . --json --no-stream # machine-readable results
275
+ agent-sdk eval --dir . --verbose # logs + reply snippets
276
+ ```
277
+
278
+ Id filters use OR semantics. Each filter selects an exact id and its
279
+ descendants. For example, `builds` selects `builds`,
280
+ `builds/checkout`, and every other case below that path. Repeated tags
281
+ also use OR semantics. When you provide both ids and tags, a case must
282
+ match both groups.
283
+
284
+ `eval` boots an ephemeral server on port 0 with a temp state root
285
+ outside the project, so cases don't inherit ambient monorepo rules and
286
+ don't write into the project state directory. Point `--url` at a running server to eval
287
+ a live agent instead:
288
+
289
+ ```bash
290
+ agent-sdk eval --dir . \
291
+ --url http://127.0.0.1:3000/weather-agent \
292
+ --bearer-token "$AGENT_TOKEN"
293
+ ```
294
+
295
+ The eval definitions still come from `--dir`; `--url` only changes the
296
+ agent that receives the turns. For a locally mounted multi-agent
297
+ directory, `--slug weather-agent` chooses the target. Use
298
+ `--state-root` to keep ephemeral session state at a chosen path,
299
+ `--timeout-ms` to override the project timeout, and `--no-stream` to
300
+ keep live progress off stderr. A TTY streams turn progress by default.
301
+ `--verbose` still writes `t.log` lines to stderr and adds reply snippets
302
+ to text results.
303
+
304
+ Model turns need a Cursor credential from `agent-sdk login` or
305
+ `CURSOR_API_KEY`.
306
+
307
+ See [CLI: eval](/docs/reference/cli.md#eval) for flags and exit codes.
308
+
309
+ ### JSON results
310
+
311
+ Use `--json --no-stream` in scripts and CI. The top-level result carries
312
+ the totals and one result per case:
313
+
314
+ ```json
315
+ {
316
+ "ok": true,
317
+ "passed": 1,
318
+ "failed": 0,
319
+ "results": [
320
+ {
321
+ "id": "readiness",
322
+ "ok": true,
323
+ "assertions": [{ "name": "succeeded", "passed": true }],
324
+ "sessionId": "ses_123",
325
+ "inputs": ["Is checkout pull request 42 ready to approve?"],
326
+ "toolCalls": [{ "toolName": "inspect_pr", "isError": false }],
327
+ "logs": [],
328
+ "durationMs": 12340
329
+ }
330
+ ]
331
+ }
332
+ ```
333
+
334
+ Each case result can also include `description`, `finalText`, `tools`,
335
+ `error`, and tool arguments or output. This shape lets CI report the
336
+ failed assertion without parsing terminal text.
337
+
338
+ ## Run evals in the playground
339
+
340
+ Start the server, open the playground, and choose **Evals**. You can run
341
+ every case or one case, watch progress, and open the resulting session
342
+ trace. The Evals tab works on a normal `serve`.
343
+
344
+ ```bash
345
+ agent-sdk serve --dir .
346
+ ```
347
+
348
+ Playground runs target the live server instead of an ephemeral one.
349
+ Their sessions appear in the session list. One eval batch can run at a
350
+ time. Persistence follows the rule under
351
+ [Configure eval runs](#configure-eval-runs). See
352
+ [Playground eval routes](/docs/reference/http-api.md#playground-eval-routes).
353
+ The start request returns `202` while cases run in the background.
354
+ Poll until the snapshot status becomes `completed`, `failed`, or `cancelled`.
355
+ Configuration errors appear on a failed snapshot.
356
+
357
+ On `--prod` / `--url`, the CLI prints the Eval ID as soon as the batch is
358
+ accepted (and a Playground deep link with `?view=evals&evalRunId=…`):
359
+
360
+ ```bash
361
+ agent-sdk eval --prod --slug vulnerability-scanner --tag deepsec
362
+ # Eval ID: evalrun_…
363
+ # Cancel: agent-sdk eval cancel evalrun_… --prod --slug vulnerability-scanner
364
+ # Playground: https://…/playground?view=evals&evalRunId=evalrun_…
365
+
366
+ agent-sdk eval cancel evalrun_… --prod --slug vulnerability-scanner
367
+ agent-sdk eval status evalrun_… --prod --slug vulnerability-scanner
368
+ ```
369
+
370
+ ## What good cases assert
371
+
372
+ Gate decisions and shape, not prose. Model wording varies run to run.
373
+ Tool choice, tool avoidance, and output structure are the stable
374
+ contract.
375
+
376
+ 1. `t.succeeded()`: always, first.
377
+ 2. The tool decision: `calledTool` for the intended path,
378
+ `notCalledTool` for the likely wrong alternative. The pair is
379
+ stronger than either alone.
380
+ 3. Output shape: a regex for the contract (`/ready|blocked/i`, a JSON
381
+ marker, a findings-block fence), never exact sentences.
382
+ 4. For structured output, parse `t.reply` and check fields with
383
+ `satisfies` instead of substring-matching JSON.
384
+
385
+ The common failure modes: asserting exact phrasing, packing more than
386
+ about five gates into one case (split it), and cases that depend on live
387
+ external state that drifts (pin the input; see fixtures).
388
+
389
+ ## Pick fixtures by agent type
390
+
391
+ The right fixture depends on the surface under test.
392
+
393
+ | Agent surface | Fixture |
394
+ | --- | --- |
395
+ | Chat / domain assistant | A canonical prompt string, chosen once and frozen |
396
+ | Tool-heavy | Run `agent-sdk call <tool>` first to pin what the tool returns, then freeze the prompt that triggers it |
397
+ | GitHub webhook | `agent-sdk github replay <pr> --events '*' --dry-run --out fixtures/github` snapshots real payloads for offline replay ([GitHub guide](/docs/guides/github.md)) |
398
+ | PR reviewer with host preparation | Diff, metadata, and gold labels pinned to commit SHAs; keep any live PR matrix small |
399
+ | Workspace-dependent | `workspaceFiles` in `t.send` options, never developer-machine paths |
400
+
401
+ Tag the fast, reliably passing core `smoke` and run `--tag smoke` in the
402
+ inner loop. Leave slow or flaky-prone cases untagged for explicit runs.
403
+
404
+ ### Materialize API-backed fixtures
405
+
406
+ An input that only points at external data, such as a pull request URL,
407
+ snapshot id, or pair of commit SHAs, is not self-contained. Fetch it
408
+ once and commit the rendered fixture before you expand the suite.
409
+
410
+ 1. Save the diff, metadata, and labels under `fixtures/` at pinned
411
+ revisions.
412
+ 2. Seed those files with `workspaceFiles`, or read them from the fixture
413
+ directory.
414
+ 3. Assert decisions and output shape against the saved evidence.
415
+ 4. Keep a small `smoke` subset for any remaining live pipeline checks.
416
+
417
+ Read committed fixtures with `@cursor/july/evals/loaders`: `loadJson`,
418
+ `loadJsonl`, and `loadYaml` resolve relative paths against the project
419
+ root the runner discovered, not the cwd the CLI was invoked from
420
+ (`resolveFixturePath` and `evalFixtureRoot` expose the same
421
+ resolution for other file formats).
422
+
423
+ `maxConcurrency` limits parallel datapoints. It does not limit model or
424
+ API fan-out inside one datapoint. Materialized fixtures prevent a large
425
+ suite from exhausting provider and GitHub rate limits. The
426
+ [evals skill](https://github.com/cursor/cursor/blob/main/packages/agent-serve/skills/evals/SKILL.md) has the full fixture workflow.
427
+
428
+ ## Keep improvements with regression evals
429
+
430
+ Every [hillclimb](/docs/hillclimbing.md) round that keeps a change must land
431
+ an eval that would have failed before the change. If you can't express
432
+ the improvement as a gate (a `calledTool` shift, a bounded
433
+ `action.result` count, an output-shape regex), the improvement is
434
+ unverified, and it'll regress silently.
435
+
436
+ The rule cuts the other way too: never weaken an existing gate to make a
437
+ round pass. That's the freeze line moving, and it turns your regression
438
+ suite into a list of checks that no longer protect anything.
439
+
440
+ ## Compare variants on live traffic
441
+
442
+ Use `defineAB` to compare variant metrics on live sessions. It is not a
443
+ test runner and has no `agent-sdk ab` command. Keep `defineEval` as the
444
+ regression ratchet. Eval sessions do not enroll or change live metrics.
445
+ See [Live A/B metrics](/docs/ab.md) for assignment, behavior, collection,
446
+ and inspection.
447
+
448
+ ## What's next
449
+
450
+ Continue with these pages:
451
+
452
+ - [Live A/B metrics](/docs/ab.md): sticky variants and cumulative metrics
453
+ on live sessions
454
+ - [Hillclimbing](/docs/hillclimbing.md): the loop evals make trustworthy
455
+ - [Building agents with agents](/docs/building-with-agents.md): have a
456
+ coding agent write the first suite
457
+ - [GitHub guide](/docs/guides/github.md): deterministic webhook fixtures
458
+ with `github replay`
459
+ - [Sessions and streaming](/docs/reference/sessions.md): the events
460
+ `t.events` contains