@cursor/july 0.1.111 → 0.1.113

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (441) hide show
  1. package/README.md +4 -0
  2. package/dist/bin/agent-serve.js +20 -4
  3. package/dist/channels/checks.d.ts +10 -0
  4. package/dist/channels/checks.d.ts.map +1 -1
  5. package/dist/channels/github/github-channel.d.ts +29 -19
  6. package/dist/channels/github/github-channel.d.ts.map +1 -1
  7. package/dist/channels/github/github-channel.js +29 -19
  8. package/dist/channels/origin/checks.d.ts +1 -1
  9. package/dist/channels/origin/checks.d.ts.map +1 -1
  10. package/dist/channels/slack/agentic-delivery.d.ts.map +1 -1
  11. package/dist/channels/slack/agentic-delivery.js +2 -0
  12. package/dist/channels/slack/dispatch.d.ts +11 -1
  13. package/dist/channels/slack/dispatch.d.ts.map +1 -1
  14. package/dist/channels/slack/dispatch.js +146 -39
  15. package/dist/channels/slack/placeholder.d.ts +26 -0
  16. package/dist/channels/slack/placeholder.d.ts.map +1 -1
  17. package/dist/channels/slack/placeholder.js +28 -0
  18. package/dist/channels/slack/slack-channel.d.ts.map +1 -1
  19. package/dist/channels/slack/slack-channel.js +9 -16
  20. package/dist/docs/404.html +2 -2
  21. package/dist/docs/assets/{app.Drol6mi6.js → app.CAeK13eM.js} +4 -4
  22. package/dist/docs/assets/building-with-agents.md.BBCx0AUo.js +9 -0
  23. package/dist/docs/assets/building-with-agents.md.BBCx0AUo.lean.js +1 -0
  24. package/dist/docs/assets/chunks/@localSearchIndexroot.Ck9E52Ls.js +1 -0
  25. package/dist/docs/assets/chunks/{VPLocalSearchBox.CdqoBZJX.js → VPLocalSearchBox.C9LbPHod.js} +1 -1
  26. package/dist/docs/assets/chunks/{arc.BVX3ycTn.js → arc.CmMq2zmS.js} +1 -1
  27. package/dist/docs/assets/chunks/{architectureDiagram-Q4EWVU46.CcChnMxX.js → architectureDiagram-Q4EWVU46.CCXB8Uj5.js} +1 -1
  28. package/dist/docs/assets/chunks/{baseUniq.xwtXO-yt.js → baseUniq.CyQo6eLe.js} +1 -1
  29. package/dist/docs/assets/chunks/{blockDiagram-DXYQGD6D.CijZ_taK.js → blockDiagram-DXYQGD6D.JYq6w91N.js} +1 -1
  30. package/dist/docs/assets/chunks/{c4Diagram-AHTNJAMY.0FBvwBvK.js → c4Diagram-AHTNJAMY.BRV8GPJJ.js} +1 -1
  31. package/dist/docs/assets/chunks/channel.BHiYmnZ4.js +1 -0
  32. package/dist/docs/assets/chunks/{chunk-4BX2VUAB.CELMmMDA.js → chunk-4BX2VUAB.Bv4ooYQR.js} +1 -1
  33. package/dist/docs/assets/chunks/{chunk-4TB4RGXK.C4bhFtSm.js → chunk-4TB4RGXK.t4JtKPcj.js} +1 -1
  34. package/dist/docs/assets/chunks/{chunk-55IACEB6.ZnQ9gRRQ.js → chunk-55IACEB6.34lCHj9Y.js} +1 -1
  35. package/dist/docs/assets/chunks/{chunk-EDXVE4YY.DPlmglG-.js → chunk-EDXVE4YY.BSwrPNrt.js} +1 -1
  36. package/dist/docs/assets/chunks/{chunk-FMBD7UC4.BrUk8pff.js → chunk-FMBD7UC4.Beeun-R-.js} +1 -1
  37. package/dist/docs/assets/chunks/{chunk-OYMX7WX6.CcmWIncu.js → chunk-OYMX7WX6.BUUFUcJc.js} +1 -1
  38. package/dist/docs/assets/chunks/{chunk-QZHKN3VN.D65-cs8I.js → chunk-QZHKN3VN.B2XjHzN_.js} +1 -1
  39. package/dist/docs/assets/chunks/{chunk-YZCP3GAM.qoXZpG9F.js → chunk-YZCP3GAM.CLYG8znk.js} +1 -1
  40. package/dist/docs/assets/chunks/classDiagram-6PBFFD2Q.Degh8l90.js +1 -0
  41. package/dist/docs/assets/chunks/classDiagram-v2-HSJHXN6E.Degh8l90.js +1 -0
  42. package/dist/docs/assets/chunks/clone.BIywbczV.js +1 -0
  43. package/dist/docs/assets/chunks/{cose-bilkent-S5V4N54A.8rYtqudO.js → cose-bilkent-S5V4N54A.DVEa6fZp.js} +1 -1
  44. package/dist/docs/assets/chunks/{dagre-KV5264BT.DrRP1fOh.js → dagre-KV5264BT.C9PZQK-S.js} +1 -1
  45. package/dist/docs/assets/chunks/{diagram-5BDNPKRD.DHj_xA2_.js → diagram-5BDNPKRD.DoN0uv3Y.js} +1 -1
  46. package/dist/docs/assets/chunks/{diagram-G4DWMVQ6.Bz_6nAKj.js → diagram-G4DWMVQ6.Czv3duqx.js} +1 -1
  47. package/dist/docs/assets/chunks/{diagram-MMDJMWI5.BWA0xSW9.js → diagram-MMDJMWI5.BinJ5kWb.js} +1 -1
  48. package/dist/docs/assets/chunks/{diagram-TYMM5635.CpTJLNJI.js → diagram-TYMM5635.DW326M4K.js} +1 -1
  49. package/dist/docs/assets/chunks/{erDiagram-SMLLAGMA.-7AWWSrP.js → erDiagram-SMLLAGMA.U2pR_OA7.js} +1 -1
  50. package/dist/docs/assets/chunks/{flowDiagram-DWJPFMVM.BTnQ742_.js → flowDiagram-DWJPFMVM.ByWJXeYK.js} +1 -1
  51. package/dist/docs/assets/chunks/framework.BNw1pucY.js +19 -0
  52. package/dist/docs/assets/chunks/{ganttDiagram-T4ZO3ILL.B5_HiiQ5.js → ganttDiagram-T4ZO3ILL.OquF0Rtg.js} +1 -1
  53. package/dist/docs/assets/chunks/{gitGraphDiagram-UUTBAWPF.CZcNTFZd.js → gitGraphDiagram-UUTBAWPF.Bpn01P7X.js} +1 -1
  54. package/dist/docs/assets/chunks/{graph.V2GLaab4.js → graph.CNRB6ETL.js} +1 -1
  55. package/dist/docs/assets/chunks/{infoDiagram-42DDH7IO.7pZOkCCU.js → infoDiagram-42DDH7IO.CqhknMWi.js} +1 -1
  56. package/dist/docs/assets/chunks/{ishikawaDiagram-UXIWVN3A.DSMo3Qa3.js → ishikawaDiagram-UXIWVN3A.C6xpR2af.js} +1 -1
  57. package/dist/docs/assets/chunks/{journeyDiagram-VCZTEJTY.BNEgWN1S.js → journeyDiagram-VCZTEJTY.Cg5f7oB3.js} +1 -1
  58. package/dist/docs/assets/chunks/{kanban-definition-6JOO6SKY.B-4d4tC7.js → kanban-definition-6JOO6SKY.Cx9YTwlU.js} +1 -1
  59. package/dist/docs/assets/chunks/{layout.Dvbn9nSb.js → layout.ljS-wFtK.js} +1 -1
  60. package/dist/docs/assets/chunks/{linear.D2GM4p4b.js → linear.jSxNrsFC.js} +1 -1
  61. package/dist/docs/assets/chunks/{min.DVBtLK-B.js → min.Cum8AlQw.js} +1 -1
  62. package/dist/docs/assets/chunks/{mindmap-definition-QFDTVHPH.CnJRI55x.js → mindmap-definition-QFDTVHPH.BLiysLpe.js} +1 -1
  63. package/dist/docs/assets/chunks/{pieDiagram-DEJITSTG.CtoaTFlA.js → pieDiagram-DEJITSTG.BoIDyuKF.js} +1 -1
  64. package/dist/docs/assets/chunks/{quadrantDiagram-34T5L4WZ.DtT24_vN.js → quadrantDiagram-34T5L4WZ.DLkpDytR.js} +1 -1
  65. package/dist/docs/assets/chunks/{requirementDiagram-MS252O5E.CjSO4o8f.js → requirementDiagram-MS252O5E.DqTVqSu2.js} +1 -1
  66. package/dist/docs/assets/chunks/{sankeyDiagram-XADWPNL6.-wWiIVWa.js → sankeyDiagram-XADWPNL6.CG_6FF7j.js} +1 -1
  67. package/dist/docs/assets/chunks/{sequenceDiagram-FGHM5R23.DSk8s4gX.js → sequenceDiagram-FGHM5R23.BIp9602K.js} +1 -1
  68. package/dist/docs/assets/chunks/{stateDiagram-FHFEXIEX.BFKAsdkN.js → stateDiagram-FHFEXIEX.COSXsD9I.js} +1 -1
  69. package/dist/docs/assets/chunks/stateDiagram-v2-QKLJ7IA2.qrxrbFsX.js +1 -0
  70. package/dist/docs/assets/chunks/{theme.C0MctGaz.js → theme.CXJ7PNwy.js} +2 -2
  71. package/dist/docs/assets/chunks/{timeline-definition-GMOUNBTQ.TUNJbAFe.js → timeline-definition-GMOUNBTQ.CXdVqkLq.js} +1 -1
  72. package/dist/docs/assets/chunks/{vennDiagram-DHZGUBPP.C-RGNnk4.js → vennDiagram-DHZGUBPP.CZxGuc4r.js} +1 -1
  73. package/dist/docs/assets/chunks/{wardley-RL74JXVD.g5efOWmT.js → wardley-RL74JXVD.3oVgfqQk.js} +1 -1
  74. package/dist/docs/assets/chunks/{wardleyDiagram-NUSXRM2D.DH4zt73Y.js → wardleyDiagram-NUSXRM2D.6_irCgGJ.js} +1 -1
  75. package/dist/docs/assets/chunks/{xychartDiagram-5P7HB3ND.B-3k6bF6.js → xychartDiagram-5P7HB3ND.TRPe92m3.js} +1 -1
  76. package/dist/docs/assets/deployment.md.D2jQZuFx.js +32 -0
  77. package/dist/docs/assets/deployment.md.D2jQZuFx.lean.js +1 -0
  78. package/dist/docs/assets/evals.md.D3Y3Aixt.js +72 -0
  79. package/dist/docs/assets/evals.md.D3Y3Aixt.lean.js +1 -0
  80. package/dist/docs/assets/guides_agent-to-agent.md.CD4T5FIl.js +41 -0
  81. package/dist/docs/assets/guides_agent-to-agent.md.CD4T5FIl.lean.js +1 -0
  82. package/dist/docs/assets/guides_bitbucket.md.mpevW-VP.js +145 -0
  83. package/dist/docs/assets/guides_bitbucket.md.mpevW-VP.lean.js +1 -0
  84. package/dist/docs/assets/guides_cloud-agents.md.Cp1O3u-X.js +15 -0
  85. package/dist/docs/assets/guides_cloud-agents.md.Cp1O3u-X.lean.js +1 -0
  86. package/dist/docs/assets/guides_convert-automation.md.CqEyfP6Y.js +43 -0
  87. package/dist/docs/assets/guides_convert-automation.md.CqEyfP6Y.lean.js +1 -0
  88. package/dist/docs/assets/guides_github.md.BwpBp3ed.js +156 -0
  89. package/dist/docs/assets/guides_github.md.BwpBp3ed.lean.js +1 -0
  90. package/dist/docs/assets/guides_gitlab.md.DaEC3nMk.js +153 -0
  91. package/dist/docs/assets/guides_gitlab.md.DaEC3nMk.lean.js +1 -0
  92. package/dist/docs/assets/guides_grokbot-agents.md.CMhZNdEU.js +16 -0
  93. package/dist/docs/assets/guides_grokbot-agents.md.CMhZNdEU.lean.js +1 -0
  94. package/dist/docs/assets/guides_improve.md.Bnp4F99w.js +22 -0
  95. package/dist/docs/assets/guides_improve.md.Bnp4F99w.lean.js +1 -0
  96. package/dist/docs/assets/guides_jev.md.F5fAkkfN.js +189 -0
  97. package/dist/docs/assets/guides_jev.md.F5fAkkfN.lean.js +1 -0
  98. package/dist/docs/assets/guides_mcp-oauth.md.bSFakfCY.js +50 -0
  99. package/dist/docs/assets/guides_mcp-oauth.md.bSFakfCY.lean.js +1 -0
  100. package/dist/docs/assets/guides_opentelemetry.md.BKDxQmmd.js +35 -0
  101. package/dist/docs/assets/guides_opentelemetry.md.BKDxQmmd.lean.js +1 -0
  102. package/dist/docs/assets/guides_slack.md.Bo96y42E.js +70 -0
  103. package/dist/docs/assets/guides_slack.md.Bo96y42E.lean.js +1 -0
  104. package/dist/docs/assets/guides_webhooks.md.1A72_VEE.js +92 -0
  105. package/dist/docs/assets/guides_webhooks.md.1A72_VEE.lean.js +1 -0
  106. package/dist/docs/assets/hillclimbing.md.D4E1o5Sa.js +7 -0
  107. package/dist/docs/assets/hillclimbing.md.D4E1o5Sa.lean.js +1 -0
  108. package/dist/docs/assets/{index.md.BFVyY2KT.js → index.md.C-t81M5J.js} +2 -2
  109. package/dist/docs/assets/{index.md.BFVyY2KT.lean.js → index.md.C-t81M5J.lean.js} +1 -1
  110. package/dist/docs/assets/{quickstart.md.D3MjSZN-.js → quickstart.md.DAvVhuuU.js} +1 -1
  111. package/dist/docs/assets/{quickstart.md.D3MjSZN-.lean.js → quickstart.md.DAvVhuuU.lean.js} +1 -1
  112. package/dist/docs/assets/{reference_agent-config.md.CfVA-LZJ.js → reference_agent-config.md.DGPyw7ms.js} +1 -1
  113. package/dist/docs/assets/{reference_agent-config.md.CfVA-LZJ.lean.js → reference_agent-config.md.DGPyw7ms.lean.js} +1 -1
  114. package/dist/docs/assets/{reference_artifacts.md.Vf7qyIZ-.js → reference_artifacts.md.Bu_4HmsD.js} +1 -1
  115. package/dist/docs/assets/{reference_artifacts.md.Vf7qyIZ-.lean.js → reference_artifacts.md.Bu_4HmsD.lean.js} +1 -1
  116. package/dist/docs/assets/{reference_channels.md.icqLKcTc.js → reference_channels.md.nFWbzAic.js} +1 -1
  117. package/dist/docs/assets/{reference_channels.md.icqLKcTc.lean.js → reference_channels.md.nFWbzAic.lean.js} +1 -1
  118. package/dist/docs/assets/{reference_cli.md.B2dBL6L8.js → reference_cli.md.DLWDz9ij.js} +3 -1
  119. package/dist/docs/assets/{reference_cli.md.B2dBL6L8.lean.js → reference_cli.md.DLWDz9ij.lean.js} +1 -1
  120. package/dist/docs/assets/{reference_connections.md.Cb3U_c8n.js → reference_connections.md.Je9dMsdd.js} +2 -2
  121. package/dist/docs/assets/{reference_connections.md.Cb3U_c8n.lean.js → reference_connections.md.Je9dMsdd.lean.js} +1 -1
  122. package/dist/docs/assets/reference_evals.md.DNJzM_yf.js +57 -0
  123. package/dist/docs/assets/reference_evals.md.DNJzM_yf.lean.js +1 -0
  124. package/dist/docs/assets/reference_extensions.md.Cv5aLCz_.js +58 -0
  125. package/dist/docs/assets/reference_extensions.md.Cv5aLCz_.lean.js +1 -0
  126. package/dist/docs/assets/{reference_hooks.md.CZuynAxj.js → reference_hooks.md.B7uzNENk.js} +2 -2
  127. package/dist/docs/assets/{reference_hooks.md.CZuynAxj.lean.js → reference_hooks.md.B7uzNENk.lean.js} +1 -1
  128. package/dist/docs/assets/{reference_http-api.md.DTKcYE6L.js → reference_http-api.md.CduHavZ2.js} +1 -1
  129. package/dist/docs/assets/{reference_http-api.md.DTKcYE6L.lean.js → reference_http-api.md.CduHavZ2.lean.js} +1 -1
  130. package/dist/docs/assets/{reference_instructions.md.D7gkckK-.js → reference_instructions.md.CU1My5My.js} +1 -1
  131. package/dist/docs/assets/{reference_instructions.md.D7gkckK-.lean.js → reference_instructions.md.CU1My5My.lean.js} +1 -1
  132. package/dist/docs/assets/{reference_playground.md.D2YExv5K.js → reference_playground.md.Ch2d0Iqi.js} +1 -1
  133. package/dist/docs/assets/{reference_playground.md.D2YExv5K.lean.js → reference_playground.md.Ch2d0Iqi.lean.js} +1 -1
  134. package/dist/docs/assets/{reference_project-layout.md.DPxbUJyt.js → reference_project-layout.md.BGhgpy9V.js} +1 -1
  135. package/dist/docs/assets/{reference_project-layout.md.DPxbUJyt.lean.js → reference_project-layout.md.BGhgpy9V.lean.js} +1 -1
  136. package/dist/docs/assets/{reference_prompt.md.BQ5uAv1F.js → reference_prompt.md.Ccp0R53H.js} +1 -1
  137. package/dist/docs/assets/{reference_prompt.md.BQ5uAv1F.lean.js → reference_prompt.md.Ccp0R53H.lean.js} +1 -1
  138. package/dist/docs/assets/{reference_schedules.md.BasfZWO-.js → reference_schedules.md.B2Nm6FaD.js} +1 -1
  139. package/dist/docs/assets/{reference_schedules.md.BasfZWO-.lean.js → reference_schedules.md.B2Nm6FaD.lean.js} +1 -1
  140. package/dist/docs/assets/{reference_sessions.md.YKvIsWAx.js → reference_sessions.md.1_6Vyv7x.js} +1 -1
  141. package/dist/docs/assets/{reference_sessions.md.YKvIsWAx.lean.js → reference_sessions.md.1_6Vyv7x.lean.js} +1 -1
  142. package/dist/docs/assets/{reference_skills.md.rNgpsGd0.js → reference_skills.md.DjQkRefx.js} +1 -1
  143. package/dist/docs/assets/{reference_skills.md.rNgpsGd0.lean.js → reference_skills.md.DjQkRefx.lean.js} +1 -1
  144. package/dist/docs/assets/{reference_subagents.md.e5qitjJt.js → reference_subagents.md.Dl16gcBj.js} +2 -2
  145. package/dist/docs/assets/{reference_subagents.md.e5qitjJt.lean.js → reference_subagents.md.Dl16gcBj.lean.js} +1 -1
  146. package/dist/docs/assets/{reference_tools.md.BdCO2aHZ.js → reference_tools.md.B1dH1lpa.js} +3 -3
  147. package/dist/docs/assets/{reference_tools.md.BdCO2aHZ.lean.js → reference_tools.md.B1dH1lpa.lean.js} +1 -1
  148. package/dist/docs/assets/{templates_agentic-owners.md.Da_AGDlH.js → templates_agentic-owners.md.9M575F5C.js} +1 -1
  149. package/dist/docs/assets/{templates_agentic-owners.md.Da_AGDlH.lean.js → templates_agentic-owners.md.9M575F5C.lean.js} +1 -1
  150. package/dist/docs/assets/{templates_pr-autofixer.md.DqxocIGh.js → templates_pr-autofixer.md.ws0DDXDy.js} +1 -1
  151. package/dist/docs/assets/{templates_pr-autofixer.md.DqxocIGh.lean.js → templates_pr-autofixer.md.ws0DDXDy.lean.js} +1 -1
  152. package/dist/docs/assets/{templates_security-reviewer.md.Bhnvd8VE.js → templates_security-reviewer.md.KEFYXzfK.js} +1 -1
  153. package/dist/docs/assets/{templates_security-reviewer.md.Bhnvd8VE.lean.js → templates_security-reviewer.md.KEFYXzfK.lean.js} +1 -1
  154. package/dist/docs/assets/templates_thermo-quality-review.md.VNJ_mohX.js +3 -0
  155. package/dist/docs/assets/templates_thermo-quality-review.md.VNJ_mohX.lean.js +1 -0
  156. package/dist/docs/assets/templates_thermo-review.md.Hi3zWOkP.js +3 -0
  157. package/dist/docs/assets/templates_thermo-review.md.Hi3zWOkP.lean.js +1 -0
  158. package/dist/docs/assets/{templates_triage.md.BdBWO9Ic.js → templates_triage.md.CVGe_FG6.js} +1 -1
  159. package/dist/docs/assets/{templates_triage.md.BdBWO9Ic.lean.js → templates_triage.md.CVGe_FG6.lean.js} +1 -1
  160. package/dist/docs/assets/troubleshooting.md.mnfFG2Em.js +1 -0
  161. package/dist/docs/assets/troubleshooting.md.mnfFG2Em.lean.js +1 -0
  162. package/dist/docs/building-with-agents.html +44 -48
  163. package/dist/docs/building-with-agents.md +94 -82
  164. package/dist/docs/deployment.html +59 -78
  165. package/dist/docs/deployment.md +117 -363
  166. package/dist/docs/evals.html +85 -224
  167. package/dist/docs/evals.md +149 -673
  168. package/dist/docs/guides/agent-to-agent.html +74 -47
  169. package/dist/docs/guides/agent-to-agent.md +128 -46
  170. package/dist/docs/guides/bitbucket.html +179 -44
  171. package/dist/docs/guides/bitbucket.md +249 -48
  172. package/dist/docs/guides/cloud-agents.html +46 -40
  173. package/dist/docs/guides/cloud-agents.md +119 -66
  174. package/dist/docs/guides/convert-automation.html +80 -49
  175. package/dist/docs/guides/convert-automation.md +156 -147
  176. package/dist/docs/guides/github.html +184 -92
  177. package/dist/docs/guides/github.md +260 -245
  178. package/dist/docs/guides/gitlab.html +184 -45
  179. package/dist/docs/guides/gitlab.md +249 -50
  180. package/dist/docs/guides/grokbot-agents.html +48 -41
  181. package/dist/docs/guides/grokbot-agents.md +99 -53
  182. package/dist/docs/guides/improve.html +48 -40
  183. package/dist/docs/guides/improve.md +111 -58
  184. package/dist/docs/guides/jev.html +248 -0
  185. package/dist/docs/guides/jev.md +348 -0
  186. package/dist/docs/guides/mcp-oauth.html +74 -52
  187. package/dist/docs/guides/mcp-oauth.md +111 -121
  188. package/dist/docs/guides/opentelemetry.html +63 -54
  189. package/dist/docs/guides/opentelemetry.md +96 -165
  190. package/dist/docs/guides/slack.html +83 -59
  191. package/dist/docs/guides/slack.md +157 -227
  192. package/dist/docs/guides/webhooks.html +97 -230
  193. package/dist/docs/guides/webhooks.md +154 -385
  194. package/dist/docs/hashmap.json +1 -1
  195. package/dist/docs/hillclimbing.html +44 -38
  196. package/dist/docs/hillclimbing.md +101 -55
  197. package/dist/docs/index.html +38 -38
  198. package/dist/docs/index.md +2 -0
  199. package/dist/docs/llms-full.txt +3490 -3146
  200. package/dist/docs/llms.txt +22 -18
  201. package/dist/docs/quickstart.html +37 -37
  202. package/dist/docs/reference/agent-config.html +37 -37
  203. package/dist/docs/reference/artifacts.html +38 -38
  204. package/dist/docs/reference/channels.html +37 -37
  205. package/dist/docs/reference/cli.html +39 -37
  206. package/dist/docs/reference/cli.md +2 -0
  207. package/dist/docs/reference/connections.html +38 -38
  208. package/dist/docs/reference/connections.md +1 -1
  209. package/dist/docs/reference/evals.html +116 -0
  210. package/dist/docs/reference/evals.md +293 -0
  211. package/dist/docs/reference/extensions.html +80 -84
  212. package/dist/docs/reference/extensions.md +131 -202
  213. package/dist/docs/reference/hooks.html +39 -39
  214. package/dist/docs/reference/hooks.md +34 -34
  215. package/dist/docs/reference/http-api.html +37 -37
  216. package/dist/docs/reference/instructions.html +37 -37
  217. package/dist/docs/reference/playground.html +37 -37
  218. package/dist/docs/reference/project-layout.html +37 -37
  219. package/dist/docs/reference/prompt.html +37 -37
  220. package/dist/docs/reference/schedules.html +37 -37
  221. package/dist/docs/reference/sessions.html +37 -37
  222. package/dist/docs/reference/skills.html +37 -37
  223. package/dist/docs/reference/subagents.html +38 -38
  224. package/dist/docs/reference/subagents.md +1 -1
  225. package/dist/docs/reference/tools.html +38 -38
  226. package/dist/docs/reference/tools.md +4 -0
  227. package/dist/docs/templates/agentic-owners.html +38 -38
  228. package/dist/docs/templates/pr-autofixer.html +37 -37
  229. package/dist/docs/templates/security-reviewer.html +38 -38
  230. package/dist/docs/templates/thermo-quality-review.html +62 -0
  231. package/dist/docs/templates/thermo-quality-review.md +75 -0
  232. package/dist/docs/templates/thermo-review.html +62 -0
  233. package/dist/docs/templates/thermo-review.md +74 -0
  234. package/dist/docs/templates/triage.html +37 -37
  235. package/dist/docs/troubleshooting.html +38 -38
  236. package/dist/docs/troubleshooting.md +97 -66
  237. package/dist/extensions/cursor-cloud-agents/skills/handoff.md +3 -7
  238. package/dist/extensions/improve/extension.d.ts +3 -1
  239. package/dist/extensions/improve/extension.d.ts.map +1 -1
  240. package/dist/extensions/improve/extension.js +4 -2
  241. package/dist/extensions/improve/skills/yourself.js +1 -1
  242. package/dist/extensions/jev/extension.d.ts +43 -0
  243. package/dist/extensions/jev/extension.d.ts.map +1 -0
  244. package/dist/extensions/jev/extension.js +47 -0
  245. package/dist/extensions/jev/lib/evaluate.d.ts +101 -0
  246. package/dist/extensions/jev/lib/evaluate.d.ts.map +1 -0
  247. package/dist/extensions/jev/lib/evaluate.js +167 -0
  248. package/dist/extensions/jev/skills/gated-write.md +25 -0
  249. package/dist/extensions/jev/skills/questions.md +33 -0
  250. package/dist/extensions/jev/tools/evaluate.d.ts +4 -0
  251. package/dist/extensions/jev/tools/evaluate.d.ts.map +1 -0
  252. package/dist/extensions/jev/tools/evaluate.js +88 -0
  253. package/dist/extensions.d.ts +1 -1
  254. package/dist/extensions.d.ts.map +1 -1
  255. package/dist/extensions.js +2 -0
  256. package/dist/filesystem.d.ts +46 -2
  257. package/dist/filesystem.d.ts.map +1 -1
  258. package/dist/filesystem.js +149 -102
  259. package/dist/index.d.ts +2 -2
  260. package/dist/index.d.ts.map +1 -1
  261. package/dist/index.js +1 -1
  262. package/dist/internal/advertise-tools.d.ts.map +1 -1
  263. package/dist/internal/advertise-tools.js +6 -0
  264. package/dist/internal/cli-ax.d.ts +6 -0
  265. package/dist/internal/cli-ax.d.ts.map +1 -1
  266. package/dist/internal/cli-ax.js +72 -14
  267. package/dist/internal/continuation-identity.d.ts.map +1 -1
  268. package/dist/internal/continuation-identity.js +1 -0
  269. package/dist/internal/cursor-agent-template.d.ts +1 -1
  270. package/dist/internal/cursor-agent-template.d.ts.map +1 -1
  271. package/dist/internal/cursor-agent-template.js +2 -0
  272. package/dist/internal/discovery/connections.d.ts.map +1 -1
  273. package/dist/internal/discovery/connections.js +18 -0
  274. package/dist/internal/discovery/extensions.d.ts.map +1 -1
  275. package/dist/internal/discovery/extensions.js +8 -4
  276. package/dist/internal/discovery/info.d.ts.map +1 -1
  277. package/dist/internal/discovery/info.js +1 -0
  278. package/dist/internal/filesystem/tools.d.ts.map +1 -1
  279. package/dist/internal/filesystem/tools.js +2 -2
  280. package/dist/internal/filesystem/walk.d.ts +7 -4
  281. package/dist/internal/filesystem/walk.d.ts.map +1 -1
  282. package/dist/internal/filesystem/walk.js +34 -11
  283. package/dist/internal/hosted-admission-adapter.d.ts +3 -0
  284. package/dist/internal/hosted-admission-adapter.d.ts.map +1 -1
  285. package/dist/internal/hosted-delivery-protocol.d.ts +20 -0
  286. package/dist/internal/hosted-delivery-protocol.d.ts.map +1 -1
  287. package/dist/internal/hosted-delivery-protocol.js +51 -1
  288. package/dist/internal/hosted-delivery.d.ts.map +1 -1
  289. package/dist/internal/hosted-delivery.js +23 -25
  290. package/dist/internal/hosted-execution-diag.d.ts +12 -4
  291. package/dist/internal/hosted-execution-diag.d.ts.map +1 -1
  292. package/dist/internal/hosted-execution-diag.js +26 -4
  293. package/dist/internal/hosted-execution-flush.d.ts +1 -0
  294. package/dist/internal/hosted-execution-flush.d.ts.map +1 -1
  295. package/dist/internal/hosted-execution-flush.js +4 -2
  296. package/dist/internal/init-project.d.ts.map +1 -1
  297. package/dist/internal/init-project.js +4 -0
  298. package/dist/internal/server.d.ts.map +1 -1
  299. package/dist/internal/server.js +111 -48
  300. package/dist/internal/session-engine.d.ts +4 -1
  301. package/dist/internal/session-engine.d.ts.map +1 -1
  302. package/dist/internal/session-engine.js +39 -9
  303. package/dist/internal/skill-catalog.d.ts +28 -0
  304. package/dist/internal/skill-catalog.d.ts.map +1 -0
  305. package/dist/internal/skill-catalog.js +44 -0
  306. package/dist/playground/assets/index-DSMAewbx.css +1 -0
  307. package/dist/playground/assets/index-De_lpFxE.js +67 -0
  308. package/dist/playground/index.html +2 -2
  309. package/dist/types.d.ts +32 -3
  310. package/dist/types.d.ts.map +1 -1
  311. package/docs/README.md +2 -0
  312. package/docs/building-with-agents.md +96 -84
  313. package/docs/deployment.md +118 -364
  314. package/docs/evals.md +149 -673
  315. package/docs/guides/agent-to-agent.md +130 -48
  316. package/docs/guides/bitbucket.md +250 -49
  317. package/docs/guides/cloud-agents.md +119 -67
  318. package/docs/guides/convert-automation.md +157 -148
  319. package/docs/guides/github.md +261 -246
  320. package/docs/guides/gitlab.md +250 -51
  321. package/docs/guides/grokbot-agents.md +100 -55
  322. package/docs/guides/improve.md +112 -59
  323. package/docs/guides/jev.md +353 -0
  324. package/docs/guides/mcp-oauth.md +112 -122
  325. package/docs/guides/opentelemetry.md +97 -166
  326. package/docs/guides/slack.md +158 -228
  327. package/docs/guides/webhooks.md +155 -386
  328. package/docs/hillclimbing.md +102 -56
  329. package/docs/reference/cli.md +2 -0
  330. package/docs/reference/connections.md +1 -1
  331. package/docs/reference/evals.md +298 -0
  332. package/docs/reference/extensions.md +132 -203
  333. package/docs/reference/hooks.md +34 -34
  334. package/docs/reference/subagents.md +1 -1
  335. package/docs/reference/tools.md +4 -0
  336. package/docs/templates/thermo-quality-review.md +80 -0
  337. package/docs/templates/thermo-review.md +79 -0
  338. package/docs/troubleshooting.md +98 -67
  339. package/package.json +8 -1
  340. package/skills/github/SKILL.md +21 -13
  341. package/src/bin/agent-serve.ts +24 -3
  342. package/src/channels/checks.ts +8 -0
  343. package/src/channels/github/github-channel.ts +29 -19
  344. package/src/channels/origin/checks.ts +3 -1
  345. package/src/channels/slack/agentic-delivery.ts +2 -0
  346. package/src/channels/slack/dispatch.ts +131 -4
  347. package/src/channels/slack/placeholder.ts +51 -0
  348. package/src/channels/slack/slack-channel.ts +9 -2
  349. package/src/extensions/cursor-cloud-agents/skills/handoff.md +3 -7
  350. package/src/extensions/improve/extension.ts +4 -2
  351. package/src/extensions/improve/skills/yourself.ts +1 -1
  352. package/src/extensions/jev/extension.ts +95 -0
  353. package/src/extensions/jev/lib/evaluate.ts +289 -0
  354. package/src/extensions/jev/skills/gated-write.md +25 -0
  355. package/src/extensions/jev/skills/questions.md +33 -0
  356. package/src/extensions/jev/tools/evaluate.ts +90 -0
  357. package/src/extensions.ts +2 -0
  358. package/src/filesystem.ts +168 -65
  359. package/src/index.ts +3 -0
  360. package/src/internal/advertise-tools.ts +6 -0
  361. package/src/internal/cli-ax.ts +88 -15
  362. package/src/internal/continuation-identity.ts +1 -0
  363. package/src/internal/cursor-agent-template.ts +2 -0
  364. package/src/internal/discovery/connections.ts +21 -0
  365. package/src/internal/discovery/extensions.ts +12 -4
  366. package/src/internal/discovery/info.ts +1 -0
  367. package/src/internal/filesystem/tools.ts +2 -0
  368. package/src/internal/filesystem/walk.ts +60 -15
  369. package/src/internal/hosted-admission-adapter.ts +3 -0
  370. package/src/internal/hosted-delivery-protocol.ts +82 -1
  371. package/src/internal/hosted-delivery.ts +23 -0
  372. package/src/internal/hosted-execution-diag.ts +33 -4
  373. package/src/internal/hosted-execution-flush.ts +4 -0
  374. package/src/internal/init-project.ts +4 -0
  375. package/src/internal/server.ts +130 -53
  376. package/src/internal/session-engine.ts +46 -8
  377. package/src/internal/skill-catalog.ts +69 -0
  378. package/src/types.ts +33 -3
  379. package/templates/thermo-quality-review/README.md +35 -0
  380. package/templates/thermo-quality-review/agent/agent.ts +8 -0
  381. package/templates/thermo-quality-review/agent/channels/github.ts +42 -0
  382. package/templates/thermo-quality-review/agent/instructions.md +43 -0
  383. package/templates/thermo-quality-review/agent/tools/post_findings.ts +70 -0
  384. package/templates/thermo-quality-review/evals/evals.config.ts +5 -0
  385. package/templates/thermo-quality-review/evals/review.eval.ts +62 -0
  386. package/templates/thermo-quality-review/package.json +18 -0
  387. package/templates/thermo-quality-review/tsconfig.json +12 -0
  388. package/templates/thermo-review/README.md +35 -0
  389. package/templates/thermo-review/agent/agent.ts +8 -0
  390. package/templates/thermo-review/agent/channels/github.ts +42 -0
  391. package/templates/thermo-review/agent/instructions.md +40 -0
  392. package/templates/thermo-review/agent/tools/post_findings.ts +70 -0
  393. package/templates/thermo-review/evals/evals.config.ts +5 -0
  394. package/templates/thermo-review/evals/review.eval.ts +53 -0
  395. package/templates/thermo-review/package.json +18 -0
  396. package/templates/thermo-review/tsconfig.json +12 -0
  397. package/dist/docs/assets/building-with-agents.md.CEGVXkmO.js +0 -13
  398. package/dist/docs/assets/building-with-agents.md.CEGVXkmO.lean.js +0 -1
  399. package/dist/docs/assets/chunks/@localSearchIndexroot.DtXk1hy-.js +0 -1
  400. package/dist/docs/assets/chunks/channel.Bkv1N-gK.js +0 -1
  401. package/dist/docs/assets/chunks/classDiagram-6PBFFD2Q.CZDco1o8.js +0 -1
  402. package/dist/docs/assets/chunks/classDiagram-v2-HSJHXN6E.CZDco1o8.js +0 -1
  403. package/dist/docs/assets/chunks/clone.YSt_40_s.js +0 -1
  404. package/dist/docs/assets/chunks/framework.dypDpWZ3.js +0 -19
  405. package/dist/docs/assets/chunks/stateDiagram-v2-QKLJ7IA2.C4CLm481.js +0 -1
  406. package/dist/docs/assets/deployment.md.Dm4Qo3hp.js +0 -51
  407. package/dist/docs/assets/deployment.md.Dm4Qo3hp.lean.js +0 -1
  408. package/dist/docs/assets/evals.md.BLDRt5LH.js +0 -211
  409. package/dist/docs/assets/evals.md.BLDRt5LH.lean.js +0 -1
  410. package/dist/docs/assets/guides_agent-to-agent.md.BI0xclmy.js +0 -14
  411. package/dist/docs/assets/guides_agent-to-agent.md.BI0xclmy.lean.js +0 -1
  412. package/dist/docs/assets/guides_bitbucket.md.CTpCl__f.js +0 -10
  413. package/dist/docs/assets/guides_bitbucket.md.CTpCl__f.lean.js +0 -1
  414. package/dist/docs/assets/guides_cloud-agents.md.lSE_l7lH.js +0 -9
  415. package/dist/docs/assets/guides_cloud-agents.md.lSE_l7lH.lean.js +0 -1
  416. package/dist/docs/assets/guides_convert-automation.md.Ck6Cr68A.js +0 -12
  417. package/dist/docs/assets/guides_convert-automation.md.Ck6Cr68A.lean.js +0 -1
  418. package/dist/docs/assets/guides_github.md.D6ER29dG.js +0 -64
  419. package/dist/docs/assets/guides_github.md.D6ER29dG.lean.js +0 -1
  420. package/dist/docs/assets/guides_gitlab.md.P-TjBnS5.js +0 -14
  421. package/dist/docs/assets/guides_gitlab.md.P-TjBnS5.lean.js +0 -1
  422. package/dist/docs/assets/guides_grokbot-agents.md.WBZIOvkz.js +0 -9
  423. package/dist/docs/assets/guides_grokbot-agents.md.WBZIOvkz.lean.js +0 -1
  424. package/dist/docs/assets/guides_improve.md.BKaDuKKK.js +0 -14
  425. package/dist/docs/assets/guides_improve.md.BKaDuKKK.lean.js +0 -1
  426. package/dist/docs/assets/guides_mcp-oauth.md.DMNMpXtO.js +0 -28
  427. package/dist/docs/assets/guides_mcp-oauth.md.DMNMpXtO.lean.js +0 -1
  428. package/dist/docs/assets/guides_opentelemetry.md._CRfDyzH.js +0 -26
  429. package/dist/docs/assets/guides_opentelemetry.md._CRfDyzH.lean.js +0 -1
  430. package/dist/docs/assets/guides_slack.md.DdT8rmsj.js +0 -46
  431. package/dist/docs/assets/guides_slack.md.DdT8rmsj.lean.js +0 -1
  432. package/dist/docs/assets/guides_webhooks.md.aQW10HRe.js +0 -225
  433. package/dist/docs/assets/guides_webhooks.md.aQW10HRe.lean.js +0 -1
  434. package/dist/docs/assets/hillclimbing.md.Dq4kkVIL.js +0 -1
  435. package/dist/docs/assets/hillclimbing.md.Dq4kkVIL.lean.js +0 -1
  436. package/dist/docs/assets/reference_extensions.md.wlFD3cUR.js +0 -62
  437. package/dist/docs/assets/reference_extensions.md.wlFD3cUR.lean.js +0 -1
  438. package/dist/docs/assets/troubleshooting.md.BcgNoYtJ.js +0 -1
  439. package/dist/docs/assets/troubleshooting.md.BcgNoYtJ.lean.js +0 -1
  440. package/dist/playground/assets/index-BLlKgZtI.css +0 -1
  441. package/dist/playground/assets/index-CDS5p9sR.js +0 -67
@@ -1,778 +1,254 @@
1
1
  # Evals
2
2
 
3
- An eval sends a fixed message to your agent and asserts over the
4
- trajectory it records: the turn completed, the right tool ran with the
5
- right input, the reply has the right shape. Evals are how you know a
6
- prompt tweak helped, a refactor didn't regress the agent, and last
7
- month's fix still holds.
8
-
9
- Nothing is mocked. The runner starts (or targets) a real agent server,
10
- drives sessions over the public API, and grades the events it gets
11
- back. The model runs and server tools execute, so
12
- [keep side effects out of eval sessions](#keep-side-effects-out-of-eval-sessions)
13
- before you point an eval at an agent that posts anywhere.
3
+ An eval sends a fixed input to the real agent and checks the resulting
4
+ trajectory: whether the turn succeeded, which tools ran, and what shape
5
+ the answer took. Use evals to protect behavior you already understand,
6
+ not to discover what the prompt should do.
14
7
 
15
- ## Evals, hooks, or hillclimbing?
8
+ Evals run real model turns and real tools. Guard external writes before
9
+ running a suite against an agent that can post, merge, or deploy.
16
10
 
17
- All three read the same session event stream. Pick by the question you
18
- are asking.
19
-
20
- | You want to | Use |
21
- | --- | --- |
22
- | Gate one fixed input's behavior, locally and in CI | Evals (this page) |
23
- | Observe every live session: metrics, audit, alerts | [Hooks](/docs/reference/hooks.md) |
24
- | Improve an agent one measured round at a time | [Hillclimbing](/docs/hillclimbing.md); each kept win lands an eval |
25
-
26
- [Hooks, channel events, or evals?](/docs/reference/hooks.md#hooks-channel-events-or-evals)
27
- has the side-by-side table.
28
-
29
- ### When not to write an eval
30
-
31
- - Test a server tool's own logic with
32
- `agent-sdk call <tool> --dir . --input '{...}'` or a unit test. No
33
- model turn, no credential.
34
- - Explore a prompt with `agent-sdk run --dir . --message "..."` and
35
- read the trajectory. Write the eval once you know which decision to
36
- gate.
37
- - Stop a bad turn while it runs with
38
- [`needsApproval`](/docs/reference/tools.md#gate-a-tool-on-human-approval)
39
- on the tool or `defineResult`. Evals grade
40
- after the fact.
41
-
42
- ## Write your first eval
43
-
44
- Evals live under the project-root `evals/` directory, a sibling of
45
- `agent/`. `agent/evals/` is silently ignored. Discovery loads every
46
- `.eval.ts` (or `.eval.js`) file under `evals/`, plus one config file.
47
-
48
- ```text
49
- my-agent/
50
- agent/
51
- agent.ts
52
- tools/inspect_pr.ts
53
- evals/
54
- evals.config.ts # required to run: maxConcurrency
55
- readiness.eval.ts # id: readiness
56
- prs.eval.ts # cases: prs/checkout, prs/search
57
- ```
11
+ ## Protect a tool decision
12
+
13
+ This smoke case asks for a PR verdict, requires the read-only inspection
14
+ tool, and fails if the agent tries to approve. The CLI reports all three
15
+ decisions together instead of stopping at the first miss.
58
16
 
59
- An eval is a single `async test(t)`. You drive the agent with `t.send`
60
- and assert on the recorded run with the same `t`:
17
+ Use the [Jev extension](/docs/guides/jev.md) when host code needs a typed
18
+ choice, score, or boolean before it writes.
61
19
 
62
20
  ```ts
63
21
  // evals/readiness.eval.ts
64
22
  import { defineEval, includes } from "@cursor/july/evals";
65
23
 
66
24
  export default defineEval({
67
- description: "Inspects a PR without approving it.",
68
25
  tags: ["smoke"],
69
- timeoutMs: 120_000,
70
26
  async test(t) {
71
27
  await t.send(
72
28
  "Is https://github.com/acme/checkout/pull/42 ready to approve?"
73
29
  );
30
+
74
31
  t.succeeded();
75
32
  t.calledTool("inspect_pr");
76
33
  t.notCalledTool("approve_pr");
77
- t.check(t.reply, includes(/ready|approve/i));
34
+ t.check(t.reply, includes(/ready|blocked|approve/i));
78
35
  },
79
36
  });
80
37
  ```
81
38
 
39
+ Evals live in the project-root `evals/` directory, beside `agent/`.
40
+ Add the required concurrency config once:
41
+
82
42
  ```ts
83
43
  // evals/evals.config.ts
84
44
  import { defineEvalConfig } from "@cursor/july/evals";
85
45
 
86
- export default defineEvalConfig({ maxConcurrency: 20 });
46
+ export default defineEvalConfig({
47
+ maxConcurrency: 20,
48
+ });
87
49
  ```
88
50
 
89
- Run it under Node 22.13 or newer (never Bun) with a Cursor credential
90
- in place; see [Credentials](#credentials):
91
-
92
51
  ```bash
93
52
  agent-sdk eval --dir . --list
94
- agent-sdk eval --dir . readiness
95
- ```
96
-
97
- ```text
98
- PASS readiness (14.2s) — Inspects a PR without approving it.
99
- ✓ succeeded
100
- ✓ calledTool(inspect_pr)
101
- ✓ notCalledTool(approve_pr)
102
- ✓ check(includes)
103
-
104
- 1 passed, 0 failed, 1 total
105
- artifacts: <project state directory>/evals/2026-09-11T15-02-11-402Z
106
- ```
107
-
108
- Every local run writes each case's assertions, inputs, tool calls, and
109
- `t.log` lines under that artifacts directory. Open
110
- `evals/<case-id>.json` there when a case fails; see
111
- [Where results land](#where-results-land).
112
-
113
- ## Name cases by path
114
-
115
- The file path is the eval's identity, so you don't author an id.
116
- `evals/builds/api.eval.ts` becomes `builds/api`. An `index` filename
117
- collapses to its directory: `evals/builds/index.eval.ts` becomes
118
- `builds`.
119
-
120
- One file can hold several datapoints through `cases`. Provide either
121
- `test` or `cases`, not both. Each case id becomes `<fileId>/<case.id>`:
122
-
123
- ```ts
124
- // evals/prs.eval.ts: prs/checkout, prs/search
125
- export default defineEval({
126
- tags: ["smoke", "prs"],
127
- cases: [
128
- {
129
- id: "checkout",
130
- description: "Checkout PR readiness.",
131
- async test(t) {
132
- await t.send(
133
- "Is https://github.com/acme/checkout/pull/42 ready to approve?"
134
- );
135
- t.succeeded();
136
- t.calledTool("inspect_pr");
137
- },
138
- },
139
- {
140
- id: "search",
141
- async test(t) {
142
- await t.send(
143
- "Check https://github.com/acme/search/pull/7 before approval."
144
- );
145
- t.succeeded();
146
- t.calledTool("inspect_pr");
147
- },
148
- },
149
- ],
150
- });
151
- ```
152
-
153
- Case ids are single path segments, unique within the file. A case can
154
- set its own `description`, `tags`, `timeoutMs`, `iterations`, `judge`,
155
- `reporters`, and `metadata`. A case-level value replaces the file-level
156
- one for that datapoint, except `metadata`, which merges with case keys
157
- winning, and `reporters`, which adds to the file's list. `metadata` is
158
- free-form data carried onto the result and every reporter.
159
-
160
- A file may instead export an array of `defineEval` calls to fan out
161
- over a dataset. Ids are then the file id plus a zero-padded index
162
- (`sql/0000`, `sql/0001`, ...); see [Load a dataset](#load-a-dataset).
163
- Prefer `cases` when datapoints are hand-written and deserve stable
164
- names.
165
-
166
- ### Iterations
167
-
168
- `iterations` (file or case, default `1`, cap `100`) runs a datapoint
169
- repeatedly. Discovery expands `iterations: 3` on case `nyc` to runnable
170
- ids `weather/nyc/1`, `weather/nyc/2`, and `weather/nyc/3`. The filter
171
- `weather/nyc` still selects all three. Each expanded case exposes
172
- `t.iteration` and `t.iterations`.
173
-
174
- `maxConcurrency` counts authored datapoints, not expanded iterations.
175
- Iterations of one datapoint share a concurrency slot and run in
176
- sequence, so a suite of 11 cases with 3 iterations each and
177
- `maxConcurrency: 20` has at most 11 cases in flight.
178
-
179
- ## Drive the agent with `t.send`
180
-
181
- `t.send(message, options?)` runs one turn and waits for it to settle:
182
- complete, park on an approval request, or fail. Several sends in one
183
- case share the session, which is how you write multi-turn evals.
184
-
185
- Each send resolves to a turn result: `message` (the assistant text),
186
- `sessionId`, `events`, `toolCalls` (tool names in order), `ok`, and
187
- `index`. The turn carries the same assertion vocabulary as `t`, scoped
188
- to that turn, so you can grade an intermediate turn before the next
189
- send overwrites `t.reply`. `turn.expectOk()` throws when the turn
190
- failed, for later steps that depend on it.
191
-
192
- Read the whole case with `t.reply` (last assistant text), `t.events`
193
- (every event so far), `t.turns` (settled turns, oldest first), and
194
- `t.sessionId`. `t.signal` aborts when the case hits its timeout; pass
195
- it to your own async work.
196
-
197
- Three options apply on the first send only, because they shape session
198
- creation:
199
-
200
- | Option | Effect |
201
- | --- | --- |
202
- | `workspaceFiles` | `{ path: contents }` seeded into the local session workspace. Prefer this over machine-local paths |
203
- | `workspaceDir` | Absolute harness cwd for the local runtime |
204
- | `cloud` | Per-session cloud options merged over the agent's static `cloud` config. Pin a fixture repo here for cloud evals instead of on the agent's default `cloud.repos`. On the cloud runtime, seeded files reach the agent as described under [Choose a runtime](/docs/reference/agent-config.md#choose-a-runtime) |
205
-
206
- ```ts
207
- await t.send("Review pr/diff.patch and post findings.", {
208
- workspaceFiles: {
209
- "pr/diff.patch": [
210
- "diff --git a/app/routes/search.ts b/app/routes/search.ts",
211
- "+res.send(`<h1>Results for ${req.query.q}</h1>`);",
212
- ].join("\n"),
213
- },
214
- });
215
- ```
216
-
217
- ## Assert over the trajectory
218
-
219
- Assertions record; they never throw. One run reports every failure
220
- instead of dying on the first. Assertions on `t` read the whole run.
221
- Assertions on a turn read only that turn.
222
-
223
- | Gate | Checks |
224
- | --- | --- |
225
- | `t.succeeded()` | the run did not fail and is not parked on an unanswered approval |
226
- | `t.parked()` | the run cleanly parked on an unanswered approval request |
227
- | `t.messageIncludes(token)` | the joined assistant text matches a string or `RegExp` |
228
- | `t.calledTool(name, matcher?)` | a matching call to `name` happened |
229
- | `t.notCalledTool(name)` | no request for `name`, in any lifecycle state |
230
- | `t.loadedSkill(name)` | the agent opened `skills/<name>/SKILL.md` (read, grep, or shell `cat`) |
231
- | `t.toolOrder(names)` | tool requests appear in this relative order; extra calls allowed |
232
- | `t.usedNoTools()` | no tool calls at all |
233
- | `t.maxToolCalls(max)` | at most `max` tool calls |
234
- | `t.noFailedActions()` | no tool call reported an error |
235
- | `t.calledSubagent(name, matcher?)` | a matching subagent delegation happened |
236
- | `t.taggedArtifact(kind?, predicate?)` | at least one [artifact](/docs/reference/artifacts.md) was tagged |
237
- | `t.event(type, matcher?)` | at least one matching [event](/docs/reference/sessions.md#which-events-can-i-stream) of `type` |
238
- | `t.notEvent(type, matcher?)` | no matching event of `type` |
239
- | `t.eventOrder(matchers)` | matching event groups occur in this relative order |
240
- | `t.eventsSatisfy(label, predicate)` | your predicate over the typed event stream |
241
- | `t.check(value, expectation)` | any value, against a [builder](#grade-values-with-expectation-builders) |
242
- | `t.score(name, value)` | a 0-1 score you computed; soft until you add a bar |
243
-
244
- Three more assertions gate and return the matched fact. They stop the
245
- test body when nothing matches, without a duplicate execution error.
246
- `t.requireToolCall(name, matcher?)` returns the call so later code can
247
- read its `input` and `output`. `t.requireInputRequest(filter?)` returns
248
- the single pending approval request. `await t.require(value, expectation)`
249
- does the same for a value check.
250
-
251
- A case with no assertions passes when at least one turn completed. Add
252
- `t.succeeded()` and behavior gates anyway. They make the contract
253
- visible in review.
254
-
255
- ### What good cases assert
256
-
257
- Gate decisions and shape, not prose. Model wording varies run to run.
258
- Tool choice, tool avoidance, and output structure are the stable
259
- contract.
260
-
261
- 1. `t.succeeded()`: always, first.
262
- 2. The tool decision: `calledTool` for the intended path,
263
- `notCalledTool` for the likely wrong alternative. The pair is
264
- stronger than either alone.
265
- 3. Output shape: a regex for the contract (`/ready|blocked/i`, a JSON
266
- marker, a findings-block fence), never exact sentences.
267
- 4. For structured output, parse `t.reply` and check fields with
268
- `matches` or `satisfies` instead of substring-matching JSON.
269
-
270
- The common failure modes: asserting exact phrasing, packing more than
271
- about five gates into one case (split it), and cases that depend on
272
- live external state that drifts (pin the input).
273
-
274
- ### Narrow tool assertions with matchers
275
-
276
- With no matcher, `calledTool` is request-based: a requested call counts
277
- even before its result arrives. A matcher narrows it:
278
-
279
- ```ts
280
- t.calledTool("inspect_pr", { status: "completed" });
281
- t.calledTool("apply_agents", { input: { verdict: "update" } });
282
- t.calledTool("bash", { input: { command: /^gh pr view/ }, count: 1 });
283
- t.calledTool("read_file", {
284
- output: (value) => String(value).includes("TODO"),
285
- });
53
+ agent-sdk eval --dir . --tag smoke
286
54
  ```
287
55
 
288
- `input`, `output`, and `count` accept a literal, a `RegExp`, or a
289
- predicate. Object literals partial-deep-match, so `{ verdict: "update" }`
290
- matches arguments that also carry other keys. `status` is one of
291
- `completed`, `failed`, `pending`, or `rejected` (a human denied the
292
- approval). `calledSubagent` takes `{ output, status, count, callId }`.
293
- `event`, `notEvent`, and `eventOrder` take `{ data, count }`.
56
+ ## Check answer shape or quality
294
57
 
295
- ### Grade values with expectation builders
296
-
297
- `t.check(value, expectation)` grades any value: `t.reply`, a parsed
298
- JSON field, a tool's output.
299
-
300
- | Builder | Checks | Severity |
301
- | --- | --- | --- |
302
- | `includes(string \| RegExp)` | substring or match; structured values are stringified first | gate |
303
- | `equals(value)` | deep equality | gate |
304
- | `matches(schema)` | a Standard Schema (Zod, Valibot, ...) or anything with `safeParse` | gate |
305
- | `similarity(expected)` | normalized text similarity, 0-1 | soft |
306
- | `satisfies(predicate, label)` | your predicate; `label` is the failure detail | gate |
58
+ Prefer a deterministic shape check when the output has a contract. The
59
+ failure tells you which field was wrong, and the case remains stable
60
+ when the model changes its wording.
307
61
 
308
62
  ```ts
309
- import { matches, satisfies } from "@cursor/july/evals";
63
+ import { matches } from "@cursor/july/evals";
310
64
  import { z } from "zod";
311
65
 
312
66
  const verdict = JSON.parse(t.reply ?? "{}");
313
67
  t.check(
314
68
  verdict,
315
- matches(z.object({ ready: z.boolean(), blockers: z.array(z.string()) }))
316
- );
317
- t.check(
318
- verdict.blockers.length,
319
- satisfies((n) => (n as number) <= 3, "at most 3 blockers")
69
+ matches(
70
+ z.object({
71
+ ready: z.boolean(),
72
+ blockers: z.array(z.string()),
73
+ })
74
+ )
320
75
  );
321
76
  ```
322
77
 
323
- `normalizedSimilarity(actual, expected)` returns the same 0-1 score as
324
- `similarity`, for use with `t.score`.
325
-
326
- ### Record without gating
327
-
328
- - `t.metric(name, value)` records a structured score or label. It shows
329
- on the CLI result, the playground case card, JUnit output, and
330
- artifacts.
331
- - `t.log(message)` records a debug line, streamed under `--verbose`.
332
- - `t.skip(reason)` ends the case as skipped. Skipped cases report
333
- separately and never change the exit code. Call it before sending
334
- messages.
335
-
336
- ## Gates, soft scores, and verdicts
337
-
338
- Every assertion returns a handle, so severity rides on the assertion
339
- instead of a separate thresholds map:
340
-
341
- ```ts
342
- t.succeeded(); // gate (default)
343
- t.calledTool("get_weather").soft(); // tracked, never fails
344
- t.check(t.reply, similarity("Sunny, 72°F")).atLeast(0.8); // soft, with a bar
345
- t.judge.closedQA("cites a source").gate(0.8); // promoted to a gate
346
- ```
347
-
348
- - `.gate(threshold?)` is hard. A miss fails the case and `eval` exits 1.
349
- - `.soft(threshold?)` is tracked. With no threshold it never fails.
350
- - `.atLeast(threshold)` is soft with a bar. A miss marks the case
351
- `scored`.
352
-
353
- Each case ends with one verdict:
354
-
355
- | Verdict | Meaning | Exit code |
356
- | --- | --- | --- |
357
- | `passed` | every gate passed and no soft bar was missed | 0 |
358
- | `failed` | a gate failed, or the test body threw | 1 |
359
- | `scored` | only soft bars were missed | 0, or 1 under `--strict` |
360
- | `skipped` | `t.skip(reason)`, or a judge with no credentials | 0 |
361
-
362
- The CLI prints soft misses as `~` and gate misses as `✗`. Start a new
363
- benchmark with `t.score("recall", recall)` and `.atLeast()` so its
364
- number reports for a while without blocking merges. Add `--strict`
365
- once the bars are trustworthy.
366
-
367
- ## Judge free-form output
368
-
369
- When wording matters and no regex captures it, `t.judge` grades with an
370
- LLM. The graders are `factuality(expected)`, `summarizes(expected)`,
371
- `closedQA(criteria)`, and `sql(expected)`. Each scores `t.reply` by
372
- default; pass `{ on }` to grade another value.
78
+ Use a judge when correctness depends on meaning that a schema, regex, or
79
+ known value cannot capture:
373
80
 
374
81
  ```ts
375
- const summary = await t.send("Why did CI fail on PR 42?");
376
- t.judge.factuality("The lint step failed on src/sidebar.ts.").atLeast(0.7);
377
- t.judge.closedQA("names the failing step", { on: summary.message }).gate(1);
82
+ t.judge
83
+ .factuality("The lint step failed on src/sidebar.ts.")
84
+ .atLeast(0.8);
378
85
  ```
379
86
 
380
- Judge assertions are soft by default, so a judge never fails a build
381
- until you give it a bar with `.atLeast()` or promote it with `.gate()`.
382
- The recorded detail names the choice the judge made and its rationale.
87
+ Judge checks are tracked scores until you give them a hard gate.
88
+ Configure the judge model in `evals.config.ts`. The
89
+ [Evals reference](/docs/reference/evals.md#judges) owns grader options and
90
+ severity rules.
383
91
 
384
- The judge model comes from `defineEvalConfig({ judge })`,
385
- `defineEval({ judge })`, a case-level `judge`, or a per-call
386
- `{ model }`. The nearest one wins. A judge call with no model
387
- configured fails the case. A judge that cannot reach a model (no
388
- credential) ends the case as `skipped`, unless a deterministic gate
389
- already failed.
92
+ ## Test a conversation or approval
390
93
 
391
- For a domain-specific judge whose verdict is not a single score,
392
- `t.judge.model(prompt)` sends a raw prompt to the same model and
393
- returns the reply. Record the parsed result with `t.score` or
394
- `t.check`. Anything derived from the agent under test is untrusted
395
- input to your prompt: wrap it with `fenceUntrusted` and include
396
- `EVAL_JUDGE_INJECTION_GUARD`, as the built-in graders do.
94
+ Several `t.send` calls in one case share a session. Assertions on a
95
+ returned turn inspect only that turn, while assertions on `t` inspect
96
+ the whole conversation.
397
97
 
398
98
  ```ts
399
- import {
400
- EVAL_JUDGE_INJECTION_GUARD,
401
- fenceUntrusted,
402
- } from "@cursor/july/evals";
403
-
404
- const gold = ["XSS in search.ts", "open redirect in login.ts"];
405
- const reply = await t.judge.model(
406
- [
407
- "For each GOLD finding, answer whether SUBMISSION reports it.",
408
- "Reply with one line per finding: <index> YES|NO.",
409
- EVAL_JUDGE_INJECTION_GUARD,
410
- fenceUntrusted("GOLD", gold.map((g, i) => `${i + 1}. ${g}`).join("\n")),
411
- fenceUntrusted("SUBMISSION", t.reply ?? ""),
412
- ].join("\n\n")
413
- );
414
- const hits = reply.match(/\bYES\b/g)?.length ?? 0;
415
- t.score("recall", hits / gold.length).atLeast(0.5);
416
- ```
417
-
418
- ## Keep side effects out of eval sessions
99
+ const first = await t.send("Inspect PR 42.");
100
+ first.calledTool("inspect_pr");
419
101
 
420
- Eval sessions run the real agent, tools included. A reviewer that
421
- comments on GitHub or posts to Slack will do so from an eval unless
422
- the tool checks the session's purpose. Eval sessions carry
423
- `purpose: "eval"`; live traffic carries `"live"`. Branch on it in the
424
- tool, hook, or result handler that actuates:
425
-
426
- ```ts
427
- // agent/tools/post_findings.ts
428
- async execute({ findings }, ctx) {
429
- if (ctx.session.purpose === "eval") {
430
- return { posted: false, reason: "eval", count: findings.length };
431
- }
432
- // post the review
433
- }
102
+ const followUp = await t.send("Now summarize only the blockers.");
103
+ followUp.usedNoTools();
104
+ t.check(followUp.message, includes(/blocker/i));
105
+ t.succeeded();
434
106
  ```
435
107
 
436
- Return a shaped result instead of throwing, so the eval can still
437
- assert `t.calledTool("post_findings", { input: ... })` on the decision.
438
- The same check belongs in [hooks](/docs/reference/hooks.md) that meter or
439
- page and in `defineResult` commits.
440
-
441
- ## Worked examples
442
-
443
- ### Multi-turn: grade each turn
108
+ For a tool that requires approval, the expected result is a parked turn
109
+ instead of a completed one:
444
110
 
445
111
  ```ts
446
- // evals/intro.eval.ts
447
- import { defineEval, includes, satisfies } from "@cursor/july/evals";
448
-
449
- export default defineEval({
450
- description: "Introduces itself once; a repeat mention gets a short ack.",
451
- async test(t) {
452
- const intro = await t.send("Meet Jenny! @Jenny introduce yourself.");
453
- intro.expectOk();
454
- t.check(intro.message, includes(/jenny/i));
455
-
456
- const repeat = await t.send("Meet, @Jenny!");
457
- t.succeeded();
458
- repeat.usedNoTools();
459
- t.check(
460
- repeat.message,
461
- satisfies((r) => (r as string).trim().length <= 280, "short ack")
462
- );
463
- t.check(
464
- repeat.message,
465
- satisfies((r) => !/what i can do/i.test(r as string), "no second intro")
466
- );
467
- },
112
+ await t.send("Apply the approved policy update.");
113
+ t.parked();
114
+ t.calledTool("apply_policy", {
115
+ input: { verdict: "update" },
468
116
  });
469
117
  ```
470
118
 
471
- `t.succeeded()` grades the whole session. `repeat.usedNoTools()` and the
472
- checks on `repeat.message` read only the second turn, even though
473
- `t.reply` now holds its text.
119
+ Pair an approval case with a safe branch that calls
120
+ `t.notCalledTool("apply_policy")`, so both sides of the decision stay
121
+ protected.
474
122
 
475
- ### Approvals: assert the parked decision
123
+ ## Pin realistic inputs
476
124
 
477
- For a tool with `needsApproval`, the turn parks instead of finishing.
478
- Gate on `t.parked()` and on the arguments the model chose:
125
+ A useful eval changes only when the agent changes. Freeze canonical chat
126
+ prompts, seed workspace evidence directly, and save webhook payloads
127
+ instead of depending on a developer's checkout or live external state.
479
128
 
480
129
  ```ts
481
- // evals/agents.eval.ts (one case; RULE and SLACK are fixture strings)
482
- {
483
- id: "update-rule",
484
- description: "A repeated billing rule parks the AGENTS.md write.",
485
- async test(t) {
486
- await t.send("Weekly AGENTS.md review. Read week/ and call apply_agents once.", {
487
- workspaceFiles: {
488
- "week/prs.md": RULE,
489
- "week/slack.md": SLACK,
490
- "week/tree/AGENTS.md.txt": "# API\n\nKeep handlers thin.\n",
491
- },
492
- });
493
- t.parked();
494
- t.calledTool("apply_agents", { input: { verdict: "update" } });
130
+ await t.send("Review pr/diff.patch and report blockers.", {
131
+ workspaceFiles: {
132
+ "pr/diff.patch": [
133
+ "diff --git a/app/routes/search.ts b/app/routes/search.ts",
134
+ "+res.send(`<h1>Results for ${req.query.q}</h1>`);",
135
+ ].join("\n"),
495
136
  },
496
- },
497
- ```
498
-
499
- `t.parked()` and `t.succeeded()` are exclusive: a parked run is a clean
500
- stop on an unanswered approval, not a completed one. Pair the parked
501
- case with a sibling that expects `verdict: "skip"` and `t.succeeded()`,
502
- so both branches stay pinned.
503
-
504
- ## Pin fixtures
505
-
506
- A fixed input is what makes an eval repeatable. Pick the fixture by the
507
- surface under test.
508
-
509
- | Agent surface | Fixture |
510
- | --- | --- |
511
- | Chat or domain assistant | One canonical prompt string, chosen once and frozen |
512
- | Tool-heavy | Run `agent-sdk call <tool>` first to pin what the tool returns, then freeze the prompt that triggers it |
513
- | GitHub webhook | `agent-sdk github replay <pr> --events '*' --dry-run --out fixtures/github` snapshots real payloads for offline replay ([GitHub guide](/docs/guides/github.md#test-with-github-replay)) |
514
- | PR reviewer with host preparation | Diff, metadata, and gold labels pinned to commit SHAs; keep any live PR matrix small |
515
- | Workspace-dependent | `workspaceFiles` on the first `t.send`, never developer-machine paths |
516
-
517
- Tag the fast, reliably passing core `smoke` and run `--tag smoke` in
518
- the inner loop. Leave slow or drift-prone cases untagged for explicit
519
- runs.
520
-
521
- ### Materialize API-backed fixtures
522
-
523
- An input that only points at external data (a pull request URL, a
524
- snapshot id, a pair of commit SHAs) is not self-contained. Fetch it once
525
- and commit the rendered fixture before you expand the suite:
526
-
527
- 1. Save the diff, metadata, and labels under `fixtures/` at pinned
528
- revisions.
529
- 2. Seed those files with `workspaceFiles`, or read them from the
530
- fixture directory.
531
- 3. Assert decisions and output shape against the saved evidence.
532
- 4. Keep a small `smoke` subset for any remaining live checks.
533
-
534
- `maxConcurrency` limits parallel datapoints, not the model or API
535
- fan-out inside one datapoint. Materialized fixtures keep a large suite
536
- from exhausting provider and GitHub rate limits.
537
-
538
- ### Load a dataset
539
-
540
- Read committed fixtures with `loadJson`, `loadJsonl`, and `loadYaml`
541
- from `@cursor/july/evals/loaders`. Relative paths resolve against the
542
- project root the runner discovered, not the cwd the CLI ran from. Eval
543
- files are ES modules, so top-level `await` can load a dataset and fan
544
- one file out over it:
545
-
546
- ```ts
547
- // evals/sql.eval.ts: sql/0000, sql/0001, ...
548
- import { defineEval, equals } from "@cursor/july/evals";
549
- import { loadYaml } from "@cursor/july/evals/loaders";
550
-
551
- const rows = await loadYaml<{ task: string; prompt: string; sql: string }[]>(
552
- "evals/data/cases.yaml"
553
- );
554
-
555
- export default rows.map((row) =>
556
- defineEval({
557
- description: row.task,
558
- async test(t) {
559
- await t.send(row.prompt);
560
- t.succeeded();
561
- t.check(t.reply, equals(row.sql));
562
- },
563
- })
564
- );
565
- ```
566
-
567
- ## Configure eval runs
568
-
569
- `evals/evals.config.ts` (or `.js`) holds project-wide defaults. It must
570
- set `maxConcurrency`. Each case issues real model requests, so
571
- concurrency is hard-capped at 200; the templates use 10.
572
- `eval --list` works without the file. Running a case does not.
573
-
574
- ```ts
575
- import { defineEvalConfig } from "@cursor/july/evals";
576
-
577
- export default defineEvalConfig({
578
- maxConcurrency: 20,
579
- timeoutMs: 180_000,
580
- judge: { model: "gpt-5.4-mini" },
581
137
  });
582
138
  ```
583
139
 
584
- | Option | Default | Meaning |
585
- | --- | --- | --- |
586
- | `maxConcurrency` | required | Datapoints in flight at once, 1-200 |
587
- | `timeoutMs` | `180_000` | Per-case timeout. Precedence: case or file `timeoutMs`, then `--timeout-ms`, then this value |
588
- | `judge` | unset | Default judge model for `t.judge.*` |
589
- | `reporters` | unset | Destinations that observe every case; `--skip-report` suppresses them |
590
- | `maxPlaygroundRuns` | `20` | Batches kept in playground history (cap 500). Counts `--prod` and `--url` batches, not the default local run |
591
-
592
- Reporters ship results somewhere; the runner still does the grading.
593
- `JUnit({ filePath, suiteName? })` writes JUnit XML and
594
- `Artifacts({ dir })` writes per-case files, both from
595
- `@cursor/july/evals/reporters`. A custom reporter is an object with any
596
- of `onRunStart`, `onEvalComplete`, and `onRunComplete`. A reporter that
597
- throws is logged and never fails the run. CI usually attaches the
598
- built-in two with `--junit` and `--artifacts` instead of `reporters`,
599
- so output paths stay with the pipeline, not the eval author.
600
-
601
- Playground batches survive restarts when the project configures
602
- durable storage. Otherwise they live in process
603
- memory until `serve` exits.
604
-
605
- ## Run evals from the CLI
140
+ For GitHub agents, snapshot a PR into committed fixtures:
606
141
 
607
142
  ```bash
608
- agent-sdk eval --dir . --list # discover only
609
- agent-sdk eval --dir . # run all
610
- agent-sdk eval --dir . builds/checkout # one datapoint
611
- agent-sdk eval --dir . builds search # several ids or prefixes
612
- agent-sdk eval --dir . --tag smoke --tag pull-request # any matching tag
613
- agent-sdk eval --dir . --verbose # t.log lines + reply snippets
614
- agent-sdk eval --dir . --timeout-ms 300000 # override the per-case timeout
143
+ agent-sdk github replay acme/checkout#42 \
144
+ --events '*' \
145
+ --dry-run \
146
+ --out fixtures/github
615
147
  ```
616
148
 
617
- Id filters use OR semantics. Each filter selects an exact id and its
618
- descendants: `builds` selects `builds`, `builds/checkout`, and every
619
- other case below that path. Repeated tags also use OR. With both ids
620
- and tags, a case must match both groups.
149
+ Keep a small, fast `smoke` set for every change. Put larger datasets and
150
+ drift-prone live cases behind an explicit selection. See
151
+ [Datasets and fixtures](/docs/reference/evals.md#datasets-and-fixtures) for
152
+ loaders and case expansion.
621
153
 
622
- By default `eval` boots a throwaway server with its own state root, so
623
- cases don't inherit your checkout's `AGENTS.md` and session state stays
624
- out of the project. Artifacts still land in the project state
625
- directory; see [Where results land](#where-results-land). `--slug`
626
- picks the target in a multi-agent directory.
154
+ ## Run and debug a regression suite
627
155
 
628
- `--url` runs the batch on a running server instead, the same way
629
- `--prod` does: that server discovers its own `evals/`, results land in
630
- its playground history, and the local-only flags (`--junit`,
631
- `--artifacts`, `--max-concurrency`, `--skip-report`) do not apply. See
632
- [Run evals in the playground or on a deployment](#run-evals-in-the-playground-or-on-a-deployment).
156
+ Run one case while iterating, then the smoke suite before review. A
157
+ failure artifact records the input, assertions, tool calls, and final
158
+ text, making it the first place to look before changing the prompt.
633
159
 
634
160
  ```bash
635
- agent-sdk eval --url http://127.0.0.1:3000/weather-agent \
636
- --bearer-token "$AGENT_TOKEN"
161
+ agent-sdk eval --dir . readiness
162
+ agent-sdk eval --dir . --tag smoke
163
+ agent-sdk eval --dir . --json
637
164
  ```
638
165
 
639
- See [CLI: eval](/docs/reference/cli.md#eval) for every flag.
640
-
641
- ### Credentials
166
+ By default, results land under `evals/<stamp>/` in the project state
167
+ directory. Open `evals/<case-id>.json` inside that run directory for the
168
+ failed case. This artifact location is independent of `--state-root`;
169
+ pass `--artifacts <dir>` to choose another destination.
642
170
 
643
- Model turns need a
644
- [Cursor credential](/docs/reference/cli.md#environment-variables):
645
- `CURSOR_API_KEY`, `CURSOR_SERVICE_ACCOUNT_KEY`, or a saved
646
- `agent-sdk login`. The judge uses the same one. `eval --list` needs
647
- none.
648
-
649
- ### Where results land
650
-
651
- Every local run writes artifacts to a timestamped directory under
652
- `evals/` in the project state directory, whatever `--state-root` says.
653
- `--artifacts <dir>` chooses the path and `--no-artifacts` skips them.
654
- The directory holds `summary.json`,
655
- `results.jsonl`, and `evals/<case-id>.json` with every assertion, the
656
- inputs, tool calls with arguments and output, the final text, and
657
- `t.log` lines. Start there when a case fails. `--out <file>` also
658
- writes the full results JSON to a path of your choice.
659
-
660
- The artifact does not include the session's event stream. Pass
661
- `--state-root <path>` to keep the ephemeral server's
662
- [session data](/docs/reference/sessions.md#where-does-the-agent-sdk-store-session-data)
663
- on disk when you need the raw events.
171
+ Model turns need a credential. Resolution checks `CURSOR_API_KEY`,
172
+ `CURSOR_API_KEY_FILE`, `CURSOR_SERVICE_ACCOUNT_KEY`, then a saved
173
+ `agent-sdk login`. `eval --list` only discovers cases and needs none.
664
174
 
665
175
  ## Run evals in CI
666
176
 
667
- Run the suite non-interactively, write JUnit for the CI annotations,
668
- and fail the job on a red gate:
177
+ Write machine-readable output and JUnit annotations, then let a failed
178
+ gate fail the job:
669
179
 
670
180
  ```bash
671
- # CURSOR_API_KEY comes from the CI secret store
181
+ # CURSOR_API_KEY comes from the CI secret store.
672
182
  agent-sdk eval --dir . --json --no-stream \
673
183
  --junit reports/evals.xml \
674
184
  --artifacts reports/evals \
675
185
  > reports/evals.json
676
186
  ```
677
187
 
678
- The exit code follows the
679
- [verdict table](#gates-soft-scores-and-verdicts); `2` means nothing
680
- matched the selection. `--max-concurrency` overrides the project
681
- setting, for example to run lower on a shared runner.
682
-
683
- The JSON on stdout carries the totals and one result per case:
684
-
685
- ```json
686
- {
687
- "ok": true,
688
- "passed": 1,
689
- "failed": 0,
690
- "scored": 0,
691
- "skipped": 0,
692
- "strict": false,
693
- "artifactsDir": "/work/my-agent/reports/evals",
694
- "results": [
695
- {
696
- "id": "readiness",
697
- "verdict": "passed",
698
- "ok": true,
699
- "assertions": [
700
- { "name": "succeeded", "passed": true },
701
- { "name": "calledTool(inspect_pr)", "passed": true }
702
- ],
703
- "sessionId": "ses_123",
704
- "inputs": ["Is https://github.com/acme/checkout/pull/42 ready to approve?"],
705
- "toolCalls": [{ "toolName": "inspect_pr", "isError": false }],
706
- "metrics": {},
707
- "logs": [],
708
- "durationMs": 12340
709
- }
710
- ]
711
- }
712
- ```
713
-
714
- Each result can also include `description`, `finalText`, `tools`,
715
- `error`, `skipReason`, `metadata`, `tags`, and tool `args` and `output`.
716
- A soft miss shows as `"severity": "soft"` with `score` and `threshold`
717
- on the assertion. This shape lets CI report the failed assertion
718
- without parsing terminal text.
188
+ Run `--tag smoke` on pull requests and the full suite on a schedule when
189
+ cost or latency makes every-push coverage impractical. Use soft scores
190
+ for a new probabilistic benchmark until its threshold is trustworthy.
719
191
 
720
- Keep CI green without weakening gates:
192
+ The [CLI reference](/docs/reference/cli.md#eval) covers selectors, hosted
193
+ runs, exit codes, and JSON output.
721
194
 
722
- - Run `--tag smoke` on every push and the full suite on a schedule.
723
- - For probabilistic behavior, use `iterations` and a soft bar instead
724
- of one hard gate.
195
+ ## Keep side effects out of eval sessions
725
196
 
726
- ## Run evals in the playground or on a deployment
197
+ Eval sessions carry `purpose: "eval"`. Check it at the deterministic
198
+ write boundary and return the decision without performing the external
199
+ action:
727
200
 
728
- Start the server, open the playground, and choose **Evals**. Run every
729
- case or one case, watch progress, and open the resulting session trace.
730
- Playground runs target the live server instead of an ephemeral one, so
731
- their sessions appear in the session list. One batch runs at a time.
201
+ ```ts
202
+ async execute({ findings }, ctx) {
203
+ if (ctx.session.purpose === "eval") {
204
+ return {
205
+ posted: false,
206
+ reason: "eval",
207
+ count: findings.length,
208
+ };
209
+ }
732
210
 
733
- ```bash
734
- agent-sdk serve --dir .
211
+ return await postFindings(findings);
212
+ }
735
213
  ```
736
214
 
737
- `--prod` (or `--url`) starts the same server-side batch on the team's
738
- hosted deployment (or the server you name), so results land in that
739
- server's playground history:
215
+ This preserves the tool call in the trajectory, so the eval can still
216
+ assert that the agent chose to post.
740
217
 
741
- ```bash
742
- agent-sdk eval --prod --slug vulnerability-scanner --tag deepsec
743
- # Eval ID: <evalId>
744
- # Playground: <deployment>/playground?view=evals&evalRunId=<evalId>
745
- agent-sdk eval status <evalId> --prod --slug vulnerability-scanner
746
- agent-sdk eval cancel <evalId> --prod --slug vulnerability-scanner
747
- ```
218
+ ## Evals, hooks, or hillclimbing?
748
219
 
749
- The CLI prints the Eval ID as soon as the batch is accepted. Pass
750
- `--no-wait` to return right away and poll with `eval status` later; it
751
- exits `3` while the batch is still running. Hosted history follows
752
- `maxPlaygroundRuns` and the persistence rule under
753
- [Configure eval runs](#configure-eval-runs). The HTTP surface is under
754
- [Playground eval routes](/docs/reference/http-api.md#playground-eval-routes).
220
+ Use an eval to gate one fixed input. Use a
221
+ [hook](/docs/reference/hooks.md) to observe every live session for metrics,
222
+ auditing, or alerts. Use [hillclimbing](/docs/hillclimbing.md) to decide
223
+ which source change improves a fixed set of inputs; every kept
224
+ hillclimb change should add an eval.
755
225
 
756
- ## Keep improvements with regression evals
226
+ ## Gates and scores
757
227
 
758
- Every [hillclimb](/docs/hillclimbing.md) round that keeps a change must
759
- land an eval that would have failed before the change. If you can't
760
- express the improvement as a gate (a `calledTool` shift, a bounded
761
- `maxToolCalls`, an output-shape check), the improvement is unverified,
762
- and it'll regress silently.
228
+ `t.succeeded()`, tool assertions, and deterministic checks are hard
229
+ gates by default. A failed gate fails the case. `.soft()` records a
230
+ measurement without blocking, while `.atLeast(threshold)` records a
231
+ score and marks a miss as `scored`.
232
+
233
+ Start with decisions and output shape: the intended tool, the tempting
234
+ wrong tool, and one structural check. Exact prose is rarely a stable
235
+ contract.
236
+
237
+ ## Keep improvements with regression evals
763
238
 
764
- The rule cuts the other way too: never weaken an existing gate to make
765
- a round pass. That's the freeze line moving, and it turns your
766
- regression suite into a list of checks that no longer protect anything.
239
+ Every kept improvement needs an eval that would have failed before the
240
+ change. A tool-choice gate, bounded `maxToolCalls`, or output-shape check
241
+ turns the improvement into a durable contract.
767
242
 
768
- ## What's next
243
+ Never weaken an existing gate to make a new implementation pass. That
244
+ moves the freeze line instead of proving the change.
769
245
 
770
- Continue with these pages:
246
+ ## Related
771
247
 
772
- - [Hillclimbing](/docs/hillclimbing.md): the loop evals make trustworthy
773
- - [Building agents with agents](/docs/building-with-agents.md): have a
774
- coding agent write the first suite
775
- - [GitHub guide](/docs/guides/github.md): deterministic webhook fixtures
776
- with `github replay`
777
- - [Sessions and streaming](/docs/reference/sessions.md): the events
778
- `t.events` contains
248
+ - [Evals reference](/docs/reference/evals.md): complete authoring and runner
249
+ contracts
250
+ - [`skills/evals/SKILL.md`](https://github.com/cursor/cursor/blob/main/packages/agent-serve/skills/evals/SKILL.md): author and seed
251
+ cases with a coding agent
252
+ - [Hillclimbing](/docs/hillclimbing.md): measure a change and lock the win
253
+ - [Sessions](/docs/reference/sessions.md): events available to trajectory
254
+ assertions