@cursor/july 0.1.111 → 0.1.113

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (441) hide show
  1. package/README.md +4 -0
  2. package/dist/bin/agent-serve.js +20 -4
  3. package/dist/channels/checks.d.ts +10 -0
  4. package/dist/channels/checks.d.ts.map +1 -1
  5. package/dist/channels/github/github-channel.d.ts +29 -19
  6. package/dist/channels/github/github-channel.d.ts.map +1 -1
  7. package/dist/channels/github/github-channel.js +29 -19
  8. package/dist/channels/origin/checks.d.ts +1 -1
  9. package/dist/channels/origin/checks.d.ts.map +1 -1
  10. package/dist/channels/slack/agentic-delivery.d.ts.map +1 -1
  11. package/dist/channels/slack/agentic-delivery.js +2 -0
  12. package/dist/channels/slack/dispatch.d.ts +11 -1
  13. package/dist/channels/slack/dispatch.d.ts.map +1 -1
  14. package/dist/channels/slack/dispatch.js +146 -39
  15. package/dist/channels/slack/placeholder.d.ts +26 -0
  16. package/dist/channels/slack/placeholder.d.ts.map +1 -1
  17. package/dist/channels/slack/placeholder.js +28 -0
  18. package/dist/channels/slack/slack-channel.d.ts.map +1 -1
  19. package/dist/channels/slack/slack-channel.js +9 -16
  20. package/dist/docs/404.html +2 -2
  21. package/dist/docs/assets/{app.Drol6mi6.js → app.CAeK13eM.js} +4 -4
  22. package/dist/docs/assets/building-with-agents.md.BBCx0AUo.js +9 -0
  23. package/dist/docs/assets/building-with-agents.md.BBCx0AUo.lean.js +1 -0
  24. package/dist/docs/assets/chunks/@localSearchIndexroot.Ck9E52Ls.js +1 -0
  25. package/dist/docs/assets/chunks/{VPLocalSearchBox.CdqoBZJX.js → VPLocalSearchBox.C9LbPHod.js} +1 -1
  26. package/dist/docs/assets/chunks/{arc.BVX3ycTn.js → arc.CmMq2zmS.js} +1 -1
  27. package/dist/docs/assets/chunks/{architectureDiagram-Q4EWVU46.CcChnMxX.js → architectureDiagram-Q4EWVU46.CCXB8Uj5.js} +1 -1
  28. package/dist/docs/assets/chunks/{baseUniq.xwtXO-yt.js → baseUniq.CyQo6eLe.js} +1 -1
  29. package/dist/docs/assets/chunks/{blockDiagram-DXYQGD6D.CijZ_taK.js → blockDiagram-DXYQGD6D.JYq6w91N.js} +1 -1
  30. package/dist/docs/assets/chunks/{c4Diagram-AHTNJAMY.0FBvwBvK.js → c4Diagram-AHTNJAMY.BRV8GPJJ.js} +1 -1
  31. package/dist/docs/assets/chunks/channel.BHiYmnZ4.js +1 -0
  32. package/dist/docs/assets/chunks/{chunk-4BX2VUAB.CELMmMDA.js → chunk-4BX2VUAB.Bv4ooYQR.js} +1 -1
  33. package/dist/docs/assets/chunks/{chunk-4TB4RGXK.C4bhFtSm.js → chunk-4TB4RGXK.t4JtKPcj.js} +1 -1
  34. package/dist/docs/assets/chunks/{chunk-55IACEB6.ZnQ9gRRQ.js → chunk-55IACEB6.34lCHj9Y.js} +1 -1
  35. package/dist/docs/assets/chunks/{chunk-EDXVE4YY.DPlmglG-.js → chunk-EDXVE4YY.BSwrPNrt.js} +1 -1
  36. package/dist/docs/assets/chunks/{chunk-FMBD7UC4.BrUk8pff.js → chunk-FMBD7UC4.Beeun-R-.js} +1 -1
  37. package/dist/docs/assets/chunks/{chunk-OYMX7WX6.CcmWIncu.js → chunk-OYMX7WX6.BUUFUcJc.js} +1 -1
  38. package/dist/docs/assets/chunks/{chunk-QZHKN3VN.D65-cs8I.js → chunk-QZHKN3VN.B2XjHzN_.js} +1 -1
  39. package/dist/docs/assets/chunks/{chunk-YZCP3GAM.qoXZpG9F.js → chunk-YZCP3GAM.CLYG8znk.js} +1 -1
  40. package/dist/docs/assets/chunks/classDiagram-6PBFFD2Q.Degh8l90.js +1 -0
  41. package/dist/docs/assets/chunks/classDiagram-v2-HSJHXN6E.Degh8l90.js +1 -0
  42. package/dist/docs/assets/chunks/clone.BIywbczV.js +1 -0
  43. package/dist/docs/assets/chunks/{cose-bilkent-S5V4N54A.8rYtqudO.js → cose-bilkent-S5V4N54A.DVEa6fZp.js} +1 -1
  44. package/dist/docs/assets/chunks/{dagre-KV5264BT.DrRP1fOh.js → dagre-KV5264BT.C9PZQK-S.js} +1 -1
  45. package/dist/docs/assets/chunks/{diagram-5BDNPKRD.DHj_xA2_.js → diagram-5BDNPKRD.DoN0uv3Y.js} +1 -1
  46. package/dist/docs/assets/chunks/{diagram-G4DWMVQ6.Bz_6nAKj.js → diagram-G4DWMVQ6.Czv3duqx.js} +1 -1
  47. package/dist/docs/assets/chunks/{diagram-MMDJMWI5.BWA0xSW9.js → diagram-MMDJMWI5.BinJ5kWb.js} +1 -1
  48. package/dist/docs/assets/chunks/{diagram-TYMM5635.CpTJLNJI.js → diagram-TYMM5635.DW326M4K.js} +1 -1
  49. package/dist/docs/assets/chunks/{erDiagram-SMLLAGMA.-7AWWSrP.js → erDiagram-SMLLAGMA.U2pR_OA7.js} +1 -1
  50. package/dist/docs/assets/chunks/{flowDiagram-DWJPFMVM.BTnQ742_.js → flowDiagram-DWJPFMVM.ByWJXeYK.js} +1 -1
  51. package/dist/docs/assets/chunks/framework.BNw1pucY.js +19 -0
  52. package/dist/docs/assets/chunks/{ganttDiagram-T4ZO3ILL.B5_HiiQ5.js → ganttDiagram-T4ZO3ILL.OquF0Rtg.js} +1 -1
  53. package/dist/docs/assets/chunks/{gitGraphDiagram-UUTBAWPF.CZcNTFZd.js → gitGraphDiagram-UUTBAWPF.Bpn01P7X.js} +1 -1
  54. package/dist/docs/assets/chunks/{graph.V2GLaab4.js → graph.CNRB6ETL.js} +1 -1
  55. package/dist/docs/assets/chunks/{infoDiagram-42DDH7IO.7pZOkCCU.js → infoDiagram-42DDH7IO.CqhknMWi.js} +1 -1
  56. package/dist/docs/assets/chunks/{ishikawaDiagram-UXIWVN3A.DSMo3Qa3.js → ishikawaDiagram-UXIWVN3A.C6xpR2af.js} +1 -1
  57. package/dist/docs/assets/chunks/{journeyDiagram-VCZTEJTY.BNEgWN1S.js → journeyDiagram-VCZTEJTY.Cg5f7oB3.js} +1 -1
  58. package/dist/docs/assets/chunks/{kanban-definition-6JOO6SKY.B-4d4tC7.js → kanban-definition-6JOO6SKY.Cx9YTwlU.js} +1 -1
  59. package/dist/docs/assets/chunks/{layout.Dvbn9nSb.js → layout.ljS-wFtK.js} +1 -1
  60. package/dist/docs/assets/chunks/{linear.D2GM4p4b.js → linear.jSxNrsFC.js} +1 -1
  61. package/dist/docs/assets/chunks/{min.DVBtLK-B.js → min.Cum8AlQw.js} +1 -1
  62. package/dist/docs/assets/chunks/{mindmap-definition-QFDTVHPH.CnJRI55x.js → mindmap-definition-QFDTVHPH.BLiysLpe.js} +1 -1
  63. package/dist/docs/assets/chunks/{pieDiagram-DEJITSTG.CtoaTFlA.js → pieDiagram-DEJITSTG.BoIDyuKF.js} +1 -1
  64. package/dist/docs/assets/chunks/{quadrantDiagram-34T5L4WZ.DtT24_vN.js → quadrantDiagram-34T5L4WZ.DLkpDytR.js} +1 -1
  65. package/dist/docs/assets/chunks/{requirementDiagram-MS252O5E.CjSO4o8f.js → requirementDiagram-MS252O5E.DqTVqSu2.js} +1 -1
  66. package/dist/docs/assets/chunks/{sankeyDiagram-XADWPNL6.-wWiIVWa.js → sankeyDiagram-XADWPNL6.CG_6FF7j.js} +1 -1
  67. package/dist/docs/assets/chunks/{sequenceDiagram-FGHM5R23.DSk8s4gX.js → sequenceDiagram-FGHM5R23.BIp9602K.js} +1 -1
  68. package/dist/docs/assets/chunks/{stateDiagram-FHFEXIEX.BFKAsdkN.js → stateDiagram-FHFEXIEX.COSXsD9I.js} +1 -1
  69. package/dist/docs/assets/chunks/stateDiagram-v2-QKLJ7IA2.qrxrbFsX.js +1 -0
  70. package/dist/docs/assets/chunks/{theme.C0MctGaz.js → theme.CXJ7PNwy.js} +2 -2
  71. package/dist/docs/assets/chunks/{timeline-definition-GMOUNBTQ.TUNJbAFe.js → timeline-definition-GMOUNBTQ.CXdVqkLq.js} +1 -1
  72. package/dist/docs/assets/chunks/{vennDiagram-DHZGUBPP.C-RGNnk4.js → vennDiagram-DHZGUBPP.CZxGuc4r.js} +1 -1
  73. package/dist/docs/assets/chunks/{wardley-RL74JXVD.g5efOWmT.js → wardley-RL74JXVD.3oVgfqQk.js} +1 -1
  74. package/dist/docs/assets/chunks/{wardleyDiagram-NUSXRM2D.DH4zt73Y.js → wardleyDiagram-NUSXRM2D.6_irCgGJ.js} +1 -1
  75. package/dist/docs/assets/chunks/{xychartDiagram-5P7HB3ND.B-3k6bF6.js → xychartDiagram-5P7HB3ND.TRPe92m3.js} +1 -1
  76. package/dist/docs/assets/deployment.md.D2jQZuFx.js +32 -0
  77. package/dist/docs/assets/deployment.md.D2jQZuFx.lean.js +1 -0
  78. package/dist/docs/assets/evals.md.D3Y3Aixt.js +72 -0
  79. package/dist/docs/assets/evals.md.D3Y3Aixt.lean.js +1 -0
  80. package/dist/docs/assets/guides_agent-to-agent.md.CD4T5FIl.js +41 -0
  81. package/dist/docs/assets/guides_agent-to-agent.md.CD4T5FIl.lean.js +1 -0
  82. package/dist/docs/assets/guides_bitbucket.md.mpevW-VP.js +145 -0
  83. package/dist/docs/assets/guides_bitbucket.md.mpevW-VP.lean.js +1 -0
  84. package/dist/docs/assets/guides_cloud-agents.md.Cp1O3u-X.js +15 -0
  85. package/dist/docs/assets/guides_cloud-agents.md.Cp1O3u-X.lean.js +1 -0
  86. package/dist/docs/assets/guides_convert-automation.md.CqEyfP6Y.js +43 -0
  87. package/dist/docs/assets/guides_convert-automation.md.CqEyfP6Y.lean.js +1 -0
  88. package/dist/docs/assets/guides_github.md.BwpBp3ed.js +156 -0
  89. package/dist/docs/assets/guides_github.md.BwpBp3ed.lean.js +1 -0
  90. package/dist/docs/assets/guides_gitlab.md.DaEC3nMk.js +153 -0
  91. package/dist/docs/assets/guides_gitlab.md.DaEC3nMk.lean.js +1 -0
  92. package/dist/docs/assets/guides_grokbot-agents.md.CMhZNdEU.js +16 -0
  93. package/dist/docs/assets/guides_grokbot-agents.md.CMhZNdEU.lean.js +1 -0
  94. package/dist/docs/assets/guides_improve.md.Bnp4F99w.js +22 -0
  95. package/dist/docs/assets/guides_improve.md.Bnp4F99w.lean.js +1 -0
  96. package/dist/docs/assets/guides_jev.md.F5fAkkfN.js +189 -0
  97. package/dist/docs/assets/guides_jev.md.F5fAkkfN.lean.js +1 -0
  98. package/dist/docs/assets/guides_mcp-oauth.md.bSFakfCY.js +50 -0
  99. package/dist/docs/assets/guides_mcp-oauth.md.bSFakfCY.lean.js +1 -0
  100. package/dist/docs/assets/guides_opentelemetry.md.BKDxQmmd.js +35 -0
  101. package/dist/docs/assets/guides_opentelemetry.md.BKDxQmmd.lean.js +1 -0
  102. package/dist/docs/assets/guides_slack.md.Bo96y42E.js +70 -0
  103. package/dist/docs/assets/guides_slack.md.Bo96y42E.lean.js +1 -0
  104. package/dist/docs/assets/guides_webhooks.md.1A72_VEE.js +92 -0
  105. package/dist/docs/assets/guides_webhooks.md.1A72_VEE.lean.js +1 -0
  106. package/dist/docs/assets/hillclimbing.md.D4E1o5Sa.js +7 -0
  107. package/dist/docs/assets/hillclimbing.md.D4E1o5Sa.lean.js +1 -0
  108. package/dist/docs/assets/{index.md.BFVyY2KT.js → index.md.C-t81M5J.js} +2 -2
  109. package/dist/docs/assets/{index.md.BFVyY2KT.lean.js → index.md.C-t81M5J.lean.js} +1 -1
  110. package/dist/docs/assets/{quickstart.md.D3MjSZN-.js → quickstart.md.DAvVhuuU.js} +1 -1
  111. package/dist/docs/assets/{quickstart.md.D3MjSZN-.lean.js → quickstart.md.DAvVhuuU.lean.js} +1 -1
  112. package/dist/docs/assets/{reference_agent-config.md.CfVA-LZJ.js → reference_agent-config.md.DGPyw7ms.js} +1 -1
  113. package/dist/docs/assets/{reference_agent-config.md.CfVA-LZJ.lean.js → reference_agent-config.md.DGPyw7ms.lean.js} +1 -1
  114. package/dist/docs/assets/{reference_artifacts.md.Vf7qyIZ-.js → reference_artifacts.md.Bu_4HmsD.js} +1 -1
  115. package/dist/docs/assets/{reference_artifacts.md.Vf7qyIZ-.lean.js → reference_artifacts.md.Bu_4HmsD.lean.js} +1 -1
  116. package/dist/docs/assets/{reference_channels.md.icqLKcTc.js → reference_channels.md.nFWbzAic.js} +1 -1
  117. package/dist/docs/assets/{reference_channels.md.icqLKcTc.lean.js → reference_channels.md.nFWbzAic.lean.js} +1 -1
  118. package/dist/docs/assets/{reference_cli.md.B2dBL6L8.js → reference_cli.md.DLWDz9ij.js} +3 -1
  119. package/dist/docs/assets/{reference_cli.md.B2dBL6L8.lean.js → reference_cli.md.DLWDz9ij.lean.js} +1 -1
  120. package/dist/docs/assets/{reference_connections.md.Cb3U_c8n.js → reference_connections.md.Je9dMsdd.js} +2 -2
  121. package/dist/docs/assets/{reference_connections.md.Cb3U_c8n.lean.js → reference_connections.md.Je9dMsdd.lean.js} +1 -1
  122. package/dist/docs/assets/reference_evals.md.DNJzM_yf.js +57 -0
  123. package/dist/docs/assets/reference_evals.md.DNJzM_yf.lean.js +1 -0
  124. package/dist/docs/assets/reference_extensions.md.Cv5aLCz_.js +58 -0
  125. package/dist/docs/assets/reference_extensions.md.Cv5aLCz_.lean.js +1 -0
  126. package/dist/docs/assets/{reference_hooks.md.CZuynAxj.js → reference_hooks.md.B7uzNENk.js} +2 -2
  127. package/dist/docs/assets/{reference_hooks.md.CZuynAxj.lean.js → reference_hooks.md.B7uzNENk.lean.js} +1 -1
  128. package/dist/docs/assets/{reference_http-api.md.DTKcYE6L.js → reference_http-api.md.CduHavZ2.js} +1 -1
  129. package/dist/docs/assets/{reference_http-api.md.DTKcYE6L.lean.js → reference_http-api.md.CduHavZ2.lean.js} +1 -1
  130. package/dist/docs/assets/{reference_instructions.md.D7gkckK-.js → reference_instructions.md.CU1My5My.js} +1 -1
  131. package/dist/docs/assets/{reference_instructions.md.D7gkckK-.lean.js → reference_instructions.md.CU1My5My.lean.js} +1 -1
  132. package/dist/docs/assets/{reference_playground.md.D2YExv5K.js → reference_playground.md.Ch2d0Iqi.js} +1 -1
  133. package/dist/docs/assets/{reference_playground.md.D2YExv5K.lean.js → reference_playground.md.Ch2d0Iqi.lean.js} +1 -1
  134. package/dist/docs/assets/{reference_project-layout.md.DPxbUJyt.js → reference_project-layout.md.BGhgpy9V.js} +1 -1
  135. package/dist/docs/assets/{reference_project-layout.md.DPxbUJyt.lean.js → reference_project-layout.md.BGhgpy9V.lean.js} +1 -1
  136. package/dist/docs/assets/{reference_prompt.md.BQ5uAv1F.js → reference_prompt.md.Ccp0R53H.js} +1 -1
  137. package/dist/docs/assets/{reference_prompt.md.BQ5uAv1F.lean.js → reference_prompt.md.Ccp0R53H.lean.js} +1 -1
  138. package/dist/docs/assets/{reference_schedules.md.BasfZWO-.js → reference_schedules.md.B2Nm6FaD.js} +1 -1
  139. package/dist/docs/assets/{reference_schedules.md.BasfZWO-.lean.js → reference_schedules.md.B2Nm6FaD.lean.js} +1 -1
  140. package/dist/docs/assets/{reference_sessions.md.YKvIsWAx.js → reference_sessions.md.1_6Vyv7x.js} +1 -1
  141. package/dist/docs/assets/{reference_sessions.md.YKvIsWAx.lean.js → reference_sessions.md.1_6Vyv7x.lean.js} +1 -1
  142. package/dist/docs/assets/{reference_skills.md.rNgpsGd0.js → reference_skills.md.DjQkRefx.js} +1 -1
  143. package/dist/docs/assets/{reference_skills.md.rNgpsGd0.lean.js → reference_skills.md.DjQkRefx.lean.js} +1 -1
  144. package/dist/docs/assets/{reference_subagents.md.e5qitjJt.js → reference_subagents.md.Dl16gcBj.js} +2 -2
  145. package/dist/docs/assets/{reference_subagents.md.e5qitjJt.lean.js → reference_subagents.md.Dl16gcBj.lean.js} +1 -1
  146. package/dist/docs/assets/{reference_tools.md.BdCO2aHZ.js → reference_tools.md.B1dH1lpa.js} +3 -3
  147. package/dist/docs/assets/{reference_tools.md.BdCO2aHZ.lean.js → reference_tools.md.B1dH1lpa.lean.js} +1 -1
  148. package/dist/docs/assets/{templates_agentic-owners.md.Da_AGDlH.js → templates_agentic-owners.md.9M575F5C.js} +1 -1
  149. package/dist/docs/assets/{templates_agentic-owners.md.Da_AGDlH.lean.js → templates_agentic-owners.md.9M575F5C.lean.js} +1 -1
  150. package/dist/docs/assets/{templates_pr-autofixer.md.DqxocIGh.js → templates_pr-autofixer.md.ws0DDXDy.js} +1 -1
  151. package/dist/docs/assets/{templates_pr-autofixer.md.DqxocIGh.lean.js → templates_pr-autofixer.md.ws0DDXDy.lean.js} +1 -1
  152. package/dist/docs/assets/{templates_security-reviewer.md.Bhnvd8VE.js → templates_security-reviewer.md.KEFYXzfK.js} +1 -1
  153. package/dist/docs/assets/{templates_security-reviewer.md.Bhnvd8VE.lean.js → templates_security-reviewer.md.KEFYXzfK.lean.js} +1 -1
  154. package/dist/docs/assets/templates_thermo-quality-review.md.VNJ_mohX.js +3 -0
  155. package/dist/docs/assets/templates_thermo-quality-review.md.VNJ_mohX.lean.js +1 -0
  156. package/dist/docs/assets/templates_thermo-review.md.Hi3zWOkP.js +3 -0
  157. package/dist/docs/assets/templates_thermo-review.md.Hi3zWOkP.lean.js +1 -0
  158. package/dist/docs/assets/{templates_triage.md.BdBWO9Ic.js → templates_triage.md.CVGe_FG6.js} +1 -1
  159. package/dist/docs/assets/{templates_triage.md.BdBWO9Ic.lean.js → templates_triage.md.CVGe_FG6.lean.js} +1 -1
  160. package/dist/docs/assets/troubleshooting.md.mnfFG2Em.js +1 -0
  161. package/dist/docs/assets/troubleshooting.md.mnfFG2Em.lean.js +1 -0
  162. package/dist/docs/building-with-agents.html +44 -48
  163. package/dist/docs/building-with-agents.md +94 -82
  164. package/dist/docs/deployment.html +59 -78
  165. package/dist/docs/deployment.md +117 -363
  166. package/dist/docs/evals.html +85 -224
  167. package/dist/docs/evals.md +149 -673
  168. package/dist/docs/guides/agent-to-agent.html +74 -47
  169. package/dist/docs/guides/agent-to-agent.md +128 -46
  170. package/dist/docs/guides/bitbucket.html +179 -44
  171. package/dist/docs/guides/bitbucket.md +249 -48
  172. package/dist/docs/guides/cloud-agents.html +46 -40
  173. package/dist/docs/guides/cloud-agents.md +119 -66
  174. package/dist/docs/guides/convert-automation.html +80 -49
  175. package/dist/docs/guides/convert-automation.md +156 -147
  176. package/dist/docs/guides/github.html +184 -92
  177. package/dist/docs/guides/github.md +260 -245
  178. package/dist/docs/guides/gitlab.html +184 -45
  179. package/dist/docs/guides/gitlab.md +249 -50
  180. package/dist/docs/guides/grokbot-agents.html +48 -41
  181. package/dist/docs/guides/grokbot-agents.md +99 -53
  182. package/dist/docs/guides/improve.html +48 -40
  183. package/dist/docs/guides/improve.md +111 -58
  184. package/dist/docs/guides/jev.html +248 -0
  185. package/dist/docs/guides/jev.md +348 -0
  186. package/dist/docs/guides/mcp-oauth.html +74 -52
  187. package/dist/docs/guides/mcp-oauth.md +111 -121
  188. package/dist/docs/guides/opentelemetry.html +63 -54
  189. package/dist/docs/guides/opentelemetry.md +96 -165
  190. package/dist/docs/guides/slack.html +83 -59
  191. package/dist/docs/guides/slack.md +157 -227
  192. package/dist/docs/guides/webhooks.html +97 -230
  193. package/dist/docs/guides/webhooks.md +154 -385
  194. package/dist/docs/hashmap.json +1 -1
  195. package/dist/docs/hillclimbing.html +44 -38
  196. package/dist/docs/hillclimbing.md +101 -55
  197. package/dist/docs/index.html +38 -38
  198. package/dist/docs/index.md +2 -0
  199. package/dist/docs/llms-full.txt +3490 -3146
  200. package/dist/docs/llms.txt +22 -18
  201. package/dist/docs/quickstart.html +37 -37
  202. package/dist/docs/reference/agent-config.html +37 -37
  203. package/dist/docs/reference/artifacts.html +38 -38
  204. package/dist/docs/reference/channels.html +37 -37
  205. package/dist/docs/reference/cli.html +39 -37
  206. package/dist/docs/reference/cli.md +2 -0
  207. package/dist/docs/reference/connections.html +38 -38
  208. package/dist/docs/reference/connections.md +1 -1
  209. package/dist/docs/reference/evals.html +116 -0
  210. package/dist/docs/reference/evals.md +293 -0
  211. package/dist/docs/reference/extensions.html +80 -84
  212. package/dist/docs/reference/extensions.md +131 -202
  213. package/dist/docs/reference/hooks.html +39 -39
  214. package/dist/docs/reference/hooks.md +34 -34
  215. package/dist/docs/reference/http-api.html +37 -37
  216. package/dist/docs/reference/instructions.html +37 -37
  217. package/dist/docs/reference/playground.html +37 -37
  218. package/dist/docs/reference/project-layout.html +37 -37
  219. package/dist/docs/reference/prompt.html +37 -37
  220. package/dist/docs/reference/schedules.html +37 -37
  221. package/dist/docs/reference/sessions.html +37 -37
  222. package/dist/docs/reference/skills.html +37 -37
  223. package/dist/docs/reference/subagents.html +38 -38
  224. package/dist/docs/reference/subagents.md +1 -1
  225. package/dist/docs/reference/tools.html +38 -38
  226. package/dist/docs/reference/tools.md +4 -0
  227. package/dist/docs/templates/agentic-owners.html +38 -38
  228. package/dist/docs/templates/pr-autofixer.html +37 -37
  229. package/dist/docs/templates/security-reviewer.html +38 -38
  230. package/dist/docs/templates/thermo-quality-review.html +62 -0
  231. package/dist/docs/templates/thermo-quality-review.md +75 -0
  232. package/dist/docs/templates/thermo-review.html +62 -0
  233. package/dist/docs/templates/thermo-review.md +74 -0
  234. package/dist/docs/templates/triage.html +37 -37
  235. package/dist/docs/troubleshooting.html +38 -38
  236. package/dist/docs/troubleshooting.md +97 -66
  237. package/dist/extensions/cursor-cloud-agents/skills/handoff.md +3 -7
  238. package/dist/extensions/improve/extension.d.ts +3 -1
  239. package/dist/extensions/improve/extension.d.ts.map +1 -1
  240. package/dist/extensions/improve/extension.js +4 -2
  241. package/dist/extensions/improve/skills/yourself.js +1 -1
  242. package/dist/extensions/jev/extension.d.ts +43 -0
  243. package/dist/extensions/jev/extension.d.ts.map +1 -0
  244. package/dist/extensions/jev/extension.js +47 -0
  245. package/dist/extensions/jev/lib/evaluate.d.ts +101 -0
  246. package/dist/extensions/jev/lib/evaluate.d.ts.map +1 -0
  247. package/dist/extensions/jev/lib/evaluate.js +167 -0
  248. package/dist/extensions/jev/skills/gated-write.md +25 -0
  249. package/dist/extensions/jev/skills/questions.md +33 -0
  250. package/dist/extensions/jev/tools/evaluate.d.ts +4 -0
  251. package/dist/extensions/jev/tools/evaluate.d.ts.map +1 -0
  252. package/dist/extensions/jev/tools/evaluate.js +88 -0
  253. package/dist/extensions.d.ts +1 -1
  254. package/dist/extensions.d.ts.map +1 -1
  255. package/dist/extensions.js +2 -0
  256. package/dist/filesystem.d.ts +46 -2
  257. package/dist/filesystem.d.ts.map +1 -1
  258. package/dist/filesystem.js +149 -102
  259. package/dist/index.d.ts +2 -2
  260. package/dist/index.d.ts.map +1 -1
  261. package/dist/index.js +1 -1
  262. package/dist/internal/advertise-tools.d.ts.map +1 -1
  263. package/dist/internal/advertise-tools.js +6 -0
  264. package/dist/internal/cli-ax.d.ts +6 -0
  265. package/dist/internal/cli-ax.d.ts.map +1 -1
  266. package/dist/internal/cli-ax.js +72 -14
  267. package/dist/internal/continuation-identity.d.ts.map +1 -1
  268. package/dist/internal/continuation-identity.js +1 -0
  269. package/dist/internal/cursor-agent-template.d.ts +1 -1
  270. package/dist/internal/cursor-agent-template.d.ts.map +1 -1
  271. package/dist/internal/cursor-agent-template.js +2 -0
  272. package/dist/internal/discovery/connections.d.ts.map +1 -1
  273. package/dist/internal/discovery/connections.js +18 -0
  274. package/dist/internal/discovery/extensions.d.ts.map +1 -1
  275. package/dist/internal/discovery/extensions.js +8 -4
  276. package/dist/internal/discovery/info.d.ts.map +1 -1
  277. package/dist/internal/discovery/info.js +1 -0
  278. package/dist/internal/filesystem/tools.d.ts.map +1 -1
  279. package/dist/internal/filesystem/tools.js +2 -2
  280. package/dist/internal/filesystem/walk.d.ts +7 -4
  281. package/dist/internal/filesystem/walk.d.ts.map +1 -1
  282. package/dist/internal/filesystem/walk.js +34 -11
  283. package/dist/internal/hosted-admission-adapter.d.ts +3 -0
  284. package/dist/internal/hosted-admission-adapter.d.ts.map +1 -1
  285. package/dist/internal/hosted-delivery-protocol.d.ts +20 -0
  286. package/dist/internal/hosted-delivery-protocol.d.ts.map +1 -1
  287. package/dist/internal/hosted-delivery-protocol.js +51 -1
  288. package/dist/internal/hosted-delivery.d.ts.map +1 -1
  289. package/dist/internal/hosted-delivery.js +23 -25
  290. package/dist/internal/hosted-execution-diag.d.ts +12 -4
  291. package/dist/internal/hosted-execution-diag.d.ts.map +1 -1
  292. package/dist/internal/hosted-execution-diag.js +26 -4
  293. package/dist/internal/hosted-execution-flush.d.ts +1 -0
  294. package/dist/internal/hosted-execution-flush.d.ts.map +1 -1
  295. package/dist/internal/hosted-execution-flush.js +4 -2
  296. package/dist/internal/init-project.d.ts.map +1 -1
  297. package/dist/internal/init-project.js +4 -0
  298. package/dist/internal/server.d.ts.map +1 -1
  299. package/dist/internal/server.js +111 -48
  300. package/dist/internal/session-engine.d.ts +4 -1
  301. package/dist/internal/session-engine.d.ts.map +1 -1
  302. package/dist/internal/session-engine.js +39 -9
  303. package/dist/internal/skill-catalog.d.ts +28 -0
  304. package/dist/internal/skill-catalog.d.ts.map +1 -0
  305. package/dist/internal/skill-catalog.js +44 -0
  306. package/dist/playground/assets/index-DSMAewbx.css +1 -0
  307. package/dist/playground/assets/index-De_lpFxE.js +67 -0
  308. package/dist/playground/index.html +2 -2
  309. package/dist/types.d.ts +32 -3
  310. package/dist/types.d.ts.map +1 -1
  311. package/docs/README.md +2 -0
  312. package/docs/building-with-agents.md +96 -84
  313. package/docs/deployment.md +118 -364
  314. package/docs/evals.md +149 -673
  315. package/docs/guides/agent-to-agent.md +130 -48
  316. package/docs/guides/bitbucket.md +250 -49
  317. package/docs/guides/cloud-agents.md +119 -67
  318. package/docs/guides/convert-automation.md +157 -148
  319. package/docs/guides/github.md +261 -246
  320. package/docs/guides/gitlab.md +250 -51
  321. package/docs/guides/grokbot-agents.md +100 -55
  322. package/docs/guides/improve.md +112 -59
  323. package/docs/guides/jev.md +353 -0
  324. package/docs/guides/mcp-oauth.md +112 -122
  325. package/docs/guides/opentelemetry.md +97 -166
  326. package/docs/guides/slack.md +158 -228
  327. package/docs/guides/webhooks.md +155 -386
  328. package/docs/hillclimbing.md +102 -56
  329. package/docs/reference/cli.md +2 -0
  330. package/docs/reference/connections.md +1 -1
  331. package/docs/reference/evals.md +298 -0
  332. package/docs/reference/extensions.md +132 -203
  333. package/docs/reference/hooks.md +34 -34
  334. package/docs/reference/subagents.md +1 -1
  335. package/docs/reference/tools.md +4 -0
  336. package/docs/templates/thermo-quality-review.md +80 -0
  337. package/docs/templates/thermo-review.md +79 -0
  338. package/docs/troubleshooting.md +98 -67
  339. package/package.json +8 -1
  340. package/skills/github/SKILL.md +21 -13
  341. package/src/bin/agent-serve.ts +24 -3
  342. package/src/channels/checks.ts +8 -0
  343. package/src/channels/github/github-channel.ts +29 -19
  344. package/src/channels/origin/checks.ts +3 -1
  345. package/src/channels/slack/agentic-delivery.ts +2 -0
  346. package/src/channels/slack/dispatch.ts +131 -4
  347. package/src/channels/slack/placeholder.ts +51 -0
  348. package/src/channels/slack/slack-channel.ts +9 -2
  349. package/src/extensions/cursor-cloud-agents/skills/handoff.md +3 -7
  350. package/src/extensions/improve/extension.ts +4 -2
  351. package/src/extensions/improve/skills/yourself.ts +1 -1
  352. package/src/extensions/jev/extension.ts +95 -0
  353. package/src/extensions/jev/lib/evaluate.ts +289 -0
  354. package/src/extensions/jev/skills/gated-write.md +25 -0
  355. package/src/extensions/jev/skills/questions.md +33 -0
  356. package/src/extensions/jev/tools/evaluate.ts +90 -0
  357. package/src/extensions.ts +2 -0
  358. package/src/filesystem.ts +168 -65
  359. package/src/index.ts +3 -0
  360. package/src/internal/advertise-tools.ts +6 -0
  361. package/src/internal/cli-ax.ts +88 -15
  362. package/src/internal/continuation-identity.ts +1 -0
  363. package/src/internal/cursor-agent-template.ts +2 -0
  364. package/src/internal/discovery/connections.ts +21 -0
  365. package/src/internal/discovery/extensions.ts +12 -4
  366. package/src/internal/discovery/info.ts +1 -0
  367. package/src/internal/filesystem/tools.ts +2 -0
  368. package/src/internal/filesystem/walk.ts +60 -15
  369. package/src/internal/hosted-admission-adapter.ts +3 -0
  370. package/src/internal/hosted-delivery-protocol.ts +82 -1
  371. package/src/internal/hosted-delivery.ts +23 -0
  372. package/src/internal/hosted-execution-diag.ts +33 -4
  373. package/src/internal/hosted-execution-flush.ts +4 -0
  374. package/src/internal/init-project.ts +4 -0
  375. package/src/internal/server.ts +130 -53
  376. package/src/internal/session-engine.ts +46 -8
  377. package/src/internal/skill-catalog.ts +69 -0
  378. package/src/types.ts +33 -3
  379. package/templates/thermo-quality-review/README.md +35 -0
  380. package/templates/thermo-quality-review/agent/agent.ts +8 -0
  381. package/templates/thermo-quality-review/agent/channels/github.ts +42 -0
  382. package/templates/thermo-quality-review/agent/instructions.md +43 -0
  383. package/templates/thermo-quality-review/agent/tools/post_findings.ts +70 -0
  384. package/templates/thermo-quality-review/evals/evals.config.ts +5 -0
  385. package/templates/thermo-quality-review/evals/review.eval.ts +62 -0
  386. package/templates/thermo-quality-review/package.json +18 -0
  387. package/templates/thermo-quality-review/tsconfig.json +12 -0
  388. package/templates/thermo-review/README.md +35 -0
  389. package/templates/thermo-review/agent/agent.ts +8 -0
  390. package/templates/thermo-review/agent/channels/github.ts +42 -0
  391. package/templates/thermo-review/agent/instructions.md +40 -0
  392. package/templates/thermo-review/agent/tools/post_findings.ts +70 -0
  393. package/templates/thermo-review/evals/evals.config.ts +5 -0
  394. package/templates/thermo-review/evals/review.eval.ts +53 -0
  395. package/templates/thermo-review/package.json +18 -0
  396. package/templates/thermo-review/tsconfig.json +12 -0
  397. package/dist/docs/assets/building-with-agents.md.CEGVXkmO.js +0 -13
  398. package/dist/docs/assets/building-with-agents.md.CEGVXkmO.lean.js +0 -1
  399. package/dist/docs/assets/chunks/@localSearchIndexroot.DtXk1hy-.js +0 -1
  400. package/dist/docs/assets/chunks/channel.Bkv1N-gK.js +0 -1
  401. package/dist/docs/assets/chunks/classDiagram-6PBFFD2Q.CZDco1o8.js +0 -1
  402. package/dist/docs/assets/chunks/classDiagram-v2-HSJHXN6E.CZDco1o8.js +0 -1
  403. package/dist/docs/assets/chunks/clone.YSt_40_s.js +0 -1
  404. package/dist/docs/assets/chunks/framework.dypDpWZ3.js +0 -19
  405. package/dist/docs/assets/chunks/stateDiagram-v2-QKLJ7IA2.C4CLm481.js +0 -1
  406. package/dist/docs/assets/deployment.md.Dm4Qo3hp.js +0 -51
  407. package/dist/docs/assets/deployment.md.Dm4Qo3hp.lean.js +0 -1
  408. package/dist/docs/assets/evals.md.BLDRt5LH.js +0 -211
  409. package/dist/docs/assets/evals.md.BLDRt5LH.lean.js +0 -1
  410. package/dist/docs/assets/guides_agent-to-agent.md.BI0xclmy.js +0 -14
  411. package/dist/docs/assets/guides_agent-to-agent.md.BI0xclmy.lean.js +0 -1
  412. package/dist/docs/assets/guides_bitbucket.md.CTpCl__f.js +0 -10
  413. package/dist/docs/assets/guides_bitbucket.md.CTpCl__f.lean.js +0 -1
  414. package/dist/docs/assets/guides_cloud-agents.md.lSE_l7lH.js +0 -9
  415. package/dist/docs/assets/guides_cloud-agents.md.lSE_l7lH.lean.js +0 -1
  416. package/dist/docs/assets/guides_convert-automation.md.Ck6Cr68A.js +0 -12
  417. package/dist/docs/assets/guides_convert-automation.md.Ck6Cr68A.lean.js +0 -1
  418. package/dist/docs/assets/guides_github.md.D6ER29dG.js +0 -64
  419. package/dist/docs/assets/guides_github.md.D6ER29dG.lean.js +0 -1
  420. package/dist/docs/assets/guides_gitlab.md.P-TjBnS5.js +0 -14
  421. package/dist/docs/assets/guides_gitlab.md.P-TjBnS5.lean.js +0 -1
  422. package/dist/docs/assets/guides_grokbot-agents.md.WBZIOvkz.js +0 -9
  423. package/dist/docs/assets/guides_grokbot-agents.md.WBZIOvkz.lean.js +0 -1
  424. package/dist/docs/assets/guides_improve.md.BKaDuKKK.js +0 -14
  425. package/dist/docs/assets/guides_improve.md.BKaDuKKK.lean.js +0 -1
  426. package/dist/docs/assets/guides_mcp-oauth.md.DMNMpXtO.js +0 -28
  427. package/dist/docs/assets/guides_mcp-oauth.md.DMNMpXtO.lean.js +0 -1
  428. package/dist/docs/assets/guides_opentelemetry.md._CRfDyzH.js +0 -26
  429. package/dist/docs/assets/guides_opentelemetry.md._CRfDyzH.lean.js +0 -1
  430. package/dist/docs/assets/guides_slack.md.DdT8rmsj.js +0 -46
  431. package/dist/docs/assets/guides_slack.md.DdT8rmsj.lean.js +0 -1
  432. package/dist/docs/assets/guides_webhooks.md.aQW10HRe.js +0 -225
  433. package/dist/docs/assets/guides_webhooks.md.aQW10HRe.lean.js +0 -1
  434. package/dist/docs/assets/hillclimbing.md.Dq4kkVIL.js +0 -1
  435. package/dist/docs/assets/hillclimbing.md.Dq4kkVIL.lean.js +0 -1
  436. package/dist/docs/assets/reference_extensions.md.wlFD3cUR.js +0 -62
  437. package/dist/docs/assets/reference_extensions.md.wlFD3cUR.lean.js +0 -1
  438. package/dist/docs/assets/troubleshooting.md.BcgNoYtJ.js +0 -1
  439. package/dist/docs/assets/troubleshooting.md.BcgNoYtJ.lean.js +0 -1
  440. package/dist/playground/assets/index-BLlKgZtI.css +0 -1
  441. package/dist/playground/assets/index-CDS5p9sR.js +0 -67
package/docs/evals.md CHANGED
@@ -1,783 +1,259 @@
1
1
  ---
2
2
  title: "Evals"
3
- description: "Write defineEval cases that gate an agent's trajectory, run them with agent-sdk eval locally and in CI, and lock every kept improvement with a regression check."
3
+ description: "Protect agent decisions with fixed inputs, trajectory assertions, and regression checks that run locally or in CI."
4
4
  ---
5
5
 
6
6
  # Evals
7
7
 
8
- An eval sends a fixed message to your agent and asserts over the
9
- trajectory it records: the turn completed, the right tool ran with the
10
- right input, the reply has the right shape. Evals are how you know a
11
- prompt tweak helped, a refactor didn't regress the agent, and last
12
- month's fix still holds.
8
+ An eval sends a fixed input to the real agent and checks the resulting
9
+ trajectory: whether the turn succeeded, which tools ran, and what shape
10
+ the answer took. Use evals to protect behavior you already understand,
11
+ not to discover what the prompt should do.
13
12
 
14
- Nothing is mocked. The runner starts (or targets) a real agent server,
15
- drives sessions over the public API, and grades the events it gets
16
- back. The model runs and server tools execute, so
17
- [keep side effects out of eval sessions](#keep-side-effects-out-of-eval-sessions)
18
- before you point an eval at an agent that posts anywhere.
13
+ Evals run real model turns and real tools. Guard external writes before
14
+ running a suite against an agent that can post, merge, or deploy.
19
15
 
20
- ## Evals, hooks, or hillclimbing?
16
+ ## Protect a tool decision
21
17
 
22
- All three read the same session event stream. Pick by the question you
23
- are asking.
24
-
25
- | You want to | Use |
26
- | --- | --- |
27
- | Gate one fixed input's behavior, locally and in CI | Evals (this page) |
28
- | Observe every live session: metrics, audit, alerts | [Hooks](./reference/hooks.md) |
29
- | Improve an agent one measured round at a time | [Hillclimbing](./hillclimbing.md); each kept win lands an eval |
30
-
31
- [Hooks, channel events, or evals?](./reference/hooks.md#hooks-channel-events-or-evals)
32
- has the side-by-side table.
33
-
34
- ### When not to write an eval
35
-
36
- - Test a server tool's own logic with
37
- `agent-sdk call <tool> --dir . --input '{...}'` or a unit test. No
38
- model turn, no credential.
39
- - Explore a prompt with `agent-sdk run --dir . --message "..."` and
40
- read the trajectory. Write the eval once you know which decision to
41
- gate.
42
- - Stop a bad turn while it runs with
43
- [`needsApproval`](./reference/tools.md#gate-a-tool-on-human-approval)
44
- on the tool or `defineResult`. Evals grade
45
- after the fact.
46
-
47
- ## Write your first eval
48
-
49
- Evals live under the project-root `evals/` directory, a sibling of
50
- `agent/`. `agent/evals/` is silently ignored. Discovery loads every
51
- `.eval.ts` (or `.eval.js`) file under `evals/`, plus one config file.
52
-
53
- ```text
54
- my-agent/
55
- agent/
56
- agent.ts
57
- tools/inspect_pr.ts
58
- evals/
59
- evals.config.ts # required to run: maxConcurrency
60
- readiness.eval.ts # id: readiness
61
- prs.eval.ts # cases: prs/checkout, prs/search
62
- ```
18
+ This smoke case asks for a PR verdict, requires the read-only inspection
19
+ tool, and fails if the agent tries to approve. The CLI reports all three
20
+ decisions together instead of stopping at the first miss.
63
21
 
64
- An eval is a single `async test(t)`. You drive the agent with `t.send`
65
- and assert on the recorded run with the same `t`:
22
+ Use the [Jev extension](./guides/jev.md) when host code needs a typed
23
+ choice, score, or boolean before it writes.
66
24
 
67
25
  ```ts
68
26
  // evals/readiness.eval.ts
69
27
  import { defineEval, includes } from "@cursor/july/evals";
70
28
 
71
29
  export default defineEval({
72
- description: "Inspects a PR without approving it.",
73
30
  tags: ["smoke"],
74
- timeoutMs: 120_000,
75
31
  async test(t) {
76
32
  await t.send(
77
33
  "Is https://github.com/acme/checkout/pull/42 ready to approve?"
78
34
  );
35
+
79
36
  t.succeeded();
80
37
  t.calledTool("inspect_pr");
81
38
  t.notCalledTool("approve_pr");
82
- t.check(t.reply, includes(/ready|approve/i));
39
+ t.check(t.reply, includes(/ready|blocked|approve/i));
83
40
  },
84
41
  });
85
42
  ```
86
43
 
44
+ Evals live in the project-root `evals/` directory, beside `agent/`.
45
+ Add the required concurrency config once:
46
+
87
47
  ```ts
88
48
  // evals/evals.config.ts
89
49
  import { defineEvalConfig } from "@cursor/july/evals";
90
50
 
91
- export default defineEvalConfig({ maxConcurrency: 20 });
51
+ export default defineEvalConfig({
52
+ maxConcurrency: 20,
53
+ });
92
54
  ```
93
55
 
94
- Run it under Node 22.13 or newer (never Bun) with a Cursor credential
95
- in place; see [Credentials](#credentials):
96
-
97
56
  ```bash
98
57
  agent-sdk eval --dir . --list
99
- agent-sdk eval --dir . readiness
100
- ```
101
-
102
- ```text
103
- PASS readiness (14.2s) — Inspects a PR without approving it.
104
- ✓ succeeded
105
- ✓ calledTool(inspect_pr)
106
- ✓ notCalledTool(approve_pr)
107
- ✓ check(includes)
108
-
109
- 1 passed, 0 failed, 1 total
110
- artifacts: <project state directory>/evals/2026-09-11T15-02-11-402Z
111
- ```
112
-
113
- Every local run writes each case's assertions, inputs, tool calls, and
114
- `t.log` lines under that artifacts directory. Open
115
- `evals/<case-id>.json` there when a case fails; see
116
- [Where results land](#where-results-land).
117
-
118
- ## Name cases by path
119
-
120
- The file path is the eval's identity, so you don't author an id.
121
- `evals/builds/api.eval.ts` becomes `builds/api`. An `index` filename
122
- collapses to its directory: `evals/builds/index.eval.ts` becomes
123
- `builds`.
124
-
125
- One file can hold several datapoints through `cases`. Provide either
126
- `test` or `cases`, not both. Each case id becomes `<fileId>/<case.id>`:
127
-
128
- ```ts
129
- // evals/prs.eval.ts: prs/checkout, prs/search
130
- export default defineEval({
131
- tags: ["smoke", "prs"],
132
- cases: [
133
- {
134
- id: "checkout",
135
- description: "Checkout PR readiness.",
136
- async test(t) {
137
- await t.send(
138
- "Is https://github.com/acme/checkout/pull/42 ready to approve?"
139
- );
140
- t.succeeded();
141
- t.calledTool("inspect_pr");
142
- },
143
- },
144
- {
145
- id: "search",
146
- async test(t) {
147
- await t.send(
148
- "Check https://github.com/acme/search/pull/7 before approval."
149
- );
150
- t.succeeded();
151
- t.calledTool("inspect_pr");
152
- },
153
- },
154
- ],
155
- });
156
- ```
157
-
158
- Case ids are single path segments, unique within the file. A case can
159
- set its own `description`, `tags`, `timeoutMs`, `iterations`, `judge`,
160
- `reporters`, and `metadata`. A case-level value replaces the file-level
161
- one for that datapoint, except `metadata`, which merges with case keys
162
- winning, and `reporters`, which adds to the file's list. `metadata` is
163
- free-form data carried onto the result and every reporter.
164
-
165
- A file may instead export an array of `defineEval` calls to fan out
166
- over a dataset. Ids are then the file id plus a zero-padded index
167
- (`sql/0000`, `sql/0001`, ...); see [Load a dataset](#load-a-dataset).
168
- Prefer `cases` when datapoints are hand-written and deserve stable
169
- names.
170
-
171
- ### Iterations
172
-
173
- `iterations` (file or case, default `1`, cap `100`) runs a datapoint
174
- repeatedly. Discovery expands `iterations: 3` on case `nyc` to runnable
175
- ids `weather/nyc/1`, `weather/nyc/2`, and `weather/nyc/3`. The filter
176
- `weather/nyc` still selects all three. Each expanded case exposes
177
- `t.iteration` and `t.iterations`.
178
-
179
- `maxConcurrency` counts authored datapoints, not expanded iterations.
180
- Iterations of one datapoint share a concurrency slot and run in
181
- sequence, so a suite of 11 cases with 3 iterations each and
182
- `maxConcurrency: 20` has at most 11 cases in flight.
183
-
184
- ## Drive the agent with `t.send`
185
-
186
- `t.send(message, options?)` runs one turn and waits for it to settle:
187
- complete, park on an approval request, or fail. Several sends in one
188
- case share the session, which is how you write multi-turn evals.
189
-
190
- Each send resolves to a turn result: `message` (the assistant text),
191
- `sessionId`, `events`, `toolCalls` (tool names in order), `ok`, and
192
- `index`. The turn carries the same assertion vocabulary as `t`, scoped
193
- to that turn, so you can grade an intermediate turn before the next
194
- send overwrites `t.reply`. `turn.expectOk()` throws when the turn
195
- failed, for later steps that depend on it.
196
-
197
- Read the whole case with `t.reply` (last assistant text), `t.events`
198
- (every event so far), `t.turns` (settled turns, oldest first), and
199
- `t.sessionId`. `t.signal` aborts when the case hits its timeout; pass
200
- it to your own async work.
201
-
202
- Three options apply on the first send only, because they shape session
203
- creation:
204
-
205
- | Option | Effect |
206
- | --- | --- |
207
- | `workspaceFiles` | `{ path: contents }` seeded into the local session workspace. Prefer this over machine-local paths |
208
- | `workspaceDir` | Absolute harness cwd for the local runtime |
209
- | `cloud` | Per-session cloud options merged over the agent's static `cloud` config. Pin a fixture repo here for cloud evals instead of on the agent's default `cloud.repos`. On the cloud runtime, seeded files reach the agent as described under [Choose a runtime](./reference/agent-config.md#choose-a-runtime) |
210
-
211
- ```ts
212
- await t.send("Review pr/diff.patch and post findings.", {
213
- workspaceFiles: {
214
- "pr/diff.patch": [
215
- "diff --git a/app/routes/search.ts b/app/routes/search.ts",
216
- "+res.send(`<h1>Results for ${req.query.q}</h1>`);",
217
- ].join("\n"),
218
- },
219
- });
220
- ```
221
-
222
- ## Assert over the trajectory
223
-
224
- Assertions record; they never throw. One run reports every failure
225
- instead of dying on the first. Assertions on `t` read the whole run.
226
- Assertions on a turn read only that turn.
227
-
228
- | Gate | Checks |
229
- | --- | --- |
230
- | `t.succeeded()` | the run did not fail and is not parked on an unanswered approval |
231
- | `t.parked()` | the run cleanly parked on an unanswered approval request |
232
- | `t.messageIncludes(token)` | the joined assistant text matches a string or `RegExp` |
233
- | `t.calledTool(name, matcher?)` | a matching call to `name` happened |
234
- | `t.notCalledTool(name)` | no request for `name`, in any lifecycle state |
235
- | `t.loadedSkill(name)` | the agent opened `skills/<name>/SKILL.md` (read, grep, or shell `cat`) |
236
- | `t.toolOrder(names)` | tool requests appear in this relative order; extra calls allowed |
237
- | `t.usedNoTools()` | no tool calls at all |
238
- | `t.maxToolCalls(max)` | at most `max` tool calls |
239
- | `t.noFailedActions()` | no tool call reported an error |
240
- | `t.calledSubagent(name, matcher?)` | a matching subagent delegation happened |
241
- | `t.taggedArtifact(kind?, predicate?)` | at least one [artifact](./reference/artifacts.md) was tagged |
242
- | `t.event(type, matcher?)` | at least one matching [event](./reference/sessions.md#which-events-can-i-stream) of `type` |
243
- | `t.notEvent(type, matcher?)` | no matching event of `type` |
244
- | `t.eventOrder(matchers)` | matching event groups occur in this relative order |
245
- | `t.eventsSatisfy(label, predicate)` | your predicate over the typed event stream |
246
- | `t.check(value, expectation)` | any value, against a [builder](#grade-values-with-expectation-builders) |
247
- | `t.score(name, value)` | a 0-1 score you computed; soft until you add a bar |
248
-
249
- Three more assertions gate and return the matched fact. They stop the
250
- test body when nothing matches, without a duplicate execution error.
251
- `t.requireToolCall(name, matcher?)` returns the call so later code can
252
- read its `input` and `output`. `t.requireInputRequest(filter?)` returns
253
- the single pending approval request. `await t.require(value, expectation)`
254
- does the same for a value check.
255
-
256
- A case with no assertions passes when at least one turn completed. Add
257
- `t.succeeded()` and behavior gates anyway. They make the contract
258
- visible in review.
259
-
260
- ### What good cases assert
261
-
262
- Gate decisions and shape, not prose. Model wording varies run to run.
263
- Tool choice, tool avoidance, and output structure are the stable
264
- contract.
265
-
266
- 1. `t.succeeded()`: always, first.
267
- 2. The tool decision: `calledTool` for the intended path,
268
- `notCalledTool` for the likely wrong alternative. The pair is
269
- stronger than either alone.
270
- 3. Output shape: a regex for the contract (`/ready|blocked/i`, a JSON
271
- marker, a findings-block fence), never exact sentences.
272
- 4. For structured output, parse `t.reply` and check fields with
273
- `matches` or `satisfies` instead of substring-matching JSON.
274
-
275
- The common failure modes: asserting exact phrasing, packing more than
276
- about five gates into one case (split it), and cases that depend on
277
- live external state that drifts (pin the input).
278
-
279
- ### Narrow tool assertions with matchers
280
-
281
- With no matcher, `calledTool` is request-based: a requested call counts
282
- even before its result arrives. A matcher narrows it:
283
-
284
- ```ts
285
- t.calledTool("inspect_pr", { status: "completed" });
286
- t.calledTool("apply_agents", { input: { verdict: "update" } });
287
- t.calledTool("bash", { input: { command: /^gh pr view/ }, count: 1 });
288
- t.calledTool("read_file", {
289
- output: (value) => String(value).includes("TODO"),
290
- });
58
+ agent-sdk eval --dir . --tag smoke
291
59
  ```
292
60
 
293
- `input`, `output`, and `count` accept a literal, a `RegExp`, or a
294
- predicate. Object literals partial-deep-match, so `{ verdict: "update" }`
295
- matches arguments that also carry other keys. `status` is one of
296
- `completed`, `failed`, `pending`, or `rejected` (a human denied the
297
- approval). `calledSubagent` takes `{ output, status, count, callId }`.
298
- `event`, `notEvent`, and `eventOrder` take `{ data, count }`.
299
-
300
- ### Grade values with expectation builders
61
+ ## Check answer shape or quality
301
62
 
302
- `t.check(value, expectation)` grades any value: `t.reply`, a parsed
303
- JSON field, a tool's output.
304
-
305
- | Builder | Checks | Severity |
306
- | --- | --- | --- |
307
- | `includes(string \| RegExp)` | substring or match; structured values are stringified first | gate |
308
- | `equals(value)` | deep equality | gate |
309
- | `matches(schema)` | a Standard Schema (Zod, Valibot, ...) or anything with `safeParse` | gate |
310
- | `similarity(expected)` | normalized text similarity, 0-1 | soft |
311
- | `satisfies(predicate, label)` | your predicate; `label` is the failure detail | gate |
63
+ Prefer a deterministic shape check when the output has a contract. The
64
+ failure tells you which field was wrong, and the case remains stable
65
+ when the model changes its wording.
312
66
 
313
67
  ```ts
314
- import { matches, satisfies } from "@cursor/july/evals";
68
+ import { matches } from "@cursor/july/evals";
315
69
  import { z } from "zod";
316
70
 
317
71
  const verdict = JSON.parse(t.reply ?? "{}");
318
72
  t.check(
319
73
  verdict,
320
- matches(z.object({ ready: z.boolean(), blockers: z.array(z.string()) }))
321
- );
322
- t.check(
323
- verdict.blockers.length,
324
- satisfies((n) => (n as number) <= 3, "at most 3 blockers")
74
+ matches(
75
+ z.object({
76
+ ready: z.boolean(),
77
+ blockers: z.array(z.string()),
78
+ })
79
+ )
325
80
  );
326
81
  ```
327
82
 
328
- `normalizedSimilarity(actual, expected)` returns the same 0-1 score as
329
- `similarity`, for use with `t.score`.
330
-
331
- ### Record without gating
332
-
333
- - `t.metric(name, value)` records a structured score or label. It shows
334
- on the CLI result, the playground case card, JUnit output, and
335
- artifacts.
336
- - `t.log(message)` records a debug line, streamed under `--verbose`.
337
- - `t.skip(reason)` ends the case as skipped. Skipped cases report
338
- separately and never change the exit code. Call it before sending
339
- messages.
340
-
341
- ## Gates, soft scores, and verdicts
342
-
343
- Every assertion returns a handle, so severity rides on the assertion
344
- instead of a separate thresholds map:
83
+ Use a judge when correctness depends on meaning that a schema, regex, or
84
+ known value cannot capture:
345
85
 
346
86
  ```ts
347
- t.succeeded(); // gate (default)
348
- t.calledTool("get_weather").soft(); // tracked, never fails
349
- t.check(t.reply, similarity("Sunny, 72°F")).atLeast(0.8); // soft, with a bar
350
- t.judge.closedQA("cites a source").gate(0.8); // promoted to a gate
87
+ t.judge
88
+ .factuality("The lint step failed on src/sidebar.ts.")
89
+ .atLeast(0.8);
351
90
  ```
352
91
 
353
- - `.gate(threshold?)` is hard. A miss fails the case and `eval` exits 1.
354
- - `.soft(threshold?)` is tracked. With no threshold it never fails.
355
- - `.atLeast(threshold)` is soft with a bar. A miss marks the case
356
- `scored`.
357
-
358
- Each case ends with one verdict:
92
+ Judge checks are tracked scores until you give them a hard gate.
93
+ Configure the judge model in `evals.config.ts`. The
94
+ [Evals reference](./reference/evals.md#judges) owns grader options and
95
+ severity rules.
359
96
 
360
- | Verdict | Meaning | Exit code |
361
- | --- | --- | --- |
362
- | `passed` | every gate passed and no soft bar was missed | 0 |
363
- | `failed` | a gate failed, or the test body threw | 1 |
364
- | `scored` | only soft bars were missed | 0, or 1 under `--strict` |
365
- | `skipped` | `t.skip(reason)`, or a judge with no credentials | 0 |
97
+ ## Test a conversation or approval
366
98
 
367
- The CLI prints soft misses as `~` and gate misses as `✗`. Start a new
368
- benchmark with `t.score("recall", recall)` and `.atLeast()` so its
369
- number reports for a while without blocking merges. Add `--strict`
370
- once the bars are trustworthy.
371
-
372
- ## Judge free-form output
373
-
374
- When wording matters and no regex captures it, `t.judge` grades with an
375
- LLM. The graders are `factuality(expected)`, `summarizes(expected)`,
376
- `closedQA(criteria)`, and `sql(expected)`. Each scores `t.reply` by
377
- default; pass `{ on }` to grade another value.
99
+ Several `t.send` calls in one case share a session. Assertions on a
100
+ returned turn inspect only that turn, while assertions on `t` inspect
101
+ the whole conversation.
378
102
 
379
103
  ```ts
380
- const summary = await t.send("Why did CI fail on PR 42?");
381
- t.judge.factuality("The lint step failed on src/sidebar.ts.").atLeast(0.7);
382
- t.judge.closedQA("names the failing step", { on: summary.message }).gate(1);
383
- ```
384
-
385
- Judge assertions are soft by default, so a judge never fails a build
386
- until you give it a bar with `.atLeast()` or promote it with `.gate()`.
387
- The recorded detail names the choice the judge made and its rationale.
388
-
389
- The judge model comes from `defineEvalConfig({ judge })`,
390
- `defineEval({ judge })`, a case-level `judge`, or a per-call
391
- `{ model }`. The nearest one wins. A judge call with no model
392
- configured fails the case. A judge that cannot reach a model (no
393
- credential) ends the case as `skipped`, unless a deterministic gate
394
- already failed.
104
+ const first = await t.send("Inspect PR 42.");
105
+ first.calledTool("inspect_pr");
395
106
 
396
- For a domain-specific judge whose verdict is not a single score,
397
- `t.judge.model(prompt)` sends a raw prompt to the same model and
398
- returns the reply. Record the parsed result with `t.score` or
399
- `t.check`. Anything derived from the agent under test is untrusted
400
- input to your prompt: wrap it with `fenceUntrusted` and include
401
- `EVAL_JUDGE_INJECTION_GUARD`, as the built-in graders do.
402
-
403
- ```ts
404
- import {
405
- EVAL_JUDGE_INJECTION_GUARD,
406
- fenceUntrusted,
407
- } from "@cursor/july/evals";
408
-
409
- const gold = ["XSS in search.ts", "open redirect in login.ts"];
410
- const reply = await t.judge.model(
411
- [
412
- "For each GOLD finding, answer whether SUBMISSION reports it.",
413
- "Reply with one line per finding: <index> YES|NO.",
414
- EVAL_JUDGE_INJECTION_GUARD,
415
- fenceUntrusted("GOLD", gold.map((g, i) => `${i + 1}. ${g}`).join("\n")),
416
- fenceUntrusted("SUBMISSION", t.reply ?? ""),
417
- ].join("\n\n")
418
- );
419
- const hits = reply.match(/\bYES\b/g)?.length ?? 0;
420
- t.score("recall", hits / gold.length).atLeast(0.5);
107
+ const followUp = await t.send("Now summarize only the blockers.");
108
+ followUp.usedNoTools();
109
+ t.check(followUp.message, includes(/blocker/i));
110
+ t.succeeded();
421
111
  ```
422
112
 
423
- ## Keep side effects out of eval sessions
424
-
425
- Eval sessions run the real agent, tools included. A reviewer that
426
- comments on GitHub or posts to Slack will do so from an eval unless
427
- the tool checks the session's purpose. Eval sessions carry
428
- `purpose: "eval"`; live traffic carries `"live"`. Branch on it in the
429
- tool, hook, or result handler that actuates:
113
+ For a tool that requires approval, the expected result is a parked turn
114
+ instead of a completed one:
430
115
 
431
116
  ```ts
432
- // agent/tools/post_findings.ts
433
- async execute({ findings }, ctx) {
434
- if (ctx.session.purpose === "eval") {
435
- return { posted: false, reason: "eval", count: findings.length };
436
- }
437
- // post the review
438
- }
439
- ```
440
-
441
- Return a shaped result instead of throwing, so the eval can still
442
- assert `t.calledTool("post_findings", { input: ... })` on the decision.
443
- The same check belongs in [hooks](./reference/hooks.md) that meter or
444
- page and in `defineResult` commits.
445
-
446
- ## Worked examples
447
-
448
- ### Multi-turn: grade each turn
449
-
450
- ```ts
451
- // evals/intro.eval.ts
452
- import { defineEval, includes, satisfies } from "@cursor/july/evals";
453
-
454
- export default defineEval({
455
- description: "Introduces itself once; a repeat mention gets a short ack.",
456
- async test(t) {
457
- const intro = await t.send("Meet Jenny! @Jenny introduce yourself.");
458
- intro.expectOk();
459
- t.check(intro.message, includes(/jenny/i));
460
-
461
- const repeat = await t.send("Meet, @Jenny!");
462
- t.succeeded();
463
- repeat.usedNoTools();
464
- t.check(
465
- repeat.message,
466
- satisfies((r) => (r as string).trim().length <= 280, "short ack")
467
- );
468
- t.check(
469
- repeat.message,
470
- satisfies((r) => !/what i can do/i.test(r as string), "no second intro")
471
- );
472
- },
117
+ await t.send("Apply the approved policy update.");
118
+ t.parked();
119
+ t.calledTool("apply_policy", {
120
+ input: { verdict: "update" },
473
121
  });
474
122
  ```
475
123
 
476
- `t.succeeded()` grades the whole session. `repeat.usedNoTools()` and the
477
- checks on `repeat.message` read only the second turn, even though
478
- `t.reply` now holds its text.
124
+ Pair an approval case with a safe branch that calls
125
+ `t.notCalledTool("apply_policy")`, so both sides of the decision stay
126
+ protected.
479
127
 
480
- ### Approvals: assert the parked decision
128
+ ## Pin realistic inputs
481
129
 
482
- For a tool with `needsApproval`, the turn parks instead of finishing.
483
- Gate on `t.parked()` and on the arguments the model chose:
130
+ A useful eval changes only when the agent changes. Freeze canonical chat
131
+ prompts, seed workspace evidence directly, and save webhook payloads
132
+ instead of depending on a developer's checkout or live external state.
484
133
 
485
134
  ```ts
486
- // evals/agents.eval.ts (one case; RULE and SLACK are fixture strings)
487
- {
488
- id: "update-rule",
489
- description: "A repeated billing rule parks the AGENTS.md write.",
490
- async test(t) {
491
- await t.send("Weekly AGENTS.md review. Read week/ and call apply_agents once.", {
492
- workspaceFiles: {
493
- "week/prs.md": RULE,
494
- "week/slack.md": SLACK,
495
- "week/tree/AGENTS.md.txt": "# API\n\nKeep handlers thin.\n",
496
- },
497
- });
498
- t.parked();
499
- t.calledTool("apply_agents", { input: { verdict: "update" } });
135
+ await t.send("Review pr/diff.patch and report blockers.", {
136
+ workspaceFiles: {
137
+ "pr/diff.patch": [
138
+ "diff --git a/app/routes/search.ts b/app/routes/search.ts",
139
+ "+res.send(`<h1>Results for ${req.query.q}</h1>`);",
140
+ ].join("\n"),
500
141
  },
501
- },
502
- ```
503
-
504
- `t.parked()` and `t.succeeded()` are exclusive: a parked run is a clean
505
- stop on an unanswered approval, not a completed one. Pair the parked
506
- case with a sibling that expects `verdict: "skip"` and `t.succeeded()`,
507
- so both branches stay pinned.
508
-
509
- ## Pin fixtures
510
-
511
- A fixed input is what makes an eval repeatable. Pick the fixture by the
512
- surface under test.
513
-
514
- | Agent surface | Fixture |
515
- | --- | --- |
516
- | Chat or domain assistant | One canonical prompt string, chosen once and frozen |
517
- | Tool-heavy | Run `agent-sdk call <tool>` first to pin what the tool returns, then freeze the prompt that triggers it |
518
- | GitHub webhook | `agent-sdk github replay <pr> --events '*' --dry-run --out fixtures/github` snapshots real payloads for offline replay ([GitHub guide](./guides/github.md#test-with-github-replay)) |
519
- | PR reviewer with host preparation | Diff, metadata, and gold labels pinned to commit SHAs; keep any live PR matrix small |
520
- | Workspace-dependent | `workspaceFiles` on the first `t.send`, never developer-machine paths |
521
-
522
- Tag the fast, reliably passing core `smoke` and run `--tag smoke` in
523
- the inner loop. Leave slow or drift-prone cases untagged for explicit
524
- runs.
525
-
526
- ### Materialize API-backed fixtures
527
-
528
- An input that only points at external data (a pull request URL, a
529
- snapshot id, a pair of commit SHAs) is not self-contained. Fetch it once
530
- and commit the rendered fixture before you expand the suite:
531
-
532
- 1. Save the diff, metadata, and labels under `fixtures/` at pinned
533
- revisions.
534
- 2. Seed those files with `workspaceFiles`, or read them from the
535
- fixture directory.
536
- 3. Assert decisions and output shape against the saved evidence.
537
- 4. Keep a small `smoke` subset for any remaining live checks.
538
-
539
- `maxConcurrency` limits parallel datapoints, not the model or API
540
- fan-out inside one datapoint. Materialized fixtures keep a large suite
541
- from exhausting provider and GitHub rate limits.
542
-
543
- ### Load a dataset
544
-
545
- Read committed fixtures with `loadJson`, `loadJsonl`, and `loadYaml`
546
- from `@cursor/july/evals/loaders`. Relative paths resolve against the
547
- project root the runner discovered, not the cwd the CLI ran from. Eval
548
- files are ES modules, so top-level `await` can load a dataset and fan
549
- one file out over it:
550
-
551
- ```ts
552
- // evals/sql.eval.ts: sql/0000, sql/0001, ...
553
- import { defineEval, equals } from "@cursor/july/evals";
554
- import { loadYaml } from "@cursor/july/evals/loaders";
555
-
556
- const rows = await loadYaml<{ task: string; prompt: string; sql: string }[]>(
557
- "evals/data/cases.yaml"
558
- );
559
-
560
- export default rows.map((row) =>
561
- defineEval({
562
- description: row.task,
563
- async test(t) {
564
- await t.send(row.prompt);
565
- t.succeeded();
566
- t.check(t.reply, equals(row.sql));
567
- },
568
- })
569
- );
570
- ```
571
-
572
- ## Configure eval runs
573
-
574
- `evals/evals.config.ts` (or `.js`) holds project-wide defaults. It must
575
- set `maxConcurrency`. Each case issues real model requests, so
576
- concurrency is hard-capped at 200; the templates use 10.
577
- `eval --list` works without the file. Running a case does not.
578
-
579
- ```ts
580
- import { defineEvalConfig } from "@cursor/july/evals";
581
-
582
- export default defineEvalConfig({
583
- maxConcurrency: 20,
584
- timeoutMs: 180_000,
585
- judge: { model: "gpt-5.4-mini" },
586
142
  });
587
143
  ```
588
144
 
589
- | Option | Default | Meaning |
590
- | --- | --- | --- |
591
- | `maxConcurrency` | required | Datapoints in flight at once, 1-200 |
592
- | `timeoutMs` | `180_000` | Per-case timeout. Precedence: case or file `timeoutMs`, then `--timeout-ms`, then this value |
593
- | `judge` | unset | Default judge model for `t.judge.*` |
594
- | `reporters` | unset | Destinations that observe every case; `--skip-report` suppresses them |
595
- | `maxPlaygroundRuns` | `20` | Batches kept in playground history (cap 500). Counts `--prod` and `--url` batches, not the default local run |
596
-
597
- Reporters ship results somewhere; the runner still does the grading.
598
- `JUnit({ filePath, suiteName? })` writes JUnit XML and
599
- `Artifacts({ dir })` writes per-case files, both from
600
- `@cursor/july/evals/reporters`. A custom reporter is an object with any
601
- of `onRunStart`, `onEvalComplete`, and `onRunComplete`. A reporter that
602
- throws is logged and never fails the run. CI usually attaches the
603
- built-in two with `--junit` and `--artifacts` instead of `reporters`,
604
- so output paths stay with the pipeline, not the eval author.
605
-
606
- Playground batches survive restarts when the project configures
607
- durable storage. Otherwise they live in process
608
- memory until `serve` exits.
609
-
610
- ## Run evals from the CLI
145
+ For GitHub agents, snapshot a PR into committed fixtures:
611
146
 
612
147
  ```bash
613
- agent-sdk eval --dir . --list # discover only
614
- agent-sdk eval --dir . # run all
615
- agent-sdk eval --dir . builds/checkout # one datapoint
616
- agent-sdk eval --dir . builds search # several ids or prefixes
617
- agent-sdk eval --dir . --tag smoke --tag pull-request # any matching tag
618
- agent-sdk eval --dir . --verbose # t.log lines + reply snippets
619
- agent-sdk eval --dir . --timeout-ms 300000 # override the per-case timeout
148
+ agent-sdk github replay acme/checkout#42 \
149
+ --events '*' \
150
+ --dry-run \
151
+ --out fixtures/github
620
152
  ```
621
153
 
622
- Id filters use OR semantics. Each filter selects an exact id and its
623
- descendants: `builds` selects `builds`, `builds/checkout`, and every
624
- other case below that path. Repeated tags also use OR. With both ids
625
- and tags, a case must match both groups.
154
+ Keep a small, fast `smoke` set for every change. Put larger datasets and
155
+ drift-prone live cases behind an explicit selection. See
156
+ [Datasets and fixtures](./reference/evals.md#datasets-and-fixtures) for
157
+ loaders and case expansion.
626
158
 
627
- By default `eval` boots a throwaway server with its own state root, so
628
- cases don't inherit your checkout's `AGENTS.md` and session state stays
629
- out of the project. Artifacts still land in the project state
630
- directory; see [Where results land](#where-results-land). `--slug`
631
- picks the target in a multi-agent directory.
159
+ ## Run and debug a regression suite
632
160
 
633
- `--url` runs the batch on a running server instead, the same way
634
- `--prod` does: that server discovers its own `evals/`, results land in
635
- its playground history, and the local-only flags (`--junit`,
636
- `--artifacts`, `--max-concurrency`, `--skip-report`) do not apply. See
637
- [Run evals in the playground or on a deployment](#run-evals-in-the-playground-or-on-a-deployment).
161
+ Run one case while iterating, then the smoke suite before review. A
162
+ failure artifact records the input, assertions, tool calls, and final
163
+ text, making it the first place to look before changing the prompt.
638
164
 
639
165
  ```bash
640
- agent-sdk eval --url http://127.0.0.1:3000/weather-agent \
641
- --bearer-token "$AGENT_TOKEN"
166
+ agent-sdk eval --dir . readiness
167
+ agent-sdk eval --dir . --tag smoke
168
+ agent-sdk eval --dir . --json
642
169
  ```
643
170
 
644
- See [CLI: eval](./reference/cli.md#eval) for every flag.
645
-
646
- ### Credentials
647
-
648
- Model turns need a
649
- [Cursor credential](./reference/cli.md#environment-variables):
650
- `CURSOR_API_KEY`, `CURSOR_SERVICE_ACCOUNT_KEY`, or a saved
651
- `agent-sdk login`. The judge uses the same one. `eval --list` needs
652
- none.
171
+ By default, results land under `evals/<stamp>/` in the project state
172
+ directory. Open `evals/<case-id>.json` inside that run directory for the
173
+ failed case. This artifact location is independent of `--state-root`;
174
+ pass `--artifacts <dir>` to choose another destination.
653
175
 
654
- ### Where results land
655
-
656
- Every local run writes artifacts to a timestamped directory under
657
- `evals/` in the project state directory, whatever `--state-root` says.
658
- `--artifacts <dir>` chooses the path and `--no-artifacts` skips them.
659
- The directory holds `summary.json`,
660
- `results.jsonl`, and `evals/<case-id>.json` with every assertion, the
661
- inputs, tool calls with arguments and output, the final text, and
662
- `t.log` lines. Start there when a case fails. `--out <file>` also
663
- writes the full results JSON to a path of your choice.
664
-
665
- The artifact does not include the session's event stream. Pass
666
- `--state-root <path>` to keep the ephemeral server's
667
- [session data](./reference/sessions.md#where-does-the-agent-sdk-store-session-data)
668
- on disk when you need the raw events.
176
+ Model turns need a credential. Resolution checks `CURSOR_API_KEY`,
177
+ `CURSOR_API_KEY_FILE`, `CURSOR_SERVICE_ACCOUNT_KEY`, then a saved
178
+ `agent-sdk login`. `eval --list` only discovers cases and needs none.
669
179
 
670
180
  ## Run evals in CI
671
181
 
672
- Run the suite non-interactively, write JUnit for the CI annotations,
673
- and fail the job on a red gate:
182
+ Write machine-readable output and JUnit annotations, then let a failed
183
+ gate fail the job:
674
184
 
675
185
  ```bash
676
- # CURSOR_API_KEY comes from the CI secret store
186
+ # CURSOR_API_KEY comes from the CI secret store.
677
187
  agent-sdk eval --dir . --json --no-stream \
678
188
  --junit reports/evals.xml \
679
189
  --artifacts reports/evals \
680
190
  > reports/evals.json
681
191
  ```
682
192
 
683
- The exit code follows the
684
- [verdict table](#gates-soft-scores-and-verdicts); `2` means nothing
685
- matched the selection. `--max-concurrency` overrides the project
686
- setting, for example to run lower on a shared runner.
687
-
688
- The JSON on stdout carries the totals and one result per case:
689
-
690
- ```json
691
- {
692
- "ok": true,
693
- "passed": 1,
694
- "failed": 0,
695
- "scored": 0,
696
- "skipped": 0,
697
- "strict": false,
698
- "artifactsDir": "/work/my-agent/reports/evals",
699
- "results": [
700
- {
701
- "id": "readiness",
702
- "verdict": "passed",
703
- "ok": true,
704
- "assertions": [
705
- { "name": "succeeded", "passed": true },
706
- { "name": "calledTool(inspect_pr)", "passed": true }
707
- ],
708
- "sessionId": "ses_123",
709
- "inputs": ["Is https://github.com/acme/checkout/pull/42 ready to approve?"],
710
- "toolCalls": [{ "toolName": "inspect_pr", "isError": false }],
711
- "metrics": {},
712
- "logs": [],
713
- "durationMs": 12340
714
- }
715
- ]
716
- }
717
- ```
193
+ Run `--tag smoke` on pull requests and the full suite on a schedule when
194
+ cost or latency makes every-push coverage impractical. Use soft scores
195
+ for a new probabilistic benchmark until its threshold is trustworthy.
718
196
 
719
- Each result can also include `description`, `finalText`, `tools`,
720
- `error`, `skipReason`, `metadata`, `tags`, and tool `args` and `output`.
721
- A soft miss shows as `"severity": "soft"` with `score` and `threshold`
722
- on the assertion. This shape lets CI report the failed assertion
723
- without parsing terminal text.
197
+ The [CLI reference](./reference/cli.md#eval) covers selectors, hosted
198
+ runs, exit codes, and JSON output.
724
199
 
725
- Keep CI green without weakening gates:
726
-
727
- - Run `--tag smoke` on every push and the full suite on a schedule.
728
- - For probabilistic behavior, use `iterations` and a soft bar instead
729
- of one hard gate.
200
+ ## Keep side effects out of eval sessions
730
201
 
731
- ## Run evals in the playground or on a deployment
202
+ Eval sessions carry `purpose: "eval"`. Check it at the deterministic
203
+ write boundary and return the decision without performing the external
204
+ action:
732
205
 
733
- Start the server, open the playground, and choose **Evals**. Run every
734
- case or one case, watch progress, and open the resulting session trace.
735
- Playground runs target the live server instead of an ephemeral one, so
736
- their sessions appear in the session list. One batch runs at a time.
206
+ ```ts
207
+ async execute({ findings }, ctx) {
208
+ if (ctx.session.purpose === "eval") {
209
+ return {
210
+ posted: false,
211
+ reason: "eval",
212
+ count: findings.length,
213
+ };
214
+ }
737
215
 
738
- ```bash
739
- agent-sdk serve --dir .
216
+ return await postFindings(findings);
217
+ }
740
218
  ```
741
219
 
742
- `--prod` (or `--url`) starts the same server-side batch on the team's
743
- hosted deployment (or the server you name), so results land in that
744
- server's playground history:
220
+ This preserves the tool call in the trajectory, so the eval can still
221
+ assert that the agent chose to post.
745
222
 
746
- ```bash
747
- agent-sdk eval --prod --slug vulnerability-scanner --tag deepsec
748
- # Eval ID: <evalId>
749
- # Playground: <deployment>/playground?view=evals&evalRunId=<evalId>
750
- agent-sdk eval status <evalId> --prod --slug vulnerability-scanner
751
- agent-sdk eval cancel <evalId> --prod --slug vulnerability-scanner
752
- ```
223
+ ## Evals, hooks, or hillclimbing?
753
224
 
754
- The CLI prints the Eval ID as soon as the batch is accepted. Pass
755
- `--no-wait` to return right away and poll with `eval status` later; it
756
- exits `3` while the batch is still running. Hosted history follows
757
- `maxPlaygroundRuns` and the persistence rule under
758
- [Configure eval runs](#configure-eval-runs). The HTTP surface is under
759
- [Playground eval routes](./reference/http-api.md#playground-eval-routes).
225
+ Use an eval to gate one fixed input. Use a
226
+ [hook](./reference/hooks.md) to observe every live session for metrics,
227
+ auditing, or alerts. Use [hillclimbing](./hillclimbing.md) to decide
228
+ which source change improves a fixed set of inputs; every kept
229
+ hillclimb change should add an eval.
760
230
 
761
- ## Keep improvements with regression evals
231
+ ## Gates and scores
762
232
 
763
- Every [hillclimb](./hillclimbing.md) round that keeps a change must
764
- land an eval that would have failed before the change. If you can't
765
- express the improvement as a gate (a `calledTool` shift, a bounded
766
- `maxToolCalls`, an output-shape check), the improvement is unverified,
767
- and it'll regress silently.
233
+ `t.succeeded()`, tool assertions, and deterministic checks are hard
234
+ gates by default. A failed gate fails the case. `.soft()` records a
235
+ measurement without blocking, while `.atLeast(threshold)` records a
236
+ score and marks a miss as `scored`.
237
+
238
+ Start with decisions and output shape: the intended tool, the tempting
239
+ wrong tool, and one structural check. Exact prose is rarely a stable
240
+ contract.
241
+
242
+ ## Keep improvements with regression evals
768
243
 
769
- The rule cuts the other way too: never weaken an existing gate to make
770
- a round pass. That's the freeze line moving, and it turns your
771
- regression suite into a list of checks that no longer protect anything.
244
+ Every kept improvement needs an eval that would have failed before the
245
+ change. A tool-choice gate, bounded `maxToolCalls`, or output-shape check
246
+ turns the improvement into a durable contract.
772
247
 
773
- ## What's next
248
+ Never weaken an existing gate to make a new implementation pass. That
249
+ moves the freeze line instead of proving the change.
774
250
 
775
- Continue with these pages:
251
+ ## Related
776
252
 
777
- - [Hillclimbing](./hillclimbing.md): the loop evals make trustworthy
778
- - [Building agents with agents](./building-with-agents.md): have a
779
- coding agent write the first suite
780
- - [GitHub guide](./guides/github.md): deterministic webhook fixtures
781
- with `github replay`
782
- - [Sessions and streaming](./reference/sessions.md): the events
783
- `t.events` contains
253
+ - [Evals reference](./reference/evals.md): complete authoring and runner
254
+ contracts
255
+ - [`skills/evals/SKILL.md`](../skills/evals/SKILL.md): author and seed
256
+ cases with a coding agent
257
+ - [Hillclimbing](./hillclimbing.md): measure a change and lock the win
258
+ - [Sessions](./reference/sessions.md): events available to trajectory
259
+ assertions