@opensearch-project/agent-health 0.3.0 → 0.5.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (518) hide show
  1. package/README.md +77 -6
  2. package/cli/dist/index.js +10072 -4502
  3. package/deployment/cloudformation/agent-health-observability.yaml +762 -0
  4. package/dist/assets/index-CCQRDlO0.js +243 -0
  5. package/dist/assets/index-CNHQVbcj.css +1 -0
  6. package/dist/index.html +2 -2
  7. package/docs/ARCHITECTURE.md +450 -0
  8. package/docs/BACKEND_JOB_QUEUE.md +405 -0
  9. package/docs/CLAUDE_CODE_TELEMETRY.md +283 -0
  10. package/docs/CLI.md +431 -0
  11. package/docs/CODING_AGENT_ANALYTICS.md +298 -0
  12. package/docs/CONFIGURATION.md +388 -0
  13. package/docs/CONNECTORS.md +536 -0
  14. package/docs/INSTRUMENT_WITH_OTEL.md +390 -0
  15. package/docs/ML-COMMONS-SETUP.md +289 -0
  16. package/docs/NPX_PACKAGING.md +195 -0
  17. package/docs/PERFORMANCE-MONITORING.md +200 -0
  18. package/docs/PERFORMANCE.md +390 -0
  19. package/docs/PI_PROFILING.md +169 -0
  20. package/docs/PLAN-non-agui-agent-support.md +525 -0
  21. package/docs/SDK.md +577 -0
  22. package/docs/SKILLS.md +264 -0
  23. package/docs/blogs/2026-02-28-opensearch-agent-health.md +200 -0
  24. package/docs/blogs/getting-started-blog.md +608 -0
  25. package/docs/diagrams/Agent-health.excalidraw +5656 -0
  26. package/docs/diagrams/architecture.png +0 -0
  27. package/docs/plans/field-redesign.md +468 -0
  28. package/docs/rfcs/001-coding-agent-analytics.md +374 -0
  29. package/docs/rfcs/002-enterprise-leaderboard.md +267 -0
  30. package/docs/rfcs/003-remote-aggregation.md +146 -0
  31. package/docs/rfcs/004-test-sdk-v2.md +599 -0
  32. package/docs/skills/AGENT_HEALTH.md +598 -0
  33. package/docs/skills/AGENT_PROFILE.md +191 -0
  34. package/docs/skills/add-connector/SKILL.md +68 -0
  35. package/docs/skills/agent-health-profile/SKILL.md +40 -0
  36. package/docs/skills/config-auth/SKILL.md +194 -0
  37. package/docs/skills/config-auth/evals/evals.json +35 -0
  38. package/docs/skills/create-pr/SKILL.md +73 -0
  39. package/docs/skills/instrument-otel/SKILL.md +84 -0
  40. package/docs/skills/write-test/SKILL.md +124 -0
  41. package/docs/ui prd.md +376 -0
  42. package/examples/README.md +53 -0
  43. package/examples/config/agent-health.config.example.ts +155 -0
  44. package/examples/connectors/echo-connector.ts +131 -0
  45. package/examples/eval-files/demo.eval.js +128 -0
  46. package/examples/eval-files/sdk-hooks-demo.eval.js +99 -0
  47. package/examples/pi-profiling/README.md +77 -0
  48. package/examples/pi-profiling/agent-health-profile.ts +417 -0
  49. package/lib/dist/lib/agentUtils.d.ts +29 -0
  50. package/lib/dist/lib/agentUtils.d.ts.map +1 -0
  51. package/lib/dist/lib/agentUtils.js +43 -0
  52. package/lib/dist/lib/agentUtils.js.map +1 -0
  53. package/lib/dist/lib/benchmarkExport.d.ts +14 -0
  54. package/lib/dist/lib/benchmarkExport.d.ts.map +1 -0
  55. package/lib/dist/lib/benchmarkExport.js +41 -0
  56. package/lib/dist/lib/benchmarkExport.js.map +1 -0
  57. package/lib/dist/lib/benchmarkVersionUtils.d.ts +37 -0
  58. package/lib/dist/lib/benchmarkVersionUtils.d.ts.map +1 -0
  59. package/lib/dist/lib/benchmarkVersionUtils.js +68 -0
  60. package/lib/dist/lib/benchmarkVersionUtils.js.map +1 -0
  61. package/lib/dist/lib/config/defineConfig.d.ts +27 -0
  62. package/lib/dist/lib/config/defineConfig.d.ts.map +1 -0
  63. package/lib/dist/lib/config/defineConfig.js +28 -0
  64. package/lib/dist/lib/config/defineConfig.js.map +1 -0
  65. package/lib/dist/lib/config/index.d.ts +9 -0
  66. package/lib/dist/lib/config/index.d.ts.map +1 -0
  67. package/lib/dist/lib/config/index.js +8 -0
  68. package/lib/dist/lib/config/index.js.map +1 -0
  69. package/lib/dist/lib/config/loader.d.ts +39 -0
  70. package/lib/dist/lib/config/loader.d.ts.map +1 -0
  71. package/lib/dist/lib/config/loader.js +258 -0
  72. package/lib/dist/lib/config/loader.js.map +1 -0
  73. package/lib/dist/lib/config/statePaths.d.ts +61 -0
  74. package/lib/dist/lib/config/statePaths.d.ts.map +1 -0
  75. package/lib/dist/lib/config/statePaths.js +188 -0
  76. package/lib/dist/lib/config/statePaths.js.map +1 -0
  77. package/lib/dist/lib/config/types.d.ts +231 -0
  78. package/lib/dist/lib/config/types.d.ts.map +1 -0
  79. package/lib/dist/lib/config/types.js +6 -0
  80. package/lib/dist/lib/config/types.js.map +1 -0
  81. package/lib/dist/lib/config.d.ts +39 -0
  82. package/lib/dist/lib/config.d.ts.map +1 -0
  83. package/lib/dist/lib/config.js +118 -0
  84. package/lib/dist/lib/config.js.map +1 -0
  85. package/lib/dist/lib/constants.d.ts +70 -0
  86. package/lib/dist/lib/constants.d.ts.map +1 -0
  87. package/lib/dist/lib/constants.js +365 -0
  88. package/lib/dist/lib/constants.js.map +1 -0
  89. package/lib/dist/lib/contextUtilization.d.ts +23 -0
  90. package/lib/dist/lib/contextUtilization.d.ts.map +1 -0
  91. package/lib/dist/lib/contextUtilization.js +72 -0
  92. package/lib/dist/lib/contextUtilization.js.map +1 -0
  93. package/lib/dist/lib/dashboardMetrics.d.ts +87 -0
  94. package/lib/dist/lib/dashboardMetrics.d.ts.map +1 -0
  95. package/lib/dist/lib/dashboardMetrics.js +242 -0
  96. package/lib/dist/lib/dashboardMetrics.js.map +1 -0
  97. package/lib/dist/lib/dataSourceConfig.d.ts +108 -0
  98. package/lib/dist/lib/dataSourceConfig.d.ts.map +1 -0
  99. package/lib/dist/lib/dataSourceConfig.js +166 -0
  100. package/lib/dist/lib/dataSourceConfig.js.map +1 -0
  101. package/lib/dist/lib/debug.d.ts +26 -0
  102. package/lib/dist/lib/debug.d.ts.map +1 -0
  103. package/lib/dist/lib/debug.js +132 -0
  104. package/lib/dist/lib/debug.js.map +1 -0
  105. package/lib/dist/lib/diagnostics.d.ts +28 -0
  106. package/lib/dist/lib/diagnostics.d.ts.map +1 -0
  107. package/lib/dist/lib/diagnostics.js +65 -0
  108. package/lib/dist/lib/diagnostics.js.map +1 -0
  109. package/lib/dist/lib/envCompat.d.ts +27 -0
  110. package/lib/dist/lib/envCompat.d.ts.map +1 -0
  111. package/lib/dist/lib/envCompat.js +73 -0
  112. package/lib/dist/lib/envCompat.js.map +1 -0
  113. package/lib/dist/lib/findPackageRoot.d.ts +7 -0
  114. package/lib/dist/lib/findPackageRoot.d.ts.map +1 -0
  115. package/lib/dist/lib/findPackageRoot.js +57 -0
  116. package/lib/dist/lib/findPackageRoot.js.map +1 -0
  117. package/lib/dist/lib/hooks.d.ts +36 -0
  118. package/lib/dist/lib/hooks.d.ts.map +1 -0
  119. package/lib/dist/lib/hooks.js +112 -0
  120. package/lib/dist/lib/hooks.js.map +1 -0
  121. package/lib/dist/lib/index.d.ts +47 -0
  122. package/lib/dist/lib/index.d.ts.map +1 -0
  123. package/lib/dist/lib/index.js +62 -0
  124. package/lib/dist/lib/index.js.map +1 -0
  125. package/lib/dist/lib/labels.d.ts +90 -0
  126. package/lib/dist/lib/labels.d.ts.map +1 -0
  127. package/lib/dist/lib/labels.js +158 -0
  128. package/lib/dist/lib/labels.js.map +1 -0
  129. package/lib/dist/lib/markdown.d.ts +16 -0
  130. package/lib/dist/lib/markdown.d.ts.map +1 -0
  131. package/lib/dist/lib/markdown.js +42 -0
  132. package/lib/dist/lib/markdown.js.map +1 -0
  133. package/lib/dist/lib/matchers/expect.d.ts +3 -0
  134. package/lib/dist/lib/matchers/expect.d.ts.map +1 -0
  135. package/lib/dist/lib/matchers/expect.js +225 -0
  136. package/lib/dist/lib/matchers/expect.js.map +1 -0
  137. package/lib/dist/lib/matchers/index.d.ts +8 -0
  138. package/lib/dist/lib/matchers/index.d.ts.map +1 -0
  139. package/lib/dist/lib/matchers/index.js +9 -0
  140. package/lib/dist/lib/matchers/index.js.map +1 -0
  141. package/lib/dist/lib/matchers/judgeAccessor.d.ts +113 -0
  142. package/lib/dist/lib/matchers/judgeAccessor.d.ts.map +1 -0
  143. package/lib/dist/lib/matchers/judgeAccessor.js +183 -0
  144. package/lib/dist/lib/matchers/judgeAccessor.js.map +1 -0
  145. package/lib/dist/lib/matchers/session.d.ts +39 -0
  146. package/lib/dist/lib/matchers/session.d.ts.map +1 -0
  147. package/lib/dist/lib/matchers/session.js +116 -0
  148. package/lib/dist/lib/matchers/session.js.map +1 -0
  149. package/lib/dist/lib/matchers/traces.d.ts +55 -0
  150. package/lib/dist/lib/matchers/traces.d.ts.map +1 -0
  151. package/lib/dist/lib/matchers/traces.js +116 -0
  152. package/lib/dist/lib/matchers/traces.js.map +1 -0
  153. package/lib/dist/lib/matchers/types.d.ts +75 -0
  154. package/lib/dist/lib/matchers/types.d.ts.map +1 -0
  155. package/lib/dist/lib/matchers/types.js +6 -0
  156. package/lib/dist/lib/matchers/types.js.map +1 -0
  157. package/lib/dist/lib/packagePaths.d.ts +29 -0
  158. package/lib/dist/lib/packagePaths.d.ts.map +1 -0
  159. package/lib/dist/lib/packagePaths.js +63 -0
  160. package/lib/dist/lib/packagePaths.js.map +1 -0
  161. package/lib/dist/lib/performance.d.ts +51 -0
  162. package/lib/dist/lib/performance.d.ts.map +1 -0
  163. package/lib/dist/lib/performance.js +159 -0
  164. package/lib/dist/lib/performance.js.map +1 -0
  165. package/lib/dist/lib/portConfig.d.ts +29 -0
  166. package/lib/dist/lib/portConfig.d.ts.map +1 -0
  167. package/lib/dist/lib/portConfig.js +64 -0
  168. package/lib/dist/lib/portConfig.js.map +1 -0
  169. package/lib/dist/lib/preferences.d.ts +63 -0
  170. package/lib/dist/lib/preferences.d.ts.map +1 -0
  171. package/lib/dist/lib/preferences.js +117 -0
  172. package/lib/dist/lib/preferences.js.map +1 -0
  173. package/lib/dist/lib/resolveAgentModel.d.ts +22 -0
  174. package/lib/dist/lib/resolveAgentModel.d.ts.map +1 -0
  175. package/lib/dist/lib/resolveAgentModel.js +37 -0
  176. package/lib/dist/lib/resolveAgentModel.js.map +1 -0
  177. package/lib/dist/lib/runStats.d.ts +92 -0
  178. package/lib/dist/lib/runStats.d.ts.map +1 -0
  179. package/lib/dist/lib/runStats.js +160 -0
  180. package/lib/dist/lib/runStats.js.map +1 -0
  181. package/lib/dist/lib/telemetry/constants.d.ts +60 -0
  182. package/lib/dist/lib/telemetry/constants.d.ts.map +1 -0
  183. package/lib/dist/lib/telemetry/constants.js +87 -0
  184. package/lib/dist/lib/telemetry/constants.js.map +1 -0
  185. package/lib/dist/lib/telemetry/evalSpans.d.ts +61 -0
  186. package/lib/dist/lib/telemetry/evalSpans.d.ts.map +1 -0
  187. package/lib/dist/lib/telemetry/evalSpans.js +254 -0
  188. package/lib/dist/lib/telemetry/evalSpans.js.map +1 -0
  189. package/lib/dist/lib/telemetry/index.d.ts +11 -0
  190. package/lib/dist/lib/telemetry/index.d.ts.map +1 -0
  191. package/lib/dist/lib/telemetry/index.js +15 -0
  192. package/lib/dist/lib/telemetry/index.js.map +1 -0
  193. package/lib/dist/lib/telemetry/opensearchExporter.d.ts +43 -0
  194. package/lib/dist/lib/telemetry/opensearchExporter.d.ts.map +1 -0
  195. package/lib/dist/lib/telemetry/opensearchExporter.js +217 -0
  196. package/lib/dist/lib/telemetry/opensearchExporter.js.map +1 -0
  197. package/lib/dist/lib/telemetry/provider.d.ts +55 -0
  198. package/lib/dist/lib/telemetry/provider.d.ts.map +1 -0
  199. package/lib/dist/lib/telemetry/provider.js +140 -0
  200. package/lib/dist/lib/telemetry/provider.js.map +1 -0
  201. package/lib/dist/lib/testCaseLabels.d.ts +34 -0
  202. package/lib/dist/lib/testCaseLabels.d.ts.map +1 -0
  203. package/lib/dist/lib/testCaseLabels.js +88 -0
  204. package/lib/dist/lib/testCaseLabels.js.map +1 -0
  205. package/lib/dist/lib/testCaseValidation.d.ts +140 -0
  206. package/lib/dist/lib/testCaseValidation.d.ts.map +1 -0
  207. package/lib/dist/lib/testCaseValidation.js +162 -0
  208. package/lib/dist/lib/testCaseValidation.js.map +1 -0
  209. package/lib/dist/lib/testCases/agentFixture.d.ts +80 -0
  210. package/lib/dist/lib/testCases/agentFixture.d.ts.map +1 -0
  211. package/lib/dist/lib/testCases/agentFixture.js +43 -0
  212. package/lib/dist/lib/testCases/agentFixture.js.map +1 -0
  213. package/lib/dist/lib/testCases/authoringSurface.d.ts +10 -0
  214. package/lib/dist/lib/testCases/authoringSurface.d.ts.map +1 -0
  215. package/lib/dist/lib/testCases/authoringSurface.js +54 -0
  216. package/lib/dist/lib/testCases/authoringSurface.js.map +1 -0
  217. package/lib/dist/lib/testCases/codemod.d.ts +13 -0
  218. package/lib/dist/lib/testCases/codemod.d.ts.map +1 -0
  219. package/lib/dist/lib/testCases/codemod.js +169 -0
  220. package/lib/dist/lib/testCases/codemod.js.map +1 -0
  221. package/lib/dist/lib/testCases/define.d.ts +114 -0
  222. package/lib/dist/lib/testCases/define.d.ts.map +1 -0
  223. package/lib/dist/lib/testCases/define.js +253 -0
  224. package/lib/dist/lib/testCases/define.js.map +1 -0
  225. package/lib/dist/lib/testCases/evaluators.d.ts +80 -0
  226. package/lib/dist/lib/testCases/evaluators.d.ts.map +1 -0
  227. package/lib/dist/lib/testCases/evaluators.js +105 -0
  228. package/lib/dist/lib/testCases/evaluators.js.map +1 -0
  229. package/lib/dist/lib/testCases/index.d.ts +14 -0
  230. package/lib/dist/lib/testCases/index.d.ts.map +1 -0
  231. package/lib/dist/lib/testCases/index.js +12 -0
  232. package/lib/dist/lib/testCases/index.js.map +1 -0
  233. package/lib/dist/lib/testCases/judge.d.ts +165 -0
  234. package/lib/dist/lib/testCases/judge.d.ts.map +1 -0
  235. package/lib/dist/lib/testCases/judge.js +359 -0
  236. package/lib/dist/lib/testCases/judge.js.map +1 -0
  237. package/lib/dist/lib/testCases/loader.d.ts +26 -0
  238. package/lib/dist/lib/testCases/loader.d.ts.map +1 -0
  239. package/lib/dist/lib/testCases/loader.js +149 -0
  240. package/lib/dist/lib/testCases/loader.js.map +1 -0
  241. package/lib/dist/lib/testCases/types.d.ts +242 -0
  242. package/lib/dist/lib/testCases/types.d.ts.map +1 -0
  243. package/lib/dist/lib/testCases/types.js +6 -0
  244. package/lib/dist/lib/testCases/types.js.map +1 -0
  245. package/lib/dist/lib/theme.d.ts +6 -0
  246. package/lib/dist/lib/theme.d.ts.map +1 -0
  247. package/lib/dist/lib/theme.js +36 -0
  248. package/lib/dist/lib/theme.js.map +1 -0
  249. package/lib/dist/lib/uiTelemetry.d.ts +7 -0
  250. package/lib/dist/lib/uiTelemetry.d.ts.map +1 -0
  251. package/lib/dist/lib/uiTelemetry.js +25 -0
  252. package/lib/dist/lib/uiTelemetry.js.map +1 -0
  253. package/lib/dist/lib/utils.d.ts +96 -0
  254. package/lib/dist/lib/utils.d.ts.map +1 -0
  255. package/lib/dist/lib/utils.js +232 -0
  256. package/lib/dist/lib/utils.js.map +1 -0
  257. package/lib/dist/lib/workflow/consolidate.d.ts +12 -0
  258. package/lib/dist/lib/workflow/consolidate.d.ts.map +1 -0
  259. package/lib/dist/lib/workflow/consolidate.js +33 -0
  260. package/lib/dist/lib/workflow/consolidate.js.map +1 -0
  261. package/lib/dist/lib/workflow/index.d.ts +13 -0
  262. package/lib/dist/lib/workflow/index.d.ts.map +1 -0
  263. package/lib/dist/lib/workflow/index.js +12 -0
  264. package/lib/dist/lib/workflow/index.js.map +1 -0
  265. package/lib/dist/lib/workflow/ledger.d.ts +30 -0
  266. package/lib/dist/lib/workflow/ledger.d.ts.map +1 -0
  267. package/lib/dist/lib/workflow/ledger.js +41 -0
  268. package/lib/dist/lib/workflow/ledger.js.map +1 -0
  269. package/lib/dist/lib/workflow/pool.d.ts +13 -0
  270. package/lib/dist/lib/workflow/pool.d.ts.map +1 -0
  271. package/lib/dist/lib/workflow/pool.js +44 -0
  272. package/lib/dist/lib/workflow/pool.js.map +1 -0
  273. package/lib/dist/lib/workflow/source.d.ts +22 -0
  274. package/lib/dist/lib/workflow/source.d.ts.map +1 -0
  275. package/lib/dist/lib/workflow/source.js +29 -0
  276. package/lib/dist/lib/workflow/source.js.map +1 -0
  277. package/lib/dist/lib/workflow/stepB.d.ts +71 -0
  278. package/lib/dist/lib/workflow/stepB.d.ts.map +1 -0
  279. package/lib/dist/lib/workflow/stepB.js +99 -0
  280. package/lib/dist/lib/workflow/stepB.js.map +1 -0
  281. package/lib/dist/lib/workflow/types.d.ts +86 -0
  282. package/lib/dist/lib/workflow/types.d.ts.map +1 -0
  283. package/lib/dist/lib/workflow/types.js +6 -0
  284. package/lib/dist/lib/workflow/types.js.map +1 -0
  285. package/lib/dist/lib/workflow/workflow.d.ts +119 -0
  286. package/lib/dist/lib/workflow/workflow.d.ts.map +1 -0
  287. package/lib/dist/lib/workflow/workflow.js +195 -0
  288. package/lib/dist/lib/workflow/workflow.js.map +1 -0
  289. package/lib/dist/services/agent/aguiConverter.d.ts +50 -0
  290. package/lib/dist/services/agent/aguiConverter.d.ts.map +1 -0
  291. package/lib/dist/services/agent/aguiConverter.js +449 -0
  292. package/lib/dist/services/agent/aguiConverter.js.map +1 -0
  293. package/lib/dist/services/agent/index.d.ts +10 -0
  294. package/lib/dist/services/agent/index.d.ts.map +1 -0
  295. package/lib/dist/services/agent/index.js +12 -0
  296. package/lib/dist/services/agent/index.js.map +1 -0
  297. package/lib/dist/services/agent/payloadBuilder.d.ts +33 -0
  298. package/lib/dist/services/agent/payloadBuilder.d.ts.map +1 -0
  299. package/lib/dist/services/agent/payloadBuilder.js +75 -0
  300. package/lib/dist/services/agent/payloadBuilder.js.map +1 -0
  301. package/lib/dist/services/agent/sseStream.d.ts +43 -0
  302. package/lib/dist/services/agent/sseStream.d.ts.map +1 -0
  303. package/lib/dist/services/agent/sseStream.js +223 -0
  304. package/lib/dist/services/agent/sseStream.js.map +1 -0
  305. package/lib/dist/services/connectors/agui/AGUIStreamingConnector.d.ts +44 -0
  306. package/lib/dist/services/connectors/agui/AGUIStreamingConnector.d.ts.map +1 -0
  307. package/lib/dist/services/connectors/agui/AGUIStreamingConnector.js +95 -0
  308. package/lib/dist/services/connectors/agui/AGUIStreamingConnector.js.map +1 -0
  309. package/lib/dist/services/connectors/base/BaseConnector.d.ts +81 -0
  310. package/lib/dist/services/connectors/base/BaseConnector.d.ts.map +1 -0
  311. package/lib/dist/services/connectors/base/BaseConnector.js +170 -0
  312. package/lib/dist/services/connectors/base/BaseConnector.js.map +1 -0
  313. package/lib/dist/services/connectors/claude-code/ClaudeCodeConnector.d.ts +116 -0
  314. package/lib/dist/services/connectors/claude-code/ClaudeCodeConnector.d.ts.map +1 -0
  315. package/lib/dist/services/connectors/claude-code/ClaudeCodeConnector.js +403 -0
  316. package/lib/dist/services/connectors/claude-code/ClaudeCodeConnector.js.map +1 -0
  317. package/lib/dist/services/connectors/index.d.ts +13 -0
  318. package/lib/dist/services/connectors/index.d.ts.map +1 -0
  319. package/lib/dist/services/connectors/index.js +32 -0
  320. package/lib/dist/services/connectors/index.js.map +1 -0
  321. package/lib/dist/services/connectors/kiro/KiroConnector.d.ts +48 -0
  322. package/lib/dist/services/connectors/kiro/KiroConnector.d.ts.map +1 -0
  323. package/lib/dist/services/connectors/kiro/KiroConnector.js +158 -0
  324. package/lib/dist/services/connectors/kiro/KiroConnector.js.map +1 -0
  325. package/lib/dist/services/connectors/langgraph/LangGraphConnector.d.ts +36 -0
  326. package/lib/dist/services/connectors/langgraph/LangGraphConnector.d.ts.map +1 -0
  327. package/lib/dist/services/connectors/langgraph/LangGraphConnector.js +175 -0
  328. package/lib/dist/services/connectors/langgraph/LangGraphConnector.js.map +1 -0
  329. package/lib/dist/services/connectors/mock/MockConnector.d.ts +37 -0
  330. package/lib/dist/services/connectors/mock/MockConnector.d.ts.map +1 -0
  331. package/lib/dist/services/connectors/mock/MockConnector.js +120 -0
  332. package/lib/dist/services/connectors/mock/MockConnector.js.map +1 -0
  333. package/lib/dist/services/connectors/openai-compatible/OpenAICompatibleConnector.d.ts +42 -0
  334. package/lib/dist/services/connectors/openai-compatible/OpenAICompatibleConnector.d.ts.map +1 -0
  335. package/lib/dist/services/connectors/openai-compatible/OpenAICompatibleConnector.js +133 -0
  336. package/lib/dist/services/connectors/openai-compatible/OpenAICompatibleConnector.js.map +1 -0
  337. package/lib/dist/services/connectors/pi/PiConnector.d.ts +87 -0
  338. package/lib/dist/services/connectors/pi/PiConnector.d.ts.map +1 -0
  339. package/lib/dist/services/connectors/pi/PiConnector.js +274 -0
  340. package/lib/dist/services/connectors/pi/PiConnector.js.map +1 -0
  341. package/lib/dist/services/connectors/registry.d.ts +57 -0
  342. package/lib/dist/services/connectors/registry.d.ts.map +1 -0
  343. package/lib/dist/services/connectors/registry.js +106 -0
  344. package/lib/dist/services/connectors/registry.js.map +1 -0
  345. package/lib/dist/services/connectors/rest/RESTConnector.d.ts +38 -0
  346. package/lib/dist/services/connectors/rest/RESTConnector.d.ts.map +1 -0
  347. package/lib/dist/services/connectors/rest/RESTConnector.js +117 -0
  348. package/lib/dist/services/connectors/rest/RESTConnector.js.map +1 -0
  349. package/lib/dist/services/connectors/server.d.ts +13 -0
  350. package/lib/dist/services/connectors/server.d.ts.map +1 -0
  351. package/lib/dist/services/connectors/server.js +34 -0
  352. package/lib/dist/services/connectors/server.js.map +1 -0
  353. package/lib/dist/services/connectors/strands/StrandsConnector.d.ts +48 -0
  354. package/lib/dist/services/connectors/strands/StrandsConnector.d.ts.map +1 -0
  355. package/lib/dist/services/connectors/strands/StrandsConnector.js +221 -0
  356. package/lib/dist/services/connectors/strands/StrandsConnector.js.map +1 -0
  357. package/lib/dist/services/connectors/subprocess/SubprocessConnector.d.ts +88 -0
  358. package/lib/dist/services/connectors/subprocess/SubprocessConnector.d.ts.map +1 -0
  359. package/lib/dist/services/connectors/subprocess/SubprocessConnector.js +418 -0
  360. package/lib/dist/services/connectors/subprocess/SubprocessConnector.js.map +1 -0
  361. package/lib/dist/services/connectors/types.d.ts +213 -0
  362. package/lib/dist/services/connectors/types.d.ts.map +1 -0
  363. package/lib/dist/services/connectors/types.js +6 -0
  364. package/lib/dist/services/connectors/types.js.map +1 -0
  365. package/lib/dist/services/evaluation/bedrockJudge.d.ts +64 -0
  366. package/lib/dist/services/evaluation/bedrockJudge.d.ts.map +1 -0
  367. package/lib/dist/services/evaluation/bedrockJudge.js +167 -0
  368. package/lib/dist/services/evaluation/bedrockJudge.js.map +1 -0
  369. package/lib/dist/services/evaluation/evaluatorError.d.ts +56 -0
  370. package/lib/dist/services/evaluation/evaluatorError.d.ts.map +1 -0
  371. package/lib/dist/services/evaluation/evaluatorError.js +56 -0
  372. package/lib/dist/services/evaluation/evaluatorError.js.map +1 -0
  373. package/lib/dist/services/evaluation/index.d.ts +106 -0
  374. package/lib/dist/services/evaluation/index.d.ts.map +1 -0
  375. package/lib/dist/services/evaluation/index.js +684 -0
  376. package/lib/dist/services/evaluation/index.js.map +1 -0
  377. package/lib/dist/services/evaluation/mockTrajectory.d.ts +3 -0
  378. package/lib/dist/services/evaluation/mockTrajectory.d.ts.map +1 -0
  379. package/lib/dist/services/evaluation/mockTrajectory.js +72 -0
  380. package/lib/dist/services/evaluation/mockTrajectory.js.map +1 -0
  381. package/lib/dist/services/opensearch/client.d.ts +26 -0
  382. package/lib/dist/services/opensearch/client.d.ts.map +1 -0
  383. package/lib/dist/services/opensearch/client.js +131 -0
  384. package/lib/dist/services/opensearch/client.js.map +1 -0
  385. package/lib/dist/services/opensearch/index.d.ts +16 -0
  386. package/lib/dist/services/opensearch/index.d.ts.map +1 -0
  387. package/lib/dist/services/opensearch/index.js +25 -0
  388. package/lib/dist/services/opensearch/index.js.map +1 -0
  389. package/lib/dist/services/storage/asyncBenchmarkStorage.d.ts +123 -0
  390. package/lib/dist/services/storage/asyncBenchmarkStorage.d.ts.map +1 -0
  391. package/lib/dist/services/storage/asyncBenchmarkStorage.js +429 -0
  392. package/lib/dist/services/storage/asyncBenchmarkStorage.js.map +1 -0
  393. package/lib/dist/services/storage/asyncRunStorage.d.ts +127 -0
  394. package/lib/dist/services/storage/asyncRunStorage.d.ts.map +1 -0
  395. package/lib/dist/services/storage/asyncRunStorage.js +448 -0
  396. package/lib/dist/services/storage/asyncRunStorage.js.map +1 -0
  397. package/lib/dist/services/storage/asyncTestCaseStorage.d.ts +156 -0
  398. package/lib/dist/services/storage/asyncTestCaseStorage.d.ts.map +1 -0
  399. package/lib/dist/services/storage/asyncTestCaseStorage.js +285 -0
  400. package/lib/dist/services/storage/asyncTestCaseStorage.js.map +1 -0
  401. package/lib/dist/services/storage/index.d.ts +17 -0
  402. package/lib/dist/services/storage/index.d.ts.map +1 -0
  403. package/lib/dist/services/storage/index.js +20 -0
  404. package/lib/dist/services/storage/index.js.map +1 -0
  405. package/lib/dist/services/storage/migration.d.ts +54 -0
  406. package/lib/dist/services/storage/migration.d.ts.map +1 -0
  407. package/lib/dist/services/storage/migration.js +296 -0
  408. package/lib/dist/services/storage/migration.js.map +1 -0
  409. package/lib/dist/services/storage/opensearchClient.d.ts +924 -0
  410. package/lib/dist/services/storage/opensearchClient.d.ts.map +1 -0
  411. package/lib/dist/services/storage/opensearchClient.js +435 -0
  412. package/lib/dist/services/storage/opensearchClient.js.map +1 -0
  413. package/lib/dist/services/traces/browserRecovery.d.ts +26 -0
  414. package/lib/dist/services/traces/browserRecovery.d.ts.map +1 -0
  415. package/lib/dist/services/traces/browserRecovery.js +81 -0
  416. package/lib/dist/services/traces/browserRecovery.js.map +1 -0
  417. package/lib/dist/services/traces/categoryStyles.d.ts +21 -0
  418. package/lib/dist/services/traces/categoryStyles.d.ts.map +1 -0
  419. package/lib/dist/services/traces/categoryStyles.js +56 -0
  420. package/lib/dist/services/traces/categoryStyles.js.map +1 -0
  421. package/lib/dist/services/traces/executionOrderTransform.d.ts +35 -0
  422. package/lib/dist/services/traces/executionOrderTransform.d.ts.map +1 -0
  423. package/lib/dist/services/traces/executionOrderTransform.js +313 -0
  424. package/lib/dist/services/traces/executionOrderTransform.js.map +1 -0
  425. package/lib/dist/services/traces/fetchSpansForRun.d.ts +86 -0
  426. package/lib/dist/services/traces/fetchSpansForRun.d.ts.map +1 -0
  427. package/lib/dist/services/traces/fetchSpansForRun.js +69 -0
  428. package/lib/dist/services/traces/fetchSpansForRun.js.map +1 -0
  429. package/lib/dist/services/traces/flowTransform.d.ts +24 -0
  430. package/lib/dist/services/traces/flowTransform.d.ts.map +1 -0
  431. package/lib/dist/services/traces/flowTransform.js +228 -0
  432. package/lib/dist/services/traces/flowTransform.js.map +1 -0
  433. package/lib/dist/services/traces/index.d.ts +121 -0
  434. package/lib/dist/services/traces/index.d.ts.map +1 -0
  435. package/lib/dist/services/traces/index.js +255 -0
  436. package/lib/dist/services/traces/index.js.map +1 -0
  437. package/lib/dist/services/traces/intentTransform.d.ts +20 -0
  438. package/lib/dist/services/traces/intentTransform.d.ts.map +1 -0
  439. package/lib/dist/services/traces/intentTransform.js +131 -0
  440. package/lib/dist/services/traces/intentTransform.js.map +1 -0
  441. package/lib/dist/services/traces/judgeAgentsHints.d.ts +63 -0
  442. package/lib/dist/services/traces/judgeAgentsHints.d.ts.map +1 -0
  443. package/lib/dist/services/traces/judgeAgentsHints.js +89 -0
  444. package/lib/dist/services/traces/judgeAgentsHints.js.map +1 -0
  445. package/lib/dist/services/traces/messageExtraction.d.ts +15 -0
  446. package/lib/dist/services/traces/messageExtraction.d.ts.map +1 -0
  447. package/lib/dist/services/traces/messageExtraction.js +251 -0
  448. package/lib/dist/services/traces/messageExtraction.js.map +1 -0
  449. package/lib/dist/services/traces/spanCategorization.d.ts +63 -0
  450. package/lib/dist/services/traces/spanCategorization.d.ts.map +1 -0
  451. package/lib/dist/services/traces/spanCategorization.js +276 -0
  452. package/lib/dist/services/traces/spanCategorization.js.map +1 -0
  453. package/lib/dist/services/traces/spanPreprocessing.d.ts +37 -0
  454. package/lib/dist/services/traces/spanPreprocessing.d.ts.map +1 -0
  455. package/lib/dist/services/traces/spanPreprocessing.js +102 -0
  456. package/lib/dist/services/traces/spanPreprocessing.js.map +1 -0
  457. package/lib/dist/services/traces/spansToTrajectory.d.ts +36 -0
  458. package/lib/dist/services/traces/spansToTrajectory.d.ts.map +1 -0
  459. package/lib/dist/services/traces/spansToTrajectory.js +387 -0
  460. package/lib/dist/services/traces/spansToTrajectory.js.map +1 -0
  461. package/lib/dist/services/traces/toolSimilarity.d.ts +35 -0
  462. package/lib/dist/services/traces/toolSimilarity.d.ts.map +1 -0
  463. package/lib/dist/services/traces/toolSimilarity.js +203 -0
  464. package/lib/dist/services/traces/toolSimilarity.js.map +1 -0
  465. package/lib/dist/services/traces/traceComparison.d.ts +31 -0
  466. package/lib/dist/services/traces/traceComparison.d.ts.map +1 -0
  467. package/lib/dist/services/traces/traceComparison.js +318 -0
  468. package/lib/dist/services/traces/traceComparison.js.map +1 -0
  469. package/lib/dist/services/traces/traceGrouping.d.ts +19 -0
  470. package/lib/dist/services/traces/traceGrouping.d.ts.map +1 -0
  471. package/lib/dist/services/traces/traceGrouping.js +107 -0
  472. package/lib/dist/services/traces/traceGrouping.js.map +1 -0
  473. package/lib/dist/services/traces/tracePoller.d.ts +84 -0
  474. package/lib/dist/services/traces/tracePoller.d.ts.map +1 -0
  475. package/lib/dist/services/traces/tracePoller.js +309 -0
  476. package/lib/dist/services/traces/tracePoller.js.map +1 -0
  477. package/lib/dist/services/traces/traceStats.d.ts +45 -0
  478. package/lib/dist/services/traces/traceStats.d.ts.map +1 -0
  479. package/lib/dist/services/traces/traceStats.js +114 -0
  480. package/lib/dist/services/traces/traceStats.js.map +1 -0
  481. package/lib/dist/services/traces/traceSummary.d.ts +47 -0
  482. package/lib/dist/services/traces/traceSummary.d.ts.map +1 -0
  483. package/lib/dist/services/traces/traceSummary.js +68 -0
  484. package/lib/dist/services/traces/traceSummary.js.map +1 -0
  485. package/lib/dist/services/traces/utils.d.ts +33 -0
  486. package/lib/dist/services/traces/utils.d.ts.map +1 -0
  487. package/lib/dist/services/traces/utils.js +114 -0
  488. package/lib/dist/services/traces/utils.js.map +1 -0
  489. package/lib/dist/types/agui.d.ts +13 -0
  490. package/lib/dist/types/agui.d.ts.map +1 -0
  491. package/lib/dist/types/agui.js +16 -0
  492. package/lib/dist/types/agui.js.map +1 -0
  493. package/lib/dist/types/index.d.ts +1175 -0
  494. package/lib/dist/types/index.d.ts.map +1 -0
  495. package/lib/dist/types/index.js +12 -0
  496. package/lib/dist/types/index.js.map +1 -0
  497. package/lib/dist/types/skills.d.ts +146 -0
  498. package/lib/dist/types/skills.d.ts.map +1 -0
  499. package/lib/dist/types/skills.js +6 -0
  500. package/lib/dist/types/skills.js.map +1 -0
  501. package/observio-sample-agent/pi-package/README.md +112 -0
  502. package/observio-sample-agent/pi-package/extensions/agent-health.ts +373 -0
  503. package/observio-sample-agent/pi-package/package.json +17 -0
  504. package/observio-sample-agent/pi-package/prompts/agent-health.md +37 -0
  505. package/observio-sample-agent/pi-package/skills/create-pr/SKILL.md +88 -0
  506. package/observio-sample-agent/pi-package/skills/fix-bug/SKILL.md +71 -0
  507. package/observio-sample-agent/pi-package/skills/implement-feature/SKILL.md +156 -0
  508. package/observio-sample-agent/pi-package/skills/instrument-otel/SKILL.md +208 -0
  509. package/observio-sample-agent/pi-package/skills/setup-collector/SKILL.md +146 -0
  510. package/observio-sample-agent/pi-package/skills/write-test/SKILL.md +115 -0
  511. package/package.json +64 -13
  512. package/server/dist/app.js +32651 -17637
  513. package/server/dist/index.js +29875 -14638
  514. package/tsconfig.lib.json +71 -0
  515. package/dist/assets/index-EvPLSTAS.js +0 -267
  516. package/dist/assets/index-RXasQKUs.css +0 -1
  517. package/lib/dist/config/index.js +0 -404
  518. package/lib/dist/index.js +0 -1665
@@ -0,0 +1,608 @@
1
+ ## Getting Started with Agent Health: A Complete Walkthrough
2
+
3
+ In our [introductory blog post](https://opensearch.org/blog/opensearch-agent-health/), we showed you what Agent Health is and why it matters. Now let's roll up our sleeves and go through a hands-on walkthrough.
4
+
5
+ This guide is progressive — you can stop at any point and come back later:
6
+
7
+ 1. **Try sample data** — explore the UI with zero setup
8
+ 2. **Try it yourself** — connect your own agent endpoint, configure a judge, and run your first evaluation (no OpenSearch required)
9
+ 3. **Add tracing** — connect or instrument OpenTelemetry traces for deep observability
10
+
11
+ **What you'll need to start:** You have an agent (any protocol), an LLM provider, and optionally tracing via OpenTelemetry to some backend. Agent Health plugs into all of these.
12
+
13
+ <!-- TODO: Diagram — high-level: Agent + LLM + (optional) OTel tracing → Agent Health -->
14
+
15
+ ### Prerequisites
16
+ * [Node.js](https://nodejs.org/) 18 or later
17
+ * [Docker](https://www.docker.com/) (optional — for the local OpenSearch stack or the Docker Compose setup)
18
+ * AWS credentials (for the Bedrock LLM judge) or an OpenAI-compatible endpoint (for LiteLLM, Ollama, etc.)
19
+
20
+ ### Launch Agent Health
21
+
22
+ ```
23
+ npx @opensearch-project/agent-health
24
+ ```
25
+
26
+ Open your browser to `http://localhost:4001`. You'll land on the Agent Health home screen with three main sections: **Traces**, **Benchmarks**, and **Compare**.
27
+
28
+ <!-- TODO: Screenshot — Agent Health home screen -->
29
+
30
+ ---
31
+
32
+ ## Part I: Try Sample Data
33
+
34
+ Before connecting your own agent, explore with the pre-loaded demo data. This gives you a feel for the interface with zero configuration.
35
+
36
+ ### Traces view
37
+
38
+ Navigate to the **Traces** tab. You'll see pre-loaded agent execution traces. Click any trace to open the detail view:
39
+ * **Timeline view** — a chronological breakdown of every span (LLM call, tool invocation, retrieval step) with durations
40
+ * **Flow view** — a visual graph of how data flows between agent components
41
+ * **Span details** — click any span to see its attributes, inputs, outputs, and metadata
42
+
43
+ <!-- TODO: Screenshot — Traces list + detail view -->
44
+
45
+ ### Benchmarks view
46
+
47
+ Navigate to **Benchmarks**. You'll find a demo benchmark called "Travel Planning Accuracy - Demo". Click into it to see the test cases, then run it to watch the LLM judge evaluate each case in real time.
48
+
49
+ <!-- TODO: Screenshot — Benchmark detail with test cases -->
50
+
51
+ ### Compare view
52
+
53
+ After running a benchmark, go to **Compare**. Select two runs to see side-by-side metrics: pass rate, latency, cost, and per-test-case diffs.
54
+
55
+ <!-- TODO: Screenshot — Compare view side-by-side -->
56
+
57
+ ---
58
+
59
+ ## Part II: Try It Yourself
60
+
61
+ Now let's connect your own agent and run a real evaluation. No OpenSearch storage required — Agent Health stores data locally on disk by default.
62
+
63
+ <!-- TODO: Diagram — build-up: Agent Health box, now adding "Your Agent" arrow -->
64
+
65
+ ### Step 1: Configure your agent endpoint
66
+
67
+ Create an `agent-health.config.ts` file in your working directory (or run `npx @opensearch-project/agent-health init` to generate one):
68
+
69
+ ```typescript
70
+ export default {
71
+ agents: [
72
+ {
73
+ key: "my-agent",
74
+ name: "My Agent",
75
+ endpoint: "http://localhost:3000/agent",
76
+ connectorType: "agui-streaming",
77
+ models: ["claude-sonnet-4"],
78
+ useTraces: false, // No OpenSearch tracing required for basic evaluation
79
+ }
80
+ ],
81
+ };
82
+ ```
83
+
84
+ **Connector types** determine how Agent Health communicates with your agent. They fall into three categories:
85
+
86
+ **Agent Protocol** — connectors that know the full request/response contract of a specific agent framework:
87
+
88
+ | Connector | Protocol |
89
+ |-----------|----------|
90
+ | `agui-streaming` | [AG-UI](https://docs.ag-ui.com) SSE streaming protocol (default) |
91
+ | `claude-code` | Claude Code CLI — NDJSON streaming with MCP tool support |
92
+
93
+ **Transport** — generic communication channels where you control the payload format via hooks or a custom connector:
94
+
95
+ | Connector | Protocol |
96
+ |-----------|----------|
97
+ | `rest` | Standard HTTP POST — Agent Health sends JSON and parses common response shapes |
98
+ | `cli` | Spawns a CLI command as a child process, captures stdout |
99
+
100
+ **LLM Protocol** — talks directly to an LLM endpoint (useful for evaluating raw model responses, not a full agent loop):
101
+
102
+ | Connector | Protocol |
103
+ |-----------|----------|
104
+ | `openai-compatible` | OpenAI Chat Completions standard: `POST /v1/chat/completions` with `messages` array. Works with LiteLLM, Ollama, vLLM, Azure OpenAI, OpenAI, etc. |
105
+
106
+ **Testing:**
107
+
108
+ | Connector | Protocol |
109
+ |-----------|----------|
110
+ | `mock` | In-memory demo trajectory, no real agent needed |
111
+
112
+ #### Writing your own connector
113
+
114
+ If none of the built-in connectors fit your agent's protocol, you can write a custom one. Create a class that extends `BaseConnector` and implement three methods: `buildPayload`, `execute`, and `parseResponse`.
115
+
116
+ ```typescript
117
+ import { BaseConnector } from '@opensearch-project/agent-health/connectors';
118
+ import type {
119
+ ConnectorAuth,
120
+ ConnectorRequest,
121
+ ConnectorResponse,
122
+ ConnectorProgressCallback,
123
+ ConnectorRawEventCallback,
124
+ } from '@opensearch-project/agent-health/connectors';
125
+ import type { TrajectoryStep } from '@opensearch-project/agent-health/types';
126
+
127
+ class MyCustomConnector extends BaseConnector {
128
+ readonly type = 'my-protocol' as const;
129
+ readonly name = 'My Custom Protocol';
130
+ readonly supportsStreaming = false;
131
+
132
+ buildPayload(request: ConnectorRequest): any {
133
+ // Transform the standard request into your agent's expected format
134
+ return {
135
+ query: request.testCase.initialPrompt,
136
+ model: request.modelId,
137
+ };
138
+ }
139
+
140
+ async execute(
141
+ endpoint: string,
142
+ request: ConnectorRequest,
143
+ auth: ConnectorAuth,
144
+ onProgress?: ConnectorProgressCallback,
145
+ onRawEvent?: ConnectorRawEventCallback
146
+ ): Promise<ConnectorResponse> {
147
+ const payload = request.payload || this.buildPayload(request);
148
+ const headers = this.buildAuthHeaders(auth);
149
+
150
+ const response = await fetch(endpoint, {
151
+ method: 'POST',
152
+ headers: { 'Content-Type': 'application/json', ...headers },
153
+ body: JSON.stringify(payload),
154
+ });
155
+
156
+ const data = await response.json();
157
+ onRawEvent?.(data);
158
+
159
+ const trajectory = this.parseResponse(data);
160
+ trajectory.forEach(step => onProgress?.(step));
161
+
162
+ return { trajectory, runId: data.id || null };
163
+ }
164
+
165
+ parseResponse(data: any): TrajectoryStep[] {
166
+ // Convert your agent's response format into TrajectoryStep array
167
+ return [
168
+ this.createStep('thinking', data.reasoning || ''),
169
+ this.createStep('response', data.answer || JSON.stringify(data)),
170
+ ];
171
+ }
172
+ }
173
+ ```
174
+
175
+ Register your connector in `agent-health.config.ts`:
176
+
177
+ ```typescript
178
+ export default {
179
+ connectors: [new MyCustomConnector()],
180
+ agents: [
181
+ {
182
+ key: "my-agent",
183
+ name: "My Agent",
184
+ endpoint: "http://localhost:3000/agent",
185
+ connectorType: "my-protocol",
186
+ models: ["claude-sonnet-4"],
187
+ useTraces: false,
188
+ }
189
+ ],
190
+ };
191
+ ```
192
+
193
+ #### Lifecycle hooks
194
+
195
+ For simpler customizations — adding auth tokens, modifying payloads, pre-creating threads — you can use lifecycle hooks without writing a full connector:
196
+
197
+ ```typescript
198
+ {
199
+ key: "my-agent",
200
+ // ...
201
+ hooks: {
202
+ beforeRequest: async ({ endpoint, payload, headers }) => {
203
+ // e.g., add auth tokens, modify payload, pre-create threads
204
+ return { endpoint, payload, headers };
205
+ },
206
+ },
207
+ }
208
+ ```
209
+
210
+ <!-- TODO: Diagram — build-up: Agent Health ↔ Your Agent (with connector arrow) -->
211
+
212
+ ### Step 2: Configure the judge
213
+
214
+ The LLM judge evaluates your agent's responses against expected outcomes. Agent Health supports two judge providers:
215
+
216
+ **Option A: AWS Bedrock (default)**
217
+
218
+ ```bash
219
+ # AWS profile (recommended)
220
+ export AWS_PROFILE=your-profile
221
+ export AWS_REGION=us-east-1
222
+
223
+ # Or explicit credentials
224
+ export AWS_ACCESS_KEY_ID=...
225
+ export AWS_SECRET_ACCESS_KEY=...
226
+ export AWS_SESSION_TOKEN=...
227
+ ```
228
+
229
+ **Option B: OpenAI-compatible endpoint** — any provider that implements the OpenAI Chat Completions standard (`POST /v1/chat/completions` with `messages` array). This includes LiteLLM, Ollama, vLLM, Azure OpenAI, or OpenAI directly.
230
+
231
+ ```bash
232
+ export OPENAI_COMPATIBLE_ENDPOINT=http://localhost:4000/v1/chat/completions
233
+ export OPENAI_COMPATIBLE_API_KEY=your-api-key # optional, depends on provider
234
+ ```
235
+
236
+ You can also configure the judge in `agent-health.config.ts`:
237
+
238
+ ```typescript
239
+ export default {
240
+ judge: {
241
+ provider: "openai-compatible", // or "bedrock" (default)
242
+ model: "gpt-4o", // model name forwarded to the endpoint
243
+ },
244
+ // ...agents, etc.
245
+ };
246
+ ```
247
+
248
+ <!-- TODO: This section will expand once judge configuration UI is implemented -->
249
+
250
+ <!-- TODO: Diagram — build-up: Agent Health ↔ Your Agent, Agent Health ↔ Judge -->
251
+
252
+ ### Step 3: Run your first evaluation
253
+
254
+ With your agent endpoint and judge configured, you can run an evaluation — no OpenSearch storage or tracing required.
255
+
256
+ #### Create test cases
257
+
258
+ <!-- TODO: Replace these sample test cases with ones based on our own shipped evaluation engine / sample agent -->
259
+
260
+ Create a file called `travel-benchmark.json` with test cases for your agent:
261
+
262
+ ```json
263
+ [
264
+ {
265
+ "name": "Basic flight search",
266
+ "description": "User asks for a simple flight search",
267
+ "labels": ["category:Travel", "difficulty:Easy"],
268
+ "initialPrompt": "Find me flights from Seattle to New York next Friday",
269
+ "expectedOutcomes": [
270
+ "Agent should call search_flights tool with origin=Seattle and destination=New York",
271
+ "Agent should present flight options with times and prices"
272
+ ]
273
+ },
274
+ {
275
+ "name": "Hotel search with dates",
276
+ "description": "User asks for hotel availability",
277
+ "labels": ["category:Travel", "difficulty:Easy"],
278
+ "initialPrompt": "Are there any hotels available in Manhattan for March 20-22?",
279
+ "expectedOutcomes": [
280
+ "Agent should call search_hotels tool with location=Manhattan",
281
+ "Agent should present hotel options with prices and ratings"
282
+ ]
283
+ },
284
+ {
285
+ "name": "Multi-step trip planning",
286
+ "description": "User asks for both flights and hotels in one query",
287
+ "labels": ["category:Travel", "difficulty:Medium"],
288
+ "initialPrompt": "Plan a trip from Seattle to New York next weekend. I need both flights and a hotel.",
289
+ "expectedOutcomes": [
290
+ "Agent should call search_flights tool",
291
+ "Agent should call search_hotels tool",
292
+ "Agent should present a combined itinerary with flight and hotel options"
293
+ ]
294
+ }
295
+ ]
296
+ ```
297
+
298
+ You can also create test cases directly in the UI via **Settings > Use Cases**.
299
+
300
+ #### Run via CLI
301
+
302
+ ```
303
+ npx @opensearch-project/agent-health benchmark \
304
+ -f travel-benchmark.json \
305
+ -a my-agent \
306
+ -v
307
+ ```
308
+
309
+ This will:
310
+ * Import the test cases
311
+ * Send each `initialPrompt` to your agent
312
+ * Have the LLM judge score agent responses against `expectedOutcomes`
313
+
314
+ #### Run via UI
315
+
316
+ Alternatively, in the Agent Health UI:
317
+ * Go to **Benchmarks**
318
+ * Click **Import JSON** and select `travel-benchmark.json`
319
+ * Click **Run** on your benchmark
320
+ * Select your agent endpoint and judge model
321
+ * Watch as each test case runs and gets evaluated in real time
322
+
323
+ <!-- TODO: Screenshot — Benchmark run in progress -->
324
+
325
+ #### Analyze results
326
+
327
+ Open the benchmark run results to see:
328
+ * **Overall pass rate** — what percentage of test cases passed
329
+ * **Per-test-case results** — each test case with its pass/fail status, the LLM judge's reasoning, and a score
330
+ * **Improvement strategies** — prioritized recommendations for what to fix first
331
+
332
+ For failed test cases, the LLM judge explains exactly _why_ it failed — for example, "the agent did not call search_hotels as expected" or "the response was missing price information."
333
+
334
+ #### Iterate and compare
335
+
336
+ Make improvements to your agent (update prompts, add tools, change models), then re-run the benchmark. Use the **Compare** view to see side-by-side how your changes affected:
337
+ * Pass rates across all test cases
338
+ * Individual test case outcomes that flipped from fail to pass (or vice versa)
339
+ * Latency and cost changes
340
+
341
+ <!-- TODO: Diagram — build-up: full eval loop: Agent Health ↔ Agent ↔ Judge, with results -->
342
+
343
+ ---
344
+
345
+ ## Part III: Add Tracing
346
+
347
+ Tracing gives you deep visibility into what your agent does internally — every LLM call, tool invocation, and reasoning step. Pick the scenario that matches your situation:
348
+
349
+ ### Scenario A: You already have traces in OpenSearch
350
+
351
+ If your agent already sends OTel traces to an OpenSearch cluster, just point Agent Health at it.
352
+
353
+ **Configure the Observability Storage connection** (pick one):
354
+
355
+ * **Via Settings UI** — Go to **Settings** and fill in the **Observability Storage** section with your OpenSearch cluster URL and credentials. Choose "Basic Auth" for username/password or "AWS SigV4" for AWS-managed clusters.
356
+ * **Via environment variables** — Add the connection details to your `.env` file:
357
+
358
+ ```bash
359
+ # Option A: Basic Auth (username/password)
360
+ OPENSEARCH_LOGS_ENDPOINT=https://your-opensearch-cluster:9200
361
+ OPENSEARCH_LOGS_USERNAME=admin
362
+ OPENSEARCH_LOGS_PASSWORD=admin
363
+
364
+ # Option B: AWS SigV4 (for AWS-managed OpenSearch or Serverless)
365
+ OPENSEARCH_LOGS_ENDPOINT=https://your-cluster.us-east-1.es.amazonaws.com
366
+ OPENSEARCH_LOGS_AUTH_TYPE=sigv4
367
+ OPENSEARCH_LOGS_AWS_REGION=us-east-1
368
+ OPENSEARCH_LOGS_AWS_SERVICE=es # 'es' for managed, 'aoss' for Serverless
369
+ # OPENSEARCH_LOGS_AWS_PROFILE=MyProfile # optional, uses default credential chain
370
+ ```
371
+
372
+ Then enable traces for your agent in `agent-health.config.ts`:
373
+
374
+ ```typescript
375
+ {
376
+ key: "my-agent",
377
+ name: "My Agent",
378
+ endpoint: "http://localhost:3000/agent",
379
+ connectorType: "rest",
380
+ useTraces: true, // Enable trace collection for this agent
381
+ models: ["claude-sonnet-4"],
382
+ }
383
+ ```
384
+
385
+ Navigate to the **Traces** tab — you should see your agent's traces immediately.
386
+
387
+ Now re-run your benchmark from Step 3. Each evaluation run will also pull in the associated traces — giving you full visibility into what the agent did internally, alongside the judge's assessment.
388
+
389
+ <!-- TODO: Screenshot — Traces tab showing real agent traces -->
390
+
391
+ ### Scenario B: Your traces go to another backend (Jaeger, Grafana, etc.)
392
+
393
+ If your agent already has OTel instrumentation but traces go to a non-OpenSearch backend, you need a local OpenSearch + OTel Collector stack to receive them. Two options:
394
+
395
+ **Option 1: Observability Stack shell script**
396
+
397
+ ```
398
+ curl -fsSL https://raw.githubusercontent.com/opensearch-project/observability-stack/main/install.sh | bash
399
+ ```
400
+
401
+ This launches:
402
+ * **OpenSearch** on port `9200` — stores traces
403
+ * **OTEL Collector** on port `4317` — receives OpenTelemetry traces
404
+
405
+ **Option 2: Docker Compose**
406
+
407
+ <!-- TODO: Add Docker Compose instructions — this may be consolidated into a single Docker Compose that ships with Agent Health itself, which would simplify both options into one -->
408
+
409
+ Both options give you the same result. Then:
410
+ 1. Swap your agent's OTEL exporter endpoint to `http://localhost:4317`
411
+ 2. Configure the Observability Storage connection in Agent Health (via Settings UI or `.env` file, as shown in Scenario A)
412
+ 3. Enable `useTraces: true` for your agent in `agent-health.config.ts`
413
+
414
+ <!-- TODO: Test shell script with Docker and add detailed instructions -->
415
+
416
+ ### Scenario C: Your agent has no OTel instrumentation yet
417
+
418
+ If your agent doesn't have OpenTelemetry instrumentation, you'll need to add it. Here's a walkthrough using a Python agent as an example.
419
+
420
+ <!-- TODO: Ship a sample agent with Agent Health that users can run to test the tracing flow end-to-end without needing their own instrumented agent -->
421
+
422
+ #### A sample agent (before instrumentation)
423
+
424
+ ```python
425
+ # travel_agent.py
426
+ import openai
427
+
428
+ client = openai.OpenAI()
429
+
430
+ tools = [
431
+ {
432
+ "type": "function",
433
+ "function": {
434
+ "name": "search_flights",
435
+ "description": "Search for available flights",
436
+ "parameters": {
437
+ "type": "object",
438
+ "properties": {
439
+ "origin": {"type": "string"},
440
+ "destination": {"type": "string"},
441
+ "date": {"type": "string"}
442
+ },
443
+ "required": ["origin", "destination", "date"]
444
+ }
445
+ }
446
+ },
447
+ {
448
+ "type": "function",
449
+ "function": {
450
+ "name": "search_hotels",
451
+ "description": "Search for available hotels",
452
+ "parameters": {
453
+ "type": "object",
454
+ "properties": {
455
+ "location": {"type": "string"},
456
+ "check_in": {"type": "string"},
457
+ "check_out": {"type": "string"}
458
+ },
459
+ "required": ["location", "check_in", "check_out"]
460
+ }
461
+ }
462
+ }
463
+ ]
464
+
465
+ def run_agent(user_query: str) -> str:
466
+ messages = [
467
+ {"role": "system", "content": "You are a helpful travel planning assistant."},
468
+ {"role": "user", "content": user_query}
469
+ ]
470
+ response = client.chat.completions.create(
471
+ model="gpt-4o-mini",
472
+ messages=messages,
473
+ tools=tools,
474
+ )
475
+ return response.choices[0].message.content
476
+
477
+ if __name__ == "__main__":
478
+ result = run_agent("Find me flights from Seattle to New York next Friday")
479
+ print(result)
480
+ ```
481
+
482
+ #### Add OpenTelemetry instrumentation
483
+
484
+ Install the required packages:
485
+
486
+ ```
487
+ pip install opentelemetry-api opentelemetry-sdk opentelemetry-exporter-otlp \
488
+ opentelemetry-instrumentation-openai
489
+ ```
490
+
491
+ Now add tracing to the agent:
492
+
493
+ ```python
494
+ # travel_agent_traced.py
495
+ from opentelemetry import trace
496
+ from opentelemetry.sdk.trace import TracerProvider
497
+ from opentelemetry.sdk.trace.export import BatchSpanProcessor
498
+ from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
499
+ from opentelemetry.sdk.resources import Resource
500
+
501
+ # Configure the tracer to send spans to the OTEL Collector
502
+ resource = Resource.create({"service.name": "travel-agent"})
503
+ provider = TracerProvider(resource=resource)
504
+ exporter = OTLPSpanExporter(endpoint="http://localhost:4317", insecure=True)
505
+ provider.add_span_processor(BatchSpanProcessor(exporter))
506
+ trace.set_tracer_provider(provider)
507
+
508
+ tracer = trace.get_tracer("travel-agent")
509
+
510
+ # Auto-instrument OpenAI calls (captures LLM spans automatically)
511
+ from opentelemetry.instrumentation.openai import OpenAIInstrumentor
512
+ OpenAIInstrumentor().instrument()
513
+
514
+ import openai
515
+
516
+ client = openai.OpenAI()
517
+
518
+ tools = [
519
+ # ... same tool definitions as above ...
520
+ ]
521
+
522
+ def search_flights(origin, destination, date):
523
+ """Simulated flight search."""
524
+ with tracer.start_as_current_span("tool.search_flights") as span:
525
+ span.set_attribute("tool.name", "search_flights")
526
+ span.set_attribute("tool.parameters.origin", origin)
527
+ span.set_attribute("tool.parameters.destination", destination)
528
+ results = [
529
+ {"flight": "AA123", "time": "8:00 AM", "price": "$350"},
530
+ {"flight": "UA456", "time": "2:00 PM", "price": "$280"},
531
+ ]
532
+ span.set_attribute("tool.result_count", len(results))
533
+ return results
534
+
535
+ def search_hotels(location, check_in, check_out):
536
+ """Simulated hotel search."""
537
+ with tracer.start_as_current_span("tool.search_hotels") as span:
538
+ span.set_attribute("tool.name", "search_hotels")
539
+ span.set_attribute("tool.parameters.location", location)
540
+ results = [
541
+ {"hotel": "Hilton Manhattan", "price": "$200/night", "rating": 4.5},
542
+ ]
543
+ span.set_attribute("tool.result_count", len(results))
544
+ return results
545
+
546
+ def run_agent(user_query: str) -> str:
547
+ with tracer.start_as_current_span("agent.run") as span:
548
+ span.set_attribute("agent.name", "travel-agent")
549
+ span.set_attribute("agent.input", user_query)
550
+
551
+ messages = [
552
+ {"role": "system", "content": "You are a helpful travel planning assistant."},
553
+ {"role": "user", "content": user_query}
554
+ ]
555
+
556
+ response = client.chat.completions.create(
557
+ model="gpt-4o-mini",
558
+ messages=messages,
559
+ tools=tools,
560
+ )
561
+
562
+ result = response.choices[0].message.content
563
+ span.set_attribute("agent.output", result or "")
564
+ return result
565
+
566
+ if __name__ == "__main__":
567
+ result = run_agent("Find me flights from Seattle to New York next Friday")
568
+ print(result)
569
+ # Flush spans before exit
570
+ provider.force_flush()
571
+ ```
572
+
573
+ #### Run the agent and view traces
574
+
575
+ ```
576
+ python travel_agent_traced.py
577
+ ```
578
+
579
+ Now go back to **Traces** in Agent Health (`http://localhost:4001`). You should see a new trace appear showing:
580
+ * The top-level `agent.run` span with your query
581
+ * Nested LLM call spans (auto-instrumented by the OpenAI instrumentor)
582
+ * Tool call spans like `tool.search_flights`
583
+
584
+ Click into the trace to explore the timeline and flow views. You can now see exactly what your agent did, how long each step took, and what data flowed between components.
585
+
586
+ <!-- TODO: Screenshot — Trace detail with timeline and flow views -->
587
+
588
+ Once traces are flowing, re-run your benchmark from Step 3 to get evaluations with full trace data.
589
+
590
+ <!-- TODO: Diagram — complete picture: Agent Health ↔ Agent ↔ Judge + OTel traces flowing in -->
591
+
592
+ ---
593
+
594
+ ## Next steps
595
+
596
+ You now have a complete Agent Health workflow — from exploring sample data, to evaluating your own agent, to adding deep observability with OTel traces. From here, you can:
597
+ * Add more test cases to cover edge cases and failure modes
598
+ * Integrate benchmark runs into your CI/CD pipeline for automated regression testing
599
+ * Explore the trace views to debug specific agent failures in detail
600
+ * Use the **Compare** view to A/B test agent configurations over time
601
+
602
+ We'll cover each of these topics in upcoming posts. Stay tuned!
603
+
604
+ ## Resources
605
+ * **GitHub Repository**: [opensearch-project/agent-health](https://github.com/opensearch-project/agent-health)
606
+ * **First blog post**: [OpenSearch Agent Health: Open-Source Observability and Evaluation for AI Agents](https://opensearch.org/blog/opensearch-agent-health/)
607
+ * **OpenTelemetry instrumentation guides**: [opentelemetry.io/docs/instrumentation](https://opentelemetry.io/docs/instrumentation/)
608
+ * **OpenSearch Observability Stack**: [opensearch-project/observability-stack](https://github.com/opensearch-project/observability-stack)