@opensearch-project/agent-health 0.3.0 → 0.5.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +77 -6
- package/cli/dist/index.js +10072 -4502
- package/deployment/cloudformation/agent-health-observability.yaml +762 -0
- package/dist/assets/index-CCQRDlO0.js +243 -0
- package/dist/assets/index-CNHQVbcj.css +1 -0
- package/dist/index.html +2 -2
- package/docs/ARCHITECTURE.md +450 -0
- package/docs/BACKEND_JOB_QUEUE.md +405 -0
- package/docs/CLAUDE_CODE_TELEMETRY.md +283 -0
- package/docs/CLI.md +431 -0
- package/docs/CODING_AGENT_ANALYTICS.md +298 -0
- package/docs/CONFIGURATION.md +388 -0
- package/docs/CONNECTORS.md +536 -0
- package/docs/INSTRUMENT_WITH_OTEL.md +390 -0
- package/docs/ML-COMMONS-SETUP.md +289 -0
- package/docs/NPX_PACKAGING.md +195 -0
- package/docs/PERFORMANCE-MONITORING.md +200 -0
- package/docs/PERFORMANCE.md +390 -0
- package/docs/PI_PROFILING.md +169 -0
- package/docs/PLAN-non-agui-agent-support.md +525 -0
- package/docs/SDK.md +577 -0
- package/docs/SKILLS.md +264 -0
- package/docs/blogs/2026-02-28-opensearch-agent-health.md +200 -0
- package/docs/blogs/getting-started-blog.md +608 -0
- package/docs/diagrams/Agent-health.excalidraw +5656 -0
- package/docs/diagrams/architecture.png +0 -0
- package/docs/plans/field-redesign.md +468 -0
- package/docs/rfcs/001-coding-agent-analytics.md +374 -0
- package/docs/rfcs/002-enterprise-leaderboard.md +267 -0
- package/docs/rfcs/003-remote-aggregation.md +146 -0
- package/docs/rfcs/004-test-sdk-v2.md +599 -0
- package/docs/skills/AGENT_HEALTH.md +598 -0
- package/docs/skills/AGENT_PROFILE.md +191 -0
- package/docs/skills/add-connector/SKILL.md +68 -0
- package/docs/skills/agent-health-profile/SKILL.md +40 -0
- package/docs/skills/config-auth/SKILL.md +194 -0
- package/docs/skills/config-auth/evals/evals.json +35 -0
- package/docs/skills/create-pr/SKILL.md +73 -0
- package/docs/skills/instrument-otel/SKILL.md +84 -0
- package/docs/skills/write-test/SKILL.md +124 -0
- package/docs/ui prd.md +376 -0
- package/examples/README.md +53 -0
- package/examples/config/agent-health.config.example.ts +155 -0
- package/examples/connectors/echo-connector.ts +131 -0
- package/examples/eval-files/demo.eval.js +128 -0
- package/examples/eval-files/sdk-hooks-demo.eval.js +99 -0
- package/examples/pi-profiling/README.md +77 -0
- package/examples/pi-profiling/agent-health-profile.ts +417 -0
- package/lib/dist/lib/agentUtils.d.ts +29 -0
- package/lib/dist/lib/agentUtils.d.ts.map +1 -0
- package/lib/dist/lib/agentUtils.js +43 -0
- package/lib/dist/lib/agentUtils.js.map +1 -0
- package/lib/dist/lib/benchmarkExport.d.ts +14 -0
- package/lib/dist/lib/benchmarkExport.d.ts.map +1 -0
- package/lib/dist/lib/benchmarkExport.js +41 -0
- package/lib/dist/lib/benchmarkExport.js.map +1 -0
- package/lib/dist/lib/benchmarkVersionUtils.d.ts +37 -0
- package/lib/dist/lib/benchmarkVersionUtils.d.ts.map +1 -0
- package/lib/dist/lib/benchmarkVersionUtils.js +68 -0
- package/lib/dist/lib/benchmarkVersionUtils.js.map +1 -0
- package/lib/dist/lib/config/defineConfig.d.ts +27 -0
- package/lib/dist/lib/config/defineConfig.d.ts.map +1 -0
- package/lib/dist/lib/config/defineConfig.js +28 -0
- package/lib/dist/lib/config/defineConfig.js.map +1 -0
- package/lib/dist/lib/config/index.d.ts +9 -0
- package/lib/dist/lib/config/index.d.ts.map +1 -0
- package/lib/dist/lib/config/index.js +8 -0
- package/lib/dist/lib/config/index.js.map +1 -0
- package/lib/dist/lib/config/loader.d.ts +39 -0
- package/lib/dist/lib/config/loader.d.ts.map +1 -0
- package/lib/dist/lib/config/loader.js +258 -0
- package/lib/dist/lib/config/loader.js.map +1 -0
- package/lib/dist/lib/config/statePaths.d.ts +61 -0
- package/lib/dist/lib/config/statePaths.d.ts.map +1 -0
- package/lib/dist/lib/config/statePaths.js +188 -0
- package/lib/dist/lib/config/statePaths.js.map +1 -0
- package/lib/dist/lib/config/types.d.ts +231 -0
- package/lib/dist/lib/config/types.d.ts.map +1 -0
- package/lib/dist/lib/config/types.js +6 -0
- package/lib/dist/lib/config/types.js.map +1 -0
- package/lib/dist/lib/config.d.ts +39 -0
- package/lib/dist/lib/config.d.ts.map +1 -0
- package/lib/dist/lib/config.js +118 -0
- package/lib/dist/lib/config.js.map +1 -0
- package/lib/dist/lib/constants.d.ts +70 -0
- package/lib/dist/lib/constants.d.ts.map +1 -0
- package/lib/dist/lib/constants.js +365 -0
- package/lib/dist/lib/constants.js.map +1 -0
- package/lib/dist/lib/contextUtilization.d.ts +23 -0
- package/lib/dist/lib/contextUtilization.d.ts.map +1 -0
- package/lib/dist/lib/contextUtilization.js +72 -0
- package/lib/dist/lib/contextUtilization.js.map +1 -0
- package/lib/dist/lib/dashboardMetrics.d.ts +87 -0
- package/lib/dist/lib/dashboardMetrics.d.ts.map +1 -0
- package/lib/dist/lib/dashboardMetrics.js +242 -0
- package/lib/dist/lib/dashboardMetrics.js.map +1 -0
- package/lib/dist/lib/dataSourceConfig.d.ts +108 -0
- package/lib/dist/lib/dataSourceConfig.d.ts.map +1 -0
- package/lib/dist/lib/dataSourceConfig.js +166 -0
- package/lib/dist/lib/dataSourceConfig.js.map +1 -0
- package/lib/dist/lib/debug.d.ts +26 -0
- package/lib/dist/lib/debug.d.ts.map +1 -0
- package/lib/dist/lib/debug.js +132 -0
- package/lib/dist/lib/debug.js.map +1 -0
- package/lib/dist/lib/diagnostics.d.ts +28 -0
- package/lib/dist/lib/diagnostics.d.ts.map +1 -0
- package/lib/dist/lib/diagnostics.js +65 -0
- package/lib/dist/lib/diagnostics.js.map +1 -0
- package/lib/dist/lib/envCompat.d.ts +27 -0
- package/lib/dist/lib/envCompat.d.ts.map +1 -0
- package/lib/dist/lib/envCompat.js +73 -0
- package/lib/dist/lib/envCompat.js.map +1 -0
- package/lib/dist/lib/findPackageRoot.d.ts +7 -0
- package/lib/dist/lib/findPackageRoot.d.ts.map +1 -0
- package/lib/dist/lib/findPackageRoot.js +57 -0
- package/lib/dist/lib/findPackageRoot.js.map +1 -0
- package/lib/dist/lib/hooks.d.ts +36 -0
- package/lib/dist/lib/hooks.d.ts.map +1 -0
- package/lib/dist/lib/hooks.js +112 -0
- package/lib/dist/lib/hooks.js.map +1 -0
- package/lib/dist/lib/index.d.ts +47 -0
- package/lib/dist/lib/index.d.ts.map +1 -0
- package/lib/dist/lib/index.js +62 -0
- package/lib/dist/lib/index.js.map +1 -0
- package/lib/dist/lib/labels.d.ts +90 -0
- package/lib/dist/lib/labels.d.ts.map +1 -0
- package/lib/dist/lib/labels.js +158 -0
- package/lib/dist/lib/labels.js.map +1 -0
- package/lib/dist/lib/markdown.d.ts +16 -0
- package/lib/dist/lib/markdown.d.ts.map +1 -0
- package/lib/dist/lib/markdown.js +42 -0
- package/lib/dist/lib/markdown.js.map +1 -0
- package/lib/dist/lib/matchers/expect.d.ts +3 -0
- package/lib/dist/lib/matchers/expect.d.ts.map +1 -0
- package/lib/dist/lib/matchers/expect.js +225 -0
- package/lib/dist/lib/matchers/expect.js.map +1 -0
- package/lib/dist/lib/matchers/index.d.ts +8 -0
- package/lib/dist/lib/matchers/index.d.ts.map +1 -0
- package/lib/dist/lib/matchers/index.js +9 -0
- package/lib/dist/lib/matchers/index.js.map +1 -0
- package/lib/dist/lib/matchers/judgeAccessor.d.ts +113 -0
- package/lib/dist/lib/matchers/judgeAccessor.d.ts.map +1 -0
- package/lib/dist/lib/matchers/judgeAccessor.js +183 -0
- package/lib/dist/lib/matchers/judgeAccessor.js.map +1 -0
- package/lib/dist/lib/matchers/session.d.ts +39 -0
- package/lib/dist/lib/matchers/session.d.ts.map +1 -0
- package/lib/dist/lib/matchers/session.js +116 -0
- package/lib/dist/lib/matchers/session.js.map +1 -0
- package/lib/dist/lib/matchers/traces.d.ts +55 -0
- package/lib/dist/lib/matchers/traces.d.ts.map +1 -0
- package/lib/dist/lib/matchers/traces.js +116 -0
- package/lib/dist/lib/matchers/traces.js.map +1 -0
- package/lib/dist/lib/matchers/types.d.ts +75 -0
- package/lib/dist/lib/matchers/types.d.ts.map +1 -0
- package/lib/dist/lib/matchers/types.js +6 -0
- package/lib/dist/lib/matchers/types.js.map +1 -0
- package/lib/dist/lib/packagePaths.d.ts +29 -0
- package/lib/dist/lib/packagePaths.d.ts.map +1 -0
- package/lib/dist/lib/packagePaths.js +63 -0
- package/lib/dist/lib/packagePaths.js.map +1 -0
- package/lib/dist/lib/performance.d.ts +51 -0
- package/lib/dist/lib/performance.d.ts.map +1 -0
- package/lib/dist/lib/performance.js +159 -0
- package/lib/dist/lib/performance.js.map +1 -0
- package/lib/dist/lib/portConfig.d.ts +29 -0
- package/lib/dist/lib/portConfig.d.ts.map +1 -0
- package/lib/dist/lib/portConfig.js +64 -0
- package/lib/dist/lib/portConfig.js.map +1 -0
- package/lib/dist/lib/preferences.d.ts +63 -0
- package/lib/dist/lib/preferences.d.ts.map +1 -0
- package/lib/dist/lib/preferences.js +117 -0
- package/lib/dist/lib/preferences.js.map +1 -0
- package/lib/dist/lib/resolveAgentModel.d.ts +22 -0
- package/lib/dist/lib/resolveAgentModel.d.ts.map +1 -0
- package/lib/dist/lib/resolveAgentModel.js +37 -0
- package/lib/dist/lib/resolveAgentModel.js.map +1 -0
- package/lib/dist/lib/runStats.d.ts +92 -0
- package/lib/dist/lib/runStats.d.ts.map +1 -0
- package/lib/dist/lib/runStats.js +160 -0
- package/lib/dist/lib/runStats.js.map +1 -0
- package/lib/dist/lib/telemetry/constants.d.ts +60 -0
- package/lib/dist/lib/telemetry/constants.d.ts.map +1 -0
- package/lib/dist/lib/telemetry/constants.js +87 -0
- package/lib/dist/lib/telemetry/constants.js.map +1 -0
- package/lib/dist/lib/telemetry/evalSpans.d.ts +61 -0
- package/lib/dist/lib/telemetry/evalSpans.d.ts.map +1 -0
- package/lib/dist/lib/telemetry/evalSpans.js +254 -0
- package/lib/dist/lib/telemetry/evalSpans.js.map +1 -0
- package/lib/dist/lib/telemetry/index.d.ts +11 -0
- package/lib/dist/lib/telemetry/index.d.ts.map +1 -0
- package/lib/dist/lib/telemetry/index.js +15 -0
- package/lib/dist/lib/telemetry/index.js.map +1 -0
- package/lib/dist/lib/telemetry/opensearchExporter.d.ts +43 -0
- package/lib/dist/lib/telemetry/opensearchExporter.d.ts.map +1 -0
- package/lib/dist/lib/telemetry/opensearchExporter.js +217 -0
- package/lib/dist/lib/telemetry/opensearchExporter.js.map +1 -0
- package/lib/dist/lib/telemetry/provider.d.ts +55 -0
- package/lib/dist/lib/telemetry/provider.d.ts.map +1 -0
- package/lib/dist/lib/telemetry/provider.js +140 -0
- package/lib/dist/lib/telemetry/provider.js.map +1 -0
- package/lib/dist/lib/testCaseLabels.d.ts +34 -0
- package/lib/dist/lib/testCaseLabels.d.ts.map +1 -0
- package/lib/dist/lib/testCaseLabels.js +88 -0
- package/lib/dist/lib/testCaseLabels.js.map +1 -0
- package/lib/dist/lib/testCaseValidation.d.ts +140 -0
- package/lib/dist/lib/testCaseValidation.d.ts.map +1 -0
- package/lib/dist/lib/testCaseValidation.js +162 -0
- package/lib/dist/lib/testCaseValidation.js.map +1 -0
- package/lib/dist/lib/testCases/agentFixture.d.ts +80 -0
- package/lib/dist/lib/testCases/agentFixture.d.ts.map +1 -0
- package/lib/dist/lib/testCases/agentFixture.js +43 -0
- package/lib/dist/lib/testCases/agentFixture.js.map +1 -0
- package/lib/dist/lib/testCases/authoringSurface.d.ts +10 -0
- package/lib/dist/lib/testCases/authoringSurface.d.ts.map +1 -0
- package/lib/dist/lib/testCases/authoringSurface.js +54 -0
- package/lib/dist/lib/testCases/authoringSurface.js.map +1 -0
- package/lib/dist/lib/testCases/codemod.d.ts +13 -0
- package/lib/dist/lib/testCases/codemod.d.ts.map +1 -0
- package/lib/dist/lib/testCases/codemod.js +169 -0
- package/lib/dist/lib/testCases/codemod.js.map +1 -0
- package/lib/dist/lib/testCases/define.d.ts +114 -0
- package/lib/dist/lib/testCases/define.d.ts.map +1 -0
- package/lib/dist/lib/testCases/define.js +253 -0
- package/lib/dist/lib/testCases/define.js.map +1 -0
- package/lib/dist/lib/testCases/evaluators.d.ts +80 -0
- package/lib/dist/lib/testCases/evaluators.d.ts.map +1 -0
- package/lib/dist/lib/testCases/evaluators.js +105 -0
- package/lib/dist/lib/testCases/evaluators.js.map +1 -0
- package/lib/dist/lib/testCases/index.d.ts +14 -0
- package/lib/dist/lib/testCases/index.d.ts.map +1 -0
- package/lib/dist/lib/testCases/index.js +12 -0
- package/lib/dist/lib/testCases/index.js.map +1 -0
- package/lib/dist/lib/testCases/judge.d.ts +165 -0
- package/lib/dist/lib/testCases/judge.d.ts.map +1 -0
- package/lib/dist/lib/testCases/judge.js +359 -0
- package/lib/dist/lib/testCases/judge.js.map +1 -0
- package/lib/dist/lib/testCases/loader.d.ts +26 -0
- package/lib/dist/lib/testCases/loader.d.ts.map +1 -0
- package/lib/dist/lib/testCases/loader.js +149 -0
- package/lib/dist/lib/testCases/loader.js.map +1 -0
- package/lib/dist/lib/testCases/types.d.ts +242 -0
- package/lib/dist/lib/testCases/types.d.ts.map +1 -0
- package/lib/dist/lib/testCases/types.js +6 -0
- package/lib/dist/lib/testCases/types.js.map +1 -0
- package/lib/dist/lib/theme.d.ts +6 -0
- package/lib/dist/lib/theme.d.ts.map +1 -0
- package/lib/dist/lib/theme.js +36 -0
- package/lib/dist/lib/theme.js.map +1 -0
- package/lib/dist/lib/uiTelemetry.d.ts +7 -0
- package/lib/dist/lib/uiTelemetry.d.ts.map +1 -0
- package/lib/dist/lib/uiTelemetry.js +25 -0
- package/lib/dist/lib/uiTelemetry.js.map +1 -0
- package/lib/dist/lib/utils.d.ts +96 -0
- package/lib/dist/lib/utils.d.ts.map +1 -0
- package/lib/dist/lib/utils.js +232 -0
- package/lib/dist/lib/utils.js.map +1 -0
- package/lib/dist/lib/workflow/consolidate.d.ts +12 -0
- package/lib/dist/lib/workflow/consolidate.d.ts.map +1 -0
- package/lib/dist/lib/workflow/consolidate.js +33 -0
- package/lib/dist/lib/workflow/consolidate.js.map +1 -0
- package/lib/dist/lib/workflow/index.d.ts +13 -0
- package/lib/dist/lib/workflow/index.d.ts.map +1 -0
- package/lib/dist/lib/workflow/index.js +12 -0
- package/lib/dist/lib/workflow/index.js.map +1 -0
- package/lib/dist/lib/workflow/ledger.d.ts +30 -0
- package/lib/dist/lib/workflow/ledger.d.ts.map +1 -0
- package/lib/dist/lib/workflow/ledger.js +41 -0
- package/lib/dist/lib/workflow/ledger.js.map +1 -0
- package/lib/dist/lib/workflow/pool.d.ts +13 -0
- package/lib/dist/lib/workflow/pool.d.ts.map +1 -0
- package/lib/dist/lib/workflow/pool.js +44 -0
- package/lib/dist/lib/workflow/pool.js.map +1 -0
- package/lib/dist/lib/workflow/source.d.ts +22 -0
- package/lib/dist/lib/workflow/source.d.ts.map +1 -0
- package/lib/dist/lib/workflow/source.js +29 -0
- package/lib/dist/lib/workflow/source.js.map +1 -0
- package/lib/dist/lib/workflow/stepB.d.ts +71 -0
- package/lib/dist/lib/workflow/stepB.d.ts.map +1 -0
- package/lib/dist/lib/workflow/stepB.js +99 -0
- package/lib/dist/lib/workflow/stepB.js.map +1 -0
- package/lib/dist/lib/workflow/types.d.ts +86 -0
- package/lib/dist/lib/workflow/types.d.ts.map +1 -0
- package/lib/dist/lib/workflow/types.js +6 -0
- package/lib/dist/lib/workflow/types.js.map +1 -0
- package/lib/dist/lib/workflow/workflow.d.ts +119 -0
- package/lib/dist/lib/workflow/workflow.d.ts.map +1 -0
- package/lib/dist/lib/workflow/workflow.js +195 -0
- package/lib/dist/lib/workflow/workflow.js.map +1 -0
- package/lib/dist/services/agent/aguiConverter.d.ts +50 -0
- package/lib/dist/services/agent/aguiConverter.d.ts.map +1 -0
- package/lib/dist/services/agent/aguiConverter.js +449 -0
- package/lib/dist/services/agent/aguiConverter.js.map +1 -0
- package/lib/dist/services/agent/index.d.ts +10 -0
- package/lib/dist/services/agent/index.d.ts.map +1 -0
- package/lib/dist/services/agent/index.js +12 -0
- package/lib/dist/services/agent/index.js.map +1 -0
- package/lib/dist/services/agent/payloadBuilder.d.ts +33 -0
- package/lib/dist/services/agent/payloadBuilder.d.ts.map +1 -0
- package/lib/dist/services/agent/payloadBuilder.js +75 -0
- package/lib/dist/services/agent/payloadBuilder.js.map +1 -0
- package/lib/dist/services/agent/sseStream.d.ts +43 -0
- package/lib/dist/services/agent/sseStream.d.ts.map +1 -0
- package/lib/dist/services/agent/sseStream.js +223 -0
- package/lib/dist/services/agent/sseStream.js.map +1 -0
- package/lib/dist/services/connectors/agui/AGUIStreamingConnector.d.ts +44 -0
- package/lib/dist/services/connectors/agui/AGUIStreamingConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/agui/AGUIStreamingConnector.js +95 -0
- package/lib/dist/services/connectors/agui/AGUIStreamingConnector.js.map +1 -0
- package/lib/dist/services/connectors/base/BaseConnector.d.ts +81 -0
- package/lib/dist/services/connectors/base/BaseConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/base/BaseConnector.js +170 -0
- package/lib/dist/services/connectors/base/BaseConnector.js.map +1 -0
- package/lib/dist/services/connectors/claude-code/ClaudeCodeConnector.d.ts +116 -0
- package/lib/dist/services/connectors/claude-code/ClaudeCodeConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/claude-code/ClaudeCodeConnector.js +403 -0
- package/lib/dist/services/connectors/claude-code/ClaudeCodeConnector.js.map +1 -0
- package/lib/dist/services/connectors/index.d.ts +13 -0
- package/lib/dist/services/connectors/index.d.ts.map +1 -0
- package/lib/dist/services/connectors/index.js +32 -0
- package/lib/dist/services/connectors/index.js.map +1 -0
- package/lib/dist/services/connectors/kiro/KiroConnector.d.ts +48 -0
- package/lib/dist/services/connectors/kiro/KiroConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/kiro/KiroConnector.js +158 -0
- package/lib/dist/services/connectors/kiro/KiroConnector.js.map +1 -0
- package/lib/dist/services/connectors/langgraph/LangGraphConnector.d.ts +36 -0
- package/lib/dist/services/connectors/langgraph/LangGraphConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/langgraph/LangGraphConnector.js +175 -0
- package/lib/dist/services/connectors/langgraph/LangGraphConnector.js.map +1 -0
- package/lib/dist/services/connectors/mock/MockConnector.d.ts +37 -0
- package/lib/dist/services/connectors/mock/MockConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/mock/MockConnector.js +120 -0
- package/lib/dist/services/connectors/mock/MockConnector.js.map +1 -0
- package/lib/dist/services/connectors/openai-compatible/OpenAICompatibleConnector.d.ts +42 -0
- package/lib/dist/services/connectors/openai-compatible/OpenAICompatibleConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/openai-compatible/OpenAICompatibleConnector.js +133 -0
- package/lib/dist/services/connectors/openai-compatible/OpenAICompatibleConnector.js.map +1 -0
- package/lib/dist/services/connectors/pi/PiConnector.d.ts +87 -0
- package/lib/dist/services/connectors/pi/PiConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/pi/PiConnector.js +274 -0
- package/lib/dist/services/connectors/pi/PiConnector.js.map +1 -0
- package/lib/dist/services/connectors/registry.d.ts +57 -0
- package/lib/dist/services/connectors/registry.d.ts.map +1 -0
- package/lib/dist/services/connectors/registry.js +106 -0
- package/lib/dist/services/connectors/registry.js.map +1 -0
- package/lib/dist/services/connectors/rest/RESTConnector.d.ts +38 -0
- package/lib/dist/services/connectors/rest/RESTConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/rest/RESTConnector.js +117 -0
- package/lib/dist/services/connectors/rest/RESTConnector.js.map +1 -0
- package/lib/dist/services/connectors/server.d.ts +13 -0
- package/lib/dist/services/connectors/server.d.ts.map +1 -0
- package/lib/dist/services/connectors/server.js +34 -0
- package/lib/dist/services/connectors/server.js.map +1 -0
- package/lib/dist/services/connectors/strands/StrandsConnector.d.ts +48 -0
- package/lib/dist/services/connectors/strands/StrandsConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/strands/StrandsConnector.js +221 -0
- package/lib/dist/services/connectors/strands/StrandsConnector.js.map +1 -0
- package/lib/dist/services/connectors/subprocess/SubprocessConnector.d.ts +88 -0
- package/lib/dist/services/connectors/subprocess/SubprocessConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/subprocess/SubprocessConnector.js +418 -0
- package/lib/dist/services/connectors/subprocess/SubprocessConnector.js.map +1 -0
- package/lib/dist/services/connectors/types.d.ts +213 -0
- package/lib/dist/services/connectors/types.d.ts.map +1 -0
- package/lib/dist/services/connectors/types.js +6 -0
- package/lib/dist/services/connectors/types.js.map +1 -0
- package/lib/dist/services/evaluation/bedrockJudge.d.ts +64 -0
- package/lib/dist/services/evaluation/bedrockJudge.d.ts.map +1 -0
- package/lib/dist/services/evaluation/bedrockJudge.js +167 -0
- package/lib/dist/services/evaluation/bedrockJudge.js.map +1 -0
- package/lib/dist/services/evaluation/evaluatorError.d.ts +56 -0
- package/lib/dist/services/evaluation/evaluatorError.d.ts.map +1 -0
- package/lib/dist/services/evaluation/evaluatorError.js +56 -0
- package/lib/dist/services/evaluation/evaluatorError.js.map +1 -0
- package/lib/dist/services/evaluation/index.d.ts +106 -0
- package/lib/dist/services/evaluation/index.d.ts.map +1 -0
- package/lib/dist/services/evaluation/index.js +684 -0
- package/lib/dist/services/evaluation/index.js.map +1 -0
- package/lib/dist/services/evaluation/mockTrajectory.d.ts +3 -0
- package/lib/dist/services/evaluation/mockTrajectory.d.ts.map +1 -0
- package/lib/dist/services/evaluation/mockTrajectory.js +72 -0
- package/lib/dist/services/evaluation/mockTrajectory.js.map +1 -0
- package/lib/dist/services/opensearch/client.d.ts +26 -0
- package/lib/dist/services/opensearch/client.d.ts.map +1 -0
- package/lib/dist/services/opensearch/client.js +131 -0
- package/lib/dist/services/opensearch/client.js.map +1 -0
- package/lib/dist/services/opensearch/index.d.ts +16 -0
- package/lib/dist/services/opensearch/index.d.ts.map +1 -0
- package/lib/dist/services/opensearch/index.js +25 -0
- package/lib/dist/services/opensearch/index.js.map +1 -0
- package/lib/dist/services/storage/asyncBenchmarkStorage.d.ts +123 -0
- package/lib/dist/services/storage/asyncBenchmarkStorage.d.ts.map +1 -0
- package/lib/dist/services/storage/asyncBenchmarkStorage.js +429 -0
- package/lib/dist/services/storage/asyncBenchmarkStorage.js.map +1 -0
- package/lib/dist/services/storage/asyncRunStorage.d.ts +127 -0
- package/lib/dist/services/storage/asyncRunStorage.d.ts.map +1 -0
- package/lib/dist/services/storage/asyncRunStorage.js +448 -0
- package/lib/dist/services/storage/asyncRunStorage.js.map +1 -0
- package/lib/dist/services/storage/asyncTestCaseStorage.d.ts +156 -0
- package/lib/dist/services/storage/asyncTestCaseStorage.d.ts.map +1 -0
- package/lib/dist/services/storage/asyncTestCaseStorage.js +285 -0
- package/lib/dist/services/storage/asyncTestCaseStorage.js.map +1 -0
- package/lib/dist/services/storage/index.d.ts +17 -0
- package/lib/dist/services/storage/index.d.ts.map +1 -0
- package/lib/dist/services/storage/index.js +20 -0
- package/lib/dist/services/storage/index.js.map +1 -0
- package/lib/dist/services/storage/migration.d.ts +54 -0
- package/lib/dist/services/storage/migration.d.ts.map +1 -0
- package/lib/dist/services/storage/migration.js +296 -0
- package/lib/dist/services/storage/migration.js.map +1 -0
- package/lib/dist/services/storage/opensearchClient.d.ts +924 -0
- package/lib/dist/services/storage/opensearchClient.d.ts.map +1 -0
- package/lib/dist/services/storage/opensearchClient.js +435 -0
- package/lib/dist/services/storage/opensearchClient.js.map +1 -0
- package/lib/dist/services/traces/browserRecovery.d.ts +26 -0
- package/lib/dist/services/traces/browserRecovery.d.ts.map +1 -0
- package/lib/dist/services/traces/browserRecovery.js +81 -0
- package/lib/dist/services/traces/browserRecovery.js.map +1 -0
- package/lib/dist/services/traces/categoryStyles.d.ts +21 -0
- package/lib/dist/services/traces/categoryStyles.d.ts.map +1 -0
- package/lib/dist/services/traces/categoryStyles.js +56 -0
- package/lib/dist/services/traces/categoryStyles.js.map +1 -0
- package/lib/dist/services/traces/executionOrderTransform.d.ts +35 -0
- package/lib/dist/services/traces/executionOrderTransform.d.ts.map +1 -0
- package/lib/dist/services/traces/executionOrderTransform.js +313 -0
- package/lib/dist/services/traces/executionOrderTransform.js.map +1 -0
- package/lib/dist/services/traces/fetchSpansForRun.d.ts +86 -0
- package/lib/dist/services/traces/fetchSpansForRun.d.ts.map +1 -0
- package/lib/dist/services/traces/fetchSpansForRun.js +69 -0
- package/lib/dist/services/traces/fetchSpansForRun.js.map +1 -0
- package/lib/dist/services/traces/flowTransform.d.ts +24 -0
- package/lib/dist/services/traces/flowTransform.d.ts.map +1 -0
- package/lib/dist/services/traces/flowTransform.js +228 -0
- package/lib/dist/services/traces/flowTransform.js.map +1 -0
- package/lib/dist/services/traces/index.d.ts +121 -0
- package/lib/dist/services/traces/index.d.ts.map +1 -0
- package/lib/dist/services/traces/index.js +255 -0
- package/lib/dist/services/traces/index.js.map +1 -0
- package/lib/dist/services/traces/intentTransform.d.ts +20 -0
- package/lib/dist/services/traces/intentTransform.d.ts.map +1 -0
- package/lib/dist/services/traces/intentTransform.js +131 -0
- package/lib/dist/services/traces/intentTransform.js.map +1 -0
- package/lib/dist/services/traces/judgeAgentsHints.d.ts +63 -0
- package/lib/dist/services/traces/judgeAgentsHints.d.ts.map +1 -0
- package/lib/dist/services/traces/judgeAgentsHints.js +89 -0
- package/lib/dist/services/traces/judgeAgentsHints.js.map +1 -0
- package/lib/dist/services/traces/messageExtraction.d.ts +15 -0
- package/lib/dist/services/traces/messageExtraction.d.ts.map +1 -0
- package/lib/dist/services/traces/messageExtraction.js +251 -0
- package/lib/dist/services/traces/messageExtraction.js.map +1 -0
- package/lib/dist/services/traces/spanCategorization.d.ts +63 -0
- package/lib/dist/services/traces/spanCategorization.d.ts.map +1 -0
- package/lib/dist/services/traces/spanCategorization.js +276 -0
- package/lib/dist/services/traces/spanCategorization.js.map +1 -0
- package/lib/dist/services/traces/spanPreprocessing.d.ts +37 -0
- package/lib/dist/services/traces/spanPreprocessing.d.ts.map +1 -0
- package/lib/dist/services/traces/spanPreprocessing.js +102 -0
- package/lib/dist/services/traces/spanPreprocessing.js.map +1 -0
- package/lib/dist/services/traces/spansToTrajectory.d.ts +36 -0
- package/lib/dist/services/traces/spansToTrajectory.d.ts.map +1 -0
- package/lib/dist/services/traces/spansToTrajectory.js +387 -0
- package/lib/dist/services/traces/spansToTrajectory.js.map +1 -0
- package/lib/dist/services/traces/toolSimilarity.d.ts +35 -0
- package/lib/dist/services/traces/toolSimilarity.d.ts.map +1 -0
- package/lib/dist/services/traces/toolSimilarity.js +203 -0
- package/lib/dist/services/traces/toolSimilarity.js.map +1 -0
- package/lib/dist/services/traces/traceComparison.d.ts +31 -0
- package/lib/dist/services/traces/traceComparison.d.ts.map +1 -0
- package/lib/dist/services/traces/traceComparison.js +318 -0
- package/lib/dist/services/traces/traceComparison.js.map +1 -0
- package/lib/dist/services/traces/traceGrouping.d.ts +19 -0
- package/lib/dist/services/traces/traceGrouping.d.ts.map +1 -0
- package/lib/dist/services/traces/traceGrouping.js +107 -0
- package/lib/dist/services/traces/traceGrouping.js.map +1 -0
- package/lib/dist/services/traces/tracePoller.d.ts +84 -0
- package/lib/dist/services/traces/tracePoller.d.ts.map +1 -0
- package/lib/dist/services/traces/tracePoller.js +309 -0
- package/lib/dist/services/traces/tracePoller.js.map +1 -0
- package/lib/dist/services/traces/traceStats.d.ts +45 -0
- package/lib/dist/services/traces/traceStats.d.ts.map +1 -0
- package/lib/dist/services/traces/traceStats.js +114 -0
- package/lib/dist/services/traces/traceStats.js.map +1 -0
- package/lib/dist/services/traces/traceSummary.d.ts +47 -0
- package/lib/dist/services/traces/traceSummary.d.ts.map +1 -0
- package/lib/dist/services/traces/traceSummary.js +68 -0
- package/lib/dist/services/traces/traceSummary.js.map +1 -0
- package/lib/dist/services/traces/utils.d.ts +33 -0
- package/lib/dist/services/traces/utils.d.ts.map +1 -0
- package/lib/dist/services/traces/utils.js +114 -0
- package/lib/dist/services/traces/utils.js.map +1 -0
- package/lib/dist/types/agui.d.ts +13 -0
- package/lib/dist/types/agui.d.ts.map +1 -0
- package/lib/dist/types/agui.js +16 -0
- package/lib/dist/types/agui.js.map +1 -0
- package/lib/dist/types/index.d.ts +1175 -0
- package/lib/dist/types/index.d.ts.map +1 -0
- package/lib/dist/types/index.js +12 -0
- package/lib/dist/types/index.js.map +1 -0
- package/lib/dist/types/skills.d.ts +146 -0
- package/lib/dist/types/skills.d.ts.map +1 -0
- package/lib/dist/types/skills.js +6 -0
- package/lib/dist/types/skills.js.map +1 -0
- package/observio-sample-agent/pi-package/README.md +112 -0
- package/observio-sample-agent/pi-package/extensions/agent-health.ts +373 -0
- package/observio-sample-agent/pi-package/package.json +17 -0
- package/observio-sample-agent/pi-package/prompts/agent-health.md +37 -0
- package/observio-sample-agent/pi-package/skills/create-pr/SKILL.md +88 -0
- package/observio-sample-agent/pi-package/skills/fix-bug/SKILL.md +71 -0
- package/observio-sample-agent/pi-package/skills/implement-feature/SKILL.md +156 -0
- package/observio-sample-agent/pi-package/skills/instrument-otel/SKILL.md +208 -0
- package/observio-sample-agent/pi-package/skills/setup-collector/SKILL.md +146 -0
- package/observio-sample-agent/pi-package/skills/write-test/SKILL.md +115 -0
- package/package.json +64 -13
- package/server/dist/app.js +32651 -17637
- package/server/dist/index.js +29875 -14638
- package/tsconfig.lib.json +71 -0
- package/dist/assets/index-EvPLSTAS.js +0 -267
- package/dist/assets/index-RXasQKUs.css +0 -1
- package/lib/dist/config/index.js +0 -404
- package/lib/dist/index.js +0 -1665
|
@@ -0,0 +1,608 @@
|
|
|
1
|
+
## Getting Started with Agent Health: A Complete Walkthrough
|
|
2
|
+
|
|
3
|
+
In our [introductory blog post](https://opensearch.org/blog/opensearch-agent-health/), we showed you what Agent Health is and why it matters. Now let's roll up our sleeves and go through a hands-on walkthrough.
|
|
4
|
+
|
|
5
|
+
This guide is progressive — you can stop at any point and come back later:
|
|
6
|
+
|
|
7
|
+
1. **Try sample data** — explore the UI with zero setup
|
|
8
|
+
2. **Try it yourself** — connect your own agent endpoint, configure a judge, and run your first evaluation (no OpenSearch required)
|
|
9
|
+
3. **Add tracing** — connect or instrument OpenTelemetry traces for deep observability
|
|
10
|
+
|
|
11
|
+
**What you'll need to start:** You have an agent (any protocol), an LLM provider, and optionally tracing via OpenTelemetry to some backend. Agent Health plugs into all of these.
|
|
12
|
+
|
|
13
|
+
<!-- TODO: Diagram — high-level: Agent + LLM + (optional) OTel tracing → Agent Health -->
|
|
14
|
+
|
|
15
|
+
### Prerequisites
|
|
16
|
+
* [Node.js](https://nodejs.org/) 18 or later
|
|
17
|
+
* [Docker](https://www.docker.com/) (optional — for the local OpenSearch stack or the Docker Compose setup)
|
|
18
|
+
* AWS credentials (for the Bedrock LLM judge) or an OpenAI-compatible endpoint (for LiteLLM, Ollama, etc.)
|
|
19
|
+
|
|
20
|
+
### Launch Agent Health
|
|
21
|
+
|
|
22
|
+
```
|
|
23
|
+
npx @opensearch-project/agent-health
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
Open your browser to `http://localhost:4001`. You'll land on the Agent Health home screen with three main sections: **Traces**, **Benchmarks**, and **Compare**.
|
|
27
|
+
|
|
28
|
+
<!-- TODO: Screenshot — Agent Health home screen -->
|
|
29
|
+
|
|
30
|
+
---
|
|
31
|
+
|
|
32
|
+
## Part I: Try Sample Data
|
|
33
|
+
|
|
34
|
+
Before connecting your own agent, explore with the pre-loaded demo data. This gives you a feel for the interface with zero configuration.
|
|
35
|
+
|
|
36
|
+
### Traces view
|
|
37
|
+
|
|
38
|
+
Navigate to the **Traces** tab. You'll see pre-loaded agent execution traces. Click any trace to open the detail view:
|
|
39
|
+
* **Timeline view** — a chronological breakdown of every span (LLM call, tool invocation, retrieval step) with durations
|
|
40
|
+
* **Flow view** — a visual graph of how data flows between agent components
|
|
41
|
+
* **Span details** — click any span to see its attributes, inputs, outputs, and metadata
|
|
42
|
+
|
|
43
|
+
<!-- TODO: Screenshot — Traces list + detail view -->
|
|
44
|
+
|
|
45
|
+
### Benchmarks view
|
|
46
|
+
|
|
47
|
+
Navigate to **Benchmarks**. You'll find a demo benchmark called "Travel Planning Accuracy - Demo". Click into it to see the test cases, then run it to watch the LLM judge evaluate each case in real time.
|
|
48
|
+
|
|
49
|
+
<!-- TODO: Screenshot — Benchmark detail with test cases -->
|
|
50
|
+
|
|
51
|
+
### Compare view
|
|
52
|
+
|
|
53
|
+
After running a benchmark, go to **Compare**. Select two runs to see side-by-side metrics: pass rate, latency, cost, and per-test-case diffs.
|
|
54
|
+
|
|
55
|
+
<!-- TODO: Screenshot — Compare view side-by-side -->
|
|
56
|
+
|
|
57
|
+
---
|
|
58
|
+
|
|
59
|
+
## Part II: Try It Yourself
|
|
60
|
+
|
|
61
|
+
Now let's connect your own agent and run a real evaluation. No OpenSearch storage required — Agent Health stores data locally on disk by default.
|
|
62
|
+
|
|
63
|
+
<!-- TODO: Diagram — build-up: Agent Health box, now adding "Your Agent" arrow -->
|
|
64
|
+
|
|
65
|
+
### Step 1: Configure your agent endpoint
|
|
66
|
+
|
|
67
|
+
Create an `agent-health.config.ts` file in your working directory (or run `npx @opensearch-project/agent-health init` to generate one):
|
|
68
|
+
|
|
69
|
+
```typescript
|
|
70
|
+
export default {
|
|
71
|
+
agents: [
|
|
72
|
+
{
|
|
73
|
+
key: "my-agent",
|
|
74
|
+
name: "My Agent",
|
|
75
|
+
endpoint: "http://localhost:3000/agent",
|
|
76
|
+
connectorType: "agui-streaming",
|
|
77
|
+
models: ["claude-sonnet-4"],
|
|
78
|
+
useTraces: false, // No OpenSearch tracing required for basic evaluation
|
|
79
|
+
}
|
|
80
|
+
],
|
|
81
|
+
};
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
**Connector types** determine how Agent Health communicates with your agent. They fall into three categories:
|
|
85
|
+
|
|
86
|
+
**Agent Protocol** — connectors that know the full request/response contract of a specific agent framework:
|
|
87
|
+
|
|
88
|
+
| Connector | Protocol |
|
|
89
|
+
|-----------|----------|
|
|
90
|
+
| `agui-streaming` | [AG-UI](https://docs.ag-ui.com) SSE streaming protocol (default) |
|
|
91
|
+
| `claude-code` | Claude Code CLI — NDJSON streaming with MCP tool support |
|
|
92
|
+
|
|
93
|
+
**Transport** — generic communication channels where you control the payload format via hooks or a custom connector:
|
|
94
|
+
|
|
95
|
+
| Connector | Protocol |
|
|
96
|
+
|-----------|----------|
|
|
97
|
+
| `rest` | Standard HTTP POST — Agent Health sends JSON and parses common response shapes |
|
|
98
|
+
| `cli` | Spawns a CLI command as a child process, captures stdout |
|
|
99
|
+
|
|
100
|
+
**LLM Protocol** — talks directly to an LLM endpoint (useful for evaluating raw model responses, not a full agent loop):
|
|
101
|
+
|
|
102
|
+
| Connector | Protocol |
|
|
103
|
+
|-----------|----------|
|
|
104
|
+
| `openai-compatible` | OpenAI Chat Completions standard: `POST /v1/chat/completions` with `messages` array. Works with LiteLLM, Ollama, vLLM, Azure OpenAI, OpenAI, etc. |
|
|
105
|
+
|
|
106
|
+
**Testing:**
|
|
107
|
+
|
|
108
|
+
| Connector | Protocol |
|
|
109
|
+
|-----------|----------|
|
|
110
|
+
| `mock` | In-memory demo trajectory, no real agent needed |
|
|
111
|
+
|
|
112
|
+
#### Writing your own connector
|
|
113
|
+
|
|
114
|
+
If none of the built-in connectors fit your agent's protocol, you can write a custom one. Create a class that extends `BaseConnector` and implement three methods: `buildPayload`, `execute`, and `parseResponse`.
|
|
115
|
+
|
|
116
|
+
```typescript
|
|
117
|
+
import { BaseConnector } from '@opensearch-project/agent-health/connectors';
|
|
118
|
+
import type {
|
|
119
|
+
ConnectorAuth,
|
|
120
|
+
ConnectorRequest,
|
|
121
|
+
ConnectorResponse,
|
|
122
|
+
ConnectorProgressCallback,
|
|
123
|
+
ConnectorRawEventCallback,
|
|
124
|
+
} from '@opensearch-project/agent-health/connectors';
|
|
125
|
+
import type { TrajectoryStep } from '@opensearch-project/agent-health/types';
|
|
126
|
+
|
|
127
|
+
class MyCustomConnector extends BaseConnector {
|
|
128
|
+
readonly type = 'my-protocol' as const;
|
|
129
|
+
readonly name = 'My Custom Protocol';
|
|
130
|
+
readonly supportsStreaming = false;
|
|
131
|
+
|
|
132
|
+
buildPayload(request: ConnectorRequest): any {
|
|
133
|
+
// Transform the standard request into your agent's expected format
|
|
134
|
+
return {
|
|
135
|
+
query: request.testCase.initialPrompt,
|
|
136
|
+
model: request.modelId,
|
|
137
|
+
};
|
|
138
|
+
}
|
|
139
|
+
|
|
140
|
+
async execute(
|
|
141
|
+
endpoint: string,
|
|
142
|
+
request: ConnectorRequest,
|
|
143
|
+
auth: ConnectorAuth,
|
|
144
|
+
onProgress?: ConnectorProgressCallback,
|
|
145
|
+
onRawEvent?: ConnectorRawEventCallback
|
|
146
|
+
): Promise<ConnectorResponse> {
|
|
147
|
+
const payload = request.payload || this.buildPayload(request);
|
|
148
|
+
const headers = this.buildAuthHeaders(auth);
|
|
149
|
+
|
|
150
|
+
const response = await fetch(endpoint, {
|
|
151
|
+
method: 'POST',
|
|
152
|
+
headers: { 'Content-Type': 'application/json', ...headers },
|
|
153
|
+
body: JSON.stringify(payload),
|
|
154
|
+
});
|
|
155
|
+
|
|
156
|
+
const data = await response.json();
|
|
157
|
+
onRawEvent?.(data);
|
|
158
|
+
|
|
159
|
+
const trajectory = this.parseResponse(data);
|
|
160
|
+
trajectory.forEach(step => onProgress?.(step));
|
|
161
|
+
|
|
162
|
+
return { trajectory, runId: data.id || null };
|
|
163
|
+
}
|
|
164
|
+
|
|
165
|
+
parseResponse(data: any): TrajectoryStep[] {
|
|
166
|
+
// Convert your agent's response format into TrajectoryStep array
|
|
167
|
+
return [
|
|
168
|
+
this.createStep('thinking', data.reasoning || ''),
|
|
169
|
+
this.createStep('response', data.answer || JSON.stringify(data)),
|
|
170
|
+
];
|
|
171
|
+
}
|
|
172
|
+
}
|
|
173
|
+
```
|
|
174
|
+
|
|
175
|
+
Register your connector in `agent-health.config.ts`:
|
|
176
|
+
|
|
177
|
+
```typescript
|
|
178
|
+
export default {
|
|
179
|
+
connectors: [new MyCustomConnector()],
|
|
180
|
+
agents: [
|
|
181
|
+
{
|
|
182
|
+
key: "my-agent",
|
|
183
|
+
name: "My Agent",
|
|
184
|
+
endpoint: "http://localhost:3000/agent",
|
|
185
|
+
connectorType: "my-protocol",
|
|
186
|
+
models: ["claude-sonnet-4"],
|
|
187
|
+
useTraces: false,
|
|
188
|
+
}
|
|
189
|
+
],
|
|
190
|
+
};
|
|
191
|
+
```
|
|
192
|
+
|
|
193
|
+
#### Lifecycle hooks
|
|
194
|
+
|
|
195
|
+
For simpler customizations — adding auth tokens, modifying payloads, pre-creating threads — you can use lifecycle hooks without writing a full connector:
|
|
196
|
+
|
|
197
|
+
```typescript
|
|
198
|
+
{
|
|
199
|
+
key: "my-agent",
|
|
200
|
+
// ...
|
|
201
|
+
hooks: {
|
|
202
|
+
beforeRequest: async ({ endpoint, payload, headers }) => {
|
|
203
|
+
// e.g., add auth tokens, modify payload, pre-create threads
|
|
204
|
+
return { endpoint, payload, headers };
|
|
205
|
+
},
|
|
206
|
+
},
|
|
207
|
+
}
|
|
208
|
+
```
|
|
209
|
+
|
|
210
|
+
<!-- TODO: Diagram — build-up: Agent Health ↔ Your Agent (with connector arrow) -->
|
|
211
|
+
|
|
212
|
+
### Step 2: Configure the judge
|
|
213
|
+
|
|
214
|
+
The LLM judge evaluates your agent's responses against expected outcomes. Agent Health supports two judge providers:
|
|
215
|
+
|
|
216
|
+
**Option A: AWS Bedrock (default)**
|
|
217
|
+
|
|
218
|
+
```bash
|
|
219
|
+
# AWS profile (recommended)
|
|
220
|
+
export AWS_PROFILE=your-profile
|
|
221
|
+
export AWS_REGION=us-east-1
|
|
222
|
+
|
|
223
|
+
# Or explicit credentials
|
|
224
|
+
export AWS_ACCESS_KEY_ID=...
|
|
225
|
+
export AWS_SECRET_ACCESS_KEY=...
|
|
226
|
+
export AWS_SESSION_TOKEN=...
|
|
227
|
+
```
|
|
228
|
+
|
|
229
|
+
**Option B: OpenAI-compatible endpoint** — any provider that implements the OpenAI Chat Completions standard (`POST /v1/chat/completions` with `messages` array). This includes LiteLLM, Ollama, vLLM, Azure OpenAI, or OpenAI directly.
|
|
230
|
+
|
|
231
|
+
```bash
|
|
232
|
+
export OPENAI_COMPATIBLE_ENDPOINT=http://localhost:4000/v1/chat/completions
|
|
233
|
+
export OPENAI_COMPATIBLE_API_KEY=your-api-key # optional, depends on provider
|
|
234
|
+
```
|
|
235
|
+
|
|
236
|
+
You can also configure the judge in `agent-health.config.ts`:
|
|
237
|
+
|
|
238
|
+
```typescript
|
|
239
|
+
export default {
|
|
240
|
+
judge: {
|
|
241
|
+
provider: "openai-compatible", // or "bedrock" (default)
|
|
242
|
+
model: "gpt-4o", // model name forwarded to the endpoint
|
|
243
|
+
},
|
|
244
|
+
// ...agents, etc.
|
|
245
|
+
};
|
|
246
|
+
```
|
|
247
|
+
|
|
248
|
+
<!-- TODO: This section will expand once judge configuration UI is implemented -->
|
|
249
|
+
|
|
250
|
+
<!-- TODO: Diagram — build-up: Agent Health ↔ Your Agent, Agent Health ↔ Judge -->
|
|
251
|
+
|
|
252
|
+
### Step 3: Run your first evaluation
|
|
253
|
+
|
|
254
|
+
With your agent endpoint and judge configured, you can run an evaluation — no OpenSearch storage or tracing required.
|
|
255
|
+
|
|
256
|
+
#### Create test cases
|
|
257
|
+
|
|
258
|
+
<!-- TODO: Replace these sample test cases with ones based on our own shipped evaluation engine / sample agent -->
|
|
259
|
+
|
|
260
|
+
Create a file called `travel-benchmark.json` with test cases for your agent:
|
|
261
|
+
|
|
262
|
+
```json
|
|
263
|
+
[
|
|
264
|
+
{
|
|
265
|
+
"name": "Basic flight search",
|
|
266
|
+
"description": "User asks for a simple flight search",
|
|
267
|
+
"labels": ["category:Travel", "difficulty:Easy"],
|
|
268
|
+
"initialPrompt": "Find me flights from Seattle to New York next Friday",
|
|
269
|
+
"expectedOutcomes": [
|
|
270
|
+
"Agent should call search_flights tool with origin=Seattle and destination=New York",
|
|
271
|
+
"Agent should present flight options with times and prices"
|
|
272
|
+
]
|
|
273
|
+
},
|
|
274
|
+
{
|
|
275
|
+
"name": "Hotel search with dates",
|
|
276
|
+
"description": "User asks for hotel availability",
|
|
277
|
+
"labels": ["category:Travel", "difficulty:Easy"],
|
|
278
|
+
"initialPrompt": "Are there any hotels available in Manhattan for March 20-22?",
|
|
279
|
+
"expectedOutcomes": [
|
|
280
|
+
"Agent should call search_hotels tool with location=Manhattan",
|
|
281
|
+
"Agent should present hotel options with prices and ratings"
|
|
282
|
+
]
|
|
283
|
+
},
|
|
284
|
+
{
|
|
285
|
+
"name": "Multi-step trip planning",
|
|
286
|
+
"description": "User asks for both flights and hotels in one query",
|
|
287
|
+
"labels": ["category:Travel", "difficulty:Medium"],
|
|
288
|
+
"initialPrompt": "Plan a trip from Seattle to New York next weekend. I need both flights and a hotel.",
|
|
289
|
+
"expectedOutcomes": [
|
|
290
|
+
"Agent should call search_flights tool",
|
|
291
|
+
"Agent should call search_hotels tool",
|
|
292
|
+
"Agent should present a combined itinerary with flight and hotel options"
|
|
293
|
+
]
|
|
294
|
+
}
|
|
295
|
+
]
|
|
296
|
+
```
|
|
297
|
+
|
|
298
|
+
You can also create test cases directly in the UI via **Settings > Use Cases**.
|
|
299
|
+
|
|
300
|
+
#### Run via CLI
|
|
301
|
+
|
|
302
|
+
```
|
|
303
|
+
npx @opensearch-project/agent-health benchmark \
|
|
304
|
+
-f travel-benchmark.json \
|
|
305
|
+
-a my-agent \
|
|
306
|
+
-v
|
|
307
|
+
```
|
|
308
|
+
|
|
309
|
+
This will:
|
|
310
|
+
* Import the test cases
|
|
311
|
+
* Send each `initialPrompt` to your agent
|
|
312
|
+
* Have the LLM judge score agent responses against `expectedOutcomes`
|
|
313
|
+
|
|
314
|
+
#### Run via UI
|
|
315
|
+
|
|
316
|
+
Alternatively, in the Agent Health UI:
|
|
317
|
+
* Go to **Benchmarks**
|
|
318
|
+
* Click **Import JSON** and select `travel-benchmark.json`
|
|
319
|
+
* Click **Run** on your benchmark
|
|
320
|
+
* Select your agent endpoint and judge model
|
|
321
|
+
* Watch as each test case runs and gets evaluated in real time
|
|
322
|
+
|
|
323
|
+
<!-- TODO: Screenshot — Benchmark run in progress -->
|
|
324
|
+
|
|
325
|
+
#### Analyze results
|
|
326
|
+
|
|
327
|
+
Open the benchmark run results to see:
|
|
328
|
+
* **Overall pass rate** — what percentage of test cases passed
|
|
329
|
+
* **Per-test-case results** — each test case with its pass/fail status, the LLM judge's reasoning, and a score
|
|
330
|
+
* **Improvement strategies** — prioritized recommendations for what to fix first
|
|
331
|
+
|
|
332
|
+
For failed test cases, the LLM judge explains exactly _why_ it failed — for example, "the agent did not call search_hotels as expected" or "the response was missing price information."
|
|
333
|
+
|
|
334
|
+
#### Iterate and compare
|
|
335
|
+
|
|
336
|
+
Make improvements to your agent (update prompts, add tools, change models), then re-run the benchmark. Use the **Compare** view to see side-by-side how your changes affected:
|
|
337
|
+
* Pass rates across all test cases
|
|
338
|
+
* Individual test case outcomes that flipped from fail to pass (or vice versa)
|
|
339
|
+
* Latency and cost changes
|
|
340
|
+
|
|
341
|
+
<!-- TODO: Diagram — build-up: full eval loop: Agent Health ↔ Agent ↔ Judge, with results -->
|
|
342
|
+
|
|
343
|
+
---
|
|
344
|
+
|
|
345
|
+
## Part III: Add Tracing
|
|
346
|
+
|
|
347
|
+
Tracing gives you deep visibility into what your agent does internally — every LLM call, tool invocation, and reasoning step. Pick the scenario that matches your situation:
|
|
348
|
+
|
|
349
|
+
### Scenario A: You already have traces in OpenSearch
|
|
350
|
+
|
|
351
|
+
If your agent already sends OTel traces to an OpenSearch cluster, just point Agent Health at it.
|
|
352
|
+
|
|
353
|
+
**Configure the Observability Storage connection** (pick one):
|
|
354
|
+
|
|
355
|
+
* **Via Settings UI** — Go to **Settings** and fill in the **Observability Storage** section with your OpenSearch cluster URL and credentials. Choose "Basic Auth" for username/password or "AWS SigV4" for AWS-managed clusters.
|
|
356
|
+
* **Via environment variables** — Add the connection details to your `.env` file:
|
|
357
|
+
|
|
358
|
+
```bash
|
|
359
|
+
# Option A: Basic Auth (username/password)
|
|
360
|
+
OPENSEARCH_LOGS_ENDPOINT=https://your-opensearch-cluster:9200
|
|
361
|
+
OPENSEARCH_LOGS_USERNAME=admin
|
|
362
|
+
OPENSEARCH_LOGS_PASSWORD=admin
|
|
363
|
+
|
|
364
|
+
# Option B: AWS SigV4 (for AWS-managed OpenSearch or Serverless)
|
|
365
|
+
OPENSEARCH_LOGS_ENDPOINT=https://your-cluster.us-east-1.es.amazonaws.com
|
|
366
|
+
OPENSEARCH_LOGS_AUTH_TYPE=sigv4
|
|
367
|
+
OPENSEARCH_LOGS_AWS_REGION=us-east-1
|
|
368
|
+
OPENSEARCH_LOGS_AWS_SERVICE=es # 'es' for managed, 'aoss' for Serverless
|
|
369
|
+
# OPENSEARCH_LOGS_AWS_PROFILE=MyProfile # optional, uses default credential chain
|
|
370
|
+
```
|
|
371
|
+
|
|
372
|
+
Then enable traces for your agent in `agent-health.config.ts`:
|
|
373
|
+
|
|
374
|
+
```typescript
|
|
375
|
+
{
|
|
376
|
+
key: "my-agent",
|
|
377
|
+
name: "My Agent",
|
|
378
|
+
endpoint: "http://localhost:3000/agent",
|
|
379
|
+
connectorType: "rest",
|
|
380
|
+
useTraces: true, // Enable trace collection for this agent
|
|
381
|
+
models: ["claude-sonnet-4"],
|
|
382
|
+
}
|
|
383
|
+
```
|
|
384
|
+
|
|
385
|
+
Navigate to the **Traces** tab — you should see your agent's traces immediately.
|
|
386
|
+
|
|
387
|
+
Now re-run your benchmark from Step 3. Each evaluation run will also pull in the associated traces — giving you full visibility into what the agent did internally, alongside the judge's assessment.
|
|
388
|
+
|
|
389
|
+
<!-- TODO: Screenshot — Traces tab showing real agent traces -->
|
|
390
|
+
|
|
391
|
+
### Scenario B: Your traces go to another backend (Jaeger, Grafana, etc.)
|
|
392
|
+
|
|
393
|
+
If your agent already has OTel instrumentation but traces go to a non-OpenSearch backend, you need a local OpenSearch + OTel Collector stack to receive them. Two options:
|
|
394
|
+
|
|
395
|
+
**Option 1: Observability Stack shell script**
|
|
396
|
+
|
|
397
|
+
```
|
|
398
|
+
curl -fsSL https://raw.githubusercontent.com/opensearch-project/observability-stack/main/install.sh | bash
|
|
399
|
+
```
|
|
400
|
+
|
|
401
|
+
This launches:
|
|
402
|
+
* **OpenSearch** on port `9200` — stores traces
|
|
403
|
+
* **OTEL Collector** on port `4317` — receives OpenTelemetry traces
|
|
404
|
+
|
|
405
|
+
**Option 2: Docker Compose**
|
|
406
|
+
|
|
407
|
+
<!-- TODO: Add Docker Compose instructions — this may be consolidated into a single Docker Compose that ships with Agent Health itself, which would simplify both options into one -->
|
|
408
|
+
|
|
409
|
+
Both options give you the same result. Then:
|
|
410
|
+
1. Swap your agent's OTEL exporter endpoint to `http://localhost:4317`
|
|
411
|
+
2. Configure the Observability Storage connection in Agent Health (via Settings UI or `.env` file, as shown in Scenario A)
|
|
412
|
+
3. Enable `useTraces: true` for your agent in `agent-health.config.ts`
|
|
413
|
+
|
|
414
|
+
<!-- TODO: Test shell script with Docker and add detailed instructions -->
|
|
415
|
+
|
|
416
|
+
### Scenario C: Your agent has no OTel instrumentation yet
|
|
417
|
+
|
|
418
|
+
If your agent doesn't have OpenTelemetry instrumentation, you'll need to add it. Here's a walkthrough using a Python agent as an example.
|
|
419
|
+
|
|
420
|
+
<!-- TODO: Ship a sample agent with Agent Health that users can run to test the tracing flow end-to-end without needing their own instrumented agent -->
|
|
421
|
+
|
|
422
|
+
#### A sample agent (before instrumentation)
|
|
423
|
+
|
|
424
|
+
```python
|
|
425
|
+
# travel_agent.py
|
|
426
|
+
import openai
|
|
427
|
+
|
|
428
|
+
client = openai.OpenAI()
|
|
429
|
+
|
|
430
|
+
tools = [
|
|
431
|
+
{
|
|
432
|
+
"type": "function",
|
|
433
|
+
"function": {
|
|
434
|
+
"name": "search_flights",
|
|
435
|
+
"description": "Search for available flights",
|
|
436
|
+
"parameters": {
|
|
437
|
+
"type": "object",
|
|
438
|
+
"properties": {
|
|
439
|
+
"origin": {"type": "string"},
|
|
440
|
+
"destination": {"type": "string"},
|
|
441
|
+
"date": {"type": "string"}
|
|
442
|
+
},
|
|
443
|
+
"required": ["origin", "destination", "date"]
|
|
444
|
+
}
|
|
445
|
+
}
|
|
446
|
+
},
|
|
447
|
+
{
|
|
448
|
+
"type": "function",
|
|
449
|
+
"function": {
|
|
450
|
+
"name": "search_hotels",
|
|
451
|
+
"description": "Search for available hotels",
|
|
452
|
+
"parameters": {
|
|
453
|
+
"type": "object",
|
|
454
|
+
"properties": {
|
|
455
|
+
"location": {"type": "string"},
|
|
456
|
+
"check_in": {"type": "string"},
|
|
457
|
+
"check_out": {"type": "string"}
|
|
458
|
+
},
|
|
459
|
+
"required": ["location", "check_in", "check_out"]
|
|
460
|
+
}
|
|
461
|
+
}
|
|
462
|
+
}
|
|
463
|
+
]
|
|
464
|
+
|
|
465
|
+
def run_agent(user_query: str) -> str:
|
|
466
|
+
messages = [
|
|
467
|
+
{"role": "system", "content": "You are a helpful travel planning assistant."},
|
|
468
|
+
{"role": "user", "content": user_query}
|
|
469
|
+
]
|
|
470
|
+
response = client.chat.completions.create(
|
|
471
|
+
model="gpt-4o-mini",
|
|
472
|
+
messages=messages,
|
|
473
|
+
tools=tools,
|
|
474
|
+
)
|
|
475
|
+
return response.choices[0].message.content
|
|
476
|
+
|
|
477
|
+
if __name__ == "__main__":
|
|
478
|
+
result = run_agent("Find me flights from Seattle to New York next Friday")
|
|
479
|
+
print(result)
|
|
480
|
+
```
|
|
481
|
+
|
|
482
|
+
#### Add OpenTelemetry instrumentation
|
|
483
|
+
|
|
484
|
+
Install the required packages:
|
|
485
|
+
|
|
486
|
+
```
|
|
487
|
+
pip install opentelemetry-api opentelemetry-sdk opentelemetry-exporter-otlp \
|
|
488
|
+
opentelemetry-instrumentation-openai
|
|
489
|
+
```
|
|
490
|
+
|
|
491
|
+
Now add tracing to the agent:
|
|
492
|
+
|
|
493
|
+
```python
|
|
494
|
+
# travel_agent_traced.py
|
|
495
|
+
from opentelemetry import trace
|
|
496
|
+
from opentelemetry.sdk.trace import TracerProvider
|
|
497
|
+
from opentelemetry.sdk.trace.export import BatchSpanProcessor
|
|
498
|
+
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
|
|
499
|
+
from opentelemetry.sdk.resources import Resource
|
|
500
|
+
|
|
501
|
+
# Configure the tracer to send spans to the OTEL Collector
|
|
502
|
+
resource = Resource.create({"service.name": "travel-agent"})
|
|
503
|
+
provider = TracerProvider(resource=resource)
|
|
504
|
+
exporter = OTLPSpanExporter(endpoint="http://localhost:4317", insecure=True)
|
|
505
|
+
provider.add_span_processor(BatchSpanProcessor(exporter))
|
|
506
|
+
trace.set_tracer_provider(provider)
|
|
507
|
+
|
|
508
|
+
tracer = trace.get_tracer("travel-agent")
|
|
509
|
+
|
|
510
|
+
# Auto-instrument OpenAI calls (captures LLM spans automatically)
|
|
511
|
+
from opentelemetry.instrumentation.openai import OpenAIInstrumentor
|
|
512
|
+
OpenAIInstrumentor().instrument()
|
|
513
|
+
|
|
514
|
+
import openai
|
|
515
|
+
|
|
516
|
+
client = openai.OpenAI()
|
|
517
|
+
|
|
518
|
+
tools = [
|
|
519
|
+
# ... same tool definitions as above ...
|
|
520
|
+
]
|
|
521
|
+
|
|
522
|
+
def search_flights(origin, destination, date):
|
|
523
|
+
"""Simulated flight search."""
|
|
524
|
+
with tracer.start_as_current_span("tool.search_flights") as span:
|
|
525
|
+
span.set_attribute("tool.name", "search_flights")
|
|
526
|
+
span.set_attribute("tool.parameters.origin", origin)
|
|
527
|
+
span.set_attribute("tool.parameters.destination", destination)
|
|
528
|
+
results = [
|
|
529
|
+
{"flight": "AA123", "time": "8:00 AM", "price": "$350"},
|
|
530
|
+
{"flight": "UA456", "time": "2:00 PM", "price": "$280"},
|
|
531
|
+
]
|
|
532
|
+
span.set_attribute("tool.result_count", len(results))
|
|
533
|
+
return results
|
|
534
|
+
|
|
535
|
+
def search_hotels(location, check_in, check_out):
|
|
536
|
+
"""Simulated hotel search."""
|
|
537
|
+
with tracer.start_as_current_span("tool.search_hotels") as span:
|
|
538
|
+
span.set_attribute("tool.name", "search_hotels")
|
|
539
|
+
span.set_attribute("tool.parameters.location", location)
|
|
540
|
+
results = [
|
|
541
|
+
{"hotel": "Hilton Manhattan", "price": "$200/night", "rating": 4.5},
|
|
542
|
+
]
|
|
543
|
+
span.set_attribute("tool.result_count", len(results))
|
|
544
|
+
return results
|
|
545
|
+
|
|
546
|
+
def run_agent(user_query: str) -> str:
|
|
547
|
+
with tracer.start_as_current_span("agent.run") as span:
|
|
548
|
+
span.set_attribute("agent.name", "travel-agent")
|
|
549
|
+
span.set_attribute("agent.input", user_query)
|
|
550
|
+
|
|
551
|
+
messages = [
|
|
552
|
+
{"role": "system", "content": "You are a helpful travel planning assistant."},
|
|
553
|
+
{"role": "user", "content": user_query}
|
|
554
|
+
]
|
|
555
|
+
|
|
556
|
+
response = client.chat.completions.create(
|
|
557
|
+
model="gpt-4o-mini",
|
|
558
|
+
messages=messages,
|
|
559
|
+
tools=tools,
|
|
560
|
+
)
|
|
561
|
+
|
|
562
|
+
result = response.choices[0].message.content
|
|
563
|
+
span.set_attribute("agent.output", result or "")
|
|
564
|
+
return result
|
|
565
|
+
|
|
566
|
+
if __name__ == "__main__":
|
|
567
|
+
result = run_agent("Find me flights from Seattle to New York next Friday")
|
|
568
|
+
print(result)
|
|
569
|
+
# Flush spans before exit
|
|
570
|
+
provider.force_flush()
|
|
571
|
+
```
|
|
572
|
+
|
|
573
|
+
#### Run the agent and view traces
|
|
574
|
+
|
|
575
|
+
```
|
|
576
|
+
python travel_agent_traced.py
|
|
577
|
+
```
|
|
578
|
+
|
|
579
|
+
Now go back to **Traces** in Agent Health (`http://localhost:4001`). You should see a new trace appear showing:
|
|
580
|
+
* The top-level `agent.run` span with your query
|
|
581
|
+
* Nested LLM call spans (auto-instrumented by the OpenAI instrumentor)
|
|
582
|
+
* Tool call spans like `tool.search_flights`
|
|
583
|
+
|
|
584
|
+
Click into the trace to explore the timeline and flow views. You can now see exactly what your agent did, how long each step took, and what data flowed between components.
|
|
585
|
+
|
|
586
|
+
<!-- TODO: Screenshot — Trace detail with timeline and flow views -->
|
|
587
|
+
|
|
588
|
+
Once traces are flowing, re-run your benchmark from Step 3 to get evaluations with full trace data.
|
|
589
|
+
|
|
590
|
+
<!-- TODO: Diagram — complete picture: Agent Health ↔ Agent ↔ Judge + OTel traces flowing in -->
|
|
591
|
+
|
|
592
|
+
---
|
|
593
|
+
|
|
594
|
+
## Next steps
|
|
595
|
+
|
|
596
|
+
You now have a complete Agent Health workflow — from exploring sample data, to evaluating your own agent, to adding deep observability with OTel traces. From here, you can:
|
|
597
|
+
* Add more test cases to cover edge cases and failure modes
|
|
598
|
+
* Integrate benchmark runs into your CI/CD pipeline for automated regression testing
|
|
599
|
+
* Explore the trace views to debug specific agent failures in detail
|
|
600
|
+
* Use the **Compare** view to A/B test agent configurations over time
|
|
601
|
+
|
|
602
|
+
We'll cover each of these topics in upcoming posts. Stay tuned!
|
|
603
|
+
|
|
604
|
+
## Resources
|
|
605
|
+
* **GitHub Repository**: [opensearch-project/agent-health](https://github.com/opensearch-project/agent-health)
|
|
606
|
+
* **First blog post**: [OpenSearch Agent Health: Open-Source Observability and Evaluation for AI Agents](https://opensearch.org/blog/opensearch-agent-health/)
|
|
607
|
+
* **OpenTelemetry instrumentation guides**: [opentelemetry.io/docs/instrumentation](https://opentelemetry.io/docs/instrumentation/)
|
|
608
|
+
* **OpenSearch Observability Stack**: [opensearch-project/observability-stack](https://github.com/opensearch-project/observability-stack)
|