@opensearch-project/agent-health 0.3.0 → 0.5.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +77 -6
- package/cli/dist/index.js +10072 -4502
- package/deployment/cloudformation/agent-health-observability.yaml +762 -0
- package/dist/assets/index-CCQRDlO0.js +243 -0
- package/dist/assets/index-CNHQVbcj.css +1 -0
- package/dist/index.html +2 -2
- package/docs/ARCHITECTURE.md +450 -0
- package/docs/BACKEND_JOB_QUEUE.md +405 -0
- package/docs/CLAUDE_CODE_TELEMETRY.md +283 -0
- package/docs/CLI.md +431 -0
- package/docs/CODING_AGENT_ANALYTICS.md +298 -0
- package/docs/CONFIGURATION.md +388 -0
- package/docs/CONNECTORS.md +536 -0
- package/docs/INSTRUMENT_WITH_OTEL.md +390 -0
- package/docs/ML-COMMONS-SETUP.md +289 -0
- package/docs/NPX_PACKAGING.md +195 -0
- package/docs/PERFORMANCE-MONITORING.md +200 -0
- package/docs/PERFORMANCE.md +390 -0
- package/docs/PI_PROFILING.md +169 -0
- package/docs/PLAN-non-agui-agent-support.md +525 -0
- package/docs/SDK.md +577 -0
- package/docs/SKILLS.md +264 -0
- package/docs/blogs/2026-02-28-opensearch-agent-health.md +200 -0
- package/docs/blogs/getting-started-blog.md +608 -0
- package/docs/diagrams/Agent-health.excalidraw +5656 -0
- package/docs/diagrams/architecture.png +0 -0
- package/docs/plans/field-redesign.md +468 -0
- package/docs/rfcs/001-coding-agent-analytics.md +374 -0
- package/docs/rfcs/002-enterprise-leaderboard.md +267 -0
- package/docs/rfcs/003-remote-aggregation.md +146 -0
- package/docs/rfcs/004-test-sdk-v2.md +599 -0
- package/docs/skills/AGENT_HEALTH.md +598 -0
- package/docs/skills/AGENT_PROFILE.md +191 -0
- package/docs/skills/add-connector/SKILL.md +68 -0
- package/docs/skills/agent-health-profile/SKILL.md +40 -0
- package/docs/skills/config-auth/SKILL.md +194 -0
- package/docs/skills/config-auth/evals/evals.json +35 -0
- package/docs/skills/create-pr/SKILL.md +73 -0
- package/docs/skills/instrument-otel/SKILL.md +84 -0
- package/docs/skills/write-test/SKILL.md +124 -0
- package/docs/ui prd.md +376 -0
- package/examples/README.md +53 -0
- package/examples/config/agent-health.config.example.ts +155 -0
- package/examples/connectors/echo-connector.ts +131 -0
- package/examples/eval-files/demo.eval.js +128 -0
- package/examples/eval-files/sdk-hooks-demo.eval.js +99 -0
- package/examples/pi-profiling/README.md +77 -0
- package/examples/pi-profiling/agent-health-profile.ts +417 -0
- package/lib/dist/lib/agentUtils.d.ts +29 -0
- package/lib/dist/lib/agentUtils.d.ts.map +1 -0
- package/lib/dist/lib/agentUtils.js +43 -0
- package/lib/dist/lib/agentUtils.js.map +1 -0
- package/lib/dist/lib/benchmarkExport.d.ts +14 -0
- package/lib/dist/lib/benchmarkExport.d.ts.map +1 -0
- package/lib/dist/lib/benchmarkExport.js +41 -0
- package/lib/dist/lib/benchmarkExport.js.map +1 -0
- package/lib/dist/lib/benchmarkVersionUtils.d.ts +37 -0
- package/lib/dist/lib/benchmarkVersionUtils.d.ts.map +1 -0
- package/lib/dist/lib/benchmarkVersionUtils.js +68 -0
- package/lib/dist/lib/benchmarkVersionUtils.js.map +1 -0
- package/lib/dist/lib/config/defineConfig.d.ts +27 -0
- package/lib/dist/lib/config/defineConfig.d.ts.map +1 -0
- package/lib/dist/lib/config/defineConfig.js +28 -0
- package/lib/dist/lib/config/defineConfig.js.map +1 -0
- package/lib/dist/lib/config/index.d.ts +9 -0
- package/lib/dist/lib/config/index.d.ts.map +1 -0
- package/lib/dist/lib/config/index.js +8 -0
- package/lib/dist/lib/config/index.js.map +1 -0
- package/lib/dist/lib/config/loader.d.ts +39 -0
- package/lib/dist/lib/config/loader.d.ts.map +1 -0
- package/lib/dist/lib/config/loader.js +258 -0
- package/lib/dist/lib/config/loader.js.map +1 -0
- package/lib/dist/lib/config/statePaths.d.ts +61 -0
- package/lib/dist/lib/config/statePaths.d.ts.map +1 -0
- package/lib/dist/lib/config/statePaths.js +188 -0
- package/lib/dist/lib/config/statePaths.js.map +1 -0
- package/lib/dist/lib/config/types.d.ts +231 -0
- package/lib/dist/lib/config/types.d.ts.map +1 -0
- package/lib/dist/lib/config/types.js +6 -0
- package/lib/dist/lib/config/types.js.map +1 -0
- package/lib/dist/lib/config.d.ts +39 -0
- package/lib/dist/lib/config.d.ts.map +1 -0
- package/lib/dist/lib/config.js +118 -0
- package/lib/dist/lib/config.js.map +1 -0
- package/lib/dist/lib/constants.d.ts +70 -0
- package/lib/dist/lib/constants.d.ts.map +1 -0
- package/lib/dist/lib/constants.js +365 -0
- package/lib/dist/lib/constants.js.map +1 -0
- package/lib/dist/lib/contextUtilization.d.ts +23 -0
- package/lib/dist/lib/contextUtilization.d.ts.map +1 -0
- package/lib/dist/lib/contextUtilization.js +72 -0
- package/lib/dist/lib/contextUtilization.js.map +1 -0
- package/lib/dist/lib/dashboardMetrics.d.ts +87 -0
- package/lib/dist/lib/dashboardMetrics.d.ts.map +1 -0
- package/lib/dist/lib/dashboardMetrics.js +242 -0
- package/lib/dist/lib/dashboardMetrics.js.map +1 -0
- package/lib/dist/lib/dataSourceConfig.d.ts +108 -0
- package/lib/dist/lib/dataSourceConfig.d.ts.map +1 -0
- package/lib/dist/lib/dataSourceConfig.js +166 -0
- package/lib/dist/lib/dataSourceConfig.js.map +1 -0
- package/lib/dist/lib/debug.d.ts +26 -0
- package/lib/dist/lib/debug.d.ts.map +1 -0
- package/lib/dist/lib/debug.js +132 -0
- package/lib/dist/lib/debug.js.map +1 -0
- package/lib/dist/lib/diagnostics.d.ts +28 -0
- package/lib/dist/lib/diagnostics.d.ts.map +1 -0
- package/lib/dist/lib/diagnostics.js +65 -0
- package/lib/dist/lib/diagnostics.js.map +1 -0
- package/lib/dist/lib/envCompat.d.ts +27 -0
- package/lib/dist/lib/envCompat.d.ts.map +1 -0
- package/lib/dist/lib/envCompat.js +73 -0
- package/lib/dist/lib/envCompat.js.map +1 -0
- package/lib/dist/lib/findPackageRoot.d.ts +7 -0
- package/lib/dist/lib/findPackageRoot.d.ts.map +1 -0
- package/lib/dist/lib/findPackageRoot.js +57 -0
- package/lib/dist/lib/findPackageRoot.js.map +1 -0
- package/lib/dist/lib/hooks.d.ts +36 -0
- package/lib/dist/lib/hooks.d.ts.map +1 -0
- package/lib/dist/lib/hooks.js +112 -0
- package/lib/dist/lib/hooks.js.map +1 -0
- package/lib/dist/lib/index.d.ts +47 -0
- package/lib/dist/lib/index.d.ts.map +1 -0
- package/lib/dist/lib/index.js +62 -0
- package/lib/dist/lib/index.js.map +1 -0
- package/lib/dist/lib/labels.d.ts +90 -0
- package/lib/dist/lib/labels.d.ts.map +1 -0
- package/lib/dist/lib/labels.js +158 -0
- package/lib/dist/lib/labels.js.map +1 -0
- package/lib/dist/lib/markdown.d.ts +16 -0
- package/lib/dist/lib/markdown.d.ts.map +1 -0
- package/lib/dist/lib/markdown.js +42 -0
- package/lib/dist/lib/markdown.js.map +1 -0
- package/lib/dist/lib/matchers/expect.d.ts +3 -0
- package/lib/dist/lib/matchers/expect.d.ts.map +1 -0
- package/lib/dist/lib/matchers/expect.js +225 -0
- package/lib/dist/lib/matchers/expect.js.map +1 -0
- package/lib/dist/lib/matchers/index.d.ts +8 -0
- package/lib/dist/lib/matchers/index.d.ts.map +1 -0
- package/lib/dist/lib/matchers/index.js +9 -0
- package/lib/dist/lib/matchers/index.js.map +1 -0
- package/lib/dist/lib/matchers/judgeAccessor.d.ts +113 -0
- package/lib/dist/lib/matchers/judgeAccessor.d.ts.map +1 -0
- package/lib/dist/lib/matchers/judgeAccessor.js +183 -0
- package/lib/dist/lib/matchers/judgeAccessor.js.map +1 -0
- package/lib/dist/lib/matchers/session.d.ts +39 -0
- package/lib/dist/lib/matchers/session.d.ts.map +1 -0
- package/lib/dist/lib/matchers/session.js +116 -0
- package/lib/dist/lib/matchers/session.js.map +1 -0
- package/lib/dist/lib/matchers/traces.d.ts +55 -0
- package/lib/dist/lib/matchers/traces.d.ts.map +1 -0
- package/lib/dist/lib/matchers/traces.js +116 -0
- package/lib/dist/lib/matchers/traces.js.map +1 -0
- package/lib/dist/lib/matchers/types.d.ts +75 -0
- package/lib/dist/lib/matchers/types.d.ts.map +1 -0
- package/lib/dist/lib/matchers/types.js +6 -0
- package/lib/dist/lib/matchers/types.js.map +1 -0
- package/lib/dist/lib/packagePaths.d.ts +29 -0
- package/lib/dist/lib/packagePaths.d.ts.map +1 -0
- package/lib/dist/lib/packagePaths.js +63 -0
- package/lib/dist/lib/packagePaths.js.map +1 -0
- package/lib/dist/lib/performance.d.ts +51 -0
- package/lib/dist/lib/performance.d.ts.map +1 -0
- package/lib/dist/lib/performance.js +159 -0
- package/lib/dist/lib/performance.js.map +1 -0
- package/lib/dist/lib/portConfig.d.ts +29 -0
- package/lib/dist/lib/portConfig.d.ts.map +1 -0
- package/lib/dist/lib/portConfig.js +64 -0
- package/lib/dist/lib/portConfig.js.map +1 -0
- package/lib/dist/lib/preferences.d.ts +63 -0
- package/lib/dist/lib/preferences.d.ts.map +1 -0
- package/lib/dist/lib/preferences.js +117 -0
- package/lib/dist/lib/preferences.js.map +1 -0
- package/lib/dist/lib/resolveAgentModel.d.ts +22 -0
- package/lib/dist/lib/resolveAgentModel.d.ts.map +1 -0
- package/lib/dist/lib/resolveAgentModel.js +37 -0
- package/lib/dist/lib/resolveAgentModel.js.map +1 -0
- package/lib/dist/lib/runStats.d.ts +92 -0
- package/lib/dist/lib/runStats.d.ts.map +1 -0
- package/lib/dist/lib/runStats.js +160 -0
- package/lib/dist/lib/runStats.js.map +1 -0
- package/lib/dist/lib/telemetry/constants.d.ts +60 -0
- package/lib/dist/lib/telemetry/constants.d.ts.map +1 -0
- package/lib/dist/lib/telemetry/constants.js +87 -0
- package/lib/dist/lib/telemetry/constants.js.map +1 -0
- package/lib/dist/lib/telemetry/evalSpans.d.ts +61 -0
- package/lib/dist/lib/telemetry/evalSpans.d.ts.map +1 -0
- package/lib/dist/lib/telemetry/evalSpans.js +254 -0
- package/lib/dist/lib/telemetry/evalSpans.js.map +1 -0
- package/lib/dist/lib/telemetry/index.d.ts +11 -0
- package/lib/dist/lib/telemetry/index.d.ts.map +1 -0
- package/lib/dist/lib/telemetry/index.js +15 -0
- package/lib/dist/lib/telemetry/index.js.map +1 -0
- package/lib/dist/lib/telemetry/opensearchExporter.d.ts +43 -0
- package/lib/dist/lib/telemetry/opensearchExporter.d.ts.map +1 -0
- package/lib/dist/lib/telemetry/opensearchExporter.js +217 -0
- package/lib/dist/lib/telemetry/opensearchExporter.js.map +1 -0
- package/lib/dist/lib/telemetry/provider.d.ts +55 -0
- package/lib/dist/lib/telemetry/provider.d.ts.map +1 -0
- package/lib/dist/lib/telemetry/provider.js +140 -0
- package/lib/dist/lib/telemetry/provider.js.map +1 -0
- package/lib/dist/lib/testCaseLabels.d.ts +34 -0
- package/lib/dist/lib/testCaseLabels.d.ts.map +1 -0
- package/lib/dist/lib/testCaseLabels.js +88 -0
- package/lib/dist/lib/testCaseLabels.js.map +1 -0
- package/lib/dist/lib/testCaseValidation.d.ts +140 -0
- package/lib/dist/lib/testCaseValidation.d.ts.map +1 -0
- package/lib/dist/lib/testCaseValidation.js +162 -0
- package/lib/dist/lib/testCaseValidation.js.map +1 -0
- package/lib/dist/lib/testCases/agentFixture.d.ts +80 -0
- package/lib/dist/lib/testCases/agentFixture.d.ts.map +1 -0
- package/lib/dist/lib/testCases/agentFixture.js +43 -0
- package/lib/dist/lib/testCases/agentFixture.js.map +1 -0
- package/lib/dist/lib/testCases/authoringSurface.d.ts +10 -0
- package/lib/dist/lib/testCases/authoringSurface.d.ts.map +1 -0
- package/lib/dist/lib/testCases/authoringSurface.js +54 -0
- package/lib/dist/lib/testCases/authoringSurface.js.map +1 -0
- package/lib/dist/lib/testCases/codemod.d.ts +13 -0
- package/lib/dist/lib/testCases/codemod.d.ts.map +1 -0
- package/lib/dist/lib/testCases/codemod.js +169 -0
- package/lib/dist/lib/testCases/codemod.js.map +1 -0
- package/lib/dist/lib/testCases/define.d.ts +114 -0
- package/lib/dist/lib/testCases/define.d.ts.map +1 -0
- package/lib/dist/lib/testCases/define.js +253 -0
- package/lib/dist/lib/testCases/define.js.map +1 -0
- package/lib/dist/lib/testCases/evaluators.d.ts +80 -0
- package/lib/dist/lib/testCases/evaluators.d.ts.map +1 -0
- package/lib/dist/lib/testCases/evaluators.js +105 -0
- package/lib/dist/lib/testCases/evaluators.js.map +1 -0
- package/lib/dist/lib/testCases/index.d.ts +14 -0
- package/lib/dist/lib/testCases/index.d.ts.map +1 -0
- package/lib/dist/lib/testCases/index.js +12 -0
- package/lib/dist/lib/testCases/index.js.map +1 -0
- package/lib/dist/lib/testCases/judge.d.ts +165 -0
- package/lib/dist/lib/testCases/judge.d.ts.map +1 -0
- package/lib/dist/lib/testCases/judge.js +359 -0
- package/lib/dist/lib/testCases/judge.js.map +1 -0
- package/lib/dist/lib/testCases/loader.d.ts +26 -0
- package/lib/dist/lib/testCases/loader.d.ts.map +1 -0
- package/lib/dist/lib/testCases/loader.js +149 -0
- package/lib/dist/lib/testCases/loader.js.map +1 -0
- package/lib/dist/lib/testCases/types.d.ts +242 -0
- package/lib/dist/lib/testCases/types.d.ts.map +1 -0
- package/lib/dist/lib/testCases/types.js +6 -0
- package/lib/dist/lib/testCases/types.js.map +1 -0
- package/lib/dist/lib/theme.d.ts +6 -0
- package/lib/dist/lib/theme.d.ts.map +1 -0
- package/lib/dist/lib/theme.js +36 -0
- package/lib/dist/lib/theme.js.map +1 -0
- package/lib/dist/lib/uiTelemetry.d.ts +7 -0
- package/lib/dist/lib/uiTelemetry.d.ts.map +1 -0
- package/lib/dist/lib/uiTelemetry.js +25 -0
- package/lib/dist/lib/uiTelemetry.js.map +1 -0
- package/lib/dist/lib/utils.d.ts +96 -0
- package/lib/dist/lib/utils.d.ts.map +1 -0
- package/lib/dist/lib/utils.js +232 -0
- package/lib/dist/lib/utils.js.map +1 -0
- package/lib/dist/lib/workflow/consolidate.d.ts +12 -0
- package/lib/dist/lib/workflow/consolidate.d.ts.map +1 -0
- package/lib/dist/lib/workflow/consolidate.js +33 -0
- package/lib/dist/lib/workflow/consolidate.js.map +1 -0
- package/lib/dist/lib/workflow/index.d.ts +13 -0
- package/lib/dist/lib/workflow/index.d.ts.map +1 -0
- package/lib/dist/lib/workflow/index.js +12 -0
- package/lib/dist/lib/workflow/index.js.map +1 -0
- package/lib/dist/lib/workflow/ledger.d.ts +30 -0
- package/lib/dist/lib/workflow/ledger.d.ts.map +1 -0
- package/lib/dist/lib/workflow/ledger.js +41 -0
- package/lib/dist/lib/workflow/ledger.js.map +1 -0
- package/lib/dist/lib/workflow/pool.d.ts +13 -0
- package/lib/dist/lib/workflow/pool.d.ts.map +1 -0
- package/lib/dist/lib/workflow/pool.js +44 -0
- package/lib/dist/lib/workflow/pool.js.map +1 -0
- package/lib/dist/lib/workflow/source.d.ts +22 -0
- package/lib/dist/lib/workflow/source.d.ts.map +1 -0
- package/lib/dist/lib/workflow/source.js +29 -0
- package/lib/dist/lib/workflow/source.js.map +1 -0
- package/lib/dist/lib/workflow/stepB.d.ts +71 -0
- package/lib/dist/lib/workflow/stepB.d.ts.map +1 -0
- package/lib/dist/lib/workflow/stepB.js +99 -0
- package/lib/dist/lib/workflow/stepB.js.map +1 -0
- package/lib/dist/lib/workflow/types.d.ts +86 -0
- package/lib/dist/lib/workflow/types.d.ts.map +1 -0
- package/lib/dist/lib/workflow/types.js +6 -0
- package/lib/dist/lib/workflow/types.js.map +1 -0
- package/lib/dist/lib/workflow/workflow.d.ts +119 -0
- package/lib/dist/lib/workflow/workflow.d.ts.map +1 -0
- package/lib/dist/lib/workflow/workflow.js +195 -0
- package/lib/dist/lib/workflow/workflow.js.map +1 -0
- package/lib/dist/services/agent/aguiConverter.d.ts +50 -0
- package/lib/dist/services/agent/aguiConverter.d.ts.map +1 -0
- package/lib/dist/services/agent/aguiConverter.js +449 -0
- package/lib/dist/services/agent/aguiConverter.js.map +1 -0
- package/lib/dist/services/agent/index.d.ts +10 -0
- package/lib/dist/services/agent/index.d.ts.map +1 -0
- package/lib/dist/services/agent/index.js +12 -0
- package/lib/dist/services/agent/index.js.map +1 -0
- package/lib/dist/services/agent/payloadBuilder.d.ts +33 -0
- package/lib/dist/services/agent/payloadBuilder.d.ts.map +1 -0
- package/lib/dist/services/agent/payloadBuilder.js +75 -0
- package/lib/dist/services/agent/payloadBuilder.js.map +1 -0
- package/lib/dist/services/agent/sseStream.d.ts +43 -0
- package/lib/dist/services/agent/sseStream.d.ts.map +1 -0
- package/lib/dist/services/agent/sseStream.js +223 -0
- package/lib/dist/services/agent/sseStream.js.map +1 -0
- package/lib/dist/services/connectors/agui/AGUIStreamingConnector.d.ts +44 -0
- package/lib/dist/services/connectors/agui/AGUIStreamingConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/agui/AGUIStreamingConnector.js +95 -0
- package/lib/dist/services/connectors/agui/AGUIStreamingConnector.js.map +1 -0
- package/lib/dist/services/connectors/base/BaseConnector.d.ts +81 -0
- package/lib/dist/services/connectors/base/BaseConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/base/BaseConnector.js +170 -0
- package/lib/dist/services/connectors/base/BaseConnector.js.map +1 -0
- package/lib/dist/services/connectors/claude-code/ClaudeCodeConnector.d.ts +116 -0
- package/lib/dist/services/connectors/claude-code/ClaudeCodeConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/claude-code/ClaudeCodeConnector.js +403 -0
- package/lib/dist/services/connectors/claude-code/ClaudeCodeConnector.js.map +1 -0
- package/lib/dist/services/connectors/index.d.ts +13 -0
- package/lib/dist/services/connectors/index.d.ts.map +1 -0
- package/lib/dist/services/connectors/index.js +32 -0
- package/lib/dist/services/connectors/index.js.map +1 -0
- package/lib/dist/services/connectors/kiro/KiroConnector.d.ts +48 -0
- package/lib/dist/services/connectors/kiro/KiroConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/kiro/KiroConnector.js +158 -0
- package/lib/dist/services/connectors/kiro/KiroConnector.js.map +1 -0
- package/lib/dist/services/connectors/langgraph/LangGraphConnector.d.ts +36 -0
- package/lib/dist/services/connectors/langgraph/LangGraphConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/langgraph/LangGraphConnector.js +175 -0
- package/lib/dist/services/connectors/langgraph/LangGraphConnector.js.map +1 -0
- package/lib/dist/services/connectors/mock/MockConnector.d.ts +37 -0
- package/lib/dist/services/connectors/mock/MockConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/mock/MockConnector.js +120 -0
- package/lib/dist/services/connectors/mock/MockConnector.js.map +1 -0
- package/lib/dist/services/connectors/openai-compatible/OpenAICompatibleConnector.d.ts +42 -0
- package/lib/dist/services/connectors/openai-compatible/OpenAICompatibleConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/openai-compatible/OpenAICompatibleConnector.js +133 -0
- package/lib/dist/services/connectors/openai-compatible/OpenAICompatibleConnector.js.map +1 -0
- package/lib/dist/services/connectors/pi/PiConnector.d.ts +87 -0
- package/lib/dist/services/connectors/pi/PiConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/pi/PiConnector.js +274 -0
- package/lib/dist/services/connectors/pi/PiConnector.js.map +1 -0
- package/lib/dist/services/connectors/registry.d.ts +57 -0
- package/lib/dist/services/connectors/registry.d.ts.map +1 -0
- package/lib/dist/services/connectors/registry.js +106 -0
- package/lib/dist/services/connectors/registry.js.map +1 -0
- package/lib/dist/services/connectors/rest/RESTConnector.d.ts +38 -0
- package/lib/dist/services/connectors/rest/RESTConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/rest/RESTConnector.js +117 -0
- package/lib/dist/services/connectors/rest/RESTConnector.js.map +1 -0
- package/lib/dist/services/connectors/server.d.ts +13 -0
- package/lib/dist/services/connectors/server.d.ts.map +1 -0
- package/lib/dist/services/connectors/server.js +34 -0
- package/lib/dist/services/connectors/server.js.map +1 -0
- package/lib/dist/services/connectors/strands/StrandsConnector.d.ts +48 -0
- package/lib/dist/services/connectors/strands/StrandsConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/strands/StrandsConnector.js +221 -0
- package/lib/dist/services/connectors/strands/StrandsConnector.js.map +1 -0
- package/lib/dist/services/connectors/subprocess/SubprocessConnector.d.ts +88 -0
- package/lib/dist/services/connectors/subprocess/SubprocessConnector.d.ts.map +1 -0
- package/lib/dist/services/connectors/subprocess/SubprocessConnector.js +418 -0
- package/lib/dist/services/connectors/subprocess/SubprocessConnector.js.map +1 -0
- package/lib/dist/services/connectors/types.d.ts +213 -0
- package/lib/dist/services/connectors/types.d.ts.map +1 -0
- package/lib/dist/services/connectors/types.js +6 -0
- package/lib/dist/services/connectors/types.js.map +1 -0
- package/lib/dist/services/evaluation/bedrockJudge.d.ts +64 -0
- package/lib/dist/services/evaluation/bedrockJudge.d.ts.map +1 -0
- package/lib/dist/services/evaluation/bedrockJudge.js +167 -0
- package/lib/dist/services/evaluation/bedrockJudge.js.map +1 -0
- package/lib/dist/services/evaluation/evaluatorError.d.ts +56 -0
- package/lib/dist/services/evaluation/evaluatorError.d.ts.map +1 -0
- package/lib/dist/services/evaluation/evaluatorError.js +56 -0
- package/lib/dist/services/evaluation/evaluatorError.js.map +1 -0
- package/lib/dist/services/evaluation/index.d.ts +106 -0
- package/lib/dist/services/evaluation/index.d.ts.map +1 -0
- package/lib/dist/services/evaluation/index.js +684 -0
- package/lib/dist/services/evaluation/index.js.map +1 -0
- package/lib/dist/services/evaluation/mockTrajectory.d.ts +3 -0
- package/lib/dist/services/evaluation/mockTrajectory.d.ts.map +1 -0
- package/lib/dist/services/evaluation/mockTrajectory.js +72 -0
- package/lib/dist/services/evaluation/mockTrajectory.js.map +1 -0
- package/lib/dist/services/opensearch/client.d.ts +26 -0
- package/lib/dist/services/opensearch/client.d.ts.map +1 -0
- package/lib/dist/services/opensearch/client.js +131 -0
- package/lib/dist/services/opensearch/client.js.map +1 -0
- package/lib/dist/services/opensearch/index.d.ts +16 -0
- package/lib/dist/services/opensearch/index.d.ts.map +1 -0
- package/lib/dist/services/opensearch/index.js +25 -0
- package/lib/dist/services/opensearch/index.js.map +1 -0
- package/lib/dist/services/storage/asyncBenchmarkStorage.d.ts +123 -0
- package/lib/dist/services/storage/asyncBenchmarkStorage.d.ts.map +1 -0
- package/lib/dist/services/storage/asyncBenchmarkStorage.js +429 -0
- package/lib/dist/services/storage/asyncBenchmarkStorage.js.map +1 -0
- package/lib/dist/services/storage/asyncRunStorage.d.ts +127 -0
- package/lib/dist/services/storage/asyncRunStorage.d.ts.map +1 -0
- package/lib/dist/services/storage/asyncRunStorage.js +448 -0
- package/lib/dist/services/storage/asyncRunStorage.js.map +1 -0
- package/lib/dist/services/storage/asyncTestCaseStorage.d.ts +156 -0
- package/lib/dist/services/storage/asyncTestCaseStorage.d.ts.map +1 -0
- package/lib/dist/services/storage/asyncTestCaseStorage.js +285 -0
- package/lib/dist/services/storage/asyncTestCaseStorage.js.map +1 -0
- package/lib/dist/services/storage/index.d.ts +17 -0
- package/lib/dist/services/storage/index.d.ts.map +1 -0
- package/lib/dist/services/storage/index.js +20 -0
- package/lib/dist/services/storage/index.js.map +1 -0
- package/lib/dist/services/storage/migration.d.ts +54 -0
- package/lib/dist/services/storage/migration.d.ts.map +1 -0
- package/lib/dist/services/storage/migration.js +296 -0
- package/lib/dist/services/storage/migration.js.map +1 -0
- package/lib/dist/services/storage/opensearchClient.d.ts +924 -0
- package/lib/dist/services/storage/opensearchClient.d.ts.map +1 -0
- package/lib/dist/services/storage/opensearchClient.js +435 -0
- package/lib/dist/services/storage/opensearchClient.js.map +1 -0
- package/lib/dist/services/traces/browserRecovery.d.ts +26 -0
- package/lib/dist/services/traces/browserRecovery.d.ts.map +1 -0
- package/lib/dist/services/traces/browserRecovery.js +81 -0
- package/lib/dist/services/traces/browserRecovery.js.map +1 -0
- package/lib/dist/services/traces/categoryStyles.d.ts +21 -0
- package/lib/dist/services/traces/categoryStyles.d.ts.map +1 -0
- package/lib/dist/services/traces/categoryStyles.js +56 -0
- package/lib/dist/services/traces/categoryStyles.js.map +1 -0
- package/lib/dist/services/traces/executionOrderTransform.d.ts +35 -0
- package/lib/dist/services/traces/executionOrderTransform.d.ts.map +1 -0
- package/lib/dist/services/traces/executionOrderTransform.js +313 -0
- package/lib/dist/services/traces/executionOrderTransform.js.map +1 -0
- package/lib/dist/services/traces/fetchSpansForRun.d.ts +86 -0
- package/lib/dist/services/traces/fetchSpansForRun.d.ts.map +1 -0
- package/lib/dist/services/traces/fetchSpansForRun.js +69 -0
- package/lib/dist/services/traces/fetchSpansForRun.js.map +1 -0
- package/lib/dist/services/traces/flowTransform.d.ts +24 -0
- package/lib/dist/services/traces/flowTransform.d.ts.map +1 -0
- package/lib/dist/services/traces/flowTransform.js +228 -0
- package/lib/dist/services/traces/flowTransform.js.map +1 -0
- package/lib/dist/services/traces/index.d.ts +121 -0
- package/lib/dist/services/traces/index.d.ts.map +1 -0
- package/lib/dist/services/traces/index.js +255 -0
- package/lib/dist/services/traces/index.js.map +1 -0
- package/lib/dist/services/traces/intentTransform.d.ts +20 -0
- package/lib/dist/services/traces/intentTransform.d.ts.map +1 -0
- package/lib/dist/services/traces/intentTransform.js +131 -0
- package/lib/dist/services/traces/intentTransform.js.map +1 -0
- package/lib/dist/services/traces/judgeAgentsHints.d.ts +63 -0
- package/lib/dist/services/traces/judgeAgentsHints.d.ts.map +1 -0
- package/lib/dist/services/traces/judgeAgentsHints.js +89 -0
- package/lib/dist/services/traces/judgeAgentsHints.js.map +1 -0
- package/lib/dist/services/traces/messageExtraction.d.ts +15 -0
- package/lib/dist/services/traces/messageExtraction.d.ts.map +1 -0
- package/lib/dist/services/traces/messageExtraction.js +251 -0
- package/lib/dist/services/traces/messageExtraction.js.map +1 -0
- package/lib/dist/services/traces/spanCategorization.d.ts +63 -0
- package/lib/dist/services/traces/spanCategorization.d.ts.map +1 -0
- package/lib/dist/services/traces/spanCategorization.js +276 -0
- package/lib/dist/services/traces/spanCategorization.js.map +1 -0
- package/lib/dist/services/traces/spanPreprocessing.d.ts +37 -0
- package/lib/dist/services/traces/spanPreprocessing.d.ts.map +1 -0
- package/lib/dist/services/traces/spanPreprocessing.js +102 -0
- package/lib/dist/services/traces/spanPreprocessing.js.map +1 -0
- package/lib/dist/services/traces/spansToTrajectory.d.ts +36 -0
- package/lib/dist/services/traces/spansToTrajectory.d.ts.map +1 -0
- package/lib/dist/services/traces/spansToTrajectory.js +387 -0
- package/lib/dist/services/traces/spansToTrajectory.js.map +1 -0
- package/lib/dist/services/traces/toolSimilarity.d.ts +35 -0
- package/lib/dist/services/traces/toolSimilarity.d.ts.map +1 -0
- package/lib/dist/services/traces/toolSimilarity.js +203 -0
- package/lib/dist/services/traces/toolSimilarity.js.map +1 -0
- package/lib/dist/services/traces/traceComparison.d.ts +31 -0
- package/lib/dist/services/traces/traceComparison.d.ts.map +1 -0
- package/lib/dist/services/traces/traceComparison.js +318 -0
- package/lib/dist/services/traces/traceComparison.js.map +1 -0
- package/lib/dist/services/traces/traceGrouping.d.ts +19 -0
- package/lib/dist/services/traces/traceGrouping.d.ts.map +1 -0
- package/lib/dist/services/traces/traceGrouping.js +107 -0
- package/lib/dist/services/traces/traceGrouping.js.map +1 -0
- package/lib/dist/services/traces/tracePoller.d.ts +84 -0
- package/lib/dist/services/traces/tracePoller.d.ts.map +1 -0
- package/lib/dist/services/traces/tracePoller.js +309 -0
- package/lib/dist/services/traces/tracePoller.js.map +1 -0
- package/lib/dist/services/traces/traceStats.d.ts +45 -0
- package/lib/dist/services/traces/traceStats.d.ts.map +1 -0
- package/lib/dist/services/traces/traceStats.js +114 -0
- package/lib/dist/services/traces/traceStats.js.map +1 -0
- package/lib/dist/services/traces/traceSummary.d.ts +47 -0
- package/lib/dist/services/traces/traceSummary.d.ts.map +1 -0
- package/lib/dist/services/traces/traceSummary.js +68 -0
- package/lib/dist/services/traces/traceSummary.js.map +1 -0
- package/lib/dist/services/traces/utils.d.ts +33 -0
- package/lib/dist/services/traces/utils.d.ts.map +1 -0
- package/lib/dist/services/traces/utils.js +114 -0
- package/lib/dist/services/traces/utils.js.map +1 -0
- package/lib/dist/types/agui.d.ts +13 -0
- package/lib/dist/types/agui.d.ts.map +1 -0
- package/lib/dist/types/agui.js +16 -0
- package/lib/dist/types/agui.js.map +1 -0
- package/lib/dist/types/index.d.ts +1175 -0
- package/lib/dist/types/index.d.ts.map +1 -0
- package/lib/dist/types/index.js +12 -0
- package/lib/dist/types/index.js.map +1 -0
- package/lib/dist/types/skills.d.ts +146 -0
- package/lib/dist/types/skills.d.ts.map +1 -0
- package/lib/dist/types/skills.js +6 -0
- package/lib/dist/types/skills.js.map +1 -0
- package/observio-sample-agent/pi-package/README.md +112 -0
- package/observio-sample-agent/pi-package/extensions/agent-health.ts +373 -0
- package/observio-sample-agent/pi-package/package.json +17 -0
- package/observio-sample-agent/pi-package/prompts/agent-health.md +37 -0
- package/observio-sample-agent/pi-package/skills/create-pr/SKILL.md +88 -0
- package/observio-sample-agent/pi-package/skills/fix-bug/SKILL.md +71 -0
- package/observio-sample-agent/pi-package/skills/implement-feature/SKILL.md +156 -0
- package/observio-sample-agent/pi-package/skills/instrument-otel/SKILL.md +208 -0
- package/observio-sample-agent/pi-package/skills/setup-collector/SKILL.md +146 -0
- package/observio-sample-agent/pi-package/skills/write-test/SKILL.md +115 -0
- package/package.json +64 -13
- package/server/dist/app.js +32651 -17637
- package/server/dist/index.js +29875 -14638
- package/tsconfig.lib.json +71 -0
- package/dist/assets/index-EvPLSTAS.js +0 -267
- package/dist/assets/index-RXasQKUs.css +0 -1
- package/lib/dist/config/index.js +0 -404
- package/lib/dist/index.js +0 -1665
|
@@ -0,0 +1,84 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: instrument-otel
|
|
3
|
+
description: Use when instrumenting an application with OpenTelemetry for Agent Health. Provides span structure, required attributes, and config setup for routing traces.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
## Task: Instrument with OpenTelemetry for Agent Health
|
|
7
|
+
|
|
8
|
+
Add OpenTelemetry instrumentation following GenAI semantic conventions so traces are compatible with Agent Health.
|
|
9
|
+
|
|
10
|
+
### Steps
|
|
11
|
+
|
|
12
|
+
1. **Install OTel dependencies** for the project's language (Python: `opentelemetry-api`, `opentelemetry-sdk`, `opentelemetry-exporter-otlp-proto-http`; Node.js: `@opentelemetry/api`, `@opentelemetry/sdk-node`, `@opentelemetry/exporter-trace-otlp-http`)
|
|
13
|
+
|
|
14
|
+
2. **Initialize tracer** with OTLP exporter pointing to `OTEL_EXPORTER_OTLP_ENDPOINT` using `http/protobuf` protocol
|
|
15
|
+
|
|
16
|
+
3. **Create spans** in this hierarchy:
|
|
17
|
+
```
|
|
18
|
+
Root: invoke_agent (gen_ai.operation.name = "invoke_agent")
|
|
19
|
+
├── chat (gen_ai.operation.name = "chat")
|
|
20
|
+
├── execute_tool (gen_ai.operation.name = "execute_tool")
|
|
21
|
+
└── ...
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
4. **Set required attributes:**
|
|
25
|
+
|
|
26
|
+
Agent root span:
|
|
27
|
+
- `gen_ai.operation.name`: `"invoke_agent"`
|
|
28
|
+
- `gen_ai.agent.name`: agent name
|
|
29
|
+
- `gen_ai.system`: provider (e.g., `"aws.bedrock"`, `"openai"`, `"anthropic"`)
|
|
30
|
+
- `gen_ai.conversation.id`: unique run/session ID (links the trace to an Agent Health run; also accepted as `agent_health.run.id`. `gen_ai.request.id` is not a registered attribute — don't use it)
|
|
31
|
+
|
|
32
|
+
LLM spans:
|
|
33
|
+
- `gen_ai.operation.name`: `"chat"`
|
|
34
|
+
- `gen_ai.request.model`: full model ID (e.g., `"anthropic.claude-sonnet-4-20250514-v1:0"`)
|
|
35
|
+
- `gen_ai.usage.input_tokens`: integer
|
|
36
|
+
- `gen_ai.usage.output_tokens`: integer
|
|
37
|
+
|
|
38
|
+
Tool spans:
|
|
39
|
+
- `gen_ai.operation.name`: `"execute_tool"`
|
|
40
|
+
- `gen_ai.tool.name`: tool name
|
|
41
|
+
|
|
42
|
+
5. **Set environment variables** to emit traces:
|
|
43
|
+
```bash
|
|
44
|
+
export OTEL_TRACES_EXPORTER=otlp
|
|
45
|
+
export OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
|
|
46
|
+
export OTEL_EXPORTER_OTLP_ENDPOINT=<YOUR_OTLP_ENDPOINT>
|
|
47
|
+
export OTEL_SERVICE_NAME=my-agent
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
6. **Configure Agent Health** to read traces from the OpenSearch domain:
|
|
51
|
+
|
|
52
|
+
**Option 1: CLI (recommended)**
|
|
53
|
+
```bash
|
|
54
|
+
npx @opensearch-project/agent-health setup-telemetry
|
|
55
|
+
```
|
|
56
|
+
This writes the observability config to `agent-health.config.json` automatically.
|
|
57
|
+
|
|
58
|
+
**Option 2: Manual JSON config**
|
|
59
|
+
Add to `agent-health.config.json`:
|
|
60
|
+
```json
|
|
61
|
+
{
|
|
62
|
+
"observability": {
|
|
63
|
+
"endpoint": "https://search-my-domain.us-west-2.es.amazonaws.com",
|
|
64
|
+
"authType": "sigv4",
|
|
65
|
+
"awsRegion": "us-west-2",
|
|
66
|
+
"awsService": "es",
|
|
67
|
+
"tracesIndex": "otel-v1-apm-span-*"
|
|
68
|
+
}
|
|
69
|
+
}
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
7. **Verify** with `npx @opensearch-project/agent-health doctor`
|
|
73
|
+
|
|
74
|
+
### Key Rules
|
|
75
|
+
|
|
76
|
+
- Token counts MUST be integers, not strings
|
|
77
|
+
- Always call `span.end()` (or use context manager)
|
|
78
|
+
- Use full model IDs (e.g., `anthropic.claude-sonnet-4-20250514-v1:0`), not short names
|
|
79
|
+
- Every span needs `gen_ai.operation.name` or it falls into "OTHER" category
|
|
80
|
+
- Set span status to ERROR on failures
|
|
81
|
+
|
|
82
|
+
### Reference
|
|
83
|
+
|
|
84
|
+
Full guide with code examples: [docs/INSTRUMENT_WITH_OTEL.md](../../../docs/INSTRUMENT_WITH_OTEL.md)
|
|
@@ -0,0 +1,124 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: write-test
|
|
3
|
+
description: Use when writing, modifying, or debugging tests. Provides project-specific test conventions, mocking patterns, import rules, and integration test cleanup requirements.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
## Test Conventions
|
|
7
|
+
|
|
8
|
+
**File naming:** `<module-name>.test.ts` in `tests/unit/<path-mirroring-source>/`
|
|
9
|
+
|
|
10
|
+
**Required header:**
|
|
11
|
+
```typescript
|
|
12
|
+
/*
|
|
13
|
+
* Copyright OpenSearch Contributors
|
|
14
|
+
* SPDX-License-Identifier: Apache-2.0
|
|
15
|
+
*/
|
|
16
|
+
```
|
|
17
|
+
|
|
18
|
+
**Basic structure:**
|
|
19
|
+
```typescript
|
|
20
|
+
import { functionToTest } from '@/path/to/module';
|
|
21
|
+
|
|
22
|
+
jest.mock('@/services/storage/opensearchClient', () => ({
|
|
23
|
+
benchmarkStorage: { getAll: jest.fn(), getById: jest.fn() },
|
|
24
|
+
}));
|
|
25
|
+
|
|
26
|
+
describe('ModuleName', () => {
|
|
27
|
+
beforeEach(() => { jest.clearAllMocks(); });
|
|
28
|
+
|
|
29
|
+
describe('functionToTest', () => {
|
|
30
|
+
it('should handle normal case', () => {
|
|
31
|
+
const result = functionToTest('test');
|
|
32
|
+
expect(result).toBe('expected');
|
|
33
|
+
});
|
|
34
|
+
});
|
|
35
|
+
});
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
## Import Conventions
|
|
39
|
+
|
|
40
|
+
**Always use `@/` path alias** — never relative paths like `../`:
|
|
41
|
+
|
|
42
|
+
| Module | Import Path |
|
|
43
|
+
|--------|-------------|
|
|
44
|
+
| Types | `@/types` |
|
|
45
|
+
| Server routes | `@/server/routes/<name>` |
|
|
46
|
+
| Server services | `@/server/services/<name>` |
|
|
47
|
+
| Frontend services | `@/services/<category>/<name>` |
|
|
48
|
+
| Lib utilities | `@/lib/<name>` |
|
|
49
|
+
| CLI commands | `@/cli/commands/<name>` |
|
|
50
|
+
|
|
51
|
+
## Mocking Patterns
|
|
52
|
+
|
|
53
|
+
**Mock external modules:**
|
|
54
|
+
```typescript
|
|
55
|
+
jest.mock('@/services/storage/opensearchClient', () => ({
|
|
56
|
+
benchmarkStorage: {
|
|
57
|
+
getAll: jest.fn().mockResolvedValue([]),
|
|
58
|
+
getById: jest.fn().mockResolvedValue(null),
|
|
59
|
+
},
|
|
60
|
+
isStorageConfigured: true,
|
|
61
|
+
}));
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
**Mock with type safety:**
|
|
65
|
+
```typescript
|
|
66
|
+
const mockStorage = benchmarkStorage as jest.Mocked<typeof benchmarkStorage>;
|
|
67
|
+
mockStorage.getAll.mockResolvedValue([mockBenchmark]);
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
**Dynamic require for module re-import:**
|
|
71
|
+
```typescript
|
|
72
|
+
jest.resetModules();
|
|
73
|
+
process.env.MY_VAR = 'new-value';
|
|
74
|
+
const { myConfig } = require('@/lib/config');
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
**Suppress console:** `jest.spyOn(console, 'log').mockImplementation(() => {});`
|
|
78
|
+
|
|
79
|
+
## Async & SSE Testing
|
|
80
|
+
|
|
81
|
+
```typescript
|
|
82
|
+
// Async
|
|
83
|
+
mockFetch.mockResolvedValue({ data: 'test' });
|
|
84
|
+
const result = await fetchData();
|
|
85
|
+
|
|
86
|
+
// Error
|
|
87
|
+
await expect(fetchData()).rejects.toThrow('Network error');
|
|
88
|
+
|
|
89
|
+
// SSE stream
|
|
90
|
+
const mockReader = {
|
|
91
|
+
read: jest.fn()
|
|
92
|
+
.mockResolvedValueOnce({ done: false, value: encoder.encode('data: {"type":"start"}\n\n') })
|
|
93
|
+
.mockResolvedValueOnce({ done: true }),
|
|
94
|
+
};
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
## Global Mocks (`__mocks__/`)
|
|
98
|
+
|
|
99
|
+
Modules using `import.meta.url` or browser-only APIs need global mocks via `__mocks__/` + `moduleNameMapper` in `jest.config.cjs`:
|
|
100
|
+
- `__mocks__/@/lib/config.ts`, `__mocks__/@/server/services/configService.ts`, `__mocks__/@/server/utils/version.ts`
|
|
101
|
+
- `__mocks__/dagre.ts`, `__mocks__/xyflow-react.ts`
|
|
102
|
+
|
|
103
|
+
## Test Categories
|
|
104
|
+
|
|
105
|
+
**Unit tests** (`tests/unit/`): Mock ALL dependencies, <100ms per test, no network/filesystem.
|
|
106
|
+
|
|
107
|
+
**Integration tests** (`tests/integration/`): Real services, `*.integration.test.ts`, 30s timeout. **Always clean up in `afterAll`:**
|
|
108
|
+
```typescript
|
|
109
|
+
const createdTestCaseIds: string[] = [];
|
|
110
|
+
const createdBenchmarkIds: string[] = [];
|
|
111
|
+
|
|
112
|
+
afterAll(async () => {
|
|
113
|
+
for (const id of createdTestCaseIds) {
|
|
114
|
+
await fetch(`${BASE_URL}/api/storage/test-cases/${encodeURIComponent(id)}`, { method: 'DELETE' }).catch(() => {});
|
|
115
|
+
}
|
|
116
|
+
for (const id of createdBenchmarkIds) {
|
|
117
|
+
await fetch(`${BASE_URL}/api/storage/benchmarks/${encodeURIComponent(id)}`, { method: 'DELETE' }).catch(() => {});
|
|
118
|
+
}
|
|
119
|
+
});
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
## Coverage Thresholds
|
|
123
|
+
|
|
124
|
+
Lines: 90%, Statements: 90%, Functions: 80%, Branches: 80%
|
package/docs/ui prd.md
ADDED
|
@@ -0,0 +1,376 @@
|
|
|
1
|
+
# Agent Evaluation Framework - UI PRD
|
|
2
|
+
|
|
3
|
+
## 1. Purpose
|
|
4
|
+
|
|
5
|
+
A framework for testing AI agents that perform multi-step tasks using tools. Users define test cases with expected outcomes, run agents against them, and receive evaluations from an LLM judge.
|
|
6
|
+
|
|
7
|
+
## 2. Core Principle
|
|
8
|
+
|
|
9
|
+
**Make creating, running, and evaluating agents intuitive.**
|
|
10
|
+
|
|
11
|
+
The user's mental model is simple:
|
|
12
|
+
```
|
|
13
|
+
Define what you want → Run the agent → See if it worked
|
|
14
|
+
```
|
|
15
|
+
|
|
16
|
+
The complexity of trajectory capture, LLM judging, and metrics should be invisible unless the user wants to dig deeper.
|
|
17
|
+
|
|
18
|
+
---
|
|
19
|
+
|
|
20
|
+
## 3. Goals
|
|
21
|
+
|
|
22
|
+
### 3.1 Primary Goals
|
|
23
|
+
|
|
24
|
+
| Goal | Description |
|
|
25
|
+
|------|-------------|
|
|
26
|
+
| **Easy test creation** | Users should define test cases by describing what they want, not by understanding framework internals |
|
|
27
|
+
| **Flexible success criteria** | Support multiple ways to define "correct" behavior (trajectory, output, natural language) |
|
|
28
|
+
| **Clear feedback** | Pass/fail should be obvious; reasoning should be human-readable |
|
|
29
|
+
| **Comparison capability** | Users need to compare agent performance across configurations and over time |
|
|
30
|
+
| **Iterative workflow** | Quick re-run with modifications, not start-from-scratch each time |
|
|
31
|
+
|
|
32
|
+
### 3.2 User Personas
|
|
33
|
+
|
|
34
|
+
**Agent Developer**: Building/improving an AI agent. Needs to test changes, identify regressions, understand why agent fails.
|
|
35
|
+
|
|
36
|
+
**QA/Evaluator**: Running systematic tests across agent versions. Needs batch execution, comparison, reporting.
|
|
37
|
+
|
|
38
|
+
**Product Owner**: Wants to understand agent quality without technical depth. Needs clear pass/fail, trends.
|
|
39
|
+
|
|
40
|
+
### 3.3 Success Metrics
|
|
41
|
+
|
|
42
|
+
| Metric | Target |
|
|
43
|
+
|--------|--------|
|
|
44
|
+
| Time to create first test case | < 5 minutes |
|
|
45
|
+
| Time to run first evaluation | < 1 minute after setup |
|
|
46
|
+
| Understand pass/fail reason | Without technical knowledge |
|
|
47
|
+
|
|
48
|
+
---
|
|
49
|
+
|
|
50
|
+
## 4. Data Model
|
|
51
|
+
|
|
52
|
+
### 4.1 TestCase
|
|
53
|
+
|
|
54
|
+
A test scenario defining what to test and how to measure success.
|
|
55
|
+
|
|
56
|
+
| Field | Type | Required | Description |
|
|
57
|
+
|-------|------|----------|-------------|
|
|
58
|
+
| id | string | auto | Unique identifier |
|
|
59
|
+
| name | string | yes | Human-readable name |
|
|
60
|
+
| description | string | no | Longer explanation of what this tests |
|
|
61
|
+
| category | string | yes | Grouping (e.g., "Query Generation", "Error Handling") |
|
|
62
|
+
| subcategory | string | no | Further grouping within category |
|
|
63
|
+
| difficulty | enum | yes | "Easy", "Medium", "Hard" |
|
|
64
|
+
| initialPrompt | string | yes | The question/task to send to the agent |
|
|
65
|
+
| context | AgentContextItem[] | no | Data provided to the agent |
|
|
66
|
+
| tools | AgentToolDefinition[] | no | Tools available to the agent |
|
|
67
|
+
| expectedOutcomes | ExpectedOutcome[] | yes | How to measure success (see 4.2) |
|
|
68
|
+
| isPromoted | boolean | no | Available for experiments (default: false) |
|
|
69
|
+
| currentVersion | number | auto | Latest version number |
|
|
70
|
+
| versions | TestCaseVersion[] | auto | Immutable history of all versions |
|
|
71
|
+
| createdAt | timestamp | auto | Creation time |
|
|
72
|
+
| updatedAt | timestamp | auto | Last modification time |
|
|
73
|
+
|
|
74
|
+
**Versioning behavior**: Each save creates a new immutable version. Previous versions are preserved for comparison and audit.
|
|
75
|
+
|
|
76
|
+
### 4.2 ExpectedOutcome
|
|
77
|
+
|
|
78
|
+
Flexible definition of success. A test case can have one or more expected outcomes of different types.
|
|
79
|
+
|
|
80
|
+
| Field | Type | Required | Description |
|
|
81
|
+
|-------|------|----------|-------------|
|
|
82
|
+
| type | enum | yes | "trajectory", "output", "criteria" |
|
|
83
|
+
| weight | number | no | Relative importance for scoring (default: 1.0) |
|
|
84
|
+
|
|
85
|
+
**Type: trajectory** — The agent should follow a specific sequence of steps
|
|
86
|
+
|
|
87
|
+
| Field | Type | Description |
|
|
88
|
+
|-------|------|-------------|
|
|
89
|
+
| steps | TrajectoryExpectation[] | Ordered list of expected steps |
|
|
90
|
+
|
|
91
|
+
Each step:
|
|
92
|
+
| Field | Type | Description |
|
|
93
|
+
|-------|------|-------------|
|
|
94
|
+
| step | number | Order in sequence (1, 2, 3...) |
|
|
95
|
+
| description | string | What should happen at this step |
|
|
96
|
+
| requiredTools | string[] | Tool(s) that must be called |
|
|
97
|
+
| optional | boolean | If true, step can be skipped without penalty |
|
|
98
|
+
|
|
99
|
+
**Type: output** — The agent should produce a specific verifiable result
|
|
100
|
+
|
|
101
|
+
| Field | Type | Description |
|
|
102
|
+
|-------|------|-------------|
|
|
103
|
+
| field | string | What to check (e.g., "pplQuery", "finalAnswer") |
|
|
104
|
+
| operator | enum | "equals", "contains", "matches" (regex), "exists" |
|
|
105
|
+
| value | string | Expected value or pattern |
|
|
106
|
+
|
|
107
|
+
**Type: criteria** — Natural language description evaluated by LLM judge
|
|
108
|
+
|
|
109
|
+
| Field | Type | Description |
|
|
110
|
+
|-------|------|-------------|
|
|
111
|
+
| description | string | What success looks like in plain language |
|
|
112
|
+
|
|
113
|
+
### 4.3 AgentContextItem
|
|
114
|
+
|
|
115
|
+
Context data passed to the agent at runtime.
|
|
116
|
+
|
|
117
|
+
| Field | Type | Description |
|
|
118
|
+
|-------|------|-------------|
|
|
119
|
+
| description | string | Human-readable label for what this context represents |
|
|
120
|
+
| value | string | The context data (JSON-stringified if complex) |
|
|
121
|
+
|
|
122
|
+
### 4.4 AgentToolDefinition
|
|
123
|
+
|
|
124
|
+
Tool available to the agent during execution.
|
|
125
|
+
|
|
126
|
+
| Field | Type | Description |
|
|
127
|
+
|-------|------|-------------|
|
|
128
|
+
| name | string | Tool identifier |
|
|
129
|
+
| description | string | What the tool does |
|
|
130
|
+
| parameters | object | JSON Schema defining the tool's parameters |
|
|
131
|
+
|
|
132
|
+
### 4.5 Agent (Optional)
|
|
133
|
+
|
|
134
|
+
Configuration for an agent endpoint. **Agent management is optional.**
|
|
135
|
+
|
|
136
|
+
In the simplest case, the user has a single agent endpoint that evolves as they develop. The endpoint URL stays the same, but the agent's behavior changes. In more complex setups, users may have multiple agent instances (different versions, A/B testing, etc.).
|
|
137
|
+
|
|
138
|
+
| Field | Type | Required | Description |
|
|
139
|
+
|-------|------|----------|-------------|
|
|
140
|
+
| key | string | yes | Unique identifier (used internally) |
|
|
141
|
+
| name | string | yes | Display name |
|
|
142
|
+
| endpoint | string | yes | URL to call |
|
|
143
|
+
| description | string | no | What this agent does |
|
|
144
|
+
| enabled | boolean | no | Can be selected for runs (default: true) |
|
|
145
|
+
| headers | Record<string, string> | no | Custom headers for requests (e.g., auth tokens) |
|
|
146
|
+
|
|
147
|
+
**Note:** If only one agent is configured, the UI should streamline the experience (e.g., skip agent selection step).
|
|
148
|
+
|
|
149
|
+
### 4.6 Model
|
|
150
|
+
|
|
151
|
+
LLM model configuration.
|
|
152
|
+
|
|
153
|
+
| Field | Type | Description |
|
|
154
|
+
|-------|------|-------------|
|
|
155
|
+
| id | string | Model identifier (e.g., "claude-3-sonnet") |
|
|
156
|
+
| displayName | string | Human-readable name |
|
|
157
|
+
| contextWindow | number | Maximum input tokens |
|
|
158
|
+
| maxOutputTokens | number | Maximum output tokens |
|
|
159
|
+
|
|
160
|
+
### 4.7 TestCaseRun
|
|
161
|
+
|
|
162
|
+
Result of running a single test case against an agent.
|
|
163
|
+
|
|
164
|
+
| Field | Type | Description |
|
|
165
|
+
|-------|------|-------------|
|
|
166
|
+
| id | string | Unique identifier |
|
|
167
|
+
| timestamp | timestamp | When the run started |
|
|
168
|
+
| testCaseId | string | Reference to TestCase |
|
|
169
|
+
| testCaseVersion | number | Which version of the test case was run |
|
|
170
|
+
| agentKey | string | Which agent was used |
|
|
171
|
+
| modelId | string | Which model was used |
|
|
172
|
+
| status | enum | "running", "completed", "failed" |
|
|
173
|
+
| passFailStatus | enum | "passed", "failed" — determined by LLM judge |
|
|
174
|
+
| trajectory | TrajectoryStep[] | Actual steps the agent took |
|
|
175
|
+
| metrics | EvaluationMetrics | Scores from the judge |
|
|
176
|
+
| llmJudgeReasoning | string | Human-readable explanation of the evaluation |
|
|
177
|
+
| improvementStrategies | ImprovementStrategy[] | Actionable suggestions |
|
|
178
|
+
| logs | OpenSearchLog[] | Debug logs (if available) |
|
|
179
|
+
| rawEvents | any[] | Raw agent protocol events |
|
|
180
|
+
|
|
181
|
+
### 4.8 TrajectoryStep
|
|
182
|
+
|
|
183
|
+
A single action captured during agent execution.
|
|
184
|
+
|
|
185
|
+
| Field | Type | Description |
|
|
186
|
+
|-------|------|-------------|
|
|
187
|
+
| id | string | Unique identifier |
|
|
188
|
+
| timestamp | number | When the step occurred (ms since epoch) |
|
|
189
|
+
| type | enum | "tool_result", "thought", "action", "response" |
|
|
190
|
+
| content | string | Human-readable description of the step |
|
|
191
|
+
| toolName | string | Tool that was called (if type involves tools) |
|
|
192
|
+
| toolArgs | object | Arguments passed to the tool |
|
|
193
|
+
| toolOutput | any | Result returned by the tool |
|
|
194
|
+
| status | enum | "SUCCESS", "FAILURE" |
|
|
195
|
+
| latencyMs | number | How long this step took |
|
|
196
|
+
|
|
197
|
+
### 4.9 EvaluationMetrics
|
|
198
|
+
|
|
199
|
+
Scores produced by the LLM judge.
|
|
200
|
+
|
|
201
|
+
| Field | Type | Range | Description |
|
|
202
|
+
|-------|------|-------|-------------|
|
|
203
|
+
| accuracy | number | 0-100 | Did the agent get the correct result? |
|
|
204
|
+
| faithfulness | number | 0-100 | Did the agent follow instructions/context? |
|
|
205
|
+
| latency_score | number | 0-100 | Was the agent efficient? |
|
|
206
|
+
| trajectory_alignment_score | number | 0-100 | Did the agent follow the expected path? |
|
|
207
|
+
|
|
208
|
+
### 4.10 ImprovementStrategy
|
|
209
|
+
|
|
210
|
+
Actionable suggestion from the judge.
|
|
211
|
+
|
|
212
|
+
| Field | Type | Description |
|
|
213
|
+
|-------|------|-------------|
|
|
214
|
+
| category | string | Area of improvement (e.g., "tool_usage", "reasoning") |
|
|
215
|
+
| issue | string | What went wrong |
|
|
216
|
+
| recommendation | string | How to fix it |
|
|
217
|
+
| priority | enum | "high", "medium", "low" |
|
|
218
|
+
|
|
219
|
+
### 4.11 Experiment
|
|
220
|
+
|
|
221
|
+
A batch of test cases run together for systematic evaluation.
|
|
222
|
+
|
|
223
|
+
| Field | Type | Description |
|
|
224
|
+
|-------|------|-------------|
|
|
225
|
+
| id | string | Unique identifier |
|
|
226
|
+
| name | string | Display name (e.g., "v2.0 Release Validation") |
|
|
227
|
+
| description | string | What this experiment is testing |
|
|
228
|
+
| useCaseIds | string[] | Which test cases are included |
|
|
229
|
+
| runs | ExperimentRun[] | Execution snapshots |
|
|
230
|
+
| createdAt | timestamp | When created |
|
|
231
|
+
| updatedAt | timestamp | Last modified |
|
|
232
|
+
|
|
233
|
+
### 4.12 ExperimentRun
|
|
234
|
+
|
|
235
|
+
A single execution of an experiment. **This is a snapshot in time.**
|
|
236
|
+
|
|
237
|
+
When an experiment run is created, it captures the agent configuration and environment at that moment. Even if the agent changes later (same endpoint, different behavior), the run record preserves what was tested.
|
|
238
|
+
|
|
239
|
+
| Field | Type | Description |
|
|
240
|
+
|-------|------|-------------|
|
|
241
|
+
| id | string | Unique identifier |
|
|
242
|
+
| name | string | Label (e.g., "Baseline", "With Fix v2", "Claude 4 Test") |
|
|
243
|
+
| description | string | What makes this run different |
|
|
244
|
+
| createdAt | timestamp | When executed |
|
|
245
|
+
| results | Record<string, RunResult> | testCaseId → result mapping |
|
|
246
|
+
|
|
247
|
+
**Snapshot fields** (captured at run time):
|
|
248
|
+
| Field | Type | Description |
|
|
249
|
+
|-------|------|-------------|
|
|
250
|
+
| agentKey | string | Agent identifier at time of run |
|
|
251
|
+
| agentEndpoint | string | Actual endpoint URL used |
|
|
252
|
+
| agentHeaders | Record<string, string> | Headers used (sanitized) |
|
|
253
|
+
| modelId | string | Model used |
|
|
254
|
+
| environmentNotes | string | Optional: user can note environment state (e.g., "commit abc123", "prod config") |
|
|
255
|
+
|
|
256
|
+
### 4.13 RunResult
|
|
257
|
+
|
|
258
|
+
Status of a single test case within an experiment run.
|
|
259
|
+
|
|
260
|
+
| Field | Type | Description |
|
|
261
|
+
|-------|------|-------------|
|
|
262
|
+
| reportId | string | Reference to TestCaseRun.id |
|
|
263
|
+
| status | enum | "pending", "running", "completed", "failed" |
|
|
264
|
+
|
|
265
|
+
---
|
|
266
|
+
|
|
267
|
+
## 5. User Capabilities
|
|
268
|
+
|
|
269
|
+
The UI must enable users to perform these actions. How these are organized into screens/flows is flexible.
|
|
270
|
+
|
|
271
|
+
### 5.1 Test Case Management
|
|
272
|
+
|
|
273
|
+
| Capability | Description |
|
|
274
|
+
|------------|-------------|
|
|
275
|
+
| **Create test case** | Define a new test with prompt, context, tools, and expected outcomes |
|
|
276
|
+
| **Edit test case** | Modify an existing test case (creates new version) |
|
|
277
|
+
| **View test case** | See all details of a test case including its definition and history |
|
|
278
|
+
| **Delete test case** | Remove a test case |
|
|
279
|
+
| **Browse test cases** | Find test cases by category, difficulty, search, or other filters |
|
|
280
|
+
| **View version history** | See how a test case changed over time, compare versions |
|
|
281
|
+
| **Promote/demote** | Control whether a test case appears in experiment selection |
|
|
282
|
+
|
|
283
|
+
### 5.2 Agent Management (Optional)
|
|
284
|
+
|
|
285
|
+
Agent management is optional. In the simplest workflow, the user provides an endpoint at run time.
|
|
286
|
+
|
|
287
|
+
| Capability | Description |
|
|
288
|
+
|------------|-------------|
|
|
289
|
+
| **Quick run with endpoint** | User can enter an endpoint URL directly without saving an agent config |
|
|
290
|
+
| **Add agent** | Save an agent configuration for reuse |
|
|
291
|
+
| **Edit agent** | Modify agent configuration |
|
|
292
|
+
| **Delete agent** | Remove an agent |
|
|
293
|
+
| **Enable/disable agent** | Control whether agent appears in run selection |
|
|
294
|
+
| **Test connection** | Verify agent endpoint is reachable |
|
|
295
|
+
|
|
296
|
+
**Streamlined mode:** When only one agent exists (or none saved), skip agent selection UI and use the available/provided endpoint directly.
|
|
297
|
+
|
|
298
|
+
### 5.3 Evaluation Execution
|
|
299
|
+
|
|
300
|
+
| Capability | Description |
|
|
301
|
+
|------------|-------------|
|
|
302
|
+
| **Run single evaluation** | Execute one test case against one agent/model |
|
|
303
|
+
| **See live progress** | Watch trajectory steps appear in real-time during execution |
|
|
304
|
+
| **Cancel running evaluation** | Stop an in-progress run |
|
|
305
|
+
| **Re-run with same config** | Quickly repeat an evaluation |
|
|
306
|
+
| **Re-run with different config** | Run same test case with different agent/model |
|
|
307
|
+
|
|
308
|
+
### 5.4 Results & Reporting
|
|
309
|
+
|
|
310
|
+
| Capability | Description |
|
|
311
|
+
|------------|-------------|
|
|
312
|
+
| **View run results** | See verdict, metrics, reasoning, trajectory for any run |
|
|
313
|
+
| **Browse run history** | Find past runs with filters (date, agent, model, pass/fail, test case) |
|
|
314
|
+
| **Drill into details** | Expand trajectory steps, view raw events, see logs |
|
|
315
|
+
| **Export results** | Download results in standard format (CSV, JSON) |
|
|
316
|
+
|
|
317
|
+
### 5.5 Experiments & Comparison
|
|
318
|
+
|
|
319
|
+
Experiments enable systematic comparison. Each run is a **snapshot** — it captures the agent config and environment at that moment, so comparisons remain valid even as the agent evolves.
|
|
320
|
+
|
|
321
|
+
| Capability | Description |
|
|
322
|
+
|------------|-------------|
|
|
323
|
+
| **Create experiment** | Define a batch of test cases to run together |
|
|
324
|
+
| **Add run to experiment** | Execute the experiment with a new configuration (captures snapshot) |
|
|
325
|
+
| **Note environment** | Optionally record environment state (commit hash, config notes, etc.) |
|
|
326
|
+
| **View experiment results** | See aggregate pass/fail and metrics per run |
|
|
327
|
+
| **Compare runs** | Side-by-side comparison of two or more runs |
|
|
328
|
+
| **Identify regressions** | See which test cases got worse between runs |
|
|
329
|
+
| **Identify improvements** | See which test cases got better between runs |
|
|
330
|
+
|
|
331
|
+
### 5.6 Configuration
|
|
332
|
+
|
|
333
|
+
| Capability | Description |
|
|
334
|
+
|------------|-------------|
|
|
335
|
+
| **Configure judge** | Set Bedrock model, endpoint for LLM evaluation |
|
|
336
|
+
| **Configure logging** | Set OpenSearch connection for debug logs |
|
|
337
|
+
| **Set preferences** | Default filters, display options |
|
|
338
|
+
|
|
339
|
+
---
|
|
340
|
+
|
|
341
|
+
## 6. Constraints & Requirements
|
|
342
|
+
|
|
343
|
+
### 6.1 Technical Constraints
|
|
344
|
+
|
|
345
|
+
- Agent communication uses AG-UI protocol (SSE streaming)
|
|
346
|
+
- LLM Judge runs via AWS Bedrock (backend proxy required)
|
|
347
|
+
- Data persisted to localStorage (no backend database)
|
|
348
|
+
- Must handle long-running evaluations (seconds to minutes)
|
|
349
|
+
|
|
350
|
+
### 6.2 UX Requirements
|
|
351
|
+
|
|
352
|
+
| Requirement | Rationale |
|
|
353
|
+
|-------------|-----------|
|
|
354
|
+
| Pass/fail must be immediately visible | Primary user question is "did it work?" |
|
|
355
|
+
| Progressive disclosure | Simple view first, details on demand |
|
|
356
|
+
| Live streaming during runs | Users need to know something is happening |
|
|
357
|
+
| Human-readable judge output | Non-technical users must understand failures |
|
|
358
|
+
| Non-destructive editing | Never lose test case history |
|
|
359
|
+
|
|
360
|
+
### 6.3 Performance Expectations
|
|
361
|
+
|
|
362
|
+
| Operation | Expected Duration |
|
|
363
|
+
|-----------|-------------------|
|
|
364
|
+
| Load test case list | < 1 second |
|
|
365
|
+
| Single evaluation | 10-60 seconds (agent + judge) |
|
|
366
|
+
| Experiment run (10 cases) | 2-10 minutes |
|
|
367
|
+
|
|
368
|
+
---
|
|
369
|
+
|
|
370
|
+
## 7. Out of Scope
|
|
371
|
+
|
|
372
|
+
- User authentication / multi-tenancy
|
|
373
|
+
- Backend database (uses localStorage)
|
|
374
|
+
- Real-time collaboration
|
|
375
|
+
- Scheduled/automated runs
|
|
376
|
+
- CI/CD integration
|
|
@@ -0,0 +1,53 @@
|
|
|
1
|
+
# Examples
|
|
2
|
+
|
|
3
|
+
Hand-written, runnable examples that ship inside the published
|
|
4
|
+
`@opensearch-project/agent-health` npm tarball.
|
|
5
|
+
|
|
6
|
+
When you install agent-health, these files land at:
|
|
7
|
+
|
|
8
|
+
```
|
|
9
|
+
node_modules/@opensearch-project/agent-health/examples/
|
|
10
|
+
```
|
|
11
|
+
|
|
12
|
+
The point of shipping them in the tarball (instead of only on GitHub) is that
|
|
13
|
+
the source you read here is **exactly** the version your `package.json`
|
|
14
|
+
resolved to — no risk of looking at `main` and being misled by changes that
|
|
15
|
+
landed after your version was cut. AI assistants debugging your project can
|
|
16
|
+
also read these files directly.
|
|
17
|
+
|
|
18
|
+
## Layout
|
|
19
|
+
|
|
20
|
+
| Directory | What's inside |
|
|
21
|
+
| ------------------ | ------------------------------------------------------- |
|
|
22
|
+
| `eval-files/` | Code-based test cases (`.eval.js` / `.eval.ts`) |
|
|
23
|
+
| `connectors/` | Custom `BaseConnector` subclasses |
|
|
24
|
+
| `config/` | Annotated `agent-health.config.ts` examples |
|
|
25
|
+
|
|
26
|
+
Each example file is self-contained, heavily commented, and runnable as-is.
|
|
27
|
+
|
|
28
|
+
## Running
|
|
29
|
+
|
|
30
|
+
For SDK examples (`eval-files/`):
|
|
31
|
+
|
|
32
|
+
```bash
|
|
33
|
+
# Copy or symlink the example into your project's evals/ dir, then run it:
|
|
34
|
+
npx @opensearch-project/agent-health benchmark \
|
|
35
|
+
--source-type code-import \
|
|
36
|
+
--files ./evals/<copied-file>.eval.js \
|
|
37
|
+
--agent <your-agent-key>
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
For connector examples, see the per-file header for the import / register
|
|
41
|
+
snippet you drop into your `agent-health.config.ts`.
|
|
42
|
+
|
|
43
|
+
## Where to look first
|
|
44
|
+
|
|
45
|
+
If you're debugging behavior that doesn't match what you expect:
|
|
46
|
+
|
|
47
|
+
1. Read the **source you imported from** under
|
|
48
|
+
`node_modules/@opensearch-project/agent-health/lib/dist/lib/` — the
|
|
49
|
+
compiled-but-readable `.js` files preserve the original JSDoc and structure
|
|
50
|
+
so you can navigate by import path.
|
|
51
|
+
2. Cross-reference against the matching `.d.ts` for the exact type contract.
|
|
52
|
+
3. If something is unclear, an example here usually demonstrates the intended
|
|
53
|
+
usage pattern.
|