agentx-python 0.8.9__tar.gz → 0.8.11__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (90) hide show
  1. {agentx_python-0.8.9/agentx_python.egg-info → agentx_python-0.8.11}/PKG-INFO +19 -28
  2. {agentx_python-0.8.9 → agentx_python-0.8.11}/README.md +18 -27
  3. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/evaluations/client.py +12 -1
  4. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/version.py +2 -2
  5. {agentx_python-0.8.9 → agentx_python-0.8.11/agentx_python.egg-info}/PKG-INFO +19 -28
  6. {agentx_python-0.8.9 → agentx_python-0.8.11}/LICENSE +0 -0
  7. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/__init__.py +0 -0
  8. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/agentx.py +0 -0
  9. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/cli.py +0 -0
  10. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/evaluations/__init__.py +0 -0
  11. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/evaluations/_term.py +0 -0
  12. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/evaluations/adapters/__init__.py +0 -0
  13. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/evaluations/adapters/http_endpoint.py +0 -0
  14. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/evaluations/adapters/precomputed.py +0 -0
  15. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/evaluations/adapters/raw.py +0 -0
  16. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/evaluations/datasets.py +0 -0
  17. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/evaluations/evaluation_settings.py +0 -0
  18. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/evaluations/models.py +0 -0
  19. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/evaluations/prompts.py +0 -0
  20. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/evaluations/reporting.py +0 -0
  21. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/evaluations/results.py +0 -0
  22. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/evaluations/runner.py +0 -0
  23. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/evaluations/tool_schemas.py +0 -0
  24. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/evaluations/tracing.py +0 -0
  25. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/exceptions.py +0 -0
  26. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/export.py +0 -0
  27. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/feedback.py +0 -0
  28. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/integrations/__init__.py +0 -0
  29. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/integrations/_traced_call.py +0 -0
  30. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/integrations/anthropic.py +0 -0
  31. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/integrations/autogen.py +0 -0
  32. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/integrations/crewai.py +0 -0
  33. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/integrations/databricks.py +0 -0
  34. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/integrations/google_adk.py +0 -0
  35. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/integrations/google_genai.py +0 -0
  36. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/integrations/langchain.py +0 -0
  37. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/integrations/litellm.py +0 -0
  38. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/integrations/llamaindex.py +0 -0
  39. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/integrations/moveworks.py +0 -0
  40. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/integrations/openai.py +0 -0
  41. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/integrations/openai_agents.py +0 -0
  42. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/monitor/__init__.py +0 -0
  43. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/monitor/agents.py +0 -0
  44. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/monitor/client.py +0 -0
  45. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/monitor/judge_scorers.py +0 -0
  46. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/monitor/models.py +0 -0
  47. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/monitor/online_evaluators.py +0 -0
  48. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/monitor/patterns.py +0 -0
  49. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/monitor/profile.py +0 -0
  50. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/monitor/review_queue.py +0 -0
  51. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/monitor/rules.py +0 -0
  52. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/monitor/scorers.py +0 -0
  53. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/monitor/sessions.py +0 -0
  54. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/monitor/signals.py +0 -0
  55. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/outcomes.py +0 -0
  56. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/projects.py +0 -0
  57. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/py.typed +0 -0
  58. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/resources/__init__.py +0 -0
  59. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/resources/agent.py +0 -0
  60. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/resources/conversation.py +0 -0
  61. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/resources/workforce.py +0 -0
  62. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/testing.py +0 -0
  63. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/traces.py +0 -0
  64. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/tracing/__init__.py +0 -0
  65. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/tracing/ci_types.py +0 -0
  66. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/tracing/eval_scope.py +0 -0
  67. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/tracing/ingest_client.py +0 -0
  68. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/tracing/tracer.py +0 -0
  69. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx/util.py +0 -0
  70. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx_python.egg-info/SOURCES.txt +0 -0
  71. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx_python.egg-info/dependency_links.txt +0 -0
  72. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx_python.egg-info/entry_points.txt +0 -0
  73. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx_python.egg-info/not-zip-safe +0 -0
  74. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx_python.egg-info/requires.txt +0 -0
  75. {agentx_python-0.8.9 → agentx_python-0.8.11}/agentx_python.egg-info/top_level.txt +0 -0
  76. {agentx_python-0.8.9 → agentx_python-0.8.11}/setup.cfg +0 -0
  77. {agentx_python-0.8.9 → agentx_python-0.8.11}/setup.py +0 -0
  78. {agentx_python-0.8.9 → agentx_python-0.8.11}/tests/test_cli_launcher.py +0 -0
  79. {agentx_python-0.8.9 → agentx_python-0.8.11}/tests/test_deep_dive_fixes.py +0 -0
  80. {agentx_python-0.8.9 → agentx_python-0.8.11}/tests/test_docs_match_sdk.py +0 -0
  81. {agentx_python-0.8.9 → agentx_python-0.8.11}/tests/test_eval_scope.py +0 -0
  82. {agentx_python-0.8.9 → agentx_python-0.8.11}/tests/test_integration.py +0 -0
  83. {agentx_python-0.8.9 → agentx_python-0.8.11}/tests/test_integrations.py +0 -0
  84. {agentx_python-0.8.9 → agentx_python-0.8.11}/tests/test_judge_scorers.py +0 -0
  85. {agentx_python-0.8.9 → agentx_python-0.8.11}/tests/test_pairwise.py +0 -0
  86. {agentx_python-0.8.9 → agentx_python-0.8.11}/tests/test_review_queue.py +0 -0
  87. {agentx_python-0.8.9 → agentx_python-0.8.11}/tests/test_runner_features.py +0 -0
  88. {agentx_python-0.8.9 → agentx_python-0.8.11}/tests/test_selfhost_analysis_fallback.py +0 -0
  89. {agentx_python-0.8.9 → agentx_python-0.8.11}/tests/test_span_tree.py +0 -0
  90. {agentx_python-0.8.9 → agentx_python-0.8.11}/tests/test_testing.py +0 -0
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: agentx-python
3
- Version: 0.8.9
3
+ Version: 0.8.11
4
4
  Summary: Official Python SDK for AgentX (https://www.agentx.so/)
5
5
  Home-page: https://github.com/AgentX-ai/AgentX-python
6
6
  Author: Robin Wang and AgentX Team
@@ -68,7 +68,7 @@ Dynamic: summary
68
68
  [![Python versions](https://img.shields.io/pypi/pyversions/agentx-python)](https://pypi.org/project/agentx-python/)
69
69
  [![License](https://img.shields.io/badge/license-Apache--2.0-blue.svg)](LICENSE)
70
70
 
71
- The official Python SDK for **[AgentX](https://app.agentx.so/)** - an evaluation, tracing, and monitoring framework for AI agents, plus a client for AgentX's own hosted agents.
71
+ The official Python SDK for **[AgentX](https://www.agentx.so/)** - an evaluation, tracing, and monitoring framework for AI agents, plus a client for AgentX's own hosted agents.
72
72
 
73
73
  Also see [SDK Developer Docs](https://developers.agentx.so), [API Reference Docs](https://docs.agentx.so/reference)
74
74
 
@@ -109,19 +109,10 @@ pip install --upgrade agentx-python
109
109
 
110
110
  Requires Python 3.9 or newer.
111
111
 
112
- ---
113
-
114
- ## Authentication
115
-
116
- Get your API key at [app.agentx.so](https://app.agentx.so), then either pass it inline or expose it as an environment variable.
117
-
118
- ```python
119
- # Option A - pass the key inline
120
- from agentx import AgentX
121
- client = AgentX(api_key="your-api-key-here")
112
+ #### Run self host eval framework locally
122
113
 
123
- # Option B - set AGENTX_API_KEY in your environment, then:
124
- client = AgentX.from_env()
114
+ ```
115
+ agentx-trace-eval --dev --update
125
116
  ```
126
117
 
127
118
  ---
@@ -179,7 +170,7 @@ print(report.recommendations) # list of prioritized, actionable fixes
179
170
 
180
171
  Ask a case's question several extra ways each run, LLM-paraphrased server-side, to catch agents that break on phrasing rather than substance, and override the judge's prompt/model per config, see [Smoke testing](EVALUATIONS.md#smoke-testing-phrasing-robustness) and [Configuring the judge](EVALUATIONS.md#configuring-the-judge) in the full guide.
181
172
 
182
- Since AgentX doesn't own your agent's code, `client.evaluations.prompts` lets AgentX become your prompt's *source of truth* instead - the same problem LangSmith's Prompt Hub and Langfuse's Prompt Management solve. Pull a version at runtime, tag your eval runs (or live traces) with it, and let a judge propose a rewrite from your real worst-rated results - a human always has to approve before it publishes:
173
+ Since AgentX doesn't own your agent's code, `client.evaluations.prompts` lets AgentX become your prompt's _source of truth_ instead - the same problem LangSmith's Prompt Hub and Langfuse's Prompt Management solve. Pull a version at runtime, tag your eval runs (or live traces) with it, and let a judge propose a rewrite from your real worst-rated results - a human always has to approve before it publishes:
183
174
 
184
175
  ```python
185
176
  prompt = client.evaluations.prompts.get("support-agent-system-prompt") # or prompt.id
@@ -193,7 +184,7 @@ client.evaluations.run(
193
184
 
194
185
  See [Prompt registry](EVALUATIONS.md#prompt-registry) in the full guide, or [self-host's docs](https://docs.agentx.so/improve/prompt-management) for the "Suggest improvement" dashboard flow (self-host only - no hosted-SaaS equivalent yet).
195
186
 
196
- On self-host, a finalized run can also **gate a CI job**: `report.gate(fail_under=7, no_regression=True)` checks the run's average rating against an absolute floor and/or the dataset's previous run, prints per-check verdicts into the CI log, and returns an exit code - `sys.exit(gate.exit_code)` blocks the merge on regression. Recorded gates appear in the dashboard's CI Gates tab. See [self-host's CI docs](https://docs.agentx.so/integrations/self-host-ci) for the GitHub Actions recipe.
187
+ On self-host, a finalized run can also **gate a CI job**: `run.gate(fail_under=7, no_regression=True)` (on the run context `.execute()` returns) checks the run's average rating against an absolute floor and/or the dataset's previous run, prints per-check verdicts into the CI log, and returns an exit code - `sys.exit(gate.exit_code)` blocks the merge on regression. Recorded gates appear in the dashboard's CI Gates tab. See [self-host's CI docs](https://docs.agentx.so/integrations/self-host-ci) for the GitHub Actions recipe.
197
188
 
198
189
  See **[EVALUATIONS.md](EVALUATIONS.md)** for the full guide - dataset builder, framework adapters, similarity metrics, smoke testing, judge configuration, prompt registry, and the complete API reference.
199
190
 
@@ -248,18 +239,18 @@ regular totals, no extra config needed. Self-host's cost estimate prices these s
248
239
  regular input token when you've set optional cache rates on that model. Install the matching
249
240
  extra:
250
241
 
251
- | Framework | Install | Integration |
252
- | --------------------- | -------------------------------------------- | ------------------------ |
253
- | LangChain | `pip install "agentx-python[langchain]"` | `AgentXCallbackHandler` |
254
- | CrewAI | `pip install "agentx-python[crewai]"` | `AgentXCrewObserver` |
255
- | OpenAI Agents SDK | `pip install "agentx-python[openai-agents]"` | `AgentXTracingProcessor` |
256
- | OpenAI (raw client) | `pip install "agentx-python[openai]"` | `patch_openai_client` |
257
- | Anthropic | `pip install "agentx-python[anthropic]"` | `patch_anthropic_client` |
258
- | Google ADK | `pip install "agentx-python[google-adk]"` | `AgentXADKPlugin` |
259
- | Google GenAI (Gemini) | `pip install "agentx-python[google-genai]"` | `patch_genai_client` |
260
- | LiteLLM | `pip install "agentx-python[litellm]"` | `AgentXLiteLLMLogger` |
261
- | LlamaIndex | `pip install "agentx-python[llamaindex]"` | `AgentXLlamaIndexHandler`|
262
- | AutoGen | `pip install "agentx-python[autogen]"` | `AgentXAutoGenObserver` |
242
+ | Framework | Install | Integration |
243
+ | --------------------- | -------------------------------------------- | ------------------------- |
244
+ | LangChain | `pip install "agentx-python[langchain]"` | `AgentXCallbackHandler` |
245
+ | CrewAI | `pip install "agentx-python[crewai]"` | `AgentXCrewObserver` |
246
+ | OpenAI Agents SDK | `pip install "agentx-python[openai-agents]"` | `AgentXTracingProcessor` |
247
+ | OpenAI (raw client) | `pip install "agentx-python[openai]"` | `patch_openai_client` |
248
+ | Anthropic | `pip install "agentx-python[anthropic]"` | `patch_anthropic_client` |
249
+ | Google ADK | `pip install "agentx-python[google-adk]"` | `AgentXADKPlugin` |
250
+ | Google GenAI (Gemini) | `pip install "agentx-python[google-genai]"` | `patch_genai_client` |
251
+ | LiteLLM | `pip install "agentx-python[litellm]"` | `AgentXLiteLLMLogger` |
252
+ | LlamaIndex | `pip install "agentx-python[llamaindex]"` | `AgentXLlamaIndexHandler` |
253
+ | AutoGen | `pip install "agentx-python[autogen]"` | `AgentXAutoGenObserver` |
263
254
 
264
255
  Or plain Python - wrap any function with `@tracer.trace(...)` and it just works, no framework required.
265
256
 
@@ -4,7 +4,7 @@
4
4
  [![Python versions](https://img.shields.io/pypi/pyversions/agentx-python)](https://pypi.org/project/agentx-python/)
5
5
  [![License](https://img.shields.io/badge/license-Apache--2.0-blue.svg)](LICENSE)
6
6
 
7
- The official Python SDK for **[AgentX](https://app.agentx.so/)** - an evaluation, tracing, and monitoring framework for AI agents, plus a client for AgentX's own hosted agents.
7
+ The official Python SDK for **[AgentX](https://www.agentx.so/)** - an evaluation, tracing, and monitoring framework for AI agents, plus a client for AgentX's own hosted agents.
8
8
 
9
9
  Also see [SDK Developer Docs](https://developers.agentx.so), [API Reference Docs](https://docs.agentx.so/reference)
10
10
 
@@ -45,19 +45,10 @@ pip install --upgrade agentx-python
45
45
 
46
46
  Requires Python 3.9 or newer.
47
47
 
48
- ---
49
-
50
- ## Authentication
51
-
52
- Get your API key at [app.agentx.so](https://app.agentx.so), then either pass it inline or expose it as an environment variable.
53
-
54
- ```python
55
- # Option A - pass the key inline
56
- from agentx import AgentX
57
- client = AgentX(api_key="your-api-key-here")
48
+ #### Run self host eval framework locally
58
49
 
59
- # Option B - set AGENTX_API_KEY in your environment, then:
60
- client = AgentX.from_env()
50
+ ```
51
+ agentx-trace-eval --dev --update
61
52
  ```
62
53
 
63
54
  ---
@@ -115,7 +106,7 @@ print(report.recommendations) # list of prioritized, actionable fixes
115
106
 
116
107
  Ask a case's question several extra ways each run, LLM-paraphrased server-side, to catch agents that break on phrasing rather than substance, and override the judge's prompt/model per config, see [Smoke testing](EVALUATIONS.md#smoke-testing-phrasing-robustness) and [Configuring the judge](EVALUATIONS.md#configuring-the-judge) in the full guide.
117
108
 
118
- Since AgentX doesn't own your agent's code, `client.evaluations.prompts` lets AgentX become your prompt's *source of truth* instead - the same problem LangSmith's Prompt Hub and Langfuse's Prompt Management solve. Pull a version at runtime, tag your eval runs (or live traces) with it, and let a judge propose a rewrite from your real worst-rated results - a human always has to approve before it publishes:
109
+ Since AgentX doesn't own your agent's code, `client.evaluations.prompts` lets AgentX become your prompt's _source of truth_ instead - the same problem LangSmith's Prompt Hub and Langfuse's Prompt Management solve. Pull a version at runtime, tag your eval runs (or live traces) with it, and let a judge propose a rewrite from your real worst-rated results - a human always has to approve before it publishes:
119
110
 
120
111
  ```python
121
112
  prompt = client.evaluations.prompts.get("support-agent-system-prompt") # or prompt.id
@@ -129,7 +120,7 @@ client.evaluations.run(
129
120
 
130
121
  See [Prompt registry](EVALUATIONS.md#prompt-registry) in the full guide, or [self-host's docs](https://docs.agentx.so/improve/prompt-management) for the "Suggest improvement" dashboard flow (self-host only - no hosted-SaaS equivalent yet).
131
122
 
132
- On self-host, a finalized run can also **gate a CI job**: `report.gate(fail_under=7, no_regression=True)` checks the run's average rating against an absolute floor and/or the dataset's previous run, prints per-check verdicts into the CI log, and returns an exit code - `sys.exit(gate.exit_code)` blocks the merge on regression. Recorded gates appear in the dashboard's CI Gates tab. See [self-host's CI docs](https://docs.agentx.so/integrations/self-host-ci) for the GitHub Actions recipe.
123
+ On self-host, a finalized run can also **gate a CI job**: `run.gate(fail_under=7, no_regression=True)` (on the run context `.execute()` returns) checks the run's average rating against an absolute floor and/or the dataset's previous run, prints per-check verdicts into the CI log, and returns an exit code - `sys.exit(gate.exit_code)` blocks the merge on regression. Recorded gates appear in the dashboard's CI Gates tab. See [self-host's CI docs](https://docs.agentx.so/integrations/self-host-ci) for the GitHub Actions recipe.
133
124
 
134
125
  See **[EVALUATIONS.md](EVALUATIONS.md)** for the full guide - dataset builder, framework adapters, similarity metrics, smoke testing, judge configuration, prompt registry, and the complete API reference.
135
126
 
@@ -184,18 +175,18 @@ regular totals, no extra config needed. Self-host's cost estimate prices these s
184
175
  regular input token when you've set optional cache rates on that model. Install the matching
185
176
  extra:
186
177
 
187
- | Framework | Install | Integration |
188
- | --------------------- | -------------------------------------------- | ------------------------ |
189
- | LangChain | `pip install "agentx-python[langchain]"` | `AgentXCallbackHandler` |
190
- | CrewAI | `pip install "agentx-python[crewai]"` | `AgentXCrewObserver` |
191
- | OpenAI Agents SDK | `pip install "agentx-python[openai-agents]"` | `AgentXTracingProcessor` |
192
- | OpenAI (raw client) | `pip install "agentx-python[openai]"` | `patch_openai_client` |
193
- | Anthropic | `pip install "agentx-python[anthropic]"` | `patch_anthropic_client` |
194
- | Google ADK | `pip install "agentx-python[google-adk]"` | `AgentXADKPlugin` |
195
- | Google GenAI (Gemini) | `pip install "agentx-python[google-genai]"` | `patch_genai_client` |
196
- | LiteLLM | `pip install "agentx-python[litellm]"` | `AgentXLiteLLMLogger` |
197
- | LlamaIndex | `pip install "agentx-python[llamaindex]"` | `AgentXLlamaIndexHandler`|
198
- | AutoGen | `pip install "agentx-python[autogen]"` | `AgentXAutoGenObserver` |
178
+ | Framework | Install | Integration |
179
+ | --------------------- | -------------------------------------------- | ------------------------- |
180
+ | LangChain | `pip install "agentx-python[langchain]"` | `AgentXCallbackHandler` |
181
+ | CrewAI | `pip install "agentx-python[crewai]"` | `AgentXCrewObserver` |
182
+ | OpenAI Agents SDK | `pip install "agentx-python[openai-agents]"` | `AgentXTracingProcessor` |
183
+ | OpenAI (raw client) | `pip install "agentx-python[openai]"` | `patch_openai_client` |
184
+ | Anthropic | `pip install "agentx-python[anthropic]"` | `patch_anthropic_client` |
185
+ | Google ADK | `pip install "agentx-python[google-adk]"` | `AgentXADKPlugin` |
186
+ | Google GenAI (Gemini) | `pip install "agentx-python[google-genai]"` | `patch_genai_client` |
187
+ | LiteLLM | `pip install "agentx-python[litellm]"` | `AgentXLiteLLMLogger` |
188
+ | LlamaIndex | `pip install "agentx-python[llamaindex]"` | `AgentXLlamaIndexHandler` |
189
+ | AutoGen | `pip install "agentx-python[autogen]"` | `AgentXAutoGenObserver` |
199
190
 
200
191
  Or plain Python - wrap any function with `@tracer.trace(...)` and it just works, no framework required.
201
192
 
@@ -37,6 +37,9 @@ _RETRY_BACKOFF = [1.0, 2.0, 4.0]
37
37
  # wait out the whole job on one connection. Matches EvaluationRunContext.analyze()'s own
38
38
  # default timeout.
39
39
  _SELF_HOST_ANALYZE_TIMEOUT = 1800
40
+ # Batch result submission scores each result synchronously inside the request (one judge call
41
+ # per result on the sync path) - a big batch on a slow judge legitimately takes minutes.
42
+ _SELF_HOST_SCORING_TIMEOUT = 900
40
43
 
41
44
 
42
45
  class AgentXEvaluationsError(Exception):
@@ -344,7 +347,15 @@ class EvaluationsClient:
344
347
  "batchId": batch_id,
345
348
  "results": [_result_to_payload(r) for r in results],
346
349
  }
347
- data = self._request("POST", f"/runs/{run_id}/results", json=payload)
350
+ # Scoring is synchronous inside this request (a judge call per result - the runner's
351
+ # spinner says "~60s+" for a reason), so the default 30s timeout + silent backoff loop
352
+ # re-POSTed the batch WHILE the first submission was still scoring, and the two raced
353
+ # into the engine's (run_id, idempotency_key) unique constraint. Long timeout, no
354
+ # transport retry - the runner's own logged retry-once is the retry layer here.
355
+ data = self._request(
356
+ "POST", f"/runs/{run_id}/results", json=payload,
357
+ timeout=_SELF_HOST_SCORING_TIMEOUT, retry=False,
358
+ )
348
359
  return BatchAppendResponse(**data)
349
360
 
350
361
  def finalize_run(self, run_id: str) -> Dict[str, Any]:
@@ -1,7 +1,7 @@
1
- VERSION = "0.8.9"
1
+ VERSION = "0.8.11"
2
2
 
3
3
  # The AgentX-trace-eval release this SDK version is tested against - what `agentx-trace-eval`
4
4
  # installs and converges to (see agentx/cli.py). Bump together with VERSION when releasing, so
5
5
  # every published SDK names a known-good engine+dashboard pair. Users can override with
6
6
  # AGENTX_TRACE_EVAL_VERSION=<tag|latest>.
7
- ENGINE_VERSION = "v0.3.1"
7
+ ENGINE_VERSION = "v0.3.2"
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: agentx-python
3
- Version: 0.8.9
3
+ Version: 0.8.11
4
4
  Summary: Official Python SDK for AgentX (https://www.agentx.so/)
5
5
  Home-page: https://github.com/AgentX-ai/AgentX-python
6
6
  Author: Robin Wang and AgentX Team
@@ -68,7 +68,7 @@ Dynamic: summary
68
68
  [![Python versions](https://img.shields.io/pypi/pyversions/agentx-python)](https://pypi.org/project/agentx-python/)
69
69
  [![License](https://img.shields.io/badge/license-Apache--2.0-blue.svg)](LICENSE)
70
70
 
71
- The official Python SDK for **[AgentX](https://app.agentx.so/)** - an evaluation, tracing, and monitoring framework for AI agents, plus a client for AgentX's own hosted agents.
71
+ The official Python SDK for **[AgentX](https://www.agentx.so/)** - an evaluation, tracing, and monitoring framework for AI agents, plus a client for AgentX's own hosted agents.
72
72
 
73
73
  Also see [SDK Developer Docs](https://developers.agentx.so), [API Reference Docs](https://docs.agentx.so/reference)
74
74
 
@@ -109,19 +109,10 @@ pip install --upgrade agentx-python
109
109
 
110
110
  Requires Python 3.9 or newer.
111
111
 
112
- ---
113
-
114
- ## Authentication
115
-
116
- Get your API key at [app.agentx.so](https://app.agentx.so), then either pass it inline or expose it as an environment variable.
117
-
118
- ```python
119
- # Option A - pass the key inline
120
- from agentx import AgentX
121
- client = AgentX(api_key="your-api-key-here")
112
+ #### Run self host eval framework locally
122
113
 
123
- # Option B - set AGENTX_API_KEY in your environment, then:
124
- client = AgentX.from_env()
114
+ ```
115
+ agentx-trace-eval --dev --update
125
116
  ```
126
117
 
127
118
  ---
@@ -179,7 +170,7 @@ print(report.recommendations) # list of prioritized, actionable fixes
179
170
 
180
171
  Ask a case's question several extra ways each run, LLM-paraphrased server-side, to catch agents that break on phrasing rather than substance, and override the judge's prompt/model per config, see [Smoke testing](EVALUATIONS.md#smoke-testing-phrasing-robustness) and [Configuring the judge](EVALUATIONS.md#configuring-the-judge) in the full guide.
181
172
 
182
- Since AgentX doesn't own your agent's code, `client.evaluations.prompts` lets AgentX become your prompt's *source of truth* instead - the same problem LangSmith's Prompt Hub and Langfuse's Prompt Management solve. Pull a version at runtime, tag your eval runs (or live traces) with it, and let a judge propose a rewrite from your real worst-rated results - a human always has to approve before it publishes:
173
+ Since AgentX doesn't own your agent's code, `client.evaluations.prompts` lets AgentX become your prompt's _source of truth_ instead - the same problem LangSmith's Prompt Hub and Langfuse's Prompt Management solve. Pull a version at runtime, tag your eval runs (or live traces) with it, and let a judge propose a rewrite from your real worst-rated results - a human always has to approve before it publishes:
183
174
 
184
175
  ```python
185
176
  prompt = client.evaluations.prompts.get("support-agent-system-prompt") # or prompt.id
@@ -193,7 +184,7 @@ client.evaluations.run(
193
184
 
194
185
  See [Prompt registry](EVALUATIONS.md#prompt-registry) in the full guide, or [self-host's docs](https://docs.agentx.so/improve/prompt-management) for the "Suggest improvement" dashboard flow (self-host only - no hosted-SaaS equivalent yet).
195
186
 
196
- On self-host, a finalized run can also **gate a CI job**: `report.gate(fail_under=7, no_regression=True)` checks the run's average rating against an absolute floor and/or the dataset's previous run, prints per-check verdicts into the CI log, and returns an exit code - `sys.exit(gate.exit_code)` blocks the merge on regression. Recorded gates appear in the dashboard's CI Gates tab. See [self-host's CI docs](https://docs.agentx.so/integrations/self-host-ci) for the GitHub Actions recipe.
187
+ On self-host, a finalized run can also **gate a CI job**: `run.gate(fail_under=7, no_regression=True)` (on the run context `.execute()` returns) checks the run's average rating against an absolute floor and/or the dataset's previous run, prints per-check verdicts into the CI log, and returns an exit code - `sys.exit(gate.exit_code)` blocks the merge on regression. Recorded gates appear in the dashboard's CI Gates tab. See [self-host's CI docs](https://docs.agentx.so/integrations/self-host-ci) for the GitHub Actions recipe.
197
188
 
198
189
  See **[EVALUATIONS.md](EVALUATIONS.md)** for the full guide - dataset builder, framework adapters, similarity metrics, smoke testing, judge configuration, prompt registry, and the complete API reference.
199
190
 
@@ -248,18 +239,18 @@ regular totals, no extra config needed. Self-host's cost estimate prices these s
248
239
  regular input token when you've set optional cache rates on that model. Install the matching
249
240
  extra:
250
241
 
251
- | Framework | Install | Integration |
252
- | --------------------- | -------------------------------------------- | ------------------------ |
253
- | LangChain | `pip install "agentx-python[langchain]"` | `AgentXCallbackHandler` |
254
- | CrewAI | `pip install "agentx-python[crewai]"` | `AgentXCrewObserver` |
255
- | OpenAI Agents SDK | `pip install "agentx-python[openai-agents]"` | `AgentXTracingProcessor` |
256
- | OpenAI (raw client) | `pip install "agentx-python[openai]"` | `patch_openai_client` |
257
- | Anthropic | `pip install "agentx-python[anthropic]"` | `patch_anthropic_client` |
258
- | Google ADK | `pip install "agentx-python[google-adk]"` | `AgentXADKPlugin` |
259
- | Google GenAI (Gemini) | `pip install "agentx-python[google-genai]"` | `patch_genai_client` |
260
- | LiteLLM | `pip install "agentx-python[litellm]"` | `AgentXLiteLLMLogger` |
261
- | LlamaIndex | `pip install "agentx-python[llamaindex]"` | `AgentXLlamaIndexHandler`|
262
- | AutoGen | `pip install "agentx-python[autogen]"` | `AgentXAutoGenObserver` |
242
+ | Framework | Install | Integration |
243
+ | --------------------- | -------------------------------------------- | ------------------------- |
244
+ | LangChain | `pip install "agentx-python[langchain]"` | `AgentXCallbackHandler` |
245
+ | CrewAI | `pip install "agentx-python[crewai]"` | `AgentXCrewObserver` |
246
+ | OpenAI Agents SDK | `pip install "agentx-python[openai-agents]"` | `AgentXTracingProcessor` |
247
+ | OpenAI (raw client) | `pip install "agentx-python[openai]"` | `patch_openai_client` |
248
+ | Anthropic | `pip install "agentx-python[anthropic]"` | `patch_anthropic_client` |
249
+ | Google ADK | `pip install "agentx-python[google-adk]"` | `AgentXADKPlugin` |
250
+ | Google GenAI (Gemini) | `pip install "agentx-python[google-genai]"` | `patch_genai_client` |
251
+ | LiteLLM | `pip install "agentx-python[litellm]"` | `AgentXLiteLLMLogger` |
252
+ | LlamaIndex | `pip install "agentx-python[llamaindex]"` | `AgentXLlamaIndexHandler` |
253
+ | AutoGen | `pip install "agentx-python[autogen]"` | `AgentXAutoGenObserver` |
263
254
 
264
255
  Or plain Python - wrap any function with `@tracer.trace(...)` and it just works, no framework required.
265
256
 
File without changes
File without changes
File without changes