agentx-python 0.6.13__tar.gz → 0.6.14__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (66) hide show
  1. {agentx_python-0.6.13/agentx_python.egg-info → agentx_python-0.6.14}/PKG-INFO +84 -109
  2. {agentx_python-0.6.13 → agentx_python-0.6.14}/README.md +83 -108
  3. agentx_python-0.6.14/agentx/version.py +1 -0
  4. {agentx_python-0.6.13 → agentx_python-0.6.14/agentx_python.egg-info}/PKG-INFO +84 -109
  5. agentx_python-0.6.13/agentx/version.py +0 -1
  6. {agentx_python-0.6.13 → agentx_python-0.6.14}/LICENSE +0 -0
  7. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/__init__.py +0 -0
  8. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/agentx.py +0 -0
  9. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/cli.py +0 -0
  10. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/evaluations/__init__.py +0 -0
  11. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/evaluations/_term.py +0 -0
  12. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/evaluations/adapters/__init__.py +0 -0
  13. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/evaluations/adapters/http_endpoint.py +0 -0
  14. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/evaluations/adapters/precomputed.py +0 -0
  15. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/evaluations/adapters/raw.py +0 -0
  16. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/evaluations/client.py +0 -0
  17. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/evaluations/datasets.py +0 -0
  18. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/evaluations/evaluation_settings.py +0 -0
  19. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/evaluations/models.py +0 -0
  20. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/evaluations/prompts.py +0 -0
  21. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/evaluations/redaction.py +0 -0
  22. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/evaluations/reporting.py +0 -0
  23. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/evaluations/results.py +0 -0
  24. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/evaluations/runner.py +0 -0
  25. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/evaluations/tracing.py +0 -0
  26. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/exceptions.py +0 -0
  27. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/integrations/__init__.py +0 -0
  28. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/integrations/_traced_call.py +0 -0
  29. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/integrations/anthropic.py +0 -0
  30. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/integrations/autogen.py +0 -0
  31. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/integrations/crewai.py +0 -0
  32. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/integrations/google_adk.py +0 -0
  33. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/integrations/google_genai.py +0 -0
  34. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/integrations/langchain.py +0 -0
  35. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/integrations/litellm.py +0 -0
  36. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/integrations/llamaindex.py +0 -0
  37. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/integrations/openai.py +0 -0
  38. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/integrations/openai_agents.py +0 -0
  39. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/monitor/__init__.py +0 -0
  40. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/monitor/client.py +0 -0
  41. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/monitor/models.py +0 -0
  42. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/monitor/online_evaluators.py +0 -0
  43. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/monitor/patterns.py +0 -0
  44. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/monitor/profile.py +0 -0
  45. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/monitor/signals.py +0 -0
  46. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/py.typed +0 -0
  47. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/resources/__init__.py +0 -0
  48. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/resources/agent.py +0 -0
  49. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/resources/conversation.py +0 -0
  50. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/resources/workforce.py +0 -0
  51. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/tracing/__init__.py +0 -0
  52. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/tracing/ci_types.py +0 -0
  53. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/tracing/ingest_client.py +0 -0
  54. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/tracing/tracer.py +0 -0
  55. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx/util.py +0 -0
  56. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx_python.egg-info/SOURCES.txt +0 -0
  57. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx_python.egg-info/dependency_links.txt +0 -0
  58. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx_python.egg-info/entry_points.txt +0 -0
  59. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx_python.egg-info/not-zip-safe +0 -0
  60. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx_python.egg-info/requires.txt +0 -0
  61. {agentx_python-0.6.13 → agentx_python-0.6.14}/agentx_python.egg-info/top_level.txt +0 -0
  62. {agentx_python-0.6.13 → agentx_python-0.6.14}/setup.cfg +0 -0
  63. {agentx_python-0.6.13 → agentx_python-0.6.14}/setup.py +0 -0
  64. {agentx_python-0.6.13 → agentx_python-0.6.14}/tests/test_integration.py +0 -0
  65. {agentx_python-0.6.13 → agentx_python-0.6.14}/tests/test_integrations.py +0 -0
  66. {agentx_python-0.6.13 → agentx_python-0.6.14}/tests/test_span_tree.py +0 -0
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: agentx-python
3
- Version: 0.6.13
3
+ Version: 0.6.14
4
4
  Summary: Official Python SDK for AgentX (https://www.agentx.so/)
5
5
  Home-page: https://github.com/AgentX-ai/AgentX-python
6
6
  Author: Robin Wang and AgentX Team
@@ -66,7 +66,7 @@ Dynamic: summary
66
66
  [![Python versions](https://img.shields.io/pypi/pyversions/agentx-python)](https://pypi.org/project/agentx-python/)
67
67
  [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
68
68
 
69
- The official Python SDK for **[AgentX](https://app.agentx.so/)** — build, chat with, orchestrate, and trace AI agents in a few lines of code.
69
+ The official Python SDK for **[AgentX](https://app.agentx.so/)** — an evaluation, tracing, and monitoring framework for AI agents, plus a client for AgentX's own hosted agents.
70
70
 
71
71
  Also see [SDK Developer Docs](https://developers.agentx.so), [API Reference Docs](https://docs.agentx.so/reference)
72
72
 
@@ -78,30 +78,24 @@ Also see [SDK Developer Docs](https://developers.agentx.so), [API Reference Docs
78
78
  - [Installation](#installation)
79
79
  - [Authentication](#authentication)
80
80
  - [Quick start](#quick-start)
81
- - [Working with agents](#working-with-agents)
82
- - [List agents](#list-agents)
83
- - [Start a conversation](#start-a-conversation)
84
- - [Chat (streaming and non-streaming)](#chat-streaming-and-non-streaming)
85
- - [Workforce (multi-agent orchestration)](#workforce-multi-agent-orchestration) — teams of agents with a designated manager
81
+ - [Custom agent evaluations](#custom-agent-evaluations) — LLM-as-a-judge, cosine / Jaccard similarity, any framework
86
82
  - [Production tracing](#production-tracing) — record live agent runs from any framework
87
83
  - [Monitor](#monitor) — automatic production monitoring, patterns and signals
88
- - [Custom agent evaluations](#custom-agent-evaluations) — LLM-as-a-judge, cosine / Jaccard similarity
89
84
  - [Self-host](#self-host) — run Trace/Evaluate/Monitor on your own machine instead of the hosted dashboard
85
+ - [Agents & conversations](#agents--conversations) — chat with and orchestrate AgentX's own hosted agents
90
86
  - [Links](#links)
91
87
 
92
88
  ---
93
89
 
94
90
  ## Why AgentX
95
91
 
96
- - **Simple mental model** — `Agent Conversation Message`.
97
- - **Chain-of-thought** is built in, no extra plumbing.
98
- - **Bring any LLM** — works across major open and closed-source vendors.
99
- - **Batteries included** — voice (ASR/TTS), image generation, document/CSV/Excel/OCR, RAG with built-in re-ranking.
100
- - **MCP support** — connect any Model Context Protocol server.
101
- - **Multi-agent orchestration** — workforces of agents with a designated manager, across LLM vendors.
102
- - **Production tracing** — one decorator or context manager records every agent run (input, output, latency, tool calls, token usage) into your workspace, for any framework.
103
- - **Agent Evaluations** — score any agent (LangChain, CrewAI, OpenAI, Anthropic, HTTP, …) with LLM-as-a-judge ratings plus optional cosine and Jaccard similarity metrics. Configurable judge prompt/model (OpenAI or Anthropic), per-question judge guidelines, and smoke testing for phrasing robustness.
104
- - **A2A** — Each agent can be published with agent-to-agent protocol compatible.
92
+ - **Agent Evaluations** — score **any** agent (LangChain, CrewAI, AutoGen, LlamaIndex, OpenAI, Anthropic, HTTP, or plain Python) against a dataset with LLM-as-a-judge ratings plus optional cosine, Jaccard, and BLEU/ROUGE similarity metrics. Configurable judge prompt/model, per-question judge guidelines, smoke testing for phrasing robustness, and a durable multi-judge analysis pass for a qualitative report.
93
+ - **Production tracing** one decorator or context manager records every agent run (input, output, latency, tool calls, token usage, and a real span tree) into your workspace, for any framework.
94
+ - **Monitor** — check live traces against detection patterns or a sampled LLM judge, and read back triage-ready signals, no dashboard setup required.
95
+ - **Prompt registry** — make AgentX the source of truth for your own agent's prompts: pull a version at runtime, tag eval runs and live traces with it, let a judge propose a rewrite from your worst-rated results.
96
+ - **Self-host** — run the whole Trace/Evaluate/Monitor stack locally, bring your own LLM keys, no account required.
97
+ - **Bring any LLM** — works across major open and closed-source vendors, for evaluation, tracing, and AgentX's own hosted agents alike.
98
+ - **AgentX's own hosted agents** — a simple `Agent Conversation Message` mental model, chain-of-thought built in, multi-agent workforces, MCP support, and A2A publishing, for when you want AgentX to run the agent too, not just evaluate/trace/monitor it.
105
99
 
106
100
  ---
107
101
 
@@ -132,80 +126,72 @@ client = AgentX.from_env()
132
126
 
133
127
  ## Quick start
134
128
 
129
+ Evaluate your own agent — any framework, or plain Python — against a dataset:
130
+
135
131
  ```python
136
132
  from agentx import AgentX
137
133
 
138
134
  client = AgentX.from_env()
139
135
 
140
- # Pick an existing agent and chat with it
141
- agent = client.list_agents()[0]
142
- conversation = agent.new_conversation()
143
- print(conversation.chat("Hello! What can you help me with?"))
136
+ def my_agent(case):
137
+ return call_my_agent(case.query) # your agent's own code, any framework
138
+
139
+ report = (
140
+ client.evaluations
141
+ .run(dataset_id="evds_…", subject={"kind": "custom_agent", "framework": "raw_python"})
142
+ .execute(my_agent)
143
+ .finalize()
144
+ .analyze()
145
+ )
146
+
147
+ print(report.average_rating) # LLM-graded score, 0–10
148
+ print(report.summary) # AI-generated narrative from .analyze()
144
149
  ```
145
150
 
146
- That's it. The remaining sections show the same primitives in more detail.
151
+ That's it. The rest of this section covers building the dataset, framework adapters, similarity metrics, and judge configuration.
147
152
 
148
153
  ---
149
154
 
150
- ## Working with agents
155
+ ## Custom agent evaluations
151
156
 
152
- ### List agents
157
+ Evaluate **any** AI agent — LangChain, CrewAI, AutoGen, LlamaIndex, OpenAI, Anthropic, HTTP endpoints, or plain Python — using AgentX as the scoring and reporting backend. Includes optional **cosine**, **Jaccard**, and **BLEU/ROUGE** similarity metrics alongside LLM-graded ratings.
153
158
 
154
159
  ```python
155
- agents = client.list_agents()
156
- print(f"You have {len(agents)} agents")
157
- ```
160
+ report = (
161
+ client.evaluations
162
+ .run(dataset_id="evds_…", subject={"kind": "custom_agent", "framework": "raw_python"})
163
+ .execute(my_agent_fn)
164
+ .finalize()
165
+ .analyze()
166
+ )
158
167
 
159
- ### Start a conversation
168
+ print(report.average_rating) # LLM-graded score, 0–10
169
+ print(report.cosine_similarity) # embedding cosine, 0–1 (None if not enabled)
170
+ print(report.jaccard_similarity) # token-set overlap, 0–1 (None if not enabled)
160
171
 
161
- ```python
162
- agent = client.get_agent(id="<agent-id>")
172
+ print(report.summary) # AI-generated narrative from .analyze()
173
+ print(report.recommendations) # list of prioritized, actionable fixes
174
+ ```
163
175
 
164
- # Either resume an existing conversation…
165
- existing = agent.list_conversations()
166
- last = existing[-1]
167
- for msg in last.list_messages():
168
- print(msg)
176
+ `.analyze()` also generates a full qualitative report (strengths, weaknesses, instruction adherence, reasoning quality, and recommendations), running the same durable, multi-judge pipeline as the dashboard's "Analyze" button. `analyze(mode=..., quality_mode=..., judges=[...])` controls how items are scored and by which models. See [AI analysis report](EVALUATIONS.md#ai-analysis-report) in the full guide for the complete field and parameter reference.
169
177
 
170
- # …or start a fresh one
171
- conversation = agent.new_conversation()
172
- ```
178
+ Ask a case's question several extra ways each run, LLM-paraphrased server-side, to catch agents that break on phrasing rather than substance, and override the judge's prompt/model per config, see [Smoke testing](EVALUATIONS.md#smoke-testing-phrasing-robustness) and [Configuring the judge](EVALUATIONS.md#configuring-the-judge) in the full guide.
173
179
 
174
- ### Chat (streaming and non-streaming)
180
+ Since AgentX doesn't own your agent's code, `client.evaluations.prompts` lets AgentX become your prompt's *source of truth* instead — the same problem LangSmith's Prompt Hub and Langfuse's Prompt Management solve. Pull a version at runtime, tag your eval runs (or live traces) with it, and let a judge propose a rewrite from your real worst-rated results — a human always has to approve before it publishes:
175
181
 
176
182
  ```python
177
- # Blocking returns the full response once it's ready
178
- response = conversation.chat("What is your name?")
179
- print(response)
183
+ prompt = client.evaluations.prompts.get("support-agent-system-prompt") # or prompt.id
184
+ # use prompt.text as your own agent's system prompt
180
185
 
181
- # Streaming — yields ChatResponse objects as the model produces them
182
- for chunk in conversation.chat_stream("Hello, what is your name?"):
183
- if chunk.text:
184
- print(chunk.text, end="")
186
+ client.evaluations.run(
187
+ dataset_id="evds_…",
188
+ subject={"kind": "custom_agent", "metadata": {"promptName": prompt.name}},
189
+ ).execute(my_agent_fn)
185
190
  ```
186
191
 
187
- Each `ChatResponse` chunk exposes the agent's `text` and, where applicable, its `cot` (chain-of-thought) reasoning, along with any retrieved references and tasks.
188
-
189
- ---
190
-
191
- ## Workforce (multi-agent orchestration)
192
-
193
- A **workforce** is a team of agents coordinated by a designated manager agent. Workforces can mix LLM vendors and route work between specialists.
194
-
195
- ```python
196
- workforces = client.list_workforces()
197
- workforce = workforces[0]
198
-
199
- print(f"Workforce: {workforce.name}")
200
- print(f"Manager: {workforce.manager.name}")
201
- print(f"Agents: {[a.name for a in workforce.agents]}")
192
+ See [Prompt registry](EVALUATIONS.md#prompt-registry) in the full guide, or [self-host's docs](https://docs.agentx.so/self-host#prompt-registry) for the "Suggest improvement" dashboard flow (self-host only no hosted-SaaS equivalent yet).
202
193
 
203
- # Chat with the workforcethe manager decides which agent(s) to delegate to
204
- conversation = workforce.new_conversation()
205
- for chunk in workforce.chat_stream(conversation.id, "How can you help me with this project?"):
206
- if chunk.text:
207
- print(chunk.text, end="")
208
- ```
194
+ See **[EVALUATIONS.md](EVALUATIONS.md)** for the full guide dataset builder, framework adapters, similarity metrics, smoke testing, judge configuration, prompt registry, and the complete API reference.
209
195
 
210
196
  ---
211
197
 
@@ -315,66 +301,55 @@ See **[TRACING.md](TRACING.md)** for the complete Monitor guide.
315
301
 
316
302
  ---
317
303
 
318
- ## Custom agent evaluations
319
-
320
- Evaluate **any** AI agent — LangChain, CrewAI, AutoGen, LlamaIndex, OpenAI, Anthropic, HTTP endpoints, or plain Python — using AgentX as the scoring and reporting backend. Includes optional **cosine** and **Jaccard** similarity metrics alongside LLM-graded ratings.
321
-
322
- ```python
323
- report = (
324
- client.evaluations
325
- .run(dataset_id="evds_…", subject={"kind": "custom_agent", "framework": "raw_python"})
326
- .execute(my_agent_fn)
327
- .finalize()
328
- .analyze()
329
- )
304
+ ## Self-host
330
305
 
331
- print(report.average_rating) # LLM-graded score, 0–10
332
- print(report.cosine_similarity) # embedding cosine, 0–1 (None if not enabled)
333
- print(report.jaccard_similarity) # token-set overlap, 0–1 (None if not enabled)
306
+ Prefer to run Trace/Evaluate/Monitor locally instead of the hosted dashboard — no account, bring your own LLM keys? This SDK ships a launcher for [AgentX-trace-eval](https://github.com/AgentX-ai/AgentX-trace-eval), a separate, portable governance engine:
334
307
 
335
- print(report.summary) # AI-generated narrative from .analyze()
336
- print(report.recommendations) # list of prioritized, actionable fixes
308
+ ```bash
309
+ agentx-trace-eval --dev
337
310
  ```
338
311
 
339
- `.analyze()` also generates a full qualitative report (strengths, weaknesses, instruction adherence, reasoning quality, and recommendations), running the same durable, multi-judge pipeline as the dashboard's "Analyze" button. `analyze(mode=..., quality_mode=..., judges=[...])` controls how items are scored and by which models. See [AI analysis report](EVALUATIONS.md#ai-analysis-report) in the full guide for the complete field and parameter reference.
312
+ The first run downloads the engine (and dashboard) into `~/.agentx/bin` and prints a local API key; every run after that just starts it. Point this SDK at it instead of the hosted API:
340
313
 
341
- Ask a case's question several extra ways each run, LLM-paraphrased server-side, to catch agents that break on phrasing rather than substance, and override the judge's prompt/model per config, see [Smoke testing](EVALUATIONS.md#smoke-testing-phrasing-robustness) and [Configuring the judge](EVALUATIONS.md#configuring-the-judge) in the full guide.
314
+ ```bash
315
+ export AGENTX_API_BASE_URL=http://localhost:4700/api/v1
316
+ export AGENTX_API_KEY=<printed by agentx-trace-eval on first run>
317
+ ```
342
318
 
343
- Since AgentX doesn't own your agent's code, `client.evaluations.prompts` lets AgentX become your prompt's *source of truth* instead the same problem LangSmith's Prompt Hub and Langfuse's Prompt Management solve. Pull a version at runtime, tag your eval runs (or live traces) with it, and let a judge propose a rewrite from your real worst-rated results a human always has to approve before it publishes:
319
+ `agentx-trace-eval` isn't this SDK's own code the engine itself is a separate, compiled binary, downloaded on demand rather than bundled into this package, so installing `agentx-python` doesn't get any heavier for the (much more common) case of just talking to the hosted AgentX API. See that repo's README for what's included, and `AGENTX_INSTALL_DIR`/`AGENTX_TRACE_EVAL_VERSION`/`AGENTX_TRACE_EVAL_SKIP_WEB` env vars to control where/what it installs.
344
320
 
345
- ```python
346
- prompt = client.evaluations.prompts.get("support-agent-system-prompt") # or prompt.id
347
- # use prompt.text as your own agent's system prompt
321
+ ---
348
322
 
349
- client.evaluations.run(
350
- dataset_id="evds_…",
351
- subject={"kind": "custom_agent", "metadata": {"promptName": prompt.name}},
352
- ).execute(my_agent_fn)
353
- ```
323
+ ## Agents & conversations
354
324
 
355
- See [Prompt registry](EVALUATIONS.md#prompt-registry) in the full guide, or [self-host's docs](https://docs.agentx.so/self-host#prompt-registry) for the "Suggest improvement" dashboard flow (self-host only no hosted-SaaS equivalent yet).
325
+ Beyond evaluation, tracing, and monitoring, this SDK is also a client for AgentX's own hosted agents build, chat with, and orchestrate them directly.
356
326
 
357
- See **[EVALUATIONS.md](EVALUATIONS.md)** for the full guide — dataset builder, framework adapters, similarity metrics, smoke testing, judge configuration, prompt registry, and the complete API reference.
327
+ ```python
328
+ agent = client.list_agents()[0]
329
+ conversation = agent.new_conversation()
358
330
 
359
- ---
331
+ # Blocking — returns the full response once it's ready
332
+ print(conversation.chat("What can you help me with?"))
360
333
 
361
- ## Self-host
334
+ # Streaming — yields ChatResponse objects as the model produces them
335
+ for chunk in conversation.chat_stream("Hello!"):
336
+ if chunk.text:
337
+ print(chunk.text, end="")
338
+ ```
362
339
 
363
- Prefer to run Trace/Evaluate/Monitor locally instead of the hosted dashboard no account, bring your own LLM keys? This SDK ships a launcher for [AgentX-trace-eval](https://github.com/AgentX-ai/AgentX-trace-eval), a separate, portable governance engine:
340
+ Each `ChatResponse` chunk exposes the agent's `text` and, where applicable, its `cot` (chain-of-thought) reasoning, along with any retrieved references and tasks. `agent.list_conversations()` / `conversation.list_messages()` resume history instead of starting fresh.
364
341
 
365
- ```bash
366
- agentx-trace-eval --dev
367
- ```
342
+ A **workforce** is a team of agents coordinated by a designated manager, mixing LLM vendors and routing work between specialists:
368
343
 
369
- The first run downloads the engine (and dashboard) into `~/.agentx/bin` and prints a local API key; every run after that just starts it. Point this SDK at it instead of the hosted API:
344
+ ```python
345
+ workforce = client.list_workforces()[0]
346
+ conversation = workforce.new_conversation()
370
347
 
371
- ```bash
372
- export AGENTX_API_BASE_URL=http://localhost:4700/api/v1
373
- export AGENTX_API_KEY=<printed by agentx-trace-eval on first run>
348
+ for chunk in workforce.chat_stream(conversation.id, "How can you help me with this project?"):
349
+ if chunk.text:
350
+ print(chunk.text, end="")
374
351
  ```
375
352
 
376
- `agentx-trace-eval` isn't this SDK's own code — the engine itself is a separate, compiled binary, downloaded on demand rather than bundled into this package, so installing `agentx-python` doesn't get any heavier for the (much more common) case of just talking to the hosted AgentX API. See that repo's README for what's included, and `AGENTX_INSTALL_DIR`/`AGENTX_TRACE_EVAL_VERSION`/`AGENTX_TRACE_EVAL_SKIP_WEB` env vars to control where/what it installs.
377
-
378
353
  ---
379
354
 
380
355
  ## Links
@@ -4,7 +4,7 @@
4
4
  [![Python versions](https://img.shields.io/pypi/pyversions/agentx-python)](https://pypi.org/project/agentx-python/)
5
5
  [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
6
6
 
7
- The official Python SDK for **[AgentX](https://app.agentx.so/)** — build, chat with, orchestrate, and trace AI agents in a few lines of code.
7
+ The official Python SDK for **[AgentX](https://app.agentx.so/)** — an evaluation, tracing, and monitoring framework for AI agents, plus a client for AgentX's own hosted agents.
8
8
 
9
9
  Also see [SDK Developer Docs](https://developers.agentx.so), [API Reference Docs](https://docs.agentx.so/reference)
10
10
 
@@ -16,30 +16,24 @@ Also see [SDK Developer Docs](https://developers.agentx.so), [API Reference Docs
16
16
  - [Installation](#installation)
17
17
  - [Authentication](#authentication)
18
18
  - [Quick start](#quick-start)
19
- - [Working with agents](#working-with-agents)
20
- - [List agents](#list-agents)
21
- - [Start a conversation](#start-a-conversation)
22
- - [Chat (streaming and non-streaming)](#chat-streaming-and-non-streaming)
23
- - [Workforce (multi-agent orchestration)](#workforce-multi-agent-orchestration) — teams of agents with a designated manager
19
+ - [Custom agent evaluations](#custom-agent-evaluations) — LLM-as-a-judge, cosine / Jaccard similarity, any framework
24
20
  - [Production tracing](#production-tracing) — record live agent runs from any framework
25
21
  - [Monitor](#monitor) — automatic production monitoring, patterns and signals
26
- - [Custom agent evaluations](#custom-agent-evaluations) — LLM-as-a-judge, cosine / Jaccard similarity
27
22
  - [Self-host](#self-host) — run Trace/Evaluate/Monitor on your own machine instead of the hosted dashboard
23
+ - [Agents & conversations](#agents--conversations) — chat with and orchestrate AgentX's own hosted agents
28
24
  - [Links](#links)
29
25
 
30
26
  ---
31
27
 
32
28
  ## Why AgentX
33
29
 
34
- - **Simple mental model** — `Agent Conversation Message`.
35
- - **Chain-of-thought** is built in, no extra plumbing.
36
- - **Bring any LLM** — works across major open and closed-source vendors.
37
- - **Batteries included** — voice (ASR/TTS), image generation, document/CSV/Excel/OCR, RAG with built-in re-ranking.
38
- - **MCP support** — connect any Model Context Protocol server.
39
- - **Multi-agent orchestration** — workforces of agents with a designated manager, across LLM vendors.
40
- - **Production tracing** — one decorator or context manager records every agent run (input, output, latency, tool calls, token usage) into your workspace, for any framework.
41
- - **Agent Evaluations** — score any agent (LangChain, CrewAI, OpenAI, Anthropic, HTTP, …) with LLM-as-a-judge ratings plus optional cosine and Jaccard similarity metrics. Configurable judge prompt/model (OpenAI or Anthropic), per-question judge guidelines, and smoke testing for phrasing robustness.
42
- - **A2A** — Each agent can be published with agent-to-agent protocol compatible.
30
+ - **Agent Evaluations** — score **any** agent (LangChain, CrewAI, AutoGen, LlamaIndex, OpenAI, Anthropic, HTTP, or plain Python) against a dataset with LLM-as-a-judge ratings plus optional cosine, Jaccard, and BLEU/ROUGE similarity metrics. Configurable judge prompt/model, per-question judge guidelines, smoke testing for phrasing robustness, and a durable multi-judge analysis pass for a qualitative report.
31
+ - **Production tracing** one decorator or context manager records every agent run (input, output, latency, tool calls, token usage, and a real span tree) into your workspace, for any framework.
32
+ - **Monitor** — check live traces against detection patterns or a sampled LLM judge, and read back triage-ready signals, no dashboard setup required.
33
+ - **Prompt registry** — make AgentX the source of truth for your own agent's prompts: pull a version at runtime, tag eval runs and live traces with it, let a judge propose a rewrite from your worst-rated results.
34
+ - **Self-host** — run the whole Trace/Evaluate/Monitor stack locally, bring your own LLM keys, no account required.
35
+ - **Bring any LLM** — works across major open and closed-source vendors, for evaluation, tracing, and AgentX's own hosted agents alike.
36
+ - **AgentX's own hosted agents** — a simple `Agent Conversation Message` mental model, chain-of-thought built in, multi-agent workforces, MCP support, and A2A publishing, for when you want AgentX to run the agent too, not just evaluate/trace/monitor it.
43
37
 
44
38
  ---
45
39
 
@@ -70,80 +64,72 @@ client = AgentX.from_env()
70
64
 
71
65
  ## Quick start
72
66
 
67
+ Evaluate your own agent — any framework, or plain Python — against a dataset:
68
+
73
69
  ```python
74
70
  from agentx import AgentX
75
71
 
76
72
  client = AgentX.from_env()
77
73
 
78
- # Pick an existing agent and chat with it
79
- agent = client.list_agents()[0]
80
- conversation = agent.new_conversation()
81
- print(conversation.chat("Hello! What can you help me with?"))
74
+ def my_agent(case):
75
+ return call_my_agent(case.query) # your agent's own code, any framework
76
+
77
+ report = (
78
+ client.evaluations
79
+ .run(dataset_id="evds_…", subject={"kind": "custom_agent", "framework": "raw_python"})
80
+ .execute(my_agent)
81
+ .finalize()
82
+ .analyze()
83
+ )
84
+
85
+ print(report.average_rating) # LLM-graded score, 0–10
86
+ print(report.summary) # AI-generated narrative from .analyze()
82
87
  ```
83
88
 
84
- That's it. The remaining sections show the same primitives in more detail.
89
+ That's it. The rest of this section covers building the dataset, framework adapters, similarity metrics, and judge configuration.
85
90
 
86
91
  ---
87
92
 
88
- ## Working with agents
93
+ ## Custom agent evaluations
89
94
 
90
- ### List agents
95
+ Evaluate **any** AI agent — LangChain, CrewAI, AutoGen, LlamaIndex, OpenAI, Anthropic, HTTP endpoints, or plain Python — using AgentX as the scoring and reporting backend. Includes optional **cosine**, **Jaccard**, and **BLEU/ROUGE** similarity metrics alongside LLM-graded ratings.
91
96
 
92
97
  ```python
93
- agents = client.list_agents()
94
- print(f"You have {len(agents)} agents")
95
- ```
98
+ report = (
99
+ client.evaluations
100
+ .run(dataset_id="evds_…", subject={"kind": "custom_agent", "framework": "raw_python"})
101
+ .execute(my_agent_fn)
102
+ .finalize()
103
+ .analyze()
104
+ )
96
105
 
97
- ### Start a conversation
106
+ print(report.average_rating) # LLM-graded score, 0–10
107
+ print(report.cosine_similarity) # embedding cosine, 0–1 (None if not enabled)
108
+ print(report.jaccard_similarity) # token-set overlap, 0–1 (None if not enabled)
98
109
 
99
- ```python
100
- agent = client.get_agent(id="<agent-id>")
110
+ print(report.summary) # AI-generated narrative from .analyze()
111
+ print(report.recommendations) # list of prioritized, actionable fixes
112
+ ```
101
113
 
102
- # Either resume an existing conversation…
103
- existing = agent.list_conversations()
104
- last = existing[-1]
105
- for msg in last.list_messages():
106
- print(msg)
114
+ `.analyze()` also generates a full qualitative report (strengths, weaknesses, instruction adherence, reasoning quality, and recommendations), running the same durable, multi-judge pipeline as the dashboard's "Analyze" button. `analyze(mode=..., quality_mode=..., judges=[...])` controls how items are scored and by which models. See [AI analysis report](EVALUATIONS.md#ai-analysis-report) in the full guide for the complete field and parameter reference.
107
115
 
108
- # …or start a fresh one
109
- conversation = agent.new_conversation()
110
- ```
116
+ Ask a case's question several extra ways each run, LLM-paraphrased server-side, to catch agents that break on phrasing rather than substance, and override the judge's prompt/model per config, see [Smoke testing](EVALUATIONS.md#smoke-testing-phrasing-robustness) and [Configuring the judge](EVALUATIONS.md#configuring-the-judge) in the full guide.
111
117
 
112
- ### Chat (streaming and non-streaming)
118
+ Since AgentX doesn't own your agent's code, `client.evaluations.prompts` lets AgentX become your prompt's *source of truth* instead — the same problem LangSmith's Prompt Hub and Langfuse's Prompt Management solve. Pull a version at runtime, tag your eval runs (or live traces) with it, and let a judge propose a rewrite from your real worst-rated results — a human always has to approve before it publishes:
113
119
 
114
120
  ```python
115
- # Blocking returns the full response once it's ready
116
- response = conversation.chat("What is your name?")
117
- print(response)
121
+ prompt = client.evaluations.prompts.get("support-agent-system-prompt") # or prompt.id
122
+ # use prompt.text as your own agent's system prompt
118
123
 
119
- # Streaming — yields ChatResponse objects as the model produces them
120
- for chunk in conversation.chat_stream("Hello, what is your name?"):
121
- if chunk.text:
122
- print(chunk.text, end="")
124
+ client.evaluations.run(
125
+ dataset_id="evds_…",
126
+ subject={"kind": "custom_agent", "metadata": {"promptName": prompt.name}},
127
+ ).execute(my_agent_fn)
123
128
  ```
124
129
 
125
- Each `ChatResponse` chunk exposes the agent's `text` and, where applicable, its `cot` (chain-of-thought) reasoning, along with any retrieved references and tasks.
126
-
127
- ---
128
-
129
- ## Workforce (multi-agent orchestration)
130
-
131
- A **workforce** is a team of agents coordinated by a designated manager agent. Workforces can mix LLM vendors and route work between specialists.
132
-
133
- ```python
134
- workforces = client.list_workforces()
135
- workforce = workforces[0]
136
-
137
- print(f"Workforce: {workforce.name}")
138
- print(f"Manager: {workforce.manager.name}")
139
- print(f"Agents: {[a.name for a in workforce.agents]}")
130
+ See [Prompt registry](EVALUATIONS.md#prompt-registry) in the full guide, or [self-host's docs](https://docs.agentx.so/self-host#prompt-registry) for the "Suggest improvement" dashboard flow (self-host only no hosted-SaaS equivalent yet).
140
131
 
141
- # Chat with the workforcethe manager decides which agent(s) to delegate to
142
- conversation = workforce.new_conversation()
143
- for chunk in workforce.chat_stream(conversation.id, "How can you help me with this project?"):
144
- if chunk.text:
145
- print(chunk.text, end="")
146
- ```
132
+ See **[EVALUATIONS.md](EVALUATIONS.md)** for the full guide dataset builder, framework adapters, similarity metrics, smoke testing, judge configuration, prompt registry, and the complete API reference.
147
133
 
148
134
  ---
149
135
 
@@ -253,66 +239,55 @@ See **[TRACING.md](TRACING.md)** for the complete Monitor guide.
253
239
 
254
240
  ---
255
241
 
256
- ## Custom agent evaluations
257
-
258
- Evaluate **any** AI agent — LangChain, CrewAI, AutoGen, LlamaIndex, OpenAI, Anthropic, HTTP endpoints, or plain Python — using AgentX as the scoring and reporting backend. Includes optional **cosine** and **Jaccard** similarity metrics alongside LLM-graded ratings.
259
-
260
- ```python
261
- report = (
262
- client.evaluations
263
- .run(dataset_id="evds_…", subject={"kind": "custom_agent", "framework": "raw_python"})
264
- .execute(my_agent_fn)
265
- .finalize()
266
- .analyze()
267
- )
242
+ ## Self-host
268
243
 
269
- print(report.average_rating) # LLM-graded score, 0–10
270
- print(report.cosine_similarity) # embedding cosine, 0–1 (None if not enabled)
271
- print(report.jaccard_similarity) # token-set overlap, 0–1 (None if not enabled)
244
+ Prefer to run Trace/Evaluate/Monitor locally instead of the hosted dashboard — no account, bring your own LLM keys? This SDK ships a launcher for [AgentX-trace-eval](https://github.com/AgentX-ai/AgentX-trace-eval), a separate, portable governance engine:
272
245
 
273
- print(report.summary) # AI-generated narrative from .analyze()
274
- print(report.recommendations) # list of prioritized, actionable fixes
246
+ ```bash
247
+ agentx-trace-eval --dev
275
248
  ```
276
249
 
277
- `.analyze()` also generates a full qualitative report (strengths, weaknesses, instruction adherence, reasoning quality, and recommendations), running the same durable, multi-judge pipeline as the dashboard's "Analyze" button. `analyze(mode=..., quality_mode=..., judges=[...])` controls how items are scored and by which models. See [AI analysis report](EVALUATIONS.md#ai-analysis-report) in the full guide for the complete field and parameter reference.
250
+ The first run downloads the engine (and dashboard) into `~/.agentx/bin` and prints a local API key; every run after that just starts it. Point this SDK at it instead of the hosted API:
278
251
 
279
- Ask a case's question several extra ways each run, LLM-paraphrased server-side, to catch agents that break on phrasing rather than substance, and override the judge's prompt/model per config, see [Smoke testing](EVALUATIONS.md#smoke-testing-phrasing-robustness) and [Configuring the judge](EVALUATIONS.md#configuring-the-judge) in the full guide.
252
+ ```bash
253
+ export AGENTX_API_BASE_URL=http://localhost:4700/api/v1
254
+ export AGENTX_API_KEY=<printed by agentx-trace-eval on first run>
255
+ ```
280
256
 
281
- Since AgentX doesn't own your agent's code, `client.evaluations.prompts` lets AgentX become your prompt's *source of truth* instead the same problem LangSmith's Prompt Hub and Langfuse's Prompt Management solve. Pull a version at runtime, tag your eval runs (or live traces) with it, and let a judge propose a rewrite from your real worst-rated results a human always has to approve before it publishes:
257
+ `agentx-trace-eval` isn't this SDK's own code the engine itself is a separate, compiled binary, downloaded on demand rather than bundled into this package, so installing `agentx-python` doesn't get any heavier for the (much more common) case of just talking to the hosted AgentX API. See that repo's README for what's included, and `AGENTX_INSTALL_DIR`/`AGENTX_TRACE_EVAL_VERSION`/`AGENTX_TRACE_EVAL_SKIP_WEB` env vars to control where/what it installs.
282
258
 
283
- ```python
284
- prompt = client.evaluations.prompts.get("support-agent-system-prompt") # or prompt.id
285
- # use prompt.text as your own agent's system prompt
259
+ ---
286
260
 
287
- client.evaluations.run(
288
- dataset_id="evds_…",
289
- subject={"kind": "custom_agent", "metadata": {"promptName": prompt.name}},
290
- ).execute(my_agent_fn)
291
- ```
261
+ ## Agents & conversations
292
262
 
293
- See [Prompt registry](EVALUATIONS.md#prompt-registry) in the full guide, or [self-host's docs](https://docs.agentx.so/self-host#prompt-registry) for the "Suggest improvement" dashboard flow (self-host only no hosted-SaaS equivalent yet).
263
+ Beyond evaluation, tracing, and monitoring, this SDK is also a client for AgentX's own hosted agents build, chat with, and orchestrate them directly.
294
264
 
295
- See **[EVALUATIONS.md](EVALUATIONS.md)** for the full guide — dataset builder, framework adapters, similarity metrics, smoke testing, judge configuration, prompt registry, and the complete API reference.
265
+ ```python
266
+ agent = client.list_agents()[0]
267
+ conversation = agent.new_conversation()
296
268
 
297
- ---
269
+ # Blocking — returns the full response once it's ready
270
+ print(conversation.chat("What can you help me with?"))
298
271
 
299
- ## Self-host
272
+ # Streaming — yields ChatResponse objects as the model produces them
273
+ for chunk in conversation.chat_stream("Hello!"):
274
+ if chunk.text:
275
+ print(chunk.text, end="")
276
+ ```
300
277
 
301
- Prefer to run Trace/Evaluate/Monitor locally instead of the hosted dashboard no account, bring your own LLM keys? This SDK ships a launcher for [AgentX-trace-eval](https://github.com/AgentX-ai/AgentX-trace-eval), a separate, portable governance engine:
278
+ Each `ChatResponse` chunk exposes the agent's `text` and, where applicable, its `cot` (chain-of-thought) reasoning, along with any retrieved references and tasks. `agent.list_conversations()` / `conversation.list_messages()` resume history instead of starting fresh.
302
279
 
303
- ```bash
304
- agentx-trace-eval --dev
305
- ```
280
+ A **workforce** is a team of agents coordinated by a designated manager, mixing LLM vendors and routing work between specialists:
306
281
 
307
- The first run downloads the engine (and dashboard) into `~/.agentx/bin` and prints a local API key; every run after that just starts it. Point this SDK at it instead of the hosted API:
282
+ ```python
283
+ workforce = client.list_workforces()[0]
284
+ conversation = workforce.new_conversation()
308
285
 
309
- ```bash
310
- export AGENTX_API_BASE_URL=http://localhost:4700/api/v1
311
- export AGENTX_API_KEY=<printed by agentx-trace-eval on first run>
286
+ for chunk in workforce.chat_stream(conversation.id, "How can you help me with this project?"):
287
+ if chunk.text:
288
+ print(chunk.text, end="")
312
289
  ```
313
290
 
314
- `agentx-trace-eval` isn't this SDK's own code — the engine itself is a separate, compiled binary, downloaded on demand rather than bundled into this package, so installing `agentx-python` doesn't get any heavier for the (much more common) case of just talking to the hosted AgentX API. See that repo's README for what's included, and `AGENTX_INSTALL_DIR`/`AGENTX_TRACE_EVAL_VERSION`/`AGENTX_TRACE_EVAL_SKIP_WEB` env vars to control where/what it installs.
315
-
316
291
  ---
317
292
 
318
293
  ## Links
@@ -0,0 +1 @@
1
+ VERSION = "0.6.14"
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: agentx-python
3
- Version: 0.6.13
3
+ Version: 0.6.14
4
4
  Summary: Official Python SDK for AgentX (https://www.agentx.so/)
5
5
  Home-page: https://github.com/AgentX-ai/AgentX-python
6
6
  Author: Robin Wang and AgentX Team
@@ -66,7 +66,7 @@ Dynamic: summary
66
66
  [![Python versions](https://img.shields.io/pypi/pyversions/agentx-python)](https://pypi.org/project/agentx-python/)
67
67
  [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
68
68
 
69
- The official Python SDK for **[AgentX](https://app.agentx.so/)** — build, chat with, orchestrate, and trace AI agents in a few lines of code.
69
+ The official Python SDK for **[AgentX](https://app.agentx.so/)** — an evaluation, tracing, and monitoring framework for AI agents, plus a client for AgentX's own hosted agents.
70
70
 
71
71
  Also see [SDK Developer Docs](https://developers.agentx.so), [API Reference Docs](https://docs.agentx.so/reference)
72
72
 
@@ -78,30 +78,24 @@ Also see [SDK Developer Docs](https://developers.agentx.so), [API Reference Docs
78
78
  - [Installation](#installation)
79
79
  - [Authentication](#authentication)
80
80
  - [Quick start](#quick-start)
81
- - [Working with agents](#working-with-agents)
82
- - [List agents](#list-agents)
83
- - [Start a conversation](#start-a-conversation)
84
- - [Chat (streaming and non-streaming)](#chat-streaming-and-non-streaming)
85
- - [Workforce (multi-agent orchestration)](#workforce-multi-agent-orchestration) — teams of agents with a designated manager
81
+ - [Custom agent evaluations](#custom-agent-evaluations) — LLM-as-a-judge, cosine / Jaccard similarity, any framework
86
82
  - [Production tracing](#production-tracing) — record live agent runs from any framework
87
83
  - [Monitor](#monitor) — automatic production monitoring, patterns and signals
88
- - [Custom agent evaluations](#custom-agent-evaluations) — LLM-as-a-judge, cosine / Jaccard similarity
89
84
  - [Self-host](#self-host) — run Trace/Evaluate/Monitor on your own machine instead of the hosted dashboard
85
+ - [Agents & conversations](#agents--conversations) — chat with and orchestrate AgentX's own hosted agents
90
86
  - [Links](#links)
91
87
 
92
88
  ---
93
89
 
94
90
  ## Why AgentX
95
91
 
96
- - **Simple mental model** — `Agent Conversation Message`.
97
- - **Chain-of-thought** is built in, no extra plumbing.
98
- - **Bring any LLM** — works across major open and closed-source vendors.
99
- - **Batteries included** — voice (ASR/TTS), image generation, document/CSV/Excel/OCR, RAG with built-in re-ranking.
100
- - **MCP support** — connect any Model Context Protocol server.
101
- - **Multi-agent orchestration** — workforces of agents with a designated manager, across LLM vendors.
102
- - **Production tracing** — one decorator or context manager records every agent run (input, output, latency, tool calls, token usage) into your workspace, for any framework.
103
- - **Agent Evaluations** — score any agent (LangChain, CrewAI, OpenAI, Anthropic, HTTP, …) with LLM-as-a-judge ratings plus optional cosine and Jaccard similarity metrics. Configurable judge prompt/model (OpenAI or Anthropic), per-question judge guidelines, and smoke testing for phrasing robustness.
104
- - **A2A** — Each agent can be published with agent-to-agent protocol compatible.
92
+ - **Agent Evaluations** — score **any** agent (LangChain, CrewAI, AutoGen, LlamaIndex, OpenAI, Anthropic, HTTP, or plain Python) against a dataset with LLM-as-a-judge ratings plus optional cosine, Jaccard, and BLEU/ROUGE similarity metrics. Configurable judge prompt/model, per-question judge guidelines, smoke testing for phrasing robustness, and a durable multi-judge analysis pass for a qualitative report.
93
+ - **Production tracing** one decorator or context manager records every agent run (input, output, latency, tool calls, token usage, and a real span tree) into your workspace, for any framework.
94
+ - **Monitor** — check live traces against detection patterns or a sampled LLM judge, and read back triage-ready signals, no dashboard setup required.
95
+ - **Prompt registry** — make AgentX the source of truth for your own agent's prompts: pull a version at runtime, tag eval runs and live traces with it, let a judge propose a rewrite from your worst-rated results.
96
+ - **Self-host** — run the whole Trace/Evaluate/Monitor stack locally, bring your own LLM keys, no account required.
97
+ - **Bring any LLM** — works across major open and closed-source vendors, for evaluation, tracing, and AgentX's own hosted agents alike.
98
+ - **AgentX's own hosted agents** — a simple `Agent Conversation Message` mental model, chain-of-thought built in, multi-agent workforces, MCP support, and A2A publishing, for when you want AgentX to run the agent too, not just evaluate/trace/monitor it.
105
99
 
106
100
  ---
107
101
 
@@ -132,80 +126,72 @@ client = AgentX.from_env()
132
126
 
133
127
  ## Quick start
134
128
 
129
+ Evaluate your own agent — any framework, or plain Python — against a dataset:
130
+
135
131
  ```python
136
132
  from agentx import AgentX
137
133
 
138
134
  client = AgentX.from_env()
139
135
 
140
- # Pick an existing agent and chat with it
141
- agent = client.list_agents()[0]
142
- conversation = agent.new_conversation()
143
- print(conversation.chat("Hello! What can you help me with?"))
136
+ def my_agent(case):
137
+ return call_my_agent(case.query) # your agent's own code, any framework
138
+
139
+ report = (
140
+ client.evaluations
141
+ .run(dataset_id="evds_…", subject={"kind": "custom_agent", "framework": "raw_python"})
142
+ .execute(my_agent)
143
+ .finalize()
144
+ .analyze()
145
+ )
146
+
147
+ print(report.average_rating) # LLM-graded score, 0–10
148
+ print(report.summary) # AI-generated narrative from .analyze()
144
149
  ```
145
150
 
146
- That's it. The remaining sections show the same primitives in more detail.
151
+ That's it. The rest of this section covers building the dataset, framework adapters, similarity metrics, and judge configuration.
147
152
 
148
153
  ---
149
154
 
150
- ## Working with agents
155
+ ## Custom agent evaluations
151
156
 
152
- ### List agents
157
+ Evaluate **any** AI agent — LangChain, CrewAI, AutoGen, LlamaIndex, OpenAI, Anthropic, HTTP endpoints, or plain Python — using AgentX as the scoring and reporting backend. Includes optional **cosine**, **Jaccard**, and **BLEU/ROUGE** similarity metrics alongside LLM-graded ratings.
153
158
 
154
159
  ```python
155
- agents = client.list_agents()
156
- print(f"You have {len(agents)} agents")
157
- ```
160
+ report = (
161
+ client.evaluations
162
+ .run(dataset_id="evds_…", subject={"kind": "custom_agent", "framework": "raw_python"})
163
+ .execute(my_agent_fn)
164
+ .finalize()
165
+ .analyze()
166
+ )
158
167
 
159
- ### Start a conversation
168
+ print(report.average_rating) # LLM-graded score, 0–10
169
+ print(report.cosine_similarity) # embedding cosine, 0–1 (None if not enabled)
170
+ print(report.jaccard_similarity) # token-set overlap, 0–1 (None if not enabled)
160
171
 
161
- ```python
162
- agent = client.get_agent(id="<agent-id>")
172
+ print(report.summary) # AI-generated narrative from .analyze()
173
+ print(report.recommendations) # list of prioritized, actionable fixes
174
+ ```
163
175
 
164
- # Either resume an existing conversation…
165
- existing = agent.list_conversations()
166
- last = existing[-1]
167
- for msg in last.list_messages():
168
- print(msg)
176
+ `.analyze()` also generates a full qualitative report (strengths, weaknesses, instruction adherence, reasoning quality, and recommendations), running the same durable, multi-judge pipeline as the dashboard's "Analyze" button. `analyze(mode=..., quality_mode=..., judges=[...])` controls how items are scored and by which models. See [AI analysis report](EVALUATIONS.md#ai-analysis-report) in the full guide for the complete field and parameter reference.
169
177
 
170
- # …or start a fresh one
171
- conversation = agent.new_conversation()
172
- ```
178
+ Ask a case's question several extra ways each run, LLM-paraphrased server-side, to catch agents that break on phrasing rather than substance, and override the judge's prompt/model per config, see [Smoke testing](EVALUATIONS.md#smoke-testing-phrasing-robustness) and [Configuring the judge](EVALUATIONS.md#configuring-the-judge) in the full guide.
173
179
 
174
- ### Chat (streaming and non-streaming)
180
+ Since AgentX doesn't own your agent's code, `client.evaluations.prompts` lets AgentX become your prompt's *source of truth* instead — the same problem LangSmith's Prompt Hub and Langfuse's Prompt Management solve. Pull a version at runtime, tag your eval runs (or live traces) with it, and let a judge propose a rewrite from your real worst-rated results — a human always has to approve before it publishes:
175
181
 
176
182
  ```python
177
- # Blocking returns the full response once it's ready
178
- response = conversation.chat("What is your name?")
179
- print(response)
183
+ prompt = client.evaluations.prompts.get("support-agent-system-prompt") # or prompt.id
184
+ # use prompt.text as your own agent's system prompt
180
185
 
181
- # Streaming — yields ChatResponse objects as the model produces them
182
- for chunk in conversation.chat_stream("Hello, what is your name?"):
183
- if chunk.text:
184
- print(chunk.text, end="")
186
+ client.evaluations.run(
187
+ dataset_id="evds_…",
188
+ subject={"kind": "custom_agent", "metadata": {"promptName": prompt.name}},
189
+ ).execute(my_agent_fn)
185
190
  ```
186
191
 
187
- Each `ChatResponse` chunk exposes the agent's `text` and, where applicable, its `cot` (chain-of-thought) reasoning, along with any retrieved references and tasks.
188
-
189
- ---
190
-
191
- ## Workforce (multi-agent orchestration)
192
-
193
- A **workforce** is a team of agents coordinated by a designated manager agent. Workforces can mix LLM vendors and route work between specialists.
194
-
195
- ```python
196
- workforces = client.list_workforces()
197
- workforce = workforces[0]
198
-
199
- print(f"Workforce: {workforce.name}")
200
- print(f"Manager: {workforce.manager.name}")
201
- print(f"Agents: {[a.name for a in workforce.agents]}")
192
+ See [Prompt registry](EVALUATIONS.md#prompt-registry) in the full guide, or [self-host's docs](https://docs.agentx.so/self-host#prompt-registry) for the "Suggest improvement" dashboard flow (self-host only no hosted-SaaS equivalent yet).
202
193
 
203
- # Chat with the workforcethe manager decides which agent(s) to delegate to
204
- conversation = workforce.new_conversation()
205
- for chunk in workforce.chat_stream(conversation.id, "How can you help me with this project?"):
206
- if chunk.text:
207
- print(chunk.text, end="")
208
- ```
194
+ See **[EVALUATIONS.md](EVALUATIONS.md)** for the full guide dataset builder, framework adapters, similarity metrics, smoke testing, judge configuration, prompt registry, and the complete API reference.
209
195
 
210
196
  ---
211
197
 
@@ -315,66 +301,55 @@ See **[TRACING.md](TRACING.md)** for the complete Monitor guide.
315
301
 
316
302
  ---
317
303
 
318
- ## Custom agent evaluations
319
-
320
- Evaluate **any** AI agent — LangChain, CrewAI, AutoGen, LlamaIndex, OpenAI, Anthropic, HTTP endpoints, or plain Python — using AgentX as the scoring and reporting backend. Includes optional **cosine** and **Jaccard** similarity metrics alongside LLM-graded ratings.
321
-
322
- ```python
323
- report = (
324
- client.evaluations
325
- .run(dataset_id="evds_…", subject={"kind": "custom_agent", "framework": "raw_python"})
326
- .execute(my_agent_fn)
327
- .finalize()
328
- .analyze()
329
- )
304
+ ## Self-host
330
305
 
331
- print(report.average_rating) # LLM-graded score, 0–10
332
- print(report.cosine_similarity) # embedding cosine, 0–1 (None if not enabled)
333
- print(report.jaccard_similarity) # token-set overlap, 0–1 (None if not enabled)
306
+ Prefer to run Trace/Evaluate/Monitor locally instead of the hosted dashboard — no account, bring your own LLM keys? This SDK ships a launcher for [AgentX-trace-eval](https://github.com/AgentX-ai/AgentX-trace-eval), a separate, portable governance engine:
334
307
 
335
- print(report.summary) # AI-generated narrative from .analyze()
336
- print(report.recommendations) # list of prioritized, actionable fixes
308
+ ```bash
309
+ agentx-trace-eval --dev
337
310
  ```
338
311
 
339
- `.analyze()` also generates a full qualitative report (strengths, weaknesses, instruction adherence, reasoning quality, and recommendations), running the same durable, multi-judge pipeline as the dashboard's "Analyze" button. `analyze(mode=..., quality_mode=..., judges=[...])` controls how items are scored and by which models. See [AI analysis report](EVALUATIONS.md#ai-analysis-report) in the full guide for the complete field and parameter reference.
312
+ The first run downloads the engine (and dashboard) into `~/.agentx/bin` and prints a local API key; every run after that just starts it. Point this SDK at it instead of the hosted API:
340
313
 
341
- Ask a case's question several extra ways each run, LLM-paraphrased server-side, to catch agents that break on phrasing rather than substance, and override the judge's prompt/model per config, see [Smoke testing](EVALUATIONS.md#smoke-testing-phrasing-robustness) and [Configuring the judge](EVALUATIONS.md#configuring-the-judge) in the full guide.
314
+ ```bash
315
+ export AGENTX_API_BASE_URL=http://localhost:4700/api/v1
316
+ export AGENTX_API_KEY=<printed by agentx-trace-eval on first run>
317
+ ```
342
318
 
343
- Since AgentX doesn't own your agent's code, `client.evaluations.prompts` lets AgentX become your prompt's *source of truth* instead the same problem LangSmith's Prompt Hub and Langfuse's Prompt Management solve. Pull a version at runtime, tag your eval runs (or live traces) with it, and let a judge propose a rewrite from your real worst-rated results a human always has to approve before it publishes:
319
+ `agentx-trace-eval` isn't this SDK's own code the engine itself is a separate, compiled binary, downloaded on demand rather than bundled into this package, so installing `agentx-python` doesn't get any heavier for the (much more common) case of just talking to the hosted AgentX API. See that repo's README for what's included, and `AGENTX_INSTALL_DIR`/`AGENTX_TRACE_EVAL_VERSION`/`AGENTX_TRACE_EVAL_SKIP_WEB` env vars to control where/what it installs.
344
320
 
345
- ```python
346
- prompt = client.evaluations.prompts.get("support-agent-system-prompt") # or prompt.id
347
- # use prompt.text as your own agent's system prompt
321
+ ---
348
322
 
349
- client.evaluations.run(
350
- dataset_id="evds_…",
351
- subject={"kind": "custom_agent", "metadata": {"promptName": prompt.name}},
352
- ).execute(my_agent_fn)
353
- ```
323
+ ## Agents & conversations
354
324
 
355
- See [Prompt registry](EVALUATIONS.md#prompt-registry) in the full guide, or [self-host's docs](https://docs.agentx.so/self-host#prompt-registry) for the "Suggest improvement" dashboard flow (self-host only no hosted-SaaS equivalent yet).
325
+ Beyond evaluation, tracing, and monitoring, this SDK is also a client for AgentX's own hosted agents build, chat with, and orchestrate them directly.
356
326
 
357
- See **[EVALUATIONS.md](EVALUATIONS.md)** for the full guide — dataset builder, framework adapters, similarity metrics, smoke testing, judge configuration, prompt registry, and the complete API reference.
327
+ ```python
328
+ agent = client.list_agents()[0]
329
+ conversation = agent.new_conversation()
358
330
 
359
- ---
331
+ # Blocking — returns the full response once it's ready
332
+ print(conversation.chat("What can you help me with?"))
360
333
 
361
- ## Self-host
334
+ # Streaming — yields ChatResponse objects as the model produces them
335
+ for chunk in conversation.chat_stream("Hello!"):
336
+ if chunk.text:
337
+ print(chunk.text, end="")
338
+ ```
362
339
 
363
- Prefer to run Trace/Evaluate/Monitor locally instead of the hosted dashboard no account, bring your own LLM keys? This SDK ships a launcher for [AgentX-trace-eval](https://github.com/AgentX-ai/AgentX-trace-eval), a separate, portable governance engine:
340
+ Each `ChatResponse` chunk exposes the agent's `text` and, where applicable, its `cot` (chain-of-thought) reasoning, along with any retrieved references and tasks. `agent.list_conversations()` / `conversation.list_messages()` resume history instead of starting fresh.
364
341
 
365
- ```bash
366
- agentx-trace-eval --dev
367
- ```
342
+ A **workforce** is a team of agents coordinated by a designated manager, mixing LLM vendors and routing work between specialists:
368
343
 
369
- The first run downloads the engine (and dashboard) into `~/.agentx/bin` and prints a local API key; every run after that just starts it. Point this SDK at it instead of the hosted API:
344
+ ```python
345
+ workforce = client.list_workforces()[0]
346
+ conversation = workforce.new_conversation()
370
347
 
371
- ```bash
372
- export AGENTX_API_BASE_URL=http://localhost:4700/api/v1
373
- export AGENTX_API_KEY=<printed by agentx-trace-eval on first run>
348
+ for chunk in workforce.chat_stream(conversation.id, "How can you help me with this project?"):
349
+ if chunk.text:
350
+ print(chunk.text, end="")
374
351
  ```
375
352
 
376
- `agentx-trace-eval` isn't this SDK's own code — the engine itself is a separate, compiled binary, downloaded on demand rather than bundled into this package, so installing `agentx-python` doesn't get any heavier for the (much more common) case of just talking to the hosted AgentX API. See that repo's README for what's included, and `AGENTX_INSTALL_DIR`/`AGENTX_TRACE_EVAL_VERSION`/`AGENTX_TRACE_EVAL_SKIP_WEB` env vars to control where/what it installs.
377
-
378
353
  ---
379
354
 
380
355
  ## Links
@@ -1 +0,0 @@
1
- VERSION = "0.6.13"
File without changes
File without changes
File without changes