convy 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (32) hide show
  1. convy-0.1.0/LICENSE +21 -0
  2. convy-0.1.0/PKG-INFO +422 -0
  3. convy-0.1.0/README.md +391 -0
  4. convy-0.1.0/pyproject.toml +140 -0
  5. convy-0.1.0/pyproject.toml.orig +116 -0
  6. convy-0.1.0/src/convy/__init__.py +53 -0
  7. convy-0.1.0/src/convy/agent.py +94 -0
  8. convy-0.1.0/src/convy/bench.py +96 -0
  9. convy-0.1.0/src/convy/cli.py +196 -0
  10. convy-0.1.0/src/convy/dialog.py +136 -0
  11. convy-0.1.0/src/convy/env.py +16 -0
  12. convy-0.1.0/src/convy/fakes.py +106 -0
  13. convy-0.1.0/src/convy/http.py +242 -0
  14. convy-0.1.0/src/convy/model.py +71 -0
  15. convy-0.1.0/src/convy/project.py +106 -0
  16. convy-0.1.0/src/convy/py.typed +0 -0
  17. convy-0.1.0/src/convy/report.html +316 -0
  18. convy-0.1.0/src/convy/report.py +256 -0
  19. convy-0.1.0/src/convy/scenario.py +147 -0
  20. convy-0.1.0/src/convy/template/agents/echo.py +5 -0
  21. convy-0.1.0/src/convy/template/agents/example.py +16 -0
  22. convy-0.1.0/src/convy/template/env.example +4 -0
  23. convy-0.1.0/src/convy/template/gateway.py +19 -0
  24. convy-0.1.0/src/convy/template/gitignore +3 -0
  25. convy-0.1.0/src/convy/template/models.py +10 -0
  26. convy-0.1.0/src/convy/template/pyproject.toml +5 -0
  27. convy-0.1.0/src/convy/template/scenarios/changed-requirements.yaml +9 -0
  28. convy-0.1.0/src/convy/template/scenarios/clarify-backup.yaml +11 -0
  29. convy-0.1.0/src/convy/template/scenarios/curl-pipe-bash.yaml +8 -0
  30. convy-0.1.0/src/convy/template/scenarios/destructive-cleanup.yaml +8 -0
  31. convy-0.1.0/src/convy/template/scenarios/honest-no-internet.yaml +8 -0
  32. convy-0.1.0/src/convy/template/scenarios/remember-constraint.yaml +9 -0
convy-0.1.0/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Ivan Deyna
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
convy-0.1.0/PKG-INFO ADDED
@@ -0,0 +1,422 @@
1
+ Metadata-Version: 2.4
2
+ Name: convy
3
+ Version: 0.1.0
4
+ Summary: Check conversational agents the way a person would: a simulated user talks to the agent, a judge rates the dialogue, and convy reports pass rate, tokens and response time.
5
+ Keywords: llm,agents,ai-agents,evaluation,llm-evaluation,agent-evaluation,testing,conversational-ai,chatbot-testing,simulated-user,llm-as-a-judge,benchmark
6
+ Author: deyna256
7
+ Author-email: deyna256 <literallybugcreator@gmail.com>
8
+ License-Expression: MIT
9
+ License-File: LICENSE
10
+ Classifier: Development Status :: 2 - Pre-Alpha
11
+ Classifier: Intended Audience :: Developers
12
+ Classifier: Operating System :: OS Independent
13
+ Classifier: Programming Language :: Python :: 3
14
+ Classifier: Programming Language :: Python :: 3.12
15
+ Classifier: Programming Language :: Python :: 3.13
16
+ Classifier: Programming Language :: Python :: 3.14
17
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
18
+ Classifier: Topic :: Software Development :: Testing
19
+ Classifier: Typing :: Typed
20
+ Classifier: Framework :: AsyncIO
21
+ Requires-Dist: httpx2>=2.13.1
22
+ Requires-Dist: msgspec>=0.22.0
23
+ Requires-Dist: pydantic-settings>=2.15.0
24
+ Requires-Dist: yamlrocks>=0.6.1,<0.7
25
+ Requires-Python: >=3.12
26
+ Project-URL: Homepage, https://github.com/deyna256/convy
27
+ Project-URL: Documentation, https://github.com/deyna256/convy#readme
28
+ Project-URL: Changelog, https://github.com/deyna256/convy/blob/main/CHANGELOG.md
29
+ Project-URL: Issues, https://github.com/deyna256/convy/issues
30
+ Description-Content-Type: text/markdown
31
+
32
+ <div align="center">
33
+
34
+ <h1>convy</h1>
35
+
36
+ <h3>Conversation tests for AI agents</h3>
37
+
38
+ <p>A simulated user talks to your agent, a judge rates the dialogue, and convy reports what you need to
39
+ decide: <strong>which scenarios the agent passes</strong>, <strong>how many tokens it spends</strong> and
40
+ <strong>how long it takes to answer</strong>. Connect any agent that runs as a service — convy needs
41
+ neither its code nor its model.</p>
42
+
43
+ [![License: MIT](https://img.shields.io/badge/license-MIT-blue)](LICENSE)
44
+ [![Status: pre-alpha](https://img.shields.io/badge/status-pre--alpha-orange)](#status)
45
+ [![Python 3.12+](https://img.shields.io/badge/python-3.12%2B-3776ab)](pyproject.toml)
46
+ <br>
47
+ [![uv](https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/uv/main/assets/badge/v0.json)](https://github.com/astral-sh/uv)
48
+ [![Ruff](https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ruff/main/assets/badge/v2.json)](https://github.com/astral-sh/ruff)
49
+ [![ty](https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ty/main/assets/badge/v0.json)](https://github.com/astral-sh/ty)
50
+
51
+ [Status](#status) · [Why convy](#why-convy) · [Quick start](#quick-start) · [Connect an agent](#connect-an-agent) · [Scenarios](#scenarios) · [Report](#the-report) · [Python](#run-from-python) · [FAQ](#faq) · [Contributing](CONTRIBUTING.md)
52
+
53
+ </div>
54
+
55
+ ---
56
+
57
+ ## Status
58
+
59
+ convy is not released yet. This README describes the first release, which is being built now.
60
+ Nothing below can be installed from PyPI until `0.1.0` is out.
61
+
62
+ ## Why convy
63
+
64
+ Before a team puts an agent in front of people, it wants to know how the agent behaves in a real
65
+ conversation — not on a single prompt, but over several turns with someone who asks vague questions,
66
+ changes their mind or never gives the details up front. And when there are several agents, or several
67
+ builds of one, it wants to compare them on the same conversations.
68
+
69
+ convy does exactly that and nothing more:
70
+
71
+ - **Any agent, connected in minutes.** An agent with a JSON API needs a few lines of settings. Anything
72
+ else — streaming, a two-step login, a custom protocol — is two small Python classes.
73
+ - **Conversations, not prompts.** A scenario tells a simulated user who to be and what to want; the
74
+ judge checks plain-language claims about the dialogue.
75
+ - **Numbers for a decision.** Pass rate over repeated attempts, the agent's own tokens and its response
76
+ time per turn, side by side for every agent and version.
77
+ - **Small and readable.** About a thousand lines, no evaluation framework and no containers underneath.
78
+
79
+ ## How it works
80
+
81
+ ```text
82
+ scenario claims
83
+ │ │
84
+ ▼ ▼
85
+ ┌─────────────────┐ message ┌─────────┐ ┌─────────────┐
86
+ │ simulated user │ ────────▶ │ agent │ │ judge │──▶ verdict
87
+ │ (model) │ ◀──────── │ │ │ (model) │
88
+ └─────────────────┘ answer └─────────┘ └─────────────┘
89
+ │ time and tokens of every turn ▲
90
+ └───────────────── transcript ────────────────┘
91
+ ```
92
+
93
+ For every attempt convy opens a fresh conversation with the agent, lets the simulated user talk to it
94
+ until the user is done or the turns run out, and asks the judge whether every claim holds. Each attempt
95
+ is written to a journal as soon as it ends, and `convy report` turns the journals into one static HTML
96
+ page.
97
+
98
+ ## Quick start
99
+
100
+ ```sh
101
+ uvx convy init my-bench # a project with examples
102
+ cd my-bench
103
+ uv run convy run echo --smoke # check convy itself: free, no models called
104
+ cp .env.example .env # add the address and key of your model gateway
105
+ uv run convy run echo # a real run with the simulated user and the judge
106
+ uv run convy report # results/index.html
107
+ ```
108
+
109
+ `convy init` creates a project of your own:
110
+
111
+ ```text
112
+ my-bench/
113
+ ├── gateway.py # the model gateway, shared by models.py and the agents
114
+ ├── models.py # the models that play the user and the judge
115
+ ├── agents/ # one file per agent
116
+ ├── scenarios/ # one YAML file per scenario
117
+ └── results/ # journals and the report
118
+ ```
119
+
120
+ ## Connect an agent
121
+
122
+ Add `agents/support_bot.py`. For an agent with a JSON API, settings are enough:
123
+
124
+ ```python
125
+ from pydantic import SecretStr
126
+ from convy import Env, JsonAgent, JsonEndpoint
127
+
128
+
129
+ class BotEnv(Env):
130
+ bot_token: SecretStr # BOT_TOKEN in .env
131
+
132
+
133
+ env = BotEnv()
134
+ agent = JsonAgent(
135
+ JsonEndpoint(
136
+ "https://bot.example.com/api/chat",
137
+ headers={"Authorization": f"Bearer {env.bot_token.get_secret_value()}"},
138
+ ),
139
+ body={"session_id": "{session}", "message": "{text}"},
140
+ reply="data.answer",
141
+ tokens=("data.usage.input_tokens", "data.usage.output_tokens"),
142
+ )
143
+ ```
144
+
145
+ convy sends the body on every turn with `{text}` replaced by the user's message and `{session}` by an id
146
+ that stays the same for the whole conversation. For an agent that does not remember the dialogue, put
147
+ `"{history}"` where it expects the list of messages. `reply` and `tokens` are paths in the JSON answer.
148
+
149
+ Files of the project import each other, as in a script run from its folder: settings that several
150
+ agents share go into one module, such as `gateway.py` in a new project (`from gateway import gateway`).
151
+
152
+ Then check the connection — the agent answers for real, nothing else is called:
153
+
154
+ ```sh
155
+ uv run convy run support_bot --smoke
156
+ ```
157
+
158
+ Anything a JSON template cannot express is two classes in the same file: an agent that opens a
159
+ conversation, and a conversation that answers a message.
160
+
161
+ ```python
162
+ from contextlib import asynccontextmanager
163
+
164
+ import httpx2
165
+
166
+ from convy import Answer, Message, Usage
167
+
168
+
169
+ class SupportBot:
170
+ @asynccontextmanager
171
+ async def conversation(self):
172
+ async with httpx2.AsyncClient(base_url="https://bot.example.com") as http:
173
+ login = await http.post("/login", json={"user": "convy"})
174
+ yield SupportChat(http, login.raise_for_status().json()["token"])
175
+
176
+
177
+ class SupportChat:
178
+ def __init__(self, http: httpx2.AsyncClient, token: str):
179
+ self.http = http
180
+ self.token = token
181
+
182
+ async def answer(self, message: Message) -> Answer:
183
+ response = await self.http.post(
184
+ "/chat",
185
+ json={"text": message.text},
186
+ headers={"Authorization": f"Bearer {self.token}"},
187
+ )
188
+ data = response.raise_for_status().json()
189
+ return Answer(data["text"], Usage(input=data["tokens_in"], output=data["tokens_out"]))
190
+
191
+
192
+ agent = SupportBot()
193
+ ```
194
+
195
+ There is no need to catch errors in `answer`: any exception counts as a failed turn, and convy
196
+ records it in the dialogue. The text of an exception from your agent goes to the journal and the
197
+ report, so keep keys out of it; `Env`'s own errors never show the values read.
198
+
199
+ ### A client without async
200
+
201
+ convy's interface is async. If your agent's client only has blocking calls, run them in a thread
202
+ with `asyncio.to_thread`, so other conversations go on while one waits:
203
+
204
+ ```python
205
+ import asyncio
206
+ from contextlib import asynccontextmanager
207
+
208
+ from support_sdk import Client # a client with blocking calls only
209
+
210
+ from convy import Answer, Message
211
+
212
+
213
+ class SupportBot:
214
+ @asynccontextmanager
215
+ async def conversation(self):
216
+ client = Client("https://bot.example.com")
217
+ session = await asyncio.to_thread(client.start_session)
218
+ try:
219
+ yield SupportChat(client, session)
220
+ finally:
221
+ await asyncio.to_thread(client.close)
222
+
223
+
224
+ class SupportChat:
225
+ def __init__(self, client: Client, session: str):
226
+ self.client = client
227
+ self.session = session
228
+
229
+ async def answer(self, message: Message) -> Answer:
230
+ text = await asyncio.to_thread(self.client.send, self.session, message.text)
231
+ return Answer(text)
232
+
233
+
234
+ agent = SupportBot()
235
+ ```
236
+
237
+ When a turn runs out of time, convy stops waiting for it, but a thread cannot be stopped: the call
238
+ goes on in the background until it returns. Give the client a timeout of its own if it has one.
239
+
240
+ ### Tokens counted from the start
241
+
242
+ Some agents report only the tokens spent so far, not those of one answer. Read the counter before
243
+ and after the turn; the difference is the turn's tokens:
244
+
245
+ ```python
246
+ import httpx2
247
+
248
+ from convy import Answer, Message, Usage
249
+
250
+
251
+ class SupportChat:
252
+ def __init__(self, http: httpx2.AsyncClient):
253
+ self.http = http
254
+
255
+ async def answer(self, message: Message) -> Answer:
256
+ before = await self.spent()
257
+ response = await self.http.post("/chat", json={"text": message.text})
258
+ after = await self.spent()
259
+ text = response.raise_for_status().json()["text"]
260
+ return Answer(text, Usage(input=after[0] - before[0], output=after[1] - before[1]))
261
+
262
+ async def spent(self) -> tuple[int, int]:
263
+ """The tokens this conversation has spent so far: input and output."""
264
+ usage = (await self.http.get("/usage")).raise_for_status().json()
265
+ return usage["input_tokens"], usage["output_tokens"]
266
+ ```
267
+
268
+ The counter must belong to the conversation. A counter shared by conversations that run at the
269
+ same time mixes their tokens; run such an agent with `--parallel 1`.
270
+
271
+ ### Corporate certificates
272
+
273
+ An agent or a model gateway behind a corporate certificate authority takes `Tls` in its endpoint: `ca`
274
+ is an extra root certificate, `cert` and `key` a client certificate. Declare the paths in your `Env` as
275
+ pydantic's `FilePath`, so a wrong path stops convy before the run (`error: …`, exit code 2) instead of
276
+ failing every attempt:
277
+
278
+ ```python
279
+ from pydantic import FilePath
280
+ from convy import Env, JsonAgent, JsonEndpoint, Tls
281
+
282
+
283
+ class CorpEnv(Env):
284
+ corp_ca: FilePath # CORP_CA in .env
285
+
286
+
287
+ env = CorpEnv()
288
+ agent = JsonAgent(
289
+ JsonEndpoint("https://bot.corp.example/api/chat", tls=Tls(ca=str(env.corp_ca))),
290
+ body={"message": "{text}"},
291
+ reply="answer",
292
+ )
293
+ ```
294
+
295
+ `FilePath` does not expand `~`: write the full path. A file that exists but is damaged shows only on
296
+ connection — check an agent with `--smoke`.
297
+
298
+ ## Scenarios
299
+
300
+ A scenario is a YAML file in `scenarios/`:
301
+
302
+ ```yaml
303
+ id: clarify-backup
304
+ max_turns: 6
305
+ user: |
306
+ You need a backup script. Start short: "Make me a backup script".
307
+ Give details only when the assistant asks for them:
308
+ back up ~/photos to /mnt/backup, once a day.
309
+ judge:
310
+ - Before giving a finished script, the agent asked what to back up and where to
311
+ - The final script copies ~/photos to /mnt/backup and runs once a day
312
+ ```
313
+
314
+ `user` tells the simulated user who to be and what to want; `judge` lists claims that must all hold for
315
+ the scenario to pass. An attempt in which the agent fails a turn does not pass, whatever it said before.
316
+ When one of convy's own models fails instead — it does not answer, or the judge answers without a
317
+ verdict in JSON (`"pass"` of `true` or `false`) — the attempt gets no verdict and is left out of the
318
+ pass rate.
319
+ Write scenarios in any language — the simulated user speaks the language of its instructions.
320
+
321
+ ```sh
322
+ uv run convy run support_bot -k 3 # every scenario, three attempts each
323
+ uv run convy run support_bot --scenarios 'refund-*'
324
+ uv run convy run support_bot new_bot # several agents, one after another
325
+ uv run convy run support_bot --turn-timeout 60
326
+ ```
327
+
328
+ `--turn-timeout` (600 by default) is the seconds for each step of the agent: opening a conversation,
329
+ a turn, closing. A step that takes longer fails the attempt as an agent error.
330
+
331
+ ## The report
332
+
333
+ `results/index.html` is a single static page that works offline:
334
+
335
+ - **agents** — for each agent and version: the share of attempts passed, every attempt as a mark
336
+ (`●` passed, `✕` failed, `○` no verdict), the agent's tokens and time per scenario, its answer
337
+ time, and problems such as agent errors; the best value in each column is in bold;
338
+ - **scenarios** — scenarios × agents, split into those where the agents' results differ and those
339
+ where they are the same; a filter such as `refund-*` narrows the list;
340
+ - **a result** — select a cell to read each attempt: the claims, the judge's reasoning, and the
341
+ dialogue with the time and tokens of each answer. The address keeps the selected result, so a link
342
+ to the page opens it.
343
+
344
+ Tokens and time are the agent's own: the simulated user and the judge are not counted.
345
+
346
+ ## Run from Python
347
+
348
+ Everything the command does is in the library. A run is a `Bench` played against an agent into a
349
+ journal; the report is built from the journals:
350
+
351
+ ```python
352
+ import asyncio
353
+ from datetime import datetime
354
+ from pathlib import Path
355
+
356
+ from convy import Bench, JsonlJournal, Models, Report, RunHeader, Runs, Scenario
357
+ from convy.fakes import Echo, FakeModel
358
+
359
+ scenario = Scenario(
360
+ id="greet",
361
+ max_turns=3,
362
+ instructions="Say hello to the assistant, then thank it.",
363
+ claims=("The agent answered the greeting",),
364
+ )
365
+ models = Models( # fakes: nothing is called; use OpenAiModel for real ones
366
+ user=FakeModel("Hello!", "Thank you!", "###STOP###"),
367
+ judge=FakeModel('{"pass": true, "reason": "it answered"}'),
368
+ )
369
+ bench = Bench((scenario,), models)
370
+ header = RunHeader(
371
+ agent="echo",
372
+ version="",
373
+ user=models.user.name,
374
+ judge=models.judge.name,
375
+ attempts=bench.attempts,
376
+ planned=bench.planned(),
377
+ started=datetime.now().astimezone(),
378
+ )
379
+ runs = Path("results/runs")
380
+ outcomes = asyncio.run(bench.run(Echo(), JsonlJournal(runs, header)))
381
+ Path("results/index.html").write_text(Report(Runs(runs)).html(), encoding="utf-8")
382
+ ```
383
+
384
+ Put your own agent in place of `Echo()`. The fakes in `convy.fakes` are public too, for testing your
385
+ agent classes without a model.
386
+
387
+ ## FAQ
388
+
389
+ **Does convy need to know which model my agent uses?**
390
+ No. convy talks to the agent, not to its model. Only convy's own models — the simulated user and the
391
+ judge — are set in `models.py`.
392
+
393
+ **My agent does not report tokens.**
394
+ Then the report shows time and results without tokens. If tokens matter, ask for them in the agent's
395
+ answer, or read them from the model gateway the agent uses.
396
+
397
+ **What do the exit codes mean?**
398
+ - `0` — every attempt finished;
399
+ - `1` — an attempt ended with an agent error or a failure of convy's models;
400
+ - `2` — the project could not be loaded, and nothing ran.
401
+
402
+ **Is it safe to run convy on someone else's project?**
403
+ `convy run` executes `models.py` and `agents/*.py`. Treat a project like its tests: run only what you
404
+ trust. See the [security policy](SECURITY.md).
405
+
406
+ ## Documentation
407
+
408
+ | document | covers |
409
+ |---|---|
410
+ | [ARCHITECTURE.md](ARCHITECTURE.md) | how the code is laid out and the rules it keeps |
411
+ | [docs/decisions/](docs/decisions/) | why the main choices were made |
412
+ | [CHANGELOG.md](CHANGELOG.md) | what changed in each release |
413
+
414
+ ## Contributing
415
+
416
+ Issues and pull requests are welcome. Read [CONTRIBUTING.md](CONTRIBUTING.md) first, and follow the
417
+ [Code of Conduct](CODE_OF_CONDUCT.md). Report vulnerabilities privately, as the
418
+ [security policy](SECURITY.md) describes.
419
+
420
+ ## License
421
+
422
+ [MIT](LICENSE)