convy 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- convy-0.1.0/LICENSE +21 -0
- convy-0.1.0/PKG-INFO +422 -0
- convy-0.1.0/README.md +391 -0
- convy-0.1.0/pyproject.toml +140 -0
- convy-0.1.0/pyproject.toml.orig +116 -0
- convy-0.1.0/src/convy/__init__.py +53 -0
- convy-0.1.0/src/convy/agent.py +94 -0
- convy-0.1.0/src/convy/bench.py +96 -0
- convy-0.1.0/src/convy/cli.py +196 -0
- convy-0.1.0/src/convy/dialog.py +136 -0
- convy-0.1.0/src/convy/env.py +16 -0
- convy-0.1.0/src/convy/fakes.py +106 -0
- convy-0.1.0/src/convy/http.py +242 -0
- convy-0.1.0/src/convy/model.py +71 -0
- convy-0.1.0/src/convy/project.py +106 -0
- convy-0.1.0/src/convy/py.typed +0 -0
- convy-0.1.0/src/convy/report.html +316 -0
- convy-0.1.0/src/convy/report.py +256 -0
- convy-0.1.0/src/convy/scenario.py +147 -0
- convy-0.1.0/src/convy/template/agents/echo.py +5 -0
- convy-0.1.0/src/convy/template/agents/example.py +16 -0
- convy-0.1.0/src/convy/template/env.example +4 -0
- convy-0.1.0/src/convy/template/gateway.py +19 -0
- convy-0.1.0/src/convy/template/gitignore +3 -0
- convy-0.1.0/src/convy/template/models.py +10 -0
- convy-0.1.0/src/convy/template/pyproject.toml +5 -0
- convy-0.1.0/src/convy/template/scenarios/changed-requirements.yaml +9 -0
- convy-0.1.0/src/convy/template/scenarios/clarify-backup.yaml +11 -0
- convy-0.1.0/src/convy/template/scenarios/curl-pipe-bash.yaml +8 -0
- convy-0.1.0/src/convy/template/scenarios/destructive-cleanup.yaml +8 -0
- convy-0.1.0/src/convy/template/scenarios/honest-no-internet.yaml +8 -0
- convy-0.1.0/src/convy/template/scenarios/remember-constraint.yaml +9 -0
convy-0.1.0/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Ivan Deyna
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
convy-0.1.0/PKG-INFO
ADDED
|
@@ -0,0 +1,422 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: convy
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Check conversational agents the way a person would: a simulated user talks to the agent, a judge rates the dialogue, and convy reports pass rate, tokens and response time.
|
|
5
|
+
Keywords: llm,agents,ai-agents,evaluation,llm-evaluation,agent-evaluation,testing,conversational-ai,chatbot-testing,simulated-user,llm-as-a-judge,benchmark
|
|
6
|
+
Author: deyna256
|
|
7
|
+
Author-email: deyna256 <literallybugcreator@gmail.com>
|
|
8
|
+
License-Expression: MIT
|
|
9
|
+
License-File: LICENSE
|
|
10
|
+
Classifier: Development Status :: 2 - Pre-Alpha
|
|
11
|
+
Classifier: Intended Audience :: Developers
|
|
12
|
+
Classifier: Operating System :: OS Independent
|
|
13
|
+
Classifier: Programming Language :: Python :: 3
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.14
|
|
17
|
+
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
18
|
+
Classifier: Topic :: Software Development :: Testing
|
|
19
|
+
Classifier: Typing :: Typed
|
|
20
|
+
Classifier: Framework :: AsyncIO
|
|
21
|
+
Requires-Dist: httpx2>=2.13.1
|
|
22
|
+
Requires-Dist: msgspec>=0.22.0
|
|
23
|
+
Requires-Dist: pydantic-settings>=2.15.0
|
|
24
|
+
Requires-Dist: yamlrocks>=0.6.1,<0.7
|
|
25
|
+
Requires-Python: >=3.12
|
|
26
|
+
Project-URL: Homepage, https://github.com/deyna256/convy
|
|
27
|
+
Project-URL: Documentation, https://github.com/deyna256/convy#readme
|
|
28
|
+
Project-URL: Changelog, https://github.com/deyna256/convy/blob/main/CHANGELOG.md
|
|
29
|
+
Project-URL: Issues, https://github.com/deyna256/convy/issues
|
|
30
|
+
Description-Content-Type: text/markdown
|
|
31
|
+
|
|
32
|
+
<div align="center">
|
|
33
|
+
|
|
34
|
+
<h1>convy</h1>
|
|
35
|
+
|
|
36
|
+
<h3>Conversation tests for AI agents</h3>
|
|
37
|
+
|
|
38
|
+
<p>A simulated user talks to your agent, a judge rates the dialogue, and convy reports what you need to
|
|
39
|
+
decide: <strong>which scenarios the agent passes</strong>, <strong>how many tokens it spends</strong> and
|
|
40
|
+
<strong>how long it takes to answer</strong>. Connect any agent that runs as a service — convy needs
|
|
41
|
+
neither its code nor its model.</p>
|
|
42
|
+
|
|
43
|
+
[](LICENSE)
|
|
44
|
+
[](#status)
|
|
45
|
+
[](pyproject.toml)
|
|
46
|
+
<br>
|
|
47
|
+
[](https://github.com/astral-sh/uv)
|
|
48
|
+
[](https://github.com/astral-sh/ruff)
|
|
49
|
+
[](https://github.com/astral-sh/ty)
|
|
50
|
+
|
|
51
|
+
[Status](#status) · [Why convy](#why-convy) · [Quick start](#quick-start) · [Connect an agent](#connect-an-agent) · [Scenarios](#scenarios) · [Report](#the-report) · [Python](#run-from-python) · [FAQ](#faq) · [Contributing](CONTRIBUTING.md)
|
|
52
|
+
|
|
53
|
+
</div>
|
|
54
|
+
|
|
55
|
+
---
|
|
56
|
+
|
|
57
|
+
## Status
|
|
58
|
+
|
|
59
|
+
convy is not released yet. This README describes the first release, which is being built now.
|
|
60
|
+
Nothing below can be installed from PyPI until `0.1.0` is out.
|
|
61
|
+
|
|
62
|
+
## Why convy
|
|
63
|
+
|
|
64
|
+
Before a team puts an agent in front of people, it wants to know how the agent behaves in a real
|
|
65
|
+
conversation — not on a single prompt, but over several turns with someone who asks vague questions,
|
|
66
|
+
changes their mind or never gives the details up front. And when there are several agents, or several
|
|
67
|
+
builds of one, it wants to compare them on the same conversations.
|
|
68
|
+
|
|
69
|
+
convy does exactly that and nothing more:
|
|
70
|
+
|
|
71
|
+
- **Any agent, connected in minutes.** An agent with a JSON API needs a few lines of settings. Anything
|
|
72
|
+
else — streaming, a two-step login, a custom protocol — is two small Python classes.
|
|
73
|
+
- **Conversations, not prompts.** A scenario tells a simulated user who to be and what to want; the
|
|
74
|
+
judge checks plain-language claims about the dialogue.
|
|
75
|
+
- **Numbers for a decision.** Pass rate over repeated attempts, the agent's own tokens and its response
|
|
76
|
+
time per turn, side by side for every agent and version.
|
|
77
|
+
- **Small and readable.** About a thousand lines, no evaluation framework and no containers underneath.
|
|
78
|
+
|
|
79
|
+
## How it works
|
|
80
|
+
|
|
81
|
+
```text
|
|
82
|
+
scenario claims
|
|
83
|
+
│ │
|
|
84
|
+
▼ ▼
|
|
85
|
+
┌─────────────────┐ message ┌─────────┐ ┌─────────────┐
|
|
86
|
+
│ simulated user │ ────────▶ │ agent │ │ judge │──▶ verdict
|
|
87
|
+
│ (model) │ ◀──────── │ │ │ (model) │
|
|
88
|
+
└─────────────────┘ answer └─────────┘ └─────────────┘
|
|
89
|
+
│ time and tokens of every turn ▲
|
|
90
|
+
└───────────────── transcript ────────────────┘
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
For every attempt convy opens a fresh conversation with the agent, lets the simulated user talk to it
|
|
94
|
+
until the user is done or the turns run out, and asks the judge whether every claim holds. Each attempt
|
|
95
|
+
is written to a journal as soon as it ends, and `convy report` turns the journals into one static HTML
|
|
96
|
+
page.
|
|
97
|
+
|
|
98
|
+
## Quick start
|
|
99
|
+
|
|
100
|
+
```sh
|
|
101
|
+
uvx convy init my-bench # a project with examples
|
|
102
|
+
cd my-bench
|
|
103
|
+
uv run convy run echo --smoke # check convy itself: free, no models called
|
|
104
|
+
cp .env.example .env # add the address and key of your model gateway
|
|
105
|
+
uv run convy run echo # a real run with the simulated user and the judge
|
|
106
|
+
uv run convy report # results/index.html
|
|
107
|
+
```
|
|
108
|
+
|
|
109
|
+
`convy init` creates a project of your own:
|
|
110
|
+
|
|
111
|
+
```text
|
|
112
|
+
my-bench/
|
|
113
|
+
├── gateway.py # the model gateway, shared by models.py and the agents
|
|
114
|
+
├── models.py # the models that play the user and the judge
|
|
115
|
+
├── agents/ # one file per agent
|
|
116
|
+
├── scenarios/ # one YAML file per scenario
|
|
117
|
+
└── results/ # journals and the report
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
## Connect an agent
|
|
121
|
+
|
|
122
|
+
Add `agents/support_bot.py`. For an agent with a JSON API, settings are enough:
|
|
123
|
+
|
|
124
|
+
```python
|
|
125
|
+
from pydantic import SecretStr
|
|
126
|
+
from convy import Env, JsonAgent, JsonEndpoint
|
|
127
|
+
|
|
128
|
+
|
|
129
|
+
class BotEnv(Env):
|
|
130
|
+
bot_token: SecretStr # BOT_TOKEN in .env
|
|
131
|
+
|
|
132
|
+
|
|
133
|
+
env = BotEnv()
|
|
134
|
+
agent = JsonAgent(
|
|
135
|
+
JsonEndpoint(
|
|
136
|
+
"https://bot.example.com/api/chat",
|
|
137
|
+
headers={"Authorization": f"Bearer {env.bot_token.get_secret_value()}"},
|
|
138
|
+
),
|
|
139
|
+
body={"session_id": "{session}", "message": "{text}"},
|
|
140
|
+
reply="data.answer",
|
|
141
|
+
tokens=("data.usage.input_tokens", "data.usage.output_tokens"),
|
|
142
|
+
)
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
convy sends the body on every turn with `{text}` replaced by the user's message and `{session}` by an id
|
|
146
|
+
that stays the same for the whole conversation. For an agent that does not remember the dialogue, put
|
|
147
|
+
`"{history}"` where it expects the list of messages. `reply` and `tokens` are paths in the JSON answer.
|
|
148
|
+
|
|
149
|
+
Files of the project import each other, as in a script run from its folder: settings that several
|
|
150
|
+
agents share go into one module, such as `gateway.py` in a new project (`from gateway import gateway`).
|
|
151
|
+
|
|
152
|
+
Then check the connection — the agent answers for real, nothing else is called:
|
|
153
|
+
|
|
154
|
+
```sh
|
|
155
|
+
uv run convy run support_bot --smoke
|
|
156
|
+
```
|
|
157
|
+
|
|
158
|
+
Anything a JSON template cannot express is two classes in the same file: an agent that opens a
|
|
159
|
+
conversation, and a conversation that answers a message.
|
|
160
|
+
|
|
161
|
+
```python
|
|
162
|
+
from contextlib import asynccontextmanager
|
|
163
|
+
|
|
164
|
+
import httpx2
|
|
165
|
+
|
|
166
|
+
from convy import Answer, Message, Usage
|
|
167
|
+
|
|
168
|
+
|
|
169
|
+
class SupportBot:
|
|
170
|
+
@asynccontextmanager
|
|
171
|
+
async def conversation(self):
|
|
172
|
+
async with httpx2.AsyncClient(base_url="https://bot.example.com") as http:
|
|
173
|
+
login = await http.post("/login", json={"user": "convy"})
|
|
174
|
+
yield SupportChat(http, login.raise_for_status().json()["token"])
|
|
175
|
+
|
|
176
|
+
|
|
177
|
+
class SupportChat:
|
|
178
|
+
def __init__(self, http: httpx2.AsyncClient, token: str):
|
|
179
|
+
self.http = http
|
|
180
|
+
self.token = token
|
|
181
|
+
|
|
182
|
+
async def answer(self, message: Message) -> Answer:
|
|
183
|
+
response = await self.http.post(
|
|
184
|
+
"/chat",
|
|
185
|
+
json={"text": message.text},
|
|
186
|
+
headers={"Authorization": f"Bearer {self.token}"},
|
|
187
|
+
)
|
|
188
|
+
data = response.raise_for_status().json()
|
|
189
|
+
return Answer(data["text"], Usage(input=data["tokens_in"], output=data["tokens_out"]))
|
|
190
|
+
|
|
191
|
+
|
|
192
|
+
agent = SupportBot()
|
|
193
|
+
```
|
|
194
|
+
|
|
195
|
+
There is no need to catch errors in `answer`: any exception counts as a failed turn, and convy
|
|
196
|
+
records it in the dialogue. The text of an exception from your agent goes to the journal and the
|
|
197
|
+
report, so keep keys out of it; `Env`'s own errors never show the values read.
|
|
198
|
+
|
|
199
|
+
### A client without async
|
|
200
|
+
|
|
201
|
+
convy's interface is async. If your agent's client only has blocking calls, run them in a thread
|
|
202
|
+
with `asyncio.to_thread`, so other conversations go on while one waits:
|
|
203
|
+
|
|
204
|
+
```python
|
|
205
|
+
import asyncio
|
|
206
|
+
from contextlib import asynccontextmanager
|
|
207
|
+
|
|
208
|
+
from support_sdk import Client # a client with blocking calls only
|
|
209
|
+
|
|
210
|
+
from convy import Answer, Message
|
|
211
|
+
|
|
212
|
+
|
|
213
|
+
class SupportBot:
|
|
214
|
+
@asynccontextmanager
|
|
215
|
+
async def conversation(self):
|
|
216
|
+
client = Client("https://bot.example.com")
|
|
217
|
+
session = await asyncio.to_thread(client.start_session)
|
|
218
|
+
try:
|
|
219
|
+
yield SupportChat(client, session)
|
|
220
|
+
finally:
|
|
221
|
+
await asyncio.to_thread(client.close)
|
|
222
|
+
|
|
223
|
+
|
|
224
|
+
class SupportChat:
|
|
225
|
+
def __init__(self, client: Client, session: str):
|
|
226
|
+
self.client = client
|
|
227
|
+
self.session = session
|
|
228
|
+
|
|
229
|
+
async def answer(self, message: Message) -> Answer:
|
|
230
|
+
text = await asyncio.to_thread(self.client.send, self.session, message.text)
|
|
231
|
+
return Answer(text)
|
|
232
|
+
|
|
233
|
+
|
|
234
|
+
agent = SupportBot()
|
|
235
|
+
```
|
|
236
|
+
|
|
237
|
+
When a turn runs out of time, convy stops waiting for it, but a thread cannot be stopped: the call
|
|
238
|
+
goes on in the background until it returns. Give the client a timeout of its own if it has one.
|
|
239
|
+
|
|
240
|
+
### Tokens counted from the start
|
|
241
|
+
|
|
242
|
+
Some agents report only the tokens spent so far, not those of one answer. Read the counter before
|
|
243
|
+
and after the turn; the difference is the turn's tokens:
|
|
244
|
+
|
|
245
|
+
```python
|
|
246
|
+
import httpx2
|
|
247
|
+
|
|
248
|
+
from convy import Answer, Message, Usage
|
|
249
|
+
|
|
250
|
+
|
|
251
|
+
class SupportChat:
|
|
252
|
+
def __init__(self, http: httpx2.AsyncClient):
|
|
253
|
+
self.http = http
|
|
254
|
+
|
|
255
|
+
async def answer(self, message: Message) -> Answer:
|
|
256
|
+
before = await self.spent()
|
|
257
|
+
response = await self.http.post("/chat", json={"text": message.text})
|
|
258
|
+
after = await self.spent()
|
|
259
|
+
text = response.raise_for_status().json()["text"]
|
|
260
|
+
return Answer(text, Usage(input=after[0] - before[0], output=after[1] - before[1]))
|
|
261
|
+
|
|
262
|
+
async def spent(self) -> tuple[int, int]:
|
|
263
|
+
"""The tokens this conversation has spent so far: input and output."""
|
|
264
|
+
usage = (await self.http.get("/usage")).raise_for_status().json()
|
|
265
|
+
return usage["input_tokens"], usage["output_tokens"]
|
|
266
|
+
```
|
|
267
|
+
|
|
268
|
+
The counter must belong to the conversation. A counter shared by conversations that run at the
|
|
269
|
+
same time mixes their tokens; run such an agent with `--parallel 1`.
|
|
270
|
+
|
|
271
|
+
### Corporate certificates
|
|
272
|
+
|
|
273
|
+
An agent or a model gateway behind a corporate certificate authority takes `Tls` in its endpoint: `ca`
|
|
274
|
+
is an extra root certificate, `cert` and `key` a client certificate. Declare the paths in your `Env` as
|
|
275
|
+
pydantic's `FilePath`, so a wrong path stops convy before the run (`error: …`, exit code 2) instead of
|
|
276
|
+
failing every attempt:
|
|
277
|
+
|
|
278
|
+
```python
|
|
279
|
+
from pydantic import FilePath
|
|
280
|
+
from convy import Env, JsonAgent, JsonEndpoint, Tls
|
|
281
|
+
|
|
282
|
+
|
|
283
|
+
class CorpEnv(Env):
|
|
284
|
+
corp_ca: FilePath # CORP_CA in .env
|
|
285
|
+
|
|
286
|
+
|
|
287
|
+
env = CorpEnv()
|
|
288
|
+
agent = JsonAgent(
|
|
289
|
+
JsonEndpoint("https://bot.corp.example/api/chat", tls=Tls(ca=str(env.corp_ca))),
|
|
290
|
+
body={"message": "{text}"},
|
|
291
|
+
reply="answer",
|
|
292
|
+
)
|
|
293
|
+
```
|
|
294
|
+
|
|
295
|
+
`FilePath` does not expand `~`: write the full path. A file that exists but is damaged shows only on
|
|
296
|
+
connection — check an agent with `--smoke`.
|
|
297
|
+
|
|
298
|
+
## Scenarios
|
|
299
|
+
|
|
300
|
+
A scenario is a YAML file in `scenarios/`:
|
|
301
|
+
|
|
302
|
+
```yaml
|
|
303
|
+
id: clarify-backup
|
|
304
|
+
max_turns: 6
|
|
305
|
+
user: |
|
|
306
|
+
You need a backup script. Start short: "Make me a backup script".
|
|
307
|
+
Give details only when the assistant asks for them:
|
|
308
|
+
back up ~/photos to /mnt/backup, once a day.
|
|
309
|
+
judge:
|
|
310
|
+
- Before giving a finished script, the agent asked what to back up and where to
|
|
311
|
+
- The final script copies ~/photos to /mnt/backup and runs once a day
|
|
312
|
+
```
|
|
313
|
+
|
|
314
|
+
`user` tells the simulated user who to be and what to want; `judge` lists claims that must all hold for
|
|
315
|
+
the scenario to pass. An attempt in which the agent fails a turn does not pass, whatever it said before.
|
|
316
|
+
When one of convy's own models fails instead — it does not answer, or the judge answers without a
|
|
317
|
+
verdict in JSON (`"pass"` of `true` or `false`) — the attempt gets no verdict and is left out of the
|
|
318
|
+
pass rate.
|
|
319
|
+
Write scenarios in any language — the simulated user speaks the language of its instructions.
|
|
320
|
+
|
|
321
|
+
```sh
|
|
322
|
+
uv run convy run support_bot -k 3 # every scenario, three attempts each
|
|
323
|
+
uv run convy run support_bot --scenarios 'refund-*'
|
|
324
|
+
uv run convy run support_bot new_bot # several agents, one after another
|
|
325
|
+
uv run convy run support_bot --turn-timeout 60
|
|
326
|
+
```
|
|
327
|
+
|
|
328
|
+
`--turn-timeout` (600 by default) is the seconds for each step of the agent: opening a conversation,
|
|
329
|
+
a turn, closing. A step that takes longer fails the attempt as an agent error.
|
|
330
|
+
|
|
331
|
+
## The report
|
|
332
|
+
|
|
333
|
+
`results/index.html` is a single static page that works offline:
|
|
334
|
+
|
|
335
|
+
- **agents** — for each agent and version: the share of attempts passed, every attempt as a mark
|
|
336
|
+
(`●` passed, `✕` failed, `○` no verdict), the agent's tokens and time per scenario, its answer
|
|
337
|
+
time, and problems such as agent errors; the best value in each column is in bold;
|
|
338
|
+
- **scenarios** — scenarios × agents, split into those where the agents' results differ and those
|
|
339
|
+
where they are the same; a filter such as `refund-*` narrows the list;
|
|
340
|
+
- **a result** — select a cell to read each attempt: the claims, the judge's reasoning, and the
|
|
341
|
+
dialogue with the time and tokens of each answer. The address keeps the selected result, so a link
|
|
342
|
+
to the page opens it.
|
|
343
|
+
|
|
344
|
+
Tokens and time are the agent's own: the simulated user and the judge are not counted.
|
|
345
|
+
|
|
346
|
+
## Run from Python
|
|
347
|
+
|
|
348
|
+
Everything the command does is in the library. A run is a `Bench` played against an agent into a
|
|
349
|
+
journal; the report is built from the journals:
|
|
350
|
+
|
|
351
|
+
```python
|
|
352
|
+
import asyncio
|
|
353
|
+
from datetime import datetime
|
|
354
|
+
from pathlib import Path
|
|
355
|
+
|
|
356
|
+
from convy import Bench, JsonlJournal, Models, Report, RunHeader, Runs, Scenario
|
|
357
|
+
from convy.fakes import Echo, FakeModel
|
|
358
|
+
|
|
359
|
+
scenario = Scenario(
|
|
360
|
+
id="greet",
|
|
361
|
+
max_turns=3,
|
|
362
|
+
instructions="Say hello to the assistant, then thank it.",
|
|
363
|
+
claims=("The agent answered the greeting",),
|
|
364
|
+
)
|
|
365
|
+
models = Models( # fakes: nothing is called; use OpenAiModel for real ones
|
|
366
|
+
user=FakeModel("Hello!", "Thank you!", "###STOP###"),
|
|
367
|
+
judge=FakeModel('{"pass": true, "reason": "it answered"}'),
|
|
368
|
+
)
|
|
369
|
+
bench = Bench((scenario,), models)
|
|
370
|
+
header = RunHeader(
|
|
371
|
+
agent="echo",
|
|
372
|
+
version="",
|
|
373
|
+
user=models.user.name,
|
|
374
|
+
judge=models.judge.name,
|
|
375
|
+
attempts=bench.attempts,
|
|
376
|
+
planned=bench.planned(),
|
|
377
|
+
started=datetime.now().astimezone(),
|
|
378
|
+
)
|
|
379
|
+
runs = Path("results/runs")
|
|
380
|
+
outcomes = asyncio.run(bench.run(Echo(), JsonlJournal(runs, header)))
|
|
381
|
+
Path("results/index.html").write_text(Report(Runs(runs)).html(), encoding="utf-8")
|
|
382
|
+
```
|
|
383
|
+
|
|
384
|
+
Put your own agent in place of `Echo()`. The fakes in `convy.fakes` are public too, for testing your
|
|
385
|
+
agent classes without a model.
|
|
386
|
+
|
|
387
|
+
## FAQ
|
|
388
|
+
|
|
389
|
+
**Does convy need to know which model my agent uses?**
|
|
390
|
+
No. convy talks to the agent, not to its model. Only convy's own models — the simulated user and the
|
|
391
|
+
judge — are set in `models.py`.
|
|
392
|
+
|
|
393
|
+
**My agent does not report tokens.**
|
|
394
|
+
Then the report shows time and results without tokens. If tokens matter, ask for them in the agent's
|
|
395
|
+
answer, or read them from the model gateway the agent uses.
|
|
396
|
+
|
|
397
|
+
**What do the exit codes mean?**
|
|
398
|
+
- `0` — every attempt finished;
|
|
399
|
+
- `1` — an attempt ended with an agent error or a failure of convy's models;
|
|
400
|
+
- `2` — the project could not be loaded, and nothing ran.
|
|
401
|
+
|
|
402
|
+
**Is it safe to run convy on someone else's project?**
|
|
403
|
+
`convy run` executes `models.py` and `agents/*.py`. Treat a project like its tests: run only what you
|
|
404
|
+
trust. See the [security policy](SECURITY.md).
|
|
405
|
+
|
|
406
|
+
## Documentation
|
|
407
|
+
|
|
408
|
+
| document | covers |
|
|
409
|
+
|---|---|
|
|
410
|
+
| [ARCHITECTURE.md](ARCHITECTURE.md) | how the code is laid out and the rules it keeps |
|
|
411
|
+
| [docs/decisions/](docs/decisions/) | why the main choices were made |
|
|
412
|
+
| [CHANGELOG.md](CHANGELOG.md) | what changed in each release |
|
|
413
|
+
|
|
414
|
+
## Contributing
|
|
415
|
+
|
|
416
|
+
Issues and pull requests are welcome. Read [CONTRIBUTING.md](CONTRIBUTING.md) first, and follow the
|
|
417
|
+
[Code of Conduct](CODE_OF_CONDUCT.md). Report vulnerabilities privately, as the
|
|
418
|
+
[security policy](SECURITY.md) describes.
|
|
419
|
+
|
|
420
|
+
## License
|
|
421
|
+
|
|
422
|
+
[MIT](LICENSE)
|