onejudge-cli 0.3.1__py3-none-win_amd64.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- onejudge_cli-0.3.1.data/scripts/onejudge.exe +0 -0
- onejudge_cli-0.3.1.dist-info/METADATA +209 -0
- onejudge_cli-0.3.1.dist-info/RECORD +6 -0
- onejudge_cli-0.3.1.dist-info/WHEEL +4 -0
- onejudge_cli-0.3.1.dist-info/licenses/LICENSE +21 -0
- onejudge_cli-0.3.1.dist-info/sboms/onejudge.cyclonedx.json +1330 -0
|
Binary file
|
|
@@ -0,0 +1,209 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: onejudge-cli
|
|
3
|
+
Version: 0.3.1
|
|
4
|
+
Classifier: Development Status :: 4 - Beta
|
|
5
|
+
Classifier: Environment :: Console
|
|
6
|
+
Classifier: Intended Audience :: Developers
|
|
7
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
8
|
+
Classifier: Operating System :: OS Independent
|
|
9
|
+
Classifier: Programming Language :: Rust
|
|
10
|
+
Classifier: Topic :: Software Development :: Testing
|
|
11
|
+
License-File: LICENSE
|
|
12
|
+
Summary: Drive a harness through a simulated conversation and score the transcript.
|
|
13
|
+
Keywords: cli,agent,evaluation,judge,harness
|
|
14
|
+
Author: Nick DeRobertis
|
|
15
|
+
License: MIT
|
|
16
|
+
Requires-Python: >=3.8
|
|
17
|
+
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
|
|
18
|
+
Project-URL: Homepage, https://github.com/nickderobertis/onejudge
|
|
19
|
+
Project-URL: Repository, https://github.com/nickderobertis/onejudge
|
|
20
|
+
|
|
21
|
+
# onejudge
|
|
22
|
+
|
|
23
|
+
A Rust library that drives a **simulated interaction and evaluation loop** on top
|
|
24
|
+
of [`oneharness`](https://github.com/nickderobertis/oneharness): take a skill or
|
|
25
|
+
agent, drive it through a multi-turn conversation with a simulated user, and score
|
|
26
|
+
the resulting transcript with natural-language (judge) verdicts and tool-event
|
|
27
|
+
queries.
|
|
28
|
+
|
|
29
|
+
It is the engine extracted from
|
|
30
|
+
[`skilltest`](https://github.com/nickderobertis/skilltest) (see
|
|
31
|
+
[nickderobertis/skilltest#31](https://github.com/nickderobertis/skilltest/issues/31)).
|
|
32
|
+
The layering:
|
|
33
|
+
|
|
34
|
+
```
|
|
35
|
+
oneharness → one harness invocation, one JSON report (pure substrate)
|
|
36
|
+
onejudge → simulated interaction + judging loop (this crate)
|
|
37
|
+
skilltest → test-framework surface: cases, evals-as-assertions, SDKs
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
Reach for onejudge when you want to "drive a harness through a simulated
|
|
41
|
+
conversation and score the transcript" without skilltest's YAML / case framing.
|
|
42
|
+
|
|
43
|
+
## Install
|
|
44
|
+
|
|
45
|
+
```sh
|
|
46
|
+
cargo add onejudge
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
Minimum supported Rust version: **1.82**.
|
|
50
|
+
|
|
51
|
+
## The `onejudge` CLI
|
|
52
|
+
|
|
53
|
+
The same engine that *tests* a skill can *drive real work*. `onejudge run` points a
|
|
54
|
+
harness at a task and lets an LLM-driven **simulated user supervise** it — pushing
|
|
55
|
+
back, asking for verification, re-prompting — until a `done_when` condition holds
|
|
56
|
+
or `max_turns` is hit. Configured by YAML; the library API is unchanged and CLI
|
|
57
|
+
deps (clap, a YAML parser) are opt-in behind the non-default `cli` feature.
|
|
58
|
+
|
|
59
|
+
**Spin up a run in three steps:**
|
|
60
|
+
|
|
61
|
+
```sh
|
|
62
|
+
cargo install onejudge --features cli # or: install.sh (prebuilt archives)
|
|
63
|
+
onejudge init # scaffold onejudge.yaml + oneharness configs
|
|
64
|
+
onejudge run # reads ./onejudge.yaml, drives to completion
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
`init` shells out to `oneharness init` (needs oneharness **0.3.20+**) to scaffold
|
|
68
|
+
`oneharness.toml` (the agent side) and `oneharness.judge.toml` (the judge side),
|
|
69
|
+
then writes a fully-commented loop-only `onejudge.yaml`. The fields that make a run
|
|
70
|
+
yours are `task` (what to do), the system framing — a `skill` (a `SKILL.md`
|
|
71
|
+
directory) and/or a `system_prompt`, both optional — and the `user` block
|
|
72
|
+
(`persona` / `done_when` / `max_turns` — omit it for a single-turn run). After
|
|
73
|
+
each nonterminal agent turn, one unified supervisor call either completes with a
|
|
74
|
+
reason or supplies the exact next user message. It sees compact normalized tool
|
|
75
|
+
summaries by default, never raw dumps; when needed it may inspect the agent-side
|
|
76
|
+
recording with `oneharness history show <session>-skill --project <worktree>
|
|
77
|
+
--format text`. Agent and judge harnesses run in that worktree, but only agent
|
|
78
|
+
runs are automatically history-recorded. **Harness
|
|
79
|
+
and model selection lives in those `oneharness.toml`
|
|
80
|
+
files, not `onejudge.yaml`.** `onejudge schema` prints the annotated config, the
|
|
81
|
+
single source of truth for every field.
|
|
82
|
+
|
|
83
|
+
Flags override the file (flags > file > defaults), so one config serves many tasks:
|
|
84
|
+
`onejudge run --task - < task.txt`, `--max-turns 8`, `--format json -o result.json`.
|
|
85
|
+
|
|
86
|
+
### Config
|
|
87
|
+
|
|
88
|
+
A run is a YAML file carrying only the loop's own concerns. The fields that make it
|
|
89
|
+
yours — `task`, the system framing (`skill` and/or `system_prompt`), and the `user`
|
|
90
|
+
block. Everything else has a default; omit `user` for a single-turn run. A minimal
|
|
91
|
+
config:
|
|
92
|
+
|
|
93
|
+
```yaml
|
|
94
|
+
system_prompt: You are a senior engineer. Complete the task and keep tests green.
|
|
95
|
+
# skill: ./skills/my-skill # optional: a SKILL.md dir; its body is appended
|
|
96
|
+
|
|
97
|
+
task: Add a --version flag to the CLI.
|
|
98
|
+
|
|
99
|
+
user: # the simulated supervisor that drives the loop
|
|
100
|
+
persona: A demanding tech lead. Do not accept "done" until you have verified it.
|
|
101
|
+
done_when: the task is complete and all tests pass
|
|
102
|
+
max_turns: 8
|
|
103
|
+
|
|
104
|
+
evals: # optional: score the finished transcript
|
|
105
|
+
- criterion: the change is well-scoped and readable
|
|
106
|
+
kind: numeric
|
|
107
|
+
scale: [1, 5]
|
|
108
|
+
assessment: Identify useful follow-up work left out of scope.
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
The harness and model come from oneharness's own config (`oneharness.toml` for the
|
|
112
|
+
agent, `oneharness.judge.toml` for the judge side) — `onejudge init` scaffolds
|
|
113
|
+
them. More keys — `provider` (`oneharness` / `command` / `split`, with the
|
|
114
|
+
oneharness `judge_config` path), `session`, boolean evals. `onejudge init` writes a
|
|
115
|
+
fully-commented starter and `onejudge schema` prints the annotated field reference
|
|
116
|
+
(the single source of truth); it is validated strictly (`deny_unknown_fields`) so a
|
|
117
|
+
typo is a loud error.
|
|
118
|
+
|
|
119
|
+
Human output is the conversation + tool actions + completion status + eval
|
|
120
|
+
verdicts; `--format json` emits the versioned [`Report`](docs/contract.md). The
|
|
121
|
+
exit code is `0` only when the task completed and every boolean eval passed, `1`
|
|
122
|
+
if it hit `max_turns` or a boolean eval failed, `2` on a bad config. Full docs:
|
|
123
|
+
**[docs/cli.md](docs/cli.md)**.
|
|
124
|
+
|
|
125
|
+
## Concepts
|
|
126
|
+
|
|
127
|
+
- **`Provider`** is the boundary — onejudge never talks to a model directly. Every
|
|
128
|
+
model call goes through `oneharness`, and harness/model *selection* lives in
|
|
129
|
+
oneharness's config files, not onejudge.
|
|
130
|
+
- **`OneharnessProvider`** (default) shells out to the `oneharness` CLI
|
|
131
|
+
(v0.3.20+): the agent side uses the discovered `oneharness.toml`, and the judge
|
|
132
|
+
side uses a separate `--config` file (default `oneharness.judge.toml`).
|
|
133
|
+
- **`CommandProvider`** speaks a small [JSON-lines protocol](docs/protocol.md),
|
|
134
|
+
for a custom backend or a deterministic test double.
|
|
135
|
+
- **`SplitProvider`** composes two providers — one that runs the skill, one that
|
|
136
|
+
judges and role-plays the user (e.g. run the skill on one harness, judge on
|
|
137
|
+
another).
|
|
138
|
+
- **`Engine`** runs a **`Conversation`** (a `Skill`, an initial input, and an
|
|
139
|
+
optional `SimulatedUser`) into a **`Transcript`**, bounded by `max_turns` /
|
|
140
|
+
`done_when` / the skill declaring itself done.
|
|
141
|
+
- **`Transcript`** carries each turn plus the normalized **`ToolEvent`**s the
|
|
142
|
+
skill took, so the judge — and a **`ToolQuery`** — can reason over *what the
|
|
143
|
+
skill did*, not just what it said.
|
|
144
|
+
- **`Report`** is onejudge's own versioned contract (`SCHEMA_VERSION`): a
|
|
145
|
+
serializable bundle of the transcript, verdicts, optional free-text
|
|
146
|
+
`assessment`, and usage that higher-level
|
|
147
|
+
frameworks compose over and re-export. See [docs/contract.md](docs/contract.md).
|
|
148
|
+
|
|
149
|
+
Two things it improves over the in-skilltest engine:
|
|
150
|
+
|
|
151
|
+
1. **The judge sees tool events.** Verdicts render the transcript with a compact,
|
|
152
|
+
token-budget-aware summary of each turn's tool calls, so a criterion like
|
|
153
|
+
"the change was committed" can be decided from the `git commit` the skill
|
|
154
|
+
actually ran — not only from what it said. `Transcript` also exposes a
|
|
155
|
+
`ToolQuery` primitive for events-backed assertions with no judge call.
|
|
156
|
+
2. **One caller-owned session name.** The engine always threads a single
|
|
157
|
+
`--session <name>` across turns instead of extracting and re-passing a native
|
|
158
|
+
id; if a harness cannot bind a session, the provider gracefully retries the call
|
|
159
|
+
without it, re-prompting the inlined transcript.
|
|
160
|
+
|
|
161
|
+
## Example
|
|
162
|
+
|
|
163
|
+
```rust
|
|
164
|
+
use onejudge::{Conversation, Engine, OneharnessProvider, Settings, SimulatedUser, Skill};
|
|
165
|
+
|
|
166
|
+
let provider = OneharnessProvider::new();
|
|
167
|
+
// Harness/model selection lives in oneharness's config files, not here; Settings
|
|
168
|
+
// carries only the loop's own concerns (turn cap, session name).
|
|
169
|
+
let settings = Settings::new();
|
|
170
|
+
let engine = Engine::new(&provider, settings);
|
|
171
|
+
|
|
172
|
+
let skill = Skill::new("greeter", "./skills/greeter", "Greet the user warmly.");
|
|
173
|
+
let user = SimulatedUser::new("A curious first-time visitor.")
|
|
174
|
+
.done_when("the assistant has answered the visitor's question")
|
|
175
|
+
.max_turns(6);
|
|
176
|
+
|
|
177
|
+
let outcome = engine.run(&Conversation::multi_turn(skill, "hi", user))?;
|
|
178
|
+
|
|
179
|
+
let verdict = engine.judge_boolean("the reply was welcoming", &outcome.transcript)?;
|
|
180
|
+
println!("{:?}: {}", verdict.value, verdict.reason);
|
|
181
|
+
# Ok::<(), onejudge::Error>(())
|
|
182
|
+
```
|
|
183
|
+
|
|
184
|
+
Drive a deterministic backend instead of a live harness by pointing a
|
|
185
|
+
`CommandProvider` at any command that speaks the [protocol](docs/protocol.md).
|
|
186
|
+
|
|
187
|
+
## Development
|
|
188
|
+
|
|
189
|
+
The command surface is a `just` recipe set; `just --list` is the index.
|
|
190
|
+
|
|
191
|
+
```sh
|
|
192
|
+
just bootstrap # clean-clone setup: toolchain + cargo tools + fetch
|
|
193
|
+
just check # the full gate: format, lint, doc, coverage-enforced tests, audit
|
|
194
|
+
just test # fast unit + integration + e2e
|
|
195
|
+
```
|
|
196
|
+
|
|
197
|
+
The gate is deterministic and offline — the model is faked by **real subprocess
|
|
198
|
+
test doubles**, never mocked. The one path that needs a real external service is
|
|
199
|
+
proven in an opt-in tier, kept out of `check`:
|
|
200
|
+
|
|
201
|
+
- **`just test-live`** — the `OneharnessProvider` path against a real harness (see
|
|
202
|
+
[docs/live-tier.md](docs/live-tier.md)).
|
|
203
|
+
|
|
204
|
+
See [AGENTS.md](AGENTS.md) for the durable contributor guide.
|
|
205
|
+
|
|
206
|
+
## License
|
|
207
|
+
|
|
208
|
+
[MIT](LICENSE).
|
|
209
|
+
|
|
@@ -0,0 +1,6 @@
|
|
|
1
|
+
onejudge_cli-0.3.1.data/scripts/onejudge.exe,sha256=OKM__4Idh_6l4bhsSO2Iia8jMKMBVGiAGFliPyd_J5A,1473536
|
|
2
|
+
onejudge_cli-0.3.1.dist-info/METADATA,sha256=OOhxQ8ziEC3CRAKzVvJN1nnm8-00VCGue4SYV-XwEwg,9537
|
|
3
|
+
onejudge_cli-0.3.1.dist-info/WHEEL,sha256=2zDlIYIdD4m4N3p5DVEG3iJhGLdhsBQgdH-FqVkAur8,94
|
|
4
|
+
onejudge_cli-0.3.1.dist-info/licenses/LICENSE,sha256=JeBGdcXIUkoTZEHhkKd-FmKgdgumkH509Ob1QWJomVE,1072
|
|
5
|
+
onejudge_cli-0.3.1.dist-info/sboms/onejudge.cyclonedx.json,sha256=yz1QIylOR266kzzhqKkgMRf-ZTcEsczj2zuO3dgx4yc,40806
|
|
6
|
+
onejudge_cli-0.3.1.dist-info/RECORD,,
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Nick DeRobertis
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|