onejudge-cli 0.3.1__py3-none-win_amd64.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,209 @@
1
+ Metadata-Version: 2.4
2
+ Name: onejudge-cli
3
+ Version: 0.3.1
4
+ Classifier: Development Status :: 4 - Beta
5
+ Classifier: Environment :: Console
6
+ Classifier: Intended Audience :: Developers
7
+ Classifier: License :: OSI Approved :: MIT License
8
+ Classifier: Operating System :: OS Independent
9
+ Classifier: Programming Language :: Rust
10
+ Classifier: Topic :: Software Development :: Testing
11
+ License-File: LICENSE
12
+ Summary: Drive a harness through a simulated conversation and score the transcript.
13
+ Keywords: cli,agent,evaluation,judge,harness
14
+ Author: Nick DeRobertis
15
+ License: MIT
16
+ Requires-Python: >=3.8
17
+ Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
18
+ Project-URL: Homepage, https://github.com/nickderobertis/onejudge
19
+ Project-URL: Repository, https://github.com/nickderobertis/onejudge
20
+
21
+ # onejudge
22
+
23
+ A Rust library that drives a **simulated interaction and evaluation loop** on top
24
+ of [`oneharness`](https://github.com/nickderobertis/oneharness): take a skill or
25
+ agent, drive it through a multi-turn conversation with a simulated user, and score
26
+ the resulting transcript with natural-language (judge) verdicts and tool-event
27
+ queries.
28
+
29
+ It is the engine extracted from
30
+ [`skilltest`](https://github.com/nickderobertis/skilltest) (see
31
+ [nickderobertis/skilltest#31](https://github.com/nickderobertis/skilltest/issues/31)).
32
+ The layering:
33
+
34
+ ```
35
+ oneharness → one harness invocation, one JSON report (pure substrate)
36
+ onejudge → simulated interaction + judging loop (this crate)
37
+ skilltest → test-framework surface: cases, evals-as-assertions, SDKs
38
+ ```
39
+
40
+ Reach for onejudge when you want to "drive a harness through a simulated
41
+ conversation and score the transcript" without skilltest's YAML / case framing.
42
+
43
+ ## Install
44
+
45
+ ```sh
46
+ cargo add onejudge
47
+ ```
48
+
49
+ Minimum supported Rust version: **1.82**.
50
+
51
+ ## The `onejudge` CLI
52
+
53
+ The same engine that *tests* a skill can *drive real work*. `onejudge run` points a
54
+ harness at a task and lets an LLM-driven **simulated user supervise** it — pushing
55
+ back, asking for verification, re-prompting — until a `done_when` condition holds
56
+ or `max_turns` is hit. Configured by YAML; the library API is unchanged and CLI
57
+ deps (clap, a YAML parser) are opt-in behind the non-default `cli` feature.
58
+
59
+ **Spin up a run in three steps:**
60
+
61
+ ```sh
62
+ cargo install onejudge --features cli # or: install.sh (prebuilt archives)
63
+ onejudge init # scaffold onejudge.yaml + oneharness configs
64
+ onejudge run # reads ./onejudge.yaml, drives to completion
65
+ ```
66
+
67
+ `init` shells out to `oneharness init` (needs oneharness **0.3.20+**) to scaffold
68
+ `oneharness.toml` (the agent side) and `oneharness.judge.toml` (the judge side),
69
+ then writes a fully-commented loop-only `onejudge.yaml`. The fields that make a run
70
+ yours are `task` (what to do), the system framing — a `skill` (a `SKILL.md`
71
+ directory) and/or a `system_prompt`, both optional — and the `user` block
72
+ (`persona` / `done_when` / `max_turns` — omit it for a single-turn run). After
73
+ each nonterminal agent turn, one unified supervisor call either completes with a
74
+ reason or supplies the exact next user message. It sees compact normalized tool
75
+ summaries by default, never raw dumps; when needed it may inspect the agent-side
76
+ recording with `oneharness history show <session>-skill --project <worktree>
77
+ --format text`. Agent and judge harnesses run in that worktree, but only agent
78
+ runs are automatically history-recorded. **Harness
79
+ and model selection lives in those `oneharness.toml`
80
+ files, not `onejudge.yaml`.** `onejudge schema` prints the annotated config, the
81
+ single source of truth for every field.
82
+
83
+ Flags override the file (flags > file > defaults), so one config serves many tasks:
84
+ `onejudge run --task - < task.txt`, `--max-turns 8`, `--format json -o result.json`.
85
+
86
+ ### Config
87
+
88
+ A run is a YAML file carrying only the loop's own concerns. The fields that make it
89
+ yours — `task`, the system framing (`skill` and/or `system_prompt`), and the `user`
90
+ block. Everything else has a default; omit `user` for a single-turn run. A minimal
91
+ config:
92
+
93
+ ```yaml
94
+ system_prompt: You are a senior engineer. Complete the task and keep tests green.
95
+ # skill: ./skills/my-skill # optional: a SKILL.md dir; its body is appended
96
+
97
+ task: Add a --version flag to the CLI.
98
+
99
+ user: # the simulated supervisor that drives the loop
100
+ persona: A demanding tech lead. Do not accept "done" until you have verified it.
101
+ done_when: the task is complete and all tests pass
102
+ max_turns: 8
103
+
104
+ evals: # optional: score the finished transcript
105
+ - criterion: the change is well-scoped and readable
106
+ kind: numeric
107
+ scale: [1, 5]
108
+ assessment: Identify useful follow-up work left out of scope.
109
+ ```
110
+
111
+ The harness and model come from oneharness's own config (`oneharness.toml` for the
112
+ agent, `oneharness.judge.toml` for the judge side) — `onejudge init` scaffolds
113
+ them. More keys — `provider` (`oneharness` / `command` / `split`, with the
114
+ oneharness `judge_config` path), `session`, boolean evals. `onejudge init` writes a
115
+ fully-commented starter and `onejudge schema` prints the annotated field reference
116
+ (the single source of truth); it is validated strictly (`deny_unknown_fields`) so a
117
+ typo is a loud error.
118
+
119
+ Human output is the conversation + tool actions + completion status + eval
120
+ verdicts; `--format json` emits the versioned [`Report`](docs/contract.md). The
121
+ exit code is `0` only when the task completed and every boolean eval passed, `1`
122
+ if it hit `max_turns` or a boolean eval failed, `2` on a bad config. Full docs:
123
+ **[docs/cli.md](docs/cli.md)**.
124
+
125
+ ## Concepts
126
+
127
+ - **`Provider`** is the boundary — onejudge never talks to a model directly. Every
128
+ model call goes through `oneharness`, and harness/model *selection* lives in
129
+ oneharness's config files, not onejudge.
130
+ - **`OneharnessProvider`** (default) shells out to the `oneharness` CLI
131
+ (v0.3.20+): the agent side uses the discovered `oneharness.toml`, and the judge
132
+ side uses a separate `--config` file (default `oneharness.judge.toml`).
133
+ - **`CommandProvider`** speaks a small [JSON-lines protocol](docs/protocol.md),
134
+ for a custom backend or a deterministic test double.
135
+ - **`SplitProvider`** composes two providers — one that runs the skill, one that
136
+ judges and role-plays the user (e.g. run the skill on one harness, judge on
137
+ another).
138
+ - **`Engine`** runs a **`Conversation`** (a `Skill`, an initial input, and an
139
+ optional `SimulatedUser`) into a **`Transcript`**, bounded by `max_turns` /
140
+ `done_when` / the skill declaring itself done.
141
+ - **`Transcript`** carries each turn plus the normalized **`ToolEvent`**s the
142
+ skill took, so the judge — and a **`ToolQuery`** — can reason over *what the
143
+ skill did*, not just what it said.
144
+ - **`Report`** is onejudge's own versioned contract (`SCHEMA_VERSION`): a
145
+ serializable bundle of the transcript, verdicts, optional free-text
146
+ `assessment`, and usage that higher-level
147
+ frameworks compose over and re-export. See [docs/contract.md](docs/contract.md).
148
+
149
+ Two things it improves over the in-skilltest engine:
150
+
151
+ 1. **The judge sees tool events.** Verdicts render the transcript with a compact,
152
+ token-budget-aware summary of each turn's tool calls, so a criterion like
153
+ "the change was committed" can be decided from the `git commit` the skill
154
+ actually ran — not only from what it said. `Transcript` also exposes a
155
+ `ToolQuery` primitive for events-backed assertions with no judge call.
156
+ 2. **One caller-owned session name.** The engine always threads a single
157
+ `--session <name>` across turns instead of extracting and re-passing a native
158
+ id; if a harness cannot bind a session, the provider gracefully retries the call
159
+ without it, re-prompting the inlined transcript.
160
+
161
+ ## Example
162
+
163
+ ```rust
164
+ use onejudge::{Conversation, Engine, OneharnessProvider, Settings, SimulatedUser, Skill};
165
+
166
+ let provider = OneharnessProvider::new();
167
+ // Harness/model selection lives in oneharness's config files, not here; Settings
168
+ // carries only the loop's own concerns (turn cap, session name).
169
+ let settings = Settings::new();
170
+ let engine = Engine::new(&provider, settings);
171
+
172
+ let skill = Skill::new("greeter", "./skills/greeter", "Greet the user warmly.");
173
+ let user = SimulatedUser::new("A curious first-time visitor.")
174
+ .done_when("the assistant has answered the visitor's question")
175
+ .max_turns(6);
176
+
177
+ let outcome = engine.run(&Conversation::multi_turn(skill, "hi", user))?;
178
+
179
+ let verdict = engine.judge_boolean("the reply was welcoming", &outcome.transcript)?;
180
+ println!("{:?}: {}", verdict.value, verdict.reason);
181
+ # Ok::<(), onejudge::Error>(())
182
+ ```
183
+
184
+ Drive a deterministic backend instead of a live harness by pointing a
185
+ `CommandProvider` at any command that speaks the [protocol](docs/protocol.md).
186
+
187
+ ## Development
188
+
189
+ The command surface is a `just` recipe set; `just --list` is the index.
190
+
191
+ ```sh
192
+ just bootstrap # clean-clone setup: toolchain + cargo tools + fetch
193
+ just check # the full gate: format, lint, doc, coverage-enforced tests, audit
194
+ just test # fast unit + integration + e2e
195
+ ```
196
+
197
+ The gate is deterministic and offline — the model is faked by **real subprocess
198
+ test doubles**, never mocked. The one path that needs a real external service is
199
+ proven in an opt-in tier, kept out of `check`:
200
+
201
+ - **`just test-live`** — the `OneharnessProvider` path against a real harness (see
202
+ [docs/live-tier.md](docs/live-tier.md)).
203
+
204
+ See [AGENTS.md](AGENTS.md) for the durable contributor guide.
205
+
206
+ ## License
207
+
208
+ [MIT](LICENSE).
209
+
@@ -0,0 +1,6 @@
1
+ onejudge_cli-0.3.1.data/scripts/onejudge.exe,sha256=OKM__4Idh_6l4bhsSO2Iia8jMKMBVGiAGFliPyd_J5A,1473536
2
+ onejudge_cli-0.3.1.dist-info/METADATA,sha256=OOhxQ8ziEC3CRAKzVvJN1nnm8-00VCGue4SYV-XwEwg,9537
3
+ onejudge_cli-0.3.1.dist-info/WHEEL,sha256=2zDlIYIdD4m4N3p5DVEG3iJhGLdhsBQgdH-FqVkAur8,94
4
+ onejudge_cli-0.3.1.dist-info/licenses/LICENSE,sha256=JeBGdcXIUkoTZEHhkKd-FmKgdgumkH509Ob1QWJomVE,1072
5
+ onejudge_cli-0.3.1.dist-info/sboms/onejudge.cyclonedx.json,sha256=yz1QIylOR266kzzhqKkgMRf-ZTcEsczj2zuO3dgx4yc,40806
6
+ onejudge_cli-0.3.1.dist-info/RECORD,,
@@ -0,0 +1,4 @@
1
+ Wheel-Version: 1.0
2
+ Generator: maturin (1.14.1)
3
+ Root-Is-Purelib: false
4
+ Tag: py3-none-win_amd64
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Nick DeRobertis
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.