miniswen 1.4.0 → 1.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +17 -2
- data/lib/miniswen/agent.rb +52 -16
- data/lib/miniswen/cli.rb +8 -3
- data/lib/miniswen/environment/docker.rb +1 -1
- data/lib/miniswen/jail.rb +1 -1
- data/lib/miniswen/local.rb +1 -1
- data/lib/miniswen/version.rb +1 -1
- metadata +1 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 2df0b59321d93fcea1ddb87d555323b3dd247eed780532bd645d1410e5a6a391
|
|
4
|
+
data.tar.gz: 07c36283c59e9a311af76aa59b46118b92ff705a5478dca847f5bd321c03c7de
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: d481729c89f35e3aae9cbb02dc2c5db59327b292c49a4a30920e81a178480411d075e72bc59dc845bc0906e9194af7cd37d2833f2a092ae135a359b7460eb066
|
|
7
|
+
data.tar.gz: 1eb136297aac2c4702efc6b2c6bc4829a53bdca70a85c582bd39150cc17e0cd7851f62adcb0acee53cb05f82461339efb9d7a1d5057657cc47fc45300aee051a
|
data/README.md
CHANGED
|
@@ -15,6 +15,7 @@ lemans is a harness for benchmarking coding agents, the Ruby way:
|
|
|
15
15
|
- Ruby 3.4+ is required to run `lemans`
|
|
16
16
|
- Daytona account (API token) or Docker (for local sandboxes)
|
|
17
17
|
- Some LLM provider/proxy credentials (e.g., OpenRouter)
|
|
18
|
+
- Bash in the sandbox image for command execution
|
|
18
19
|
- iproute2 in the sandbox image, so a jailed agent keeps loopback while cut off the internet
|
|
19
20
|
|
|
20
21
|
## Getting started
|
|
@@ -102,6 +103,8 @@ verifier:
|
|
|
102
103
|
# restore: [test, bin] # [optional] paths restored from the pre-agent snapshot, so grading can't be tampered with
|
|
103
104
|
```
|
|
104
105
|
|
|
106
|
+
On the Docker backend, `allowlist` sends traffic through a proxy container. Only tools that use `http_proxy`/`https_proxy` can reach the listed hosts. Host names must match exactly, and IP ranges are not supported. See [the proxy README](lib/lemans/environments/docker/proxy/README.md).
|
|
107
|
+
|
|
105
108
|
A minimal task example—checking whether an agent can write "Hello, world" into a file:
|
|
106
109
|
|
|
107
110
|
- `instruction.md`: the task itself (note the preamble about the training corpora and the frontmatter)
|
|
@@ -231,6 +234,17 @@ Each trial writes a flat `runs/<model>/<task>__<id>/` directory:
|
|
|
231
234
|
|
|
232
235
|
Every `result.json` is stamped with the `lemans_version` that wrote it, and newer lemans keeps reading runs produced by older releases. `lemans run --resume` skips trials that already have a scored result for the same agent and model.
|
|
233
236
|
|
|
237
|
+
`lemans restart <run>` continues a multistep run that failed for reasons outside the model, such as a provider outage. Give it the run directories or the trial ids: `lemans restart RUN [RUN...]` restarts every run under the same options, `-c` at once, and runs none if any is refused. The new run replays the agent patches of the settled steps in a fresh sandbox and starts at the next step. It copies the settled steps' usage, phases, and artifacts, and records the source in `restarted_from`. The failed run stays as it is. It restarts even if the task or the bench changed since; the new run records the current digests.
|
|
238
|
+
|
|
239
|
+
Two modes restart inside a step:
|
|
240
|
+
|
|
241
|
+
- `--recover` continues the failed step's agent session. The new run also replays the step's partial patch, and the agent goes on from the saved `agent.result.json`. The session keeps the steps, tokens, cost, and time it already spent. Only `miniswen` and `miniswen-installed` can recover.
|
|
242
|
+
- `--reverify` runs the last verification again with the current tests. Use it after you change the grading or the harness. The new run replays the graded step's patch and does not run its agent. If the verification was an intermediate gate and it now passes, the trial goes on to the next step. A reverification accepts scored runs.
|
|
243
|
+
|
|
244
|
+
A scored run restarts in the other modes only with `--allow-scored`.
|
|
245
|
+
|
|
246
|
+
Only the files that git tracks come back: ignored files (databases, logs, installed dependencies) are not in the patches.
|
|
247
|
+
|
|
234
248
|
You can also run `lemans report` with various flags to see aggregated results, e.g.:
|
|
235
249
|
|
|
236
250
|
```sh
|
|
@@ -256,8 +270,9 @@ gpt-5.6-luna ar-archive-book-access 2/2 2m 23s $0.0132 12.5 156905
|
|
|
256
270
|
| `lemans init` | Scaffold a new bench directory: an annotated `bench.yml` and two example tasks |
|
|
257
271
|
| `lemans tasks` | List the tasks in a bench (`--tag` to filter) |
|
|
258
272
|
| `lemans run` | Run tasks and grade them (`--task`, `--tag`, `--agent`, `--model`, `--max-output-tokens`, `-k`, `-c`, `--resume`) |
|
|
259
|
-
| `lemans
|
|
260
|
-
| `lemans
|
|
273
|
+
| `lemans restart <run>...` | Continue failed multistep runs from their last settled step in new runs (`-c`, `--recover` to continue the failed step's session, `--reverify` to grade again, `--allow-scored`, `--backend`, `--max-output-tokens`) |
|
|
274
|
+
| `lemans report [RUNS_DIR]` | Summarize `runs/` (or `RUNS_DIR`) as a table or CSV (`--task`, `--tag`, `--metadata key:value` to filter, `-A [task-agent-model]` to aggregate, `-S <column>` to sort); repeated attempts add pass@k per model × task, fractional grading a `credit` column |
|
|
275
|
+
| `lemans clobber [RUNS_DIR]` | Delete run results under `runs/` (or `RUNS_DIR`) (`--task`, `--ttl 10m\|2h\|1d`, `--invalid`, `-f` to skip the confirmation) |
|
|
261
276
|
|
|
262
277
|
## miniswen
|
|
263
278
|
|
data/lib/miniswen/agent.rb
CHANGED
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
# frozen_string_literal: true
|
|
2
2
|
|
|
3
3
|
require "json"
|
|
4
|
+
require "time"
|
|
4
5
|
require "miniswen/version"
|
|
5
6
|
require "miniswen/ruby_llm"
|
|
6
7
|
|
|
@@ -261,25 +262,17 @@ module Miniswen
|
|
|
261
262
|
@reporter = reporter
|
|
262
263
|
end
|
|
263
264
|
|
|
264
|
-
|
|
265
|
-
|
|
266
|
-
|
|
267
|
-
|
|
268
|
-
|
|
269
|
-
|
|
270
|
-
|
|
271
|
-
|
|
272
|
-
]
|
|
273
|
-
|
|
274
|
-
@steps = 0
|
|
275
|
-
@cost = 0.0
|
|
276
|
-
|
|
277
|
-
@totals = { input_tokens: 0, output_tokens: 0, cached_tokens: 0, thinking_tokens: 0 }
|
|
265
|
+
# With a `history` (the Result of an interrupted run), the session goes on
|
|
266
|
+
# where it stopped, under the budget the earlier part already spent.
|
|
267
|
+
def run(instruction = nil, history: nil)
|
|
268
|
+
if history
|
|
269
|
+
resume(history)
|
|
270
|
+
else
|
|
271
|
+
start(instruction)
|
|
272
|
+
end
|
|
278
273
|
|
|
279
|
-
@cost_known = true
|
|
280
274
|
@consecutive_format_errors = 0
|
|
281
275
|
@refused_turns = 0
|
|
282
|
-
@started_at = @clock.call
|
|
283
276
|
|
|
284
277
|
loop do
|
|
285
278
|
(status = limit_reached) and return finish(status)
|
|
@@ -329,6 +322,49 @@ module Miniswen
|
|
|
329
322
|
|
|
330
323
|
private
|
|
331
324
|
|
|
325
|
+
def start(instruction)
|
|
326
|
+
uname = execute("uname -srvm").output.to_s.strip
|
|
327
|
+
@messages = [
|
|
328
|
+
{ role: "system", content: SYSTEM_TEMPLATE },
|
|
329
|
+
{ role: "user", content: format(INSTANCE_TEMPLATE,
|
|
330
|
+
instruction: instruction,
|
|
331
|
+
system_information: uname,
|
|
332
|
+
macos_sed_note: uname.start_with?("Darwin") ? "\n#{MACOS_SED_NOTE}" : "") }
|
|
333
|
+
]
|
|
334
|
+
|
|
335
|
+
@steps = 0
|
|
336
|
+
@cost = 0.0
|
|
337
|
+
@cost_known = true
|
|
338
|
+
@totals = { input_tokens: 0, output_tokens: 0, cached_tokens: 0, thinking_tokens: 0 }
|
|
339
|
+
@started_at = @clock.call
|
|
340
|
+
end
|
|
341
|
+
|
|
342
|
+
def resume(history)
|
|
343
|
+
@messages = answered_messages(history.messages)
|
|
344
|
+
|
|
345
|
+
@steps = history.steps.to_i
|
|
346
|
+
@cost = history.cost_usd.to_f
|
|
347
|
+
@cost_known = !history.cost_usd.nil?
|
|
348
|
+
@totals = { input_tokens: history.input_tokens.to_i, output_tokens: history.output_tokens.to_i,
|
|
349
|
+
cached_tokens: history.cached_tokens.to_i, thinking_tokens: history.thinking_tokens.to_i }
|
|
350
|
+
@started_at = @clock.call - elapsed(history.messages)
|
|
351
|
+
end
|
|
352
|
+
|
|
353
|
+
# A turn whose calls did not all get a result cannot go back to a provider:
|
|
354
|
+
# it is dropped, and the model decides again from the tree it left.
|
|
355
|
+
def answered_messages(messages)
|
|
356
|
+
turn = messages.rindex { it[:role] == "assistant" && it[:tool_calls] }
|
|
357
|
+
return messages.dup unless turn
|
|
358
|
+
|
|
359
|
+
answered = messages[(turn + 1)..].filter_map { it[:tool_call_id] if it[:role] == "tool" }
|
|
360
|
+
messages[turn][:tool_calls].all? { answered.include?(it[:id]) } ? messages.dup : messages[0...turn]
|
|
361
|
+
end
|
|
362
|
+
|
|
363
|
+
def elapsed(messages)
|
|
364
|
+
times = messages.filter_map { Time.parse(it[:timestamp]) if it[:role] == "assistant" && it[:timestamp] }
|
|
365
|
+
times.empty? ? 0.0 : times.max - times.min
|
|
366
|
+
end
|
|
367
|
+
|
|
332
368
|
def execute(command)
|
|
333
369
|
environment.exec(command, timeout: exec_timeout, env: EXEC_ENV)
|
|
334
370
|
end
|
data/lib/miniswen/cli.rb
CHANGED
|
@@ -25,6 +25,7 @@ module Miniswen
|
|
|
25
25
|
@jail = false
|
|
26
26
|
@allowed_hosts = nil
|
|
27
27
|
@workdir = nil
|
|
28
|
+
@history = nil
|
|
28
29
|
end
|
|
29
30
|
|
|
30
31
|
def run
|
|
@@ -56,7 +57,7 @@ module Miniswen
|
|
|
56
57
|
agent = Agent.new(model:, reporter:, environment:, **options)
|
|
57
58
|
|
|
58
59
|
begin
|
|
59
|
-
result = agent.run(instruction)
|
|
60
|
+
result = agent.run(instruction, history: @history && Agent::Result.from_h(JSON.parse(@history)))
|
|
60
61
|
rescue StandardError => e
|
|
61
62
|
write_results(agent.partial_result(error_message(e)))
|
|
62
63
|
raise
|
|
@@ -105,7 +106,7 @@ module Miniswen
|
|
|
105
106
|
|
|
106
107
|
def parse_args!
|
|
107
108
|
parser = OptionParser.new do |opts|
|
|
108
|
-
opts.banner = "Usage: miniswen -m MODEL -p INSTRUCTION [...options]"
|
|
109
|
+
opts.banner = "Usage: miniswen -m MODEL (-p INSTRUCTION | --continue-from PATH) [...options]"
|
|
109
110
|
|
|
110
111
|
opts.on("-m MODEL", "--model=MODEL", String,
|
|
111
112
|
"LLM to use (litellm format, e.g.: openrouter/openai/gpt-5.6-luna") do |v|
|
|
@@ -116,6 +117,10 @@ module Miniswen
|
|
|
116
117
|
@instruction = File.file?(v) ? File.read(v) : v
|
|
117
118
|
end
|
|
118
119
|
|
|
120
|
+
opts.on("--continue-from=PATH", String, "Continue the session saved by --results-path at PATH (no -p needed)") do |v|
|
|
121
|
+
@history = File.read(v)
|
|
122
|
+
end
|
|
123
|
+
|
|
119
124
|
opts.on("--max-steps=STEPS", Integer, "Max steps count") do |v|
|
|
120
125
|
options[:max_steps] = v
|
|
121
126
|
end
|
|
@@ -197,7 +202,7 @@ module Miniswen
|
|
|
197
202
|
return if @refresh_registry
|
|
198
203
|
|
|
199
204
|
raise "Use -m to specify the model" unless @model
|
|
200
|
-
raise "Please, provide instructions via -p option" unless @instruction
|
|
205
|
+
raise "Please, provide instructions via -p option" unless @instruction || @history
|
|
201
206
|
end
|
|
202
207
|
end
|
|
203
208
|
end
|
|
@@ -20,7 +20,7 @@ module Miniswen
|
|
|
20
20
|
env&.each { |key, value| argv += [ "--env", "#{key}=#{value}" ] }
|
|
21
21
|
argv << id
|
|
22
22
|
argv += [ "timeout", timeout.ceil.to_s ] if timeout&.positive?
|
|
23
|
-
argv += [ "
|
|
23
|
+
argv += [ "bash", "-c", command ]
|
|
24
24
|
|
|
25
25
|
Open3.popen2e(*argv) do |stdin, pipe, wait|
|
|
26
26
|
stdin.close
|
data/lib/miniswen/jail.rb
CHANGED
|
@@ -61,7 +61,7 @@ module Miniswen
|
|
|
61
61
|
def spawn_arguments(command, env)
|
|
62
62
|
[ command_env(env),
|
|
63
63
|
"nsenter", "--target", @holder.pid.to_s, "--net", "--mount", "--pid=/proc/#{@holder.pid}/ns/pid_for_children", "--wd=#{@workdir}",
|
|
64
|
-
"--", "setpriv", "--bounding-set=-all", "--inh-caps=-all", "--no-new-privs", "--", "
|
|
64
|
+
"--", "setpriv", "--bounding-set=-all", "--inh-caps=-all", "--no-new-privs", "--", "bash", "-c", command ]
|
|
65
65
|
end
|
|
66
66
|
|
|
67
67
|
def spawn_options = super.merge(unsetenv_others: true)
|
data/lib/miniswen/local.rb
CHANGED
data/lib/miniswen/version.rb
CHANGED