miniswen 1.4.0 → 1.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 49e7b2a4a062867c6b85476f94ebda50ce7d356add5500d40b2c5c513155f7ca
4
- data.tar.gz: 4b5dfc61e161b9e2776b2547e8a10862e785fedcbd5fff1a1621054657ed21fb
3
+ metadata.gz: 2df0b59321d93fcea1ddb87d555323b3dd247eed780532bd645d1410e5a6a391
4
+ data.tar.gz: 07c36283c59e9a311af76aa59b46118b92ff705a5478dca847f5bd321c03c7de
5
5
  SHA512:
6
- metadata.gz: b6e67519d94b5e32469173068f8121f8682dd5f16398c38c118f4516931d9e8b2e068e89f612bdc3d543df0a10cc9d8279bc425de404b1e63774085f77f8fab1
7
- data.tar.gz: addc1fbee5c18257a2e0f115316a33a4f0f5e9233bb6a7f47ed43d244f1e49ec0a4fd912519e9442edb149a01c25d21cebe3f84cf8d4d6142d170bc30266c80b
6
+ metadata.gz: d481729c89f35e3aae9cbb02dc2c5db59327b292c49a4a30920e81a178480411d075e72bc59dc845bc0906e9194af7cd37d2833f2a092ae135a359b7460eb066
7
+ data.tar.gz: 1eb136297aac2c4702efc6b2c6bc4829a53bdca70a85c582bd39150cc17e0cd7851f62adcb0acee53cb05f82461339efb9d7a1d5057657cc47fc45300aee051a
data/README.md CHANGED
@@ -15,6 +15,7 @@ lemans is a harness for benchmarking coding agents, the Ruby way:
15
15
  - Ruby 3.4+ is required to run `lemans`
16
16
  - Daytona account (API token) or Docker (for local sandboxes)
17
17
  - Some LLM provider/proxy credentials (e.g., OpenRouter)
18
+ - Bash in the sandbox image for command execution
18
19
  - iproute2 in the sandbox image, so a jailed agent keeps loopback while cut off the internet
19
20
 
20
21
  ## Getting started
@@ -102,6 +103,8 @@ verifier:
102
103
  # restore: [test, bin] # [optional] paths restored from the pre-agent snapshot, so grading can't be tampered with
103
104
  ```
104
105
 
106
+ On the Docker backend, `allowlist` sends traffic through a proxy container. Only tools that use `http_proxy`/`https_proxy` can reach the listed hosts. Host names must match exactly, and IP ranges are not supported. See [the proxy README](lib/lemans/environments/docker/proxy/README.md).
107
+
105
108
  A minimal task example—checking whether an agent can write "Hello, world" into a file:
106
109
 
107
110
  - `instruction.md`: the task itself (note the preamble about the training corpora and the frontmatter)
@@ -231,6 +234,17 @@ Each trial writes a flat `runs/<model>/<task>__<id>/` directory:
231
234
 
232
235
  Every `result.json` is stamped with the `lemans_version` that wrote it, and newer lemans keeps reading runs produced by older releases. `lemans run --resume` skips trials that already have a scored result for the same agent and model.
233
236
 
237
+ `lemans restart <run>` continues a multistep run that failed for reasons outside the model, such as a provider outage. Give it the run directories or the trial ids: `lemans restart RUN [RUN...]` restarts every run under the same options, `-c` at once, and runs none if any is refused. The new run replays the agent patches of the settled steps in a fresh sandbox and starts at the next step. It copies the settled steps' usage, phases, and artifacts, and records the source in `restarted_from`. The failed run stays as it is. It restarts even if the task or the bench changed since; the new run records the current digests.
238
+
239
+ Two modes restart inside a step:
240
+
241
+ - `--recover` continues the failed step's agent session. The new run also replays the step's partial patch, and the agent goes on from the saved `agent.result.json`. The session keeps the steps, tokens, cost, and time it already spent. Only `miniswen` and `miniswen-installed` can recover.
242
+ - `--reverify` runs the last verification again with the current tests. Use it after you change the grading or the harness. The new run replays the graded step's patch and does not run its agent. If the verification was an intermediate gate and it now passes, the trial goes on to the next step. A reverification accepts scored runs.
243
+
244
+ A scored run restarts in the other modes only with `--allow-scored`.
245
+
246
+ Only the files that git tracks come back: ignored files (databases, logs, installed dependencies) are not in the patches.
247
+
234
248
  You can also run `lemans report` with various flags to see aggregated results, e.g.:
235
249
 
236
250
  ```sh
@@ -256,8 +270,9 @@ gpt-5.6-luna ar-archive-book-access 2/2 2m 23s $0.0132 12.5 156905
256
270
  | `lemans init` | Scaffold a new bench directory: an annotated `bench.yml` and two example tasks |
257
271
  | `lemans tasks` | List the tasks in a bench (`--tag` to filter) |
258
272
  | `lemans run` | Run tasks and grade them (`--task`, `--tag`, `--agent`, `--model`, `--max-output-tokens`, `-k`, `-c`, `--resume`) |
259
- | `lemans report` | Summarize `runs/` as a table or CSV (`--task`, `--tag`, `--metadata key:value` to filter, `-A [task-agent-model]` to aggregate, `-S <column>` to sort); repeated attempts add pass@k per model × task, fractional grading a `credit` column |
260
- | `lemans clobber` | Delete run results (`--task`, `--ttl 10m\|2h\|1d`, `--invalid`, `-f` to skip the confirmation) |
273
+ | `lemans restart <run>...` | Continue failed multistep runs from their last settled step in new runs (`-c`, `--recover` to continue the failed step's session, `--reverify` to grade again, `--allow-scored`, `--backend`, `--max-output-tokens`) |
274
+ | `lemans report [RUNS_DIR]` | Summarize `runs/` (or `RUNS_DIR`) as a table or CSV (`--task`, `--tag`, `--metadata key:value` to filter, `-A [task-agent-model]` to aggregate, `-S <column>` to sort); repeated attempts add pass@k per model × task, fractional grading a `credit` column |
275
+ | `lemans clobber [RUNS_DIR]` | Delete run results under `runs/` (or `RUNS_DIR`) (`--task`, `--ttl 10m\|2h\|1d`, `--invalid`, `-f` to skip the confirmation) |
261
276
 
262
277
  ## miniswen
263
278
 
@@ -1,6 +1,7 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  require "json"
4
+ require "time"
4
5
  require "miniswen/version"
5
6
  require "miniswen/ruby_llm"
6
7
 
@@ -261,25 +262,17 @@ module Miniswen
261
262
  @reporter = reporter
262
263
  end
263
264
 
264
- def run(instruction)
265
- uname = execute("uname -srvm").output.to_s.strip
266
- @messages = [
267
- { role: "system", content: SYSTEM_TEMPLATE },
268
- { role: "user", content: format(INSTANCE_TEMPLATE,
269
- instruction: instruction,
270
- system_information: uname,
271
- macos_sed_note: uname.start_with?("Darwin") ? "\n#{MACOS_SED_NOTE}" : "") }
272
- ]
273
-
274
- @steps = 0
275
- @cost = 0.0
276
-
277
- @totals = { input_tokens: 0, output_tokens: 0, cached_tokens: 0, thinking_tokens: 0 }
265
+ # With a `history` (the Result of an interrupted run), the session goes on
266
+ # where it stopped, under the budget the earlier part already spent.
267
+ def run(instruction = nil, history: nil)
268
+ if history
269
+ resume(history)
270
+ else
271
+ start(instruction)
272
+ end
278
273
 
279
- @cost_known = true
280
274
  @consecutive_format_errors = 0
281
275
  @refused_turns = 0
282
- @started_at = @clock.call
283
276
 
284
277
  loop do
285
278
  (status = limit_reached) and return finish(status)
@@ -329,6 +322,49 @@ module Miniswen
329
322
 
330
323
  private
331
324
 
325
+ def start(instruction)
326
+ uname = execute("uname -srvm").output.to_s.strip
327
+ @messages = [
328
+ { role: "system", content: SYSTEM_TEMPLATE },
329
+ { role: "user", content: format(INSTANCE_TEMPLATE,
330
+ instruction: instruction,
331
+ system_information: uname,
332
+ macos_sed_note: uname.start_with?("Darwin") ? "\n#{MACOS_SED_NOTE}" : "") }
333
+ ]
334
+
335
+ @steps = 0
336
+ @cost = 0.0
337
+ @cost_known = true
338
+ @totals = { input_tokens: 0, output_tokens: 0, cached_tokens: 0, thinking_tokens: 0 }
339
+ @started_at = @clock.call
340
+ end
341
+
342
+ def resume(history)
343
+ @messages = answered_messages(history.messages)
344
+
345
+ @steps = history.steps.to_i
346
+ @cost = history.cost_usd.to_f
347
+ @cost_known = !history.cost_usd.nil?
348
+ @totals = { input_tokens: history.input_tokens.to_i, output_tokens: history.output_tokens.to_i,
349
+ cached_tokens: history.cached_tokens.to_i, thinking_tokens: history.thinking_tokens.to_i }
350
+ @started_at = @clock.call - elapsed(history.messages)
351
+ end
352
+
353
+ # A turn whose calls did not all get a result cannot go back to a provider:
354
+ # it is dropped, and the model decides again from the tree it left.
355
+ def answered_messages(messages)
356
+ turn = messages.rindex { it[:role] == "assistant" && it[:tool_calls] }
357
+ return messages.dup unless turn
358
+
359
+ answered = messages[(turn + 1)..].filter_map { it[:tool_call_id] if it[:role] == "tool" }
360
+ messages[turn][:tool_calls].all? { answered.include?(it[:id]) } ? messages.dup : messages[0...turn]
361
+ end
362
+
363
+ def elapsed(messages)
364
+ times = messages.filter_map { Time.parse(it[:timestamp]) if it[:role] == "assistant" && it[:timestamp] }
365
+ times.empty? ? 0.0 : times.max - times.min
366
+ end
367
+
332
368
  def execute(command)
333
369
  environment.exec(command, timeout: exec_timeout, env: EXEC_ENV)
334
370
  end
data/lib/miniswen/cli.rb CHANGED
@@ -25,6 +25,7 @@ module Miniswen
25
25
  @jail = false
26
26
  @allowed_hosts = nil
27
27
  @workdir = nil
28
+ @history = nil
28
29
  end
29
30
 
30
31
  def run
@@ -56,7 +57,7 @@ module Miniswen
56
57
  agent = Agent.new(model:, reporter:, environment:, **options)
57
58
 
58
59
  begin
59
- result = agent.run(instruction)
60
+ result = agent.run(instruction, history: @history && Agent::Result.from_h(JSON.parse(@history)))
60
61
  rescue StandardError => e
61
62
  write_results(agent.partial_result(error_message(e)))
62
63
  raise
@@ -105,7 +106,7 @@ module Miniswen
105
106
 
106
107
  def parse_args!
107
108
  parser = OptionParser.new do |opts|
108
- opts.banner = "Usage: miniswen -m MODEL -p INSTRUCTION [...options]"
109
+ opts.banner = "Usage: miniswen -m MODEL (-p INSTRUCTION | --continue-from PATH) [...options]"
109
110
 
110
111
  opts.on("-m MODEL", "--model=MODEL", String,
111
112
  "LLM to use (litellm format, e.g.: openrouter/openai/gpt-5.6-luna") do |v|
@@ -116,6 +117,10 @@ module Miniswen
116
117
  @instruction = File.file?(v) ? File.read(v) : v
117
118
  end
118
119
 
120
+ opts.on("--continue-from=PATH", String, "Continue the session saved by --results-path at PATH (no -p needed)") do |v|
121
+ @history = File.read(v)
122
+ end
123
+
119
124
  opts.on("--max-steps=STEPS", Integer, "Max steps count") do |v|
120
125
  options[:max_steps] = v
121
126
  end
@@ -197,7 +202,7 @@ module Miniswen
197
202
  return if @refresh_registry
198
203
 
199
204
  raise "Use -m to specify the model" unless @model
200
- raise "Please, provide instructions via -p option" unless @instruction
205
+ raise "Please, provide instructions via -p option" unless @instruction || @history
201
206
  end
202
207
  end
203
208
  end
@@ -20,7 +20,7 @@ module Miniswen
20
20
  env&.each { |key, value| argv += [ "--env", "#{key}=#{value}" ] }
21
21
  argv << id
22
22
  argv += [ "timeout", timeout.ceil.to_s ] if timeout&.positive?
23
- argv += [ "sh", "-c", command ]
23
+ argv += [ "bash", "-c", command ]
24
24
 
25
25
  Open3.popen2e(*argv) do |stdin, pipe, wait|
26
26
  stdin.close
data/lib/miniswen/jail.rb CHANGED
@@ -61,7 +61,7 @@ module Miniswen
61
61
  def spawn_arguments(command, env)
62
62
  [ command_env(env),
63
63
  "nsenter", "--target", @holder.pid.to_s, "--net", "--mount", "--pid=/proc/#{@holder.pid}/ns/pid_for_children", "--wd=#{@workdir}",
64
- "--", "setpriv", "--bounding-set=-all", "--inh-caps=-all", "--no-new-privs", "--", "sh", "-c", command ]
64
+ "--", "setpriv", "--bounding-set=-all", "--inh-caps=-all", "--no-new-privs", "--", "bash", "-c", command ]
65
65
  end
66
66
 
67
67
  def spawn_options = super.merge(unsetenv_others: true)
@@ -28,7 +28,7 @@ module Miniswen
28
28
 
29
29
  private
30
30
 
31
- def spawn_arguments(command, env) = [ env || {}, "sh", "-c", command ]
31
+ def spawn_arguments(command, env) = [ env || {}, "bash", "-c", command ]
32
32
 
33
33
  def spawn_options = { pgroup: true }
34
34
 
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module Miniswen
4
- VERSION = "1.4.0"
4
+ VERSION = "1.5.0"
5
5
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: miniswen
3
3
  version: !ruby/object:Gem::Version
4
- version: 1.4.0
4
+ version: 1.5.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Svyatoslav Kryukov