lemans 1.4.0 → 1.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 7ec0760a162ddb4b0fe13d4ddeb4027f934c1ffae9c7acbfda330da9ff68d91f
4
- data.tar.gz: 85cd135708fade8a06610c4cd26e4f4556e8d90ec566d3a5e12e02d13e46d351
3
+ metadata.gz: 4383d276c7923cc0853a1eb42bc179cc753769010893e61fd63dab4304aad2ef
4
+ data.tar.gz: 6805d08e1946fcc7bc7d06adb41c721daa2aeb184f2093c720a8affa2881b40c
5
5
  SHA512:
6
- metadata.gz: eb13268e0b98f9e8ca4f82f3e50b0a454366f00f0030fcaed4d03bee6d67eb214370e1317f9ffbba7701beefdb4d4e8e6c5f399cdaad90425208bdbc2b950fdc
7
- data.tar.gz: f0a8acc7bc79be0fd546fdc0c59db8e0cfa354aba6806bb09f416bf2d61b898c895a0b94a71b27c122a6a7f65854f51d07fd32f2fd0ce66d91e9e4568163bdeb
6
+ metadata.gz: d9a0993a5a5ec87224ea241ee0001c4d382b3ee4da2d401840e00275c66952844f25d2f86d504e9fdbd967c115bb59bed8f6366bb6fea33ffd3fe5ca54b0c013
7
+ data.tar.gz: 291121a4b9cafca127af023db929e3c1a5a162f197e5d306d1e45f6e26aebb63032b56d47f944189392c2cd7a9b18f563d7558a57e07719132b0555b685e0b35
data/CHANGELOG.md CHANGED
@@ -1,5 +1,15 @@
1
1
  ## [Unreleased]
2
2
 
3
+ ## [1.5.0] - 2026-10-05
4
+
5
+ - `lemans restart --recover` continues the failed step's agent session
6
+ - `lemans restart --reverify` grades the last verified step again and goes on from there; `--allow-scored` restarts scored runs.
7
+ - miniswen: `--continue-from PATH` continues a session from its `--results-path` file.
8
+ - Docker: support the `allowlist` network mode through a proxy container on an internal network.
9
+ - Support `lemans report path/to/runs`, `lemans clobber path/to/runs`, and `lemans regrade path/to/runs` (in addition to `--runs-dir path/to/runs`)
10
+ - Add `lemans restart <run>` to continue a failed multistep run from its last completed step as a new run.
11
+ - Fix(miniswen): run Bash tool commands through `bash -c` in Local, Jail, and Docker; sandbox images must include Bash.
12
+
3
13
  ## [1.4.0] - 2026-09-24
4
14
 
5
15
  - miniswen: add built-in proxy and the `--allow-hosts` switch.
data/README.md CHANGED
@@ -15,6 +15,7 @@ lemans is a harness for benchmarking coding agents, the Ruby way:
15
15
  - Ruby 3.4+ is required to run `lemans`
16
16
  - Daytona account (API token) or Docker (for local sandboxes)
17
17
  - Some LLM provider/proxy credentials (e.g., OpenRouter)
18
+ - Bash in the sandbox image for command execution
18
19
  - iproute2 in the sandbox image, so a jailed agent keeps loopback while cut off the internet
19
20
 
20
21
  ## Getting started
@@ -102,6 +103,8 @@ verifier:
102
103
  # restore: [test, bin] # [optional] paths restored from the pre-agent snapshot, so grading can't be tampered with
103
104
  ```
104
105
 
106
+ On the Docker backend, `allowlist` sends traffic through a proxy container. Only tools that use `http_proxy`/`https_proxy` can reach the listed hosts. Host names must match exactly, and IP ranges are not supported. See [the proxy README](lib/lemans/environments/docker/proxy/README.md).
107
+
105
108
  A minimal task example—checking whether an agent can write "Hello, world" into a file:
106
109
 
107
110
  - `instruction.md`: the task itself (note the preamble about the training corpora and the frontmatter)
@@ -231,6 +234,17 @@ Each trial writes a flat `runs/<model>/<task>__<id>/` directory:
231
234
 
232
235
  Every `result.json` is stamped with the `lemans_version` that wrote it, and newer lemans keeps reading runs produced by older releases. `lemans run --resume` skips trials that already have a scored result for the same agent and model.
233
236
 
237
+ `lemans restart <run>` continues a multistep run that failed for reasons outside the model, such as a provider outage. Give it the run directories or the trial ids: `lemans restart RUN [RUN...]` restarts every run under the same options, `-c` at once, and runs none if any is refused. The new run replays the agent patches of the settled steps in a fresh sandbox and starts at the next step. It copies the settled steps' usage, phases, and artifacts, and records the source in `restarted_from`. The failed run stays as it is. It restarts even if the task or the bench changed since; the new run records the current digests.
238
+
239
+ Two modes restart inside a step:
240
+
241
+ - `--recover` continues the failed step's agent session. The new run also replays the step's partial patch, and the agent goes on from the saved `agent.result.json`. The session keeps the steps, tokens, cost, and time it already spent. Only `miniswen` and `miniswen-installed` can recover.
242
+ - `--reverify` runs the last verification again with the current tests. Use it after you change the grading or the harness. The new run replays the graded step's patch and does not run its agent. If the verification was an intermediate gate and it now passes, the trial goes on to the next step. A reverification accepts scored runs.
243
+
244
+ A scored run restarts in the other modes only with `--allow-scored`.
245
+
246
+ Only the files that git tracks come back: ignored files (databases, logs, installed dependencies) are not in the patches.
247
+
234
248
  You can also run `lemans report` with various flags to see aggregated results, e.g.:
235
249
 
236
250
  ```sh
@@ -256,8 +270,9 @@ gpt-5.6-luna ar-archive-book-access 2/2 2m 23s $0.0132 12.5 156905
256
270
  | `lemans init` | Scaffold a new bench directory: an annotated `bench.yml` and two example tasks |
257
271
  | `lemans tasks` | List the tasks in a bench (`--tag` to filter) |
258
272
  | `lemans run` | Run tasks and grade them (`--task`, `--tag`, `--agent`, `--model`, `--max-output-tokens`, `-k`, `-c`, `--resume`) |
259
- | `lemans report` | Summarize `runs/` as a table or CSV (`--task`, `--tag`, `--metadata key:value` to filter, `-A [task-agent-model]` to aggregate, `-S <column>` to sort); repeated attempts add pass@k per model × task, fractional grading a `credit` column |
260
- | `lemans clobber` | Delete run results (`--task`, `--ttl 10m\|2h\|1d`, `--invalid`, `-f` to skip the confirmation) |
273
+ | `lemans restart <run>...` | Continue failed multistep runs from their last settled step in new runs (`-c`, `--recover` to continue the failed step's session, `--reverify` to grade again, `--allow-scored`, `--backend`, `--max-output-tokens`) |
274
+ | `lemans report [RUNS_DIR]` | Summarize `runs/` (or `RUNS_DIR`) as a table or CSV (`--task`, `--tag`, `--metadata key:value` to filter, `-A [task-agent-model]` to aggregate, `-S <column>` to sort); repeated attempts add pass@k per model × task, fractional grading a `credit` column |
275
+ | `lemans clobber [RUNS_DIR]` | Delete run results under `runs/` (or `RUNS_DIR`) (`--task`, `--ttl 10m\|2h\|1d`, `--invalid`, `-f` to skip the confirmation) |
261
276
 
262
277
  ## miniswen
263
278
 
data/exe/lemans-remote CHANGED
@@ -29,6 +29,13 @@
29
29
  #
30
30
  # exe/lemans-remote run --bench ../ai-evals --retry-runs=./runs
31
31
  #
32
+ # Restart local runs remotely, with the options of `lemans restart` — one
33
+ # sandbox per run (--run-in-band for a single sandbox, --sync to wait and
34
+ # download). The runs travel with the bench; the new runs come back via
35
+ # pull-runs (the shipped copies do not):
36
+ #
37
+ # exe/lemans-remote restart --bench ../ai-evals runs/b/model/task__AbC1234 task__XyZ5678 --reverify
38
+ #
32
39
  # Then watch, fetch, and clean up:
33
40
  #
34
41
  # exe/lemans-remote status [--history] [--running | --complete] [-W [INTERVAL]]
@@ -81,6 +88,7 @@ module LemansRemote # :nodoc: all
81
88
  REMOTE_BENCH_DIR = "/task"
82
89
  REMOTE_BENCH_ARCHIVE = "/tmp/bench.tgz"
83
90
  REMOTE_RUNS_ARCHIVE = "/tmp/runs.tgz"
91
+ REMOTE_SOURCES_ARCHIVE = "/tmp/sources.tgz"
84
92
  REMOTE_LOG = "/tmp/lemans-remote-run.log"
85
93
  REMOTE_STATUS = "/tmp/lemans-remote-run.status"
86
94
  REMOTE_WRAPPER = "/tmp/lemans-remote-wrapper.sh"
@@ -546,8 +554,11 @@ module LemansRemote # :nodoc: all
546
554
  TIMED_OUT = 124
547
555
  CREATE_ATTEMPTS = 3
548
556
 
557
+ # `command` replaces the `lemans run` built from tasks/models/args; `sources`
558
+ # are run dirs (relative to `runs_dir`) shipped into the sandbox's runs/
559
+ # for the command to read, and dropped again before the results leave it.
549
560
  def initialize(bench:, snapshot:, run_id:, tasks:, models:, extra_args:, extra_env:,
550
- timeout_sec:, runs_dir:, keep:, sync:, shell:, launch_interval: 0)
561
+ timeout_sec:, runs_dir:, keep:, sync:, shell:, launch_interval: 0, command: nil, sources: [])
551
562
  @bench = bench
552
563
  @snapshot = snapshot
553
564
  @run_id = run_id
@@ -561,6 +572,8 @@ module LemansRemote # :nodoc: all
561
572
  @sync = sync
562
573
  @shell = shell
563
574
  @launch_interval = launch_interval
575
+ @command = command
576
+ @sources = sources
564
577
  @started_at = Time.now.utc.iso8601
565
578
  end
566
579
 
@@ -666,6 +679,31 @@ module LemansRemote # :nodoc: all
666
679
  exec!(sandbox, "mkdir -p #{REMOTE_BENCH_DIR} && " \
667
680
  "tar -xzf #{REMOTE_BENCH_ARCHIVE} -C #{REMOTE_BENCH_DIR} && " \
668
681
  "rm -f #{REMOTE_BENCH_ARCHIVE}")
682
+ upload_sources(sandbox) if @sources.any?
683
+ end
684
+
685
+ def upload_sources(sandbox)
686
+ say :upload, "#{@sources.size} run(s) from #{@runs_dir}"
687
+ Dir.mktmpdir("lemans-remote") do |tmp|
688
+ archive = File.join(tmp, "sources.tgz")
689
+ _, err, status = Open3.capture3("tar", "-czf", archive, "-C", @runs_dir.to_s, *@sources)
690
+ raise "could not pack the runs: #{err}" unless status.success?
691
+
692
+ File.open(archive, "rb") do |file|
693
+ sandbox.fs.upload_file_stream(file, REMOTE_SOURCES_ARCHIVE, timeout: TRANSFER_TIMEOUT_SEC)
694
+ end
695
+ end
696
+ exec!(sandbox, "mkdir -p #{REMOTE_BENCH_DIR}/runs && " \
697
+ "tar -xzf #{REMOTE_SOURCES_ARCHIVE} -C #{REMOTE_BENCH_DIR}/runs && " \
698
+ "rm -f #{REMOTE_SOURCES_ARCHIVE}")
699
+ end
700
+
701
+ # The shipped runs are copies of local ones: archiving them back would
702
+ # resurrect whatever was clobbered locally in the meantime.
703
+ def drop_sources_command
704
+ return "true" if @sources.empty?
705
+
706
+ "(cd #{REMOTE_BENCH_DIR}/runs && rm -rf -- #{@sources.shelljoin})"
669
707
  end
670
708
 
671
709
  def execute(sandbox)
@@ -683,6 +721,8 @@ module LemansRemote # :nodoc: all
683
721
  end
684
722
 
685
723
  def remote_command
724
+ return @command if @command
725
+
686
726
  command = %w[lemans run --bench .]
687
727
  @tasks.each { command += [ "--task", it ] }
688
728
  @models.each { command += [ "--model", it ] }
@@ -718,6 +758,7 @@ module LemansRemote # :nodoc: all
718
758
  "tasks" => @tasks,
719
759
  "models" => @models,
720
760
  "args" => @extra_args,
761
+ "command" => remote_command,
721
762
  "lemans_version" => Lemans::VERSION,
722
763
  "snapshot" => @snapshot,
723
764
  "sandbox_id" => sandbox.id,
@@ -735,6 +776,7 @@ module LemansRemote # :nodoc: all
735
776
  (cd #{REMOTE_BENCH_DIR} && #{remote_command}) > #{REMOTE_LOG} 2>&1
736
777
  status=$?
737
778
  cp #{REMOTE_LOG} #{vault_dir}/run.log || true
779
+ #{drop_sources_command}
738
780
  if [ -d #{REMOTE_BENCH_DIR}/runs ]; then
739
781
  tar -czf #{vault_dir}/runs.tar.gz -C #{REMOTE_BENCH_DIR}/runs .
740
782
  fi
@@ -827,7 +869,7 @@ module LemansRemote # :nodoc: all
827
869
  return
828
870
  end
829
871
 
830
- exec!(sandbox, "tar -czf #{REMOTE_RUNS_ARCHIVE} -C #{REMOTE_BENCH_DIR}/runs .")
872
+ exec!(sandbox, "#{drop_sources_command} && tar -czf #{REMOTE_RUNS_ARCHIVE} -C #{REMOTE_BENCH_DIR}/runs .")
831
873
  FileUtils.mkdir_p(@runs_dir)
832
874
  Dir.mktmpdir("lemans-remote") do |tmp|
833
875
  local = File.join(tmp, "runs.tgz")
@@ -975,6 +1017,67 @@ module LemansRemote # :nodoc: all
975
1017
  raise Thor::Error, "lemans-remote: #{e.message}"
976
1018
  end
977
1019
 
1020
+ desc "restart RUN...", "Restart local runs on remote Daytona sandboxes, the way `lemans restart` does (one sandbox per run)"
1021
+ option :bench, default: ".", desc: "Directory holding bench.yml"
1022
+ option :runs_dir, default: "runs", desc: "Local directory holding the runs (sync mode: and receiving the new ones)"
1023
+ option :recover, type: :boolean, default: false, desc: "Continue the failed step's agent session from its saved history"
1024
+ option :reverify, type: :boolean, default: false,
1025
+ desc: "Grade the last verified step again (with the current tests) and go on from there"
1026
+ option :allow_scored, type: :boolean, default: false, desc: "Restart a scored run (--reverify always may)"
1027
+ option :max_output_tokens, type: :numeric, banner: "TOKENS", desc: "Cap the agent's output per model call"
1028
+ option :concurrency, type: :numeric, aliases: "-c", desc: "Restarts in flight at once inside a sandbox (with --run-in-band)"
1029
+ option :timeout, default: "6h", desc: "Give up on the remote run after this long"
1030
+ option :sync, type: :boolean, default: false, desc: "Restart every run in one sandbox, wait, and download the results"
1031
+ option :run_in_band, type: :boolean, default: false, aliases: %w[--runInBand],
1032
+ desc: "Async mode: restart every run in one sandbox instead of one sandbox per run"
1033
+ option :keep, type: :boolean, default: false, desc: "Sync mode: leave the sandbox around for debugging"
1034
+ option :env, repeatable: true, desc: "Forward an extra host ENV variable by name"
1035
+ option :launch_interval, type: :numeric, default: 3,
1036
+ desc: "Async mode: seconds between consecutive launches (0 to disable)"
1037
+ def restart(*runs)
1038
+ raise Thor::Error, "lemans-remote: name the run(s) to restart" if runs.empty?
1039
+ raise Thor::Error, "lemans-remote: --recover and --reverify exclude each other" if options[:recover] && options[:reverify]
1040
+ raise Thor::Error, "lemans-remote: --launch-interval must not be negative" if options[:launch_interval].to_f.negative?
1041
+
1042
+ bench = Lemans::Config.load_file(options[:bench])
1043
+ runs_root = Pathname(options[:runs_dir])
1044
+ store = Lemans::Stores::FS.new(runs_root)
1045
+ ids = runs.map { File.basename(it) }.uniq
1046
+ found = store.fetch.select { ids.include?(it.id) }.to_h { [ it.id, it ] }
1047
+ missing = ids - found.keys
1048
+ raise Thor::Error, "lemans-remote: no run #{missing.join(", ")} under #{runs_root}" if missing.any?
1049
+
1050
+ sources = found.values_at(*ids)
1051
+
1052
+ # The sandbox would refuse them anyway: refuse here, before any sandbox starts
1053
+ bench.load_options(model: sources.map(&:model).uniq)
1054
+ Lemans::Runner.new(bench, bench.tasks, store:, restarts: sources, restart_mode: restart_mode,
1055
+ allow_scored: options[:allow_scored]).attempts
1056
+
1057
+ provisioner = build_provisioner
1058
+ raise Thor::Error, "lemans-remote: snapshot #{provisioner.name} not found — run `lemans-remote provision` first" unless provisioner.provisioned?
1059
+
1060
+ jobs = options[:sync] || options[:run_in_band] ? [ sources ] : sources.map { [ it ] }
1061
+ Vault.ensure unless options[:sync]
1062
+
1063
+ say_status :fanout, "#{jobs.size} sandboxes, one per run, one every #{options[:launch_interval]}s" if jobs.size > 1
1064
+ failures = []
1065
+ jobs.each_with_index do |job, index|
1066
+ sleep options[:launch_interval].to_f if index.positive? && options[:launch_interval].to_f.positive?
1067
+ exit_code = launch_restart(bench, job, runs_root, provisioner)
1068
+ exit exit_code if options[:sync] && !exit_code.zero?
1069
+ rescue StandardError => e
1070
+ raise if options[:sync]
1071
+
1072
+ failures << "#{job.map(&:id).join(",")} (#{e.message})"
1073
+ end
1074
+
1075
+ say_status :detached, "`lemans-remote status` to watch, `lemans-remote pull-runs` to fetch results" unless options[:sync]
1076
+ raise Thor::Error, "lemans-remote: failed to launch: #{failures.join("; ")}" if failures.any?
1077
+ rescue Lemans::ConfigError, RuntimeError => e
1078
+ raise Thor::Error, "lemans-remote: #{e.message}"
1079
+ end
1080
+
978
1081
  desc "status", "List lemans-remote runs (--history adds completed runs whose sandboxes are gone)"
979
1082
  option :history, type: :boolean, default: false, desc: "Also read the vault manifests (spins a short-lived helper sandbox)"
980
1083
  option :running, type: :boolean, default: false, desc: "Only show runs still in flight"
@@ -1282,8 +1385,43 @@ module LemansRemote # :nodoc: all
1282
1385
  [ failures, launched ]
1283
1386
  end
1284
1387
 
1285
- def build_run_id(bench, tasks, models, attempt = nil)
1388
+ def restart_mode = (:recover if options[:recover]) || (:reverify if options[:reverify])
1389
+
1390
+ def launch_restart(bench, sources, runs_root, provisioner)
1391
+ tasks = sources.map(&:task).uniq
1392
+ models = sources.map(&:model).uniq
1393
+ Runner.new(
1394
+ bench: bench,
1395
+ snapshot: provisioner.name,
1396
+ run_id: build_run_id(bench, tasks, models, kind: "restart"),
1397
+ tasks: tasks,
1398
+ models: models,
1399
+ extra_args: "",
1400
+ extra_env: options[:env] || [],
1401
+ timeout_sec: seconds!(options[:timeout]).to_i,
1402
+ runs_dir: runs_root.to_s,
1403
+ keep: options[:keep],
1404
+ sync: options[:sync],
1405
+ shell: shell,
1406
+ launch_interval: options[:launch_interval].to_f,
1407
+ command: restart_command(sources),
1408
+ sources: sources.map { |source| runs_root.glob("**/#{source.id}").find(&:directory?).relative_path_from(runs_root).to_s }
1409
+ ).call
1410
+ end
1411
+
1412
+ def restart_command(sources)
1413
+ argv = [ "lemans", "restart", *sources.map(&:id), "--bench", ".", "--runs-dir", "runs" ]
1414
+ argv << "--recover" if options[:recover]
1415
+ argv << "--reverify" if options[:reverify]
1416
+ argv << "--allow-scored" if options[:allow_scored]
1417
+ argv += [ "--max-output-tokens", options[:max_output_tokens].to_i.to_s ] if options[:max_output_tokens]
1418
+ argv += [ "-c", options[:concurrency].to_i.to_s ] if options[:concurrency]
1419
+ argv.shelljoin
1420
+ end
1421
+
1422
+ def build_run_id(bench, tasks, models, attempt = nil, kind: nil)
1286
1423
  parts = [ bench.root.expand_path.basename.to_s ]
1424
+ parts << kind if kind
1287
1425
  parts << (tasks.size == 1 ? tasks.first : "#{tasks.size}tasks") if tasks.any?
1288
1426
  if models.any?
1289
1427
  short = models.first.split("/").last