miniswen 1.0.0.rc.1 → 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +10 -9
- data/lib/miniswen/version.rb +1 -1
- metadata +1 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 85e11349e68be75bb5c676a1ad629639d663cf117fd1bc27555522c9b2b9e4c7
|
|
4
|
+
data.tar.gz: a1c34a07733af4791e348f856326847e01183cd4f7e78d03bca08c665c86bb04
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 86f44e426ca5c3e66cfccf4defd18489c2019a54844505026e497ae0481dfe11fa68c32131fef6cc303ca968d1b423fe1e44ac34f4bc84c618e356f974033168
|
|
7
|
+
data.tar.gz: 33f3b1333f54a4a9c3f642e26ac09ac3717a8906b56e82248fcd1555e545bafca024793bd5eac05cd181baffd2f2adae5bc906a5443e13a892b4833f1a1420e3
|
data/README.md
CHANGED
|
@@ -1,11 +1,11 @@
|
|
|
1
|
-
#
|
|
1
|
+
# lemans
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
lemans is a harness for benchmarking coding agents, the Ruby way:
|
|
4
4
|
|
|
5
5
|
- **CLI-first**: the `lemans` command is all you (or your agent) need for running tasks and generating reports
|
|
6
6
|
- **Conventional**, aka _boilerplate-free_: an instruction, a Docker environment, and a test script — that's all you need to describe an eval
|
|
7
7
|
- **Trustworthy**: a grade can't be gamed by the agent, an infrastructure failure can't masquerade as a model failure, and every result carries enough digests to prove what it actually measured
|
|
8
|
-
- Powered by **
|
|
8
|
+
- Powered by **miniswen**, a Ruby version of mini-swe-agent (powered by [RubyLLM](https://rubyllm.com/), so it works with any LLM) and [Daytona](https://www.daytona.io) sandboxes.
|
|
9
9
|
|
|
10
10
|
> [!TIP]
|
|
11
11
|
> Check [Rails AI Evals](https://github.com/rails/ai-evals) for a full-featured example.
|
|
@@ -13,14 +13,14 @@ Lemans is a harness for benchmarking coding agents, the Ruby way:
|
|
|
13
13
|
## Prerequisites
|
|
14
14
|
|
|
15
15
|
- Ruby 3.4+ is required to run `lemans`
|
|
16
|
-
- Daytona account (API token)
|
|
16
|
+
- Daytona account (API token) or Docker (for local sandboxes)
|
|
17
17
|
- Some LLM provider/proxy credentials (e.g., OpenRouter)
|
|
18
18
|
|
|
19
19
|
## Getting started
|
|
20
20
|
|
|
21
21
|
### 1. Install lemans
|
|
22
22
|
|
|
23
|
-
Install
|
|
23
|
+
Install lemans CLI:
|
|
24
24
|
|
|
25
25
|
```bash
|
|
26
26
|
gem install lemans
|
|
@@ -65,6 +65,7 @@ version: 1
|
|
|
65
65
|
# A bare list is a commands shorthand: `setup: [bin/sandbox-setup]`
|
|
66
66
|
|
|
67
67
|
environment:
|
|
68
|
+
# backend: "daytona" # or "docker"
|
|
68
69
|
# dockerfile: environment/Dockerfile # the default — or pin a published image instead:
|
|
69
70
|
# image: ghcr.io/acme/my-bench@sha256:...
|
|
70
71
|
resources: { cpus: 2, memory: 2GB, storage: 5GB }
|
|
@@ -163,7 +164,7 @@ An `environment.patch` next to `instruction.md` is always applied, declared or n
|
|
|
163
164
|
### 3. Set credentials
|
|
164
165
|
|
|
165
166
|
```bash
|
|
166
|
-
export DAYTONA_API_KEY=... # or DAYTONA_TOKEN
|
|
167
|
+
export DAYTONA_API_KEY=... # or DAYTONA_TOKEN (if using Daytona)
|
|
167
168
|
export OPENROUTER_API_KEY=... # or ANTHROPIC_API_KEY, OPENAI_API_KEY, ... — matching your model
|
|
168
169
|
|
|
169
170
|
export LEMANS_PROVIDER_ORDER="Chutes" # [optional] Pin an OpenRouter model to named backends
|
|
@@ -244,9 +245,9 @@ gpt-5.6-luna ar-archive-book-access 2/2 2m 23s $0.0132 12.5 156905
|
|
|
244
245
|
| `lemans report` | Summarize `runs/` as a table or CSV (`--tag`, `-A [task-agent-model]` to aggregate, `-S <column>` to sort); repeated attempts add pass@k per model × task |
|
|
245
246
|
| `lemans clobber` | Delete run results (`--task`, `--ttl 10m\|2h\|1d`, `--invalid`, `-f` to skip the confirmation) |
|
|
246
247
|
|
|
247
|
-
##
|
|
248
|
+
## miniswen
|
|
248
249
|
|
|
249
|
-
|
|
250
|
+
miniswen can be used independently of lemans as a basic coding agent:
|
|
250
251
|
|
|
251
252
|
```sh
|
|
252
253
|
$ gem install miniswen
|
|
@@ -256,7 +257,7 @@ $ miniswen --model 'openrouter/openai/gpt-5.6-luna' --prompt 'Write hello-world
|
|
|
256
257
|
...
|
|
257
258
|
```
|
|
258
259
|
|
|
259
|
-
Currently, it's a one-shot agent that doesn't ask any questions. It's mostly useful for playing with eval ideas before encoding them as
|
|
260
|
+
Currently, it's a one-shot agent that doesn't ask any questions. It's mostly useful for playing with eval ideas before encoding them as lemans tasks.
|
|
260
261
|
|
|
261
262
|
## Development
|
|
262
263
|
|
data/lib/miniswen/version.rb
CHANGED