RubyGems - ruby-skill-bench - Versions diffs - 1.1.0 → 1.2.0 - Mend

ruby-skill-bench 1.1.0 → 1.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.

Files changed (60) hide show

checksums.yaml +4 -4
data/README.md +166 -35
data/docs/architecture.md +3 -1
data/docs/first-eval-guide.md +7 -7
data/docs/testing-guide.md +1 -1
data/lib/skill_bench/agent/react_agent/loop_runner.rb +44 -9
data/lib/skill_bench/agent/react_agent/step.rb +7 -1
data/lib/skill_bench/cli/batch_result_printer.rb +45 -0
data/lib/skill_bench/cli/eval/eval_options.rb +4 -0
data/lib/skill_bench/cli/help_printer.rb +10 -2
data/lib/skill_bench/cli/init_command.rb +2 -1
data/lib/skill_bench/cli/result_printer.rb +1 -1
data/lib/skill_bench/cli/run_command.rb +47 -9
data/lib/skill_bench/cli/validate_command.rb +242 -0
data/lib/skill_bench/cli.rb +3 -0
data/lib/skill_bench/client.rb +43 -1
data/lib/skill_bench/clients/all.rb +2 -0
data/lib/skill_bench/clients/base_client.rb +12 -1
data/lib/skill_bench/clients/base_url_validator.rb +105 -0
data/lib/skill_bench/clients/provider_config.rb +34 -1
data/lib/skill_bench/clients/provider_schemas.rb +4 -0
data/lib/skill_bench/clients/providers/mistral.rb +47 -0
data/lib/skill_bench/commands/init.rb +5 -0
data/lib/skill_bench/commands/skill_new.rb +3 -1
data/lib/skill_bench/config/applier.rb +2 -0
data/lib/skill_bench/config/defaults.rb +2 -0
data/lib/skill_bench/config/facade_readers.rb +7 -0
data/lib/skill_bench/config/facade_writers.rb +17 -0
data/lib/skill_bench/config/json_loader.rb +1 -1
data/lib/skill_bench/config/store.rb +29 -0
data/lib/skill_bench/config.rb +18 -0
data/lib/skill_bench/evaluation/runner.rb +20 -3
data/lib/skill_bench/execution/context_hydrator.rb +52 -11
data/lib/skill_bench/execution/sandbox.rb +58 -11
data/lib/skill_bench/judge/judge.rb +4 -0
data/lib/skill_bench/judge/prompt.rb +42 -6
data/lib/skill_bench/models/config.rb +32 -0
data/lib/skill_bench/output_formatter.rb +60 -1
data/lib/skill_bench/package_verifier.rb +1 -1
data/lib/skill_bench/rails/skill_templates.rb +19 -5
data/lib/skill_bench/services/agent_spawner_service.rb +7 -3
data/lib/skill_bench/services/batch_runner_service.rb +111 -0
data/lib/skill_bench/services/compare_option_parser.rb +1 -0
data/lib/skill_bench/services/cost_calculator.rb +91 -0
data/lib/skill_bench/services/html_formatter.rb +289 -0
data/lib/skill_bench/services/json_formatter.rb +19 -1
data/lib/skill_bench/services/junit_formatter.rb +74 -24
data/lib/skill_bench/services/provider_resolver.rb +5 -2
data/lib/skill_bench/services/response_cache.rb +130 -0
data/lib/skill_bench/services/runner_service.rb +88 -4
data/lib/skill_bench/services/summary_formatter.rb +90 -0
data/lib/skill_bench/services/template_registry.rb +43 -9
data/lib/skill_bench/services/trend_recorder_service.rb +29 -2
data/lib/skill_bench/tools/registry.rb +29 -3
data/lib/skill_bench/tools/run_command.rb +171 -19
data/lib/skill_bench/trend_tracker/persistence.rb +27 -10
data/lib/skill_bench/trend_tracker.rb +5 -5
data/lib/skill_bench/version.rb +1 -1
data/lib/skill_bench.rb +2 -3
metadata +17 -36

checksums.yaml CHANGED Viewed

@@ -1,7 +1,7 @@
 ---
 SHA256:
-  metadata.gz: d2ad524e13bc006a56f0197d07b3ba7b0ce2f99f60b61f0739c3d5bc0d75a687
-  data.tar.gz: a920c473148b52584653acbb1e91cb3973791c09de6c4df994a77c097eabc476
+  metadata.gz: f47976b55f6f8c147adb4ed784ce04ba52ff71f805f8e35d797ba776021641c4
+  data.tar.gz: c2febaadbdeb7e149041661258ce84e41499121445cf726cefece642e60174a4
 SHA512:
-  metadata.gz: c1f131af9bcde90e7fc3a7e6bef7f3770edfa4e2826ee19c3aabf5c210d6d3b6e5bdd460778a87f6fdc77b5b99bc17b2225e1b79de31674ad4acfe1bbc89f862
-  data.tar.gz: d8e3791c91242b25779afa3a21c57daeb06995bc4c65b01ca6f378a69491aeec95c5844ea541e5dcaf18b46d8e7f153ffb38ff2bb060ab5dacd653c8c1026bcd
+  metadata.gz: 5ab3082fa715a0776455a88b28e2d990d3bb7e52fbc5f0cc47176e1d44e76cdc430531bc0401abab350cb4c2321d0b625ce78d381094fe404cdedd7e61b27227
+  data.tar.gz: 3d5f67b876457691e003e62a8ba57fc04ac21e002bb5c6a84d1ae7954cf7dd7136f05a658d7ee4d416ab2db19cb0ad90e51dd0d920df9ffa998cb8391db65df5

data/README.md CHANGED Viewed

@@ -30,7 +30,7 @@ See the [Ecosystem Overview](https://github.com/igmarin/agent-mcp-runtime/blob/m
 - **Isolated Git Sandboxes**: Every run operates in a temporary repo. Clean diffs, zero side-effects, 100% reproducibility.
 - **Blind Judging with Dimensions**: LLM judge scores baseline and context independently across 5 canonical dimensions (Correctness, Skill Adherence, Code Quality, Test Coverage, Documentation). Eval authors configure weights and thresholds via `criteria.json`.
 - **Sophisticated ReAct Loop**: Employs a robust `Thought → Tool → Observation` loop to handle complex, multi-step engineering tasks.
-- **Multi-Provider Ecosystem**: Native support for **OpenAI**, **Anthropic**, **Google Gemini**, **Azure OpenAI**, **Ollama**, **Groq**, **DeepSeek**, and **OpenCode**.
+- **Multi-Provider Ecosystem**: Native support for **OpenAI**, **Anthropic**, **Google Gemini**, **Azure OpenAI**, **Ollama**, **Groq**, **DeepSeek**, **Mistral**, and **OpenCode**.
 - **Standardized Intelligence**: Consistent reporting format regardless of the underlying LLM provider.
 ---
@@ -64,11 +64,14 @@ CLI / API → RunnerService → Sandbox + ReAct Agent → LLM Client Layer → P
 | **Ollama** | — | `:ollama` |
 | **Groq** | `SKILL_BENCH_GROQ_API_KEY` | `:groq` |
 | **DeepSeek** | `SKILL_BENCH_DEEPSEEK_API_KEY` | `:deepseek` |
+| **Mistral** | `SKILL_BENCH_MISTRAL_API_KEY` | `:mistral` |
 | **OpenCode** | `SKILL_BENCH_OPENCODE_API_KEY`, `SKILL_BENCH_OPENCODE_BASE_URL` | `:opencode` |
 > **Note:** Environment variables are loaded automatically. You can also configure provider settings in `skill-bench.json` (created by `skill-bench init`).
 >
 > **OpenCode requires a custom `base_url`:** OpenCode does not host a public LLM API. You must provide your own OpenAI-compatible endpoint (e.g. a LiteLLM proxy, self-hosted vLLM, or company gateway) via the `base_url` config key. Without it, the provider will fail with "Base URL not set for Opencode".
+>
+> **Mistral** uses Mistral's OpenAI-compatible chat completions API (default model `mistral-large-latest`). Set `SKILL_BENCH_MISTRAL_API_KEY` and scaffold it with `skill-bench init --mistral`.
 ### Command Allowlist
@@ -79,6 +82,7 @@ By default, no shell commands are permitted. You must configure `allowed_command
   "provider": "openai",
   "max_execution_time": 30,
   "allowed_commands": ["rspec", "bundle", "ruby", "git"],
+  "allow_host_execution": false,
   "config": {
     "api_key": null,
     "model": "gpt-4o"
@@ -87,6 +91,8 @@ By default, no shell commands are permitted. You must configure `allowed_command
 ```
 > **Security:** The agent can only execute commands on this list. Dangerous commands (bash, curl, sudo, etc.) are always blocked regardless of configuration.
+>
+> **Where commands run:** Allowed commands run inside a temporary git **sandbox directory** on the host — a copy of your eval files, not your project. True container isolation (Docker) is **not yet shipped**, so the sandbox directory is the only boundary. Because of this, host execution **fails closed**: it is disabled by default and must be explicitly enabled with `"allow_host_execution": true`. With it disabled (the default), `run_command` refuses to execute and returns an error instead of running un-isolated. Enable it only when you accept that allowed commands run directly on your machine.
 ### Configuration Hierarchy
@@ -137,7 +143,9 @@ skill-bench init --openai
 }
 ```
-**Available providers:** `--openai`, `--anthropic`, `--gemini`, `--ollama`, `--azure`, `--groq`, `--deepseek`, `--opencode`
+**Available providers:** `--openai`, `--anthropic`, `--gemini`, `--ollama`, `--azure`, `--groq`, `--deepseek`, `--mistral`, `--opencode`
+**Zero-config offline path:** `skill-bench init --mock` scaffolds a minimal offline config that needs no API key and no network — `{"provider":"mock","max_execution_time":30}`. Use it to try the full flow (and run the bundled examples) before wiring up a real provider.
 Use `--force` to overwrite an existing config.
@@ -338,7 +346,7 @@ skill-bench run my-first-eval --skill=my-service
 3. **Context run** — Agent receives `task.md` + `SKILL.md` as prompt → produces output B
 4. **Blind judging** — LLM judge scores output A and output B independently across the dimensions defined in `criteria.json`
 5. **Delta computation** — Compare scores, compute deltas, apply pass/fail logic
-6. **History recording** — Store result in `.skill-bench-history.json` for trend tracking
+6. **History recording** — Store result in `.skill-bench-trends.json` for trend tracking
 Provider is read from `skill-bench.json` — no `--provider` flag needed.
@@ -350,11 +358,54 @@ skill-bench run my-first-eval --skill=skill-a --skill=skill-b
 Both skill contexts are concatenated and sent to the agent. The judge evaluates whether the combined context improves results.
-**Output Formats:**
+**Output Formats:** `--format human` (default), `json`, `junit`, or `html`.
+- Human-readable (default) — full delta table, iteration timeline, and a `Tokens: N | Est. Cost: $X.XXXX` line.
+- JSON: `--format json` — machine-readable, including top-level `tokens` and `cost` fields.
+- JUnit XML: `--format junit` — for CI test reporting.
+- HTML: `--format html` — a self-contained, shareable report (styles inlined, no external assets) with the delta table and iteration timeline. Redirect it to a file:
+  ```bash
+  skill-bench run my-first-eval --skill=my-service --format html > report.html
+  ```
+---
+## Pre-flight Checks: `validate` / `doctor`
+Before spending tokens on a run, sanity-check your setup. `skill-bench validate` (aliased as `doctor`) runs read-only pre-flight checks — it never runs an eval and never makes a network call:
+```bash
+skill-bench validate
+# or, identically:
+skill-bench doctor
+```
+It runs three checks and prints a `PASS` / `FAIL` / `SKIP` line for each:
+1. **criteria** — validates the criteria JSON (default `criteria.json`, override with `--criteria PATH`). Skipped if the default file is absent.
+2. **config** — schema-checks `skill-bench.json` (default, override with `--config PATH`): `provider` is required and must be a known provider, `max_execution_time` must be a positive integer, and `config` (when present) must be an object.
+3. **provider key** — reports whether the configured provider's API key is present (the `mock` provider needs none).
-- Human-readable (default)
-- JSON: `--format json`
-- JUnit XML: `--format junit`
+A passing report exits `0`:
+```text
+skill-bench validate
+[PASS] criteria      criteria.json is valid
+[PASS] config        skill-bench.json matches the expected shape
+[PASS] provider key  openai credentials present
+All checks passed.
+```
+A failure exits non-zero and names what is wrong:
+```text
+[FAIL] provider key  openai is missing: api_key
+1 check(s) failed.
+```
 ---
@@ -427,6 +478,22 @@ The `--variant` spec supports two forms:
 - `pack:<name>` — resolve via registry manifest
 - `/absolute/path` or `relative/path` — use a direct path
+### Response Caching (opt-in, `--cache`)
+LLM responses can be cached so identical requests reuse a previous result instead of calling the provider again. Caching is **off by default**. Enable it per run with `--cache`, or set the `SKILL_BENCH_CACHE` environment variable to a truthy value (`1`, `true`, `yes`, or `on`):
+```bash
+# Per-run flag
+skill-bench run my-first-eval --skill=my-service --cache
+# Or via the environment
+SKILL_BENCH_CACHE=1 skill-bench run my-first-eval --skill=my-service
+```
+The cache is in-memory (process-lifetime) and content-addressed: the key is a SHA-256 digest of the provider, model, system prompt, messages, tools, and temperature, so only truly identical requests share an entry. The `mock` and null providers are never cached.
+This pays off most with `compare`, which runs the skill-less baseline twice with identical inputs — with caching enabled, the repeated baseline reuses the cached response instead of making a second call.
 ---
 ## File Reference: What Lives on Disk
@@ -446,6 +513,7 @@ SkillBench creates and manages three files in your project. Understanding them h
   "provider": "openai",
   "max_execution_time": 300,
   "allowed_commands": ["rspec", "bundle", "ruby", "git"],
+  "allow_host_execution": false,
   "config": {
     "api_key": "sk-...",
     "model": "gpt-4o",
@@ -458,10 +526,11 @@ SkillBench creates and manages three files in your project. Understanding them h
 - Configuration is loaded in this order: **code defaults** → `~/.skill-bench.json` (user-wide) → `./skill-bench.json` (local) → **environment variables**. Later sources override earlier ones.
 - If `api_key` is `null`, SkillBench looks for the matching environment variable (e.g. `SKILL_BENCH_OPENAI_API_KEY`).
 - `allowed_commands` is a **safeguard**, not a convenience. By default the agent cannot run *any* shell command. Add only what your evals need.
+- `allow_host_execution` (default `false`) gates whether `run_command` may run on the host when no container isolation is active. Since container isolation is not yet shipped, leaving it `false` means `run_command` **fails closed** (refuses to execute). Set it to `true` only if you accept that allowed commands run directly on your machine inside the temporary sandbox directory.
 ---
-### `.skill-bench-history.json` — Evaluation History (Auto-Generated)
+### `.skill-bench-trends.json` — Evaluation History (Auto-Generated)
 **What it is:** A JSON array that records every successful eval run. SkillBench appends to it automatically. It stores the timestamp, eval name, skill names, scores, and deltas so you can track improvement over time.
@@ -497,13 +566,13 @@ TREND: baseline ↑ (+2), context ↑ (+7)
 The trend compares the current run against the *previous run of the same eval + skill*. This tells you at a glance whether your latest skill edit made things better or worse.
-**Pro tip:** Commit `.skill-bench-history.json` to git if you want to share trend data with your team. Add it to `.gitignore` if you prefer to keep scores private.
+**Pro tip:** `.skill-bench-trends.json` is git-ignored by default (via the `.skill-bench-trends.json*` line in `.gitignore`). If you want to share trend data with your team, remove that line so the file can be committed.
 ---
-### `.skill-bench-history.json.bak` — Backup (Auto-Generated)
+### `.skill-bench-trends.json.bak` — Backup (Auto-Generated)
-**What it is:** A copy of `.skill-bench-history.json` created every time SkillBench writes a new entry. If the main file gets corrupted (e.g. you kill the process mid-write), SkillBench automatically falls back to the `.bak` file.
+**What it is:** A snapshot of the *previous* good version of `.skill-bench-trends.json`, copied just before each new write. (The first run has no prior version yet, so it creates no `.bak`.) If the main file gets corrupted (e.g. you kill the process mid-write), SkillBench automatically falls back to the `.bak` file.
 **Who edits it:** Nobody. It is a safety net.
@@ -541,7 +610,7 @@ Read the output carefully. Look at **two things:**
 ### Step 3: Inspect the History
 ```bash
-cat .skill-bench-history.json | jq '.[-1]'
+cat .skill-bench-trends.json | jq '.[-1]'
 ```
 This shows the latest entry. Focus on the dimension with the smallest delta — that is where your skill is weakest.
@@ -729,6 +798,7 @@ These 5 dimensions are **mandatory** in every `criteria.json`. You can add custo
   Eval: my-first-eval
   Skill: my-service
   Provider: openai
+  Tokens: 18432  |  Est. Cost: $0.0934
 ═══════════════════════════════════════════════════════
   === BASELINE ITERATIONS ===
@@ -774,8 +844,9 @@ These 5 dimensions are **mandatory** in every `criteria.json`. You can add custo
 - **CONTEXT:** The agent's score *with* the skill. This is the "aided" performance.
 - **DELTA:** `CONTEXT - BASELINE`. How much the skill helped.
 - **TOTAL:** Sum of all dimension scores. Max possible is 100.
-- **TREND:** Comparison against the previous run of the same eval + skill (from `.skill-bench-history.json`). Shows whether scores are improving over time.
+- **TREND:** Comparison against the previous run of the same eval + skill (from `.skill-bench-trends.json`). Shows whether scores are improving over time.
 - **VERDICT:** `PASS` only if `CONTEXT >= pass_threshold` AND `DELTA >= minimum_delta`.
+- **Tokens / Est. Cost:** The header shows total tokens used across the run and an estimated USD cost as `Tokens: N | Est. Cost: $X.XXXX`. The cost is approximate — it comes from a built-in per-model price table (`Services::CostCalculator`) and shows `—` when the model isn't in that table. JSON output (`--format json`) exposes the same data as top-level `tokens` and `cost` fields.
 **Iteration timeline:**
@@ -827,7 +898,7 @@ Your eval result depends on **both** conditions. Here is every scenario:
 ## Reliability & Security
-- **Safe-by-Design**: No code execution occurs on the host system; everything happens in the sandbox.
+- **Allowlist-Gated Execution**: The agent can only run commands you add to `allowed_commands`; with an empty allowlist it can run nothing. Commands run inside a temporary git sandbox **directory** (a copy of the eval files) on the host — container isolation is not yet shipped, so host execution is **disabled by default** and must be explicitly opted into with `allow_host_execution: true`.
 - **Command Blocklist**: Dangerous commands (`bash`, `sh`, `python`, `curl`, etc.) are always blocked, even if listed in `allowed_commands`.
 - **Path Validation**: Eval paths are validated to prevent directory traversal attacks.
 - **Atomic History Writes**: Benchmark history uses file locking to prevent corruption from concurrent writes.
@@ -836,7 +907,7 @@ Your eval result depends on **both** conditions. Here is every scenario:
 - **Traceability**: Every thought and tool call is logged with full backtrace for post-mortem analysis.
 - **Robust Error Recovery**: Handles provider outages and rate limits gracefully with standardized error logging.
 - **XML-Safe Output**: JUnit XML output is properly escaped to prevent injection attacks.
-- **Test Coverage**: 373+ tests covering core engine, CLI commands, and all provider clients.
+- **Test Coverage**: 700+ tests covering core engine, CLI commands, and all provider clients. Run `bundle exec rake test` to see the current count.
 ## Testing
@@ -855,9 +926,18 @@ bundle exec ruby -Itest test/integration_test.rb
 **Test Structure:**
-- `test/evaluator/` — Core evaluation engine tests
-- `test/agent_eval/` — CLI, models, and service tests
+- `test/agent/` — Agent runtime tests
+- `test/agent_eval/` — Agent evaluation tests
+- `test/cli/` — CLI command tests
 - `test/clients/` — Provider client tests
+- `test/evaluator/` — Core evaluation engine tests
+- `test/history_recorder/` — Benchmark history persistence tests
+- `test/models/` — Domain model tests
+- `test/registry/` — Skill/eval registry tests
+- `test/services/` — Service layer tests
+- `test/skills/` — Skill loading tests
+- `test/tools/` — Agent tool tests
+- Plus several top-level `test/*_test.rb` files (e.g. `integration_test.rb`, `evaluation_runner_test.rb`, `trend_tracker_test.rb`).
 ---
@@ -886,11 +966,17 @@ Ruby Skill Bench is designed with security as a primary concern. The system exec
 - **Command Allowlist:** Only explicitly allowed commands can be executed
 - **Dangerous Commands Blocklist:** Dangerous commands (bash, curl, sudo, etc.) are always blocked
 - **Shell Tokenization:** Commands are tokenized before execution to prevent shell injection
-- **Docker Isolation:** Commands can be executed in isolated Docker containers with hardened security settings
+- **Fail-Closed Host Execution:** Container isolation is not yet active, so commands run on the host inside a temporary sandbox directory. To match this reality, `run_command` refuses to execute unless `allow_host_execution: true` is set; it is **disabled by default**.
+> **The allowlist is the only real authorization control — and it only checks the base command.** `run_command` authorizes by the first token of the command (`rake`, `find`, `git`, …); it does **not** inspect arguments. Shell tokenization stops metacharacter injection, but it does **not** sandbox what an allowlisted binary can do. Because many common tools are general-purpose execution wrappers, **allowlisting any one of them is equivalent to granting arbitrary host code execution** — for example `rake -e '...'`, `rspec -e`, `make` (arbitrary recipes), `find . -exec ...`, or `git` (hooks, `-c core.fsmonitor=...`, `! ...` aliases). Combined with the fail-closed model above (`run_command` refuses to run on the host unless `allow_host_execution` is explicitly enabled — see `HOST_EXECUTION_REFUSED` in `run_command.rb`), the practical guidance is: **keep `allowed_commands` as minimal as possible — empty for untrusted skills** — and treat every entry as if you were handing the skill a shell.
+>
+> An **optional, default-off** `command_argument_constraints` setting can refuse commands whose arguments contain configured substrings/flags (for example blocking `-e` or `-exec`). It is a defense-in-depth speed bump, **not** a sandbox, and is unset by default; the allowlist remains the control that matters.
+#### Docker Security Hardening (Planned — Not Yet Active)
-#### Docker Security Hardening
+> **Status:** The container isolation model described below is **planned, not shipped**. No Docker build context is packaged, so containers are never launched today — `run_command` runs on the host gated by the allowlist and `allow_host_execution`. The settings below document the intended hardened model for when container isolation lands.
-When Docker is available, containers are launched with hardened security settings:
+When container isolation is enabled in a future release, containers are intended to launch with hardened security settings:
 - **Non-root User:** Containers run as a non-root user
 - **Privilege Prevention:** `--security-opt no-new-privileges` prevents privilege escalation
@@ -922,12 +1008,10 @@ When Docker is available, containers are launched with hardened security setting
 ### Reporting Security Issues
-If you discover a security vulnerability:
-1. **Do Not Open a Public Issue:** Send a private email to the maintainers
-2. **Provide Details:** Include steps to reproduce and potential impact
-3. **Allow Time for Fix:** Give maintainers time to address the issue before disclosure
-4. **Follow Responsible Disclosure:** Follow responsible disclosure practices
+To report a security vulnerability, please follow the process in
+[SECURITY.md](SECURITY.md). **Do not open a public issue** — use GitHub's
+private vulnerability reporting (Security tab) or email the maintainer at
+[ismael.marin@gmail.com](mailto:ismael.marin@gmail.com).
 ---
@@ -955,9 +1039,9 @@ If you discover a security vulnerability:
 - **Solution:** Increase `max_execution_time` in your `skill-bench.json` or simplify the task
 - **Check:** Verify the command isn't hanging or waiting for input
-**Problem:** "Docker container failed to start"
-- **Solution:** Ensure Docker is running and you have permissions to run Docker commands
-- **Check:** Run `docker info` to verify Docker daemon is accessible
+**Problem:** "Command execution refused: no sandbox isolation is active and 'allow_host_execution' is not enabled"
+- **Cause:** Container isolation is not yet shipped, so commands would run on the host. SkillBench fails closed by default rather than run un-isolated.
+- **Solution:** Set `"allow_host_execution": true` in `skill-bench.json` to permit allowed commands to run directly on the host (inside the temporary sandbox directory). Enable it only when you accept that trade-off.
 **Problem:** "Context hydration failed"
 - **Solution:** Verify the source path exists and is a directory
@@ -1006,15 +1090,62 @@ If you encounter issues not covered here:
 ## CI/CD Integration
-GitHub Actions workflow included (`.github/workflows/ci.yml`):
+### Batch Runs
+Run every eval at once instead of one at a time:
+```bash
+# Every eval under the default evals/ directory
+skill-bench run --all --skill=my-service
+# Or point at a specific directory
+skill-bench run --evals-dir path/to/evals --skill=my-service
+```
+A batch run exits `0` only when **every** eval passes and non-zero if any fail, so the process exit code is itself a CI gate. Two formats are built for batch consumption:
+- `--summary` emits an aggregate JSON gate — `passed` / `failed` / `total` counts, summed `tokens` and `cost`, and the `worst_delta` eval (the smallest context-minus-baseline delta in the batch). Archive it as a single machine-readable artifact:
+  ```bash
+  skill-bench run --all --skill=my-service --summary
+  ```
+- `--format junit` aggregates the batch into one JUnit document with **one `<testcase>` per eval** (a `<failure>` child for each failing eval), so test reporters show per-eval results:
+  ```bash
+  skill-bench run --all --skill=my-service --format junit > junit.xml
+  ```
+### GitHub Action
+Downstream repos can gate a skill change on every push or PR with the bundled composite action. Add a step that references `igmarin/ruby-skill-bench@v1`:
+```yaml
+# .github/workflows/skill-bench.yml
+name: skill-bench
+on: [pull_request]
+jobs:
+  skill-bench:
+    runs-on: ubuntu-latest
+    steps:
+      - uses: actions/checkout@v4
+      - uses: igmarin/ruby-skill-bench@v1
+        with:
+          evals-dir: evals          # directory scanned for evals (default: evals)
+          skill: skills/my-service  # skill applied to every eval (default: "")
+          format: junit             # human | json | junit | html (default: junit)
+          ruby-version: "3.3"       # Ruby for ruby/setup-ruby (default: 3.3)
+          args: --summary           # extra flags appended verbatim (e.g. --summary, --pack NAME)
+```
+The action installs the gem and runs `skill-bench run --all --evals-dir <evals-dir> --format <format>` (adding `--skill` when set and appending `args` verbatim). The run step's exit code is the gate. For a full copy-paste workflow template, see [`examples/ci/`](examples/ci/).
+> The gem's own repository CI (`.github/workflows/ci.yml`) runs the test suite — rubocop, reek, and minitest against Ruby 3.3 and 3.4, on push and pull requests — and is separate from the reusable action above.
-- Runs on push and pull requests
-- Tests against Ruby 3.3 and 3.4
-- Executes rubocop, reek, and minitest
-- Outputs JUnit XML for test reporting
+To preview the machine-readable output locally:
 ```bash
-# Run locally with CI output
 skill-bench run my-eval --skill=my-skill --format json
 ```

data/docs/architecture.md CHANGED Viewed

@@ -172,9 +172,11 @@ project-root/
 │   └── my-first-eval/
 │       ├── task.md               # Agent prompt
 │       └── criteria.json         # Scoring rules
-└── .skill-bench-history.json     # Benchmark history (auto-generated)
+└── .skill-bench-trends.json      # Benchmark history (auto-generated)
 ```
+A `.skill-bench-trends.json.bak` file is created automatically as a backup of the trend file.
 ### Skill Discovery
 Skills are discovered recursively. These are all valid:

data/docs/first-eval-guide.md CHANGED Viewed

@@ -268,7 +268,7 @@ Provider is read from `skill-bench.json` — no `--provider` flag needed.
 2. Agent runs **with** skill context → produces context output
 3. Judge scores both independently → per-dimension scores
 4. Engine computes deltas → applies pass/fail logic
-5. Result is recorded in `.skill-bench-history.json` for trend tracking
+5. Result is recorded in `.skill-bench-trends.json` for trend tracking
 **Run with multiple skills:**
@@ -346,7 +346,7 @@ Both skill contexts are concatenated. The judge evaluates whether the combined c
 | **BASELINE** | Score without skill (unaided performance). Think: "How well does the AI do on its own?" |
 | **CONTEXT** | Score with skill (aided performance). Think: "How well does the AI do when it reads my skill?" |
 | **DELTA** | Improvement = CONTEXT - BASELINE. Think: "How many points did my skill add?" |
-| **TREND** | Change since the *previous* run of this exact eval + skill. Stored in `.skill-bench-history.json`. |
+| **TREND** | Change since the *previous* run of this exact eval + skill. Stored in `.skill-bench-trends.json`. |
 | **VERDICT** | PASS only if CONTEXT >= threshold AND DELTA >= minimum_delta. Both must be true. |
 | **Iterations** | ReAct loop steps for each run: thought → tools → observation. Helps you understand *how* the agent worked. |
 | **What went well** | Dimensions scoring ≥ 80% of max, with judge reasoning. Strengths of your skill. |
@@ -417,10 +417,10 @@ Your first run probably will not pass. That is normal. Here is how to improve.
 ### Use the History File
-After each run, SkillBench appends to `.skill-bench-history.json`. You can read it to track progress:
+After each run, SkillBench appends to `.skill-bench-trends.json`. You can read it to track progress:
 ```bash
-cat .skill-bench-history.json | jq '.[-1]'
+cat .skill-bench-trends.json | jq '.[-1]'
 ```
 Look at the dimension with the **smallest delta**. That is where your skill is weakest. Open `SKILL.md` and add a concrete rule targeting that dimension.
@@ -467,7 +467,7 @@ Created by `skill-bench init`. Stores provider, API key, model, timeout, and all
 }
 ```
-### `.skill-bench-history.json` — Evaluation History (Auto-Generated)
+### `.skill-bench-trends.json` — Evaluation History (Auto-Generated)
 A JSON array recording every successful eval run. SkillBench writes it automatically. It stores timestamps, eval names, skill names, scores, and deltas. This powers the **TREND** line in your output.
@@ -487,9 +487,9 @@ A JSON array recording every successful eval run. SkillBench writes it automatic
 **Tip:** Commit this file to git if you want to share trend data with your team.
-### `.skill-bench-history.json.bak` — Backup (Auto-Generated)
+### `.skill-bench-trends.json.bak` — Backup (Auto-Generated)
-A safety copy of the history file. If the main file gets corrupted, SkillBench recovers from this backup automatically. You never need to touch it.
+A snapshot of the previous good version of the history file, copied just before each new write. If the main file gets corrupted, SkillBench recovers from this backup automatically. You never need to touch it.
 ---

data/docs/testing-guide.md CHANGED Viewed

@@ -273,7 +273,7 @@ Both must be true. This prevents two failure modes:
 TREND: baseline ↑ (+2), context ↑ (+7)
 ```
-This compares the current run against the **previous run of the same eval + skill** (stored in `.skill-bench-history.json`).
+This compares the current run against the **previous run of the same eval + skill** (stored in `.skill-bench-trends.json`).
 - `↑` = improved since last run
 - `↓` = regressed since last run

data/lib/skill_bench/agent/react_agent/loop_runner.rb CHANGED Viewed

@@ -16,6 +16,7 @@ module SkillBench
         def self.call(initial_prompt, max_iterations, config)
           messages = [{ role: 'user', content: initial_prompt }]
           iterations_log = []
+          total_usage = empty_usage
           step_count = 0
           while step_count < max_iterations
@@ -24,24 +25,27 @@ module SkillBench
             step_result = Step.call(messages, config)
             iteration = step_result[:iteration]
             iterations_log << attach_step_number(iteration, step_count) if iteration
+            total_usage = add_usage(total_usage, step_result[:usage])
             unless step_result[:continue]
               final_result = step_result[:result] || { success: false, response: { error: { message: 'Step returned no result' } } }
-              return merge_iterations(final_result, iterations_log)
+              return finalize(final_result, iterations_log, total_usage)
             end
             messages = step_result[:messages]
           end
-          merge_iterations(
+          finalize(
             { success: false, response: { error: { message: Agent::ReactAgent::MAX_ITERATIONS_REACHED } } },
-            iterations_log
+            iterations_log,
+            total_usage
           )
         rescue StandardError => e
           SkillBench::ErrorLogger.log_error(e, 'ReactAgent Error')
-          merge_iterations(
+          finalize(
             { success: false, response: { error: { message: e.message } } },
-            iterations_log
+            iterations_log,
+            total_usage
           )
         end
@@ -54,14 +58,45 @@ module SkillBench
           iteration.merge(step_number: step_count)
         end
-        # Merges the collected iterations into the result response.
+        # Merges the collected iterations and accumulated usage into the response.
         #
         # @param result [Hash] The final result hash from the loop.
         # @param iterations_log [Array<Hash>] Collected iteration metadata.
-        # @return [Hash] The result with :iterations injected into :response.
-        def self.merge_iterations(result, iterations_log)
+        # @param total_usage [Hash] Summed token usage across all iterations.
+        # @return [Hash] The result with :iterations and :usage injected into :response.
+        def self.finalize(result, iterations_log, total_usage)
           response = result[:response] || {}
-          result.merge(response: response.merge(iterations: iterations_log))
+          result.merge(response: response.merge(iterations: iterations_log, usage: total_usage))
+        end
+        # A zeroed token-usage accumulator.
+        #
+        # @return [Hash] Usage hash with prompt/completion/total token counts set to zero.
+        def self.empty_usage
+          { prompt_tokens: 0, completion_tokens: 0, total_tokens: 0 }
+        end
+        # Adds a single step's usage onto a running total.
+        #
+        # @param total [Hash] The running usage total.
+        # @param usage [Hash, nil] A step's usage hash (may be nil or empty).
+        # @return [Hash] A new summed usage hash.
+        def self.add_usage(total, usage)
+          usage ||= {}
+          {
+            prompt_tokens: total[:prompt_tokens] + token_count(usage, :prompt_tokens),
+            completion_tokens: total[:completion_tokens] + token_count(usage, :completion_tokens),
+            total_tokens: total[:total_tokens] + token_count(usage, :total_tokens)
+          }
+        end
+        # Reads a token count from a usage hash, tolerating string keys.
+        #
+        # @param usage [Hash] The usage hash.
+        # @param key [Symbol] The usage key (e.g. :prompt_tokens).
+        # @return [Integer] The token count, or zero when absent.
+        def self.token_count(usage, key)
+          (usage[key] || usage[key.to_s] || 0).to_i
         end
       end
     end

data/lib/skill_bench/agent/react_agent/step.rb CHANGED Viewed

@@ -12,7 +12,8 @@ module SkillBench
         #
         # @param messages [Array<Hash>] The conversation history.
         # @param config [Hash] Configuration for this step (client params, system prompt, working dir).
-        # @return [Hash] Step outcome containing :continue (boolean), :result (hash, if finished), and :messages.
+        # @return [Hash] Step outcome containing :continue (boolean), :result (hash, if finished),
+        #   :usage (token usage for this step), and :messages.
         def self.call(messages, config)
           messages = messages.dup
           client_result = Client.call(
@@ -21,12 +22,14 @@ module SkillBench
             tools: Tools.definitions,
             **config[:client_params]
           )
+          usage = client_result[:usage] || {}
           unless client_result[:success]
             error_msg = client_result.dig(:response, :error, :message) || 'Unknown error'
             return {
               continue: false,
               result: client_result,
+              usage: usage,
               iteration: build_iteration(thought: '', tools_used: [], observation_summary: error_msg)
             }
           end
@@ -36,6 +39,7 @@ module SkillBench
             return {
               continue: false,
               result: { success: false, response: { error: { message: 'Empty response from LLM' } } },
+              usage: usage,
               iteration: build_iteration(thought: '', tools_used: [], observation_summary: 'Empty response from LLM')
             }
           end
@@ -51,6 +55,7 @@ module SkillBench
             return {
               continue: false,
               result: { success: true, response: { content: content } },
+              usage: usage,
               iteration: build_iteration(thought: thought, tools_used: [], observation_summary: '')
             }
           end
@@ -69,6 +74,7 @@ module SkillBench
           {
             continue: true,
             messages: messages,
+            usage: usage,
             iteration: build_iteration(thought: thought, tools_used: tools_used, observation_summary: observation_summary)
           }
         end

data/lib/skill_bench/cli/batch_result_printer.rb ADDED Viewed

@@ -0,0 +1,45 @@
+# frozen_string_literal: true
+require_relative '../output_formatter'
+require_relative '../services/summary_formatter'
+module SkillBench
+  module Cli
+    # Prints the aggregate result of a batch `skill-bench run --all` command.
+    #
+    # Defaults to the human-readable batch summary, but can instead emit a
+    # JUnit document (`format: :junit`) or a JSON gate (`summary: true`). The
+    # returned exit code is always {OutputFormatter.batch_exit_code}, so CI
+    # gating works identically across every output mode.
+    class BatchResultPrinter
+      # Prints the aggregate summary and returns the appropriate exit code.
+      #
+      # @param aggregate [Hash] Aggregate envelope from BatchRunnerService.
+      # @param format [Symbol, nil] Output format (:junit for JUnit XML, else human).
+      # @param summary [Boolean] When true, print the JSON summary gate instead.
+      # @return [Integer] Exit code (0 when all pass, 1 when any fails).
+      def self.call(aggregate, format: nil, summary: false)
+        puts batch_output(aggregate, format: format, summary: summary)
+        OutputFormatter.batch_exit_code(aggregate)
+      end
+      # Selects the rendered batch output for the requested mode.
+      #
+      # `:junit` and `:json` produce machine-readable batch output; `:json` maps
+      # to the same JSON gate as `summary: true`. `:html` (and any other format)
+      # falls back to the human batch summary, since there is no batch HTML report.
+      #
+      # @param aggregate [Hash] Aggregate envelope from BatchRunnerService.
+      # @param format [Symbol, nil] Output format (:junit, :json, else human).
+      # @param summary [Boolean] When true, render the JSON summary gate.
+      # @return [String] The formatted batch output.
+      def self.batch_output(aggregate, format:, summary:)
+        return Services::SummaryFormatter.format(aggregate) if summary || format == :json
+        return Services::JUnitFormatter.format_batch(aggregate) if format == :junit
+        OutputFormatter.format_batch(aggregate)
+      end
+      private_class_method :batch_output
+    end
+  end
+end