activeagent 1.7.1 → 1.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 14e5b0cec7ba170c446d1ed32ad5c9d0705ac8eae4fdc6e4e4a7d17bdae502a2
4
- data.tar.gz: 28e12f0ef3727a97fc231e5f7ace15ad4cb020b379fc76c757c2206c89f2b2e6
3
+ metadata.gz: 3642068294052d483decf5260b5d7dd1793647f06920713ca658f7b89ba948b0
4
+ data.tar.gz: 1d0811bee23dc646e916d17abc711637c1e2d41055633ace8bb1c067615d26b4
5
5
  SHA512:
6
- metadata.gz: 9cd766d568ab510b043e4902c1c5fa73ce877a268761168509f2772190ab0f2d6f259fe62e4654aa2a002d20228f876fce82c7c2f5f3e1018cd6cb01935aed27
7
- data.tar.gz: 197cab65243e11561c7216f08e9b30369819e3ef73627d0f6e72092a7fc74c16a278638f5044e11f292149a1a39dbe11cab9593a41e7289049a3a90c6232a0a0
6
+ metadata.gz: 6d4492797c3d4b238bf86c49c48b9ec12e5876bbc9ef075c4cae7f369a3e3274b3db30781b9367643740bd9fa365eb8fe8ad81308b125f6843f3f828ccf5f339
7
+ data.tar.gz: 601676a4590808eef1469f3cfafc5136c43bccbd22e7b600c15940eb7a1fc957479c568676d699d7d41c10836fa621eb072a3ff5fd20d5f4cfa22a517e982416
data/CHANGELOG.md CHANGED
@@ -7,6 +7,214 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [1.8.0] - 2026-09-29
11
+
12
+ Releases `activeagent` and `actionagent` 1.8.0 from one tag. A minor release.
13
+ Settings -> Integrations connects GitHub and Claude Code. A connected
14
+ repository's checkout boots as a sandbox, on a developer's machine with the new
15
+ `:local` backend, where Claude Code sessions run and an evaluation can run
16
+ against the checkout without the agent being edited. The dashboard's MCP
17
+ server gains evaluation and telemetry tools for a developer's own coding
18
+ harness. Ollama hosts can be tested and can be remote, with an optional Bearer
19
+ API key. Comparison runs lead with a per-model table and filter the fix list
20
+ by model. A publisher can check its collector before a run.
21
+
22
+ Upgrading: run `bin/rails generate action_agent:install --skip` and
23
+ `bin/rails db:migrate`. The generator adds what an install lacks:
24
+ `provider_keys.api_key` and the `github_connections` and `code_sessions`
25
+ tables. An app on the RubyLLM provider needs ruby_llm 1.16 or later; 2.x
26
+ works too.
27
+
28
+ The Claude Code connection stores Anthropic API keys only. Anthropic does not
29
+ let third-party products collect, store or route requests through Claude.ai
30
+ subscription credentials
31
+ ([Claude Code legal and compliance](https://code.claude.com/docs/en/legal-and-compliance.md)),
32
+ so a `claude setup-token` token (`sk-ant-oat…`) is refused. An install that
33
+ stored one while running a pre-release build never hands it to a session: its
34
+ owner sees Claude Code as needing an API key (`needs_replacing: true` in
35
+ `GET /api/provider_keys`) until they paste one.
36
+ `bin/rails action_agent:claude_code:purge_subscription_tokens` deletes the
37
+ stored tokens and prints how many it removed. A developer who wants sessions on
38
+ their own Claude login sets `config.claude_code_auth = :local_login` with the
39
+ `:local` backend and runs `claude /login` on that machine instead.
40
+
41
+ ### Added
42
+
43
+ - **Evaluation and telemetry tools on the MCP facade** (`actionagent`). The
44
+ dashboard's MCP server now offers `evaluations_list`, `evaluations_get`,
45
+ `evaluations_run`, `evaluation_runs_get`, `evaluation_runs_compare`,
46
+ `traces_search` and `traces_get`, so a developer's own coding harness can
47
+ run an agent's evaluations, read the fix items and failing traces, and
48
+ iterate on the agent in its own checkout without the dashboard holding a
49
+ model login. The tools read under the API key's owner exactly as the JSON
50
+ API reads under the signed-in owner, run through the same execution,
51
+ quota and sandbox checks as `POST /api/evaluations/:id/run`, bound their
52
+ output, and mask the owner's credentials. Their names cannot collide with
53
+ schema tools or agent tools. `ActionAgent.mcp_dashboard_tools = false`
54
+ turns them off.
55
+
56
+ - **Local checkout sandboxes and Claude Code sessions** (`actionagent`, #489).
57
+ A new `:local` sandbox backend (`config.sandbox_service = :local`) makes
58
+ **Start sandbox** work on a developer's machine without containers. It
59
+ clones the repository under `tmp/action_agent/sandboxes`, runs the setup
60
+ the checkout's optional `.activeagents/sandbox.yml` names (`env`, `setup`,
61
+ `manifest`, `start`), and boots the app on `127.0.0.1`. It then registers
62
+ the app's MCP facade as the `sandbox:<session_id>` server. Every process
63
+ starts from a sanitized copy of the dashboard's environment: without its
64
+ database and Redis URLs, Rails keys and environment, Bundler and Ruby
65
+ settings, git repository and config variables, `SSH_AUTH_SOCK`,
66
+ model-provider and Claude Code settings, variables named like a secret, or
67
+ URLs carrying credentials. The GitHub token reaches
68
+ only the fetch, and the Claude Code API key reaches only Claude Code.
69
+ `:local` runs the owner's code with the dashboard's privileges, so it is off
70
+ outside development and test unless
71
+ `ActionAgent.local_sandboxes_enabled = true`.
72
+ - `bin/rails action_agent:sandbox:manifest` writes the booted app's
73
+ `{mcp_path, mcp_token}` for the backend.
74
+ - `bin/rails action_agent:sandbox:reap` expires overdue sandboxes and stops
75
+ them.
76
+ - A ready sandbox runs headless Claude Code sessions (`--permission-mode
77
+ acceptEdits` by default). You start them from Settings -> Integrations, or
78
+ through `/api/sandboxes/:session_id/code_sessions`. Their events stream
79
+ into the dashboard, and the checkout's diff follows.
80
+ - New options: `local_sandboxes_enabled`, `local_sandbox_root`,
81
+ `local_sandbox_boot_timeout`, `claude_code_command`,
82
+ `claude_code_permission_mode`, `claude_code_max_turns`,
83
+ `claude_code_timeout` and `claude_code_auth`.
84
+ - `claude_code_auth = :local_login` runs sessions on the machine's own
85
+ Claude Code login (`claude /login`), with no stored key: the backend
86
+ passes no credential and no `CLAUDE_CONFIG_DIR`, so `claude` uses the
87
+ dashboard user's own `~/.claude` or keychain, which the dashboard never
88
+ reads. `LocalSandboxBackend.claude_login_status` asks
89
+ `claude auth status --json` (cached for a minute) and keeps only
90
+ `loggedIn` and the login method. `GET /api/sandboxes` reports
91
+ `claude_code_auth` and, in this mode, `claude_code_login`
92
+ (`{ logged_in, auth_method }`), and the assistant's
93
+ `connections.claude_code` the same as `auth` and `login`. Other backends
94
+ refuse sessions in this mode (`code_sessions_supported: false`, and a
95
+ `422` naming the reason).
96
+ - `app_runtime` sandboxes now provision in the background and last 2 hours.
97
+ - Run `rails g action_agent:install` to add the
98
+ `create_active_agent_code_sessions` migration.
99
+ - This repository's own `.activeagents/sandbox.yml` boots `test/dummy`.
100
+ - Every `:local` sandbox boots on databases of its own, so a checkout of
101
+ the dashboard's own app no longer migrates the developer's development
102
+ database. The backend reads the adapter from the checkout's
103
+ `config/database.yml` without running its ERB, and sets `DATABASE_URL`
104
+ and `<NAME>_DATABASE_URL` (`QUEUE_DATABASE_URL`, `CACHE_DATABASE_URL`):
105
+ SQLite files in the workspace, or `<database>_sandbox_<id>` on
106
+ PostgreSQL and MySQL, which terminate drops with the checkout's
107
+ Rails database tasks restricted to the names and URLs recorded at boot.
108
+ Setting a variable in `sandbox.yml`'s `env` overrides it and excludes
109
+ that database from cleanup. Replica mappings follow their own writer;
110
+ ambiguous mappings require an explicit URL instead of guessing.
111
+ - The Claude Code panel has a **Model** select: Claude Code's own
112
+ default, the `sonnet`, `opus` and `haiku` aliases, or any model id under
113
+ *Other…*. It remembers the last choice per browser, and each session
114
+ shows the model it ran on.
115
+ - A run can use a checkout sandbox without the agent being edited:
116
+ `sandbox_id` on `POST /api/evaluations/:id/run` (and on the runner's
117
+ `/api/agents/:id/execute` and `/test`) gives that run's tool dispatcher
118
+ the sandbox's `sandbox:<session_id>` runtime, as if the agent listed it.
119
+ The sandbox must be the caller's, a ready `app_runtime` sandbox, and the
120
+ agent owner's; anything else is a `422`. The run records which sandbox
121
+ it used (`run.sandbox`), and a scenario suite's **Run against sandbox**
122
+ select, its Runs list and the run report show it. The selected sandbox
123
+ takes precedence for matching tool names, with one schema per name.
124
+ Queued agent runs fail if their selected sandbox stops, and failed
125
+ sandbox discovery never silently falls back to the original tools.
126
+ - **Claude Code connection** (`actionagent`, #478). Settings -> Integrations
127
+ stores an Anthropic API key (`sk-ant-api…`, from the Claude Console) as the
128
+ `claude_code` provider key. It is encrypted, write-only, and not an agent
129
+ provider. `SandboxSession#runtime_environment` hands it to an
130
+ `app_runtime` backend as `ANTHROPIC_API_KEY`, so the checkout can run
131
+ Claude Code sessions. Claude subscription tokens (`claude setup-token`)
132
+ are refused, as Anthropic's terms require (see the upgrading note above).
133
+ `/api/provider_keys` rows now carry `kind` (`key`, `host` or
134
+ `connection`) and `needs_replacing`.
135
+ - **GitHub connections and checkout sandboxes** (`actionagent`, #477).
136
+ Settings -> Integrations connects GitHub over OAuth
137
+ (`ActionAgent.github_client_id` / `github_client_secret`, or
138
+ `GITHUB_CLIENT_ID` / `GITHUB_CLIENT_SECRET`), stores the token encrypted,
139
+ and lets the owner choose which repositories the workspace may use. Only
140
+ repositories GitHub lists for the token can be selected. A new
141
+ `app_runtime` sandbox type checks out one of them: the backend receives
142
+ `sandbox_session.checkout_spec` (repository, ref, clone URL, token),
143
+ boots the app, and returns `mcp_url` / `mcp_token` from `create_sandbox`.
144
+ The session is then an MCP server keyed `sandbox:<session_id>`. An agent
145
+ that lists that key in `mcp_servers` runs, and is evaluated, with the
146
+ checkout's own tools. Lookups are scoped to the agent's owner. Run
147
+ `rails g action_agent:install` to add the
148
+ `create_active_agent_github_connections` migration.
149
+
150
+ - **Ollama hosts are testable and can be remote** (`actionagent`). Settings ->
151
+ Provider API Keys gains a **Test connection** for Ollama that reports
152
+ whether the server is reachable, the round-trip time and the models it
153
+ serves, before or after saving (`POST <mount>/api/provider_keys/test`,
154
+ read-only). The host is accepted as a bare server address
155
+ (`http://localhost:11434`; the OpenAI-compatible `/v1` path is added) and
156
+ an optional **API key** is stored beside it and sent as a Bearer token,
157
+ for a server behind an authenticating proxy or Ollama Cloud. The agent
158
+ builder's live Ollama model list uses the same probe and key. When no host
159
+ is configured the card shows the host app's `config/active_agent.yml`
160
+ default. The install generator emits a guarded `add_provider_key_api_key`
161
+ migration for existing installs; re-run
162
+ `bin/rails generate action_agent:install --skip` and `bin/rails db:migrate`.
163
+ - **A model comparison table on comparison runs** (`actionagent`,
164
+ `activeagent`). A run over several models now leads its Models section
165
+ with one row per model, best first: passed, mean score, average latency,
166
+ average tokens per scenario, cost (and per scenario), and the model's
167
+ typical fault — its most frequent one with the diagnosis of a result that
168
+ carries it. The scenario suite panel, the sampling run detail and the
169
+ standalone HTML report (`ActiveAgent::Evals::ReportHtml`) all render it.
170
+ - **What to fix, filtered by model** (`actionagent`, `activeagent`). On a
171
+ comparison run the fix list takes a model chip, narrowing to the items
172
+ attributed to that model and counting what that model alone produced,
173
+ since one model may need more instruction than another. The standalone
174
+ report filters through radio chips and stylesheet rules — it still ships
175
+ no script.
176
+
177
+ - **A publisher can check its collector before a run** (`activeagent`).
178
+ `ActiveAgent::Evals::Publisher#verify!` asks the collector whether it is up
179
+ and accepts the key before a run is paid for. It posts an empty JSON object,
180
+ which a compatible collector refuses with a 422 naming `version`, without
181
+ storing anything; anything else raises `Publisher::Error` with a delivery's
182
+ status, detail and guidance. `Publisher#endpoint` returns the collector URL.
183
+
184
+ ### Changed
185
+
186
+ - **Collector rejections say what the status means** (`activeagent`). A
187
+ `Publisher::Error` for a 401, 403, 404, 415 or 501 rejection names a refused
188
+ key, an account an operator must act on, an endpoint that is not a
189
+ collector, a rewritten `Content-Type`, or an install with no evaluation
190
+ store, in place of the generic guidance.
191
+
192
+ ### Fixed
193
+
194
+ - **A nested scenario expectation written as one value** (`activeagent`).
195
+ `ScenarioParser` now stores `{ expectations: { contains: "30" } }` as a
196
+ list of one, the shape the persisted scenario and the dashboard's matrix
197
+ read; an object-list import with a lone value used to break the suite
198
+ panel. The matrix also tolerates scenarios persisted before this.
199
+
200
+ ## [1.7.2] - 2026-09-29
201
+
202
+ Releases `activeagent` and `actionagent` 1.7.2 from one tag. A patch on 1.7.1:
203
+ the RubyLLM provider runs on ruby_llm 1.16 and 2.x, where 1.7.1 required 1.x.
204
+ Apps on the RubyLLM provider need ruby_llm 1.16 or later; earlier 1.x
205
+ releases lack the APIs the provider calls. No migrations.
206
+
207
+ ### Changed
208
+
209
+ - **The RubyLLM provider supports ruby_llm 1.16 and 2.x** (`activeagent`).
210
+ The adapter handles 2.x's tool interface, token limits, embedding model
211
+ objects, usage and finish reasons while keeping the 1.16 API working.
212
+ Requiring `>= 1.16, < 3` prevents Bundler from selecting an older 1.x
213
+ release without the APIs the adapter calls. Rails main can now resolve
214
+ RubyLLM 2.x alongside Active Storage's Marcel 2 dependency. CI also runs
215
+ the full suite with RubyLLM 1.16 to retain coverage of that version.
216
+ Unsupported-version errors name the loaded version and both bounds (#508).
217
+
10
218
  ## [1.7.1] - 2026-09-29
11
219
 
12
220
  Releases `activeagent` and `actionagent` 1.7.1 from one tag. A patch on 1.7.0:
@@ -11,7 +11,8 @@ module ActiveAgent
11
11
  # Publishes a completed report without replaying the agent. The caller must
12
12
  # retain run_id when retrying: compatible collectors treat that identity as
13
13
  # immutable within the authenticated account. Delivery is blocking and does
14
- # not follow redirects with the account's bearer credential.
14
+ # not follow redirects with the account's bearer credential. +verify!+ asks
15
+ # the collector whether it is up and accepts the key before a run is paid for.
15
16
  #
16
17
  # Every failure to deliver raises Error. Invalid arguments raise
17
18
  # ArgumentError before anything is sent.
@@ -20,6 +21,9 @@ module ActiveAgent
20
21
  MAX_BYTES = 2 * 1024 * 1024
21
22
  DETAIL_LIMIT = 200
22
23
 
24
+ # @return [String] the collector URL reports go to
25
+ attr_reader :endpoint
26
+
23
27
  # Raised for every failed delivery. Only a network failure keeps the
24
28
  # underlying error as its +cause+, so neither the response nor the
25
29
  # report reaches a log through the exception chain.
@@ -48,10 +52,15 @@ module ActiveAgent
48
52
  # Whether each rejection status is retryable, and what the caller should
49
53
  # do about it. Other statuses fall back to the rules in +rejection+.
50
54
  REJECTIONS = {
55
+ 401 => [ false, "the collector refused the API key; check the key against the collector's account" ],
56
+ 403 => [ false, "the account may not store this report until an operator acts, for example on a cap on observed agents, evaluations or scenarios; resolve that before retrying" ],
57
+ 404 => [ false, "nothing at the endpoint takes evaluation reports; check that it is a collector's /v1/evaluations or <mount>/api/evaluation_reports URL" ],
51
58
  409 => [ false, "the collector already holds a different report under this run_id; never retry this report with the same run_id" ],
52
59
  413 => [ false, "the report exceeds the collector's size limit; publish a smaller selection" ],
60
+ 415 => [ false, "the collector did not receive application/json; check anything between the publisher and the collector that rewrites the Content-Type" ],
53
61
  422 => [ false, "correct what the collector refused before retrying" ],
54
- 429 => [ true, "the account is over its quota or rate limit; retain the report and run_id and retry later" ]
62
+ 429 => [ true, "the account is over its quota or rate limit; retain the report and run_id and retry later" ],
63
+ 501 => [ false, "the collector has no evaluation store; migrate the install, or publish to one generated with evaluation tables" ]
55
64
  }.freeze
56
65
 
57
66
  # The key is sent as a bearer token and filtered from the collector's
@@ -65,6 +74,7 @@ module ActiveAgent
65
74
  unless @uri.scheme == "https" || %w[localhost 127.0.0.1 ::1].include?(@uri.hostname)
66
75
  raise ArgumentError, "Evaluation endpoint requires HTTPS except on loopback hosts"
67
76
  end
77
+ @endpoint = @uri.to_s
68
78
 
69
79
  @api_key = api_key.to_s.strip
70
80
  raise ArgumentError, "Evaluation API key is required" if @api_key.empty?
@@ -91,6 +101,52 @@ module ActiveAgent
91
101
  body = encode(identities.merge("version" => 1, "report" => report_hash(report)))
92
102
  raise Error, "Evaluation report exceeds the 2 MiB delivery limit; publish a smaller selection" if body.bytesize > MAX_BYTES
93
103
 
104
+ deliver("retain the report and run_id for retry") do
105
+ response = post(body)
106
+ raise rejection(response) unless %w[200 201].include?(response.code)
107
+
108
+ receipt = JSON.parse(response.body.to_s)
109
+ unless receipt.is_a?(Hash) && receipt["run_id"] == run_id && receipt["status"] == "complete" && receipt["id"] && receipt["evaluation_id"]
110
+ raise Error.new("Evaluation collector returned an invalid completion receipt; retain the report and run_id for retry", retryable: true)
111
+ end
112
+ receipt
113
+ end
114
+ end
115
+
116
+ # Returns true when the collector is up and accepts the API key, without
117
+ # storing anything. It posts an empty JSON object, which a compatible
118
+ # collector authenticates, parses, and then refuses with a 422 whose
119
+ # +error+ names +version+, as not a version-1 report. That refusal, and
120
+ # only that one, is the ready answer. Call it before an expensive run, so
121
+ # a stopped collector or a refused key fails before the first model call
122
+ # rather than after the last one.
123
+ #
124
+ # Raises Error for anything else: a rejection other than 422, with the
125
+ # status, detail and guidance +call+ would carry (401 for a refused key,
126
+ # 404 when nothing at the endpoint takes reports); a retryable delivery
127
+ # failure when the collector cannot be reached; a 422 that says anything
128
+ # else; or a collector that stores the empty object. The last two mean the
129
+ # endpoint is not a compatible collector.
130
+ def verify!
131
+ deliver("start the collector or check the endpoint, then verify again") do
132
+ response = post("{}")
133
+ if response.code == "422"
134
+ detail = collector_detail(response.body)
135
+ next true if detail&.match?(/\bversion\b/)
136
+
137
+ raise Error.new("Evaluation collector answered HTTP 422 without refusing the empty envelope as a version-1 report, " \
138
+ "so it is not a compatible collector; check the endpoint", status: 422, detail: detail)
139
+ end
140
+ raise rejection(response) unless %w[200 201].include?(response.code)
141
+
142
+ raise Error, "Evaluation collector stored an empty report, so it is not a compatible collector; check the endpoint"
143
+ end
144
+ end
145
+
146
+ private
147
+
148
+ # Posts +body+ to the endpoint with the bearer key and the delivery timeouts.
149
+ def post(body)
94
150
  http = Net::HTTP.new(@uri.hostname, @uri.port)
95
151
  http.use_ssl = @uri.scheme == "https"
96
152
  http.open_timeout = @open_timeout
@@ -101,26 +157,25 @@ module ActiveAgent
101
157
  request["Content-Type"] = "application/json"
102
158
  request["Accept"] = "application/json"
103
159
  request.body = body
104
- response = http.request(request)
105
- raise rejection(response) unless %w[200 201].include?(response.code)
160
+ http.request(request)
161
+ end
106
162
 
107
- receipt = JSON.parse(response.body.to_s)
108
- unless receipt.is_a?(Hash) && receipt["run_id"] == run_id && receipt["status"] == "complete" && receipt["id"] && receipt["evaluation_id"]
109
- raise Error.new("Evaluation collector returned an invalid completion receipt; retain the report and run_id for retry", retryable: true)
110
- end
111
- receipt
163
+ # Runs one exchange with the collector, turning every failure to reach it
164
+ # or to read its answer into a retryable Error that quotes nothing from
165
+ # the response and ends with +guidance+, what the caller should do next.
166
+ # An Error the block raises passes through unchanged.
167
+ def deliver(guidance)
168
+ yield
112
169
  rescue JSON::ParserError
113
170
  # The parser's message quotes the body.
114
- raise Error.new("Evaluation collector returned invalid JSON; retain the report and run_id for retry", retryable: true), cause: nil
171
+ raise Error.new("Evaluation collector returned invalid JSON; #{guidance}", retryable: true), cause: nil
115
172
  rescue Net::HTTPBadResponse, Net::HTTPHeaderSyntaxError, Zlib::Error => e
116
173
  # These messages can quote the response's status line, headers or body.
117
- raise Error.new("Evaluation collector returned a malformed response (#{e.class}); retain the report and run_id for retry", retryable: true), cause: nil
174
+ raise Error.new("Evaluation collector returned a malformed response (#{e.class}); #{guidance}", retryable: true), cause: nil
118
175
  rescue IOError, SocketError, SystemCallError, Timeout::Error, OpenSSL::SSL::SSLError => e
119
- raise Error.new("Evaluation delivery failed (#{e.class}); retain the report and run_id for retry", retryable: true)
176
+ raise Error.new("Evaluation delivery failed (#{e.class}); #{guidance}", retryable: true)
120
177
  end
121
178
 
122
- private
123
-
124
179
  def report_hash(report)
125
180
  hash = report.to_h if report.respond_to?(:to_h) && !report.nil? && !report.is_a?(Array)
126
181
  raise ArgumentError, "report must be a Report or its saved JSON hash" unless hash.is_a?(Hash)
@@ -178,6 +178,9 @@ module ActiveAgent
178
178
  DesignTokens.css(scope: ":root.theme-dark", tokens: DesignTokens::DARK, color_scheme: "dark"),
179
179
  STYLES,
180
180
  ".mx { grid-template-columns: minmax(240px, 1.6fr) 150px repeat(#{@models.size}, minmax(170px, 1fr)); }",
181
+ # One rule per model: with that chip checked, hide every fix card
182
+ # attributed to other models (cards attributed to none stay).
183
+ *@models.each_index.map { |i| ".fix-section:has(input[value=\"m#{i}\"]:checked) .fix[data-models]:not([data-models~=\"m#{i}\"]) { display: none; }" },
181
184
  ".matrix .inner { min-width: #{390 + 185 * @models.size}px; }"
182
185
  ].join("\n")
183
186
  end
@@ -233,12 +236,78 @@ module ActiveAgent
233
236
  <<~PANEL
234
237
  <div class="panel">
235
238
  <div class="panel-head"><span class="micro">Models</span><span class="right">judged by #{h(judged_by)}</span></div>
239
+ #{html_comparison_table if comparing?}
236
240
  #{blocks.join}
237
241
  #{verdict_row}
238
242
  </div>
239
243
  PANEL
240
244
  end
241
245
 
246
+ # The comparison read across: one row per model, best first (pass rate,
247
+ # then mean score) — passed, mean score, average latency, average
248
+ # tokens per scenario, cost, and the model's typical fault. The blocks
249
+ # under it carry the same figures per model with bars and every fault.
250
+ def html_comparison_table
251
+ rows = summary_by_model.sort_by do |label, stats|
252
+ total = stats["scenarios"].to_i
253
+ [ total.positive? ? -stats["passed"].to_f / total : 0.0, -(stats["avg_score"] || -1).to_f, @models.index(model_by_label(label)).to_i ]
254
+ end
255
+
256
+ <<~TABLE
257
+ <div class="compare"><table>
258
+ <thead><tr><th>Model</th><th class="num">Passed</th><th class="num">Mean score</th><th class="num">Avg latency</th><th class="num" title="Average input + output tokens per scenario">Avg tokens</th><th class="num" title="Cohort spend, and per scenario">Cost</th><th class="fault">Typical fault</th></tr></thead>
259
+ <tbody>#{rows.map { |label, stats| html_comparison_row(label, stats) }.join}</tbody>
260
+ </table></div>
261
+ TABLE
262
+ end
263
+
264
+ def html_comparison_row(label, stats)
265
+ short, provider = split_label(model_by_label(label))
266
+ total = stats["scenarios"].to_i
267
+ ratio = total.positive? ? stats["passed"].to_f / total : 0.0
268
+ pick = comparing? && winner == label ? %(<span class="pick" title="picked by the judge">★ pick</span>) : ""
269
+ per = ->(value) { value.nil? || total.zero? ? nil : value.to_f / total }
270
+ avg_tokens = per.call(stats["input_tokens"].to_i + stats["output_tokens"].to_i)
271
+ tokens_cell = avg_tokens ? h(fmt_k(avg_tokens.round)) : "—"
272
+ tokens_title = avg_tokens ? %( title="#{per.call(stats['input_tokens']).to_f.round} in · #{per.call(stats['output_tokens']).to_f.round} out per scenario") : ""
273
+ per_cost = per.call(stats["cost"])
274
+ cost_cell = stats["cost"].nil? ? "—" : h(fmt_cost(stats["cost"]))
275
+ cost_cell += "<span class=\"per\">#{h(fmt_cost(per_cost))}/scenario</span>" if per_cost
276
+
277
+ <<~ROW
278
+ <tr>
279
+ <td class="model-cell"><span class="name">#{h(short)}</span>#{pick}<span class="provider">#{h(provider)}</span></td>
280
+ <td class="num ratio tone-#{tone_for(ratio)}">#{total.positive? ? "#{stats['passed']}/#{total}" : '—'}</td>
281
+ <td class="num">#{h(fmt_mean_score(stats['avg_score']))}</td>
282
+ <td class="num">#{h(fmt_ms(stats['avg_duration_ms']))}</td>
283
+ <td class="num"#{tokens_title}>#{tokens_cell}</td>
284
+ <td class="num">#{cost_cell}</td>
285
+ <td class="fault">#{typical_fault_text(label, stats)}</td>
286
+ </tr>
287
+ ROW
288
+ end
289
+
290
+ # "missing content ×2 · refund_window: The answer is missing expected
291
+ # content: 30." — the model's most frequent fault, and the diagnosis of
292
+ # the first result that carries it; "no faults" for a clean cohort.
293
+ def typical_fault_text(label, stats)
294
+ tally = stats["faults"] || {}
295
+ return %(<span class="clean">no faults</span>) if tally.empty?
296
+
297
+ mine = @results.select { |result| result.label == label }
298
+ example_of = ->(fault) { mine.find { |result| result.fault == fault } }
299
+ # Most frequent first; between equals, a fault a result can explain,
300
+ # then a specific fault over the judge's catch-all, then the name.
301
+ fault, count = tally.min_by { |name, n| [ -n, example_of.call(name) ? 0 : 1, name == "low_quality" ? 1 : 0, name ] }
302
+ example = example_of.call(fault)
303
+ head = "#{fault_name(fault)} ×#{count}"
304
+ return h(head) unless example&.summary.present?
305
+
306
+ detail = "#{example.scenario.key}: #{example.summary}"
307
+ detail = "#{detail[0, 119]}…" if detail.length > 120
308
+ "#{h(head)} <span class=\"detail\">· #{h(detail)}</span>"
309
+ end
310
+
242
311
  def html_model_block(label, stats)
243
312
  short, provider = split_label(model_by_label(label))
244
313
  total = stats["scenarios"]
@@ -277,13 +346,36 @@ module ActiveAgent
277
346
  end
278
347
 
279
348
  <<~FIXES
280
- <section class="section" aria-label="Recommendations">
349
+ <section class="section fix-section" aria-label="Recommendations">
281
350
  <div class="section-head"><span class="micro">What to fix</span><span class="meta">#{h(meta)}</span></div>
351
+ #{html_fix_filter(items) if comparing? && items.any?}
282
352
  #{body}
283
353
  </section>
284
354
  FIXES
285
355
  end
286
356
 
357
+ # A model filter for the fix cards — a fault one model keeps making is
358
+ # that model's to fix, so the list narrows to what was attributed to
359
+ # it. Radio chips and stylesheet rules alone (the page carries no
360
+ # script): each card names its models in data-models, and a checked
361
+ # model hides every card that does not name it. Cards attributed to no
362
+ # model (an older run) stay under every filter.
363
+ def html_fix_filter(items)
364
+ chips = [ %(<label class="chip pick-model"><input type="radio" name="fix-model" value="all" checked><span>all models #{items.size}</span></label>) ]
365
+ @models.each_with_index do |spec, index|
366
+ count = items.count { |item| Array(item["models"]).empty? || item["models"].include?(spec.label) }
367
+ chips << %(<label class="chip pick-model"><input type="radio" name="fix-model" value="m#{index}"><span>#{h(short_name(spec))} #{count}</span></label>)
368
+ end
369
+ %(<div class="fix-filter"><span class="micro sm">for</span>#{chips.join}</div>)
370
+ end
371
+
372
+ def fix_model_tokens(item)
373
+ labels = Array(item["models"])
374
+ return "" if labels.empty?
375
+
376
+ labels.filter_map { |label| (index = @models.index(model_by_label(label))) && "m#{index}" }.join(" ")
377
+ end
378
+
287
379
  def html_fix_card(item)
288
380
  tone = item["kind"] == "instruction" ? "info" : "error"
289
381
  glyph = tone == "info" ? "[i]" : "[!]"
@@ -297,7 +389,8 @@ module ActiveAgent
297
389
  parts << html_fix_server(item["server"]) if item["server"]
298
390
  parts << %(<div class="note">#{h(item['note'])}</div>) if item["note"].present?
299
391
  parts << html_fix_action(item["action"]) if item["action"]
300
- %(<div class="fix">#{parts.join}</div>)
392
+ models = fix_model_tokens(item)
393
+ %(<div class="fix"#{%( data-models="#{models}") if models.present?}>#{parts.join}</div>)
301
394
  end
302
395
 
303
396
  def html_fix_tools(item)
@@ -557,6 +650,24 @@ module ActiveAgent
557
650
  .tok .in { color: var(--color-token-in); }
558
651
  .tok .out { color: var(--color-token-out); }
559
652
  .faults { display: flex; gap: 6px; flex-wrap: wrap; }
653
+ .compare { overflow-x: auto; border-top: 1px solid var(--color-border-light); }
654
+ .compare table { width: 100%; border-collapse: collapse; }
655
+ .compare th { padding: 8px 12px; text-align: left; vertical-align: bottom; white-space: nowrap; font-family: var(--font-mono); font-size: 10px; font-weight: 600; letter-spacing: 0.05em; text-transform: uppercase; color: var(--color-text-muted); background: var(--color-muted); }
656
+ .compare td { padding: 9px 12px; vertical-align: top; border-top: 1px solid var(--color-border-light); font-size: 13px; color: var(--color-text-cell); }
657
+ .compare th.num, .compare td.num { text-align: right; }
658
+ .compare td.num { font-family: var(--font-mono); font-size: 12px; white-space: nowrap; }
659
+ .compare td.ratio { font-weight: 600; }
660
+ .compare .per { display: block; font-weight: 400; color: var(--color-text-muted); }
661
+ .compare .model-cell { white-space: nowrap; }
662
+ .compare .model-cell .name { font-family: var(--font-mono); font-size: 12px; font-weight: 600; color: var(--color-text-primary); }
663
+ .compare .model-cell .provider { display: block; font-family: var(--font-mono); font-size: 11px; color: var(--color-text-muted); }
664
+ .compare .pick { margin-left: 6px; font-family: var(--font-mono); font-size: 10px; font-weight: 700; color: var(--color-warning-text); }
665
+ .compare th.fault, .compare td.fault { width: 34%; }
666
+ .compare td.fault .detail { color: var(--color-text-secondary); }
667
+ .fix-filter { display: flex; align-items: center; gap: 6px; flex-wrap: wrap; }
668
+ .pick-model { cursor: pointer; border: 1px solid var(--color-border); background: var(--color-card); }
669
+ .pick-model input { position: absolute; opacity: 0; width: 0; height: 0; }
670
+ .pick-model:has(input:checked) { border-color: var(--color-accent-ui); background: var(--color-accent-ui-tint); color: var(--color-accent-ui); }
560
671
  .clean { font-family: var(--font-mono); font-size: 11px; color: var(--color-success-text); }
561
672
  .verdict { padding: 10px 12px; border-top: 1px solid var(--color-border-light); font-size: 12px; line-height: 18px; color: var(--color-text-cell); }
562
673
  .verdict .micro { margin-right: 8px; }
@@ -125,6 +125,10 @@ module ActiveAgent
125
125
  expectations = (entry["expectations"] || entry["expect"] || {}).to_h.stringify_keys
126
126
  %w[tools contains not_contains].each do |field|
127
127
  expectations[field] = Array(entry[field]) if entry.key?(field)
128
+ # A nested expectation written as one value ({ contains: "30" })
129
+ # is a list of one: the persisted scenario and the dashboard's
130
+ # matrix read each field as an array.
131
+ expectations[field] = Array(expectations[field]) if expectations.key?(field)
128
132
  end
129
133
 
130
134
  scenario(
@@ -10,8 +10,8 @@ require_relative "concerns/tool_choice_clearing"
10
10
  GEM_LOADERS = {
11
11
  anthropic: [ "anthropic", "~> 1.12", "anthropic" ],
12
12
  openai: [ "openai", "~> 0.34", "openai" ],
13
- # ruby_llm 2.0 renamed the APIs RubyLLMProvider calls.
14
- ruby_llm: [ "ruby_llm", "~> 1.0", "ruby_llm" ]
13
+ # Keep a tested floor: 1.2 lacks the provider APIs this adapter uses.
14
+ ruby_llm: [ "ruby_llm", [ ">= 1.16", "< 3" ], "ruby_llm" ]
15
15
  }
16
16
 
17
17
  # Requires a provider's gem dependency.
@@ -23,16 +23,17 @@ GEM_LOADERS = {
23
23
  # version is outside the supported range
24
24
  def require_gem!(type, file_name)
25
25
  gem_name, requirement, package_name = GEM_LOADERS.fetch(type)
26
+ requirements = Array(requirement)
26
27
  provider_name = file_name.split("/").last.delete_suffix(".rb").camelize
27
28
 
28
29
  begin
29
- gem(gem_name, requirement)
30
+ gem(gem_name, *requirements)
30
31
  require(package_name)
31
32
  rescue LoadError
32
33
  loaded = Gem.loaded_specs[gem_name]
33
34
  if loaded && !Gem::Requirement.new(requirement).satisfied_by?(loaded.version)
34
- raise LoadError, "#{provider_name} supports the '#{gem_name}' gem #{requirement}, but #{loaded.version} is loaded. " \
35
- "Add `gem \"#{gem_name}\", \"#{requirement}\"` to your Gemfile and run `bundle update #{gem_name}`."
35
+ raise LoadError, "#{provider_name} supports the '#{gem_name}' gem #{requirements.join(', ')}, but #{loaded.version} is loaded. " \
36
+ "Add `gem \"#{gem_name}\", #{requirements.map(&:inspect).join(', ')}` to your Gemfile and run `bundle update #{gem_name}`."
36
37
  end
37
38
 
38
39
  raise LoadError, "The '#{gem_name}' gem is required for #{provider_name}. Please add it to your Gemfile and run `bundle install`."
@@ -27,6 +27,11 @@ module ActiveAgent
27
27
  {}
28
28
  end
29
29
 
30
+ # RubyLLM 2 names for the same tool interface.
31
+ alias_method :parameters_schema, :params_schema
32
+ alias_method :declared_parameters, :parameters
33
+ alias_method :provider_options, :provider_params
34
+
30
35
  private
31
36
 
32
37
  def deep_stringify(obj)
@@ -62,14 +62,19 @@ module ActiveAgent
62
62
  schema = ruby_llm_schema(parameters[:response_format])
63
63
  kwargs[:schema] = schema if schema
64
64
 
65
- # Pass extra params (max_tokens, etc.) via RubyLLM's params: deep-merge
65
+ # RubyLLM 2 renamed params: and exposes a provider-neutral token limit.
66
66
  max_tokens = parameters[:max_tokens] || options.max_tokens
67
67
  if max_tokens
68
- kwargs[:params] = { max_tokens: max_tokens }
68
+ if ruby_llm_v2?
69
+ kwargs[:max_output_tokens] = max_tokens
70
+ else
71
+ kwargs[:params] = { max_tokens: max_tokens }
72
+ end
69
73
  end
70
74
 
71
75
  if parameters[:stream]
72
76
  stream_proc = parameters[:stream]
77
+ @stream_tool_calls = {}
73
78
 
74
79
  # For streaming, pass a block that forwards chunks
75
80
  @ruby_llm_provider.complete(messages, **kwargs) do |chunk|
@@ -95,7 +100,8 @@ module ActiveAgent
95
100
  inputs = input.is_a?(Array) ? input : [ input ]
96
101
 
97
102
  data = inputs.map.with_index do |text, index|
98
- embedding = @ruby_llm_provider.embed(text, model: model_id, dimensions: parameters[:dimensions])
103
+ embedding_model = ruby_llm_v2? ? @ruby_llm_model : model_id
104
+ embedding = @ruby_llm_provider.embed(text, model: embedding_model, dimensions: parameters[:dimensions])
99
105
 
100
106
  {
101
107
  object: "embedding",
@@ -138,12 +144,15 @@ module ActiveAgent
138
144
  # Handle tool calls in chunk
139
145
  if chunk.tool_calls&.any?
140
146
  message[:tool_calls] ||= []
141
- chunk.tool_calls.each do |_id, tool_call|
142
- existing = message[:tool_calls].find { |tc| tc[:id] == tool_call.id }
147
+ chunk.tool_calls.each do |key, tool_call|
148
+ # RubyLLM 2 keys OpenAI deltas by index; only the first delta
149
+ # includes the call ID and name. Keep each index tied to its call.
150
+ existing = message[:tool_calls].find { |tc| tool_call.id && tc[:id] == tool_call.id }
151
+ existing ||= @stream_tool_calls[key]
143
152
  if existing
144
153
  existing[:function][:arguments] += tool_call.arguments.to_s if tool_call.arguments
145
154
  else
146
- message[:tool_calls] << {
155
+ existing = {
147
156
  id: tool_call.id,
148
157
  type: "function",
149
158
  function: {
@@ -151,7 +160,9 @@ module ActiveAgent
151
160
  arguments: tool_call.arguments.to_s
152
161
  }
153
162
  }
163
+ message[:tool_calls] << existing
154
164
  end
165
+ @stream_tool_calls[key] = existing
155
166
  end
156
167
  end
157
168
 
@@ -254,6 +265,10 @@ module ActiveAgent
254
265
 
255
266
  private
256
267
 
268
+ def ruby_llm_v2?
269
+ Gem.loaded_specs.fetch("ruby_llm").version.segments.first >= 2
270
+ end
271
+
257
272
  # Resolves and caches the RubyLLM provider for the given model.
258
273
  #
259
274
  # Reuses the cached provider if the model hasn't changed (e.g., during
@@ -447,6 +462,8 @@ module ActiveAgent
447
462
  # Add stop_reason if available
448
463
  if response.respond_to?(:stop_reason) && response.stop_reason
449
464
  hash[:stop_reason] = response.stop_reason
465
+ elsif response.respond_to?(:finish_reason) && response.finish_reason
466
+ hash[:stop_reason] = { stop: "end_turn", tool_calls: "tool_use" }.fetch(response.finish_reason, response.finish_reason.to_s)
450
467
  elsif response.tool_calls&.any?
451
468
  hash[:stop_reason] = "tool_use"
452
469
  else
@@ -454,7 +471,10 @@ module ActiveAgent
454
471
  end
455
472
 
456
473
  # Add usage info if available
457
- if response.respond_to?(:input_tokens) && response.input_tokens
474
+ if response.respond_to?(:tokens)
475
+ tokens = response.tokens
476
+ hash[:usage] = { input_tokens: tokens.input, output_tokens: tokens.output } if tokens.input
477
+ elsif response.respond_to?(:input_tokens) && response.input_tokens
458
478
  hash[:usage] = {
459
479
  input_tokens: response.input_tokens,
460
480
  output_tokens: response.output_tokens
@@ -1,3 +1,3 @@
1
1
  module ActiveAgent
2
- VERSION = "1.7.1"
2
+ VERSION = "1.8.0"
3
3
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: activeagent
3
3
  version: !ruby/object:Gem::Version
4
- version: 1.7.1
4
+ version: 1.8.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Justin Bowen
@@ -204,16 +204,22 @@ dependencies:
204
204
  name: ruby_llm
205
205
  requirement: !ruby/object:Gem::Requirement
206
206
  requirements:
207
- - - "~>"
207
+ - - ">="
208
+ - !ruby/object:Gem::Version
209
+ version: '1.16'
210
+ - - "<"
208
211
  - !ruby/object:Gem::Version
209
- version: '1.0'
212
+ version: '3'
210
213
  type: :development
211
214
  prerelease: false
212
215
  version_requirements: !ruby/object:Gem::Requirement
213
216
  requirements:
214
- - - "~>"
217
+ - - ">="
218
+ - !ruby/object:Gem::Version
219
+ version: '1.16'
220
+ - - "<"
215
221
  - !ruby/object:Gem::Version
216
- version: '1.0'
222
+ version: '3'
217
223
  - !ruby/object:Gem::Dependency
218
224
  name: capybara
219
225
  requirement: !ruby/object:Gem::Requirement