activeagent 1.7.2 → 1.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 881ca4e27488bd719dda01c80e6241bf6d17082876039033c22f9265d1e559aa
4
- data.tar.gz: 62526259e8a48d5243a480cf3ac6f795c1672b20e8569f7804fcf13af4d25ec4
3
+ metadata.gz: 3642068294052d483decf5260b5d7dd1793647f06920713ca658f7b89ba948b0
4
+ data.tar.gz: 1d0811bee23dc646e916d17abc711637c1e2d41055633ace8bb1c067615d26b4
5
5
  SHA512:
6
- metadata.gz: 0d135cc35623442cdf38df2d588e323fec82f7e16c7d9c8fcf90d2add91bbb55476e07fe548c70ddb2b22cc709e684ab8ccc53dd337c0650800f9d6dd268f547
7
- data.tar.gz: 76a0b25ef5cf125e99399ff61cef131d2fe6ae6dade18173601f73f0ec7ecedde6bfc0738af99ac490444b1a067cd136dc2caac34cc1f1cd6c01996cbdddf841
6
+ metadata.gz: 6d4492797c3d4b238bf86c49c48b9ec12e5876bbc9ef075c4cae7f369a3e3274b3db30781b9367643740bd9fa365eb8fe8ad81308b125f6843f3f828ccf5f339
7
+ data.tar.gz: 601676a4590808eef1469f3cfafc5136c43bccbd22e7b600c15940eb7a1fc957479c568676d699d7d41c10836fa621eb072a3ff5fd20d5f4cfa22a517e982416
data/CHANGELOG.md CHANGED
@@ -7,6 +7,196 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [1.8.0] - 2026-09-29
11
+
12
+ Releases `activeagent` and `actionagent` 1.8.0 from one tag. A minor release.
13
+ Settings -> Integrations connects GitHub and Claude Code. A connected
14
+ repository's checkout boots as a sandbox, on a developer's machine with the new
15
+ `:local` backend, where Claude Code sessions run and an evaluation can run
16
+ against the checkout without the agent being edited. The dashboard's MCP
17
+ server gains evaluation and telemetry tools for a developer's own coding
18
+ harness. Ollama hosts can be tested and can be remote, with an optional Bearer
19
+ API key. Comparison runs lead with a per-model table and filter the fix list
20
+ by model. A publisher can check its collector before a run.
21
+
22
+ Upgrading: run `bin/rails generate action_agent:install --skip` and
23
+ `bin/rails db:migrate`. The generator adds what an install lacks:
24
+ `provider_keys.api_key` and the `github_connections` and `code_sessions`
25
+ tables. An app on the RubyLLM provider needs ruby_llm 1.16 or later; 2.x
26
+ works too.
27
+
28
+ The Claude Code connection stores Anthropic API keys only. Anthropic does not
29
+ let third-party products collect, store or route requests through Claude.ai
30
+ subscription credentials
31
+ ([Claude Code legal and compliance](https://code.claude.com/docs/en/legal-and-compliance.md)),
32
+ so a `claude setup-token` token (`sk-ant-oat…`) is refused. An install that
33
+ stored one while running a pre-release build never hands it to a session: its
34
+ owner sees Claude Code as needing an API key (`needs_replacing: true` in
35
+ `GET /api/provider_keys`) until they paste one.
36
+ `bin/rails action_agent:claude_code:purge_subscription_tokens` deletes the
37
+ stored tokens and prints how many it removed. A developer who wants sessions on
38
+ their own Claude login sets `config.claude_code_auth = :local_login` with the
39
+ `:local` backend and runs `claude /login` on that machine instead.
40
+
41
+ ### Added
42
+
43
+ - **Evaluation and telemetry tools on the MCP facade** (`actionagent`). The
44
+ dashboard's MCP server now offers `evaluations_list`, `evaluations_get`,
45
+ `evaluations_run`, `evaluation_runs_get`, `evaluation_runs_compare`,
46
+ `traces_search` and `traces_get`, so a developer's own coding harness can
47
+ run an agent's evaluations, read the fix items and failing traces, and
48
+ iterate on the agent in its own checkout without the dashboard holding a
49
+ model login. The tools read under the API key's owner exactly as the JSON
50
+ API reads under the signed-in owner, run through the same execution,
51
+ quota and sandbox checks as `POST /api/evaluations/:id/run`, bound their
52
+ output, and mask the owner's credentials. Their names cannot collide with
53
+ schema tools or agent tools. `ActionAgent.mcp_dashboard_tools = false`
54
+ turns them off.
55
+
56
+ - **Local checkout sandboxes and Claude Code sessions** (`actionagent`, #489).
57
+ A new `:local` sandbox backend (`config.sandbox_service = :local`) makes
58
+ **Start sandbox** work on a developer's machine without containers. It
59
+ clones the repository under `tmp/action_agent/sandboxes`, runs the setup
60
+ the checkout's optional `.activeagents/sandbox.yml` names (`env`, `setup`,
61
+ `manifest`, `start`), and boots the app on `127.0.0.1`. It then registers
62
+ the app's MCP facade as the `sandbox:<session_id>` server. Every process
63
+ starts from a sanitized copy of the dashboard's environment: without its
64
+ database and Redis URLs, Rails keys and environment, Bundler and Ruby
65
+ settings, git repository and config variables, `SSH_AUTH_SOCK`,
66
+ model-provider and Claude Code settings, variables named like a secret, or
67
+ URLs carrying credentials. The GitHub token reaches
68
+ only the fetch, and the Claude Code API key reaches only Claude Code.
69
+ `:local` runs the owner's code with the dashboard's privileges, so it is off
70
+ outside development and test unless
71
+ `ActionAgent.local_sandboxes_enabled = true`.
72
+ - `bin/rails action_agent:sandbox:manifest` writes the booted app's
73
+ `{mcp_path, mcp_token}` for the backend.
74
+ - `bin/rails action_agent:sandbox:reap` expires overdue sandboxes and stops
75
+ them.
76
+ - A ready sandbox runs headless Claude Code sessions (`--permission-mode
77
+ acceptEdits` by default). You start them from Settings -> Integrations, or
78
+ through `/api/sandboxes/:session_id/code_sessions`. Their events stream
79
+ into the dashboard, and the checkout's diff follows.
80
+ - New options: `local_sandboxes_enabled`, `local_sandbox_root`,
81
+ `local_sandbox_boot_timeout`, `claude_code_command`,
82
+ `claude_code_permission_mode`, `claude_code_max_turns`,
83
+ `claude_code_timeout` and `claude_code_auth`.
84
+ - `claude_code_auth = :local_login` runs sessions on the machine's own
85
+ Claude Code login (`claude /login`), with no stored key: the backend
86
+ passes no credential and no `CLAUDE_CONFIG_DIR`, so `claude` uses the
87
+ dashboard user's own `~/.claude` or keychain, which the dashboard never
88
+ reads. `LocalSandboxBackend.claude_login_status` asks
89
+ `claude auth status --json` (cached for a minute) and keeps only
90
+ `loggedIn` and the login method. `GET /api/sandboxes` reports
91
+ `claude_code_auth` and, in this mode, `claude_code_login`
92
+ (`{ logged_in, auth_method }`), and the assistant's
93
+ `connections.claude_code` the same as `auth` and `login`. Other backends
94
+ refuse sessions in this mode (`code_sessions_supported: false`, and a
95
+ `422` naming the reason).
96
+ - `app_runtime` sandboxes now provision in the background and last 2 hours.
97
+ - Run `rails g action_agent:install` to add the
98
+ `create_active_agent_code_sessions` migration.
99
+ - This repository's own `.activeagents/sandbox.yml` boots `test/dummy`.
100
+ - Every `:local` sandbox boots on databases of its own, so a checkout of
101
+ the dashboard's own app no longer migrates the developer's development
102
+ database. The backend reads the adapter from the checkout's
103
+ `config/database.yml` without running its ERB, and sets `DATABASE_URL`
104
+ and `<NAME>_DATABASE_URL` (`QUEUE_DATABASE_URL`, `CACHE_DATABASE_URL`):
105
+ SQLite files in the workspace, or `<database>_sandbox_<id>` on
106
+ PostgreSQL and MySQL, which terminate drops with the checkout's
107
+ Rails database tasks restricted to the names and URLs recorded at boot.
108
+ Setting a variable in `sandbox.yml`'s `env` overrides it and excludes
109
+ that database from cleanup. Replica mappings follow their own writer;
110
+ ambiguous mappings require an explicit URL instead of guessing.
111
+ - The Claude Code panel has a **Model** select: Claude Code's own
112
+ default, the `sonnet`, `opus` and `haiku` aliases, or any model id under
113
+ *Other…*. It remembers the last choice per browser, and each session
114
+ shows the model it ran on.
115
+ - A run can use a checkout sandbox without the agent being edited:
116
+ `sandbox_id` on `POST /api/evaluations/:id/run` (and on the runner's
117
+ `/api/agents/:id/execute` and `/test`) gives that run's tool dispatcher
118
+ the sandbox's `sandbox:<session_id>` runtime, as if the agent listed it.
119
+ The sandbox must be the caller's, a ready `app_runtime` sandbox, and the
120
+ agent owner's; anything else is a `422`. The run records which sandbox
121
+ it used (`run.sandbox`), and a scenario suite's **Run against sandbox**
122
+ select, its Runs list and the run report show it. The selected sandbox
123
+ takes precedence for matching tool names, with one schema per name.
124
+ Queued agent runs fail if their selected sandbox stops, and failed
125
+ sandbox discovery never silently falls back to the original tools.
126
+ - **Claude Code connection** (`actionagent`, #478). Settings -> Integrations
127
+ stores an Anthropic API key (`sk-ant-api…`, from the Claude Console) as the
128
+ `claude_code` provider key. It is encrypted, write-only, and not an agent
129
+ provider. `SandboxSession#runtime_environment` hands it to an
130
+ `app_runtime` backend as `ANTHROPIC_API_KEY`, so the checkout can run
131
+ Claude Code sessions. Claude subscription tokens (`claude setup-token`)
132
+ are refused, as Anthropic's terms require (see the upgrading note above).
133
+ `/api/provider_keys` rows now carry `kind` (`key`, `host` or
134
+ `connection`) and `needs_replacing`.
135
+ - **GitHub connections and checkout sandboxes** (`actionagent`, #477).
136
+ Settings -> Integrations connects GitHub over OAuth
137
+ (`ActionAgent.github_client_id` / `github_client_secret`, or
138
+ `GITHUB_CLIENT_ID` / `GITHUB_CLIENT_SECRET`), stores the token encrypted,
139
+ and lets the owner choose which repositories the workspace may use. Only
140
+ repositories GitHub lists for the token can be selected. A new
141
+ `app_runtime` sandbox type checks out one of them: the backend receives
142
+ `sandbox_session.checkout_spec` (repository, ref, clone URL, token),
143
+ boots the app, and returns `mcp_url` / `mcp_token` from `create_sandbox`.
144
+ The session is then an MCP server keyed `sandbox:<session_id>`. An agent
145
+ that lists that key in `mcp_servers` runs, and is evaluated, with the
146
+ checkout's own tools. Lookups are scoped to the agent's owner. Run
147
+ `rails g action_agent:install` to add the
148
+ `create_active_agent_github_connections` migration.
149
+
150
+ - **Ollama hosts are testable and can be remote** (`actionagent`). Settings ->
151
+ Provider API Keys gains a **Test connection** for Ollama that reports
152
+ whether the server is reachable, the round-trip time and the models it
153
+ serves, before or after saving (`POST <mount>/api/provider_keys/test`,
154
+ read-only). The host is accepted as a bare server address
155
+ (`http://localhost:11434`; the OpenAI-compatible `/v1` path is added) and
156
+ an optional **API key** is stored beside it and sent as a Bearer token,
157
+ for a server behind an authenticating proxy or Ollama Cloud. The agent
158
+ builder's live Ollama model list uses the same probe and key. When no host
159
+ is configured the card shows the host app's `config/active_agent.yml`
160
+ default. The install generator emits a guarded `add_provider_key_api_key`
161
+ migration for existing installs; re-run
162
+ `bin/rails generate action_agent:install --skip` and `bin/rails db:migrate`.
163
+ - **A model comparison table on comparison runs** (`actionagent`,
164
+ `activeagent`). A run over several models now leads its Models section
165
+ with one row per model, best first: passed, mean score, average latency,
166
+ average tokens per scenario, cost (and per scenario), and the model's
167
+ typical fault — its most frequent one with the diagnosis of a result that
168
+ carries it. The scenario suite panel, the sampling run detail and the
169
+ standalone HTML report (`ActiveAgent::Evals::ReportHtml`) all render it.
170
+ - **What to fix, filtered by model** (`actionagent`, `activeagent`). On a
171
+ comparison run the fix list takes a model chip, narrowing to the items
172
+ attributed to that model and counting what that model alone produced,
173
+ since one model may need more instruction than another. The standalone
174
+ report filters through radio chips and stylesheet rules — it still ships
175
+ no script.
176
+
177
+ - **A publisher can check its collector before a run** (`activeagent`).
178
+ `ActiveAgent::Evals::Publisher#verify!` asks the collector whether it is up
179
+ and accepts the key before a run is paid for. It posts an empty JSON object,
180
+ which a compatible collector refuses with a 422 naming `version`, without
181
+ storing anything; anything else raises `Publisher::Error` with a delivery's
182
+ status, detail and guidance. `Publisher#endpoint` returns the collector URL.
183
+
184
+ ### Changed
185
+
186
+ - **Collector rejections say what the status means** (`activeagent`). A
187
+ `Publisher::Error` for a 401, 403, 404, 415 or 501 rejection names a refused
188
+ key, an account an operator must act on, an endpoint that is not a
189
+ collector, a rewritten `Content-Type`, or an install with no evaluation
190
+ store, in place of the generic guidance.
191
+
192
+ ### Fixed
193
+
194
+ - **A nested scenario expectation written as one value** (`activeagent`).
195
+ `ScenarioParser` now stores `{ expectations: { contains: "30" } }` as a
196
+ list of one, the shape the persisted scenario and the dashboard's matrix
197
+ read; an object-list import with a lone value used to break the suite
198
+ panel. The matrix also tolerates scenarios persisted before this.
199
+
10
200
  ## [1.7.2] - 2026-09-29
11
201
 
12
202
  Releases `activeagent` and `actionagent` 1.7.2 from one tag. A patch on 1.7.1:
@@ -11,7 +11,8 @@ module ActiveAgent
11
11
  # Publishes a completed report without replaying the agent. The caller must
12
12
  # retain run_id when retrying: compatible collectors treat that identity as
13
13
  # immutable within the authenticated account. Delivery is blocking and does
14
- # not follow redirects with the account's bearer credential.
14
+ # not follow redirects with the account's bearer credential. +verify!+ asks
15
+ # the collector whether it is up and accepts the key before a run is paid for.
15
16
  #
16
17
  # Every failure to deliver raises Error. Invalid arguments raise
17
18
  # ArgumentError before anything is sent.
@@ -20,6 +21,9 @@ module ActiveAgent
20
21
  MAX_BYTES = 2 * 1024 * 1024
21
22
  DETAIL_LIMIT = 200
22
23
 
24
+ # @return [String] the collector URL reports go to
25
+ attr_reader :endpoint
26
+
23
27
  # Raised for every failed delivery. Only a network failure keeps the
24
28
  # underlying error as its +cause+, so neither the response nor the
25
29
  # report reaches a log through the exception chain.
@@ -48,10 +52,15 @@ module ActiveAgent
48
52
  # Whether each rejection status is retryable, and what the caller should
49
53
  # do about it. Other statuses fall back to the rules in +rejection+.
50
54
  REJECTIONS = {
55
+ 401 => [ false, "the collector refused the API key; check the key against the collector's account" ],
56
+ 403 => [ false, "the account may not store this report until an operator acts, for example on a cap on observed agents, evaluations or scenarios; resolve that before retrying" ],
57
+ 404 => [ false, "nothing at the endpoint takes evaluation reports; check that it is a collector's /v1/evaluations or <mount>/api/evaluation_reports URL" ],
51
58
  409 => [ false, "the collector already holds a different report under this run_id; never retry this report with the same run_id" ],
52
59
  413 => [ false, "the report exceeds the collector's size limit; publish a smaller selection" ],
60
+ 415 => [ false, "the collector did not receive application/json; check anything between the publisher and the collector that rewrites the Content-Type" ],
53
61
  422 => [ false, "correct what the collector refused before retrying" ],
54
- 429 => [ true, "the account is over its quota or rate limit; retain the report and run_id and retry later" ]
62
+ 429 => [ true, "the account is over its quota or rate limit; retain the report and run_id and retry later" ],
63
+ 501 => [ false, "the collector has no evaluation store; migrate the install, or publish to one generated with evaluation tables" ]
55
64
  }.freeze
56
65
 
57
66
  # The key is sent as a bearer token and filtered from the collector's
@@ -65,6 +74,7 @@ module ActiveAgent
65
74
  unless @uri.scheme == "https" || %w[localhost 127.0.0.1 ::1].include?(@uri.hostname)
66
75
  raise ArgumentError, "Evaluation endpoint requires HTTPS except on loopback hosts"
67
76
  end
77
+ @endpoint = @uri.to_s
68
78
 
69
79
  @api_key = api_key.to_s.strip
70
80
  raise ArgumentError, "Evaluation API key is required" if @api_key.empty?
@@ -91,6 +101,52 @@ module ActiveAgent
91
101
  body = encode(identities.merge("version" => 1, "report" => report_hash(report)))
92
102
  raise Error, "Evaluation report exceeds the 2 MiB delivery limit; publish a smaller selection" if body.bytesize > MAX_BYTES
93
103
 
104
+ deliver("retain the report and run_id for retry") do
105
+ response = post(body)
106
+ raise rejection(response) unless %w[200 201].include?(response.code)
107
+
108
+ receipt = JSON.parse(response.body.to_s)
109
+ unless receipt.is_a?(Hash) && receipt["run_id"] == run_id && receipt["status"] == "complete" && receipt["id"] && receipt["evaluation_id"]
110
+ raise Error.new("Evaluation collector returned an invalid completion receipt; retain the report and run_id for retry", retryable: true)
111
+ end
112
+ receipt
113
+ end
114
+ end
115
+
116
+ # Returns true when the collector is up and accepts the API key, without
117
+ # storing anything. It posts an empty JSON object, which a compatible
118
+ # collector authenticates, parses, and then refuses with a 422 whose
119
+ # +error+ names +version+, as not a version-1 report. That refusal, and
120
+ # only that one, is the ready answer. Call it before an expensive run, so
121
+ # a stopped collector or a refused key fails before the first model call
122
+ # rather than after the last one.
123
+ #
124
+ # Raises Error for anything else: a rejection other than 422, with the
125
+ # status, detail and guidance +call+ would carry (401 for a refused key,
126
+ # 404 when nothing at the endpoint takes reports); a retryable delivery
127
+ # failure when the collector cannot be reached; a 422 that says anything
128
+ # else; or a collector that stores the empty object. The last two mean the
129
+ # endpoint is not a compatible collector.
130
+ def verify!
131
+ deliver("start the collector or check the endpoint, then verify again") do
132
+ response = post("{}")
133
+ if response.code == "422"
134
+ detail = collector_detail(response.body)
135
+ next true if detail&.match?(/\bversion\b/)
136
+
137
+ raise Error.new("Evaluation collector answered HTTP 422 without refusing the empty envelope as a version-1 report, " \
138
+ "so it is not a compatible collector; check the endpoint", status: 422, detail: detail)
139
+ end
140
+ raise rejection(response) unless %w[200 201].include?(response.code)
141
+
142
+ raise Error, "Evaluation collector stored an empty report, so it is not a compatible collector; check the endpoint"
143
+ end
144
+ end
145
+
146
+ private
147
+
148
+ # Posts +body+ to the endpoint with the bearer key and the delivery timeouts.
149
+ def post(body)
94
150
  http = Net::HTTP.new(@uri.hostname, @uri.port)
95
151
  http.use_ssl = @uri.scheme == "https"
96
152
  http.open_timeout = @open_timeout
@@ -101,26 +157,25 @@ module ActiveAgent
101
157
  request["Content-Type"] = "application/json"
102
158
  request["Accept"] = "application/json"
103
159
  request.body = body
104
- response = http.request(request)
105
- raise rejection(response) unless %w[200 201].include?(response.code)
160
+ http.request(request)
161
+ end
106
162
 
107
- receipt = JSON.parse(response.body.to_s)
108
- unless receipt.is_a?(Hash) && receipt["run_id"] == run_id && receipt["status"] == "complete" && receipt["id"] && receipt["evaluation_id"]
109
- raise Error.new("Evaluation collector returned an invalid completion receipt; retain the report and run_id for retry", retryable: true)
110
- end
111
- receipt
163
+ # Runs one exchange with the collector, turning every failure to reach it
164
+ # or to read its answer into a retryable Error that quotes nothing from
165
+ # the response and ends with +guidance+, what the caller should do next.
166
+ # An Error the block raises passes through unchanged.
167
+ def deliver(guidance)
168
+ yield
112
169
  rescue JSON::ParserError
113
170
  # The parser's message quotes the body.
114
- raise Error.new("Evaluation collector returned invalid JSON; retain the report and run_id for retry", retryable: true), cause: nil
171
+ raise Error.new("Evaluation collector returned invalid JSON; #{guidance}", retryable: true), cause: nil
115
172
  rescue Net::HTTPBadResponse, Net::HTTPHeaderSyntaxError, Zlib::Error => e
116
173
  # These messages can quote the response's status line, headers or body.
117
- raise Error.new("Evaluation collector returned a malformed response (#{e.class}); retain the report and run_id for retry", retryable: true), cause: nil
174
+ raise Error.new("Evaluation collector returned a malformed response (#{e.class}); #{guidance}", retryable: true), cause: nil
118
175
  rescue IOError, SocketError, SystemCallError, Timeout::Error, OpenSSL::SSL::SSLError => e
119
- raise Error.new("Evaluation delivery failed (#{e.class}); retain the report and run_id for retry", retryable: true)
176
+ raise Error.new("Evaluation delivery failed (#{e.class}); #{guidance}", retryable: true)
120
177
  end
121
178
 
122
- private
123
-
124
179
  def report_hash(report)
125
180
  hash = report.to_h if report.respond_to?(:to_h) && !report.nil? && !report.is_a?(Array)
126
181
  raise ArgumentError, "report must be a Report or its saved JSON hash" unless hash.is_a?(Hash)
@@ -178,6 +178,9 @@ module ActiveAgent
178
178
  DesignTokens.css(scope: ":root.theme-dark", tokens: DesignTokens::DARK, color_scheme: "dark"),
179
179
  STYLES,
180
180
  ".mx { grid-template-columns: minmax(240px, 1.6fr) 150px repeat(#{@models.size}, minmax(170px, 1fr)); }",
181
+ # One rule per model: with that chip checked, hide every fix card
182
+ # attributed to other models (cards attributed to none stay).
183
+ *@models.each_index.map { |i| ".fix-section:has(input[value=\"m#{i}\"]:checked) .fix[data-models]:not([data-models~=\"m#{i}\"]) { display: none; }" },
181
184
  ".matrix .inner { min-width: #{390 + 185 * @models.size}px; }"
182
185
  ].join("\n")
183
186
  end
@@ -233,12 +236,78 @@ module ActiveAgent
233
236
  <<~PANEL
234
237
  <div class="panel">
235
238
  <div class="panel-head"><span class="micro">Models</span><span class="right">judged by #{h(judged_by)}</span></div>
239
+ #{html_comparison_table if comparing?}
236
240
  #{blocks.join}
237
241
  #{verdict_row}
238
242
  </div>
239
243
  PANEL
240
244
  end
241
245
 
246
+ # The comparison read across: one row per model, best first (pass rate,
247
+ # then mean score) — passed, mean score, average latency, average
248
+ # tokens per scenario, cost, and the model's typical fault. The blocks
249
+ # under it carry the same figures per model with bars and every fault.
250
+ def html_comparison_table
251
+ rows = summary_by_model.sort_by do |label, stats|
252
+ total = stats["scenarios"].to_i
253
+ [ total.positive? ? -stats["passed"].to_f / total : 0.0, -(stats["avg_score"] || -1).to_f, @models.index(model_by_label(label)).to_i ]
254
+ end
255
+
256
+ <<~TABLE
257
+ <div class="compare"><table>
258
+ <thead><tr><th>Model</th><th class="num">Passed</th><th class="num">Mean score</th><th class="num">Avg latency</th><th class="num" title="Average input + output tokens per scenario">Avg tokens</th><th class="num" title="Cohort spend, and per scenario">Cost</th><th class="fault">Typical fault</th></tr></thead>
259
+ <tbody>#{rows.map { |label, stats| html_comparison_row(label, stats) }.join}</tbody>
260
+ </table></div>
261
+ TABLE
262
+ end
263
+
264
+ def html_comparison_row(label, stats)
265
+ short, provider = split_label(model_by_label(label))
266
+ total = stats["scenarios"].to_i
267
+ ratio = total.positive? ? stats["passed"].to_f / total : 0.0
268
+ pick = comparing? && winner == label ? %(<span class="pick" title="picked by the judge">★ pick</span>) : ""
269
+ per = ->(value) { value.nil? || total.zero? ? nil : value.to_f / total }
270
+ avg_tokens = per.call(stats["input_tokens"].to_i + stats["output_tokens"].to_i)
271
+ tokens_cell = avg_tokens ? h(fmt_k(avg_tokens.round)) : "—"
272
+ tokens_title = avg_tokens ? %( title="#{per.call(stats['input_tokens']).to_f.round} in · #{per.call(stats['output_tokens']).to_f.round} out per scenario") : ""
273
+ per_cost = per.call(stats["cost"])
274
+ cost_cell = stats["cost"].nil? ? "—" : h(fmt_cost(stats["cost"]))
275
+ cost_cell += "<span class=\"per\">#{h(fmt_cost(per_cost))}/scenario</span>" if per_cost
276
+
277
+ <<~ROW
278
+ <tr>
279
+ <td class="model-cell"><span class="name">#{h(short)}</span>#{pick}<span class="provider">#{h(provider)}</span></td>
280
+ <td class="num ratio tone-#{tone_for(ratio)}">#{total.positive? ? "#{stats['passed']}/#{total}" : '—'}</td>
281
+ <td class="num">#{h(fmt_mean_score(stats['avg_score']))}</td>
282
+ <td class="num">#{h(fmt_ms(stats['avg_duration_ms']))}</td>
283
+ <td class="num"#{tokens_title}>#{tokens_cell}</td>
284
+ <td class="num">#{cost_cell}</td>
285
+ <td class="fault">#{typical_fault_text(label, stats)}</td>
286
+ </tr>
287
+ ROW
288
+ end
289
+
290
+ # "missing content ×2 · refund_window: The answer is missing expected
291
+ # content: 30." — the model's most frequent fault, and the diagnosis of
292
+ # the first result that carries it; "no faults" for a clean cohort.
293
+ def typical_fault_text(label, stats)
294
+ tally = stats["faults"] || {}
295
+ return %(<span class="clean">no faults</span>) if tally.empty?
296
+
297
+ mine = @results.select { |result| result.label == label }
298
+ example_of = ->(fault) { mine.find { |result| result.fault == fault } }
299
+ # Most frequent first; between equals, a fault a result can explain,
300
+ # then a specific fault over the judge's catch-all, then the name.
301
+ fault, count = tally.min_by { |name, n| [ -n, example_of.call(name) ? 0 : 1, name == "low_quality" ? 1 : 0, name ] }
302
+ example = example_of.call(fault)
303
+ head = "#{fault_name(fault)} ×#{count}"
304
+ return h(head) unless example&.summary.present?
305
+
306
+ detail = "#{example.scenario.key}: #{example.summary}"
307
+ detail = "#{detail[0, 119]}…" if detail.length > 120
308
+ "#{h(head)} <span class=\"detail\">· #{h(detail)}</span>"
309
+ end
310
+
242
311
  def html_model_block(label, stats)
243
312
  short, provider = split_label(model_by_label(label))
244
313
  total = stats["scenarios"]
@@ -277,13 +346,36 @@ module ActiveAgent
277
346
  end
278
347
 
279
348
  <<~FIXES
280
- <section class="section" aria-label="Recommendations">
349
+ <section class="section fix-section" aria-label="Recommendations">
281
350
  <div class="section-head"><span class="micro">What to fix</span><span class="meta">#{h(meta)}</span></div>
351
+ #{html_fix_filter(items) if comparing? && items.any?}
282
352
  #{body}
283
353
  </section>
284
354
  FIXES
285
355
  end
286
356
 
357
+ # A model filter for the fix cards — a fault one model keeps making is
358
+ # that model's to fix, so the list narrows to what was attributed to
359
+ # it. Radio chips and stylesheet rules alone (the page carries no
360
+ # script): each card names its models in data-models, and a checked
361
+ # model hides every card that does not name it. Cards attributed to no
362
+ # model (an older run) stay under every filter.
363
+ def html_fix_filter(items)
364
+ chips = [ %(<label class="chip pick-model"><input type="radio" name="fix-model" value="all" checked><span>all models #{items.size}</span></label>) ]
365
+ @models.each_with_index do |spec, index|
366
+ count = items.count { |item| Array(item["models"]).empty? || item["models"].include?(spec.label) }
367
+ chips << %(<label class="chip pick-model"><input type="radio" name="fix-model" value="m#{index}"><span>#{h(short_name(spec))} #{count}</span></label>)
368
+ end
369
+ %(<div class="fix-filter"><span class="micro sm">for</span>#{chips.join}</div>)
370
+ end
371
+
372
+ def fix_model_tokens(item)
373
+ labels = Array(item["models"])
374
+ return "" if labels.empty?
375
+
376
+ labels.filter_map { |label| (index = @models.index(model_by_label(label))) && "m#{index}" }.join(" ")
377
+ end
378
+
287
379
  def html_fix_card(item)
288
380
  tone = item["kind"] == "instruction" ? "info" : "error"
289
381
  glyph = tone == "info" ? "[i]" : "[!]"
@@ -297,7 +389,8 @@ module ActiveAgent
297
389
  parts << html_fix_server(item["server"]) if item["server"]
298
390
  parts << %(<div class="note">#{h(item['note'])}</div>) if item["note"].present?
299
391
  parts << html_fix_action(item["action"]) if item["action"]
300
- %(<div class="fix">#{parts.join}</div>)
392
+ models = fix_model_tokens(item)
393
+ %(<div class="fix"#{%( data-models="#{models}") if models.present?}>#{parts.join}</div>)
301
394
  end
302
395
 
303
396
  def html_fix_tools(item)
@@ -557,6 +650,24 @@ module ActiveAgent
557
650
  .tok .in { color: var(--color-token-in); }
558
651
  .tok .out { color: var(--color-token-out); }
559
652
  .faults { display: flex; gap: 6px; flex-wrap: wrap; }
653
+ .compare { overflow-x: auto; border-top: 1px solid var(--color-border-light); }
654
+ .compare table { width: 100%; border-collapse: collapse; }
655
+ .compare th { padding: 8px 12px; text-align: left; vertical-align: bottom; white-space: nowrap; font-family: var(--font-mono); font-size: 10px; font-weight: 600; letter-spacing: 0.05em; text-transform: uppercase; color: var(--color-text-muted); background: var(--color-muted); }
656
+ .compare td { padding: 9px 12px; vertical-align: top; border-top: 1px solid var(--color-border-light); font-size: 13px; color: var(--color-text-cell); }
657
+ .compare th.num, .compare td.num { text-align: right; }
658
+ .compare td.num { font-family: var(--font-mono); font-size: 12px; white-space: nowrap; }
659
+ .compare td.ratio { font-weight: 600; }
660
+ .compare .per { display: block; font-weight: 400; color: var(--color-text-muted); }
661
+ .compare .model-cell { white-space: nowrap; }
662
+ .compare .model-cell .name { font-family: var(--font-mono); font-size: 12px; font-weight: 600; color: var(--color-text-primary); }
663
+ .compare .model-cell .provider { display: block; font-family: var(--font-mono); font-size: 11px; color: var(--color-text-muted); }
664
+ .compare .pick { margin-left: 6px; font-family: var(--font-mono); font-size: 10px; font-weight: 700; color: var(--color-warning-text); }
665
+ .compare th.fault, .compare td.fault { width: 34%; }
666
+ .compare td.fault .detail { color: var(--color-text-secondary); }
667
+ .fix-filter { display: flex; align-items: center; gap: 6px; flex-wrap: wrap; }
668
+ .pick-model { cursor: pointer; border: 1px solid var(--color-border); background: var(--color-card); }
669
+ .pick-model input { position: absolute; opacity: 0; width: 0; height: 0; }
670
+ .pick-model:has(input:checked) { border-color: var(--color-accent-ui); background: var(--color-accent-ui-tint); color: var(--color-accent-ui); }
560
671
  .clean { font-family: var(--font-mono); font-size: 11px; color: var(--color-success-text); }
561
672
  .verdict { padding: 10px 12px; border-top: 1px solid var(--color-border-light); font-size: 12px; line-height: 18px; color: var(--color-text-cell); }
562
673
  .verdict .micro { margin-right: 8px; }
@@ -125,6 +125,10 @@ module ActiveAgent
125
125
  expectations = (entry["expectations"] || entry["expect"] || {}).to_h.stringify_keys
126
126
  %w[tools contains not_contains].each do |field|
127
127
  expectations[field] = Array(entry[field]) if entry.key?(field)
128
+ # A nested expectation written as one value ({ contains: "30" })
129
+ # is a list of one: the persisted scenario and the dashboard's
130
+ # matrix read each field as an array.
131
+ expectations[field] = Array(expectations[field]) if expectations.key?(field)
128
132
  end
129
133
 
130
134
  scenario(
@@ -1,3 +1,3 @@
1
1
  module ActiveAgent
2
- VERSION = "1.7.2"
2
+ VERSION = "1.8.0"
3
3
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: activeagent
3
3
  version: !ruby/object:Gem::Version
4
- version: 1.7.2
4
+ version: 1.8.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Justin Bowen