activeagent 1.7.2 → 1.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +190 -0
- data/lib/active_agent/evals/publisher.rb +69 -14
- data/lib/active_agent/evals/report_html.rb +113 -2
- data/lib/active_agent/evals/scenario_parser.rb +4 -0
- data/lib/active_agent/version.rb +1 -1
- metadata +1 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 3642068294052d483decf5260b5d7dd1793647f06920713ca658f7b89ba948b0
|
|
4
|
+
data.tar.gz: 1d0811bee23dc646e916d17abc711637c1e2d41055633ace8bb1c067615d26b4
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 6d4492797c3d4b238bf86c49c48b9ec12e5876bbc9ef075c4cae7f369a3e3274b3db30781b9367643740bd9fa365eb8fe8ad81308b125f6843f3f828ccf5f339
|
|
7
|
+
data.tar.gz: 601676a4590808eef1469f3cfafc5136c43bccbd22e7b600c15940eb7a1fc957479c568676d699d7d41c10836fa621eb072a3ff5fd20d5f4cfa22a517e982416
|
data/CHANGELOG.md
CHANGED
|
@@ -7,6 +7,196 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
7
7
|
|
|
8
8
|
## [Unreleased]
|
|
9
9
|
|
|
10
|
+
## [1.8.0] - 2026-09-29
|
|
11
|
+
|
|
12
|
+
Releases `activeagent` and `actionagent` 1.8.0 from one tag. A minor release.
|
|
13
|
+
Settings -> Integrations connects GitHub and Claude Code. A connected
|
|
14
|
+
repository's checkout boots as a sandbox, on a developer's machine with the new
|
|
15
|
+
`:local` backend, where Claude Code sessions run and an evaluation can run
|
|
16
|
+
against the checkout without the agent being edited. The dashboard's MCP
|
|
17
|
+
server gains evaluation and telemetry tools for a developer's own coding
|
|
18
|
+
harness. Ollama hosts can be tested and can be remote, with an optional Bearer
|
|
19
|
+
API key. Comparison runs lead with a per-model table and filter the fix list
|
|
20
|
+
by model. A publisher can check its collector before a run.
|
|
21
|
+
|
|
22
|
+
Upgrading: run `bin/rails generate action_agent:install --skip` and
|
|
23
|
+
`bin/rails db:migrate`. The generator adds what an install lacks:
|
|
24
|
+
`provider_keys.api_key` and the `github_connections` and `code_sessions`
|
|
25
|
+
tables. An app on the RubyLLM provider needs ruby_llm 1.16 or later; 2.x
|
|
26
|
+
works too.
|
|
27
|
+
|
|
28
|
+
The Claude Code connection stores Anthropic API keys only. Anthropic does not
|
|
29
|
+
let third-party products collect, store or route requests through Claude.ai
|
|
30
|
+
subscription credentials
|
|
31
|
+
([Claude Code legal and compliance](https://code.claude.com/docs/en/legal-and-compliance.md)),
|
|
32
|
+
so a `claude setup-token` token (`sk-ant-oat…`) is refused. An install that
|
|
33
|
+
stored one while running a pre-release build never hands it to a session: its
|
|
34
|
+
owner sees Claude Code as needing an API key (`needs_replacing: true` in
|
|
35
|
+
`GET /api/provider_keys`) until they paste one.
|
|
36
|
+
`bin/rails action_agent:claude_code:purge_subscription_tokens` deletes the
|
|
37
|
+
stored tokens and prints how many it removed. A developer who wants sessions on
|
|
38
|
+
their own Claude login sets `config.claude_code_auth = :local_login` with the
|
|
39
|
+
`:local` backend and runs `claude /login` on that machine instead.
|
|
40
|
+
|
|
41
|
+
### Added
|
|
42
|
+
|
|
43
|
+
- **Evaluation and telemetry tools on the MCP facade** (`actionagent`). The
|
|
44
|
+
dashboard's MCP server now offers `evaluations_list`, `evaluations_get`,
|
|
45
|
+
`evaluations_run`, `evaluation_runs_get`, `evaluation_runs_compare`,
|
|
46
|
+
`traces_search` and `traces_get`, so a developer's own coding harness can
|
|
47
|
+
run an agent's evaluations, read the fix items and failing traces, and
|
|
48
|
+
iterate on the agent in its own checkout without the dashboard holding a
|
|
49
|
+
model login. The tools read under the API key's owner exactly as the JSON
|
|
50
|
+
API reads under the signed-in owner, run through the same execution,
|
|
51
|
+
quota and sandbox checks as `POST /api/evaluations/:id/run`, bound their
|
|
52
|
+
output, and mask the owner's credentials. Their names cannot collide with
|
|
53
|
+
schema tools or agent tools. `ActionAgent.mcp_dashboard_tools = false`
|
|
54
|
+
turns them off.
|
|
55
|
+
|
|
56
|
+
- **Local checkout sandboxes and Claude Code sessions** (`actionagent`, #489).
|
|
57
|
+
A new `:local` sandbox backend (`config.sandbox_service = :local`) makes
|
|
58
|
+
**Start sandbox** work on a developer's machine without containers. It
|
|
59
|
+
clones the repository under `tmp/action_agent/sandboxes`, runs the setup
|
|
60
|
+
the checkout's optional `.activeagents/sandbox.yml` names (`env`, `setup`,
|
|
61
|
+
`manifest`, `start`), and boots the app on `127.0.0.1`. It then registers
|
|
62
|
+
the app's MCP facade as the `sandbox:<session_id>` server. Every process
|
|
63
|
+
starts from a sanitized copy of the dashboard's environment: without its
|
|
64
|
+
database and Redis URLs, Rails keys and environment, Bundler and Ruby
|
|
65
|
+
settings, git repository and config variables, `SSH_AUTH_SOCK`,
|
|
66
|
+
model-provider and Claude Code settings, variables named like a secret, or
|
|
67
|
+
URLs carrying credentials. The GitHub token reaches
|
|
68
|
+
only the fetch, and the Claude Code API key reaches only Claude Code.
|
|
69
|
+
`:local` runs the owner's code with the dashboard's privileges, so it is off
|
|
70
|
+
outside development and test unless
|
|
71
|
+
`ActionAgent.local_sandboxes_enabled = true`.
|
|
72
|
+
- `bin/rails action_agent:sandbox:manifest` writes the booted app's
|
|
73
|
+
`{mcp_path, mcp_token}` for the backend.
|
|
74
|
+
- `bin/rails action_agent:sandbox:reap` expires overdue sandboxes and stops
|
|
75
|
+
them.
|
|
76
|
+
- A ready sandbox runs headless Claude Code sessions (`--permission-mode
|
|
77
|
+
acceptEdits` by default). You start them from Settings -> Integrations, or
|
|
78
|
+
through `/api/sandboxes/:session_id/code_sessions`. Their events stream
|
|
79
|
+
into the dashboard, and the checkout's diff follows.
|
|
80
|
+
- New options: `local_sandboxes_enabled`, `local_sandbox_root`,
|
|
81
|
+
`local_sandbox_boot_timeout`, `claude_code_command`,
|
|
82
|
+
`claude_code_permission_mode`, `claude_code_max_turns`,
|
|
83
|
+
`claude_code_timeout` and `claude_code_auth`.
|
|
84
|
+
- `claude_code_auth = :local_login` runs sessions on the machine's own
|
|
85
|
+
Claude Code login (`claude /login`), with no stored key: the backend
|
|
86
|
+
passes no credential and no `CLAUDE_CONFIG_DIR`, so `claude` uses the
|
|
87
|
+
dashboard user's own `~/.claude` or keychain, which the dashboard never
|
|
88
|
+
reads. `LocalSandboxBackend.claude_login_status` asks
|
|
89
|
+
`claude auth status --json` (cached for a minute) and keeps only
|
|
90
|
+
`loggedIn` and the login method. `GET /api/sandboxes` reports
|
|
91
|
+
`claude_code_auth` and, in this mode, `claude_code_login`
|
|
92
|
+
(`{ logged_in, auth_method }`), and the assistant's
|
|
93
|
+
`connections.claude_code` the same as `auth` and `login`. Other backends
|
|
94
|
+
refuse sessions in this mode (`code_sessions_supported: false`, and a
|
|
95
|
+
`422` naming the reason).
|
|
96
|
+
- `app_runtime` sandboxes now provision in the background and last 2 hours.
|
|
97
|
+
- Run `rails g action_agent:install` to add the
|
|
98
|
+
`create_active_agent_code_sessions` migration.
|
|
99
|
+
- This repository's own `.activeagents/sandbox.yml` boots `test/dummy`.
|
|
100
|
+
- Every `:local` sandbox boots on databases of its own, so a checkout of
|
|
101
|
+
the dashboard's own app no longer migrates the developer's development
|
|
102
|
+
database. The backend reads the adapter from the checkout's
|
|
103
|
+
`config/database.yml` without running its ERB, and sets `DATABASE_URL`
|
|
104
|
+
and `<NAME>_DATABASE_URL` (`QUEUE_DATABASE_URL`, `CACHE_DATABASE_URL`):
|
|
105
|
+
SQLite files in the workspace, or `<database>_sandbox_<id>` on
|
|
106
|
+
PostgreSQL and MySQL, which terminate drops with the checkout's
|
|
107
|
+
Rails database tasks restricted to the names and URLs recorded at boot.
|
|
108
|
+
Setting a variable in `sandbox.yml`'s `env` overrides it and excludes
|
|
109
|
+
that database from cleanup. Replica mappings follow their own writer;
|
|
110
|
+
ambiguous mappings require an explicit URL instead of guessing.
|
|
111
|
+
- The Claude Code panel has a **Model** select: Claude Code's own
|
|
112
|
+
default, the `sonnet`, `opus` and `haiku` aliases, or any model id under
|
|
113
|
+
*Other…*. It remembers the last choice per browser, and each session
|
|
114
|
+
shows the model it ran on.
|
|
115
|
+
- A run can use a checkout sandbox without the agent being edited:
|
|
116
|
+
`sandbox_id` on `POST /api/evaluations/:id/run` (and on the runner's
|
|
117
|
+
`/api/agents/:id/execute` and `/test`) gives that run's tool dispatcher
|
|
118
|
+
the sandbox's `sandbox:<session_id>` runtime, as if the agent listed it.
|
|
119
|
+
The sandbox must be the caller's, a ready `app_runtime` sandbox, and the
|
|
120
|
+
agent owner's; anything else is a `422`. The run records which sandbox
|
|
121
|
+
it used (`run.sandbox`), and a scenario suite's **Run against sandbox**
|
|
122
|
+
select, its Runs list and the run report show it. The selected sandbox
|
|
123
|
+
takes precedence for matching tool names, with one schema per name.
|
|
124
|
+
Queued agent runs fail if their selected sandbox stops, and failed
|
|
125
|
+
sandbox discovery never silently falls back to the original tools.
|
|
126
|
+
- **Claude Code connection** (`actionagent`, #478). Settings -> Integrations
|
|
127
|
+
stores an Anthropic API key (`sk-ant-api…`, from the Claude Console) as the
|
|
128
|
+
`claude_code` provider key. It is encrypted, write-only, and not an agent
|
|
129
|
+
provider. `SandboxSession#runtime_environment` hands it to an
|
|
130
|
+
`app_runtime` backend as `ANTHROPIC_API_KEY`, so the checkout can run
|
|
131
|
+
Claude Code sessions. Claude subscription tokens (`claude setup-token`)
|
|
132
|
+
are refused, as Anthropic's terms require (see the upgrading note above).
|
|
133
|
+
`/api/provider_keys` rows now carry `kind` (`key`, `host` or
|
|
134
|
+
`connection`) and `needs_replacing`.
|
|
135
|
+
- **GitHub connections and checkout sandboxes** (`actionagent`, #477).
|
|
136
|
+
Settings -> Integrations connects GitHub over OAuth
|
|
137
|
+
(`ActionAgent.github_client_id` / `github_client_secret`, or
|
|
138
|
+
`GITHUB_CLIENT_ID` / `GITHUB_CLIENT_SECRET`), stores the token encrypted,
|
|
139
|
+
and lets the owner choose which repositories the workspace may use. Only
|
|
140
|
+
repositories GitHub lists for the token can be selected. A new
|
|
141
|
+
`app_runtime` sandbox type checks out one of them: the backend receives
|
|
142
|
+
`sandbox_session.checkout_spec` (repository, ref, clone URL, token),
|
|
143
|
+
boots the app, and returns `mcp_url` / `mcp_token` from `create_sandbox`.
|
|
144
|
+
The session is then an MCP server keyed `sandbox:<session_id>`. An agent
|
|
145
|
+
that lists that key in `mcp_servers` runs, and is evaluated, with the
|
|
146
|
+
checkout's own tools. Lookups are scoped to the agent's owner. Run
|
|
147
|
+
`rails g action_agent:install` to add the
|
|
148
|
+
`create_active_agent_github_connections` migration.
|
|
149
|
+
|
|
150
|
+
- **Ollama hosts are testable and can be remote** (`actionagent`). Settings ->
|
|
151
|
+
Provider API Keys gains a **Test connection** for Ollama that reports
|
|
152
|
+
whether the server is reachable, the round-trip time and the models it
|
|
153
|
+
serves, before or after saving (`POST <mount>/api/provider_keys/test`,
|
|
154
|
+
read-only). The host is accepted as a bare server address
|
|
155
|
+
(`http://localhost:11434`; the OpenAI-compatible `/v1` path is added) and
|
|
156
|
+
an optional **API key** is stored beside it and sent as a Bearer token,
|
|
157
|
+
for a server behind an authenticating proxy or Ollama Cloud. The agent
|
|
158
|
+
builder's live Ollama model list uses the same probe and key. When no host
|
|
159
|
+
is configured the card shows the host app's `config/active_agent.yml`
|
|
160
|
+
default. The install generator emits a guarded `add_provider_key_api_key`
|
|
161
|
+
migration for existing installs; re-run
|
|
162
|
+
`bin/rails generate action_agent:install --skip` and `bin/rails db:migrate`.
|
|
163
|
+
- **A model comparison table on comparison runs** (`actionagent`,
|
|
164
|
+
`activeagent`). A run over several models now leads its Models section
|
|
165
|
+
with one row per model, best first: passed, mean score, average latency,
|
|
166
|
+
average tokens per scenario, cost (and per scenario), and the model's
|
|
167
|
+
typical fault — its most frequent one with the diagnosis of a result that
|
|
168
|
+
carries it. The scenario suite panel, the sampling run detail and the
|
|
169
|
+
standalone HTML report (`ActiveAgent::Evals::ReportHtml`) all render it.
|
|
170
|
+
- **What to fix, filtered by model** (`actionagent`, `activeagent`). On a
|
|
171
|
+
comparison run the fix list takes a model chip, narrowing to the items
|
|
172
|
+
attributed to that model and counting what that model alone produced,
|
|
173
|
+
since one model may need more instruction than another. The standalone
|
|
174
|
+
report filters through radio chips and stylesheet rules — it still ships
|
|
175
|
+
no script.
|
|
176
|
+
|
|
177
|
+
- **A publisher can check its collector before a run** (`activeagent`).
|
|
178
|
+
`ActiveAgent::Evals::Publisher#verify!` asks the collector whether it is up
|
|
179
|
+
and accepts the key before a run is paid for. It posts an empty JSON object,
|
|
180
|
+
which a compatible collector refuses with a 422 naming `version`, without
|
|
181
|
+
storing anything; anything else raises `Publisher::Error` with a delivery's
|
|
182
|
+
status, detail and guidance. `Publisher#endpoint` returns the collector URL.
|
|
183
|
+
|
|
184
|
+
### Changed
|
|
185
|
+
|
|
186
|
+
- **Collector rejections say what the status means** (`activeagent`). A
|
|
187
|
+
`Publisher::Error` for a 401, 403, 404, 415 or 501 rejection names a refused
|
|
188
|
+
key, an account an operator must act on, an endpoint that is not a
|
|
189
|
+
collector, a rewritten `Content-Type`, or an install with no evaluation
|
|
190
|
+
store, in place of the generic guidance.
|
|
191
|
+
|
|
192
|
+
### Fixed
|
|
193
|
+
|
|
194
|
+
- **A nested scenario expectation written as one value** (`activeagent`).
|
|
195
|
+
`ScenarioParser` now stores `{ expectations: { contains: "30" } }` as a
|
|
196
|
+
list of one, the shape the persisted scenario and the dashboard's matrix
|
|
197
|
+
read; an object-list import with a lone value used to break the suite
|
|
198
|
+
panel. The matrix also tolerates scenarios persisted before this.
|
|
199
|
+
|
|
10
200
|
## [1.7.2] - 2026-09-29
|
|
11
201
|
|
|
12
202
|
Releases `activeagent` and `actionagent` 1.7.2 from one tag. A patch on 1.7.1:
|
|
@@ -11,7 +11,8 @@ module ActiveAgent
|
|
|
11
11
|
# Publishes a completed report without replaying the agent. The caller must
|
|
12
12
|
# retain run_id when retrying: compatible collectors treat that identity as
|
|
13
13
|
# immutable within the authenticated account. Delivery is blocking and does
|
|
14
|
-
# not follow redirects with the account's bearer credential.
|
|
14
|
+
# not follow redirects with the account's bearer credential. +verify!+ asks
|
|
15
|
+
# the collector whether it is up and accepts the key before a run is paid for.
|
|
15
16
|
#
|
|
16
17
|
# Every failure to deliver raises Error. Invalid arguments raise
|
|
17
18
|
# ArgumentError before anything is sent.
|
|
@@ -20,6 +21,9 @@ module ActiveAgent
|
|
|
20
21
|
MAX_BYTES = 2 * 1024 * 1024
|
|
21
22
|
DETAIL_LIMIT = 200
|
|
22
23
|
|
|
24
|
+
# @return [String] the collector URL reports go to
|
|
25
|
+
attr_reader :endpoint
|
|
26
|
+
|
|
23
27
|
# Raised for every failed delivery. Only a network failure keeps the
|
|
24
28
|
# underlying error as its +cause+, so neither the response nor the
|
|
25
29
|
# report reaches a log through the exception chain.
|
|
@@ -48,10 +52,15 @@ module ActiveAgent
|
|
|
48
52
|
# Whether each rejection status is retryable, and what the caller should
|
|
49
53
|
# do about it. Other statuses fall back to the rules in +rejection+.
|
|
50
54
|
REJECTIONS = {
|
|
55
|
+
401 => [ false, "the collector refused the API key; check the key against the collector's account" ],
|
|
56
|
+
403 => [ false, "the account may not store this report until an operator acts, for example on a cap on observed agents, evaluations or scenarios; resolve that before retrying" ],
|
|
57
|
+
404 => [ false, "nothing at the endpoint takes evaluation reports; check that it is a collector's /v1/evaluations or <mount>/api/evaluation_reports URL" ],
|
|
51
58
|
409 => [ false, "the collector already holds a different report under this run_id; never retry this report with the same run_id" ],
|
|
52
59
|
413 => [ false, "the report exceeds the collector's size limit; publish a smaller selection" ],
|
|
60
|
+
415 => [ false, "the collector did not receive application/json; check anything between the publisher and the collector that rewrites the Content-Type" ],
|
|
53
61
|
422 => [ false, "correct what the collector refused before retrying" ],
|
|
54
|
-
429 => [ true, "the account is over its quota or rate limit; retain the report and run_id and retry later" ]
|
|
62
|
+
429 => [ true, "the account is over its quota or rate limit; retain the report and run_id and retry later" ],
|
|
63
|
+
501 => [ false, "the collector has no evaluation store; migrate the install, or publish to one generated with evaluation tables" ]
|
|
55
64
|
}.freeze
|
|
56
65
|
|
|
57
66
|
# The key is sent as a bearer token and filtered from the collector's
|
|
@@ -65,6 +74,7 @@ module ActiveAgent
|
|
|
65
74
|
unless @uri.scheme == "https" || %w[localhost 127.0.0.1 ::1].include?(@uri.hostname)
|
|
66
75
|
raise ArgumentError, "Evaluation endpoint requires HTTPS except on loopback hosts"
|
|
67
76
|
end
|
|
77
|
+
@endpoint = @uri.to_s
|
|
68
78
|
|
|
69
79
|
@api_key = api_key.to_s.strip
|
|
70
80
|
raise ArgumentError, "Evaluation API key is required" if @api_key.empty?
|
|
@@ -91,6 +101,52 @@ module ActiveAgent
|
|
|
91
101
|
body = encode(identities.merge("version" => 1, "report" => report_hash(report)))
|
|
92
102
|
raise Error, "Evaluation report exceeds the 2 MiB delivery limit; publish a smaller selection" if body.bytesize > MAX_BYTES
|
|
93
103
|
|
|
104
|
+
deliver("retain the report and run_id for retry") do
|
|
105
|
+
response = post(body)
|
|
106
|
+
raise rejection(response) unless %w[200 201].include?(response.code)
|
|
107
|
+
|
|
108
|
+
receipt = JSON.parse(response.body.to_s)
|
|
109
|
+
unless receipt.is_a?(Hash) && receipt["run_id"] == run_id && receipt["status"] == "complete" && receipt["id"] && receipt["evaluation_id"]
|
|
110
|
+
raise Error.new("Evaluation collector returned an invalid completion receipt; retain the report and run_id for retry", retryable: true)
|
|
111
|
+
end
|
|
112
|
+
receipt
|
|
113
|
+
end
|
|
114
|
+
end
|
|
115
|
+
|
|
116
|
+
# Returns true when the collector is up and accepts the API key, without
|
|
117
|
+
# storing anything. It posts an empty JSON object, which a compatible
|
|
118
|
+
# collector authenticates, parses, and then refuses with a 422 whose
|
|
119
|
+
# +error+ names +version+, as not a version-1 report. That refusal, and
|
|
120
|
+
# only that one, is the ready answer. Call it before an expensive run, so
|
|
121
|
+
# a stopped collector or a refused key fails before the first model call
|
|
122
|
+
# rather than after the last one.
|
|
123
|
+
#
|
|
124
|
+
# Raises Error for anything else: a rejection other than 422, with the
|
|
125
|
+
# status, detail and guidance +call+ would carry (401 for a refused key,
|
|
126
|
+
# 404 when nothing at the endpoint takes reports); a retryable delivery
|
|
127
|
+
# failure when the collector cannot be reached; a 422 that says anything
|
|
128
|
+
# else; or a collector that stores the empty object. The last two mean the
|
|
129
|
+
# endpoint is not a compatible collector.
|
|
130
|
+
def verify!
|
|
131
|
+
deliver("start the collector or check the endpoint, then verify again") do
|
|
132
|
+
response = post("{}")
|
|
133
|
+
if response.code == "422"
|
|
134
|
+
detail = collector_detail(response.body)
|
|
135
|
+
next true if detail&.match?(/\bversion\b/)
|
|
136
|
+
|
|
137
|
+
raise Error.new("Evaluation collector answered HTTP 422 without refusing the empty envelope as a version-1 report, " \
|
|
138
|
+
"so it is not a compatible collector; check the endpoint", status: 422, detail: detail)
|
|
139
|
+
end
|
|
140
|
+
raise rejection(response) unless %w[200 201].include?(response.code)
|
|
141
|
+
|
|
142
|
+
raise Error, "Evaluation collector stored an empty report, so it is not a compatible collector; check the endpoint"
|
|
143
|
+
end
|
|
144
|
+
end
|
|
145
|
+
|
|
146
|
+
private
|
|
147
|
+
|
|
148
|
+
# Posts +body+ to the endpoint with the bearer key and the delivery timeouts.
|
|
149
|
+
def post(body)
|
|
94
150
|
http = Net::HTTP.new(@uri.hostname, @uri.port)
|
|
95
151
|
http.use_ssl = @uri.scheme == "https"
|
|
96
152
|
http.open_timeout = @open_timeout
|
|
@@ -101,26 +157,25 @@ module ActiveAgent
|
|
|
101
157
|
request["Content-Type"] = "application/json"
|
|
102
158
|
request["Accept"] = "application/json"
|
|
103
159
|
request.body = body
|
|
104
|
-
|
|
105
|
-
|
|
160
|
+
http.request(request)
|
|
161
|
+
end
|
|
106
162
|
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
163
|
+
# Runs one exchange with the collector, turning every failure to reach it
|
|
164
|
+
# or to read its answer into a retryable Error that quotes nothing from
|
|
165
|
+
# the response and ends with +guidance+, what the caller should do next.
|
|
166
|
+
# An Error the block raises passes through unchanged.
|
|
167
|
+
def deliver(guidance)
|
|
168
|
+
yield
|
|
112
169
|
rescue JSON::ParserError
|
|
113
170
|
# The parser's message quotes the body.
|
|
114
|
-
raise Error.new("Evaluation collector returned invalid JSON;
|
|
171
|
+
raise Error.new("Evaluation collector returned invalid JSON; #{guidance}", retryable: true), cause: nil
|
|
115
172
|
rescue Net::HTTPBadResponse, Net::HTTPHeaderSyntaxError, Zlib::Error => e
|
|
116
173
|
# These messages can quote the response's status line, headers or body.
|
|
117
|
-
raise Error.new("Evaluation collector returned a malformed response (#{e.class});
|
|
174
|
+
raise Error.new("Evaluation collector returned a malformed response (#{e.class}); #{guidance}", retryable: true), cause: nil
|
|
118
175
|
rescue IOError, SocketError, SystemCallError, Timeout::Error, OpenSSL::SSL::SSLError => e
|
|
119
|
-
raise Error.new("Evaluation delivery failed (#{e.class});
|
|
176
|
+
raise Error.new("Evaluation delivery failed (#{e.class}); #{guidance}", retryable: true)
|
|
120
177
|
end
|
|
121
178
|
|
|
122
|
-
private
|
|
123
|
-
|
|
124
179
|
def report_hash(report)
|
|
125
180
|
hash = report.to_h if report.respond_to?(:to_h) && !report.nil? && !report.is_a?(Array)
|
|
126
181
|
raise ArgumentError, "report must be a Report or its saved JSON hash" unless hash.is_a?(Hash)
|
|
@@ -178,6 +178,9 @@ module ActiveAgent
|
|
|
178
178
|
DesignTokens.css(scope: ":root.theme-dark", tokens: DesignTokens::DARK, color_scheme: "dark"),
|
|
179
179
|
STYLES,
|
|
180
180
|
".mx { grid-template-columns: minmax(240px, 1.6fr) 150px repeat(#{@models.size}, minmax(170px, 1fr)); }",
|
|
181
|
+
# One rule per model: with that chip checked, hide every fix card
|
|
182
|
+
# attributed to other models (cards attributed to none stay).
|
|
183
|
+
*@models.each_index.map { |i| ".fix-section:has(input[value=\"m#{i}\"]:checked) .fix[data-models]:not([data-models~=\"m#{i}\"]) { display: none; }" },
|
|
181
184
|
".matrix .inner { min-width: #{390 + 185 * @models.size}px; }"
|
|
182
185
|
].join("\n")
|
|
183
186
|
end
|
|
@@ -233,12 +236,78 @@ module ActiveAgent
|
|
|
233
236
|
<<~PANEL
|
|
234
237
|
<div class="panel">
|
|
235
238
|
<div class="panel-head"><span class="micro">Models</span><span class="right">judged by #{h(judged_by)}</span></div>
|
|
239
|
+
#{html_comparison_table if comparing?}
|
|
236
240
|
#{blocks.join}
|
|
237
241
|
#{verdict_row}
|
|
238
242
|
</div>
|
|
239
243
|
PANEL
|
|
240
244
|
end
|
|
241
245
|
|
|
246
|
+
# The comparison read across: one row per model, best first (pass rate,
|
|
247
|
+
# then mean score) — passed, mean score, average latency, average
|
|
248
|
+
# tokens per scenario, cost, and the model's typical fault. The blocks
|
|
249
|
+
# under it carry the same figures per model with bars and every fault.
|
|
250
|
+
def html_comparison_table
|
|
251
|
+
rows = summary_by_model.sort_by do |label, stats|
|
|
252
|
+
total = stats["scenarios"].to_i
|
|
253
|
+
[ total.positive? ? -stats["passed"].to_f / total : 0.0, -(stats["avg_score"] || -1).to_f, @models.index(model_by_label(label)).to_i ]
|
|
254
|
+
end
|
|
255
|
+
|
|
256
|
+
<<~TABLE
|
|
257
|
+
<div class="compare"><table>
|
|
258
|
+
<thead><tr><th>Model</th><th class="num">Passed</th><th class="num">Mean score</th><th class="num">Avg latency</th><th class="num" title="Average input + output tokens per scenario">Avg tokens</th><th class="num" title="Cohort spend, and per scenario">Cost</th><th class="fault">Typical fault</th></tr></thead>
|
|
259
|
+
<tbody>#{rows.map { |label, stats| html_comparison_row(label, stats) }.join}</tbody>
|
|
260
|
+
</table></div>
|
|
261
|
+
TABLE
|
|
262
|
+
end
|
|
263
|
+
|
|
264
|
+
def html_comparison_row(label, stats)
|
|
265
|
+
short, provider = split_label(model_by_label(label))
|
|
266
|
+
total = stats["scenarios"].to_i
|
|
267
|
+
ratio = total.positive? ? stats["passed"].to_f / total : 0.0
|
|
268
|
+
pick = comparing? && winner == label ? %(<span class="pick" title="picked by the judge">★ pick</span>) : ""
|
|
269
|
+
per = ->(value) { value.nil? || total.zero? ? nil : value.to_f / total }
|
|
270
|
+
avg_tokens = per.call(stats["input_tokens"].to_i + stats["output_tokens"].to_i)
|
|
271
|
+
tokens_cell = avg_tokens ? h(fmt_k(avg_tokens.round)) : "—"
|
|
272
|
+
tokens_title = avg_tokens ? %( title="#{per.call(stats['input_tokens']).to_f.round} in · #{per.call(stats['output_tokens']).to_f.round} out per scenario") : ""
|
|
273
|
+
per_cost = per.call(stats["cost"])
|
|
274
|
+
cost_cell = stats["cost"].nil? ? "—" : h(fmt_cost(stats["cost"]))
|
|
275
|
+
cost_cell += "<span class=\"per\">#{h(fmt_cost(per_cost))}/scenario</span>" if per_cost
|
|
276
|
+
|
|
277
|
+
<<~ROW
|
|
278
|
+
<tr>
|
|
279
|
+
<td class="model-cell"><span class="name">#{h(short)}</span>#{pick}<span class="provider">#{h(provider)}</span></td>
|
|
280
|
+
<td class="num ratio tone-#{tone_for(ratio)}">#{total.positive? ? "#{stats['passed']}/#{total}" : '—'}</td>
|
|
281
|
+
<td class="num">#{h(fmt_mean_score(stats['avg_score']))}</td>
|
|
282
|
+
<td class="num">#{h(fmt_ms(stats['avg_duration_ms']))}</td>
|
|
283
|
+
<td class="num"#{tokens_title}>#{tokens_cell}</td>
|
|
284
|
+
<td class="num">#{cost_cell}</td>
|
|
285
|
+
<td class="fault">#{typical_fault_text(label, stats)}</td>
|
|
286
|
+
</tr>
|
|
287
|
+
ROW
|
|
288
|
+
end
|
|
289
|
+
|
|
290
|
+
# "missing content ×2 · refund_window: The answer is missing expected
|
|
291
|
+
# content: 30." — the model's most frequent fault, and the diagnosis of
|
|
292
|
+
# the first result that carries it; "no faults" for a clean cohort.
|
|
293
|
+
def typical_fault_text(label, stats)
|
|
294
|
+
tally = stats["faults"] || {}
|
|
295
|
+
return %(<span class="clean">no faults</span>) if tally.empty?
|
|
296
|
+
|
|
297
|
+
mine = @results.select { |result| result.label == label }
|
|
298
|
+
example_of = ->(fault) { mine.find { |result| result.fault == fault } }
|
|
299
|
+
# Most frequent first; between equals, a fault a result can explain,
|
|
300
|
+
# then a specific fault over the judge's catch-all, then the name.
|
|
301
|
+
fault, count = tally.min_by { |name, n| [ -n, example_of.call(name) ? 0 : 1, name == "low_quality" ? 1 : 0, name ] }
|
|
302
|
+
example = example_of.call(fault)
|
|
303
|
+
head = "#{fault_name(fault)} ×#{count}"
|
|
304
|
+
return h(head) unless example&.summary.present?
|
|
305
|
+
|
|
306
|
+
detail = "#{example.scenario.key}: #{example.summary}"
|
|
307
|
+
detail = "#{detail[0, 119]}…" if detail.length > 120
|
|
308
|
+
"#{h(head)} <span class=\"detail\">· #{h(detail)}</span>"
|
|
309
|
+
end
|
|
310
|
+
|
|
242
311
|
def html_model_block(label, stats)
|
|
243
312
|
short, provider = split_label(model_by_label(label))
|
|
244
313
|
total = stats["scenarios"]
|
|
@@ -277,13 +346,36 @@ module ActiveAgent
|
|
|
277
346
|
end
|
|
278
347
|
|
|
279
348
|
<<~FIXES
|
|
280
|
-
<section class="section" aria-label="Recommendations">
|
|
349
|
+
<section class="section fix-section" aria-label="Recommendations">
|
|
281
350
|
<div class="section-head"><span class="micro">What to fix</span><span class="meta">#{h(meta)}</span></div>
|
|
351
|
+
#{html_fix_filter(items) if comparing? && items.any?}
|
|
282
352
|
#{body}
|
|
283
353
|
</section>
|
|
284
354
|
FIXES
|
|
285
355
|
end
|
|
286
356
|
|
|
357
|
+
# A model filter for the fix cards — a fault one model keeps making is
|
|
358
|
+
# that model's to fix, so the list narrows to what was attributed to
|
|
359
|
+
# it. Radio chips and stylesheet rules alone (the page carries no
|
|
360
|
+
# script): each card names its models in data-models, and a checked
|
|
361
|
+
# model hides every card that does not name it. Cards attributed to no
|
|
362
|
+
# model (an older run) stay under every filter.
|
|
363
|
+
def html_fix_filter(items)
|
|
364
|
+
chips = [ %(<label class="chip pick-model"><input type="radio" name="fix-model" value="all" checked><span>all models #{items.size}</span></label>) ]
|
|
365
|
+
@models.each_with_index do |spec, index|
|
|
366
|
+
count = items.count { |item| Array(item["models"]).empty? || item["models"].include?(spec.label) }
|
|
367
|
+
chips << %(<label class="chip pick-model"><input type="radio" name="fix-model" value="m#{index}"><span>#{h(short_name(spec))} #{count}</span></label>)
|
|
368
|
+
end
|
|
369
|
+
%(<div class="fix-filter"><span class="micro sm">for</span>#{chips.join}</div>)
|
|
370
|
+
end
|
|
371
|
+
|
|
372
|
+
def fix_model_tokens(item)
|
|
373
|
+
labels = Array(item["models"])
|
|
374
|
+
return "" if labels.empty?
|
|
375
|
+
|
|
376
|
+
labels.filter_map { |label| (index = @models.index(model_by_label(label))) && "m#{index}" }.join(" ")
|
|
377
|
+
end
|
|
378
|
+
|
|
287
379
|
def html_fix_card(item)
|
|
288
380
|
tone = item["kind"] == "instruction" ? "info" : "error"
|
|
289
381
|
glyph = tone == "info" ? "[i]" : "[!]"
|
|
@@ -297,7 +389,8 @@ module ActiveAgent
|
|
|
297
389
|
parts << html_fix_server(item["server"]) if item["server"]
|
|
298
390
|
parts << %(<div class="note">#{h(item['note'])}</div>) if item["note"].present?
|
|
299
391
|
parts << html_fix_action(item["action"]) if item["action"]
|
|
300
|
-
|
|
392
|
+
models = fix_model_tokens(item)
|
|
393
|
+
%(<div class="fix"#{%( data-models="#{models}") if models.present?}>#{parts.join}</div>)
|
|
301
394
|
end
|
|
302
395
|
|
|
303
396
|
def html_fix_tools(item)
|
|
@@ -557,6 +650,24 @@ module ActiveAgent
|
|
|
557
650
|
.tok .in { color: var(--color-token-in); }
|
|
558
651
|
.tok .out { color: var(--color-token-out); }
|
|
559
652
|
.faults { display: flex; gap: 6px; flex-wrap: wrap; }
|
|
653
|
+
.compare { overflow-x: auto; border-top: 1px solid var(--color-border-light); }
|
|
654
|
+
.compare table { width: 100%; border-collapse: collapse; }
|
|
655
|
+
.compare th { padding: 8px 12px; text-align: left; vertical-align: bottom; white-space: nowrap; font-family: var(--font-mono); font-size: 10px; font-weight: 600; letter-spacing: 0.05em; text-transform: uppercase; color: var(--color-text-muted); background: var(--color-muted); }
|
|
656
|
+
.compare td { padding: 9px 12px; vertical-align: top; border-top: 1px solid var(--color-border-light); font-size: 13px; color: var(--color-text-cell); }
|
|
657
|
+
.compare th.num, .compare td.num { text-align: right; }
|
|
658
|
+
.compare td.num { font-family: var(--font-mono); font-size: 12px; white-space: nowrap; }
|
|
659
|
+
.compare td.ratio { font-weight: 600; }
|
|
660
|
+
.compare .per { display: block; font-weight: 400; color: var(--color-text-muted); }
|
|
661
|
+
.compare .model-cell { white-space: nowrap; }
|
|
662
|
+
.compare .model-cell .name { font-family: var(--font-mono); font-size: 12px; font-weight: 600; color: var(--color-text-primary); }
|
|
663
|
+
.compare .model-cell .provider { display: block; font-family: var(--font-mono); font-size: 11px; color: var(--color-text-muted); }
|
|
664
|
+
.compare .pick { margin-left: 6px; font-family: var(--font-mono); font-size: 10px; font-weight: 700; color: var(--color-warning-text); }
|
|
665
|
+
.compare th.fault, .compare td.fault { width: 34%; }
|
|
666
|
+
.compare td.fault .detail { color: var(--color-text-secondary); }
|
|
667
|
+
.fix-filter { display: flex; align-items: center; gap: 6px; flex-wrap: wrap; }
|
|
668
|
+
.pick-model { cursor: pointer; border: 1px solid var(--color-border); background: var(--color-card); }
|
|
669
|
+
.pick-model input { position: absolute; opacity: 0; width: 0; height: 0; }
|
|
670
|
+
.pick-model:has(input:checked) { border-color: var(--color-accent-ui); background: var(--color-accent-ui-tint); color: var(--color-accent-ui); }
|
|
560
671
|
.clean { font-family: var(--font-mono); font-size: 11px; color: var(--color-success-text); }
|
|
561
672
|
.verdict { padding: 10px 12px; border-top: 1px solid var(--color-border-light); font-size: 12px; line-height: 18px; color: var(--color-text-cell); }
|
|
562
673
|
.verdict .micro { margin-right: 8px; }
|
|
@@ -125,6 +125,10 @@ module ActiveAgent
|
|
|
125
125
|
expectations = (entry["expectations"] || entry["expect"] || {}).to_h.stringify_keys
|
|
126
126
|
%w[tools contains not_contains].each do |field|
|
|
127
127
|
expectations[field] = Array(entry[field]) if entry.key?(field)
|
|
128
|
+
# A nested expectation written as one value ({ contains: "30" })
|
|
129
|
+
# is a list of one: the persisted scenario and the dashboard's
|
|
130
|
+
# matrix read each field as an array.
|
|
131
|
+
expectations[field] = Array(expectations[field]) if expectations.key?(field)
|
|
128
132
|
end
|
|
129
133
|
|
|
130
134
|
scenario(
|
data/lib/active_agent/version.rb
CHANGED