activeagent 1.7.2 → 1.8.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 881ca4e27488bd719dda01c80e6241bf6d17082876039033c22f9265d1e559aa
4
- data.tar.gz: 62526259e8a48d5243a480cf3ac6f795c1672b20e8569f7804fcf13af4d25ec4
3
+ metadata.gz: 1c7be2b31907ff2779ea3818516d887afecc6c540c89670c33f7ec1be242c01a
4
+ data.tar.gz: 0d7327bb8f94d2573f6e87b9191afbc7758b2fdb7d42a2ea896fbd25dfd71b17
5
5
  SHA512:
6
- metadata.gz: 0d135cc35623442cdf38df2d588e323fec82f7e16c7d9c8fcf90d2add91bbb55476e07fe548c70ddb2b22cc709e684ab8ccc53dd337c0650800f9d6dd268f547
7
- data.tar.gz: 76a0b25ef5cf125e99399ff61cef131d2fe6ae6dade18173601f73f0ec7ecedde6bfc0738af99ac490444b1a067cd136dc2caac34cc1f1cd6c01996cbdddf841
6
+ metadata.gz: fd0f4bbab5439daabe85170c3fb57fe5c94260e0573fc35944e80cc8301bd35803fb5d89f2a6de9c6061c44752e67f125db9ad17d116058b444e92785c3819c6
7
+ data.tar.gz: 453ecb45269cc816fbea41378591e642a460a2a188e052745b87eed1fa9f7777281449a7adfe19ebef17e12e242009ad46ee6a773f0cc33660b161f6d4585197
data/CHANGELOG.md CHANGED
@@ -7,6 +7,345 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [1.8.1] - 2026-10-01
11
+
12
+ ### Added
13
+
14
+ - **Every evaluation cost is priced, and says how** (`actionagent`,
15
+ `activeagent`). A scenario result with no cost is priced down a chain —
16
+ the cost the publishing application reported, the estimate the engine
17
+ stored, the result's tokens × its model's rate, the tokens of the trace
18
+ it links to, or its text at four characters a token as a lower bound —
19
+ and a result that recorded no tokens costs `$0.00` (`cost_source:
20
+ no_usage`); only a result with no tokens, no trace and no text stays
21
+ unpriced. Results carry `cost` (the effective figure), `reported_cost`,
22
+ `cost_source`, `cost_rate` and `judge_usage`; a run's `usage` adds
23
+ `reported`, `estimated`, `cost_basis` and `total`; `scores._models` adds
24
+ per model `reported`, `estimated`, `judge_cost` and `judge_calls`; and
25
+ `GET /api/evaluations/:id/runs/:run_id` adds `costs` per scenario (judge
26
+ apart) and for the run (`ActionAgent::EvaluationRunCost`, cached per
27
+ finished run). The judge's spend is found the same way — the engine's
28
+ meter, the application's figures (`result.judge_usage`,
29
+ `report.judge_usage.run`, accepted by the report import) or the judge
30
+ traces priced on input and output tokens, never thinking tokens — within
31
+ the run's tenant. `ModelPricing` looks rates up under the provider the
32
+ model ran on, strips gateway and vendor prefixes and date suffixes, tries
33
+ dots and dashes both ways, prices `claude-sonnet-5` and the `gpt-5`
34
+ family by exact rows ahead of the family patterns, and reports where a
35
+ rate came from (`estimate_detailed`, `rate_detail`, `fingerprint`). The
36
+ engine's judge meter counts Anthropic's cached prompt tokens.
37
+ - **One display format for passes, scores and costs** (`activeagent`
38
+ `ActiveAgent::Evals::Format`). A fraction always carries its percent
39
+ (`14/16 · 88%`; `14/16 (88%)` in Markdown; `—` for nothing scored), every
40
+ 0..1 score reads as a whole percent (`93%`, `pass ≥ 70%`), and money
41
+ reads `$0.0243` when reported and `~$0.0243` when any part was estimated,
42
+ with one legend per surface. The HTML report gains a Cost tile, a Judge
43
+ column and per-model judge line, a trailing matrix Cost column with
44
+ group subtotals, a cost line per cell and in each result's details, a
45
+ judge chip with its calls and cost, a release chip, the pass mark and
46
+ the cost in its footer; the Markdown report gains a Judge column, a
47
+ matrix Cost column, the cost per answer and a total line. Diagnosis
48
+ text reads `Task completion scored 60% against a pass threshold of 70%`.
49
+ - **Costs from replay metadata** (`activeagent`). A `Replay`'s metadata may
50
+ carry `cost_source`, `cost_rate` and `judge_usage`; `Report#summary_by_model`
51
+ adds `reported`, `estimated`, `judge_cost` and `judge_calls`,
52
+ `Report#scenario_costs` gives each scenario's cost across models, and
53
+ `Report#judge_usage` sums every result's judge calls with the run-level
54
+ part passed as `Report.new(judge_usage:)`. `Report.new(release:)` (and
55
+ `Runner.new(release:)`) names the release the run scored, in
56
+ `to_h["release"]` and a header chip. `to_h` is unchanged when neither is
57
+ given.
58
+ - **An evaluation's standing against the agent as it is now**
59
+ (`actionagent`). Each evaluation reports its `headline_run_id` (its
60
+ newest complete run; a newer pending or failed run shows beside it), its
61
+ `standing` — `current`, `stale`, `unrecorded`, `archived` or `none`
62
+ (`ActionAgent::EvaluationStanding`) — its `archived_at` and `per_model`
63
+ passes, and every run its `agent_version` and `version_state`. Only
64
+ model-facing edits (instructions, action prompts, tools, MCP servers,
65
+ model config, response format) make a run stale. A published report's
66
+ `report.release` pins the run to that release, recorded as a version when
67
+ the dashboard has not seen the digest (`Agent#find_or_record_release!`,
68
+ which never moves a deploy's `release_digest` backwards); a report with
69
+ no release leaves the run unrecorded. `PATCH /api/evaluations/:id` with
70
+ `evaluation: { archived: true | false }` archives an evaluation or brings
71
+ it back; `GET /api/evaluations` leaves archived evaluations out before
72
+ its 50-row limit unless `?archived=1`, and returns `archived_count`. A
73
+ new run or a published report brings an archived evaluation back.
74
+
75
+ - Codex code sessions in checkout sandboxes. Connect an OpenAI API key under
76
+ Settings → Integrations and select Codex in the code-session panel. The local
77
+ backend runs `codex exec` with JSONL events, workspace-write sandboxing, stdin
78
+ prompts, per-sandbox configuration, cancellation, timeout and diff capture.
79
+ - Explicit code-runner capability checks for host backends. Existing adapters
80
+ continue to support Claude Code without implicitly receiving Codex credentials.
81
+
82
+ ### Changed
83
+
84
+ - **The dashboard reads passes, scores and costs one way** (`actionagent`
85
+ frontend). Every fraction carries its percent (`14/16 · 88%`), every 0..1
86
+ score reads as a whole percent, and every cost reads `$0.0243` when
87
+ reported or `~$0.0243` when any part was estimated, with the legend once
88
+ per surface and the tokens × rate working in the figure's tooltip. The
89
+ Evaluations page's tiles pool only the headline run of each current or
90
+ unrecorded evaluation, show a pass line per model, and say how many
91
+ evaluations were left out as stale or archived; a card carries its
92
+ standing, an archive/unarchive control and, for a newer run still
93
+ pending or failed, that run's badge beside the headline's; *Show
94
+ archived (n)* lists the archived ones. The suite panel gains the spend
95
+ strip between Runs and Models, the matrix a cost line per cell
96
+ (`~$0.0243 · judge ~$0.0015`) and a trailing Cost column with group
97
+ subtotals, the model comparison a Judge column, the runs list a version
98
+ chip and a same-version / new-version / release-not-recorded line under
99
+ each delta, and the spend strip reads the judge's cost from the traces
100
+ and says "rules only" only when there is no judge. The agent cards'
101
+ Eval tile is the pooled pass rate with its fraction in the title.
102
+ - **The agent card's Eval tile is the pooled pass rate** (`actionagent`
103
+ `AgentScorecard`) over the headline runs of the agent's current and
104
+ unrecorded evaluations — never a stale suite's or an archived one's —
105
+ rather than the mean criterion score of whichever run was last.
106
+ `eval_runs` and `eval_not_counted` say what was pooled and what was left
107
+ out.
108
+ - **The partial-cost notes are retired** (`actionagent`, `activeagent`).
109
+ The `*` marker and the "k of n priced" notes of 1.8.0's partial-cost fix
110
+ give way to the `~` mark and its legend: a cost that covers only some of
111
+ a model's replays is a lower bound and reads as an estimate. The verdict
112
+ rationale reads `Passed 2 of 2 scenarios (100%) with a mean score of 100%
113
+ at ~$0.0010 (estimated)`, and the judge ruling on a comparison is told
114
+ each model's agent cost alone — never the judge's own spend, which never
115
+ enters the ranking either. The HTML report's header shows a chip per
116
+ scalar metadata value only: an array or object (the judge's trace ids)
117
+ is no longer rendered as one.
118
+ - **What to fix comes after the scenario results** (`actionagent`,
119
+ `activeagent`). A run report now reads models, then the scenario results,
120
+ then What to fix. The scenario suite panel moves What to fix below the
121
+ scenario matrix. The standalone HTML report
122
+ (`ActiveAgent::Evals::ReportHtml`) moves its fix cards below the matrix and
123
+ the per-scenario details. `Report#to_markdown` moves its Recommendations
124
+ below its Answers. The sampling run detail already read in this order.
125
+ Each section's content is unchanged.
126
+
127
+ ### Fixed
128
+
129
+ - Inherited Codex settings and credentials are removed from sandbox process
130
+ environments. Each Codex run receives only its owner's selected connection.
131
+
132
+ - **An evaluation's cost when some interactions carried no cost estimate**
133
+ (`actionagent`, `activeagent`). A run's cost sums only the interactions that
134
+ were priced, but its per-interaction rate divided that partial sum by every
135
+ interaction, and nothing said part of the run was unpriced. The rate is now
136
+ over the priced interactions, and a run's `usage` counts them: `priced` and
137
+ `unpriced` beside `replays` (or `samples`). A run where nothing was priced
138
+ reports no `cost` or `per_interaction`, as before, with `priced: 0`. The
139
+ per-model summaries count them too: `Report#summary_by_model` adds `priced`
140
+ per model and a sampling run's `_cohorts` add `priced` per cohort. A
141
+ scenario run recorded before that has its `_models` counted from its
142
+ results when the API serves it (`EvaluationRun#model_summaries`); a
143
+ sampling cohort recorded before that reads as fully priced when it has a
144
+ cost. The dashboard shows a partial cost as an estimate — "estimated, 3 of
145
+ 5 replays priced" — on the run's spend strip and footer, the model
146
+ scorecards and the Evaluations page's cost-per-interaction tile, and marks
147
+ it `*` in the runs list, the model comparison table and the spend strip's
148
+ total, whose titles give the count. The HTML
149
+ and Markdown reports and the pass-rate verdict name the priced count beside
150
+ a partial cost, the judge ruling on a comparison is told it, and a
151
+ pass-rate verdict breaks a tie on cost per priced scenario rather than on
152
+ the partial sum, so an unpriced replay no longer makes a model look
153
+ cheaper.
154
+
155
+ Upgrade both gems together, then run `bin/rails generate action_agent:install
156
+ --skip` and `bin/rails db:migrate`. The migration adds runner identity to code
157
+ sessions; existing sessions remain Claude Code sessions.
158
+
159
+ ## [1.8.0] - 2026-09-29
160
+
161
+ Releases `activeagent` and `actionagent` 1.8.0 from one tag. A minor release.
162
+ Settings -> Integrations connects GitHub and Claude Code. A connected
163
+ repository's checkout boots as a sandbox, on a developer's machine with the new
164
+ `:local` backend, where Claude Code sessions run and an evaluation can run
165
+ against the checkout without the agent being edited. The dashboard's MCP
166
+ server gains evaluation and telemetry tools for a developer's own coding
167
+ harness. Ollama hosts can be tested and can be remote, with an optional Bearer
168
+ API key. Comparison runs lead with a per-model table and filter the fix list
169
+ by model. A publisher can check its collector before a run.
170
+
171
+ Upgrading: run `bin/rails generate action_agent:install --skip` and
172
+ `bin/rails db:migrate`. The generator adds what an install lacks:
173
+ `provider_keys.api_key` and the `github_connections` and `code_sessions`
174
+ tables. An app on the RubyLLM provider needs ruby_llm 1.16 or later; 2.x
175
+ works too.
176
+
177
+ The Claude Code connection stores Anthropic API keys only. Anthropic does not
178
+ let third-party products collect, store or route requests through Claude.ai
179
+ subscription credentials
180
+ ([Claude Code legal and compliance](https://code.claude.com/docs/en/legal-and-compliance.md)),
181
+ so a `claude setup-token` token (`sk-ant-oat…`) is refused. An install that
182
+ stored one while running a pre-release build never hands it to a session: its
183
+ owner sees Claude Code as needing an API key (`needs_replacing: true` in
184
+ `GET /api/provider_keys`) until they paste one.
185
+ `bin/rails action_agent:claude_code:purge_subscription_tokens` deletes the
186
+ stored tokens and prints how many it removed. A developer who wants sessions on
187
+ their own Claude login sets `config.claude_code_auth = :local_login` with the
188
+ `:local` backend and runs `claude /login` on that machine instead.
189
+
190
+ ### Added
191
+
192
+ - **Evaluation and telemetry tools on the MCP facade** (`actionagent`). The
193
+ dashboard's MCP server now offers `evaluations_list`, `evaluations_get`,
194
+ `evaluations_run`, `evaluation_runs_get`, `evaluation_runs_compare`,
195
+ `traces_search` and `traces_get`, so a developer's own coding harness can
196
+ run an agent's evaluations, read the fix items and failing traces, and
197
+ iterate on the agent in its own checkout without the dashboard holding a
198
+ model login. The tools read under the API key's owner exactly as the JSON
199
+ API reads under the signed-in owner, run through the same execution,
200
+ quota and sandbox checks as `POST /api/evaluations/:id/run`, bound their
201
+ output, and mask the owner's credentials. Their names cannot collide with
202
+ schema tools or agent tools. `ActionAgent.mcp_dashboard_tools = false`
203
+ turns them off.
204
+
205
+ - **Local checkout sandboxes and Claude Code sessions** (`actionagent`, #489).
206
+ A new `:local` sandbox backend (`config.sandbox_service = :local`) makes
207
+ **Start sandbox** work on a developer's machine without containers. It
208
+ clones the repository under `tmp/action_agent/sandboxes`, runs the setup
209
+ the checkout's optional `.activeagents/sandbox.yml` names (`env`, `setup`,
210
+ `manifest`, `start`), and boots the app on `127.0.0.1`. It then registers
211
+ the app's MCP facade as the `sandbox:<session_id>` server. Every process
212
+ starts from a sanitized copy of the dashboard's environment: without its
213
+ database and Redis URLs, Rails keys and environment, Bundler and Ruby
214
+ settings, git repository and config variables, `SSH_AUTH_SOCK`,
215
+ model-provider and Claude Code settings, variables named like a secret, or
216
+ URLs carrying credentials. The GitHub token reaches
217
+ only the fetch, and the Claude Code API key reaches only Claude Code.
218
+ `:local` runs the owner's code with the dashboard's privileges, so it is off
219
+ outside development and test unless
220
+ `ActionAgent.local_sandboxes_enabled = true`.
221
+ - `bin/rails action_agent:sandbox:manifest` writes the booted app's
222
+ `{mcp_path, mcp_token}` for the backend.
223
+ - `bin/rails action_agent:sandbox:reap` expires overdue sandboxes and stops
224
+ them.
225
+ - A ready sandbox runs headless Claude Code sessions (`--permission-mode
226
+ acceptEdits` by default). You start them from Settings -> Integrations, or
227
+ through `/api/sandboxes/:session_id/code_sessions`. Their events stream
228
+ into the dashboard, and the checkout's diff follows.
229
+ - New options: `local_sandboxes_enabled`, `local_sandbox_root`,
230
+ `local_sandbox_boot_timeout`, `claude_code_command`,
231
+ `claude_code_permission_mode`, `claude_code_max_turns`,
232
+ `claude_code_timeout` and `claude_code_auth`.
233
+ - `claude_code_auth = :local_login` runs sessions on the machine's own
234
+ Claude Code login (`claude /login`), with no stored key: the backend
235
+ passes no credential and no `CLAUDE_CONFIG_DIR`, so `claude` uses the
236
+ dashboard user's own `~/.claude` or keychain, which the dashboard never
237
+ reads. `LocalSandboxBackend.claude_login_status` asks
238
+ `claude auth status --json` (cached for a minute) and keeps only
239
+ `loggedIn` and the login method. `GET /api/sandboxes` reports
240
+ `claude_code_auth` and, in this mode, `claude_code_login`
241
+ (`{ logged_in, auth_method }`), and the assistant's
242
+ `connections.claude_code` the same as `auth` and `login`. Other backends
243
+ refuse sessions in this mode (`code_sessions_supported: false`, and a
244
+ `422` naming the reason).
245
+ - `app_runtime` sandboxes now provision in the background and last 2 hours.
246
+ - Run `rails g action_agent:install` to add the
247
+ `create_active_agent_code_sessions` migration.
248
+ - This repository's own `.activeagents/sandbox.yml` boots `test/dummy`.
249
+ - Every `:local` sandbox boots on databases of its own, so a checkout of
250
+ the dashboard's own app no longer migrates the developer's development
251
+ database. The backend reads the adapter from the checkout's
252
+ `config/database.yml` without running its ERB, and sets `DATABASE_URL`
253
+ and `<NAME>_DATABASE_URL` (`QUEUE_DATABASE_URL`, `CACHE_DATABASE_URL`):
254
+ SQLite files in the workspace, or `<database>_sandbox_<id>` on
255
+ PostgreSQL and MySQL, which terminate drops with the checkout's
256
+ Rails database tasks restricted to the names and URLs recorded at boot.
257
+ Setting a variable in `sandbox.yml`'s `env` overrides it and excludes
258
+ that database from cleanup. Replica mappings follow their own writer;
259
+ ambiguous mappings require an explicit URL instead of guessing.
260
+ - The Claude Code panel has a **Model** select: Claude Code's own
261
+ default, the `sonnet`, `opus` and `haiku` aliases, or any model id under
262
+ *Other…*. It remembers the last choice per browser, and each session
263
+ shows the model it ran on.
264
+ - A run can use a checkout sandbox without the agent being edited:
265
+ `sandbox_id` on `POST /api/evaluations/:id/run` (and on the runner's
266
+ `/api/agents/:id/execute` and `/test`) gives that run's tool dispatcher
267
+ the sandbox's `sandbox:<session_id>` runtime, as if the agent listed it.
268
+ The sandbox must be the caller's, a ready `app_runtime` sandbox, and the
269
+ agent owner's; anything else is a `422`. The run records which sandbox
270
+ it used (`run.sandbox`), and a scenario suite's **Run against sandbox**
271
+ select, its Runs list and the run report show it. The selected sandbox
272
+ takes precedence for matching tool names, with one schema per name.
273
+ Queued agent runs fail if their selected sandbox stops, and failed
274
+ sandbox discovery never silently falls back to the original tools.
275
+ - **Claude Code connection** (`actionagent`, #478). Settings -> Integrations
276
+ stores an Anthropic API key (`sk-ant-api…`, from the Claude Console) as the
277
+ `claude_code` provider key. It is encrypted, write-only, and not an agent
278
+ provider. `SandboxSession#runtime_environment` hands it to an
279
+ `app_runtime` backend as `ANTHROPIC_API_KEY`, so the checkout can run
280
+ Claude Code sessions. Claude subscription tokens (`claude setup-token`)
281
+ are refused, as Anthropic's terms require (see the upgrading note above).
282
+ `/api/provider_keys` rows now carry `kind` (`key`, `host` or
283
+ `connection`) and `needs_replacing`.
284
+ - **GitHub connections and checkout sandboxes** (`actionagent`, #477).
285
+ Settings -> Integrations connects GitHub over OAuth
286
+ (`ActionAgent.github_client_id` / `github_client_secret`, or
287
+ `GITHUB_CLIENT_ID` / `GITHUB_CLIENT_SECRET`), stores the token encrypted,
288
+ and lets the owner choose which repositories the workspace may use. Only
289
+ repositories GitHub lists for the token can be selected. A new
290
+ `app_runtime` sandbox type checks out one of them: the backend receives
291
+ `sandbox_session.checkout_spec` (repository, ref, clone URL, token),
292
+ boots the app, and returns `mcp_url` / `mcp_token` from `create_sandbox`.
293
+ The session is then an MCP server keyed `sandbox:<session_id>`. An agent
294
+ that lists that key in `mcp_servers` runs, and is evaluated, with the
295
+ checkout's own tools. Lookups are scoped to the agent's owner. Run
296
+ `rails g action_agent:install` to add the
297
+ `create_active_agent_github_connections` migration.
298
+
299
+ - **Ollama hosts are testable and can be remote** (`actionagent`). Settings ->
300
+ Provider API Keys gains a **Test connection** for Ollama that reports
301
+ whether the server is reachable, the round-trip time and the models it
302
+ serves, before or after saving (`POST <mount>/api/provider_keys/test`,
303
+ read-only). The host is accepted as a bare server address
304
+ (`http://localhost:11434`; the OpenAI-compatible `/v1` path is added) and
305
+ an optional **API key** is stored beside it and sent as a Bearer token,
306
+ for a server behind an authenticating proxy or Ollama Cloud. The agent
307
+ builder's live Ollama model list uses the same probe and key. When no host
308
+ is configured the card shows the host app's `config/active_agent.yml`
309
+ default. The install generator emits a guarded `add_provider_key_api_key`
310
+ migration for existing installs; re-run
311
+ `bin/rails generate action_agent:install --skip` and `bin/rails db:migrate`.
312
+ - **A model comparison table on comparison runs** (`actionagent`,
313
+ `activeagent`). A run over several models now leads its Models section
314
+ with one row per model, best first: passed, mean score, average latency,
315
+ average tokens per scenario, cost (and per scenario), and the model's
316
+ typical fault — its most frequent one with the diagnosis of a result that
317
+ carries it. The scenario suite panel, the sampling run detail and the
318
+ standalone HTML report (`ActiveAgent::Evals::ReportHtml`) all render it.
319
+ - **What to fix, filtered by model** (`actionagent`, `activeagent`). On a
320
+ comparison run the fix list takes a model chip, narrowing to the items
321
+ attributed to that model and counting what that model alone produced,
322
+ since one model may need more instruction than another. The standalone
323
+ report filters through radio chips and stylesheet rules — it still ships
324
+ no script.
325
+
326
+ - **A publisher can check its collector before a run** (`activeagent`).
327
+ `ActiveAgent::Evals::Publisher#verify!` asks the collector whether it is up
328
+ and accepts the key before a run is paid for. It posts an empty JSON object,
329
+ which a compatible collector refuses with a 422 naming `version`, without
330
+ storing anything; anything else raises `Publisher::Error` with a delivery's
331
+ status, detail and guidance. `Publisher#endpoint` returns the collector URL.
332
+
333
+ ### Changed
334
+
335
+ - **Collector rejections say what the status means** (`activeagent`). A
336
+ `Publisher::Error` for a 401, 403, 404, 415 or 501 rejection names a refused
337
+ key, an account an operator must act on, an endpoint that is not a
338
+ collector, a rewritten `Content-Type`, or an install with no evaluation
339
+ store, in place of the generic guidance.
340
+
341
+ ### Fixed
342
+
343
+ - **A nested scenario expectation written as one value** (`activeagent`).
344
+ `ScenarioParser` now stores `{ expectations: { contains: "30" } }` as a
345
+ list of one, the shape the persisted scenario and the dashboard's matrix
346
+ read; an object-list import with a lone value used to break the suite
347
+ panel. The matrix also tolerates scenarios persisted before this.
348
+
10
349
  ## [1.7.2] - 2026-09-29
11
350
 
12
351
  Releases `activeagent` and `actionagent` 1.7.2 from one tag. A patch on 1.7.1:
@@ -285,11 +285,11 @@ module ActiveAgent
285
285
 
286
286
  weakest = (failed_grade ? grades : @scores.compact).min_by { |_, value| value }
287
287
  summary = if failed_grade
288
- "#{graded_label(grades)} scored #{grade.round(2)} against a pass threshold of #{@threshold}"
288
+ "#{graded_label(grades)} scored #{Format.score(grade)} against a pass threshold of #{Format.percent(@threshold)}"
289
289
  else
290
- "Scored #{@score.round(2)} against a pass threshold of #{@threshold}"
290
+ "Scored #{Format.score(@score)} against a pass threshold of #{Format.percent(@threshold)}"
291
291
  end
292
- summary += ", weakest on #{weakest.first} (#{weakest.last.round(2)})" if weakest
292
+ summary += ", weakest on #{weakest.first} (#{Format.score(weakest.last)})" if weakest
293
293
  recommendation =
294
294
  if weakest
295
295
  "Read the answer against the #{weakest.first.to_s.humanize.downcase} criterion and adjust the " \
@@ -0,0 +1,115 @@
1
+ # frozen_string_literal: true
2
+
3
+ module ActiveAgent
4
+ module Evals
5
+ # How every evaluation surface writes a pass count, a score and a cost, so
6
+ # the HTML report, the Markdown report, a diagnosis and the dashboard read
7
+ # alike:
8
+ #
9
+ # - a pass/fail fraction carries its percentage: "14/16 · 88%" (in
10
+ # Markdown "14/16 (88%)"); nothing scored (0/0) reads "—"
11
+ # - a 0..1 score — a mean score, a criterion score, task completion, a
12
+ # judge's confidence, the pass threshold — reads as a whole percent:
13
+ # "93%", "pass ≥ 70%"
14
+ # - money has four decimals, six below $0.001, "$0.00" for an explicit
15
+ # zero, and a leading "~" when any part of it was estimated from
16
+ # tokens × model rates rather than reported; LEGEND explains the mark
17
+ #
18
+ # Percentages round half up to a whole number. Stored values stay 0..1
19
+ # and USD; only their display changes here. The dashboard's
20
+ # actionagent/frontend/utils/evalFormat.mjs is the JavaScript copy of
21
+ # these rules, so a change here needs the same change there.
22
+ module Format
23
+ EMPTY = "—"
24
+ # Shown once per surface where a "~" figure is visible.
25
+ LEGEND = "~ estimated from tokens × model rates"
26
+ # Rate sources that are a guess at the model's price rather than a
27
+ # catalog entry for it (see the dashboard's ModelPricing).
28
+ FALLBACK_RATE_SOURCES = %w[pattern default].freeze
29
+
30
+ module_function
31
+
32
+ # 0.875 → "88%". "—" for a missing value.
33
+ def percent(fraction)
34
+ return EMPTY unless finite?(fraction)
35
+
36
+ "#{whole(fraction.to_f * 100)}%"
37
+ end
38
+
39
+ # 14 of 16 → "14/16 · 88%", or "14/16 (88%)" with `style: :markdown`,
40
+ # where " · " would read as a table cell separator next to a pipe.
41
+ # "—" when nothing was scored.
42
+ def passes(passed, total, style: :text)
43
+ count = total.to_i
44
+ return EMPTY unless count.positive?
45
+
46
+ done = passed.to_i
47
+ pct = "#{whole(done * 100.0 / count)}%"
48
+ style == :markdown ? "#{done}/#{count} (#{pct})" : "#{done}/#{count} · #{pct}"
49
+ end
50
+
51
+ # A 0..1 score as a percent: 0.93 → "93%".
52
+ def score(value)
53
+ percent(value)
54
+ end
55
+
56
+ # 0.7 → "pass ≥ 70%".
57
+ def threshold(value)
58
+ "pass ≥ #{percent(value)}"
59
+ end
60
+
61
+ # 0.0243 → "$0.0243", 0.000697 → "$0.000697", 0 → "$0.00"; with
62
+ # `estimated`, "~$0.0243". "—" for a missing value.
63
+ def money(value, estimated: false)
64
+ return EMPTY unless finite?(value)
65
+
66
+ amount = value.to_f
67
+ digits = if amount.zero? then 2
68
+ elsif amount.abs < 0.001 then 6
69
+ else 4
70
+ end
71
+ "#{'~' if estimated}$#{format("%.#{digits}f", amount)}"
72
+ end
73
+
74
+ # A $/M token rate with at least two decimals and no trailing noise:
75
+ # 5 → "$5.00/M", 0.075 → "$0.075/M".
76
+ def per_million(rate)
77
+ whole, fraction = format("%.4f", rate.to_f).split(".")
78
+ fraction = fraction.sub(/0+\z/, "")
79
+ "$#{whole}.#{fraction.ljust(2, '0')}/M"
80
+ end
81
+
82
+ # The tooltip of an estimated figure: how it was worked out, e.g.
83
+ # "estimated: 2,328 in × $5.00/M + 423 out × $30.00/M · catalog rate".
84
+ # `rate` is `{ "input", "output", "source" }` in $ per million tokens;
85
+ # a rate from the name-pattern table or the default appends
86
+ # "(fallback rate)". Without a rate it names the method only.
87
+ def cost_title(input_tokens: nil, output_tokens: nil, rate: nil)
88
+ rate = rate.to_h.transform_keys(&:to_s) if rate.respond_to?(:to_h)
89
+ return "estimated from tokens × model rates" unless rate.is_a?(Hash) && finite?(rate["input"]) && finite?(rate["output"])
90
+
91
+ source = rate["source"].presence || "catalog"
92
+ title = "estimated: #{delimited(input_tokens)} in × #{per_million(rate['input'])} + " \
93
+ "#{delimited(output_tokens)} out × #{per_million(rate['output'])} · #{source} rate"
94
+ title += " (fallback rate)" if FALLBACK_RATE_SOURCES.include?(source.to_s)
95
+ title
96
+ end
97
+
98
+ # Half up, on the decimal value: 14.5 → 15, even when the float arrived
99
+ # as 14.499999999999998.
100
+ def whole(value)
101
+ value.to_f.round(9).round
102
+ end
103
+
104
+ def delimited(count)
105
+ count.to_i.to_s.gsub(/(\d)(?=(\d{3})+\z)/, '\1,')
106
+ end
107
+
108
+ def finite?(value)
109
+ return false if value.nil? || value == ""
110
+
111
+ Float(value, exception: false)&.finite? || false
112
+ end
113
+ end
114
+ end
115
+ end
@@ -158,6 +158,8 @@ module ActiveAgent
158
158
  def verdict(summaries, instructions: nil)
159
159
  lines = summaries.map do |label, stats|
160
160
  faults = (stats["faults"] || {}).map { |fault, count| "#{fault}×#{count}" }.join(", ")
161
+ # The agent's cost only: what the judge itself spent scoring a
162
+ # cohort says nothing about the model under comparison.
161
163
  "#{label}: pass rate #{stats['pass_rate']}%, mean score #{stats['avg_score'] || 'n/a'}, " \
162
164
  "avg latency #{stats['avg_duration_ms'] || 'n/a'}ms, cost $#{stats['cost'] || 'n/a'}" \
163
165
  "#{", faults: #{faults}" if faults.present?}"
@@ -11,7 +11,8 @@ module ActiveAgent
11
11
  # Publishes a completed report without replaying the agent. The caller must
12
12
  # retain run_id when retrying: compatible collectors treat that identity as
13
13
  # immutable within the authenticated account. Delivery is blocking and does
14
- # not follow redirects with the account's bearer credential.
14
+ # not follow redirects with the account's bearer credential. +verify!+ asks
15
+ # the collector whether it is up and accepts the key before a run is paid for.
15
16
  #
16
17
  # Every failure to deliver raises Error. Invalid arguments raise
17
18
  # ArgumentError before anything is sent.
@@ -20,6 +21,9 @@ module ActiveAgent
20
21
  MAX_BYTES = 2 * 1024 * 1024
21
22
  DETAIL_LIMIT = 200
22
23
 
24
+ # @return [String] the collector URL reports go to
25
+ attr_reader :endpoint
26
+
23
27
  # Raised for every failed delivery. Only a network failure keeps the
24
28
  # underlying error as its +cause+, so neither the response nor the
25
29
  # report reaches a log through the exception chain.
@@ -48,10 +52,15 @@ module ActiveAgent
48
52
  # Whether each rejection status is retryable, and what the caller should
49
53
  # do about it. Other statuses fall back to the rules in +rejection+.
50
54
  REJECTIONS = {
55
+ 401 => [ false, "the collector refused the API key; check the key against the collector's account" ],
56
+ 403 => [ false, "the account may not store this report until an operator acts, for example on a cap on observed agents, evaluations or scenarios; resolve that before retrying" ],
57
+ 404 => [ false, "nothing at the endpoint takes evaluation reports; check that it is a collector's /v1/evaluations or <mount>/api/evaluation_reports URL" ],
51
58
  409 => [ false, "the collector already holds a different report under this run_id; never retry this report with the same run_id" ],
52
59
  413 => [ false, "the report exceeds the collector's size limit; publish a smaller selection" ],
60
+ 415 => [ false, "the collector did not receive application/json; check anything between the publisher and the collector that rewrites the Content-Type" ],
53
61
  422 => [ false, "correct what the collector refused before retrying" ],
54
- 429 => [ true, "the account is over its quota or rate limit; retain the report and run_id and retry later" ]
62
+ 429 => [ true, "the account is over its quota or rate limit; retain the report and run_id and retry later" ],
63
+ 501 => [ false, "the collector has no evaluation store; migrate the install, or publish to one generated with evaluation tables" ]
55
64
  }.freeze
56
65
 
57
66
  # The key is sent as a bearer token and filtered from the collector's
@@ -65,6 +74,7 @@ module ActiveAgent
65
74
  unless @uri.scheme == "https" || %w[localhost 127.0.0.1 ::1].include?(@uri.hostname)
66
75
  raise ArgumentError, "Evaluation endpoint requires HTTPS except on loopback hosts"
67
76
  end
77
+ @endpoint = @uri.to_s
68
78
 
69
79
  @api_key = api_key.to_s.strip
70
80
  raise ArgumentError, "Evaluation API key is required" if @api_key.empty?
@@ -91,6 +101,52 @@ module ActiveAgent
91
101
  body = encode(identities.merge("version" => 1, "report" => report_hash(report)))
92
102
  raise Error, "Evaluation report exceeds the 2 MiB delivery limit; publish a smaller selection" if body.bytesize > MAX_BYTES
93
103
 
104
+ deliver("retain the report and run_id for retry") do
105
+ response = post(body)
106
+ raise rejection(response) unless %w[200 201].include?(response.code)
107
+
108
+ receipt = JSON.parse(response.body.to_s)
109
+ unless receipt.is_a?(Hash) && receipt["run_id"] == run_id && receipt["status"] == "complete" && receipt["id"] && receipt["evaluation_id"]
110
+ raise Error.new("Evaluation collector returned an invalid completion receipt; retain the report and run_id for retry", retryable: true)
111
+ end
112
+ receipt
113
+ end
114
+ end
115
+
116
+ # Returns true when the collector is up and accepts the API key, without
117
+ # storing anything. It posts an empty JSON object, which a compatible
118
+ # collector authenticates, parses, and then refuses with a 422 whose
119
+ # +error+ names +version+, as not a version-1 report. That refusal, and
120
+ # only that one, is the ready answer. Call it before an expensive run, so
121
+ # a stopped collector or a refused key fails before the first model call
122
+ # rather than after the last one.
123
+ #
124
+ # Raises Error for anything else: a rejection other than 422, with the
125
+ # status, detail and guidance +call+ would carry (401 for a refused key,
126
+ # 404 when nothing at the endpoint takes reports); a retryable delivery
127
+ # failure when the collector cannot be reached; a 422 that says anything
128
+ # else; or a collector that stores the empty object. The last two mean the
129
+ # endpoint is not a compatible collector.
130
+ def verify!
131
+ deliver("start the collector or check the endpoint, then verify again") do
132
+ response = post("{}")
133
+ if response.code == "422"
134
+ detail = collector_detail(response.body)
135
+ next true if detail&.match?(/\bversion\b/)
136
+
137
+ raise Error.new("Evaluation collector answered HTTP 422 without refusing the empty envelope as a version-1 report, " \
138
+ "so it is not a compatible collector; check the endpoint", status: 422, detail: detail)
139
+ end
140
+ raise rejection(response) unless %w[200 201].include?(response.code)
141
+
142
+ raise Error, "Evaluation collector stored an empty report, so it is not a compatible collector; check the endpoint"
143
+ end
144
+ end
145
+
146
+ private
147
+
148
+ # Posts +body+ to the endpoint with the bearer key and the delivery timeouts.
149
+ def post(body)
94
150
  http = Net::HTTP.new(@uri.hostname, @uri.port)
95
151
  http.use_ssl = @uri.scheme == "https"
96
152
  http.open_timeout = @open_timeout
@@ -101,26 +157,25 @@ module ActiveAgent
101
157
  request["Content-Type"] = "application/json"
102
158
  request["Accept"] = "application/json"
103
159
  request.body = body
104
- response = http.request(request)
105
- raise rejection(response) unless %w[200 201].include?(response.code)
160
+ http.request(request)
161
+ end
106
162
 
107
- receipt = JSON.parse(response.body.to_s)
108
- unless receipt.is_a?(Hash) && receipt["run_id"] == run_id && receipt["status"] == "complete" && receipt["id"] && receipt["evaluation_id"]
109
- raise Error.new("Evaluation collector returned an invalid completion receipt; retain the report and run_id for retry", retryable: true)
110
- end
111
- receipt
163
+ # Runs one exchange with the collector, turning every failure to reach it
164
+ # or to read its answer into a retryable Error that quotes nothing from
165
+ # the response and ends with +guidance+, what the caller should do next.
166
+ # An Error the block raises passes through unchanged.
167
+ def deliver(guidance)
168
+ yield
112
169
  rescue JSON::ParserError
113
170
  # The parser's message quotes the body.
114
- raise Error.new("Evaluation collector returned invalid JSON; retain the report and run_id for retry", retryable: true), cause: nil
171
+ raise Error.new("Evaluation collector returned invalid JSON; #{guidance}", retryable: true), cause: nil
115
172
  rescue Net::HTTPBadResponse, Net::HTTPHeaderSyntaxError, Zlib::Error => e
116
173
  # These messages can quote the response's status line, headers or body.
117
- raise Error.new("Evaluation collector returned a malformed response (#{e.class}); retain the report and run_id for retry", retryable: true), cause: nil
174
+ raise Error.new("Evaluation collector returned a malformed response (#{e.class}); #{guidance}", retryable: true), cause: nil
118
175
  rescue IOError, SocketError, SystemCallError, Timeout::Error, OpenSSL::SSL::SSLError => e
119
- raise Error.new("Evaluation delivery failed (#{e.class}); retain the report and run_id for retry", retryable: true)
176
+ raise Error.new("Evaluation delivery failed (#{e.class}); #{guidance}", retryable: true)
120
177
  end
121
178
 
122
- private
123
-
124
179
  def report_hash(report)
125
180
  hash = report.to_h if report.respond_to?(:to_h) && !report.nil? && !report.is_a?(Array)
126
181
  raise ArgumentError, "report must be a Report or its saved JSON hash" unless hash.is_a?(Hash)