activeagent 1.7.2 → 1.8.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +339 -0
- data/lib/active_agent/evals/diagnosis.rb +3 -3
- data/lib/active_agent/evals/format.rb +115 -0
- data/lib/active_agent/evals/judge.rb +2 -0
- data/lib/active_agent/evals/publisher.rb +69 -14
- data/lib/active_agent/evals/report.rb +232 -31
- data/lib/active_agent/evals/report_html.rb +393 -115
- data/lib/active_agent/evals/result.rb +40 -0
- data/lib/active_agent/evals/runner.rb +6 -2
- data/lib/active_agent/evals/scenario_parser.rb +4 -0
- data/lib/active_agent/evals.rb +1 -0
- data/lib/active_agent/version.rb +1 -1
- metadata +3 -2
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 1c7be2b31907ff2779ea3818516d887afecc6c540c89670c33f7ec1be242c01a
|
|
4
|
+
data.tar.gz: 0d7327bb8f94d2573f6e87b9191afbc7758b2fdb7d42a2ea896fbd25dfd71b17
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: fd0f4bbab5439daabe85170c3fb57fe5c94260e0573fc35944e80cc8301bd35803fb5d89f2a6de9c6061c44752e67f125db9ad17d116058b444e92785c3819c6
|
|
7
|
+
data.tar.gz: 453ecb45269cc816fbea41378591e642a460a2a188e052745b87eed1fa9f7777281449a7adfe19ebef17e12e242009ad46ee6a773f0cc33660b161f6d4585197
|
data/CHANGELOG.md
CHANGED
|
@@ -7,6 +7,345 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
7
7
|
|
|
8
8
|
## [Unreleased]
|
|
9
9
|
|
|
10
|
+
## [1.8.1] - 2026-10-01
|
|
11
|
+
|
|
12
|
+
### Added
|
|
13
|
+
|
|
14
|
+
- **Every evaluation cost is priced, and says how** (`actionagent`,
|
|
15
|
+
`activeagent`). A scenario result with no cost is priced down a chain —
|
|
16
|
+
the cost the publishing application reported, the estimate the engine
|
|
17
|
+
stored, the result's tokens × its model's rate, the tokens of the trace
|
|
18
|
+
it links to, or its text at four characters a token as a lower bound —
|
|
19
|
+
and a result that recorded no tokens costs `$0.00` (`cost_source:
|
|
20
|
+
no_usage`); only a result with no tokens, no trace and no text stays
|
|
21
|
+
unpriced. Results carry `cost` (the effective figure), `reported_cost`,
|
|
22
|
+
`cost_source`, `cost_rate` and `judge_usage`; a run's `usage` adds
|
|
23
|
+
`reported`, `estimated`, `cost_basis` and `total`; `scores._models` adds
|
|
24
|
+
per model `reported`, `estimated`, `judge_cost` and `judge_calls`; and
|
|
25
|
+
`GET /api/evaluations/:id/runs/:run_id` adds `costs` per scenario (judge
|
|
26
|
+
apart) and for the run (`ActionAgent::EvaluationRunCost`, cached per
|
|
27
|
+
finished run). The judge's spend is found the same way — the engine's
|
|
28
|
+
meter, the application's figures (`result.judge_usage`,
|
|
29
|
+
`report.judge_usage.run`, accepted by the report import) or the judge
|
|
30
|
+
traces priced on input and output tokens, never thinking tokens — within
|
|
31
|
+
the run's tenant. `ModelPricing` looks rates up under the provider the
|
|
32
|
+
model ran on, strips gateway and vendor prefixes and date suffixes, tries
|
|
33
|
+
dots and dashes both ways, prices `claude-sonnet-5` and the `gpt-5`
|
|
34
|
+
family by exact rows ahead of the family patterns, and reports where a
|
|
35
|
+
rate came from (`estimate_detailed`, `rate_detail`, `fingerprint`). The
|
|
36
|
+
engine's judge meter counts Anthropic's cached prompt tokens.
|
|
37
|
+
- **One display format for passes, scores and costs** (`activeagent`
|
|
38
|
+
`ActiveAgent::Evals::Format`). A fraction always carries its percent
|
|
39
|
+
(`14/16 · 88%`; `14/16 (88%)` in Markdown; `—` for nothing scored), every
|
|
40
|
+
0..1 score reads as a whole percent (`93%`, `pass ≥ 70%`), and money
|
|
41
|
+
reads `$0.0243` when reported and `~$0.0243` when any part was estimated,
|
|
42
|
+
with one legend per surface. The HTML report gains a Cost tile, a Judge
|
|
43
|
+
column and per-model judge line, a trailing matrix Cost column with
|
|
44
|
+
group subtotals, a cost line per cell and in each result's details, a
|
|
45
|
+
judge chip with its calls and cost, a release chip, the pass mark and
|
|
46
|
+
the cost in its footer; the Markdown report gains a Judge column, a
|
|
47
|
+
matrix Cost column, the cost per answer and a total line. Diagnosis
|
|
48
|
+
text reads `Task completion scored 60% against a pass threshold of 70%`.
|
|
49
|
+
- **Costs from replay metadata** (`activeagent`). A `Replay`'s metadata may
|
|
50
|
+
carry `cost_source`, `cost_rate` and `judge_usage`; `Report#summary_by_model`
|
|
51
|
+
adds `reported`, `estimated`, `judge_cost` and `judge_calls`,
|
|
52
|
+
`Report#scenario_costs` gives each scenario's cost across models, and
|
|
53
|
+
`Report#judge_usage` sums every result's judge calls with the run-level
|
|
54
|
+
part passed as `Report.new(judge_usage:)`. `Report.new(release:)` (and
|
|
55
|
+
`Runner.new(release:)`) names the release the run scored, in
|
|
56
|
+
`to_h["release"]` and a header chip. `to_h` is unchanged when neither is
|
|
57
|
+
given.
|
|
58
|
+
- **An evaluation's standing against the agent as it is now**
|
|
59
|
+
(`actionagent`). Each evaluation reports its `headline_run_id` (its
|
|
60
|
+
newest complete run; a newer pending or failed run shows beside it), its
|
|
61
|
+
`standing` — `current`, `stale`, `unrecorded`, `archived` or `none`
|
|
62
|
+
(`ActionAgent::EvaluationStanding`) — its `archived_at` and `per_model`
|
|
63
|
+
passes, and every run its `agent_version` and `version_state`. Only
|
|
64
|
+
model-facing edits (instructions, action prompts, tools, MCP servers,
|
|
65
|
+
model config, response format) make a run stale. A published report's
|
|
66
|
+
`report.release` pins the run to that release, recorded as a version when
|
|
67
|
+
the dashboard has not seen the digest (`Agent#find_or_record_release!`,
|
|
68
|
+
which never moves a deploy's `release_digest` backwards); a report with
|
|
69
|
+
no release leaves the run unrecorded. `PATCH /api/evaluations/:id` with
|
|
70
|
+
`evaluation: { archived: true | false }` archives an evaluation or brings
|
|
71
|
+
it back; `GET /api/evaluations` leaves archived evaluations out before
|
|
72
|
+
its 50-row limit unless `?archived=1`, and returns `archived_count`. A
|
|
73
|
+
new run or a published report brings an archived evaluation back.
|
|
74
|
+
|
|
75
|
+
- Codex code sessions in checkout sandboxes. Connect an OpenAI API key under
|
|
76
|
+
Settings → Integrations and select Codex in the code-session panel. The local
|
|
77
|
+
backend runs `codex exec` with JSONL events, workspace-write sandboxing, stdin
|
|
78
|
+
prompts, per-sandbox configuration, cancellation, timeout and diff capture.
|
|
79
|
+
- Explicit code-runner capability checks for host backends. Existing adapters
|
|
80
|
+
continue to support Claude Code without implicitly receiving Codex credentials.
|
|
81
|
+
|
|
82
|
+
### Changed
|
|
83
|
+
|
|
84
|
+
- **The dashboard reads passes, scores and costs one way** (`actionagent`
|
|
85
|
+
frontend). Every fraction carries its percent (`14/16 · 88%`), every 0..1
|
|
86
|
+
score reads as a whole percent, and every cost reads `$0.0243` when
|
|
87
|
+
reported or `~$0.0243` when any part was estimated, with the legend once
|
|
88
|
+
per surface and the tokens × rate working in the figure's tooltip. The
|
|
89
|
+
Evaluations page's tiles pool only the headline run of each current or
|
|
90
|
+
unrecorded evaluation, show a pass line per model, and say how many
|
|
91
|
+
evaluations were left out as stale or archived; a card carries its
|
|
92
|
+
standing, an archive/unarchive control and, for a newer run still
|
|
93
|
+
pending or failed, that run's badge beside the headline's; *Show
|
|
94
|
+
archived (n)* lists the archived ones. The suite panel gains the spend
|
|
95
|
+
strip between Runs and Models, the matrix a cost line per cell
|
|
96
|
+
(`~$0.0243 · judge ~$0.0015`) and a trailing Cost column with group
|
|
97
|
+
subtotals, the model comparison a Judge column, the runs list a version
|
|
98
|
+
chip and a same-version / new-version / release-not-recorded line under
|
|
99
|
+
each delta, and the spend strip reads the judge's cost from the traces
|
|
100
|
+
and says "rules only" only when there is no judge. The agent cards'
|
|
101
|
+
Eval tile is the pooled pass rate with its fraction in the title.
|
|
102
|
+
- **The agent card's Eval tile is the pooled pass rate** (`actionagent`
|
|
103
|
+
`AgentScorecard`) over the headline runs of the agent's current and
|
|
104
|
+
unrecorded evaluations — never a stale suite's or an archived one's —
|
|
105
|
+
rather than the mean criterion score of whichever run was last.
|
|
106
|
+
`eval_runs` and `eval_not_counted` say what was pooled and what was left
|
|
107
|
+
out.
|
|
108
|
+
- **The partial-cost notes are retired** (`actionagent`, `activeagent`).
|
|
109
|
+
The `*` marker and the "k of n priced" notes of 1.8.0's partial-cost fix
|
|
110
|
+
give way to the `~` mark and its legend: a cost that covers only some of
|
|
111
|
+
a model's replays is a lower bound and reads as an estimate. The verdict
|
|
112
|
+
rationale reads `Passed 2 of 2 scenarios (100%) with a mean score of 100%
|
|
113
|
+
at ~$0.0010 (estimated)`, and the judge ruling on a comparison is told
|
|
114
|
+
each model's agent cost alone — never the judge's own spend, which never
|
|
115
|
+
enters the ranking either. The HTML report's header shows a chip per
|
|
116
|
+
scalar metadata value only: an array or object (the judge's trace ids)
|
|
117
|
+
is no longer rendered as one.
|
|
118
|
+
- **What to fix comes after the scenario results** (`actionagent`,
|
|
119
|
+
`activeagent`). A run report now reads models, then the scenario results,
|
|
120
|
+
then What to fix. The scenario suite panel moves What to fix below the
|
|
121
|
+
scenario matrix. The standalone HTML report
|
|
122
|
+
(`ActiveAgent::Evals::ReportHtml`) moves its fix cards below the matrix and
|
|
123
|
+
the per-scenario details. `Report#to_markdown` moves its Recommendations
|
|
124
|
+
below its Answers. The sampling run detail already read in this order.
|
|
125
|
+
Each section's content is unchanged.
|
|
126
|
+
|
|
127
|
+
### Fixed
|
|
128
|
+
|
|
129
|
+
- Inherited Codex settings and credentials are removed from sandbox process
|
|
130
|
+
environments. Each Codex run receives only its owner's selected connection.
|
|
131
|
+
|
|
132
|
+
- **An evaluation's cost when some interactions carried no cost estimate**
|
|
133
|
+
(`actionagent`, `activeagent`). A run's cost sums only the interactions that
|
|
134
|
+
were priced, but its per-interaction rate divided that partial sum by every
|
|
135
|
+
interaction, and nothing said part of the run was unpriced. The rate is now
|
|
136
|
+
over the priced interactions, and a run's `usage` counts them: `priced` and
|
|
137
|
+
`unpriced` beside `replays` (or `samples`). A run where nothing was priced
|
|
138
|
+
reports no `cost` or `per_interaction`, as before, with `priced: 0`. The
|
|
139
|
+
per-model summaries count them too: `Report#summary_by_model` adds `priced`
|
|
140
|
+
per model and a sampling run's `_cohorts` add `priced` per cohort. A
|
|
141
|
+
scenario run recorded before that has its `_models` counted from its
|
|
142
|
+
results when the API serves it (`EvaluationRun#model_summaries`); a
|
|
143
|
+
sampling cohort recorded before that reads as fully priced when it has a
|
|
144
|
+
cost. The dashboard shows a partial cost as an estimate — "estimated, 3 of
|
|
145
|
+
5 replays priced" — on the run's spend strip and footer, the model
|
|
146
|
+
scorecards and the Evaluations page's cost-per-interaction tile, and marks
|
|
147
|
+
it `*` in the runs list, the model comparison table and the spend strip's
|
|
148
|
+
total, whose titles give the count. The HTML
|
|
149
|
+
and Markdown reports and the pass-rate verdict name the priced count beside
|
|
150
|
+
a partial cost, the judge ruling on a comparison is told it, and a
|
|
151
|
+
pass-rate verdict breaks a tie on cost per priced scenario rather than on
|
|
152
|
+
the partial sum, so an unpriced replay no longer makes a model look
|
|
153
|
+
cheaper.
|
|
154
|
+
|
|
155
|
+
Upgrade both gems together, then run `bin/rails generate action_agent:install
|
|
156
|
+
--skip` and `bin/rails db:migrate`. The migration adds runner identity to code
|
|
157
|
+
sessions; existing sessions remain Claude Code sessions.
|
|
158
|
+
|
|
159
|
+
## [1.8.0] - 2026-09-29
|
|
160
|
+
|
|
161
|
+
Releases `activeagent` and `actionagent` 1.8.0 from one tag. A minor release.
|
|
162
|
+
Settings -> Integrations connects GitHub and Claude Code. A connected
|
|
163
|
+
repository's checkout boots as a sandbox, on a developer's machine with the new
|
|
164
|
+
`:local` backend, where Claude Code sessions run and an evaluation can run
|
|
165
|
+
against the checkout without the agent being edited. The dashboard's MCP
|
|
166
|
+
server gains evaluation and telemetry tools for a developer's own coding
|
|
167
|
+
harness. Ollama hosts can be tested and can be remote, with an optional Bearer
|
|
168
|
+
API key. Comparison runs lead with a per-model table and filter the fix list
|
|
169
|
+
by model. A publisher can check its collector before a run.
|
|
170
|
+
|
|
171
|
+
Upgrading: run `bin/rails generate action_agent:install --skip` and
|
|
172
|
+
`bin/rails db:migrate`. The generator adds what an install lacks:
|
|
173
|
+
`provider_keys.api_key` and the `github_connections` and `code_sessions`
|
|
174
|
+
tables. An app on the RubyLLM provider needs ruby_llm 1.16 or later; 2.x
|
|
175
|
+
works too.
|
|
176
|
+
|
|
177
|
+
The Claude Code connection stores Anthropic API keys only. Anthropic does not
|
|
178
|
+
let third-party products collect, store or route requests through Claude.ai
|
|
179
|
+
subscription credentials
|
|
180
|
+
([Claude Code legal and compliance](https://code.claude.com/docs/en/legal-and-compliance.md)),
|
|
181
|
+
so a `claude setup-token` token (`sk-ant-oat…`) is refused. An install that
|
|
182
|
+
stored one while running a pre-release build never hands it to a session: its
|
|
183
|
+
owner sees Claude Code as needing an API key (`needs_replacing: true` in
|
|
184
|
+
`GET /api/provider_keys`) until they paste one.
|
|
185
|
+
`bin/rails action_agent:claude_code:purge_subscription_tokens` deletes the
|
|
186
|
+
stored tokens and prints how many it removed. A developer who wants sessions on
|
|
187
|
+
their own Claude login sets `config.claude_code_auth = :local_login` with the
|
|
188
|
+
`:local` backend and runs `claude /login` on that machine instead.
|
|
189
|
+
|
|
190
|
+
### Added
|
|
191
|
+
|
|
192
|
+
- **Evaluation and telemetry tools on the MCP facade** (`actionagent`). The
|
|
193
|
+
dashboard's MCP server now offers `evaluations_list`, `evaluations_get`,
|
|
194
|
+
`evaluations_run`, `evaluation_runs_get`, `evaluation_runs_compare`,
|
|
195
|
+
`traces_search` and `traces_get`, so a developer's own coding harness can
|
|
196
|
+
run an agent's evaluations, read the fix items and failing traces, and
|
|
197
|
+
iterate on the agent in its own checkout without the dashboard holding a
|
|
198
|
+
model login. The tools read under the API key's owner exactly as the JSON
|
|
199
|
+
API reads under the signed-in owner, run through the same execution,
|
|
200
|
+
quota and sandbox checks as `POST /api/evaluations/:id/run`, bound their
|
|
201
|
+
output, and mask the owner's credentials. Their names cannot collide with
|
|
202
|
+
schema tools or agent tools. `ActionAgent.mcp_dashboard_tools = false`
|
|
203
|
+
turns them off.
|
|
204
|
+
|
|
205
|
+
- **Local checkout sandboxes and Claude Code sessions** (`actionagent`, #489).
|
|
206
|
+
A new `:local` sandbox backend (`config.sandbox_service = :local`) makes
|
|
207
|
+
**Start sandbox** work on a developer's machine without containers. It
|
|
208
|
+
clones the repository under `tmp/action_agent/sandboxes`, runs the setup
|
|
209
|
+
the checkout's optional `.activeagents/sandbox.yml` names (`env`, `setup`,
|
|
210
|
+
`manifest`, `start`), and boots the app on `127.0.0.1`. It then registers
|
|
211
|
+
the app's MCP facade as the `sandbox:<session_id>` server. Every process
|
|
212
|
+
starts from a sanitized copy of the dashboard's environment: without its
|
|
213
|
+
database and Redis URLs, Rails keys and environment, Bundler and Ruby
|
|
214
|
+
settings, git repository and config variables, `SSH_AUTH_SOCK`,
|
|
215
|
+
model-provider and Claude Code settings, variables named like a secret, or
|
|
216
|
+
URLs carrying credentials. The GitHub token reaches
|
|
217
|
+
only the fetch, and the Claude Code API key reaches only Claude Code.
|
|
218
|
+
`:local` runs the owner's code with the dashboard's privileges, so it is off
|
|
219
|
+
outside development and test unless
|
|
220
|
+
`ActionAgent.local_sandboxes_enabled = true`.
|
|
221
|
+
- `bin/rails action_agent:sandbox:manifest` writes the booted app's
|
|
222
|
+
`{mcp_path, mcp_token}` for the backend.
|
|
223
|
+
- `bin/rails action_agent:sandbox:reap` expires overdue sandboxes and stops
|
|
224
|
+
them.
|
|
225
|
+
- A ready sandbox runs headless Claude Code sessions (`--permission-mode
|
|
226
|
+
acceptEdits` by default). You start them from Settings -> Integrations, or
|
|
227
|
+
through `/api/sandboxes/:session_id/code_sessions`. Their events stream
|
|
228
|
+
into the dashboard, and the checkout's diff follows.
|
|
229
|
+
- New options: `local_sandboxes_enabled`, `local_sandbox_root`,
|
|
230
|
+
`local_sandbox_boot_timeout`, `claude_code_command`,
|
|
231
|
+
`claude_code_permission_mode`, `claude_code_max_turns`,
|
|
232
|
+
`claude_code_timeout` and `claude_code_auth`.
|
|
233
|
+
- `claude_code_auth = :local_login` runs sessions on the machine's own
|
|
234
|
+
Claude Code login (`claude /login`), with no stored key: the backend
|
|
235
|
+
passes no credential and no `CLAUDE_CONFIG_DIR`, so `claude` uses the
|
|
236
|
+
dashboard user's own `~/.claude` or keychain, which the dashboard never
|
|
237
|
+
reads. `LocalSandboxBackend.claude_login_status` asks
|
|
238
|
+
`claude auth status --json` (cached for a minute) and keeps only
|
|
239
|
+
`loggedIn` and the login method. `GET /api/sandboxes` reports
|
|
240
|
+
`claude_code_auth` and, in this mode, `claude_code_login`
|
|
241
|
+
(`{ logged_in, auth_method }`), and the assistant's
|
|
242
|
+
`connections.claude_code` the same as `auth` and `login`. Other backends
|
|
243
|
+
refuse sessions in this mode (`code_sessions_supported: false`, and a
|
|
244
|
+
`422` naming the reason).
|
|
245
|
+
- `app_runtime` sandboxes now provision in the background and last 2 hours.
|
|
246
|
+
- Run `rails g action_agent:install` to add the
|
|
247
|
+
`create_active_agent_code_sessions` migration.
|
|
248
|
+
- This repository's own `.activeagents/sandbox.yml` boots `test/dummy`.
|
|
249
|
+
- Every `:local` sandbox boots on databases of its own, so a checkout of
|
|
250
|
+
the dashboard's own app no longer migrates the developer's development
|
|
251
|
+
database. The backend reads the adapter from the checkout's
|
|
252
|
+
`config/database.yml` without running its ERB, and sets `DATABASE_URL`
|
|
253
|
+
and `<NAME>_DATABASE_URL` (`QUEUE_DATABASE_URL`, `CACHE_DATABASE_URL`):
|
|
254
|
+
SQLite files in the workspace, or `<database>_sandbox_<id>` on
|
|
255
|
+
PostgreSQL and MySQL, which terminate drops with the checkout's
|
|
256
|
+
Rails database tasks restricted to the names and URLs recorded at boot.
|
|
257
|
+
Setting a variable in `sandbox.yml`'s `env` overrides it and excludes
|
|
258
|
+
that database from cleanup. Replica mappings follow their own writer;
|
|
259
|
+
ambiguous mappings require an explicit URL instead of guessing.
|
|
260
|
+
- The Claude Code panel has a **Model** select: Claude Code's own
|
|
261
|
+
default, the `sonnet`, `opus` and `haiku` aliases, or any model id under
|
|
262
|
+
*Other…*. It remembers the last choice per browser, and each session
|
|
263
|
+
shows the model it ran on.
|
|
264
|
+
- A run can use a checkout sandbox without the agent being edited:
|
|
265
|
+
`sandbox_id` on `POST /api/evaluations/:id/run` (and on the runner's
|
|
266
|
+
`/api/agents/:id/execute` and `/test`) gives that run's tool dispatcher
|
|
267
|
+
the sandbox's `sandbox:<session_id>` runtime, as if the agent listed it.
|
|
268
|
+
The sandbox must be the caller's, a ready `app_runtime` sandbox, and the
|
|
269
|
+
agent owner's; anything else is a `422`. The run records which sandbox
|
|
270
|
+
it used (`run.sandbox`), and a scenario suite's **Run against sandbox**
|
|
271
|
+
select, its Runs list and the run report show it. The selected sandbox
|
|
272
|
+
takes precedence for matching tool names, with one schema per name.
|
|
273
|
+
Queued agent runs fail if their selected sandbox stops, and failed
|
|
274
|
+
sandbox discovery never silently falls back to the original tools.
|
|
275
|
+
- **Claude Code connection** (`actionagent`, #478). Settings -> Integrations
|
|
276
|
+
stores an Anthropic API key (`sk-ant-api…`, from the Claude Console) as the
|
|
277
|
+
`claude_code` provider key. It is encrypted, write-only, and not an agent
|
|
278
|
+
provider. `SandboxSession#runtime_environment` hands it to an
|
|
279
|
+
`app_runtime` backend as `ANTHROPIC_API_KEY`, so the checkout can run
|
|
280
|
+
Claude Code sessions. Claude subscription tokens (`claude setup-token`)
|
|
281
|
+
are refused, as Anthropic's terms require (see the upgrading note above).
|
|
282
|
+
`/api/provider_keys` rows now carry `kind` (`key`, `host` or
|
|
283
|
+
`connection`) and `needs_replacing`.
|
|
284
|
+
- **GitHub connections and checkout sandboxes** (`actionagent`, #477).
|
|
285
|
+
Settings -> Integrations connects GitHub over OAuth
|
|
286
|
+
(`ActionAgent.github_client_id` / `github_client_secret`, or
|
|
287
|
+
`GITHUB_CLIENT_ID` / `GITHUB_CLIENT_SECRET`), stores the token encrypted,
|
|
288
|
+
and lets the owner choose which repositories the workspace may use. Only
|
|
289
|
+
repositories GitHub lists for the token can be selected. A new
|
|
290
|
+
`app_runtime` sandbox type checks out one of them: the backend receives
|
|
291
|
+
`sandbox_session.checkout_spec` (repository, ref, clone URL, token),
|
|
292
|
+
boots the app, and returns `mcp_url` / `mcp_token` from `create_sandbox`.
|
|
293
|
+
The session is then an MCP server keyed `sandbox:<session_id>`. An agent
|
|
294
|
+
that lists that key in `mcp_servers` runs, and is evaluated, with the
|
|
295
|
+
checkout's own tools. Lookups are scoped to the agent's owner. Run
|
|
296
|
+
`rails g action_agent:install` to add the
|
|
297
|
+
`create_active_agent_github_connections` migration.
|
|
298
|
+
|
|
299
|
+
- **Ollama hosts are testable and can be remote** (`actionagent`). Settings ->
|
|
300
|
+
Provider API Keys gains a **Test connection** for Ollama that reports
|
|
301
|
+
whether the server is reachable, the round-trip time and the models it
|
|
302
|
+
serves, before or after saving (`POST <mount>/api/provider_keys/test`,
|
|
303
|
+
read-only). The host is accepted as a bare server address
|
|
304
|
+
(`http://localhost:11434`; the OpenAI-compatible `/v1` path is added) and
|
|
305
|
+
an optional **API key** is stored beside it and sent as a Bearer token,
|
|
306
|
+
for a server behind an authenticating proxy or Ollama Cloud. The agent
|
|
307
|
+
builder's live Ollama model list uses the same probe and key. When no host
|
|
308
|
+
is configured the card shows the host app's `config/active_agent.yml`
|
|
309
|
+
default. The install generator emits a guarded `add_provider_key_api_key`
|
|
310
|
+
migration for existing installs; re-run
|
|
311
|
+
`bin/rails generate action_agent:install --skip` and `bin/rails db:migrate`.
|
|
312
|
+
- **A model comparison table on comparison runs** (`actionagent`,
|
|
313
|
+
`activeagent`). A run over several models now leads its Models section
|
|
314
|
+
with one row per model, best first: passed, mean score, average latency,
|
|
315
|
+
average tokens per scenario, cost (and per scenario), and the model's
|
|
316
|
+
typical fault — its most frequent one with the diagnosis of a result that
|
|
317
|
+
carries it. The scenario suite panel, the sampling run detail and the
|
|
318
|
+
standalone HTML report (`ActiveAgent::Evals::ReportHtml`) all render it.
|
|
319
|
+
- **What to fix, filtered by model** (`actionagent`, `activeagent`). On a
|
|
320
|
+
comparison run the fix list takes a model chip, narrowing to the items
|
|
321
|
+
attributed to that model and counting what that model alone produced,
|
|
322
|
+
since one model may need more instruction than another. The standalone
|
|
323
|
+
report filters through radio chips and stylesheet rules — it still ships
|
|
324
|
+
no script.
|
|
325
|
+
|
|
326
|
+
- **A publisher can check its collector before a run** (`activeagent`).
|
|
327
|
+
`ActiveAgent::Evals::Publisher#verify!` asks the collector whether it is up
|
|
328
|
+
and accepts the key before a run is paid for. It posts an empty JSON object,
|
|
329
|
+
which a compatible collector refuses with a 422 naming `version`, without
|
|
330
|
+
storing anything; anything else raises `Publisher::Error` with a delivery's
|
|
331
|
+
status, detail and guidance. `Publisher#endpoint` returns the collector URL.
|
|
332
|
+
|
|
333
|
+
### Changed
|
|
334
|
+
|
|
335
|
+
- **Collector rejections say what the status means** (`activeagent`). A
|
|
336
|
+
`Publisher::Error` for a 401, 403, 404, 415 or 501 rejection names a refused
|
|
337
|
+
key, an account an operator must act on, an endpoint that is not a
|
|
338
|
+
collector, a rewritten `Content-Type`, or an install with no evaluation
|
|
339
|
+
store, in place of the generic guidance.
|
|
340
|
+
|
|
341
|
+
### Fixed
|
|
342
|
+
|
|
343
|
+
- **A nested scenario expectation written as one value** (`activeagent`).
|
|
344
|
+
`ScenarioParser` now stores `{ expectations: { contains: "30" } }` as a
|
|
345
|
+
list of one, the shape the persisted scenario and the dashboard's matrix
|
|
346
|
+
read; an object-list import with a lone value used to break the suite
|
|
347
|
+
panel. The matrix also tolerates scenarios persisted before this.
|
|
348
|
+
|
|
10
349
|
## [1.7.2] - 2026-09-29
|
|
11
350
|
|
|
12
351
|
Releases `activeagent` and `actionagent` 1.7.2 from one tag. A patch on 1.7.1:
|
|
@@ -285,11 +285,11 @@ module ActiveAgent
|
|
|
285
285
|
|
|
286
286
|
weakest = (failed_grade ? grades : @scores.compact).min_by { |_, value| value }
|
|
287
287
|
summary = if failed_grade
|
|
288
|
-
"#{graded_label(grades)} scored #{
|
|
288
|
+
"#{graded_label(grades)} scored #{Format.score(grade)} against a pass threshold of #{Format.percent(@threshold)}"
|
|
289
289
|
else
|
|
290
|
-
"Scored #{@score
|
|
290
|
+
"Scored #{Format.score(@score)} against a pass threshold of #{Format.percent(@threshold)}"
|
|
291
291
|
end
|
|
292
|
-
summary += ", weakest on #{weakest.first} (#{weakest.last
|
|
292
|
+
summary += ", weakest on #{weakest.first} (#{Format.score(weakest.last)})" if weakest
|
|
293
293
|
recommendation =
|
|
294
294
|
if weakest
|
|
295
295
|
"Read the answer against the #{weakest.first.to_s.humanize.downcase} criterion and adjust the " \
|
|
@@ -0,0 +1,115 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module ActiveAgent
|
|
4
|
+
module Evals
|
|
5
|
+
# How every evaluation surface writes a pass count, a score and a cost, so
|
|
6
|
+
# the HTML report, the Markdown report, a diagnosis and the dashboard read
|
|
7
|
+
# alike:
|
|
8
|
+
#
|
|
9
|
+
# - a pass/fail fraction carries its percentage: "14/16 · 88%" (in
|
|
10
|
+
# Markdown "14/16 (88%)"); nothing scored (0/0) reads "—"
|
|
11
|
+
# - a 0..1 score — a mean score, a criterion score, task completion, a
|
|
12
|
+
# judge's confidence, the pass threshold — reads as a whole percent:
|
|
13
|
+
# "93%", "pass ≥ 70%"
|
|
14
|
+
# - money has four decimals, six below $0.001, "$0.00" for an explicit
|
|
15
|
+
# zero, and a leading "~" when any part of it was estimated from
|
|
16
|
+
# tokens × model rates rather than reported; LEGEND explains the mark
|
|
17
|
+
#
|
|
18
|
+
# Percentages round half up to a whole number. Stored values stay 0..1
|
|
19
|
+
# and USD; only their display changes here. The dashboard's
|
|
20
|
+
# actionagent/frontend/utils/evalFormat.mjs is the JavaScript copy of
|
|
21
|
+
# these rules, so a change here needs the same change there.
|
|
22
|
+
module Format
|
|
23
|
+
EMPTY = "—"
|
|
24
|
+
# Shown once per surface where a "~" figure is visible.
|
|
25
|
+
LEGEND = "~ estimated from tokens × model rates"
|
|
26
|
+
# Rate sources that are a guess at the model's price rather than a
|
|
27
|
+
# catalog entry for it (see the dashboard's ModelPricing).
|
|
28
|
+
FALLBACK_RATE_SOURCES = %w[pattern default].freeze
|
|
29
|
+
|
|
30
|
+
module_function
|
|
31
|
+
|
|
32
|
+
# 0.875 → "88%". "—" for a missing value.
|
|
33
|
+
def percent(fraction)
|
|
34
|
+
return EMPTY unless finite?(fraction)
|
|
35
|
+
|
|
36
|
+
"#{whole(fraction.to_f * 100)}%"
|
|
37
|
+
end
|
|
38
|
+
|
|
39
|
+
# 14 of 16 → "14/16 · 88%", or "14/16 (88%)" with `style: :markdown`,
|
|
40
|
+
# where " · " would read as a table cell separator next to a pipe.
|
|
41
|
+
# "—" when nothing was scored.
|
|
42
|
+
def passes(passed, total, style: :text)
|
|
43
|
+
count = total.to_i
|
|
44
|
+
return EMPTY unless count.positive?
|
|
45
|
+
|
|
46
|
+
done = passed.to_i
|
|
47
|
+
pct = "#{whole(done * 100.0 / count)}%"
|
|
48
|
+
style == :markdown ? "#{done}/#{count} (#{pct})" : "#{done}/#{count} · #{pct}"
|
|
49
|
+
end
|
|
50
|
+
|
|
51
|
+
# A 0..1 score as a percent: 0.93 → "93%".
|
|
52
|
+
def score(value)
|
|
53
|
+
percent(value)
|
|
54
|
+
end
|
|
55
|
+
|
|
56
|
+
# 0.7 → "pass ≥ 70%".
|
|
57
|
+
def threshold(value)
|
|
58
|
+
"pass ≥ #{percent(value)}"
|
|
59
|
+
end
|
|
60
|
+
|
|
61
|
+
# 0.0243 → "$0.0243", 0.000697 → "$0.000697", 0 → "$0.00"; with
|
|
62
|
+
# `estimated`, "~$0.0243". "—" for a missing value.
|
|
63
|
+
def money(value, estimated: false)
|
|
64
|
+
return EMPTY unless finite?(value)
|
|
65
|
+
|
|
66
|
+
amount = value.to_f
|
|
67
|
+
digits = if amount.zero? then 2
|
|
68
|
+
elsif amount.abs < 0.001 then 6
|
|
69
|
+
else 4
|
|
70
|
+
end
|
|
71
|
+
"#{'~' if estimated}$#{format("%.#{digits}f", amount)}"
|
|
72
|
+
end
|
|
73
|
+
|
|
74
|
+
# A $/M token rate with at least two decimals and no trailing noise:
|
|
75
|
+
# 5 → "$5.00/M", 0.075 → "$0.075/M".
|
|
76
|
+
def per_million(rate)
|
|
77
|
+
whole, fraction = format("%.4f", rate.to_f).split(".")
|
|
78
|
+
fraction = fraction.sub(/0+\z/, "")
|
|
79
|
+
"$#{whole}.#{fraction.ljust(2, '0')}/M"
|
|
80
|
+
end
|
|
81
|
+
|
|
82
|
+
# The tooltip of an estimated figure: how it was worked out, e.g.
|
|
83
|
+
# "estimated: 2,328 in × $5.00/M + 423 out × $30.00/M · catalog rate".
|
|
84
|
+
# `rate` is `{ "input", "output", "source" }` in $ per million tokens;
|
|
85
|
+
# a rate from the name-pattern table or the default appends
|
|
86
|
+
# "(fallback rate)". Without a rate it names the method only.
|
|
87
|
+
def cost_title(input_tokens: nil, output_tokens: nil, rate: nil)
|
|
88
|
+
rate = rate.to_h.transform_keys(&:to_s) if rate.respond_to?(:to_h)
|
|
89
|
+
return "estimated from tokens × model rates" unless rate.is_a?(Hash) && finite?(rate["input"]) && finite?(rate["output"])
|
|
90
|
+
|
|
91
|
+
source = rate["source"].presence || "catalog"
|
|
92
|
+
title = "estimated: #{delimited(input_tokens)} in × #{per_million(rate['input'])} + " \
|
|
93
|
+
"#{delimited(output_tokens)} out × #{per_million(rate['output'])} · #{source} rate"
|
|
94
|
+
title += " (fallback rate)" if FALLBACK_RATE_SOURCES.include?(source.to_s)
|
|
95
|
+
title
|
|
96
|
+
end
|
|
97
|
+
|
|
98
|
+
# Half up, on the decimal value: 14.5 → 15, even when the float arrived
|
|
99
|
+
# as 14.499999999999998.
|
|
100
|
+
def whole(value)
|
|
101
|
+
value.to_f.round(9).round
|
|
102
|
+
end
|
|
103
|
+
|
|
104
|
+
def delimited(count)
|
|
105
|
+
count.to_i.to_s.gsub(/(\d)(?=(\d{3})+\z)/, '\1,')
|
|
106
|
+
end
|
|
107
|
+
|
|
108
|
+
def finite?(value)
|
|
109
|
+
return false if value.nil? || value == ""
|
|
110
|
+
|
|
111
|
+
Float(value, exception: false)&.finite? || false
|
|
112
|
+
end
|
|
113
|
+
end
|
|
114
|
+
end
|
|
115
|
+
end
|
|
@@ -158,6 +158,8 @@ module ActiveAgent
|
|
|
158
158
|
def verdict(summaries, instructions: nil)
|
|
159
159
|
lines = summaries.map do |label, stats|
|
|
160
160
|
faults = (stats["faults"] || {}).map { |fault, count| "#{fault}×#{count}" }.join(", ")
|
|
161
|
+
# The agent's cost only: what the judge itself spent scoring a
|
|
162
|
+
# cohort says nothing about the model under comparison.
|
|
161
163
|
"#{label}: pass rate #{stats['pass_rate']}%, mean score #{stats['avg_score'] || 'n/a'}, " \
|
|
162
164
|
"avg latency #{stats['avg_duration_ms'] || 'n/a'}ms, cost $#{stats['cost'] || 'n/a'}" \
|
|
163
165
|
"#{", faults: #{faults}" if faults.present?}"
|
|
@@ -11,7 +11,8 @@ module ActiveAgent
|
|
|
11
11
|
# Publishes a completed report without replaying the agent. The caller must
|
|
12
12
|
# retain run_id when retrying: compatible collectors treat that identity as
|
|
13
13
|
# immutable within the authenticated account. Delivery is blocking and does
|
|
14
|
-
# not follow redirects with the account's bearer credential.
|
|
14
|
+
# not follow redirects with the account's bearer credential. +verify!+ asks
|
|
15
|
+
# the collector whether it is up and accepts the key before a run is paid for.
|
|
15
16
|
#
|
|
16
17
|
# Every failure to deliver raises Error. Invalid arguments raise
|
|
17
18
|
# ArgumentError before anything is sent.
|
|
@@ -20,6 +21,9 @@ module ActiveAgent
|
|
|
20
21
|
MAX_BYTES = 2 * 1024 * 1024
|
|
21
22
|
DETAIL_LIMIT = 200
|
|
22
23
|
|
|
24
|
+
# @return [String] the collector URL reports go to
|
|
25
|
+
attr_reader :endpoint
|
|
26
|
+
|
|
23
27
|
# Raised for every failed delivery. Only a network failure keeps the
|
|
24
28
|
# underlying error as its +cause+, so neither the response nor the
|
|
25
29
|
# report reaches a log through the exception chain.
|
|
@@ -48,10 +52,15 @@ module ActiveAgent
|
|
|
48
52
|
# Whether each rejection status is retryable, and what the caller should
|
|
49
53
|
# do about it. Other statuses fall back to the rules in +rejection+.
|
|
50
54
|
REJECTIONS = {
|
|
55
|
+
401 => [ false, "the collector refused the API key; check the key against the collector's account" ],
|
|
56
|
+
403 => [ false, "the account may not store this report until an operator acts, for example on a cap on observed agents, evaluations or scenarios; resolve that before retrying" ],
|
|
57
|
+
404 => [ false, "nothing at the endpoint takes evaluation reports; check that it is a collector's /v1/evaluations or <mount>/api/evaluation_reports URL" ],
|
|
51
58
|
409 => [ false, "the collector already holds a different report under this run_id; never retry this report with the same run_id" ],
|
|
52
59
|
413 => [ false, "the report exceeds the collector's size limit; publish a smaller selection" ],
|
|
60
|
+
415 => [ false, "the collector did not receive application/json; check anything between the publisher and the collector that rewrites the Content-Type" ],
|
|
53
61
|
422 => [ false, "correct what the collector refused before retrying" ],
|
|
54
|
-
429 => [ true, "the account is over its quota or rate limit; retain the report and run_id and retry later" ]
|
|
62
|
+
429 => [ true, "the account is over its quota or rate limit; retain the report and run_id and retry later" ],
|
|
63
|
+
501 => [ false, "the collector has no evaluation store; migrate the install, or publish to one generated with evaluation tables" ]
|
|
55
64
|
}.freeze
|
|
56
65
|
|
|
57
66
|
# The key is sent as a bearer token and filtered from the collector's
|
|
@@ -65,6 +74,7 @@ module ActiveAgent
|
|
|
65
74
|
unless @uri.scheme == "https" || %w[localhost 127.0.0.1 ::1].include?(@uri.hostname)
|
|
66
75
|
raise ArgumentError, "Evaluation endpoint requires HTTPS except on loopback hosts"
|
|
67
76
|
end
|
|
77
|
+
@endpoint = @uri.to_s
|
|
68
78
|
|
|
69
79
|
@api_key = api_key.to_s.strip
|
|
70
80
|
raise ArgumentError, "Evaluation API key is required" if @api_key.empty?
|
|
@@ -91,6 +101,52 @@ module ActiveAgent
|
|
|
91
101
|
body = encode(identities.merge("version" => 1, "report" => report_hash(report)))
|
|
92
102
|
raise Error, "Evaluation report exceeds the 2 MiB delivery limit; publish a smaller selection" if body.bytesize > MAX_BYTES
|
|
93
103
|
|
|
104
|
+
deliver("retain the report and run_id for retry") do
|
|
105
|
+
response = post(body)
|
|
106
|
+
raise rejection(response) unless %w[200 201].include?(response.code)
|
|
107
|
+
|
|
108
|
+
receipt = JSON.parse(response.body.to_s)
|
|
109
|
+
unless receipt.is_a?(Hash) && receipt["run_id"] == run_id && receipt["status"] == "complete" && receipt["id"] && receipt["evaluation_id"]
|
|
110
|
+
raise Error.new("Evaluation collector returned an invalid completion receipt; retain the report and run_id for retry", retryable: true)
|
|
111
|
+
end
|
|
112
|
+
receipt
|
|
113
|
+
end
|
|
114
|
+
end
|
|
115
|
+
|
|
116
|
+
# Returns true when the collector is up and accepts the API key, without
|
|
117
|
+
# storing anything. It posts an empty JSON object, which a compatible
|
|
118
|
+
# collector authenticates, parses, and then refuses with a 422 whose
|
|
119
|
+
# +error+ names +version+, as not a version-1 report. That refusal, and
|
|
120
|
+
# only that one, is the ready answer. Call it before an expensive run, so
|
|
121
|
+
# a stopped collector or a refused key fails before the first model call
|
|
122
|
+
# rather than after the last one.
|
|
123
|
+
#
|
|
124
|
+
# Raises Error for anything else: a rejection other than 422, with the
|
|
125
|
+
# status, detail and guidance +call+ would carry (401 for a refused key,
|
|
126
|
+
# 404 when nothing at the endpoint takes reports); a retryable delivery
|
|
127
|
+
# failure when the collector cannot be reached; a 422 that says anything
|
|
128
|
+
# else; or a collector that stores the empty object. The last two mean the
|
|
129
|
+
# endpoint is not a compatible collector.
|
|
130
|
+
def verify!
|
|
131
|
+
deliver("start the collector or check the endpoint, then verify again") do
|
|
132
|
+
response = post("{}")
|
|
133
|
+
if response.code == "422"
|
|
134
|
+
detail = collector_detail(response.body)
|
|
135
|
+
next true if detail&.match?(/\bversion\b/)
|
|
136
|
+
|
|
137
|
+
raise Error.new("Evaluation collector answered HTTP 422 without refusing the empty envelope as a version-1 report, " \
|
|
138
|
+
"so it is not a compatible collector; check the endpoint", status: 422, detail: detail)
|
|
139
|
+
end
|
|
140
|
+
raise rejection(response) unless %w[200 201].include?(response.code)
|
|
141
|
+
|
|
142
|
+
raise Error, "Evaluation collector stored an empty report, so it is not a compatible collector; check the endpoint"
|
|
143
|
+
end
|
|
144
|
+
end
|
|
145
|
+
|
|
146
|
+
private
|
|
147
|
+
|
|
148
|
+
# Posts +body+ to the endpoint with the bearer key and the delivery timeouts.
|
|
149
|
+
def post(body)
|
|
94
150
|
http = Net::HTTP.new(@uri.hostname, @uri.port)
|
|
95
151
|
http.use_ssl = @uri.scheme == "https"
|
|
96
152
|
http.open_timeout = @open_timeout
|
|
@@ -101,26 +157,25 @@ module ActiveAgent
|
|
|
101
157
|
request["Content-Type"] = "application/json"
|
|
102
158
|
request["Accept"] = "application/json"
|
|
103
159
|
request.body = body
|
|
104
|
-
|
|
105
|
-
|
|
160
|
+
http.request(request)
|
|
161
|
+
end
|
|
106
162
|
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
163
|
+
# Runs one exchange with the collector, turning every failure to reach it
|
|
164
|
+
# or to read its answer into a retryable Error that quotes nothing from
|
|
165
|
+
# the response and ends with +guidance+, what the caller should do next.
|
|
166
|
+
# An Error the block raises passes through unchanged.
|
|
167
|
+
def deliver(guidance)
|
|
168
|
+
yield
|
|
112
169
|
rescue JSON::ParserError
|
|
113
170
|
# The parser's message quotes the body.
|
|
114
|
-
raise Error.new("Evaluation collector returned invalid JSON;
|
|
171
|
+
raise Error.new("Evaluation collector returned invalid JSON; #{guidance}", retryable: true), cause: nil
|
|
115
172
|
rescue Net::HTTPBadResponse, Net::HTTPHeaderSyntaxError, Zlib::Error => e
|
|
116
173
|
# These messages can quote the response's status line, headers or body.
|
|
117
|
-
raise Error.new("Evaluation collector returned a malformed response (#{e.class});
|
|
174
|
+
raise Error.new("Evaluation collector returned a malformed response (#{e.class}); #{guidance}", retryable: true), cause: nil
|
|
118
175
|
rescue IOError, SocketError, SystemCallError, Timeout::Error, OpenSSL::SSL::SSLError => e
|
|
119
|
-
raise Error.new("Evaluation delivery failed (#{e.class});
|
|
176
|
+
raise Error.new("Evaluation delivery failed (#{e.class}); #{guidance}", retryable: true)
|
|
120
177
|
end
|
|
121
178
|
|
|
122
|
-
private
|
|
123
|
-
|
|
124
179
|
def report_hash(report)
|
|
125
180
|
hash = report.to_h if report.respond_to?(:to_h) && !report.nil? && !report.is_a?(Array)
|
|
126
181
|
raise ArgumentError, "report must be a Report or its saved JSON hash" unless hash.is_a?(Hash)
|