activeagent 1.8.0 → 1.8.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +149 -0
- data/lib/active_agent/evals/diagnosis.rb +3 -3
- data/lib/active_agent/evals/format.rb +115 -0
- data/lib/active_agent/evals/judge.rb +2 -0
- data/lib/active_agent/evals/report.rb +232 -31
- data/lib/active_agent/evals/report_html.rb +316 -149
- data/lib/active_agent/evals/result.rb +40 -0
- data/lib/active_agent/evals/runner.rb +6 -2
- data/lib/active_agent/evals.rb +1 -0
- data/lib/active_agent/version.rb +1 -1
- metadata +3 -2
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 1c7be2b31907ff2779ea3818516d887afecc6c540c89670c33f7ec1be242c01a
|
|
4
|
+
data.tar.gz: 0d7327bb8f94d2573f6e87b9191afbc7758b2fdb7d42a2ea896fbd25dfd71b17
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: fd0f4bbab5439daabe85170c3fb57fe5c94260e0573fc35944e80cc8301bd35803fb5d89f2a6de9c6061c44752e67f125db9ad17d116058b444e92785c3819c6
|
|
7
|
+
data.tar.gz: 453ecb45269cc816fbea41378591e642a460a2a188e052745b87eed1fa9f7777281449a7adfe19ebef17e12e242009ad46ee6a773f0cc33660b161f6d4585197
|
data/CHANGELOG.md
CHANGED
|
@@ -7,6 +7,155 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
7
7
|
|
|
8
8
|
## [Unreleased]
|
|
9
9
|
|
|
10
|
+
## [1.8.1] - 2026-10-01
|
|
11
|
+
|
|
12
|
+
### Added
|
|
13
|
+
|
|
14
|
+
- **Every evaluation cost is priced, and says how** (`actionagent`,
|
|
15
|
+
`activeagent`). A scenario result with no cost is priced down a chain —
|
|
16
|
+
the cost the publishing application reported, the estimate the engine
|
|
17
|
+
stored, the result's tokens × its model's rate, the tokens of the trace
|
|
18
|
+
it links to, or its text at four characters a token as a lower bound —
|
|
19
|
+
and a result that recorded no tokens costs `$0.00` (`cost_source:
|
|
20
|
+
no_usage`); only a result with no tokens, no trace and no text stays
|
|
21
|
+
unpriced. Results carry `cost` (the effective figure), `reported_cost`,
|
|
22
|
+
`cost_source`, `cost_rate` and `judge_usage`; a run's `usage` adds
|
|
23
|
+
`reported`, `estimated`, `cost_basis` and `total`; `scores._models` adds
|
|
24
|
+
per model `reported`, `estimated`, `judge_cost` and `judge_calls`; and
|
|
25
|
+
`GET /api/evaluations/:id/runs/:run_id` adds `costs` per scenario (judge
|
|
26
|
+
apart) and for the run (`ActionAgent::EvaluationRunCost`, cached per
|
|
27
|
+
finished run). The judge's spend is found the same way — the engine's
|
|
28
|
+
meter, the application's figures (`result.judge_usage`,
|
|
29
|
+
`report.judge_usage.run`, accepted by the report import) or the judge
|
|
30
|
+
traces priced on input and output tokens, never thinking tokens — within
|
|
31
|
+
the run's tenant. `ModelPricing` looks rates up under the provider the
|
|
32
|
+
model ran on, strips gateway and vendor prefixes and date suffixes, tries
|
|
33
|
+
dots and dashes both ways, prices `claude-sonnet-5` and the `gpt-5`
|
|
34
|
+
family by exact rows ahead of the family patterns, and reports where a
|
|
35
|
+
rate came from (`estimate_detailed`, `rate_detail`, `fingerprint`). The
|
|
36
|
+
engine's judge meter counts Anthropic's cached prompt tokens.
|
|
37
|
+
- **One display format for passes, scores and costs** (`activeagent`
|
|
38
|
+
`ActiveAgent::Evals::Format`). A fraction always carries its percent
|
|
39
|
+
(`14/16 · 88%`; `14/16 (88%)` in Markdown; `—` for nothing scored), every
|
|
40
|
+
0..1 score reads as a whole percent (`93%`, `pass ≥ 70%`), and money
|
|
41
|
+
reads `$0.0243` when reported and `~$0.0243` when any part was estimated,
|
|
42
|
+
with one legend per surface. The HTML report gains a Cost tile, a Judge
|
|
43
|
+
column and per-model judge line, a trailing matrix Cost column with
|
|
44
|
+
group subtotals, a cost line per cell and in each result's details, a
|
|
45
|
+
judge chip with its calls and cost, a release chip, the pass mark and
|
|
46
|
+
the cost in its footer; the Markdown report gains a Judge column, a
|
|
47
|
+
matrix Cost column, the cost per answer and a total line. Diagnosis
|
|
48
|
+
text reads `Task completion scored 60% against a pass threshold of 70%`.
|
|
49
|
+
- **Costs from replay metadata** (`activeagent`). A `Replay`'s metadata may
|
|
50
|
+
carry `cost_source`, `cost_rate` and `judge_usage`; `Report#summary_by_model`
|
|
51
|
+
adds `reported`, `estimated`, `judge_cost` and `judge_calls`,
|
|
52
|
+
`Report#scenario_costs` gives each scenario's cost across models, and
|
|
53
|
+
`Report#judge_usage` sums every result's judge calls with the run-level
|
|
54
|
+
part passed as `Report.new(judge_usage:)`. `Report.new(release:)` (and
|
|
55
|
+
`Runner.new(release:)`) names the release the run scored, in
|
|
56
|
+
`to_h["release"]` and a header chip. `to_h` is unchanged when neither is
|
|
57
|
+
given.
|
|
58
|
+
- **An evaluation's standing against the agent as it is now**
|
|
59
|
+
(`actionagent`). Each evaluation reports its `headline_run_id` (its
|
|
60
|
+
newest complete run; a newer pending or failed run shows beside it), its
|
|
61
|
+
`standing` — `current`, `stale`, `unrecorded`, `archived` or `none`
|
|
62
|
+
(`ActionAgent::EvaluationStanding`) — its `archived_at` and `per_model`
|
|
63
|
+
passes, and every run its `agent_version` and `version_state`. Only
|
|
64
|
+
model-facing edits (instructions, action prompts, tools, MCP servers,
|
|
65
|
+
model config, response format) make a run stale. A published report's
|
|
66
|
+
`report.release` pins the run to that release, recorded as a version when
|
|
67
|
+
the dashboard has not seen the digest (`Agent#find_or_record_release!`,
|
|
68
|
+
which never moves a deploy's `release_digest` backwards); a report with
|
|
69
|
+
no release leaves the run unrecorded. `PATCH /api/evaluations/:id` with
|
|
70
|
+
`evaluation: { archived: true | false }` archives an evaluation or brings
|
|
71
|
+
it back; `GET /api/evaluations` leaves archived evaluations out before
|
|
72
|
+
its 50-row limit unless `?archived=1`, and returns `archived_count`. A
|
|
73
|
+
new run or a published report brings an archived evaluation back.
|
|
74
|
+
|
|
75
|
+
- Codex code sessions in checkout sandboxes. Connect an OpenAI API key under
|
|
76
|
+
Settings → Integrations and select Codex in the code-session panel. The local
|
|
77
|
+
backend runs `codex exec` with JSONL events, workspace-write sandboxing, stdin
|
|
78
|
+
prompts, per-sandbox configuration, cancellation, timeout and diff capture.
|
|
79
|
+
- Explicit code-runner capability checks for host backends. Existing adapters
|
|
80
|
+
continue to support Claude Code without implicitly receiving Codex credentials.
|
|
81
|
+
|
|
82
|
+
### Changed
|
|
83
|
+
|
|
84
|
+
- **The dashboard reads passes, scores and costs one way** (`actionagent`
|
|
85
|
+
frontend). Every fraction carries its percent (`14/16 · 88%`), every 0..1
|
|
86
|
+
score reads as a whole percent, and every cost reads `$0.0243` when
|
|
87
|
+
reported or `~$0.0243` when any part was estimated, with the legend once
|
|
88
|
+
per surface and the tokens × rate working in the figure's tooltip. The
|
|
89
|
+
Evaluations page's tiles pool only the headline run of each current or
|
|
90
|
+
unrecorded evaluation, show a pass line per model, and say how many
|
|
91
|
+
evaluations were left out as stale or archived; a card carries its
|
|
92
|
+
standing, an archive/unarchive control and, for a newer run still
|
|
93
|
+
pending or failed, that run's badge beside the headline's; *Show
|
|
94
|
+
archived (n)* lists the archived ones. The suite panel gains the spend
|
|
95
|
+
strip between Runs and Models, the matrix a cost line per cell
|
|
96
|
+
(`~$0.0243 · judge ~$0.0015`) and a trailing Cost column with group
|
|
97
|
+
subtotals, the model comparison a Judge column, the runs list a version
|
|
98
|
+
chip and a same-version / new-version / release-not-recorded line under
|
|
99
|
+
each delta, and the spend strip reads the judge's cost from the traces
|
|
100
|
+
and says "rules only" only when there is no judge. The agent cards'
|
|
101
|
+
Eval tile is the pooled pass rate with its fraction in the title.
|
|
102
|
+
- **The agent card's Eval tile is the pooled pass rate** (`actionagent`
|
|
103
|
+
`AgentScorecard`) over the headline runs of the agent's current and
|
|
104
|
+
unrecorded evaluations — never a stale suite's or an archived one's —
|
|
105
|
+
rather than the mean criterion score of whichever run was last.
|
|
106
|
+
`eval_runs` and `eval_not_counted` say what was pooled and what was left
|
|
107
|
+
out.
|
|
108
|
+
- **The partial-cost notes are retired** (`actionagent`, `activeagent`).
|
|
109
|
+
The `*` marker and the "k of n priced" notes of 1.8.0's partial-cost fix
|
|
110
|
+
give way to the `~` mark and its legend: a cost that covers only some of
|
|
111
|
+
a model's replays is a lower bound and reads as an estimate. The verdict
|
|
112
|
+
rationale reads `Passed 2 of 2 scenarios (100%) with a mean score of 100%
|
|
113
|
+
at ~$0.0010 (estimated)`, and the judge ruling on a comparison is told
|
|
114
|
+
each model's agent cost alone — never the judge's own spend, which never
|
|
115
|
+
enters the ranking either. The HTML report's header shows a chip per
|
|
116
|
+
scalar metadata value only: an array or object (the judge's trace ids)
|
|
117
|
+
is no longer rendered as one.
|
|
118
|
+
- **What to fix comes after the scenario results** (`actionagent`,
|
|
119
|
+
`activeagent`). A run report now reads models, then the scenario results,
|
|
120
|
+
then What to fix. The scenario suite panel moves What to fix below the
|
|
121
|
+
scenario matrix. The standalone HTML report
|
|
122
|
+
(`ActiveAgent::Evals::ReportHtml`) moves its fix cards below the matrix and
|
|
123
|
+
the per-scenario details. `Report#to_markdown` moves its Recommendations
|
|
124
|
+
below its Answers. The sampling run detail already read in this order.
|
|
125
|
+
Each section's content is unchanged.
|
|
126
|
+
|
|
127
|
+
### Fixed
|
|
128
|
+
|
|
129
|
+
- Inherited Codex settings and credentials are removed from sandbox process
|
|
130
|
+
environments. Each Codex run receives only its owner's selected connection.
|
|
131
|
+
|
|
132
|
+
- **An evaluation's cost when some interactions carried no cost estimate**
|
|
133
|
+
(`actionagent`, `activeagent`). A run's cost sums only the interactions that
|
|
134
|
+
were priced, but its per-interaction rate divided that partial sum by every
|
|
135
|
+
interaction, and nothing said part of the run was unpriced. The rate is now
|
|
136
|
+
over the priced interactions, and a run's `usage` counts them: `priced` and
|
|
137
|
+
`unpriced` beside `replays` (or `samples`). A run where nothing was priced
|
|
138
|
+
reports no `cost` or `per_interaction`, as before, with `priced: 0`. The
|
|
139
|
+
per-model summaries count them too: `Report#summary_by_model` adds `priced`
|
|
140
|
+
per model and a sampling run's `_cohorts` add `priced` per cohort. A
|
|
141
|
+
scenario run recorded before that has its `_models` counted from its
|
|
142
|
+
results when the API serves it (`EvaluationRun#model_summaries`); a
|
|
143
|
+
sampling cohort recorded before that reads as fully priced when it has a
|
|
144
|
+
cost. The dashboard shows a partial cost as an estimate — "estimated, 3 of
|
|
145
|
+
5 replays priced" — on the run's spend strip and footer, the model
|
|
146
|
+
scorecards and the Evaluations page's cost-per-interaction tile, and marks
|
|
147
|
+
it `*` in the runs list, the model comparison table and the spend strip's
|
|
148
|
+
total, whose titles give the count. The HTML
|
|
149
|
+
and Markdown reports and the pass-rate verdict name the priced count beside
|
|
150
|
+
a partial cost, the judge ruling on a comparison is told it, and a
|
|
151
|
+
pass-rate verdict breaks a tie on cost per priced scenario rather than on
|
|
152
|
+
the partial sum, so an unpriced replay no longer makes a model look
|
|
153
|
+
cheaper.
|
|
154
|
+
|
|
155
|
+
Upgrade both gems together, then run `bin/rails generate action_agent:install
|
|
156
|
+
--skip` and `bin/rails db:migrate`. The migration adds runner identity to code
|
|
157
|
+
sessions; existing sessions remain Claude Code sessions.
|
|
158
|
+
|
|
10
159
|
## [1.8.0] - 2026-09-29
|
|
11
160
|
|
|
12
161
|
Releases `activeagent` and `actionagent` 1.8.0 from one tag. A minor release.
|
|
@@ -285,11 +285,11 @@ module ActiveAgent
|
|
|
285
285
|
|
|
286
286
|
weakest = (failed_grade ? grades : @scores.compact).min_by { |_, value| value }
|
|
287
287
|
summary = if failed_grade
|
|
288
|
-
"#{graded_label(grades)} scored #{
|
|
288
|
+
"#{graded_label(grades)} scored #{Format.score(grade)} against a pass threshold of #{Format.percent(@threshold)}"
|
|
289
289
|
else
|
|
290
|
-
"Scored #{@score
|
|
290
|
+
"Scored #{Format.score(@score)} against a pass threshold of #{Format.percent(@threshold)}"
|
|
291
291
|
end
|
|
292
|
-
summary += ", weakest on #{weakest.first} (#{weakest.last
|
|
292
|
+
summary += ", weakest on #{weakest.first} (#{Format.score(weakest.last)})" if weakest
|
|
293
293
|
recommendation =
|
|
294
294
|
if weakest
|
|
295
295
|
"Read the answer against the #{weakest.first.to_s.humanize.downcase} criterion and adjust the " \
|
|
@@ -0,0 +1,115 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module ActiveAgent
|
|
4
|
+
module Evals
|
|
5
|
+
# How every evaluation surface writes a pass count, a score and a cost, so
|
|
6
|
+
# the HTML report, the Markdown report, a diagnosis and the dashboard read
|
|
7
|
+
# alike:
|
|
8
|
+
#
|
|
9
|
+
# - a pass/fail fraction carries its percentage: "14/16 · 88%" (in
|
|
10
|
+
# Markdown "14/16 (88%)"); nothing scored (0/0) reads "—"
|
|
11
|
+
# - a 0..1 score — a mean score, a criterion score, task completion, a
|
|
12
|
+
# judge's confidence, the pass threshold — reads as a whole percent:
|
|
13
|
+
# "93%", "pass ≥ 70%"
|
|
14
|
+
# - money has four decimals, six below $0.001, "$0.00" for an explicit
|
|
15
|
+
# zero, and a leading "~" when any part of it was estimated from
|
|
16
|
+
# tokens × model rates rather than reported; LEGEND explains the mark
|
|
17
|
+
#
|
|
18
|
+
# Percentages round half up to a whole number. Stored values stay 0..1
|
|
19
|
+
# and USD; only their display changes here. The dashboard's
|
|
20
|
+
# actionagent/frontend/utils/evalFormat.mjs is the JavaScript copy of
|
|
21
|
+
# these rules, so a change here needs the same change there.
|
|
22
|
+
module Format
|
|
23
|
+
EMPTY = "—"
|
|
24
|
+
# Shown once per surface where a "~" figure is visible.
|
|
25
|
+
LEGEND = "~ estimated from tokens × model rates"
|
|
26
|
+
# Rate sources that are a guess at the model's price rather than a
|
|
27
|
+
# catalog entry for it (see the dashboard's ModelPricing).
|
|
28
|
+
FALLBACK_RATE_SOURCES = %w[pattern default].freeze
|
|
29
|
+
|
|
30
|
+
module_function
|
|
31
|
+
|
|
32
|
+
# 0.875 → "88%". "—" for a missing value.
|
|
33
|
+
def percent(fraction)
|
|
34
|
+
return EMPTY unless finite?(fraction)
|
|
35
|
+
|
|
36
|
+
"#{whole(fraction.to_f * 100)}%"
|
|
37
|
+
end
|
|
38
|
+
|
|
39
|
+
# 14 of 16 → "14/16 · 88%", or "14/16 (88%)" with `style: :markdown`,
|
|
40
|
+
# where " · " would read as a table cell separator next to a pipe.
|
|
41
|
+
# "—" when nothing was scored.
|
|
42
|
+
def passes(passed, total, style: :text)
|
|
43
|
+
count = total.to_i
|
|
44
|
+
return EMPTY unless count.positive?
|
|
45
|
+
|
|
46
|
+
done = passed.to_i
|
|
47
|
+
pct = "#{whole(done * 100.0 / count)}%"
|
|
48
|
+
style == :markdown ? "#{done}/#{count} (#{pct})" : "#{done}/#{count} · #{pct}"
|
|
49
|
+
end
|
|
50
|
+
|
|
51
|
+
# A 0..1 score as a percent: 0.93 → "93%".
|
|
52
|
+
def score(value)
|
|
53
|
+
percent(value)
|
|
54
|
+
end
|
|
55
|
+
|
|
56
|
+
# 0.7 → "pass ≥ 70%".
|
|
57
|
+
def threshold(value)
|
|
58
|
+
"pass ≥ #{percent(value)}"
|
|
59
|
+
end
|
|
60
|
+
|
|
61
|
+
# 0.0243 → "$0.0243", 0.000697 → "$0.000697", 0 → "$0.00"; with
|
|
62
|
+
# `estimated`, "~$0.0243". "—" for a missing value.
|
|
63
|
+
def money(value, estimated: false)
|
|
64
|
+
return EMPTY unless finite?(value)
|
|
65
|
+
|
|
66
|
+
amount = value.to_f
|
|
67
|
+
digits = if amount.zero? then 2
|
|
68
|
+
elsif amount.abs < 0.001 then 6
|
|
69
|
+
else 4
|
|
70
|
+
end
|
|
71
|
+
"#{'~' if estimated}$#{format("%.#{digits}f", amount)}"
|
|
72
|
+
end
|
|
73
|
+
|
|
74
|
+
# A $/M token rate with at least two decimals and no trailing noise:
|
|
75
|
+
# 5 → "$5.00/M", 0.075 → "$0.075/M".
|
|
76
|
+
def per_million(rate)
|
|
77
|
+
whole, fraction = format("%.4f", rate.to_f).split(".")
|
|
78
|
+
fraction = fraction.sub(/0+\z/, "")
|
|
79
|
+
"$#{whole}.#{fraction.ljust(2, '0')}/M"
|
|
80
|
+
end
|
|
81
|
+
|
|
82
|
+
# The tooltip of an estimated figure: how it was worked out, e.g.
|
|
83
|
+
# "estimated: 2,328 in × $5.00/M + 423 out × $30.00/M · catalog rate".
|
|
84
|
+
# `rate` is `{ "input", "output", "source" }` in $ per million tokens;
|
|
85
|
+
# a rate from the name-pattern table or the default appends
|
|
86
|
+
# "(fallback rate)". Without a rate it names the method only.
|
|
87
|
+
def cost_title(input_tokens: nil, output_tokens: nil, rate: nil)
|
|
88
|
+
rate = rate.to_h.transform_keys(&:to_s) if rate.respond_to?(:to_h)
|
|
89
|
+
return "estimated from tokens × model rates" unless rate.is_a?(Hash) && finite?(rate["input"]) && finite?(rate["output"])
|
|
90
|
+
|
|
91
|
+
source = rate["source"].presence || "catalog"
|
|
92
|
+
title = "estimated: #{delimited(input_tokens)} in × #{per_million(rate['input'])} + " \
|
|
93
|
+
"#{delimited(output_tokens)} out × #{per_million(rate['output'])} · #{source} rate"
|
|
94
|
+
title += " (fallback rate)" if FALLBACK_RATE_SOURCES.include?(source.to_s)
|
|
95
|
+
title
|
|
96
|
+
end
|
|
97
|
+
|
|
98
|
+
# Half up, on the decimal value: 14.5 → 15, even when the float arrived
|
|
99
|
+
# as 14.499999999999998.
|
|
100
|
+
def whole(value)
|
|
101
|
+
value.to_f.round(9).round
|
|
102
|
+
end
|
|
103
|
+
|
|
104
|
+
def delimited(count)
|
|
105
|
+
count.to_i.to_s.gsub(/(\d)(?=(\d{3})+\z)/, '\1,')
|
|
106
|
+
end
|
|
107
|
+
|
|
108
|
+
def finite?(value)
|
|
109
|
+
return false if value.nil? || value == ""
|
|
110
|
+
|
|
111
|
+
Float(value, exception: false)&.finite? || false
|
|
112
|
+
end
|
|
113
|
+
end
|
|
114
|
+
end
|
|
115
|
+
end
|
|
@@ -158,6 +158,8 @@ module ActiveAgent
|
|
|
158
158
|
def verdict(summaries, instructions: nil)
|
|
159
159
|
lines = summaries.map do |label, stats|
|
|
160
160
|
faults = (stats["faults"] || {}).map { |fault, count| "#{fault}×#{count}" }.join(", ")
|
|
161
|
+
# The agent's cost only: what the judge itself spent scoring a
|
|
162
|
+
# cohort says nothing about the model under comparison.
|
|
161
163
|
"#{label}: pass rate #{stats['pass_rate']}%, mean score #{stats['avg_score'] || 'n/a'}, " \
|
|
162
164
|
"avg latency #{stats['avg_duration_ms'] || 'n/a'}ms, cost $#{stats['cost'] || 'n/a'}" \
|
|
163
165
|
"#{", faults: #{faults}" if faults.present?}"
|