activeagent 1.8.0 → 1.8.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 3642068294052d483decf5260b5d7dd1793647f06920713ca658f7b89ba948b0
4
- data.tar.gz: 1d0811bee23dc646e916d17abc711637c1e2d41055633ace8bb1c067615d26b4
3
+ metadata.gz: 1c7be2b31907ff2779ea3818516d887afecc6c540c89670c33f7ec1be242c01a
4
+ data.tar.gz: 0d7327bb8f94d2573f6e87b9191afbc7758b2fdb7d42a2ea896fbd25dfd71b17
5
5
  SHA512:
6
- metadata.gz: 6d4492797c3d4b238bf86c49c48b9ec12e5876bbc9ef075c4cae7f369a3e3274b3db30781b9367643740bd9fa365eb8fe8ad81308b125f6843f3f828ccf5f339
7
- data.tar.gz: 601676a4590808eef1469f3cfafc5136c43bccbd22e7b600c15940eb7a1fc957479c568676d699d7d41c10836fa621eb072a3ff5fd20d5f4cfa22a517e982416
6
+ metadata.gz: fd0f4bbab5439daabe85170c3fb57fe5c94260e0573fc35944e80cc8301bd35803fb5d89f2a6de9c6061c44752e67f125db9ad17d116058b444e92785c3819c6
7
+ data.tar.gz: 453ecb45269cc816fbea41378591e642a460a2a188e052745b87eed1fa9f7777281449a7adfe19ebef17e12e242009ad46ee6a773f0cc33660b161f6d4585197
data/CHANGELOG.md CHANGED
@@ -7,6 +7,155 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [1.8.1] - 2026-10-01
11
+
12
+ ### Added
13
+
14
+ - **Every evaluation cost is priced, and says how** (`actionagent`,
15
+ `activeagent`). A scenario result with no cost is priced down a chain —
16
+ the cost the publishing application reported, the estimate the engine
17
+ stored, the result's tokens × its model's rate, the tokens of the trace
18
+ it links to, or its text at four characters a token as a lower bound —
19
+ and a result that recorded no tokens costs `$0.00` (`cost_source:
20
+ no_usage`); only a result with no tokens, no trace and no text stays
21
+ unpriced. Results carry `cost` (the effective figure), `reported_cost`,
22
+ `cost_source`, `cost_rate` and `judge_usage`; a run's `usage` adds
23
+ `reported`, `estimated`, `cost_basis` and `total`; `scores._models` adds
24
+ per model `reported`, `estimated`, `judge_cost` and `judge_calls`; and
25
+ `GET /api/evaluations/:id/runs/:run_id` adds `costs` per scenario (judge
26
+ apart) and for the run (`ActionAgent::EvaluationRunCost`, cached per
27
+ finished run). The judge's spend is found the same way — the engine's
28
+ meter, the application's figures (`result.judge_usage`,
29
+ `report.judge_usage.run`, accepted by the report import) or the judge
30
+ traces priced on input and output tokens, never thinking tokens — within
31
+ the run's tenant. `ModelPricing` looks rates up under the provider the
32
+ model ran on, strips gateway and vendor prefixes and date suffixes, tries
33
+ dots and dashes both ways, prices `claude-sonnet-5` and the `gpt-5`
34
+ family by exact rows ahead of the family patterns, and reports where a
35
+ rate came from (`estimate_detailed`, `rate_detail`, `fingerprint`). The
36
+ engine's judge meter counts Anthropic's cached prompt tokens.
37
+ - **One display format for passes, scores and costs** (`activeagent`
38
+ `ActiveAgent::Evals::Format`). A fraction always carries its percent
39
+ (`14/16 · 88%`; `14/16 (88%)` in Markdown; `—` for nothing scored), every
40
+ 0..1 score reads as a whole percent (`93%`, `pass ≥ 70%`), and money
41
+ reads `$0.0243` when reported and `~$0.0243` when any part was estimated,
42
+ with one legend per surface. The HTML report gains a Cost tile, a Judge
43
+ column and per-model judge line, a trailing matrix Cost column with
44
+ group subtotals, a cost line per cell and in each result's details, a
45
+ judge chip with its calls and cost, a release chip, the pass mark and
46
+ the cost in its footer; the Markdown report gains a Judge column, a
47
+ matrix Cost column, the cost per answer and a total line. Diagnosis
48
+ text reads `Task completion scored 60% against a pass threshold of 70%`.
49
+ - **Costs from replay metadata** (`activeagent`). A `Replay`'s metadata may
50
+ carry `cost_source`, `cost_rate` and `judge_usage`; `Report#summary_by_model`
51
+ adds `reported`, `estimated`, `judge_cost` and `judge_calls`,
52
+ `Report#scenario_costs` gives each scenario's cost across models, and
53
+ `Report#judge_usage` sums every result's judge calls with the run-level
54
+ part passed as `Report.new(judge_usage:)`. `Report.new(release:)` (and
55
+ `Runner.new(release:)`) names the release the run scored, in
56
+ `to_h["release"]` and a header chip. `to_h` is unchanged when neither is
57
+ given.
58
+ - **An evaluation's standing against the agent as it is now**
59
+ (`actionagent`). Each evaluation reports its `headline_run_id` (its
60
+ newest complete run; a newer pending or failed run shows beside it), its
61
+ `standing` — `current`, `stale`, `unrecorded`, `archived` or `none`
62
+ (`ActionAgent::EvaluationStanding`) — its `archived_at` and `per_model`
63
+ passes, and every run its `agent_version` and `version_state`. Only
64
+ model-facing edits (instructions, action prompts, tools, MCP servers,
65
+ model config, response format) make a run stale. A published report's
66
+ `report.release` pins the run to that release, recorded as a version when
67
+ the dashboard has not seen the digest (`Agent#find_or_record_release!`,
68
+ which never moves a deploy's `release_digest` backwards); a report with
69
+ no release leaves the run unrecorded. `PATCH /api/evaluations/:id` with
70
+ `evaluation: { archived: true | false }` archives an evaluation or brings
71
+ it back; `GET /api/evaluations` leaves archived evaluations out before
72
+ its 50-row limit unless `?archived=1`, and returns `archived_count`. A
73
+ new run or a published report brings an archived evaluation back.
74
+
75
+ - Codex code sessions in checkout sandboxes. Connect an OpenAI API key under
76
+ Settings → Integrations and select Codex in the code-session panel. The local
77
+ backend runs `codex exec` with JSONL events, workspace-write sandboxing, stdin
78
+ prompts, per-sandbox configuration, cancellation, timeout and diff capture.
79
+ - Explicit code-runner capability checks for host backends. Existing adapters
80
+ continue to support Claude Code without implicitly receiving Codex credentials.
81
+
82
+ ### Changed
83
+
84
+ - **The dashboard reads passes, scores and costs one way** (`actionagent`
85
+ frontend). Every fraction carries its percent (`14/16 · 88%`), every 0..1
86
+ score reads as a whole percent, and every cost reads `$0.0243` when
87
+ reported or `~$0.0243` when any part was estimated, with the legend once
88
+ per surface and the tokens × rate working in the figure's tooltip. The
89
+ Evaluations page's tiles pool only the headline run of each current or
90
+ unrecorded evaluation, show a pass line per model, and say how many
91
+ evaluations were left out as stale or archived; a card carries its
92
+ standing, an archive/unarchive control and, for a newer run still
93
+ pending or failed, that run's badge beside the headline's; *Show
94
+ archived (n)* lists the archived ones. The suite panel gains the spend
95
+ strip between Runs and Models, the matrix a cost line per cell
96
+ (`~$0.0243 · judge ~$0.0015`) and a trailing Cost column with group
97
+ subtotals, the model comparison a Judge column, the runs list a version
98
+ chip and a same-version / new-version / release-not-recorded line under
99
+ each delta, and the spend strip reads the judge's cost from the traces
100
+ and says "rules only" only when there is no judge. The agent cards'
101
+ Eval tile is the pooled pass rate with its fraction in the title.
102
+ - **The agent card's Eval tile is the pooled pass rate** (`actionagent`
103
+ `AgentScorecard`) over the headline runs of the agent's current and
104
+ unrecorded evaluations — never a stale suite's or an archived one's —
105
+ rather than the mean criterion score of whichever run was last.
106
+ `eval_runs` and `eval_not_counted` say what was pooled and what was left
107
+ out.
108
+ - **The partial-cost notes are retired** (`actionagent`, `activeagent`).
109
+ The `*` marker and the "k of n priced" notes of 1.8.0's partial-cost fix
110
+ give way to the `~` mark and its legend: a cost that covers only some of
111
+ a model's replays is a lower bound and reads as an estimate. The verdict
112
+ rationale reads `Passed 2 of 2 scenarios (100%) with a mean score of 100%
113
+ at ~$0.0010 (estimated)`, and the judge ruling on a comparison is told
114
+ each model's agent cost alone — never the judge's own spend, which never
115
+ enters the ranking either. The HTML report's header shows a chip per
116
+ scalar metadata value only: an array or object (the judge's trace ids)
117
+ is no longer rendered as one.
118
+ - **What to fix comes after the scenario results** (`actionagent`,
119
+ `activeagent`). A run report now reads models, then the scenario results,
120
+ then What to fix. The scenario suite panel moves What to fix below the
121
+ scenario matrix. The standalone HTML report
122
+ (`ActiveAgent::Evals::ReportHtml`) moves its fix cards below the matrix and
123
+ the per-scenario details. `Report#to_markdown` moves its Recommendations
124
+ below its Answers. The sampling run detail already read in this order.
125
+ Each section's content is unchanged.
126
+
127
+ ### Fixed
128
+
129
+ - Inherited Codex settings and credentials are removed from sandbox process
130
+ environments. Each Codex run receives only its owner's selected connection.
131
+
132
+ - **An evaluation's cost when some interactions carried no cost estimate**
133
+ (`actionagent`, `activeagent`). A run's cost sums only the interactions that
134
+ were priced, but its per-interaction rate divided that partial sum by every
135
+ interaction, and nothing said part of the run was unpriced. The rate is now
136
+ over the priced interactions, and a run's `usage` counts them: `priced` and
137
+ `unpriced` beside `replays` (or `samples`). A run where nothing was priced
138
+ reports no `cost` or `per_interaction`, as before, with `priced: 0`. The
139
+ per-model summaries count them too: `Report#summary_by_model` adds `priced`
140
+ per model and a sampling run's `_cohorts` add `priced` per cohort. A
141
+ scenario run recorded before that has its `_models` counted from its
142
+ results when the API serves it (`EvaluationRun#model_summaries`); a
143
+ sampling cohort recorded before that reads as fully priced when it has a
144
+ cost. The dashboard shows a partial cost as an estimate — "estimated, 3 of
145
+ 5 replays priced" — on the run's spend strip and footer, the model
146
+ scorecards and the Evaluations page's cost-per-interaction tile, and marks
147
+ it `*` in the runs list, the model comparison table and the spend strip's
148
+ total, whose titles give the count. The HTML
149
+ and Markdown reports and the pass-rate verdict name the priced count beside
150
+ a partial cost, the judge ruling on a comparison is told it, and a
151
+ pass-rate verdict breaks a tie on cost per priced scenario rather than on
152
+ the partial sum, so an unpriced replay no longer makes a model look
153
+ cheaper.
154
+
155
+ Upgrade both gems together, then run `bin/rails generate action_agent:install
156
+ --skip` and `bin/rails db:migrate`. The migration adds runner identity to code
157
+ sessions; existing sessions remain Claude Code sessions.
158
+
10
159
  ## [1.8.0] - 2026-09-29
11
160
 
12
161
  Releases `activeagent` and `actionagent` 1.8.0 from one tag. A minor release.
@@ -285,11 +285,11 @@ module ActiveAgent
285
285
 
286
286
  weakest = (failed_grade ? grades : @scores.compact).min_by { |_, value| value }
287
287
  summary = if failed_grade
288
- "#{graded_label(grades)} scored #{grade.round(2)} against a pass threshold of #{@threshold}"
288
+ "#{graded_label(grades)} scored #{Format.score(grade)} against a pass threshold of #{Format.percent(@threshold)}"
289
289
  else
290
- "Scored #{@score.round(2)} against a pass threshold of #{@threshold}"
290
+ "Scored #{Format.score(@score)} against a pass threshold of #{Format.percent(@threshold)}"
291
291
  end
292
- summary += ", weakest on #{weakest.first} (#{weakest.last.round(2)})" if weakest
292
+ summary += ", weakest on #{weakest.first} (#{Format.score(weakest.last)})" if weakest
293
293
  recommendation =
294
294
  if weakest
295
295
  "Read the answer against the #{weakest.first.to_s.humanize.downcase} criterion and adjust the " \
@@ -0,0 +1,115 @@
1
+ # frozen_string_literal: true
2
+
3
+ module ActiveAgent
4
+ module Evals
5
+ # How every evaluation surface writes a pass count, a score and a cost, so
6
+ # the HTML report, the Markdown report, a diagnosis and the dashboard read
7
+ # alike:
8
+ #
9
+ # - a pass/fail fraction carries its percentage: "14/16 · 88%" (in
10
+ # Markdown "14/16 (88%)"); nothing scored (0/0) reads "—"
11
+ # - a 0..1 score — a mean score, a criterion score, task completion, a
12
+ # judge's confidence, the pass threshold — reads as a whole percent:
13
+ # "93%", "pass ≥ 70%"
14
+ # - money has four decimals, six below $0.001, "$0.00" for an explicit
15
+ # zero, and a leading "~" when any part of it was estimated from
16
+ # tokens × model rates rather than reported; LEGEND explains the mark
17
+ #
18
+ # Percentages round half up to a whole number. Stored values stay 0..1
19
+ # and USD; only their display changes here. The dashboard's
20
+ # actionagent/frontend/utils/evalFormat.mjs is the JavaScript copy of
21
+ # these rules, so a change here needs the same change there.
22
+ module Format
23
+ EMPTY = "—"
24
+ # Shown once per surface where a "~" figure is visible.
25
+ LEGEND = "~ estimated from tokens × model rates"
26
+ # Rate sources that are a guess at the model's price rather than a
27
+ # catalog entry for it (see the dashboard's ModelPricing).
28
+ FALLBACK_RATE_SOURCES = %w[pattern default].freeze
29
+
30
+ module_function
31
+
32
+ # 0.875 → "88%". "—" for a missing value.
33
+ def percent(fraction)
34
+ return EMPTY unless finite?(fraction)
35
+
36
+ "#{whole(fraction.to_f * 100)}%"
37
+ end
38
+
39
+ # 14 of 16 → "14/16 · 88%", or "14/16 (88%)" with `style: :markdown`,
40
+ # where " · " would read as a table cell separator next to a pipe.
41
+ # "—" when nothing was scored.
42
+ def passes(passed, total, style: :text)
43
+ count = total.to_i
44
+ return EMPTY unless count.positive?
45
+
46
+ done = passed.to_i
47
+ pct = "#{whole(done * 100.0 / count)}%"
48
+ style == :markdown ? "#{done}/#{count} (#{pct})" : "#{done}/#{count} · #{pct}"
49
+ end
50
+
51
+ # A 0..1 score as a percent: 0.93 → "93%".
52
+ def score(value)
53
+ percent(value)
54
+ end
55
+
56
+ # 0.7 → "pass ≥ 70%".
57
+ def threshold(value)
58
+ "pass ≥ #{percent(value)}"
59
+ end
60
+
61
+ # 0.0243 → "$0.0243", 0.000697 → "$0.000697", 0 → "$0.00"; with
62
+ # `estimated`, "~$0.0243". "—" for a missing value.
63
+ def money(value, estimated: false)
64
+ return EMPTY unless finite?(value)
65
+
66
+ amount = value.to_f
67
+ digits = if amount.zero? then 2
68
+ elsif amount.abs < 0.001 then 6
69
+ else 4
70
+ end
71
+ "#{'~' if estimated}$#{format("%.#{digits}f", amount)}"
72
+ end
73
+
74
+ # A $/M token rate with at least two decimals and no trailing noise:
75
+ # 5 → "$5.00/M", 0.075 → "$0.075/M".
76
+ def per_million(rate)
77
+ whole, fraction = format("%.4f", rate.to_f).split(".")
78
+ fraction = fraction.sub(/0+\z/, "")
79
+ "$#{whole}.#{fraction.ljust(2, '0')}/M"
80
+ end
81
+
82
+ # The tooltip of an estimated figure: how it was worked out, e.g.
83
+ # "estimated: 2,328 in × $5.00/M + 423 out × $30.00/M · catalog rate".
84
+ # `rate` is `{ "input", "output", "source" }` in $ per million tokens;
85
+ # a rate from the name-pattern table or the default appends
86
+ # "(fallback rate)". Without a rate it names the method only.
87
+ def cost_title(input_tokens: nil, output_tokens: nil, rate: nil)
88
+ rate = rate.to_h.transform_keys(&:to_s) if rate.respond_to?(:to_h)
89
+ return "estimated from tokens × model rates" unless rate.is_a?(Hash) && finite?(rate["input"]) && finite?(rate["output"])
90
+
91
+ source = rate["source"].presence || "catalog"
92
+ title = "estimated: #{delimited(input_tokens)} in × #{per_million(rate['input'])} + " \
93
+ "#{delimited(output_tokens)} out × #{per_million(rate['output'])} · #{source} rate"
94
+ title += " (fallback rate)" if FALLBACK_RATE_SOURCES.include?(source.to_s)
95
+ title
96
+ end
97
+
98
+ # Half up, on the decimal value: 14.5 → 15, even when the float arrived
99
+ # as 14.499999999999998.
100
+ def whole(value)
101
+ value.to_f.round(9).round
102
+ end
103
+
104
+ def delimited(count)
105
+ count.to_i.to_s.gsub(/(\d)(?=(\d{3})+\z)/, '\1,')
106
+ end
107
+
108
+ def finite?(value)
109
+ return false if value.nil? || value == ""
110
+
111
+ Float(value, exception: false)&.finite? || false
112
+ end
113
+ end
114
+ end
115
+ end
@@ -158,6 +158,8 @@ module ActiveAgent
158
158
  def verdict(summaries, instructions: nil)
159
159
  lines = summaries.map do |label, stats|
160
160
  faults = (stats["faults"] || {}).map { |fault, count| "#{fault}×#{count}" }.join(", ")
161
+ # The agent's cost only: what the judge itself spent scoring a
162
+ # cohort says nothing about the model under comparison.
161
163
  "#{label}: pass rate #{stats['pass_rate']}%, mean score #{stats['avg_score'] || 'n/a'}, " \
162
164
  "avg latency #{stats['avg_duration_ms'] || 'n/a'}ms, cost $#{stats['cost'] || 'n/a'}" \
163
165
  "#{", faults: #{faults}" if faults.present?}"