activeagent 1.6.0 → 1.6.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +222 -1
- data/lib/active_agent/base.rb +2 -0
- data/lib/active_agent/concerns/release.rb +171 -0
- data/lib/active_agent/evals/correlation.rb +178 -0
- data/lib/active_agent/evals/judge.rb +44 -6
- data/lib/active_agent/evals.rb +3 -1
- data/lib/active_agent/schema_tools.rb +108 -6
- data/lib/active_agent/telemetry/configuration.rb +11 -0
- data/lib/active_agent/telemetry/instrumentation.rb +26 -1
- data/lib/active_agent/telemetry/tracer.rb +4 -1
- data/lib/active_agent/telemetry.rb +22 -0
- data/lib/active_agent/version.rb +1 -1
- metadata +8 -3
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 1611f71e0ff6a30b864e797b01cd20242a46acad9df3888c836799f0d297e3d1
|
|
4
|
+
data.tar.gz: 5f93c9a8c43a98cf7eb599d9efb37f2915c1d2e5ef6f4138464c33cd1f4d3af1
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 7e27caea807dc1d32744260bf157086160d73935fba48aea9c4c59e3234c26a5aee707eb650104a76f2fdf2e59497b0408b6a80392b83ceeab2770b4c8eef3c0
|
|
7
|
+
data.tar.gz: e44dc91193d40a6fafa9bc79c3cfc3a8c4f2acc215fc0fae0229424218a8c1bc6ca5b96c5803553bb610f89b86eb519b2f0c6061018117b65a38d77eba61100b
|
data/CHANGELOG.md
CHANGED
|
@@ -7,6 +7,228 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
7
7
|
|
|
8
8
|
## [Unreleased]
|
|
9
9
|
|
|
10
|
+
## [1.6.3] - 2026-09-18
|
|
11
|
+
|
|
12
|
+
Releases `activeagent` and `actionagent` 1.6.3 from one tag.
|
|
13
|
+
|
|
14
|
+
A release about telling the truth on the screens that report what happened.
|
|
15
|
+
The context meter now divides the provider's own `prompt_tokens` among its
|
|
16
|
+
segments instead of subtracting estimates from it, and sizes each piece —
|
|
17
|
+
tool schemas, MCP schemas, instructions and the transcript — before the span
|
|
18
|
+
clips it for storage; previously a trace with dense tool schemas showed a
|
|
19
|
+
large message history that was never sent. An adapted replay is metered as
|
|
20
|
+
one execution like any other, so a host that supplies its own runtime is no
|
|
21
|
+
longer silently uncounted, and a spec naming a provider the agent cannot
|
|
22
|
+
serve now fails before the replay rather than reaching it.
|
|
23
|
+
|
|
24
|
+
Three seams hosts were reaching around become API. `ActiveAgent::Evals::Correlation`
|
|
25
|
+
joins `Runner`'s `around_evaluation:` hook to a telemetry backend's trace
|
|
26
|
+
scope, so a report row links back to the conversation behind it. A `Judge`
|
|
27
|
+
block that accepts `kind:` is told whether it is scoring, recommending or
|
|
28
|
+
writing the verdict, instead of matching on the gem's own instruction prose.
|
|
29
|
+
`Agent#generations` replaces the polymorphic join hosts were copying out of a
|
|
30
|
+
private service method.
|
|
31
|
+
|
|
32
|
+
Upgrading: no migration, and nothing that already worked changes. The judge
|
|
33
|
+
keyword reaches only a block that asks for it, so existing judges are
|
|
34
|
+
untouched; `Evaluation#replace_scenarios!` keeps `:destroy` as its default.
|
|
35
|
+
Adapters should drop any `ActionAgent.record_usage` call of their own, which
|
|
36
|
+
now double-counts, and any provider allow-list check of their own, which is
|
|
37
|
+
now dead code.
|
|
38
|
+
|
|
39
|
+
### Added
|
|
40
|
+
|
|
41
|
+
- The prompt span records how large the tool schemas actually are, as
|
|
42
|
+
`prompt.input.tools.tokens`, `prompt.input.mcp_tools.tokens`,
|
|
43
|
+
`prompt.input.instructions.tokens` and `prompt.input.messages.tokens`. The
|
|
44
|
+
transcript's size is measured before the span trims the history to the turns
|
|
45
|
+
that fit, the others before their content is clipped. The content attributes
|
|
46
|
+
beside them are previews clipped for storage — and on the SDK path the tool
|
|
47
|
+
attribute is a roster of names and parameter keys, several times smaller than
|
|
48
|
+
the schema the model is sent — so a reader that sized the context from one
|
|
49
|
+
understated tool pressure badly.
|
|
50
|
+
- MCP tool schemas are attributed apart from the toolbox's, so the context meter
|
|
51
|
+
can name which half fills the window.
|
|
52
|
+
- `Evaluation#replace_scenarios!` takes `on_removed:` — `:destroy` (the
|
|
53
|
+
default, unchanged) or `:disable`, which keeps a scenario the suite no
|
|
54
|
+
longer names as `enabled: false` so earlier runs' results still resolve.
|
|
55
|
+
- **An evaluation's traces link back to the result that caused them.**
|
|
56
|
+
`ActiveAgent::Evals::Correlation` joins two APIs the module already had but
|
|
57
|
+
never connected: `Runner`'s `around_evaluation:` hook and its `metadata:`
|
|
58
|
+
run identity, and a telemetry backend's per-block agent scope. A run mints a
|
|
59
|
+
`run_id`, each evaluation a `result_id`, and both ride every trace opened
|
|
60
|
+
inside them as `eval.`-prefixed attributes; the trace ids travel the other
|
|
61
|
+
way onto `result.replay.metadata` — `trace_id` for the replay,
|
|
62
|
+
`judge_trace_ids` for the judge calls that graded it, with a run-level
|
|
63
|
+
verdict landing on the run metadata the Report carries rather than on
|
|
64
|
+
whichever result was evaluated last. The tracer is injected, so the module
|
|
65
|
+
takes on no telemetry dependency and `require "active_agent/evals"` still
|
|
66
|
+
loads on its own. Hand the object to `Runner.new(around_evaluation:)`
|
|
67
|
+
directly; a plain lambda there keeps working unchanged.
|
|
68
|
+
- A `Judge` block that accepts `kind:` is told which of the judge's three calls
|
|
69
|
+
it is serving — `:score`, `:recommend` or `:verdict` — so a host can trace,
|
|
70
|
+
budget or model them separately. Previously the only signal was the
|
|
71
|
+
`instructions` string, so hosts matched against the gem's own
|
|
72
|
+
`RECOMMEND_INSTRUCTIONS` / `VERDICT_INSTRUCTIONS` constants; rewording one
|
|
73
|
+
then sent every such host quietly down its `else` branch, mislabelling traces
|
|
74
|
+
rather than failing. The keyword reaches only a block that names it or
|
|
75
|
+
collects `**`, so judges taking `instructions:` and `prompt:` are unaffected.
|
|
76
|
+
(#462)
|
|
77
|
+
- `Agent#generations` reads the generations recorded against an agent, with
|
|
78
|
+
`Agent#agent_contexts` beside it. Generations hang off `AgentContext`
|
|
79
|
+
polymorphically, so reaching them meant hand-writing that join — the engine
|
|
80
|
+
did it itself in a private service method a host could not reuse, which now
|
|
81
|
+
uses the association instead. Destroying an agent still leaves its contexts
|
|
82
|
+
alone, as it always has. (#464)
|
|
83
|
+
|
|
84
|
+
### Fixed
|
|
85
|
+
|
|
86
|
+
- The dashboard's context meter divides the provider's own `prompt_tokens`
|
|
87
|
+
among its segments instead of subtracting its estimates from it. Charging the
|
|
88
|
+
difference to one segment made "Messages" absorb the whole approximation
|
|
89
|
+
error, so a trace with dense JSON tool schemas read as a large message history
|
|
90
|
+
that was never sent. The transcript is one of the divided segments: dividing
|
|
91
|
+
only the rest would hand its share to the segments that remained, so a long
|
|
92
|
+
conversation reported an enormous system prompt and no history at all.
|
|
93
|
+
- A host that supplies a `scenario_evaluation_adapter_resolver` now has its
|
|
94
|
+
replays metered as executions, one per scenario x model, the same unit the
|
|
95
|
+
default path records. Previously an adapted replay was counted only if the
|
|
96
|
+
host remembered to call `ActionAgent.record_usage` itself.
|
|
97
|
+
- A scenario evaluation whose selected models name a provider the agent cannot
|
|
98
|
+
serve fails with `ArgumentError` before any replay runs. A spec handed back
|
|
99
|
+
as a Hash naming both `provider` and `model` bypassed the `providers:`
|
|
100
|
+
allow-list, so the run reached the replay with a provider nothing serves.
|
|
101
|
+
## [1.6.2] - 2026-09-16
|
|
102
|
+
|
|
103
|
+
Releases `activeagent` and `actionagent` 1.6.2 from one tag.
|
|
104
|
+
|
|
105
|
+
Agents gain releases: a digest of everything the model is given, cut on
|
|
106
|
+
deploy and pinned to every trace, run and evaluation run, so a score is a
|
|
107
|
+
statement about a specific release and a regression is attributable to the
|
|
108
|
+
change that caused it. Around it, five dashboard fixes: an evaluation
|
|
109
|
+
created on MySQL can be run, the Tools tab reads the same `agent.tools` the
|
|
110
|
+
runner does, a container-valued query parameter is coerced instead of
|
|
111
|
+
raising, a recording's detail response no longer carries the visitor's
|
|
112
|
+
cookies and web storage, and the MCP endpoint answers an unsupported method with 405
|
|
113
|
+
instead of the dashboard page. `sign_in_path` and `sign_out_path` are now
|
|
114
|
+
documented.
|
|
115
|
+
|
|
116
|
+
Upgrading: the install generator emits a new `add_agent_releases` migration
|
|
117
|
+
(guarded column by column); run it. Cutting a release is
|
|
118
|
+
`rake action_agent:agents:release[REVISION]` in the deploy.
|
|
119
|
+
|
|
120
|
+
### Added
|
|
121
|
+
|
|
122
|
+
- **Agents have releases, and every trace, run and evaluation says which one
|
|
123
|
+
it ran under.** `ActiveAgent::Release` gives each agent class a digest of
|
|
124
|
+
what the model is given — provider and model, generation options minus
|
|
125
|
+
credentials, the actions, the prompt templates on disk, and the tools and
|
|
126
|
+
delegations it declares — so two deploys of the same agent share a digest
|
|
127
|
+
and any change to those inputs is a new one, with no number to bump.
|
|
128
|
+
`ActiveAgent::Release.revision` carries the deploy alongside (a git SHA;
|
|
129
|
+
read from `SERVICE_VERSION`, `GIT_SHA`, `KAMAL_VERSION` and friends when
|
|
130
|
+
not set). The instrumentation stamps `agent.version` and `agent.revision`
|
|
131
|
+
on every generation's root span, and every trace gets `service.version`.
|
|
132
|
+
In the dashboard, `rake action_agent:agents:release[REVISION]` cuts an
|
|
133
|
+
`AgentVersion` for each agent whose code changed since the last release —
|
|
134
|
+
idempotent, so it belongs in the deploy — `rake action_agent:agents:versions`
|
|
135
|
+
lists them, and `Agent#record_release!` is the call behind both for a host
|
|
136
|
+
that syncs agents its own way. Traces are pinned to the release their root
|
|
137
|
+
span names, runs and evaluation runs to the version current when they
|
|
138
|
+
started (`agent_version_id` on all three; the install generator emits the
|
|
139
|
+
migration). A version's JSON carries `release`, `release_digest` and
|
|
140
|
+
`revision`, so the Versions tab tells a deploy from an edit. For that to
|
|
141
|
+
reach a host's own agents, a trace from a class the host mirrors into the
|
|
142
|
+
dashboard is now attributed to that mirror — the registrar matched only on
|
|
143
|
+
service, class *and* action, so every code-path trace registered an
|
|
144
|
+
observed per-action twin beside the synced record and could never be
|
|
145
|
+
pinned to its release.
|
|
146
|
+
|
|
147
|
+
### Fixed
|
|
148
|
+
|
|
149
|
+
- **An evaluation created on MySQL can be run.** MySQL cannot give a JSON
|
|
150
|
+
column a default, so an evaluation saved there without `config` read it
|
|
151
|
+
back as `nil`, and `compare_models` raised before the runner did anything
|
|
152
|
+
else. `config` and `criteria` now read as the empty value their column
|
|
153
|
+
default supplies on other databases. (#417)
|
|
154
|
+
- **The Tools tab now says which schema tools an agent is offered, and
|
|
155
|
+
lets you change it.** The editor listed every schema tool as enabled and
|
|
156
|
+
read-only whatever `agent.tools` held — *"a checkbox that cannot add or
|
|
157
|
+
remove the tool is a control that changes nothing"* — while evaluations,
|
|
158
|
+
dashboard runs and the MCP facade offered exactly what that column named.
|
|
159
|
+
An agent whose roster had been emptied over the API ran a suite with no
|
|
160
|
+
tools (1/8, `expected tool not called ×6`) under a tab reading "12
|
|
161
|
+
enabled". A schema tool's row now reads the roster and is switchable, and
|
|
162
|
+
every schema tool the host declares has a row, off unless the roster names
|
|
163
|
+
it — any agent may enable any of them, and a tool switched off has to keep
|
|
164
|
+
its row to be switched back on. A tool the agent class declares in code is
|
|
165
|
+
still reported rather than selected: the class offers it, and no checkbox
|
|
166
|
+
could change that.
|
|
167
|
+
- **A container-valued query parameter no longer 500s the dashboard API.**
|
|
168
|
+
`page`, `per_page`, `days`, `minutes`, `limit` and `after_sequence` were
|
|
169
|
+
read with `to_i`, which neither an Array (`minutes[]=1&minutes[]=2`) nor a
|
|
170
|
+
nested object (`page[x]=1`) answers. `Api::BaseController` now coerces
|
|
171
|
+
them: a multi-valued parameter means its first value, a nested object falls
|
|
172
|
+
back to the default, and the clamps that bounded the number still apply.
|
|
173
|
+
`sandboxes#compare` answers a `providers` value that is not a list of
|
|
174
|
+
names with a 400 instead of a `NoMethodError`.
|
|
175
|
+
- **A session recording's `show` no longer returns the visitor's cookies and
|
|
176
|
+
web storage.** Every other read path redacted the handoff state, but the
|
|
177
|
+
detail response carried `cookies`, `session_storage` and `local_storage`
|
|
178
|
+
unscrubbed, both as its own key and nested inside `metadata`. Both are now
|
|
179
|
+
stripped; only `#handoff` returns them, to the recording's owner. (#456)
|
|
180
|
+
|
|
181
|
+
## [1.6.1] - 2026-09-16
|
|
182
|
+
|
|
183
|
+
Releases `activeagent` and `actionagent` 1.6.1 from one tag.
|
|
184
|
+
|
|
185
|
+
A patch for two defects that share a failure mode: each one turns a broken
|
|
186
|
+
run into a plausible-looking success rather than an error. A date filter that
|
|
187
|
+
matched nothing reported zero instead of raising, and an agent reported that
|
|
188
|
+
zero as fact; telemetry that was enabled but never instrumented wrote no
|
|
189
|
+
traces while every configuration signal read healthy. Neither surfaced in a
|
|
190
|
+
test suite, because neither produces a failure — only a confident wrong
|
|
191
|
+
answer and an empty table.
|
|
192
|
+
|
|
193
|
+
No new public surface and no behaviour change for anything that was already
|
|
194
|
+
working, so a patch under semver. Suites that filter on a date column will
|
|
195
|
+
report different — correct — numbers after upgrading; read the first run as a
|
|
196
|
+
corrected baseline.
|
|
197
|
+
|
|
198
|
+
### Fixed
|
|
199
|
+
|
|
200
|
+
- **A range filter on a `SchemaTools` column no longer matches nothing and
|
|
201
|
+
reports zero.** `permitted_filters!` validated the column against the
|
|
202
|
+
allowlist but passed the value through untouched, so a range hash reached
|
|
203
|
+
`where` unrecognized and Rails compiled `where(due_date: {"before" => x})`
|
|
204
|
+
to `due_date = NULL` — a predicate that matches no row. The tool returned
|
|
205
|
+
`{count: 0}` with no error and the model read it as a truthful empty
|
|
206
|
+
answer: "0 overdue tickets" against a database holding four. Equality
|
|
207
|
+
filters were unaffected, which is why this went unnoticed. Comparisons are
|
|
208
|
+
now built through Arel with the column's own type cast, under the operators
|
|
209
|
+
`before`, `after`, `lt`, `lte`, `gt`, `gte`, `on_or_before` and
|
|
210
|
+
`on_or_after`; two bounds may be given together to express a window; and an
|
|
211
|
+
operator outside that set raises `UnpermittedAttribute` rather than
|
|
212
|
+
returning zero, consistent with how an undeclared column is already
|
|
213
|
+
rejected. Ranges are offered for date, datetime, time and numeric columns
|
|
214
|
+
only — a lexical `>` on a name column answers a question nobody asked.
|
|
215
|
+
- **A range filter is now discoverable.** `filter_properties` described a date
|
|
216
|
+
column as a bare `{type: "string", format: "date"}`, so the tool surface
|
|
217
|
+
could not express "before today" at all and a model asking the question
|
|
218
|
+
correctly still had no way to ask it. Comparable columns are now offered as
|
|
219
|
+
`anyOf: [scalar, range object]`, with the operator roster in the schema.
|
|
220
|
+
- **Telemetry enabled from a host app's initializer now installs
|
|
221
|
+
instrumentation.** The railtie prepended `GenerationInstrumentation` only
|
|
222
|
+
when `Telemetry.enabled?` was already true as railties ran — before
|
|
223
|
+
`config/initializers/*.rb`. An app that configures telemetry in its own
|
|
224
|
+
initializer, which is what the documentation shows, was therefore never
|
|
225
|
+
instrumented: `enabled?` answered true, `local_storage` was on, the trace
|
|
226
|
+
model resolved and the store lambda worked when called directly, and no
|
|
227
|
+
generation ever produced a span to store. `configure` now installs as well
|
|
228
|
+
when the resulting configuration is enabled; `instrument_telemetry!` is
|
|
229
|
+
idempotent, so the railtie path and the configure path cannot
|
|
230
|
+
double-prepend and initializer order stops mattering.
|
|
231
|
+
|
|
10
232
|
## [1.6.0] - 2026-09-14
|
|
11
233
|
|
|
12
234
|
Releases `activeagent` and `actionagent` 1.6.0 from one tag.
|
|
@@ -142,7 +364,6 @@ still correct: 1.6.0 satisfies it.
|
|
|
142
364
|
hash naming both `provider` and `model` is now rebuilt as it was; a bare
|
|
143
365
|
label is still parsed. The dashboard's "re-run" of a saved selection is
|
|
144
366
|
the path this fixes.
|
|
145
|
-
|
|
146
367
|
## [1.5.2] - 2026-09-11
|
|
147
368
|
|
|
148
369
|
Releases `activeagent` and `actionagent` 1.5.2 from one tag.
|
data/lib/active_agent/base.rb
CHANGED
|
@@ -12,6 +12,7 @@ require "active_agent/concerns/parameterized"
|
|
|
12
12
|
require "active_agent/concerns/preview"
|
|
13
13
|
require "active_agent/concerns/provider"
|
|
14
14
|
require "active_agent/concerns/queueing"
|
|
15
|
+
require "active_agent/concerns/release"
|
|
15
16
|
require "active_agent/concerns/rescue"
|
|
16
17
|
require "active_agent/concerns/streaming"
|
|
17
18
|
require "active_agent/concerns/tooling"
|
|
@@ -49,6 +50,7 @@ module ActiveAgent
|
|
|
49
50
|
include Parameterized
|
|
50
51
|
include Provider
|
|
51
52
|
include Queueing
|
|
53
|
+
include Release
|
|
52
54
|
include Rescue
|
|
53
55
|
include Streaming
|
|
54
56
|
include Tooling
|
|
@@ -0,0 +1,171 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require "digest"
|
|
4
|
+
require "json"
|
|
5
|
+
|
|
6
|
+
module ActiveAgent
|
|
7
|
+
# A release of an agent is what the model is given: the provider and model,
|
|
8
|
+
# the generation options, the prompt templates on disk, the actions and the
|
|
9
|
+
# tools the class declares. {ClassMethods#release_digest} names that
|
|
10
|
+
# deterministically, so two deploys that ship the same agent share a digest
|
|
11
|
+
# and a change to any of those inputs yields a new one — without anyone
|
|
12
|
+
# bumping a number by hand.
|
|
13
|
+
#
|
|
14
|
+
# The digest is what telemetry stamps on every generation (`agent.version`)
|
|
15
|
+
# and what a dashboard cuts an AgentVersion from on deploy, so a trace, a
|
|
16
|
+
# run and an evaluation can all say which release of the agent produced
|
|
17
|
+
# them. {Release.revision} carries the deploy itself (a git SHA or a release
|
|
18
|
+
# label) alongside, when the host knows it.
|
|
19
|
+
#
|
|
20
|
+
# Class-level and memoized: in development a reload replaces the class, so
|
|
21
|
+
# the next reference recomputes it.
|
|
22
|
+
module Release
|
|
23
|
+
extend ActiveSupport::Concern
|
|
24
|
+
|
|
25
|
+
# Option keys that never belong in a manifest: credentials, and per-call
|
|
26
|
+
# state the class does not own.
|
|
27
|
+
EXCLUDED_OPTION_KEYS = %i[
|
|
28
|
+
api_key access_token secret password token trace_id messages message instructions
|
|
29
|
+
].freeze
|
|
30
|
+
SECRET_KEY_PATTERN = /key|token|secret|password|credential/i
|
|
31
|
+
|
|
32
|
+
# How the deploy identifies itself, when it does. A host sets
|
|
33
|
+
# `ActiveAgent::Release.revision = ENV["GIT_SHA"]` (or a proc) from an
|
|
34
|
+
# initializer; otherwise the conventional deploy variables are read.
|
|
35
|
+
class << self
|
|
36
|
+
attr_writer :revision
|
|
37
|
+
|
|
38
|
+
# @return [String, nil]
|
|
39
|
+
def revision
|
|
40
|
+
value = @revision.respond_to?(:call) ? @revision.call : @revision
|
|
41
|
+
value = value.presence || ENV.values_at("SERVICE_VERSION", "GIT_SHA", "KAMAL_VERSION", "SOURCE_VERSION", "HEROKU_SLUG_COMMIT").find(&:present?)
|
|
42
|
+
value&.to_s
|
|
43
|
+
end
|
|
44
|
+
|
|
45
|
+
# Canonical JSON: sorted keys at every level, so the digest does not
|
|
46
|
+
# depend on the order anything was declared in.
|
|
47
|
+
# @api private
|
|
48
|
+
def canonical(value)
|
|
49
|
+
case value
|
|
50
|
+
when Hash then value.map { |k, v| [ k.to_s, canonical(v) ] }.sort_by(&:first).to_h
|
|
51
|
+
when Array then value.map { |v| canonical(v) }
|
|
52
|
+
when Symbol then value.to_s
|
|
53
|
+
else value
|
|
54
|
+
end
|
|
55
|
+
end
|
|
56
|
+
end
|
|
57
|
+
|
|
58
|
+
class_methods do
|
|
59
|
+
# Everything about this class that shapes a generation, as data.
|
|
60
|
+
#
|
|
61
|
+
# @return [Hash]
|
|
62
|
+
def release_manifest
|
|
63
|
+
@release_manifest ||= Release.canonical(
|
|
64
|
+
agent: name,
|
|
65
|
+
provider: release_provider,
|
|
66
|
+
model: prompt_options&.dig(:model),
|
|
67
|
+
options: release_options,
|
|
68
|
+
actions: release_actions,
|
|
69
|
+
templates: release_templates,
|
|
70
|
+
tools: release_tools,
|
|
71
|
+
delegations: release_delegations
|
|
72
|
+
)
|
|
73
|
+
end
|
|
74
|
+
|
|
75
|
+
# A short, stable identifier for {#release_manifest}: the first twelve
|
|
76
|
+
# hex characters of its SHA-256.
|
|
77
|
+
#
|
|
78
|
+
# @return [String]
|
|
79
|
+
def release_digest
|
|
80
|
+
@release_digest ||= Digest::SHA256.hexdigest(JSON.generate(release_manifest))[0, 12]
|
|
81
|
+
end
|
|
82
|
+
|
|
83
|
+
# Forgets the memoized manifest and digest — for a host that edits
|
|
84
|
+
# templates at runtime, and for tests.
|
|
85
|
+
# @return [void]
|
|
86
|
+
def reset_release!
|
|
87
|
+
@release_manifest = nil
|
|
88
|
+
@release_digest = nil
|
|
89
|
+
end
|
|
90
|
+
|
|
91
|
+
private
|
|
92
|
+
|
|
93
|
+
def release_provider
|
|
94
|
+
provider = respond_to?(:prompt_provider) ? prompt_provider : nil
|
|
95
|
+
(provider || prompt_options&.dig(:service))&.to_s
|
|
96
|
+
end
|
|
97
|
+
|
|
98
|
+
# Generation options minus credentials and per-call state. Nested
|
|
99
|
+
# hashes are walked so a token under `options: { headers: … }` is
|
|
100
|
+
# dropped too.
|
|
101
|
+
def release_options
|
|
102
|
+
strip_secrets((prompt_options || {}).except(*EXCLUDED_OPTION_KEYS, :model, :service))
|
|
103
|
+
end
|
|
104
|
+
|
|
105
|
+
def strip_secrets(value)
|
|
106
|
+
case value
|
|
107
|
+
when Hash
|
|
108
|
+
value.each_with_object({}) do |(key, inner), kept|
|
|
109
|
+
next if EXCLUDED_OPTION_KEYS.include?(key.to_sym) || key.to_s.match?(SECRET_KEY_PATTERN)
|
|
110
|
+
|
|
111
|
+
kept[key] = strip_secrets(inner)
|
|
112
|
+
end
|
|
113
|
+
when Array then value.map { |inner| strip_secrets(inner) }
|
|
114
|
+
else value
|
|
115
|
+
end
|
|
116
|
+
end
|
|
117
|
+
|
|
118
|
+
# The public actions — the prompts a caller can invoke.
|
|
119
|
+
def release_actions
|
|
120
|
+
respond_to?(:action_methods) ? action_methods.to_a.sort : []
|
|
121
|
+
end
|
|
122
|
+
|
|
123
|
+
# Every template file under this agent's view prefixes, keyed by its
|
|
124
|
+
# path relative to the view root, with a digest of its contents. The
|
|
125
|
+
# prefixes mirror View#_prefixes without an action: `app/views/<agent>/`
|
|
126
|
+
# and `app/views/agents/<agent without suffix>/`.
|
|
127
|
+
def release_templates
|
|
128
|
+
return {} if anonymous? || !respond_to?(:view_paths)
|
|
129
|
+
|
|
130
|
+
base = name.underscore
|
|
131
|
+
prefixes = [ base, "agents/#{base.delete_suffix("_agent")}" ]
|
|
132
|
+
roots = Array(view_paths).map { |path| path.respond_to?(:to_path) ? path.to_path : path.to_s }
|
|
133
|
+
|
|
134
|
+
roots.each_with_object({}) do |root, templates|
|
|
135
|
+
prefixes.each do |prefix|
|
|
136
|
+
Dir.glob(File.join(root, prefix, "**", "*")).sort.each do |file|
|
|
137
|
+
next unless File.file?(file)
|
|
138
|
+
|
|
139
|
+
relative = file.delete_prefix("#{root}/")
|
|
140
|
+
templates[relative] = Digest::SHA256.hexdigest(File.binread(file))[0, 12]
|
|
141
|
+
end
|
|
142
|
+
end
|
|
143
|
+
end
|
|
144
|
+
end
|
|
145
|
+
|
|
146
|
+
# Tool definitions the class declares itself (a host convention such as
|
|
147
|
+
# schema-derived rosters), reduced to what identifies them.
|
|
148
|
+
def release_tools
|
|
149
|
+
return [] unless respond_to?(:tool_definitions)
|
|
150
|
+
|
|
151
|
+
Array(tool_definitions).map do |definition|
|
|
152
|
+
next definition.to_s unless definition.respond_to?(:to_h)
|
|
153
|
+
|
|
154
|
+
hash = definition.to_h
|
|
155
|
+
{
|
|
156
|
+
name: (hash[:name] || hash["name"]).to_s,
|
|
157
|
+
description: (hash[:description] || hash["description"]).to_s,
|
|
158
|
+
parameters: hash[:parameters] || hash["parameters"]
|
|
159
|
+
}
|
|
160
|
+
end.sort_by { |tool| tool.is_a?(Hash) ? tool[:name] : tool }
|
|
161
|
+
end
|
|
162
|
+
|
|
163
|
+
# The delegations this class declares, by tool name.
|
|
164
|
+
def release_delegations
|
|
165
|
+
return [] unless respond_to?(:delegations)
|
|
166
|
+
|
|
167
|
+
Array(delegations).map { |tool_name, _definition| tool_name.to_s }.sort
|
|
168
|
+
end
|
|
169
|
+
end
|
|
170
|
+
end
|
|
171
|
+
end
|
|
@@ -0,0 +1,178 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require "securerandom"
|
|
4
|
+
require "active_support/isolated_execution_state"
|
|
5
|
+
|
|
6
|
+
module ActiveAgent
|
|
7
|
+
module Evals
|
|
8
|
+
# Used to tie the traces an evaluation produces back to the run and the
|
|
9
|
+
# result that caused them, so a report row links to the exact conversation
|
|
10
|
+
# behind it.
|
|
11
|
+
#
|
|
12
|
+
# A run mints a `run_id`, each evaluation mints a `result_id`, and both ride
|
|
13
|
+
# every trace opened inside them as `eval.`-prefixed attributes. The trace
|
|
14
|
+
# ids travel the other way: a replay's trace id lands on
|
|
15
|
+
# `result.replay.metadata["trace_id"]`, and every judge call made while
|
|
16
|
+
# scoring that result appends to its `"judge_trace_ids"`. The verdict — a
|
|
17
|
+
# judge call made outside any evaluation — appends to the run metadata
|
|
18
|
+
# instead, which is the same Hash a Report carries as its `metadata`.
|
|
19
|
+
#
|
|
20
|
+
# A tracer is `(name, action:, attributes:, on_trace:) { ... }`: it opens a
|
|
21
|
+
# trace named for the agent, and calls `on_trace` with something answering
|
|
22
|
+
# to `#trace_id` once the trace is known.
|
|
23
|
+
#
|
|
24
|
+
# correlation = ActiveAgent::Evals::Correlation.new(
|
|
25
|
+
# agent_name: "SupportAgent",
|
|
26
|
+
# judge_name: "SupportAgentJudge",
|
|
27
|
+
# tracer: ->(name, action:, attributes:, on_trace:, &block) {
|
|
28
|
+
# MyTelemetry.with_agent(name, action: action, attributes: attributes,
|
|
29
|
+
# on_trace: on_trace, synchronous: true, &block)
|
|
30
|
+
# }
|
|
31
|
+
# )
|
|
32
|
+
#
|
|
33
|
+
# correlation.with_run("suite" => "support") do |metadata|
|
|
34
|
+
# Runner.new(
|
|
35
|
+
# scenarios: scenarios, models: models, metadata: metadata,
|
|
36
|
+
# replay: ->(scenario, spec) { correlation.replay { agent.run(scenario.prompt) } },
|
|
37
|
+
# judge: Judge.new(label: "judge-model") { |instructions:, prompt:|
|
|
38
|
+
# correlation.judge("score") { chat.with_instructions(instructions).ask(prompt).content }
|
|
39
|
+
# },
|
|
40
|
+
# around_evaluation: correlation
|
|
41
|
+
# ).call
|
|
42
|
+
# end
|
|
43
|
+
#
|
|
44
|
+
# Without a tracer the correlation still mints ids and merges metadata. The
|
|
45
|
+
# blocks run untraced.
|
|
46
|
+
class Correlation
|
|
47
|
+
STATE_KEY = :active_agent_evals_correlation
|
|
48
|
+
|
|
49
|
+
# DEFAULT_TRACE_KEYS names the correlation metadata that rides a trace as
|
|
50
|
+
# `eval.`-prefixed attributes. Anything else a caller puts in the run
|
|
51
|
+
# metadata (a tenant, a role) stays on the report but off the traces.
|
|
52
|
+
DEFAULT_TRACE_KEYS = %w[run_id result_id suite scenario_key model_label model provider].freeze
|
|
53
|
+
|
|
54
|
+
attr_reader :agent_name, :judge_name, :trace_keys
|
|
55
|
+
|
|
56
|
+
# @param agent_name [String] the trace name for the agent under evaluation
|
|
57
|
+
# @param judge_name [String] the trace name for judge traffic, kept distinct
|
|
58
|
+
# so grading calls do not read as the agent's own traffic
|
|
59
|
+
# @param tracer [#call, nil] `(name, action:, attributes:, on_trace:) { ... }`;
|
|
60
|
+
# nil runs every block untraced
|
|
61
|
+
# @param replay_action [String] the action name recorded for a replay trace
|
|
62
|
+
# @param trace_keys [Array<String>] which correlation keys become attributes
|
|
63
|
+
def initialize(agent_name:, judge_name: "#{agent_name}Judge", tracer: nil, replay_action: "eval",
|
|
64
|
+
trace_keys: DEFAULT_TRACE_KEYS)
|
|
65
|
+
@agent_name = agent_name
|
|
66
|
+
@judge_name = judge_name
|
|
67
|
+
@tracer = tracer
|
|
68
|
+
@replay_action = replay_action
|
|
69
|
+
@trace_keys = trace_keys.map(&:to_s)
|
|
70
|
+
end
|
|
71
|
+
|
|
72
|
+
# Opens a run. Mints `run_id` unless `metadata` carries one, and yields the
|
|
73
|
+
# metadata hash the traces will be correlated against — pass that same hash
|
|
74
|
+
# to `Runner.new(metadata:)` so the Report carries the run's identity and
|
|
75
|
+
# collects the verdict's trace id.
|
|
76
|
+
#
|
|
77
|
+
# The hash yielded is the caller's own, mutated in place, so a run
|
|
78
|
+
# reopened around a later verdict accumulates onto the metadata a Report
|
|
79
|
+
# already carries.
|
|
80
|
+
#
|
|
81
|
+
# @param metadata [Hash] opaque run metadata. A non-Hash is coerced with
|
|
82
|
+
# `#to_h`, and non-String keys are stringified.
|
|
83
|
+
# @yieldparam metadata [Hash]
|
|
84
|
+
def with_run(metadata = {})
|
|
85
|
+
run = metadata.is_a?(Hash) ? metadata : metadata.to_h
|
|
86
|
+
run.transform_keys!(&:to_s) unless run.keys.all?(String)
|
|
87
|
+
run["run_id"] ||= SecureRandom.uuid
|
|
88
|
+
with_context({ run: run, result: nil }) { yield run }
|
|
89
|
+
end
|
|
90
|
+
|
|
91
|
+
# Wraps one evaluation, in the shape `Runner.new(around_evaluation:)` calls:
|
|
92
|
+
# `(scenario, spec) { ... } → Result`. Mints a `result_id`, merges the
|
|
93
|
+
# correlation onto the Result's replay metadata, and returns the Result.
|
|
94
|
+
def around_evaluation(scenario, spec)
|
|
95
|
+
result_metadata = {
|
|
96
|
+
"run_id" => run_metadata["run_id"],
|
|
97
|
+
"result_id" => SecureRandom.uuid,
|
|
98
|
+
"scenario_key" => scenario.key,
|
|
99
|
+
"model_label" => spec.label,
|
|
100
|
+
"model" => spec.model,
|
|
101
|
+
"provider" => spec.provider
|
|
102
|
+
}.compact
|
|
103
|
+
|
|
104
|
+
with_context(run: run_metadata, result: result_metadata) do
|
|
105
|
+
yield.tap { |result| result.replay.metadata.merge!(result_metadata) }
|
|
106
|
+
end
|
|
107
|
+
end
|
|
108
|
+
|
|
109
|
+
# Delegates to `around_evaluation`, so the object satisfies
|
|
110
|
+
# `Runner.new(around_evaluation:)` directly.
|
|
111
|
+
def call(scenario, spec, &)
|
|
112
|
+
around_evaluation(scenario, spec, &)
|
|
113
|
+
end
|
|
114
|
+
|
|
115
|
+
# Traces one replay of the agent under evaluation. The trace id lands on
|
|
116
|
+
# the current result's metadata, so a report row links to the conversation.
|
|
117
|
+
def replay(action = @replay_action, &)
|
|
118
|
+
trace(@agent_name, action, judge: false, &)
|
|
119
|
+
end
|
|
120
|
+
|
|
121
|
+
# Traces one judge call. Appends to the current result's `judge_trace_ids`,
|
|
122
|
+
# or the run's when no evaluation is open (the verdict).
|
|
123
|
+
def judge(action = "score", &)
|
|
124
|
+
trace(@judge_name, action, judge: true, &)
|
|
125
|
+
end
|
|
126
|
+
|
|
127
|
+
# The correlation metadata in scope, or nil outside a run. A result's
|
|
128
|
+
# values win over the run's.
|
|
129
|
+
def current
|
|
130
|
+
context = ActiveSupport::IsolatedExecutionState[STATE_KEY]
|
|
131
|
+
return nil unless context
|
|
132
|
+
|
|
133
|
+
context.fetch(:run, {}).merge(context[:result] || {})
|
|
134
|
+
end
|
|
135
|
+
|
|
136
|
+
private
|
|
137
|
+
|
|
138
|
+
def run_metadata
|
|
139
|
+
context = ActiveSupport::IsolatedExecutionState[STATE_KEY]
|
|
140
|
+
context&.fetch(:run, nil) || {}
|
|
141
|
+
end
|
|
142
|
+
|
|
143
|
+
def with_context(context)
|
|
144
|
+
previous = ActiveSupport::IsolatedExecutionState[STATE_KEY]
|
|
145
|
+
ActiveSupport::IsolatedExecutionState[STATE_KEY] = context
|
|
146
|
+
yield
|
|
147
|
+
ensure
|
|
148
|
+
ActiveSupport::IsolatedExecutionState[STATE_KEY] = previous
|
|
149
|
+
end
|
|
150
|
+
|
|
151
|
+
def trace(name, action, judge:, &block)
|
|
152
|
+
return block.call unless @tracer
|
|
153
|
+
|
|
154
|
+
context = ActiveSupport::IsolatedExecutionState[STATE_KEY] || {}
|
|
155
|
+
correlation = context.fetch(:run, {}).merge(context[:result] || {})
|
|
156
|
+
attributes = correlation.slice(*@trace_keys).transform_keys { |key| "eval.#{key}" }
|
|
157
|
+
# A replay belongs to the evaluation that opened it and nowhere else, so
|
|
158
|
+
# it records no trace id when called outside one. A judge call outside an
|
|
159
|
+
# evaluation is the verdict, which belongs to the run.
|
|
160
|
+
target = judge ? (context[:result] || context[:run]) : context[:result]
|
|
161
|
+
|
|
162
|
+
@tracer.call(name, action: action, attributes: attributes, on_trace: recorder(target, judge: judge), &block)
|
|
163
|
+
end
|
|
164
|
+
|
|
165
|
+
def recorder(target, judge:)
|
|
166
|
+
lambda do |trace|
|
|
167
|
+
next unless target
|
|
168
|
+
|
|
169
|
+
if judge
|
|
170
|
+
(target["judge_trace_ids"] ||= []) << trace.trace_id
|
|
171
|
+
else
|
|
172
|
+
target["trace_id"] = trace.trace_id
|
|
173
|
+
end
|
|
174
|
+
end
|
|
175
|
+
end
|
|
176
|
+
end
|
|
177
|
+
end
|
|
178
|
+
end
|
|
@@ -10,6 +10,21 @@ module ActiveAgent
|
|
|
10
10
|
# RubyLLM.chat(model: "claude-opus-5").with_instructions(instructions).ask(prompt).content
|
|
11
11
|
# end
|
|
12
12
|
#
|
|
13
|
+
# A judge serves three different calls, and a block that accepts `kind:` is
|
|
14
|
+
# told which one it is serving — `:score`, `:recommend` or `:verdict` — so a
|
|
15
|
+
# host can trace them apart, budget them apart, or score with a cheaper model
|
|
16
|
+
# than it writes the verdict with:
|
|
17
|
+
#
|
|
18
|
+
# Judge.new(label: "claude-opus-5") do |instructions:, prompt:, kind:|
|
|
19
|
+
# model = kind == :score ? "claude-haiku-4-5" : "claude-opus-5"
|
|
20
|
+
# RubyLLM.chat(model: model).with_instructions(instructions).ask(prompt).content
|
|
21
|
+
# end
|
|
22
|
+
#
|
|
23
|
+
# The keyword is passed only to a block that names it (or collects `**`), so
|
|
24
|
+
# a two-keyword block written before this is unaffected. Without it the only
|
|
25
|
+
# signal is the `instructions` string, which means matching on the gem's own
|
|
26
|
+
# prose — and a reworded constant then mislabels silently instead of failing.
|
|
27
|
+
#
|
|
13
28
|
# Every method returns nil when the judge fails or answers unusably, so an
|
|
14
29
|
# evaluation degrades to rule scoring rather than aborting.
|
|
15
30
|
class Judge
|
|
@@ -23,6 +38,8 @@ module ActiveAgent
|
|
|
23
38
|
# @param label [String] how reports name the judge (usually its model)
|
|
24
39
|
# @yieldparam instructions [String] the system prompt
|
|
25
40
|
# @yieldparam prompt [String] the user prompt
|
|
41
|
+
# @yieldparam kind [Symbol] which call this is — `:score`, `:recommend` or
|
|
42
|
+
# `:verdict`. Passed only to a block that accepts it.
|
|
26
43
|
# @yieldreturn [String] the completion text
|
|
27
44
|
# How much of a scenario's notes the judge reads. Where a suite's notes
|
|
28
45
|
# are its grading rubric, a "Must not…" clause tends to come last, and a
|
|
@@ -41,7 +58,7 @@ module ActiveAgent
|
|
|
41
58
|
return nil if answer.blank?
|
|
42
59
|
|
|
43
60
|
guidance = criterion.dig("config", "prompt").presence || criterion["key"].to_s.humanize
|
|
44
|
-
parse_score(ask(SCORE_INSTRUCTIONS, <<~PROMPT))
|
|
61
|
+
parse_score(ask(SCORE_INSTRUCTIONS, <<~PROMPT, :score))
|
|
45
62
|
Criterion: #{guidance}
|
|
46
63
|
|
|
47
64
|
The user asked:
|
|
@@ -64,7 +81,7 @@ module ActiveAgent
|
|
|
64
81
|
def score_task(scenario:, answer:)
|
|
65
82
|
return nil if answer.blank?
|
|
66
83
|
|
|
67
|
-
parse_score(ask(SCORE_INSTRUCTIONS, <<~PROMPT))
|
|
84
|
+
parse_score(ask(SCORE_INSTRUCTIONS, <<~PROMPT, :score))
|
|
68
85
|
A user asked an assistant:
|
|
69
86
|
---
|
|
70
87
|
#{scenario.prompt}
|
|
@@ -90,7 +107,7 @@ module ActiveAgent
|
|
|
90
107
|
"- #{call['name']}#{' (errored)' if call['error']}: #{call['arguments'].to_json.truncate(200)}"
|
|
91
108
|
end.join("\n")
|
|
92
109
|
|
|
93
|
-
parsed = parse_object(ask(RECOMMEND_INSTRUCTIONS, <<~PROMPT))
|
|
110
|
+
parsed = parse_object(ask(RECOMMEND_INSTRUCTIONS, <<~PROMPT, :recommend))
|
|
94
111
|
An AI agent failed one evaluation scenario. Recommend the fix.
|
|
95
112
|
|
|
96
113
|
Agent instructions:
|
|
@@ -146,7 +163,7 @@ module ActiveAgent
|
|
|
146
163
|
"#{", faults: #{faults}" if faults.present?}"
|
|
147
164
|
end
|
|
148
165
|
|
|
149
|
-
parsed = parse_object(ask(VERDICT_INSTRUCTIONS, <<~PROMPT))
|
|
166
|
+
parsed = parse_object(ask(VERDICT_INSTRUCTIONS, <<~PROMPT, :verdict))
|
|
150
167
|
An AI agent ran the same scenarios under several models. Its goals:
|
|
151
168
|
---
|
|
152
169
|
#{instructions.to_s.truncate(1_000).presence || '(no instructions configured)'}
|
|
@@ -180,13 +197,34 @@ module ActiveAgent
|
|
|
180
197
|
end
|
|
181
198
|
end
|
|
182
199
|
|
|
183
|
-
def ask(instructions, prompt)
|
|
184
|
-
@generate.call(instructions
|
|
200
|
+
def ask(instructions, prompt, kind)
|
|
201
|
+
@generate.call(**ask_arguments(instructions, prompt, kind)).to_s
|
|
185
202
|
rescue StandardError => e
|
|
186
203
|
warn_failure(e)
|
|
187
204
|
nil
|
|
188
205
|
end
|
|
189
206
|
|
|
207
|
+
# The block signature is public API, and every judge written before `kind:`
|
|
208
|
+
# existed takes exactly `instructions:` and `prompt:` — passing a third
|
|
209
|
+
# keyword to one of those raises ArgumentError, which `ask` would swallow
|
|
210
|
+
# as a judge failure, degrading the run to rule scoring. So the kind goes
|
|
211
|
+
# only to a block that asked for it.
|
|
212
|
+
def ask_arguments(instructions, prompt, kind)
|
|
213
|
+
arguments = { instructions: instructions, prompt: prompt }
|
|
214
|
+
arguments[:kind] = kind if generate_accepts_kind?
|
|
215
|
+
arguments
|
|
216
|
+
end
|
|
217
|
+
|
|
218
|
+
# True for a block naming `kind:` or collecting `**`. Memoized because the
|
|
219
|
+
# answer cannot change for a given judge and `ask` runs per scored result.
|
|
220
|
+
def generate_accepts_kind?
|
|
221
|
+
return @generate_accepts_kind if defined?(@generate_accepts_kind)
|
|
222
|
+
|
|
223
|
+
@generate_accepts_kind = @generate.parameters.any? do |type, name|
|
|
224
|
+
type == :keyrest || (name == :kind && (type == :key || type == :keyreq))
|
|
225
|
+
end
|
|
226
|
+
end
|
|
227
|
+
|
|
190
228
|
def warn_failure(error)
|
|
191
229
|
message = "[ActiveAgent::Evals] judge #{label} failed: #{error.class}: #{error.message}"
|
|
192
230
|
if defined?(Rails) && Rails.respond_to?(:logger) && Rails.logger
|
data/lib/active_agent/evals.rb
CHANGED
|
@@ -25,6 +25,7 @@ require_relative "evals/design_tokens"
|
|
|
25
25
|
require_relative "evals/report_html"
|
|
26
26
|
require_relative "evals/report"
|
|
27
27
|
require_relative "evals/runner"
|
|
28
|
+
require_relative "evals/correlation"
|
|
28
29
|
require_relative "evals/publisher"
|
|
29
30
|
|
|
30
31
|
# Scenario evaluations for agents that answer with tools.
|
|
@@ -40,7 +41,8 @@ require_relative "evals/publisher"
|
|
|
40
41
|
# fell short and what would fix it (Diagnosis, refined by an optional Judge),
|
|
41
42
|
# and rolling everything up per model (Report). Runner ties them together
|
|
42
43
|
# around one callable you supply: given a scenario and a model, run the
|
|
43
|
-
# agent and return a Replay.
|
|
44
|
+
# agent and return a Replay. Correlation is optional plumbing on top: it
|
|
45
|
+
# links the traces a run emits back to the result that caused them.
|
|
44
46
|
#
|
|
45
47
|
# scenarios = ActiveAgent::Evals::ScenarioParser.scenarios(pasted_text)
|
|
46
48
|
# models = ActiveAgent::Evals::ModelSpec.parse_all(%w[gpt-5-mini qwen3:8b], default_provider: "openai")
|
|
@@ -68,6 +68,21 @@ module ActiveAgent
|
|
|
68
68
|
# ask for 10_000; this is what stops that from becoming the prompt.
|
|
69
69
|
MAX_LIMIT = 100
|
|
70
70
|
|
|
71
|
+
# Column types a range comparison is offered for. Strings and booleans
|
|
72
|
+
# are deliberately absent: a lexical `>` on a name column answers a
|
|
73
|
+
# question nobody asked.
|
|
74
|
+
RANGE_FILTERABLE_TYPES = %i[date datetime time integer float decimal].freeze
|
|
75
|
+
|
|
76
|
+
# The comparison operators a range filter may use, mapped to the Arel
|
|
77
|
+
# predicate that builds them. Names are the ones models reach for
|
|
78
|
+
# unprompted (`before`/`after` for dates, `lt`/`gte` for numbers), so a
|
|
79
|
+
# reasonable guess resolves instead of erroring.
|
|
80
|
+
RANGE_OPERATORS = {
|
|
81
|
+
"before" => :lt, "after" => :gt,
|
|
82
|
+
"lt" => :lt, "lte" => :lteq, "gt" => :gt, "gte" => :gteq,
|
|
83
|
+
"on_or_before" => :lteq, "on_or_after" => :gteq
|
|
84
|
+
}.freeze
|
|
85
|
+
|
|
71
86
|
# Raised when a tool call names a column outside the declared allowlists,
|
|
72
87
|
# or is otherwise outside the declared boundary.
|
|
73
88
|
class UnpermittedAttribute < ArgumentError; end
|
|
@@ -300,8 +315,13 @@ module ActiveAgent
|
|
|
300
315
|
|
|
301
316
|
# Validates and normalizes a filter hash against the allowlist.
|
|
302
317
|
#
|
|
318
|
+
# A filter value is normally matched for equality. A Hash value instead
|
|
319
|
+
# declares a range — `{ "before" => "2026-01-01" }`, `{ "gte" => 10 }` —
|
|
320
|
+
# and may carry two bounds at once to express a window.
|
|
321
|
+
#
|
|
303
322
|
# @api private
|
|
304
|
-
# @raise [UnpermittedAttribute] if any key is not declared filterable
|
|
323
|
+
# @raise [UnpermittedAttribute] if any key is not declared filterable,
|
|
324
|
+
# or a range names an operator that does not exist
|
|
305
325
|
def permitted_filters!(arguments)
|
|
306
326
|
filters = arguments.each_with_object({}) do |(key, value), memo|
|
|
307
327
|
next if value.nil?
|
|
@@ -318,6 +338,54 @@ module ActiveAgent
|
|
|
318
338
|
filters
|
|
319
339
|
end
|
|
320
340
|
|
|
341
|
+
# Splits filters into equality pairs and range predicates.
|
|
342
|
+
#
|
|
343
|
+
# Kept separate from {.permitted_filters!} because the two halves are
|
|
344
|
+
# applied differently: equality goes to `where(hash)`, ranges have to be
|
|
345
|
+
# built through Arel.
|
|
346
|
+
#
|
|
347
|
+
# @api private
|
|
348
|
+
# @return [Array(Hash, Array<Arel::Nodes::Node>)]
|
|
349
|
+
def partition_filters!(filters)
|
|
350
|
+
equality = {}
|
|
351
|
+
ranges = []
|
|
352
|
+
|
|
353
|
+
filters.each do |column, value|
|
|
354
|
+
if value.is_a?(Hash)
|
|
355
|
+
ranges.concat(range_predicates!(column, value))
|
|
356
|
+
else
|
|
357
|
+
equality[column] = value
|
|
358
|
+
end
|
|
359
|
+
end
|
|
360
|
+
|
|
361
|
+
[ equality, ranges ]
|
|
362
|
+
end
|
|
363
|
+
|
|
364
|
+
# Builds Arel predicates for one column's range hash.
|
|
365
|
+
#
|
|
366
|
+
# Rails silently turns `where(col: { "before" => x })` into `col = NULL`,
|
|
367
|
+
# which matches nothing and reports zero rather than failing — the worst
|
|
368
|
+
# outcome for an agent, which reads it as a truthful empty answer. So an
|
|
369
|
+
# unknown operator is rejected loudly here instead.
|
|
370
|
+
#
|
|
371
|
+
# @api private
|
|
372
|
+
# @raise [UnpermittedAttribute] on an unknown operator
|
|
373
|
+
def range_predicates!(column, value)
|
|
374
|
+
arel = @model.arel_table[column]
|
|
375
|
+
type = @model.type_for_attribute(column)
|
|
376
|
+
|
|
377
|
+
value.map do |operator, operand|
|
|
378
|
+
predicate = RANGE_OPERATORS[operator.to_s]
|
|
379
|
+
unless predicate
|
|
380
|
+
raise UnpermittedAttribute,
|
|
381
|
+
"`#{operator}` is not a valid comparison for `#{column}`. " \
|
|
382
|
+
"Allowed comparisons: #{RANGE_OPERATORS.keys.join(", ")}"
|
|
383
|
+
end
|
|
384
|
+
|
|
385
|
+
arel.public_send(predicate, type.cast(operand))
|
|
386
|
+
end
|
|
387
|
+
end
|
|
388
|
+
|
|
321
389
|
# Projects a record down to the declared return columns.
|
|
322
390
|
#
|
|
323
391
|
# The projection happens in SQL (+select+) as well as here, but the Ruby
|
|
@@ -386,7 +454,40 @@ module ActiveAgent
|
|
|
386
454
|
)
|
|
387
455
|
properties = schema[:schema][:properties]
|
|
388
456
|
|
|
389
|
-
filterable.index_with
|
|
457
|
+
filterable.index_with do |column|
|
|
458
|
+
scalar = (properties[column] || { type: "string" }).deep_dup
|
|
459
|
+
range_filterable?(column) ? with_range_form(column, scalar) : scalar
|
|
460
|
+
end
|
|
461
|
+
end
|
|
462
|
+
|
|
463
|
+
# Dates, times and numbers are the columns a question like "overdue" or
|
|
464
|
+
# "more than 10" actually needs a comparison on.
|
|
465
|
+
def range_filterable?(column)
|
|
466
|
+
RANGE_FILTERABLE_TYPES.include?(@model.type_for_attribute(column).type)
|
|
467
|
+
end
|
|
468
|
+
|
|
469
|
+
# Offers a column as either a scalar (equality) or a range object.
|
|
470
|
+
#
|
|
471
|
+
# Without this the range form works but is undiscoverable: a model shown
|
|
472
|
+
# only `{type: "string", format: "date"}` has no way to know it may ask
|
|
473
|
+
# for `before`, and answers date questions with an equality match or no
|
|
474
|
+
# filter at all.
|
|
475
|
+
def with_range_form(column, scalar)
|
|
476
|
+
operand = scalar.slice(:type, :format)
|
|
477
|
+
description = scalar[:description]
|
|
478
|
+
|
|
479
|
+
{
|
|
480
|
+
description: [ description, "Accepts an exact value, or a range object such as " \
|
|
481
|
+
"{\"before\": ...} / {\"gte\": ...} (#{RANGE_OPERATORS.keys.join(", ")})." ].compact.join(" "),
|
|
482
|
+
anyOf: [
|
|
483
|
+
scalar.except(:description),
|
|
484
|
+
{
|
|
485
|
+
type: "object",
|
|
486
|
+
properties: RANGE_OPERATORS.keys.index_with { operand.dup },
|
|
487
|
+
additionalProperties: false
|
|
488
|
+
}
|
|
489
|
+
]
|
|
490
|
+
}
|
|
390
491
|
end
|
|
391
492
|
|
|
392
493
|
def resource_name
|
|
@@ -415,10 +516,10 @@ module ActiveAgent
|
|
|
415
516
|
)
|
|
416
517
|
|
|
417
518
|
define_singleton_method(name) do |actor: nil, limit: nil, **arguments|
|
|
418
|
-
|
|
519
|
+
equality, ranges = partition_filters!(permitted_filters!(arguments))
|
|
419
520
|
capped = normalize_limit(limit)
|
|
420
521
|
|
|
421
|
-
relation = relation_for(actor).where(
|
|
522
|
+
relation = ranges.reduce(relation_for(actor).where(equality)) { |rel, p| rel.where(p) }
|
|
422
523
|
# One extra row distinguishes "exactly at the limit" from "more than
|
|
423
524
|
# the limit", without a second COUNT query.
|
|
424
525
|
records = relation.limit(capped + 1).to_a
|
|
@@ -443,9 +544,10 @@ module ActiveAgent
|
|
|
443
544
|
)
|
|
444
545
|
|
|
445
546
|
define_singleton_method(name) do |actor: nil, **arguments|
|
|
446
|
-
|
|
547
|
+
equality, ranges = partition_filters!(permitted_filters!(arguments))
|
|
548
|
+
relation = ranges.reduce(relation_for(actor).where(equality)) { |rel, p| rel.where(p) }
|
|
447
549
|
|
|
448
|
-
{ count:
|
|
550
|
+
{ count: relation.count }
|
|
449
551
|
end
|
|
450
552
|
end
|
|
451
553
|
|
|
@@ -24,6 +24,17 @@ module ActiveAgent
|
|
|
24
24
|
# @return [Boolean] Whether to store traces in the app's own database
|
|
25
25
|
attr_reader :local_storage
|
|
26
26
|
|
|
27
|
+
# The deploy every trace is stamped with (`service.version`). Unset, it
|
|
28
|
+
# is whatever ActiveAgent::Release.revision resolves — a git SHA from
|
|
29
|
+
# the conventional deploy variables — so traces from two deploys of
|
|
30
|
+
# the same service can be told apart without any host configuration.
|
|
31
|
+
attr_writer :service_version
|
|
32
|
+
|
|
33
|
+
def service_version
|
|
34
|
+
value = @service_version.respond_to?(:call) ? @service_version.call : @service_version
|
|
35
|
+
(value.presence || ActiveAgent::Release.revision)&.to_s
|
|
36
|
+
end
|
|
37
|
+
|
|
27
38
|
def initialize
|
|
28
39
|
super
|
|
29
40
|
# The framework predates the shared gem and has always been opt-in;
|
|
@@ -58,6 +58,15 @@ module ActiveAgent
|
|
|
58
58
|
span.set_attribute("agent.action", action_name.to_s)
|
|
59
59
|
span.set_attribute("agent.provider", provider_name)
|
|
60
60
|
span.set_attribute("agent.model", model_name)
|
|
61
|
+
# Which release of the agent ran: the digest of what the model was
|
|
62
|
+
# given (ActiveAgent::Release), so a dashboard can pin this trace
|
|
63
|
+
# to the version it cut on deploy.
|
|
64
|
+
if self.class.respond_to?(:release_digest)
|
|
65
|
+
span.set_attribute("agent.version", self.class.release_digest)
|
|
66
|
+
if (revision = ActiveAgent::Release.revision).present?
|
|
67
|
+
span.set_attribute("agent.revision", revision)
|
|
68
|
+
end
|
|
69
|
+
end
|
|
61
70
|
|
|
62
71
|
# Add prompt span, carrying the prompt contents (instructions +
|
|
63
72
|
# outbound messages) so dashboards can show what was sent.
|
|
@@ -92,7 +101,23 @@ module ActiveAgent
|
|
|
92
101
|
parameters: properties.is_a?(Hash) ? properties.keys : []
|
|
93
102
|
}.compact
|
|
94
103
|
}
|
|
95
|
-
|
|
104
|
+
if roster.any?
|
|
105
|
+
prompt_span.set_attribute("prompt.input.tools", JSON.generate(roster))
|
|
106
|
+
# The roster is a reading aid: names, a truncated description
|
|
107
|
+
# and parameter keys. What the model is sent is the full JSON
|
|
108
|
+
# Schema — types, enums, nested objects, anyOf branches —
|
|
109
|
+
# which for a twelve-tool agent is several times the roster,
|
|
110
|
+
# so a context meter sizing tool pressure from the roster
|
|
111
|
+
# understates it badly. The schemas come from the host and
|
|
112
|
+
# need not be serializable; a size is worth less than the
|
|
113
|
+
# generation it would otherwise take down.
|
|
114
|
+
tools_size = begin
|
|
115
|
+
JSON.generate(tools).length
|
|
116
|
+
rescue StandardError
|
|
117
|
+
nil
|
|
118
|
+
end
|
|
119
|
+
prompt_span.set_attribute("prompt.input.tools.tokens", (tools_size / 4.0).round) if tools_size
|
|
120
|
+
end
|
|
96
121
|
end
|
|
97
122
|
|
|
98
123
|
# prompt_options[:messages] holds the turns a caller passed
|
|
@@ -169,12 +169,15 @@ module ActiveAgent
|
|
|
169
169
|
|
|
170
170
|
# Returns default attributes for all spans.
|
|
171
171
|
def default_attributes
|
|
172
|
-
{
|
|
172
|
+
attributes = {
|
|
173
173
|
"service.name" => configuration.resolved_service_name,
|
|
174
174
|
"service.environment" => configuration.resolved_environment,
|
|
175
175
|
"telemetry.sdk.name" => "activeagent",
|
|
176
176
|
"telemetry.sdk.version" => ActiveAgent::VERSION
|
|
177
177
|
}
|
|
178
|
+
version = configuration.respond_to?(:service_version) ? configuration.service_version : nil
|
|
179
|
+
attributes["service.version"] = version if version.present?
|
|
180
|
+
attributes
|
|
178
181
|
end
|
|
179
182
|
end
|
|
180
183
|
end
|
|
@@ -78,9 +78,31 @@ module ActiveAgent
|
|
|
78
78
|
# end
|
|
79
79
|
def configure
|
|
80
80
|
yield configuration if block_given?
|
|
81
|
+
install_instrumentation! if enabled?
|
|
81
82
|
configuration
|
|
82
83
|
end
|
|
83
84
|
|
|
85
|
+
# Installs generation instrumentation on ActiveAgent::Base.
|
|
86
|
+
#
|
|
87
|
+
# The railtie also does this at boot, but only for configuration already
|
|
88
|
+
# loaded by then (activeagent.yml, config.active_agent.telemetry). A host
|
|
89
|
+
# app that enables telemetry from its own initializer runs *after*
|
|
90
|
+
# railties, so that check has already seen `enabled? == false` and
|
|
91
|
+
# skipped the install — leaving telemetry enabled but nothing
|
|
92
|
+
# instrumented, and so no traces despite a valid local_store. Calling it
|
|
93
|
+
# from {.configure} as well makes the install order-independent;
|
|
94
|
+
# `instrument_telemetry!` is idempotent, so the two paths cannot
|
|
95
|
+
# double-prepend.
|
|
96
|
+
#
|
|
97
|
+
# @api private
|
|
98
|
+
# @return [void]
|
|
99
|
+
def install_instrumentation!
|
|
100
|
+
return unless defined?(ActiveAgent::Base)
|
|
101
|
+
|
|
102
|
+
ActiveAgent::Base.include(Instrumentation)
|
|
103
|
+
ActiveAgent::Base.instrument_telemetry!
|
|
104
|
+
end
|
|
105
|
+
|
|
84
106
|
# Resets the configuration to defaults.
|
|
85
107
|
#
|
|
86
108
|
# @return [Configuration] New default configuration
|
data/lib/active_agent/version.rb
CHANGED
metadata
CHANGED
|
@@ -1,13 +1,14 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: activeagent
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 1.6.
|
|
4
|
+
version: 1.6.3
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Justin Bowen
|
|
8
|
+
autorequire:
|
|
8
9
|
bindir: bin
|
|
9
10
|
cert_chain: []
|
|
10
|
-
date:
|
|
11
|
+
date: 2026-09-18 00:00:00.000000000 Z
|
|
11
12
|
dependencies:
|
|
12
13
|
- !ruby/object:Gem::Dependency
|
|
13
14
|
name: actionpack
|
|
@@ -417,6 +418,7 @@ files:
|
|
|
417
418
|
- lib/active_agent/concerns/preview.rb
|
|
418
419
|
- lib/active_agent/concerns/provider.rb
|
|
419
420
|
- lib/active_agent/concerns/queueing.rb
|
|
421
|
+
- lib/active_agent/concerns/release.rb
|
|
420
422
|
- lib/active_agent/concerns/rescue.rb
|
|
421
423
|
- lib/active_agent/concerns/streaming.rb
|
|
422
424
|
- lib/active_agent/concerns/tooling.rb
|
|
@@ -432,6 +434,7 @@ files:
|
|
|
432
434
|
- lib/active_agent/delegation/schema.rb
|
|
433
435
|
- lib/active_agent/deprecator.rb
|
|
434
436
|
- lib/active_agent/evals.rb
|
|
437
|
+
- lib/active_agent/evals/correlation.rb
|
|
435
438
|
- lib/active_agent/evals/design_tokens.rb
|
|
436
439
|
- lib/active_agent/evals/diagnosis.rb
|
|
437
440
|
- lib/active_agent/evals/judge.rb
|
|
@@ -605,6 +608,7 @@ metadata:
|
|
|
605
608
|
documentation_uri: https://docs.activeagents.ai
|
|
606
609
|
source_code_uri: https://github.com/activeagents/activeagent
|
|
607
610
|
rubygems_mfa_required: 'true'
|
|
611
|
+
post_install_message:
|
|
608
612
|
rdoc_options: []
|
|
609
613
|
require_paths:
|
|
610
614
|
- lib
|
|
@@ -619,7 +623,8 @@ required_rubygems_version: !ruby/object:Gem::Requirement
|
|
|
619
623
|
- !ruby/object:Gem::Version
|
|
620
624
|
version: '0'
|
|
621
625
|
requirements: []
|
|
622
|
-
rubygems_version: 3.
|
|
626
|
+
rubygems_version: 3.5.22
|
|
627
|
+
signing_key:
|
|
623
628
|
specification_version: 4
|
|
624
629
|
summary: Rails AI Agents Framework
|
|
625
630
|
test_files: []
|