activeagent 1.6.2 → 1.6.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: b14c71e437e0c5701b6300861adb5d8b79eba328602e255b38928e6030584a2b
4
- data.tar.gz: c5eeeedce5730d88609b5131304571f9d7ecc12d3177ada460de7a4dd56ad61d
3
+ metadata.gz: a0332669c4ffee1ccbea3e1645e126acac6b4e957b980000e5d96baedc908638
4
+ data.tar.gz: 1d6062b56a6e7fc0e76b965d308df706c3be563292344cdf68b500046e838784
5
5
  SHA512:
6
- metadata.gz: d76b024171c0a84e4ecc6f6618ccb92d1e94a215f15e15839e8b78b37c8bcd6a312626b8dda996cca23bb2ebcab325fc08fab289689d66143855fa0c372ffb40
7
- data.tar.gz: 02605c62c4d22bc001462ee13451fce5f578cb43adb8f9380c8f116ef589629cafb175e10ce019fa073b9086382c2e8dd09bb03d9aab52813f75852a27e43aeb
6
+ metadata.gz: c0a940227737d94497c19e8760cf1241ea02f6078987f8cd27f78b0b1a1fb22ac90c8b924dff87e85a4bf8f13c952f6a36eec388dd806f56d230f5b38c9fe59b
7
+ data.tar.gz: df01fdb757be0bcfa549086dfe0082db0ada892fc552600087d07fe9c1299a4aa4d4c897987cbda897832d11df50f19368261413647f63db4e636a4bb5bb98e5
data/CHANGELOG.md CHANGED
@@ -7,6 +7,118 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [1.6.4] - 2026-09-22
11
+
12
+ Releases `activeagent` and `actionagent` 1.6.4 from one tag. A patch on 1.6.3
13
+ carrying one fix to `SchemaTools`, for a filter that answered confidently and
14
+ wrongly instead of failing.
15
+
16
+ ### Fixed
17
+
18
+ - A Rails enum is offered to the model as its names (`{type: "string", enum:
19
+ [...]}`) instead of the integer backing it. `SchemaGenerator` reads enums from
20
+ inclusion validators and never consulted `defined_enums`, so a `status` column
21
+ reached the model as a bare integer with no labels.
22
+ - A filter value outside an enum — alone or inside an IN list — is rejected,
23
+ naming the valid values, instead of matching no rows. `status: "pending"`
24
+ returned `{count: 0}`, which an agent reports as a fact, indistinguishable
25
+ from "none match". Same reasoning as the unknown-operator rejection in
26
+ `range_predicates!`.
27
+ - An enum is no longer offered the range form. Its integer backing is a
28
+ declaration-order artefact, so `status: {gt: 1}` was a meaningless filter that
29
+ still returned a confident count.
30
+
31
+ ## [1.6.3] - 2026-09-18
32
+
33
+ Releases `activeagent` and `actionagent` 1.6.3 from one tag.
34
+
35
+ A release about telling the truth on the screens that report what happened.
36
+ The context meter now divides the provider's own `prompt_tokens` among its
37
+ segments instead of subtracting estimates from it, and sizes each piece —
38
+ tool schemas, MCP schemas, instructions and the transcript — before the span
39
+ clips it for storage; previously a trace with dense tool schemas showed a
40
+ large message history that was never sent. An adapted replay is metered as
41
+ one execution like any other, so a host that supplies its own runtime is no
42
+ longer silently uncounted, and a spec naming a provider the agent cannot
43
+ serve now fails before the replay rather than reaching it.
44
+
45
+ Three seams hosts were reaching around become API. `ActiveAgent::Evals::Correlation`
46
+ joins `Runner`'s `around_evaluation:` hook to a telemetry backend's trace
47
+ scope, so a report row links back to the conversation behind it. A `Judge`
48
+ block that accepts `kind:` is told whether it is scoring, recommending or
49
+ writing the verdict, instead of matching on the gem's own instruction prose.
50
+ `Agent#generations` replaces the polymorphic join hosts were copying out of a
51
+ private service method.
52
+
53
+ Upgrading: no migration, and nothing that already worked changes. The judge
54
+ keyword reaches only a block that asks for it, so existing judges are
55
+ untouched; `Evaluation#replace_scenarios!` keeps `:destroy` as its default.
56
+ Adapters should drop any `ActionAgent.record_usage` call of their own, which
57
+ now double-counts, and any provider allow-list check of their own, which is
58
+ now dead code.
59
+
60
+ ### Added
61
+
62
+ - The prompt span records how large the tool schemas actually are, as
63
+ `prompt.input.tools.tokens`, `prompt.input.mcp_tools.tokens`,
64
+ `prompt.input.instructions.tokens` and `prompt.input.messages.tokens`. The
65
+ transcript's size is measured before the span trims the history to the turns
66
+ that fit, the others before their content is clipped. The content attributes
67
+ beside them are previews clipped for storage — and on the SDK path the tool
68
+ attribute is a roster of names and parameter keys, several times smaller than
69
+ the schema the model is sent — so a reader that sized the context from one
70
+ understated tool pressure badly.
71
+ - MCP tool schemas are attributed apart from the toolbox's, so the context meter
72
+ can name which half fills the window.
73
+ - `Evaluation#replace_scenarios!` takes `on_removed:` — `:destroy` (the
74
+ default, unchanged) or `:disable`, which keeps a scenario the suite no
75
+ longer names as `enabled: false` so earlier runs' results still resolve.
76
+ - **An evaluation's traces link back to the result that caused them.**
77
+ `ActiveAgent::Evals::Correlation` joins two APIs the module already had but
78
+ never connected: `Runner`'s `around_evaluation:` hook and its `metadata:`
79
+ run identity, and a telemetry backend's per-block agent scope. A run mints a
80
+ `run_id`, each evaluation a `result_id`, and both ride every trace opened
81
+ inside them as `eval.`-prefixed attributes; the trace ids travel the other
82
+ way onto `result.replay.metadata` — `trace_id` for the replay,
83
+ `judge_trace_ids` for the judge calls that graded it, with a run-level
84
+ verdict landing on the run metadata the Report carries rather than on
85
+ whichever result was evaluated last. The tracer is injected, so the module
86
+ takes on no telemetry dependency and `require "active_agent/evals"` still
87
+ loads on its own. Hand the object to `Runner.new(around_evaluation:)`
88
+ directly; a plain lambda there keeps working unchanged.
89
+ - A `Judge` block that accepts `kind:` is told which of the judge's three calls
90
+ it is serving — `:score`, `:recommend` or `:verdict` — so a host can trace,
91
+ budget or model them separately. Previously the only signal was the
92
+ `instructions` string, so hosts matched against the gem's own
93
+ `RECOMMEND_INSTRUCTIONS` / `VERDICT_INSTRUCTIONS` constants; rewording one
94
+ then sent every such host quietly down its `else` branch, mislabelling traces
95
+ rather than failing. The keyword reaches only a block that names it or
96
+ collects `**`, so judges taking `instructions:` and `prompt:` are unaffected.
97
+ (#462)
98
+ - `Agent#generations` reads the generations recorded against an agent, with
99
+ `Agent#agent_contexts` beside it. Generations hang off `AgentContext`
100
+ polymorphically, so reaching them meant hand-writing that join — the engine
101
+ did it itself in a private service method a host could not reuse, which now
102
+ uses the association instead. Destroying an agent still leaves its contexts
103
+ alone, as it always has. (#464)
104
+
105
+ ### Fixed
106
+
107
+ - The dashboard's context meter divides the provider's own `prompt_tokens`
108
+ among its segments instead of subtracting its estimates from it. Charging the
109
+ difference to one segment made "Messages" absorb the whole approximation
110
+ error, so a trace with dense JSON tool schemas read as a large message history
111
+ that was never sent. The transcript is one of the divided segments: dividing
112
+ only the rest would hand its share to the segments that remained, so a long
113
+ conversation reported an enormous system prompt and no history at all.
114
+ - A host that supplies a `scenario_evaluation_adapter_resolver` now has its
115
+ replays metered as executions, one per scenario x model, the same unit the
116
+ default path records. Previously an adapted replay was counted only if the
117
+ host remembered to call `ActionAgent.record_usage` itself.
118
+ - A scenario evaluation whose selected models name a provider the agent cannot
119
+ serve fails with `ArgumentError` before any replay runs. A spec handed back
120
+ as a Hash naming both `provider` and `model` bypassed the `providers:`
121
+ allow-list, so the run reached the replay with a provider nothing serves.
10
122
  ## [1.6.2] - 2026-09-16
11
123
 
12
124
  Releases `activeagent` and `actionagent` 1.6.2 from one tag.
@@ -0,0 +1,178 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "securerandom"
4
+ require "active_support/isolated_execution_state"
5
+
6
+ module ActiveAgent
7
+ module Evals
8
+ # Used to tie the traces an evaluation produces back to the run and the
9
+ # result that caused them, so a report row links to the exact conversation
10
+ # behind it.
11
+ #
12
+ # A run mints a `run_id`, each evaluation mints a `result_id`, and both ride
13
+ # every trace opened inside them as `eval.`-prefixed attributes. The trace
14
+ # ids travel the other way: a replay's trace id lands on
15
+ # `result.replay.metadata["trace_id"]`, and every judge call made while
16
+ # scoring that result appends to its `"judge_trace_ids"`. The verdict — a
17
+ # judge call made outside any evaluation — appends to the run metadata
18
+ # instead, which is the same Hash a Report carries as its `metadata`.
19
+ #
20
+ # A tracer is `(name, action:, attributes:, on_trace:) { ... }`: it opens a
21
+ # trace named for the agent, and calls `on_trace` with something answering
22
+ # to `#trace_id` once the trace is known.
23
+ #
24
+ # correlation = ActiveAgent::Evals::Correlation.new(
25
+ # agent_name: "SupportAgent",
26
+ # judge_name: "SupportAgentJudge",
27
+ # tracer: ->(name, action:, attributes:, on_trace:, &block) {
28
+ # MyTelemetry.with_agent(name, action: action, attributes: attributes,
29
+ # on_trace: on_trace, synchronous: true, &block)
30
+ # }
31
+ # )
32
+ #
33
+ # correlation.with_run("suite" => "support") do |metadata|
34
+ # Runner.new(
35
+ # scenarios: scenarios, models: models, metadata: metadata,
36
+ # replay: ->(scenario, spec) { correlation.replay { agent.run(scenario.prompt) } },
37
+ # judge: Judge.new(label: "judge-model") { |instructions:, prompt:|
38
+ # correlation.judge("score") { chat.with_instructions(instructions).ask(prompt).content }
39
+ # },
40
+ # around_evaluation: correlation
41
+ # ).call
42
+ # end
43
+ #
44
+ # Without a tracer the correlation still mints ids and merges metadata. The
45
+ # blocks run untraced.
46
+ class Correlation
47
+ STATE_KEY = :active_agent_evals_correlation
48
+
49
+ # DEFAULT_TRACE_KEYS names the correlation metadata that rides a trace as
50
+ # `eval.`-prefixed attributes. Anything else a caller puts in the run
51
+ # metadata (a tenant, a role) stays on the report but off the traces.
52
+ DEFAULT_TRACE_KEYS = %w[run_id result_id suite scenario_key model_label model provider].freeze
53
+
54
+ attr_reader :agent_name, :judge_name, :trace_keys
55
+
56
+ # @param agent_name [String] the trace name for the agent under evaluation
57
+ # @param judge_name [String] the trace name for judge traffic, kept distinct
58
+ # so grading calls do not read as the agent's own traffic
59
+ # @param tracer [#call, nil] `(name, action:, attributes:, on_trace:) { ... }`;
60
+ # nil runs every block untraced
61
+ # @param replay_action [String] the action name recorded for a replay trace
62
+ # @param trace_keys [Array<String>] which correlation keys become attributes
63
+ def initialize(agent_name:, judge_name: "#{agent_name}Judge", tracer: nil, replay_action: "eval",
64
+ trace_keys: DEFAULT_TRACE_KEYS)
65
+ @agent_name = agent_name
66
+ @judge_name = judge_name
67
+ @tracer = tracer
68
+ @replay_action = replay_action
69
+ @trace_keys = trace_keys.map(&:to_s)
70
+ end
71
+
72
+ # Opens a run. Mints `run_id` unless `metadata` carries one, and yields the
73
+ # metadata hash the traces will be correlated against — pass that same hash
74
+ # to `Runner.new(metadata:)` so the Report carries the run's identity and
75
+ # collects the verdict's trace id.
76
+ #
77
+ # The hash yielded is the caller's own, mutated in place, so a run
78
+ # reopened around a later verdict accumulates onto the metadata a Report
79
+ # already carries.
80
+ #
81
+ # @param metadata [Hash] opaque run metadata. A non-Hash is coerced with
82
+ # `#to_h`, and non-String keys are stringified.
83
+ # @yieldparam metadata [Hash]
84
+ def with_run(metadata = {})
85
+ run = metadata.is_a?(Hash) ? metadata : metadata.to_h
86
+ run.transform_keys!(&:to_s) unless run.keys.all?(String)
87
+ run["run_id"] ||= SecureRandom.uuid
88
+ with_context({ run: run, result: nil }) { yield run }
89
+ end
90
+
91
+ # Wraps one evaluation, in the shape `Runner.new(around_evaluation:)` calls:
92
+ # `(scenario, spec) { ... } → Result`. Mints a `result_id`, merges the
93
+ # correlation onto the Result's replay metadata, and returns the Result.
94
+ def around_evaluation(scenario, spec)
95
+ result_metadata = {
96
+ "run_id" => run_metadata["run_id"],
97
+ "result_id" => SecureRandom.uuid,
98
+ "scenario_key" => scenario.key,
99
+ "model_label" => spec.label,
100
+ "model" => spec.model,
101
+ "provider" => spec.provider
102
+ }.compact
103
+
104
+ with_context(run: run_metadata, result: result_metadata) do
105
+ yield.tap { |result| result.replay.metadata.merge!(result_metadata) }
106
+ end
107
+ end
108
+
109
+ # Delegates to `around_evaluation`, so the object satisfies
110
+ # `Runner.new(around_evaluation:)` directly.
111
+ def call(scenario, spec, &)
112
+ around_evaluation(scenario, spec, &)
113
+ end
114
+
115
+ # Traces one replay of the agent under evaluation. The trace id lands on
116
+ # the current result's metadata, so a report row links to the conversation.
117
+ def replay(action = @replay_action, &)
118
+ trace(@agent_name, action, judge: false, &)
119
+ end
120
+
121
+ # Traces one judge call. Appends to the current result's `judge_trace_ids`,
122
+ # or the run's when no evaluation is open (the verdict).
123
+ def judge(action = "score", &)
124
+ trace(@judge_name, action, judge: true, &)
125
+ end
126
+
127
+ # The correlation metadata in scope, or nil outside a run. A result's
128
+ # values win over the run's.
129
+ def current
130
+ context = ActiveSupport::IsolatedExecutionState[STATE_KEY]
131
+ return nil unless context
132
+
133
+ context.fetch(:run, {}).merge(context[:result] || {})
134
+ end
135
+
136
+ private
137
+
138
+ def run_metadata
139
+ context = ActiveSupport::IsolatedExecutionState[STATE_KEY]
140
+ context&.fetch(:run, nil) || {}
141
+ end
142
+
143
+ def with_context(context)
144
+ previous = ActiveSupport::IsolatedExecutionState[STATE_KEY]
145
+ ActiveSupport::IsolatedExecutionState[STATE_KEY] = context
146
+ yield
147
+ ensure
148
+ ActiveSupport::IsolatedExecutionState[STATE_KEY] = previous
149
+ end
150
+
151
+ def trace(name, action, judge:, &block)
152
+ return block.call unless @tracer
153
+
154
+ context = ActiveSupport::IsolatedExecutionState[STATE_KEY] || {}
155
+ correlation = context.fetch(:run, {}).merge(context[:result] || {})
156
+ attributes = correlation.slice(*@trace_keys).transform_keys { |key| "eval.#{key}" }
157
+ # A replay belongs to the evaluation that opened it and nowhere else, so
158
+ # it records no trace id when called outside one. A judge call outside an
159
+ # evaluation is the verdict, which belongs to the run.
160
+ target = judge ? (context[:result] || context[:run]) : context[:result]
161
+
162
+ @tracer.call(name, action: action, attributes: attributes, on_trace: recorder(target, judge: judge), &block)
163
+ end
164
+
165
+ def recorder(target, judge:)
166
+ lambda do |trace|
167
+ next unless target
168
+
169
+ if judge
170
+ (target["judge_trace_ids"] ||= []) << trace.trace_id
171
+ else
172
+ target["trace_id"] = trace.trace_id
173
+ end
174
+ end
175
+ end
176
+ end
177
+ end
178
+ end
@@ -10,6 +10,21 @@ module ActiveAgent
10
10
  # RubyLLM.chat(model: "claude-opus-5").with_instructions(instructions).ask(prompt).content
11
11
  # end
12
12
  #
13
+ # A judge serves three different calls, and a block that accepts `kind:` is
14
+ # told which one it is serving — `:score`, `:recommend` or `:verdict` — so a
15
+ # host can trace them apart, budget them apart, or score with a cheaper model
16
+ # than it writes the verdict with:
17
+ #
18
+ # Judge.new(label: "claude-opus-5") do |instructions:, prompt:, kind:|
19
+ # model = kind == :score ? "claude-haiku-4-5" : "claude-opus-5"
20
+ # RubyLLM.chat(model: model).with_instructions(instructions).ask(prompt).content
21
+ # end
22
+ #
23
+ # The keyword is passed only to a block that names it (or collects `**`), so
24
+ # a two-keyword block written before this is unaffected. Without it the only
25
+ # signal is the `instructions` string, which means matching on the gem's own
26
+ # prose — and a reworded constant then mislabels silently instead of failing.
27
+ #
13
28
  # Every method returns nil when the judge fails or answers unusably, so an
14
29
  # evaluation degrades to rule scoring rather than aborting.
15
30
  class Judge
@@ -23,6 +38,8 @@ module ActiveAgent
23
38
  # @param label [String] how reports name the judge (usually its model)
24
39
  # @yieldparam instructions [String] the system prompt
25
40
  # @yieldparam prompt [String] the user prompt
41
+ # @yieldparam kind [Symbol] which call this is — `:score`, `:recommend` or
42
+ # `:verdict`. Passed only to a block that accepts it.
26
43
  # @yieldreturn [String] the completion text
27
44
  # How much of a scenario's notes the judge reads. Where a suite's notes
28
45
  # are its grading rubric, a "Must not…" clause tends to come last, and a
@@ -41,7 +58,7 @@ module ActiveAgent
41
58
  return nil if answer.blank?
42
59
 
43
60
  guidance = criterion.dig("config", "prompt").presence || criterion["key"].to_s.humanize
44
- parse_score(ask(SCORE_INSTRUCTIONS, <<~PROMPT))
61
+ parse_score(ask(SCORE_INSTRUCTIONS, <<~PROMPT, :score))
45
62
  Criterion: #{guidance}
46
63
 
47
64
  The user asked:
@@ -64,7 +81,7 @@ module ActiveAgent
64
81
  def score_task(scenario:, answer:)
65
82
  return nil if answer.blank?
66
83
 
67
- parse_score(ask(SCORE_INSTRUCTIONS, <<~PROMPT))
84
+ parse_score(ask(SCORE_INSTRUCTIONS, <<~PROMPT, :score))
68
85
  A user asked an assistant:
69
86
  ---
70
87
  #{scenario.prompt}
@@ -90,7 +107,7 @@ module ActiveAgent
90
107
  "- #{call['name']}#{' (errored)' if call['error']}: #{call['arguments'].to_json.truncate(200)}"
91
108
  end.join("\n")
92
109
 
93
- parsed = parse_object(ask(RECOMMEND_INSTRUCTIONS, <<~PROMPT))
110
+ parsed = parse_object(ask(RECOMMEND_INSTRUCTIONS, <<~PROMPT, :recommend))
94
111
  An AI agent failed one evaluation scenario. Recommend the fix.
95
112
 
96
113
  Agent instructions:
@@ -146,7 +163,7 @@ module ActiveAgent
146
163
  "#{", faults: #{faults}" if faults.present?}"
147
164
  end
148
165
 
149
- parsed = parse_object(ask(VERDICT_INSTRUCTIONS, <<~PROMPT))
166
+ parsed = parse_object(ask(VERDICT_INSTRUCTIONS, <<~PROMPT, :verdict))
150
167
  An AI agent ran the same scenarios under several models. Its goals:
151
168
  ---
152
169
  #{instructions.to_s.truncate(1_000).presence || '(no instructions configured)'}
@@ -180,13 +197,34 @@ module ActiveAgent
180
197
  end
181
198
  end
182
199
 
183
- def ask(instructions, prompt)
184
- @generate.call(instructions: instructions, prompt: prompt).to_s
200
+ def ask(instructions, prompt, kind)
201
+ @generate.call(**ask_arguments(instructions, prompt, kind)).to_s
185
202
  rescue StandardError => e
186
203
  warn_failure(e)
187
204
  nil
188
205
  end
189
206
 
207
+ # The block signature is public API, and every judge written before `kind:`
208
+ # existed takes exactly `instructions:` and `prompt:` — passing a third
209
+ # keyword to one of those raises ArgumentError, which `ask` would swallow
210
+ # as a judge failure, degrading the run to rule scoring. So the kind goes
211
+ # only to a block that asked for it.
212
+ def ask_arguments(instructions, prompt, kind)
213
+ arguments = { instructions: instructions, prompt: prompt }
214
+ arguments[:kind] = kind if generate_accepts_kind?
215
+ arguments
216
+ end
217
+
218
+ # True for a block naming `kind:` or collecting `**`. Memoized because the
219
+ # answer cannot change for a given judge and `ask` runs per scored result.
220
+ def generate_accepts_kind?
221
+ return @generate_accepts_kind if defined?(@generate_accepts_kind)
222
+
223
+ @generate_accepts_kind = @generate.parameters.any? do |type, name|
224
+ type == :keyrest || (name == :kind && (type == :key || type == :keyreq))
225
+ end
226
+ end
227
+
190
228
  def warn_failure(error)
191
229
  message = "[ActiveAgent::Evals] judge #{label} failed: #{error.class}: #{error.message}"
192
230
  if defined?(Rails) && Rails.respond_to?(:logger) && Rails.logger
@@ -25,6 +25,7 @@ require_relative "evals/design_tokens"
25
25
  require_relative "evals/report_html"
26
26
  require_relative "evals/report"
27
27
  require_relative "evals/runner"
28
+ require_relative "evals/correlation"
28
29
  require_relative "evals/publisher"
29
30
 
30
31
  # Scenario evaluations for agents that answer with tools.
@@ -40,7 +41,8 @@ require_relative "evals/publisher"
40
41
  # fell short and what would fix it (Diagnosis, refined by an optional Judge),
41
42
  # and rolling everything up per model (Report). Runner ties them together
42
43
  # around one callable you supply: given a scenario and a model, run the
43
- # agent and return a Replay.
44
+ # agent and return a Replay. Correlation is optional plumbing on top: it
45
+ # links the traces a run emits back to the result that caused them.
44
46
  #
45
47
  # scenarios = ActiveAgent::Evals::ScenarioParser.scenarios(pasted_text)
46
48
  # models = ActiveAgent::Evals::ModelSpec.parse_all(%w[gpt-5-mini qwen3:8b], default_provider: "openai")
@@ -332,6 +332,8 @@ module ActiveAgent
332
332
  "`#{key}` is not a filterable attribute. Allowed filters: #{filterable.join(", ")}"
333
333
  end
334
334
 
335
+ validate_enum_value!(column, value)
336
+
335
337
  memo[column] = value
336
338
  end
337
339
 
@@ -386,6 +388,40 @@ module ActiveAgent
386
388
  end
387
389
  end
388
390
 
391
+ # A value outside a Rails enum reaches `where` as an unmatched name and
392
+ # returns zero rows — the same silent-nothing that {.range_predicates!}
393
+ # rejects for an unknown operator, and just as bad here: an agent reads
394
+ # "0 tickets under review" as a fact rather than a mistyped filter.
395
+ #
396
+ # The range form is not offered on an enum (see {.range_filterable?}), so
397
+ # a Hash here is a filter that cannot mean anything.
398
+ #
399
+ # @api private
400
+ # @raise [UnpermittedAttribute] on a value the enum does not define
401
+ def validate_enum_value!(column, value)
402
+ values = enum_values_for(column)
403
+ return if values.blank?
404
+
405
+ if value.is_a?(Hash)
406
+ raise UnpermittedAttribute,
407
+ "`#{column}` is an enum and cannot be compared as a range. " \
408
+ "Allowed values: #{values.keys.join(", ")}"
409
+ end
410
+
411
+ # An Array is an IN filter, which `where` silently narrows to its
412
+ # defined members just as it drops a lone undefined one.
413
+ Array(value).each do |member|
414
+ # The schema offers names only, so an integer here came from a
415
+ # model ignoring it. Accepted anyway: it is unambiguous, and a host
416
+ # calling the tool directly in Ruby reasonably passes the backing
417
+ # value.
418
+ next if values.key?(member.to_s) || values.value?(member)
419
+
420
+ raise UnpermittedAttribute,
421
+ "`#{member}` is not a valid `#{column}`. Allowed values: #{values.keys.join(", ")}"
422
+ end
423
+ end
424
+
389
425
  # Projects a record down to the declared return columns.
390
426
  #
391
427
  # The projection happens in SQL (+select+) as well as here, but the Ruby
@@ -455,14 +491,44 @@ module ActiveAgent
455
491
  properties = schema[:schema][:properties]
456
492
 
457
493
  filterable.index_with do |column|
494
+ next enum_property(column) if enum_column?(column)
495
+
458
496
  scalar = (properties[column] || { type: "string" }).deep_dup
459
497
  range_filterable?(column) ? with_range_form(column, scalar) : scalar
460
498
  end
461
499
  end
462
500
 
501
+ # A Rails enum is declared on the model, not through the inclusion
502
+ # validator SchemaGenerator reads, so without this a `status` column
503
+ # reaches the model as a bare integer: no labels, no constraint. The
504
+ # names are the model's own API (`Ticket.open`, `status: "open"`), so
505
+ # they are what the tool offers.
506
+ #
507
+ # @api private
508
+ def enum_property(column)
509
+ { type: "string", enum: enum_values_for(column).keys, description: "#{column.to_s.humanize} field" }
510
+ end
511
+
512
+ # @api private
513
+ def enum_column?(column)
514
+ enum_values_for(column).present?
515
+ end
516
+
517
+ # @api private
518
+ def enum_values_for(column)
519
+ @model.defined_enums[column.to_s] || {}
520
+ end
521
+
463
522
  # Dates, times and numbers are the columns a question like "overdue" or
464
523
  # "more than 10" actually needs a comparison on.
524
+ #
525
+ # An enum is backed by an integer but is not ordered in any sense a
526
+ # question can use: `status > 1` is an artefact of declaration order, so
527
+ # offering the range form on one invites a meaningless filter that still
528
+ # returns a confident count.
465
529
  def range_filterable?(column)
530
+ return false if enum_column?(column)
531
+
466
532
  RANGE_FILTERABLE_TYPES.include?(@model.type_for_attribute(column).type)
467
533
  end
468
534
 
@@ -101,7 +101,23 @@ module ActiveAgent
101
101
  parameters: properties.is_a?(Hash) ? properties.keys : []
102
102
  }.compact
103
103
  }
104
- prompt_span.set_attribute("prompt.input.tools", JSON.generate(roster)) if roster.any?
104
+ if roster.any?
105
+ prompt_span.set_attribute("prompt.input.tools", JSON.generate(roster))
106
+ # The roster is a reading aid: names, a truncated description
107
+ # and parameter keys. What the model is sent is the full JSON
108
+ # Schema — types, enums, nested objects, anyOf branches —
109
+ # which for a twelve-tool agent is several times the roster,
110
+ # so a context meter sizing tool pressure from the roster
111
+ # understates it badly. The schemas come from the host and
112
+ # need not be serializable; a size is worth less than the
113
+ # generation it would otherwise take down.
114
+ tools_size = begin
115
+ JSON.generate(tools).length
116
+ rescue StandardError
117
+ nil
118
+ end
119
+ prompt_span.set_attribute("prompt.input.tools.tokens", (tools_size / 4.0).round) if tools_size
120
+ end
105
121
  end
106
122
 
107
123
  # prompt_options[:messages] holds the turns a caller passed
@@ -1,3 +1,3 @@
1
1
  module ActiveAgent
2
- VERSION = "1.6.2"
2
+ VERSION = "1.6.4"
3
3
  end
metadata CHANGED
@@ -1,14 +1,14 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: activeagent
3
3
  version: !ruby/object:Gem::Version
4
- version: 1.6.2
4
+ version: 1.6.4
5
5
  platform: ruby
6
6
  authors:
7
7
  - Justin Bowen
8
8
  autorequire:
9
9
  bindir: bin
10
10
  cert_chain: []
11
- date: 2026-09-17 00:00:00.000000000 Z
11
+ date: 2026-09-23 00:00:00.000000000 Z
12
12
  dependencies:
13
13
  - !ruby/object:Gem::Dependency
14
14
  name: actionpack
@@ -434,6 +434,7 @@ files:
434
434
  - lib/active_agent/delegation/schema.rb
435
435
  - lib/active_agent/deprecator.rb
436
436
  - lib/active_agent/evals.rb
437
+ - lib/active_agent/evals/correlation.rb
437
438
  - lib/active_agent/evals/design_tokens.rb
438
439
  - lib/active_agent/evals/diagnosis.rb
439
440
  - lib/active_agent/evals/judge.rb