activeagent 1.7.0 → 1.7.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +80 -0
- data/README.md +3 -0
- data/lib/active_agent/evals/model_spec.rb +4 -0
- data/lib/active_agent/evals/runner.rb +2 -2
- data/lib/active_agent/evals/scenario.rb +3 -3
- data/lib/active_agent/evals/scenario_parser.rb +3 -3
- data/lib/active_agent/evals/suite.rb +8 -8
- data/lib/active_agent/evals.rb +1 -1
- data/lib/active_agent/providers/_base_provider.rb +10 -2
- data/lib/active_agent/providers/ruby_llm_provider.rb +53 -3
- data/lib/active_agent/version.rb +1 -1
- metadata +4 -4
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 14e5b0cec7ba170c446d1ed32ad5c9d0705ac8eae4fdc6e4e4a7d17bdae502a2
|
|
4
|
+
data.tar.gz: 28e12f0ef3727a97fc231e5f7ace15ad4cb020b379fc76c757c2206c89f2b2e6
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 9cd766d568ab510b043e4902c1c5fa73ce877a268761168509f2772190ab0f2d6f259fe62e4654aa2a002d20228f876fce82c7c2f5f3e1018cd6cb01935aed27
|
|
7
|
+
data.tar.gz: 197cab65243e11561c7216f08e9b30369819e3ef73627d0f6e72092a7fc74c16a278638f5044e11f292149a1a39dbe11cab9593a41e7289049a3a90c6232a0a0
|
data/CHANGELOG.md
CHANGED
|
@@ -7,6 +7,86 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
7
7
|
|
|
8
8
|
## [Unreleased]
|
|
9
9
|
|
|
10
|
+
## [1.7.1] - 2026-09-29
|
|
11
|
+
|
|
12
|
+
Releases `activeagent` and `actionagent` 1.7.1 from one tag. A patch on 1.7.0:
|
|
13
|
+
the Evaluations page picks the judge and compared models from the provider
|
|
14
|
+
catalogs, `GET /api/provider_models` lists the host's RubyLLM registry, and
|
|
15
|
+
the RubyLLM provider requires ruby_llm 1.x and sends tool calls and structured
|
|
16
|
+
output in the shape ruby_llm reads. No migrations.
|
|
17
|
+
|
|
18
|
+
### Added
|
|
19
|
+
|
|
20
|
+
- **Model pickers on the Evaluations page** (`actionagent`). The judge model
|
|
21
|
+
and the models to compare, on the new-evaluation form and on a scenario
|
|
22
|
+
suite, are chosen from type-ahead suggestions, and a model the suggestions
|
|
23
|
+
lack can still be typed. The compare fields hold one removable chip per
|
|
24
|
+
model and still submit the same comma-separated list.
|
|
25
|
+
- With scenarios, the suggestions are the catalogs the agent builder uses,
|
|
26
|
+
for the providers the owner's runs have credentials for. A model whose
|
|
27
|
+
name alone would run elsewhere is offered with its provider in front, as
|
|
28
|
+
`ModelSpec.parse` reads it: `ollama/llama3.2`, or
|
|
29
|
+
`openrouter/anthropic/claude-sonnet-4.5` for OpenRouter's copy of an
|
|
30
|
+
Anthropic model.
|
|
31
|
+
- Without scenarios, a compared model selects the generations recorded
|
|
32
|
+
under its name, usually the provider's dated id, so the suggestions are
|
|
33
|
+
the names the agent's generations were recorded under, from the new
|
|
34
|
+
`GET /api/agents/:id/recorded_models`. Adding or clearing the scenarios
|
|
35
|
+
renames the catalog models already chosen to match.
|
|
36
|
+
- The judge field suggests the models of the provider the judge runs on
|
|
37
|
+
and says which provider that is, or that the credentials deciding it
|
|
38
|
+
could not be read.
|
|
39
|
+
|
|
40
|
+
`GET /api/evaluations` reports that provider as `judge_provider` and the
|
|
41
|
+
providers runs can use as `model_providers`
|
|
42
|
+
(`AgentExecutionService.available_providers`). Runs and their judge use
|
|
43
|
+
the evaluated agent's owner's credentials, so both fields describe that
|
|
44
|
+
owner's when `agent_id` scopes the list, and the signed-in owner's
|
|
45
|
+
otherwise. `judge_provider` is null when no provider has credentials, and
|
|
46
|
+
also when reading them raised, which `judge_provider_error: true` marks.
|
|
47
|
+
`model_providers` leaves out a provider whose credentials cannot be read.
|
|
48
|
+
Either way the list still loads.
|
|
49
|
+
- **`GET /api/provider_models` lists the host's RubyLLM registry**
|
|
50
|
+
(`actionagent`). When the host app loads RubyLLM, the chat models its
|
|
51
|
+
registry lists for the provider that take and return text (RubyLLM's
|
|
52
|
+
bundled catalog, or the host's own model table) follow the live or curated
|
|
53
|
+
list, each once, so the builder's preselected default is unchanged. That
|
|
54
|
+
leaves out the speech, transcription, moderation and completion-only
|
|
55
|
+
models RubyLLM counts as chat models, which its registry lists with no
|
|
56
|
+
modalities or as taking audio, and with them the few chat models it lists
|
|
57
|
+
with no modalities, which can still be typed. A registry that raises is
|
|
58
|
+
logged and leaves the list as it was.
|
|
59
|
+
|
|
60
|
+
### Changed
|
|
61
|
+
|
|
62
|
+
- **The RubyLLM provider requires ruby_llm 1.x** (`activeagent`). ruby_llm
|
|
63
|
+
2.0 renamed the APIs `RubyLLMProvider` calls, and the open `>= 1.0`
|
|
64
|
+
requirement let `bundle update` install it. The provider now requires
|
|
65
|
+
`~> 1.0` until it supports 2.0 (#502). Loading it with an unsupported
|
|
66
|
+
version names the supported range and the loaded version, instead of
|
|
67
|
+
asking for a gem that is already in the Gemfile. Pin
|
|
68
|
+
`gem "ruby_llm", "~> 1.0"` if your bundle resolved 2.0.
|
|
69
|
+
|
|
70
|
+
### Fixed
|
|
71
|
+
|
|
72
|
+
- **Tool calls sent back through the RubyLLM provider** (`activeagent`).
|
|
73
|
+
After a tool ran, the follow-up request repeated the model's tool call
|
|
74
|
+
with its arguments as a JSON string where ruby_llm expects a Hash: OpenAI
|
|
75
|
+
received them JSON-encoded twice, and Anthropic received a string for
|
|
76
|
+
`tool_use.input`, which its API requires to be an object. The same
|
|
77
|
+
happened when a stored conversation containing a tool call was replayed.
|
|
78
|
+
The provider now hands ruby_llm the parsed arguments (#501).
|
|
79
|
+
- **Structured output through the RubyLLM provider** (`activeagent`). A
|
|
80
|
+
`json_schema` response_format reached ruby_llm unchanged, but ruby_llm
|
|
81
|
+
reads `{ name:, schema:, strict: }`, so OpenAI received a schema with a
|
|
82
|
+
null name and body, and the Anthropic request raised inside ruby_llm
|
|
83
|
+
before it was sent. The provider now converts it, naming the schema
|
|
84
|
+
`response` and making it strict unless the format says otherwise, as
|
|
85
|
+
ruby_llm's own `with_schema` does. A `text` format asks for plain text;
|
|
86
|
+
`json_object`, which ruby_llm has no mode for, and a `json_schema`
|
|
87
|
+
without a schema now raise `ArgumentError` instead of sending a request
|
|
88
|
+
the API rejects (#501).
|
|
89
|
+
|
|
10
90
|
## [1.7.0] - 2026-09-24
|
|
11
91
|
|
|
12
92
|
Releases `activeagent` and `actionagent` 1.7.0 from one tag. A minor release:
|
data/README.md
CHANGED
|
@@ -20,6 +20,10 @@ module ActiveAgent
|
|
|
20
20
|
class ModelSpec
|
|
21
21
|
DEFAULT_PROVIDERS = %w[openai anthropic ollama openrouter].freeze
|
|
22
22
|
|
|
23
|
+
# The dashboard's model pickers resolve names with a JavaScript copy of
|
|
24
|
+
# these rules and of .parse (actionagent/frontend/utils/modelOptions.mjs).
|
|
25
|
+
# actionagent/test/fixtures/model_spec_cases.json holds the cases both
|
|
26
|
+
# are tested against, so a change here needs the same change there.
|
|
23
27
|
DEFAULT_INFERENCE_RULES = [
|
|
24
28
|
[ /\Aclaude/i, "anthropic" ],
|
|
25
29
|
[ /\A(gpt-|o\d|chatgpt|text-embedding)/i, "openai" ],
|
|
@@ -16,10 +16,10 @@ module ActiveAgent
|
|
|
16
16
|
# faults in `refine_faults` (up to `judge_limit` calls per run).
|
|
17
17
|
#
|
|
18
18
|
# Runner.new(
|
|
19
|
-
# scenarios: suite.scenarios(groups: %w[
|
|
19
|
+
# scenarios: suite.scenarios(groups: %w[history]),
|
|
20
20
|
# models: ModelSpec.parse_all(%w[gpt-5-mini claude-haiku-4-5], default_provider: "openai"),
|
|
21
21
|
# criteria: [{ "key" => "response_present", "type" => "response_present" }],
|
|
22
|
-
# available_tools: { "
|
|
22
|
+
# available_tools: { "lookup_order" => "Find an order by number" },
|
|
23
23
|
# instructions: agent.instructions,
|
|
24
24
|
# judge: Judge.new(label: "claude-opus-5") { |instructions:, prompt:| ... },
|
|
25
25
|
# replay: ->(scenario, spec) { ... },
|
|
@@ -7,11 +7,11 @@ module ActiveAgent
|
|
|
7
7
|
# answer is expected to do.
|
|
8
8
|
#
|
|
9
9
|
# @!attribute key
|
|
10
|
-
# @return [String] stable within a suite ("
|
|
10
|
+
# @return [String] stable within a suite ("history_3"), so results line up across runs
|
|
11
11
|
# @!attribute group
|
|
12
|
-
# @return [String, nil] the group key ("
|
|
12
|
+
# @return [String, nil] the group key ("history")
|
|
13
13
|
# @!attribute group_name
|
|
14
|
-
# @return [String, nil] the group's display name ("
|
|
14
|
+
# @return [String, nil] the group's display name ("History / audit")
|
|
15
15
|
# @!attribute expected_tools
|
|
16
16
|
# @return [Array<String>] tool names a passing answer calls (any one of them)
|
|
17
17
|
# @!attribute expected_patterns
|
|
@@ -9,8 +9,8 @@ module ActiveAgent
|
|
|
9
9
|
# - `# Heading` or `**Heading**` lines, which start a group; so does a
|
|
10
10
|
# short unmarked line ending in a colon (`Find records:`)
|
|
11
11
|
# - a message in backticks at the start of the line, followed by notes,
|
|
12
|
-
# as in an issue
|
|
13
|
-
# `` 3. `Show me all
|
|
12
|
+
# as in a list copied from an issue:
|
|
13
|
+
# `` 3. `Show me all tickets with no assignee` — 12 in the sample data ``
|
|
14
14
|
# - trailing ` | tools: a, b | contains: x | not_contains: y | key: k`
|
|
15
15
|
# options on a line
|
|
16
16
|
# - a JSON array of strings, or of objects with `prompt` (or `message`),
|
|
@@ -19,7 +19,7 @@ module ActiveAgent
|
|
|
19
19
|
# group names, stable keys, notes and production-only flags
|
|
20
20
|
#
|
|
21
21
|
# Every scenario gets a key unique within the paste, derived from its group
|
|
22
|
-
# and position ("
|
|
22
|
+
# and position ("history_3"), unless the line names one. The result is an
|
|
23
23
|
# array of string-keyed hashes; `Scenario.from_hash` builds the structs.
|
|
24
24
|
class ScenarioParser
|
|
25
25
|
class ParseError < ArgumentError; end
|
|
@@ -8,17 +8,17 @@ module ActiveAgent
|
|
|
8
8
|
# suite. That is how an app keeps a shared suite and lets a deployment add or
|
|
9
9
|
# reword questions.
|
|
10
10
|
#
|
|
11
|
-
# suite:
|
|
12
|
-
# description:
|
|
11
|
+
# suite: support_desk
|
|
12
|
+
# description: Questions the support team asks every week
|
|
13
13
|
# groups:
|
|
14
|
-
# - key:
|
|
15
|
-
# name:
|
|
14
|
+
# - key: open_tickets
|
|
15
|
+
# name: Open tickets
|
|
16
16
|
# scenarios:
|
|
17
|
-
# - key:
|
|
18
|
-
# prompt: Which
|
|
17
|
+
# - key: open_tickets_1
|
|
18
|
+
# prompt: Which open tickets mention a refund?
|
|
19
19
|
# expect:
|
|
20
|
-
# tools: [
|
|
21
|
-
# notes:
|
|
20
|
+
# tools: [find_tickets, count_tickets]
|
|
21
|
+
# notes: The sample data has three.
|
|
22
22
|
# production_only: false
|
|
23
23
|
class Suite
|
|
24
24
|
class NotFound < StandardError; end
|
data/lib/active_agent/evals.rb
CHANGED
|
@@ -50,7 +50,7 @@ require_relative "evals/publisher"
|
|
|
50
50
|
# report = ActiveAgent::Evals::Runner.new(
|
|
51
51
|
# scenarios: scenarios,
|
|
52
52
|
# models: models,
|
|
53
|
-
# available_tools: { "
|
|
53
|
+
# available_tools: { "lookup_order" => "Find an order by number" },
|
|
54
54
|
# replay: ->(scenario, spec) { my_agent.run(scenario.prompt, model: spec.model, provider: spec.provider) }
|
|
55
55
|
# ).call
|
|
56
56
|
#
|
|
@@ -10,7 +10,8 @@ require_relative "concerns/tool_choice_clearing"
|
|
|
10
10
|
GEM_LOADERS = {
|
|
11
11
|
anthropic: [ "anthropic", "~> 1.12", "anthropic" ],
|
|
12
12
|
openai: [ "openai", "~> 0.34", "openai" ],
|
|
13
|
-
|
|
13
|
+
# ruby_llm 2.0 renamed the APIs RubyLLMProvider calls.
|
|
14
|
+
ruby_llm: [ "ruby_llm", "~> 1.0", "ruby_llm" ]
|
|
14
15
|
}
|
|
15
16
|
|
|
16
17
|
# Requires a provider's gem dependency.
|
|
@@ -18,7 +19,8 @@ GEM_LOADERS = {
|
|
|
18
19
|
# @param type [Symbol] provider type (:anthropic, :openai)
|
|
19
20
|
# @param file_name [String] for error context
|
|
20
21
|
# @return [void]
|
|
21
|
-
# @raise [LoadError] when
|
|
22
|
+
# @raise [LoadError] when the gem is not installed, or when the loaded
|
|
23
|
+
# version is outside the supported range
|
|
22
24
|
def require_gem!(type, file_name)
|
|
23
25
|
gem_name, requirement, package_name = GEM_LOADERS.fetch(type)
|
|
24
26
|
provider_name = file_name.split("/").last.delete_suffix(".rb").camelize
|
|
@@ -27,6 +29,12 @@ def require_gem!(type, file_name)
|
|
|
27
29
|
gem(gem_name, requirement)
|
|
28
30
|
require(package_name)
|
|
29
31
|
rescue LoadError
|
|
32
|
+
loaded = Gem.loaded_specs[gem_name]
|
|
33
|
+
if loaded && !Gem::Requirement.new(requirement).satisfied_by?(loaded.version)
|
|
34
|
+
raise LoadError, "#{provider_name} supports the '#{gem_name}' gem #{requirement}, but #{loaded.version} is loaded. " \
|
|
35
|
+
"Add `gem \"#{gem_name}\", \"#{requirement}\"` to your Gemfile and run `bundle update #{gem_name}`."
|
|
36
|
+
end
|
|
37
|
+
|
|
30
38
|
raise LoadError, "The '#{gem_name}' gem is required for #{provider_name}. Please add it to your Gemfile and run `bundle install`."
|
|
31
39
|
end
|
|
32
40
|
end
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
require_relative "_base_provider"
|
|
4
4
|
|
|
5
|
-
require_gem!(:ruby_llm, __FILE__)
|
|
5
|
+
require_gem!(:ruby_llm, __FILE__)
|
|
6
6
|
|
|
7
7
|
require_relative "ruby_llm/_types"
|
|
8
8
|
require_relative "ruby_llm/tool_proxy"
|
|
@@ -59,7 +59,8 @@ module ActiveAgent
|
|
|
59
59
|
tools: tools || {},
|
|
60
60
|
temperature: parameters[:temperature]
|
|
61
61
|
}
|
|
62
|
-
|
|
62
|
+
schema = ruby_llm_schema(parameters[:response_format])
|
|
63
|
+
kwargs[:schema] = schema if schema
|
|
63
64
|
|
|
64
65
|
# Pass extra params (max_tokens, etc.) via RubyLLM's params: deep-merge
|
|
65
66
|
max_tokens = parameters[:max_tokens] || options.max_tokens
|
|
@@ -345,12 +346,61 @@ module ActiveAgent
|
|
|
345
346
|
call = ::RubyLLM::ToolCall.new(
|
|
346
347
|
id: id,
|
|
347
348
|
name: tc.dig(:function, :name) || tc[:name],
|
|
348
|
-
arguments: tc.dig(:function, :arguments) || tc[:input]
|
|
349
|
+
arguments: ruby_llm_tool_arguments(tc.dig(:function, :arguments) || tc[:input])
|
|
349
350
|
)
|
|
350
351
|
hash[id] = call
|
|
351
352
|
end
|
|
352
353
|
end
|
|
353
354
|
|
|
355
|
+
# ActiveAgent keeps a tool call's arguments as the JSON string the
|
|
356
|
+
# model sent; RubyLLM takes them as a Hash and renders them itself
|
|
357
|
+
# (JSON-encoded for OpenAI, as the input object for Anthropic).
|
|
358
|
+
#
|
|
359
|
+
# A string that is not valid JSON, such as arguments a model cut off
|
|
360
|
+
# mid-way in a stored conversation, is passed through unchanged.
|
|
361
|
+
#
|
|
362
|
+
# @param arguments [String, Hash, nil]
|
|
363
|
+
# @return [Hash, String]
|
|
364
|
+
def ruby_llm_tool_arguments(arguments)
|
|
365
|
+
case arguments
|
|
366
|
+
when Hash
|
|
367
|
+
arguments.deep_stringify_keys
|
|
368
|
+
when String
|
|
369
|
+
arguments.blank? ? {} : JSON.parse(arguments)
|
|
370
|
+
else
|
|
371
|
+
{}
|
|
372
|
+
end
|
|
373
|
+
rescue JSON::ParserError
|
|
374
|
+
arguments
|
|
375
|
+
end
|
|
376
|
+
|
|
377
|
+
# Converts ActiveAgent's response_format to the schema RubyLLM's
|
|
378
|
+
# complete takes, { name:, schema:, strict: }, filling the name and
|
|
379
|
+
# strict flag in the way RubyLLM::Chat#with_schema does.
|
|
380
|
+
#
|
|
381
|
+
# @param response_format [Hash, Symbol, String, nil] ActiveAgent common format
|
|
382
|
+
# @return [Hash, nil] nil when the response is plain text
|
|
383
|
+
# @raise [ArgumentError] for json_schema without a schema, and for any
|
|
384
|
+
# other type, including json_object, which RubyLLM has no mode for
|
|
385
|
+
def ruby_llm_schema(response_format)
|
|
386
|
+
return nil if response_format.nil?
|
|
387
|
+
|
|
388
|
+
format = response_format.is_a?(Hash) ? response_format : { type: response_format.to_s }
|
|
389
|
+
|
|
390
|
+
case format[:type].to_s
|
|
391
|
+
when "text"
|
|
392
|
+
nil
|
|
393
|
+
when "json_schema"
|
|
394
|
+
json_schema = format[:json_schema] || {}
|
|
395
|
+
raise ArgumentError, "RubyLLMProvider needs a schema for a json_schema response_format" unless json_schema[:schema]
|
|
396
|
+
|
|
397
|
+
{ name: json_schema[:name] || "response", schema: json_schema[:schema], strict: json_schema[:strict] != false }
|
|
398
|
+
else
|
|
399
|
+
raise ArgumentError, "RubyLLMProvider supports a json_schema or text response_format, not #{format[:type].inspect}; " \
|
|
400
|
+
"ruby_llm has no JSON object mode, so give json_schema a schema instead"
|
|
401
|
+
end
|
|
402
|
+
end
|
|
403
|
+
|
|
354
404
|
# Converts ActiveAgent tool definitions to RubyLLM ToolProxy objects.
|
|
355
405
|
#
|
|
356
406
|
# @param tools [Array<Hash>, nil] ActiveAgent tool definitions
|
data/lib/active_agent/version.rb
CHANGED
metadata
CHANGED
|
@@ -1,14 +1,14 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: activeagent
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 1.7.
|
|
4
|
+
version: 1.7.1
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Justin Bowen
|
|
8
8
|
autorequire:
|
|
9
9
|
bindir: bin
|
|
10
10
|
cert_chain: []
|
|
11
|
-
date: 2026-09-
|
|
11
|
+
date: 2026-09-30 00:00:00.000000000 Z
|
|
12
12
|
dependencies:
|
|
13
13
|
- !ruby/object:Gem::Dependency
|
|
14
14
|
name: actionpack
|
|
@@ -204,14 +204,14 @@ dependencies:
|
|
|
204
204
|
name: ruby_llm
|
|
205
205
|
requirement: !ruby/object:Gem::Requirement
|
|
206
206
|
requirements:
|
|
207
|
-
- - "
|
|
207
|
+
- - "~>"
|
|
208
208
|
- !ruby/object:Gem::Version
|
|
209
209
|
version: '1.0'
|
|
210
210
|
type: :development
|
|
211
211
|
prerelease: false
|
|
212
212
|
version_requirements: !ruby/object:Gem::Requirement
|
|
213
213
|
requirements:
|
|
214
|
-
- - "
|
|
214
|
+
- - "~>"
|
|
215
215
|
- !ruby/object:Gem::Version
|
|
216
216
|
version: '1.0'
|
|
217
217
|
- !ruby/object:Gem::Dependency
|