ruby_llm-llm_judge 0.1.1 → 0.1.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +11 -0
- data/README.md +12 -9
- data/lib/ruby_llm/llm_judge/engine.rb +46 -17
- data/lib/ruby_llm/llm_judge/legacy.rb +60 -14
- data/lib/ruby_llm/llm_judge/version.rb +1 -1
- data/lib/ruby_llm/llm_judge.rb +3 -5
- metadata +1 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 71f6bd771e7f6bec3a9dcdf7036ee1238110c63ee530b15d32c6dadcfb0d0eaa
|
|
4
|
+
data.tar.gz: e54b13efadda477c3d3713c66184dc4ce01061b740e06ab52e95df0b672ae7e9
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 4e29ff7f1b47222dfc71f6f096cd058615d2517b7e16b29f9c2a68f094f2af95be8a99232d5ba72cc1d6b9a932fa5f9bf290e701a9744ee1a241fc6a0c710834
|
|
7
|
+
data.tar.gz: 942f4a7da0c27d1741d3e0e2b80ae719335123d6ddf276b84b0e4d200ed315d5f839cfc1ea998e708b6b54f83d7b7e97c718f6d4a070e17d2bfe0be23e9b3541
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,16 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.1.2 (2026-09-28)
|
|
4
|
+
|
|
5
|
+
- When a rating call fails, no new calls start and the error is raised once calls in flight finish. Previously the remaining calls kept running, and were billed, after `judge` raised.
|
|
6
|
+
- `chat_provider_options` is merged over the defaults instead of replacing them. Setting one option no longer re-enables reasoning or drops `store: false` for Luna.
|
|
7
|
+
- One-call JSON inside a Markdown code fence is accepted. Claude Haiku 4.5 now works with `strategy: :single_request`.
|
|
8
|
+
- A one-call Choice whose top probabilities tie gets the corrective retry, then raises, instead of selecting the first option. `tie_breaker: :first` keeps the old behavior.
|
|
9
|
+
- New `max_workers:` option (default 6) for concurrent rating calls.
|
|
10
|
+
- Unknown provider options are named in the error.
|
|
11
|
+
- A one-option Choice reports confidence 1.0.
|
|
12
|
+
- RubyLLM 1.13 through 1.16: question validation matches RubyLLM 2.1 (duplicate or empty option names, duplicate probability outcomes, description types), and token totals keep cached and thinking tokens.
|
|
13
|
+
|
|
3
14
|
## 0.1.1 (2026-09-28)
|
|
4
15
|
|
|
5
16
|
- Fix invalid scoring responses raising `NoMethodError` on RubyLLM 1.13, which also skipped the corrective retry. They now raise `RubyLLM::LLMJudge::Error`, a `RubyLLM::Error`, on every supported RubyLLM version.
|
data/README.md
CHANGED
|
@@ -19,7 +19,7 @@ The gem also runs on RubyLLM 1.13 through 1.16, which predate the Judge API. The
|
|
|
19
19
|
What differs before 2.1:
|
|
20
20
|
|
|
21
21
|
- Results are `RubyLLM::LLMJudge::Legacy` objects, not `RubyLLM::Judgment`, `RubyLLM::Choice`, and so on. Avoid class checks until you upgrade.
|
|
22
|
-
- There is no `RubyLLM.judge`, `RubyLLM::Judge` class DSL, `cost`,
|
|
22
|
+
- There is no `RubyLLM.judge`, `RubyLLM::Judge` class DSL, `cost`, instrumentation, or `metadata:`. Passing `metadata:` raises. `context:` is accepted and supplies its configuration.
|
|
23
23
|
- `scoring_protocol` accepts only `:chat_completions`, the API RubyLLM 1.x uses for OpenAI.
|
|
24
24
|
- The output limit and `chat_provider_options` are sent with `with_params`, using the field each provider expects (`max_completion_tokens` for OpenAI and Azure, `generationConfig.maxOutputTokens` for Gemini and Vertex AI, `inferenceConfig.maxTokens` for Bedrock, and `max_tokens` otherwise).
|
|
25
25
|
|
|
@@ -70,20 +70,23 @@ TicketTriage.judge('Please refund today.').urgent.probability
|
|
|
70
70
|
|
|
71
71
|
## Choose a model
|
|
72
72
|
|
|
73
|
-
Pass the chat model as `model:` and its provider as `scoring_provider
|
|
73
|
+
Pass the chat model as `model:` and its provider as `scoring_provider`. For example, Claude Haiku 4.5 through OpenRouter:
|
|
74
74
|
|
|
75
75
|
```ruby
|
|
76
76
|
result = RubyLLM::LLMJudge.judge(
|
|
77
77
|
'Please refund the duplicate charge today.',
|
|
78
|
-
model: 'claude-haiku-4
|
|
78
|
+
model: 'anthropic/claude-haiku-4.5',
|
|
79
79
|
questions: { urgent: { type: :probability, instructions: 'Does this need attention today?' } },
|
|
80
80
|
provider_options: {
|
|
81
|
-
scoring_provider: :
|
|
81
|
+
scoring_provider: :openrouter,
|
|
82
|
+
temperature: 0
|
|
82
83
|
}
|
|
83
84
|
)
|
|
84
85
|
```
|
|
85
86
|
|
|
86
|
-
|
|
87
|
+
Both strategies return valid judgments with this model. Its one-call answers arrive in a Markdown code fence, which the gem removes before parsing.
|
|
88
|
+
|
|
89
|
+
The default Luna rating request uses Chat Completions, temperature zero, reasoning disabled, `store: false`, and a four-token output limit. Tune requests with `scoring_protocol`, `temperature`, `max_output_tokens`, and `chat_provider_options`. `chat_provider_options` is merged over those defaults, so setting one key, such as `service_tier: 'flex'`, keeps reasoning disabled; set a key explicitly to change a default.
|
|
87
90
|
|
|
88
91
|
## One-call typed decisions
|
|
89
92
|
|
|
@@ -107,7 +110,7 @@ result.urgent.probability
|
|
|
107
110
|
result.department.probabilities
|
|
108
111
|
```
|
|
109
112
|
|
|
110
|
-
The model returns one JSON object containing a distribution for each question. The gem checks that every question and option is present, validates each probability, normalizes each distribution, and builds RubyLLM's typed answers. `result.raw[:reported_probabilities]` and `result.raw[:reported_totals]` preserve the model's original numbers for inspection. With an OpenAI key configured, this example calls Luna directly through Chat Completions with reasoning disabled, temperature zero, and `store: false`. The 1024-token limit suits small judgments; omit it for large question sets to use the 8192-token default. It makes one corrective retry for malformed JSON
|
|
113
|
+
The model returns one JSON object containing a distribution for each question. The gem checks that every question and option is present, validates each probability, normalizes each distribution, and builds RubyLLM's typed answers. `result.raw[:reported_probabilities]` and `result.raw[:reported_totals]` preserve the model's original numbers for inspection. With an OpenAI key configured, this example calls Luna directly through Chat Completions with reasoning disabled, temperature zero, and `store: false`. The 1024-token limit suits small judgments; omit it for large question sets to use the 8192-token default. It makes one corrective retry for malformed JSON, missing fields, or a Choice whose top probabilities tie, counting both calls in `result.tokens` and `result.raw[:attempts]`. A tie that remains raises instead of selecting the first option; set `tie_breaker: :first` to accept the first option. Set `malformed_retries: 0` to disable that retry. An invalid response after retries raises an error.
|
|
111
114
|
|
|
112
115
|
If you use OpenRouter for Luna, prioritize the provider with the lowest observed latency:
|
|
113
116
|
|
|
@@ -139,11 +142,11 @@ OpenRouter's [latency sorting](https://openrouter.ai/docs/guides/routing/provide
|
|
|
139
142
|
|
|
140
143
|
## How scoring works
|
|
141
144
|
|
|
142
|
-
The default `:ratings` strategy asks the model to rate every declared answer from 0 to 9. It makes one call per answer, with up to six calls running at once, then applies softmax to the ratings. It retries a malformed digit once. When the highest ratings tie on a Choice question, it makes a one-call JSON judgment for that question and uses its distribution. If that call also ties, it raises an error instead of choosing whichever option came first. `result.raw[:tie_breaks]` records each resolution, and token usage includes the extra call. Set `tie_breaker: :first` to use the original first-option rule, or `tie_break_max_output_tokens:` to change the tie-break response limit (default 8192).
|
|
145
|
+
The default `:ratings` strategy asks the model to rate every declared answer from 0 to 9. It makes one call per answer, with up to six calls running at once (`max_workers:` changes this), then applies softmax to the ratings. If a call fails after its retry, no new calls start, calls in flight finish, and the error is raised. It retries a malformed digit once. When the highest ratings tie on a Choice question, it makes a one-call JSON judgment for that question and uses its distribution. If that call also ties, it raises an error instead of choosing whichever option came first. `result.raw[:tie_breaks]` records each resolution, and token usage includes the extra call. Set `tie_breaker: :first` to use the original first-option rule, or `tie_break_max_output_tokens:` to change the tie-break response limit (default 8192).
|
|
143
146
|
|
|
144
147
|
The `:single_request` strategy asks for all distributions in one call. Both strategies build the same RubyLLM answer types: a `choice` returns the selected answer and distribution; a `score` returns the distribution and its weighted level; a `probability` returns the positive answer's share. The gem supports Judge's 1–255 choice options, 2–10 score levels, and multiple questions in one judgment.
|
|
145
148
|
|
|
146
|
-
These probabilities compare the answers you supplied. The `:single_request` values are reported by the chat model; the default ratings values come from softmax over its 0–9 ratings. Neither strategy establishes calibration by itself. Include an `other` or `escalate` choice when the named answers may not cover the input. `confidence` measures how concentrated the returned distribution is. Use labeled examples to set any automation thresholds.
|
|
149
|
+
These probabilities compare the answers you supplied. The `:single_request` values are reported by the chat model; the default ratings values come from softmax over its 0–9 ratings. Neither strategy establishes calibration by itself. Include an `other` or `escalate` choice when the named answers may not cover the input. `confidence` measures how concentrated the returned distribution is: 1.0 for a single option, 0.0 for a tie. Use labeled examples to set any automation thresholds.
|
|
147
150
|
|
|
148
151
|
Set optional `max_arms` and `max_input_bytes` in `provider_options` to cap work per judgment. For example, `{ max_arms: 12, max_input_bytes: 32_768 }` rejects larger requests before scoring. Token usage is aggregated across calls, and `raw` contains the ratings and model metadata.
|
|
149
152
|
|
|
@@ -211,7 +214,7 @@ Latency sorting reduced median time by 20% on AG News and 28% on SST-2. It was f
|
|
|
211
214
|
|
|
212
215
|
## Development
|
|
213
216
|
|
|
214
|
-
The gem uses RubyLLM's Judge API when it is present and the `Legacy` path otherwise. `test/llm_judge_test.rb` covers the Judge API and `test/legacy_test.rb` covers RubyLLM 1.13 through 1.16; each skips on the other side.
|
|
217
|
+
The gem uses RubyLLM's Judge API when it is present and the `Legacy` path otherwise. `test/llm_judge_test.rb` covers the Judge API and `test/legacy_test.rb` covers RubyLLM 1.13 through 1.16; each skips on the other side. `test/engine_test.rb` and `test/http_test.rb` run on every version; the HTTP tests send real RubyLLM requests to WebMock stubs and check their bodies.
|
|
215
218
|
|
|
216
219
|
```bash
|
|
217
220
|
bundle exec rake test # the latest released RubyLLM
|
|
@@ -17,8 +17,10 @@ module RubyLLM
|
|
|
17
17
|
@config = config
|
|
18
18
|
options = provider_options.transform_keys(&:to_sym)
|
|
19
19
|
allowed = %i[scoring_provider scoring_protocol temperature max_output_tokens chat_provider_options
|
|
20
|
-
max_arms max_input_bytes strategy malformed_retries tie_breaker tie_break_max_output_tokens
|
|
21
|
-
|
|
20
|
+
max_arms max_input_bytes strategy malformed_retries tie_breaker tie_break_max_output_tokens
|
|
21
|
+
max_workers]
|
|
22
|
+
unknown = options.keys - allowed
|
|
23
|
+
raise ArgumentError, "Unknown LLMJudge provider options: #{unknown.join(', ')}" unless unknown.empty?
|
|
22
24
|
|
|
23
25
|
@strategy = options.fetch(:strategy, :ratings).to_sym
|
|
24
26
|
raise ArgumentError, 'strategy must be :ratings or :single_request' unless %i[ratings single_request].include?(@strategy)
|
|
@@ -38,11 +40,16 @@ module RubyLLM
|
|
|
38
40
|
@temperature = options.fetch(:temperature, luna_defaults ? 0 : nil)
|
|
39
41
|
@max_output_tokens = options.fetch(:max_output_tokens, @strategy == :single_request ? 8192 : 4)
|
|
40
42
|
@tie_break_max_output_tokens = options.fetch(:tie_break_max_output_tokens, 8192)
|
|
41
|
-
|
|
43
|
+
chat_provider_options = options.fetch(:chat_provider_options, {})
|
|
44
|
+
raise ArgumentError, 'chat_provider_options must be a Hash' unless chat_provider_options.is_a?(Hash)
|
|
45
|
+
|
|
46
|
+
# Merged over the defaults, so tuning one option keeps the others (such as disabled reasoning).
|
|
47
|
+
@chat_provider_options = deep_merge(default_chat_options(luna_defaults), chat_provider_options)
|
|
42
48
|
@max_arms = options[:max_arms]
|
|
43
49
|
@max_input_bytes = options[:max_input_bytes]
|
|
44
50
|
@malformed_retries = options.fetch(:malformed_retries, 1)
|
|
45
|
-
|
|
51
|
+
@max_workers = options.fetch(:max_workers, MAX_WORKERS)
|
|
52
|
+
raise ArgumentError, 'max_workers must be a positive integer' unless @max_workers.is_a?(Integer) && @max_workers.positive?
|
|
46
53
|
raise ArgumentError, 'max_output_tokens must be positive' unless @max_output_tokens.is_a?(Integer) && @max_output_tokens.positive?
|
|
47
54
|
unless @tie_break_max_output_tokens.is_a?(Integer) && @tie_break_max_output_tokens.positive?
|
|
48
55
|
raise ArgumentError, 'tie_break_max_output_tokens must be positive'
|
|
@@ -70,12 +77,9 @@ module RubyLLM
|
|
|
70
77
|
}
|
|
71
78
|
chosen = answer(question, softmax(ratings))
|
|
72
79
|
if question.type == :choice && ratings.count(ratings.max) > 1 && @tie_breaker == :single_request
|
|
80
|
+
# Raises if the one-call judgment also ties, after its corrective retry.
|
|
73
81
|
judgment = judge_single_request(input, { question.name => question }, model)
|
|
74
82
|
chosen = judgment.answers.fetch(question.name)
|
|
75
|
-
probabilities = chosen.probabilities.values
|
|
76
|
-
if probabilities.count(probabilities.max) > 1
|
|
77
|
-
raise Error, "Scoring model could not resolve the tie for #{question.name}"
|
|
78
|
-
end
|
|
79
83
|
tie_breaks << { question: question.name, ratings:, judgment: }
|
|
80
84
|
end
|
|
81
85
|
[question.name, chosen]
|
|
@@ -127,7 +131,7 @@ module RubyLLM
|
|
|
127
131
|
end
|
|
128
132
|
|
|
129
133
|
def parse_single_request(content, specs)
|
|
130
|
-
parsed = JSON.parse(content)
|
|
134
|
+
parsed = JSON.parse(strip_code_fence(content))
|
|
131
135
|
raise Error, 'Scoring model returned a non-object response' unless parsed.is_a?(Hash)
|
|
132
136
|
|
|
133
137
|
reported = parsed.fetch('answers')
|
|
@@ -154,6 +158,12 @@ module RubyLLM
|
|
|
154
158
|
total = values.sum.to_f
|
|
155
159
|
raise Error, "Scoring model returned zero probability for #{question.name}" unless total.positive?
|
|
156
160
|
|
|
161
|
+
# A tied Choice would otherwise select whichever option was listed first.
|
|
162
|
+
if question.type == :choice && @tie_breaker == :single_request && values.count(values.max) > 1
|
|
163
|
+
raise Error, "Scoring model could not resolve the tie for #{question.name}; " \
|
|
164
|
+
'one option must have the highest probability'
|
|
165
|
+
end
|
|
166
|
+
|
|
157
167
|
normalization[question.name.to_s] = total
|
|
158
168
|
[question.name, answer(question, values.map { |value| value / total })]
|
|
159
169
|
end
|
|
@@ -230,28 +240,46 @@ module RubyLLM
|
|
|
230
240
|
value.is_a?(String) ? value : JSON.generate(value)
|
|
231
241
|
end
|
|
232
242
|
|
|
243
|
+
# Runs jobs on up to @max_workers threads. After a job fails, no new jobs
|
|
244
|
+
# start; calls already in flight finish, then the first error is raised.
|
|
233
245
|
def run(jobs)
|
|
234
246
|
queue = Queue.new
|
|
235
247
|
jobs.each_with_index { |job, index| queue << [index, job] }
|
|
236
248
|
results = Array.new(jobs.size)
|
|
237
|
-
|
|
249
|
+
errors = Queue.new
|
|
250
|
+
workers = Array.new([jobs.size, @max_workers].min) do
|
|
238
251
|
Thread.new do
|
|
239
|
-
|
|
240
|
-
loop do
|
|
252
|
+
while errors.empty?
|
|
241
253
|
begin
|
|
242
254
|
index, job = queue.pop(true)
|
|
243
255
|
rescue ThreadError
|
|
244
256
|
break
|
|
245
257
|
end
|
|
246
|
-
|
|
258
|
+
begin
|
|
259
|
+
results[index] = yield job
|
|
260
|
+
rescue StandardError => error
|
|
261
|
+
errors << error
|
|
262
|
+
end
|
|
247
263
|
end
|
|
248
264
|
end
|
|
249
265
|
end
|
|
250
266
|
workers.each(&:join)
|
|
251
|
-
|
|
267
|
+
raise errors.pop unless errors.empty?
|
|
268
|
+
|
|
252
269
|
results
|
|
253
270
|
end
|
|
254
271
|
|
|
272
|
+
# Some models wrap JSON in a Markdown code fence despite the instructions.
|
|
273
|
+
def strip_code_fence(content)
|
|
274
|
+
content.to_s.strip.sub(/\A```(?:json)?[ \t]*\n?/i, '').sub(/\n?```\z/, '')
|
|
275
|
+
end
|
|
276
|
+
|
|
277
|
+
def deep_merge(base, overrides)
|
|
278
|
+
base.merge(overrides) do |_key, old, new|
|
|
279
|
+
old.is_a?(Hash) && new.is_a?(Hash) ? deep_merge(old, new) : new
|
|
280
|
+
end
|
|
281
|
+
end
|
|
282
|
+
|
|
255
283
|
def default_chat_options(luna_defaults)
|
|
256
284
|
return {} unless @scoring_provider == :openai
|
|
257
285
|
|
|
@@ -292,7 +320,7 @@ module RubyLLM
|
|
|
292
320
|
chat.with_max_output_tokens(max_output_tokens)
|
|
293
321
|
chat.with_provider_options(@chat_provider_options)
|
|
294
322
|
else
|
|
295
|
-
chat.with_params(**
|
|
323
|
+
chat.with_params(**deep_merge(max_output_tokens_param(max_output_tokens), @chat_provider_options))
|
|
296
324
|
end
|
|
297
325
|
chat
|
|
298
326
|
end
|
|
@@ -331,11 +359,12 @@ module RubyLLM
|
|
|
331
359
|
def softmax(ratings)
|
|
332
360
|
peak = ratings.max
|
|
333
361
|
weights = ratings.map { |rating| Math.exp(rating - peak) }
|
|
334
|
-
|
|
362
|
+
total = weights.sum
|
|
363
|
+
weights.map { |weight| weight / total }
|
|
335
364
|
end
|
|
336
365
|
|
|
337
366
|
def concentration(probabilities)
|
|
338
|
-
return
|
|
367
|
+
return 1.0 if probabilities.size == 1
|
|
339
368
|
return 0.0 if probabilities.count(probabilities.max) > 1
|
|
340
369
|
|
|
341
370
|
entropy = -probabilities.sum { |probability| probability.zero? ? 0.0 : probability * Math.log(probability) }
|
|
@@ -26,7 +26,10 @@ module RubyLLM
|
|
|
26
26
|
end
|
|
27
27
|
|
|
28
28
|
def initialize(name, type:, instructions:, criteria:)
|
|
29
|
-
|
|
29
|
+
unless name.is_a?(String) || name.is_a?(Symbol)
|
|
30
|
+
raise ArgumentError, 'A question name must be a String or Symbol'
|
|
31
|
+
end
|
|
32
|
+
raise ArgumentError, 'A question name cannot be empty' if name.to_s.empty?
|
|
30
33
|
|
|
31
34
|
@name = name
|
|
32
35
|
@type = type
|
|
@@ -38,20 +41,58 @@ module RubyLLM
|
|
|
38
41
|
|
|
39
42
|
private
|
|
40
43
|
|
|
44
|
+
# Mirrors RubyLLM 2.1's Judge::Question validation.
|
|
41
45
|
def validate!
|
|
46
|
+
raise ArgumentError, 'Question instructions must be text, a Hash, an Array, or nil' unless description?(instructions)
|
|
47
|
+
|
|
42
48
|
case type
|
|
43
|
-
when :probability
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
49
|
+
when :probability then validate_probability!
|
|
50
|
+
when :choice then validate_choice!
|
|
51
|
+
when :score then validate_score!
|
|
52
|
+
end
|
|
53
|
+
end
|
|
54
|
+
|
|
55
|
+
def validate_probability!
|
|
56
|
+
return if criteria.nil?
|
|
57
|
+
|
|
58
|
+
unless criteria.is_a?(Hash) && (criteria.keys.map(&:to_s) - %w[yes no true false]).empty?
|
|
59
|
+
raise ArgumentError, 'Probability criteria must describe yes and no'
|
|
54
60
|
end
|
|
61
|
+
|
|
62
|
+
positive = criteria.keys.map { |key| %w[yes true].include?(key.to_s) }
|
|
63
|
+
raise ArgumentError, 'Probability criteria contain duplicate outcomes' unless positive.uniq.size == positive.size
|
|
64
|
+
|
|
65
|
+
validate_descriptions!(criteria.values)
|
|
66
|
+
end
|
|
67
|
+
|
|
68
|
+
def validate_choice!
|
|
69
|
+
raise ArgumentError, 'A choice needs a nonempty Hash of options' unless criteria.is_a?(Hash) && !criteria.empty?
|
|
70
|
+
unless criteria.keys.all? { |key| (key.is_a?(String) || key.is_a?(Symbol)) && !key.to_s.empty? }
|
|
71
|
+
raise ArgumentError, 'Choice options must have nonempty String or Symbol names'
|
|
72
|
+
end
|
|
73
|
+
|
|
74
|
+
duplicate = criteria.keys.map(&:to_s).tally.find { |_, count| count > 1 }&.first
|
|
75
|
+
raise ArgumentError, "Duplicate judgment key: #{duplicate}" if duplicate
|
|
76
|
+
|
|
77
|
+
validate_descriptions!(criteria.values)
|
|
78
|
+
end
|
|
79
|
+
|
|
80
|
+
def validate_score!
|
|
81
|
+
unless criteria.is_a?(Array) && criteria.size >= 2 && criteria.none?(&:nil?)
|
|
82
|
+
raise ArgumentError, 'A score needs at least two non-nil levels'
|
|
83
|
+
end
|
|
84
|
+
|
|
85
|
+
validate_descriptions!(criteria)
|
|
86
|
+
end
|
|
87
|
+
|
|
88
|
+
def validate_descriptions!(values)
|
|
89
|
+
return if values.all? { |value| description?(value) }
|
|
90
|
+
|
|
91
|
+
raise ArgumentError, 'Descriptions must be text, a Hash, an Array, or nil'
|
|
92
|
+
end
|
|
93
|
+
|
|
94
|
+
def description?(value)
|
|
95
|
+
value.nil? || value.is_a?(String) || value.is_a?(Hash) || value.is_a?(Array)
|
|
55
96
|
end
|
|
56
97
|
end
|
|
57
98
|
|
|
@@ -140,10 +181,15 @@ module RubyLLM
|
|
|
140
181
|
|
|
141
182
|
module_function
|
|
142
183
|
|
|
184
|
+
TOKEN_FIELDS = %i[input output cached cache_creation thinking].freeze
|
|
185
|
+
|
|
186
|
+
# Sums each field across calls; a field no call reported stays nil.
|
|
143
187
|
def aggregate_tokens(tokens)
|
|
144
188
|
tokens = tokens.compact
|
|
145
|
-
|
|
146
|
-
|
|
189
|
+
RubyLLM::Tokens.new(**TOKEN_FIELDS.to_h do |field|
|
|
190
|
+
values = tokens.filter_map { |token| token.public_send(field) }
|
|
191
|
+
[field, values.empty? ? nil : values.sum]
|
|
192
|
+
end)
|
|
147
193
|
end
|
|
148
194
|
end
|
|
149
195
|
end
|
data/lib/ruby_llm/llm_judge.rb
CHANGED
|
@@ -14,12 +14,10 @@ module RubyLLM
|
|
|
14
14
|
# Raised for invalid scoring responses. RubyLLM::Error takes (response, message)
|
|
15
15
|
# before 1.16 and (message, response:) on main, so build it for either.
|
|
16
16
|
class Error < RubyLLM::Error
|
|
17
|
+
MESSAGE_FIRST = RubyLLM::Error.instance_method(:initialize).parameters.include?(%i[key response])
|
|
18
|
+
|
|
17
19
|
def initialize(message = nil)
|
|
18
|
-
|
|
19
|
-
super(message)
|
|
20
|
-
else
|
|
21
|
-
super(nil, message)
|
|
22
|
-
end
|
|
20
|
+
MESSAGE_FIRST ? super(message) : super(nil, message)
|
|
23
21
|
end
|
|
24
22
|
end
|
|
25
23
|
end
|