ruby_llm-llm_judge 0.1.1 → 0.1.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 451771cadd6d75e7ea6361b43814d186bc0b9e3aeb0dea5f128fdc3c4a68e982
4
- data.tar.gz: ca97b903a28238275417a96551abbf440198eb2dea95d52196cfb2e5d1b0d173
3
+ metadata.gz: c438b1693068b5c47ada88cf13b23ed8786c597ed4c349d33d8c012f8edaedd2
4
+ data.tar.gz: ca228933f0f0772c7f523c788d9391edf012138b7f0faec84466da3026b1eb9b
5
5
  SHA512:
6
- metadata.gz: ff30bd52e7b0dcb79f8804458e622d725ad3570351e00d180efb9ed040d4bb07c8551c37024f1730ca876125f52513e0e28a557fb9ed8f9d22211b286266aa6d
7
- data.tar.gz: 25259a68bba8ffc259d343f076b7962788f236143ca8dabe43c03e626181f9bac24bfb50a98b50951130478be94951c56944775072614732c1341aa3218626eb
6
+ metadata.gz: b50f0831b1e86781c9d0aaac5d9d96b69284b0e527841119df6a73f7379bee194b67b7e09bad5e72778ea35851749bff9e0787664c815f90ff83e3d653a4f6a9
7
+ data.tar.gz: c49adc022376e4b959c73d48704b293a536fd6e8a9580eaf8b495676b93c236ad7c02d24f748814c5394edb26edc7668bfee91acaf1fec03a43cbabb39153a18
data/CHANGELOG.md CHANGED
@@ -1,5 +1,23 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.1.3 (2026-09-29)
4
+
5
+ - Setting a `chat_provider_options` key to `nil` removes that field, including the Luna defaults and, before RubyLLM 2.1, the output limit field. A nested hash emptied this way is removed too.
6
+ - One-call answers given as a list of single-question objects are accepted when each question appears once. Claude Haiku 4.5 returns this shape on its first attempt, which previously cost a corrective retry on nearly every judgment.
7
+ - The one-call instructions now say that `answers` is an object keyed by question ID. Haiku no longer needs corrective retries (59 of 96 benchmark cases before), and Luna scored at least as well in a back-to-back comparison. See [the Haiku benchmark](benchmarks/haiku.md).
8
+ - Luna defaults apply to every OpenAI Luna model ID (`gpt-6-luna`, `gpt-5.6-luna`, `gpt-luna-latest`) and to OpenRouter's `openai/` Luna IDs, not only `gpt-6-luna` on OpenAI. On OpenRouter they send temperature zero and `reasoning: { enabled: false }`. With reasoning on, Luna's ratings often returned no digit. Pro variants and prefixed gateway IDs, such as `us.openai.gpt-5.6-luna` through `:openai`, are excluded.
9
+
10
+ ## 0.1.2 (2026-09-28)
11
+
12
+ - When a rating call fails, no new calls start and the error is raised once calls in flight finish. Previously the remaining calls kept running, and were billed, after `judge` raised.
13
+ - `chat_provider_options` is merged over the defaults instead of replacing them. Setting one option no longer re-enables reasoning or drops `store: false` for Luna.
14
+ - One-call JSON inside a Markdown code fence is accepted. Claude Haiku 4.5 now works with `strategy: :single_request`.
15
+ - A one-call Choice whose top probabilities tie gets the corrective retry, then raises, instead of selecting the first option. `tie_breaker: :first` keeps the old behavior.
16
+ - New `max_workers:` option (default 6) for concurrent rating calls.
17
+ - Unknown provider options are named in the error.
18
+ - A one-option Choice reports confidence 1.0.
19
+ - RubyLLM 1.13 through 1.16: question validation matches RubyLLM 2.1 (duplicate or empty option names, duplicate probability outcomes, description types), and token totals keep cached and thinking tokens.
20
+
3
21
  ## 0.1.1 (2026-09-28)
4
22
 
5
23
  - Fix invalid scoring responses raising `NoMethodError` on RubyLLM 1.13, which also skipped the corrective retry. They now raise `RubyLLM::LLMJudge::Error`, a `RubyLLM::Error`, on every supported RubyLLM version.
data/README.md CHANGED
@@ -19,7 +19,7 @@ The gem also runs on RubyLLM 1.13 through 1.16, which predate the Judge API. The
19
19
  What differs before 2.1:
20
20
 
21
21
  - Results are `RubyLLM::LLMJudge::Legacy` objects, not `RubyLLM::Judgment`, `RubyLLM::Choice`, and so on. Avoid class checks until you upgrade.
22
- - There is no `RubyLLM.judge`, `RubyLLM::Judge` class DSL, `cost`, `context:` instrumentation, or `metadata:`. Passing `metadata:` raises.
22
+ - There is no `RubyLLM.judge`, `RubyLLM::Judge` class DSL, `cost`, instrumentation, or `metadata:`. Passing `metadata:` raises. `context:` is accepted and supplies its configuration.
23
23
  - `scoring_protocol` accepts only `:chat_completions`, the API RubyLLM 1.x uses for OpenAI.
24
24
  - The output limit and `chat_provider_options` are sent with `with_params`, using the field each provider expects (`max_completion_tokens` for OpenAI and Azure, `generationConfig.maxOutputTokens` for Gemini and Vertex AI, `inferenceConfig.maxTokens` for Bedrock, and `max_tokens` otherwise).
25
25
 
@@ -70,20 +70,25 @@ TicketTriage.judge('Please refund today.').urgent.probability
70
70
 
71
71
  ## Choose a model
72
72
 
73
- Pass the chat model as `model:` and its provider as `scoring_provider`:
73
+ Pass the chat model as `model:` and its provider as `scoring_provider`. For example, Claude Haiku 4.5 through OpenRouter:
74
74
 
75
75
  ```ruby
76
76
  result = RubyLLM::LLMJudge.judge(
77
77
  'Please refund the duplicate charge today.',
78
- model: 'claude-haiku-4-5',
78
+ model: 'anthropic/claude-haiku-4.5',
79
79
  questions: { urgent: { type: :probability, instructions: 'Does this need attention today?' } },
80
80
  provider_options: {
81
- scoring_provider: :anthropic
81
+ scoring_provider: :openrouter,
82
+ temperature: 0
82
83
  }
83
84
  )
84
85
  ```
85
86
 
86
- The default Luna rating request uses Chat Completions, temperature zero, reasoning disabled, `store: false`, and a four-token output limit. You can tune requests with `scoring_protocol`, `temperature`, `max_output_tokens`, and `chat_provider_options`.
87
+ Both strategies return valid judgments with this model. Its one-call answers arrive in a Markdown code fence, which the gem removes before parsing. On the benchmark cases it scored 56–57/64 on AG News and 29–30/32 on SST-2 at about 1.1 s median, a little below Luna. [Haiku results](benchmarks/haiku.md).
88
+
89
+ The default Luna rating request uses Chat Completions, temperature zero, reasoning disabled, `store: false`, and a four-token output limit. Tune requests with `scoring_protocol`, `temperature`, `max_output_tokens`, and `chat_provider_options`. `chat_provider_options` is merged over those defaults, so setting one key, such as `service_tier: 'flex'`, keeps reasoning disabled; set a key explicitly to change a default. Set a key to `nil` to remove that field from the request, whether it is a default or, before RubyLLM 2.1, the output limit field. For example, an app that routes RubyLLM 1.x OpenAI chats through the Responses API can send `chat_provider_options: { max_completion_tokens: nil, max_output_tokens: 256 }`.
90
+
91
+ Luna defaults apply to each provider's own Luna model IDs: `gpt-6-luna`, `gpt-5.6-luna`, or `gpt-luna-latest` on `:openai`, which sends temperature zero, `reasoning_effort: 'none'`, and `store: false`; and `openai/gpt-6-luna` or `openai/gpt-5.6-luna` on `:openrouter`, which sends temperature zero and `reasoning: { enabled: false }`. Pro variants, other providers such as `:bedrock`, and prefixed IDs such as `us.openai.gpt-5.6-luna` sent through `:openai` to an OpenAI-compatible gateway get no Luna defaults, because their request format may differ; set `temperature:` and disable reasoning in that endpoint's format yourself. With reasoning left on, the ratings strategy's four-token output limit is spent on reasoning: on eight AG News cases, `gpt-6-luna` through OpenRouter returned no digit for all eight, and `gpt-5.6-luna` for four. RubyLLM before 2.1 sends temperature 1.0 for any OpenAI model whose ID starts with `gpt-5`, so `gpt-5.6-luna` on `:openai` runs at 1.0 there unless you pass `chat_provider_options: { temperature: 0 }`.
87
92
 
88
93
  ## One-call typed decisions
89
94
 
@@ -107,7 +112,7 @@ result.urgent.probability
107
112
  result.department.probabilities
108
113
  ```
109
114
 
110
- The model returns one JSON object containing a distribution for each question. The gem checks that every question and option is present, validates each probability, normalizes each distribution, and builds RubyLLM's typed answers. `result.raw[:reported_probabilities]` and `result.raw[:reported_totals]` preserve the model's original numbers for inspection. With an OpenAI key configured, this example calls Luna directly through Chat Completions with reasoning disabled, temperature zero, and `store: false`. The 1024-token limit suits small judgments; omit it for large question sets to use the 8192-token default. It makes one corrective retry for malformed JSON or missing fields, counting both calls in `result.tokens` and `result.raw[:attempts]`. Set `malformed_retries: 0` to disable that retry. An invalid response after retries raises an error.
115
+ The model returns one JSON object containing a distribution for each question. The gem checks that every question and option is present, validates each probability, normalizes each distribution, and builds RubyLLM's typed answers. `result.raw[:reported_probabilities]` and `result.raw[:reported_totals]` preserve the model's original numbers for inspection. With an OpenAI key configured, this example calls Luna directly through Chat Completions with reasoning disabled, temperature zero, and `store: false`. The 1024-token limit suits small judgments; omit it for large question sets to use the 8192-token default. It makes one corrective retry for malformed JSON, missing fields, or a Choice whose top probabilities tie, counting both calls in `result.tokens` and `result.raw[:attempts]`. A tie that remains raises instead of selecting the first option; set `tie_breaker: :first` to accept the first option. Set `malformed_retries: 0` to disable that retry. An invalid response after retries raises an error.
111
116
 
112
117
  If you use OpenRouter for Luna, prioritize the provider with the lowest observed latency:
113
118
 
@@ -139,11 +144,11 @@ OpenRouter's [latency sorting](https://openrouter.ai/docs/guides/routing/provide
139
144
 
140
145
  ## How scoring works
141
146
 
142
- The default `:ratings` strategy asks the model to rate every declared answer from 0 to 9. It makes one call per answer, with up to six calls running at once, then applies softmax to the ratings. It retries a malformed digit once. When the highest ratings tie on a Choice question, it makes a one-call JSON judgment for that question and uses its distribution. If that call also ties, it raises an error instead of choosing whichever option came first. `result.raw[:tie_breaks]` records each resolution, and token usage includes the extra call. Set `tie_breaker: :first` to use the original first-option rule, or `tie_break_max_output_tokens:` to change the tie-break response limit (default 8192).
147
+ The default `:ratings` strategy asks the model to rate every declared answer from 0 to 9. It makes one call per answer, with up to six calls running at once (`max_workers:` changes this), then applies softmax to the ratings. If a call fails after its retry, no new calls start, calls in flight finish, and the error is raised. It retries a malformed digit once. When the highest ratings tie on a Choice question, it makes a one-call JSON judgment for that question and uses its distribution. If that call also ties, it raises an error instead of choosing whichever option came first. `result.raw[:tie_breaks]` records each resolution, and token usage includes the extra call. Set `tie_breaker: :first` to use the original first-option rule, or `tie_break_max_output_tokens:` to change the tie-break response limit (default 8192).
143
148
 
144
149
  The `:single_request` strategy asks for all distributions in one call. Both strategies build the same RubyLLM answer types: a `choice` returns the selected answer and distribution; a `score` returns the distribution and its weighted level; a `probability` returns the positive answer's share. The gem supports Judge's 1–255 choice options, 2–10 score levels, and multiple questions in one judgment.
145
150
 
146
- These probabilities compare the answers you supplied. The `:single_request` values are reported by the chat model; the default ratings values come from softmax over its 0–9 ratings. Neither strategy establishes calibration by itself. Include an `other` or `escalate` choice when the named answers may not cover the input. `confidence` measures how concentrated the returned distribution is. Use labeled examples to set any automation thresholds.
151
+ These probabilities compare the answers you supplied. The `:single_request` values are reported by the chat model; the default ratings values come from softmax over its 0–9 ratings. Neither strategy establishes calibration by itself. Include an `other` or `escalate` choice when the named answers may not cover the input. `confidence` measures how concentrated the returned distribution is: 1.0 for a single option, 0.0 for a tie. Use labeled examples to set any automation thresholds.
147
152
 
148
153
  Set optional `max_arms` and `max_input_bytes` in `provider_options` to cap work per judgment. For example, `{ max_arms: 12, max_input_bytes: 32_768 }` rejects larger requests before scoring. Token usage is aggregated across calls, and `raw` contains the ratings and model metadata.
149
154
 
@@ -211,7 +216,7 @@ Latency sorting reduced median time by 20% on AG News and 28% on SST-2. It was f
211
216
 
212
217
  ## Development
213
218
 
214
- The gem uses RubyLLM's Judge API when it is present and the `Legacy` path otherwise. `test/llm_judge_test.rb` covers the Judge API and `test/legacy_test.rb` covers RubyLLM 1.13 through 1.16; each skips on the other side.
219
+ The gem uses RubyLLM's Judge API when it is present and the `Legacy` path otherwise. `test/llm_judge_test.rb` covers the Judge API and `test/legacy_test.rb` covers RubyLLM 1.13 through 1.16; each skips on the other side. `test/engine_test.rb` and `test/http_test.rb` run on every version; the HTTP tests send real RubyLLM requests to WebMock stubs and check their bodies.
215
220
 
216
221
  ```bash
217
222
  bundle exec rake test # the latest released RubyLLM
@@ -6,19 +6,30 @@ module RubyLLM
6
6
  module LLMJudge
7
7
  class Engine
8
8
  MAX_WORKERS = 6
9
+ # Luna model IDs by provider, which get temperature zero and reasoning disabled in that provider's
10
+ # format. Only each provider's own IDs match: an OpenAI-compatible gateway such as Bedrock, reached
11
+ # through :openai with an ID like "us.openai.gpt-5.6-luna", may use another request format. Pro
12
+ # variants are excluded: they exist to reason more.
13
+ LUNA_IDS = {
14
+ openai: /\Agpt-(?:\d+(?:\.\d+)?-)?luna(?:-latest)?\z/,
15
+ openrouter: %r{\A~?openai/gpt-(?:\d+(?:\.\d+)?-)?luna(?:-latest)?(?::\w+)?\z}
16
+ }.freeze
9
17
  SYSTEM_INSTRUCTIONS = 'Return exactly one ASCII digit 0-9. 9 means very likely to be the correct answer; 0 means very unlikely. No explanation.'
10
18
  SINGLE_REQUEST_INSTRUCTIONS = 'Answer all questions using only a JSON object with an "answers" field. ' \
11
19
  'For each question, return an object mapping every supplied option ID to a ' \
12
20
  'probability between 0 and 1. Include every option exactly once and make ' \
13
21
  'each question\'s probabilities sum to 1. Evaluate each question independently ' \
14
- 'against the same state. Return no explanations or markdown.'
22
+ 'against the same state. Return no explanations or markdown. The "answers" value ' \
23
+ 'must be an object whose keys are the question IDs.'
15
24
 
16
25
  def initialize(config:, model: DEFAULT_MODEL, provider_options: {}, scorer: nil, responder: nil)
17
26
  @config = config
18
27
  options = provider_options.transform_keys(&:to_sym)
19
28
  allowed = %i[scoring_provider scoring_protocol temperature max_output_tokens chat_provider_options
20
- max_arms max_input_bytes strategy malformed_retries tie_breaker tie_break_max_output_tokens]
21
- raise ArgumentError, 'Unknown LLMJudge provider options' unless (options.keys - allowed).empty?
29
+ max_arms max_input_bytes strategy malformed_retries tie_breaker tie_break_max_output_tokens
30
+ max_workers]
31
+ unknown = options.keys - allowed
32
+ raise ArgumentError, "Unknown LLMJudge provider options: #{unknown.join(', ')}" unless unknown.empty?
22
33
 
23
34
  @strategy = options.fetch(:strategy, :ratings).to_sym
24
35
  raise ArgumentError, 'strategy must be :ratings or :single_request' unless %i[ratings single_request].include?(@strategy)
@@ -29,8 +40,9 @@ module RubyLLM
29
40
  @scoring_model = model
30
41
  raise ArgumentError, 'A model is required' unless @scoring_model.is_a?(String) && !@scoring_model.empty?
31
42
 
32
- luna_defaults = @scoring_provider == :openai && @scoring_model == DEFAULT_MODEL
33
- @scoring_protocol = options.fetch(:scoring_protocol, luna_defaults ? :chat_completions : nil)
43
+ luna_defaults = LUNA_IDS.fetch(@scoring_provider, nil)&.match?(@scoring_model) || false
44
+ @scoring_protocol = options.fetch(:scoring_protocol,
45
+ luna_defaults && @scoring_provider == :openai ? :chat_completions : nil)
34
46
  # Before 2.1, RubyLLM chats use each provider's one chat API (Chat Completions for OpenAI).
35
47
  if !NATIVE && ![nil, :chat_completions].include?(@scoring_protocol&.to_sym)
36
48
  raise ArgumentError, 'scoring_protocol requires RubyLLM 2.1'
@@ -38,11 +50,16 @@ module RubyLLM
38
50
  @temperature = options.fetch(:temperature, luna_defaults ? 0 : nil)
39
51
  @max_output_tokens = options.fetch(:max_output_tokens, @strategy == :single_request ? 8192 : 4)
40
52
  @tie_break_max_output_tokens = options.fetch(:tie_break_max_output_tokens, 8192)
41
- @chat_provider_options = options.fetch(:chat_provider_options, default_chat_options(luna_defaults))
53
+ chat_provider_options = options.fetch(:chat_provider_options, {})
54
+ raise ArgumentError, 'chat_provider_options must be a Hash' unless chat_provider_options.is_a?(Hash)
55
+
56
+ # Merged over the defaults, so tuning one option keeps the others (such as disabled reasoning).
57
+ @chat_provider_options = deep_merge(default_chat_options(luna_defaults), chat_provider_options)
42
58
  @max_arms = options[:max_arms]
43
59
  @max_input_bytes = options[:max_input_bytes]
44
60
  @malformed_retries = options.fetch(:malformed_retries, 1)
45
- raise ArgumentError, 'chat_provider_options must be a Hash' unless @chat_provider_options.is_a?(Hash)
61
+ @max_workers = options.fetch(:max_workers, MAX_WORKERS)
62
+ raise ArgumentError, 'max_workers must be a positive integer' unless @max_workers.is_a?(Integer) && @max_workers.positive?
46
63
  raise ArgumentError, 'max_output_tokens must be positive' unless @max_output_tokens.is_a?(Integer) && @max_output_tokens.positive?
47
64
  unless @tie_break_max_output_tokens.is_a?(Integer) && @tie_break_max_output_tokens.positive?
48
65
  raise ArgumentError, 'tie_break_max_output_tokens must be positive'
@@ -70,12 +87,9 @@ module RubyLLM
70
87
  }
71
88
  chosen = answer(question, softmax(ratings))
72
89
  if question.type == :choice && ratings.count(ratings.max) > 1 && @tie_breaker == :single_request
90
+ # Raises if the one-call judgment also ties, after its corrective retry.
73
91
  judgment = judge_single_request(input, { question.name => question }, model)
74
92
  chosen = judgment.answers.fetch(question.name)
75
- probabilities = chosen.probabilities.values
76
- if probabilities.count(probabilities.max) > 1
77
- raise Error, "Scoring model could not resolve the tie for #{question.name}"
78
- end
79
93
  tie_breaks << { question: question.name, ratings:, judgment: }
80
94
  end
81
95
  [question.name, chosen]
@@ -127,10 +141,15 @@ module RubyLLM
127
141
  end
128
142
 
129
143
  def parse_single_request(content, specs)
130
- parsed = JSON.parse(content)
144
+ parsed = JSON.parse(strip_code_fence(content))
131
145
  raise Error, 'Scoring model returned a non-object response' unless parsed.is_a?(Hash)
132
146
 
133
147
  reported = parsed.fetch('answers')
148
+ # Some models mirror the prompt's question list: [{"id": {...}}, ...].
149
+ if reported.is_a?(Array) && reported.all? { |item| item.is_a?(Hash) }
150
+ ids = reported.flat_map(&:keys)
151
+ reported = reported.reduce({}, :merge) if ids.uniq.size == ids.size
152
+ end
134
153
  expected_questions = specs.map { |question, _| question.name.to_s }
135
154
  unless reported.is_a?(Hash) && reported.keys.sort == expected_questions.sort
136
155
  actual = reported.is_a?(Hash) ? reported.keys : reported.class.name
@@ -154,6 +173,12 @@ module RubyLLM
154
173
  total = values.sum.to_f
155
174
  raise Error, "Scoring model returned zero probability for #{question.name}" unless total.positive?
156
175
 
176
+ # A tied Choice would otherwise select whichever option was listed first.
177
+ if question.type == :choice && @tie_breaker == :single_request && values.count(values.max) > 1
178
+ raise Error, "Scoring model could not resolve the tie for #{question.name}; " \
179
+ 'one option must have the highest probability'
180
+ end
181
+
157
182
  normalization[question.name.to_s] = total
158
183
  [question.name, answer(question, values.map { |value| value / total })]
159
184
  end
@@ -230,32 +255,53 @@ module RubyLLM
230
255
  value.is_a?(String) ? value : JSON.generate(value)
231
256
  end
232
257
 
258
+ # Runs jobs on up to @max_workers threads. After a job fails, no new jobs
259
+ # start; calls already in flight finish, then the first error is raised.
233
260
  def run(jobs)
234
261
  queue = Queue.new
235
262
  jobs.each_with_index { |job, index| queue << [index, job] }
236
263
  results = Array.new(jobs.size)
237
- workers = Array.new([jobs.size, MAX_WORKERS].min) do
264
+ errors = Queue.new
265
+ workers = Array.new([jobs.size, @max_workers].min) do
238
266
  Thread.new do
239
- Thread.current.report_on_exception = false
240
- loop do
267
+ while errors.empty?
241
268
  begin
242
269
  index, job = queue.pop(true)
243
270
  rescue ThreadError
244
271
  break
245
272
  end
246
- results[index] = yield job
273
+ begin
274
+ results[index] = yield job
275
+ rescue StandardError => error
276
+ errors << error
277
+ end
247
278
  end
248
279
  end
249
280
  end
250
281
  workers.each(&:join)
251
- workers.each(&:value)
282
+ raise errors.pop unless errors.empty?
283
+
252
284
  results
253
285
  end
254
286
 
255
- def default_chat_options(luna_defaults)
256
- return {} unless @scoring_provider == :openai
287
+ # Some models wrap JSON in a Markdown code fence despite the instructions.
288
+ def strip_code_fence(content)
289
+ content.to_s.strip.sub(/\A```(?:json)?[ \t]*\n?/i, '').sub(/\n?```\z/, '')
290
+ end
257
291
 
258
- luna_defaults ? { store: false, reasoning_effort: 'none' } : { store: false }
292
+ def deep_merge(base, overrides)
293
+ base.merge(overrides) do |_key, old, new|
294
+ old.is_a?(Hash) && new.is_a?(Hash) ? deep_merge(old, new) : new
295
+ end
296
+ end
297
+
298
+ # Luna rates poorly with reasoning on: the output limit is spent before the answer.
299
+ def default_chat_options(luna_defaults)
300
+ case @scoring_provider
301
+ when :openai then luna_defaults ? { store: false, reasoning_effort: 'none' } : { store: false }
302
+ when :openrouter then luna_defaults ? { reasoning: { enabled: false } } : {}
303
+ else {}
304
+ end
259
305
  end
260
306
 
261
307
  def score_with_model(prompt)
@@ -290,13 +336,27 @@ module RubyLLM
290
336
  chat.with_temperature(@temperature) unless @temperature.nil?
291
337
  if NATIVE
292
338
  chat.with_max_output_tokens(max_output_tokens)
293
- chat.with_provider_options(@chat_provider_options)
339
+ chat.with_provider_options(deep_compact(@chat_provider_options))
294
340
  else
295
- chat.with_params(**RubyLLM::Utils.deep_merge(max_output_tokens_param(max_output_tokens), @chat_provider_options))
341
+ chat.with_params(**deep_compact(deep_merge(max_output_tokens_param(max_output_tokens), @chat_provider_options)))
296
342
  end
297
343
  chat
298
344
  end
299
345
 
346
+ # A nil in chat_provider_options removes that field: a default, or before 2.1
347
+ # the output limit field. A hash emptied by removal is removed too.
348
+ def deep_compact(hash)
349
+ hash.each_with_object({}) do |(key, value), compacted|
350
+ next if value.nil?
351
+
352
+ if value.is_a?(Hash) && !value.empty?
353
+ value = deep_compact(value)
354
+ next if value.empty?
355
+ end
356
+ compacted[key] = value
357
+ end
358
+ end
359
+
300
360
  # The request field each provider reads for the output limit, as RubyLLM 2.1 renders it.
301
361
  def max_output_tokens_param(limit)
302
362
  case @scoring_provider
@@ -331,11 +391,12 @@ module RubyLLM
331
391
  def softmax(ratings)
332
392
  peak = ratings.max
333
393
  weights = ratings.map { |rating| Math.exp(rating - peak) }
334
- weights.map { |weight| weight / weights.sum }
394
+ total = weights.sum
395
+ weights.map { |weight| weight / total }
335
396
  end
336
397
 
337
398
  def concentration(probabilities)
338
- return 0.0 if probabilities.one?
399
+ return 1.0 if probabilities.size == 1
339
400
  return 0.0 if probabilities.count(probabilities.max) > 1
340
401
 
341
402
  entropy = -probabilities.sum { |probability| probability.zero? ? 0.0 : probability * Math.log(probability) }
@@ -26,7 +26,10 @@ module RubyLLM
26
26
  end
27
27
 
28
28
  def initialize(name, type:, instructions:, criteria:)
29
- raise ArgumentError, 'A question name must be a String or Symbol' unless name.is_a?(String) || name.is_a?(Symbol)
29
+ unless name.is_a?(String) || name.is_a?(Symbol)
30
+ raise ArgumentError, 'A question name must be a String or Symbol'
31
+ end
32
+ raise ArgumentError, 'A question name cannot be empty' if name.to_s.empty?
30
33
 
31
34
  @name = name
32
35
  @type = type
@@ -38,20 +41,58 @@ module RubyLLM
38
41
 
39
42
  private
40
43
 
44
+ # Mirrors RubyLLM 2.1's Judge::Question validation.
41
45
  def validate!
46
+ raise ArgumentError, 'Question instructions must be text, a Hash, an Array, or nil' unless description?(instructions)
47
+
42
48
  case type
43
- when :probability
44
- return if criteria.nil?
45
- unless criteria.is_a?(Hash) && (criteria.keys.map(&:to_s) - %w[yes no true false]).empty?
46
- raise ArgumentError, 'Probability criteria must describe yes and no'
47
- end
48
- when :choice
49
- raise ArgumentError, 'A choice needs a nonempty Hash of options' unless criteria.is_a?(Hash) && !criteria.empty?
50
- when :score
51
- unless criteria.is_a?(Array) && criteria.size >= 2 && criteria.none?(&:nil?)
52
- raise ArgumentError, 'A score needs at least two non-nil levels'
53
- end
49
+ when :probability then validate_probability!
50
+ when :choice then validate_choice!
51
+ when :score then validate_score!
52
+ end
53
+ end
54
+
55
+ def validate_probability!
56
+ return if criteria.nil?
57
+
58
+ unless criteria.is_a?(Hash) && (criteria.keys.map(&:to_s) - %w[yes no true false]).empty?
59
+ raise ArgumentError, 'Probability criteria must describe yes and no'
54
60
  end
61
+
62
+ positive = criteria.keys.map { |key| %w[yes true].include?(key.to_s) }
63
+ raise ArgumentError, 'Probability criteria contain duplicate outcomes' unless positive.uniq.size == positive.size
64
+
65
+ validate_descriptions!(criteria.values)
66
+ end
67
+
68
+ def validate_choice!
69
+ raise ArgumentError, 'A choice needs a nonempty Hash of options' unless criteria.is_a?(Hash) && !criteria.empty?
70
+ unless criteria.keys.all? { |key| (key.is_a?(String) || key.is_a?(Symbol)) && !key.to_s.empty? }
71
+ raise ArgumentError, 'Choice options must have nonempty String or Symbol names'
72
+ end
73
+
74
+ duplicate = criteria.keys.map(&:to_s).tally.find { |_, count| count > 1 }&.first
75
+ raise ArgumentError, "Duplicate judgment key: #{duplicate}" if duplicate
76
+
77
+ validate_descriptions!(criteria.values)
78
+ end
79
+
80
+ def validate_score!
81
+ unless criteria.is_a?(Array) && criteria.size >= 2 && criteria.none?(&:nil?)
82
+ raise ArgumentError, 'A score needs at least two non-nil levels'
83
+ end
84
+
85
+ validate_descriptions!(criteria)
86
+ end
87
+
88
+ def validate_descriptions!(values)
89
+ return if values.all? { |value| description?(value) }
90
+
91
+ raise ArgumentError, 'Descriptions must be text, a Hash, an Array, or nil'
92
+ end
93
+
94
+ def description?(value)
95
+ value.nil? || value.is_a?(String) || value.is_a?(Hash) || value.is_a?(Array)
55
96
  end
56
97
  end
57
98
 
@@ -140,10 +181,15 @@ module RubyLLM
140
181
 
141
182
  module_function
142
183
 
184
+ TOKEN_FIELDS = %i[input output cached cache_creation thinking].freeze
185
+
186
+ # Sums each field across calls; a field no call reported stays nil.
143
187
  def aggregate_tokens(tokens)
144
188
  tokens = tokens.compact
145
- sum = ->(reader) { tokens.sum { |token| token.public_send(reader) || 0 } }
146
- RubyLLM::Tokens.new(input: sum.(:input), output: sum.(:output))
189
+ RubyLLM::Tokens.new(**TOKEN_FIELDS.to_h do |field|
190
+ values = tokens.filter_map { |token| token.public_send(field) }
191
+ [field, values.empty? ? nil : values.sum]
192
+ end)
147
193
  end
148
194
  end
149
195
  end
@@ -2,6 +2,6 @@
2
2
 
3
3
  module RubyLLM
4
4
  module LLMJudge
5
- VERSION = '0.1.1'
5
+ VERSION = '0.1.3'
6
6
  end
7
7
  end
@@ -14,12 +14,10 @@ module RubyLLM
14
14
  # Raised for invalid scoring responses. RubyLLM::Error takes (response, message)
15
15
  # before 1.16 and (message, response:) on main, so build it for either.
16
16
  class Error < RubyLLM::Error
17
+ MESSAGE_FIRST = RubyLLM::Error.instance_method(:initialize).parameters.include?(%i[key response])
18
+
17
19
  def initialize(message = nil)
18
- if RubyLLM::Error.instance_method(:initialize).parameters.include?(%i[key response])
19
- super(message)
20
- else
21
- super(nil, message)
22
- end
20
+ MESSAGE_FIRST ? super(message) : super(nil, message)
23
21
  end
24
22
  end
25
23
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: ruby_llm-llm_judge
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.1.1
4
+ version: 0.1.3
5
5
  platform: ruby
6
6
  authors:
7
7
  - JP Camara