ruby_llm-llm_judge 0.1.2 → 0.1.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 71f6bd771e7f6bec3a9dcdf7036ee1238110c63ee530b15d32c6dadcfb0d0eaa
4
- data.tar.gz: e54b13efadda477c3d3713c66184dc4ce01061b740e06ab52e95df0b672ae7e9
3
+ metadata.gz: c438b1693068b5c47ada88cf13b23ed8786c597ed4c349d33d8c012f8edaedd2
4
+ data.tar.gz: ca228933f0f0772c7f523c788d9391edf012138b7f0faec84466da3026b1eb9b
5
5
  SHA512:
6
- metadata.gz: 4e29ff7f1b47222dfc71f6f096cd058615d2517b7e16b29f9c2a68f094f2af95be8a99232d5ba72cc1d6b9a932fa5f9bf290e701a9744ee1a241fc6a0c710834
7
- data.tar.gz: 942f4a7da0c27d1741d3e0e2b80ae719335123d6ddf276b84b0e4d200ed315d5f839cfc1ea998e708b6b54f83d7b7e97c718f6d4a070e17d2bfe0be23e9b3541
6
+ metadata.gz: b50f0831b1e86781c9d0aaac5d9d96b69284b0e527841119df6a73f7379bee194b67b7e09bad5e72778ea35851749bff9e0787664c815f90ff83e3d653a4f6a9
7
+ data.tar.gz: c49adc022376e4b959c73d48704b293a536fd6e8a9580eaf8b495676b93c236ad7c02d24f748814c5394edb26edc7668bfee91acaf1fec03a43cbabb39153a18
data/CHANGELOG.md CHANGED
@@ -1,5 +1,12 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.1.3 (2026-09-29)
4
+
5
+ - Setting a `chat_provider_options` key to `nil` removes that field, including the Luna defaults and, before RubyLLM 2.1, the output limit field. A nested hash emptied this way is removed too.
6
+ - One-call answers given as a list of single-question objects are accepted when each question appears once. Claude Haiku 4.5 returns this shape on its first attempt, which previously cost a corrective retry on nearly every judgment.
7
+ - The one-call instructions now say that `answers` is an object keyed by question ID. Haiku no longer needs corrective retries (59 of 96 benchmark cases before), and Luna scored at least as well in a back-to-back comparison. See [the Haiku benchmark](benchmarks/haiku.md).
8
+ - Luna defaults apply to every OpenAI Luna model ID (`gpt-6-luna`, `gpt-5.6-luna`, `gpt-luna-latest`) and to OpenRouter's `openai/` Luna IDs, not only `gpt-6-luna` on OpenAI. On OpenRouter they send temperature zero and `reasoning: { enabled: false }`. With reasoning on, Luna's ratings often returned no digit. Pro variants and prefixed gateway IDs, such as `us.openai.gpt-5.6-luna` through `:openai`, are excluded.
9
+
3
10
  ## 0.1.2 (2026-09-28)
4
11
 
5
12
  - When a rating call fails, no new calls start and the error is raised once calls in flight finish. Previously the remaining calls kept running, and were billed, after `judge` raised.
data/README.md CHANGED
@@ -84,9 +84,11 @@ result = RubyLLM::LLMJudge.judge(
84
84
  )
85
85
  ```
86
86
 
87
- Both strategies return valid judgments with this model. Its one-call answers arrive in a Markdown code fence, which the gem removes before parsing.
87
+ Both strategies return valid judgments with this model. Its one-call answers arrive in a Markdown code fence, which the gem removes before parsing. On the benchmark cases it scored 56–57/64 on AG News and 29–30/32 on SST-2 at about 1.1 s median, a little below Luna. [Haiku results](benchmarks/haiku.md).
88
88
 
89
- The default Luna rating request uses Chat Completions, temperature zero, reasoning disabled, `store: false`, and a four-token output limit. Tune requests with `scoring_protocol`, `temperature`, `max_output_tokens`, and `chat_provider_options`. `chat_provider_options` is merged over those defaults, so setting one key, such as `service_tier: 'flex'`, keeps reasoning disabled; set a key explicitly to change a default.
89
+ The default Luna rating request uses Chat Completions, temperature zero, reasoning disabled, `store: false`, and a four-token output limit. Tune requests with `scoring_protocol`, `temperature`, `max_output_tokens`, and `chat_provider_options`. `chat_provider_options` is merged over those defaults, so setting one key, such as `service_tier: 'flex'`, keeps reasoning disabled; set a key explicitly to change a default. Set a key to `nil` to remove that field from the request, whether it is a default or, before RubyLLM 2.1, the output limit field. For example, an app that routes RubyLLM 1.x OpenAI chats through the Responses API can send `chat_provider_options: { max_completion_tokens: nil, max_output_tokens: 256 }`.
90
+
91
+ Luna defaults apply to each provider's own Luna model IDs: `gpt-6-luna`, `gpt-5.6-luna`, or `gpt-luna-latest` on `:openai`, which sends temperature zero, `reasoning_effort: 'none'`, and `store: false`; and `openai/gpt-6-luna` or `openai/gpt-5.6-luna` on `:openrouter`, which sends temperature zero and `reasoning: { enabled: false }`. Pro variants, other providers such as `:bedrock`, and prefixed IDs such as `us.openai.gpt-5.6-luna` sent through `:openai` to an OpenAI-compatible gateway get no Luna defaults, because their request format may differ; set `temperature:` and disable reasoning in that endpoint's format yourself. With reasoning left on, the ratings strategy's four-token output limit is spent on reasoning: on eight AG News cases, `gpt-6-luna` through OpenRouter returned no digit for all eight, and `gpt-5.6-luna` for four. RubyLLM before 2.1 sends temperature 1.0 for any OpenAI model whose ID starts with `gpt-5`, so `gpt-5.6-luna` on `:openai` runs at 1.0 there unless you pass `chat_provider_options: { temperature: 0 }`.
90
92
 
91
93
  ## One-call typed decisions
92
94
 
@@ -6,12 +6,21 @@ module RubyLLM
6
6
  module LLMJudge
7
7
  class Engine
8
8
  MAX_WORKERS = 6
9
+ # Luna model IDs by provider, which get temperature zero and reasoning disabled in that provider's
10
+ # format. Only each provider's own IDs match: an OpenAI-compatible gateway such as Bedrock, reached
11
+ # through :openai with an ID like "us.openai.gpt-5.6-luna", may use another request format. Pro
12
+ # variants are excluded: they exist to reason more.
13
+ LUNA_IDS = {
14
+ openai: /\Agpt-(?:\d+(?:\.\d+)?-)?luna(?:-latest)?\z/,
15
+ openrouter: %r{\A~?openai/gpt-(?:\d+(?:\.\d+)?-)?luna(?:-latest)?(?::\w+)?\z}
16
+ }.freeze
9
17
  SYSTEM_INSTRUCTIONS = 'Return exactly one ASCII digit 0-9. 9 means very likely to be the correct answer; 0 means very unlikely. No explanation.'
10
18
  SINGLE_REQUEST_INSTRUCTIONS = 'Answer all questions using only a JSON object with an "answers" field. ' \
11
19
  'For each question, return an object mapping every supplied option ID to a ' \
12
20
  'probability between 0 and 1. Include every option exactly once and make ' \
13
21
  'each question\'s probabilities sum to 1. Evaluate each question independently ' \
14
- 'against the same state. Return no explanations or markdown.'
22
+ 'against the same state. Return no explanations or markdown. The "answers" value ' \
23
+ 'must be an object whose keys are the question IDs.'
15
24
 
16
25
  def initialize(config:, model: DEFAULT_MODEL, provider_options: {}, scorer: nil, responder: nil)
17
26
  @config = config
@@ -31,8 +40,9 @@ module RubyLLM
31
40
  @scoring_model = model
32
41
  raise ArgumentError, 'A model is required' unless @scoring_model.is_a?(String) && !@scoring_model.empty?
33
42
 
34
- luna_defaults = @scoring_provider == :openai && @scoring_model == DEFAULT_MODEL
35
- @scoring_protocol = options.fetch(:scoring_protocol, luna_defaults ? :chat_completions : nil)
43
+ luna_defaults = LUNA_IDS.fetch(@scoring_provider, nil)&.match?(@scoring_model) || false
44
+ @scoring_protocol = options.fetch(:scoring_protocol,
45
+ luna_defaults && @scoring_provider == :openai ? :chat_completions : nil)
36
46
  # Before 2.1, RubyLLM chats use each provider's one chat API (Chat Completions for OpenAI).
37
47
  if !NATIVE && ![nil, :chat_completions].include?(@scoring_protocol&.to_sym)
38
48
  raise ArgumentError, 'scoring_protocol requires RubyLLM 2.1'
@@ -135,6 +145,11 @@ module RubyLLM
135
145
  raise Error, 'Scoring model returned a non-object response' unless parsed.is_a?(Hash)
136
146
 
137
147
  reported = parsed.fetch('answers')
148
+ # Some models mirror the prompt's question list: [{"id": {...}}, ...].
149
+ if reported.is_a?(Array) && reported.all? { |item| item.is_a?(Hash) }
150
+ ids = reported.flat_map(&:keys)
151
+ reported = reported.reduce({}, :merge) if ids.uniq.size == ids.size
152
+ end
138
153
  expected_questions = specs.map { |question, _| question.name.to_s }
139
154
  unless reported.is_a?(Hash) && reported.keys.sort == expected_questions.sort
140
155
  actual = reported.is_a?(Hash) ? reported.keys : reported.class.name
@@ -280,10 +295,13 @@ module RubyLLM
280
295
  end
281
296
  end
282
297
 
298
+ # Luna rates poorly with reasoning on: the output limit is spent before the answer.
283
299
  def default_chat_options(luna_defaults)
284
- return {} unless @scoring_provider == :openai
285
-
286
- luna_defaults ? { store: false, reasoning_effort: 'none' } : { store: false }
300
+ case @scoring_provider
301
+ when :openai then luna_defaults ? { store: false, reasoning_effort: 'none' } : { store: false }
302
+ when :openrouter then luna_defaults ? { reasoning: { enabled: false } } : {}
303
+ else {}
304
+ end
287
305
  end
288
306
 
289
307
  def score_with_model(prompt)
@@ -318,13 +336,27 @@ module RubyLLM
318
336
  chat.with_temperature(@temperature) unless @temperature.nil?
319
337
  if NATIVE
320
338
  chat.with_max_output_tokens(max_output_tokens)
321
- chat.with_provider_options(@chat_provider_options)
339
+ chat.with_provider_options(deep_compact(@chat_provider_options))
322
340
  else
323
- chat.with_params(**deep_merge(max_output_tokens_param(max_output_tokens), @chat_provider_options))
341
+ chat.with_params(**deep_compact(deep_merge(max_output_tokens_param(max_output_tokens), @chat_provider_options)))
324
342
  end
325
343
  chat
326
344
  end
327
345
 
346
+ # A nil in chat_provider_options removes that field: a default, or before 2.1
347
+ # the output limit field. A hash emptied by removal is removed too.
348
+ def deep_compact(hash)
349
+ hash.each_with_object({}) do |(key, value), compacted|
350
+ next if value.nil?
351
+
352
+ if value.is_a?(Hash) && !value.empty?
353
+ value = deep_compact(value)
354
+ next if value.empty?
355
+ end
356
+ compacted[key] = value
357
+ end
358
+ end
359
+
328
360
  # The request field each provider reads for the output limit, as RubyLLM 2.1 renders it.
329
361
  def max_output_tokens_param(limit)
330
362
  case @scoring_provider
@@ -2,6 +2,6 @@
2
2
 
3
3
  module RubyLLM
4
4
  module LLMJudge
5
- VERSION = '0.1.2'
5
+ VERSION = '0.1.3'
6
6
  end
7
7
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: ruby_llm-llm_judge
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.1.2
4
+ version: 0.1.3
5
5
  platform: ruby
6
6
  authors:
7
7
  - JP Camara