ruby_llm-llm_judge 0.1.2 → 0.1.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +7 -0
- data/README.md +4 -2
- data/lib/ruby_llm/llm_judge/engine.rb +40 -8
- data/lib/ruby_llm/llm_judge/version.rb +1 -1
- metadata +1 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: c438b1693068b5c47ada88cf13b23ed8786c597ed4c349d33d8c012f8edaedd2
|
|
4
|
+
data.tar.gz: ca228933f0f0772c7f523c788d9391edf012138b7f0faec84466da3026b1eb9b
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: b50f0831b1e86781c9d0aaac5d9d96b69284b0e527841119df6a73f7379bee194b67b7e09bad5e72778ea35851749bff9e0787664c815f90ff83e3d653a4f6a9
|
|
7
|
+
data.tar.gz: c49adc022376e4b959c73d48704b293a536fd6e8a9580eaf8b495676b93c236ad7c02d24f748814c5394edb26edc7668bfee91acaf1fec03a43cbabb39153a18
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,12 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.1.3 (2026-09-29)
|
|
4
|
+
|
|
5
|
+
- Setting a `chat_provider_options` key to `nil` removes that field, including the Luna defaults and, before RubyLLM 2.1, the output limit field. A nested hash emptied this way is removed too.
|
|
6
|
+
- One-call answers given as a list of single-question objects are accepted when each question appears once. Claude Haiku 4.5 returns this shape on its first attempt, which previously cost a corrective retry on nearly every judgment.
|
|
7
|
+
- The one-call instructions now say that `answers` is an object keyed by question ID. Haiku no longer needs corrective retries (59 of 96 benchmark cases before), and Luna scored at least as well in a back-to-back comparison. See [the Haiku benchmark](benchmarks/haiku.md).
|
|
8
|
+
- Luna defaults apply to every OpenAI Luna model ID (`gpt-6-luna`, `gpt-5.6-luna`, `gpt-luna-latest`) and to OpenRouter's `openai/` Luna IDs, not only `gpt-6-luna` on OpenAI. On OpenRouter they send temperature zero and `reasoning: { enabled: false }`. With reasoning on, Luna's ratings often returned no digit. Pro variants and prefixed gateway IDs, such as `us.openai.gpt-5.6-luna` through `:openai`, are excluded.
|
|
9
|
+
|
|
3
10
|
## 0.1.2 (2026-09-28)
|
|
4
11
|
|
|
5
12
|
- When a rating call fails, no new calls start and the error is raised once calls in flight finish. Previously the remaining calls kept running, and were billed, after `judge` raised.
|
data/README.md
CHANGED
|
@@ -84,9 +84,11 @@ result = RubyLLM::LLMJudge.judge(
|
|
|
84
84
|
)
|
|
85
85
|
```
|
|
86
86
|
|
|
87
|
-
Both strategies return valid judgments with this model. Its one-call answers arrive in a Markdown code fence, which the gem removes before parsing.
|
|
87
|
+
Both strategies return valid judgments with this model. Its one-call answers arrive in a Markdown code fence, which the gem removes before parsing. On the benchmark cases it scored 56–57/64 on AG News and 29–30/32 on SST-2 at about 1.1 s median, a little below Luna. [Haiku results](benchmarks/haiku.md).
|
|
88
88
|
|
|
89
|
-
The default Luna rating request uses Chat Completions, temperature zero, reasoning disabled, `store: false`, and a four-token output limit. Tune requests with `scoring_protocol`, `temperature`, `max_output_tokens`, and `chat_provider_options`. `chat_provider_options` is merged over those defaults, so setting one key, such as `service_tier: 'flex'`, keeps reasoning disabled; set a key explicitly to change a default.
|
|
89
|
+
The default Luna rating request uses Chat Completions, temperature zero, reasoning disabled, `store: false`, and a four-token output limit. Tune requests with `scoring_protocol`, `temperature`, `max_output_tokens`, and `chat_provider_options`. `chat_provider_options` is merged over those defaults, so setting one key, such as `service_tier: 'flex'`, keeps reasoning disabled; set a key explicitly to change a default. Set a key to `nil` to remove that field from the request, whether it is a default or, before RubyLLM 2.1, the output limit field. For example, an app that routes RubyLLM 1.x OpenAI chats through the Responses API can send `chat_provider_options: { max_completion_tokens: nil, max_output_tokens: 256 }`.
|
|
90
|
+
|
|
91
|
+
Luna defaults apply to each provider's own Luna model IDs: `gpt-6-luna`, `gpt-5.6-luna`, or `gpt-luna-latest` on `:openai`, which sends temperature zero, `reasoning_effort: 'none'`, and `store: false`; and `openai/gpt-6-luna` or `openai/gpt-5.6-luna` on `:openrouter`, which sends temperature zero and `reasoning: { enabled: false }`. Pro variants, other providers such as `:bedrock`, and prefixed IDs such as `us.openai.gpt-5.6-luna` sent through `:openai` to an OpenAI-compatible gateway get no Luna defaults, because their request format may differ; set `temperature:` and disable reasoning in that endpoint's format yourself. With reasoning left on, the ratings strategy's four-token output limit is spent on reasoning: on eight AG News cases, `gpt-6-luna` through OpenRouter returned no digit for all eight, and `gpt-5.6-luna` for four. RubyLLM before 2.1 sends temperature 1.0 for any OpenAI model whose ID starts with `gpt-5`, so `gpt-5.6-luna` on `:openai` runs at 1.0 there unless you pass `chat_provider_options: { temperature: 0 }`.
|
|
90
92
|
|
|
91
93
|
## One-call typed decisions
|
|
92
94
|
|
|
@@ -6,12 +6,21 @@ module RubyLLM
|
|
|
6
6
|
module LLMJudge
|
|
7
7
|
class Engine
|
|
8
8
|
MAX_WORKERS = 6
|
|
9
|
+
# Luna model IDs by provider, which get temperature zero and reasoning disabled in that provider's
|
|
10
|
+
# format. Only each provider's own IDs match: an OpenAI-compatible gateway such as Bedrock, reached
|
|
11
|
+
# through :openai with an ID like "us.openai.gpt-5.6-luna", may use another request format. Pro
|
|
12
|
+
# variants are excluded: they exist to reason more.
|
|
13
|
+
LUNA_IDS = {
|
|
14
|
+
openai: /\Agpt-(?:\d+(?:\.\d+)?-)?luna(?:-latest)?\z/,
|
|
15
|
+
openrouter: %r{\A~?openai/gpt-(?:\d+(?:\.\d+)?-)?luna(?:-latest)?(?::\w+)?\z}
|
|
16
|
+
}.freeze
|
|
9
17
|
SYSTEM_INSTRUCTIONS = 'Return exactly one ASCII digit 0-9. 9 means very likely to be the correct answer; 0 means very unlikely. No explanation.'
|
|
10
18
|
SINGLE_REQUEST_INSTRUCTIONS = 'Answer all questions using only a JSON object with an "answers" field. ' \
|
|
11
19
|
'For each question, return an object mapping every supplied option ID to a ' \
|
|
12
20
|
'probability between 0 and 1. Include every option exactly once and make ' \
|
|
13
21
|
'each question\'s probabilities sum to 1. Evaluate each question independently ' \
|
|
14
|
-
'against the same state. Return no explanations or markdown.'
|
|
22
|
+
'against the same state. Return no explanations or markdown. The "answers" value ' \
|
|
23
|
+
'must be an object whose keys are the question IDs.'
|
|
15
24
|
|
|
16
25
|
def initialize(config:, model: DEFAULT_MODEL, provider_options: {}, scorer: nil, responder: nil)
|
|
17
26
|
@config = config
|
|
@@ -31,8 +40,9 @@ module RubyLLM
|
|
|
31
40
|
@scoring_model = model
|
|
32
41
|
raise ArgumentError, 'A model is required' unless @scoring_model.is_a?(String) && !@scoring_model.empty?
|
|
33
42
|
|
|
34
|
-
luna_defaults = @scoring_provider
|
|
35
|
-
@scoring_protocol = options.fetch(:scoring_protocol,
|
|
43
|
+
luna_defaults = LUNA_IDS.fetch(@scoring_provider, nil)&.match?(@scoring_model) || false
|
|
44
|
+
@scoring_protocol = options.fetch(:scoring_protocol,
|
|
45
|
+
luna_defaults && @scoring_provider == :openai ? :chat_completions : nil)
|
|
36
46
|
# Before 2.1, RubyLLM chats use each provider's one chat API (Chat Completions for OpenAI).
|
|
37
47
|
if !NATIVE && ![nil, :chat_completions].include?(@scoring_protocol&.to_sym)
|
|
38
48
|
raise ArgumentError, 'scoring_protocol requires RubyLLM 2.1'
|
|
@@ -135,6 +145,11 @@ module RubyLLM
|
|
|
135
145
|
raise Error, 'Scoring model returned a non-object response' unless parsed.is_a?(Hash)
|
|
136
146
|
|
|
137
147
|
reported = parsed.fetch('answers')
|
|
148
|
+
# Some models mirror the prompt's question list: [{"id": {...}}, ...].
|
|
149
|
+
if reported.is_a?(Array) && reported.all? { |item| item.is_a?(Hash) }
|
|
150
|
+
ids = reported.flat_map(&:keys)
|
|
151
|
+
reported = reported.reduce({}, :merge) if ids.uniq.size == ids.size
|
|
152
|
+
end
|
|
138
153
|
expected_questions = specs.map { |question, _| question.name.to_s }
|
|
139
154
|
unless reported.is_a?(Hash) && reported.keys.sort == expected_questions.sort
|
|
140
155
|
actual = reported.is_a?(Hash) ? reported.keys : reported.class.name
|
|
@@ -280,10 +295,13 @@ module RubyLLM
|
|
|
280
295
|
end
|
|
281
296
|
end
|
|
282
297
|
|
|
298
|
+
# Luna rates poorly with reasoning on: the output limit is spent before the answer.
|
|
283
299
|
def default_chat_options(luna_defaults)
|
|
284
|
-
|
|
285
|
-
|
|
286
|
-
luna_defaults ? {
|
|
300
|
+
case @scoring_provider
|
|
301
|
+
when :openai then luna_defaults ? { store: false, reasoning_effort: 'none' } : { store: false }
|
|
302
|
+
when :openrouter then luna_defaults ? { reasoning: { enabled: false } } : {}
|
|
303
|
+
else {}
|
|
304
|
+
end
|
|
287
305
|
end
|
|
288
306
|
|
|
289
307
|
def score_with_model(prompt)
|
|
@@ -318,13 +336,27 @@ module RubyLLM
|
|
|
318
336
|
chat.with_temperature(@temperature) unless @temperature.nil?
|
|
319
337
|
if NATIVE
|
|
320
338
|
chat.with_max_output_tokens(max_output_tokens)
|
|
321
|
-
chat.with_provider_options(@chat_provider_options)
|
|
339
|
+
chat.with_provider_options(deep_compact(@chat_provider_options))
|
|
322
340
|
else
|
|
323
|
-
chat.with_params(**deep_merge(max_output_tokens_param(max_output_tokens), @chat_provider_options))
|
|
341
|
+
chat.with_params(**deep_compact(deep_merge(max_output_tokens_param(max_output_tokens), @chat_provider_options)))
|
|
324
342
|
end
|
|
325
343
|
chat
|
|
326
344
|
end
|
|
327
345
|
|
|
346
|
+
# A nil in chat_provider_options removes that field: a default, or before 2.1
|
|
347
|
+
# the output limit field. A hash emptied by removal is removed too.
|
|
348
|
+
def deep_compact(hash)
|
|
349
|
+
hash.each_with_object({}) do |(key, value), compacted|
|
|
350
|
+
next if value.nil?
|
|
351
|
+
|
|
352
|
+
if value.is_a?(Hash) && !value.empty?
|
|
353
|
+
value = deep_compact(value)
|
|
354
|
+
next if value.empty?
|
|
355
|
+
end
|
|
356
|
+
compacted[key] = value
|
|
357
|
+
end
|
|
358
|
+
end
|
|
359
|
+
|
|
328
360
|
# The request field each provider reads for the output limit, as RubyLLM 2.1 renders it.
|
|
329
361
|
def max_output_tokens_param(limit)
|
|
330
362
|
case @scoring_provider
|