ruby_llm-llm_judge 0.1.1 → 0.1.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +18 -0
- data/README.md +14 -9
- data/lib/ruby_llm/llm_judge/engine.rb +85 -24
- data/lib/ruby_llm/llm_judge/legacy.rb +60 -14
- data/lib/ruby_llm/llm_judge/version.rb +1 -1
- data/lib/ruby_llm/llm_judge.rb +3 -5
- metadata +1 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: c438b1693068b5c47ada88cf13b23ed8786c597ed4c349d33d8c012f8edaedd2
|
|
4
|
+
data.tar.gz: ca228933f0f0772c7f523c788d9391edf012138b7f0faec84466da3026b1eb9b
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: b50f0831b1e86781c9d0aaac5d9d96b69284b0e527841119df6a73f7379bee194b67b7e09bad5e72778ea35851749bff9e0787664c815f90ff83e3d653a4f6a9
|
|
7
|
+
data.tar.gz: c49adc022376e4b959c73d48704b293a536fd6e8a9580eaf8b495676b93c236ad7c02d24f748814c5394edb26edc7668bfee91acaf1fec03a43cbabb39153a18
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,23 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.1.3 (2026-09-29)
|
|
4
|
+
|
|
5
|
+
- Setting a `chat_provider_options` key to `nil` removes that field, including the Luna defaults and, before RubyLLM 2.1, the output limit field. A nested hash emptied this way is removed too.
|
|
6
|
+
- One-call answers given as a list of single-question objects are accepted when each question appears once. Claude Haiku 4.5 returns this shape on its first attempt, which previously cost a corrective retry on nearly every judgment.
|
|
7
|
+
- The one-call instructions now say that `answers` is an object keyed by question ID. Haiku no longer needs corrective retries (59 of 96 benchmark cases before), and Luna scored at least as well in a back-to-back comparison. See [the Haiku benchmark](benchmarks/haiku.md).
|
|
8
|
+
- Luna defaults apply to every OpenAI Luna model ID (`gpt-6-luna`, `gpt-5.6-luna`, `gpt-luna-latest`) and to OpenRouter's `openai/` Luna IDs, not only `gpt-6-luna` on OpenAI. On OpenRouter they send temperature zero and `reasoning: { enabled: false }`. With reasoning on, Luna's ratings often returned no digit. Pro variants and prefixed gateway IDs, such as `us.openai.gpt-5.6-luna` through `:openai`, are excluded.
|
|
9
|
+
|
|
10
|
+
## 0.1.2 (2026-09-28)
|
|
11
|
+
|
|
12
|
+
- When a rating call fails, no new calls start and the error is raised once calls in flight finish. Previously the remaining calls kept running, and were billed, after `judge` raised.
|
|
13
|
+
- `chat_provider_options` is merged over the defaults instead of replacing them. Setting one option no longer re-enables reasoning or drops `store: false` for Luna.
|
|
14
|
+
- One-call JSON inside a Markdown code fence is accepted. Claude Haiku 4.5 now works with `strategy: :single_request`.
|
|
15
|
+
- A one-call Choice whose top probabilities tie gets the corrective retry, then raises, instead of selecting the first option. `tie_breaker: :first` keeps the old behavior.
|
|
16
|
+
- New `max_workers:` option (default 6) for concurrent rating calls.
|
|
17
|
+
- Unknown provider options are named in the error.
|
|
18
|
+
- A one-option Choice reports confidence 1.0.
|
|
19
|
+
- RubyLLM 1.13 through 1.16: question validation matches RubyLLM 2.1 (duplicate or empty option names, duplicate probability outcomes, description types), and token totals keep cached and thinking tokens.
|
|
20
|
+
|
|
3
21
|
## 0.1.1 (2026-09-28)
|
|
4
22
|
|
|
5
23
|
- Fix invalid scoring responses raising `NoMethodError` on RubyLLM 1.13, which also skipped the corrective retry. They now raise `RubyLLM::LLMJudge::Error`, a `RubyLLM::Error`, on every supported RubyLLM version.
|
data/README.md
CHANGED
|
@@ -19,7 +19,7 @@ The gem also runs on RubyLLM 1.13 through 1.16, which predate the Judge API. The
|
|
|
19
19
|
What differs before 2.1:
|
|
20
20
|
|
|
21
21
|
- Results are `RubyLLM::LLMJudge::Legacy` objects, not `RubyLLM::Judgment`, `RubyLLM::Choice`, and so on. Avoid class checks until you upgrade.
|
|
22
|
-
- There is no `RubyLLM.judge`, `RubyLLM::Judge` class DSL, `cost`,
|
|
22
|
+
- There is no `RubyLLM.judge`, `RubyLLM::Judge` class DSL, `cost`, instrumentation, or `metadata:`. Passing `metadata:` raises. `context:` is accepted and supplies its configuration.
|
|
23
23
|
- `scoring_protocol` accepts only `:chat_completions`, the API RubyLLM 1.x uses for OpenAI.
|
|
24
24
|
- The output limit and `chat_provider_options` are sent with `with_params`, using the field each provider expects (`max_completion_tokens` for OpenAI and Azure, `generationConfig.maxOutputTokens` for Gemini and Vertex AI, `inferenceConfig.maxTokens` for Bedrock, and `max_tokens` otherwise).
|
|
25
25
|
|
|
@@ -70,20 +70,25 @@ TicketTriage.judge('Please refund today.').urgent.probability
|
|
|
70
70
|
|
|
71
71
|
## Choose a model
|
|
72
72
|
|
|
73
|
-
Pass the chat model as `model:` and its provider as `scoring_provider
|
|
73
|
+
Pass the chat model as `model:` and its provider as `scoring_provider`. For example, Claude Haiku 4.5 through OpenRouter:
|
|
74
74
|
|
|
75
75
|
```ruby
|
|
76
76
|
result = RubyLLM::LLMJudge.judge(
|
|
77
77
|
'Please refund the duplicate charge today.',
|
|
78
|
-
model: 'claude-haiku-4
|
|
78
|
+
model: 'anthropic/claude-haiku-4.5',
|
|
79
79
|
questions: { urgent: { type: :probability, instructions: 'Does this need attention today?' } },
|
|
80
80
|
provider_options: {
|
|
81
|
-
scoring_provider: :
|
|
81
|
+
scoring_provider: :openrouter,
|
|
82
|
+
temperature: 0
|
|
82
83
|
}
|
|
83
84
|
)
|
|
84
85
|
```
|
|
85
86
|
|
|
86
|
-
|
|
87
|
+
Both strategies return valid judgments with this model. Its one-call answers arrive in a Markdown code fence, which the gem removes before parsing. On the benchmark cases it scored 56–57/64 on AG News and 29–30/32 on SST-2 at about 1.1 s median, a little below Luna. [Haiku results](benchmarks/haiku.md).
|
|
88
|
+
|
|
89
|
+
The default Luna rating request uses Chat Completions, temperature zero, reasoning disabled, `store: false`, and a four-token output limit. Tune requests with `scoring_protocol`, `temperature`, `max_output_tokens`, and `chat_provider_options`. `chat_provider_options` is merged over those defaults, so setting one key, such as `service_tier: 'flex'`, keeps reasoning disabled; set a key explicitly to change a default. Set a key to `nil` to remove that field from the request, whether it is a default or, before RubyLLM 2.1, the output limit field. For example, an app that routes RubyLLM 1.x OpenAI chats through the Responses API can send `chat_provider_options: { max_completion_tokens: nil, max_output_tokens: 256 }`.
|
|
90
|
+
|
|
91
|
+
Luna defaults apply to each provider's own Luna model IDs: `gpt-6-luna`, `gpt-5.6-luna`, or `gpt-luna-latest` on `:openai`, which sends temperature zero, `reasoning_effort: 'none'`, and `store: false`; and `openai/gpt-6-luna` or `openai/gpt-5.6-luna` on `:openrouter`, which sends temperature zero and `reasoning: { enabled: false }`. Pro variants, other providers such as `:bedrock`, and prefixed IDs such as `us.openai.gpt-5.6-luna` sent through `:openai` to an OpenAI-compatible gateway get no Luna defaults, because their request format may differ; set `temperature:` and disable reasoning in that endpoint's format yourself. With reasoning left on, the ratings strategy's four-token output limit is spent on reasoning: on eight AG News cases, `gpt-6-luna` through OpenRouter returned no digit for all eight, and `gpt-5.6-luna` for four. RubyLLM before 2.1 sends temperature 1.0 for any OpenAI model whose ID starts with `gpt-5`, so `gpt-5.6-luna` on `:openai` runs at 1.0 there unless you pass `chat_provider_options: { temperature: 0 }`.
|
|
87
92
|
|
|
88
93
|
## One-call typed decisions
|
|
89
94
|
|
|
@@ -107,7 +112,7 @@ result.urgent.probability
|
|
|
107
112
|
result.department.probabilities
|
|
108
113
|
```
|
|
109
114
|
|
|
110
|
-
The model returns one JSON object containing a distribution for each question. The gem checks that every question and option is present, validates each probability, normalizes each distribution, and builds RubyLLM's typed answers. `result.raw[:reported_probabilities]` and `result.raw[:reported_totals]` preserve the model's original numbers for inspection. With an OpenAI key configured, this example calls Luna directly through Chat Completions with reasoning disabled, temperature zero, and `store: false`. The 1024-token limit suits small judgments; omit it for large question sets to use the 8192-token default. It makes one corrective retry for malformed JSON
|
|
115
|
+
The model returns one JSON object containing a distribution for each question. The gem checks that every question and option is present, validates each probability, normalizes each distribution, and builds RubyLLM's typed answers. `result.raw[:reported_probabilities]` and `result.raw[:reported_totals]` preserve the model's original numbers for inspection. With an OpenAI key configured, this example calls Luna directly through Chat Completions with reasoning disabled, temperature zero, and `store: false`. The 1024-token limit suits small judgments; omit it for large question sets to use the 8192-token default. It makes one corrective retry for malformed JSON, missing fields, or a Choice whose top probabilities tie, counting both calls in `result.tokens` and `result.raw[:attempts]`. A tie that remains raises instead of selecting the first option; set `tie_breaker: :first` to accept the first option. Set `malformed_retries: 0` to disable that retry. An invalid response after retries raises an error.
|
|
111
116
|
|
|
112
117
|
If you use OpenRouter for Luna, prioritize the provider with the lowest observed latency:
|
|
113
118
|
|
|
@@ -139,11 +144,11 @@ OpenRouter's [latency sorting](https://openrouter.ai/docs/guides/routing/provide
|
|
|
139
144
|
|
|
140
145
|
## How scoring works
|
|
141
146
|
|
|
142
|
-
The default `:ratings` strategy asks the model to rate every declared answer from 0 to 9. It makes one call per answer, with up to six calls running at once, then applies softmax to the ratings. It retries a malformed digit once. When the highest ratings tie on a Choice question, it makes a one-call JSON judgment for that question and uses its distribution. If that call also ties, it raises an error instead of choosing whichever option came first. `result.raw[:tie_breaks]` records each resolution, and token usage includes the extra call. Set `tie_breaker: :first` to use the original first-option rule, or `tie_break_max_output_tokens:` to change the tie-break response limit (default 8192).
|
|
147
|
+
The default `:ratings` strategy asks the model to rate every declared answer from 0 to 9. It makes one call per answer, with up to six calls running at once (`max_workers:` changes this), then applies softmax to the ratings. If a call fails after its retry, no new calls start, calls in flight finish, and the error is raised. It retries a malformed digit once. When the highest ratings tie on a Choice question, it makes a one-call JSON judgment for that question and uses its distribution. If that call also ties, it raises an error instead of choosing whichever option came first. `result.raw[:tie_breaks]` records each resolution, and token usage includes the extra call. Set `tie_breaker: :first` to use the original first-option rule, or `tie_break_max_output_tokens:` to change the tie-break response limit (default 8192).
|
|
143
148
|
|
|
144
149
|
The `:single_request` strategy asks for all distributions in one call. Both strategies build the same RubyLLM answer types: a `choice` returns the selected answer and distribution; a `score` returns the distribution and its weighted level; a `probability` returns the positive answer's share. The gem supports Judge's 1–255 choice options, 2–10 score levels, and multiple questions in one judgment.
|
|
145
150
|
|
|
146
|
-
These probabilities compare the answers you supplied. The `:single_request` values are reported by the chat model; the default ratings values come from softmax over its 0–9 ratings. Neither strategy establishes calibration by itself. Include an `other` or `escalate` choice when the named answers may not cover the input. `confidence` measures how concentrated the returned distribution is. Use labeled examples to set any automation thresholds.
|
|
151
|
+
These probabilities compare the answers you supplied. The `:single_request` values are reported by the chat model; the default ratings values come from softmax over its 0–9 ratings. Neither strategy establishes calibration by itself. Include an `other` or `escalate` choice when the named answers may not cover the input. `confidence` measures how concentrated the returned distribution is: 1.0 for a single option, 0.0 for a tie. Use labeled examples to set any automation thresholds.
|
|
147
152
|
|
|
148
153
|
Set optional `max_arms` and `max_input_bytes` in `provider_options` to cap work per judgment. For example, `{ max_arms: 12, max_input_bytes: 32_768 }` rejects larger requests before scoring. Token usage is aggregated across calls, and `raw` contains the ratings and model metadata.
|
|
149
154
|
|
|
@@ -211,7 +216,7 @@ Latency sorting reduced median time by 20% on AG News and 28% on SST-2. It was f
|
|
|
211
216
|
|
|
212
217
|
## Development
|
|
213
218
|
|
|
214
|
-
The gem uses RubyLLM's Judge API when it is present and the `Legacy` path otherwise. `test/llm_judge_test.rb` covers the Judge API and `test/legacy_test.rb` covers RubyLLM 1.13 through 1.16; each skips on the other side.
|
|
219
|
+
The gem uses RubyLLM's Judge API when it is present and the `Legacy` path otherwise. `test/llm_judge_test.rb` covers the Judge API and `test/legacy_test.rb` covers RubyLLM 1.13 through 1.16; each skips on the other side. `test/engine_test.rb` and `test/http_test.rb` run on every version; the HTTP tests send real RubyLLM requests to WebMock stubs and check their bodies.
|
|
215
220
|
|
|
216
221
|
```bash
|
|
217
222
|
bundle exec rake test # the latest released RubyLLM
|
|
@@ -6,19 +6,30 @@ module RubyLLM
|
|
|
6
6
|
module LLMJudge
|
|
7
7
|
class Engine
|
|
8
8
|
MAX_WORKERS = 6
|
|
9
|
+
# Luna model IDs by provider, which get temperature zero and reasoning disabled in that provider's
|
|
10
|
+
# format. Only each provider's own IDs match: an OpenAI-compatible gateway such as Bedrock, reached
|
|
11
|
+
# through :openai with an ID like "us.openai.gpt-5.6-luna", may use another request format. Pro
|
|
12
|
+
# variants are excluded: they exist to reason more.
|
|
13
|
+
LUNA_IDS = {
|
|
14
|
+
openai: /\Agpt-(?:\d+(?:\.\d+)?-)?luna(?:-latest)?\z/,
|
|
15
|
+
openrouter: %r{\A~?openai/gpt-(?:\d+(?:\.\d+)?-)?luna(?:-latest)?(?::\w+)?\z}
|
|
16
|
+
}.freeze
|
|
9
17
|
SYSTEM_INSTRUCTIONS = 'Return exactly one ASCII digit 0-9. 9 means very likely to be the correct answer; 0 means very unlikely. No explanation.'
|
|
10
18
|
SINGLE_REQUEST_INSTRUCTIONS = 'Answer all questions using only a JSON object with an "answers" field. ' \
|
|
11
19
|
'For each question, return an object mapping every supplied option ID to a ' \
|
|
12
20
|
'probability between 0 and 1. Include every option exactly once and make ' \
|
|
13
21
|
'each question\'s probabilities sum to 1. Evaluate each question independently ' \
|
|
14
|
-
'against the same state. Return no explanations or markdown.'
|
|
22
|
+
'against the same state. Return no explanations or markdown. The "answers" value ' \
|
|
23
|
+
'must be an object whose keys are the question IDs.'
|
|
15
24
|
|
|
16
25
|
def initialize(config:, model: DEFAULT_MODEL, provider_options: {}, scorer: nil, responder: nil)
|
|
17
26
|
@config = config
|
|
18
27
|
options = provider_options.transform_keys(&:to_sym)
|
|
19
28
|
allowed = %i[scoring_provider scoring_protocol temperature max_output_tokens chat_provider_options
|
|
20
|
-
max_arms max_input_bytes strategy malformed_retries tie_breaker tie_break_max_output_tokens
|
|
21
|
-
|
|
29
|
+
max_arms max_input_bytes strategy malformed_retries tie_breaker tie_break_max_output_tokens
|
|
30
|
+
max_workers]
|
|
31
|
+
unknown = options.keys - allowed
|
|
32
|
+
raise ArgumentError, "Unknown LLMJudge provider options: #{unknown.join(', ')}" unless unknown.empty?
|
|
22
33
|
|
|
23
34
|
@strategy = options.fetch(:strategy, :ratings).to_sym
|
|
24
35
|
raise ArgumentError, 'strategy must be :ratings or :single_request' unless %i[ratings single_request].include?(@strategy)
|
|
@@ -29,8 +40,9 @@ module RubyLLM
|
|
|
29
40
|
@scoring_model = model
|
|
30
41
|
raise ArgumentError, 'A model is required' unless @scoring_model.is_a?(String) && !@scoring_model.empty?
|
|
31
42
|
|
|
32
|
-
luna_defaults = @scoring_provider
|
|
33
|
-
@scoring_protocol = options.fetch(:scoring_protocol,
|
|
43
|
+
luna_defaults = LUNA_IDS.fetch(@scoring_provider, nil)&.match?(@scoring_model) || false
|
|
44
|
+
@scoring_protocol = options.fetch(:scoring_protocol,
|
|
45
|
+
luna_defaults && @scoring_provider == :openai ? :chat_completions : nil)
|
|
34
46
|
# Before 2.1, RubyLLM chats use each provider's one chat API (Chat Completions for OpenAI).
|
|
35
47
|
if !NATIVE && ![nil, :chat_completions].include?(@scoring_protocol&.to_sym)
|
|
36
48
|
raise ArgumentError, 'scoring_protocol requires RubyLLM 2.1'
|
|
@@ -38,11 +50,16 @@ module RubyLLM
|
|
|
38
50
|
@temperature = options.fetch(:temperature, luna_defaults ? 0 : nil)
|
|
39
51
|
@max_output_tokens = options.fetch(:max_output_tokens, @strategy == :single_request ? 8192 : 4)
|
|
40
52
|
@tie_break_max_output_tokens = options.fetch(:tie_break_max_output_tokens, 8192)
|
|
41
|
-
|
|
53
|
+
chat_provider_options = options.fetch(:chat_provider_options, {})
|
|
54
|
+
raise ArgumentError, 'chat_provider_options must be a Hash' unless chat_provider_options.is_a?(Hash)
|
|
55
|
+
|
|
56
|
+
# Merged over the defaults, so tuning one option keeps the others (such as disabled reasoning).
|
|
57
|
+
@chat_provider_options = deep_merge(default_chat_options(luna_defaults), chat_provider_options)
|
|
42
58
|
@max_arms = options[:max_arms]
|
|
43
59
|
@max_input_bytes = options[:max_input_bytes]
|
|
44
60
|
@malformed_retries = options.fetch(:malformed_retries, 1)
|
|
45
|
-
|
|
61
|
+
@max_workers = options.fetch(:max_workers, MAX_WORKERS)
|
|
62
|
+
raise ArgumentError, 'max_workers must be a positive integer' unless @max_workers.is_a?(Integer) && @max_workers.positive?
|
|
46
63
|
raise ArgumentError, 'max_output_tokens must be positive' unless @max_output_tokens.is_a?(Integer) && @max_output_tokens.positive?
|
|
47
64
|
unless @tie_break_max_output_tokens.is_a?(Integer) && @tie_break_max_output_tokens.positive?
|
|
48
65
|
raise ArgumentError, 'tie_break_max_output_tokens must be positive'
|
|
@@ -70,12 +87,9 @@ module RubyLLM
|
|
|
70
87
|
}
|
|
71
88
|
chosen = answer(question, softmax(ratings))
|
|
72
89
|
if question.type == :choice && ratings.count(ratings.max) > 1 && @tie_breaker == :single_request
|
|
90
|
+
# Raises if the one-call judgment also ties, after its corrective retry.
|
|
73
91
|
judgment = judge_single_request(input, { question.name => question }, model)
|
|
74
92
|
chosen = judgment.answers.fetch(question.name)
|
|
75
|
-
probabilities = chosen.probabilities.values
|
|
76
|
-
if probabilities.count(probabilities.max) > 1
|
|
77
|
-
raise Error, "Scoring model could not resolve the tie for #{question.name}"
|
|
78
|
-
end
|
|
79
93
|
tie_breaks << { question: question.name, ratings:, judgment: }
|
|
80
94
|
end
|
|
81
95
|
[question.name, chosen]
|
|
@@ -127,10 +141,15 @@ module RubyLLM
|
|
|
127
141
|
end
|
|
128
142
|
|
|
129
143
|
def parse_single_request(content, specs)
|
|
130
|
-
parsed = JSON.parse(content)
|
|
144
|
+
parsed = JSON.parse(strip_code_fence(content))
|
|
131
145
|
raise Error, 'Scoring model returned a non-object response' unless parsed.is_a?(Hash)
|
|
132
146
|
|
|
133
147
|
reported = parsed.fetch('answers')
|
|
148
|
+
# Some models mirror the prompt's question list: [{"id": {...}}, ...].
|
|
149
|
+
if reported.is_a?(Array) && reported.all? { |item| item.is_a?(Hash) }
|
|
150
|
+
ids = reported.flat_map(&:keys)
|
|
151
|
+
reported = reported.reduce({}, :merge) if ids.uniq.size == ids.size
|
|
152
|
+
end
|
|
134
153
|
expected_questions = specs.map { |question, _| question.name.to_s }
|
|
135
154
|
unless reported.is_a?(Hash) && reported.keys.sort == expected_questions.sort
|
|
136
155
|
actual = reported.is_a?(Hash) ? reported.keys : reported.class.name
|
|
@@ -154,6 +173,12 @@ module RubyLLM
|
|
|
154
173
|
total = values.sum.to_f
|
|
155
174
|
raise Error, "Scoring model returned zero probability for #{question.name}" unless total.positive?
|
|
156
175
|
|
|
176
|
+
# A tied Choice would otherwise select whichever option was listed first.
|
|
177
|
+
if question.type == :choice && @tie_breaker == :single_request && values.count(values.max) > 1
|
|
178
|
+
raise Error, "Scoring model could not resolve the tie for #{question.name}; " \
|
|
179
|
+
'one option must have the highest probability'
|
|
180
|
+
end
|
|
181
|
+
|
|
157
182
|
normalization[question.name.to_s] = total
|
|
158
183
|
[question.name, answer(question, values.map { |value| value / total })]
|
|
159
184
|
end
|
|
@@ -230,32 +255,53 @@ module RubyLLM
|
|
|
230
255
|
value.is_a?(String) ? value : JSON.generate(value)
|
|
231
256
|
end
|
|
232
257
|
|
|
258
|
+
# Runs jobs on up to @max_workers threads. After a job fails, no new jobs
|
|
259
|
+
# start; calls already in flight finish, then the first error is raised.
|
|
233
260
|
def run(jobs)
|
|
234
261
|
queue = Queue.new
|
|
235
262
|
jobs.each_with_index { |job, index| queue << [index, job] }
|
|
236
263
|
results = Array.new(jobs.size)
|
|
237
|
-
|
|
264
|
+
errors = Queue.new
|
|
265
|
+
workers = Array.new([jobs.size, @max_workers].min) do
|
|
238
266
|
Thread.new do
|
|
239
|
-
|
|
240
|
-
loop do
|
|
267
|
+
while errors.empty?
|
|
241
268
|
begin
|
|
242
269
|
index, job = queue.pop(true)
|
|
243
270
|
rescue ThreadError
|
|
244
271
|
break
|
|
245
272
|
end
|
|
246
|
-
|
|
273
|
+
begin
|
|
274
|
+
results[index] = yield job
|
|
275
|
+
rescue StandardError => error
|
|
276
|
+
errors << error
|
|
277
|
+
end
|
|
247
278
|
end
|
|
248
279
|
end
|
|
249
280
|
end
|
|
250
281
|
workers.each(&:join)
|
|
251
|
-
|
|
282
|
+
raise errors.pop unless errors.empty?
|
|
283
|
+
|
|
252
284
|
results
|
|
253
285
|
end
|
|
254
286
|
|
|
255
|
-
|
|
256
|
-
|
|
287
|
+
# Some models wrap JSON in a Markdown code fence despite the instructions.
|
|
288
|
+
def strip_code_fence(content)
|
|
289
|
+
content.to_s.strip.sub(/\A```(?:json)?[ \t]*\n?/i, '').sub(/\n?```\z/, '')
|
|
290
|
+
end
|
|
257
291
|
|
|
258
|
-
|
|
292
|
+
def deep_merge(base, overrides)
|
|
293
|
+
base.merge(overrides) do |_key, old, new|
|
|
294
|
+
old.is_a?(Hash) && new.is_a?(Hash) ? deep_merge(old, new) : new
|
|
295
|
+
end
|
|
296
|
+
end
|
|
297
|
+
|
|
298
|
+
# Luna rates poorly with reasoning on: the output limit is spent before the answer.
|
|
299
|
+
def default_chat_options(luna_defaults)
|
|
300
|
+
case @scoring_provider
|
|
301
|
+
when :openai then luna_defaults ? { store: false, reasoning_effort: 'none' } : { store: false }
|
|
302
|
+
when :openrouter then luna_defaults ? { reasoning: { enabled: false } } : {}
|
|
303
|
+
else {}
|
|
304
|
+
end
|
|
259
305
|
end
|
|
260
306
|
|
|
261
307
|
def score_with_model(prompt)
|
|
@@ -290,13 +336,27 @@ module RubyLLM
|
|
|
290
336
|
chat.with_temperature(@temperature) unless @temperature.nil?
|
|
291
337
|
if NATIVE
|
|
292
338
|
chat.with_max_output_tokens(max_output_tokens)
|
|
293
|
-
chat.with_provider_options(@chat_provider_options)
|
|
339
|
+
chat.with_provider_options(deep_compact(@chat_provider_options))
|
|
294
340
|
else
|
|
295
|
-
chat.with_params(**
|
|
341
|
+
chat.with_params(**deep_compact(deep_merge(max_output_tokens_param(max_output_tokens), @chat_provider_options)))
|
|
296
342
|
end
|
|
297
343
|
chat
|
|
298
344
|
end
|
|
299
345
|
|
|
346
|
+
# A nil in chat_provider_options removes that field: a default, or before 2.1
|
|
347
|
+
# the output limit field. A hash emptied by removal is removed too.
|
|
348
|
+
def deep_compact(hash)
|
|
349
|
+
hash.each_with_object({}) do |(key, value), compacted|
|
|
350
|
+
next if value.nil?
|
|
351
|
+
|
|
352
|
+
if value.is_a?(Hash) && !value.empty?
|
|
353
|
+
value = deep_compact(value)
|
|
354
|
+
next if value.empty?
|
|
355
|
+
end
|
|
356
|
+
compacted[key] = value
|
|
357
|
+
end
|
|
358
|
+
end
|
|
359
|
+
|
|
300
360
|
# The request field each provider reads for the output limit, as RubyLLM 2.1 renders it.
|
|
301
361
|
def max_output_tokens_param(limit)
|
|
302
362
|
case @scoring_provider
|
|
@@ -331,11 +391,12 @@ module RubyLLM
|
|
|
331
391
|
def softmax(ratings)
|
|
332
392
|
peak = ratings.max
|
|
333
393
|
weights = ratings.map { |rating| Math.exp(rating - peak) }
|
|
334
|
-
|
|
394
|
+
total = weights.sum
|
|
395
|
+
weights.map { |weight| weight / total }
|
|
335
396
|
end
|
|
336
397
|
|
|
337
398
|
def concentration(probabilities)
|
|
338
|
-
return
|
|
399
|
+
return 1.0 if probabilities.size == 1
|
|
339
400
|
return 0.0 if probabilities.count(probabilities.max) > 1
|
|
340
401
|
|
|
341
402
|
entropy = -probabilities.sum { |probability| probability.zero? ? 0.0 : probability * Math.log(probability) }
|
|
@@ -26,7 +26,10 @@ module RubyLLM
|
|
|
26
26
|
end
|
|
27
27
|
|
|
28
28
|
def initialize(name, type:, instructions:, criteria:)
|
|
29
|
-
|
|
29
|
+
unless name.is_a?(String) || name.is_a?(Symbol)
|
|
30
|
+
raise ArgumentError, 'A question name must be a String or Symbol'
|
|
31
|
+
end
|
|
32
|
+
raise ArgumentError, 'A question name cannot be empty' if name.to_s.empty?
|
|
30
33
|
|
|
31
34
|
@name = name
|
|
32
35
|
@type = type
|
|
@@ -38,20 +41,58 @@ module RubyLLM
|
|
|
38
41
|
|
|
39
42
|
private
|
|
40
43
|
|
|
44
|
+
# Mirrors RubyLLM 2.1's Judge::Question validation.
|
|
41
45
|
def validate!
|
|
46
|
+
raise ArgumentError, 'Question instructions must be text, a Hash, an Array, or nil' unless description?(instructions)
|
|
47
|
+
|
|
42
48
|
case type
|
|
43
|
-
when :probability
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
49
|
+
when :probability then validate_probability!
|
|
50
|
+
when :choice then validate_choice!
|
|
51
|
+
when :score then validate_score!
|
|
52
|
+
end
|
|
53
|
+
end
|
|
54
|
+
|
|
55
|
+
def validate_probability!
|
|
56
|
+
return if criteria.nil?
|
|
57
|
+
|
|
58
|
+
unless criteria.is_a?(Hash) && (criteria.keys.map(&:to_s) - %w[yes no true false]).empty?
|
|
59
|
+
raise ArgumentError, 'Probability criteria must describe yes and no'
|
|
54
60
|
end
|
|
61
|
+
|
|
62
|
+
positive = criteria.keys.map { |key| %w[yes true].include?(key.to_s) }
|
|
63
|
+
raise ArgumentError, 'Probability criteria contain duplicate outcomes' unless positive.uniq.size == positive.size
|
|
64
|
+
|
|
65
|
+
validate_descriptions!(criteria.values)
|
|
66
|
+
end
|
|
67
|
+
|
|
68
|
+
def validate_choice!
|
|
69
|
+
raise ArgumentError, 'A choice needs a nonempty Hash of options' unless criteria.is_a?(Hash) && !criteria.empty?
|
|
70
|
+
unless criteria.keys.all? { |key| (key.is_a?(String) || key.is_a?(Symbol)) && !key.to_s.empty? }
|
|
71
|
+
raise ArgumentError, 'Choice options must have nonempty String or Symbol names'
|
|
72
|
+
end
|
|
73
|
+
|
|
74
|
+
duplicate = criteria.keys.map(&:to_s).tally.find { |_, count| count > 1 }&.first
|
|
75
|
+
raise ArgumentError, "Duplicate judgment key: #{duplicate}" if duplicate
|
|
76
|
+
|
|
77
|
+
validate_descriptions!(criteria.values)
|
|
78
|
+
end
|
|
79
|
+
|
|
80
|
+
def validate_score!
|
|
81
|
+
unless criteria.is_a?(Array) && criteria.size >= 2 && criteria.none?(&:nil?)
|
|
82
|
+
raise ArgumentError, 'A score needs at least two non-nil levels'
|
|
83
|
+
end
|
|
84
|
+
|
|
85
|
+
validate_descriptions!(criteria)
|
|
86
|
+
end
|
|
87
|
+
|
|
88
|
+
def validate_descriptions!(values)
|
|
89
|
+
return if values.all? { |value| description?(value) }
|
|
90
|
+
|
|
91
|
+
raise ArgumentError, 'Descriptions must be text, a Hash, an Array, or nil'
|
|
92
|
+
end
|
|
93
|
+
|
|
94
|
+
def description?(value)
|
|
95
|
+
value.nil? || value.is_a?(String) || value.is_a?(Hash) || value.is_a?(Array)
|
|
55
96
|
end
|
|
56
97
|
end
|
|
57
98
|
|
|
@@ -140,10 +181,15 @@ module RubyLLM
|
|
|
140
181
|
|
|
141
182
|
module_function
|
|
142
183
|
|
|
184
|
+
TOKEN_FIELDS = %i[input output cached cache_creation thinking].freeze
|
|
185
|
+
|
|
186
|
+
# Sums each field across calls; a field no call reported stays nil.
|
|
143
187
|
def aggregate_tokens(tokens)
|
|
144
188
|
tokens = tokens.compact
|
|
145
|
-
|
|
146
|
-
|
|
189
|
+
RubyLLM::Tokens.new(**TOKEN_FIELDS.to_h do |field|
|
|
190
|
+
values = tokens.filter_map { |token| token.public_send(field) }
|
|
191
|
+
[field, values.empty? ? nil : values.sum]
|
|
192
|
+
end)
|
|
147
193
|
end
|
|
148
194
|
end
|
|
149
195
|
end
|
data/lib/ruby_llm/llm_judge.rb
CHANGED
|
@@ -14,12 +14,10 @@ module RubyLLM
|
|
|
14
14
|
# Raised for invalid scoring responses. RubyLLM::Error takes (response, message)
|
|
15
15
|
# before 1.16 and (message, response:) on main, so build it for either.
|
|
16
16
|
class Error < RubyLLM::Error
|
|
17
|
+
MESSAGE_FIRST = RubyLLM::Error.instance_method(:initialize).parameters.include?(%i[key response])
|
|
18
|
+
|
|
17
19
|
def initialize(message = nil)
|
|
18
|
-
|
|
19
|
-
super(message)
|
|
20
|
-
else
|
|
21
|
-
super(nil, message)
|
|
22
|
-
end
|
|
20
|
+
MESSAGE_FIRST ? super(message) : super(nil, message)
|
|
23
21
|
end
|
|
24
22
|
end
|
|
25
23
|
end
|