ruby_llm-llm_judge 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml ADDED
@@ -0,0 +1,7 @@
1
+ ---
2
+ SHA256:
3
+ metadata.gz: '009274010f75c4b80a06f6b13b94fe9de590c7adf61ee78f57a771b8b115a742'
4
+ data.tar.gz: e801d20e964135b0e17a3de5ad6c0d121974afb4f61aa888035cb96f30a7a2d8
5
+ SHA512:
6
+ metadata.gz: 59f5ca2d328b42993d30168ebc528d58fb98f468b0d7600e4d9d059eacf7b3332d205ffe003175d05ebb6078cf2eda09f5d779309b17f5609b57e8a2363ae522
7
+ data.tar.gz: 9e0f82133b0b3fb1c295ac4b96a00b838c5d685df04964da90d49ffdd037ef156545cf3c2bf942d0a7e708f405ca0fecb07f810f71800e3015314e2c6f18f6da
data/CHANGELOG.md ADDED
@@ -0,0 +1,12 @@
1
+ # Changelog
2
+
3
+ ## 0.1.0 (2026-09-28)
4
+
5
+ Initial release.
6
+
7
+ - `RubyLLM::LLMJudge.judge` answers `probability`, `choice`, and `score` questions with any RubyLLM chat model. The default model is `gpt-6-luna`.
8
+ - Registers the `:llm_judge` provider for RubyLLM 2.1's Judge API, so `RubyLLM.judge` and `RubyLLM::Judge` classes return native `Judgment` and typed answers.
9
+ - Runs on RubyLLM 1.13 through 1.16, which predate the Judge API. `RubyLLM::LLMJudge.judge` returns `Legacy` answers with the same readers.
10
+ - `:ratings` strategy (default): one 0–9 rating per answer, up to six at once, turned into a distribution with softmax. Retries a malformed digit once and breaks tied Choice ratings with a one-call judgment.
11
+ - `:single_request` strategy: one call returns a distribution for every question. Validates and normalizes the distributions and makes one corrective retry for malformed JSON.
12
+ - `max_arms` and `max_input_bytes` limits reject oversized judgments before any call.
data/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 JP Camara
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
data/README.md ADDED
@@ -0,0 +1,223 @@
1
+ # ruby_llm-llm_judge
2
+
3
+ Use a RubyLLM chat model to answer [RubyLLM Judge](https://rubyllm.com/next/judgments/) questions. Define `probability`, `choice`, or `score` questions and receive RubyLLM's native `Judgment` and typed answers. Choose between parallel per-answer ratings and one-call typed distributions.
4
+
5
+ ## Install
6
+
7
+ ```ruby
8
+ gem 'ruby_llm-llm_judge'
9
+ ```
10
+
11
+ The gem works with RubyLLM 1.13 and later. On RubyLLM 2.1 it plugs into the Judge API; on earlier versions see [RubyLLM 1.13 to 1.16](#rubyllm-113-to-116).
12
+
13
+ Configure your chat provider's credentials through RubyLLM. The default scoring model is `gpt-6-luna`; you can choose another RubyLLM chat model for each judgment.
14
+
15
+ ### RubyLLM 1.13 to 1.16
16
+
17
+ The gem also runs on RubyLLM 1.13 through 1.16, which predate the Judge API. There, `RubyLLM::LLMJudge.judge` runs the same scoring directly and accepts the same questions and `provider_options`. Answers have the same readers as RubyLLM's (`probability`, `choice`, `probabilities`, `confidence`, `score`, `levels`, and `model`, `tokens`, `raw`, `[]` and `fetch` on the judgment), so call sites keep working after you upgrade to 2.1.
18
+
19
+ What differs before 2.1:
20
+
21
+ - Results are `RubyLLM::LLMJudge::Legacy` objects, not `RubyLLM::Judgment`, `RubyLLM::Choice`, and so on. Avoid class checks until you upgrade.
22
+ - There is no `RubyLLM.judge`, `RubyLLM::Judge` class DSL, `cost`, `context:` instrumentation, or `metadata:`. Passing `metadata:` raises.
23
+ - `scoring_protocol` accepts only `:chat_completions`, the API RubyLLM 1.x uses for OpenAI.
24
+ - The output limit and `chat_provider_options` are sent with `with_params`, using the field each provider expects (`max_completion_tokens` for OpenAI and Azure, `generationConfig.maxOutputTokens` for Gemini and Vertex AI, `inferenceConfig.maxTokens` for Bedrock, and `max_tokens` otherwise).
25
+
26
+ ## Use
27
+
28
+ ```ruby
29
+ require 'ruby_llm/llm_judge'
30
+
31
+ result = RubyLLM::LLMJudge.judge(
32
+ 'Please refund the duplicate charge today.',
33
+ questions: {
34
+ urgent: {
35
+ type: :probability,
36
+ instructions: 'Does this need attention today?',
37
+ criteria: { yes: 'An explicit deadline today', no: 'No deadline today' }
38
+ },
39
+ department: {
40
+ type: :choice,
41
+ instructions: 'Which team should handle this?',
42
+ options: { billing: 'Payments and refunds', technical: 'Bugs and integrations' }
43
+ },
44
+ frustration: {
45
+ type: :score,
46
+ instructions: 'How frustrated is the customer?',
47
+ levels: ['Calm', 'Frustrated', 'Angry']
48
+ }
49
+ }
50
+ )
51
+
52
+ result.urgent.probability
53
+ result.department.choice
54
+ result.department.probabilities
55
+ result.frustration.score
56
+ result.model # => "gpt-6-luna"
57
+ result.tokens
58
+ ```
59
+
60
+ You can also use a reusable Judge class:
61
+
62
+ ```ruby
63
+ class TicketTriage < RubyLLM::Judge
64
+ model 'gpt-6-luna', provider: :llm_judge, assume_model_exists: true
65
+ probability :urgent, 'Does this need attention today?'
66
+ end
67
+
68
+ TicketTriage.judge('Please refund today.').urgent.probability
69
+ ```
70
+
71
+ ## Choose a model
72
+
73
+ Pass the chat model as `model:` and its provider as `scoring_provider`:
74
+
75
+ ```ruby
76
+ result = RubyLLM::LLMJudge.judge(
77
+ 'Please refund the duplicate charge today.',
78
+ model: 'claude-haiku-4-5',
79
+ questions: { urgent: { type: :probability, instructions: 'Does this need attention today?' } },
80
+ provider_options: {
81
+ scoring_provider: :anthropic
82
+ }
83
+ )
84
+ ```
85
+
86
+ The default Luna rating request uses Chat Completions, temperature zero, reasoning disabled, `store: false`, and a four-token output limit. You can tune requests with `scoring_protocol`, `temperature`, `max_output_tokens`, and `chat_provider_options`.
87
+
88
+ ## One-call typed decisions
89
+
90
+ Set `strategy: :single_request` to send the state and every Judge question to the model in one request:
91
+
92
+ ```ruby
93
+ result = RubyLLM::LLMJudge.judge(
94
+ 'Please refund the duplicate charge today.',
95
+ questions: {
96
+ urgent: { type: :probability, instructions: 'Does this need attention today?' },
97
+ department: {
98
+ type: :choice,
99
+ instructions: 'Which team should handle this?',
100
+ options: { billing: 'Payments and refunds', technical: 'Bugs and integrations' }
101
+ }
102
+ },
103
+ provider_options: { strategy: :single_request, max_output_tokens: 1024 }
104
+ )
105
+
106
+ result.urgent.probability
107
+ result.department.probabilities
108
+ ```
109
+
110
+ The model returns one JSON object containing a distribution for each question. The gem checks that every question and option is present, validates each probability, normalizes each distribution, and builds RubyLLM's typed answers. `result.raw[:reported_probabilities]` and `result.raw[:reported_totals]` preserve the model's original numbers for inspection. With an OpenAI key configured, this example calls Luna directly through Chat Completions with reasoning disabled, temperature zero, and `store: false`. The 1024-token limit suits small judgments; omit it for large question sets to use the 8192-token default. It makes one corrective retry for malformed JSON or missing fields, counting both calls in `result.tokens` and `result.raw[:attempts]`. Set `malformed_retries: 0` to disable that retry. An invalid response after retries raises an error.
111
+
112
+ If you use OpenRouter for Luna, prioritize the provider with the lowest observed latency:
113
+
114
+ ```ruby
115
+ result = RubyLLM::LLMJudge.judge(
116
+ 'Please refund the duplicate charge today.',
117
+ model: 'openai/gpt-6-luna',
118
+ questions: {
119
+ department: {
120
+ type: :choice,
121
+ instructions: 'Which team should handle this?',
122
+ options: { billing: 'Payments and refunds', technical: 'Bugs and integrations' }
123
+ }
124
+ },
125
+ provider_options: {
126
+ strategy: :single_request,
127
+ scoring_provider: :openrouter,
128
+ temperature: 0,
129
+ max_output_tokens: 1024,
130
+ chat_provider_options: {
131
+ reasoning: { enabled: false },
132
+ provider: { sort: 'latency' }
133
+ }
134
+ }
135
+ )
136
+ ```
137
+
138
+ OpenRouter's [latency sorting](https://openrouter.ai/docs/guides/routing/provider-selection) chooses among its available providers using their observed response times. It sped up OpenRouter in our test, while direct OpenAI was faster still. The measured speed and answer-quality tradeoffs are below.
139
+
140
+ ## How scoring works
141
+
142
+ The default `:ratings` strategy asks the model to rate every declared answer from 0 to 9. It makes one call per answer, with up to six calls running at once, then applies softmax to the ratings. It retries a malformed digit once. When the highest ratings tie on a Choice question, it makes a one-call JSON judgment for that question and uses its distribution. If that call also ties, it raises an error instead of choosing whichever option came first. `result.raw[:tie_breaks]` records each resolution, and token usage includes the extra call. Set `tie_breaker: :first` to use the original first-option rule, or `tie_break_max_output_tokens:` to change the tie-break response limit (default 8192).
143
+
144
+ The `:single_request` strategy asks for all distributions in one call. Both strategies build the same RubyLLM answer types: a `choice` returns the selected answer and distribution; a `score` returns the distribution and its weighted level; a `probability` returns the positive answer's share. The gem supports Judge's 1–255 choice options, 2–10 score levels, and multiple questions in one judgment.
145
+
146
+ These probabilities compare the answers you supplied. The `:single_request` values are reported by the chat model; the default ratings values come from softmax over its 0–9 ratings. Neither strategy establishes calibration by itself. Include an `other` or `escalate` choice when the named answers may not cover the input. `confidence` measures how concentrated the returned distribution is. Use labeled examples to set any automation thresholds.
147
+
148
+ Set optional `max_arms` and `max_input_bytes` in `provider_options` to cap work per judgment. For example, `{ max_arms: 12, max_input_bytes: 32_768 }` rejects larger requests before scoring. Token usage is aggregated across calls, and `raw` contains the ratings and model metadata.
149
+
150
+ The one-digit-per-answer approach was inspired by [@burkov's post on Jev](https://x.com/burkov/status/2102438687392797045). The one-call strategy follows the typed-question pattern in [TypeSafe's System One adapter](https://github.com/typesafe-ai/system-one-adapter-python).
151
+
152
+ ## Benchmarks
153
+
154
+ On September 23, 2026, we tested GPT-6 Luna (OpenRouter), DeepSeek V4.1 Flash (Fireworks), [Celeris-1](https://docs.celeris.ai/making-requests), and Jev 1.13.0 on 64 balanced [AG News](https://huggingface.co/datasets/fancyzhx/ag_news) articles with four choices and 32 balanced [SST-2](https://huggingface.co/datasets/stanfordnlp/sst2) sentences. LLM runs used temperature zero with reasoning disabled. **Correct** counts all cases; **Usable** counts cases that returned a valid judgment. Brier scores and latency use usable cases only. Latency is the full client round trip, including retries.
155
+
156
+ | AG News model | Method | Correct / 64 | Usable / 64 | Brier ↓ | Median / p90 latency |
157
+ | --- | --- | ---: | ---: | ---: | ---: |
158
+ | Luna | One call | 59 | 64 | 0.1444 | 1,320 / 1,708 ms |
159
+ | Luna | Parallel ratings | 53 | 64 | 0.1578 | 1,718 / 2,753 ms |
160
+ | DeepSeek | One call | 59 | 64 | 0.1453 | 1,001 / 2,197 ms |
161
+ | DeepSeek | Parallel ratings | 56 | 64 | 0.1513 | 2,577 / 3,531 ms |
162
+ | Celeris | Plain one call | 54 | 58 | 0.1284 | 428 / 744 ms |
163
+ | Celeris | Original ratings | 42 | 54 | 0.1998 | 505 / 861 ms |
164
+ | Celeris | Original ratings, rerun | 47 | 58 | 0.1754 | 537 / 889 ms |
165
+ | Celeris | Ratings + retry/tie-break, run 1 | 57 | 61 | 0.1284 | 545 / 1,137 ms |
166
+ | Celeris | Ratings + retry/tie-break, run 2 | 56 | 59 | 0.1153 | 545 / 1,009 ms |
167
+ | Celeris | JSON one call + retry, run 1 | 60 | 63 | 0.0958 | 453 / 852 ms |
168
+ | Celeris | JSON one call + retry, run 2 | 59 | 62 | 0.1107 | 401 / 776 ms |
169
+ | Jev | Luna paired run | 59 | 64 | 0.1032 | 365 / 586 ms |
170
+ | Jev | DeepSeek paired run | 59 | 64 | 0.1040 | 505 / 1,060 ms |
171
+
172
+ | SST-2 model | Method | Choice / 32 | Probability / 32 | Score / 32 | Usable / 32 | Median latency |
173
+ | --- | --- | ---: | ---: | ---: | ---: | ---: |
174
+ | Luna | One call | 29 | 29 | 29 | 32 | 1,858 ms |
175
+ | Luna | Parallel ratings | 28 | 30 | 28 | 32 | 3,437 ms |
176
+ | DeepSeek | One call | 29 | 29 | 29 | 32 | 1,443 ms |
177
+ | DeepSeek | Parallel ratings | 30 | 26 | 28 | 32 | 3,009 ms |
178
+ | Celeris | Plain one call | 29 | 29 | 29 | 32 | 282 ms |
179
+ | Celeris | Original ratings | 28 | 24 | 26 | 30 | 530 ms |
180
+ | Celeris | Ratings + retry/tie-break, run 1 | 28 | 24 | 28 | 32 | 577 ms |
181
+ | Celeris | JSON one call + retry, run 1 | 28 | 28 | 28 | 32 | 248 ms |
182
+ | Celeris | JSON one call + retry, run 2 | 28 | 28 | 28 | 32 | 264 ms |
183
+ | Jev | Luna paired run | 30 | 30 | 30 | 32 | 522 ms |
184
+ | Jev | DeepSeek paired run | 30 | 29 | 30 | 32 | 308 ms |
185
+
186
+ The original 42/64 Celeris ratings result was not a one-off: the same behavior scored 47/64 in a fresh control run. In the original run, 10 cases returned no usable judgment and nine of the 12 incorrect usable choices had tied top ratings. With malformed-digit retry and Choice tie resolution, two runs scored 57/64 and 56/64. Those runs resolved nine of ten and seven of seven top ties to the reference label. Their observed median and p90 latencies were higher; tie-break calls add work. The control and revised run 2 covered only the 64 news cases; revised run 1 covered all 112 requests. Jev rows are saved earlier paired runs on the same cases, while Celeris was measured later.
187
+
188
+ ### Luna latency paths
189
+
190
+ In a paired API-level run using the gem's one-call prompt, direct OpenAI was faster than latency-sorted OpenRouter. Both used temperature zero, reasoning disabled, a 1024-token output limit, and one corrective retry for malformed JSON. Four direct cases were also verified through the gem. All final judgments were usable.
191
+
192
+ | Dataset | Luna path | Correct | Median | p90 | Brier ↓ |
193
+ | --- | --- | ---: | ---: | ---: | ---: |
194
+ | AG News (64) | Direct OpenAI | 58/64 | 823 ms | 1,241 ms | 0.1405 |
195
+ | AG News (64) | OpenRouter latency sort | 59/64 | 1,379 ms | 1,938 ms | 0.1265 |
196
+ | SST-2 (32) | Direct OpenAI | 28/32 | 843 ms | 1,745 ms | 0.1042 |
197
+ | SST-2 (32) | OpenRouter latency sort | 29/32 | 1,423 ms | 3,829 ms | 0.0762 |
198
+
199
+ In that paired run, direct OpenAI cut median latency by about 40% on both datasets, but lost one correct label on each and had worse Brier scores. Its SST-2 calls needed five corrective retries versus one through OpenRouter. A later direct-only repeat scored **59/64 news at 868 ms median** and **28/32 SST-2 at 1,022 ms median**, all usable. Just one news label and two sentiment labels changed across direct runs. OpenRouter was not rerun simultaneously with the repeat. [Methods and both runs' case-level results](benchmarks/luna-direct.md).
200
+
201
+ A separate paired gem run sent the same one-call Luna judgments through OpenRouter's default route and `provider: { sort: 'latency' }`, rotating request order across cases. All responses were usable. Each AG News request asked one four-choice question; each SST-2 request asked Choice, Probability, and Score together. Both routes used temperature zero, reasoning disabled, and a 1024-token output limit. These results are separate from the runs above.
202
+
203
+ | Dataset | OpenRouter route | Correct | Median | p90 | Brier ↓ |
204
+ | --- | --- | ---: | ---: | ---: | ---: |
205
+ | AG News (64) | Default | 61/64 | 1,569 ms | 2,805 ms | 0.0857 |
206
+ | AG News (64) | Lowest latency | 60/64 | 1,260 ms | 1,812 ms | 0.1165 |
207
+ | SST-2 (32) | Default | 29/32 | 2,093 ms | 8,693 ms | 0.0775 |
208
+ | SST-2 (32) | Lowest latency | 30/32 | 1,507 ms | 1,944 ms | 0.0695 |
209
+
210
+ Latency sorting reduced median time by 20% on AG News and 28% on SST-2. It was faster on 50/64 and 21/32 matched cases, respectively. AG News lost one correct label and its Brier score worsened; SST-2 gained one correct label. Check this route on your own labeled decisions before relying on its probabilities. [Methods and case-level data](benchmarks/luna-routing.md).
211
+
212
+ ## Development
213
+
214
+ The gem uses RubyLLM's Judge API when it is present and the `Legacy` path otherwise. `test/llm_judge_test.rb` covers the Judge API and `test/legacy_test.rb` covers RubyLLM 1.13 through 1.16; each skips on the other side.
215
+
216
+ ```bash
217
+ bundle exec rake test # the latest released RubyLLM
218
+ RUBY_LLM_VERSION=1.13.2 bundle exec rake test # a specific release
219
+ RUBY_LLM_VERSION=main bundle exec rake test # RubyLLM's main branch, with the Judge API
220
+ RUBY_LLM_PATH=../ruby_llm bundle exec rake test # a local checkout
221
+ ```
222
+
223
+ CI runs Ruby 3.2 through 4.0 against RubyLLM 1.13.2 and 1.16.0, and Ruby 3.4 against RubyLLM main.
@@ -0,0 +1,346 @@
1
+ # frozen_string_literal: true
2
+
3
+ require 'json'
4
+
5
+ module RubyLLM
6
+ module LLMJudge
7
+ class Engine
8
+ MAX_WORKERS = 6
9
+ SYSTEM_INSTRUCTIONS = 'Return exactly one ASCII digit 0-9. 9 means very likely to be the correct answer; 0 means very unlikely. No explanation.'
10
+ SINGLE_REQUEST_INSTRUCTIONS = 'Answer all questions using only a JSON object with an "answers" field. ' \
11
+ 'For each question, return an object mapping every supplied option ID to a ' \
12
+ 'probability between 0 and 1. Include every option exactly once and make ' \
13
+ 'each question\'s probabilities sum to 1. Evaluate each question independently ' \
14
+ 'against the same state. Return no explanations or markdown.'
15
+
16
+ def initialize(config:, model: DEFAULT_MODEL, provider_options: {}, scorer: nil, responder: nil)
17
+ @config = config
18
+ options = provider_options.transform_keys(&:to_sym)
19
+ allowed = %i[scoring_provider scoring_protocol temperature max_output_tokens chat_provider_options
20
+ max_arms max_input_bytes strategy malformed_retries tie_breaker tie_break_max_output_tokens]
21
+ raise ArgumentError, 'Unknown LLMJudge provider options' unless (options.keys - allowed).empty?
22
+
23
+ @strategy = options.fetch(:strategy, :ratings).to_sym
24
+ raise ArgumentError, 'strategy must be :ratings or :single_request' unless %i[ratings single_request].include?(@strategy)
25
+ @tie_breaker = options.fetch(:tie_breaker, :single_request).to_sym
26
+ raise ArgumentError, 'tie_breaker must be :single_request or :first' unless %i[single_request first].include?(@tie_breaker)
27
+
28
+ @scoring_provider = options.fetch(:scoring_provider, :openai).to_sym
29
+ @scoring_model = model
30
+ raise ArgumentError, 'A model is required' unless @scoring_model.is_a?(String) && !@scoring_model.empty?
31
+
32
+ luna_defaults = @scoring_provider == :openai && @scoring_model == DEFAULT_MODEL
33
+ @scoring_protocol = options.fetch(:scoring_protocol, luna_defaults ? :chat_completions : nil)
34
+ # Before 2.1, RubyLLM chats use each provider's one chat API (Chat Completions for OpenAI).
35
+ if !NATIVE && ![nil, :chat_completions].include?(@scoring_protocol&.to_sym)
36
+ raise ArgumentError, 'scoring_protocol requires RubyLLM 2.1'
37
+ end
38
+ @temperature = options.fetch(:temperature, luna_defaults ? 0 : nil)
39
+ @max_output_tokens = options.fetch(:max_output_tokens, @strategy == :single_request ? 8192 : 4)
40
+ @tie_break_max_output_tokens = options.fetch(:tie_break_max_output_tokens, 8192)
41
+ @chat_provider_options = options.fetch(:chat_provider_options, default_chat_options(luna_defaults))
42
+ @max_arms = options[:max_arms]
43
+ @max_input_bytes = options[:max_input_bytes]
44
+ @malformed_retries = options.fetch(:malformed_retries, 1)
45
+ raise ArgumentError, 'chat_provider_options must be a Hash' unless @chat_provider_options.is_a?(Hash)
46
+ raise ArgumentError, 'max_output_tokens must be positive' unless @max_output_tokens.is_a?(Integer) && @max_output_tokens.positive?
47
+ unless @tie_break_max_output_tokens.is_a?(Integer) && @tie_break_max_output_tokens.positive?
48
+ raise ArgumentError, 'tie_break_max_output_tokens must be positive'
49
+ end
50
+ unless [@max_arms, @max_input_bytes].all? { |limit| limit.nil? || (limit.is_a?(Integer) && limit.positive?) }
51
+ raise ArgumentError, 'LLMJudge limits must be positive integers'
52
+ end
53
+ unless @malformed_retries.is_a?(Integer) && @malformed_retries >= 0
54
+ raise ArgumentError, 'malformed_retries must be a nonnegative integer'
55
+ end
56
+
57
+ @scorer = scorer || method(:score_with_model)
58
+ @responder = responder || method(:respond_with_model)
59
+ end
60
+
61
+ def judge(input, questions:, model:)
62
+ return judge_single_request(input, questions, model) if @strategy == :single_request
63
+
64
+ jobs = build_jobs(input, questions)
65
+ results = run(jobs) { |job| @scorer.call(job[:prompt]) }
66
+ tie_breaks = []
67
+ answers = questions.values.to_h do |question|
68
+ ratings = jobs.each_index.filter_map { |index|
69
+ results[index].fetch(:digit) if jobs[index][:question] == question
70
+ }
71
+ chosen = answer(question, softmax(ratings))
72
+ if question.type == :choice && ratings.count(ratings.max) > 1 && @tie_breaker == :single_request
73
+ judgment = judge_single_request(input, { question.name => question }, model)
74
+ chosen = judgment.answers.fetch(question.name)
75
+ probabilities = chosen.probabilities.values
76
+ if probabilities.count(probabilities.max) > 1
77
+ raise RubyLLM::Error, "Scoring model could not resolve the tie for #{question.name}"
78
+ end
79
+ tie_breaks << { question: question.name, ratings:, judgment: }
80
+ end
81
+ [question.name, chosen]
82
+ end
83
+ tokens = aggregate_tokens(results.map { |result| result.fetch(:tokens) } +
84
+ tie_breaks.map { |item| item[:judgment].tokens })
85
+ Types::Judgment.new(
86
+ answers:, model: model.id, tokens:,
87
+ raw: { method: 'parallel_0_to_9_softmax', scoring_provider: @scoring_provider,
88
+ scoring_model: @scoring_model,
89
+ tie_breaks: tie_breaks.map { |item|
90
+ { question: item[:question], ratings: item[:ratings],
91
+ probabilities: item[:judgment].answers.fetch(item[:question]).probabilities,
92
+ attempts: item[:judgment].raw[:attempts] }
93
+ },
94
+ ratings: jobs.each_with_index.map { |job, index|
95
+ { question: job[:question].name, option: job[:name], digit: results[index][:digit] }
96
+ } }
97
+ )
98
+ end
99
+
100
+ private
101
+
102
+ def judge_single_request(input, questions, model)
103
+ specs = validated_options(input, questions)
104
+ original_prompt = single_request_prompt(input, specs)
105
+ prompt = original_prompt
106
+ attempts = []
107
+ loop do
108
+ result = @responder.call(prompt)
109
+ attempts << result
110
+ begin
111
+ answers, reported, normalization = parse_single_request(result.fetch(:content), specs)
112
+ return Types::Judgment.new(
113
+ answers:, model: model.id,
114
+ tokens: aggregate_tokens(attempts.map { |attempt| attempt.fetch(:tokens) }),
115
+ raw: { method: 'single_request_probabilities', scoring_provider: @scoring_provider,
116
+ scoring_model: @scoring_model, reported_probabilities: reported,
117
+ reported_totals: normalization, attempts: attempts.size }
118
+ )
119
+ rescue RubyLLM::Error => error
120
+ raise if attempts.size > @malformed_retries
121
+
122
+ prompt = "#{original_prompt}\n\nThe previous answer was invalid: #{error.message}. " \
123
+ "Previous answer: #{result.fetch(:content).to_s[0, 4000]}\n" \
124
+ 'Return a complete corrected JSON object with every question and option.'
125
+ end
126
+ end
127
+ end
128
+
129
+ def parse_single_request(content, specs)
130
+ parsed = JSON.parse(content)
131
+ raise RubyLLM::Error, 'Scoring model returned a non-object response' unless parsed.is_a?(Hash)
132
+
133
+ reported = parsed.fetch('answers')
134
+ expected_questions = specs.map { |question, _| question.name.to_s }
135
+ unless reported.is_a?(Hash) && reported.keys.sort == expected_questions.sort
136
+ actual = reported.is_a?(Hash) ? reported.keys : reported.class.name
137
+ raise RubyLLM::Error, "Question IDs must be #{expected_questions.inspect}; got #{actual.inspect}"
138
+ end
139
+
140
+ normalization = {}
141
+ answers = specs.to_h do |question, options|
142
+ distribution = reported.fetch(question.name.to_s)
143
+ keys = options.map { |name, _| name.to_s }
144
+ unless distribution.is_a?(Hash) && distribution.keys.sort == keys.sort
145
+ actual = distribution.is_a?(Hash) ? distribution.keys : distribution.class.name
146
+ raise RubyLLM::Error, "Option IDs for #{question.name} must be #{keys.inspect}; got #{actual.inspect}"
147
+ end
148
+
149
+ values = keys.map { |key| distribution.fetch(key) }
150
+ unless values.all? { |value| value.is_a?(Numeric) && value.finite? && (0..1).cover?(value) }
151
+ raise RubyLLM::Error, "Scoring model returned invalid probabilities for #{question.name}"
152
+ end
153
+
154
+ total = values.sum.to_f
155
+ raise RubyLLM::Error, "Scoring model returned zero probability for #{question.name}" unless total.positive?
156
+
157
+ normalization[question.name.to_s] = total
158
+ [question.name, answer(question, values.map { |value| value / total })]
159
+ end
160
+ [answers, reported, normalization]
161
+ rescue JSON::ParserError, KeyError, TypeError => error
162
+ raise RubyLLM::Error, "Scoring model returned invalid JSON probabilities: #{error.message}"
163
+ end
164
+
165
+ def single_request_prompt(input, specs)
166
+ request = {
167
+ state: input,
168
+ questions: specs.map do |question, options|
169
+ { id: question.name, type: question.type, instructions: question.instructions,
170
+ options: options.map { |name, description| { id: name.to_s, description: } } }
171
+ end
172
+ }
173
+ "Evaluate this decision request. Return only the requested JSON object.\n#{JSON.generate(request)}"
174
+ end
175
+
176
+ def build_jobs(input, questions)
177
+ jobs = validated_options(input, questions).flat_map do |question, options|
178
+ options.map do |name, description|
179
+ { question:, name:, description:, prompt: prompt(input, question.instructions, name, description) }
180
+ end
181
+ end
182
+ jobs
183
+ end
184
+
185
+ def validated_options(input, questions)
186
+ if @max_input_bytes
187
+ input_bytes = serialize(input).bytesize + questions.values.sum do |question|
188
+ serialize(question.instructions).bytesize + serialize(question.criteria).bytesize
189
+ end
190
+ raise ArgumentError, "LLMJudge input exceeds #{@max_input_bytes} bytes" if input_bytes > @max_input_bytes
191
+ end
192
+
193
+ specs = questions.values.map { |question| [question, options_for(question)] }
194
+ arms = specs.sum { |_, options| options.size }
195
+ raise ArgumentError, "LLMJudge exceeds #{@max_arms} scoring arms" if @max_arms && arms > @max_arms
196
+
197
+ specs
198
+ end
199
+
200
+ def options_for(question)
201
+ case question.type
202
+ when :probability
203
+ criteria = question.criteria || {}
204
+ yes_key = criteria.keys.find { |key| %w[yes true].include?(key.to_s) }
205
+ no_key = criteria.keys.find { |key| %w[no false].include?(key.to_s) }
206
+ yes = yes_key ? criteria[yes_key] : 'Yes, the condition holds'
207
+ no = no_key ? criteria[no_key] : 'No, the condition does not hold'
208
+ [['true', yes], ['false', no]]
209
+ when :choice
210
+ raise ArgumentError, 'A choice needs 1–255 options' unless (1..255).cover?(question.criteria.size)
211
+
212
+ question.criteria.to_a
213
+ when :score
214
+ raise ArgumentError, 'A score needs 2–10 levels' unless (2..10).cover?(question.criteria.size)
215
+
216
+ question.criteria.each_with_index.map { |description, index| [index, description] }
217
+ else
218
+ raise ArgumentError, "Unsupported question type: #{question.type}"
219
+ end
220
+ end
221
+
222
+ def prompt(input, instructions, name, description)
223
+ "Question: #{serialize(instructions || 'Which answer best applies?')}\n" \
224
+ "State: #{serialize(input)}\n\n" \
225
+ 'Output a single number between 0 and 9 to tell how likely it is that the correct answer is: ' \
226
+ "#{name} — #{serialize(description)}"
227
+ end
228
+
229
+ def serialize(value)
230
+ value.is_a?(String) ? value : JSON.generate(value)
231
+ end
232
+
233
+ def run(jobs)
234
+ queue = Queue.new
235
+ jobs.each_with_index { |job, index| queue << [index, job] }
236
+ results = Array.new(jobs.size)
237
+ workers = Array.new([jobs.size, MAX_WORKERS].min) do
238
+ Thread.new do
239
+ Thread.current.report_on_exception = false
240
+ loop do
241
+ begin
242
+ index, job = queue.pop(true)
243
+ rescue ThreadError
244
+ break
245
+ end
246
+ results[index] = yield job
247
+ end
248
+ end
249
+ end
250
+ workers.each(&:join)
251
+ workers.each(&:value)
252
+ results
253
+ end
254
+
255
+ def default_chat_options(luna_defaults)
256
+ return {} unless @scoring_provider == :openai
257
+
258
+ luna_defaults ? { store: false, reasoning_effort: 'none' } : { store: false }
259
+ end
260
+
261
+ def score_with_model(prompt)
262
+ attempts = []
263
+ loop do
264
+ request = attempts.empty? ? prompt : "#{prompt}\n\nYour previous answer was invalid. Return exactly one ASCII digit 0-9."
265
+ response = scoring_chat.ask(request)
266
+ attempts << response
267
+ digit = response.content.to_s.strip
268
+ if /\A[0-9]\z/.match?(digit)
269
+ return { digit: digit.to_i, tokens: aggregate_tokens(attempts.map(&:tokens)) }
270
+ end
271
+ raise RubyLLM::Error, 'Scoring model returned no single digit' if attempts.size > @malformed_retries
272
+ end
273
+ end
274
+
275
+ def respond_with_model(prompt)
276
+ limit = @strategy == :ratings ? @tie_break_max_output_tokens : @max_output_tokens
277
+ response = scoring_chat(instructions: SINGLE_REQUEST_INSTRUCTIONS, max_output_tokens: limit).ask(prompt)
278
+ { content: response.content.to_s, tokens: response.tokens }
279
+ end
280
+
281
+ def scoring_chat(instructions: SYSTEM_INSTRUCTIONS, max_output_tokens: @max_output_tokens)
282
+ context = RubyLLM::Context.new(@config)
283
+ chat = if NATIVE
284
+ context.chat(model: @scoring_model, provider: @scoring_provider,
285
+ protocol: @scoring_protocol, assume_model_exists: true)
286
+ else
287
+ context.chat(model: @scoring_model, provider: @scoring_provider, assume_model_exists: true)
288
+ end
289
+ chat.with_instructions(instructions)
290
+ chat.with_temperature(@temperature) unless @temperature.nil?
291
+ if NATIVE
292
+ chat.with_max_output_tokens(max_output_tokens)
293
+ chat.with_provider_options(@chat_provider_options)
294
+ else
295
+ chat.with_params(**RubyLLM::Utils.deep_merge(max_output_tokens_param(max_output_tokens), @chat_provider_options))
296
+ end
297
+ chat
298
+ end
299
+
300
+ # The request field each provider reads for the output limit, as RubyLLM 2.1 renders it.
301
+ def max_output_tokens_param(limit)
302
+ case @scoring_provider
303
+ when :openai, :azure then { max_completion_tokens: limit }
304
+ when :gemini, :vertexai then { generationConfig: { maxOutputTokens: limit } }
305
+ when :bedrock then { inferenceConfig: { maxTokens: limit } }
306
+ else { max_tokens: limit }
307
+ end
308
+ end
309
+
310
+ def aggregate_tokens(tokens)
311
+ NATIVE ? RubyLLM::Tokens.aggregate(tokens) : Legacy.aggregate_tokens(tokens)
312
+ end
313
+
314
+ def answer(question, probabilities)
315
+ case question.type
316
+ when :probability
317
+ Types::Probability.new(probability: probabilities.first)
318
+ when :choice
319
+ names = question.criteria.keys
320
+ distribution = names.zip(probabilities).to_h
321
+ Types::Choice.new(choice: names[probabilities.index(probabilities.max)], probabilities: distribution,
322
+ confidence: concentration(probabilities))
323
+ when :score
324
+ Types::Score.new(score: probabilities.each_with_index.sum { |probability, index| probability * index },
325
+ levels: question.criteria,
326
+ probabilities: probabilities.each_with_index.to_h { |probability, index| [index, probability] },
327
+ confidence: concentration(probabilities))
328
+ end
329
+ end
330
+
331
+ def softmax(ratings)
332
+ peak = ratings.max
333
+ weights = ratings.map { |rating| Math.exp(rating - peak) }
334
+ weights.map { |weight| weight / weights.sum }
335
+ end
336
+
337
+ def concentration(probabilities)
338
+ return 0.0 if probabilities.one?
339
+ return 0.0 if probabilities.count(probabilities.max) > 1
340
+
341
+ entropy = -probabilities.sum { |probability| probability.zero? ? 0.0 : probability * Math.log(probability) }
342
+ [[1.0 - entropy / Math.log(probabilities.size), 0.0].max, 1.0].min
343
+ end
344
+ end
345
+ end
346
+ end
@@ -0,0 +1,150 @@
1
+ # frozen_string_literal: true
2
+
3
+ module RubyLLM
4
+ module LLMJudge
5
+ # Stand-ins for RubyLLM 2.1's Judge types, used when RubyLLM predates the
6
+ # Judge API. They expose the same readers, so code written against
7
+ # RubyLLM::LLMJudge.judge keeps working after upgrading.
8
+ module Legacy
9
+ Model = Data.define(:id)
10
+
11
+ class Question
12
+ CRITERIA_KEYS = { probability: :criteria, choice: :options, score: :levels }.freeze
13
+
14
+ attr_reader :name, :type, :instructions, :criteria
15
+
16
+ def self.from_h(name, definition)
17
+ raise ArgumentError, 'Each question must be a Hash' unless definition.is_a?(Hash)
18
+
19
+ definition = definition.transform_keys(&:to_sym)
20
+ type = definition[:type]&.to_sym
21
+ key = CRITERIA_KEYS.fetch(type) { raise ArgumentError, "Unknown judgment type: #{type.inspect}" }
22
+ extra = definition.keys - [:type, :instructions, key]
23
+ raise ArgumentError, "Unknown question options: #{extra.join(', ')}" unless extra.empty?
24
+
25
+ new(name, type:, instructions: definition[:instructions], criteria: definition[key])
26
+ end
27
+
28
+ def initialize(name, type:, instructions:, criteria:)
29
+ raise ArgumentError, 'A question name must be a String or Symbol' unless name.is_a?(String) || name.is_a?(Symbol)
30
+
31
+ @name = name
32
+ @type = type
33
+ @instructions = instructions
34
+ @criteria = criteria
35
+ validate!
36
+ freeze
37
+ end
38
+
39
+ private
40
+
41
+ def validate!
42
+ case type
43
+ when :probability
44
+ return if criteria.nil?
45
+ unless criteria.is_a?(Hash) && (criteria.keys.map(&:to_s) - %w[yes no true false]).empty?
46
+ raise ArgumentError, 'Probability criteria must describe yes and no'
47
+ end
48
+ when :choice
49
+ raise ArgumentError, 'A choice needs a nonempty Hash of options' unless criteria.is_a?(Hash) && !criteria.empty?
50
+ when :score
51
+ unless criteria.is_a?(Array) && criteria.size >= 2 && criteria.none?(&:nil?)
52
+ raise ArgumentError, 'A score needs at least two non-nil levels'
53
+ end
54
+ end
55
+ end
56
+ end
57
+
58
+ class Probability
59
+ attr_reader :probability
60
+
61
+ def initialize(probability:)
62
+ @probability = probability
63
+ freeze
64
+ end
65
+
66
+ def type = :probability
67
+ def to_h = { type:, probability: }
68
+ end
69
+
70
+ class Choice
71
+ attr_reader :choice, :probabilities, :confidence
72
+
73
+ def initialize(choice:, probabilities:, confidence:)
74
+ @choice = choice
75
+ @probabilities = probabilities.dup.freeze
76
+ @confidence = confidence
77
+ freeze
78
+ end
79
+
80
+ def type = :choice
81
+ def to_h = { type:, choice:, probabilities:, confidence: }
82
+ end
83
+
84
+ class Score
85
+ attr_reader :score, :levels, :probabilities, :confidence
86
+
87
+ def initialize(score:, levels:, probabilities:, confidence:)
88
+ @score = score
89
+ @levels = levels
90
+ @probabilities = probabilities.dup.freeze
91
+ @confidence = confidence
92
+ freeze
93
+ end
94
+
95
+ def type = :score
96
+ def to_h = { type:, score:, levels:, probabilities:, confidence: }
97
+ end
98
+
99
+ class Judgment
100
+ include Enumerable
101
+
102
+ attr_reader :answers, :model, :tokens, :raw
103
+
104
+ def initialize(answers:, model:, tokens:, raw: nil)
105
+ @answers = answers.dup.freeze
106
+ @answer_keys = answers.keys.to_h { |key| [key.to_s, key] }.freeze
107
+ @model = model
108
+ @tokens = tokens
109
+ @raw = raw
110
+ end
111
+
112
+ def [](name)
113
+ answers[@answer_keys[name.to_s]]
114
+ end
115
+
116
+ def fetch(name)
117
+ answers.fetch(@answer_keys.fetch(name.to_s))
118
+ end
119
+
120
+ def each(&)
121
+ answers.each(&)
122
+ end
123
+
124
+ def to_h
125
+ { model:, answers: answers.transform_values(&:to_h), tokens: tokens.to_h }
126
+ end
127
+
128
+ private
129
+
130
+ def method_missing(name, *args, &block)
131
+ return super unless args.empty? && !block && @answer_keys.key?(name.to_s)
132
+
133
+ self[name]
134
+ end
135
+
136
+ def respond_to_missing?(name, include_private = false)
137
+ @answer_keys.key?(name.to_s) || super
138
+ end
139
+ end
140
+
141
+ module_function
142
+
143
+ def aggregate_tokens(tokens)
144
+ tokens = tokens.compact
145
+ sum = ->(reader) { tokens.sum { |token| token.public_send(reader) || 0 } }
146
+ RubyLLM::Tokens.new(input: sum.(:input), output: sum.(:output))
147
+ end
148
+ end
149
+ end
150
+ end
@@ -0,0 +1,19 @@
1
+ # frozen_string_literal: true
2
+
3
+ module RubyLLM
4
+ module LLMJudge
5
+ class Provider < RubyLLM::Provider
6
+ def api_base
7
+ 'http://127.0.0.1'
8
+ end
9
+
10
+ def judge(input, questions:, model:, provider_options: {})
11
+ Engine.new(config:, model: model.id, provider_options:).judge(input, questions:, model:)
12
+ end
13
+
14
+ def self.assume_models_exist?
15
+ true
16
+ end
17
+ end
18
+ end
19
+ end
@@ -0,0 +1,7 @@
1
+ # frozen_string_literal: true
2
+
3
+ module RubyLLM
4
+ module LLMJudge
5
+ VERSION = '0.1.0'
6
+ end
7
+ end
@@ -0,0 +1,42 @@
1
+ # frozen_string_literal: true
2
+
3
+ require 'ruby_llm'
4
+ require_relative 'llm_judge/version'
5
+
6
+ module RubyLLM
7
+ module LLMJudge
8
+ DEFAULT_MODEL = 'gpt-6-luna'
9
+
10
+ # RubyLLM 2.1 added the Judge API. Earlier versions run the engine directly
11
+ # and return Legacy stand-ins with the same readers.
12
+ NATIVE = defined?(RubyLLM::Judge) ? true : false
13
+ end
14
+ end
15
+
16
+ require_relative 'llm_judge/legacy' unless RubyLLM::LLMJudge::NATIVE
17
+ require_relative 'llm_judge/engine'
18
+ require_relative 'llm_judge/provider' if RubyLLM::LLMJudge::NATIVE
19
+
20
+ module RubyLLM
21
+ module LLMJudge
22
+ # Answer and judgment classes: RubyLLM's own on 2.1+, Legacy otherwise.
23
+ Types = NATIVE ? RubyLLM : Legacy
24
+
25
+ def self.judge(input, questions:, model: DEFAULT_MODEL, **options)
26
+ if NATIVE
27
+ return RubyLLM.judge(input, questions:, model:, provider: :llm_judge,
28
+ assume_model_exists: true, **options)
29
+ end
30
+
31
+ provider_options = options.delete(:provider_options) || {}
32
+ context = options.delete(:context)
33
+ raise ArgumentError, "Unsupported before RubyLLM 2.1: #{options.keys.join(', ')}" unless options.empty?
34
+
35
+ built = questions.to_h { |name, definition| [name.to_s, Legacy::Question.from_h(name, definition)] }
36
+ Engine.new(config: context&.config || RubyLLM.config, model:, provider_options:)
37
+ .judge(input, questions: built, model: Legacy::Model.new(id: model))
38
+ end
39
+ end
40
+ end
41
+
42
+ RubyLLM::Provider.register(:llm_judge, RubyLLM::LLMJudge::Provider) if RubyLLM::LLMJudge::NATIVE
metadata ADDED
@@ -0,0 +1,72 @@
1
+ --- !ruby/object:Gem::Specification
2
+ name: ruby_llm-llm_judge
3
+ version: !ruby/object:Gem::Version
4
+ version: 0.1.0
5
+ platform: ruby
6
+ authors:
7
+ - JP Camara
8
+ bindir: bin
9
+ cert_chain: []
10
+ date: 1980-01-02 00:00:00.000000000 Z
11
+ dependencies:
12
+ - !ruby/object:Gem::Dependency
13
+ name: ruby_llm
14
+ requirement: !ruby/object:Gem::Requirement
15
+ requirements:
16
+ - - ">="
17
+ - !ruby/object:Gem::Version
18
+ version: 1.13.0
19
+ - - "<"
20
+ - !ruby/object:Gem::Version
21
+ version: '3.0'
22
+ type: :runtime
23
+ prerelease: false
24
+ version_requirements: !ruby/object:Gem::Requirement
25
+ requirements:
26
+ - - ">="
27
+ - !ruby/object:Gem::Version
28
+ version: 1.13.0
29
+ - - "<"
30
+ - !ruby/object:Gem::Version
31
+ version: '3.0'
32
+ description: Answers probability, choice, and score questions with any RubyLLM chat
33
+ model, using parallel per-answer ratings or one-call typed distributions. Plugs
34
+ into RubyLLM 2.1's Judge API and also runs on RubyLLM 1.13 through 1.16.
35
+ executables: []
36
+ extensions: []
37
+ extra_rdoc_files: []
38
+ files:
39
+ - CHANGELOG.md
40
+ - LICENSE
41
+ - README.md
42
+ - lib/ruby_llm/llm_judge.rb
43
+ - lib/ruby_llm/llm_judge/engine.rb
44
+ - lib/ruby_llm/llm_judge/legacy.rb
45
+ - lib/ruby_llm/llm_judge/provider.rb
46
+ - lib/ruby_llm/llm_judge/version.rb
47
+ homepage: https://github.com/jpcamara/ruby_llm-llm_judge
48
+ licenses:
49
+ - MIT
50
+ metadata:
51
+ source_code_uri: https://github.com/jpcamara/ruby_llm-llm_judge
52
+ changelog_uri: https://github.com/jpcamara/ruby_llm-llm_judge/blob/main/CHANGELOG.md
53
+ bug_tracker_uri: https://github.com/jpcamara/ruby_llm-llm_judge/issues
54
+ rubygems_mfa_required: 'true'
55
+ rdoc_options: []
56
+ require_paths:
57
+ - lib
58
+ required_ruby_version: !ruby/object:Gem::Requirement
59
+ requirements:
60
+ - - ">="
61
+ - !ruby/object:Gem::Version
62
+ version: '3.2'
63
+ required_rubygems_version: !ruby/object:Gem::Requirement
64
+ requirements:
65
+ - - ">="
66
+ - !ruby/object:Gem::Version
67
+ version: '0'
68
+ requirements: []
69
+ rubygems_version: 3.6.9
70
+ specification_version: 4
71
+ summary: Use RubyLLM chat models to answer RubyLLM Judge questions
72
+ test_files: []