ruby_llm-llm_judge 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +7 -0
- data/CHANGELOG.md +12 -0
- data/LICENSE +21 -0
- data/README.md +223 -0
- data/lib/ruby_llm/llm_judge/engine.rb +346 -0
- data/lib/ruby_llm/llm_judge/legacy.rb +150 -0
- data/lib/ruby_llm/llm_judge/provider.rb +19 -0
- data/lib/ruby_llm/llm_judge/version.rb +7 -0
- data/lib/ruby_llm/llm_judge.rb +42 -0
- metadata +72 -0
checksums.yaml
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
1
|
+
---
|
|
2
|
+
SHA256:
|
|
3
|
+
metadata.gz: '009274010f75c4b80a06f6b13b94fe9de590c7adf61ee78f57a771b8b115a742'
|
|
4
|
+
data.tar.gz: e801d20e964135b0e17a3de5ad6c0d121974afb4f61aa888035cb96f30a7a2d8
|
|
5
|
+
SHA512:
|
|
6
|
+
metadata.gz: 59f5ca2d328b42993d30168ebc528d58fb98f468b0d7600e4d9d059eacf7b3332d205ffe003175d05ebb6078cf2eda09f5d779309b17f5609b57e8a2363ae522
|
|
7
|
+
data.tar.gz: 9e0f82133b0b3fb1c295ac4b96a00b838c5d685df04964da90d49ffdd037ef156545cf3c2bf942d0a7e708f405ca0fecb07f810f71800e3015314e2c6f18f6da
|
data/CHANGELOG.md
ADDED
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 0.1.0 (2026-09-28)
|
|
4
|
+
|
|
5
|
+
Initial release.
|
|
6
|
+
|
|
7
|
+
- `RubyLLM::LLMJudge.judge` answers `probability`, `choice`, and `score` questions with any RubyLLM chat model. The default model is `gpt-6-luna`.
|
|
8
|
+
- Registers the `:llm_judge` provider for RubyLLM 2.1's Judge API, so `RubyLLM.judge` and `RubyLLM::Judge` classes return native `Judgment` and typed answers.
|
|
9
|
+
- Runs on RubyLLM 1.13 through 1.16, which predate the Judge API. `RubyLLM::LLMJudge.judge` returns `Legacy` answers with the same readers.
|
|
10
|
+
- `:ratings` strategy (default): one 0–9 rating per answer, up to six at once, turned into a distribution with softmax. Retries a malformed digit once and breaks tied Choice ratings with a one-call judgment.
|
|
11
|
+
- `:single_request` strategy: one call returns a distribution for every question. Validates and normalizes the distributions and makes one corrective retry for malformed JSON.
|
|
12
|
+
- `max_arms` and `max_input_bytes` limits reject oversized judgments before any call.
|
data/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 JP Camara
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
data/README.md
ADDED
|
@@ -0,0 +1,223 @@
|
|
|
1
|
+
# ruby_llm-llm_judge
|
|
2
|
+
|
|
3
|
+
Use a RubyLLM chat model to answer [RubyLLM Judge](https://rubyllm.com/next/judgments/) questions. Define `probability`, `choice`, or `score` questions and receive RubyLLM's native `Judgment` and typed answers. Choose between parallel per-answer ratings and one-call typed distributions.
|
|
4
|
+
|
|
5
|
+
## Install
|
|
6
|
+
|
|
7
|
+
```ruby
|
|
8
|
+
gem 'ruby_llm-llm_judge'
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
The gem works with RubyLLM 1.13 and later. On RubyLLM 2.1 it plugs into the Judge API; on earlier versions see [RubyLLM 1.13 to 1.16](#rubyllm-113-to-116).
|
|
12
|
+
|
|
13
|
+
Configure your chat provider's credentials through RubyLLM. The default scoring model is `gpt-6-luna`; you can choose another RubyLLM chat model for each judgment.
|
|
14
|
+
|
|
15
|
+
### RubyLLM 1.13 to 1.16
|
|
16
|
+
|
|
17
|
+
The gem also runs on RubyLLM 1.13 through 1.16, which predate the Judge API. There, `RubyLLM::LLMJudge.judge` runs the same scoring directly and accepts the same questions and `provider_options`. Answers have the same readers as RubyLLM's (`probability`, `choice`, `probabilities`, `confidence`, `score`, `levels`, and `model`, `tokens`, `raw`, `[]` and `fetch` on the judgment), so call sites keep working after you upgrade to 2.1.
|
|
18
|
+
|
|
19
|
+
What differs before 2.1:
|
|
20
|
+
|
|
21
|
+
- Results are `RubyLLM::LLMJudge::Legacy` objects, not `RubyLLM::Judgment`, `RubyLLM::Choice`, and so on. Avoid class checks until you upgrade.
|
|
22
|
+
- There is no `RubyLLM.judge`, `RubyLLM::Judge` class DSL, `cost`, `context:` instrumentation, or `metadata:`. Passing `metadata:` raises.
|
|
23
|
+
- `scoring_protocol` accepts only `:chat_completions`, the API RubyLLM 1.x uses for OpenAI.
|
|
24
|
+
- The output limit and `chat_provider_options` are sent with `with_params`, using the field each provider expects (`max_completion_tokens` for OpenAI and Azure, `generationConfig.maxOutputTokens` for Gemini and Vertex AI, `inferenceConfig.maxTokens` for Bedrock, and `max_tokens` otherwise).
|
|
25
|
+
|
|
26
|
+
## Use
|
|
27
|
+
|
|
28
|
+
```ruby
|
|
29
|
+
require 'ruby_llm/llm_judge'
|
|
30
|
+
|
|
31
|
+
result = RubyLLM::LLMJudge.judge(
|
|
32
|
+
'Please refund the duplicate charge today.',
|
|
33
|
+
questions: {
|
|
34
|
+
urgent: {
|
|
35
|
+
type: :probability,
|
|
36
|
+
instructions: 'Does this need attention today?',
|
|
37
|
+
criteria: { yes: 'An explicit deadline today', no: 'No deadline today' }
|
|
38
|
+
},
|
|
39
|
+
department: {
|
|
40
|
+
type: :choice,
|
|
41
|
+
instructions: 'Which team should handle this?',
|
|
42
|
+
options: { billing: 'Payments and refunds', technical: 'Bugs and integrations' }
|
|
43
|
+
},
|
|
44
|
+
frustration: {
|
|
45
|
+
type: :score,
|
|
46
|
+
instructions: 'How frustrated is the customer?',
|
|
47
|
+
levels: ['Calm', 'Frustrated', 'Angry']
|
|
48
|
+
}
|
|
49
|
+
}
|
|
50
|
+
)
|
|
51
|
+
|
|
52
|
+
result.urgent.probability
|
|
53
|
+
result.department.choice
|
|
54
|
+
result.department.probabilities
|
|
55
|
+
result.frustration.score
|
|
56
|
+
result.model # => "gpt-6-luna"
|
|
57
|
+
result.tokens
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
You can also use a reusable Judge class:
|
|
61
|
+
|
|
62
|
+
```ruby
|
|
63
|
+
class TicketTriage < RubyLLM::Judge
|
|
64
|
+
model 'gpt-6-luna', provider: :llm_judge, assume_model_exists: true
|
|
65
|
+
probability :urgent, 'Does this need attention today?'
|
|
66
|
+
end
|
|
67
|
+
|
|
68
|
+
TicketTriage.judge('Please refund today.').urgent.probability
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
## Choose a model
|
|
72
|
+
|
|
73
|
+
Pass the chat model as `model:` and its provider as `scoring_provider`:
|
|
74
|
+
|
|
75
|
+
```ruby
|
|
76
|
+
result = RubyLLM::LLMJudge.judge(
|
|
77
|
+
'Please refund the duplicate charge today.',
|
|
78
|
+
model: 'claude-haiku-4-5',
|
|
79
|
+
questions: { urgent: { type: :probability, instructions: 'Does this need attention today?' } },
|
|
80
|
+
provider_options: {
|
|
81
|
+
scoring_provider: :anthropic
|
|
82
|
+
}
|
|
83
|
+
)
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
The default Luna rating request uses Chat Completions, temperature zero, reasoning disabled, `store: false`, and a four-token output limit. You can tune requests with `scoring_protocol`, `temperature`, `max_output_tokens`, and `chat_provider_options`.
|
|
87
|
+
|
|
88
|
+
## One-call typed decisions
|
|
89
|
+
|
|
90
|
+
Set `strategy: :single_request` to send the state and every Judge question to the model in one request:
|
|
91
|
+
|
|
92
|
+
```ruby
|
|
93
|
+
result = RubyLLM::LLMJudge.judge(
|
|
94
|
+
'Please refund the duplicate charge today.',
|
|
95
|
+
questions: {
|
|
96
|
+
urgent: { type: :probability, instructions: 'Does this need attention today?' },
|
|
97
|
+
department: {
|
|
98
|
+
type: :choice,
|
|
99
|
+
instructions: 'Which team should handle this?',
|
|
100
|
+
options: { billing: 'Payments and refunds', technical: 'Bugs and integrations' }
|
|
101
|
+
}
|
|
102
|
+
},
|
|
103
|
+
provider_options: { strategy: :single_request, max_output_tokens: 1024 }
|
|
104
|
+
)
|
|
105
|
+
|
|
106
|
+
result.urgent.probability
|
|
107
|
+
result.department.probabilities
|
|
108
|
+
```
|
|
109
|
+
|
|
110
|
+
The model returns one JSON object containing a distribution for each question. The gem checks that every question and option is present, validates each probability, normalizes each distribution, and builds RubyLLM's typed answers. `result.raw[:reported_probabilities]` and `result.raw[:reported_totals]` preserve the model's original numbers for inspection. With an OpenAI key configured, this example calls Luna directly through Chat Completions with reasoning disabled, temperature zero, and `store: false`. The 1024-token limit suits small judgments; omit it for large question sets to use the 8192-token default. It makes one corrective retry for malformed JSON or missing fields, counting both calls in `result.tokens` and `result.raw[:attempts]`. Set `malformed_retries: 0` to disable that retry. An invalid response after retries raises an error.
|
|
111
|
+
|
|
112
|
+
If you use OpenRouter for Luna, prioritize the provider with the lowest observed latency:
|
|
113
|
+
|
|
114
|
+
```ruby
|
|
115
|
+
result = RubyLLM::LLMJudge.judge(
|
|
116
|
+
'Please refund the duplicate charge today.',
|
|
117
|
+
model: 'openai/gpt-6-luna',
|
|
118
|
+
questions: {
|
|
119
|
+
department: {
|
|
120
|
+
type: :choice,
|
|
121
|
+
instructions: 'Which team should handle this?',
|
|
122
|
+
options: { billing: 'Payments and refunds', technical: 'Bugs and integrations' }
|
|
123
|
+
}
|
|
124
|
+
},
|
|
125
|
+
provider_options: {
|
|
126
|
+
strategy: :single_request,
|
|
127
|
+
scoring_provider: :openrouter,
|
|
128
|
+
temperature: 0,
|
|
129
|
+
max_output_tokens: 1024,
|
|
130
|
+
chat_provider_options: {
|
|
131
|
+
reasoning: { enabled: false },
|
|
132
|
+
provider: { sort: 'latency' }
|
|
133
|
+
}
|
|
134
|
+
}
|
|
135
|
+
)
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
OpenRouter's [latency sorting](https://openrouter.ai/docs/guides/routing/provider-selection) chooses among its available providers using their observed response times. It sped up OpenRouter in our test, while direct OpenAI was faster still. The measured speed and answer-quality tradeoffs are below.
|
|
139
|
+
|
|
140
|
+
## How scoring works
|
|
141
|
+
|
|
142
|
+
The default `:ratings` strategy asks the model to rate every declared answer from 0 to 9. It makes one call per answer, with up to six calls running at once, then applies softmax to the ratings. It retries a malformed digit once. When the highest ratings tie on a Choice question, it makes a one-call JSON judgment for that question and uses its distribution. If that call also ties, it raises an error instead of choosing whichever option came first. `result.raw[:tie_breaks]` records each resolution, and token usage includes the extra call. Set `tie_breaker: :first` to use the original first-option rule, or `tie_break_max_output_tokens:` to change the tie-break response limit (default 8192).
|
|
143
|
+
|
|
144
|
+
The `:single_request` strategy asks for all distributions in one call. Both strategies build the same RubyLLM answer types: a `choice` returns the selected answer and distribution; a `score` returns the distribution and its weighted level; a `probability` returns the positive answer's share. The gem supports Judge's 1–255 choice options, 2–10 score levels, and multiple questions in one judgment.
|
|
145
|
+
|
|
146
|
+
These probabilities compare the answers you supplied. The `:single_request` values are reported by the chat model; the default ratings values come from softmax over its 0–9 ratings. Neither strategy establishes calibration by itself. Include an `other` or `escalate` choice when the named answers may not cover the input. `confidence` measures how concentrated the returned distribution is. Use labeled examples to set any automation thresholds.
|
|
147
|
+
|
|
148
|
+
Set optional `max_arms` and `max_input_bytes` in `provider_options` to cap work per judgment. For example, `{ max_arms: 12, max_input_bytes: 32_768 }` rejects larger requests before scoring. Token usage is aggregated across calls, and `raw` contains the ratings and model metadata.
|
|
149
|
+
|
|
150
|
+
The one-digit-per-answer approach was inspired by [@burkov's post on Jev](https://x.com/burkov/status/2102438687392797045). The one-call strategy follows the typed-question pattern in [TypeSafe's System One adapter](https://github.com/typesafe-ai/system-one-adapter-python).
|
|
151
|
+
|
|
152
|
+
## Benchmarks
|
|
153
|
+
|
|
154
|
+
On September 23, 2026, we tested GPT-6 Luna (OpenRouter), DeepSeek V4.1 Flash (Fireworks), [Celeris-1](https://docs.celeris.ai/making-requests), and Jev 1.13.0 on 64 balanced [AG News](https://huggingface.co/datasets/fancyzhx/ag_news) articles with four choices and 32 balanced [SST-2](https://huggingface.co/datasets/stanfordnlp/sst2) sentences. LLM runs used temperature zero with reasoning disabled. **Correct** counts all cases; **Usable** counts cases that returned a valid judgment. Brier scores and latency use usable cases only. Latency is the full client round trip, including retries.
|
|
155
|
+
|
|
156
|
+
| AG News model | Method | Correct / 64 | Usable / 64 | Brier ↓ | Median / p90 latency |
|
|
157
|
+
| --- | --- | ---: | ---: | ---: | ---: |
|
|
158
|
+
| Luna | One call | 59 | 64 | 0.1444 | 1,320 / 1,708 ms |
|
|
159
|
+
| Luna | Parallel ratings | 53 | 64 | 0.1578 | 1,718 / 2,753 ms |
|
|
160
|
+
| DeepSeek | One call | 59 | 64 | 0.1453 | 1,001 / 2,197 ms |
|
|
161
|
+
| DeepSeek | Parallel ratings | 56 | 64 | 0.1513 | 2,577 / 3,531 ms |
|
|
162
|
+
| Celeris | Plain one call | 54 | 58 | 0.1284 | 428 / 744 ms |
|
|
163
|
+
| Celeris | Original ratings | 42 | 54 | 0.1998 | 505 / 861 ms |
|
|
164
|
+
| Celeris | Original ratings, rerun | 47 | 58 | 0.1754 | 537 / 889 ms |
|
|
165
|
+
| Celeris | Ratings + retry/tie-break, run 1 | 57 | 61 | 0.1284 | 545 / 1,137 ms |
|
|
166
|
+
| Celeris | Ratings + retry/tie-break, run 2 | 56 | 59 | 0.1153 | 545 / 1,009 ms |
|
|
167
|
+
| Celeris | JSON one call + retry, run 1 | 60 | 63 | 0.0958 | 453 / 852 ms |
|
|
168
|
+
| Celeris | JSON one call + retry, run 2 | 59 | 62 | 0.1107 | 401 / 776 ms |
|
|
169
|
+
| Jev | Luna paired run | 59 | 64 | 0.1032 | 365 / 586 ms |
|
|
170
|
+
| Jev | DeepSeek paired run | 59 | 64 | 0.1040 | 505 / 1,060 ms |
|
|
171
|
+
|
|
172
|
+
| SST-2 model | Method | Choice / 32 | Probability / 32 | Score / 32 | Usable / 32 | Median latency |
|
|
173
|
+
| --- | --- | ---: | ---: | ---: | ---: | ---: |
|
|
174
|
+
| Luna | One call | 29 | 29 | 29 | 32 | 1,858 ms |
|
|
175
|
+
| Luna | Parallel ratings | 28 | 30 | 28 | 32 | 3,437 ms |
|
|
176
|
+
| DeepSeek | One call | 29 | 29 | 29 | 32 | 1,443 ms |
|
|
177
|
+
| DeepSeek | Parallel ratings | 30 | 26 | 28 | 32 | 3,009 ms |
|
|
178
|
+
| Celeris | Plain one call | 29 | 29 | 29 | 32 | 282 ms |
|
|
179
|
+
| Celeris | Original ratings | 28 | 24 | 26 | 30 | 530 ms |
|
|
180
|
+
| Celeris | Ratings + retry/tie-break, run 1 | 28 | 24 | 28 | 32 | 577 ms |
|
|
181
|
+
| Celeris | JSON one call + retry, run 1 | 28 | 28 | 28 | 32 | 248 ms |
|
|
182
|
+
| Celeris | JSON one call + retry, run 2 | 28 | 28 | 28 | 32 | 264 ms |
|
|
183
|
+
| Jev | Luna paired run | 30 | 30 | 30 | 32 | 522 ms |
|
|
184
|
+
| Jev | DeepSeek paired run | 30 | 29 | 30 | 32 | 308 ms |
|
|
185
|
+
|
|
186
|
+
The original 42/64 Celeris ratings result was not a one-off: the same behavior scored 47/64 in a fresh control run. In the original run, 10 cases returned no usable judgment and nine of the 12 incorrect usable choices had tied top ratings. With malformed-digit retry and Choice tie resolution, two runs scored 57/64 and 56/64. Those runs resolved nine of ten and seven of seven top ties to the reference label. Their observed median and p90 latencies were higher; tie-break calls add work. The control and revised run 2 covered only the 64 news cases; revised run 1 covered all 112 requests. Jev rows are saved earlier paired runs on the same cases, while Celeris was measured later.
|
|
187
|
+
|
|
188
|
+
### Luna latency paths
|
|
189
|
+
|
|
190
|
+
In a paired API-level run using the gem's one-call prompt, direct OpenAI was faster than latency-sorted OpenRouter. Both used temperature zero, reasoning disabled, a 1024-token output limit, and one corrective retry for malformed JSON. Four direct cases were also verified through the gem. All final judgments were usable.
|
|
191
|
+
|
|
192
|
+
| Dataset | Luna path | Correct | Median | p90 | Brier ↓ |
|
|
193
|
+
| --- | --- | ---: | ---: | ---: | ---: |
|
|
194
|
+
| AG News (64) | Direct OpenAI | 58/64 | 823 ms | 1,241 ms | 0.1405 |
|
|
195
|
+
| AG News (64) | OpenRouter latency sort | 59/64 | 1,379 ms | 1,938 ms | 0.1265 |
|
|
196
|
+
| SST-2 (32) | Direct OpenAI | 28/32 | 843 ms | 1,745 ms | 0.1042 |
|
|
197
|
+
| SST-2 (32) | OpenRouter latency sort | 29/32 | 1,423 ms | 3,829 ms | 0.0762 |
|
|
198
|
+
|
|
199
|
+
In that paired run, direct OpenAI cut median latency by about 40% on both datasets, but lost one correct label on each and had worse Brier scores. Its SST-2 calls needed five corrective retries versus one through OpenRouter. A later direct-only repeat scored **59/64 news at 868 ms median** and **28/32 SST-2 at 1,022 ms median**, all usable. Just one news label and two sentiment labels changed across direct runs. OpenRouter was not rerun simultaneously with the repeat. [Methods and both runs' case-level results](benchmarks/luna-direct.md).
|
|
200
|
+
|
|
201
|
+
A separate paired gem run sent the same one-call Luna judgments through OpenRouter's default route and `provider: { sort: 'latency' }`, rotating request order across cases. All responses were usable. Each AG News request asked one four-choice question; each SST-2 request asked Choice, Probability, and Score together. Both routes used temperature zero, reasoning disabled, and a 1024-token output limit. These results are separate from the runs above.
|
|
202
|
+
|
|
203
|
+
| Dataset | OpenRouter route | Correct | Median | p90 | Brier ↓ |
|
|
204
|
+
| --- | --- | ---: | ---: | ---: | ---: |
|
|
205
|
+
| AG News (64) | Default | 61/64 | 1,569 ms | 2,805 ms | 0.0857 |
|
|
206
|
+
| AG News (64) | Lowest latency | 60/64 | 1,260 ms | 1,812 ms | 0.1165 |
|
|
207
|
+
| SST-2 (32) | Default | 29/32 | 2,093 ms | 8,693 ms | 0.0775 |
|
|
208
|
+
| SST-2 (32) | Lowest latency | 30/32 | 1,507 ms | 1,944 ms | 0.0695 |
|
|
209
|
+
|
|
210
|
+
Latency sorting reduced median time by 20% on AG News and 28% on SST-2. It was faster on 50/64 and 21/32 matched cases, respectively. AG News lost one correct label and its Brier score worsened; SST-2 gained one correct label. Check this route on your own labeled decisions before relying on its probabilities. [Methods and case-level data](benchmarks/luna-routing.md).
|
|
211
|
+
|
|
212
|
+
## Development
|
|
213
|
+
|
|
214
|
+
The gem uses RubyLLM's Judge API when it is present and the `Legacy` path otherwise. `test/llm_judge_test.rb` covers the Judge API and `test/legacy_test.rb` covers RubyLLM 1.13 through 1.16; each skips on the other side.
|
|
215
|
+
|
|
216
|
+
```bash
|
|
217
|
+
bundle exec rake test # the latest released RubyLLM
|
|
218
|
+
RUBY_LLM_VERSION=1.13.2 bundle exec rake test # a specific release
|
|
219
|
+
RUBY_LLM_VERSION=main bundle exec rake test # RubyLLM's main branch, with the Judge API
|
|
220
|
+
RUBY_LLM_PATH=../ruby_llm bundle exec rake test # a local checkout
|
|
221
|
+
```
|
|
222
|
+
|
|
223
|
+
CI runs Ruby 3.2 through 4.0 against RubyLLM 1.13.2 and 1.16.0, and Ruby 3.4 against RubyLLM main.
|
|
@@ -0,0 +1,346 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require 'json'
|
|
4
|
+
|
|
5
|
+
module RubyLLM
|
|
6
|
+
module LLMJudge
|
|
7
|
+
class Engine
|
|
8
|
+
MAX_WORKERS = 6
|
|
9
|
+
SYSTEM_INSTRUCTIONS = 'Return exactly one ASCII digit 0-9. 9 means very likely to be the correct answer; 0 means very unlikely. No explanation.'
|
|
10
|
+
SINGLE_REQUEST_INSTRUCTIONS = 'Answer all questions using only a JSON object with an "answers" field. ' \
|
|
11
|
+
'For each question, return an object mapping every supplied option ID to a ' \
|
|
12
|
+
'probability between 0 and 1. Include every option exactly once and make ' \
|
|
13
|
+
'each question\'s probabilities sum to 1. Evaluate each question independently ' \
|
|
14
|
+
'against the same state. Return no explanations or markdown.'
|
|
15
|
+
|
|
16
|
+
def initialize(config:, model: DEFAULT_MODEL, provider_options: {}, scorer: nil, responder: nil)
|
|
17
|
+
@config = config
|
|
18
|
+
options = provider_options.transform_keys(&:to_sym)
|
|
19
|
+
allowed = %i[scoring_provider scoring_protocol temperature max_output_tokens chat_provider_options
|
|
20
|
+
max_arms max_input_bytes strategy malformed_retries tie_breaker tie_break_max_output_tokens]
|
|
21
|
+
raise ArgumentError, 'Unknown LLMJudge provider options' unless (options.keys - allowed).empty?
|
|
22
|
+
|
|
23
|
+
@strategy = options.fetch(:strategy, :ratings).to_sym
|
|
24
|
+
raise ArgumentError, 'strategy must be :ratings or :single_request' unless %i[ratings single_request].include?(@strategy)
|
|
25
|
+
@tie_breaker = options.fetch(:tie_breaker, :single_request).to_sym
|
|
26
|
+
raise ArgumentError, 'tie_breaker must be :single_request or :first' unless %i[single_request first].include?(@tie_breaker)
|
|
27
|
+
|
|
28
|
+
@scoring_provider = options.fetch(:scoring_provider, :openai).to_sym
|
|
29
|
+
@scoring_model = model
|
|
30
|
+
raise ArgumentError, 'A model is required' unless @scoring_model.is_a?(String) && !@scoring_model.empty?
|
|
31
|
+
|
|
32
|
+
luna_defaults = @scoring_provider == :openai && @scoring_model == DEFAULT_MODEL
|
|
33
|
+
@scoring_protocol = options.fetch(:scoring_protocol, luna_defaults ? :chat_completions : nil)
|
|
34
|
+
# Before 2.1, RubyLLM chats use each provider's one chat API (Chat Completions for OpenAI).
|
|
35
|
+
if !NATIVE && ![nil, :chat_completions].include?(@scoring_protocol&.to_sym)
|
|
36
|
+
raise ArgumentError, 'scoring_protocol requires RubyLLM 2.1'
|
|
37
|
+
end
|
|
38
|
+
@temperature = options.fetch(:temperature, luna_defaults ? 0 : nil)
|
|
39
|
+
@max_output_tokens = options.fetch(:max_output_tokens, @strategy == :single_request ? 8192 : 4)
|
|
40
|
+
@tie_break_max_output_tokens = options.fetch(:tie_break_max_output_tokens, 8192)
|
|
41
|
+
@chat_provider_options = options.fetch(:chat_provider_options, default_chat_options(luna_defaults))
|
|
42
|
+
@max_arms = options[:max_arms]
|
|
43
|
+
@max_input_bytes = options[:max_input_bytes]
|
|
44
|
+
@malformed_retries = options.fetch(:malformed_retries, 1)
|
|
45
|
+
raise ArgumentError, 'chat_provider_options must be a Hash' unless @chat_provider_options.is_a?(Hash)
|
|
46
|
+
raise ArgumentError, 'max_output_tokens must be positive' unless @max_output_tokens.is_a?(Integer) && @max_output_tokens.positive?
|
|
47
|
+
unless @tie_break_max_output_tokens.is_a?(Integer) && @tie_break_max_output_tokens.positive?
|
|
48
|
+
raise ArgumentError, 'tie_break_max_output_tokens must be positive'
|
|
49
|
+
end
|
|
50
|
+
unless [@max_arms, @max_input_bytes].all? { |limit| limit.nil? || (limit.is_a?(Integer) && limit.positive?) }
|
|
51
|
+
raise ArgumentError, 'LLMJudge limits must be positive integers'
|
|
52
|
+
end
|
|
53
|
+
unless @malformed_retries.is_a?(Integer) && @malformed_retries >= 0
|
|
54
|
+
raise ArgumentError, 'malformed_retries must be a nonnegative integer'
|
|
55
|
+
end
|
|
56
|
+
|
|
57
|
+
@scorer = scorer || method(:score_with_model)
|
|
58
|
+
@responder = responder || method(:respond_with_model)
|
|
59
|
+
end
|
|
60
|
+
|
|
61
|
+
def judge(input, questions:, model:)
|
|
62
|
+
return judge_single_request(input, questions, model) if @strategy == :single_request
|
|
63
|
+
|
|
64
|
+
jobs = build_jobs(input, questions)
|
|
65
|
+
results = run(jobs) { |job| @scorer.call(job[:prompt]) }
|
|
66
|
+
tie_breaks = []
|
|
67
|
+
answers = questions.values.to_h do |question|
|
|
68
|
+
ratings = jobs.each_index.filter_map { |index|
|
|
69
|
+
results[index].fetch(:digit) if jobs[index][:question] == question
|
|
70
|
+
}
|
|
71
|
+
chosen = answer(question, softmax(ratings))
|
|
72
|
+
if question.type == :choice && ratings.count(ratings.max) > 1 && @tie_breaker == :single_request
|
|
73
|
+
judgment = judge_single_request(input, { question.name => question }, model)
|
|
74
|
+
chosen = judgment.answers.fetch(question.name)
|
|
75
|
+
probabilities = chosen.probabilities.values
|
|
76
|
+
if probabilities.count(probabilities.max) > 1
|
|
77
|
+
raise RubyLLM::Error, "Scoring model could not resolve the tie for #{question.name}"
|
|
78
|
+
end
|
|
79
|
+
tie_breaks << { question: question.name, ratings:, judgment: }
|
|
80
|
+
end
|
|
81
|
+
[question.name, chosen]
|
|
82
|
+
end
|
|
83
|
+
tokens = aggregate_tokens(results.map { |result| result.fetch(:tokens) } +
|
|
84
|
+
tie_breaks.map { |item| item[:judgment].tokens })
|
|
85
|
+
Types::Judgment.new(
|
|
86
|
+
answers:, model: model.id, tokens:,
|
|
87
|
+
raw: { method: 'parallel_0_to_9_softmax', scoring_provider: @scoring_provider,
|
|
88
|
+
scoring_model: @scoring_model,
|
|
89
|
+
tie_breaks: tie_breaks.map { |item|
|
|
90
|
+
{ question: item[:question], ratings: item[:ratings],
|
|
91
|
+
probabilities: item[:judgment].answers.fetch(item[:question]).probabilities,
|
|
92
|
+
attempts: item[:judgment].raw[:attempts] }
|
|
93
|
+
},
|
|
94
|
+
ratings: jobs.each_with_index.map { |job, index|
|
|
95
|
+
{ question: job[:question].name, option: job[:name], digit: results[index][:digit] }
|
|
96
|
+
} }
|
|
97
|
+
)
|
|
98
|
+
end
|
|
99
|
+
|
|
100
|
+
private
|
|
101
|
+
|
|
102
|
+
def judge_single_request(input, questions, model)
|
|
103
|
+
specs = validated_options(input, questions)
|
|
104
|
+
original_prompt = single_request_prompt(input, specs)
|
|
105
|
+
prompt = original_prompt
|
|
106
|
+
attempts = []
|
|
107
|
+
loop do
|
|
108
|
+
result = @responder.call(prompt)
|
|
109
|
+
attempts << result
|
|
110
|
+
begin
|
|
111
|
+
answers, reported, normalization = parse_single_request(result.fetch(:content), specs)
|
|
112
|
+
return Types::Judgment.new(
|
|
113
|
+
answers:, model: model.id,
|
|
114
|
+
tokens: aggregate_tokens(attempts.map { |attempt| attempt.fetch(:tokens) }),
|
|
115
|
+
raw: { method: 'single_request_probabilities', scoring_provider: @scoring_provider,
|
|
116
|
+
scoring_model: @scoring_model, reported_probabilities: reported,
|
|
117
|
+
reported_totals: normalization, attempts: attempts.size }
|
|
118
|
+
)
|
|
119
|
+
rescue RubyLLM::Error => error
|
|
120
|
+
raise if attempts.size > @malformed_retries
|
|
121
|
+
|
|
122
|
+
prompt = "#{original_prompt}\n\nThe previous answer was invalid: #{error.message}. " \
|
|
123
|
+
"Previous answer: #{result.fetch(:content).to_s[0, 4000]}\n" \
|
|
124
|
+
'Return a complete corrected JSON object with every question and option.'
|
|
125
|
+
end
|
|
126
|
+
end
|
|
127
|
+
end
|
|
128
|
+
|
|
129
|
+
def parse_single_request(content, specs)
|
|
130
|
+
parsed = JSON.parse(content)
|
|
131
|
+
raise RubyLLM::Error, 'Scoring model returned a non-object response' unless parsed.is_a?(Hash)
|
|
132
|
+
|
|
133
|
+
reported = parsed.fetch('answers')
|
|
134
|
+
expected_questions = specs.map { |question, _| question.name.to_s }
|
|
135
|
+
unless reported.is_a?(Hash) && reported.keys.sort == expected_questions.sort
|
|
136
|
+
actual = reported.is_a?(Hash) ? reported.keys : reported.class.name
|
|
137
|
+
raise RubyLLM::Error, "Question IDs must be #{expected_questions.inspect}; got #{actual.inspect}"
|
|
138
|
+
end
|
|
139
|
+
|
|
140
|
+
normalization = {}
|
|
141
|
+
answers = specs.to_h do |question, options|
|
|
142
|
+
distribution = reported.fetch(question.name.to_s)
|
|
143
|
+
keys = options.map { |name, _| name.to_s }
|
|
144
|
+
unless distribution.is_a?(Hash) && distribution.keys.sort == keys.sort
|
|
145
|
+
actual = distribution.is_a?(Hash) ? distribution.keys : distribution.class.name
|
|
146
|
+
raise RubyLLM::Error, "Option IDs for #{question.name} must be #{keys.inspect}; got #{actual.inspect}"
|
|
147
|
+
end
|
|
148
|
+
|
|
149
|
+
values = keys.map { |key| distribution.fetch(key) }
|
|
150
|
+
unless values.all? { |value| value.is_a?(Numeric) && value.finite? && (0..1).cover?(value) }
|
|
151
|
+
raise RubyLLM::Error, "Scoring model returned invalid probabilities for #{question.name}"
|
|
152
|
+
end
|
|
153
|
+
|
|
154
|
+
total = values.sum.to_f
|
|
155
|
+
raise RubyLLM::Error, "Scoring model returned zero probability for #{question.name}" unless total.positive?
|
|
156
|
+
|
|
157
|
+
normalization[question.name.to_s] = total
|
|
158
|
+
[question.name, answer(question, values.map { |value| value / total })]
|
|
159
|
+
end
|
|
160
|
+
[answers, reported, normalization]
|
|
161
|
+
rescue JSON::ParserError, KeyError, TypeError => error
|
|
162
|
+
raise RubyLLM::Error, "Scoring model returned invalid JSON probabilities: #{error.message}"
|
|
163
|
+
end
|
|
164
|
+
|
|
165
|
+
def single_request_prompt(input, specs)
|
|
166
|
+
request = {
|
|
167
|
+
state: input,
|
|
168
|
+
questions: specs.map do |question, options|
|
|
169
|
+
{ id: question.name, type: question.type, instructions: question.instructions,
|
|
170
|
+
options: options.map { |name, description| { id: name.to_s, description: } } }
|
|
171
|
+
end
|
|
172
|
+
}
|
|
173
|
+
"Evaluate this decision request. Return only the requested JSON object.\n#{JSON.generate(request)}"
|
|
174
|
+
end
|
|
175
|
+
|
|
176
|
+
def build_jobs(input, questions)
|
|
177
|
+
jobs = validated_options(input, questions).flat_map do |question, options|
|
|
178
|
+
options.map do |name, description|
|
|
179
|
+
{ question:, name:, description:, prompt: prompt(input, question.instructions, name, description) }
|
|
180
|
+
end
|
|
181
|
+
end
|
|
182
|
+
jobs
|
|
183
|
+
end
|
|
184
|
+
|
|
185
|
+
def validated_options(input, questions)
|
|
186
|
+
if @max_input_bytes
|
|
187
|
+
input_bytes = serialize(input).bytesize + questions.values.sum do |question|
|
|
188
|
+
serialize(question.instructions).bytesize + serialize(question.criteria).bytesize
|
|
189
|
+
end
|
|
190
|
+
raise ArgumentError, "LLMJudge input exceeds #{@max_input_bytes} bytes" if input_bytes > @max_input_bytes
|
|
191
|
+
end
|
|
192
|
+
|
|
193
|
+
specs = questions.values.map { |question| [question, options_for(question)] }
|
|
194
|
+
arms = specs.sum { |_, options| options.size }
|
|
195
|
+
raise ArgumentError, "LLMJudge exceeds #{@max_arms} scoring arms" if @max_arms && arms > @max_arms
|
|
196
|
+
|
|
197
|
+
specs
|
|
198
|
+
end
|
|
199
|
+
|
|
200
|
+
def options_for(question)
|
|
201
|
+
case question.type
|
|
202
|
+
when :probability
|
|
203
|
+
criteria = question.criteria || {}
|
|
204
|
+
yes_key = criteria.keys.find { |key| %w[yes true].include?(key.to_s) }
|
|
205
|
+
no_key = criteria.keys.find { |key| %w[no false].include?(key.to_s) }
|
|
206
|
+
yes = yes_key ? criteria[yes_key] : 'Yes, the condition holds'
|
|
207
|
+
no = no_key ? criteria[no_key] : 'No, the condition does not hold'
|
|
208
|
+
[['true', yes], ['false', no]]
|
|
209
|
+
when :choice
|
|
210
|
+
raise ArgumentError, 'A choice needs 1–255 options' unless (1..255).cover?(question.criteria.size)
|
|
211
|
+
|
|
212
|
+
question.criteria.to_a
|
|
213
|
+
when :score
|
|
214
|
+
raise ArgumentError, 'A score needs 2–10 levels' unless (2..10).cover?(question.criteria.size)
|
|
215
|
+
|
|
216
|
+
question.criteria.each_with_index.map { |description, index| [index, description] }
|
|
217
|
+
else
|
|
218
|
+
raise ArgumentError, "Unsupported question type: #{question.type}"
|
|
219
|
+
end
|
|
220
|
+
end
|
|
221
|
+
|
|
222
|
+
def prompt(input, instructions, name, description)
|
|
223
|
+
"Question: #{serialize(instructions || 'Which answer best applies?')}\n" \
|
|
224
|
+
"State: #{serialize(input)}\n\n" \
|
|
225
|
+
'Output a single number between 0 and 9 to tell how likely it is that the correct answer is: ' \
|
|
226
|
+
"#{name} — #{serialize(description)}"
|
|
227
|
+
end
|
|
228
|
+
|
|
229
|
+
def serialize(value)
|
|
230
|
+
value.is_a?(String) ? value : JSON.generate(value)
|
|
231
|
+
end
|
|
232
|
+
|
|
233
|
+
def run(jobs)
|
|
234
|
+
queue = Queue.new
|
|
235
|
+
jobs.each_with_index { |job, index| queue << [index, job] }
|
|
236
|
+
results = Array.new(jobs.size)
|
|
237
|
+
workers = Array.new([jobs.size, MAX_WORKERS].min) do
|
|
238
|
+
Thread.new do
|
|
239
|
+
Thread.current.report_on_exception = false
|
|
240
|
+
loop do
|
|
241
|
+
begin
|
|
242
|
+
index, job = queue.pop(true)
|
|
243
|
+
rescue ThreadError
|
|
244
|
+
break
|
|
245
|
+
end
|
|
246
|
+
results[index] = yield job
|
|
247
|
+
end
|
|
248
|
+
end
|
|
249
|
+
end
|
|
250
|
+
workers.each(&:join)
|
|
251
|
+
workers.each(&:value)
|
|
252
|
+
results
|
|
253
|
+
end
|
|
254
|
+
|
|
255
|
+
def default_chat_options(luna_defaults)
|
|
256
|
+
return {} unless @scoring_provider == :openai
|
|
257
|
+
|
|
258
|
+
luna_defaults ? { store: false, reasoning_effort: 'none' } : { store: false }
|
|
259
|
+
end
|
|
260
|
+
|
|
261
|
+
def score_with_model(prompt)
|
|
262
|
+
attempts = []
|
|
263
|
+
loop do
|
|
264
|
+
request = attempts.empty? ? prompt : "#{prompt}\n\nYour previous answer was invalid. Return exactly one ASCII digit 0-9."
|
|
265
|
+
response = scoring_chat.ask(request)
|
|
266
|
+
attempts << response
|
|
267
|
+
digit = response.content.to_s.strip
|
|
268
|
+
if /\A[0-9]\z/.match?(digit)
|
|
269
|
+
return { digit: digit.to_i, tokens: aggregate_tokens(attempts.map(&:tokens)) }
|
|
270
|
+
end
|
|
271
|
+
raise RubyLLM::Error, 'Scoring model returned no single digit' if attempts.size > @malformed_retries
|
|
272
|
+
end
|
|
273
|
+
end
|
|
274
|
+
|
|
275
|
+
def respond_with_model(prompt)
|
|
276
|
+
limit = @strategy == :ratings ? @tie_break_max_output_tokens : @max_output_tokens
|
|
277
|
+
response = scoring_chat(instructions: SINGLE_REQUEST_INSTRUCTIONS, max_output_tokens: limit).ask(prompt)
|
|
278
|
+
{ content: response.content.to_s, tokens: response.tokens }
|
|
279
|
+
end
|
|
280
|
+
|
|
281
|
+
def scoring_chat(instructions: SYSTEM_INSTRUCTIONS, max_output_tokens: @max_output_tokens)
|
|
282
|
+
context = RubyLLM::Context.new(@config)
|
|
283
|
+
chat = if NATIVE
|
|
284
|
+
context.chat(model: @scoring_model, provider: @scoring_provider,
|
|
285
|
+
protocol: @scoring_protocol, assume_model_exists: true)
|
|
286
|
+
else
|
|
287
|
+
context.chat(model: @scoring_model, provider: @scoring_provider, assume_model_exists: true)
|
|
288
|
+
end
|
|
289
|
+
chat.with_instructions(instructions)
|
|
290
|
+
chat.with_temperature(@temperature) unless @temperature.nil?
|
|
291
|
+
if NATIVE
|
|
292
|
+
chat.with_max_output_tokens(max_output_tokens)
|
|
293
|
+
chat.with_provider_options(@chat_provider_options)
|
|
294
|
+
else
|
|
295
|
+
chat.with_params(**RubyLLM::Utils.deep_merge(max_output_tokens_param(max_output_tokens), @chat_provider_options))
|
|
296
|
+
end
|
|
297
|
+
chat
|
|
298
|
+
end
|
|
299
|
+
|
|
300
|
+
# The request field each provider reads for the output limit, as RubyLLM 2.1 renders it.
|
|
301
|
+
def max_output_tokens_param(limit)
|
|
302
|
+
case @scoring_provider
|
|
303
|
+
when :openai, :azure then { max_completion_tokens: limit }
|
|
304
|
+
when :gemini, :vertexai then { generationConfig: { maxOutputTokens: limit } }
|
|
305
|
+
when :bedrock then { inferenceConfig: { maxTokens: limit } }
|
|
306
|
+
else { max_tokens: limit }
|
|
307
|
+
end
|
|
308
|
+
end
|
|
309
|
+
|
|
310
|
+
def aggregate_tokens(tokens)
|
|
311
|
+
NATIVE ? RubyLLM::Tokens.aggregate(tokens) : Legacy.aggregate_tokens(tokens)
|
|
312
|
+
end
|
|
313
|
+
|
|
314
|
+
def answer(question, probabilities)
|
|
315
|
+
case question.type
|
|
316
|
+
when :probability
|
|
317
|
+
Types::Probability.new(probability: probabilities.first)
|
|
318
|
+
when :choice
|
|
319
|
+
names = question.criteria.keys
|
|
320
|
+
distribution = names.zip(probabilities).to_h
|
|
321
|
+
Types::Choice.new(choice: names[probabilities.index(probabilities.max)], probabilities: distribution,
|
|
322
|
+
confidence: concentration(probabilities))
|
|
323
|
+
when :score
|
|
324
|
+
Types::Score.new(score: probabilities.each_with_index.sum { |probability, index| probability * index },
|
|
325
|
+
levels: question.criteria,
|
|
326
|
+
probabilities: probabilities.each_with_index.to_h { |probability, index| [index, probability] },
|
|
327
|
+
confidence: concentration(probabilities))
|
|
328
|
+
end
|
|
329
|
+
end
|
|
330
|
+
|
|
331
|
+
def softmax(ratings)
|
|
332
|
+
peak = ratings.max
|
|
333
|
+
weights = ratings.map { |rating| Math.exp(rating - peak) }
|
|
334
|
+
weights.map { |weight| weight / weights.sum }
|
|
335
|
+
end
|
|
336
|
+
|
|
337
|
+
def concentration(probabilities)
|
|
338
|
+
return 0.0 if probabilities.one?
|
|
339
|
+
return 0.0 if probabilities.count(probabilities.max) > 1
|
|
340
|
+
|
|
341
|
+
entropy = -probabilities.sum { |probability| probability.zero? ? 0.0 : probability * Math.log(probability) }
|
|
342
|
+
[[1.0 - entropy / Math.log(probabilities.size), 0.0].max, 1.0].min
|
|
343
|
+
end
|
|
344
|
+
end
|
|
345
|
+
end
|
|
346
|
+
end
|
|
@@ -0,0 +1,150 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module RubyLLM
|
|
4
|
+
module LLMJudge
|
|
5
|
+
# Stand-ins for RubyLLM 2.1's Judge types, used when RubyLLM predates the
|
|
6
|
+
# Judge API. They expose the same readers, so code written against
|
|
7
|
+
# RubyLLM::LLMJudge.judge keeps working after upgrading.
|
|
8
|
+
module Legacy
|
|
9
|
+
Model = Data.define(:id)
|
|
10
|
+
|
|
11
|
+
class Question
|
|
12
|
+
CRITERIA_KEYS = { probability: :criteria, choice: :options, score: :levels }.freeze
|
|
13
|
+
|
|
14
|
+
attr_reader :name, :type, :instructions, :criteria
|
|
15
|
+
|
|
16
|
+
def self.from_h(name, definition)
|
|
17
|
+
raise ArgumentError, 'Each question must be a Hash' unless definition.is_a?(Hash)
|
|
18
|
+
|
|
19
|
+
definition = definition.transform_keys(&:to_sym)
|
|
20
|
+
type = definition[:type]&.to_sym
|
|
21
|
+
key = CRITERIA_KEYS.fetch(type) { raise ArgumentError, "Unknown judgment type: #{type.inspect}" }
|
|
22
|
+
extra = definition.keys - [:type, :instructions, key]
|
|
23
|
+
raise ArgumentError, "Unknown question options: #{extra.join(', ')}" unless extra.empty?
|
|
24
|
+
|
|
25
|
+
new(name, type:, instructions: definition[:instructions], criteria: definition[key])
|
|
26
|
+
end
|
|
27
|
+
|
|
28
|
+
def initialize(name, type:, instructions:, criteria:)
|
|
29
|
+
raise ArgumentError, 'A question name must be a String or Symbol' unless name.is_a?(String) || name.is_a?(Symbol)
|
|
30
|
+
|
|
31
|
+
@name = name
|
|
32
|
+
@type = type
|
|
33
|
+
@instructions = instructions
|
|
34
|
+
@criteria = criteria
|
|
35
|
+
validate!
|
|
36
|
+
freeze
|
|
37
|
+
end
|
|
38
|
+
|
|
39
|
+
private
|
|
40
|
+
|
|
41
|
+
def validate!
|
|
42
|
+
case type
|
|
43
|
+
when :probability
|
|
44
|
+
return if criteria.nil?
|
|
45
|
+
unless criteria.is_a?(Hash) && (criteria.keys.map(&:to_s) - %w[yes no true false]).empty?
|
|
46
|
+
raise ArgumentError, 'Probability criteria must describe yes and no'
|
|
47
|
+
end
|
|
48
|
+
when :choice
|
|
49
|
+
raise ArgumentError, 'A choice needs a nonempty Hash of options' unless criteria.is_a?(Hash) && !criteria.empty?
|
|
50
|
+
when :score
|
|
51
|
+
unless criteria.is_a?(Array) && criteria.size >= 2 && criteria.none?(&:nil?)
|
|
52
|
+
raise ArgumentError, 'A score needs at least two non-nil levels'
|
|
53
|
+
end
|
|
54
|
+
end
|
|
55
|
+
end
|
|
56
|
+
end
|
|
57
|
+
|
|
58
|
+
class Probability
|
|
59
|
+
attr_reader :probability
|
|
60
|
+
|
|
61
|
+
def initialize(probability:)
|
|
62
|
+
@probability = probability
|
|
63
|
+
freeze
|
|
64
|
+
end
|
|
65
|
+
|
|
66
|
+
def type = :probability
|
|
67
|
+
def to_h = { type:, probability: }
|
|
68
|
+
end
|
|
69
|
+
|
|
70
|
+
class Choice
|
|
71
|
+
attr_reader :choice, :probabilities, :confidence
|
|
72
|
+
|
|
73
|
+
def initialize(choice:, probabilities:, confidence:)
|
|
74
|
+
@choice = choice
|
|
75
|
+
@probabilities = probabilities.dup.freeze
|
|
76
|
+
@confidence = confidence
|
|
77
|
+
freeze
|
|
78
|
+
end
|
|
79
|
+
|
|
80
|
+
def type = :choice
|
|
81
|
+
def to_h = { type:, choice:, probabilities:, confidence: }
|
|
82
|
+
end
|
|
83
|
+
|
|
84
|
+
class Score
|
|
85
|
+
attr_reader :score, :levels, :probabilities, :confidence
|
|
86
|
+
|
|
87
|
+
def initialize(score:, levels:, probabilities:, confidence:)
|
|
88
|
+
@score = score
|
|
89
|
+
@levels = levels
|
|
90
|
+
@probabilities = probabilities.dup.freeze
|
|
91
|
+
@confidence = confidence
|
|
92
|
+
freeze
|
|
93
|
+
end
|
|
94
|
+
|
|
95
|
+
def type = :score
|
|
96
|
+
def to_h = { type:, score:, levels:, probabilities:, confidence: }
|
|
97
|
+
end
|
|
98
|
+
|
|
99
|
+
class Judgment
|
|
100
|
+
include Enumerable
|
|
101
|
+
|
|
102
|
+
attr_reader :answers, :model, :tokens, :raw
|
|
103
|
+
|
|
104
|
+
def initialize(answers:, model:, tokens:, raw: nil)
|
|
105
|
+
@answers = answers.dup.freeze
|
|
106
|
+
@answer_keys = answers.keys.to_h { |key| [key.to_s, key] }.freeze
|
|
107
|
+
@model = model
|
|
108
|
+
@tokens = tokens
|
|
109
|
+
@raw = raw
|
|
110
|
+
end
|
|
111
|
+
|
|
112
|
+
def [](name)
|
|
113
|
+
answers[@answer_keys[name.to_s]]
|
|
114
|
+
end
|
|
115
|
+
|
|
116
|
+
def fetch(name)
|
|
117
|
+
answers.fetch(@answer_keys.fetch(name.to_s))
|
|
118
|
+
end
|
|
119
|
+
|
|
120
|
+
def each(&)
|
|
121
|
+
answers.each(&)
|
|
122
|
+
end
|
|
123
|
+
|
|
124
|
+
def to_h
|
|
125
|
+
{ model:, answers: answers.transform_values(&:to_h), tokens: tokens.to_h }
|
|
126
|
+
end
|
|
127
|
+
|
|
128
|
+
private
|
|
129
|
+
|
|
130
|
+
def method_missing(name, *args, &block)
|
|
131
|
+
return super unless args.empty? && !block && @answer_keys.key?(name.to_s)
|
|
132
|
+
|
|
133
|
+
self[name]
|
|
134
|
+
end
|
|
135
|
+
|
|
136
|
+
def respond_to_missing?(name, include_private = false)
|
|
137
|
+
@answer_keys.key?(name.to_s) || super
|
|
138
|
+
end
|
|
139
|
+
end
|
|
140
|
+
|
|
141
|
+
module_function
|
|
142
|
+
|
|
143
|
+
def aggregate_tokens(tokens)
|
|
144
|
+
tokens = tokens.compact
|
|
145
|
+
sum = ->(reader) { tokens.sum { |token| token.public_send(reader) || 0 } }
|
|
146
|
+
RubyLLM::Tokens.new(input: sum.(:input), output: sum.(:output))
|
|
147
|
+
end
|
|
148
|
+
end
|
|
149
|
+
end
|
|
150
|
+
end
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module RubyLLM
|
|
4
|
+
module LLMJudge
|
|
5
|
+
class Provider < RubyLLM::Provider
|
|
6
|
+
def api_base
|
|
7
|
+
'http://127.0.0.1'
|
|
8
|
+
end
|
|
9
|
+
|
|
10
|
+
def judge(input, questions:, model:, provider_options: {})
|
|
11
|
+
Engine.new(config:, model: model.id, provider_options:).judge(input, questions:, model:)
|
|
12
|
+
end
|
|
13
|
+
|
|
14
|
+
def self.assume_models_exist?
|
|
15
|
+
true
|
|
16
|
+
end
|
|
17
|
+
end
|
|
18
|
+
end
|
|
19
|
+
end
|
|
@@ -0,0 +1,42 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require 'ruby_llm'
|
|
4
|
+
require_relative 'llm_judge/version'
|
|
5
|
+
|
|
6
|
+
module RubyLLM
|
|
7
|
+
module LLMJudge
|
|
8
|
+
DEFAULT_MODEL = 'gpt-6-luna'
|
|
9
|
+
|
|
10
|
+
# RubyLLM 2.1 added the Judge API. Earlier versions run the engine directly
|
|
11
|
+
# and return Legacy stand-ins with the same readers.
|
|
12
|
+
NATIVE = defined?(RubyLLM::Judge) ? true : false
|
|
13
|
+
end
|
|
14
|
+
end
|
|
15
|
+
|
|
16
|
+
require_relative 'llm_judge/legacy' unless RubyLLM::LLMJudge::NATIVE
|
|
17
|
+
require_relative 'llm_judge/engine'
|
|
18
|
+
require_relative 'llm_judge/provider' if RubyLLM::LLMJudge::NATIVE
|
|
19
|
+
|
|
20
|
+
module RubyLLM
|
|
21
|
+
module LLMJudge
|
|
22
|
+
# Answer and judgment classes: RubyLLM's own on 2.1+, Legacy otherwise.
|
|
23
|
+
Types = NATIVE ? RubyLLM : Legacy
|
|
24
|
+
|
|
25
|
+
def self.judge(input, questions:, model: DEFAULT_MODEL, **options)
|
|
26
|
+
if NATIVE
|
|
27
|
+
return RubyLLM.judge(input, questions:, model:, provider: :llm_judge,
|
|
28
|
+
assume_model_exists: true, **options)
|
|
29
|
+
end
|
|
30
|
+
|
|
31
|
+
provider_options = options.delete(:provider_options) || {}
|
|
32
|
+
context = options.delete(:context)
|
|
33
|
+
raise ArgumentError, "Unsupported before RubyLLM 2.1: #{options.keys.join(', ')}" unless options.empty?
|
|
34
|
+
|
|
35
|
+
built = questions.to_h { |name, definition| [name.to_s, Legacy::Question.from_h(name, definition)] }
|
|
36
|
+
Engine.new(config: context&.config || RubyLLM.config, model:, provider_options:)
|
|
37
|
+
.judge(input, questions: built, model: Legacy::Model.new(id: model))
|
|
38
|
+
end
|
|
39
|
+
end
|
|
40
|
+
end
|
|
41
|
+
|
|
42
|
+
RubyLLM::Provider.register(:llm_judge, RubyLLM::LLMJudge::Provider) if RubyLLM::LLMJudge::NATIVE
|
metadata
ADDED
|
@@ -0,0 +1,72 @@
|
|
|
1
|
+
--- !ruby/object:Gem::Specification
|
|
2
|
+
name: ruby_llm-llm_judge
|
|
3
|
+
version: !ruby/object:Gem::Version
|
|
4
|
+
version: 0.1.0
|
|
5
|
+
platform: ruby
|
|
6
|
+
authors:
|
|
7
|
+
- JP Camara
|
|
8
|
+
bindir: bin
|
|
9
|
+
cert_chain: []
|
|
10
|
+
date: 1980-01-02 00:00:00.000000000 Z
|
|
11
|
+
dependencies:
|
|
12
|
+
- !ruby/object:Gem::Dependency
|
|
13
|
+
name: ruby_llm
|
|
14
|
+
requirement: !ruby/object:Gem::Requirement
|
|
15
|
+
requirements:
|
|
16
|
+
- - ">="
|
|
17
|
+
- !ruby/object:Gem::Version
|
|
18
|
+
version: 1.13.0
|
|
19
|
+
- - "<"
|
|
20
|
+
- !ruby/object:Gem::Version
|
|
21
|
+
version: '3.0'
|
|
22
|
+
type: :runtime
|
|
23
|
+
prerelease: false
|
|
24
|
+
version_requirements: !ruby/object:Gem::Requirement
|
|
25
|
+
requirements:
|
|
26
|
+
- - ">="
|
|
27
|
+
- !ruby/object:Gem::Version
|
|
28
|
+
version: 1.13.0
|
|
29
|
+
- - "<"
|
|
30
|
+
- !ruby/object:Gem::Version
|
|
31
|
+
version: '3.0'
|
|
32
|
+
description: Answers probability, choice, and score questions with any RubyLLM chat
|
|
33
|
+
model, using parallel per-answer ratings or one-call typed distributions. Plugs
|
|
34
|
+
into RubyLLM 2.1's Judge API and also runs on RubyLLM 1.13 through 1.16.
|
|
35
|
+
executables: []
|
|
36
|
+
extensions: []
|
|
37
|
+
extra_rdoc_files: []
|
|
38
|
+
files:
|
|
39
|
+
- CHANGELOG.md
|
|
40
|
+
- LICENSE
|
|
41
|
+
- README.md
|
|
42
|
+
- lib/ruby_llm/llm_judge.rb
|
|
43
|
+
- lib/ruby_llm/llm_judge/engine.rb
|
|
44
|
+
- lib/ruby_llm/llm_judge/legacy.rb
|
|
45
|
+
- lib/ruby_llm/llm_judge/provider.rb
|
|
46
|
+
- lib/ruby_llm/llm_judge/version.rb
|
|
47
|
+
homepage: https://github.com/jpcamara/ruby_llm-llm_judge
|
|
48
|
+
licenses:
|
|
49
|
+
- MIT
|
|
50
|
+
metadata:
|
|
51
|
+
source_code_uri: https://github.com/jpcamara/ruby_llm-llm_judge
|
|
52
|
+
changelog_uri: https://github.com/jpcamara/ruby_llm-llm_judge/blob/main/CHANGELOG.md
|
|
53
|
+
bug_tracker_uri: https://github.com/jpcamara/ruby_llm-llm_judge/issues
|
|
54
|
+
rubygems_mfa_required: 'true'
|
|
55
|
+
rdoc_options: []
|
|
56
|
+
require_paths:
|
|
57
|
+
- lib
|
|
58
|
+
required_ruby_version: !ruby/object:Gem::Requirement
|
|
59
|
+
requirements:
|
|
60
|
+
- - ">="
|
|
61
|
+
- !ruby/object:Gem::Version
|
|
62
|
+
version: '3.2'
|
|
63
|
+
required_rubygems_version: !ruby/object:Gem::Requirement
|
|
64
|
+
requirements:
|
|
65
|
+
- - ">="
|
|
66
|
+
- !ruby/object:Gem::Version
|
|
67
|
+
version: '0'
|
|
68
|
+
requirements: []
|
|
69
|
+
rubygems_version: 3.6.9
|
|
70
|
+
specification_version: 4
|
|
71
|
+
summary: Use RubyLLM chat models to answer RubyLLM Judge questions
|
|
72
|
+
test_files: []
|