judge_rails 0.0.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +7 -0
- data/ADVANCED.md +136 -0
- data/BENCHMARK.md +885 -0
- data/CHANGELOG.md +163 -0
- data/LICENSE.txt +21 -0
- data/README.md +588 -0
- data/lib/generators/judge/attribute_generator.rb +149 -0
- data/lib/generators/judge/install_generator.rb +36 -0
- data/lib/generators/judge/templates/initializer.rb.tt +18 -0
- data/lib/generators/judge/templates/migration.rb.tt +9 -0
- data/lib/judge/adapter.rb +69 -0
- data/lib/judge/client.rb +253 -0
- data/lib/judge/configuration.rb +103 -0
- data/lib/judge/decision.rb +24 -0
- data/lib/judge/errors.rb +36 -0
- data/lib/judge/facade.rb +75 -0
- data/lib/judge/pool.rb +97 -0
- data/lib/judge/question/choice.rb +36 -0
- data/lib/judge/question/noul.rb +22 -0
- data/lib/judge/question/score.rb +36 -0
- data/lib/judge/question.rb +93 -0
- data/lib/judge/rails/attributes.rb +161 -0
- data/lib/judge/rails/definition.rb +156 -0
- data/lib/judge/rails/jobs.rb +156 -0
- data/lib/judge/rails/locale/en.yml +6 -0
- data/lib/judge/rails/migration.rb +115 -0
- data/lib/judge/rails/refresh.rb +37 -0
- data/lib/judge/rails/refresh_job.rb +15 -0
- data/lib/judge/rails/registry.rb +57 -0
- data/lib/judge/rails/relation.rb +132 -0
- data/lib/judge/rails/scopes.rb +124 -0
- data/lib/judge/rails/storage.rb +104 -0
- data/lib/judge/rails/validator.rb +246 -0
- data/lib/judge/rails.rb +45 -0
- data/lib/judge/result.rb +173 -0
- data/lib/judge/result_set.rb +75 -0
- data/lib/judge/version.rb +5 -0
- data/lib/judge.rb +41 -0
- data/lib/judge_rails.rb +3 -0
- metadata +126 -0
data/README.md
ADDED
|
@@ -0,0 +1,588 @@
|
|
|
1
|
+
# judge_rails
|
|
2
|
+
|
|
3
|
+
**Semantic judgments as ordinary ActiveRecord attributes.**
|
|
4
|
+
|
|
5
|
+
A judgment model answers typed questions about text and returns a calibrated probability. This gem
|
|
6
|
+
turns that into a column on your model: indexable, sortable, paginable, and kept up to date for you.
|
|
7
|
+
It is also a plain Ruby client, with no Rails and nothing outside the standard library.
|
|
8
|
+
|
|
9
|
+
It talks to [TypeSafe Jev](https://docs.typesafe.ai) out of the box, and to anything else through
|
|
10
|
+
one small seam if you ever need it.
|
|
11
|
+
|
|
12
|
+
```ruby
|
|
13
|
+
class Ticket < ApplicationRecord
|
|
14
|
+
judge_source { [subject, body] }
|
|
15
|
+
|
|
16
|
+
judge_attribute :urgency, Judge.noul("Does this need a human within the hour?")
|
|
17
|
+
judge_attribute :intent, Judge.choice("What is this about?", %w[billing technical sales])
|
|
18
|
+
judge_attribute :frustration, Judge.score("How frustrated is the customer?", ["Calm", "Frustrated", "Very angry"])
|
|
19
|
+
end
|
|
20
|
+
|
|
21
|
+
Ticket.urgency_above(0.8).intent_is("billing").order(frustration: :desc).limit(20)
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
Three judgments cost one API call per record. That query costs none.
|
|
25
|
+
|
|
26
|
+
## Install
|
|
27
|
+
|
|
28
|
+
```ruby
|
|
29
|
+
gem "judge_rails"
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
```sh
|
|
33
|
+
bin/rails generate judge:install
|
|
34
|
+
bin/rails generate judge:attribute Ticket urgency:noul intent:choice frustration:score
|
|
35
|
+
bin/rails db:migrate
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
Set `JEV_API_KEY` in the environment, or `judge.api_key` in Rails credentials. `TYPESAFE_API_KEY`
|
|
39
|
+
works too.
|
|
40
|
+
|
|
41
|
+
**That is the whole setup.** There is nothing to choose and nothing to wire: the default adapter is
|
|
42
|
+
TypeSafe Jev, and it is used unless you say otherwise. [Providers](#providers) is there if you ever
|
|
43
|
+
need something else, and you can ignore it until then.
|
|
44
|
+
|
|
45
|
+
### Two things worth knowing first
|
|
46
|
+
|
|
47
|
+
**A key is not automatic.** TypeSafe Jev is reached through an early-access waitlist, so
|
|
48
|
+
`bundle install` will not by itself get you a working gem. You can still evaluate the idea today:
|
|
49
|
+
`judge_rails_demo` boots with **250 support tickets already judged**, seeded from a committed file,
|
|
50
|
+
and needs no key at all.
|
|
51
|
+
|
|
52
|
+
```sh
|
|
53
|
+
bin/rails db:prepare db:seed # 250 judged tickets, no API call
|
|
54
|
+
bin/rails server
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
**Pin `json` below 3 if you are on activesupport 8.1.** It calls `::JSON.parse(json, options)` with a
|
|
58
|
+
positional hash, which json 3.x rejects. It only surfaces on a jsonb column with a non-nil default,
|
|
59
|
+
which is exactly what the sidecar is, so it looks like a bug in this gem and is not:
|
|
60
|
+
|
|
61
|
+
```ruby
|
|
62
|
+
gem "json", "< 3"
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
## Questions
|
|
66
|
+
|
|
67
|
+
Three types. Each one is a frozen value object you can hold, pass around and compare.
|
|
68
|
+
|
|
69
|
+
```ruby
|
|
70
|
+
Judge.noul("Does this convey urgency?")
|
|
71
|
+
Judge.noul("Does this convey urgency?", { true: "Time-sensitive", false: "Can wait" })
|
|
72
|
+
|
|
73
|
+
Judge.choice("Which team?", %w[billing technical sales])
|
|
74
|
+
Judge.choice("Which team?", { billing: "Refunds, invoices", technical: "Bugs, outages" })
|
|
75
|
+
|
|
76
|
+
Judge.score("How frustrated?", ["Calm", "Mildly annoyed", "Frustrated", "Very angry"])
|
|
77
|
+
Judge.score("How urgent?", 1..5)
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
Criteria are optional on a noul and sharpen the judgment. An entry can also be an object, the form
|
|
81
|
+
TypeSafe documents for drawing a boundary between options:
|
|
82
|
+
|
|
83
|
+
```ruby
|
|
84
|
+
Judge.choice("Which team?", {
|
|
85
|
+
billing: { what: "Refunds, invoices", not_for: "Bugs", examples: ["I was charged twice"] },
|
|
86
|
+
technical: { what: "Bugs, outages", not_for: "Charges or refunds", examples: ["The API returns 500"] }
|
|
87
|
+
})
|
|
88
|
+
Judge.noul("Does this need a human within the hour?", {
|
|
89
|
+
true: { what: "Money is stuck or a service is down", examples: ["Checkout has failed for an hour"] },
|
|
90
|
+
false: { what: "Anything that can wait until tomorrow", examples: ["How do I export invoices?"] }
|
|
91
|
+
})
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
Worth doing wherever two options can be confused. On the demo's 250 tickets this raised agreement with
|
|
95
|
+
reference labels by about four points on both questions, for about 290 more input tokens per ticket,
|
|
96
|
+
which is a thousandth of a cent. The demo now declares its questions this way. The labels come from
|
|
97
|
+
two model annotators, not from people. `BENCHMARK.md` (in French), arm 8, has the numbers and their
|
|
98
|
+
limits.
|
|
99
|
+
|
|
100
|
+
A question fingerprints itself, which is what makes invalidation automatic later.
|
|
101
|
+
|
|
102
|
+
```ruby
|
|
103
|
+
question = Judge.choice("Which team?", %w[billing technical sales spam])
|
|
104
|
+
question.digest # => "1c3be3cf24579b81" over type, wording and criteria
|
|
105
|
+
question.options # => ["billing", "technical", "sales", "spam"]
|
|
106
|
+
|
|
107
|
+
Judge.score("How frustrated?", ["Calm", "Mildly annoyed", "Frustrated", "Very angry"]).max_level # => 3
|
|
108
|
+
```
|
|
109
|
+
|
|
110
|
+
## Asking
|
|
111
|
+
|
|
112
|
+
One question in, one `Result` out. Many questions in, one `ResultSet` out, and one HTTP request.
|
|
113
|
+
|
|
114
|
+
```ruby
|
|
115
|
+
Judge.ask("does this sound angry?", text: ticket.body)
|
|
116
|
+
# => #<Judge::Result :answer noul value=0.96 p=0.96>
|
|
117
|
+
|
|
118
|
+
results = Judge.ask({
|
|
119
|
+
urgency: Judge.noul("Does this need a human within the hour?"),
|
|
120
|
+
intent: Judge.choice("What is this about?", %w[billing technical sales spam]),
|
|
121
|
+
frustration: Judge.score("How frustrated?", ["Calm", "Mildly annoyed", "Frustrated", "Very angry"])
|
|
122
|
+
}, text: ticket.body)
|
|
123
|
+
# => #<Judge::ResultSet [:urgency, :intent, :frustration] model="jev-1.13.0" latency=0.69>
|
|
124
|
+
```
|
|
125
|
+
|
|
126
|
+
A bare String is treated as a noul. An Array works too, named by each question or positionally.
|
|
127
|
+
|
|
128
|
+
## Reading a result
|
|
129
|
+
|
|
130
|
+
```ruby
|
|
131
|
+
results[:urgency].value # => 0.96 a noul is its own probability
|
|
132
|
+
results[:urgency].true?(0.8) # => true
|
|
133
|
+
|
|
134
|
+
results[:intent].value # => "technical"
|
|
135
|
+
results[:intent].confidence # => 1.0
|
|
136
|
+
results[:intent].probabilities # => {"technical" => 1.0, "billing" => 0.0, ...}
|
|
137
|
+
results[:intent].probability_of(:billing)
|
|
138
|
+
|
|
139
|
+
results[:frustration].value # => 2.01 continuous
|
|
140
|
+
results[:frustration].level # => 2 rounded
|
|
141
|
+
results[:frustration].label # => "Frustrated"
|
|
142
|
+
|
|
143
|
+
results.model # => "jev-1.13.0"
|
|
144
|
+
results.usage.input_tokens # => 430
|
|
145
|
+
results.usage.total # => 503
|
|
146
|
+
results.latency # => 0.69
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
The calibrated probability is what you are paying for, so act on the band rather than a boolean.
|
|
150
|
+
|
|
151
|
+
```ruby
|
|
152
|
+
case results[:urgency].decide(above: 0.9, below: 0.1)
|
|
153
|
+
when :yes then escalate!
|
|
154
|
+
when :no then queue_normally!
|
|
155
|
+
when :unsure then assign_to_human!
|
|
156
|
+
end
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
## Attributes
|
|
160
|
+
|
|
161
|
+
`judge_attribute` stores a judgment as a real column, plus a `<name>_judge` jsonb sidecar holding its
|
|
162
|
+
provenance. `judge_source` sets the text once, so every attribute sharing it travels in one call per record.
|
|
163
|
+
|
|
164
|
+
```ruby
|
|
165
|
+
class Ticket < ApplicationRecord
|
|
166
|
+
judge_source { [subject, body] }
|
|
167
|
+
|
|
168
|
+
judge_attribute :urgency, Judge.noul("Does this need a human within the hour?")
|
|
169
|
+
judge_attribute :intent, Judge.choice("What is this about?", %w[billing technical sales spam])
|
|
170
|
+
judge_attribute :frustration, Judge.score("How frustrated?", ["Calm", "Mildly annoyed", "Frustrated", "Very angry"])
|
|
171
|
+
|
|
172
|
+
judge_attribute :spam, Judge.noul("Is this spam?"), source: :body
|
|
173
|
+
end
|
|
174
|
+
```
|
|
175
|
+
|
|
176
|
+
Every attribute gets readers for its value and its provenance. A noul also gets a threshold predicate.
|
|
177
|
+
|
|
178
|
+
```ruby
|
|
179
|
+
ticket.urgency # => 0.96
|
|
180
|
+
ticket.urgency?(0.8) # => true
|
|
181
|
+
ticket.urgency_probability # => 0.96
|
|
182
|
+
ticket.intent_confidence # => 1.0
|
|
183
|
+
ticket.urgency_computed_at # => 2026-09-20 11:42:10 UTC
|
|
184
|
+
ticket.urgency_stale? # => false
|
|
185
|
+
ticket.urgency_judge_meta
|
|
186
|
+
# => {"digest" => ..., "state_digest" => ..., "computed_at" => ..., "probability" => 0.96,
|
|
187
|
+
# "probabilities" => {...}, "model" => "jev-1.13.0", "latency" => 0.69}
|
|
188
|
+
|
|
189
|
+
ticket.judge_decide(:urgency, above: 0.9, below: 0.1) # => :yes
|
|
190
|
+
```
|
|
191
|
+
|
|
192
|
+
## Invalidation
|
|
193
|
+
|
|
194
|
+
Two things make a judgment stale: the source text changing, and the question changing. Both are detected
|
|
195
|
+
by comparing digests, which costs no API call.
|
|
196
|
+
|
|
197
|
+
```ruby
|
|
198
|
+
ticket.judge_stale? # => false
|
|
199
|
+
ticket.body = "actually, all sorted, thanks"
|
|
200
|
+
ticket.judge_stale? # => true
|
|
201
|
+
ticket.judge_pending # => [:urgency, :intent, :frustration, :spam]
|
|
202
|
+
|
|
203
|
+
ticket.judge_refresh # recompute in memory, returns the names it touched
|
|
204
|
+
ticket.judge_refresh! # or recompute and store the judgment columns
|
|
205
|
+
ticket.judge_refresh(:urgency, force: true) # or one attribute, even if it is fresh
|
|
206
|
+
```
|
|
207
|
+
|
|
208
|
+
`judge_refresh!` writes only the judgment columns and `updated_at`, with `update_columns`: no
|
|
209
|
+
validations and no save callbacks, so a refresh cannot trigger another one, and fragment caches keyed
|
|
210
|
+
on the record's version see the new judgment. Other unsaved edits on the record stay unsaved. If the
|
|
211
|
+
row was deleted meanwhile, it raises `ActiveRecord::RecordNotFound`.
|
|
212
|
+
|
|
213
|
+
On a record that was never saved, `judge_refresh!` creates it. If one of its requests fails, the row is
|
|
214
|
+
still created from the answers already paid for, without asking again, and then the error is raised.
|
|
215
|
+
Reload before retrying, or the retry creates a second row. To react to a
|
|
216
|
+
new judgment, for a broadcast say, use `after_judge_refresh`:
|
|
217
|
+
|
|
218
|
+
```ruby
|
|
219
|
+
after_judge_refresh { broadcast_replace_later_to :tickets }
|
|
220
|
+
```
|
|
221
|
+
|
|
222
|
+
Editing the wording of a question moves its digest, so every stored judgment for it goes stale on its own.
|
|
223
|
+
There is no version number to remember to bump.
|
|
224
|
+
|
|
225
|
+
Pinning a model does the same. Changing the pin makes every judgment made under the old one stale, and
|
|
226
|
+
two attributes on the same text but different models travel in two calls instead of one. An attribute
|
|
227
|
+
with no pin follows `config.model`, so changing that makes its judgments stale too.
|
|
228
|
+
|
|
229
|
+
```ruby
|
|
230
|
+
judge_attribute :urgency, Judge.noul("..."), model: "jev-1.13.0" # the rest follow config.model
|
|
231
|
+
```
|
|
232
|
+
|
|
233
|
+
Blank source text has nothing to judge, and no call is made for it. An automatic attribute clears its
|
|
234
|
+
value and sidecar on save. A `callbacks: false` one shows as stale until `judge_refresh` clears it.
|
|
235
|
+
|
|
236
|
+
A source should depend only on content. One that reads `updated_at` or a column a callback rewrites
|
|
237
|
+
makes every ordinary save look like new text, and shows as stale right after its own refresh. The
|
|
238
|
+
refresh runs no callbacks, so it cannot feed a loop, but each of your saves will enqueue one more
|
|
239
|
+
judgment.
|
|
240
|
+
|
|
241
|
+
## When it runs
|
|
242
|
+
|
|
243
|
+
```ruby
|
|
244
|
+
judge_attribute :urgency, Judge.noul("...") # async after_commit (default)
|
|
245
|
+
judge_attribute :urgency, Judge.noul("..."), sync: true # inline, during the save
|
|
246
|
+
judge_attribute :urgency, Judge.noul("..."), callbacks: false # manual only
|
|
247
|
+
```
|
|
248
|
+
|
|
249
|
+
Async needs ActiveJob, or an enqueuer of your own (see [ADVANCED.md](ADVANCED.md)). Without either, the
|
|
250
|
+
first save raises `Judge::ConfigurationError` instead of silently skipping the judgment.
|
|
251
|
+
|
|
252
|
+
Async is the default on purpose. An HTTP call inside a save holds a pooled database connection for the
|
|
253
|
+
whole request, so a slow vendor exhausts the pool and takes down more than the feature. A save that
|
|
254
|
+
changes nothing relevant enqueues nothing, and ActiveJob is optional: the enqueuer is injectable.
|
|
255
|
+
|
|
256
|
+
```ruby
|
|
257
|
+
ticket.judge_refresh_later(:urgency)
|
|
258
|
+
```
|
|
259
|
+
|
|
260
|
+
`sync: true` buys one thing, the right to block a save when the judgment fails. That, `if_condition`
|
|
261
|
+
and bulk backfill are in [ADVANCED.md](ADVANCED.md).
|
|
262
|
+
|
|
263
|
+
## Querying
|
|
264
|
+
|
|
265
|
+
Judgments are real columns, so this is plain indexed SQL and no API call.
|
|
266
|
+
|
|
267
|
+
```ruby
|
|
268
|
+
Ticket.urgency_above(0.8) # >=
|
|
269
|
+
Ticket.urgency_below(0.2) # <=
|
|
270
|
+
Ticket.urgency_between(0.2, 0.8) # the strict middle
|
|
271
|
+
Ticket.urgency_unknown # never judged
|
|
272
|
+
|
|
273
|
+
Ticket.intent_is("billing", "technical")
|
|
274
|
+
Ticket.intent_not("spam") # raises on an option you never declared
|
|
275
|
+
|
|
276
|
+
Ticket.frustration_at_least("Frustrated") # by name
|
|
277
|
+
Ticket.frustration_at_most(1) # or by index
|
|
278
|
+
Ticket.frustration_level(2) # on numeric levels like 1..5, an Integer is the label
|
|
279
|
+
|
|
280
|
+
Ticket.order_by_urgency(:desc) # NULLs where your database puts them: first on PostgreSQL
|
|
281
|
+
Ticket.judge_computed
|
|
282
|
+
Ticket.judge_uncomputed
|
|
283
|
+
```
|
|
284
|
+
|
|
285
|
+
The bands partition the table exactly the way `judge_decide` does, so SQL and Ruby never disagree at the
|
|
286
|
+
threshold.
|
|
287
|
+
|
|
288
|
+
```ruby
|
|
289
|
+
Ticket.urgency_above(0.8).count +
|
|
290
|
+
Ticket.urgency_between(0.2, 0.8).count +
|
|
291
|
+
Ticket.urgency_below(0.2).count +
|
|
292
|
+
Ticket.urgency_unknown.count == Ticket.count # => true
|
|
293
|
+
```
|
|
294
|
+
|
|
295
|
+
For a question you never declared, there is an ad-hoc path. It loads records, judges them and returns an
|
|
296
|
+
Array rather than a Relation, because the work has already happened.
|
|
297
|
+
|
|
298
|
+
```ruby
|
|
299
|
+
Ticket.where(channel: "chat").judge_filter("mentions a chargeback", limit: 500) # => [Ticket, ...]
|
|
300
|
+
Ticket.judge_map("how angry is this?", limit: 200) # => {ticket => Result}
|
|
301
|
+
Ticket.judge_sort("most likely to churn", limit: 200, dir: :desc) # => [Ticket, ...]
|
|
302
|
+
|
|
303
|
+
Ticket.judge_filter("mentions a chargeback")
|
|
304
|
+
# ArgumentError: judge_filter requires limit:. It makes one API call per row.
|
|
305
|
+
```
|
|
306
|
+
|
|
307
|
+
A choice or a score question needs a target, because "the answer is confident" says nothing about which
|
|
308
|
+
answer won. A missing target raises before any call is made.
|
|
309
|
+
|
|
310
|
+
```ruby
|
|
311
|
+
team = Judge.choice("Which team?", %w[billing technical sales])
|
|
312
|
+
Ticket.judge_filter(team, option: "billing", limit: 100) # P(billing) >= threshold
|
|
313
|
+
Ticket.judge_sort(team, option: "billing", limit: 100) # ranked by P(billing)
|
|
314
|
+
|
|
315
|
+
mood = Judge.score("How frustrated?", ["Calm", "Frustrated", "Very angry"])
|
|
316
|
+
Ticket.judge_filter(mood, at_least: "Frustrated", limit: 100)
|
|
317
|
+
```
|
|
318
|
+
|
|
319
|
+
Rows sharing the same text are asked once.
|
|
320
|
+
|
|
321
|
+
One row per request is not an accident, it is the only shape that keeps the judgment intact: Jev
|
|
322
|
+
scores each question against the whole state, so putting several records in one request makes every
|
|
323
|
+
answer a judgment about a mostly irrelevant document. Measured on 250 tickets, packing subjects cost
|
|
324
|
+
nine to fourteen points of decision agreement. `BENCHMARK.md` has the protocol and the numbers.
|
|
325
|
+
|
|
326
|
+
What is free is running those requests at the same time. The request is unchanged byte for byte, so
|
|
327
|
+
nothing is traded for the speed.
|
|
328
|
+
|
|
329
|
+
```ruby
|
|
330
|
+
Ticket.judge_filter("mentions a chargeback", limit: 500, concurrency: 16)
|
|
331
|
+
Judge.configure { |c| c.concurrency = 16 } # or set the default once
|
|
332
|
+
```
|
|
333
|
+
|
|
334
|
+
Measured against the live API on 100 records: 6.9x faster at the default of 8 threads, 18.4x at 32,
|
|
335
|
+
with input tokens identical to the unit at every level.
|
|
336
|
+
|
|
337
|
+
TypeSafe documents a limit of 1,200 requests per minute and says it can change without notice. The
|
|
338
|
+
default of 8 threads runs at about 29 requests a second, above that limit. On 2026-09-23 one key
|
|
339
|
+
sustained 1,758 requests a minute at 8 threads and 6,696 at 32 without a single 429, but nothing
|
|
340
|
+
promises that tomorrow. A 429 is retried twice, honouring `Retry-After` up to `max_retry_wait`.
|
|
341
|
+
|
|
342
|
+
## Validations
|
|
343
|
+
|
|
344
|
+
An ordinary ActiveModel validation that happens to ask a model.
|
|
345
|
+
|
|
346
|
+
```ruby
|
|
347
|
+
class Ticket < ApplicationRecord
|
|
348
|
+
validates :body, judge: { refute: "contains a phone number or email address" }
|
|
349
|
+
validates :body, judge: { assert: "is written in English", threshold: 0.8, message: "must be in English" }
|
|
350
|
+
validates :body, judge: { refute: "is spam", on_error: :fail }, if: -> { channel == "web" }
|
|
351
|
+
end
|
|
352
|
+
|
|
353
|
+
ticket.valid?
|
|
354
|
+
ticket.errors.full_messages # => ["Body matched \"contains a phone number or email address\""]
|
|
355
|
+
ticket.judge_validation_results # the judgment behind each verdict, never persisted
|
|
356
|
+
ticket.errors.of_kind?(:body, :judge_refuted) # also :judge_unmatched, and :judge_unavailable on :fail
|
|
357
|
+
```
|
|
358
|
+
|
|
359
|
+
The messages are I18n keys under `errors.messages` (`judge_refuted`, `judge_unmatched`,
|
|
360
|
+
`judge_unavailable`), with the instruction as `%{instruction}`. A locale without them falls back to the
|
|
361
|
+
English text rather than "Translation missing". `strict:`, `on:`, `except_on:`, `if:` and
|
|
362
|
+
`unless:` behave as they do on any Rails validation.
|
|
363
|
+
|
|
364
|
+
Several judge validations on the same attribute travel in one call. A blank attribute costs nothing.
|
|
365
|
+
A record remembers the judgment of its current text, so saving it again without changing the text costs
|
|
366
|
+
nothing. The memory is per Ruby object: a record freshly loaded from the database pays one call per judge
|
|
367
|
+
validation on its first save. `if: :will_save_change_to_body?` avoids that when the text is all you check.
|
|
368
|
+
`reload` forgets it. A skipped `on_error: :pass` check is logged at warn level.
|
|
369
|
+
|
|
370
|
+
It works on a plain `ActiveModel::Model` form object too, one call per validation.
|
|
371
|
+
|
|
372
|
+
`on_error` defaults to `:pass`, so a TypeSafe outage cannot stop your users saving. **That default is
|
|
373
|
+
wrong for moderation**: a check that blocks spam or personal data and then fails open lets through
|
|
374
|
+
exactly what it exists to catch. Use `on_error: :fail` there.
|
|
375
|
+
|
|
376
|
+
## Migrations and generators
|
|
377
|
+
|
|
378
|
+
```ruby
|
|
379
|
+
class AddJudgeToTickets < ActiveRecord::Migration[8.0]
|
|
380
|
+
def change
|
|
381
|
+
judge_attribute :tickets, :urgency, :noul # float + urgency_judge + an index on urgency
|
|
382
|
+
judge_attribute :tickets, :intent, :choice # string + intent_judge
|
|
383
|
+
judge_attribute :tickets, :frustration, :score # float + frustration_judge
|
|
384
|
+
end
|
|
385
|
+
end
|
|
386
|
+
|
|
387
|
+
create_table :tickets do |t|
|
|
388
|
+
t.judge_attribute :urgency, :noul
|
|
389
|
+
end
|
|
390
|
+
|
|
391
|
+
change_table :tickets do |t|
|
|
392
|
+
t.judge_attribute :intent, :choice
|
|
393
|
+
end
|
|
394
|
+
```
|
|
395
|
+
|
|
396
|
+
Built from `add_column` and `add_index`, so it is reversible inside `change`. The sidecar is `jsonb` with
|
|
397
|
+
no index on PostgreSQL and `json` elsewhere, chosen from the connection actually running the migration.
|
|
398
|
+
Nothing in the gem queries inside the sidecar, so an index there would only slow every write. The value
|
|
399
|
+
column always allows NULL: it is empty until judged, so `null: false` raises.
|
|
400
|
+
CI runs the suite on SQLite and MySQL, and the demo runs on PostgreSQL. MySQL rejects a default on a
|
|
401
|
+
JSON column, so there the sidecar is nullable with no default, and a nil sidecar reads as `{}`.
|
|
402
|
+
|
|
403
|
+
```sh
|
|
404
|
+
bin/rails generate judge:install
|
|
405
|
+
bin/rails generate judge:attribute Ticket urgency:noul intent:choice frustration:score
|
|
406
|
+
bin/rails generate judge:attribute Ticket spam:noul --database=secondary # multi-database apps
|
|
407
|
+
```
|
|
408
|
+
|
|
409
|
+
The attribute generator writes live declarations with TODO questions, so replace the wording before the
|
|
410
|
+
first save: each save is judged, and billed, with whatever the question says. When the model has no
|
|
411
|
+
`judge_source`, it writes one from the table's `text` columns only, never its string columns, so a
|
|
412
|
+
`User` table does not send `encrypted_password` anywhere. With no `text` column it writes an empty one
|
|
413
|
+
that judges nothing until you fill it in. Later runs add their declarations below it, and `bin/rails destroy judge:attribute`
|
|
414
|
+
removes them.
|
|
415
|
+
|
|
416
|
+
## What it costs
|
|
417
|
+
|
|
418
|
+
A judged record is a billed network call. The gem's job is to make that number predictable.
|
|
419
|
+
|
|
420
|
+
| | |
|
|
421
|
+
|---|---|
|
|
422
|
+
| Per record, however many questions | **one** call. `judge_source` groups every attribute sharing a text |
|
|
423
|
+
| A save that changes nothing relevant | **zero** calls from `judge_attribute`: digests are compared first. A judge validation pays once per freshly loaded record |
|
|
424
|
+
| Any query over judged columns | **zero** calls. They are ordinary indexed columns |
|
|
425
|
+
| `judge_filter` / `judge_map` / `judge_sort` | **one call per row**, which is why `limit:` is mandatory |
|
|
426
|
+
| A request, before it carries anything | ~326 input tokens of fixed overhead |
|
|
427
|
+
| Each additional question in a request | ~24 input tokens |
|
|
428
|
+
| The demo's three questions, structured criteria, on a ~200-character ticket | ~825 input tokens, about 0.00003 USD |
|
|
429
|
+
|
|
430
|
+
Measured on 2026-09-21 and 2026-09-23 against the live model, at TypeSafe's published 0.042 USD per
|
|
431
|
+
million input tokens, output free; `BENCHMARK.md` carries the method. Those last two
|
|
432
|
+
lines are the whole argument for `judge_source`: three separate calls pay the overhead three times and
|
|
433
|
+
answer exactly the same thing.
|
|
434
|
+
|
|
435
|
+
Timeouts and retries are yours to set, and the defaults are deliberately short:
|
|
436
|
+
|
|
437
|
+
```ruby
|
|
438
|
+
Judge.configure do |config|
|
|
439
|
+
config.timeout = 10.0 # seconds, per request
|
|
440
|
+
config.open_timeout = 5.0
|
|
441
|
+
config.max_retries = 2 # 429, 5xx and transport errors, jittered
|
|
442
|
+
config.max_retry_wait = 10.0 # a longer Retry-After raises RateLimitError instead of sleeping
|
|
443
|
+
end
|
|
444
|
+
```
|
|
445
|
+
|
|
446
|
+
An async refresh that fails on a 429, a 5xx or a transport error raises out of `RefreshJob`, so
|
|
447
|
+
ActiveJob retries it with backoff, five attempts. Any other failure is logged and dropped.
|
|
448
|
+
|
|
449
|
+
Every request emits `request.judge` through `ActiveSupport::Notifications` when it is loaded, so an APM
|
|
450
|
+
sees them without any wiring:
|
|
451
|
+
|
|
452
|
+
```ruby
|
|
453
|
+
ActiveSupport::Notifications.subscribe("request.judge") do |*args|
|
|
454
|
+
event = ActiveSupport::Notifications::Event.new(*args)
|
|
455
|
+
event.payload # => {model:, questions:, request_bytes:, latency:, input_tokens:, output_tokens:}
|
|
456
|
+
end
|
|
457
|
+
```
|
|
458
|
+
|
|
459
|
+
`model`, `questions` and `request_bytes` are always there. `latency` and the two token counts are
|
|
460
|
+
added once the response parses, so a **failed** request carries only the first three, plus the
|
|
461
|
+
`:exception` pair ActiveSupport adds itself. Read them with `dig`, not `fetch`.
|
|
462
|
+
|
|
463
|
+
One event spans the whole call, retries included, not one per HTTP attempt.
|
|
464
|
+
|
|
465
|
+
The event carries sizes and counts. It never carries the state or the key.
|
|
466
|
+
|
|
467
|
+
## Client and errors
|
|
468
|
+
|
|
469
|
+
```ruby
|
|
470
|
+
Judge.configure do |config|
|
|
471
|
+
config.api_key = "..." # read from JEV_API_KEY or TYPESAFE_API_KEY by default
|
|
472
|
+
config.model = "jev-latest"
|
|
473
|
+
config.timeout = 10.0
|
|
474
|
+
config.open_timeout = 5.0
|
|
475
|
+
config.max_retries = 2
|
|
476
|
+
config.max_retry_wait = 10.0
|
|
477
|
+
config.logger = Rails.logger
|
|
478
|
+
end
|
|
479
|
+
```
|
|
480
|
+
|
|
481
|
+
`Net::HTTP`, one persistent connection per fiber keyed on URL and timeouts, never reused across a fork.
|
|
482
|
+
Jittered exponential backoff on 429, 5xx and connection failures, `Retry-After` honoured on 429 and 503 up
|
|
483
|
+
to `max_retry_wait`, no retry on any other 4xx. A read timeout is not re-sent, because the server may
|
|
484
|
+
already have judged (and billed) the request. Nothing about the key or the payload is ever logged, and
|
|
485
|
+
`inspect` on a configuration or a client hides the key.
|
|
486
|
+
|
|
487
|
+
```ruby
|
|
488
|
+
Judge::Error
|
|
489
|
+
├── Judge::ConfigurationError # no usable key, or async attributes without a way to enqueue
|
|
490
|
+
├── Judge::TransportError # timeout, reset, DNS, TLS
|
|
491
|
+
├── Judge::InvalidResponseError # body or an answer was not what the API promises
|
|
492
|
+
└── Judge::APIError # carries #status and #body
|
|
493
|
+
├── Judge::AuthenticationError # 401, 403
|
|
494
|
+
├── Judge::InvalidRequestError # 400, 404, 422
|
|
495
|
+
├── Judge::PayloadTooLargeError # 413
|
|
496
|
+
├── Judge::RateLimitError # 429, carries #retry_after
|
|
497
|
+
└── Judge::ServerError # 5xx
|
|
498
|
+
```
|
|
499
|
+
|
|
500
|
+
A per-request client, for a key you do not want to keep:
|
|
501
|
+
|
|
502
|
+
```ruby
|
|
503
|
+
config = Judge::Configuration.new
|
|
504
|
+
config.api_key = params[:api_key]
|
|
505
|
+
Judge.ask(questions, text: text, adapter: Judge::Client.new(config: config))
|
|
506
|
+
```
|
|
507
|
+
|
|
508
|
+
## Providers
|
|
509
|
+
|
|
510
|
+
**Skip this unless you need it.** TypeSafe Jev is the default and needs no configuration: set a key
|
|
511
|
+
and everything above works. This section is the escape hatch.
|
|
512
|
+
|
|
513
|
+
Everything above goes through one seam. An adapter is any object answering a single method:
|
|
514
|
+
|
|
515
|
+
```ruby
|
|
516
|
+
call(state:, questions:, model:) -> Judge::ResultSet
|
|
517
|
+
```
|
|
518
|
+
|
|
519
|
+
`questions` is a Hash of `{name => Judge::Question}`. The adapter reaches a provider however it
|
|
520
|
+
likes and builds each answer from typed values, so **no adapter ever writes or reads a wire
|
|
521
|
+
format**:
|
|
522
|
+
|
|
523
|
+
```ruby
|
|
524
|
+
class LayaAdapter
|
|
525
|
+
def call(state:, questions:, model: nil)
|
|
526
|
+
answers = MyLayaService.judge(state, questions.transform_values(&:to_payload))
|
|
527
|
+
|
|
528
|
+
results = questions.map do |name, question|
|
|
529
|
+
answer = answers.fetch(name)
|
|
530
|
+
Judge::Result.from_values(name: name, question: question, type: question.type,
|
|
531
|
+
value: answer.value, confidence: answer.confidence,
|
|
532
|
+
probabilities: answer.distribution, legend: answer.legend)
|
|
533
|
+
end
|
|
534
|
+
Judge::ResultSet.new(results, model: "laya-1")
|
|
535
|
+
end
|
|
536
|
+
end
|
|
537
|
+
|
|
538
|
+
Judge::Adapter.register(:laya) { LayaAdapter.new }
|
|
539
|
+
Judge.configure { |c| c.adapter = :laya }
|
|
540
|
+
```
|
|
541
|
+
|
|
542
|
+
The default is `:jev` and stays `:jev` until you change it. `JUDGE_ADAPTER` picks one from the
|
|
543
|
+
environment, and `adapter:` overrides it for a single call, which is how the test suite installs a
|
|
544
|
+
recorder.
|
|
545
|
+
|
|
546
|
+
### You may not need an adapter at all
|
|
547
|
+
|
|
548
|
+
The portable thing is the protocol, not this registry. Anything already serving this shape works
|
|
549
|
+
with the default adapter and a changed `base_url`, with no code:
|
|
550
|
+
|
|
551
|
+
```
|
|
552
|
+
POST <base_url>
|
|
553
|
+
Authorization: Bearer <key>
|
|
554
|
+
|
|
555
|
+
{"state": "...",
|
|
556
|
+
"model": "...",
|
|
557
|
+
"questions": {"urgency": {"type": "noul", "instructions": "...", "criteria": {...}}}}
|
|
558
|
+
```
|
|
559
|
+
|
|
560
|
+
```json
|
|
561
|
+
{"model": "...",
|
|
562
|
+
"answers": {"urgency": {"type": "noul", "noul": 0.96}},
|
|
563
|
+
"usage": {"input_tokens": 430, "output_tokens": 73}}
|
|
564
|
+
```
|
|
565
|
+
|
|
566
|
+
`choice` answers carry `choice`, `confidence` and `probabilities`; `score` answers carry `score`,
|
|
567
|
+
`confidence`, `legend` and `probabilities`.
|
|
568
|
+
|
|
569
|
+
## Without Rails
|
|
570
|
+
|
|
571
|
+
```ruby
|
|
572
|
+
require "judge" # questions, results, client, thread pool. Loads nothing from Rails
|
|
573
|
+
require "judge/rails" # the ActiveRecord layer
|
|
574
|
+
```
|
|
575
|
+
|
|
576
|
+
`require "judge"` pulls in nothing but the standard library: no ActiveRecord, no ActiveSupport, and
|
|
577
|
+
`net/http` is not loaded until you make a call. The gemspec declares `activerecord` and
|
|
578
|
+
`activesupport` because `require "judge/rails"` needs them, which is the whole dependency list.
|
|
579
|
+
|
|
580
|
+
## Demo
|
|
581
|
+
|
|
582
|
+
`judge_rails_demo` is a Rails app with 250 support tickets judged offline, a datatable built on these
|
|
583
|
+
scopes, and six pages explaining the escalation band, one-call batching, self-invalidation, per-request
|
|
584
|
+
keys, ad-hoc filtering and validation failure modes.
|
|
585
|
+
|
|
586
|
+
## Licence
|
|
587
|
+
|
|
588
|
+
MIT.
|