judge_rails 0.0.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +7 -0
- data/ADVANCED.md +136 -0
- data/BENCHMARK.md +885 -0
- data/CHANGELOG.md +163 -0
- data/LICENSE.txt +21 -0
- data/README.md +588 -0
- data/lib/generators/judge/attribute_generator.rb +149 -0
- data/lib/generators/judge/install_generator.rb +36 -0
- data/lib/generators/judge/templates/initializer.rb.tt +18 -0
- data/lib/generators/judge/templates/migration.rb.tt +9 -0
- data/lib/judge/adapter.rb +69 -0
- data/lib/judge/client.rb +253 -0
- data/lib/judge/configuration.rb +103 -0
- data/lib/judge/decision.rb +24 -0
- data/lib/judge/errors.rb +36 -0
- data/lib/judge/facade.rb +75 -0
- data/lib/judge/pool.rb +97 -0
- data/lib/judge/question/choice.rb +36 -0
- data/lib/judge/question/noul.rb +22 -0
- data/lib/judge/question/score.rb +36 -0
- data/lib/judge/question.rb +93 -0
- data/lib/judge/rails/attributes.rb +161 -0
- data/lib/judge/rails/definition.rb +156 -0
- data/lib/judge/rails/jobs.rb +156 -0
- data/lib/judge/rails/locale/en.yml +6 -0
- data/lib/judge/rails/migration.rb +115 -0
- data/lib/judge/rails/refresh.rb +37 -0
- data/lib/judge/rails/refresh_job.rb +15 -0
- data/lib/judge/rails/registry.rb +57 -0
- data/lib/judge/rails/relation.rb +132 -0
- data/lib/judge/rails/scopes.rb +124 -0
- data/lib/judge/rails/storage.rb +104 -0
- data/lib/judge/rails/validator.rb +246 -0
- data/lib/judge/rails.rb +45 -0
- data/lib/judge/result.rb +173 -0
- data/lib/judge/result_set.rb +75 -0
- data/lib/judge/version.rb +5 -0
- data/lib/judge.rb +41 -0
- data/lib/judge_rails.rb +3 -0
- metadata +126 -0
data/BENCHMARK.md
ADDED
|
@@ -0,0 +1,885 @@
|
|
|
1
|
+
# BENCHMARK
|
|
2
|
+
|
|
3
|
+
Pre-registered protocol for the `Judge::Batch` batching engine, written before anything ran.
|
|
4
|
+
|
|
5
|
+
**Status on 2026-09-21: the campaign is finished and all five arms have run. The verdict is in
|
|
6
|
+
section 7: batching fails the criterion and is not wired in.** Every figure in this document is
|
|
7
|
+
either measured or arithmetic that is labeled as such.
|
|
8
|
+
|
|
9
|
+
**Rename note.** The gem was called `jev-in-rails` when these measurements were taken, and has been
|
|
10
|
+
`judge_rails` since 2026-09-23. Paths and identifiers in this document were updated so it stays
|
|
11
|
+
reproducible. The measurements themselves are unchanged. `wiki/queries/why-the-gem-was-renamed.md`
|
|
12
|
+
explains why.
|
|
13
|
+
|
|
14
|
+
The protocol comes first because the house rule is "verify, do not assert", and a protocol fixed
|
|
15
|
+
after seeing the results proves nothing.
|
|
16
|
+
|
|
17
|
+
## 1. The question
|
|
18
|
+
|
|
19
|
+
`judge_filter` makes one network call per row, in series. pg_jev judges 2000 rows in 3.5 seconds by
|
|
20
|
+
packing 20 subjects into one `state` with one question per subject, and keeping 16 requests in
|
|
21
|
+
flight. Can batching carry over here without degrading the judgment, and by how much?
|
|
22
|
+
|
|
23
|
+
The API has no "subjects" dimension. It takes `{state, model, questions}`. Batching therefore means
|
|
24
|
+
encoding the subjects inside the `state` and asking one anchored question per subject. For the
|
|
25
|
+
model that is a different input regime. The transport stays the same.
|
|
26
|
+
|
|
27
|
+
## 2. Measured offline, free, reproducible
|
|
28
|
+
|
|
29
|
+
Everything below is read from a file in the workspace, with no network call.
|
|
30
|
+
|
|
31
|
+
### 2.1 The baseline
|
|
32
|
+
|
|
33
|
+
`judge_rails_demo/db/seed_judgments.rb`, 250 judgments produced by the unbatched path against the
|
|
34
|
+
real model.
|
|
35
|
+
|
|
36
|
+
| Quantity | Value |
|
|
37
|
+
|---|---|
|
|
38
|
+
| Judgments | 250 |
|
|
39
|
+
| Model | `jev-1.13.0`, only one |
|
|
40
|
+
| Distinct question digests | 3 (`9e08ebfeb6e0e0fa`, `5f3c0bad775abc1c`, `d3f59e90f1ce9b41`) |
|
|
41
|
+
| Mean latency | 0.309 s |
|
|
42
|
+
| Latency p50 / min / max | 0.286 s / 0.228 s / 0.773 s |
|
|
43
|
+
|
|
44
|
+
Each row carries the noul probability, the full `probabilities` distribution of the choice, and the
|
|
45
|
+
value, level and distribution of the score. Agreement is measured against this baseline.
|
|
46
|
+
|
|
47
|
+
### 2.2 The inputs
|
|
48
|
+
|
|
49
|
+
`judge_rails_demo/db/seed_tickets.rb`, where the `state` is `[subject, body]` joined by two newlines
|
|
50
|
+
(`app/models/ticket.rb:14`).
|
|
51
|
+
|
|
52
|
+
| Quantity | Value |
|
|
53
|
+
|---|---|
|
|
54
|
+
| Tickets | 250 |
|
|
55
|
+
| `state` in characters | mean 208, p50 156, max 851 |
|
|
56
|
+
| `state` total | 52,114 characters |
|
|
57
|
+
| The 3 questions + criteria | ~620 characters |
|
|
58
|
+
| Rows that fit in 24,000 characters | ~115 |
|
|
59
|
+
|
|
60
|
+
On this dataset the accuracy ceiling (20 to 25 rows) bites about five times earlier than the `state`
|
|
61
|
+
size ceiling. The demo will never reach the character budget.
|
|
62
|
+
|
|
63
|
+
### 2.3 The current paths, read from the code
|
|
64
|
+
|
|
65
|
+
| Path | File | Requests | Parallelism |
|
|
66
|
+
|---|---|---|---|
|
|
67
|
+
| `judge_filter` / `map` / `sort` | `lib/judge/rails/relation.rb:41-44` | 1 per row | **1 thread** |
|
|
68
|
+
| backfill by job | `lib/judge/rails/jobs.rb:56-61` | 1 per record | depends on the queue |
|
|
69
|
+
| single computation | `lib/judge/rails/storage.rb:26-34` | 1 per record | n/a |
|
|
70
|
+
| `judge:demo:precompute` | `judge_rails_demo/lib/tasks/judge_demo.rake:8-30` | 1 per ticket | **8 threads** |
|
|
71
|
+
|
|
72
|
+
None of these paths puts two subjects in one request. `Jobs.batch` (`jobs.rb:39-48`) groups ids to
|
|
73
|
+
reduce the number of jobs. The number of requests stays the same.
|
|
74
|
+
|
|
75
|
+
## 3. Projected: arithmetic, not measured
|
|
76
|
+
|
|
77
|
+
Labeled separately because nothing here was observed.
|
|
78
|
+
|
|
79
|
+
### 3.1 Requests to judge 250 tickets on 3 questions
|
|
80
|
+
|
|
81
|
+
| Shape | Requests | Calculation |
|
|
82
|
+
|---|---|---|
|
|
83
|
+
| Current | 250 | 1 per ticket, 3 questions in the request |
|
|
84
|
+
| A, 1 question × 20 rows | 39 | 3 definitions × ceil(250/20) |
|
|
85
|
+
| B, 3 questions × 20 rows | 13 | ceil(250/20) |
|
|
86
|
+
|
|
87
|
+
### 3.2 Tokens per row, estimated at 4 characters per token
|
|
88
|
+
|
|
89
|
+
| Shape | Tokens / row | Detail |
|
|
90
|
+
|---|---|---|
|
|
91
|
+
| A | ~400 | the ticket `state` is sent 3 times, once per definition |
|
|
92
|
+
| B | ~296 | the `state` is sent once, with 3 questions per row |
|
|
93
|
+
|
|
94
|
+
Estimated difference: **B is about 26% cheaper than A**, far from a factor of 3. With short rows the
|
|
95
|
+
per-row instructions dominate the cost, more than the `state` does.
|
|
96
|
+
|
|
97
|
+
For reference, pg_jev documents ~175 input tokens per row in batches of 20, against ~435 for a single
|
|
98
|
+
row. Our questions carry richer criteria than its conditions, hence the higher figure.
|
|
99
|
+
|
|
100
|
+
### 3.3 Campaign cost
|
|
101
|
+
|
|
102
|
+
~430 requests, ~710,000 input tokens, so **~$0.03** at the $0.042 per million input tokens quoted in
|
|
103
|
+
the pg_jev documentation. **That price is borrowed and was never checked with Jev.** The real cost
|
|
104
|
+
will show up in `ResultSet#usage`, which the client already fills in (`result_set.rb:34`).
|
|
105
|
+
|
|
106
|
+
## 4. What the protocol cannot establish
|
|
107
|
+
|
|
108
|
+
- **Absolute calibration.** We measure agreement with the unbatched path, with no human labels to
|
|
109
|
+
measure correctness against. This workspace has no human-labeled dataset.
|
|
110
|
+
- **Generalization beyond this dataset.** 250 short, synthetic support tickets in English. Long
|
|
111
|
+
documents and multilingual data are not covered.
|
|
112
|
+
- **Behavior near the 32k limit.** No input in the dataset comes close.
|
|
113
|
+
|
|
114
|
+
## 5. Design decisions that define the experiment
|
|
115
|
+
|
|
116
|
+
Taken before the run, so that the arms of the protocol mean something.
|
|
117
|
+
|
|
118
|
+
| # | Decision |
|
|
119
|
+
|---|---|
|
|
120
|
+
| D1 | Batching is opt-in. `Judge.config.batch_rows` and `Judge.config.concurrency`, `nil` by default |
|
|
121
|
+
| D2 | Only the bulk paths read the setting. The sync `before_save` cannot be batched, because a `before_save` has to finish before its own save. The validator is excluded because it fails closed |
|
|
122
|
+
| D3 | `state` = strict JSON `{condition, rows:[{id, text}]}`. One question per row, named `row0`, `row1`, with instructions prefixed by an anchor that names the id and says to treat the text as data |
|
|
123
|
+
| D4 | The batch path never substitutes the anchored question into the `Definition`. The sidecar carries the digest of the original question, otherwise `stale?` loops |
|
|
124
|
+
| D5 | Incomplete response: bisection. The batch is split in two and retried, depth capped at 5, and the one-row base case goes through the single path |
|
|
125
|
+
| D6 | The character budget is a guard that raises with the measured size in the message. Packing counts rows, not characters |
|
|
126
|
+
| D7 | This document decides the chosen shape (A or B) and the value of `batch_rows`. Nothing before it does |
|
|
127
|
+
|
|
128
|
+
## 6. Protocol
|
|
129
|
+
|
|
130
|
+
### Arm 0. Latency as a function of the number of questions. RUN on 2026-09-21
|
|
131
|
+
|
|
132
|
+
**This arm decides whether the rest is worth anything.** The vendor claims that response time barely
|
|
133
|
+
moves when you add questions. The wiki carried that claim at `confidence: medium`, never checked
|
|
134
|
+
here.
|
|
135
|
+
|
|
136
|
+
`judge_rails/bin/arm0_latency`. One 280-character state, the same question repeated N times, 3
|
|
137
|
+
repetitions per point, 18 requests against the real model.
|
|
138
|
+
|
|
139
|
+
| N questions | Median latency | Spread | Input tokens | Tokens / question | vs N=1 |
|
|
140
|
+
|---|---|---|---|---|---|
|
|
141
|
+
| 1 | 0.280 s | 0.428 s | 350 | 350 | 1.00x |
|
|
142
|
+
| 2 | 0.264 s | 0.032 s | 373 | 186 | 0.94x |
|
|
143
|
+
| 5 | 0.242 s | 0.031 s | 442 | 88 | 0.86x |
|
|
144
|
+
| 10 | 0.234 s | 0.014 s | 557 | 56 | 0.84x |
|
|
145
|
+
| 20 | 0.242 s | 0.020 s | 797 | 40 | 0.86x |
|
|
146
|
+
| 40 | 0.246 s | 0.024 s | 1,277 | 32 | 0.88x |
|
|
147
|
+
|
|
148
|
+
**Latency is flat.** 40 questions in one request take the same time as one, at 0.88x. The go/no-go
|
|
149
|
+
criterion asked for less than 10x at N=20, and we measured 0.86x. The 0.428 s spread at N=1 is the
|
|
150
|
+
first request on a cold connection, and it does not recur at any other point.
|
|
151
|
+
|
|
152
|
+
The vendor's claim is **verified**, and the wiki can move `[[typesafe-jev]]` from `medium` to `high`
|
|
153
|
+
on this specific point.
|
|
154
|
+
|
|
155
|
+
**Cost structure, measured.** The fixed overhead of a request is about **326 tokens** (350 at N=1
|
|
156
|
+
minus the marginal cost of one question), and a marginal question costs
|
|
157
|
+
**(1277 - 350) / 39 ≈ 24 tokens**. Batching amortizes this fixed overhead. pg_jev documents the same
|
|
158
|
+
mechanism without putting a number on it for us.
|
|
159
|
+
|
|
160
|
+
Arithmetic consequence for the demo: judging 250 tickets without batching pays the overhead 250
|
|
161
|
+
times, ~81,500 tokens of pure structure. In batches of 20 it pays 13 times, ~4,200. **~77,000 tokens
|
|
162
|
+
saved on structure alone.**
|
|
163
|
+
|
|
164
|
+
**This arm refutes arm 4's conclusion on cost.** The byte proxy predicted +24.6% for batching. Real
|
|
165
|
+
billing says 8.75x fewer tokens per question at N=20. Bytes on the wire are not the bill: the ~326
|
|
166
|
+
token overhead has no visible counterpart in the payload. The proxy is dropped as a cost instrument
|
|
167
|
+
and `ResultSet#usage` replaces it.
|
|
168
|
+
|
|
169
|
+
**Caveat.** This arm keeps the state fixed and varies N. In a real batch the state grows with the
|
|
170
|
+
rows. The arm shows that latency does not depend on N and that the per-request overhead is real. The
|
|
171
|
+
total cost of a batch is an output of arms 2 and 3.
|
|
172
|
+
|
|
173
|
+
### Arm 1. Control: model determinism. RUN on 2026-09-21
|
|
174
|
+
|
|
175
|
+
`bin/rails judge:bench:control SAMPLE=50`. The unbatched path replayed on 50 tickets and compared to
|
|
176
|
+
the baseline. This is the noise floor, and no other arm can be read without it.
|
|
177
|
+
|
|
178
|
+
| Question | n | Decision agreement | Mean \|Δ\| | p95 \|Δ\| |
|
|
179
|
+
|---|---|---|---|---|
|
|
180
|
+
| urgency (noul) | 50 | **100.0%** | 0.0118 | 0.0300 |
|
|
181
|
+
| intent (choice) | 50 | **100.0%** | 0.0082 | 0.0400 |
|
|
182
|
+
| frustration (score) | 50 | **96.0%** | 0.0108 | 0.0400 |
|
|
183
|
+
|
|
184
|
+
50 requests, 26,492 input tokens, 13.52 s.
|
|
185
|
+
|
|
186
|
+
The model is close to deterministic. Mean drift is ~0.01 and decisions are stable, except for one
|
|
187
|
+
score level in 25 that flips. **The noise floor is therefore ~0.01 drift and 96 to 100% agreement.**
|
|
188
|
+
Everything that follows is read against those three numbers.
|
|
189
|
+
|
|
190
|
+
### Arms 2 and 3. Shapes A and B, four batch sizes. RUN on 2026-09-21
|
|
191
|
+
|
|
192
|
+
`bin/rails judge:bench:campaign SHAPE=a|b BATCH=n`, 250 tickets, compared to the baseline.
|
|
193
|
+
|
|
194
|
+
**Decision agreement, in percent. Control on the first row.**
|
|
195
|
+
|
|
196
|
+
| Batch | Shape | urgency | intent | frustration | Requests | Tokens |
|
|
197
|
+
|---|---|---|---|---|---|---|
|
|
198
|
+
| 1 | control | **100.0** | **100.0** | **96.0** | 250* | 132,460* |
|
|
199
|
+
| 2 | A | 91.6 | 90.4 | 90.0 | 375 | 228,021 |
|
|
200
|
+
| 2 | B | 93.2 | 89.2 | 89.2 | 125 | 129,882 |
|
|
201
|
+
| 5 | A | 91.2 | 89.6 | 84.4 | 150 | 165,999 |
|
|
202
|
+
| 5 | B | 91.2 | 90.4 | 86.4 | 50 | 110,083 |
|
|
203
|
+
| 10 | A | 88.0 | 87.6 | 83.6 | 75 | 145,324 |
|
|
204
|
+
| 10 | B | 88.4 | 88.0 | 85.6 | 25 | 103,483 |
|
|
205
|
+
| 20 | A | 86.4 | 87.6 | 84.4 | 39 | 136,120 |
|
|
206
|
+
| 20 | B | 85.6 | 89.6 | 86.8 | 13 | 100,795 |
|
|
207
|
+
| 40 | A | 83.2 | 87.6 | 82.4 | 21 | 131,518 |
|
|
208
|
+
| 40 | B | 85.2 | 86.8 | 84.4 | 7 | 99,451 |
|
|
209
|
+
|
|
210
|
+
\* control extrapolated from 50 to 250 tickets for the comparison.
|
|
211
|
+
|
|
212
|
+
**Mean \|Δ\| drift, control at ~0.01.**
|
|
213
|
+
|
|
214
|
+
| Batch | Shape | urgency | intent | frustration |
|
|
215
|
+
|---|---|---|---|---|
|
|
216
|
+
| 1 | control | 0.0118 | 0.0082 | 0.0108 |
|
|
217
|
+
| 2 | B | 0.0532 | 0.0804 | 0.1110 |
|
|
218
|
+
| 5 | B | 0.0730 | 0.0937 | 0.1394 |
|
|
219
|
+
| 10 | B | 0.0816 | 0.1042 | 0.1543 |
|
|
220
|
+
| 20 | B | 0.0819 | 0.1097 | 0.1587 |
|
|
221
|
+
| 40 | B | 0.0874 | 0.1146 | 0.1780 |
|
|
222
|
+
|
|
223
|
+
### Arm 5. 2x2 isolation, one row per request. RUN on 2026-09-21
|
|
224
|
+
|
|
225
|
+
`bin/rails judge:bench:isolate VARIANT=... SAMPLE=150`. With one row per request there is no
|
|
226
|
+
neighbor, so this square separates the two other causes: the JSON envelope around the `state`, and
|
|
227
|
+
the anchor prefix on the instructions.
|
|
228
|
+
|
|
229
|
+
| Variant | `state` | instructions | urgency | intent | frustration |
|
|
230
|
+
|---|---|---|---|---|---|
|
|
231
|
+
| `raw_plain` | raw text | original | **98.7** | **100.0** | **98.7** |
|
|
232
|
+
| `json_plain` | JSON `{rows:[...]}` | original | 96.0 | 98.0 | 95.3 |
|
|
233
|
+
| `raw_anchored` | raw text | prefixed | 98.0 | 94.0 | 95.3 |
|
|
234
|
+
| `json_anchored` | JSON | prefixed | 96.0 | 96.7 | 96.0 |
|
|
235
|
+
|
|
236
|
+
600 requests, ~351,000 tokens, ~22 s in total.
|
|
237
|
+
|
|
238
|
+
`raw_plain` reproduces the control, which validates the harness. Each of the two transformations
|
|
239
|
+
costs about 2 to 4 points on its own, and the two together cost about **3 points**.
|
|
240
|
+
|
|
241
|
+
### Breaking down the gap
|
|
242
|
+
|
|
243
|
+
The batch of 20 lost 10 to 14 points against the control. Arm 5 says where they come from:
|
|
244
|
+
|
|
245
|
+
| Cause | Points lost | Avoidable? |
|
|
246
|
+
|---|---|---|
|
|
247
|
+
| JSON envelope + anchor prefix | **~3** | maybe, by changing the shape |
|
|
248
|
+
| Neighbors present in the `state` | **~9** | no, that is batching itself |
|
|
249
|
+
|
|
250
|
+
The dominant cause is the neighbors, more than the JSON or the prefix.
|
|
251
|
+
|
|
252
|
+
**Corrected on 2026-09-23 by arm 7.** This split is a residual, not a measurement, and arm 7 refutes
|
|
253
|
+
it at a batch of 2: there the neighbor's content costs -0.3 points, the object shape -5.6 and a
|
|
254
|
+
second key -3.2. See arm 7.
|
|
255
|
+
|
|
256
|
+
### What the TypeSafe documentation says, and how it explains the result
|
|
257
|
+
|
|
258
|
+
The [Parallel Questions](https://docs.typesafe.ai/cookbooks/parallel_questions.md) cookbook measures
|
|
259
|
+
13 batched questions against 13 separate questions on a 54,000-character article: **12.2x cheaper,
|
|
260
|
+
10.0x faster, no change in answers**, with a standard deviation of exactly 0.0 on most questions. It
|
|
261
|
+
gives the reason:
|
|
262
|
+
|
|
263
|
+
> "each question is scored on its own against the document, so its answer doesn't depend on what
|
|
264
|
+
> else is in the request."
|
|
265
|
+
|
|
266
|
+
A question does not depend on the other questions. It depends entirely on the **document**. Putting
|
|
267
|
+
20 tickets in the `state` leaves the questions alone and changes the document each one is scored
|
|
268
|
+
against. Every ticket is judged in the middle of 19 texts that have nothing to do with it.
|
|
269
|
+
|
|
270
|
+
The [Re-Ranking](https://docs.typesafe.ai/cookbooks/rerank_typesafe.md) cookbook is the same problem
|
|
271
|
+
as ours, 3,565 passages to judge independently, and TypeSafe solves it with a
|
|
272
|
+
`{query_excerpt, candidate_passage}` state, **one candidate per request**, and **1,200 concurrent
|
|
273
|
+
calls in a thread pool**, 1,536,002 input tokens for $0.0645. Batching by subject does not appear.
|
|
274
|
+
|
|
275
|
+
**Corrected on 2026-09-23.** The cookbook does not say "1,200 concurrent calls". It runs 40 queries ×
|
|
276
|
+
30 candidates preselected by BM25 out of the 3,565 passages, 1,200 calls in total, in a
|
|
277
|
+
`ThreadPoolExecutor(max_workers=12)`. The lesson we kept (one candidate per request) still holds. The
|
|
278
|
+
concurrency figure never existed. Re-read in `wiki/raw/transcripts/vendor-docs-check-2026-09-23.md`.
|
|
279
|
+
|
|
280
|
+
**Conclusion.** Batching by question is the documented route and the gem already follows it: three
|
|
281
|
+
`judge_attribute` on a ticket go out in one request (`storage.rb:26-34`). Batching by subject is not
|
|
282
|
+
a documented route, and the measurement shows why.
|
|
283
|
+
|
|
284
|
+
### The diagnosis behind the curve
|
|
285
|
+
|
|
286
|
+
The degradation is **already complete at a batch of 2** and moves little up to 40. Going from 1 to 2
|
|
287
|
+
rows costs about 8 points of agreement. Going from 2 to 40 costs only 5 more.
|
|
288
|
+
|
|
289
|
+
**The dominant cause is the change of shape, more than the number of neighbors.** Two things change
|
|
290
|
+
at once between the control and a batch of 2:
|
|
291
|
+
|
|
292
|
+
1. the `state` goes from raw text to a JSON object `{condition, rows:[{id, text}]}`,
|
|
293
|
+
2. the question instructions are rewritten with an anchor prefix.
|
|
294
|
+
|
|
295
|
+
The protocol does not separate these two causes. That is the next test to run, and it is cheap.
|
|
296
|
+
|
|
297
|
+
### Shape A against shape B
|
|
298
|
+
|
|
299
|
+
Agreement is the same for both shapes, within the noise. B is strictly cheaper: at a batch of 20, 13
|
|
300
|
+
requests against 39 and 100,795 tokens against 136,120, so **26% fewer tokens for the same
|
|
301
|
+
accuracy**. If batching were adopted, it would take shape B. That settles A against B, though it
|
|
302
|
+
does not make batching acceptable.
|
|
303
|
+
|
|
304
|
+
### What batching would buy, and at what price
|
|
305
|
+
|
|
306
|
+
| Axis | Control | Shape B, batch 20 | Effect |
|
|
307
|
+
|---|---|---|---|
|
|
308
|
+
| Wall time per ticket | 0.270 s | 0.0055 s | **49x faster** |
|
|
309
|
+
| Tokens per ticket | 530 | 403 | **-24%** |
|
|
310
|
+
| urgency agreement | 100.0% | 85.6% | **-14.4 points** |
|
|
311
|
+
| intent agreement | 100.0% | 89.6% | **-10.4 points** |
|
|
312
|
+
| frustration agreement | 96.0% | 86.8% | **-9.2 points** |
|
|
313
|
+
|
|
314
|
+
The speed gain is real and large. The token saving is real and modest, far from the 8.75x that arm 0
|
|
315
|
+
suggested: once the rows are in the `state`, the `state` dominates, and the amortized fixed overhead
|
|
316
|
+
no longer weighs much.
|
|
317
|
+
|
|
318
|
+
### Arm 4. Offline throughput, no API. RUN on 2026-09-21
|
|
319
|
+
|
|
320
|
+
`judge_rails/bin/bench_batch`, against `FakeJev` with **0.300 s** of injected latency, the mean
|
|
321
|
+
measured in 2.1. 250 rows of a size comparable to the demo's. It measures the thread pool and the
|
|
322
|
+
packing. The model is not involved.
|
|
323
|
+
|
|
324
|
+
| Configuration | Requests | Wall time | Bytes sent | vs sequential 1 thread |
|
|
325
|
+
|---|---|---|---|---|
|
|
326
|
+
| sequential, 1 thread | 250 | 76.11 s | 127,480 | 1.0x |
|
|
327
|
+
| sequential, 8 threads | 250 | 9.75 s | 127,480 | 7.8x |
|
|
328
|
+
| batch 20, 8 threads | 13 | 0.62 s | 158,858 | **122.8x** |
|
|
329
|
+
| batch 20, 16 threads | 13 | 0.32 s | 158,858 | **241.4x** |
|
|
330
|
+
|
|
331
|
+
Reproducible: `cd judge_rails && bundle exec ruby bin/bench_batch`. `ROWS` and `LATENCY` can be
|
|
332
|
+
overridden.
|
|
333
|
+
|
|
334
|
+
**A caveat that governs the whole table.** `FakeJev` answers in a fixed time whatever the number of
|
|
335
|
+
questions in the request. The 122x and 241x figures therefore assume, **by construction**, that
|
|
336
|
+
latency does not grow with the number of questions. Arm 0 exists to test exactly that assumption.
|
|
337
|
+
Until arm 0 has run, these two figures are a theoretical ceiling, not a measurement of the real gain.
|
|
338
|
+
|
|
339
|
+
**A result that contradicts a starting assumption.** Batching sends **more** bytes: 158,858 against
|
|
340
|
+
127,480, **+24.6%**, or 635 bytes per row against 510. There are two causes, both measured here: the
|
|
341
|
+
anchor prefix is sent once per row instead of once per request, and the JSON wrapping of each row
|
|
342
|
+
(`{"id":n,"text":...}`) adds bytes the single path does not have.
|
|
343
|
+
|
|
344
|
+
So **the offline bench cannot settle the cost axis**, contrary to what section 3.2 projected. The
|
|
345
|
+
token saving pg_jev documents (175 against 435 per row) comes from amortizing a per-request overhead
|
|
346
|
+
of about 270 tokens, which is billed on the server and does not appear in the payload. Only the
|
|
347
|
+
`ResultSet#usage` returned by the real API can decide, which makes the token count an output of arms
|
|
348
|
+
1 to 3 and not of arm 4.
|
|
349
|
+
|
|
350
|
+
### Arm 4b, in the test suite
|
|
351
|
+
|
|
352
|
+
`test/batch_throughput_test.rb`, 6 tests, 0.1 s of injected latency, small volumes. It checks in
|
|
353
|
+
under a second that the pool really overlaps requests and that batching reduces the request count.
|
|
354
|
+
The faithful version of arm 4 takes 76 seconds, so it has no place in `rake test`.
|
|
355
|
+
|
|
356
|
+
### Arm 6. Concurrency sweep on the real API. RUN on 2026-09-21
|
|
357
|
+
|
|
358
|
+
`bin/rails judge:bench:isolate VARIANT=raw_plain SAMPLE=100 CONCURRENCY=n`. The exact production
|
|
359
|
+
shape: `state` in raw text, 3 questions per ticket, one ticket per request. Only the thread count
|
|
360
|
+
changes.
|
|
361
|
+
|
|
362
|
+
| Threads | Wall time | Speedup | Efficiency | urgency | intent | frustration | Tokens |
|
|
363
|
+
|---|---|---|---|---|---|---|---|
|
|
364
|
+
| 1 | 25.01 s | 1.0x | 100% | 99.0 | 99.0 | 98.0 | 52,859 |
|
|
365
|
+
| 2 | 13.02 s | 1.92x | 96% | 100.0 | 99.0 | 100.0 | 52,859 |
|
|
366
|
+
| 4 | 6.87 s | 3.64x | 91% | 99.0 | 99.0 | 99.0 | 52,859 |
|
|
367
|
+
| 8 | 3.60 s | **6.95x** | 87% | 99.0 | 99.0 | 100.0 | 52,859 |
|
|
368
|
+
| 16 | 2.13 s | **11.7x** | 73% | 100.0 | 99.0 | 99.0 | 52,859 |
|
|
369
|
+
| 32 | 1.36 s | **18.4x** | 57% | 99.0 | 99.0 | 98.0 | 52,859 |
|
|
370
|
+
|
|
371
|
+
What the table shows:
|
|
372
|
+
|
|
373
|
+
1. **No ceiling observed up to 32.** No errors, no 429s, and wall time keeps dropping. Efficiency
|
|
374
|
+
degrades (57% at 32) but the speedup stays monotonic.
|
|
375
|
+
**Caveat added on 2026-09-23:** `models.md` documents 1,200 requests per minute. 100 requests per
|
|
376
|
+
level never fill a minute, so this arm could not see the limit. See arm 9.
|
|
377
|
+
2. **Accuracy does not move.** 98 to 100% at every level, the same as in series. That is expected,
|
|
378
|
+
and it is the basic difference from batching: the request is **byte for byte the same**, and only
|
|
379
|
+
its scheduling changes.
|
|
380
|
+
3. **Tokens are identical to the unit**, 52,859 at every level. Concurrency costs nothing.
|
|
381
|
+
|
|
382
|
+
**Default chosen: 8.** 6.95x at 87% efficiency, and 8 connections per process is reasonable for a
|
|
383
|
+
gem that runs inside an application server. The sweep ran on one machine, with one key and 100
|
|
384
|
+
requests. It says nothing about the quota of a shared key under real load. `Judge.config.concurrency`
|
|
385
|
+
and the `concurrency:` keyword exist for anyone whose own measurement says otherwise.
|
|
386
|
+
|
|
387
|
+
### Arm 7. Object `state` with named keys. PRE-REGISTERED, then RUN on 2026-09-23
|
|
388
|
+
|
|
389
|
+
**Why this arm.** Re-read on 2026-09-23, arm 5 does not measure what the "breaking down the gap"
|
|
390
|
+
section attributes to it. The ~9 points "from the neighbors" are a residual, not a measurement.
|
|
391
|
+
Between one JSON row and a batch of 20, several things change at once: the presence of foreign text
|
|
392
|
+
(dilution), designating the subject by a numeric `id` in an array (addressing), the number of
|
|
393
|
+
questions and the length of the `state`. The `raw_anchored` cell is inconsistent: it mentions a
|
|
394
|
+
"rows array" that a text `state` does not contain. The sample for arms 1 and 5 (`first(n)` by id)
|
|
395
|
+
contains no spam ticket.
|
|
396
|
+
|
|
397
|
+
The TypeSafe documentation (`concepts/state.md`, `primitives/advanced.md`) describes an **object**
|
|
398
|
+
`state` where each part has a name, and instructions that name the part they target. A manual test
|
|
399
|
+
in the playground, `{"State_1": "Shut up !", "State_2": "Shut up ?"}` with "Is State_1 a question ?",
|
|
400
|
+
returns 2% and 98%: the API accepts the shape and addressing holds on a minimal pair. This arm
|
|
401
|
+
measures whether addressing by named key recovers the accuracy that `rows[id]` lost.
|
|
402
|
+
|
|
403
|
+
**Shape.** `state` = a real JSON object (not a string), keys `ticket_1` to `ticket_N`. One question
|
|
404
|
+
per key and per definition, named `ticket_k_<name>`, whose original instruction is rewritten to name
|
|
405
|
+
the key ("Does ticket_1 need a human...", "Which team should handle ticket_1?", "How frustrated does
|
|
406
|
+
the customer in ticket_1 sound?"). No anchor prefix. Criteria unchanged.
|
|
407
|
+
|
|
408
|
+
**Variants, all on the 250 tickets**, compared to the `db/seed_judgments.rb` baseline:
|
|
409
|
+
|
|
410
|
+
| Variant | `state` | Requests | Isolates |
|
|
411
|
+
|---|---|---|---|
|
|
412
|
+
| `control` | raw text, original questions | 250 | noise floor on the same population |
|
|
413
|
+
| `named_1` | `{ticket_1}` | 250 | cost of the object shape alone |
|
|
414
|
+
| `named_same` | `{ticket_1, ticket_2}`, the same ticket twice | 250 | cost of a second key, with no foreign content |
|
|
415
|
+
| `named_2` | two different tickets | 125 | addressing + dilution, where the gap appeared |
|
|
416
|
+
| `named_20` | twenty different tickets | 13 | comparable to shape B at batch 20 |
|
|
417
|
+
|
|
418
|
+
Tickets are shuffled with a fixed seed (`SEED=7`) before grouping, so neighbors do not share a
|
|
419
|
+
category by construction. Concurrency 8. Each judged row is written as JSONL to `tmp/bench/`, with
|
|
420
|
+
the model that served it (`ResultSet#model`) and the neighbor ids.
|
|
421
|
+
|
|
422
|
+
**Criterion, fixed before the run.** A variant passes if, on each of the three questions, its
|
|
423
|
+
agreement is at least that of `control` minus 2 points (the range observed in arm 6 over six
|
|
424
|
+
repetitions of the production shape).
|
|
425
|
+
|
|
426
|
+
**Reading, fixed before the run.**
|
|
427
|
+
|
|
428
|
+
- `named_1` passes: the object shape is free.
|
|
429
|
+
- `named_same` passes and `named_2` fails: the cause is foreign content, and dilution is confirmed.
|
|
430
|
+
- `named_same` fails: a second key has a cost of its own, whatever its content.
|
|
431
|
+
- `named_2` and `named_20` pass: the 9 points of arms 2 and 3 came from the `rows[id]` shape, and
|
|
432
|
+
batching by subject is back on the table.
|
|
433
|
+
- Contamination: among the disagreements in `named_2`, the share where the answer equals the
|
|
434
|
+
neighbor's baseline answer, compared to chance.
|
|
435
|
+
|
|
436
|
+
Budget: 888 requests, ~500,000 input tokens, ~$0.02 at the borrowed price.
|
|
437
|
+
|
|
438
|
+
Command: `bin/rails judge:bench:named VARIANT=control|named_1|named_same|named_2|named_20`.
|
|
439
|
+
`DRY_RUN=1` prints the first request and calls nothing.
|
|
440
|
+
|
|
441
|
+
#### Results, 2026-09-23
|
|
442
|
+
|
|
443
|
+
888 requests, 668,120 input tokens, no errors. All 1,250 judgments were served by `jev-1.13.0`, the
|
|
444
|
+
baseline model, checked on `ResultSet#model` rather than assumed.
|
|
445
|
+
|
|
446
|
+
| Variant | urgency | intent | frustration | Mean | Requests | Tokens | Wall time |
|
|
447
|
+
|---|---|---|---|---|---|---|---|
|
|
448
|
+
| `control` | **99.2** | **98.8** | **97.6** | 98.5 | 250 | 134,664 | 9.55 s |
|
|
449
|
+
| `named_1` | 93.2 | 96.0 | 89.6 | 92.9 | 250 | 138,164 | 9.07 s |
|
|
450
|
+
| `named_same` | 92.8 | 94.0 | 82.4 | 89.7 | 250 | 211,828 | 9.85 s |
|
|
451
|
+
| `named_2` | 91.6 | 90.8 | 86.0 | 89.5 | 125 | 105,914 | 4.96 s |
|
|
452
|
+
| `named_20` | 84.4 | 84.0 | 85.2 | 84.5 | 13 | 77,550 | 0.80 s |
|
|
453
|
+
|
|
454
|
+
Pass threshold: 97.2 / 96.8 / 95.6. **No named variant passes, not even `named_1`.** `control`
|
|
455
|
+
reproduces arm 1 on the whole population, spam included, which validates the harness.
|
|
456
|
+
|
|
457
|
+
**Direction of the disagreements**, band and level moving up or down relative to the baseline:
|
|
458
|
+
|
|
459
|
+
| Variant | urgency up / down | frustration up / down | Dominant intent flip |
|
|
460
|
+
|---|---|---|---|
|
|
461
|
+
| `control` | 0 / 2 | 3 / 3 | none, 3 isolated flips |
|
|
462
|
+
| `named_1` | 3 / 14 | 1 / 25 | none |
|
|
463
|
+
| `named_same` | 2 / 16 | 1 / 43 | sales to spam x4 |
|
|
464
|
+
| `named_2` | 8 / 13 | 4 / 31 | billing to technical x9 |
|
|
465
|
+
| `named_20` | 3 / 36 | 10 / 27 | billing to technical x18 |
|
|
466
|
+
|
|
467
|
+
In the control the disagreements are symmetric, which is noise. In every named variant they go one
|
|
468
|
+
way: **a ticket placed as a named field of an object reads calmer and less urgent.** That is a
|
|
469
|
+
systematic calibration shift, with no sign of random scatter.
|
|
470
|
+
|
|
471
|
+
**Contamination in `named_2`**, disagreements equal to the neighbor's baseline against chance:
|
|
472
|
+
urgency 7 against 8.0, intent 2 against 6.1, frustration 7 against 10.1. **Wrong answers do not copy
|
|
473
|
+
the neighbor.** Agreement by position in `named_2`: 89.3% for `ticket_1`, 89.6% for `ticket_2`, no
|
|
474
|
+
position effect.
|
|
475
|
+
|
|
476
|
+
#### Reading, against the grid fixed before the run
|
|
477
|
+
|
|
478
|
+
| Step | Mean | Cost | What changes |
|
|
479
|
+
|---|---|---|---|
|
|
480
|
+
| `control` | 98.5 | - | - |
|
|
481
|
+
| `named_1` | 92.9 | **-5.6** | the shape: a one-key object, an instruction that names the key |
|
|
482
|
+
| `named_same` | 89.7 | **-3.2** | a second key, with no new content at all |
|
|
483
|
+
| `named_2` | 89.5 | **-0.3** | the neighbor's content |
|
|
484
|
+
| `named_20` | 84.5 | -5.0 | 18 more tickets |
|
|
485
|
+
|
|
486
|
+
1. **`named_1` fails: the object shape is not free.** It is the biggest measured step, and it is
|
|
487
|
+
there with no neighbor at all.
|
|
488
|
+
2. **`named_same` fails: a second key has a cost of its own**, even though it only carries the same
|
|
489
|
+
text. Frustration drops the most here, 43 of 44 disagreements.
|
|
490
|
+
3. **Foreign content costs almost nothing at a batch of 2**: -0.3 points between `named_same` and
|
|
491
|
+
`named_2`, and contamination at or below chance. **"Dilution", in the sense of a neighbor bleeding
|
|
492
|
+
into the answer, is not observed.**
|
|
493
|
+
4. **`named_20` fails** at the level of shape B at batch 20 (85.6 / 89.6 / 86.8). Batching by subject
|
|
494
|
+
stays closed, whatever the addressing scheme.
|
|
495
|
+
|
|
496
|
+
**What this corrects.** The explanation written after arm 5, "3 points of shape, 9 points of
|
|
497
|
+
neighbors", has the wrong split. At a batch of 2, nearly all of the gap comes from leaving the
|
|
498
|
+
production shape (-5.6) and from the presence of a second key (-3.2). What the neighbor says barely
|
|
499
|
+
matters. At a batch of 20, the extra 5 points remain unattributed: `state` length, number of
|
|
500
|
+
questions and foreign content all change together.
|
|
501
|
+
|
|
502
|
+
**What this arm cannot say.** Agreement is measured against a baseline produced from raw text. A
|
|
503
|
+
shift toward "calmer" is a disagreement with the production shape, and not necessarily an error:
|
|
504
|
+
without human labels we cannot tell which of the two readings is more accurate. We do know they are
|
|
505
|
+
not interchangeable, and that the demo's thresholds (0.8 and 0.2) are tuned on the text shape.
|
|
506
|
+
`named_1` also conflates two changes, the object and the rewritten instruction ("this message"
|
|
507
|
+
becomes "ticket_1"), and this arm does not separate them. There is one run per variant. At n=250 one
|
|
508
|
+
ticket is worth 0.4 points, and the differences we keep are 8 to 40 tickets.
|
|
509
|
+
|
|
510
|
+
**Decision unchanged.** One subject per request, raw text, original questions. The thread pool is
|
|
511
|
+
still the way forward. Arm 7 does not reopen batching. It replaces the reason for rejecting it.
|
|
512
|
+
|
|
513
|
+
Raw data: `judge_rails_demo/tmp/bench/arm7_*.jsonl`, snapshotted in
|
|
514
|
+
`wiki/raw/transcripts/named-state-arm7-2026-09-23.md`.
|
|
515
|
+
|
|
516
|
+
### Reference labels. PRE-REGISTERED on 2026-09-23
|
|
517
|
+
|
|
518
|
+
**Why.** Every previous arm measures agreement with raw text, never correctness.
|
|
519
|
+
`db/seed_tickets.rb` carries no labels. Without a reference, the "calmer, less urgent" shift in arm 7
|
|
520
|
+
cannot be settled.
|
|
521
|
+
|
|
522
|
+
**Shape.** Two independent annotators label the 250 tickets on the three questions, using the exact
|
|
523
|
+
definitions in `app/models/ticket.rb`: urgency true or false, intent among the four teams,
|
|
524
|
+
frustration from 0 to 3. They read only `db/seed_tickets.rb`, never a Jev judgment. One reads in
|
|
525
|
+
order, the other in reverse. Each annotator flags the calls they consider borderline.
|
|
526
|
+
|
|
527
|
+
**Caveat, written before seeing the result.** Both annotators are models (Claude Opus), not humans.
|
|
528
|
+
The reference is their **consensus**: a ticket counts for a question only if both agree.
|
|
529
|
+
Disagreements are listed for human review. Any figure read against this reference is written as
|
|
530
|
+
"agreement with the annotator consensus", never as "correctness".
|
|
531
|
+
|
|
532
|
+
Storage: `db/ticket_labels.json`, keyed by `subject|customer_name` like `db/seed_judgments.rb`.
|
|
533
|
+
|
|
534
|
+
### Analyses without requests. PRE-REGISTERED on 2026-09-23
|
|
535
|
+
|
|
536
|
+
`bin/rails judge:eval:offline`. Reads `db/seed_judgments.rb` and `db/ticket_labels.json` and calls
|
|
537
|
+
nothing.
|
|
538
|
+
|
|
539
|
+
1. **Baseline correctness** against the consensus. urgency: decision at p >= 0.5 against the label,
|
|
540
|
+
plus the Brier score. intent: argmax, plus the multiclass Brier. frustration: rounded level, plus
|
|
541
|
+
the mean absolute error.
|
|
542
|
+
2. **Confidence bands.** The share of tickets in the noul `:unsure` band (0.2 < p < 0.8), and the
|
|
543
|
+
share of intents below 0.6 confidence. For each band, the agreement with the consensus. If
|
|
544
|
+
low-confidence answers are no less correct than the others, confidence tells us nothing here.
|
|
545
|
+
3. **Recalibration.** One temperature per question, fitted by likelihood on a random half (fixed
|
|
546
|
+
seed) and tested on the other. Criterion: it is adopted only if the Brier on the test half drops
|
|
547
|
+
by at least 10% against the identity. Otherwise Jev is declared calibrated on this dataset, to
|
|
548
|
+
the precision of 125 tickets.
|
|
549
|
+
|
|
550
|
+
### Arm 7b. Is the named-key shift a constant? PRE-REGISTERED on 2026-09-23
|
|
551
|
+
|
|
552
|
+
The arm 7 JSONL files keep only the decisions, not the probabilities. `named_1` is replayed once (250
|
|
553
|
+
requests, seed 7) keeping the raw values. On one half, we estimate the mean logit shift between
|
|
554
|
+
`named_1` and the baseline, per question (urgency: logit of p; frustration: difference in continuous
|
|
555
|
+
value). We subtract it on the other half.
|
|
556
|
+
|
|
557
|
+
**Fixed reading.** If the test half's agreement with the baseline climbs back to within 2 points of
|
|
558
|
+
the control, the shift is a correctable constant. Otherwise it depends on the ticket. Either way,
|
|
559
|
+
the annotator consensus says which of the two readings is closer.
|
|
560
|
+
|
|
561
|
+
### Arm 8. Structured criteria. PRE-REGISTERED on 2026-09-23
|
|
562
|
+
|
|
563
|
+
**Why.** `docs.typesafe.ai/primitives/advanced.md` documents Choice options as
|
|
564
|
+
`{what, not_for, examples}` objects and Noul `true` and `false` criteria as objects, "to sharpen the
|
|
565
|
+
boundary". No figures are published. Until now the gem converted every entry to a string. It now
|
|
566
|
+
accepts objects, without changing the digest of a string criterion.
|
|
567
|
+
|
|
568
|
+
**Shape.** One subject per request, raw text, three questions, concurrency 8. Only the urgency and
|
|
569
|
+
intent criteria change. The `what` repeats the original text word for word, and `not_for` and two
|
|
570
|
+
generic `examples` are added. No example is taken from the 250 tickets. frustration does not change:
|
|
571
|
+
it is the control question and must not move.
|
|
572
|
+
|
|
573
|
+
| Variant | Criteria | Requests |
|
|
574
|
+
|---|---|---|
|
|
575
|
+
| `current` | those in `ticket.rb` | 250 |
|
|
576
|
+
| `structured` | objects | 250 |
|
|
577
|
+
| `structured_replay` | objects, first 50 tickets replayed | 50 |
|
|
578
|
+
|
|
579
|
+
**Criterion, fixed before the run**, against the annotator consensus, `structured` against
|
|
580
|
+
`current` from the same day:
|
|
581
|
+
|
|
582
|
+
- adopted if, on urgency and on intent, agreement drops by no more than one point and rises by at
|
|
583
|
+
least 2 points on one of the two, or if the Brier drops by at least 10% with no loss of agreement;
|
|
584
|
+
- rejected otherwise. The extra token cost is reported, not weighed: at the confirmed price it is
|
|
585
|
+
negligible.
|
|
586
|
+
|
|
587
|
+
`structured_replay` gives its noise floor. frustration must stay within the noise of arm 1.
|
|
588
|
+
|
|
589
|
+
Command: `bin/rails judge:eval:criteria VARIANT=current|structured|structured_replay`.
|
|
590
|
+
Since the demo switched over, `current` is called `plain`: the model's criteria are structured, and
|
|
591
|
+
`plain` rebuilds the original strings, including the digests `9e08ebfeb6e0e0fa` and
|
|
592
|
+
`5f3c0bad775abc1c`.
|
|
593
|
+
|
|
594
|
+
### Arm 9. Sustained rate limit. PRE-REGISTERED on 2026-09-23
|
|
595
|
+
|
|
596
|
+
**Why.** `docs.typesafe.ai/models.md` documents 1,200 requests per minute and 250,000 tokens per
|
|
597
|
+
second. Arm 6 sent 100 requests per level and could not fill a minute. At 8 threads the pool runs at
|
|
598
|
+
~28 requests per second, above the documented limit.
|
|
599
|
+
|
|
600
|
+
**Shape.** The production shape, looping over the 250 tickets for 75 s, with a client that never
|
|
601
|
+
retries (`max_retries = 0`) so every raw 429 and its `Retry-After` are visible. Levels run in
|
|
602
|
+
sequence: 8, then 16, then 32 only if 16 saw no 429. Two minutes of pause between levels.
|
|
603
|
+
|
|
604
|
+
**Fixed reading.**
|
|
605
|
+
|
|
606
|
+
- 429s at 8: the limit applies and the gem's default exceeds it. We need a client-side limiter or a
|
|
607
|
+
`max_retry_wait` that covers the observed `Retry-After`.
|
|
608
|
+
- No 429 at 32 over 75 s: the documented limit is not enforced on this key today. The default of 8
|
|
609
|
+
stays, and the limit is written down in the README.
|
|
610
|
+
- In between: the observed threshold is reported, and the default does not exceed the last level
|
|
611
|
+
without a 429.
|
|
612
|
+
|
|
613
|
+
Budget: at most ~9,000 requests, ~4.8M tokens, ~$0.20 at the confirmed price.
|
|
614
|
+
|
|
615
|
+
Command: `bin/rails judge:eval:rate CONCURRENCY=8 DURATION=75`.
|
|
616
|
+
|
|
617
|
+
### Results of the 2026-09-23 arms (labels, 7b, 8, 9)
|
|
618
|
+
|
|
619
|
+
Everything was served by `jev-1.13.0`, checked on `ResultSet#model`. Total spend: 8,462,588 input
|
|
620
|
+
tokens, **~$0.36** at the confirmed price. Arm 9 cost 7.94M tokens against the 4.8M planned:
|
|
621
|
+
throughput doubled with each thread level, like everything else, and the budget underestimated it.
|
|
622
|
+
|
|
623
|
+
#### Labels
|
|
624
|
+
|
|
625
|
+
| Question | Inter-annotator agreement | kappa | Consensus kept |
|
|
626
|
+
|---|---|---|---|
|
|
627
|
+
| urgency | 96.4% | 0.88 | 241 |
|
|
628
|
+
| intent | 88.4% | 0.82 | 221 |
|
|
629
|
+
| frustration | 84.8% | 0.75 | 212 |
|
|
630
|
+
|
|
631
|
+
The main disagreement is telling: 15 thank-you messages are `technical` for one annotator and
|
|
632
|
+
`spam` for the other, because the spam definition includes "anything not a genuine support request".
|
|
633
|
+
The demo's criteria have no place for a message that is neither a request nor spam.
|
|
634
|
+
|
|
635
|
+
#### Analyses without requests
|
|
636
|
+
|
|
637
|
+
| Question | Baseline against consensus | Brier |
|
|
638
|
+
|---|---|---|
|
|
639
|
+
| urgency, decision at 0.5 | 82.6% | 0.1141 |
|
|
640
|
+
| intent | 84.2% | 0.2285 |
|
|
641
|
+
| frustration, rounded level | 60.8% | mean error 0.399 |
|
|
642
|
+
|
|
643
|
+
**The errors go one way.** Jev reads tickets as more urgent (41 false positives, 1 false negative at
|
|
644
|
+
0.5) and more frustrated (83 levels too high, 0 too low) than both annotators. For intent, it says
|
|
645
|
+
`billing` where the consensus says `technical` 18 times.
|
|
646
|
+
|
|
647
|
+
**Bands.** Confident urgency (p <= 0.2 or >= 0.8): 97.7% agreement on 131 tickets. `:unsure` band:
|
|
648
|
+
64.5% on 110. intent at confidence >= 0.6: 88.7% on 194; below 0.6: 51.9% on 27. Confidence sorts
|
|
649
|
+
the answers well, so routing the uncertain band makes sense here.
|
|
650
|
+
|
|
651
|
+
**Temperature.** urgency +0.2%, intent -0.2%, frustration -9.3% of Brier on the test half. None
|
|
652
|
+
crosses the 10% threshold: **no temperature recalibration.**
|
|
653
|
+
|
|
654
|
+
**Exploratory, not pre-registered, tested on a held-out half.** For frustration, subtracting 0.5
|
|
655
|
+
before rounding, which amounts to taking the integer part, raises agreement from 62.7% to 83.3%.
|
|
656
|
+
For urgency, the 0.5 threshold gives 80.8%, and the demo's 0.8 threshold gives 92.5%, on par with
|
|
657
|
+
the best learned threshold (0.75). The cost comes from how the demo reads the number: rounding and
|
|
658
|
+
the median threshold.
|
|
659
|
+
|
|
660
|
+
#### Arm 7b
|
|
661
|
+
|
|
662
|
+
Shift learned on 125 tickets: -0.164 in logit on urgency, -0.083 in value on frustration. The
|
|
663
|
+
per-ticket standard deviation is 0.188, the same order as the mean. The correction raises the test
|
|
664
|
+
half's agreement with the baseline from 115 to 121 out of 125 on urgency, and from 110 to 113 on
|
|
665
|
+
frustration. **The 2-point criterion is not met: the shift is not a constant.**
|
|
666
|
+
|
|
667
|
+
Against the consensus, however, `named_1` beats raw text on all three questions:
|
|
668
|
+
|
|
669
|
+
| | urgency | intent | frustration |
|
|
670
|
+
|---|---|---|---|
|
|
671
|
+
| raw text | 82.6% (Brier 0.114) | 84.2% (0.229) | 60.8% |
|
|
672
|
+
| `named_1` | 85.9% (0.099) | 86.4% (0.213) | 70.3% |
|
|
673
|
+
|
|
674
|
+
The "calmer, less urgent" shift from arm 7 moves **toward** the labels. Arm 7 measured distance from
|
|
675
|
+
raw text, and raw text is the reading furthest from the consensus. That still does not reopen
|
|
676
|
+
batching by subject: `named_20` was not re-read against the labels.
|
|
677
|
+
|
|
678
|
+
#### Arm 8
|
|
679
|
+
|
|
680
|
+
| Variant | urgency | intent | frustration | Tokens / ticket |
|
|
681
|
+
|---|---|---|---|---|
|
|
682
|
+
| `current` | 82.6% (Brier 0.1141) | 85.1% (0.2271) | 62.3% | 539 |
|
|
683
|
+
| `structured` | **86.7%** (0.0897) | **89.6%** (0.1608) | 61.8% | 825 |
|
|
684
|
+
|
|
685
|
+
Against the consensus. `current` reproduces the baseline (98.8 / 99.2 / 98.8%). `structured` against
|
|
686
|
+
its own replay on 50 tickets: 50/50, 50/50, 49/50, mean drift 0.008. frustration, the control
|
|
687
|
+
question, stays within the noise.
|
|
688
|
+
|
|
689
|
+
**Criterion met: adopted.** +4.1 and +4.5 points, Brier -21% and -29%. `billing` errors that should
|
|
690
|
+
have been `technical` drop from 18 to 12. Extra cost: +286 tokens per ticket, +53%, ~$0.000012.
|
|
691
|
+
|
|
692
|
+
A caveat noticed afterwards: at the demo's escalation threshold (0.8) rather than 0.5, urgency goes
|
|
693
|
+
from 91.3% to 90.0%, three tickets. The urgency gain rests on the Brier and the median threshold,
|
|
694
|
+
while the intent gain holds everywhere. The demo is not switched over: changing its criteria
|
|
695
|
+
invalidates the 250 committed judgments, and the reference is still a consensus of models.
|
|
696
|
+
|
|
697
|
+
#### Arm 9
|
|
698
|
+
|
|
699
|
+
| Threads | Requests | Throughput | Max over 60 s | 429 | Latency p50 / p95 |
|
|
700
|
+
|---|---|---|---|---|---|
|
|
701
|
+
| 8 | 2,187 | 29.2 /s | 1,758 | 0 | 0.268 / 0.336 s |
|
|
702
|
+
| 16 | 4,214 | 56.2 /s | 3,393 | 0 | 0.277 / 0.361 s |
|
|
703
|
+
| 32 | 8,355 | 111.4 /s | 6,696 | 0 | 0.279 / 0.371 s |
|
|
704
|
+
|
|
705
|
+
No 429s and no errors over 14,756 requests without retries. At 32 threads that is 5.6 times the
|
|
706
|
+
documented per-minute limit and ~60,000 tokens per second, a quarter of the token limit. Latency
|
|
707
|
+
does not move. **Fixed reading: the documented limit is not enforced on this key today.** The
|
|
708
|
+
default of 8 stays. The limit and this measurement are written in the gem's README, with the
|
|
709
|
+
provider's caveat ("can change without notice").
|
|
710
|
+
|
|
711
|
+
Raw data: `judge_rails_demo/tmp/bench/arm7bis_named_1.jsonl`, `arm8_*.jsonl`, `arm9_c*.jsonl`,
|
|
712
|
+
snapshotted in `wiki/raw/transcripts/labels-and-arms-8-9-2026-09-23.md`.
|
|
713
|
+
|
|
714
|
+
## 7. Acceptance criterion and verdict
|
|
715
|
+
|
|
716
|
+
### VERDICT, 2026-09-21: batching by subject fails, and the cause is structural.
|
|
717
|
+
|
|
718
|
+
The hard criterion required 100% agreement on the choice label, the noul band and the score level.
|
|
719
|
+
**It fails at every batch size, including 2.** Read relative to arm 1, as the section below
|
|
720
|
+
prescribes, the gap is still 9 to 15 points of agreement and the drift is 5 to 16 times the
|
|
721
|
+
control's noise floor.
|
|
722
|
+
|
|
723
|
+
Batching as implemented trades 10 to 15 points of decision accuracy for 49x the speed. For a gem
|
|
724
|
+
whose product is a calibrated probability, that is not an acceptable default trade. This protocol
|
|
725
|
+
was written before the result was known precisely so the result could not be rationalized
|
|
726
|
+
afterwards.
|
|
727
|
+
|
|
728
|
+
`Judge::Batch` stays in `lib/`, tested and not wired in. `Judge.config.batch_rows` stays `nil`. No
|
|
729
|
+
path in the gem batches.
|
|
730
|
+
|
|
731
|
+
### Next steps, decided by arm 5
|
|
732
|
+
|
|
733
|
+
**Corrected on 2026-09-23.** The "3 from shape, 9 from neighbors" split below is refuted by arm 7,
|
|
734
|
+
and arm 7b shows that the named shape is closer to the labels than raw text. The recommendation (the
|
|
735
|
+
thread pool) holds for other reasons, written up in arm 7.
|
|
736
|
+
|
|
737
|
+
Arm 5 ran and it closes the question: of the 12 points lost, 3 come from the shape and 9 from the
|
|
738
|
+
neighbors. Fixing the shape would still leave 9 points on the table, so no version of batching by
|
|
739
|
+
subject passes the hard criterion again.
|
|
740
|
+
|
|
741
|
+
**What to build instead, backed by the measurements:**
|
|
742
|
+
|
|
743
|
+
1. **A thread pool on `judge_filter`, with no batching.** Same request shape as today, so **zero
|
|
744
|
+
accuracy loss**, and it is the solution the Re-Ranking cookbook applies to 3,565 passages. Arm 4
|
|
745
|
+
measures 7.8x for 8 threads, and arm 5 confirms it on the real API: 150 tickets in 5.44 s at 8
|
|
746
|
+
threads, 0.036 s per ticket against 0.270 s in series.
|
|
747
|
+
2. **Do not wire in `Judge::Batch`.** It stays in `lib/`, tested, with this document as the written
|
|
748
|
+
reason.
|
|
749
|
+
3. **Move `[[typesafe-jev]]` to `high`** on flat latency as a function of the number of questions,
|
|
750
|
+
which is now measured.
|
|
751
|
+
|
|
752
|
+
### The criterion, as it was fixed
|
|
753
|
+
|
|
754
|
+
A double budget, fixed before the run.
|
|
755
|
+
|
|
756
|
+
**Hard: must pass, or batching does not ship at that batch size:**
|
|
757
|
+
|
|
758
|
+
- 100% agreement on the choice label
|
|
759
|
+
- 100% agreement on the noul's `judge_decide` band, at the demo's thresholds
|
|
760
|
+
- 100% agreement on the integer level of the score
|
|
761
|
+
|
|
762
|
+
**Soft: read as a curve, not as a threshold:**
|
|
763
|
+
|
|
764
|
+
- mean and p95 `|Δp|` per batch size and per question type
|
|
765
|
+
- total variation distance on the choice distribution
|
|
766
|
+
- continuous value difference on the score
|
|
767
|
+
|
|
768
|
+
Both budgets are read **relative to arm 1**, never in absolute terms. If the unbatched control
|
|
769
|
+
itself drifts by 0.03 in mean `|Δp|`, a batch at 0.03 has degraded nothing.
|
|
770
|
+
|
|
771
|
+
The default `batch_rows` would be the largest batch size that passes the hard budget and whose soft
|
|
772
|
+
drift stays at the control's level. Not 20 just because pg_jev says 20.
|
|
773
|
+
|
|
774
|
+
## 8. What still needs protecting
|
|
775
|
+
|
|
776
|
+
- The current baseline must be snapshotted in `wiki/raw/transcripts/` with its sha256 **before** any
|
|
777
|
+
regeneration of `db/seed_judgments.rb`, otherwise the baseline is lost.
|
|
778
|
+
- Arms 0 to 3 never run in `rake test`. The house rule is never to call the real API from a test.
|
|
779
|
+
They will be a manual task. Arm 4 lives in `bin/bench_batch`, outside the suite, and only its
|
|
780
|
+
reduced version is a test.
|
|
781
|
+
- Batching widens the blast radius of a prompt injection from 1 to `batch_rows`, since the rows share
|
|
782
|
+
a `state` and the vendor's documentation describes the model as steerable by injected
|
|
783
|
+
instructions. This protocol does not measure it. An adversarial arm would have to be written
|
|
784
|
+
separately.
|
|
785
|
+
|
|
786
|
+
## 8b. The optimal request recipe, as measured
|
|
787
|
+
|
|
788
|
+
Three rules, each backed by an arm.
|
|
789
|
+
|
|
790
|
+
| Rule | Arm | Gain | Accuracy cost |
|
|
791
|
+
|---|---|---|---|
|
|
792
|
+
| **One subject per request** | 2, 3, 5, 7 | - | leaving raw text costs ~6 points of agreement with the baseline, a neighbor ~0.3 (arm 7). Against the labels, the named shape does better (arm 7b) |
|
|
793
|
+
| **All of the subject's questions in the same request** | 0 | 326-token overhead amortized, flat latency up to 40 questions | **zero** |
|
|
794
|
+
| **Thread fan-out over subjects** | 6 | 6.95x at 8 threads, 18.4x at 32 | **zero** |
|
|
795
|
+
|
|
796
|
+
The gem now applies all three: `storage.rb:26-34` groups a record's questions into one call, and
|
|
797
|
+
`relation.rb` fans the records out over a pool.
|
|
798
|
+
|
|
799
|
+
## 9. Log
|
|
800
|
+
|
|
801
|
+
### 2026-09-21, phase A
|
|
802
|
+
|
|
803
|
+
Written: `lib/judge/batch.rb` (L2 engine, stdlib only), `Judge::PayloadTooLargeError` in the error
|
|
804
|
+
taxonomy, `batch_rows` and `concurrency` on the configuration at `nil`, `test/batch_test.rb` (22
|
|
805
|
+
tests), `test/batch_throughput_test.rb` (6 tests), `bin/bench_batch`.
|
|
806
|
+
|
|
807
|
+
Measured: 200 tests green against 172 before, 580 assertions, zero rubocop offenses, and the Rails
|
|
808
|
+
7.2, 8.0 and 8.1 gemfiles all green. Arm 4 ran, table above.
|
|
809
|
+
|
|
810
|
+
Two corrections to the protocol, made while running it:
|
|
811
|
+
|
|
812
|
+
1. Arm 4 cannot live in the test suite. At the measured latency of 0.300 s, 250 rows in sequence
|
|
813
|
+
take 76 seconds.
|
|
814
|
+
2. `rows:` carries no accuracy ceiling. The previous version of this plan set one at 25, which would
|
|
815
|
+
have stopped arms 2 and 3 from measuring a batch of 40. The measured curve is the guard, not a
|
|
816
|
+
guessed constant.
|
|
817
|
+
|
|
818
|
+
Naming decision: the questions are called `row0`, `row1` on the wire, in normalcase, because `row_0`
|
|
819
|
+
trips `Naming/VariableNumber` and the only alternatives were to disable the cop or to modify a
|
|
820
|
+
pre-existing test file that did not belong to this work.
|
|
821
|
+
|
|
822
|
+
### 2026-09-21, phase B
|
|
823
|
+
|
|
824
|
+
Written: `judge_rails/bin/arm0_latency`, `judge_rails_demo/lib/tasks/judge_bench.rake`,
|
|
825
|
+
`wiki/raw/transcripts/seed-judgments-baseline-2026-09-21.md` (sha256
|
|
826
|
+
`0b358b56125d79dfef0b42c351d63c5c603c99066263b78f287fe643457a5839`). `Judge::Batch.judge` extended to
|
|
827
|
+
accept a Hash of questions, which arm 3 required. 206 tests green, zero offenses.
|
|
828
|
+
|
|
829
|
+
Fixed along the way: the demo's `config/initializers/judge.rb` only read `JEV_API_KEY`, while
|
|
830
|
+
`Judge::Configuration` has always accepted `TYPESAFE_API_KEY` as a fallback. The demo now accepts
|
|
831
|
+
both.
|
|
832
|
+
|
|
833
|
+
Actual campaign spend: about 1,000 requests and ~1.40 million input tokens, so **~$0.06** at the
|
|
834
|
+
borrowed price of $0.042/M. The projected budget was $0.03 for 430 requests. The extra diagnosis at
|
|
835
|
+
a batch of 2 doubled the volume.
|
|
836
|
+
|
|
837
|
+
No database writes. `db/seed_judgments.rb` is intact, and the batched path never touched it.
|
|
838
|
+
|
|
839
|
+
### 2026-09-21, phase C revised
|
|
840
|
+
|
|
841
|
+
Batching by subject is not wired in and will not be. What was wired in instead:
|
|
842
|
+
|
|
843
|
+
- `lib/judge/pool.rb`, new. `Judge::Pool.map(items, concurrency:)` preserves input order, runs
|
|
844
|
+
concurrency 1 inline with no thread, and re-raises the first error after all workers have
|
|
845
|
+
stopped. `Judge::Batch` builds on it, so the threading code exists only once.
|
|
846
|
+
- `lib/judge/rails/relation.rb`. `judge_filter`, `judge_map` and `judge_sort` take `concurrency:` and
|
|
847
|
+
fan out over the pool. The `state` values are built on the calling thread before the fan-out, so
|
|
848
|
+
no worker touches an ActiveRecord connection.
|
|
849
|
+
- Default: explicit `concurrency:`, otherwise `Judge.config.concurrency`, otherwise 8.
|
|
850
|
+
|
|
851
|
+
219 tests green against 172 at the start, zero offenses, Rails 7.2, 8.0 and 8.1 green.
|
|
852
|
+
|
|
853
|
+
### 2026-09-23, labels and arms 7b, 8, 9
|
|
854
|
+
|
|
855
|
+
Written: `db/ticket_labels.json` (two model annotators, consensus), `lib/tasks/judge_eval.rake`
|
|
856
|
+
(`judge:eval:offline`, `named_values`, `criteria`, `rate`). In the gem, `Question::Choice` and
|
|
857
|
+
`Question::Noul` accept object or array entries. A string entry keeps its digest, checked against
|
|
858
|
+
the three baseline digests. 263 tests green on Rails 7.2, 8.0 and 8.1, zero offenses. Demo: 42 tests
|
|
859
|
+
green, zero offenses.
|
|
860
|
+
|
|
861
|
+
Fixed on re-reading: the gem's README showed `question.digest # => "9e08ebfeb6e0e0fa"` for a choice.
|
|
862
|
+
That is the digest of the demo's urgency question. The real value is `1c3be3cf24579b81`.
|
|
863
|
+
|
|
864
|
+
Pre-registered before any request, run afterwards. The key is read from the shell environment on
|
|
865
|
+
each command and never written into the project.
|
|
866
|
+
|
|
867
|
+
### 2026-09-23, the demo switches to structured criteria
|
|
868
|
+
|
|
869
|
+
Decided by the owner after arm 8. `app/models/ticket.rb` declares urgency and intent with the arm 8
|
|
870
|
+
criteria, word for word. The benches follow without changes, since they read the model's
|
|
871
|
+
definitions: `judge_bench.rake` (arms 1 to 7) and `judge_eval.rake`. The gem's two benches,
|
|
872
|
+
`bin/arm0_latency` and `bin/bench_batch`, declare the same shape.
|
|
873
|
+
|
|
874
|
+
**The baseline has changed.** `db/seed_judgments.rb` is regenerated: 250 requests, `jev-1.13.0` on
|
|
875
|
+
all 250, urgency digest `4f1911289ec33ca2` and intent digest `4234fcadf9f795b9`, frustration
|
|
876
|
+
unchanged at `d3f59e90f1ce9b41`. Against the structured replay from arm 8: 246/250, 250/250 and
|
|
877
|
+
243/250, within arm 1's noise. The old baseline is snapshotted in
|
|
878
|
+
`wiki/raw/transcripts/seed-judgments-plain-2026-09-23.md`. Any arm replayed from here on is read
|
|
879
|
+
against the structured baseline, not the one behind the figures above.
|
|
880
|
+
|
|
881
|
+
New baseline against the consensus: urgency 88.4% (Brier 0.090), intent 89.6% (0.162), frustration
|
|
882
|
+
62.3%. Confident bands: urgency 99.2% on 133, intent 93.9% on 196. No temperature crosses 10% yet
|
|
883
|
+
(urgency -6.7%, frustration -9.5%).
|
|
884
|
+
|
|
885
|
+
`db:seed` on an empty database restores 250 judgments, 0 stale. Demo: 42 tests green, zero offenses.
|