judge_rails 0.0.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
data/BENCHMARK.md ADDED
@@ -0,0 +1,885 @@
1
+ # BENCHMARK
2
+
3
+ Pre-registered protocol for the `Judge::Batch` batching engine, written before anything ran.
4
+
5
+ **Status on 2026-09-21: the campaign is finished and all five arms have run. The verdict is in
6
+ section 7: batching fails the criterion and is not wired in.** Every figure in this document is
7
+ either measured or arithmetic that is labeled as such.
8
+
9
+ **Rename note.** The gem was called `jev-in-rails` when these measurements were taken, and has been
10
+ `judge_rails` since 2026-09-23. Paths and identifiers in this document were updated so it stays
11
+ reproducible. The measurements themselves are unchanged. `wiki/queries/why-the-gem-was-renamed.md`
12
+ explains why.
13
+
14
+ The protocol comes first because the house rule is "verify, do not assert", and a protocol fixed
15
+ after seeing the results proves nothing.
16
+
17
+ ## 1. The question
18
+
19
+ `judge_filter` makes one network call per row, in series. pg_jev judges 2000 rows in 3.5 seconds by
20
+ packing 20 subjects into one `state` with one question per subject, and keeping 16 requests in
21
+ flight. Can batching carry over here without degrading the judgment, and by how much?
22
+
23
+ The API has no "subjects" dimension. It takes `{state, model, questions}`. Batching therefore means
24
+ encoding the subjects inside the `state` and asking one anchored question per subject. For the
25
+ model that is a different input regime. The transport stays the same.
26
+
27
+ ## 2. Measured offline, free, reproducible
28
+
29
+ Everything below is read from a file in the workspace, with no network call.
30
+
31
+ ### 2.1 The baseline
32
+
33
+ `judge_rails_demo/db/seed_judgments.rb`, 250 judgments produced by the unbatched path against the
34
+ real model.
35
+
36
+ | Quantity | Value |
37
+ |---|---|
38
+ | Judgments | 250 |
39
+ | Model | `jev-1.13.0`, only one |
40
+ | Distinct question digests | 3 (`9e08ebfeb6e0e0fa`, `5f3c0bad775abc1c`, `d3f59e90f1ce9b41`) |
41
+ | Mean latency | 0.309 s |
42
+ | Latency p50 / min / max | 0.286 s / 0.228 s / 0.773 s |
43
+
44
+ Each row carries the noul probability, the full `probabilities` distribution of the choice, and the
45
+ value, level and distribution of the score. Agreement is measured against this baseline.
46
+
47
+ ### 2.2 The inputs
48
+
49
+ `judge_rails_demo/db/seed_tickets.rb`, where the `state` is `[subject, body]` joined by two newlines
50
+ (`app/models/ticket.rb:14`).
51
+
52
+ | Quantity | Value |
53
+ |---|---|
54
+ | Tickets | 250 |
55
+ | `state` in characters | mean 208, p50 156, max 851 |
56
+ | `state` total | 52,114 characters |
57
+ | The 3 questions + criteria | ~620 characters |
58
+ | Rows that fit in 24,000 characters | ~115 |
59
+
60
+ On this dataset the accuracy ceiling (20 to 25 rows) bites about five times earlier than the `state`
61
+ size ceiling. The demo will never reach the character budget.
62
+
63
+ ### 2.3 The current paths, read from the code
64
+
65
+ | Path | File | Requests | Parallelism |
66
+ |---|---|---|---|
67
+ | `judge_filter` / `map` / `sort` | `lib/judge/rails/relation.rb:41-44` | 1 per row | **1 thread** |
68
+ | backfill by job | `lib/judge/rails/jobs.rb:56-61` | 1 per record | depends on the queue |
69
+ | single computation | `lib/judge/rails/storage.rb:26-34` | 1 per record | n/a |
70
+ | `judge:demo:precompute` | `judge_rails_demo/lib/tasks/judge_demo.rake:8-30` | 1 per ticket | **8 threads** |
71
+
72
+ None of these paths puts two subjects in one request. `Jobs.batch` (`jobs.rb:39-48`) groups ids to
73
+ reduce the number of jobs. The number of requests stays the same.
74
+
75
+ ## 3. Projected: arithmetic, not measured
76
+
77
+ Labeled separately because nothing here was observed.
78
+
79
+ ### 3.1 Requests to judge 250 tickets on 3 questions
80
+
81
+ | Shape | Requests | Calculation |
82
+ |---|---|---|
83
+ | Current | 250 | 1 per ticket, 3 questions in the request |
84
+ | A, 1 question × 20 rows | 39 | 3 definitions × ceil(250/20) |
85
+ | B, 3 questions × 20 rows | 13 | ceil(250/20) |
86
+
87
+ ### 3.2 Tokens per row, estimated at 4 characters per token
88
+
89
+ | Shape | Tokens / row | Detail |
90
+ |---|---|---|
91
+ | A | ~400 | the ticket `state` is sent 3 times, once per definition |
92
+ | B | ~296 | the `state` is sent once, with 3 questions per row |
93
+
94
+ Estimated difference: **B is about 26% cheaper than A**, far from a factor of 3. With short rows the
95
+ per-row instructions dominate the cost, more than the `state` does.
96
+
97
+ For reference, pg_jev documents ~175 input tokens per row in batches of 20, against ~435 for a single
98
+ row. Our questions carry richer criteria than its conditions, hence the higher figure.
99
+
100
+ ### 3.3 Campaign cost
101
+
102
+ ~430 requests, ~710,000 input tokens, so **~$0.03** at the $0.042 per million input tokens quoted in
103
+ the pg_jev documentation. **That price is borrowed and was never checked with Jev.** The real cost
104
+ will show up in `ResultSet#usage`, which the client already fills in (`result_set.rb:34`).
105
+
106
+ ## 4. What the protocol cannot establish
107
+
108
+ - **Absolute calibration.** We measure agreement with the unbatched path, with no human labels to
109
+ measure correctness against. This workspace has no human-labeled dataset.
110
+ - **Generalization beyond this dataset.** 250 short, synthetic support tickets in English. Long
111
+ documents and multilingual data are not covered.
112
+ - **Behavior near the 32k limit.** No input in the dataset comes close.
113
+
114
+ ## 5. Design decisions that define the experiment
115
+
116
+ Taken before the run, so that the arms of the protocol mean something.
117
+
118
+ | # | Decision |
119
+ |---|---|
120
+ | D1 | Batching is opt-in. `Judge.config.batch_rows` and `Judge.config.concurrency`, `nil` by default |
121
+ | D2 | Only the bulk paths read the setting. The sync `before_save` cannot be batched, because a `before_save` has to finish before its own save. The validator is excluded because it fails closed |
122
+ | D3 | `state` = strict JSON `{condition, rows:[{id, text}]}`. One question per row, named `row0`, `row1`, with instructions prefixed by an anchor that names the id and says to treat the text as data |
123
+ | D4 | The batch path never substitutes the anchored question into the `Definition`. The sidecar carries the digest of the original question, otherwise `stale?` loops |
124
+ | D5 | Incomplete response: bisection. The batch is split in two and retried, depth capped at 5, and the one-row base case goes through the single path |
125
+ | D6 | The character budget is a guard that raises with the measured size in the message. Packing counts rows, not characters |
126
+ | D7 | This document decides the chosen shape (A or B) and the value of `batch_rows`. Nothing before it does |
127
+
128
+ ## 6. Protocol
129
+
130
+ ### Arm 0. Latency as a function of the number of questions. RUN on 2026-09-21
131
+
132
+ **This arm decides whether the rest is worth anything.** The vendor claims that response time barely
133
+ moves when you add questions. The wiki carried that claim at `confidence: medium`, never checked
134
+ here.
135
+
136
+ `judge_rails/bin/arm0_latency`. One 280-character state, the same question repeated N times, 3
137
+ repetitions per point, 18 requests against the real model.
138
+
139
+ | N questions | Median latency | Spread | Input tokens | Tokens / question | vs N=1 |
140
+ |---|---|---|---|---|---|
141
+ | 1 | 0.280 s | 0.428 s | 350 | 350 | 1.00x |
142
+ | 2 | 0.264 s | 0.032 s | 373 | 186 | 0.94x |
143
+ | 5 | 0.242 s | 0.031 s | 442 | 88 | 0.86x |
144
+ | 10 | 0.234 s | 0.014 s | 557 | 56 | 0.84x |
145
+ | 20 | 0.242 s | 0.020 s | 797 | 40 | 0.86x |
146
+ | 40 | 0.246 s | 0.024 s | 1,277 | 32 | 0.88x |
147
+
148
+ **Latency is flat.** 40 questions in one request take the same time as one, at 0.88x. The go/no-go
149
+ criterion asked for less than 10x at N=20, and we measured 0.86x. The 0.428 s spread at N=1 is the
150
+ first request on a cold connection, and it does not recur at any other point.
151
+
152
+ The vendor's claim is **verified**, and the wiki can move `[[typesafe-jev]]` from `medium` to `high`
153
+ on this specific point.
154
+
155
+ **Cost structure, measured.** The fixed overhead of a request is about **326 tokens** (350 at N=1
156
+ minus the marginal cost of one question), and a marginal question costs
157
+ **(1277 - 350) / 39 ≈ 24 tokens**. Batching amortizes this fixed overhead. pg_jev documents the same
158
+ mechanism without putting a number on it for us.
159
+
160
+ Arithmetic consequence for the demo: judging 250 tickets without batching pays the overhead 250
161
+ times, ~81,500 tokens of pure structure. In batches of 20 it pays 13 times, ~4,200. **~77,000 tokens
162
+ saved on structure alone.**
163
+
164
+ **This arm refutes arm 4's conclusion on cost.** The byte proxy predicted +24.6% for batching. Real
165
+ billing says 8.75x fewer tokens per question at N=20. Bytes on the wire are not the bill: the ~326
166
+ token overhead has no visible counterpart in the payload. The proxy is dropped as a cost instrument
167
+ and `ResultSet#usage` replaces it.
168
+
169
+ **Caveat.** This arm keeps the state fixed and varies N. In a real batch the state grows with the
170
+ rows. The arm shows that latency does not depend on N and that the per-request overhead is real. The
171
+ total cost of a batch is an output of arms 2 and 3.
172
+
173
+ ### Arm 1. Control: model determinism. RUN on 2026-09-21
174
+
175
+ `bin/rails judge:bench:control SAMPLE=50`. The unbatched path replayed on 50 tickets and compared to
176
+ the baseline. This is the noise floor, and no other arm can be read without it.
177
+
178
+ | Question | n | Decision agreement | Mean \|Δ\| | p95 \|Δ\| |
179
+ |---|---|---|---|---|
180
+ | urgency (noul) | 50 | **100.0%** | 0.0118 | 0.0300 |
181
+ | intent (choice) | 50 | **100.0%** | 0.0082 | 0.0400 |
182
+ | frustration (score) | 50 | **96.0%** | 0.0108 | 0.0400 |
183
+
184
+ 50 requests, 26,492 input tokens, 13.52 s.
185
+
186
+ The model is close to deterministic. Mean drift is ~0.01 and decisions are stable, except for one
187
+ score level in 25 that flips. **The noise floor is therefore ~0.01 drift and 96 to 100% agreement.**
188
+ Everything that follows is read against those three numbers.
189
+
190
+ ### Arms 2 and 3. Shapes A and B, four batch sizes. RUN on 2026-09-21
191
+
192
+ `bin/rails judge:bench:campaign SHAPE=a|b BATCH=n`, 250 tickets, compared to the baseline.
193
+
194
+ **Decision agreement, in percent. Control on the first row.**
195
+
196
+ | Batch | Shape | urgency | intent | frustration | Requests | Tokens |
197
+ |---|---|---|---|---|---|---|
198
+ | 1 | control | **100.0** | **100.0** | **96.0** | 250* | 132,460* |
199
+ | 2 | A | 91.6 | 90.4 | 90.0 | 375 | 228,021 |
200
+ | 2 | B | 93.2 | 89.2 | 89.2 | 125 | 129,882 |
201
+ | 5 | A | 91.2 | 89.6 | 84.4 | 150 | 165,999 |
202
+ | 5 | B | 91.2 | 90.4 | 86.4 | 50 | 110,083 |
203
+ | 10 | A | 88.0 | 87.6 | 83.6 | 75 | 145,324 |
204
+ | 10 | B | 88.4 | 88.0 | 85.6 | 25 | 103,483 |
205
+ | 20 | A | 86.4 | 87.6 | 84.4 | 39 | 136,120 |
206
+ | 20 | B | 85.6 | 89.6 | 86.8 | 13 | 100,795 |
207
+ | 40 | A | 83.2 | 87.6 | 82.4 | 21 | 131,518 |
208
+ | 40 | B | 85.2 | 86.8 | 84.4 | 7 | 99,451 |
209
+
210
+ \* control extrapolated from 50 to 250 tickets for the comparison.
211
+
212
+ **Mean \|Δ\| drift, control at ~0.01.**
213
+
214
+ | Batch | Shape | urgency | intent | frustration |
215
+ |---|---|---|---|---|
216
+ | 1 | control | 0.0118 | 0.0082 | 0.0108 |
217
+ | 2 | B | 0.0532 | 0.0804 | 0.1110 |
218
+ | 5 | B | 0.0730 | 0.0937 | 0.1394 |
219
+ | 10 | B | 0.0816 | 0.1042 | 0.1543 |
220
+ | 20 | B | 0.0819 | 0.1097 | 0.1587 |
221
+ | 40 | B | 0.0874 | 0.1146 | 0.1780 |
222
+
223
+ ### Arm 5. 2x2 isolation, one row per request. RUN on 2026-09-21
224
+
225
+ `bin/rails judge:bench:isolate VARIANT=... SAMPLE=150`. With one row per request there is no
226
+ neighbor, so this square separates the two other causes: the JSON envelope around the `state`, and
227
+ the anchor prefix on the instructions.
228
+
229
+ | Variant | `state` | instructions | urgency | intent | frustration |
230
+ |---|---|---|---|---|---|
231
+ | `raw_plain` | raw text | original | **98.7** | **100.0** | **98.7** |
232
+ | `json_plain` | JSON `{rows:[...]}` | original | 96.0 | 98.0 | 95.3 |
233
+ | `raw_anchored` | raw text | prefixed | 98.0 | 94.0 | 95.3 |
234
+ | `json_anchored` | JSON | prefixed | 96.0 | 96.7 | 96.0 |
235
+
236
+ 600 requests, ~351,000 tokens, ~22 s in total.
237
+
238
+ `raw_plain` reproduces the control, which validates the harness. Each of the two transformations
239
+ costs about 2 to 4 points on its own, and the two together cost about **3 points**.
240
+
241
+ ### Breaking down the gap
242
+
243
+ The batch of 20 lost 10 to 14 points against the control. Arm 5 says where they come from:
244
+
245
+ | Cause | Points lost | Avoidable? |
246
+ |---|---|---|
247
+ | JSON envelope + anchor prefix | **~3** | maybe, by changing the shape |
248
+ | Neighbors present in the `state` | **~9** | no, that is batching itself |
249
+
250
+ The dominant cause is the neighbors, more than the JSON or the prefix.
251
+
252
+ **Corrected on 2026-09-23 by arm 7.** This split is a residual, not a measurement, and arm 7 refutes
253
+ it at a batch of 2: there the neighbor's content costs -0.3 points, the object shape -5.6 and a
254
+ second key -3.2. See arm 7.
255
+
256
+ ### What the TypeSafe documentation says, and how it explains the result
257
+
258
+ The [Parallel Questions](https://docs.typesafe.ai/cookbooks/parallel_questions.md) cookbook measures
259
+ 13 batched questions against 13 separate questions on a 54,000-character article: **12.2x cheaper,
260
+ 10.0x faster, no change in answers**, with a standard deviation of exactly 0.0 on most questions. It
261
+ gives the reason:
262
+
263
+ > "each question is scored on its own against the document, so its answer doesn't depend on what
264
+ > else is in the request."
265
+
266
+ A question does not depend on the other questions. It depends entirely on the **document**. Putting
267
+ 20 tickets in the `state` leaves the questions alone and changes the document each one is scored
268
+ against. Every ticket is judged in the middle of 19 texts that have nothing to do with it.
269
+
270
+ The [Re-Ranking](https://docs.typesafe.ai/cookbooks/rerank_typesafe.md) cookbook is the same problem
271
+ as ours, 3,565 passages to judge independently, and TypeSafe solves it with a
272
+ `{query_excerpt, candidate_passage}` state, **one candidate per request**, and **1,200 concurrent
273
+ calls in a thread pool**, 1,536,002 input tokens for $0.0645. Batching by subject does not appear.
274
+
275
+ **Corrected on 2026-09-23.** The cookbook does not say "1,200 concurrent calls". It runs 40 queries ×
276
+ 30 candidates preselected by BM25 out of the 3,565 passages, 1,200 calls in total, in a
277
+ `ThreadPoolExecutor(max_workers=12)`. The lesson we kept (one candidate per request) still holds. The
278
+ concurrency figure never existed. Re-read in `wiki/raw/transcripts/vendor-docs-check-2026-09-23.md`.
279
+
280
+ **Conclusion.** Batching by question is the documented route and the gem already follows it: three
281
+ `judge_attribute` on a ticket go out in one request (`storage.rb:26-34`). Batching by subject is not
282
+ a documented route, and the measurement shows why.
283
+
284
+ ### The diagnosis behind the curve
285
+
286
+ The degradation is **already complete at a batch of 2** and moves little up to 40. Going from 1 to 2
287
+ rows costs about 8 points of agreement. Going from 2 to 40 costs only 5 more.
288
+
289
+ **The dominant cause is the change of shape, more than the number of neighbors.** Two things change
290
+ at once between the control and a batch of 2:
291
+
292
+ 1. the `state` goes from raw text to a JSON object `{condition, rows:[{id, text}]}`,
293
+ 2. the question instructions are rewritten with an anchor prefix.
294
+
295
+ The protocol does not separate these two causes. That is the next test to run, and it is cheap.
296
+
297
+ ### Shape A against shape B
298
+
299
+ Agreement is the same for both shapes, within the noise. B is strictly cheaper: at a batch of 20, 13
300
+ requests against 39 and 100,795 tokens against 136,120, so **26% fewer tokens for the same
301
+ accuracy**. If batching were adopted, it would take shape B. That settles A against B, though it
302
+ does not make batching acceptable.
303
+
304
+ ### What batching would buy, and at what price
305
+
306
+ | Axis | Control | Shape B, batch 20 | Effect |
307
+ |---|---|---|---|
308
+ | Wall time per ticket | 0.270 s | 0.0055 s | **49x faster** |
309
+ | Tokens per ticket | 530 | 403 | **-24%** |
310
+ | urgency agreement | 100.0% | 85.6% | **-14.4 points** |
311
+ | intent agreement | 100.0% | 89.6% | **-10.4 points** |
312
+ | frustration agreement | 96.0% | 86.8% | **-9.2 points** |
313
+
314
+ The speed gain is real and large. The token saving is real and modest, far from the 8.75x that arm 0
315
+ suggested: once the rows are in the `state`, the `state` dominates, and the amortized fixed overhead
316
+ no longer weighs much.
317
+
318
+ ### Arm 4. Offline throughput, no API. RUN on 2026-09-21
319
+
320
+ `judge_rails/bin/bench_batch`, against `FakeJev` with **0.300 s** of injected latency, the mean
321
+ measured in 2.1. 250 rows of a size comparable to the demo's. It measures the thread pool and the
322
+ packing. The model is not involved.
323
+
324
+ | Configuration | Requests | Wall time | Bytes sent | vs sequential 1 thread |
325
+ |---|---|---|---|---|
326
+ | sequential, 1 thread | 250 | 76.11 s | 127,480 | 1.0x |
327
+ | sequential, 8 threads | 250 | 9.75 s | 127,480 | 7.8x |
328
+ | batch 20, 8 threads | 13 | 0.62 s | 158,858 | **122.8x** |
329
+ | batch 20, 16 threads | 13 | 0.32 s | 158,858 | **241.4x** |
330
+
331
+ Reproducible: `cd judge_rails && bundle exec ruby bin/bench_batch`. `ROWS` and `LATENCY` can be
332
+ overridden.
333
+
334
+ **A caveat that governs the whole table.** `FakeJev` answers in a fixed time whatever the number of
335
+ questions in the request. The 122x and 241x figures therefore assume, **by construction**, that
336
+ latency does not grow with the number of questions. Arm 0 exists to test exactly that assumption.
337
+ Until arm 0 has run, these two figures are a theoretical ceiling, not a measurement of the real gain.
338
+
339
+ **A result that contradicts a starting assumption.** Batching sends **more** bytes: 158,858 against
340
+ 127,480, **+24.6%**, or 635 bytes per row against 510. There are two causes, both measured here: the
341
+ anchor prefix is sent once per row instead of once per request, and the JSON wrapping of each row
342
+ (`{"id":n,"text":...}`) adds bytes the single path does not have.
343
+
344
+ So **the offline bench cannot settle the cost axis**, contrary to what section 3.2 projected. The
345
+ token saving pg_jev documents (175 against 435 per row) comes from amortizing a per-request overhead
346
+ of about 270 tokens, which is billed on the server and does not appear in the payload. Only the
347
+ `ResultSet#usage` returned by the real API can decide, which makes the token count an output of arms
348
+ 1 to 3 and not of arm 4.
349
+
350
+ ### Arm 4b, in the test suite
351
+
352
+ `test/batch_throughput_test.rb`, 6 tests, 0.1 s of injected latency, small volumes. It checks in
353
+ under a second that the pool really overlaps requests and that batching reduces the request count.
354
+ The faithful version of arm 4 takes 76 seconds, so it has no place in `rake test`.
355
+
356
+ ### Arm 6. Concurrency sweep on the real API. RUN on 2026-09-21
357
+
358
+ `bin/rails judge:bench:isolate VARIANT=raw_plain SAMPLE=100 CONCURRENCY=n`. The exact production
359
+ shape: `state` in raw text, 3 questions per ticket, one ticket per request. Only the thread count
360
+ changes.
361
+
362
+ | Threads | Wall time | Speedup | Efficiency | urgency | intent | frustration | Tokens |
363
+ |---|---|---|---|---|---|---|---|
364
+ | 1 | 25.01 s | 1.0x | 100% | 99.0 | 99.0 | 98.0 | 52,859 |
365
+ | 2 | 13.02 s | 1.92x | 96% | 100.0 | 99.0 | 100.0 | 52,859 |
366
+ | 4 | 6.87 s | 3.64x | 91% | 99.0 | 99.0 | 99.0 | 52,859 |
367
+ | 8 | 3.60 s | **6.95x** | 87% | 99.0 | 99.0 | 100.0 | 52,859 |
368
+ | 16 | 2.13 s | **11.7x** | 73% | 100.0 | 99.0 | 99.0 | 52,859 |
369
+ | 32 | 1.36 s | **18.4x** | 57% | 99.0 | 99.0 | 98.0 | 52,859 |
370
+
371
+ What the table shows:
372
+
373
+ 1. **No ceiling observed up to 32.** No errors, no 429s, and wall time keeps dropping. Efficiency
374
+ degrades (57% at 32) but the speedup stays monotonic.
375
+ **Caveat added on 2026-09-23:** `models.md` documents 1,200 requests per minute. 100 requests per
376
+ level never fill a minute, so this arm could not see the limit. See arm 9.
377
+ 2. **Accuracy does not move.** 98 to 100% at every level, the same as in series. That is expected,
378
+ and it is the basic difference from batching: the request is **byte for byte the same**, and only
379
+ its scheduling changes.
380
+ 3. **Tokens are identical to the unit**, 52,859 at every level. Concurrency costs nothing.
381
+
382
+ **Default chosen: 8.** 6.95x at 87% efficiency, and 8 connections per process is reasonable for a
383
+ gem that runs inside an application server. The sweep ran on one machine, with one key and 100
384
+ requests. It says nothing about the quota of a shared key under real load. `Judge.config.concurrency`
385
+ and the `concurrency:` keyword exist for anyone whose own measurement says otherwise.
386
+
387
+ ### Arm 7. Object `state` with named keys. PRE-REGISTERED, then RUN on 2026-09-23
388
+
389
+ **Why this arm.** Re-read on 2026-09-23, arm 5 does not measure what the "breaking down the gap"
390
+ section attributes to it. The ~9 points "from the neighbors" are a residual, not a measurement.
391
+ Between one JSON row and a batch of 20, several things change at once: the presence of foreign text
392
+ (dilution), designating the subject by a numeric `id` in an array (addressing), the number of
393
+ questions and the length of the `state`. The `raw_anchored` cell is inconsistent: it mentions a
394
+ "rows array" that a text `state` does not contain. The sample for arms 1 and 5 (`first(n)` by id)
395
+ contains no spam ticket.
396
+
397
+ The TypeSafe documentation (`concepts/state.md`, `primitives/advanced.md`) describes an **object**
398
+ `state` where each part has a name, and instructions that name the part they target. A manual test
399
+ in the playground, `{"State_1": "Shut up !", "State_2": "Shut up ?"}` with "Is State_1 a question ?",
400
+ returns 2% and 98%: the API accepts the shape and addressing holds on a minimal pair. This arm
401
+ measures whether addressing by named key recovers the accuracy that `rows[id]` lost.
402
+
403
+ **Shape.** `state` = a real JSON object (not a string), keys `ticket_1` to `ticket_N`. One question
404
+ per key and per definition, named `ticket_k_<name>`, whose original instruction is rewritten to name
405
+ the key ("Does ticket_1 need a human...", "Which team should handle ticket_1?", "How frustrated does
406
+ the customer in ticket_1 sound?"). No anchor prefix. Criteria unchanged.
407
+
408
+ **Variants, all on the 250 tickets**, compared to the `db/seed_judgments.rb` baseline:
409
+
410
+ | Variant | `state` | Requests | Isolates |
411
+ |---|---|---|---|
412
+ | `control` | raw text, original questions | 250 | noise floor on the same population |
413
+ | `named_1` | `{ticket_1}` | 250 | cost of the object shape alone |
414
+ | `named_same` | `{ticket_1, ticket_2}`, the same ticket twice | 250 | cost of a second key, with no foreign content |
415
+ | `named_2` | two different tickets | 125 | addressing + dilution, where the gap appeared |
416
+ | `named_20` | twenty different tickets | 13 | comparable to shape B at batch 20 |
417
+
418
+ Tickets are shuffled with a fixed seed (`SEED=7`) before grouping, so neighbors do not share a
419
+ category by construction. Concurrency 8. Each judged row is written as JSONL to `tmp/bench/`, with
420
+ the model that served it (`ResultSet#model`) and the neighbor ids.
421
+
422
+ **Criterion, fixed before the run.** A variant passes if, on each of the three questions, its
423
+ agreement is at least that of `control` minus 2 points (the range observed in arm 6 over six
424
+ repetitions of the production shape).
425
+
426
+ **Reading, fixed before the run.**
427
+
428
+ - `named_1` passes: the object shape is free.
429
+ - `named_same` passes and `named_2` fails: the cause is foreign content, and dilution is confirmed.
430
+ - `named_same` fails: a second key has a cost of its own, whatever its content.
431
+ - `named_2` and `named_20` pass: the 9 points of arms 2 and 3 came from the `rows[id]` shape, and
432
+ batching by subject is back on the table.
433
+ - Contamination: among the disagreements in `named_2`, the share where the answer equals the
434
+ neighbor's baseline answer, compared to chance.
435
+
436
+ Budget: 888 requests, ~500,000 input tokens, ~$0.02 at the borrowed price.
437
+
438
+ Command: `bin/rails judge:bench:named VARIANT=control|named_1|named_same|named_2|named_20`.
439
+ `DRY_RUN=1` prints the first request and calls nothing.
440
+
441
+ #### Results, 2026-09-23
442
+
443
+ 888 requests, 668,120 input tokens, no errors. All 1,250 judgments were served by `jev-1.13.0`, the
444
+ baseline model, checked on `ResultSet#model` rather than assumed.
445
+
446
+ | Variant | urgency | intent | frustration | Mean | Requests | Tokens | Wall time |
447
+ |---|---|---|---|---|---|---|---|
448
+ | `control` | **99.2** | **98.8** | **97.6** | 98.5 | 250 | 134,664 | 9.55 s |
449
+ | `named_1` | 93.2 | 96.0 | 89.6 | 92.9 | 250 | 138,164 | 9.07 s |
450
+ | `named_same` | 92.8 | 94.0 | 82.4 | 89.7 | 250 | 211,828 | 9.85 s |
451
+ | `named_2` | 91.6 | 90.8 | 86.0 | 89.5 | 125 | 105,914 | 4.96 s |
452
+ | `named_20` | 84.4 | 84.0 | 85.2 | 84.5 | 13 | 77,550 | 0.80 s |
453
+
454
+ Pass threshold: 97.2 / 96.8 / 95.6. **No named variant passes, not even `named_1`.** `control`
455
+ reproduces arm 1 on the whole population, spam included, which validates the harness.
456
+
457
+ **Direction of the disagreements**, band and level moving up or down relative to the baseline:
458
+
459
+ | Variant | urgency up / down | frustration up / down | Dominant intent flip |
460
+ |---|---|---|---|
461
+ | `control` | 0 / 2 | 3 / 3 | none, 3 isolated flips |
462
+ | `named_1` | 3 / 14 | 1 / 25 | none |
463
+ | `named_same` | 2 / 16 | 1 / 43 | sales to spam x4 |
464
+ | `named_2` | 8 / 13 | 4 / 31 | billing to technical x9 |
465
+ | `named_20` | 3 / 36 | 10 / 27 | billing to technical x18 |
466
+
467
+ In the control the disagreements are symmetric, which is noise. In every named variant they go one
468
+ way: **a ticket placed as a named field of an object reads calmer and less urgent.** That is a
469
+ systematic calibration shift, with no sign of random scatter.
470
+
471
+ **Contamination in `named_2`**, disagreements equal to the neighbor's baseline against chance:
472
+ urgency 7 against 8.0, intent 2 against 6.1, frustration 7 against 10.1. **Wrong answers do not copy
473
+ the neighbor.** Agreement by position in `named_2`: 89.3% for `ticket_1`, 89.6% for `ticket_2`, no
474
+ position effect.
475
+
476
+ #### Reading, against the grid fixed before the run
477
+
478
+ | Step | Mean | Cost | What changes |
479
+ |---|---|---|---|
480
+ | `control` | 98.5 | - | - |
481
+ | `named_1` | 92.9 | **-5.6** | the shape: a one-key object, an instruction that names the key |
482
+ | `named_same` | 89.7 | **-3.2** | a second key, with no new content at all |
483
+ | `named_2` | 89.5 | **-0.3** | the neighbor's content |
484
+ | `named_20` | 84.5 | -5.0 | 18 more tickets |
485
+
486
+ 1. **`named_1` fails: the object shape is not free.** It is the biggest measured step, and it is
487
+ there with no neighbor at all.
488
+ 2. **`named_same` fails: a second key has a cost of its own**, even though it only carries the same
489
+ text. Frustration drops the most here, 43 of 44 disagreements.
490
+ 3. **Foreign content costs almost nothing at a batch of 2**: -0.3 points between `named_same` and
491
+ `named_2`, and contamination at or below chance. **"Dilution", in the sense of a neighbor bleeding
492
+ into the answer, is not observed.**
493
+ 4. **`named_20` fails** at the level of shape B at batch 20 (85.6 / 89.6 / 86.8). Batching by subject
494
+ stays closed, whatever the addressing scheme.
495
+
496
+ **What this corrects.** The explanation written after arm 5, "3 points of shape, 9 points of
497
+ neighbors", has the wrong split. At a batch of 2, nearly all of the gap comes from leaving the
498
+ production shape (-5.6) and from the presence of a second key (-3.2). What the neighbor says barely
499
+ matters. At a batch of 20, the extra 5 points remain unattributed: `state` length, number of
500
+ questions and foreign content all change together.
501
+
502
+ **What this arm cannot say.** Agreement is measured against a baseline produced from raw text. A
503
+ shift toward "calmer" is a disagreement with the production shape, and not necessarily an error:
504
+ without human labels we cannot tell which of the two readings is more accurate. We do know they are
505
+ not interchangeable, and that the demo's thresholds (0.8 and 0.2) are tuned on the text shape.
506
+ `named_1` also conflates two changes, the object and the rewritten instruction ("this message"
507
+ becomes "ticket_1"), and this arm does not separate them. There is one run per variant. At n=250 one
508
+ ticket is worth 0.4 points, and the differences we keep are 8 to 40 tickets.
509
+
510
+ **Decision unchanged.** One subject per request, raw text, original questions. The thread pool is
511
+ still the way forward. Arm 7 does not reopen batching. It replaces the reason for rejecting it.
512
+
513
+ Raw data: `judge_rails_demo/tmp/bench/arm7_*.jsonl`, snapshotted in
514
+ `wiki/raw/transcripts/named-state-arm7-2026-09-23.md`.
515
+
516
+ ### Reference labels. PRE-REGISTERED on 2026-09-23
517
+
518
+ **Why.** Every previous arm measures agreement with raw text, never correctness.
519
+ `db/seed_tickets.rb` carries no labels. Without a reference, the "calmer, less urgent" shift in arm 7
520
+ cannot be settled.
521
+
522
+ **Shape.** Two independent annotators label the 250 tickets on the three questions, using the exact
523
+ definitions in `app/models/ticket.rb`: urgency true or false, intent among the four teams,
524
+ frustration from 0 to 3. They read only `db/seed_tickets.rb`, never a Jev judgment. One reads in
525
+ order, the other in reverse. Each annotator flags the calls they consider borderline.
526
+
527
+ **Caveat, written before seeing the result.** Both annotators are models (Claude Opus), not humans.
528
+ The reference is their **consensus**: a ticket counts for a question only if both agree.
529
+ Disagreements are listed for human review. Any figure read against this reference is written as
530
+ "agreement with the annotator consensus", never as "correctness".
531
+
532
+ Storage: `db/ticket_labels.json`, keyed by `subject|customer_name` like `db/seed_judgments.rb`.
533
+
534
+ ### Analyses without requests. PRE-REGISTERED on 2026-09-23
535
+
536
+ `bin/rails judge:eval:offline`. Reads `db/seed_judgments.rb` and `db/ticket_labels.json` and calls
537
+ nothing.
538
+
539
+ 1. **Baseline correctness** against the consensus. urgency: decision at p >= 0.5 against the label,
540
+ plus the Brier score. intent: argmax, plus the multiclass Brier. frustration: rounded level, plus
541
+ the mean absolute error.
542
+ 2. **Confidence bands.** The share of tickets in the noul `:unsure` band (0.2 < p < 0.8), and the
543
+ share of intents below 0.6 confidence. For each band, the agreement with the consensus. If
544
+ low-confidence answers are no less correct than the others, confidence tells us nothing here.
545
+ 3. **Recalibration.** One temperature per question, fitted by likelihood on a random half (fixed
546
+ seed) and tested on the other. Criterion: it is adopted only if the Brier on the test half drops
547
+ by at least 10% against the identity. Otherwise Jev is declared calibrated on this dataset, to
548
+ the precision of 125 tickets.
549
+
550
+ ### Arm 7b. Is the named-key shift a constant? PRE-REGISTERED on 2026-09-23
551
+
552
+ The arm 7 JSONL files keep only the decisions, not the probabilities. `named_1` is replayed once (250
553
+ requests, seed 7) keeping the raw values. On one half, we estimate the mean logit shift between
554
+ `named_1` and the baseline, per question (urgency: logit of p; frustration: difference in continuous
555
+ value). We subtract it on the other half.
556
+
557
+ **Fixed reading.** If the test half's agreement with the baseline climbs back to within 2 points of
558
+ the control, the shift is a correctable constant. Otherwise it depends on the ticket. Either way,
559
+ the annotator consensus says which of the two readings is closer.
560
+
561
+ ### Arm 8. Structured criteria. PRE-REGISTERED on 2026-09-23
562
+
563
+ **Why.** `docs.typesafe.ai/primitives/advanced.md` documents Choice options as
564
+ `{what, not_for, examples}` objects and Noul `true` and `false` criteria as objects, "to sharpen the
565
+ boundary". No figures are published. Until now the gem converted every entry to a string. It now
566
+ accepts objects, without changing the digest of a string criterion.
567
+
568
+ **Shape.** One subject per request, raw text, three questions, concurrency 8. Only the urgency and
569
+ intent criteria change. The `what` repeats the original text word for word, and `not_for` and two
570
+ generic `examples` are added. No example is taken from the 250 tickets. frustration does not change:
571
+ it is the control question and must not move.
572
+
573
+ | Variant | Criteria | Requests |
574
+ |---|---|---|
575
+ | `current` | those in `ticket.rb` | 250 |
576
+ | `structured` | objects | 250 |
577
+ | `structured_replay` | objects, first 50 tickets replayed | 50 |
578
+
579
+ **Criterion, fixed before the run**, against the annotator consensus, `structured` against
580
+ `current` from the same day:
581
+
582
+ - adopted if, on urgency and on intent, agreement drops by no more than one point and rises by at
583
+ least 2 points on one of the two, or if the Brier drops by at least 10% with no loss of agreement;
584
+ - rejected otherwise. The extra token cost is reported, not weighed: at the confirmed price it is
585
+ negligible.
586
+
587
+ `structured_replay` gives its noise floor. frustration must stay within the noise of arm 1.
588
+
589
+ Command: `bin/rails judge:eval:criteria VARIANT=current|structured|structured_replay`.
590
+ Since the demo switched over, `current` is called `plain`: the model's criteria are structured, and
591
+ `plain` rebuilds the original strings, including the digests `9e08ebfeb6e0e0fa` and
592
+ `5f3c0bad775abc1c`.
593
+
594
+ ### Arm 9. Sustained rate limit. PRE-REGISTERED on 2026-09-23
595
+
596
+ **Why.** `docs.typesafe.ai/models.md` documents 1,200 requests per minute and 250,000 tokens per
597
+ second. Arm 6 sent 100 requests per level and could not fill a minute. At 8 threads the pool runs at
598
+ ~28 requests per second, above the documented limit.
599
+
600
+ **Shape.** The production shape, looping over the 250 tickets for 75 s, with a client that never
601
+ retries (`max_retries = 0`) so every raw 429 and its `Retry-After` are visible. Levels run in
602
+ sequence: 8, then 16, then 32 only if 16 saw no 429. Two minutes of pause between levels.
603
+
604
+ **Fixed reading.**
605
+
606
+ - 429s at 8: the limit applies and the gem's default exceeds it. We need a client-side limiter or a
607
+ `max_retry_wait` that covers the observed `Retry-After`.
608
+ - No 429 at 32 over 75 s: the documented limit is not enforced on this key today. The default of 8
609
+ stays, and the limit is written down in the README.
610
+ - In between: the observed threshold is reported, and the default does not exceed the last level
611
+ without a 429.
612
+
613
+ Budget: at most ~9,000 requests, ~4.8M tokens, ~$0.20 at the confirmed price.
614
+
615
+ Command: `bin/rails judge:eval:rate CONCURRENCY=8 DURATION=75`.
616
+
617
+ ### Results of the 2026-09-23 arms (labels, 7b, 8, 9)
618
+
619
+ Everything was served by `jev-1.13.0`, checked on `ResultSet#model`. Total spend: 8,462,588 input
620
+ tokens, **~$0.36** at the confirmed price. Arm 9 cost 7.94M tokens against the 4.8M planned:
621
+ throughput doubled with each thread level, like everything else, and the budget underestimated it.
622
+
623
+ #### Labels
624
+
625
+ | Question | Inter-annotator agreement | kappa | Consensus kept |
626
+ |---|---|---|---|
627
+ | urgency | 96.4% | 0.88 | 241 |
628
+ | intent | 88.4% | 0.82 | 221 |
629
+ | frustration | 84.8% | 0.75 | 212 |
630
+
631
+ The main disagreement is telling: 15 thank-you messages are `technical` for one annotator and
632
+ `spam` for the other, because the spam definition includes "anything not a genuine support request".
633
+ The demo's criteria have no place for a message that is neither a request nor spam.
634
+
635
+ #### Analyses without requests
636
+
637
+ | Question | Baseline against consensus | Brier |
638
+ |---|---|---|
639
+ | urgency, decision at 0.5 | 82.6% | 0.1141 |
640
+ | intent | 84.2% | 0.2285 |
641
+ | frustration, rounded level | 60.8% | mean error 0.399 |
642
+
643
+ **The errors go one way.** Jev reads tickets as more urgent (41 false positives, 1 false negative at
644
+ 0.5) and more frustrated (83 levels too high, 0 too low) than both annotators. For intent, it says
645
+ `billing` where the consensus says `technical` 18 times.
646
+
647
+ **Bands.** Confident urgency (p <= 0.2 or >= 0.8): 97.7% agreement on 131 tickets. `:unsure` band:
648
+ 64.5% on 110. intent at confidence >= 0.6: 88.7% on 194; below 0.6: 51.9% on 27. Confidence sorts
649
+ the answers well, so routing the uncertain band makes sense here.
650
+
651
+ **Temperature.** urgency +0.2%, intent -0.2%, frustration -9.3% of Brier on the test half. None
652
+ crosses the 10% threshold: **no temperature recalibration.**
653
+
654
+ **Exploratory, not pre-registered, tested on a held-out half.** For frustration, subtracting 0.5
655
+ before rounding, which amounts to taking the integer part, raises agreement from 62.7% to 83.3%.
656
+ For urgency, the 0.5 threshold gives 80.8%, and the demo's 0.8 threshold gives 92.5%, on par with
657
+ the best learned threshold (0.75). The cost comes from how the demo reads the number: rounding and
658
+ the median threshold.
659
+
660
+ #### Arm 7b
661
+
662
+ Shift learned on 125 tickets: -0.164 in logit on urgency, -0.083 in value on frustration. The
663
+ per-ticket standard deviation is 0.188, the same order as the mean. The correction raises the test
664
+ half's agreement with the baseline from 115 to 121 out of 125 on urgency, and from 110 to 113 on
665
+ frustration. **The 2-point criterion is not met: the shift is not a constant.**
666
+
667
+ Against the consensus, however, `named_1` beats raw text on all three questions:
668
+
669
+ | | urgency | intent | frustration |
670
+ |---|---|---|---|
671
+ | raw text | 82.6% (Brier 0.114) | 84.2% (0.229) | 60.8% |
672
+ | `named_1` | 85.9% (0.099) | 86.4% (0.213) | 70.3% |
673
+
674
+ The "calmer, less urgent" shift from arm 7 moves **toward** the labels. Arm 7 measured distance from
675
+ raw text, and raw text is the reading furthest from the consensus. That still does not reopen
676
+ batching by subject: `named_20` was not re-read against the labels.
677
+
678
+ #### Arm 8
679
+
680
+ | Variant | urgency | intent | frustration | Tokens / ticket |
681
+ |---|---|---|---|---|
682
+ | `current` | 82.6% (Brier 0.1141) | 85.1% (0.2271) | 62.3% | 539 |
683
+ | `structured` | **86.7%** (0.0897) | **89.6%** (0.1608) | 61.8% | 825 |
684
+
685
+ Against the consensus. `current` reproduces the baseline (98.8 / 99.2 / 98.8%). `structured` against
686
+ its own replay on 50 tickets: 50/50, 50/50, 49/50, mean drift 0.008. frustration, the control
687
+ question, stays within the noise.
688
+
689
+ **Criterion met: adopted.** +4.1 and +4.5 points, Brier -21% and -29%. `billing` errors that should
690
+ have been `technical` drop from 18 to 12. Extra cost: +286 tokens per ticket, +53%, ~$0.000012.
691
+
692
+ A caveat noticed afterwards: at the demo's escalation threshold (0.8) rather than 0.5, urgency goes
693
+ from 91.3% to 90.0%, three tickets. The urgency gain rests on the Brier and the median threshold,
694
+ while the intent gain holds everywhere. The demo is not switched over: changing its criteria
695
+ invalidates the 250 committed judgments, and the reference is still a consensus of models.
696
+
697
+ #### Arm 9
698
+
699
+ | Threads | Requests | Throughput | Max over 60 s | 429 | Latency p50 / p95 |
700
+ |---|---|---|---|---|---|
701
+ | 8 | 2,187 | 29.2 /s | 1,758 | 0 | 0.268 / 0.336 s |
702
+ | 16 | 4,214 | 56.2 /s | 3,393 | 0 | 0.277 / 0.361 s |
703
+ | 32 | 8,355 | 111.4 /s | 6,696 | 0 | 0.279 / 0.371 s |
704
+
705
+ No 429s and no errors over 14,756 requests without retries. At 32 threads that is 5.6 times the
706
+ documented per-minute limit and ~60,000 tokens per second, a quarter of the token limit. Latency
707
+ does not move. **Fixed reading: the documented limit is not enforced on this key today.** The
708
+ default of 8 stays. The limit and this measurement are written in the gem's README, with the
709
+ provider's caveat ("can change without notice").
710
+
711
+ Raw data: `judge_rails_demo/tmp/bench/arm7bis_named_1.jsonl`, `arm8_*.jsonl`, `arm9_c*.jsonl`,
712
+ snapshotted in `wiki/raw/transcripts/labels-and-arms-8-9-2026-09-23.md`.
713
+
714
+ ## 7. Acceptance criterion and verdict
715
+
716
+ ### VERDICT, 2026-09-21: batching by subject fails, and the cause is structural.
717
+
718
+ The hard criterion required 100% agreement on the choice label, the noul band and the score level.
719
+ **It fails at every batch size, including 2.** Read relative to arm 1, as the section below
720
+ prescribes, the gap is still 9 to 15 points of agreement and the drift is 5 to 16 times the
721
+ control's noise floor.
722
+
723
+ Batching as implemented trades 10 to 15 points of decision accuracy for 49x the speed. For a gem
724
+ whose product is a calibrated probability, that is not an acceptable default trade. This protocol
725
+ was written before the result was known precisely so the result could not be rationalized
726
+ afterwards.
727
+
728
+ `Judge::Batch` stays in `lib/`, tested and not wired in. `Judge.config.batch_rows` stays `nil`. No
729
+ path in the gem batches.
730
+
731
+ ### Next steps, decided by arm 5
732
+
733
+ **Corrected on 2026-09-23.** The "3 from shape, 9 from neighbors" split below is refuted by arm 7,
734
+ and arm 7b shows that the named shape is closer to the labels than raw text. The recommendation (the
735
+ thread pool) holds for other reasons, written up in arm 7.
736
+
737
+ Arm 5 ran and it closes the question: of the 12 points lost, 3 come from the shape and 9 from the
738
+ neighbors. Fixing the shape would still leave 9 points on the table, so no version of batching by
739
+ subject passes the hard criterion again.
740
+
741
+ **What to build instead, backed by the measurements:**
742
+
743
+ 1. **A thread pool on `judge_filter`, with no batching.** Same request shape as today, so **zero
744
+ accuracy loss**, and it is the solution the Re-Ranking cookbook applies to 3,565 passages. Arm 4
745
+ measures 7.8x for 8 threads, and arm 5 confirms it on the real API: 150 tickets in 5.44 s at 8
746
+ threads, 0.036 s per ticket against 0.270 s in series.
747
+ 2. **Do not wire in `Judge::Batch`.** It stays in `lib/`, tested, with this document as the written
748
+ reason.
749
+ 3. **Move `[[typesafe-jev]]` to `high`** on flat latency as a function of the number of questions,
750
+ which is now measured.
751
+
752
+ ### The criterion, as it was fixed
753
+
754
+ A double budget, fixed before the run.
755
+
756
+ **Hard: must pass, or batching does not ship at that batch size:**
757
+
758
+ - 100% agreement on the choice label
759
+ - 100% agreement on the noul's `judge_decide` band, at the demo's thresholds
760
+ - 100% agreement on the integer level of the score
761
+
762
+ **Soft: read as a curve, not as a threshold:**
763
+
764
+ - mean and p95 `|Δp|` per batch size and per question type
765
+ - total variation distance on the choice distribution
766
+ - continuous value difference on the score
767
+
768
+ Both budgets are read **relative to arm 1**, never in absolute terms. If the unbatched control
769
+ itself drifts by 0.03 in mean `|Δp|`, a batch at 0.03 has degraded nothing.
770
+
771
+ The default `batch_rows` would be the largest batch size that passes the hard budget and whose soft
772
+ drift stays at the control's level. Not 20 just because pg_jev says 20.
773
+
774
+ ## 8. What still needs protecting
775
+
776
+ - The current baseline must be snapshotted in `wiki/raw/transcripts/` with its sha256 **before** any
777
+ regeneration of `db/seed_judgments.rb`, otherwise the baseline is lost.
778
+ - Arms 0 to 3 never run in `rake test`. The house rule is never to call the real API from a test.
779
+ They will be a manual task. Arm 4 lives in `bin/bench_batch`, outside the suite, and only its
780
+ reduced version is a test.
781
+ - Batching widens the blast radius of a prompt injection from 1 to `batch_rows`, since the rows share
782
+ a `state` and the vendor's documentation describes the model as steerable by injected
783
+ instructions. This protocol does not measure it. An adversarial arm would have to be written
784
+ separately.
785
+
786
+ ## 8b. The optimal request recipe, as measured
787
+
788
+ Three rules, each backed by an arm.
789
+
790
+ | Rule | Arm | Gain | Accuracy cost |
791
+ |---|---|---|---|
792
+ | **One subject per request** | 2, 3, 5, 7 | - | leaving raw text costs ~6 points of agreement with the baseline, a neighbor ~0.3 (arm 7). Against the labels, the named shape does better (arm 7b) |
793
+ | **All of the subject's questions in the same request** | 0 | 326-token overhead amortized, flat latency up to 40 questions | **zero** |
794
+ | **Thread fan-out over subjects** | 6 | 6.95x at 8 threads, 18.4x at 32 | **zero** |
795
+
796
+ The gem now applies all three: `storage.rb:26-34` groups a record's questions into one call, and
797
+ `relation.rb` fans the records out over a pool.
798
+
799
+ ## 9. Log
800
+
801
+ ### 2026-09-21, phase A
802
+
803
+ Written: `lib/judge/batch.rb` (L2 engine, stdlib only), `Judge::PayloadTooLargeError` in the error
804
+ taxonomy, `batch_rows` and `concurrency` on the configuration at `nil`, `test/batch_test.rb` (22
805
+ tests), `test/batch_throughput_test.rb` (6 tests), `bin/bench_batch`.
806
+
807
+ Measured: 200 tests green against 172 before, 580 assertions, zero rubocop offenses, and the Rails
808
+ 7.2, 8.0 and 8.1 gemfiles all green. Arm 4 ran, table above.
809
+
810
+ Two corrections to the protocol, made while running it:
811
+
812
+ 1. Arm 4 cannot live in the test suite. At the measured latency of 0.300 s, 250 rows in sequence
813
+ take 76 seconds.
814
+ 2. `rows:` carries no accuracy ceiling. The previous version of this plan set one at 25, which would
815
+ have stopped arms 2 and 3 from measuring a batch of 40. The measured curve is the guard, not a
816
+ guessed constant.
817
+
818
+ Naming decision: the questions are called `row0`, `row1` on the wire, in normalcase, because `row_0`
819
+ trips `Naming/VariableNumber` and the only alternatives were to disable the cop or to modify a
820
+ pre-existing test file that did not belong to this work.
821
+
822
+ ### 2026-09-21, phase B
823
+
824
+ Written: `judge_rails/bin/arm0_latency`, `judge_rails_demo/lib/tasks/judge_bench.rake`,
825
+ `wiki/raw/transcripts/seed-judgments-baseline-2026-09-21.md` (sha256
826
+ `0b358b56125d79dfef0b42c351d63c5c603c99066263b78f287fe643457a5839`). `Judge::Batch.judge` extended to
827
+ accept a Hash of questions, which arm 3 required. 206 tests green, zero offenses.
828
+
829
+ Fixed along the way: the demo's `config/initializers/judge.rb` only read `JEV_API_KEY`, while
830
+ `Judge::Configuration` has always accepted `TYPESAFE_API_KEY` as a fallback. The demo now accepts
831
+ both.
832
+
833
+ Actual campaign spend: about 1,000 requests and ~1.40 million input tokens, so **~$0.06** at the
834
+ borrowed price of $0.042/M. The projected budget was $0.03 for 430 requests. The extra diagnosis at
835
+ a batch of 2 doubled the volume.
836
+
837
+ No database writes. `db/seed_judgments.rb` is intact, and the batched path never touched it.
838
+
839
+ ### 2026-09-21, phase C revised
840
+
841
+ Batching by subject is not wired in and will not be. What was wired in instead:
842
+
843
+ - `lib/judge/pool.rb`, new. `Judge::Pool.map(items, concurrency:)` preserves input order, runs
844
+ concurrency 1 inline with no thread, and re-raises the first error after all workers have
845
+ stopped. `Judge::Batch` builds on it, so the threading code exists only once.
846
+ - `lib/judge/rails/relation.rb`. `judge_filter`, `judge_map` and `judge_sort` take `concurrency:` and
847
+ fan out over the pool. The `state` values are built on the calling thread before the fan-out, so
848
+ no worker touches an ActiveRecord connection.
849
+ - Default: explicit `concurrency:`, otherwise `Judge.config.concurrency`, otherwise 8.
850
+
851
+ 219 tests green against 172 at the start, zero offenses, Rails 7.2, 8.0 and 8.1 green.
852
+
853
+ ### 2026-09-23, labels and arms 7b, 8, 9
854
+
855
+ Written: `db/ticket_labels.json` (two model annotators, consensus), `lib/tasks/judge_eval.rake`
856
+ (`judge:eval:offline`, `named_values`, `criteria`, `rate`). In the gem, `Question::Choice` and
857
+ `Question::Noul` accept object or array entries. A string entry keeps its digest, checked against
858
+ the three baseline digests. 263 tests green on Rails 7.2, 8.0 and 8.1, zero offenses. Demo: 42 tests
859
+ green, zero offenses.
860
+
861
+ Fixed on re-reading: the gem's README showed `question.digest # => "9e08ebfeb6e0e0fa"` for a choice.
862
+ That is the digest of the demo's urgency question. The real value is `1c3be3cf24579b81`.
863
+
864
+ Pre-registered before any request, run afterwards. The key is read from the shell environment on
865
+ each command and never written into the project.
866
+
867
+ ### 2026-09-23, the demo switches to structured criteria
868
+
869
+ Decided by the owner after arm 8. `app/models/ticket.rb` declares urgency and intent with the arm 8
870
+ criteria, word for word. The benches follow without changes, since they read the model's
871
+ definitions: `judge_bench.rake` (arms 1 to 7) and `judge_eval.rake`. The gem's two benches,
872
+ `bin/arm0_latency` and `bin/bench_batch`, declare the same shape.
873
+
874
+ **The baseline has changed.** `db/seed_judgments.rb` is regenerated: 250 requests, `jev-1.13.0` on
875
+ all 250, urgency digest `4f1911289ec33ca2` and intent digest `4234fcadf9f795b9`, frustration
876
+ unchanged at `d3f59e90f1ce9b41`. Against the structured replay from arm 8: 246/250, 250/250 and
877
+ 243/250, within arm 1's noise. The old baseline is snapshotted in
878
+ `wiki/raw/transcripts/seed-judgments-plain-2026-09-23.md`. Any arm replayed from here on is read
879
+ against the structured baseline, not the one behind the figures above.
880
+
881
+ New baseline against the consensus: urgency 88.4% (Brier 0.090), intent 89.6% (0.162), frustration
882
+ 62.3%. Confident bands: urgency 99.2% on 133, intent 93.9% on 196. No temperature crosses 10% yet
883
+ (urgency -6.7%, frustration -9.5%).
884
+
885
+ `db:seed` on an empty database restores 250 judgments, 0 stale. Demo: 42 tests green, zero offenses.