logit-classifier 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (56) hide show
  1. logit_classifier-0.1.0/.gitignore +28 -0
  2. logit_classifier-0.1.0/FINDINGS.md +334 -0
  3. logit_classifier-0.1.0/LICENSE +674 -0
  4. logit_classifier-0.1.0/PKG-INFO +354 -0
  5. logit_classifier-0.1.0/README.md +321 -0
  6. logit_classifier-0.1.0/examples/compare_models.py +76 -0
  7. logit_classifier-0.1.0/examples/http_client.py +72 -0
  8. logit_classifier-0.1.0/examples/images.py +76 -0
  9. logit_classifier-0.1.0/examples/own_backend.py +122 -0
  10. logit_classifier-0.1.0/examples/question_types.py +83 -0
  11. logit_classifier-0.1.0/examples/quickstart.py +87 -0
  12. logit_classifier-0.1.0/pyproject.toml +178 -0
  13. logit_classifier-0.1.0/src/logit_classifier/__init__.py +90 -0
  14. logit_classifier-0.1.0/src/logit_classifier/__main__.py +10 -0
  15. logit_classifier-0.1.0/src/logit_classifier/backends/__init__.py +5 -0
  16. logit_classifier-0.1.0/src/logit_classifier/backends/base.py +125 -0
  17. logit_classifier-0.1.0/src/logit_classifier/backends/hf.py +412 -0
  18. logit_classifier-0.1.0/src/logit_classifier/calibrate.py +141 -0
  19. logit_classifier-0.1.0/src/logit_classifier/classifier.py +329 -0
  20. logit_classifier-0.1.0/src/logit_classifier/cli.py +51 -0
  21. logit_classifier-0.1.0/src/logit_classifier/config.py +187 -0
  22. logit_classifier-0.1.0/src/logit_classifier/deps.py +30 -0
  23. logit_classifier-0.1.0/src/logit_classifier/errors.py +16 -0
  24. logit_classifier-0.1.0/src/logit_classifier/labels.py +91 -0
  25. logit_classifier-0.1.0/src/logit_classifier/prompt.py +154 -0
  26. logit_classifier-0.1.0/src/logit_classifier/py.typed +0 -0
  27. logit_classifier-0.1.0/src/logit_classifier/schema.py +285 -0
  28. logit_classifier-0.1.0/src/logit_classifier/scoring.py +102 -0
  29. logit_classifier-0.1.0/src/logit_classifier/service.py +174 -0
  30. logit_classifier-0.1.0/src/logit_classifier/vision.py +95 -0
  31. logit_classifier-0.1.0/src/logit_classifier/web/index.html +411 -0
  32. logit_classifier-0.1.0/tests/fixtures/banking77_test.json +1 -0
  33. logit_classifier-0.1.0/tests/fixtures/eval_set.json +41 -0
  34. logit_classifier-0.1.0/tests/fixtures/many_options_request.json +100 -0
  35. logit_classifier-0.1.0/tests/fixtures/quickstart_request.json +11 -0
  36. logit_classifier-0.1.0/tests/test_model.py +200 -0
  37. logit_classifier-0.1.0/tests/test_unit.py +1293 -0
  38. logit_classifier-0.1.0/tests-AB/ab_branch_packing.py +227 -0
  39. logit_classifier-0.1.0/tests-AB/ab_determinism_scope.py +210 -0
  40. logit_classifier-0.1.0/tests-AB/ab_env.py +86 -0
  41. logit_classifier-0.1.0/tests-AB/ab_math_sdp_reduction.py +71 -0
  42. logit_classifier-0.1.0/tests-AB/ab_multi_label.py +131 -0
  43. logit_classifier-0.1.0/tests-AB/ab_noul_wording.py +85 -0
  44. logit_classifier-0.1.0/tests-AB/ab_temperature.py +127 -0
  45. logit_classifier-0.1.0/tests-AB/banking77.py +94 -0
  46. logit_classifier-0.1.0/tests-AB/benchmark.py +134 -0
  47. logit_classifier-0.1.0/tests-AB/evaluate.py +170 -0
  48. logit_classifier-0.1.0/tests-AB/inputs/_full.png +0 -0
  49. logit_classifier-0.1.0/tests-AB/inputs/_sheet.png +0 -0
  50. logit_classifier-0.1.0/tests-AB/inputs/low-L.png +0 -0
  51. logit_classifier-0.1.0/tests-AB/inputs/low-R.png +0 -0
  52. logit_classifier-0.1.0/tests-AB/inputs/mid-L.png +0 -0
  53. logit_classifier-0.1.0/tests-AB/inputs/mid-R.png +0 -0
  54. logit_classifier-0.1.0/tests-AB/inputs/top-L.png +0 -0
  55. logit_classifier-0.1.0/tests-AB/inputs/top-R.png +0 -0
  56. logit_classifier-0.1.0/tests-AB/tune_groups.py +91 -0
@@ -0,0 +1,28 @@
1
+ # Python-generated files
2
+ __pycache__/
3
+ *.py[oc]
4
+ build/
5
+ dist/
6
+ wheels/
7
+ *.egg-info
8
+
9
+ # Virtual environments
10
+ .venv
11
+
12
+ .hf-cache/
13
+ *.pyc
14
+ calibration.json
15
+ .pytest_cache/
16
+ banking77_prior.json
17
+ tests-AB/cache/
18
+ .ruff_cache/
19
+ .mypy_cache/
20
+
21
+ # Local Claude state, the project guide it reads, and the build plan
22
+ .claude/
23
+ CLAUDE.md
24
+ TODO.md
25
+
26
+ # Local machine paths, set per install from the README
27
+ .env
28
+ models/
@@ -0,0 +1,334 @@
1
+ # Findings
2
+
3
+ What changed, why, and what was measured. Every number here came from a script in
4
+ `tests-AB/`, on an RTX 3090 Ti. Reproduce any row by running the script named beside it.
5
+
6
+ ## The pipeline every finding attaches to
7
+
8
+ ```
9
+ request JSON
10
+ → parse validate, lift any image out of the state
11
+ → label give each option one letter, A to Z then a to z
12
+ → assemble build the shared prefix once, one suffix per branch
13
+ → prefill append "Answer: (" to the rendered chat template
14
+ → score encode the prefix once, broadcast its cache, one batched forward
15
+ → read gather the logits at the candidate letter ids
16
+ → calibrate subtract the learned label prior, divide by the fitted temperature
17
+ → normalize softmax over those letters only, in float64
18
+ → map letters back to option names
19
+ response JSON
20
+ ```
21
+
22
+ A branch is one prompt whose last token position carries a distribution. A question needs
23
+ one branch, or several when it has levels or more than 52 options.
24
+
25
+ ## Model
26
+
27
+ | | Qwen3-4B-Instruct-2507 | Qwen3-VL-4B-Instruct |
28
+ |---|---|---|
29
+ | Banking77, 462 held out rows | 0.554 | **0.626** |
30
+ | eval set, 39 items | 0.897 | **0.949** |
31
+ | eval set, score questions | 0.750 | **1.000** |
32
+ | eval set, noul questions | **1.000** | 0.917 |
33
+ | median latency | 171 ms | **169 ms** |
34
+ | reads images | no | **yes** |
35
+
36
+ The vision model is better at text as well as capable of images, at the same speed, so it
37
+ is the default. `ab_temperature.py`, `evaluate.py`.
38
+
39
+ Noul is the one regression, down 8 points, and noul is what image tagging uses most.
40
+
41
+ ## Temperature belongs to the model
42
+
43
+ Fitted on the held out half of a 924 row Banking77 sample, chosen on calibration error.
44
+
45
+ | Model | fitted T | ECE at its own T | ECE at the other model's T |
46
+ |---|---|---|---|
47
+ | Qwen3-VL-4B-Instruct | 1.25 | 0.088 | 0.565 |
48
+ | Qwen3-4B-Instruct-2507 | 6.0 | 0.029 | 0.088 |
49
+
50
+ Nearly a 5x difference between two models of the same size and family, so a single constant
51
+ was wrong. `config.py` holds `FITTED_TEMPERATURES`, and an unfitted model gets 2.5.
52
+
53
+ The fit follows calibration error, not accuracy. Accuracy climbs with temperature above 52
54
+ options, but that is the split path defect below, not a gain. `ab_temperature.py`.
55
+
56
+ ## The prior corrects letters, not fixed positions
57
+
58
+ The running prior subtracts the model's standing preference for one letter over another.
59
+ Two positions carry a fixed meaning instead. The escape label always means none of these,
60
+ and a score's letters always name the same rungs. Their mean mass follows the traffic, so
61
+ correcting them moves one request's answer with unrelated requests.
62
+
63
+ Banking77, 462 held out rows, Qwen3-VL-4B-Instruct at its fitted temperature.
64
+
65
+ | Prior | Accuracy | ECE |
66
+ |---|---|---|
67
+ | off | 0.606 | 0.083 |
68
+ | escape corrected, score shares the choice bucket | 0.621 | 0.114 |
69
+ | escape left alone, score in its own bucket | **0.630** | 0.103 |
70
+
71
+ The fix does not regress calibration. The accuracy gain is four rows in 462, near the noise
72
+ of this sample. The prior still costs calibration against no prior, because the shipped
73
+ temperature was fitted with the prior off. `banking77.py`, with and without `--no-prior`.
74
+
75
+ ## Option order helps above 52 options and nowhere else
76
+
77
+ | Options | Letterings | Accuracy | ECE |
78
+ |---|---|---|---|
79
+ | 10, one branch | 1 | 0.881 | 0.049 |
80
+ | 10, one branch | 4 | 0.881 | 0.025 |
81
+ | 77, split | 1 | 0.554 | 0.097 |
82
+ | 77, split | 4 | **0.693** | 0.288 |
83
+
84
+ OpenJev ships the same idea as `READOUT_PERMS`, default off, with a code comment claiming
85
+ about 16 points. The gain reproduces, but it is group assignment rather than letter
86
+ position. Ten options fit one branch, where relettering changes nothing measurable.
87
+
88
+ Shipped as `LOGIT_PERMUTATIONS`, default 1. Calibration gets worse as accuracy improves, and
89
+ no temperature between 0.5 and 50 repaired it.
90
+
91
+ ## Abstain
92
+
93
+ Every choice question carries a `none of these` label, reported as `abstain` outside
94
+ `probabilities`. Measured against rows whose correct label was deleted from the list, so
95
+ declining was the only right answer.
96
+
97
+ | Setup | AUC |
98
+ |---|---|
99
+ | one branch, 10 options | **0.878** |
100
+ | split path, 77 options | 0.639 |
101
+
102
+ The split path is weak because a group that does not hold the answer declines whether or not
103
+ the answer sits in another group.
104
+
105
+ Two alternatives were tested and rejected. Mass on letters nobody offered reaches AUC 0.806
106
+ alone, but adding it to the escape label drops the pair to 0.850. Combining across groups by
107
+ product scored 0.616 against 0.639 for the weakest decline, and its scale shrinks with the
108
+ group count, so five honest half declines would read as 0.03. The weakest decline ships.
109
+
110
+ Offering the label does not measurably change accuracy. It moved by two to three examples in
111
+ 462, inside the noise of that sample. `ab_abstain` measurements, `ab_temperature.py`.
112
+
113
+ ## Abstain and confidence answer different questions
114
+
115
+ | What to catch | Signal | AUC |
116
+ |---|---|---|
117
+ | nothing in the list fits | `abstain` | 0.878 |
118
+ | the tag is probably wrong | `confidence` | 0.864 |
119
+
120
+ They correlate at -0.64, so they overlap, but each wins its own job. A tagging pipeline
121
+ should check both.
122
+
123
+ ## Several true answers need one noul each
124
+
125
+ | Method | AUC across tiles | AUC within a tile | best F1 |
126
+ |---|---|---|---|
127
+ | one `noul` per fragment | **0.963** | **1.000** | **0.909** |
128
+ | one `choice` over all fragments | 0.667 | 0.685 | 0.516 |
129
+
130
+ A choice is normalized and sums to 1. On a tile holding four fragments it gave the winner
131
+ 1.00 and the other three 0.00. The published result that symbol scoring beats independent
132
+ scoring by 9.7 points covers single label tasks, and does not transfer here.
133
+ `ab_multi_label.py`.
134
+
135
+ ## The wording of a boolean does not matter
136
+
137
+ Yes and no against true and false, correct and incorrect, present and absent, visible and
138
+ hidden, agree and disagree. AUC spanned 0.946 to 0.957 on 462 balanced text statements and
139
+ 0.960 to 0.969 on the tiles. That is inside noise at these sample sizes.
140
+
141
+ The model writes "no" more readily than "yes" when generating. The readout never asks it to
142
+ write anything, so both words are equally reachable at the one position we read.
143
+ `ab_noul_wording.py`.
144
+
145
+ ## Confidence formulas
146
+
147
+ TypeSafe publishes two, and we were using the choice one for score questions. A score is
148
+ ordinal, so `[0, 0.5, 0.5, 0, 0]` and `[0.5, 0, 0, 0, 0.5]` are different answers that the
149
+ peak height formula scores identically at 0.375. The official formula gives 0.583 and 0.0.
150
+
151
+ Verified by reading `system_one_adapter._utils.confidence_metrics` from TypeSafe's own PyPI
152
+ package, not by trusting a third party's description of it.
153
+
154
+ ## Which torch globals decide bit-exactness
155
+
156
+ `backends/hf.py` holds a set of process-global torch settings around each forward pass. On
157
+ the CUDA attention path, only one of them changes a logit on either shipped model.
158
+
159
+ | setting flipped away from pinned | Qwen3-VL-4B-Instruct | Qwen3-4B-Instruct-2507 |
160
+ |---|---|---|
161
+ | `cudnn.benchmark` | 0.0 | 0.0 |
162
+ | `cudnn.deterministic` | 0.0 | 0.0 |
163
+ | `allow_bf16_reduced_precision_reduction` | **0.5** | **0.75** |
164
+ | float32 matmul precision | 0.0 | 0.0 |
165
+
166
+ Max absolute logit shift against the fully pinned run, window removed so the process
167
+ globals reach the forward pass. The forward runs in bfloat16, so the fp32 matmul setting
168
+ governs matmuls this graph does not have, and the cudnn pair governs convolutions that
169
+ only the vision tower's patch embed uses. The three inert settings are kept because
170
+ `cudnn.benchmark` picks convolution algorithms by timing, which is specific to the card
171
+ and the model, and this was measured on one card with two models.
172
+
173
+ Scoping the settings to the forward pass moved nothing. The same request returns the same
174
+ logits, and a run under a hostile host state matches a run under a pinned one, both at a
175
+ gap of 0.0 on both models.
176
+
177
+ `allow_fp16_bf16_reduction_math_sdp` is held for a different reason and the sweep above
178
+ could not reach it. `sdpa_kernel` turns the math backend **off** while it holds the CUDA
179
+ backends, so this setting is live only when `_usable_attention_backends` returns nothing and
180
+ no `sdpa_kernel` block is entered. ComfyUI turns it **on** at import, at
181
+ `comfy/model_management.py:569`, so inside ComfyUI that path would run with reduced
182
+ precision reductions.
183
+
184
+ Forcing the math backend measures it on the CPU, with no card or weights.
185
+
186
+ | bfloat16 attention, math backend | bitwise equal | max abs diff |
187
+ |---|---|---|
188
+ | `(1, 4, 256, 128)` | no | 0.00977 |
189
+ | `(1, 4, 1024, 128)` | no | 0.00391 |
190
+ | `(8, 4, 256, 128)` | no | 0.00781 |
191
+
192
+ That is far above a low bit, so the math path needs the setting held.
193
+ `ab_math_sdp_reduction.py`.
194
+
195
+ Three more settings were considered and left to the host. `cudnn.enabled` is off on hosts
196
+ where cuDNN does not work, so forcing it on would break the machine that turned it off.
197
+ `cudnn.conv.fp32_precision` and `cudnn.allow_tf32` govern fp32 convolutions, and the
198
+ forward runs in bfloat16. `torch.use_deterministic_algorithms` moves toward determinism, so
199
+ a host that sets it costs us nothing.
200
+
201
+ When the probe finds no usable kernel the pass pins flash, cudnn, mem efficient and math in
202
+ that order instead of leaving the choice to the host's enable flags. Without that pin, two
203
+ hosts could dispatch one request to two different kernels.
204
+
205
+ `set_float32_matmul_precision` writes the cuda and the mkldnn matmul slots together, and
206
+ its getter reports both `none` and `ieee` as `highest`. So restoring through that setter
207
+ alone leaves the mkldnn slot changed, which is why both slots are saved raw.
208
+
209
+ `ab_determinism_scope.py`.
210
+
211
+ ## Packing branches by suffix length
212
+
213
+ Every row in a chunk is left-padded to that chunk's longest suffix, and each padded
214
+ token costs a full forward plus attention over the whole prefix. Taking branches in the
215
+ order the caller sent them drags short branches to the longest width.
216
+
217
+ | Request shape | ms | peak GB | suffix tokens | chunks | |
218
+ |---|---|---|---|---|---|
219
+ | mixed, four choices plus 12 nouls plus 4 scores | 2055.6 | 13.66 | 14600 | 1 | request order |
220
+ | | **623.9** | **10.66** | **2614** | 3 | packed |
221
+ | one 45 option choice beside 20 nouls | 2152.9 | 13.92 | 15330 | 1 | request order |
222
+ | | **425.2** | **10.94** | **1310** | 2 | packed |
223
+ | 256 nouls at the Jev cap | 1370 | 8.75 | 7584 | 8 | both, bitwise identical |
224
+ | one 77 option choice, split | 272 | 8.61 | 1282 | 1 | both, bitwise identical |
225
+ | 24 tile fragments | 198 | 8.63 | 696 | 1 | both, bitwise identical |
226
+ | 30 three option choices | 396 | 8.99 | 2160 | 1 | both, bitwise identical |
227
+ | 12 twenty option choices | 586 | 9.24 | 3960 | 1 | both, bitwise identical |
228
+
229
+ 3.29x and 5.06x on the two shapes that mix widths, each giving back about 3 GB. Nothing
230
+ at all on the five that do not, which is the point. The sort is stable, so equal lengths
231
+ keep the caller's order and the grouping stays a pure function of the request.
232
+
233
+ The two mixed shapes move low bits, by at most 0.375 and 0.5. Changing the padded width
234
+ or the batch composition of a bf16 forward does that. The shapes whose bits move are the
235
+ shapes that were three to five times too slow.
236
+
237
+ **The token ceiling is checked only when a branch widens the chunk it joins.** A request
238
+ whose branches are all one width has no padding to remove, so the ceiling would only add
239
+ chunks. Checking it unconditionally re-chunked any uniform request above 64 tokens of
240
+ width, which is every choice question, and moved its logits for no gain.
241
+
242
+ | ceiling | mixed shape, 800 token prefix | mixed widths, long prefix |
243
+ |---|---|---|
244
+ | 512 | 610.9 ms, 9.91 GB, 5 chunks | 2015.7 ms, 5 chunks |
245
+ | 1024 | 608.3 ms, 10.47 GB, 3 chunks | 2014.9 ms, 5 chunks |
246
+ | **2048** | **626.0 ms, 10.66 GB, 3 chunks** | **2015.8 ms, 5 chunks** |
247
+ | 4096 | 793.5 ms, 11.03 GB, 2 chunks | 2286.5 ms, 4 chunks |
248
+ | 8192 | 1119.4 ms, 11.84 GB, 2 chunks | 2289.5 ms, 4 chunks |
249
+ | none | 2062.2 ms, 13.66 GB, 1 chunk | 2299.1 ms, 4 chunks |
250
+
251
+ Anything from 512 to 2048 measured the same on both, and 4096 upward is clearly worse.
252
+ The value is not a knife edge. 2048 sits in the middle of the flat band and leaves the
253
+ most room before two near-equal wide branches are split apart.
254
+
255
+ `ab_branch_packing.py`.
256
+
257
+ ## Throughput and where it falls off
258
+
259
+ Warm model, RTX 3090 Ti. Adding questions to a record is close to free up to about 8,
260
+ because the fixed cost of two forward passes dominates a short record.
261
+
262
+ | Questions per record | Total ms | ms per tag | Records per hour |
263
+ |---|---|---|---|
264
+ | 1 | 138 | 138 | 26,100 |
265
+ | 4 | 136 | 34 | 26,500 |
266
+ | 8 | 140 | 17.5 | 25,800 |
267
+ | 32 | 278 | 8.7 | 13,000 |
268
+
269
+ The state is prefilled once per record, so a longer record costs more.
270
+
271
+ | State tokens | Total ms, 8 questions | ms per tag | Records per hour |
272
+ |---|---|---|---|
273
+ | 172 | 140 | 17.5 | 25,700 |
274
+ | 556 | 185 | 23 | 19,500 |
275
+ | 2,092 | 451 | 56 | 8,000 |
276
+ | 8,236 | 1,575 | 197 | 2,300 |
277
+ | 32,812 | 41,515 | 5,189 | 87 |
278
+
279
+ Prefill runs at 46 to 48 TFLOP/s from about 2,000 tokens up, near this card's practical
280
+ BF16 ceiling. Below that the time is almost all kernel launch rather than arithmetic, so a
281
+ 428 token request spends 0.14 ms on the math. The last row falls off the ceiling for the
282
+ reason under Known weaknesses. `benchmark.py`.
283
+
284
+ ## The attention backend has to be pinned on Windows
285
+
286
+ Torch 2.11 ships no FlashAttention kernel for Windows. Left to choose for itself the
287
+ dispatcher reaches the math backend, which builds the full attention matrix. Measured on an
288
+ 8,192 token prefill.
289
+
290
+ | Backend selection | Time | Peak memory |
291
+ |---|---|---|
292
+ | left to the dispatcher | 147,221 ms | 27.6 GB |
293
+ | pinned to cudnn and efficient | 1,291 ms | 9.2 GB |
294
+
295
+ `HFBackend` probes the backends at startup, at the model's own dtype, and pins the working
296
+ ones. `benchmark.py`.
297
+
298
+ ## Bugs found and fixed
299
+
300
+ | Bug | Effect | Fix |
301
+ |---|---|---|
302
+ | score used the choice confidence formula | wrong confidence on every score answer | ported both official formulas |
303
+ | attention probe ran in bfloat16 whatever the model dtype | any other dtype crashed on the first forward | probe at the model's dtype |
304
+ | attention probe used one call shape | approved a backend that then failed | probe causal and masked, with grouped query attention |
305
+ | abstain combined groups by product | shrinks with group count, never fires at scale | weakest decline |
306
+ | image source detected by string length | an 8x8 PNG encodes shorter than the threshold and was read as a path | check whether the file exists |
307
+ | M-RoPE offset dropped on branch rows | Qwen3-VL raises rather than misplacing tokens | read `rope_deltas` after the prefix pass |
308
+ | prior learned and corrected the escape label | traffic that often declines taught the prior to undo a decline | prior over the lettered options only |
309
+ | a joint score shared a prior bucket with a choice of the same width | a score's top rung learned the choice traffic's escape mass | score has its own bucket |
310
+
311
+ ## Known weaknesses
312
+
313
+ **The split path.** Three defects share one cause, that groups above 52 options cannot see
314
+ each other. Group assignment bias costs 13.9 points, temperature leaks into accuracy, and
315
+ the abstain score falls from 0.878 to 0.639. Parked by choice. Candidates are OpenJev's
316
+ chunk winners tournament and a second pass over each group's winner. Neither is measured.
317
+
318
+ **Noul temperature is unfitted.** The shipped 1.25 was fitted on choice questions. Booleans
319
+ now saturate near 0 and 1, so the best threshold sits at 0.9999 rather than near 0.5.
320
+
321
+ **Noul regressed 8 points** moving to the vision model, and image tagging leans on noul.
322
+
323
+ **Very long states fall off the prefill ceiling.** At 33,068 prefilled tokens the card runs
324
+ at 7.1 TFLOP/s against 47.9 at 8,492. The prefix key value cache is about 4.9 GB at that
325
+ length and every chunk copies it, which leaves a 24 GB card no room. Expanding the cache
326
+ once and cropping between chunks was measured and is worse, at 12,732 ms against 1,710 ms
327
+ and 23.66 GB against 15.12 GB, because a cropped cache is a non-contiguous view. Unsolved.
328
+ `benchmark.py`.
329
+
330
+ ## Layout
331
+
332
+ `tests/` holds the pytest suite, 146 unit tests and 14 that need the GPU. `tests-AB/` holds
333
+ the measurement harnesses, matching ComfyUI-ContextAnchoredTileRefine. Fixtures live in
334
+ `tests/fixtures/` and `tests-AB/ab_env.py` points at them.