logit-classifier 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- logit_classifier-0.1.0/.gitignore +28 -0
- logit_classifier-0.1.0/FINDINGS.md +334 -0
- logit_classifier-0.1.0/LICENSE +674 -0
- logit_classifier-0.1.0/PKG-INFO +354 -0
- logit_classifier-0.1.0/README.md +321 -0
- logit_classifier-0.1.0/examples/compare_models.py +76 -0
- logit_classifier-0.1.0/examples/http_client.py +72 -0
- logit_classifier-0.1.0/examples/images.py +76 -0
- logit_classifier-0.1.0/examples/own_backend.py +122 -0
- logit_classifier-0.1.0/examples/question_types.py +83 -0
- logit_classifier-0.1.0/examples/quickstart.py +87 -0
- logit_classifier-0.1.0/pyproject.toml +178 -0
- logit_classifier-0.1.0/src/logit_classifier/__init__.py +90 -0
- logit_classifier-0.1.0/src/logit_classifier/__main__.py +10 -0
- logit_classifier-0.1.0/src/logit_classifier/backends/__init__.py +5 -0
- logit_classifier-0.1.0/src/logit_classifier/backends/base.py +125 -0
- logit_classifier-0.1.0/src/logit_classifier/backends/hf.py +412 -0
- logit_classifier-0.1.0/src/logit_classifier/calibrate.py +141 -0
- logit_classifier-0.1.0/src/logit_classifier/classifier.py +329 -0
- logit_classifier-0.1.0/src/logit_classifier/cli.py +51 -0
- logit_classifier-0.1.0/src/logit_classifier/config.py +187 -0
- logit_classifier-0.1.0/src/logit_classifier/deps.py +30 -0
- logit_classifier-0.1.0/src/logit_classifier/errors.py +16 -0
- logit_classifier-0.1.0/src/logit_classifier/labels.py +91 -0
- logit_classifier-0.1.0/src/logit_classifier/prompt.py +154 -0
- logit_classifier-0.1.0/src/logit_classifier/py.typed +0 -0
- logit_classifier-0.1.0/src/logit_classifier/schema.py +285 -0
- logit_classifier-0.1.0/src/logit_classifier/scoring.py +102 -0
- logit_classifier-0.1.0/src/logit_classifier/service.py +174 -0
- logit_classifier-0.1.0/src/logit_classifier/vision.py +95 -0
- logit_classifier-0.1.0/src/logit_classifier/web/index.html +411 -0
- logit_classifier-0.1.0/tests/fixtures/banking77_test.json +1 -0
- logit_classifier-0.1.0/tests/fixtures/eval_set.json +41 -0
- logit_classifier-0.1.0/tests/fixtures/many_options_request.json +100 -0
- logit_classifier-0.1.0/tests/fixtures/quickstart_request.json +11 -0
- logit_classifier-0.1.0/tests/test_model.py +200 -0
- logit_classifier-0.1.0/tests/test_unit.py +1293 -0
- logit_classifier-0.1.0/tests-AB/ab_branch_packing.py +227 -0
- logit_classifier-0.1.0/tests-AB/ab_determinism_scope.py +210 -0
- logit_classifier-0.1.0/tests-AB/ab_env.py +86 -0
- logit_classifier-0.1.0/tests-AB/ab_math_sdp_reduction.py +71 -0
- logit_classifier-0.1.0/tests-AB/ab_multi_label.py +131 -0
- logit_classifier-0.1.0/tests-AB/ab_noul_wording.py +85 -0
- logit_classifier-0.1.0/tests-AB/ab_temperature.py +127 -0
- logit_classifier-0.1.0/tests-AB/banking77.py +94 -0
- logit_classifier-0.1.0/tests-AB/benchmark.py +134 -0
- logit_classifier-0.1.0/tests-AB/evaluate.py +170 -0
- logit_classifier-0.1.0/tests-AB/inputs/_full.png +0 -0
- logit_classifier-0.1.0/tests-AB/inputs/_sheet.png +0 -0
- logit_classifier-0.1.0/tests-AB/inputs/low-L.png +0 -0
- logit_classifier-0.1.0/tests-AB/inputs/low-R.png +0 -0
- logit_classifier-0.1.0/tests-AB/inputs/mid-L.png +0 -0
- logit_classifier-0.1.0/tests-AB/inputs/mid-R.png +0 -0
- logit_classifier-0.1.0/tests-AB/inputs/top-L.png +0 -0
- logit_classifier-0.1.0/tests-AB/inputs/top-R.png +0 -0
- logit_classifier-0.1.0/tests-AB/tune_groups.py +91 -0
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
# Python-generated files
|
|
2
|
+
__pycache__/
|
|
3
|
+
*.py[oc]
|
|
4
|
+
build/
|
|
5
|
+
dist/
|
|
6
|
+
wheels/
|
|
7
|
+
*.egg-info
|
|
8
|
+
|
|
9
|
+
# Virtual environments
|
|
10
|
+
.venv
|
|
11
|
+
|
|
12
|
+
.hf-cache/
|
|
13
|
+
*.pyc
|
|
14
|
+
calibration.json
|
|
15
|
+
.pytest_cache/
|
|
16
|
+
banking77_prior.json
|
|
17
|
+
tests-AB/cache/
|
|
18
|
+
.ruff_cache/
|
|
19
|
+
.mypy_cache/
|
|
20
|
+
|
|
21
|
+
# Local Claude state, the project guide it reads, and the build plan
|
|
22
|
+
.claude/
|
|
23
|
+
CLAUDE.md
|
|
24
|
+
TODO.md
|
|
25
|
+
|
|
26
|
+
# Local machine paths, set per install from the README
|
|
27
|
+
.env
|
|
28
|
+
models/
|
|
@@ -0,0 +1,334 @@
|
|
|
1
|
+
# Findings
|
|
2
|
+
|
|
3
|
+
What changed, why, and what was measured. Every number here came from a script in
|
|
4
|
+
`tests-AB/`, on an RTX 3090 Ti. Reproduce any row by running the script named beside it.
|
|
5
|
+
|
|
6
|
+
## The pipeline every finding attaches to
|
|
7
|
+
|
|
8
|
+
```
|
|
9
|
+
request JSON
|
|
10
|
+
→ parse validate, lift any image out of the state
|
|
11
|
+
→ label give each option one letter, A to Z then a to z
|
|
12
|
+
→ assemble build the shared prefix once, one suffix per branch
|
|
13
|
+
→ prefill append "Answer: (" to the rendered chat template
|
|
14
|
+
→ score encode the prefix once, broadcast its cache, one batched forward
|
|
15
|
+
→ read gather the logits at the candidate letter ids
|
|
16
|
+
→ calibrate subtract the learned label prior, divide by the fitted temperature
|
|
17
|
+
→ normalize softmax over those letters only, in float64
|
|
18
|
+
→ map letters back to option names
|
|
19
|
+
response JSON
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
A branch is one prompt whose last token position carries a distribution. A question needs
|
|
23
|
+
one branch, or several when it has levels or more than 52 options.
|
|
24
|
+
|
|
25
|
+
## Model
|
|
26
|
+
|
|
27
|
+
| | Qwen3-4B-Instruct-2507 | Qwen3-VL-4B-Instruct |
|
|
28
|
+
|---|---|---|
|
|
29
|
+
| Banking77, 462 held out rows | 0.554 | **0.626** |
|
|
30
|
+
| eval set, 39 items | 0.897 | **0.949** |
|
|
31
|
+
| eval set, score questions | 0.750 | **1.000** |
|
|
32
|
+
| eval set, noul questions | **1.000** | 0.917 |
|
|
33
|
+
| median latency | 171 ms | **169 ms** |
|
|
34
|
+
| reads images | no | **yes** |
|
|
35
|
+
|
|
36
|
+
The vision model is better at text as well as capable of images, at the same speed, so it
|
|
37
|
+
is the default. `ab_temperature.py`, `evaluate.py`.
|
|
38
|
+
|
|
39
|
+
Noul is the one regression, down 8 points, and noul is what image tagging uses most.
|
|
40
|
+
|
|
41
|
+
## Temperature belongs to the model
|
|
42
|
+
|
|
43
|
+
Fitted on the held out half of a 924 row Banking77 sample, chosen on calibration error.
|
|
44
|
+
|
|
45
|
+
| Model | fitted T | ECE at its own T | ECE at the other model's T |
|
|
46
|
+
|---|---|---|---|
|
|
47
|
+
| Qwen3-VL-4B-Instruct | 1.25 | 0.088 | 0.565 |
|
|
48
|
+
| Qwen3-4B-Instruct-2507 | 6.0 | 0.029 | 0.088 |
|
|
49
|
+
|
|
50
|
+
Nearly a 5x difference between two models of the same size and family, so a single constant
|
|
51
|
+
was wrong. `config.py` holds `FITTED_TEMPERATURES`, and an unfitted model gets 2.5.
|
|
52
|
+
|
|
53
|
+
The fit follows calibration error, not accuracy. Accuracy climbs with temperature above 52
|
|
54
|
+
options, but that is the split path defect below, not a gain. `ab_temperature.py`.
|
|
55
|
+
|
|
56
|
+
## The prior corrects letters, not fixed positions
|
|
57
|
+
|
|
58
|
+
The running prior subtracts the model's standing preference for one letter over another.
|
|
59
|
+
Two positions carry a fixed meaning instead. The escape label always means none of these,
|
|
60
|
+
and a score's letters always name the same rungs. Their mean mass follows the traffic, so
|
|
61
|
+
correcting them moves one request's answer with unrelated requests.
|
|
62
|
+
|
|
63
|
+
Banking77, 462 held out rows, Qwen3-VL-4B-Instruct at its fitted temperature.
|
|
64
|
+
|
|
65
|
+
| Prior | Accuracy | ECE |
|
|
66
|
+
|---|---|---|
|
|
67
|
+
| off | 0.606 | 0.083 |
|
|
68
|
+
| escape corrected, score shares the choice bucket | 0.621 | 0.114 |
|
|
69
|
+
| escape left alone, score in its own bucket | **0.630** | 0.103 |
|
|
70
|
+
|
|
71
|
+
The fix does not regress calibration. The accuracy gain is four rows in 462, near the noise
|
|
72
|
+
of this sample. The prior still costs calibration against no prior, because the shipped
|
|
73
|
+
temperature was fitted with the prior off. `banking77.py`, with and without `--no-prior`.
|
|
74
|
+
|
|
75
|
+
## Option order helps above 52 options and nowhere else
|
|
76
|
+
|
|
77
|
+
| Options | Letterings | Accuracy | ECE |
|
|
78
|
+
|---|---|---|---|
|
|
79
|
+
| 10, one branch | 1 | 0.881 | 0.049 |
|
|
80
|
+
| 10, one branch | 4 | 0.881 | 0.025 |
|
|
81
|
+
| 77, split | 1 | 0.554 | 0.097 |
|
|
82
|
+
| 77, split | 4 | **0.693** | 0.288 |
|
|
83
|
+
|
|
84
|
+
OpenJev ships the same idea as `READOUT_PERMS`, default off, with a code comment claiming
|
|
85
|
+
about 16 points. The gain reproduces, but it is group assignment rather than letter
|
|
86
|
+
position. Ten options fit one branch, where relettering changes nothing measurable.
|
|
87
|
+
|
|
88
|
+
Shipped as `LOGIT_PERMUTATIONS`, default 1. Calibration gets worse as accuracy improves, and
|
|
89
|
+
no temperature between 0.5 and 50 repaired it.
|
|
90
|
+
|
|
91
|
+
## Abstain
|
|
92
|
+
|
|
93
|
+
Every choice question carries a `none of these` label, reported as `abstain` outside
|
|
94
|
+
`probabilities`. Measured against rows whose correct label was deleted from the list, so
|
|
95
|
+
declining was the only right answer.
|
|
96
|
+
|
|
97
|
+
| Setup | AUC |
|
|
98
|
+
|---|---|
|
|
99
|
+
| one branch, 10 options | **0.878** |
|
|
100
|
+
| split path, 77 options | 0.639 |
|
|
101
|
+
|
|
102
|
+
The split path is weak because a group that does not hold the answer declines whether or not
|
|
103
|
+
the answer sits in another group.
|
|
104
|
+
|
|
105
|
+
Two alternatives were tested and rejected. Mass on letters nobody offered reaches AUC 0.806
|
|
106
|
+
alone, but adding it to the escape label drops the pair to 0.850. Combining across groups by
|
|
107
|
+
product scored 0.616 against 0.639 for the weakest decline, and its scale shrinks with the
|
|
108
|
+
group count, so five honest half declines would read as 0.03. The weakest decline ships.
|
|
109
|
+
|
|
110
|
+
Offering the label does not measurably change accuracy. It moved by two to three examples in
|
|
111
|
+
462, inside the noise of that sample. `ab_abstain` measurements, `ab_temperature.py`.
|
|
112
|
+
|
|
113
|
+
## Abstain and confidence answer different questions
|
|
114
|
+
|
|
115
|
+
| What to catch | Signal | AUC |
|
|
116
|
+
|---|---|---|
|
|
117
|
+
| nothing in the list fits | `abstain` | 0.878 |
|
|
118
|
+
| the tag is probably wrong | `confidence` | 0.864 |
|
|
119
|
+
|
|
120
|
+
They correlate at -0.64, so they overlap, but each wins its own job. A tagging pipeline
|
|
121
|
+
should check both.
|
|
122
|
+
|
|
123
|
+
## Several true answers need one noul each
|
|
124
|
+
|
|
125
|
+
| Method | AUC across tiles | AUC within a tile | best F1 |
|
|
126
|
+
|---|---|---|---|
|
|
127
|
+
| one `noul` per fragment | **0.963** | **1.000** | **0.909** |
|
|
128
|
+
| one `choice` over all fragments | 0.667 | 0.685 | 0.516 |
|
|
129
|
+
|
|
130
|
+
A choice is normalized and sums to 1. On a tile holding four fragments it gave the winner
|
|
131
|
+
1.00 and the other three 0.00. The published result that symbol scoring beats independent
|
|
132
|
+
scoring by 9.7 points covers single label tasks, and does not transfer here.
|
|
133
|
+
`ab_multi_label.py`.
|
|
134
|
+
|
|
135
|
+
## The wording of a boolean does not matter
|
|
136
|
+
|
|
137
|
+
Yes and no against true and false, correct and incorrect, present and absent, visible and
|
|
138
|
+
hidden, agree and disagree. AUC spanned 0.946 to 0.957 on 462 balanced text statements and
|
|
139
|
+
0.960 to 0.969 on the tiles. That is inside noise at these sample sizes.
|
|
140
|
+
|
|
141
|
+
The model writes "no" more readily than "yes" when generating. The readout never asks it to
|
|
142
|
+
write anything, so both words are equally reachable at the one position we read.
|
|
143
|
+
`ab_noul_wording.py`.
|
|
144
|
+
|
|
145
|
+
## Confidence formulas
|
|
146
|
+
|
|
147
|
+
TypeSafe publishes two, and we were using the choice one for score questions. A score is
|
|
148
|
+
ordinal, so `[0, 0.5, 0.5, 0, 0]` and `[0.5, 0, 0, 0, 0.5]` are different answers that the
|
|
149
|
+
peak height formula scores identically at 0.375. The official formula gives 0.583 and 0.0.
|
|
150
|
+
|
|
151
|
+
Verified by reading `system_one_adapter._utils.confidence_metrics` from TypeSafe's own PyPI
|
|
152
|
+
package, not by trusting a third party's description of it.
|
|
153
|
+
|
|
154
|
+
## Which torch globals decide bit-exactness
|
|
155
|
+
|
|
156
|
+
`backends/hf.py` holds a set of process-global torch settings around each forward pass. On
|
|
157
|
+
the CUDA attention path, only one of them changes a logit on either shipped model.
|
|
158
|
+
|
|
159
|
+
| setting flipped away from pinned | Qwen3-VL-4B-Instruct | Qwen3-4B-Instruct-2507 |
|
|
160
|
+
|---|---|---|
|
|
161
|
+
| `cudnn.benchmark` | 0.0 | 0.0 |
|
|
162
|
+
| `cudnn.deterministic` | 0.0 | 0.0 |
|
|
163
|
+
| `allow_bf16_reduced_precision_reduction` | **0.5** | **0.75** |
|
|
164
|
+
| float32 matmul precision | 0.0 | 0.0 |
|
|
165
|
+
|
|
166
|
+
Max absolute logit shift against the fully pinned run, window removed so the process
|
|
167
|
+
globals reach the forward pass. The forward runs in bfloat16, so the fp32 matmul setting
|
|
168
|
+
governs matmuls this graph does not have, and the cudnn pair governs convolutions that
|
|
169
|
+
only the vision tower's patch embed uses. The three inert settings are kept because
|
|
170
|
+
`cudnn.benchmark` picks convolution algorithms by timing, which is specific to the card
|
|
171
|
+
and the model, and this was measured on one card with two models.
|
|
172
|
+
|
|
173
|
+
Scoping the settings to the forward pass moved nothing. The same request returns the same
|
|
174
|
+
logits, and a run under a hostile host state matches a run under a pinned one, both at a
|
|
175
|
+
gap of 0.0 on both models.
|
|
176
|
+
|
|
177
|
+
`allow_fp16_bf16_reduction_math_sdp` is held for a different reason and the sweep above
|
|
178
|
+
could not reach it. `sdpa_kernel` turns the math backend **off** while it holds the CUDA
|
|
179
|
+
backends, so this setting is live only when `_usable_attention_backends` returns nothing and
|
|
180
|
+
no `sdpa_kernel` block is entered. ComfyUI turns it **on** at import, at
|
|
181
|
+
`comfy/model_management.py:569`, so inside ComfyUI that path would run with reduced
|
|
182
|
+
precision reductions.
|
|
183
|
+
|
|
184
|
+
Forcing the math backend measures it on the CPU, with no card or weights.
|
|
185
|
+
|
|
186
|
+
| bfloat16 attention, math backend | bitwise equal | max abs diff |
|
|
187
|
+
|---|---|---|
|
|
188
|
+
| `(1, 4, 256, 128)` | no | 0.00977 |
|
|
189
|
+
| `(1, 4, 1024, 128)` | no | 0.00391 |
|
|
190
|
+
| `(8, 4, 256, 128)` | no | 0.00781 |
|
|
191
|
+
|
|
192
|
+
That is far above a low bit, so the math path needs the setting held.
|
|
193
|
+
`ab_math_sdp_reduction.py`.
|
|
194
|
+
|
|
195
|
+
Three more settings were considered and left to the host. `cudnn.enabled` is off on hosts
|
|
196
|
+
where cuDNN does not work, so forcing it on would break the machine that turned it off.
|
|
197
|
+
`cudnn.conv.fp32_precision` and `cudnn.allow_tf32` govern fp32 convolutions, and the
|
|
198
|
+
forward runs in bfloat16. `torch.use_deterministic_algorithms` moves toward determinism, so
|
|
199
|
+
a host that sets it costs us nothing.
|
|
200
|
+
|
|
201
|
+
When the probe finds no usable kernel the pass pins flash, cudnn, mem efficient and math in
|
|
202
|
+
that order instead of leaving the choice to the host's enable flags. Without that pin, two
|
|
203
|
+
hosts could dispatch one request to two different kernels.
|
|
204
|
+
|
|
205
|
+
`set_float32_matmul_precision` writes the cuda and the mkldnn matmul slots together, and
|
|
206
|
+
its getter reports both `none` and `ieee` as `highest`. So restoring through that setter
|
|
207
|
+
alone leaves the mkldnn slot changed, which is why both slots are saved raw.
|
|
208
|
+
|
|
209
|
+
`ab_determinism_scope.py`.
|
|
210
|
+
|
|
211
|
+
## Packing branches by suffix length
|
|
212
|
+
|
|
213
|
+
Every row in a chunk is left-padded to that chunk's longest suffix, and each padded
|
|
214
|
+
token costs a full forward plus attention over the whole prefix. Taking branches in the
|
|
215
|
+
order the caller sent them drags short branches to the longest width.
|
|
216
|
+
|
|
217
|
+
| Request shape | ms | peak GB | suffix tokens | chunks | |
|
|
218
|
+
|---|---|---|---|---|---|
|
|
219
|
+
| mixed, four choices plus 12 nouls plus 4 scores | 2055.6 | 13.66 | 14600 | 1 | request order |
|
|
220
|
+
| | **623.9** | **10.66** | **2614** | 3 | packed |
|
|
221
|
+
| one 45 option choice beside 20 nouls | 2152.9 | 13.92 | 15330 | 1 | request order |
|
|
222
|
+
| | **425.2** | **10.94** | **1310** | 2 | packed |
|
|
223
|
+
| 256 nouls at the Jev cap | 1370 | 8.75 | 7584 | 8 | both, bitwise identical |
|
|
224
|
+
| one 77 option choice, split | 272 | 8.61 | 1282 | 1 | both, bitwise identical |
|
|
225
|
+
| 24 tile fragments | 198 | 8.63 | 696 | 1 | both, bitwise identical |
|
|
226
|
+
| 30 three option choices | 396 | 8.99 | 2160 | 1 | both, bitwise identical |
|
|
227
|
+
| 12 twenty option choices | 586 | 9.24 | 3960 | 1 | both, bitwise identical |
|
|
228
|
+
|
|
229
|
+
3.29x and 5.06x on the two shapes that mix widths, each giving back about 3 GB. Nothing
|
|
230
|
+
at all on the five that do not, which is the point. The sort is stable, so equal lengths
|
|
231
|
+
keep the caller's order and the grouping stays a pure function of the request.
|
|
232
|
+
|
|
233
|
+
The two mixed shapes move low bits, by at most 0.375 and 0.5. Changing the padded width
|
|
234
|
+
or the batch composition of a bf16 forward does that. The shapes whose bits move are the
|
|
235
|
+
shapes that were three to five times too slow.
|
|
236
|
+
|
|
237
|
+
**The token ceiling is checked only when a branch widens the chunk it joins.** A request
|
|
238
|
+
whose branches are all one width has no padding to remove, so the ceiling would only add
|
|
239
|
+
chunks. Checking it unconditionally re-chunked any uniform request above 64 tokens of
|
|
240
|
+
width, which is every choice question, and moved its logits for no gain.
|
|
241
|
+
|
|
242
|
+
| ceiling | mixed shape, 800 token prefix | mixed widths, long prefix |
|
|
243
|
+
|---|---|---|
|
|
244
|
+
| 512 | 610.9 ms, 9.91 GB, 5 chunks | 2015.7 ms, 5 chunks |
|
|
245
|
+
| 1024 | 608.3 ms, 10.47 GB, 3 chunks | 2014.9 ms, 5 chunks |
|
|
246
|
+
| **2048** | **626.0 ms, 10.66 GB, 3 chunks** | **2015.8 ms, 5 chunks** |
|
|
247
|
+
| 4096 | 793.5 ms, 11.03 GB, 2 chunks | 2286.5 ms, 4 chunks |
|
|
248
|
+
| 8192 | 1119.4 ms, 11.84 GB, 2 chunks | 2289.5 ms, 4 chunks |
|
|
249
|
+
| none | 2062.2 ms, 13.66 GB, 1 chunk | 2299.1 ms, 4 chunks |
|
|
250
|
+
|
|
251
|
+
Anything from 512 to 2048 measured the same on both, and 4096 upward is clearly worse.
|
|
252
|
+
The value is not a knife edge. 2048 sits in the middle of the flat band and leaves the
|
|
253
|
+
most room before two near-equal wide branches are split apart.
|
|
254
|
+
|
|
255
|
+
`ab_branch_packing.py`.
|
|
256
|
+
|
|
257
|
+
## Throughput and where it falls off
|
|
258
|
+
|
|
259
|
+
Warm model, RTX 3090 Ti. Adding questions to a record is close to free up to about 8,
|
|
260
|
+
because the fixed cost of two forward passes dominates a short record.
|
|
261
|
+
|
|
262
|
+
| Questions per record | Total ms | ms per tag | Records per hour |
|
|
263
|
+
|---|---|---|---|
|
|
264
|
+
| 1 | 138 | 138 | 26,100 |
|
|
265
|
+
| 4 | 136 | 34 | 26,500 |
|
|
266
|
+
| 8 | 140 | 17.5 | 25,800 |
|
|
267
|
+
| 32 | 278 | 8.7 | 13,000 |
|
|
268
|
+
|
|
269
|
+
The state is prefilled once per record, so a longer record costs more.
|
|
270
|
+
|
|
271
|
+
| State tokens | Total ms, 8 questions | ms per tag | Records per hour |
|
|
272
|
+
|---|---|---|---|
|
|
273
|
+
| 172 | 140 | 17.5 | 25,700 |
|
|
274
|
+
| 556 | 185 | 23 | 19,500 |
|
|
275
|
+
| 2,092 | 451 | 56 | 8,000 |
|
|
276
|
+
| 8,236 | 1,575 | 197 | 2,300 |
|
|
277
|
+
| 32,812 | 41,515 | 5,189 | 87 |
|
|
278
|
+
|
|
279
|
+
Prefill runs at 46 to 48 TFLOP/s from about 2,000 tokens up, near this card's practical
|
|
280
|
+
BF16 ceiling. Below that the time is almost all kernel launch rather than arithmetic, so a
|
|
281
|
+
428 token request spends 0.14 ms on the math. The last row falls off the ceiling for the
|
|
282
|
+
reason under Known weaknesses. `benchmark.py`.
|
|
283
|
+
|
|
284
|
+
## The attention backend has to be pinned on Windows
|
|
285
|
+
|
|
286
|
+
Torch 2.11 ships no FlashAttention kernel for Windows. Left to choose for itself the
|
|
287
|
+
dispatcher reaches the math backend, which builds the full attention matrix. Measured on an
|
|
288
|
+
8,192 token prefill.
|
|
289
|
+
|
|
290
|
+
| Backend selection | Time | Peak memory |
|
|
291
|
+
|---|---|---|
|
|
292
|
+
| left to the dispatcher | 147,221 ms | 27.6 GB |
|
|
293
|
+
| pinned to cudnn and efficient | 1,291 ms | 9.2 GB |
|
|
294
|
+
|
|
295
|
+
`HFBackend` probes the backends at startup, at the model's own dtype, and pins the working
|
|
296
|
+
ones. `benchmark.py`.
|
|
297
|
+
|
|
298
|
+
## Bugs found and fixed
|
|
299
|
+
|
|
300
|
+
| Bug | Effect | Fix |
|
|
301
|
+
|---|---|---|
|
|
302
|
+
| score used the choice confidence formula | wrong confidence on every score answer | ported both official formulas |
|
|
303
|
+
| attention probe ran in bfloat16 whatever the model dtype | any other dtype crashed on the first forward | probe at the model's dtype |
|
|
304
|
+
| attention probe used one call shape | approved a backend that then failed | probe causal and masked, with grouped query attention |
|
|
305
|
+
| abstain combined groups by product | shrinks with group count, never fires at scale | weakest decline |
|
|
306
|
+
| image source detected by string length | an 8x8 PNG encodes shorter than the threshold and was read as a path | check whether the file exists |
|
|
307
|
+
| M-RoPE offset dropped on branch rows | Qwen3-VL raises rather than misplacing tokens | read `rope_deltas` after the prefix pass |
|
|
308
|
+
| prior learned and corrected the escape label | traffic that often declines taught the prior to undo a decline | prior over the lettered options only |
|
|
309
|
+
| a joint score shared a prior bucket with a choice of the same width | a score's top rung learned the choice traffic's escape mass | score has its own bucket |
|
|
310
|
+
|
|
311
|
+
## Known weaknesses
|
|
312
|
+
|
|
313
|
+
**The split path.** Three defects share one cause, that groups above 52 options cannot see
|
|
314
|
+
each other. Group assignment bias costs 13.9 points, temperature leaks into accuracy, and
|
|
315
|
+
the abstain score falls from 0.878 to 0.639. Parked by choice. Candidates are OpenJev's
|
|
316
|
+
chunk winners tournament and a second pass over each group's winner. Neither is measured.
|
|
317
|
+
|
|
318
|
+
**Noul temperature is unfitted.** The shipped 1.25 was fitted on choice questions. Booleans
|
|
319
|
+
now saturate near 0 and 1, so the best threshold sits at 0.9999 rather than near 0.5.
|
|
320
|
+
|
|
321
|
+
**Noul regressed 8 points** moving to the vision model, and image tagging leans on noul.
|
|
322
|
+
|
|
323
|
+
**Very long states fall off the prefill ceiling.** At 33,068 prefilled tokens the card runs
|
|
324
|
+
at 7.1 TFLOP/s against 47.9 at 8,492. The prefix key value cache is about 4.9 GB at that
|
|
325
|
+
length and every chunk copies it, which leaves a 24 GB card no room. Expanding the cache
|
|
326
|
+
once and cropping between chunks was measured and is worse, at 12,732 ms against 1,710 ms
|
|
327
|
+
and 23.66 GB against 15.12 GB, because a cropped cache is a non-contiguous view. Unsolved.
|
|
328
|
+
`benchmark.py`.
|
|
329
|
+
|
|
330
|
+
## Layout
|
|
331
|
+
|
|
332
|
+
`tests/` holds the pytest suite, 146 unit tests and 14 that need the GPU. `tests-AB/` holds
|
|
333
|
+
the measurement harnesses, matching ComfyUI-ContextAnchoredTileRefine. Fixtures live in
|
|
334
|
+
`tests/fixtures/` and `tests-AB/ab_env.py` points at them.
|