poseaudit 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,8 @@
1
+ .venv/
2
+ __pycache__/
3
+ *.egg-info/
4
+ dist/
5
+ .pytest_cache/
6
+ .ruff_cache/
7
+ examples/coco_elbow/data/
8
+ .hypothesis/
@@ -0,0 +1,22 @@
1
+ # Changelog
2
+
3
+ ## 0.1.0
4
+
5
+ First release.
6
+
7
+ - Measures: tilt, angle, length and ratio, optionally relative to the rest of
8
+ each image (median or a percentile, of values or of their absolute values,
9
+ leaving each reading out of its own baseline).
10
+ - Inputs: COCO annotations and results, Ultralytics YOLO pose, supervision
11
+ `sv.KeyPoints` (with `sv.Detections` for boxes), or arrays; class filters
12
+ and class-aware matching.
13
+ - Agreement: error, large-error rate, bias, percentile and normal limits of
14
+ agreement, repeated-measures limits, ICC(A,1), CCC, Pearson r, error by size
15
+ of the measured part.
16
+ - Gain and squashing: least-squares and robust (Theil-Sen) gain, Bland-Altman
17
+ and Deming slopes, and the gain keypoint jitter alone gives, with the gap's
18
+ interval and a one-sided p(squash).
19
+ - Decisions at thresholds, with a separate or count-matched predicted
20
+ threshold.
21
+ - Intervals resample whole images or named clusters.
22
+ - Output: summary, markdown report, CSV, JSON and a two-panel plot.
@@ -0,0 +1,8 @@
1
+ cff-version: 1.2.0
2
+ message: "If you use poseaudit in your work, please cite it as below."
3
+ title: "poseaudit: auditing pose models used as measuring instruments"
4
+ authors:
5
+ - alias: 8rulerstar
6
+ license: MIT
7
+ repository-code: "https://github.com/8rulerstar/poseaudit"
8
+ version: 0.1.0
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 8rulerstar
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,487 @@
1
+ Metadata-Version: 2.5
2
+ Name: poseaudit
3
+ Version: 0.1.0
4
+ Summary: Audit a pose model used as a measuring instrument: how wrong are the angles, tilts and lengths it reads off keypoints?
5
+ Project-URL: Homepage, https://github.com/8rulerstar/poseaudit
6
+ Author: 8rulerstar
7
+ License-Expression: MIT
8
+ License-File: LICENSE
9
+ Keywords: computer vision,evaluation,keypoints,measurement,pose estimation
10
+ Classifier: Development Status :: 3 - Alpha
11
+ Classifier: Intended Audience :: Developers
12
+ Classifier: Intended Audience :: Science/Research
13
+ Classifier: Operating System :: OS Independent
14
+ Classifier: Programming Language :: Python :: 3
15
+ Classifier: Programming Language :: Python :: 3.10
16
+ Classifier: Programming Language :: Python :: 3.11
17
+ Classifier: Programming Language :: Python :: 3.12
18
+ Classifier: Programming Language :: Python :: 3.13
19
+ Classifier: Topic :: Scientific/Engineering :: Image Recognition
20
+ Classifier: Typing :: Typed
21
+ Requires-Python: >=3.10
22
+ Requires-Dist: numpy>=1.23
23
+ Provides-Extra: images
24
+ Requires-Dist: pillow>=12.3; extra == 'images'
25
+ Provides-Extra: plot
26
+ Requires-Dist: matplotlib>=3.6; extra == 'plot'
27
+ Provides-Extra: supervision
28
+ Requires-Dist: supervision>=0.30.6; extra == 'supervision'
29
+ Description-Content-Type: text/markdown
30
+
31
+ # poseaudit
32
+
33
+ **Your pose model has a good mAP. How far off are the angles it measures?**
34
+
35
+ When keypoints are used to *measure* something (a joint angle, a tilt, a length),
36
+ position metrics such as OKS, PCK or pixel error do not say how far off the
37
+ measurement is. `poseaudit` reads the measurement off the ground truth and off
38
+ the prediction and reports the difference the way a measuring instrument is
39
+ judged: how large, how often large, where, biased which way, and how sure you
40
+ can be of each figure. The statistics are the familiar agreement ones (Bland-Altman
41
+ limits, ICC, CCC, Deming and Theil-Sen slopes); what `poseaudit` adds is the
42
+ path to them from the formats computer vision evaluates in (COCO, YOLO,
43
+ supervision): matching, counting what could not be read and why, intervals
44
+ that respect images or subjects, and a check of whether a slope below 1 is
45
+ more than keypoint jitter.
46
+
47
+ ```text
48
+ $ poseaudit audit --format coco --gt gt_200.json --pred pred_yolo11n.json \
49
+ --angle 5,7,9 --big-error 15 --size-bands 30,60
50
+ angle (5, 7, 9): read 322 of 380 labelled instances
51
+ not read no matching prediction 35, prediction lacked a point 23, unmeasurable 0
52
+ not counted 206 with a point unlabelled in the truth, 79 unmatched predictions
53
+ |error| mean 19.55° [17.13 to 22.16], median 12.64°, 95th pct 68.99°
54
+ >= 15° 43.5% [38.2% to 48.9%]
55
+ by size 0-30 px 58% (n 134), 30-60 px 37% (n 112), 60+ px 28% (n 76)
56
+ bias +2.34° [-1.16 to +6.22], median +0.52°
57
+ limits -50.90° to +79.99° (2.5th to 97.5th percentile of errors)
58
+ gain 0.731 [0.639 to 0.814]; robust 0.784
59
+ vs jitter 0.871 from keypoint jitter alone; gap -0.140 [-0.199 to -0.077], p(squash) <= 0.002*
60
+ * swaps and gross failures lower the gain too, and noisy
61
+ labels make p small for an honest model: read the gap
62
+ ! 2.5% of errors fall below the normal limits and 4.3% above them, against 2.5% each for a normal error: use the percentile limits.
63
+ ```
64
+
65
+ That is the left elbow angle (COCO keypoints 5, 7, 9: shoulder, elbow, wrist;
66
+ indices count from 0) of `yolo11n-pose` against COCO's human labels on 200
67
+ val2017 images. The data is in
68
+ [`examples/coco_elbow`](https://github.com/8rulerstar/poseaudit/tree/main/examples/coco_elbow);
69
+ run the command from that folder.
70
+
71
+ These figures are disagreement between the model and **one human label**, not
72
+ the model's error against the world. COCO's own annotators disagree: placing
73
+ the three points with the spread COCO publishes for repeated labels (its OKS
74
+ sigmas) moves the elbow angle by 11 to 16° on average on these same arms,
75
+ depending on how those sigmas are read (a rough estimate that treats each
76
+ label's points as independent; [`validation/label_noise.py`](https://github.com/8rulerstar/poseaudit/blob/main/validation/label_noise.py)).
77
+ A good part of the 19.55° may be the labels'. With a reference better than
78
+ the model (motion capture, careful relabelling), the same figures describe
79
+ the model.
80
+
81
+ What a few lines of NumPy would miss:
82
+
83
+ - **What was not read, and why.** Unlabelled points, unmatched people and
84
+ missing predicted points are counted apart, so a model that skips the hard
85
+ cases does not look better for it.
86
+ - **Intervals that respect the data.** Readings from one image are resampled
87
+ together, and with `--cluster` so are the frames of one video or the photos
88
+ of one subject (left out, frames count as independent).
89
+ - **The sorting trap.** When the labels are noisy, bias in bands sorted by the
90
+ truth leans toward squashing even for an unbiased model, the more so the
91
+ noisier they are ([below](#which-way-you-sort-decides-the-story)).
92
+ - **Decisions.** Whether the prediction lands on the same side of a threshold
93
+ as the truth, which is what a pass or fail rule depends on.
94
+
95
+ ## Install
96
+
97
+ ```bash
98
+ pip install poseaudit # needs only numpy
99
+ pip install "poseaudit[plot]" # plots (matplotlib)
100
+ pip install "poseaudit[images]" # --images: read image sizes (Pillow)
101
+ pip install "poseaudit[supervision]" # from_supervision
102
+ ```
103
+
104
+ ## Quick start
105
+
106
+ ```bash
107
+ poseaudit audit --format coco --gt annotations.json --pred results.json \
108
+ --min-conf 0.5 --angle 5,7,9 --big-error 15 \
109
+ --report report.md --plot panels.png --csv readings.csv
110
+ ```
111
+
112
+ `--big-error` is required: the error, in the measure's unit, that counts as
113
+ large. `report.md` explains every line of the summary. In Python:
114
+
115
+ ```python
116
+ import poseaudit as pa
117
+
118
+ truth = pa.load_coco("person_keypoints_val2017.json")
119
+ predicted = pa.load_coco_results(
120
+ "results.json", "person_keypoints_val2017.json", min_confidence=0.5
121
+ )
122
+
123
+ result = pa.audit(pa.pair(truth, predicted), pa.angle(5, 7, 9), big_error=15)
124
+ print(result.summary()) # summary(full=True) adds every agreement statistic
125
+ result.to_markdown("report.md") # sizes, bands, decisions, largest errors
126
+ result.to_csv("readings.csv") # one row per reading, to recompute anything
127
+ result.plot("panels.png") # predicted against true, and a Bland-Altman plot
128
+ ```
129
+
130
+ From your own arrays: `xy` of shape (K, 2) in pixel coordinates (not 0 to 1)
131
+ and a boolean `visible` of shape (K,) per object, here a 33-point skeleton:
132
+
133
+ ```python
134
+ import numpy as np
135
+
136
+ truth = {"frame_001": [pa.Instance.from_keypoints(xy_true, np.ones(33, bool))]}
137
+ predicted = {"frame_001": [pa.Instance.from_keypoints(xy_pred, visible_pred)]}
138
+ result = pa.audit(pa.pair(truth, predicted), pa.angle(23, 25, 27), big_error=10)
139
+ ```
140
+
141
+ ## What the demo shows
142
+
143
+ **Disagreement grows as the arms get smaller.** When the arm segments average
144
+ under 30 px, 58% of elbows are off by 15° or more; above 60 px, 28%. The
145
+ signed bias of +2° hides all of this, and so does a single mAP. The link with
146
+ size is an association; size is not the whole cause. Some of it is geometry:
147
+ a pixel of error turns a short segment further than a long one, for the label
148
+ as much as for the model, and an arm pointing toward the camera looks short
149
+ and is hard to read even on a large person. Small people in COCO are also
150
+ more often occluded, blurred and loosely labelled.
151
+
152
+ | arm segment | n | mean abs error | off by 15° or more | gain (no jitter reference per band) |
153
+ |---|---|---|---|---|
154
+ | 0 to 30 px | 134 | 25.55° | 58.2% [49.7% to 66.2%] | 0.526 |
155
+ | 30 to 60 px | 112 | 16.57° | 36.6% [28.3% to 45.8%] | 0.791 |
156
+ | 60+ px | 76 | 13.36° | 27.6% [18.8% to 38.6%] | 0.911 |
157
+
158
+ ![Left: predicted against true angle. Right: Bland-Altman plot](https://raw.githubusercontent.com/8rulerstar/poseaudit/main/docs/panels.png)
159
+
160
+ ## Which way you sort decides the story
161
+
162
+ The natural next check is the bias in bands of angle. Sort the same readings
163
+ three ways and you get three stories:
164
+
165
+ ![Mean error per band when sorting by the truth, the prediction, or their mean](https://raw.githubusercontent.com/8rulerstar/poseaudit/main/docs/trap.png)
166
+
167
+ Whichever reading you sort by, its own noise puts its extreme values in the
168
+ extreme bands, where they read as bias (regression to the mean). Trust the
169
+ sort by the truth when the truth is much less noisy than the model (the
170
+ default, `--band-by truth`); sort by the mean of both (`--band-by mean`) and
171
+ read the Bland-Altman slope when the two are about equally noisy. Under the
172
+ wrong assumption, either view misleads: with a model much noisier than its
173
+ labels, the Bland-Altman slope calls an unbiased model expanding.
174
+
175
+ ## Options
176
+
177
+ ### Decisions at a threshold
178
+
179
+ When a measurement feeds a rule (flag a tilt of 10° or more, an elbow bent past
180
+ 90°), what matters is whether the prediction lands on the same side of it. The
181
+ "truth" side is the same rule applied to the truth keypoints, not a defect
182
+ label: an inspector's pass or fail verdict can differ from any angle rule. A
183
+ model that reads angles differently may need its own threshold: give one with
184
+ `--pred-threshold`, and `poseaudit` also finds the one that flags as many
185
+ instances as the truth does. That one is fitted on the data it is scored on,
186
+ so its figures are optimistic until checked on other data. Only readings
187
+ count: a labelled instance that was not read is neither caught nor missed.
188
+
189
+ ```text
190
+ $ poseaudit audit ... --threshold 90 --side below --pred-threshold 100
191
+ below 90°: caught 53 of 84 (63% [52% to 73%]), false alarms 11 of 238
192
+ below 90°, predicted 100°: caught 63 of 84 (75% [65% to 83%]), false alarms 23 of 238
193
+ below 90°, predicted 98.40° (same count, fitted here): caught 61 of 84 (73% [62% to 81%]), false alarms 23 of 238
194
+ ```
195
+
196
+ ### Tilts relative to the rest of the photo
197
+
198
+ If a rule compares each object with the others in its photo, a tilted camera
199
+ moves every tilt together. `--relative-to median` reads each tilt against the
200
+ median tilt of the other objects in the image, on the truth and the prediction
201
+ separately, so a camera roll drops out. `--relative-to p20 --relative-abs`
202
+ reads how much more each object leans than the 20th percentile of the others'
203
+ absolute tilts. Each object's baseline leaves the object itself out, so none
204
+ is scored against itself. With `--relative-abs` a value below 0 means
205
+ straighter than the baseline, so band edges should start below 0. This is for
206
+ tilts; an angle inside a limb does not change when the camera rolls, and the
207
+ report says so.
208
+
209
+ ### Several readings of one subject
210
+
211
+ Frames of one video, or photos of one subject, are not independent. Name the
212
+ group with `--cluster REGEX` (its first capture group in the image name) or
213
+ `cluster=` in Python, for example `cluster=lambda name: name.split("_")[0]`
214
+ for names such as `athlete3_trial2_f0041`. Intervals then resample whole
215
+ groups, and limits of agreement for repeated readings (Bland and Altman 2007)
216
+ are added. Leave the groups out and every frame counts as independent: the
217
+ intervals come out far too narrow, and nothing warns about it.
218
+
219
+ By default the group is the image. With fewer than 20 groups the intervals are
220
+ too narrow, and the report says so; nested groups (subject, trial, frame) are
221
+ not supported. With named groups the jitter reference gives no p: it treats
222
+ readings as independent, and when a subject carries the same error from
223
+ reading to reading it would flag honest models as squashed far more often than
224
+ it says. Read the gap's interval, which resamples whole groups.
225
+
226
+ ## Does the model squash large angles?
227
+
228
+ The **gain** is the least-squares slope of the predicted angle on the true one:
229
+ 0.73 means that, on average over these readings, a 10° difference comes out as
230
+ about 7°. Keypoint jitter alone pulls a gain below 1 too: a straight arm can
231
+ only be read as more bent, a folded one only as more open. The **jitter
232
+ reference** (`vs jitter`) is the gain of predictions rebuilt from the truth
233
+ plus this model's own point displacements ([how](#the-jitter-reference)):
234
+ 0.87 here. The model's gain is 0.140 lower (interval 0.077 to 0.199, paired
235
+ within resamples). With the default seed none of the 500 rebuilds comes out
236
+ as low as the model's (`p(squash) <= 0.002`); with some other seeds one does
237
+ (0.004).
238
+ The model reads angle differences as smaller than its own scatter explains.
239
+ What that is, the check cannot say:
240
+
241
+ - **Gross failures and swaps.** Gross failures tied to the true angle (a
242
+ straight arm read as folded) and left and right swapped on one side lower
243
+ the gain just as squashing does; failures in random directions are part of
244
+ the rebuilds and do not widen the gap. Leaving out the 31 readings off by
245
+ 45° or more leaves a gap of -0.069 [-0.109 to -0.028] (about half the gap),
246
+ and leaving out the 61 off by 30° or more, -0.038 [-0.073 to -0.003]. Those
247
+ subsets are chosen by the outcome, so they describe where the gap comes from
248
+ rather than test it, and trimming by the outcome changes a gap by itself, so
249
+ what is left is not a clean measure of squashing either. Some of the largest errors may be
250
+ the labels' (a point on the wrong limb), and the 58 instances not read may
251
+ not be a random part of the rest.
252
+ - **Noise in the truth.** The rebuild takes its displacements against the
253
+ labels, so label noise is already in its scatter. On these same arms, honest
254
+ models with labels as noisy as COCO's annotators fell 0.02 to 0.04 below the
255
+ reference, far less than 0.140
256
+ ([`validation/label_noise.py`](https://github.com/8rulerstar/poseaudit/blob/main/validation/label_noise.py), with independent
257
+ normal noise on each point). Under that model label noise does not explain
258
+ this gap, but it does trip the p: those honest models got
259
+ `p(squash) <= 0.05` in 17 to 40% of 30 runs each. With noisy labels, read
260
+ the size of the gap, not the p.
261
+ - **Squashing** is what remains, and it is not separated from the two above.
262
+ The slopes that suit equal noise, the Bland-Altman slope (-0.049 [-0.119 to
263
+ +0.017]) and Deming at a noise ratio of 1 (0.945 [0.870 to 1.019]), include
264
+ no squashing, but do not rule out a small one. Which to trust
265
+ depends on the noise ratio, which this data does not pin down. Deming for
266
+ other ratios (`--noise-ratio`, the variance of prediction noise over that
267
+ of label noise):
268
+
269
+ | noise ratio | 1 | 4 | 10 |
270
+ |---|---|---|---|
271
+ | Deming slope | 0.945 [0.870 to 1.019] | 0.797 [0.708 to 0.878] | 0.758 [0.667 to 0.842] |
272
+
273
+ On these arms honest models give a Deming slope at a ratio of 1 near 1
274
+ whether the labels are clean or as noisy as the model, with independent
275
+ normal noise and no gross failures
276
+ ([`validation/label_noise.py`](https://github.com/8rulerstar/poseaudit/blob/main/validation/label_noise.py)): there, jitter
277
+ at the ends of the range does not drag it down the way it drags the gain.
278
+
279
+ For this demo, then: a real gap, about half of it in the largest errors.
280
+ Whether the rest is squashing of a few percent or nothing, this data cannot
281
+ tell.
282
+
283
+ A few gross failures move a least-squares slope a lot. The **robust gain**, a
284
+ Theil-Sen slope that they barely move, is 0.78 here against the least-squares
285
+ 0.73. Noise that grows with the angle, or on smaller arms, also separates the
286
+ two, so their difference alone does not say how much comes from gross
287
+ failures, and a gross failure can also be a label on the wrong limb.
288
+
289
+ ## Measures
290
+
291
+ | measure | points | reads |
292
+ |---|---|---|
293
+ | `tilt(a, b)` | 2 | degrees of the axis through a and b from vertical, -90 to 90, positive when the upper end leans right; the axis has no direction, so errors wrap (89° vs -89° is 2° off) and a part read upside down is not an error |
294
+ | `angle(a, b, c)` | 3 | interior angle at `b`, 0 to 180° (180 is straight; a flexion angle is 180 minus this, which flips the sign of the bias); unsigned, so an angle read bending the other way is not an error |
295
+ | `length(a, b)` | 2 | pixels |
296
+ | `ratio(a, b, c, d)` | 4 | \|a-b\| / \|c-d\|, independent of scale |
297
+
298
+ The size of a reading is the mean length of its segments in the truth. There
299
+ is no directed angle (0 to 360) and no fixed axis other than vertical yet;
300
+ `--relative-to` reads tilts against the rest of the image instead.
301
+
302
+ ## Inputs
303
+
304
+ | source | loader |
305
+ |---|---|
306
+ | COCO keypoint annotations | `load_coco(path, classes)` (crowd regions skipped) |
307
+ | COCO results format | `load_coco_results(path, annotations_path, min_confidence, classes)` |
308
+ | Ultralytics YOLO pose labels or predictions | `load_yolo(folder, image_size, num_keypoints, classes, min_confidence)` |
309
+ | `sv.KeyPoints` (supervision 0.30.6+) | `from_supervision(keypoints_by_image, detections_by_image=None)` |
310
+ | anything else | `Instance.from_keypoints(xy, visible)` per object, in a dict of image name to list |
311
+
312
+ COCO annotations with no labelled keypoint (`num_keypoints` 0) and crowd
313
+ regions are skipped as they load, so they appear in no count.
314
+
315
+ - **Bad values.** NaN or infinite coordinates, scores or boxes stop the load
316
+ with the file and the row.
317
+ - **Confidence.** In a results file or a YOLO prediction the third value of a
318
+ point is a confidence. Without a threshold every returned point counts as
319
+ seen, and a warning says so.
320
+ - **Missing points.** A YOLO point written as `0 0` is missing, as in
321
+ supervision.
322
+ - **Pixels.** Angles are measured on the coordinates given. YOLO coordinates
323
+ are fractions of width and height, so an angle measured on them is
324
+ distorted on any non-square image. Give one size, a function of the image
325
+ name, or `--images DIR` to read each image's size. A size given in the
326
+ wrong order (height, width) passes every check and quietly distorts the
327
+ figures (on upright parts it made them look better); `--images` avoids that.
328
+ Arrays passed to `Instance.from_keypoints` must be in pixels too.
329
+ - **Classes.** Instances of different classes are not paired, and an audit
330
+ whose readings mix classes stops with an error: parts of different kinds can
331
+ hide each other's errors. Filter with `--classes` (any format), or allow it
332
+ with `--mixed-classes`. When the two sides number classes differently (COCO
333
+ calls a person 1, Ultralytics 0) they are matched ignoring classes, with a
334
+ warning. A class filter that keeps nothing names the classes that are
335
+ there.
336
+ - **Wrong keypoint counts.** A YOLO load stops on coordinates outside the
337
+ image, which usually means `--keypoints` is wrong, and on a missing folder.
338
+ If every row could also be read as fewer points in x y v form with flags 0,
339
+ 1 and 2, a warning asks you to check `--keypoints`.
340
+ - **Skeletons** of different sizes on the two sides are refused.
341
+ - **supervision.**
342
+ - The container's `visible` mask decides visibility: filter with
343
+ `kp.visible = kp.keypoint_confidence > 0.5`, not with a 2D mask such as
344
+ `kp[mask]`, which drops columns and so changes what each keypoint index
345
+ means.
346
+ - `from_ultralytics` drops the model's boxes; to match by box, pass
347
+ `sv.Detections.from_ultralytics` of the same result, row for row (not
348
+ `kp.as_detections()`, which drops rows with no visible point). Boxes a
349
+ model keeps in `kp.data["xyxy"]` (RF-DETR) are used when no detections
350
+ are given; RF-DETR also fills `visible` with confidence > 0, so set your
351
+ own threshold on it.
352
+ - When both carry the model's scores, rows whose scores rank differently
353
+ stop the load. Without scores only geometry is left: two rows whose
354
+ keypoints clearly fit each other's boxes better than their own give a
355
+ warning, which crowded human labels can also trip.
356
+
357
+ ## How pairs, counts and intervals are made
358
+
359
+ - Predictions are matched to the truth within each image, by image name and
360
+ never by list position (a file extension is ignored when only that differs,
361
+ with a warning). Within an image the most similar pairs are matched first:
362
+ by box IoU (`--min-iou`, default 0.3), or by keypoint closeness when a side
363
+ has no box (`--min-similarity`, default 0.5). Detection scores are ignored
364
+ in matching: drop low-scoring detections first (`--min-score`, or
365
+ `min_score=` in the loaders), or a stray box that overlaps better can take
366
+ the pair.
367
+ - A prediction too far off to reach the matching threshold counts as missed,
368
+ not as an error. When many such misses overlap a prediction that was left
369
+ over, a warning names the threshold to lower.
370
+ - A reading needs every point of the measure in both. The report separates what
371
+ was not read: points **unlabelled in the truth** (not the model's doing, and
372
+ out of the denominator), no matching prediction, a predicted point missing,
373
+ or geometry that cannot be measured.
374
+ - Intervals are percentile bootstraps over whole clusters (by default images),
375
+ 2,000 resamples with a fixed seed. Rates of large errors and of decisions use
376
+ Wilson intervals, which treat readings as independent, unless clusters are
377
+ named; then the overall rate and the decision rates resample whole clusters
378
+ too (never narrower than Wilson's). Rates per size band stay Wilson.
379
+
380
+ ## The jitter reference
381
+
382
+ Each rebuilt arm takes all its point shifts from one arm with a similar true
383
+ angle (one of up to six equal-count groups; possibly itself), scaled to its
384
+ own size. Shifts are expressed along and across each point's own segment, and
385
+ mirrored for arms that bend the other way, so an error along the limb stays
386
+ along it and one toward the inside of the bend stays there. The shift a whole
387
+ group shares is removed first, because that shared shift is the tendency
388
+ being tested. Drawing each point from a different arm instead would break the
389
+ correlation within an arm (a whole limb shifting at once barely changes its
390
+ angle) and set the reference too low.
391
+
392
+ ## Statistics used
393
+
394
+ | figure | definition |
395
+ |---|---|
396
+ | gain | least-squares slope of predicted on truth |
397
+ | robust gain | Theil-Sen slope: median of the slopes between pairs of readings |
398
+ | vs jitter | median gain over 500 rebuilds (see [The jitter reference](#the-jitter-reference)); the gap's interval averages 10 rebuilds per resample; p(squash) = (1 + rebuilds with gain at or below the model's) / 501 |
399
+ | BA slope | slope of the error on the mean of both readings (Bland and Altman 1999) |
400
+ | Deming | slope of predicted on truth with a known ratio of noise variances (Deming 1943; Linnet 1993); the ratio is prediction over label |
401
+ | limits | 2.5th and 97.5th percentiles of the error; normal limits are bias ± 1.96 SD (Bland and Altman 1986) |
402
+ | repeated limits | bias ± 1.96 √(between + within variance) (Bland and Altman 2007) |
403
+ | ICC(A,1) | two-way, absolute agreement, single rating (McGraw and Wong 1996); matches pingouin; its interval is the bootstrap, not the F interval |
404
+ | CCC | Lin's concordance correlation (Lin 1989) |
405
+ | intervals | percentile bootstrap over clusters (Davison and Hinkley 1997); Wilson score intervals for rates without named clusters (Wilson 1927) |
406
+
407
+ The robust gain uses every pair of readings up to about 1,000 readings and
408
+ 500,000 random pairs above that.
409
+
410
+ ### References
411
+
412
+ - Bland JM, Altman DG (1986). Statistical methods for assessing agreement between two methods of clinical measurement. *Lancet* 327(8476):307-310.
413
+ - Bland JM, Altman DG (1999). Measuring agreement in method comparison studies. *Statistical Methods in Medical Research* 8(2):135-160.
414
+ - Bland JM, Altman DG (2007). Agreement between methods of measurement with multiple observations per individual. *Journal of Biopharmaceutical Statistics* 17(4):571-582.
415
+ - Davison AC, Hinkley DV (1997). *Bootstrap Methods and their Application*. Cambridge University Press.
416
+ - Deming WE (1943). *Statistical Adjustment of Data*. Wiley.
417
+ - Lin LI (1989). A concordance correlation coefficient to evaluate reproducibility. *Biometrics* 45(1):255-268.
418
+ - Linnet K (1993). Evaluation of regression procedures for methods comparison studies. *Clinical Chemistry* 39(3):424-432.
419
+ - McGraw KO, Wong SP (1996). Forming inferences about some intraclass correlation coefficients. *Psychological Methods* 1(1):30-46.
420
+ - Sen PK (1968). Estimates of the regression coefficient based on Kendall's tau. *Journal of the American Statistical Association* 63(324):1379-1389.
421
+ - Theil H (1950). A rank-invariant method of linear and polynomial regression analysis. *Indagationes Mathematicae* 12:85-91.
422
+ - Wilson EB (1927). Probable inference, the law of succession, and statistical inference. *Journal of the American Statistical Association* 22(158):209-212.
423
+
424
+ ## Output for machines
425
+
426
+ `--json` writes every figure, with `poseaudit` (the version) and `schema`
427
+ (1) at the top and the settings used. NaN and infinity become `null`. Note
428
+ that `limits` there are the normal limits; the percentile limits the summary
429
+ prints are `empirical_limits`. Warnings, including those raised while
430
+ loading, are in `warnings`. `--csv` writes one row per reading:
431
+ `image, truth_index, predicted_index, class_id, cluster, size, truth,
432
+ predicted, error, mean`. Same inputs and seed, same bytes.
433
+
434
+ The command exits 0 on success (warnings included); 1 on bad input or
435
+ settings, a failed write, or nothing read (the JSON and CSV are still
436
+ written); and 2 when the arguments cannot be parsed. There is no pass or fail threshold built in; gate on the JSON, for
437
+ example `jq -e '.mean_abs_error_ci[1] <= 3 and .n / .measurable >= 0.9'`.
438
+ `--resamples` and `--jitter-repeats` trade precision for speed: 50,000
439
+ readings take about three minutes at the defaults.
440
+
441
+ ## Limitations
442
+
443
+ - The figures describe agreement with the ground truth, not with the world:
444
+ the truth's own error is inside them, and 2D angles are not 3D joint angles.
445
+ - Every reading counts once: frames pooled from a few trials inflate ICC and
446
+ CCC, and there is no per-trial or per-subject summary (peak angle, range of
447
+ motion) yet.
448
+ - Tilts near horizontal wrap from +90 to -90; the report warns when many are
449
+ close.
450
+ - The jitter reference is new, not a published method, and was checked by
451
+ simulation only ([`validation/jitter_reference.py`](https://github.com/8rulerstar/poseaudit/blob/main/validation/jitter_reference.py),
452
+ seeded). For models with no tendency and clean labels, with 300 arms and
453
+ 600 runs per case, the one-sided p was 0.05 or less in 2.3 to 5.7% of runs,
454
+ and the gap's interval lay wholly below 0 in 2.2 to 4.3% with the 200
455
+ resamples that table uses, and in 1.3 and 2.7% for the two cases rerun at
456
+ the tool's defaults (150 runs each; 2.5% nominal). With 40 arms and noise
457
+ that grows on straight arms the p ran to 8%, hence the warning under 60
458
+ readings (where exactly to draw that line was not tested). **With labels as
459
+ noisy as the model, honest models were flagged in about 35% of runs** (37%
460
+ at 200 resamples, 35% at the defaults): the check assumes labels much
461
+ cleaner than the model. A 7% squash was caught in every run, with point
462
+ noise of 8% of the arm; with noisier points it is caught less often (not
463
+ measured here). It assumes independent objects whose errors, apart from a
464
+ shift shared by similar angles, depend only on the true value and the
465
+ object's own frames. The p and the interval can disagree near the edge.
466
+ - Matching is greedy and ignores scores (filter them first with `--min-score`);
467
+ in dense crowds an optimal assignment may pair differently.
468
+ - One measure per run, one level of clusters, no time series.
469
+
470
+ ## Roadmap
471
+
472
+ - a label-noise reference beside the jitter one (from annotator spread or
473
+ repeated labels)
474
+ - left-right swap detection
475
+ - several measures and several models per run, on the readings they share
476
+ - auditing angles given directly (filtered or from inverse kinematics)
477
+ - directed angles (0 to 360), signed joint angles, a chosen reference axis
478
+ - regression-based limits of agreement; per-subject summaries
479
+
480
+ ## License and data
481
+
482
+ Code: MIT. The demo's annotations are a subset of the
483
+ [COCO](https://cocodataset.org) 2017 keypoint annotations (keypoint fields
484
+ only), © COCO Consortium, licensed under
485
+ [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/); the images are not
486
+ included. The demo's predictions were made with Ultralytics `yolo11n-pose`
487
+ (AGPL-3.0), which this package does not depend on or include.