bayesbin 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,29 @@
1
+ cff-version: 1.2.0
2
+ message: "If you use this software, please cite the paper below."
3
+ title: bayesbin
4
+ abstract: Exact Bayesian binning of rates in NumPy/SciPy.
5
+ type: software
6
+ license: BSD-3-Clause
7
+ repository-code: https://github.com/petfold/bayesbin
8
+ authors:
9
+ - family-names: Foldiak
10
+ given-names: Peter
11
+ preferred-citation:
12
+ type: conference-paper
13
+ title: "Bayesian binning beats approximate alternatives: estimating peri-stimulus time histograms"
14
+ authors:
15
+ - family-names: Endres
16
+ given-names: Dominik
17
+ - family-names: Oram
18
+ given-names: Mike
19
+ - family-names: Schindelin
20
+ given-names: Johannes
21
+ - family-names: Földiák
22
+ given-names: Peter
23
+ collection-title: Advances in Neural Information Processing Systems 20
24
+ publisher:
25
+ name: MIT Press
26
+ start: 393
27
+ end: 400
28
+ year: 2008
29
+ url: https://papers.nips.cc/paper_files/paper/2007/hash/b73ce398c39f506af761d2277d853a92-Abstract.html
bayesbin-0.1.0/LICENSE ADDED
@@ -0,0 +1,28 @@
1
+ BSD 3-Clause License
2
+
3
+ Copyright (c) 2026, Peter Foldiak
4
+
5
+ Redistribution and use in source and binary forms, with or without
6
+ modification, are permitted provided that the following conditions are met:
7
+
8
+ 1. Redistributions of source code must retain the above copyright notice, this
9
+ list of conditions and the following disclaimer.
10
+
11
+ 2. Redistributions in binary form must reproduce the above copyright notice,
12
+ this list of conditions and the following disclaimer in the documentation
13
+ and/or other materials provided with the distribution.
14
+
15
+ 3. Neither the name of the copyright holder nor the names of its
16
+ contributors may be used to endorse or promote products derived from
17
+ this software without specific prior written permission.
18
+
19
+ THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
20
+ AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
21
+ IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
22
+ DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
23
+ FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
24
+ DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
25
+ SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
26
+ CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
27
+ OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
28
+ OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
@@ -0,0 +1,11 @@
1
+ # The source distribution: the Python package (BSD-3-Clause) with what its tests need.
2
+ # The C++ (cpp/, reference/, GPL-2.0-or-later) and the paper stay in the repository only;
3
+ # the tests against the C++ skip without it.
4
+ include LICENSE README.md CITATION.cff llms.txt
5
+ recursive-include docs *.md *.png
6
+ recursive-include tests *.py *.txt
7
+ include tools/longdouble_reference.py
8
+ prune cpp
9
+ prune reference
10
+ prune paper
11
+ exclude tools/make_testdata.py
@@ -0,0 +1,289 @@
1
+ Metadata-Version: 2.4
2
+ Name: bayesbin
3
+ Version: 0.1.0
4
+ Summary: Exact Bayesian binning of rates (Endres, Oram, Schindelin & Földiák 2008) in NumPy/SciPy
5
+ Author: Peter Foldiak
6
+ License-Expression: BSD-3-Clause
7
+ Project-URL: Homepage, https://github.com/petfold/bayesbin
8
+ Project-URL: Paper, https://papers.nips.cc/paper_files/paper/2007/hash/b73ce398c39f506af761d2277d853a92-Abstract.html
9
+ Keywords: bayesian,binning,histogram,psth,spike trains,change points,rate estimation,poisson,segmentation
10
+ Classifier: Development Status :: 4 - Beta
11
+ Classifier: Intended Audience :: Science/Research
12
+ Classifier: Operating System :: OS Independent
13
+ Classifier: Programming Language :: Python :: 3
14
+ Classifier: Topic :: Scientific/Engineering :: Mathematics
15
+ Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
16
+ Requires-Python: >=3.10
17
+ Description-Content-Type: text/markdown
18
+ License-File: LICENSE
19
+ Requires-Dist: numpy>=1.24
20
+ Requires-Dist: scipy>=1.10
21
+ Provides-Extra: fast
22
+ Requires-Dist: numba>=0.60; extra == "fast"
23
+ Requires-Dist: threadpoolctl>=3.1; extra == "fast"
24
+ Provides-Extra: test
25
+ Requires-Dist: pytest>=7; extra == "test"
26
+ Dynamic: license-file
27
+
28
+ # bayesbin
29
+
30
+ [![tests](https://github.com/petfold/bayesbin/actions/workflows/tests.yml/badge.svg)](https://github.com/petfold/bayesbin/actions/workflows/tests.yml)
31
+ [![license](https://img.shields.io/badge/license-BSD--3--Clause-blue)](https://github.com/petfold/bayesbin/blob/main/LICENSE)
32
+ [![status](https://img.shields.io/badge/status-beta-yellow)](#status-and-limits)
33
+
34
+ Exact Bayesian binning of rates in NumPy/SciPy, after
35
+
36
+ > D. Endres, M. Oram, J. Schindelin, P. Földiák (2008). *Bayesian binning beats
37
+ > approximate alternatives: estimating peri-stimulus time histograms.*
38
+ > Advances in Neural Information Processing Systems 20, 393–400. MIT Press.
39
+ > ([NeurIPS page](https://papers.nips.cc/paper_files/paper/2007/hash/b73ce398c39f506af761d2277d853a92-Abstract.html); a copy in [paper/](https://github.com/petfold/bayesbin/tree/main/paper))
40
+
41
+ A rate on T ordered intervals is modelled as piecewise constant with M bin
42
+ boundaries. Boundary positions, per-bin rates (conjugate priors) and M itself
43
+ are all integrated out exactly: one forward dynamic programme gives the
44
+ evidence of every M in O(M·T²). A matching backward programme gives the
45
+ posterior of every candidate bin at once, from which the predictive rate, its
46
+ error bars and the posterior over boundary positions follow.
47
+
48
+ **New to this? Start with the [User Guide](https://github.com/petfold/bayesbin/blob/main/docs/USER_GUIDE.md)**: why fixed-width
49
+ bins mislead, the assumptions in plain words, four worked examples (spike
50
+ trains, counts with exposure, success rates, a daily profile) and the pitfalls.
51
+ For AI coding assistants there is a compact [llms.txt](https://github.com/petfold/bayesbin/blob/main/llms.txt).
52
+
53
+ ```sh
54
+ pip install bayesbin # NumPy/SciPy only
55
+ pip install "bayesbin[fast]" # + numba kernels: ~2x faster, all cores
56
+ ```
57
+
58
+ ```python
59
+ from bayesbin import BernoulliModel, PoissonModel, fit, spike_counts
60
+
61
+ # spike trains, as in the paper: one list of integer spike times per trial
62
+ s, g = spike_counts(trials, t_start=-100, t_end=499)
63
+ r = fit(BernoulliModel(s, g, sigma=1.0, gamma=32.0), max_boundaries=10)
64
+ r.rate, r.rate_std # predictive firing probability per interval, ± 1 sd
65
+ r.m_posterior # P(M | data)
66
+ r.boundary_posterior # P(a bin ends at interval k | data)
67
+
68
+ # counts per window, several events per window allowed, with exposure
69
+ r = fit(PoissonModel.weak_prior(counts, exposure), max_boundaries=20)
70
+ ```
71
+
72
+ By default predictions average over every M, as the paper recommends;
73
+ `m_mass=0.9` restricts them to the credible range of M, which is what the
74
+ original program does.
75
+
76
+ NumPy and SciPy are all it needs. With the `fast` extra (`pip install
77
+ bayesbin[fast]`: numba and threadpoolctl) the work runs in fused kernels on all
78
+ cores: about 2× faster on one core, and scaling to about 3× more on four. The
79
+ first call in a new environment compiles them (about a minute, once; cached
80
+ afterwards), and `BAYESBIN_NUMBA=0` switches them off. Both paths give the same
81
+ results to rounding, and the fused ones the same bits on any number of threads.
82
+
83
+ Threads: numba's, `NUMBA_NUM_THREADS` or `numba.set_num_threads(n)`; the
84
+ default is every logical CPU, and hyperthreads gain nothing here, so set it to
85
+ the number of physical cores for the best time. While a fit runs, the fused
86
+ path holds BLAS to one thread and splits the matrix products over numba's
87
+ threads itself. The kernels release the GIL, so many separate fits (many
88
+ short series) can also run in parallel from a thread pool, each on one numba
89
+ thread (under numba's TBB or OpenMP threading layer).
90
+
91
+ ## What it is for
92
+
93
+ Any rate that varies along an ordered axis and is observed as events per
94
+ interval, where you want the rate *and* its uncertainty without choosing bin
95
+ widths by hand:
96
+
97
+ - **Peri-stimulus time histograms** (the paper's case): spike trains over
98
+ repeated trials, the firing probability per millisecond, with error bars,
99
+ from a few dozen trials.
100
+ - **Event counts per window**: arrivals, requests, incidents, photon or
101
+ particle counts, cases per week — with an exposure per window (observation
102
+ time, population, detector area) when windows differ.
103
+ - **Proportions along an axis**: successes out of trials per interval
104
+ (conversion or failure rates by time of day, by age, by dose), with the
105
+ Bernoulli model.
106
+ - **Change points**: `boundary_posterior` is the posterior probability that the
107
+ rate changes after each interval, averaged over every segmentation.
108
+ - **Periodic profiles**: daily or weekly shapes, with the days (or weeks) as
109
+ trials and the time of day as the axis.
110
+
111
+ Nothing is fitted by optimisation and nothing is sampled: every segmentation
112
+ into up to M + 1 bins, and every M, is summed exactly. Compared with
113
+ Bayesian Blocks (Scargle et al. 2013, ApJ 764:167; `astropy.stats.bayesian_blocks`),
114
+ which finds the single best segmentation under a penalty per block, this
115
+ averages over all segmentations, so the rate is smooth where the data do not
116
+ decide where a step is, and comes with error bars.
117
+
118
+ The axis is 1-D; see [docs/NOTES.md](https://github.com/petfold/bayesbin/blob/main/docs/NOTES.md) for how far this extends
119
+ to two dimensions.
120
+
121
+ ## Models
122
+
123
+ | model | per interval | per-bin prior | use |
124
+ |---|---|---|---|
125
+ | `BernoulliModel(s, g, sigma, gamma)` | s trials with an event, g without | f ~ Beta(σ, γ) | the paper's PSTH |
126
+ | `PoissonModel(y, alpha, beta, e)` | count y over exposure e | λ ~ Gamma(α, β) | event counts per window |
127
+
128
+ `BernoulliModel` defaults to σ = 1, γ = 32, the original program's default.
129
+
130
+ ## Verification
131
+
132
+ `pytest` (41 tests; 5 need `cpp/binsdfc-fb` built, 7 need the `fast` extra):
133
+
134
+ - **Against the original C++ program** (`binsdfc` 0.1, in [reference/](https://github.com/petfold/bayesbin/tree/main/reference)),
135
+ on a seeded dataset in its own input format (`tools/make_testdata.py`):
136
+ - log P(D | M) for M = 0..10 and the marginal likelihood agree to every
137
+ printed digit;
138
+ - the predictive rate agrees within 2 × 10⁻⁵ relative;
139
+ - its standard deviation agrees within 5 × 10⁻⁵. The original adds in log
140
+ space through an interpolated lookup table, which shows at that level.
141
+ - **Against brute-force enumeration** of every boundary configuration, for both
142
+ models: evidence and predictive rate to 10⁻¹⁰.
143
+ - **Against the paper's own device** (§4): P(spike | k) as the ratio of
144
+ evidences with and without a virtual spike at k equals the forward–backward
145
+ result for every k.
146
+ - The fast paths against the exact log-space ones, on data whose evidences span
147
+ thousands of nats, with the underflow fallback forced; the fused kernels
148
+ against the NumPy path and the exact one, for both models (constant and
149
+ varying exposure).
150
+ - **Against a long-double reference** (`tools/longdouble_reference.py`: the
151
+ whole computation in 80-bit arithmetic, the variance in its stable form),
152
+ on steps strong enough that the sd, √(E[f²] − E[f]²), cancels: every path's
153
+ rate to 10⁻¹², its sd to 10⁻⁹. (Both moments are divided by the computed
154
+ coverage, the posterior of the bins covering each interval, which is 1 in
155
+ exact arithmetic; its rounding error would otherwise reach the sd amplified
156
+ by rate²/var.)
157
+ - The User Guide's examples run and print what the guide says they print.
158
+ - One-bin evidences against direct numerical integration; every interval
159
+ covered by exactly one bin; the simulated response onset recovered.
160
+
161
+ ## Status and limits
162
+
163
+ - Cost: O(M·T²) time, O(T·M) memory. Nothing of size T×T is formed on the
164
+ default path: the models give bin evidences and posterior moments block by
165
+ block (`bin_block`, from prefix sums and `lgamma` tables); the forward and
166
+ backward programmes run in blocks of 256 columns, each block's slice of
167
+ exponentiated gains made once from the upper triangle and serving all M
168
+ steps (the rows before a block, final for every step, as one matrix
169
+ product; only the rows inside it step by step); the bin posterior is
170
+ accumulated in tiles of 256 × 1024 bins, as scaled matrix products, into the
171
+ rates and the boundary posterior. The factors of every matrix product are
172
+ flushed to 0 below a threshold chosen so that no product is subnormal
173
+ (subnormal arithmetic is ~100× slower on x86, and BLAS runs without
174
+ flush-to-zero); the error bounds count the flushed terms as lost.
175
+ `keep_bins=True` (the whole bin posterior) and `exact=True` use T×T arrays. Wherever underflow could cost
176
+ more than 10⁻¹³ (relative, evidences) or 10⁻¹⁴ (absolute, bin posterior),
177
+ that entry is recomputed exactly in log space, and `exact=True` does
178
+ everything that way. See the timings below.
179
+ - Not yet ported from the original: latency posteriors, signal separation
180
+ levels, hyperparameter optimisation (`-P`), bin-boundary position posteriors
181
+ for a fixed M (`-p`).
182
+ - Planned: a release on PyPI, so that `pip install bayesbin` works (steps in
183
+ [docs/NOTES.md](https://github.com/petfold/bayesbin/blob/main/docs/NOTES.md#plan)); cyclic profiles (a bin may wrap round
184
+ the end of a day or week); 2-D via recursive partitions (see
185
+ [docs/NOTES.md](https://github.com/petfold/bayesbin/blob/main/docs/NOTES.md)).
186
+
187
+ ## Speed against the original
188
+
189
+ binsdfc 0.1 unmodified (`g++ -O2 -fopenmp`; its own flags `-march=native
190
+ -ffast-math` made no real difference), against bayesbin with NumPy 2.5,
191
+ OpenBLAS and numba 0.67. Intel i7-3612QM (4 cores, 2 hyperthreads each; AVX,
192
+ no AVX2 or FMA); "1 core" means one physical core, and "4 threads" four
193
+ separate physical cores (`taskset`; numba, OpenMP and OpenBLAS thread counts
194
+ set to match). binsdfc is timed as a process, bayesbin in-process without the
195
+ import and after a warm-up call (the fused kernels' cache load, once per
196
+ process).
197
+
198
+ | case | binsdfc, 1 core | binsdfc, 4 threads | binsdfc-fb, 1 core | binsdfc-fb, 4 threads | bayesbin (NumPy), 1 core | bayesbin + numba, 1 core | bayesbin + numba, 4 threads |
199
+ |---|---|---|---|---|---|---|---|
200
+ | T=300, M≤10, rate ± sd | 0.98 s | 0.27 s | | | 0.018 s | 0.008 s | 0.006 s |
201
+ | T=600, M≤10, rate ± sd | 50.0 s | 12.4 s | 0.033 s | 0.021 s | 0.050 s | 0.023 s | 0.012 s |
202
+ | T=600, M≤10, evidence only | 0.065 s | — | 0.020 s | 0.017 s | 0.014 s | 0.008 s | 0.005 s |
203
+ | T=2016, M≤30, evidence only | 2.0 s | — | 0.14 s | 0.088 s | 0.14 s | 0.083 s | 0.035 s |
204
+ | T=2016, M≤30, rate ± sd | stopped after 26 min | | 0.33 s | 0.17 s | 0.51 s | 0.26 s | 0.099 s |
205
+
206
+ `binsdfc-fb` is the original with the forward–backward SDF, the matrix-vector
207
+ central iteration and table-driven bin evidences added ([cpp/binsdfc-fb/](https://github.com/petfold/bayesbin/blob/main/cpp/binsdfc-fb/README.fb.md)); best
208
+ of 5 runs. (binsdfc itself always runs 4 threads.)
209
+
210
+ - The evidences use the same dynamic programme in both. binsdfc's triple loop
211
+ is the slowest; binsdfc-fb and bayesbin both run its central iteration in
212
+ column blocks, bayesbin with the rows before each block as one matrix
213
+ product for all M steps (BLAS-3), which binsdfc-fb does not have.
214
+ - For the rate and its error bars binsdfc uses the paper's virtual-spike
215
+ device: for every time point it reruns the whole programme twice (rate and
216
+ second moment), O(M·T³). bayesbin gets every time point from one backward
217
+ pass, O(M·T²). The gap is the algorithm, not the language.
218
+ - binsdfc fixes 4 OpenMP threads over time points (`omp_set_num_threads(4)`)
219
+ and scales almost 4×. bayesbin's NumPy path runs on one thread outside BLAS;
220
+ its fused path runs every step on numba's threads (the matrix products split
221
+ into fixed parts) except the steps inside each 256-column block of the
222
+ dynamic programme, which are sequential in M, and scales about 2.9× on 4
223
+ cores.
224
+
225
+ ### Larger problems
226
+
227
+ 30 trials of a synthetic daily profile (5-minute slots: night, morning ramp,
228
+ day, evening peak, plus a 2-hour burst each week; `tools/make_longdata.py T`),
229
+ rate ± sd, most probable M only (`-l 0`, `m_mass=0.0`), peak memory from
230
+ `/usr/bin/time` and `getrusage`. All but the 12-week row run back to back, one
231
+ run each; the laptop was thermally throttled.
232
+
233
+ | T | M ≤ | binsdfc-fb, 1 core | binsdfc-fb, 4 threads | binsdfc-fb memory | bayesbin (NumPy), 1 core | bayesbin + numba, 1 core | bayesbin + numba, 4 threads |
234
+ |---|---|---|---|---|---|---|---|
235
+ | 2016 (1 week) | 30 | 0.25 s | 0.10 s | 12 MB | 0.54 s, 88 MB | 0.27 s, 175 MB | 0.11 s, 177 MB |
236
+ | 4032 (2 weeks) | 60 | 1.44 s | 0.45 s | 19 MB | 1.99 s, 110 MB | 1.18 s, 189 MB | 0.53 s, 192 MB |
237
+ | 8064 (4 weeks) | 120 | 9.2 s | 2.9 s | 44 MB | 8.7 s, 179 MB | 5.8 s, 235 MB | 2.0 s, 242 MB |
238
+ | 12096 (6 weeks) | 120 | 20.3 s | 6.3 s | 61 MB | 18.6 s, 250 MB | 13.0 s, 271 MB | 4.3 s, 280 MB |
239
+ | 24192 (12 weeks) | 120 | | 86 s | 116 MB | | | |
240
+
241
+ - binsdfc-fb memory is O(T·M): no T×T array is kept. (Before: ≈14·T²
242
+ bytes, 2.1 GB at 6 weeks; 12 weeks would have needed ≈8 GB.) Output is
243
+ byte-identical to the T² version at 2 and 6 weeks. bayesbin is O(T·M) too
244
+ now (its figures include ≈60 MB of Python and NumPy, and ≈90 MB more for
245
+ numba and LLVM); before, ≈70·T² bytes (1.1 GB at 2 weeks).
246
+ - Time grows about as T² and linearly in M; the original needs 8.5 s for the
247
+ evidences alone at T=4032.
248
+ - binsdfc-fb and bayesbin agree to the 6 printed digits at T=4032.
249
+ - A periodic series needs M to grow with its length: the most probable M was
250
+ 106 at 2 weeks (M ≤ 120) and 297 at 6 weeks (M ≤ 400), the same daily shape
251
+ re-learnt every day. For daily or weekly profiles, fold the series instead
252
+ (days as trials, time of day as the axis): T = 288, a few milliseconds.
253
+
254
+ ## Licence
255
+
256
+ Two licences, by directory:
257
+
258
+ - **BSD-3-Clause** ([LICENSE](https://github.com/petfold/bayesbin/blob/main/LICENSE)): the Python package `src/bayesbin` (all
259
+ that `pip` installs), its tests, `tools/bench_vs_binsdfc.py` and the
260
+ documentation. The package was written from the paper; the C++ program was
261
+ used only as a reference to test against, and none of its code is in it.
262
+ - **GPL-2.0-or-later**: [reference/binsdfc-0.1/](https://github.com/petfold/bayesbin/tree/main/reference) (Dominik Endres's
263
+ original, unmodified, with its provenance in [reference/README.md](https://github.com/petfold/bayesbin/blob/main/reference/README.md)),
264
+ [cpp/binsdfc-fb/](https://github.com/petfold/bayesbin/blob/main/cpp/binsdfc-fb/README.fb.md) (that program with a
265
+ forward–backward SDF and the speed-ups described there, each changed file
266
+ marked) and `tools/make_testdata.py` (a port of its test-data script). Their
267
+ licence text is in [cpp/COPYING](https://github.com/petfold/bayesbin/blob/main/cpp/COPYING) and
268
+ [reference/COPYING](https://github.com/petfold/bayesbin/blob/main/reference/COPYING).
269
+ - The paper in [paper/](https://github.com/petfold/bayesbin/tree/main/paper) is copyright its authors, not under either
270
+ licence.
271
+
272
+ ## Citing
273
+
274
+ If you use this, please cite the paper:
275
+
276
+ ```bibtex
277
+ @inproceedings{endres2008bayesian,
278
+ title = {Bayesian binning beats approximate alternatives: estimating peri-stimulus time histograms},
279
+ author = {Endres, Dominik and Oram, Mike and Schindelin, Johannes and F{\"o}ldi{\'a}k, Peter},
280
+ booktitle = {Advances in Neural Information Processing Systems 20},
281
+ pages = {393--400},
282
+ publisher = {MIT Press},
283
+ year = {2008}
284
+ }
285
+ ```
286
+
287
+ The method is Endres, Oram, Schindelin and Földiák's; binsdfc, the original
288
+ C++ implementation, is Dominik Endres's; bayesbin and the binsdfc-fb changes
289
+ are by Peter Foldiak.
@@ -0,0 +1,262 @@
1
+ # bayesbin
2
+
3
+ [![tests](https://github.com/petfold/bayesbin/actions/workflows/tests.yml/badge.svg)](https://github.com/petfold/bayesbin/actions/workflows/tests.yml)
4
+ [![license](https://img.shields.io/badge/license-BSD--3--Clause-blue)](https://github.com/petfold/bayesbin/blob/main/LICENSE)
5
+ [![status](https://img.shields.io/badge/status-beta-yellow)](#status-and-limits)
6
+
7
+ Exact Bayesian binning of rates in NumPy/SciPy, after
8
+
9
+ > D. Endres, M. Oram, J. Schindelin, P. Földiák (2008). *Bayesian binning beats
10
+ > approximate alternatives: estimating peri-stimulus time histograms.*
11
+ > Advances in Neural Information Processing Systems 20, 393–400. MIT Press.
12
+ > ([NeurIPS page](https://papers.nips.cc/paper_files/paper/2007/hash/b73ce398c39f506af761d2277d853a92-Abstract.html); a copy in [paper/](https://github.com/petfold/bayesbin/tree/main/paper))
13
+
14
+ A rate on T ordered intervals is modelled as piecewise constant with M bin
15
+ boundaries. Boundary positions, per-bin rates (conjugate priors) and M itself
16
+ are all integrated out exactly: one forward dynamic programme gives the
17
+ evidence of every M in O(M·T²). A matching backward programme gives the
18
+ posterior of every candidate bin at once, from which the predictive rate, its
19
+ error bars and the posterior over boundary positions follow.
20
+
21
+ **New to this? Start with the [User Guide](https://github.com/petfold/bayesbin/blob/main/docs/USER_GUIDE.md)**: why fixed-width
22
+ bins mislead, the assumptions in plain words, four worked examples (spike
23
+ trains, counts with exposure, success rates, a daily profile) and the pitfalls.
24
+ For AI coding assistants there is a compact [llms.txt](https://github.com/petfold/bayesbin/blob/main/llms.txt).
25
+
26
+ ```sh
27
+ pip install bayesbin # NumPy/SciPy only
28
+ pip install "bayesbin[fast]" # + numba kernels: ~2x faster, all cores
29
+ ```
30
+
31
+ ```python
32
+ from bayesbin import BernoulliModel, PoissonModel, fit, spike_counts
33
+
34
+ # spike trains, as in the paper: one list of integer spike times per trial
35
+ s, g = spike_counts(trials, t_start=-100, t_end=499)
36
+ r = fit(BernoulliModel(s, g, sigma=1.0, gamma=32.0), max_boundaries=10)
37
+ r.rate, r.rate_std # predictive firing probability per interval, ± 1 sd
38
+ r.m_posterior # P(M | data)
39
+ r.boundary_posterior # P(a bin ends at interval k | data)
40
+
41
+ # counts per window, several events per window allowed, with exposure
42
+ r = fit(PoissonModel.weak_prior(counts, exposure), max_boundaries=20)
43
+ ```
44
+
45
+ By default predictions average over every M, as the paper recommends;
46
+ `m_mass=0.9` restricts them to the credible range of M, which is what the
47
+ original program does.
48
+
49
+ NumPy and SciPy are all it needs. With the `fast` extra (`pip install
50
+ bayesbin[fast]`: numba and threadpoolctl) the work runs in fused kernels on all
51
+ cores: about 2× faster on one core, and scaling to about 3× more on four. The
52
+ first call in a new environment compiles them (about a minute, once; cached
53
+ afterwards), and `BAYESBIN_NUMBA=0` switches them off. Both paths give the same
54
+ results to rounding, and the fused ones the same bits on any number of threads.
55
+
56
+ Threads: numba's, `NUMBA_NUM_THREADS` or `numba.set_num_threads(n)`; the
57
+ default is every logical CPU, and hyperthreads gain nothing here, so set it to
58
+ the number of physical cores for the best time. While a fit runs, the fused
59
+ path holds BLAS to one thread and splits the matrix products over numba's
60
+ threads itself. The kernels release the GIL, so many separate fits (many
61
+ short series) can also run in parallel from a thread pool, each on one numba
62
+ thread (under numba's TBB or OpenMP threading layer).
63
+
64
+ ## What it is for
65
+
66
+ Any rate that varies along an ordered axis and is observed as events per
67
+ interval, where you want the rate *and* its uncertainty without choosing bin
68
+ widths by hand:
69
+
70
+ - **Peri-stimulus time histograms** (the paper's case): spike trains over
71
+ repeated trials, the firing probability per millisecond, with error bars,
72
+ from a few dozen trials.
73
+ - **Event counts per window**: arrivals, requests, incidents, photon or
74
+ particle counts, cases per week — with an exposure per window (observation
75
+ time, population, detector area) when windows differ.
76
+ - **Proportions along an axis**: successes out of trials per interval
77
+ (conversion or failure rates by time of day, by age, by dose), with the
78
+ Bernoulli model.
79
+ - **Change points**: `boundary_posterior` is the posterior probability that the
80
+ rate changes after each interval, averaged over every segmentation.
81
+ - **Periodic profiles**: daily or weekly shapes, with the days (or weeks) as
82
+ trials and the time of day as the axis.
83
+
84
+ Nothing is fitted by optimisation and nothing is sampled: every segmentation
85
+ into up to M + 1 bins, and every M, is summed exactly. Compared with
86
+ Bayesian Blocks (Scargle et al. 2013, ApJ 764:167; `astropy.stats.bayesian_blocks`),
87
+ which finds the single best segmentation under a penalty per block, this
88
+ averages over all segmentations, so the rate is smooth where the data do not
89
+ decide where a step is, and comes with error bars.
90
+
91
+ The axis is 1-D; see [docs/NOTES.md](https://github.com/petfold/bayesbin/blob/main/docs/NOTES.md) for how far this extends
92
+ to two dimensions.
93
+
94
+ ## Models
95
+
96
+ | model | per interval | per-bin prior | use |
97
+ |---|---|---|---|
98
+ | `BernoulliModel(s, g, sigma, gamma)` | s trials with an event, g without | f ~ Beta(σ, γ) | the paper's PSTH |
99
+ | `PoissonModel(y, alpha, beta, e)` | count y over exposure e | λ ~ Gamma(α, β) | event counts per window |
100
+
101
+ `BernoulliModel` defaults to σ = 1, γ = 32, the original program's default.
102
+
103
+ ## Verification
104
+
105
+ `pytest` (41 tests; 5 need `cpp/binsdfc-fb` built, 7 need the `fast` extra):
106
+
107
+ - **Against the original C++ program** (`binsdfc` 0.1, in [reference/](https://github.com/petfold/bayesbin/tree/main/reference)),
108
+ on a seeded dataset in its own input format (`tools/make_testdata.py`):
109
+ - log P(D | M) for M = 0..10 and the marginal likelihood agree to every
110
+ printed digit;
111
+ - the predictive rate agrees within 2 × 10⁻⁵ relative;
112
+ - its standard deviation agrees within 5 × 10⁻⁵. The original adds in log
113
+ space through an interpolated lookup table, which shows at that level.
114
+ - **Against brute-force enumeration** of every boundary configuration, for both
115
+ models: evidence and predictive rate to 10⁻¹⁰.
116
+ - **Against the paper's own device** (§4): P(spike | k) as the ratio of
117
+ evidences with and without a virtual spike at k equals the forward–backward
118
+ result for every k.
119
+ - The fast paths against the exact log-space ones, on data whose evidences span
120
+ thousands of nats, with the underflow fallback forced; the fused kernels
121
+ against the NumPy path and the exact one, for both models (constant and
122
+ varying exposure).
123
+ - **Against a long-double reference** (`tools/longdouble_reference.py`: the
124
+ whole computation in 80-bit arithmetic, the variance in its stable form),
125
+ on steps strong enough that the sd, √(E[f²] − E[f]²), cancels: every path's
126
+ rate to 10⁻¹², its sd to 10⁻⁹. (Both moments are divided by the computed
127
+ coverage, the posterior of the bins covering each interval, which is 1 in
128
+ exact arithmetic; its rounding error would otherwise reach the sd amplified
129
+ by rate²/var.)
130
+ - The User Guide's examples run and print what the guide says they print.
131
+ - One-bin evidences against direct numerical integration; every interval
132
+ covered by exactly one bin; the simulated response onset recovered.
133
+
134
+ ## Status and limits
135
+
136
+ - Cost: O(M·T²) time, O(T·M) memory. Nothing of size T×T is formed on the
137
+ default path: the models give bin evidences and posterior moments block by
138
+ block (`bin_block`, from prefix sums and `lgamma` tables); the forward and
139
+ backward programmes run in blocks of 256 columns, each block's slice of
140
+ exponentiated gains made once from the upper triangle and serving all M
141
+ steps (the rows before a block, final for every step, as one matrix
142
+ product; only the rows inside it step by step); the bin posterior is
143
+ accumulated in tiles of 256 × 1024 bins, as scaled matrix products, into the
144
+ rates and the boundary posterior. The factors of every matrix product are
145
+ flushed to 0 below a threshold chosen so that no product is subnormal
146
+ (subnormal arithmetic is ~100× slower on x86, and BLAS runs without
147
+ flush-to-zero); the error bounds count the flushed terms as lost.
148
+ `keep_bins=True` (the whole bin posterior) and `exact=True` use T×T arrays. Wherever underflow could cost
149
+ more than 10⁻¹³ (relative, evidences) or 10⁻¹⁴ (absolute, bin posterior),
150
+ that entry is recomputed exactly in log space, and `exact=True` does
151
+ everything that way. See the timings below.
152
+ - Not yet ported from the original: latency posteriors, signal separation
153
+ levels, hyperparameter optimisation (`-P`), bin-boundary position posteriors
154
+ for a fixed M (`-p`).
155
+ - Planned: a release on PyPI, so that `pip install bayesbin` works (steps in
156
+ [docs/NOTES.md](https://github.com/petfold/bayesbin/blob/main/docs/NOTES.md#plan)); cyclic profiles (a bin may wrap round
157
+ the end of a day or week); 2-D via recursive partitions (see
158
+ [docs/NOTES.md](https://github.com/petfold/bayesbin/blob/main/docs/NOTES.md)).
159
+
160
+ ## Speed against the original
161
+
162
+ binsdfc 0.1 unmodified (`g++ -O2 -fopenmp`; its own flags `-march=native
163
+ -ffast-math` made no real difference), against bayesbin with NumPy 2.5,
164
+ OpenBLAS and numba 0.67. Intel i7-3612QM (4 cores, 2 hyperthreads each; AVX,
165
+ no AVX2 or FMA); "1 core" means one physical core, and "4 threads" four
166
+ separate physical cores (`taskset`; numba, OpenMP and OpenBLAS thread counts
167
+ set to match). binsdfc is timed as a process, bayesbin in-process without the
168
+ import and after a warm-up call (the fused kernels' cache load, once per
169
+ process).
170
+
171
+ | case | binsdfc, 1 core | binsdfc, 4 threads | binsdfc-fb, 1 core | binsdfc-fb, 4 threads | bayesbin (NumPy), 1 core | bayesbin + numba, 1 core | bayesbin + numba, 4 threads |
172
+ |---|---|---|---|---|---|---|---|
173
+ | T=300, M≤10, rate ± sd | 0.98 s | 0.27 s | | | 0.018 s | 0.008 s | 0.006 s |
174
+ | T=600, M≤10, rate ± sd | 50.0 s | 12.4 s | 0.033 s | 0.021 s | 0.050 s | 0.023 s | 0.012 s |
175
+ | T=600, M≤10, evidence only | 0.065 s | — | 0.020 s | 0.017 s | 0.014 s | 0.008 s | 0.005 s |
176
+ | T=2016, M≤30, evidence only | 2.0 s | — | 0.14 s | 0.088 s | 0.14 s | 0.083 s | 0.035 s |
177
+ | T=2016, M≤30, rate ± sd | stopped after 26 min | | 0.33 s | 0.17 s | 0.51 s | 0.26 s | 0.099 s |
178
+
179
+ `binsdfc-fb` is the original with the forward–backward SDF, the matrix-vector
180
+ central iteration and table-driven bin evidences added ([cpp/binsdfc-fb/](https://github.com/petfold/bayesbin/blob/main/cpp/binsdfc-fb/README.fb.md)); best
181
+ of 5 runs. (binsdfc itself always runs 4 threads.)
182
+
183
+ - The evidences use the same dynamic programme in both. binsdfc's triple loop
184
+ is the slowest; binsdfc-fb and bayesbin both run its central iteration in
185
+ column blocks, bayesbin with the rows before each block as one matrix
186
+ product for all M steps (BLAS-3), which binsdfc-fb does not have.
187
+ - For the rate and its error bars binsdfc uses the paper's virtual-spike
188
+ device: for every time point it reruns the whole programme twice (rate and
189
+ second moment), O(M·T³). bayesbin gets every time point from one backward
190
+ pass, O(M·T²). The gap is the algorithm, not the language.
191
+ - binsdfc fixes 4 OpenMP threads over time points (`omp_set_num_threads(4)`)
192
+ and scales almost 4×. bayesbin's NumPy path runs on one thread outside BLAS;
193
+ its fused path runs every step on numba's threads (the matrix products split
194
+ into fixed parts) except the steps inside each 256-column block of the
195
+ dynamic programme, which are sequential in M, and scales about 2.9× on 4
196
+ cores.
197
+
198
+ ### Larger problems
199
+
200
+ 30 trials of a synthetic daily profile (5-minute slots: night, morning ramp,
201
+ day, evening peak, plus a 2-hour burst each week; `tools/make_longdata.py T`),
202
+ rate ± sd, most probable M only (`-l 0`, `m_mass=0.0`), peak memory from
203
+ `/usr/bin/time` and `getrusage`. All but the 12-week row run back to back, one
204
+ run each; the laptop was thermally throttled.
205
+
206
+ | T | M ≤ | binsdfc-fb, 1 core | binsdfc-fb, 4 threads | binsdfc-fb memory | bayesbin (NumPy), 1 core | bayesbin + numba, 1 core | bayesbin + numba, 4 threads |
207
+ |---|---|---|---|---|---|---|---|
208
+ | 2016 (1 week) | 30 | 0.25 s | 0.10 s | 12 MB | 0.54 s, 88 MB | 0.27 s, 175 MB | 0.11 s, 177 MB |
209
+ | 4032 (2 weeks) | 60 | 1.44 s | 0.45 s | 19 MB | 1.99 s, 110 MB | 1.18 s, 189 MB | 0.53 s, 192 MB |
210
+ | 8064 (4 weeks) | 120 | 9.2 s | 2.9 s | 44 MB | 8.7 s, 179 MB | 5.8 s, 235 MB | 2.0 s, 242 MB |
211
+ | 12096 (6 weeks) | 120 | 20.3 s | 6.3 s | 61 MB | 18.6 s, 250 MB | 13.0 s, 271 MB | 4.3 s, 280 MB |
212
+ | 24192 (12 weeks) | 120 | | 86 s | 116 MB | | | |
213
+
214
+ - binsdfc-fb memory is O(T·M): no T×T array is kept. (Before: ≈14·T²
215
+ bytes, 2.1 GB at 6 weeks; 12 weeks would have needed ≈8 GB.) Output is
216
+ byte-identical to the T² version at 2 and 6 weeks. bayesbin is O(T·M) too
217
+ now (its figures include ≈60 MB of Python and NumPy, and ≈90 MB more for
218
+ numba and LLVM); before, ≈70·T² bytes (1.1 GB at 2 weeks).
219
+ - Time grows about as T² and linearly in M; the original needs 8.5 s for the
220
+ evidences alone at T=4032.
221
+ - binsdfc-fb and bayesbin agree to the 6 printed digits at T=4032.
222
+ - A periodic series needs M to grow with its length: the most probable M was
223
+ 106 at 2 weeks (M ≤ 120) and 297 at 6 weeks (M ≤ 400), the same daily shape
224
+ re-learnt every day. For daily or weekly profiles, fold the series instead
225
+ (days as trials, time of day as the axis): T = 288, a few milliseconds.
226
+
227
+ ## Licence
228
+
229
+ Two licences, by directory:
230
+
231
+ - **BSD-3-Clause** ([LICENSE](https://github.com/petfold/bayesbin/blob/main/LICENSE)): the Python package `src/bayesbin` (all
232
+ that `pip` installs), its tests, `tools/bench_vs_binsdfc.py` and the
233
+ documentation. The package was written from the paper; the C++ program was
234
+ used only as a reference to test against, and none of its code is in it.
235
+ - **GPL-2.0-or-later**: [reference/binsdfc-0.1/](https://github.com/petfold/bayesbin/tree/main/reference) (Dominik Endres's
236
+ original, unmodified, with its provenance in [reference/README.md](https://github.com/petfold/bayesbin/blob/main/reference/README.md)),
237
+ [cpp/binsdfc-fb/](https://github.com/petfold/bayesbin/blob/main/cpp/binsdfc-fb/README.fb.md) (that program with a
238
+ forward–backward SDF and the speed-ups described there, each changed file
239
+ marked) and `tools/make_testdata.py` (a port of its test-data script). Their
240
+ licence text is in [cpp/COPYING](https://github.com/petfold/bayesbin/blob/main/cpp/COPYING) and
241
+ [reference/COPYING](https://github.com/petfold/bayesbin/blob/main/reference/COPYING).
242
+ - The paper in [paper/](https://github.com/petfold/bayesbin/tree/main/paper) is copyright its authors, not under either
243
+ licence.
244
+
245
+ ## Citing
246
+
247
+ If you use this, please cite the paper:
248
+
249
+ ```bibtex
250
+ @inproceedings{endres2008bayesian,
251
+ title = {Bayesian binning beats approximate alternatives: estimating peri-stimulus time histograms},
252
+ author = {Endres, Dominik and Oram, Mike and Schindelin, Johannes and F{\"o}ldi{\'a}k, Peter},
253
+ booktitle = {Advances in Neural Information Processing Systems 20},
254
+ pages = {393--400},
255
+ publisher = {MIT Press},
256
+ year = {2008}
257
+ }
258
+ ```
259
+
260
+ The method is Endres, Oram, Schindelin and Földiák's; binsdfc, the original
261
+ C++ implementation, is Dominik Endres's; bayesbin and the binsdfc-fb changes
262
+ are by Peter Foldiak.