bayesbin 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- bayesbin-0.1.0/CITATION.cff +29 -0
- bayesbin-0.1.0/LICENSE +28 -0
- bayesbin-0.1.0/MANIFEST.in +11 -0
- bayesbin-0.1.0/PKG-INFO +289 -0
- bayesbin-0.1.0/README.md +262 -0
- bayesbin-0.1.0/docs/NOTES.md +150 -0
- bayesbin-0.1.0/docs/USER_GUIDE.md +397 -0
- bayesbin-0.1.0/docs/img/bins-dark.png +0 -0
- bayesbin-0.1.0/docs/img/bins-light.png +0 -0
- bayesbin-0.1.0/docs/img/boundaries-dark.png +0 -0
- bayesbin-0.1.0/docs/img/boundaries-light.png +0 -0
- bayesbin-0.1.0/llms.txt +68 -0
- bayesbin-0.1.0/pyproject.toml +38 -0
- bayesbin-0.1.0/setup.cfg +4 -0
- bayesbin-0.1.0/src/bayesbin/__init__.py +24 -0
- bayesbin-0.1.0/src/bayesbin/_fast.py +452 -0
- bayesbin-0.1.0/src/bayesbin/core.py +749 -0
- bayesbin-0.1.0/src/bayesbin.egg-info/PKG-INFO +289 -0
- bayesbin-0.1.0/src/bayesbin.egg-info/SOURCES.txt +28 -0
- bayesbin-0.1.0/src/bayesbin.egg-info/dependency_links.txt +1 -0
- bayesbin-0.1.0/src/bayesbin.egg-info/requires.txt +9 -0
- bayesbin-0.1.0/src/bayesbin.egg-info/top_level.txt +1 -0
- bayesbin-0.1.0/tests/data/binsdfc_seed1_marginal.txt +14 -0
- bayesbin-0.1.0/tests/data/binsdfc_seed1_sdf_l0.9.txt +603 -0
- bayesbin-0.1.0/tests/data/binsdfc_seed1_sdf_l0.txt +603 -0
- bayesbin-0.1.0/tests/data/strong_400trials.txt +400 -0
- bayesbin-0.1.0/tests/data/testdata_seed1.txt +30 -0
- bayesbin-0.1.0/tests/test_bayesbin.py +518 -0
- bayesbin-0.1.0/tests/test_user_guide.py +34 -0
- bayesbin-0.1.0/tools/longdouble_reference.py +79 -0
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
cff-version: 1.2.0
|
|
2
|
+
message: "If you use this software, please cite the paper below."
|
|
3
|
+
title: bayesbin
|
|
4
|
+
abstract: Exact Bayesian binning of rates in NumPy/SciPy.
|
|
5
|
+
type: software
|
|
6
|
+
license: BSD-3-Clause
|
|
7
|
+
repository-code: https://github.com/petfold/bayesbin
|
|
8
|
+
authors:
|
|
9
|
+
- family-names: Foldiak
|
|
10
|
+
given-names: Peter
|
|
11
|
+
preferred-citation:
|
|
12
|
+
type: conference-paper
|
|
13
|
+
title: "Bayesian binning beats approximate alternatives: estimating peri-stimulus time histograms"
|
|
14
|
+
authors:
|
|
15
|
+
- family-names: Endres
|
|
16
|
+
given-names: Dominik
|
|
17
|
+
- family-names: Oram
|
|
18
|
+
given-names: Mike
|
|
19
|
+
- family-names: Schindelin
|
|
20
|
+
given-names: Johannes
|
|
21
|
+
- family-names: Földiák
|
|
22
|
+
given-names: Peter
|
|
23
|
+
collection-title: Advances in Neural Information Processing Systems 20
|
|
24
|
+
publisher:
|
|
25
|
+
name: MIT Press
|
|
26
|
+
start: 393
|
|
27
|
+
end: 400
|
|
28
|
+
year: 2008
|
|
29
|
+
url: https://papers.nips.cc/paper_files/paper/2007/hash/b73ce398c39f506af761d2277d853a92-Abstract.html
|
bayesbin-0.1.0/LICENSE
ADDED
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
BSD 3-Clause License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026, Peter Foldiak
|
|
4
|
+
|
|
5
|
+
Redistribution and use in source and binary forms, with or without
|
|
6
|
+
modification, are permitted provided that the following conditions are met:
|
|
7
|
+
|
|
8
|
+
1. Redistributions of source code must retain the above copyright notice, this
|
|
9
|
+
list of conditions and the following disclaimer.
|
|
10
|
+
|
|
11
|
+
2. Redistributions in binary form must reproduce the above copyright notice,
|
|
12
|
+
this list of conditions and the following disclaimer in the documentation
|
|
13
|
+
and/or other materials provided with the distribution.
|
|
14
|
+
|
|
15
|
+
3. Neither the name of the copyright holder nor the names of its
|
|
16
|
+
contributors may be used to endorse or promote products derived from
|
|
17
|
+
this software without specific prior written permission.
|
|
18
|
+
|
|
19
|
+
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
|
20
|
+
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
|
21
|
+
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
|
22
|
+
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
|
23
|
+
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
|
24
|
+
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
|
25
|
+
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
|
26
|
+
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
|
27
|
+
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
|
28
|
+
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
|
@@ -0,0 +1,11 @@
|
|
|
1
|
+
# The source distribution: the Python package (BSD-3-Clause) with what its tests need.
|
|
2
|
+
# The C++ (cpp/, reference/, GPL-2.0-or-later) and the paper stay in the repository only;
|
|
3
|
+
# the tests against the C++ skip without it.
|
|
4
|
+
include LICENSE README.md CITATION.cff llms.txt
|
|
5
|
+
recursive-include docs *.md *.png
|
|
6
|
+
recursive-include tests *.py *.txt
|
|
7
|
+
include tools/longdouble_reference.py
|
|
8
|
+
prune cpp
|
|
9
|
+
prune reference
|
|
10
|
+
prune paper
|
|
11
|
+
exclude tools/make_testdata.py
|
bayesbin-0.1.0/PKG-INFO
ADDED
|
@@ -0,0 +1,289 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: bayesbin
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Exact Bayesian binning of rates (Endres, Oram, Schindelin & Földiák 2008) in NumPy/SciPy
|
|
5
|
+
Author: Peter Foldiak
|
|
6
|
+
License-Expression: BSD-3-Clause
|
|
7
|
+
Project-URL: Homepage, https://github.com/petfold/bayesbin
|
|
8
|
+
Project-URL: Paper, https://papers.nips.cc/paper_files/paper/2007/hash/b73ce398c39f506af761d2277d853a92-Abstract.html
|
|
9
|
+
Keywords: bayesian,binning,histogram,psth,spike trains,change points,rate estimation,poisson,segmentation
|
|
10
|
+
Classifier: Development Status :: 4 - Beta
|
|
11
|
+
Classifier: Intended Audience :: Science/Research
|
|
12
|
+
Classifier: Operating System :: OS Independent
|
|
13
|
+
Classifier: Programming Language :: Python :: 3
|
|
14
|
+
Classifier: Topic :: Scientific/Engineering :: Mathematics
|
|
15
|
+
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
|
|
16
|
+
Requires-Python: >=3.10
|
|
17
|
+
Description-Content-Type: text/markdown
|
|
18
|
+
License-File: LICENSE
|
|
19
|
+
Requires-Dist: numpy>=1.24
|
|
20
|
+
Requires-Dist: scipy>=1.10
|
|
21
|
+
Provides-Extra: fast
|
|
22
|
+
Requires-Dist: numba>=0.60; extra == "fast"
|
|
23
|
+
Requires-Dist: threadpoolctl>=3.1; extra == "fast"
|
|
24
|
+
Provides-Extra: test
|
|
25
|
+
Requires-Dist: pytest>=7; extra == "test"
|
|
26
|
+
Dynamic: license-file
|
|
27
|
+
|
|
28
|
+
# bayesbin
|
|
29
|
+
|
|
30
|
+
[](https://github.com/petfold/bayesbin/actions/workflows/tests.yml)
|
|
31
|
+
[](https://github.com/petfold/bayesbin/blob/main/LICENSE)
|
|
32
|
+
[](#status-and-limits)
|
|
33
|
+
|
|
34
|
+
Exact Bayesian binning of rates in NumPy/SciPy, after
|
|
35
|
+
|
|
36
|
+
> D. Endres, M. Oram, J. Schindelin, P. Földiák (2008). *Bayesian binning beats
|
|
37
|
+
> approximate alternatives: estimating peri-stimulus time histograms.*
|
|
38
|
+
> Advances in Neural Information Processing Systems 20, 393–400. MIT Press.
|
|
39
|
+
> ([NeurIPS page](https://papers.nips.cc/paper_files/paper/2007/hash/b73ce398c39f506af761d2277d853a92-Abstract.html); a copy in [paper/](https://github.com/petfold/bayesbin/tree/main/paper))
|
|
40
|
+
|
|
41
|
+
A rate on T ordered intervals is modelled as piecewise constant with M bin
|
|
42
|
+
boundaries. Boundary positions, per-bin rates (conjugate priors) and M itself
|
|
43
|
+
are all integrated out exactly: one forward dynamic programme gives the
|
|
44
|
+
evidence of every M in O(M·T²). A matching backward programme gives the
|
|
45
|
+
posterior of every candidate bin at once, from which the predictive rate, its
|
|
46
|
+
error bars and the posterior over boundary positions follow.
|
|
47
|
+
|
|
48
|
+
**New to this? Start with the [User Guide](https://github.com/petfold/bayesbin/blob/main/docs/USER_GUIDE.md)**: why fixed-width
|
|
49
|
+
bins mislead, the assumptions in plain words, four worked examples (spike
|
|
50
|
+
trains, counts with exposure, success rates, a daily profile) and the pitfalls.
|
|
51
|
+
For AI coding assistants there is a compact [llms.txt](https://github.com/petfold/bayesbin/blob/main/llms.txt).
|
|
52
|
+
|
|
53
|
+
```sh
|
|
54
|
+
pip install bayesbin # NumPy/SciPy only
|
|
55
|
+
pip install "bayesbin[fast]" # + numba kernels: ~2x faster, all cores
|
|
56
|
+
```
|
|
57
|
+
|
|
58
|
+
```python
|
|
59
|
+
from bayesbin import BernoulliModel, PoissonModel, fit, spike_counts
|
|
60
|
+
|
|
61
|
+
# spike trains, as in the paper: one list of integer spike times per trial
|
|
62
|
+
s, g = spike_counts(trials, t_start=-100, t_end=499)
|
|
63
|
+
r = fit(BernoulliModel(s, g, sigma=1.0, gamma=32.0), max_boundaries=10)
|
|
64
|
+
r.rate, r.rate_std # predictive firing probability per interval, ± 1 sd
|
|
65
|
+
r.m_posterior # P(M | data)
|
|
66
|
+
r.boundary_posterior # P(a bin ends at interval k | data)
|
|
67
|
+
|
|
68
|
+
# counts per window, several events per window allowed, with exposure
|
|
69
|
+
r = fit(PoissonModel.weak_prior(counts, exposure), max_boundaries=20)
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
By default predictions average over every M, as the paper recommends;
|
|
73
|
+
`m_mass=0.9` restricts them to the credible range of M, which is what the
|
|
74
|
+
original program does.
|
|
75
|
+
|
|
76
|
+
NumPy and SciPy are all it needs. With the `fast` extra (`pip install
|
|
77
|
+
bayesbin[fast]`: numba and threadpoolctl) the work runs in fused kernels on all
|
|
78
|
+
cores: about 2× faster on one core, and scaling to about 3× more on four. The
|
|
79
|
+
first call in a new environment compiles them (about a minute, once; cached
|
|
80
|
+
afterwards), and `BAYESBIN_NUMBA=0` switches them off. Both paths give the same
|
|
81
|
+
results to rounding, and the fused ones the same bits on any number of threads.
|
|
82
|
+
|
|
83
|
+
Threads: numba's, `NUMBA_NUM_THREADS` or `numba.set_num_threads(n)`; the
|
|
84
|
+
default is every logical CPU, and hyperthreads gain nothing here, so set it to
|
|
85
|
+
the number of physical cores for the best time. While a fit runs, the fused
|
|
86
|
+
path holds BLAS to one thread and splits the matrix products over numba's
|
|
87
|
+
threads itself. The kernels release the GIL, so many separate fits (many
|
|
88
|
+
short series) can also run in parallel from a thread pool, each on one numba
|
|
89
|
+
thread (under numba's TBB or OpenMP threading layer).
|
|
90
|
+
|
|
91
|
+
## What it is for
|
|
92
|
+
|
|
93
|
+
Any rate that varies along an ordered axis and is observed as events per
|
|
94
|
+
interval, where you want the rate *and* its uncertainty without choosing bin
|
|
95
|
+
widths by hand:
|
|
96
|
+
|
|
97
|
+
- **Peri-stimulus time histograms** (the paper's case): spike trains over
|
|
98
|
+
repeated trials, the firing probability per millisecond, with error bars,
|
|
99
|
+
from a few dozen trials.
|
|
100
|
+
- **Event counts per window**: arrivals, requests, incidents, photon or
|
|
101
|
+
particle counts, cases per week — with an exposure per window (observation
|
|
102
|
+
time, population, detector area) when windows differ.
|
|
103
|
+
- **Proportions along an axis**: successes out of trials per interval
|
|
104
|
+
(conversion or failure rates by time of day, by age, by dose), with the
|
|
105
|
+
Bernoulli model.
|
|
106
|
+
- **Change points**: `boundary_posterior` is the posterior probability that the
|
|
107
|
+
rate changes after each interval, averaged over every segmentation.
|
|
108
|
+
- **Periodic profiles**: daily or weekly shapes, with the days (or weeks) as
|
|
109
|
+
trials and the time of day as the axis.
|
|
110
|
+
|
|
111
|
+
Nothing is fitted by optimisation and nothing is sampled: every segmentation
|
|
112
|
+
into up to M + 1 bins, and every M, is summed exactly. Compared with
|
|
113
|
+
Bayesian Blocks (Scargle et al. 2013, ApJ 764:167; `astropy.stats.bayesian_blocks`),
|
|
114
|
+
which finds the single best segmentation under a penalty per block, this
|
|
115
|
+
averages over all segmentations, so the rate is smooth where the data do not
|
|
116
|
+
decide where a step is, and comes with error bars.
|
|
117
|
+
|
|
118
|
+
The axis is 1-D; see [docs/NOTES.md](https://github.com/petfold/bayesbin/blob/main/docs/NOTES.md) for how far this extends
|
|
119
|
+
to two dimensions.
|
|
120
|
+
|
|
121
|
+
## Models
|
|
122
|
+
|
|
123
|
+
| model | per interval | per-bin prior | use |
|
|
124
|
+
|---|---|---|---|
|
|
125
|
+
| `BernoulliModel(s, g, sigma, gamma)` | s trials with an event, g without | f ~ Beta(σ, γ) | the paper's PSTH |
|
|
126
|
+
| `PoissonModel(y, alpha, beta, e)` | count y over exposure e | λ ~ Gamma(α, β) | event counts per window |
|
|
127
|
+
|
|
128
|
+
`BernoulliModel` defaults to σ = 1, γ = 32, the original program's default.
|
|
129
|
+
|
|
130
|
+
## Verification
|
|
131
|
+
|
|
132
|
+
`pytest` (41 tests; 5 need `cpp/binsdfc-fb` built, 7 need the `fast` extra):
|
|
133
|
+
|
|
134
|
+
- **Against the original C++ program** (`binsdfc` 0.1, in [reference/](https://github.com/petfold/bayesbin/tree/main/reference)),
|
|
135
|
+
on a seeded dataset in its own input format (`tools/make_testdata.py`):
|
|
136
|
+
- log P(D | M) for M = 0..10 and the marginal likelihood agree to every
|
|
137
|
+
printed digit;
|
|
138
|
+
- the predictive rate agrees within 2 × 10⁻⁵ relative;
|
|
139
|
+
- its standard deviation agrees within 5 × 10⁻⁵. The original adds in log
|
|
140
|
+
space through an interpolated lookup table, which shows at that level.
|
|
141
|
+
- **Against brute-force enumeration** of every boundary configuration, for both
|
|
142
|
+
models: evidence and predictive rate to 10⁻¹⁰.
|
|
143
|
+
- **Against the paper's own device** (§4): P(spike | k) as the ratio of
|
|
144
|
+
evidences with and without a virtual spike at k equals the forward–backward
|
|
145
|
+
result for every k.
|
|
146
|
+
- The fast paths against the exact log-space ones, on data whose evidences span
|
|
147
|
+
thousands of nats, with the underflow fallback forced; the fused kernels
|
|
148
|
+
against the NumPy path and the exact one, for both models (constant and
|
|
149
|
+
varying exposure).
|
|
150
|
+
- **Against a long-double reference** (`tools/longdouble_reference.py`: the
|
|
151
|
+
whole computation in 80-bit arithmetic, the variance in its stable form),
|
|
152
|
+
on steps strong enough that the sd, √(E[f²] − E[f]²), cancels: every path's
|
|
153
|
+
rate to 10⁻¹², its sd to 10⁻⁹. (Both moments are divided by the computed
|
|
154
|
+
coverage, the posterior of the bins covering each interval, which is 1 in
|
|
155
|
+
exact arithmetic; its rounding error would otherwise reach the sd amplified
|
|
156
|
+
by rate²/var.)
|
|
157
|
+
- The User Guide's examples run and print what the guide says they print.
|
|
158
|
+
- One-bin evidences against direct numerical integration; every interval
|
|
159
|
+
covered by exactly one bin; the simulated response onset recovered.
|
|
160
|
+
|
|
161
|
+
## Status and limits
|
|
162
|
+
|
|
163
|
+
- Cost: O(M·T²) time, O(T·M) memory. Nothing of size T×T is formed on the
|
|
164
|
+
default path: the models give bin evidences and posterior moments block by
|
|
165
|
+
block (`bin_block`, from prefix sums and `lgamma` tables); the forward and
|
|
166
|
+
backward programmes run in blocks of 256 columns, each block's slice of
|
|
167
|
+
exponentiated gains made once from the upper triangle and serving all M
|
|
168
|
+
steps (the rows before a block, final for every step, as one matrix
|
|
169
|
+
product; only the rows inside it step by step); the bin posterior is
|
|
170
|
+
accumulated in tiles of 256 × 1024 bins, as scaled matrix products, into the
|
|
171
|
+
rates and the boundary posterior. The factors of every matrix product are
|
|
172
|
+
flushed to 0 below a threshold chosen so that no product is subnormal
|
|
173
|
+
(subnormal arithmetic is ~100× slower on x86, and BLAS runs without
|
|
174
|
+
flush-to-zero); the error bounds count the flushed terms as lost.
|
|
175
|
+
`keep_bins=True` (the whole bin posterior) and `exact=True` use T×T arrays. Wherever underflow could cost
|
|
176
|
+
more than 10⁻¹³ (relative, evidences) or 10⁻¹⁴ (absolute, bin posterior),
|
|
177
|
+
that entry is recomputed exactly in log space, and `exact=True` does
|
|
178
|
+
everything that way. See the timings below.
|
|
179
|
+
- Not yet ported from the original: latency posteriors, signal separation
|
|
180
|
+
levels, hyperparameter optimisation (`-P`), bin-boundary position posteriors
|
|
181
|
+
for a fixed M (`-p`).
|
|
182
|
+
- Planned: a release on PyPI, so that `pip install bayesbin` works (steps in
|
|
183
|
+
[docs/NOTES.md](https://github.com/petfold/bayesbin/blob/main/docs/NOTES.md#plan)); cyclic profiles (a bin may wrap round
|
|
184
|
+
the end of a day or week); 2-D via recursive partitions (see
|
|
185
|
+
[docs/NOTES.md](https://github.com/petfold/bayesbin/blob/main/docs/NOTES.md)).
|
|
186
|
+
|
|
187
|
+
## Speed against the original
|
|
188
|
+
|
|
189
|
+
binsdfc 0.1 unmodified (`g++ -O2 -fopenmp`; its own flags `-march=native
|
|
190
|
+
-ffast-math` made no real difference), against bayesbin with NumPy 2.5,
|
|
191
|
+
OpenBLAS and numba 0.67. Intel i7-3612QM (4 cores, 2 hyperthreads each; AVX,
|
|
192
|
+
no AVX2 or FMA); "1 core" means one physical core, and "4 threads" four
|
|
193
|
+
separate physical cores (`taskset`; numba, OpenMP and OpenBLAS thread counts
|
|
194
|
+
set to match). binsdfc is timed as a process, bayesbin in-process without the
|
|
195
|
+
import and after a warm-up call (the fused kernels' cache load, once per
|
|
196
|
+
process).
|
|
197
|
+
|
|
198
|
+
| case | binsdfc, 1 core | binsdfc, 4 threads | binsdfc-fb, 1 core | binsdfc-fb, 4 threads | bayesbin (NumPy), 1 core | bayesbin + numba, 1 core | bayesbin + numba, 4 threads |
|
|
199
|
+
|---|---|---|---|---|---|---|---|
|
|
200
|
+
| T=300, M≤10, rate ± sd | 0.98 s | 0.27 s | | | 0.018 s | 0.008 s | 0.006 s |
|
|
201
|
+
| T=600, M≤10, rate ± sd | 50.0 s | 12.4 s | 0.033 s | 0.021 s | 0.050 s | 0.023 s | 0.012 s |
|
|
202
|
+
| T=600, M≤10, evidence only | 0.065 s | — | 0.020 s | 0.017 s | 0.014 s | 0.008 s | 0.005 s |
|
|
203
|
+
| T=2016, M≤30, evidence only | 2.0 s | — | 0.14 s | 0.088 s | 0.14 s | 0.083 s | 0.035 s |
|
|
204
|
+
| T=2016, M≤30, rate ± sd | stopped after 26 min | | 0.33 s | 0.17 s | 0.51 s | 0.26 s | 0.099 s |
|
|
205
|
+
|
|
206
|
+
`binsdfc-fb` is the original with the forward–backward SDF, the matrix-vector
|
|
207
|
+
central iteration and table-driven bin evidences added ([cpp/binsdfc-fb/](https://github.com/petfold/bayesbin/blob/main/cpp/binsdfc-fb/README.fb.md)); best
|
|
208
|
+
of 5 runs. (binsdfc itself always runs 4 threads.)
|
|
209
|
+
|
|
210
|
+
- The evidences use the same dynamic programme in both. binsdfc's triple loop
|
|
211
|
+
is the slowest; binsdfc-fb and bayesbin both run its central iteration in
|
|
212
|
+
column blocks, bayesbin with the rows before each block as one matrix
|
|
213
|
+
product for all M steps (BLAS-3), which binsdfc-fb does not have.
|
|
214
|
+
- For the rate and its error bars binsdfc uses the paper's virtual-spike
|
|
215
|
+
device: for every time point it reruns the whole programme twice (rate and
|
|
216
|
+
second moment), O(M·T³). bayesbin gets every time point from one backward
|
|
217
|
+
pass, O(M·T²). The gap is the algorithm, not the language.
|
|
218
|
+
- binsdfc fixes 4 OpenMP threads over time points (`omp_set_num_threads(4)`)
|
|
219
|
+
and scales almost 4×. bayesbin's NumPy path runs on one thread outside BLAS;
|
|
220
|
+
its fused path runs every step on numba's threads (the matrix products split
|
|
221
|
+
into fixed parts) except the steps inside each 256-column block of the
|
|
222
|
+
dynamic programme, which are sequential in M, and scales about 2.9× on 4
|
|
223
|
+
cores.
|
|
224
|
+
|
|
225
|
+
### Larger problems
|
|
226
|
+
|
|
227
|
+
30 trials of a synthetic daily profile (5-minute slots: night, morning ramp,
|
|
228
|
+
day, evening peak, plus a 2-hour burst each week; `tools/make_longdata.py T`),
|
|
229
|
+
rate ± sd, most probable M only (`-l 0`, `m_mass=0.0`), peak memory from
|
|
230
|
+
`/usr/bin/time` and `getrusage`. All but the 12-week row run back to back, one
|
|
231
|
+
run each; the laptop was thermally throttled.
|
|
232
|
+
|
|
233
|
+
| T | M ≤ | binsdfc-fb, 1 core | binsdfc-fb, 4 threads | binsdfc-fb memory | bayesbin (NumPy), 1 core | bayesbin + numba, 1 core | bayesbin + numba, 4 threads |
|
|
234
|
+
|---|---|---|---|---|---|---|---|
|
|
235
|
+
| 2016 (1 week) | 30 | 0.25 s | 0.10 s | 12 MB | 0.54 s, 88 MB | 0.27 s, 175 MB | 0.11 s, 177 MB |
|
|
236
|
+
| 4032 (2 weeks) | 60 | 1.44 s | 0.45 s | 19 MB | 1.99 s, 110 MB | 1.18 s, 189 MB | 0.53 s, 192 MB |
|
|
237
|
+
| 8064 (4 weeks) | 120 | 9.2 s | 2.9 s | 44 MB | 8.7 s, 179 MB | 5.8 s, 235 MB | 2.0 s, 242 MB |
|
|
238
|
+
| 12096 (6 weeks) | 120 | 20.3 s | 6.3 s | 61 MB | 18.6 s, 250 MB | 13.0 s, 271 MB | 4.3 s, 280 MB |
|
|
239
|
+
| 24192 (12 weeks) | 120 | | 86 s | 116 MB | | | |
|
|
240
|
+
|
|
241
|
+
- binsdfc-fb memory is O(T·M): no T×T array is kept. (Before: ≈14·T²
|
|
242
|
+
bytes, 2.1 GB at 6 weeks; 12 weeks would have needed ≈8 GB.) Output is
|
|
243
|
+
byte-identical to the T² version at 2 and 6 weeks. bayesbin is O(T·M) too
|
|
244
|
+
now (its figures include ≈60 MB of Python and NumPy, and ≈90 MB more for
|
|
245
|
+
numba and LLVM); before, ≈70·T² bytes (1.1 GB at 2 weeks).
|
|
246
|
+
- Time grows about as T² and linearly in M; the original needs 8.5 s for the
|
|
247
|
+
evidences alone at T=4032.
|
|
248
|
+
- binsdfc-fb and bayesbin agree to the 6 printed digits at T=4032.
|
|
249
|
+
- A periodic series needs M to grow with its length: the most probable M was
|
|
250
|
+
106 at 2 weeks (M ≤ 120) and 297 at 6 weeks (M ≤ 400), the same daily shape
|
|
251
|
+
re-learnt every day. For daily or weekly profiles, fold the series instead
|
|
252
|
+
(days as trials, time of day as the axis): T = 288, a few milliseconds.
|
|
253
|
+
|
|
254
|
+
## Licence
|
|
255
|
+
|
|
256
|
+
Two licences, by directory:
|
|
257
|
+
|
|
258
|
+
- **BSD-3-Clause** ([LICENSE](https://github.com/petfold/bayesbin/blob/main/LICENSE)): the Python package `src/bayesbin` (all
|
|
259
|
+
that `pip` installs), its tests, `tools/bench_vs_binsdfc.py` and the
|
|
260
|
+
documentation. The package was written from the paper; the C++ program was
|
|
261
|
+
used only as a reference to test against, and none of its code is in it.
|
|
262
|
+
- **GPL-2.0-or-later**: [reference/binsdfc-0.1/](https://github.com/petfold/bayesbin/tree/main/reference) (Dominik Endres's
|
|
263
|
+
original, unmodified, with its provenance in [reference/README.md](https://github.com/petfold/bayesbin/blob/main/reference/README.md)),
|
|
264
|
+
[cpp/binsdfc-fb/](https://github.com/petfold/bayesbin/blob/main/cpp/binsdfc-fb/README.fb.md) (that program with a
|
|
265
|
+
forward–backward SDF and the speed-ups described there, each changed file
|
|
266
|
+
marked) and `tools/make_testdata.py` (a port of its test-data script). Their
|
|
267
|
+
licence text is in [cpp/COPYING](https://github.com/petfold/bayesbin/blob/main/cpp/COPYING) and
|
|
268
|
+
[reference/COPYING](https://github.com/petfold/bayesbin/blob/main/reference/COPYING).
|
|
269
|
+
- The paper in [paper/](https://github.com/petfold/bayesbin/tree/main/paper) is copyright its authors, not under either
|
|
270
|
+
licence.
|
|
271
|
+
|
|
272
|
+
## Citing
|
|
273
|
+
|
|
274
|
+
If you use this, please cite the paper:
|
|
275
|
+
|
|
276
|
+
```bibtex
|
|
277
|
+
@inproceedings{endres2008bayesian,
|
|
278
|
+
title = {Bayesian binning beats approximate alternatives: estimating peri-stimulus time histograms},
|
|
279
|
+
author = {Endres, Dominik and Oram, Mike and Schindelin, Johannes and F{\"o}ldi{\'a}k, Peter},
|
|
280
|
+
booktitle = {Advances in Neural Information Processing Systems 20},
|
|
281
|
+
pages = {393--400},
|
|
282
|
+
publisher = {MIT Press},
|
|
283
|
+
year = {2008}
|
|
284
|
+
}
|
|
285
|
+
```
|
|
286
|
+
|
|
287
|
+
The method is Endres, Oram, Schindelin and Földiák's; binsdfc, the original
|
|
288
|
+
C++ implementation, is Dominik Endres's; bayesbin and the binsdfc-fb changes
|
|
289
|
+
are by Peter Foldiak.
|
bayesbin-0.1.0/README.md
ADDED
|
@@ -0,0 +1,262 @@
|
|
|
1
|
+
# bayesbin
|
|
2
|
+
|
|
3
|
+
[](https://github.com/petfold/bayesbin/actions/workflows/tests.yml)
|
|
4
|
+
[](https://github.com/petfold/bayesbin/blob/main/LICENSE)
|
|
5
|
+
[](#status-and-limits)
|
|
6
|
+
|
|
7
|
+
Exact Bayesian binning of rates in NumPy/SciPy, after
|
|
8
|
+
|
|
9
|
+
> D. Endres, M. Oram, J. Schindelin, P. Földiák (2008). *Bayesian binning beats
|
|
10
|
+
> approximate alternatives: estimating peri-stimulus time histograms.*
|
|
11
|
+
> Advances in Neural Information Processing Systems 20, 393–400. MIT Press.
|
|
12
|
+
> ([NeurIPS page](https://papers.nips.cc/paper_files/paper/2007/hash/b73ce398c39f506af761d2277d853a92-Abstract.html); a copy in [paper/](https://github.com/petfold/bayesbin/tree/main/paper))
|
|
13
|
+
|
|
14
|
+
A rate on T ordered intervals is modelled as piecewise constant with M bin
|
|
15
|
+
boundaries. Boundary positions, per-bin rates (conjugate priors) and M itself
|
|
16
|
+
are all integrated out exactly: one forward dynamic programme gives the
|
|
17
|
+
evidence of every M in O(M·T²). A matching backward programme gives the
|
|
18
|
+
posterior of every candidate bin at once, from which the predictive rate, its
|
|
19
|
+
error bars and the posterior over boundary positions follow.
|
|
20
|
+
|
|
21
|
+
**New to this? Start with the [User Guide](https://github.com/petfold/bayesbin/blob/main/docs/USER_GUIDE.md)**: why fixed-width
|
|
22
|
+
bins mislead, the assumptions in plain words, four worked examples (spike
|
|
23
|
+
trains, counts with exposure, success rates, a daily profile) and the pitfalls.
|
|
24
|
+
For AI coding assistants there is a compact [llms.txt](https://github.com/petfold/bayesbin/blob/main/llms.txt).
|
|
25
|
+
|
|
26
|
+
```sh
|
|
27
|
+
pip install bayesbin # NumPy/SciPy only
|
|
28
|
+
pip install "bayesbin[fast]" # + numba kernels: ~2x faster, all cores
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
```python
|
|
32
|
+
from bayesbin import BernoulliModel, PoissonModel, fit, spike_counts
|
|
33
|
+
|
|
34
|
+
# spike trains, as in the paper: one list of integer spike times per trial
|
|
35
|
+
s, g = spike_counts(trials, t_start=-100, t_end=499)
|
|
36
|
+
r = fit(BernoulliModel(s, g, sigma=1.0, gamma=32.0), max_boundaries=10)
|
|
37
|
+
r.rate, r.rate_std # predictive firing probability per interval, ± 1 sd
|
|
38
|
+
r.m_posterior # P(M | data)
|
|
39
|
+
r.boundary_posterior # P(a bin ends at interval k | data)
|
|
40
|
+
|
|
41
|
+
# counts per window, several events per window allowed, with exposure
|
|
42
|
+
r = fit(PoissonModel.weak_prior(counts, exposure), max_boundaries=20)
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
By default predictions average over every M, as the paper recommends;
|
|
46
|
+
`m_mass=0.9` restricts them to the credible range of M, which is what the
|
|
47
|
+
original program does.
|
|
48
|
+
|
|
49
|
+
NumPy and SciPy are all it needs. With the `fast` extra (`pip install
|
|
50
|
+
bayesbin[fast]`: numba and threadpoolctl) the work runs in fused kernels on all
|
|
51
|
+
cores: about 2× faster on one core, and scaling to about 3× more on four. The
|
|
52
|
+
first call in a new environment compiles them (about a minute, once; cached
|
|
53
|
+
afterwards), and `BAYESBIN_NUMBA=0` switches them off. Both paths give the same
|
|
54
|
+
results to rounding, and the fused ones the same bits on any number of threads.
|
|
55
|
+
|
|
56
|
+
Threads: numba's, `NUMBA_NUM_THREADS` or `numba.set_num_threads(n)`; the
|
|
57
|
+
default is every logical CPU, and hyperthreads gain nothing here, so set it to
|
|
58
|
+
the number of physical cores for the best time. While a fit runs, the fused
|
|
59
|
+
path holds BLAS to one thread and splits the matrix products over numba's
|
|
60
|
+
threads itself. The kernels release the GIL, so many separate fits (many
|
|
61
|
+
short series) can also run in parallel from a thread pool, each on one numba
|
|
62
|
+
thread (under numba's TBB or OpenMP threading layer).
|
|
63
|
+
|
|
64
|
+
## What it is for
|
|
65
|
+
|
|
66
|
+
Any rate that varies along an ordered axis and is observed as events per
|
|
67
|
+
interval, where you want the rate *and* its uncertainty without choosing bin
|
|
68
|
+
widths by hand:
|
|
69
|
+
|
|
70
|
+
- **Peri-stimulus time histograms** (the paper's case): spike trains over
|
|
71
|
+
repeated trials, the firing probability per millisecond, with error bars,
|
|
72
|
+
from a few dozen trials.
|
|
73
|
+
- **Event counts per window**: arrivals, requests, incidents, photon or
|
|
74
|
+
particle counts, cases per week — with an exposure per window (observation
|
|
75
|
+
time, population, detector area) when windows differ.
|
|
76
|
+
- **Proportions along an axis**: successes out of trials per interval
|
|
77
|
+
(conversion or failure rates by time of day, by age, by dose), with the
|
|
78
|
+
Bernoulli model.
|
|
79
|
+
- **Change points**: `boundary_posterior` is the posterior probability that the
|
|
80
|
+
rate changes after each interval, averaged over every segmentation.
|
|
81
|
+
- **Periodic profiles**: daily or weekly shapes, with the days (or weeks) as
|
|
82
|
+
trials and the time of day as the axis.
|
|
83
|
+
|
|
84
|
+
Nothing is fitted by optimisation and nothing is sampled: every segmentation
|
|
85
|
+
into up to M + 1 bins, and every M, is summed exactly. Compared with
|
|
86
|
+
Bayesian Blocks (Scargle et al. 2013, ApJ 764:167; `astropy.stats.bayesian_blocks`),
|
|
87
|
+
which finds the single best segmentation under a penalty per block, this
|
|
88
|
+
averages over all segmentations, so the rate is smooth where the data do not
|
|
89
|
+
decide where a step is, and comes with error bars.
|
|
90
|
+
|
|
91
|
+
The axis is 1-D; see [docs/NOTES.md](https://github.com/petfold/bayesbin/blob/main/docs/NOTES.md) for how far this extends
|
|
92
|
+
to two dimensions.
|
|
93
|
+
|
|
94
|
+
## Models
|
|
95
|
+
|
|
96
|
+
| model | per interval | per-bin prior | use |
|
|
97
|
+
|---|---|---|---|
|
|
98
|
+
| `BernoulliModel(s, g, sigma, gamma)` | s trials with an event, g without | f ~ Beta(σ, γ) | the paper's PSTH |
|
|
99
|
+
| `PoissonModel(y, alpha, beta, e)` | count y over exposure e | λ ~ Gamma(α, β) | event counts per window |
|
|
100
|
+
|
|
101
|
+
`BernoulliModel` defaults to σ = 1, γ = 32, the original program's default.
|
|
102
|
+
|
|
103
|
+
## Verification
|
|
104
|
+
|
|
105
|
+
`pytest` (41 tests; 5 need `cpp/binsdfc-fb` built, 7 need the `fast` extra):
|
|
106
|
+
|
|
107
|
+
- **Against the original C++ program** (`binsdfc` 0.1, in [reference/](https://github.com/petfold/bayesbin/tree/main/reference)),
|
|
108
|
+
on a seeded dataset in its own input format (`tools/make_testdata.py`):
|
|
109
|
+
- log P(D | M) for M = 0..10 and the marginal likelihood agree to every
|
|
110
|
+
printed digit;
|
|
111
|
+
- the predictive rate agrees within 2 × 10⁻⁵ relative;
|
|
112
|
+
- its standard deviation agrees within 5 × 10⁻⁵. The original adds in log
|
|
113
|
+
space through an interpolated lookup table, which shows at that level.
|
|
114
|
+
- **Against brute-force enumeration** of every boundary configuration, for both
|
|
115
|
+
models: evidence and predictive rate to 10⁻¹⁰.
|
|
116
|
+
- **Against the paper's own device** (§4): P(spike | k) as the ratio of
|
|
117
|
+
evidences with and without a virtual spike at k equals the forward–backward
|
|
118
|
+
result for every k.
|
|
119
|
+
- The fast paths against the exact log-space ones, on data whose evidences span
|
|
120
|
+
thousands of nats, with the underflow fallback forced; the fused kernels
|
|
121
|
+
against the NumPy path and the exact one, for both models (constant and
|
|
122
|
+
varying exposure).
|
|
123
|
+
- **Against a long-double reference** (`tools/longdouble_reference.py`: the
|
|
124
|
+
whole computation in 80-bit arithmetic, the variance in its stable form),
|
|
125
|
+
on steps strong enough that the sd, √(E[f²] − E[f]²), cancels: every path's
|
|
126
|
+
rate to 10⁻¹², its sd to 10⁻⁹. (Both moments are divided by the computed
|
|
127
|
+
coverage, the posterior of the bins covering each interval, which is 1 in
|
|
128
|
+
exact arithmetic; its rounding error would otherwise reach the sd amplified
|
|
129
|
+
by rate²/var.)
|
|
130
|
+
- The User Guide's examples run and print what the guide says they print.
|
|
131
|
+
- One-bin evidences against direct numerical integration; every interval
|
|
132
|
+
covered by exactly one bin; the simulated response onset recovered.
|
|
133
|
+
|
|
134
|
+
## Status and limits
|
|
135
|
+
|
|
136
|
+
- Cost: O(M·T²) time, O(T·M) memory. Nothing of size T×T is formed on the
|
|
137
|
+
default path: the models give bin evidences and posterior moments block by
|
|
138
|
+
block (`bin_block`, from prefix sums and `lgamma` tables); the forward and
|
|
139
|
+
backward programmes run in blocks of 256 columns, each block's slice of
|
|
140
|
+
exponentiated gains made once from the upper triangle and serving all M
|
|
141
|
+
steps (the rows before a block, final for every step, as one matrix
|
|
142
|
+
product; only the rows inside it step by step); the bin posterior is
|
|
143
|
+
accumulated in tiles of 256 × 1024 bins, as scaled matrix products, into the
|
|
144
|
+
rates and the boundary posterior. The factors of every matrix product are
|
|
145
|
+
flushed to 0 below a threshold chosen so that no product is subnormal
|
|
146
|
+
(subnormal arithmetic is ~100× slower on x86, and BLAS runs without
|
|
147
|
+
flush-to-zero); the error bounds count the flushed terms as lost.
|
|
148
|
+
`keep_bins=True` (the whole bin posterior) and `exact=True` use T×T arrays. Wherever underflow could cost
|
|
149
|
+
more than 10⁻¹³ (relative, evidences) or 10⁻¹⁴ (absolute, bin posterior),
|
|
150
|
+
that entry is recomputed exactly in log space, and `exact=True` does
|
|
151
|
+
everything that way. See the timings below.
|
|
152
|
+
- Not yet ported from the original: latency posteriors, signal separation
|
|
153
|
+
levels, hyperparameter optimisation (`-P`), bin-boundary position posteriors
|
|
154
|
+
for a fixed M (`-p`).
|
|
155
|
+
- Planned: a release on PyPI, so that `pip install bayesbin` works (steps in
|
|
156
|
+
[docs/NOTES.md](https://github.com/petfold/bayesbin/blob/main/docs/NOTES.md#plan)); cyclic profiles (a bin may wrap round
|
|
157
|
+
the end of a day or week); 2-D via recursive partitions (see
|
|
158
|
+
[docs/NOTES.md](https://github.com/petfold/bayesbin/blob/main/docs/NOTES.md)).
|
|
159
|
+
|
|
160
|
+
## Speed against the original
|
|
161
|
+
|
|
162
|
+
binsdfc 0.1 unmodified (`g++ -O2 -fopenmp`; its own flags `-march=native
|
|
163
|
+
-ffast-math` made no real difference), against bayesbin with NumPy 2.5,
|
|
164
|
+
OpenBLAS and numba 0.67. Intel i7-3612QM (4 cores, 2 hyperthreads each; AVX,
|
|
165
|
+
no AVX2 or FMA); "1 core" means one physical core, and "4 threads" four
|
|
166
|
+
separate physical cores (`taskset`; numba, OpenMP and OpenBLAS thread counts
|
|
167
|
+
set to match). binsdfc is timed as a process, bayesbin in-process without the
|
|
168
|
+
import and after a warm-up call (the fused kernels' cache load, once per
|
|
169
|
+
process).
|
|
170
|
+
|
|
171
|
+
| case | binsdfc, 1 core | binsdfc, 4 threads | binsdfc-fb, 1 core | binsdfc-fb, 4 threads | bayesbin (NumPy), 1 core | bayesbin + numba, 1 core | bayesbin + numba, 4 threads |
|
|
172
|
+
|---|---|---|---|---|---|---|---|
|
|
173
|
+
| T=300, M≤10, rate ± sd | 0.98 s | 0.27 s | | | 0.018 s | 0.008 s | 0.006 s |
|
|
174
|
+
| T=600, M≤10, rate ± sd | 50.0 s | 12.4 s | 0.033 s | 0.021 s | 0.050 s | 0.023 s | 0.012 s |
|
|
175
|
+
| T=600, M≤10, evidence only | 0.065 s | — | 0.020 s | 0.017 s | 0.014 s | 0.008 s | 0.005 s |
|
|
176
|
+
| T=2016, M≤30, evidence only | 2.0 s | — | 0.14 s | 0.088 s | 0.14 s | 0.083 s | 0.035 s |
|
|
177
|
+
| T=2016, M≤30, rate ± sd | stopped after 26 min | | 0.33 s | 0.17 s | 0.51 s | 0.26 s | 0.099 s |
|
|
178
|
+
|
|
179
|
+
`binsdfc-fb` is the original with the forward–backward SDF, the matrix-vector
|
|
180
|
+
central iteration and table-driven bin evidences added ([cpp/binsdfc-fb/](https://github.com/petfold/bayesbin/blob/main/cpp/binsdfc-fb/README.fb.md)); best
|
|
181
|
+
of 5 runs. (binsdfc itself always runs 4 threads.)
|
|
182
|
+
|
|
183
|
+
- The evidences use the same dynamic programme in both. binsdfc's triple loop
|
|
184
|
+
is the slowest; binsdfc-fb and bayesbin both run its central iteration in
|
|
185
|
+
column blocks, bayesbin with the rows before each block as one matrix
|
|
186
|
+
product for all M steps (BLAS-3), which binsdfc-fb does not have.
|
|
187
|
+
- For the rate and its error bars binsdfc uses the paper's virtual-spike
|
|
188
|
+
device: for every time point it reruns the whole programme twice (rate and
|
|
189
|
+
second moment), O(M·T³). bayesbin gets every time point from one backward
|
|
190
|
+
pass, O(M·T²). The gap is the algorithm, not the language.
|
|
191
|
+
- binsdfc fixes 4 OpenMP threads over time points (`omp_set_num_threads(4)`)
|
|
192
|
+
and scales almost 4×. bayesbin's NumPy path runs on one thread outside BLAS;
|
|
193
|
+
its fused path runs every step on numba's threads (the matrix products split
|
|
194
|
+
into fixed parts) except the steps inside each 256-column block of the
|
|
195
|
+
dynamic programme, which are sequential in M, and scales about 2.9× on 4
|
|
196
|
+
cores.
|
|
197
|
+
|
|
198
|
+
### Larger problems
|
|
199
|
+
|
|
200
|
+
30 trials of a synthetic daily profile (5-minute slots: night, morning ramp,
|
|
201
|
+
day, evening peak, plus a 2-hour burst each week; `tools/make_longdata.py T`),
|
|
202
|
+
rate ± sd, most probable M only (`-l 0`, `m_mass=0.0`), peak memory from
|
|
203
|
+
`/usr/bin/time` and `getrusage`. All but the 12-week row run back to back, one
|
|
204
|
+
run each; the laptop was thermally throttled.
|
|
205
|
+
|
|
206
|
+
| T | M ≤ | binsdfc-fb, 1 core | binsdfc-fb, 4 threads | binsdfc-fb memory | bayesbin (NumPy), 1 core | bayesbin + numba, 1 core | bayesbin + numba, 4 threads |
|
|
207
|
+
|---|---|---|---|---|---|---|---|
|
|
208
|
+
| 2016 (1 week) | 30 | 0.25 s | 0.10 s | 12 MB | 0.54 s, 88 MB | 0.27 s, 175 MB | 0.11 s, 177 MB |
|
|
209
|
+
| 4032 (2 weeks) | 60 | 1.44 s | 0.45 s | 19 MB | 1.99 s, 110 MB | 1.18 s, 189 MB | 0.53 s, 192 MB |
|
|
210
|
+
| 8064 (4 weeks) | 120 | 9.2 s | 2.9 s | 44 MB | 8.7 s, 179 MB | 5.8 s, 235 MB | 2.0 s, 242 MB |
|
|
211
|
+
| 12096 (6 weeks) | 120 | 20.3 s | 6.3 s | 61 MB | 18.6 s, 250 MB | 13.0 s, 271 MB | 4.3 s, 280 MB |
|
|
212
|
+
| 24192 (12 weeks) | 120 | | 86 s | 116 MB | | | |
|
|
213
|
+
|
|
214
|
+
- binsdfc-fb memory is O(T·M): no T×T array is kept. (Before: ≈14·T²
|
|
215
|
+
bytes, 2.1 GB at 6 weeks; 12 weeks would have needed ≈8 GB.) Output is
|
|
216
|
+
byte-identical to the T² version at 2 and 6 weeks. bayesbin is O(T·M) too
|
|
217
|
+
now (its figures include ≈60 MB of Python and NumPy, and ≈90 MB more for
|
|
218
|
+
numba and LLVM); before, ≈70·T² bytes (1.1 GB at 2 weeks).
|
|
219
|
+
- Time grows about as T² and linearly in M; the original needs 8.5 s for the
|
|
220
|
+
evidences alone at T=4032.
|
|
221
|
+
- binsdfc-fb and bayesbin agree to the 6 printed digits at T=4032.
|
|
222
|
+
- A periodic series needs M to grow with its length: the most probable M was
|
|
223
|
+
106 at 2 weeks (M ≤ 120) and 297 at 6 weeks (M ≤ 400), the same daily shape
|
|
224
|
+
re-learnt every day. For daily or weekly profiles, fold the series instead
|
|
225
|
+
(days as trials, time of day as the axis): T = 288, a few milliseconds.
|
|
226
|
+
|
|
227
|
+
## Licence
|
|
228
|
+
|
|
229
|
+
Two licences, by directory:
|
|
230
|
+
|
|
231
|
+
- **BSD-3-Clause** ([LICENSE](https://github.com/petfold/bayesbin/blob/main/LICENSE)): the Python package `src/bayesbin` (all
|
|
232
|
+
that `pip` installs), its tests, `tools/bench_vs_binsdfc.py` and the
|
|
233
|
+
documentation. The package was written from the paper; the C++ program was
|
|
234
|
+
used only as a reference to test against, and none of its code is in it.
|
|
235
|
+
- **GPL-2.0-or-later**: [reference/binsdfc-0.1/](https://github.com/petfold/bayesbin/tree/main/reference) (Dominik Endres's
|
|
236
|
+
original, unmodified, with its provenance in [reference/README.md](https://github.com/petfold/bayesbin/blob/main/reference/README.md)),
|
|
237
|
+
[cpp/binsdfc-fb/](https://github.com/petfold/bayesbin/blob/main/cpp/binsdfc-fb/README.fb.md) (that program with a
|
|
238
|
+
forward–backward SDF and the speed-ups described there, each changed file
|
|
239
|
+
marked) and `tools/make_testdata.py` (a port of its test-data script). Their
|
|
240
|
+
licence text is in [cpp/COPYING](https://github.com/petfold/bayesbin/blob/main/cpp/COPYING) and
|
|
241
|
+
[reference/COPYING](https://github.com/petfold/bayesbin/blob/main/reference/COPYING).
|
|
242
|
+
- The paper in [paper/](https://github.com/petfold/bayesbin/tree/main/paper) is copyright its authors, not under either
|
|
243
|
+
licence.
|
|
244
|
+
|
|
245
|
+
## Citing
|
|
246
|
+
|
|
247
|
+
If you use this, please cite the paper:
|
|
248
|
+
|
|
249
|
+
```bibtex
|
|
250
|
+
@inproceedings{endres2008bayesian,
|
|
251
|
+
title = {Bayesian binning beats approximate alternatives: estimating peri-stimulus time histograms},
|
|
252
|
+
author = {Endres, Dominik and Oram, Mike and Schindelin, Johannes and F{\"o}ldi{\'a}k, Peter},
|
|
253
|
+
booktitle = {Advances in Neural Information Processing Systems 20},
|
|
254
|
+
pages = {393--400},
|
|
255
|
+
publisher = {MIT Press},
|
|
256
|
+
year = {2008}
|
|
257
|
+
}
|
|
258
|
+
```
|
|
259
|
+
|
|
260
|
+
The method is Endres, Oram, Schindelin and Földiák's; binsdfc, the original
|
|
261
|
+
C++ implementation, is Dominik Endres's; bayesbin and the binsdfc-fb changes
|
|
262
|
+
are by Peter Foldiak.
|