phonesim 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- phonesim-0.1.0/CHANGELOG.md +42 -0
- phonesim-0.1.0/CONTRIBUTING.md +12 -0
- phonesim-0.1.0/LICENSE +21 -0
- phonesim-0.1.0/MANIFEST.in +3 -0
- phonesim-0.1.0/PKG-INFO +473 -0
- phonesim-0.1.0/README.md +435 -0
- phonesim-0.1.0/SECURITY.md +6 -0
- phonesim-0.1.0/examples/degrade.py +58 -0
- phonesim-0.1.0/phonesim/__init__.py +83 -0
- phonesim-0.1.0/phonesim/analysis.py +331 -0
- phonesim-0.1.0/phonesim/cli.py +225 -0
- phonesim-0.1.0/phonesim/config.py +130 -0
- phonesim-0.1.0/phonesim/core.py +298 -0
- phonesim-0.1.0/phonesim/dsp.py +228 -0
- phonesim-0.1.0/phonesim/ffmpeg_backend.py +304 -0
- phonesim-0.1.0/phonesim/io_utils.py +70 -0
- phonesim-0.1.0/phonesim/opus_backend.py +248 -0
- phonesim-0.1.0/phonesim/plc.py +206 -0
- phonesim-0.1.0/phonesim/profiles/__init__.py +289 -0
- phonesim-0.1.0/phonesim/simulator.py +204 -0
- phonesim-0.1.0/phonesim/stages/__init__.py +40 -0
- phonesim-0.1.0/phonesim/stages/ambient.py +45 -0
- phonesim-0.1.0/phonesim/stages/channel_edge.py +109 -0
- phonesim-0.1.0/phonesim/stages/codec.py +212 -0
- phonesim-0.1.0/phonesim/stages/companding.py +93 -0
- phonesim-0.1.0/phonesim/stages/filtering.py +60 -0
- phonesim-0.1.0/phonesim/stages/gain.py +114 -0
- phonesim-0.1.0/phonesim/stages/level.py +77 -0
- phonesim-0.1.0/phonesim/stages/misc.py +75 -0
- phonesim-0.1.0/phonesim/stages/noise.py +60 -0
- phonesim-0.1.0/phonesim/stages/packet.py +180 -0
- phonesim-0.1.0/phonesim/stages/resample.py +43 -0
- phonesim-0.1.0/phonesim/stages/timing.py +140 -0
- phonesim-0.1.0/phonesim.egg-info/PKG-INFO +473 -0
- phonesim-0.1.0/phonesim.egg-info/SOURCES.txt +53 -0
- phonesim-0.1.0/phonesim.egg-info/dependency_links.txt +1 -0
- phonesim-0.1.0/phonesim.egg-info/entry_points.txt +2 -0
- phonesim-0.1.0/phonesim.egg-info/requires.txt +26 -0
- phonesim-0.1.0/phonesim.egg-info/top_level.txt +1 -0
- phonesim-0.1.0/pyproject.toml +50 -0
- phonesim-0.1.0/setup.cfg +4 -0
- phonesim-0.1.0/tests/__init__.py +0 -0
- phonesim-0.1.0/tests/data/profile_fingerprints.json +879 -0
- phonesim-0.1.0/tests/data/pstn_narrowband@1.npy +0 -0
- phonesim-0.1.0/tests/data/run_logs.json +85 -0
- phonesim-0.1.0/tests/regen_goldens.py +80 -0
- phonesim-0.1.0/tests/test_cli_rates.py +231 -0
- phonesim-0.1.0/tests/test_config.py +107 -0
- phonesim-0.1.0/tests/test_ffmpeg_erasure.py +129 -0
- phonesim-0.1.0/tests/test_g711.py +54 -0
- phonesim-0.1.0/tests/test_opus_backend.py +199 -0
- phonesim-0.1.0/tests/test_phonesim.py +989 -0
- phonesim-0.1.0/tests/test_plc.py +130 -0
- phonesim-0.1.0/tests/test_timing.py +88 -0
- phonesim-0.1.0/tests/test_versions.py +166 -0
|
@@ -0,0 +1,42 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 0.1.0
|
|
4
|
+
|
|
5
|
+
- Seven profiles, each a seeded model of one call path: `pstn_narrowband`
|
|
6
|
+
(analogue loop into a G.711 exchange), `pstn_g726` (PSTN over G.726 ADPCM),
|
|
7
|
+
`voip_opus_wideband` (WebRTC/VoIP over Opus), `voip_g722_wideband` (SIP HD
|
|
8
|
+
voice over G.722), `voip_to_cellular_wideband` (VoIP into an AMR-WB HD-voice
|
|
9
|
+
call), `voip_to_cellular_narrowband` (VoIP into an ordinary AMR-NB mobile
|
|
10
|
+
call; the default) and `stress_multi_transcode` (a synthetic stress chain,
|
|
11
|
+
not a real route).
|
|
12
|
+
- Real codecs only: G.711 with the ITU-T segmented coder in-process; G.722,
|
|
13
|
+
G.726, AMR-NB and AMR-WB through ffmpeg; Opus through ffmpeg or the libopus
|
|
14
|
+
shared library. AMR and Opus decode with `libopencore_amrnb`,
|
|
15
|
+
`libopencore_amrwb` and `libopus`; algorithmic delay is compensated. A
|
|
16
|
+
profile whose codec this machine cannot run raises `CodecUnavailableError`
|
|
17
|
+
when built.
|
|
18
|
+
- Frame erasures in `CodecStage`, concealed on the receiving side. AMR and
|
|
19
|
+
Opus frames are removed from the coded stream before the decoder: AMR runs
|
|
20
|
+
its decoder's error concealment, which carries the error into the following
|
|
21
|
+
frames; libopus decodes from the next packet's in-band FEC when it carries
|
|
22
|
+
one, otherwise its PLC. G.711, G.722 and G.726 are decoded in full and the
|
|
23
|
+
erased frames are replaced in the decoded PCM by the ITU-T G.711 Appendix I
|
|
24
|
+
waveform substitution (`phonesim.plc`).
|
|
25
|
+
- Send-side speech level and ambient noise, codec- and handset-defined band
|
|
26
|
+
edges, adaptive playout, clock drift in ppm, and the stress chain's
|
|
27
|
+
packet-loss, jitter-buffer, AGC, noise, clipping, speed-drift and
|
|
28
|
+
time-offset stages.
|
|
29
|
+
- Profile versions: `name@N`; a bare name is the latest. A version fixes the
|
|
30
|
+
stage chain and its parameters; the codec build and decoder are recorded in
|
|
31
|
+
the run log, not versioned. `phonesim info` lists versions.
|
|
32
|
+
- Reproducibility: a seed drives one `torch.Generator`; the same input, seed,
|
|
33
|
+
profile version and codec build give the same output on one machine and
|
|
34
|
+
torch thread count. `per_example=True` gives each row of a batch its own
|
|
35
|
+
call, seeded by `row_seeds`. The run log has one line per stage.
|
|
36
|
+
- `analyze_channel` (aligned SNR, band energies, high-frequency energy,
|
|
37
|
+
optional PESQ and STOI, caller-supplied `metrics=`) and `plot_channel`.
|
|
38
|
+
- `phonesim` CLI (`info`, `run`, `analyze`, `batch`) and YAML/JSON pipeline
|
|
39
|
+
configs.
|
|
40
|
+
- Tests pin every profile version's stage chain, parameters and run log; CI
|
|
41
|
+
runs them on a distro ffmpeg (AMR tests skip) and on a static AMR-capable
|
|
42
|
+
build pinned by sha256.
|
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
# Contributing
|
|
2
|
+
|
|
3
|
+
- Open a pull request against `main`; every change is reviewed.
|
|
4
|
+
- Run `pytest -q` before pushing, with an AMR-capable ffmpeg on `PATH` and libopus
|
|
5
|
+
installed if you can (see README); without them the AMR and Opus-erasure tests
|
|
6
|
+
skip.
|
|
7
|
+
- A change to a profile's stage chain or parameters is a new profile version,
|
|
8
|
+
never an edit in place; `pytest` enforces it against
|
|
9
|
+
`tests/data/profile_fingerprints.json`, and `python tests/regen_goldens.py`
|
|
10
|
+
adds the entries for a new version. The codec implementation (ffmpeg build,
|
|
11
|
+
decoder) is logged, not versioned.
|
|
12
|
+
- Claims about realism need a measurement in the PR, not an adjective.
|
phonesim-0.1.0/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 DeepMark Inc.
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
phonesim-0.1.0/PKG-INFO
ADDED
|
@@ -0,0 +1,473 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: phonesim
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Provider-free phone-call audio-degradation simulator with real telephony codecs
|
|
5
|
+
Author: DeepMark Inc.
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/deepmark/phonesim
|
|
8
|
+
Project-URL: Issues, https://github.com/deepmark/phonesim/issues
|
|
9
|
+
Keywords: audio,telephony,codec,simulation,dsp
|
|
10
|
+
Classifier: Programming Language :: Python :: 3
|
|
11
|
+
Classifier: Topic :: Multimedia :: Sound/Audio
|
|
12
|
+
Classifier: Topic :: Scientific/Engineering
|
|
13
|
+
Requires-Python: >=3.10
|
|
14
|
+
Description-Content-Type: text/markdown
|
|
15
|
+
License-File: LICENSE
|
|
16
|
+
Requires-Dist: numpy>=1.21
|
|
17
|
+
Requires-Dist: torch>=2.0
|
|
18
|
+
Requires-Dist: soundfile>=0.11
|
|
19
|
+
Provides-Extra: yaml
|
|
20
|
+
Requires-Dist: pyyaml>=5.4; extra == "yaml"
|
|
21
|
+
Provides-Extra: plot
|
|
22
|
+
Requires-Dist: matplotlib>=3.4; extra == "plot"
|
|
23
|
+
Provides-Extra: metrics
|
|
24
|
+
Requires-Dist: pesq>=0.0.4; extra == "metrics"
|
|
25
|
+
Requires-Dist: pystoi>=0.3; extra == "metrics"
|
|
26
|
+
Provides-Extra: test
|
|
27
|
+
Requires-Dist: pytest>=7.0; extra == "test"
|
|
28
|
+
Requires-Dist: pyyaml>=5.4; extra == "test"
|
|
29
|
+
Requires-Dist: pystoi>=0.3; extra == "test"
|
|
30
|
+
Requires-Dist: matplotlib>=3.4; extra == "test"
|
|
31
|
+
Provides-Extra: full
|
|
32
|
+
Requires-Dist: pyyaml>=5.4; extra == "full"
|
|
33
|
+
Requires-Dist: matplotlib>=3.4; extra == "full"
|
|
34
|
+
Requires-Dist: pesq>=0.0.4; extra == "full"
|
|
35
|
+
Requires-Dist: pystoi>=0.3; extra == "full"
|
|
36
|
+
Requires-Dist: pytest>=7.0; extra == "full"
|
|
37
|
+
Dynamic: license-file
|
|
38
|
+
|
|
39
|
+
# phonesim
|
|
40
|
+
|
|
41
|
+
A provider-free phone-call audio-degradation simulator: a local, seeded model
|
|
42
|
+
of the signal path of a telephone call, running the real telephony codecs.
|
|
43
|
+
|
|
44
|
+
`phonesim` recreates, locally and with no third-party telephony services (no
|
|
45
|
+
Twilio / Telnyx / Vonage), the chain of distortions a signal accumulates when it
|
|
46
|
+
travels through a real phone call: VoIP/WebRTC transport, PSTN and cellular
|
|
47
|
+
interconnects, mobile voice codecs, packet loss, jitter, automatic gain control,
|
|
48
|
+
level control, background noise, and clock drift. **The codecs are real**:
|
|
49
|
+
G.722, G.726 and (with an AMR-capable ffmpeg build) AMR-NB and AMR-WB run
|
|
50
|
+
through ffmpeg, Opus through ffmpeg or the libopus library; G.711 companding is
|
|
51
|
+
the ITU-T segmented coder, in-process.
|
|
52
|
+
No codec is approximated: a profile whose codec this machine cannot run fails
|
|
53
|
+
when built, with the install hint.
|
|
54
|
+
|
|
55
|
+
Input and output default to **24 kHz**; both rates are parameters. Batches
|
|
56
|
+
(`[B, T]`) can be processed as one call or, with `per_example=True`, as one
|
|
57
|
+
independent call per row, which is the mode for data generation.
|
|
58
|
+
|
|
59
|
+
---
|
|
60
|
+
|
|
61
|
+
## Why this exists
|
|
62
|
+
|
|
63
|
+
A phone call band-limits the signal to a few kHz, compresses it with a lossy
|
|
64
|
+
speech codec (often more than once when the call crosses network boundaries),
|
|
65
|
+
chops it into packets that can be lost or delayed, re-levels it, and re-digitises
|
|
66
|
+
it at the far end. Any audio system that has to work over calls (speech
|
|
67
|
+
recognition, speaker verification, watermark detection, enhancement) needs to be
|
|
68
|
+
measured and trained against that channel.
|
|
69
|
+
|
|
70
|
+
Doing so with real calls means a telephony provider, cost and no
|
|
71
|
+
reproducibility. `phonesim` is a local, seeded model of the same signal path,
|
|
72
|
+
so you can measure over representative paths, generate degraded data at scale,
|
|
73
|
+
and reproduce a result from a seed.
|
|
74
|
+
|
|
75
|
+
---
|
|
76
|
+
|
|
77
|
+
## What it does and does not claim to do
|
|
78
|
+
|
|
79
|
+
**It models** the *signal-level* transformations of a call path: resampling,
|
|
80
|
+
the channel's band edges, real codec encode/decode, frame erasures with
|
|
81
|
+
receiver-side concealment, playout-buffer behaviour, send-side level control, ambient
|
|
82
|
+
noise, multi-transcode chains, and clock drift.
|
|
83
|
+
|
|
84
|
+
**Every codec is real.** A stock ffmpeg covers G.711, G.722, G.726 and Opus; an
|
|
85
|
+
**AMR-capable ffmpeg build** (libopencore-amr + libvo-amrwbenc) adds the real
|
|
86
|
+
**AMR-NB and AMR-WB** cellular codecs, the ones that carry mobile voice. AMR
|
|
87
|
+
and Opus are decoded with `libopencore_amrnb` / `libopencore_amrwb` / `libopus`
|
|
88
|
+
rather than ffmpeg's own decoders, and the run log records the ffmpeg version,
|
|
89
|
+
decoder and mode that ran. There is no approximation to fall back to: building
|
|
90
|
+
a profile whose codec is missing raises `CodecUnavailableError`. EVS has no
|
|
91
|
+
open encoder and ffmpeg has no G.729 encoder, so neither is offered.
|
|
92
|
+
|
|
93
|
+
---
|
|
94
|
+
|
|
95
|
+
## Installation
|
|
96
|
+
|
|
97
|
+
```bash
|
|
98
|
+
# from the repo root:
|
|
99
|
+
pip install -e . # core: numpy, torch, soundfile
|
|
100
|
+
pip install -e ".[full]" # + pyyaml, matplotlib, pesq, pystoi, pytest
|
|
101
|
+
|
|
102
|
+
# codec profiles need ffmpeg on the PATH:
|
|
103
|
+
# apt-get install ffmpeg (or: brew install ffmpeg)
|
|
104
|
+
# the default profile also needs the AMR encoders and libopus (both below)
|
|
105
|
+
```
|
|
106
|
+
|
|
107
|
+
This installs the package (so `import phonesim` works from anywhere) and a
|
|
108
|
+
`phonesim` console command. For GPU use, install the torch build matching your
|
|
109
|
+
CUDA version from <https://pytorch.org>. Quick check:
|
|
110
|
+
|
|
111
|
+
```python
|
|
112
|
+
import phonesim
|
|
113
|
+
print(phonesim.list_profiles())
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
### Real codecs via ffmpeg
|
|
117
|
+
|
|
118
|
+
`phonesim` probes which codecs the ffmpeg on `PATH` (or the one named by
|
|
119
|
+
`PHONESIM_FFMPEG`) can round-trip:
|
|
120
|
+
|
|
121
|
+
```python
|
|
122
|
+
from phonesim import ffmpeg_backend
|
|
123
|
+
print(sorted(ffmpeg_backend.available_codecs()))
|
|
124
|
+
# stock ffmpeg → ['g711_alaw', 'g711_ulaw', 'g722', 'g726', 'opus']
|
|
125
|
+
# AMR-capable → … plus 'amr_nb' and 'amr_wb'
|
|
126
|
+
```
|
|
127
|
+
|
|
128
|
+
A **stock distro ffmpeg cannot run AMR**: it has neither the encoders
|
|
129
|
+
(`libopencore-amrnb`, `libvo-amrwbenc`) nor the OpenCORE decoders. The
|
|
130
|
+
no-compile path on Linux is a static GPL build:
|
|
131
|
+
|
|
132
|
+
```bash
|
|
133
|
+
base=https://johnvansickle.com/ffmpeg/releases/ffmpeg-release-amd64-static.tar.xz
|
|
134
|
+
curl -fL -o ffmpeg-release-amd64-static.tar.xz "$base"
|
|
135
|
+
curl -fL -o ffmpeg-release-amd64-static.tar.xz.md5 "$base.md5"
|
|
136
|
+
md5sum -c ffmpeg-release-amd64-static.tar.xz.md5
|
|
137
|
+
tar xf ffmpeg-release-amd64-static.tar.xz
|
|
138
|
+
install ffmpeg-*-amd64-static/ffmpeg ~/.local/bin/ffmpeg # ~/.local/bin on PATH
|
|
139
|
+
|
|
140
|
+
ffmpeg -hide_banner -encoders | grep -i amr
|
|
141
|
+
# A....D libopencore_amrnb OpenCORE AMR-NB ... (codec amr_nb)
|
|
142
|
+
# A....D libvo_amrwbenc Android VisualOn AMR-WB ... (codec amr_wb)
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
macOS: Homebrew's `ffmpeg` is built without AMR; build ffmpeg with
|
|
146
|
+
`--enable-version3 --enable-libopencore-amrnb --enable-libopencore-amrwb --enable-libvo-amrwbenc`.
|
|
147
|
+
|
|
148
|
+
Licensing: phonesim is MIT and links nothing at build time; it calls the `ffmpeg`
|
|
149
|
+
binary at runtime. The AMR libraries (opencore-amr, vo-amrwbenc) are
|
|
150
|
+
Apache-2.0; ffmpeg accepts them only with `--enable-version3`, which makes the
|
|
151
|
+
ffmpeg build LGPLv3, or GPLv3 when `--enable-gpl` is also set (the static build
|
|
152
|
+
above sets both, so it is GPLv3). Distributing such a binary together with your
|
|
153
|
+
software carries that licence's obligations; the AMR codecs are also subject to
|
|
154
|
+
patents in some jurisdictions.
|
|
155
|
+
|
|
156
|
+
Without AMR, the `voip_to_cellular_*` and `stress_multi_transcode` profiles
|
|
157
|
+
raise `CodecUnavailableError` when built; the PSTN and G.722 profiles work with
|
|
158
|
+
a stock ffmpeg, `voip_opus_wideband` also needs libopus (next section).
|
|
159
|
+
|
|
160
|
+
### Opus erasures need libopus
|
|
161
|
+
|
|
162
|
+
Erasing Opus frames runs Opus in-process through the libopus shared library
|
|
163
|
+
(`apt-get install libopus0`, `brew install opus`; system and Homebrew paths are
|
|
164
|
+
searched, `PHONESIM_LIBOPUS` names any other file), so the decoder's own
|
|
165
|
+
concealment and in-band FEC apply. Without it, `voip_opus_wideband` and the
|
|
166
|
+
two `voip_to_cellular_*` profiles raise `CodecUnavailableError`.
|
|
167
|
+
|
|
168
|
+
> **Device support:** CPU and NVIDIA CUDA; torch tensors are processed on the
|
|
169
|
+
> device they arrive on. Codecs always run on the CPU (ffmpeg, libopus or the
|
|
170
|
+
> in-process G.711 coder).
|
|
171
|
+
|
|
172
|
+
---
|
|
173
|
+
|
|
174
|
+
## Quick start
|
|
175
|
+
|
|
176
|
+
### Degrade an audio file
|
|
177
|
+
|
|
178
|
+
```python
|
|
179
|
+
from phonesim import PhoneCallSimulator, load_audio, save_audio
|
|
180
|
+
|
|
181
|
+
x, sr = load_audio("input.wav", sr=24000) # loads & resamples to 24 kHz
|
|
182
|
+
sim = PhoneCallSimulator(profile="voip_to_cellular_narrowband")
|
|
183
|
+
y = sim(x, seed=1234) # reproducible degraded audio
|
|
184
|
+
save_audio("degraded.wav", y, sr=24000)
|
|
185
|
+
```
|
|
186
|
+
|
|
187
|
+
`sim(x)` accepts and returns a NumPy array **or** a torch tensor, of rank
|
|
188
|
+
`[T]`, `[B, T]`, or `[B, C, T]`, and gives back the same type and rank at 24 kHz.
|
|
189
|
+
`sim(x, seed=1234, per_example=True)` gives each row of a batch its own call.
|
|
190
|
+
|
|
191
|
+
---
|
|
192
|
+
|
|
193
|
+
|
|
194
|
+
## The built-in profiles
|
|
195
|
+
|
|
196
|
+
A *profile* is a named combination of stages modelling one call path. All
|
|
197
|
+
start and end at 24 kHz by default.
|
|
198
|
+
|
|
199
|
+
Profiles are versioned: `name` resolves to the latest version, `name@N` pins
|
|
200
|
+
version `N`. A version fixes the stage chain and its parameters; a change to
|
|
201
|
+
either ships as a new version. The codec implementation is not part of the
|
|
202
|
+
version: the ffmpeg or libopus build and the decoder that ran are recorded in
|
|
203
|
+
the run log, and output is reproducible for a given seed, version and codec
|
|
204
|
+
build on one machine and torch thread count. `phonesim info` lists the versions.
|
|
205
|
+
|
|
206
|
+
| Profile | Models | Internal rate | Path |
|
|
207
|
+
|---|---|---|---|
|
|
208
|
+
| `pstn_narrowband` | Analogue loop into a G.711 exchange | 8 kHz | send-side level + ambient noise, ~300 Hz low edge, roll-off from 3.4 kHz to 4 kHz, G.711, clock drift |
|
|
209
|
+
| `pstn_g726` | PSTN over G.726 ADPCM | 8 kHz | as above with G.726 (32 kbit/s) |
|
|
210
|
+
| `voip_opus_wideband` | WebRTC / VoIP | 16 kHz | send-side level + ambient noise, ~50–80 Hz low edge, roll-off from ~7 kHz, Opus (libopus) with erasures up to 5 % (FEC + PLC), adaptive playout, clock drift |
|
|
211
|
+
| `voip_g722_wideband` | SIP HD voice over G.722 | 16 kHz | send-side level + ambient noise, ~50–80 Hz low edge, roll-off from ~7 kHz, G.722 with erasures up to 1 % (G.711 App. I PLC), adaptive playout, clock drift |
|
|
212
|
+
| `voip_to_cellular_wideband` | VoIP into mobile **HD** voice | 16 kHz | send-side level + ambient noise, Opus hop with erasures up to 1 % → wideband edges → AMR-WB with radio erasures up to 1 %, optional AMR-NB second transcode, adaptive playout, clock drift |
|
|
213
|
+
| `voip_to_cellular_narrowband` | VoIP into a **regular (non-HD)** mobile call | 8 kHz | send-side level + ambient noise, Opus hop with erasures up to 1 % → 8 kHz, ~100 Hz low edge, roll-off from 3.4 kHz to 4 kHz → AMR-NB with radio erasures up to 1 %, adaptive playout, clock drift; `bitrate` selects the AMR-NB mode |
|
|
214
|
+
| `stress_multi_transcode` | Synthetic stress chain (not a real route) | 16/8/16 kHz | brick-wall band-pass, Opus → packet loss → AMR-NB → G.711 → packet loss → AMR-WB, jitter buffer, AGC, noise, clipping, speed drift, time offset |
|
|
215
|
+
|
|
216
|
+
The six call-path profiles share one structure. On the send side the active
|
|
217
|
+
speech level is set and ambient noise added before the first encoder, as a
|
|
218
|
+
handset or platform does. Band edges are the ones the codec and handset
|
|
219
|
+
define: a 2nd-order low edge (250–320 Hz on the analogue loop, ~100 Hz on
|
|
220
|
+
the digital narrowband path, 50–80 Hz on the wideband ones) and a roll-off
|
|
221
|
+
from the codec's passband to the channel Nyquist. The four packetised
|
|
222
|
+
profiles erase frames and conceal them on the receiving side. AMR and Opus
|
|
223
|
+
frames are removed from the coded stream before the decoder: AMR runs its
|
|
224
|
+
decoder's error concealment, which carries the error into the following
|
|
225
|
+
frames; libopus decodes the frame from the next packet's in-band FEC when it
|
|
226
|
+
carries one and otherwise runs its PLC. G.722 is decoded in full and the
|
|
227
|
+
erased frames are replaced in the decoded PCM by the ITU-T G.711 Appendix I
|
|
228
|
+
waveform substitution (`phonesim.plc`), scaled to 16 kHz. Network loss and
|
|
229
|
+
late arrivals are one erasure process per hop, in bursts: up to 1 % of frames
|
|
230
|
+
on a managed trunk or a radio leg (the LTE conversational-voice loss target),
|
|
231
|
+
up to 5 % on the public Internet (`voip_opus_wideband`). The playout buffer
|
|
232
|
+
adapts by expanding or dropping single frames; clock drift is within ±50 ppm.
|
|
233
|
+
`stress_multi_transcode` instead uses brick-wall band-passes, packet loss and
|
|
234
|
+
a jitter buffer in the decoded signal, AGC, added noise, clipping, speed drift
|
|
235
|
+
and a start offset.
|
|
236
|
+
|
|
237
|
+
Use `phonesim.list_profiles()` to enumerate profiles and
|
|
238
|
+
`PhoneCallSimulator(profile=...).describe()` to print the exact stage chain with
|
|
239
|
+
the backend and mode of each codec.
|
|
240
|
+
|
|
241
|
+
**Wideband vs narrowband cellular.** `voip_to_cellular_wideband` models an **HD
|
|
242
|
+
voice** call (AMR-WB, ~7 kHz), which the model assumes is negotiated along the
|
|
243
|
+
whole path. `voip_to_cellular_narrowband` models a call delivered **narrowband**
|
|
244
|
+
(AMR-NB, ~3.4 kHz), the assumption for a call that does not negotiate HD end to
|
|
245
|
+
end. Its `bitrate` argument selects the AMR-NB mode (`12.2k`, the default, down
|
|
246
|
+
to `4.75k`); lower modes model poorer radio conditions, e.g.
|
|
247
|
+
`profile_params={"bitrate": "7.4k"}`.
|
|
248
|
+
|
|
249
|
+
---
|
|
250
|
+
|
|
251
|
+
## The signal path, stage by stage
|
|
252
|
+
|
|
253
|
+
The package is built from small `Stage` modules (each an `nn.Module`) composed
|
|
254
|
+
into a `Pipeline`. The stages, grouped by what they model:
|
|
255
|
+
|
|
256
|
+
- **Rate & bandwidth**: `ResampleStage` (polyphase resampling
|
|
257
|
+
between rates), `ChannelEdgeStage` (2nd-order low edge plus a roll-off to the
|
|
258
|
+
channel Nyquist; the edges of a digital channel), `BandlimitStage` (brick-wall
|
|
259
|
+
FIR band-pass; `stress_multi_transcode` and custom chains).
|
|
260
|
+
- **Send side**: `AmbientNoiseStage` (room noise at an SNR relative to the active
|
|
261
|
+
speech level, entering before the encoder), `SpeechLevelStage` (P.56-style
|
|
262
|
+
active speech level), `LimiterStage` (peak limiter, identity below the knee).
|
|
263
|
+
- **Codecs**: `CodecStage`, three backends: `ffmpeg` (G.711/G.722/G.726/Opus,
|
|
264
|
+
plus AMR-NB/AMR-WB when built in; AMR and Opus through their reference
|
|
265
|
+
decoders, delay-compensated), `native` (G.711 via `CompandingStage`, the
|
|
266
|
+
exact segmented coder; the default for G.711, used by `pstn_narrowband` and
|
|
267
|
+
the G.711 hop of `stress_multi_transcode`, while `pstn_g726` runs G.726
|
|
268
|
+
through `ffmpeg`) and `libopus` (Opus in-process). `erasure_rate` erases
|
|
269
|
+
20 ms frames in bursts. AMR and Opus frames are removed from the coded
|
|
270
|
+
stream before the decoder (AMR: the slot becomes a NO_DATA frame and the
|
|
271
|
+
decoder runs its error concealment, which carries the error into the
|
|
272
|
+
following frames; Opus: libopus decodes from the next packet's in-band FEC
|
|
273
|
+
when it carries one, otherwise runs its PLC). G.711, G.722 and G.726 are
|
|
274
|
+
decoded in full and the erased frames are replaced in the decoded PCM by
|
|
275
|
+
the ITU-T G.711 Appendix I waveform substitution (`phonesim.plc`).
|
|
276
|
+
- **Packetization / transport**: `PlayoutBufferStage` (per-frame under/over-run
|
|
277
|
+
with overlap-add; bursty late arrivals concealed), `PacketLossStage` (bursty
|
|
278
|
+
loss with repeat-and-fade in the decoded signal), `JitterBufferStage` (late
|
|
279
|
+
frames and under-runs in the decoded signal); the packetised profiles erase
|
|
280
|
+
frames in `CodecStage` and use `PlayoutBufferStage`, `stress_multi_transcode`
|
|
281
|
+
and custom chains use the other two.
|
|
282
|
+
- **Level & nonlinearity**: `AGCStage`, `GainStage`, `ClipStage`.
|
|
283
|
+
- **Noise**: `NoiseStage` (white / pink / mains-hum at a target SNR).
|
|
284
|
+
- **Timing**: `ClockDriftStage` (clock mismatch in ppm, windowed-sinc
|
|
285
|
+
fractional resampling), `SpeedDriftStage` (drift as a linear-interpolation
|
|
286
|
+
resample), `TimeOffsetStage` (recording start offset).
|
|
287
|
+
|
|
288
|
+
A final resample returns the signal to 24 kHz regardless of the internal path.
|
|
289
|
+
|
|
290
|
+
---
|
|
291
|
+
|
|
292
|
+
## Signal analysis
|
|
293
|
+
|
|
294
|
+
`phonesim.analyze_channel(clean, degraded, sample_rate=24000, metrics=None)`
|
|
295
|
+
returns a dictionary of metrics: alignment-corrected SNR, per-band
|
|
296
|
+
energy ratios, high-frequency energy above 4 kHz / 8 kHz, PESQ and STOI (if those
|
|
297
|
+
packages are installed), and, if you pass `metrics={name: fn}` with
|
|
298
|
+
`fn(audio, sr) -> float`, each of your own task metrics on clean vs degraded
|
|
299
|
+
audio.
|
|
300
|
+
|
|
301
|
+
`phonesim.plot_channel(clean, degraded, path="channel_analysis.png")` renders a
|
|
302
|
+
side-by-side comparison (waveforms, spectrograms, log-mel, frequency response and
|
|
303
|
+
band energy) to a PNG.
|
|
304
|
+
|
|
305
|
+
```python
|
|
306
|
+
from phonesim import analyze_channel
|
|
307
|
+
report = analyze_channel(clean, degraded, sample_rate=24000,
|
|
308
|
+
metrics={"my_score": my_metric})
|
|
309
|
+
print(report["snr_db"], report["my_score_clean"], report["my_score_degraded"])
|
|
310
|
+
```
|
|
311
|
+
|
|
312
|
+
---
|
|
313
|
+
|
|
314
|
+
## Command-line interface
|
|
315
|
+
|
|
316
|
+
```bash
|
|
317
|
+
python -m phonesim.cli info # list profiles & available codecs
|
|
318
|
+
python -m phonesim.cli run --in a.wav --out b.wav --profile pstn_narrowband --seed 7
|
|
319
|
+
python -m phonesim.cli analyze --clean a.wav --degraded b.wav --plot out.png --json metrics.json
|
|
320
|
+
python -m phonesim.cli batch --in-dir clips/ --out-dir degraded/ --profile voip_opus_wideband
|
|
321
|
+
```
|
|
322
|
+
|
|
323
|
+
`run` and `batch` build the pipeline from `--profile` (default
|
|
324
|
+
`voip_to_cellular_narrowband`) or from `--config config.yaml`, which replaces
|
|
325
|
+
`--profile`: the file names a profile or lists the stages itself. With
|
|
326
|
+
`--config` the file's `input_sr` / `output_sr` (default 24000) set the rates
|
|
327
|
+
and `--sr` / `--out-sr` may only repeat them; a different value exits with
|
|
328
|
+
status 2. Without `--config`, `--sr` / `--out-sr` set the rates (default
|
|
329
|
+
24000). `--deterministic` uses the nominal (non-random) parameters.
|
|
330
|
+
`analyze --plot` needs `pip install "phonesim[plot]"`.
|
|
331
|
+
|
|
332
|
+
---
|
|
333
|
+
|
|
334
|
+
## Configuration files
|
|
335
|
+
|
|
336
|
+
Pipelines can be described in YAML/JSON instead of code, either by naming a
|
|
337
|
+
profile or by listing stages explicitly:
|
|
338
|
+
|
|
339
|
+
```yaml
|
|
340
|
+
input_sr: 24000
|
|
341
|
+
output_sr: 24000
|
|
342
|
+
stages:
|
|
343
|
+
- {type: ResampleStage, from_sr: 24000, to_sr: 16000}
|
|
344
|
+
- {type: BandlimitStage, low_hz: 50, high_hz: 7000}
|
|
345
|
+
- {type: CodecStage, codec: amr_wb, bitrate: 12.65k}
|
|
346
|
+
- {type: PacketLossStage, loss_rate: 0.02, burst_probability: 0.2}
|
|
347
|
+
- {type: NoiseStage, snr_db: [25, 40]}
|
|
348
|
+
- {type: ResampleStage, from_sr: 16000, to_sr: 24000}
|
|
349
|
+
```
|
|
350
|
+
|
|
351
|
+
```python
|
|
352
|
+
sim = PhoneCallSimulator.from_config("config.yaml")
|
|
353
|
+
```
|
|
354
|
+
|
|
355
|
+
`input_sr` and `output_sr` are optional and default to 24000.
|
|
356
|
+
|
|
357
|
+
---
|
|
358
|
+
|
|
359
|
+
## Reproducibility
|
|
360
|
+
|
|
361
|
+
Every call takes an optional `seed`. The simulator seeds a dedicated
|
|
362
|
+
`torch.Generator`, so the same input, seed, profile version and codec build
|
|
363
|
+
(ffmpeg, libopus) yield identical output on one machine and torch thread count.
|
|
364
|
+
Across machines the G.711 path reproduces to better than 60 dB SNR (the test
|
|
365
|
+
suite pins its output); the adaptive coders (G.726, G.722, Opus, AMR) turn
|
|
366
|
+
last-bit differences in their input into different bitstreams. The run log
|
|
367
|
+
has one line per stage with the parameters drawn and, for each codec, the
|
|
368
|
+
build, decoder and mode:
|
|
369
|
+
|
|
370
|
+
```python
|
|
371
|
+
y, log = sim(x, seed=1234, return_log=True)
|
|
372
|
+
```
|
|
373
|
+
|
|
374
|
+
With `per_example=True`, row 0 of a batch uses `seed` and the other rows use
|
|
375
|
+
seeds drawn from a generator seeded with it (`phonesim.simulator.row_seeds`).
|
|
376
|
+
|
|
377
|
+
---
|
|
378
|
+
|
|
379
|
+
## Shapes at a glance
|
|
380
|
+
|
|
381
|
+
- **Internal tensor convention**: `[B, C, T]`. The public API accepts `[T]`,
|
|
382
|
+
`[B, T]`, `[B, C, T]`, NumPy or torch, and restores the original rank/type.
|
|
383
|
+
- **Output rate**: always `output_sample_rate` (24 kHz by default).
|
|
384
|
+
|
|
385
|
+
---
|
|
386
|
+
|
|
387
|
+
## Project layout
|
|
388
|
+
|
|
389
|
+
```
|
|
390
|
+
phonesim/
|
|
391
|
+
phonesim/
|
|
392
|
+
core.py # Stage / Pipeline / SimContext, shape & dtype handling
|
|
393
|
+
dsp.py # resampling, FIR filters, level helpers
|
|
394
|
+
ffmpeg_backend.py # codec round-trips through ffmpeg (G.711/G.722/G.726/Opus + AMR-NB/WB if built in)
|
|
395
|
+
opus_backend.py # Opus through the libopus shared library (PLC, in-band FEC)
|
|
396
|
+
plc.py # ITU-T G.711 Appendix I packet loss concealment
|
|
397
|
+
stages/ # all signal-path stages
|
|
398
|
+
profiles/ # named profiles + registry
|
|
399
|
+
simulator.py # PhoneCallSimulator / PhoneCallPipeline
|
|
400
|
+
analysis.py # metrics + plotting
|
|
401
|
+
config.py # YAML/JSON pipeline loading
|
|
402
|
+
io_utils.py # load_audio / save_audio
|
|
403
|
+
cli.py # command-line interface
|
|
404
|
+
tests/ # pytest suite
|
|
405
|
+
examples/ # runnable scripts
|
|
406
|
+
pyproject.toml # packaging / install / console script
|
|
407
|
+
README.md # this documentation
|
|
408
|
+
```
|
|
409
|
+
|
|
410
|
+
## Effective bandwidth per profile
|
|
411
|
+
|
|
412
|
+
Share of the *output* energy per band, and the −3 dB band relative to 1 kHz,
|
|
413
|
+
for a 5 s white-noise probe through each profile (24 kHz in/out, deterministic
|
|
414
|
+
parameters, 4096-point Welch spectra):
|
|
415
|
+
|
|
416
|
+
| Profile | 0–4 kHz | 4–8 kHz | 8–12 kHz | −3 dB band |
|
|
417
|
+
|---|---|---|---|---|
|
|
418
|
+
| `pstn_narrowband` | 1.00 | ~0 | ~0 | ~300 Hz – 3.4 kHz |
|
|
419
|
+
| `pstn_g726` | 1.00 | ~0 | ~0 | ~300 Hz – 3.4 kHz |
|
|
420
|
+
| `voip_to_cellular_narrowband` | 1.00 | ~0 | ~0 | ~120 Hz – 3.3 kHz, roll-off to 4 kHz |
|
|
421
|
+
| `stress_multi_transcode` | 1.00 | ~0 | ~0 | ~350 Hz – 3 kHz |
|
|
422
|
+
| `voip_to_cellular_wideband` | 0.68 | 0.32 | ~0 | ~80 Hz – 6 kHz¹ |
|
|
423
|
+
| `voip_opus_wideband` | 0.64 | 0.36 | ~0 | ~80 Hz – 6.9 kHz |
|
|
424
|
+
| `voip_g722_wideband` | 0.59 | 0.41 | ~0 | ~70 Hz – 7 kHz |
|
|
425
|
+
|
|
426
|
+
¹ AMR-WB at 12.65 kbit/s; `profile_params={"second_transcode": True}` adds an
|
|
427
|
+
AMR-NB interconnect hop that collapses the call to the narrowband path.
|
|
428
|
+
|
|
429
|
+
If a downstream system depends on fine spectral detail above 3.4 kHz, only the
|
|
430
|
+
wideband profiles keep it; the mobile narrowband path, the common case for a
|
|
431
|
+
call to an ordinary number, keeps the telephone band only.
|
|
432
|
+
|
|
433
|
+
## Assumptions and limitations
|
|
434
|
+
|
|
435
|
+
- **No EVS, no G.729.** EVS has no open encoder and ffmpeg has no G.729 encoder
|
|
436
|
+
(an open one exists outside ffmpeg), so neither is offered; VoLTE with EVS is
|
|
437
|
+
not modelled.
|
|
438
|
+
- **G.722 and G.726 conceal with the G.711 Appendix I algorithm**, scaled to the
|
|
439
|
+
codec's rate; G.722's own Appendix III/IV PLC is not implemented. AMR and Opus
|
|
440
|
+
conceal with their own decoders.
|
|
441
|
+
- **G.711, G.722 and G.726 decoder state is never disturbed by an erasure.**
|
|
442
|
+
These codecs are decoded in full and the erased frames are replaced in the
|
|
443
|
+
decoded PCM, so the post-erasure divergence a real ADPCM receiver (G.722,
|
|
444
|
+
G.726) shows is not modelled; frames after an erasure equal a loss-free
|
|
445
|
+
decode.
|
|
446
|
+
- **`native` and `ffmpeg` G.711 are not bit-identical.** Same decoder; ffmpeg's
|
|
447
|
+
encoder table rounds differently at the decision levels, so 512 µ-law and 964
|
|
448
|
+
A-law of the 65536 int16 inputs take the adjacent code, one level apart.
|
|
449
|
+
- **Transport effects are statistical, not protocol-accurate.** Erasures follow
|
|
450
|
+
a two-state (Gilbert-Elliott) chain per codec hop; jitter-buffer adaptation
|
|
451
|
+
and level control are parametric models, not RTP/WebRTC implementations.
|
|
452
|
+
- **No acoustic path is modeled.** Room reverberation, speaker/handset
|
|
453
|
+
acoustics, acoustic echo, and the receiver's re-recording/re-digitization step
|
|
454
|
+
are out of scope; add them upstream if needed.
|
|
455
|
+
|
|
456
|
+
## Checking a profile against your own path
|
|
457
|
+
|
|
458
|
+
Each profile is a physically motivated model of its path; no comparison against
|
|
459
|
+
recorded calls ships with this release. To check a profile against your own path, record the same clips through a real
|
|
460
|
+
call, run `analyze_channel(clean, received)` on both the real and the simulated
|
|
461
|
+
output, and compare the distributions. Prefer changing a parameter for a physical
|
|
462
|
+
reason over tuning it to a handful of recordings.
|
|
463
|
+
|
|
464
|
+
## Further reading
|
|
465
|
+
|
|
466
|
+
- `examples/degrade.py` — degrade a WAV (or a synthetic tone) through a
|
|
467
|
+
profile and print channel metrics.
|
|
468
|
+
- `tests/test_phonesim.py` — executable specification: output rate, shape/type
|
|
469
|
+
preservation, determinism, length preservation, band-limiting behaviour, the
|
|
470
|
+
physics of each stage and batch processing are asserted there and double as
|
|
471
|
+
usage examples.
|
|
472
|
+
- Each stage and the high-level classes carry docstrings; e.g.
|
|
473
|
+
`help(phonesim.PhoneCallSimulator)`.
|