phonesim 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (55) hide show
  1. phonesim-0.1.0/CHANGELOG.md +42 -0
  2. phonesim-0.1.0/CONTRIBUTING.md +12 -0
  3. phonesim-0.1.0/LICENSE +21 -0
  4. phonesim-0.1.0/MANIFEST.in +3 -0
  5. phonesim-0.1.0/PKG-INFO +473 -0
  6. phonesim-0.1.0/README.md +435 -0
  7. phonesim-0.1.0/SECURITY.md +6 -0
  8. phonesim-0.1.0/examples/degrade.py +58 -0
  9. phonesim-0.1.0/phonesim/__init__.py +83 -0
  10. phonesim-0.1.0/phonesim/analysis.py +331 -0
  11. phonesim-0.1.0/phonesim/cli.py +225 -0
  12. phonesim-0.1.0/phonesim/config.py +130 -0
  13. phonesim-0.1.0/phonesim/core.py +298 -0
  14. phonesim-0.1.0/phonesim/dsp.py +228 -0
  15. phonesim-0.1.0/phonesim/ffmpeg_backend.py +304 -0
  16. phonesim-0.1.0/phonesim/io_utils.py +70 -0
  17. phonesim-0.1.0/phonesim/opus_backend.py +248 -0
  18. phonesim-0.1.0/phonesim/plc.py +206 -0
  19. phonesim-0.1.0/phonesim/profiles/__init__.py +289 -0
  20. phonesim-0.1.0/phonesim/simulator.py +204 -0
  21. phonesim-0.1.0/phonesim/stages/__init__.py +40 -0
  22. phonesim-0.1.0/phonesim/stages/ambient.py +45 -0
  23. phonesim-0.1.0/phonesim/stages/channel_edge.py +109 -0
  24. phonesim-0.1.0/phonesim/stages/codec.py +212 -0
  25. phonesim-0.1.0/phonesim/stages/companding.py +93 -0
  26. phonesim-0.1.0/phonesim/stages/filtering.py +60 -0
  27. phonesim-0.1.0/phonesim/stages/gain.py +114 -0
  28. phonesim-0.1.0/phonesim/stages/level.py +77 -0
  29. phonesim-0.1.0/phonesim/stages/misc.py +75 -0
  30. phonesim-0.1.0/phonesim/stages/noise.py +60 -0
  31. phonesim-0.1.0/phonesim/stages/packet.py +180 -0
  32. phonesim-0.1.0/phonesim/stages/resample.py +43 -0
  33. phonesim-0.1.0/phonesim/stages/timing.py +140 -0
  34. phonesim-0.1.0/phonesim.egg-info/PKG-INFO +473 -0
  35. phonesim-0.1.0/phonesim.egg-info/SOURCES.txt +53 -0
  36. phonesim-0.1.0/phonesim.egg-info/dependency_links.txt +1 -0
  37. phonesim-0.1.0/phonesim.egg-info/entry_points.txt +2 -0
  38. phonesim-0.1.0/phonesim.egg-info/requires.txt +26 -0
  39. phonesim-0.1.0/phonesim.egg-info/top_level.txt +1 -0
  40. phonesim-0.1.0/pyproject.toml +50 -0
  41. phonesim-0.1.0/setup.cfg +4 -0
  42. phonesim-0.1.0/tests/__init__.py +0 -0
  43. phonesim-0.1.0/tests/data/profile_fingerprints.json +879 -0
  44. phonesim-0.1.0/tests/data/pstn_narrowband@1.npy +0 -0
  45. phonesim-0.1.0/tests/data/run_logs.json +85 -0
  46. phonesim-0.1.0/tests/regen_goldens.py +80 -0
  47. phonesim-0.1.0/tests/test_cli_rates.py +231 -0
  48. phonesim-0.1.0/tests/test_config.py +107 -0
  49. phonesim-0.1.0/tests/test_ffmpeg_erasure.py +129 -0
  50. phonesim-0.1.0/tests/test_g711.py +54 -0
  51. phonesim-0.1.0/tests/test_opus_backend.py +199 -0
  52. phonesim-0.1.0/tests/test_phonesim.py +989 -0
  53. phonesim-0.1.0/tests/test_plc.py +130 -0
  54. phonesim-0.1.0/tests/test_timing.py +88 -0
  55. phonesim-0.1.0/tests/test_versions.py +166 -0
@@ -0,0 +1,42 @@
1
+ # Changelog
2
+
3
+ ## 0.1.0
4
+
5
+ - Seven profiles, each a seeded model of one call path: `pstn_narrowband`
6
+ (analogue loop into a G.711 exchange), `pstn_g726` (PSTN over G.726 ADPCM),
7
+ `voip_opus_wideband` (WebRTC/VoIP over Opus), `voip_g722_wideband` (SIP HD
8
+ voice over G.722), `voip_to_cellular_wideband` (VoIP into an AMR-WB HD-voice
9
+ call), `voip_to_cellular_narrowband` (VoIP into an ordinary AMR-NB mobile
10
+ call; the default) and `stress_multi_transcode` (a synthetic stress chain,
11
+ not a real route).
12
+ - Real codecs only: G.711 with the ITU-T segmented coder in-process; G.722,
13
+ G.726, AMR-NB and AMR-WB through ffmpeg; Opus through ffmpeg or the libopus
14
+ shared library. AMR and Opus decode with `libopencore_amrnb`,
15
+ `libopencore_amrwb` and `libopus`; algorithmic delay is compensated. A
16
+ profile whose codec this machine cannot run raises `CodecUnavailableError`
17
+ when built.
18
+ - Frame erasures in `CodecStage`, concealed on the receiving side. AMR and
19
+ Opus frames are removed from the coded stream before the decoder: AMR runs
20
+ its decoder's error concealment, which carries the error into the following
21
+ frames; libopus decodes from the next packet's in-band FEC when it carries
22
+ one, otherwise its PLC. G.711, G.722 and G.726 are decoded in full and the
23
+ erased frames are replaced in the decoded PCM by the ITU-T G.711 Appendix I
24
+ waveform substitution (`phonesim.plc`).
25
+ - Send-side speech level and ambient noise, codec- and handset-defined band
26
+ edges, adaptive playout, clock drift in ppm, and the stress chain's
27
+ packet-loss, jitter-buffer, AGC, noise, clipping, speed-drift and
28
+ time-offset stages.
29
+ - Profile versions: `name@N`; a bare name is the latest. A version fixes the
30
+ stage chain and its parameters; the codec build and decoder are recorded in
31
+ the run log, not versioned. `phonesim info` lists versions.
32
+ - Reproducibility: a seed drives one `torch.Generator`; the same input, seed,
33
+ profile version and codec build give the same output on one machine and
34
+ torch thread count. `per_example=True` gives each row of a batch its own
35
+ call, seeded by `row_seeds`. The run log has one line per stage.
36
+ - `analyze_channel` (aligned SNR, band energies, high-frequency energy,
37
+ optional PESQ and STOI, caller-supplied `metrics=`) and `plot_channel`.
38
+ - `phonesim` CLI (`info`, `run`, `analyze`, `batch`) and YAML/JSON pipeline
39
+ configs.
40
+ - Tests pin every profile version's stage chain, parameters and run log; CI
41
+ runs them on a distro ffmpeg (AMR tests skip) and on a static AMR-capable
42
+ build pinned by sha256.
@@ -0,0 +1,12 @@
1
+ # Contributing
2
+
3
+ - Open a pull request against `main`; every change is reviewed.
4
+ - Run `pytest -q` before pushing, with an AMR-capable ffmpeg on `PATH` and libopus
5
+ installed if you can (see README); without them the AMR and Opus-erasure tests
6
+ skip.
7
+ - A change to a profile's stage chain or parameters is a new profile version,
8
+ never an edit in place; `pytest` enforces it against
9
+ `tests/data/profile_fingerprints.json`, and `python tests/regen_goldens.py`
10
+ adds the entries for a new version. The codec implementation (ffmpeg build,
11
+ decoder) is logged, not versioned.
12
+ - Claims about realism need a measurement in the PR, not an adjective.
phonesim-0.1.0/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 DeepMark Inc.
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,3 @@
1
+ include LICENSE CHANGELOG.md CONTRIBUTING.md SECURITY.md
2
+ recursive-include examples *.py
3
+ recursive-include tests *.py *.json *.npy
@@ -0,0 +1,473 @@
1
+ Metadata-Version: 2.4
2
+ Name: phonesim
3
+ Version: 0.1.0
4
+ Summary: Provider-free phone-call audio-degradation simulator with real telephony codecs
5
+ Author: DeepMark Inc.
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/deepmark/phonesim
8
+ Project-URL: Issues, https://github.com/deepmark/phonesim/issues
9
+ Keywords: audio,telephony,codec,simulation,dsp
10
+ Classifier: Programming Language :: Python :: 3
11
+ Classifier: Topic :: Multimedia :: Sound/Audio
12
+ Classifier: Topic :: Scientific/Engineering
13
+ Requires-Python: >=3.10
14
+ Description-Content-Type: text/markdown
15
+ License-File: LICENSE
16
+ Requires-Dist: numpy>=1.21
17
+ Requires-Dist: torch>=2.0
18
+ Requires-Dist: soundfile>=0.11
19
+ Provides-Extra: yaml
20
+ Requires-Dist: pyyaml>=5.4; extra == "yaml"
21
+ Provides-Extra: plot
22
+ Requires-Dist: matplotlib>=3.4; extra == "plot"
23
+ Provides-Extra: metrics
24
+ Requires-Dist: pesq>=0.0.4; extra == "metrics"
25
+ Requires-Dist: pystoi>=0.3; extra == "metrics"
26
+ Provides-Extra: test
27
+ Requires-Dist: pytest>=7.0; extra == "test"
28
+ Requires-Dist: pyyaml>=5.4; extra == "test"
29
+ Requires-Dist: pystoi>=0.3; extra == "test"
30
+ Requires-Dist: matplotlib>=3.4; extra == "test"
31
+ Provides-Extra: full
32
+ Requires-Dist: pyyaml>=5.4; extra == "full"
33
+ Requires-Dist: matplotlib>=3.4; extra == "full"
34
+ Requires-Dist: pesq>=0.0.4; extra == "full"
35
+ Requires-Dist: pystoi>=0.3; extra == "full"
36
+ Requires-Dist: pytest>=7.0; extra == "full"
37
+ Dynamic: license-file
38
+
39
+ # phonesim
40
+
41
+ A provider-free phone-call audio-degradation simulator: a local, seeded model
42
+ of the signal path of a telephone call, running the real telephony codecs.
43
+
44
+ `phonesim` recreates, locally and with no third-party telephony services (no
45
+ Twilio / Telnyx / Vonage), the chain of distortions a signal accumulates when it
46
+ travels through a real phone call: VoIP/WebRTC transport, PSTN and cellular
47
+ interconnects, mobile voice codecs, packet loss, jitter, automatic gain control,
48
+ level control, background noise, and clock drift. **The codecs are real**:
49
+ G.722, G.726 and (with an AMR-capable ffmpeg build) AMR-NB and AMR-WB run
50
+ through ffmpeg, Opus through ffmpeg or the libopus library; G.711 companding is
51
+ the ITU-T segmented coder, in-process.
52
+ No codec is approximated: a profile whose codec this machine cannot run fails
53
+ when built, with the install hint.
54
+
55
+ Input and output default to **24 kHz**; both rates are parameters. Batches
56
+ (`[B, T]`) can be processed as one call or, with `per_example=True`, as one
57
+ independent call per row, which is the mode for data generation.
58
+
59
+ ---
60
+
61
+ ## Why this exists
62
+
63
+ A phone call band-limits the signal to a few kHz, compresses it with a lossy
64
+ speech codec (often more than once when the call crosses network boundaries),
65
+ chops it into packets that can be lost or delayed, re-levels it, and re-digitises
66
+ it at the far end. Any audio system that has to work over calls (speech
67
+ recognition, speaker verification, watermark detection, enhancement) needs to be
68
+ measured and trained against that channel.
69
+
70
+ Doing so with real calls means a telephony provider, cost and no
71
+ reproducibility. `phonesim` is a local, seeded model of the same signal path,
72
+ so you can measure over representative paths, generate degraded data at scale,
73
+ and reproduce a result from a seed.
74
+
75
+ ---
76
+
77
+ ## What it does and does not claim to do
78
+
79
+ **It models** the *signal-level* transformations of a call path: resampling,
80
+ the channel's band edges, real codec encode/decode, frame erasures with
81
+ receiver-side concealment, playout-buffer behaviour, send-side level control, ambient
82
+ noise, multi-transcode chains, and clock drift.
83
+
84
+ **Every codec is real.** A stock ffmpeg covers G.711, G.722, G.726 and Opus; an
85
+ **AMR-capable ffmpeg build** (libopencore-amr + libvo-amrwbenc) adds the real
86
+ **AMR-NB and AMR-WB** cellular codecs, the ones that carry mobile voice. AMR
87
+ and Opus are decoded with `libopencore_amrnb` / `libopencore_amrwb` / `libopus`
88
+ rather than ffmpeg's own decoders, and the run log records the ffmpeg version,
89
+ decoder and mode that ran. There is no approximation to fall back to: building
90
+ a profile whose codec is missing raises `CodecUnavailableError`. EVS has no
91
+ open encoder and ffmpeg has no G.729 encoder, so neither is offered.
92
+
93
+ ---
94
+
95
+ ## Installation
96
+
97
+ ```bash
98
+ # from the repo root:
99
+ pip install -e . # core: numpy, torch, soundfile
100
+ pip install -e ".[full]" # + pyyaml, matplotlib, pesq, pystoi, pytest
101
+
102
+ # codec profiles need ffmpeg on the PATH:
103
+ # apt-get install ffmpeg (or: brew install ffmpeg)
104
+ # the default profile also needs the AMR encoders and libopus (both below)
105
+ ```
106
+
107
+ This installs the package (so `import phonesim` works from anywhere) and a
108
+ `phonesim` console command. For GPU use, install the torch build matching your
109
+ CUDA version from <https://pytorch.org>. Quick check:
110
+
111
+ ```python
112
+ import phonesim
113
+ print(phonesim.list_profiles())
114
+ ```
115
+
116
+ ### Real codecs via ffmpeg
117
+
118
+ `phonesim` probes which codecs the ffmpeg on `PATH` (or the one named by
119
+ `PHONESIM_FFMPEG`) can round-trip:
120
+
121
+ ```python
122
+ from phonesim import ffmpeg_backend
123
+ print(sorted(ffmpeg_backend.available_codecs()))
124
+ # stock ffmpeg → ['g711_alaw', 'g711_ulaw', 'g722', 'g726', 'opus']
125
+ # AMR-capable → … plus 'amr_nb' and 'amr_wb'
126
+ ```
127
+
128
+ A **stock distro ffmpeg cannot run AMR**: it has neither the encoders
129
+ (`libopencore-amrnb`, `libvo-amrwbenc`) nor the OpenCORE decoders. The
130
+ no-compile path on Linux is a static GPL build:
131
+
132
+ ```bash
133
+ base=https://johnvansickle.com/ffmpeg/releases/ffmpeg-release-amd64-static.tar.xz
134
+ curl -fL -o ffmpeg-release-amd64-static.tar.xz "$base"
135
+ curl -fL -o ffmpeg-release-amd64-static.tar.xz.md5 "$base.md5"
136
+ md5sum -c ffmpeg-release-amd64-static.tar.xz.md5
137
+ tar xf ffmpeg-release-amd64-static.tar.xz
138
+ install ffmpeg-*-amd64-static/ffmpeg ~/.local/bin/ffmpeg # ~/.local/bin on PATH
139
+
140
+ ffmpeg -hide_banner -encoders | grep -i amr
141
+ # A....D libopencore_amrnb OpenCORE AMR-NB ... (codec amr_nb)
142
+ # A....D libvo_amrwbenc Android VisualOn AMR-WB ... (codec amr_wb)
143
+ ```
144
+
145
+ macOS: Homebrew's `ffmpeg` is built without AMR; build ffmpeg with
146
+ `--enable-version3 --enable-libopencore-amrnb --enable-libopencore-amrwb --enable-libvo-amrwbenc`.
147
+
148
+ Licensing: phonesim is MIT and links nothing at build time; it calls the `ffmpeg`
149
+ binary at runtime. The AMR libraries (opencore-amr, vo-amrwbenc) are
150
+ Apache-2.0; ffmpeg accepts them only with `--enable-version3`, which makes the
151
+ ffmpeg build LGPLv3, or GPLv3 when `--enable-gpl` is also set (the static build
152
+ above sets both, so it is GPLv3). Distributing such a binary together with your
153
+ software carries that licence's obligations; the AMR codecs are also subject to
154
+ patents in some jurisdictions.
155
+
156
+ Without AMR, the `voip_to_cellular_*` and `stress_multi_transcode` profiles
157
+ raise `CodecUnavailableError` when built; the PSTN and G.722 profiles work with
158
+ a stock ffmpeg, `voip_opus_wideband` also needs libopus (next section).
159
+
160
+ ### Opus erasures need libopus
161
+
162
+ Erasing Opus frames runs Opus in-process through the libopus shared library
163
+ (`apt-get install libopus0`, `brew install opus`; system and Homebrew paths are
164
+ searched, `PHONESIM_LIBOPUS` names any other file), so the decoder's own
165
+ concealment and in-band FEC apply. Without it, `voip_opus_wideband` and the
166
+ two `voip_to_cellular_*` profiles raise `CodecUnavailableError`.
167
+
168
+ > **Device support:** CPU and NVIDIA CUDA; torch tensors are processed on the
169
+ > device they arrive on. Codecs always run on the CPU (ffmpeg, libopus or the
170
+ > in-process G.711 coder).
171
+
172
+ ---
173
+
174
+ ## Quick start
175
+
176
+ ### Degrade an audio file
177
+
178
+ ```python
179
+ from phonesim import PhoneCallSimulator, load_audio, save_audio
180
+
181
+ x, sr = load_audio("input.wav", sr=24000) # loads & resamples to 24 kHz
182
+ sim = PhoneCallSimulator(profile="voip_to_cellular_narrowband")
183
+ y = sim(x, seed=1234) # reproducible degraded audio
184
+ save_audio("degraded.wav", y, sr=24000)
185
+ ```
186
+
187
+ `sim(x)` accepts and returns a NumPy array **or** a torch tensor, of rank
188
+ `[T]`, `[B, T]`, or `[B, C, T]`, and gives back the same type and rank at 24 kHz.
189
+ `sim(x, seed=1234, per_example=True)` gives each row of a batch its own call.
190
+
191
+ ---
192
+
193
+
194
+ ## The built-in profiles
195
+
196
+ A *profile* is a named combination of stages modelling one call path. All
197
+ start and end at 24 kHz by default.
198
+
199
+ Profiles are versioned: `name` resolves to the latest version, `name@N` pins
200
+ version `N`. A version fixes the stage chain and its parameters; a change to
201
+ either ships as a new version. The codec implementation is not part of the
202
+ version: the ffmpeg or libopus build and the decoder that ran are recorded in
203
+ the run log, and output is reproducible for a given seed, version and codec
204
+ build on one machine and torch thread count. `phonesim info` lists the versions.
205
+
206
+ | Profile | Models | Internal rate | Path |
207
+ |---|---|---|---|
208
+ | `pstn_narrowband` | Analogue loop into a G.711 exchange | 8 kHz | send-side level + ambient noise, ~300 Hz low edge, roll-off from 3.4 kHz to 4 kHz, G.711, clock drift |
209
+ | `pstn_g726` | PSTN over G.726 ADPCM | 8 kHz | as above with G.726 (32 kbit/s) |
210
+ | `voip_opus_wideband` | WebRTC / VoIP | 16 kHz | send-side level + ambient noise, ~50–80 Hz low edge, roll-off from ~7 kHz, Opus (libopus) with erasures up to 5 % (FEC + PLC), adaptive playout, clock drift |
211
+ | `voip_g722_wideband` | SIP HD voice over G.722 | 16 kHz | send-side level + ambient noise, ~50–80 Hz low edge, roll-off from ~7 kHz, G.722 with erasures up to 1 % (G.711 App. I PLC), adaptive playout, clock drift |
212
+ | `voip_to_cellular_wideband` | VoIP into mobile **HD** voice | 16 kHz | send-side level + ambient noise, Opus hop with erasures up to 1 % → wideband edges → AMR-WB with radio erasures up to 1 %, optional AMR-NB second transcode, adaptive playout, clock drift |
213
+ | `voip_to_cellular_narrowband` | VoIP into a **regular (non-HD)** mobile call | 8 kHz | send-side level + ambient noise, Opus hop with erasures up to 1 % → 8 kHz, ~100 Hz low edge, roll-off from 3.4 kHz to 4 kHz → AMR-NB with radio erasures up to 1 %, adaptive playout, clock drift; `bitrate` selects the AMR-NB mode |
214
+ | `stress_multi_transcode` | Synthetic stress chain (not a real route) | 16/8/16 kHz | brick-wall band-pass, Opus → packet loss → AMR-NB → G.711 → packet loss → AMR-WB, jitter buffer, AGC, noise, clipping, speed drift, time offset |
215
+
216
+ The six call-path profiles share one structure. On the send side the active
217
+ speech level is set and ambient noise added before the first encoder, as a
218
+ handset or platform does. Band edges are the ones the codec and handset
219
+ define: a 2nd-order low edge (250–320 Hz on the analogue loop, ~100 Hz on
220
+ the digital narrowband path, 50–80 Hz on the wideband ones) and a roll-off
221
+ from the codec's passband to the channel Nyquist. The four packetised
222
+ profiles erase frames and conceal them on the receiving side. AMR and Opus
223
+ frames are removed from the coded stream before the decoder: AMR runs its
224
+ decoder's error concealment, which carries the error into the following
225
+ frames; libopus decodes the frame from the next packet's in-band FEC when it
226
+ carries one and otherwise runs its PLC. G.722 is decoded in full and the
227
+ erased frames are replaced in the decoded PCM by the ITU-T G.711 Appendix I
228
+ waveform substitution (`phonesim.plc`), scaled to 16 kHz. Network loss and
229
+ late arrivals are one erasure process per hop, in bursts: up to 1 % of frames
230
+ on a managed trunk or a radio leg (the LTE conversational-voice loss target),
231
+ up to 5 % on the public Internet (`voip_opus_wideband`). The playout buffer
232
+ adapts by expanding or dropping single frames; clock drift is within ±50 ppm.
233
+ `stress_multi_transcode` instead uses brick-wall band-passes, packet loss and
234
+ a jitter buffer in the decoded signal, AGC, added noise, clipping, speed drift
235
+ and a start offset.
236
+
237
+ Use `phonesim.list_profiles()` to enumerate profiles and
238
+ `PhoneCallSimulator(profile=...).describe()` to print the exact stage chain with
239
+ the backend and mode of each codec.
240
+
241
+ **Wideband vs narrowband cellular.** `voip_to_cellular_wideband` models an **HD
242
+ voice** call (AMR-WB, ~7 kHz), which the model assumes is negotiated along the
243
+ whole path. `voip_to_cellular_narrowband` models a call delivered **narrowband**
244
+ (AMR-NB, ~3.4 kHz), the assumption for a call that does not negotiate HD end to
245
+ end. Its `bitrate` argument selects the AMR-NB mode (`12.2k`, the default, down
246
+ to `4.75k`); lower modes model poorer radio conditions, e.g.
247
+ `profile_params={"bitrate": "7.4k"}`.
248
+
249
+ ---
250
+
251
+ ## The signal path, stage by stage
252
+
253
+ The package is built from small `Stage` modules (each an `nn.Module`) composed
254
+ into a `Pipeline`. The stages, grouped by what they model:
255
+
256
+ - **Rate & bandwidth**: `ResampleStage` (polyphase resampling
257
+ between rates), `ChannelEdgeStage` (2nd-order low edge plus a roll-off to the
258
+ channel Nyquist; the edges of a digital channel), `BandlimitStage` (brick-wall
259
+ FIR band-pass; `stress_multi_transcode` and custom chains).
260
+ - **Send side**: `AmbientNoiseStage` (room noise at an SNR relative to the active
261
+ speech level, entering before the encoder), `SpeechLevelStage` (P.56-style
262
+ active speech level), `LimiterStage` (peak limiter, identity below the knee).
263
+ - **Codecs**: `CodecStage`, three backends: `ffmpeg` (G.711/G.722/G.726/Opus,
264
+ plus AMR-NB/AMR-WB when built in; AMR and Opus through their reference
265
+ decoders, delay-compensated), `native` (G.711 via `CompandingStage`, the
266
+ exact segmented coder; the default for G.711, used by `pstn_narrowband` and
267
+ the G.711 hop of `stress_multi_transcode`, while `pstn_g726` runs G.726
268
+ through `ffmpeg`) and `libopus` (Opus in-process). `erasure_rate` erases
269
+ 20 ms frames in bursts. AMR and Opus frames are removed from the coded
270
+ stream before the decoder (AMR: the slot becomes a NO_DATA frame and the
271
+ decoder runs its error concealment, which carries the error into the
272
+ following frames; Opus: libopus decodes from the next packet's in-band FEC
273
+ when it carries one, otherwise runs its PLC). G.711, G.722 and G.726 are
274
+ decoded in full and the erased frames are replaced in the decoded PCM by
275
+ the ITU-T G.711 Appendix I waveform substitution (`phonesim.plc`).
276
+ - **Packetization / transport**: `PlayoutBufferStage` (per-frame under/over-run
277
+ with overlap-add; bursty late arrivals concealed), `PacketLossStage` (bursty
278
+ loss with repeat-and-fade in the decoded signal), `JitterBufferStage` (late
279
+ frames and under-runs in the decoded signal); the packetised profiles erase
280
+ frames in `CodecStage` and use `PlayoutBufferStage`, `stress_multi_transcode`
281
+ and custom chains use the other two.
282
+ - **Level & nonlinearity**: `AGCStage`, `GainStage`, `ClipStage`.
283
+ - **Noise**: `NoiseStage` (white / pink / mains-hum at a target SNR).
284
+ - **Timing**: `ClockDriftStage` (clock mismatch in ppm, windowed-sinc
285
+ fractional resampling), `SpeedDriftStage` (drift as a linear-interpolation
286
+ resample), `TimeOffsetStage` (recording start offset).
287
+
288
+ A final resample returns the signal to 24 kHz regardless of the internal path.
289
+
290
+ ---
291
+
292
+ ## Signal analysis
293
+
294
+ `phonesim.analyze_channel(clean, degraded, sample_rate=24000, metrics=None)`
295
+ returns a dictionary of metrics: alignment-corrected SNR, per-band
296
+ energy ratios, high-frequency energy above 4 kHz / 8 kHz, PESQ and STOI (if those
297
+ packages are installed), and, if you pass `metrics={name: fn}` with
298
+ `fn(audio, sr) -> float`, each of your own task metrics on clean vs degraded
299
+ audio.
300
+
301
+ `phonesim.plot_channel(clean, degraded, path="channel_analysis.png")` renders a
302
+ side-by-side comparison (waveforms, spectrograms, log-mel, frequency response and
303
+ band energy) to a PNG.
304
+
305
+ ```python
306
+ from phonesim import analyze_channel
307
+ report = analyze_channel(clean, degraded, sample_rate=24000,
308
+ metrics={"my_score": my_metric})
309
+ print(report["snr_db"], report["my_score_clean"], report["my_score_degraded"])
310
+ ```
311
+
312
+ ---
313
+
314
+ ## Command-line interface
315
+
316
+ ```bash
317
+ python -m phonesim.cli info # list profiles & available codecs
318
+ python -m phonesim.cli run --in a.wav --out b.wav --profile pstn_narrowband --seed 7
319
+ python -m phonesim.cli analyze --clean a.wav --degraded b.wav --plot out.png --json metrics.json
320
+ python -m phonesim.cli batch --in-dir clips/ --out-dir degraded/ --profile voip_opus_wideband
321
+ ```
322
+
323
+ `run` and `batch` build the pipeline from `--profile` (default
324
+ `voip_to_cellular_narrowband`) or from `--config config.yaml`, which replaces
325
+ `--profile`: the file names a profile or lists the stages itself. With
326
+ `--config` the file's `input_sr` / `output_sr` (default 24000) set the rates
327
+ and `--sr` / `--out-sr` may only repeat them; a different value exits with
328
+ status 2. Without `--config`, `--sr` / `--out-sr` set the rates (default
329
+ 24000). `--deterministic` uses the nominal (non-random) parameters.
330
+ `analyze --plot` needs `pip install "phonesim[plot]"`.
331
+
332
+ ---
333
+
334
+ ## Configuration files
335
+
336
+ Pipelines can be described in YAML/JSON instead of code, either by naming a
337
+ profile or by listing stages explicitly:
338
+
339
+ ```yaml
340
+ input_sr: 24000
341
+ output_sr: 24000
342
+ stages:
343
+ - {type: ResampleStage, from_sr: 24000, to_sr: 16000}
344
+ - {type: BandlimitStage, low_hz: 50, high_hz: 7000}
345
+ - {type: CodecStage, codec: amr_wb, bitrate: 12.65k}
346
+ - {type: PacketLossStage, loss_rate: 0.02, burst_probability: 0.2}
347
+ - {type: NoiseStage, snr_db: [25, 40]}
348
+ - {type: ResampleStage, from_sr: 16000, to_sr: 24000}
349
+ ```
350
+
351
+ ```python
352
+ sim = PhoneCallSimulator.from_config("config.yaml")
353
+ ```
354
+
355
+ `input_sr` and `output_sr` are optional and default to 24000.
356
+
357
+ ---
358
+
359
+ ## Reproducibility
360
+
361
+ Every call takes an optional `seed`. The simulator seeds a dedicated
362
+ `torch.Generator`, so the same input, seed, profile version and codec build
363
+ (ffmpeg, libopus) yield identical output on one machine and torch thread count.
364
+ Across machines the G.711 path reproduces to better than 60 dB SNR (the test
365
+ suite pins its output); the adaptive coders (G.726, G.722, Opus, AMR) turn
366
+ last-bit differences in their input into different bitstreams. The run log
367
+ has one line per stage with the parameters drawn and, for each codec, the
368
+ build, decoder and mode:
369
+
370
+ ```python
371
+ y, log = sim(x, seed=1234, return_log=True)
372
+ ```
373
+
374
+ With `per_example=True`, row 0 of a batch uses `seed` and the other rows use
375
+ seeds drawn from a generator seeded with it (`phonesim.simulator.row_seeds`).
376
+
377
+ ---
378
+
379
+ ## Shapes at a glance
380
+
381
+ - **Internal tensor convention**: `[B, C, T]`. The public API accepts `[T]`,
382
+ `[B, T]`, `[B, C, T]`, NumPy or torch, and restores the original rank/type.
383
+ - **Output rate**: always `output_sample_rate` (24 kHz by default).
384
+
385
+ ---
386
+
387
+ ## Project layout
388
+
389
+ ```
390
+ phonesim/
391
+ phonesim/
392
+ core.py # Stage / Pipeline / SimContext, shape & dtype handling
393
+ dsp.py # resampling, FIR filters, level helpers
394
+ ffmpeg_backend.py # codec round-trips through ffmpeg (G.711/G.722/G.726/Opus + AMR-NB/WB if built in)
395
+ opus_backend.py # Opus through the libopus shared library (PLC, in-band FEC)
396
+ plc.py # ITU-T G.711 Appendix I packet loss concealment
397
+ stages/ # all signal-path stages
398
+ profiles/ # named profiles + registry
399
+ simulator.py # PhoneCallSimulator / PhoneCallPipeline
400
+ analysis.py # metrics + plotting
401
+ config.py # YAML/JSON pipeline loading
402
+ io_utils.py # load_audio / save_audio
403
+ cli.py # command-line interface
404
+ tests/ # pytest suite
405
+ examples/ # runnable scripts
406
+ pyproject.toml # packaging / install / console script
407
+ README.md # this documentation
408
+ ```
409
+
410
+ ## Effective bandwidth per profile
411
+
412
+ Share of the *output* energy per band, and the −3 dB band relative to 1 kHz,
413
+ for a 5 s white-noise probe through each profile (24 kHz in/out, deterministic
414
+ parameters, 4096-point Welch spectra):
415
+
416
+ | Profile | 0–4 kHz | 4–8 kHz | 8–12 kHz | −3 dB band |
417
+ |---|---|---|---|---|
418
+ | `pstn_narrowband` | 1.00 | ~0 | ~0 | ~300 Hz – 3.4 kHz |
419
+ | `pstn_g726` | 1.00 | ~0 | ~0 | ~300 Hz – 3.4 kHz |
420
+ | `voip_to_cellular_narrowband` | 1.00 | ~0 | ~0 | ~120 Hz – 3.3 kHz, roll-off to 4 kHz |
421
+ | `stress_multi_transcode` | 1.00 | ~0 | ~0 | ~350 Hz – 3 kHz |
422
+ | `voip_to_cellular_wideband` | 0.68 | 0.32 | ~0 | ~80 Hz – 6 kHz¹ |
423
+ | `voip_opus_wideband` | 0.64 | 0.36 | ~0 | ~80 Hz – 6.9 kHz |
424
+ | `voip_g722_wideband` | 0.59 | 0.41 | ~0 | ~70 Hz – 7 kHz |
425
+
426
+ ¹ AMR-WB at 12.65 kbit/s; `profile_params={"second_transcode": True}` adds an
427
+ AMR-NB interconnect hop that collapses the call to the narrowband path.
428
+
429
+ If a downstream system depends on fine spectral detail above 3.4 kHz, only the
430
+ wideband profiles keep it; the mobile narrowband path, the common case for a
431
+ call to an ordinary number, keeps the telephone band only.
432
+
433
+ ## Assumptions and limitations
434
+
435
+ - **No EVS, no G.729.** EVS has no open encoder and ffmpeg has no G.729 encoder
436
+ (an open one exists outside ffmpeg), so neither is offered; VoLTE with EVS is
437
+ not modelled.
438
+ - **G.722 and G.726 conceal with the G.711 Appendix I algorithm**, scaled to the
439
+ codec's rate; G.722's own Appendix III/IV PLC is not implemented. AMR and Opus
440
+ conceal with their own decoders.
441
+ - **G.711, G.722 and G.726 decoder state is never disturbed by an erasure.**
442
+ These codecs are decoded in full and the erased frames are replaced in the
443
+ decoded PCM, so the post-erasure divergence a real ADPCM receiver (G.722,
444
+ G.726) shows is not modelled; frames after an erasure equal a loss-free
445
+ decode.
446
+ - **`native` and `ffmpeg` G.711 are not bit-identical.** Same decoder; ffmpeg's
447
+ encoder table rounds differently at the decision levels, so 512 µ-law and 964
448
+ A-law of the 65536 int16 inputs take the adjacent code, one level apart.
449
+ - **Transport effects are statistical, not protocol-accurate.** Erasures follow
450
+ a two-state (Gilbert-Elliott) chain per codec hop; jitter-buffer adaptation
451
+ and level control are parametric models, not RTP/WebRTC implementations.
452
+ - **No acoustic path is modeled.** Room reverberation, speaker/handset
453
+ acoustics, acoustic echo, and the receiver's re-recording/re-digitization step
454
+ are out of scope; add them upstream if needed.
455
+
456
+ ## Checking a profile against your own path
457
+
458
+ Each profile is a physically motivated model of its path; no comparison against
459
+ recorded calls ships with this release. To check a profile against your own path, record the same clips through a real
460
+ call, run `analyze_channel(clean, received)` on both the real and the simulated
461
+ output, and compare the distributions. Prefer changing a parameter for a physical
462
+ reason over tuning it to a handful of recordings.
463
+
464
+ ## Further reading
465
+
466
+ - `examples/degrade.py` — degrade a WAV (or a synthetic tone) through a
467
+ profile and print channel metrics.
468
+ - `tests/test_phonesim.py` — executable specification: output rate, shape/type
469
+ preservation, determinism, length preservation, band-limiting behaviour, the
470
+ physics of each stage and batch processing are asserted there and double as
471
+ usage examples.
472
+ - Each stage and the high-level classes carry docstrings; e.g.
473
+ `help(phonesim.PhoneCallSimulator)`.