speechdsp 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- speechdsp-0.1.0/LICENSE +21 -0
- speechdsp-0.1.0/PKG-INFO +550 -0
- speechdsp-0.1.0/README.md +513 -0
- speechdsp-0.1.0/pyproject.toml +82 -0
- speechdsp-0.1.0/setup.cfg +4 -0
- speechdsp-0.1.0/src/speechdsp/__init__.py +89 -0
- speechdsp-0.1.0/src/speechdsp/enhance.py +223 -0
- speechdsp-0.1.0/src/speechdsp/features.py +473 -0
- speechdsp-0.1.0/src/speechdsp/framing.py +268 -0
- speechdsp-0.1.0/src/speechdsp/io.py +302 -0
- speechdsp-0.1.0/src/speechdsp/metrics.py +341 -0
- speechdsp-0.1.0/src/speechdsp/py.typed +0 -0
- speechdsp-0.1.0/src/speechdsp/spectral.py +283 -0
- speechdsp-0.1.0/src/speechdsp/vad.py +252 -0
- speechdsp-0.1.0/src/speechdsp.egg-info/PKG-INFO +550 -0
- speechdsp-0.1.0/src/speechdsp.egg-info/SOURCES.txt +25 -0
- speechdsp-0.1.0/src/speechdsp.egg-info/dependency_links.txt +1 -0
- speechdsp-0.1.0/src/speechdsp.egg-info/requires.txt +13 -0
- speechdsp-0.1.0/src/speechdsp.egg-info/top_level.txt +1 -0
- speechdsp-0.1.0/tests/test_enhance.py +109 -0
- speechdsp-0.1.0/tests/test_features.py +216 -0
- speechdsp-0.1.0/tests/test_framing.py +104 -0
- speechdsp-0.1.0/tests/test_io.py +123 -0
- speechdsp-0.1.0/tests/test_metrics.py +167 -0
- speechdsp-0.1.0/tests/test_package.py +75 -0
- speechdsp-0.1.0/tests/test_spectral.py +102 -0
- speechdsp-0.1.0/tests/test_vad.py +107 -0
speechdsp-0.1.0/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 RL
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
speechdsp-0.1.0/PKG-INFO
ADDED
|
@@ -0,0 +1,550 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: speechdsp
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Digital signal processing, speech feature extraction and evaluation toolkit.
|
|
5
|
+
Author: RL
|
|
6
|
+
Maintainer: RL
|
|
7
|
+
License-Expression: MIT
|
|
8
|
+
Project-URL: Homepage, https://github.com/recklight/SpeechDsp
|
|
9
|
+
Project-URL: Repository, https://github.com/recklight/SpeechDsp
|
|
10
|
+
Project-URL: Issues, https://github.com/recklight/SpeechDsp/issues
|
|
11
|
+
Project-URL: Changelog, https://github.com/recklight/SpeechDsp/blob/master/CHANGELOG.md
|
|
12
|
+
Keywords: signal-processing,speech,mfcc,stft,speech-enhancement,voice-activity-detection,evaluation
|
|
13
|
+
Classifier: Development Status :: 4 - Beta
|
|
14
|
+
Classifier: Intended Audience :: Science/Research
|
|
15
|
+
Classifier: Operating System :: OS Independent
|
|
16
|
+
Classifier: Programming Language :: Python :: 3
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
20
|
+
Classifier: Topic :: Multimedia :: Sound/Audio :: Analysis
|
|
21
|
+
Classifier: Topic :: Scientific/Engineering :: Information Analysis
|
|
22
|
+
Classifier: Typing :: Typed
|
|
23
|
+
Requires-Python: >=3.10
|
|
24
|
+
Description-Content-Type: text/markdown
|
|
25
|
+
License-File: LICENSE
|
|
26
|
+
Requires-Dist: numpy>=1.24
|
|
27
|
+
Requires-Dist: scipy>=1.10
|
|
28
|
+
Requires-Dist: scikit-learn>=1.2
|
|
29
|
+
Provides-Extra: plot
|
|
30
|
+
Requires-Dist: matplotlib>=3.6; extra == "plot"
|
|
31
|
+
Provides-Extra: dl
|
|
32
|
+
Requires-Dist: torch>=2.2; extra == "dl"
|
|
33
|
+
Provides-Extra: dev
|
|
34
|
+
Requires-Dist: pytest>=7.2; extra == "dev"
|
|
35
|
+
Requires-Dist: ruff>=0.4; extra == "dev"
|
|
36
|
+
Dynamic: license-file
|
|
37
|
+
|
|
38
|
+
# SpeechDsp
|
|
39
|
+
|
|
40
|
+
[](https://pypi.org/project/speechdsp/)
|
|
41
|
+
[](https://github.com/recklight/SpeechDsp/actions/workflows/ci.yml)
|
|
42
|
+
[](https://www.python.org/downloads/)
|
|
43
|
+
[](LICENSE)
|
|
44
|
+
|
|
45
|
+
> 語音訊號處理、特徵擷取與模型評估的共用工具箱
|
|
46
|
+
|
|
47
|
+
`speechdsp` 是一套以 NumPy / SciPy / scikit-learn 為基礎、從零撰寫的 Python 套件,
|
|
48
|
+
提供語音與聲學研究常用的前端處理流程:音檔與特徵檔讀寫、分幀、短時傅立葉轉換、
|
|
49
|
+
MFCC 與差量特徵、端點偵測、語音增強,以及適合類別不平衡資料的評估指標。
|
|
50
|
+
|
|
51
|
+
本套件刻意不依賴 `librosa`、`soundfile` 或任何深度學習框架,所有演算法都以
|
|
52
|
+
`numpy` + `scipy` 自行實作,方便部署在沒有額外套件的運算環境(例如叢集節點或
|
|
53
|
+
乾淨的 conda 環境)。
|
|
54
|
+
|
|
55
|
+
- **作者**:RL
|
|
56
|
+
- **授權**:MIT
|
|
57
|
+
- **支援版本**:Python 3.10 以上
|
|
58
|
+
|
|
59
|
+
---
|
|
60
|
+
|
|
61
|
+
## 目錄
|
|
62
|
+
|
|
63
|
+
1. [安裝](#安裝)
|
|
64
|
+
2. [快速開始](#快速開始)
|
|
65
|
+
3. [模組說明](#模組說明)
|
|
66
|
+
- [speechdsp.io — 音檔與特徵檔存取](#speechdspio--音檔與特徵檔存取)
|
|
67
|
+
- [speechdsp.framing — 分幀與重疊相加](#speechdspframing--分幀與重疊相加)
|
|
68
|
+
- [speechdsp.spectral — 短時頻譜分析](#speechdspspectral--短時頻譜分析)
|
|
69
|
+
- [speechdsp.features — 倒頻譜特徵](#speechdspfeatures--倒頻譜特徵)
|
|
70
|
+
- [speechdsp.vad — 端點偵測](#speechdspvad--端點偵測)
|
|
71
|
+
- [speechdsp.enhance — 語音增強](#speechdspenhance--語音增強)
|
|
72
|
+
- [speechdsp.metrics — 評估指標](#speechdspmetrics--評估指標)
|
|
73
|
+
4. [API 速查表](#api-速查表)
|
|
74
|
+
5. [資料路徑說明](#資料路徑說明)
|
|
75
|
+
6. [測試與程式碼檢查](#測試與程式碼檢查)
|
|
76
|
+
7. [設計原則](#設計原則)
|
|
77
|
+
8. [References](#references)
|
|
78
|
+
9. [參與開發與引用](#參與開發與引用)
|
|
79
|
+
|
|
80
|
+
---
|
|
81
|
+
|
|
82
|
+
## 安裝
|
|
83
|
+
|
|
84
|
+
從 PyPI 安裝:
|
|
85
|
+
|
|
86
|
+
```bash
|
|
87
|
+
pip install speechdsp
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
若要修改原始碼,建議在虛擬環境中以可編輯模式安裝:
|
|
91
|
+
|
|
92
|
+
```bash
|
|
93
|
+
git clone https://github.com/recklight/SpeechDsp.git
|
|
94
|
+
cd SpeechDsp
|
|
95
|
+
python -m pip install -e .
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
若要一併安裝開發與測試工具:
|
|
99
|
+
|
|
100
|
+
```bash
|
|
101
|
+
python -m pip install -e ".[dev]"
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
### 相依套件
|
|
105
|
+
|
|
106
|
+
| 類別 | 套件 | 版本下限 |
|
|
107
|
+
| --- | --- | --- |
|
|
108
|
+
| 核心 | `numpy` | 1.24 |
|
|
109
|
+
| 核心 | `scipy` | 1.10 |
|
|
110
|
+
| 核心 | `scikit-learn` | 1.2 |
|
|
111
|
+
| 選用 `[plot]` | `matplotlib` | 3.6 |
|
|
112
|
+
| 選用 `[dl]` | `torch` | 2.2 |
|
|
113
|
+
| 選用 `[dev]` | `pytest`、`ruff` | — |
|
|
114
|
+
|
|
115
|
+
核心相依只有三個科學運算套件。`torch` 只是選用的額外項目,套件本身在任何情況下
|
|
116
|
+
都不會在載入時 import 它。
|
|
117
|
+
|
|
118
|
+
### 在其他專案中引用
|
|
119
|
+
|
|
120
|
+
其他專案可以用 path dependency 的方式直接指向本資料夾,例如在該專案的
|
|
121
|
+
`pyproject.toml` 中:
|
|
122
|
+
|
|
123
|
+
```toml
|
|
124
|
+
[project]
|
|
125
|
+
dependencies = ["speechdsp"]
|
|
126
|
+
|
|
127
|
+
[tool.uv.sources]
|
|
128
|
+
speechdsp = { path = "../SpeechDsp", editable = true }
|
|
129
|
+
```
|
|
130
|
+
|
|
131
|
+
或者最簡單的做法,先把本套件安裝進同一個環境即可:
|
|
132
|
+
|
|
133
|
+
```bash
|
|
134
|
+
python -m pip install -e /path/to/SpeechDsp
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
---
|
|
138
|
+
|
|
139
|
+
## 快速開始
|
|
140
|
+
|
|
141
|
+
```python
|
|
142
|
+
import numpy as np
|
|
143
|
+
import speechdsp
|
|
144
|
+
|
|
145
|
+
# 讀取音檔:回傳 float64 單聲道訊號(範圍 -1 ~ 1)與取樣率
|
|
146
|
+
x, sr = speechdsp.read_wav("example.wav")
|
|
147
|
+
|
|
148
|
+
# 去除頭尾靜音
|
|
149
|
+
x = speechdsp.trim_silence(x, sr)
|
|
150
|
+
|
|
151
|
+
# 語音增強(log-MMSE)
|
|
152
|
+
x = speechdsp.log_mmse(x, sr, noise_frames=6)
|
|
153
|
+
|
|
154
|
+
# 擷取 39 維 MFCC(13 維靜態 + 一階差量 + 二階差量)
|
|
155
|
+
feats = speechdsp.mfcc_with_deltas(x, sr, n_mfcc=13)
|
|
156
|
+
|
|
157
|
+
# 倒頻譜平均變異數正規化
|
|
158
|
+
feats = speechdsp.cmvn(feats)
|
|
159
|
+
|
|
160
|
+
print(feats.shape) # (幀數, 39)
|
|
161
|
+
```
|
|
162
|
+
|
|
163
|
+
---
|
|
164
|
+
|
|
165
|
+
## 模組說明
|
|
166
|
+
|
|
167
|
+
### speechdsp.io — 音檔與特徵檔存取
|
|
168
|
+
|
|
169
|
+
處理 WAV 與 HTK 特徵檔的讀寫。不論磁碟上是 16-bit PCM、32-bit PCM 或浮點格式,
|
|
170
|
+
`read_wav` 一律回傳 `float64` 的單聲道訊號,多聲道會自動平均混成單聲道。
|
|
171
|
+
|
|
172
|
+
```python
|
|
173
|
+
from speechdsp.io import read_wav, write_wav, read_htk, write_htk
|
|
174
|
+
|
|
175
|
+
x, sr = read_wav("input.wav") # x: float64, 範圍 -1 ~ 1
|
|
176
|
+
write_wav("output.wav", x, sr) # 以 16-bit PCM 寫出
|
|
177
|
+
|
|
178
|
+
# HTK 特徵檔(.mfc)
|
|
179
|
+
feats, header = read_htk("utt001.mfc")
|
|
180
|
+
print(header)
|
|
181
|
+
# {'n_samples': 312, 'samp_period': 100000, 'samp_size': 52,
|
|
182
|
+
# 'parm_kind': 6, 'n_dim': 13, 'format': 'binary'}
|
|
183
|
+
|
|
184
|
+
write_htk("utt001.mfc", feats, period_100ns=100000, kind=6)
|
|
185
|
+
```
|
|
186
|
+
|
|
187
|
+
HTK 二進位格式的表頭共 12 個位元組,全部使用**大端序(big-endian)**:
|
|
188
|
+
|
|
189
|
+
| 欄位 | 型別 | 說明 |
|
|
190
|
+
| --- | --- | --- |
|
|
191
|
+
| `nSamples` | int32 | 幀數 |
|
|
192
|
+
| `sampPeriod` | int32 | 幀週期,單位 100 奈秒(10 ms 即 100000) |
|
|
193
|
+
| `sampSize` | int16 | 每幀的位元組數(維度 × 4) |
|
|
194
|
+
| `parmKind` | int16 | 特徵種類代碼,`6` 代表 MFCC |
|
|
195
|
+
|
|
196
|
+
表頭之後是大端序的 `float32` 資料。部分工具會把特徵輸出成純文字的數值矩陣,
|
|
197
|
+
`read_htk` 會自動偵測並改用文字模式解析,此時回傳的 `header['format']` 為
|
|
198
|
+
`'text'`。
|
|
199
|
+
|
|
200
|
+
### speechdsp.framing — 分幀與重疊相加
|
|
201
|
+
|
|
202
|
+
短時分析的共用基礎。全套件的分幀慣例都定在這裡:第 `t` 幀從第 `t * hop` 個取樣點
|
|
203
|
+
開始、長度為 `frame_len`,尾端不足一幀的取樣點會被捨棄。
|
|
204
|
+
|
|
205
|
+
```python
|
|
206
|
+
from speechdsp.framing import enframe, overlap_add, frame_to_sample, frame_time
|
|
207
|
+
|
|
208
|
+
frames = enframe(x, frame_len=400, hop=160, window="hamming") # (幀數, 400)
|
|
209
|
+
y = overlap_add(frames, hop=160) # 疊回一維訊號
|
|
210
|
+
|
|
211
|
+
frame_to_sample(10, frame_len=400, hop=160) # 1600,第 10 幀的起始取樣點
|
|
212
|
+
frame_time(100, 400, 160, sr=16000) # 每一幀的中心時間(秒)
|
|
213
|
+
```
|
|
214
|
+
|
|
215
|
+
`enframe` 使用 `sliding_window_view` 建立跨步視圖後一次複製,不用 Python 迴圈,
|
|
216
|
+
長訊號也很快。
|
|
217
|
+
|
|
218
|
+
### speechdsp.spectral — 短時頻譜分析
|
|
219
|
+
|
|
220
|
+
STFT 與其反轉換。合成時會除以「窗函數平方的重疊相加包絡線」,而不是假設包絡線
|
|
221
|
+
是常數,因此只要窗函數滿足 NOLA 條件(預設的週期性 Hann 窗在 `hop <= n_fft // 2`
|
|
222
|
+
時都滿足),`istft(stft(x))` 就能把訊號完整還原,誤差在浮點精度等級。
|
|
223
|
+
|
|
224
|
+
```python
|
|
225
|
+
from speechdsp.spectral import stft, istft, spectrogram_db, spectrogram_image
|
|
226
|
+
|
|
227
|
+
S = stft(x, n_fft=512, hop=128) # (幀數, 257) 複數陣列
|
|
228
|
+
y = istft(S, n_fft=512, hop=128) # 完美重建,誤差 < 1e-10
|
|
229
|
+
|
|
230
|
+
db = spectrogram_db(x, sr, n_fft=512, hop=160, top_db=80.0) # 值域 -80 ~ 0 dB
|
|
231
|
+
|
|
232
|
+
img = spectrogram_image(x, sr, shape=(40, 98)) # uint8 灰階圖,固定 40x98
|
|
233
|
+
```
|
|
234
|
+
|
|
235
|
+
`spectrogram_image` 會把時間軸用線性內插重新取樣成固定寬度,無論語句長短都輸出
|
|
236
|
+
同一個尺寸,可直接餵給 CNN。第 0 列是最低的 mel 頻帶,第 0 行是最早的一幀。
|
|
237
|
+
|
|
238
|
+
### speechdsp.features — 倒頻譜特徵
|
|
239
|
+
|
|
240
|
+
MFCC 的完整流程:預強調 → 分幀加 Hamming 窗 → FFT 功率譜 → mel 三角濾波器組 →
|
|
241
|
+
取對數 → DCT-II → 倒頻譜提升(liftering)。
|
|
242
|
+
|
|
243
|
+
```python
|
|
244
|
+
from speechdsp.features import mfcc, mfcc_with_deltas, delta, mel_filterbank, cmvn
|
|
245
|
+
|
|
246
|
+
feats = mfcc(x, sr, n_mfcc=13, n_mels=26,
|
|
247
|
+
frame_ms=25.0, hop_ms=10.0, n_fft=512,
|
|
248
|
+
preemph=0.97, lifter=22, append_energy=True)
|
|
249
|
+
|
|
250
|
+
feats39 = mfcc_with_deltas(x, sr, n_mfcc=13) # [靜態 | 一階差量 | 二階差量]
|
|
251
|
+
feats39 = cmvn(feats39) # 平均與變異數正規化
|
|
252
|
+
|
|
253
|
+
fb = mel_filterbank(sr, n_fft=512, n_mels=26) # (26, 257)
|
|
254
|
+
```
|
|
255
|
+
|
|
256
|
+
幾個實作細節:
|
|
257
|
+
|
|
258
|
+
- **mel 轉換**使用標準公式 `m = 2595 * log10(1 + f / 700)`,三角濾波器的峰值
|
|
259
|
+
正規化為 1.0(HTK 慣例),因此濾波器組輸出的能量與輸入功率譜同單位。
|
|
260
|
+
- **`append_energy=True`** 時,第 0 維係數會換成該幀總能量的自然對數,這比 DCT 的
|
|
261
|
+
直流項更穩定,也是多數 ASR 前端的預設作法。
|
|
262
|
+
- **`n_fft` 自動調整**:若取樣率較高使得一幀的長度超過 `n_fft`(例如 44.1 kHz 的
|
|
263
|
+
25 ms 幀共 1102 點),會自動提升到下一個 2 的冪次並記錄一則警告,不會丟掉取樣點。
|
|
264
|
+
- **`delta`** 使用標準迴歸式,分母為 `2 * sum(n^2)`,邊界以複製方式補幀。對一段
|
|
265
|
+
線性斜坡輸入,內部區域會精確得到該斜率。
|
|
266
|
+
|
|
267
|
+
### speechdsp.vad — 端點偵測
|
|
268
|
+
|
|
269
|
+
以短時能量與過零率為基礎的雙門檻端點偵測。
|
|
270
|
+
|
|
271
|
+
```python
|
|
272
|
+
from speechdsp.vad import endpoint_detect, trim_silence, frame_energy, zero_crossing_rate
|
|
273
|
+
|
|
274
|
+
start, end = endpoint_detect(x, sr, frame_ms=25.0, hop_ms=10.0)
|
|
275
|
+
speech = x[start:end]
|
|
276
|
+
|
|
277
|
+
speech = trim_silence(x, sr) # 等同於上面兩行
|
|
278
|
+
```
|
|
279
|
+
|
|
280
|
+
演算法流程:先用開頭幾幀估計背景雜訊能量,再據此推出上門檻 `ITU` 與下門檻 `ITL`;
|
|
281
|
+
掃描出第一個超過 `ITU` 的幀之後,往回退到最後一個高於 `ITL` 的幀,尾端同理。最後
|
|
282
|
+
用過零率往外延伸,把能量偏低但過零率偏高的清擦音(例如 `/s/`、`/f/`)也納進來。
|
|
283
|
+
|
|
284
|
+
門檻全部由訊號自己推得,所以偵測結果不受錄音音量絕對值影響。過零率的延伸步驟另外
|
|
285
|
+
加了能量閘門,避免寬頻背景雜訊(過零率本來就接近 0.5)把邊界一路往外拖。
|
|
286
|
+
|
|
287
|
+
### speechdsp.enhance — 語音增強
|
|
288
|
+
|
|
289
|
+
兩種單通道頻域增強方法,都假設訊號開頭幾幀只有背景雜訊。
|
|
290
|
+
|
|
291
|
+
```python
|
|
292
|
+
from speechdsp.enhance import log_mmse, spectral_subtraction
|
|
293
|
+
|
|
294
|
+
y1 = log_mmse(x, sr, noise_frames=6, alpha=0.98)
|
|
295
|
+
y2 = spectral_subtraction(x, sr, noise_frames=6, over_sub=2.0, floor=0.002)
|
|
296
|
+
```
|
|
297
|
+
|
|
298
|
+
- **`log_mmse`**:log-MMSE 短時頻譜振幅估測。增益為
|
|
299
|
+
`G = xi / (1 + xi) * exp(0.5 * E1(v))`,其中 `xi` 是以 decision-directed 方式
|
|
300
|
+
遞迴估計的先驗訊雜比,`E1` 是指數積分(`scipy.special.exp1`)。相較於單純的
|
|
301
|
+
MMSE-STSA,它最小化的是**對數**頻譜振幅的誤差,比較貼近聽覺感受,殘留的
|
|
302
|
+
musical noise 也明顯較少。
|
|
303
|
+
- **`spectral_subtraction`**:功率頻譜相減,附帶過度相減係數與頻譜地板。
|
|
304
|
+
|
|
305
|
+
兩者都保留原始相位,輸出長度與輸入相同。
|
|
306
|
+
|
|
307
|
+
> **注意**:`noise_frames` 指的是開頭用來估計雜訊的幀數,這幾幀必須確實只有背景
|
|
308
|
+
> 雜訊。若音檔開頭沒有靜音段,請先自行補上,或改用其他雜訊估計方式。
|
|
309
|
+
|
|
310
|
+
### speechdsp.metrics — 評估指標
|
|
311
|
+
|
|
312
|
+
臨床或病理語音資料幾乎都有類別不平衡的問題,單看正確率(accuracy)會產生誤導:
|
|
313
|
+
一個永遠只猜多數類別的模型也可以有很漂亮的數字。因此本模組以
|
|
314
|
+
**UAR(unweighted average recall,各類別召回率的算術平均)**
|
|
315
|
+
與 **敏感度/特異度** 為主要指標。
|
|
316
|
+
|
|
317
|
+
```python
|
|
318
|
+
from speechdsp.metrics import uar, sensitivity_specificity, confusion_report, cross_val_report
|
|
319
|
+
from sklearn.svm import SVC
|
|
320
|
+
|
|
321
|
+
uar(y_true, y_pred) # 0.625
|
|
322
|
+
sensitivity_specificity(y_true, y_pred, pos_label=1) # (0.5, 0.75)
|
|
323
|
+
|
|
324
|
+
report = confusion_report(y_true, y_pred)
|
|
325
|
+
print(report["confusion_matrix"], report["uar"], report["macro_f1"])
|
|
326
|
+
|
|
327
|
+
# 五折交叉驗證,固定亂數種子,結果完全可重現
|
|
328
|
+
cv = cross_val_report(SVC(kernel="rbf"), X, y, n_splits=5, seed=0)
|
|
329
|
+
print(cv["mean"]["uar"], cv["std"]["uar"])
|
|
330
|
+
print(cv["pooled"]["confusion_matrix"])
|
|
331
|
+
```
|
|
332
|
+
|
|
333
|
+
`cross_val_report` 回傳的字典包含:
|
|
334
|
+
|
|
335
|
+
| 鍵值 | 內容 |
|
|
336
|
+
| --- | --- |
|
|
337
|
+
| `folds` | 每一折的完整報告(含混淆矩陣、UAR、正確率、敏感度、特異度、訓練/測試樣本數) |
|
|
338
|
+
| `mean` / `std` | 各折指標的平均與標準差 |
|
|
339
|
+
| `pooled` | 把所有折的 out-of-fold 預測合併後的整體報告 |
|
|
340
|
+
| `labels` | 類別順序 |
|
|
341
|
+
|
|
342
|
+
若同一位受試者有多筆錄音,請傳入 `groups` 參數,此時會改用
|
|
343
|
+
`StratifiedGroupKFold`,確保同一組樣本不會同時出現在訓練集與測試集,否則模型可能
|
|
344
|
+
學到的是「認人」而不是「辨識症狀」,分數會過度樂觀。
|
|
345
|
+
|
|
346
|
+
每一折都會以 `sklearn.base.clone` 複製一份新的估計器,傳進去的物件不會被修改。
|
|
347
|
+
|
|
348
|
+
---
|
|
349
|
+
|
|
350
|
+
## API 速查表
|
|
351
|
+
|
|
352
|
+
### `speechdsp.io`
|
|
353
|
+
|
|
354
|
+
| 函式 | 說明 |
|
|
355
|
+
| --- | --- |
|
|
356
|
+
| `read_wav(path) -> (np.ndarray, int)` | 讀 WAV,回傳 float64 單聲道訊號與取樣率 |
|
|
357
|
+
| `write_wav(path, x, sr) -> None` | 以 16-bit PCM 寫出 WAV |
|
|
358
|
+
| `read_htk(path) -> (np.ndarray, dict)` | 讀 HTK 特徵檔,自動偵測二進位/文字格式 |
|
|
359
|
+
| `write_htk(path, feats, period_100ns=100000, kind=6) -> None` | 寫出二進位 HTK 特徵檔 |
|
|
360
|
+
|
|
361
|
+
### `speechdsp.framing`
|
|
362
|
+
|
|
363
|
+
| 函式 | 說明 |
|
|
364
|
+
| --- | --- |
|
|
365
|
+
| `enframe(x, frame_len, hop, window=None)` | 分幀,回傳 `(幀數, frame_len)` |
|
|
366
|
+
| `overlap_add(frames, hop, window=None)` | 重疊相加,疊回一維訊號 |
|
|
367
|
+
| `num_frames(n_samples, frame_len, hop)` | 依分幀慣例計算的幀數 |
|
|
368
|
+
| `frame_to_sample(frame_idx, frame_len, hop)` | 幀索引 → 起始取樣點 |
|
|
369
|
+
| `frame_time(n_frames, frame_len, hop, sr)` | 各幀的中心時間(秒) |
|
|
370
|
+
|
|
371
|
+
### `speechdsp.spectral`
|
|
372
|
+
|
|
373
|
+
| 函式 | 說明 |
|
|
374
|
+
| --- | --- |
|
|
375
|
+
| `stft(x, n_fft, hop, window='hann', center=True)` | 短時傅立葉轉換,回傳複數矩陣 |
|
|
376
|
+
| `istft(S, n_fft, hop, window='hann', center=True)` | 反短時傅立葉轉換,可完美重建 |
|
|
377
|
+
| `power_spectrum(x, n_fft, hop)` | 每幀的功率譜 |
|
|
378
|
+
| `spectrogram_db(x, sr, n_fft, hop, top_db=80.0)` | 相對峰值的分貝頻譜圖 |
|
|
379
|
+
| `spectrogram_image(x, sr, shape=(40, 98), n_fft=512, hop=None)` | 固定尺寸 uint8 灰階頻譜圖 |
|
|
380
|
+
|
|
381
|
+
### `speechdsp.features`
|
|
382
|
+
|
|
383
|
+
| 函式 | 說明 |
|
|
384
|
+
| --- | --- |
|
|
385
|
+
| `preemphasis(x, coeff=0.97)` | 一階預強調濾波器 |
|
|
386
|
+
| `hz_to_mel(f)` / `mel_to_hz(m)` | Hz 與 mel 互換 |
|
|
387
|
+
| `mel_filterbank(sr, n_fft, n_mels=26, fmin=0.0, fmax=None)` | mel 三角濾波器組 |
|
|
388
|
+
| `mfcc(x, sr, n_mfcc=13, ...)` | MFCC 特徵 |
|
|
389
|
+
| `delta(feats, width=9)` | 迴歸式差量係數 |
|
|
390
|
+
| `mfcc_with_deltas(x, sr, **kw)` | `[靜態 \| 一階 \| 二階]` 串接 |
|
|
391
|
+
| `cmn(X)` / `cvn(X)` / `cmvn(X)` | 倒頻譜平均/變異數/兩者正規化 |
|
|
392
|
+
|
|
393
|
+
### `speechdsp.vad`
|
|
394
|
+
|
|
395
|
+
| 函式 | 說明 |
|
|
396
|
+
| --- | --- |
|
|
397
|
+
| `frame_energy(frames)` | 每幀的短時能量 |
|
|
398
|
+
| `zero_crossing_rate(frames)` | 每幀的過零率(0 ~ 1) |
|
|
399
|
+
| `endpoint_detect(x, sr, frame_ms=25.0, hop_ms=10.0)` | 回傳 `(起始取樣點, 結束取樣點)` |
|
|
400
|
+
| `trim_silence(x, sr, **kw)` | 直接切掉頭尾靜音 |
|
|
401
|
+
|
|
402
|
+
### `speechdsp.enhance`
|
|
403
|
+
|
|
404
|
+
| 函式 | 說明 |
|
|
405
|
+
| --- | --- |
|
|
406
|
+
| `log_mmse(x, sr, noise_frames=6, alpha=0.98)` | log-MMSE 頻譜振幅估測 |
|
|
407
|
+
| `spectral_subtraction(x, sr, noise_frames=6, over_sub=2.0, floor=0.002)` | 頻譜相減法 |
|
|
408
|
+
|
|
409
|
+
### `speechdsp.metrics`
|
|
410
|
+
|
|
411
|
+
| 函式 | 說明 |
|
|
412
|
+
| --- | --- |
|
|
413
|
+
| `uar(y_true, y_pred)` | 各類別召回率的平均 |
|
|
414
|
+
| `sensitivity_specificity(y_true, y_pred, pos_label=1)` | 敏感度與特異度 |
|
|
415
|
+
| `confusion_report(y_true, y_pred, labels=None)` | 混淆矩陣與各項指標的字典 |
|
|
416
|
+
| `cross_val_report(estimator, X, y, groups=None, n_splits=5, seed=0)` | 分層交叉驗證報告 |
|
|
417
|
+
|
|
418
|
+
---
|
|
419
|
+
|
|
420
|
+
## 資料路徑說明
|
|
421
|
+
|
|
422
|
+
**本套件不內含、也不會複製任何語料或音檔。** 所有資料路徑都應該由使用者自己的
|
|
423
|
+
設定檔提供,套件本身只接受已經載入記憶體的陣列或明確傳入的檔案路徑。
|
|
424
|
+
|
|
425
|
+
在你自己的專案中,建議把資料位置集中在一處設定,例如:
|
|
426
|
+
|
|
427
|
+
```python
|
|
428
|
+
from pathlib import Path
|
|
429
|
+
|
|
430
|
+
DATA_ROOT = Path("請改成你自己的資料夾路徑") # 預設值只是佔位字串
|
|
431
|
+
|
|
432
|
+
for wav_path in sorted(DATA_ROOT.glob("*.wav")):
|
|
433
|
+
x, sr = speechdsp.read_wav(wav_path)
|
|
434
|
+
...
|
|
435
|
+
```
|
|
436
|
+
|
|
437
|
+
`tests/` 目錄下的測試完全使用程式合成的小型訊號(正弦波、白雜訊、線性斜坡),
|
|
438
|
+
不需要任何外部檔案就能執行。
|
|
439
|
+
|
|
440
|
+
---
|
|
441
|
+
|
|
442
|
+
## 測試與程式碼檢查
|
|
443
|
+
|
|
444
|
+
```bash
|
|
445
|
+
# 執行測試(pyproject.toml 已設定好 pythonpath,不需另外設環境變數)
|
|
446
|
+
python -m pytest -q
|
|
447
|
+
|
|
448
|
+
# 程式碼風格與靜態檢查
|
|
449
|
+
python -m ruff check .
|
|
450
|
+
python -m ruff format --check .
|
|
451
|
+
```
|
|
452
|
+
|
|
453
|
+
CI 會在 Ubuntu 與 Windows 上,以 Python 3.10 至 3.13 執行上述檢查。
|
|
454
|
+
|
|
455
|
+
測試涵蓋的數值驗證包括:
|
|
456
|
+
|
|
457
|
+
- 已知頻率的正弦波,其 STFT 峰值落在正確的頻率 bin
|
|
458
|
+
- `istft(stft(x))` 的重建誤差小於 `1e-10`(多種窗函數與 hop 組合)
|
|
459
|
+
- `enframe` / `overlap_add` 往返一致
|
|
460
|
+
- mel 濾波器組的峰值增益、中心頻率單調遞增、全頻帶覆蓋
|
|
461
|
+
- MFCC 對常數訊號、數位靜音、不同取樣率的行為
|
|
462
|
+
- 線性斜坡訊號的 `delta` 應等於斜率
|
|
463
|
+
- HTK 檔案寫入再讀回完全一致,且表頭確實為大端序
|
|
464
|
+
- `log_mmse` 與 `spectral_subtraction` 對加噪訊號能提升訊雜比
|
|
465
|
+
- `endpoint_detect` 對「靜音-語音-靜音」合成訊號抓到正確邊界
|
|
466
|
+
- `uar` 與 `sensitivity_specificity` 對照手算的小例子
|
|
467
|
+
|
|
468
|
+
### 相容性
|
|
469
|
+
|
|
470
|
+
程式碼以 **Python 3.10** 為目標版本撰寫(Ubuntu 22.04、Google Colab 與多數 conda
|
|
471
|
+
環境的預設版本),同時相容 NumPy 1.x 與 2.x。若要確認語法沒有用到 3.11 以後才有的
|
|
472
|
+
特性:
|
|
473
|
+
|
|
474
|
+
```bash
|
|
475
|
+
python -c "import ast,pathlib; [ast.parse(f.read_text(encoding='utf-8'), filename=str(f), feature_version=(3,10)) for d in ('src','tests') for f in pathlib.Path(d).rglob('*.py')]"
|
|
476
|
+
```
|
|
477
|
+
|
|
478
|
+
---
|
|
479
|
+
|
|
480
|
+
## 設計原則
|
|
481
|
+
|
|
482
|
+
1. **向量化優先。** 能用 NumPy 廣播解決的就不寫 Python 迴圈。唯一保留迴圈的地方是
|
|
483
|
+
log-MMSE 的 decision-directed 遞迴(本質上必須逐幀進行),但即使如此,每一次
|
|
484
|
+
迭代仍然是對整個頻率軸一次算完。
|
|
485
|
+
2. **慣例集中管理。** 分幀的定義只寫在 `speechdsp.framing` 一處,其他模組全部沿用,
|
|
486
|
+
避免不同模組對「第幾幀從哪裡開始」有不同解讀。
|
|
487
|
+
3. **明確的錯誤訊息。** 尺寸不合、參數超出範圍時直接丟出帶有實際數值的
|
|
488
|
+
`ValueError`,不要讓錯誤在後面幾層才以奇怪的形式爆出來。
|
|
489
|
+
4. **完整的型別註解與 numpy-style docstring。** 每個公開函式都標註了
|
|
490
|
+
Parameters / Returns / References,方便 IDE 與文件工具正確呈現。
|
|
491
|
+
5. **不強制深度學習框架。** `torch` 只是選用額外項目,套件本身不會 import 它,
|
|
492
|
+
沒有安裝也不影響套件載入與測試(`tests/test_package.py` 有對應檢查)。
|
|
493
|
+
|
|
494
|
+
---
|
|
495
|
+
|
|
496
|
+
## References
|
|
497
|
+
|
|
498
|
+
演算法實作參考下列文獻:
|
|
499
|
+
|
|
500
|
+
1. S. B. Davis and P. Mermelstein, "Comparison of parametric representations for
|
|
501
|
+
monosyllabic word recognition in continuously spoken sentences," *IEEE
|
|
502
|
+
Transactions on Acoustics, Speech, and Signal Processing*, vol. 28, no. 4,
|
|
503
|
+
pp. 357–366, 1980.
|
|
504
|
+
2. S. S. Stevens, J. Volkmann, and E. B. Newman, "A scale for the measurement of
|
|
505
|
+
the psychological magnitude pitch," *The Journal of the Acoustical Society of
|
|
506
|
+
America*, vol. 8, no. 3, pp. 185–190, 1937.
|
|
507
|
+
3. Y. Ephraim and D. Malah, "Speech enhancement using a minimum mean-square error
|
|
508
|
+
short-time spectral amplitude estimator," *IEEE Transactions on Acoustics,
|
|
509
|
+
Speech, and Signal Processing*, vol. 32, no. 6, pp. 1109–1121, 1984.
|
|
510
|
+
4. Y. Ephraim and D. Malah, "Speech enhancement using a minimum mean-square error
|
|
511
|
+
log-spectral amplitude estimator," *IEEE Transactions on Acoustics, Speech, and
|
|
512
|
+
Signal Processing*, vol. 33, no. 2, pp. 443–445, 1985.
|
|
513
|
+
5. S. F. Boll, "Suppression of acoustic noise in speech using spectral
|
|
514
|
+
subtraction," *IEEE Transactions on Acoustics, Speech, and Signal Processing*,
|
|
515
|
+
vol. 27, no. 2, pp. 113–120, 1979.
|
|
516
|
+
6. M. Berouti, R. Schwartz, and J. Makhoul, "Enhancement of speech corrupted by
|
|
517
|
+
acoustic noise," *ICASSP*, vol. 4, pp. 208–211, 1979.
|
|
518
|
+
7. L. R. Rabiner and M. R. Sambur, "An algorithm for determining the endpoints of
|
|
519
|
+
isolated utterances," *Bell System Technical Journal*, vol. 54, no. 2,
|
|
520
|
+
pp. 297–315, 1975.
|
|
521
|
+
8. L. R. Rabiner and R. W. Schafer, *Digital Processing of Speech Signals*,
|
|
522
|
+
Prentice-Hall, 1978.
|
|
523
|
+
9. J. B. Allen and L. R. Rabiner, "A unified approach to short-time Fourier
|
|
524
|
+
analysis and synthesis," *Proceedings of the IEEE*, vol. 65, no. 11,
|
|
525
|
+
pp. 1558–1564, 1977.
|
|
526
|
+
10. D. W. Griffin and J. S. Lim, "Signal estimation from modified short-time
|
|
527
|
+
Fourier transform," *IEEE Transactions on Acoustics, Speech, and Signal
|
|
528
|
+
Processing*, vol. 32, no. 2, pp. 236–243, 1984.
|
|
529
|
+
11. S. Young et al., *The HTK Book (version 3.4)*, Cambridge University
|
|
530
|
+
Engineering Department, 2006.
|
|
531
|
+
12. A. Rosenberg, "Classifying skewed data: Importance weighting to optimize
|
|
532
|
+
average recall," *Interspeech*, 2012.
|
|
533
|
+
13. T. Fawcett, "An introduction to ROC analysis," *Pattern Recognition Letters*,
|
|
534
|
+
vol. 27, no. 8, pp. 861–874, 2006.
|
|
535
|
+
|
|
536
|
+
---
|
|
537
|
+
|
|
538
|
+
## 參與開發與引用
|
|
539
|
+
|
|
540
|
+
- 開發環境、提交前檢查與程式碼風格:見 [CONTRIBUTING.md](CONTRIBUTING.md)
|
|
541
|
+
- 各版本的變更:見 [CHANGELOG.md](CHANGELOG.md)
|
|
542
|
+
- 在研究中使用本套件時,引用資訊見 [CITATION.cff](CITATION.cff)
|
|
543
|
+
(GitHub 頁面右側的「Cite this repository」也會讀取這個檔案)
|
|
544
|
+
- 錯誤回報與功能建議請使用 [Issues](https://github.com/recklight/SpeechDsp/issues)
|
|
545
|
+
|
|
546
|
+
---
|
|
547
|
+
|
|
548
|
+
## 授權
|
|
549
|
+
|
|
550
|
+
MIT License, Copyright (c) RL。詳見 [LICENSE](LICENSE)。
|