decibri 4.4.2 → 5.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +33 -3
- package/MIGRATION.md +70 -0
- package/README.md +73 -8
- package/examples/README.md +4 -3
- package/examples/decibri.browser.js +18 -6
- package/index.d.ts +126 -0
- package/index.js +53 -52
- package/models/README.md +149 -0
- package/models/fastenhancer_t.onnx +0 -0
- package/package.json +5 -5
- package/src/browser/decibri-browser.js +33 -13
- package/src/browser/index.d.ts +30 -18
- package/src/decibri.d.ts +326 -17
- package/src/decibri.js +573 -56
- package/src/errors.js +16 -1
package/CHANGELOG.md
CHANGED
|
@@ -9,7 +9,37 @@ For other decibri packages, see:
|
|
|
9
9
|
- Rust crate: [crates/decibri/CHANGELOG.md](../../crates/decibri/CHANGELOG.md)
|
|
10
10
|
- Python wheel: [bindings/python/CHANGELOG.md](../../bindings/python/CHANGELOG.md)
|
|
11
11
|
|
|
12
|
-
## [
|
|
12
|
+
## [5.1.0] - 2026-07-17
|
|
13
|
+
|
|
14
|
+
### Added
|
|
15
|
+
|
|
16
|
+
- `File`, an offline source: the same conditioning options as `Microphone` over a WAV file (`new File(path)` synchronously, `await File.open(path)` off the event loop) or a `Float32Array` of samples (`File.buffer(samples, { inputRate })`; a raw `Buffer` of bytes is rejected as ambiguous), delivered as a finite Readable stream of conditioned chunks.
|
|
17
|
+
- Whole-file speech analysis: `file.analyze()` (also spelled `file.analyse()`) resolves to a `VadReport` with per-window `scores` (`{ start, end, vadScore, isSpeech }`) and merged speech `segments` (`{ start, end }`), all in seconds of file time. Requires `vad: 'silero'`; a `File` opened without `vad` rejects with `analysis requires VAD`.
|
|
18
|
+
- Per-chunk VAD on files: `vadScore` and the `speech` / `silence` events alongside the stream, with the speaking holdoff measured in file time (sample positions) rather than wall-clock time, so processing speed never changes the reported events.
|
|
19
|
+
|
|
20
|
+
## [5.0.0] - 2026-06-24
|
|
21
|
+
|
|
22
|
+
### Added
|
|
23
|
+
|
|
24
|
+
- A device or driver failure during streaming now surfaces as a `DecibriError` with the dedicated `code` `'DEVICE_FAILED'`, and a non-ORT ONNX backend failure as `'ONNX_BACKEND_FAILED'`, instead of the generic `'DECIBRI_ERROR'`. Both are catchable by branching on `err.code`; the message text is unchanged.
|
|
25
|
+
- `dcRemoval`: an opt-in `Microphone` option that removes a constant (DC) offset from the captured audio with a one-pole DC-blocking high-pass. Set `true` to enable it; omit it or set `false` to leave it off (the default), which keeps the capture path byte-identical. Pure DSP: no bundled file or download is needed. It runs first in the chain, before denoise, and is same-length with no added latency, so `vadScore` and the `speech` / `silence` events are unaffected.
|
|
26
|
+
- `highpass`: an opt-in `Microphone` option that applies a high-pass filter to the captured audio, removing low-frequency rumble below the voice band. The accepted values are the cutoff in Hz, `80` (an 80 Hz second-order Butterworth high-pass) or `100` (a 100 Hz one); omit it to leave the high-pass off (the default), which keeps the capture path full-range. Pure DSP: no bundled file or download is needed. It runs after denoise in the chain and adds no latency, so `vadScore` and the `speech` / `silence` events are unaffected. An out-of-set value throws a `RangeError`. The closed cutoff set is designed to grow (further cutoffs are additive).
|
|
27
|
+
- `agc`: an opt-in `Microphone` option that applies automatic gain control to the captured audio, driving the running level toward a target with a smoothed, rate-limited gain. The value is a target level in dBFS (an integer in -40 to -3, typical -18); omit it to leave AGC off (the default), which keeps the level untouched. Pure DSP: no bundled file or download is needed. It runs after the high-pass step and adds no latency, so `vadScore` and the `speech` / `silence` events are unaffected. The opening is delivered at its natural level and reaches the target within tens of milliseconds (no opening window of wrong gain). An out-of-range value throws a `RangeError`.
|
|
28
|
+
- `limiter`: an opt-in `Microphone` option that applies a peak limiter to the captured audio, holding the signal at or below a sample-peak ceiling, the safety net that catches a transient the AGC's gain would let through. The value is a ceiling in dBFS (a number in -3.0 to 0.0, typical -1.0); omit it to leave the limiter off (the default), which keeps the level untouched. Pure DSP: no bundled file or download is needed. It runs last in the chain, after the AGC step, and adds no latency, so `vadScore` and the `speech` / `silence` events are unaffected. No output sample ever exceeds the ceiling, even on an instantaneous transient. An out-of-range value throws a `RangeError`.
|
|
29
|
+
- `denoise`: an opt-in `Microphone` option that runs a bundled single-channel speech-enhancement model over the captured audio. The only accepted value is `'fastenhancer-t'`; omit it to leave denoise off (the default), which keeps the capture path unchanged. The model ships with the package, so no path or download is needed. The delivered `'data'` chunks carry the enhanced audio; VAD reads the pre-enhancement signal, so `vadScore` and the `speech` / `silence` events are unaffected. An unknown model name throws a `TypeError`. A denoise model-load failure raises a dedicated `MODEL_LOAD_FAILED` `OrtError` via `wrapNativeError`; note that a failure surfaced from `start()` is delivered on the `'error'` event as the raw native error without the decibri `code` attached (as with device-open failures), so match it by message there.
|
|
30
|
+
|
|
31
|
+
### Changed
|
|
32
|
+
|
|
33
|
+
- A microphone whose native sample rate differs from the configured `sampleRate` is now resampled to the requested rate in the engine, so capture delivers audio at exactly the configured rate on every device. The device is opened at its native rate and decibri resamples (anti-aliased polyphase) to the target; a device already at the requested rate is unchanged (no resample). Chunk size and the reported rate are unchanged (still `framesPerBuffer * outputChannels * bytesPerSample` at the configured rate). Previously the requested rate was handed to the platform audio backend, which delivered it only when the device or OS could. Breaking for consumers whose device's native rate differs from the configured rate: the delivered samples are now decibri-resampled.
|
|
34
|
+
- Microphone capture is now mono only, narrowing a capability that shipped in 4.x. decibri 4.x delivered interleaved multichannel audio when a microphone was opened with `channels` greater than `1`; 5.0 captures mono, and the `channels` option accepts only `1` (the default). A value greater than `1` now throws a `RangeError` (`multichannel capture is not supported; channels must be 1 (mono)`), a clear error rather than a silent substitution of mono. The `channels` option is retained and the emitted `'data'` chunks keep their channel-general shape (`framesPerBuffer * bytesPerSample`, mono today), so multichannel can return later as an additive change rather than a further break: a future release may accept a value greater than `1` by delivering true interleaved multichannel. The intended longer-term multichannel direction is array ingest (consuming several channels internally for processing such as beamforming or array noise reduction while still delivering one conditioned stream), which is distinct from raw multichannel delivery. The default single-channel capture is unchanged. Breaking for consumers that opened a microphone with more than one channel: pass `channels: 1` or omit it.
|
|
35
|
+
- Capture now emits a final, possibly-shorter `'data'` chunk at stream close (on `stop()`), carrying the buffered tail, before `'end'`. Steady-state chunks are unchanged (still exactly `framesPerBuffer * channels * bytesPerSample`); only the last chunk before `'end'` may be shorter, and no captured audio is dropped. Internally the binding now re-blocks through the Rust core instead of a binding-side accumulator, and `stop()` defers the stream's end-of-stream signal by one tick so the flushed tail is delivered before `'end'` rather than dropped.
|
|
36
|
+
- The `vad` option now accepts a config object `{ model, threshold, holdoffMs }` (a `VadOptions`) alongside the `'silero'` / `'energy'` shorthand, and the flat `vadThreshold` / `vadHoldoff` options are removed. The shorthand is unchanged and keeps the default threshold (0.5 for `'silero'`, 0.01 for `'energy'`) and 300 ms holdoff; only code that tuned the threshold or holdoff migrates, by passing them on the object (`vad: { model: 'silero', threshold: 0.6, holdoffMs: 200 }`). An out-of-range threshold or a negative holdoff throws a `RangeError`, an unknown model a `TypeError`. Passing the removed `vadThreshold` or `vadHoldoff` throws a `TypeError`. The same object form applies to the browser build, where `model` is `'energy'`. Breaking for consumers that set `vadThreshold` or `vadHoldoff`; detection, the `vadScore` getter, and the `speech` / `silence` events are unchanged.
|
|
37
|
+
|
|
38
|
+
### Fixed
|
|
39
|
+
|
|
40
|
+
- Energy-mode VAD (`vad: 'energy'`) now reads the signal before the opt-in capture enhancement, so enabling an enhancement step (`denoise`, `highpass`, `agc`, `limiter`) no longer changes `vadScore` or the `speech` / `silence` events, matching the guarantee Silero mode already had. The energy RMS is computed natively on the pre-enhancement signal; previously it was computed on the delivered (post-enhancement) audio, so an enhancement step shifted the score (`agc` most of all, since it drives the delivered level toward its target, which could defeat energy endpointing). Energy detection with no enhancement enabled is unchanged.
|
|
41
|
+
- Resampled capture (a device whose native rate differs from the configured `sampleRate`) no longer drops the resampler's group-delay tail at stream close: the final, possibly-shorter `'data'` chunk(s) before `'end'` now carry it, so the complete resampled signal is delivered. A device already at the requested rate (no resample) is unchanged.
|
|
42
|
+
- On macOS, the microphone-permission error message now reads "System Settings > Privacy & Security" (the modern macOS wording) instead of the pre-Ventura "System Preferences > Security & Privacy".
|
|
13
43
|
|
|
14
44
|
## [4.4.2] - 2026-06-12
|
|
15
45
|
|
|
@@ -137,7 +167,7 @@ Error message wording on shipped `DecibriError` variants has historically been s
|
|
|
137
167
|
|
|
138
168
|
## [3.3.0] - 2026-04-23
|
|
139
169
|
|
|
140
|
-
Groundwork release for upcoming
|
|
170
|
+
Groundwork release for upcoming Python bindings. Adds a stable-ID form for audio device selection (`DeviceSelector::Id`), fixes a long-standing direction bug in `DecibriError::DeviceNotFound` when resolving output devices, exposes both in the Node binding, and extends the reference documentation with a Cargo feature flag guide plus additional crate-level rustdoc. No Node.js or browser API break. Direct Rust crate consumers pattern-matching on `DeviceSelector` or struct-literal-constructing `DeviceInfo` / `OutputDeviceInfo` need to update for the new `#[non_exhaustive]` attributes (see Migration notes below).
|
|
141
171
|
|
|
142
172
|
### Changed
|
|
143
173
|
|
|
@@ -160,7 +190,7 @@ Groundwork release for upcoming P3 Python bindings. Adds a stable-ID form for au
|
|
|
160
190
|
### Internal
|
|
161
191
|
|
|
162
192
|
- `DeviceDirection` trait gains a `not_found_error(String) -> DecibriError` method so `resolve_device_generic`'s `Name` and `Id` arms produce direction-correct errors via the `Input` / `Output` impls.
|
|
163
|
-
- Unit tests for `Arc<Mutex<CaptureStream>>` confirming the wrapping is `Send + Sync` (compile-time assertion) and serializes concurrent access across two threads (runtime test with `Barrier`). Documents the wrapping strategy the
|
|
193
|
+
- Unit tests for `Arc<Mutex<CaptureStream>>` confirming the wrapping is `Send + Sync` (compile-time assertion) and serializes concurrent access across two threads (runtime test with `Barrier`). Documents the wrapping strategy the Python binding will apply to share `!Sync` capture streams across Python threads.
|
|
164
194
|
- Crate-level rustdoc additions in `lib.rs`: a section on ORT error construction FFI side effects (the `ortsys![CreateStatus]` dylib-load trigger that motivates the `OrtPathInvalid` split from `OrtLoadFailed`) and a section on fork safety (guidance for Python `multiprocessing` consumers to use `spawn` start method).
|
|
165
195
|
- `lib.rs` rustdoc "Feature flags" section cross-references `docs/features.md` for consumers wanting the deep-dive reference.
|
|
166
196
|
- Em-dash cleanup across 19 code locations in `lib.rs`, `capture.rs`, `output.rs`, `vad.rs`, `error.rs`, `vad_integration.rs`, and `vad_ort_load_failure.rs`. Per CLAUDE.md, the codebase forbids em dashes; these were pre-existing violations.
|
package/MIGRATION.md
CHANGED
|
@@ -5,6 +5,76 @@ vocabulary that matches the Rust and Python packages, and tidies several option
|
|
|
5
5
|
and return shapes. This guide lists every breaking change with before and after
|
|
6
6
|
code.
|
|
7
7
|
|
|
8
|
+
## Breaking in 5.0.0: the vad config object
|
|
9
|
+
|
|
10
|
+
decibri 5.0.0 replaces the flat `vadThreshold` and `vadHoldoff` options with a
|
|
11
|
+
single `vad` config object that carries the model selector and its policy. The
|
|
12
|
+
`vad: 'silero'` and `vad: 'energy'` shorthand is unchanged and still uses the
|
|
13
|
+
default threshold (0.5 for `'silero'`, 0.01 for `'energy'`) and a 300 ms
|
|
14
|
+
holdoff, so only code that tuned the threshold or holdoff needs to migrate.
|
|
15
|
+
Move those values onto the object.
|
|
16
|
+
|
|
17
|
+
Before:
|
|
18
|
+
|
|
19
|
+
```js
|
|
20
|
+
new Microphone({ vad: 'silero', vadThreshold: 0.6, vadHoldoff: 200 });
|
|
21
|
+
new Microphone({ vad: 'energy', vadThreshold: 0.02 });
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
After:
|
|
25
|
+
|
|
26
|
+
```js
|
|
27
|
+
new Microphone({ vad: { model: 'silero', threshold: 0.6, holdoffMs: 200 } });
|
|
28
|
+
new Microphone({ vad: { model: 'energy', threshold: 0.02 } });
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
`vadThreshold` and `vadHoldoff` are no longer accepted; passing either throws a
|
|
32
|
+
`TypeError`. The shorthand without tuning is unchanged:
|
|
33
|
+
|
|
34
|
+
```js
|
|
35
|
+
new Microphone({ vad: 'silero' }); // unchanged
|
|
36
|
+
new Microphone({ vad: false }); // unchanged; the default
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
The same realignment applies to the browser build, where the object's `model`
|
|
40
|
+
is `'energy'` (the only browser detector): `vad: { model: 'energy', threshold: 0.02, holdoffMs: 200 }`.
|
|
41
|
+
|
|
42
|
+
## Breaking in 5.0.0: mono-only capture
|
|
43
|
+
|
|
44
|
+
decibri 4.x delivered interleaved multichannel audio when a microphone was
|
|
45
|
+
opened with `channels` greater than `1`. decibri 5.0.0 narrows capture to mono:
|
|
46
|
+
the `channels` option accepts only `1` (the default), and a value greater than
|
|
47
|
+
`1` throws a `RangeError` instead of being captured. If your code opened a
|
|
48
|
+
microphone with more than one channel, pass `channels: 1` (or omit it).
|
|
49
|
+
|
|
50
|
+
Before:
|
|
51
|
+
|
|
52
|
+
```js
|
|
53
|
+
new Microphone({ channels: 2 }); // 4.x: interleaved stereo capture
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
After:
|
|
57
|
+
|
|
58
|
+
```js
|
|
59
|
+
new Microphone({ channels: 1 }); // 5.0: mono (or omit channels entirely)
|
|
60
|
+
new Microphone(); // unchanged; mono is the default
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
A value greater than `1` now throws:
|
|
64
|
+
|
|
65
|
+
```js
|
|
66
|
+
new Microphone({ channels: 2 });
|
|
67
|
+
// RangeError: multichannel capture is not supported; channels must be 1 (mono)
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
The `channels` option and the channel-general `'data'` chunk shape are kept, so
|
|
71
|
+
multichannel may return later as an additive change (a future release accepting
|
|
72
|
+
a value greater than `1` by delivering true interleaved multichannel) rather
|
|
73
|
+
than a further break. The intended longer-term multichannel direction is array
|
|
74
|
+
ingest (consuming several channels internally for processing such as
|
|
75
|
+
beamforming or array noise reduction while still delivering one conditioned
|
|
76
|
+
stream), which is distinct from raw multichannel delivery.
|
|
77
|
+
|
|
8
78
|
## New in 4.2.0 (additive, nothing to migrate)
|
|
9
79
|
|
|
10
80
|
decibri 4.2.0 is a browser-only, additive release. Code written for 4.1.0 keeps
|
package/README.md
CHANGED
|
@@ -30,6 +30,22 @@ mic.on('data', (chunk) => { /* Buffer of Int16 PCM samples */ });
|
|
|
30
30
|
setTimeout(() => mic.stop(), 5000);
|
|
31
31
|
```
|
|
32
32
|
|
|
33
|
+
### Condition and analyze a file
|
|
34
|
+
|
|
35
|
+
```javascript
|
|
36
|
+
const { File } = require('decibri');
|
|
37
|
+
|
|
38
|
+
// The same conditioning chain as the live microphone, over a WAV file.
|
|
39
|
+
const file = await File.open('clip.wav', { denoise: 'fastenhancer-t', highpass: 80 });
|
|
40
|
+
file.on('data', (chunk) => { /* Buffer of conditioned Int16 PCM */ });
|
|
41
|
+
file.on('end', () => console.log('done'));
|
|
42
|
+
|
|
43
|
+
// Whole-file speech analysis (a live stream cannot do this).
|
|
44
|
+
const f = await File.open('clip.wav', { vad: 'silero' });
|
|
45
|
+
const report = await f.analyze();
|
|
46
|
+
for (const s of report.segments) console.log(s.start, s.end); // seconds
|
|
47
|
+
```
|
|
48
|
+
|
|
33
49
|
### Play audio
|
|
34
50
|
|
|
35
51
|
```javascript
|
|
@@ -82,18 +98,21 @@ Creates a Readable stream that captures from the microphone.
|
|
|
82
98
|
| Option | Type | Default | Description |
|
|
83
99
|
| --- | --- | --- | --- |
|
|
84
100
|
| `sampleRate` | number | 16000 | Samples per second (1000 to 384000) |
|
|
85
|
-
| `channels` | number | 1 |
|
|
101
|
+
| `channels` | number | 1 | Mono only: the only accepted value is `1`; a value greater than `1` throws a `RangeError` |
|
|
86
102
|
| `framesPerBuffer` | number | 1600 | Frames per chunk (64 to 65536). At 16kHz mono, 1600 = 100ms = 3200 bytes |
|
|
87
103
|
| `device` | number, string, or `{ id: string }` | system default | Device index, case-insensitive name substring, or stable per-host ID |
|
|
88
104
|
| `dtype` | `'int16'` \| `'float32'` | `'int16'` | Sample encoding |
|
|
89
|
-
| `vad` | `false` \| `'silero'` \| `'energy'` | `false` | Voice activity detection: disabled, the Silero ML model,
|
|
90
|
-
| `vadThreshold` | number | 0.5 / 0.01 | Speech threshold. Default is 0.5 for `'silero'`, 0.01 for `'energy'` |
|
|
91
|
-
| `vadHoldoff` | number | 300 | Silence holdoff in ms |
|
|
105
|
+
| `vad` | `false` \| `'silero'` \| `'energy'` \| `VadOptions` | `false` | Voice activity detection: disabled, the Silero ML model, an RMS energy threshold, or a config object `{ model, threshold, holdoffMs }` to tune the policy |
|
|
92
106
|
| `modelPath` | string | bundled model | Path to the Silero model. Only used when `vad` is `'silero'` |
|
|
107
|
+
| `dcRemoval` | boolean | off | Remove a constant (DC) offset with a one-pole DC-blocking high-pass. Runs first in the chain; same-length, no added latency |
|
|
108
|
+
| `denoise` | `'fastenhancer-t'` | off | Single-channel speech-enhancement (denoise) model. The model ships in the package; no path or download needed |
|
|
109
|
+
| `highpass` | `80` \| `100` | off | High-pass cutoff in Hz (second-order Butterworth) that removes low-frequency rumble. Runs after denoise. Out-of-set values throw a `RangeError` |
|
|
110
|
+
| `agc` | number | off | AGC target level in dBFS, an integer in -40 to -3 (typical -18). Runs after the high-pass. Out-of-range throws a `RangeError` |
|
|
111
|
+
| `limiter` | number | off | Peak limiter ceiling in dBFS, a number in -3.0 to 0.0 (typical -1.0). Runs last. Out-of-range throws a `RangeError` |
|
|
93
112
|
|
|
94
113
|
Standard `ReadableOptions` (e.g. `highWaterMark`) are also accepted.
|
|
95
114
|
|
|
96
|
-
`vad: true` is not accepted; pass the mode explicitly as `vad: 'silero'` or `vad: 'energy'`.
|
|
115
|
+
`vad: true` is not accepted; pass the mode explicitly as `vad: 'silero'` or `vad: 'energy'`. The string shorthand uses the default threshold (0.5 for `'silero'`, 0.01 for `'energy'`) and a 300 ms holdoff; to tune them pass a config object: `vad: { model: 'silero', threshold: 0.6, holdoffMs: 200 }`. The `threshold` (`0`–`1`) and `holdoffMs` fields are optional and fall back to those defaults.
|
|
97
116
|
|
|
98
117
|
### Methods
|
|
99
118
|
|
|
@@ -120,7 +139,7 @@ The module-level `inputDevices()` and `version()` free functions are equivalent
|
|
|
120
139
|
| `'data'` | Buffer | Audio chunk (Int16 LE or Float32 LE) |
|
|
121
140
|
| `'backpressure'` | - | Internal buffer full, consumer too slow |
|
|
122
141
|
| `'speech'` | - | VAD: audio crosses threshold |
|
|
123
|
-
| `'silence'` | - | VAD: audio below threshold for
|
|
142
|
+
| `'silence'` | - | VAD: audio below threshold for the holdoff period |
|
|
124
143
|
| `'end'` | - | Stream ended |
|
|
125
144
|
| `'error'` | Error | An error occurred |
|
|
126
145
|
|
|
@@ -267,7 +286,7 @@ To verify playback in real browsers, open `examples/browser-speaker-test.html` (
|
|
|
267
286
|
Lightweight RMS energy threshold. No model required.
|
|
268
287
|
|
|
269
288
|
```javascript
|
|
270
|
-
const mic = new Microphone({ vad: 'energy',
|
|
289
|
+
const mic = new Microphone({ vad: { model: 'energy', threshold: 0.01 } });
|
|
271
290
|
mic.on('speech', () => console.log('speaking'));
|
|
272
291
|
mic.on('silence', () => console.log('silent'));
|
|
273
292
|
```
|
|
@@ -277,13 +296,59 @@ mic.on('silence', () => console.log('silent'));
|
|
|
277
296
|
ML-based detection using the Silero VAD v5 model. More accurate than energy mode, especially in noisy environments.
|
|
278
297
|
|
|
279
298
|
```javascript
|
|
280
|
-
const mic = new Microphone({ vad: 'silero',
|
|
299
|
+
const mic = new Microphone({ vad: { model: 'silero', threshold: 0.5 } });
|
|
281
300
|
mic.on('speech', () => console.log('speaking'));
|
|
282
301
|
mic.on('silence', () => console.log('silent'));
|
|
283
302
|
```
|
|
284
303
|
|
|
285
304
|
The Silero model (~2MB) ships inside the npm package. No downloads or API keys required. Silero mode is Node.js only.
|
|
286
305
|
|
|
306
|
+
## decibri ACE (audio conditioning)
|
|
307
|
+
|
|
308
|
+
decibri ACE (Audio Capture Engine) is decibri's opt-in audio front-end for speech. It is a conditioning chain that runs on the captured audio before it reaches your `'data'` handler. Every stage is off by default, runs on-device, and needs no API key. With nothing enabled the capture path is byte-identical to plain capture.
|
|
309
|
+
|
|
310
|
+
The stages run in a fixed order. Enable any subset:
|
|
311
|
+
|
|
312
|
+
| Stage | Option | Range |
|
|
313
|
+
| --- | --- | --- |
|
|
314
|
+
| DC removal | `dcRemoval: true` | boolean |
|
|
315
|
+
| Denoise | `denoise: 'fastenhancer-t'` | the one bundled model |
|
|
316
|
+
| High-pass | `highpass: 80` or `100` | Hz |
|
|
317
|
+
| AGC | `agc: -18` | dBFS, -40 to -3 |
|
|
318
|
+
| Limiter | `limiter: -1.0` | dBFS, -3.0 to 0.0 |
|
|
319
|
+
|
|
320
|
+
```javascript
|
|
321
|
+
const { Microphone } = require('decibri');
|
|
322
|
+
|
|
323
|
+
const mic = new Microphone({
|
|
324
|
+
sampleRate: 16000,
|
|
325
|
+
denoise: 'fastenhancer-t', // bundled speech-enhancement model
|
|
326
|
+
highpass: 80, // remove low-frequency rumble
|
|
327
|
+
agc: -18, // target level in dBFS
|
|
328
|
+
limiter: -1.0, // peak ceiling in dBFS
|
|
329
|
+
vad: { model: 'silero', threshold: 0.5 },
|
|
330
|
+
});
|
|
331
|
+
|
|
332
|
+
mic.on('data', (chunk) => { /* Buffer of conditioned Int16 PCM */ });
|
|
333
|
+
mic.on('speech', () => console.log('speech'));
|
|
334
|
+
mic.on('silence', () => console.log('silence'));
|
|
335
|
+
setTimeout(() => mic.stop(), 5000);
|
|
336
|
+
```
|
|
337
|
+
|
|
338
|
+
VAD reads the signal before the chain, so `vadScore` and the `'speech'` / `'silence'` events are unaffected by which conditioning stages you enable. The conditioning chain runs in the native Node.js capture path; the browser build does not include it.
|
|
339
|
+
|
|
340
|
+
## API: File (offline source)
|
|
341
|
+
|
|
342
|
+
Everything a `Microphone` does to live audio, `File` does to audio you already have: the same conditioning options, the same Readable stream of conditioned chunks (finite: it ends at EOF), and the same opt-in `vad`. Because a `File` is a complete recording, it can also analyze the whole recording for speech.
|
|
343
|
+
|
|
344
|
+
- `await File.open(path, options?)`: read a WAV off the event loop (recommended, like `Microphone.open`).
|
|
345
|
+
- `new File(path, options?)`: the same result, synchronous (blocks on disk I/O; fine for scripts).
|
|
346
|
+
- `File.buffer(samples, options)`: wrap a `Float32Array` of samples you already hold. `options.inputRate` is required (raw samples carry no header); a raw `Buffer` of bytes is rejected as ambiguous.
|
|
347
|
+
- `await file.analyze()` (also spelled `analyse()`): consume the source and resolve to a `VadReport` of per-window `scores` (`{ start, end, vadScore, isSpeech }`) and merged speech `segments` (`{ start, end }`), in seconds of file time. Requires `vad: 'silero'`; a `File` opened without `vad` rejects with `analysis requires VAD`.
|
|
348
|
+
- `file.vadScore`, `'speech'` / `'silence'` events: per-chunk VAD alongside the stream, with the holdoff measured in FILE time (sample positions), never wall-clock time, so processing speed does not change the reported events.
|
|
349
|
+
|
|
350
|
+
Options mirror `Microphone` (`sampleRate`, `dtype`, `vad`, `dcRemoval`, `denoise`, `highpass`, `agc`, `limiter`); the live-capture options (`device`, `channels`, `framesPerBuffer`) do not apply. Iteration and analysis are separate single passes: construct one `File` per operation. Note: Node also has a global `File` (the web File API); import decibri's explicitly to avoid shadowing surprises.
|
|
351
|
+
|
|
287
352
|
## Device Selection
|
|
288
353
|
|
|
289
354
|
```javascript
|
package/examples/README.md
CHANGED
|
@@ -71,9 +71,10 @@ npx localtunnel --port 8080
|
|
|
71
71
|
### Regenerating the browser bundle
|
|
72
72
|
|
|
73
73
|
`decibri.browser.js` is generated from the package's browser entry
|
|
74
|
-
(`../src/browser/index.js`)
|
|
75
|
-
|
|
74
|
+
(`../src/browser/index.js`) with rolldown, the bundler used to produce the
|
|
75
|
+
checked-in build. Regenerate it after changing the browser source so the
|
|
76
|
+
bundle stays in sync:
|
|
76
77
|
|
|
77
78
|
```bash
|
|
78
|
-
npx
|
|
79
|
+
npx rolldown ../src/browser/index.js --format iife --name decibri --file decibri.browser.js
|
|
79
80
|
```
|
|
@@ -70,7 +70,7 @@ var decibri = (function() {
|
|
|
70
70
|
var require_decibri_browser = /* @__PURE__ */ __commonJSMin(((exports, module) => {
|
|
71
71
|
const { Emitter } = require_emitter();
|
|
72
72
|
const { WORKLET_SOURCE } = require_worklet_inline();
|
|
73
|
-
const VERSION = "
|
|
73
|
+
const VERSION = "5.1.0";
|
|
74
74
|
/**
|
|
75
75
|
* Browser microphone capture.
|
|
76
76
|
*
|
|
@@ -97,13 +97,27 @@ var decibri = (function() {
|
|
|
97
97
|
this._started = false;
|
|
98
98
|
this._starting = null;
|
|
99
99
|
this._stopRequested = false;
|
|
100
|
+
if (options.vadThreshold !== void 0 || options.vadHoldoff !== void 0) throw new TypeError("vadThreshold and vadHoldoff are no longer supported. Pass them on the vad config object: vad: { model: 'energy', threshold: 0.01, holdoffMs: 300 }.");
|
|
100
101
|
const vad = options.vad ?? false;
|
|
102
|
+
let vadThreshold = .01;
|
|
103
|
+
let vadHoldoff = 300;
|
|
101
104
|
if (vad === false) this._vad = false;
|
|
102
105
|
else if (vad === true) throw new TypeError("vad: true is no longer supported. Specify the mode explicitly: vad: 'energy'.");
|
|
103
106
|
else if (vad === "energy") this._vad = true;
|
|
104
|
-
else
|
|
105
|
-
|
|
106
|
-
|
|
107
|
+
else if (vad !== null && typeof vad === "object" && !Array.isArray(vad)) {
|
|
108
|
+
if (vad.model !== "energy") throw new TypeError(`Invalid vad model: ${JSON.stringify(vad.model)}. Expected 'energy'.`);
|
|
109
|
+
this._vad = true;
|
|
110
|
+
if (vad.threshold !== void 0) {
|
|
111
|
+
if (vad.threshold < 0 || vad.threshold > 1) throw new TypeError(`threshold must be between 0 and 1, got ${vad.threshold}`);
|
|
112
|
+
vadThreshold = vad.threshold;
|
|
113
|
+
}
|
|
114
|
+
if (vad.holdoffMs !== void 0) {
|
|
115
|
+
if (vad.holdoffMs < 0) throw new TypeError(`holdoffMs must be >= 0, got ${vad.holdoffMs}`);
|
|
116
|
+
vadHoldoff = vad.holdoffMs;
|
|
117
|
+
}
|
|
118
|
+
} else throw new TypeError(`Invalid vad value: ${JSON.stringify(vad)}. Expected false, 'energy', or a config object { model, threshold, holdoffMs }.`);
|
|
119
|
+
this._vadThreshold = vadThreshold;
|
|
120
|
+
this._vadHoldoff = vadHoldoff;
|
|
107
121
|
this._vadScore = 0;
|
|
108
122
|
this._isSpeaking = false;
|
|
109
123
|
this._silenceTimer = null;
|
|
@@ -119,8 +133,6 @@ var decibri = (function() {
|
|
|
119
133
|
if (this._channels < 1 || this._channels > 32) throw new TypeError(`channels must be between 1 and 32, got ${this._channels}`);
|
|
120
134
|
if (this._framesPerBuffer < 64 || this._framesPerBuffer > 65536) throw new TypeError(`frames per buffer must be between 64 and 65536, got ${this._framesPerBuffer}`);
|
|
121
135
|
if (this._dtype !== "int16" && this._dtype !== "float32") throw new TypeError("dtype must be 'int16' or 'float32'");
|
|
122
|
-
if (this._vadThreshold < 0 || this._vadThreshold > 1) throw new TypeError(`vadThreshold must be between 0 and 1, got ${this._vadThreshold}`);
|
|
123
|
-
if (this._vadHoldoff < 0) throw new TypeError(`vadHoldoff must be >= 0, got ${this._vadHoldoff}`);
|
|
124
136
|
}
|
|
125
137
|
/**
|
|
126
138
|
* Start microphone capture.
|
package/index.d.ts
CHANGED
|
@@ -78,6 +78,52 @@ export declare class DecibriOutputBridge {
|
|
|
78
78
|
static version(): VersionInfoJs
|
|
79
79
|
}
|
|
80
80
|
|
|
81
|
+
/**
|
|
82
|
+
* Native offline-source handle exposed to Node.js via napi-rs. The public
|
|
83
|
+
* `File` Readable lives in the JS wrapper; consumers construct that, not
|
|
84
|
+
* this handle, directly.
|
|
85
|
+
*/
|
|
86
|
+
export declare class FileHandle {
|
|
87
|
+
/**
|
|
88
|
+
* Open a WAV path as an offline source, synchronously (blocks on disk
|
|
89
|
+
* I/O; the JS wrapper's async `File.open` uses `openAsync` instead).
|
|
90
|
+
*/
|
|
91
|
+
static open(path: string, options?: FileOptions | undefined | null): FileHandle
|
|
92
|
+
/**
|
|
93
|
+
* Open a WAV path without blocking the JS event loop: the disk read,
|
|
94
|
+
* WAV parse, and chain construction run on the libuv thread pool.
|
|
95
|
+
*/
|
|
96
|
+
static openAsync(path: string, options?: FileOptions | undefined | null): Promise<unknown>
|
|
97
|
+
/**
|
|
98
|
+
* Wrap in-memory samples as an offline source. `samples` are mono f32 in
|
|
99
|
+
* [-1.0, 1.0]; `inputRate` is their native rate (raw samples carry no
|
|
100
|
+
* header). No I/O, so construction is synchronous.
|
|
101
|
+
*/
|
|
102
|
+
static buffer(samples: Float32Array, inputRate: number, options?: FileOptions | undefined | null): FileHandle
|
|
103
|
+
/**
|
|
104
|
+
* Pull the next conditioned chunk, advancing the per-chunk VAD score on
|
|
105
|
+
* the pre-conditioning feed. Returns `null` once the source is fully
|
|
106
|
+
* delivered (after the end-of-stream tail) or already consumed.
|
|
107
|
+
*/
|
|
108
|
+
readChunk(): Buffer | null
|
|
109
|
+
/**
|
|
110
|
+
* Consume the source with the core's whole-recording analysis, off the
|
|
111
|
+
* JS event loop. Resolves to the `VadReport`; a `File` built without VAD
|
|
112
|
+
* rejects with the core's typed error, never a silently constructed
|
|
113
|
+
* detector.
|
|
114
|
+
*/
|
|
115
|
+
analyze(): Promise<unknown>
|
|
116
|
+
/** Release the source. Idempotent; a closed File reads as ended. */
|
|
117
|
+
close(): void
|
|
118
|
+
/**
|
|
119
|
+
* Most recent per-chunk VAD score (0.0 to 1.0), computed on the
|
|
120
|
+
* pre-conditioning feed. 0.0 before the first chunk or with VAD off.
|
|
121
|
+
*/
|
|
122
|
+
get vadProbability(): number
|
|
123
|
+
/** The target output rate every delivered chunk carries. */
|
|
124
|
+
get sampleRate(): number
|
|
125
|
+
}
|
|
126
|
+
|
|
81
127
|
/**
|
|
82
128
|
* Options passed from JS constructor.
|
|
83
129
|
*
|
|
@@ -111,6 +157,40 @@ export interface DecibriOptions {
|
|
|
111
157
|
device?: any
|
|
112
158
|
vadMode?: string
|
|
113
159
|
modelPath?: string
|
|
160
|
+
/**
|
|
161
|
+
* Capture DC-removal toggle. When `true`, removes a constant (DC) offset
|
|
162
|
+
* from the captured audio with a one-pole DC-blocking high-pass, applied
|
|
163
|
+
* first in the transform chain (before denoise). Absent or `false` leaves
|
|
164
|
+
* it off (the default), a byte-identical no-op. Pure DSP: no bundled file
|
|
165
|
+
* and no model path, like `highpass`.
|
|
166
|
+
*/
|
|
167
|
+
dcRemoval?: boolean
|
|
168
|
+
/**
|
|
169
|
+
* Capture denoise model selector. The only accepted value is
|
|
170
|
+
* `'fastenhancer-t'`; absent leaves denoise off. The JS wrapper resolves
|
|
171
|
+
* the bundled model file and passes its path through `denoise_model_path`.
|
|
172
|
+
*/
|
|
173
|
+
denoise?: string
|
|
174
|
+
/**
|
|
175
|
+
* Capture high-pass filter cutoff in Hz. The accepted values are `80` (an
|
|
176
|
+
* 80 Hz second-order Butterworth high-pass) and `100` (a 100 Hz one);
|
|
177
|
+
* absent leaves the high-pass off. Pure DSP: no bundled file and no model
|
|
178
|
+
* path, unlike `denoise`.
|
|
179
|
+
*/
|
|
180
|
+
highpass?: number
|
|
181
|
+
/**
|
|
182
|
+
* Capture AGC target level in dBFS: an integer in `-40..=-3` (typical -18);
|
|
183
|
+
* absent leaves AGC off. Drives the captured level toward the target. Pure
|
|
184
|
+
* DSP: no bundled file and no model path, like `highpass`.
|
|
185
|
+
*/
|
|
186
|
+
agc?: number
|
|
187
|
+
/**
|
|
188
|
+
* Capture limiter ceiling in dBFS (sample-peak): a number in `-3.0..=0.0`
|
|
189
|
+
* (typical -1.0); absent leaves the limiter off. Holds the captured signal
|
|
190
|
+
* at or below the ceiling, catching a peak the AGC would let through. Pure
|
|
191
|
+
* DSP: no bundled file and no model path, like `agc`.
|
|
192
|
+
*/
|
|
193
|
+
limiter?: number
|
|
114
194
|
}
|
|
115
195
|
|
|
116
196
|
/** Options passed from JS constructor for output. */
|
|
@@ -136,6 +216,26 @@ export interface DeviceInfoJs {
|
|
|
136
216
|
isDefault: boolean
|
|
137
217
|
}
|
|
138
218
|
|
|
219
|
+
/**
|
|
220
|
+
* Options passed from the JS `File` wrapper. The conditioning fields mirror
|
|
221
|
+
* `DecibriOptions` exactly; the live-capture-only fields (device, channels,
|
|
222
|
+
* framesPerBuffer) do not apply to an offline source. `vadThreshold` and
|
|
223
|
+
* `vadHoldoffMs` are internal plumbing (the user passes them on the `vad`
|
|
224
|
+
* config object; the wrapper resolves them), hidden from the generated
|
|
225
|
+
* TypeScript like `ortLibraryPath`.
|
|
226
|
+
*/
|
|
227
|
+
export interface FileOptions {
|
|
228
|
+
sampleRate?: number
|
|
229
|
+
format?: string
|
|
230
|
+
vadMode?: string
|
|
231
|
+
modelPath?: string
|
|
232
|
+
dcRemoval?: boolean
|
|
233
|
+
denoise?: string
|
|
234
|
+
highpass?: number
|
|
235
|
+
agc?: number
|
|
236
|
+
limiter?: number
|
|
237
|
+
}
|
|
238
|
+
|
|
139
239
|
/** Output device info returned to JS. */
|
|
140
240
|
export interface OutputDeviceInfoJs {
|
|
141
241
|
index: number
|
|
@@ -150,6 +250,32 @@ export interface OutputDeviceInfoJs {
|
|
|
150
250
|
isDefault: boolean
|
|
151
251
|
}
|
|
152
252
|
|
|
253
|
+
/** One merged speech region of a recording, in seconds of file time. */
|
|
254
|
+
export interface Segment {
|
|
255
|
+
start: number
|
|
256
|
+
end: number
|
|
257
|
+
}
|
|
258
|
+
|
|
259
|
+
/**
|
|
260
|
+
* The whole-recording analysis `File.analyze()` resolves to: per-window
|
|
261
|
+
* scores and merged speech segments, in file order.
|
|
262
|
+
*/
|
|
263
|
+
export interface VadReport {
|
|
264
|
+
scores: Array<VadWindow>
|
|
265
|
+
segments: Array<Segment>
|
|
266
|
+
}
|
|
267
|
+
|
|
268
|
+
/**
|
|
269
|
+
* One scored voice-activity window of a recording: `start` / `end` in
|
|
270
|
+
* seconds of file time, the speech probability, and the raw threshold test.
|
|
271
|
+
*/
|
|
272
|
+
export interface VadWindow {
|
|
273
|
+
start: number
|
|
274
|
+
end: number
|
|
275
|
+
vadScore: number
|
|
276
|
+
isSpeech: boolean
|
|
277
|
+
}
|
|
278
|
+
|
|
153
279
|
/** Version info returned to JS. */
|
|
154
280
|
export interface VersionInfoJs {
|
|
155
281
|
decibri: string
|