decibri 4.4.1 → 5.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +35 -3
- package/MIGRATION.md +70 -0
- package/README.md +45 -8
- package/examples/README.md +4 -3
- package/examples/decibri.browser.js +18 -6
- package/index.d.ts +41 -0
- package/index.js +52 -52
- package/models/README.md +149 -0
- package/models/fastenhancer_t.onnx +0 -0
- package/package.json +5 -5
- package/src/browser/decibri-browser.js +33 -13
- package/src/browser/index.d.ts +30 -18
- package/src/decibri.d.ts +104 -17
- package/src/decibri.js +197 -55
- package/src/errors.js +14 -0
package/CHANGELOG.md
CHANGED
|
@@ -9,7 +9,39 @@ For other decibri packages, see:
|
|
|
9
9
|
- Rust crate: [crates/decibri/CHANGELOG.md](../../crates/decibri/CHANGELOG.md)
|
|
10
10
|
- Python wheel: [bindings/python/CHANGELOG.md](../../bindings/python/CHANGELOG.md)
|
|
11
11
|
|
|
12
|
-
## [
|
|
12
|
+
## [5.0.0] - 2026-06-24
|
|
13
|
+
|
|
14
|
+
### Added
|
|
15
|
+
|
|
16
|
+
- A device or driver failure during streaming now surfaces as a `DecibriError` with the dedicated `code` `'DEVICE_FAILED'`, and a non-ORT ONNX backend failure as `'ONNX_BACKEND_FAILED'`, instead of the generic `'DECIBRI_ERROR'`. Both are catchable by branching on `err.code`; the message text is unchanged.
|
|
17
|
+
- `dcRemoval`: an opt-in `Microphone` option that removes a constant (DC) offset from the captured audio with a one-pole DC-blocking high-pass. Set `true` to enable it; omit it or set `false` to leave it off (the default), which keeps the capture path byte-identical. Pure DSP: no bundled file or download is needed. It runs first in the chain, before denoise, and is same-length with no added latency, so `vadScore` and the `speech` / `silence` events are unaffected.
|
|
18
|
+
- `highpass`: an opt-in `Microphone` option that applies a high-pass filter to the captured audio, removing low-frequency rumble below the voice band. The accepted values are the cutoff in Hz, `80` (an 80 Hz second-order Butterworth high-pass) or `100` (a 100 Hz one); omit it to leave the high-pass off (the default), which keeps the capture path full-range. Pure DSP: no bundled file or download is needed. It runs after denoise in the chain and adds no latency, so `vadScore` and the `speech` / `silence` events are unaffected. An out-of-set value throws a `RangeError`. The closed cutoff set is designed to grow (further cutoffs are additive).
|
|
19
|
+
- `agc`: an opt-in `Microphone` option that applies automatic gain control to the captured audio, driving the running level toward a target with a smoothed, rate-limited gain. The value is a target level in dBFS (an integer in -40 to -3, typical -18); omit it to leave AGC off (the default), which keeps the level untouched. Pure DSP: no bundled file or download is needed. It runs after the high-pass step and adds no latency, so `vadScore` and the `speech` / `silence` events are unaffected. The opening is delivered at its natural level and reaches the target within tens of milliseconds (no opening window of wrong gain). An out-of-range value throws a `RangeError`.
|
|
20
|
+
- `limiter`: an opt-in `Microphone` option that applies a peak limiter to the captured audio, holding the signal at or below a sample-peak ceiling, the safety net that catches a transient the AGC's gain would let through. The value is a ceiling in dBFS (a number in -3.0 to 0.0, typical -1.0); omit it to leave the limiter off (the default), which keeps the level untouched. Pure DSP: no bundled file or download is needed. It runs last in the chain, after the AGC step, and adds no latency, so `vadScore` and the `speech` / `silence` events are unaffected. No output sample ever exceeds the ceiling, even on an instantaneous transient. An out-of-range value throws a `RangeError`.
|
|
21
|
+
- `denoise`: an opt-in `Microphone` option that runs a bundled single-channel speech-enhancement model over the captured audio. The only accepted value is `'fastenhancer-t'`; omit it to leave denoise off (the default), which keeps the capture path unchanged. The model ships with the package, so no path or download is needed. The delivered `'data'` chunks carry the enhanced audio; VAD reads the pre-enhancement signal, so `vadScore` and the `speech` / `silence` events are unaffected. An unknown model name throws a `TypeError`. A denoise model-load failure raises a dedicated `MODEL_LOAD_FAILED` `OrtError` via `wrapNativeError`; note that a failure surfaced from `start()` is delivered on the `'error'` event as the raw native error without the decibri `code` attached (as with device-open failures), so match it by message there.
|
|
22
|
+
|
|
23
|
+
### Changed
|
|
24
|
+
|
|
25
|
+
- A microphone whose native sample rate differs from the configured `sampleRate` is now resampled to the requested rate in the engine, so capture delivers audio at exactly the configured rate on every device. The device is opened at its native rate and decibri resamples (anti-aliased polyphase) to the target; a device already at the requested rate is unchanged (no resample). Chunk size and the reported rate are unchanged (still `framesPerBuffer * outputChannels * bytesPerSample` at the configured rate). Previously the requested rate was handed to the platform audio backend, which delivered it only when the device or OS could. Breaking for consumers whose device's native rate differs from the configured rate: the delivered samples are now decibri-resampled.
|
|
26
|
+
- Microphone capture is now mono only, narrowing a capability that shipped in 4.x. decibri 4.x delivered interleaved multichannel audio when a microphone was opened with `channels` greater than `1`; 5.0 captures mono, and the `channels` option accepts only `1` (the default). A value greater than `1` now throws a `RangeError` (`multichannel capture is not supported; channels must be 1 (mono)`), a clear error rather than a silent substitution of mono. The `channels` option is retained and the emitted `'data'` chunks keep their channel-general shape (`framesPerBuffer * bytesPerSample`, mono today), so multichannel can return later as an additive change rather than a further break: a future release may accept a value greater than `1` by delivering true interleaved multichannel. The intended longer-term multichannel direction is array ingest (consuming several channels internally for processing such as beamforming or array noise reduction while still delivering one conditioned stream), which is distinct from raw multichannel delivery. The default single-channel capture is unchanged. Breaking for consumers that opened a microphone with more than one channel: pass `channels: 1` or omit it.
|
|
27
|
+
- Capture now emits a final, possibly-shorter `'data'` chunk at stream close (on `stop()`), carrying the buffered tail, before `'end'`. Steady-state chunks are unchanged (still exactly `framesPerBuffer * channels * bytesPerSample`); only the last chunk before `'end'` may be shorter, and no captured audio is dropped. Internally the binding now re-blocks through the Rust core instead of a binding-side accumulator, and `stop()` defers the stream's end-of-stream signal by one tick so the flushed tail is delivered before `'end'` rather than dropped.
|
|
28
|
+
- The `vad` option now accepts a config object `{ model, threshold, holdoffMs }` (a `VadOptions`) alongside the `'silero'` / `'energy'` shorthand, and the flat `vadThreshold` / `vadHoldoff` options are removed. The shorthand is unchanged and keeps the default threshold (0.5 for `'silero'`, 0.01 for `'energy'`) and 300 ms holdoff; only code that tuned the threshold or holdoff migrates, by passing them on the object (`vad: { model: 'silero', threshold: 0.6, holdoffMs: 200 }`). An out-of-range threshold or a negative holdoff throws a `RangeError`, an unknown model a `TypeError`. Passing the removed `vadThreshold` or `vadHoldoff` throws a `TypeError`. The same object form applies to the browser build, where `model` is `'energy'`. Breaking for consumers that set `vadThreshold` or `vadHoldoff`; detection, the `vadScore` getter, and the `speech` / `silence` events are unchanged.
|
|
29
|
+
|
|
30
|
+
### Fixed
|
|
31
|
+
|
|
32
|
+
- Energy-mode VAD (`vad: 'energy'`) now reads the signal before the opt-in capture enhancement, so enabling an enhancement step (`denoise`, `highpass`, `agc`, `limiter`) no longer changes `vadScore` or the `speech` / `silence` events, matching the guarantee Silero mode already had. The energy RMS is computed natively on the pre-enhancement signal; previously it was computed on the delivered (post-enhancement) audio, so an enhancement step shifted the score (`agc` most of all, since it drives the delivered level toward its target, which could defeat energy endpointing). Energy detection with no enhancement enabled is unchanged.
|
|
33
|
+
- Resampled capture (a device whose native rate differs from the configured `sampleRate`) no longer drops the resampler's group-delay tail at stream close: the final, possibly-shorter `'data'` chunk(s) before `'end'` now carry it, so the complete resampled signal is delivered. A device already at the requested rate (no resample) is unchanged.
|
|
34
|
+
- On macOS, the microphone-permission error message now reads "System Settings > Privacy & Security" (the modern macOS wording) instead of the pre-Ventura "System Preferences > Security & Privacy".
|
|
35
|
+
|
|
36
|
+
## [4.4.2] - 2026-06-12
|
|
37
|
+
|
|
38
|
+
### Fixed
|
|
39
|
+
|
|
40
|
+
- Multichannel capture is now downmixed to mono before the Silero VAD, so `vadScore` and the `speech` / `silence` events are correct on devices opened with more than one channel. Previously the interleaved multichannel audio reached the VAD as if it were mono, so it scored garbled input. Only affected explicit multichannel capture with VAD enabled (the default is one channel); emitted audio is unchanged.
|
|
41
|
+
|
|
42
|
+
### Added
|
|
43
|
+
|
|
44
|
+
- `overrunCount`: a read-only accessor on `Microphone` exposing the core stream's dropped-buffer counter. 0 while the consumer keeps pace; a rising value means audio is being dropped to bound memory.
|
|
13
45
|
|
|
14
46
|
## [4.4.1] - 2026-06-11
|
|
15
47
|
|
|
@@ -127,7 +159,7 @@ Error message wording on shipped `DecibriError` variants has historically been s
|
|
|
127
159
|
|
|
128
160
|
## [3.3.0] - 2026-04-23
|
|
129
161
|
|
|
130
|
-
Groundwork release for upcoming
|
|
162
|
+
Groundwork release for upcoming Python bindings. Adds a stable-ID form for audio device selection (`DeviceSelector::Id`), fixes a long-standing direction bug in `DecibriError::DeviceNotFound` when resolving output devices, exposes both in the Node binding, and extends the reference documentation with a Cargo feature flag guide plus additional crate-level rustdoc. No Node.js or browser API break. Direct Rust crate consumers pattern-matching on `DeviceSelector` or struct-literal-constructing `DeviceInfo` / `OutputDeviceInfo` need to update for the new `#[non_exhaustive]` attributes (see Migration notes below).
|
|
131
163
|
|
|
132
164
|
### Changed
|
|
133
165
|
|
|
@@ -150,7 +182,7 @@ Groundwork release for upcoming P3 Python bindings. Adds a stable-ID form for au
|
|
|
150
182
|
### Internal
|
|
151
183
|
|
|
152
184
|
- `DeviceDirection` trait gains a `not_found_error(String) -> DecibriError` method so `resolve_device_generic`'s `Name` and `Id` arms produce direction-correct errors via the `Input` / `Output` impls.
|
|
153
|
-
- Unit tests for `Arc<Mutex<CaptureStream>>` confirming the wrapping is `Send + Sync` (compile-time assertion) and serializes concurrent access across two threads (runtime test with `Barrier`). Documents the wrapping strategy the
|
|
185
|
+
- Unit tests for `Arc<Mutex<CaptureStream>>` confirming the wrapping is `Send + Sync` (compile-time assertion) and serializes concurrent access across two threads (runtime test with `Barrier`). Documents the wrapping strategy the Python binding will apply to share `!Sync` capture streams across Python threads.
|
|
154
186
|
- Crate-level rustdoc additions in `lib.rs`: a section on ORT error construction FFI side effects (the `ortsys![CreateStatus]` dylib-load trigger that motivates the `OrtPathInvalid` split from `OrtLoadFailed`) and a section on fork safety (guidance for Python `multiprocessing` consumers to use `spawn` start method).
|
|
155
187
|
- `lib.rs` rustdoc "Feature flags" section cross-references `docs/features.md` for consumers wanting the deep-dive reference.
|
|
156
188
|
- Em-dash cleanup across 19 code locations in `lib.rs`, `capture.rs`, `output.rs`, `vad.rs`, `error.rs`, `vad_integration.rs`, and `vad_ort_load_failure.rs`. Per CLAUDE.md, the codebase forbids em dashes; these were pre-existing violations.
|
package/MIGRATION.md
CHANGED
|
@@ -5,6 +5,76 @@ vocabulary that matches the Rust and Python packages, and tidies several option
|
|
|
5
5
|
and return shapes. This guide lists every breaking change with before and after
|
|
6
6
|
code.
|
|
7
7
|
|
|
8
|
+
## Breaking in 5.0.0: the vad config object
|
|
9
|
+
|
|
10
|
+
decibri 5.0.0 replaces the flat `vadThreshold` and `vadHoldoff` options with a
|
|
11
|
+
single `vad` config object that carries the model selector and its policy. The
|
|
12
|
+
`vad: 'silero'` and `vad: 'energy'` shorthand is unchanged and still uses the
|
|
13
|
+
default threshold (0.5 for `'silero'`, 0.01 for `'energy'`) and a 300 ms
|
|
14
|
+
holdoff, so only code that tuned the threshold or holdoff needs to migrate.
|
|
15
|
+
Move those values onto the object.
|
|
16
|
+
|
|
17
|
+
Before:
|
|
18
|
+
|
|
19
|
+
```js
|
|
20
|
+
new Microphone({ vad: 'silero', vadThreshold: 0.6, vadHoldoff: 200 });
|
|
21
|
+
new Microphone({ vad: 'energy', vadThreshold: 0.02 });
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
After:
|
|
25
|
+
|
|
26
|
+
```js
|
|
27
|
+
new Microphone({ vad: { model: 'silero', threshold: 0.6, holdoffMs: 200 } });
|
|
28
|
+
new Microphone({ vad: { model: 'energy', threshold: 0.02 } });
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
`vadThreshold` and `vadHoldoff` are no longer accepted; passing either throws a
|
|
32
|
+
`TypeError`. The shorthand without tuning is unchanged:
|
|
33
|
+
|
|
34
|
+
```js
|
|
35
|
+
new Microphone({ vad: 'silero' }); // unchanged
|
|
36
|
+
new Microphone({ vad: false }); // unchanged; the default
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
The same realignment applies to the browser build, where the object's `model`
|
|
40
|
+
is `'energy'` (the only browser detector): `vad: { model: 'energy', threshold: 0.02, holdoffMs: 200 }`.
|
|
41
|
+
|
|
42
|
+
## Breaking in 5.0.0: mono-only capture
|
|
43
|
+
|
|
44
|
+
decibri 4.x delivered interleaved multichannel audio when a microphone was
|
|
45
|
+
opened with `channels` greater than `1`. decibri 5.0.0 narrows capture to mono:
|
|
46
|
+
the `channels` option accepts only `1` (the default), and a value greater than
|
|
47
|
+
`1` throws a `RangeError` instead of being captured. If your code opened a
|
|
48
|
+
microphone with more than one channel, pass `channels: 1` (or omit it).
|
|
49
|
+
|
|
50
|
+
Before:
|
|
51
|
+
|
|
52
|
+
```js
|
|
53
|
+
new Microphone({ channels: 2 }); // 4.x: interleaved stereo capture
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
After:
|
|
57
|
+
|
|
58
|
+
```js
|
|
59
|
+
new Microphone({ channels: 1 }); // 5.0: mono (or omit channels entirely)
|
|
60
|
+
new Microphone(); // unchanged; mono is the default
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
A value greater than `1` now throws:
|
|
64
|
+
|
|
65
|
+
```js
|
|
66
|
+
new Microphone({ channels: 2 });
|
|
67
|
+
// RangeError: multichannel capture is not supported; channels must be 1 (mono)
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
The `channels` option and the channel-general `'data'` chunk shape are kept, so
|
|
71
|
+
multichannel may return later as an additive change (a future release accepting
|
|
72
|
+
a value greater than `1` by delivering true interleaved multichannel) rather
|
|
73
|
+
than a further break. The intended longer-term multichannel direction is array
|
|
74
|
+
ingest (consuming several channels internally for processing such as
|
|
75
|
+
beamforming or array noise reduction while still delivering one conditioned
|
|
76
|
+
stream), which is distinct from raw multichannel delivery.
|
|
77
|
+
|
|
8
78
|
## New in 4.2.0 (additive, nothing to migrate)
|
|
9
79
|
|
|
10
80
|
decibri 4.2.0 is a browser-only, additive release. Code written for 4.1.0 keeps
|
package/README.md
CHANGED
|
@@ -82,18 +82,21 @@ Creates a Readable stream that captures from the microphone.
|
|
|
82
82
|
| Option | Type | Default | Description |
|
|
83
83
|
| --- | --- | --- | --- |
|
|
84
84
|
| `sampleRate` | number | 16000 | Samples per second (1000 to 384000) |
|
|
85
|
-
| `channels` | number | 1 |
|
|
85
|
+
| `channels` | number | 1 | Mono only: the only accepted value is `1`; a value greater than `1` throws a `RangeError` |
|
|
86
86
|
| `framesPerBuffer` | number | 1600 | Frames per chunk (64 to 65536). At 16kHz mono, 1600 = 100ms = 3200 bytes |
|
|
87
87
|
| `device` | number, string, or `{ id: string }` | system default | Device index, case-insensitive name substring, or stable per-host ID |
|
|
88
88
|
| `dtype` | `'int16'` \| `'float32'` | `'int16'` | Sample encoding |
|
|
89
|
-
| `vad` | `false` \| `'silero'` \| `'energy'` | `false` | Voice activity detection: disabled, the Silero ML model,
|
|
90
|
-
| `vadThreshold` | number | 0.5 / 0.01 | Speech threshold. Default is 0.5 for `'silero'`, 0.01 for `'energy'` |
|
|
91
|
-
| `vadHoldoff` | number | 300 | Silence holdoff in ms |
|
|
89
|
+
| `vad` | `false` \| `'silero'` \| `'energy'` \| `VadOptions` | `false` | Voice activity detection: disabled, the Silero ML model, an RMS energy threshold, or a config object `{ model, threshold, holdoffMs }` to tune the policy |
|
|
92
90
|
| `modelPath` | string | bundled model | Path to the Silero model. Only used when `vad` is `'silero'` |
|
|
91
|
+
| `dcRemoval` | boolean | off | Remove a constant (DC) offset with a one-pole DC-blocking high-pass. Runs first in the chain; same-length, no added latency |
|
|
92
|
+
| `denoise` | `'fastenhancer-t'` | off | Single-channel speech-enhancement (denoise) model. The model ships in the package; no path or download needed |
|
|
93
|
+
| `highpass` | `80` \| `100` | off | High-pass cutoff in Hz (second-order Butterworth) that removes low-frequency rumble. Runs after denoise. Out-of-set values throw a `RangeError` |
|
|
94
|
+
| `agc` | number | off | AGC target level in dBFS, an integer in -40 to -3 (typical -18). Runs after the high-pass. Out-of-range throws a `RangeError` |
|
|
95
|
+
| `limiter` | number | off | Peak limiter ceiling in dBFS, a number in -3.0 to 0.0 (typical -1.0). Runs last. Out-of-range throws a `RangeError` |
|
|
93
96
|
|
|
94
97
|
Standard `ReadableOptions` (e.g. `highWaterMark`) are also accepted.
|
|
95
98
|
|
|
96
|
-
`vad: true` is not accepted; pass the mode explicitly as `vad: 'silero'` or `vad: 'energy'`.
|
|
99
|
+
`vad: true` is not accepted; pass the mode explicitly as `vad: 'silero'` or `vad: 'energy'`. The string shorthand uses the default threshold (0.5 for `'silero'`, 0.01 for `'energy'`) and a 300 ms holdoff; to tune them pass a config object: `vad: { model: 'silero', threshold: 0.6, holdoffMs: 200 }`. The `threshold` (`0`–`1`) and `holdoffMs` fields are optional and fall back to those defaults.
|
|
97
100
|
|
|
98
101
|
### Methods
|
|
99
102
|
|
|
@@ -120,7 +123,7 @@ The module-level `inputDevices()` and `version()` free functions are equivalent
|
|
|
120
123
|
| `'data'` | Buffer | Audio chunk (Int16 LE or Float32 LE) |
|
|
121
124
|
| `'backpressure'` | - | Internal buffer full, consumer too slow |
|
|
122
125
|
| `'speech'` | - | VAD: audio crosses threshold |
|
|
123
|
-
| `'silence'` | - | VAD: audio below threshold for
|
|
126
|
+
| `'silence'` | - | VAD: audio below threshold for the holdoff period |
|
|
124
127
|
| `'end'` | - | Stream ended |
|
|
125
128
|
| `'error'` | Error | An error occurred |
|
|
126
129
|
|
|
@@ -267,7 +270,7 @@ To verify playback in real browsers, open `examples/browser-speaker-test.html` (
|
|
|
267
270
|
Lightweight RMS energy threshold. No model required.
|
|
268
271
|
|
|
269
272
|
```javascript
|
|
270
|
-
const mic = new Microphone({ vad: 'energy',
|
|
273
|
+
const mic = new Microphone({ vad: { model: 'energy', threshold: 0.01 } });
|
|
271
274
|
mic.on('speech', () => console.log('speaking'));
|
|
272
275
|
mic.on('silence', () => console.log('silent'));
|
|
273
276
|
```
|
|
@@ -277,13 +280,47 @@ mic.on('silence', () => console.log('silent'));
|
|
|
277
280
|
ML-based detection using the Silero VAD v5 model. More accurate than energy mode, especially in noisy environments.
|
|
278
281
|
|
|
279
282
|
```javascript
|
|
280
|
-
const mic = new Microphone({ vad: 'silero',
|
|
283
|
+
const mic = new Microphone({ vad: { model: 'silero', threshold: 0.5 } });
|
|
281
284
|
mic.on('speech', () => console.log('speaking'));
|
|
282
285
|
mic.on('silence', () => console.log('silent'));
|
|
283
286
|
```
|
|
284
287
|
|
|
285
288
|
The Silero model (~2MB) ships inside the npm package. No downloads or API keys required. Silero mode is Node.js only.
|
|
286
289
|
|
|
290
|
+
## decibri ACE (audio conditioning)
|
|
291
|
+
|
|
292
|
+
decibri ACE (Audio Capture Engine) is decibri's opt-in audio front-end for speech. It is a conditioning chain that runs on the captured audio before it reaches your `'data'` handler. Every stage is off by default, runs on-device, and needs no API key. With nothing enabled the capture path is byte-identical to plain capture.
|
|
293
|
+
|
|
294
|
+
The stages run in a fixed order. Enable any subset:
|
|
295
|
+
|
|
296
|
+
| Stage | Option | Range |
|
|
297
|
+
| --- | --- | --- |
|
|
298
|
+
| DC removal | `dcRemoval: true` | boolean |
|
|
299
|
+
| Denoise | `denoise: 'fastenhancer-t'` | the one bundled model |
|
|
300
|
+
| High-pass | `highpass: 80` or `100` | Hz |
|
|
301
|
+
| AGC | `agc: -18` | dBFS, -40 to -3 |
|
|
302
|
+
| Limiter | `limiter: -1.0` | dBFS, -3.0 to 0.0 |
|
|
303
|
+
|
|
304
|
+
```javascript
|
|
305
|
+
const { Microphone } = require('decibri');
|
|
306
|
+
|
|
307
|
+
const mic = new Microphone({
|
|
308
|
+
sampleRate: 16000,
|
|
309
|
+
denoise: 'fastenhancer-t', // bundled speech-enhancement model
|
|
310
|
+
highpass: 80, // remove low-frequency rumble
|
|
311
|
+
agc: -18, // target level in dBFS
|
|
312
|
+
limiter: -1.0, // peak ceiling in dBFS
|
|
313
|
+
vad: { model: 'silero', threshold: 0.5 },
|
|
314
|
+
});
|
|
315
|
+
|
|
316
|
+
mic.on('data', (chunk) => { /* Buffer of conditioned Int16 PCM */ });
|
|
317
|
+
mic.on('speech', () => console.log('speech'));
|
|
318
|
+
mic.on('silence', () => console.log('silence'));
|
|
319
|
+
setTimeout(() => mic.stop(), 5000);
|
|
320
|
+
```
|
|
321
|
+
|
|
322
|
+
VAD reads the signal before the chain, so `vadScore` and the `'speech'` / `'silence'` events are unaffected by which conditioning stages you enable. The conditioning chain runs in the native Node.js capture path; the browser build does not include it.
|
|
323
|
+
|
|
287
324
|
## Device Selection
|
|
288
325
|
|
|
289
326
|
```javascript
|
package/examples/README.md
CHANGED
|
@@ -71,9 +71,10 @@ npx localtunnel --port 8080
|
|
|
71
71
|
### Regenerating the browser bundle
|
|
72
72
|
|
|
73
73
|
`decibri.browser.js` is generated from the package's browser entry
|
|
74
|
-
(`../src/browser/index.js`)
|
|
75
|
-
|
|
74
|
+
(`../src/browser/index.js`) with rolldown, the bundler used to produce the
|
|
75
|
+
checked-in build. Regenerate it after changing the browser source so the
|
|
76
|
+
bundle stays in sync:
|
|
76
77
|
|
|
77
78
|
```bash
|
|
78
|
-
npx
|
|
79
|
+
npx rolldown ../src/browser/index.js --format iife --name decibri --file decibri.browser.js
|
|
79
80
|
```
|
|
@@ -70,7 +70,7 @@ var decibri = (function() {
|
|
|
70
70
|
var require_decibri_browser = /* @__PURE__ */ __commonJSMin(((exports, module) => {
|
|
71
71
|
const { Emitter } = require_emitter();
|
|
72
72
|
const { WORKLET_SOURCE } = require_worklet_inline();
|
|
73
|
-
const VERSION = "
|
|
73
|
+
const VERSION = "5.0.0";
|
|
74
74
|
/**
|
|
75
75
|
* Browser microphone capture.
|
|
76
76
|
*
|
|
@@ -97,13 +97,27 @@ var decibri = (function() {
|
|
|
97
97
|
this._started = false;
|
|
98
98
|
this._starting = null;
|
|
99
99
|
this._stopRequested = false;
|
|
100
|
+
if (options.vadThreshold !== void 0 || options.vadHoldoff !== void 0) throw new TypeError("vadThreshold and vadHoldoff are no longer supported. Pass them on the vad config object: vad: { model: 'energy', threshold: 0.01, holdoffMs: 300 }.");
|
|
100
101
|
const vad = options.vad ?? false;
|
|
102
|
+
let vadThreshold = .01;
|
|
103
|
+
let vadHoldoff = 300;
|
|
101
104
|
if (vad === false) this._vad = false;
|
|
102
105
|
else if (vad === true) throw new TypeError("vad: true is no longer supported. Specify the mode explicitly: vad: 'energy'.");
|
|
103
106
|
else if (vad === "energy") this._vad = true;
|
|
104
|
-
else
|
|
105
|
-
|
|
106
|
-
|
|
107
|
+
else if (vad !== null && typeof vad === "object" && !Array.isArray(vad)) {
|
|
108
|
+
if (vad.model !== "energy") throw new TypeError(`Invalid vad model: ${JSON.stringify(vad.model)}. Expected 'energy'.`);
|
|
109
|
+
this._vad = true;
|
|
110
|
+
if (vad.threshold !== void 0) {
|
|
111
|
+
if (vad.threshold < 0 || vad.threshold > 1) throw new TypeError(`threshold must be between 0 and 1, got ${vad.threshold}`);
|
|
112
|
+
vadThreshold = vad.threshold;
|
|
113
|
+
}
|
|
114
|
+
if (vad.holdoffMs !== void 0) {
|
|
115
|
+
if (vad.holdoffMs < 0) throw new TypeError(`holdoffMs must be >= 0, got ${vad.holdoffMs}`);
|
|
116
|
+
vadHoldoff = vad.holdoffMs;
|
|
117
|
+
}
|
|
118
|
+
} else throw new TypeError(`Invalid vad value: ${JSON.stringify(vad)}. Expected false, 'energy', or a config object { model, threshold, holdoffMs }.`);
|
|
119
|
+
this._vadThreshold = vadThreshold;
|
|
120
|
+
this._vadHoldoff = vadHoldoff;
|
|
107
121
|
this._vadScore = 0;
|
|
108
122
|
this._isSpeaking = false;
|
|
109
123
|
this._silenceTimer = null;
|
|
@@ -119,8 +133,6 @@ var decibri = (function() {
|
|
|
119
133
|
if (this._channels < 1 || this._channels > 32) throw new TypeError(`channels must be between 1 and 32, got ${this._channels}`);
|
|
120
134
|
if (this._framesPerBuffer < 64 || this._framesPerBuffer > 65536) throw new TypeError(`frames per buffer must be between 64 and 65536, got ${this._framesPerBuffer}`);
|
|
121
135
|
if (this._dtype !== "int16" && this._dtype !== "float32") throw new TypeError("dtype must be 'int16' or 'float32'");
|
|
122
|
-
if (this._vadThreshold < 0 || this._vadThreshold > 1) throw new TypeError(`vadThreshold must be between 0 and 1, got ${this._vadThreshold}`);
|
|
123
|
-
if (this._vadHoldoff < 0) throw new TypeError(`vadHoldoff must be >= 0, got ${this._vadHoldoff}`);
|
|
124
136
|
}
|
|
125
137
|
/**
|
|
126
138
|
* Start microphone capture.
|
package/index.d.ts
CHANGED
|
@@ -22,6 +22,13 @@ export declare class DecibriBridge {
|
|
|
22
22
|
* Returns 0.0 if VAD is not active.
|
|
23
23
|
*/
|
|
24
24
|
get vadProbability(): number
|
|
25
|
+
/**
|
|
26
|
+
* Number of capture buffers dropped because the consumer could not keep
|
|
27
|
+
* pace (the core stream's overrun counter). Returns 0 while the consumer
|
|
28
|
+
* keeps up or when no stream is active. Returned as f64 (an exact JS
|
|
29
|
+
* number for any realistic count) to match the `vadProbability` getter.
|
|
30
|
+
*/
|
|
31
|
+
get overrunCount(): number
|
|
25
32
|
/** List all available audio input devices. */
|
|
26
33
|
static devices(): Array<DeviceInfoJs>
|
|
27
34
|
/** Version information. */
|
|
@@ -104,6 +111,40 @@ export interface DecibriOptions {
|
|
|
104
111
|
device?: any
|
|
105
112
|
vadMode?: string
|
|
106
113
|
modelPath?: string
|
|
114
|
+
/**
|
|
115
|
+
* Capture DC-removal toggle. When `true`, removes a constant (DC) offset
|
|
116
|
+
* from the captured audio with a one-pole DC-blocking high-pass, applied
|
|
117
|
+
* first in the transform chain (before denoise). Absent or `false` leaves
|
|
118
|
+
* it off (the default), a byte-identical no-op. Pure DSP: no bundled file
|
|
119
|
+
* and no model path, like `highpass`.
|
|
120
|
+
*/
|
|
121
|
+
dcRemoval?: boolean
|
|
122
|
+
/**
|
|
123
|
+
* Capture denoise model selector. The only accepted value is
|
|
124
|
+
* `'fastenhancer-t'`; absent leaves denoise off. The JS wrapper resolves
|
|
125
|
+
* the bundled model file and passes its path through `denoise_model_path`.
|
|
126
|
+
*/
|
|
127
|
+
denoise?: string
|
|
128
|
+
/**
|
|
129
|
+
* Capture high-pass filter cutoff in Hz. The accepted values are `80` (an
|
|
130
|
+
* 80 Hz second-order Butterworth high-pass) and `100` (a 100 Hz one);
|
|
131
|
+
* absent leaves the high-pass off. Pure DSP: no bundled file and no model
|
|
132
|
+
* path, unlike `denoise`.
|
|
133
|
+
*/
|
|
134
|
+
highpass?: number
|
|
135
|
+
/**
|
|
136
|
+
* Capture AGC target level in dBFS: an integer in `-40..=-3` (typical -18);
|
|
137
|
+
* absent leaves AGC off. Drives the captured level toward the target. Pure
|
|
138
|
+
* DSP: no bundled file and no model path, like `highpass`.
|
|
139
|
+
*/
|
|
140
|
+
agc?: number
|
|
141
|
+
/**
|
|
142
|
+
* Capture limiter ceiling in dBFS (sample-peak): a number in `-3.0..=0.0`
|
|
143
|
+
* (typical -1.0); absent leaves the limiter off. Holds the captured signal
|
|
144
|
+
* at or below the ceiling, catching a peak the AGC would let through. Pure
|
|
145
|
+
* DSP: no bundled file and no model path, like `agc`.
|
|
146
|
+
*/
|
|
147
|
+
limiter?: number
|
|
107
148
|
}
|
|
108
149
|
|
|
109
150
|
/** Options passed from JS constructor for output. */
|