decibri 4.4.2 → 5.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -9,7 +9,37 @@ For other decibri packages, see:
9
9
  - Rust crate: [crates/decibri/CHANGELOG.md](../../crates/decibri/CHANGELOG.md)
10
10
  - Python wheel: [bindings/python/CHANGELOG.md](../../bindings/python/CHANGELOG.md)
11
11
 
12
- ## [Unreleased]
12
+ ## [5.1.0] - 2026-07-17
13
+
14
+ ### Added
15
+
16
+ - `File`, an offline source: the same conditioning options as `Microphone` over a WAV file (`new File(path)` synchronously, `await File.open(path)` off the event loop) or a `Float32Array` of samples (`File.buffer(samples, { inputRate })`; a raw `Buffer` of bytes is rejected as ambiguous), delivered as a finite Readable stream of conditioned chunks.
17
+ - Whole-file speech analysis: `file.analyze()` (also spelled `file.analyse()`) resolves to a `VadReport` with per-window `scores` (`{ start, end, vadScore, isSpeech }`) and merged speech `segments` (`{ start, end }`), all in seconds of file time. Requires `vad: 'silero'`; a `File` opened without `vad` rejects with `analysis requires VAD`.
18
+ - Per-chunk VAD on files: `vadScore` and the `speech` / `silence` events alongside the stream, with the speaking holdoff measured in file time (sample positions) rather than wall-clock time, so processing speed never changes the reported events.
19
+
20
+ ## [5.0.0] - 2026-06-24
21
+
22
+ ### Added
23
+
24
+ - A device or driver failure during streaming now surfaces as a `DecibriError` with the dedicated `code` `'DEVICE_FAILED'`, and a non-ORT ONNX backend failure as `'ONNX_BACKEND_FAILED'`, instead of the generic `'DECIBRI_ERROR'`. Both are catchable by branching on `err.code`; the message text is unchanged.
25
+ - `dcRemoval`: an opt-in `Microphone` option that removes a constant (DC) offset from the captured audio with a one-pole DC-blocking high-pass. Set `true` to enable it; omit it or set `false` to leave it off (the default), which keeps the capture path byte-identical. Pure DSP: no bundled file or download is needed. It runs first in the chain, before denoise, and is same-length with no added latency, so `vadScore` and the `speech` / `silence` events are unaffected.
26
+ - `highpass`: an opt-in `Microphone` option that applies a high-pass filter to the captured audio, removing low-frequency rumble below the voice band. The accepted values are the cutoff in Hz, `80` (an 80 Hz second-order Butterworth high-pass) or `100` (a 100 Hz one); omit it to leave the high-pass off (the default), which keeps the capture path full-range. Pure DSP: no bundled file or download is needed. It runs after denoise in the chain and adds no latency, so `vadScore` and the `speech` / `silence` events are unaffected. An out-of-set value throws a `RangeError`. The closed cutoff set is designed to grow (further cutoffs are additive).
27
+ - `agc`: an opt-in `Microphone` option that applies automatic gain control to the captured audio, driving the running level toward a target with a smoothed, rate-limited gain. The value is a target level in dBFS (an integer in -40 to -3, typical -18); omit it to leave AGC off (the default), which keeps the level untouched. Pure DSP: no bundled file or download is needed. It runs after the high-pass step and adds no latency, so `vadScore` and the `speech` / `silence` events are unaffected. The opening is delivered at its natural level and reaches the target within tens of milliseconds (no opening window of wrong gain). An out-of-range value throws a `RangeError`.
28
+ - `limiter`: an opt-in `Microphone` option that applies a peak limiter to the captured audio, holding the signal at or below a sample-peak ceiling, the safety net that catches a transient the AGC's gain would let through. The value is a ceiling in dBFS (a number in -3.0 to 0.0, typical -1.0); omit it to leave the limiter off (the default), which keeps the level untouched. Pure DSP: no bundled file or download is needed. It runs last in the chain, after the AGC step, and adds no latency, so `vadScore` and the `speech` / `silence` events are unaffected. No output sample ever exceeds the ceiling, even on an instantaneous transient. An out-of-range value throws a `RangeError`.
29
+ - `denoise`: an opt-in `Microphone` option that runs a bundled single-channel speech-enhancement model over the captured audio. The only accepted value is `'fastenhancer-t'`; omit it to leave denoise off (the default), which keeps the capture path unchanged. The model ships with the package, so no path or download is needed. The delivered `'data'` chunks carry the enhanced audio; VAD reads the pre-enhancement signal, so `vadScore` and the `speech` / `silence` events are unaffected. An unknown model name throws a `TypeError`. A denoise model-load failure raises a dedicated `MODEL_LOAD_FAILED` `OrtError` via `wrapNativeError`; note that a failure surfaced from `start()` is delivered on the `'error'` event as the raw native error without the decibri `code` attached (as with device-open failures), so match it by message there.
30
+
31
+ ### Changed
32
+
33
+ - A microphone whose native sample rate differs from the configured `sampleRate` is now resampled to the requested rate in the engine, so capture delivers audio at exactly the configured rate on every device. The device is opened at its native rate and decibri resamples (anti-aliased polyphase) to the target; a device already at the requested rate is unchanged (no resample). Chunk size and the reported rate are unchanged (still `framesPerBuffer * outputChannels * bytesPerSample` at the configured rate). Previously the requested rate was handed to the platform audio backend, which delivered it only when the device or OS could. Breaking for consumers whose device's native rate differs from the configured rate: the delivered samples are now decibri-resampled.
34
+ - Microphone capture is now mono only, narrowing a capability that shipped in 4.x. decibri 4.x delivered interleaved multichannel audio when a microphone was opened with `channels` greater than `1`; 5.0 captures mono, and the `channels` option accepts only `1` (the default). A value greater than `1` now throws a `RangeError` (`multichannel capture is not supported; channels must be 1 (mono)`), a clear error rather than a silent substitution of mono. The `channels` option is retained and the emitted `'data'` chunks keep their channel-general shape (`framesPerBuffer * bytesPerSample`, mono today), so multichannel can return later as an additive change rather than a further break: a future release may accept a value greater than `1` by delivering true interleaved multichannel. The intended longer-term multichannel direction is array ingest (consuming several channels internally for processing such as beamforming or array noise reduction while still delivering one conditioned stream), which is distinct from raw multichannel delivery. The default single-channel capture is unchanged. Breaking for consumers that opened a microphone with more than one channel: pass `channels: 1` or omit it.
35
+ - Capture now emits a final, possibly-shorter `'data'` chunk at stream close (on `stop()`), carrying the buffered tail, before `'end'`. Steady-state chunks are unchanged (still exactly `framesPerBuffer * channels * bytesPerSample`); only the last chunk before `'end'` may be shorter, and no captured audio is dropped. Internally the binding now re-blocks through the Rust core instead of a binding-side accumulator, and `stop()` defers the stream's end-of-stream signal by one tick so the flushed tail is delivered before `'end'` rather than dropped.
36
+ - The `vad` option now accepts a config object `{ model, threshold, holdoffMs }` (a `VadOptions`) alongside the `'silero'` / `'energy'` shorthand, and the flat `vadThreshold` / `vadHoldoff` options are removed. The shorthand is unchanged and keeps the default threshold (0.5 for `'silero'`, 0.01 for `'energy'`) and 300 ms holdoff; only code that tuned the threshold or holdoff migrates, by passing them on the object (`vad: { model: 'silero', threshold: 0.6, holdoffMs: 200 }`). An out-of-range threshold or a negative holdoff throws a `RangeError`, an unknown model a `TypeError`. Passing the removed `vadThreshold` or `vadHoldoff` throws a `TypeError`. The same object form applies to the browser build, where `model` is `'energy'`. Breaking for consumers that set `vadThreshold` or `vadHoldoff`; detection, the `vadScore` getter, and the `speech` / `silence` events are unchanged.
37
+
38
+ ### Fixed
39
+
40
+ - Energy-mode VAD (`vad: 'energy'`) now reads the signal before the opt-in capture enhancement, so enabling an enhancement step (`denoise`, `highpass`, `agc`, `limiter`) no longer changes `vadScore` or the `speech` / `silence` events, matching the guarantee Silero mode already had. The energy RMS is computed natively on the pre-enhancement signal; previously it was computed on the delivered (post-enhancement) audio, so an enhancement step shifted the score (`agc` most of all, since it drives the delivered level toward its target, which could defeat energy endpointing). Energy detection with no enhancement enabled is unchanged.
41
+ - Resampled capture (a device whose native rate differs from the configured `sampleRate`) no longer drops the resampler's group-delay tail at stream close: the final, possibly-shorter `'data'` chunk(s) before `'end'` now carry it, so the complete resampled signal is delivered. A device already at the requested rate (no resample) is unchanged.
42
+ - On macOS, the microphone-permission error message now reads "System Settings > Privacy & Security" (the modern macOS wording) instead of the pre-Ventura "System Preferences > Security & Privacy".
13
43
 
14
44
  ## [4.4.2] - 2026-06-12
15
45
 
@@ -137,7 +167,7 @@ Error message wording on shipped `DecibriError` variants has historically been s
137
167
 
138
168
  ## [3.3.0] - 2026-04-23
139
169
 
140
- Groundwork release for upcoming P3 Python bindings. Adds a stable-ID form for audio device selection (`DeviceSelector::Id`), fixes a long-standing direction bug in `DecibriError::DeviceNotFound` when resolving output devices, exposes both in the Node binding, and extends the reference documentation with a Cargo feature flag guide plus additional crate-level rustdoc. No Node.js or browser API break. Direct Rust crate consumers pattern-matching on `DeviceSelector` or struct-literal-constructing `DeviceInfo` / `OutputDeviceInfo` need to update for the new `#[non_exhaustive]` attributes (see Migration notes below).
170
+ Groundwork release for upcoming Python bindings. Adds a stable-ID form for audio device selection (`DeviceSelector::Id`), fixes a long-standing direction bug in `DecibriError::DeviceNotFound` when resolving output devices, exposes both in the Node binding, and extends the reference documentation with a Cargo feature flag guide plus additional crate-level rustdoc. No Node.js or browser API break. Direct Rust crate consumers pattern-matching on `DeviceSelector` or struct-literal-constructing `DeviceInfo` / `OutputDeviceInfo` need to update for the new `#[non_exhaustive]` attributes (see Migration notes below).
141
171
 
142
172
  ### Changed
143
173
 
@@ -160,7 +190,7 @@ Groundwork release for upcoming P3 Python bindings. Adds a stable-ID form for au
160
190
  ### Internal
161
191
 
162
192
  - `DeviceDirection` trait gains a `not_found_error(String) -> DecibriError` method so `resolve_device_generic`'s `Name` and `Id` arms produce direction-correct errors via the `Input` / `Output` impls.
163
- - Unit tests for `Arc<Mutex<CaptureStream>>` confirming the wrapping is `Send + Sync` (compile-time assertion) and serializes concurrent access across two threads (runtime test with `Barrier`). Documents the wrapping strategy the P3 Python binding will apply to share `!Sync` capture streams across Python threads.
193
+ - Unit tests for `Arc<Mutex<CaptureStream>>` confirming the wrapping is `Send + Sync` (compile-time assertion) and serializes concurrent access across two threads (runtime test with `Barrier`). Documents the wrapping strategy the Python binding will apply to share `!Sync` capture streams across Python threads.
164
194
  - Crate-level rustdoc additions in `lib.rs`: a section on ORT error construction FFI side effects (the `ortsys![CreateStatus]` dylib-load trigger that motivates the `OrtPathInvalid` split from `OrtLoadFailed`) and a section on fork safety (guidance for Python `multiprocessing` consumers to use `spawn` start method).
165
195
  - `lib.rs` rustdoc "Feature flags" section cross-references `docs/features.md` for consumers wanting the deep-dive reference.
166
196
  - Em-dash cleanup across 19 code locations in `lib.rs`, `capture.rs`, `output.rs`, `vad.rs`, `error.rs`, `vad_integration.rs`, and `vad_ort_load_failure.rs`. Per CLAUDE.md, the codebase forbids em dashes; these were pre-existing violations.
package/MIGRATION.md CHANGED
@@ -5,6 +5,76 @@ vocabulary that matches the Rust and Python packages, and tidies several option
5
5
  and return shapes. This guide lists every breaking change with before and after
6
6
  code.
7
7
 
8
+ ## Breaking in 5.0.0: the vad config object
9
+
10
+ decibri 5.0.0 replaces the flat `vadThreshold` and `vadHoldoff` options with a
11
+ single `vad` config object that carries the model selector and its policy. The
12
+ `vad: 'silero'` and `vad: 'energy'` shorthand is unchanged and still uses the
13
+ default threshold (0.5 for `'silero'`, 0.01 for `'energy'`) and a 300 ms
14
+ holdoff, so only code that tuned the threshold or holdoff needs to migrate.
15
+ Move those values onto the object.
16
+
17
+ Before:
18
+
19
+ ```js
20
+ new Microphone({ vad: 'silero', vadThreshold: 0.6, vadHoldoff: 200 });
21
+ new Microphone({ vad: 'energy', vadThreshold: 0.02 });
22
+ ```
23
+
24
+ After:
25
+
26
+ ```js
27
+ new Microphone({ vad: { model: 'silero', threshold: 0.6, holdoffMs: 200 } });
28
+ new Microphone({ vad: { model: 'energy', threshold: 0.02 } });
29
+ ```
30
+
31
+ `vadThreshold` and `vadHoldoff` are no longer accepted; passing either throws a
32
+ `TypeError`. The shorthand without tuning is unchanged:
33
+
34
+ ```js
35
+ new Microphone({ vad: 'silero' }); // unchanged
36
+ new Microphone({ vad: false }); // unchanged; the default
37
+ ```
38
+
39
+ The same realignment applies to the browser build, where the object's `model`
40
+ is `'energy'` (the only browser detector): `vad: { model: 'energy', threshold: 0.02, holdoffMs: 200 }`.
41
+
42
+ ## Breaking in 5.0.0: mono-only capture
43
+
44
+ decibri 4.x delivered interleaved multichannel audio when a microphone was
45
+ opened with `channels` greater than `1`. decibri 5.0.0 narrows capture to mono:
46
+ the `channels` option accepts only `1` (the default), and a value greater than
47
+ `1` throws a `RangeError` instead of being captured. If your code opened a
48
+ microphone with more than one channel, pass `channels: 1` (or omit it).
49
+
50
+ Before:
51
+
52
+ ```js
53
+ new Microphone({ channels: 2 }); // 4.x: interleaved stereo capture
54
+ ```
55
+
56
+ After:
57
+
58
+ ```js
59
+ new Microphone({ channels: 1 }); // 5.0: mono (or omit channels entirely)
60
+ new Microphone(); // unchanged; mono is the default
61
+ ```
62
+
63
+ A value greater than `1` now throws:
64
+
65
+ ```js
66
+ new Microphone({ channels: 2 });
67
+ // RangeError: multichannel capture is not supported; channels must be 1 (mono)
68
+ ```
69
+
70
+ The `channels` option and the channel-general `'data'` chunk shape are kept, so
71
+ multichannel may return later as an additive change (a future release accepting
72
+ a value greater than `1` by delivering true interleaved multichannel) rather
73
+ than a further break. The intended longer-term multichannel direction is array
74
+ ingest (consuming several channels internally for processing such as
75
+ beamforming or array noise reduction while still delivering one conditioned
76
+ stream), which is distinct from raw multichannel delivery.
77
+
8
78
  ## New in 4.2.0 (additive, nothing to migrate)
9
79
 
10
80
  decibri 4.2.0 is a browser-only, additive release. Code written for 4.1.0 keeps
package/README.md CHANGED
@@ -30,6 +30,22 @@ mic.on('data', (chunk) => { /* Buffer of Int16 PCM samples */ });
30
30
  setTimeout(() => mic.stop(), 5000);
31
31
  ```
32
32
 
33
+ ### Condition and analyze a file
34
+
35
+ ```javascript
36
+ const { File } = require('decibri');
37
+
38
+ // The same conditioning chain as the live microphone, over a WAV file.
39
+ const file = await File.open('clip.wav', { denoise: 'fastenhancer-t', highpass: 80 });
40
+ file.on('data', (chunk) => { /* Buffer of conditioned Int16 PCM */ });
41
+ file.on('end', () => console.log('done'));
42
+
43
+ // Whole-file speech analysis (a live stream cannot do this).
44
+ const f = await File.open('clip.wav', { vad: 'silero' });
45
+ const report = await f.analyze();
46
+ for (const s of report.segments) console.log(s.start, s.end); // seconds
47
+ ```
48
+
33
49
  ### Play audio
34
50
 
35
51
  ```javascript
@@ -82,18 +98,21 @@ Creates a Readable stream that captures from the microphone.
82
98
  | Option | Type | Default | Description |
83
99
  | --- | --- | --- | --- |
84
100
  | `sampleRate` | number | 16000 | Samples per second (1000 to 384000) |
85
- | `channels` | number | 1 | Input channels (1 to 32) |
101
+ | `channels` | number | 1 | Mono only: the only accepted value is `1`; a value greater than `1` throws a `RangeError` |
86
102
  | `framesPerBuffer` | number | 1600 | Frames per chunk (64 to 65536). At 16kHz mono, 1600 = 100ms = 3200 bytes |
87
103
  | `device` | number, string, or `{ id: string }` | system default | Device index, case-insensitive name substring, or stable per-host ID |
88
104
  | `dtype` | `'int16'` \| `'float32'` | `'int16'` | Sample encoding |
89
- | `vad` | `false` \| `'silero'` \| `'energy'` | `false` | Voice activity detection: disabled, the Silero ML model, or an RMS energy threshold |
90
- | `vadThreshold` | number | 0.5 / 0.01 | Speech threshold. Default is 0.5 for `'silero'`, 0.01 for `'energy'` |
91
- | `vadHoldoff` | number | 300 | Silence holdoff in ms |
105
+ | `vad` | `false` \| `'silero'` \| `'energy'` \| `VadOptions` | `false` | Voice activity detection: disabled, the Silero ML model, an RMS energy threshold, or a config object `{ model, threshold, holdoffMs }` to tune the policy |
92
106
  | `modelPath` | string | bundled model | Path to the Silero model. Only used when `vad` is `'silero'` |
107
+ | `dcRemoval` | boolean | off | Remove a constant (DC) offset with a one-pole DC-blocking high-pass. Runs first in the chain; same-length, no added latency |
108
+ | `denoise` | `'fastenhancer-t'` | off | Single-channel speech-enhancement (denoise) model. The model ships in the package; no path or download needed |
109
+ | `highpass` | `80` \| `100` | off | High-pass cutoff in Hz (second-order Butterworth) that removes low-frequency rumble. Runs after denoise. Out-of-set values throw a `RangeError` |
110
+ | `agc` | number | off | AGC target level in dBFS, an integer in -40 to -3 (typical -18). Runs after the high-pass. Out-of-range throws a `RangeError` |
111
+ | `limiter` | number | off | Peak limiter ceiling in dBFS, a number in -3.0 to 0.0 (typical -1.0). Runs last. Out-of-range throws a `RangeError` |
93
112
 
94
113
  Standard `ReadableOptions` (e.g. `highWaterMark`) are also accepted.
95
114
 
96
- `vad: true` is not accepted; pass the mode explicitly as `vad: 'silero'` or `vad: 'energy'`.
115
+ `vad: true` is not accepted; pass the mode explicitly as `vad: 'silero'` or `vad: 'energy'`. The string shorthand uses the default threshold (0.5 for `'silero'`, 0.01 for `'energy'`) and a 300 ms holdoff; to tune them pass a config object: `vad: { model: 'silero', threshold: 0.6, holdoffMs: 200 }`. The `threshold` (`0`–`1`) and `holdoffMs` fields are optional and fall back to those defaults.
97
116
 
98
117
  ### Methods
99
118
 
@@ -120,7 +139,7 @@ The module-level `inputDevices()` and `version()` free functions are equivalent
120
139
  | `'data'` | Buffer | Audio chunk (Int16 LE or Float32 LE) |
121
140
  | `'backpressure'` | - | Internal buffer full, consumer too slow |
122
141
  | `'speech'` | - | VAD: audio crosses threshold |
123
- | `'silence'` | - | VAD: audio below threshold for `vadHoldoff` ms |
142
+ | `'silence'` | - | VAD: audio below threshold for the holdoff period |
124
143
  | `'end'` | - | Stream ended |
125
144
  | `'error'` | Error | An error occurred |
126
145
 
@@ -267,7 +286,7 @@ To verify playback in real browsers, open `examples/browser-speaker-test.html` (
267
286
  Lightweight RMS energy threshold. No model required.
268
287
 
269
288
  ```javascript
270
- const mic = new Microphone({ vad: 'energy', vadThreshold: 0.01 });
289
+ const mic = new Microphone({ vad: { model: 'energy', threshold: 0.01 } });
271
290
  mic.on('speech', () => console.log('speaking'));
272
291
  mic.on('silence', () => console.log('silent'));
273
292
  ```
@@ -277,13 +296,59 @@ mic.on('silence', () => console.log('silent'));
277
296
  ML-based detection using the Silero VAD v5 model. More accurate than energy mode, especially in noisy environments.
278
297
 
279
298
  ```javascript
280
- const mic = new Microphone({ vad: 'silero', vadThreshold: 0.5 });
299
+ const mic = new Microphone({ vad: { model: 'silero', threshold: 0.5 } });
281
300
  mic.on('speech', () => console.log('speaking'));
282
301
  mic.on('silence', () => console.log('silent'));
283
302
  ```
284
303
 
285
304
  The Silero model (~2MB) ships inside the npm package. No downloads or API keys required. Silero mode is Node.js only.
286
305
 
306
+ ## decibri ACE (audio conditioning)
307
+
308
+ decibri ACE (Audio Capture Engine) is decibri's opt-in audio front-end for speech. It is a conditioning chain that runs on the captured audio before it reaches your `'data'` handler. Every stage is off by default, runs on-device, and needs no API key. With nothing enabled the capture path is byte-identical to plain capture.
309
+
310
+ The stages run in a fixed order. Enable any subset:
311
+
312
+ | Stage | Option | Range |
313
+ | --- | --- | --- |
314
+ | DC removal | `dcRemoval: true` | boolean |
315
+ | Denoise | `denoise: 'fastenhancer-t'` | the one bundled model |
316
+ | High-pass | `highpass: 80` or `100` | Hz |
317
+ | AGC | `agc: -18` | dBFS, -40 to -3 |
318
+ | Limiter | `limiter: -1.0` | dBFS, -3.0 to 0.0 |
319
+
320
+ ```javascript
321
+ const { Microphone } = require('decibri');
322
+
323
+ const mic = new Microphone({
324
+ sampleRate: 16000,
325
+ denoise: 'fastenhancer-t', // bundled speech-enhancement model
326
+ highpass: 80, // remove low-frequency rumble
327
+ agc: -18, // target level in dBFS
328
+ limiter: -1.0, // peak ceiling in dBFS
329
+ vad: { model: 'silero', threshold: 0.5 },
330
+ });
331
+
332
+ mic.on('data', (chunk) => { /* Buffer of conditioned Int16 PCM */ });
333
+ mic.on('speech', () => console.log('speech'));
334
+ mic.on('silence', () => console.log('silence'));
335
+ setTimeout(() => mic.stop(), 5000);
336
+ ```
337
+
338
+ VAD reads the signal before the chain, so `vadScore` and the `'speech'` / `'silence'` events are unaffected by which conditioning stages you enable. The conditioning chain runs in the native Node.js capture path; the browser build does not include it.
339
+
340
+ ## API: File (offline source)
341
+
342
+ Everything a `Microphone` does to live audio, `File` does to audio you already have: the same conditioning options, the same Readable stream of conditioned chunks (finite: it ends at EOF), and the same opt-in `vad`. Because a `File` is a complete recording, it can also analyze the whole recording for speech.
343
+
344
+ - `await File.open(path, options?)`: read a WAV off the event loop (recommended, like `Microphone.open`).
345
+ - `new File(path, options?)`: the same result, synchronous (blocks on disk I/O; fine for scripts).
346
+ - `File.buffer(samples, options)`: wrap a `Float32Array` of samples you already hold. `options.inputRate` is required (raw samples carry no header); a raw `Buffer` of bytes is rejected as ambiguous.
347
+ - `await file.analyze()` (also spelled `analyse()`): consume the source and resolve to a `VadReport` of per-window `scores` (`{ start, end, vadScore, isSpeech }`) and merged speech `segments` (`{ start, end }`), in seconds of file time. Requires `vad: 'silero'`; a `File` opened without `vad` rejects with `analysis requires VAD`.
348
+ - `file.vadScore`, `'speech'` / `'silence'` events: per-chunk VAD alongside the stream, with the holdoff measured in FILE time (sample positions), never wall-clock time, so processing speed does not change the reported events.
349
+
350
+ Options mirror `Microphone` (`sampleRate`, `dtype`, `vad`, `dcRemoval`, `denoise`, `highpass`, `agc`, `limiter`); the live-capture options (`device`, `channels`, `framesPerBuffer`) do not apply. Iteration and analysis are separate single passes: construct one `File` per operation. Note: Node also has a global `File` (the web File API); import decibri's explicitly to avoid shadowing surprises.
351
+
287
352
  ## Device Selection
288
353
 
289
354
  ```javascript
@@ -71,9 +71,10 @@ npx localtunnel --port 8080
71
71
  ### Regenerating the browser bundle
72
72
 
73
73
  `decibri.browser.js` is generated from the package's browser entry
74
- (`../src/browser/index.js`). Regenerate it after changing the browser source
75
- with any bundler that resolves the package's `browser` entry, for example:
74
+ (`../src/browser/index.js`) with rolldown, the bundler used to produce the
75
+ checked-in build. Regenerate it after changing the browser source so the
76
+ bundle stays in sync:
76
77
 
77
78
  ```bash
78
- npx esbuild ../src/browser/index.js --bundle --format=iife --global-name=decibri --outfile=decibri.browser.js
79
+ npx rolldown ../src/browser/index.js --format iife --name decibri --file decibri.browser.js
79
80
  ```
@@ -70,7 +70,7 @@ var decibri = (function() {
70
70
  var require_decibri_browser = /* @__PURE__ */ __commonJSMin(((exports, module) => {
71
71
  const { Emitter } = require_emitter();
72
72
  const { WORKLET_SOURCE } = require_worklet_inline();
73
- const VERSION = "4.2.0";
73
+ const VERSION = "5.1.0";
74
74
  /**
75
75
  * Browser microphone capture.
76
76
  *
@@ -97,13 +97,27 @@ var decibri = (function() {
97
97
  this._started = false;
98
98
  this._starting = null;
99
99
  this._stopRequested = false;
100
+ if (options.vadThreshold !== void 0 || options.vadHoldoff !== void 0) throw new TypeError("vadThreshold and vadHoldoff are no longer supported. Pass them on the vad config object: vad: { model: 'energy', threshold: 0.01, holdoffMs: 300 }.");
100
101
  const vad = options.vad ?? false;
102
+ let vadThreshold = .01;
103
+ let vadHoldoff = 300;
101
104
  if (vad === false) this._vad = false;
102
105
  else if (vad === true) throw new TypeError("vad: true is no longer supported. Specify the mode explicitly: vad: 'energy'.");
103
106
  else if (vad === "energy") this._vad = true;
104
- else throw new TypeError(`Invalid vad value: ${JSON.stringify(vad)}. Expected false or 'energy'.`);
105
- this._vadThreshold = options.vadThreshold ?? .01;
106
- this._vadHoldoff = options.vadHoldoff ?? 300;
107
+ else if (vad !== null && typeof vad === "object" && !Array.isArray(vad)) {
108
+ if (vad.model !== "energy") throw new TypeError(`Invalid vad model: ${JSON.stringify(vad.model)}. Expected 'energy'.`);
109
+ this._vad = true;
110
+ if (vad.threshold !== void 0) {
111
+ if (vad.threshold < 0 || vad.threshold > 1) throw new TypeError(`threshold must be between 0 and 1, got ${vad.threshold}`);
112
+ vadThreshold = vad.threshold;
113
+ }
114
+ if (vad.holdoffMs !== void 0) {
115
+ if (vad.holdoffMs < 0) throw new TypeError(`holdoffMs must be >= 0, got ${vad.holdoffMs}`);
116
+ vadHoldoff = vad.holdoffMs;
117
+ }
118
+ } else throw new TypeError(`Invalid vad value: ${JSON.stringify(vad)}. Expected false, 'energy', or a config object { model, threshold, holdoffMs }.`);
119
+ this._vadThreshold = vadThreshold;
120
+ this._vadHoldoff = vadHoldoff;
107
121
  this._vadScore = 0;
108
122
  this._isSpeaking = false;
109
123
  this._silenceTimer = null;
@@ -119,8 +133,6 @@ var decibri = (function() {
119
133
  if (this._channels < 1 || this._channels > 32) throw new TypeError(`channels must be between 1 and 32, got ${this._channels}`);
120
134
  if (this._framesPerBuffer < 64 || this._framesPerBuffer > 65536) throw new TypeError(`frames per buffer must be between 64 and 65536, got ${this._framesPerBuffer}`);
121
135
  if (this._dtype !== "int16" && this._dtype !== "float32") throw new TypeError("dtype must be 'int16' or 'float32'");
122
- if (this._vadThreshold < 0 || this._vadThreshold > 1) throw new TypeError(`vadThreshold must be between 0 and 1, got ${this._vadThreshold}`);
123
- if (this._vadHoldoff < 0) throw new TypeError(`vadHoldoff must be >= 0, got ${this._vadHoldoff}`);
124
136
  }
125
137
  /**
126
138
  * Start microphone capture.
package/index.d.ts CHANGED
@@ -78,6 +78,52 @@ export declare class DecibriOutputBridge {
78
78
  static version(): VersionInfoJs
79
79
  }
80
80
 
81
+ /**
82
+ * Native offline-source handle exposed to Node.js via napi-rs. The public
83
+ * `File` Readable lives in the JS wrapper; consumers construct that, not
84
+ * this handle, directly.
85
+ */
86
+ export declare class FileHandle {
87
+ /**
88
+ * Open a WAV path as an offline source, synchronously (blocks on disk
89
+ * I/O; the JS wrapper's async `File.open` uses `openAsync` instead).
90
+ */
91
+ static open(path: string, options?: FileOptions | undefined | null): FileHandle
92
+ /**
93
+ * Open a WAV path without blocking the JS event loop: the disk read,
94
+ * WAV parse, and chain construction run on the libuv thread pool.
95
+ */
96
+ static openAsync(path: string, options?: FileOptions | undefined | null): Promise<unknown>
97
+ /**
98
+ * Wrap in-memory samples as an offline source. `samples` are mono f32 in
99
+ * [-1.0, 1.0]; `inputRate` is their native rate (raw samples carry no
100
+ * header). No I/O, so construction is synchronous.
101
+ */
102
+ static buffer(samples: Float32Array, inputRate: number, options?: FileOptions | undefined | null): FileHandle
103
+ /**
104
+ * Pull the next conditioned chunk, advancing the per-chunk VAD score on
105
+ * the pre-conditioning feed. Returns `null` once the source is fully
106
+ * delivered (after the end-of-stream tail) or already consumed.
107
+ */
108
+ readChunk(): Buffer | null
109
+ /**
110
+ * Consume the source with the core's whole-recording analysis, off the
111
+ * JS event loop. Resolves to the `VadReport`; a `File` built without VAD
112
+ * rejects with the core's typed error, never a silently constructed
113
+ * detector.
114
+ */
115
+ analyze(): Promise<unknown>
116
+ /** Release the source. Idempotent; a closed File reads as ended. */
117
+ close(): void
118
+ /**
119
+ * Most recent per-chunk VAD score (0.0 to 1.0), computed on the
120
+ * pre-conditioning feed. 0.0 before the first chunk or with VAD off.
121
+ */
122
+ get vadProbability(): number
123
+ /** The target output rate every delivered chunk carries. */
124
+ get sampleRate(): number
125
+ }
126
+
81
127
  /**
82
128
  * Options passed from JS constructor.
83
129
  *
@@ -111,6 +157,40 @@ export interface DecibriOptions {
111
157
  device?: any
112
158
  vadMode?: string
113
159
  modelPath?: string
160
+ /**
161
+ * Capture DC-removal toggle. When `true`, removes a constant (DC) offset
162
+ * from the captured audio with a one-pole DC-blocking high-pass, applied
163
+ * first in the transform chain (before denoise). Absent or `false` leaves
164
+ * it off (the default), a byte-identical no-op. Pure DSP: no bundled file
165
+ * and no model path, like `highpass`.
166
+ */
167
+ dcRemoval?: boolean
168
+ /**
169
+ * Capture denoise model selector. The only accepted value is
170
+ * `'fastenhancer-t'`; absent leaves denoise off. The JS wrapper resolves
171
+ * the bundled model file and passes its path through `denoise_model_path`.
172
+ */
173
+ denoise?: string
174
+ /**
175
+ * Capture high-pass filter cutoff in Hz. The accepted values are `80` (an
176
+ * 80 Hz second-order Butterworth high-pass) and `100` (a 100 Hz one);
177
+ * absent leaves the high-pass off. Pure DSP: no bundled file and no model
178
+ * path, unlike `denoise`.
179
+ */
180
+ highpass?: number
181
+ /**
182
+ * Capture AGC target level in dBFS: an integer in `-40..=-3` (typical -18);
183
+ * absent leaves AGC off. Drives the captured level toward the target. Pure
184
+ * DSP: no bundled file and no model path, like `highpass`.
185
+ */
186
+ agc?: number
187
+ /**
188
+ * Capture limiter ceiling in dBFS (sample-peak): a number in `-3.0..=0.0`
189
+ * (typical -1.0); absent leaves the limiter off. Holds the captured signal
190
+ * at or below the ceiling, catching a peak the AGC would let through. Pure
191
+ * DSP: no bundled file and no model path, like `agc`.
192
+ */
193
+ limiter?: number
114
194
  }
115
195
 
116
196
  /** Options passed from JS constructor for output. */
@@ -136,6 +216,26 @@ export interface DeviceInfoJs {
136
216
  isDefault: boolean
137
217
  }
138
218
 
219
+ /**
220
+ * Options passed from the JS `File` wrapper. The conditioning fields mirror
221
+ * `DecibriOptions` exactly; the live-capture-only fields (device, channels,
222
+ * framesPerBuffer) do not apply to an offline source. `vadThreshold` and
223
+ * `vadHoldoffMs` are internal plumbing (the user passes them on the `vad`
224
+ * config object; the wrapper resolves them), hidden from the generated
225
+ * TypeScript like `ortLibraryPath`.
226
+ */
227
+ export interface FileOptions {
228
+ sampleRate?: number
229
+ format?: string
230
+ vadMode?: string
231
+ modelPath?: string
232
+ dcRemoval?: boolean
233
+ denoise?: string
234
+ highpass?: number
235
+ agc?: number
236
+ limiter?: number
237
+ }
238
+
139
239
  /** Output device info returned to JS. */
140
240
  export interface OutputDeviceInfoJs {
141
241
  index: number
@@ -150,6 +250,32 @@ export interface OutputDeviceInfoJs {
150
250
  isDefault: boolean
151
251
  }
152
252
 
253
+ /** One merged speech region of a recording, in seconds of file time. */
254
+ export interface Segment {
255
+ start: number
256
+ end: number
257
+ }
258
+
259
+ /**
260
+ * The whole-recording analysis `File.analyze()` resolves to: per-window
261
+ * scores and merged speech segments, in file order.
262
+ */
263
+ export interface VadReport {
264
+ scores: Array<VadWindow>
265
+ segments: Array<Segment>
266
+ }
267
+
268
+ /**
269
+ * One scored voice-activity window of a recording: `start` / `end` in
270
+ * seconds of file time, the speech probability, and the raw threshold test.
271
+ */
272
+ export interface VadWindow {
273
+ start: number
274
+ end: number
275
+ vadScore: number
276
+ isSpeech: boolean
277
+ }
278
+
153
279
  /** Version info returned to JS. */
154
280
  export interface VersionInfoJs {
155
281
  decibri: string