decibri 5.4.0 → 5.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -9,6 +9,71 @@ For other decibri packages, see:
9
9
  - Rust core: [crates/decibri/CHANGELOG.md](../../crates/decibri/CHANGELOG.md)
10
10
  - Python package: [bindings/python/CHANGELOG.md](../../bindings/python/CHANGELOG.md)
11
11
 
12
+ ## [5.6.0] - 2026-08-16
13
+
14
+ ### Breaking changes
15
+
16
+ - **`Microphone` opens the device at its native channel count, and decibri performs the collapse to the delivered mono.** The device was opened at the configured count of 1, so on a multichannel device the operating system collapsed the channels before decibri saw them, and it did so differently on each platform: the Windows audio engine applies an unpublished mixing matrix, macOS's default channel map selects device channel zero rather than mixing, and on Linux the result depends on which PCM the device name resolved to (a `plug` device routes under a named policy, a raw `hw` device refuses the open). The capture stream now receives every device channel and decibri averages them into the delivered audio, the same documented average `File` applies to a multichannel container. The delivered audio changes for any device whose native count is above one: on macOS it changes from device channel zero to the average of every channel; on Windows the engine's stereo mix is replaced by the average, which is close to it but not guaranteed identical; above two channels every platform's mix differs from a plain average. A device already delivering one channel is unaffected, byte for byte. The `channels` option now names the DELIVERED count, exactly as `sampleRate` names the delivered rate.
17
+ - **The browser `Microphone` receives every channel the browser grants, and decibri performs the collapse to the delivered mono.** The capture requested a single channel from `getUserMedia` and read the first channel of the granted track, so on a track granted with more than one channel the delivered audio was the first channel alone. The capture now requests the 32-channel count the Web Audio specification requires an implementation to support, with ideal semantics, reads the granted track's own channel count, and averages every granted channel into the single delivered channel, the same documented average the Node entry delivers. The delivered audio changes for any track granted with more than one channel: the first granted channel is replaced by the average of every granted channel. A track granted with a single channel is unchanged, byte for byte.
18
+ - **`Microphone` accepts a `channels` count above 1, and a request for one is honoured rather than refused.** The constructor threw `RangeError: multichannel capture is not supported; channels must be 1 (mono)` for every count above 1, so the request failed before the device was touched. It now constructs, opens, and emits chunks carrying that many interleaved channels. Code that used the throw as a guard, catching it to fall back to a single-channel request, no longer has one: it receives interleaved multichannel audio where it received an error, and a consumer that reads a `'data'` chunk as a run of mono samples reads interleaved frames instead. Pass `channels: 1` to keep the previous delivery, which is unchanged byte for byte.
19
+ - **The browser `Microphone` accepts a `channels` count above 1, and delivers that many channels.** It threw `RangeError: multichannel capture is not supported; channels must be 1 (mono)` for every count above 1. It now delivers the requested channels interleaved, from the granted track, exactly as the Node entry does, so the two entries accept and refuse the same inputs with the same classes and messages. Code that relied on the throw loses the same guard the Node entry's callers lose. `channels: 1` is unchanged, byte for byte, and so is every other browser `Microphone` option.
20
+ - **The `Invalid vad value:` `TypeError` names `source` among the config object's keys.** The message reads `Invalid vad value: <value>. Expected false, 'silero', 'energy', or a config object { model, threshold, holdoffMs, source }.`, and the browser entry's copy names the same key list with its own accepted set of `false`, `'energy'`, and the config object. Code matching the full message text exactly stops matching and has to be updated to the new text; code matching the `Invalid vad value:` prefix is unaffected. The class, the inputs that throw it, the accepted `vad` values, and every other message are unchanged.
21
+ - **`pushAecReference` refuses a typed array whose sample dtype is not the microphone's own.** Any `ArrayBufferView` was read as raw bytes in the configured `dtype`, so a `Float32Array` pushed on an `'int16'` capture fed the canceller a byte reinterpretation of its samples, with nothing said. A `Float32Array` on an `'int16'` capture and an `Int16Array` on a `'float32'` capture now throw a `TypeError` naming both dtypes and both remedies (`dtype 'int16' configured but Float32Array samples were pushed; convert to Int16Array or construct Microphone with dtype: 'float32'`, and the mirror for the other direction), and every other sample-dtype view (`Float64Array`, `Int32Array`, and the rest) throws the existing `pushAecReference requires a Buffer, TypedArray, or DataView of PCM samples in the configured dtype` message, whatever the capture state, the refusal the Python surface raises for the same condition. A call site that relied on byte reinterpretation passes its bytes as a `Buffer`, `Uint8Array`, or `DataView` instead, which remain format-agnostic byte carriers, read exactly as before. `Int16Array` on `'int16'` and `Float32Array` on `'float32'` remain accepted unchanged.
22
+ - **`aecMetrics().referenceDropped` counts reference samples pushed while capture is not running.** A push made between construction (or `await Microphone.open()`) and the first consumer engaging the stream, or after `stop()`, was discarded with no counter moved, so the reference for audio played at startup vanished with nothing to show for it. Those samples now add to `referenceDropped`, readable once capture runs. A `referenceDropped` of zero no longer certifies only that every push fit the queue's bound; it also certifies that none arrived while capture was down, so code alerting on a nonzero figure now fires for pre-start pushes as well. The push itself still never blocks and never throws on any capture state, a push with `aec` unset is still an uncounted no-op, `aecMetrics()` still reads `null` while capture is not running, and the Python surface's counters are unchanged.
23
+ - **`AudioWriter` accepts a `channels` count above 1, and writes that many interleaved channels.** The constructor threw `RangeError: multichannel write is not supported; channels must be 1 (mono)` for every count above 1; that message is no longer thrown, and code that relied on the throw as a guard no longer has one. The incoming bytes are read as interleaved frames at the declared count and the file's header carries it; the stream's total sample count must divide into whole frames, refused when the stream finishes otherwise. Each container's own channel ceiling applies at the write, reported as a `DecibriError` with code `AUDIO_FORMAT_UNSUPPORTED` carrying the container layer's own text: a FLAC frame carries at most 8 channels, and a WAV format chunk's `nBlockAlign` field holds at most 32767 channels at 16-bit samples. decibri enforces no ceiling of its own. `channels: 1` (still the default) writes the identical file it wrote before, byte for byte, and `channels: 0` throws `RangeError: channels must be at least 1`. `sampleRate`, `dtype`, `format`, `compression` and the `report` are unchanged.
24
+
25
+ ### Added
26
+
27
+ - `channelMap` on the `Microphone` options: an optional array of 0-based device channel indices selecting which device channels feed the delivered channels, so `channelMap: [1]` delivers the device's second channel alone. The length must equal `channels`, and entries may repeat and may appear in any order, so a map both selects and permutes, and may name more delivered channels than the device has; a length that does not match, or a malformed array, throws a `RangeError` or `TypeError` from the constructor. Absent delivers the average of every opened channel, the previous behaviour. The same shape as CoreAudio AUHAL's channel map (an array of device channel indices, one entry per client channel), not miniaudio's `channelMap`, which names a spatial layout. Entries are checked against the resolved device's own report when the stream starts; the device's report is the only ceiling, and no fixed maximum exists.
28
+ - Error code `CHANNEL_MAP_OUT_OF_RANGE` on `DecibriError`, emitted on the `'error'` event when the channel map names a device channel the device does not have. The message names the offending entry and the count the device reports, the figure `MicrophoneInfo.maxInputChannels` carries.
29
+ - `channelMap` on the browser `Microphone` options, mirroring the Node entry: an optional array of 0-based device channel indices selecting which granted channels feed the delivered channels, so `channelMap: [1]` delivers the granted track's second channel alone. The length must equal `channels`, and entries may repeat and may appear in any order, so a map both selects and permutes; a length that does not match, or a malformed array, throws the same `RangeError` or `TypeError` as the Node entry, with the same message, from the constructor. Absent delivers the average of every granted channel. Entries are checked against the granted track's own report: where the browser reports the granted channel count, `start()` rejects with an `Error` whose message names the entry and the granted count, the same message the Node entry carries for the same condition; where it does not, the same `Error` is emitted on the `'error'` event and the capture stops as soon as the audio graph reports its true channel count. The granted report is the only ceiling, and no fixed maximum exists.
30
+ - Multichannel capture on the `Microphone`: `channels` names how many interleaved channels each `'data'` chunk carries, bounded below by the constructor and above by the resolved device alone, with no fixed maximum. With a `channelMap`, delivered channel `j` carries the device channel the map names. Without one, two derivations are accepted: 1 delivers the average of every device channel, and the device's own count delivers every device channel in device order. A chunk's samples are interleaved frames of the delivered width, so a chunk holds `framesPerBuffer * channels` samples.
31
+ - Multichannel capture on the browser `Microphone`, the same rules read against the granted track instead of a device: a `channelMap` names granted channels, an unmapped count of 1 averages every granted channel, and an unmapped count equal to the grant delivers every granted channel in granted order. Above one delivered channel `vadScore` is the RMS of the per-frame average of the delivered channels rather than of the samples as they lie, which is what the Node entry's detector reads for the same capture; at one delivered channel it is unchanged.
32
+ - Error code `MICROPHONE_CHANNELS_UNSUPPORTED` on `DecibriError`, emitted on the `'error'` event when the delivered count exceeds the device's own. The message names both figures, the second being the one `MicrophoneInfo.maxInputChannels` carries. The browser entry rejects `start()` with the same message.
33
+ - Error code `CHANNEL_SELECTION_AMBIGUOUS` on `DecibriError`, emitted on the `'error'` event when the delivered count is above 1 and below the device's own with no `channelMap` set. The message names both figures. Set a `channelMap` naming which channels to deliver. The browser entry rejects `start()` with the same message.
34
+ - Echo cancellation on every delivered channel: with `aec` set and `channels` above 1, one canceller engine runs per delivered channel, each fed the same pushed reference and each finding its own channel's echo delay. `pushAecReference` is unchanged: one push serves every channel. The processing and memory cost scale with the delivered count, and a capture that outruns the machine shows up as `overrunCount` climbing; no capacity ceiling exists. A single-channel capture with the canceller is unchanged byte for byte.
35
+ - `channels` on the object `aecMetrics()` returns: every delivered channel's canceller report in delivered order, one entry per channel, each carrying `delaySamples`, `erleDb`, `doubleTalk`, `referenceStarved`, `acquisitionParked` and `referenceReanchors`. The top-level fields keep their names and report the first delivered channel's engine (the queue counters `referenceDropped` and `referenceSilence` describe the shared queue), so a single-channel stream reads as before with a one-entry array alongside. `delaySamples` is each engine's alignment offset from the reference frontier, not a room measurement, and `erleDb` is not a quality ranking across channels: it rises with echo distance, so a far microphone reports a higher figure while removing less echo in absolute terms.
36
+ - `channels` and `channelMap` on the `File` options, the `Microphone`'s channel vocabulary on the offline source, with the source's own channel count (the file's header, or `File.buffer`'s `inputChannels`) standing where the device's report stands. `channels: 1` (the default) delivers the documented average of every source channel, the previous behaviour, byte for byte; a count equal to the source's own delivers every source channel in source order, interleaved frame by frame in the emitted chunks. `channelMap` is an array of 0-based source channel indices, one entry per delivered channel; entries may repeat and may appear in any order, so a map both selects and permutes, and may name more delivered channels than the source has. Checked at construction, where the source's count is known; the source's count is the only ceiling, and no fixed maximum exists.
37
+ - Error codes `FILE_CHANNELS_UNSUPPORTED`, `FILE_CHANNEL_SELECTION_AMBIGUOUS` and `FILE_CHANNEL_MAP_OUT_OF_RANGE` on `DecibriError`, thrown from the `File` constructors: an unmapped delivered count above the source's own, an unmapped count above 1 and below the source's own (set a `channelMap` naming which source channels to deliver), and a map entry the source does not have. Each message names the figures in question. A map length that differs from `channels` keeps the shared `RangeError`.
38
+ - `inputChannels` on the `File.buffer` options, the channel counterpart of `inputRate`: the interleave of the caller's own samples, 1 (mono) by default, up to 65535. A sample count that is not a whole number of frames at the declared count throws the same `RangeError` the live block-size refusal carries. The option applies to `File.buffer` alone and throws a `TypeError` on the open path, where the file's own header answers.
39
+ - `File.save` writes the delivered channel count: the saved file carries the same interleaved layout the stream emits, in every container, with each container's own ceiling reported exactly as the `AudioWriter` entry above describes.
40
+ - `source` on the `vad` config object, on the `Microphone` and `File` options alike: the 0-based DELIVERED channel the detector reads, the position within the delivered interleaved frames after any `channelMap` is applied (a `channelMap` names device or source channels; `source` names the delivered position, so a channel a map delivers at two positions is named by position). Absent feeds the detector the frame average of every delivered channel, the previous behaviour. The value must be an integer (`TypeError: vad source must be an integer` otherwise), in 0 to 65535 (`RangeError: vad source must be between 0 and 65535` otherwise), and below the delivered channel count, refused from the constructor with a `RangeError` whose message is `the detector source names delivered channel <i>; the delivered channel count is <n>`; the delivered count is the only ceiling, and no fixed maximum exists. Selecting a source changes which samples the detector reads and nothing else: the emitted audio, `threshold`, `holdoffMs`, and the `'speech'` / `'silence'` mechanics are untouched.
41
+ - `source` on the browser `Microphone`'s `vad` config object, mirroring the Node entry: `vadScore` becomes the RMS of the named delivered channel's samples rather than of the per-frame average, and the same values are refused with the same classes and messages from the constructor.
42
+
43
+ ### Changed
44
+
45
+ - The capture queue's fixed 64-buffer bound now holds buffers at the device's native channel width, so its memory scales with the device's channel count: roughly 123 KB per channel at a 48 kHz native rate with a typical 10 millisecond driver period.
46
+
47
+ ### Removed
48
+
49
+ - The internal classification row for the message `multichannel capture is not supported; channels must be 1 (mono)`, which nothing produces. The message classified to a built-in `RangeError` with no `code`, so no thrown class, code, message or option changes.
50
+
51
+ ## [5.5.0] - 2026-08-10
52
+
53
+ ### Breaking changes
54
+
55
+ - **The `RangeError` thrown for a `channels` count below 1 carries the message `channels must be at least 1`.** It carried `channels must be between 1 and 32`, naming an upper bound that neither `Microphone` nor `Speaker` enforces. Code that matches the message text exactly stops matching and has to be updated to the new text. Nothing else about the error moves: it is still a `RangeError`, thrown for the same input, `channels: 0` on either constructor or `open()`.
56
+ - **`Speaker` accepts a `channels` count above 32.** It threw `RangeError: channels must be between 1 and 32` from the constructor. It now accepts any count above zero, offers the count to the device when playback starts, and a device that cannot serve it emits a `DecibriError` with code `SPEAKER_CHANNELS_UNSUPPORTED` on the `'error'` event, or rejects `writeAsync()` with it. Code that relied on the constructor throwing to reject an unsupported output channel count has to handle the failure asynchronously instead, on `'error'` or the rejection. `new Speaker({ channels: 0 })` still throws a `RangeError` from the constructor, and a count above 16383 still carries `STREAM_OPEN_FAILED` naming that limit. How many channels a device serves depends on the `sampleRate` asked for as well as the count.
57
+ - **An output channel count above the device's reported figure carries `SPEAKER_CHANNELS_UNSUPPORTED`.** It carried `STREAM_OPEN_FAILED`. Code branching on `STREAM_OPEN_FAILED` to detect an unsupported output channel count has to match `SPEAKER_CHANNELS_UNSUPPORTED` as well, and cannot drop `STREAM_OPEN_FAILED`, because a device that reports no figure at all keeps `STREAM_OPEN_FAILED` for the same condition. `STREAM_OPEN_FAILED` is otherwise unchanged and remains the code for an output open that failed for any other reason.
58
+ - **The browser `Microphone` rejects a `channels` count above 1.** It accepted 1 to 32 and delivered the first channel only, with nothing to indicate the rest were dropped. A count above 1 now throws `RangeError: multichannel capture is not supported; channels must be 1 (mono)`, and a count below 1 throws `RangeError: channels must be at least 1` where it threw a `TypeError`. Code passing a count above 1 to the browser entry has to pass 1 and mix down its own sources, or read the single channel it was already receiving. `channels: 1` is unchanged, as is every other browser `Microphone` option. Both class and message now match the Node entry's for the same values, so the two entries reject the same input the same way.
59
+ - The browser `Speaker` is unchanged. It still accepts 1 to 32 channels, the range the Web Audio specification requires an implementation to support, and keeps its own `channels must be between 1 and 32, got <n>` message, which that range does enforce.
60
+
61
+ ### Added
62
+
63
+ - Error code `SPEAKER_CHANNELS_UNSUPPORTED` on `DecibriError`, for an output device that cannot serve the requested `channels`. The message names the count asked for, the count the device reports (the figure `SpeakerInfo.maxOutputChannels` carries) and the platform's own message.
64
+ - `referenceChannels` on the `aec` option object: the channel count of the far-end reference pushed through `pushAecReference`. Default 1 (mono). With a count above 1 the pushed samples are read as interleaved frames and each frame is averaged to one mono sample before the canceller sees it. The collapse is opt-in: a caller pushing a multichannel reference must declare the count, and an undeclared multichannel push keeps its current behaviour, cancelling nothing and reporting no error. The declared count must match the pushed buffer: a mismatch is not detected and raises no error, and shows up only as `aecMetrics().delaySamples` staying `null` with no fault reported. A count below 1 throws a `RangeError`; the only ceiling is the option's own 16-bit carrier. A mono reference against playback through more than one loudspeaker has a cancellation ceiling: the canceller models one room response applied to the channel average, so a placement where the per-loudspeaker echo paths differ leaves a residual that adaptation does not remove.
65
+ - Each of the four platform packages (`@decibri/decibri-win32-x64-msvc`, `@decibri/decibri-darwin-arm64`, `@decibri/decibri-linux-x64-gnu`, `@decibri/decibri-linux-arm64-gnu`) now includes `THIRD-PARTY-NOTICES.md`, the third-party license notices for the ONNX Runtime dynamic libraries the package carries and for the third-party material incorporated into them, with a source-availability statement for the MPL-2.0-licensed Eigen code they contain.
66
+ - `models/THIRD-PARTY-NOTICES.md`, carrying the origin, version and license text for the two bundled ONNX models together with the training-data attribution the denoise checkpoint requires. It ships beside the weights it covers. `models/README.md` alongside it documents each model's tensor interface and points at the notice.
67
+
68
+ ### Changed
69
+
70
+ - ONNX Runtime telemetry is disabled on the environment decibri commits when it initializes the runtime, where it was left at ONNX Runtime's own default of enabled. This covers every path that reaches the runtime: `vad: 'silero'` and the ACE `denoise` stage. Set `DECIBRI_ORT_TELEMETRY=1` in the environment before first use to leave it enabled; every other value, an empty value, and an absent variable leave it disabled. Two limits apply on Windows and decibri can close neither. ONNX Runtime logs one process-information event while the environment is being created, before the setting is applied, and logs it once per process, so that event is emitted whichever way the setting is left. The runtime also assigns its telemetry state from the Windows tracing session through an ETW callback, so the platform can re-enable telemetry after decibri has disabled it. On other platforms ONNX Runtime's telemetry provider does nothing. The browser build has no ONNX Runtime and is unaffected. No option, event, method or error text changes.
71
+
72
+ ### Fixed
73
+
74
+ - The bundled Silero VAD model is documented as v6.2, the version that ships. The `model` field on `VadOptions`, the `vad` option on `MicrophoneOptions`, the README and the notice beside the model named v5. Which model file ships is unchanged.
75
+ - The Silero VAD tensor specification in `models/README.md` names the tensors the model exposes: `input`, `state` and `sr` in, `output` and `stateN` out. It described a four-input form carrying separate `h` and `c` LSTM tensors.
76
+
12
77
  ## [5.4.0] - 2026-08-04
13
78
 
14
79
  ### Added
@@ -36,7 +101,7 @@ For other decibri packages, see:
36
101
  ### Added
37
102
 
38
103
  - `aec` option on `Microphone`: acoustic echo cancellation on the capture path. The short form names the model (`aec: 'tau'`); the object form takes `{ model, tailMs, suppression, referenceSampleRate }`. It runs before the detector tap, so `vadScore` and the `speech` / `silence` events read the echo-removed signal, and it requires `sampleRate` in 8000 to 48000. Native capture only: the browser entry keeps the platform's own `echoCancellation` constraint.
39
- - `Microphone.pushAecReference(data)`, which queues the far-end audio the canceller cancels against: the same input shapes `Speaker.write` accepts, mono, in played order, at the declared `referenceSampleRate`. It never blocks and never throws on a full queue. When it is pushed does not have to match when it plays: the queue is read at the rate the capture consumes it, so a greeting pushed before the first `data` event, or a whole utterance handed over in one call, is read out over the capture it echoes into and every sample of it is cancelled against. The queue holds two seconds, which bounds how far ahead of its own capture a caller may run.
104
+ - `Microphone.pushAecReference(data)`, which queues the far-end audio the canceller cancels against: the same input shapes `Speaker.write` accepts, mono, in played order, at the declared `referenceSampleRate`. It never blocks and never throws on a full queue. Push the reference as it plays rather than ahead of it: the canceller acquires its delay from the reference already standing in front of the consumer's first read, so a push that runs far enough ahead of the capture it echoes into leaves the delay unacquired for the rest of the session, and the capture is then delivered uncancelled with no error reported and `aecMetrics().delaySamples` staying `null`. It does not recover on its own, and how large a lead is too large depends on the rest of the capture chain. The queue holds two seconds at the declared rate, and a push beyond that is dropped and counted in `aecMetrics().referenceDropped`.
40
105
  - `Microphone.aecMetrics()`, the canceller's transport and cancellation metrics merged with the reference queue's counters, or `null` while echo cancellation is off or capture is not running.
41
106
 
42
107
  Echo cancellation joins automatic gain control as a stage that can drive captured samples above full scale when the limiter is off, because it subtracts its estimate of the echo from the capture and exceeds the capture wherever that estimate is wrong in phase. The limiter runs after it and bounds the output to its ceiling; the `int16` sample format clamps, so an over-scale sample arrives as full scale rather than wrapping, and a `float32` consumer without the limiter should clamp its own output.
package/MIGRATION.md CHANGED
@@ -42,10 +42,10 @@ is `'energy'` (the only browser detector): `vad: { model: 'energy', threshold: 0
42
42
  ## Breaking in 5.0.0: mono-only capture
43
43
 
44
44
  decibri 4.x delivered interleaved multichannel audio when a microphone was
45
- opened with `channels` greater than `1`. decibri 5.0.0 narrows capture to mono:
46
- the `channels` option accepts only `1` (the default), and a value greater than
47
- `1` throws a `RangeError` instead of being captured. If your code opened a
48
- microphone with more than one channel, pass `channels: 1` (or omit it).
45
+ opened with `channels` greater than `1`. decibri 5.0.0 narrowed capture to
46
+ mono: the `channels` option accepted only `1` (the default), and a value
47
+ greater than `1` threw a `RangeError` instead of being captured. To run 4.x
48
+ code on the 5.0.0 through 5.5.0 releases, pass `channels: 1` (or omit it).
49
49
 
50
50
  Before:
51
51
 
@@ -60,20 +60,16 @@ new Microphone({ channels: 1 }); // 5.0: mono (or omit channels entirely)
60
60
  new Microphone(); // unchanged; mono is the default
61
61
  ```
62
62
 
63
- A value greater than `1` now throws:
63
+ On those releases a value greater than `1` threw:
64
64
 
65
65
  ```js
66
66
  new Microphone({ channels: 2 });
67
67
  // RangeError: multichannel capture is not supported; channels must be 1 (mono)
68
68
  ```
69
69
 
70
- The `channels` option and the channel-general `'data'` chunk shape are kept, so
71
- multichannel may return later as an additive change (a future release accepting
72
- a value greater than `1` by delivering true interleaved multichannel) rather
73
- than a further break. The intended longer-term multichannel direction is array
74
- ingest (consuming several channels internally for processing such as
75
- beamforming or array noise reduction while still delivering one conditioned
76
- stream), which is distinct from raw multichannel delivery.
70
+ That restriction has since been removed: `channels` above `1` is accepted again
71
+ and delivers interleaved multichannel audio. See
72
+ [CHANGELOG.md](./CHANGELOG.md) for the current `channels` semantics.
77
73
 
78
74
  ## New in 4.2.0 (additive, nothing to migrate)
79
75
 
package/README.md CHANGED
@@ -1,6 +1,6 @@
1
1
  # decibri
2
2
 
3
- Cross-platform audio capture, playback, and voice activity detection for Node.js and browsers.
3
+ Cross-platform audio capture, conditioning, and playback for Node.js and browsers, with voice activity detection on live and recorded audio.
4
4
 
5
5
  ## Installation
6
6
 
@@ -98,11 +98,12 @@ Creates a Readable stream that captures from the microphone.
98
98
  | Option | Type | Default | Description |
99
99
  | --- | --- | --- | --- |
100
100
  | `sampleRate` | number | 16000 | Samples per second (1000 to 384000) |
101
- | `channels` | number | 1 | Mono only: the only accepted value is `1`; a value greater than `1` throws a `RangeError` |
101
+ | `channels` | number | 1 | Delivered channels per frame, interleaved. `1` (the default) delivers the average of every device channel; the device's own count delivers every device channel in device order; a `channelMap` names any other selection. The device's own report is the only ceiling; no fixed maximum exists |
102
+ | `channelMap` | number[] | none | 0-based device channel indices choosing which device channels feed the delivered channels: delivered channel `j` carries device channel `channelMap[j]`. The length must equal `channels`; entries may repeat and may appear in any order, so a map selects, permutes, and duplicates |
102
103
  | `framesPerBuffer` | number | 1600 | Frames per chunk (64 to 65536). At 16kHz mono, 1600 = 100ms = 3200 bytes |
103
104
  | `device` | number, string, or `{ id: string }` | system default | Device index, case-insensitive name substring, or stable per-host ID |
104
105
  | `dtype` | `'int16'` \| `'float32'` | `'int16'` | Sample encoding |
105
- | `vad` | `false` \| `'silero'` \| `'energy'` \| `VadOptions` | `false` | Voice activity detection: disabled, the Silero ML model, an RMS energy threshold, or a config object `{ model, threshold, holdoffMs }` to tune the policy |
106
+ | `vad` | `false` \| `'silero'` \| `'energy'` \| `VadOptions` | `false` | Voice activity detection: disabled, the Silero ML model, an RMS energy threshold, or a config object `{ model, threshold, holdoffMs, source }` to tune the policy and the detector source |
106
107
  | `modelPath` | string | bundled model | Path to the Silero model. Only used when `vad` is `'silero'` |
107
108
  | `dcRemoval` | boolean | off | Remove a constant (DC) offset with a one-pole DC-blocking high-pass. Runs first in the chain; same-length, no added latency |
108
109
  | `denoise` | `'fastenhancer-t'` | off | Single-channel speech-enhancement (denoise) model. The model ships in the package; no path or download needed |
@@ -112,7 +113,9 @@ Creates a Readable stream that captures from the microphone.
112
113
 
113
114
  Standard `ReadableOptions` (e.g. `highWaterMark`) are also accepted.
114
115
 
115
- `vad: true` is not accepted; pass the mode explicitly as `vad: 'silero'` or `vad: 'energy'`. The string shorthand uses the default threshold (0.5 for `'silero'`, 0.01 for `'energy'`) and a 300 ms holdoff; to tune them pass a config object: `vad: { model: 'silero', threshold: 0.6, holdoffMs: 200 }`. The `threshold` (`0`–`1`) and `holdoffMs` fields are optional and fall back to those defaults.
116
+ `vad: true` is not accepted; pass the mode explicitly as `vad: 'silero'` or `vad: 'energy'`. The string shorthand uses the default threshold (0.5 for `'silero'`, 0.01 for `'energy'`) and a 300 ms holdoff; to tune them pass a config object: `vad: { model: 'silero', threshold: 0.6, holdoffMs: 200 }`. The `threshold` (`0`–`1`) and `holdoffMs` fields are optional and fall back to those defaults. The optional `source` field names the 0-based DELIVERED channel the detector reads: the position within the delivered interleaved frames after any `channelMap` is applied (a `channelMap` names device channels; `source` names a delivered position). Absent, the detector reads the frame average of every delivered channel. Selecting a source changes which samples the detector reads and nothing else; the delivered audio is untouched.
117
+
118
+ Multichannel capture delivers interleaved frames: each chunk holds `framesPerBuffer * channels` samples, and each frame carries the delivered channels in order. Without a `channelMap`, a `channels` count above `1` must equal the device's own: a count above the device's report fails with a `DecibriError` carrying the code `'MICROPHONE_CHANNELS_UNSUPPORTED'`, and a count above `1` and below the device's report fails with `'CHANNEL_SELECTION_AMBIGUOUS'`, because which channels it means has no single answer; a `channelMap` names them. A map entry the device does not have fails with `'CHANNEL_MAP_OUT_OF_RANGE'`, naming the entry and the count the device reports.
116
119
 
117
120
  ### Methods
118
121
 
@@ -152,7 +155,7 @@ Creates a Writable stream for speaker playback.
152
155
  | Option | Type | Default | Description |
153
156
  | --- | --- | --- | --- |
154
157
  | `sampleRate` | number | 16000 | Playback sample rate (1000 to 384000) |
155
- | `channels` | number | 1 | Output channels (1 to 32) |
158
+ | `channels` | number | 1 | Output channels (1 or more, up to what the device supports) |
156
159
  | `dtype` | `'int16'` \| `'float32'` | `'int16'` | Sample encoding of incoming data |
157
160
  | `device` | number, string, or `{ id: string }` | system default | Output device index, case-insensitive name substring, or stable per-host ID |
158
161
 
@@ -235,7 +238,7 @@ Same options as Node.js, plus:
235
238
  | `noiseSuppression` | boolean | true | Browser noise suppression |
236
239
  | `workletUrl` | string | inline blob | Custom worklet URL for strict CSP |
237
240
 
238
- The browser runs energy-mode VAD only, so its `vad` option accepts `false` or `'energy'`. The browser `version()` returns `{ decibri }` only: the browser build has no native core, so `decibri` reports the installed package version. In Node, `version().decibri` reports the native core version.
241
+ The browser runs energy-mode VAD only, so its `vad` option accepts `false`, `'energy'`, or a config object `{ model: 'energy', threshold, holdoffMs, source }`. The browser `version()` returns `{ decibri }` only: the browser build has no native core, so `decibri` reports the installed package version. In Node, `version().decibri` reports the native core version.
239
242
 
240
243
  ### Key differences from Node.js
241
244
 
@@ -293,7 +296,7 @@ mic.on('silence', () => console.log('silent'));
293
296
 
294
297
  ### Silero mode
295
298
 
296
- ML-based detection using the Silero VAD v5 model. More accurate than energy mode, especially in noisy environments.
299
+ ML-based detection using the Silero VAD v6.2 model. More accurate than energy mode, especially in noisy environments.
297
300
 
298
301
  ```javascript
299
302
  const mic = new Microphone({ vad: { model: 'silero', threshold: 0.5 } });
@@ -303,9 +306,9 @@ mic.on('silence', () => console.log('silent'));
303
306
 
304
307
  The Silero model (~2MB) ships inside the npm package. No downloads or API keys required. Silero mode is Node.js only.
305
308
 
306
- ## decibri ACE (audio conditioning)
309
+ ## Decibri ACE (audio conditioning)
307
310
 
308
- decibri ACE (Audio Capture Engine) is decibri's opt-in audio front-end for speech. It is a conditioning chain that runs on the captured audio before it reaches your `'data'` handler. Every stage is off by default, runs on-device, and needs no API key. With nothing enabled the capture path is byte-identical to plain capture.
311
+ Decibri ACE (Audio Conditioning Engine) is decibri's opt-in audio front-end for speech. It is a conditioning chain that runs on the captured audio before it reaches your `'data'` handler. Every stage is off by default, runs on-device, and needs no API key. With nothing enabled the capture path is byte-identical to plain capture.
309
312
 
310
313
  The stages run in a fixed order. Enable any subset:
311
314
 
@@ -337,17 +340,25 @@ setTimeout(() => mic.stop(), 5000);
337
340
 
338
341
  VAD reads the signal before the chain, so `vadScore` and the `'speech'` / `'silence'` events are unaffected by which conditioning stages you enable. The conditioning chain runs in the native Node.js capture path; the browser build does not include it.
339
342
 
343
+ ## ONNX Runtime telemetry
344
+
345
+ Silero mode (`vad: 'silero'`) and the ACE `denoise` stage run on ONNX Runtime, which carries its own telemetry, separate from anything decibri does. Decibri disables it on the environment it commits when it initializes the runtime. Set `DECIBRI_ORT_TELEMETRY=1` in the environment before first use to leave it enabled; every other value, an empty value, and an absent variable leave it disabled.
346
+
347
+ Two limits apply on Windows and decibri can close neither, so decibri does not claim that no telemetry is emitted. ONNX Runtime logs one process-information event while the environment is being created, before the setting is applied, and logs it once per process, so that event is emitted whichever way the setting is left. The runtime also assigns its telemetry state from the Windows tracing session through an ETW callback, so the platform can re-enable telemetry after decibri has disabled it. On other platforms ONNX Runtime's telemetry provider does nothing.
348
+
349
+ Neither ONNX Runtime nor this setting applies to the browser build, which has no ONNX Runtime.
350
+
340
351
  ## API: File (offline source)
341
352
 
342
353
  Everything a `Microphone` does to live audio, `File` does to audio you already have: the same conditioning options, the same Readable stream of conditioned chunks (finite: it ends at EOF), and the same opt-in `vad`. Because a `File` is a complete recording, it can also analyze the whole recording for speech.
343
354
 
344
355
  - `await File.open(path, options?)`: read a file off the event loop (recommended, like `Microphone.open`). Reads WAV, AIFF, AIFF-C and FLAC, identified from the file's own bytes rather than its extension.
345
356
  - `new File(path, options?)`: the same result, synchronous (blocks on disk I/O; fine for scripts).
346
- - `File.buffer(samples, options)`: wrap a `Float32Array` of samples you already hold. `options.inputRate` is required (raw samples carry no header); a raw `Buffer` of bytes is rejected as ambiguous.
357
+ - `File.buffer(samples, options)`: wrap a `Float32Array` of samples you already hold. `options.inputRate` is required (raw samples carry no header); a raw `Buffer` of bytes is rejected as ambiguous. `options.inputChannels` states the interleave of the samples, `1` by default, the channel counterpart of `inputRate`.
347
358
  - `await file.analyze()` (also spelled `analyse()`): consume the source and resolve to a `VadReport` of per-window `scores` (`{ start, end, vadScore, isSpeech }`) and merged speech `segments` (`{ start, end }`), in seconds of file time. Requires `vad: 'silero'`; a `File` opened without `vad` rejects with `analysis requires VAD`.
348
359
  - `file.vadScore`, `'speech'` / `'silence'` events: per-chunk VAD alongside the stream, with the holdoff measured in FILE time (sample positions), never wall-clock time, so processing speed does not change the reported events.
349
360
 
350
- Options mirror `Microphone` (`sampleRate`, `dtype`, `vad`, `dcRemoval`, `denoise`, `highpass`, `agc`, `limiter`); the live-capture options (`device`, `channels`, `framesPerBuffer`) do not apply. Iteration and analysis are separate single passes: construct one `File` per operation. Note: Node also has a global `File` (the web File API); import decibri's explicitly to avoid shadowing surprises.
361
+ Options mirror `Microphone` (`sampleRate`, `channels`, `channelMap`, `dtype`, `vad`, `dcRemoval`, `denoise`, `highpass`, `agc`, `limiter`), with `channels` and `channelMap` read against the source's own channel count (the file's header, or `inputChannels` for `File.buffer`) where the live path reads the device's report, and the vad `source` naming a delivered channel exactly as on `Microphone`; the live-capture options (`device`, `framesPerBuffer`) do not apply. Iteration and analysis are separate single passes: construct one `File` per operation. Note: Node also has a global `File` (the web File API); import decibri's explicitly to avoid shadowing surprises.
351
362
 
352
363
  ## Device Selection
353
364
 
@@ -63,19 +63,21 @@ var decibri = (function() {
63
63
  //#endregion
64
64
  //#region npm/decibri/src/browser/worklet-inline.js
65
65
  var require_worklet_inline = /* @__PURE__ */ __commonJSMin(((exports, module) => {
66
- module.exports = { WORKLET_SOURCE: "var u=class extends AudioWorkletProcessor{constructor(r){super();let e=r.processorOptions;this.framesPerBuffer=e.framesPerBuffer,this.format=e.format,this.ratio=e.nativeSampleRate/e.targetSampleRate,this.needsResample=e.nativeSampleRate!==e.targetSampleRate,this.position=0,this.buffer=new Float32Array(this.framesPerBuffer),this.bufferIndex=0}process(r,e,s){let t=r[0]?.[0];if(!t||t.length===0)return!0;let f;this.needsResample?f=this.resample(t):f=t;let a=0;for(;a<f.length;){let i=this.framesPerBuffer-this.bufferIndex,o=f.length-a,n=Math.min(i,o);this.buffer.set(f.subarray(a,a+n),this.bufferIndex),this.bufferIndex+=n,a+=n,this.bufferIndex>=this.framesPerBuffer&&this.flush()}return!0}resample(r){let e=r.length,s=0,t=this.position;for(;t<e-1;)s++,t+=this.ratio;let f=new Float32Array(s);t=this.position;for(let a=0;a<s;a++){let i=Math.floor(t),o=t-i;f[a]=r[i]*(1-o)+r[i+1]*o,t+=this.ratio}return this.position=Math.max(0,t-e),f}flush(){let r;if(this.format===\"int16\"){let e=new Int16Array(this.framesPerBuffer);for(let s=0;s<this.framesPerBuffer;s++)e[s]=Math.max(-32768,Math.min(32767,Math.round(this.buffer[s]*32768)));r=e.buffer}else r=this.buffer.slice(0,this.framesPerBuffer).buffer;this.port.postMessage(r,[r]),this.buffer=new Float32Array(this.framesPerBuffer),this.bufferIndex=0}};registerProcessor(\"decibri-processor\",u);\n" };
66
+ module.exports = { WORKLET_SOURCE: "var e=class extends AudioWorkletProcessor{constructor(e){super();let t=e.processorOptions;this.framesPerBuffer=t.framesPerBuffer,this.format=t.format,this.ratio=t.nativeSampleRate/t.targetSampleRate,this.needsResample=t.nativeSampleRate!==t.targetSampleRate,this.channelMap=t.channelMap??null,this.channels=this.channelMap?this.channelMap.length:t.channels??1,this.channelError=!1,this.position=0,this.samplesPerChunk=this.framesPerBuffer*this.channels,this.buffer=new Float32Array(this.samplesPerChunk),this.bufferIndex=0}process(e,t,n){let r=e[0];if(!r||r.length===0||!r[0]||r[0].length===0)return!0;if(this.channelError)return!1;let i=r.length,a;if(this.channelMap){for(let e=0;e<this.channelMap.length;e++)if(this.channelMap[e]>=i)return this.refuse(`the channel map names device channel `+this.channelMap[e]+`; the device reports `+i+` input channels`);a=[];for(let e=0;e<this.channelMap.length;e++)a.push(r[this.channelMap[e]])}else if(this.channels===1)if(i===1)a=[r[0]];else{let e=r[0].length,t=new Float32Array(e);for(let n=0;n<e;n++){let e=0;for(let t=0;t<i;t++)e=Math.fround(e+r[t][n]);t[n]=e/i}a=[t]}else if(this.channels===i){a=[];for(let e=0;e<i;e++)a.push(r[e])}else if(this.channels>i)return this.refuse(`the input device does not support `+this.channels+` delivered channels; it reports `+i);else return this.refuse(`a channel map is required to deliver `+this.channels+` of the device's `+i+` input channels`);this.needsResample&&(a=this.resample(a));let o=a[0].length;for(let e=0;e<o;e++){for(let t=0;t<this.channels;t++)this.buffer[this.bufferIndex++]=a[t][e];this.bufferIndex>=this.samplesPerChunk&&this.flush()}return!0}refuse(e){return this.channelError=!0,this.port.postMessage({type:`error`,message:e}),!1}resample(e){let t=e[0].length,n=0,r=this.position;for(;r<t-1;)n++,r+=this.ratio;let i=e.map(()=>new Float32Array(n));r=this.position;for(let t=0;t<n;t++){let n=Math.floor(r),a=r-n;for(let r=0;r<e.length;r++)i[r][t]=e[r][n]*(1-a)+e[r][n+1]*a;r+=this.ratio}return this.position=Math.max(0,r-t),i}flush(){let e;if(this.format===`int16`){let t=new Int16Array(this.samplesPerChunk);for(let e=0;e<this.samplesPerChunk;e++)t[e]=Math.max(-32768,Math.min(32767,Math.round(this.buffer[e]*32768)));e=t.buffer}else e=this.buffer.slice(0,this.samplesPerChunk).buffer;this.port.postMessage(e,[e]),this.buffer=new Float32Array(this.samplesPerChunk),this.bufferIndex=0}};registerProcessor(`decibri-processor`,e);" };
67
67
  }));
68
68
  //#endregion
69
69
  //#region npm/decibri/src/browser/decibri-browser.js
70
70
  var require_decibri_browser = /* @__PURE__ */ __commonJSMin(((exports, module) => {
71
71
  const { Emitter } = require_emitter();
72
72
  const { WORKLET_SOURCE } = require_worklet_inline();
73
- const VERSION = "5.4.0";
73
+ const VERSION = "5.6.0";
74
74
  /**
75
75
  * Browser microphone capture.
76
76
  *
77
77
  * Uses getUserMedia + AudioWorklet for real-time audio capture in browsers.
78
- * Emits 'data' events with Int16Array or Float32Array chunks.
78
+ * Emits 'data' events with Int16Array or Float32Array chunks holding
79
+ * framesPerBuffer frames of the delivered channel count, interleaved frame
80
+ * by frame.
79
81
  *
80
82
  * Ported from decibri-web decibri.ts. Logic identical, types removed.
81
83
  *
@@ -101,6 +103,7 @@ var decibri = (function() {
101
103
  const vad = options.vad ?? false;
102
104
  let vadThreshold = .01;
103
105
  let vadHoldoff = 300;
106
+ let vadSource;
104
107
  if (vad === false) this._vad = false;
105
108
  else if (vad === true) throw new TypeError("vad: true is no longer supported. Specify the mode explicitly: vad: 'energy'.");
106
109
  else if (vad === "energy") this._vad = true;
@@ -117,14 +120,24 @@ var decibri = (function() {
117
120
  if (vad.holdoffMs < 0) throw new RangeError("vad holdoffMs must be non-negative");
118
121
  vadHoldoff = vad.holdoffMs;
119
122
  }
120
- } else throw new TypeError(`Invalid vad value: ${JSON.stringify(vad)}. Expected false, 'energy', or a config object { model, threshold, holdoffMs }.`);
123
+ if (vad.source !== void 0) {
124
+ const source = vad.source;
125
+ if (typeof source !== "number" || !Number.isInteger(source)) throw new TypeError("vad source must be an integer");
126
+ if (source < 0 || source > 65535) throw new RangeError("vad source must be between 0 and 65535");
127
+ const deliveredChannels = options.channels ?? 1;
128
+ if (deliveredChannels >= 1 && source >= deliveredChannels) throw new RangeError(`the detector source names delivered channel ${source}; the delivered channel count is ${deliveredChannels}`);
129
+ vadSource = source;
130
+ }
131
+ } else throw new TypeError(`Invalid vad value: ${JSON.stringify(vad)}. Expected false, 'energy', or a config object { model, threshold, holdoffMs, source }.`);
121
132
  this._vadThreshold = vadThreshold;
122
133
  this._vadHoldoff = vadHoldoff;
134
+ this._vadSource = vadSource;
123
135
  this._vadScore = 0;
124
136
  this._isSpeaking = false;
125
137
  this._silenceTimer = null;
126
138
  this._sampleRate = options.sampleRate ?? 16e3;
127
139
  this._channels = options.channels ?? 1;
140
+ this._channelMap = options.channelMap;
128
141
  this._framesPerBuffer = options.framesPerBuffer ?? 1600;
129
142
  this._device = options.device;
130
143
  this._dtype = options.dtype ?? "int16";
@@ -132,7 +145,16 @@ var decibri = (function() {
132
145
  this._noiseSuppression = options.noiseSuppression ?? true;
133
146
  this._workletUrl = options.workletUrl;
134
147
  if (this._sampleRate < 1e3 || this._sampleRate > 384e3) throw new RangeError("sample rate must be between 1000 and 384000");
135
- if (this._channels < 1 || this._channels > 32) throw new TypeError(`channels must be between 1 and 32, got ${this._channels}`);
148
+ if (this._channels < 1) throw new RangeError("channels must be at least 1");
149
+ const channelMap = this._channelMap;
150
+ if (channelMap !== void 0) {
151
+ if (!Array.isArray(channelMap)) throw new TypeError(`Invalid channelMap value: ${JSON.stringify(channelMap)}. Expected an array of 0-based device channel indices, such as [0].`);
152
+ for (const entry of channelMap) {
153
+ if (typeof entry !== "number" || !Number.isInteger(entry)) throw new TypeError("channelMap entries must be integers");
154
+ if (entry < 0 || entry > 65535) throw new RangeError("channelMap entries must be between 0 and 65535");
155
+ }
156
+ if (channelMap.length !== this._channels) throw new RangeError("channelMap must have exactly one entry per channel");
157
+ }
136
158
  if (this._framesPerBuffer < 64 || this._framesPerBuffer > 65536) throw new TypeError(`frames per buffer must be between 64 and 65536, got ${this._framesPerBuffer}`);
137
159
  if (this._dtype !== "int16" && this._dtype !== "float32") throw new TypeError("dtype must be 'int16' or 'float32'");
138
160
  }
@@ -188,6 +210,9 @@ var decibri = (function() {
188
210
  /**
189
211
  * Most recent VAD score: the normalized RMS of the last chunk in `'energy'`
190
212
  * mode, or 0 when VAD is disabled or before the first chunk is processed.
213
+ * A chunk carrying more than one channel is collapsed to the average of
214
+ * its channels, or to the one delivered channel a `vad: { source }` names,
215
+ * before the RMS, so the score reflects one channel's level.
191
216
  * @returns {number}
192
217
  */
193
218
  get vadScore() {
@@ -213,7 +238,7 @@ var decibri = (function() {
213
238
  const nativeSampleRate = this._audioContext.sampleRate;
214
239
  await this._audioContext.resume();
215
240
  const audioConstraints = {
216
- channelCount: this._channels,
241
+ channelCount: { ideal: 32 },
217
242
  echoCancellation: this._echoCancellation,
218
243
  noiseSuppression: this._noiseSuppression
219
244
  };
@@ -227,6 +252,29 @@ var decibri = (function() {
227
252
  this.emit("error", error);
228
253
  throw error;
229
254
  }
255
+ if (this._channelMap !== void 0 || this._channels > 1) {
256
+ const track = this._stream.getAudioTracks()[0];
257
+ const granted = (track && typeof track.getSettings === "function" ? track.getSettings() : {}).channelCount;
258
+ if (typeof granted === "number") {
259
+ let message = null;
260
+ if (this._channelMap !== void 0) {
261
+ for (const entry of this._channelMap) if (entry >= granted) {
262
+ message = `the channel map names device channel ${entry}; the device reports ${granted} input channels`;
263
+ break;
264
+ }
265
+ } else if (this._channels > granted) message = `the input device does not support ${this._channels} delivered channels; it reports ${granted}`;
266
+ else if (this._channels < granted) message = `a channel map is required to deliver ${this._channels} of the device's ${granted} input channels`;
267
+ if (message !== null) {
268
+ this._stream.getTracks().forEach((t) => t.stop());
269
+ this._stream = null;
270
+ await this._audioContext.close();
271
+ this._audioContext = null;
272
+ const error = new Error(message);
273
+ this.emit("error", error);
274
+ throw error;
275
+ }
276
+ }
277
+ }
230
278
  let blobUrl = null;
231
279
  const workletUrl = this._workletUrl ?? (blobUrl = this._createBlobUrl());
232
280
  try {
@@ -247,11 +295,21 @@ var decibri = (function() {
247
295
  framesPerBuffer: this._framesPerBuffer,
248
296
  format: this._dtype,
249
297
  nativeSampleRate,
250
- targetSampleRate: this._sampleRate
298
+ targetSampleRate: this._sampleRate,
299
+ channels: this._channels,
300
+ channelMap: this._channelMap ?? null
251
301
  } });
252
302
  this._workletNode.port.onmessage = (event) => {
253
- const buffer = event.data;
254
- const chunk = this._dtype === "int16" ? new Int16Array(buffer) : new Float32Array(buffer);
303
+ const data = event.data;
304
+ if (!(data instanceof ArrayBuffer)) {
305
+ if (data && data.type === "error") {
306
+ const error = new Error(data.message);
307
+ this.emit("error", error);
308
+ this.stop();
309
+ }
310
+ return;
311
+ }
312
+ const chunk = this._dtype === "int16" ? new Int16Array(data) : new Float32Array(data);
255
313
  this.emit("data", chunk);
256
314
  if (this._vad) this._processVad(chunk);
257
315
  };
@@ -293,10 +351,35 @@ var decibri = (function() {
293
351
  }, this._vadHoldoff);
294
352
  }
295
353
  _computeRms(chunk) {
296
- let sum = 0;
297
354
  const n = chunk.length;
298
355
  if (n === 0) return 0;
299
- if (chunk instanceof Float32Array) for (let i = 0; i < n; i++) sum += chunk[i] * chunk[i];
356
+ const channels = this._channels;
357
+ const isFloat = chunk instanceof Float32Array;
358
+ if (channels > 1) {
359
+ const frames = Math.floor(n / channels);
360
+ if (frames === 0) return 0;
361
+ const source = this._vadSource;
362
+ let sum = 0;
363
+ if (source !== void 0) {
364
+ for (let f = 0; f < frames; f++) {
365
+ const s = isFloat ? chunk[f * channels + source] : chunk[f * channels + source] / 32768;
366
+ sum += s * s;
367
+ }
368
+ return Math.sqrt(sum / frames);
369
+ }
370
+ for (let f = 0; f < frames; f++) {
371
+ let acc = 0;
372
+ for (let c = 0; c < channels; c++) {
373
+ const s = isFloat ? chunk[f * channels + c] : chunk[f * channels + c] / 32768;
374
+ acc = Math.fround(acc + s);
375
+ }
376
+ const mono = Math.fround(acc / channels);
377
+ sum += mono * mono;
378
+ }
379
+ return Math.sqrt(sum / frames);
380
+ }
381
+ let sum = 0;
382
+ if (isFloat) for (let i = 0; i < n; i++) sum += chunk[i] * chunk[i];
300
383
  else for (let i = 0; i < n; i++) {
301
384
  const s = chunk[i] / 32768;
302
385
  sum += s * s;