decibri 3.4.2 → 4.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,6 +1,6 @@
1
1
  # decibri
2
2
 
3
- Cross-platform audio capture, output, and voice activity detection for Node.js and browsers.
3
+ Cross-platform audio capture, playback, and voice activity detection for Node.js and browsers.
4
4
 
5
5
  ## Installation
6
6
 
@@ -8,18 +8,24 @@ Cross-platform audio capture, output, and voice activity detection for Node.js a
8
8
  npm install decibri
9
9
  ```
10
10
 
11
- One package for Node.js and browsers. Node.js gets a native addon (Rust via napi-rs). Browsers get a JavaScript AudioWorklet implementation. Platform-specific binaries are installed automatically.
11
+ One package for Node.js and browsers. Node.js gets a prebuilt native addon. Browsers get a JavaScript AudioWorklet implementation. Platform-specific binaries are installed automatically.
12
12
 
13
13
  Requires Node.js >= 18. TypeScript definitions are bundled.
14
14
 
15
+ The package uses named exports:
16
+
17
+ ```javascript
18
+ const { Microphone, Speaker, inputDevices, outputDevices, version } = require('decibri');
19
+ ```
20
+
15
21
  ## Quick Start
16
22
 
17
23
  ### Capture audio
18
24
 
19
25
  ```javascript
20
- const Decibri = require('decibri');
26
+ const { Microphone } = require('decibri');
21
27
 
22
- const mic = new Decibri({ sampleRate: 16000, channels: 1 });
28
+ const mic = new Microphone({ sampleRate: 16000, channels: 1 });
23
29
  mic.on('data', (chunk) => { /* Buffer of Int16 PCM samples */ });
24
30
  setTimeout(() => mic.stop(), 5000);
25
31
  ```
@@ -27,9 +33,9 @@ setTimeout(() => mic.stop(), 5000);
27
33
  ### Play audio
28
34
 
29
35
  ```javascript
30
- const { DecibriOutput } = require('decibri');
36
+ const { Speaker } = require('decibri');
31
37
 
32
- const speaker = new DecibriOutput({ sampleRate: 16000, channels: 1 });
38
+ const speaker = new Speaker({ sampleRate: 16000, channels: 1 });
33
39
  speaker.write(pcmBuffer);
34
40
  speaker.end();
35
41
  ```
@@ -37,9 +43,9 @@ speaker.end();
37
43
  ### Browser capture
38
44
 
39
45
  ```javascript
40
- import { Decibri } from 'decibri'; // browser entry via conditional export
46
+ import { Microphone } from 'decibri'; // browser entry via conditional export
41
47
 
42
- const mic = new Decibri({ sampleRate: 16000 });
48
+ const mic = new Microphone({ sampleRate: 16000 });
43
49
  mic.on('data', (chunk) => { /* Int16Array of PCM samples */ });
44
50
  await mic.start(); // requires user gesture in Safari
45
51
  ```
@@ -47,14 +53,16 @@ await mic.start(); // requires user gesture in Safari
47
53
  ### Pipe capture to playback (echo)
48
54
 
49
55
  ```javascript
50
- const mic = new Decibri({ sampleRate: 16000, channels: 1 });
51
- const speaker = new DecibriOutput({ sampleRate: 16000, channels: 1 });
56
+ const { Microphone, Speaker } = require('decibri');
57
+
58
+ const mic = new Microphone({ sampleRate: 16000, channels: 1 });
59
+ const speaker = new Speaker({ sampleRate: 16000, channels: 1 });
52
60
  mic.pipe(speaker);
53
61
  ```
54
62
 
55
- ## API: Decibri (Capture)
63
+ ## API: Microphone (Capture)
56
64
 
57
- ### `new Decibri(options?)`
65
+ ### `new Microphone(options?)`
58
66
 
59
67
  Creates a Readable stream that captures from the microphone.
60
68
 
@@ -64,27 +72,33 @@ Creates a Readable stream that captures from the microphone.
64
72
  | `channels` | number | 1 | Input channels (1 to 32) |
65
73
  | `framesPerBuffer` | number | 1600 | Frames per chunk (64 to 65536). At 16kHz mono, 1600 = 100ms = 3200 bytes |
66
74
  | `device` | number, string, or `{ id: string }` | system default | Device index, case-insensitive name substring, or stable per-host ID |
67
- | `format` | `'int16'` \| `'float32'` | `'int16'` | Sample encoding |
68
- | `vad` | boolean | false | Enable voice activity detection |
69
- | `vadMode` | `'energy'` \| `'silero'` | `'energy'` | VAD engine: RMS threshold or Silero ML model |
70
- | `vadThreshold` | number | 0.01 / 0.5 | Speech threshold. Default depends on `vadMode` |
75
+ | `dtype` | `'int16'` \| `'float32'` | `'int16'` | Sample encoding |
76
+ | `vad` | `false` \| `'silero'` \| `'energy'` | `false` | Voice activity detection: disabled, the Silero ML model, or an RMS energy threshold |
77
+ | `vadThreshold` | number | 0.5 / 0.01 | Speech threshold. Default is 0.5 for `'silero'`, 0.01 for `'energy'` |
71
78
  | `vadHoldoff` | number | 300 | Silence holdoff in ms |
79
+ | `modelPath` | string | bundled model | Path to the Silero model. Only used when `vad` is `'silero'` |
72
80
 
73
81
  Standard `ReadableOptions` (e.g. `highWaterMark`) are also accepted.
74
82
 
83
+ `vad: true` is not accepted; pass the mode explicitly as `vad: 'silero'` or `vad: 'energy'`.
84
+
75
85
  ### Methods
76
86
 
77
87
  | Method | Description |
78
88
  | --- | --- |
79
89
  | `mic.stop()` | Stop capture and end stream. Safe to call multiple times |
80
- | `Decibri.devices()` | List available input devices |
81
- | `Decibri.version()` | Version info |
90
+ | `Microphone.open(options?)` | Construct without blocking the event loop. Returns a `Promise<Microphone>`. See [Non-blocking API](#non-blocking-api) |
91
+ | `Microphone.devices()` | List available input devices |
92
+ | `Microphone.version()` | Version info: `{ decibri, audioBackend, binding }` |
93
+
94
+ The module-level `inputDevices()` and `version()` free functions are equivalent to the static methods.
82
95
 
83
96
  ### Properties
84
97
 
85
98
  | Property | Type | Description |
86
99
  | --- | --- | --- |
87
100
  | `mic.isOpen` | boolean | `true` while capturing |
101
+ | `mic.vadScore` | number | Latest VAD score for the active mode (Silero probability or normalized RMS); 0 when disabled |
88
102
 
89
103
  ### Events
90
104
 
@@ -97,9 +111,9 @@ Standard `ReadableOptions` (e.g. `highWaterMark`) are also accepted.
97
111
  | `'end'` | - | Stream ended |
98
112
  | `'error'` | Error | An error occurred |
99
113
 
100
- ## API: DecibriOutput (Playback)
114
+ ## API: Speaker (Playback)
101
115
 
102
- ### `new DecibriOutput(options?)`
116
+ ### `new Speaker(options?)`
103
117
 
104
118
  Creates a Writable stream for speaker playback.
105
119
 
@@ -107,7 +121,7 @@ Creates a Writable stream for speaker playback.
107
121
  | --- | --- | --- | --- |
108
122
  | `sampleRate` | number | 16000 | Playback sample rate (1000 to 384000) |
109
123
  | `channels` | number | 1 | Output channels (1 to 32) |
110
- | `format` | `'int16'` \| `'float32'` | `'int16'` | Sample encoding of incoming data |
124
+ | `dtype` | `'int16'` \| `'float32'` | `'int16'` | Sample encoding of incoming data |
111
125
  | `device` | number, string, or `{ id: string }` | system default | Output device index, case-insensitive name substring, or stable per-host ID |
112
126
 
113
127
  Standard `WritableOptions` (e.g. `highWaterMark`) are also accepted.
@@ -115,27 +129,82 @@ Standard `WritableOptions` (e.g. `highWaterMark`) are also accepted.
115
129
  | Method / Property | Description |
116
130
  | --- | --- |
117
131
  | `speaker.write(chunk)` | Write PCM data for playback |
132
+ | `speaker.writeAsync(chunk)` | Write without blocking the event loop. Returns a `Promise`. See [Non-blocking API](#non-blocking-api) |
118
133
  | `speaker.end()` | Signal end. Drains remaining audio, then emits `'finish'` |
134
+ | `speaker.drainAsync()` | Wait for queued audio to finish without blocking the event loop. Returns a `Promise` |
119
135
  | `speaker.stop()` | Immediate stop. Discards remaining audio |
120
136
  | `speaker.isPlaying` | `true` while audio is being output |
121
- | `DecibriOutput.devices()` | List available output devices |
122
- | `DecibriOutput.version()` | Same as `Decibri.version()` |
137
+ | `Speaker.open(options?)` | Construct without blocking the event loop. Returns a `Promise<Speaker>` |
138
+ | `Speaker.devices()` | List available output devices |
139
+ | `Speaker.version()` | Same as `Microphone.version()` |
140
+
141
+ The module-level `outputDevices()` free function is equivalent to `Speaker.devices()`.
142
+
143
+ ## Non-blocking API
144
+
145
+ The synchronous constructors and `write` / `drain` do their work on the event loop, which is fine for most apps. For event-loop-sensitive code (servers, real-time voice pipelines), 4.1.0 adds async variants that perform the blocking work without stalling the event loop. They are additive: the synchronous API is unchanged, and you opt in only where you need it.
146
+
147
+ ### Non-blocking construction
148
+
149
+ `Microphone.open(options?)` and `Speaker.open(options?)` are async factories that return a Promise of a ready instance. They take the same options as the constructors. For a microphone with Silero VAD, the model load that the constructor does inline runs without blocking the event loop.
150
+
151
+ ```javascript
152
+ const { Microphone, Speaker } = require('decibri');
153
+
154
+ const mic = await Microphone.open({ sampleRate: 16000, vad: 'silero' });
155
+ const speaker = await Speaker.open({ sampleRate: 16000, channels: 1 });
156
+ ```
157
+
158
+ The synchronous `new Microphone(...)` and `new Speaker(...)` still work unchanged. A failed open rejects the Promise with the same typed error a failed constructor throws.
159
+
160
+ ### Non-blocking playback
161
+
162
+ `speaker.writeAsync(chunk)` resolves once the audio is queued, performing the backpressure wait (when the playback buffer is full) without blocking the event loop. `speaker.drainAsync()` resolves when all queued audio has finished playing, again without blocking.
163
+
164
+ ```javascript
165
+ const speaker = await Speaker.open({ sampleRate: 16000, channels: 1 });
166
+
167
+ await speaker.writeAsync(pcmBuffer);
168
+ await speaker.drainAsync(); // resolves when playback finishes
169
+ ```
170
+
171
+ These are a direct alternative to the synchronous `write()` / `pipe()` / `end()` stream interface, which is unchanged. Use one path per instance (the stream methods or the async methods, not both at once), and await calls in sequence to keep samples in order.
172
+
173
+ ## Errors
174
+
175
+ Construction errors come as typed classes you can catch:
176
+
177
+ ```javascript
178
+ const { Microphone, DecibriError, DeviceError } = require('decibri');
179
+
180
+ try {
181
+ new Microphone({ device: 'no such device' });
182
+ } catch (err) {
183
+ if (err instanceof DeviceError) {
184
+ console.log(err.code); // e.g. 'MICROPHONE_NOT_FOUND'
185
+ }
186
+ }
187
+ ```
188
+
189
+ `DeviceError`, `OrtError`, and `OrtPathError` extend `DecibriError`, which extends `Error`. Each carries a stable `code` string. Argument validation (bad `sampleRate`, `channels`, `dtype`, or `vad`) throws a built-in `RangeError` or `TypeError`.
123
190
 
124
191
  ## API: Browser
125
192
 
126
193
  The browser API uses `getUserMedia` and `AudioWorklet`. It differs from the Node.js API because browser audio is fundamentally async.
127
194
 
128
- ### `new Decibri(options?)` (browser)
195
+ ### `new Microphone(options?)` (browser)
129
196
 
130
197
  Same options as Node.js, plus:
131
198
 
132
199
  | Option | Type | Default | Description |
133
200
  | --- | --- | --- | --- |
134
- | `device` | string | system default | Device ID from `Decibri.devices()` (not index) |
201
+ | `device` | string | system default | Device ID from `Microphone.devices()` (not index) |
135
202
  | `echoCancellation` | boolean | true | Browser echo cancellation |
136
203
  | `noiseSuppression` | boolean | true | Browser noise suppression |
137
204
  | `workletUrl` | string | inline blob | Custom worklet URL for strict CSP |
138
205
 
206
+ The browser runs energy-mode VAD only, so its `vad` option accepts `false` or `'energy'`. The browser `version()` returns `{ decibri }` only.
207
+
139
208
  ### Key differences from Node.js
140
209
 
141
210
  | Aspect | Node.js | Browser |
@@ -144,49 +213,52 @@ Same options as Node.js, plus:
144
213
  | Base class | Readable stream | Custom Emitter |
145
214
  | Data type | Buffer | Int16Array / Float32Array |
146
215
  | `devices()` | Sync, returns array | Async, returns Promise |
147
- | Sample rate | Direct via cpal | Resampled from native rate |
216
+ | Sample rate | Native device rate | Resampled from native rate |
217
+ | VAD | `'silero'` or `'energy'` | `'energy'` only |
148
218
 
149
219
  ## Voice Activity Detection
150
220
 
151
- ### Energy mode (default)
221
+ ### Energy mode
152
222
 
153
223
  Lightweight RMS energy threshold. No model required.
154
224
 
155
225
  ```javascript
156
- const mic = new Decibri({ vad: true, vadThreshold: 0.01 });
226
+ const mic = new Microphone({ vad: 'energy', vadThreshold: 0.01 });
157
227
  mic.on('speech', () => console.log('speaking'));
158
228
  mic.on('silence', () => console.log('silent'));
159
229
  ```
160
230
 
161
231
  ### Silero mode
162
232
 
163
- ML-based detection using the Silero VAD v5 ONNX model. More accurate than energy mode, especially in noisy environments.
233
+ ML-based detection using the Silero VAD v5 model. More accurate than energy mode, especially in noisy environments.
164
234
 
165
235
  ```javascript
166
- const mic = new Decibri({ vad: true, vadMode: 'silero', vadThreshold: 0.5 });
236
+ const mic = new Microphone({ vad: 'silero', vadThreshold: 0.5 });
167
237
  mic.on('speech', () => console.log('speaking'));
168
238
  mic.on('silence', () => console.log('silent'));
169
239
  ```
170
240
 
171
- The Silero model (~2MB) ships inside the npm package. No downloads or API keys required.
241
+ The Silero model (~2MB) ships inside the npm package. No downloads or API keys required. Silero mode is Node.js only.
172
242
 
173
243
  ## Device Selection
174
244
 
175
245
  ```javascript
246
+ const { Microphone } = require('decibri');
247
+
176
248
  // System default
177
- const mic = new Decibri();
249
+ const mic = new Microphone();
178
250
 
179
251
  // By name (case-insensitive substring match)
180
- const mic = new Decibri({ device: 'USB' });
252
+ const mic = new Microphone({ device: 'USB' });
181
253
 
182
254
  // By index
183
- const devices = Decibri.devices();
184
- const mic = new Decibri({ device: devices[1].index });
255
+ const devices = Microphone.devices();
256
+ const mic = new Microphone({ device: devices[1].index });
185
257
 
186
258
  // By stable per-host ID (survives across enumerations)
187
- const mic = new Decibri({ device: { id: devices[1].id } });
259
+ const mic = new Microphone({ device: { id: devices[1].id } });
188
260
 
189
- Decibri.devices();
261
+ Microphone.devices();
190
262
  // [
191
263
  // { index: 0, name: 'Microphone', id: '{0.0.1.00000000}.{...}', maxInputChannels: 2, defaultSampleRate: 48000, isDefault: true },
192
264
  // { index: 1, name: 'USB Headset', id: '{0.0.1.00000000}.{...}', maxInputChannels: 1, defaultSampleRate: 44100, isDefault: false }
@@ -206,6 +278,10 @@ node node_modules/decibri/examples/websocket-server.js # terminal 1
206
278
  node node_modules/decibri/examples/websocket-stream.js # terminal 2
207
279
  ```
208
280
 
281
+ ## Migrating from 3.x
282
+
283
+ decibri 4.0.0 renames the API to a microphone and speaker vocabulary and switches to named exports. See [MIGRATION.md](./MIGRATION.md) for a complete before-and-after guide.
284
+
209
285
  ## Platform Support
210
286
 
211
287
  | Platform | Architecture | Audio Backend |
@@ -218,9 +294,9 @@ node node_modules/decibri/examples/websocket-stream.js # terminal 2
218
294
 
219
295
  ## How It Works
220
296
 
221
- decibri is a Rust library using cpal for cross-platform audio I/O. The Rust core compiles to a Node.js native addon via napi-rs and ships pre-built binaries for each platform. Browser support uses a JavaScript AudioWorklet implementation with the same event-driven API.
297
+ decibri compiles a Rust audio core to a Node.js native addon and ships prebuilt binaries for each platform, so there is no build step on install. Browser support uses a JavaScript AudioWorklet implementation with the same event-driven API.
222
298
 
223
- Audio flows from the OS audio device through cpal's callback, into a crossbeam channel, through frame-exact buffering (guarantees consistent chunk sizes), and into Node.js via a threadsafe function. The JavaScript layer wraps this in a standard Readable stream.
299
+ On Node.js, audio flows from the OS audio device through frame-exact buffering (which guarantees consistent chunk sizes) and into a standard Readable stream. In the browser, audio is captured and resampled in an AudioWorklet and delivered through the same `'data'` event interface.
224
300
 
225
301
  ## Documentation
226
302
 
@@ -5,14 +5,14 @@
5
5
  // Usage (from repo clone): node npm/decibri/examples/wav-capture.js
6
6
 
7
7
  const fs = require('fs');
8
- const Decibri = require('decibri');
8
+ const { Microphone } = require('decibri');
9
9
 
10
10
  const SAMPLE_RATE = 16000;
11
11
  const CHANNELS = 1;
12
12
  const BITS_PER_SAMPLE = 16;
13
13
  const DURATION_MS = 5000;
14
14
 
15
- const mic = new Decibri({ sampleRate: SAMPLE_RATE, channels: CHANNELS });
15
+ const mic = new Microphone({ sampleRate: SAMPLE_RATE, channels: CHANNELS });
16
16
  const chunks = [];
17
17
 
18
18
  mic.on('data', (chunk) => chunks.push(chunk));
@@ -5,11 +5,11 @@
5
5
  // Usage (from repo clone): node npm/decibri/examples/websocket-stream.js [ws://localhost:8080]
6
6
 
7
7
  const { WebSocket } = require('ws');
8
- const Decibri = require('decibri');
8
+ const { Microphone } = require('decibri');
9
9
 
10
10
  const url = process.argv[2] || 'ws://localhost:8080';
11
11
  const ws = new WebSocket(url);
12
- const mic = new Decibri({ sampleRate: 16000, channels: 1 });
12
+ const mic = new Microphone({ sampleRate: 16000, channels: 1 });
13
13
 
14
14
  ws.on('open', () => {
15
15
  console.log(`Connected to ${url}, streaming audio. Ctrl+C to stop.`);
package/index.d.ts CHANGED
@@ -3,6 +3,14 @@
3
3
  /** Native bridge class exposed to Node.js via napi-rs. */
4
4
  export declare class DecibriBridge {
5
5
  constructor(options?: DecibriOptions | undefined | null)
6
+ /**
7
+ * Construct a microphone bridge without blocking the JS event loop. The
8
+ * device resolution and Silero model load run on the libuv thread pool;
9
+ * the returned Promise resolves to a fully constructed bridge, or rejects
10
+ * with the matching error. The synchronous `new` remains available and
11
+ * unchanged.
12
+ */
13
+ static openAsync(options?: DecibriOptions | undefined | null): Promise<unknown>
6
14
  /** Start capturing audio. The callback receives `(err, chunk)` for each buffer. */
7
15
  start(callback: (err: Error | null, chunk: Buffer) => void): void
8
16
  /** Stop capturing audio. */
@@ -23,13 +31,36 @@ export declare class DecibriBridge {
23
31
  /** Native bridge class for audio output, exposed to Node.js via napi-rs. */
24
32
  export declare class DecibriOutputBridge {
25
33
  constructor(options?: DecibriOutputOptions | undefined | null)
34
+ /**
35
+ * Construct a speaker bridge without blocking the JS event loop. The device
36
+ * resolution runs on the libuv thread pool; the returned Promise resolves
37
+ * to a constructed bridge, or rejects with the matching error. The
38
+ * synchronous `new` remains available and unchanged.
39
+ */
40
+ static openAsync(options?: DecibriOutputOptions | undefined | null): Promise<unknown>
26
41
  /**
27
42
  * Write PCM data for playback. Starts the output stream on first call.
28
43
  * Empty buffers are a no-op.
29
44
  */
30
45
  write(buffer: Buffer): void
46
+ /**
47
+ * Non-blocking write: convert the samples and start the stream on the JS
48
+ * thread (a fast device open, same as the synchronous first write), then
49
+ * perform the blocking channel `send` (which stalls under backpressure when
50
+ * the queue is full) on the libuv thread pool. The returned Promise
51
+ * resolves when the samples are queued, or rejects with the matching error.
52
+ * Empty buffers resolve immediately. The synchronous `write` is unchanged.
53
+ */
54
+ writeAsync(buffer: Buffer): Promise<unknown>
31
55
  /** Graceful drain: blocks until all queued samples have been played. */
32
56
  drain(): void
57
+ /**
58
+ * Non-blocking drain: the poll loop that waits for the cpal callback to play
59
+ * everything queued runs on the libuv thread pool instead of the event loop.
60
+ * The returned Promise resolves when the buffer has drained. With no stream
61
+ * yet created it resolves immediately. The synchronous `drain` is unchanged.
62
+ */
63
+ drainAsync(): Promise<unknown>
33
64
  /** Immediate stop. Discards remaining samples. */
34
65
  stop(): void
35
66
  /** Whether audio is currently being output. */
@@ -115,5 +146,5 @@ export interface OutputDeviceInfoJs {
115
146
  /** Version info returned to JS. */
116
147
  export interface VersionInfoJs {
117
148
  decibri: string
118
- portaudio: string
149
+ audioBackend: string
119
150
  }