decibri 4.0.0 → 4.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -11,6 +11,19 @@ For other decibri packages, see:
11
11
 
12
12
  ## [Unreleased]
13
13
 
14
+ ## [4.2.0] - 2026-05-31
15
+
16
+ ### Added
17
+
18
+ - Browser `Speaker` for audio playback through the Web Audio API: `start()`, async `write(chunk)`, async `drain()`, `stop()`, and an `isPlaying` getter, with `int16` and `float32` input and resampling from the source rate to the output rate. Playback is started from a user gesture, as browsers require. This adds playback to the browser build alongside the existing browser `Microphone` capture. The release is browser-only and additive: the Node.js API and behavior are unchanged.
19
+
20
+ ## [4.1.0] - 2026-05-31
21
+
22
+ ### Added
23
+
24
+ - Async factories `Microphone.open(options)` and `Speaker.open(options)`, each returning a `Promise` that resolves to a constructed instance. They perform the blocking open work (the Silero VAD model load for the microphone; device resolution for both) on the native thread pool instead of the event loop, so latency-sensitive callers do not stall during construction. A failed open rejects with the matching error (`RangeError` / `TypeError` for invalid options, `DeviceError` / `OrtError` / `OrtPathError` for native failures). The synchronous `new Microphone(...)` and `new Speaker(...)` constructors are unchanged; the factories are an additive, non-blocking alternative that mirrors the Python `AsyncMicrophone.open()` / `AsyncSpeaker.open()` surface.
25
+ - Non-blocking `Speaker.writeAsync(chunk)` and `Speaker.drainAsync()` methods, each returning a `Promise`. They run the blocking parts of playback off the event loop: `writeAsync` performs the backpressure wait when the native playback queue is full, and `drainAsync` performs the wait for queued audio to finish playing. The audio stream stays on its own thread; only the thread-safe sample channel and drain state are used off the event loop. A failed write or drain rejects with the matching error class. The synchronous `write()` / `pipe()` / `end()` Writable interface is unchanged; the async methods are an additive, direct alternative (do not interleave the two paths on one instance).
26
+
14
27
  ## [4.0.0] - 2026-05-30
15
28
 
16
29
  ### Changed
package/MIGRATION.md CHANGED
@@ -5,6 +5,37 @@ vocabulary that matches the Rust and Python packages, and tidies several option
5
5
  and return shapes. This guide lists every breaking change with before and after
6
6
  code.
7
7
 
8
+ ## New in 4.2.0 (additive, nothing to migrate)
9
+
10
+ decibri 4.2.0 is a browser-only, additive release. Code written for 4.1.0 keeps
11
+ working unchanged; there is nothing to migrate. The release adds audio playback
12
+ to the browser build:
13
+
14
+ - A browser `Speaker` that plays audio through the Web Audio API: `start()`,
15
+ async `write(chunk)`, async `drain()`, `stop()`, and `isPlaying`, with int16
16
+ and float32 input and resampling from the source rate to the output rate. It
17
+ is started from a user gesture, as browsers require. The browser `Microphone`
18
+ (capture) and the entire Node.js API are unchanged.
19
+
20
+ See the browser Speaker section of the README for examples.
21
+
22
+ ## New in 4.1.0 (additive, nothing to migrate)
23
+
24
+ decibri 4.1.0 is a non-breaking, additive release. Code written for 4.0.0 keeps
25
+ working unchanged; there is nothing to migrate. The release adds an opt-in
26
+ non-blocking API for event-loop-sensitive code:
27
+
28
+ - `Microphone.open(options)` and `Speaker.open(options)`: async factories that
29
+ construct an instance without blocking the event loop and resolve to a ready
30
+ instance. The synchronous `new Microphone(...)` and `new Speaker(...)`
31
+ constructors are unchanged.
32
+ - `speaker.writeAsync(chunk)` and `speaker.drainAsync()`: write and drain
33
+ without blocking the event loop. The synchronous `write()` / `pipe()` /
34
+ `end()` interface is unchanged.
35
+
36
+ See the Non-blocking API section of the README for examples. The rest of this
37
+ guide covers the 4.0.0 changes from 3.x.
38
+
8
39
  ## Named exports
9
40
 
10
41
  The package no longer has a single default export. Destructure what you need.
package/README.md CHANGED
@@ -50,6 +50,19 @@ mic.on('data', (chunk) => { /* Int16Array of PCM samples */ });
50
50
  await mic.start(); // requires user gesture in Safari
51
51
  ```
52
52
 
53
+ ### Browser playback
54
+
55
+ ```javascript
56
+ import { Speaker } from 'decibri'; // browser entry via conditional export
57
+
58
+ const speaker = new Speaker({ sampleRate: 16000 });
59
+ playButton.onclick = async () => {
60
+ await speaker.write(int16Chunk); // Int16Array of PCM samples
61
+ await speaker.drain(); // resolves when playback finishes
62
+ speaker.stop();
63
+ };
64
+ ```
65
+
53
66
  ### Pipe capture to playback (echo)
54
67
 
55
68
  ```javascript
@@ -87,6 +100,7 @@ Standard `ReadableOptions` (e.g. `highWaterMark`) are also accepted.
87
100
  | Method | Description |
88
101
  | --- | --- |
89
102
  | `mic.stop()` | Stop capture and end stream. Safe to call multiple times |
103
+ | `Microphone.open(options?)` | Construct without blocking the event loop. Returns a `Promise<Microphone>`. See [Non-blocking API](#non-blocking-api) |
90
104
  | `Microphone.devices()` | List available input devices |
91
105
  | `Microphone.version()` | Version info: `{ decibri, audioBackend, binding }` |
92
106
 
@@ -128,14 +142,47 @@ Standard `WritableOptions` (e.g. `highWaterMark`) are also accepted.
128
142
  | Method / Property | Description |
129
143
  | --- | --- |
130
144
  | `speaker.write(chunk)` | Write PCM data for playback |
145
+ | `speaker.writeAsync(chunk)` | Write without blocking the event loop. Returns a `Promise`. See [Non-blocking API](#non-blocking-api) |
131
146
  | `speaker.end()` | Signal end. Drains remaining audio, then emits `'finish'` |
147
+ | `speaker.drainAsync()` | Wait for queued audio to finish without blocking the event loop. Returns a `Promise` |
132
148
  | `speaker.stop()` | Immediate stop. Discards remaining audio |
133
149
  | `speaker.isPlaying` | `true` while audio is being output |
150
+ | `Speaker.open(options?)` | Construct without blocking the event loop. Returns a `Promise<Speaker>` |
134
151
  | `Speaker.devices()` | List available output devices |
135
152
  | `Speaker.version()` | Same as `Microphone.version()` |
136
153
 
137
154
  The module-level `outputDevices()` free function is equivalent to `Speaker.devices()`.
138
155
 
156
+ ## Non-blocking API
157
+
158
+ The synchronous constructors and `write` / `drain` do their work on the event loop, which is fine for most apps. For event-loop-sensitive code (servers, real-time voice pipelines), 4.1.0 adds async variants that perform the blocking work without stalling the event loop. They are additive: the synchronous API is unchanged, and you opt in only where you need it.
159
+
160
+ ### Non-blocking construction
161
+
162
+ `Microphone.open(options?)` and `Speaker.open(options?)` are async factories that return a Promise of a ready instance. They take the same options as the constructors. For a microphone with Silero VAD, the model load that the constructor does inline runs without blocking the event loop.
163
+
164
+ ```javascript
165
+ const { Microphone, Speaker } = require('decibri');
166
+
167
+ const mic = await Microphone.open({ sampleRate: 16000, vad: 'silero' });
168
+ const speaker = await Speaker.open({ sampleRate: 16000, channels: 1 });
169
+ ```
170
+
171
+ The synchronous `new Microphone(...)` and `new Speaker(...)` still work unchanged. A failed open rejects the Promise with the same typed error a failed constructor throws.
172
+
173
+ ### Non-blocking playback
174
+
175
+ `speaker.writeAsync(chunk)` resolves once the audio is queued, performing the backpressure wait (when the playback buffer is full) without blocking the event loop. `speaker.drainAsync()` resolves when all queued audio has finished playing, again without blocking.
176
+
177
+ ```javascript
178
+ const speaker = await Speaker.open({ sampleRate: 16000, channels: 1 });
179
+
180
+ await speaker.writeAsync(pcmBuffer);
181
+ await speaker.drainAsync(); // resolves when playback finishes
182
+ ```
183
+
184
+ These are a direct alternative to the synchronous `write()` / `pipe()` / `end()` stream interface, which is unchanged. Use one path per instance (the stream methods or the async methods, not both at once), and await calls in sequence to keep samples in order.
185
+
139
186
  ## Errors
140
187
 
141
188
  Construction errors come as typed classes you can catch:
@@ -182,6 +229,37 @@ The browser runs energy-mode VAD only, so its `vad` option accepts `false` or `'
182
229
  | Sample rate | Native device rate | Resampled from native rate |
183
230
  | VAD | `'silero'` or `'energy'` | `'energy'` only |
184
231
 
232
+ ### `new Speaker(options?)` (browser)
233
+
234
+ Browser audio playback through the Web Audio API. Playback is async (Promise based) and must be started from a user gesture so the browser allows audio.
235
+
236
+ ```javascript
237
+ import { Speaker } from 'decibri'; // browser entry via conditional export
238
+
239
+ const speaker = new Speaker({ sampleRate: 16000 });
240
+
241
+ playButton.onclick = async () => {
242
+ await speaker.write(int16Chunk); // Int16Array of PCM samples
243
+ await speaker.drain(); // resolves when playback finishes
244
+ speaker.stop();
245
+ };
246
+ ```
247
+
248
+ | Option | Type | Default | Description |
249
+ | --- | --- | --- | --- |
250
+ | `sampleRate` | number | 16000 | Sample rate of the audio you write (resampled to the output rate) |
251
+ | `channels` | number | 1 | Output channels (a mono stream plays on every channel) |
252
+ | `dtype` | `'int16'` \| `'float32'` | `'int16'` | Encoding of the samples you write |
253
+ | `workletUrl` | string | inline blob | Custom worklet URL for strict CSP |
254
+
255
+ - `start()` creates and resumes the audio output. Optional: `write()` starts it on the first call. Either must run in a user gesture (a click or tap); a context blocked by the autoplay policy surfaces a clear error.
256
+ - `write(chunk)` resolves when the samples are queued. It waits when the buffer is full, so awaiting it paces playback. Await calls sequentially to preserve order.
257
+ - `drain()` resolves when the queued audio has finished playing, immediately if nothing is queued.
258
+ - `stop()` halts immediately and discards anything queued.
259
+ - `isPlaying` reports whether audio is currently queued and playing.
260
+
261
+ To verify playback in real browsers, open `examples/browser-speaker-test.html` (see `examples/README.md`).
262
+
185
263
  ## Voice Activity Detection
186
264
 
187
265
  ### Energy mode
@@ -244,6 +322,8 @@ node node_modules/decibri/examples/websocket-server.js # terminal 1
244
322
  node node_modules/decibri/examples/websocket-stream.js # terminal 2
245
323
  ```
246
324
 
325
+ For the browser, `examples/browser-speaker-test.html` is a page for manually verifying audio playback in each browser. See `examples/README.md` for how to serve it on desktop and mobile.
326
+
247
327
  ## Migrating from 3.x
248
328
 
249
329
  decibri 4.0.0 renames the API to a microphone and speaker vocabulary and switches to named exports. See [MIGRATION.md](./MIGRATION.md) for a complete before-and-after guide.
@@ -0,0 +1,79 @@
1
+ # decibri examples
2
+
3
+ Runnable examples for decibri. The Node.js examples run with `node`; the browser
4
+ example is an HTML page you open in a browser.
5
+
6
+ ## Node.js examples
7
+
8
+ | File | What it does | Notes |
9
+ | --- | --- | --- |
10
+ | `wav-capture.js` | Captures audio and writes a valid WAV file | No extra dependencies |
11
+ | `websocket-server.js` | Receives raw PCM chunks and logs byte counts | Requires `npm install ws` |
12
+ | `websocket-stream.js` | Streams raw PCM audio to a WebSocket server | Requires `npm install ws` |
13
+
14
+ ```bash
15
+ node wav-capture.js
16
+ node websocket-server.js # terminal 1
17
+ node websocket-stream.js # terminal 2
18
+ ```
19
+
20
+ ## Browser audio playback test (`browser-speaker-test.html`)
21
+
22
+ A page for manually verifying that the browser Speaker plays real audio in each
23
+ browser. It loads the real browser build (`decibri.browser.js`, a bundle of the
24
+ package's browser entry) and plays generated tones so you can listen for a clean
25
+ signal, glitches, correct pitch after resampling, and immediate stop.
26
+
27
+ Browser audio needs a secure context, so the page must be served over
28
+ `localhost` or `https`. Opening it as a `file://` URL will not load the audio
29
+ engine.
30
+
31
+ ### Desktop
32
+
33
+ From this `examples` directory, serve it over localhost and open the printed URL
34
+ in each browser you want to test:
35
+
36
+ ```bash
37
+ npx serve .
38
+ # or
39
+ npx http-server . -p 8080
40
+ ```
41
+
42
+ Then open `http://localhost:<port>/browser-speaker-test.html` in Chrome, Firefox,
43
+ Edge, and Safari, and tap the buttons.
44
+
45
+ ### Mobile (same WiFi)
46
+
47
+ Serve on your computer as above, find your machine's LAN IP, and open
48
+ `http://<machine-ip>:<port>/browser-speaker-test.html` on the phone (on the same
49
+ network).
50
+
51
+ iOS Safari treats a plain `http://<LAN-IP>` address as an insecure context and
52
+ will refuse to start audio. For iOS, expose the local server over `https` with a
53
+ tunnel and open the tunnel URL on the phone:
54
+
55
+ ```bash
56
+ npx localtunnel --port 8080
57
+ # open the printed https://... URL on the phone, then add /browser-speaker-test.html
58
+ ```
59
+
60
+ ### What to check in each browser
61
+
62
+ - 440 Hz tone: a clean, steady tone with no buzz or distortion.
63
+ - Continuous (3 s): smooth, with no gaps, clicks, or dropouts.
64
+ - Resampled from 16 kHz: the same pitch as the 440 Hz tone button (confirms
65
+ resampling does not shift pitch).
66
+ - Stop: playback halts immediately when tapped mid-playback.
67
+ - Gesture requirement: nothing plays until you tap a button; if the browser
68
+ blocks audio, a clear error appears in the on-page log.
69
+ - Both formats: repeat with the int16 and float32 toggle.
70
+
71
+ ### Regenerating the browser bundle
72
+
73
+ `decibri.browser.js` is generated from the package's browser entry
74
+ (`../src/browser/index.js`). Regenerate it after changing the browser source
75
+ with any bundler that resolves the package's `browser` entry, for example:
76
+
77
+ ```bash
78
+ npx esbuild ../src/browser/index.js --bundle --format=iife --global-name=decibri --outfile=decibri.browser.js
79
+ ```
@@ -0,0 +1,159 @@
1
+ <!DOCTYPE html>
2
+ <html lang="en">
3
+ <head>
4
+ <meta charset="UTF-8">
5
+ <meta name="viewport" content="width=device-width, initial-scale=1.0">
6
+ <title>decibri browser audio playback test</title>
7
+ <style>
8
+ * { box-sizing: border-box; margin: 0; padding: 0; }
9
+ body { font-family: system-ui, sans-serif; max-width: 680px; margin: 40px auto; padding: 0 20px; color: #222; }
10
+ h1 { font-size: 1.4rem; margin-bottom: 8px; }
11
+ p.intro { font-size: 0.9rem; color: #555; margin-bottom: 16px; line-height: 1.5; }
12
+ .format-toggle { margin-bottom: 16px; font-size: 0.9rem; }
13
+ .format-toggle label { margin-right: 12px; cursor: pointer; }
14
+ .controls { display: flex; gap: 8px; flex-wrap: wrap; margin-bottom: 16px; }
15
+ button { padding: 10px 16px; font-size: 0.9rem; border: 1px solid #ccc; border-radius: 4px; cursor: pointer; background: #f8f8f8; }
16
+ button:hover { background: #e8e8e8; }
17
+ button.stop { border-color: #f44336; color: #c62828; }
18
+ .hint { display: block; font-size: 0.75rem; color: #888; margin-top: 2px; }
19
+ .stats { font-size: 0.85rem; color: #666; margin-bottom: 12px; }
20
+ #log { background: #1a1a1a; color: #0f0; font-family: monospace; font-size: 0.8rem; padding: 12px; border-radius: 4px; height: 320px; overflow-y: auto; white-space: pre-wrap; word-break: break-all; }
21
+ </style>
22
+ </head>
23
+ <body>
24
+ <h1>decibri browser audio playback test</h1>
25
+ <p class="intro">
26
+ Tap a button to play audio through the browser Speaker, then listen and watch the log.
27
+ Playback must start from a tap (the browser only allows audio after a user gesture), so
28
+ nothing plays until you tap. Serve this page over <strong>localhost</strong> or
29
+ <strong>https</strong>; opening it as a <code>file://</code> URL will not work. See
30
+ <code>examples/README.md</code> for serve and mobile instructions.
31
+ </p>
32
+
33
+ <div class="format-toggle">
34
+ <strong>Format:</strong>
35
+ <label><input type="radio" name="dtype" value="int16" checked> int16</label>
36
+ <label><input type="radio" name="dtype" value="float32"> float32</label>
37
+ </div>
38
+
39
+ <div class="controls">
40
+ <div>
41
+ <button id="btn-tone">Play 440 Hz tone</button>
42
+ <span class="hint">about 1.5 s. Should be a clean, steady tone.</span>
43
+ </div>
44
+ <div>
45
+ <button id="btn-continuous">Play 3 s continuous</button>
46
+ <span class="hint">Listen for any gaps, clicks, or dropouts.</span>
47
+ </div>
48
+ <div>
49
+ <button id="btn-resampled">Play 440 Hz from 16 kHz</button>
50
+ <span class="hint">Forces resampling. Should match the first button's pitch.</span>
51
+ </div>
52
+ <div>
53
+ <button id="btn-stop" class="stop">Stop</button>
54
+ <span class="hint">Should halt immediately.</span>
55
+ </div>
56
+ </div>
57
+
58
+ <div class="stats" id="stats">Idle.</div>
59
+ <div id="log"></div>
60
+
61
+ <!--
62
+ This page loads the real browser Speaker from examples/decibri.browser.js, a
63
+ generated bundle of the package's browser entry (src/browser/index.js). It is
64
+ the actual shipped code, not a copy. See examples/README.md to regenerate it.
65
+ -->
66
+ <script src="./decibri.browser.js"></script>
67
+ <script>
68
+ const { Speaker } = decibri;
69
+
70
+ const logEl = document.getElementById('log');
71
+ const statsEl = document.getElementById('stats');
72
+
73
+ let current = null;
74
+
75
+ function log(msg) {
76
+ const ts = new Date().toISOString().slice(11, 23);
77
+ logEl.textContent += `[${ts}] ${msg}\n`;
78
+ logEl.scrollTop = logEl.scrollHeight;
79
+ }
80
+
81
+ function setStats(text) {
82
+ const playing = current ? current.isPlaying : false;
83
+ statsEl.textContent = `${text} | isPlaying: ${playing}`;
84
+ }
85
+
86
+ function getDtype() {
87
+ return document.querySelector('input[name="dtype"]:checked').value;
88
+ }
89
+
90
+ // Generate a sine tone at `rate` Hz with short fades to avoid edge clicks.
91
+ function genTone(freq, seconds, rate, dtype) {
92
+ const n = Math.floor(rate * seconds);
93
+ const amp = 0.25;
94
+ const fade = Math.min(Math.floor(rate * 0.005), Math.floor(n / 2));
95
+ const f32 = new Float32Array(n);
96
+ for (let i = 0; i < n; i++) {
97
+ let env = 1;
98
+ if (i < fade) env = i / fade;
99
+ else if (i >= n - fade) env = (n - 1 - i) / fade;
100
+ f32[i] = Math.sin(2 * Math.PI * freq * i / rate) * amp * env;
101
+ }
102
+ if (dtype === 'float32') return f32;
103
+ const i16 = new Int16Array(n);
104
+ for (let i = 0; i < n; i++) {
105
+ i16[i] = Math.max(-32768, Math.min(32767, Math.round(f32[i] * 32767)));
106
+ }
107
+ return i16;
108
+ }
109
+
110
+ function stopCurrent() {
111
+ if (current) {
112
+ current.stop();
113
+ log('stop()');
114
+ current = null;
115
+ setStats('Stopped.');
116
+ }
117
+ }
118
+
119
+ async function play(label, freq, seconds, srcRate) {
120
+ stopCurrent();
121
+ const dtype = getDtype();
122
+ log(`${label}: ${seconds}s of ${freq} Hz, source ${srcRate} Hz, dtype=${dtype}`);
123
+ setStats('Starting...');
124
+ const speaker = new Speaker({ sampleRate: srcRate, dtype });
125
+ current = speaker;
126
+ try {
127
+ const tone = genTone(freq, seconds, srcRate, dtype);
128
+ await speaker.write(tone);
129
+ log(`queued ${tone.length} samples, waiting for playback to finish...`);
130
+ setStats('Playing...');
131
+ await speaker.drain();
132
+ log('drained (playback finished)');
133
+ setStats('Finished.');
134
+ // Only tear down if a newer play has not replaced this one.
135
+ if (current === speaker) {
136
+ speaker.stop();
137
+ current = null;
138
+ }
139
+ } catch (err) {
140
+ log('ERROR: ' + (err && err.message ? err.message : String(err)));
141
+ setStats('Error (see log).');
142
+ }
143
+ }
144
+
145
+ document.getElementById('btn-tone').addEventListener('click', () => {
146
+ play('440 Hz tone', 440, 1.5, 44100);
147
+ });
148
+ document.getElementById('btn-continuous').addEventListener('click', () => {
149
+ play('continuous', 440, 3, 44100);
150
+ });
151
+ document.getElementById('btn-resampled').addEventListener('click', () => {
152
+ play('resampled from 16 kHz', 440, 1.5, 16000);
153
+ });
154
+ document.getElementById('btn-stop').addEventListener('click', stopCurrent);
155
+
156
+ log('Ready. Tap a button to play.');
157
+ </script>
158
+ </body>
159
+ </html>