nixamp 0.26.2 → 0.26.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -558,7 +558,7 @@ A rolling 15-second audio window advances every 5 seconds. Speaker labels are
558
558
  reconciled using overlapping timestamps, with different voices assigned to
559
559
  separate speakers. Voices are picked from the available stock catalogue without
560
560
  inferring a person's gender from pitch; each detected speaker gets an unused
561
- voice until the catalogue is exhausted. **Speaker voices** folds away optional
561
+ voice until the catalogue is exhausted. **Audio options** folds away optional
562
562
  individual overrides. A speaker returning after leaving the rolling context may
563
563
  receive a new label. Simultaneous speech and noisy crowds can still confuse
564
564
  recognition. Native captions never translate to English as
@@ -568,7 +568,19 @@ Processing has one active request and only the latest pending window per
568
568
  listener; speech queues and response sizes are bounded. Old transcript history
569
569
  is never spoken. Pause, seek, source changes, and disabling the feature cancel
570
570
  queued speech; errors restore the original audio. This is a delayed live
571
- interpreter, not a promise of exact lip sync or background-music separation.
571
+ interpreter, not a promise of exact lip sync.
572
+
573
+ Translated playback keeps an approximate version of the original background
574
+ sound. FastEnhancer Web's Tiny model estimates speech locally in a dedicated
575
+ browser worker; the player subtracts that estimate from the aligned source in
576
+ each stereo channel and mixes the remainder with translated voices. Background
577
+ processing adds no API calls or provider charges. It stops with translation;
578
+ recognition starts without waiting for it. If the device cannot keep up or load the model, that
579
+ background branch is silenced while translated speech continues. Separation can
580
+ leave some original speech or remove parts of music and crowd noise; disable
581
+ **Keep background sound** under **Audio options** when needed. This uses
582
+ [FastEnhancer Web](https://github.com/ryyr-ry/fastenhancer-web), under the MIT
583
+ license.
572
584
 
573
585
  The account server needs `ELEVENLABS_API_KEY`; `NIXAMP_DUBBING=off` disables
574
586
  this feature. The key stays on the server. Sign-in is required for speaker
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "nixamp",
3
- "version": "0.26.2",
3
+ "version": "0.26.4",
4
4
  "description": "It really whips the terminal's ass. A Winamp-shaped audio player for your terminal.",
5
5
  "license": "MIT",
6
6
  "type": "module",