nixamp 0.26.2 → 0.26.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +14 -2
- package/package.json +1 -1
- package/web/dist/assets/background.worker-D9ejvk_o.js +547 -0
- package/web/dist/assets/{hls-3VKVEQE3-c1jFw5OF.js → hls-3VKVEQE3-BJCEkhkX.js} +1 -1
- package/web/dist/assets/index-BZ57qkpe.js +1 -0
- package/web/dist/assets/{mpegts-Tk8uQ0GA.js → mpegts-DaLDFdZV.js} +1 -1
- package/web/dist/assets/{mpegts-LO6RVLD6-DWFBhjSm.js → mpegts-LO6RVLD6-DLNvQAAS.js} +1 -1
- package/web/dist/index.html +4 -2
- package/web/dist/licenses/fastenhancer-web.txt +21 -0
- package/web/dist/sw.js +7 -5
- package/web/dist/assets/index-BhAjOmgV.js +0 -1
package/README.md
CHANGED
|
@@ -558,7 +558,7 @@ A rolling 15-second audio window advances every 5 seconds. Speaker labels are
|
|
|
558
558
|
reconciled using overlapping timestamps, with different voices assigned to
|
|
559
559
|
separate speakers. Voices are picked from the available stock catalogue without
|
|
560
560
|
inferring a person's gender from pitch; each detected speaker gets an unused
|
|
561
|
-
voice until the catalogue is exhausted. **
|
|
561
|
+
voice until the catalogue is exhausted. **Audio options** folds away optional
|
|
562
562
|
individual overrides. A speaker returning after leaving the rolling context may
|
|
563
563
|
receive a new label. Simultaneous speech and noisy crowds can still confuse
|
|
564
564
|
recognition. Native captions never translate to English as
|
|
@@ -568,7 +568,19 @@ Processing has one active request and only the latest pending window per
|
|
|
568
568
|
listener; speech queues and response sizes are bounded. Old transcript history
|
|
569
569
|
is never spoken. Pause, seek, source changes, and disabling the feature cancel
|
|
570
570
|
queued speech; errors restore the original audio. This is a delayed live
|
|
571
|
-
interpreter, not a promise of exact lip sync
|
|
571
|
+
interpreter, not a promise of exact lip sync.
|
|
572
|
+
|
|
573
|
+
Translated playback keeps an approximate version of the original background
|
|
574
|
+
sound. FastEnhancer Web's Tiny model estimates speech locally in a dedicated
|
|
575
|
+
browser worker; the player subtracts that estimate from the aligned source in
|
|
576
|
+
each stereo channel and mixes the remainder with translated voices. Background
|
|
577
|
+
processing adds no API calls or provider charges. It stops with translation;
|
|
578
|
+
recognition starts without waiting for it. If the device cannot keep up or load the model, that
|
|
579
|
+
background branch is silenced while translated speech continues. Separation can
|
|
580
|
+
leave some original speech or remove parts of music and crowd noise; disable
|
|
581
|
+
**Keep background sound** under **Audio options** when needed. This uses
|
|
582
|
+
[FastEnhancer Web](https://github.com/ryyr-ry/fastenhancer-web), under the MIT
|
|
583
|
+
license.
|
|
572
584
|
|
|
573
585
|
The account server needs `ELEVENLABS_API_KEY`; `NIXAMP_DUBBING=off` disables
|
|
574
586
|
this feature. The key stays on the server. Sign-in is required for speaker
|