echogarden 1.0.4 → 1.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +26 -23
- package/data/schemas/options.json +177 -36
- package/dist/alignment/SpeechAlignment.d.ts +1 -1
- package/dist/alignment/SpeechAlignment.js +1 -1
- package/dist/alignment/SpeechAlignment.js.map +1 -1
- package/dist/api/API.d.ts +1 -0
- package/dist/api/API.js +1 -0
- package/dist/api/API.js.map +1 -1
- package/dist/api/APIOptions.d.ts +1 -0
- package/dist/api/Alignment.d.ts +3 -3
- package/dist/api/Alignment.js +5 -10
- package/dist/api/Alignment.js.map +1 -1
- package/dist/api/LanguageDetection.d.ts +5 -7
- package/dist/api/LanguageDetection.js +3 -2
- package/dist/api/LanguageDetection.js.map +1 -1
- package/dist/api/Recognition.d.ts +4 -5
- package/dist/api/Recognition.js +5 -8
- package/dist/api/Recognition.js.map +1 -1
- package/dist/api/SourceSeparation.d.ts +2 -0
- package/dist/api/SourceSeparation.js +4 -2
- package/dist/api/SourceSeparation.js.map +1 -1
- package/dist/api/Synthesis.d.ts +3 -1
- package/dist/api/Synthesis.js +9 -10
- package/dist/api/Synthesis.js.map +1 -1
- package/dist/api/Translation.d.ts +1 -1
- package/dist/api/Translation.js +4 -8
- package/dist/api/Translation.js.map +1 -1
- package/dist/api/TranslationAlignment.d.ts +31 -0
- package/dist/api/TranslationAlignment.js +121 -0
- package/dist/api/TranslationAlignment.js.map +1 -0
- package/dist/api/VoiceActivityDetection.d.ts +5 -1
- package/dist/api/VoiceActivityDetection.js +38 -2
- package/dist/api/VoiceActivityDetection.js.map +1 -1
- package/dist/audio/AudioPlayer.js +6 -1
- package/dist/audio/AudioPlayer.js.map +1 -1
- package/dist/cli/CLI.js +85 -0
- package/dist/cli/CLI.js.map +1 -1
- package/dist/dsp/FFT.js.map +1 -1
- package/dist/math/MedianFilter.d.ts +5 -0
- package/dist/math/MedianFilter.js +102 -0
- package/dist/math/MedianFilter.js.map +1 -0
- package/dist/math/VectorMath.d.ts +0 -2
- package/dist/math/VectorMath.js +1 -25
- package/dist/math/VectorMath.js.map +1 -1
- package/dist/recognition/OpenAICloudSTT.d.ts +1 -1
- package/dist/recognition/OpenAICloudSTT.js.map +1 -1
- package/dist/recognition/SileroSTT.d.ts +22 -1
- package/dist/recognition/SileroSTT.js +122 -95
- package/dist/recognition/SileroSTT.js.map +1 -1
- package/dist/recognition/WhisperCppSTT.js +1 -1
- package/dist/recognition/WhisperCppSTT.js.map +1 -1
- package/dist/recognition/WhisperSTT.d.ts +52 -19
- package/dist/recognition/WhisperSTT.js +645 -494
- package/dist/recognition/WhisperSTT.js.map +1 -1
- package/dist/server/Server.js.map +1 -1
- package/dist/source-separation/MDXNetSourceSeparation.d.ts +5 -3
- package/dist/source-separation/MDXNetSourceSeparation.js +26 -19
- package/dist/source-separation/MDXNetSourceSeparation.js.map +1 -1
- package/dist/speech-language-detection/SileroLanguageDetection.d.ts +15 -9
- package/dist/speech-language-detection/SileroLanguageDetection.js +23 -16
- package/dist/speech-language-detection/SileroLanguageDetection.js.map +1 -1
- package/dist/synthesis/EspeakTTS.js +4 -0
- package/dist/synthesis/EspeakTTS.js.map +1 -1
- package/dist/synthesis/GoogleCloudTTS.js.map +1 -1
- package/dist/synthesis/VitsTTS.d.ts +8 -6
- package/dist/synthesis/VitsTTS.js +36 -31
- package/dist/synthesis/VitsTTS.js.map +1 -1
- package/dist/tests/Test.js.map +1 -1
- package/dist/utilities/OnnxUtilities.d.ts +14 -0
- package/dist/utilities/OnnxUtilities.js +43 -0
- package/dist/utilities/OnnxUtilities.js.map +1 -0
- package/dist/utilities/Utilities.d.ts +4 -8
- package/dist/utilities/Utilities.js +35 -58
- package/dist/utilities/Utilities.js.map +1 -1
- package/dist/voice-activity-detection/SileroVAD.d.ts +5 -3
- package/dist/voice-activity-detection/SileroVAD.js +9 -11
- package/dist/voice-activity-detection/SileroVAD.js.map +1 -1
- package/docs/API.md +54 -34
- package/docs/CLI.md +25 -13
- package/docs/Contributing.md +4 -2
- package/docs/Engines.md +43 -32
- package/docs/Licenses.md +3 -4
- package/docs/Options.md +47 -11
- package/docs/Releases.md +4 -0
- package/docs/Server.md +8 -6
- package/docs/Tasklist.md +39 -52
- package/docs/Technical.md +1 -1
- package/package.json +8 -12
- package/src/alignment/SpeechAlignment.ts +1 -1
- package/src/api/API.ts +1 -0
- package/src/api/APIOptions.ts +1 -0
- package/src/api/Alignment.ts +10 -14
- package/src/api/LanguageDetection.ts +14 -10
- package/src/api/Recognition.ts +17 -10
- package/src/api/SourceSeparation.ts +7 -2
- package/src/api/Synthesis.ts +26 -11
- package/src/api/Translation.ts +14 -8
- package/src/api/TranslationAlignment.ts +213 -0
- package/src/api/VoiceActivityDetection.ts +66 -3
- package/src/audio/AudioPlayer.ts +6 -2
- package/src/cli/CLI.ts +121 -2
- package/src/dsp/FFT.ts +3 -0
- package/src/math/MedianFilter.ts +124 -0
- package/src/math/VectorMath.ts +1 -36
- package/src/recognition/OpenAICloudSTT.ts +27 -27
- package/src/recognition/SileroSTT.ts +149 -102
- package/src/recognition/WhisperCppSTT.ts +1 -1
- package/src/recognition/WhisperSTT.ts +961 -684
- package/src/server/Server.ts +1 -1
- package/src/source-separation/MDXNetSourceSeparation.ts +35 -19
- package/src/speech-language-detection/SileroLanguageDetection.ts +53 -33
- package/src/synthesis/EspeakTTS.ts +8 -0
- package/src/synthesis/GoogleCloudTTS.ts +12 -1
- package/src/synthesis/VitsTTS.ts +57 -46
- package/src/tests/Test.ts +1 -1
- package/src/utilities/OnnxUtilities.ts +68 -0
- package/src/utilities/Utilities.ts +38 -66
- package/src/voice-activity-detection/SileroVAD.ts +15 -15
- package/dist/utilities/NdArrayUtilities.d.ts +0 -3
- package/dist/utilities/NdArrayUtilities.js +0 -23
- package/dist/utilities/NdArrayUtilities.js.map +0 -1
- package/src/utilities/NdArrayUtilities.ts +0 -31
package/docs/Engines.md
CHANGED
|
@@ -5,17 +5,17 @@
|
|
|
5
5
|
|
|
6
6
|
**Offline**:
|
|
7
7
|
|
|
8
|
-
* [VITS](https://github.com/jaywalnut310/vits) (`vits`):
|
|
9
|
-
* [SVOX Pico](https://github.com/naggety/picotts) (`pico`): a legacy diphone-based synthesis engine. Supports English (US, UK), Spanish, Italian, French, and German
|
|
10
|
-
* [Flite](https://github.com/festvox/flite) (`flite`): a legacy diphone-based synthesis engine. Supports English (US, Scottish), and several Indic languages: Hindi, Bengali, Marathi, Telugu, Tamil, Gujarati, Kannada and Punjabi
|
|
11
|
-
* [eSpeak-NG](https://github.com/espeak-ng/espeak-ng/) (`espeak`): a lightweight "robot" sounding formant-based synthesizer. Supports 100+ languages.
|
|
12
|
-
* [SAM (Software Automatic Mouth)](https://github.com/discordier/sam) (`sam`): a classic "robot" speech synthesizer from 1982. English only
|
|
8
|
+
* [VITS](https://github.com/jaywalnut310/vits) (`vits`): end-to-end neural speech synthesis architecture. Available models were trained by Michael Hansen as part of his [Piper speech synthesis system](https://github.com/rhasspy/piper). Currently, there are 117 voices, in a range of languages, including English (US, UK), Spanish (ES, MX), Portuguese (PT, BR), Italian, French, German, Dutch (NL, BE), Swedish, Norwegian, Danish, Finnish, Polish, Greek, Romanian, Serbian, Czech, Hungarian, Slovak, Slovenian, Turkish, Arabic, Farsi, Russian, Ukrainian, Catalan, Luxembourgish, Icelandic, Swahili, Kazakh, Georgian, Nepali, Vietnamese and Chinese. You can listen to audio samples of all voices and languages in [Piper's samples page](https://rhasspy.github.io/piper-samples/)
|
|
9
|
+
* [SVOX Pico](https://github.com/naggety/picotts) (`pico`): a legacy diphone-based synthesis engine. Supports English (US, UK), Spanish, Italian, French, and German
|
|
10
|
+
* [Flite](https://github.com/festvox/flite) (`flite`): a legacy diphone-based synthesis engine. Supports English (US, Scottish), and several Indic languages: Hindi, Bengali, Marathi, Telugu, Tamil, Gujarati, Kannada and Punjabi
|
|
11
|
+
* [eSpeak-NG](https://github.com/espeak-ng/espeak-ng/) (`espeak`): a lightweight "robot" sounding formant-based synthesizer. Supports 100+ languages. Used internally for speech alignment, phonemization, and other internal tasks
|
|
12
|
+
* [SAM (Software Automatic Mouth)](https://github.com/discordier/sam) (`sam`): a classic "robot" speech synthesizer from 1982. English only
|
|
13
13
|
|
|
14
14
|
**Offline, Windows only**:
|
|
15
15
|
|
|
16
|
-
* [SAPI](https://en.wikipedia.org/wiki/Microsoft_Speech_API) (`sapi`): Microsoft Speech API. Supports the system's language voices, as well as legacy voices produced by third-party vendors, like Ivona, NeoSpeech, Acapela, Cepstral, CereProc, Nuance, AT&T, Loquendo, ScanSoft and others (note that only 64-bit SAPI voices are supported, which makes it incompatible with a significant portion of older voices)
|
|
16
|
+
* [SAPI](https://en.wikipedia.org/wiki/Microsoft_Speech_API) (`sapi`): Microsoft Speech API. Supports the system's language voices, as well as legacy voices produced by third-party vendors, like Ivona, NeoSpeech, Acapela, Cepstral, CereProc, Nuance, AT&T, Loquendo, ScanSoft and others (note that only 64-bit SAPI voices are supported, which makes it incompatible with a significant portion of older voices)
|
|
17
17
|
|
|
18
|
-
* [Microsoft Speech Platform](https://www.microsoft.com/en-us/download/details.aspx?id=27225) (`msspeech`): Microsoft Server Speech API. Requires [installing a runtime (2.6MB)](https://www.microsoft.com/en-us/download/details.aspx?id=27225). Supports 28 dialects, which can be individually downloaded via [freely available installers](https://www.microsoft.com/en-us/download/details.aspx?id=27224), or, for convenience, bundled as [a single 358MB zip file](https://drive.google.com/u/0/uc?id=1uQdFNxLzUxpaEwVVKhMawys8cIh3F21T&export=download). Has voices for English (US, UK, AU, CA), Spanish (ES, MX), Portuguese (BR, PT), German, French (FR, CA), Italian, Norwegian, Dutch, Russian, Swedish, Danish, Catalan, Finnish, Japanese, Korean and Chinese (ZH, HK, TW). All voices are female
|
|
18
|
+
* [Microsoft Speech Platform](https://www.microsoft.com/en-us/download/details.aspx?id=27225) (`msspeech`): Microsoft Server Speech API. Requires [installing a runtime (2.6MB)](https://www.microsoft.com/en-us/download/details.aspx?id=27225). Supports 28 dialects, which can be individually downloaded via [freely available installers](https://www.microsoft.com/en-us/download/details.aspx?id=27224), or, for convenience, bundled as [a single 358MB zip file](https://drive.google.com/u/0/uc?id=1uQdFNxLzUxpaEwVVKhMawys8cIh3F21T&export=download). Has voices for English (US, UK, AU, CA), Spanish (ES, MX), Portuguese (BR, PT), German, French (FR, CA), Italian, Norwegian, Dutch, Russian, Swedish, Danish, Catalan, Finnish, Japanese, Korean and Chinese (ZH, HK, TW). All voices are female
|
|
19
19
|
|
|
20
20
|
**Note**: both these engines require manually installing the [`winax` npm package](https://www.npmjs.com/package/winax) by running `npm install winax -g`.
|
|
21
21
|
|
|
@@ -30,71 +30,82 @@
|
|
|
30
30
|
These are commercial services that require a subscription and an API key to use:
|
|
31
31
|
|
|
32
32
|
* [Google Cloud](https://cloud.google.com/text-to-speech) (`google-cloud`)
|
|
33
|
-
* [Azure Cognitive Services](https://azure.microsoft.com/en-us/products/
|
|
33
|
+
* [Azure Cognitive Services](https://azure.microsoft.com/en-us/products/ai-services/text-to-speech/) (`microsoft-azure`)
|
|
34
34
|
* [Amazon Polly](https://aws.amazon.com/polly/) (`amazon-polly`)
|
|
35
|
-
* [OpenAI Cloud](https://platform.openai.com/) (`openai-cloud`)
|
|
35
|
+
* [OpenAI Cloud Platform](https://platform.openai.com/) (`openai-cloud`)
|
|
36
36
|
* [Elevenlabs](https://elevenlabs.io/) (`elevenlabs`)
|
|
37
37
|
|
|
38
38
|
**Cloud services (unofficial)**:
|
|
39
39
|
|
|
40
40
|
These cloud-based engines connect to public cloud APIs that are not officially publicized by their operators. They are included for educational purposes only, and may be removed in the future:
|
|
41
41
|
|
|
42
|
-
* Google Translate (`google-translate`): used by the [Google Translate web UI](https://translate.google.com/) to speak written text in any one of its supported languages. Offers a single voice for each language (usually female)
|
|
43
|
-
* Microsoft Edge (`microsoft-edge`): subset of the Azure Cognitive Services cloud TTS API used by the Microsoft Edge browser as part of its support for the [Web Speech API](https://developer.mozilla.org/en-US/docs/Web/API/Web_Speech_API) and its [Read Aloud](https://www.microsoft.com/en-us/edge/features/read-aloud?form=MT00D8) feature. Using this engine requires a special token, which should be passed via the `microsoftEdge.trustedClientToken` option
|
|
44
|
-
* Streamlabs Polly (`streamlabs-polly`): a public REST API by Streamlabs, primarily intended for generating speech for TTS donations. It includes a few English (US, UK, AU, IN) voices, which are similar to some of the non-neural (Ivona-based) voices offered by Amazon Polly (**Note**: as of April 2024, the public Streamlabs Polly REST API doesn't seem to be accessible anymore)
|
|
42
|
+
* Google Translate (`google-translate`): used by the [Google Translate web UI](https://translate.google.com/) to speak written text in any one of its supported languages. Offers a single voice for each language (usually female)
|
|
43
|
+
* Microsoft Edge (`microsoft-edge`): subset of the Azure Cognitive Services cloud TTS API used by the Microsoft Edge browser as part of its support for the [Web Speech API](https://developer.mozilla.org/en-US/docs/Web/API/Web_Speech_API) and its [Read Aloud](https://www.microsoft.com/en-us/edge/features/read-aloud?form=MT00D8) feature. Using this engine requires a special token, which should be passed via the `microsoftEdge.trustedClientToken` option
|
|
44
|
+
* Streamlabs Polly (`streamlabs-polly`): a public REST API by Streamlabs, primarily intended for generating speech for TTS donations. It includes a few English (US, UK, AU, IN) voices, which are similar to some of the non-neural (Ivona-based) voices offered by Amazon Polly (**Note**: as of April 2024, the public Streamlabs Polly REST API doesn't seem to be accessible anymore)
|
|
45
45
|
|
|
46
46
|
## Speech-to-text
|
|
47
47
|
|
|
48
48
|
**Offline**:
|
|
49
|
-
* [OpenAI Whisper](https://github.com/openai/whisper) (`whisper`): high
|
|
50
|
-
* [Whisper.cpp](https://github.com/ggerganov/whisper.cpp) (`whisper.cpp`): a port of the Whisper architecture
|
|
51
|
-
* [Vosk](https://github.com/alphacep/vosk-api) (`vosk`): models available for 25+ languages. **Note**: the Vosk package is not included in the default installation, but you can add support for it using `npm install @echogarden/vosk -g`. Then, you'll need to manually [download a model](https://alphacephei.com/vosk/models) and specify its directory path via the `vosk.modelPath` option
|
|
52
|
-
* [Silero](https://github.com/snakers4/silero-models) (`silero`): models available for English, Spanish, German and Ukrainian. For [non-commercial use only](https://github.com/snakers4/silero-models/blob/master/LICENSE)
|
|
49
|
+
* [OpenAI Whisper](https://github.com/openai/whisper) (`whisper`): high-accuracy transformer-based speech recognition architecture. TypeScript implementation, with inference done via the [ONNX runtime](https://onnxruntime.ai/). Supports [98 languages](https://platform.openai.com/docs/guides/speech-to-text/supported-languages). There are several models of different sizes, some are multilingual, and some are English only: `tiny`, `tiny.en`, `base`, `base.en`, `small`, `small.en`, `medium`, `medium.en`, `large`, `large-v1` and `large-v2`, `large-v3`. **Note**: large models are not currently supported by `onnxruntime-node` due to model size restrictions
|
|
50
|
+
* [Whisper.cpp](https://github.com/ggerganov/whisper.cpp) (`whisper.cpp`): a C++ port of the Whisper architecture by Georgi Gerganov. Supports all Whisper models, including several quantized ones (see full model list in the [options reference](docs/Options.md)). Has various builds, including CUDA and OpenCL for GPU support
|
|
51
|
+
* [Vosk](https://github.com/alphacep/vosk-api) (`vosk`): models available for 25+ languages. **Note**: the Vosk package is not included in the default installation, but you can add support for it using `npm install @echogarden/vosk -g`. Then, you'll need to manually [download a model](https://alphacephei.com/vosk/models) and specify its directory path via the `vosk.modelPath` option
|
|
52
|
+
* [Silero](https://github.com/snakers4/silero-models) (`silero`): models available for English, Spanish, German and Ukrainian. For [non-commercial use only](https://github.com/snakers4/silero-models/blob/master/LICENSE)
|
|
53
53
|
|
|
54
54
|
**Cloud services**:
|
|
55
55
|
|
|
56
56
|
These are commercial services that require a subscription and an API key to use:
|
|
57
57
|
|
|
58
58
|
* [Google Cloud](https://cloud.google.com/speech-to-text) (`google-cloud`)
|
|
59
|
-
* [Azure Cognitive Services](https://azure.microsoft.com/en-us/products/
|
|
59
|
+
* [Azure Cognitive Services](https://azure.microsoft.com/en-us/products/ai-services/speech-to-text/) (`microsoft-azure`)
|
|
60
60
|
* [Amazon Transcribe](https://aws.amazon.com/transcribe/) (`amazon-transcribe`)
|
|
61
|
-
* [OpenAI Cloud](https://platform.openai.com/) (`openai-cloud`)
|
|
61
|
+
* [OpenAI Cloud Platform](https://platform.openai.com/) (`openai-cloud`): runs the `large-v2` Whisper model on the cloud
|
|
62
62
|
|
|
63
63
|
## Speech-to-transcript alignment
|
|
64
64
|
|
|
65
65
|
These engines' goal is to match (or "align") a given spoken recording with a given transcript as closely as possible. They will annotate each word in the transcript with approximate start and end timestamps:
|
|
66
66
|
|
|
67
|
-
* Dynamic Time Warping (`dtw`): transcript is first synthesized using the eSpeak engine, then the [DTW](https://en.wikipedia.org/wiki/Dynamic_time_warping) alignment algorithm is applied to find the best mapping between the synthesized and original audio frames
|
|
68
|
-
* Dynamic Time Warping with Recognition Assist (`dtw-ra`): recognition is applied to the audio (any recognition engine can be used), then both the ground-truth transcript and the recognized transcript are synthesized using eSpeak. Then, the best mapping is found between the two synthesized waveforms, and the result is
|
|
69
|
-
* Whisper-based alignment (`whisper`): transcript is tokenized
|
|
67
|
+
* Dynamic Time Warping (`dtw`): transcript is first synthesized using the eSpeak engine, then the [DTW](https://en.wikipedia.org/wiki/Dynamic_time_warping) sequence alignment algorithm is applied to find the best mapping between the synthesized and original audio frames
|
|
68
|
+
* Dynamic Time Warping with Recognition Assist (`dtw-ra`): recognition is applied to the audio (any recognition engine can be used), then both the ground-truth transcript and the recognized transcript are synthesized using eSpeak. Then, the best mapping is found between the two synthesized waveforms, using the DTW algorithm, and the result is remapped back to the original audio using the timing information produced by the recognizer
|
|
69
|
+
* Whisper-based alignment (`whisper`): transcript is first tokenized, then, its tokens are decoded, in order, with a guided approach, using the Whisper model. The resulting token timestamps are then used to derive the timing for each word
|
|
70
|
+
|
|
70
71
|
|
|
71
72
|
## Speech-to-text translation
|
|
72
73
|
|
|
73
|
-
|
|
74
|
+
**Offline**:
|
|
75
|
+
* [Whisper](https://github.com/openai/whisper) (`whisper`): the Whisper model can recognize speech in any one of its supported languages and output a transcript directly translated to English. Other languages are not supported as targets
|
|
74
76
|
* [Whisper.cpp](https://github.com/ggerganov/whisper.cpp) (`whisper.cpp`): supports translation to English
|
|
75
|
-
|
|
77
|
+
|
|
78
|
+
**Cloud services**:
|
|
79
|
+
* [OpenAI Cloud Platform](https://platform.openai.com/) (`openai-cloud`): runs the `large-v2` Whisper model on the cloud. Only supports English as target
|
|
80
|
+
|
|
81
|
+
## Speech-to-translated-transcript alignment
|
|
82
|
+
|
|
83
|
+
These goal here is to match (or "align") a given spoken recording in one language, with a given translated transcript in a different language, as closely as possible.
|
|
84
|
+
|
|
85
|
+
* `whisper`: given a spoken recording in any of the [98 languages](https://platform.openai.com/docs/guides/speech-to-text/supported-languages) supported by Whisper, and an English translation of its transcript, the translated transcript is tokenized and then decoded, in order, using a guided approach, with any multilingual Whisper model, set to its `translate` task mode. In this way, the approximate mapping between the spoken audio and each word of the translation is estimated
|
|
86
|
+
|
|
76
87
|
|
|
77
88
|
## Language detection
|
|
78
89
|
|
|
79
90
|
**Spoken language detection**:
|
|
80
|
-
* [
|
|
81
|
-
* [
|
|
91
|
+
* [Whisper](https://github.com/openai/whisper) (`whisper`): uses the language token produced by the `whisper` speech recognition model to generate a set of probabilities for the [98 languages](https://platform.openai.com/docs/guides/speech-to-text/supported-languages) it has been trained on
|
|
92
|
+
* [Silero Language Classifier](https://github.com/snakers4/silero-vad/wiki/Other-Models) (`silero`): a speech language classification model by Silero
|
|
82
93
|
|
|
83
94
|
**Text language detection**:
|
|
84
|
-
* [TinyLD](https://www.npmjs.com/package/tinyld) (`tinyld`): a simple language detection library
|
|
85
|
-
* [FastText](https://github.com/facebookresearch/fastText) (`fasttext`): a library for word representations and sentence classification by Facebook research
|
|
95
|
+
* [TinyLD](https://www.npmjs.com/package/tinyld) (`tinyld`): a simple language detection library
|
|
96
|
+
* [FastText](https://github.com/facebookresearch/fastText) (`fasttext`): a library for word representations and sentence classification by Facebook research
|
|
86
97
|
|
|
87
98
|
## Voice activity detection
|
|
88
99
|
|
|
89
|
-
* [WebRTC VAD](https://github.com/dpirch/libfvad) (`webrtc`): a voice activity detector. Originally from the Chromium browser source code
|
|
100
|
+
* [WebRTC VAD](https://github.com/dpirch/libfvad) (`webrtc`): a voice activity detector. Originally from the Chromium browser source code
|
|
90
101
|
* [Silero VAD](https://github.com/snakers4/silero-vad) (`silero`): a voice activity detection model by Silero.
|
|
91
|
-
* [RNNoise](https://github.com/xiph/rnnoise) (`rnnoise`): uses RNNoise's speech probabilities output for each audio frame as a VAD metric
|
|
92
|
-
* Adaptive Gate (`adaptive-gate`): uses a band-limited adaptive gate to identify activity in the lower voice frequencies. Reliable, but will often pass non-vocal sounds if they are loud enough. Good for clean speech and a cappella singing, where most non-vocal segments are quiet
|
|
102
|
+
* [RNNoise](https://github.com/xiph/rnnoise) (`rnnoise`): uses RNNoise's speech probabilities output for each audio frame as a VAD metric
|
|
103
|
+
* Adaptive Gate (`adaptive-gate`): uses a band-limited adaptive gate to identify activity in the lower voice frequencies. Reliable, but will often pass non-vocal sounds if they are loud enough. Good for clean speech and a cappella singing, where most non-vocal segments are quiet
|
|
93
104
|
|
|
94
105
|
## Speech denoising
|
|
95
106
|
|
|
96
|
-
* [RNNoise](https://github.com/xiph/rnnoise) (`rnnoise`): a noise suppression library based on a recurrent neural network
|
|
107
|
+
* [RNNoise](https://github.com/xiph/rnnoise) (`rnnoise`): a noise suppression library based on a recurrent neural network
|
|
97
108
|
|
|
98
109
|
## Source separation
|
|
99
110
|
|
|
100
|
-
* [MDX-NET](https://github.com/kuielab/mdx-net) (`mdx-net`):
|
|
111
|
+
* [MDX-NET](https://github.com/kuielab/mdx-net) (`mdx-net`): deep learning source separation architecture by [KUIELAB (Korea University)](https://kuielab.github.io/)
|
package/docs/Licenses.md
CHANGED
|
@@ -1,9 +1,8 @@
|
|
|
1
1
|
# Echogarden components licensing
|
|
2
2
|
|
|
3
|
-
## Engines and libraries
|
|
3
|
+
## Engines and libraries
|
|
4
4
|
|
|
5
5
|
* `onnxruntime-node`: [MIT License](https://github.com/microsoft/onnxruntime/blob/main/LICENSE)
|
|
6
|
-
* `whisper.cpp`: [MIT License](https://github.com/ggerganov/whisper.cpp/blob/master/LICENSE)
|
|
7
6
|
* `espeak`: [GNU GPL v3](https://github.com/espeak-ng/espeak-ng/blob/master/COPYING)
|
|
8
7
|
* `flite`: [BSD License](https://github.com/festvox/flite/blob/master/COPYING)
|
|
9
8
|
* `pico`: [Apache License 2.0](https://github.com/gmn/nanotts/blob/master/LICENSE)
|
|
@@ -31,11 +30,11 @@ All are freely distributable, with varying licenses:
|
|
|
31
30
|
* SVOX Pico resources (`pico-`): [Apache License 2.0](https://github.com/gmn/nanotts/blob/master/LICENSE)
|
|
32
31
|
* Silero VAD (`silero-vad`) and Silero language classifier (`silero-lang-classifier-95`): [MIT License](https://github.com/snakers4/silero-vad/blob/master/LICENSE)
|
|
33
32
|
* Silero speech recognition models (`silero-en-`, `silero-de-`, `silero-es-`, `silero-ua-`): [BY-NC-SA](https://github.com/snakers4/silero-models/blob/master/LICENSE)
|
|
34
|
-
* VITS pre-trained models (`vits-`): licensed under various creative commons licenses: [CC0](https://creativecommons.org/share-your-work/public-domain/cc0/), [CC-BY](https://creativecommons.org/licenses/by/4.0/) and [BY-NC-SA](https://creativecommons.org/licenses/by-nc-sa/4.0/), and few are public domain
|
|
33
|
+
* VITS pre-trained models (`vits-`): licensed under various creative commons licenses: [CC0](https://creativecommons.org/share-your-work/public-domain/cc0/), [CC-BY](https://creativecommons.org/licenses/by/4.0/) and [BY-NC-SA](https://creativecommons.org/licenses/by-nc-sa/4.0/), and few are public domain. You can view the individual license for each model in the model cards on the [Piper samples page](https://rhasspy.github.io/piper-samples/)
|
|
35
34
|
* Whisper pre-trained models (`whisper-`): [MIT License](https://github.com/openai/whisper/blob/main/LICENSE)
|
|
36
35
|
* MDX-NET source separation models (`mdxnet-`): [MIT License](https://github.com/kuielab/mdx-net/blob/main/LICENSE)
|
|
37
36
|
|
|
38
|
-
Tool binary distributions
|
|
37
|
+
Tool binary distributions
|
|
39
38
|
* FFmpeg: [LGPL, GPL v2 and GPL v3 Licenses](https://github.com/FFmpeg/FFmpeg)
|
|
40
39
|
* SoX: [GPL v2 License](https://github.com/chirlu/sox/blob/master/LICENSE.GPL)
|
|
41
40
|
* whisper.cpp: [MIT License](https://github.com/ggerganov/whisper.cpp/blob/master/LICENSE)
|
package/docs/Options.md
CHANGED
|
@@ -4,7 +4,7 @@ Here's a detailed reference for all the options accepted by the Echogarden CLI a
|
|
|
4
4
|
|
|
5
5
|
**Related pages**:
|
|
6
6
|
* [List of all supported engines](Engines.md)
|
|
7
|
-
* [Quick guide
|
|
7
|
+
* [Quick guide to the command line interface](CLI.md)
|
|
8
8
|
* [Node.js API reference](API.md)
|
|
9
9
|
|
|
10
10
|
## Text-to-speech
|
|
@@ -46,7 +46,8 @@ Applies to CLI operations: `speak`, `speak-file`, `speak-url`, `speak-wikipedia`
|
|
|
46
46
|
* `outputAudioFormat.bitrate`: Custom bitrate for encoding, applies only to `mp3`, `opus`, `m4a`, `ogg`. By default, bitrates are selected between 48Kbps and 64Kbps, to provide a good speech quality while minimizing file size. Optional
|
|
47
47
|
|
|
48
48
|
**VITS**:
|
|
49
|
-
* `vits.speakerId`: speaker ID, for VITS models that support multiple speakers.
|
|
49
|
+
* `vits.speakerId`: speaker ID, for VITS models that support multiple speakers. Defaults to `0`
|
|
50
|
+
* `vits.provider`: ONNX execution provider to use. Can be `cpu` or `dml` (https://microsoft.github.io/DirectML/)-based GPU acceleration - Windows only). Using GPU acceleration for VITS may or may not be faster than CPU, depending on your hardware. Defaults to `cpu`
|
|
50
51
|
|
|
51
52
|
**eSpeak**:
|
|
52
53
|
* `espeak.rate`: speech rate, in eSpeak units. Overrides `speed` when set
|
|
@@ -108,7 +109,7 @@ Applies to CLI operations: `speak`, `speak-file`, `speak-url`, `speak-wikipedia`
|
|
|
108
109
|
* `microsoftEdge.trustedClientToken`: trusted client token (required). A special token required to use the service
|
|
109
110
|
* `microsoftEdge.pitchDeltaHz`: pitch delta in Hz. Overrides `pitch` when set
|
|
110
111
|
|
|
111
|
-
|
|
112
|
+
### Voice list request
|
|
112
113
|
|
|
113
114
|
Applies to CLI operation: `list-voices`, API method: `requestVoiceList`
|
|
114
115
|
|
|
@@ -146,9 +147,11 @@ Applies to CLI operation: `transcribe`, API method: `recognize`
|
|
|
146
147
|
* `whisper.topCandidateCount`: the number of top candidate tokens to consider. Defaults to `5`
|
|
147
148
|
* `whisper.punctuationThreshold`: the minimal probability for a punctuation token, included in the top candidates, to be chosen unconditionally. A lower threshold encourages the model to output more punctuation symbols. Defaults to `0.2`
|
|
148
149
|
* `whisper.autoPromptParts`: use previous part's recognized text as prompt for the next part. Disabling this may help to prevent repetition carrying over between parts, in some cases. Defaults to `true`
|
|
149
|
-
* `whisper.maxTokensPerPart`: maximum number of tokens to decode for each
|
|
150
|
-
* `whisper.suppressRepetition`: attempt to suppress decoding repeating token patterns. Defaults to `true`
|
|
151
|
-
* `whisper.decodeTimestampTokens`: enable/disable decoding of timestamp tokens
|
|
150
|
+
* `whisper.maxTokensPerPart`: maximum number of tokens to decode for each audio part. Defaults to `250`
|
|
151
|
+
* `whisper.suppressRepetition`: attempt to suppress decoding of repeating token patterns. Defaults to `true`
|
|
152
|
+
* `whisper.decodeTimestampTokens`: enable/disable decoding of timestamp tokens. Setting to `false` can reduce the occurrence of hallucinations and token repetition loops, possibly due to the overall reduction in the number of tokens decoded. This has no impact on the accuracy of timestamps, since they are derived independently using cross-attention weights. However, there are cases where this can cause the model to end a part prematurely, especially in singing and less speech-like voice segments, or when there are multiple speakers. Defaults to `true`
|
|
153
|
+
* `whisper.encoderProvider`: identifier for the ONNX execution provider to use with the encoder model. Can be `cpu` or `dml` ([DirectML](https://microsoft.github.io/DirectML/)-based GPU acceleration - Windows only). In general, GPU-based encoding should be significantly faster. Defaults to `cpu`, or `dml` if available
|
|
154
|
+
* `whisper.decoderProvider`: identifier for the ONNX execution provider to use with the decoder model. Can be `cpu` or `dml` (Windows only). Using GPU acceleration for the decoder may be faster than CPU, especially for larger models, but that depends on your particular combination of CPU and GPU. Defaults to `cpu`
|
|
152
155
|
* `whisper.seed`: provide a custom random seed for token selection when temperature is greater than 0. Uses a constant seed by default to ensure reproducibility
|
|
153
156
|
|
|
154
157
|
**Whisper.cpp**:
|
|
@@ -157,7 +160,7 @@ Applies to CLI operation: `transcribe`, API method: `recognize`
|
|
|
157
160
|
* `whisperCpp.build`: type of `whisper.cpp` build to use. Can be set `cpu`, `cublas-11.8.0`, `cublas-12.4.0`. By default, builds are auto-selected and downloaded for Windows x64 (`cpu`, `cublas-11.8.0`, `cublas-12.4.0`) and Linux x64 (`cpu`). Using other builds requires providing a custom `executablePath`
|
|
158
161
|
* `whisperCpp.threadCount`: number of threads to use, defaults to `4`
|
|
159
162
|
* `whisperCpp.splitCount`: number of splits of the audio data to process in parallel (called `--processors` in the `whisper.cpp` CLI). A value greater than `1` can increase memory use significantly, reduce timing accuracy, and slow down execution in some cases. Defaults to `1` (highly recommended)
|
|
160
|
-
* `whisperCpp.enableGPU`: enable GPU processing. Defaults to `true`
|
|
163
|
+
* `whisperCpp.enableGPU`: enable GPU processing. Setting to `true` will try to use a CUDA build, if available for your system. Defaults to `true` when a CUDA-enabled build is selected via `whisperCpp.build`, otherwise `false`
|
|
161
164
|
* `whisperCpp.topCandidateCount`: the number of top candidate tokens to consider. Defaults to `5`
|
|
162
165
|
* `whisperCpp.beamCount`: the number of decoding paths to use during beam search. Defaults to `5`
|
|
163
166
|
* `whisperCpp.repetitionThreshold`: minimal repetition / compressibility score to cause a decoded segment to be discarded. Defaults to `2.4`
|
|
@@ -170,6 +173,7 @@ Applies to CLI operation: `transcribe`, API method: `recognize`
|
|
|
170
173
|
|
|
171
174
|
**Silero**:
|
|
172
175
|
* `silero.modelPath`: path to a Silero model. Note that latest `en`, `de`, `fr` and `uk` models are automatically installed when needed based on the selected language. This should only be used to manually specify a different model, otherwise specify `language` instead
|
|
176
|
+
* `silero.provider`: ONNX execution provider to use. Can be `cpu` or `dml` (Windows only). Defaults to `cpu`, or `dml` if available
|
|
173
177
|
|
|
174
178
|
**Google Cloud**:
|
|
175
179
|
* `googleCloud.apiKey`: Google Cloud API key (required)
|
|
@@ -192,7 +196,7 @@ Applies to CLI operation: `transcribe`, API method: `recognize`
|
|
|
192
196
|
* `openAICloud.model`: model to use. Can only be `whisper-1`
|
|
193
197
|
* `openAICloud.organization`: organization identifier. Optional
|
|
194
198
|
* `openAICloud.baseURL`: override the default base URL used by the API. Optional
|
|
195
|
-
* `openAICloud.temperature`: temperature. Choosing `0` uses a dynamic temperature approach. Defaults to `0
|
|
199
|
+
* `openAICloud.temperature`: temperature. Choosing `0` uses a dynamic temperature approach. Defaults to `0`
|
|
196
200
|
* `openAICloud.prompt`: initial prompt for the model. Optional
|
|
197
201
|
* `openAICloud.timeout`: request timeout. Optional
|
|
198
202
|
* `openAICloud.maxRetries`: maximum retries on failure. Defaults to 10
|
|
@@ -220,12 +224,18 @@ Applies to CLI operation: `align`, API method: `align`
|
|
|
220
224
|
* `dtw.granularity`: adjusts the MFCC frame width and hop size based on the profile selected. Can be set to either `auto` (auto-selected based on audio duration and task), `xx-low` (400ms width, 160ms hop), `x-low` (200ms width, 80ms hop), `low` (100ms width, 40ms hop), `medium` (50ms width, 20ms hop), `high` (25ms width, 10ms hop), `x-high` (20ms width, 5ms hop). For multi-pass processing, multiple granularities can be provided, like `dtw.granularity=['low','high']`. Defaults to `auto`.
|
|
221
225
|
* `dtw.windowDuration`: maximum duration (in seconds) of the Sakoe-Chiba window when performing DTW alignment. Higher values consume quadratically larger amounts of memory. The estimated memory requirement is shown in the log before alignment starts. Recommended to be set to at least 10% - 20% of total audio duration. For multi-pass processing, multiple durations can be provided, like `dtw.windowDuration=[240,20]`. Auto-selected by default
|
|
222
226
|
|
|
223
|
-
**DTW-RA
|
|
227
|
+
**DTW-RA**:
|
|
224
228
|
* `recognition`: prefix to provide recognition options when using `dtw-ra` method, for example: setting `recognition.engine = whisper` and `recognition.whisper.model = base.en`
|
|
225
229
|
* `dtw.phoneAlignmentMethod`: algorithm to use when aligning phones: can either be set to `dtw` or `interpolation`. Defaults to `dtw`
|
|
226
230
|
|
|
227
|
-
**Whisper
|
|
228
|
-
|
|
231
|
+
**Whisper**:
|
|
232
|
+
|
|
233
|
+
Applies to the `whisper` engine only. To provide Whisper options for `dtw-ra`, use `recognition.whisper` instead.
|
|
234
|
+
|
|
235
|
+
* `whisper.model`: Whisper model to use. Defaults to `tiny` or `tiny.en`
|
|
236
|
+
* `whisper.endTokenThreshold`: minimal probability to accept an end-of-text token for a recognized part. The probability is measured via the softmax between the end-of-text token's logit and the second highest logit. You can try to adjust this threshold in cases the model is ending a part with too few, or many tokens decoded. Defaults to `0.9`. On the last audio part, it is always effectively set to `Infinity`, to ensure the remaining transcript tokens are decoded in full
|
|
237
|
+
* `whisper.encoderProvider`: encoder ONNX provider. See details in recognition section above
|
|
238
|
+
* `whisper.decoderProvider`: decoder ONNX provider. See details in recognition section above
|
|
229
239
|
|
|
230
240
|
|
|
231
241
|
## Speech-to-text translation
|
|
@@ -254,6 +264,25 @@ Applies to CLI operation: `translate-speech`, API method: `translateSpeech`
|
|
|
254
264
|
|
|
255
265
|
* `openAICloud`: prefix to provide options for OpenAI cloud. Same options as detailed in the recognition section above
|
|
256
266
|
|
|
267
|
+
## Speech-to-translated-transcript alignment
|
|
268
|
+
|
|
269
|
+
Applies to CLI operation: `align-translation`, API method: `alignTranslation`
|
|
270
|
+
|
|
271
|
+
**General**:
|
|
272
|
+
* `engine`: alignment algorithm to use, can only be `whisper`. Defaults to `whisper`
|
|
273
|
+
* `language`: language code for the source audio ([ISO 639-1](https://en.wikipedia.org/wiki/List_of_ISO_639-1_codes)), like `en`, `fr`, `zh`, etc. Auto-detected from audio if not set
|
|
274
|
+
* `crop`: crop to active parts using voice activity detection before starting. Defaults to `true`
|
|
275
|
+
* `isolate`: apply source separation to isolate voice before starting alignment. Defaults to `false`
|
|
276
|
+
* `subtitles`: prefix to provide options for subtitles. Options detailed in section for subtitles
|
|
277
|
+
* `vad`: prefix to provide options for voice activity detection when `crop` is set to `true`. Options detailed in section for voice activity detection
|
|
278
|
+
* `sourceSeparation`: prefix to provide options for source separation when `isolate` is set to `true`. Options detailed in section for source separation
|
|
279
|
+
|
|
280
|
+
**Whisper**:
|
|
281
|
+
* `whisper.model`: Whisper model to use. Only multilingual models can be used. Defaults to `tiny`
|
|
282
|
+
* `whisper.endTokenThreshold`: see details in the alignment section above
|
|
283
|
+
* `whisper.encoderProvider`: encoder ONNX execution provider. See details in recognition section above
|
|
284
|
+
* `whisper.decoderProvider`: decoder ONNX execution provider. See details in recognition section above
|
|
285
|
+
|
|
257
286
|
## Language detection
|
|
258
287
|
|
|
259
288
|
### Speech language detection
|
|
@@ -270,6 +299,11 @@ Applies to CLI operation: `detect-speech-langauge`, API method: `detectSpeechLan
|
|
|
270
299
|
**Whisper**:
|
|
271
300
|
* `whisper.model`: Whisper model to use. See model list in the recognition section
|
|
272
301
|
* `whisper.temperature`: impacts the distribution of candidate languages when applying the softmax function to compute language probabilities over the model output. Higher temperature causes the distribution to be more uniform, while lower temperature causes it to be more strongly weighted towards the best scoring candidates. Defaults to `1.0`
|
|
302
|
+
* `whisper.encoderProvider`: encoder ONNX execution provider. See details in recognition section above
|
|
303
|
+
* `whisper.decoderProvider`: decoder ONNX execution provider. See details in recognition section above
|
|
304
|
+
|
|
305
|
+
**Silero**:
|
|
306
|
+
* `silero.provider`: ONNX execution provider to use. Can be `cpu` or `dml` (Windows only). Using GPU may be faster, but the initialization overhead is larger. **Note**: `dml` provider seems to be unstable at the moment for this model. Defaults to `cpu`
|
|
273
307
|
|
|
274
308
|
### Text language detection
|
|
275
309
|
|
|
@@ -294,6 +328,7 @@ Applies to CLI operation: `detect-voice-activity`, API method: `detectVoiceActiv
|
|
|
294
328
|
|
|
295
329
|
**Silero**:
|
|
296
330
|
* `silero.frameDuration`: Silero frame duration (ms). Can be `30`, `60` or `90`. Defaults to `90`
|
|
331
|
+
* `silero.provider`: ONNX provider to use. Can be `cpu` or `dml` (Windows only). Using GPU is likely to be slower than CPU due to inference being independently executed on each audio frame. Defaults to `cpu` (recommended)
|
|
297
332
|
|
|
298
333
|
## Speech denoising
|
|
299
334
|
|
|
@@ -319,6 +354,7 @@ Applies to CLI operation: `isolate`, API method: `isolate`
|
|
|
319
354
|
**MDX-NET**:
|
|
320
355
|
|
|
321
356
|
* `mdxNet.model`: model to use. Currently available models are `UVR_MDXNET_1_9703`, `UVR_MDXNET_2_9682`, `UVR_MDXNET_3_9662`, `UVR_MDXNET_KARA`. Defaults to `UVR_MDXNET_1_9703`
|
|
357
|
+
* `mdxNet.provider`: ONNX execution provider to use. Can be `cpu` or `dml` ([DirectML](https://microsoft.github.io/DirectML/), Windows only). **Note**: `dml` provider seems to be unstable with MDX-NET models at the moment. Defaults to `cpu`
|
|
322
358
|
|
|
323
359
|
# Common options
|
|
324
360
|
|
package/docs/Releases.md
CHANGED
package/docs/Server.md
CHANGED
|
@@ -1,6 +1,8 @@
|
|
|
1
|
-
#
|
|
1
|
+
# WebSocket server API reference
|
|
2
2
|
|
|
3
|
-
This is a guide
|
|
3
|
+
This is a guide to the WebSocket server protocol.
|
|
4
|
+
|
|
5
|
+
**Note**: The protocol is still in early development and may change in future releases. Many features are currently missing, and the server hasn't been thoroughly tested.
|
|
4
6
|
|
|
5
7
|
## Starting the server
|
|
6
8
|
|
|
@@ -36,10 +38,10 @@ ws.on("open", async () => {
|
|
|
36
38
|
})
|
|
37
39
|
```
|
|
38
40
|
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
41
|
+
**TODO**:
|
|
42
|
+
* Separate the client to an independent, lightweight, Node.js package, with browser compatibility
|
|
43
|
+
* Add support for cancellation signals
|
|
44
|
+
* Document how to use with a background worker
|
|
43
45
|
|
|
44
46
|
## Protocol details
|
|
45
47
|
|
package/docs/Tasklist.md
CHANGED
|
@@ -4,13 +4,13 @@
|
|
|
4
4
|
|
|
5
5
|
### Alignment
|
|
6
6
|
|
|
7
|
-
* In DTW-RA, recognition transcript including something like "Question 2.What does Juan", where "2.What" has a point in the middle, is breaking playback of the timeline
|
|
8
|
-
* DTW-RA will not work correctly with Polish language texts, due to issues with the eSpeak engine pronouncing `|` characters, which are intended to be used as separators and ignored by all other eSpeak languages
|
|
7
|
+
* In DTW-RA, recognition transcript including something like "Question 2.What does Juan", where "2.What" has a point in the middle, is breaking playback of the timeline
|
|
8
|
+
* DTW-RA will not work correctly with Polish language texts, due to issues with the eSpeak engine pronouncing `|` characters, which are intended to be used as separators and ignored by all other eSpeak languages
|
|
9
9
|
|
|
10
10
|
### Synthesis
|
|
11
11
|
|
|
12
12
|
### Phoneme processing
|
|
13
|
-
* IPA -> Kirshenbaum translation is still not completely similar to what is output by eSpeak. Also, in rare situations, it outputs characters that are not accepted by eSpeak and eSpeak errors. Investigate when that happens and how to improve on this
|
|
13
|
+
* IPA -> Kirshenbaum translation is still not completely similar to what is output by eSpeak. Also, in rare situations, it outputs characters that are not accepted by eSpeak and eSpeak errors. Investigate when that happens and how to improve on this
|
|
14
14
|
|
|
15
15
|
### Browser extension
|
|
16
16
|
* Investigate why WebSpeech events sometimes completely stop working in the middle of an utterance for no apparent reason. Sometimes this is permanent, until the extension is restarted. Is this a browser issue?
|
|
@@ -34,17 +34,9 @@
|
|
|
34
34
|
|
|
35
35
|
## Features and enhancements
|
|
36
36
|
|
|
37
|
-
### Server
|
|
38
|
-
* Option to allow or disallow local file paths as arguments to API methods (as a security safeguard)
|
|
39
|
-
|
|
40
|
-
### Worker
|
|
41
|
-
* Add cancellation checks in more operations
|
|
42
|
-
* Support more operations
|
|
43
|
-
|
|
44
37
|
### CLI
|
|
45
38
|
* Show names of files written do disk. This is useful for cases where a file is auto-renamed to prevent overwriting existing data
|
|
46
39
|
* Restrict input media file extensions to ensure that invalid files are not passed to FFmpeg
|
|
47
|
-
* Mode to print IPA words when speaking
|
|
48
40
|
* Consider what to do with non-supported templates like `[hello]`
|
|
49
41
|
* Show a message when a new version is available
|
|
50
42
|
* Figure out which terminal outputs should go to stdout, or if that's a good idea at all
|
|
@@ -55,7 +47,8 @@
|
|
|
55
47
|
* Suggest possible correction on the error of not using `=`, e.g. `speed 0.9` instead of `speed=0.9`
|
|
56
48
|
* Multiple configuration files in `--config=..` taking precedence by order
|
|
57
49
|
* Generate JSON configuration file schema
|
|
58
|
-
* Use a file type detector like `file-type` that uses magic numbers to detect the type of a binary file regardless of its extension. This would help to give better error messages when the given file type is wrong
|
|
50
|
+
* Use a file type detector like `file-type` that uses magic numbers to detect the type of a binary file regardless of its extension. This would help to give better error messages when the given file type is wrong
|
|
51
|
+
* Mode to print IPA words when speaking
|
|
59
52
|
|
|
60
53
|
### CLI / playback
|
|
61
54
|
* Option to set audio output device for playback
|
|
@@ -64,7 +57,7 @@
|
|
|
64
57
|
* Add phone playback support
|
|
65
58
|
|
|
66
59
|
### CLI / `speak`
|
|
67
|
-
* Add support for sentence templates, like `echogarden speak-file text.txt /parts/[sentence].wav
|
|
60
|
+
* Add support for sentence templates, like `echogarden speak-file text.txt /parts/[sentence].wav`
|
|
68
61
|
|
|
69
62
|
### CLI / `speak-wikipedia`
|
|
70
63
|
* Correctly detect language when a Wikipedia URL is passed instead of an article name
|
|
@@ -84,7 +77,7 @@
|
|
|
84
77
|
* `play-with-timeline`: Preview timeline in terminal
|
|
85
78
|
* `subtitles-to-text`, `subtitles-to-timeline`, `srt-to-vtt`, `vtt-to-srt`
|
|
86
79
|
* `text-to-ipa`, `arpabet-to-ipa`, `ipa-to-arpabet`
|
|
87
|
-
* `phonemize
|
|
80
|
+
* `phonemize`
|
|
88
81
|
* `normalize-text`
|
|
89
82
|
* `transcribe-youtube`: Transcribe the audio in a YouTube video (requires fetching the audio somehow - which can't be done using the normal YouTube API)
|
|
90
83
|
* `speak-youtube-subtitles`: To speak the subtitles of a YouTube video
|
|
@@ -98,7 +91,7 @@
|
|
|
98
91
|
* Accept voice list caching options in `SynthesisOptions`
|
|
99
92
|
|
|
100
93
|
### Package manager
|
|
101
|
-
* Better error message when package is not found remotely. Currently, it just gives a `404 not found` without any other information
|
|
94
|
+
* Better error message when package is not found remotely. Currently, it just gives a `404 not found` without any other information
|
|
102
95
|
* Retry on network failure
|
|
103
96
|
|
|
104
97
|
### Speech language detection
|
|
@@ -111,7 +104,7 @@
|
|
|
111
104
|
* See if it's possible to reliably use eSpeak as a segmentation engine
|
|
112
105
|
|
|
113
106
|
### Subtitles
|
|
114
|
-
* Split long words if needed. This is especially important for Chinese
|
|
107
|
+
* Split long words if needed. This is especially important for Chinese and Japanese
|
|
115
108
|
* If a subtitle is too short and at the end of the audio, try to extend it back if possible (for example, if the previous subtitle is already extended, take back from it)
|
|
116
109
|
* Decide how many punctuation characters to allow before breaking to a new line (currently it's infinite)
|
|
117
110
|
* Add more clause separators, for even more special cases
|
|
@@ -119,26 +112,26 @@
|
|
|
119
112
|
* Parse VTT's language
|
|
120
113
|
|
|
121
114
|
### Synthesis
|
|
122
|
-
* Option to disable alignment (only for some engines). Alternative: use a low granularity setting that is very fast to compute
|
|
115
|
+
* Option to disable alignment (only for some engines). Alternative: use a low granularity DTW setting that is very fast to compute
|
|
123
116
|
* Find places to add commas (",") to improve speech fluency. VITS voices don't normally add speech breaks if there is no punctuation
|
|
124
|
-
* An isolated dash " - " can be converted to a " , " to ensure there's a break in the speech
|
|
117
|
+
* An isolated dash " - " can be converted to a " , " to ensure there's a break in the speech
|
|
125
118
|
* Ensure abbreviations like "Ph.d" or names like are segmented and read correctly (does `cldr` treat it as a word? Maybe eSpeak doesn't recognize it as a word). "C#" and ".NET" as well
|
|
126
119
|
* Find way to manually reset voice list cache
|
|
127
120
|
* When synthesized text isn't pre-split to sentences, apply sentence splits by using the existing method to convert the output of word timelines to sentence/segment timelines
|
|
128
121
|
* Some `sapi` voices and `msspeech` languages output phones that are converted to Microsoft alphabet, not IPA symbols. Try to see if these can be translated to IPA
|
|
129
122
|
* Decide whether asterisk `*` should be spoken when using `speak-url` or `speak-wikipedia`
|
|
130
|
-
* Add partial SSML support for all engines. In particular, allow changing language or voice using the `<voice>` and `<lang>` tags, `<say-as>` and `<phoneme>` where possible
|
|
131
|
-
* Try to remove reliance on `()` after `.` character hack in `EspeakTTS.synthesizeFragments
|
|
132
|
-
* eSpeak IPA output puts stress marks on vowels, not syllables - which is the standard for IPA. Consider how to make a conversion to and from these two approaches (possibly detect it automatically)
|
|
123
|
+
* Add partial SSML support for all engines. In particular, allow changing language or voice using the `<voice>` and `<lang>` tags, `<say-as>` and `<phoneme>` where possible
|
|
124
|
+
* Try to remove reliance on `()` after `.` character hack in `EspeakTTS.synthesizeFragments`
|
|
125
|
+
* eSpeak IPA output puts stress marks on vowels, not syllables - which is the standard for IPA. Consider how to make a conversion to and from these two approaches (possibly detect it automatically)
|
|
133
126
|
* Decide if `msspeech` engine should be selected if available. This would require attempting to load a matching voice, and falling back if it is not installed
|
|
134
127
|
* Speaker-specific voice option
|
|
135
128
|
* Use VAD on the synthesized audio file to get more accurate sentence or word segmentation
|
|
136
|
-
* When `splitToSentences` is set to `false`, the timeline doesn't include proper sentences. Find a way to pass larger sections to the TTS, but still have proper sentences in the timeline
|
|
129
|
+
* When `splitToSentences` is set to `false`, the timeline doesn't include proper sentences. Find a way to pass larger sections to the TTS, but still have proper sentences in the timeline
|
|
137
130
|
|
|
138
131
|
### Synthesis / preprocessing
|
|
139
132
|
* Extend the heteronyms JSON document with additional words like "conducts", "survey", "protest", "transport", "abuse", "combat", "combats", "affect", "contest", "detail", "marked", "contrast", "construct", "constructs", "console", "recall", "permit", "permits", "prospect", "prospects", "proceed", "proceeds", "invite", "reject", "deserts", "transcript", "transcripts", "compact", "impact", "impacts"
|
|
140
133
|
* Full date normalization (e.g. `21 August 2023`, `21 Aug 2023`, `August 21, 2023`)
|
|
141
|
-
* Add support for capitalized-only rules, and possibly also all uppercase / all lowercase rules
|
|
134
|
+
* Add support for capitalized-only rules, and possibly also all uppercase / all lowercase rules
|
|
142
135
|
* Add support for multiple words in `precededBy` and `succeededBy`
|
|
143
136
|
* Support substituting to graphemes in lexicons, not only phonemes
|
|
144
137
|
* Cache lexicons to avoid parsing the JSON each time it is loaded (this may not be needed for if the file is relatively small)
|
|
@@ -160,30 +153,33 @@
|
|
|
160
153
|
### Recognition
|
|
161
154
|
* Recognized word entries that span VAD segment boundaries can be split
|
|
162
155
|
* Show alternatives when playing in the CLI. Clear current line and rewrite already printed text for alternatives during the speech recognition process
|
|
163
|
-
* Option to split recognized audio to segments or sentences, as is done with synthesized audio
|
|
164
|
-
* Try to exclude the timing for trailing punctuation tokens in words that contain them. This can help narrow down the end timestamp to cover the word more tightly
|
|
165
156
|
|
|
166
157
|
### Recognition / Whisper
|
|
167
|
-
*
|
|
158
|
+
* Whisper's Chinese and Japanese output can be split to words in a more accurate way. Consider using a dedicated segmentation library to perform the segmentation in character sequences that have no spaces within them
|
|
168
159
|
* Automatically disable using previous section recognized transcript as prompt for the next section when lots of repetition occurred in previous section
|
|
169
160
|
* Cache last model (if enough memory available)
|
|
170
|
-
* Bring back option to use eSpeak DTW based alignment on segments, as an alternative approach
|
|
171
161
|
* The segment output can be used to split to segments, otherwise it is possible to try to guess using pause lengths or voice activity detection
|
|
172
|
-
*
|
|
173
|
-
* Way to specify model size only, such that the English-only/multilingual variant would be automatically selected for sizes other than `tiny`?
|
|
174
|
-
* Timestamps extracted from cross-attention are still not as accurate as what the official Python implementation gets. Try to see if you can make them better
|
|
175
|
-
* Whisper's Chinese output can be split to words in a more accurate way. Consider using a dedicated segmentation library to perform the segmentation in character sequences that have no spaces within them
|
|
162
|
+
* Bring back option to use eSpeak DTW based alignment on segments, as an alternative approach
|
|
176
163
|
|
|
177
164
|
### Alignment
|
|
178
165
|
|
|
179
166
|
* Aligned words entries that span VAD boundaries may be split
|
|
180
167
|
|
|
181
168
|
### Alignment / DTW-RA
|
|
182
|
-
* Remove emojis and other special characters, that are not likely to be pronounced in the speech, from the transcript timeline before it is synthesized. For example Whisper may produce 'note' emojis when it detects singing or music. Pronouncing them reduces the accuracy of the alignment
|
|
183
|
-
* Optional mode to pass Whisper a special vocabulary of tokens that can appear in the transcript. All other tokens would be suppressed
|
|
184
169
|
|
|
185
170
|
### Alignment / Whisper
|
|
186
|
-
|
|
171
|
+
|
|
172
|
+
### Source separation / MDX-NET
|
|
173
|
+
* Since MDX-NET requires FFT with large window sizes, the FFT computation overhead currently acts as a bottleneck, especially when GPU is used for inference. Currently it uses a WASM port of KissFFT, running on a single thread, which is still relatively fast. To get higher performance, try to (optionally) use native, SIMD optimized FFT like FFTW3 via a NAPI addon, with multi-threading enabled
|
|
174
|
+
* Option to customize overlap
|
|
175
|
+
* Add more models
|
|
176
|
+
|
|
177
|
+
### Server
|
|
178
|
+
* Option to allow or disallow local file paths as arguments to API methods (as a security safeguard)
|
|
179
|
+
|
|
180
|
+
### Worker
|
|
181
|
+
* Add cancellation checks in more operations
|
|
182
|
+
* Support more operations
|
|
187
183
|
|
|
188
184
|
### Browser extension
|
|
189
185
|
* Options UI
|
|
@@ -212,7 +208,7 @@
|
|
|
212
208
|
* See if the installation of `winax` can be automated and only initiate if it is in a Windows environment
|
|
213
209
|
* Ensure that all modules have no internal state other than caching
|
|
214
210
|
* Start thinking about some modules being available in the browser. Which node core APIs the use? Which of them can be polyfilled, and which cannot?
|
|
215
|
-
* Change all the Emscripten WASM modules to use the `EXPORT_ES6=1` flag and rebuild them. Support for node.js modules was
|
|
211
|
+
* Change all the Emscripten WASM modules to use the `EXPORT_ES6=1` flag and rebuild them. Support for node.js modules was added in September 2022 (https://github.com/emscripten-core/emscripten/pull/17915)
|
|
216
212
|
* Remove built-in voices from `flite` to reduce size?
|
|
217
213
|
* Slim down `kuromoji` package to reduce base installation size
|
|
218
214
|
|
|
@@ -224,7 +220,6 @@
|
|
|
224
220
|
* Test everything's fine on macOS
|
|
225
221
|
* Test that cloud services all still work correctly, especially with SSML inputs
|
|
226
222
|
|
|
227
|
-
|
|
228
223
|
## Future features and enhancements
|
|
229
224
|
|
|
230
225
|
### CLI
|
|
@@ -259,18 +254,10 @@
|
|
|
259
254
|
* Investigate exporting Whisper models to 16-bit quantized ONNX or a mix of 16-bit and 32-bit
|
|
260
255
|
|
|
261
256
|
### Alignment
|
|
262
|
-
* Implement alignment with speech-to-text translation assistance, which would enable multilingual subtitle replacement for translated subtitles
|
|
263
257
|
* Method to align audio file to audio file
|
|
264
|
-
*
|
|
265
|
-
* Predict timing for individual letters (graphemes) based on phoneme timestamps
|
|
258
|
+
* Allow `dtw` mode work with more speech synthesizers to produce its reference
|
|
259
|
+
* Predict timing for individual letters (graphemes) based on phoneme timestamps (especially useful for Chinese and Japanese)
|
|
266
260
|
|
|
267
|
-
### Voice activity detection
|
|
268
|
-
|
|
269
|
-
* Whisper-based VAD. Use Whisper's 'no speech' token to determine if the audio contains speech
|
|
270
|
-
|
|
271
|
-
### Source separation
|
|
272
|
-
* Option to customize overlap
|
|
273
|
-
* Add more MDX-NET models
|
|
274
261
|
|
|
275
262
|
## Possible new engines or platforms
|
|
276
263
|
|
|
@@ -280,13 +267,13 @@
|
|
|
280
267
|
* Coqui STT server connection
|
|
281
268
|
* [MarbleNet VAD](https://github.com/NVIDIA/NeMo/blob/main/tutorials/asr/Online_Offline_Microphone_VAD_Demo.ipynb), included of the NVIDIA NeMo framework, can be exported to ONNX
|
|
282
269
|
* Silero text enhancement engine can be ported to ONNX
|
|
283
|
-
* See what can be done
|
|
284
|
-
* Figure out how to support `julius` speech recognition via WASM
|
|
270
|
+
* See what can be done for supporting WinRT speech: in particular `windows.media.speechsynthesis` and `windows.media.speechrecognition` support, possibly using NodeRT or some other method
|
|
271
|
+
* Figure out how to support `julius` speech recognition via WASM
|
|
285
272
|
* Any way to support RHVoice?
|
|
286
273
|
|
|
287
274
|
## Maybe?
|
|
288
275
|
|
|
289
|
-
* Using a machine translation model to provide speech translation to languages other than English?
|
|
276
|
+
* Using a machine translation model to provide speech translation to languages other than English? How would the timing be determined?
|
|
290
277
|
* Is it possible to get sentence boundaries without punctuation using NLP techniques like part of speech tagging?
|
|
291
278
|
|
|
292
279
|
## May or may not be good ideas
|
|
@@ -298,10 +285,10 @@
|
|
|
298
285
|
|
|
299
286
|
* Support alignment of EPUB 3 eBooks with corresponding audiobook
|
|
300
287
|
* Voice cloning
|
|
301
|
-
* Speech
|
|
302
|
-
* Speech-to-speech translation
|
|
288
|
+
* Speech-to-speech voice conversion
|
|
289
|
+
* Speech-to-speech translation
|
|
303
290
|
* HTML generator, that includes text and audio, with playback and word highlighting
|
|
304
291
|
* Video generator
|
|
305
292
|
* Desktop app that uses the tool to transcribe the PC audio output
|
|
306
|
-
* Special method to use time stretching to project between different utterances of the same text
|
|
293
|
+
* Special method to use time stretching to project between different aligned utterances of the same text
|
|
307
294
|
* Is it possible to combine the Silero speech recognizer and a language model and try to perform Viterbi decoding to find alignments?
|
package/docs/Technical.md
CHANGED
|
@@ -26,7 +26,7 @@ The base installed (uncompressed) size, including dependencies, is around 270MB.
|
|
|
26
26
|
|
|
27
27
|
Currently, the largest contributors to the size are:
|
|
28
28
|
|
|
29
|
-
* `onnxruntime-node` (NAPI):
|
|
29
|
+
* `onnxruntime-node` (NAPI): 133MB
|
|
30
30
|
* `kuromoji` (JavaScript) 40MB
|
|
31
31
|
* `flite-wasi` (WASI): 20MB
|
|
32
32
|
* `espeak-ng-emscripten` (WASM): 18MB
|