echogarden 1.0.4 → 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (122) hide show
  1. package/README.md +26 -23
  2. package/data/schemas/options.json +177 -36
  3. package/dist/alignment/SpeechAlignment.d.ts +1 -1
  4. package/dist/alignment/SpeechAlignment.js +1 -1
  5. package/dist/alignment/SpeechAlignment.js.map +1 -1
  6. package/dist/api/API.d.ts +1 -0
  7. package/dist/api/API.js +1 -0
  8. package/dist/api/API.js.map +1 -1
  9. package/dist/api/APIOptions.d.ts +1 -0
  10. package/dist/api/Alignment.d.ts +3 -3
  11. package/dist/api/Alignment.js +5 -10
  12. package/dist/api/Alignment.js.map +1 -1
  13. package/dist/api/LanguageDetection.d.ts +5 -7
  14. package/dist/api/LanguageDetection.js +3 -2
  15. package/dist/api/LanguageDetection.js.map +1 -1
  16. package/dist/api/Recognition.d.ts +4 -5
  17. package/dist/api/Recognition.js +5 -8
  18. package/dist/api/Recognition.js.map +1 -1
  19. package/dist/api/SourceSeparation.d.ts +2 -0
  20. package/dist/api/SourceSeparation.js +4 -2
  21. package/dist/api/SourceSeparation.js.map +1 -1
  22. package/dist/api/Synthesis.d.ts +3 -1
  23. package/dist/api/Synthesis.js +9 -10
  24. package/dist/api/Synthesis.js.map +1 -1
  25. package/dist/api/Translation.d.ts +1 -1
  26. package/dist/api/Translation.js +4 -8
  27. package/dist/api/Translation.js.map +1 -1
  28. package/dist/api/TranslationAlignment.d.ts +31 -0
  29. package/dist/api/TranslationAlignment.js +121 -0
  30. package/dist/api/TranslationAlignment.js.map +1 -0
  31. package/dist/api/VoiceActivityDetection.d.ts +5 -1
  32. package/dist/api/VoiceActivityDetection.js +38 -2
  33. package/dist/api/VoiceActivityDetection.js.map +1 -1
  34. package/dist/audio/AudioPlayer.js +6 -1
  35. package/dist/audio/AudioPlayer.js.map +1 -1
  36. package/dist/cli/CLI.js +85 -0
  37. package/dist/cli/CLI.js.map +1 -1
  38. package/dist/dsp/FFT.js.map +1 -1
  39. package/dist/math/MedianFilter.d.ts +5 -0
  40. package/dist/math/MedianFilter.js +102 -0
  41. package/dist/math/MedianFilter.js.map +1 -0
  42. package/dist/math/VectorMath.d.ts +0 -2
  43. package/dist/math/VectorMath.js +1 -25
  44. package/dist/math/VectorMath.js.map +1 -1
  45. package/dist/recognition/OpenAICloudSTT.d.ts +1 -1
  46. package/dist/recognition/OpenAICloudSTT.js.map +1 -1
  47. package/dist/recognition/SileroSTT.d.ts +22 -1
  48. package/dist/recognition/SileroSTT.js +122 -95
  49. package/dist/recognition/SileroSTT.js.map +1 -1
  50. package/dist/recognition/WhisperCppSTT.js +1 -1
  51. package/dist/recognition/WhisperCppSTT.js.map +1 -1
  52. package/dist/recognition/WhisperSTT.d.ts +52 -19
  53. package/dist/recognition/WhisperSTT.js +645 -494
  54. package/dist/recognition/WhisperSTT.js.map +1 -1
  55. package/dist/server/Server.js.map +1 -1
  56. package/dist/source-separation/MDXNetSourceSeparation.d.ts +5 -3
  57. package/dist/source-separation/MDXNetSourceSeparation.js +26 -19
  58. package/dist/source-separation/MDXNetSourceSeparation.js.map +1 -1
  59. package/dist/speech-language-detection/SileroLanguageDetection.d.ts +15 -9
  60. package/dist/speech-language-detection/SileroLanguageDetection.js +23 -16
  61. package/dist/speech-language-detection/SileroLanguageDetection.js.map +1 -1
  62. package/dist/synthesis/EspeakTTS.js +4 -0
  63. package/dist/synthesis/EspeakTTS.js.map +1 -1
  64. package/dist/synthesis/GoogleCloudTTS.js.map +1 -1
  65. package/dist/synthesis/VitsTTS.d.ts +8 -6
  66. package/dist/synthesis/VitsTTS.js +36 -31
  67. package/dist/synthesis/VitsTTS.js.map +1 -1
  68. package/dist/tests/Test.js.map +1 -1
  69. package/dist/utilities/OnnxUtilities.d.ts +14 -0
  70. package/dist/utilities/OnnxUtilities.js +43 -0
  71. package/dist/utilities/OnnxUtilities.js.map +1 -0
  72. package/dist/utilities/Utilities.d.ts +4 -8
  73. package/dist/utilities/Utilities.js +35 -58
  74. package/dist/utilities/Utilities.js.map +1 -1
  75. package/dist/voice-activity-detection/SileroVAD.d.ts +5 -3
  76. package/dist/voice-activity-detection/SileroVAD.js +9 -11
  77. package/dist/voice-activity-detection/SileroVAD.js.map +1 -1
  78. package/docs/API.md +54 -34
  79. package/docs/CLI.md +25 -13
  80. package/docs/Contributing.md +4 -2
  81. package/docs/Engines.md +43 -32
  82. package/docs/Licenses.md +3 -4
  83. package/docs/Options.md +47 -11
  84. package/docs/Releases.md +4 -0
  85. package/docs/Server.md +8 -6
  86. package/docs/Tasklist.md +39 -52
  87. package/docs/Technical.md +1 -1
  88. package/package.json +8 -12
  89. package/src/alignment/SpeechAlignment.ts +1 -1
  90. package/src/api/API.ts +1 -0
  91. package/src/api/APIOptions.ts +1 -0
  92. package/src/api/Alignment.ts +10 -14
  93. package/src/api/LanguageDetection.ts +14 -10
  94. package/src/api/Recognition.ts +17 -10
  95. package/src/api/SourceSeparation.ts +7 -2
  96. package/src/api/Synthesis.ts +26 -11
  97. package/src/api/Translation.ts +14 -8
  98. package/src/api/TranslationAlignment.ts +213 -0
  99. package/src/api/VoiceActivityDetection.ts +66 -3
  100. package/src/audio/AudioPlayer.ts +6 -2
  101. package/src/cli/CLI.ts +121 -2
  102. package/src/dsp/FFT.ts +3 -0
  103. package/src/math/MedianFilter.ts +124 -0
  104. package/src/math/VectorMath.ts +1 -36
  105. package/src/recognition/OpenAICloudSTT.ts +27 -27
  106. package/src/recognition/SileroSTT.ts +149 -102
  107. package/src/recognition/WhisperCppSTT.ts +1 -1
  108. package/src/recognition/WhisperSTT.ts +961 -684
  109. package/src/server/Server.ts +1 -1
  110. package/src/source-separation/MDXNetSourceSeparation.ts +35 -19
  111. package/src/speech-language-detection/SileroLanguageDetection.ts +53 -33
  112. package/src/synthesis/EspeakTTS.ts +8 -0
  113. package/src/synthesis/GoogleCloudTTS.ts +12 -1
  114. package/src/synthesis/VitsTTS.ts +57 -46
  115. package/src/tests/Test.ts +1 -1
  116. package/src/utilities/OnnxUtilities.ts +68 -0
  117. package/src/utilities/Utilities.ts +38 -66
  118. package/src/voice-activity-detection/SileroVAD.ts +15 -15
  119. package/dist/utilities/NdArrayUtilities.d.ts +0 -3
  120. package/dist/utilities/NdArrayUtilities.js +0 -23
  121. package/dist/utilities/NdArrayUtilities.js.map +0 -1
  122. package/src/utilities/NdArrayUtilities.ts +0 -31
package/docs/Engines.md CHANGED
@@ -5,17 +5,17 @@
5
5
 
6
6
  **Offline**:
7
7
 
8
- * [VITS](https://github.com/jaywalnut310/vits) (`vits`): a high-quality end-to-end neural speech synthesis architecture. Available models were trained by Michael Hansen as part of his [Piper speech synthesis system](https://github.com/rhasspy/piper). Currently, there are 117 voices, in a range of languages, including English (US, UK), Spanish (ES, MX), Portuguese (PT, BR), Italian, French, German, Dutch (NL, BE), Swedish, Norwegian, Danish, Finnish, Polish, Greek, Romanian, Serbian, Czech, Hungarian, Slovak, Slovenian, Turkish, Arabic, Farsi, Russian, Ukrainian, Catalan, Luxembourgish, Icelandic, Swahili, Kazakh, Georgian, Nepali, Vietnamese and Chinese. You can listen to audio samples of all voices and languages in [Piper's samples page](https://rhasspy.github.io/piper-samples/).
9
- * [SVOX Pico](https://github.com/naggety/picotts) (`pico`): a legacy diphone-based synthesis engine. Supports English (US, UK), Spanish, Italian, French, and German.
10
- * [Flite](https://github.com/festvox/flite) (`flite`): a legacy diphone-based synthesis engine. Supports English (US, Scottish), and several Indic languages: Hindi, Bengali, Marathi, Telugu, Tamil, Gujarati, Kannada and Punjabi.
11
- * [eSpeak-NG](https://github.com/espeak-ng/espeak-ng/) (`espeak`): a lightweight "robot" sounding formant-based synthesizer. Supports 100+ languages. Extensively used internally for speech alignment, phonemization, and other internal tasks.
12
- * [SAM (Software Automatic Mouth)](https://github.com/discordier/sam) (`sam`): a classic "robot" speech synthesizer from 1982. English only.
8
+ * [VITS](https://github.com/jaywalnut310/vits) (`vits`): end-to-end neural speech synthesis architecture. Available models were trained by Michael Hansen as part of his [Piper speech synthesis system](https://github.com/rhasspy/piper). Currently, there are 117 voices, in a range of languages, including English (US, UK), Spanish (ES, MX), Portuguese (PT, BR), Italian, French, German, Dutch (NL, BE), Swedish, Norwegian, Danish, Finnish, Polish, Greek, Romanian, Serbian, Czech, Hungarian, Slovak, Slovenian, Turkish, Arabic, Farsi, Russian, Ukrainian, Catalan, Luxembourgish, Icelandic, Swahili, Kazakh, Georgian, Nepali, Vietnamese and Chinese. You can listen to audio samples of all voices and languages in [Piper's samples page](https://rhasspy.github.io/piper-samples/)
9
+ * [SVOX Pico](https://github.com/naggety/picotts) (`pico`): a legacy diphone-based synthesis engine. Supports English (US, UK), Spanish, Italian, French, and German
10
+ * [Flite](https://github.com/festvox/flite) (`flite`): a legacy diphone-based synthesis engine. Supports English (US, Scottish), and several Indic languages: Hindi, Bengali, Marathi, Telugu, Tamil, Gujarati, Kannada and Punjabi
11
+ * [eSpeak-NG](https://github.com/espeak-ng/espeak-ng/) (`espeak`): a lightweight "robot" sounding formant-based synthesizer. Supports 100+ languages. Used internally for speech alignment, phonemization, and other internal tasks
12
+ * [SAM (Software Automatic Mouth)](https://github.com/discordier/sam) (`sam`): a classic "robot" speech synthesizer from 1982. English only
13
13
 
14
14
  **Offline, Windows only**:
15
15
 
16
- * [SAPI](https://en.wikipedia.org/wiki/Microsoft_Speech_API) (`sapi`): Microsoft Speech API. Supports the system's language voices, as well as legacy voices produced by third-party vendors, like Ivona, NeoSpeech, Acapela, Cepstral, CereProc, Nuance, AT&T, Loquendo, ScanSoft and others (note that only 64-bit SAPI voices are supported, which makes it incompatible with a significant portion of older voices).
16
+ * [SAPI](https://en.wikipedia.org/wiki/Microsoft_Speech_API) (`sapi`): Microsoft Speech API. Supports the system's language voices, as well as legacy voices produced by third-party vendors, like Ivona, NeoSpeech, Acapela, Cepstral, CereProc, Nuance, AT&T, Loquendo, ScanSoft and others (note that only 64-bit SAPI voices are supported, which makes it incompatible with a significant portion of older voices)
17
17
 
18
- * [Microsoft Speech Platform](https://www.microsoft.com/en-us/download/details.aspx?id=27225) (`msspeech`): Microsoft Server Speech API. Requires [installing a runtime (2.6MB)](https://www.microsoft.com/en-us/download/details.aspx?id=27225). Supports 28 dialects, which can be individually downloaded via [freely available installers](https://www.microsoft.com/en-us/download/details.aspx?id=27224), or, for convenience, bundled as [a single 358MB zip file](https://drive.google.com/u/0/uc?id=1uQdFNxLzUxpaEwVVKhMawys8cIh3F21T&export=download). Has voices for English (US, UK, AU, CA), Spanish (ES, MX), Portuguese (BR, PT), German, French (FR, CA), Italian, Norwegian, Dutch, Russian, Swedish, Danish, Catalan, Finnish, Japanese, Korean and Chinese (ZH, HK, TW). All voices are female.
18
+ * [Microsoft Speech Platform](https://www.microsoft.com/en-us/download/details.aspx?id=27225) (`msspeech`): Microsoft Server Speech API. Requires [installing a runtime (2.6MB)](https://www.microsoft.com/en-us/download/details.aspx?id=27225). Supports 28 dialects, which can be individually downloaded via [freely available installers](https://www.microsoft.com/en-us/download/details.aspx?id=27224), or, for convenience, bundled as [a single 358MB zip file](https://drive.google.com/u/0/uc?id=1uQdFNxLzUxpaEwVVKhMawys8cIh3F21T&export=download). Has voices for English (US, UK, AU, CA), Spanish (ES, MX), Portuguese (BR, PT), German, French (FR, CA), Italian, Norwegian, Dutch, Russian, Swedish, Danish, Catalan, Finnish, Japanese, Korean and Chinese (ZH, HK, TW). All voices are female
19
19
 
20
20
  **Note**: both these engines require manually installing the [`winax` npm package](https://www.npmjs.com/package/winax) by running `npm install winax -g`.
21
21
 
@@ -30,71 +30,82 @@
30
30
  These are commercial services that require a subscription and an API key to use:
31
31
 
32
32
  * [Google Cloud](https://cloud.google.com/text-to-speech) (`google-cloud`)
33
- * [Azure Cognitive Services](https://azure.microsoft.com/en-us/products/cognitive-services/text-to-speech/) (`microsoft-azure`)
33
+ * [Azure Cognitive Services](https://azure.microsoft.com/en-us/products/ai-services/text-to-speech/) (`microsoft-azure`)
34
34
  * [Amazon Polly](https://aws.amazon.com/polly/) (`amazon-polly`)
35
- * [OpenAI Cloud](https://platform.openai.com/) (`openai-cloud`)
35
+ * [OpenAI Cloud Platform](https://platform.openai.com/) (`openai-cloud`)
36
36
  * [Elevenlabs](https://elevenlabs.io/) (`elevenlabs`)
37
37
 
38
38
  **Cloud services (unofficial)**:
39
39
 
40
40
  These cloud-based engines connect to public cloud APIs that are not officially publicized by their operators. They are included for educational purposes only, and may be removed in the future:
41
41
 
42
- * Google Translate (`google-translate`): used by the [Google Translate web UI](https://translate.google.com/) to speak written text in any one of its supported languages. Offers a single voice for each language (usually female).
43
- * Microsoft Edge (`microsoft-edge`): subset of the Azure Cognitive Services cloud TTS API used by the Microsoft Edge browser as part of its support for the [Web Speech API](https://developer.mozilla.org/en-US/docs/Web/API/Web_Speech_API) and its [Read Aloud](https://www.microsoft.com/en-us/edge/features/read-aloud?form=MT00D8) feature. Using this engine requires a special token, which should be passed via the `microsoftEdge.trustedClientToken` option.
44
- * Streamlabs Polly (`streamlabs-polly`): a public REST API by Streamlabs, primarily intended for generating speech for TTS donations. It includes a few English (US, UK, AU, IN) voices, which are similar to some of the non-neural (Ivona-based) voices offered by Amazon Polly (**Note**: as of April 2024, the public Streamlabs Polly REST API doesn't seem to be accessible anymore).
42
+ * Google Translate (`google-translate`): used by the [Google Translate web UI](https://translate.google.com/) to speak written text in any one of its supported languages. Offers a single voice for each language (usually female)
43
+ * Microsoft Edge (`microsoft-edge`): subset of the Azure Cognitive Services cloud TTS API used by the Microsoft Edge browser as part of its support for the [Web Speech API](https://developer.mozilla.org/en-US/docs/Web/API/Web_Speech_API) and its [Read Aloud](https://www.microsoft.com/en-us/edge/features/read-aloud?form=MT00D8) feature. Using this engine requires a special token, which should be passed via the `microsoftEdge.trustedClientToken` option
44
+ * Streamlabs Polly (`streamlabs-polly`): a public REST API by Streamlabs, primarily intended for generating speech for TTS donations. It includes a few English (US, UK, AU, IN) voices, which are similar to some of the non-neural (Ivona-based) voices offered by Amazon Polly (**Note**: as of April 2024, the public Streamlabs Polly REST API doesn't seem to be accessible anymore)
45
45
 
46
46
  ## Speech-to-text
47
47
 
48
48
  **Offline**:
49
- * [OpenAI Whisper](https://github.com/openai/whisper) (`whisper`): high accuracy transformer-based speech recognition architecture. Supports 99 languages. There are several models of different sizes, some are multilingual, and some are English only (`.en`): `tiny`, `tiny.en`, `base`, `base.en`, `small`, `small.en`, `medium`, `medium.en`, `large`, `large-v1` and `large-v2`, `large-v3`. **Note**: large models are not currently supported by `onnxruntime-node` due to model size restrictions.
50
- * [Whisper.cpp](https://github.com/ggerganov/whisper.cpp) (`whisper.cpp`): a port of the Whisper architecture to C++, by Georgi Gerganov. Supports all Whisper models, including several quantized ones (see full list in the options page). Has various different builds, including CUDA and OpenCL for GPU support.
51
- * [Vosk](https://github.com/alphacep/vosk-api) (`vosk`): models available for 25+ languages. **Note**: the Vosk package is not included in the default installation, but you can add support for it using `npm install @echogarden/vosk -g`. Then, you'll need to manually [download a model](https://alphacephei.com/vosk/models) and specify its directory path via the `vosk.modelPath` option.
52
- * [Silero](https://github.com/snakers4/silero-models) (`silero`): models available for English, Spanish, German and Ukrainian. For [non-commercial use only](https://github.com/snakers4/silero-models/blob/master/LICENSE).
49
+ * [OpenAI Whisper](https://github.com/openai/whisper) (`whisper`): high-accuracy transformer-based speech recognition architecture. TypeScript implementation, with inference done via the [ONNX runtime](https://onnxruntime.ai/). Supports [98 languages](https://platform.openai.com/docs/guides/speech-to-text/supported-languages). There are several models of different sizes, some are multilingual, and some are English only: `tiny`, `tiny.en`, `base`, `base.en`, `small`, `small.en`, `medium`, `medium.en`, `large`, `large-v1` and `large-v2`, `large-v3`. **Note**: large models are not currently supported by `onnxruntime-node` due to model size restrictions
50
+ * [Whisper.cpp](https://github.com/ggerganov/whisper.cpp) (`whisper.cpp`): a C++ port of the Whisper architecture by Georgi Gerganov. Supports all Whisper models, including several quantized ones (see full model list in the [options reference](docs/Options.md)). Has various builds, including CUDA and OpenCL for GPU support
51
+ * [Vosk](https://github.com/alphacep/vosk-api) (`vosk`): models available for 25+ languages. **Note**: the Vosk package is not included in the default installation, but you can add support for it using `npm install @echogarden/vosk -g`. Then, you'll need to manually [download a model](https://alphacephei.com/vosk/models) and specify its directory path via the `vosk.modelPath` option
52
+ * [Silero](https://github.com/snakers4/silero-models) (`silero`): models available for English, Spanish, German and Ukrainian. For [non-commercial use only](https://github.com/snakers4/silero-models/blob/master/LICENSE)
53
53
 
54
54
  **Cloud services**:
55
55
 
56
56
  These are commercial services that require a subscription and an API key to use:
57
57
 
58
58
  * [Google Cloud](https://cloud.google.com/speech-to-text) (`google-cloud`)
59
- * [Azure Cognitive Services](https://azure.microsoft.com/en-us/products/cognitive-services/speech-to-text/) (`microsoft-azure`)
59
+ * [Azure Cognitive Services](https://azure.microsoft.com/en-us/products/ai-services/speech-to-text/) (`microsoft-azure`)
60
60
  * [Amazon Transcribe](https://aws.amazon.com/transcribe/) (`amazon-transcribe`)
61
- * [OpenAI Cloud](https://platform.openai.com/) (`openai-cloud`)
61
+ * [OpenAI Cloud Platform](https://platform.openai.com/) (`openai-cloud`): runs the `large-v2` Whisper model on the cloud
62
62
 
63
63
  ## Speech-to-transcript alignment
64
64
 
65
65
  These engines' goal is to match (or "align") a given spoken recording with a given transcript as closely as possible. They will annotate each word in the transcript with approximate start and end timestamps:
66
66
 
67
- * Dynamic Time Warping (`dtw`): transcript is first synthesized using the eSpeak engine, then the [DTW](https://en.wikipedia.org/wiki/Dynamic_time_warping) alignment algorithm is applied to find the best mapping between the synthesized and original audio frames.
68
- * Dynamic Time Warping with Recognition Assist (`dtw-ra`): recognition is applied to the audio (any recognition engine can be used), then both the ground-truth transcript and the recognized transcript are synthesized using eSpeak. Then, the best mapping is found between the two synthesized waveforms, and the result is mapped back to the original audio using the timing information produced by the recognizer.
69
- * Whisper-based alignment (`whisper`): transcript is tokenized and decoded along with the audio using the Whisper model, then timestamps are extracted from the internal state of the model. **Note**: currently, only supports audio inputs that are 30 seconds or less. Some words or special characters may fail to tokenize due to limitations of the tokenizer used by Whisper.
67
+ * Dynamic Time Warping (`dtw`): transcript is first synthesized using the eSpeak engine, then the [DTW](https://en.wikipedia.org/wiki/Dynamic_time_warping) sequence alignment algorithm is applied to find the best mapping between the synthesized and original audio frames
68
+ * Dynamic Time Warping with Recognition Assist (`dtw-ra`): recognition is applied to the audio (any recognition engine can be used), then both the ground-truth transcript and the recognized transcript are synthesized using eSpeak. Then, the best mapping is found between the two synthesized waveforms, using the DTW algorithm, and the result is remapped back to the original audio using the timing information produced by the recognizer
69
+ * Whisper-based alignment (`whisper`): transcript is first tokenized, then, its tokens are decoded, in order, with a guided approach, using the Whisper model. The resulting token timestamps are then used to derive the timing for each word
70
+
70
71
 
71
72
  ## Speech-to-text translation
72
73
 
73
- * [Whisper](https://github.com/openai/whisper) (`whisper`): the Whisper model can recognize speech in any one of its supported languages and output a transcript directly translated to English. Other languages are not supported as targets.
74
+ **Offline**:
75
+ * [Whisper](https://github.com/openai/whisper) (`whisper`): the Whisper model can recognize speech in any one of its supported languages and output a transcript directly translated to English. Other languages are not supported as targets
74
76
  * [Whisper.cpp](https://github.com/ggerganov/whisper.cpp) (`whisper.cpp`): supports translation to English
75
- * [OpenAI Cloud](https://platform.openai.com/) (`openai-cloud`): OpenAI cloud service. Only translates to English.
77
+
78
+ **Cloud services**:
79
+ * [OpenAI Cloud Platform](https://platform.openai.com/) (`openai-cloud`): runs the `large-v2` Whisper model on the cloud. Only supports English as target
80
+
81
+ ## Speech-to-translated-transcript alignment
82
+
83
+ These goal here is to match (or "align") a given spoken recording in one language, with a given translated transcript in a different language, as closely as possible.
84
+
85
+ * `whisper`: given a spoken recording in any of the [98 languages](https://platform.openai.com/docs/guides/speech-to-text/supported-languages) supported by Whisper, and an English translation of its transcript, the translated transcript is tokenized and then decoded, in order, using a guided approach, with any multilingual Whisper model, set to its `translate` task mode. In this way, the approximate mapping between the spoken audio and each word of the translation is estimated
86
+
76
87
 
77
88
  ## Language detection
78
89
 
79
90
  **Spoken language detection**:
80
- * [Silero Language Classifier](https://github.com/snakers4/silero-vad/wiki/Other-Models) (`silero`): a speech language classification model by Silero.
81
- * [Whisper](https://github.com/openai/whisper) (`whisper`): uses the language token produced by the `whisper` speech recognition model to generate a set of probabilities for the 99 languages it has been trained on.
91
+ * [Whisper](https://github.com/openai/whisper) (`whisper`): uses the language token produced by the `whisper` speech recognition model to generate a set of probabilities for the [98 languages](https://platform.openai.com/docs/guides/speech-to-text/supported-languages) it has been trained on
92
+ * [Silero Language Classifier](https://github.com/snakers4/silero-vad/wiki/Other-Models) (`silero`): a speech language classification model by Silero
82
93
 
83
94
  **Text language detection**:
84
- * [TinyLD](https://www.npmjs.com/package/tinyld) (`tinyld`): a simple language detection library.
85
- * [FastText](https://github.com/facebookresearch/fastText) (`fasttext`): a library for word representations and sentence classification by Facebook research.
95
+ * [TinyLD](https://www.npmjs.com/package/tinyld) (`tinyld`): a simple language detection library
96
+ * [FastText](https://github.com/facebookresearch/fastText) (`fasttext`): a library for word representations and sentence classification by Facebook research
86
97
 
87
98
  ## Voice activity detection
88
99
 
89
- * [WebRTC VAD](https://github.com/dpirch/libfvad) (`webrtc`): a voice activity detector. Originally from the Chromium browser source code.
100
+ * [WebRTC VAD](https://github.com/dpirch/libfvad) (`webrtc`): a voice activity detector. Originally from the Chromium browser source code
90
101
  * [Silero VAD](https://github.com/snakers4/silero-vad) (`silero`): a voice activity detection model by Silero.
91
- * [RNNoise](https://github.com/xiph/rnnoise) (`rnnoise`): uses RNNoise's speech probabilities output for each audio frame as a VAD metric.
92
- * Adaptive Gate (`adaptive-gate`): uses a band-limited adaptive gate to identify activity in the lower voice frequencies. Reliable, but will often pass non-vocal sounds if they are loud enough. Good for clean speech and a cappella singing, where most non-vocal segments are quiet.
102
+ * [RNNoise](https://github.com/xiph/rnnoise) (`rnnoise`): uses RNNoise's speech probabilities output for each audio frame as a VAD metric
103
+ * Adaptive Gate (`adaptive-gate`): uses a band-limited adaptive gate to identify activity in the lower voice frequencies. Reliable, but will often pass non-vocal sounds if they are loud enough. Good for clean speech and a cappella singing, where most non-vocal segments are quiet
93
104
 
94
105
  ## Speech denoising
95
106
 
96
- * [RNNoise](https://github.com/xiph/rnnoise) (`rnnoise`): a noise suppression library based on a recurrent neural network.
107
+ * [RNNoise](https://github.com/xiph/rnnoise) (`rnnoise`): a noise suppression library based on a recurrent neural network
97
108
 
98
109
  ## Source separation
99
110
 
100
- * [MDX-NET](https://github.com/kuielab/mdx-net) (`mdx-net`): Deep learning source separation architecture by [KUIELAB (Korea University)](https://kuielab.github.io/).
111
+ * [MDX-NET](https://github.com/kuielab/mdx-net) (`mdx-net`): deep learning source separation architecture by [KUIELAB (Korea University)](https://kuielab.github.io/)
package/docs/Licenses.md CHANGED
@@ -1,9 +1,8 @@
1
1
  # Echogarden components licensing
2
2
 
3
- ## Engines and libraries:
3
+ ## Engines and libraries
4
4
 
5
5
  * `onnxruntime-node`: [MIT License](https://github.com/microsoft/onnxruntime/blob/main/LICENSE)
6
- * `whisper.cpp`: [MIT License](https://github.com/ggerganov/whisper.cpp/blob/master/LICENSE)
7
6
  * `espeak`: [GNU GPL v3](https://github.com/espeak-ng/espeak-ng/blob/master/COPYING)
8
7
  * `flite`: [BSD License](https://github.com/festvox/flite/blob/master/COPYING)
9
8
  * `pico`: [Apache License 2.0](https://github.com/gmn/nanotts/blob/master/LICENSE)
@@ -31,11 +30,11 @@ All are freely distributable, with varying licenses:
31
30
  * SVOX Pico resources (`pico-`): [Apache License 2.0](https://github.com/gmn/nanotts/blob/master/LICENSE)
32
31
  * Silero VAD (`silero-vad`) and Silero language classifier (`silero-lang-classifier-95`): [MIT License](https://github.com/snakers4/silero-vad/blob/master/LICENSE)
33
32
  * Silero speech recognition models (`silero-en-`, `silero-de-`, `silero-es-`, `silero-ua-`): [BY-NC-SA](https://github.com/snakers4/silero-models/blob/master/LICENSE)
34
- * VITS pre-trained models (`vits-`): licensed under various creative commons licenses: [CC0](https://creativecommons.org/share-your-work/public-domain/cc0/), [CC-BY](https://creativecommons.org/licenses/by/4.0/) and [BY-NC-SA](https://creativecommons.org/licenses/by-nc-sa/4.0/), and few are public domain (you can view the individual license for each model in the model cards on the [samples page](https://rhasspy.github.io/piper-samples/)). The [Piper system](https://github.com/rhasspy/piper) itself is published under the [MIT License](https://github.com/rhasspy/piper/blob/master/LICENSE.md)
33
+ * VITS pre-trained models (`vits-`): licensed under various creative commons licenses: [CC0](https://creativecommons.org/share-your-work/public-domain/cc0/), [CC-BY](https://creativecommons.org/licenses/by/4.0/) and [BY-NC-SA](https://creativecommons.org/licenses/by-nc-sa/4.0/), and few are public domain. You can view the individual license for each model in the model cards on the [Piper samples page](https://rhasspy.github.io/piper-samples/)
35
34
  * Whisper pre-trained models (`whisper-`): [MIT License](https://github.com/openai/whisper/blob/main/LICENSE)
36
35
  * MDX-NET source separation models (`mdxnet-`): [MIT License](https://github.com/kuielab/mdx-net/blob/main/LICENSE)
37
36
 
38
- Tool binary distributions:
37
+ Tool binary distributions
39
38
  * FFmpeg: [LGPL, GPL v2 and GPL v3 Licenses](https://github.com/FFmpeg/FFmpeg)
40
39
  * SoX: [GPL v2 License](https://github.com/chirlu/sox/blob/master/LICENSE.GPL)
41
40
  * whisper.cpp: [MIT License](https://github.com/ggerganov/whisper.cpp/blob/master/LICENSE)
package/docs/Options.md CHANGED
@@ -4,7 +4,7 @@ Here's a detailed reference for all the options accepted by the Echogarden CLI a
4
4
 
5
5
  **Related pages**:
6
6
  * [List of all supported engines](Engines.md)
7
- * [Quick guide for the command line interface](CLI.md)
7
+ * [Quick guide to the command line interface](CLI.md)
8
8
  * [Node.js API reference](API.md)
9
9
 
10
10
  ## Text-to-speech
@@ -46,7 +46,8 @@ Applies to CLI operations: `speak`, `speak-file`, `speak-url`, `speak-wikipedia`
46
46
  * `outputAudioFormat.bitrate`: Custom bitrate for encoding, applies only to `mp3`, `opus`, `m4a`, `ogg`. By default, bitrates are selected between 48Kbps and 64Kbps, to provide a good speech quality while minimizing file size. Optional
47
47
 
48
48
  **VITS**:
49
- * `vits.speakerId`: speaker ID, for VITS models that support multiple speakers. Optional
49
+ * `vits.speakerId`: speaker ID, for VITS models that support multiple speakers. Defaults to `0`
50
+ * `vits.provider`: ONNX execution provider to use. Can be `cpu` or `dml` (https://microsoft.github.io/DirectML/)-based GPU acceleration - Windows only). Using GPU acceleration for VITS may or may not be faster than CPU, depending on your hardware. Defaults to `cpu`
50
51
 
51
52
  **eSpeak**:
52
53
  * `espeak.rate`: speech rate, in eSpeak units. Overrides `speed` when set
@@ -108,7 +109,7 @@ Applies to CLI operations: `speak`, `speak-file`, `speak-url`, `speak-wikipedia`
108
109
  * `microsoftEdge.trustedClientToken`: trusted client token (required). A special token required to use the service
109
110
  * `microsoftEdge.pitchDeltaHz`: pitch delta in Hz. Overrides `pitch` when set
110
111
 
111
- ## Voice list request
112
+ ### Voice list request
112
113
 
113
114
  Applies to CLI operation: `list-voices`, API method: `requestVoiceList`
114
115
 
@@ -146,9 +147,11 @@ Applies to CLI operation: `transcribe`, API method: `recognize`
146
147
  * `whisper.topCandidateCount`: the number of top candidate tokens to consider. Defaults to `5`
147
148
  * `whisper.punctuationThreshold`: the minimal probability for a punctuation token, included in the top candidates, to be chosen unconditionally. A lower threshold encourages the model to output more punctuation symbols. Defaults to `0.2`
148
149
  * `whisper.autoPromptParts`: use previous part's recognized text as prompt for the next part. Disabling this may help to prevent repetition carrying over between parts, in some cases. Defaults to `true`
149
- * `whisper.maxTokensPerPart`: maximum number of tokens to decode for each 30 second audio part. Defaults to `250`
150
- * `whisper.suppressRepetition`: attempt to suppress decoding repeating token patterns. Defaults to `true`
151
- * `whisper.decodeTimestampTokens`: enable/disable decoding of timestamp tokens, since more accurate timing is already extracted via cross-attention weight alignment. For unclear reasons, setting to `false` can significantly reduce the occurrence of hallucinations and token repetition loops, and increases word timestamp accuracy. However, there are cases where this can cause the model to end a part prematurely, especially in singing and less speech-like voice segments, or when there are multiple speakers. Defaults to `true`
150
+ * `whisper.maxTokensPerPart`: maximum number of tokens to decode for each audio part. Defaults to `250`
151
+ * `whisper.suppressRepetition`: attempt to suppress decoding of repeating token patterns. Defaults to `true`
152
+ * `whisper.decodeTimestampTokens`: enable/disable decoding of timestamp tokens. Setting to `false` can reduce the occurrence of hallucinations and token repetition loops, possibly due to the overall reduction in the number of tokens decoded. This has no impact on the accuracy of timestamps, since they are derived independently using cross-attention weights. However, there are cases where this can cause the model to end a part prematurely, especially in singing and less speech-like voice segments, or when there are multiple speakers. Defaults to `true`
153
+ * `whisper.encoderProvider`: identifier for the ONNX execution provider to use with the encoder model. Can be `cpu` or `dml` ([DirectML](https://microsoft.github.io/DirectML/)-based GPU acceleration - Windows only). In general, GPU-based encoding should be significantly faster. Defaults to `cpu`, or `dml` if available
154
+ * `whisper.decoderProvider`: identifier for the ONNX execution provider to use with the decoder model. Can be `cpu` or `dml` (Windows only). Using GPU acceleration for the decoder may be faster than CPU, especially for larger models, but that depends on your particular combination of CPU and GPU. Defaults to `cpu`
152
155
  * `whisper.seed`: provide a custom random seed for token selection when temperature is greater than 0. Uses a constant seed by default to ensure reproducibility
153
156
 
154
157
  **Whisper.cpp**:
@@ -157,7 +160,7 @@ Applies to CLI operation: `transcribe`, API method: `recognize`
157
160
  * `whisperCpp.build`: type of `whisper.cpp` build to use. Can be set `cpu`, `cublas-11.8.0`, `cublas-12.4.0`. By default, builds are auto-selected and downloaded for Windows x64 (`cpu`, `cublas-11.8.0`, `cublas-12.4.0`) and Linux x64 (`cpu`). Using other builds requires providing a custom `executablePath`
158
161
  * `whisperCpp.threadCount`: number of threads to use, defaults to `4`
159
162
  * `whisperCpp.splitCount`: number of splits of the audio data to process in parallel (called `--processors` in the `whisper.cpp` CLI). A value greater than `1` can increase memory use significantly, reduce timing accuracy, and slow down execution in some cases. Defaults to `1` (highly recommended)
160
- * `whisperCpp.enableGPU`: enable GPU processing. Defaults to `true` on CUDA-enabled builds, otherwise `false`
163
+ * `whisperCpp.enableGPU`: enable GPU processing. Setting to `true` will try to use a CUDA build, if available for your system. Defaults to `true` when a CUDA-enabled build is selected via `whisperCpp.build`, otherwise `false`
161
164
  * `whisperCpp.topCandidateCount`: the number of top candidate tokens to consider. Defaults to `5`
162
165
  * `whisperCpp.beamCount`: the number of decoding paths to use during beam search. Defaults to `5`
163
166
  * `whisperCpp.repetitionThreshold`: minimal repetition / compressibility score to cause a decoded segment to be discarded. Defaults to `2.4`
@@ -170,6 +173,7 @@ Applies to CLI operation: `transcribe`, API method: `recognize`
170
173
 
171
174
  **Silero**:
172
175
  * `silero.modelPath`: path to a Silero model. Note that latest `en`, `de`, `fr` and `uk` models are automatically installed when needed based on the selected language. This should only be used to manually specify a different model, otherwise specify `language` instead
176
+ * `silero.provider`: ONNX execution provider to use. Can be `cpu` or `dml` (Windows only). Defaults to `cpu`, or `dml` if available
173
177
 
174
178
  **Google Cloud**:
175
179
  * `googleCloud.apiKey`: Google Cloud API key (required)
@@ -192,7 +196,7 @@ Applies to CLI operation: `transcribe`, API method: `recognize`
192
196
  * `openAICloud.model`: model to use. Can only be `whisper-1`
193
197
  * `openAICloud.organization`: organization identifier. Optional
194
198
  * `openAICloud.baseURL`: override the default base URL used by the API. Optional
195
- * `openAICloud.temperature`: temperature. Choosing `0` uses a dynamic temperature approach. Defaults to `0.0`
199
+ * `openAICloud.temperature`: temperature. Choosing `0` uses a dynamic temperature approach. Defaults to `0`
196
200
  * `openAICloud.prompt`: initial prompt for the model. Optional
197
201
  * `openAICloud.timeout`: request timeout. Optional
198
202
  * `openAICloud.maxRetries`: maximum retries on failure. Defaults to 10
@@ -220,12 +224,18 @@ Applies to CLI operation: `align`, API method: `align`
220
224
  * `dtw.granularity`: adjusts the MFCC frame width and hop size based on the profile selected. Can be set to either `auto` (auto-selected based on audio duration and task), `xx-low` (400ms width, 160ms hop), `x-low` (200ms width, 80ms hop), `low` (100ms width, 40ms hop), `medium` (50ms width, 20ms hop), `high` (25ms width, 10ms hop), `x-high` (20ms width, 5ms hop). For multi-pass processing, multiple granularities can be provided, like `dtw.granularity=['low','high']`. Defaults to `auto`.
221
225
  * `dtw.windowDuration`: maximum duration (in seconds) of the Sakoe-Chiba window when performing DTW alignment. Higher values consume quadratically larger amounts of memory. The estimated memory requirement is shown in the log before alignment starts. Recommended to be set to at least 10% - 20% of total audio duration. For multi-pass processing, multiple durations can be provided, like `dtw.windowDuration=[240,20]`. Auto-selected by default
222
226
 
223
- **DTW-RA only**:
227
+ **DTW-RA**:
224
228
  * `recognition`: prefix to provide recognition options when using `dtw-ra` method, for example: setting `recognition.engine = whisper` and `recognition.whisper.model = base.en`
225
229
  * `dtw.phoneAlignmentMethod`: algorithm to use when aligning phones: can either be set to `dtw` or `interpolation`. Defaults to `dtw`
226
230
 
227
- **Whisper alignment only**:
228
- * `whisper`: prefix to provide Whisper options when the `whisper` alignment engine is used (does not apply to `dtw-ra` when `whisper` engine is used, for that use `recognition.whisper` prefix instead).
231
+ **Whisper**:
232
+
233
+ Applies to the `whisper` engine only. To provide Whisper options for `dtw-ra`, use `recognition.whisper` instead.
234
+
235
+ * `whisper.model`: Whisper model to use. Defaults to `tiny` or `tiny.en`
236
+ * `whisper.endTokenThreshold`: minimal probability to accept an end-of-text token for a recognized part. The probability is measured via the softmax between the end-of-text token's logit and the second highest logit. You can try to adjust this threshold in cases the model is ending a part with too few, or many tokens decoded. Defaults to `0.9`. On the last audio part, it is always effectively set to `Infinity`, to ensure the remaining transcript tokens are decoded in full
237
+ * `whisper.encoderProvider`: encoder ONNX provider. See details in recognition section above
238
+ * `whisper.decoderProvider`: decoder ONNX provider. See details in recognition section above
229
239
 
230
240
 
231
241
  ## Speech-to-text translation
@@ -254,6 +264,25 @@ Applies to CLI operation: `translate-speech`, API method: `translateSpeech`
254
264
 
255
265
  * `openAICloud`: prefix to provide options for OpenAI cloud. Same options as detailed in the recognition section above
256
266
 
267
+ ## Speech-to-translated-transcript alignment
268
+
269
+ Applies to CLI operation: `align-translation`, API method: `alignTranslation`
270
+
271
+ **General**:
272
+ * `engine`: alignment algorithm to use, can only be `whisper`. Defaults to `whisper`
273
+ * `language`: language code for the source audio ([ISO 639-1](https://en.wikipedia.org/wiki/List_of_ISO_639-1_codes)), like `en`, `fr`, `zh`, etc. Auto-detected from audio if not set
274
+ * `crop`: crop to active parts using voice activity detection before starting. Defaults to `true`
275
+ * `isolate`: apply source separation to isolate voice before starting alignment. Defaults to `false`
276
+ * `subtitles`: prefix to provide options for subtitles. Options detailed in section for subtitles
277
+ * `vad`: prefix to provide options for voice activity detection when `crop` is set to `true`. Options detailed in section for voice activity detection
278
+ * `sourceSeparation`: prefix to provide options for source separation when `isolate` is set to `true`. Options detailed in section for source separation
279
+
280
+ **Whisper**:
281
+ * `whisper.model`: Whisper model to use. Only multilingual models can be used. Defaults to `tiny`
282
+ * `whisper.endTokenThreshold`: see details in the alignment section above
283
+ * `whisper.encoderProvider`: encoder ONNX execution provider. See details in recognition section above
284
+ * `whisper.decoderProvider`: decoder ONNX execution provider. See details in recognition section above
285
+
257
286
  ## Language detection
258
287
 
259
288
  ### Speech language detection
@@ -270,6 +299,11 @@ Applies to CLI operation: `detect-speech-langauge`, API method: `detectSpeechLan
270
299
  **Whisper**:
271
300
  * `whisper.model`: Whisper model to use. See model list in the recognition section
272
301
  * `whisper.temperature`: impacts the distribution of candidate languages when applying the softmax function to compute language probabilities over the model output. Higher temperature causes the distribution to be more uniform, while lower temperature causes it to be more strongly weighted towards the best scoring candidates. Defaults to `1.0`
302
+ * `whisper.encoderProvider`: encoder ONNX execution provider. See details in recognition section above
303
+ * `whisper.decoderProvider`: decoder ONNX execution provider. See details in recognition section above
304
+
305
+ **Silero**:
306
+ * `silero.provider`: ONNX execution provider to use. Can be `cpu` or `dml` (Windows only). Using GPU may be faster, but the initialization overhead is larger. **Note**: `dml` provider seems to be unstable at the moment for this model. Defaults to `cpu`
273
307
 
274
308
  ### Text language detection
275
309
 
@@ -294,6 +328,7 @@ Applies to CLI operation: `detect-voice-activity`, API method: `detectVoiceActiv
294
328
 
295
329
  **Silero**:
296
330
  * `silero.frameDuration`: Silero frame duration (ms). Can be `30`, `60` or `90`. Defaults to `90`
331
+ * `silero.provider`: ONNX provider to use. Can be `cpu` or `dml` (Windows only). Using GPU is likely to be slower than CPU due to inference being independently executed on each audio frame. Defaults to `cpu` (recommended)
297
332
 
298
333
  ## Speech denoising
299
334
 
@@ -319,6 +354,7 @@ Applies to CLI operation: `isolate`, API method: `isolate`
319
354
  **MDX-NET**:
320
355
 
321
356
  * `mdxNet.model`: model to use. Currently available models are `UVR_MDXNET_1_9703`, `UVR_MDXNET_2_9682`, `UVR_MDXNET_3_9662`, `UVR_MDXNET_KARA`. Defaults to `UVR_MDXNET_1_9703`
357
+ * `mdxNet.provider`: ONNX execution provider to use. Can be `cpu` or `dml` ([DirectML](https://microsoft.github.io/DirectML/), Windows only). **Note**: `dml` provider seems to be unstable with MDX-NET models at the moment. Defaults to `cpu`
322
358
 
323
359
  # Common options
324
360
 
package/docs/Releases.md CHANGED
@@ -1,5 +1,9 @@
1
1
  # Release notes
2
2
 
3
+ ## Releases after `1.0.0`
4
+
5
+ For releases after `1.0.0`, see the [GitHub releases page](https://github.com/echogarden-project/echogarden/releases).
6
+
3
7
  ## `1.0.0` (April 12, 2024)
4
8
 
5
9
  **New features**:
package/docs/Server.md CHANGED
@@ -1,6 +1,8 @@
1
- # Starting and interfacing with the WebSocket server
1
+ # WebSocket server API reference
2
2
 
3
- This is a guide for the WebSocket server protocol. The protocol is still in an early development stage and may change on future releases.
3
+ This is a guide to the WebSocket server protocol.
4
+
5
+ **Note**: The protocol is still in early development and may change in future releases. Many features are currently missing, and the server hasn't been thoroughly tested.
4
6
 
5
7
  ## Starting the server
6
8
 
@@ -36,10 +38,10 @@ ws.on("open", async () => {
36
38
  })
37
39
  ```
38
40
 
39
- In the future, this module may be separated to an independent lightweight package.
40
-
41
- **TODO**: Document using the client class with a background worker.
42
- **TODO**: Add support for cancellation signals in the client class.
41
+ **TODO**:
42
+ * Separate the client to an independent, lightweight, Node.js package, with browser compatibility
43
+ * Add support for cancellation signals
44
+ * Document how to use with a background worker
43
45
 
44
46
  ## Protocol details
45
47
 
package/docs/Tasklist.md CHANGED
@@ -4,13 +4,13 @@
4
4
 
5
5
  ### Alignment
6
6
 
7
- * In DTW-RA, recognition transcript including something like "Question 2.What does Juan", where "2.What" has a point in the middle, is breaking playback of the timeline.
8
- * DTW-RA will not work correctly with Polish language texts, due to issues with the eSpeak engine pronouncing `|` characters, which are intended to be used as separators and ignored by all other eSpeak languages.
7
+ * In DTW-RA, recognition transcript including something like "Question 2.What does Juan", where "2.What" has a point in the middle, is breaking playback of the timeline
8
+ * DTW-RA will not work correctly with Polish language texts, due to issues with the eSpeak engine pronouncing `|` characters, which are intended to be used as separators and ignored by all other eSpeak languages
9
9
 
10
10
  ### Synthesis
11
11
 
12
12
  ### Phoneme processing
13
- * IPA -> Kirshenbaum translation is still not completely similar to what is output by eSpeak. Also, in rare situations, it outputs characters that are not accepted by eSpeak and eSpeak errors. Investigate when that happens and how to improve on this.
13
+ * IPA -> Kirshenbaum translation is still not completely similar to what is output by eSpeak. Also, in rare situations, it outputs characters that are not accepted by eSpeak and eSpeak errors. Investigate when that happens and how to improve on this
14
14
 
15
15
  ### Browser extension
16
16
  * Investigate why WebSpeech events sometimes completely stop working in the middle of an utterance for no apparent reason. Sometimes this is permanent, until the extension is restarted. Is this a browser issue?
@@ -34,17 +34,9 @@
34
34
 
35
35
  ## Features and enhancements
36
36
 
37
- ### Server
38
- * Option to allow or disallow local file paths as arguments to API methods (as a security safeguard)
39
-
40
- ### Worker
41
- * Add cancellation checks in more operations
42
- * Support more operations
43
-
44
37
  ### CLI
45
38
  * Show names of files written do disk. This is useful for cases where a file is auto-renamed to prevent overwriting existing data
46
39
  * Restrict input media file extensions to ensure that invalid files are not passed to FFmpeg
47
- * Mode to print IPA words when speaking
48
40
  * Consider what to do with non-supported templates like `[hello]`
49
41
  * Show a message when a new version is available
50
42
  * Figure out which terminal outputs should go to stdout, or if that's a good idea at all
@@ -55,7 +47,8 @@
55
47
  * Suggest possible correction on the error of not using `=`, e.g. `speed 0.9` instead of `speed=0.9`
56
48
  * Multiple configuration files in `--config=..` taking precedence by order
57
49
  * Generate JSON configuration file schema
58
- * Use a file type detector like `file-type` that uses magic numbers to detect the type of a binary file regardless of its extension. This would help to give better error messages when the given file type is wrong.
50
+ * Use a file type detector like `file-type` that uses magic numbers to detect the type of a binary file regardless of its extension. This would help to give better error messages when the given file type is wrong
51
+ * Mode to print IPA words when speaking
59
52
 
60
53
  ### CLI / playback
61
54
  * Option to set audio output device for playback
@@ -64,7 +57,7 @@
64
57
  * Add phone playback support
65
58
 
66
59
  ### CLI / `speak`
67
- * Add support for sentence templates, like `echogarden speak-file text.txt /parts/[sentence].wav`.
60
+ * Add support for sentence templates, like `echogarden speak-file text.txt /parts/[sentence].wav`
68
61
 
69
62
  ### CLI / `speak-wikipedia`
70
63
  * Correctly detect language when a Wikipedia URL is passed instead of an article name
@@ -84,7 +77,7 @@
84
77
  * `play-with-timeline`: Preview timeline in terminal
85
78
  * `subtitles-to-text`, `subtitles-to-timeline`, `srt-to-vtt`, `vtt-to-srt`
86
79
  * `text-to-ipa`, `arpabet-to-ipa`, `ipa-to-arpabet`
87
- * `phonemize-text`
80
+ * `phonemize`
88
81
  * `normalize-text`
89
82
  * `transcribe-youtube`: Transcribe the audio in a YouTube video (requires fetching the audio somehow - which can't be done using the normal YouTube API)
90
83
  * `speak-youtube-subtitles`: To speak the subtitles of a YouTube video
@@ -98,7 +91,7 @@
98
91
  * Accept voice list caching options in `SynthesisOptions`
99
92
 
100
93
  ### Package manager
101
- * Better error message when package is not found remotely. Currently, it just gives a `404 not found` without any other information.
94
+ * Better error message when package is not found remotely. Currently, it just gives a `404 not found` without any other information
102
95
  * Retry on network failure
103
96
 
104
97
  ### Speech language detection
@@ -111,7 +104,7 @@
111
104
  * See if it's possible to reliably use eSpeak as a segmentation engine
112
105
 
113
106
  ### Subtitles
114
- * Split long words if needed. This is especially important for Chinese
107
+ * Split long words if needed. This is especially important for Chinese and Japanese
115
108
  * If a subtitle is too short and at the end of the audio, try to extend it back if possible (for example, if the previous subtitle is already extended, take back from it)
116
109
  * Decide how many punctuation characters to allow before breaking to a new line (currently it's infinite)
117
110
  * Add more clause separators, for even more special cases
@@ -119,26 +112,26 @@
119
112
  * Parse VTT's language
120
113
 
121
114
  ### Synthesis
122
- * Option to disable alignment (only for some engines). Alternative: use a low granularity setting that is very fast to compute
115
+ * Option to disable alignment (only for some engines). Alternative: use a low granularity DTW setting that is very fast to compute
123
116
  * Find places to add commas (",") to improve speech fluency. VITS voices don't normally add speech breaks if there is no punctuation
124
- * An isolated dash " - " can be converted to a " , " to ensure there's a break in the speech.
117
+ * An isolated dash " - " can be converted to a " , " to ensure there's a break in the speech
125
118
  * Ensure abbreviations like "Ph.d" or names like are segmented and read correctly (does `cldr` treat it as a word? Maybe eSpeak doesn't recognize it as a word). "C#" and ".NET" as well
126
119
  * Find way to manually reset voice list cache
127
120
  * When synthesized text isn't pre-split to sentences, apply sentence splits by using the existing method to convert the output of word timelines to sentence/segment timelines
128
121
  * Some `sapi` voices and `msspeech` languages output phones that are converted to Microsoft alphabet, not IPA symbols. Try to see if these can be translated to IPA
129
122
  * Decide whether asterisk `*` should be spoken when using `speak-url` or `speak-wikipedia`
130
- * Add partial SSML support for all engines. In particular, allow changing language or voice using the `<voice>` and `<lang>` tags, `<say-as>` and `<phoneme>` where possible.
131
- * Try to remove reliance on `()` after `.` character hack in `EspeakTTS.synthesizeFragments`.
132
- * eSpeak IPA output puts stress marks on vowels, not syllables - which is the standard for IPA. Consider how to make a conversion to and from these two approaches (possibly detect it automatically).
123
+ * Add partial SSML support for all engines. In particular, allow changing language or voice using the `<voice>` and `<lang>` tags, `<say-as>` and `<phoneme>` where possible
124
+ * Try to remove reliance on `()` after `.` character hack in `EspeakTTS.synthesizeFragments`
125
+ * eSpeak IPA output puts stress marks on vowels, not syllables - which is the standard for IPA. Consider how to make a conversion to and from these two approaches (possibly detect it automatically)
133
126
  * Decide if `msspeech` engine should be selected if available. This would require attempting to load a matching voice, and falling back if it is not installed
134
127
  * Speaker-specific voice option
135
128
  * Use VAD on the synthesized audio file to get more accurate sentence or word segmentation
136
- * When `splitToSentences` is set to `false`, the timeline doesn't include proper sentences. Find a way to pass larger sections to the TTS, but still have proper sentences in the timeline.
129
+ * When `splitToSentences` is set to `false`, the timeline doesn't include proper sentences. Find a way to pass larger sections to the TTS, but still have proper sentences in the timeline
137
130
 
138
131
  ### Synthesis / preprocessing
139
132
  * Extend the heteronyms JSON document with additional words like "conducts", "survey", "protest", "transport", "abuse", "combat", "combats", "affect", "contest", "detail", "marked", "contrast", "construct", "constructs", "console", "recall", "permit", "permits", "prospect", "prospects", "proceed", "proceeds", "invite", "reject", "deserts", "transcript", "transcripts", "compact", "impact", "impacts"
140
133
  * Full date normalization (e.g. `21 August 2023`, `21 Aug 2023`, `August 21, 2023`)
141
- * Add support for capitalized-only rules, and possibly also all uppercase / all lowercase rules.
134
+ * Add support for capitalized-only rules, and possibly also all uppercase / all lowercase rules
142
135
  * Add support for multiple words in `precededBy` and `succeededBy`
143
136
  * Support substituting to graphemes in lexicons, not only phonemes
144
137
  * Cache lexicons to avoid parsing the JSON each time it is loaded (this may not be needed for if the file is relatively small)
@@ -160,30 +153,33 @@
160
153
  ### Recognition
161
154
  * Recognized word entries that span VAD segment boundaries can be split
162
155
  * Show alternatives when playing in the CLI. Clear current line and rewrite already printed text for alternatives during the speech recognition process
163
- * Option to split recognized audio to segments or sentences, as is done with synthesized audio
164
- * Try to exclude the timing for trailing punctuation tokens in words that contain them. This can help narrow down the end timestamp to cover the word more tightly
165
156
 
166
157
  ### Recognition / Whisper
167
- * May get stuck in a token repeat loop when silence or non-speech segment encountered in audio. Decide what to do
158
+ * Whisper's Chinese and Japanese output can be split to words in a more accurate way. Consider using a dedicated segmentation library to perform the segmentation in character sequences that have no spaces within them
168
159
  * Automatically disable using previous section recognized transcript as prompt for the next section when lots of repetition occurred in previous section
169
160
  * Cache last model (if enough memory available)
170
- * Bring back option to use eSpeak DTW based alignment on segments, as an alternative approach
171
161
  * The segment output can be used to split to segments, otherwise it is possible to try to guess using pause lengths or voice activity detection
172
- * Use compression ratios on the decoded tokens of individual segments and discard if too much repetition detected
173
- * Way to specify model size only, such that the English-only/multilingual variant would be automatically selected for sizes other than `tiny`?
174
- * Timestamps extracted from cross-attention are still not as accurate as what the official Python implementation gets. Try to see if you can make them better
175
- * Whisper's Chinese output can be split to words in a more accurate way. Consider using a dedicated segmentation library to perform the segmentation in character sequences that have no spaces within them
162
+ * Bring back option to use eSpeak DTW based alignment on segments, as an alternative approach
176
163
 
177
164
  ### Alignment
178
165
 
179
166
  * Aligned words entries that span VAD boundaries may be split
180
167
 
181
168
  ### Alignment / DTW-RA
182
- * Remove emojis and other special characters, that are not likely to be pronounced in the speech, from the transcript timeline before it is synthesized. For example Whisper may produce 'note' emojis when it detects singing or music. Pronouncing them reduces the accuracy of the alignment
183
- * Optional mode to pass Whisper a special vocabulary of tokens that can appear in the transcript. All other tokens would be suppressed
184
169
 
185
170
  ### Alignment / Whisper
186
- * New mode to decode the transcript tokens in order using a more standard decoding approach (updating the KV cache at each step). This would allow audio inputs longer than 30 seconds. See if this produces better results
171
+
172
+ ### Source separation / MDX-NET
173
+ * Since MDX-NET requires FFT with large window sizes, the FFT computation overhead currently acts as a bottleneck, especially when GPU is used for inference. Currently it uses a WASM port of KissFFT, running on a single thread, which is still relatively fast. To get higher performance, try to (optionally) use native, SIMD optimized FFT like FFTW3 via a NAPI addon, with multi-threading enabled
174
+ * Option to customize overlap
175
+ * Add more models
176
+
177
+ ### Server
178
+ * Option to allow or disallow local file paths as arguments to API methods (as a security safeguard)
179
+
180
+ ### Worker
181
+ * Add cancellation checks in more operations
182
+ * Support more operations
187
183
 
188
184
  ### Browser extension
189
185
  * Options UI
@@ -212,7 +208,7 @@
212
208
  * See if the installation of `winax` can be automated and only initiate if it is in a Windows environment
213
209
  * Ensure that all modules have no internal state other than caching
214
210
  * Start thinking about some modules being available in the browser. Which node core APIs the use? Which of them can be polyfilled, and which cannot?
215
- * Change all the Emscripten WASM modules to use the `EXPORT_ES6=1` flag and rebuild them. Support for node.js modules was only added in September 2022 (https://github.com/emscripten-core/emscripten/pull/17915).
211
+ * Change all the Emscripten WASM modules to use the `EXPORT_ES6=1` flag and rebuild them. Support for node.js modules was added in September 2022 (https://github.com/emscripten-core/emscripten/pull/17915)
216
212
  * Remove built-in voices from `flite` to reduce size?
217
213
  * Slim down `kuromoji` package to reduce base installation size
218
214
 
@@ -224,7 +220,6 @@
224
220
  * Test everything's fine on macOS
225
221
  * Test that cloud services all still work correctly, especially with SSML inputs
226
222
 
227
-
228
223
  ## Future features and enhancements
229
224
 
230
225
  ### CLI
@@ -259,18 +254,10 @@
259
254
  * Investigate exporting Whisper models to 16-bit quantized ONNX or a mix of 16-bit and 32-bit
260
255
 
261
256
  ### Alignment
262
- * Implement alignment with speech-to-text translation assistance, which would enable multilingual subtitle replacement for translated subtitles
263
257
  * Method to align audio file to audio file
264
- * Make `dtw` mode work with more speech synthesizers to produce its reference
265
- * Predict timing for individual letters (graphemes) based on phoneme timestamps
258
+ * Allow `dtw` mode work with more speech synthesizers to produce its reference
259
+ * Predict timing for individual letters (graphemes) based on phoneme timestamps (especially useful for Chinese and Japanese)
266
260
 
267
- ### Voice activity detection
268
-
269
- * Whisper-based VAD. Use Whisper's 'no speech' token to determine if the audio contains speech
270
-
271
- ### Source separation
272
- * Option to customize overlap
273
- * Add more MDX-NET models
274
261
 
275
262
  ## Possible new engines or platforms
276
263
 
@@ -280,13 +267,13 @@
280
267
  * Coqui STT server connection
281
268
  * [MarbleNet VAD](https://github.com/NVIDIA/NeMo/blob/main/tutorials/asr/Online_Offline_Microphone_VAD_Demo.ipynb), included of the NVIDIA NeMo framework, can be exported to ONNX
282
269
  * Silero text enhancement engine can be ported to ONNX
283
- * See what can be done on supporting WinRT speech: in particular `windows.media.speechsynthesis` and `windows.media.speechrecognition` support, possibly using NodeRT or some other method.
284
- * Figure out how to support `julius` speech recognition via WASM.
270
+ * See what can be done for supporting WinRT speech: in particular `windows.media.speechsynthesis` and `windows.media.speechrecognition` support, possibly using NodeRT or some other method
271
+ * Figure out how to support `julius` speech recognition via WASM
285
272
  * Any way to support RHVoice?
286
273
 
287
274
  ## Maybe?
288
275
 
289
- * Using a machine translation model to provide speech translation to languages other than English?
276
+ * Using a machine translation model to provide speech translation to languages other than English? How would the timing be determined?
290
277
  * Is it possible to get sentence boundaries without punctuation using NLP techniques like part of speech tagging?
291
278
 
292
279
  ## May or may not be good ideas
@@ -298,10 +285,10 @@
298
285
 
299
286
  * Support alignment of EPUB 3 eBooks with corresponding audiobook
300
287
  * Voice cloning
301
- * Speech to speech voice conversion
302
- * Speech-to-speech translation (need to find a good model)
288
+ * Speech-to-speech voice conversion
289
+ * Speech-to-speech translation
303
290
  * HTML generator, that includes text and audio, with playback and word highlighting
304
291
  * Video generator
305
292
  * Desktop app that uses the tool to transcribe the PC audio output
306
- * Special method to use time stretching to project between different utterances of the same text
293
+ * Special method to use time stretching to project between different aligned utterances of the same text
307
294
  * Is it possible to combine the Silero speech recognizer and a language model and try to perform Viterbi decoding to find alignments?
package/docs/Technical.md CHANGED
@@ -26,7 +26,7 @@ The base installed (uncompressed) size, including dependencies, is around 270MB.
26
26
 
27
27
  Currently, the largest contributors to the size are:
28
28
 
29
- * `onnxruntime-node` (NAPI): 92MB
29
+ * `onnxruntime-node` (NAPI): 133MB
30
30
  * `kuromoji` (JavaScript) 40MB
31
31
  * `flite-wasi` (WASI): 20MB
32
32
  * `espeak-ng-emscripten` (WASM): 18MB