jaketts 1.0.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- jaketts-1.0.0/LICENSE +21 -0
- jaketts-1.0.0/PKG-INFO +18 -0
- jaketts-1.0.0/README.md +98 -0
- jaketts-1.0.0/jaketts.egg-info/PKG-INFO +18 -0
- jaketts-1.0.0/jaketts.egg-info/SOURCES.txt +11 -0
- jaketts-1.0.0/jaketts.egg-info/dependency_links.txt +1 -0
- jaketts-1.0.0/jaketts.egg-info/entry_points.txt +3 -0
- jaketts-1.0.0/jaketts.egg-info/requires.txt +6 -0
- jaketts-1.0.0/jaketts.egg-info/top_level.txt +1 -0
- jaketts-1.0.0/jaketts.py +421 -0
- jaketts-1.0.0/pyproject.toml +3 -0
- jaketts-1.0.0/setup.cfg +4 -0
- jaketts-1.0.0/setup.py +24 -0
jaketts-1.0.0/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Jake
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
jaketts-1.0.0/PKG-INFO
ADDED
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: jaketts
|
|
3
|
+
Version: 1.0.0
|
|
4
|
+
Summary: Jake's Local CLI Text-to-Speech tool powered by Kokoro-82M
|
|
5
|
+
Author: Jake
|
|
6
|
+
Requires-Python: >=3.10,<3.13
|
|
7
|
+
License-File: LICENSE
|
|
8
|
+
Requires-Dist: kokoro>=0.7.0
|
|
9
|
+
Requires-Dist: sounddevice>=0.4.0
|
|
10
|
+
Requires-Dist: soundfile>=0.4.0
|
|
11
|
+
Requires-Dist: numpy<2.0.0,>=1.20.0
|
|
12
|
+
Requires-Dist: torch>=2.0.0
|
|
13
|
+
Requires-Dist: tqdm>=4.65.0
|
|
14
|
+
Dynamic: author
|
|
15
|
+
Dynamic: license-file
|
|
16
|
+
Dynamic: requires-dist
|
|
17
|
+
Dynamic: requires-python
|
|
18
|
+
Dynamic: summary
|
jaketts-1.0.0/README.md
ADDED
|
@@ -0,0 +1,98 @@
|
|
|
1
|
+
# π jaketts
|
|
2
|
+
|
|
3
|
+
`jaketts` (and its short alias `jtts`) is a lightning-fast, ultra-realistic command-line text-to-speech utility for macOS. Powered by the open-weight **Kokoro-82M** neural engine, it synthesizes highly natural, human-like narration directly in your terminalβcompletely locally, completely offline, and without requiring any API keys.
|
|
4
|
+
|
|
5
|
+
By default, it features a rich, deep British voice (`bm_george`) optimized for storytelling and audiobook-style ingestion.
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## β¨ Features
|
|
10
|
+
* π **Live Playback:** Streams synthesized speech directly through your default Mac speakers using system memory.
|
|
11
|
+
* πΎ **Audio Export:** Automatically stiches sentence fragments together to output crisp, high-fidelity `.wav` files via the `-o` or `--output` flags.
|
|
12
|
+
* π **Smart Input Handling:** Accepts raw text strings or directly parses `.txt` files seamlessly.
|
|
13
|
+
* π **Global Execution:** Runs from any directory on your machine after running a single setup script.
|
|
14
|
+
|
|
15
|
+
---
|
|
16
|
+
|
|
17
|
+
## π οΈ Prerequisites & Installation
|
|
18
|
+
|
|
19
|
+
### 1. Install System Dependencies
|
|
20
|
+
`jaketts` relies on `espeak-ng` for its phonetic mapping backend. Install it via Homebrew:
|
|
21
|
+
```bash
|
|
22
|
+
brew install espeak-ng
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
### 2. Clone the Repository & Configure Python
|
|
26
|
+
Because underlying text processing tools require pre-compiled binaries, we **strongly recommend** running this tool using a stable **Python 3.11 or 3.12** environment (managed easily via `mise` or native virtual environments).
|
|
27
|
+
|
|
28
|
+
```bash
|
|
29
|
+
# Clone the workspace
|
|
30
|
+
git clone https://github.com
|
|
31
|
+
cd jaketts
|
|
32
|
+
|
|
33
|
+
# (Optional) If using mise, lock this folder to Python 3.12
|
|
34
|
+
mise use python@3.12
|
|
35
|
+
|
|
36
|
+
# Initialize your local application sandbox
|
|
37
|
+
python -m venv .venv
|
|
38
|
+
source .venv/bin/activate
|
|
39
|
+
|
|
40
|
+
# Install requirements
|
|
41
|
+
pip install --upgrade pip setuptools wheel
|
|
42
|
+
pip install -r requirements.txt
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
### 3. Make the Command Global
|
|
46
|
+
Run the included installation script to register `jaketts` and `jtts` directly inside your machine's system execution paths:
|
|
47
|
+
```bash
|
|
48
|
+
chmod +x install.sh
|
|
49
|
+
./install.sh
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
---
|
|
53
|
+
|
|
54
|
+
## π Usage Examples
|
|
55
|
+
|
|
56
|
+
*Note: The very first time you execute a command, the app will automatically download the ~340MB neural model weights. Subsequent runs will process completely offline and execute instantly.*
|
|
57
|
+
|
|
58
|
+
### Basic Shorthand Execution (Live Playback)
|
|
59
|
+
Pass any text string directly to your shorter alias:
|
|
60
|
+
```bash
|
|
61
|
+
jtts "Three Rings for the Elven-kings under the sky, Seven for the Dwarf-lords in their halls of stone."
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
### Reading from Text Files
|
|
65
|
+
Pass a text file to read it straight out loud over your speakers:
|
|
66
|
+
```bash
|
|
67
|
+
jtts story.txt
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
### Saving Audio Files
|
|
71
|
+
To bypass the speakers and save directly to an output file, use the `-o` or `--output` flag:
|
|
72
|
+
```bash
|
|
73
|
+
# Saves the track using the default name 'output.wav'
|
|
74
|
+
jtts -o story.txt
|
|
75
|
+
|
|
76
|
+
# Saves the track using a custom file name
|
|
77
|
+
jtts -o fantasy_intro.wav "Deep in the land of Mordor where the Shadows lie."
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
### Overriding the Default Voice
|
|
81
|
+
If you want to shift away from the deep British male voice (`bm_george`), you can explicitly pass an alternative voice ID using the `-v` flag:
|
|
82
|
+
```bash
|
|
83
|
+
# Switch to a conversational American Female profile
|
|
84
|
+
jtts -v af_sarah "Hello from a clear American voice configuration."
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
---
|
|
88
|
+
|
|
89
|
+
## ποΈ Voice Directory Reference
|
|
90
|
+
You can swap to any of Kokoro's built-in regional presets using the `-v` flag. Highly recommended profiles include:
|
|
91
|
+
|
|
92
|
+
| Voice ID | Accent | Gender | Best Used For... |
|
|
93
|
+
| :--- | :--- | :--- | :--- |
|
|
94
|
+
| `bm_george` | British | Male | **[Default]** Deep, steady, and clear audiobook tone. |
|
|
95
|
+
| `bm_lewis` | British | Male | Warm, documentary-ready cinematic narrator. |
|
|
96
|
+
| `bf_emma` | British | Female | Crisp, highly professional narrative spacing. |
|
|
97
|
+
| `af_heart` | American | Female | Highly melodic, emotional, and expressive. |
|
|
98
|
+
| `am_adam` | American | Male | Deep, classic American broadcast tone. |
|
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: jaketts
|
|
3
|
+
Version: 1.0.0
|
|
4
|
+
Summary: Jake's Local CLI Text-to-Speech tool powered by Kokoro-82M
|
|
5
|
+
Author: Jake
|
|
6
|
+
Requires-Python: >=3.10,<3.13
|
|
7
|
+
License-File: LICENSE
|
|
8
|
+
Requires-Dist: kokoro>=0.7.0
|
|
9
|
+
Requires-Dist: sounddevice>=0.4.0
|
|
10
|
+
Requires-Dist: soundfile>=0.4.0
|
|
11
|
+
Requires-Dist: numpy<2.0.0,>=1.20.0
|
|
12
|
+
Requires-Dist: torch>=2.0.0
|
|
13
|
+
Requires-Dist: tqdm>=4.65.0
|
|
14
|
+
Dynamic: author
|
|
15
|
+
Dynamic: license-file
|
|
16
|
+
Dynamic: requires-dist
|
|
17
|
+
Dynamic: requires-python
|
|
18
|
+
Dynamic: summary
|
|
@@ -0,0 +1,11 @@
|
|
|
1
|
+
LICENSE
|
|
2
|
+
README.md
|
|
3
|
+
jaketts.py
|
|
4
|
+
pyproject.toml
|
|
5
|
+
setup.py
|
|
6
|
+
jaketts.egg-info/PKG-INFO
|
|
7
|
+
jaketts.egg-info/SOURCES.txt
|
|
8
|
+
jaketts.egg-info/dependency_links.txt
|
|
9
|
+
jaketts.egg-info/entry_points.txt
|
|
10
|
+
jaketts.egg-info/requires.txt
|
|
11
|
+
jaketts.egg-info/top_level.txt
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
jaketts
|
jaketts-1.0.0/jaketts.py
ADDED
|
@@ -0,0 +1,421 @@
|
|
|
1
|
+
#!/usr/bin/env python3
|
|
2
|
+
import sys
|
|
3
|
+
import os
|
|
4
|
+
import argparse
|
|
5
|
+
import re
|
|
6
|
+
import numpy as np
|
|
7
|
+
import soundfile as sf
|
|
8
|
+
import sounddevice as sd
|
|
9
|
+
|
|
10
|
+
# Silence torch, tokenizer, and huggingface framework warnings completely
|
|
11
|
+
os.environ["TOKENIZERS_PARALLELISM"] = "false"
|
|
12
|
+
os.environ["HF_HUB_DISABLE_SYMLINKS_WARNING"] = "1"
|
|
13
|
+
import warnings
|
|
14
|
+
|
|
15
|
+
warnings.filterwarnings("ignore")
|
|
16
|
+
|
|
17
|
+
try:
|
|
18
|
+
import logging
|
|
19
|
+
|
|
20
|
+
logging.getLogger("huggingface_hub").setLevel(logging.ERROR)
|
|
21
|
+
logging.getLogger("huggingface_hub.utils._validators").setLevel(logging.ERROR)
|
|
22
|
+
logging.getLogger("huggingface_hub.hub_mixin").setLevel(logging.ERROR)
|
|
23
|
+
except:
|
|
24
|
+
pass
|
|
25
|
+
|
|
26
|
+
from kokoro import KPipeline
|
|
27
|
+
|
|
28
|
+
try:
|
|
29
|
+
from tqdm import tqdm
|
|
30
|
+
except ImportError:
|
|
31
|
+
|
|
32
|
+
def tqdm(iterable, *args, **kwargs):
|
|
33
|
+
return iterable
|
|
34
|
+
|
|
35
|
+
|
|
36
|
+
# --- BATCH PROCESSING PIPELINE ---
|
|
37
|
+
def process_text_in_batches(pipeline, text, voice, speed, pbar=None):
|
|
38
|
+
"""
|
|
39
|
+
Safely breaks text into smaller paragraph blocks to prevent
|
|
40
|
+
memory exhaustion on massive inputs.
|
|
41
|
+
"""
|
|
42
|
+
paragraphs = [p.strip() for p in re.split(r"\n+", text) if p.strip()]
|
|
43
|
+
if not paragraphs:
|
|
44
|
+
paragraphs = [text]
|
|
45
|
+
|
|
46
|
+
audio_chunks = []
|
|
47
|
+
|
|
48
|
+
for para in paragraphs:
|
|
49
|
+
try:
|
|
50
|
+
generator = pipeline(para, voice=voice, speed=speed)
|
|
51
|
+
for _, _, audio in generator:
|
|
52
|
+
if audio is not None:
|
|
53
|
+
audio_chunks.append(audio)
|
|
54
|
+
if pbar:
|
|
55
|
+
pbar.update(1)
|
|
56
|
+
except Exception as e:
|
|
57
|
+
print(f"\nβ οΈ Warning: Failed to synthesize segment: {e}")
|
|
58
|
+
continue
|
|
59
|
+
|
|
60
|
+
return audio_chunks
|
|
61
|
+
|
|
62
|
+
|
|
63
|
+
# --- NATIVE DESKTOP GUI APP ---
|
|
64
|
+
def launch_desktop_gui():
|
|
65
|
+
"""
|
|
66
|
+
Launches a modern, themed Tkinter desktop interface.
|
|
67
|
+
Only runs if the user executes the command with zero parameters.
|
|
68
|
+
"""
|
|
69
|
+
import tkinter as tk
|
|
70
|
+
from tkinter import ttk, filedialog, messagebox
|
|
71
|
+
import threading
|
|
72
|
+
|
|
73
|
+
root = tk.Tk()
|
|
74
|
+
root.title("π jaketts - Text to Speech")
|
|
75
|
+
root.geometry("600x500")
|
|
76
|
+
root.minsize(500, 400)
|
|
77
|
+
|
|
78
|
+
style = ttk.Style(root)
|
|
79
|
+
style.theme_use("clam" if sys.platform == "darwin" else "default")
|
|
80
|
+
|
|
81
|
+
main_frame = ttk.Frame(root, padding="15")
|
|
82
|
+
main_frame.pack(fill=tk.BOTH, expand=True)
|
|
83
|
+
|
|
84
|
+
ttk.Label(
|
|
85
|
+
main_frame, text="Enter Text or Load a File:", font=("Helvetica", 12, "bold")
|
|
86
|
+
).pack(anchor=tk.W, pady=(0, 5))
|
|
87
|
+
|
|
88
|
+
container = ttk.Frame(main_frame)
|
|
89
|
+
container.pack(fill=tk.BOTH, expand=True)
|
|
90
|
+
|
|
91
|
+
text_scroll = ttk.Scrollbar(container)
|
|
92
|
+
text_scroll.pack(side=tk.RIGHT, fill=tk.Y)
|
|
93
|
+
|
|
94
|
+
text_box = tk.Text(
|
|
95
|
+
container,
|
|
96
|
+
yscrollcommand=text_scroll.set,
|
|
97
|
+
wrap=tk.WORD,
|
|
98
|
+
font=("Helvetica", 11),
|
|
99
|
+
height=10,
|
|
100
|
+
)
|
|
101
|
+
text_box.pack(fill=tk.BOTH, expand=True, side=tk.LEFT)
|
|
102
|
+
text_scroll.config(command=text_box.yview)
|
|
103
|
+
|
|
104
|
+
control_frame = ttk.Frame(main_frame, padding="10")
|
|
105
|
+
control_frame.pack(fill=tk.X, pady=10)
|
|
106
|
+
|
|
107
|
+
ttk.Label(control_frame, text="Voice:").grid(row=0, column=0, sticky=tk.W, padx=5)
|
|
108
|
+
voice_var = tk.StringVar(value="bm_george")
|
|
109
|
+
voice_dropdown = ttk.Combobox(
|
|
110
|
+
control_frame, textvariable=voice_var, state="readonly", width=12
|
|
111
|
+
)
|
|
112
|
+
voice_dropdown["values"] = (
|
|
113
|
+
"[en-us] af_heart",
|
|
114
|
+
"[en-us] af_sarah",
|
|
115
|
+
"[en-us] af_bella",
|
|
116
|
+
"[en-us] af_nicole",
|
|
117
|
+
"[en-us] af_sky",
|
|
118
|
+
"[en-us] af_alloy",
|
|
119
|
+
"[en-us] af_aoede",
|
|
120
|
+
"[en-us] af_jessica",
|
|
121
|
+
"[en-us] af_river",
|
|
122
|
+
"[en-us] am_adam",
|
|
123
|
+
"[en-us] am_michael",
|
|
124
|
+
"[en-us] am_echo",
|
|
125
|
+
"[en-us] am_eric",
|
|
126
|
+
"[en-us] am_fenrir",
|
|
127
|
+
"[en-us] am_liam",
|
|
128
|
+
"[en-us] am_onizuka",
|
|
129
|
+
"[en-us] am_puck",
|
|
130
|
+
"[en-us] am_santa",
|
|
131
|
+
"[en-gb] bm_george",
|
|
132
|
+
"[en-gb] bm_lewis",
|
|
133
|
+
"[en-gb] bf_emma",
|
|
134
|
+
"[en-gb] bf_isabella",
|
|
135
|
+
"[en-gb] bm_fable",
|
|
136
|
+
"[en-gb] bm_daniel",
|
|
137
|
+
"[en-gb] bf_alice",
|
|
138
|
+
"[en-gb] bf_lily",
|
|
139
|
+
"[es] ef_dora",
|
|
140
|
+
"[es] em_alex",
|
|
141
|
+
"[fr] ff_sixtine",
|
|
142
|
+
"[fr] fm_julien",
|
|
143
|
+
"[hi] hf_ananya",
|
|
144
|
+
"[hi] hf_kavya",
|
|
145
|
+
"[hi] hm_anshul",
|
|
146
|
+
"[hi] hm_shiwani",
|
|
147
|
+
"[it] if_sara",
|
|
148
|
+
"[it] im_nicola",
|
|
149
|
+
"[ja] jf_alpha",
|
|
150
|
+
"[ja] jf_glowing",
|
|
151
|
+
"[ja] jf_neutral",
|
|
152
|
+
"[ja] jf_reader",
|
|
153
|
+
"[ja] jm_kanta",
|
|
154
|
+
"[pt] pf_doris",
|
|
155
|
+
"[pt] pm_ramon",
|
|
156
|
+
"[zh] zf_xiaobei",
|
|
157
|
+
"[zh] zf_xiaoni",
|
|
158
|
+
"[zh] zf_xiaoxiao",
|
|
159
|
+
"[zh] zf_xiaoyi",
|
|
160
|
+
"[zh] zm_yunjian",
|
|
161
|
+
"[zh] zm_yunxi",
|
|
162
|
+
"[zh] zm_yunxia",
|
|
163
|
+
"[zh] zm_yunyang",
|
|
164
|
+
)
|
|
165
|
+
voice_dropdown.grid(row=0, column=1, padx=5, sticky=tk.W)
|
|
166
|
+
|
|
167
|
+
ttk.Label(control_frame, text="Speed:").grid(row=0, column=2, sticky=tk.W, padx=15)
|
|
168
|
+
speed_var = tk.DoubleVar(value=1.0)
|
|
169
|
+
speed_scale = ttk.Scale(
|
|
170
|
+
control_frame,
|
|
171
|
+
from_=0.5,
|
|
172
|
+
to=2.0,
|
|
173
|
+
variable=speed_var,
|
|
174
|
+
orient=tk.HORIZONTAL,
|
|
175
|
+
length=120,
|
|
176
|
+
)
|
|
177
|
+
speed_scale.grid(row=0, column=3, padx=5, sticky=tk.W)
|
|
178
|
+
|
|
179
|
+
speed_label = ttk.Label(control_frame, text="1.0x")
|
|
180
|
+
speed_label.grid(row=0, column=4, padx=2)
|
|
181
|
+
|
|
182
|
+
def update_speed_label(*args):
|
|
183
|
+
speed_label.config(text=f"{speed_var.get():.2f}x")
|
|
184
|
+
|
|
185
|
+
speed_var.trace_add("write", update_speed_label)
|
|
186
|
+
|
|
187
|
+
progress_bar = ttk.Progressbar(main_frame, orient=tk.HORIZONTAL, mode="determinate")
|
|
188
|
+
progress_bar.pack(fill=tk.X, pady=5)
|
|
189
|
+
|
|
190
|
+
status_label = ttk.Label(main_frame, text="Ready", font=("Helvetica", 10, "italic"))
|
|
191
|
+
status_label.pack(anchor=tk.W, pady=2)
|
|
192
|
+
|
|
193
|
+
def open_text_file():
|
|
194
|
+
file_path = filedialog.askopenfilename(
|
|
195
|
+
filetypes=[("Text Files", "*.txt"), ("All Files", "*.*")]
|
|
196
|
+
)
|
|
197
|
+
if file_path:
|
|
198
|
+
try:
|
|
199
|
+
with open(file_path, "r", encoding="utf-8") as f:
|
|
200
|
+
content = f.read()
|
|
201
|
+
text_box.delete("1.0", tk.END)
|
|
202
|
+
text_box.insert("1.0", content)
|
|
203
|
+
status_label.config(
|
|
204
|
+
text=f"π Loaded file: {os.path.basename(file_path)}"
|
|
205
|
+
)
|
|
206
|
+
except Exception as e:
|
|
207
|
+
messagebox.showerror("Error", f"Failed to read text file:\n{e}")
|
|
208
|
+
|
|
209
|
+
def run_synthesis(action_type):
|
|
210
|
+
input_text = text_box.get("1.0", tk.END).strip()
|
|
211
|
+
if not input_text:
|
|
212
|
+
messagebox.showwarning(
|
|
213
|
+
"Warning", "Please provide or load some text content first!"
|
|
214
|
+
)
|
|
215
|
+
return
|
|
216
|
+
|
|
217
|
+
dropdown_selection = voice_var.get()
|
|
218
|
+
# Extracts the raw voice name from the end (e.g., "af_heart")
|
|
219
|
+
voice = dropdown_selection.split()[-1]
|
|
220
|
+
|
|
221
|
+
# Pulls the prefix inside the brackets to set the language code (a, b, e, f, h, i, j, p, z)
|
|
222
|
+
prefix = dropdown_selection.split("]")[0].replace("[", "").strip()
|
|
223
|
+
if "-" in prefix:
|
|
224
|
+
# Safely converts 'en-us' to 'a' and 'en-gb' to 'b'
|
|
225
|
+
lang_code = "a" if prefix.endswith("us") else "b"
|
|
226
|
+
else:
|
|
227
|
+
# Takes the first letter for other locales ('es' -> 'e', 'ja' -> 'j')
|
|
228
|
+
lang_code = prefix[0]
|
|
229
|
+
|
|
230
|
+
speed = speed_var.get()
|
|
231
|
+
|
|
232
|
+
def worker():
|
|
233
|
+
try:
|
|
234
|
+
status_label.config(text="π€ Loading AI Neural Engine Checkpoints...")
|
|
235
|
+
progress_bar.config(mode="indet")
|
|
236
|
+
progress_bar.start(10)
|
|
237
|
+
|
|
238
|
+
pipeline = KPipeline(lang_code=lang_code, repo_id="hexgrad/Kokoro-82M")
|
|
239
|
+
|
|
240
|
+
paragraphs = [p for p in input_text.split("\n") if p.strip()]
|
|
241
|
+
total_p = len(paragraphs) if paragraphs else 1
|
|
242
|
+
|
|
243
|
+
progress_bar.stop()
|
|
244
|
+
progress_bar.config(mode="determinate", maximum=total_p, value=0)
|
|
245
|
+
status_label.config(text="π£οΈ Generating Speech Array Channels...")
|
|
246
|
+
|
|
247
|
+
if action_type == "play":
|
|
248
|
+
for para in paragraphs:
|
|
249
|
+
generator = pipeline(para, voice=voice, speed=speed)
|
|
250
|
+
for _, _, audio in generator:
|
|
251
|
+
if audio is not None:
|
|
252
|
+
sd.play(audio, samplerate=24000)
|
|
253
|
+
sd.wait()
|
|
254
|
+
progress_bar.step(1)
|
|
255
|
+
status_label.config(text="β¨ Finished live playback successfully!")
|
|
256
|
+
else:
|
|
257
|
+
audio_chunks = []
|
|
258
|
+
for para in paragraphs:
|
|
259
|
+
generator = pipeline(para, voice=voice, speed=speed)
|
|
260
|
+
for _, _, audio in generator:
|
|
261
|
+
if audio is not None:
|
|
262
|
+
audio_chunks.append(audio)
|
|
263
|
+
progress_bar.step(1)
|
|
264
|
+
|
|
265
|
+
if audio_chunks:
|
|
266
|
+
combined = np.concatenate(audio_chunks)
|
|
267
|
+
root.after(0, lambda: save_audio_file(combined))
|
|
268
|
+
else:
|
|
269
|
+
status_label.config(
|
|
270
|
+
text="β Generation produced no audio data."
|
|
271
|
+
)
|
|
272
|
+
|
|
273
|
+
except Exception as e:
|
|
274
|
+
root.after(
|
|
275
|
+
0, lambda: messagebox.showerror("Error", f"Synthesis broken:\n{e}")
|
|
276
|
+
)
|
|
277
|
+
status_label.config(text="β Process Failed.")
|
|
278
|
+
finally:
|
|
279
|
+
root.after(0, lambda: progress_bar.config(value=0))
|
|
280
|
+
|
|
281
|
+
threading.Thread(target=worker, daemon=True).start()
|
|
282
|
+
|
|
283
|
+
def save_audio_file(audio_data):
|
|
284
|
+
save_path = filedialog.asksaveasfilename(
|
|
285
|
+
defaultextension=".wav", filetypes=[("WAV Audio", "*.wav")]
|
|
286
|
+
)
|
|
287
|
+
if save_path:
|
|
288
|
+
try:
|
|
289
|
+
sf.write(save_path, audio_data, 24000)
|
|
290
|
+
status_label.config(
|
|
291
|
+
text=f"β¨ Saved track to: {os.path.basename(save_path)}"
|
|
292
|
+
)
|
|
293
|
+
messagebox.showinfo(
|
|
294
|
+
"Success", f"Audio track exported successfully to:\n{save_path}"
|
|
295
|
+
)
|
|
296
|
+
except Exception as e:
|
|
297
|
+
messagebox.showerror("Error", f"Failed to save file:\n{e}")
|
|
298
|
+
else:
|
|
299
|
+
status_label.config(text="β οΈ File export cancelled.")
|
|
300
|
+
|
|
301
|
+
btn_frame = ttk.Frame(main_frame)
|
|
302
|
+
btn_frame.pack(fill=tk.X, pady=10)
|
|
303
|
+
|
|
304
|
+
ttk.Button(btn_frame, text="π Open File", command=open_text_file).pack(
|
|
305
|
+
side=tk.LEFT, padx=5
|
|
306
|
+
)
|
|
307
|
+
ttk.Button(
|
|
308
|
+
btn_frame, text="π Play Speech", command=lambda: run_synthesis("play")
|
|
309
|
+
).pack(side=tk.RIGHT, padx=5)
|
|
310
|
+
ttk.Button(
|
|
311
|
+
btn_frame, text="πΎ Save WAV File", command=lambda: run_synthesis("save")
|
|
312
|
+
).pack(side=tk.RIGHT, padx=5)
|
|
313
|
+
|
|
314
|
+
root.mainloop()
|
|
315
|
+
|
|
316
|
+
|
|
317
|
+
# --- PRIMARY COMMAND LINE INTERFACE ROUTING ENGINE ---
|
|
318
|
+
def main():
|
|
319
|
+
if len(sys.argv) == 1:
|
|
320
|
+
launch_desktop_gui()
|
|
321
|
+
sys.exit(0)
|
|
322
|
+
|
|
323
|
+
if len(sys.argv) == 3 and sys.argv[1] in ["-o", "--output"]:
|
|
324
|
+
potential_file = sys.argv[2]
|
|
325
|
+
if os.path.isfile(potential_file) and not potential_file.lower().endswith(
|
|
326
|
+
".wav"
|
|
327
|
+
):
|
|
328
|
+
sys.argv = [sys.argv[0], "-o", "output.wav", potential_file]
|
|
329
|
+
|
|
330
|
+
parser = argparse.ArgumentParser(
|
|
331
|
+
description="π Jake's Text-to-Speech CLI utility powered by Kokoro."
|
|
332
|
+
)
|
|
333
|
+
parser.add_argument(
|
|
334
|
+
"-o",
|
|
335
|
+
"--output",
|
|
336
|
+
nargs="?",
|
|
337
|
+
const="output.wav",
|
|
338
|
+
default=None,
|
|
339
|
+
help="Output filename.",
|
|
340
|
+
)
|
|
341
|
+
parser.add_argument(
|
|
342
|
+
"-v", "--voice", default="bm_george", help="Voice profile ID selection."
|
|
343
|
+
)
|
|
344
|
+
parser.add_argument(
|
|
345
|
+
"-s", "--speed", type=float, default=1.0, help="Vocal speed modifier parameter."
|
|
346
|
+
)
|
|
347
|
+
parser.add_argument(
|
|
348
|
+
"text_input", help="The text string or path to a .txt file input source."
|
|
349
|
+
)
|
|
350
|
+
|
|
351
|
+
args = parser.parse_args()
|
|
352
|
+
|
|
353
|
+
final_text = args.text_input
|
|
354
|
+
if os.path.isfile(args.text_input):
|
|
355
|
+
print(f"π Reading text from file: {os.path.abspath(args.text_input)}")
|
|
356
|
+
try:
|
|
357
|
+
with open(args.text_input, "r", encoding="utf-8") as f:
|
|
358
|
+
final_text = f.read().strip()
|
|
359
|
+
except Exception as e:
|
|
360
|
+
print(f"β Error reading file: {e}")
|
|
361
|
+
sys.exit(1)
|
|
362
|
+
|
|
363
|
+
if not final_text:
|
|
364
|
+
print("β Error: No text content found to synthesize.")
|
|
365
|
+
sys.exit(1)
|
|
366
|
+
|
|
367
|
+
lang_code = args.voice.lower()
|
|
368
|
+
if lang_code not in ["a", "b", "e", "f", "h", "i", "j", "p", "z"]:
|
|
369
|
+
lang_code = "b" if lang_code.startswith("b") else "a"
|
|
370
|
+
|
|
371
|
+
print(f"π€ Initializing Kokoro Engine (Locale: {lang_code})...")
|
|
372
|
+
try:
|
|
373
|
+
pipeline = KPipeline(lang_code=lang_code, repo_id="hexgrad/Kokoro-82M")
|
|
374
|
+
except Exception as e:
|
|
375
|
+
print(f"β Failed to load pipeline: {e}")
|
|
376
|
+
sys.exit(1)
|
|
377
|
+
|
|
378
|
+
print(f"π£οΈ Synthesizing text via voice '{args.voice}' (Speed: {args.speed}x)...")
|
|
379
|
+
|
|
380
|
+
if args.output is not None:
|
|
381
|
+
print(f"πΎ Gathering audio tracks for batch file output...")
|
|
382
|
+
paragraphs = [p for p in final_text.split("\n") if p.strip()]
|
|
383
|
+
total_chunks = len(paragraphs) if paragraphs else 1
|
|
384
|
+
|
|
385
|
+
pbar = tqdm(
|
|
386
|
+
total=total_chunks, desc="Processing Paragraph Blocks", unit="chunk"
|
|
387
|
+
)
|
|
388
|
+
audio_chunks = process_text_in_batches(
|
|
389
|
+
pipeline, final_text, args.voice, args.speed, pbar=pbar
|
|
390
|
+
)
|
|
391
|
+
|
|
392
|
+
pbar.n = pbar.total
|
|
393
|
+
pbar.refresh()
|
|
394
|
+
pbar.close()
|
|
395
|
+
|
|
396
|
+
if not audio_chunks:
|
|
397
|
+
print("β No audio data generated.")
|
|
398
|
+
sys.exit(1)
|
|
399
|
+
|
|
400
|
+
combined_audio = np.concatenate(audio_chunks)
|
|
401
|
+
sf.write(args.output, combined_audio, 24000)
|
|
402
|
+
print(f"β¨ Success! Audio file written to: {os.path.abspath(args.output)}")
|
|
403
|
+
else:
|
|
404
|
+
print("π Playing audio directly through your speakers...")
|
|
405
|
+
paragraphs = [p for p in final_text.split("\n") if p.strip()]
|
|
406
|
+
if not paragraphs:
|
|
407
|
+
paragraphs = [final_text]
|
|
408
|
+
|
|
409
|
+
for para in paragraphs:
|
|
410
|
+
try:
|
|
411
|
+
generator = pipeline(para, voice=args.voice, speed=args.speed)
|
|
412
|
+
for _, _, audio in generator:
|
|
413
|
+
if audio is not None:
|
|
414
|
+
sd.play(audio, samplerate=24000)
|
|
415
|
+
sd.wait()
|
|
416
|
+
except Exception as e:
|
|
417
|
+
print(f"\nβ οΈ Error speaking segment: {e}")
|
|
418
|
+
|
|
419
|
+
|
|
420
|
+
if __name__ == "__main__":
|
|
421
|
+
main()
|
jaketts-1.0.0/setup.cfg
ADDED
jaketts-1.0.0/setup.py
ADDED
|
@@ -0,0 +1,24 @@
|
|
|
1
|
+
from setuptools import setup
|
|
2
|
+
|
|
3
|
+
setup(
|
|
4
|
+
name="jaketts",
|
|
5
|
+
version="1.0.0",
|
|
6
|
+
description="Jake's Local CLI Text-to-Speech tool powered by Kokoro-82M",
|
|
7
|
+
author="Jake",
|
|
8
|
+
py_modules=["jaketts"],
|
|
9
|
+
install_requires=[
|
|
10
|
+
"kokoro>=0.7.0",
|
|
11
|
+
"sounddevice>=0.4.0",
|
|
12
|
+
"soundfile>=0.4.0",
|
|
13
|
+
"numpy>=1.20.0,<2.0.0",
|
|
14
|
+
"torch>=2.0.0",
|
|
15
|
+
"tqdm>=4.65.0",
|
|
16
|
+
],
|
|
17
|
+
entry_points={
|
|
18
|
+
"console_scripts": [
|
|
19
|
+
"jaketts=jaketts:main",
|
|
20
|
+
"jtts=jaketts:main",
|
|
21
|
+
],
|
|
22
|
+
},
|
|
23
|
+
python_requires=">=3.10,<3.13",
|
|
24
|
+
)
|