wfloat 2.0.0__py3-none-win_amd64.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,251 @@
1
+ Metadata-Version: 2.4
2
+ Name: wfloat
3
+ Version: 2.0.0
4
+ Summary: Wfloat Python SDK for running models locally
5
+ Home-page: https://github.com/wfloat/wfloat-python
6
+ Author: wfloat
7
+ License: MIT
8
+ Classifier: Programming Language :: Python :: 3
9
+ Classifier: Operating System :: Microsoft :: Windows
10
+ Classifier: Operating System :: POSIX :: Linux
11
+ Classifier: Operating System :: MacOS :: MacOS X
12
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
13
+ Requires-Python: >=3.9
14
+ Description-Content-Type: text/markdown
15
+ Dynamic: author
16
+ Dynamic: classifier
17
+ Dynamic: description
18
+ Dynamic: description-content-type
19
+ Dynamic: home-page
20
+ Dynamic: license
21
+ Dynamic: requires-python
22
+ Dynamic: summary
23
+
24
+ # wfloat
25
+
26
+ `wfloat` is the Python package for Wfloat's on-device model families, including
27
+ TTS, STT, VAD, and LLM.
28
+
29
+ It runs inference locally in Python instead of calling a hosted inference API.
30
+ The TTS model supports 20 voices with emotion and intensity control.
31
+
32
+ If you're building for the browser, use
33
+ [`@wfloat/wfloat-web`](https://github.com/wfloat/wfloat-web). If you're
34
+ building for React Native, use
35
+ [`@wfloat/react-native-wfloat`](https://github.com/wfloat/react-native-wfloat).
36
+
37
+ Browser demo to hear how it sounds: https://wfloat.com/demo
38
+
39
+ ## Install
40
+
41
+ ```bash
42
+ pip install wfloat
43
+ ```
44
+
45
+ ## Usage
46
+
47
+ ```python
48
+ import wfloat
49
+
50
+ tts = wfloat.load_tts_model("wfloat/wfloat-tts")
51
+
52
+ result = tts.synthesize(
53
+ text="No, no, that's not possible.",
54
+ voice="mad_scientist_woman",
55
+ emotion="surprise",
56
+ intensity=0.7,
57
+ )
58
+
59
+ print(result.model_id)
60
+ print(result.timeline.chunks[0].text)
61
+ ```
62
+
63
+ For multi-speaker dialogue:
64
+
65
+ ```python
66
+ import wfloat
67
+
68
+ tts = wfloat.load_tts_model("wfloat/wfloat-tts")
69
+
70
+ result = tts.synthesize_dialogue(
71
+ segments=[
72
+ {
73
+ "voice": "wise_elder_man",
74
+ "text": "Rain taps against the tavern shutters as you step inside.",
75
+ "emotion": "neutral",
76
+ "intensity": 0.5,
77
+ },
78
+ {
79
+ "voice": "strong_hero_man",
80
+ "text": "You're late. Two bandits stole the king's map over three hours ago.",
81
+ "emotion": "fear",
82
+ "intensity": 0.6,
83
+ },
84
+ {
85
+ "voice": "strong_hero_man",
86
+ "text": "They fled north, up into the woods.",
87
+ "emotion": "neutral",
88
+ "intensity": 0.5,
89
+ },
90
+ ],
91
+ silence_between_segments_sec=0.35,
92
+ )
93
+
94
+ result.audio.save("dialogue.wav")
95
+ ```
96
+
97
+ The older `load(...)`, `generate(...)`, and `generate_dialogue(...)` names are
98
+ still available as compatibility aliases.
99
+
100
+ ## STT Usage
101
+
102
+ The shared Python entrypoint is `load_stt_model(...)`, with
103
+ `load_whisper_tiny_en(...)` as a convenience wrapper for offline STT:
104
+
105
+ ```python
106
+ import wfloat
107
+
108
+ stt = wfloat.load_stt_model(
109
+ "openai/whisper-tiny-en",
110
+ )
111
+
112
+ result = stt.transcribe(audio="/path/to/audio.wav")
113
+ print(result.text)
114
+ ```
115
+
116
+ The loader accepts canonical built-in model IDs and resolves Wfloat-hosted
117
+ registry assets internally.
118
+
119
+ Streaming-capable STT families also expose a separate session path instead of
120
+ overloading `transcribe(...)`:
121
+
122
+ ```python
123
+ stt = wfloat.load_stt_model("k2-fsa/streaming-zipformer-en")
124
+ session = stt.create_session()
125
+
126
+ session.push(audio_chunk, sample_rate=16000)
127
+ partial = session.get_result()
128
+ final_result = session.finish()
129
+ session.close()
130
+ ```
131
+
132
+ ## VAD Usage
133
+
134
+ Python also exposes the same one-shot VAD model shape as the web and React
135
+ Native packages. It is intentionally file/buffer based for now; there is no
136
+ Python live microphone/session helper. The Python VAD path uses `wfloat-core`,
137
+ matching the TTS, STT, and LLM backend boundary.
138
+
139
+ ```python
140
+ vad = wfloat.load_vad_model(
141
+ "snakers4/silero-vad",
142
+ threshold=0.5,
143
+ min_silence_duration_sec=0.5,
144
+ min_speech_duration_sec=0.25,
145
+ max_speech_duration_sec=20.0,
146
+ )
147
+
148
+ result = vad.detect(audio="/path/to/mono-16khz.wav")
149
+
150
+ for segment in result.segments:
151
+ print(segment.start_sec, segment.duration_sec)
152
+ ```
153
+
154
+ VAD currently expects mono 16 kHz audio.
155
+
156
+ ## LLM Usage
157
+
158
+ The Python LLM path loads local GGUF artifacts through `wfloat-core`:
159
+
160
+ ```python
161
+ import wfloat
162
+
163
+ llm = wfloat.load_llm_model("HuggingFaceTB/SmolLM2-360M-Instruct")
164
+ result = llm.generate("Write one calm sentence about local inference.", seed=0)
165
+ print(result.text)
166
+ llm.close()
167
+ ```
168
+
169
+ ## CLI Usage
170
+
171
+ You can also generate a WAV from the command line:
172
+
173
+ ```bash
174
+ wfloat generate \
175
+ --text "Hello world!" \
176
+ --out out.wav \
177
+ --voice-id mad_scientist_woman \
178
+ --emotion surprise \
179
+ --intensity 0.7 \
180
+ --silence-padding-sec 0
181
+ ```
182
+
183
+ For the full CLI help:
184
+
185
+ ```bash
186
+ wfloat generate --help
187
+ ```
188
+
189
+ The first load downloads the model assets. After that, the package uses the
190
+ cached local copy.
191
+
192
+ ## Native Backend
193
+
194
+ Python TTS, STT, VAD, and LLM use the `wfloat-core` native runtime. Release
195
+ wheels bundle the platform-specific shared library inside the `wfloat` package.
196
+
197
+ Inside this monorepo, local development can point at an explicit build artifact
198
+ with:
199
+
200
+ ```bash
201
+ export WFLOAT_CORE_LIBRARY=/abs/path/to/libwfloat-core.so
202
+ ```
203
+
204
+ ## Speaker IDs
205
+
206
+ Use `voice_id` string names or numeric `sid` values:
207
+
208
+ | Speaker | SID |
209
+ | --- | ---: |
210
+ | `skilled_hero_man` | 0 |
211
+ | `skilled_hero_woman` | 1 |
212
+ | `fun_hero_man` | 2 |
213
+ | `fun_hero_woman` | 3 |
214
+ | `strong_hero_man` | 4 |
215
+ | `strong_hero_woman` | 5 |
216
+ | `mad_scientist_man` | 6 |
217
+ | `mad_scientist_woman` | 7 |
218
+ | `clever_villain_man` | 8 |
219
+ | `clever_villain_woman` | 9 |
220
+ | `narrator_man` | 10 |
221
+ | `narrator_woman` | 11 |
222
+ | `wise_elder_man` | 12 |
223
+ | `wise_elder_woman` | 13 |
224
+ | `outgoing_anime_man` | 14 |
225
+ | `outgoing_anime_woman` | 15 |
226
+ | `scary_villain_man` | 16 |
227
+ | `scary_villain_woman` | 17 |
228
+ | `news_reporter_man` | 18 |
229
+ | `news_reporter_woman` | 19 |
230
+
231
+ ## Emotions
232
+
233
+ Supported emotion labels:
234
+
235
+ - `neutral`
236
+ - `joy`
237
+ - `sadness`
238
+ - `anger`
239
+ - `fear`
240
+ - `surprise`
241
+ - `dismissive`
242
+ - `confusion`
243
+
244
+ `intensity` must be between `0.0` and `1.0`.
245
+
246
+ ## More
247
+
248
+ - Docs: https://docs.wfloat.com
249
+ - Model card, voices, emotions, and samples: https://huggingface.co/Wfloat/wfloat-tts
250
+ - Web package: https://github.com/wfloat/wfloat-web
251
+ - React Native package: https://github.com/wfloat/react-native-wfloat
@@ -0,0 +1,28 @@
1
+ wfloat/__init__.py,sha256=pcvCRVgzrCtofPqKoV---3gDvTEcPnm2vLzW5u0-swQ,1447
2
+ wfloat/__main__.py,sha256=NUh9D_rWfVdbNziDDz-ls3KZY5OJcI__l4OE-Chri1Y,106
3
+ wfloat/_assets.py,sha256=g3PxLasu10X5dTZMwEF91CsnAHQJbpga1cQZmmk2gtI,15961
4
+ wfloat/_cache.py,sha256=kU47uO93Yp3OBQDwzaFPEJoFw_ZJroL1Ep_0mDZ0W4Y,5985
5
+ wfloat/_cli.py,sha256=iCr-CbLD58j1e93vx0r5eHECU17fwnprW7AnKsxSL0w,2372
6
+ wfloat/_constants.py,sha256=OXh4Lrhp0MIR--ZR6bmCV1egRci4cDMY0JgSon-1dGo,3752
7
+ wfloat/_core.py,sha256=Nn4Eckc0VNbGtOp19TcsVX_7t_-FIoN_cqq9VGQhiYE,62148
8
+ wfloat/_download.py,sha256=6l-OTBF0ns0FmKrB9CErxs4wcUWIbYbBF6K4LGRc4kw,4166
9
+ wfloat/_generated_model_urls.py,sha256=8WGhUjdcfkfXoUW-KA3vyqlm-hAFfENAUfDeE0HN1Ps,4543
10
+ wfloat/_llm.py,sha256=HO0iRAXawg7U9Kr5acyusFIEgeJ0jg3Bnuik9WdBPTE,2433
11
+ wfloat/_llm_assets.py,sha256=8HyMuyFX1qrukxgOQNEXAnaaMX3yfB5Xl7t03mkKZJw,4326
12
+ wfloat/_llm_load.py,sha256=5hFFgIdm1GdQf6uOU-IfCdP668ERS5etBx_DMIT5JdM,1941
13
+ wfloat/_model.py,sha256=CxEyquiKQ1cM8PERR9IFPzWfPv6tKtsostVJ0BeiFe8,13370
14
+ wfloat/_native.py,sha256=g4jKupO18KwMolryHv4dBGpQ6c62pgZEhCCHiq5yJhU,369
15
+ wfloat/_results.py,sha256=HNAToKeLAiAyJaQFWL_lln2zMAcX3khWa22RaIVvaEc,3578
16
+ wfloat/_stt.py,sha256=6OUqXOQKAIIPF1D2OCnxWx2raY6m5k-OSilWf4yOphA,4123
17
+ wfloat/_stt_assets.py,sha256=nV2momO8m3aaq7amssUdzfmrS56m0C5PjU0Kxy3yOiI,4519
18
+ wfloat/_stt_load.py,sha256=FW7SaEkM7WBoJeRHj7MSVREtzMVQb-wXCLS9TfGChew,2296
19
+ wfloat/_vad.py,sha256=6sYhlBnaMvkioub_ih4fIb0POV_xYuIcQj98OGUDg7k,4342
20
+ wfloat/_vad_assets.py,sha256=nDJfM5nLm-zHgzToFjSpUKrQUlX0-q-0LP8I4T7VXP8,3792
21
+ wfloat/_vad_load.py,sha256=GOled83VUGk_dB13U6eDgeg0NuJLxX_PfKGzbnpBqvA,3335
22
+ wfloat/_version.py,sha256=HLJc82fZ6h65qdEwfCWx0U1_51VFQHVHv3kddt9t2Yk,23
23
+ wfloat/native/wfloat-core.dll,sha256=2L-ceUI5d4uMOM4o04BVBTS8kPK_VpS9Il8XK82WDZQ,21232640
24
+ wfloat-2.0.0.dist-info/METADATA,sha256=r8EQgvJ3m0s4R3W83c5wSnI1jt0hF3-NRBhSI5DWjBY,6334
25
+ wfloat-2.0.0.dist-info/WHEEL,sha256=Zx98gwb_dQKckJK3HEOVKnxxD52EfU_PXDG_usMJ2ng,98
26
+ wfloat-2.0.0.dist-info/entry_points.txt,sha256=lWQgaOC_jw8m10sdYieAD0q9D_kewtVt2mrJyn3nyt4,44
27
+ wfloat-2.0.0.dist-info/top_level.txt,sha256=uWWyRL0PFWuaufMp4oEELfXtvmKTJkOtbApfN4XOvtk,7
28
+ wfloat-2.0.0.dist-info/RECORD,,
@@ -0,0 +1,5 @@
1
+ Wheel-Version: 1.0
2
+ Generator: setuptools (84.0.0)
3
+ Root-Is-Purelib: false
4
+ Tag: py3-none-win_amd64
5
+
@@ -0,0 +1,2 @@
1
+ [console_scripts]
2
+ wfloat = wfloat._cli:main
@@ -0,0 +1 @@
1
+ wfloat