turkish-stt 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,79 @@
1
+ mihu Community License
2
+ Version 1.0
3
+
4
+ These terms govern all use of the turkish-stt materials released by mihu
5
+ ("mihu", "Licensor"). By downloading, accessing, or otherwise using the Materials, you
6
+ ("Licensee") accept and agree to be bound by this License.
7
+
8
+ 1. DEFINITIONS
9
+ 1.1 "Materials" — the turkish-stt model weights, model card, and any
10
+ accompanying code, documentation, or related artifacts released under this License.
11
+ 1.2 "Outputs" — anything produced by running the Materials.
12
+ 1.3 "Derivative" — any work created from, adapted from, or built upon the Materials.
13
+ 1.4 "Annual Revenue" — the combined worldwide gross revenue of Licensee together with
14
+ its affiliates over the preceding twelve months.
15
+
16
+ 2. COMMUNITY GRANT
17
+ Provided Licensee's Annual Revenue stays below USD 500,000, and subject to every term
18
+ below, Licensor grants Licensee a worldwide, non-exclusive, non-transferable,
19
+ royalty-free, no-cost right to use, reproduce, and make Derivatives of the Materials.
20
+ This covers, without limitation:
21
+ (a) research and academic work;
22
+ (b) personal and other non-commercial use;
23
+ (c) commercial use within the revenue limit above, once any commercial-use
24
+ registration that Licensor requires has been completed.
25
+
26
+ 3. ABOVE THE THRESHOLD
27
+ Once Licensee's Annual Revenue reaches USD 500,000 or more, no use is permitted under
28
+ the Community Grant. Licensee must first secure a separate commercial agreement from
29
+ mihu. Write to support@mihu.ai to arrange terms.
30
+
31
+ 4. ATTRIBUTION
32
+ Licensee shall (a) keep this License intact, together with any copyright or attribution
33
+ notices, in a "NOTICE" file shipped alongside the Materials or any Derivative, and
34
+ (b) show "Powered by mihu" clearly in any product, service, interface, or documentation
35
+ that builds on the Materials or their Outputs.
36
+
37
+ 5. RESTRICTIONS
38
+ Licensee shall not:
39
+ (a) use the Materials or Outputs to train, build, or refine any competing speech,
40
+ foundational, or general-purpose AI model;
41
+ (b) strip, hide, or modify any attribution or proprietary notice;
42
+ (c) use the mihu name, logo, or marks beyond the attribution required in Section 4
43
+ (no trademark rights are granted here);
44
+ (d) use the Materials unlawfully or contrary to any Acceptable Use Policy that
45
+ Licensor may publish;
46
+ (e) sublicense, sell, or redistribute the Materials except where this License
47
+ expressly allows it.
48
+
49
+ 6. OWNERSHIP
50
+ Between the parties, Licensor keeps all right, title, and interest in the Materials.
51
+ Licensee holds the Derivatives it lawfully makes and the Outputs it generates, always
52
+ subject to Licensor's rights in the underlying Materials. Nothing beyond the rights
53
+ stated here is granted, by implication or otherwise.
54
+
55
+ 7. NO WARRANTY
56
+ THE MATERIALS ARE SUPPLIED "AS IS", WITH NO WARRANTY OF ANY KIND, WHETHER EXPRESS OR
57
+ IMPLIED, INCLUDING ANY IMPLIED WARRANTY OF MERCHANTABILITY, FITNESS FOR A PARTICULAR
58
+ PURPOSE, OR NON-INFRINGEMENT.
59
+
60
+ 8. LIMITATION OF LIABILITY
61
+ TO THE FULLEST EXTENT THE LAW ALLOWS, LICENSOR IS NOT LIABLE FOR ANY INDIRECT,
62
+ INCIDENTAL, SPECIAL, CONSEQUENTIAL, OR EXEMPLARY DAMAGES CONNECTED TO THE MATERIALS OR
63
+ THIS LICENSE.
64
+
65
+ 9. TERMINATION
66
+ This License ends automatically if (a) Licensee's Annual Revenue reaches or passes
67
+ USD 500,000 without a commercial agreement in place, (b) Licensee breaches any term, or
68
+ (c) Licensee brings litigation or a patent claim against Licensor over the Materials.
69
+ On termination, Licensee must stop all use and delete every copy of the Materials. To
70
+ keep using them, arrange a commercial agreement via support@mihu.ai.
71
+
72
+ 10. GENERAL
73
+ This License is the complete agreement covering the Materials. It is governed by the
74
+ laws of Licensor's principal place of business, disregarding conflict-of-law rules.
75
+
76
+ © mihu. All rights reserved.
77
+
78
+ turkish-stt and the mihu Community License are provided by Hunters AI, operating
79
+ under the mihu brand. "mihu" is a trademark of Hunters AI.
@@ -0,0 +1,198 @@
1
+ Metadata-Version: 2.4
2
+ Name: turkish-stt
3
+ Version: 0.1.0
4
+ Summary: Real-time, low-latency Turkish speech-to-text (STT) that runs on the CPU.
5
+ Author: mihu · Speech Systems
6
+ Maintainer-email: Deniz Gökdünek <deniz@mihu.ai>
7
+ License: mihu Community License
8
+ Version 1.0
9
+
10
+ These terms govern all use of the turkish-stt materials released by mihu
11
+ ("mihu", "Licensor"). By downloading, accessing, or otherwise using the Materials, you
12
+ ("Licensee") accept and agree to be bound by this License.
13
+
14
+ 1. DEFINITIONS
15
+ 1.1 "Materials" — the turkish-stt model weights, model card, and any
16
+ accompanying code, documentation, or related artifacts released under this License.
17
+ 1.2 "Outputs" — anything produced by running the Materials.
18
+ 1.3 "Derivative" — any work created from, adapted from, or built upon the Materials.
19
+ 1.4 "Annual Revenue" — the combined worldwide gross revenue of Licensee together with
20
+ its affiliates over the preceding twelve months.
21
+
22
+ 2. COMMUNITY GRANT
23
+ Provided Licensee's Annual Revenue stays below USD 500,000, and subject to every term
24
+ below, Licensor grants Licensee a worldwide, non-exclusive, non-transferable,
25
+ royalty-free, no-cost right to use, reproduce, and make Derivatives of the Materials.
26
+ This covers, without limitation:
27
+ (a) research and academic work;
28
+ (b) personal and other non-commercial use;
29
+ (c) commercial use within the revenue limit above, once any commercial-use
30
+ registration that Licensor requires has been completed.
31
+
32
+ 3. ABOVE THE THRESHOLD
33
+ Once Licensee's Annual Revenue reaches USD 500,000 or more, no use is permitted under
34
+ the Community Grant. Licensee must first secure a separate commercial agreement from
35
+ mihu. Write to support@mihu.ai to arrange terms.
36
+
37
+ 4. ATTRIBUTION
38
+ Licensee shall (a) keep this License intact, together with any copyright or attribution
39
+ notices, in a "NOTICE" file shipped alongside the Materials or any Derivative, and
40
+ (b) show "Powered by mihu" clearly in any product, service, interface, or documentation
41
+ that builds on the Materials or their Outputs.
42
+
43
+ 5. RESTRICTIONS
44
+ Licensee shall not:
45
+ (a) use the Materials or Outputs to train, build, or refine any competing speech,
46
+ foundational, or general-purpose AI model;
47
+ (b) strip, hide, or modify any attribution or proprietary notice;
48
+ (c) use the mihu name, logo, or marks beyond the attribution required in Section 4
49
+ (no trademark rights are granted here);
50
+ (d) use the Materials unlawfully or contrary to any Acceptable Use Policy that
51
+ Licensor may publish;
52
+ (e) sublicense, sell, or redistribute the Materials except where this License
53
+ expressly allows it.
54
+
55
+ 6. OWNERSHIP
56
+ Between the parties, Licensor keeps all right, title, and interest in the Materials.
57
+ Licensee holds the Derivatives it lawfully makes and the Outputs it generates, always
58
+ subject to Licensor's rights in the underlying Materials. Nothing beyond the rights
59
+ stated here is granted, by implication or otherwise.
60
+
61
+ 7. NO WARRANTY
62
+ THE MATERIALS ARE SUPPLIED "AS IS", WITH NO WARRANTY OF ANY KIND, WHETHER EXPRESS OR
63
+ IMPLIED, INCLUDING ANY IMPLIED WARRANTY OF MERCHANTABILITY, FITNESS FOR A PARTICULAR
64
+ PURPOSE, OR NON-INFRINGEMENT.
65
+
66
+ 8. LIMITATION OF LIABILITY
67
+ TO THE FULLEST EXTENT THE LAW ALLOWS, LICENSOR IS NOT LIABLE FOR ANY INDIRECT,
68
+ INCIDENTAL, SPECIAL, CONSEQUENTIAL, OR EXEMPLARY DAMAGES CONNECTED TO THE MATERIALS OR
69
+ THIS LICENSE.
70
+
71
+ 9. TERMINATION
72
+ This License ends automatically if (a) Licensee's Annual Revenue reaches or passes
73
+ USD 500,000 without a commercial agreement in place, (b) Licensee breaches any term, or
74
+ (c) Licensee brings litigation or a patent claim against Licensor over the Materials.
75
+ On termination, Licensee must stop all use and delete every copy of the Materials. To
76
+ keep using them, arrange a commercial agreement via support@mihu.ai.
77
+
78
+ 10. GENERAL
79
+ This License is the complete agreement covering the Materials. It is governed by the
80
+ laws of Licensor's principal place of business, disregarding conflict-of-law rules.
81
+
82
+ © mihu. All rights reserved.
83
+
84
+ turkish-stt and the mihu Community License are provided by Hunters AI, operating
85
+ under the mihu brand. "mihu" is a trademark of Hunters AI.
86
+
87
+ Project-URL: Model weights, https://huggingface.co/mihuai/turkish-stt
88
+ Keywords: speech-recognition,asr,streaming,turkish,on-device,cpu,voice-agent
89
+ Classifier: Programming Language :: Python :: 3
90
+ Classifier: License :: Other/Proprietary License
91
+ Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
92
+ Classifier: Operating System :: OS Independent
93
+ Requires-Python: >=3.9
94
+ Description-Content-Type: text/markdown
95
+ License-File: LICENSE
96
+ Requires-Dist: sherpa-onnx
97
+ Requires-Dist: numpy
98
+ Requires-Dist: soundfile
99
+ Requires-Dist: soxr
100
+ Requires-Dist: huggingface_hub
101
+ Provides-Extra: eval
102
+ Requires-Dist: jiwer; extra == "eval"
103
+ Requires-Dist: num2words; extra == "eval"
104
+ Dynamic: license-file
105
+
106
+ # turkish-stt
107
+
108
+ **Real-time, low-latency speech-to-text (STT) for Turkish — on the CPU.**
109
+
110
+ Turkish-only · streaming · CPU-native · runs on the edge. A compact **66 M-parameter** causal streaming model,
111
+ built for voice agents, call centers, and on-device transcription where a GPU is not an option — no API key, no
112
+ cloud, no GPU.
113
+
114
+ - 🇹🇷 **Turkish-first** — every parameter serves one language, not ~100
115
+ - ⚡ **Streaming** — partial results as you speak (~320 ms), no 30-second windows
116
+ - 🖥️ **CPU-native** — **RTF 0.128** on a single CPU thread (~8× faster than real time)
117
+ - 🔌 **Tiny** — ~68 MB; runs on a Raspberry-Pi-class device
118
+ - 🏷️ **Entity/brand aware** — tuned for Turkish brands, models, and numbers
119
+
120
+ ## Installation
121
+
122
+ ```bash
123
+ pip install turkish-stt
124
+ ```
125
+
126
+ ## Usage
127
+
128
+ ```python
129
+ import mihu_stt
130
+
131
+ # one-shot: transcribe a file (16 kHz mono, or anything soundfile reads)
132
+ text = mihu_stt.transcribe("audio.wav")
133
+ print(text)
134
+ ```
135
+
136
+ Streaming — feed audio as it arrives and read an updating transcript, with turn-end detection built in:
137
+
138
+ ```python
139
+ import mihu_stt
140
+
141
+ stt = mihu_stt.StreamingSTT()
142
+ # pcm16 = 16-bit PCM bytes from your mic / phone line
143
+ text, is_final = stt.accept_pcm16(pcm16, sample_rate=16000)
144
+ # `text` updates while speech continues; `is_final` turns True at end-of-utterance
145
+ ```
146
+
147
+ Model weights download automatically from Hugging Face
148
+ ([`mihuai/turkish-stt`](https://huggingface.co/mihuai/turkish-stt)) the first time and are cached.
149
+
150
+ ## Integrations
151
+
152
+ Drop-in adapters for the common voice-agent stacks (see [`examples/`](examples)):
153
+
154
+ **LiveKit Agents** — [`examples/livekit_stt.py`](examples/livekit_stt.py):
155
+ ```python
156
+ from livekit.agents import AgentSession
157
+ from livekit_stt import MihuSTT
158
+ session = AgentSession(stt=MihuSTT(), llm=..., tts=...)
159
+ ```
160
+
161
+ **Pipecat** — [`examples/pipecat_stt.py`](examples/pipecat_stt.py):
162
+ ```python
163
+ from pipecat_stt import MihuSTTService
164
+ pipeline = Pipeline([transport.input(), vad, MihuSTTService(), llm, tts, transport.output()])
165
+ ```
166
+
167
+ ## Benchmarks
168
+
169
+ Measured in-house under one matched Turkish normalization. Real-human held-out **FLEURS-TR** (743 clips) and a
170
+ second real-human set **MediaSpeech-TR** (200-clip subset), both audited absent from training. Latency on a
171
+ single CPU thread (AMD EPYC 7282).
172
+
173
+ | System | Class | Params | FLEURS WER | MediaSpeech WER | CPU RTF |
174
+ |---|---|---:|---:|---:|---:|
175
+ | **turkish-stt** | **streaming · CPU** | **66 M** | **13.90%** | **15.89%** | **0.128** |
176
+ | Vosk-TR | streaming · CPU | ~50 M | 34.09% | 31.02% | 0.138 |
177
+ | Whisper small | offline · GPU | 244 M | 15.12% | — | 1.098 |
178
+ | Whisper large-v3 | offline · GPU | 1.5 B | 7.08% | — | 6.095 |
179
+
180
+ Among **streaming, CPU-deployable** systems turkish-stt more than halves Vosk's error on both real-human sets,
181
+ matches the four-times-larger offline Whisper-small, and is the only option that is real-time on CPU — it also
182
+ degrades more gracefully than Vosk under additive noise. A detailed technical report — full methodology, ablations,
183
+ latency/concurrency study, and the leakage audits — is being prepared and will be shared.
184
+
185
+ ## Reproduce the numbers
186
+
187
+ The evaluation set, references, and scripts are in [`eval/`](eval) — including the text (n-gram) and acoustic
188
+ (chromaprint) contamination audits used to verify the held-out sets are absent from training.
189
+
190
+ ```bash
191
+ pip install "turkish-stt[eval]"
192
+ python eval/run_eval.py --data-dir eval/tts-synthetic
193
+ ```
194
+
195
+ ## License
196
+
197
+ **mihu Community License** — free for research, personal, and limited commercial use. See [`LICENSE`](LICENSE)
198
+ for the full terms. Weights are released under the same license on Hugging Face.
@@ -0,0 +1,93 @@
1
+ # turkish-stt
2
+
3
+ **Real-time, low-latency speech-to-text (STT) for Turkish — on the CPU.**
4
+
5
+ Turkish-only · streaming · CPU-native · runs on the edge. A compact **66 M-parameter** causal streaming model,
6
+ built for voice agents, call centers, and on-device transcription where a GPU is not an option — no API key, no
7
+ cloud, no GPU.
8
+
9
+ - 🇹🇷 **Turkish-first** — every parameter serves one language, not ~100
10
+ - ⚡ **Streaming** — partial results as you speak (~320 ms), no 30-second windows
11
+ - 🖥️ **CPU-native** — **RTF 0.128** on a single CPU thread (~8× faster than real time)
12
+ - 🔌 **Tiny** — ~68 MB; runs on a Raspberry-Pi-class device
13
+ - 🏷️ **Entity/brand aware** — tuned for Turkish brands, models, and numbers
14
+
15
+ ## Installation
16
+
17
+ ```bash
18
+ pip install turkish-stt
19
+ ```
20
+
21
+ ## Usage
22
+
23
+ ```python
24
+ import mihu_stt
25
+
26
+ # one-shot: transcribe a file (16 kHz mono, or anything soundfile reads)
27
+ text = mihu_stt.transcribe("audio.wav")
28
+ print(text)
29
+ ```
30
+
31
+ Streaming — feed audio as it arrives and read an updating transcript, with turn-end detection built in:
32
+
33
+ ```python
34
+ import mihu_stt
35
+
36
+ stt = mihu_stt.StreamingSTT()
37
+ # pcm16 = 16-bit PCM bytes from your mic / phone line
38
+ text, is_final = stt.accept_pcm16(pcm16, sample_rate=16000)
39
+ # `text` updates while speech continues; `is_final` turns True at end-of-utterance
40
+ ```
41
+
42
+ Model weights download automatically from Hugging Face
43
+ ([`mihuai/turkish-stt`](https://huggingface.co/mihuai/turkish-stt)) the first time and are cached.
44
+
45
+ ## Integrations
46
+
47
+ Drop-in adapters for the common voice-agent stacks (see [`examples/`](examples)):
48
+
49
+ **LiveKit Agents** — [`examples/livekit_stt.py`](examples/livekit_stt.py):
50
+ ```python
51
+ from livekit.agents import AgentSession
52
+ from livekit_stt import MihuSTT
53
+ session = AgentSession(stt=MihuSTT(), llm=..., tts=...)
54
+ ```
55
+
56
+ **Pipecat** — [`examples/pipecat_stt.py`](examples/pipecat_stt.py):
57
+ ```python
58
+ from pipecat_stt import MihuSTTService
59
+ pipeline = Pipeline([transport.input(), vad, MihuSTTService(), llm, tts, transport.output()])
60
+ ```
61
+
62
+ ## Benchmarks
63
+
64
+ Measured in-house under one matched Turkish normalization. Real-human held-out **FLEURS-TR** (743 clips) and a
65
+ second real-human set **MediaSpeech-TR** (200-clip subset), both audited absent from training. Latency on a
66
+ single CPU thread (AMD EPYC 7282).
67
+
68
+ | System | Class | Params | FLEURS WER | MediaSpeech WER | CPU RTF |
69
+ |---|---|---:|---:|---:|---:|
70
+ | **turkish-stt** | **streaming · CPU** | **66 M** | **13.90%** | **15.89%** | **0.128** |
71
+ | Vosk-TR | streaming · CPU | ~50 M | 34.09% | 31.02% | 0.138 |
72
+ | Whisper small | offline · GPU | 244 M | 15.12% | — | 1.098 |
73
+ | Whisper large-v3 | offline · GPU | 1.5 B | 7.08% | — | 6.095 |
74
+
75
+ Among **streaming, CPU-deployable** systems turkish-stt more than halves Vosk's error on both real-human sets,
76
+ matches the four-times-larger offline Whisper-small, and is the only option that is real-time on CPU — it also
77
+ degrades more gracefully than Vosk under additive noise. A detailed technical report — full methodology, ablations,
78
+ latency/concurrency study, and the leakage audits — is being prepared and will be shared.
79
+
80
+ ## Reproduce the numbers
81
+
82
+ The evaluation set, references, and scripts are in [`eval/`](eval) — including the text (n-gram) and acoustic
83
+ (chromaprint) contamination audits used to verify the held-out sets are absent from training.
84
+
85
+ ```bash
86
+ pip install "turkish-stt[eval]"
87
+ python eval/run_eval.py --data-dir eval/tts-synthetic
88
+ ```
89
+
90
+ ## License
91
+
92
+ **mihu Community License** — free for research, personal, and limited commercial use. See [`LICENSE`](LICENSE)
93
+ for the full terms. Weights are released under the same license on Hugging Face.
@@ -0,0 +1,120 @@
1
+ # -*- coding: utf-8 -*-
2
+ """
3
+ turkish-stt — real-time Turkish speech-to-text, on the CPU.
4
+
5
+ import mihu_stt
6
+
7
+ # one-shot
8
+ text = mihu_stt.transcribe("audio.wav")
9
+
10
+ # streaming (voice agents): feed audio as it arrives
11
+ stt = mihu_stt.StreamingSTT()
12
+ text, is_final = stt.accept_pcm16(pcm16_bytes, sample_rate=16000)
13
+
14
+ Weights download once from Hugging Face (mihuai/turkish-stt) and are cached locally. No API key,
15
+ no GPU, no cloud — Turkish-only, runs faster than real time on a single CPU thread.
16
+ """
17
+ import numpy as np
18
+
19
+ __version__ = "0.1.0"
20
+ _HF_REPO = "mihuai/turkish-stt"
21
+ _cache = {}
22
+
23
+
24
+ def _model_files(model_dir=None, quantized=True):
25
+ # int8 encoder is the default (smaller/faster). The fp32 encoder is a byte-exact
26
+ # twin of the same model — quality is effectively identical (~0.1% WER), it just
27
+ # trades ~260 MB for no quantization rounding. Kept as an opt-in for anyone who wants it.
28
+ encoder = "encoder.int8.onnx" if quantized else "encoder.onnx"
29
+ if model_dir:
30
+ import os
31
+ j = lambda f: os.path.join(model_dir, f)
32
+ return j("tokens.txt"), j(encoder), j("decoder.onnx"), j("joiner.int8.onnx")
33
+ from huggingface_hub import hf_hub_download
34
+ dl = lambda f: hf_hub_download(_HF_REPO, f)
35
+ return dl("tokens.txt"), dl(encoder), dl("decoder.onnx"), dl("joiner.int8.onnx")
36
+
37
+
38
+ def _build(model_dir=None, num_threads=2, endpoint=True, quantized=True):
39
+ import sherpa_onnx
40
+ tokens, encoder, decoder, joiner = _model_files(model_dir, quantized)
41
+ kw = dict(
42
+ tokens=tokens, encoder=encoder, decoder=decoder, joiner=joiner,
43
+ num_threads=num_threads, decoding_method="modified_beam_search", provider="cpu",
44
+ )
45
+ if endpoint:
46
+ kw.update(enable_endpoint_detection=True, rule1_min_trailing_silence=2.4,
47
+ rule2_min_trailing_silence=1.2, rule3_min_utterance_length=300)
48
+ return sherpa_onnx.OnlineRecognizer.from_transducer(**kw)
49
+
50
+
51
+ def _to_16k_float(audio, sample_rate):
52
+ if isinstance(audio, str):
53
+ import soundfile as sf
54
+ audio, sample_rate = sf.read(audio, dtype="float32")
55
+ audio = np.asarray(audio, dtype=np.float32)
56
+ if audio.ndim > 1:
57
+ audio = audio.mean(axis=1)
58
+ if audio.dtype.kind == "i":
59
+ audio = audio.astype(np.float32) / 32768.0
60
+ if sample_rate != 16000:
61
+ import soxr
62
+ audio = soxr.resample(audio, sample_rate, 16000)
63
+ return audio
64
+
65
+
66
+ def transcribe(audio, sample_rate=16000, model_dir=None, quantized=True):
67
+ """Transcribe a wav path or a mono numpy array (float32 [-1,1] or int16). Returns Turkish text.
68
+
69
+ quantized=True -> int8 encoder (default): smaller, faster.
70
+ quantized=False -> fp32 encoder: same accuracy, larger download.
71
+ """
72
+ rec = _cache.get(("batch", model_dir, quantized))
73
+ if rec is None:
74
+ rec = _cache[("batch", model_dir, quantized)] = _build(model_dir, endpoint=False, quantized=quantized)
75
+ x = _to_16k_float(audio, sample_rate)
76
+ s = rec.create_stream()
77
+ s.accept_waveform(16000, x)
78
+ s.accept_waveform(16000, np.zeros(8000, dtype=np.float32))
79
+ s.input_finished()
80
+ while rec.is_ready(s):
81
+ rec.decode_stream(s)
82
+ r = rec.get_result(s)
83
+ return (r if isinstance(r, str) else r.text).strip()
84
+
85
+
86
+ class StreamingSTT:
87
+ """
88
+ Streaming recognizer with turn-end (endpoint) detection.
89
+
90
+ stt = StreamingSTT()
91
+ text, is_final = stt.accept_pcm16(pcm16_bytes, sample_rate)
92
+
93
+ `is_final` becomes True when the model detects end-of-utterance; the stream then resets so the
94
+ next turn starts clean. Ideal for LiveKit / Pipecat and other real-time voice pipelines.
95
+ """
96
+ def __init__(self, model_dir=None, num_threads=2, quantized=True):
97
+ self._rec = _build(model_dir, num_threads=num_threads, endpoint=True, quantized=quantized)
98
+ self._stream = self._rec.create_stream()
99
+
100
+ def _text(self):
101
+ r = self._rec.get_result(self._stream)
102
+ return (r if isinstance(r, str) else r.text).strip()
103
+
104
+ def accept_pcm16(self, pcm_bytes: bytes, sample_rate: int = 16000):
105
+ x = np.frombuffer(pcm_bytes, dtype=np.int16).astype(np.float32) / 32768.0
106
+ if sample_rate != 16000:
107
+ import soxr
108
+ x = soxr.resample(x, sample_rate, 16000)
109
+ self._stream.accept_waveform(16000, x)
110
+ while self._rec.is_ready(self._stream):
111
+ self._rec.decode_stream(self._stream)
112
+ text = self._text()
113
+ is_final = self._rec.is_endpoint(self._stream)
114
+ if is_final:
115
+ self._rec.reset(self._stream)
116
+ return text, is_final
117
+
118
+ def accept_float32(self, samples, sample_rate: int = 16000):
119
+ pcm = (np.clip(np.asarray(samples, dtype=np.float32), -1, 1) * 32768).astype(np.int16).tobytes()
120
+ return self.accept_pcm16(pcm, sample_rate)
@@ -0,0 +1,36 @@
1
+ [build-system]
2
+ requires = ["setuptools>=61"]
3
+ build-backend = "setuptools.build_meta"
4
+
5
+ [project]
6
+ name = "turkish-stt"
7
+ version = "0.1.0"
8
+ description = "Real-time, low-latency Turkish speech-to-text (STT) that runs on the CPU."
9
+ readme = "README.md"
10
+ requires-python = ">=3.9"
11
+ license = { file = "LICENSE" }
12
+ authors = [{ name = "mihu · Speech Systems" }]
13
+ maintainers = [{ name = "Deniz Gökdünek", email = "deniz@mihu.ai" }]
14
+ keywords = ["speech-recognition", "asr", "streaming", "turkish", "on-device", "cpu", "voice-agent"]
15
+ classifiers = [
16
+ "Programming Language :: Python :: 3",
17
+ "License :: Other/Proprietary License",
18
+ "Topic :: Multimedia :: Sound/Audio :: Speech",
19
+ "Operating System :: OS Independent",
20
+ ]
21
+ dependencies = [
22
+ "sherpa-onnx",
23
+ "numpy",
24
+ "soundfile",
25
+ "soxr",
26
+ "huggingface_hub",
27
+ ]
28
+
29
+ [project.optional-dependencies]
30
+ eval = ["jiwer", "num2words"]
31
+
32
+ [project.urls]
33
+ "Model weights" = "https://huggingface.co/mihuai/turkish-stt"
34
+
35
+ [tool.setuptools.packages.find]
36
+ include = ["mihu_stt*"]
@@ -0,0 +1,4 @@
1
+ [egg_info]
2
+ tag_build =
3
+ tag_date = 0
4
+
@@ -0,0 +1,198 @@
1
+ Metadata-Version: 2.4
2
+ Name: turkish-stt
3
+ Version: 0.1.0
4
+ Summary: Real-time, low-latency Turkish speech-to-text (STT) that runs on the CPU.
5
+ Author: mihu · Speech Systems
6
+ Maintainer-email: Deniz Gökdünek <deniz@mihu.ai>
7
+ License: mihu Community License
8
+ Version 1.0
9
+
10
+ These terms govern all use of the turkish-stt materials released by mihu
11
+ ("mihu", "Licensor"). By downloading, accessing, or otherwise using the Materials, you
12
+ ("Licensee") accept and agree to be bound by this License.
13
+
14
+ 1. DEFINITIONS
15
+ 1.1 "Materials" — the turkish-stt model weights, model card, and any
16
+ accompanying code, documentation, or related artifacts released under this License.
17
+ 1.2 "Outputs" — anything produced by running the Materials.
18
+ 1.3 "Derivative" — any work created from, adapted from, or built upon the Materials.
19
+ 1.4 "Annual Revenue" — the combined worldwide gross revenue of Licensee together with
20
+ its affiliates over the preceding twelve months.
21
+
22
+ 2. COMMUNITY GRANT
23
+ Provided Licensee's Annual Revenue stays below USD 500,000, and subject to every term
24
+ below, Licensor grants Licensee a worldwide, non-exclusive, non-transferable,
25
+ royalty-free, no-cost right to use, reproduce, and make Derivatives of the Materials.
26
+ This covers, without limitation:
27
+ (a) research and academic work;
28
+ (b) personal and other non-commercial use;
29
+ (c) commercial use within the revenue limit above, once any commercial-use
30
+ registration that Licensor requires has been completed.
31
+
32
+ 3. ABOVE THE THRESHOLD
33
+ Once Licensee's Annual Revenue reaches USD 500,000 or more, no use is permitted under
34
+ the Community Grant. Licensee must first secure a separate commercial agreement from
35
+ mihu. Write to support@mihu.ai to arrange terms.
36
+
37
+ 4. ATTRIBUTION
38
+ Licensee shall (a) keep this License intact, together with any copyright or attribution
39
+ notices, in a "NOTICE" file shipped alongside the Materials or any Derivative, and
40
+ (b) show "Powered by mihu" clearly in any product, service, interface, or documentation
41
+ that builds on the Materials or their Outputs.
42
+
43
+ 5. RESTRICTIONS
44
+ Licensee shall not:
45
+ (a) use the Materials or Outputs to train, build, or refine any competing speech,
46
+ foundational, or general-purpose AI model;
47
+ (b) strip, hide, or modify any attribution or proprietary notice;
48
+ (c) use the mihu name, logo, or marks beyond the attribution required in Section 4
49
+ (no trademark rights are granted here);
50
+ (d) use the Materials unlawfully or contrary to any Acceptable Use Policy that
51
+ Licensor may publish;
52
+ (e) sublicense, sell, or redistribute the Materials except where this License
53
+ expressly allows it.
54
+
55
+ 6. OWNERSHIP
56
+ Between the parties, Licensor keeps all right, title, and interest in the Materials.
57
+ Licensee holds the Derivatives it lawfully makes and the Outputs it generates, always
58
+ subject to Licensor's rights in the underlying Materials. Nothing beyond the rights
59
+ stated here is granted, by implication or otherwise.
60
+
61
+ 7. NO WARRANTY
62
+ THE MATERIALS ARE SUPPLIED "AS IS", WITH NO WARRANTY OF ANY KIND, WHETHER EXPRESS OR
63
+ IMPLIED, INCLUDING ANY IMPLIED WARRANTY OF MERCHANTABILITY, FITNESS FOR A PARTICULAR
64
+ PURPOSE, OR NON-INFRINGEMENT.
65
+
66
+ 8. LIMITATION OF LIABILITY
67
+ TO THE FULLEST EXTENT THE LAW ALLOWS, LICENSOR IS NOT LIABLE FOR ANY INDIRECT,
68
+ INCIDENTAL, SPECIAL, CONSEQUENTIAL, OR EXEMPLARY DAMAGES CONNECTED TO THE MATERIALS OR
69
+ THIS LICENSE.
70
+
71
+ 9. TERMINATION
72
+ This License ends automatically if (a) Licensee's Annual Revenue reaches or passes
73
+ USD 500,000 without a commercial agreement in place, (b) Licensee breaches any term, or
74
+ (c) Licensee brings litigation or a patent claim against Licensor over the Materials.
75
+ On termination, Licensee must stop all use and delete every copy of the Materials. To
76
+ keep using them, arrange a commercial agreement via support@mihu.ai.
77
+
78
+ 10. GENERAL
79
+ This License is the complete agreement covering the Materials. It is governed by the
80
+ laws of Licensor's principal place of business, disregarding conflict-of-law rules.
81
+
82
+ © mihu. All rights reserved.
83
+
84
+ turkish-stt and the mihu Community License are provided by Hunters AI, operating
85
+ under the mihu brand. "mihu" is a trademark of Hunters AI.
86
+
87
+ Project-URL: Model weights, https://huggingface.co/mihuai/turkish-stt
88
+ Keywords: speech-recognition,asr,streaming,turkish,on-device,cpu,voice-agent
89
+ Classifier: Programming Language :: Python :: 3
90
+ Classifier: License :: Other/Proprietary License
91
+ Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
92
+ Classifier: Operating System :: OS Independent
93
+ Requires-Python: >=3.9
94
+ Description-Content-Type: text/markdown
95
+ License-File: LICENSE
96
+ Requires-Dist: sherpa-onnx
97
+ Requires-Dist: numpy
98
+ Requires-Dist: soundfile
99
+ Requires-Dist: soxr
100
+ Requires-Dist: huggingface_hub
101
+ Provides-Extra: eval
102
+ Requires-Dist: jiwer; extra == "eval"
103
+ Requires-Dist: num2words; extra == "eval"
104
+ Dynamic: license-file
105
+
106
+ # turkish-stt
107
+
108
+ **Real-time, low-latency speech-to-text (STT) for Turkish — on the CPU.**
109
+
110
+ Turkish-only · streaming · CPU-native · runs on the edge. A compact **66 M-parameter** causal streaming model,
111
+ built for voice agents, call centers, and on-device transcription where a GPU is not an option — no API key, no
112
+ cloud, no GPU.
113
+
114
+ - 🇹🇷 **Turkish-first** — every parameter serves one language, not ~100
115
+ - ⚡ **Streaming** — partial results as you speak (~320 ms), no 30-second windows
116
+ - 🖥️ **CPU-native** — **RTF 0.128** on a single CPU thread (~8× faster than real time)
117
+ - 🔌 **Tiny** — ~68 MB; runs on a Raspberry-Pi-class device
118
+ - 🏷️ **Entity/brand aware** — tuned for Turkish brands, models, and numbers
119
+
120
+ ## Installation
121
+
122
+ ```bash
123
+ pip install turkish-stt
124
+ ```
125
+
126
+ ## Usage
127
+
128
+ ```python
129
+ import mihu_stt
130
+
131
+ # one-shot: transcribe a file (16 kHz mono, or anything soundfile reads)
132
+ text = mihu_stt.transcribe("audio.wav")
133
+ print(text)
134
+ ```
135
+
136
+ Streaming — feed audio as it arrives and read an updating transcript, with turn-end detection built in:
137
+
138
+ ```python
139
+ import mihu_stt
140
+
141
+ stt = mihu_stt.StreamingSTT()
142
+ # pcm16 = 16-bit PCM bytes from your mic / phone line
143
+ text, is_final = stt.accept_pcm16(pcm16, sample_rate=16000)
144
+ # `text` updates while speech continues; `is_final` turns True at end-of-utterance
145
+ ```
146
+
147
+ Model weights download automatically from Hugging Face
148
+ ([`mihuai/turkish-stt`](https://huggingface.co/mihuai/turkish-stt)) the first time and are cached.
149
+
150
+ ## Integrations
151
+
152
+ Drop-in adapters for the common voice-agent stacks (see [`examples/`](examples)):
153
+
154
+ **LiveKit Agents** — [`examples/livekit_stt.py`](examples/livekit_stt.py):
155
+ ```python
156
+ from livekit.agents import AgentSession
157
+ from livekit_stt import MihuSTT
158
+ session = AgentSession(stt=MihuSTT(), llm=..., tts=...)
159
+ ```
160
+
161
+ **Pipecat** — [`examples/pipecat_stt.py`](examples/pipecat_stt.py):
162
+ ```python
163
+ from pipecat_stt import MihuSTTService
164
+ pipeline = Pipeline([transport.input(), vad, MihuSTTService(), llm, tts, transport.output()])
165
+ ```
166
+
167
+ ## Benchmarks
168
+
169
+ Measured in-house under one matched Turkish normalization. Real-human held-out **FLEURS-TR** (743 clips) and a
170
+ second real-human set **MediaSpeech-TR** (200-clip subset), both audited absent from training. Latency on a
171
+ single CPU thread (AMD EPYC 7282).
172
+
173
+ | System | Class | Params | FLEURS WER | MediaSpeech WER | CPU RTF |
174
+ |---|---|---:|---:|---:|---:|
175
+ | **turkish-stt** | **streaming · CPU** | **66 M** | **13.90%** | **15.89%** | **0.128** |
176
+ | Vosk-TR | streaming · CPU | ~50 M | 34.09% | 31.02% | 0.138 |
177
+ | Whisper small | offline · GPU | 244 M | 15.12% | — | 1.098 |
178
+ | Whisper large-v3 | offline · GPU | 1.5 B | 7.08% | — | 6.095 |
179
+
180
+ Among **streaming, CPU-deployable** systems turkish-stt more than halves Vosk's error on both real-human sets,
181
+ matches the four-times-larger offline Whisper-small, and is the only option that is real-time on CPU — it also
182
+ degrades more gracefully than Vosk under additive noise. A detailed technical report — full methodology, ablations,
183
+ latency/concurrency study, and the leakage audits — is being prepared and will be shared.
184
+
185
+ ## Reproduce the numbers
186
+
187
+ The evaluation set, references, and scripts are in [`eval/`](eval) — including the text (n-gram) and acoustic
188
+ (chromaprint) contamination audits used to verify the held-out sets are absent from training.
189
+
190
+ ```bash
191
+ pip install "turkish-stt[eval]"
192
+ python eval/run_eval.py --data-dir eval/tts-synthetic
193
+ ```
194
+
195
+ ## License
196
+
197
+ **mihu Community License** — free for research, personal, and limited commercial use. See [`LICENSE`](LICENSE)
198
+ for the full terms. Weights are released under the same license on Hugging Face.
@@ -0,0 +1,9 @@
1
+ LICENSE
2
+ README.md
3
+ pyproject.toml
4
+ mihu_stt/__init__.py
5
+ turkish_stt.egg-info/PKG-INFO
6
+ turkish_stt.egg-info/SOURCES.txt
7
+ turkish_stt.egg-info/dependency_links.txt
8
+ turkish_stt.egg-info/requires.txt
9
+ turkish_stt.egg-info/top_level.txt
@@ -0,0 +1,9 @@
1
+ sherpa-onnx
2
+ numpy
3
+ soundfile
4
+ soxr
5
+ huggingface_hub
6
+
7
+ [eval]
8
+ jiwer
9
+ num2words
@@ -0,0 +1 @@
1
+ mihu_stt