myvoiceai 0.0.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,217 @@
1
+ Metadata-Version: 2.4
2
+ Name: myvoiceai
3
+ Version: 0.0.1
4
+ Summary: Voice AI agent – real-time WebSocket voice session pipeline
5
+ Author: Your Name
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/yourorg/myvoiceai
8
+ Project-URL: Repository, https://github.com/yourorg/myvoiceai
9
+ Classifier: Development Status :: 3 - Alpha
10
+ Classifier: Intended Audience :: Developers
11
+ Classifier: Programming Language :: Python :: 3
12
+ Classifier: Programming Language :: Python :: 3.8
13
+ Classifier: Programming Language :: Python :: 3.9
14
+ Classifier: Programming Language :: Python :: 3.10
15
+ Classifier: Programming Language :: Python :: 3.11
16
+ Classifier: Programming Language :: Python :: 3.12
17
+ Requires-Python: >=3.8
18
+ Description-Content-Type: text/markdown
19
+ Requires-Dist: fastapi>=0.100
20
+ Requires-Dist: uvicorn[standard]>=0.23
21
+ Requires-Dist: websockets>=12.0
22
+ Requires-Dist: httpx>=0.24
23
+ Requires-Dist: python-dotenv>=1.0
24
+ Requires-Dist: litellm>=0.13
25
+ Provides-Extra: observability
26
+ Requires-Dist: opentelemetry-sdk>=1.20; extra == "observability"
27
+ Requires-Dist: opentelemetry-exporter-otlp>=1.20; extra == "observability"
28
+
29
+ # myvoiceai
30
+
31
+ A reusable voice AI session package for real-time WebSocket voice pipelines.
32
+
33
+ ## Installation
34
+
35
+ Install the package from PyPI:
36
+
37
+ ```bash
38
+ pip install myvoiceai
39
+ ```
40
+
41
+ Install with optional observability support:
42
+
43
+ ```bash
44
+ pip install myvoiceai[observability]
45
+ ```
46
+
47
+ ## Quickstart
48
+
49
+ Use the `run_voice_session()` function to start a voice session. Pass API keys and session config as arguments (no required `.env`):
50
+
51
+ ```python
52
+ from myvoiceai import run_voice_session
53
+
54
+ await run_voice_session(
55
+ websocket=client_websocket,
56
+ system_prompt="You are a helpful assistant.",
57
+ greeting_message="Hello! How can I help you today?",
58
+ session_id="session-001",
59
+ model="gemini-2.5-flash",
60
+ llm_provider_api_key="your-gemini-key",
61
+ deepgram_api_key="your-deepgram-key",
62
+ tracing=False,
63
+ endpointing=1200,
64
+ utterance_end=2500,
65
+ stable_interim_secs=1.5,
66
+ stable_interim_secs_no_punct=3.0,
67
+ inactivity_timeout_seconds=30,
68
+ )
69
+ ```
70
+
71
+ ## Custom Tools
72
+
73
+ Pass only one tool (`echo`) with an OpenAI-style `function` schema:
74
+
75
+ ```python
76
+ from myvoiceai.tools import ToolRegistry
77
+
78
+ my_registry = ToolRegistry()
79
+
80
+ async def echo(**kwargs):
81
+ return kwargs.get("text", "")
82
+
83
+ my_registry.register(
84
+ name="echo",
85
+ schema={
86
+ "type": "function",
87
+ "function": {
88
+ "name": "echo",
89
+ "description": "Echo back the text provided by the LLM.",
90
+ "parameters": {
91
+ "type": "object",
92
+ "properties": {
93
+ "text": {"type": "string", "description": "Text to echo back"}
94
+ },
95
+ "required": ["text"],
96
+ },
97
+ },
98
+ },
99
+ impl=echo,
100
+ )
101
+
102
+ await run_voice_session(
103
+ websocket=ws,
104
+ tool_registry=my_registry,
105
+ llm_provider_api_key="your-gemini-key",
106
+ deepgram_api_key="your-deepgram-key",
107
+ )
108
+ ```
109
+
110
+ ## Example Server
111
+
112
+ Run the FastAPI example server (keys passed via `.env` in the example, but they can also be hardcoded or passed through config):
113
+
114
+ ```bash
115
+ uvicorn myvoiceai.example.fastapi_app:app --host 0.0.0.0 --port 8000
116
+ ```
117
+
118
+ Connect WebSocket clients to `ws://localhost:8000/ws/voice`.
119
+
120
+ ## Key Configuration
121
+
122
+ Keys are passed directly to `run_voice_session()` or `CustomVoiceAgent`:
123
+
124
+ - `api_key` — LLM provider API key (e.g., Gemini / OpenAI)
125
+ - `deepgram_api_key` — Deepgram STT/TTS key
126
+ - `session_id` — Session identifier for tracing/logging (replaces old `interview_id`)
127
+ - `model` — Model identifier (e.g., `gemini-2.5-flash`)
128
+ - `tracing` — `True` only if `[observability]` extras are installed; default `False`
129
+
130
+ A `.env.example` is included for local convenience but is optional.
131
+
132
+ ## API Reference
133
+
134
+ ### `run_voice_session(websocket, ...)`
135
+ Main entry point. Parameter reference:
136
+
137
+ | Parameter | Use / Purpose |
138
+ |---|---|
139
+ | `websocket` | WebSocket connection from FastAPI / client |
140
+ | `system_prompt` | LLM system instruction (default: concise voice assistant) |
141
+ | `greeting_message` | First spoken message to user |
142
+ | `session_id` | Session identifier for tracing/logs (replaces old `interview_id`) |
143
+ | `model` | LLM model ID (e.g., `gemini/gemini-2.5-flash`) |
144
+ | `llm_provider_api_key` | LLM provider API key (required) |
145
+ | `deepgram_api_key` | Deepgram STT/TTS API key (required) |
146
+ | `stt_model` | Deepgram STT model (default `nova-2`) |
147
+ | `tts_model` | Deepgram TTS voice model (default `aura-asteria-en`) |
148
+ | `tracing` | `True` only with `[observability]` installed; sends traces to OTLP |
149
+ | `tool_registry` | `ToolRegistry` with custom `function` schemas |
150
+ | `endpointing` | Pause duration (ms) after endpointing |
151
+ | `utterance_end` | Stop recording (ms) when utterance ends |
152
+ | `stable_interim_secs` | Stable interim result delay (secs) |
153
+ | `stable_interim_secs_no_punct` | Same, when no punctuation detected |
154
+ | `inactivity_timeout_seconds` | Close session after silence (default 10) |
155
+ | `max_session_seconds` | Hard session time cap (`CustomVoiceAgent` only) |
156
+
157
+ ### `CustomVoiceAgent`
158
+ Core pipeline class. Same params as `run_voice_session` plus `client_websocket`, `max_session_seconds`, `max_duration_message`, `inactivity_message`.
159
+
160
+ ### `CustomVoiceAgent`
161
+ Core pipeline class. Initialize with `client_websocket` and optional `system_prompt`, `greeting_message`, `session_id`, `model`, `api_key`, `deepgram_api_key`, `tracing`, `tool_registry`.
162
+
163
+ ### `ToolRegistry`
164
+ Dynamic registry for callable tools. Use `.register()` with full `function` schema, `.lookup()` to retrieve, `.schemas()` to get LLM-ready tool definitions.
165
+
166
+ ## Observability / Tracing
167
+
168
+ Install extras and run the local collector stack to enable tracing:
169
+
170
+ ```bash
171
+ pip install myvoiceai[observability]
172
+ ```
173
+
174
+ Example infrastructure files are included in `example/`:
175
+
176
+ - `docker-compose.yml` — Jaeger (UI at `localhost:16686`) + OpenTelemetry Collector (`4318`/`4319`)
177
+ - `otel-collector-config.yaml` — Collector pipeline: OTLP → batch → Jaeger + Prometheus metrics
178
+ - `prometheus.yml` — Scrapes collector metrics at `localhost:8889`
179
+
180
+ Start the stack:
181
+
182
+ ```bash
183
+ cd example
184
+ # Set OTLP endpoint in your app to http://localhost:4318 (HTTP) or localhost:4317 (gRPC)
185
+ docker-compose up -d
186
+ ```
187
+
188
+ Then run the session with tracing enabled:
189
+
190
+ ```python
191
+ await run_voice_session(
192
+ websocket=ws,
193
+ tracing=True,
194
+ session_id="session-001",
195
+ api_key="...",
196
+ deepgram_api_key="...",
197
+ )
198
+ ```
199
+
200
+ Jaeger UI: `http://localhost:16686` (search by `session.id`).
201
+
202
+ ## WebSocket Messages
203
+
204
+ Server sends JSON over the WebSocket:
205
+
206
+ - `transcript_chunk` — incremental partial transcript (`role`: `user` or `assistant`; `turn_id`; `text`; optional `replace`)
207
+ - `transcript` — finalized transcript (`role`; `text`; `turn_id`)
208
+ - `turn` — completed turn (`turn_id`; `timestamp`; `user`; `assistant`; `interrupted`: bool)
209
+
210
+ Client should listen for these to update UI/state.
211
+
212
+ ## Notes
213
+
214
+ - The package does **not** include infrastructure files like `docker-compose.yml` or `prometheus.yml`. These are maintained separately.
215
+ - Observability (OpenTelemetry) is optional via `pip install myvoiceai[observability]`. Without it, `tracing=False` runs normally.
216
+ - `.env` is excluded from the package for security. Only `.env.example` ships for reference.
217
+ - All tool schemas must use the OpenAI-style `{"type": "function", "function": {...}}` format.
@@ -0,0 +1,189 @@
1
+ # myvoiceai
2
+
3
+ A reusable voice AI session package for real-time WebSocket voice pipelines.
4
+
5
+ ## Installation
6
+
7
+ Install the package from PyPI:
8
+
9
+ ```bash
10
+ pip install myvoiceai
11
+ ```
12
+
13
+ Install with optional observability support:
14
+
15
+ ```bash
16
+ pip install myvoiceai[observability]
17
+ ```
18
+
19
+ ## Quickstart
20
+
21
+ Use the `run_voice_session()` function to start a voice session. Pass API keys and session config as arguments (no required `.env`):
22
+
23
+ ```python
24
+ from myvoiceai import run_voice_session
25
+
26
+ await run_voice_session(
27
+ websocket=client_websocket,
28
+ system_prompt="You are a helpful assistant.",
29
+ greeting_message="Hello! How can I help you today?",
30
+ session_id="session-001",
31
+ model="gemini-2.5-flash",
32
+ llm_provider_api_key="your-gemini-key",
33
+ deepgram_api_key="your-deepgram-key",
34
+ tracing=False,
35
+ endpointing=1200,
36
+ utterance_end=2500,
37
+ stable_interim_secs=1.5,
38
+ stable_interim_secs_no_punct=3.0,
39
+ inactivity_timeout_seconds=30,
40
+ )
41
+ ```
42
+
43
+ ## Custom Tools
44
+
45
+ Pass only one tool (`echo`) with an OpenAI-style `function` schema:
46
+
47
+ ```python
48
+ from myvoiceai.tools import ToolRegistry
49
+
50
+ my_registry = ToolRegistry()
51
+
52
+ async def echo(**kwargs):
53
+ return kwargs.get("text", "")
54
+
55
+ my_registry.register(
56
+ name="echo",
57
+ schema={
58
+ "type": "function",
59
+ "function": {
60
+ "name": "echo",
61
+ "description": "Echo back the text provided by the LLM.",
62
+ "parameters": {
63
+ "type": "object",
64
+ "properties": {
65
+ "text": {"type": "string", "description": "Text to echo back"}
66
+ },
67
+ "required": ["text"],
68
+ },
69
+ },
70
+ },
71
+ impl=echo,
72
+ )
73
+
74
+ await run_voice_session(
75
+ websocket=ws,
76
+ tool_registry=my_registry,
77
+ llm_provider_api_key="your-gemini-key",
78
+ deepgram_api_key="your-deepgram-key",
79
+ )
80
+ ```
81
+
82
+ ## Example Server
83
+
84
+ Run the FastAPI example server (keys passed via `.env` in the example, but they can also be hardcoded or passed through config):
85
+
86
+ ```bash
87
+ uvicorn myvoiceai.example.fastapi_app:app --host 0.0.0.0 --port 8000
88
+ ```
89
+
90
+ Connect WebSocket clients to `ws://localhost:8000/ws/voice`.
91
+
92
+ ## Key Configuration
93
+
94
+ Keys are passed directly to `run_voice_session()` or `CustomVoiceAgent`:
95
+
96
+ - `api_key` — LLM provider API key (e.g., Gemini / OpenAI)
97
+ - `deepgram_api_key` — Deepgram STT/TTS key
98
+ - `session_id` — Session identifier for tracing/logging (replaces old `interview_id`)
99
+ - `model` — Model identifier (e.g., `gemini-2.5-flash`)
100
+ - `tracing` — `True` only if `[observability]` extras are installed; default `False`
101
+
102
+ A `.env.example` is included for local convenience but is optional.
103
+
104
+ ## API Reference
105
+
106
+ ### `run_voice_session(websocket, ...)`
107
+ Main entry point. Parameter reference:
108
+
109
+ | Parameter | Use / Purpose |
110
+ |---|---|
111
+ | `websocket` | WebSocket connection from FastAPI / client |
112
+ | `system_prompt` | LLM system instruction (default: concise voice assistant) |
113
+ | `greeting_message` | First spoken message to user |
114
+ | `session_id` | Session identifier for tracing/logs (replaces old `interview_id`) |
115
+ | `model` | LLM model ID (e.g., `gemini/gemini-2.5-flash`) |
116
+ | `llm_provider_api_key` | LLM provider API key (required) |
117
+ | `deepgram_api_key` | Deepgram STT/TTS API key (required) |
118
+ | `stt_model` | Deepgram STT model (default `nova-2`) |
119
+ | `tts_model` | Deepgram TTS voice model (default `aura-asteria-en`) |
120
+ | `tracing` | `True` only with `[observability]` installed; sends traces to OTLP |
121
+ | `tool_registry` | `ToolRegistry` with custom `function` schemas |
122
+ | `endpointing` | Pause duration (ms) after endpointing |
123
+ | `utterance_end` | Stop recording (ms) when utterance ends |
124
+ | `stable_interim_secs` | Stable interim result delay (secs) |
125
+ | `stable_interim_secs_no_punct` | Same, when no punctuation detected |
126
+ | `inactivity_timeout_seconds` | Close session after silence (default 10) |
127
+ | `max_session_seconds` | Hard session time cap (`CustomVoiceAgent` only) |
128
+
129
+ ### `CustomVoiceAgent`
130
+ Core pipeline class. Same params as `run_voice_session` plus `client_websocket`, `max_session_seconds`, `max_duration_message`, `inactivity_message`.
131
+
132
+ ### `CustomVoiceAgent`
133
+ Core pipeline class. Initialize with `client_websocket` and optional `system_prompt`, `greeting_message`, `session_id`, `model`, `api_key`, `deepgram_api_key`, `tracing`, `tool_registry`.
134
+
135
+ ### `ToolRegistry`
136
+ Dynamic registry for callable tools. Use `.register()` with full `function` schema, `.lookup()` to retrieve, `.schemas()` to get LLM-ready tool definitions.
137
+
138
+ ## Observability / Tracing
139
+
140
+ Install extras and run the local collector stack to enable tracing:
141
+
142
+ ```bash
143
+ pip install myvoiceai[observability]
144
+ ```
145
+
146
+ Example infrastructure files are included in `example/`:
147
+
148
+ - `docker-compose.yml` — Jaeger (UI at `localhost:16686`) + OpenTelemetry Collector (`4318`/`4319`)
149
+ - `otel-collector-config.yaml` — Collector pipeline: OTLP → batch → Jaeger + Prometheus metrics
150
+ - `prometheus.yml` — Scrapes collector metrics at `localhost:8889`
151
+
152
+ Start the stack:
153
+
154
+ ```bash
155
+ cd example
156
+ # Set OTLP endpoint in your app to http://localhost:4318 (HTTP) or localhost:4317 (gRPC)
157
+ docker-compose up -d
158
+ ```
159
+
160
+ Then run the session with tracing enabled:
161
+
162
+ ```python
163
+ await run_voice_session(
164
+ websocket=ws,
165
+ tracing=True,
166
+ session_id="session-001",
167
+ api_key="...",
168
+ deepgram_api_key="...",
169
+ )
170
+ ```
171
+
172
+ Jaeger UI: `http://localhost:16686` (search by `session.id`).
173
+
174
+ ## WebSocket Messages
175
+
176
+ Server sends JSON over the WebSocket:
177
+
178
+ - `transcript_chunk` — incremental partial transcript (`role`: `user` or `assistant`; `turn_id`; `text`; optional `replace`)
179
+ - `transcript` — finalized transcript (`role`; `text`; `turn_id`)
180
+ - `turn` — completed turn (`turn_id`; `timestamp`; `user`; `assistant`; `interrupted`: bool)
181
+
182
+ Client should listen for these to update UI/state.
183
+
184
+ ## Notes
185
+
186
+ - The package does **not** include infrastructure files like `docker-compose.yml` or `prometheus.yml`. These are maintained separately.
187
+ - Observability (OpenTelemetry) is optional via `pip install myvoiceai[observability]`. Without it, `tracing=False` runs normally.
188
+ - `.env` is excluded from the package for security. Only `.env.example` ships for reference.
189
+ - All tool schemas must use the OpenAI-style `{"type": "function", "function": {...}}` format.
@@ -0,0 +1,12 @@
1
+ """
2
+ myvoiceai package - main entry point
3
+ """
4
+ from .agent import CustomVoiceAgent, run_voice_session
5
+ from .tools import ToolRegistry, get_default_registry
6
+
7
+ __all__ = [
8
+ "CustomVoiceAgent",
9
+ "run_voice_session",
10
+ "ToolRegistry",
11
+ "get_default_registry",
12
+ ]