@mastra/voice-google-gemini-live 0.14.8-alpha.0 → 0.14.8

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -8,78 +8,13 @@ Google Gemini Live API integration for Mastra, providing real-time multimodal vo
8
8
  npm install @mastra/voice-google-gemini-live
9
9
  ```
10
10
 
11
- ## Configuration
12
-
13
- The module supports two authentication methods:
14
-
15
- ### Option 1: Gemini API (Recommended for development)
16
-
17
- Use an API key from [Google AI Studio](https://makersuite.google.com/app/apikey):
18
-
19
- ```bash
20
- # Set environment variable
21
- GOOGLE_API_KEY=your_api_key
22
- ```
23
-
24
- ### Option 2: Vertex AI (Recommended for production)
25
-
26
- Use OAuth authentication with Google Cloud Platform. There are multiple ways to authenticate:
27
-
28
- #### Application Default Credentials (ADC)
29
-
30
- ```bash
31
- # Install gcloud CLI and authenticate
32
- gcloud auth application-default login
33
-
34
- # Set project ID
35
- GOOGLE_CLOUD_PROJECT=your_project_id
36
- ```
37
-
38
- #### Service Account Key File
39
-
40
- ```bash
41
- # Set path to service account JSON
42
- GOOGLE_APPLICATION_CREDENTIALS=path/to/service-account.json
43
- GOOGLE_CLOUD_PROJECT=your_project_id
44
- ```
45
-
46
- #### Service Account in Code
47
-
48
- ```typescript
49
- const voice = new GeminiLiveVoice({
50
- vertexAI: true,
51
- project: 'your-gcp-project',
52
- location: 'us-central1',
53
- serviceAccountKeyFile: '/path/to/service-account.json',
54
- // OR use service account email for impersonation
55
- serviceAccountEmail: 'service-account@project.iam.gserviceaccount.com',
56
- });
57
- ```
58
-
59
- ### Required Permissions for Vertex AI
60
-
61
- When using Vertex AI, ensure your service account or user has these IAM roles:
62
-
63
- - `aiplatform.user` or specific permissions:
64
- - `aiplatform.endpoints.predict`
65
- - `aiplatform.models.predict`
66
-
67
11
  ## Usage
68
12
 
69
13
  ```typescript
70
14
  import { GeminiLiveVoice } from '@mastra/voice-google-gemini-live';
71
15
 
72
- // Initialize with Gemini API
73
- const voice = new GeminiLiveVoice({
74
- apiKey: 'your-api-key', // Optional, can use GOOGLE_API_KEY env var
75
- model: 'gemini-2.0-flash-live-001',
76
- speaker: 'Puck', // Default voice
77
- });
78
-
79
- // OR initialize with Vertex AI (recommended for production)
80
16
  const voice = new GeminiLiveVoice({
81
- vertexAI: true,
82
- project: 'your-project-id',
17
+ apiKey: process.env.GOOGLE_API_KEY,
83
18
  model: 'gemini-2.0-flash-live-001',
84
19
  speaker: 'Puck',
85
20
  });
@@ -125,356 +60,14 @@ await voice.send(microphoneStream);
125
60
  voice.disconnect();
126
61
  ```
127
62
 
128
- ## API Reference
129
-
130
- ### Constructor
131
-
132
- **`new GeminiLiveVoice(options?: GeminiLiveVoiceConfig)`**
133
-
134
- Creates a new GeminiLiveVoice instance.
135
-
136
- **Parameters:**
137
-
138
- - `options` (optional): Configuration object
139
- - `apiKey?: string` - Google API key (falls back to GOOGLE_API_KEY env var)
140
- - `model?: GeminiVoiceModel` - Model to use (default: 'gemini-2.0-flash-exp')
141
- - `speaker?: GeminiVoiceName` - Voice to use (default: 'Puck')
142
- - `vertexAI?: boolean` - Use Vertex AI instead of Gemini API
143
- - `project?: string` - Google Cloud project ID (required for Vertex AI)
144
- - `location?: string` - Google Cloud region (default: 'us-central1')
145
- - `serviceAccountKeyFile?: string` - Path to service account JSON key file
146
- - `serviceAccountEmail?: string` - Service account email for impersonation
147
- - `instructions?: string` - System instructions for the model
148
- - `tools?: GeminiToolConfig[]` - Tools available to the model
149
- - `sessionConfig?: GeminiSessionConfig` - Session configuration
150
- - `audioConfig?: Partial<AudioConfig>` - Audio configuration
151
- - `debug?: boolean` - Enable debug logging
152
-
153
- ### Connection Management
154
-
155
- **`async connect(): Promise<void>`**
156
-
157
- Establishes connection to the Gemini Live API. Must be called before using other methods.
158
-
159
- **Returns:** Promise that resolves when connection is established
160
-
161
- **Throws:** Error if connection fails or authentication is invalid
162
-
163
- ---
164
-
165
- **`async disconnect(): Promise<void>`**
166
-
167
- Disconnects from the Gemini Live API and cleans up resources.
168
-
169
- **Returns:** Promise that resolves when disconnection is complete
170
-
171
- ---
172
-
173
- **`getConnectionState(): 'disconnected' | 'connected'`**
174
-
175
- Gets the current connection state.
176
-
177
- **Returns:** Current connection state
178
-
179
- ---
180
-
181
- **`isConnected(): boolean`**
182
-
183
- Checks if currently connected to the API.
184
-
185
- **Returns:** true if connected, false otherwise
186
-
187
- ---
188
-
189
- Connection lifecycle transitions such as "connecting", "disconnecting", and "updated" are emitted via the `session` event:
190
-
191
- ```ts
192
- voice.on('session', data => {
193
- // data.state is one of: 'connecting' | 'connected' | 'disconnected' | 'disconnecting' | 'updated'
194
- });
195
- ```
196
-
197
- ### Audio and Speech
198
-
199
- **`async speak(input: string | NodeJS.ReadableStream, options?: GeminiLiveVoiceOptions): Promise<void>`**
200
-
201
- Converts text to speech and sends it to the model.
202
-
203
- **Parameters:**
204
-
205
- - `input: string | NodeJS.ReadableStream` - Text to convert to speech
206
- - `options?: GeminiLiveVoiceOptions` - Optional speech options
207
- - `speaker?: GeminiVoiceName` - Override the default speaker
208
- - `languageCode?: string` - Language code for the response
209
- - `responseModalities?: ('AUDIO' | 'TEXT')[]` - Response modalities
210
-
211
- **Returns:** Promise<void> (responses are emitted via `speaker` and `writing` events)
212
-
213
- **Throws:** Error if not connected or input is empty
214
-
215
- ---
216
-
217
- **`async send(audioData: NodeJS.ReadableStream | Int16Array): Promise<void>`**
218
-
219
- Sends audio data for real-time processing.
220
-
221
- **Parameters:**
222
-
223
- - `audioData: NodeJS.ReadableStream | Int16Array` - Audio data to send
224
-
225
- **Returns:** Promise that resolves when audio is sent
226
-
227
- **Throws:** Error if not connected or audio format is invalid
228
-
229
- ---
230
-
231
- **`async listen(audioStream: NodeJS.ReadableStream, options?: GeminiLiveVoiceOptions): Promise<string>`**
232
-
233
- Processes audio stream for speech-to-text transcription.
234
-
235
- **Parameters:**
236
-
237
- - `audioStream: NodeJS.ReadableStream` - Audio stream to transcribe
238
- - `options?: GeminiLiveVoiceOptions` - Optional transcription options
239
-
240
- **Returns:** Promise that resolves to transcribed text
241
-
242
- **Throws:** Error if not connected, audio format is invalid, or transcription fails
243
-
244
- ---
245
-
246
- **`getCurrentSpeakerStream(): NodeJS.ReadableStream | null`**
247
-
248
- Gets the current concatenated audio stream for the active response.
249
-
250
- **Returns:** ReadableStream of concatenated audio chunks, or null if no active stream
251
-
252
- ### Session Management
253
-
254
- **`async updateSessionConfig(config: Partial<GeminiLiveVoiceConfig>): Promise<void>`**
255
-
256
- Updates session configuration during an active session.
257
-
258
- **Parameters:**
259
-
260
- - `config: Partial<GeminiLiveVoiceConfig>` - Configuration to update
261
- - `speaker?: GeminiVoiceName` - Change voice/speaker
262
- - `instructions?: string` - Update system instructions
263
- - `tools?: GeminiToolConfig[]` - Update available tools
264
- - `sessionConfig?: GeminiSessionConfig` - Update session settings (e.g. `vad`, `interrupts`, `contextCompression`)
265
-
266
- **Returns:** Promise that resolves when configuration is updated
267
-
268
- **Throws:** Error if not connected or update fails
269
-
270
- ---
271
-
272
- **`async resumeSession(handle: string): Promise<void>`**
273
-
274
- Resumes a previous session using a session handle.
275
-
276
- **Parameters:**
277
-
278
- - `handle: string` - Session handle from previous session
279
-
280
- **Returns:** Promise that resolves when session is resumed
281
-
282
- **Note:** Session resumption is not yet fully implemented for Gemini Live API
283
-
284
- ---
285
-
286
- **`getSessionHandle(): string | undefined`**
287
-
288
- Gets the current session handle for resumption.
289
-
290
- **Returns:** Session handle string, or undefined if not available
291
-
292
- **Note:** Session handles are not yet fully supported by Gemini Live API
293
-
294
- ### Voice and Model Information
295
-
296
- **`async getSpeakers(): Promise<Array<{ voiceId: string; description?: string }>>`**
297
-
298
- Gets available speakers/voices.
299
-
300
- **Returns:** Promise that resolves to array of available voices with descriptions
301
-
302
- ---
303
-
304
- **`async getListener(): Promise<{ enabled: boolean }>`**
305
-
306
- Checks if listening capabilities are enabled.
307
-
308
- **Returns:** Promise that resolves to listening status
309
-
310
- **Note:** Inherits default implementation from MastraVoice base class
311
-
312
- ### Event Handling
313
-
314
- **`on<E extends VoiceEventType>(event: E, callback: (data: E extends keyof GeminiLiveEventMap ? GeminiLiveEventMap[E] : unknown) => void): void`**
315
-
316
- Registers an event listener.
317
-
318
- **Parameters:**
319
-
320
- - `event: E` - Event name to listen for
321
- - `callback: (data) => void` - Function to call when event occurs
322
-
323
- **Available Events:**
324
-
325
- - `'speaking'` - Audio response from model
326
- - `'speaker'` - Readable stream of concatenated audio for the active response
327
- - `'writing'` - Transcribed text. Callback receives `{ text, role: 'user' | 'assistant' }`. On native-audio models the assistant transcript is driven by the server's `output_audio_transcription` channel
328
- - `'thinking'` - Model chain-of-thought / reasoning text on native-audio models. Callback receives `{ text }`. Does not fire on non-native-audio models, where reasoning is not surfaced separately
329
- - `'error'` - Error events
330
- - `'session'` - Session state changes
331
- - `'toolCall'` - Tool calls from model
332
- - `'vad'` - Voice activity detection events
333
- - `'interrupt'` - Emitted on barge-in when the user starts speaking over an in-flight model response. Callback receives `{ type: 'user', timestamp }`
334
- - `'usage'` - Token usage information
335
- - `'sessionHandle'` - Session resumption handle
336
- - `'turnComplete'` - Turn completion for the current model response
337
-
338
- #### Native-audio models
339
-
340
- Native-audio models (any model whose ID contains `native-audio`, e.g. `gemini-2.5-flash-native-audio-preview-12-2025`) split text output across two channels:
341
-
342
- - The model's spoken reply is delivered as audio plus an `output_audio_transcription` transcript — surfaced as `writing` with `role: 'assistant'`.
343
- - The model's internal reasoning is delivered as `modelTurn.parts.text` — surfaced as `thinking`.
344
-
345
- On non-native-audio models there is no `output_audio_transcription` channel; `modelTurn.parts.text` is the spoken response itself and is emitted as `writing` (so `thinking` will not fire). Transcription and barge-in detection are enabled automatically in the setup payload — no extra configuration is required.
346
-
347
- ### Tools
348
-
349
- Add tools with `addTools()` using either `@mastra/core/tools` or a plain object matching `ToolsInput`.
350
-
351
- Using `createTool`:
352
-
353
- ```ts
354
- import { createTool } from '@mastra/core/tools';
355
- import { z } from 'zod';
356
-
357
- const searchTool = createTool({
358
- id: 'search',
359
- description: 'Search the web',
360
- inputSchema: z.object({ query: z.string() }),
361
- execute: async inputData => {
362
- const { query } = inputData;
363
- // ... perform search
364
- return { results: [] };
365
- },
366
- });
367
-
368
- voice.addTools({ search: searchTool });
369
- ```
370
-
371
- Using a plain object (ensure each tool has an `id`):
372
-
373
- ```ts
374
- voice.addTools({
375
- search: {
376
- id: 'search',
377
- description: 'Search the web',
378
- inputSchema: { type: 'object', properties: { query: { type: 'string' } } },
379
- execute: async (inputData, context) => ({ results: [] }),
380
- },
381
- });
382
- ```
383
-
384
- Tool call events from the model are emitted as:
385
-
386
- ```ts
387
- voice.on('toolCall', ({ name, args, id }) => {
388
- // name: string, args: Record<string, any>, id: string
389
- });
390
- ```
391
-
392
- ---
393
-
394
- **`off<E extends VoiceEventType>(event: E, callback: (data: E extends keyof GeminiLiveEventMap ? GeminiLiveEventMap[E] : unknown) => void): void`**
395
-
396
- Removes an event listener.
397
-
398
- **Parameters:**
399
-
400
- - `event: E` - Event name to stop listening to
401
- - `callback: (data) => void` - Specific callback function to remove
402
-
403
- ### Configuration Types
404
-
405
- **`GeminiLiveVoiceConfig`**
406
-
407
- ```typescript
408
- interface GeminiLiveVoiceConfig {
409
- apiKey?: string;
410
- model?: GeminiVoiceModel;
411
- speaker?: GeminiVoiceName;
412
- vertexAI?: boolean;
413
- project?: string;
414
- location?: string;
415
- serviceAccountKeyFile?: string;
416
- serviceAccountEmail?: string;
417
- instructions?: string;
418
- tools?: GeminiToolConfig[];
419
- sessionConfig?: GeminiSessionConfig;
420
- audioConfig?: Partial<AudioConfig>;
421
- debug?: boolean;
422
- }
423
- ```
424
-
425
- **`GeminiLiveVoiceOptions`**
426
-
427
- ```typescript
428
- interface GeminiLiveVoiceOptions {
429
- speaker?: GeminiVoiceName;
430
- languageCode?: string;
431
- responseModalities?: ('AUDIO' | 'TEXT')[];
432
- }
433
- ```
434
-
435
- **`GeminiSessionConfig`**
436
-
437
- ```typescript
438
- interface GeminiSessionConfig {
439
- enableResumption?: boolean;
440
- maxDuration?: string;
441
- contextCompression?: boolean;
442
- vad?: {
443
- enabled?: boolean;
444
- sensitivity?: number;
445
- silenceDurationMs?: number;
446
- };
447
- interrupts?: {
448
- enabled?: boolean;
449
- allowUserInterruption?: boolean;
450
- };
451
- }
452
- ```
453
-
454
- ## Features
455
-
456
- - **Real-time bidirectional audio streaming**
457
- - **Multimodal input support** (audio, video, text)
458
- - **Built-in Voice Activity Detection (VAD)**
459
- - **Interrupt handling** - Natural conversation flow
460
- - **Session management** - Resume conversations after network interruptions
461
- - **Tool calling support** - Integrate with external APIs and functions
462
- - **Live transcription** - Real-time speech-to-text
463
- - **Multiple voice options** - Choose from various voice personalities
464
- - **Multilingual support** - Support for 30+ languages
63
+ ## Documentation
465
64
 
466
- ## Voice Options
65
+ - [@mastra/voice-google-gemini-live documentation](https://mastra.ai/integrations/voice/google)
467
66
 
468
- - **Puck** - Conversational, friendly
469
- - **Charon** - Deep, authoritative
470
- - **Kore** - Neutral, professional
471
- - **Fenrir** - Warm, approachable
67
+ ## Changelog
472
68
 
473
- ## Model Options
69
+ See the [package changelog](https://github.com/mastra-ai/mastra/blob/main/voice/google-gemini-live-api/CHANGELOG.md) for version history and release notes.
474
70
 
475
- - `gemini-2.0-flash-exp` - Default model
476
- - `gemini-2.0-flash-live-001` - Latest production model
477
- - `gemini-2.5-flash-preview-native-audio-dialog` - Preview with native audio
478
- - `gemini-live-2.5-flash-preview` - Half-cascade architecture
71
+ ## Support
479
72
 
480
- For detailed API documentation, visit [Google's Gemini Live API docs](https://ai.google.dev/gemini-api/docs/live).
73
+ We have an [open community Discord](https://discord.gg/mastra-ai). Come and say hello and let us know if you have any questions or need any help getting things running.
@@ -3,7 +3,7 @@ name: mastra-voice-google-gemini-live
3
3
  description: Documentation for @mastra/voice-google-gemini-live. Use when working with @mastra/voice-google-gemini-live APIs, configuration, or implementation.
4
4
  metadata:
5
5
  package: "@mastra/voice-google-gemini-live"
6
- version: "0.14.8-alpha.0"
6
+ version: "0.14.8"
7
7
  ---
8
8
 
9
9
  ## When to use
@@ -1,5 +1,5 @@
1
1
  {
2
- "version": "0.14.8-alpha.0",
2
+ "version": "0.14.8",
3
3
  "package": "@mastra/voice-google-gemini-live",
4
4
  "exports": {},
5
5
  "modules": {}
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@mastra/voice-google-gemini-live",
3
- "version": "0.14.8-alpha.0",
3
+ "version": "0.14.8",
4
4
  "description": "Mastra Google Gemini Live API integration",
5
5
  "type": "module",
6
6
  "files": [
@@ -26,7 +26,7 @@
26
26
  "@google/genai": "^1.52.0",
27
27
  "google-auth-library": "^10.9.1",
28
28
  "ws": "^8.21.0",
29
- "@mastra/schema-compat": "1.3.8-alpha.0"
29
+ "@mastra/schema-compat": "1.3.8"
30
30
  },
31
31
  "devDependencies": {
32
32
  "@types/node": "22.20.1",
@@ -39,10 +39,10 @@
39
39
  "typescript": "^7.0.2",
40
40
  "vitest": "4.1.10",
41
41
  "zod": "^4.4.3",
42
- "@internal/lint": "0.0.129",
43
- "@internal/test-utils": "0.0.65",
44
- "@internal/types-builder": "0.0.104",
45
- "@internal/voice": "0.0.27"
42
+ "@internal/test-utils": "0.0.66",
43
+ "@internal/types-builder": "0.0.105",
44
+ "@internal/voice": "0.0.28",
45
+ "@internal/lint": "0.0.130"
46
46
  },
47
47
  "homepage": "https://mastra.ai",
48
48
  "repository": {