@mastra/voice-aws-nova-sonic 0.2.0 → 0.2.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,3 +1,5 @@
1
+ > Discover all available pages from the documentation index: https://mastra.ai/llms.txt
2
+
1
3
  # Speech-to-Speech capabilities in Mastra
2
4
 
3
5
  ## Introduction
@@ -8,7 +10,7 @@ Speech-to-Speech (STS) in Mastra provides a standardized interface for real-time
8
10
 
9
11
  - **`apiKey`**: Your OpenAI API key. Falls back to the `OPENAI_API_KEY` environment variable.
10
12
  - **`model`**: The model ID to use for real-time voice interactions (e.g., `gpt-5.1-realtime`).
11
- - **`speaker`**: The default voice ID for speech synthesis. This allows you to specify which voice to use for the speech output.
13
+ - **`speaker`**: The default voice ID for speech synthesis. You can specify which voice to use for the speech output.
12
14
 
13
15
  ```typescript
14
16
  const voice = new OpenAIRealtimeVoice({
@@ -32,7 +34,7 @@ const agent = new Agent({
32
34
  id: 'agent',
33
35
  name: 'OpenAI Realtime Agent',
34
36
  instructions: `You are a helpful assistant with real-time voice capabilities.`,
35
- model: 'openai/gpt-5.5',
37
+ model: 'openai/gpt-5.6-sol',
36
38
  voice: new OpenAIRealtimeVoice(),
37
39
  })
38
40
 
@@ -52,7 +54,87 @@ const micStream = getMicrophoneStream()
52
54
  await agent.voice.send(micStream)
53
55
  ```
54
56
 
55
- For integrating Speech-to-Speech capabilities with agents, refer to the [Adding Voice to Agents](https://mastra.ai/docs/agents/adding-voice) documentation.
57
+ For a broader overview of voice providers on agents, see [Voice in Mastra](https://mastra.ai/guides/voice/overview).
58
+
59
+ ## Use tools in realtime sessions
60
+
61
+ Realtime voice providers can use tools configured on the agent. Add the tools to the `Agent` definition, then connect and send audio through the voice provider:
62
+
63
+ ```typescript
64
+ import { Agent } from '@mastra/core/agent'
65
+ import { OpenAIRealtimeVoice } from '@mastra/voice-openai-realtime'
66
+ import { calculate, search } from '../tools'
67
+
68
+ export const agent = new Agent({
69
+ id: 'speech-to-speech-agent',
70
+ name: 'Speech-to-Speech Agent',
71
+ instructions: 'You are a helpful assistant with speech-to-speech capabilities.',
72
+ model: 'openai/gpt-5.6-sol',
73
+ tools: {
74
+ search,
75
+ calculate,
76
+ },
77
+ voice: new OpenAIRealtimeVoice(),
78
+ })
79
+ ```
80
+
81
+ ## Listen for realtime events
82
+
83
+ Realtime voice providers emit events you can use to update your UI, play assistant audio, log transcriptions, and handle errors:
84
+
85
+ ```typescript
86
+ agent.voice.on('speaking', ({ audio }) => {
87
+ playAudio(audio)
88
+ })
89
+
90
+ agent.voice.on('writing', ({ text, role }) => {
91
+ console.log(`${role}: ${text}`)
92
+ })
93
+
94
+ agent.voice.on('error', error => {
95
+ console.error('Voice error:', error)
96
+ })
97
+ ```
98
+
99
+ Event names and payloads vary by provider. Check the provider section below or the provider reference for the full event list.
100
+
101
+ ## Per-session voice instances
102
+
103
+ A static `voice` instance is shared across every request. This works for one-shot text-to-speech, but real-time and speech-to-speech providers store session state such as the WebSocket connection, tools, instructions, and request context. If one agent handles several live sessions at once, a shared instance can let one session overwrite another session's state.
104
+
105
+ Provide `voice` as a resolver when each live session needs its own voice instance. Mastra runs the resolver on each `getVoice()` call and returns a fresh instance for that request context:
106
+
107
+ ```typescript
108
+ import { Agent } from '@mastra/core/agent'
109
+ import { RequestContext } from '@mastra/core/request-context'
110
+ import { OpenAIRealtimeVoice } from '@mastra/voice-openai-realtime'
111
+
112
+ export const agent = new Agent({
113
+ id: 'support-line',
114
+ name: 'Support Line',
115
+ instructions: ({ requestContext }) => `Help user ${requestContext.get('user')}.`,
116
+ model: 'openai/gpt-5.6-sol',
117
+ voice: ({ requestContext }) =>
118
+ new OpenAIRealtimeVoice({
119
+ apiKey: requestContext.get('apiKey'),
120
+ }),
121
+ })
122
+
123
+ const requestContext = new RequestContext()
124
+ requestContext.set('user', 'user-123')
125
+ requestContext.set('apiKey', process.env.OPENAI_API_KEY)
126
+
127
+ const voice = await agent.getVoice({ requestContext })
128
+ await voice.connect()
129
+ ```
130
+
131
+ When you use a resolver:
132
+
133
+ - Each call to `getVoice()` returns a new instance, so concurrent sessions don't share state.
134
+ - Mastra doesn't add tools or instructions to a resolver instance. Configure them inside the resolver or on the provider.
135
+ - You own the returned instance lifecycle, so call `disconnect()` or `close()` when the session ends.
136
+
137
+ The `agent.voice` getter has no request context, so it throws when `voice` is a resolver. Use `agent.getVoice({ requestContext })` instead.
56
138
 
57
139
  ## Google Gemini Live (Realtime)
58
140
 
@@ -66,7 +148,7 @@ const agent = new Agent({
66
148
  name: 'Gemini Live Agent',
67
149
  instructions: 'You are a helpful assistant with real-time voice capabilities.',
68
150
  // Model used for text generation; voice provider handles realtime audio
69
- model: 'openai/gpt-5.5',
151
+ model: 'openai/gpt-5.6-sol',
70
152
  voice: new GeminiLiveVoice({
71
153
  apiKey: process.env.GOOGLE_API_KEY,
72
154
  model: 'gemini-2.0-flash-exp',
@@ -113,7 +195,7 @@ const agent = new Agent({
113
195
  name: 'Nova Sonic Agent',
114
196
  instructions: 'You are a helpful assistant with real-time voice capabilities.',
115
197
  // Model used for text generation; voice provider handles realtime audio
116
- model: 'openai/gpt-5.5',
198
+ model: 'openai/gpt-5.6-sol',
117
199
  voice: new NovaSonicVoice({
118
200
  region: 'us-east-1',
119
201
  speaker: 'matthew',
@@ -157,7 +239,7 @@ const agent = new Agent({
157
239
  name: 'Inworld Realtime Agent',
158
240
  instructions: 'You are a helpful assistant with real-time voice capabilities.',
159
241
  // Model used for text generation; voice provider handles realtime audio
160
- model: 'openai/gpt-5.5',
242
+ model: 'openai/gpt-5.6-sol',
161
243
  voice: new InworldRealtimeVoice({
162
244
  apiKey: process.env.INWORLD_API_KEY,
163
245
  model: 'inworld/models/gemma-4-26b-a4b-it',
@@ -190,8 +272,8 @@ await agent.voice.send(micStream)
190
272
 
191
273
  Note:
192
274
 
193
- - Requires `INWORLD_API_KEY`. Inworld API keys ship pre-Basic-encoded — paste them verbatim.
275
+ - Requires `INWORLD_API_KEY`. Inworld API keys are pre-Basic-encoded: paste them verbatim.
194
276
  - The WebSocket URL appends a client-generated `?key=...&protocol=realtime`. The model is configured via the initial `session.update`, not in the URL.
195
277
  - Inworld's wire protocol is the OpenAI Realtime GA spec, so event names match `@mastra/voice-openai-realtime`.
196
- - Typed Inworld realtime knobs (MCP tool routing, semantic VAD eagerness, playback speed, transcription model, output modalities, …) are exposed through the `session` constructor field; an untyped `providerData` escape hatch is also deep-merged for forward compatibility with new Inworld features.
278
+ - Typed Inworld realtime knobs (MCP tool routing, semantic VAD eagerness, playback speed, transcription model, output modalities, …) are exposed through the `session` constructor field. An untyped `providerData` escape hatch is also deep-merged for forward compatibility with new Inworld features.
197
279
  - Events: `speaker` (PCM audio stream), `speaking` (audio Buffer per delta), `writing` (text), `conversation.item.added`, `conversation.item.done`, `function_call.arguments`, `tool-call-start`, `tool-call-result`, and `error`.
@@ -1,4 +1,6 @@
1
- # AWS Nova Sonic voice
1
+ > Discover all available pages from the documentation index: https://mastra.ai/llms.txt
2
+
3
+ # AWS Nova Sonic
2
4
 
3
5
  The `NovaSonicVoice` class provides real-time speech-to-speech capabilities backed by [AWS Bedrock Nova 2 Sonic](https://docs.aws.amazon.com/nova/latest/userguide/speech.html). It opens a bidirectional stream to the model and emits events for assistant audio, transcribed text, tool calls, turn boundaries, and interruptions.
4
6
 
@@ -216,7 +218,7 @@ Registers and removes event listeners. See [Voice events](https://mastra.ai/refe
216
218
 
217
219
  ## Available voices
218
220
 
219
- Nova 2 Sonic ships voices in ten locales. Tiffany and Matthew are polyglot and can speak any supported language.
221
+ Nova 2 Sonic provides voices in ten locales. Tiffany and Matthew are polyglot and can speak any supported language.
220
222
 
221
223
  | Voice ID | Name | Language | Locale | Gender | Polyglot |
222
224
  | ---------- | -------- | ---------- | ------ | --------- | -------- |