@mastra/voice-aws-nova-sonic 0.2.0 → 0.2.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +16 -0
- package/dist/_types/@internal_voice/dist/_types/@internal_ai-sdk-v5/dist/index.d.ts +110 -221
- package/dist/_types/@internal_voice/dist/_types/@internal_core/dist/base/index.d.ts +27 -27
- package/dist/_types/@internal_voice/dist/_types/@internal_core/dist/index-S1lgaKO7.d.ts +218 -0
- package/dist/_types/@internal_voice/dist/_types/@internal_core/dist/request-context/index.d.ts +122 -95
- package/dist/_types/@internal_voice/dist/_types/@internal_core/dist/types/index.d.ts +4 -2
- package/dist/docs/SKILL.md +6 -6
- package/dist/docs/assets/SOURCE_MAP.json +1 -1
- package/dist/docs/references/{docs-voice-overview.md → guides-voice-overview.md} +90 -134
- package/dist/docs/references/{docs-voice-speech-to-speech.md → guides-voice-speech-to-speech.md} +90 -8
- package/dist/docs/references/{reference-voice-aws-nova-sonic.md → integrations-voice-aws-nova-sonic.md} +4 -2
- package/dist/index.cjs +1679 -1860
- package/dist/index.cjs.map +1 -1
- package/dist/index.js +1678 -1858
- package/dist/index.js.map +1 -1
- package/package.json +17 -16
- package/dist/_types/@internal_voice/dist/_types/@internal_core/dist/logger/index.d.ts +0 -217
package/dist/docs/references/{docs-voice-speech-to-speech.md → guides-voice-speech-to-speech.md}
RENAMED
|
@@ -1,3 +1,5 @@
|
|
|
1
|
+
> Discover all available pages from the documentation index: https://mastra.ai/llms.txt
|
|
2
|
+
|
|
1
3
|
# Speech-to-Speech capabilities in Mastra
|
|
2
4
|
|
|
3
5
|
## Introduction
|
|
@@ -8,7 +10,7 @@ Speech-to-Speech (STS) in Mastra provides a standardized interface for real-time
|
|
|
8
10
|
|
|
9
11
|
- **`apiKey`**: Your OpenAI API key. Falls back to the `OPENAI_API_KEY` environment variable.
|
|
10
12
|
- **`model`**: The model ID to use for real-time voice interactions (e.g., `gpt-5.1-realtime`).
|
|
11
|
-
- **`speaker`**: The default voice ID for speech synthesis.
|
|
13
|
+
- **`speaker`**: The default voice ID for speech synthesis. You can specify which voice to use for the speech output.
|
|
12
14
|
|
|
13
15
|
```typescript
|
|
14
16
|
const voice = new OpenAIRealtimeVoice({
|
|
@@ -32,7 +34,7 @@ const agent = new Agent({
|
|
|
32
34
|
id: 'agent',
|
|
33
35
|
name: 'OpenAI Realtime Agent',
|
|
34
36
|
instructions: `You are a helpful assistant with real-time voice capabilities.`,
|
|
35
|
-
model: 'openai/gpt-5.
|
|
37
|
+
model: 'openai/gpt-5.6-sol',
|
|
36
38
|
voice: new OpenAIRealtimeVoice(),
|
|
37
39
|
})
|
|
38
40
|
|
|
@@ -52,7 +54,87 @@ const micStream = getMicrophoneStream()
|
|
|
52
54
|
await agent.voice.send(micStream)
|
|
53
55
|
```
|
|
54
56
|
|
|
55
|
-
For
|
|
57
|
+
For a broader overview of voice providers on agents, see [Voice in Mastra](https://mastra.ai/guides/voice/overview).
|
|
58
|
+
|
|
59
|
+
## Use tools in realtime sessions
|
|
60
|
+
|
|
61
|
+
Realtime voice providers can use tools configured on the agent. Add the tools to the `Agent` definition, then connect and send audio through the voice provider:
|
|
62
|
+
|
|
63
|
+
```typescript
|
|
64
|
+
import { Agent } from '@mastra/core/agent'
|
|
65
|
+
import { OpenAIRealtimeVoice } from '@mastra/voice-openai-realtime'
|
|
66
|
+
import { calculate, search } from '../tools'
|
|
67
|
+
|
|
68
|
+
export const agent = new Agent({
|
|
69
|
+
id: 'speech-to-speech-agent',
|
|
70
|
+
name: 'Speech-to-Speech Agent',
|
|
71
|
+
instructions: 'You are a helpful assistant with speech-to-speech capabilities.',
|
|
72
|
+
model: 'openai/gpt-5.6-sol',
|
|
73
|
+
tools: {
|
|
74
|
+
search,
|
|
75
|
+
calculate,
|
|
76
|
+
},
|
|
77
|
+
voice: new OpenAIRealtimeVoice(),
|
|
78
|
+
})
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
## Listen for realtime events
|
|
82
|
+
|
|
83
|
+
Realtime voice providers emit events you can use to update your UI, play assistant audio, log transcriptions, and handle errors:
|
|
84
|
+
|
|
85
|
+
```typescript
|
|
86
|
+
agent.voice.on('speaking', ({ audio }) => {
|
|
87
|
+
playAudio(audio)
|
|
88
|
+
})
|
|
89
|
+
|
|
90
|
+
agent.voice.on('writing', ({ text, role }) => {
|
|
91
|
+
console.log(`${role}: ${text}`)
|
|
92
|
+
})
|
|
93
|
+
|
|
94
|
+
agent.voice.on('error', error => {
|
|
95
|
+
console.error('Voice error:', error)
|
|
96
|
+
})
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
Event names and payloads vary by provider. Check the provider section below or the provider reference for the full event list.
|
|
100
|
+
|
|
101
|
+
## Per-session voice instances
|
|
102
|
+
|
|
103
|
+
A static `voice` instance is shared across every request. This works for one-shot text-to-speech, but real-time and speech-to-speech providers store session state such as the WebSocket connection, tools, instructions, and request context. If one agent handles several live sessions at once, a shared instance can let one session overwrite another session's state.
|
|
104
|
+
|
|
105
|
+
Provide `voice` as a resolver when each live session needs its own voice instance. Mastra runs the resolver on each `getVoice()` call and returns a fresh instance for that request context:
|
|
106
|
+
|
|
107
|
+
```typescript
|
|
108
|
+
import { Agent } from '@mastra/core/agent'
|
|
109
|
+
import { RequestContext } from '@mastra/core/request-context'
|
|
110
|
+
import { OpenAIRealtimeVoice } from '@mastra/voice-openai-realtime'
|
|
111
|
+
|
|
112
|
+
export const agent = new Agent({
|
|
113
|
+
id: 'support-line',
|
|
114
|
+
name: 'Support Line',
|
|
115
|
+
instructions: ({ requestContext }) => `Help user ${requestContext.get('user')}.`,
|
|
116
|
+
model: 'openai/gpt-5.6-sol',
|
|
117
|
+
voice: ({ requestContext }) =>
|
|
118
|
+
new OpenAIRealtimeVoice({
|
|
119
|
+
apiKey: requestContext.get('apiKey'),
|
|
120
|
+
}),
|
|
121
|
+
})
|
|
122
|
+
|
|
123
|
+
const requestContext = new RequestContext()
|
|
124
|
+
requestContext.set('user', 'user-123')
|
|
125
|
+
requestContext.set('apiKey', process.env.OPENAI_API_KEY)
|
|
126
|
+
|
|
127
|
+
const voice = await agent.getVoice({ requestContext })
|
|
128
|
+
await voice.connect()
|
|
129
|
+
```
|
|
130
|
+
|
|
131
|
+
When you use a resolver:
|
|
132
|
+
|
|
133
|
+
- Each call to `getVoice()` returns a new instance, so concurrent sessions don't share state.
|
|
134
|
+
- Mastra doesn't add tools or instructions to a resolver instance. Configure them inside the resolver or on the provider.
|
|
135
|
+
- You own the returned instance lifecycle, so call `disconnect()` or `close()` when the session ends.
|
|
136
|
+
|
|
137
|
+
The `agent.voice` getter has no request context, so it throws when `voice` is a resolver. Use `agent.getVoice({ requestContext })` instead.
|
|
56
138
|
|
|
57
139
|
## Google Gemini Live (Realtime)
|
|
58
140
|
|
|
@@ -66,7 +148,7 @@ const agent = new Agent({
|
|
|
66
148
|
name: 'Gemini Live Agent',
|
|
67
149
|
instructions: 'You are a helpful assistant with real-time voice capabilities.',
|
|
68
150
|
// Model used for text generation; voice provider handles realtime audio
|
|
69
|
-
model: 'openai/gpt-5.
|
|
151
|
+
model: 'openai/gpt-5.6-sol',
|
|
70
152
|
voice: new GeminiLiveVoice({
|
|
71
153
|
apiKey: process.env.GOOGLE_API_KEY,
|
|
72
154
|
model: 'gemini-2.0-flash-exp',
|
|
@@ -113,7 +195,7 @@ const agent = new Agent({
|
|
|
113
195
|
name: 'Nova Sonic Agent',
|
|
114
196
|
instructions: 'You are a helpful assistant with real-time voice capabilities.',
|
|
115
197
|
// Model used for text generation; voice provider handles realtime audio
|
|
116
|
-
model: 'openai/gpt-5.
|
|
198
|
+
model: 'openai/gpt-5.6-sol',
|
|
117
199
|
voice: new NovaSonicVoice({
|
|
118
200
|
region: 'us-east-1',
|
|
119
201
|
speaker: 'matthew',
|
|
@@ -157,7 +239,7 @@ const agent = new Agent({
|
|
|
157
239
|
name: 'Inworld Realtime Agent',
|
|
158
240
|
instructions: 'You are a helpful assistant with real-time voice capabilities.',
|
|
159
241
|
// Model used for text generation; voice provider handles realtime audio
|
|
160
|
-
model: 'openai/gpt-5.
|
|
242
|
+
model: 'openai/gpt-5.6-sol',
|
|
161
243
|
voice: new InworldRealtimeVoice({
|
|
162
244
|
apiKey: process.env.INWORLD_API_KEY,
|
|
163
245
|
model: 'inworld/models/gemma-4-26b-a4b-it',
|
|
@@ -190,8 +272,8 @@ await agent.voice.send(micStream)
|
|
|
190
272
|
|
|
191
273
|
Note:
|
|
192
274
|
|
|
193
|
-
- Requires `INWORLD_API_KEY`. Inworld API keys
|
|
275
|
+
- Requires `INWORLD_API_KEY`. Inworld API keys are pre-Basic-encoded: paste them verbatim.
|
|
194
276
|
- The WebSocket URL appends a client-generated `?key=...&protocol=realtime`. The model is configured via the initial `session.update`, not in the URL.
|
|
195
277
|
- Inworld's wire protocol is the OpenAI Realtime GA spec, so event names match `@mastra/voice-openai-realtime`.
|
|
196
|
-
- Typed Inworld realtime knobs (MCP tool routing, semantic VAD eagerness, playback speed, transcription model, output modalities, …) are exposed through the `session` constructor field
|
|
278
|
+
- Typed Inworld realtime knobs (MCP tool routing, semantic VAD eagerness, playback speed, transcription model, output modalities, …) are exposed through the `session` constructor field. An untyped `providerData` escape hatch is also deep-merged for forward compatibility with new Inworld features.
|
|
197
279
|
- Events: `speaker` (PCM audio stream), `speaking` (audio Buffer per delta), `writing` (text), `conversation.item.added`, `conversation.item.done`, `function_call.arguments`, `tool-call-start`, `tool-call-result`, and `error`.
|
|
@@ -1,4 +1,6 @@
|
|
|
1
|
-
|
|
1
|
+
> Discover all available pages from the documentation index: https://mastra.ai/llms.txt
|
|
2
|
+
|
|
3
|
+
# AWS Nova Sonic
|
|
2
4
|
|
|
3
5
|
The `NovaSonicVoice` class provides real-time speech-to-speech capabilities backed by [AWS Bedrock Nova 2 Sonic](https://docs.aws.amazon.com/nova/latest/userguide/speech.html). It opens a bidirectional stream to the model and emits events for assistant audio, transcribed text, tool calls, turn boundaries, and interruptions.
|
|
4
6
|
|
|
@@ -216,7 +218,7 @@ Registers and removes event listeners. See [Voice events](https://mastra.ai/refe
|
|
|
216
218
|
|
|
217
219
|
## Available voices
|
|
218
220
|
|
|
219
|
-
Nova 2 Sonic
|
|
221
|
+
Nova 2 Sonic provides voices in ten locales. Tiffany and Matthew are polyglot and can speak any supported language.
|
|
220
222
|
|
|
221
223
|
| Voice ID | Name | Language | Locale | Gender | Polyglot |
|
|
222
224
|
| ---------- | -------- | ---------- | ------ | --------- | -------- |
|