@mastra/voice-aws-nova-sonic 0.2.2-alpha.0 → 0.2.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -2,97 +2,10 @@
2
2
 
3
3
  Mastra integration for AWS Nova 2 Sonic, providing real-time bidirectional speech-to-speech capabilities using Amazon Bedrock's bidirectional streaming API.
4
4
 
5
- ## Features
6
-
7
- - **Real-time bidirectional streaming**: Continuous audio streaming in both directions
8
- - **Multilingual support**: Supports English, French, Italian, German, Spanish, Portuguese, and Hindi
9
- - **Polyglot voices**: Voices that can speak multiple languages within the same session
10
- - **Barge-in support**: Users can interrupt the assistant mid-speech; handled server-side by Nova Sonic
11
- - **Tool/function calling**: Support for agentic workflows and async tool execution
12
- - **Cross-modal input**: Support for both audio and text inputs in the same conversation
13
- - **Natural turn-taking**: Intelligent voice activity detection and turn management
14
- - **Robust error handling**: Comprehensive error handling with detailed error codes
15
-
16
5
  ## Installation
17
6
 
18
7
  ```bash
19
8
  npm install @mastra/voice-aws-nova-sonic
20
- # or
21
- pnpm add @mastra/voice-aws-nova-sonic
22
- # or
23
- yarn add @mastra/voice-aws-nova-sonic
24
- ```
25
-
26
- ## Prerequisites
27
-
28
- - Node.js >= 22.13.0
29
- - AWS account with access to Amazon Bedrock
30
- - AWS credentials configured (see [AWS Setup](#aws-setup))
31
- - Access to Nova 2 Sonic model in your AWS region
32
-
33
- ## AWS Setup
34
-
35
- ### 1. Enable Nova 2 Sonic in Amazon Bedrock
36
-
37
- 1. Go to the [Amazon Bedrock Console](https://console.aws.amazon.com/bedrock/)
38
- 2. Navigate to "Model access" in the left sidebar
39
- 3. Request access to "Amazon Nova 2 Sonic" model
40
- 4. Wait for approval (usually instant)
41
-
42
- ### 2. Configure AWS Credentials
43
-
44
- You can configure AWS credentials in several ways:
45
-
46
- **Option 1: Environment Variables**
47
-
48
- ```bash
49
- export AWS_ACCESS_KEY_ID=your-access-key-id
50
- export AWS_SECRET_ACCESS_KEY=your-secret-access-key
51
- export AWS_REGION=us-east-1
52
- ```
53
-
54
- **Option 2: AWS Credentials File**
55
-
56
- ```ini
57
- # ~/.aws/credentials
58
- [default]
59
- aws_access_key_id = your-access-key-id
60
- aws_secret_access_key = your-secret-access-key
61
- ```
62
-
63
- **Option 3: IAM Role** (for EC2/Lambda)
64
-
65
- - Attach an IAM role with Bedrock permissions to your EC2 instance or Lambda function
66
-
67
- **Option 4: Explicit Credentials in Code**
68
-
69
- ```typescript
70
- import { NovaSonicVoice } from '@mastra/voice-aws-nova-sonic';
71
-
72
- const voice = new NovaSonicVoice({
73
- region: 'us-east-1',
74
- credentials: {
75
- accessKeyId: 'your-access-key-id',
76
- secretAccessKey: 'your-secret-access-key',
77
- },
78
- });
79
- ```
80
-
81
- ### 3. IAM Permissions
82
-
83
- Your AWS credentials need the following IAM permissions:
84
-
85
- ```json
86
- {
87
- "Version": "2012-10-17",
88
- "Statement": [
89
- {
90
- "Effect": "Allow",
91
- "Action": ["bedrock:InvokeModel", "bedrock:InvokeModelWithBidirectionalStream"],
92
- "Resource": "arn:aws:bedrock:*::foundation-model/amazon.nova-2-sonic-v1:0"
93
- }
94
- ]
95
- }
96
9
  ```
97
10
 
98
11
  ## Usage
@@ -132,253 +45,14 @@ agent.voice.on('writing', ({ text, role, generationStage }) => {
132
45
  await agent.voice.send(microphoneStream);
133
46
  ```
134
47
 
135
- ### Advanced Configuration
136
-
137
- ```typescript
138
- import { NovaSonicVoice } from '@mastra/voice-aws-nova-sonic';
139
-
140
- const voice = new NovaSonicVoice({
141
- region: 'us-east-1', // or 'us-west-2', 'ap-northeast-1'
142
- model: 'amazon.nova-2-sonic-v1:0',
143
- speaker: 'matthew', // or 'tiffany', 'amy', etc.
144
- languageCode: 'en-US',
145
- instructions: 'You are a helpful assistant.',
146
- sessionConfig: {
147
- tools: [
148
- {
149
- name: 'search',
150
- description: 'Search the web',
151
- inputSchema: {
152
- type: 'object',
153
- properties: {
154
- query: { type: 'string' },
155
- },
156
- required: ['query'],
157
- },
158
- },
159
- ],
160
- turnDetectionConfiguration: {
161
- // HIGH = fastest (1.5s pause), MEDIUM = balanced (1.75s), LOW = slowest (2s)
162
- endpointingSensitivity: 'MEDIUM',
163
- },
164
- },
165
- debug: true,
166
- });
167
-
168
- await voice.connect();
169
- ```
170
-
171
- ### With Tools
172
-
173
- ```typescript
174
- import { Agent } from '@mastra/core/agent';
175
- import { NovaSonicVoice } from '@mastra/voice-aws-nova-sonic';
176
- import { createTool } from '@mastra/core/tools';
177
- import { z } from 'zod';
178
-
179
- const weatherTool = createTool({
180
- id: 'weather',
181
- description: 'Get weather information',
182
- inputSchema: z.object({
183
- location: z.string(),
184
- }),
185
- execute: async ({ context }) => {
186
- // Fetch weather data
187
- return { temperature: 72, condition: 'sunny' };
188
- },
189
- });
190
-
191
- const agent = new Agent({
192
- name: 'Weather Agent',
193
- instructions: 'You help users get weather information.',
194
- model: 'openai/gpt-4o',
195
- tools: {
196
- weather: weatherTool,
197
- },
198
- voice: new NovaSonicVoice({
199
- region: 'us-east-1',
200
- }),
201
- });
202
-
203
- await agent.voice.connect();
204
- // Tools are automatically available to the voice model
205
- ```
206
-
207
- ### Cross-Modal Text Input
208
-
209
- Send text messages during an active voice session:
210
-
211
- ```typescript
212
- // After connecting and starting audio streaming
213
- await agent.voice.speak('What is the weather in New York?');
214
- ```
215
-
216
- ## API Reference
217
-
218
- ### Constructor
219
-
220
- ```typescript
221
- new NovaSonicVoice(config?: NovaSonicVoiceConfig)
222
- ```
223
-
224
- **Configuration Options:**
225
-
226
- - `region` (string, optional): AWS region. Default: `'us-east-1'`. Supported: `'us-east-1'`, `'us-west-2'`, `'ap-northeast-1'`
227
- - `model` (string, optional): Model ID. Default: `'amazon.nova-2-sonic-v1:0'`
228
- - `credentials` (Credentials, optional): AWS credentials. If not provided, uses default credential chain
229
- - `speaker` (string, optional): Voice name/identifier (e.g., `'matthew'`, `'tiffany'`, `'amy'`)
230
- - `languageCode` (string, optional): Language code (e.g., `'en-US'`, `'fr-FR'`)
231
- - `instructions` (string, optional): System instructions for the model
232
- - `tools` (array, optional): Tool definitions
233
- - `sessionConfig` (object, optional): Session configuration including `turnDetectionConfiguration`, `tools`, `inferenceConfiguration`
234
- - `debug` (boolean, optional): Enable debug logging. Default: `false`
235
-
236
- ### Methods
237
-
238
- #### `connect(options?)`
239
-
240
- Establishes connection to AWS Bedrock. Must be called before using other methods.
241
-
242
- ```typescript
243
- await voice.connect();
244
- ```
245
-
246
- #### `speak(input, options?)`
247
-
248
- Send cross-modal text input during an active voice session. Nova Sonic processes it and responds with audio.
249
-
250
- ```typescript
251
- await voice.speak('Hello, world!');
252
- ```
253
-
254
- #### `listen(audioStream, options?)`
255
-
256
- Stream audio input for transcription. For Nova Sonic, this is equivalent to `send()`.
257
-
258
- ```typescript
259
- await voice.listen(audioStream);
260
- ```
261
-
262
- #### `send(audioData)`
263
-
264
- Stream audio data in real-time. Accepts a `NodeJS.ReadableStream` (PCM16 audio) or an `Int16Array`.
265
-
266
- ```typescript
267
- // Stream from a ReadableStream
268
- await voice.send(audioStream);
269
-
270
- // Or with Int16Array
271
- const audioArray = new Int16Array([...]);
272
- await voice.send(audioArray);
273
- ```
274
-
275
- #### `close()`
276
-
277
- Disconnect and cleanup resources.
278
-
279
- ```typescript
280
- voice.close();
281
- ```
282
-
283
- #### `on(event, callback)`
284
-
285
- Register an event listener.
286
-
287
- ```typescript
288
- voice.on('speaking', ({ audio }) => {
289
- // audio is a base64-encoded string of PCM audio
290
- });
291
-
292
- voice.on('writing', ({ text, role, generationStage }) => {
293
- // generationStage: 'SPECULATIVE' (preview) or 'FINAL' (actual transcript)
294
- console.log(`${role}: ${text}`);
295
- });
296
-
297
- voice.on('error', ({ message, code }) => {
298
- console.error(`Error: ${message} (${code})`);
299
- });
300
- ```
301
-
302
- #### `off(event, callback)`
303
-
304
- Remove an event listener.
305
-
306
- ```typescript
307
- voice.off('speaking', callback);
308
- ```
309
-
310
- ### Events
311
-
312
- - **`speaker`**: Audio stream (`NodeJS.ReadableStream`) for the full response
313
- - **`speaking`**: Audio chunk `{ audio: string, audioData: Buffer, response_id?: string }`
314
- - **`writing`**: Text transcription `{ text: string, role: 'assistant' | 'user', generationStage?: 'SPECULATIVE' | 'FINAL' }`
315
- - **`error`**: Error event `{ message: string, code?: string, details?: unknown }`
316
- - **`toolCall`**: Tool invocation `{ name: string, args: Record<string, any>, id: string }`
317
- - **`turnComplete`**: Turn completion `{ timestamp: number }`
318
- - **`interrupt`**: Barge-in detected `{ type: string, timestamp: number }`
319
- - **`contentStart`**: Content block started (raw Nova Sonic event)
320
- - **`contentEnd`**: Content block ended (raw Nova Sonic event)
321
- - **`usage`**: Token usage `{ inputTokens: number, outputTokens: number, totalTokens: number }`
322
-
323
- ## Supported Regions
324
-
325
- - `us-east-1` (US East - N. Virginia)
326
- - `us-west-2` (US West - Oregon)
327
- - `ap-northeast-1` (Asia Pacific - Tokyo)
328
-
329
- ## Supported Languages
330
-
331
- - English (US, UK, India, Australia)
332
- - French
333
- - Italian
334
- - German
335
- - Spanish
336
- - Portuguese
337
- - Hindi
338
-
339
- ## Error Handling
340
-
341
- The package provides error handling with specific error codes:
342
-
343
- ```typescript
344
- import { NovaSonicError, NovaSonicErrorCode } from '@mastra/voice-aws-nova-sonic';
345
-
346
- voice.on('error', ({ message, code, details }) => {
347
- if (code === NovaSonicErrorCode.CONNECTION_FAILED) {
348
- // Handle connection error
349
- } else if (code === NovaSonicErrorCode.CREDENTIALS_MISSING) {
350
- // Handle credentials error
351
- }
352
- });
353
- ```
354
-
355
- ## Troubleshooting
356
-
357
- ### Connection Issues
358
-
359
- - Verify AWS credentials are configured correctly
360
- - Check that Nova 2 Sonic is enabled in your AWS Bedrock console
361
- - Ensure your IAM role/user has the required permissions
362
- - Verify the region supports Nova 2 Sonic
363
-
364
- ### Audio Issues
365
-
366
- - Ensure audio format is compatible (PCM, 16-bit, 16kHz)
367
- - Check sample rate matches expected format
368
- - Verify audio stream is not empty
369
-
370
- ### Authentication Issues
48
+ ## Documentation
371
49
 
372
- - Check AWS credentials are valid
373
- - Verify IAM permissions include Bedrock access
374
- - Ensure region is correct
50
+ - [@mastra/voice-aws-nova-sonic documentation](https://mastra.ai/integrations/voice/aws-nova-sonic)
375
51
 
376
- ## License
52
+ ## Changelog
377
53
 
378
- Apache-2.0
54
+ See the [package changelog](https://github.com/mastra-ai/mastra/blob/main/voice/aws-nova-sonic/CHANGELOG.md) for version history and release notes.
379
55
 
380
- ## Links
56
+ ## Support
381
57
 
382
- - [Mastra Documentation](https://mastra.ai)
383
- - [AWS Nova 2 Sonic Documentation](https://docs.aws.amazon.com/nova/latest/nova2-userguide/using-conversational-speech.html)
384
- - [Amazon Bedrock Documentation](https://docs.aws.amazon.com/bedrock/)
58
+ We have an [open community Discord](https://discord.gg/mastra-ai). Come and say hello and let us know if you have any questions or need any help getting things running.
@@ -3,7 +3,7 @@ name: mastra-voice-aws-nova-sonic
3
3
  description: Documentation for @mastra/voice-aws-nova-sonic. Use when working with @mastra/voice-aws-nova-sonic APIs, configuration, or implementation.
4
4
  metadata:
5
5
  package: "@mastra/voice-aws-nova-sonic"
6
- version: "0.2.2-alpha.0"
6
+ version: "0.2.2"
7
7
  ---
8
8
 
9
9
  ## When to use
@@ -1,5 +1,5 @@
1
1
  {
2
- "version": "0.2.2-alpha.0",
2
+ "version": "0.2.2",
3
3
  "package": "@mastra/voice-aws-nova-sonic",
4
4
  "exports": {},
5
5
  "modules": {}
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@mastra/voice-aws-nova-sonic",
3
- "version": "0.2.2-alpha.0",
3
+ "version": "0.2.2",
4
4
  "description": "Mastra AWS Nova 2 Sonic voice integration",
5
5
  "type": "module",
6
6
  "main": "dist/index.js",
@@ -38,9 +38,9 @@
38
38
  "typescript": "^7.0.2",
39
39
  "vitest": "4.1.10",
40
40
  "zod": "^4.4.3",
41
- "@internal/types-builder": "0.0.104",
42
- "@internal/lint": "0.0.129",
43
- "@internal/voice": "0.0.27"
41
+ "@internal/lint": "0.0.130",
42
+ "@internal/types-builder": "0.0.105",
43
+ "@internal/voice": "0.0.28"
44
44
  },
45
45
  "peerDependencies": {
46
46
  "zod": "^3.25.0 || ^4.0.0"