@mastra/livekit 0.3.1-alpha.1 → 0.3.1-alpha.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +19 -419
- package/package.json +3 -3
package/README.md
CHANGED
|
@@ -3,446 +3,46 @@
|
|
|
3
3
|
Realtime voice for [Mastra](https://mastra.ai) agents and workflows, powered by [LiveKit Agents](https://docs.livekit.io/agents/).
|
|
4
4
|
|
|
5
5
|
LiveKit's agents framework owns the **audio loop** — WebRTC transport, voice activity detection (VAD), streaming speech-to-text (STT), semantic turn detection, barge-in, and text-to-speech (TTS). This package bridges **reply generation** to Mastra, so each detected user turn is answered by a Mastra **agent** (`agent.stream()`) or **workflow** — with your tools, memory, processors, and model routing all running inside Mastra.
|
|
6
|
-
|
|
7
|
-
```
|
|
8
|
-
caller speaks ─▶ VAD ─▶ STT ─▶ turn detection ─▶ [ Mastra agent / workflow ] ─▶ TTS ─▶ caller hears
|
|
9
|
-
(LiveKit owns the audio loop) (this package bridges replies)
|
|
10
|
-
```
|
|
11
|
-
|
|
12
|
-
## What's in the box
|
|
13
|
-
|
|
14
|
-
- **Two reply paths** — answer turns with a Mastra **agent** (the default, richest path) or a Mastra **workflow** (run-to-completion per turn, e.g. deterministic intent routing). A low-level `generate` escape hatch accepts any custom reply generator.
|
|
15
|
-
- **Full speech stack, pluggable** — STT/TTS as LiveKit inference model strings (`'deepgram/nova-3'`, `'cartesia/sonic-3'`) or your own plugin instances; Silero VAD and LiveKit multilingual/English turn detection; barge-in cancels in-flight generation automatically.
|
|
16
|
-
- **Memory, scoped to the call** — `thread` = call, `resource` = caller, so a returning caller is recognized across calls. Up-front thread creation and greeting persistence keep the saved thread a faithful transcript. Works on the agent path and the workflow path (via `memoryInstance`).
|
|
17
|
-
- **Lifecycle hooks** — `toolFeedback` (speak filler while a tool runs), `onTurnComplete` (post-turn, fire-and-forget, off the audio path), and `onCallEnd` (end-of-call, awaited within LiveKit's shutdown window — the place to summarize the finished call with `memory.summarizeThread()`).
|
|
18
|
-
- **Compliance controls, grouped under `configuration`** — AI-disclosure greeting that can't be barged over (with a per-tenant resolver), periodic re-disclosure on long calls (`repeatEvery`), an extensible named consent model (`consentPolicy` declared on the worker, captured at runtime with `createConsentTool`), and agent-initiated hang-up (`endCall` + `createEndCallTool`) that waits for the goodbye to play out before disconnecting.
|
|
19
|
-
- **Observability** — one `voice call` trace per session with LiveKit pipeline metrics and every Mastra run nested under it.
|
|
20
|
-
- **Connection + dispatch helpers** — `liveKitConnectionRoute` mints tokens and dispatches the worker so a frontend can join.
|
|
21
|
-
|
|
22
6
|
## Installation
|
|
23
7
|
|
|
24
8
|
```bash
|
|
25
|
-
npm install @mastra/livekit
|
|
9
|
+
npm install @mastra/livekit
|
|
26
10
|
```
|
|
27
11
|
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
| Package | Needed for |
|
|
31
|
-
| -------------------------------- | -------------------------------------------- |
|
|
32
|
-
| `@mastra/core` | the Mastra agent/workflow you bridge to |
|
|
33
|
-
| `@livekit/agents` | the audio loop runtime |
|
|
34
|
-
| `@livekit/agents-plugin-silero` | the default `vad: 'silero'` |
|
|
35
|
-
| `@livekit/agents-plugin-livekit` | `turnDetection: 'multilingual' \| 'english'` |
|
|
36
|
-
|
|
37
|
-
The package has three entry points:
|
|
38
|
-
|
|
39
|
-
- `@mastra/livekit` — server-side helpers (`liveKitConnectionRoute`, `dispatchVoiceSession`, `pipeAgentReplyToWriter`) and the agent tool factories (`createConsentTool`, `createEndCallTool` — they go on agents defined in server/shared code). Safe to import from Mastra server code; never loads the `@livekit/agents` runtime.
|
|
40
|
-
- `@mastra/livekit/worker` — the worker runtime (`createLiveKitWorker`, `runLiveKitWorker`). Import it only from the worker entry file.
|
|
41
|
-
- `@mastra/livekit/plugin` — the `MastraLLM` plugin, for customers who own their `voice.AgentSession` and want a Mastra agent in the `llm` slot (a standard LiveKit `llm.LLM`; in-process agent, remote Mastra server, or custom generator). Also loads the `@livekit/agents` runtime — keep it out of Mastra server code.
|
|
42
|
-
|
|
43
|
-
## Quick start
|
|
44
|
-
|
|
45
|
-
A worker is a standalone Node process that connects to LiveKit and answers sessions. Define it with `createLiveKitWorker` and run it with `runLiveKitWorker`:
|
|
46
|
-
|
|
47
|
-
```typescript
|
|
48
|
-
// src/mastra/voice-worker.ts
|
|
49
|
-
import { fileURLToPath } from 'node:url';
|
|
50
|
-
import { createLiveKitWorker, runLiveKitWorker } from '@mastra/livekit/worker';
|
|
51
|
-
import { mastra } from './index';
|
|
52
|
-
|
|
53
|
-
export default createLiveKitWorker({
|
|
54
|
-
mastra,
|
|
55
|
-
agent: 'support', // a Mastra agent key/id (or a resolver, or use `workflow` instead)
|
|
56
|
-
stt: 'deepgram/nova-3',
|
|
57
|
-
tts: 'cartesia/sonic-3',
|
|
58
|
-
turnDetection: 'multilingual',
|
|
59
|
-
configuration: {
|
|
60
|
-
greeting: { text: 'Thanks for calling. How can I help?' },
|
|
61
|
-
},
|
|
62
|
-
});
|
|
63
|
-
|
|
64
|
-
if (process.argv[1] === fileURLToPath(import.meta.url)) {
|
|
65
|
-
runLiveKitWorker({ entry: import.meta.url, agentName: 'mastra-voice' });
|
|
66
|
-
}
|
|
67
|
-
```
|
|
12
|
+
## Usage
|
|
68
13
|
|
|
69
|
-
Add a connection
|
|
14
|
+
Set `LIVEKIT_URL`, `LIVEKIT_API_KEY`, and `LIVEKIT_API_SECRET`. Add a connection route to your Mastra server so clients can receive a room token and dispatch the configured agent.
|
|
70
15
|
|
|
71
16
|
```typescript
|
|
72
|
-
|
|
17
|
+
import { Agent } from '@mastra/core/agent';
|
|
73
18
|
import { Mastra } from '@mastra/core/mastra';
|
|
74
19
|
import { liveKitConnectionRoute } from '@mastra/livekit';
|
|
75
20
|
|
|
21
|
+
const supportAgent = new Agent({
|
|
22
|
+
id: 'support',
|
|
23
|
+
name: 'Support agent',
|
|
24
|
+
instructions: 'Keep voice responses short and conversational.',
|
|
25
|
+
model: 'openai/gpt-5-mini',
|
|
26
|
+
});
|
|
27
|
+
|
|
76
28
|
export const mastra = new Mastra({
|
|
29
|
+
agents: { supportAgent },
|
|
77
30
|
server: {
|
|
78
31
|
apiRoutes: [liveKitConnectionRoute({ agentName: 'mastra-voice' })],
|
|
79
32
|
},
|
|
80
33
|
});
|
|
81
34
|
```
|
|
82
35
|
|
|
83
|
-
Run the worker
|
|
84
|
-
|
|
85
|
-
```bash
|
|
86
|
-
npx livekit-agents download-files # one-time: turn-detection + VAD model files
|
|
87
|
-
npx tsx src/mastra/voice-worker.ts dev
|
|
88
|
-
```
|
|
89
|
-
|
|
90
|
-
The model strings (`deepgram/nova-3`, `cartesia/sonic-3`) route through **LiveKit Cloud inference**, so with a LiveKit Cloud project you don't need separate Deepgram/Cartesia accounts — only your `LIVEKIT_URL` / `LIVEKIT_API_KEY` / `LIVEKIT_API_SECRET` (plus whatever key your Mastra model needs). To bring your own providers, pass plugin instances to `stt` / `tts` instead of strings.
|
|
91
|
-
|
|
92
|
-
## Reply paths
|
|
93
|
-
|
|
94
|
-
### Agent (default)
|
|
95
|
-
|
|
96
|
-
Pass `agent` (a key/id, an `Agent` instance, or a resolver). The agent runs its full loop each turn — model, tools, memory, processors — and streams its text deltas to TTS. Barge-in cancels the in-flight `agent.stream()`.
|
|
97
|
-
|
|
98
|
-
### Workflow
|
|
99
|
-
|
|
100
|
-
Pass `workflow` + `workflowInput` instead of `agent` (mutually exclusive). LiveKit owns the turn boundary, so the workflow runs **once to completion per turn** — no suspend/resume. Use it for deterministic per-turn structure (e.g. classify intent, then reply).
|
|
101
|
-
|
|
102
|
-
```typescript
|
|
103
|
-
import { createLiveKitWorker, chatContextToMessages } from '@mastra/livekit/worker';
|
|
104
|
-
|
|
105
|
-
export default createLiveKitWorker({
|
|
106
|
-
mastra,
|
|
107
|
-
workflow: 'phoneConversation',
|
|
108
|
-
workflowInput: ({ messages, memory }) => ({ turn: messages, memory: memory || undefined }),
|
|
109
|
-
replyStep: 'generateResponse', // only stream text from this step (optional)
|
|
110
|
-
stt: 'deepgram/nova-3',
|
|
111
|
-
tts: 'cartesia/sonic-3',
|
|
112
|
-
turnDetection: 'multilingual',
|
|
113
|
-
});
|
|
114
|
-
```
|
|
115
|
-
|
|
116
|
-
In the reply-producing step, use **`pipeAgentReplyToWriter`** to forward the agent's reply into the step `writer`. It streams text deltas (so TTS starts early) **and** tool-call chunks (so `toolFeedback` fires and `onTurnComplete` sees the tool list) — unlike piping only `.textStream`, which silently drops tool calls:
|
|
117
|
-
|
|
118
|
-
```typescript
|
|
119
|
-
import { pipeAgentReplyToWriter } from '@mastra/livekit';
|
|
120
|
-
|
|
121
|
-
const generateResponse = createStep({
|
|
122
|
-
id: 'generateResponse',
|
|
123
|
-
execute: async ({ inputData, mastra, writer, abortSignal }) => {
|
|
124
|
-
const stream = await mastra.getAgent('support').stream(inputData.turn, {
|
|
125
|
-
memory: inputData.memory, // engages working memory, recall, etc.
|
|
126
|
-
abortSignal, // lets barge-in stop generation promptly
|
|
127
|
-
});
|
|
128
|
-
const reply = await pipeAgentReplyToWriter(stream, writer);
|
|
129
|
-
return { reply };
|
|
130
|
-
},
|
|
131
|
-
});
|
|
132
|
-
```
|
|
133
|
-
|
|
134
|
-
A step that writes no text stays silent unless you pass `resultText` to derive the reply from the final run result.
|
|
135
|
-
|
|
136
|
-
### Custom (`generate`)
|
|
137
|
-
|
|
138
|
-
For full control, pass a `generate` function — any `VoiceReplyGenerator` that turns a turn into a `ReadableStream<string>` (a remote bridge, a bespoke pipeline, …).
|
|
139
|
-
|
|
140
|
-
## `createLiveKitWorker` options
|
|
141
|
-
|
|
142
|
-
| Option | Type | Notes |
|
|
143
|
-
| --------------------------------------------------- | ----------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
144
|
-
| `mastra` | `Mastra` | **Required.** The instance whose agents/workflows answer sessions. |
|
|
145
|
-
| **Reply generation** (pick one) | | |
|
|
146
|
-
| `agent` | `string \| Agent \| (args) => …` | The agent that answers. Defaults to `metadata.agentId`. |
|
|
147
|
-
| `workflow` | `string \| Workflow \| (args) => string` | Answer with a workflow instead. Requires `workflowInput`. |
|
|
148
|
-
| `workflowInput` | `(ctx & { metadata }) => inputData` | Maps a turn into the workflow's `inputData`. |
|
|
149
|
-
| `replyStep` | `string` | Only stream text from this workflow step id. |
|
|
150
|
-
| `resultText` | `(result) => string` | Fallback reply text when the workflow streams nothing. |
|
|
151
|
-
| `generate` | `VoiceReplyGenerator` | Lowest-level escape hatch. |
|
|
152
|
-
| **Speech stack** | | |
|
|
153
|
-
| `stt` | plugin or `'provider/model'` | Speech-to-text. |
|
|
154
|
-
| `tts` | plugin or `'provider/model'` | Text-to-speech. |
|
|
155
|
-
| `vad` | `VAD \| 'silero' \| false` | Voice activity detection. Defaults to `'silero'`. |
|
|
156
|
-
| `turnDetection` | `'multilingual' \| 'english' \| …` | End-of-turn detection. |
|
|
157
|
-
| `turnHandling` | `AgentSessionOptions['turnHandling']` | Endpointing delays, interruption sensitivity, preemptive generation. |
|
|
158
|
-
| `sessionOptions` / `inputOptions` / `outputOptions` | partial LiveKit options | Merged over what the helper builds. |
|
|
159
|
-
| **Memory** | | |
|
|
160
|
-
| `memory` | `false \| (args) => { thread, resource }` | Memory mapping. Defaults to `{ thread: metadata.threadId ?? room, resource: metadata.resourceId ?? thread }` when the agent has memory. |
|
|
161
|
-
| `memoryInstance` | `Memory \| (args) => Memory` | The `Memory` used to bootstrap the thread + persist the greeting on the **workflow/custom** path (no agent to source it from). Mastra storage is injected if the `Memory` has none. |
|
|
162
|
-
| `configuration` | `{ greeting?, consentPolicy?, endCall?, stt?, tts? }` | Grouped conversation & compliance config (see [Configuration](#configuration)). `greeting` (`text` — a string or a per-tenant resolver — plus `allowInterruptions`, `awaitPlayout`, `persist`, `repeatEvery`/`repeatText` for periodic AI re-disclosure), `consentPolicy` (extensible consent set, e.g. `summaryStorage`, surfaced on `onCallEnd`), `endCall` (agent-initiated hang-up; pair with `createEndCallTool`), and per-call `stt` / `tts` resolvers that pick this call's transcriber and voice (fall back to the top-level options). |
|
|
163
|
-
| **Lifecycle hooks** | | |
|
|
164
|
-
| `toolFeedback` | `(toolCall) => string \| void` | Speak filler while a tool runs (agent + workflow). |
|
|
165
|
-
| `onTurnComplete` | `VoiceTurnCompleteHook` | After each turn streams, **fire-and-forget**, off the audio path. |
|
|
166
|
-
| `onCallEnd` | `VoiceCallEndHook` | When the call ends, **awaited** within LiveKit's shutdown window. |
|
|
167
|
-
| `onSessionStart` | `(args) => …` | After the session starts — attach listeners, trigger replies, etc. |
|
|
168
|
-
| **Other** | | |
|
|
169
|
-
| `observability` | `boolean` | Voice-pipeline tracing. Defaults to `true`. |
|
|
170
|
-
|
|
171
|
-
## Configuration
|
|
172
|
-
|
|
173
|
-
`configuration` groups related conversation & compliance knobs in one place, so they don't each
|
|
174
|
-
become a top-level worker option — and it's where further compliance controls land as they ship.
|
|
175
|
-
|
|
176
|
-
```typescript
|
|
177
|
-
createLiveKitWorker({
|
|
178
|
-
mastra,
|
|
179
|
-
agent: 'support',
|
|
180
|
-
configuration: {
|
|
181
|
-
greeting: {
|
|
182
|
-
// Spoken via TTS at call start (no model round-trip). Doubles as a required AI disclosure.
|
|
183
|
-
text: 'You are speaking with an AI assistant. This call may be recorded. How can I help?',
|
|
184
|
-
allowInterruptions: false, // caller can't barge over the disclosure (EU AI Act Art. 50)
|
|
185
|
-
awaitPlayout: true, // hold post-greeting work until the disclosure finishes
|
|
186
|
-
persist: true, // save it to the memory thread (default true)
|
|
187
|
-
// Periodic re-disclosure on long calls: once this interval elapses, the NEXT turn's reply is
|
|
188
|
-
// prefixed with a short "you're speaking with an AI" reminder (spoken at the turn boundary,
|
|
189
|
-
// never mid-turn). Omit to disable.
|
|
190
|
-
repeatEvery: 3 * 60_000, // ~every 3 minutes (California SB 243 and similar)
|
|
191
|
-
repeatText: 'Quick reminder — you are speaking with an AI assistant.', // optional; has a default
|
|
192
|
-
},
|
|
193
|
-
// Consent requirements — a named, extensible set. Each item is independently required and
|
|
194
|
-
// independently granted, so new items are added without one global "consented" flag.
|
|
195
|
-
consentPolicy: {
|
|
196
|
-
summaryStorage: true, // or { required: true, purpose: 'storing a summary of this call' }
|
|
197
|
-
},
|
|
198
|
-
// Let the agent end the call itself. Pair with a `createEndCallTool` tool on the agent; the
|
|
199
|
-
// worker waits for the agent's closing words to play out, then hangs up (running onCallEnd).
|
|
200
|
-
endCall: {
|
|
201
|
-
message: 'Thanks for calling. Goodbye!', // optional non-interruptible sign-off before hangup
|
|
202
|
-
},
|
|
203
|
-
},
|
|
204
|
-
});
|
|
205
|
-
```
|
|
206
|
-
|
|
207
|
-
**Greeting / AI disclosure.** `text` is spoken at call start; `allowInterruptions: false` makes a
|
|
208
|
-
required disclosure play through; `awaitPlayout: true` waits for it before anything else runs;
|
|
209
|
-
`repeatEvery` re-discloses periodically on long calls. Under the EU AI Act (Art. 50) a person must be
|
|
210
|
-
told they're interacting with an AI at the first interaction.
|
|
211
|
-
|
|
212
|
-
**Per-tenant greeting.** `greeting.text` also takes a resolver — a function called once per call
|
|
213
|
-
(post-connect) with the call context (`metadata`, `requestContext`, `roomName`, `ctx`). Return a
|
|
214
|
-
greeting keyed off the dispatch metadata so one multi-tenant agent opens differently per tenant (the
|
|
215
|
-
disclosure options still apply to whatever it returns; return `undefined` for no greeting):
|
|
216
|
-
|
|
217
|
-
```typescript
|
|
218
|
-
configuration: {
|
|
219
|
-
greeting: {
|
|
220
|
-
text: ({ requestContext }) => {
|
|
221
|
-
const tenant = TENANTS[requestContext?.tenantId as string];
|
|
222
|
-
return `Thanks for calling ${tenant?.name ?? 'us'}. You're speaking with an AI assistant.`;
|
|
223
|
-
},
|
|
224
|
-
allowInterruptions: false, // the disclosure still can't be barged over
|
|
225
|
-
},
|
|
226
|
-
},
|
|
227
|
-
```
|
|
228
|
-
|
|
229
|
-
**Consent.** `consentPolicy` _declares_ which consents the call needs (starting with `summaryStorage`
|
|
230
|
-
— consent to store a summary of the call). Capture the caller's decision at runtime with
|
|
231
|
-
**`createConsentTool`** (exported from `@mastra/livekit`) — add it to your agent, and it reads the
|
|
232
|
-
caller identity from the tool context and hands each decision to your store:
|
|
233
|
-
|
|
234
|
-
```typescript
|
|
235
|
-
import { createConsentTool } from '@mastra/livekit';
|
|
236
|
-
|
|
237
|
-
// in your agent's tools:
|
|
238
|
-
recordConsent: createConsentTool({
|
|
239
|
-
items: ['summaryStorage'],
|
|
240
|
-
onGrant: async ({ item, granted, resourceId }) => {
|
|
241
|
-
if (resourceId) await db.saveConsent(resourceId, item, granted); // your system of record
|
|
242
|
-
},
|
|
243
|
-
}),
|
|
244
|
-
```
|
|
245
|
-
|
|
246
|
-
Then _enforce_ it: the requirements are surfaced on the `onCallEnd` hook (`args.configuration`), so
|
|
247
|
-
you only run the consent-gated action — e.g. the `memory.summarizeThread()` call that stores the
|
|
248
|
-
summary — when it isn't required or the caller granted it. Whether to gate at all is your policy
|
|
249
|
-
call: a permissive deployment skips `consentPolicy` entirely and always summarizes, while a
|
|
250
|
-
regulated one declares → captures → enforces (the runnable example ships both flavors). Further
|
|
251
|
-
controls (recording notice, data retention, human handoff) are planned to land here too.
|
|
252
|
-
|
|
253
|
-
**Agent-initiated hang-up.** `endCall` lets the agent end the call itself — say goodbye, then hang up.
|
|
254
|
-
Enable it under `configuration`, and add a matching tool to the agent with **`createEndCallTool`**
|
|
255
|
-
(both default to the tool name `'endCall'`). The tool only _signals_ intent — from inside
|
|
256
|
-
`agent.stream()` it can't reach the room — so the worker owns the hang-up: on each turn it watches for
|
|
257
|
-
the tool, waits for the agent's closing words to finish playing, holds a short drain (`drainMs`,
|
|
258
|
-
default 800ms) so audio still buffered at the caller isn't clipped, then disconnects, running
|
|
259
|
-
`onCallEnd` on the way out exactly as a caller hang-up does. It works on the
|
|
260
|
-
agent and workflow reply paths.
|
|
261
|
-
|
|
262
|
-
```typescript
|
|
263
|
-
import { createEndCallTool } from '@mastra/livekit';
|
|
264
|
-
|
|
265
|
-
// in your agent's tools:
|
|
266
|
-
endCall: createEndCallTool({
|
|
267
|
-
// optional bookkeeping — the tool reads the caller identity from its context
|
|
268
|
-
onEndCall: ({ reason, resourceId }) => log.info('agent ended call', { reason, resourceId }),
|
|
269
|
-
}),
|
|
270
|
-
```
|
|
271
|
-
|
|
272
|
-
Instruct the agent to say its goodbye and then call `endCall` as its final action. `endCall.message`
|
|
273
|
-
adds a guaranteed non-interruptible sign-off spoken right before hanging up; `endCall.reason` sets the
|
|
274
|
-
shutdown reason in LiveKit logs; `endCall.maxWaitMs` caps how long to wait for the closing words
|
|
275
|
-
(default 30s). This is the AI-oversight companion to a future first-class human `handoff`.
|
|
276
|
-
|
|
277
|
-
**Backwards compatibility.** The previous top-level `greeting` (string) and `persistGreeting` options
|
|
278
|
-
still work — they're deprecated aliases for `configuration.greeting.text` and
|
|
279
|
-
`configuration.greeting.persist`. If both are set, `configuration.greeting` wins field-by-field, so
|
|
280
|
-
existing worker configs keep running unchanged while you migrate.
|
|
281
|
-
|
|
282
|
-
## Lifecycle hooks
|
|
283
|
-
|
|
284
|
-
Three hooks let you do work around a turn without adding to the caller's latency:
|
|
285
|
-
|
|
286
|
-
```typescript
|
|
287
|
-
createLiveKitWorker({
|
|
288
|
-
mastra,
|
|
289
|
-
agent: 'support',
|
|
290
|
-
|
|
291
|
-
// 1. In-turn: speak a short phrase while a tool runs, so the caller isn't left in silence.
|
|
292
|
-
toolFeedback: ({ toolName }) => (toolName === 'lookupOrder' ? 'Let me pull that up.' : undefined),
|
|
293
|
-
|
|
294
|
-
// 2. Post-turn: fire-and-forget AFTER the reply has streamed — the worker never awaits it, so it
|
|
295
|
-
// can't delay the caller or the next turn. Carries the produced reply + the memory mapping.
|
|
296
|
-
onTurnComplete: async ({ result, memory }) => {
|
|
297
|
-
if (memory) await crm.logContact(memory.resource, result.text); // result.text/toolCalls/interrupted
|
|
298
|
-
},
|
|
299
|
-
|
|
300
|
-
// 3. End-of-call: runs when the caller hangs up, AWAITED within LiveKit's shutdown grace window
|
|
301
|
-
// (so it finishes before the process exits). The place for end-of-call work — e.g. summarize
|
|
302
|
-
// the finished call into your own records. `memory.summarizeThread()` (@mastra/memory) runs a
|
|
303
|
-
// one-shot summarization + structured extraction over the whole call, outside observational
|
|
304
|
-
// memory's lifecycle — nothing is written back to memory; you decide where the result goes
|
|
305
|
-
// (return value and/or each Extractor's `onExtracted` hook).
|
|
306
|
-
onCallEnd: async ({ memory }) => {
|
|
307
|
-
if (!memory) return;
|
|
308
|
-
await myMemory.summarizeThread({
|
|
309
|
-
// your app's `Memory` (from @mastra/memory)
|
|
310
|
-
model: 'openai/gpt-4.1-mini',
|
|
311
|
-
threadId: memory.thread,
|
|
312
|
-
resourceId: memory.resource,
|
|
313
|
-
instructions: 'Summarize this call for the business owner.',
|
|
314
|
-
});
|
|
315
|
-
},
|
|
316
|
-
});
|
|
317
|
-
```
|
|
318
|
-
|
|
319
|
-
`onTurnComplete` and `toolFeedback` work on the **workflow** path too (the reply step must surface tool calls via `pipeAgentReplyToWriter`).
|
|
320
|
-
|
|
321
|
-
## Joining a call: `liveKitConnectionRoute`
|
|
322
|
-
|
|
323
|
-
Mounts an API route on your Mastra server that mints a LiveKit token and dispatches the worker by `agentName`. Frontends `POST` to it to get connection details.
|
|
324
|
-
|
|
325
|
-
| Option | Default | Notes |
|
|
326
|
-
| ------------------------------------ | -------------------------------------------------------- | --------------------------------------------------- |
|
|
327
|
-
| `path` | `/voice/livekit/connection-details` | Must not start with `/api`. |
|
|
328
|
-
| `serverUrl` / `apiKey` / `apiSecret` | `LIVEKIT_URL` / `LIVEKIT_API_KEY` / `LIVEKIT_API_SECRET` | LiveKit credentials. |
|
|
329
|
-
| `agentName` | — | Must match the worker's `agentName`. |
|
|
330
|
-
| `ttl` | `'15m'` | Token lifetime. |
|
|
331
|
-
| `requiresAuth` | `true` | Mastra custom routes require auth unless opted out. |
|
|
332
|
-
| `roomName` / `participantIdentity` | generated | String or `(args) => string`. |
|
|
333
|
-
| `metadata` | passes `agentId`/`threadId`/`resourceId` | Session metadata delivered to the worker. |
|
|
334
|
-
|
|
335
|
-
For programmatic dispatch (no HTTP), use `dispatchVoiceSession`.
|
|
336
|
-
|
|
337
|
-
## Running the worker: `runLiveKitWorker`
|
|
338
|
-
|
|
339
|
-
Starts the LiveKit agent worker CLI (`dev` / `start` / `connect`) for your entry file.
|
|
340
|
-
|
|
341
|
-
| Option | Default | Notes |
|
|
342
|
-
| --------------- | ---------------- | --------------------------------------------------- |
|
|
343
|
-
| `entry` | — | The worker module; pass `import.meta.url`. |
|
|
344
|
-
| `agentName` | `'mastra-voice'` | Dispatch name; must match `liveKitConnectionRoute`. |
|
|
345
|
-
| `serverOptions` | — | Extra LiveKit `ServerOptions`. |
|
|
346
|
-
|
|
347
|
-
## Observability
|
|
348
|
-
|
|
349
|
-
When the Mastra instance has observability configured, the worker opens one `voice call` span per session, nests every turn's Mastra run under it, and adds child spans for LiveKit pipeline metrics — STT, TTS, end-of-utterance, VAD, and LLM time-to-first-token — closing with a per-model token/character/audio usage roll-up. On by default; pass `observability: false` to disable.
|
|
350
|
-
|
|
351
|
-
## Deployment
|
|
36
|
+
Run a separate LiveKit worker to own the audio pipeline and call the agent for each detected turn. See the quickstart for the worker setup and required LiveKit plugins.
|
|
352
37
|
|
|
353
|
-
|
|
354
|
-
|
|
355
|
-
| Component | What runs it | Notes |
|
|
356
|
-
| ------------------------ | ------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------- |
|
|
357
|
-
| **LiveKit media server** | LiveKit Cloud, or self-hosted `livekit-server` + Redis | WebRTC transport. Both processes below need its `LIVEKIT_URL` / `LIVEKIT_API_KEY` / `LIVEKIT_API_SECRET`. |
|
|
358
|
-
| **Mastra HTTP server** | `mastra build` → `node .mastra/output/index.mjs` | Your agents/workflows + `liveKitConnectionRoute` (mints tokens, dispatches the worker). |
|
|
359
|
-
| **LiveKit worker** | your worker entry → `runLiveKitWorker` (LiveKit Agents CLI `start`) | Connects **outbound** to the media server and answers calls. Not an HTTP server. |
|
|
360
|
-
|
|
361
|
-
The worker is a **separate process**: `mastra build` bundles only your `Mastra` instance and `src/mastra/tools/**` — never the worker entry, because nothing imports it. `liveKitConnectionRoute`, by contrast, is an `apiRoute`, so it ships inside the server build automatically. The server and worker therefore build and run independently and can live on different hosts, as long as they share a LiveKit project and the same `agentName`.
|
|
362
|
-
|
|
363
|
-
> Note: `mastra worker build` is unrelated — it bundles Mastra's own pubsub/scheduler workflow workers (`mastra.startWorkers()`), not this LiveKit worker.
|
|
364
|
-
|
|
365
|
-
### Building the worker
|
|
366
|
-
|
|
367
|
-
The worker imports `@livekit/agents` (plus the optional plugins) and your `Mastra` instance, then runs the LiveKit Agents CLI. Two build-time essentials:
|
|
368
|
-
|
|
369
|
-
- **Run `start`, not `dev`, in production** — `dev` is hot-reload only.
|
|
370
|
-
- **Pre-download the model files** so they're baked into the image instead of fetched on cold start:
|
|
371
|
-
|
|
372
|
-
```bash
|
|
373
|
-
node --import tsx src/mastra/voice-worker.ts download-files # Silero VAD + turn-detector ONNX
|
|
374
|
-
```
|
|
375
|
-
|
|
376
|
-
Running the worker with `tsx` against source avoids bundling LiveKit's native deps (onnxruntime, …). If you do, make `tsx` and the `@livekit/agents*` packages **real** dependencies (not devDependencies) in the deployed image.
|
|
377
|
-
|
|
378
|
-
### Docker
|
|
379
|
-
|
|
380
|
-
Build the server with `mastra build`, bake the model files, then run the two processes. **Recommended: one image, two services** — workers scale by call volume and the HTTP server scales by request volume, so keep them independent:
|
|
381
|
-
|
|
382
|
-
```dockerfile
|
|
383
|
-
FROM node:22-slim AS build
|
|
384
|
-
WORKDIR /app
|
|
385
|
-
COPY package.json package-lock.json ./
|
|
386
|
-
RUN npm ci
|
|
387
|
-
COPY . .
|
|
388
|
-
RUN npx mastra build
|
|
389
|
-
RUN node --import tsx src/mastra/voice-worker.ts download-files
|
|
390
|
-
|
|
391
|
-
FROM node:22-slim
|
|
392
|
-
WORKDIR /app
|
|
393
|
-
COPY --from=build /app /app
|
|
394
|
-
ENV NODE_ENV=production
|
|
395
|
-
EXPOSE 4111
|
|
396
|
-
# server service: CMD ["node", ".mastra/output/index.mjs"]
|
|
397
|
-
# worker service: CMD ["node", "--import", "tsx", "src/mastra/voice-worker.ts", "start"]
|
|
398
|
-
```
|
|
399
|
-
|
|
400
|
-
To run **both in one container** (single-tenant boxes, demos), supervise them with `bash` so the container exits — and is restarted by the orchestrator — if either dies:
|
|
401
|
-
|
|
402
|
-
```bash
|
|
403
|
-
#!/usr/bin/env bash
|
|
404
|
-
set -euo pipefail
|
|
405
|
-
node .mastra/output/index.mjs &
|
|
406
|
-
node --import tsx src/mastra/voice-worker.ts start &
|
|
407
|
-
wait -n
|
|
408
|
-
exit 1
|
|
409
|
-
```
|
|
410
|
-
|
|
411
|
-
This is simpler, but it couples two processes that have opposite scaling curves and no independent autoscaling — prefer the split for anything beyond a demo.
|
|
412
|
-
|
|
413
|
-
### Managed platforms (Mastra Cloud, Railway, Cloud Run, …)
|
|
414
|
-
|
|
415
|
-
Single-process HTTP hosts run the **Mastra server** as-is (`node .mastra/output/index.mjs`). They can't host the worker — it isn't an HTTP server, it isn't in the build output, and it's a long-lived outbound connection. Deploy the **hybrid**: the server on the managed platform (it still mints tokens and dispatches via the bundled connection route), and the worker on any plain process host (a dedicated Railway/Fly/Render service, a VM, a Kubernetes `Deployment`, ECS, …). Point both at the same LiveKit project, keep ≥1 always-on worker instance (no scale-to-zero, so it stays registered), and inject the shared `LIVEKIT_*` env into both.
|
|
416
|
-
|
|
417
|
-
## Runnable example
|
|
418
|
-
|
|
419
|
-
A complete, runnable reference lives in the Mastra monorepo at **[`examples/voice-agent`](https://github.com/mastra-ai/mastra/tree/main/examples/voice-agent)** — a trades-contractor front-desk voice agent that exercises nearly every feature here:
|
|
420
|
-
|
|
421
|
-
- **Three workers, one agent name** (run one at a time): the default **agent** worker (`pnpm worker`) and **workflow** worker (`pnpm worker:workflow`, deterministic intent routing → memory-backed reply) are deliberately **permissive** — no consent friction, the end-of-call summary always runs. The **regulated** worker (`pnpm worker:regulated`, a "Northwind Financial" line) demonstrates every compliance control at once: non-interruptible AI disclosure, 45-second re-disclosure, a four-item runtime consent sweep with an audit ledger, agent-initiated hang-up with a compliance sign-off, and a consent-gated summary (no consent → no stored summary).
|
|
422
|
-
- **End-of-call summarization** — every finished call is distilled into a structured record (summary, sentiment, requested services) via `memory.summarizeThread()` + an `Extractor` whose `onExtracted` hook writes to the app's own store, from the `onCallEnd` hook.
|
|
423
|
-
- **Three memory layers** — working memory, semantic recall, and observational memory, all scoped to the caller.
|
|
424
|
-
- **Tools + deterministic reconciliation**, a tenant-context input processor, the `toolFeedback` / `onTurnComplete` / `onCallEnd` hooks, and full observability.
|
|
425
|
-
|
|
426
|
-
To run it:
|
|
427
|
-
|
|
428
|
-
```bash
|
|
429
|
-
git clone https://github.com/mastra-ai/mastra
|
|
430
|
-
cd mastra && pnpm install && pnpm build:packages # build the workspace packages
|
|
38
|
+
## Documentation
|
|
431
39
|
|
|
432
|
-
|
|
433
|
-
cp .env.example .env # add LiveKit Cloud creds + your model key (e.g. OPENAI_API_KEY)
|
|
434
|
-
pnpm install
|
|
435
|
-
pnpm worker:download-files # one-time model download
|
|
40
|
+
- [@mastra/livekit documentation](https://mastra.ai/integrations/voice/livekit)
|
|
436
41
|
|
|
437
|
-
|
|
438
|
-
pnpm dev # Mastra server + Studio at http://localhost:4111
|
|
439
|
-
pnpm worker # the voice worker (or `pnpm worker:workflow` / `pnpm worker:regulated`)
|
|
440
|
-
```
|
|
42
|
+
## Changelog
|
|
441
43
|
|
|
442
|
-
See the
|
|
44
|
+
See the [package changelog](https://github.com/mastra-ai/mastra/blob/main/integrations/livekit/CHANGELOG.md) for version history and release notes.
|
|
443
45
|
|
|
444
|
-
##
|
|
46
|
+
## Support
|
|
445
47
|
|
|
446
|
-
|
|
447
|
-
- [`@mastra/livekit` reference](https://mastra.ai/reference/voice/livekit)
|
|
448
|
-
- [LiveKit Agents docs](https://docs.livekit.io/agents/)
|
|
48
|
+
We have an [open community Discord](https://discord.gg/mastra-ai). Come and say hello and let us know if you have any questions or need any help getting things running.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@mastra/livekit",
|
|
3
|
-
"version": "0.3.1-alpha.
|
|
3
|
+
"version": "0.3.1-alpha.2",
|
|
4
4
|
"description": "LiveKit voice integration for Mastra agents — realtime voice with semantic turn detection and barge-in",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "./dist/index.js",
|
|
@@ -92,9 +92,9 @@
|
|
|
92
92
|
"typescript": "^7.0.2",
|
|
93
93
|
"vitest": "4.1.10",
|
|
94
94
|
"zod": "^4.4.3",
|
|
95
|
-
"@internal/types-builder": "0.0.104",
|
|
96
95
|
"@internal/lint": "0.0.129",
|
|
97
|
-
"@
|
|
96
|
+
"@internal/types-builder": "0.0.104",
|
|
97
|
+
"@mastra/core": "1.64.0-alpha.7"
|
|
98
98
|
},
|
|
99
99
|
"scripts": {
|
|
100
100
|
"build:lib": "tsdown --silent --config tsdown.config.ts",
|