realtime-voice-agents 2.5.3 → 2.5.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +104 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -75,7 +75,110 @@ app.register(async (i) => {
|
|
|
75
75
|
await app.listen({ port: 3000 });
|
|
76
76
|
```
|
|
77
77
|
|
|
78
|
-
Point your Twilio number's Voice webhook at `POST /twilio/voice`. That's a working agent. See [examples/fastify](examples/fastify) for the full tour (handoffs, strategies, approvals, outbound calls) and [examples/express-ws](examples/express-ws) for the minimal version.
|
|
78
|
+
Point your Twilio number's Voice webhook at `POST /twilio/voice`. That's a working agent. See [examples/fastify](examples/fastify) for the full tour (handoffs, strategies, approvals, outbound calls) and [examples/express-ws](examples/express-ws) for the minimal version. Running on **GPT-Live**? Read [Using GPT-Live](#using-gpt-live-full-duplex) next — the prompting is different.
|
|
79
|
+
|
|
80
|
+
## Using GPT-Live (full-duplex)
|
|
81
|
+
|
|
82
|
+
[GPT-Live](https://developers.openai.com/api/docs/guides/live) is OpenAI's full-duplex voice API: the voice model listens while it speaks and owns turn-taking, and a separate **backend** Responses model does the reasoning and runs your tools while the conversation keeps going. Same `Agent` / `tool()` / bridge as above — the difference is how you prompt it. This is the pattern we run in production:
|
|
83
|
+
|
|
84
|
+
```ts
|
|
85
|
+
import { Agent, TwilioRealtimeBridge, tool } from 'realtime-voice-agents';
|
|
86
|
+
import { gptLive } from 'realtime-voice-agents/gpt-live';
|
|
87
|
+
|
|
88
|
+
const saveMessage = tool({
|
|
89
|
+
name: 'save_service_message',
|
|
90
|
+
description: 'Save a service message after the caller confirmed name, callback phone and subject.',
|
|
91
|
+
parameters: z.object({ customerName: z.string(), phone: z.string(), subject: z.string() }),
|
|
92
|
+
backgroundAudio: false, // the voice keeps the caller company while this runs — no hold loop needed
|
|
93
|
+
execute: async (args) => ({ success: true, ...args }),
|
|
94
|
+
});
|
|
95
|
+
|
|
96
|
+
// 1. The VOICE prompt lives on the Agent: style + three policies. Nothing about procedures.
|
|
97
|
+
const voiceInstructions = `
|
|
98
|
+
You are Dana, a voice agent at Acme's service desk. Speak naturally, calm pace, short sentences.
|
|
99
|
+
Goal: take a message — name, callback phone, subject. Read the phone number back and ask for confirmation.
|
|
100
|
+
|
|
101
|
+
Backchannel policy: minimal. No "uh-huh" while the caller is talking; acknowledge briefly after they finish.
|
|
102
|
+
Interruption policy: if the caller talks over you, stop at once and listen. Finish the opening sentence first.
|
|
103
|
+
|
|
104
|
+
Delegation policy:
|
|
105
|
+
Backend tools: save_service_message (stores the message); finish_call (hangs up).
|
|
106
|
+
Delegate to the backend when: the caller confirmed every detail; or the caller says goodbye — delegate finish_call on the first "bye", then say a short goodbye.
|
|
107
|
+
Do not delegate when: details are still missing or unconfirmed; the caller only greets or asks you to repeat.
|
|
108
|
+
While the backend works, tell the caller you are saving the message and keep listening. Never guess the
|
|
109
|
+
result — only after the backend confirms, say the message was saved.
|
|
110
|
+
`.trim();
|
|
111
|
+
|
|
112
|
+
// 2. The BACKEND prompt lives on the provider: procedures and tool rules. Short, one-sentence replies.
|
|
113
|
+
const backendInstructions = `
|
|
114
|
+
You are the backend for a service-message desk. The voice agent delegates to you only once the caller
|
|
115
|
+
has confirmed all details. Call save_service_message exactly once per confirmed message, with the details
|
|
116
|
+
as stated. When the caller says goodbye, call finish_call. Reply in one short sentence the voice agent can
|
|
117
|
+
relay. Do not ask the caller new questions.
|
|
118
|
+
`.trim();
|
|
119
|
+
|
|
120
|
+
const bridge = new TwilioRealtimeBridge({
|
|
121
|
+
agent: new Agent({ name: 'Dana', instructions: voiceInstructions, tools: [saveMessage] }),
|
|
122
|
+
provider: gptLive({
|
|
123
|
+
voice: 'marin',
|
|
124
|
+
delegation: { model: 'gpt-5.6-terra', instructions: backendInstructions },
|
|
125
|
+
}),
|
|
126
|
+
session: {
|
|
127
|
+
// Delivered as session.commentary.append — the model speaks the quoted line verbatim.
|
|
128
|
+
greeting: {
|
|
129
|
+
mode: 'agent-initiates',
|
|
130
|
+
instructions: 'Open the call with exactly: "Hi, you have reached Acme. What is your name, please?"',
|
|
131
|
+
},
|
|
132
|
+
},
|
|
133
|
+
builtinTools: { finishCall: true },
|
|
134
|
+
twilio: { accountSid, authToken },
|
|
135
|
+
});
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
What to know before your first call:
|
|
139
|
+
|
|
140
|
+
- **Two prompts, two jobs.** The voice prompt is *how to talk* (style, backchannel, interruption, when to delegate). The backend prompt is *what to do* (procedures, tool rules, reply format). Putting procedures in the voice prompt makes the voice narrate tools it cannot call; putting style in the backend prompt does nothing.
|
|
141
|
+
- **Delegate first, announce after.** The voice model cannot call tools — only the backend can. A voice that says "I'm transferring you" or "saving that now" without delegating leaves the caller waiting for nothing. Every voice prompt we ship carries the line: *never promise a transfer or a result in words without delegating first.*
|
|
142
|
+
- **Tools don't pause the voice.** Results reach the backend the moment they are ready (`toolResultDelivery` is ignored), and the voice relays them in its own words. Set `backgroundAudio: false` on tools — hold music over a voice that is still talking sounds broken.
|
|
143
|
+
- **Barge-in belongs to the model.** `interruptions`, `vad`, `noiseAdaptiveVad` and `session.interrupt()` are no-ops here, and `user.speech.*` never fires. Protect the greeting through the prompt ("finish the opening sentence first"), not through guards. First-turn deafness defaults to off.
|
|
144
|
+
- **Goodbyes work as usual.** `finish_call` is delivered as a commentary append, marks confirm the farewell played out, then the leg completes via REST.
|
|
145
|
+
|
|
146
|
+
### Multi-agent with GPT-Live: a voice per agent
|
|
147
|
+
|
|
148
|
+
Sessions are immutable (instructions, voice), so every handoff is a close-and-reopen with the attributed transcript seeded into the new session. That gives each agent its own voice for free — and a real gap to cover:
|
|
149
|
+
|
|
150
|
+
```ts
|
|
151
|
+
const billing = new Agent({
|
|
152
|
+
name: 'Michal', id: 'billing', voice: 'coral',
|
|
153
|
+
instructions: billingVoicePrompt, // "You were transferred this caller — acknowledge, don't re-greet."
|
|
154
|
+
handoffDescription: 'Transfer for balance, charges, invoices and payments.',
|
|
155
|
+
tools: [checkBalance],
|
|
156
|
+
providerOptions: { delegation: { responses: { instructions: billingBackendPrompt } } }, // per-agent backend prompt
|
|
157
|
+
});
|
|
158
|
+
const support = new Agent({ name: 'Ido', id: 'support', voice: 'cedar', /* … */ handoffs: [billing] });
|
|
159
|
+
billing.handoffs.push(support); // specialists can hand back and forth
|
|
160
|
+
|
|
161
|
+
const bridge = new TwilioRealtimeBridge({
|
|
162
|
+
agent: new Agent({
|
|
163
|
+
name: 'Rotem', id: 'reception', voice: 'marin',
|
|
164
|
+
instructions: receptionVoicePrompt, // "Only find out billing vs. support, then delegate — the backend transfers."
|
|
165
|
+
handoffs: [billing, support],
|
|
166
|
+
}),
|
|
167
|
+
provider: gptLive({
|
|
168
|
+
voice: 'marin',
|
|
169
|
+
delegation: { model: 'gpt-5.6-terra', instructions: 'Call exactly one transfer_to_* tool per delegation.' },
|
|
170
|
+
}),
|
|
171
|
+
session: { handoffHold: { spec: 'elevator-jazz' } }, // covers the reopen
|
|
172
|
+
builtinTools: { finishCall: true },
|
|
173
|
+
});
|
|
174
|
+
```
|
|
175
|
+
|
|
176
|
+
- **The transfer is a backend tool.** Each specialist's voice prompt says so explicitly: *"the transfer is done by the backend (transfer_to_support) — you cannot transfer yourself; delegate first, then say 'transferring you'."*
|
|
177
|
+
- **The bridge waits for the sentence.** The backend's transfer lands while the voice is still announcing it; the handoff is held until that utterance has played (plus one sentence gap, capped at 8 s), so nothing is cut mid-word, and `handoffHold` audio covers the reopen.
|
|
178
|
+
- **Loops are structurally impossible.** An agent that just took over cannot transfer again until the caller speaks (see [handoffs](#multi-agent-handoffs-swarm)).
|
|
179
|
+
- **Tell incoming agents they were transferred.** Their voice prompt opens with "the caller was transferred to you — acknowledge briefly, do not greet again as if this were a new call"; the seeded transcript gives them the context.
|
|
180
|
+
|
|
181
|
+
Useful events for a test call: `agent.speech.started` / `agent.speech.ended` (utterance boundaries synthesized from the stream), `playback.started` / `playback.finished`, `transcript.user` / `transcript.agent`, `agent.handoff` / `agent.handoff.blocked`, `tool.started` / `tool.completed`, and `usage.updated` (`audioSeconds` — billing is per second of session — plus backend tokens). The full list of GPT-Live specifics is in [GPT-Live: full-duplex, two prompts](#gpt-live-full-duplex-two-prompts).
|
|
79
182
|
|
|
80
183
|
## Providers
|
|
81
184
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "realtime-voice-agents",
|
|
3
|
-
"version": "2.5.
|
|
3
|
+
"version": "2.5.4",
|
|
4
4
|
"description": "Provider-agnostic bridge between Twilio Media Streams and realtime speech-to-speech AI APIs (OpenAI Realtime, OpenAI GPT-Live full-duplex, xAI Grok Voice, Gemini Live). Multi-agent handoffs, Zod tools with execution strategies, mark-based playback tracking, interruption guards, and hold audio — for Node.js voice agents over the phone.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"twilio",
|