realtime-voice-agents 2.5.2 → 2.5.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +105 -2
- package/dist/gpt-live.cjs +13 -4
- package/dist/gpt-live.mjs +13 -4
- package/dist/index.cjs +7 -2
- package/dist/index.mjs +7 -2
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -75,7 +75,110 @@ app.register(async (i) => {
|
|
|
75
75
|
await app.listen({ port: 3000 });
|
|
76
76
|
```
|
|
77
77
|
|
|
78
|
-
Point your Twilio number's Voice webhook at `POST /twilio/voice`. That's a working agent. See [examples/fastify](examples/fastify) for the full tour (handoffs, strategies, approvals, outbound calls) and [examples/express-ws](examples/express-ws) for the minimal version.
|
|
78
|
+
Point your Twilio number's Voice webhook at `POST /twilio/voice`. That's a working agent. See [examples/fastify](examples/fastify) for the full tour (handoffs, strategies, approvals, outbound calls) and [examples/express-ws](examples/express-ws) for the minimal version. Running on **GPT-Live**? Read [Using GPT-Live](#using-gpt-live-full-duplex) next — the prompting is different.
|
|
79
|
+
|
|
80
|
+
## Using GPT-Live (full-duplex)
|
|
81
|
+
|
|
82
|
+
[GPT-Live](https://developers.openai.com/api/docs/guides/live) is OpenAI's full-duplex voice API: the voice model listens while it speaks and owns turn-taking, and a separate **backend** Responses model does the reasoning and runs your tools while the conversation keeps going. Same `Agent` / `tool()` / bridge as above — the difference is how you prompt it. This is the pattern we run in production:
|
|
83
|
+
|
|
84
|
+
```ts
|
|
85
|
+
import { Agent, TwilioRealtimeBridge, tool } from 'realtime-voice-agents';
|
|
86
|
+
import { gptLive } from 'realtime-voice-agents/gpt-live';
|
|
87
|
+
|
|
88
|
+
const saveMessage = tool({
|
|
89
|
+
name: 'save_service_message',
|
|
90
|
+
description: 'Save a service message after the caller confirmed name, callback phone and subject.',
|
|
91
|
+
parameters: z.object({ customerName: z.string(), phone: z.string(), subject: z.string() }),
|
|
92
|
+
backgroundAudio: false, // the voice keeps the caller company while this runs — no hold loop needed
|
|
93
|
+
execute: async (args) => ({ success: true, ...args }),
|
|
94
|
+
});
|
|
95
|
+
|
|
96
|
+
// 1. The VOICE prompt lives on the Agent: style + three policies. Nothing about procedures.
|
|
97
|
+
const voiceInstructions = `
|
|
98
|
+
You are Dana, a voice agent at Acme's service desk. Speak naturally, calm pace, short sentences.
|
|
99
|
+
Goal: take a message — name, callback phone, subject. Read the phone number back and ask for confirmation.
|
|
100
|
+
|
|
101
|
+
Backchannel policy: minimal. No "uh-huh" while the caller is talking; acknowledge briefly after they finish.
|
|
102
|
+
Interruption policy: if the caller talks over you, stop at once and listen. Finish the opening sentence first.
|
|
103
|
+
|
|
104
|
+
Delegation policy:
|
|
105
|
+
Backend tools: save_service_message (stores the message); finish_call (hangs up).
|
|
106
|
+
Delegate to the backend when: the caller confirmed every detail; or the caller says goodbye — delegate finish_call on the first "bye", then say a short goodbye.
|
|
107
|
+
Do not delegate when: details are still missing or unconfirmed; the caller only greets or asks you to repeat.
|
|
108
|
+
While the backend works, tell the caller you are saving the message and keep listening. Never guess the
|
|
109
|
+
result — only after the backend confirms, say the message was saved.
|
|
110
|
+
`.trim();
|
|
111
|
+
|
|
112
|
+
// 2. The BACKEND prompt lives on the provider: procedures and tool rules. Short, one-sentence replies.
|
|
113
|
+
const backendInstructions = `
|
|
114
|
+
You are the backend for a service-message desk. The voice agent delegates to you only once the caller
|
|
115
|
+
has confirmed all details. Call save_service_message exactly once per confirmed message, with the details
|
|
116
|
+
as stated. When the caller says goodbye, call finish_call. Reply in one short sentence the voice agent can
|
|
117
|
+
relay. Do not ask the caller new questions.
|
|
118
|
+
`.trim();
|
|
119
|
+
|
|
120
|
+
const bridge = new TwilioRealtimeBridge({
|
|
121
|
+
agent: new Agent({ name: 'Dana', instructions: voiceInstructions, tools: [saveMessage] }),
|
|
122
|
+
provider: gptLive({
|
|
123
|
+
voice: 'marin',
|
|
124
|
+
delegation: { model: 'gpt-5.6-terra', instructions: backendInstructions },
|
|
125
|
+
}),
|
|
126
|
+
session: {
|
|
127
|
+
// Delivered as session.commentary.append — the model speaks the quoted line verbatim.
|
|
128
|
+
greeting: {
|
|
129
|
+
mode: 'agent-initiates',
|
|
130
|
+
instructions: 'Open the call with exactly: "Hi, you have reached Acme. What is your name, please?"',
|
|
131
|
+
},
|
|
132
|
+
},
|
|
133
|
+
builtinTools: { finishCall: true },
|
|
134
|
+
twilio: { accountSid, authToken },
|
|
135
|
+
});
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
What to know before your first call:
|
|
139
|
+
|
|
140
|
+
- **Two prompts, two jobs.** The voice prompt is *how to talk* (style, backchannel, interruption, when to delegate). The backend prompt is *what to do* (procedures, tool rules, reply format). Putting procedures in the voice prompt makes the voice narrate tools it cannot call; putting style in the backend prompt does nothing.
|
|
141
|
+
- **Delegate first, announce after.** The voice model cannot call tools — only the backend can. A voice that says "I'm transferring you" or "saving that now" without delegating leaves the caller waiting for nothing. Every voice prompt we ship carries the line: *never promise a transfer or a result in words without delegating first.*
|
|
142
|
+
- **Tools don't pause the voice.** Results reach the backend the moment they are ready (`toolResultDelivery` is ignored), and the voice relays them in its own words. Set `backgroundAudio: false` on tools — hold music over a voice that is still talking sounds broken.
|
|
143
|
+
- **Barge-in belongs to the model.** `interruptions`, `vad`, `noiseAdaptiveVad` and `session.interrupt()` are no-ops here, and `user.speech.*` never fires. Protect the greeting through the prompt ("finish the opening sentence first"), not through guards. First-turn deafness defaults to off.
|
|
144
|
+
- **Goodbyes work as usual.** `finish_call` is delivered as a commentary append, marks confirm the farewell played out, then the leg completes via REST.
|
|
145
|
+
|
|
146
|
+
### Multi-agent with GPT-Live: a voice per agent
|
|
147
|
+
|
|
148
|
+
Sessions are immutable (instructions, voice), so every handoff is a close-and-reopen with the attributed transcript seeded into the new session. That gives each agent its own voice for free — and a real gap to cover:
|
|
149
|
+
|
|
150
|
+
```ts
|
|
151
|
+
const billing = new Agent({
|
|
152
|
+
name: 'Michal', id: 'billing', voice: 'coral',
|
|
153
|
+
instructions: billingVoicePrompt, // "You were transferred this caller — acknowledge, don't re-greet."
|
|
154
|
+
handoffDescription: 'Transfer for balance, charges, invoices and payments.',
|
|
155
|
+
tools: [checkBalance],
|
|
156
|
+
providerOptions: { delegation: { responses: { instructions: billingBackendPrompt } } }, // per-agent backend prompt
|
|
157
|
+
});
|
|
158
|
+
const support = new Agent({ name: 'Ido', id: 'support', voice: 'cedar', /* … */ handoffs: [billing] });
|
|
159
|
+
billing.handoffs.push(support); // specialists can hand back and forth
|
|
160
|
+
|
|
161
|
+
const bridge = new TwilioRealtimeBridge({
|
|
162
|
+
agent: new Agent({
|
|
163
|
+
name: 'Rotem', id: 'reception', voice: 'marin',
|
|
164
|
+
instructions: receptionVoicePrompt, // "Only find out billing vs. support, then delegate — the backend transfers."
|
|
165
|
+
handoffs: [billing, support],
|
|
166
|
+
}),
|
|
167
|
+
provider: gptLive({
|
|
168
|
+
voice: 'marin',
|
|
169
|
+
delegation: { model: 'gpt-5.6-terra', instructions: 'Call exactly one transfer_to_* tool per delegation.' },
|
|
170
|
+
}),
|
|
171
|
+
session: { handoffHold: { spec: 'elevator-jazz' } }, // covers the reopen
|
|
172
|
+
builtinTools: { finishCall: true },
|
|
173
|
+
});
|
|
174
|
+
```
|
|
175
|
+
|
|
176
|
+
- **The transfer is a backend tool.** Each specialist's voice prompt says so explicitly: *"the transfer is done by the backend (transfer_to_support) — you cannot transfer yourself; delegate first, then say 'transferring you'."*
|
|
177
|
+
- **The bridge waits for the sentence.** The backend's transfer lands while the voice is still announcing it; the handoff is held until that utterance has played (plus one sentence gap, capped at 8 s), so nothing is cut mid-word, and `handoffHold` audio covers the reopen.
|
|
178
|
+
- **Loops are structurally impossible.** An agent that just took over cannot transfer again until the caller speaks (see [handoffs](#multi-agent-handoffs-swarm)).
|
|
179
|
+
- **Tell incoming agents they were transferred.** Their voice prompt opens with "the caller was transferred to you — acknowledge briefly, do not greet again as if this were a new call"; the seeded transcript gives them the context.
|
|
180
|
+
|
|
181
|
+
Useful events for a test call: `agent.speech.started` / `agent.speech.ended` (utterance boundaries synthesized from the stream), `playback.started` / `playback.finished`, `transcript.user` / `transcript.agent`, `agent.handoff` / `agent.handoff.blocked`, `tool.started` / `tool.completed`, and `usage.updated` (`audioSeconds` — billing is per second of session — plus backend tokens). The full list of GPT-Live specifics is in [GPT-Live: full-duplex, two prompts](#gpt-live-full-duplex-two-prompts).
|
|
79
182
|
|
|
80
183
|
## Providers
|
|
81
184
|
|
|
@@ -119,7 +222,7 @@ One `SessionOptions` surface configures all four; where a provider can't honor a
|
|
|
119
222
|
- **Tools never pause the voice.** Results are delivered the moment they are ready regardless of `toolResultDelivery`; an interruption does not cancel a running tool, and its result still reaches the backend. Results are relayed in the model's own words — use exact wording only through the voice prompt.
|
|
120
223
|
- **Greetings, nudges and goodbyes** (`greeting.instructions`, `idle.prompts`, `finish_call`) are delivered as `session.commentary.append` — the append that reliably produces speech on demand. Keypad entries and deferred results are `session.thinking.append`; runtime instructions are `session.instructions.append`. Each append is capped at 500 tokens (long texts are split).
|
|
121
224
|
- **Immutable session.** Instructions, voice and audio format cannot change after start, so handoffs and reconnects open a fresh session and seed the attributed transcript through `session.input` (≤ 128 messages) — the anti-loop replay is preserved. Sessions expire after 120 minutes; an expiry reconnects the same way.
|
|
122
|
-
- **Transfers wait for the sentence.** A handoff here is a close-and-reopen, and the backend's transfer lands while the voice is still announcing it — the bridge holds the handoff until that utterance has played out (plus one sentence gap, capped at
|
|
225
|
+
- **Transfers wait for the sentence.** A handoff here is a close-and-reopen, and the backend's transfer lands while the voice is still announcing it — the bridge holds the handoff until that utterance has played out (plus one sentence gap, capped at 8 s), so nothing is cut mid-word and `session.handoffHold` audio covers the reopen. Prompt the voice to *delegate first, announce after*: a transfer or tool the voice announces without delegating never happens.
|
|
123
226
|
- **Deafness feeds silence.** The model's session clock runs on input audio, so `deafness` options replace caller audio with silence instead of dropping frames. `ignoreUserAudioUntilFirstTurnDone` therefore defaults to **off** here — the model handles talk-over itself; set it explicitly to keep the greeting deaf.
|
|
124
227
|
- **Real-time stream, 200 ms of cushion.** The voice arrives at exactly real-time pace, so Twilio's buffer never runs ahead of playout and every delivery hiccup between OpenAI, your server and Twilio would be an audible gap (Realtime generates faster than real time, so it never has this problem). The provider holds the first 200 ms of each utterance — the last idle delta included, so soft onsets are not clipped — then streams through. `gptLive({ playoutLeadMs })` tunes it, `0` disables; the cost is that much latency on each turn's first word.
|
|
125
228
|
- **Billing is per second** of session (plus backend tokens). `session.usage.audioSeconds` carries the running total; backend token usage is summed from `response.completed`. The provider sends `session.close` on teardown and waits for `session.closed`, so a hung-up call never keeps billing.
|
package/dist/gpt-live.cjs
CHANGED
|
@@ -689,18 +689,26 @@ var GptLiveProvider = class extends require_BaseRealtimeProvider.BaseRealtimePro
|
|
|
689
689
|
const wasOpen = this.gate.isOpen;
|
|
690
690
|
const events = this.gate.feed(bytes);
|
|
691
691
|
let forwardId = wasOpen ? this.currentUtteranceId : null;
|
|
692
|
-
let
|
|
693
|
-
for (const gateEvent of events)
|
|
692
|
+
let pendingClose = false;
|
|
693
|
+
for (const gateEvent of events) {
|
|
694
|
+
if (gateEvent.type === "close") {
|
|
695
|
+
pendingClose = true;
|
|
696
|
+
continue;
|
|
697
|
+
}
|
|
698
|
+
if (pendingClose) {
|
|
699
|
+
this.endUtterance();
|
|
700
|
+
pendingClose = false;
|
|
701
|
+
}
|
|
694
702
|
this.beginUtterance();
|
|
695
703
|
forwardId = this.currentUtteranceId;
|
|
696
704
|
this.armPlayoutLead();
|
|
697
|
-
}
|
|
705
|
+
}
|
|
698
706
|
if (forwardId) this.forwardAudio(delta, bytes.length / MULAW_BYTES_PER_MS, forwardId);
|
|
699
707
|
else this.preRoll = {
|
|
700
708
|
delta,
|
|
701
709
|
ms: bytes.length / MULAW_BYTES_PER_MS
|
|
702
710
|
};
|
|
703
|
-
if (
|
|
711
|
+
if (pendingClose) this.endUtterance();
|
|
704
712
|
else if (this.gate.isOpen) this.armGateStall();
|
|
705
713
|
}
|
|
706
714
|
/**
|
|
@@ -779,6 +787,7 @@ var GptLiveProvider = class extends require_BaseRealtimeProvider.BaseRealtimePro
|
|
|
779
787
|
});
|
|
780
788
|
}
|
|
781
789
|
beginUtterance() {
|
|
790
|
+
if (this.currentUtteranceId) this.endUtterance();
|
|
782
791
|
this.currentUtteranceId = `live_utt_${++this.utteranceCounter}`;
|
|
783
792
|
this.lastUtteranceId = this.currentUtteranceId;
|
|
784
793
|
this.emit("responseStarted", { responseId: this.currentUtteranceId });
|
package/dist/gpt-live.mjs
CHANGED
|
@@ -686,18 +686,26 @@ var GptLiveProvider = class extends BaseRealtimeProvider {
|
|
|
686
686
|
const wasOpen = this.gate.isOpen;
|
|
687
687
|
const events = this.gate.feed(bytes);
|
|
688
688
|
let forwardId = wasOpen ? this.currentUtteranceId : null;
|
|
689
|
-
let
|
|
690
|
-
for (const gateEvent of events)
|
|
689
|
+
let pendingClose = false;
|
|
690
|
+
for (const gateEvent of events) {
|
|
691
|
+
if (gateEvent.type === "close") {
|
|
692
|
+
pendingClose = true;
|
|
693
|
+
continue;
|
|
694
|
+
}
|
|
695
|
+
if (pendingClose) {
|
|
696
|
+
this.endUtterance();
|
|
697
|
+
pendingClose = false;
|
|
698
|
+
}
|
|
691
699
|
this.beginUtterance();
|
|
692
700
|
forwardId = this.currentUtteranceId;
|
|
693
701
|
this.armPlayoutLead();
|
|
694
|
-
}
|
|
702
|
+
}
|
|
695
703
|
if (forwardId) this.forwardAudio(delta, bytes.length / MULAW_BYTES_PER_MS, forwardId);
|
|
696
704
|
else this.preRoll = {
|
|
697
705
|
delta,
|
|
698
706
|
ms: bytes.length / MULAW_BYTES_PER_MS
|
|
699
707
|
};
|
|
700
|
-
if (
|
|
708
|
+
if (pendingClose) this.endUtterance();
|
|
701
709
|
else if (this.gate.isOpen) this.armGateStall();
|
|
702
710
|
}
|
|
703
711
|
/**
|
|
@@ -776,6 +784,7 @@ var GptLiveProvider = class extends BaseRealtimeProvider {
|
|
|
776
784
|
});
|
|
777
785
|
}
|
|
778
786
|
beginUtterance() {
|
|
787
|
+
if (this.currentUtteranceId) this.endUtterance();
|
|
779
788
|
this.currentUtteranceId = `live_utt_${++this.utteranceCounter}`;
|
|
780
789
|
this.lastUtteranceId = this.currentUtteranceId;
|
|
781
790
|
this.emit("responseStarted", { responseId: this.currentUtteranceId });
|
package/dist/index.cjs
CHANGED
|
@@ -2631,8 +2631,13 @@ const PREGREETING_MARK = "pre:greeting";
|
|
|
2631
2631
|
/** Longer than the speech gate's quiet window plus the mark round trip (0.8 s + ~0.3 s). */
|
|
2632
2632
|
/** Model-owned turn-taking: the gate can split one sentence at a pause — wait one gap for the next. */
|
|
2633
2633
|
const SENTENCE_GRACE_MS = 1500;
|
|
2634
|
-
/**
|
|
2635
|
-
|
|
2634
|
+
/**
|
|
2635
|
+
* A deferred handoff runs no later than this after the transfer landed, even
|
|
2636
|
+
* mid-utterance. One announce + a stray utterance + the sentence grace lands
|
|
2637
|
+
* right at 5 s in the field (Sept 2026); the cap is for a voice that never
|
|
2638
|
+
* stops, and the caller hears the agent meanwhile, so it errs long.
|
|
2639
|
+
*/
|
|
2640
|
+
const HANDOFF_DEFER_MAX_MS = 8e3;
|
|
2636
2641
|
/** 400ms per frame — matches production burst-write implementations. */
|
|
2637
2642
|
const PREGREETING_CHUNK_BYTES = 3200;
|
|
2638
2643
|
function toError(value) {
|
package/dist/index.mjs
CHANGED
|
@@ -2627,8 +2627,13 @@ const PREGREETING_MARK = "pre:greeting";
|
|
|
2627
2627
|
/** Longer than the speech gate's quiet window plus the mark round trip (0.8 s + ~0.3 s). */
|
|
2628
2628
|
/** Model-owned turn-taking: the gate can split one sentence at a pause — wait one gap for the next. */
|
|
2629
2629
|
const SENTENCE_GRACE_MS = 1500;
|
|
2630
|
-
/**
|
|
2631
|
-
|
|
2630
|
+
/**
|
|
2631
|
+
* A deferred handoff runs no later than this after the transfer landed, even
|
|
2632
|
+
* mid-utterance. One announce + a stray utterance + the sentence grace lands
|
|
2633
|
+
* right at 5 s in the field (Sept 2026); the cap is for a voice that never
|
|
2634
|
+
* stops, and the caller hears the agent meanwhile, so it errs long.
|
|
2635
|
+
*/
|
|
2636
|
+
const HANDOFF_DEFER_MAX_MS = 8e3;
|
|
2632
2637
|
/** 400ms per frame — matches production burst-write implementations. */
|
|
2633
2638
|
const PREGREETING_CHUNK_BYTES = 3200;
|
|
2634
2639
|
function toError(value) {
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "realtime-voice-agents",
|
|
3
|
-
"version": "2.5.
|
|
3
|
+
"version": "2.5.4",
|
|
4
4
|
"description": "Provider-agnostic bridge between Twilio Media Streams and realtime speech-to-speech AI APIs (OpenAI Realtime, OpenAI GPT-Live full-duplex, xAI Grok Voice, Gemini Live). Multi-agent handoffs, Zod tools with execution strategies, mark-based playback tracking, interruption guards, and hold audio — for Node.js voice agents over the phone.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"twilio",
|