@mindstudio-ai/remy 0.1.276 → 0.1.277

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -6,114 +6,54 @@ when: Before authoring `src/interfaces/voice.md`, choosing a voice model or pipe
6
6
 
7
7
  # Building Voice Interfaces
8
8
 
9
- A voice interface is the app's agent as a live phone-call-quality conversation: the user speaks, the
10
- agent answers in speech, and the app's methods are its tools. It is a **sibling of the agent
11
- interface, not a mode of it** — the two share a philosophy (an LLM projecting the backend contract
12
- into conversation; load the `agentInterfaces` skill for that shared ground), but everything you
13
- author differs. The persona is written for the ear, not the screen. The toolset is smaller and
14
- curated for conversational latency. And every tool declares how the agent should handle the wait,
15
- because in a live call, silence reads as a dropped line.
16
-
17
- The platform owns the hard parts — realtime audio transport, turn detection, interruption handling,
18
- transcripts, session limits, per-user auth on every tool call. Your job is the spec
19
- (`src/interfaces/voice.md`) and its compilation into `dist/interfaces/voice/`.
9
+ A voice interface is the app's agent as a live phone-call-quality conversation: the user speaks, the agent answers in speech, and the app's methods are its tools. It is a **sibling of the agent interface, not a mode of it** — the two share a philosophy (an LLM projecting the backend contract into conversation; load the `agentInterfaces` skill for that shared ground), but everything you author differs. The persona is written for the ear, not the screen. The toolset is smaller and curated for conversational latency. And every tool declares how the agent should handle the wait, because in a live call, silence reads as a dropped line.
10
+
11
+ The platform owns the hard parts — realtime audio transport, turn detection, interruption handling, transcripts, session limits, per-user auth on every tool call. Your job is the spec (`src/interfaces/voice.md`) and its compilation into `dist/interfaces/voice/`.
20
12
 
21
13
  ## Voice Agent Design
22
14
 
23
15
  ### Written for the ear
24
16
 
25
- Everything the agent produces gets spoken aloud. That inverts several habits that are correct
26
- everywhere else, and the compiled system prompt must carry them explicitly:
27
-
28
- - **No visual formatting, ever.** No markdown, no lists, no tables, no emoji, no URLs read as
29
- punctuation soup. If a tool returns a link, say what it is and where it will be, don't recite it.
30
- - **Spoken-form values.** "Forty-two fifty," not "$42.50". "Two fifteen in the afternoon," not
31
- "14:15". Read email addresses and confirmation codes character by character, and read them *back*
32
- for confirmation before acting on them — mishearing one digit of a phone number is the classic
33
- voice failure. Collect one value per turn; two asked together blend when spoken.
34
- - **Brevity is a hard rule, not a style preference.** One to two sentences per turn, one question at
35
- a time. A paragraph that reads fine in chat is a monologue on a call.
36
- - **Vary the phrasing.** Repeated openers and acknowledgments sound convincing once and robotic by
37
- the third turn — give the prompt an explicit variety rule, and treat any sample phrases as
38
- anchors, never scripts.
39
- - **Handle unclear audio explicitly.** Give the prompt a rule for it: respond only to clear audio;
40
- if it's noisy or ambiguous, ask the user to repeat — never guess, never call a tool on input the
41
- agent isn't sure it heard, and don't reuse the same clarification line twice in a row.
42
- - **Pin the language.** State the response language in the prompt; don't let the model infer it from
43
- an accent. If the app's domain has brand names or terms with non-obvious pronunciations, give
44
- them a line ("pronounce SQL as 'sequel'").
45
-
46
- Beyond the mechanics, the persona itself should be *of the ear*: pacing, warmth, how it handles
47
- being interrupted, what it says when it needs a second. This is the fun part, same as the agent
48
- interface — a distinct character beats a generic assistant, and voice makes character land harder
49
- than any other surface.
17
+ Everything the agent produces gets spoken aloud. That inverts several habits that are correct everywhere else, and the compiled system prompt must carry them explicitly:
18
+
19
+ - **No visual formatting, ever.** No markdown, no lists, no tables, no emoji, no URLs read as punctuation soup. If a tool returns a link, say what it is and where it will be, don't recite it.
20
+ - **Spoken-form values.** "Forty-two fifty," not "$42.50". "Two fifteen in the afternoon," not "14:15". Read email addresses and confirmation codes character by character, and read them *back* for confirmation before acting on them — mishearing one digit of a phone number is the classic voice failure. Collect one value per turn; two asked together blend when spoken.
21
+ - **Brevity is a hard rule, not a style preference.** One to two sentences per turn, one question at a time. A paragraph that reads fine in chat is a monologue on a call.
22
+ - **Vary the phrasing.** Repeated openers and acknowledgments sound convincing once and robotic by the third turn — give the prompt an explicit variety rule, and treat any sample phrases as anchors, never scripts.
23
+ - **Handle unclear audio explicitly.** Give the prompt a rule for it: respond only to clear audio; if it's noisy or ambiguous, ask the user to repeat — never guess, never call a tool on input the agent isn't sure it heard, and don't reuse the same clarification line twice in a row.
24
+ - **Pin the language.** State the response language in the prompt; don't let the model infer it from an accent. If the app's domain has brand names or terms with non-obvious pronunciations, give them a line ("pronounce SQL as 'sequel'").
25
+
26
+ Beyond the mechanics, the persona itself should be *of the ear*: pacing, warmth, how it handles being interrupted, what it says when it needs a second. This is the fun part, same as the agent interface — a distinct character beats a generic assistant, and voice makes character land harder than any other surface.
50
27
 
51
28
  ### The shape of `system.md`
52
29
 
53
- Structure the compiled prompt as short **labeled sections** — Role & Objective, Personality & Tone,
54
- Rules, and (when the app has a real call flow) Conversation Flow — with bullets over paragraphs;
55
- realtime models find and follow sectioned rules far more reliably than prose. Scope rules
56
- precisely; blanket `always`/`never` makes the agent rigid and unable to handle reasonable exceptions.
57
- And start minimal: state the role, the boundaries, and the voice mechanics above, then add rules only for
58
- behaviors that actually misfire in test calls (the transcripts in the call log are the feedback
59
- loop — `mindstudio-prod voice sessions get` reads a call verbatim) rather than front-loading a
60
- policy manual.
30
+ Structure the compiled prompt as short **labeled sections** — Role & Objective, Personality & Tone, Rules, and (when the app has a real call flow) Conversation Flow — with bullets over paragraphs; realtime models find and follow sectioned rules far more reliably than prose. Scope rules precisely; blanket `always`/`never` makes the agent rigid and unable to handle reasonable exceptions. And start minimal: state the role, the boundaries, and the voice mechanics above, then add rules only for behaviors that actually misfire in test calls (the transcripts in the call log are the feedback loop — `mindstudio-prod voice sessions get` reads a call verbatim) rather than front-loading a policy manual.
61
31
 
62
32
  ### The latency classes
63
33
 
64
- Every tool in the spec declares one of three classes. This is the voice-specific discipline — get it
65
- right and tool use feels like talking to a competent person; get it wrong and every action is an
66
- awkward pause.
34
+ Every tool in the spec declares one of three classes. This is the voice-specific discipline — get it right and tool use feels like talking to a competent person; get it wrong and every action is an awkward pause.
67
35
 
68
- - **`fast`** — sub-second reads: lookups, availability checks, small queries. The agent calls
69
- silently; announcing a sub-second call adds more delay than the call itself.
70
- - **`slow`** — a noticeable wait, roughly one to three seconds: writes, searches, anything that does
71
- real work. The agent speaks a one-line preamble ("Let me get that booked") generated in parallel
72
- with the call, so the line never goes quiet.
73
- - **`background`** — long-running work: reports, enrichment, bulk operations. The agent
74
- acknowledges, keeps conversing, and reports the result when it lands. Background tools are
75
- cancellable — if the user changes course mid-run, the work stops.
36
+ - **`fast`** — sub-second reads: lookups, availability checks, small queries. The agent calls silently; announcing a sub-second call adds more delay than the call itself.
37
+ - **`slow`** — a noticeable wait, roughly one to three seconds: writes, searches, anything that does real work. The agent speaks a one-line preamble ("Let me get that booked") generated in parallel with the call, so the line never goes quiet.
38
+ - **`background`** — long-running work: reports, enrichment, bulk operations. The agent acknowledges, keeps conversing, and reports the result when it lands. Background tools are cancellable — if the user changes course mid-run, the work stops.
76
39
 
77
- Classify by how the method actually behaves, not by what it is named. A "lookup" that fans out to an
78
- external service is `slow`. When in doubt between `fast` and `slow`, pick `slow` — a needless
79
- preamble is mildly chatty; an unexplained silence feels broken.
40
+ Classify by how the method actually behaves, not by what it is named. A "lookup" that fans out to an external service is `slow`. When in doubt between `fast` and `slow`, pick `slow` — a needless preamble is mildly chatty; an unexplained silence feels broken.
80
41
 
81
42
  ### Tool results reach the screen
82
43
 
83
- Every successful tool call delivers its raw return value to the session's browser on the SDK's
84
- `toolCall` event (`result` field, on `done`) — so the UI can render what the agent just did (the
85
- citation it found, the record it pulled up, the booking it made) in lockstep with the spoken
86
- answer. No flag, no polling, no key-threading, no model involvement: delivery to the invoking
87
- session's own client is the same security context as the invocation itself (an RPC response), and
88
- it's scoped to that one session.
89
-
90
- Consequence for authoring: **a tool's return value is user-visible by definition.** Return what
91
- the user may see — no internal fields, keys, or diagnostics you wouldn't put on screen (the same
92
- discipline as agent-interface tools, whose results render in chat). Payloads over ~32KB serialized
93
- arrive as `resultTruncated: true` with no data — keep returns compact, or have the UI fetch big
94
- data itself. Failed calls deliver nothing to the client (the model gets the `{ error }` and speaks
95
- a decline).
96
-
97
- For backend-side correlation (writing results to a table keyed by the call, custom channels), the
98
- method itself can read `session.voiceSessionId` / `session.visitorId` from the agent SDK
99
- (`import { session } from '@mindstudio-ai/agent'`) — the same id the browser holds as
100
- `session.sessionId`, guaranteed by the platform rather than echoed by the model.
101
-
102
- Voice sessions also carry `session.medium` (`'web' | 'phone-in' | 'phone-out'`) and, on phone
103
- calls, `session.sip` (`{ to, fromNumber }`) — use `medium` in tool methods and the session-context
104
- method to branch web vs phone behavior (what to prefetch, how to phrase context). Treat
105
- `session.sip.fromNumber` as context only, never identity: caller ID is spoofable, so don't gate
106
- data or roles on it — in-call verification is the auth rail.
107
-
108
- Ordinary web method calls carry the same context (`session.channel === 'web'` with `visitorId`),
109
- so correlating a browser's web activity with its voice calls can key on `session.visitorId` on
110
- both sides — one browser, one visitor id, web and voice alike. `session.voiceSessionId` remains
111
- the per-call key when you need to distinguish individual calls.
44
+ Every successful tool call delivers its raw return value to the session's browser on the SDK's `toolCall` event (`result` field, on `done`) — so the UI can render what the agent just did (the citation it found, the record it pulled up, the booking it made) in lockstep with the spoken answer. No flag, no polling, no key-threading, no model involvement: delivery to the invoking session's own client is the same security context as the invocation itself (an RPC response), and it's scoped to that one session.
45
+
46
+ Consequence for authoring: **a tool's return value is user-visible by definition.** Return what the user may see — no internal fields, keys, or diagnostics you wouldn't put on screen (the same discipline as agent-interface tools, whose results render in chat). Payloads over ~32KB serialized arrive as `resultTruncated: true` with no data — keep returns compact, or have the UI fetch big data itself. Failed calls deliver nothing to the client (the model gets the `{ error }` and speaks a decline).
47
+
48
+ For backend-side correlation (writing results to a table keyed by the call, custom channels), the method itself can read `session.voiceSessionId` / `session.visitorId` from the agent SDK (`import { session } from '@mindstudio-ai/agent'`) — the same id the browser holds as `session.sessionId`, guaranteed by the platform rather than echoed by the model.
49
+
50
+ Voice sessions also carry `session.medium` (`'web' | 'phone-in' | 'phone-out'`) and, on phone calls, `session.sip` (`{ to, fromNumber }`) — use `medium` in tool methods and the session-context method to branch web vs phone behavior (what to prefetch, how to phrase context). Treat `session.sip.fromNumber` as context only, never identity: caller ID is spoofable, so don't gate data or roles on it — in-call verification is the auth rail.
51
+
52
+ Ordinary web method calls carry the same context (`session.channel === 'web'` with `visitorId`), so correlating a browser's web activity with its voice calls can key on `session.visitorId` on both sides — one browser, one visitor id, web and voice alike. `session.voiceSessionId` remains the per-call key when you need to distinguish individual calls.
112
53
 
113
54
  ### Client tools: actions that happen on screen (`target: "client"`)
114
55
 
115
- A tool whose effect belongs in the browser — open the verification sheet, navigate to a page,
116
- highlight a record — is declared with `target: "client"` instead of a `method`:
56
+ A tool whose effect belongs in the browser — open the verification sheet, navigate to a page, highlight a record — is declared with `target: "client"` instead of a `method`:
117
57
 
118
58
  ```json
119
59
  {
@@ -124,15 +64,10 @@ highlight a record — is declared with `target: "client"` instead of a `method`
124
64
  }
125
65
  ```
126
66
 
127
- The platform never touches the backend for these: the agent's invocation is delivered to the
128
- session's browser, the app's registered handler runs, and the handler's **return value goes back
129
- to the agent as the tool result** — a real request/response, so the agent knows the sheet
130
- actually opened (or that the user dismissed it) and speaks accordingly. Rules:
67
+ The platform never touches the backend for these: the agent's invocation is delivered to the session's browser, the app's registered handler runs, and the handler's **return value goes back to the agent as the tool result** — a real request/response, so the agent knows the sheet actually opened (or that the user dismissed it) and speaks accordingly. Rules:
131
68
 
132
- - `name` instead of `method`; must not collide with any backend method id. No latency class —
133
- the agent holds the turn while the browser responds (up to ~30s, then a timeout error).
134
- - `inputSchema` is authored inline (an object schema) — there's no method contract to derive
135
- it from. Keep it small; these are UI directives, not data payloads.
69
+ - `name` instead of `method`; must not collide with any backend method id. No latency class — the agent holds the turn while the browser responds (up to ~30s, then a timeout error).
70
+ - `inputSchema` is authored inline (an object schema) — there's no method contract to derive it from. Keep it small; these are UI directives, not data payloads.
136
71
  - The frontend must register a handler, or invocations fail as `unhandled_client_tool`:
137
72
 
138
73
  ```js
@@ -142,128 +77,67 @@ session.registerClientTool('showVerification', async ({ reason }) => {
142
77
  });
143
78
  ```
144
79
 
145
- - One client tool runs at a time per session; the description should tell the agent when to use
146
- it and what to say while it's on screen. Throwing from the handler (or returning nothing)
147
- becomes an error/ack the agent can speak around.
148
- - The progressive-auth pattern above is the canonical use: make the verification sheet a client
149
- tool and the agent opens it deliberately instead of the frontend inferring it from tool events.
150
- - Phone sessions never see client tools — there is no browser on a call, so the platform drops
151
- them from the toolset and tells the agent it's on a phone call with no screen. Verification
152
- branches by itself: on a phone the agent gets the platform's in-call verify tools instead of
153
- the app's sheet. Nothing to author; just don't make a client tool the only path to something
154
- phone callers need.
80
+ - One client tool runs at a time per session; the description should tell the agent when to use it and what to say while it's on screen. Throwing from the handler (or returning nothing) becomes an error/ack the agent can speak around.
81
+ - The progressive-auth pattern above is the canonical use: make the verification sheet a client tool and the agent opens it deliberately instead of the frontend inferring it from tool events.
82
+ - Phone sessions never see client tools — there is no browser on a call, so the platform drops them from the toolset and tells the agent it's on a phone call with no screen. Verification branches by itself: on a phone the agent gets the platform's in-call verify tools instead of the app's sheet. Nothing to author; just don't make a client tool the only path to something phone callers need.
155
83
 
156
84
  ### Tool descriptions say results out loud
157
85
 
158
- Follow the agent-interface principles for tool descriptions (when to use and when not, parameter
159
- guidance, what comes back) — plus one voice-specific layer: **how to speak the result.** A tool that
160
- returns a booking record needs its description to say what the confirmation sounds like ("You're all
161
- set for Tuesday at two") and what never gets read aloud (internal ids, timestamps, enum values).
86
+ Follow the agent-interface principles for tool descriptions (when to use and when not, parameter guidance, what comes back) — plus one voice-specific layer: **how to speak the result.** A tool that returns a booking record needs its description to say what the confirmation sounds like ("You're all set for Tuesday at two") and what never gets read aloud (internal ids, timestamps, enum values).
162
87
 
163
- Curate harder than you would for chat. A voice agent with four excellent tools outperforms one with
164
- twelve adequate ones — every tool the model considers is a beat of hesitation. Skip batch
165
- operations, admin utilities, and anything whose output can't be said in a breath or two. Note role
166
- restrictions in the description so the agent declines gracefully in character instead of surfacing a
167
- rejection.
88
+ Curate harder than you would for chat. A voice agent with four excellent tools outperforms one with twelve adequate ones — every tool the model considers is a beat of hesitation. Skip batch operations, admin utilities, and anything whose output can't be said in a breath or two. Note role restrictions in the description so the agent declines gracefully in character instead of surfacing a rejection.
168
89
 
169
90
  ### Confirmation scales with risk
170
91
 
171
- Bake the policy into the system prompt: read-only tools — just call them. Writes — summarize what's
172
- about to happen and get a yes. Anything destructive or financial — read the details back first,
173
- piece by piece. In voice there is no confirmation dialog to lean on; the conversation *is* the
174
- confirmation UI.
92
+ Bake the policy into the system prompt: read-only tools — just call them. Writes — summarize what's about to happen and get a yes. Anything destructive or financial — read the details back first, piece by piece. In voice there is no confirmation dialog to lean on; the conversation *is* the confirmation UI.
175
93
 
176
- And give failure a script: never speak a raw error. When a lookup misses or a tool fails, read back
177
- the value it used ("I couldn't find an order ending three-one-two-five — did I get part of that
178
- wrong?"), offer one retry, then move to an alternate path — in character, without blaming the
179
- caller.
94
+ And give failure a script: never speak a raw error. When a lookup misses or a tool fails, read back the value it used ("I couldn't find an order ending three-one-two-five — did I get part of that wrong?"), offer one retry, then move to an alternate path — in character, without blaming the caller.
180
95
 
181
96
  ### Choosing the model
182
97
 
183
- Two shapes, one `model` field. **Use native speech-to-speech unless the user specifically asks
184
- for a cascaded pipeline** — one realtime model hears and speaks: lowest latency, most natural
185
- prosody, hears tone and hesitation.
98
+ Two shapes, one `model` field. **Use native speech-to-speech unless the user specifically asks for a cascaded pipeline** — one realtime model hears and speaks: lowest latency, most natural prosody, hears tone and hesitation.
186
99
 
187
- **Native** (`{"model": ..., "voice": ...}`) — **default to `gpt-realtime-2.1` with voice
188
- `marin`.**
100
+ **Native** (`{"model": ..., "voice": ...}`) — **default to `gpt-realtime-2.1` with voice `marin`.**
189
101
 
190
- - `gpt-realtime-2.1` — the default. Voices: `marin` (default), `cedar`, `alloy`, `ash`,
191
- `ballad`, `coral`, `echo`, `sage`, `shimmer`, `verse`.
102
+ - `gpt-realtime-2.1` — the default. Voices: `marin` (default), `cedar`, `alloy`, `ash`, `ballad`, `coral`, `echo`, `sage`, `shimmer`, `verse`.
192
103
  - `gpt-realtime-2.1-mini` — the same family, lighter; same voices.
193
- - `gemini-2.5-flash-native-audio-preview-12-2025` — the Gemini pick, with a large expressive
194
- roster: `Puck` (default, upbeat), `Zephyr` (bright), `Charon` (informative), `Kore` (firm),
195
- `Fenrir` (excitable), `Leda` (youthful), `Orus` (firm), `Aoede` (breezy), `Callirrhoe`
196
- (easy-going), `Autonoe` (bright), `Enceladus` (breathy), `Iapetus` (clear), `Umbriel`
197
- (easy-going), `Algieba` (smooth), `Despina` (smooth), `Erinome` (clear), `Algenib` (gravelly),
198
- `Rasalgethi` (informative), `Laomedeia` (upbeat), `Achernar` (soft), `Alnilam` (firm),
199
- `Schedar` (even), `Gacrux` (mature), `Pulcherrima` (forward), `Achird` (friendly),
200
- `Zubenelgenubi` (casual), `Vindemiatrix` (gentle), `Sadachbia` (lively), `Sadaltager`
201
- (knowledgeable), `Sulafat` (warm).
202
- - `gemini-3.1-flash-live-preview` — newer Gemini, same voices as 2.5, but currently can't speak
203
- an opening greeting or take mid-call prompt updates (a plugin limitation expected to resolve
204
- upstream) — prefer 2.5 until then.
205
- - `grok-voice-think-fast-2.0` — a distinct personality register. Voices: `eve` (default),
206
- `altair`, `ara`, `atlas`, `aurora`, `carina`, `castor`, `celeste`, `cosmo`, `helios`, `helix`,
207
- `iris`, `kepler`, `leo`, `liora`, `lumen`, `luna`, `lux`, `naksh`, `orion`, `perseus`, `rex`,
208
- `rigel`, `sal`, `sirius`, `ursa`, `zagan`, `zenith`.
209
-
210
- **Cascaded** (`{"llm": ..., "stt": ..., "tts": ..., "voice": ...}`) — streaming transcription
211
- into any chat model in the catalog, streaming speech out. Slightly higher latency; reach for it
212
- only when the user wants it or the app's reasoning demands a specific chat model (e.g. the agent
213
- interface already uses one and the voice should think identically). Slots: `stt` is
214
- `deepgram-nova-3`; `tts` is `cartesia-sonic-3` (voices are per-account Cartesia UUIDs — see
215
- play.cartesia.ai) or `elevenlabs-tts` (the account's ElevenLabs voice library); `llm` is any chat
216
- model — ask `askMindStudioSdk` for chat model ids. One nuance: cascaded engines speak the
217
- `greeting` verbatim (they have a real TTS); speech-to-speech engines have the model say it, so it
218
- may paraphrase slightly.
219
-
220
- The model and voice ids above are current and maintained with the platform — use them as written
221
- (they are MindStudio ids, not vendor ids). The user's UI has a picker for changing the model
222
- later, so validate only when you set it.
104
+ - `gemini-2.5-flash-native-audio-preview-12-2025` — the Gemini pick, with a large expressive roster: `Puck` (default, upbeat), `Zephyr` (bright), `Charon` (informative), `Kore` (firm), `Fenrir` (excitable), `Leda` (youthful), `Orus` (firm), `Aoede` (breezy), `Callirrhoe` (easy-going), `Autonoe` (bright), `Enceladus` (breathy), `Iapetus` (clear), `Umbriel` (easy-going), `Algieba` (smooth), `Despina` (smooth), `Erinome` (clear), `Algenib` (gravelly), `Rasalgethi` (informative), `Laomedeia` (upbeat), `Achernar` (soft), `Alnilam` (firm), `Schedar` (even), `Gacrux` (mature), `Pulcherrima` (forward), `Achird` (friendly), `Zubenelgenubi` (casual), `Vindemiatrix` (gentle), `Sadachbia` (lively), `Sadaltager` (knowledgeable), `Sulafat` (warm).
105
+ - `gemini-3.1-flash-live-preview` — newer Gemini, same voices as 2.5, but currently can't speak an opening greeting or take mid-call prompt updates (a plugin limitation expected to resolve upstream) — prefer 2.5 until then.
106
+ - `grok-voice-think-fast-2.0` — a distinct personality register. Voices: `eve` (default), `altair`, `ara`, `atlas`, `aurora`, `carina`, `castor`, `celeste`, `cosmo`, `helios`, `helix`, `iris`, `kepler`, `leo`, `liora`, `lumen`, `luna`, `lux`, `naksh`, `orion`, `perseus`, `rex`, `rigel`, `sal`, `sirius`, `ursa`, `zagan`, `zenith`.
107
+
108
+ **Cascaded** (`{"llm": ..., "stt": ..., "tts": ..., "voice": ...}`) — streaming transcription into any chat model in the catalog, streaming speech out. Slightly higher latency; reach for it only when the user wants it or the app's reasoning demands a specific chat model (e.g. the agent interface already uses one and the voice should think identically). Slots: `stt` is `deepgram-nova-3`; `tts` is `cartesia-sonic-3` (voices are per-account Cartesia UUIDs — see play.cartesia.ai) or `elevenlabs-tts` (the account's ElevenLabs voice library); `llm` is any chat model — ask `askMindStudioSdk` for chat model ids. One nuance: cascaded engines speak the `greeting` verbatim (they have a real TTS); speech-to-speech engines have the model say it, so it may paraphrase slightly.
109
+
110
+ The model and voice ids above are current and maintained with the platform — use them as written (they are MindStudio ids, not vendor ids). The user's UI has a picker for changing the model later, so validate only when you set it.
223
111
 
224
112
  ### Seeding from an existing agent
225
113
 
226
- If the app already has an agent interface, start from it: same character, same values, same
227
- terminology — then rewrite for the ear (shorter, spoken-form, no formatting) and re-curate the
228
- toolset for latency. Don't copy `agent.md`'s prose wholesale; a chat persona read aloud sounds like
229
- someone reading chat aloud.
114
+ If the app already has an agent interface, start from it: same character, same values, same terminology — then rewrite for the ear (shorter, spoken-form, no formatting) and re-curate the toolset for latency. Don't copy `agent.md`'s prose wholesale; a chat persona read aloud sounds like someone reading chat aloud.
230
115
 
231
116
  ### Anti-patterns
232
117
 
233
118
  - Prose that would render fine in chat — bullet lists, headers, or markdown anywhere in `system.md`.
234
119
  - A tool description that explains what to display instead of what to say.
235
120
  - Exposing the whole method surface. Voice is the most curated interface the app has.
236
- - A generic greeting ("Hello! How can I assist you today?"). The greeting is the first thing anyone
237
- hears; make it the character's.
238
- - Writing your own current-user placeholder — the platform appends a `## Current User` block
239
- (email, phone, roles) to every system prompt at runtime.
121
+ - A generic greeting ("Hello! How can I assist you today?"). The greeting is the first thing anyone hears; make it the character's.
122
+ - Writing your own current-user placeholder — the platform appends a `## Current User` block (email, phone, roles) to every system prompt at runtime.
240
123
 
241
124
  ## Compiling the Voice Spec
242
125
 
243
- When building `dist/interfaces/voice/`, consider the spec, the app, and the `@brand/` guidelines —
244
- the voice agent should be unmistakably the same product as the web UI, projected into sound. Output:
126
+ When building `dist/interfaces/voice/`, consider the spec, the app, and the `@brand/` guidelines — the voice agent should be unmistakably the same product as the web UI, projected into sound. Output:
245
127
 
246
- **`system.md`** — the persona compiled for the ear. Character first, then the mandatory carries from
247
- "Written for the ear" above (spoken-form rules, brevity, unclear-audio handling, language pinning,
248
- confirmation-by-risk), then any preamble phrasing guidance for `slow` tools so the fillers sound like
249
- the character too.
128
+ **`system.md`** — the persona compiled for the ear. Character first, then the mandatory carries from "Written for the ear" above (spoken-form rules, brevity, unclear-audio handling, language pinning, confirmation-by-risk), then any preamble phrasing guidance for `slow` tools so the fillers sound like the character too.
250
129
 
251
- **`tools/*.md`** — one per tool: when to use, parameter guidance, how to say the result, role
252
- restrictions.
130
+ **`tools/*.md`** — one per tool: when to use, parameter guidance, how to say the result, role restrictions.
253
131
 
254
132
  **`interface.json`** — the config tying it together. Full shape in "The wiring" below.
255
133
 
256
134
  ## Voice UI
257
135
 
258
- When the app has a web interface, voice arrives as a **layer over it**, not a separate page: a
259
- persistent affordance (a button, an orb in a corner) that starts a session in place, with the app
260
- still visible and usable. A dedicated full-screen voice mode is the immersive option for apps where
261
- the conversation *is* the product — earn it, don't default to it.
136
+ When the app has a web interface, voice arrives as a **layer over it**, not a separate page: a persistent affordance (a button, an orb in a corner) that starts a session in place, with the app still visible and usable. A dedicated full-screen voice mode is the immersive option for apps where the conversation *is* the product — earn it, don't default to it.
262
137
 
263
138
  ### Frontend SDK: `createVoiceClient()`
264
139
 
265
- Ships as a subpath of the interface SDK so apps that never use voice pay nothing for it. All voice
266
- UIs go through it — never hand-roll audio capture or transport.
140
+ Ships as a subpath of the interface SDK so apps that never use voice pay nothing for it. All voice UIs go through it — never hand-roll audio capture or transport.
267
141
 
268
142
  ```ts
269
143
  import { createVoiceClient } from '@mindstudio-ai/interface/voice';
@@ -293,49 +167,37 @@ await session.refreshIdentity(); // after in-app verification — upgrade t
293
167
  session.end();
294
168
  ```
295
169
 
296
- Agent audio playback is handled inside the SDK (a hidden autoplaying element) — never create audio
297
- elements for the agent. `startSession()` throws `MindStudioInterfaceError` with code
298
- `microphone_denied` when mic access is refused (surface that state gently in the UI),
299
- `voice_concurrency_limit` / `voice_visitor_limit` when the app's session limits are hit, and
300
- `auth_required` (401) / `role_required` (403) when the interface's `auth` block denies the caller
301
- (route those to the app's login flow).
170
+ Agent audio playback is handled inside the SDK (a hidden autoplaying element) — never create audio elements for the agent. `startSession()` throws `MindStudioInterfaceError` with code `microphone_denied` when mic access is refused (surface that state gently in the UI), `voice_concurrency_limit` / `voice_visitor_limit` when the app's session limits are hit, and `auth_required` (401) / `role_required` (403) when the interface's `auth` block denies the caller (route those to the app's login flow).
302
171
 
303
- Past sessions are call records with transcripts: `voice.listSessions()` /
304
- `voice.getSession(id)` — the material for a history view if the app wants one.
172
+ Past sessions are call records with transcripts: `voice.listSessions()` / `voice.getSession(id)` — the material for a history view if the app wants one.
305
173
 
306
- ### The state machine, made visible
174
+ ### The voice agent, on screen
307
175
 
308
- One audio-reactive element carries the session: idle → connecting → listening → thinking → speaking.
309
- Always pair it with a **text state label** — never signal state by color or motion alone. Calm at
310
- idle, responsive to actual audio levels while listening and speaking. Respect
311
- `prefers-reduced-motion` with a static-but-labeled variant.
176
+ One audio-reactive element carries the session state machine: idle → connecting → listening → thinking → speaking. It is on screen for the entire conversation and it is the most visual thing about a voice app — treat it as a **first-class design deliverable**, not a widget. Bring in `visualDesignExpert` for the session experience as a whole — the centerpiece, the captions, how tool activity surfaces, and the controls, composed as one scene (it has a dedicated craft reference for exactly this) — and implement what it prescribes. The bar is a real-time computed piece — WebGL, structured, alive, derived from the app's brand (a sampled point-cloud object, an instrument/meter, whatever the domain suggests). A generic gradient sphere or a pulsing CSS circle is a defaulted artifact, not a designed one.
177
+
178
+ Whatever the designer specs, the mechanics you own: always pair the piece with a **text state label** — never signal state by color or motion alone; calm at idle, responsive to actual audio levels while listening and speaking. And build the performance budget in from day one: cap `devicePixelRatio` at 2, reduce density on mobile, pause the render loop when the element is off-screen or the tab is hidden, and ship a static-but-labeled fallback for `prefers-reduced-motion` and no-WebGL.
312
179
 
313
180
  ### Live captions
314
181
 
315
- Stream `transcript` events as captions — both sides of the conversation. Captions make the agent
316
- feel accurate, catch mishearings early, and are the accessibility story. User-side transcripts
317
- arrive as recognition output and can lag or differ slightly from what the model heard; render them
318
- as captions, never treat them as input to app logic.
182
+ Stream `transcript` events as captions — both sides of the conversation. Captions make the agent feel accurate, catch mishearings early, and are the accessibility story. User-side transcripts arrive as recognition output and can lag or differ slightly from what the model heard; render them as captions, never treat them as input to app logic.
183
+
184
+ Captions must be **layout-stable**: reserve a fixed-height caption region so arriving text never shifts the layout around it (a streaming segment that reflows the page on every event reads as jank, and it fights the agent's visual for attention). Upsert each segment in place by `segmentId` — the event carries the segment's full text, so replacing is free — cap the visible lines, and fade old lines out rather than pushing content down.
319
185
 
320
186
  ### Controls that must exist
321
187
 
322
- **Mute** and **end call**, always visible, always working. `sendText` earns its place the moment the
323
- conversation needs an exact string — an address, a code, an email — typing it beats spelling it
324
- aloud three times. Show tool activity as a compact inline status from `toolCall` events, in the
325
- app's voice ("Booking your appointment…"), never raw names or JSON.
188
+ **Mute** and **end call**, always visible, always working. `sendText` earns its place the moment the conversation needs an exact string — an address, a code, an email — typing it beats spelling it aloud three times. Show tool activity as a compact inline status from `toolCall` events, in the app's voice ("Booking your appointment…"), never raw names or JSON.
326
189
 
327
190
  ### Anti-patterns
328
191
 
329
192
  - Blocking the whole UI behind the session — voice is a layer, the app stays usable.
193
+ - A generic gradient sphere or pulsing CSS circle as the voice agent's visual. The visual is designed (by the design expert), not defaulted.
330
194
  - An orb with no label, or state changes conveyed only by color.
331
195
  - Rendering user-side captions as authoritative ("you said X") — they're recognition output.
332
196
  - Auto-starting a session on page load. Microphone access is always a deliberate user action.
333
197
 
334
198
  ## Outbound calls (`voice.call`)
335
199
 
336
- The agent can call the user. Backend methods (and crons) place outbound phone calls with the
337
- agent SDK's `voice` namespace — the platform dials the number and connects the callee to this
338
- app's voice agent (same persona, engine, and tools as the web sessions):
200
+ The agent can call the user. Backend methods (and crons) place outbound phone calls with the agent SDK's `voice` namespace — the platform dials the number and connects the callee to this app's voice agent (same persona, engine, and tools as the web sessions):
339
201
 
340
202
  ```ts
341
203
  import { voice, auth } from '@mindstudio-ai/agent';
@@ -347,84 +209,36 @@ export async function callMeAboutMyOrder(input: { phone: string }) {
347
209
  }
348
210
  ```
349
211
 
350
- - **The method is the authorization gate.** The voice interface's `auth` block does not apply to
351
- calls the backend places deliberately — gate the *method* with `auth.requireRole(...)` exactly
352
- as you would any sensitive action.
353
- - **`assumeIdentity: true`** runs the call as the user who invoked the method: the agent knows
354
- who it's talking to (Current User block) and every tool call carries their roles — regardless
355
- of which number was dialed (the user types any number into a field; identity comes from their
356
- session, not the phone). Omitted/false → anonymous call; role-gated tools decline.
357
- System/cron invocations have no human identity and always run anonymously. Anonymous outbound
358
- calls (deployed) get the same in-call verification flow as inbound — the callee proves
359
- possession of the number that was dialed, or verifies by email — so an anonymous call can
360
- still upgrade to a known user mid-conversation.
361
- - **Production needs a dedicated phone number.** The app owner attaches one ($1/month) via the
362
- dashboard or `mindstudio-prod voice numbers` (see "The voice CLI"
363
- below) — it becomes the caller ID for every call, in dev sessions too, so users always see
364
- the same number. Without one, deployed calls throw `phone_out_requires_dedicated_number`, and
365
- dev sessions fall back to a shared platform test number that varies per call (tighter limits
366
- apply on the shared pool).
367
- - **Outcome is on the call record**, not the return value: `voice.call` returns as soon as
368
- dialing starts (`{ sessionId, status: 'dialing', from, to }`); answered/busy/no-answer land on
369
- the session in the app's call log (`voice.listSessions()` / the dashboard).
370
- - **Limits**: the app's concurrent-session policy, a daily outbound-call cap, a per-call
371
- duration ceiling, and one active call per callee number (`voice_callee_busy`).
372
- - **Compliance**: automated calls require prior consent. Call your own users who opted in to
373
- calls from this app, honor reasonable calling hours, never dial purchased or cold lists —
374
- design the consent moment into the product (a "call me" button IS consent; a scraped list is
375
- not).
212
+ - **The method is the authorization gate.** The voice interface's `auth` block does not apply to calls the backend places deliberately — gate the *method* with `auth.requireRole(...)` exactly as you would any sensitive action.
213
+ - **`assumeIdentity: true`** runs the call as the user who invoked the method: the agent knows who it's talking to (Current User block) and every tool call carries their roles — regardless of which number was dialed (the user types any number into a field; identity comes from their session, not the phone). Omitted/false → anonymous call; role-gated tools decline. System/cron invocations have no human identity and always run anonymously. Anonymous outbound calls (deployed) get the same in-call verification flow as inbound — the callee proves possession of the number that was dialed, or verifies by email — so an anonymous call can still upgrade to a known user mid-conversation.
214
+ - **Production needs a dedicated phone number.** The app owner attaches one ($1/month) via the dashboard or `mindstudio-prod voice numbers` (see "The voice CLI" below) — it becomes the caller ID for every call, in dev sessions too, so users always see the same number. Without one, deployed calls throw `phone_out_requires_dedicated_number`, and dev sessions fall back to a shared platform test number that varies per call (tighter limits apply on the shared pool).
215
+ - **Outcome is on the call record**, not the return value: `voice.call` returns as soon as dialing starts (`{ sessionId, status: 'dialing', from, to }`); answered/busy/no-answer land on the session in the app's call log (`voice.listSessions()` / the dashboard).
216
+ - **Limits**: the app's concurrent-session policy, a daily outbound-call cap, a per-call duration ceiling, and one active call per callee number (`voice_callee_busy`).
217
+ - **Compliance**: automated calls require prior consent. Call your own users who opted in to calls from this app, honor reasonable calling hours, never dial purchased or cold lists — design the consent moment into the product (a "call me" button IS consent; a scraped list is not).
376
218
 
377
219
  ## Inbound calls
378
220
 
379
- Once the app has a dedicated phone number, people can call it — the same voice agent answers
380
- (same persona, engine, and tools). Nothing extra to author for the basic case; the number in the
381
- app's settings is the whole switch.
221
+ Once the app has a dedicated phone number, people can call it — the same voice agent answers (same persona, engine, and tools). Nothing extra to author for the basic case; the number in the app's settings is the whole switch.
382
222
 
383
223
  How answering works:
384
224
 
385
- - **Inbound always runs the live release.** There is no dev inbound — test the agent over the
386
- normal WebRTC session in the editor; the phone is the same interface with a different
387
- transport. An app with no live voice interface (or at its concurrency limit) doesn't answer.
388
- - **Callers are anonymous until verified.** The `auth` block still applies, but a phone call
389
- can't show a login page — so the platform answers first, and `requireUser` becomes an
390
- in-call verification flow. The agent can serve whatever anonymous callers are allowed, and
391
- offers verification when the caller wants something account-bound.
392
- - **Verification uses the app's own auth methods** (`sms-code` / `email-code` from the
393
- manifest):
394
- - SMS: a code is texted to the phone number on the call (no other number is possible by
395
- design), and confirming it signs the caller in — **creating their account if they're new**,
396
- the same find-or-create policy `sms-code` has on web (enabling the method is what enables
397
- sign-up; there is no separate toggle on either surface). After a first-time caller verifies,
398
- the session-context method re-fires with the fresh identity — that's the hook for seeding a
399
- new account with data.
400
- - Email: existing accounts only — the caller says their address; the platform matches it
401
- against the app's users (transcription-tolerant — no letter-by-letter spelling ceremony) and
402
- emails the account's stored address a code. There is no sign-up by email over the phone: a
403
- call can't reliably capture a verbatim never-seen address, so new callers sign up by text
404
- instead.
405
- - The email flow never confirms or denies that an account exists — a code is "sent if an
406
- account matches", always phrased that neutrally. The persona should offer verification
407
- naturally when it unlocks something, never as a robotic gate.
408
- - **Verified mid-call, upgraded mid-call**: once the code checks out, the session becomes that
409
- user's — Current User block, roles on every tool call — without redialing.
225
+ - **Inbound always runs the live release.** There is no dev inbound — test the agent over the normal WebRTC session in the editor; the phone is the same interface with a different transport. An app with no live voice interface (or at its concurrency limit) doesn't answer.
226
+ - **Callers are anonymous until verified.** The `auth` block still applies, but a phone call can't show a login page — so the platform answers first, and `requireUser` becomes an in-call verification flow. The agent can serve whatever anonymous callers are allowed, and offers verification when the caller wants something account-bound.
227
+ - **Verification uses the app's own auth methods** (`sms-code` / `email-code` from the manifest):
228
+ - SMS: a code is texted to the phone number on the call (no other number is possible by design), and confirming it signs the caller in — **creating their account if they're new**, the same find-or-create policy `sms-code` has on web (enabling the method is what enables sign-up; there is no separate toggle on either surface). After a first-time caller verifies, the session-context method re-fires with the fresh identity — that's the hook for seeding a new account with data.
229
+ - Email: existing accounts only — the caller says their address; the platform matches it against the app's users (transcription-tolerant — no letter-by-letter spelling ceremony) and emails the account's stored address a code. There is no sign-up by email over the phone: a call can't reliably capture a verbatim never-seen address, so new callers sign up by text instead.
230
+ - The email flow never confirms or denies that an account exists — a code is "sent if an account matches", always phrased that neutrally. The persona should offer verification naturally when it unlocks something, never as a robotic gate.
231
+ - **Verified mid-call, upgraded mid-call**: once the code checks out, the session becomes that user's — Current User block, roles on every tool call — without redialing.
410
232
 
411
233
  ### `phone.trustCallerId`
412
234
 
413
- For apps whose users are known by phone number, the interface config may opt into treating
414
- caller ID as identity:
235
+ For apps whose users are known by phone number, the interface config may opt into treating caller ID as identity:
415
236
 
416
237
  ```json
417
238
  "phone": { "trustCallerId": true }
418
239
  ```
419
240
 
420
- A caller whose number exactly matches an app user's phone starts the call already verified —
421
- no code. This is a real security tradeoff: **caller ID can be spoofed**, so a motivated
422
- attacker who knows a user's phone number can impersonate them to this agent. Before enabling
423
- it, you MUST surface that risk to the user and get their explicit confirmation — it's the
424
- right call for convenience-first, low-stakes apps (a family assistant, a status line), and the
425
- wrong one wherever the agent's tools can move money, reveal sensitive records, or take
426
- destructive actions. It lives in the interface config deliberately: enabling it is a code
427
- change, visible in review and auditable via deploys, not a dashboard toggle.
241
+ A caller whose number exactly matches an app user's phone starts the call already verified — no code. This is a real security tradeoff: **caller ID can be spoofed**, so a motivated attacker who knows a user's phone number can impersonate them to this agent. Before enabling it, you MUST surface that risk to the user and get their explicit confirmation — it's the right call for convenience-first, low-stakes apps (a family assistant, a status line), and the wrong one wherever the agent's tools can move money, reveal sensitive records, or take destructive actions. It lives in the interface config deliberately: enabling it is a code change, visible in review and auditable via deploys, not a dashboard toggle.
428
242
 
429
243
  ## The voice CLI
430
244
 
@@ -438,18 +252,11 @@ mindstudio-prod voice sessions list --limit 10 # call log: web / phone-o
438
252
  mindstudio-prod voice sessions get <sessionId> # full transcript + cost breakdown
439
253
  ```
440
254
 
441
- Also `voice numbers list`, `voice numbers set-name` (outbound caller-ID display
442
- name; 12-72h carrier propagation), `voice settings get`/`set` (concurrency, per-visitor,
443
- max duration — `set` merges: only the settings you pass change). `--help` for flags.
255
+ Also `voice numbers list`, `voice numbers set-name` (outbound caller-ID display name; 12-72h carrier propagation), `voice settings get`/`set` (concurrency, per-visitor, max duration — `set` merges: only the settings you pass change). `--help` for flags.
444
256
 
445
- **Never buy a number without the user's explicit confirmation** — it starts a recurring
446
- $1/month workspace charge. Search first, present the options with the price, and only run
447
- `numbers buy` after they've picked one and said yes.
257
+ **Never buy a number without the user's explicit confirmation** — it starts a recurring $1/month workspace charge. Search first, present the options with the price, and only run `numbers buy` after they've picked one and said yes.
448
258
 
449
- Transcripts are how you iterate on a voice persona: after the user test-calls the agent, read
450
- `voice sessions get` for what was actually said — misheard input, interruptions, tools declining
451
- — and fix the spec from evidence rather than guesses. (Dev-session test calls carry a
452
- `devSessionId` in the list, so you can tell them from live traffic.)
259
+ Transcripts are how you iterate on a voice persona: after the user test-calls the agent, read `voice sessions get` for what was actually said — misheard input, interruptions, tools declining — and fix the spec from evidence rather than guesses. (Dev-session test calls carry a `devSessionId` in the list, so you can tell them from live traffic.)
453
260
 
454
261
  ---
455
262
 
@@ -457,8 +264,7 @@ Transcripts are how you iterate on a voice persona: after the user test-calls th
457
264
 
458
265
  ## Spec: `src/interfaces/voice.md`
459
266
 
460
- Frontmatter holds the structured fields; the body is the persona plus an explicit `## Tools`
461
- section.
267
+ Frontmatter holds the structured fields; the body is the persona plus an explicit `## Tools` section.
462
268
 
463
269
  ```yaml
464
270
  ---
@@ -475,15 +281,9 @@ Frontmatter fields:
475
281
 
476
282
  - `name` — display name
477
283
  - `description` — one-liner for listings
478
- - `model` — JSON string, two shapes: native speech-to-speech `{"model": <realtime model id>,
479
- "voice": <voice id>}`, or cascaded `{"llm": <chat model id>, "stt": <transcription model id>,
480
- "tts": <speech model id>, "voice": <voice id>}`. Ids via `askMindStudioSdk`.
481
- - `turnDetection` — optional; `{"eagerness": "low" | "medium" | "high"}` — how quickly the platform
482
- decides the user finished speaking. High is snappier; low is more patient (users dictating
483
- numbers or addresses). Default `medium`.
484
- - `greeting` — optional spoken opener, delivered on session start. Omit and the agent waits for the
485
- user to speak first. Verbatim on cascaded engines; model-spoken (may paraphrase) on
486
- speech-to-speech.
284
+ - `model` — JSON string, two shapes: native speech-to-speech `{"model": <realtime model id>, "voice": <voice id>}`, or cascaded `{"llm": <chat model id>, "stt": <transcription model id>, "tts": <speech model id>, "voice": <voice id>}`. Ids via `askMindStudioSdk`.
285
+ - `turnDetection` — optional; `{"eagerness": "low" | "medium" | "high"}` — how quickly the platform decides the user finished speaking. High is snappier; low is more patient (users dictating numbers or addresses). Default `medium`.
286
+ - `greeting` — optional spoken opener, delivered on session start. Omit and the agent waits for the user to speak first. Verbatim on cascaded engines; model-spoken (may paraphrase) on speech-to-speech.
487
287
 
488
288
  Body: persona prose (voice register), then the toolset:
489
289
 
@@ -501,8 +301,7 @@ booking id aloud unless asked.
501
301
  ~~~
502
302
  ```
503
303
 
504
- `latency` is one of `fast` / `slow` / `background` (semantics in "The latency classes" above).
505
- Don't hand-author input schemas — the platform derives them from the method contract.
304
+ `latency` is one of `fast` / `slow` / `background` (semantics in "The latency classes" above). Don't hand-author input schemas — the platform derives them from the method contract.
506
305
 
507
306
  ## Compiled Output: `dist/interfaces/voice/`
508
307
 
@@ -560,95 +359,47 @@ Declare it in `mindstudio.json`:
560
359
 
561
360
  ## Session context (auto-loaded)
562
361
 
563
- When the config declares `"context": { "method": "session-context" }`, the platform fires that
564
- backend method automatically when a session starts — in the background, so the greeting is
565
- never delayed — and appends its return to the system prompt as a `## Session Context` block.
566
- Use it for situational state that should color every turn: the caller's open orders, account
567
- standing, where they left off. Timing: the method runs while the greeting audio plays, so the
568
- agent has the context by roughly the first exchange and is guaranteed to have it shortly
569
- after — it is NOT guaranteed for the literal first utterance. On inbound phone calls it
570
- re-fires after the caller verifies mid-call, so the context recomputes for the now-known user.
362
+ When the config declares `"context": { "method": "session-context" }`, the platform fires that backend method automatically when a session starts — in the background, so the greeting is never delayed — and appends its return to the system prompt as a `## Session Context` block. Use it for situational state that should color every turn: the caller's open orders, account standing, where they left off. Timing: the method runs while the greeting audio plays, so the agent has the context by roughly the first exchange and is guaranteed to have it shortly after — it is NOT guaranteed for the literal first utterance. On inbound phone calls it re-fires after the caller verifies mid-call, so the context recomputes for the now-known user.
571
363
 
572
364
  The method contract:
573
- - Runs as the session's user (same identity/RBAC as a tool call); anonymous sessions run it
574
- anonymously — return generic or empty content for them.
575
- - Return a short markdown **string** (a few lines). Results are capped at 4,000 characters;
576
- keep it situational context, not documents — deep or on-demand data belongs in tools or
577
- data sources.
578
- - It is never a model-visible tool, and failures degrade silently to the generic prompt —
579
- never make correctness depend on it.
580
-
581
- Rule of thumb: `context` for always-relevant state the agent should just know; tools for
582
- anything looked up on demand. Identity itself (name, roles) is already injected via the
583
- Current User block — don't re-fetch it in the context method.
365
+ - Runs as the session's user (same identity/RBAC as a tool call); anonymous sessions run it anonymously — return generic or empty content for them.
366
+ - Return a short markdown **string** (a few lines). Results are capped at 4,000 characters; keep it situational context, not documents — deep or on-demand data belongs in tools or data sources.
367
+ - It is never a model-visible tool, and failures degrade silently to the generic prompt — never make correctness depend on it.
368
+
369
+ Rule of thumb: `context` for always-relevant state the agent should just know; tools for anything looked up on demand. Identity itself (name, roles) is already injected via the Current User block — don't re-fetch it in the context method.
584
370
 
585
371
  ## Platform Behavior
586
372
 
587
373
  - Input schemas are derived from each method's contract — never hand-written.
588
- - The platform appends a `## Current User` block (email, phone, roles) to the system prompt at
589
- runtime; never author a placeholder for it.
590
- - Turn detection, barge-in (interruption truncates the agent's context to the audio the user
591
- actually heard), and background-noise handling are platform-managed; `turnDetection.eagerness` is
592
- the only knob (not yet wired on Gemini realtime engines — it's a no-op there).
593
- - Sessions have a per-app concurrency limit and a maximum duration, both configurable in the app's
594
- settings; an idle session is ended gracefully after a prompt. Voice minutes and model usage are
595
- metered.
596
- - Every session persists as a call record with a transcript, visible in the dashboard and readable
597
- from the frontend via `voice.listSessions()` / `voice.getSession(id)`.
374
+ - The platform appends a `## Current User` block (email, phone, roles) to the system prompt at runtime; never author a placeholder for it.
375
+ - Turn detection, barge-in (interruption truncates the agent's context to the audio the user actually heard), and background-noise handling are platform-managed; `turnDetection.eagerness` is the only knob (not yet wired on Gemini realtime engines — it's a no-op there).
376
+ - Sessions have a per-app concurrency limit and a maximum duration, both configurable in the app's settings; an idle session is ended gracefully after a prompt. Voice minutes and model usage are metered.
377
+ - Every session persists as a call record with a transcript, visible in the dashboard and readable from the frontend via `voice.listSessions()` / `voice.getSession(id)`.
598
378
 
599
379
  ## Auth
600
380
 
601
- **Every voice config declares an `auth` block.** A voice session spends the owner's money for its
602
- entire duration without necessarily touching a backend method, so the platform gates session
603
- creation itself:
381
+ **Every voice config declares an `auth` block.** A voice session spends the owner's money for its entire duration without necessarily touching a backend method, so the platform gates session creation itself:
604
382
 
605
383
  ```json
606
384
  "auth": { "requireUser": true, "requireRole": ["member"] }
607
385
  ```
608
386
 
609
- - `requireUser: true` — only authenticated app users may start a session; `false` — anyone,
610
- including anonymous visitors. Most apps want `true`; choose `false` deliberately (a public
611
- front-desk line).
612
- - `requireRole` (optional) — the user must hold **at least one** of the listed manifest role ids
613
- (OR semantics, same as the backend `auth.requireRole(...)`). Omit or leave empty for no role
614
- gate. Requires `requireUser: true`. Unknown role ids fail the build.
387
+ - `requireUser: true` — only authenticated app users may start a session; `false` — anyone, including anonymous visitors. Most apps want `true`; choose `false` deliberately (a public front-desk line).
388
+ - `requireRole` (optional) — the user must hold **at least one** of the listed manifest role ids (OR semantics, same as the backend `auth.requireRole(...)`). Omit or leave empty for no role gate. Requires `requireUser: true`. Unknown role ids fail the build.
615
389
  - Denials reject `startSession()` with code `auth_required` (401) or `role_required` (403).
616
- - On the phone channel there is no login page to bounce to, so `requireUser` becomes
617
- answer-then-verify — see "Inbound calls" above.
390
+ - On the phone channel there is no login page to bounce to, so `requireUser` becomes answer-then-verify — see "Inbound calls" above.
618
391
  - Dev preview is exempt — the builder is never locked out while testing.
619
- - Older compiled apps without the block fall back to the manifest's `auth.enabled` (auth-enabled →
620
- users only; no auth → public). New configs always declare it explicitly.
392
+ - Older compiled apps without the block fall back to the manifest's `auth.enabled` (auth-enabled → users only; no auth → public). New configs always declare it explicitly.
621
393
 
622
- Once inside, voice sessions run as the **authenticated user** — every tool call carries that
623
- user's roles, so a method gated with `auth.requireRole` behaves exactly as it would from the web
624
- frontend or the agent interface. Anonymous sessions (when allowed) have no user and no roles:
625
- gated methods reject, and the caller's history is scoped to their browser's visitor identity.
626
- That's why role restrictions belong in the tool descriptions — the agent should decline in
627
- character, not relay a rejection.
394
+ Once inside, voice sessions run as the **authenticated user** — every tool call carries that user's roles, so a method gated with `auth.requireRole` behaves exactly as it would from the web frontend or the agent interface. Anonymous sessions (when allowed) have no user and no roles: gated methods reject, and the caller's history is scoped to their browser's visitor identity. That's why role restrictions belong in the tool descriptions — the agent should decline in character, not relay a rejection.
628
395
 
629
396
  ### Progressive auth: verify mid-call without dropping the conversation
630
397
 
631
- The best pattern for apps that allow anonymous sessions (`requireUser: false`): let visitors
632
- explore by voice, and verify only when they hit an account-bound action — without killing the
633
- live call. Four pieces, all platform rails:
634
-
635
- 1. **Account-gated tools return a standard not-verified shape** instead of doing the work:
636
- `{ verified: false, message: 'The caller is not verified. Offer to verify them before sharing
637
- account details.' }`. The agent speaks the offer in character (reinforce tone in the system
638
- prompt's verification section). Check with the agent SDK's `auth.userId` inside the method.
639
- 2. **The frontend opens its verification sheet off the same signal.** It already receives every
640
- tool's `toolCall` event (and the tool's return in `result`) — when an account tool fires (or
641
- returns `verified: false`) while the app has no signed-in user, open the sheet.
642
- 3. **The sheet runs the platform's auth rails** — `auth.sendSmsCode()` / `auth.verifySmsCode()`
643
- (or the email pair) from `@mindstudio-ai/interface`. On success the app's session becomes the
644
- verified user.
645
- 4. **Hand the verified session back to the live call**: `await session.refreshIdentity()`. The
646
- platform upgrades the running voice session in place — subsequent tool calls carry the user's
647
- identity and roles, and the agent's Current User context refreshes — no teardown, no lost
648
- conversation. (Phone calls don't need this: they verify through the agent's built-in flow.)
649
-
650
- `refreshIdentity()` is upgrade-only (anonymous → signed-in; an already-identified session rejects
651
- with `already_identified`) and requires the session to have been started by this same browser. If
652
- it fails, ending and restarting the session is the graceful fallback. The chat sibling for agent
653
- interfaces is `claimThread(threadId)` — anonymous threads become unreachable after login until
654
- claimed.
398
+ The best pattern for apps that allow anonymous sessions (`requireUser: false`): let visitors explore by voice, and verify only when they hit an account-bound action — without killing the live call. Four pieces, all platform rails:
399
+
400
+ 1. **Account-gated tools return a standard not-verified shape** instead of doing the work: `{ verified: false, message: 'The caller is not verified. Offer to verify them before sharing account details.' }`. The agent speaks the offer in character (reinforce tone in the system prompt's verification section). Check with the agent SDK's `auth.userId` inside the method.
401
+ 2. **The frontend opens its verification sheet off the same signal.** It already receives every tool's `toolCall` event (and the tool's return in `result`) — when an account tool fires (or returns `verified: false`) while the app has no signed-in user, open the sheet.
402
+ 3. **The sheet runs the platform's auth rails** — `auth.sendSmsCode()` / `auth.verifySmsCode()` (or the email pair) from `@mindstudio-ai/interface`. On success the app's session becomes the verified user.
403
+ 4. **Hand the verified session back to the live call**: `await session.refreshIdentity()`. The platform upgrades the running voice session in place — subsequent tool calls carry the user's identity and roles, and the agent's Current User context refreshes — no teardown, no lost conversation. (Phone calls don't need this: they verify through the agent's built-in flow.)
404
+
405
+ `refreshIdentity()` is upgrade-only (anonymous → signed-in; an already-identified session rejects with `already_identified`) and requires the session to have been started by this same browser. If it fails, ending and restarting the session is the graceful fallback. The chat sibling for agent interfaces is `claimThread(threadId)` — anonymous threads become unreachable after login until claimed.