@mindstudio-ai/remy 0.1.276 → 0.1.277
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/headless.js +112 -33
- package/dist/index.js +135 -35
- package/dist/prompt/skills/agentInterfaces.md +13 -39
- package/dist/prompt/skills/auth.md +1 -7
- package/dist/prompt/skills/dataSources.md +20 -66
- package/dist/prompt/skills/files.md +17 -47
- package/dist/prompt/skills/inboundEmail.md +11 -32
- package/dist/prompt/skills/mcpInterfaces.md +21 -63
- package/dist/prompt/skills/restApi.md +7 -21
- package/dist/prompt/skills/scenarios.md +3 -7
- package/dist/prompt/skills/scheduledJobs.md +3 -8
- package/dist/prompt/skills/voiceInterfaces.md +117 -366
- package/dist/prompt/skills/webhooks.md +10 -31
- package/dist/prompt/static/authoring.md +1 -2
- package/dist/subagents/browserAutomation/prompt.md +6 -22
- package/dist/subagents/codeSanityCheck/prompt.md +1 -2
- package/dist/subagents/designExpert/prompt.md +0 -3
- package/dist/subagents/designExpert/prompts/ui-patterns.md +7 -21
- package/dist/subagents/designExpert/skills/authExperience.md +60 -0
- package/dist/subagents/designExpert/skills/chatExperience.md +59 -0
- package/dist/subagents/designExpert/skills/dataViz.md +82 -0
- package/dist/subagents/designExpert/{prompts → skills}/images.md +6 -0
- package/dist/subagents/designExpert/skills/voiceExperience.md +103 -0
- package/dist/subagents/productVision/prompt.md +2 -0
- package/package.json +1 -1
|
@@ -6,114 +6,54 @@ when: Before authoring `src/interfaces/voice.md`, choosing a voice model or pipe
|
|
|
6
6
|
|
|
7
7
|
# Building Voice Interfaces
|
|
8
8
|
|
|
9
|
-
A voice interface is the app's agent as a live phone-call-quality conversation: the user speaks, the
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
into conversation; load the `agentInterfaces` skill for that shared ground), but everything you
|
|
13
|
-
author differs. The persona is written for the ear, not the screen. The toolset is smaller and
|
|
14
|
-
curated for conversational latency. And every tool declares how the agent should handle the wait,
|
|
15
|
-
because in a live call, silence reads as a dropped line.
|
|
16
|
-
|
|
17
|
-
The platform owns the hard parts — realtime audio transport, turn detection, interruption handling,
|
|
18
|
-
transcripts, session limits, per-user auth on every tool call. Your job is the spec
|
|
19
|
-
(`src/interfaces/voice.md`) and its compilation into `dist/interfaces/voice/`.
|
|
9
|
+
A voice interface is the app's agent as a live phone-call-quality conversation: the user speaks, the agent answers in speech, and the app's methods are its tools. It is a **sibling of the agent interface, not a mode of it** — the two share a philosophy (an LLM projecting the backend contract into conversation; load the `agentInterfaces` skill for that shared ground), but everything you author differs. The persona is written for the ear, not the screen. The toolset is smaller and curated for conversational latency. And every tool declares how the agent should handle the wait, because in a live call, silence reads as a dropped line.
|
|
10
|
+
|
|
11
|
+
The platform owns the hard parts — realtime audio transport, turn detection, interruption handling, transcripts, session limits, per-user auth on every tool call. Your job is the spec (`src/interfaces/voice.md`) and its compilation into `dist/interfaces/voice/`.
|
|
20
12
|
|
|
21
13
|
## Voice Agent Design
|
|
22
14
|
|
|
23
15
|
### Written for the ear
|
|
24
16
|
|
|
25
|
-
Everything the agent produces gets spoken aloud. That inverts several habits that are correct
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
- **
|
|
29
|
-
|
|
30
|
-
- **
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
a time. A paragraph that reads fine in chat is a monologue on a call.
|
|
36
|
-
- **Vary the phrasing.** Repeated openers and acknowledgments sound convincing once and robotic by
|
|
37
|
-
the third turn — give the prompt an explicit variety rule, and treat any sample phrases as
|
|
38
|
-
anchors, never scripts.
|
|
39
|
-
- **Handle unclear audio explicitly.** Give the prompt a rule for it: respond only to clear audio;
|
|
40
|
-
if it's noisy or ambiguous, ask the user to repeat — never guess, never call a tool on input the
|
|
41
|
-
agent isn't sure it heard, and don't reuse the same clarification line twice in a row.
|
|
42
|
-
- **Pin the language.** State the response language in the prompt; don't let the model infer it from
|
|
43
|
-
an accent. If the app's domain has brand names or terms with non-obvious pronunciations, give
|
|
44
|
-
them a line ("pronounce SQL as 'sequel'").
|
|
45
|
-
|
|
46
|
-
Beyond the mechanics, the persona itself should be *of the ear*: pacing, warmth, how it handles
|
|
47
|
-
being interrupted, what it says when it needs a second. This is the fun part, same as the agent
|
|
48
|
-
interface — a distinct character beats a generic assistant, and voice makes character land harder
|
|
49
|
-
than any other surface.
|
|
17
|
+
Everything the agent produces gets spoken aloud. That inverts several habits that are correct everywhere else, and the compiled system prompt must carry them explicitly:
|
|
18
|
+
|
|
19
|
+
- **No visual formatting, ever.** No markdown, no lists, no tables, no emoji, no URLs read as punctuation soup. If a tool returns a link, say what it is and where it will be, don't recite it.
|
|
20
|
+
- **Spoken-form values.** "Forty-two fifty," not "$42.50". "Two fifteen in the afternoon," not "14:15". Read email addresses and confirmation codes character by character, and read them *back* for confirmation before acting on them — mishearing one digit of a phone number is the classic voice failure. Collect one value per turn; two asked together blend when spoken.
|
|
21
|
+
- **Brevity is a hard rule, not a style preference.** One to two sentences per turn, one question at a time. A paragraph that reads fine in chat is a monologue on a call.
|
|
22
|
+
- **Vary the phrasing.** Repeated openers and acknowledgments sound convincing once and robotic by the third turn — give the prompt an explicit variety rule, and treat any sample phrases as anchors, never scripts.
|
|
23
|
+
- **Handle unclear audio explicitly.** Give the prompt a rule for it: respond only to clear audio; if it's noisy or ambiguous, ask the user to repeat — never guess, never call a tool on input the agent isn't sure it heard, and don't reuse the same clarification line twice in a row.
|
|
24
|
+
- **Pin the language.** State the response language in the prompt; don't let the model infer it from an accent. If the app's domain has brand names or terms with non-obvious pronunciations, give them a line ("pronounce SQL as 'sequel'").
|
|
25
|
+
|
|
26
|
+
Beyond the mechanics, the persona itself should be *of the ear*: pacing, warmth, how it handles being interrupted, what it says when it needs a second. This is the fun part, same as the agent interface — a distinct character beats a generic assistant, and voice makes character land harder than any other surface.
|
|
50
27
|
|
|
51
28
|
### The shape of `system.md`
|
|
52
29
|
|
|
53
|
-
Structure the compiled prompt as short **labeled sections** — Role & Objective, Personality & Tone,
|
|
54
|
-
Rules, and (when the app has a real call flow) Conversation Flow — with bullets over paragraphs;
|
|
55
|
-
realtime models find and follow sectioned rules far more reliably than prose. Scope rules
|
|
56
|
-
precisely; blanket `always`/`never` makes the agent rigid and unable to handle reasonable exceptions.
|
|
57
|
-
And start minimal: state the role, the boundaries, and the voice mechanics above, then add rules only for
|
|
58
|
-
behaviors that actually misfire in test calls (the transcripts in the call log are the feedback
|
|
59
|
-
loop — `mindstudio-prod voice sessions get` reads a call verbatim) rather than front-loading a
|
|
60
|
-
policy manual.
|
|
30
|
+
Structure the compiled prompt as short **labeled sections** — Role & Objective, Personality & Tone, Rules, and (when the app has a real call flow) Conversation Flow — with bullets over paragraphs; realtime models find and follow sectioned rules far more reliably than prose. Scope rules precisely; blanket `always`/`never` makes the agent rigid and unable to handle reasonable exceptions. And start minimal: state the role, the boundaries, and the voice mechanics above, then add rules only for behaviors that actually misfire in test calls (the transcripts in the call log are the feedback loop — `mindstudio-prod voice sessions get` reads a call verbatim) rather than front-loading a policy manual.
|
|
61
31
|
|
|
62
32
|
### The latency classes
|
|
63
33
|
|
|
64
|
-
Every tool in the spec declares one of three classes. This is the voice-specific discipline — get it
|
|
65
|
-
right and tool use feels like talking to a competent person; get it wrong and every action is an
|
|
66
|
-
awkward pause.
|
|
34
|
+
Every tool in the spec declares one of three classes. This is the voice-specific discipline — get it right and tool use feels like talking to a competent person; get it wrong and every action is an awkward pause.
|
|
67
35
|
|
|
68
|
-
- **`fast`** — sub-second reads: lookups, availability checks, small queries. The agent calls
|
|
69
|
-
|
|
70
|
-
- **`
|
|
71
|
-
real work. The agent speaks a one-line preamble ("Let me get that booked") generated in parallel
|
|
72
|
-
with the call, so the line never goes quiet.
|
|
73
|
-
- **`background`** — long-running work: reports, enrichment, bulk operations. The agent
|
|
74
|
-
acknowledges, keeps conversing, and reports the result when it lands. Background tools are
|
|
75
|
-
cancellable — if the user changes course mid-run, the work stops.
|
|
36
|
+
- **`fast`** — sub-second reads: lookups, availability checks, small queries. The agent calls silently; announcing a sub-second call adds more delay than the call itself.
|
|
37
|
+
- **`slow`** — a noticeable wait, roughly one to three seconds: writes, searches, anything that does real work. The agent speaks a one-line preamble ("Let me get that booked") generated in parallel with the call, so the line never goes quiet.
|
|
38
|
+
- **`background`** — long-running work: reports, enrichment, bulk operations. The agent acknowledges, keeps conversing, and reports the result when it lands. Background tools are cancellable — if the user changes course mid-run, the work stops.
|
|
76
39
|
|
|
77
|
-
Classify by how the method actually behaves, not by what it is named. A "lookup" that fans out to an
|
|
78
|
-
external service is `slow`. When in doubt between `fast` and `slow`, pick `slow` — a needless
|
|
79
|
-
preamble is mildly chatty; an unexplained silence feels broken.
|
|
40
|
+
Classify by how the method actually behaves, not by what it is named. A "lookup" that fans out to an external service is `slow`. When in doubt between `fast` and `slow`, pick `slow` — a needless preamble is mildly chatty; an unexplained silence feels broken.
|
|
80
41
|
|
|
81
42
|
### Tool results reach the screen
|
|
82
43
|
|
|
83
|
-
Every successful tool call delivers its raw return value to the session's browser on the SDK's
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
session
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
the
|
|
92
|
-
discipline as agent-interface tools, whose results render in chat). Payloads over ~32KB serialized
|
|
93
|
-
arrive as `resultTruncated: true` with no data — keep returns compact, or have the UI fetch big
|
|
94
|
-
data itself. Failed calls deliver nothing to the client (the model gets the `{ error }` and speaks
|
|
95
|
-
a decline).
|
|
96
|
-
|
|
97
|
-
For backend-side correlation (writing results to a table keyed by the call, custom channels), the
|
|
98
|
-
method itself can read `session.voiceSessionId` / `session.visitorId` from the agent SDK
|
|
99
|
-
(`import { session } from '@mindstudio-ai/agent'`) — the same id the browser holds as
|
|
100
|
-
`session.sessionId`, guaranteed by the platform rather than echoed by the model.
|
|
101
|
-
|
|
102
|
-
Voice sessions also carry `session.medium` (`'web' | 'phone-in' | 'phone-out'`) and, on phone
|
|
103
|
-
calls, `session.sip` (`{ to, fromNumber }`) — use `medium` in tool methods and the session-context
|
|
104
|
-
method to branch web vs phone behavior (what to prefetch, how to phrase context). Treat
|
|
105
|
-
`session.sip.fromNumber` as context only, never identity: caller ID is spoofable, so don't gate
|
|
106
|
-
data or roles on it — in-call verification is the auth rail.
|
|
107
|
-
|
|
108
|
-
Ordinary web method calls carry the same context (`session.channel === 'web'` with `visitorId`),
|
|
109
|
-
so correlating a browser's web activity with its voice calls can key on `session.visitorId` on
|
|
110
|
-
both sides — one browser, one visitor id, web and voice alike. `session.voiceSessionId` remains
|
|
111
|
-
the per-call key when you need to distinguish individual calls.
|
|
44
|
+
Every successful tool call delivers its raw return value to the session's browser on the SDK's `toolCall` event (`result` field, on `done`) — so the UI can render what the agent just did (the citation it found, the record it pulled up, the booking it made) in lockstep with the spoken answer. No flag, no polling, no key-threading, no model involvement: delivery to the invoking session's own client is the same security context as the invocation itself (an RPC response), and it's scoped to that one session.
|
|
45
|
+
|
|
46
|
+
Consequence for authoring: **a tool's return value is user-visible by definition.** Return what the user may see — no internal fields, keys, or diagnostics you wouldn't put on screen (the same discipline as agent-interface tools, whose results render in chat). Payloads over ~32KB serialized arrive as `resultTruncated: true` with no data — keep returns compact, or have the UI fetch big data itself. Failed calls deliver nothing to the client (the model gets the `{ error }` and speaks a decline).
|
|
47
|
+
|
|
48
|
+
For backend-side correlation (writing results to a table keyed by the call, custom channels), the method itself can read `session.voiceSessionId` / `session.visitorId` from the agent SDK (`import { session } from '@mindstudio-ai/agent'`) — the same id the browser holds as `session.sessionId`, guaranteed by the platform rather than echoed by the model.
|
|
49
|
+
|
|
50
|
+
Voice sessions also carry `session.medium` (`'web' | 'phone-in' | 'phone-out'`) and, on phone calls, `session.sip` (`{ to, fromNumber }`) — use `medium` in tool methods and the session-context method to branch web vs phone behavior (what to prefetch, how to phrase context). Treat `session.sip.fromNumber` as context only, never identity: caller ID is spoofable, so don't gate data or roles on it — in-call verification is the auth rail.
|
|
51
|
+
|
|
52
|
+
Ordinary web method calls carry the same context (`session.channel === 'web'` with `visitorId`), so correlating a browser's web activity with its voice calls can key on `session.visitorId` on both sides — one browser, one visitor id, web and voice alike. `session.voiceSessionId` remains the per-call key when you need to distinguish individual calls.
|
|
112
53
|
|
|
113
54
|
### Client tools: actions that happen on screen (`target: "client"`)
|
|
114
55
|
|
|
115
|
-
A tool whose effect belongs in the browser — open the verification sheet, navigate to a page,
|
|
116
|
-
highlight a record — is declared with `target: "client"` instead of a `method`:
|
|
56
|
+
A tool whose effect belongs in the browser — open the verification sheet, navigate to a page, highlight a record — is declared with `target: "client"` instead of a `method`:
|
|
117
57
|
|
|
118
58
|
```json
|
|
119
59
|
{
|
|
@@ -124,15 +64,10 @@ highlight a record — is declared with `target: "client"` instead of a `method`
|
|
|
124
64
|
}
|
|
125
65
|
```
|
|
126
66
|
|
|
127
|
-
The platform never touches the backend for these: the agent's invocation is delivered to the
|
|
128
|
-
session's browser, the app's registered handler runs, and the handler's **return value goes back
|
|
129
|
-
to the agent as the tool result** — a real request/response, so the agent knows the sheet
|
|
130
|
-
actually opened (or that the user dismissed it) and speaks accordingly. Rules:
|
|
67
|
+
The platform never touches the backend for these: the agent's invocation is delivered to the session's browser, the app's registered handler runs, and the handler's **return value goes back to the agent as the tool result** — a real request/response, so the agent knows the sheet actually opened (or that the user dismissed it) and speaks accordingly. Rules:
|
|
131
68
|
|
|
132
|
-
- `name` instead of `method`; must not collide with any backend method id. No latency class —
|
|
133
|
-
|
|
134
|
-
- `inputSchema` is authored inline (an object schema) — there's no method contract to derive
|
|
135
|
-
it from. Keep it small; these are UI directives, not data payloads.
|
|
69
|
+
- `name` instead of `method`; must not collide with any backend method id. No latency class — the agent holds the turn while the browser responds (up to ~30s, then a timeout error).
|
|
70
|
+
- `inputSchema` is authored inline (an object schema) — there's no method contract to derive it from. Keep it small; these are UI directives, not data payloads.
|
|
136
71
|
- The frontend must register a handler, or invocations fail as `unhandled_client_tool`:
|
|
137
72
|
|
|
138
73
|
```js
|
|
@@ -142,128 +77,67 @@ session.registerClientTool('showVerification', async ({ reason }) => {
|
|
|
142
77
|
});
|
|
143
78
|
```
|
|
144
79
|
|
|
145
|
-
- One client tool runs at a time per session; the description should tell the agent when to use
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
- The progressive-auth pattern above is the canonical use: make the verification sheet a client
|
|
149
|
-
tool and the agent opens it deliberately instead of the frontend inferring it from tool events.
|
|
150
|
-
- Phone sessions never see client tools — there is no browser on a call, so the platform drops
|
|
151
|
-
them from the toolset and tells the agent it's on a phone call with no screen. Verification
|
|
152
|
-
branches by itself: on a phone the agent gets the platform's in-call verify tools instead of
|
|
153
|
-
the app's sheet. Nothing to author; just don't make a client tool the only path to something
|
|
154
|
-
phone callers need.
|
|
80
|
+
- One client tool runs at a time per session; the description should tell the agent when to use it and what to say while it's on screen. Throwing from the handler (or returning nothing) becomes an error/ack the agent can speak around.
|
|
81
|
+
- The progressive-auth pattern above is the canonical use: make the verification sheet a client tool and the agent opens it deliberately instead of the frontend inferring it from tool events.
|
|
82
|
+
- Phone sessions never see client tools — there is no browser on a call, so the platform drops them from the toolset and tells the agent it's on a phone call with no screen. Verification branches by itself: on a phone the agent gets the platform's in-call verify tools instead of the app's sheet. Nothing to author; just don't make a client tool the only path to something phone callers need.
|
|
155
83
|
|
|
156
84
|
### Tool descriptions say results out loud
|
|
157
85
|
|
|
158
|
-
Follow the agent-interface principles for tool descriptions (when to use and when not, parameter
|
|
159
|
-
guidance, what comes back) — plus one voice-specific layer: **how to speak the result.** A tool that
|
|
160
|
-
returns a booking record needs its description to say what the confirmation sounds like ("You're all
|
|
161
|
-
set for Tuesday at two") and what never gets read aloud (internal ids, timestamps, enum values).
|
|
86
|
+
Follow the agent-interface principles for tool descriptions (when to use and when not, parameter guidance, what comes back) — plus one voice-specific layer: **how to speak the result.** A tool that returns a booking record needs its description to say what the confirmation sounds like ("You're all set for Tuesday at two") and what never gets read aloud (internal ids, timestamps, enum values).
|
|
162
87
|
|
|
163
|
-
Curate harder than you would for chat. A voice agent with four excellent tools outperforms one with
|
|
164
|
-
twelve adequate ones — every tool the model considers is a beat of hesitation. Skip batch
|
|
165
|
-
operations, admin utilities, and anything whose output can't be said in a breath or two. Note role
|
|
166
|
-
restrictions in the description so the agent declines gracefully in character instead of surfacing a
|
|
167
|
-
rejection.
|
|
88
|
+
Curate harder than you would for chat. A voice agent with four excellent tools outperforms one with twelve adequate ones — every tool the model considers is a beat of hesitation. Skip batch operations, admin utilities, and anything whose output can't be said in a breath or two. Note role restrictions in the description so the agent declines gracefully in character instead of surfacing a rejection.
|
|
168
89
|
|
|
169
90
|
### Confirmation scales with risk
|
|
170
91
|
|
|
171
|
-
Bake the policy into the system prompt: read-only tools — just call them. Writes — summarize what's
|
|
172
|
-
about to happen and get a yes. Anything destructive or financial — read the details back first,
|
|
173
|
-
piece by piece. In voice there is no confirmation dialog to lean on; the conversation *is* the
|
|
174
|
-
confirmation UI.
|
|
92
|
+
Bake the policy into the system prompt: read-only tools — just call them. Writes — summarize what's about to happen and get a yes. Anything destructive or financial — read the details back first, piece by piece. In voice there is no confirmation dialog to lean on; the conversation *is* the confirmation UI.
|
|
175
93
|
|
|
176
|
-
And give failure a script: never speak a raw error. When a lookup misses or a tool fails, read back
|
|
177
|
-
the value it used ("I couldn't find an order ending three-one-two-five — did I get part of that
|
|
178
|
-
wrong?"), offer one retry, then move to an alternate path — in character, without blaming the
|
|
179
|
-
caller.
|
|
94
|
+
And give failure a script: never speak a raw error. When a lookup misses or a tool fails, read back the value it used ("I couldn't find an order ending three-one-two-five — did I get part of that wrong?"), offer one retry, then move to an alternate path — in character, without blaming the caller.
|
|
180
95
|
|
|
181
96
|
### Choosing the model
|
|
182
97
|
|
|
183
|
-
Two shapes, one `model` field. **Use native speech-to-speech unless the user specifically asks
|
|
184
|
-
for a cascaded pipeline** — one realtime model hears and speaks: lowest latency, most natural
|
|
185
|
-
prosody, hears tone and hesitation.
|
|
98
|
+
Two shapes, one `model` field. **Use native speech-to-speech unless the user specifically asks for a cascaded pipeline** — one realtime model hears and speaks: lowest latency, most natural prosody, hears tone and hesitation.
|
|
186
99
|
|
|
187
|
-
**Native** (`{"model": ..., "voice": ...}`) — **default to `gpt-realtime-2.1` with voice
|
|
188
|
-
`marin`.**
|
|
100
|
+
**Native** (`{"model": ..., "voice": ...}`) — **default to `gpt-realtime-2.1` with voice `marin`.**
|
|
189
101
|
|
|
190
|
-
- `gpt-realtime-2.1` — the default. Voices: `marin` (default), `cedar`, `alloy`, `ash`,
|
|
191
|
-
`ballad`, `coral`, `echo`, `sage`, `shimmer`, `verse`.
|
|
102
|
+
- `gpt-realtime-2.1` — the default. Voices: `marin` (default), `cedar`, `alloy`, `ash`, `ballad`, `coral`, `echo`, `sage`, `shimmer`, `verse`.
|
|
192
103
|
- `gpt-realtime-2.1-mini` — the same family, lighter; same voices.
|
|
193
|
-
- `gemini-2.5-flash-native-audio-preview-12-2025` — the Gemini pick, with a large expressive
|
|
194
|
-
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
|
|
200
|
-
`Zubenelgenubi` (casual), `Vindemiatrix` (gentle), `Sadachbia` (lively), `Sadaltager`
|
|
201
|
-
(knowledgeable), `Sulafat` (warm).
|
|
202
|
-
- `gemini-3.1-flash-live-preview` — newer Gemini, same voices as 2.5, but currently can't speak
|
|
203
|
-
an opening greeting or take mid-call prompt updates (a plugin limitation expected to resolve
|
|
204
|
-
upstream) — prefer 2.5 until then.
|
|
205
|
-
- `grok-voice-think-fast-2.0` — a distinct personality register. Voices: `eve` (default),
|
|
206
|
-
`altair`, `ara`, `atlas`, `aurora`, `carina`, `castor`, `celeste`, `cosmo`, `helios`, `helix`,
|
|
207
|
-
`iris`, `kepler`, `leo`, `liora`, `lumen`, `luna`, `lux`, `naksh`, `orion`, `perseus`, `rex`,
|
|
208
|
-
`rigel`, `sal`, `sirius`, `ursa`, `zagan`, `zenith`.
|
|
209
|
-
|
|
210
|
-
**Cascaded** (`{"llm": ..., "stt": ..., "tts": ..., "voice": ...}`) — streaming transcription
|
|
211
|
-
into any chat model in the catalog, streaming speech out. Slightly higher latency; reach for it
|
|
212
|
-
only when the user wants it or the app's reasoning demands a specific chat model (e.g. the agent
|
|
213
|
-
interface already uses one and the voice should think identically). Slots: `stt` is
|
|
214
|
-
`deepgram-nova-3`; `tts` is `cartesia-sonic-3` (voices are per-account Cartesia UUIDs — see
|
|
215
|
-
play.cartesia.ai) or `elevenlabs-tts` (the account's ElevenLabs voice library); `llm` is any chat
|
|
216
|
-
model — ask `askMindStudioSdk` for chat model ids. One nuance: cascaded engines speak the
|
|
217
|
-
`greeting` verbatim (they have a real TTS); speech-to-speech engines have the model say it, so it
|
|
218
|
-
may paraphrase slightly.
|
|
219
|
-
|
|
220
|
-
The model and voice ids above are current and maintained with the platform — use them as written
|
|
221
|
-
(they are MindStudio ids, not vendor ids). The user's UI has a picker for changing the model
|
|
222
|
-
later, so validate only when you set it.
|
|
104
|
+
- `gemini-2.5-flash-native-audio-preview-12-2025` — the Gemini pick, with a large expressive roster: `Puck` (default, upbeat), `Zephyr` (bright), `Charon` (informative), `Kore` (firm), `Fenrir` (excitable), `Leda` (youthful), `Orus` (firm), `Aoede` (breezy), `Callirrhoe` (easy-going), `Autonoe` (bright), `Enceladus` (breathy), `Iapetus` (clear), `Umbriel` (easy-going), `Algieba` (smooth), `Despina` (smooth), `Erinome` (clear), `Algenib` (gravelly), `Rasalgethi` (informative), `Laomedeia` (upbeat), `Achernar` (soft), `Alnilam` (firm), `Schedar` (even), `Gacrux` (mature), `Pulcherrima` (forward), `Achird` (friendly), `Zubenelgenubi` (casual), `Vindemiatrix` (gentle), `Sadachbia` (lively), `Sadaltager` (knowledgeable), `Sulafat` (warm).
|
|
105
|
+
- `gemini-3.1-flash-live-preview` — newer Gemini, same voices as 2.5, but currently can't speak an opening greeting or take mid-call prompt updates (a plugin limitation expected to resolve upstream) — prefer 2.5 until then.
|
|
106
|
+
- `grok-voice-think-fast-2.0` — a distinct personality register. Voices: `eve` (default), `altair`, `ara`, `atlas`, `aurora`, `carina`, `castor`, `celeste`, `cosmo`, `helios`, `helix`, `iris`, `kepler`, `leo`, `liora`, `lumen`, `luna`, `lux`, `naksh`, `orion`, `perseus`, `rex`, `rigel`, `sal`, `sirius`, `ursa`, `zagan`, `zenith`.
|
|
107
|
+
|
|
108
|
+
**Cascaded** (`{"llm": ..., "stt": ..., "tts": ..., "voice": ...}`) — streaming transcription into any chat model in the catalog, streaming speech out. Slightly higher latency; reach for it only when the user wants it or the app's reasoning demands a specific chat model (e.g. the agent interface already uses one and the voice should think identically). Slots: `stt` is `deepgram-nova-3`; `tts` is `cartesia-sonic-3` (voices are per-account Cartesia UUIDs — see play.cartesia.ai) or `elevenlabs-tts` (the account's ElevenLabs voice library); `llm` is any chat model — ask `askMindStudioSdk` for chat model ids. One nuance: cascaded engines speak the `greeting` verbatim (they have a real TTS); speech-to-speech engines have the model say it, so it may paraphrase slightly.
|
|
109
|
+
|
|
110
|
+
The model and voice ids above are current and maintained with the platform — use them as written (they are MindStudio ids, not vendor ids). The user's UI has a picker for changing the model later, so validate only when you set it.
|
|
223
111
|
|
|
224
112
|
### Seeding from an existing agent
|
|
225
113
|
|
|
226
|
-
If the app already has an agent interface, start from it: same character, same values, same
|
|
227
|
-
terminology — then rewrite for the ear (shorter, spoken-form, no formatting) and re-curate the
|
|
228
|
-
toolset for latency. Don't copy `agent.md`'s prose wholesale; a chat persona read aloud sounds like
|
|
229
|
-
someone reading chat aloud.
|
|
114
|
+
If the app already has an agent interface, start from it: same character, same values, same terminology — then rewrite for the ear (shorter, spoken-form, no formatting) and re-curate the toolset for latency. Don't copy `agent.md`'s prose wholesale; a chat persona read aloud sounds like someone reading chat aloud.
|
|
230
115
|
|
|
231
116
|
### Anti-patterns
|
|
232
117
|
|
|
233
118
|
- Prose that would render fine in chat — bullet lists, headers, or markdown anywhere in `system.md`.
|
|
234
119
|
- A tool description that explains what to display instead of what to say.
|
|
235
120
|
- Exposing the whole method surface. Voice is the most curated interface the app has.
|
|
236
|
-
- A generic greeting ("Hello! How can I assist you today?"). The greeting is the first thing anyone
|
|
237
|
-
|
|
238
|
-
- Writing your own current-user placeholder — the platform appends a `## Current User` block
|
|
239
|
-
(email, phone, roles) to every system prompt at runtime.
|
|
121
|
+
- A generic greeting ("Hello! How can I assist you today?"). The greeting is the first thing anyone hears; make it the character's.
|
|
122
|
+
- Writing your own current-user placeholder — the platform appends a `## Current User` block (email, phone, roles) to every system prompt at runtime.
|
|
240
123
|
|
|
241
124
|
## Compiling the Voice Spec
|
|
242
125
|
|
|
243
|
-
When building `dist/interfaces/voice/`, consider the spec, the app, and the `@brand/` guidelines —
|
|
244
|
-
the voice agent should be unmistakably the same product as the web UI, projected into sound. Output:
|
|
126
|
+
When building `dist/interfaces/voice/`, consider the spec, the app, and the `@brand/` guidelines — the voice agent should be unmistakably the same product as the web UI, projected into sound. Output:
|
|
245
127
|
|
|
246
|
-
**`system.md`** — the persona compiled for the ear. Character first, then the mandatory carries from
|
|
247
|
-
"Written for the ear" above (spoken-form rules, brevity, unclear-audio handling, language pinning,
|
|
248
|
-
confirmation-by-risk), then any preamble phrasing guidance for `slow` tools so the fillers sound like
|
|
249
|
-
the character too.
|
|
128
|
+
**`system.md`** — the persona compiled for the ear. Character first, then the mandatory carries from "Written for the ear" above (spoken-form rules, brevity, unclear-audio handling, language pinning, confirmation-by-risk), then any preamble phrasing guidance for `slow` tools so the fillers sound like the character too.
|
|
250
129
|
|
|
251
|
-
**`tools/*.md`** — one per tool: when to use, parameter guidance, how to say the result, role
|
|
252
|
-
restrictions.
|
|
130
|
+
**`tools/*.md`** — one per tool: when to use, parameter guidance, how to say the result, role restrictions.
|
|
253
131
|
|
|
254
132
|
**`interface.json`** — the config tying it together. Full shape in "The wiring" below.
|
|
255
133
|
|
|
256
134
|
## Voice UI
|
|
257
135
|
|
|
258
|
-
When the app has a web interface, voice arrives as a **layer over it**, not a separate page: a
|
|
259
|
-
persistent affordance (a button, an orb in a corner) that starts a session in place, with the app
|
|
260
|
-
still visible and usable. A dedicated full-screen voice mode is the immersive option for apps where
|
|
261
|
-
the conversation *is* the product — earn it, don't default to it.
|
|
136
|
+
When the app has a web interface, voice arrives as a **layer over it**, not a separate page: a persistent affordance (a button, an orb in a corner) that starts a session in place, with the app still visible and usable. A dedicated full-screen voice mode is the immersive option for apps where the conversation *is* the product — earn it, don't default to it.
|
|
262
137
|
|
|
263
138
|
### Frontend SDK: `createVoiceClient()`
|
|
264
139
|
|
|
265
|
-
Ships as a subpath of the interface SDK so apps that never use voice pay nothing for it. All voice
|
|
266
|
-
UIs go through it — never hand-roll audio capture or transport.
|
|
140
|
+
Ships as a subpath of the interface SDK so apps that never use voice pay nothing for it. All voice UIs go through it — never hand-roll audio capture or transport.
|
|
267
141
|
|
|
268
142
|
```ts
|
|
269
143
|
import { createVoiceClient } from '@mindstudio-ai/interface/voice';
|
|
@@ -293,49 +167,37 @@ await session.refreshIdentity(); // after in-app verification — upgrade t
|
|
|
293
167
|
session.end();
|
|
294
168
|
```
|
|
295
169
|
|
|
296
|
-
Agent audio playback is handled inside the SDK (a hidden autoplaying element) — never create audio
|
|
297
|
-
elements for the agent. `startSession()` throws `MindStudioInterfaceError` with code
|
|
298
|
-
`microphone_denied` when mic access is refused (surface that state gently in the UI),
|
|
299
|
-
`voice_concurrency_limit` / `voice_visitor_limit` when the app's session limits are hit, and
|
|
300
|
-
`auth_required` (401) / `role_required` (403) when the interface's `auth` block denies the caller
|
|
301
|
-
(route those to the app's login flow).
|
|
170
|
+
Agent audio playback is handled inside the SDK (a hidden autoplaying element) — never create audio elements for the agent. `startSession()` throws `MindStudioInterfaceError` with code `microphone_denied` when mic access is refused (surface that state gently in the UI), `voice_concurrency_limit` / `voice_visitor_limit` when the app's session limits are hit, and `auth_required` (401) / `role_required` (403) when the interface's `auth` block denies the caller (route those to the app's login flow).
|
|
302
171
|
|
|
303
|
-
Past sessions are call records with transcripts: `voice.listSessions()` /
|
|
304
|
-
`voice.getSession(id)` — the material for a history view if the app wants one.
|
|
172
|
+
Past sessions are call records with transcripts: `voice.listSessions()` / `voice.getSession(id)` — the material for a history view if the app wants one.
|
|
305
173
|
|
|
306
|
-
### The
|
|
174
|
+
### The voice agent, on screen
|
|
307
175
|
|
|
308
|
-
One audio-reactive element carries the session: idle → connecting → listening → thinking → speaking.
|
|
309
|
-
|
|
310
|
-
idle, responsive to actual audio levels while listening and speaking.
|
|
311
|
-
`prefers-reduced-motion` with a static-but-labeled variant.
|
|
176
|
+
One audio-reactive element carries the session state machine: idle → connecting → listening → thinking → speaking. It is on screen for the entire conversation and it is the most visual thing about a voice app — treat it as a **first-class design deliverable**, not a widget. Bring in `visualDesignExpert` for the session experience as a whole — the centerpiece, the captions, how tool activity surfaces, and the controls, composed as one scene (it has a dedicated craft reference for exactly this) — and implement what it prescribes. The bar is a real-time computed piece — WebGL, structured, alive, derived from the app's brand (a sampled point-cloud object, an instrument/meter, whatever the domain suggests). A generic gradient sphere or a pulsing CSS circle is a defaulted artifact, not a designed one.
|
|
177
|
+
|
|
178
|
+
Whatever the designer specs, the mechanics you own: always pair the piece with a **text state label** — never signal state by color or motion alone; calm at idle, responsive to actual audio levels while listening and speaking. And build the performance budget in from day one: cap `devicePixelRatio` at 2, reduce density on mobile, pause the render loop when the element is off-screen or the tab is hidden, and ship a static-but-labeled fallback for `prefers-reduced-motion` and no-WebGL.
|
|
312
179
|
|
|
313
180
|
### Live captions
|
|
314
181
|
|
|
315
|
-
Stream `transcript` events as captions — both sides of the conversation. Captions make the agent
|
|
316
|
-
|
|
317
|
-
|
|
318
|
-
as captions, never treat them as input to app logic.
|
|
182
|
+
Stream `transcript` events as captions — both sides of the conversation. Captions make the agent feel accurate, catch mishearings early, and are the accessibility story. User-side transcripts arrive as recognition output and can lag or differ slightly from what the model heard; render them as captions, never treat them as input to app logic.
|
|
183
|
+
|
|
184
|
+
Captions must be **layout-stable**: reserve a fixed-height caption region so arriving text never shifts the layout around it (a streaming segment that reflows the page on every event reads as jank, and it fights the agent's visual for attention). Upsert each segment in place by `segmentId` — the event carries the segment's full text, so replacing is free — cap the visible lines, and fade old lines out rather than pushing content down.
|
|
319
185
|
|
|
320
186
|
### Controls that must exist
|
|
321
187
|
|
|
322
|
-
**Mute** and **end call**, always visible, always working. `sendText` earns its place the moment the
|
|
323
|
-
conversation needs an exact string — an address, a code, an email — typing it beats spelling it
|
|
324
|
-
aloud three times. Show tool activity as a compact inline status from `toolCall` events, in the
|
|
325
|
-
app's voice ("Booking your appointment…"), never raw names or JSON.
|
|
188
|
+
**Mute** and **end call**, always visible, always working. `sendText` earns its place the moment the conversation needs an exact string — an address, a code, an email — typing it beats spelling it aloud three times. Show tool activity as a compact inline status from `toolCall` events, in the app's voice ("Booking your appointment…"), never raw names or JSON.
|
|
326
189
|
|
|
327
190
|
### Anti-patterns
|
|
328
191
|
|
|
329
192
|
- Blocking the whole UI behind the session — voice is a layer, the app stays usable.
|
|
193
|
+
- A generic gradient sphere or pulsing CSS circle as the voice agent's visual. The visual is designed (by the design expert), not defaulted.
|
|
330
194
|
- An orb with no label, or state changes conveyed only by color.
|
|
331
195
|
- Rendering user-side captions as authoritative ("you said X") — they're recognition output.
|
|
332
196
|
- Auto-starting a session on page load. Microphone access is always a deliberate user action.
|
|
333
197
|
|
|
334
198
|
## Outbound calls (`voice.call`)
|
|
335
199
|
|
|
336
|
-
The agent can call the user. Backend methods (and crons) place outbound phone calls with the
|
|
337
|
-
agent SDK's `voice` namespace — the platform dials the number and connects the callee to this
|
|
338
|
-
app's voice agent (same persona, engine, and tools as the web sessions):
|
|
200
|
+
The agent can call the user. Backend methods (and crons) place outbound phone calls with the agent SDK's `voice` namespace — the platform dials the number and connects the callee to this app's voice agent (same persona, engine, and tools as the web sessions):
|
|
339
201
|
|
|
340
202
|
```ts
|
|
341
203
|
import { voice, auth } from '@mindstudio-ai/agent';
|
|
@@ -347,84 +209,36 @@ export async function callMeAboutMyOrder(input: { phone: string }) {
|
|
|
347
209
|
}
|
|
348
210
|
```
|
|
349
211
|
|
|
350
|
-
- **The method is the authorization gate.** The voice interface's `auth` block does not apply to
|
|
351
|
-
|
|
352
|
-
|
|
353
|
-
-
|
|
354
|
-
|
|
355
|
-
|
|
356
|
-
session, not the phone). Omitted/false → anonymous call; role-gated tools decline.
|
|
357
|
-
System/cron invocations have no human identity and always run anonymously. Anonymous outbound
|
|
358
|
-
calls (deployed) get the same in-call verification flow as inbound — the callee proves
|
|
359
|
-
possession of the number that was dialed, or verifies by email — so an anonymous call can
|
|
360
|
-
still upgrade to a known user mid-conversation.
|
|
361
|
-
- **Production needs a dedicated phone number.** The app owner attaches one ($1/month) via the
|
|
362
|
-
dashboard or `mindstudio-prod voice numbers` (see "The voice CLI"
|
|
363
|
-
below) — it becomes the caller ID for every call, in dev sessions too, so users always see
|
|
364
|
-
the same number. Without one, deployed calls throw `phone_out_requires_dedicated_number`, and
|
|
365
|
-
dev sessions fall back to a shared platform test number that varies per call (tighter limits
|
|
366
|
-
apply on the shared pool).
|
|
367
|
-
- **Outcome is on the call record**, not the return value: `voice.call` returns as soon as
|
|
368
|
-
dialing starts (`{ sessionId, status: 'dialing', from, to }`); answered/busy/no-answer land on
|
|
369
|
-
the session in the app's call log (`voice.listSessions()` / the dashboard).
|
|
370
|
-
- **Limits**: the app's concurrent-session policy, a daily outbound-call cap, a per-call
|
|
371
|
-
duration ceiling, and one active call per callee number (`voice_callee_busy`).
|
|
372
|
-
- **Compliance**: automated calls require prior consent. Call your own users who opted in to
|
|
373
|
-
calls from this app, honor reasonable calling hours, never dial purchased or cold lists —
|
|
374
|
-
design the consent moment into the product (a "call me" button IS consent; a scraped list is
|
|
375
|
-
not).
|
|
212
|
+
- **The method is the authorization gate.** The voice interface's `auth` block does not apply to calls the backend places deliberately — gate the *method* with `auth.requireRole(...)` exactly as you would any sensitive action.
|
|
213
|
+
- **`assumeIdentity: true`** runs the call as the user who invoked the method: the agent knows who it's talking to (Current User block) and every tool call carries their roles — regardless of which number was dialed (the user types any number into a field; identity comes from their session, not the phone). Omitted/false → anonymous call; role-gated tools decline. System/cron invocations have no human identity and always run anonymously. Anonymous outbound calls (deployed) get the same in-call verification flow as inbound — the callee proves possession of the number that was dialed, or verifies by email — so an anonymous call can still upgrade to a known user mid-conversation.
|
|
214
|
+
- **Production needs a dedicated phone number.** The app owner attaches one ($1/month) via the dashboard or `mindstudio-prod voice numbers` (see "The voice CLI" below) — it becomes the caller ID for every call, in dev sessions too, so users always see the same number. Without one, deployed calls throw `phone_out_requires_dedicated_number`, and dev sessions fall back to a shared platform test number that varies per call (tighter limits apply on the shared pool).
|
|
215
|
+
- **Outcome is on the call record**, not the return value: `voice.call` returns as soon as dialing starts (`{ sessionId, status: 'dialing', from, to }`); answered/busy/no-answer land on the session in the app's call log (`voice.listSessions()` / the dashboard).
|
|
216
|
+
- **Limits**: the app's concurrent-session policy, a daily outbound-call cap, a per-call duration ceiling, and one active call per callee number (`voice_callee_busy`).
|
|
217
|
+
- **Compliance**: automated calls require prior consent. Call your own users who opted in to calls from this app, honor reasonable calling hours, never dial purchased or cold lists — design the consent moment into the product (a "call me" button IS consent; a scraped list is not).
|
|
376
218
|
|
|
377
219
|
## Inbound calls
|
|
378
220
|
|
|
379
|
-
Once the app has a dedicated phone number, people can call it — the same voice agent answers
|
|
380
|
-
(same persona, engine, and tools). Nothing extra to author for the basic case; the number in the
|
|
381
|
-
app's settings is the whole switch.
|
|
221
|
+
Once the app has a dedicated phone number, people can call it — the same voice agent answers (same persona, engine, and tools). Nothing extra to author for the basic case; the number in the app's settings is the whole switch.
|
|
382
222
|
|
|
383
223
|
How answering works:
|
|
384
224
|
|
|
385
|
-
- **Inbound always runs the live release.** There is no dev inbound — test the agent over the
|
|
386
|
-
|
|
387
|
-
|
|
388
|
-
- **
|
|
389
|
-
can't
|
|
390
|
-
|
|
391
|
-
|
|
392
|
-
- **Verification uses the app's own auth methods** (`sms-code` / `email-code` from the
|
|
393
|
-
manifest):
|
|
394
|
-
- SMS: a code is texted to the phone number on the call (no other number is possible by
|
|
395
|
-
design), and confirming it signs the caller in — **creating their account if they're new**,
|
|
396
|
-
the same find-or-create policy `sms-code` has on web (enabling the method is what enables
|
|
397
|
-
sign-up; there is no separate toggle on either surface). After a first-time caller verifies,
|
|
398
|
-
the session-context method re-fires with the fresh identity — that's the hook for seeding a
|
|
399
|
-
new account with data.
|
|
400
|
-
- Email: existing accounts only — the caller says their address; the platform matches it
|
|
401
|
-
against the app's users (transcription-tolerant — no letter-by-letter spelling ceremony) and
|
|
402
|
-
emails the account's stored address a code. There is no sign-up by email over the phone: a
|
|
403
|
-
call can't reliably capture a verbatim never-seen address, so new callers sign up by text
|
|
404
|
-
instead.
|
|
405
|
-
- The email flow never confirms or denies that an account exists — a code is "sent if an
|
|
406
|
-
account matches", always phrased that neutrally. The persona should offer verification
|
|
407
|
-
naturally when it unlocks something, never as a robotic gate.
|
|
408
|
-
- **Verified mid-call, upgraded mid-call**: once the code checks out, the session becomes that
|
|
409
|
-
user's — Current User block, roles on every tool call — without redialing.
|
|
225
|
+
- **Inbound always runs the live release.** There is no dev inbound — test the agent over the normal WebRTC session in the editor; the phone is the same interface with a different transport. An app with no live voice interface (or at its concurrency limit) doesn't answer.
|
|
226
|
+
- **Callers are anonymous until verified.** The `auth` block still applies, but a phone call can't show a login page — so the platform answers first, and `requireUser` becomes an in-call verification flow. The agent can serve whatever anonymous callers are allowed, and offers verification when the caller wants something account-bound.
|
|
227
|
+
- **Verification uses the app's own auth methods** (`sms-code` / `email-code` from the manifest):
|
|
228
|
+
- SMS: a code is texted to the phone number on the call (no other number is possible by design), and confirming it signs the caller in — **creating their account if they're new**, the same find-or-create policy `sms-code` has on web (enabling the method is what enables sign-up; there is no separate toggle on either surface). After a first-time caller verifies, the session-context method re-fires with the fresh identity — that's the hook for seeding a new account with data.
|
|
229
|
+
- Email: existing accounts only — the caller says their address; the platform matches it against the app's users (transcription-tolerant — no letter-by-letter spelling ceremony) and emails the account's stored address a code. There is no sign-up by email over the phone: a call can't reliably capture a verbatim never-seen address, so new callers sign up by text instead.
|
|
230
|
+
- The email flow never confirms or denies that an account exists — a code is "sent if an account matches", always phrased that neutrally. The persona should offer verification naturally when it unlocks something, never as a robotic gate.
|
|
231
|
+
- **Verified mid-call, upgraded mid-call**: once the code checks out, the session becomes that user's — Current User block, roles on every tool call — without redialing.
|
|
410
232
|
|
|
411
233
|
### `phone.trustCallerId`
|
|
412
234
|
|
|
413
|
-
For apps whose users are known by phone number, the interface config may opt into treating
|
|
414
|
-
caller ID as identity:
|
|
235
|
+
For apps whose users are known by phone number, the interface config may opt into treating caller ID as identity:
|
|
415
236
|
|
|
416
237
|
```json
|
|
417
238
|
"phone": { "trustCallerId": true }
|
|
418
239
|
```
|
|
419
240
|
|
|
420
|
-
A caller whose number exactly matches an app user's phone starts the call already verified —
|
|
421
|
-
no code. This is a real security tradeoff: **caller ID can be spoofed**, so a motivated
|
|
422
|
-
attacker who knows a user's phone number can impersonate them to this agent. Before enabling
|
|
423
|
-
it, you MUST surface that risk to the user and get their explicit confirmation — it's the
|
|
424
|
-
right call for convenience-first, low-stakes apps (a family assistant, a status line), and the
|
|
425
|
-
wrong one wherever the agent's tools can move money, reveal sensitive records, or take
|
|
426
|
-
destructive actions. It lives in the interface config deliberately: enabling it is a code
|
|
427
|
-
change, visible in review and auditable via deploys, not a dashboard toggle.
|
|
241
|
+
A caller whose number exactly matches an app user's phone starts the call already verified — no code. This is a real security tradeoff: **caller ID can be spoofed**, so a motivated attacker who knows a user's phone number can impersonate them to this agent. Before enabling it, you MUST surface that risk to the user and get their explicit confirmation — it's the right call for convenience-first, low-stakes apps (a family assistant, a status line), and the wrong one wherever the agent's tools can move money, reveal sensitive records, or take destructive actions. It lives in the interface config deliberately: enabling it is a code change, visible in review and auditable via deploys, not a dashboard toggle.
|
|
428
242
|
|
|
429
243
|
## The voice CLI
|
|
430
244
|
|
|
@@ -438,18 +252,11 @@ mindstudio-prod voice sessions list --limit 10 # call log: web / phone-o
|
|
|
438
252
|
mindstudio-prod voice sessions get <sessionId> # full transcript + cost breakdown
|
|
439
253
|
```
|
|
440
254
|
|
|
441
|
-
Also `voice numbers list`, `voice numbers set-name` (outbound caller-ID display
|
|
442
|
-
name; 12-72h carrier propagation), `voice settings get`/`set` (concurrency, per-visitor,
|
|
443
|
-
max duration — `set` merges: only the settings you pass change). `--help` for flags.
|
|
255
|
+
Also `voice numbers list`, `voice numbers set-name` (outbound caller-ID display name; 12-72h carrier propagation), `voice settings get`/`set` (concurrency, per-visitor, max duration — `set` merges: only the settings you pass change). `--help` for flags.
|
|
444
256
|
|
|
445
|
-
**Never buy a number without the user's explicit confirmation** — it starts a recurring
|
|
446
|
-
$1/month workspace charge. Search first, present the options with the price, and only run
|
|
447
|
-
`numbers buy` after they've picked one and said yes.
|
|
257
|
+
**Never buy a number without the user's explicit confirmation** — it starts a recurring $1/month workspace charge. Search first, present the options with the price, and only run `numbers buy` after they've picked one and said yes.
|
|
448
258
|
|
|
449
|
-
Transcripts are how you iterate on a voice persona: after the user test-calls the agent, read
|
|
450
|
-
`voice sessions get` for what was actually said — misheard input, interruptions, tools declining
|
|
451
|
-
— and fix the spec from evidence rather than guesses. (Dev-session test calls carry a
|
|
452
|
-
`devSessionId` in the list, so you can tell them from live traffic.)
|
|
259
|
+
Transcripts are how you iterate on a voice persona: after the user test-calls the agent, read `voice sessions get` for what was actually said — misheard input, interruptions, tools declining — and fix the spec from evidence rather than guesses. (Dev-session test calls carry a `devSessionId` in the list, so you can tell them from live traffic.)
|
|
453
260
|
|
|
454
261
|
---
|
|
455
262
|
|
|
@@ -457,8 +264,7 @@ Transcripts are how you iterate on a voice persona: after the user test-calls th
|
|
|
457
264
|
|
|
458
265
|
## Spec: `src/interfaces/voice.md`
|
|
459
266
|
|
|
460
|
-
Frontmatter holds the structured fields; the body is the persona plus an explicit `## Tools`
|
|
461
|
-
section.
|
|
267
|
+
Frontmatter holds the structured fields; the body is the persona plus an explicit `## Tools` section.
|
|
462
268
|
|
|
463
269
|
```yaml
|
|
464
270
|
---
|
|
@@ -475,15 +281,9 @@ Frontmatter fields:
|
|
|
475
281
|
|
|
476
282
|
- `name` — display name
|
|
477
283
|
- `description` — one-liner for listings
|
|
478
|
-
- `model` — JSON string, two shapes: native speech-to-speech `{"model": <realtime model id>,
|
|
479
|
-
|
|
480
|
-
|
|
481
|
-
- `turnDetection` — optional; `{"eagerness": "low" | "medium" | "high"}` — how quickly the platform
|
|
482
|
-
decides the user finished speaking. High is snappier; low is more patient (users dictating
|
|
483
|
-
numbers or addresses). Default `medium`.
|
|
484
|
-
- `greeting` — optional spoken opener, delivered on session start. Omit and the agent waits for the
|
|
485
|
-
user to speak first. Verbatim on cascaded engines; model-spoken (may paraphrase) on
|
|
486
|
-
speech-to-speech.
|
|
284
|
+
- `model` — JSON string, two shapes: native speech-to-speech `{"model": <realtime model id>, "voice": <voice id>}`, or cascaded `{"llm": <chat model id>, "stt": <transcription model id>, "tts": <speech model id>, "voice": <voice id>}`. Ids via `askMindStudioSdk`.
|
|
285
|
+
- `turnDetection` — optional; `{"eagerness": "low" | "medium" | "high"}` — how quickly the platform decides the user finished speaking. High is snappier; low is more patient (users dictating numbers or addresses). Default `medium`.
|
|
286
|
+
- `greeting` — optional spoken opener, delivered on session start. Omit and the agent waits for the user to speak first. Verbatim on cascaded engines; model-spoken (may paraphrase) on speech-to-speech.
|
|
487
287
|
|
|
488
288
|
Body: persona prose (voice register), then the toolset:
|
|
489
289
|
|
|
@@ -501,8 +301,7 @@ booking id aloud unless asked.
|
|
|
501
301
|
~~~
|
|
502
302
|
```
|
|
503
303
|
|
|
504
|
-
`latency` is one of `fast` / `slow` / `background` (semantics in "The latency classes" above).
|
|
505
|
-
Don't hand-author input schemas — the platform derives them from the method contract.
|
|
304
|
+
`latency` is one of `fast` / `slow` / `background` (semantics in "The latency classes" above). Don't hand-author input schemas — the platform derives them from the method contract.
|
|
506
305
|
|
|
507
306
|
## Compiled Output: `dist/interfaces/voice/`
|
|
508
307
|
|
|
@@ -560,95 +359,47 @@ Declare it in `mindstudio.json`:
|
|
|
560
359
|
|
|
561
360
|
## Session context (auto-loaded)
|
|
562
361
|
|
|
563
|
-
When the config declares `"context": { "method": "session-context" }`, the platform fires that
|
|
564
|
-
backend method automatically when a session starts — in the background, so the greeting is
|
|
565
|
-
never delayed — and appends its return to the system prompt as a `## Session Context` block.
|
|
566
|
-
Use it for situational state that should color every turn: the caller's open orders, account
|
|
567
|
-
standing, where they left off. Timing: the method runs while the greeting audio plays, so the
|
|
568
|
-
agent has the context by roughly the first exchange and is guaranteed to have it shortly
|
|
569
|
-
after — it is NOT guaranteed for the literal first utterance. On inbound phone calls it
|
|
570
|
-
re-fires after the caller verifies mid-call, so the context recomputes for the now-known user.
|
|
362
|
+
When the config declares `"context": { "method": "session-context" }`, the platform fires that backend method automatically when a session starts — in the background, so the greeting is never delayed — and appends its return to the system prompt as a `## Session Context` block. Use it for situational state that should color every turn: the caller's open orders, account standing, where they left off. Timing: the method runs while the greeting audio plays, so the agent has the context by roughly the first exchange and is guaranteed to have it shortly after — it is NOT guaranteed for the literal first utterance. On inbound phone calls it re-fires after the caller verifies mid-call, so the context recomputes for the now-known user.
|
|
571
363
|
|
|
572
364
|
The method contract:
|
|
573
|
-
- Runs as the session's user (same identity/RBAC as a tool call); anonymous sessions run it
|
|
574
|
-
|
|
575
|
-
-
|
|
576
|
-
|
|
577
|
-
|
|
578
|
-
- It is never a model-visible tool, and failures degrade silently to the generic prompt —
|
|
579
|
-
never make correctness depend on it.
|
|
580
|
-
|
|
581
|
-
Rule of thumb: `context` for always-relevant state the agent should just know; tools for
|
|
582
|
-
anything looked up on demand. Identity itself (name, roles) is already injected via the
|
|
583
|
-
Current User block — don't re-fetch it in the context method.
|
|
365
|
+
- Runs as the session's user (same identity/RBAC as a tool call); anonymous sessions run it anonymously — return generic or empty content for them.
|
|
366
|
+
- Return a short markdown **string** (a few lines). Results are capped at 4,000 characters; keep it situational context, not documents — deep or on-demand data belongs in tools or data sources.
|
|
367
|
+
- It is never a model-visible tool, and failures degrade silently to the generic prompt — never make correctness depend on it.
|
|
368
|
+
|
|
369
|
+
Rule of thumb: `context` for always-relevant state the agent should just know; tools for anything looked up on demand. Identity itself (name, roles) is already injected via the Current User block — don't re-fetch it in the context method.
|
|
584
370
|
|
|
585
371
|
## Platform Behavior
|
|
586
372
|
|
|
587
373
|
- Input schemas are derived from each method's contract — never hand-written.
|
|
588
|
-
- The platform appends a `## Current User` block (email, phone, roles) to the system prompt at
|
|
589
|
-
|
|
590
|
-
-
|
|
591
|
-
|
|
592
|
-
the only knob (not yet wired on Gemini realtime engines — it's a no-op there).
|
|
593
|
-
- Sessions have a per-app concurrency limit and a maximum duration, both configurable in the app's
|
|
594
|
-
settings; an idle session is ended gracefully after a prompt. Voice minutes and model usage are
|
|
595
|
-
metered.
|
|
596
|
-
- Every session persists as a call record with a transcript, visible in the dashboard and readable
|
|
597
|
-
from the frontend via `voice.listSessions()` / `voice.getSession(id)`.
|
|
374
|
+
- The platform appends a `## Current User` block (email, phone, roles) to the system prompt at runtime; never author a placeholder for it.
|
|
375
|
+
- Turn detection, barge-in (interruption truncates the agent's context to the audio the user actually heard), and background-noise handling are platform-managed; `turnDetection.eagerness` is the only knob (not yet wired on Gemini realtime engines — it's a no-op there).
|
|
376
|
+
- Sessions have a per-app concurrency limit and a maximum duration, both configurable in the app's settings; an idle session is ended gracefully after a prompt. Voice minutes and model usage are metered.
|
|
377
|
+
- Every session persists as a call record with a transcript, visible in the dashboard and readable from the frontend via `voice.listSessions()` / `voice.getSession(id)`.
|
|
598
378
|
|
|
599
379
|
## Auth
|
|
600
380
|
|
|
601
|
-
**Every voice config declares an `auth` block.** A voice session spends the owner's money for its
|
|
602
|
-
entire duration without necessarily touching a backend method, so the platform gates session
|
|
603
|
-
creation itself:
|
|
381
|
+
**Every voice config declares an `auth` block.** A voice session spends the owner's money for its entire duration without necessarily touching a backend method, so the platform gates session creation itself:
|
|
604
382
|
|
|
605
383
|
```json
|
|
606
384
|
"auth": { "requireUser": true, "requireRole": ["member"] }
|
|
607
385
|
```
|
|
608
386
|
|
|
609
|
-
- `requireUser: true` — only authenticated app users may start a session; `false` — anyone,
|
|
610
|
-
|
|
611
|
-
front-desk line).
|
|
612
|
-
- `requireRole` (optional) — the user must hold **at least one** of the listed manifest role ids
|
|
613
|
-
(OR semantics, same as the backend `auth.requireRole(...)`). Omit or leave empty for no role
|
|
614
|
-
gate. Requires `requireUser: true`. Unknown role ids fail the build.
|
|
387
|
+
- `requireUser: true` — only authenticated app users may start a session; `false` — anyone, including anonymous visitors. Most apps want `true`; choose `false` deliberately (a public front-desk line).
|
|
388
|
+
- `requireRole` (optional) — the user must hold **at least one** of the listed manifest role ids (OR semantics, same as the backend `auth.requireRole(...)`). Omit or leave empty for no role gate. Requires `requireUser: true`. Unknown role ids fail the build.
|
|
615
389
|
- Denials reject `startSession()` with code `auth_required` (401) or `role_required` (403).
|
|
616
|
-
- On the phone channel there is no login page to bounce to, so `requireUser` becomes
|
|
617
|
-
answer-then-verify — see "Inbound calls" above.
|
|
390
|
+
- On the phone channel there is no login page to bounce to, so `requireUser` becomes answer-then-verify — see "Inbound calls" above.
|
|
618
391
|
- Dev preview is exempt — the builder is never locked out while testing.
|
|
619
|
-
- Older compiled apps without the block fall back to the manifest's `auth.enabled` (auth-enabled →
|
|
620
|
-
users only; no auth → public). New configs always declare it explicitly.
|
|
392
|
+
- Older compiled apps without the block fall back to the manifest's `auth.enabled` (auth-enabled → users only; no auth → public). New configs always declare it explicitly.
|
|
621
393
|
|
|
622
|
-
Once inside, voice sessions run as the **authenticated user** — every tool call carries that
|
|
623
|
-
user's roles, so a method gated with `auth.requireRole` behaves exactly as it would from the web
|
|
624
|
-
frontend or the agent interface. Anonymous sessions (when allowed) have no user and no roles:
|
|
625
|
-
gated methods reject, and the caller's history is scoped to their browser's visitor identity.
|
|
626
|
-
That's why role restrictions belong in the tool descriptions — the agent should decline in
|
|
627
|
-
character, not relay a rejection.
|
|
394
|
+
Once inside, voice sessions run as the **authenticated user** — every tool call carries that user's roles, so a method gated with `auth.requireRole` behaves exactly as it would from the web frontend or the agent interface. Anonymous sessions (when allowed) have no user and no roles: gated methods reject, and the caller's history is scoped to their browser's visitor identity. That's why role restrictions belong in the tool descriptions — the agent should decline in character, not relay a rejection.
|
|
628
395
|
|
|
629
396
|
### Progressive auth: verify mid-call without dropping the conversation
|
|
630
397
|
|
|
631
|
-
The best pattern for apps that allow anonymous sessions (`requireUser: false`): let visitors
|
|
632
|
-
|
|
633
|
-
|
|
634
|
-
|
|
635
|
-
|
|
636
|
-
|
|
637
|
-
|
|
638
|
-
|
|
639
|
-
2. **The frontend opens its verification sheet off the same signal.** It already receives every
|
|
640
|
-
tool's `toolCall` event (and the tool's return in `result`) — when an account tool fires (or
|
|
641
|
-
returns `verified: false`) while the app has no signed-in user, open the sheet.
|
|
642
|
-
3. **The sheet runs the platform's auth rails** — `auth.sendSmsCode()` / `auth.verifySmsCode()`
|
|
643
|
-
(or the email pair) from `@mindstudio-ai/interface`. On success the app's session becomes the
|
|
644
|
-
verified user.
|
|
645
|
-
4. **Hand the verified session back to the live call**: `await session.refreshIdentity()`. The
|
|
646
|
-
platform upgrades the running voice session in place — subsequent tool calls carry the user's
|
|
647
|
-
identity and roles, and the agent's Current User context refreshes — no teardown, no lost
|
|
648
|
-
conversation. (Phone calls don't need this: they verify through the agent's built-in flow.)
|
|
649
|
-
|
|
650
|
-
`refreshIdentity()` is upgrade-only (anonymous → signed-in; an already-identified session rejects
|
|
651
|
-
with `already_identified`) and requires the session to have been started by this same browser. If
|
|
652
|
-
it fails, ending and restarting the session is the graceful fallback. The chat sibling for agent
|
|
653
|
-
interfaces is `claimThread(threadId)` — anonymous threads become unreachable after login until
|
|
654
|
-
claimed.
|
|
398
|
+
The best pattern for apps that allow anonymous sessions (`requireUser: false`): let visitors explore by voice, and verify only when they hit an account-bound action — without killing the live call. Four pieces, all platform rails:
|
|
399
|
+
|
|
400
|
+
1. **Account-gated tools return a standard not-verified shape** instead of doing the work: `{ verified: false, message: 'The caller is not verified. Offer to verify them before sharing account details.' }`. The agent speaks the offer in character (reinforce tone in the system prompt's verification section). Check with the agent SDK's `auth.userId` inside the method.
|
|
401
|
+
2. **The frontend opens its verification sheet off the same signal.** It already receives every tool's `toolCall` event (and the tool's return in `result`) — when an account tool fires (or returns `verified: false`) while the app has no signed-in user, open the sheet.
|
|
402
|
+
3. **The sheet runs the platform's auth rails** — `auth.sendSmsCode()` / `auth.verifySmsCode()` (or the email pair) from `@mindstudio-ai/interface`. On success the app's session becomes the verified user.
|
|
403
|
+
4. **Hand the verified session back to the live call**: `await session.refreshIdentity()`. The platform upgrades the running voice session in place — subsequent tool calls carry the user's identity and roles, and the agent's Current User context refreshes — no teardown, no lost conversation. (Phone calls don't need this: they verify through the agent's built-in flow.)
|
|
404
|
+
|
|
405
|
+
`refreshIdentity()` is upgrade-only (anonymous → signed-in; an already-identified session rejects with `already_identified`) and requires the session to have been started by this same browser. If it fails, ending and restarting the session is the graceful fallback. The chat sibling for agent interfaces is `claimThread(threadId)` — anonymous threads become unreachable after login until claimed.
|