@mindstudio-ai/remy 0.1.265 → 0.1.267

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -210,7 +210,7 @@ export async function createPurchaseOrder(input: {
210
210
 
211
211
  ## Fire-and-Forget Background Tasks
212
212
 
213
- A method can return immediately while kicking off slow work (like `runTask()`) that continues in the background. Don't await the slow call — use `.then()` / `.catch()` to update the record when it completes, and return an early result to the caller. The frontend polls the record's status to track progress.
213
+ A method can return immediately while kicking off slow work (like `runTask()`) that continues in the background. Don't await the slow call — use `.then()` / `.catch()` to update the record when it completes, and return an early result to the caller. The frontend polls the record's status to track progress. Wrap background chains other than `runTask()` in `mindstudio.waitUntil(...)` so the platform keeps the sandbox alive for them and records an interruption if they're cut short — `runTask()` registers itself automatically.
214
214
 
215
215
  The example below shows the fire-and-forget shape, not a complete `runTask()` call. Load the `taskAgents` skill before writing one — configuring its tools, validating the output, and handling failures are all there, and none of them are visible here.
216
216
 
@@ -103,6 +103,30 @@ const { threads, nextCursor } = await chat.listThreads();
103
103
  const full = await chat.getThread(thread.id);
104
104
  await chat.updateThread(thread.id, 'New title');
105
105
  await chat.deleteThread(thread.id);
106
+
107
+ // Progressive auth: threads started anonymously become unreachable after the
108
+ // user signs in (login replaces the session and its visitor identity). Claim
109
+ // them right after your verification/login succeeds so the conversation
110
+ // survives — the client remembers each thread's pre-login token automatically
111
+ // for threads touched this page session.
112
+ await chat.claimThread(thread.id);
113
+ ```
114
+
115
+ **Client tools** — a tool whose effect happens in the browser (open a sheet, navigate, highlight)
116
+ is declared with `target: "client"` and a `name` + inline `inputSchema` instead of a `method`
117
+ (names must not collide with method ids; the schema is authored — there's no method contract to
118
+ derive it from). The agent's invocation arrives as the `client_tool_call` stream event / the
119
+ `onClientToolCall` callback on `sendMessage`; run the action there. Fire-and-forget on this
120
+ surface: the agent is told the action was displayed and keeps going — the user's next message
121
+ closes the loop.
122
+
123
+ ```js
124
+ await chat.sendMessage(thread.id, text, {
125
+ onText: (delta) => append(delta),
126
+ onClientToolCall: (name, input) => {
127
+ if (name === 'showVerification') openVerifySheet(input);
128
+ },
129
+ });
106
130
  ```
107
131
 
108
132
  **Sending messages (streaming):**
@@ -18,7 +18,7 @@ This is one of the most powerful pieces of the MindStudio SDK, and it can turn a
18
18
 
19
19
  This is the tool to reach for whenever a feature would be dramatically more compelling if the app could autonomously research, enrich, or create on behalf of the user. Think about the difference between "user enters a restaurant name and it gets saved" vs. "user enters a restaurant name and gets back a fully researched, illustrated card." Task agents close that gap.
20
20
 
21
- Run tasks in the background — depending on complexity they can take time to complete. Return an early partial result to the user and upsert later with the final result when the agent finishes.
21
+ Run tasks in the background — depending on complexity they can take time to complete. Return an early partial result to the user and upsert later with the final result when the agent finishes. The exception is cron and email triggers: there is no user waiting there, so `await` the task instead — awaiting is what surfaces its failures in the run's own result.
22
22
 
23
23
  - **Research and enrichment:** "Given this email, find the person's LinkedIn, role, company, and a headshot" — the model searches, scrapes, extracts, and assembles structured data.
24
24
  - **Content creation pipelines:** "Write SEO copy for this product in 3 languages, generate a hero image, extract keywords" — the model calls text generation, image generation, and analysis actions as needed.
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: Voice Interfaces
3
3
  what: Realtime voice conversation as a first-class interface — the user talks to the app and its voice agent talks back in sub-second, interruptible speech, calling the app's methods mid-conversation as the authenticated user. The platform handles the media transport, turn-taking, barge-in, and transcripts, so the work is authorship — a persona written for the ear, a small toolset where every tool carries a latency class, and descriptions that say results out loud. Any app whose methods do something interesting can pick up a voice, and it is often the most impressive surface it has.
4
- when: Before authoring `src/interfaces/voice.md`, choosing a voice model or pipeline, deciding which methods a voice agent gets, or building the voice UI with `createVoiceClient()`.
4
+ when: Before authoring `src/interfaces/voice.md`, choosing a voice model or pipeline, deciding which methods a voice agent gets, building the voice UI with `createVoiceClient()`, or working out why a voice agent behaved the way it did on a call.
5
5
  ---
6
6
 
7
7
  # Building Voice Interfaces
@@ -53,11 +53,11 @@ than any other surface.
53
53
  Structure the compiled prompt as short **labeled sections** — Role & Objective, Personality & Tone,
54
54
  Rules, and (when the app has a real call flow) Conversation Flow — with bullets over paragraphs;
55
55
  realtime models find and follow sectioned rules far more reliably than prose. Scope rules
56
- precisely: "confirm before any tool that changes data," not "always confirm everything" — blanket
57
- `always`/`never` makes the agent rigid and unable to handle reasonable exceptions. And start
58
- minimal: state the role, the boundaries, and the voice mechanics above, then add rules only for
56
+ precisely; blanket `always`/`never` makes the agent rigid and unable to handle reasonable exceptions.
57
+ And start minimal: state the role, the boundaries, and the voice mechanics above, then add rules only for
59
58
  behaviors that actually misfire in test calls (the transcripts in the call log are the feedback
60
- loop) rather than front-loading a policy manual.
59
+ loop — `mindstudio-prod voice sessions get` reads a call verbatim) rather than front-loading a
60
+ policy manual.
61
61
 
62
62
  ### The latency classes
63
63
 
@@ -78,25 +78,65 @@ Classify by how the method actually behaves, not by what it is named. A "lookup"
78
78
  external service is `slow`. When in doubt between `fast` and `slow`, pick `slow` — a needless
79
79
  preamble is mildly chatty; an unexplained silence feels broken.
80
80
 
81
- ### Forwarding results to the screen (`forwardResult`)
81
+ ### Tool results reach the screen
82
82
 
83
- A tool block may declare `forwardResult: true`. On completion, the platform then delivers the
84
- tool's raw return value to the session's browser on the SDK's `toolCall` event (`result` field) —
85
- so the UI can render what the agent just did (the citation it found, the record it pulled up, the
86
- booking it made) in lockstep with the spoken answer. No polling, no key-threading, no model
87
- involvement: the correlation is platform-guaranteed and scoped to that one session's client.
83
+ Every successful tool call delivers its raw return value to the session's browser on the SDK's
84
+ `toolCall` event (`result` field, on `done`) — so the UI can render what the agent just did (the
85
+ citation it found, the record it pulled up, the booking it made) in lockstep with the spoken
86
+ answer. No flag, no polling, no key-threading, no model involvement: delivery to the invoking
87
+ session's own client is the same security context as the invocation itself (an RPC response), and
88
+ it's scoped to that one session.
88
89
 
89
- Opt in deliberately, per tool. The forwarded payload is the method's raw return — the same data
90
- the model sees — so only enable it on tools whose returns are safe to render for the user in the
91
- call (no internal fields you wouldn't show on screen). Payloads over ~32KB serialized arrive as
92
- `resultTruncated: true` with no data — keep forwarded returns compact, or have the UI fetch big
93
- data itself. Failed calls never forward anything.
90
+ Consequence for authoring: **a tool's return value is user-visible by definition.** Return what
91
+ the user may see — no internal fields, keys, or diagnostics you wouldn't put on screen (the same
92
+ discipline as agent-interface tools, whose results render in chat). Payloads over ~32KB serialized
93
+ arrive as `resultTruncated: true` with no data — keep returns compact, or have the UI fetch big
94
+ data itself. Failed calls deliver nothing to the client (the model gets the `{ error }` and speaks
95
+ a decline).
94
96
 
95
97
  For backend-side correlation (writing results to a table keyed by the call, custom channels), the
96
98
  method itself can read `session.voiceSessionId` / `session.visitorId` from the agent SDK
97
99
  (`import { session } from '@mindstudio-ai/agent'`) — the same id the browser holds as
98
100
  `session.sessionId`, guaranteed by the platform rather than echoed by the model.
99
101
 
102
+ ### Client tools: actions that happen on screen (`target: "client"`)
103
+
104
+ A tool whose effect belongs in the browser — open the verification sheet, navigate to a page,
105
+ highlight a record — is declared with `target: "client"` instead of a `method`:
106
+
107
+ ```json
108
+ {
109
+ "target": "client",
110
+ "name": "showVerification",
111
+ "description": "tools/showVerification.md",
112
+ "inputSchema": { "type": "object", "properties": { "reason": { "type": "string" } } }
113
+ }
114
+ ```
115
+
116
+ The platform never touches the backend for these: the agent's invocation is delivered to the
117
+ session's browser, the app's registered handler runs, and the handler's **return value goes back
118
+ to the agent as the tool result** — a real request/response, so the agent knows the sheet
119
+ actually opened (or that the user dismissed it) and speaks accordingly. Rules:
120
+
121
+ - `name` instead of `method`; must not collide with any backend method id. No latency class —
122
+ the agent holds the turn while the browser responds (up to ~30s, then a timeout error).
123
+ - `inputSchema` is authored inline (an object schema) — there's no method contract to derive
124
+ it from. Keep it small; these are UI directives, not data payloads.
125
+ - The frontend must register a handler, or invocations fail as `unhandled_client_tool`:
126
+
127
+ ```js
128
+ session.registerClientTool('showVerification', async ({ reason }) => {
129
+ openVerifySheet(reason);
130
+ return { opened: true }; // what the agent hears back
131
+ });
132
+ ```
133
+
134
+ - One client tool runs at a time per session; the description should tell the agent when to use
135
+ it and what to say while it's on screen. Throwing from the handler (or returning nothing)
136
+ becomes an error/ack the agent can speak around.
137
+ - The progressive-auth pattern above is the canonical use: make the verification sheet a client
138
+ tool and the agent opens it deliberately instead of the frontend inferring it from tool events.
139
+
100
140
  ### Tool descriptions say results out loud
101
141
 
102
142
  Follow the agent-interface principles for tool descriptions (when to use and when not, parameter
@@ -124,24 +164,45 @@ caller.
124
164
 
125
165
  ### Choosing the model
126
166
 
127
- Two shapes, one `model` field:
128
-
129
- - **Native speech-to-speech** (`{"model": ..., "voice": ...}`) — one realtime model hears and
130
- speaks. Lowest latency, most natural prosody, hears tone and hesitation. The default for
131
- personality-forward, conversational apps.
132
- - **Cascaded** (`{"llm": ..., "stt": ..., "tts": ..., "voice": ...}`) — streaming transcription
133
- into any chat model in the catalog, streaming speech out. Slightly higher latency, but the brain
134
- can be *any* chat model — the right choice when the app's reasoning demands a specific model, or
135
- when the agent interface already uses one and the voice should think identically. The blessed
136
- streaming pairing is `"stt": "deepgram-nova-3", "tts": "cartesia-sonic-3"` — the lowest-latency
137
- combination the platform wires; prefer it unless there's a reason not to (ElevenLabs TTS,
138
- `"tts": "elevenlabs-tts"`, is also wired when its voice library fits better). One nuance: cascaded
139
- engines speak the `greeting` verbatim (they have a real TTS); speech-to-speech engines have the
140
- model say it, so it may paraphrase slightly.
141
-
142
- Ask `askMindStudioSdk` for available ids — realtime, transcription, and speech models are separate
143
- catalogs, and MindStudio ids don't match vendor ids, so treat ids in this document as illustrative.
144
- Voice ids are model-specific; query for those too. The user's UI has a picker for changing the model
167
+ Two shapes, one `model` field. **Use native speech-to-speech unless the user specifically asks
168
+ for a cascaded pipeline** — one realtime model hears and speaks: lowest latency, most natural
169
+ prosody, hears tone and hesitation.
170
+
171
+ **Native** (`{"model": ..., "voice": ...}`) — **default to `gpt-realtime-2.1` with voice
172
+ `marin`.**
173
+
174
+ - `gpt-realtime-2.1` — the default. Voices: `marin` (default), `cedar`, `alloy`, `ash`,
175
+ `ballad`, `coral`, `echo`, `sage`, `shimmer`, `verse`.
176
+ - `gpt-realtime-2.1-mini` — the same family, lighter; same voices.
177
+ - `gemini-2.5-flash-native-audio-preview-12-2025` — the Gemini pick, with a large expressive
178
+ roster: `Puck` (default, upbeat), `Zephyr` (bright), `Charon` (informative), `Kore` (firm),
179
+ `Fenrir` (excitable), `Leda` (youthful), `Orus` (firm), `Aoede` (breezy), `Callirrhoe`
180
+ (easy-going), `Autonoe` (bright), `Enceladus` (breathy), `Iapetus` (clear), `Umbriel`
181
+ (easy-going), `Algieba` (smooth), `Despina` (smooth), `Erinome` (clear), `Algenib` (gravelly),
182
+ `Rasalgethi` (informative), `Laomedeia` (upbeat), `Achernar` (soft), `Alnilam` (firm),
183
+ `Schedar` (even), `Gacrux` (mature), `Pulcherrima` (forward), `Achird` (friendly),
184
+ `Zubenelgenubi` (casual), `Vindemiatrix` (gentle), `Sadachbia` (lively), `Sadaltager`
185
+ (knowledgeable), `Sulafat` (warm).
186
+ - `gemini-3.1-flash-live-preview` — newer Gemini, same voices as 2.5, but currently can't speak
187
+ an opening greeting or take mid-call prompt updates (a plugin limitation expected to resolve
188
+ upstream) — prefer 2.5 until then.
189
+ - `grok-voice-think-fast-2.0` — a distinct personality register. Voices: `eve` (default),
190
+ `altair`, `ara`, `atlas`, `aurora`, `carina`, `castor`, `celeste`, `cosmo`, `helios`, `helix`,
191
+ `iris`, `kepler`, `leo`, `liora`, `lumen`, `luna`, `lux`, `naksh`, `orion`, `perseus`, `rex`,
192
+ `rigel`, `sal`, `sirius`, `ursa`, `zagan`, `zenith`.
193
+
194
+ **Cascaded** (`{"llm": ..., "stt": ..., "tts": ..., "voice": ...}`) — streaming transcription
195
+ into any chat model in the catalog, streaming speech out. Slightly higher latency; reach for it
196
+ only when the user wants it or the app's reasoning demands a specific chat model (e.g. the agent
197
+ interface already uses one and the voice should think identically). Slots: `stt` is
198
+ `deepgram-nova-3`; `tts` is `cartesia-sonic-3` (voices are per-account Cartesia UUIDs — see
199
+ play.cartesia.ai) or `elevenlabs-tts` (the account's ElevenLabs voice library); `llm` is any chat
200
+ model — ask `askMindStudioSdk` for chat model ids. One nuance: cascaded engines speak the
201
+ `greeting` verbatim (they have a real TTS); speech-to-speech engines have the model say it, so it
202
+ may paraphrase slightly.
203
+
204
+ The model and voice ids above are current and maintained with the platform — use them as written
205
+ (they are MindStudio ids, not vendor ids). The user's UI has a picker for changing the model
145
206
  later, so validate only when you set it.
146
207
 
147
208
  ### Seeding from an existing agent
@@ -204,13 +265,15 @@ session.on('stateChange', (state) => { }); // on() returns an unsubscri
204
265
  // far (never a delta) — render by upserting on segmentId, not appending.
205
266
  session.on('transcript', ({ role, segmentId, text, final }) => { });
206
267
 
207
- // status: 'running' | 'done' | 'failed'. Tools declared with `forwardResult: true`
208
- // carry their return value in `result` on 'done' (or `resultTruncated: true` if >~32KB).
268
+ // status: 'running' | 'done' | 'failed'. Every 'done' carries the tool's raw
269
+ // return value in `result` (or `resultTruncated: true` if >~32KB serialized).
209
270
  session.on('toolCall', ({ method, status, result }) => { });
210
271
  session.on('error', (err) => { });
211
272
 
212
273
  session.mute(); session.unmute(); session.isMuted;
213
274
  session.sendText('123 Main Street'); // inject text into the live conversation
275
+ await session.refreshIdentity(); // after in-app verification — upgrade the
276
+ // live session anonymous → signed-in in place
214
277
  session.end();
215
278
  ```
216
279
 
@@ -277,7 +340,7 @@ export async function callMeAboutMyOrder(input: { phone: string }) {
277
340
  session, not the phone). Omitted/false → anonymous call; role-gated tools decline.
278
341
  System/cron invocations have no human identity and always run anonymously.
279
342
  - **Production needs a dedicated phone number.** The app owner attaches one ($1/month) via the
280
- dashboard or `mindstudio-prod voice numbers` (see "Managing the phone side from the CLI"
343
+ dashboard or `mindstudio-prod voice numbers` (see "The voice CLI"
281
344
  below) — it becomes the caller ID for every call, in dev sessions too, so users always see
282
345
  the same number. Without one, deployed calls throw `phone_out_requires_dedicated_number`, and
283
346
  dev sessions fall back to a shared platform test number that varies per call (tighter limits
@@ -338,7 +401,7 @@ wrong one wherever the agent's tools can move money, reveal sensitive records, o
338
401
  destructive actions. It lives in the interface config deliberately: enabling it is a code
339
402
  change, visible in review and auditable via deploys, not a dashboard toggle.
340
403
 
341
- ## Managing the phone side from the CLI
404
+ ## The voice CLI
342
405
 
343
406
  The `mindstudio-prod voice` family covers numbers, the call log, and voice policy:
344
407
 
@@ -377,7 +440,7 @@ section.
377
440
  name: Front Desk
378
441
  description: Books appointments and answers questions by voice.
379
442
  type: interface/voice
380
- model: {"model": "gpt-realtime-mini", "voice": "marin"}
443
+ model: {"model": "gpt-realtime-2.1", "voice": "marin"}
381
444
  turnDetection: {"eagerness": "medium"}
382
445
  greeting: Hey! I can help you book, reschedule, or answer questions — what do you need?
383
446
  ---
@@ -435,7 +498,7 @@ The top-level key must match the interface type (`voice`):
435
498
  "voice": {
436
499
  "name": "Front Desk",
437
500
  "description": "Books appointments and answers questions by voice.",
438
- "model": "gpt-realtime-mini",
501
+ "model": "gpt-realtime-2.1",
439
502
  "voice": "marin",
440
503
  "turnDetection": { "eagerness": "medium" },
441
504
  "greeting": "Hey! I can help you book, reschedule, or answer questions — what do you need?",
@@ -537,3 +600,30 @@ frontend or the agent interface. Anonymous sessions (when allowed) have no user
537
600
  gated methods reject, and the caller's history is scoped to their browser's visitor identity.
538
601
  That's why role restrictions belong in the tool descriptions — the agent should decline in
539
602
  character, not relay a rejection.
603
+
604
+ ### Progressive auth: verify mid-call without dropping the conversation
605
+
606
+ The best pattern for apps that allow anonymous sessions (`requireUser: false`): let visitors
607
+ explore by voice, and verify only when they hit an account-bound action — without killing the
608
+ live call. Four pieces, all platform rails:
609
+
610
+ 1. **Account-gated tools return a standard not-verified shape** instead of doing the work:
611
+ `{ verified: false, message: 'The caller is not verified. Offer to verify them before sharing
612
+ account details.' }`. The agent speaks the offer in character (reinforce tone in the system
613
+ prompt's verification section). Check with the agent SDK's `auth.userId` inside the method.
614
+ 2. **The frontend opens its verification sheet off the same signal.** It already receives every
615
+ tool's `toolCall` event (and the tool's return in `result`) — when an account tool fires (or
616
+ returns `verified: false`) while the app has no signed-in user, open the sheet.
617
+ 3. **The sheet runs the platform's auth rails** — `auth.sendSmsCode()` / `auth.verifySmsCode()`
618
+ (or the email pair) from `@mindstudio-ai/interface`. On success the app's session becomes the
619
+ verified user.
620
+ 4. **Hand the verified session back to the live call**: `await session.refreshIdentity()`. The
621
+ platform upgrades the running voice session in place — subsequent tool calls carry the user's
622
+ identity and roles, and the agent's Current User context refreshes — no teardown, no lost
623
+ conversation. (Phone calls don't need this: they verify through the agent's built-in flow.)
624
+
625
+ `refreshIdentity()` is upgrade-only (anonymous → signed-in; an already-identified session rejects
626
+ with `already_identified`) and requires the session to have been started by this same browser. If
627
+ it fails, ending and restarting the session is the graceful fallback. The chat sibling for agent
628
+ interfaces is `claimThread(threadId)` — anonymous threads become unreachable after login until
629
+ claimed.
@@ -6,7 +6,7 @@
6
6
 
7
7
  ## Principles
8
8
  - The spec in `src/` is the source of truth for what the app does: consult it before making behavioral code changes, and keep it in sync as the app evolves by dispatching periodic requests to the specSync tool after meaningful changes. Some amount of drift is fine and inevitable - your priority is working efficiently with the user, let the specSync tool handle things in the background.
9
- - Keep `src/overview.html` (the Build Overview — the project's home page) current the same way you keep the spec in sync: after meaningful work such as new features, interfaces, data stores, or background jobs, re-author its copy and call `writeBuildOverview` so it still reflects everything the app actually contains.
9
+ - The Build Overview (`src/overview.html` — the project's home page) stays current via specSync: after a deploy or a large milestone, set `refreshBuildOverview: true` on your specSync dispatch and it re-authors the overview from the updated spec. Call `writeBuildOverview` directly only for the initial end-of-build generation or when the user explicitly asks.
10
10
  - Change only what the task requires. Match existing styles. Keep solutions simple.
11
11
  - Read files before editing them. Understand the context before making changes.
12
12
  - When the user asks you to make a change, execute it fully — all steps, no pausing for confirmation. Use `confirmDestructiveAction` to gate before destructive or irreversible actions (e.g., deleting data, resetting the database). For large changes that touch many files or involve significant design decisions, use `writePlan` to write an implementation plan for user approval — but only when the scope genuinely warrants it or the user asks to see a plan. The plan is saved to `.remy-plan.md` and the user can review, discuss, and refine it before approving. Do not begin implementation until the plan is approved. Most work should be done autonomously without a plan.
@@ -44,7 +44,9 @@ Your editor — a design expert for words. Hand it any user-facing copy — an e
44
44
 
45
45
  Your spec keeper. Once the app is built and you're iterating on it, whenever you make code changes that alter what the app does, hand it a brief, plain-language description of what you changed and why (prefer bullet points, batch multiple changes into one invocation). It finds the affected sections of the spec in `src/` and updates them to match what has been built. It reads the spec itself and decides what to touch, so you don't need to name files or locations.
46
46
 
47
- It always runs in the background: it returns immediately and you keep working while it reconciles, and you'll get notified when the updated spec lands later. So don't hunt through spec files to sync them yourself and don't wait on it — hand off and move on. You decide when the spec has drifted enough to be worth a hand-off (after a meaningful change, or a batch of them - you do not need to invoke this after every change or conversation turn - some amount of drift between code and spec is completely normal and acceptable).
47
+ It always runs in the background: it returns immediately and you keep working while it reconciles; it completes silently and the outcome appears as an automated note at the start of a later turn — it never wakes you. So don't hunt through spec files to sync them yourself and don't wait on it — hand off and move on. You decide when the spec has drifted enough to be worth a hand-off (after a meaningful change, or a batch of them - you do not need to invoke this after every change or conversation turn - some amount of drift between code and spec is completely normal and acceptable).
48
+
49
+ It can also keep the Build Overview current: set `refreshBuildOverview: true` and, after reconciling, it re-authors the overview copy from the updated spec and re-renders `src/overview.html`. Set it after a deploy or a large milestone; leave it off for routine syncs.
48
50
 
49
51
  ### QA (`runAutomatedBrowserTest`)
50
52
 
@@ -13,9 +13,11 @@ When the content you need to test is behind authentication, use the `setupBrowse
13
13
 
14
14
  If you need to test the login/signup flow itself (e.g., verifying the UI, error states, or the verification code input), navigate it manually: use `remy@mindstudio.ai` for email and `+15551234567` for phone. In the dev environment, verification codes are bypassed for this email and any 555-prefixed phone number — enter any 6-digit code (e.g., `123456`).
15
15
 
16
+ To test as a **signed-out visitor** (public pages, landing/join links), call `setupBrowser` with NO `auth` — it clears the auth cookie and reloads at the given path, giving you a clean unauthenticated session. Combine with `navigate` + `fresh: true` when you need a fresh-document view of an entry page mid-run.
17
+
16
18
  ## Browser Commands
17
19
 
18
- Your session always starts on the app root / in a logged out/unauthenticated state. Use `setupBrowser` to authenticate before testing protected pages.
20
+ Your session always starts on the app root / in a logged out/unauthenticated state, on a freshly reloaded page running the current code — any changes made since the last run are already picked up. Never restart the dev server (or reload manually) to clear a "stale bundle"; that staleness cannot survive the start-of-run refresh. Use `setupBrowser` to authenticate before testing protected pages.
19
21
 
20
22
  ### Snapshot format
21
23
 
@@ -40,13 +42,43 @@ Note: the snapshot concatenates inline text and strips whitespace. If you need t
40
42
  - `type`: Type text into an input. Characters appear one at a time. Set `clear: true` to clear the field first.
41
43
  - `select`: Select a dropdown option by text. Target the `<select>` element, set `option` to the option text.
42
44
  - `wait`: Wait for an element to appear (polls every 100ms, default 5s timeout). Also waits for network to settle after the element is found.
43
- - `navigate`: Navigate to a new URL within the app. Waits for the new page to load before continuing with subsequent steps. Use this instead of evaluate with `window.location.href` when you need to navigate and then continue interacting with the new page. Steps after navigate execute on the new page automatically.
45
+ - `navigate`: Navigate to a new URL within the app. Waits for the route to load before continuing with subsequent steps. Use this instead of evaluate with `window.location.href` when you need to navigate and then continue interacting with the new page. Steps after navigate execute on the new page automatically. Same-origin navigation is a soft in-app route change (like clicking a link in an SPA — in-memory app state survives); set `fresh: true` to force a real full page load with a fresh document instead. Use `fresh: true` when the test is about what a user sees on *entry* — landing pages, join/invite links, "what does a signed-out visitor see" — where reusing the SPA's in-memory state would test the wrong thing. The result reports the URL the page actually landed on, so if the app redirected you (e.g. an auth wall bounced you off a public page), you'll see the real destination — check it instead of assuming the navigation stuck.
44
46
  - `evaluate`: Run arbitrary JavaScript in the page and return the result.
45
47
  - `styles`: Read computed CSS styles from page elements. Pass a `properties` array with camelCase CSS property names (e.g., `["backgroundColor", "borderRadius", "fontSize"]`). Omit `properties` for a default set covering colors, typography, spacing, borders, shadows, dimensions, and layout. Uses the same targeting as click/type (ref, text, role, label, selector). Omit the target to get styles for all elements from the last snapshot.
46
48
  - `screenshotFullPage`: Take a screenshot of the whole page, top to bottom. Returns CDN url with full text analysis and dimensions. Use for overall composition or content past the fold.
47
49
  - `screenshotViewport`: Take a screenshot of the visible viewport. Returns CDN url with full text analysis and dimensions. To capture a specific section, set `scrollToSelector` (a CSS selector) — or `scrollY` (an absolute offset) — on this same step; it scrolls the target into view and captures it atomically, so you do NOT need a separate scroll step. Do not use if you can get what you need with other tools - only use when you need to visually see the viewport.
48
50
  - `setViewport`: Switch the browser between desktop and mobile rendering. Set `mode` to `"desktop"` or `"mobile"`. Mobile emulates a phone (390-wide, touch, device pixel ratio 2); desktop is the standard wide viewport. This reloads the page so media queries, responsive layouts, and `matchMedia` re-evaluate — the reload clears in-page state, so switch before you set up the state you want to inspect. The mode persists across navigations within a run. Each run starts in the app's default mode, so only use this when you need to check the other one.
49
51
 
52
+ ### Voice interfaces
53
+
54
+ Apps with a voice interface are testable end to end — the UI layer included. The sandbox browser
55
+ auto-grants a (silent) microphone, and while a session is live the SDK publishes a handle at
56
+ `window.__MS_VOICE__` so you can converse by text: the agent treats injected text exactly like
57
+ user speech (interrupts and replies), backend tools run for real, and client tools render their
58
+ real UI (cards, sheets) in the page.
59
+
60
+ The loop:
61
+
62
+ 1. Start a session through the app's real UI — `click` its voice affordance (orb/button). No mic
63
+ prompt appears. Then `wait` briefly and confirm the session is live:
64
+ `evaluate: window.__MS_VOICE__?.state` (undefined means no session started — report that,
65
+ don't improvise).
66
+ 2. Speak by injection: `evaluate: window.__MS_VOICE__.sendText("I'd like to book Tuesday at 2")`.
67
+ 3. Give the agent a few seconds to respond (replies are generated speech — slower than chat).
68
+ `wait` for the UI you expect (client-tool cards appear via the app's real handlers), and read
69
+ the conversation: `evaluate: window.__MS_VOICE__.transcript` (one entry per utterance, both
70
+ sides, `final` marks settled ones) and `window.__MS_VOICE__.toolCalls` (which tools ran;
71
+ `done` entries carry the tool's return value).
72
+ 4. Verify visuals with `screenshotViewport` like any other flow.
73
+ 5. Read `transcript`/`toolCalls` BEFORE ending — then `evaluate: window.__MS_VOICE__.end()` (the
74
+ handle is removed when the session ends).
75
+
76
+ Voice sessions are the most expensive thing you can run — real voice-model minutes are metered,
77
+ and the agent speaks its replies out loud even when you type at it. Keep voice tests short and
78
+ purposeful: a handful of turns that exercise the target behavior, then end the session. What you
79
+ cannot test is the audio layer itself (mishearing, interruptions, pronunciation) — never attempt
80
+ to simulate audio; report that scope limit instead.
81
+
50
82
  ### Element targeting (tried in order)
51
83
 
52
84
  1. `ref`: From the last snapshot. Most reliable.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@mindstudio-ai/remy",
3
- "version": "0.1.265",
3
+ "version": "0.1.267",
4
4
  "description": "Remy coding agent",
5
5
  "repository": {
6
6
  "type": "git",