@mindstudio-ai/remy 0.1.265 → 0.1.267
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +13 -2
- package/dist/automatedActions/publish.md +1 -0
- package/dist/headless.d.ts +17 -1
- package/dist/headless.js +1397 -954
- package/dist/index.js +1424 -977
- package/dist/prompt/compiled/methods.md +1 -1
- package/dist/prompt/skills/agentInterfaces.md +24 -0
- package/dist/prompt/skills/taskAgents.md +1 -1
- package/dist/prompt/skills/voiceInterfaces.md +130 -40
- package/dist/prompt/static/instructions.md +1 -1
- package/dist/prompt/static/team.md +3 -1
- package/dist/subagents/browserAutomation/prompt.md +34 -2
- package/package.json +1 -1
|
@@ -210,7 +210,7 @@ export async function createPurchaseOrder(input: {
|
|
|
210
210
|
|
|
211
211
|
## Fire-and-Forget Background Tasks
|
|
212
212
|
|
|
213
|
-
A method can return immediately while kicking off slow work (like `runTask()`) that continues in the background. Don't await the slow call — use `.then()` / `.catch()` to update the record when it completes, and return an early result to the caller. The frontend polls the record's status to track progress.
|
|
213
|
+
A method can return immediately while kicking off slow work (like `runTask()`) that continues in the background. Don't await the slow call — use `.then()` / `.catch()` to update the record when it completes, and return an early result to the caller. The frontend polls the record's status to track progress. Wrap background chains other than `runTask()` in `mindstudio.waitUntil(...)` so the platform keeps the sandbox alive for them and records an interruption if they're cut short — `runTask()` registers itself automatically.
|
|
214
214
|
|
|
215
215
|
The example below shows the fire-and-forget shape, not a complete `runTask()` call. Load the `taskAgents` skill before writing one — configuring its tools, validating the output, and handling failures are all there, and none of them are visible here.
|
|
216
216
|
|
|
@@ -103,6 +103,30 @@ const { threads, nextCursor } = await chat.listThreads();
|
|
|
103
103
|
const full = await chat.getThread(thread.id);
|
|
104
104
|
await chat.updateThread(thread.id, 'New title');
|
|
105
105
|
await chat.deleteThread(thread.id);
|
|
106
|
+
|
|
107
|
+
// Progressive auth: threads started anonymously become unreachable after the
|
|
108
|
+
// user signs in (login replaces the session and its visitor identity). Claim
|
|
109
|
+
// them right after your verification/login succeeds so the conversation
|
|
110
|
+
// survives — the client remembers each thread's pre-login token automatically
|
|
111
|
+
// for threads touched this page session.
|
|
112
|
+
await chat.claimThread(thread.id);
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
**Client tools** — a tool whose effect happens in the browser (open a sheet, navigate, highlight)
|
|
116
|
+
is declared with `target: "client"` and a `name` + inline `inputSchema` instead of a `method`
|
|
117
|
+
(names must not collide with method ids; the schema is authored — there's no method contract to
|
|
118
|
+
derive it from). The agent's invocation arrives as the `client_tool_call` stream event / the
|
|
119
|
+
`onClientToolCall` callback on `sendMessage`; run the action there. Fire-and-forget on this
|
|
120
|
+
surface: the agent is told the action was displayed and keeps going — the user's next message
|
|
121
|
+
closes the loop.
|
|
122
|
+
|
|
123
|
+
```js
|
|
124
|
+
await chat.sendMessage(thread.id, text, {
|
|
125
|
+
onText: (delta) => append(delta),
|
|
126
|
+
onClientToolCall: (name, input) => {
|
|
127
|
+
if (name === 'showVerification') openVerifySheet(input);
|
|
128
|
+
},
|
|
129
|
+
});
|
|
106
130
|
```
|
|
107
131
|
|
|
108
132
|
**Sending messages (streaming):**
|
|
@@ -18,7 +18,7 @@ This is one of the most powerful pieces of the MindStudio SDK, and it can turn a
|
|
|
18
18
|
|
|
19
19
|
This is the tool to reach for whenever a feature would be dramatically more compelling if the app could autonomously research, enrich, or create on behalf of the user. Think about the difference between "user enters a restaurant name and it gets saved" vs. "user enters a restaurant name and gets back a fully researched, illustrated card." Task agents close that gap.
|
|
20
20
|
|
|
21
|
-
Run tasks in the background — depending on complexity they can take time to complete. Return an early partial result to the user and upsert later with the final result when the agent finishes.
|
|
21
|
+
Run tasks in the background — depending on complexity they can take time to complete. Return an early partial result to the user and upsert later with the final result when the agent finishes. The exception is cron and email triggers: there is no user waiting there, so `await` the task instead — awaiting is what surfaces its failures in the run's own result.
|
|
22
22
|
|
|
23
23
|
- **Research and enrichment:** "Given this email, find the person's LinkedIn, role, company, and a headshot" — the model searches, scrapes, extracts, and assembles structured data.
|
|
24
24
|
- **Content creation pipelines:** "Write SEO copy for this product in 3 languages, generate a hero image, extract keywords" — the model calls text generation, image generation, and analysis actions as needed.
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: Voice Interfaces
|
|
3
3
|
what: Realtime voice conversation as a first-class interface — the user talks to the app and its voice agent talks back in sub-second, interruptible speech, calling the app's methods mid-conversation as the authenticated user. The platform handles the media transport, turn-taking, barge-in, and transcripts, so the work is authorship — a persona written for the ear, a small toolset where every tool carries a latency class, and descriptions that say results out loud. Any app whose methods do something interesting can pick up a voice, and it is often the most impressive surface it has.
|
|
4
|
-
when: Before authoring `src/interfaces/voice.md`, choosing a voice model or pipeline, deciding which methods a voice agent gets,
|
|
4
|
+
when: Before authoring `src/interfaces/voice.md`, choosing a voice model or pipeline, deciding which methods a voice agent gets, building the voice UI with `createVoiceClient()`, or working out why a voice agent behaved the way it did on a call.
|
|
5
5
|
---
|
|
6
6
|
|
|
7
7
|
# Building Voice Interfaces
|
|
@@ -53,11 +53,11 @@ than any other surface.
|
|
|
53
53
|
Structure the compiled prompt as short **labeled sections** — Role & Objective, Personality & Tone,
|
|
54
54
|
Rules, and (when the app has a real call flow) Conversation Flow — with bullets over paragraphs;
|
|
55
55
|
realtime models find and follow sectioned rules far more reliably than prose. Scope rules
|
|
56
|
-
precisely
|
|
57
|
-
|
|
58
|
-
minimal: state the role, the boundaries, and the voice mechanics above, then add rules only for
|
|
56
|
+
precisely; blanket `always`/`never` makes the agent rigid and unable to handle reasonable exceptions.
|
|
57
|
+
And start minimal: state the role, the boundaries, and the voice mechanics above, then add rules only for
|
|
59
58
|
behaviors that actually misfire in test calls (the transcripts in the call log are the feedback
|
|
60
|
-
loop) rather than front-loading a
|
|
59
|
+
loop — `mindstudio-prod voice sessions get` reads a call verbatim) rather than front-loading a
|
|
60
|
+
policy manual.
|
|
61
61
|
|
|
62
62
|
### The latency classes
|
|
63
63
|
|
|
@@ -78,25 +78,65 @@ Classify by how the method actually behaves, not by what it is named. A "lookup"
|
|
|
78
78
|
external service is `slow`. When in doubt between `fast` and `slow`, pick `slow` — a needless
|
|
79
79
|
preamble is mildly chatty; an unexplained silence feels broken.
|
|
80
80
|
|
|
81
|
-
###
|
|
81
|
+
### Tool results reach the screen
|
|
82
82
|
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
83
|
+
Every successful tool call delivers its raw return value to the session's browser on the SDK's
|
|
84
|
+
`toolCall` event (`result` field, on `done`) — so the UI can render what the agent just did (the
|
|
85
|
+
citation it found, the record it pulled up, the booking it made) in lockstep with the spoken
|
|
86
|
+
answer. No flag, no polling, no key-threading, no model involvement: delivery to the invoking
|
|
87
|
+
session's own client is the same security context as the invocation itself (an RPC response), and
|
|
88
|
+
it's scoped to that one session.
|
|
88
89
|
|
|
89
|
-
|
|
90
|
-
the
|
|
91
|
-
|
|
92
|
-
`resultTruncated: true` with no data — keep
|
|
93
|
-
data itself. Failed calls
|
|
90
|
+
Consequence for authoring: **a tool's return value is user-visible by definition.** Return what
|
|
91
|
+
the user may see — no internal fields, keys, or diagnostics you wouldn't put on screen (the same
|
|
92
|
+
discipline as agent-interface tools, whose results render in chat). Payloads over ~32KB serialized
|
|
93
|
+
arrive as `resultTruncated: true` with no data — keep returns compact, or have the UI fetch big
|
|
94
|
+
data itself. Failed calls deliver nothing to the client (the model gets the `{ error }` and speaks
|
|
95
|
+
a decline).
|
|
94
96
|
|
|
95
97
|
For backend-side correlation (writing results to a table keyed by the call, custom channels), the
|
|
96
98
|
method itself can read `session.voiceSessionId` / `session.visitorId` from the agent SDK
|
|
97
99
|
(`import { session } from '@mindstudio-ai/agent'`) — the same id the browser holds as
|
|
98
100
|
`session.sessionId`, guaranteed by the platform rather than echoed by the model.
|
|
99
101
|
|
|
102
|
+
### Client tools: actions that happen on screen (`target: "client"`)
|
|
103
|
+
|
|
104
|
+
A tool whose effect belongs in the browser — open the verification sheet, navigate to a page,
|
|
105
|
+
highlight a record — is declared with `target: "client"` instead of a `method`:
|
|
106
|
+
|
|
107
|
+
```json
|
|
108
|
+
{
|
|
109
|
+
"target": "client",
|
|
110
|
+
"name": "showVerification",
|
|
111
|
+
"description": "tools/showVerification.md",
|
|
112
|
+
"inputSchema": { "type": "object", "properties": { "reason": { "type": "string" } } }
|
|
113
|
+
}
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
The platform never touches the backend for these: the agent's invocation is delivered to the
|
|
117
|
+
session's browser, the app's registered handler runs, and the handler's **return value goes back
|
|
118
|
+
to the agent as the tool result** — a real request/response, so the agent knows the sheet
|
|
119
|
+
actually opened (or that the user dismissed it) and speaks accordingly. Rules:
|
|
120
|
+
|
|
121
|
+
- `name` instead of `method`; must not collide with any backend method id. No latency class —
|
|
122
|
+
the agent holds the turn while the browser responds (up to ~30s, then a timeout error).
|
|
123
|
+
- `inputSchema` is authored inline (an object schema) — there's no method contract to derive
|
|
124
|
+
it from. Keep it small; these are UI directives, not data payloads.
|
|
125
|
+
- The frontend must register a handler, or invocations fail as `unhandled_client_tool`:
|
|
126
|
+
|
|
127
|
+
```js
|
|
128
|
+
session.registerClientTool('showVerification', async ({ reason }) => {
|
|
129
|
+
openVerifySheet(reason);
|
|
130
|
+
return { opened: true }; // what the agent hears back
|
|
131
|
+
});
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
- One client tool runs at a time per session; the description should tell the agent when to use
|
|
135
|
+
it and what to say while it's on screen. Throwing from the handler (or returning nothing)
|
|
136
|
+
becomes an error/ack the agent can speak around.
|
|
137
|
+
- The progressive-auth pattern above is the canonical use: make the verification sheet a client
|
|
138
|
+
tool and the agent opens it deliberately instead of the frontend inferring it from tool events.
|
|
139
|
+
|
|
100
140
|
### Tool descriptions say results out loud
|
|
101
141
|
|
|
102
142
|
Follow the agent-interface principles for tool descriptions (when to use and when not, parameter
|
|
@@ -124,24 +164,45 @@ caller.
|
|
|
124
164
|
|
|
125
165
|
### Choosing the model
|
|
126
166
|
|
|
127
|
-
Two shapes, one `model` field
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
167
|
+
Two shapes, one `model` field. **Use native speech-to-speech unless the user specifically asks
|
|
168
|
+
for a cascaded pipeline** — one realtime model hears and speaks: lowest latency, most natural
|
|
169
|
+
prosody, hears tone and hesitation.
|
|
170
|
+
|
|
171
|
+
**Native** (`{"model": ..., "voice": ...}`) — **default to `gpt-realtime-2.1` with voice
|
|
172
|
+
`marin`.**
|
|
173
|
+
|
|
174
|
+
- `gpt-realtime-2.1` — the default. Voices: `marin` (default), `cedar`, `alloy`, `ash`,
|
|
175
|
+
`ballad`, `coral`, `echo`, `sage`, `shimmer`, `verse`.
|
|
176
|
+
- `gpt-realtime-2.1-mini` — the same family, lighter; same voices.
|
|
177
|
+
- `gemini-2.5-flash-native-audio-preview-12-2025` — the Gemini pick, with a large expressive
|
|
178
|
+
roster: `Puck` (default, upbeat), `Zephyr` (bright), `Charon` (informative), `Kore` (firm),
|
|
179
|
+
`Fenrir` (excitable), `Leda` (youthful), `Orus` (firm), `Aoede` (breezy), `Callirrhoe`
|
|
180
|
+
(easy-going), `Autonoe` (bright), `Enceladus` (breathy), `Iapetus` (clear), `Umbriel`
|
|
181
|
+
(easy-going), `Algieba` (smooth), `Despina` (smooth), `Erinome` (clear), `Algenib` (gravelly),
|
|
182
|
+
`Rasalgethi` (informative), `Laomedeia` (upbeat), `Achernar` (soft), `Alnilam` (firm),
|
|
183
|
+
`Schedar` (even), `Gacrux` (mature), `Pulcherrima` (forward), `Achird` (friendly),
|
|
184
|
+
`Zubenelgenubi` (casual), `Vindemiatrix` (gentle), `Sadachbia` (lively), `Sadaltager`
|
|
185
|
+
(knowledgeable), `Sulafat` (warm).
|
|
186
|
+
- `gemini-3.1-flash-live-preview` — newer Gemini, same voices as 2.5, but currently can't speak
|
|
187
|
+
an opening greeting or take mid-call prompt updates (a plugin limitation expected to resolve
|
|
188
|
+
upstream) — prefer 2.5 until then.
|
|
189
|
+
- `grok-voice-think-fast-2.0` — a distinct personality register. Voices: `eve` (default),
|
|
190
|
+
`altair`, `ara`, `atlas`, `aurora`, `carina`, `castor`, `celeste`, `cosmo`, `helios`, `helix`,
|
|
191
|
+
`iris`, `kepler`, `leo`, `liora`, `lumen`, `luna`, `lux`, `naksh`, `orion`, `perseus`, `rex`,
|
|
192
|
+
`rigel`, `sal`, `sirius`, `ursa`, `zagan`, `zenith`.
|
|
193
|
+
|
|
194
|
+
**Cascaded** (`{"llm": ..., "stt": ..., "tts": ..., "voice": ...}`) — streaming transcription
|
|
195
|
+
into any chat model in the catalog, streaming speech out. Slightly higher latency; reach for it
|
|
196
|
+
only when the user wants it or the app's reasoning demands a specific chat model (e.g. the agent
|
|
197
|
+
interface already uses one and the voice should think identically). Slots: `stt` is
|
|
198
|
+
`deepgram-nova-3`; `tts` is `cartesia-sonic-3` (voices are per-account Cartesia UUIDs — see
|
|
199
|
+
play.cartesia.ai) or `elevenlabs-tts` (the account's ElevenLabs voice library); `llm` is any chat
|
|
200
|
+
model — ask `askMindStudioSdk` for chat model ids. One nuance: cascaded engines speak the
|
|
201
|
+
`greeting` verbatim (they have a real TTS); speech-to-speech engines have the model say it, so it
|
|
202
|
+
may paraphrase slightly.
|
|
203
|
+
|
|
204
|
+
The model and voice ids above are current and maintained with the platform — use them as written
|
|
205
|
+
(they are MindStudio ids, not vendor ids). The user's UI has a picker for changing the model
|
|
145
206
|
later, so validate only when you set it.
|
|
146
207
|
|
|
147
208
|
### Seeding from an existing agent
|
|
@@ -204,13 +265,15 @@ session.on('stateChange', (state) => { }); // on() returns an unsubscri
|
|
|
204
265
|
// far (never a delta) — render by upserting on segmentId, not appending.
|
|
205
266
|
session.on('transcript', ({ role, segmentId, text, final }) => { });
|
|
206
267
|
|
|
207
|
-
// status: 'running' | 'done' | 'failed'.
|
|
208
|
-
//
|
|
268
|
+
// status: 'running' | 'done' | 'failed'. Every 'done' carries the tool's raw
|
|
269
|
+
// return value in `result` (or `resultTruncated: true` if >~32KB serialized).
|
|
209
270
|
session.on('toolCall', ({ method, status, result }) => { });
|
|
210
271
|
session.on('error', (err) => { });
|
|
211
272
|
|
|
212
273
|
session.mute(); session.unmute(); session.isMuted;
|
|
213
274
|
session.sendText('123 Main Street'); // inject text into the live conversation
|
|
275
|
+
await session.refreshIdentity(); // after in-app verification — upgrade the
|
|
276
|
+
// live session anonymous → signed-in in place
|
|
214
277
|
session.end();
|
|
215
278
|
```
|
|
216
279
|
|
|
@@ -277,7 +340,7 @@ export async function callMeAboutMyOrder(input: { phone: string }) {
|
|
|
277
340
|
session, not the phone). Omitted/false → anonymous call; role-gated tools decline.
|
|
278
341
|
System/cron invocations have no human identity and always run anonymously.
|
|
279
342
|
- **Production needs a dedicated phone number.** The app owner attaches one ($1/month) via the
|
|
280
|
-
dashboard or `mindstudio-prod voice numbers` (see "
|
|
343
|
+
dashboard or `mindstudio-prod voice numbers` (see "The voice CLI"
|
|
281
344
|
below) — it becomes the caller ID for every call, in dev sessions too, so users always see
|
|
282
345
|
the same number. Without one, deployed calls throw `phone_out_requires_dedicated_number`, and
|
|
283
346
|
dev sessions fall back to a shared platform test number that varies per call (tighter limits
|
|
@@ -338,7 +401,7 @@ wrong one wherever the agent's tools can move money, reveal sensitive records, o
|
|
|
338
401
|
destructive actions. It lives in the interface config deliberately: enabling it is a code
|
|
339
402
|
change, visible in review and auditable via deploys, not a dashboard toggle.
|
|
340
403
|
|
|
341
|
-
##
|
|
404
|
+
## The voice CLI
|
|
342
405
|
|
|
343
406
|
The `mindstudio-prod voice` family covers numbers, the call log, and voice policy:
|
|
344
407
|
|
|
@@ -377,7 +440,7 @@ section.
|
|
|
377
440
|
name: Front Desk
|
|
378
441
|
description: Books appointments and answers questions by voice.
|
|
379
442
|
type: interface/voice
|
|
380
|
-
model: {"model": "gpt-realtime-
|
|
443
|
+
model: {"model": "gpt-realtime-2.1", "voice": "marin"}
|
|
381
444
|
turnDetection: {"eagerness": "medium"}
|
|
382
445
|
greeting: Hey! I can help you book, reschedule, or answer questions — what do you need?
|
|
383
446
|
---
|
|
@@ -435,7 +498,7 @@ The top-level key must match the interface type (`voice`):
|
|
|
435
498
|
"voice": {
|
|
436
499
|
"name": "Front Desk",
|
|
437
500
|
"description": "Books appointments and answers questions by voice.",
|
|
438
|
-
"model": "gpt-realtime-
|
|
501
|
+
"model": "gpt-realtime-2.1",
|
|
439
502
|
"voice": "marin",
|
|
440
503
|
"turnDetection": { "eagerness": "medium" },
|
|
441
504
|
"greeting": "Hey! I can help you book, reschedule, or answer questions — what do you need?",
|
|
@@ -537,3 +600,30 @@ frontend or the agent interface. Anonymous sessions (when allowed) have no user
|
|
|
537
600
|
gated methods reject, and the caller's history is scoped to their browser's visitor identity.
|
|
538
601
|
That's why role restrictions belong in the tool descriptions — the agent should decline in
|
|
539
602
|
character, not relay a rejection.
|
|
603
|
+
|
|
604
|
+
### Progressive auth: verify mid-call without dropping the conversation
|
|
605
|
+
|
|
606
|
+
The best pattern for apps that allow anonymous sessions (`requireUser: false`): let visitors
|
|
607
|
+
explore by voice, and verify only when they hit an account-bound action — without killing the
|
|
608
|
+
live call. Four pieces, all platform rails:
|
|
609
|
+
|
|
610
|
+
1. **Account-gated tools return a standard not-verified shape** instead of doing the work:
|
|
611
|
+
`{ verified: false, message: 'The caller is not verified. Offer to verify them before sharing
|
|
612
|
+
account details.' }`. The agent speaks the offer in character (reinforce tone in the system
|
|
613
|
+
prompt's verification section). Check with the agent SDK's `auth.userId` inside the method.
|
|
614
|
+
2. **The frontend opens its verification sheet off the same signal.** It already receives every
|
|
615
|
+
tool's `toolCall` event (and the tool's return in `result`) — when an account tool fires (or
|
|
616
|
+
returns `verified: false`) while the app has no signed-in user, open the sheet.
|
|
617
|
+
3. **The sheet runs the platform's auth rails** — `auth.sendSmsCode()` / `auth.verifySmsCode()`
|
|
618
|
+
(or the email pair) from `@mindstudio-ai/interface`. On success the app's session becomes the
|
|
619
|
+
verified user.
|
|
620
|
+
4. **Hand the verified session back to the live call**: `await session.refreshIdentity()`. The
|
|
621
|
+
platform upgrades the running voice session in place — subsequent tool calls carry the user's
|
|
622
|
+
identity and roles, and the agent's Current User context refreshes — no teardown, no lost
|
|
623
|
+
conversation. (Phone calls don't need this: they verify through the agent's built-in flow.)
|
|
624
|
+
|
|
625
|
+
`refreshIdentity()` is upgrade-only (anonymous → signed-in; an already-identified session rejects
|
|
626
|
+
with `already_identified`) and requires the session to have been started by this same browser. If
|
|
627
|
+
it fails, ending and restarting the session is the graceful fallback. The chat sibling for agent
|
|
628
|
+
interfaces is `claimThread(threadId)` — anonymous threads become unreachable after login until
|
|
629
|
+
claimed.
|
|
@@ -6,7 +6,7 @@
|
|
|
6
6
|
|
|
7
7
|
## Principles
|
|
8
8
|
- The spec in `src/` is the source of truth for what the app does: consult it before making behavioral code changes, and keep it in sync as the app evolves by dispatching periodic requests to the specSync tool after meaningful changes. Some amount of drift is fine and inevitable - your priority is working efficiently with the user, let the specSync tool handle things in the background.
|
|
9
|
-
-
|
|
9
|
+
- The Build Overview (`src/overview.html` — the project's home page) stays current via specSync: after a deploy or a large milestone, set `refreshBuildOverview: true` on your specSync dispatch and it re-authors the overview from the updated spec. Call `writeBuildOverview` directly only for the initial end-of-build generation or when the user explicitly asks.
|
|
10
10
|
- Change only what the task requires. Match existing styles. Keep solutions simple.
|
|
11
11
|
- Read files before editing them. Understand the context before making changes.
|
|
12
12
|
- When the user asks you to make a change, execute it fully — all steps, no pausing for confirmation. Use `confirmDestructiveAction` to gate before destructive or irreversible actions (e.g., deleting data, resetting the database). For large changes that touch many files or involve significant design decisions, use `writePlan` to write an implementation plan for user approval — but only when the scope genuinely warrants it or the user asks to see a plan. The plan is saved to `.remy-plan.md` and the user can review, discuss, and refine it before approving. Do not begin implementation until the plan is approved. Most work should be done autonomously without a plan.
|
|
@@ -44,7 +44,9 @@ Your editor — a design expert for words. Hand it any user-facing copy — an e
|
|
|
44
44
|
|
|
45
45
|
Your spec keeper. Once the app is built and you're iterating on it, whenever you make code changes that alter what the app does, hand it a brief, plain-language description of what you changed and why (prefer bullet points, batch multiple changes into one invocation). It finds the affected sections of the spec in `src/` and updates them to match what has been built. It reads the spec itself and decides what to touch, so you don't need to name files or locations.
|
|
46
46
|
|
|
47
|
-
It always runs in the background: it returns immediately and you keep working while it reconciles
|
|
47
|
+
It always runs in the background: it returns immediately and you keep working while it reconciles; it completes silently and the outcome appears as an automated note at the start of a later turn — it never wakes you. So don't hunt through spec files to sync them yourself and don't wait on it — hand off and move on. You decide when the spec has drifted enough to be worth a hand-off (after a meaningful change, or a batch of them - you do not need to invoke this after every change or conversation turn - some amount of drift between code and spec is completely normal and acceptable).
|
|
48
|
+
|
|
49
|
+
It can also keep the Build Overview current: set `refreshBuildOverview: true` and, after reconciling, it re-authors the overview copy from the updated spec and re-renders `src/overview.html`. Set it after a deploy or a large milestone; leave it off for routine syncs.
|
|
48
50
|
|
|
49
51
|
### QA (`runAutomatedBrowserTest`)
|
|
50
52
|
|
|
@@ -13,9 +13,11 @@ When the content you need to test is behind authentication, use the `setupBrowse
|
|
|
13
13
|
|
|
14
14
|
If you need to test the login/signup flow itself (e.g., verifying the UI, error states, or the verification code input), navigate it manually: use `remy@mindstudio.ai` for email and `+15551234567` for phone. In the dev environment, verification codes are bypassed for this email and any 555-prefixed phone number — enter any 6-digit code (e.g., `123456`).
|
|
15
15
|
|
|
16
|
+
To test as a **signed-out visitor** (public pages, landing/join links), call `setupBrowser` with NO `auth` — it clears the auth cookie and reloads at the given path, giving you a clean unauthenticated session. Combine with `navigate` + `fresh: true` when you need a fresh-document view of an entry page mid-run.
|
|
17
|
+
|
|
16
18
|
## Browser Commands
|
|
17
19
|
|
|
18
|
-
Your session always starts on the app root / in a logged out/unauthenticated state. Use `setupBrowser` to authenticate before testing protected pages.
|
|
20
|
+
Your session always starts on the app root / in a logged out/unauthenticated state, on a freshly reloaded page running the current code — any changes made since the last run are already picked up. Never restart the dev server (or reload manually) to clear a "stale bundle"; that staleness cannot survive the start-of-run refresh. Use `setupBrowser` to authenticate before testing protected pages.
|
|
19
21
|
|
|
20
22
|
### Snapshot format
|
|
21
23
|
|
|
@@ -40,13 +42,43 @@ Note: the snapshot concatenates inline text and strips whitespace. If you need t
|
|
|
40
42
|
- `type`: Type text into an input. Characters appear one at a time. Set `clear: true` to clear the field first.
|
|
41
43
|
- `select`: Select a dropdown option by text. Target the `<select>` element, set `option` to the option text.
|
|
42
44
|
- `wait`: Wait for an element to appear (polls every 100ms, default 5s timeout). Also waits for network to settle after the element is found.
|
|
43
|
-
- `navigate`: Navigate to a new URL within the app. Waits for the
|
|
45
|
+
- `navigate`: Navigate to a new URL within the app. Waits for the route to load before continuing with subsequent steps. Use this instead of evaluate with `window.location.href` when you need to navigate and then continue interacting with the new page. Steps after navigate execute on the new page automatically. Same-origin navigation is a soft in-app route change (like clicking a link in an SPA — in-memory app state survives); set `fresh: true` to force a real full page load with a fresh document instead. Use `fresh: true` when the test is about what a user sees on *entry* — landing pages, join/invite links, "what does a signed-out visitor see" — where reusing the SPA's in-memory state would test the wrong thing. The result reports the URL the page actually landed on, so if the app redirected you (e.g. an auth wall bounced you off a public page), you'll see the real destination — check it instead of assuming the navigation stuck.
|
|
44
46
|
- `evaluate`: Run arbitrary JavaScript in the page and return the result.
|
|
45
47
|
- `styles`: Read computed CSS styles from page elements. Pass a `properties` array with camelCase CSS property names (e.g., `["backgroundColor", "borderRadius", "fontSize"]`). Omit `properties` for a default set covering colors, typography, spacing, borders, shadows, dimensions, and layout. Uses the same targeting as click/type (ref, text, role, label, selector). Omit the target to get styles for all elements from the last snapshot.
|
|
46
48
|
- `screenshotFullPage`: Take a screenshot of the whole page, top to bottom. Returns CDN url with full text analysis and dimensions. Use for overall composition or content past the fold.
|
|
47
49
|
- `screenshotViewport`: Take a screenshot of the visible viewport. Returns CDN url with full text analysis and dimensions. To capture a specific section, set `scrollToSelector` (a CSS selector) — or `scrollY` (an absolute offset) — on this same step; it scrolls the target into view and captures it atomically, so you do NOT need a separate scroll step. Do not use if you can get what you need with other tools - only use when you need to visually see the viewport.
|
|
48
50
|
- `setViewport`: Switch the browser between desktop and mobile rendering. Set `mode` to `"desktop"` or `"mobile"`. Mobile emulates a phone (390-wide, touch, device pixel ratio 2); desktop is the standard wide viewport. This reloads the page so media queries, responsive layouts, and `matchMedia` re-evaluate — the reload clears in-page state, so switch before you set up the state you want to inspect. The mode persists across navigations within a run. Each run starts in the app's default mode, so only use this when you need to check the other one.
|
|
49
51
|
|
|
52
|
+
### Voice interfaces
|
|
53
|
+
|
|
54
|
+
Apps with a voice interface are testable end to end — the UI layer included. The sandbox browser
|
|
55
|
+
auto-grants a (silent) microphone, and while a session is live the SDK publishes a handle at
|
|
56
|
+
`window.__MS_VOICE__` so you can converse by text: the agent treats injected text exactly like
|
|
57
|
+
user speech (interrupts and replies), backend tools run for real, and client tools render their
|
|
58
|
+
real UI (cards, sheets) in the page.
|
|
59
|
+
|
|
60
|
+
The loop:
|
|
61
|
+
|
|
62
|
+
1. Start a session through the app's real UI — `click` its voice affordance (orb/button). No mic
|
|
63
|
+
prompt appears. Then `wait` briefly and confirm the session is live:
|
|
64
|
+
`evaluate: window.__MS_VOICE__?.state` (undefined means no session started — report that,
|
|
65
|
+
don't improvise).
|
|
66
|
+
2. Speak by injection: `evaluate: window.__MS_VOICE__.sendText("I'd like to book Tuesday at 2")`.
|
|
67
|
+
3. Give the agent a few seconds to respond (replies are generated speech — slower than chat).
|
|
68
|
+
`wait` for the UI you expect (client-tool cards appear via the app's real handlers), and read
|
|
69
|
+
the conversation: `evaluate: window.__MS_VOICE__.transcript` (one entry per utterance, both
|
|
70
|
+
sides, `final` marks settled ones) and `window.__MS_VOICE__.toolCalls` (which tools ran;
|
|
71
|
+
`done` entries carry the tool's return value).
|
|
72
|
+
4. Verify visuals with `screenshotViewport` like any other flow.
|
|
73
|
+
5. Read `transcript`/`toolCalls` BEFORE ending — then `evaluate: window.__MS_VOICE__.end()` (the
|
|
74
|
+
handle is removed when the session ends).
|
|
75
|
+
|
|
76
|
+
Voice sessions are the most expensive thing you can run — real voice-model minutes are metered,
|
|
77
|
+
and the agent speaks its replies out loud even when you type at it. Keep voice tests short and
|
|
78
|
+
purposeful: a handful of turns that exercise the target behavior, then end the session. What you
|
|
79
|
+
cannot test is the audio layer itself (mishearing, interruptions, pronunciation) — never attempt
|
|
80
|
+
to simulate audio; report that scope limit instead.
|
|
81
|
+
|
|
50
82
|
### Element targeting (tried in order)
|
|
51
83
|
|
|
52
84
|
1. `ref`: From the last snapshot. Most reliable.
|