@mindstudio-ai/remy 0.1.276 → 0.1.277

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -6,18 +6,11 @@ when: Before adding a `webhook` interface or writing a method that receives a pr
6
6
 
7
7
  # Webhook Interfaces
8
8
 
9
- Inbound HTTP endpoints that invoke a method directly and synchronously — the caller waits for the
10
- method to finish. Use for receiving webhooks from external services (Stripe, GitHub, Shopify, Slack,
11
- Twilio). Direct inbound webhooks with signature verification work natively; do **not** build
12
- confirmation-token or polling workarounds.
9
+ Inbound HTTP endpoints that invoke a method directly and synchronously — the caller waits for the method to finish. Use for receiving webhooks from external services (Stripe, GitHub, Shopify, Slack, Twilio). Direct inbound webhooks with signature verification work natively; do **not** build confirmation-token or polling workarounds.
13
10
 
14
- The other native path for inbound HTTP is the API interface, which uses bearer auth and exposes the raw
15
- body at `input._request.rawBody` instead of at the top level. Use that one when the caller can send an
16
- `Authorization` header and you want a documented REST surface; use this one for provider callbacks.
17
- Load the `restApi` skill if that's the direction.
11
+ The other native path for inbound HTTP is the API interface, which uses bearer auth and exposes the raw body at `input._request.rawBody` instead of at the top level. Use that one when the caller can send an `Authorization` header and you want a documented REST surface; use this one for provider callbacks. Load the `restApi` skill if that's the direction.
18
12
 
19
- Webhook secrets are configured at the project level by the user through the Remy platform. Your job is
20
- the `interface.json` and the handling method.
13
+ Webhook secrets are configured at the project level by the user through the Remy platform. Your job is the `interface.json` and the handling method.
21
14
 
22
15
  ## Config (`interface.json`)
23
16
 
@@ -38,10 +31,7 @@ The top-level key must match the interface type (`webhook`):
38
31
  ```
39
32
 
40
33
  - `method` — the id of a method in `methods[]` to invoke.
41
- - `secret` — a developer-chosen opaque token that is **both the routing key and the access guard**. It
42
- is stable across deploys (compilation is a passthrough — redeploying never rotates it), so a URL you
43
- register with Stripe/GitHub stays valid. Generate one long random value per endpoint and keep it
44
- constant.
34
+ - `secret` — a developer-chosen opaque token that is **both the routing key and the access guard**. It is stable across deploys (compilation is a passthrough — redeploying never rotates it), so a URL you register with Stripe/GitHub stays valid. Generate one long random value per endpoint and keep it constant.
45
35
  - Declare multiple endpoints if needed; each `secret` maps to one method.
46
36
 
47
37
  Declare it in `mindstudio.json`:
@@ -58,9 +48,7 @@ Register this with the external service:
58
48
  https://{app-host}/_/webhook/{secret}
59
49
  ```
60
50
 
61
- `{app-host}` is any host the app is served on: its `custom_subdomain` host (e.g.
62
- `myapp.madewithremy.com`), a custom domain if configured, or the UUID host
63
- (`<appId>.madewithremy.com` / `.msagent.ai`). All HTTP verbs are accepted.
51
+ `{app-host}` is any host the app is served on: its `custom_subdomain` host (e.g. `myapp.madewithremy.com`), a custom domain if configured, or the UUID host (`<appId>.madewithremy.com` / `.msagent.ai`). All HTTP verbs are accepted.
64
52
 
65
53
  ## Input
66
54
 
@@ -76,8 +64,7 @@ The method receives:
76
64
  }
77
65
  ```
78
66
 
79
- For signature verification **always use `rawBody`, never `body`** — providers (Stripe, GitHub, Shopify,
80
- Slack) HMAC the raw payload, and a re-serialized `body` will not match:
67
+ For signature verification **always use `rawBody`, never `body`** — providers (Stripe, GitHub, Shopify, Slack) HMAC the raw payload, and a re-serialized `body` will not match:
81
68
 
82
69
  ```typescript
83
70
  const event = stripe.webhooks.constructEvent(
@@ -87,22 +74,14 @@ const event = stripe.webhooks.constructEvent(
87
74
  );
88
75
  ```
89
76
 
90
- `rawBody` is populated for `application/json` and `application/x-www-form-urlencoded` bodies (what these
91
- providers send).
77
+ `rawBody` is populated for `application/json` and `application/x-www-form-urlencoded` bodies (what these providers send).
92
78
 
93
79
  ## Response
94
80
 
95
- Whatever the method returns as output is sent back to the caller as JSON; if it returns no output, the
96
- platform responds `204`. A wrong/unknown secret returns `401`; an app with no live release returns
97
- `404`.
81
+ Whatever the method returns as output is sent back to the caller as JSON; if it returns no output, the platform responds `204`. A wrong/unknown secret returns `401`; an app with no live release returns `404`.
98
82
 
99
83
  ## Auth
100
84
 
101
- Methods invoked through this interface run with `auth.roles: ['system']` — the platform is calling, not
102
- a user session, so there's no user to impersonate. Use `auth.requireRole('system')` to gate methods that
103
- should only be reachable via a platform trigger. The auth reference in your system prompt covers the
104
- system role in full.
85
+ Methods invoked through this interface run with `auth.roles: ['system']` — the platform is calling, not a user session, so there's no user to impersonate. Use `auth.requireRole('system')` to gate methods that should only be reachable via a platform trigger. The auth reference in your system prompt covers the system role in full.
105
86
 
106
- Note that the URL secret is the only access control on the endpoint itself, which is why it needs to be
107
- long and random. Signature verification is a second, independent check that the payload really came from
108
- the provider — do both.
87
+ Note that the URL secret is the only access control on the endpoint itself, which is why it needs to be long and random. Signature verification is a second, independent check that the payload really came from the provider — do both.
@@ -2,8 +2,7 @@
2
2
 
3
3
  The spec is the application. It defines what the app does — the data, the workflows, the roles, the edge cases — and how it looks and feels. Code is derived from it. Your job is to help the user build a spec that's complete enough to compile into a working app.
4
4
 
5
- **Writing the first draft:**
6
- After intake, write the spec immediately. Do not ask "ready for me to start?" or wait for confirmation — just start writing. The first draft should cover the full shape of the app — it's better to have every section roughed in than to have one section perfect and the rest missing.
5
+ **Writing the first draft:** After intake, write the spec immediately. Do not ask "ready for me to start?" or wait for confirmation — just start writing. The first draft should cover the full shape of the app — it's better to have every section roughed in than to have one section perfect and the rest missing.
7
6
 
8
7
  - Make concrete decisions rather than leaving things vague. The user can change a decision; they can't react to vagueness.
9
8
  - Flag assumptions you made during intake so the user can confirm or correct them.
@@ -51,33 +51,17 @@ Note: the snapshot concatenates inline text and strips whitespace. If you need t
51
51
 
52
52
  ### Voice interfaces
53
53
 
54
- Apps with a voice interface are testable end to end — the UI layer included. The sandbox browser
55
- auto-grants a (silent) microphone, and while a session is live the SDK publishes a handle at
56
- `window.__MS_VOICE__` so you can converse by text: the agent treats injected text exactly like
57
- user speech (interrupts and replies), backend tools run for real, and client tools render their
58
- real UI (cards, sheets) in the page.
54
+ Apps with a voice interface are testable end to end — the UI layer included. The sandbox browser auto-grants a (silent) microphone, and while a session is live the SDK publishes a handle at `window.__MS_VOICE__` so you can converse by text: the agent treats injected text exactly like user speech (interrupts and replies), backend tools run for real, and client tools render their real UI (cards, sheets) in the page.
59
55
 
60
56
  The loop:
61
57
 
62
- 1. Start a session through the app's real UI — `click` its voice affordance (orb/button). No mic
63
- prompt appears. Then `wait` briefly and confirm the session is live:
64
- `evaluate: window.__MS_VOICE__?.state` (undefined means no session started — report that,
65
- don't improvise).
58
+ 1. Start a session through the app's real UI — `click` its voice affordance (orb/button). No mic prompt appears. Then `wait` briefly and confirm the session is live: `evaluate: window.__MS_VOICE__?.state` (undefined means no session started — report that, don't improvise).
66
59
  2. Speak by injection: `evaluate: window.__MS_VOICE__.sendText("I'd like to book Tuesday at 2")`.
67
- 3. Give the agent a few seconds to respond (replies are generated speech — slower than chat).
68
- `wait` for the UI you expect (client-tool cards appear via the app's real handlers), and read
69
- the conversation: `evaluate: window.__MS_VOICE__.transcript` (one entry per utterance, both
70
- sides, `final` marks settled ones) and `window.__MS_VOICE__.toolCalls` (which tools ran;
71
- `done` entries carry the tool's return value).
60
+ 3. Give the agent a few seconds to respond (replies are generated speech — slower than chat). `wait` for the UI you expect (client-tool cards appear via the app's real handlers), and read the conversation: `evaluate: window.__MS_VOICE__.transcript` (one entry per utterance, both sides, `final` marks settled ones) and `window.__MS_VOICE__.toolCalls` (which tools ran; `done` entries carry the tool's return value).
72
61
  4. Verify visuals with `screenshotViewport` like any other flow.
73
- 5. Read `transcript`/`toolCalls` BEFORE ending — then `evaluate: window.__MS_VOICE__.end()` (the
74
- handle is removed when the session ends).
75
-
76
- Voice sessions are the most expensive thing you can run — real voice-model minutes are metered,
77
- and the agent speaks its replies out loud even when you type at it. Keep voice tests short and
78
- purposeful: a handful of turns that exercise the target behavior, then end the session. What you
79
- cannot test is the audio layer itself (mishearing, interruptions, pronunciation) — never attempt
80
- to simulate audio; report that scope limit instead.
62
+ 5. Read `transcript`/`toolCalls` BEFORE ending — then `evaluate: window.__MS_VOICE__.end()` (the handle is removed when the session ends).
63
+
64
+ Voice sessions are the most expensive thing you can run — real voice-model minutes are metered, and the agent speaks its replies out loud even when you type at it. Keep voice tests short and purposeful: a handful of turns that exercise the target behavior, then end the session. What you cannot test is the audio layer itself (mishearing, interruptions, pronunciation) — never attempt to simulate audio; report that scope limit instead.
81
65
 
82
66
  ### Element targeting (tried in order)
83
67
 
@@ -21,8 +21,7 @@ These are things we already know about and have decided to accept:
21
21
  - **`dist/` is where code lives.** Remy apps use `dist/` for all code (methods, interfaces, tables) and `src/` for natural language specs. This is NOT the conventional "dist is build output" pattern. Never flag code being in `dist/` as wrong.
22
22
  - The raw request body for webhook signature verification (Stripe, GitHub, etc.) is available natively — under `input._request.rawBody` on API interface methods, and at top-level `input.rawBody` on Webhook interface methods. Do NOT suggest external proxies or workarounds.
23
23
 
24
- - Ignore limited browser support for `oklch` gradients using `in <colorspace>` syntax — we accept the compatibility tradeoff for better color quality
25
- -Ignore limited browser support for CSS scroll-driven animations (`animation-timeline: scroll()` / `view()`) - we accept this tradeoff
24
+ - Ignore limited browser support for `oklch` gradients using `in <colorspace>` syntax — we accept the compatibility tradeoff for better color quality -Ignore limited browser support for CSS scroll-driven animations (`animation-timeline: scroll()` / `view()`) - we accept this tradeoff
26
25
  - Trust your knowledge about Platform SDKs (these are the core of every app) - for our purposes, assume they're always current and stable:
27
26
  - `@mindstudio-ai/interface` — frontend SDK. `createClient<T>()` gives typed RPC to backend methods (no raw fetch). `auth` handles auth state (`auth.currentUser`, `auth.onAuthStateChanged(cb)`, verification flows, logout). `platform.uploadFile()` handles signed S3 uploads and returns permanent CDN URLs with query-string resizing for images and auto-thumbnails for videos.
28
27
  - `@mindstudio-ai/agent` — backend SDK. `db.defineTable<T>()` gives a typed ORM with Query (chainable reads) and direct writes. `auth` gives `auth.userId`, `auth.roles`, `auth.requireRole()`, `auth.hasRole()`. Also provides 200+ managed actions for AI models, email/SMS, third-party APIs, media processing.
@@ -32,9 +32,6 @@
32
32
  <animation>
33
33
  {{prompts/animation.md}}
34
34
  </animation>
35
- <images>
36
- {{prompts/images.md}}
37
- </images>
38
35
  <ui_patterns>
39
36
  {{prompts/ui-patterns.md}}
40
37
  </ui_patterns>
@@ -14,27 +14,14 @@ The design should look like it could be an Apple iOS/macOS app of the year winne
14
14
 
15
15
  When specifying sheets, drawers, modals, or any surface that slides/fades into view, always include the interaction and motion details. The developer will build the minimal static version if you don't. Be explicit about: how it enters (direction, easing, duration), how it's dismissed (drag-to-dismiss threshold, swipe velocity, tap-outside), how the backdrop behaves (opacity, blur, tap to close), and any spring/bounce physics. These details are the difference between "functional" and "feels like a real app."
16
16
 
17
- ### Notes for Designing Auth Flows
17
+ ### Surfaces With Dedicated Craft References
18
18
 
19
- Login and signup screens set the tone for the user's entire experience with the app and are important to get right - they should feel like exciting entry points into the next level of the user journy. A janky login form with misaligned inputs and no feedback dminishes excitement and undermines trust before the user even gets in.
19
+ Some surfaces are deep enough to carry their own craft reference in <available_skills> — load the matching skill with loadSkill *before* designing or reviewing one of these, not after. What stays resident here is only the invariant that applies even when the skill isn't loaded:
20
20
 
21
- Authentication moments must feel natural and intuitive - they should not feel jarring or surprising. Take care to integrate them into the entire experience when building. Remy apps support SMS code verification, email verification, delegated "Sign in with Remy", or a combination, depending on how the app is configured.
22
-
23
- **"Continue with {Org}" (delegated sign-in):** Some apps are internal business tools that are owned by an organization and sign members in through the platform — a single "Continue with {Org}" button, no verification-code step at all. For these apps this button is often the primary (sometimes only) path, so give it real presence in the branded login moment rather than tucking it away, and label it with the organization's actual name. If the app also offers code methods, lead with the delegated button and place the code form beneath. This scheme should ONLY be used for internal apps AND when it is explicitly enabled in <org_context> - it should not be used for public-facing apps.
24
-
25
- **Verification code input:** The 6-digit code entry is the critical moment. Prefer to design it as individual digit boxes (not a single text input), with auto-advance between digits, auto-submit on paste, and clear visual feedback. The boxes should be large enough to tap easily on mobile. Show a subtle animation on successful verification. Error states should be inline and immediate, not a separate alert.
26
-
27
- **The send/resend flow:** After the user enters their email or phone and taps "Send code," show clear confirmation that the code was sent ("Check your email" with the address displayed). Include a resend option with a cooldown timer (e.g., "Resend in 30s"). The transition from "enter email" to "enter code" should feel smooth, not like a page reload.
28
-
29
- **The overall login page:** This is a branding moment. Use the app's full visual identity — colors, typography, any hero imagery or illustration. A centered card on a branded background is a classic pattern. Don't make it look like a generic SaaS login template. The login page should feel like it belongs to this specific app.
30
-
31
- **Post-login transition:** After successful verification, the transition into the app should feel seamless. Avoid a blank loading screen — if data needs to load, show the app shell with skeleton states.
32
-
33
- ### Notes for Designing AI Chat Interfaces
34
-
35
- If the app includes an AI chat interface, take care to make it beautiful and intentional. A good chat interface feels like magic, a bad one feels like a broken customer service bot that will leave the user frustrated and annoyed.
36
-
37
- Pay close attention to text streaming when the AI replies - it should feel natural, smooth, and beautiful. There must never be any abrupt layout shift for tool use or new messages, and scrolling should feel natural - like you are in a well-designed iOS chat app. Make sure to specify styles, layouts, animations, and remind the developer of things to watch out for. Reference chat apps you know are well-designed, this is not the place to re-invent the wheel. Users have expectations about how chat works and we should meet them and surpass them.
21
+ - **Data visualization** (`dataViz`) — charts, dashboards, metric tiles, and tables of numbers are hand-designed, never assembled from a chart library's defaults, and every column of figures gets `tabular-nums`.
22
+ - **Voice experiences** (`voiceExperience`) — the on-screen experience of a voice agent (centerpiece, captions, tool activity, call controls) is a first-class design surface, and the centerpiece is a real-time computed piece, never a defaulted CSS gradient blob.
23
+ - **Auth flows** (`authExperience`) — auth is the app's front door, usually the first designed thing a user sees. Remy apps are passwordless — verification codes, plus (rarely, only when <org_context> explicitly enables it) delegated "Continue with {Org}" — never design a password field.
24
+ - **AI chat** (`chatExperience`) — users arrive fluent in chat with expectations set by the best chat products; this is not the place to re-invent the wheel — meet those expectations, then surpass them.
38
25
 
39
26
  ### Wireframes
40
27
 
@@ -46,8 +33,7 @@ Wireframes isolate one small piece: a single card, a button animation, a transit
46
33
 
47
34
  Wireframes render in a small transparent iframe. Set a background color and shadow on the component's container (not the body) so it's visible against the transparent background. Center it in the viewport. No annotations or labels inside the wireframe. Put notes in the surrounding markdown. For interactive wireframes with states or animations, include a play/reset control. No images.
48
35
 
49
- Wireframes are vanilla HTML/CSS/JS (no React). For animations beyond CSS, use GSAP via CDN:
50
- `<script src="https://cdn.jsdelivr.net/npm/gsap@3/dist/gsap.min.js"></script>`
36
+ Wireframes are vanilla HTML/CSS/JS (no React). For animations beyond CSS, use GSAP via CDN: `<script src="https://cdn.jsdelivr.net/npm/gsap@3/dist/gsap.min.js"></script>`
51
37
 
52
38
  Quick skeleton wireframe (grey boxes, just showing layout and hierarchy):
53
39
 
@@ -0,0 +1,60 @@
1
+ ---
2
+ name: Auth Experience
3
+ what: The app's front door — login and signup screens, verification-code entry, delegated org sign-in, mid-experience sign-in gates, and the post-login transition. Auth is the one surface every user passes through and usually the first designed thing they see, and it is where trust is won or lost before the product gets a chance - a first-class design deliverable and a branding moment, never a generic SaaS template. This reference carries the full craft recipe — the branded login moment, code-entry mechanics, send/resend flow, progressive-auth gates, transitions, and the failure states that make or break it.
4
+ when: Before designing (or reviewing) any auth surface — a login or signup screen, verification-code entry, a "Continue with {Org}" moment, an in-app sign-in gate or verification sheet, or the post-login transition.
5
+ ---
6
+
7
+ # Auth Experience
8
+
9
+ Auth is the front door: the one surface every user passes through, usually before they've seen anything else you designed. A janky login form with misaligned inputs and no feedback undermines trust before the user is even inside — and a beautiful one sets the expectation that everything past it is built to the same standard. Users arrive fluent in login flows, so like chat this is craft within convention: a familiar shape, executed like the entrance to something worth entering.
10
+
11
+ One platform fact shapes everything here: **Remy apps are passwordless.** Sign-in is a verification code (SMS or email), a delegated "Continue with {Org}" button, or a combination — there is no password anywhere in the system. Never design a password field, a "forgot password" link, or a password-strength meter; their presence marks the design as a template pasted onto the wrong product.
12
+
13
+ ## The branded login moment
14
+
15
+ The login screen is a branding moment — the app's full visual identity working: its palette, its type, its imagery or illustration if the brand has one. A centered card on a branded field is a classic, reliable composition; what's forbidden is the generic SaaS template look that could belong to any product. The screen should feel like the front door of *this specific app*, and it should feel like an exciting entry point to the next level of the user's journey, not a checkpoint.
16
+
17
+ Method hierarchy is part of the composition:
18
+
19
+ - **"Continue with {Org}" (delegated sign-in)** — a rare configuration used by a small minority of apps: internal business tools whose members sign in through the platform with a single button, no code step. **The default assumption is that this method does not exist.** It is available ONLY when <org_context> explicitly enables it or the user has explicitly asked for it — when neither is true, ignore it entirely: do not design the button, do not offer it as an option, do not leave space for it in the composition. Never design it into a public-facing app. On the rare app where it IS enabled, it is often the *only* path — give it real presence in the branded moment (a primary, confident button labeled with the organization's actual name), and when code methods coexist, the delegated button leads with the code form beneath.
20
+ - **Code methods (email / SMS)** — one clear input, one clear action. If both methods exist, pick a primary and make switching quiet (a text link, not dueling forms).
21
+
22
+ ## Code entry
23
+
24
+ The 6-digit code is the critical moment of the flow — design it precisely:
25
+
26
+ - **Individual digit boxes**, not a single text input: auto-advance between digits, full-code paste handled (auto-submit on a complete paste), backspace moving backward naturally.
27
+ - **Sized for thumbs.** The boxes are large enough to tap easily on mobile, and the numeric keyboard is requested (`inputmode="numeric"`).
28
+ - **Success is felt**: a subtle confirmation animation on verification — a settle, a check, the boxes tinting to the brand's positive tone — before the transition begins.
29
+ - **Errors are inline and immediate**: the boxes shake or tint, a one-line message appears in place, the code clears for retry with focus back on the first box. Never a browser alert, never a separate error page.
30
+
31
+ ## The send and resend flow
32
+
33
+ After the user enters their address and requests a code, confirm it plainly: "Check your email" with the actual address displayed (mistyped addresses are the most common failure, and showing it is the fix — pair it with a quiet "edit" affordance to go back). Include resend with a visible cooldown ("Resend in 30s") so the user is never staring at a dead link. The transition from enter-address to enter-code is a designed motion within the same composition — a slide or crossfade — never a page reload.
34
+
35
+ Design the unhappy paths with the same care: a code that never arrives (the resend plus a check-spam hint, or the other method as fallback), an expired code (say so, offer resend — don't make the user diagnose it), rate-limited resends (the cooldown communicates it).
36
+
37
+ ## Mid-experience gates and progressive auth
38
+
39
+ Not every sign-in happens at the front door. Apps that allow anonymous use gate specific actions — and chat or voice agents open verification sheets mid-conversation. Design these gates as invitations rather than walls:
40
+
41
+ - **Keep the context visible.** The gate renders over the live experience (a sheet or modal with the app dimmed behind it), so the user can see the thing they're unlocking. Never navigate away to a full login page mid-flow.
42
+ - **Say why.** One line connecting the gate to the action ("Verify your number to save this booking") outperforms a bare "Sign in required."
43
+ - **Return the user precisely.** After verification, the user lands exactly where they were, with the gated action completed or one tap away — the conversation, the cart, the draft all intact.
44
+
45
+ ## The post-login transition
46
+
47
+ Entering the app is the payoff — make it seamless. No blank loading screens: if data needs to load, show the app shell immediately with skeleton states. The moment of transition can carry a small piece of brand motion (the login card releasing into the app), but it must be fast; the best post-login transition is barely noticed. Returning users with a live session skip the front door entirely — make sure the signed-in landing carries the same polish, since for them *that's* the entry.
48
+
49
+ ## The register's exclusions
50
+
51
+ Each of these reads as a template or an afterthought — never ship them: a password field or "forgot password" link in any form; the generic SaaS login card (gray page, blue button, product name in plain text where a brand should be); social sign-in buttons for providers the app doesn't have; a single bare text input for the verification code; full page reloads between auth steps; browser alerts for errors; "Welcome back!" boilerplate copy in place of the app's actual voice; CAPTCHAs or terms-checkbox clutter the platform never asked for.
52
+
53
+ ## Your deliverable: art direction, not suggestions
54
+
55
+ You art-direct this surface end-to-end. The developer has a terrible sense of design and will fill any gap you leave with a default — and on this surface the default is the generic template that undermines the app before it opens. Deliver an implementation-ready specification:
56
+
57
+ - **Exact values everywhere.** The login composition's dimensions and breakpoints, the card's radius/border/shadow, type sizes, the digit boxes' size/gap/radius, every transition duration and easing, all colors as hexes from the brand.
58
+ - **The flow, state by state.** Enter-address / code-sent / entering-code / verifying / success / error / resend-cooldown — what appears, what animates, where focus lands — plus the delegated path and the mid-experience gate variant when the app has them.
59
+ - **One answer per question.** If you would accept either of two options, pick one and prescribe it. "Something like," "roughly," and "consider" are how implementations go generic; the only tolerances that exist are the ones you state numerically.
60
+ - **A verification checklist.** End with the specific things to screenshot-check after implementation — the login screen on mobile and desktop, the code boxes mid-entry and in the error state, the send-confirmation with the address shown, the post-login skeleton — so the developer can prove the direction landed rather than assume it did.
@@ -0,0 +1,59 @@
1
+ ---
2
+ name: Agent Chat Experience
3
+ what: The holistic experience of an app's chat agent — message design, streaming behavior, thinking and tool activity, the composer, the empty state, and threads, composed as one designed surface. Chat is often the app's most-used interface and the one users judge hardest, because they arrive fluent in it - a first-class design deliverable end to end, never a bolted-on widget. This reference carries the full craft recipe — the modern message shape, streaming and transition rules, tool-activity presentation, composer and empty-state patterns — that separates a product-grade conversation from a generic chatbot.
4
+ when: Before designing (or reviewing) an agent chat UI — its layout and placement, message design, streaming behavior, thinking/tool presentation, composer, empty state, or thread navigation.
5
+ ---
6
+
7
+ # Agent Chat Experience
8
+
9
+ An app's chat agent is one designed surface, and you own it end-to-end: the messages, the way text streams in, the quiet rows that show the agent working, the composer, the first screen, the thread history. Chat is frequently the most-used interface an app has, and the one users judge hardest — everyone arrives already fluent in chat, holding expectations set by the best chat products in the world. That fluency is the design constraint that makes this surface different: **the win is craft within convention, not novel form.** Where a voice agent's centerpiece rewards invention, a chat UI rewards a familiar shape executed to a standard users didn't expect from this app. Meet the expectations, then surpass them.
10
+
11
+ Placement is itself a design decision, made deliberately: a full-page conversation when chat is a primary surface; a side panel when the agent works alongside content; inline, next to the thing the agent acts on, when the conversation is about one artifact. A floating bubble in the corner is the support-widget pattern — it tells the user this agent is an afterthought.
12
+
13
+ ## The register
14
+
15
+ - **Modern document-flow conversations (Claude, ChatGPT in their current form)** — the agent's turns are a flowing *document*, not a bubble: generous measure, real markdown hierarchy set in the brand's type, comfortable line height, no container fighting the text. The user's turns are compact right-aligned chips. That asymmetry is the modern shape — the agent writes, the user interjects — and it's what makes long, substantive answers readable.
16
+ - **iMessage** — the physics: scroll that sticks to the bottom while new content arrives and releases the moment the user scrolls up, momentum that feels native, rhythm and breathing room between turns. People know in one flick whether a chat scroll is right.
17
+ - **Linear** — the finish: exact alignment, restrained color, subtle borders over heavy shadows, motion only where it informs. Nothing in the transcript exists for decoration.
18
+
19
+ And the register's explicit exclusions — each of these reads as a dated messenger or a bolted-on bot; never ship any of them: symmetrical bubbles on both sides (the 2015 messaging-app look); the SaaS support-widget (floating launcher, branded header bar, "We typically reply in a few minutes"); avatar-vs-avatar rows on every message; sparkle-emoji or gradient "AI" badges and buttons; robotic empty states ("Hello! I'm your AI assistant. How can I help you today?"); raw JSON, method names, or tool ids in the transcript; unstyled gray-box code blocks.
20
+
21
+ ## Messages are typography
22
+
23
+ A conversation is 95% type, so the brand's typography does almost all of the work — treat message design as editorial design. Prescribe it exactly: the measure of the agent's column (comfortable reading width, not full-bleed), the type scale for body and for markdown headings inside agent turns, the spacing rhythm between turns versus within a turn, the chip treatment for user messages (background, radius, max-width, alignment). Timestamps stay quiet — visible on demand or at conversation breaks, never shouting on every row. No avatars unless they carry real meaning (a multi-agent app, a human-handoff surface); a two-party conversation doesn't need faces to say who's talking.
24
+
25
+ Code blocks, tables, and lists inside agent turns are part of the brand too: styled, syntax-lit, copy-affordanced — an unstyled gray block in the middle of a designed transcript reads as a bug.
26
+
27
+ ## Streaming craft
28
+
29
+ Streaming is the heartbeat of the surface, and it must feel poured, not stuttered. Design the message lifecycle — thinking → streaming → complete — as one set of continuous, layout-shift-free transitions:
30
+
31
+ - **Thinking** shows as a compact, in-character indicator the moment the user sends (the optimistic send is non-negotiable: the user's message appears instantly, the indicator with it). If the model emits visible reasoning, give it a collapsed-but-present treatment the user can expand — never a wall of gray text pushed above the answer.
32
+ - **Streaming** renders batched (~50–100ms per paint, not per token) so arrival reads as pouring; give the streaming edge a treatment — a soft cursor or shimmer — so "still writing" is legible at a glance.
33
+ - **Completion** is a settle, not a jump: the cursor fades, actions (copy, retry) ease in. No element of the transcript moves except by growing downward.
34
+ - The transition between these states never reflows what's already on screen — reserve space for indicators instead of inserting them.
35
+
36
+ ## Tool activity
37
+
38
+ When the agent calls the app's methods, the transcript should show it working the way the product would say it: a compact status row in the app's voice ("Booking your appointment…", "Searching your library…"), appearing when the call starts and resolving in place when it lands — never raw method names, spinners without labels, or JSON. Design the row as a real element of the transcript: aligned to the agent's column, quiet, layout-stable.
39
+
40
+ Results deserve more than prose when they're structural: the record the agent pulled up, the item it created, the rows it found can render as real UI inside the turn — a card, a compact list, the app's own components — so the agent visibly operates the same product the user sees. Decide which tools earn a rendered result and prescribe what it looks like. When the agent runs several tools in a row, collapse them into one grouped status that expands on demand; a stack of six status rows reads as noise.
41
+
42
+ ## The composer
43
+
44
+ The composer is a designed object, not a bare input: an auto-growing textarea with a clear send affordance, a visible stop control while the agent is streaming, and an attachment affordance when the app supports uploads. The placeholder is the agent's personality in six words — never "Type a message...". On mobile, the composer and the virtual keyboard are the layout: design for the reduced viewport instead of letting the keyboard crop the transcript.
45
+
46
+ ## The empty state and threads
47
+
48
+ The first screen is the character's opening move, in the agent's voice and the brand's type: a greeting that sounds like *this* agent, a few suggested prompts as designed elements (real questions this app answers, tappable, skippable), or a concise line about what it can do. The user must always be able to just start typing.
49
+
50
+ Thread history is real navigation, not an afterthought dropdown: design where past conversations live (a sidebar on desktop, a sheet or screen on mobile), how titles read, and how a new thread starts. If the app allows anonymous chat, account for the sign-in moment — the conversation should visibly survive it.
51
+
52
+ ## Your deliverable: art direction, not suggestions
53
+
54
+ You art-direct this surface end-to-end. The developer has a terrible sense of design and will fill any gap you leave with a default — and defaults are how a designed conversation decays into a generic chatbot. Deliver an implementation-ready specification:
55
+
56
+ - **Exact values everywhere.** The agent column's measure, type sizes and line heights for body and markdown levels, chip radius and max-width, spacing between and within turns, indicator dimensions, streaming batch interval, transition durations and easings, all colors as hexes from the brand.
57
+ - **The message lifecycle, state by state.** Sending / thinking / streaming / complete / error — what appears, where space is reserved, what animates — plus the tool-row lifecycle (start, running, resolved, grouped), so no transition is left to improvisation.
58
+ - **One answer per question.** If you would accept either of two options, pick one and prescribe it. "Something like," "roughly," and "consider" are how implementations go generic; the only tolerances that exist are the ones you state numerically.
59
+ - **A verification checklist.** End with the specific things to screenshot-check after implementation — no layout shift while streaming (capture mid-stream), tool rows appearing and resolving cleanly, the empty state, code blocks inside a turn, the mobile viewport with the keyboard up — so the developer can prove the direction landed rather than assume it did.
@@ -0,0 +1,82 @@
1
+ ---
2
+ name: Data Visualization
3
+ what: Charts, dashboards, metrics, and tables of numbers — the surfaces where an app shows its data. Most remy apps have one, and it is where designs most reliably go lazy - default chart-library styling, chart junk, and fake-looking demo curves. The craft is brand-derived and hand-built: models are genuinely strong at authoring SVG charts directly (the default medium — no library), with visx as the escalation for heavyweight interaction, plus the typography of numbers, color-for-data discipline, and honest seed data that make a dashboard read as a real instrument.
4
+ when: Before designing (or reviewing) any chart, graph, dashboard, metric tile, sparkline, or data-heavy table — anything that visualizes the app's data.
5
+ ---
6
+
7
+ # Data Visualization
8
+
9
+ Data surfaces are where an app earns trust: a dashboard that reads like a real instrument makes the whole product feel serious, and one assembled from chart-library defaults makes it feel like a template. This is also where design most reliably goes lazy — three failure modes account for almost all of it: **default library styling** (the recognizable palette and tooltip of a chart package), **chart junk** (decoration that encodes nothing), and **fake demo data** (perfect curves that mark the dashboard as a mockup). Everything in this reference exists to defeat those three.
10
+
11
+ ## The medium: hand-authored SVG first
12
+
13
+ **The default medium is SVG you author directly — no charting library.** You (and the developer) are genuinely good at this: a line chart is a `<path>`, a bar chart is rects, a sparkline is a 20-line component — and direct authorship gives brand-exact control over every hairline, tick, and easing with zero dependency weight. Most dashboards need nothing more: lines, bars, areas, donuts, sparklines, small multiples, and metric tiles are all comfortably hand-built.
14
+
15
+ - **Escalate to visx** for genuinely heavyweight work: brush-and-zoom, force layouts, complex time scales with smart tick generation, streaming series, hierarchies, geographic projections. visx is the right escalation because it is d3's math wrapped in unstyled React primitives — you keep authoring the visual layer yourself; the library only does scales and layout. Prefer it over raw d3, and never adopt a styled chart package (Chart.js, Recharts, ECharts and kin) — their baked-in palettes, fonts, and tooltips are precisely the "defaulted" look this reference exists to prevent.
16
+ - **Canvas is the escape hatch for scale**, not style: reach for it only past roughly ten thousand rendered points (dense scatter, long high-frequency series), where SVG's DOM cost becomes real. The visual rules below apply identically.
17
+
18
+ ## The register
19
+
20
+ Touchstones, stated as qualities to reproduce rather than moods:
21
+
22
+ - **New Relic** — density done right. Sparklines carrying real trend at tiny sizes; the billboard pattern (a large current value, its delta, and a small trend chart composed as one tile); color used as *series identity*, never decoration; a dark ground with luminous series; density achieved through small multiples and rhythm, never through clutter.
23
+ - **Bloomberg / TradingView** — the terminal register. Information density as the aesthetic: monospaced/tabular figures everywhere, thin crosshairs, precise session gridlines, color reserved almost entirely for signal (up/down, in/out of range), instant legibility at a glance across dozens of numbers. Right for finance-adjacent and power-user tools where density *is* the brand.
24
+ - **Stripe's dashboard** — the light editorial register. Hairline charts, generous whitespace, restrained single-accent palettes, numbers doing the talking with charts in a supporting role. Right for business tools that want calm authority instead of mission-control density.
25
+ - **Observable** — the craft canon. Direct labeling instead of legends, considered axes, every mark earning its place; what a chart looks like when someone thought about it.
26
+ - **Apple Health / Fitness** — the consumer register. Friendly, chunky, glanceable: rounded bars, large type, one number per view that matters, charts you can read at arm's length. Right for consumer wellness/habit/lifestyle apps where the terminal look would feel hostile.
27
+
28
+ Pick the register the app's brand and audience call for — density and calm are both excellent when chosen deliberately.
29
+
30
+ ## Numbers are typography
31
+
32
+ The cheapest, highest-leverage move on any data surface, and the one most often skipped:
33
+
34
+ - **`font-variant-numeric: tabular-nums`** on every column of figures, every ticking value, every timer — proportional figures wobble as values change and never align in columns.
35
+ - **Right-align numeric columns** in tables; align to the decimal when precision varies.
36
+ - **Precision discipline**: two or three significant figures for display ("4.2K", "12.4%", "1.3s") — never raw floats with a tail of decimals. Full precision belongs in the tooltip.
37
+ - **Units styled as units**: smaller, lighter, or muted next to the value ("142 ms", "$4,210"), not baked into the number's own weight.
38
+ - **Big numbers are typography first**: a metric tile is a type specimen — the brand's face at display scale, the delta and label set quietly around it. The chart under it is context, not the star.
39
+
40
+ ## Color for data
41
+
42
+ Derive every data color from the app's brand, then apply data discipline on top:
43
+
44
+ - **Categorical**: at most six distinguishable hues, separated enough in lightness to survive colorblindness and grayscale; assign each series a color once and keep it consistent across every chart and view in the app.
45
+ - **Sequential and diverging ramps** for intensity and for above/below-baseline — built from brand hues, perceptually ordered (lightness does the work), not rainbow.
46
+ - **Reserve semantic color.** If the app uses green/red (or any pair) for good/bad, those hues never moonlight as ordinary series colors.
47
+ - **Highlight by contrast**: the focused series saturated, everything else muted to the same quiet gray-toned family — not seven series all shouting.
48
+ - Gridlines, axes, and labels sit in the muted end of the palette; the data gets the color.
49
+
50
+ ## Anatomy restraint
51
+
52
+ - Gridlines quiet or absent — a few horizontal hairlines at most; never a full grid cage.
53
+ - Few ticks, no redundant axis lines, no boxed borders around the plot.
54
+ - **Direct labeling over legends** wherever the layout allows — label the line's endpoint, not a color key the eye has to shuttle to.
55
+ - **Honest scales**: bars start at zero; a truncated line-chart axis is acceptable for trend but never for magnitude comparison; time axes have honest, even intervals.
56
+ - The chart-junk ban, explicitly: no 3D, no drop shadows on marks, no gradient-filled bars, no exploded pies, no dual y-axes (two charts beat one lie), no decorative icons inside plots.
57
+
58
+ ## Chart-type judgment
59
+
60
+ Line for change over time; bar for comparison across categories; area only when the quantity is cumulative or the filled mass means something; a **table is often the correct visualization** (sortable, tabular-numeric, with sparklines inline if trend matters); small multiples over one overloaded chart; donut/pie only for a single part-of-whole at a glance, never a six-slice pie; scatter for correlation with a fitted context line when it helps. When a chart needs a paragraph to explain, the chart is wrong.
61
+
62
+ ## Interaction and life
63
+
64
+ - **Design the tooltip** — a browser-default tooltip on a designed chart is a slop marker. Spec its card: background, radius, type scale, the full-precision values, and the crosshair or point-highlight that anchors it.
65
+ - **Crosshair on time series**; hover states that lift the focused series and mute the rest.
66
+ - **Draw-in animation on first render only** — a line drawing in or bars rising once is a designed moment; replaying it on every data refresh is noise.
67
+ - **Live updates tween.** Apps here poll routinely — when a value or series updates, animate the transition (a few hundred ms, eased) instead of snapping; ticking numbers count toward their new value with tabular figures so nothing shifts.
68
+ - **Loading and empty are designed states**: skeletons shaped like the chart they precede, and a zero-state that reads as an invitation ("data appears after your first sync") rather than an empty axis frame.
69
+ - **Responsive means recomposed**, not shrunk: on mobile, reduce tick counts, drop to sparkline-plus-number tiles, stack small multiples — an illegible miniature of the desktop chart is a failure.
70
+
71
+ ## Honest seed data
72
+
73
+ Demo and scenario data must look like it came from a real system: plausible noise, weekday/ weekend rhythm, the occasional gap or spike, series that don't move in lockstep. A perfect sine wave or a smooth exponential marks the whole dashboard as fake in one glance — it's the data equivalent of lorem ipsum. Specify realistic shapes when you hand seed guidance to the developer (ranges, trend, noise character, anomalies worth including), and make sure the axis ranges are set from the data's real domain, not defaulted.
74
+
75
+ ## Your deliverable: art direction, not suggestions
76
+
77
+ You art-direct this surface end-to-end. The developer has a terrible sense of design and will fill any gap you leave with a library default — and on this surface the defaults are instantly recognizable as no design at all. Deliver an implementation-ready specification:
78
+
79
+ - **Exact values everywhere.** The data palette as hexes with role mapping (each series, the ramps, semantic colors, gridline/label grays), stroke widths, tick counts per breakpoint, tile dimensions, type scale for values/labels/deltas, tooltip card spec, animation durations and easings.
80
+ - **Per-chart engine calls.** For each visualization: hand-SVG or visx (or canvas past the point threshold), the chart type and why, its axes/scales, and its responsive recomposition.
81
+ - **One answer per question.** If you would accept either of two options, pick one and prescribe it. "Something like," "roughly," and "consider" are how implementations go generic; the only tolerances that exist are the ones you state numerically.
82
+ - **A verification checklist.** End with the specific things to screenshot-check after implementation — tabular alignment in columns, the designed tooltip, first-render vs refresh animation, the empty and loading states, the mobile recomposition, and the seed data's realism — so the developer can prove the direction landed rather than assume it did.
@@ -1,3 +1,9 @@
1
+ ---
2
+ name: Images & Visual Assets
3
+ what: Generating, editing, and browser-rendering production imagery for the app — editorial photography, illustrations, icons and logos, Open Graph share images, transparent assets, and token-exact rendered graphics, all delivered as ready-to-use CDN URLs. The image tools (generateImages, editImages, renderImage) are always available; this reference is how to use them well — prompt craft, the engine split between model and browser, icon/OG recipes, CDN transforms — and when imagery is genuinely additive versus shoehorned decoration.
4
+ when: Before generating, editing, or rendering any image (generateImages / editImages / renderImage), or before specifying imagery, icons, logos, or share images in a design.
5
+ ---
6
+
1
7
  ## Photo and Image Guidelines
2
8
 
3
9
  Important: All images used in the app might be high resolution and high quality. If serving them via the mindstudio cdn, make sure to specify the ?dpr=3 param for retina displays.
@@ -0,0 +1,103 @@
1
+ ---
2
+ name: Voice Agent Experience
3
+ what: The holistic on-screen experience of a live voice agent — the audio-reactive centerpiece, the streaming captions, the way tool activity surfaces, and the call controls, composed as one scene. It is the most visual thing a voice app has and a first-class design deliverable end to end - a real-time computed centerpiece (WebGL; three.js when a dedicated voice-first app should wow), layout-stable captions, in-character tool status, controls that belong to the composition. This reference carries the full craft recipe — technique families, motion language, palette discipline, caption and control patterns, and the performance budget that separate a living instrument from a janky blob with widgets around it.
4
+ when: Before designing (or reviewing) the UI of a voice interface — the agent's centerpiece, captions, tool-activity presentation, or call controls.
5
+ ---
6
+
7
+ # Voice Agent Experience
8
+
9
+ A voice agent, live on screen, is one composed scene, and you own the whole scene: the centerpiece the user watches, the captions streaming beneath it, the quiet line that says what the agent is doing, and the controls that frame it. Every other surface of a voice app is transient — this scene is on screen from the first "listening" to the last "ended," and it *is* the product's face while the product is speaking. It deserves the same design investment as the brand itself, and it is the single most common place voice apps ship something embarrassing: a pulsing CSS circle with default buttons scattered around it.
10
+
11
+ Design the scene as a composition with hierarchy — centerpiece dominant, captions supporting, tool status quiet, controls present but calm — not as four widgets that happen to share a screen. And decide the framing deliberately: when the app has a web interface, voice is a layer over it (the app stays visible and usable); a full-screen voice mode is the immersive option for apps where the conversation is the product — earn it, don't default to it.
12
+
13
+ ## The register
14
+
15
+ Whatever the app, this surface pulls from one aesthetic register: the 2026-and-beyond language of ambient intelligence — future-ish, clean, quietly beautiful. The app's brand supplies the hues, the type, and the object; this register supplies the bearing. Touchstones worth drawing on, stated as qualities to reproduce rather than moods:
16
+
17
+ - **Siri / Apple Intelligence** — intelligence rendered as *light behaving intentionally*: a luminous, iridescent presence on a restrained neutral ground, hardware-grade polish, color that glows from within the form rather than being painted onto it. Note what it never does: no mascot, no face, no skeuomorphic microphone.
18
+ - **Arrival (the film)** — monolithic calm. An immense, precise thing communicating slowly: vast negative space, a restrained near-monochrome ground, organic form emerging from exact structure, motion measured in tens of seconds. The awe comes from patience and scale, not activity.
19
+ - **Precision instruments** — the confidence of something measured: exact alignment, real signal driving every moving element, monospaced labels earning their place, nothing decorative that doesn't encode state.
20
+
21
+ And the register's explicit exclusions — each of these reads as costume sci-fi or AI slop; never ship any of them: the hologram-cockpit HUD (fake crosshairs, orbiting rings, scattered random digits, wireframe globes, corner brackets); the generic AI-assistant look (an indigo-to-violet gradient blob with an outer glow); chrome and lens flares; "digital rain"; a robot or assistant mascot in any form.
22
+
23
+ ## The centerpiece
24
+
25
+ **Computed, structured, alive, brand-derived.** The piece is rendered in real time, every frame — a real-time WebGL rendering, not a video, not a GIF, not a CSS transform on a blurred div. Raw WebGL or a small helper library is right for a voice layer over an app; three.js is justified when the app is voice-first and the visual is the product's hero. What makes the difference between "computed" and "janky" is structure: the good ones read as something *measured or sampled* — an instrument, a scan, a constellation — never as smoke, lava, or a screensaver.
26
+
27
+ The shape is not prescribed. A sphere is one option among many, and often the least interesting. Let the domain pick the object: a vector-search product wants an embedding constellation or a query lighting up its neighbors; a terminal-flavored dev tool wants a levels meter or scope; a wellness app might want something botanical built from the same sampled medium. Whatever the object, it should be unmistakably *this app's*.
28
+
29
+ ### Archetype families
30
+
31
+ - **Computed point clouds** — thousands of individually lit points forming one recognizable object. Placement is algorithmic and even (a Fibonacci / golden-angle lattice for surfaces, low-discrepancy sampling for volumes) so coverage is uniform with no clumping — even sampling is most of why these read "computed" rather than "smoke." Objects: spheres, toruses, terrains, constellations, node graphs, lattices.
32
+ - **The instrument family** — meters, scopes, waveform lattices, spectrum bars with real signal behind them. Terminal-adjacent, precise, monospaced-label energy. Often the right call for developer tools and utilitarian brands where an orb would feel like costume jewelry.
33
+ - **Whatever the brand suggests** — the families above are starting points, not a menu. The test is the same either way: does it look designed for this app, and does it look *computed*?
34
+
35
+ ### Craft techniques
36
+
37
+ These are the moves that make the difference, stated as implementable direction. Prescribe them concretely to the developer — point counts, timings, blend modes — not as vibes.
38
+
39
+ - **Even, algorithmic placement.** Points and elements sit on a computed lattice, never `Math.random()` scatter. Uniform coverage is what makes the object read as sampled and sharp.
40
+ - **Sharp opaque elements, not additive glow.** Each point is a small, depth-tested, round sprite with a solid core and a thin feather, on normal blending. Counterintuitively, *avoiding* additive glow is what keeps the piece crisp and gives it real volume; additive blending is how you get the hazy nebula.
41
+ - **Depth is the shading language.** Front-facing elements are brighter, larger, and more saturated; back elements recede — dimmer, desaturated, smaller. This single trick turns a flat particle set into something with genuine dimension.
42
+ - **The whole spectrum, present at once.** Map hue to a *spatial axis* through the object — not per-point random, not a time cycle. Every color exists across the surface simultaneously, and a slowly precessing spectral axis makes the color reorganize and flow without ever fading or strobing. This is the signature move people respond to.
43
+ - **Slow, layered, organic motion.** Think tens of seconds: a full rotation around ~50s, a gentle breathe (±1–2% scale on a ~7s sine), simplex-noise undulation, an occasional slow scan band washing across the object. Everything calm; nothing flashy. Layered slow motions read as alive; one fast motion reads as a loading spinner.
44
+ - **Give the light something to bloom against.** Sit the piece on a dark ground with a soft, blurred caustic or ambient pool behind it in the anchor hue, so the saturated points have depth behind them instead of floating on flat black.
45
+
46
+ ### Palette discipline
47
+
48
+ Derive the hues from the app's brand — never ship a stock palette. The discipline that keeps an iridescent palette from turning to mud:
49
+
50
+ - **One anchor hue** that owns the piece (and matches the brand's identity).
51
+ - **Warm as a deliberate minority** — a coral or amber presence, never half the wheel.
52
+ - **A bridge hue** between the warm end and the cool end so the loop never blends into brown.
53
+ - **A near-white** reserved for a handful of hot sparkle points and pulse crests.
54
+ - Bake the ramp into a small gradient texture (e.g. 256×1) and sample it in the shader — cheap, consistent, and easy to swap when the brand evolves.
55
+
56
+ ### State mapping
57
+
58
+ The centerpiece carries the agent's state machine — idle → connecting → listening → thinking → speaking → ended — and reacts to live audio amplitude while listening and speaking. Design a distinct-but-related behavior for each state (at rest it should settle into the hero look), and **always pair the piece with a text state label**: state is never conveyed by color or motion alone. Specify each state's behavior explicitly — the amplitude response while listening, what "thinking" looks like (often the scan band's moment), how speaking differs from listening — so the developer isn't left to invent transitions.
59
+
60
+ ## Captions
61
+
62
+ Captions are typography — design them like any other type in the app, with the brand's faces and scale, not a default sans in a gray box. Both sides of the conversation stream as captions; they are what makes the agent feel accurate and they are the accessibility story. The rules that make them feel engineered rather than jittery:
63
+
64
+ - **Layout-stable, always.** Reserve a fixed-height caption region so arriving text never shifts the composition around it. Streaming segments update in place (each event carries the segment's full text — replace, never append), the visible line count is capped, and old lines fade out rather than pushing content down. A caption region that reflows the page on every event reads as jank and fights the centerpiece for attention.
65
+ - **Smooth arrival.** Ease new words in (a fast opacity ramp is enough); never let text pop hard at stream rate. The target feel is a broadcast lower-third, not a log tail.
66
+ - **Differentiate the voices quietly.** User and agent captions need distinct treatments — weight, color, or alignment — legible at a glance without reading like a chat transcript. The agent's caption can carry slightly more presence; it's the one speaking.
67
+ - **Place them in the composition.** Captions sit in the centerpiece's orbit (typically below), sized to support rather than compete. In a layer-over-app treatment they stay compact — one or two lines — so the app remains usable behind them.
68
+
69
+ ## Tool activity
70
+
71
+ When the agent calls the app's methods mid-conversation, the screen should acknowledge it the way the agent's voice does: quietly, in character, in the app's language. Design this layer — don't let it default to nothing (the agent looks frozen during a slow tool) or to raw output.
72
+
73
+ - **A compact inline status** near the captions or centerpiece: "Booking your appointment…", "Looking that up…" — the app's voice, never raw method names, spinners with no label, or JSON.
74
+ - **Results can render.** Successful tool results are delivered to the caller's browser in lockstep with the spoken answer — so the record the agent just pulled up, the booking it made, or the citation it found can appear as a real UI element (a card, a highlighted row) while the agent says it. Decide which tools deserve a visual result and what it looks like; this is the moment the voice agent proves it's operating the same product the user sees.
75
+ - **Transient by default.** Status lines appear, resolve, and clear; results persist only when they're genuinely useful to keep on screen. The scene should end each exchange as calm as it started.
76
+
77
+ ## Controls and chrome
78
+
79
+ Controls are part of the composition, not an afterthought row of default buttons.
80
+
81
+ - **Mute and end call: always visible, always working**, styled as first-class elements of the scene. These two are non-negotiable; everything else is optional chrome.
82
+ - **A text input, when exactness matters.** The platform supports injecting typed text into the live conversation — design the affordance for apps where users will need to hand over an address, a code, or an email (typing beats spelling it aloud three times). Keep it secondary: a small "type instead" affordance, not a chat box competing with the mic.
83
+ - **Design the awkward states.** Mic permission denied (a gentle, non-blaming explanation with a path to fix it — never a dead end), connecting (the centerpiece's connecting behavior plus its label), reconnecting/ended (what the scene settles into). These states are where a defaulted UI most obviously falls apart.
84
+ - **The state label** (listening / thinking / speaking) lives with the centerpiece — set it in the brand's type, and treat it as chrome that's always legible over whatever the piece is doing.
85
+
86
+ ## Performance and fallbacks — from day one
87
+
88
+ These are why the good ones perform, and they are design requirements, not optimizations to defer:
89
+
90
+ - Cap `devicePixelRatio` at 2.
91
+ - Reduce density on mobile (roughly 40% of the desktop element count).
92
+ - Pause the render loop entirely when the piece scrolls off-screen or the tab is hidden.
93
+ - Ship a static-but-labeled fallback for `prefers-reduced-motion` and for no-WebGL.
94
+
95
+ ## Your deliverable: art direction, not suggestions
96
+
97
+ You art-direct this surface end-to-end. The developer has a terrible sense of design and will fill any gap you leave with a default — and defaults are how this surface dies. Deliver an implementation-ready specification:
98
+
99
+ - **Exact values everywhere.** Element counts (desktop and mobile), sizes in px, every duration in seconds, easing curves, blend modes, sprite core/feather proportions, amplitude-response ranges, the full color ramp as ordered hexes, dimensions for the caption region and controls.
100
+ - **State by state.** Cover idle / connecting / listening / thinking / speaking / ended — what the centerpiece does, what the label reads, how captions and controls behave — and the transitions between states, so none are left to improvisation.
101
+ - **Pseudocode or shader math where prose is ambiguous.** The lattice formula, the hue-to-axis mapping, the noise parameters — a few lines of code communicate what a paragraph can't.
102
+ - **One answer per question.** If you would accept either of two options, pick one and prescribe it. "Something like," "roughly," and "consider" are how implementations go generic; the only tolerances that exist are the ones you state numerically.
103
+ - **A verification checklist.** End with the specific things to screenshot-check after implementation — sprite sharpness, caption stability while streaming, each awkward state, the fallbacks — so the developer can prove the direction landed rather than assume it did.
@@ -76,6 +76,8 @@ Common operations:
76
76
 
77
77
  **Seeding the initial roadmap:** Write an MVP item (slug "mvp") capturing what's being built, then generate future roadmap ideas. Think big — what would the team build in the next quarter? Six months? Year? The self-check: would a user be excited showing this roadmap to a friend? Create 10-15 roadmap items for the initial seeding. At least 3 items should be large effort. At least 2 lanes should extend beyond the current product scope into genuinely new territory. Write the roadmap index with lane structure and narratives. Generate the pitch deck.
78
78
 
79
+ When the product is consumer-facing — something whose deployed link will be shared with the public — the first item after the MVP should usually be a proper marketing landing page. The MVP deliberately spends all of its energy on the product itself, so visitors otherwise arrive at the sign-in screen; a real landing page (hero, story, imagery, a reason to sign up) is the natural foundation of a growth lane and the highest-leverage first follow-up. Use judgment: for internal tools, team utilities, and enterprise workflows a landing page is dead weight — the pitch deck carries that story instead.
80
+
79
81
  **Adding items:** The user or the coding agent wants to add something to the roadmap. Create the item, add it to the appropriate lane in the index, and update the index.
80
82
 
81
83
  **Marking items complete:** Update the status to `done` and append a history entry. Consider whether the completed feature unlocks or changes other roadmap items. Update the index if lane structure changed. When the assistant mark items complete, it takes a look at the rest of the roadmap and make sure the remaining items all still make sense. It sakes any adjustments it needs in order to keep everything holistic and synced, and also think about new items that the completed work makes possible. If there are new items, add them! Consider refining the pitch deck if the product's story has meaningfully evolved.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@mindstudio-ai/remy",
3
- "version": "0.1.276",
3
+ "version": "0.1.277",
4
4
  "description": "Remy coding agent",
5
5
  "repository": {
6
6
  "type": "git",