@mindstudio-ai/remy 0.1.256 → 0.1.258
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/headless.js +514 -315
- package/dist/index.js +517 -288
- package/dist/prompt/compiled/files.md +17 -3
- package/dist/prompt/compiled/interfaces.md +26 -472
- package/dist/prompt/compiled/methods.md +10 -15
- package/dist/prompt/compiled/scenarios.md +16 -0
- package/dist/prompt/compiled/sdk-actions.md +1 -2
- package/dist/prompt/{compiled/agent-interfaces.md → skills/agentInterfaces.md} +118 -10
- package/dist/prompt/skills/dataSources.md +131 -0
- package/dist/prompt/skills/inboundEmail.md +116 -0
- package/dist/prompt/skills/mcpInterfaces.md +280 -0
- package/dist/prompt/skills/restApi.md +149 -0
- package/dist/prompt/skills/scheduledJobs.md +51 -0
- package/dist/prompt/{compiled/task-agents.md → skills/taskAgents.md} +10 -8
- package/dist/prompt/skills/webhooks.md +108 -0
- package/dist/prompt/static/authoring.md +1 -1
- package/dist/prompt/static/instructions.md +1 -1
- package/dist/subagents/designExpert/prompts/images.md +1 -1
- package/package.json +2 -2
- package/dist/prompt/.notes.md +0 -194
- package/dist/prompt/compiled/README.md +0 -100
- package/dist/prompt/compiled/mcp-interfaces.md +0 -34
- package/dist/prompt/compiled/media-cdn.md +0 -51
- package/dist/prompt/sources/llms.txt +0 -1618
- package/dist/subagents/.notes-background-agents.md +0 -64
- package/dist/subagents/codeSanityCheck/.notes.md +0 -44
- package/dist/subagents/designExpert/.notes.md +0 -265
- package/dist/subagents/productVision/.notes.md +0 -79
|
@@ -1,64 +0,0 @@
|
|
|
1
|
-
# Background Agent Execution
|
|
2
|
-
|
|
3
|
-
## How it works
|
|
4
|
-
|
|
5
|
-
The parent agent decides at dispatch time whether a sub-agent runs in background by passing `background: true` in the tool input. The sub-agent doesn't know or care — it runs identically to foreground.
|
|
6
|
-
|
|
7
|
-
### Dispatch
|
|
8
|
-
|
|
9
|
-
```
|
|
10
|
-
visualDesignExpert({ task: "...", background: true })
|
|
11
|
-
productVision({ task: "...", background: true })
|
|
12
|
-
```
|
|
13
|
-
|
|
14
|
-
### Split lifecycle (runner.ts)
|
|
15
|
-
|
|
16
|
-
When `background: true`:
|
|
17
|
-
|
|
18
|
-
1. The runner creates its own AbortController (detached from the parent turn signal)
|
|
19
|
-
2. The sub-agent's LLM runs normally — streaming, thinking, tool calls
|
|
20
|
-
3. After the first LLM turn that produces text content, the runner **resolves the parent's promise early** with that text + `backgrounded: true`
|
|
21
|
-
4. The sub-agent loop continues in the background — more tool calls, more LLM turns
|
|
22
|
-
5. On completion, `onBackgroundComplete` fires, pushing the result to the notification queue
|
|
23
|
-
|
|
24
|
-
The parent agent gets the initial response immediately and continues its turn. The background agent keeps working.
|
|
25
|
-
|
|
26
|
-
### Result delivery (headless.ts)
|
|
27
|
-
|
|
28
|
-
A notification queue collects background completions. Delivery:
|
|
29
|
-
|
|
30
|
-
- **If Remy is idle** — deliver immediately as a hidden automated message
|
|
31
|
-
- **If Remy is mid-turn** — queue and flush after `turn_done`
|
|
32
|
-
- **Multiple completions** — batched into one message
|
|
33
|
-
|
|
34
|
-
Format (hidden, XML-tagged):
|
|
35
|
-
```
|
|
36
|
-
@@automated::background_results@@
|
|
37
|
-
<background_results>
|
|
38
|
-
<tool_result id="toolu_abc" name="visualDesignExpert">
|
|
39
|
-
Result text here...
|
|
40
|
-
</tool_result>
|
|
41
|
-
</background_results>
|
|
42
|
-
```
|
|
43
|
-
|
|
44
|
-
### Events
|
|
45
|
-
|
|
46
|
-
The `tool_start` event includes `background: true` when a tool is backgrounded. The frontend knows every subsequent event with that `parentToolId` is background work — no need to flag every individual event.
|
|
47
|
-
|
|
48
|
-
### Process management
|
|
49
|
-
|
|
50
|
-
Background agents stay in the tool registry after the parent's promise settles. The existing `stop_tool` and `restart_tool` stdin commands work on them. Stopping a background agent via `stop_tool` is how users cancel dangling work.
|
|
51
|
-
|
|
52
|
-
### Which sub-agents support this?
|
|
53
|
-
|
|
54
|
-
- **designExpert** — return font/color/layout recommendations immediately, generate images in background
|
|
55
|
-
- **productVision** — return initial plan immediately, write roadmap files in background
|
|
56
|
-
- **codeSanityCheck** — NOT a candidate, Remy needs the advice before proceeding
|
|
57
|
-
- **browserAutomation** — NOT a candidate, results inform Remy's next action
|
|
58
|
-
|
|
59
|
-
## Future considerations
|
|
60
|
-
|
|
61
|
-
- **Resource budgets** — token/cost ceilings for background agents running unattended
|
|
62
|
-
- **Checkpoint/resume** — serialized state for surviving process restarts
|
|
63
|
-
- **Speculative execution** — start work optimistically, cancel if the parent's reasoning goes a different direction
|
|
64
|
-
- **Fan-out** — dispatch multiple background agents in parallel, collect results
|
|
@@ -1,44 +0,0 @@
|
|
|
1
|
-
# Code Sanity Check Sub-Agent — Design Notes & Decisions
|
|
2
|
-
|
|
3
|
-
Notes from the initial build (March 2026).
|
|
4
|
-
|
|
5
|
-
## Purpose
|
|
6
|
-
|
|
7
|
-
A lightweight, readonly pre-build advisor. Reviews an approach before the main agent starts building and flags anything that would cause real pain. Most of the time, responds "lgtm."
|
|
8
|
-
|
|
9
|
-
## Why a sub-agent and not a prompt rule?
|
|
10
|
-
|
|
11
|
-
Two categories of issues this catches:
|
|
12
|
-
|
|
13
|
-
1. **Package freshness** — the main agent reaches for packages from training data that may be outdated or superseded. A prompt rule can't fix this because the agent needs to actually search the web to check. The sub-agent has `searchGoogle` and `fetchUrl`.
|
|
14
|
-
|
|
15
|
-
2. **Architecture/organization** — the main agent is optimized for execution velocity, which means it won't pause to reconsider file structure or schema design. A prompt rule saying "think about organization" would slow it down on every task. The sub-agent is called at specific moments (before big builds) so the cost is targeted.
|
|
16
|
-
|
|
17
|
-
## "Checked-out staff eng" personality
|
|
18
|
-
|
|
19
|
-
The prompt deliberately cultivates low-energy, high-signal energy. The failure mode we're avoiding: a conscientious reviewer that flags everything and becomes a bottleneck. The prompt explicitly says:
|
|
20
|
-
|
|
21
|
-
- "lgtm" is a complete response, use it often
|
|
22
|
-
- Tech debt is normal and sometimes useful
|
|
23
|
-
- Let nits, style preferences, and minor code smells slide
|
|
24
|
-
- A few sentences is ideal, never an essay
|
|
25
|
-
|
|
26
|
-
This prevents scope creep from the reviewer itself.
|
|
27
|
-
|
|
28
|
-
## Tools
|
|
29
|
-
|
|
30
|
-
Reuses the main agent's tools via `executeTool` passthrough — no duplicate tool implementations. Seven readonly tools: readFile, grep, glob, searchGoogle, fetchUrl, askMindStudioSdk, bash.
|
|
31
|
-
|
|
32
|
-
Bash is included for complex read operations (git history, package.json inspection, analysis scripts) with prompt guidance to use it for reading only.
|
|
33
|
-
|
|
34
|
-
## Second-person prompt (not third-person)
|
|
35
|
-
|
|
36
|
-
Unlike the product vision agent which uses third-person RP framing to break out of RLHF constraints, this agent benefits from the model's base personality: pragmatic, responsible, grounded. The second-person "you are" framing keeps it in its natural problem-solving mode, which is exactly what a sanity checker should be.
|
|
37
|
-
|
|
38
|
-
## "We/us" language in team.md
|
|
39
|
-
|
|
40
|
-
The team.md description of this agent uses collaborative language ("a schema decision that'll paint us into a corner") rather than corrective language ("you're about to make a mistake"). This avoids triggering ego-defensiveness in the main agent and frames the sanity check as a teammate, not a gatekeeper.
|
|
41
|
-
|
|
42
|
-
## Spec context
|
|
43
|
-
|
|
44
|
-
Spec files are injected into the system prompt via `loadSpecContext()` from `subagents/common/context.ts`. The agent sees the project's domain, data model, and design direction without the main agent needing to summarize.
|
|
@@ -1,265 +0,0 @@
|
|
|
1
|
-
# Design Research Agent — Design Notes & Decisions
|
|
2
|
-
|
|
3
|
-
Notes from the initial design of the design research sub-agent (March 2026).
|
|
4
|
-
|
|
5
|
-
## Purpose
|
|
6
|
-
|
|
7
|
-
The design research agent is Remy's design consultant. Any design question or decision gets delegated here. It handles both exploratory research ("give me three visual directions for a fintech app") and directed tasks ("the spec says Midnight #000000 and Snow #F5F5F7, what gradient would work for a hero section?"). It returns concrete, actionable output — URLs, hex values, font names with CSS URLs — that Remy can interpret and use however it sees fit.
|
|
8
|
-
|
|
9
|
-
## Scope
|
|
10
|
-
|
|
11
|
-
The agent covers five areas:
|
|
12
|
-
|
|
13
|
-
1. **Typography** — font selection and pairings from curated sources
|
|
14
|
-
2. **Color palettes** — brand color generation from seed colors, domain context, or reference sites; including modern CSS gradients
|
|
15
|
-
3. **Stock photography / placeholder images** — finding relevant imagery via curated stock sites
|
|
16
|
-
4. **Layout & responsive design** — research real products for layout patterns, propose interesting/non-generic compositions, optionally generate wireframe concepts via image generation. This is a major use case — Remy defaults to boring layouts without this.
|
|
17
|
-
5. **Visual reference analysis** — fetch + screenshot sites, analyze them for design insights
|
|
18
|
-
|
|
19
|
-
Icons and brand SVGs are deferred for now.
|
|
20
|
-
|
|
21
|
-
## No modes, no structured output format
|
|
22
|
-
|
|
23
|
-
The agent doesn't have separate "research mode" vs "proposal mode." The task description from Remy carries the context. Sometimes Remy asks for options to present to the user during intake; sometimes it asks for specific answers while building from a spec. The agent handles both naturally.
|
|
24
|
-
|
|
25
|
-
Output should include concrete resources (URLs, hex values, font names with CSS links) but doesn't need rigid YAML structure. Remy interprets the results.
|
|
26
|
-
|
|
27
|
-
## Architecture: static docs, runtime-sampled data, and tools
|
|
28
|
-
|
|
29
|
-
The agent's prompt is assembled from three layers:
|
|
30
|
-
|
|
31
|
-
### 1. Static documents (baked into prompt)
|
|
32
|
-
|
|
33
|
-
Guidelines and knowledge that don't change between invocations:
|
|
34
|
-
|
|
35
|
-
- **Gradient techniques** — what's current/dated, CSS techniques (oklch, color-mix, mesh gradients, grain overlays), reference sites
|
|
36
|
-
- **Animation patterns** — what's current/dated, libraries (Motion as default), performance rules
|
|
37
|
-
- **Color theory** — HSL rotation formulas, oklch color space, color-mix(), tint/shade derivation. No third-party API needed for palette generation.
|
|
38
|
-
- **Design philosophy** — distinctiveness, typography-first thinking, brand-level output, the AI defaults problem
|
|
39
|
-
- **Layout guidance** — what makes layouts interesting (asymmetry, creative whitespace, varied compositions), common anti-patterns to avoid. This is critical because layout is where AI-generated interfaces are weakest.
|
|
40
|
-
- **Visual reference analysis framework** — the consistent analysis structure (mood, color, type, layout, what's distinctive)
|
|
41
|
-
|
|
42
|
-
### 2. Runtime-sampled JSON arrays (injected at invocation time)
|
|
43
|
-
|
|
44
|
-
Data arrays that are sampled randomly each time the agent is invoked, to prevent bias and keep results fresh:
|
|
45
|
-
|
|
46
|
-
**Design inspiration screenshots:**
|
|
47
|
-
- Godly website thumbnails (URLs grabbed from Chrome dev tools, stored as JSON array)
|
|
48
|
-
- Other gallery thumbnails (Awwwards, Figma Community, etc.) as we collect them
|
|
49
|
-
- Just a big array of image URLs; the prompt loader samples ~3-5 per invocation
|
|
50
|
-
- The agent can use `analyzeDesignReference` on these to understand what it's looking at
|
|
51
|
-
|
|
52
|
-
**Font catalog + pairings:**
|
|
53
|
-
- Fontshare catalog (~100 fonts with metadata: name, slug, category, weights, tags)
|
|
54
|
-
- Fontshare curated pairings (59 heading/body combinations with recommended weights)
|
|
55
|
-
- Open Foundry catalog if we compile it
|
|
56
|
-
- The prompt loader samples ~10-15 fonts per invocation to prevent the agent from always reaching for the same favorites
|
|
57
|
-
- Fonts are provided with CSS URLs: `https://api.fontshare.com/v2/css?f[]={slug}@{weights}&display=swap`
|
|
58
|
-
- The agent should provide a URL to a font and let Remy figure out how to load it. No need for Google Fonts-specific guidance.
|
|
59
|
-
|
|
60
|
-
**Sampling mechanism:** The prompt loader (in `index.ts` or `prompt.ts`) reads the JSON arrays at startup. Each invocation of the design research tool picks a random subset and injects them into the system prompt before passing it to `runSubAgent`. The agent sees fresh examples every time without managing randomness itself.
|
|
61
|
-
|
|
62
|
-
### 3. Tools (live, query-dependent)
|
|
63
|
-
|
|
64
|
-
**`searchGoogle`** — general design research: font recommendations, "best finance app design 2026," real products in the user's domain
|
|
65
|
-
|
|
66
|
-
**`fetchUrl`** (with optional screenshot) — analyze reference sites, brand sites users share, font specimen pages
|
|
67
|
-
|
|
68
|
-
**`analyzeImage`** — general-purpose vision analysis. Takes a prompt and an image URL. The agent crafts the prompt based on what it needs to know ("what colors dominate this image?", "describe the typography choices", etc.)
|
|
69
|
-
|
|
70
|
-
**`analyzeDesignReference`** — specialized wrapper around `analyzeImage` with a pre-built analysis prompt. Takes just an image URL. Returns a consistent analysis: mood/aesthetic, color palette with approximate hex values, typography style, layout composition, and what makes it distinctive. Use this when analyzing screenshots for design inspiration rather than crafting a custom prompt each time.
|
|
71
|
-
|
|
72
|
-
**Stock photo search** — search Pexels, Unsplash, and Pixabay via MindStudio SDK CLI. No Google Images — too much junk/noise for design work. These are the only sources for stock photography.
|
|
73
|
-
|
|
74
|
-
**`generateImage`** (via MindStudio SDK) — for generating rough wireframe/layout concepts. TBD on prompt engineering for this.
|
|
75
|
-
|
|
76
|
-
All tools shell out to the `mindstudio` CLI. No external tool resolution needed.
|
|
77
|
-
|
|
78
|
-
## Visual reference analysis framework
|
|
79
|
-
|
|
80
|
-
When analyzing a screenshot or design image, the agent should consistently assess:
|
|
81
|
-
|
|
82
|
-
1. **Mood/aesthetic** — minimal, bold, editorial, playful, corporate, etc.
|
|
83
|
-
2. **Color** — dominant colors with approximate hex values, palette strategy (monochromatic, complementary, etc.), use of gradients
|
|
84
|
-
3. **Typography** — serif/sans/display, weight, size hierarchy, distinctive choices
|
|
85
|
-
4. **Layout** — composition (symmetric/asymmetric), grid structure, whitespace usage, content density
|
|
86
|
-
5. **What makes it distinctive** — the specific design choices that make this stand out from generic AI-generated interfaces
|
|
87
|
-
|
|
88
|
-
This framework is baked into the `analyzeDesignReference` tool so analyses are consistent. The general `analyzeImage` tool is available for ad-hoc questions.
|
|
89
|
-
|
|
90
|
-
## Randomization for inspiration browsing
|
|
91
|
-
|
|
92
|
-
When searching for design inspiration via tools (not the pre-sampled arrays), the agent tends to see the same top/featured results. Strategies to vary:
|
|
93
|
-
|
|
94
|
-
- **Varied search queries** — add style adjectives and the current year ("dashboard design minimal dark 2026")
|
|
95
|
-
- **Pagination/offset** — browse page 3 instead of page 1
|
|
96
|
-
- **Domain-specific searches** — search for real products in the user's space, not generic "design inspiration"
|
|
97
|
-
- **Multiple sources per task** — hit 2-3 different sources with different query strategies
|
|
98
|
-
|
|
99
|
-
For layout and UX patterns, research real products in the user's domain. For aesthetic inspiration, browse design galleries with varied search terms.
|
|
100
|
-
|
|
101
|
-
## Design philosophy notes
|
|
102
|
-
|
|
103
|
-
### The AI defaults problem
|
|
104
|
-
|
|
105
|
-
LLMs converge on statistical medians from training data. For design this means: Inter/DM Sans/Space Grotesk fonts, purple/indigo accents, three-boxes-with-icons layouts, safe neutral palettes. The design research agent exists partly to break this pattern by forcing the model to look at real-world references and curated sources rather than generating from priors. Runtime sampling of fonts and inspiration images further prevents the agent from developing its own biases.
|
|
106
|
-
|
|
107
|
-
### Color palettes should be UI-functional
|
|
108
|
-
|
|
109
|
-
"Pretty" palettes from generators (Coolors trending, etc.) are not necessarily good UI palettes. A UI palette needs: a background color, a readable text color on that background, an accent for interactive elements. Deriving palettes from real products or from color theory applied to brand colors produces better results than aesthetic palette generators.
|
|
110
|
-
|
|
111
|
-
### Typography is the highest-impact decision
|
|
112
|
-
|
|
113
|
-
Font selection has more impact on perceived quality than any other single design decision. The agent should spend proportionally more effort here — looking at specimens, considering personality match, checking that the font has the weights and styles needed for a complete type hierarchy.
|
|
114
|
-
|
|
115
|
-
### Brand-level, not implementation-level
|
|
116
|
-
|
|
117
|
-
The agent's output should be at the brand style guide level: "Midnight #000000 for dark surfaces, Snow #F5F5F7 for text." Not at the CSS variable level: "Background: #000000, BorderFocus: #3A3A3C." The coding agent (Remy) derives implementation details like accessibility contrast adjustments, dark mode palettes, and CSS variables from the brand-level output.
|
|
118
|
-
|
|
119
|
-
### Layout is where AI is weakest
|
|
120
|
-
|
|
121
|
-
This is the primary reason the design research agent exists for layout work. Without outside input, Remy will produce the same centered-content, three-column, card-grid layouts every time. The design research agent should push for: asymmetry, varied column widths, creative negative space, unexpected compositions, full-bleed elements, strong visual hierarchy through scale contrast. Analyzing real sites (via the sampled inspiration images and live research) is the best way to inject layout creativity.
|
|
122
|
-
|
|
123
|
-
## Modern gradients (2026)
|
|
124
|
-
|
|
125
|
-
### What's current
|
|
126
|
-
|
|
127
|
-
**Mesh / aurora gradients** — the dominant look. Multiple layered `radial-gradient()`s with `filter: blur()` over dark backgrounds. The Stripe/Linear/Vercel aesthetic. Creates organic, atmospheric backgrounds.
|
|
128
|
-
|
|
129
|
-
**Grain/noise overlays** — SVG `feTurbulence` filters layered under gradients. Combats color banding on long subtle gradients and adds tactile warmth.
|
|
130
|
-
|
|
131
|
-
**Glassmorphism (matured)** — subtle `backdrop-filter: blur()` with gradient tints. Used sparingly as an accent, not the whole design language.
|
|
132
|
-
|
|
133
|
-
**Animated gradient blobs** — hero sections with continuously morphing gradients. CSS `@keyframes` animating `background-position` on oversized gradients for simple cases; WebGL (like Stripe's minigl) for more complex effects.
|
|
134
|
-
|
|
135
|
-
### CSS techniques that matter
|
|
136
|
-
|
|
137
|
-
- **`oklch` color space in gradients**: `linear-gradient(to right in oklch, blue, green)` avoids the muddy gray zone that RGB/HSL produce. Perceptually uniform, vibrant transitions. Production-ready in all modern browsers. This is the single biggest upgrade for gradient quality.
|
|
138
|
-
- **`color-mix()`**: `color-mix(in oklch, #3b82f6 70%, white)` for generating tints/shades programmatically. Essential for design systems deriving UI colors from brand colors.
|
|
139
|
-
- **Relative color syntax**: `oklch(from var(--brand) calc(l * 1.25) c h)` derives lighter/darker/desaturated variants from a single token. Powerful for theme layers.
|
|
140
|
-
- **Stacked radial-gradients**: Multiple `radial-gradient()` layers with different positions/sizes to create mesh-like effects without canvas/WebGL.
|
|
141
|
-
- **Conic gradients**: Useful for pie charts, color wheels, angular shading.
|
|
142
|
-
|
|
143
|
-
### What looks dated
|
|
144
|
-
|
|
145
|
-
- Simple two-color linear gradients (the 2018 "purple to blue" hero)
|
|
146
|
-
- Instagram-style gradient borders as a primary design element
|
|
147
|
-
- Overly saturated, uniform gradients without texture or depth
|
|
148
|
-
- Flat gradient cards without noise, blur, or layering
|
|
149
|
-
|
|
150
|
-
### Reference sites
|
|
151
|
-
|
|
152
|
-
- **Stripe** — WebGL aurora gradient hero, flowing blues/purples/pinks/oranges. The gold standard.
|
|
153
|
-
- **Vercel** — prism/light-refraction effects, animated lines through gradients, grainy textures.
|
|
154
|
-
- **Linear** — dark backgrounds with precise, subtle gradient accents. Minimalist, purposeful.
|
|
155
|
-
|
|
156
|
-
### Implications for the design research agent
|
|
157
|
-
|
|
158
|
-
The agent can recommend gradient CSS alongside color palette proposals. Modern gradients are color theory applied through `oklch` — the seed color + HSL rotation math we're already baking in produces the inputs, and the agent can suggest specific gradient techniques (mesh, grain overlay, animated blob) appropriate to the app's aesthetic. The `color-mix()` and relative color syntax are how the coding agent (Remy) derives the implementation — the design research agent just needs to know these exist when making recommendations.
|
|
159
|
-
|
|
160
|
-
## Modern animations (2026)
|
|
161
|
-
|
|
162
|
-
### What's current
|
|
163
|
-
|
|
164
|
-
**CSS scroll-driven animations** — the biggest shift. `animation-timeline: scroll()` and `animation-timeline: view()` tie animations to scroll position purely in CSS, running off the main thread for smooth 60fps. ~85% browser support. Replaces Intersection Observer for entrance animations.
|
|
165
|
-
|
|
166
|
-
**Scroll-triggered animations** — new in Chrome 145 (2026): time-based animations that trigger at specific scroll offsets, declaratively in CSS. No JS scroll listeners needed.
|
|
167
|
-
|
|
168
|
-
**View Transitions API** — morphing between page states with cinematic transitions. Page navigations in SPAs, expanding cards to detail views. Growing browser support.
|
|
169
|
-
|
|
170
|
-
**Spring physics** — natural-feeling motion with spring-based easing rather than cubic-bezier. Motion (Framer Motion) has this built-in.
|
|
171
|
-
|
|
172
|
-
**Purposeful micro-interactions** — subtle hover/click feedback: scaling, color shifts, depth changes. The philosophy is "guide, confirm, smooth" — not "show off."
|
|
173
|
-
|
|
174
|
-
### Animation libraries for React
|
|
175
|
-
|
|
176
|
-
| Library | Size | Best for |
|
|
177
|
-
|---------|------|----------|
|
|
178
|
-
| CSS-native (scroll-timeline, @keyframes) | 0KB | Simple entrances, scroll effects, hover states |
|
|
179
|
-
| Motion (fka Framer Motion) | ~85KB | Complex state-driven animations, layout transitions, spring physics. The React default. |
|
|
180
|
-
| GSAP | ~78KB | Complex sequenced timelines, scroll-driven narratives, raw performance with many simultaneous tweens |
|
|
181
|
-
|
|
182
|
-
**Motion is the right default for most React app work.** Only reach for GSAP if building complex sequenced timelines.
|
|
183
|
-
|
|
184
|
-
### Performance rules
|
|
185
|
-
|
|
186
|
-
- Only animate `transform`, `opacity`, `filter` — these skip layout and paint (GPU-composited).
|
|
187
|
-
- `will-change` sparingly — promotes to GPU layer but overuse causes excessive memory.
|
|
188
|
-
- CSS scroll-driven animations are inherently performant (off main thread).
|
|
189
|
-
- Never animate `width`, `height`, `top`, `left`, `margin`, `padding` — triggers layout recalculation.
|
|
190
|
-
|
|
191
|
-
### What looks dated
|
|
192
|
-
|
|
193
|
-
- Parallax scrolling as a primary design pattern
|
|
194
|
-
- Manual `nth-child` stagger delays (use `sibling-index()` + `calc()` now)
|
|
195
|
-
- JS scroll event listeners for scroll animations
|
|
196
|
-
- Heavy animation libraries for simple UI toggles
|
|
197
|
-
- Loading spinners everywhere (skeleton screens preferred)
|
|
198
|
-
- Flashy transitions that don't serve UX purpose
|
|
199
|
-
- Bounce/elastic easing
|
|
200
|
-
|
|
201
|
-
### Reference sites
|
|
202
|
-
|
|
203
|
-
- **Apple.com** — scroll-driven storytelling. Information fades in/out, never overwhelming.
|
|
204
|
-
- **Linear.app** — subtle, purposeful micro-interactions. Every motion serves the interface.
|
|
205
|
-
- **Vercel.com** — smooth page transitions, gradient animations, scroll-triggered reveals.
|
|
206
|
-
- **Stripe.com** — gradient animations double as motion design; smooth, continuous, non-distracting.
|
|
207
|
-
|
|
208
|
-
### Implications for the design research agent
|
|
209
|
-
|
|
210
|
-
Animation guidance is mostly about restraint — "be purposeful, not decorative." The agent should recommend specific animation patterns appropriate to the app's complexity (a simple CRUD app needs entrance animations and hover states, not scroll-driven narratives). When recommending layouts that involve motion (hero sections, staggered card reveals), the agent should specify the technique (CSS scroll-timeline, Motion spring, etc.) so Remy can implement correctly.
|
|
211
|
-
|
|
212
|
-
The existing design.md animation section in the main Remy prompt covers the "what not to do" well. The design research agent adds the "what to do" with specific modern techniques.
|
|
213
|
-
|
|
214
|
-
## Available APIs (no auth required)
|
|
215
|
-
|
|
216
|
-
| API | URL | What it does |
|
|
217
|
-
|---|---|---|
|
|
218
|
-
| Fontshare | `api.fontshare.com/v2/fonts` | Curated font catalog with metadata |
|
|
219
|
-
| The Color API | `thecolorapi.com` | Color schemes from seed color (analogic, complement, triad, etc.) |
|
|
220
|
-
| Colormind | `colormind.io/api/` | AI-generated 5-color UI palettes, lock known colors |
|
|
221
|
-
| theSVG.org | `thesvg.org/api/` | 4k+ brand logos with hex brand colors, 7 variants |
|
|
222
|
-
| SVGL | `api.svgl.app` | 300+ brand logos with light/dark variants |
|
|
223
|
-
| Iconify | `api.iconify.design` | 275k+ icons from Tabler, Lucide, Phosphor, etc. |
|
|
224
|
-
|
|
225
|
-
Note: Icons and brand SVG APIs are documented here for future reference but not currently wired up as tools.
|
|
226
|
-
|
|
227
|
-
## Available via MindStudio SDK (no separate API keys needed)
|
|
228
|
-
|
|
229
|
-
| Action | What it does |
|
|
230
|
-
|---|---|
|
|
231
|
-
| Pexels search | Stock photos with avg_color field |
|
|
232
|
-
| Unsplash search | High-quality stock photos with color metadata |
|
|
233
|
-
| Pixabay search | 5.6M+ images/videos/vectors |
|
|
234
|
-
| Google search | General web search for design research |
|
|
235
|
-
| Image analysis | Vision model for analyzing screenshots |
|
|
236
|
-
| Image generation | For wireframe/layout concept generation |
|
|
237
|
-
|
|
238
|
-
Google Images is explicitly excluded — too much noise/junk for design work.
|
|
239
|
-
|
|
240
|
-
## What's done
|
|
241
|
-
|
|
242
|
-
- **Font catalog** — 105 fonts (80 Fontshare + 14 Google Fonts + 11 Open Foundry) with 51 curated pairings, compiled in `data/fonts.json`. Runtime-sampled per invocation.
|
|
243
|
-
- **Design inspiration** — Godly screenshots rehosted on MindStudio CDN, pre-analyzed via vision model, compiled in `data/inspiration.json`. Runtime-sampled per invocation. Compilation script at `data/compile-inspiration.sh`.
|
|
244
|
-
- **Runtime sampling** — `prompt.ts` samples 15 fonts + 5 pairings + 5 inspiration images per invocation.
|
|
245
|
-
- **Tools** — searchGoogle, fetchUrl, analyzeReferenceImageOrUrl, screenshot, generateImages.
|
|
246
|
-
- **Prompt** — split into files in `prompts/` (identity, color, animation, layout, icons, images, resources, instructions, frontend-design-notes), assembled via template includes.
|
|
247
|
-
- **Spec context** — automatically injected via `loadSpecContext()` from `subagents/common/context.ts`. Agent sees the full project spec without Remy summarizing.
|
|
248
|
-
- **Image generation** — Seedream 4.5 via `generateImages` tool. Prompt guidance emphasizes: style/medium first, then subject; avoid hex codes (rendered as text); generate visual ingredients not UI components; default to real subjects over abstract.
|
|
249
|
-
- **Consolidated team guidance** — "when to use the design expert" guidance lives in `static/team.md`, not in the tool description. Tool description stays concise "what I do."
|
|
250
|
-
|
|
251
|
-
## Changes from initial design
|
|
252
|
-
|
|
253
|
-
- Removed stock photo search (Pexels) and image editing tools — AI generation produces better bespoke results.
|
|
254
|
-
- Removed Google Images for design inspiration — too much noise.
|
|
255
|
-
- Consolidated `analyzeImage`, `analyzeDesignReference`, and `screenshotAndAnalyze` into single `analyzeReferenceImageOrUrl` tool that auto-detects image URLs vs website URLs.
|
|
256
|
-
- Added `screenshot` tool (external, sandbox-resolved) so the agent can capture the app preview directly.
|
|
257
|
-
- Font examples in all Remy prompts changed from Google Fonts (DM Sans) to Fontshare (Satoshi) to avoid reinforcing AI defaults.
|
|
258
|
-
- Media CDN guidance changed from "Google Fonts via CDN" to "load fonts from CDNs" to avoid biasing toward Google Fonts.
|
|
259
|
-
|
|
260
|
-
## What's not done
|
|
261
|
-
|
|
262
|
-
- **Wireframe generation** — need to figure out good prompts for `generateImage` that produce useful layout wireframe concepts.
|
|
263
|
-
- **More inspiration sources** — currently only Godly. Could add Awwwards, Figma Community, or other galleries to the inspiration pool.
|
|
264
|
-
- **Open Foundry fonts without Google Fonts hosting** — 11 fonts marked "(self-host required)" have no CSS URL. Would need to host font files on CDN to make them usable.
|
|
265
|
-
- **Icon integration** — Iconify API and theSVG.org are documented but not wired up as tools yet. Deferred.
|
|
@@ -1,79 +0,0 @@
|
|
|
1
|
-
# Product Vision Sub-Agent — Design Notes & Decisions
|
|
2
|
-
|
|
3
|
-
Notes from the initial build (March 2026).
|
|
4
|
-
|
|
5
|
-
## Purpose
|
|
6
|
-
|
|
7
|
-
Generates ambitious, creative roadmap items at the end of spec authoring. Called once per project, reads the completed spec files, and writes 10-15 roadmap items directly to `src/roadmap/`. The main agent passes a brief task description; the sub-agent reads the full spec context from disk automatically.
|
|
8
|
-
|
|
9
|
-
## Why a separate sub-agent?
|
|
10
|
-
|
|
11
|
-
The main agent (Remy) is optimized for execution: understanding requests, writing specs, generating code, verifying results. That execution mindset makes it too conservative when imagining future features. It anchors on what the user asked for and proposes incremental add-ons ("dark mode", "better onboarding") rather than transformative ideas ("build the actual product the landing page is selling").
|
|
12
|
-
|
|
13
|
-
A dedicated sub-agent with a different personality can think freely without the constraints of the builder's pragmatism. It never interacts with the user directly and never sees the user's original messages — it only sees the spec output. This separation is intentional.
|
|
14
|
-
|
|
15
|
-
## Why third-person prompt framing
|
|
16
|
-
|
|
17
|
-
The system prompt uses third-person construction ("The role of the assistant is to act as...") instead of the usual second-person ("You are..."). This is a deliberate choice to work around RLHF training patterns.
|
|
18
|
-
|
|
19
|
-
Modern LLMs are fine-tuned to be responsive to user requests — stay on topic, don't overstep, respect stated scope. That's exactly the opposite of what we want here. We need the agent to say "actually, think bigger" and propose things the user didn't ask for.
|
|
20
|
-
|
|
21
|
-
The third-person framing creates a performance/role-play context rather than an identity context. The model is playing a character (a product visionary), and characters can do things the model's base personality wouldn't. It's the same mechanism that makes models more creative when writing dialogue for a fictional character than when answering directly.
|
|
22
|
-
|
|
23
|
-
Combined with three layers of separation from the user's original messages (sub-agent sees spec files, not user chat; gets a brief task from Remy, not the user's words; operates in RP mode), this makes the agent much more willing to go beyond the stated scope.
|
|
24
|
-
|
|
25
|
-
## Why it reads spec files from disk
|
|
26
|
-
|
|
27
|
-
The product vision tool injects all spec files from `src/` (excluding `src/roadmap/`) into the system prompt as XML-tagged file contents. This means:
|
|
28
|
-
|
|
29
|
-
1. Remy doesn't waste tokens summarizing the spec in the task description — it just passes a brief "this is a dating app for Gen Z" and the sub-agent has the full context.
|
|
30
|
-
2. The sub-agent sees the complete spec as written, not Remy's interpretation of it.
|
|
31
|
-
3. The sub-agent is further distanced from the user's original framing, since it reads the spec output rather than the conversation.
|
|
32
|
-
|
|
33
|
-
## Why it has a tool (writeRoadmapItem)
|
|
34
|
-
|
|
35
|
-
Originally this was a tool-less sub-agent that returned YAML for Remy to parse and write into files. This was changed to give the sub-agent a `writeRoadmapItem` tool that writes directly to `src/roadmap/`. Benefits:
|
|
36
|
-
|
|
37
|
-
- The sub-agent writes files in parallel (batching all tool calls in one turn)
|
|
38
|
-
- No parsing/serialization step for Remy
|
|
39
|
-
- The MVP item gets `status: in-progress` automatically (hardcoded in the executor for slug "mvp")
|
|
40
|
-
- Each file is a proper MSFM document with frontmatter
|
|
41
|
-
|
|
42
|
-
## Key prompt engineering decisions
|
|
43
|
-
|
|
44
|
-
**"Dream big" framing.** The prompt explicitly says: safe/boring roadmap is worse than no roadmap. At least 3 items must be large effort. At least 2 lanes must extend beyond the current product scope. The self-check asks "would a user be excited showing this to a friend?"
|
|
45
|
-
|
|
46
|
-
**"Lanes, not lists" structure.** Instead of a flat list of features, the prompt asks for 3-5 distinct growth directions (lanes), each with depth and dependencies. Like a skill tree in a game. This produces coherent product narratives rather than grab bags of ideas.
|
|
47
|
-
|
|
48
|
-
**User-facing language.** Names and descriptions must be written for non-developers. No library names, no technical jargon. "Interactive Personality Quiz" not "Multi-step React form with state machine."
|
|
49
|
-
|
|
50
|
-
**Structured body format.** Each roadmap item body follows: elevator pitch → "What it looks like" → "Key details" → technical annotation. This prevents the narrative rambling that the agent initially produced and makes each item a scannable mini-spec.
|
|
51
|
-
|
|
52
|
-
**Cap at 15 items.** Without this constraint the agent gets carried away (which is the right problem to have — better than being too conservative). The cap forces quality and depth over quantity.
|
|
53
|
-
|
|
54
|
-
## What worked well
|
|
55
|
-
|
|
56
|
-
The third-person framing + spec-from-disk + separation from user messages produced dramatically better results than having Remy generate roadmap items directly. The agent went from "dark mode toggle, waitlist email drip" to "build the actual dating app, AI conversation coach, match intelligence engine" in one change.
|
|
57
|
-
|
|
58
|
-
## Generalization to roadmap owner (March 2026)
|
|
59
|
-
|
|
60
|
-
The agent was initially a one-shot idea generator. It has since been generalized to own the entire roadmap lifecycle:
|
|
61
|
-
|
|
62
|
-
- **Three tools:** `writeRoadmapItem` (create), `updateRoadmapItem` (modify status, append history, change fields), `deleteRoadmapItem` (remove).
|
|
63
|
-
- **Roadmap context injected:** The system prompt now includes both `<spec_files>` and `<current_roadmap>` so the agent sees the full picture before making changes.
|
|
64
|
-
- **Multiple operations:** Seeding initial ideas, marking items done, adding features from user requests, removing irrelevant items, reorganizing after builds, answering strategic product questions.
|
|
65
|
-
- **Not always writing files:** The prompt explicitly says "not every task requires tool calls" — sometimes it just answers a question about product direction.
|
|
66
|
-
|
|
67
|
-
File structure split into: `index.ts` (tool definition), `prompt.ts` (prompt assembly with context injection), `prompt.md` (personality/rules), `tools.ts` (tool definitions), `executor.ts` (filesystem operations).
|
|
68
|
-
|
|
69
|
-
Context loaders (`loadSpecContext`, `loadRoadmapContext`) extracted to `subagents/common/context.ts` for sharing with other sub-agents.
|
|
70
|
-
|
|
71
|
-
### Ego-conscious language in team.md
|
|
72
|
-
|
|
73
|
-
The guidance in `team.md` about the product vision agent uses "we/us" language ("a schema decision that'll paint us into a corner") rather than "you" language. This avoids triggering ego-defensiveness in the main agent — it reads as collaborative rather than corrective. Same principle applied to the code sanity check agent.
|
|
74
|
-
|
|
75
|
-
## What to watch for
|
|
76
|
-
|
|
77
|
-
- The agent may still occasionally produce developer-facing language in item names. The prompt guards against this but it's a tendency to monitor.
|
|
78
|
-
- Very simple projects may get roadmap items that feel like a stretch. That's by design — the agent is supposed to dream big — but some ideas may not land well with all users.
|
|
79
|
-
- The spec + roadmap file injection adds to the system prompt size. For very large specs this could be significant. Currently not a problem since specs are typically a few KB total.
|