@mindstudio-ai/remy 0.1.255 → 0.1.257

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -85,9 +85,23 @@ const { key, url } = await platform.upload(token, file, { onProgress: (f) => set
85
85
 
86
86
  ## Public assets + image resizing
87
87
 
88
- Public files are world-readable, on the app's domain, and **images resize via query params**
89
- (`?w=&h=&fit=&crop=&fm=&dpr=&q=&blur=&sharpen=` — same vocabulary as the image CDN; set `dpr=2/3` for
90
- retina). Request the size you need rather than CSS-scaling a full-res original.
88
+ Public files are world-readable, served on the app's own domain, and **images resize via query
89
+ params** — request the size you need rather than CSS-scaling a full-res original. Always set `dpr=2`
90
+ or `3` when sizing so images stay crisp on Retina displays.
91
+
92
+ | Param | Example | Effect |
93
+ |-------|---------|--------|
94
+ | `w` | `?w=400` | Max width in pixels |
95
+ | `h` | `?h=300` | Max height in pixels |
96
+ | `fit` | `?fit=crop` | Resize mode: `scale-down`, `contain`, `cover`, `crop`, `pad` |
97
+ | `crop` | `?crop=face` | Face-aware crop (with `fit=crop`) |
98
+ | `fm` | `?fm=webp` | Output format: `avif`, `webp`, `jpeg`, `auto` |
99
+ | `dpr` | `?dpr=2` | Device pixel ratio |
100
+ | `q` | `?q=80` | Quality (1–100) |
101
+ | `blur` | `?blur=10` | Blur radius |
102
+ | `sharpen` | `?sharpen=1` | Sharpen amount |
103
+
104
+ Combine freely: `…/hero.jpg?w=200&h=200&fit=crop&fm=avif`.
91
105
 
92
106
  Lightweight-config pattern: a public store + a stable `key` is a file the frontend can `fetch` with no
93
107
  DB hit and the backend can overwrite (`Config.put(json, { key: 'config/latest.json' })`).
@@ -59,19 +59,16 @@ const api = createClient<{
59
59
  const { vendorId } = await api.submitVendorRequest({ name: 'Acme' });
60
60
  const { vendors } = await api.listVendors();
61
61
 
62
- // File upload (returns CDN URL)
63
- const url = await platform.uploadFile(file);
62
+ // File upload → client-direct to the app's file store (see Files & Storage).
63
+ // A backend method mints an upload token; the browser uploads straight to storage.
64
+ const token = await api.getUploadSlot({ filename: file.name, contentType: file.type });
65
+ const { key, url } = await platform.upload(token, file);
64
66
 
65
- // With progress tracking
66
- const url = await platform.uploadFile(file, {
67
- onProgress: (fraction) => setProgress(fraction), // 0 to 1
68
- });
69
-
70
- // With abort support
67
+ // With progress + abort
71
68
  const controller = new AbortController();
72
- const url = await platform.uploadFile(file, {
69
+ const { url } = await platform.upload(token, file, {
73
70
  signal: controller.signal,
74
- onProgress: (f) => setProgress(f),
71
+ onProgress: (fraction) => setProgress(fraction), // 0 to 1
75
72
  });
76
73
  controller.abort(); // cancels the upload
77
74
 
@@ -419,7 +416,7 @@ It supports the full MCP surface:
419
416
  - **Prompts** — reusable, parameterized prompt templates the server offers.
420
417
  - **Instructions** — server-level guidance shown to the calling agent (the toolset's "system prompt").
421
418
 
422
- The platform hosts the server, handles auth like the API interface (optional — keyed or anonymous), and derives every tool's input schema from the method contract. Because the consumer is an external agent with no knowledge of your app, **the descriptions are the product** — see "Building MCP Interfaces" for how to write them.
419
+ The platform hosts the server, handles auth like the API interface (optional — keyed or anonymous), and derives every tool's input schema from the method contract. Because the consumer is an external agent with no knowledge of your app, **the descriptions are the product** — load the `mcpInterfaces` skill for how to write them.
423
420
 
424
421
  ### Spec: `src/interfaces/mcp.md`
425
422
 
@@ -566,6 +563,8 @@ There is no `inputSchema` field — the platform derives each tool's schema from
566
563
 
567
564
  A conversational interface where an LLM has access to the app's methods as tools. Unlike MCP (which exposes methods for external agents), the agent interface IS the agent — it has its own personality, system prompt, and model config, and orchestrates tool calls against the app's methods internally.
568
565
 
566
+ This section is the wiring. Load the `agentInterfaces` skill before authoring the spec body or building the chat UI — what belongs in the agent's system prompt, how to write its tool descriptions, and the `createAgentChatClient()` frontend API are all there.
567
+
569
568
  ### Spec: `src/interfaces/agent.md`
570
569
 
571
570
  The human-readable spec. Frontmatter contains structured fields; the prose body is the behavioral spec — voice, personality, capabilities, rules — written in MSFM.
@@ -97,11 +97,8 @@ await mindstudio.sendEmail({
97
97
  cc: input.cc, // reply-all
98
98
  });
99
99
 
100
- // Upload files
101
- const { url } = await mindstudio.uploadFile({
102
- data: buffer,
103
- fileName: 'report.pdf',
104
- });
100
+ // Store a file → returns a stable URL (define the store at module scope; see Files & Storage)
101
+ const { url } = await Reports.put(buffer, { contentType: 'application/pdf', filename: 'report.pdf' });
105
102
 
106
103
  // Web scraping
107
104
  const { markdown } = await mindstudio.scrapeUrl({
@@ -221,6 +218,8 @@ export async function createPurchaseOrder(input: {
221
218
 
222
219
  A method can return immediately while kicking off slow work (like `runTask()`) that continues in the background. Don't await the slow call — use `.then()` / `.catch()` to update the record when it completes, and return an early result to the caller. The frontend polls the record's status to track progress.
223
220
 
221
+ The example below shows the fire-and-forget shape, not a complete `runTask()` call. Load the `taskAgents` skill before writing one — configuring its tools, validating the output, and handling failures are all there, and none of them are visible here.
222
+
224
223
  ```typescript
225
224
  export async function enrichRestaurant(input: { id: string; name: string }) {
226
225
  await Restaurants.update(input.id, { status: 'enriching' });
@@ -104,6 +104,22 @@ This is deterministic — same scenario always produces the same state.
104
104
 
105
105
  Scenarios are useful for seeding initial app state after build for testing, as well as to give the user a first impression of an app that is already filled with data and looks and feels usable. The user can choose to run further scenarios after initial build by clicking the Scenarios tab and selecting a scenario to run.
106
106
 
107
+ ## What scenarios don't touch
108
+
109
+ **Scenarios seed database tables and nothing else.** They do not touch file stores or data sources —
110
+ deliberately: both are durable and shared across dev and prod, with no per-release copy to reset, so
111
+ there is nothing to truncate.
112
+
113
+ Don't try to seed documents into a data source from a scenario, and don't write `clear()`-style reset
114
+ helpers for one. Load a test corpus once from the CLI instead:
115
+
116
+ ```bash
117
+ mindstudio-prod datasources add --source policies --wait fixtures/*.pdf
118
+ ```
119
+
120
+ Re-running it is free (documents are content-addressed), so it's safe to keep in a setup script
121
+ beside your scenarios.
122
+
107
123
  ## Scenario Data
108
124
 
109
125
  Align scenario data to the vibe of the app - construct data that feels like it fits.
@@ -83,7 +83,6 @@ result.$billingCost; // cost in credits (if applicable)
83
83
 
84
84
  | Action | What it does |
85
85
  |--------|-------------|
86
- | `uploadFile` | Upload a file to CDN |
87
86
  | `downloadVideo` | Download a video URL |
88
87
  | `getMediaMetadata` | Get dimensions, duration, etc. |
89
88
  | `convertPdfToImages` | PDF pages to PNG images |
@@ -154,4 +153,4 @@ Consider the ways in which AI can be incorporated into backend methods to solve
154
153
 
155
154
  ### Task Agents
156
155
 
157
- For multi-step tasks where the model needs to autonomously compose actions (research + scrape + generate, enrichment pipelines, content creation), use `runTask()` instead of chaining actions manually. It runs an agent loop and returns structured JSON. Its tools can include SDK actions as well as your app's own methods, so the agent can read your data to decide what to do next and write results back itself. See the task agents reference for full details.
156
+ For multi-step tasks where the model needs to autonomously compose actions (research + scrape + generate, enrichment pipelines, content creation), use `runTask()` instead of chaining actions manually. It runs an agent loop and returns structured JSON. Its tools can include SDK actions as well as your app's own methods, so the agent can read your data to decide what to do next and write results back itself. Load the `taskAgents` skill before writing one — it is the full reference.
@@ -1,3 +1,9 @@
1
+ ---
2
+ name: Agent Interfaces
3
+ what: Conversational AI as a first-class interface to the app — an LLM with authenticated, per-user access to the app's methods as tools, paired with a streaming chat UI. The platform handles auth, tool dispatch, threads, and streaming, so the work is authorship: who the agent is, which methods it can reach, and how each one is described to it. Any app whose methods do something interesting can be projected into a conversation this way, often as its most compelling surface. This reference covers both halves — writing the agent spec, and building the chat frontend with the SDK's `createAgentChatClient()`.
4
+ when: Before authoring `src/interfaces/agent.md`, compiling `dist/interfaces/agent/`, or building an agent's chat UI. The `<interfaces>` platform doc has the wiring; this is how to author one well.
5
+ ---
6
+
1
7
  # Building Agent Interfaces
2
8
 
3
9
  Guidance for designing conversational AI agents and their frontends. An agent interface pairs an LLM (with per-user-scoped/authenticated access to app methods as tools, handled by platform automatically) with a chat UI. The developer authors the agent's character in MSFM (`src/interfaces/agent.md`); you compile it into a system prompt and tool descriptions (`dist/interfaces/agent/`).
@@ -10,21 +16,32 @@ A good system prompt establishes who the agent is — personality, tone, judgmen
10
16
 
11
17
  Short and opinionated beats long and comprehensive. "Sounds like a sharp, organized friend — brief by default" gives the model more to work with than a page of behavioral rules. Define constraints through character, not checklists. Let the model's judgment work.
12
18
 
13
- #### System Prompt Specifics
14
- Always include a note like "## Tool Usage
15
- - When multiple tool calls are independent, make them all in a single turn. Searching for three different products, or fetching two reference sites: batch them instead of doing one per turn." to help the model know it can run tools in parallel
16
- - The user's name and current role(s) at the time of message, if any, will be automatically appended to the end of every system prompt at runtime like:
19
+ Three things every compiled system prompt should carry, on top of the character:
20
+
21
+ **Parallel tool use.** The model won't batch independent calls unless told it can. Include a section like:
17
22
 
23
+ ```markdown
24
+ ## Tool Usage
25
+
26
+ When multiple tool calls are independent, make them all in a single turn.
27
+ Searching for three different products, or fetching two reference sites:
28
+ batch them instead of doing one per turn.
18
29
  ```
30
+
31
+ **Markdown and house style.** Unless the user says otherwise, tell the agent it can use markdown (the chat UI renders it) and to avoid em dashes and emojis.
32
+
33
+ **The current user is appended for you.** At runtime the platform appends the user's name and roles to the end of every system prompt, so don't write your own placeholder for it:
34
+
35
+ ```markdown
19
36
  ## Current User
37
+
20
38
  Name: Jane Smith
21
39
  Roles: editor
22
40
  ```
23
- - Unless the user specifies otherwise, always include a note that the agent can use markdown in responses (since the chat UI renders it) and should avoid using em dashes and emojis in its responses.
24
41
 
25
42
  ### Tool descriptions are the most important artifact
26
43
 
27
- The system prompt says *who* the agent is. The tool descriptions say *what it can do*. A great tool description means the agent uses the tool correctly without explicit instruction. Do not be overly precise or micromanage. Your goal with tool descriptions is to provide context and faming- trust that the model is intelligent enough to fill in the gaps.. Each `tools/*.md` file should cover:
44
+ The system prompt says *who* the agent is. The tool descriptions say *what it can do*. A great tool description means the agent uses the tool correctly without explicit instruction. Do not be overly precise or micromanage. Your goal with tool descriptions is to provide context and framing — trust that the model is intelligent enough to fill in the gaps. Each `tools/*.md` file should cover:
28
45
 
29
46
  - **When to use** this tool (and when NOT to — e.g. "NOT for marking complete, use toggle-todo")
30
47
  - **Parameter guidance** beyond the schema — what makes a good value, when to include optional fields, what to skip
@@ -116,10 +133,12 @@ response.abort();
116
133
 
117
134
  **Attachments:**
118
135
 
119
- Send images or documents alongside a message. Upload via `platform.uploadFile()` first, then pass CDN URLs as the 4th argument:
136
+ Send images or documents alongside a message. Upload to the app's file store first (see Files & Storage), then pass the returned URLs as the 4th argument:
120
137
 
121
138
  ```ts
122
- const url = await platform.uploadFile(file);
139
+ // backend method mints a token; the browser uploads straight to storage
140
+ const token = await api.getUploadSlot({ filename: file.name, contentType: file.type });
141
+ const { url } = await platform.upload(token, file);
123
142
 
124
143
  chat.sendMessage(threadId, "What's in this document?", {
125
144
  onText: (delta) => setText((prev) => prev + delta),
@@ -128,7 +147,7 @@ chat.sendMessage(threadId, "What's in this document?", {
128
147
  });
129
148
  ```
130
149
 
131
- Images (`i.mscdn.ai`) are sent as vision input. Documents (`f.mscdn.ai`) have text extracted server-side and included in context. Attachments are preserved in thread history.
150
+ Images are sent as vision input; documents have their text extracted server-side and included in context. Attachments are preserved in thread history.
132
151
 
133
152
  **Key points:**
134
153
  - `onText` and `onThinking` receive deltas (append to state, don't replace)
@@ -174,7 +193,7 @@ The first screen should invite conversation. A greeting from the agent, a few su
174
193
 
175
194
  ### Mobile
176
195
 
177
- Chat is inherently mobile-friendly — lean into it. Pay attention to viewport sizing on mobile as the virtual keyboard changes the available height.
196
+ Chat is inherently mobile-friendly — lean into it. Pay attention to viewport sizing on mobile as the virtual keyboard changes the available height.
178
197
 
179
198
  ### Respect the brand
180
199
 
@@ -0,0 +1,131 @@
1
+ ---
2
+ name: Data Sources
3
+ what: A managed retrieval system for document corpora, not a bolt-on keyword search. Documents are chunked and embedded, candidate hits are re-scored by a reranking model, semantic search runs alongside exact keyword matching so part numbers and error codes still land, images inside documents are described by a vision model and made searchable, and every hit returns a citation that links to its source page. Chunking and embedding settings are versioned — a rebuilt corpus can be compared against the live one and promoted without downtime. All of that applies to unstructured documents queried by meaning, and to nothing else.
4
+ when: Only for unstructured documents queried by meaning — "find the clause about early termination". Structured data belongs in `db`: if the question can be expressed as a filter, it is not a search problem, and a `WHERE` clause is faster, cheaper and exact. Load before defining or querying a data source.
5
+ ---
6
+
7
+ # Data Sources (Search Over Documents)
8
+
9
+ Per-app searchable document corpora: upload documents, ask in plain language, get back the passages
10
+ that answer it with a citation to the source.
11
+
12
+ **Most apps should not use one — check this before reaching for it.** Structured data (rows with
13
+ fields you filter on) belongs in `db`; a `WHERE` clause is faster, cheaper and exact. A data source
14
+ earns its cost only for **unstructured documents queried by meaning**. "Find the clause about early
15
+ termination" is a data source. "Find contracts signed after March" is a `db` query. If the question
16
+ can be expressed as a filter, it is not a search problem.
17
+
18
+ Don't use one when: the data is structured (`db`), you only need to store files (`files` — nobody is
19
+ searching the contents), the requirement is exact lookup by identifier, or the corpus is a handful of
20
+ short docs that fit in a prompt.
21
+
22
+ ## Behaviour (read before the API)
23
+
24
+ - **One corpus shared across dev and prod** — like a file store, not a table. No dev copy, no
25
+ per-release isolation. A document added while building is already live.
26
+ - **Scenarios never reset a data source.** Don't write `clear()`-style reset helpers.
27
+ - **Re-adding the same bytes is free** — content-addressed, so ingest scripts are safe to re-run.
28
+ - **Ingest is async.** `add()` returns once queued; poll `documents()`, or use `--wait` from the CLI.
29
+ - **Reprocessing costs real money**, so changing how a corpus is built is always explicit.
30
+ - **Limits apply**: 25 data sources per app, 5,000 documents per source, 10,000 chunks per
31
+ document, 300 searches/minute. Well clear of normal use — but **source names must be fixed, not
32
+ computed per user or per request**, since referencing one creates it. Partition inside a source
33
+ with a metadata filter instead.
34
+
35
+ ## Defining and searching
36
+
37
+ ```typescript
38
+ import { dataSources } from '@mindstudio-ai/agent';
39
+ export const Policies = dataSources.defineDataSource('policies'); // lowercase [a-z0-9_-], ≤64
40
+
41
+ const { results } = await Policies.search('what are the payment terms?', { topK: 5 });
42
+ const context = results.map((r) => r.text).join('\n\n');
43
+ ```
44
+
45
+ Hits are `{ score, text, citation }` with
46
+ `citation: { documentId, filename, pageNumber, chunkIndex, headingPath, boundingBox?, url }`, plus
47
+ `retrievalRank`/`retrievalScore` — the position before reranking, so you can show what reranking did.
48
+
49
+ **Always render the citation.** `citation.url` is a stable on-domain link — put it in an `<a href>`
50
+ beside the answer. Retrieval is approximate; a user who can click through can judge for themselves.
51
+ An answer with no citation is an assertion.
52
+
53
+ Created on first use, so searching a source the build hasn't populated returns no results rather than
54
+ throwing. `search` options: `topK` (default 5, max 50), `scoreThreshold`, `rerank`, `hybrid`.
55
+
56
+ Search is deterministic for a fixed corpus and configuration, so eval sets and regression checks are
57
+ meaningful — key them on `(documentId, chunkIndex)` rather than on chunk text.
58
+
59
+ **Debugging retrieval.** Two opt-in options, neither of which changes the results or their order:
60
+ `explain: true` adds `explain.{dense, lexical, matchedVia}` (which half of hybrid found each hit;
61
+ costs two extra round trips), and `expand: 1` adds `neighbors.{before, after}` for surrounding
62
+ context. When a document never comes back at all, `Policies.stats()` reports the config actually in
63
+ effect and `Policies.chunks(documentId)` shows exactly how it was split.
64
+
65
+ **Configuration is not declared in code** — chunking and embedding settings live on the corpus and are
66
+ set with the CLI, so code and reality can't drift.
67
+
68
+ ## Loading documents — normally at build time, from the CLI
69
+
70
+ ```bash
71
+ mindstudio-prod datasources add --source policies --wait docs/*.pdf
72
+ mindstudio-prod datasources search --source policies "what are the payment terms?" # sanity-check
73
+ mindstudio-prod datasources delete --source policies # whole source; --source is required, never defaulted
74
+ ```
75
+
76
+ `--wait` blocks until processing finishes and exits non-zero on failure. Also `datasources list`,
77
+ `status` (per-document state + ingest errors), `rm --document <id>`. `--help` for flags.
78
+
79
+ **Seeding a test corpus:** scenarios don't touch data sources, so load fixtures with the same command
80
+ in a setup script — `datasources add --source <slug> --wait fixtures/*.pdf`. Re-running is free, so
81
+ it needs no guard.
82
+
83
+ Use the SDK's `add()` only when *users* upload documents that must become searchable:
84
+
85
+ ```typescript
86
+ await Policies.add(buffer, { filename: 'policy.pdf', contentType: 'application/pdf' });
87
+ const docs = await Policies.documents(); // 'processing' | 'done' | 'error'
88
+ await Policies.remove(documentId);
89
+ ```
90
+
91
+ Formats: pdf, docx, pptx, xlsx, odt, rtf, epub, images, txt, md, json, csv, tsv, log, html.
92
+
93
+ ## Answering from results
94
+
95
+ Retrieve → join passages as context → have a model answer *from that context* → render citations.
96
+ Never paste raw chunks at the user; they're fragments. For agentic flows, give the model `search` as a
97
+ tool so it can query repeatedly and refine, rather than retrieving once up front.
98
+
99
+ ## Tuning — two kinds of setting
100
+
101
+ | Kind | Settings | Cost |
102
+ |---|---|---|
103
+ | **Free** (ranking) | `--rerank`, `--rerank-model`, `--hybrid`, `--top-k` | none, next search |
104
+ | **Rebuild** (how docs become vectors) | `--max-chars`, `--min-chars`, `--drop-blocks`, `--contextual`, `--describe-images`, `--embedding-model`, `--extraction-model` | every document reprocessed |
105
+
106
+ Images inside documents are described by a vision model and the description substituted into the
107
+ searchable text (`--describe-images`, on by default) — without it a chart contributes nothing to
108
+ search at all. Documents with no images cost nothing.
109
+
110
+ `rerank` and `hybrid` default on and are usually right — reranking is the biggest quality lever, and
111
+ hybrid is what finds part numbers, error codes and proper nouns a semantic model never learned. Both
112
+ are also per-query (`search(q, { rerank: false })`) for a latency-sensitive path.
113
+
114
+ ```bash
115
+ mindstudio-prod datasources config --source policies # show
116
+ mindstudio-prod datasources config --source policies --top-k 8 # free, immediate
117
+ ```
118
+
119
+ **A rebuild-class change on a populated corpus is rejected** — you're told what it would invalidate
120
+ and what it costs. To make it, build a new version alongside the live one:
121
+
122
+ ```bash
123
+ mindstudio-prod datasources revectorize --source policies --max-chars 900 --wait
124
+ mindstudio-prod datasources search --source policies --candidate "payment terms" # compare
125
+ mindstudio-prod datasources promote --source policies # go live
126
+ ```
127
+
128
+ Search serves the current version throughout, so nothing degrades while the new one builds.
129
+ `datasources drop` discards an unwanted candidate.
130
+
131
+ For anything deeper on the SDK, ask `askMindStudioSdk` rather than guessing at an API.
@@ -0,0 +1,80 @@
1
+ ---
2
+ name: MCP Interfaces
3
+ what: Ships the app as an MCP server, so external AI agents — Claude Desktop, Cursor, anyone's agent — can drive it as a tool surface. The platform hosts the server, handles auth, and derives every tool's input schema from the method contract, so there is no protocol code to write: the work is choosing which methods an outsider should see and describing them well enough for a stranger to use correctly. Cheap to add to an app that already has methods, and it puts the app inside the tools its users already work in.
4
+ when: Before authoring `src/interfaces/mcp.md` or deciding which of the app's methods an external agent gets to see.
5
+ ---
6
+
7
+ # Building MCP Interfaces
8
+
9
+ Guidance for exposing an app as an MCP server — a tool / resource / prompt surface for *external* AI agents (Claude Desktop, Cursor, anyone's agent). The contract (spec format, compiled output, `interface.json`) is in the `<interfaces>` platform doc in your system prompt; this is how to author one well. Unlike the agent interface, there's no LLM, personality, or UI to design — the entire product is the descriptions and the shape of what you expose.
10
+
11
+ ## The descriptions are the product
12
+
13
+ The calling agent is a stranger with no knowledge of your app. It decides what to invoke entirely from the names, descriptions, and annotations you ship. Follow the same principles as the agent interface's tool descriptions (load the `agentInterfaces` skill for those — when to use and when not, parameter guidance beyond the schema, what the tool returns) — but write them **self-contained**. An in-app agent tool can lean on the app's framing; an MCP tool can't, because the caller has no context. Spell out what an outsider wouldn't know.
14
+
15
+ What that looks like in practice — the same method, described twice:
16
+
17
+ ```markdown
18
+ <!-- Weak: restates the schema, assumes the caller knows the app -->
19
+ Updates a vendor. Takes a vendorId and the fields to change.
20
+ ```
21
+
22
+ ```markdown
23
+ <!-- Strong: says when to call it, what an outsider can't infer, what comes back -->
24
+ Update an existing vendor's details. Call `listVendors` or `getVendor` first —
25
+ vendorId is the app's internal id, not a name, and there is no lookup by name.
26
+ Only the fields you pass are changed; omitted fields are left alone. Returns the
27
+ full updated vendor. Editors and admins only; other roles are rejected. For a
28
+ vendor that doesn't exist yet, use `createVendor`.
29
+ ```
30
+
31
+ The weak one is what a schema already tells the caller. The strong one carries the three things a schema can't: the prerequisite, the partial-update semantics, and the alternative when this isn't the right tool.
32
+
33
+ ## Curate — not every method is a tool
34
+
35
+ Expose what an outside agent would actually use. Skip internal helpers, admin-only methods, and batch operations. A focused set of well-described tools beats a large set of thin ones. Note role restrictions in the description — gated tools are listed but reject unauthorized calls at runtime, so set expectations rather than surfacing a raw error.
36
+
37
+ ## Annotations
38
+
39
+ Annotations are machine-readable hints clients use to decide whether to auto-call a tool or ask the user first. Set them honestly:
40
+
41
+ - `readOnly` — the tool only reads, never mutates. The highest-value hint: clients auto-call reads without prompting, so set it on every pure read.
42
+ - `destructive` — the tool can delete or overwrite. Clients gate these behind confirmation.
43
+ - `idempotent` — calling twice with the same arguments has the same effect as calling once.
44
+ - `openWorld` — the tool reaches outside the app (external web/services) rather than operating only on app data.
45
+
46
+ The judgement is per tool, and getting `readOnly` right is what makes a toolset feel responsive rather than nagging:
47
+
48
+ ```jsonc
49
+ "getVendor": { "readOnly": true, "idempotent": true }
50
+ "searchVendors": { "readOnly": true, "idempotent": true }
51
+ "updateVendor": { "idempotent": true } // repeatable, but it writes
52
+ "deleteVendor": { "destructive": true, "idempotent": true }
53
+ "enrichVendorFromWeb": { "openWorld": true } // calls out to the internet
54
+ ```
55
+
56
+ Set them honestly rather than defensively. Marking a read `destructive` to be safe means the caller's user gets a confirmation prompt for looking something up, and they will stop reading the prompts.
57
+
58
+ ## Tools vs. resources
59
+
60
+ A **tool** is an action the agent *invokes*; a **resource** is data the agent *reads into context*. A read-only method can be either — expose it as a tool if the agent will call it as a step, as a resource if it's reference data the agent should pull in, and as both when both fit.
61
+
62
+ Resources are method-backed: a read invokes the method. Use a static `uri` for a fixed collection (`app://vendors`) and a `uriTemplate` when the read takes parameters (`app://vendors/{id}`, where `{id}` maps to the method's input). Keep URIs stable and human-legible.
63
+
64
+ ## Prompts
65
+
66
+ Prompts are reusable, parameterized templates the server offers to clients — e.g. a "draft a vendor email" starter. Author the template body with `{{arg}}` placeholders and declare its arguments. Offer a prompt when there's a recurring task worth packaging; skip it if a tool already covers the need.
67
+
68
+ ## Server instructions
69
+
70
+ The spec's intro prose becomes the server `instructions` — toolset-level guidance returned to the calling agent at connect time (its "system prompt"). Put *cross-cutting* guidance here: how the tools fit together, ordering or prerequisites ("read a vendor before updating it"), and norms that apply across the whole toolset. Keep per-tool specifics in the tool descriptions; instructions are for the toolset as a whole.
71
+
72
+ ```markdown
73
+ This server exposes a procurement app. Vendors are the central record and
74
+ purchase orders reference them, so a vendor generally has to exist before
75
+ anything else is useful. Ids are internal — resolve a name to an id with a
76
+ search tool before calling anything that takes one. Search results are capped
77
+ at 50; page with the returned cursor rather than broadening the query.
78
+ ```
79
+
80
+ That's four sentences doing what no individual tool description could: it explains the shape of the domain, so the calling agent's first move is a reasonable one.
@@ -1,3 +1,9 @@
1
+ ---
2
+ name: Task Agents
3
+ what: A full autonomous agent loop callable from any method. Give it a prompt, a set of tools, and an example of the output shape; the platform runs the model until it produces that shape — searching, scraping, generating images, retrying approaches that failed, and calling your app's own methods to read and write data as it goes. Tools can be any of the 1000+ SDK actions and your own methods in any combination, which is what makes it part of the app rather than a detached research bot. This is the difference between a feature that saves what the user typed and one that researches, enriches, and creates on their behalf, and it is one of the most powerful things the platform can do. Consider it whenever a feature would be dramatically more compelling if the app could do real work autonomously.
4
+ when: Before writing any `mindstudio.runTask()` call — background enrichment, research-and-generate, anything where the model decides its own next step.
5
+ ---
6
+
1
7
  # Task Agents (`mindstudio.runTask`)
2
8
 
3
9
  A user types the name of a restaurant into your app, or uploads a photo of a storefront. The API call returns early, and in the background, a task agent searches Google, finds the official website, scrapes the address, gets the official social media accounts, and generates a stylized watercolor postcard of the exterior from images it found online. The user gets back a rich, illustrated card with the canonical name, website, address, and a custom image. A few tool calls (some in parallel), fully autonomous.
@@ -6,7 +12,7 @@ A user types the name of a restaurant into your app, or uploads a photo of a sto
6
12
 
7
13
  Tools are **SDK actions** (`searchGoogle`, `generateImage`, …) and **your own app's methods** (`{ appMethod: 'saveVendor' }`), in any combination. That second half is what makes a task agent part of your app rather than a detached research bot: it can read your tables to decide what to do next, and write results back itself instead of handing them to you to persist.
8
14
 
9
- This is one of the most powerful pieces of the MindStudio SDK and can make turn apps from amazing into truly magical. Use `askMindStudioSdk` to help construct the perfect agent for a task.
15
+ This is one of the most powerful pieces of the MindStudio SDK, and it can turn an app from amazing into truly magical. Use `askMindStudioSdk` to help construct the right agent for a task — including which model to give it.
10
16
 
11
17
  ## When to Use
12
18
 
@@ -62,7 +68,7 @@ const result = await mindstudio.runTask<{
62
68
  photoUrl: 'https://cdn.mindstudio.ai/...',
63
69
  },
64
70
 
65
- model: 'claude-5-sonnet',
71
+ model: 'claude-5-sonnet', // ask askMindStudioSdk — don't copy this one blind
66
72
  maxTurns: 15,
67
73
  });
68
74
 
@@ -136,7 +142,7 @@ Keep them short and task-specific. Say when to reach for it and when not to, sin
136
142
 
137
143
  ## Voice & Tone in Prompts
138
144
 
139
- When a task agent produces user-facing text, the prompt must include a note voice and tone constraints. Make sure to specify no emojis, em dashes, and other "ai-isms" in the prompt, as well as the desired tone and voice of the output.
145
+ When a task agent produces user-facing text, the prompt must state the voice and tone it should write in. Specify the desired voice explicitly, and rule out emojis, em dashes, and other "ai-isms" — the output goes straight to the user, so nothing downstream will catch them.
140
146
 
141
147
  ## Options
142
148
 
@@ -146,14 +152,10 @@ When a task agent produces user-facing text, the prompt must include a note voic
146
152
  | `input` | Yes | — | Structured input (passed as user message) |
147
153
  | `tools` | Yes | — | SDK action names and/or `{ appMethod, description }` entries, each with optional `defaults` |
148
154
  | `structuredOutputExample` | Yes | — | Object or JSON string showing expected output shape. Use realistic example values, not placeholders like `'string'` |
149
- | `model` | Yes | — | Model ID (must support tool use) |
155
+ | `model` | Yes | — | Model ID (must support tool use). Ask `askMindStudioSdk` for the right one — MindStudio's ids don't match vendor ids, so a plausible-looking guess is usually wrong |
150
156
  | `maxTurns` | No | 20 | Max loop iterations (capped at 100) |
151
157
  | `onEvent` | No | — | SSE event callback for real-time streaming |
152
158
 
153
- ## Models
154
-
155
- Use `askMindStudioSdk` for appropriate models given the task and its complexity.
156
-
157
159
  ## Return Value
158
160
 
159
161
  ```typescript
@@ -28,7 +28,7 @@ The user can already see your tool calls, so most of your work is visible withou
28
28
  Skip the rest: narrating what you're about to do, restating what the user asked, explaining tool calls they can already see.
29
29
 
30
30
  ### User attachments
31
- When a user uploads a file (PDF, Word doc, image, etc.), it is automatically saved to `src/.user-uploads/` in the project directory. The message includes the local file path, the CDN URL, and for documents with extractable text, a `.txt` sidecar with the extracted content. Use `readFile` on the sidecar to access document contents. The CDN URL can be used directly in code and specs without any upload step. If a raw file from `src/` needs to be served by the web interface, copy it to `dist/interfaces/web/public/`. These files persist across the conversation — they survive compaction and session restarts. Do not ask the user to re-upload a document that has already been saved. Voice messages are not saved to disk — their transcripts appear inline in the message.
31
+ When a user uploads a file (PDF, Word doc, image, etc.), it is automatically saved to `src/.user-uploads/` in the project directory. The message includes the local file path, and for documents with extractable text, a `.txt` sidecar with the extracted content. Use `readFile` on the sidecar to access document contents. Pass the file path itself to tools that take an image — `screenshot`, and the design expert's `analyzeImage` / `analyzeDesign` / `editImages` — and they host the file and hand back a URL you can reuse or embed in a spec. If a raw file from `src/` needs to be served by the web interface, copy it to `dist/interfaces/web/public/`. These files persist across the conversation — they survive compaction and session restarts. Do not ask the user to re-upload a document that has already been saved. Voice messages are not saved to disk — their transcripts appear inline in the message.
32
32
 
33
33
  ### Automated messages
34
34
  You will occasionally receive automated messages prefixed with `@@automated_message@@` - these are triggered by things like background agents returning their work, or by the user clicking a button in the UI (e.g., the user might click a "Build Feature" button in the product roadmap UI, and you will receive a message detailing what they want to build). You will be able to see these messages in your chat history but the user will not see them, so acknowledge them appropriately and then perform the requested work.
@@ -43,6 +43,7 @@ Note: the snapshot concatenates inline text and strips whitespace. If you need t
43
43
  - `navigate`: Navigate to a new URL within the app. Waits for the new page to load before continuing with subsequent steps. Use this instead of evaluate with `window.location.href` when you need to navigate and then continue interacting with the new page. Steps after navigate execute on the new page automatically.
44
44
  - `evaluate`: Run arbitrary JavaScript in the page and return the result.
45
45
  - `styles`: Read computed CSS styles from page elements. Pass a `properties` array with camelCase CSS property names (e.g., `["backgroundColor", "borderRadius", "fontSize"]`). Omit `properties` for a default set covering colors, typography, spacing, borders, shadows, dimensions, and layout. Uses the same targeting as click/type (ref, text, role, label, selector). Omit the target to get styles for all elements from the last snapshot.
46
+ - `screenshotFullPage`: Take a screenshot of the whole page, top to bottom. Returns CDN url with full text analysis and dimensions. Use for overall composition or content past the fold.
46
47
  - `screenshotViewport`: Take a screenshot of the visible viewport. Returns CDN url with full text analysis and dimensions. To capture a specific section, set `scrollToSelector` (a CSS selector) — or `scrollY` (an absolute offset) — on this same step; it scrolls the target into view and captures it atomically, so you do NOT need a separate scroll step. Do not use if you can get what you need with other tools - only use when you need to visually see the viewport.
47
48
  - `setViewport`: Switch the browser between desktop and mobile rendering. Set `mode` to `"desktop"` or `"mobile"`. Mobile emulates a phone (390-wide, touch, device pixel ratio 2); desktop is the standard wide viewport. This reloads the page so media queries, responsive layouts, and `matchMedia` re-evaluate — the reload clears in-page state, so switch before you set up the state you want to inspect. The mode persists across navigations within a run. Each run starts in the app's default mode, so only use this when you need to check the other one.
48
49
 
@@ -161,7 +162,7 @@ Check a count with evaluate:
161
162
 
162
163
  ### Final Screenshot
163
164
  How you take the final screenshot depends on what the task asked for:
164
- - **Whole page** → use the standalone `screenshotFullPage` tool. It takes a full-height screenshot of the current page and returns the URL plus a full-text description.
165
+ - **Whole page** → use a `browserCommand` batch ending in a `screenshotFullPage` step. Returns the URL plus a full-text description.
165
166
  - **A specific section / viewport** → use a `browserCommand` batch ending in a `screenshotViewport` step with `scrollToSelector` set to the section (e.g. `{ "command": "screenshotViewport", "scrollToSelector": "#pricing" }`). This scrolls the section into view and captures it in one atomic step. Do this rather than a separate scroll step followed by a capture — capturing the viewport is only reliable when the scroll and the shot are in the same step.
166
167
 
167
168
  <rules>
@@ -20,7 +20,7 @@ Generated images are production assets, not mockups or concepts — they are hos
20
20
 
21
21
  ### Image editing
22
22
 
23
- Use `editImages` to transform or build on existing images. Provide one or more source image URLs and a prompt describing the desired result. The source images act as reference material — the model uses them as anchors for style, subject, or composition. Think about image editing as part of a pipeline for generating a final asset from constituent pieces.
23
+ Use `editImages` to transform or build on existing images. Provide one or more source images — a URL, or the path of an image file on disk such as a logo the user uploaded — and a prompt describing the desired result. The source images act as reference material — the model uses them as anchors for style, subject, or composition. Think about image editing as part of a pipeline for generating a final asset from constituent pieces.
24
24
 
25
25
  Good use cases for editing:
26
26
  - Incorporating a logo or brand mark into a product mockup or scene
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@mindstudio-ai/remy",
3
- "version": "0.1.255",
3
+ "version": "0.1.257",
4
4
  "description": "Remy coding agent",
5
5
  "repository": {
6
6
  "type": "git",
@@ -20,7 +20,7 @@
20
20
  },
21
21
  "scripts": {
22
22
  "build": "tsup",
23
- "dev": "tsup --watch",
23
+ "dev": "tsup --watch src",
24
24
  "typecheck": "tsc --noEmit",
25
25
  "lint:fix": "prettier --write ./src",
26
26
  "local-update": "npm run build && npm link",