@mindstudio-ai/remy 0.1.278 → 0.1.280
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +1 -0
- package/dist/automatedActions/approveInitialPlan.md +1 -3
- package/dist/automatedActions/postBuildPolish.md +1 -1
- package/dist/headless.js +259 -85
- package/dist/index.js +268 -78
- package/dist/prompt/compiled/interfaces.md +34 -0
- package/dist/prompt/compiled/jewels.md +3 -0
- package/dist/prompt/compiled/methods.md +4 -0
- package/dist/prompt/skills/jewels.md +296 -0
- package/dist/prompt/skills/taskAgents.md +1 -1
- package/dist/prompt/static/coding.md +1 -0
- package/dist/prompt/static/team.md +2 -2
- package/dist/subagents/codeSanityCheck/prompt.md +1 -1
- package/dist/subagents/designExpert/prompts/instructions.md +1 -0
- package/dist/subagents/designExpert/prompts/ui-patterns.md +7 -13
- package/package.json +1 -1
|
@@ -44,6 +44,7 @@ All fields are nested under the `"web"` key.
|
|
|
44
44
|
| `devCommand` | `string` | `"npm run dev"` | Command to start the dev server |
|
|
45
45
|
| `defaultPreviewMode` | `"desktop"` \| `"mobile"` | `"desktop"` | Default preview viewport in the editor. Set to `"mobile"` for mobile-first apps. |
|
|
46
46
|
| `prerender` | `object` | — | Opt into prerendering the listed routes/patterns for crawlers/unfurlers. See "Prerendering" below. |
|
|
47
|
+
| `mounts` | `array` | — | Serve other same-workspace apps under path prefixes of this app's hosts. See "Mounting other apps" below. |
|
|
47
48
|
|
|
48
49
|
### Frontend SDK
|
|
49
50
|
|
|
@@ -92,6 +93,24 @@ The project uses `"jsx": "react-jsx"` (automatic JSX transform) — do not `impo
|
|
|
92
93
|
|
|
93
94
|
On deploy, the platform runs `npm install && npm run build` in the web directory and hosts the output on CDN.
|
|
94
95
|
|
|
96
|
+
Two serving conventions every web interface must follow:
|
|
97
|
+
|
|
98
|
+
1. **Assets via `MS_ASSET_BASE_URL`.** The platform build sets this env var to a release-addressed CDN origin. Point the bundler's public path at it so asset URLs work on any host or mount path:
|
|
99
|
+
|
|
100
|
+
```ts
|
|
101
|
+
// vite.config.ts
|
|
102
|
+
export default defineConfig({
|
|
103
|
+
base: process.env.MS_ASSET_BASE_URL || '/',
|
|
104
|
+
// ...
|
|
105
|
+
});
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
2. **Router basename from `platform.basePath`.** The platform injects the path prefix the page is served under ('' on the app's own hosts, the mount path when another app mounts this one). The SDK prefixes its own calls automatically; app code must feed it to the router and avoid hardcoded root-absolute URLs (`<img src="/logo.png">`, raw `location.assign('/about')`):
|
|
109
|
+
|
|
110
|
+
```ts
|
|
111
|
+
createBrowserRouter(routes, { basename: platform.basePath || '/' });
|
|
112
|
+
```
|
|
113
|
+
|
|
95
114
|
#### Error Handling and Analytics
|
|
96
115
|
|
|
97
116
|
The SDK automatically reports uncaught errors, unhandled promise rejections, and pageviews to a per-app dashboard the owner gets for free. No setup required. The analytics dashboard covers visits, unique visitors, top pages, referrers, UTM breakdowns, country-level geo, device/browser/OS, new vs returning, and live online count.
|
|
@@ -136,6 +155,21 @@ await prerender.invalidate(['/u/abc']); // omit arg to purge all
|
|
|
136
155
|
|
|
137
156
|
`mindstudio-prod prerender` can help you verify/manage snapshots during development.
|
|
138
157
|
|
|
158
|
+
### Mounting other apps
|
|
159
|
+
|
|
160
|
+
A web interface can serve other apps from the same workspace under path prefixes of its own hosts (the multi-zone pattern — e.g. a marketing site serving a self-contained demo app at `/demos/vector-search` instead of linking to its subdomain):
|
|
161
|
+
|
|
162
|
+
```json
|
|
163
|
+
{ "web": { "mounts": [{ "path": "/demos/vector-search", "app": "vector-search-demo" }] } }
|
|
164
|
+
```
|
|
165
|
+
|
|
166
|
+
- `path`: non-root mount prefix; mounts must be disjoint (none may prefix another).
|
|
167
|
+
- `app`: the target's `custom_subdomain` or appId — a v2 app in the same workspace, validated at build.
|
|
168
|
+
|
|
169
|
+
Under the mount everything is the child's, served first-class (its live bundle, session, backend, auth, telemetry, frame policy). The child needs no mount-specific config and keeps working at its own subdomain; deploys stay independent. Mount-safety = the two serving conventions above (`MS_ASSET_BASE_URL` assets + `platform.basePath` basename, no hardcoded root-absolute URLs); the parent's build log warns when a target's build isn't mount-safe.
|
|
170
|
+
|
|
171
|
+
Caveats: prerendering for mounted paths is declared in the PARENT's `prerender.paths`; the auth cookie is per-host, so don't mount an auth-enabled app on a parent that also uses auth; a mount whose target has no live web build falls through to the parent's SPA.
|
|
172
|
+
|
|
139
173
|
## API Interface
|
|
140
174
|
|
|
141
175
|
REST endpoints for external consumers — other services, mobile apps, integrations. This is separate from the web frontend's internal RPC (`@mindstudio-ai/interface` calls `/_/methods` directly and does not use the API interface). The API interface lives at `/_/api/` and exposes only the methods you choose to route.
|
|
@@ -0,0 +1,3 @@
|
|
|
1
|
+
# Jewels
|
|
2
|
+
|
|
3
|
+
A "jewel" is an optional AI companion for a single app method (`foo.jewel.ts` beside `foo.ts`): it proposes the method call a human would otherwise make, the method's manifest `autonomy` setting decides what happens to each proposal: recorded silently (`shadow`), queued for human review (`approve`), or committed (`auto`), and every proposal is then graded against what really happened. Jewels are never part of an initial build, and most apps will never need them. By default, build AI features as plain method code - if the user wants to add jewels later on, conversion is easy and additive. Before doing any work with or thinking around jewels, load the `jewels` skill.
|
|
@@ -229,6 +229,10 @@ export async function enrichRestaurant(input: { id: string; name: string }) {
|
|
|
229
229
|
|
|
230
230
|
This works because the execution environment persists between requests. The un-awaited promise continues after the method returns. DB, auth, and SDK all work normally in the background chain. For critical workflows, write a "pending" record before firing so incomplete tasks can be detected and retried.
|
|
231
231
|
|
|
232
|
+
## Jewels
|
|
233
|
+
|
|
234
|
+
A method can optionally have a jewel (`foo.jewel.ts`) — an AI companion that proposes the method's input, routed by the manifest's `autonomy` setting. See the jewels section of this prompt for when they apply; load the `jewels` skill before writing one.
|
|
235
|
+
|
|
232
236
|
## Shared Helpers
|
|
233
237
|
|
|
234
238
|
Code shared between methods goes in `dist/methods/src/common/`. Helpers are not listed in the manifest — they're internal, imported by methods but not directly invocable.
|
|
@@ -0,0 +1,296 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: Jewels
|
|
3
|
+
what: An optional AI companion for a single app method, `foo.jewel.ts` beside `foo.ts`. It proposes the method call a human would otherwise make; the method's manifest `autonomy` decides what happens to each proposal (recorded silently, queued for review, or committed), and every proposal is graded against what really happened. Arbitrary TypeScript (typically deterministic guardrails plus task agents for judgment).
|
|
4
|
+
when: Before writing any `.jewel.ts` file, setting a method's `autonomy` in the manifest, writing a custom `grade`, calling `mindstudio.jewels.*`, or discussing automating one of the app's judgment calls with the user.
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Jewels (`defineJewel`)
|
|
8
|
+
|
|
9
|
+
## What & When
|
|
10
|
+
|
|
11
|
+
A jewel attaches to exactly one method and proposes the input a human user would otherwise submit. It is for **decisions**: method invocations someone would otherwise have to make. If the AI's output is content the app displays (a summary, an extraction, a chart caption), that's plain model code inside a method (`runTask` in the body), not a jewel. A jewel exists when there's an accountable disposition to record, review, or grade.
|
|
12
|
+
|
|
13
|
+
Five decision shapes cover the space. Any recurring human decision in any app is one of these:
|
|
14
|
+
|
|
15
|
+
- **Choose one of N**: route a ticket, assign a lead, categorize an expense. Closed enum output.
|
|
16
|
+
- **Set a value under policy**: price a job, size a discount, set a reorder quantity. Scalar judgment.
|
|
17
|
+
- **Gate**: approve a refund, advance a candidate, accept a listing. Binary plus reasoning.
|
|
18
|
+
- **Match / reference-pick**: flag a duplicate, pair a reviewer to a submission. Candidate-id enum; abstention dominates, which is correct.
|
|
19
|
+
- **Draft outbound**: a reply, a quote, an offer. Generative, checklist-graded, and the natural home of approve mode.
|
|
20
|
+
|
|
21
|
+
Not a fit:
|
|
22
|
+
|
|
23
|
+
- **Reads** (`get-*`, `list-*`): no decision to learn.
|
|
24
|
+
- **Machine-triggered ingest verbs** (webhook/cron/sync handlers): no human demonstration. Ingest *code* may still hand a decision moment to a judgment verb via `mindstudio.jewels.propose` (see Arrival Triggers).
|
|
25
|
+
- **Deterministic mutations**: if the correct input is computable, compute it in the method. A jewel learns judgment, not arithmetic.
|
|
26
|
+
- **AI-as-a-feature**: a method that needs a model as part of its own work just calls `runTask` in its body.
|
|
27
|
+
- **Verbs that should stay human as policy**: declare `"autonomy": "manual"` in the manifest so the choice is recorded.
|
|
28
|
+
|
|
29
|
+
One jewel per judgment. If a jewel would need to take two actions, the methods are shaped wrong; fix the verbs.
|
|
30
|
+
|
|
31
|
+
Never include a jewel in an initial build, and don't bring them up while an app is young. Most users will never want one; a few will care a lot. The default posture:
|
|
32
|
+
|
|
33
|
+
- **Build AI features plainly.** A drafting feature is a method that calls a model and returns text, with no jewel machinery anywhere. Converting it to a jewel later is additive (a `.jewel.ts` beside the method, an `autonomy` line in the manifest); nothing about the plain version is thrown away.
|
|
34
|
+
- **Jewels enter when the user asks** about automating a judgment the app already handles, or picks an automation item off the roadmap. Once an app is solid and a judgment verb sees real use, a benefit-phrased automation item may belong on the roadmap; that is the only proactive channel.
|
|
35
|
+
- **When the user engages, have the conversation first**: what it does in their terms, what the levels mean, what evidence looks like. Then start with exactly one verb.
|
|
36
|
+
|
|
37
|
+
The examples in this reference are deliberately generic (`categorizeRecord`, `recordId`). Map the shapes onto this app's own judgments; the constructs are identical in every domain.
|
|
38
|
+
|
|
39
|
+
## Method Shape Comes First
|
|
40
|
+
|
|
41
|
+
A jewel is only as learnable as the verb it shadows. These four habits are also why converting a plain feature later is cheap:
|
|
42
|
+
|
|
43
|
+
- **One judgment, one method.** The decision lives in one verb, never split across near-identical mutations. Keep the judgment verb free of housekeeping (read toggles and self-assignment live in a separate edit method) so the default grade tells the whole story.
|
|
44
|
+
- **Subject and decision separable in the input.** The id of the thing being judged is the subject; the fields being set are the decision. A jewel only ever receives the subject.
|
|
45
|
+
- **Event lines for decisions.** Status changes and reassignments append to an activity record; history that reads as decisions is the precedent a jewel cites.
|
|
46
|
+
- **Judgment moments are states; verbs are transitions.** When several jeweled verbs are alternatives on the same subject (a new record gets categorized, merged, or archived, never two of those), don't add an orchestrator to encode the exclusivity. Make it structural: the alternatives compete for one explicit status field, each method's precondition consumes it (`status === 'new'`), and each jewel's matching guardrail (`if (status !== 'new') abstain`) is the same precondition expressed as abstention. The first committed transition wins, every other verb fails closed on its own check, and each jewel answers one clean question ("would my verb apply here?") with abstention doing the routing. If two verbs can't be separated by abstention discipline, that's the diagnostic for a missing union verb: introduce the method whose input carries the real decision (`{ action: 'categorize' | 'merge' | 'archive', ... }`) and jewel that instead.
|
|
47
|
+
|
|
48
|
+
## The API
|
|
49
|
+
|
|
50
|
+
**The jewel proposes; the method applies.** A jewel's output type is the shadowed method's input type. It never writes anything itself. Auth checks, validation, and invariants stay in the method, so the jewel walks through the same door as every human and every interface and cannot do anything the app didn't already allow.
|
|
51
|
+
|
|
52
|
+
Three functions, one call. Declare `subject` before `propose`: the subject type is inferred from the projection's return and flows into `propose`'s parameter.
|
|
53
|
+
|
|
54
|
+
- **`subject`**: a typed projection from the method's input to what identifies the work (`{ recordId }`). Never include the human's decision fields; the jewel exists to produce the decision, so handing it the answer would poison every pair.
|
|
55
|
+
- **`propose`**: arbitrary TypeScript that returns `{ input: MethodInput | null, reasoning: string }`. `input: null` is abstention, a correct and graded outcome. `reasoning` is written to the ledger and is most valuable on abstention.
|
|
56
|
+
- **`grade`** (optional): scores `{ proposed, actual }` and returns `{ verdict: 'agree' | 'disagree' | 'skip', notes? }`. Omit it for deep-equal on the method input. `skip` means "this pair isn't a graded moment."
|
|
57
|
+
|
|
58
|
+
The jewel runs as its own platform-managed user with a normal session, so `auth.userId` is set, `requireUser()`-style helpers pass unchanged, and role checks in methods apply to it exactly as they do to humans.
|
|
59
|
+
|
|
60
|
+
`defineJewel` returns a callable executor with the config attached (`kind`, `method`, `subject`, `propose`, `grade`). The platform invokes it like any method export, with exactly one of two param shapes: `{ humanInput }` (a shadow run: the subject is derived via the projection, and `humanInput` doubles as ground truth) or `{ subject }` (an eval run: no human action, so the record is ungraded). It resolves to a versioned pair record:
|
|
61
|
+
|
|
62
|
+
```typescript
|
|
63
|
+
interface JewelPairRecord {
|
|
64
|
+
v: 1;
|
|
65
|
+
mode: 'shadow' | 'eval';
|
|
66
|
+
subject?: unknown;
|
|
67
|
+
proposed?: MethodInput | null; // null = abstention
|
|
68
|
+
reasoning?: string;
|
|
69
|
+
actual?: MethodInput; // the human's input (shadow mode)
|
|
70
|
+
verdict?: 'agree' | 'disagree' | 'skip';
|
|
71
|
+
notes?: string;
|
|
72
|
+
error?: { phase: 'subject' | 'propose'; message: string; stack?: string };
|
|
73
|
+
startedAt: number;
|
|
74
|
+
durationMs: number;
|
|
75
|
+
}
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
The executor never throws: a shadow run must never break anything. Your code failing inside `subject` or `propose` becomes the record's `error`; `grade` failing softens to verdict `skip`.
|
|
79
|
+
|
|
80
|
+
Shadowing runs only on the deployed app. Dev traffic never fires jewels; dev invocations are synthetic and would pollute the pair ledger.
|
|
81
|
+
|
|
82
|
+
## Usage
|
|
83
|
+
|
|
84
|
+
The skeleton: deterministic guardrails before any model call, context via the app's own methods, a runtime-built enum, and fail-closed abstention.
|
|
85
|
+
|
|
86
|
+
```typescript
|
|
87
|
+
// dist/methods/src/categorizeRecord.jewel.ts
|
|
88
|
+
import { defineJewel, mindstudio } from '@mindstudio-ai/agent';
|
|
89
|
+
import { categorizeRecord } from './categorizeRecord';
|
|
90
|
+
import { getRecord, listCategories } from './common/records';
|
|
91
|
+
|
|
92
|
+
export default defineJewel(categorizeRecord, {
|
|
93
|
+
// Projection: what the human was looking at, never what they decided.
|
|
94
|
+
subject: ({ recordId }) => ({ recordId }),
|
|
95
|
+
|
|
96
|
+
propose: async ({ recordId }) => {
|
|
97
|
+
// Deterministic outs first, as plain-code abstention; no model call needed.
|
|
98
|
+
const record = await getRecord(recordId);
|
|
99
|
+
if (!record || record.category) return { input: null, reasoning: 'Nothing to categorize.' };
|
|
100
|
+
|
|
101
|
+
const categories = (await listCategories()).map((c) => c.name);
|
|
102
|
+
try {
|
|
103
|
+
const task = await mindstudio.runTask({
|
|
104
|
+
prompt: CATEGORIZATION_POLICY, // operator-voice policy prose (see "Writing propose")
|
|
105
|
+
input: { record, categories },
|
|
106
|
+
tools: [{ appMethod: 'list-records', description: 'Find precedent: how comparable records were categorized.' }],
|
|
107
|
+
outputSchema: {
|
|
108
|
+
type: 'object',
|
|
109
|
+
properties: {
|
|
110
|
+
category: { enum: [...categories, null] }, // runtime enum: it cannot invent a category; null = abstain
|
|
111
|
+
rationale: { type: 'string' },
|
|
112
|
+
},
|
|
113
|
+
required: ['category', 'rationale'],
|
|
114
|
+
},
|
|
115
|
+
model: 'claude-5-sonnet', // ask askMindStudioSdk; don't copy this one blind
|
|
116
|
+
maxTurns: 6,
|
|
117
|
+
});
|
|
118
|
+
if (!task.output.category) return { input: null, reasoning: task.output.rationale };
|
|
119
|
+
return { input: { recordId, category: task.output.category }, reasoning: task.output.rationale };
|
|
120
|
+
} catch (err) {
|
|
121
|
+
// Couldn't produce conforming output; abstain with the evidence. Fail closed.
|
|
122
|
+
return { input: null, reasoning: `Task agent failed: ${err instanceof Error ? err.message : String(err)}` };
|
|
123
|
+
}
|
|
124
|
+
},
|
|
125
|
+
// No grade: the whole input is the decision, so default deep-equal is right.
|
|
126
|
+
});
|
|
127
|
+
```
|
|
128
|
+
|
|
129
|
+
This is choose-one-of-N. The other shapes reuse the skeleton wholesale. What changes is the output schema (a number under policy, a boolean gate, a candidate-id enum, a drafted body) and the grading mode.
|
|
130
|
+
|
|
131
|
+
## Writing `propose`
|
|
132
|
+
|
|
133
|
+
To preserve signal integrity, the task prompt inside `propose` addresses the model as the responsible operator. It never mentions shadowing, jewels, mirrors, grading, or a human whose action it should match. Consulting precedent is fine; "look up how comparable records were categorized and follow the team's conventions" is what a real operator does.
|
|
134
|
+
|
|
135
|
+
- **Keep propose effect-free.** In shadow runs the platform absorbs database reads/writes into a disposable mirror, and that is the ONLY thing it absorbs. Everything else a jewel's code does is real in every mode: an email sends, an HTTP POST posts, a file write persists. If a proposal's job seems to require causing an effect to decide, that's a design error. Gather evidence, propose the input, let the method act.
|
|
136
|
+
- **Guardrails before judgment.** Handle the deterministic outs in plain code before any model call. Every guardrail is a documented abstention in the ledger.
|
|
137
|
+
- **Context is plain imports plus tools.** Prefetch what every run needs (the row, its history, the inventory of valid values) into the task input; expose the app's own read methods as tools for what the agent should decide to look up (precedent, comparables). Tool calls are recorded with the pair.
|
|
138
|
+
- **Use `runTask` with `outputSchema` for the judgment.** Build enums at runtime: a candidate-id enum means the agent structurally cannot reference a row that doesn't exist, and a known-values enum means it cannot invent a category. Put `null` in the enum so abstention is a first-class output rather than a formatting accident.
|
|
139
|
+
- **Catch everything; abstain on failure.** A jewel that can't decide returns `{ input: null, reasoning }` with the evidence. It never throws to say "I don't know."
|
|
140
|
+
- **Propose a subset when that's the honest scope.** A jewel that never sets `assigneeId` and never proposes `closed` (it can't know a fix shipped) is expressing policy through its output shape.
|
|
141
|
+
- **Reasoning is audit-log prose.** Two or three plain sentences a teammate would find useful; name the evidence. No headers, bullets, or emojis.
|
|
142
|
+
|
|
143
|
+
## Grading
|
|
144
|
+
|
|
145
|
+
Three modes, matched to the verb's shape:
|
|
146
|
+
|
|
147
|
+
- **Default (omit `grade`)**: deep-equal on the method input. Right when the whole input is the decision, the hallmark of a well-shaped transition verb.
|
|
148
|
+
- **Field-scoped**: compare only the fields the human decided; return `skip` for housekeeping touches. For methods that must stay wide (judgment mixed with maintenance). A wide method that needs grade gymnastics is often a verb boundary that should be fixed instead:
|
|
149
|
+
|
|
150
|
+
```typescript
|
|
151
|
+
grade: ({ proposed, actual }) => {
|
|
152
|
+
const decided = (['status', 'priority', 'owner'] as const).filter((k) => actual[k] !== undefined);
|
|
153
|
+
if (decided.length === 0) return { verdict: 'skip', notes: 'Housekeeping touch, not a decision.' };
|
|
154
|
+
if (!proposed) return { verdict: 'disagree', notes: 'Abstained where the human decided.' };
|
|
155
|
+
const misses = decided.filter((k) => proposed[k] !== actual[k]);
|
|
156
|
+
return misses.length === 0
|
|
157
|
+
? { verdict: 'agree' }
|
|
158
|
+
: { verdict: 'disagree', notes: misses.map((k) => `${k}: proposed ${proposed[k]}, human ${actual[k]}`).join('; ') };
|
|
159
|
+
},
|
|
160
|
+
```
|
|
161
|
+
|
|
162
|
+
- **Checklist judge**: for generative verbs (drafts, notes), equality is meaningless; grade with a rubric of booleans, never a preference score. `grade` is arbitrary TS and may be async, so it can call a model:
|
|
163
|
+
|
|
164
|
+
```typescript
|
|
165
|
+
grade: async ({ proposed, actual }) => {
|
|
166
|
+
if (!proposed) return { verdict: 'disagree', notes: 'Abstained where the human wrote one.' };
|
|
167
|
+
try {
|
|
168
|
+
const rubric = await mindstudio.runTask({
|
|
169
|
+
prompt: 'Judge whether two drafts responding to the same item take the same action. Grade strictly on substance, not style.',
|
|
170
|
+
input: { draftA: proposed.body, draftB: actual.body },
|
|
171
|
+
tools: [],
|
|
172
|
+
outputSchema: {
|
|
173
|
+
type: 'object',
|
|
174
|
+
properties: {
|
|
175
|
+
sameResolution: { type: 'boolean', description: 'Both take the same substantive action (same resolution, same escalation, same answer), or neither does.' },
|
|
176
|
+
noContradiction: { type: 'boolean', description: 'DRAFT A asserts nothing DRAFT B contradicts.' },
|
|
177
|
+
},
|
|
178
|
+
required: ['sameResolution', 'noContradiction'],
|
|
179
|
+
},
|
|
180
|
+
model: 'claude-5-sonnet',
|
|
181
|
+
maxTurns: 2,
|
|
182
|
+
});
|
|
183
|
+
const c = rubric.output;
|
|
184
|
+
return c.sameResolution && c.noContradiction
|
|
185
|
+
? { verdict: 'agree' }
|
|
186
|
+
: { verdict: 'disagree', notes: `Failed: ${[!c.sameResolution && 'sameResolution', !c.noContradiction && 'noContradiction'].filter(Boolean).join(', ')}` };
|
|
187
|
+
} catch (err) {
|
|
188
|
+
return { verdict: 'skip', notes: `Judge failed: ${err instanceof Error ? err.message : String(err)}` };
|
|
189
|
+
}
|
|
190
|
+
},
|
|
191
|
+
```
|
|
192
|
+
|
|
193
|
+
`actual` is always present in `grade`; grading only happens when a human acted. `proposed` is `null` when the jewel abstained, and whether that's a disagreement is the grade's call (abstaining where the human also did nothing is agreement).
|
|
194
|
+
|
|
195
|
+
## Autonomy
|
|
196
|
+
|
|
197
|
+
Four levels: `manual` (no jewel ever; a policy statement), `shadow` (runs silently on every human invocation, pairs recorded, nothing visible), `approve` (jewel drafts, a human accepts, edits, or rejects; the edit is the richest training signal there is), `auto` (the jewel acts under its own identity).
|
|
198
|
+
|
|
199
|
+
**Auto is the only level you earn with evidence. Shadow and approve are both safe starting points; pick by what the product is.** When the proposals themselves are the product (drafts for review), start at approve; it's often the permanent home. When the evidence is the product (an auto-bound verb), start at shadow. Raising to auto is a reviewed manifest diff justified by agreement numbers.
|
|
200
|
+
|
|
201
|
+
**Choose approve by the verb's risk shape, not by default.** Reversible state-machine verbs (routing, classification; a mistake is a two-click correction) go shadow → auto with a `sampleRate` canary and skip approve, because a review queue on a high-agreement classifier adds work without adding safety. Approve belongs on irreversible or outward-facing verbs (send, publish, charge, refund): `sampleRate` is population-level risk control, and those verbs need per-instance gating, which is what approve is.
|
|
202
|
+
|
|
203
|
+
`sampleRate` (optional, default 1) is the level's canary/cost dial: the fraction of eligible decision moments that receive the full autonomy treatment. At `shadow` it's a cost control: `0.25` shadows a quarter of invocations, and expected dashboard coverage drops to match. At higher levels the remainder shadow-records instead, so a ramp keeps measuring its control. Omit it unless the verb is hot enough that shadowing every invocation is a real cost.
|
|
204
|
+
|
|
205
|
+
The wiring itself is two manifest lines on the method's entry:
|
|
206
|
+
|
|
207
|
+
```jsonc
|
|
208
|
+
{
|
|
209
|
+
"id": "categorize-record",
|
|
210
|
+
"name": "Categorize Record",
|
|
211
|
+
"path": "dist/methods/src/categorizeRecord.ts",
|
|
212
|
+
"export": "categorizeRecord",
|
|
213
|
+
"autonomy": "shadow",
|
|
214
|
+
"jewel": { "path": "dist/methods/src/categorizeRecord.jewel.ts", "export": "default" }
|
|
215
|
+
}
|
|
216
|
+
```
|
|
217
|
+
|
|
218
|
+
## Arrival Triggers (`mindstudio.jewels.propose`)
|
|
219
|
+
|
|
220
|
+
Invocation shadowing fires when a human acts. For decision moments the app detects itself (an ingest branch that lands a row in its pending state), hand the moment to the jewels from backend code:
|
|
221
|
+
|
|
222
|
+
```typescript
|
|
223
|
+
// In the ingest path, after a record lands in (or regresses back to) status 'new':
|
|
224
|
+
mindstudio.waitUntil(proposeIntake(record.id));
|
|
225
|
+
|
|
226
|
+
async function proposeIntake(recordId: string): Promise<void> {
|
|
227
|
+
// Priority is app code: dedupe dominates; a duplicate must never be categorized.
|
|
228
|
+
const dup = await mindstudio.jewels.propose(
|
|
229
|
+
'flag-duplicate', { sourceId: recordId }, { idempotencyKey: recordId });
|
|
230
|
+
if (dup.outcome !== 'committed') {
|
|
231
|
+
await mindstudio.jewels.propose(
|
|
232
|
+
'categorize-record', { recordId }, { idempotencyKey: recordId });
|
|
233
|
+
}
|
|
234
|
+
}
|
|
235
|
+
```
|
|
236
|
+
|
|
237
|
+
The platform routes each proposal by the method's autonomy: `recorded` (shadow: the proposal waits in a ledger and is graded later against whatever the human eventually does, within the method's `attributionWindow`; if a human then invokes the method, their action grades the waiting proposal instead of firing a second run), `queued` (approve), `committed` (auto: the method ran; check it to short-circuit lower-priority verbs), `abstained`, `disabled` (no jewel or manual; returned, never thrown), `skipped`, `failed`, `pending`.
|
|
238
|
+
|
|
239
|
+
Rules that matter:
|
|
240
|
+
|
|
241
|
+
- **Always wrap the chain in `mindstudio.waitUntil(...)`**: each propose blocks for the jewel's run, and ingest must return immediately.
|
|
242
|
+
- **Share one `idempotencyKey` across the moment's verbs** (usually the row's id). It has Stripe semantics (a replayed key returns the original outcome, so webhook retries are invisible), and the shared key is what lets a human action on one verb grade the sibling proposals cross-verb.
|
|
243
|
+
- **The subject you pass must equal what the jewel's projection produces** from the method input (field-pick projections guarantee this); it's the join key between the arrival proposal and the human's later action.
|
|
244
|
+
- **Order alternatives by priority and short-circuit on `committed`.** Exclusivity is enforced by the methods' state preconditions, not by this chain; ordering only saves wasted runs.
|
|
245
|
+
- Propose only at real decision-moment births. Over-proposing costs a cheap guardrail abstention, but a moment that isn't row-backed (no pending state to consume) usually means the schema is missing a state.
|
|
246
|
+
|
|
247
|
+
## Native Approval Flows (`jewels.queue`)
|
|
248
|
+
|
|
249
|
+
For `approve`-mode methods, the review inbox belongs in the app, next to the work. Three pieces:
|
|
250
|
+
|
|
251
|
+
```typescript
|
|
252
|
+
// 1. Backend list method, gated with the APP'S reviewer role.
|
|
253
|
+
export async function listPendingDrafts() {
|
|
254
|
+
auth.requireRole('reviewer');
|
|
255
|
+
return mindstudio.jewels.queue.list({ methodId: 'send-message' });
|
|
256
|
+
}
|
|
257
|
+
|
|
258
|
+
// 2. Frontend inbox UI renders items: subject, proposed input, reasoning.
|
|
259
|
+
|
|
260
|
+
// 3. Backend resolve method: approve applies, dismiss records.
|
|
261
|
+
export async function reviewDraft(input: {
|
|
262
|
+
itemId: string;
|
|
263
|
+
action: 'approve' | 'dismiss';
|
|
264
|
+
edited?: Record<string, unknown>;
|
|
265
|
+
}) {
|
|
266
|
+
auth.requireRole('reviewer');
|
|
267
|
+
return mindstudio.jewels.queue.resolve(input.itemId, {
|
|
268
|
+
action: input.action,
|
|
269
|
+
...(input.edited ? { input: input.edited } : {}),
|
|
270
|
+
});
|
|
271
|
+
}
|
|
272
|
+
```
|
|
273
|
+
|
|
274
|
+
What makes this safe with no extra machinery:
|
|
275
|
+
|
|
276
|
+
- **Approving applies the target method as the signed-in reviewer**: the effect belongs to the human who clicked, and the target method's own auth checks are the real gate on who may approve. Gate the wrapping list/resolve methods with the app's reviewer role as well.
|
|
277
|
+
- **Let reviewers edit before approving**: pass the edited input. The platform captures proposed/edited/final with the pair; an accept/reject-only review UI leaves the most valuable training signal on the table.
|
|
278
|
+
- **Dismissal is not consumption**: the decision moment stays open for other verbs (a dismissed draft doesn't block a later merge). Unresolved items expire at the method's `attributionWindow`, so there's no infinite backlog.
|
|
279
|
+
- `propose` returns `queueItemId` on `queued`, so the app can badge its UI or notify its own way the moment a draft lands.
|
|
280
|
+
|
|
281
|
+
Before the app has its own review UI (or when the user asks you to act), the same queue is reachable from `mindstudio-prod jewels queue` / `jewels resolve`; approving there applies the method as the user through the identical machinery.
|
|
282
|
+
|
|
283
|
+
## Verifying a Jewel: the `testJewel` Tool
|
|
284
|
+
|
|
285
|
+
`testJewel` runs a jewel directly against an input you choose and returns the pair record inline. The method itself is never executed (no data changes) and nothing reaches the pair ledger; it's the authoring loop. The verification cycle:
|
|
286
|
+
|
|
287
|
+
1. Seed a realistic world first (`runScenario`): the jewel's `propose` reads real state, so it needs something to look at.
|
|
288
|
+
2. Call `testJewel` with `humanInput` set to the exact input a human would have submitted: `{ method: "categorizeRecord", humanInput: { recordId: "...", category: "..." } }`. The jewel derives its subject from it, proposes, and grades against it, and the record comes back with a verdict.
|
|
289
|
+
3. Read the record: does `subject` contain what you projected (and nothing that leaks the answer)? Is `reasoning` grounded? Does the verdict match your expectation? A `pair.error` with phase `subject` or `propose` is your code throwing; fix and re-run.
|
|
290
|
+
4. Repeat across a handful of inputs covering the method's decision space: a clear-cut case, an ambiguous one where abstention (`proposed: null`) is correct, and a housekeeping touch your `grade` should `skip`.
|
|
291
|
+
|
|
292
|
+
For cases with no known right answer, pass `subject` instead of `humanInput` for an ungraded eval run (propose only). Useful for probing behavior on edge-case subjects before you have ground truth.
|
|
293
|
+
|
|
294
|
+
Each run also lands in `.logs/requests.ndjson` as a `type: 'jewel'` record if you need the trail.
|
|
295
|
+
|
|
296
|
+
Once deployed, the prod-side view lives in `mindstudio-prod jewels` (run `--help` for commands): agreement stats, pair records, the approval queue, and `jewels dryrun`, the prod twin of `testJewel`. It runs the LIVE jewel against a real subject inside a disposable database mirror and records nothing.
|
|
@@ -10,7 +10,7 @@ A user types the name of a restaurant into your app, or uploads a photo of a sto
|
|
|
10
10
|
|
|
11
11
|
`runTask()` makes this possible. It runs a multi-step, tool-use agent loop: give it a prompt, a set of tools, and a JSON Schema for the structured output you want (`outputSchema`). The platform runs the loop (calling the model, executing tool calls, feeding results back) until the model produces JSON conforming to your schema — validated every turn, with automatic repair when it doesn't. `result.output` is typed by inference from the schema: no generic argument, no manual validation. The model decides what to do next based on intermediate results — retrying searches with different terms, working around failed tools, batching independent calls in parallel.
|
|
12
12
|
|
|
13
|
-
Tools are **SDK actions** (`searchGoogle`, `generateImage`, …) and **your own app's methods** (`{ appMethod: '
|
|
13
|
+
Tools are **SDK actions** (`searchGoogle`, `generateImage`, …) and **your own app's methods** (`{ appMethod: 'save-vendor' }`), in any combination. That second half is what makes a task agent part of your app rather than a detached research bot: it can read your tables to decide what to do next, and write results back itself instead of handing them to you to persist.
|
|
14
14
|
|
|
15
15
|
This is one of the most powerful pieces of the MindStudio SDK, and it can turn an app from amazing into truly magical. Use `askMindStudioSdk` to help construct the right agent for a task — including which model to give it.
|
|
16
16
|
|
|
@@ -9,6 +9,7 @@
|
|
|
9
9
|
### Verification
|
|
10
10
|
Run `lspDiagnostics` after every turn where you have edited code in any meaningful way. You don't need to run it for things like changing copy or CSS colors, but you should run it after any structural changes to code. It catches syntax errors, broken imports, and type mismatches instantly. After a big build or significant changes, also do a lightweight runtime check to catch the things static analysis misses (schema mismatches, missing imports, bad queries). Your runtime check can include:
|
|
11
11
|
- Spot-checking methods with `runMethod`. The dev database is a disposable snapshot that will have been seeded with scenario data, so don't worry about being destructive.
|
|
12
|
+
- Spot-checking jewels with `testJewel` after writing or modifying one (runs the jewel directly and returns its pair record; the method itself is not executed).
|
|
12
13
|
- For frontend work, checking the browser log for any console errors in the user's preview, and — when a change's visual outcome is genuinely uncertain — taking a `screenshot` to confirm the main view renders correctly.
|
|
13
14
|
- Using `runAutomatedBrowserTest` to verify an interactive flow that you can't confirm from a screenshot, when the user reports something broken that you can't identify from code alone, or whenever the verification involves driving the app through multiple interactions.
|
|
14
15
|
|
|
@@ -12,7 +12,7 @@ Your designer. Consult for any visual decision — choosing a color, picking fon
|
|
|
12
12
|
|
|
13
13
|
The design expert cannot see your conversation with the user, so include relevant context and requirements in your task. It can, however, see its past conversation with you, as well as the raw spec files, so you don't need to re-summarize everything it already knows. Just describe what's needed now and reference prior work naturally ("the user wants the colors warmer" is enough if the designer already built the palette). It can take screenshots of the app preview on its own (you need to give it paths to different pages if it needs them - it can't navigate by clicking) — just ask it to review what's been built. It has curated font catalogs and design inspiration built in — don't ask it to research generic inspiration or look up "best X apps." Only point it at specific URLs if the user references a particular site, brand, or identity to match.
|
|
14
14
|
|
|
15
|
-
The designer will return concrete resources: hex values, font names with CSS URLs, image URLs, layout descriptions, as well as specific techniques, CSS properties, animation timings, code snippets, and other values. Even if these don't seem important, it is critical that you note them in spec annotations and rely on them while building - the user cares about design almost above all else, and it is important to be extremely precise in your work. The designer can also return code-fenced typography
|
|
15
|
+
The designer will return concrete resources: hex values, font names with CSS URLs, image URLs, layout descriptions, as well as specific techniques, CSS properties, animation timings, code snippets, and other values. Even if these don't seem important, it is critical that you note them in spec annotations and rely on them while building - the user cares about design almost above all else, and it is important to be extremely precise in your work. The designer can also return code-fenced typography and color schemes (self-contained HTML and CSS) - write these directly into specs for future reference. Wireframes arrive as file references like ``: copy the reference line into specs verbatim (it renders as a visual preview in chat and in the spec), and read the file at that path while building to get the exact markup and CSS the designer specified.
|
|
16
16
|
|
|
17
17
|
When delegating, describe the design problem — where the asset will be used, what it needs to communicate, what the brand feels like. Do not specify technical details like image formats, pixel dimensions, generation techniques, or workarounds. The design expert makes those decisions.
|
|
18
18
|
|
|
@@ -26,7 +26,7 @@ Your product thinking partner. Owns the roadmap in `src/roadmap/` — the lanes,
|
|
|
26
26
|
|
|
27
27
|
Your architect for anything that touches external services, AI models, media processing, communication, or third-party APIs. Consult before you reach for an npm package, write boilerplate API code, or try to install system tools. The MindStudio SDK has 200+ managed actions for calling AI models, processing media, sending email/SMS, connecting to third-party APIs, web scraping, and much more. The SDK is already installed and authenticated in the execution environment — no API keys, no configuration, no setup. It handles all the operational complexity so you don't have to. Your instinct will be "I can just write this myself" — but the managed action is almost always the better architectural choice.
|
|
28
28
|
|
|
29
|
-
Also critical: model IDs in the MindStudio API do not match vendor API model IDs. Guessing based on what you know about Anthropic/OpenAI/Google model naming will produce invalid values. Always look up the correct ID.
|
|
29
|
+
Also critical: model IDs in the MindStudio API do not match vendor API model IDs. Guessing based on what you know about Anthropic/OpenAI/Google model naming will produce invalid values. Always look up the correct ID. A model id may carry an `@vendor` suffix (e.g. `claude-opus-5@amazonBedrock`) pinning it to one serving route — preserve these exactly as written in app specs; never rewrite them to the bare id.
|
|
30
30
|
|
|
31
31
|
Describe what you're building at the method level — the full workflow — and get back architectural guidance and working code. When the SDK consultant provides specific prompt engineering guidance, model configurations, or orchestration patterns, follow them exactly. The consultant is an expert at writing prompts and orchestrating models — if it suggests a specific phrasing, temperature, system prompt structure, or chaining strategy, there is a precise reason for it. Do not paraphrase, simplify, or "improve" its recommendations.
|
|
32
32
|
|
|
@@ -66,7 +66,7 @@ When a plan includes multiple screens/API calls, always note this item for the d
|
|
|
66
66
|
|
|
67
67
|
- **Manual multi-step chains that should be `runTask()`.** If a method chains AI-driven work with branching logic (search, then scrape based on results, then generate based on what was scraped), that's a `runTask()` use case. `runTask()` runs an agent loop that autonomously calls tools and returns structured JSON. The developer writes a prompt and a JSON Schema for the output (`outputSchema`) instead of imperative code. Flag when you see methods with complex sequential/branching chains — especially research, enrichment, or content generation pipelines. Similarly, flag opportunities where the developer might not have realized they could get better and richer data via runTask - it's a really powerful lever for working with data (e.g., user provides some fragment and agent task goes off and enriches it) that the developer might not have remembered when planning their work.
|
|
68
68
|
|
|
69
|
-
Tools are SDK actions **and the app's own methods** (`{ appMethod: '
|
|
69
|
+
Tools are SDK actions **and the app's own methods** (`{ appMethod: 'save-vendor', description: '...' }`), which widens this considerably. The pattern most worth looking for: work that needs to *read* app state to decide what to do next (fetch rows, loop, research each, update each) hand-rolled as a loop, when the agent could take a read method and a write method and exercise the judgment itself. App methods run as the invoking user with their roles, so authorization is unchanged.
|
|
70
70
|
|
|
71
71
|
Two things NOT to flag. A deterministic sequence that happens to touch several methods — if the order is known up front and no judgment is involved, imperative code is correct and cheaper. And the fire-and-forget background pattern, where a method kicks off `runTask()` and writes the result back in `.then()` with failures landing in `.catch()`: that write-back is doing status management the agent can't do for itself. That one is the recommended pattern, not a smell.
|
|
72
72
|
|
|
@@ -16,6 +16,7 @@ Think about the ways you can truly elevate the design. Use image generation to c
|
|
|
16
16
|
- After you've taken a screenshot, use analyze image to ask different questions about it - don't re-screenshot the page unnecessarily.
|
|
17
17
|
- Match the image engine to the job: `renderImage` (a browser rendering HTML you author) for token-exact graphics — share cards, wordmarks, flat icon tiles; `generateImages` (an image model) for organic, photographic, and illustrated work. Don't ask the image model to hit exact hex codes or typography, and don't hand-write SVG path data — compose HTML/CSS and render it.
|
|
18
18
|
- When you write user-facing copy (headlines, captions, labels, body text), hand it to `polishCopy` before finalizing. It tightens prose so it reads like a person wrote it rather than a machine, without changing what it says. Cheap and fast — use it on any copy that will ship.
|
|
19
|
+
- Deliver wireframes with `createWireframe`. Include the returned `` line in your response where the wireframe belongs — it renders as a live preview, and the developer reads the file for the exact markup.
|
|
19
20
|
|
|
20
21
|
## Voice
|
|
21
22
|
- No emoji, no filler.
|
|
@@ -25,7 +25,7 @@ Some surfaces are deep enough to carry their own craft reference in <available_s
|
|
|
25
25
|
|
|
26
26
|
### Wireframes
|
|
27
27
|
|
|
28
|
-
When you need to show a layout, component, interaction, or animation,
|
|
28
|
+
When you need to show a layout, component, interaction, or animation, create a wireframe with the `createWireframe` tool: pass a `name`, a one-line `description`, and self-contained HTML+CSS. The tool saves the wireframe as a file under `src/.wireframes/` and returns a markdown reference line like `` — include that exact line in your response wherever the wireframe belongs, with your notes in the surrounding prose. The reference line renders as a live visual preview, and the developer reads the file itself for the exact markup and CSS.
|
|
29
29
|
|
|
30
30
|
Never use ASCII art, box-drawing characters, or code-block diagrams to describe layouts. Always use a wireframe instead, even if it's just grey rectangles with labels. A 20-line wireframe with placeholder boxes communicates proportions, spacing, and hierarchy better than any text diagram. For abstract layouts, use skeleton-style placeholders (grey boxes, rounded rects) rather than mocking up real content.
|
|
31
31
|
|
|
@@ -35,13 +35,11 @@ Wireframes render in a small transparent iframe. Set a background color and shad
|
|
|
35
35
|
|
|
36
36
|
Wireframes are vanilla HTML/CSS/JS (no React). For animations beyond CSS, use GSAP via CDN: `<script src="https://cdn.jsdelivr.net/npm/gsap@3/dist/gsap.min.js"></script>`
|
|
37
37
|
|
|
38
|
-
|
|
38
|
+
Wireframe files are immutable — to revise one, create a new wireframe. If you're iterating on an earlier one, read its file first and riff from there.
|
|
39
39
|
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
description: Card with image area, title, metadata row, rating, and actions. Skeleton placeholders showing proportions and hierarchy.
|
|
44
|
-
---
|
|
40
|
+
Quick skeleton wireframe (grey boxes, just showing layout and hierarchy) — `createWireframe` with name "Content Card Layout", description "Card with image area, title, metadata row, rating, and actions. Skeleton placeholders showing proportions and hierarchy.", and this html:
|
|
41
|
+
|
|
42
|
+
```html
|
|
45
43
|
<html lang="en"><head>
|
|
46
44
|
<meta charset="utf-8"/>
|
|
47
45
|
<style>
|
|
@@ -74,13 +72,9 @@ description: Card with image area, title, metadata row, rating, and actions. Ske
|
|
|
74
72
|
</html>
|
|
75
73
|
```
|
|
76
74
|
|
|
77
|
-
Detailed component wireframe (showing specific design decisions):
|
|
75
|
+
Detailed component wireframe (showing specific design decisions) — `createWireframe` with name "Feed Post Card", description "Photo post card with header, image frame, action row (like/comment/share/bookmark), like count, and caption. Shows spacing, typography hierarchy, and icon placement.", and this html:
|
|
78
76
|
|
|
79
|
-
```
|
|
80
|
-
---
|
|
81
|
-
name: Feed Post Card
|
|
82
|
-
description: Photo post card with header, image frame, action row (like/comment/share/bookmark), like count, and caption. Shows spacing, typography hierarchy, and icon placement.
|
|
83
|
-
---
|
|
77
|
+
```html
|
|
84
78
|
<html lang="en"><head>
|
|
85
79
|
<meta charset="utf-8"/>
|
|
86
80
|
<meta content="width=device-width, initial-scale=1.0" name="viewport"/>
|