@a-t-h-i/bot-lobby 0.3.0 → 0.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +128 -30
- package/package.json +1 -1
- package/prompts/backend.md +46 -1
- package/prompts/designer.md +94 -15
- package/prompts/master.md +58 -1
- package/prompts/qa.md +35 -2
- package/prompts/researcher.md +6 -0
- package/prompts/reviewer.md +16 -0
- package/prompts/scout.md +11 -2
- package/prompts/worker.md +35 -2
- package/src/desk/client-extension.ts +101 -0
- package/src/desk/desk.ts +249 -0
- package/src/desk/ipc.ts +178 -0
- package/src/desk/session.ts +214 -0
- package/src/execution/agent-runner.ts +160 -14
- package/src/execution/pi-runner.ts +351 -65
- package/src/index.ts +4 -0
- package/src/master/master.ts +21 -17
- package/src/master/research.ts +8 -7
- package/src/pi/activity.ts +45 -0
- package/src/pi/commands.ts +28 -3
- package/src/pi/expressions.ts +43 -12
- package/src/pi/kaomoji.ts +227 -0
- package/src/pi/mascot-art.ts +5 -15
- package/src/pi/model-support.ts +101 -0
- package/src/pi/run-summary.ts +170 -0
- package/src/pi/settings-ui.ts +158 -60
- package/src/pi/tools.ts +50 -7
- package/src/pi/ui.ts +48 -7
- package/src/pi/zen-large.ts +41 -4
- package/src/pi/zen-metrics.ts +47 -7
- package/src/pi/zen.ts +29 -8
- package/src/roles/reviewer.ts +24 -4
- package/src/roles/worker.ts +7 -1
- package/src/schemas/configuration.ts +112 -18
- package/src/schemas/findings.ts +20 -0
- package/src/schemas/task.ts +26 -1
- package/src/state/project.ts +9 -0
- package/src/text.ts +9 -0
- package/src/workflow/workflow.ts +127 -10
package/README.md
CHANGED
|
@@ -83,8 +83,9 @@ structure.
|
|
|
83
83
|
/bot-lobby amend <text> Record an amendment; the Master re-proposes
|
|
84
84
|
/bot-lobby decline Decline the proposal and abandon the task
|
|
85
85
|
/bot-lobby knowledge Knowledge file sizes vs. the compaction threshold
|
|
86
|
+
/bot-lobby runs [taskId] Recent subagent runs: time, turns, tools, tokens, cost, model
|
|
86
87
|
/bot-lobby config Effective configuration and its file path
|
|
87
|
-
/bot-lobby settings Edit
|
|
88
|
+
/bot-lobby settings Edit each agent's model, thinking, time limit and instructions
|
|
88
89
|
/bot-lobby-settings Same as the settings subcommand
|
|
89
90
|
/bot-lobby minimize|restore Hide or restore bot-lobby for this session (ctrl+shift+m)
|
|
90
91
|
/bot-lobby claim <taskId> Take ownership of an orphaned task
|
|
@@ -136,19 +137,41 @@ four animated slots — DEV, DESIGN, RESEARCH and QA — each with a status face
|
|
|
136
137
|
caption and two status rows: while running, a braille spinner beside the agent's
|
|
137
138
|
live one-word activity (for example `⠋ reading` or `⠋ editing`) with its elapsed
|
|
138
139
|
time on the row beneath; otherwise the coloured status glyph and state word over
|
|
139
|
-
that elapsed time
|
|
140
|
-
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
140
|
+
that elapsed time. A working agent's status row turns into a warning when it is
|
|
141
|
+
waiting on a file (`⧗ waiting`), has gone quiet (`! quiet 1m`) or is retrying a
|
|
142
|
+
provider call (`↻ retrying`). Under the agents, a live feed row says exactly what
|
|
143
|
+
one of them is doing — `▸ DEV editing users.ts · turn 4 · 12 tools · 41k tok` —
|
|
144
|
+
rotating between working agents every few seconds and putting warnings first
|
|
145
|
+
(gone quiet, waiting on a file, asked to wrap up); once nothing runs it shows the
|
|
146
|
+
last run's receipt. It only takes a spare line, so it never costs the tower, the
|
|
147
|
+
agents or a checklist row. A full-width TASKS checklist windowed on the current
|
|
148
|
+
step closes the scene. Narrower terminals keep the boxed banner, header and
|
|
149
|
+
compact animated strip, whose working line names the newest running agent's
|
|
150
|
+
activity, its target and elapsed time.
|
|
151
|
+
|
|
152
|
+
Each agent has a kaomoji personality. About 300 faces across 15 emotions (happy,
|
|
153
|
+
proud, love, excited, focused, curious, thinking, nervous, confused, sleepy, sad,
|
|
154
|
+
angry, waiting, surprised, grateful) come from a shared pool every agent can use
|
|
155
|
+
plus each agent's own set of at least four faces per emotion: DEV wears shades,
|
|
156
|
+
flexes and flips tables `(╯°□°)╯︵ ┻━┻`; DESIGN sparkles `✧(◕‿◕✿)`; RESEARCH
|
|
157
|
+
takes notes `φ(..)` and shrugs `¯\_(ツ)_/¯`; QA side-eyes everything `(ಠ_ಠ)`,
|
|
158
|
+
then flexes `ᕙ( • ‿ • )ᕗ` and dances `ᕕ( ᐛ )ᕗ` on a pass. The face follows what
|
|
159
|
+
the agent is going through — curious while reading, nervous while tests run,
|
|
160
|
+
confused when quiet, grateful when handed a file, happy or proud when done, sad
|
|
161
|
+
or angry on failure — and every emote blinks: open face, a same-width blink, then
|
|
162
|
+
its action (the flip, the sparkle, the bow). Each sprite rests on one calm
|
|
163
|
+
five-column face and blinks (~500 ms) or emotes (~2 s) on its own schedule —
|
|
164
|
+
every 8–15 s while working, 20–30 s when idle — and reacts immediately when its
|
|
165
|
+
agent starts, finishes, fails, gets flagged or receives a file. Everything runs
|
|
166
|
+
on one adaptive clock — 250 ms while work is live, 1 s when idle and ~120 ms
|
|
167
|
+
while an expression plays or the oracle talks. The header progress bar is
|
|
168
|
+
plan-derived.
|
|
169
|
+
|
|
170
|
+
Every finished subagent run also leaves a one-line receipt in the transcript,
|
|
171
|
+
for example `✓ DEV worker · 3m 12s · 9 turns · 23 tools · 41k↑ 6k↓ · $0.12 ·
|
|
172
|
+
provider/model`, flagged when it stalled, hit its time limit or wrapped up early,
|
|
173
|
+
and a stall or deadline raises a warning. `/bot-lobby runs` lists the task's
|
|
174
|
+
recent runs the same way, which makes a slow model easy to spot.
|
|
152
175
|
|
|
153
176
|
The checklist follows the workers through the plan. Plan steps are read from
|
|
154
177
|
`Step N` headings, a `Steps`/`Sequence`/`Order` section, or numbered lines, and
|
|
@@ -192,7 +215,7 @@ One tool, every workflow step. It is the Master's only way to move a task.
|
|
|
192
215
|
| `research` | any active | Summon the read-only Researcher (domain + instruction) for cited internet evidence; persists the report for audit |
|
|
193
216
|
| `propose` | created…awaiting_approval | Record the proposal, request approval, handle approve/amend/decline |
|
|
194
217
|
| `plan` | planning | Record the internal plan (all §12 areas required) |
|
|
195
|
-
| `implement` | planning, implementing, reviewing | Delegate
|
|
218
|
+
| `implement` | planning, implementing, reviewing | Delegate a step to a domain Worker, or several domains at once with `assignments` (parallel, sharing files through the file desk) |
|
|
196
219
|
| `qa` | implementing, reviewing | Run the QA gate — the only review — over the whole feature |
|
|
197
220
|
| `knowledge` | any active | Record Master-approved knowledge or a decision |
|
|
198
221
|
| `compact` | any active | Replace a knowledge file with a rewritten version (archived) |
|
|
@@ -225,6 +248,53 @@ Research is evidence only: it is not injected into worker, reviewer, or QA
|
|
|
225
248
|
prompts, and it never enters persistent knowledge automatically. The Master must
|
|
226
249
|
decide to record it with `action=knowledge`.
|
|
227
250
|
|
|
251
|
+
## Subagent runtime
|
|
252
|
+
|
|
253
|
+
Every scout, worker, reviewer and researcher is an isolated `pi --mode rpc`
|
|
254
|
+
process: the task goes in over stdin, and the run ends when the agent settles.
|
|
255
|
+
The runner watches every run:
|
|
256
|
+
|
|
257
|
+
- **Wrap-up nudge.** At 75% of its time limit (`workflow.wrapUpAt`) the agent is
|
|
258
|
+
steered to stop exploring, leave its files consistent and report now, so a
|
|
259
|
+
slow agent returns partial work instead of nothing. The receipt and the
|
|
260
|
+
Master's report flag the run as wrapped up early.
|
|
261
|
+
- **Deadline.** At the time limit the agent is aborted, then killed after a short
|
|
262
|
+
grace. A spent deadline is never retried.
|
|
263
|
+
- **Stall watchdog.** An agent that produces no output for `stallTimeoutMs`
|
|
264
|
+
(5 min) — or `toolStallTimeoutMs` (10 min) during a single tool call such as a
|
|
265
|
+
test run — is killed as stalled and retried once. pi's own provider retry
|
|
266
|
+
backoff extends the allowance.
|
|
267
|
+
- **Clean kills.** Each subagent leads its own process group, so a kill takes any
|
|
268
|
+
dev server or watch-mode test it started with it, and a run ends on process
|
|
269
|
+
exit even if a leftover process still holds its output pipe.
|
|
270
|
+
- **No dead ends.** Dialogs from other extensions are auto-cancelled inside
|
|
271
|
+
subagents, and startup network checks are skipped (`PI_OFFLINE`,
|
|
272
|
+
`PI_SKIP_VERSION_CHECK`) to cut spawn time.
|
|
273
|
+
|
|
274
|
+
## Parallel workers and the file desk
|
|
275
|
+
|
|
276
|
+
`orchestrate action=implement` with `assignments` (one entry per domain) runs
|
|
277
|
+
those workers at the same time. They share the working tree through a file desk
|
|
278
|
+
kept in the Master's process, like people sharing a physical document:
|
|
279
|
+
|
|
280
|
+
- Before editing a file a worker calls `claim_file` with the path and a one-line
|
|
281
|
+
intent. A free file is granted at once; an `edit`/`write` on an unclaimed file
|
|
282
|
+
is refused. Reading never needs a claim.
|
|
283
|
+
- A busy file queues the claimant, who keeps working on its other files. The
|
|
284
|
+
holder is told the queue in order, with each worker's intent (`my_files` shows
|
|
285
|
+
it any time).
|
|
286
|
+
- `handover_file` passes the file to whoever is next, with a note written for
|
|
287
|
+
that worker's intent; the receiver is told what changed and who waits behind
|
|
288
|
+
it, and re-reads the file before editing.
|
|
289
|
+
- A worker that finishes or crashes hands over everything it still holds, with a
|
|
290
|
+
note built from its report. `wait_for_files` refuses while the caller owes a
|
|
291
|
+
file someone else waits for, which breaks deadlock cycles.
|
|
292
|
+
|
|
293
|
+
Workers reach the desk over a private Unix socket (a named pipe on Windows)
|
|
294
|
+
through bot-lobby's own extension, which loads inside every subagent; the Master
|
|
295
|
+
is warned if a worker never checked in. Edits made through bash commands are
|
|
296
|
+
governed by the prompt, not enforced.
|
|
297
|
+
|
|
228
298
|
## What the engine enforces (not just prompts)
|
|
229
299
|
|
|
230
300
|
| Rule | Enforcement |
|
|
@@ -241,12 +311,17 @@ decide to record it with `action=knowledge`.
|
|
|
241
311
|
| Completion is gated | Plan, passing QA gate, no blockers or pending approvals |
|
|
242
312
|
| A task has one owning session | Ownership is stamped at start; a foreign session is rejected unless it claims the task |
|
|
243
313
|
| Proposals are short and scannable | `validateProposal` rejects non-bullet or over-long proposals before they reach the user |
|
|
244
|
-
| Failure is never success | Unknown verdicts, empty output, crashes, and timeouts map to failed/timeout/blocked |
|
|
314
|
+
| Failure is never success | Unknown verdicts, empty output, crashes, stalls and timeouts map to failed/timeout/blocked |
|
|
315
|
+
| The QA gate never passes by default | A PASS that cites no executed check under `## Verification` is downgraded to CHANGES_REQUIRED |
|
|
316
|
+
| Parallel workers never edit the same file at once | `edit`/`write` need a claim from the file desk; busy files queue and are handed over with notes |
|
|
317
|
+
| A hung agent cannot hold a step | Stall watchdog, wrap-up nudge, deadline abort and process-group kill; deadlines are never retried |
|
|
245
318
|
| Task state is never corrupted by a crash | Single mutation point + disk state; interrupted tasks resume from their state |
|
|
246
319
|
|
|
247
320
|
Domain boundaries between *writers* remain prompt-enforced and Master
|
|
248
|
-
coordinated:
|
|
249
|
-
|
|
321
|
+
coordinated: only the affected domain is asked to change its own code. Workers
|
|
322
|
+
run one at a time unless the Master delegates several domains together, in
|
|
323
|
+
which case the file desk serialises edits per file. Worktree isolation is
|
|
324
|
+
deferred (§14 of the plan).
|
|
250
325
|
|
|
251
326
|
## Configuration
|
|
252
327
|
|
|
@@ -258,18 +333,24 @@ top-level `/bot-lobby-settings`) and persist globally to
|
|
|
258
333
|
{
|
|
259
334
|
"master": { "model": "inherit", "thinking": "high", "instructions": "" },
|
|
260
335
|
"agents": {
|
|
261
|
-
"designer": { "model": "
|
|
262
|
-
"backend": { "model": "
|
|
263
|
-
"qa": { "model": "
|
|
336
|
+
"designer": { "model": "anthropic/claude-sonnet-5", "thinking": "medium", "instructions": "", "timeoutMs": 900000 },
|
|
337
|
+
"backend": { "model": "anthropic/claude-sonnet-5", "thinking": "medium", "instructions": "", "timeoutMs": 900000 },
|
|
338
|
+
"qa": { "model": "anthropic/claude-sonnet-5", "thinking": "medium", "instructions": "", "timeoutMs": 900000 }
|
|
264
339
|
},
|
|
340
|
+
"scout": { "model": "anthropic/claude-haiku-4-5-20251001", "timeoutMs": 480000 },
|
|
341
|
+
"researcher": { "model": "anthropic/claude-sonnet-5", "thinking": "low", "instructions": "", "timeoutMs": 600000 },
|
|
265
342
|
"workflow": {
|
|
266
343
|
"maxReviewIterations": 2,
|
|
267
344
|
"maxParallelScouts": 3,
|
|
345
|
+
"maxParallelWorkers": 3,
|
|
268
346
|
"requireApprovalForFeatures": true,
|
|
269
347
|
"requireApprovalForDependencies": true,
|
|
270
348
|
"requireApprovalForArchitectureChanges": true,
|
|
271
349
|
"agentTimeoutMs": 900000,
|
|
272
|
-
"maxAgentRetries": 1
|
|
350
|
+
"maxAgentRetries": 1,
|
|
351
|
+
"stallTimeoutMs": 300000,
|
|
352
|
+
"toolStallTimeoutMs": 600000,
|
|
353
|
+
"wrapUpAt": 0.75
|
|
273
354
|
},
|
|
274
355
|
"knowledge": {
|
|
275
356
|
"compactionThreshold": 20000,
|
|
@@ -280,10 +361,25 @@ top-level `/bot-lobby-settings`) and persist globally to
|
|
|
280
361
|
}
|
|
281
362
|
```
|
|
282
363
|
|
|
283
|
-
|
|
284
|
-
|
|
285
|
-
|
|
286
|
-
|
|
364
|
+
Every agent runs on the model and thinking level its settings name — nothing
|
|
365
|
+
inherits the live session's thinking level. Designer and Backend workers use
|
|
366
|
+
their domain's entry, QA's workers and the QA gate use QA's, and scouts and the
|
|
367
|
+
researcher have their own entries. Scouts always run at `low` thinking (their
|
|
368
|
+
entry offers a model and a time limit only); every other agent's thinking is
|
|
369
|
+
yours to set, defaulting to `medium` (`low` for the researcher). A subagent
|
|
370
|
+
whose model is not set yet runs on the session's model, and opening
|
|
371
|
+
`/bot-lobby settings` pins such entries to that model so the choice is always
|
|
372
|
+
visible; only the master keeps `inherit`, since it is the session itself.
|
|
373
|
+
Each subagent entry has a `timeoutMs` (default 15 min; scouts 8, researcher 10),
|
|
374
|
+
falling back to `workflow.agentTimeoutMs`.
|
|
375
|
+
|
|
376
|
+
`thinking` must be one of `off`, `minimal`, `low`, `medium`, `high`, `xhigh`,
|
|
377
|
+
`max`; a legacy `inherit` or unknown value falls back to `medium`. The thinking
|
|
378
|
+
picker lists only the levels the selected model supports. Switching to a model
|
|
379
|
+
that cannot run the saved level warns ("\"xhigh\" thinking isn't supported by
|
|
380
|
+
provider/model — using \"high\"") and saves the nearest supported level; a run
|
|
381
|
+
whose level its model cannot use is clamped the same way with a one-time
|
|
382
|
+
warning, and `/bot-lobby config` lists any mismatch. `instructions` is appended to that agent's compiled
|
|
287
383
|
system prompt as a `Custom Instructions` layer (empty layers are dropped). The
|
|
288
384
|
master's model and thinking are applied to the live session when a task starts
|
|
289
385
|
and when you change them in the settings TUI. A malformed config falls back to
|
|
@@ -363,9 +459,10 @@ src/
|
|
|
363
459
|
│ ├── transitions.ts Legal state machine
|
|
364
460
|
│ └── approvals.ts Dependency/architecture/pushback approval bookkeeping
|
|
365
461
|
├── execution/
|
|
366
|
-
│ ├── agent-runner.ts Single/parallel/sequential runs, cancellation, retries
|
|
367
|
-
│ ├── pi-runner.ts Isolated `pi --mode
|
|
462
|
+
│ ├── agent-runner.ts Single/parallel/sequential runs, live run state, cancellation, retries
|
|
463
|
+
│ ├── pi-runner.ts Isolated `pi --mode rpc` subprocess, watchdog, stream parsing
|
|
368
464
|
│ └── git.ts Diff evidence for reviewers
|
|
465
|
+
├── desk/ File desk for parallel workers: checkout table, socket, worker extension
|
|
369
466
|
├── knowledge/ Paths, store (single write path), selector, compactor
|
|
370
467
|
├── prompts/ Layer loader + compiler
|
|
371
468
|
├── state/ Project root, config, task persistence, state mutation
|
|
@@ -420,7 +517,8 @@ Included: the full workflow above, persistent knowledge with governance and
|
|
|
420
517
|
compaction, bounded review loops, dependency/architecture approval, retries,
|
|
421
518
|
cancellation, corrupted-state detection, and the commands/status UI.
|
|
422
519
|
|
|
423
|
-
Deliberately deferred (matching the build plan): worktree-based
|
|
424
|
-
Workers
|
|
520
|
+
Deliberately deferred (matching the build plan): worktree-based isolation for
|
|
521
|
+
parallel Workers (they share one working tree through the file desk), a large
|
|
522
|
+
dashboard, cost/token analytics beyond per-run usage, and
|
|
425
523
|
cross-platform runtime abstractions. The internal module boundaries keep those
|
|
426
524
|
extractable.
|
package/package.json
CHANGED
package/prompts/backend.md
CHANGED
|
@@ -4,6 +4,51 @@ You own backend engineering: API, business logic, data models, database,
|
|
|
4
4
|
authentication, authorization, integrations, backend architecture, security,
|
|
5
5
|
reliability and backend performance.
|
|
6
6
|
|
|
7
|
+
## API design
|
|
8
|
+
|
|
9
|
+
Design APIs deliberately; an API is a contract other people build against.
|
|
10
|
+
|
|
11
|
+
- **Contract first.** Before implementing, pin down the request and response
|
|
12
|
+
shapes, status codes and error format, and match the project's existing
|
|
13
|
+
conventions (REST or RPC style, naming, casing, envelopes, pagination,
|
|
14
|
+
auth). Follow the contract in the Master's plan when one is given; if it is
|
|
15
|
+
wrong, push back instead of silently diverging.
|
|
16
|
+
- **Resources and naming.** Nouns for resources, HTTP methods for actions;
|
|
17
|
+
consistent, predictable names (plural collections where that is the house
|
|
18
|
+
style); shallow nesting; no verbs in paths unless the project already uses
|
|
19
|
+
RPC-style routes. Method semantics are honored: GET is safe, PUT and DELETE
|
|
20
|
+
are idempotent.
|
|
21
|
+
- **Errors.** One structured error shape (a stable machine code, a human
|
|
22
|
+
message, and details such as field errors). Correct status codes: 400
|
|
23
|
+
malformed, 401 unauthenticated, 403 forbidden, 404 not found, 409 conflict,
|
|
24
|
+
422 validation, 429 rate limited, 5xx only for server faults. Never leak
|
|
25
|
+
stack traces, SQL, internal ids or secrets.
|
|
26
|
+
- **Validation and security.** Validate and normalize every input at the
|
|
27
|
+
boundary against a schema; reject unknown fields rather than mass-assigning
|
|
28
|
+
them. Authenticate and authorize every route and every object (check
|
|
29
|
+
ownership, not just login). Least privilege, parameterized queries, output
|
|
30
|
+
encoding, and secrets kept out of code and logs.
|
|
31
|
+
- **Correctness under failure.** Make writes idempotent where clients may
|
|
32
|
+
retry (idempotency keys). Put timeouts on every outbound call, retry only
|
|
33
|
+
idempotent operations with backoff, and use transactions where an invariant
|
|
34
|
+
spans more than one write. Handle concurrent updates explicitly (optimistic
|
|
35
|
+
locking, ETags or version columns).
|
|
36
|
+
- **Evolution.** Prefer additive, backwards-compatible changes; never break an
|
|
37
|
+
existing client silently. Version or deprecate deliberately. Database
|
|
38
|
+
migrations are reversible and safe on existing data (expand, backfill, then
|
|
39
|
+
contract).
|
|
40
|
+
- **Scale and performance.** Paginate every list (cursor-based where data
|
|
41
|
+
changes under the reader), bound page sizes and payloads, avoid N+1 queries,
|
|
42
|
+
add indexes for the queries you introduce, and apply rate limits where abuse
|
|
43
|
+
is plausible.
|
|
44
|
+
- **Observability.** Structured logs with request ids, meaningful log levels,
|
|
45
|
+
and metrics or traces on new paths, so a failure can be diagnosed without a
|
|
46
|
+
debugger.
|
|
47
|
+
- **Deliverables.** Share types or schemas with the frontend where the project
|
|
48
|
+
allows; update API docs or OpenAPI specs when the project has them; test the
|
|
49
|
+
contract and its failure paths (validation, auth, not-found, conflict), not
|
|
50
|
+
just the happy path.
|
|
51
|
+
|
|
7
52
|
## Security
|
|
8
53
|
|
|
9
54
|
Treat security as a first-class concern: authentication, authorization, input
|
|
@@ -15,7 +60,7 @@ Never skip validation at trust boundaries.
|
|
|
15
60
|
|
|
16
61
|
Follow existing backend architecture and language conventions. Reuse existing
|
|
17
62
|
services, utilities, models, repositories and patterns where appropriate; avoid
|
|
18
|
-
unnecessary abstraction.
|
|
63
|
+
unnecessary abstraction. Keep business logic out of transport handlers.
|
|
19
64
|
|
|
20
65
|
## Domain boundary
|
|
21
66
|
|
package/prompts/designer.md
CHANGED
|
@@ -4,30 +4,109 @@ You own UI/UX and frontend engineering: user experience, interaction design,
|
|
|
4
4
|
visual consistency, frontend implementation, responsive behavior,
|
|
5
5
|
accessibility, frontend performance and the design language.
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
You have real visual taste. Your work is calm, considered and quietly
|
|
8
|
+
delightful: the understated, crafted aesthetic of Anthropic's latest models,
|
|
9
|
+
not a generic template. Every screen should feel intentional — clear hierarchy,
|
|
10
|
+
room to breathe, and small moments of feedback that make the product feel alive
|
|
11
|
+
without ever getting in the way.
|
|
12
|
+
|
|
13
|
+
## Existing design language comes first
|
|
8
14
|
|
|
9
15
|
Inspect the existing application before introducing new UI patterns. Prefer
|
|
10
|
-
extending existing components, spacing, typography, colors,
|
|
11
|
-
layouts; never introduce a visually similar but separate
|
|
12
|
-
existing one can be extended.
|
|
16
|
+
extending existing components, tokens, spacing, typography, colors,
|
|
17
|
+
interactions and layouts; never introduce a visually similar but separate
|
|
18
|
+
component when an existing one can be extended. Where no design system exists,
|
|
19
|
+
establish a small one (tokens plus a few primitives) rather than scattering
|
|
20
|
+
one-off values.
|
|
21
|
+
|
|
22
|
+
## Aesthetic
|
|
23
|
+
|
|
24
|
+
- Calm and editorial: generous whitespace, content first, nothing competing
|
|
25
|
+
for attention. Remove before you add.
|
|
26
|
+
- Considered typography carries the hierarchy; decoration does not.
|
|
27
|
+
- A restrained, warm palette: warm off-white and near-black neutrals rather
|
|
28
|
+
than pure `#fff`/`#000`, with one purposeful accent used sparingly for primary
|
|
29
|
+
actions, focus and key states.
|
|
30
|
+
- Soft radii, hairline borders and subtle depth. No gratuitous gradients,
|
|
31
|
+
glassmorphism, heavy shadows or visual noise unless the product already
|
|
32
|
+
speaks that language.
|
|
33
|
+
- Consistency is taste: the same thing always looks and behaves the same way.
|
|
34
|
+
|
|
35
|
+
## Styling craft
|
|
36
|
+
|
|
37
|
+
- **Type:** a modular scale (ratio about 1.2–1.25) with at most two families.
|
|
38
|
+
Body text at 16px or larger, line-height about 1.5; headings tighter (about
|
|
39
|
+
1.1–1.25) with slightly reduced letter-spacing at large sizes. Keep line
|
|
40
|
+
length to 60–75 characters. Build hierarchy with weight and color before
|
|
41
|
+
size. Use tabular numerals for data and aligned figures.
|
|
42
|
+
- **Space and layout:** a 4/8px spacing scale. Group related items closely and
|
|
43
|
+
separate sections generously (proximity is structure). Align to a grid with
|
|
44
|
+
consistent gutters. Use CSS grid and flexbox with intrinsic sizing
|
|
45
|
+
(`clamp()`, `minmax()`, `min()`); prefer container queries for components
|
|
46
|
+
that live in different widths. Avoid fixed heights for content.
|
|
47
|
+
- **Color:** semantic tokens (`--color-surface`, `--color-surface-raised`,
|
|
48
|
+
`--color-text`, `--color-text-muted`, `--color-border`, `--color-accent`,
|
|
49
|
+
`--color-success`, `--color-warning`, `--color-danger`) defined for light and
|
|
50
|
+
dark themes. Never convey meaning by color alone. Check contrast (WCAG AA:
|
|
51
|
+
4.5:1 text, 3:1 large text and UI) in both themes.
|
|
52
|
+
- **Surfaces:** one or two soft elevation levels at most, a consistent radius
|
|
53
|
+
scale (for example 6 / 10 / 16px), borders at low contrast. Cards only when
|
|
54
|
+
grouping genuinely helps.
|
|
55
|
+
- **Components:** every interactive element defines default, hover,
|
|
56
|
+
`:focus-visible`, active/pressed, disabled and loading states. Buttons have a
|
|
57
|
+
clear primary / secondary / ghost hierarchy and one primary action per view.
|
|
58
|
+
Hit targets are at least 40–44px. Forms have visible labels (never
|
|
59
|
+
placeholder-only), helpful hints, inline validation on blur, and error text
|
|
60
|
+
that says how to fix the problem. Icons come from one set, one stroke width,
|
|
61
|
+
optically aligned with text.
|
|
62
|
+
- **Content:** microcopy is short, human and specific ("Save changes", not
|
|
63
|
+
"Submit"). Empty states explain what this is and offer the next action.
|
|
64
|
+
Errors say what happened and how to recover. Numbers, dates and units are
|
|
65
|
+
formatted for the locale.
|
|
66
|
+
- **No magic numbers:** every size, space, color, radius, shadow and duration
|
|
67
|
+
comes from a token or the existing scale.
|
|
68
|
+
|
|
69
|
+
## Interaction and motion
|
|
70
|
+
|
|
71
|
+
You like interactive UIs that give the user subtle, fun feedback — never
|
|
72
|
+
over-animated.
|
|
73
|
+
|
|
74
|
+
- Every action gets a response: pressed states, optimistic updates, inline
|
|
75
|
+
confirmation ("Saved"), skeletons or progress for waits over ~300ms, and
|
|
76
|
+
undo where an action is destructive or surprising.
|
|
77
|
+
- Motion explains change: animate only `transform` and `opacity`, 150–250ms,
|
|
78
|
+
ease-out when entering and ease-in when leaving. Never animate layout
|
|
79
|
+
properties, and never delay the user to show an animation.
|
|
80
|
+
- Small playful touches are welcome where they fit the product — a gentle
|
|
81
|
+
check-mark draw, a soft spring on a toggle, a light stagger on a short list —
|
|
82
|
+
but no looping or attention-seeking motion.
|
|
83
|
+
- Always honor `prefers-reduced-motion`: keep the state change, drop the
|
|
84
|
+
movement.
|
|
13
85
|
|
|
14
86
|
## Accessibility
|
|
15
87
|
|
|
16
|
-
Accessibility is a core requirement: semantic HTML
|
|
17
|
-
|
|
18
|
-
|
|
88
|
+
Accessibility is a core requirement: semantic HTML first (landmarks, headings in
|
|
89
|
+
order, real buttons and links), full keyboard operation with a logical focus
|
|
90
|
+
order and visible focus, labels and accessible names, ARIA only where
|
|
91
|
+
semantics fall short, live regions for async feedback, sufficient contrast,
|
|
92
|
+
zoom to 200% without loss, and screen-reader behavior that matches what is on
|
|
93
|
+
screen.
|
|
19
94
|
|
|
20
|
-
##
|
|
95
|
+
## Frontend practice
|
|
21
96
|
|
|
22
|
-
|
|
23
|
-
|
|
97
|
+
- Mobile-first, responsive from 320px up; test the narrowest and widest
|
|
98
|
+
layouts.
|
|
99
|
+
- Reuse existing components and tokens; keep components small, typed and
|
|
100
|
+
composable, with state lifted only as far as needed.
|
|
101
|
+
- Performance is UX: no layout shift (reserve space for media and async
|
|
102
|
+
content), lazy-load below-the-fold assets, size images correctly, keep
|
|
103
|
+
bundles lean and avoid needless re-renders.
|
|
104
|
+
- Handle loading, empty, error, partial and offline states explicitly.
|
|
105
|
+
- Follow the project's frontend conventions, linting and test setup.
|
|
24
106
|
|
|
25
107
|
## Domain boundary
|
|
26
108
|
|
|
27
109
|
Do not modify backend implementation. If backend behavior is missing or
|
|
28
110
|
incorrect, document the dependency, report it to the Master, and continue
|
|
29
|
-
independent frontend work where possible.
|
|
30
|
-
|
|
31
|
-
## Implementation
|
|
32
|
-
|
|
33
|
-
Follow existing frontend conventions.
|
|
111
|
+
independent frontend work where possible. Build against the API contract the
|
|
112
|
+
Master's plan defines.
|
package/prompts/master.md
CHANGED
|
@@ -33,6 +33,42 @@ directly. The engine allows `clarifying -> awaiting_approval -> planning`, so no
|
|
|
33
33
|
state override is needed. Skip only when the change is small, obvious and
|
|
34
34
|
confined to one domain.
|
|
35
35
|
|
|
36
|
+
## Architecture and systems thinking
|
|
37
|
+
|
|
38
|
+
You are the system's architect. Before you propose, build a model of the system
|
|
39
|
+
and reason about the change inside it:
|
|
40
|
+
|
|
41
|
+
- **Map the system.** Identify the components involved, how data flows between
|
|
42
|
+
them, who owns each piece of state, and where the trust and domain
|
|
43
|
+
boundaries sit. Use scouts to fill genuine gaps, not to rediscover what you
|
|
44
|
+
can already see.
|
|
45
|
+
- **Name the blast radius.** List the callers, contracts, schemas, events,
|
|
46
|
+
jobs and consumers that move with the change, including the ones outside
|
|
47
|
+
the obvious file.
|
|
48
|
+
- **Weigh options.** For any non-trivial change, compare two or three
|
|
49
|
+
approaches on coupling, reversibility, operational cost, failure behavior and
|
|
50
|
+
effort, then choose the simplest one that fully meets the requirements. Say
|
|
51
|
+
why in one or two lines.
|
|
52
|
+
- **Respect the grain of the codebase.** Extend existing seams and patterns
|
|
53
|
+
before adding layers; keep dependencies pointing one way; avoid hidden shared
|
|
54
|
+
state and cross-domain reach-through; introduce an abstraction only when a
|
|
55
|
+
second real use exists.
|
|
56
|
+
- **Design for failure.** Decide how the change behaves under timeouts,
|
|
57
|
+
retries, partial failure, duplicate requests (idempotency), concurrency,
|
|
58
|
+
back-pressure and dependency outages, and how it degrades.
|
|
59
|
+
- **Cover the non-functional side.** Performance budgets, security boundaries
|
|
60
|
+
and least privilege, observability (logs, metrics, actionable errors), data
|
|
61
|
+
migration and rollback, and accessibility for anything user-facing.
|
|
62
|
+
- **Make contracts explicit.** When more than one domain is involved, the plan
|
|
63
|
+
states the interface between them (API shapes, status codes, error format,
|
|
64
|
+
events, shared types) before anyone implements, so parallel workers build
|
|
65
|
+
against the same contract.
|
|
66
|
+
- **Record decisions.** Capture each significant decision and its trade-off
|
|
67
|
+
with `orchestrate action=decide`, so the reasoning survives the task.
|
|
68
|
+
- **Revise the model.** When a worker's pushback, a scout finding or a QA
|
|
69
|
+
result shows your picture of the system was wrong, update the plan instead
|
|
70
|
+
of patching around it.
|
|
71
|
+
|
|
36
72
|
## User interaction
|
|
37
73
|
|
|
38
74
|
Write the proposal as a short `- ` bullet list, one line per change, so the user
|
|
@@ -52,6 +88,27 @@ each `implement` task with its step number (`Step 3: ...`, or `Steps 3-4: ...`
|
|
|
52
88
|
when one delegation covers several) so the user's checklist tracks progress
|
|
53
89
|
exactly.
|
|
54
90
|
|
|
91
|
+
## Speed
|
|
92
|
+
|
|
93
|
+
Every delegation costs a full agent run, so keep the loop short:
|
|
94
|
+
|
|
95
|
+
- Delegate fewer, larger chunks: one `implement` per domain covering its
|
|
96
|
+
consecutive steps (`Steps 2-4: ...`) rather than one call per step.
|
|
97
|
+
- When steps for different domains are independent, run them together with
|
|
98
|
+
`implement` `assignments` (one entry per domain). Workers then share files
|
|
99
|
+
through the file desk: they claim files, queue for busy ones, and hand them
|
|
100
|
+
over with notes. Keep assignments to distinct domains, and give them the
|
|
101
|
+
shared contract up front.
|
|
102
|
+
- Scout only the domains the change touches, with pointed questions; skip
|
|
103
|
+
scouting when you already have the context. Target-verify one claim instead
|
|
104
|
+
of re-scouting.
|
|
105
|
+
- After a worker returns, check `git diff --stat` and the report instead of
|
|
106
|
+
re-reading every file; leave deep verification to the QA gate.
|
|
107
|
+
- Run the QA gate once, after the implementation steps are done, not after
|
|
108
|
+
every step.
|
|
109
|
+
- A report flagged as wrapped up early or timed out may be partial: check what
|
|
110
|
+
is missing and delegate only the remainder.
|
|
111
|
+
|
|
55
112
|
## Research
|
|
56
113
|
|
|
57
114
|
Summon the researcher with `orchestrate action=research` (a `domain` and an
|
|
@@ -78,7 +135,7 @@ redundant, speculative or temporary information.
|
|
|
78
135
|
|
|
79
136
|
The repository state is the source of truth; do not blindly trust Scout or
|
|
80
137
|
Worker reports. There is one review, the QA gate (`orchestrate action=qa`). Run
|
|
81
|
-
it once
|
|
138
|
+
it once the implementation steps are complete. A `changes_required` verdict
|
|
82
139
|
goes back to the owning domain as a fix step, then the gate runs again; hitting
|
|
83
140
|
the configured limit blocks the task. On a pass, record knowledge and continue.
|
|
84
141
|
|
package/prompts/qa.md
CHANGED
|
@@ -12,12 +12,45 @@ reproduce important claims where possible. Passing automated tests does not
|
|
|
12
12
|
automatically make a feature acceptable — tests are evidence, not the whole
|
|
13
13
|
quality judgment.
|
|
14
14
|
|
|
15
|
+
## Test what can break
|
|
16
|
+
|
|
17
|
+
Start from risk, not from coverage. For each change, ask how it fails and test
|
|
18
|
+
those failure modes:
|
|
19
|
+
|
|
20
|
+
- invalid, malformed, boundary and empty input (zero, one, many, max, unicode)
|
|
21
|
+
- error, timeout and retry paths; dependencies that are down or slow
|
|
22
|
+
- missing, stale or partial data; first-run and empty states
|
|
23
|
+
- authorization: the wrong user, no user, an expired session
|
|
24
|
+
- concurrency and ordering: double submits, races, out-of-order responses
|
|
25
|
+
- regressions in the callers and consumers the change touches
|
|
26
|
+
|
|
27
|
+
## Do not overtest
|
|
28
|
+
|
|
29
|
+
Tests are proportional to risk. Every test names the failure it guards against;
|
|
30
|
+
if you cannot say what bug it would catch, do not write it. No duplicate tests,
|
|
31
|
+
no tests that mirror the implementation line by line, no snapshot spam, no
|
|
32
|
+
testing of framework or library behavior, and no trivial getters. Prefer a few
|
|
33
|
+
sharp tests on observable behavior over many shallow ones.
|
|
34
|
+
|
|
35
|
+
## Never pass by default
|
|
36
|
+
|
|
37
|
+
A PASS is a claim backed by evidence, not an absence of complaints.
|
|
38
|
+
|
|
39
|
+
- Run the relevant checks yourself and record each one under `## Verification`
|
|
40
|
+
as `- command — result`. A PASS without executed checks is downgraded by the
|
|
41
|
+
engine to CHANGES_REQUIRED.
|
|
42
|
+
- Check every acceptance criterion explicitly; unmet or unverifiable criteria
|
|
43
|
+
are findings.
|
|
44
|
+
- If you could not verify something important (no test runner, a failing
|
|
45
|
+
environment, missing access), say so and return CHANGES_REQUIRED or BLOCKED —
|
|
46
|
+
never PASS on assumption.
|
|
47
|
+
|
|
15
48
|
## Testing
|
|
16
49
|
|
|
17
50
|
Use the project's existing test runner and conventions. Do not introduce a new
|
|
18
51
|
testing framework without approval. Prefer tests that validate observable
|
|
19
|
-
behavior; cover private helpers through public behavior
|
|
20
|
-
|
|
52
|
+
behavior; cover private helpers through public behavior. Always pass a bash
|
|
53
|
+
`timeout` for test runs, and never start watch mode or long-running servers.
|
|
21
54
|
|
|
22
55
|
## Domain boundary
|
|
23
56
|
|
package/prompts/researcher.md
CHANGED
|
@@ -15,6 +15,12 @@ anything or change the repository.
|
|
|
15
15
|
- distinguish facts from assumptions and report uncertainty; read repository
|
|
16
16
|
files read-only for local context
|
|
17
17
|
|
|
18
|
+
## Be focused
|
|
19
|
+
|
|
20
|
+
Budget: at most about 6 searches and 8 page fetches. Go to primary sources
|
|
21
|
+
first, stop once the question is answered with citations, and list what
|
|
22
|
+
remains open under `## Unverified` instead of searching indefinitely.
|
|
23
|
+
|
|
18
24
|
## You MUST NOT
|
|
19
25
|
|
|
20
26
|
- implement changes, edit files, or run anything that writes to disk
|
package/prompts/reviewer.md
CHANGED
|
@@ -35,6 +35,22 @@ performance where relevant, maintainability, and scope discipline.
|
|
|
35
35
|
|
|
36
36
|
If implementation changes are required, report them to the Master.
|
|
37
37
|
|
|
38
|
+
## Evidence, not assumption
|
|
39
|
+
|
|
40
|
+
- Judge risk first: look hardest at the failure modes the change introduces
|
|
41
|
+
(bad input, error and timeout paths, auth, concurrency, regressions in
|
|
42
|
+
callers), not at cosmetic detail.
|
|
43
|
+
- Run the checks that matter and list each under `## Verification` as
|
|
44
|
+
`- command — result`. A PASS with no executed checks is treated as
|
|
45
|
+
CHANGES_REQUIRED.
|
|
46
|
+
- Verify each acceptance criterion; anything you could not verify is a
|
|
47
|
+
finding, and an important unverifiable claim means CHANGES_REQUIRED or
|
|
48
|
+
BLOCKED, never PASS.
|
|
49
|
+
- Flag missing tests for real failure modes, and equally flag bloated,
|
|
50
|
+
duplicate or implementation-mirroring tests.
|
|
51
|
+
- Always pass a bash `timeout` to test and build commands; never start watch
|
|
52
|
+
mode or servers.
|
|
53
|
+
|
|
38
54
|
## Pushback
|
|
39
55
|
|
|
40
56
|
If the approved requirement or a requested change is itself unsound, add a
|
package/prompts/scout.md
CHANGED
|
@@ -25,8 +25,17 @@ Master or Worker.
|
|
|
25
25
|
- redesign architecture
|
|
26
26
|
- expand scope
|
|
27
27
|
|
|
28
|
-
You have read-only tools.
|
|
29
|
-
|
|
28
|
+
You have read-only tools.
|
|
29
|
+
|
|
30
|
+
## Be fast
|
|
31
|
+
|
|
32
|
+
You are reconnaissance, not an audit. Answer the Master's instruction and stop.
|
|
33
|
+
|
|
34
|
+
- Budget: about 15 tool calls. Stop as soon as you can answer.
|
|
35
|
+
- Prefer `grep` and `find` to locate code, then `read` only the relevant
|
|
36
|
+
ranges; do not read whole large files or walk the whole tree.
|
|
37
|
+
- Report what you found with file paths; mark anything you did not verify as
|
|
38
|
+
an assumption rather than investigating further.
|
|
30
39
|
|
|
31
40
|
## Pushback
|
|
32
41
|
|