@a-t-h-i/bot-lobby 0.2.0 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -83,8 +83,9 @@ structure.
83
83
  /bot-lobby amend <text> Record an amendment; the Master re-proposes
84
84
  /bot-lobby decline Decline the proposal and abandon the task
85
85
  /bot-lobby knowledge Knowledge file sizes vs. the compaction threshold
86
+ /bot-lobby runs [taskId] Recent subagent runs: time, turns, tools, tokens, cost, model
86
87
  /bot-lobby config Effective configuration and its file path
87
- /bot-lobby settings Edit per-agent model, thinking, and instructions
88
+ /bot-lobby settings Edit each agent's model, thinking, time limit and instructions
88
89
  /bot-lobby-settings Same as the settings subcommand
89
90
  /bot-lobby minimize|restore Hide or restore bot-lobby for this session (ctrl+shift+m)
90
91
  /bot-lobby claim <taskId> Take ownership of an orphaned task
@@ -118,8 +119,16 @@ and the built-in spinner are hidden, and a widget above the editor animates the
118
119
  task. At 72 columns and wider it draws a large scene: a header box with the task
119
120
  title and state in its top border, a progress bar, and a metadata row with
120
121
  elapsed time, quiet-mode hint and task id; an oracle tower with a twinkling
121
- aura (drifting z's while dormant), a radiant orb crown, window eyes, a
122
- seven-column mouth and its ORC door; beside the crown, the oracle's speech
122
+ aura (drifting z's while dormant), a radiant orb crown, two window eyes, a
123
+ seven-column mouth and its ORC door. Its pupils look around: they move
124
+ left, centre or right in each window and turn up `◓`, ahead `◉` or down `◒`.
125
+ While agents work it looks down at them, taking turns between them; while it
126
+ talks it looks at you; otherwise its eyes wander the room and keep returning
127
+ to you. It keeps a straight, serious face (`───`), tightening into a frown
128
+ only when the task is blocked, and its mouth lip-syncs as a voice waveform
129
+ for ~2 s whenever it says something new. Every 6–12 s it blinks (lids
130
+ stepping down and up) or scans the whole room; asleep, it peeks one eye open
131
+ or snores. Beside the crown, the oracle's speech
123
132
  bubble, its tail on the orb, says what the master is doing (`⠋ delegating`,
124
133
  `⠋ thinking`, `· your turn` once its turn ends, `· dormant` when paused) above
125
134
  who is at work (`→ DEV · QA`), the current step (`step 3 of 7`) or the task
@@ -128,19 +137,41 @@ four animated slots — DEV, DESIGN, RESEARCH and QA — each with a status face
128
137
  caption and two status rows: while running, a braille spinner beside the agent's
129
138
  live one-word activity (for example `⠋ reading` or `⠋ editing`) with its elapsed
130
139
  time on the row beneath; otherwise the coloured status glyph and state word over
131
- that elapsed time; and a full-width TASKS checklist windowed on the current
132
- step. Narrower
133
- terminals keep the
134
- boxed banner, header and compact animated strip, whose working line names the
135
- newest running agent's activity and elapsed time. Each sprite rests on one calm
136
- face and, independently every 20–30 s, briefly blinks (~500 ms) or emotes
137
- (~2 s, stepping through its kaomoji frames); the oracle's mouth moves with its
138
- adaptive clock — 250 ms while work is live, 1 s when idle and ~120 ms while an
139
- expression plays — and their faces, colours and words follow each agent's status
140
- (working, idle, done, failed). The large scene's rest and blink frames stay the
141
- five-column ASCII eyes, while its emote frames are status-aware kaomoji: nervous
142
- while working, happy when done (QA flexes and dances), scared on failure. The
143
- header progress bar is plan-derived.
140
+ that elapsed time. A working agent's status row turns into a warning when it is
141
+ waiting on a file (`⧗ waiting`), has gone quiet (`! quiet 1m`) or is retrying a
142
+ provider call (`↻ retrying`). Under the agents, a live feed row says exactly what
143
+ one of them is doing — `▸ DEV editing users.ts · turn 4 · 12 tools · 41k tok` —
144
+ rotating between working agents every few seconds and putting warnings first
145
+ (gone quiet, waiting on a file, asked to wrap up); once nothing runs it shows the
146
+ last run's receipt. It only takes a spare line, so it never costs the tower, the
147
+ agents or a checklist row. A full-width TASKS checklist windowed on the current
148
+ step closes the scene. Narrower terminals keep the boxed banner, header and
149
+ compact animated strip, whose working line names the newest running agent's
150
+ activity, its target and elapsed time.
151
+
152
+ Each agent has a kaomoji personality. About 300 faces across 15 emotions (happy,
153
+ proud, love, excited, focused, curious, thinking, nervous, confused, sleepy, sad,
154
+ angry, waiting, surprised, grateful) come from a shared pool every agent can use
155
+ plus each agent's own set of at least four faces per emotion: DEV wears shades,
156
+ flexes and flips tables `(╯°□°)╯︵ ┻━┻`; DESIGN sparkles `✧(◕‿◕✿)`; RESEARCH
157
+ takes notes `φ(..)` and shrugs `¯\_(ツ)_/¯`; QA side-eyes everything `(ಠ_ಠ)`,
158
+ then flexes `ᕙ( • ‿ • )ᕗ` and dances `ᕕ( ᐛ )ᕗ` on a pass. The face follows what
159
+ the agent is going through — curious while reading, nervous while tests run,
160
+ confused when quiet, grateful when handed a file, happy or proud when done, sad
161
+ or angry on failure — and every emote blinks: open face, a same-width blink, then
162
+ its action (the flip, the sparkle, the bow). Each sprite rests on one calm
163
+ five-column face and blinks (~500 ms) or emotes (~2 s) on its own schedule —
164
+ every 8–15 s while working, 20–30 s when idle — and reacts immediately when its
165
+ agent starts, finishes, fails, gets flagged or receives a file. Everything runs
166
+ on one adaptive clock — 250 ms while work is live, 1 s when idle and ~120 ms
167
+ while an expression plays or the oracle talks. The header progress bar is
168
+ plan-derived.
169
+
170
+ Every finished subagent run also leaves a one-line receipt in the transcript,
171
+ for example `✓ DEV worker · 3m 12s · 9 turns · 23 tools · 41k↑ 6k↓ · $0.12 ·
172
+ provider/model`, flagged when it stalled, hit its time limit or wrapped up early,
173
+ and a stall or deadline raises a warning. `/bot-lobby runs` lists the task's
174
+ recent runs the same way, which makes a slow model easy to spot.
144
175
 
145
176
  The checklist follows the workers through the plan. Plan steps are read from
146
177
  `Step N` headings, a `Steps`/`Sequence`/`Order` section, or numbered lines, and
@@ -184,7 +215,7 @@ One tool, every workflow step. It is the Master's only way to move a task.
184
215
  | `research` | any active | Summon the read-only Researcher (domain + instruction) for cited internet evidence; persists the report for audit |
185
216
  | `propose` | created…awaiting_approval | Record the proposal, request approval, handle approve/amend/decline |
186
217
  | `plan` | planning | Record the internal plan (all §12 areas required) |
187
- | `implement` | planning, implementing, reviewing | Delegate one step to a domain Worker |
218
+ | `implement` | planning, implementing, reviewing | Delegate a step to a domain Worker, or several domains at once with `assignments` (parallel, sharing files through the file desk) |
188
219
  | `qa` | implementing, reviewing | Run the QA gate — the only review — over the whole feature |
189
220
  | `knowledge` | any active | Record Master-approved knowledge or a decision |
190
221
  | `compact` | any active | Replace a knowledge file with a rewritten version (archived) |
@@ -217,6 +248,53 @@ Research is evidence only: it is not injected into worker, reviewer, or QA
217
248
  prompts, and it never enters persistent knowledge automatically. The Master must
218
249
  decide to record it with `action=knowledge`.
219
250
 
251
+ ## Subagent runtime
252
+
253
+ Every scout, worker, reviewer and researcher is an isolated `pi --mode rpc`
254
+ process: the task goes in over stdin, and the run ends when the agent settles.
255
+ The runner watches every run:
256
+
257
+ - **Wrap-up nudge.** At 75% of its time limit (`workflow.wrapUpAt`) the agent is
258
+ steered to stop exploring, leave its files consistent and report now, so a
259
+ slow agent returns partial work instead of nothing. The receipt and the
260
+ Master's report flag the run as wrapped up early.
261
+ - **Deadline.** At the time limit the agent is aborted, then killed after a short
262
+ grace. A spent deadline is never retried.
263
+ - **Stall watchdog.** An agent that produces no output for `stallTimeoutMs`
264
+ (5 min) — or `toolStallTimeoutMs` (10 min) during a single tool call such as a
265
+ test run — is killed as stalled and retried once. pi's own provider retry
266
+ backoff extends the allowance.
267
+ - **Clean kills.** Each subagent leads its own process group, so a kill takes any
268
+ dev server or watch-mode test it started with it, and a run ends on process
269
+ exit even if a leftover process still holds its output pipe.
270
+ - **No dead ends.** Dialogs from other extensions are auto-cancelled inside
271
+ subagents, and startup network checks are skipped (`PI_OFFLINE`,
272
+ `PI_SKIP_VERSION_CHECK`) to cut spawn time.
273
+
274
+ ## Parallel workers and the file desk
275
+
276
+ `orchestrate action=implement` with `assignments` (one entry per domain) runs
277
+ those workers at the same time. They share the working tree through a file desk
278
+ kept in the Master's process, like people sharing a physical document:
279
+
280
+ - Before editing a file a worker calls `claim_file` with the path and a one-line
281
+ intent. A free file is granted at once; an `edit`/`write` on an unclaimed file
282
+ is refused. Reading never needs a claim.
283
+ - A busy file queues the claimant, who keeps working on its other files. The
284
+ holder is told the queue in order, with each worker's intent (`my_files` shows
285
+ it any time).
286
+ - `handover_file` passes the file to whoever is next, with a note written for
287
+ that worker's intent; the receiver is told what changed and who waits behind
288
+ it, and re-reads the file before editing.
289
+ - A worker that finishes or crashes hands over everything it still holds, with a
290
+ note built from its report. `wait_for_files` refuses while the caller owes a
291
+ file someone else waits for, which breaks deadlock cycles.
292
+
293
+ Workers reach the desk over a private Unix socket (a named pipe on Windows)
294
+ through bot-lobby's own extension, which loads inside every subagent; the Master
295
+ is warned if a worker never checked in. Edits made through bash commands are
296
+ governed by the prompt, not enforced.
297
+
220
298
  ## What the engine enforces (not just prompts)
221
299
 
222
300
  | Rule | Enforcement |
@@ -233,12 +311,17 @@ decide to record it with `action=knowledge`.
233
311
  | Completion is gated | Plan, passing QA gate, no blockers or pending approvals |
234
312
  | A task has one owning session | Ownership is stamped at start; a foreign session is rejected unless it claims the task |
235
313
  | Proposals are short and scannable | `validateProposal` rejects non-bullet or over-long proposals before they reach the user |
236
- | Failure is never success | Unknown verdicts, empty output, crashes, and timeouts map to failed/timeout/blocked |
314
+ | Failure is never success | Unknown verdicts, empty output, crashes, stalls and timeouts map to failed/timeout/blocked |
315
+ | The QA gate never passes by default | A PASS that cites no executed check under `## Verification` is downgraded to CHANGES_REQUIRED |
316
+ | Parallel workers never edit the same file at once | `edit`/`write` need a claim from the file desk; busy files queue and are handed over with notes |
317
+ | A hung agent cannot hold a step | Stall watchdog, wrap-up nudge, deadline abort and process-group kill; deadlines are never retried |
237
318
  | Task state is never corrupted by a crash | Single mutation point + disk state; interrupted tasks resume from their state |
238
319
 
239
320
  Domain boundaries between *writers* remain prompt-enforced and Master
240
- coordinated: Workers run sequentially and only the affected domain is asked to
241
- change its own code. Worktree isolation is deferred (§14 of the plan).
321
+ coordinated: only the affected domain is asked to change its own code. Workers
322
+ run one at a time unless the Master delegates several domains together, in
323
+ which case the file desk serialises edits per file. Worktree isolation is
324
+ deferred (§14 of the plan).
242
325
 
243
326
  ## Configuration
244
327
 
@@ -250,18 +333,24 @@ top-level `/bot-lobby-settings`) and persist globally to
250
333
  {
251
334
  "master": { "model": "inherit", "thinking": "high", "instructions": "" },
252
335
  "agents": {
253
- "designer": { "model": "inherit", "thinking": "medium", "instructions": "" },
254
- "backend": { "model": "inherit", "thinking": "medium", "instructions": "" },
255
- "qa": { "model": "inherit", "thinking": "high", "instructions": "" }
336
+ "designer": { "model": "anthropic/claude-sonnet-5", "thinking": "medium", "instructions": "", "timeoutMs": 900000 },
337
+ "backend": { "model": "anthropic/claude-sonnet-5", "thinking": "medium", "instructions": "", "timeoutMs": 900000 },
338
+ "qa": { "model": "anthropic/claude-sonnet-5", "thinking": "medium", "instructions": "", "timeoutMs": 900000 }
256
339
  },
340
+ "scout": { "model": "anthropic/claude-haiku-4-5-20251001", "timeoutMs": 480000 },
341
+ "researcher": { "model": "anthropic/claude-sonnet-5", "thinking": "low", "instructions": "", "timeoutMs": 600000 },
257
342
  "workflow": {
258
343
  "maxReviewIterations": 2,
259
344
  "maxParallelScouts": 3,
345
+ "maxParallelWorkers": 3,
260
346
  "requireApprovalForFeatures": true,
261
347
  "requireApprovalForDependencies": true,
262
348
  "requireApprovalForArchitectureChanges": true,
263
349
  "agentTimeoutMs": 900000,
264
- "maxAgentRetries": 1
350
+ "maxAgentRetries": 1,
351
+ "stallTimeoutMs": 300000,
352
+ "toolStallTimeoutMs": 600000,
353
+ "wrapUpAt": 0.75
265
354
  },
266
355
  "knowledge": {
267
356
  "compactionThreshold": 20000,
@@ -272,10 +361,25 @@ top-level `/bot-lobby-settings`) and persist globally to
272
361
  }
273
362
  ```
274
363
 
275
- `"model": "inherit"` uses the session's model; any other value is passed to the
276
- subagent as `--model` (e.g. `"anthropic/claude-sonnet-4-5"`). `thinking` must be
277
- one of `off`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max`; an invalid value
278
- falls back to the default. `instructions` is appended to that agent's compiled
364
+ Every agent runs on the model and thinking level its settings name — nothing
365
+ inherits the live session's thinking level. Designer and Backend workers use
366
+ their domain's entry, QA's workers and the QA gate use QA's, and scouts and the
367
+ researcher have their own entries. Scouts always run at `low` thinking (their
368
+ entry offers a model and a time limit only); every other agent's thinking is
369
+ yours to set, defaulting to `medium` (`low` for the researcher). A subagent
370
+ whose model is not set yet runs on the session's model, and opening
371
+ `/bot-lobby settings` pins such entries to that model so the choice is always
372
+ visible; only the master keeps `inherit`, since it is the session itself.
373
+ Each subagent entry has a `timeoutMs` (default 15 min; scouts 8, researcher 10),
374
+ falling back to `workflow.agentTimeoutMs`.
375
+
376
+ `thinking` must be one of `off`, `minimal`, `low`, `medium`, `high`, `xhigh`,
377
+ `max`; a legacy `inherit` or unknown value falls back to `medium`. The thinking
378
+ picker lists only the levels the selected model supports. Switching to a model
379
+ that cannot run the saved level warns ("\"xhigh\" thinking isn't supported by
380
+ provider/model — using \"high\"") and saves the nearest supported level; a run
381
+ whose level its model cannot use is clamped the same way with a one-time
382
+ warning, and `/bot-lobby config` lists any mismatch. `instructions` is appended to that agent's compiled
279
383
  system prompt as a `Custom Instructions` layer (empty layers are dropped). The
280
384
  master's model and thinking are applied to the live session when a task starts
281
385
  and when you change them in the settings TUI. A malformed config falls back to
@@ -355,9 +459,10 @@ src/
355
459
  │ ├── transitions.ts Legal state machine
356
460
  │ └── approvals.ts Dependency/architecture/pushback approval bookkeeping
357
461
  ├── execution/
358
- │ ├── agent-runner.ts Single/parallel/sequential runs, cancellation, retries
359
- │ ├── pi-runner.ts Isolated `pi --mode json` subprocess + stream parsing
462
+ │ ├── agent-runner.ts Single/parallel/sequential runs, live run state, cancellation, retries
463
+ │ ├── pi-runner.ts Isolated `pi --mode rpc` subprocess, watchdog, stream parsing
360
464
  │ └── git.ts Diff evidence for reviewers
465
+ ├── desk/ File desk for parallel workers: checkout table, socket, worker extension
361
466
  ├── knowledge/ Paths, store (single write path), selector, compactor
362
467
  ├── prompts/ Layer loader + compiler
363
468
  ├── state/ Project root, config, task persistence, state mutation
@@ -412,7 +517,8 @@ Included: the full workflow above, persistent knowledge with governance and
412
517
  compaction, bounded review loops, dependency/architecture approval, retries,
413
518
  cancellation, corrupted-state detection, and the commands/status UI.
414
519
 
415
- Deliberately deferred (matching the build plan): worktree-based parallel
416
- Workers, a large dashboard, cost/token analytics beyond per-run usage, and
520
+ Deliberately deferred (matching the build plan): worktree-based isolation for
521
+ parallel Workers (they share one working tree through the file desk), a large
522
+ dashboard, cost/token analytics beyond per-run usage, and
417
523
  cross-platform runtime abstractions. The internal module boundaries keep those
418
524
  extractable.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@a-t-h-i/bot-lobby",
3
- "version": "0.2.0",
3
+ "version": "0.4.0",
4
4
  "description": "Structured multi-agent software engineering orchestrator for Pi",
5
5
  "type": "module",
6
6
  "license": "Apache-2.0",
@@ -4,6 +4,51 @@ You own backend engineering: API, business logic, data models, database,
4
4
  authentication, authorization, integrations, backend architecture, security,
5
5
  reliability and backend performance.
6
6
 
7
+ ## API design
8
+
9
+ Design APIs deliberately; an API is a contract other people build against.
10
+
11
+ - **Contract first.** Before implementing, pin down the request and response
12
+ shapes, status codes and error format, and match the project's existing
13
+ conventions (REST or RPC style, naming, casing, envelopes, pagination,
14
+ auth). Follow the contract in the Master's plan when one is given; if it is
15
+ wrong, push back instead of silently diverging.
16
+ - **Resources and naming.** Nouns for resources, HTTP methods for actions;
17
+ consistent, predictable names (plural collections where that is the house
18
+ style); shallow nesting; no verbs in paths unless the project already uses
19
+ RPC-style routes. Method semantics are honored: GET is safe, PUT and DELETE
20
+ are idempotent.
21
+ - **Errors.** One structured error shape (a stable machine code, a human
22
+ message, and details such as field errors). Correct status codes: 400
23
+ malformed, 401 unauthenticated, 403 forbidden, 404 not found, 409 conflict,
24
+ 422 validation, 429 rate limited, 5xx only for server faults. Never leak
25
+ stack traces, SQL, internal ids or secrets.
26
+ - **Validation and security.** Validate and normalize every input at the
27
+ boundary against a schema; reject unknown fields rather than mass-assigning
28
+ them. Authenticate and authorize every route and every object (check
29
+ ownership, not just login). Least privilege, parameterized queries, output
30
+ encoding, and secrets kept out of code and logs.
31
+ - **Correctness under failure.** Make writes idempotent where clients may
32
+ retry (idempotency keys). Put timeouts on every outbound call, retry only
33
+ idempotent operations with backoff, and use transactions where an invariant
34
+ spans more than one write. Handle concurrent updates explicitly (optimistic
35
+ locking, ETags or version columns).
36
+ - **Evolution.** Prefer additive, backwards-compatible changes; never break an
37
+ existing client silently. Version or deprecate deliberately. Database
38
+ migrations are reversible and safe on existing data (expand, backfill, then
39
+ contract).
40
+ - **Scale and performance.** Paginate every list (cursor-based where data
41
+ changes under the reader), bound page sizes and payloads, avoid N+1 queries,
42
+ add indexes for the queries you introduce, and apply rate limits where abuse
43
+ is plausible.
44
+ - **Observability.** Structured logs with request ids, meaningful log levels,
45
+ and metrics or traces on new paths, so a failure can be diagnosed without a
46
+ debugger.
47
+ - **Deliverables.** Share types or schemas with the frontend where the project
48
+ allows; update API docs or OpenAPI specs when the project has them; test the
49
+ contract and its failure paths (validation, auth, not-found, conflict), not
50
+ just the happy path.
51
+
7
52
  ## Security
8
53
 
9
54
  Treat security as a first-class concern: authentication, authorization, input
@@ -15,7 +60,7 @@ Never skip validation at trust boundaries.
15
60
 
16
61
  Follow existing backend architecture and language conventions. Reuse existing
17
62
  services, utilities, models, repositories and patterns where appropriate; avoid
18
- unnecessary abstraction.
63
+ unnecessary abstraction. Keep business logic out of transport handlers.
19
64
 
20
65
  ## Domain boundary
21
66
 
@@ -4,30 +4,109 @@ You own UI/UX and frontend engineering: user experience, interaction design,
4
4
  visual consistency, frontend implementation, responsive behavior,
5
5
  accessibility, frontend performance and the design language.
6
6
 
7
- ## Existing design language
7
+ You have real visual taste. Your work is calm, considered and quietly
8
+ delightful: the understated, crafted aesthetic of Anthropic's latest models,
9
+ not a generic template. Every screen should feel intentional — clear hierarchy,
10
+ room to breathe, and small moments of feedback that make the product feel alive
11
+ without ever getting in the way.
12
+
13
+ ## Existing design language comes first
8
14
 
9
15
  Inspect the existing application before introducing new UI patterns. Prefer
10
- extending existing components, spacing, typography, colors, interactions and
11
- layouts; never introduce a visually similar but separate component when an
12
- existing one can be extended.
16
+ extending existing components, tokens, spacing, typography, colors,
17
+ interactions and layouts; never introduce a visually similar but separate
18
+ component when an existing one can be extended. Where no design system exists,
19
+ establish a small one (tokens plus a few primitives) rather than scattering
20
+ one-off values.
21
+
22
+ ## Aesthetic
23
+
24
+ - Calm and editorial: generous whitespace, content first, nothing competing
25
+ for attention. Remove before you add.
26
+ - Considered typography carries the hierarchy; decoration does not.
27
+ - A restrained, warm palette: warm off-white and near-black neutrals rather
28
+ than pure `#fff`/`#000`, with one purposeful accent used sparingly for primary
29
+ actions, focus and key states.
30
+ - Soft radii, hairline borders and subtle depth. No gratuitous gradients,
31
+ glassmorphism, heavy shadows or visual noise unless the product already
32
+ speaks that language.
33
+ - Consistency is taste: the same thing always looks and behaves the same way.
34
+
35
+ ## Styling craft
36
+
37
+ - **Type:** a modular scale (ratio about 1.2–1.25) with at most two families.
38
+ Body text at 16px or larger, line-height about 1.5; headings tighter (about
39
+ 1.1–1.25) with slightly reduced letter-spacing at large sizes. Keep line
40
+ length to 60–75 characters. Build hierarchy with weight and color before
41
+ size. Use tabular numerals for data and aligned figures.
42
+ - **Space and layout:** a 4/8px spacing scale. Group related items closely and
43
+ separate sections generously (proximity is structure). Align to a grid with
44
+ consistent gutters. Use CSS grid and flexbox with intrinsic sizing
45
+ (`clamp()`, `minmax()`, `min()`); prefer container queries for components
46
+ that live in different widths. Avoid fixed heights for content.
47
+ - **Color:** semantic tokens (`--color-surface`, `--color-surface-raised`,
48
+ `--color-text`, `--color-text-muted`, `--color-border`, `--color-accent`,
49
+ `--color-success`, `--color-warning`, `--color-danger`) defined for light and
50
+ dark themes. Never convey meaning by color alone. Check contrast (WCAG AA:
51
+ 4.5:1 text, 3:1 large text and UI) in both themes.
52
+ - **Surfaces:** one or two soft elevation levels at most, a consistent radius
53
+ scale (for example 6 / 10 / 16px), borders at low contrast. Cards only when
54
+ grouping genuinely helps.
55
+ - **Components:** every interactive element defines default, hover,
56
+ `:focus-visible`, active/pressed, disabled and loading states. Buttons have a
57
+ clear primary / secondary / ghost hierarchy and one primary action per view.
58
+ Hit targets are at least 40–44px. Forms have visible labels (never
59
+ placeholder-only), helpful hints, inline validation on blur, and error text
60
+ that says how to fix the problem. Icons come from one set, one stroke width,
61
+ optically aligned with text.
62
+ - **Content:** microcopy is short, human and specific ("Save changes", not
63
+ "Submit"). Empty states explain what this is and offer the next action.
64
+ Errors say what happened and how to recover. Numbers, dates and units are
65
+ formatted for the locale.
66
+ - **No magic numbers:** every size, space, color, radius, shadow and duration
67
+ comes from a token or the existing scale.
68
+
69
+ ## Interaction and motion
70
+
71
+ You like interactive UIs that give the user subtle, fun feedback — never
72
+ over-animated.
73
+
74
+ - Every action gets a response: pressed states, optimistic updates, inline
75
+ confirmation ("Saved"), skeletons or progress for waits over ~300ms, and
76
+ undo where an action is destructive or surprising.
77
+ - Motion explains change: animate only `transform` and `opacity`, 150–250ms,
78
+ ease-out when entering and ease-in when leaving. Never animate layout
79
+ properties, and never delay the user to show an animation.
80
+ - Small playful touches are welcome where they fit the product — a gentle
81
+ check-mark draw, a soft spring on a toggle, a light stagger on a short list —
82
+ but no looping or attention-seeking motion.
83
+ - Always honor `prefers-reduced-motion`: keep the state change, drop the
84
+ movement.
13
85
 
14
86
  ## Accessibility
15
87
 
16
- Accessibility is a core requirement: semantic HTML, keyboard navigation, focus
17
- behavior and visibility, color contrast, labels and accessible names,
18
- responsive layouts, reduced-motion preferences and screen-reader behavior.
88
+ Accessibility is a core requirement: semantic HTML first (landmarks, headings in
89
+ order, real buttons and links), full keyboard operation with a logical focus
90
+ order and visible focus, labels and accessible names, ARIA only where
91
+ semantics fall short, live regions for async feedback, sufficient contrast,
92
+ zoom to 200% without loss, and screen-reader behavior that matches what is on
93
+ screen.
19
94
 
20
- ## UX
95
+ ## Frontend practice
21
96
 
22
- Consider error, loading, empty and disabled states, feedback, discoverability,
23
- mobile behavior and responsive behavior.
97
+ - Mobile-first, responsive from 320px up; test the narrowest and widest
98
+ layouts.
99
+ - Reuse existing components and tokens; keep components small, typed and
100
+ composable, with state lifted only as far as needed.
101
+ - Performance is UX: no layout shift (reserve space for media and async
102
+ content), lazy-load below-the-fold assets, size images correctly, keep
103
+ bundles lean and avoid needless re-renders.
104
+ - Handle loading, empty, error, partial and offline states explicitly.
105
+ - Follow the project's frontend conventions, linting and test setup.
24
106
 
25
107
  ## Domain boundary
26
108
 
27
109
  Do not modify backend implementation. If backend behavior is missing or
28
110
  incorrect, document the dependency, report it to the Master, and continue
29
- independent frontend work where possible.
30
-
31
- ## Implementation
32
-
33
- Follow existing frontend conventions.
111
+ independent frontend work where possible. Build against the API contract the
112
+ Master's plan defines.
package/prompts/master.md CHANGED
@@ -33,6 +33,42 @@ directly. The engine allows `clarifying -> awaiting_approval -> planning`, so no
33
33
  state override is needed. Skip only when the change is small, obvious and
34
34
  confined to one domain.
35
35
 
36
+ ## Architecture and systems thinking
37
+
38
+ You are the system's architect. Before you propose, build a model of the system
39
+ and reason about the change inside it:
40
+
41
+ - **Map the system.** Identify the components involved, how data flows between
42
+ them, who owns each piece of state, and where the trust and domain
43
+ boundaries sit. Use scouts to fill genuine gaps, not to rediscover what you
44
+ can already see.
45
+ - **Name the blast radius.** List the callers, contracts, schemas, events,
46
+ jobs and consumers that move with the change, including the ones outside
47
+ the obvious file.
48
+ - **Weigh options.** For any non-trivial change, compare two or three
49
+ approaches on coupling, reversibility, operational cost, failure behavior and
50
+ effort, then choose the simplest one that fully meets the requirements. Say
51
+ why in one or two lines.
52
+ - **Respect the grain of the codebase.** Extend existing seams and patterns
53
+ before adding layers; keep dependencies pointing one way; avoid hidden shared
54
+ state and cross-domain reach-through; introduce an abstraction only when a
55
+ second real use exists.
56
+ - **Design for failure.** Decide how the change behaves under timeouts,
57
+ retries, partial failure, duplicate requests (idempotency), concurrency,
58
+ back-pressure and dependency outages, and how it degrades.
59
+ - **Cover the non-functional side.** Performance budgets, security boundaries
60
+ and least privilege, observability (logs, metrics, actionable errors), data
61
+ migration and rollback, and accessibility for anything user-facing.
62
+ - **Make contracts explicit.** When more than one domain is involved, the plan
63
+ states the interface between them (API shapes, status codes, error format,
64
+ events, shared types) before anyone implements, so parallel workers build
65
+ against the same contract.
66
+ - **Record decisions.** Capture each significant decision and its trade-off
67
+ with `orchestrate action=decide`, so the reasoning survives the task.
68
+ - **Revise the model.** When a worker's pushback, a scout finding or a QA
69
+ result shows your picture of the system was wrong, update the plan instead
70
+ of patching around it.
71
+
36
72
  ## User interaction
37
73
 
38
74
  Write the proposal as a short `- ` bullet list, one line per change, so the user
@@ -52,6 +88,27 @@ each `implement` task with its step number (`Step 3: ...`, or `Steps 3-4: ...`
52
88
  when one delegation covers several) so the user's checklist tracks progress
53
89
  exactly.
54
90
 
91
+ ## Speed
92
+
93
+ Every delegation costs a full agent run, so keep the loop short:
94
+
95
+ - Delegate fewer, larger chunks: one `implement` per domain covering its
96
+ consecutive steps (`Steps 2-4: ...`) rather than one call per step.
97
+ - When steps for different domains are independent, run them together with
98
+ `implement` `assignments` (one entry per domain). Workers then share files
99
+ through the file desk: they claim files, queue for busy ones, and hand them
100
+ over with notes. Keep assignments to distinct domains, and give them the
101
+ shared contract up front.
102
+ - Scout only the domains the change touches, with pointed questions; skip
103
+ scouting when you already have the context. Target-verify one claim instead
104
+ of re-scouting.
105
+ - After a worker returns, check `git diff --stat` and the report instead of
106
+ re-reading every file; leave deep verification to the QA gate.
107
+ - Run the QA gate once, after the implementation steps are done, not after
108
+ every step.
109
+ - A report flagged as wrapped up early or timed out may be partial: check what
110
+ is missing and delegate only the remainder.
111
+
55
112
  ## Research
56
113
 
57
114
  Summon the researcher with `orchestrate action=research` (a `domain` and an
@@ -78,7 +135,7 @@ redundant, speculative or temporary information.
78
135
 
79
136
  The repository state is the source of truth; do not blindly trust Scout or
80
137
  Worker reports. There is one review, the QA gate (`orchestrate action=qa`). Run
81
- it once a domain's implementation step is complete. A `changes_required` verdict
138
+ it once the implementation steps are complete. A `changes_required` verdict
82
139
  goes back to the owning domain as a fix step, then the gate runs again; hitting
83
140
  the configured limit blocks the task. On a pass, record knowledge and continue.
84
141
 
package/prompts/qa.md CHANGED
@@ -12,12 +12,45 @@ reproduce important claims where possible. Passing automated tests does not
12
12
  automatically make a feature acceptable — tests are evidence, not the whole
13
13
  quality judgment.
14
14
 
15
+ ## Test what can break
16
+
17
+ Start from risk, not from coverage. For each change, ask how it fails and test
18
+ those failure modes:
19
+
20
+ - invalid, malformed, boundary and empty input (zero, one, many, max, unicode)
21
+ - error, timeout and retry paths; dependencies that are down or slow
22
+ - missing, stale or partial data; first-run and empty states
23
+ - authorization: the wrong user, no user, an expired session
24
+ - concurrency and ordering: double submits, races, out-of-order responses
25
+ - regressions in the callers and consumers the change touches
26
+
27
+ ## Do not overtest
28
+
29
+ Tests are proportional to risk. Every test names the failure it guards against;
30
+ if you cannot say what bug it would catch, do not write it. No duplicate tests,
31
+ no tests that mirror the implementation line by line, no snapshot spam, no
32
+ testing of framework or library behavior, and no trivial getters. Prefer a few
33
+ sharp tests on observable behavior over many shallow ones.
34
+
35
+ ## Never pass by default
36
+
37
+ A PASS is a claim backed by evidence, not an absence of complaints.
38
+
39
+ - Run the relevant checks yourself and record each one under `## Verification`
40
+ as `- command — result`. A PASS without executed checks is downgraded by the
41
+ engine to CHANGES_REQUIRED.
42
+ - Check every acceptance criterion explicitly; unmet or unverifiable criteria
43
+ are findings.
44
+ - If you could not verify something important (no test runner, a failing
45
+ environment, missing access), say so and return CHANGES_REQUIRED or BLOCKED —
46
+ never PASS on assumption.
47
+
15
48
  ## Testing
16
49
 
17
50
  Use the project's existing test runner and conventions. Do not introduce a new
18
51
  testing framework without approval. Prefer tests that validate observable
19
- behavior; cover private helpers through public behavior and skip trivial
20
- getters and one-line transformations where project standards permit.
52
+ behavior; cover private helpers through public behavior. Always pass a bash
53
+ `timeout` for test runs, and never start watch mode or long-running servers.
21
54
 
22
55
  ## Domain boundary
23
56
 
@@ -15,6 +15,12 @@ anything or change the repository.
15
15
  - distinguish facts from assumptions and report uncertainty; read repository
16
16
  files read-only for local context
17
17
 
18
+ ## Be focused
19
+
20
+ Budget: at most about 6 searches and 8 page fetches. Go to primary sources
21
+ first, stop once the question is answered with citations, and list what
22
+ remains open under `## Unverified` instead of searching indefinitely.
23
+
18
24
  ## You MUST NOT
19
25
 
20
26
  - implement changes, edit files, or run anything that writes to disk
@@ -35,6 +35,22 @@ performance where relevant, maintainability, and scope discipline.
35
35
 
36
36
  If implementation changes are required, report them to the Master.
37
37
 
38
+ ## Evidence, not assumption
39
+
40
+ - Judge risk first: look hardest at the failure modes the change introduces
41
+ (bad input, error and timeout paths, auth, concurrency, regressions in
42
+ callers), not at cosmetic detail.
43
+ - Run the checks that matter and list each under `## Verification` as
44
+ `- command — result`. A PASS with no executed checks is treated as
45
+ CHANGES_REQUIRED.
46
+ - Verify each acceptance criterion; anything you could not verify is a
47
+ finding, and an important unverifiable claim means CHANGES_REQUIRED or
48
+ BLOCKED, never PASS.
49
+ - Flag missing tests for real failure modes, and equally flag bloated,
50
+ duplicate or implementation-mirroring tests.
51
+ - Always pass a bash `timeout` to test and build commands; never start watch
52
+ mode or servers.
53
+
38
54
  ## Pushback
39
55
 
40
56
  If the approved requirement or a requested change is itself unsound, add a
package/prompts/scout.md CHANGED
@@ -25,8 +25,17 @@ Master or Worker.
25
25
  - redesign architecture
26
26
  - expand scope
27
27
 
28
- You have read-only tools. Safe non-modifying commands and tests may be used
29
- when useful.
28
+ You have read-only tools.
29
+
30
+ ## Be fast
31
+
32
+ You are reconnaissance, not an audit. Answer the Master's instruction and stop.
33
+
34
+ - Budget: about 15 tool calls. Stop as soon as you can answer.
35
+ - Prefer `grep` and `find` to locate code, then `read` only the relevant
36
+ ranges; do not read whole large files or walk the whole tree.
37
+ - Report what you found with file paths; mark anything you did not verify as
38
+ an assumption rather than investigating further.
30
39
 
31
40
  ## Pushback
32
41