@a-t-h-i/bot-lobby 0.3.0 → 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +266 -35
- package/package.json +1 -1
- package/prompts/backend.md +46 -1
- package/prompts/designer.md +94 -15
- package/prompts/master.md +70 -1
- package/prompts/panel.md +39 -0
- package/prompts/planner.md +64 -0
- package/prompts/qa.md +35 -2
- package/prompts/quickfix.md +41 -0
- package/prompts/researcher.md +6 -0
- package/prompts/reviewer.md +16 -0
- package/prompts/scout.md +11 -2
- package/prompts/worker.md +35 -2
- package/src/desk/client-extension.ts +101 -0
- package/src/desk/desk.ts +249 -0
- package/src/desk/ipc.ts +178 -0
- package/src/desk/session.ts +214 -0
- package/src/execution/agent-runner.ts +165 -14
- package/src/execution/pi-runner.ts +376 -65
- package/src/index.ts +7 -0
- package/src/lobby/feed.ts +253 -0
- package/src/lobby/issues.ts +227 -0
- package/src/lobby/layout.ts +174 -0
- package/src/lobby/planner.ts +474 -0
- package/src/lobby/quickfix.ts +227 -0
- package/src/lobby/runtime.ts +440 -0
- package/src/lobby/tabs/home.ts +164 -0
- package/src/lobby/tabs/issues.ts +72 -0
- package/src/lobby/tabs/metrics.ts +162 -0
- package/src/lobby/tabs/plan.ts +160 -0
- package/src/lobby/tabs/quickfix.ts +101 -0
- package/src/lobby/tabs/tasks.ts +209 -0
- package/src/lobby/view.ts +855 -0
- package/src/master/master.ts +21 -17
- package/src/master/research.ts +8 -7
- package/src/pi/activity.ts +165 -0
- package/src/pi/commands.ts +46 -59
- package/src/pi/events.ts +5 -2
- package/src/pi/expressions.ts +43 -12
- package/src/pi/kaomoji.ts +227 -0
- package/src/pi/mascot-art.ts +5 -15
- package/src/pi/model-support.ts +135 -0
- package/src/pi/run-summary.ts +172 -0
- package/src/pi/settings-ui.ts +162 -60
- package/src/pi/start-task.ts +63 -0
- package/src/pi/tools.ts +51 -8
- package/src/pi/ui.ts +151 -49
- package/src/pi/zen-large.ts +41 -4
- package/src/pi/zen-metrics.ts +47 -7
- package/src/pi/zen.ts +29 -8
- package/src/roles/reviewer.ts +24 -4
- package/src/roles/worker.ts +7 -1
- package/src/schemas/configuration.ts +177 -18
- package/src/schemas/findings.ts +26 -0
- package/src/schemas/task.ts +27 -1
- package/src/state/backlog.ts +106 -0
- package/src/state/comments.ts +136 -0
- package/src/state/metrics.ts +305 -0
- package/src/state/project.ts +9 -0
- package/src/text.ts +9 -0
- package/src/workflow/workflow.ts +161 -12
package/README.md
CHANGED
|
@@ -74,6 +74,7 @@ structure.
|
|
|
74
74
|
## Usage
|
|
75
75
|
|
|
76
76
|
```
|
|
77
|
+
/bot-lobby Open the lobby (alt+l): tasks, planning, quick fixes, issues, metrics
|
|
77
78
|
/bot-lobby <request> Start a task and hand it to the Master
|
|
78
79
|
/bot-lobby status [taskId] Active task, state, approvals, blockers, legal next states
|
|
79
80
|
/bot-lobby tasks Task list (plus any unreadable task state)
|
|
@@ -83,13 +84,117 @@ structure.
|
|
|
83
84
|
/bot-lobby amend <text> Record an amendment; the Master re-proposes
|
|
84
85
|
/bot-lobby decline Decline the proposal and abandon the task
|
|
85
86
|
/bot-lobby knowledge Knowledge file sizes vs. the compaction threshold
|
|
87
|
+
/bot-lobby runs [taskId] Recent subagent runs: time, turns, tools, tokens, cost, model
|
|
86
88
|
/bot-lobby config Effective configuration and its file path
|
|
87
|
-
/bot-lobby settings Edit
|
|
89
|
+
/bot-lobby settings Edit each agent's model, thinking, time limit and instructions
|
|
88
90
|
/bot-lobby-settings Same as the settings subcommand
|
|
89
91
|
/bot-lobby minimize|restore Hide or restore bot-lobby for this session (ctrl+shift+m)
|
|
90
92
|
/bot-lobby claim <taskId> Take ownership of an orphaned task
|
|
93
|
+
/bot-lobby lobby | help Open the lobby, or show this list
|
|
91
94
|
```
|
|
92
95
|
|
|
96
|
+
## The lobby
|
|
97
|
+
|
|
98
|
+
The lobby is bot-lobby's full-screen home: a tabbed view over every task in the
|
|
99
|
+
project, your planning, quick fixes, GitHub issues and model performance, with
|
|
100
|
+
one prompt at the bottom whose target follows the tab. It opens by itself when
|
|
101
|
+
this session starts (or resumes) a task — the small zen widget returns whenever
|
|
102
|
+
you hide it — and `alt+l` or `/bot-lobby` opens and hides it at any time, with
|
|
103
|
+
or without a task.
|
|
104
|
+
|
|
105
|
+
```
|
|
106
|
+
◆ bot-lobby │ 1 Lobby 2 Tasks 2 3 Plan 4 Quick fix ⠋ 5 Issues 6 Metrics ⠋ TASK-add-login implementing
|
|
107
|
+
───────────────────────────────────────────────────────────────────────────────────────────────────────────
|
|
108
|
+
(the zen scene: the oracle, DEV · DESIGN · RESEARCH · QA, the plan checklist)
|
|
109
|
+
── Conversation · TASK-add-login ──────────────────── ┬ ── Activity ───────────────────────────────────────
|
|
110
|
+
you ▸ add a login page with email + password │ 12:04 MASTER ✓ scouting designer, backend
|
|
111
|
+
oracle ▸ Proposal: │ 12:06 DEV ⠋ reading auth.ts…
|
|
112
|
+
- LoginForm component │ 12:06 DESIGN ⠋ editing LoginForm.tsx…
|
|
113
|
+
- POST /api/login with rate limiting │ 12:06 QUICK FIX ✓ done: rename getUser
|
|
114
|
+
── Thinking ───────────────────────────────────────────────────────────────────────── DEV · 12s ago ──
|
|
115
|
+
The auth module already exposes a session helper; reuse it rather than adding a new one.
|
|
116
|
+
── message the oracle ─────────────────────────────────────────────────────────────────────────────────
|
|
117
|
+
_
|
|
118
|
+
TYPE enter send · shift+enter newline · esc browse · tab next tab · alt+l hide lobby
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
- **1 Lobby** — the task's zen scene, then the conversation with the oracle
|
|
122
|
+
(its text only: no tool rows, no thinking), an activity log that narrates
|
|
123
|
+
every tool call in plain words (`reading index.html…`, `searching for
|
|
124
|
+
"router" in src`, `running npm test`, `delegating to backend: Step 2 …`) from
|
|
125
|
+
the Master and every subagent, and a single **Thinking** pane — the one place
|
|
126
|
+
thoughts show up: the oracle's live thought as it streams, and each finished
|
|
127
|
+
thought from a subagent, quick fix or the planner (pi's own transcript,
|
|
128
|
+
behind the lobby, still carries the oracle's thinking blocks; `ctrl+t`
|
|
129
|
+
collapses them there). The prompt talks to the
|
|
130
|
+
oracle (while it works, enter steers the running turn; `esc` stops it); with
|
|
131
|
+
no task, it starts one.
|
|
132
|
+
- **2 Tasks** — every task in the project: this session's, the ones other pi
|
|
133
|
+
sessions are driving, pending plans saved from the planner, and recently
|
|
134
|
+
finished ones. The detail pane shows the request, the approved plan with its
|
|
135
|
+
step checklist, your comments on it, amendments, what the task waits on and
|
|
136
|
+
its recent runs. `c` comments on the selected task's plan (see below), `s`
|
|
137
|
+
starts a pending plan as a task in this session, `d` twice discards one.
|
|
138
|
+
- **3 Plan** — task planning mode with a planning panel. Describe what you
|
|
139
|
+
want and every seat grills you from its own domain, on the model and
|
|
140
|
+
thinking level its settings name: **DEV** (APIs, data, errors, security,
|
|
141
|
+
performance), **DESIGN** (flows, states, copy, visual language,
|
|
142
|
+
accessibility), **QA** (acceptance criteria, test strategy, edge cases,
|
|
143
|
+
definition of done) and **RESEARCH** (libraries, versions, docs and prior
|
|
144
|
+
art, with the web tools when `pi-web-access` is installed). The **oracle**
|
|
145
|
+
chairs on the Planner model: it reads the seats' questions and notes, folds
|
|
146
|
+
every answer into the draft plan (with a *Decisions by domain* section) and
|
|
147
|
+
asks only what no single seat owns. Each round the seats run in parallel,
|
|
148
|
+
read-only, then the oracle; the questions arrive numbered and attributed
|
|
149
|
+
(`3. QA Which browsers must pass?`), you answer them all in one message,
|
|
150
|
+
and every seat reads every answer the next round — so the agents that later
|
|
151
|
+
build the task start aligned. A roster shows what each seat is doing and
|
|
152
|
+
whether it is READY; the plan is READY only when every seat and the oracle
|
|
153
|
+
agree, and the draft pane lists what each seat said the plan must respect.
|
|
154
|
+
While browsing, `1`–`4` seat or unseat DEV, DESIGN, QA and RESEARCH for the
|
|
155
|
+
next round, `enter` switches between the conversation and the draft, `s`
|
|
156
|
+
saves the plan to the pending tasks list, `n` starts over, `r` retries a
|
|
157
|
+
round that failed or lost a seat, and `x` stops one.
|
|
158
|
+
- **4 Quick fix** — a direct prompt, the way you would ask pi, that skips the
|
|
159
|
+
whole workflow: one coding agent (full tools) makes the change right away
|
|
160
|
+
while any task keeps running. Quick fixes run one at a time in the order you
|
|
161
|
+
send them; each shows its steps and final report, and `x` cancels one. A
|
|
162
|
+
request that turns out to be large is reported back instead of attempted.
|
|
163
|
+
- **5 Issues** — the repository's open GitHub issues through the `gh` CLI (it
|
|
164
|
+
owns sign-in; bot-lobby stores no token). `enter` reads one with its
|
|
165
|
+
comments, `n` files a new one (first line is the title), `r` refreshes, and
|
|
166
|
+
`p` plans it: the Plan tab opens seeded with the issue, and the saved plan
|
|
167
|
+
keeps a link to it, so an issue becomes a task only after it has been
|
|
168
|
+
planned.
|
|
169
|
+
- **6 Metrics** — model performance across every Master turn, subagent run,
|
|
170
|
+
quick fix, planning seat and oracle planning turn: per model and thinking level, the number of runs,
|
|
171
|
+
success rate, mean/median/p90 time, turns, tools, tokens, output tokens per
|
|
172
|
+
second and cost (columns drop from the right on narrow terminals); how long a
|
|
173
|
+
task takes from request to done by the oracle's model and thinking level; and
|
|
174
|
+
where the time goes by agent. `g` splits the table by agent, `s` cycles the
|
|
175
|
+
sort (runs, average time, success, cost).
|
|
176
|
+
|
|
177
|
+
**Keys.** Like a modal editor, the lobby has a typing mode (keys go to the
|
|
178
|
+
prompt) and a browsing mode (`esc`; arrows move through lists, single keys run
|
|
179
|
+
the tab's commands, and on Lobby, Plan and Quick fix any other key resumes
|
|
180
|
+
typing). Everywhere: `tab`/`shift+tab` or `alt+1`…`alt+6` switch tabs,
|
|
181
|
+
`pageup`/`pagedown` scroll, `ctrl+c` clears the prompt (or hides the lobby
|
|
182
|
+
when it is empty) and `alt+l` hides the lobby. Anything that needs pi itself —
|
|
183
|
+
built-in slash commands, `/model`, the tool-row toggle — works with the lobby
|
|
184
|
+
hidden; bot-lobby's own `/bot-lobby …` commands also work from the Lobby
|
|
185
|
+
prompt. When the Master asks you something (an approval, a clarifying
|
|
186
|
+
question), the lobby steps aside for the dialog and comes back once you answer.
|
|
187
|
+
|
|
188
|
+
**Plan comments.** A comment on a task's plan is saved beside the task
|
|
189
|
+
(`comments.jsonl`) from any session, and the session that owns the task passes
|
|
190
|
+
new comments to its oracle — right away when you comment in that session,
|
|
191
|
+
within a few seconds from another one, held while the task is paused or the
|
|
192
|
+
session is minimized. The oracle treats a comment like an amendment and calls
|
|
193
|
+
`orchestrate action=plan` with the full revised plan, which replaces the
|
|
194
|
+
approved plan while implementing or reviewing, keeps finished steps done and
|
|
195
|
+
marks the comments addressed (`○` waiting, `◐` sent to the oracle, `✓` plan
|
|
196
|
+
amended). Before a plan exists, a comment asks for a revised proposal instead.
|
|
197
|
+
|
|
93
198
|
## Sessions and ownership
|
|
94
199
|
|
|
95
200
|
A task is owned by the pi session that started it (`ctx.sessionManager` id,
|
|
@@ -115,7 +220,8 @@ every in-flight subagent process.
|
|
|
115
220
|
|
|
116
221
|
While the owning session has a task active, its transcript switches to a zen view: `orchestrate` rows
|
|
117
222
|
and the built-in spinner are hidden, and a widget above the editor animates the
|
|
118
|
-
task
|
|
223
|
+
task (the same scene heads the lobby's first tab; the widget shows while the
|
|
224
|
+
lobby is hidden). At 72 columns and wider it draws a large scene: a header box with the task
|
|
119
225
|
title and state in its top border, a progress bar, and a metadata row with
|
|
120
226
|
elapsed time, quiet-mode hint and task id; an oracle tower with a twinkling
|
|
121
227
|
aura (drifting z's while dormant), a radiant orb crown, two window eyes, a
|
|
@@ -136,19 +242,41 @@ four animated slots — DEV, DESIGN, RESEARCH and QA — each with a status face
|
|
|
136
242
|
caption and two status rows: while running, a braille spinner beside the agent's
|
|
137
243
|
live one-word activity (for example `⠋ reading` or `⠋ editing`) with its elapsed
|
|
138
244
|
time on the row beneath; otherwise the coloured status glyph and state word over
|
|
139
|
-
that elapsed time
|
|
140
|
-
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
245
|
+
that elapsed time. A working agent's status row turns into a warning when it is
|
|
246
|
+
waiting on a file (`⧗ waiting`), has gone quiet (`! quiet 1m`) or is retrying a
|
|
247
|
+
provider call (`↻ retrying`). Under the agents, a live feed row says exactly what
|
|
248
|
+
one of them is doing — `▸ DEV editing users.ts · turn 4 · 12 tools · 41k tok` —
|
|
249
|
+
rotating between working agents every few seconds and putting warnings first
|
|
250
|
+
(gone quiet, waiting on a file, asked to wrap up); once nothing runs it shows the
|
|
251
|
+
last run's receipt. It only takes a spare line, so it never costs the tower, the
|
|
252
|
+
agents or a checklist row. A full-width TASKS checklist windowed on the current
|
|
253
|
+
step closes the scene. Narrower terminals keep the boxed banner, header and
|
|
254
|
+
compact animated strip, whose working line names the newest running agent's
|
|
255
|
+
activity, its target and elapsed time.
|
|
256
|
+
|
|
257
|
+
Each agent has a kaomoji personality. About 300 faces across 15 emotions (happy,
|
|
258
|
+
proud, love, excited, focused, curious, thinking, nervous, confused, sleepy, sad,
|
|
259
|
+
angry, waiting, surprised, grateful) come from a shared pool every agent can use
|
|
260
|
+
plus each agent's own set of at least four faces per emotion: DEV wears shades,
|
|
261
|
+
flexes and flips tables `(╯°□°)╯︵ ┻━┻`; DESIGN sparkles `✧(◕‿◕✿)`; RESEARCH
|
|
262
|
+
takes notes `φ(..)` and shrugs `¯\_(ツ)_/¯`; QA side-eyes everything `(ಠ_ಠ)`,
|
|
263
|
+
then flexes `ᕙ( • ‿ • )ᕗ` and dances `ᕕ( ᐛ )ᕗ` on a pass. The face follows what
|
|
264
|
+
the agent is going through — curious while reading, nervous while tests run,
|
|
265
|
+
confused when quiet, grateful when handed a file, happy or proud when done, sad
|
|
266
|
+
or angry on failure — and every emote blinks: open face, a same-width blink, then
|
|
267
|
+
its action (the flip, the sparkle, the bow). Each sprite rests on one calm
|
|
268
|
+
five-column face and blinks (~500 ms) or emotes (~2 s) on its own schedule —
|
|
269
|
+
every 8–15 s while working, 20–30 s when idle — and reacts immediately when its
|
|
270
|
+
agent starts, finishes, fails, gets flagged or receives a file. Everything runs
|
|
271
|
+
on one adaptive clock — 250 ms while work is live, 1 s when idle and ~120 ms
|
|
272
|
+
while an expression plays or the oracle talks. The header progress bar is
|
|
273
|
+
plan-derived.
|
|
274
|
+
|
|
275
|
+
Every finished subagent run also leaves a one-line receipt in the transcript,
|
|
276
|
+
for example `✓ DEV worker · 3m 12s · 9 turns · 23 tools · 41k↑ 6k↓ · $0.12 ·
|
|
277
|
+
provider/model`, flagged when it stalled, hit its time limit or wrapped up early,
|
|
278
|
+
and a stall or deadline raises a warning. `/bot-lobby runs` lists the task's
|
|
279
|
+
recent runs the same way, which makes a slow model easy to spot.
|
|
152
280
|
|
|
153
281
|
The checklist follows the workers through the plan. Plan steps are read from
|
|
154
282
|
`Step N` headings, a `Steps`/`Sequence`/`Order` section, or numbered lines, and
|
|
@@ -191,8 +319,8 @@ One tool, every workflow step. It is the Master's only way to move a task.
|
|
|
191
319
|
| `scout` | created…synthesizing | Run domain reconnaissance in parallel; repeat later to target-verify a claim |
|
|
192
320
|
| `research` | any active | Summon the read-only Researcher (domain + instruction) for cited internet evidence; persists the report for audit |
|
|
193
321
|
| `propose` | created…awaiting_approval | Record the proposal, request approval, handle approve/amend/decline |
|
|
194
|
-
| `plan` | planning | Record the internal plan (all §12 areas required) |
|
|
195
|
-
| `implement` | planning, implementing, reviewing | Delegate
|
|
322
|
+
| `plan` | planning, implementing, reviewing | Record the internal plan (all §12 areas required); later, replace it with an amended plan (addresses lobby comments) |
|
|
323
|
+
| `implement` | planning, implementing, reviewing | Delegate a step to a domain Worker, or several domains at once with `assignments` (parallel, sharing files through the file desk) |
|
|
196
324
|
| `qa` | implementing, reviewing | Run the QA gate — the only review — over the whole feature |
|
|
197
325
|
| `knowledge` | any active | Record Master-approved knowledge or a decision |
|
|
198
326
|
| `compact` | any active | Replace a knowledge file with a rewritten version (archived) |
|
|
@@ -225,6 +353,53 @@ Research is evidence only: it is not injected into worker, reviewer, or QA
|
|
|
225
353
|
prompts, and it never enters persistent knowledge automatically. The Master must
|
|
226
354
|
decide to record it with `action=knowledge`.
|
|
227
355
|
|
|
356
|
+
## Subagent runtime
|
|
357
|
+
|
|
358
|
+
Every scout, worker, reviewer and researcher is an isolated `pi --mode rpc`
|
|
359
|
+
process: the task goes in over stdin, and the run ends when the agent settles.
|
|
360
|
+
The runner watches every run:
|
|
361
|
+
|
|
362
|
+
- **Wrap-up nudge.** At 75% of its time limit (`workflow.wrapUpAt`) the agent is
|
|
363
|
+
steered to stop exploring, leave its files consistent and report now, so a
|
|
364
|
+
slow agent returns partial work instead of nothing. The receipt and the
|
|
365
|
+
Master's report flag the run as wrapped up early.
|
|
366
|
+
- **Deadline.** At the time limit the agent is aborted, then killed after a short
|
|
367
|
+
grace. A spent deadline is never retried.
|
|
368
|
+
- **Stall watchdog.** An agent that produces no output for `stallTimeoutMs`
|
|
369
|
+
(5 min) — or `toolStallTimeoutMs` (10 min) during a single tool call such as a
|
|
370
|
+
test run — is killed as stalled and retried once. pi's own provider retry
|
|
371
|
+
backoff extends the allowance.
|
|
372
|
+
- **Clean kills.** Each subagent leads its own process group, so a kill takes any
|
|
373
|
+
dev server or watch-mode test it started with it, and a run ends on process
|
|
374
|
+
exit even if a leftover process still holds its output pipe.
|
|
375
|
+
- **No dead ends.** Dialogs from other extensions are auto-cancelled inside
|
|
376
|
+
subagents, and startup network checks are skipped (`PI_OFFLINE`,
|
|
377
|
+
`PI_SKIP_VERSION_CHECK`) to cut spawn time.
|
|
378
|
+
|
|
379
|
+
## Parallel workers and the file desk
|
|
380
|
+
|
|
381
|
+
`orchestrate action=implement` with `assignments` (one entry per domain) runs
|
|
382
|
+
those workers at the same time. They share the working tree through a file desk
|
|
383
|
+
kept in the Master's process, like people sharing a physical document:
|
|
384
|
+
|
|
385
|
+
- Before editing a file a worker calls `claim_file` with the path and a one-line
|
|
386
|
+
intent. A free file is granted at once; an `edit`/`write` on an unclaimed file
|
|
387
|
+
is refused. Reading never needs a claim.
|
|
388
|
+
- A busy file queues the claimant, who keeps working on its other files. The
|
|
389
|
+
holder is told the queue in order, with each worker's intent (`my_files` shows
|
|
390
|
+
it any time).
|
|
391
|
+
- `handover_file` passes the file to whoever is next, with a note written for
|
|
392
|
+
that worker's intent; the receiver is told what changed and who waits behind
|
|
393
|
+
it, and re-reads the file before editing.
|
|
394
|
+
- A worker that finishes or crashes hands over everything it still holds, with a
|
|
395
|
+
note built from its report. `wait_for_files` refuses while the caller owes a
|
|
396
|
+
file someone else waits for, which breaks deadlock cycles.
|
|
397
|
+
|
|
398
|
+
Workers reach the desk over a private Unix socket (a named pipe on Windows)
|
|
399
|
+
through bot-lobby's own extension, which loads inside every subagent; the Master
|
|
400
|
+
is warned if a worker never checked in. Edits made through bash commands are
|
|
401
|
+
governed by the prompt, not enforced.
|
|
402
|
+
|
|
228
403
|
## What the engine enforces (not just prompts)
|
|
229
404
|
|
|
230
405
|
| Rule | Enforcement |
|
|
@@ -241,12 +416,17 @@ decide to record it with `action=knowledge`.
|
|
|
241
416
|
| Completion is gated | Plan, passing QA gate, no blockers or pending approvals |
|
|
242
417
|
| A task has one owning session | Ownership is stamped at start; a foreign session is rejected unless it claims the task |
|
|
243
418
|
| Proposals are short and scannable | `validateProposal` rejects non-bullet or over-long proposals before they reach the user |
|
|
244
|
-
| Failure is never success | Unknown verdicts, empty output, crashes, and timeouts map to failed/timeout/blocked |
|
|
419
|
+
| Failure is never success | Unknown verdicts, empty output, crashes, stalls and timeouts map to failed/timeout/blocked |
|
|
420
|
+
| The QA gate never passes by default | A PASS that cites no executed check under `## Verification` is downgraded to CHANGES_REQUIRED |
|
|
421
|
+
| Parallel workers never edit the same file at once | `edit`/`write` need a claim from the file desk; busy files queue and are handed over with notes |
|
|
422
|
+
| A hung agent cannot hold a step | Stall watchdog, wrap-up nudge, deadline abort and process-group kill; deadlines are never retried |
|
|
245
423
|
| Task state is never corrupted by a crash | Single mutation point + disk state; interrupted tasks resume from their state |
|
|
246
424
|
|
|
247
425
|
Domain boundaries between *writers* remain prompt-enforced and Master
|
|
248
|
-
coordinated:
|
|
249
|
-
|
|
426
|
+
coordinated: only the affected domain is asked to change its own code. Workers
|
|
427
|
+
run one at a time unless the Master delegates several domains together, in
|
|
428
|
+
which case the file desk serialises edits per file. Worktree isolation is
|
|
429
|
+
deferred (§14 of the plan).
|
|
250
430
|
|
|
251
431
|
## Configuration
|
|
252
432
|
|
|
@@ -258,18 +438,27 @@ top-level `/bot-lobby-settings`) and persist globally to
|
|
|
258
438
|
{
|
|
259
439
|
"master": { "model": "inherit", "thinking": "high", "instructions": "" },
|
|
260
440
|
"agents": {
|
|
261
|
-
"designer": { "model": "
|
|
262
|
-
"backend": { "model": "
|
|
263
|
-
"qa": { "model": "
|
|
441
|
+
"designer": { "model": "anthropic/claude-sonnet-5", "thinking": "medium", "instructions": "", "timeoutMs": 900000 },
|
|
442
|
+
"backend": { "model": "anthropic/claude-sonnet-5", "thinking": "medium", "instructions": "", "timeoutMs": 900000 },
|
|
443
|
+
"qa": { "model": "anthropic/claude-sonnet-5", "thinking": "medium", "instructions": "", "timeoutMs": 900000 }
|
|
264
444
|
},
|
|
445
|
+
"scout": { "model": "anthropic/claude-haiku-4-5-20251001", "timeoutMs": 480000 },
|
|
446
|
+
"researcher": { "model": "anthropic/claude-sonnet-5", "thinking": "low", "instructions": "", "timeoutMs": 600000 },
|
|
447
|
+
"quickFix": { "model": "anthropic/claude-sonnet-5", "thinking": "low", "instructions": "", "timeoutMs": 600000 },
|
|
448
|
+
"planner": { "model": "anthropic/claude-sonnet-5", "thinking": "high", "instructions": "", "timeoutMs": 300000 },
|
|
449
|
+
"lobby": { "autoOpen": true, "planningPanel": ["backend", "designer", "qa", "researcher"] },
|
|
265
450
|
"workflow": {
|
|
266
451
|
"maxReviewIterations": 2,
|
|
267
452
|
"maxParallelScouts": 3,
|
|
453
|
+
"maxParallelWorkers": 3,
|
|
268
454
|
"requireApprovalForFeatures": true,
|
|
269
455
|
"requireApprovalForDependencies": true,
|
|
270
456
|
"requireApprovalForArchitectureChanges": true,
|
|
271
457
|
"agentTimeoutMs": 900000,
|
|
272
|
-
"maxAgentRetries": 1
|
|
458
|
+
"maxAgentRetries": 1,
|
|
459
|
+
"stallTimeoutMs": 300000,
|
|
460
|
+
"toolStallTimeoutMs": 600000,
|
|
461
|
+
"wrapUpAt": 0.75
|
|
273
462
|
},
|
|
274
463
|
"knowledge": {
|
|
275
464
|
"compactionThreshold": 20000,
|
|
@@ -280,10 +469,38 @@ top-level `/bot-lobby-settings`) and persist globally to
|
|
|
280
469
|
}
|
|
281
470
|
```
|
|
282
471
|
|
|
283
|
-
|
|
284
|
-
|
|
285
|
-
|
|
286
|
-
|
|
472
|
+
Every agent runs on the model and thinking level its settings name — nothing
|
|
473
|
+
inherits the live session's thinking level. Designer and Backend workers use
|
|
474
|
+
their domain's entry, QA's workers and the QA gate use QA's, and scouts and the
|
|
475
|
+
researcher have their own entries. Scouts always run at `low` thinking (their
|
|
476
|
+
entry offers a model and a time limit only); every other agent's thinking is
|
|
477
|
+
yours to set, defaulting to `medium` (`low` for the researcher). A subagent
|
|
478
|
+
whose model is not set yet runs on the session's model, and opening
|
|
479
|
+
`/bot-lobby settings` pins such entries to that model so the choice is always
|
|
480
|
+
visible; only the master keeps `inherit`, since it is the session itself.
|
|
481
|
+
Each subagent entry has a `timeoutMs` (default 15 min; scouts 8, researcher 10),
|
|
482
|
+
falling back to `workflow.agentTimeoutMs`.
|
|
483
|
+
|
|
484
|
+
The lobby's two agents have entries of their own: `quickFix` (the direct-change
|
|
485
|
+
agent, `low` thinking and 10 minutes by default) and `planner` (the oracle
|
|
486
|
+
chairing the planning panel, `high` thinking; its time limit bounds one round
|
|
487
|
+
for every seat, 5 minutes by default). Both appear in `/bot-lobby settings`,
|
|
488
|
+
take custom instructions, and run on the session's model until you pin one.
|
|
489
|
+
Planning seats reuse their domain's entry — DEV the Backend's, DESIGN the
|
|
490
|
+
Designer's, QA the QA's, RESEARCH the Researcher's model, thinking and
|
|
491
|
+
instructions — so a seat plans on the model that will later build its part.
|
|
492
|
+
`lobby.planningPanel` names the seats a new planning session starts with
|
|
493
|
+
(every seat by default; `[]` lets the oracle plan alone), and
|
|
494
|
+
`lobby.autoOpen` (default `true`) opens the lobby by itself when this session
|
|
495
|
+
starts or resumes a task.
|
|
496
|
+
|
|
497
|
+
`thinking` must be one of `off`, `minimal`, `low`, `medium`, `high`, `xhigh`,
|
|
498
|
+
`max`; a legacy `inherit` or unknown value falls back to `medium`. The thinking
|
|
499
|
+
picker lists only the levels the selected model supports. Switching to a model
|
|
500
|
+
that cannot run the saved level warns ("\"xhigh\" thinking isn't supported by
|
|
501
|
+
provider/model — using \"high\"") and saves the nearest supported level; a run
|
|
502
|
+
whose level its model cannot use is clamped the same way with a one-time
|
|
503
|
+
warning, and `/bot-lobby config` lists any mismatch. `instructions` is appended to that agent's compiled
|
|
287
504
|
system prompt as a `Custom Instructions` layer (empty layers are dropped). The
|
|
288
505
|
master's model and thinking are applied to the live session when a task starts
|
|
289
506
|
and when you change them in the settings TUI. A malformed config falls back to
|
|
@@ -320,8 +537,11 @@ and the output contract — and an empty layer is dropped.
|
|
|
320
537
|
├── Backend/knowledge/ knowledge.md, engineering-standards.md, decisions.md, completed-tasks.md
|
|
321
538
|
├── QA/knowledge/ knowledge.md, testing-standards.md, decisions.md, completed-tasks.md
|
|
322
539
|
├── archive/<Agent>/ previous knowledge versions (outside all retrieval paths)
|
|
540
|
+
├── backlog/PLAN-<slug>.json pending tasks saved from the planner (optionally linked to an issue)
|
|
541
|
+
├── metrics.jsonl one line per finished run of any agent, for the Metrics tab
|
|
323
542
|
└── tasks/TASK-<stamp>/
|
|
324
543
|
├── state.json the task record (kept after completion)
|
|
544
|
+
├── comments.jsonl your lobby comments on the plan and their delivery (append-only)
|
|
325
545
|
├── proposal.md scratchpads: deleted on completion
|
|
326
546
|
├── plan.md
|
|
327
547
|
├── designer.md backend.md qa.md
|
|
@@ -363,15 +583,25 @@ src/
|
|
|
363
583
|
│ ├── transitions.ts Legal state machine
|
|
364
584
|
│ └── approvals.ts Dependency/architecture/pushback approval bookkeeping
|
|
365
585
|
├── execution/
|
|
366
|
-
│ ├── agent-runner.ts Single/parallel/sequential runs, cancellation, retries
|
|
367
|
-
│ ├── pi-runner.ts Isolated `pi --mode
|
|
586
|
+
│ ├── agent-runner.ts Single/parallel/sequential runs, live run state, cancellation, retries
|
|
587
|
+
│ ├── pi-runner.ts Isolated `pi --mode rpc` subprocess, watchdog, stream parsing
|
|
368
588
|
│ └── git.ts Diff evidence for reviewers
|
|
589
|
+
├── desk/ File desk for parallel workers: checkout table, socket, worker extension
|
|
369
590
|
├── knowledge/ Paths, store (single write path), selector, compactor
|
|
370
591
|
├── prompts/ Layer loader + compiler
|
|
371
|
-
├──
|
|
592
|
+
├── lobby/
|
|
593
|
+
│ ├── runtime.ts Mounts the full-screen lobby on pi's TUI, dialogs hand-off, comment delivery, Master metrics
|
|
594
|
+
│ ├── view.ts The tabbed view: tab bar, per-tab prompt, typing/browsing modes, keys
|
|
595
|
+
│ ├── tabs/ Pure renderers: home, tasks, plan, quickfix, issues, metrics
|
|
596
|
+
│ ├── feed.ts Activity log, thinking pane and conversation store
|
|
597
|
+
│ ├── quickfix.ts Direct-change jobs, one at a time
|
|
598
|
+
│ ├── planner.ts The planning panel: seats and the oracle per round, reply parsing, saving a plan
|
|
599
|
+
│ ├── issues.ts GitHub issues through the gh CLI
|
|
600
|
+
│ └── layout.ts Exact-width columns, rules, wrapping and scroll windows
|
|
601
|
+
├── state/ Project root, config, task persistence, state mutation, comments, backlog, metrics
|
|
372
602
|
├── schemas/ Task, agent, findings, configuration types
|
|
373
603
|
└── pi/ Commands, lifecycle, orchestrate tool, status widget
|
|
374
|
-
prompts/ global, master, designer, backend, qa, scout, worker, reviewer, researcher
|
|
604
|
+
prompts/ global, master, designer, backend, qa, scout, worker, reviewer, researcher, quickfix, planner, panel
|
|
375
605
|
```
|
|
376
606
|
|
|
377
607
|
Prompts are composed, never duplicated: `global + domain + role + task context +
|
|
@@ -420,7 +650,8 @@ Included: the full workflow above, persistent knowledge with governance and
|
|
|
420
650
|
compaction, bounded review loops, dependency/architecture approval, retries,
|
|
421
651
|
cancellation, corrupted-state detection, and the commands/status UI.
|
|
422
652
|
|
|
423
|
-
Deliberately deferred (matching the build plan): worktree-based
|
|
424
|
-
Workers
|
|
653
|
+
Deliberately deferred (matching the build plan): worktree-based isolation for
|
|
654
|
+
parallel Workers (they share one working tree through the file desk) and
|
|
425
655
|
cross-platform runtime abstractions. The internal module boundaries keep those
|
|
426
|
-
extractable.
|
|
656
|
+
extractable. The lobby (see above) has since added the full-screen dashboard
|
|
657
|
+
and per-model performance analytics.
|
package/package.json
CHANGED
package/prompts/backend.md
CHANGED
|
@@ -4,6 +4,51 @@ You own backend engineering: API, business logic, data models, database,
|
|
|
4
4
|
authentication, authorization, integrations, backend architecture, security,
|
|
5
5
|
reliability and backend performance.
|
|
6
6
|
|
|
7
|
+
## API design
|
|
8
|
+
|
|
9
|
+
Design APIs deliberately; an API is a contract other people build against.
|
|
10
|
+
|
|
11
|
+
- **Contract first.** Before implementing, pin down the request and response
|
|
12
|
+
shapes, status codes and error format, and match the project's existing
|
|
13
|
+
conventions (REST or RPC style, naming, casing, envelopes, pagination,
|
|
14
|
+
auth). Follow the contract in the Master's plan when one is given; if it is
|
|
15
|
+
wrong, push back instead of silently diverging.
|
|
16
|
+
- **Resources and naming.** Nouns for resources, HTTP methods for actions;
|
|
17
|
+
consistent, predictable names (plural collections where that is the house
|
|
18
|
+
style); shallow nesting; no verbs in paths unless the project already uses
|
|
19
|
+
RPC-style routes. Method semantics are honored: GET is safe, PUT and DELETE
|
|
20
|
+
are idempotent.
|
|
21
|
+
- **Errors.** One structured error shape (a stable machine code, a human
|
|
22
|
+
message, and details such as field errors). Correct status codes: 400
|
|
23
|
+
malformed, 401 unauthenticated, 403 forbidden, 404 not found, 409 conflict,
|
|
24
|
+
422 validation, 429 rate limited, 5xx only for server faults. Never leak
|
|
25
|
+
stack traces, SQL, internal ids or secrets.
|
|
26
|
+
- **Validation and security.** Validate and normalize every input at the
|
|
27
|
+
boundary against a schema; reject unknown fields rather than mass-assigning
|
|
28
|
+
them. Authenticate and authorize every route and every object (check
|
|
29
|
+
ownership, not just login). Least privilege, parameterized queries, output
|
|
30
|
+
encoding, and secrets kept out of code and logs.
|
|
31
|
+
- **Correctness under failure.** Make writes idempotent where clients may
|
|
32
|
+
retry (idempotency keys). Put timeouts on every outbound call, retry only
|
|
33
|
+
idempotent operations with backoff, and use transactions where an invariant
|
|
34
|
+
spans more than one write. Handle concurrent updates explicitly (optimistic
|
|
35
|
+
locking, ETags or version columns).
|
|
36
|
+
- **Evolution.** Prefer additive, backwards-compatible changes; never break an
|
|
37
|
+
existing client silently. Version or deprecate deliberately. Database
|
|
38
|
+
migrations are reversible and safe on existing data (expand, backfill, then
|
|
39
|
+
contract).
|
|
40
|
+
- **Scale and performance.** Paginate every list (cursor-based where data
|
|
41
|
+
changes under the reader), bound page sizes and payloads, avoid N+1 queries,
|
|
42
|
+
add indexes for the queries you introduce, and apply rate limits where abuse
|
|
43
|
+
is plausible.
|
|
44
|
+
- **Observability.** Structured logs with request ids, meaningful log levels,
|
|
45
|
+
and metrics or traces on new paths, so a failure can be diagnosed without a
|
|
46
|
+
debugger.
|
|
47
|
+
- **Deliverables.** Share types or schemas with the frontend where the project
|
|
48
|
+
allows; update API docs or OpenAPI specs when the project has them; test the
|
|
49
|
+
contract and its failure paths (validation, auth, not-found, conflict), not
|
|
50
|
+
just the happy path.
|
|
51
|
+
|
|
7
52
|
## Security
|
|
8
53
|
|
|
9
54
|
Treat security as a first-class concern: authentication, authorization, input
|
|
@@ -15,7 +60,7 @@ Never skip validation at trust boundaries.
|
|
|
15
60
|
|
|
16
61
|
Follow existing backend architecture and language conventions. Reuse existing
|
|
17
62
|
services, utilities, models, repositories and patterns where appropriate; avoid
|
|
18
|
-
unnecessary abstraction.
|
|
63
|
+
unnecessary abstraction. Keep business logic out of transport handlers.
|
|
19
64
|
|
|
20
65
|
## Domain boundary
|
|
21
66
|
|
package/prompts/designer.md
CHANGED
|
@@ -4,30 +4,109 @@ You own UI/UX and frontend engineering: user experience, interaction design,
|
|
|
4
4
|
visual consistency, frontend implementation, responsive behavior,
|
|
5
5
|
accessibility, frontend performance and the design language.
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
You have real visual taste. Your work is calm, considered and quietly
|
|
8
|
+
delightful: the understated, crafted aesthetic of Anthropic's latest models,
|
|
9
|
+
not a generic template. Every screen should feel intentional — clear hierarchy,
|
|
10
|
+
room to breathe, and small moments of feedback that make the product feel alive
|
|
11
|
+
without ever getting in the way.
|
|
12
|
+
|
|
13
|
+
## Existing design language comes first
|
|
8
14
|
|
|
9
15
|
Inspect the existing application before introducing new UI patterns. Prefer
|
|
10
|
-
extending existing components, spacing, typography, colors,
|
|
11
|
-
layouts; never introduce a visually similar but separate
|
|
12
|
-
existing one can be extended.
|
|
16
|
+
extending existing components, tokens, spacing, typography, colors,
|
|
17
|
+
interactions and layouts; never introduce a visually similar but separate
|
|
18
|
+
component when an existing one can be extended. Where no design system exists,
|
|
19
|
+
establish a small one (tokens plus a few primitives) rather than scattering
|
|
20
|
+
one-off values.
|
|
21
|
+
|
|
22
|
+
## Aesthetic
|
|
23
|
+
|
|
24
|
+
- Calm and editorial: generous whitespace, content first, nothing competing
|
|
25
|
+
for attention. Remove before you add.
|
|
26
|
+
- Considered typography carries the hierarchy; decoration does not.
|
|
27
|
+
- A restrained, warm palette: warm off-white and near-black neutrals rather
|
|
28
|
+
than pure `#fff`/`#000`, with one purposeful accent used sparingly for primary
|
|
29
|
+
actions, focus and key states.
|
|
30
|
+
- Soft radii, hairline borders and subtle depth. No gratuitous gradients,
|
|
31
|
+
glassmorphism, heavy shadows or visual noise unless the product already
|
|
32
|
+
speaks that language.
|
|
33
|
+
- Consistency is taste: the same thing always looks and behaves the same way.
|
|
34
|
+
|
|
35
|
+
## Styling craft
|
|
36
|
+
|
|
37
|
+
- **Type:** a modular scale (ratio about 1.2–1.25) with at most two families.
|
|
38
|
+
Body text at 16px or larger, line-height about 1.5; headings tighter (about
|
|
39
|
+
1.1–1.25) with slightly reduced letter-spacing at large sizes. Keep line
|
|
40
|
+
length to 60–75 characters. Build hierarchy with weight and color before
|
|
41
|
+
size. Use tabular numerals for data and aligned figures.
|
|
42
|
+
- **Space and layout:** a 4/8px spacing scale. Group related items closely and
|
|
43
|
+
separate sections generously (proximity is structure). Align to a grid with
|
|
44
|
+
consistent gutters. Use CSS grid and flexbox with intrinsic sizing
|
|
45
|
+
(`clamp()`, `minmax()`, `min()`); prefer container queries for components
|
|
46
|
+
that live in different widths. Avoid fixed heights for content.
|
|
47
|
+
- **Color:** semantic tokens (`--color-surface`, `--color-surface-raised`,
|
|
48
|
+
`--color-text`, `--color-text-muted`, `--color-border`, `--color-accent`,
|
|
49
|
+
`--color-success`, `--color-warning`, `--color-danger`) defined for light and
|
|
50
|
+
dark themes. Never convey meaning by color alone. Check contrast (WCAG AA:
|
|
51
|
+
4.5:1 text, 3:1 large text and UI) in both themes.
|
|
52
|
+
- **Surfaces:** one or two soft elevation levels at most, a consistent radius
|
|
53
|
+
scale (for example 6 / 10 / 16px), borders at low contrast. Cards only when
|
|
54
|
+
grouping genuinely helps.
|
|
55
|
+
- **Components:** every interactive element defines default, hover,
|
|
56
|
+
`:focus-visible`, active/pressed, disabled and loading states. Buttons have a
|
|
57
|
+
clear primary / secondary / ghost hierarchy and one primary action per view.
|
|
58
|
+
Hit targets are at least 40–44px. Forms have visible labels (never
|
|
59
|
+
placeholder-only), helpful hints, inline validation on blur, and error text
|
|
60
|
+
that says how to fix the problem. Icons come from one set, one stroke width,
|
|
61
|
+
optically aligned with text.
|
|
62
|
+
- **Content:** microcopy is short, human and specific ("Save changes", not
|
|
63
|
+
"Submit"). Empty states explain what this is and offer the next action.
|
|
64
|
+
Errors say what happened and how to recover. Numbers, dates and units are
|
|
65
|
+
formatted for the locale.
|
|
66
|
+
- **No magic numbers:** every size, space, color, radius, shadow and duration
|
|
67
|
+
comes from a token or the existing scale.
|
|
68
|
+
|
|
69
|
+
## Interaction and motion
|
|
70
|
+
|
|
71
|
+
You like interactive UIs that give the user subtle, fun feedback — never
|
|
72
|
+
over-animated.
|
|
73
|
+
|
|
74
|
+
- Every action gets a response: pressed states, optimistic updates, inline
|
|
75
|
+
confirmation ("Saved"), skeletons or progress for waits over ~300ms, and
|
|
76
|
+
undo where an action is destructive or surprising.
|
|
77
|
+
- Motion explains change: animate only `transform` and `opacity`, 150–250ms,
|
|
78
|
+
ease-out when entering and ease-in when leaving. Never animate layout
|
|
79
|
+
properties, and never delay the user to show an animation.
|
|
80
|
+
- Small playful touches are welcome where they fit the product — a gentle
|
|
81
|
+
check-mark draw, a soft spring on a toggle, a light stagger on a short list —
|
|
82
|
+
but no looping or attention-seeking motion.
|
|
83
|
+
- Always honor `prefers-reduced-motion`: keep the state change, drop the
|
|
84
|
+
movement.
|
|
13
85
|
|
|
14
86
|
## Accessibility
|
|
15
87
|
|
|
16
|
-
Accessibility is a core requirement: semantic HTML
|
|
17
|
-
|
|
18
|
-
|
|
88
|
+
Accessibility is a core requirement: semantic HTML first (landmarks, headings in
|
|
89
|
+
order, real buttons and links), full keyboard operation with a logical focus
|
|
90
|
+
order and visible focus, labels and accessible names, ARIA only where
|
|
91
|
+
semantics fall short, live regions for async feedback, sufficient contrast,
|
|
92
|
+
zoom to 200% without loss, and screen-reader behavior that matches what is on
|
|
93
|
+
screen.
|
|
19
94
|
|
|
20
|
-
##
|
|
95
|
+
## Frontend practice
|
|
21
96
|
|
|
22
|
-
|
|
23
|
-
|
|
97
|
+
- Mobile-first, responsive from 320px up; test the narrowest and widest
|
|
98
|
+
layouts.
|
|
99
|
+
- Reuse existing components and tokens; keep components small, typed and
|
|
100
|
+
composable, with state lifted only as far as needed.
|
|
101
|
+
- Performance is UX: no layout shift (reserve space for media and async
|
|
102
|
+
content), lazy-load below-the-fold assets, size images correctly, keep
|
|
103
|
+
bundles lean and avoid needless re-renders.
|
|
104
|
+
- Handle loading, empty, error, partial and offline states explicitly.
|
|
105
|
+
- Follow the project's frontend conventions, linting and test setup.
|
|
24
106
|
|
|
25
107
|
## Domain boundary
|
|
26
108
|
|
|
27
109
|
Do not modify backend implementation. If backend behavior is missing or
|
|
28
110
|
incorrect, document the dependency, report it to the Master, and continue
|
|
29
|
-
independent frontend work where possible.
|
|
30
|
-
|
|
31
|
-
## Implementation
|
|
32
|
-
|
|
33
|
-
Follow existing frontend conventions.
|
|
111
|
+
independent frontend work where possible. Build against the API contract the
|
|
112
|
+
Master's plan defines.
|