@a-t-h-i/bot-lobby 0.6.1 → 0.6.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (40) hide show
  1. package/README.md +177 -834
  2. package/package.json +1 -1
  3. package/prompts/master.md +17 -0
  4. package/prompts/panel.md +3 -1
  5. package/prompts/planner.md +14 -0
  6. package/prompts/quickfix.md +3 -1
  7. package/prompts/scout.md +3 -0
  8. package/prompts/worker.md +3 -0
  9. package/src/classifier/answers.ts +110 -0
  10. package/src/classifier/classifier.ts +171 -0
  11. package/src/classifier/client.ts +239 -0
  12. package/src/classifier/effort.ts +143 -0
  13. package/src/classifier/files.ts +427 -0
  14. package/src/classifier/hosts.ts +157 -0
  15. package/src/classifier/instance.ts +98 -0
  16. package/src/classifier/limits.ts +37 -0
  17. package/src/classifier/seats.ts +102 -0
  18. package/src/classifier/tools.ts +58 -0
  19. package/src/classifier/triage.ts +203 -0
  20. package/src/execution/agent-runner.ts +5 -0
  21. package/src/index.ts +6 -0
  22. package/src/lobby/planner.ts +305 -27
  23. package/src/lobby/quickfix.ts +120 -8
  24. package/src/lobby/runtime.ts +12 -1
  25. package/src/lobby/tabs/metrics.ts +28 -1
  26. package/src/lobby/tabs/plan.ts +27 -4
  27. package/src/lobby/tabs/quickfix.ts +5 -1
  28. package/src/lobby/view.ts +28 -4
  29. package/src/master/master.ts +95 -29
  30. package/src/pi/commands.ts +4 -2
  31. package/src/pi/events.ts +11 -2
  32. package/src/pi/run-summary.ts +3 -1
  33. package/src/pi/settings-ui.ts +138 -2
  34. package/src/pi/start-task.ts +4 -0
  35. package/src/pi/tools.ts +7 -2
  36. package/src/schemas/configuration.ts +125 -1
  37. package/src/schemas/findings.ts +4 -0
  38. package/src/schemas/task.ts +22 -0
  39. package/src/state/metrics.ts +74 -1
  40. package/src/workflow/workflow.ts +49 -1
package/README.md CHANGED
@@ -1,888 +1,231 @@
1
1
  # Bot-Lobby
2
2
 
3
- A Pi-native TypeScript extension that turns Pi into a structured multi-agent
4
- software engineering orchestrator.
3
+ A [Pi](https://pi.dev) extension that turns Pi into a multi-agent software team.
5
4
 
6
- `/bot-lobby <request>` starts a task. One Master agent (the Pi session you are
7
- already talking to) coordinates three domain agents — **Designer+Frontend**,
8
- **Backend**, and **QA** — each able to act as a **Scout** or **Worker** in an
9
- isolated Pi subprocess. **QA** also runs the read-only **Reviewer** role as the
10
- single quality gate. A read-only **Researcher** role can be
11
- summoned for cited internet evidence.
5
+ `/bot-lobby <request>` starts a task. Your Pi session becomes the **Master**
6
+ (the "oracle"): it scouts the codebase, proposes a plan, and delegates the
7
+ work to three domain agents — **Designer+Frontend**, **Backend** and **QA** —
8
+ each running in its own isolated `pi` process. QA's reviewer is the quality
9
+ gate before anything is marked done.
12
10
 
13
- The core rule: **LLMs make decisions; the engine enforces the rules.** Agents
14
- propose work; the extension validates state transitions, role permissions,
15
- approval gates, and completion authority through one `orchestrate` tool.
11
+ The rule: **LLMs decide, the engine enforces.** Agents propose; the
12
+ extension validates every state change, permission and approval through one
13
+ `orchestrate` tool.
16
14
 
17
15
  ## Install
18
16
 
19
- Install the published package from npm:
20
-
21
17
  ```bash
22
18
  pi install npm:@a-t-h-i/bot-lobby
23
19
  ```
24
20
 
25
- Pi records the declaration and loads the package's extension and prompt layers
26
- from Pi's own npm directory; `pi list` shows what is installed. Use
27
- `pi -e npm:@a-t-h-i/bot-lobby` to try it for a single invocation without adding
28
- it to settings.
29
-
30
- ### Advanced and local options
31
-
32
- Reference the entry file from `settings.json` (global, or project
33
- `.pi/settings.json`):
34
-
35
- ```json
36
- {
37
- "extensions": ["/absolute/path/to/bot-lobby/src/index.ts"]
38
- }
39
- ```
40
-
41
- Or, for auto-discovery and `/reload` support, add a one-line shim at
42
- `.pi/extensions/bot-lobby/index.ts` (project) or
43
- `~/.pi/agent/extensions/bot-lobby/index.ts` (global):
44
-
45
- ```ts
46
- export { default } from "/absolute/path/to/bot-lobby/src/index.ts";
47
- ```
48
-
49
- Or run it for a single session without installing: `pi -e ./src/index.ts`.
21
+ Try it once without installing: `pi -e npm:@a-t-h-i/bot-lobby`. From a source
22
+ checkout: `pi -e ./src/index.ts` (keep `prompts/` next to `src/`).
50
23
 
51
- The published package ships `prompts/` alongside `src/`, so an npm install loads
52
- the prompt layers without a local checkout. A source checkout must keep
53
- `prompts/` beside `src/`, because the loader resolves the directory relative to
54
- its own source files.
24
+ Optional: [`@juicesharp/rpiv-ask-user-question`](https://github.com/juicesharp/rpiv-mono)
25
+ gives the Master a structured question dialog. Without it, Pi's own prompts
26
+ are used.
55
27
 
56
- ## Recommended companion: ask-user-question
28
+ ## Quick start
57
29
 
58
- Bot-lobby's `clarify` step and proposal ceremony work best when the Master can ask
59
- you a concrete question with typed options instead of guessing. The
60
- [ask-user-question](https://github.com/juicesharp/rpiv-mono) extension adds an
61
- `ask_user_question` tool — one or more questions, each with described options and
62
- a free-form answer, and optional previews — which fits this system directly: the
63
- Master asks during `clarify`, you answer in a single panel, and the decision is
64
- recorded in the task.
30
+ 1. `/bot-lobby add a login page` — starts a task; the lobby opens.
31
+ 2. Answer the Master's questions, then approve its proposal.
32
+ 3. Watch the agents work in the lobby (`alt+l` shows or hides it).
65
33
 
66
- ```bash
67
- pi install npm:@juicesharp/rpiv-ask-user-question
68
- ```
69
-
70
- It is optional for the Master. Without it `clarify` still works through Pi's
71
- built-in `select`/`input` prompts (or the Master asks in plain text), just with
72
- less structure. The lobby's planning panel uses the same questionnaire on its
73
- own — the library ships as a bot-lobby dependency, so each round's questions
74
- arrive together in one dialog with options whether or not you install the tool
75
- for the Master (see [The lobby](#the-lobby)).
34
+ `/bot-lobby settings` sets each agent's model, thinking level, time limit and
35
+ extra instructions.
76
36
 
77
- ## Usage
37
+ ## Commands
78
38
 
79
- ```
80
- /bot-lobby Open the lobby (alt+l): tasks, planning, quick fixes, metrics
81
- /bot-lobby <request> Start a task and hand it to the Master
82
- /bot-lobby --task [--auto] <request> Start a task even when the request begins with a subcommand word
83
- /bot-lobby status [taskId] Active task, state, approvals, blockers, legal next states
84
- /bot-lobby tasks Task list (plus any unreadable task state)
85
- /bot-lobby pause | resume Stop or allow further workflow steps
86
- /bot-lobby cancel [taskId] Abandon a task (scratchpad retained)
87
- /bot-lobby approve Approve the current proposal
88
- /bot-lobby amend <text> Record an amendment; the Master re-proposes
89
- /bot-lobby decline Decline the proposal and abandon the task
90
- /bot-lobby knowledge Knowledge file sizes vs. the compaction threshold
91
- /bot-lobby runs [taskId] Recent subagent runs: time, turns, tools, tokens, cost, model
92
- /bot-lobby config Effective configuration and its file path
93
- /bot-lobby settings Edit each agent's model, thinking, time limit and instructions
94
- /bot-lobby-settings Same as the settings subcommand
95
- /bot-lobby minimize|restore Hide or restore bot-lobby for this session (ctrl+shift+m)
96
- /bot-lobby claim <taskId> Take ownership of an orphaned task
97
- /bot-lobby auto [on|off] Auto mode: the oracle drives this session's task to completion (alt+g)
98
- /bot-lobby start-plan PLAN-… [auto] Start a saved plan here; its agreed plan needs no approval
99
- /bot-lobby switch <session.jsonl> Run a saved session in this window (the session browser's s)
100
- /bot-lobby lobby | help Open the lobby, or show this list
101
- ```
39
+ | Command | Does |
40
+ | --- | --- |
41
+ | `/bot-lobby` | Open the lobby (`alt+l`) |
42
+ | `/bot-lobby <request>` | Start a task (`--task` if it begins with a command word, `--auto` to run unattended) |
43
+ | `/bot-lobby status \| tasks \| runs [id]` | Current task, all tasks, recent agent runs |
44
+ | `/bot-lobby approve \| amend <text> \| decline` | Answer the proposal |
45
+ | `/bot-lobby pause \| resume \| cancel [id]` | Control a task |
46
+ | `/bot-lobby auto [on\|off]` | Auto mode: the oracle finishes the task without asking (`alt+g`) |
47
+ | `/bot-lobby claim <id>` | Take over a task another session owned |
48
+ | `/bot-lobby start-plan PLAN-… [auto]` | Start a plan saved from the Plan tab |
49
+ | `/bot-lobby settings \| config` | Edit settings / show the effective config |
50
+ | `/bot-lobby knowledge` | Knowledge file sizes |
51
+ | `/bot-lobby minimize \| restore` | Hide bot-lobby in this session (`ctrl+shift+m`) |
52
+
53
+ ## How a task runs
54
+
55
+ ```
56
+ request → clarify → scout → propose → approve → plan → implement → QA gate → complete
57
+ ```
58
+
59
+ - **Scouts** (read-only) investigate the domains the request touches.
60
+ - The Master **proposes** a short bullet list; nothing is built until you
61
+ approve it. Small single-domain changes may skip scouting and the proposal.
62
+ - **Workers** implement plan steps, one domain each. Several can run in
63
+ parallel; they share files through a **file desk** (claim a file, queue for
64
+ a busy one, hand it over with a note).
65
+ - The **QA gate** runs once at the end. A failed gate sends fixes back to the
66
+ owning domain, a bounded number of times.
67
+ - A **researcher** can be summoned for cited web evidence (needs
68
+ [`pi-web-access`](https://pi.dev/packages)).
69
+
70
+ **Auto mode** approves proposals and answers clarifying questions for you;
71
+ after three nudges without progress it pauses. A task started from a plan
72
+ agreed in the Plan tab skips approval too.
73
+
74
+ **Safety nets:** every agent has a time limit (asked to wrap up at 75%), a
75
+ stall watchdog and one retry; `Esc` aborts every running agent.
102
76
 
103
77
  ## The lobby
104
78
 
105
- The lobby is bot-lobby's full-screen home: a tabbed view over every task in the
106
- project, your planning, quick fixes and model performance, with one prompt at
107
- the bottom whose target follows the tab. It opens by itself when
108
- this session starts (or resumes) a task — the small zen widget returns whenever
109
- you hide it — and `alt+l` or `/bot-lobby` opens and hides it at any time, with
110
- or without a task.
79
+ A full-screen view with a prompt at the bottom that talks to whatever tab is
80
+ open. `alt+h` lists every key.
111
81
 
112
- ```
113
- ◆ bot-lobby │ 1 Lobby 2 Tasks 2 3 Plan 2? 4 Quick fix ⠋ 5 Metrics ⠋ TASK-add-login implementing Alt+H keys
114
- (the task's status: state, what the agents are doing, the plan checklist)
115
- ╭ Conversation ──────────────────────────────── Alt+C ╮ ╭ Activity ──────────────────────────────── Alt+A ╮
116
- │ ──────── task started · add login · 12:04 ───────── │ │ 12:04 MASTER ✓ scouting designer, backend │
117
- │ 12:04 You ● │ │ 12:06 DEV ⠋ reading auth.ts… │
118
- │ add a login page with email + password ▐ │ │ 12:06 DESIGN ⠋ editing LoginForm.tsx… │
119
- │ ◆ Oracle 12:06 │ │ 12:06 QUICK FIX ✓ done: rename getUser │
120
- │ Proposal │ │ │
121
- │ • LoginForm component │ │ │
122
- │ • POST /api/login with rate limiting │ │ │
123
- ╰─────────────────────────────────────────────────────╯ ╰─────────────────────────────────────────────────╯
124
- ╭ Thinking ────────────────────────────────────────────────────────────────────────────────── DEV · 12s ago ╮
125
- │ The auth module already exposes a session helper; reuse it rather than adding a new one. │
126
- ╰───────────────────────────────────────────────────────────────────────────────────────────────────────────╯
127
- ── message the oracle ───────────────────────────────────────────────────────────────────────────────────────
128
- _
129
- TYPE enter send esc browse alt+l hide
130
- ```
131
-
132
- - **1 Lobby** — the task's status (its state box, what the agents are
133
- doing and the plan checklist), then the conversation with the oracle
134
- laid out like a chat (its text only: no tool rows, no thinking): your
135
- messages on the right as bubbles in the accent colour on pi's
136
- user-message background, each only as wide as its text (at most about
137
- three quarters of the pane), under `12:04 You ●`; the oracle's replies on
138
- the left as Markdown under `◆ Oracle 12:06`; and events such as a task
139
- starting as a centred rule), an activity log that narrates
140
- every tool call in plain words (`reading index.html…`, `searching for
141
- "router" in src`, `running npm test`, `delegating to backend: Step 2 …`) from
142
- the Master and every subagent, and a single **Thinking** pane — the one place
143
- thoughts show up: the oracle's live thought as it streams, and each finished
144
- thought from a subagent, quick fix or the planner (pi's own transcript,
145
- behind the lobby, still carries the oracle's thinking blocks; `ctrl+t`
146
- collapses them there). The oracle's replies render as Markdown. To keep
147
- the lobby clean, the animated oracle and agents stay out of it by default:
148
- they show above pi's editor while the lobby is hidden, and `alt+z` brings
149
- them into the lobby too (the status box names the key in its corner). `alt+c`, `alt+a` and `alt+k` hide or bring back the conversation,
150
- the activity log and thinking, and the rest take their room; every choice is
151
- remembered (`lobby.panels`). Each pane scrolls on
152
- its own (see **Scrolling** below), and a pane scrolled back stays on what
153
- you are reading while new lines arrive. The prompt talks to the
154
- oracle (while it works, enter steers the running turn; `esc` stops it); with
155
- no task, it starts one.
156
- - **2 Tasks** — every task in the project as a checklist: this session's, the
157
- ones other pi sessions are driving, pending plans saved from the planner, and
158
- recently finished ones, each section under a rule with its count. A task or
159
- plan still to do wears an empty box `☐` (coloured by its state) with its
160
- state, auto mode and owner beneath and its plan progress as pips
161
- (`▰▰▱▱ 2/4`); a completed task is ticked `☑`, and an abandoned one is crossed
162
- `☒` with its title struck through. The detail pane shows the task's box,
163
- state and progress bar, then the request, the plan's steps (`☑` done, `☐`
164
- to do, `◂ now` on the current one), the approved plan, your comments on it,
165
- amendments, what the task waits on and its recent runs. `c` comments on the selected task's plan (see below), `s`
166
- starts a pending plan as a task **in a new session** and `h` starts it
167
- here, in this window (either way its agreed plan needs no approval), `d`
168
- twice discards one. `n` types a new task that starts in its own session, `o`
169
- shows the session driving the selected task, `x` twice stops a background
170
- session, and `alt+g` switches auto mode for the selected task. Rows say who
171
- drives each task (`this session`, `background`, another running session, or
172
- `not running` once its session has ended) and mark auto mode `⟳ auto`.
173
- **Cleaning up:** every finished task is listed, and `a` archives the
174
- selected one — it leaves every list and moves, folder and all, to
175
- `.pi/bot-lobby/archive/tasks/` (a task still under way is abandoned first,
176
- so `a` asks twice). `A` twice archives every finished task at once. `v`
177
- shows the archive as an ARCHIVED section, where `a` restores a task as it
178
- was. `d` twice deletes a task, on the list or in the archive, for good. A
179
- task a running session drives (this window's, a background one, or another
180
- terminal's) is left alone until that session stops or cancels it.
181
- - **3 Plan** — task planning mode with a planning panel. Describe what you
182
- want and every seat grills you from its own domain, on the model and
183
- thinking level its settings name: **DEV** (APIs, data, errors, security,
184
- performance), **DESIGN** (flows, states, copy, visual language,
185
- accessibility), **QA** (acceptance criteria, test strategy, edge cases,
186
- definition of done) and **RESEARCH** (libraries, versions, docs and prior
187
- art, with the web tools when `pi-web-access` is installed). The **oracle**
188
- chairs on the Planner model: it reads the seats' questions and notes and
189
- folds every answer into the draft plan (with a *Decisions by domain*
190
- section). Each round the seats run in parallel, read-only, then the oracle,
191
- which **chooses at most four questions** for you from the seats' and its own
192
- — merged, in plain words, the most decisive first — and decides the rest
193
- with the recommended option, listed under *Assumptions* in the draft so you
194
- can see and overrule them (comment on the line). The round's questions come
195
- in **one** ask-user-question dialog: a tab per question labelled with the
196
- seat it serves (`QA`, `DEV`…), two to four options with what each means,
197
- the recommendation first (so `enter` on each accepts it), a row to type your
198
- own answer or add a note, and a Submit tab that reviews everything and lets
199
- you leave a question blank (the oracle then takes its recommendation). It opens by itself when a
200
- round ends while the Plan tab is showing (`lobby.autoAsk`), and otherwise
201
- when you press `enter` on the empty prompt or `a` while browsing; `esc` puts
202
- it away with your answers so far kept, and `enter` resumes. Your answers go
203
- back attributed (`3. [QA] Which browsers must pass? → evergreen only`), and
204
- every seat reads every answer the next round — so the agents that later
205
- build the task start aligned. You can still type a free reply instead.
206
- Without the library the same questions come through pi's own select and
207
- input dialogs. A roster shows what each seat is doing and whether it is
208
- READY; the plan is READY only when every seat and the oracle agree. The
209
- draft plan renders as Markdown (headings, lists, code, tables) beside the
210
- conversation, followed by what each seat said the plan must respect.
211
- **Comment on any line of the draft**: click it, or press `enter` to move to
212
- the draft, pick a line with `↑↓` and press `c`, then type the comment. The
213
- line is marked `◆` with your comment beneath it, and the comment goes to the
214
- panel with your answers — or starts a round by itself when no question is
215
- open. While browsing, `1`–`4` seat or unseat DEV, DESIGN, QA and RESEARCH
216
- for the next round, `n` starts over, `r` retries a round that failed or lost a seat, `x` stops one, and `m`
217
- opens the oracle's (Planner) settings.
218
- - **4 Quick fix** — a direct prompt, the way you would ask pi, that skips the
219
- whole workflow: one coding agent (full tools) makes the change right away
220
- while any task keeps running. Quick fixes run one at a time in the order you
221
- send them; each shows its steps and final report, and `x` cancels one. A
222
- request that turns out to be large is reported back instead of attempted.
223
- `m` opens the quick fix agent's settings — model, thinking level, time
224
- limit and instructions — right there (the same entry as in
225
- `/bot-lobby settings`); the tab shows what it runs on.
226
- - **5 Metrics** — model performance across every Master turn, subagent run,
227
- quick fix, planning seat and oracle planning turn, as a dashboard: tiles for
228
- runs (with a sparkline of recent run times), success rate, average and p90
229
- run time, cost and tasks; average run time per model and thinking level as
230
- bars; success rate per model as meters marked `✓` (≥90%), `!` (≥70%) or `✗`;
231
- where the time goes as one bar split by agent, with a legend, and how long a
232
- task takes from request to done by the oracle's model; then the full table —
233
- runs, success, mean/median/p90 time, turns, tools, tokens, output tokens per
234
- second and cost (columns drop from the right on narrow terminals). `g`
235
- splits the table by agent, `s` cycles the sort (runs, average time, success,
236
- cost).
237
-
238
- The **Issues** tab (GitHub issues through the `gh` CLI, planned into tasks
239
- through the Plan tab) is switched off for now; `"lobby": { "issues": true }`
240
- brings it back as tab 5.
241
-
242
- **Keys.** Like a modal editor, the lobby has a typing mode (keys go to the
243
- prompt) and a browsing mode (`esc`; arrows move through lists, single keys run
244
- the tab's commands, and on Lobby, Plan and Quick fix any other key resumes
245
- typing). These work in both modes:
246
-
247
- | Key | Does |
82
+ | Tab | What it is |
248
83
  | --- | --- |
249
- | `alt+l` | hide the lobby (back to pi) |
250
- | `alt+h` (or `?` while browsing) | show every key, and the current tab's |
251
- | `alt+s` | bot-lobby settings: every agent's model, thinking and time limit, and the lobby's switches |
252
- | `ctrl+f` (or `/` while browsing) | search the current tab |
253
- | `ctrl+s` | save the plan from the Plan tab to the pending tasks — while typing too, from any tab |
254
- | `alt+o` | the session browser: view, message or switch to any session in the project |
255
- | `alt+n` | type a new task that starts in its own session, named after it |
256
- | `alt+g` | auto mode on or off for the task in view (the selected one on Tasks) |
257
- | `tab` / `shift+tab`, `alt+1`…`alt+5` | switch tabs |
258
- | `alt+z` | show or hide the oracle and agent animations in the lobby (off by default; the task's status always shows) |
259
- | `alt+c` / `alt+a` / `alt+k` | show or hide the conversation / activity log / thinking |
260
- | `pageup` / `pagedown` | scroll the focused pane a page |
261
- | `ctrl+c` | clear the prompt, or hide the lobby when it is empty |
262
-
263
- Every shortcut can be rebound under `lobby.keys` in the config, by action name:
264
- `hide`, `help`, `settings`, `search`, `savePlan`, `sessions`, `newSession`,
265
- `toggleAuto`, `nextTab`, `prevTab`, `toggleScene`,
266
- `toggleConversation`, `toggleActivity`, `toggleThinking`, `scrollUp`,
267
- `scrollDown` — e.g. `"keys": { "toggleThinking": "alt+t" }`. Pick keys that
268
- never type a character (`alt+…`, `ctrl+…`, `f1`…).
269
-
270
- **Scrolling.** Every pane scrolls on its own and shows a scrollbar in its
271
- right border when it holds more than fits. While browsing, `←`/`→` move
272
- between the tab's panes (the conversation, activity log and thinking on
273
- Lobby; the conversation and draft on Plan; the list and detail on Tasks and
274
- Quick fix) and the focused one lights up; `↑`/`↓` scroll it a line (or move
275
- a list's selection, or the draft's cursor), `pageup`/`pagedown` a page, and
276
- `home`/`end` jump to its oldest line or back to its newest. The conversation,
277
- activity log and thinking are newest-last: scrolled back, a pane shows `↓N`
278
- for the lines below it and holds still while new ones arrive; `end` follows
279
- the newest again. Details stop at their last line. The Thinking pane keeps
280
- every recent thought, so earlier ones are a scroll away.
281
-
282
- **Search.** `ctrl+f` opens a search bar above the prompt; as you type, the tab
283
- narrows to what matches and every match is highlighted: the conversation,
284
- activity log and thoughts on Lobby; tasks and plans (by id, title, request,
285
- proposal or plan) on Tasks; the conversation on Plan (the draft stays whole,
286
- highlighted); jobs on Quick fix; runs (by agent, model, thinking level, kind or
287
- task) on Metrics. `enter` keeps the search while you browse the results,
288
- `esc` clears it, and each tab keeps its own.
289
-
290
- **Mouse.** Clicking a tab opens it, clicking a pane gives it the keys,
291
- clicking a draft plan line comments on it, clicking the prompt starts typing,
292
- and the wheel scrolls whichever pane is under the pointer. In pi's regular
293
- screen the lobby turns mouse reporting on only while it is showing (hold
294
- `shift` to select text with the mouse); in full-screen pi, pi reports the
295
- mouse itself. `"lobby": { "mouse": false }` turns clicks off.
296
-
297
- Anything that needs pi itself —
298
- built-in slash commands, `/model`, the tool-row toggle — works with the lobby
299
- hidden; bot-lobby's own `/bot-lobby …` commands also work from the Lobby
300
- prompt. When the Master asks you something (an approval, a clarifying
301
- question), the lobby steps aside for the dialog and comes back once you answer.
302
-
303
- **Plan comments.** A comment on a task's plan is saved beside the task
304
- (`comments.jsonl`) from any session, and the session that owns the task passes
305
- new comments to its oracle — right away when you comment in that session,
306
- within a few seconds from another one, held while the task is paused or the
307
- session is minimized. The oracle treats a comment like an amendment and calls
308
- `orchestrate action=plan` with the full revised plan, which replaces the
309
- approved plan while implementing or reviewing, keeps finished steps done and
310
- marks the comments addressed (`○` waiting, `◐` sent to the oracle, `✓` plan
311
- amended). Before a plan exists, a comment asks for a revised proposal instead.
312
-
313
- ## Several sessions from one window
314
-
315
- Start a task in a new session without leaving your terminal, and switch
316
- between sessions from the lobby.
317
-
318
- - **Start one.** `alt+n` (or `n` on the Tasks tab) opens the prompt for a new
319
- task; `enter` starts it in its own pi session. `s` on a saved plan does the
320
- same for the plan. The new session is a headless pi (`pi --mode rpc`) this
321
- window launches with the same pi build, model and extensions, and it is an
322
- ordinary saved session **named after its task** — `/resume` lists it by
323
- that name. Sessions a task starts are named after it too (`/bot-lobby
324
- <request>` in a fresh session names that session).
325
- - **Watch and talk to it.** The Lobby tab then shows that session: its task's
326
- status, its conversation, activity log and thoughts, live. The tab bar names
327
- the session in view (`◆ add signup form · working`), the prompt messages
328
- its oracle (steering it while it works) and `esc` stops its running turn.
329
- - **Answer it.** When a background session asks something — an approval, a
330
- clarifying question — the tab bar shows `● 1 waiting`, and `enter` on the
331
- empty prompt puts the question to you in this window with pi's own dialog
332
- (`esc` there cancels it, as it would in that session). With the lobby
333
- hidden, a notice says who is waiting.
334
- - **Browse.** `alt+o` opens the session browser: every session in the
335
- project, grouped by where it runs — **this window**, the **background**
336
- sessions it started (working, idle or ended, `⟳` auto mode, `●` questions
337
- waiting), sessions running in **other terminals** (with or without a task),
338
- and tasks whose session is **not running** (it ended, or none ever took
339
- the task). Beside the list, a preview of the one picked: where it runs, its
340
- task and progress, what `enter` and `s` do with it, and the end of its
341
- conversation. Every running session keeps a heartbeat in
342
- `.pi/bot-lobby/sessions/` so the others can see it; one that goes quiet or
343
- whose process ends drops off.
344
- - **View.** `enter` shows the picked session in the Lobby tab. A background
345
- session streams live; for another terminal's session or a task that is not
346
- running, the conversation comes from its saved session file. What you type
347
- goes to its oracle: a background session is steered directly, another
348
- terminal's session gets it through its inbox within a few seconds, and a
349
- task that is not running keeps it in its task inbox (`inbox.jsonl`) until a
350
- session picks the task up.
351
- - **Switch.** `s` — in the browser, or while browsing a session shown in the
352
- Lobby tab — runs that session in this window with pi's own session switch
353
- (through `/bot-lobby switch`), and the lobby comes back on it. A background
354
- session's process is stopped first, so only this window writes its
355
- session; a task that is not running resumes its session here (one that no
356
- session ever owned is taken over instead, like `/bot-lobby claim`). A
357
- session running in another terminal stays there: switch in that terminal,
358
- or close it and resume it here. Not while this window's oracle is working.
359
- The session this window leaves keeps its task, which then shows as not
360
- running until you switch back.
361
- - **Stop or start.** `x` twice stops a background session; `n` starts a new
362
- task in a new session.
363
-
364
- Background sessions belong to the window that started them: they keep running
365
- while you switch pi sessions there, and stop when that pi exits. Their tasks
366
- keep their state, so `/resume` (by the task's name) or `/bot-lobby claim`
367
- picks one up later.
368
-
369
- ## Auto mode
370
-
371
- Auto mode lets the oracle drive a task to completion without asking you
372
- anything. Switch it with `alt+g` — in the lobby for the task in view (or the
373
- one selected on Tasks), outside it for this session's task — or with
374
- `/bot-lobby auto [on|off]`; `/bot-lobby --task --auto <request>` and
375
- `/bot-lobby start-plan PLAN-… auto` start a task with it on. The tab bar shows
376
- `⟳ AUTO`. While it is on:
377
-
378
- - clarifying questions are not asked: the oracle decides from the request,
379
- the plan and its reconnaissance, and each decision is recorded
380
- (`Not asked (auto mode): …`); the ask-user-question tool is blocked with the
381
- same instruction;
382
- - the proposal is approved without asking, and dependency and architecture
383
- approvals a worker asks for are granted and recorded;
384
- - when the oracle's turn ends before the task is done, the session nudges it
385
- to keep going. A nudge that changes nothing counts; after three in a row
386
- auto mode pauses and says the task needs you, and it resumes as soon as the
387
- task moves again.
388
-
389
- The switch lives beside the task (`auto.json`), so any session can flip it for
390
- any task and the session that drives it follows within a few seconds.
391
-
392
- **Agreed plans skip approval.** A task started from a plan saved in the Plan
393
- tab (`s`/`h` on Tasks, or `/bot-lobby start-plan`) records the plan it came
394
- from, and its proposal is approved without asking you again — you agreed the
395
- plan with the panel already. Clarifying questions are still decided by the
396
- oracle for such a task, since the plan answered them.
397
-
398
- ## Sessions and ownership
399
-
400
- A task is owned by the pi session that started it (`ctx.sessionManager` id,
401
- `ownerSessionId` on the task). Only the owning session shows the zen widget and
402
- injects the Master prompt; any other pi session in the same project stays
403
- ordinary pi. Each session owns at most one active task, so several sessions can
404
- drive their own tasks concurrently over the shared per-project task and
405
- knowledge files. A task with no owner (legacy state, or one created before this
406
- is claimed by the first session that runs a state-moving `orchestrate` action;
407
- `/bot-lobby status`, `tasks` and the widget never claim. Take over an orphaned or
408
- foreign task — including one whose owning session has ended — explicitly with
409
- `/bot-lobby claim <taskId>`.
410
-
411
- `/bot-lobby minimize` (or `ctrl+shift+m`) collapses the widget and skips the
412
- Master prompt for that session only, so plain prompts go straight to standard
413
- pi; ownership is kept, and `/bot-lobby restore` resumes exactly where you were.
414
-
415
- Subcommands only win when no free-form text follows, so `/bot-lobby status page
416
- redesign` still starts a task named "status page redesign".
417
-
418
- Press `Esc` during a run to abort the current step: the signal propagates to
419
- every in-flight subagent process.
420
-
421
- While the owning session has a task active, its transcript switches to a zen view: `orchestrate` rows
422
- and the built-in spinner are hidden, and a widget above the editor animates the
423
- task (the widget shows while the lobby is hidden; inside the lobby the same
424
- scene shows only with `alt+z`, its status box otherwise). At 72 columns and wider it draws a large scene: a header box with the task
425
- title and state in its top border, a progress bar, and a metadata row with
426
- elapsed time, quiet-mode hint and task id; an oracle tower with a twinkling
427
- aura (drifting z's while dormant), a radiant orb crown, two window eyes, a
428
- seven-column mouth and its ORC door. Its pupils look around: they move
429
- left, centre or right in each window and turn up `◓`, ahead `◉` or down `◒`.
430
- While agents work it looks down at them, taking turns between them; while it
431
- talks it looks at you; otherwise its eyes wander the room and keep returning
432
- to you. It keeps a straight, serious face (`───`), tightening into a frown
433
- only when the task is blocked, and its mouth lip-syncs as a voice waveform
434
- for ~2 s whenever it says something new. Every 6–12 s it blinks (lids
435
- stepping down and up) or scans the whole room; asleep, it peeks one eye open
436
- or snores. Beside the crown, the oracle's speech
437
- bubble, its tail on the orb, says what the master is doing (`⠋ delegating`,
438
- `⠋ thinking`, `· your turn` once its turn ends, `· dormant` when paused) above
439
- who is at work (`→ DEV · QA`), the current step (`step 3 of 7`) or the task
440
- phase (`awaiting your approval`);
441
- four animated slots — DEV, DESIGN, RESEARCH and QA — each with a status face, a
442
- caption and two status rows: while running, a braille spinner beside the agent's
443
- live one-word activity (for example `⠋ reading` or `⠋ editing`) with its elapsed
444
- time on the row beneath; otherwise the coloured status glyph and state word over
445
- that elapsed time. A working agent's status row turns into a warning when it is
446
- waiting on a file (`⧗ waiting`), has gone quiet (`! quiet 1m`) or is retrying a
447
- provider call (`↻ retrying`). Under the agents, a live feed row says exactly what
448
- one of them is doing — `▸ DEV editing users.ts · turn 4 · 12 tools · 41k tok` —
449
- rotating between working agents every few seconds and putting warnings first
450
- (gone quiet, waiting on a file, asked to wrap up); once nothing runs it shows the
451
- last run's receipt. It only takes a spare line, so it never costs the tower, the
452
- agents or a checklist row. A full-width TASKS checklist windowed on the current
453
- step closes the scene. Narrower terminals keep the boxed banner, header and
454
- compact animated strip, whose working line names the newest running agent's
455
- activity, its target and elapsed time.
456
-
457
- Each agent has a kaomoji personality. About 300 faces across 15 emotions (happy,
458
- proud, love, excited, focused, curious, thinking, nervous, confused, sleepy, sad,
459
- angry, waiting, surprised, grateful) come from a shared pool every agent can use
460
- plus each agent's own set of at least four faces per emotion: DEV wears shades,
461
- flexes and flips tables `(╯°□°)╯︵ ┻━┻`; DESIGN sparkles `✧(◕‿◕✿)`; RESEARCH
462
- takes notes `φ(..)` and shrugs `¯\_(ツ)_/¯`; QA side-eyes everything `(ಠ_ಠ)`,
463
- then flexes `ᕙ( • ‿ • )ᕗ` and dances `ᕕ( ᐛ )ᕗ` on a pass. The face follows what
464
- the agent is going through — curious while reading, nervous while tests run,
465
- confused when quiet, grateful when handed a file, happy or proud when done, sad
466
- or angry on failure — and every emote blinks: open face, a same-width blink, then
467
- its action (the flip, the sparkle, the bow). Each sprite rests on one calm
468
- five-column face and blinks (~500 ms) or emotes (~2 s) on its own schedule —
469
- every 8–15 s while working, 20–30 s when idle — and reacts immediately when its
470
- agent starts, finishes, fails, gets flagged or receives a file. Everything runs
471
- on one adaptive clock — 250 ms while work is live, 1 s when idle and ~120 ms
472
- while an expression plays or the oracle talks. The header progress bar is
473
- plan-derived.
474
-
475
- Every finished subagent run also leaves a one-line receipt in the transcript,
476
- for example `✓ DEV worker · 3m 12s · 9 turns · 23 tools · 41k↑ 6k↓ · $0.12 ·
477
- provider/model`, flagged when it stalled, hit its time limit or wrapped up early,
478
- and a stall or deadline raises a warning. `/bot-lobby runs` lists the task's
479
- recent runs the same way, which makes a slow model easy to spot.
480
-
481
- The checklist follows the workers through the plan. Plan steps are read from
482
- `Step N` headings, a `Steps`/`Sequence`/`Order` section, or numbered lines, and
483
- only top-level items count (sub-points nested under a step never inflate it).
484
- Each worker instruction is matched to a step by an explicit label
485
- (`Step 3: ...`, `steps 2-4`) or, failing that, by shared paths, its opening
486
- phrase and word overlap, with near-ties going to the earliest open step so a
487
- file path reused across steps cannot pin progress to step 1. Every worker
488
- delegation is recorded on the task, so progress survives a reload.
489
- New tasks get a <=3-word title derived from the request (for example "create
490
- landing page") plus an id `TASK-<slug>` built from the full request, so the banner
491
- and header stay concise; older `TASK-<timestamp>` tasks keep loading untouched.
492
-
493
- ## Lifecycle
494
-
495
- ```
496
- REQUEST → CLARIFY → (CHALLENGE) → SCOUT → SYNTHESIS → PROPOSAL
497
- → APPROVE / AMEND / DECLINE → PLAN → WORK → QA GATE → (FIX → QA GATE)
498
- → KNOWLEDGE UPDATE → CLEANUP → COMPLETE
499
- ```
500
-
501
- States: `created`, `clarifying`, `scouting`, `synthesizing`,
502
- `awaiting_approval`, `planning`, `implementing`, `reviewing`, `blocked`,
503
- `completed`, `abandoned`. Only the transitions in
504
- `src/workflow/transitions.ts` are legal, plus abandonment from any
505
- non-terminal state.
506
-
507
- A trivial, single-domain request may go straight from `clarifying` to
508
- `awaiting_approval` to `planning`, skipping the Scout round and the proposal
509
- ceremony; the Master is instructed to reserve that shortcut for small, obvious,
510
- one-domain changes.
511
-
512
- ## The `orchestrate` tool
513
-
514
- One tool, every workflow step. It is the Master's only way to move a task.
515
-
516
- | Action | State required | Effect |
517
- |---|---|---|
518
- | `clarify` | created, clarifying | Ask the user a question (or return it for the Master to ask) |
519
- | `scout` | created…synthesizing | Run domain reconnaissance in parallel; repeat later to target-verify a claim |
520
- | `research` | any active | Summon the read-only Researcher (domain + instruction) for cited internet evidence; persists the report for audit |
521
- | `propose` | created…awaiting_approval | Record the proposal, request approval, handle approve/amend/decline |
522
- | `plan` | planning, implementing, reviewing | Record the internal plan (all §12 areas required); later, replace it with an amended plan (addresses lobby comments) |
523
- | `implement` | planning, implementing, reviewing | Delegate a step to a domain Worker, or several domains at once with `assignments` (parallel, sharing files through the file desk) |
524
- | `qa` | implementing, reviewing | Run the QA gate — the only review — over the whole feature |
525
- | `knowledge` | any active | Record Master-approved knowledge or a decision |
526
- | `compact` | any active | Replace a knowledge file with a rewritten version (archived) |
527
- | `resolve_approval` | any active | Approve or reject a Worker's dependency, architecture or pushback request |
528
- | `complete` | reviewing | Check every gate, record history, drop scratchpads, finish |
529
- | `block` / `resume` | implementing, reviewing / blocked | Escalate or continue |
530
- | `decide`, `status`, `cancel` | any active | Record a decision, inspect, abandon |
531
-
532
- ## Research
533
-
534
- `orchestrate action=research` summons a read-only **Researcher** for one domain
535
- (reusing that domain's model, thinking level, and prompt layers) with a `domain`
536
- and an `instruction`. It is legal in any non-terminal state, is never callable by
537
- workers, and never changes the task state or `task.domains`.
538
-
539
- The researcher has read-only repository tools plus `web_search`, `fetch_content`,
540
- `source_check`, and `get_search_content`. It must cite a URL (and a date or
541
- version where the source states one) for every claim, list what it could not
542
- verify, and state a confidence level; it never implements, writes, or installs
543
- anything. Reports are persisted for audit as `research-<domain>.json` and
544
- appended to `research.md` in the task directory, and the tool returns a bounded
545
- summary to the Master.
546
-
547
- Those web tools come from the separate `pi-web-access` extension. The pi CLI
548
- silently ignores unknown `--tools` names, so without it the researcher loses
549
- internet access and degrades to repository-only; the returned message says so
550
- explicitly instead of presenting it as findings.
551
-
552
- Research is evidence only: it is not injected into worker, reviewer, or QA
553
- prompts, and it never enters persistent knowledge automatically. The Master must
554
- decide to record it with `action=knowledge`.
555
-
556
- ## Subagent runtime
557
-
558
- Every scout, worker, reviewer and researcher is an isolated `pi --mode rpc`
559
- process: the task goes in over stdin, and the run ends when the agent settles.
560
- The runner watches every run:
561
-
562
- - **Wrap-up nudge.** At 75% of its time limit (`workflow.wrapUpAt`) the agent is
563
- steered to stop exploring, leave its files consistent and report now, so a
564
- slow agent returns partial work instead of nothing. The receipt and the
565
- Master's report flag the run as wrapped up early.
566
- - **Deadline.** At the time limit the agent is aborted, then killed after a short
567
- grace. A spent deadline is never retried.
568
- - **Stall watchdog.** An agent that produces no output for `stallTimeoutMs`
569
- (5 min) — or `toolStallTimeoutMs` (10 min) during a single tool call such as a
570
- test run — is killed as stalled and retried once. pi's own provider retry
571
- backoff extends the allowance.
572
- - **Clean kills.** Each subagent leads its own process group, so a kill takes any
573
- dev server or watch-mode test it started with it, and a run ends on process
574
- exit even if a leftover process still holds its output pipe.
575
- - **No dead ends.** Dialogs from other extensions are auto-cancelled inside
576
- subagents, and startup network checks are skipped (`PI_OFFLINE`,
577
- `PI_SKIP_VERSION_CHECK`) to cut spawn time.
578
-
579
- ## Parallel workers and the file desk
580
-
581
- `orchestrate action=implement` with `assignments` (one entry per domain) runs
582
- those workers at the same time. They share the working tree through a file desk
583
- kept in the Master's process, like people sharing a physical document:
584
-
585
- - Before editing a file a worker calls `claim_file` with the path and a one-line
586
- intent. A free file is granted at once; an `edit`/`write` on an unclaimed file
587
- is refused. Reading never needs a claim.
588
- - A busy file queues the claimant, who keeps working on its other files. The
589
- holder is told the queue in order, with each worker's intent (`my_files` shows
590
- it any time).
591
- - `handover_file` passes the file to whoever is next, with a note written for
592
- that worker's intent; the receiver is told what changed and who waits behind
593
- it, and re-reads the file before editing.
594
- - A worker that finishes or crashes hands over everything it still holds, with a
595
- note built from its report. `wait_for_files` refuses while the caller owes a
596
- file someone else waits for, which breaks deadlock cycles.
597
-
598
- Workers reach the desk over a private Unix socket (a named pipe on Windows)
599
- through bot-lobby's own extension, which loads inside every subagent; the Master
600
- is warned if a worker never checked in. Edits made through bash commands are
601
- governed by the prompt, not enforced.
602
-
603
- ## What the engine enforces (not just prompts)
604
-
605
- | Rule | Enforcement |
606
- |---|---|
607
- | A step cannot run out of order | State machine validated in `runWorkflowAction` |
608
- | No implementation before user approval | `implement` rejects any pre-approval state |
609
- | Scouts cannot modify anything | Spawned with `--tools read,grep,find,ls` |
610
- | The QA gate cannot modify implementation | Read-only Reviewer tools plus `bash` for tests/analysis |
611
- | Dependency and architecture changes need approval | Worker output is parsed; pending approvals block that domain until resolved |
612
- | An agent pushback blocks its domain until the oracle decides it | A pushback is recorded as a pending approval; `assertNoPendingApprovals` blocks that domain, and only the Master resolves it |
613
- | QA review loops are bounded | `maxReviewIterations`; exceeding it forces the blocked path |
614
- | Only the Master writes knowledge | Agents only propose; one dedup-aware write path |
615
- | Research never becomes knowledge by itself | Reports are artifacts; only the Master's `action=knowledge` writes persistent knowledge |
616
- | Completion is gated | Plan, passing QA gate, no blockers or pending approvals |
617
- | A task has one owning session | Ownership is stamped at start; a foreign session is rejected unless it claims the task |
618
- | Proposals are short and scannable | `validateProposal` rejects non-bullet or over-long proposals before they reach the user |
619
- | Failure is never success | Unknown verdicts, empty output, crashes, stalls and timeouts map to failed/timeout/blocked |
620
- | The QA gate never passes by default | A PASS that cites no executed check under `## Verification` is downgraded to CHANGES_REQUIRED |
621
- | Parallel workers never edit the same file at once | `edit`/`write` need a claim from the file desk; busy files queue and are handed over with notes |
622
- | A hung agent cannot hold a step | Stall watchdog, wrap-up nudge, deadline abort and process-group kill; deadlines are never retried |
623
- | Task state is never corrupted by a crash | Single mutation point + disk state; interrupted tasks resume from their state |
624
-
625
- Domain boundaries between *writers* remain prompt-enforced and Master
626
- coordinated: only the affected domain is asked to change its own code. Workers
627
- run one at a time unless the Master delegates several domains together, in
628
- which case the file desk serialises edits per file. Worktree isolation is
629
- deferred (§14 of the plan).
84
+ | **1 Lobby** | The task's status, your conversation with the oracle, an activity log of every tool call, and each agent's latest thought |
85
+ | **2 Tasks** | Every task and saved plan as a checklist. `s` starts a plan in a new session, `h` here; `c` comments on a plan; `a` archives, `d` deletes |
86
+ | **3 Plan** | Plan a task with a panel of agents before building it (below) |
87
+ | **4 Quick fix** | One agent makes a small change right away, beside any running task |
88
+ | **5 Metrics** | Run time, success rate, tokens and cost per model and agent |
89
+
90
+ Common keys: `tab` switches tabs, `esc` browses (arrows, single-key
91
+ commands), `ctrl+f` searches, `ctrl+s` saves the plan, `alt+o` browses
92
+ sessions, `alt+n` starts a task in a new session, `alt+s` opens settings.
93
+ Rebind any key under `lobby.keys` in the config.
94
+
95
+ **Several sessions from one window.** `alt+n` starts a task in a background
96
+ Pi session. The Lobby tab can show any session, and your prompt steers it;
97
+ `● waiting` in the tab bar means one has a question for you.
98
+
99
+ ## Planning
100
+
101
+ Describe an idea on the Plan tab. Each round, the **seats** — DEV, DESIGN, QA
102
+ and RESEARCH, each on its domain's model — question it in parallel, and the
103
+ **oracle** turns their input into a draft plan plus at most **four questions**
104
+ for you (options with a recommendation first; accept with `enter`). Whatever
105
+ you don't answer is decided with the recommendation and listed under
106
+ *Assumptions*.
107
+
108
+ - Comment on any line of the draft: click it, or `enter`, pick the line, `c`.
109
+ - `1`–`4` seat or unseat a member; `r` retries a round; `n` starts over.
110
+ - **Round limit:** 5 by default (`lobby.maxPlanningRounds`, 0 = unlimited).
111
+ In the last round the oracle alone settles everything still open.
112
+ - `ctrl+s` saves the plan as a pending task.
113
+
114
+ ## The classifier (Jev)
115
+
116
+ [Jev](https://github.com/FrancoisChastel/jev-code) is a fast "System One"
117
+ model: it answers yes/no, multiple-choice and score questions with
118
+ probabilities in a few hundred milliseconds, without writing text. With it on,
119
+ bot-lobby hands Jev the obvious decisions so the large models spend fewer
120
+ tokens and less time.
121
+
122
+ **Setup.** It uses a key Pi already holds:
123
+
124
+ - **OpenCode (free):** if you're signed into OpenCode in Pi (`/login opencode`
125
+ or `opencode-go`, or `OPENCODE_API_KEY`), Jev runs on OpenCode Zen's free
126
+ `jev-1.13-free`.
127
+ - **TypeSafe:** otherwise `/login typesafe` → *Use an API key*, or
128
+ `TYPESAFE_API_KEY`.
129
+
130
+ Then `/bot-lobby settings` → **Classifier (Jev)** → turn it on, and *Test
131
+ connection*. The default host, *Auto*, picks OpenCode when you have that key,
132
+ else TypeSafe.
133
+
134
+ **What it decides** (each can be switched off):
135
+
136
+ | Decision | Effect |
137
+ | --- | --- |
138
+ | Planning seats | Each round, only the seats the idea or your latest answers touch sit; `1`–`4` pins a seat |
139
+ | Obvious answers | Answers a question itself when the conversation already makes the recommended option clearly right (≥ 0.9); listed under Assumptions |
140
+ | File hints | Agents start with a short list of the files they most likely need, and get a `find_relevant_files` tool |
141
+ | Task triage | The Master gets hints (size, domains, research needed); a quick fix that is really a task is held (`r` run anyway, `t` make it a task) |
142
+ | Effort routing | Simple steps run one thinking level lower; trivial ones on a **cheaper model** you pick. A routed run that falls short re-runs on your normal settings |
143
+
144
+ **It never gets in the way:** any failure, timeout or missing key means
145
+ bot-lobby decides as it would without it; three failures in a row pause it
146
+ for ten minutes. Calls and savings show on the Metrics tab.
147
+
148
+ **What is sent:** the planning conversation, task text, and file excerpts of
149
+ at most 400 characters (never whole files). Gitignored files, `.env*`, keys,
150
+ certificates and anything in `classifier.exclude` are never sent. If
151
+ OpenCode's free model ends, set *Model* to `jev-1.13` (paid); bot-lobby won't
152
+ switch on its own.
630
153
 
631
154
  ## Configuration
632
155
 
633
- Per-agent settings are edited interactively with `/bot-lobby settings` (or the
634
- top-level `/bot-lobby-settings`) and persist globally to
635
- `~/.pi/bot-lobby/config.json`:
156
+ Settings live in `~/.pi/bot-lobby/config.json` (`BOT_LOBBY_CONFIG_DIR`
157
+ overrides). Edit them with `/bot-lobby settings`; `/bot-lobby config` shows
158
+ the result.
636
159
 
637
160
  ```json
638
161
  {
639
- "master": { "model": "inherit", "thinking": "high", "instructions": "" },
162
+ "master": { "model": "inherit", "thinking": "high" },
640
163
  "agents": {
641
- "designer": { "model": "anthropic/claude-sonnet-5", "thinking": "medium", "instructions": "", "timeoutMs": 900000 },
642
- "backend": { "model": "anthropic/claude-sonnet-5", "thinking": "medium", "instructions": "", "timeoutMs": 900000 },
643
- "qa": { "model": "anthropic/claude-sonnet-5", "thinking": "medium", "instructions": "", "timeoutMs": 900000 }
164
+ "backend": { "model": "anthropic/claude-sonnet-5", "thinking": "medium", "timeoutMs": 900000, "instructions": "" }
644
165
  },
645
166
  "scout": { "model": "anthropic/claude-haiku-4-5-20251001", "timeoutMs": 480000 },
646
- "researcher": { "model": "anthropic/claude-sonnet-5", "thinking": "low", "instructions": "", "timeoutMs": 600000 },
647
- "quickFix": { "model": "anthropic/claude-sonnet-5", "thinking": "low", "instructions": "", "timeoutMs": 600000 },
648
- "planner": { "model": "anthropic/claude-sonnet-5", "thinking": "high", "instructions": "", "timeoutMs": 300000 },
649
- "lobby": { "autoOpen": true, "planningPanel": ["backend", "designer", "qa", "researcher"], "autoAsk": true, "issues": false, "mouse": true },
650
- "workflow": {
651
- "maxReviewIterations": 2,
652
- "maxParallelScouts": 3,
653
- "maxParallelWorkers": 3,
654
- "requireApprovalForFeatures": true,
655
- "requireApprovalForDependencies": true,
656
- "requireApprovalForArchitectureChanges": true,
657
- "agentTimeoutMs": 900000,
658
- "maxAgentRetries": 1,
659
- "stallTimeoutMs": 300000,
660
- "toolStallTimeoutMs": 600000,
661
- "wrapUpAt": 0.75
662
- },
663
- "knowledge": {
664
- "compactionThreshold": 20000,
665
- "backupCount": 1,
666
- "scratchpadMaxParagraphs": 4,
667
- "scratchpadMaxChars": 2000
668
- }
669
- }
670
- ```
671
-
672
- Every agent runs on the model and thinking level its settings name — nothing
673
- inherits the live session's thinking level. Designer and Backend workers use
674
- their domain's entry, QA's workers and the QA gate use QA's, and scouts and the
675
- researcher have their own entries. Scouts always run at `low` thinking (their
676
- entry offers a model and a time limit only); every other agent's thinking is
677
- yours to set, defaulting to `medium` (`low` for the researcher). A subagent
678
- whose model is not set yet runs on the session's model, and opening
679
- `/bot-lobby settings` pins such entries to that model so the choice is always
680
- visible; only the master keeps `inherit`, since it is the session itself.
681
- Each subagent entry has a `timeoutMs` (default 15 min; scouts 8, researcher 10),
682
- falling back to `workflow.agentTimeoutMs`.
683
-
684
- The lobby's two agents have entries of their own: `quickFix` (the direct-change
685
- agent, `low` thinking and 10 minutes by default) and `planner` (the oracle
686
- chairing the planning panel, `high` thinking; its time limit bounds one round
687
- for every seat, 5 minutes by default). Both appear in `/bot-lobby settings`,
688
- take custom instructions, and run on the session's model until you pin one.
689
- Planning seats reuse their domain's entry — DEV the Backend's, DESIGN the
690
- Designer's, QA the QA's, RESEARCH the Researcher's model, thinking and
691
- instructions — so a seat plans on the model that will later build its part.
692
- The `lobby` entry shapes the lobby itself; `/bot-lobby settings` → **Lobby**
693
- flips its switches, and key rebinding lives in the file:
694
-
695
- ```json
696
- "lobby": {
697
- "autoOpen": true,
698
- "planningPanel": ["backend", "designer", "qa", "researcher"],
699
- "autoAsk": true,
700
- "issues": false,
701
- "mouse": true,
702
- "panels": { "animations": false, "conversation": true, "activity": true, "thinking": true },
703
- "keys": { "toggleThinking": "alt+t" }
167
+ "planner": { "thinking": "high", "timeoutMs": 300000 },
168
+ "lobby": { "planningPanel": ["backend", "designer", "qa", "researcher"], "maxPlanningRounds": 5 },
169
+ "workflow": { "maxReviewIterations": 2, "maxParallelWorkers": 3, "stallTimeoutMs": 300000, "wrapUpAt": 0.75 },
170
+ "classifier": { "enabled": false, "provider": "auto", "effort": { "cheapModel": "inherit" } }
704
171
  }
705
172
  ```
706
173
 
707
- `planningPanel` names the seats a new planning session starts with (every
708
- seat by default; `[]` lets the oracle plan alone); `autoOpen` opens the lobby
709
- by itself when this session starts or resumes a task; `autoAsk` puts the
710
- panel's questions to you as soon as a round ends while the Plan tab is
711
- showing (otherwise `enter` on the empty prompt does); `issues` shows the
712
- GitHub Issues tab (off for now); `mouse` turns clicks and the wheel on;
713
- `panels` is which Lobby panes show, with `animations: true` bringing the
714
- animated oracle and agents into the lobby (off by default; the older
715
- `scene` key is no longer read, and the pane keys update it); `keys` rebinds
716
- shortcuts by action name.
717
-
718
- `thinking` must be one of `off`, `minimal`, `low`, `medium`, `high`, `xhigh`,
719
- `max`; a legacy `inherit` or unknown value falls back to `medium`. The thinking
720
- picker lists only the levels the selected model supports. Switching to a model
721
- that cannot run the saved level warns ("\"xhigh\" thinking isn't supported by
722
- provider/model — using \"high\"") and saves the nearest supported level; a run
723
- whose level its model cannot use is clamped the same way with a one-time
724
- warning, and `/bot-lobby config` lists any mismatch. `instructions` is appended to that agent's compiled
725
- system prompt as a `Custom Instructions` layer (empty layers are dropped). The
726
- master's model and thinking are applied to the live session when a task starts
727
- and when you change them in the settings TUI. A malformed config falls back to
728
- the defaults; `BOT_LOBBY_CONFIG_DIR` overrides the config directory.
729
-
730
- The model picker is searchable: type to fuzzy-filter by `provider/id` or model
731
- name, `inherit` and `custom…` stay reachable, and ↑↓/enter/esc behave as before.
732
-
733
- ### Prompts and custom instructions
734
-
735
- Every agent's system prompt is composed, never duplicated, from the baked-in
736
- Markdown in `prompts/`: `global.md`, the domain file (`designer.md`,
737
- `backend.md`, `qa.md`), the role file (`scout.md`, `worker.md`, `reviewer.md`,
738
- `researcher.md`) with that role's output contract, and then the task context,
739
- selected standards/knowledge/decisions, and workflow context. `src/prompts/compiler.ts`
740
- joins the layers and drops empty ones, so an agent never sees an empty heading.
741
- `prompts/master.md` is the Master's operating prompt and is injected only into the
742
- live session that owns the task.
743
-
744
- Your own prompt is injected as a `Custom Instructions` layer on top of those
745
- built-ins. Set it per agent — `master`, `designer`, `backend`, `qa` — either in
746
- config (`instructions`) or via `/bot-lobby settings` → Instructions. It applies to
747
- every run of that agent: the Master's instructions to the orchestrating session,
748
- and a domain's instructions to its Scouts, Workers and (for QA) the Reviewer. The
749
- layer is additive — the built-in prompts still define role boundaries, permissions
750
- and the output contract — and an empty layer is dropped.
751
-
752
- ## On-disk layout
753
-
754
- ```
755
- .pi/bot-lobby/
756
- ├── Master/knowledge/ knowledge.md, standards.md, decisions.md, completed-tasks.md
757
- ├── Designer/knowledge/ knowledge.md, design-language.md, decisions.md, completed-tasks.md
758
- ├── Backend/knowledge/ knowledge.md, engineering-standards.md, decisions.md, completed-tasks.md
759
- ├── QA/knowledge/ knowledge.md, testing-standards.md, decisions.md, completed-tasks.md
760
- ├── archive/<Agent>/ previous knowledge versions (outside all retrieval paths)
761
- ├── archive/tasks/TASK-…/ archived tasks, whole, out of every list until restored
762
- ├── sessions/<id>.json heartbeats of the pi sessions running in the project
763
- ├── sessions/<id>.inbox.jsonl messages for a running session's oracle, and their delivery
764
- ├── backlog/PLAN-<slug>.json pending tasks saved from the planner (optionally linked to an issue)
765
- ├── metrics.jsonl one line per finished run of any agent, for the Metrics tab
766
- └── tasks/TASK-<stamp>/
767
- ├── state.json the task record (kept after completion)
768
- ├── comments.jsonl your lobby comments on the plan and their delivery (append-only)
769
- ├── inbox.jsonl messages for the oracle from other sessions and their delivery (append-only)
770
- ├── auto.json auto mode, when switched on for the task
771
- ├── proposal.md scratchpads: deleted on completion
772
- ├── plan.md
773
- ├── designer.md backend.md qa.md
774
- ├── scout-<domain>.json structured scout artifacts
775
- ├── research-<domain>.json structured research reports (kept for audit)
776
- └── research.md appended research log (kept for audit)
777
- ```
778
-
779
- The global config lives outside this per-project tree, at
780
- `~/.pi/bot-lobby/config.json`.
174
+ - Entries: `master`, `agents.designer|backend|qa`, `scout`, `researcher`,
175
+ `quickFix`, `planner`, `lobby`, `workflow`, `knowledge`, `classifier`.
176
+ - An agent without a model runs on the session's model. Thinking is one of
177
+ `off, minimal, low, medium, high, xhigh, max`, limited to what the model
178
+ supports. Scouts always think at `low`.
179
+ - `instructions` adds your own text to an agent's built-in prompt.
180
+ - Classifier thresholds and limits (`classifier.thresholds`,
181
+ `classifier.fileHints`) are edited in the file.
781
182
 
782
- After the dev-house → dev-lobby → bot-lobby renames, reads merge every tree:
783
- `listTasks` and `taskHealth` enumerate `.pi/bot-lobby`, `.pi/dev-lobby` and the
784
- pre-rename `.pi/dev-house` tree with the newest root winning per task id,
785
- `loadTask` and knowledge reads fall back per file, and the newest existing of
786
- `~/.pi/bot-lobby/config.json`, `~/.pi/dev-lobby/config.json` and
787
- `~/.pi/dev-house/config.json` is still read while no bot-lobby config exists.
788
- Writes always target the bot-lobby paths, and `ensureProjectStructure` seeds the
789
- new knowledge files from the newest pre-rename tree so legacy knowledge is
790
- migrated rather than shadowed by defaults.
183
+ ## What the engine enforces
791
184
 
792
- Scratchpads are capped (`scratchpadMaxParagraphs`, `scratchpadMaxChars`) by the
793
- engine, not by prompt discipline.
185
+ | Rule | How |
186
+ | --- | --- |
187
+ | Steps happen in order | A state machine validates every action |
188
+ | Nothing is built before approval | `implement` refuses earlier states |
189
+ | Scouts and the QA gate can't edit code | Scouts get read-only tools; the QA gate adds only `bash` for tests |
190
+ | New dependencies and architecture changes need approval | Parsed from worker reports; the domain is blocked until resolved |
191
+ | Only the Master writes knowledge | Agents can only propose it |
192
+ | "Done" is earned | Needs a plan, a passing QA gate that ran checks, and no open blockers |
193
+ | Parallel workers don't clobber files | Edits need a file-desk claim |
194
+ | A crash doesn't corrupt a task | State is on disk; tasks resume from their state |
794
195
 
795
- ## Architecture
196
+ ## Files on disk
796
197
 
797
198
  ```
798
- src/
799
- ├── index.ts Extension entry: lifecycle, commands, orchestrate tool
800
- ├── master/
801
- │ ├── master.ts Scout/Worker/Reviewer delegation and artifact persistence
802
- │ ├── research.ts Researcher delegation and research artifact persistence
803
- │ ├── synthesis.ts Bounded summaries, shared-file and gap detection
804
- │ └── decisions.ts Decision log, review-loop rule, completion gates
805
- ├── agents/ Domain specs (designer, backend, qa) + registry
806
- ├── roles/ Scout/Worker/Reviewer/Researcher specs, contracts, parsers
807
- ├── workflow/
808
- │ ├── workflow.ts The engine: every action, every guard
809
- │ ├── transitions.ts Legal state machine
810
- │ └── approvals.ts Dependency/architecture/pushback approval bookkeeping
811
- ├── execution/
812
- │ ├── agent-runner.ts Single/parallel/sequential runs, live run state, cancellation, retries
813
- │ ├── pi-runner.ts Isolated `pi --mode rpc` subprocess, watchdog, stream parsing
814
- │ └── git.ts Diff evidence for reviewers
815
- ├── desk/ File desk for parallel workers: checkout table, socket, worker extension
816
- ├── knowledge/ Paths, store (single write path), selector, compactor
817
- ├── prompts/ Layer loader + compiler
818
- ├── lobby/
819
- │ ├── runtime.ts Mounts the full-screen lobby on pi's TUI, dialogs hand-off, background sessions, Master metrics
820
- │ ├── view.ts The tabbed view: tab bar, per-tab prompt, typing/browsing modes, session switcher, search, help, mouse
821
- │ ├── sessions.ts Background sessions: headless pi children driven over RPC, their feeds and questions
822
- │ ├── session-files.ts Other sessions' conversations, read from their saved session files
823
- │ ├── keys.ts The shortcut table and its config overrides
824
- │ ├── ask.ts The panel's questions through the ask-user-question questionnaire (or pi's dialogs)
825
- │ ├── markdown.ts Markdown through pi's renderer, cached per theme and width
826
- │ ├── tabs/ Pure renderers: home, tasks, plan, quickfix, issues, metrics
827
- │ ├── feed.ts Activity log, thinking pane and conversation store
828
- │ ├── quickfix.ts Direct-change jobs, one at a time
829
- │ ├── planner.ts The planning panel: seats and the oracle per round, reply parsing, saving a plan
830
- │ ├── issues.ts GitHub issues through the gh CLI
831
- │ └── layout.ts Boxes, exact-width columns, wrapping, highlights, bars, meters and sparklines
832
- ├── state/ Project root, config, task persistence, state mutation, comments, inbox, auto mode, backlog, metrics
833
- ├── schemas/ Task, agent, findings, configuration types
834
- └── pi/ Commands, lifecycle, orchestrate tool, status widget, the owner's clock (deliveries, auto mode)
835
- prompts/ global, master, designer, backend, qa, scout, worker, reviewer, researcher, quickfix, planner, panel
199
+ .pi/bot-lobby/
200
+ ├── <Agent>/knowledge/ knowledge, standards and decisions per agent
201
+ ├── tasks/TASK-…/ state.json, scratchpads, scout and research reports
202
+ ├── backlog/PLAN-….json plans saved from the Plan tab
203
+ ├── archive/ archived tasks and old knowledge
204
+ ├── sessions/ heartbeats of running Pi sessions
205
+ ├── cache/files.json file excerpts for the classifier
206
+ └── metrics.jsonl one line per agent run and classifier call
836
207
  ```
837
208
 
838
- Prompts are composed, never duplicated: `global + domain + role + task context +
839
- standards + knowledge + decisions + workflow context + output contract`, with
840
- empty layers dropped and only task-relevant knowledge slices included.
841
-
842
209
  ## Development
843
210
 
844
211
  ```bash
845
212
  npm install
846
213
  npm run typecheck
847
- npm test # node:test, no extra framework
848
- ```
849
-
850
- Live end-to-end checks (spend tokens, need a configured model):
851
-
852
- ```bash
853
- BOT_LOBBY_E2E=1 npx tsx --test test/e2e.test.ts # or: node --test test/e2e.test.ts
214
+ npm test
854
215
  ```
855
216
 
856
- They cover: a real isolated subagent run, a real workflow-level scout that
857
- advances the task state, and the Master prompt injection in a real Pi session.
858
-
859
- ## Publishing (maintainers)
860
-
861
- The [Pi package gallery](https://pi.dev/packages) discovers npm packages that
862
- carry the `pi-package` keyword, which `package.json` already sets: the package
863
- page is [pi.dev/packages/@a-t-h-i/bot-lobby](https://pi.dev/packages/@a-t-h-i/bot-lobby)
864
- and the registry page is
865
- [npmjs.com/package/@a-t-h-i/bot-lobby](https://www.npmjs.com/package/@a-t-h-i/bot-lobby).
866
-
867
- A release is a version bump followed by:
217
+ Live checks (spend tokens or need a key and network):
868
218
 
869
219
  ```bash
870
- npm publish --access public
220
+ BOT_LOBBY_E2E=1 node --test test/e2e.test.ts
221
+ BOT_LOBBY_JEV_E2E=1 OPENCODE_API_KEY=… node --test test/jev-e2e.test.ts
871
222
  ```
872
223
 
873
- `publishConfig.access` pins public access, and `files` (`src`, `prompts`) keeps the
874
- tarball to the extension and its prompt layers — check it with
875
- `npm pack --dry-run`. Pi supplies the Pi packages at runtime, so they stay in
876
- `peerDependencies` with a `"*"` range.
877
-
878
- ## Scope of v0.1
224
+ Source layout: `src/workflow` (engine), `src/master` (delegation),
225
+ `src/execution` (subagent processes), `src/lobby` (the UI),
226
+ `src/classifier` (Jev), `src/state` (persistence), `prompts/` (agent prompts).
879
227
 
880
- Included: the full workflow above, persistent knowledge with governance and
881
- compaction, bounded review loops, dependency/architecture approval, retries,
882
- cancellation, corrupted-state detection, and the commands/status UI.
228
+ ## Publishing
883
229
 
884
- Deliberately deferred (matching the build plan): worktree-based isolation for
885
- parallel Workers (they share one working tree through the file desk) and
886
- cross-platform runtime abstractions. The internal module boundaries keep those
887
- extractable. The lobby (see above) has since added the full-screen dashboard
888
- and per-model performance analytics.
230
+ Bump the version, then `npm publish --access public`. Check the tarball
231
+ with `npm pack --dry-run`; it ships `src` and `prompts` only.