omp-conductor 0.15.10 → 0.15.12

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (92) hide show
  1. package/README.md +273 -2544
  2. package/REFERENCE.md +2680 -0
  3. package/package.json +3 -2
  4. package/schema/config.schema.json +11 -23
  5. package/src/arm-challenge.ts +112 -0
  6. package/src/ask.ts +434 -0
  7. package/src/board.ts +81 -15
  8. package/src/brief-upgrade.ts +114 -8
  9. package/src/briefs/orchestrator.md +113 -34
  10. package/src/briefs/policy.md +33 -8
  11. package/src/briefs/worker.md +18 -9
  12. package/src/chain-check.ts +1 -1
  13. package/src/check-trailing-newlines.ts +82 -0
  14. package/src/cli.ts +225 -1406
  15. package/src/commands/arm.ts +21 -0
  16. package/src/commands/board.ts +23 -0
  17. package/src/commands/brief-upgrade.ts +186 -0
  18. package/src/commands/context.ts +150 -0
  19. package/src/commands/daemon.ts +71 -0
  20. package/src/commands/dashboard.ts +74 -0
  21. package/src/commands/decision.ts +103 -0
  22. package/src/commands/disarm.ts +21 -0
  23. package/src/commands/doctor.ts +100 -0
  24. package/src/commands/event.ts +62 -0
  25. package/src/commands/extend.ts +64 -0
  26. package/src/commands/friction.ts +56 -0
  27. package/src/commands/help.ts +9 -0
  28. package/src/commands/hold.ts +26 -0
  29. package/src/commands/intake.ts +155 -0
  30. package/src/commands/ledger.ts +69 -0
  31. package/src/commands/message.ts +96 -0
  32. package/src/commands/report.ts +206 -0
  33. package/src/commands/restart.ts +76 -0
  34. package/src/commands/resume.ts +58 -0
  35. package/src/commands/setup.ts +143 -0
  36. package/src/commands/start.ts +23 -0
  37. package/src/commands/stats.ts +131 -0
  38. package/src/commands/status.ts +48 -0
  39. package/src/commands/stop.ts +51 -0
  40. package/src/commands/tail.ts +109 -0
  41. package/src/commands/unblock.ts +39 -0
  42. package/src/commands/upgrade-install.ts +31 -0
  43. package/src/commands/upgrade-rollback.ts +32 -0
  44. package/src/commands/upgrade.ts +25 -0
  45. package/src/commands/verb.ts +83 -0
  46. package/src/commands/version.ts +30 -0
  47. package/src/commands/worker.ts +100 -0
  48. package/src/config-schema.ts +47 -1
  49. package/src/config.ts +45 -4
  50. package/src/daemon.ts +972 -101
  51. package/src/dashboard/app.js +459 -0
  52. package/src/dashboard/index.html +61 -0
  53. package/src/dashboard/server.ts +481 -0
  54. package/src/dashboard/style.css +348 -0
  55. package/src/decisions.ts +39 -14
  56. package/src/diff-flags.ts +131 -241
  57. package/src/doctor.ts +932 -0
  58. package/src/escalate.ts +2 -2
  59. package/src/failure-class.ts +66 -3
  60. package/src/fleet.ts +58 -1
  61. package/src/gitops.ts +157 -0
  62. package/src/graph-health.ts +1 -1
  63. package/src/label-projection.ts +1 -1
  64. package/src/lifecycle.ts +198 -2
  65. package/src/model-fallback.ts +177 -0
  66. package/src/notices.ts +9 -0
  67. package/src/omp.ts +93 -13
  68. package/src/orchestrator-tick.ts +414 -20
  69. package/src/orchestrator.ts +4 -4
  70. package/src/privileged.ts +10 -0
  71. package/src/release-policy.ts +342 -30
  72. package/src/reports.ts +19 -5
  73. package/src/session-host.ts +11 -5
  74. package/src/setup-host.ts +663 -25
  75. package/src/setup-install.ts +292 -28
  76. package/src/setup-wizard.ts +255 -74
  77. package/src/setup.ts +156 -35
  78. package/src/stats.ts +331 -0
  79. package/src/store.ts +219 -22
  80. package/src/tracker/github.ts +74 -4
  81. package/src/types.ts +215 -31
  82. package/src/unblock.ts +55 -11
  83. package/src/upgrade-journal.ts +220 -0
  84. package/src/upgrade-verify.ts +506 -0
  85. package/src/upgrade.ts +385 -58
  86. package/src/verbs/actions.ts +73 -1
  87. package/src/verbs/protocol.ts +29 -4
  88. package/src/verbs/server.ts +183 -20
  89. package/src/worker.ts +3 -3
  90. package/systemd/omp-conductor-recover.sh +433 -0
  91. package/systemd/omp-conductor.service.example +14 -3
  92. package/systemd/recover-unit-test.sh +428 -0
package/REFERENCE.md ADDED
@@ -0,0 +1,2680 @@
1
+ # omp-conductor — Reference
2
+
3
+ The complete reference for `omp-conductor`, every surface and every internal. For
4
+ the 10-minute operator guide — what it is, install, quick start, operating the
5
+ fleet, and multi-project operation — read [`README.md`](README.md).
6
+
7
+ The sections below are the reference half of the documentation. Each guide
8
+ section links back here for the full detail it trims.
9
+
10
+ ## Install prerequisites
11
+
12
+ The guide's [Install](README.md#install) section links here. What the host needs
13
+ before anything runs:
14
+
15
+ `@oh-my-pi/pi-coding-agent` (`>=17.1.4`) is a **peer dependency** and must already
16
+ be present. If you run omp, it is.
17
+
18
+ Also required on the host:
19
+
20
+ - `bun`: the CLI and the daemon run on it (`Bun.serve` backs `/healthz`).
21
+ - **A model credential the *daemon's own account* can reach.** Sessions are child
22
+ processes of the daemon and inherit its environment and `$HOME` unmodified, so a
23
+ session authenticates with exactly what the daemon authenticates with — nothing
24
+ is injected and nothing is scrubbed. Any shape the harness itself understands
25
+ works, including the ordinary one:
26
+ - an **OAuth login** already recorded for that account (`omp` login state under
27
+ its `~/.omp/agent`). This is the common case and needs no configuration.
28
+ - a model **API key** in the daemon's environment (`ANTHROPIC_API_KEY`,
29
+ `OPENAI_API_KEY`, …). Note systemd starts the service with a clean
30
+ environment, so it has to be an `Environment=` line on the unit, not something
31
+ exported in your shell.
32
+ - the harness's **auth broker**, configured in that account's
33
+ `~/.omp/agent/config.yml`.
34
+
35
+ The account matters more than the shape: a login recorded under a *different*
36
+ account is invisible to the service. A unit running `User=fleet` cannot see
37
+ `root`'s login, and workers then die at turn 0 with `No model selected`. Such a
38
+ run is classified `env-start-failure` and charges **neither** the failure budget
39
+ nor a continuation — an environment fault is not a failed implementation, and a
40
+ run that recorded no turn, no commit and no error did not attempt anything — but
41
+ nothing dispatches successfully until the credential is reachable.
42
+ - `gh`, already authenticated: every tracker operation shells out to it, so the
43
+ daemon never handles a GitHub token itself.
44
+ - `git`: mirrors and worktrees.
45
+ - **[omp-telegram](https://www.npmjs.com/package/omp-telegram)**, for the
46
+ escalation channel. It is a separate package and is not vendored here.
47
+
48
+ Two different things depend on it, and they need different amounts of it:
49
+
50
+ - **Tier-2 paging** needs only its bot token. This package reads
51
+ `TELEGRAM_BOT_TOKEN` out of `$OMP_TELEGRAM_STATE_DIR/.env` (default
52
+ `~/.omp/agent/telegram/.env`) and posts to the chat id you configure. No
53
+ pairing required, and no token ever passes through this package's own config.
54
+ - **The interactive channel** — replying to an escalation, approving a brief
55
+ amendment from your phone — needs omp-telegram actually paired, which is what
56
+ writes `access.json`. The fleet heartbeat also reads that file and refuses to
57
+ tick unless exactly one owner is paired, on the grounds that unattended
58
+ dispatch is only defensible while a tier-2 page can reach a person.
59
+ - **Approving a Learning-loop amendment from a heartbeat tick** needs one more
60
+ setting than pairing: a notify destination. `/telegram notify` writes
61
+ `notifyMode` and `notifyChat` into the same `access.json`. omp-telegram
62
+ mounts its `telegram_ask` tool only for a turn that resolves a notify target,
63
+ and a locally injected tick resolves one only through that setting — so
64
+ without it the orchestrator can page you but cannot put a yes/no question in
65
+ front of you, which is the one thing the Learning loop's approval step
66
+ requires. `omp-conductor status` reports this on the `telegram` row, and a
67
+ tick that cannot ask says so in its own prompt and falls back to
68
+ `telegram_send`.
69
+
70
+ Set `notifyChat` even on a forum fleet. `/telegram topics` routes to a topic
71
+ this session claims at runtime, and that claim is not visible in
72
+ `access.json` — so a file carrying only `topicsChat` is reported as
73
+ unconfigured rather than guessed at, on the grounds that a health row which
74
+ reads green over a broken contract is worse than one that overstates a fault.
75
+
76
+ - **A fleet that answers instead of narrating** needs one key, set once:
77
+ `/telegram set profile daemon` (omp-telegram 0.11.0 or newer). Without it the
78
+ bridge behaves as it does on a laptop: it finalizes a real Telegram message
79
+ per assistant turn for as long as a conversation is active — so one answer
80
+ arrives as several messages, and a message that lands mid-tick keeps relaying
81
+ that tick's internal turns — and it posts every local run's closing text to
82
+ `notifyChat`, which on a host whose runs are heartbeat ticks means each tick's
83
+ working prose. The profile switches all of it off at the transport: text
84
+ reaches Telegram only through `telegram_send` / `telegram_ask`, the idle post
85
+ is suppressed, and `telegram_ask` stays mounted and aimed at the paired owner
86
+ on every turn — including a locally injected tick, so it also removes the need
87
+ for `notifyMode` above. Approval and blocked-input pings still fire; those
88
+ mean a human is needed, which is the point of the channel.
89
+
90
+ `omp-conductor status` reports an interactive profile on the `telegram` row,
91
+ and every tick composed on one carries a prompt line saying so. Note the two
92
+ settings pull against each other before the profile exists: setting
93
+ `notifyMode` to make `telegram_ask` mountable is exactly what arms the idle
94
+ post, so the correctly askable fleet was also the loud one.
95
+
96
+ With neither, tier 2 degrades to a comment on the issue. Nothing is broken in
97
+ that configuration: it is supported, just slower to reach you.
98
+
99
+
100
+ ## Updating
101
+
102
+ Run one command from a shell outside the target `herdr-fleet.service`:
103
+
104
+ ```bash
105
+ omp-conductor upgrade
106
+ ```
107
+
108
+ The fleet can also install its own fixes: with the `install` release shape
109
+ granted to the orchestrator (`omp-conductor setup authority`, gate **install**),
110
+ an orchestrator requests `conductor_install --arg version=X.Y.Z`. The daemon
111
+ refuses a version npm does not expose, then starts a transient systemd unit
112
+ (`omp-conductor-upgrade-install-<version>`) that runs the same transaction
113
+ outside the pane and the daemon, journaling every surface it touches to the
114
+ state-dir upgrade journal. The first tick after the restart reads that journal
115
+ and verifies independently — installed version, `/healthz`, ticks, pane,
116
+ `doctor` — then restores dispatch and reports through the durable outbox; any
117
+ gap triggers the detached rollback and a tier-2 page. The unit never declares
118
+ its own success, so a fleet is never left half-upgraded and quiet.
119
+
120
+ It resolves the latest published npm release, pauses new claims, drains active
121
+ workers, and pins that exact release across the Bun-global CLI, omp plugin, and
122
+ Herdr plugin. It also recomposes the conductor-owned brief floor, restarts Herdr
123
+ and the daemon, waits for pane recovery, verifies the installed identities and
124
+ layered fleet status twice, then restores the original dispatch state.
125
+
126
+ Use `omp-conductor upgrade --to X.Y.Z` for an explicit published version. The
127
+ command exits without changing anything when all three surfaces already use that
128
+ release, the Herdr plugin is pinned to its exact `gitHead`, and the brief is
129
+ current.
130
+
131
+ The upgrade is host-wide, because everything it replaces is: one daemon serves
132
+ every configured project, so its restart lands on all of them at once. A bare
133
+ `omp-conductor upgrade` therefore drains **every** project's workers and
134
+ refreshes **every** project's brief, and pauses them with the fleet-wide
135
+ sentinel — which leaves any per-project `hold` you set standing when it
136
+ restores dispatch.
137
+
138
+ `--project` narrows that only when it is truthful to do so. When a live daemon
139
+ recorded a single project, the command drains and restarts that project, and an
140
+ explicit `--project` naming a different one is rejected before pause or
141
+ installation. When the daemon serves every configured project and there is more
142
+ than one, `--project` is rejected too: draining one queue and then restarting
143
+ the shared daemon would kill another queue's workers without ever counting
144
+ them. Re-run without the flag.
145
+
146
+ Ticks remain in their existing armed or disarmed state, so an ordinary update
147
+ does not halt the exact pane or require another Telegram arm challenge. Progress
148
+ names the Bun-global CLI, omp plugin, Herdr plugin, brief, reloads, and both
149
+ verification passes separately.
150
+
151
+ An installation, brief, reload, or verification failure pauses dispatch and
152
+ attempts to restore the exact CLI/plugin identities that were present before the
153
+ command. If rollback also fails, the error names every failed restoration and
154
+ keeps dispatch paused; it never brings a known mixed fleet back into service.
155
+ The command never publishes npm, edits an install root, or delegates lifecycle
156
+ steps to an AI session. It refuses to run inside a Herdr-managed session because
157
+ an updater that restarts itself cannot verify the result.
158
+
159
+ After every upgrade — and after the initial install — run the mechanical host
160
+ verification:
161
+
162
+ ```bash
163
+ omp-conductor doctor
164
+ ```
165
+
166
+ It is read-only and exits 0 only when nothing failed. Each finding is one
167
+ deployment fault that has already cost a debugging session: expired `gh` auth
168
+ or scope gaps, case-mismatched routing/state labels (silent by design), the
169
+ installed systemd units drifting from the staged render, runtime-dir ownership,
170
+ config.json validation and backup freshness, sqlite `PRAGMA integrity_check`,
171
+ spend telemetry reading $0.00 on every completed run, invalid reporting
172
+ timezones, and Telegram delivery health. `--probe-telegram` is the one opt-in
173
+ side effect: one self-identified test message through the report transport.
174
+
175
+ The unit-drift check compares the installed units against the staged render,
176
+ excluding the herdr pane-shell `SHELL=` pin: that value is host state resolved
177
+ at staging time, and `doctor`'s own process resolves it differently (or not at
178
+ all) in a non-login/tick context, so it is reported as not-comparable rather
179
+ than as drift (#511). It still flags genuine drift in `User=`, `ExecStart=`,
180
+ other `Environment=` lines, `Restart=` and `SuccessExitStatus=`, and the unit
181
+ itself keeps pinning `SHELL=` so panes do not fall back to dash (#463).
182
+
183
+ ## Onboarding
184
+
185
+ `omp-conductor setup` is the whole of it. One command, in a plain terminal, doing
186
+ the two jobs onboarding has always had:
187
+
188
+ | Half | What it does |
189
+ | --- | --- |
190
+ | **The interview** | Asks the judgment no amount of repo reading produces, then writes it into `POLICY.md` as prose you own and can edit. |
191
+ | **The probes** | Reads your repos and *proposes* the rest — real CI gates, the project context, the release procedure — each a default you edit or a draft you confirm. |
192
+
193
+ The split matters because the two halves fail differently. A wrong config value is
194
+ a run that errors on the next tick; a wrong release boundary is a fleet that
195
+ publishes something at 03:00. The first is worth a validated prompt. The second is
196
+ worth being asked properly, which is why it is asked and never guessed.
197
+
198
+ ### What only you can answer
199
+
200
+ Always asked: **where the roadmap lives and what the current priority is.** A
201
+ tracker shows what is *open*, never what *matters*, and an orchestrator that cannot
202
+ rank work grooms by recency — which is how a stale issue outranks the thing you are
203
+ shipping this month.
204
+
205
+ Asked only when you grant the orchestrator a release shape, because a
206
+ humans-release fleet has no boundary to draw:
207
+
208
+ - **Where the orchestrator's leg ENDS**, in one sentence. If it cannot be said in
209
+ one sentence it is not a boundary, and a vague release mandate is what eventually
210
+ publishes something at 03:00.
211
+ - **What** may be released and from which branch; **when** — batched how, after
212
+ which *named* checks; **what proof** must be held first, results actually read
213
+ rather than an impression; **what must still be asked** every time; and **what
214
+ stays permanently forbidden**.
215
+ - **What makes a release worth cutting** — a sprint, an epic's children all closed,
216
+ N merged issues. Without it the orchestrator either releases per merge, a stream
217
+ of meaningless versions burning shared runners, or never releases at all.
218
+ - **Who owns the rollback.** Name a person and setup says so plainly: that person
219
+ already owns the release, so the honest configuration ends the agent's leg
220
+ *before* the irreversible step. It offers to move the boundary there; declining
221
+ is a choice, not a mistake.
222
+
223
+ Grant every release shape and setup pushes back once — credentials sitting in the
224
+ environment of a session that runs unattended for weeks, and a 03:00 rollback being
225
+ a judgement call under time pressure with partial information — then records what
226
+ you decide. It is your fleet.
227
+
228
+ ### What setup reads for you
229
+
230
+ Setup discovers factual defaults before it asks for them. `git remote get-url
231
+ origin` supplies the tracker repo and single-repo routing key; bounded `gh` calls
232
+ supply the default branch, existing queue/state labels, branch-protection checks,
233
+ environments, an unambiguous npm package name, and GitHub Projects/open milestones.
234
+ Each discovered value is shown with its evidence and remains editable at the same
235
+ prompt. Discovery only seeds a fresh interview: a re-run starts from the saved
236
+ project, so an operator-edited value is never guessed again.
237
+
238
+ Judgment and prose still belong to the existing confined model probes, and **every
239
+ answer is a proposal**:
240
+
241
+ - **Gates.** Reads each routing repo's CI workflows, `package.json` scripts and
242
+ `Makefile`/`justfile`, then pre-fills the [gates](#configuration) prompt with the
243
+ exact commands and the `cwd` each runs from, so your gates match what CI runs. It
244
+ reports the evidence it used, and an honest "this repo has no cheap local check"
245
+ is a real answer rather than an invented `npm test`.
246
+ - **Project context** and **the release procedure.** Drafted across *every* routing
247
+ repo — which repo owns which concern, which ship together, where the release
248
+ machinery actually lives — then shown to you in full and kept **only if you
249
+ confirm**. `POLICY.md` is re-read on every tick, so a paragraph you never read
250
+ would become an instruction the orchestrator follows all week.
251
+
252
+ A model probe has **no shell, no editor and no verbs**: it reads files and answers,
253
+ and a tool it was not given is refused rather than allowed. It is not a sandbox —
254
+ it runs as your own user and reads what you can read — which is why it is pointed
255
+ only at repos you configured yourself.
256
+
257
+ **Setup never fails because discovery or a probe did.** No `gh`, no auth or
258
+ network, a private repo, no omp peer, a clone failure, a timeout, or a malformed
259
+ reply each produces a warning and leaves the typed default in place. Every question
260
+ is still asked.
261
+
262
+ Skip the reading half entirely with `--no-ai`:
263
+
264
+ ```bash
265
+ omp-conductor setup --no-ai
266
+ ```
267
+
268
+ To fill in or revise just the brief later — the two `POLICY.md` sections above — run
269
+ the `brief` area, which re-asks the judgment questions and re-runs the probes:
270
+
271
+ ```bash
272
+ omp-conductor setup brief
273
+ ```
274
+
275
+ ### Changing one setting
276
+
277
+ `config.json` is wizard-written, so changing a value means running the wizard —
278
+ and a wizard that re-asks twenty questions to add one key is a wizard people edit
279
+ the file behind instead. So a re-run against a project that is already configured
280
+ opens with one question:
281
+
282
+ ```text
283
+ "platform" is already configured — what would you like to do?
284
+ > Change one area
285
+ asks one area's questions; every other answer is carried through from the saved config
286
+ Walk every question again
287
+ the full interview, every prompt pre-filled with what is configured now
288
+ Add another project
289
+ full interview for a new project; existing projects stay as they are
290
+ ```
291
+
292
+ Amending is the default. Pick it and the eight areas are listed with what each one
293
+ says right now, so the row you want is the row you can see:
294
+
295
+ ```text
296
+ Which area? Each row shows what it says now
297
+ tracker & repos — acme/platform, queue "ready-for-agent", "repo:" → platform, api, web, worker
298
+ gates — platform: bun run check; api: ruff check . @ backend; web: pnpm lint…
299
+ caps & worker model — 2 workers, 120 turns, 90m, $25/day, 2 attempts (all defaults) — harness default model
300
+ code graph — not configured — workers grep
301
+ authority — merge=orchestrator, release=orchestrator
302
+ escalation & triage — tier 2 pages Telegram 123456789, comments too, triage external
303
+ reporting scope — material — escalations, plus green PRs, second failures, and anything that stops the fleet
304
+ orchestrator brief — none at ~/.omp/conductor/worktrees/ORCHESTRATOR.md
305
+ ```
306
+
307
+ Only that area's questions are asked. Every other answer is read back out of
308
+ `config.json` and written again unchanged — the same answers, the same builder,
309
+ the same single confirm, so there is still exactly one thing in this package that
310
+ writes a config, and it still writes nothing before you agree. The consent screen
311
+ leads with the delta and then shows the whole project as it would be written:
312
+
313
+ ```text
314
+ amending code graph — project platform
315
+ was not configured — workers grep
316
+ now ~/.cache/conductor-graph/acme — 4 clone(s): platform, api, web, worker
317
+ carried over tracker & repos, gates, caps & worker model, authority, escalation & triage, reporting scope, orchestrator brief
318
+ read back from ~/.omp/conductor/config.json and rewritten unchanged
319
+ ```
320
+
321
+ A first run, or a project name this config has never seen, never sees either
322
+ question: there is nothing to amend, so it is the full interview exactly as
323
+ before. Choosing *Walk every question again* asks once to confirm the replace
324
+ (so a silent overwrite cannot happen from muscle-memory Enter) — every prompt
325
+ pre-filled with what is configured, Enter to keep it — with one wrinkle worth
326
+ knowing: the two authority confirms and the orchestrator-session confirm cannot
327
+ start on "yes", so Entering through the full interview **revokes** a delegation
328
+ rather than renewing it. Amending the `authority` area names the current grant in
329
+ the question, which is the safer way to leave one alone.
330
+
331
+ ### Adding another project
332
+
333
+ One daemon serves every configured project. To put a second (or third) fleet on
334
+ the same host without touching the first:
335
+
336
+ ```bash
337
+ omp-conductor setup --project second
338
+ # or pick "Add another project" from the re-run chooser
339
+ ```
340
+
341
+ That is a full interview for the new name only. Defaults land under
342
+ `~/.omp/conductor/projects/<name>/{worktrees,mirrors}` so two fleets never share
343
+ a cwd; the first project's existing flat `worktrees`/`mirrors` paths are never
344
+ migrated. A `workspaceRoot` that collides with another project is refused with
345
+ both names in the error. Re-using an existing name asks amend-or-replace before
346
+ anything is written.
347
+
348
+ After apply, setup provisions labels, brief, tick config (with `project` +
349
+ `agentName`), topic binding, smoke, and arm for the new project only, then prints:
350
+
351
+ - `omp-conductor restart --now` — the running daemon picks up the new project
352
+ only after reload (printed, not auto-run when workers are live)
353
+ - a copy-pasteable **herdr handoff**: `herdr --session conductor workspace create
354
+ --cwd <workspaceRoot> --label <project> --no-focus`, then `herdr --session
355
+ conductor agent start <project> --kind omp --pane <pane-id>` into that empty
356
+ pane (never into a live orchestrator), plus the `FLEET_CWDS` list for
357
+ multi-fleet recovery
358
+
359
+ `omp-conductor setup gates --project second` (and every other area) still amends
360
+ only that project.
361
+
362
+ ### Keeping a brief current
363
+
364
+ The standing prompt is three layers:
365
+
366
+ | Layer | File | Updates how? |
367
+ | --- | --- | --- |
368
+ | Package floor | `src/briefs/orchestrator.md` | Every tick recomposes it into `ORCHESTRATOR.md` from the installed package. Upgrade the package in this host's existing install root + restart is enough. |
369
+ | Shared policy | `$OMP_CONDUCTOR_HOME/SHARED_POLICY.md` (default `~/.omp/conductor/SHARED_POLICY.md`) | Optional, host-wide: one file applying to **every** project. Create it by hand beside `config.json`; it is re-read each tick like `POLICY.md`, so an edit binds the next heartbeat. The per-project `POLICY.md` overrides it where the two conflict. |
370
+ | Fleet policy | `POLICY.md` | Yours. Setup writes the scaffold once; the Learning loop edits only this file. |
371
+ | Composed view | `ORCHESTRATOR.md` | Regenerated from floor + `SHARED_POLICY.md` (when present) + `POLICY.md` on each tick (and at setup). Do not hand-amend it for durable policy. |
372
+ | Worker brief | `src/briefs/worker.md` | Read per run from the package. |
373
+
374
+ ```bash
375
+ omp-conductor brief-upgrade # report overlay / legacy state
376
+ omp-conductor brief-upgrade --migrate # dry-run: bannered ORCHESTRATOR.md → POLICY.md
377
+ omp-conductor brief-upgrade --migrate --apply
378
+ omp-conductor brief-upgrade --retrofit # #20: propose YOURS TO EDIT cut on a hand-written brief
379
+ omp-conductor brief-upgrade --retrofit --apply
380
+ ```
381
+
382
+ - **Overlay already active** (`POLICY.md` present): protocol updates need no brief-upgrade.
383
+ `brief-upgrade` reports which layers compose — floor, optional shared policy,
384
+ and `POLICY.md` — and the composed banner names all three when the shared one
385
+ exists, so a reader can tell which layer a paragraph came from.
386
+ - **Legacy bannered brief**: `--migrate` lifts the owned half into `POLICY.md`
387
+ and recomposes. Previous brief and policy versions go to
388
+ `$OMP_CONDUCTOR_HOME/backups/briefs/` (default
389
+ `~/.omp/conductor/backups/briefs/`), named with their source filename and
390
+ timestamp.
391
+ - **Hand-written brief** (no banner): `--retrofit` inserts the banner before the first Releases / Project context / Reporting / Amendments heading; then `--migrate`.
392
+ - **The legacy single-file merge is gone** (0.4.3). A bare `--apply` exits `2`
393
+ naming the two paths that remain, rather than rewriting a brief nobody asked
394
+ it to. `--migrate` is the cross-version ABI: an upgrade keeps calling it, and
395
+ an overlay fleet tolerates its absence because the floor recomposes each tick.
396
+ - **`status` names the layout**, so a fleet still on a legacy brief is visible
397
+ where an operator already looks: `brief overlay (package floor + POLICY.md)`,
398
+ or `brief legacy-bannered — run omp-conductor brief-upgrade`.
399
+ - **Existing sidecars**: a composed refresh relocates conductor-generated
400
+ `ORCHESTRATOR.md.bak-<timestamp>` and `POLICY.md.bak-<timestamp>` files into
401
+ that backup directory. Other `.bak` files stay untouched.
402
+
403
+ `--file PATH` checks a brief that is not where the wizard would have put it.
404
+
405
+ The **Learning loop** proposes diffs against `POLICY.md` for you to approve over
406
+ Telegram. It also learns from repeated operational friction. The daemon
407
+ automatically rolls up repairable admission holds; the orchestrator records
408
+ judgments code cannot make with:
409
+
410
+ ```bash
411
+ omp-conductor friction escalation-digest --detail "routine retry belonged in the digest" [--issue N]
412
+ omp-conductor friction report-noise --detail "green status repeated with no operator action"
413
+ omp-conductor friction report-surprise --detail "a material failure was missing from the report"
414
+ ```
415
+
416
+ Three observations within seven days make a bounded signal eligible for one
417
+ tick. After it is surfaced, that signal cools down for seven days. A signal is
418
+ evidence to investigate, never an automatic policy edit: the existing one-at-a-
419
+ time Telegram approval, `POLICY.md`-only edit, Hard-boundary prohibition, and
420
+ **Amendments** log still apply.
421
+
422
+ ## How one tick works
423
+
424
+ Per tick, for the daemon's project:
425
+
426
+ 1. **Verify and settle pushed PRs.** For every run in `pushed-pending`, repeat
427
+ the independent head/check verification; green → `pushed-green`, red →
428
+ `failed`, and still pending stays occupied. For every verified
429
+ `pushed-green` run, ask what became of its PR. Merged → `merged`; closed
430
+ without merging → `failed`. Unknown answers leave the row unchanged. Every row
431
+ that settles also loses its `agent:in-progress` label. This
432
+ maintenance runs even while dispatch is paused or workers are active, so
433
+ status converges on the five-minute tick cadence. It also runs above admission
434
+ so a row settled here frees its issue in the same tick. See
435
+ [what settles a green PR](#what-settles-a-green-pr).
436
+ 2. **Paused?** If the pause sentinel exists, the tick claims nothing and returns.
437
+ Settlement has already run, but no queue or admission work occurs. This makes
438
+ `omp-conductor hold` take effect without signalling the process.
439
+ 3. **List the queue.** Open issues in `tracker.repo` labelled `queueLabel`.
440
+ 4. **Filter and route.** An issue is eligible only if it carries the queue label
441
+ and none of the three state labels (`inProgress`, `blocked`, `failed`). Eligible
442
+ issues are partitioned into routable and unroutable.
443
+ 5. **Escalate the unroutable** at Tier 1, quoting the repo labels actually seen and
444
+ the configured repo names. These are never dispatched.
445
+ 6. **Check spend.** If spend since local midnight has reached `dailySpendUsd`, the
446
+ daemon **pauses itself**, pages at Tier 2, and returns.
447
+ 7. **Check capacity.** `maxConcurrentWorkers` minus *live* workers (runs in
448
+ `claimed` or `running`) gives the free slots. A pending or green PR occupies
449
+ its issue but not a slot: its worker is finished, and counting pushed PRs
450
+ would let two completed workers stop the fleet.
451
+ If no slot is free, the tick logs and returns.
452
+ 8. **Check the plan allowance.** If `caps.planUsage` names a window, the daemon
453
+ reads it (cached, see [Caps](#caps)) and holds *every* candidate under
454
+ `plan-usage-cap` when the window is at or over its threshold. Unlike the
455
+ spend cap this does **not** pause the daemon: the window resets on the
456
+ provider's clock, so dispatch resumes by itself.
457
+ 9. **Admit issues** up to the free slots, skipping any issue that already has
458
+ an active run — including a pending or green PR, so a second attempt cannot
459
+ land on a live PR. Repeated implementation failures consume
460
+ `maxAttemptsPerIssue`; cap kills, daemon orphans and answered blocks consume
461
+ the independent `maxContinuationsPerIssue`. Exhausting either escalates.
462
+ 10. **Ask the tracker whether the work already exists.** For each candidate that
463
+ survived step 9 — so at most one API call per free slot, never one per queued
464
+ issue — the daemon asks whether an **open** PR already closes it. An open PR
465
+ normally holds the issue. One narrow exception permits a routed continuation:
466
+ the latest run must be terminal, and the open PR must be that run's retained
467
+ work — either the PR URL it recorded or a PR opened on the branch it retained.
468
+ The branch half matters because a run can be cap-killed before its worker ever
469
+ opens a PR, leaving a retained branch and no recorded URL; a PR pushed to that
470
+ branch afterwards is still the continuation target. Drafts count because their
471
+ branch can hold the only copy of the work.
472
+ The tracker also finds work missing from a new, moved, restored, or cleared
473
+ store. If the check fails, the candidate is **held**, not admitted, and
474
+ retried next tick: the cost of holding is five minutes, the cost of admitting
475
+ on an unknown is a burned attempt and a duplicate PR. Only that candidate is
476
+ held, so a flaky API cannot stall the rest of the queue.
477
+ 11. **Record the pass.** Persist ready/routed/admitted counts and group every hold
478
+ under a stable reason code with at most five sample issue numbers. Tracker
479
+ failures mark the summary `DEGRADED`; capacity, sibling, open-PR and budget
480
+ holds remain normal policy state.
481
+ 12. **Dispatch** the admitted issues concurrently.
482
+
483
+ Then, per admitted issue:
484
+
485
+ 1. **Create the run row (`claimed`) — before any worktree or session exists.**
486
+ This ordering is the whole crash-safety story: the *row*, not a label, is the
487
+ guard against dispatching the same issue twice. It is local and written before
488
+ anything that can fail; if the process dies at any later point, the startup
489
+ orphan sweep marks the left-behind row `orphaned` and the orchestrator's drain
490
+ duty triages it (see below). The `agent:in-progress` label is a write-behind
491
+ projection of that row — enqueued in the same breath and applied to the tracker
492
+ by the projector with retry — so even a tracker that refuses the write cannot
493
+ recreate a duplicate PR while the guard is unavailable.
494
+ 2. Run the tick's post-admission projection pass, which applies freshly enqueued
495
+ label ops on the healthy path.
496
+ 3. Clear any stale tree for this issue, then add a fresh worktree at
497
+ `<workspaceRoot>/<issue>` cut from the bare mirror at `<mirrorRoot>/<repo>.git`,
498
+ on the run's branch off the repo's default branch.
499
+ 4. Allocate a session transcript under `<state dir>/sessions/`, one per attempt,
500
+ and move the run to `running`. The run record keeps the exact path and a
501
+ failure escalation quotes it, so you can read what the worker actually did.
502
+ 5. Run one omp session with the rendered brief, under the turn and wall-clock caps.
503
+ 6. Record the outcome:
504
+
505
+ | Outcome | Labels | Worktree | Escalation |
506
+ | --- | --- | --- | --- |
507
+ | `pushed-pending` | `agent:in-progress` stays while the daemon rechecks GitHub | removed | none |
508
+ | `pushed-green` | `agent:in-progress` stays while the PR is open | removed | none |
509
+ | `blocked` | swapped to `agent:blocked` | dirty tree committed to the branch, then removed | Tier 1 |
510
+ | `failed` / `killed` | swapped to `agent:failed` | dirty tree committed to the branch, then retained until the PR or issue is terminal | Tier 1 |
511
+ | unexpected error | swapped to `agent:failed` | same | Tier 1 |
512
+
513
+ `pushed-pending` and `pushed-green` are not the end of the row: later ticks
514
+ verify outstanding checks and settle the PR once it resolves, and a row that
515
+ settles gives up its `agent:in-progress` label. See
516
+ [what settles a green PR](#what-settles-a-green-pr).
517
+
518
+ Every label write the dispatcher makes goes through the label projection
519
+ outbox and is applied to the tracker by the projector with retry, strictly in
520
+ the order it was enqueued per issue — a later op for one issue never lands
521
+ before an earlier one that is still owed. A state-label swap on a dispatch
522
+ outcome enqueues the new label ahead of the old one's removal, so the issue
523
+ is never briefly bare (the shape eligibility reads as fresh work), while a
524
+ requeue (`swapToQueue`, `unblock`) enqueues its removals ahead of the queue
525
+ add for the same reason in reverse: the issue must not look claimable before
526
+ its stale state label is gone. A refused write is deferred with backoff
527
+ instead of dropped, and `status` shows any un-applied lag on a
528
+ `labels projection` row.
529
+
530
+ **Every continuable end salvages the tree first.** A turns-cap kill, a
531
+ wall-clock kill, a crash and a graceful block all leave a tree the next
532
+ attempt removes `--force` — only the run's *branch* is preserved across
533
+ attempts. So before the escalation is written, a dirty tree is committed to
534
+ the run's own branch as `wip(#<issue>): attempt <n> <ending> — auto-salvaged`
535
+ (everything, including files git has never seen) and pushed, and the
536
+ escalation says where it went: `WIP committed to <branch> @ <sha>`. A push
537
+ that is refused leaves the commit in this host's mirror and says so.
538
+
539
+ Blocking was excluded from this until #118, on the argument that a worker
540
+ which stops on purpose has turns left to commit for itself. It cost a
541
+ 34-file refactor: the worker blocked to ask whether a failing test was
542
+ obsolete — which is precisely a worker declining to commit a half-migrated
543
+ tree — and the daemon removed the tree seconds later, leaving the run branch
544
+ and `origin/main` on the same commit. A `pushed-green` or `pushed-pending`
545
+ run is now the only end that does not salvage: its deliverable is already on
546
+ a remote branch, and appending a WIP commit would turn the PR the daemon
547
+ just verified red.
548
+
549
+ **A salvage that fails keeps the tree and stops the issue.** If git refuses
550
+ the commit, the worktree is the only copy in existence, so it is retained
551
+ whatever the run's outcome was, the row records the failure, and the issue
552
+ is held out of dispatch with the `unsalvaged-wip` reason — because claiming
553
+ it is what would finally destroy the tree. `status` shows it under `wip` as
554
+ `UNSALVAGED`, and `omp-conductor unblock <n>` refuses. Recover the tree by
555
+ hand, then `unblock <n> --force` records that you accepted it and releases
556
+ the hold.
557
+
558
+ **A preserved tip is named to the next worker.** The sha is written to the
559
+ run row, shown by `status` and the board, and the continuation brief tells
560
+ the resuming worker the exact commit it is building on and that it is the
561
+ only copy.
562
+
563
+ Later ticks reap retained failure trees in bounded batches after the tracker
564
+ proves their PR merged/closed or their issue closed, provided no live run or
565
+ queued continuation owns the issue. Cleanup fetches remote refs first and
566
+ keeps any dirty tree or branch with uniquely local commits. Only then does it
567
+ remove the physical tree, prune registrations, and delete the obsolete local
568
+ mirror branch. Unknown tracker, network, repo, or git state is a no-op.
569
+
570
+ ### What a restart does to runs that were in flight
571
+
572
+ A `claimed` or `running` row is a promise that a worker process exists, and a
573
+ daemon that just started knows that promise is broken: its workers died with the
574
+ previous process. At startup — unless another daemon is alive, so a foreground
575
+ `daemon --once` cannot orphan a running daemon's real workers — every such row is
576
+ **salvaged first** (dirty tree → `wip(#N): attempt N killed by a daemon restart —
577
+ auto-salvaged` on the run's branch, same path as a turns-cap kill), then moved to
578
+ `orphaned`, with a log line naming the issue, the attempt and the worktree. That
579
+ frees the slots immediately; a fleet must never resume as deadlocked as it
580
+ crashed, and uncommitted edits must not wait for a human with `bun -e`.
581
+
582
+ Only the rows change after salvage. The issue keeps `agent:in-progress` — the
583
+ label is the crash guard against double-dispatch — and deciding what the dead
584
+ worker's remains are worth is the orchestrator's drain-duty judgement, spelled
585
+ out in its brief: an open green PR goes to the merge path, a salvaged sha is a
586
+ continuation hand-off, and a clean orphan has its label released so the next
587
+ tick re-claims it. Orphans consume `maxContinuationsPerIssue`, not failed
588
+ implementation attempts, so crashes cannot starve the retry needed for a real
589
+ code or CI failure — and a crash loop still escalates.
590
+
591
+ ### Deploying a new package onto a busy fleet
592
+
593
+ `systemctl restart` / `omp-conductor restart` is safe for **work product** once
594
+ this version is installed: startup salvage commits dirty trees before orphaning
595
+ rows, and salvage rewrites the mirror's managed `info/exclude` to the package's
596
+ current list before `git add` so a narrowed ignore cannot hide deliverables.
597
+
598
+ It is still disruptive for **in-flight sessions** because the worker process dies
599
+ and the attempt is spent. Update through the lifecycle command rather than
600
+ hand-installing or restarting individual surfaces:
601
+
602
+ ```bash
603
+ omp-conductor upgrade
604
+ ```
605
+
606
+ It pauses new claims, drains workers, installs one pinned release across all
607
+ surfaces, restarts, verifies twice, and resumes only if dispatch was initially
608
+ running. If an install, restart, or verification step fails, dispatch remains
609
+ paused and the command exits nonzero.
610
+
611
+ Do **not** edit files under the running install and expect the daemon to keep
612
+ dispatching — the integrity tripwire pauses and pages. Upgrade by whole release
613
+ so the new process records a fresh baseline.
614
+
615
+ ### What settles a green PR
616
+
617
+ The worker watches CI, then reports the PR URL and the exact remote head SHA it
618
+ observed. The daemon independently reads the PR again and requires it to be open,
619
+ non-draft, still at that head, and backed by a non-empty check rollup in which
620
+ every check succeeded or was skipped. Missing or nonterminal checks become
621
+ `pushed-pending` and are rechecked on later ticks; red or cancelled checks become
622
+ `failed` with a bounded job/log digest. Only verified evidence becomes
623
+ `pushed-green`.
624
+ If a failed report mentions PRs only in prose, the daemon retains the last URL
625
+ whose `owner/repo` matches the run's repository; links to other repositories are
626
+ ignored. This preserves the continuation target without trusting an unrelated
627
+ PR mentioned in the same report.
628
+
629
+ What happens after verification is a human's decision, taken minutes to days
630
+ later and never announced to the daemon — so every tick asks the tracker about
631
+ every pushed PR it is still holding:
632
+
633
+ | PR | Row becomes | Why |
634
+ | --- | --- | --- |
635
+ | merged | `merged` | The work landed. This is the state `merged` was reserved for. |
636
+ | closed without merging | `failed`, class `returned-for-revision`, with the PR in `lastError` | A human read the work and asked for another pass. Leaving it `pushed-green` strands the issue forever behind a PR nobody will merge, and calling it `merged` is a lie about code that is not on the base branch. `failed` is the honest row state; the class preserves the review decision and releases the issue so a re-queue can be attempted again. |
637
+ | still open | unchanged | The normal steady state. Its issue must stay occupied, or a second attempt lands on the live PR. |
638
+ | could not be determined | unchanged | A flaky network, a revoked token, a deleted PR. An unknown answer never settles a row; the next tick asks again for free. |
639
+
640
+ Run history is untouched. A PR closed without merging consumes one continuation,
641
+ not a failed implementation attempt; a merge spends neither budget. A settled
642
+ row also loses `agent:in-progress` from its issue: the row transition and label
643
+ removal are one fact, and a terminal answer about the PR proves no worker
644
+ process owns the issue,
645
+ so the duplicate-dispatch guard it exists for is spent. The removal is *enqueued
646
+ on the label projection outbox in the same breath as the terminal write* — a
647
+ durable local write that cannot fail on the tracker — so the row settles at once
648
+ and the projector applies the label with retry. A tracker that refuses the
649
+ removal can no longer strand it: the op stays owed, eligibility overlays the
650
+ pending removal so the issue is not held back by a label that is already decided
651
+ gone, and `status` shows the lag. Anything beyond that one release — a re-queue,
652
+ a `blocked` marker — is still the orchestrator's drain-duty judgement. One
653
+ unreachable PR costs its own row and nothing else; the rest of the sweep still
654
+ settles.
655
+
656
+ Until this existed, nothing ever revisited a `pushed-green` row: the startup
657
+ reconciler only settles rows that held a process, and `merged` went unwritten. On
658
+ 2026-08-07 the reference fleet reported three active runs whose PRs were all
659
+ merged and whose issues were all closed, through two daemon restarts — and because
660
+ the active set *is* the busy set, those three issues were permanently unclaimable.
661
+ A status page that has stopped being evidence is worse than no status page.
662
+
663
+ The label half of that outlived the row half by two days. On 2026-08-09 a merged
664
+ PR and a closed-unmerged one both settled their rows correctly and both left their
665
+ issues carrying `agent:in-progress`, which eligibility reads as "a worker owns
666
+ this" — with the brief forbidding the orchestrator from editing a state label and
667
+ `unblock` declining to clear that one, neither issue could ever be claimed again.
668
+
669
+ ### Base-branch health after merge
670
+
671
+ The daemon records two different facts after a merge:
672
+
673
+ - The **post-merge audit** attributes a regression to one merge. For up to 24
674
+ hours, it checks only push-triggered workflow runs for the exact merge SHA and
675
+ base branch. A new red result adds `base-branch-red` evidence and escalates.
676
+ - **Current health** drives `status` and release policy. On every sweep, the
677
+ daemon resolves the live head of each branch it merged into during the last
678
+ seven days, then reads only push-triggered runs for that head and branch. The
679
+ status row includes the head SHA and run count.
680
+
681
+ Current health is `green` only when every observed run completed successfully,
682
+ neutrally, or skipped. A failing conclusion is `red`; an in-progress or
683
+ unrecognised conclusion is `pending`; and a head with no push-triggered run is
684
+ `unknown`, never green. Pending and unknown heads are rechecked. Green and red
685
+ heads are read again when the branch moves, so an old verdict cannot describe a
686
+ new commit. If GitHub cannot return the head or its runs, the daemon keeps the
687
+ last honest row instead of replacing evidence with a network failure.
688
+
689
+ The `base-branch-green` release requirement reads this current live-head row for
690
+ the repository being released. Red, pending, unknown, and absent evidence all
691
+ refuse the release.
692
+
693
+ ### The settlement audit
694
+
695
+ Two lines of the worker report template were once taken on faith. `state:
696
+ pushed-green` stopped being believed in #85: a claim is now verified against
697
+ GitHub before a run settles. The `changed:` line stopped being *asked for* in
698
+ #488: the file list is derived at settlement from the pull request's own diff,
699
+ so a worker that writes no file list still settles with a correct one, and a
700
+ report whose list contradicts the diff is rewritten to say what the diff says.
701
+ The worker's narrative — what changed and why — is untouched; only the file list
702
+ stops being hand-authored.
703
+
704
+ So at settlement the daemon fetches the pull request's diff, derives the report's
705
+ `changed:` line from it, and checks the diff itself for weakened tests. What it
706
+ finds is a **settlement audit flag** — advisory, never a gate. A flagged run
707
+ settles exactly as an unflagged one does; nothing here can change a run's state,
708
+ hold a merge, or spend an attempt.
709
+
710
+ | Flag | Raised when |
711
+ | --- | --- |
712
+ | `changed-line-missing` | The settlement could not read the PR's diff, so no `changed:` file list was derived. The one disclosure fault left, and it is a tree-read failure, not a worker's. |
713
+ | `test-file-deleted` | A test file left the tree with no rename to account for it. |
714
+ | `test-disabled` | A `.skip` / `.only` / `xit` / `@pytest.mark.skip` / `t.Skip` marker appears on a line the PR added. |
715
+ | `assertions-removed` | An assertion was commented out, or a test file lost more assertions than it gained. |
716
+ | `test-timeout-raised` | A named timeout in a test file went up — compared against its own previous value, so a brand-new timeout is not a finding. |
717
+
718
+ A flag on a test file the dispatching issue never names is marked
719
+ `[unattributed]`: that is the "don't weaken tests you didn't write" case, and it
720
+ is the one worth reading first.
721
+
722
+ Flags reach you three ways: appended to the settlement report, stored on the run
723
+ row and shown under the run in `omp-conductor status` for as long as its PR is
724
+ open, and — once per flagged settlement — as a tier-1 escalation to the
725
+ orchestrator, whose brief says what judgement each flag invites.
726
+
727
+ **A clean, accurately reported PR produces nothing.** That is a design
728
+ constraint, not an aspiration: an audit that fires on honest work gets muted, and
729
+ a muted audit is worse than none because the fleet still believes it is being
730
+ checked. Every rule resolves ambiguity towards silence, and each accepts a named
731
+ blind spot to stay quiet — a renamed test file is not a deleted one (even when
732
+ git did not detect the rename, matched by basename), a `.skip` inside a string
733
+ literal or a recorded fixture is not a skip, and a rewritten test that keeps its
734
+ coverage is not a weakening.
735
+
736
+ The analyser is pure: it takes a parsed diff, the report and the issue text, and
737
+ returns flags. Only `Tracker.prDiff` touches the network, and a diff it cannot
738
+ read produces no flags *and says so* — silence about a diff nobody read is not a
739
+ clean bill.
740
+
741
+ ### Continuation runs
742
+
743
+ When a worktree is provisioned onto a branch that already exists in the mirror
744
+ (reattach after a prior attempt, orphan, or turns-cap auto-requeue), the worker
745
+ brief includes a **Continuation** section: read `git log` / `git diff` against
746
+ the default branch first, and do not recreate work already on the branch.
747
+
748
+ A **turns-cap kill with attempts remaining** salvages the tree, puts the queue
749
+ label back on, and skips the failed label so the next tick reclaims as a
750
+ continuation automatically.
751
+
752
+ ### Branch names
753
+
754
+ `<type>/<slug>`, where the type is `fix` when any label's last segment (after `:`
755
+ or `/`) is `bug`, and `feat` otherwise. The slug is the issue title folded to
756
+ `[a-z0-9-]`, and the whole ref is capped at 60 characters. It is computed from the
757
+ issue alone, so a retried run recomputes the same branch and finds its own work
758
+ instead of forking a second one.
759
+
760
+ ## Routing
761
+
762
+ An issue must carry **exactly one** `repo:<name>` label naming a repo in
763
+ `routing.repos`. The prefix is `routing.labelPrefix` and defaults to `repo:`.
764
+
765
+ Routing never guesses. An issue it cannot resolve to a single configured checkout
766
+ is handed back as unroutable:
767
+
768
+ | Reason | Condition |
769
+ | --- | --- |
770
+ | `no-repo-label` | The issue carries no label starting with the prefix. |
771
+ | `multiple-repo-labels` | It carries two or more distinct prefixed labels. A repeated identical label is deduplicated, not treated as an ambiguity. |
772
+ | `unknown-repo` | Its single prefixed label names a repo that is not in `routing.repos`. |
773
+
774
+ In all three cases the issue is **escalated at Tier 1 and never dispatched**. The
775
+ fix is always the same, and the escalation says so: put exactly one
776
+ `repo:<name>` label on the issue.
777
+
778
+ This is deliberate. A request that spans two repos, taken whole by one worker, is
779
+ the precise failure this guard exists to prevent: the worker cannot open a PR
780
+ against two checkouts, so it improvises — it vendors a copy, edits the wrong repo,
781
+ or produces a PR that cannot be merged without the other half. Splitting a
782
+ multi-repo request is a human decision about contracts; it is not something to
783
+ infer from a label. Sending the issue back costs a label edit; guessing costs a
784
+ bad merge.
785
+
786
+ ## Host sizing and memory
787
+
788
+ Workers are **child processes** of the daemon (plus one long-lived orchestrator
789
+ session), each its own pid, talking back over a unix socket. They are still
790
+ inside the service's cgroup, so systemd's Memory peak for
791
+ `omp-conductor.service` is daemon + every live worker + the orchestrator + any
792
+ MCP stdio children those sessions mount. `MemoryMax=` governs that whole total,
793
+ not one process.
794
+
795
+ The generated unit starts `omp-conductor daemon --port 8787` without a
796
+ `--project` filter, so one daemon serves every configured project. Its automatic
797
+ `MemoryMax=` tier uses the sum of resolved `maxConcurrentWorkers` values across
798
+ all projects: `3G` for a total of one worker, otherwise `5G`.
799
+
800
+ On the reference deploy that produced [issue #51](https://github.com/TerrifiedBug/conductor/issues/51):
801
+
802
+ | Shape | Observed |
803
+ | --- | --- |
804
+ | Idle / workers restarting | ~430 MB RSS for the daemon alone |
805
+ | Two workers + orchestrator, busy | **3.2–4.2 GB** Memory peak for the unit; up to ~800 MB swap |
806
+
807
+ That peak is **expected for concurrent SDK sessions**, not evidence of a
808
+ conductor-side leak: the SQLite store is disk-backed, admission state is
809
+ per-tick, and worker sessions are disposed when a run ends. What grows is the
810
+ session heap (conversation + tool output); a single graph-assisted run has been
811
+ measured in the hundreds of thousands of characters of tool output.
812
+
813
+ **Practical guidance**
814
+
815
+ - Prefer **≥16 GiB RAM** for the default `maxConcurrentWorkers: 2`, and do **not**
816
+ co-locate ClickHouse / other multi-GB services beside that fleet on an ≤8 GiB
817
+ box.
818
+ - On hosts under ~16 GiB, keep the **sum** of every project's
819
+ `maxConcurrentWorkers` at **1**. `omp-conductor setup` chooses that default
820
+ for a new project when it can read host RAM and warns before applying a
821
+ configuration whose combined capacity exceeds the host recommendation.
822
+ - Supervise the daemon with a unit that sets `SuccessExitStatus=0 143`; setup
823
+ renders `MemoryMax=3G` for one configured worker and `MemoryMax=5G` for two
824
+ or more.
825
+ - `omp-conductor status` prints daemon `rss` from `/healthz` when the process is
826
+ up, so you can see pressure without scraping journald.
827
+
828
+ ## Caps
829
+
830
+ Caps resolve per project: the global `defaults` block, then the project's own
831
+ `caps` layered on field by field, so a project that pins one cap still inherits the
832
+ rest. `0` is a real value (a hard stop), not "unset".
833
+
834
+ | Cap | Default | What it protects |
835
+ | --- | --- | --- |
836
+ | `maxConcurrentWorkers` | `2` (setup may write `1` on &lt;16 GiB hosts) | Parallel omp sessions, each a child process of the daemon and all inside its cgroup. Two, because **CI runner slots, not model tokens, are the usual throughput ceiling** — a third worker would starve its own PR checks on a small self-hosted runner pool. On hosts under ~16 GiB RAM, prefer `1` so the unit stays out of swap ([host sizing](#host-sizing-and-memory)). Raise it only if you actually have the runners *and* the RAM. |
837
+ | `maxConcurrentWorkersPerRepo` | `1` | Max live workers in the **same repo**. The mirror, branch-protection staleness and shared CI egress are all per-repo collision domains, so extra slots should land on other repos. Raise it only when a repo genuinely needs two workers at once. |
838
+ | `dailySpendUsd` | `25` | Rolling-day spend ceiling in USD, or `null` for no spend gate. `0` is a hard stop. Metered from assistant `usage.cost.total`. |
839
+ | `planUsage` | `null` (unmetered) | Subscription/plan allowance guard: `{ "windowId": "anthropic:7d", "maxUsedFraction": 0.85 }`, or `null` for no plan gate. Independent of `dailySpendUsd` — see [Plan allowance](#plan-allowance-planusage) below. |
840
+ | `workerMaxTurns` | `120` | Base ceiling for each new worker. Catches a session looping without converging; use `omp-conductor extend` to raise one live run or one issue's next attempt without changing this default. |
841
+ | `workerMaxTurnsCeiling` | `240` (twice the effective `workerMaxTurns` when omitted) | Upper bound for per-issue turn extensions. Prevents the loopback control from granting an unbounded worker budget. |
842
+ | `workerWallClockMs` | `5400000` (90 minutes) | Wall-clock ceiling for one worker. A session that is merely stuck spends no turns, so turns alone cannot detect it. |
843
+ | `maxAttemptsPerIssue` | `2` | Failed implementation or CI attempts before escalation. Operational stops do not consume this budget, so salvage can continue without stealing the retry needed for a real failure. |
844
+ | `maxContinuationsPerIssue` | `2` | Cap-kill, daemon-orphan and answered-block resumes before escalation. This independently bounds crash/resume loops. |
845
+
846
+ Days are counted from **local midnight**, matching how a human reads "today".
847
+
848
+ Set `dailySpendUsd` to `null` (wizard: blank) for no money gate — turns and wall-clock still apply. Hitting a numeric `dailySpendUsd` is not the same as hitting the other caps. A concurrency
849
+ limit simply defers work to a later tick. The spend cap **pauses the daemon and
850
+ pages at Tier 2**: a loop that is burning money has to halt itself, because
851
+ waiting for someone to notice tomorrow is how a runaway becomes expensive.
852
+ Work resumes only after `omp-conductor resume`.
853
+
854
+ `workerMaxTurns` and `workerWallClockMs` are enforced inside the session driver.
855
+ The daemon reads a live run's effective turn ceiling at every turn boundary. Use
856
+ `omp-conductor extend <issue> --turns N [--project NAME]` to raise it without
857
+ restarting or reconstructing the session. For a live worker, extension is
858
+ monotonic: equal or lower values are refused. If the latest run is failed,
859
+ killed, orphaned, or blocked, the command instead stores a one-shot ceiling for
860
+ that issue's next attempt. A next-attempt ceiling must exceed the effective
861
+ project base, and every extension must stay at or below
862
+ `workerMaxTurnsCeiling`. `status` shows both active ceilings and pending
863
+ next-attempt overrides. The store consumes an override atomically when it claims
864
+ the next run, so later attempts return to the project base. Config edits change
865
+ that base on the next tick but do not change workers already in flight. A cap
866
+ that fires aborts the run, records it as `killed`, and names the ceiling in the
867
+ escalation.
868
+
869
+ Pause one live worker cooperatively with
870
+ `omp-conductor worker pause <issue> [--project NAME]`. The daemon aborts the
871
+ active turn to an idle harness state, freezes the remaining wall-clock budget,
872
+ and keeps the run in the Running lane. `status` overlays `paused`/`pausing` from `/healthz` on that active-run line while the SQLite row stays `running`. `omp-conductor worker resume <issue>`
873
+ continues the same session with a prompt to re-check its last action before
874
+ proceeding. To end that run instead, use
875
+ `omp-conductor worker stop <issue> --reason TEXT [--project NAME]`. Stop works
876
+ from running or paused, salvages dirty work before removing the worktree, records
877
+ the distinct terminal `stopped` state, and removes `agent:in-progress` through
878
+ the label outbox. A salvage failure keeps the only copy in place and reports its
879
+ path. Stopped runs consume neither implementation-failure nor continuation
880
+ budget. Repeating stop reports the already-terminal state. These worker controls
881
+ are separate from fleet-level `pause`, which stops new claims.
882
+
883
+ ### Plan allowance (`planUsage`)
884
+
885
+ `dailySpendUsd` meters money, which is the only thing an API-billed account can
886
+ run out of. A fixed-price subscription cannot be expressed that way: the real
887
+ ceiling is a **provider allowance** — a weekly token window whose marginal
888
+ dollar cost is zero and whose exhaustion stops every session on the host.
889
+ Pricing that into the dollar meter would mean inventing a number.
890
+
891
+ `planUsage` is the second, independent guard. It reads `omp usage --json` — the
892
+ structured form of the harness `/usage` view — and holds new claims while the
893
+ named window is at or over its threshold. Running workers finish normally, and
894
+ **the daemon is not paused**: the window resets on the provider's clock, so
895
+ dispatch resumes by itself once a fresh reading is below the threshold. Nothing
896
+ estimates a plan quota from conductor's own transcript token counts.
897
+
898
+ ```json
899
+ "caps": {
900
+ "planUsage": { "windowId": "anthropic:7d", "maxUsedFraction": 0.85 }
901
+ }
902
+ ```
903
+
904
+ **Naming the window.** `limits` in the payload is a *list*, not a single
905
+ number: one Anthropic account reports `anthropic:5h`, `anthropic:7d` and the
906
+ tier-scoped `anthropic:7d:fable` at the same time, and other providers add
907
+ their own. So the cap names its window rather than taking whichever entry came
908
+ first. Run this on the fleet host and copy an `id`:
909
+
910
+ ```bash
911
+ omp usage --json | jq -r '.reports[].limits[] | "\(.id) \(.amount.usedFraction) \(.amount.unit)"'
912
+ ```
913
+
914
+ A bare window key (`"7d"`) also works, but **only** when exactly one reported
915
+ allowance carries it. On an Anthropic account `7d` matches two, and the guard
916
+ refuses to guess.
917
+
918
+ **`maxUsedFraction` is a fraction, not a percentage.** `0.85` holds at 85%.
919
+ A value outside `0`–`1` is rejected at config load, because `85` would mean
920
+ "hold at 8500% consumed" — a guard that reads as configured and can never fire.
921
+ Comparison always goes through the provider's `usedFraction`, never a raw
922
+ count: `unit` is `percent` for Anthropic and `unknown` with raw counts for
923
+ `xai-oauth`, so a threshold compared against `used` misreads any non-percent
924
+ provider by orders of magnitude.
925
+
926
+ **Availability policy.** The guard never displays a number it did not read, and
927
+ never shows a fabricated `0% used`. What each situation does:
928
+
929
+ | Situation | `status` / `board` | New claims |
930
+ | --- | --- | --- |
931
+ | `planUsage: null` | `unmetered` | admitted |
932
+ | Window below threshold | `5% / 85% of anthropic:7d used · resets in 6d 2h` | admitted |
933
+ | Window at or over threshold | same, plus `holding new claims` | **held** (`plan-usage-cap`, Tier 1) |
934
+ | No provider reports a readable allowance, or `omp usage --json` fails | `unavailable — <reason>` | admitted for up to 30 minutes, then **held** and paged at Tier 2 |
935
+ | `windowId` names a window the reading does not contain | `window "<id>" is not in this reading — Reported: …` | **held**, Tier 2 |
936
+ | `windowId` matches more than one allowance | `window "<id>" matches …` | **held**, Tier 2 |
937
+ | The window reports nothing a fraction can be derived from | `window "<id>" reports no comparable fraction …` | **held**, Tier 2 |
938
+
939
+ The split is deliberate. A *read error* is transient — a token refresh, a
940
+ provider 502, `omp` briefly absent mid-upgrade — and stalling a fleet on one
941
+ would cost more than admitting through it, since spend, turns, wall clock and
942
+ concurrency are all still enforced. Half an hour of continuous failure is not
943
+ an outage, it is a broken meter, and a plan-capped fleet running on a broken
944
+ meter is how the allowance gets spent to zero unnoticed. A *successful* read
945
+ that does not contain the configured window is not a read error at all: the
946
+ source answered, and it says the config names something that is not there. That
947
+ fails closed immediately, like every other config fault in this package, and
948
+ recovers by itself as soon as a reading contains the window again.
949
+
950
+ Readings are cached for 60 seconds (15 for a failure) so one tick costs one
951
+ provider call rather than one per candidate, and a cached reading is dropped
952
+ the moment its own `resetsAt` passes — that is what makes admission resume at
953
+ the rollover instead of a TTL later. `omp usage invalidate` clears omp's own
954
+ cache; conductor picks the change up at its next read.
955
+
956
+ Both controls are shown separately, never folded together — `status` prints a
957
+ `spend today` row and a `plan usage` row, and the board's admission line ends
958
+ with `spend $2.40/$25.00 | plan 5%/85%`.
959
+
960
+ ## Worker model
961
+
962
+ `workerModel` on a project pins the model its workers run on, as a pattern in
963
+ omp's own model/role syntax (whatever `/model` accepts). It sits beside `caps`
964
+ rather than inside them, because it is not a ceiling:
965
+
966
+ ```json
967
+ "workerModel": "smol"
968
+ ```
969
+
970
+ Omit it and the harness picks, which is the right answer until you have a reason.
971
+ The pattern is passed through unresolved: omp resolves it after its extensions
972
+ load, so a name this package has never heard of still works. If the harness cannot
973
+ honour the pattern it says so, and the daemon logs that per run:
974
+
975
+ ```text
976
+ #412 model fallback: <what the harness substituted>
977
+ ```
978
+
979
+ Worth reading the log for. A run that quietly used a weaker model than you chose
980
+ otherwise looks like a run that was merely unlucky.
981
+
982
+ ## Code-graph discovery
983
+
984
+ Optional, off unless you answer yes in the wizard, and worth answering yes to for
985
+ one measured reason: **workers spend most of a run finding code, not changing it.**
986
+ On the dogfood fleet a single run typically spends 30–62 `read` calls and 32–69
987
+ `bash` calls against 9–24 edits — 215–390k characters of tool output, roughly four
988
+ fifths of a 120-turn budget — and the runs that died at the turns cap died with
989
+ the work unfinished. A code graph answers "who calls this" and "where is this
990
+ defined" in one call instead of twenty greps.
991
+
992
+ ### Two things this package does not do for you
993
+
994
+ `omp-conductor` never installs, starts, imports, or depends on the indexer for
995
+ dispatch. With `graphProject` unset, nothing about dispatch, caps, escalation, or
996
+ status changes. A fresh host needs both of these before an index is worth
997
+ anything, and `setup graph` reports them as step 0:
998
+
999
+ 1. **`codebase-memory-mcp` on PATH** — a separate project,
1000
+ [DeusData/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp).
1001
+ 2. **Mounted as an MCP server** in `~/.omp/agent/mcp.json`, on the account the
1002
+ daemon runs as. Miss this and the failure is silent: every index builds
1003
+ correctly, no worker session can read any of them, so workers fall back to
1004
+ grepping and the feature looks like a no-op. `setup graph` prints the entry.
1005
+
1006
+ Say yes and the wizard asks for one root, then derives one clone per routed repo
1007
+ underneath it (default `~/.cache/conductor-graph/<org>/<repo>`) and writes it to
1008
+ each repo's [`graphProject`](#configuration). The only automatic interaction is
1009
+ a bounded, read-only health query; this package never clones, fetches, builds an
1010
+ index, or changes systemd. Dispatch, caps and escalation do not depend on graph
1011
+ health. On a fleet configured before this key existed, `omp-conductor setup` and
1012
+ the `code graph` area add it in two prompts — see
1013
+ [Changing one setting](#changing-one-setting).
1014
+
1015
+ ### Why the clone, and not your checkout or the worktree
1016
+
1017
+ This is the part that decides whether the feature helps or hurts, so it is worth
1018
+ being blunt about all three candidates.
1019
+
1020
+ | Directory | Why not |
1021
+ | --- | --- |
1022
+ | **The worker's worktree** | An index is keyed by the realpath of the directory it was built from, and has no git-worktree awareness. A run's `worktrees/<issue>` path is therefore *always* an empty project — a worker that queried its own cwd would get silence, conclude there is no graph, and spend the run grepping. This is why `graphProject` is an absolute path in the config and not something derived at run time. |
1023
+ | **Your own checkout** | Refreshing an index means resetting the clone to its default branch. In a directory you work in, that either destroys uncommitted work or — if it is made safe instead — indexes whatever feature branch you left checked out, so the fleet orients against your WIP. |
1024
+ | **A conductor mirror** | The daemon's mirrors are bare. There is no working tree to index. |
1025
+
1026
+ So `graphProject` names a fourth thing: a clone that exists only to be indexed,
1027
+ that nothing human ever edits, and that is therefore safe to `git reset --hard`
1028
+ every night. The worker brief names that path, tells the session to match it
1029
+ against `list_projects`' `root_path` and query by the `name` beside it, and says
1030
+ plainly that the graph is a snapshot which does **not** contain the worker's own
1031
+ edits — orient with it, then read the real file before changing it.
1032
+
1033
+ ### Creating and refreshing them
1034
+
1035
+ ```bash
1036
+ omp-conductor setup graph --print # print the plan: clones, index commands, units
1037
+ omp-conductor setup graph # run it: clone, install, enable, seed, verify
1038
+ ```
1039
+
1040
+ `setup graph --print` prints a `git clone` for every clone that does not exist yet, the
1041
+ one-shot index command per repo, and a `cbm-reindex.service` + `cbm-reindex.timer`
1042
+ pair built from the project's own repos and branches. `--write` stages all three
1043
+ in the state directory and prints the two `sudo` lines that install and enable
1044
+ them; it never runs `systemctl`.
1045
+
1046
+ **Run it as the account the fleet runs as, never under `sudo`** — it refuses if
1047
+ you try. Everything it derives resolves per-account: the config it loads, the
1048
+ state directory it stages into, and the `HOME`/`User=` it bakes into the unit.
1049
+ Under root you get a timer that goes green while writing indexes into
1050
+ `/root/.cache`, where no worker session looks — silent, and indistinguishable
1051
+ from the feature simply not helping. Only installing the units needs root, which
1052
+ is why that is two separate printed commands.
1053
+
1054
+ Two properties of the generated unit are deliberate:
1055
+
1056
+ - **It is a timer, not the server's own watcher.** That watcher lives inside a
1057
+ connected MCP session and dies with it, so an ephemeral worker session keeps
1058
+ nothing fresh. The refresh has to come from outside the fleet.
1059
+ - **It fails loudly.** The refresh is `set -euo pipefail`, then per repo
1060
+ `git fetch --prune origin` and `git reset --hard origin/<its own defaultBranch>`
1061
+ before indexing. Nothing is `|| true`-ed, so a fetch that has been broken for a
1062
+ week turns the unit red instead of quietly re-indexing a stale tree and exiting
1063
+ `0` — a green timer serving a month-old graph is worse than no graph at all.
1064
+
1065
+ The unit spells out `HOME` and an explicit `PATH`, because systemd supplies
1066
+ neither usefully: the indexer resolves its store from `HOME`, systemd's default
1067
+ `PATH` has no `~/.local/bin`, and the indexer shells out to `git`. Both are the
1068
+ user that ran `setup graph`; the unit sets no `User=`, so check them if that is
1069
+ not the account the timer runs as.
1070
+
1071
+ ### Seeing whether the graph is usable
1072
+
1073
+ When at least one routed repo has `graphProject`, `omp-conductor status` adds a
1074
+ `code graph` block. It proves the indexer is on `PATH`, the worker MCP config
1075
+ mounts it, every configured clone exists and exactly matches an indexed
1076
+ `root_path`, the refresh timer is enabled and active, and the last service run
1077
+ succeeded within 45 minutes. A running daemon refreshes this evidence every
1078
+ minute and publishes the cached result through `/healthz`; status probes the host
1079
+ directly when that cache is unavailable. Every command is read-only, runs with a
1080
+ one-second timeout, and graph degradation never changes `/healthz.ok` or blocks
1081
+ dispatch. Unconfigured projects omit the block entirely.
1082
+
1083
+ ## Escalation tiers
1084
+
1085
+ | Tier | Meaning | Raised by | Delivered to |
1086
+ | --- | --- | --- | --- |
1087
+ | 1 | "Not a human's problem yet" — the run is parked and safe. | Unroutable issue, blocked run, failed or killed run, dispatch error, attempts exhausted. | The orchestrator session, as an injected prompt. Falls back to an issue comment when no orchestrator is running, or when it will not accept the injection. |
1088
+ | 2 | "The fleet is stopped until you look." | Daily spend cap reached; the installed package changed under the running daemon. Either way the project is already paused. | Telegram, when `escalation.telegramChatId` is set and a bot token is readable; otherwise it falls back to the issue comment. |
1089
+
1090
+ **The orchestrator** is one persistent, file-backed session per daemon run, resumed
1091
+ across restarts so it remembers what it has already handled. Its `cwd` is the state
1092
+ directory, deliberately not a checkout. Delivery resolves when the harness *accepts*
1093
+ the prompt, not when the model answers it, so a tick never parks behind a model; an
1094
+ injection arriving mid-thought queues as a follow-up instead of interrupting the
1095
+ turn in flight. Its standing orders are explicit: re-brief the worker, file or
1096
+ comment on issues, or promote to tier 2, and never edit product code or push a
1097
+ branch. Merging is the one line worded from config — see [`authority`](#configuration).
1098
+ If it fails to start, the daemon logs a warning and runs on, with tier-1
1099
+ escalations degraded to issue comments.
1100
+
1101
+ **Or no orchestrator at all.** Set `escalation.orchestrator` to `"external"` when
1102
+ you already run your own supervising session — a visible TUI session in a pane,
1103
+ typically. The daemon then starts none of its own and every tier-1 escalation
1104
+ posts as an issue comment, which is what that session drains. One brain, and it
1105
+ is the one you can watch.
1106
+
1107
+ **Answering a tier 1 is only half of it.** A blocked or failed run leaves its state
1108
+ label on the issue, and eligibility reads any state label as disqualifying, so an
1109
+ answered issue that keeps one is never re-claimed and the answer is inert — nothing
1110
+ fails, the issue just stops existing as far as dispatch is concerned.
1111
+ [`omp-conductor unblock <issue>`](#cli-reference) is the way back: it clears the
1112
+ label through the same tracker the dispatcher writes with, including
1113
+ `agent:in-progress` when the newest recorded run is terminal, since a terminal row
1114
+ is proof the worker process is gone. The brief tells the
1115
+ orchestrator to run that verb rather than edit the label itself, and that is not a
1116
+ formality — orphan detection works by comparing `agent:in-progress` labels against
1117
+ live runs, and it is only trustworthy while every state label on the tracker was
1118
+ written by this package.
1119
+
1120
+ Tier 2 borrows the bot token that `omp-telegram` already owns, at
1121
+ `~/.omp/agent/telegram/.env` (or `$OMP_TELEGRAM_STATE_DIR/.env`). If you run that
1122
+ bot, Tier 2 needs no extra configuration beyond the chat id. If the token is
1123
+ absent, Tier 2 degrades to the issue comment instead of failing. The token is
1124
+ never logged, and it is redacted out of any error text that could reach a public
1125
+ issue comment.
1126
+
1127
+ **Escalations are deduplicated.** The dispatcher re-notices the same unroutable
1128
+ issue on every poll, so a ledger in the store — keyed by project, issue, tier and
1129
+ summary — makes a recurring condition page **once** and suppresses the five-minute
1130
+ repeats. The marker is recorded only on successful delivery, so a page that could
1131
+ not be delivered is retried on the next tick instead of being written off as sent.
1132
+ The spend-cap and integrity-tripwire summaries carry the date, so the same
1133
+ condition pages again tomorrow but only once per day.
1134
+
1135
+ If `fallbackToIssueComment` is off and no Telegram transport is configured,
1136
+ delivery throws instead of dropping silently. The failure is logged and retried,
1137
+ because a swallowed escalation looks exactly like a healthy fleet.
1138
+
1139
+ ## Report delivery (the outbox)
1140
+
1141
+ Escalations are the daemon's. **Reports** — the material events and the daily
1142
+ digest your [`reporting.scope`](README.md#your-workflow-vs-the-package) asks for — are
1143
+ written by the orchestrator, and until v0.3.26 they were also *delivered* by it:
1144
+ a report reached you only if the model remembered to call `telegram_send`. On
1145
+ 2026-08-06 a suite release and two tier-2 escalations were written that way and
1146
+ none of the three arrived, and nothing anywhere recorded that fact — an undelivered
1147
+ report and a quiet tick look identical.
1148
+
1149
+ ### Material events survive the session
1150
+
1151
+ A deferred digest does not use the session transcript as its source of truth.
1152
+ Record each ordinary outcome when it happens:
1153
+
1154
+ ```bash
1155
+ omp-conductor event record \
1156
+ --category merge \
1157
+ --summary "#42 merged" \
1158
+ --evidence "https://github.com/acme/api/pull/42"
1159
+ ```
1160
+
1161
+ `--category` is a short lowercase slug. `--summary` states the outcome, and
1162
+ `--evidence` names the issue, PR, release, run, commit, or URL that proves it.
1163
+ Use `--occurred-at <ISO timestamp>` when the event happened earlier; otherwise,
1164
+ the command uses the current time. The command writes one row to SQLite and
1165
+ sends nothing. The row survives later ticks, session compaction, session
1166
+ replacement, and daemon restarts.
1167
+
1168
+ When a digest is due, its tick prompt lists a bounded, oldest-first set of
1169
+ owed material events and deferred escalations. Each line includes its ledger id.
1170
+ The prompt gives the exact handoff shape:
1171
+
1172
+ ```bash
1173
+ omp-conductor report \
1174
+ --kind digest \
1175
+ --events EVENT_ID_1,EVENT_ID_2 \
1176
+ --notices NOTICE_ID_1,NOTICE_ID_2 \
1177
+ --text "<the whole digest>"
1178
+ ```
1179
+
1180
+ Remove the id of any row you did not use. Omitted rows stay owed. The report row
1181
+ and the named ledger rows are associated in one SQLite transaction. If the
1182
+ handoff fails, no row is consumed. If a daily report deduplicates against a
1183
+ daily report already queued that day, newly named rows also stay owed. If
1184
+ delivery exhausts its retry budget and the report becomes `failed`, its rows
1185
+ return to the owed backlog, where a replacement digest can claim them.
1186
+ `omp-conductor status` always shows the
1187
+ material-event and held-escalation backlog counts, including the age of the
1188
+ oldest row when one exists.
1189
+
1190
+ This accumulator does not poll GitHub and does not infer outcomes from tracker
1191
+ state. The orchestrator still decides what is material and records the evidence.
1192
+ The mechanism only makes that decision durable until a non-failed digest owns it.
1193
+
1194
+ Authorship still needs judgement the daemon does not have, so it stays with the
1195
+ model. Delivery does not, so it moved:
1196
+
1197
+ ```bash
1198
+ omp-conductor report --text "<the whole report>" # immediate report, when policy permits
1199
+ omp-conductor report \
1200
+ --text "<the whole digest>" --kind digest \
1201
+ --events EVENT_IDS --notices NOTICE_IDS # use row ids from its tick
1202
+ ```
1203
+
1204
+ The command persists the text before anything is sent and prints a durable
1205
+ handoff id. An immediate report admitted during quiet hours is stored as a held
1206
+ notice and prints that id; the daemon includes it in the next digest or in a
1207
+ catch-up report when the configured window opens. Otherwise it writes a
1208
+ `reports` row and prints its report id. Both survive the session being
1209
+ compacted, interrupted or restarted,
1210
+ and the daemon being restarted under it. The daemon delivers over the same bot
1211
+ token tier 2 uses, with bounded retries, and `omp-conductor status` lists
1212
+ anything it still owes. If availability closes after a material report was
1213
+ queued but before its first attempt, the outbox atomically converts that row to
1214
+ the same held-notice path instead of leaking the update through quiet hours.
1215
+
1216
+ ### Answering a person, in the thread they wrote in
1217
+
1218
+ A report is an update; an answer is a conversation, and it goes back where the
1219
+ question came from. `telegram_send` keeps the active forum topic **only while it
1220
+ names no chat** — `thread_id` defaults to the active topic when `chat_id` is
1221
+ omitted — so an orchestrator that helpfully supplied `chat_id` (and nothing
1222
+ else) answered three topic messages in the main chat instead (#366). The
1223
+ orchestrator floor now says to name **neither** target or **both**, and the
1224
+ dispatcher refuses `chat_id` without `thread_id` for a project that configured
1225
+ `escalation.telegramTopicId`. A flat-chat project is unaffected. When a pinned
1226
+ topic id has gone stale — the bridge re-claims pane topics across restarts — the
1227
+ live claim for the project's herdr space is used instead, falling back to a
1228
+ claim titled for the project, so a restart does not quietly move every page into
1229
+ the main chat (#407, #412).
1230
+
1231
+ A locally injected tick has no inbound message to inherit a topic from, so a
1232
+ bare `telegram_send` there has nothing to preserve. That turn addresses the
1233
+ operator from config instead:
1234
+
1235
+ ```bash
1236
+ omp-conductor message --text "<the message>" # this project's chat and topic
1237
+ omp-conductor message --category tier2 --text "cut 0.16.0 tonight?" # blocking question: declared category + open decision row
1238
+ omp-conductor message --text "QUESTION: do we release?" # marker form still maps to decision-needed
1239
+ ```
1240
+
1241
+ A question — a declared `--category` that is not `material`, or text beginning
1242
+ `QUESTION:` — is an ask, not a send: before delivery the command records an
1243
+ open decision row for the question (parked on silence: "nobody answered" keeps
1244
+ the row open and pending, re-surfaced in every tick until answered or the
1245
+ seven-day expiry, never an approval) and prints the id to resolve it later
1246
+ with `omp-conductor decision resolve <id> --answer "…"`, so the orchestrator
1247
+ never has to remember a separate `decision open` (#520).
1248
+
1249
+ It is not a bypass of the interrupt policy: the same availability decision an
1250
+ autonomous Telegram tool call gets is applied, so a message whose category the
1251
+ policy defers is durably held for the digest or the working-hours catch-up and
1252
+ the command prints that held-notice id instead of claiming delivery — a
1253
+ *blocking* question must therefore declare a category the fleet interrupts on,
1254
+ not default to the marker. It is also not a report — it
1255
+ leaves no `reports` row, and nothing retries it.
1256
+
1257
+ ### Delivery is at-least-once, and the docs will not pretend otherwise
1258
+
1259
+ The Telegram Bot API accepts no client-supplied idempotency key and offers the
1260
+ bot no readable record of what it has already sent. There is nothing to replay a
1261
+ request against and nothing to reconcile with, so **exactly-once delivery cannot
1262
+ be built on this transport** and this package does not claim it. `delivered` is
1263
+ proof that Telegram accepted *an* attempt, never proof that exactly one message
1264
+ exists.
1265
+
1266
+ What it does instead is make the ambiguity explicit and always resolve it in the
1267
+ direction of the duplicate:
1268
+
1269
+ | State | Meaning | What you do |
1270
+ | --- | --- | --- |
1271
+ | `pending` | Nothing is in flight. Never attempted, or the last attempt failed **definitively** — see below. Nobody has this report. | Nothing. It retries on a bounded backoff (30s doubling to a 15-minute floor) and `status` shows the error. |
1272
+ | `sending` | A request left this host and its outcome was never learned: the daemon died, or the request was cut off after the bytes went out. Telegram may be holding the message. | Nothing, but expect a possible duplicate. Check the chat if you want to know now. |
1273
+ | `delivered` | Telegram answered `ok: true`. The row records the message id from the response **body**, not the HTTP status. | Nothing. |
1274
+ | `failed` | The retry budget ran out — six attempts, roughly half an hour. | Fix the transport. This state pages tier 2 in its own right — see below. |
1275
+
1276
+ The row is written `sending`, with the id of the attempt about to run, **before**
1277
+ the request is made. A crash in that window therefore leaves an explicitly
1278
+ ambiguous row rather than a silently lost one. The next daemon start sweeps every
1279
+ `sending` row, retries it, and the retried message carries the report id and a
1280
+ plain-English line saying it may already be in the chat. A duplicate you can spot
1281
+ by its report id is much the cheaper of the two mistakes; a silently dropped
1282
+ report is the entire reason this exists.
1283
+
1284
+ #### Two kinds of failure, and only one of them is quiet
1285
+
1286
+ A failed send is classified where the socket is watched, not by the caller, and
1287
+ the two classes are treated differently on purpose:
1288
+
1289
+ | Outcome | What happened | Row | Retry says |
1290
+ | --- | --- | --- | --- |
1291
+ | **Definitive** — nobody has it | Telegram answered and refused it (`{"ok":false}` under any status, or a non-2xx status), or the connection never opened at all (refused, DNS failure) so the request provably never left. | back to `pending`, backoff, attempt counted | nothing special — it *is* a first attempt |
1292
+ | **Outcome unknown** — Telegram might have it | The request was cut off after it left: timeout, abort, socket reset, `EPIPE`. Or the POST came back `200` and the **response body could not be read** — Telegram had already decided and the answer was lost coming back. | stays `sending`, flagged as a possible repeat | `POSSIBLE REPEAT`, with the report id to compare against |
1293
+
1294
+ Anything that cannot be classified confidently is treated as **outcome unknown**.
1295
+ That default is deliberate and is the safe direction: the worst case is a
1296
+ duplicate you were warned about, against a delivered report re-posted as though
1297
+ it were new, with nothing anywhere saying it might be a second copy.
1298
+
1299
+ The half of "never double-post" that *is* achievable is enforced: a report cannot
1300
+ be **concurrently** in flight twice. Claiming a report is a conditional update,
1301
+ so only one attempt can move a `pending` row, and every terminal transition names
1302
+ the attempt it is settling — a request that answers after its row was reclaimed
1303
+ is discarded rather than allowed to overwrite a newer attempt's outcome. That is
1304
+ what stops a retry storm.
1305
+
1306
+ ### A report nobody can deliver is itself news
1307
+
1308
+ A report that exhausts its retries is marked `failed` **and escalates as tier 2**.
1309
+ This rides the transport that just failed, which is deliberate and accepted: the
1310
+ common failure is a wrong chat id or a bot kicked from the chat, not a global
1311
+ Telegram outage, and in both of those the page reaches an operator who is
1312
+ otherwise being told nothing at all. If the whole channel is down the page
1313
+ degrades to a line in `daemon.log` and the `reports` block in `status`, which is
1314
+ then the only surface — a report has no tracker issue, so there is no issue
1315
+ comment to fall back to. The page goes through the ordinary escalation ledger and
1316
+ carries the report id, so one undeliverable report pages exactly once.
1317
+
1318
+ ### Daily digests are deduplicated from the ledger
1319
+
1320
+ With `digest.cadence: "daily"`, `--kind digest` is accepted at most once per
1321
+ **local** day, per project. The second hand-over on the same day is refused and
1322
+ told which report already holds the slot, including when that report has already
1323
+ been delivered. This is decided from the `reports` table, not from the model's
1324
+ memory of the last tick — a restarted or compacted session cannot send a second
1325
+ daily digest by forgetting the first. A `per-tick` digest carries no daily key,
1326
+ so later ticks can hand off newly accumulated rows. Material reports carry no
1327
+ dedupe key either: two events in a day are two events.
1328
+
1329
+ ### What `status` shows
1330
+
1331
+ ```text
1332
+ reports 1 pending · 1 sending · 0 failed (delivery is at-least-once — a retry may duplicate)
1333
+ 9f2c1ab0d3e4 pending material 12m old attempt 2/6, retry in 1m (telegram sendMessage rejected: {"ok":false,…)
1334
+ 4b7c1ad9e001 SENDING digest 3m old attempt 1, outcome unknown — the process that sent it never said; a daemon start retries it and the message will say it may be a repeat
1335
+ ```
1336
+
1337
+ `pending` and `sending` are printed differently because they ask different things
1338
+ of you, and every row carries its age — "1 report pending since 08:15Z" is the
1339
+ signal that was missing when the reports went nowhere. Delivered reports leave
1340
+ the block: it is a list of what you are still owed, not a log.
1341
+
1342
+ Delivery keeps running while the fleet is **paused**. Pause stops claiming, not
1343
+ your right to hear about work that already happened. It runs on its own
1344
+ thirty-second timer rather than the five-minute dispatch tick, so a report does
1345
+ not sit in the outbox for the length of a poll interval.
1346
+
1347
+ The tier-2 escalation ledger (`notifications`) is untouched by all of this. It is
1348
+ a bare dedupe key by design — its primary key *is* the key — which is exactly why
1349
+ reports needed a separate table rather than an extension of that one.
1350
+
1351
+ ## The decision ledger (#136)
1352
+
1353
+ The outbox above fixed reports the orchestrator sends. This fixes the ones it is
1354
+ **waiting on**. A question put to you — an amendment, a tier-2 decision, "do I
1355
+ ship this tonight?" — lived in exactly one place: the model's context. A
1356
+ compaction, a restart, or a tick that ran long lost the question *and* the fact
1357
+ that one was owed, after which the session either asked again (you answer twice)
1358
+ or dropped it silently (the decision never lands, and nothing anywhere says one
1359
+ is outstanding).
1360
+
1361
+ So questions are written down, and every tick's prompt carries what is still
1362
+ open — read from the store, never from what the session remembers asking:
1363
+
1364
+ ```bash
1365
+ omp-conductor decision open --question "ship 0.4.3 tonight?" \
1366
+ --blocks "the release" --resolves-when issue-closed:132
1367
+ omp-conductor decision list
1368
+ omp-conductor decision resolve <id> --answer "yes, after #132 lands"
1369
+ omp-conductor decision withdraw <id> --reason "the release slipped a week"
1370
+ ```
1371
+
1372
+ **`--resolves-when` is the part that makes a parked question wake up.** Six
1373
+ conditions, each one something this package can check without asking you:
1374
+
1375
+ | Condition | Met when |
1376
+ | --- | --- |
1377
+ | `pr-merged:<https url>` | `gh` reports that pull request merged. |
1378
+ | `pr-checks-green:<https url>` | Every check on that pull request has a green verdict (a non-empty list, all `success`/`neutral`); a failing or still-pending check is not met. |
1379
+ | `pr-mergeable:<https url>` | The pull request is mergeable (`clean`, not `unknown` or conflicting). |
1380
+ | `issue-closed:<number>` | That issue is closed on the tracker. |
1381
+ | `npm-version:<pkg>@<version>` | `npm view <pkg>@<version> version` succeeds — the version is published. |
1382
+ | `rate-limit-reset:github` | GraphQL quota on `github` has any remaining capacity again. |
1383
+
1384
+ The daemon evaluates them beside each tick, fire-and-forget: a hanging registry
1385
+ costs one unevaluated condition, never the tick. A row that transitions
1386
+ false→true also writes the same `.conductor-tick-requested` poke recover uses,
1387
+ so the orchestrator heartbeat fires promptly (mid-interval poll, still gated by
1388
+ arm/channel/pending single-flight) instead of waiting a full interval. The poke
1389
+ reason and the digest flag `[CONDITION MET — act on this now]` both surface the
1390
+ wake so the session acts when the answer becomes actionable. Repeated sweeps
1391
+ while the condition stays true do nothing further — the store marks the
1392
+ transition once. After a green-but-behind PR is updated through
1393
+ `conductor_pr_update_branch`, open a fresh `pr-checks-green` watch on the new
1394
+ head so the next green transition can wake merge review the same way; nothing
1395
+ here merges on its own.
1396
+
1397
+ Anything else exits `2` and lists the six forms. An unparseable condition on an
1398
+ existing row is *listed and never treated as met*: a grammar a future release
1399
+ adds must not make an old row unloadable, and a question must never be hidden by
1400
+ a condition nobody can check.
1401
+
1402
+ **Expiry is enforced, not remembered.** An unanswered question closes itself
1403
+ after seven days — the deadline the floor's parked-amendment protocol already
1404
+ promised — so the digest stays a list of live questions instead of a graveyard.
1405
+ Answering or withdrawing is explicit, and a second resolution of the same id is
1406
+ refused rather than overwriting the first answer.
1407
+
1408
+ `omp-conductor status` carries one row, `decisions`, reported whether or not
1409
+ anything is open: `decisions 2 open (oldest 26h)`, or `decisions none open`. A
1410
+ row that appeared only when something was outstanding would leave "it forgot to
1411
+ record the question" and "there genuinely is none" looking identical, which is
1412
+ the ambiguity this table exists to remove.
1413
+
1414
+ ## Failure classes and recovery by class (#132)
1415
+
1416
+ Every run that did not reach a merged PR used to end at a human. The
1417
+ orchestrator re-derived the same triage on each tick — read the row, read the
1418
+ PR's checks, decide whether to requeue, re-run, settle or escalate — and then
1419
+ threw the conclusion away. Measured on this fleet's own history: **half the
1420
+ spend produced no merged PR**, and a large share of it was not implementation
1421
+ failure at all but daemon restarts, cancelled runners and a base branch moving
1422
+ under a green PR.
1423
+
1424
+ So the daemon classifies each terminal non-success before the next dispatch,
1425
+ persists the class on the row, and performs the one recovery that class names.
1426
+
1427
+ | Class | Signals | Recovery | Budget |
1428
+ | --- | --- | --- | --- |
1429
+ | `env-start-failure` | turn 0 plus an explicit harness start error (`No model selected`, a rejected key) | escalate — the session never read the issue | none |
1430
+ | `settlement-stuck` | a row carrying a PR that has since merged | settle: release the label, mark the row merged | none |
1431
+ | `returned-for-revision` | a `pushed-green` or `pushed-pending` PR was closed without merging | none — preserve the review decision for a human re-queue | continuation |
1432
+ | `merge-conflict` | `pushed-green`, PR open, GitHub reports conflicting | requeue for a rebase continuation | continuation |
1433
+ | `question` | the worker stopped to ask something (`blocked`) | escalate, carrying the worker's own report as evidence | none |
1434
+ | `orphan-dirty` | orphaned with a failed salvage and no operator ack | hold — recorded only; the tree is the only copy | none |
1435
+ | `orphan-clean` | orphaned with nothing uncommitted | requeue | continuation |
1436
+ | `turn-cap-progress` | at the turn ceiling **with** a PR, head or salvage commit | continue from the branch | continuation |
1437
+ | `turn-cap-spinning` | at the ceiling with no PR and no commits | escalate with the last tool calls the transcript recorded — and the completion path deliberately does **not** requeue it | none |
1438
+ | `admin-kill` | killed *below* its own ceiling — a restart or a drain | requeue | none |
1439
+ | `ci-infra` | PR open, every unresolved check cancelled / timed out / stale | re-run the failed jobs | none |
1440
+ | `ci-deterministic` | PR open, a check genuinely reports `FAILURE` | escalate with the failing check names and links | failed attempt |
1441
+ | `dispatch-infra` | the conductor's own Git path failed before the worker's first turn | requeue, bounded by per-class strikes | none |
1442
+ | `provider-credit` | the provider refused the run for credit (HTTP 402, or its own out-of-credit text read off the transcript) | pause the fleet and require `omp-conductor resume` once the provider has credit | none |
1443
+ | `provider-transient` | the provider aborted a request stream before the run produced a verdict | requeue, bounded by per-class strikes | none |
1444
+ | `unknown` | anything unrecognised | escalate | as recorded |
1445
+
1446
+ **Unknown escalates; it never silently retries.** A shape this table does not
1447
+ recognise is a gap in the table, and a quiet requeue would spend a budget on a
1448
+ cause nobody has named — the behaviour this exists to end.
1449
+
1450
+ ### The budgets follow the cause
1451
+
1452
+ `failuresFor` (implementation attempts) excludes `ci-infra`, `settlement-stuck`,
1453
+ `env-start-failure`, `dispatch-infra`, `provider-credit`, `provider-transient`
1454
+ and `returned-for-revision`. `continuationsFor` excludes `admin-kill`,
1455
+ `settlement-stuck`, `env-start-failure`, `dispatch-infra`, `provider-credit`
1456
+ and `provider-transient`, but explicitly counts a failed
1457
+ `returned-for-revision` row. Environment, dispatch and provider faults charge
1458
+ neither budget because the issue did not receive a valid implementation
1459
+ attempt. A merge conflict and a returned review both charge a continuation:
1460
+ each asks for more work, but neither is a failed implementation attempt.
1461
+
1462
+ An **unclassified** row (every row written before 0.4.3) counts exactly as it
1463
+ did before classification existed. Upgrading therefore changes no existing
1464
+ budget: the columns are additive and nullable, and a pre-0.4.3 `conductor.db`
1465
+ opens unchanged.
1466
+
1467
+ ### Stale labels are reconciled
1468
+
1469
+ On 2026-08-09 four issues carried `agent:failed` while every one of them was
1470
+ already complete — residue of a turns-cap kill two days earlier that nothing in
1471
+ the loop ever revisited. The board counted four phantom failures while the
1472
+ genuinely stuck issues were invisible.
1473
+
1474
+ Each tick now reconciles the three state labels against the tracker:
1475
+
1476
+ - A **closed** issue never keeps an `agent:*` label.
1477
+ - An **open** issue carrying `failed` whose sub-issues have *all* closed loses
1478
+ the label and gets one comment naming them, deduplicated through the same
1479
+ notifications ledger escalations use.
1480
+
1481
+ Positive evidence only: a tracker that cannot list answers empty, and an empty
1482
+ answer removes nothing — the label is the interlock that keeps two workers off
1483
+ one issue.
1484
+
1485
+ ### Where you see it
1486
+
1487
+ - `omp-conductor status` grows a `failure classes (unrecovered)` block, counting
1488
+ only rows whose recovery has *not* run. Classes rather than row states,
1489
+ because a row state is not an issue state.
1490
+ - The board appends `[<class>]` to a card whose newest run carries one.
1491
+ - The tick prompt carries one line — `Auto-recovered since last tick: 3
1492
+ (merge-conflict #365, admin-kill #82, …) — already handled, do not re-triage
1493
+ these.` — so the orchestrator stops writing that paragraph by re-deriving it.
1494
+
1495
+ ## Configuration
1496
+
1497
+ The config lives at `$OMP_CONDUCTOR_HOME/config.json`, or
1498
+ `~/.omp/conductor/config.json` when that variable is unset. It is written with mode
1499
+ `0600` in a directory created `0700`, because it carries chat ids and clone URLs.
1500
+ That same directory holds the SQLite store (`conductor.db`), the `paused` sentinel,
1501
+ the `sessions/` worker transcripts, the `orchestrator/` session directory,
1502
+ `backups/briefs/` for timestamped brief and policy safety copies, and
1503
+ `release-policy-blocks.jsonl`, the append-only audit of mechanically rejected
1504
+ release/deploy calls.
1505
+
1506
+ Runtime state lives elsewhere, under `$OMP_CONDUCTOR_RUNTIME_DIR` (default
1507
+ `~/.omp/run/daemons/omp-conductor`): `daemon.json`, a mode-`0600` pidfile written
1508
+ atomically, and `daemon.log`, appended across every boot so the previous failure is
1509
+ still there when you go looking. It is kept apart from the config directory because
1510
+ it is meaningless after a reboot, and the pidfile's liveness is probed on every
1511
+ read — a stale one never blocks a `start`. Both `start` and a bare `daemon` write
1512
+ the pidfile, so a daemon run in the foreground under systemd is as visible to
1513
+ `status` as a backgrounded one; `daemon --once` writes nothing, because that drill
1514
+ is exactly what the orphan-reconciliation guard reads the pidfile to protect.
1515
+
1516
+ The file is validated on every read. A malformed config produces one readable error
1517
+ listing every fault, and the daemon refuses to start rather than running with half
1518
+ a project.
1519
+
1520
+ The same vocabulary the loader enforces ships as a JSON Schema at
1521
+ `schema/config.schema.json` in the installed package (draft 2020-12). Anything
1522
+ `saveConfig` writes carries a top-level `"$schema"` reference to that installed
1523
+ copy (resolved from the package's own location, so it points at a real file),
1524
+ which lets an editor that understands JSON Schema validate a hand-edited config
1525
+ as you type; a config without the key is just as valid. Regenerate the shipped
1526
+ schema from `ConfigSchema` (`src/config-schema.ts`) with:
1527
+
1528
+ ```sh
1529
+ bun run schema
1530
+ ```
1531
+
1532
+ and commit the resulting `schema/config.schema.json`. CI's `bun test` fails if the
1533
+ checked-in schema drifts from what the code renders, so you cannot forget the step.
1534
+
1535
+ `omp-conductor setup` is the only thing here that writes this file, and on a project
1536
+ it already knows it can rewrite one area of it without re-asking the rest — see
1537
+ [Changing one setting](#changing-one-setting).
1538
+
1539
+ `version` is `2`. A `version: 1` file still loads: caps it names that this build no
1540
+ longer enforces are dropped rather than treated as typos, and the next save writes
1541
+ it back as `2`. In a `version: 2` file an unrecognised cap key **is** an error,
1542
+ because there is nothing left to retire — a mistyped `dailySpendUSD` would
1543
+ otherwise read as configured while the real ceiling stayed the default.
1544
+
1545
+ A complete, valid config for one project with two target repos:
1546
+
1547
+ ```json
1548
+ {
1549
+ "version": 2,
1550
+ "defaults": {
1551
+ "maxConcurrentWorkers": 2,
1552
+ "dailySpendUsd": 25,
1553
+ "planUsage": { "windowId": "anthropic:7d", "maxUsedFraction": 0.85 },
1554
+ "workerMaxTurns": 120,
1555
+ "workerWallClockMs": 5400000,
1556
+ "maxAttemptsPerIssue": 2,
1557
+ "maxContinuationsPerIssue": 2
1558
+ },
1559
+ "projects": [
1560
+ {
1561
+ "name": "demo",
1562
+ "tracker": { "kind": "github", "repo": "acme/planning" },
1563
+ "queueLabel": "ready-for-agent",
1564
+ "stateLabels": {
1565
+ "inProgress": "agent:in-progress",
1566
+ "blocked": "agent:blocked",
1567
+ "failed": "agent:failed"
1568
+ },
1569
+ "routing": {
1570
+ "labelPrefix": "repo:",
1571
+ "repos": {
1572
+ "api": {
1573
+ "name": "api",
1574
+ "cloneUrl": "git@github.com:acme/api.git",
1575
+ "defaultBranch": "main",
1576
+ "gates": [
1577
+ { "cmd": "bun run lint", "cwd": "." },
1578
+ { "cmd": "bun test", "cwd": "." }
1579
+ ],
1580
+ "graphProject": "~/.cache/conductor-graph/acme/api",
1581
+ "migrations": { "dir": "backend/alembic/versions" },
1582
+ "release": { "versionFile": "omp/package.json" }
1583
+ },
1584
+ "worker": {
1585
+ "name": "worker",
1586
+ "cloneUrl": "git@github.com:acme/worker.git",
1587
+ "defaultBranch": "main",
1588
+ "gates": [
1589
+ { "cmd": "ruff check .", "cwd": "." },
1590
+ { "cmd": "pytest -q", "cwd": "backend" }
1591
+ ]
1592
+ }
1593
+ }
1594
+ },
1595
+ "caps": {
1596
+ "maxConcurrentWorkers": 1,
1597
+ "dailySpendUsd": 15
1598
+ },
1599
+ "workerModel": "smol",
1600
+ "escalation": {
1601
+ "telegramChatId": "123456789",
1602
+ "telegramTopicId": 8713,
1603
+ "fallbackToIssueComment": true,
1604
+ "orchestrator": "embedded"
1605
+ },
1606
+ "authority": {
1607
+ "merge": "human",
1608
+ "release": "human"
1609
+ },
1610
+ "releasePolicy": {
1611
+ "version-bump-pr": "human",
1612
+ "git-tag": "human",
1613
+ "git-push-tags": "human",
1614
+ "package-publish": "human",
1615
+ "github-release": "human",
1616
+ "deploy": "human"
1617
+ },
1618
+ "policy": {
1619
+ "merge": {
1620
+ "requiredChecks": ["build", "lint"],
1621
+ "baseFreshness": "up-to-date",
1622
+ "drafts": "block",
1623
+ "whenBehindBase": "update-branch"
1624
+ },
1625
+ "release": {
1626
+ "requires": ["runs-settled", "no-open-prs"],
1627
+ "requiredChecks": ["release"],
1628
+ "artefacts": ["@acme/sdk"],
1629
+ "environments": ["staging"]
1630
+ }
1631
+ },
1632
+ "recoveryMerges": [
1633
+ {
1634
+ "prUrl": "https://github.com/acme/api/pull/381",
1635
+ "headSha": "9a783d8f17071d63f2d5d764d43a29837c365920",
1636
+ "reason": "operator-instructed"
1637
+ }
1638
+ ],
1639
+ "reporting": {
1640
+ "scope": "material"
1641
+ },
1642
+ "workspaceRoot": "~/.omp/conductor/worktrees",
1643
+ "mirrorRoot": "~/.omp/conductor/mirrors"
1644
+ }
1645
+ ]
1646
+ }
1647
+ ```
1648
+
1649
+ Field notes:
1650
+
1651
+ | Field | Notes |
1652
+ | --- | --- |
1653
+ | `version` | Must be `2`. A `version: 1` file still loads, drops the caps this build no longer enforces, and is rewritten as `2` on the next save. Present from day one so a format change can be migrated instead of silently misread. |
1654
+ | `defaults` | Every `Caps` field. Anything omitted falls back to the built-in default. |
1655
+ | `tracker.repo` | `owner/repo`. `tracker.kind` may be omitted; `"github"` is the only accepted value. |
1656
+ | `queueLabel` | The one label meaning "a human has signed this off as agent-ready". Matched exactly, case-sensitively. |
1657
+ | `release.versionFile` | Optional, per repo: a repo-relative JSON file with a top-level string `version`, such as `omp/package.json`. Declares that tags must match the version already landed on the live default branch. A delegated `git-tag` for such a repo requires delegated `version-bump-pr` too; otherwise config loading fails with the missing preparation path instead of granting an impossible release. Absolute paths and `..` are refused. |
1658
+ | `groomBelow` | Optional; default `4`. Routable candidates below this count make the orchestrator's tick prompt say the queue is running low and to groom it (Duty 2). An integer ≥ 1; anything else degrades to the default. |
1659
+ | `stateLabels` | Optional; defaults to `agent:in-progress`, `agent:blocked`, `agent:failed`. |
1660
+ | `routing.labelPrefix` | Optional; defaults to `repo:`. |
1661
+ | `routing.repos` | At least one entry, or nothing can be routed. `name` defaults to the map key, `defaultBranch` to `main`. |
1662
+ | `gates` | The exact cheap commands CI also runs, each with the `cwd` it runs from (`cwd` defaults to `.`). Running the real gate locally is what makes an unattended push safe — a subset lets an error outside the source dir reach the runners. |
1663
+ | `graphProject` | Optional, per repo. Absolute path of the **index-only clone** whose code graph this repo's workers query — conductor's own disposable clone, pinned to the repo's default branch, never a checkout you work in and never a worker's worktree. Written by the wizard; `~` is expanded, and a relative path is an error rather than something resolved against whichever cwd happened to read the file. Absent means this repo has no graph and its briefs say nothing about one. See [Code-graph discovery](#code-graph-discovery). |
1664
+ | `migrations` | Optional, per repo: `{ "dir": "backend/alembic/versions" }`. Names the repo-relative directory of an Alembic-style ordered migration chain (`revision` / `down_revision` in `*.py`). When set, `conductor_pr_merge` **refuses** a merge that would corrupt the chain at the base tip: reusing a revision id another file already declares, deleting a published migration, or a merge that would leave the combined graph with more than one head (so a stale parent is refused, and a fork-repair merge migration that unifies the heads passes). Absent means the repo opts out of the chain check entirely. Repo-relative only: a leading `/` or `..` is an error. |
1665
+ | `caps` | Per-project overrides; omit it or pin only the fields you want to change. |
1666
+ | `escalation.fallbackToIssueComment` | Defaults to `true`. Absent means "yes, still tell me". |
1667
+ | `escalation.telegramTopicId` | Optional forum topic for everything conductor sends: tier-2 pages, reports, digests, arm challenges, `omp-conductor message`. Setup offers the topics omp-telegram has claimed, naming each one's herdr space. The bridge re-claims a pane's topic across restarts — including the restarts `upgrade` and `restart` perform — so a pinned id that is no longer claimed is replaced at send time by the live claim whose **herdr space** is this project, falling back to one titled for the project, logged without ids (#407, #412). The space is read first because the bridge titles a topic `ownAgentName ?? basename(cwd)`, and a multi-project host whose panes sit under one state directory gives every claim the same title. An identity two claims share is treated as no match at all rather than a guess. A pin that is still claimed always wins, so a deliberately separate topic is never hijacked. Absent keeps flat-chat behaviour. |
1668
+ | `escalation.orchestrator` | Optional; `"embedded"` (default) or `"external"`. `external` means an orchestrator session already runs elsewhere: the daemon starts none, and tier-1 escalations post as issue comments for that session to drain. Any other value is an error. |
1669
+ | `authority` | Optional; `{ "merge": …, "release": … }`, each `"human"` (default) or `"orchestrator"`. It grants nothing to the daemon — it words the orchestrator's standing orders and the Releases paragraph of the rendered brief, so the config and the prompt cannot disagree about who holds the merge button. Unknown keys and any other value are errors, never folded to the default. |
1670
+ | `releasePolicy` | Optional; a per-shape map whose values are `"human"` (default) or `"orchestrator"`. Shapes are `version-bump-pr`, `git-tag`, `git-push-tags`, `package-publish`, `github-release`, and `deploy`. The legacy `"none"` denies every shape; legacy `"operator-brief"` grants the artifact-producing shapes, including reviewed version preparation, but keeps deploy human-owned. The in-session tripwire blocks recognised raw release/deploy calls before execution. Every rejection is written to `release-policy-blocks.jsonl`; the heartbeat carries that day's count into the daily digest. This is the mechanical gate; `authority.release` still says who owns the decision. |
1671
+ | `recoveryMerges` | Optional, hand-edited recovery authority for a PR that has no conductor run record. Each entry is an exact `{ prUrl, headSha, reason: "operator-instructed" }` tuple. When all three values match, `conductor_pr_merge` may merge that one PR even while the fleet is held and even when standing merge authority is `"human"`. It still requires a routed project repo, an open PR at that exact live head, green checks, the migration-chain guard, and the single-flight lock—the same safety path as an ordinary merge. Duplicate PR URLs and malformed values make config loading fail closed. Setup preserves entries but never creates them. Remove an entry after the recovery is complete. |
1672
+ | `reporting` | Optional; a **legacy scope preset** (`reporting.scope` — `"material"` default, `"decisions"`, `"escalations"`) or the **explicit form** `{ "interruptOn": [...], "digest": { ... }, "availability": { ... } }`. The preset writes which categories may page the operator (`interruptOn`) and when the rollup happens (`digest.cadence`); the explicit form sets both directly and may add a weekly operator-availability window. The two forms are mutually exclusive in one config. See [Reporting policy](#reporting-policy-reporting). |
1673
+ | `orchestratorReadPaths` | **Retired in 0.4.3.** Still accepted in a config and ignored, so a fleet carrying it upgrades without an edit. It widened the orchestrator's file-tool allowlist; there is no allowlist any more — the orchestrator is [unconfined by design](#the-orchestrator-is-unconfined-deliberately). |
1674
+ | `policy` | Optional; the gating conditions a merge or a release must satisfy, in two sections — `policy.merge` and `policy.release`. Any member may be omitted and the loader fills it from the strict default; an unknown key in either section, or a value outside its vocabulary, is an error naming the field, never a silent downgrade. See [Merge and release preconditions](#merge-and-release-preconditions-policy). |
1675
+ | `workspaceRoot` / `mirrorRoot` | Optional; default to `worktrees/` and `mirrors/` under the state directory. `~` is expanded. |
1676
+
1677
+ Prefer an SSH `cloneUrl`, or an https URL backed by a credential helper. A clone URL
1678
+ with credentials embedded is persisted into the mirror's git config, exactly as it
1679
+ would be for a hand-run clone.
1680
+
1681
+ ### Reporting policy (`reporting`)
1682
+
1683
+ What may interrupt the operator's phone, and when the daily rollup happens. Two
1684
+ spellings, mutually exclusive in one config (the loader rejects a `scope` next to
1685
+ `interruptOn`/`digest`):
1686
+
1687
+ - **Preset** — `reporting.scope`, the three legacy values, mapped verbatim:
1688
+ - `material` (default) → `interruptOn: [tier2, decision-needed, fleet-stopped, confirmed-failure, material]`, digest `per-tick`.
1689
+ - `decisions` → `interruptOn: [tier2, decision-needed, fleet-stopped]`, digest `per-tick`.
1690
+ - `escalations` → `interruptOn: [tier2, fleet-stopped]`, digest `daily` (model-timed).
1691
+ - **Explicit** — `reporting: { "interruptOn": ["tier2", "fleet-stopped", ...], "digest": { "cadence": "none" | "per-tick" | "daily" } }`.
1692
+ `interruptOn` must be a non-empty array of known categories (`tier2`, `decision-needed`, `fleet-stopped`, `confirmed-failure`, `material`), each an escalation's tier-2 category. `daily` may add `at` (`HH:MM`, 24h) and `timezone` (a known IANA zone, defaulting to the host zone) — both only valid with `daily`.
1693
+
1694
+ The explicit form may add a weekly local-time window:
1695
+
1696
+ ```json
1697
+ {
1698
+ "reporting": {
1699
+ "interruptOn": ["tier2", "fleet-stopped"],
1700
+ "digest": { "cadence": "daily", "at": "17:00", "timezone": "Europe/London" },
1701
+ "availability": {
1702
+ "timezone": "Europe/London",
1703
+ "days": ["mon", "tue", "wed", "thu", "fri"],
1704
+ "start": "09:00",
1705
+ "end": "17:00",
1706
+ "bypass": ["fleet-stopped"]
1707
+ }
1708
+ }
1709
+ }
1710
+ ```
1711
+
1712
+ `timezone` must be a known IANA zone. For a daily digest, its timezone defaults
1713
+ to this value and must match it when both are set.
1714
+
1715
+ `days` is a non-empty set of `mon` through `sun`; `start` is inclusive and
1716
+ `end` is exclusive. A start later than the end defines an overnight window on
1717
+ the day it opens. `bypass` is an explicit list of known interrupt categories
1718
+ that may still page outside the window; it may be empty. For ordinary notices,
1719
+ a bypass has no effect on a category omitted from `interruptOn`. Urgent recovery
1720
+ notices may bypass category batching when the digest loop itself is unavailable,
1721
+ but they still require the configured availability bypass outside the window.
1722
+
1723
+ The setup wizard offers this as **Weekly availability window** and asks for the
1724
+ zone, days, start/end, bypass categories, and digest schedule: every tick,
1725
+ model-timed daily, disabled, or a fixed daily `HH:MM`. Re-running setup or
1726
+ amending reporting preselects and preserves the configured `none`, `per-tick`,
1727
+ or `daily` cadence. Choosing **Continuous (24-hour interrupts)** is the explicit
1728
+ opt-out and preserves the behavior of every existing config; an absent
1729
+ `availability` key also means continuous operation.
1730
+
1731
+ Outside the window, an otherwise interruptible escalation is stored durably
1732
+ instead of sent. A daily digest may consume it first. Otherwise the daemon
1733
+ atomically queues one working-hours catch-up report when the window opens,
1734
+ including after downtime; associating the held rows before delivery prevents a
1735
+ later tick from authoring a duplicate. Each heartbeat prompt names the
1736
+ mechanically computed current mode and next transition. `status` shows the same
1737
+ state plus the next digest opportunity (`due now`, every tick, disabled, or its
1738
+ next operator-local timestamp). Config, escalation routing, and report transport
1739
+ are re-read at tick or send time, so changing the window or Telegram target
1740
+ does not require a daemon restart.
1741
+
1742
+ Attachment-bearing autonomous Telegram sends cannot be replayed by the text
1743
+ digest, so they are blocked with an explicit “nothing sent or held” error rather
1744
+ than silently dropping their files.
1745
+
1746
+ A tier-2 escalation whose category is **not** in `interruptOn` is not dropped: it
1747
+ is held (`held_notices`) and the next accepted digest is its delivery authority.
1748
+ A `daily` digest is at-most-once per local day (`digest:<YYYY-MM-DD>` in the
1749
+ configured zone), which remains the delivery authority across restarts.
1750
+ `per-tick` digests are not daily-deduplicated, so a later tick can claim newly
1751
+ accumulated rows. A scheduled `daily` digest is only sent on a day it has not
1752
+ already run, once the local clock has passed `at`; a restart after `at` still
1753
+ sends today's (one catch-up), and a fully missed day is skipped, never sent late.
1754
+
1755
+ **What the scope does:** the [orchestrator heartbeat](#orchestrator-tick) appends
1756
+ the current constraint to every tick it sends, so the reporting contract arrives
1757
+ with the prompt instead of only in a brief the session read hours ago. Explicit
1758
+ policies name their actual interrupt categories and digest cadence. The policy is
1759
+ re-read from `~/.omp/conductor/config.json` on **every** tick, and escalation,
1760
+ direct Telegram, and durable report paths apply it mechanically. Turning the
1761
+ volume up or down — `omp-conductor setup` again, or an edit to the file — therefore
1762
+ binds the next tick without restarting the session.
1763
+
1764
+ No config, an unreadable or invalid config, or several unnamed projects fall
1765
+ back to the legacy `material` scope for the heartbeat and log the reason once.
1766
+ An invalid live availability policy blocks autonomous Telegram fail-closed; it
1767
+ does not guess that the operator is awake.
1768
+
1769
+ Changing the key later does not rewrite an `ORCHESTRATOR.md` you already have.
1770
+ The generated `POLICY.md` describes every scope without pinning the current
1771
+ choice; the tick constraint remains derived from live config. Keep any
1772
+ operator-owned reporting additions in `ORCHESTRATOR.md` consistent with it.
1773
+
1774
+ ### Merge and release preconditions (`policy`)
1775
+
1776
+ These used to be sentences in your `POLICY.md`: when a PR may be merged, what
1777
+ must be green, what a release requires. Prose cannot be checked, so every tick
1778
+ re-decided them by reading and interpreting them again. They are configuration
1779
+ now, `POLICY.md` keeps only judgement, and the rendered brief *describes* the
1780
+ policy instead of restating it — no threshold lives in two places.
1781
+
1782
+ `policy.merge`:
1783
+
1784
+ | Field | Values | Default | Means |
1785
+ | --- | --- | --- | --- |
1786
+ | `requiredChecks` | any check names | `[]` | Checks that must have concluded successfully. **Empty is the strict answer** — it means every check the PR reports, not "no checks". |
1787
+ | `baseFreshness` | `up-to-date`, `any` | `up-to-date` | Whether the head must be level with the base branch. `any` accepts a verdict produced against an older base. |
1788
+ | `drafts` | `block`, `allow` | `block` | Whether a draft PR can be merged at all. |
1789
+ | `whenBehindBase` | `update-branch`, `hold`, `escalate` | `update-branch` | What to do with a green PR that fell behind. `update-branch` runs `gh pr update-branch` and waits for the fresh run. Closing it and an admin bypass are not spellable. |
1790
+
1791
+ `policy.release`:
1792
+
1793
+ | Field | Values | Default | Means |
1794
+ | --- | --- | --- | --- |
1795
+ | `requires` | `runs-settled`, `no-open-prs`, `queue-drained`, `base-branch-green`, `epic-children-closed` | `["runs-settled"]` | What must already have landed. `runs-settled` reads each active run's PR fact at release time: a pushed run whose PR has merged counts as settled even when the settle sweep has not yet written the terminal row — so a hold-drained release does not wait an extra tick the operator reached the gate by holding. Live workers and unmerged/unknown PRs still refuse, and the message names which is which. `base-branch-green` requires the current live head's push-triggered workflow verdict for that routed repository to be green; pending, unknown, red, or no observation refuses release. Order and duplicates do not matter; the loader canonicalises. |
1796
+ | `requiredChecks` | any check names | `[]` | Checks that must be green on the branch being released. Empty means every check it reports. |
1797
+ | `artefacts` | any names | `[]` | The packages or images this project releases. **Empty denies**: nothing has been authorised to ship. |
1798
+ | `environments` | any names | `[]` | Deploy targets. **Empty denies** every environment. |
1799
+
1800
+ A project with no `policy` block loads as the whole default above, which is the
1801
+ strictest reading of the prose it replaced. `omp-conductor setup` asks for all of
1802
+ it under the **merge & release preconditions** area, so changing one condition
1803
+ costs eight prompts rather than a hand-edit — see
1804
+ [Changing one setting](#changing-one-setting).
1805
+
1806
+ This key grants nothing. Who *may* merge or release is
1807
+ [`authority`](#configuration), and which release tool calls are mechanically
1808
+ permitted is [`releasePolicy`](#configuration). `policy` says what must be true
1809
+ before the act, whoever is doing it.
1810
+
1811
+ #### Reasons are a closed vocabulary
1812
+
1813
+ Where an automated verb takes a `reason`, the argument is one value out of a
1814
+ fixed set, not free text — a reason a rule matches on is a reason that decides,
1815
+ and a decision made out of a model's own wording is one no two runs spell the
1816
+ same way. A reason outside its set is refused, and the refusal names every
1817
+ accepted value.
1818
+
1819
+ | Verb | Accepted reasons |
1820
+ | --- | --- |
1821
+ | merge | `preconditions-met`, `behind-base-refreshed`, `operator-instructed`, `release-blocking` |
1822
+ | release | `batch-complete`, `epic-closed`, `hotfix`, `operator-instructed` |
1823
+ | label change | `promoted-to-queue`, `re-briefed`, `needs-human`, `duplicate`, `superseded`, `out-of-scope` |
1824
+
1825
+ Free-form rationale still has a home: it rides alongside as a separate
1826
+ `rationale` field, is written into the audit trail verbatim, and is never
1827
+ parsed or matched by anything.
1828
+
1829
+ ## Orchestrator tick
1830
+
1831
+ The escalation path above assumes an orchestrator session that is actually
1832
+ running its loop. A 24/7 omp session with a standing brief and nobody typing into
1833
+ it never gets prompted, so it never runs anything. Installing
1834
+ `omp plugin install omp-conductor` also installs a heartbeat that prompts it.
1835
+
1836
+ The heartbeat is **inert unless the session cwd contains
1837
+ `.conductor-tick.json`**, so an ordinary session has no timer. `omp-conductor setup`
1838
+ writes this file for external orchestration. A manual configuration has this form:
1839
+
1840
+ ```json
1841
+ {
1842
+ "intervalSeconds": 900,
1843
+ "project": "fleet",
1844
+ "armedFile": "/home/fleet/.omp/conductor/armed-fleet",
1845
+ "accessFile": "/home/fleet/.omp/agent/telegram/access.json",
1846
+ "message": "Run your standing loop from ORCHESTRATOR.md now."
1847
+ }
1848
+ ```
1849
+
1850
+ | Key | Required | Default | Notes |
1851
+ | --- | --- | --- | --- |
1852
+ | `intervalSeconds` | yes | — | Whole seconds between ticks, minimum `60`. A tick costs a full turn of a frontier model, so a sub-minute period is refused rather than obeyed. |
1853
+ | `project` | no | the only configured project | Which conductor project this fleet session ticks for. `setup host` stamps it, one tick config per fleet cwd, and it is what lets a host with several configured projects resolve *this* fleet's brief, reporting policy, digest ledger and release grants. Omitting it is the pre-multi-project spelling: correct on a single-project host, and on a host with two or more it degrades every tick to the default reporting scope with no release grants — `status` and the tick log then name the one fix (`re-run omp-conductor setup host`). A name no configured project has degrades the same way. |
1854
+ | `budgetSeconds` | no | `600` | Seconds a turn may run before the tick guard refuses its remaining tool calls (#189), and before a queued operator message preempts them. An integer ≥ 60; anything else degrades to the default. |
1855
+ | `askTimeoutSeconds` | no | `300` | Seconds one `conductor_ask` call waits for the operator before its declared `on-timeout` outcome fires (#438). An integer 60–3600; anything else degrades to the default. The resolved ceiling is always capped at the turn budget, so an ask can never outlive the turn it runs in. |
1856
+ | `armedFile` | no | none — the gate passes | Path to the arm marker. A tick does nothing while the file is missing. **Re-read from disk on every tick**, so a `setup host` restamp onto `armed-<project>` binds on the next heartbeat instead of leaving a live pane watching the path it captured at session start. Relative paths resolve against the session cwd, so `state/armed` means `<cwd>/state/armed`. `setup host` writes `<state dir>/armed-<project>`, one marker per project, so `arm --project A` cannot arm B. A value it did not generate is left alone as your own choice. |
1857
+ | `accessFile` | no | none — the gate passes | Path to the Telegram bridge's `access.json`. Every tick re-reads it and requires `enabled: true` with exactly one entry in `allowFrom`. Relative paths resolve against the session cwd. **Configure this on any fleet deploy** — see below. |
1858
+ | `message` | no | `Tick <ISO timestamp>: re-read <workspaceRoot>/ORCHESTRATOR.md from disk, then run your standing loop from it.`, then the reporting-policy line, delivery rule, and mechanical availability state | When set, this text replaces the ordinary reporting-policy line and delivery rule, but the runtime-owned availability state is still appended: a custom prompt cannot infer whether the operator may be interrupted. Re-read from disk on **every** tick, so rewording it binds the next heartbeat instead of waiting for a session restart; a re-read that fails — caught mid-edit, removed, or invalid — keeps the value read at session start rather than stopping the heartbeat. `intervalSeconds` is *not* re-read: rescheduling a live timer still needs a restart. The default *orders* the session to re-read its brief, naming the path resolved from the project's `workspaceRoot`, because a standing prompt drifts out of a long-lived session's context while the file on disk does not. |
1859
+ | `agentName` | no | the project name, else `fleet` | The herdr agent name the orchestrator's pane is registered under. Under herdr this is the whole of the identity check below. `setup host` writes the project name, so two fleets in one herdr session are distinguishable; when no tick config names one, the fallback matches `AGENT_NAME=${AGENT_NAME:-fleet}` in the recovery plugin's `recover.sh`, so both halves key on one name. Rename the agent and set this to match. |
1860
+
1861
+ #### Bounded operator asks (#438)
1862
+
1863
+ On 2026-08-16 an unanswered `telegram_ask` blocked the orchestrator turn for
1864
+ six hours; every duty behind the question stopped with it. `telegram_ask` is
1865
+ omp-telegram's tool and it waits as long as the answer takes, so on a locally
1866
+ injected tick the extension refuses it and mounts its own bounded surface,
1867
+ `conductor_ask`:
1868
+
1869
+ - The call carries `question`, `on-timeout` (`auto-proceed` or `park`, required),
1870
+ and optionally `timeoutSeconds`, `blocks`, `recommended`, `options` and
1871
+ `category`.
1872
+ - The tool records a decision row first (durable, seven-day expiry), then
1873
+ delivers the question through the same path `omp-conductor message` uses —
1874
+ immediately when the reporting policy permits, durably held otherwise.
1875
+ - It waits at most the ceiling: `timeoutSeconds` if the ask names one, else the
1876
+ tick config's `askTimeoutSeconds`, else 300 seconds — always capped at the
1877
+ turn budget, so the ask can never outlive the turn it runs in. An ask issued
1878
+ without a timeout gets the default all the same.
1879
+ - On timeout, `auto-proceed` applies the recommended option and resolves the row
1880
+ with `"<option> (auto-applied on ask timeout)"` so the record never reads as a
1881
+ human choice; `park` leaves the row open and pending — re-surfaced in every
1882
+ tick prompt until answered or the seven-day expiry — and the caller takes the
1883
+ blocked work out of the claimable queue. A timeout is "nobody answered yet",
1884
+ never a cancelled/errored ask and never an operator no.
1885
+
1886
+ The refusal of the raw tool is mechanical (the tool-call gate), not a prompt
1887
+ reminder: a model cannot wait unbounded on a local tick even by omitting the
1888
+ timeout argument.
1889
+
1890
+ #### Upgrading from one shared arm marker
1891
+
1892
+ Before per-project markers every project was given the same `<state dir>/armed`,
1893
+ so arming one fleet armed all of them. `setup host` rewrites that value — and
1894
+ only that value — to `armed-<project>`. The old bare marker is honoured for one
1895
+ more cycle on a **single-project** host, so the upgrade never silently disarms a
1896
+ live fleet, and the next `arm` or `disarm` retires it. On a host with **two or
1897
+ more** projects it arms nothing: `status` reports `legacy global arm marker —
1898
+ re-run setup host, then arm per project`, and every tick stays disarmed until each
1899
+ project is armed on its own marker.
1900
+
1901
+ The same restamp renames the identity this pane ticks under: an `agentName` of
1902
+ `fleet` — the value every project used to be given — becomes the project name.
1903
+ **Under herdr that is an operator step, not a no-op.** Ownership is proved against
1904
+ the pane's registered herdr agent, so after re-running `setup host` the live fleet
1905
+ pane needs the new name.
1906
+
1907
+ **Rename the agent herdr already detects — do not `agent start`.** `herdr agent
1908
+ start` submits omp *into* the pane's existing shell and requires a pane sitting at
1909
+ a shell prompt with no agent on it; the live orchestrator pane is neither, so it
1910
+ is refused at best and starts a second omp in that pane at worst. `rename` touches
1911
+ no process and keeps the session as it is:
1912
+
1913
+ ```sh
1914
+ herdr --session <session> agent list # find the fleet's pane_id
1915
+ herdr --session <session> agent rename <pane-id> <project>
1916
+ ```
1917
+
1918
+ If the name cannot be reassigned in place, stop and resume rather than starting a
1919
+ second orchestrator — the same shape `recover.sh` uses, so the omp session is
1920
+ preserved rather than replaced:
1921
+
1922
+ ```sh
1923
+ herdr --session <session> agent get <pane-id> # note agent_session.value — the session ref
1924
+ # exit omp in that pane (/exit) so the pane is back at a shell prompt, then:
1925
+ herdr --session <session> agent start <project> --kind omp --pane <pane-id> -- --resume=<ref>
1926
+ ```
1927
+
1928
+ Until the pane carries the new name it declines to tick and logs which agent it
1929
+ actually is versus the one the tick config names, with the `rename` command in the
1930
+ line — the heartbeat fails closed and says so rather than letting two fleets both
1931
+ answer to `fleet`. Set `agentName` explicitly if you would rather keep the old
1932
+ name; a value that is not the shared default is never rewritten.
1933
+
1934
+ The marker itself needs no restart: `armedFile` is re-read from disk every tick, so
1935
+ the restamped path binds on the next heartbeat. Before 0.15.2 it was read once at
1936
+ session start, and a restamp under a live pane left that pane watching a path the
1937
+ restamp had just replaced — `status` reported `armed` from the file while the
1938
+ heartbeat skipped silently as "not armed", which writes no stall marker.
1939
+
1940
+ Recovery is fail-closed across that window. A restamped `agentName` moves the
1941
+ recovery plugin's own state files to per-agent paths that do not exist yet, and
1942
+ the live pane is still saved under `fleet`, so the snapshot offers no candidate
1943
+ for the new name. `herdr-conductor` treats the pre-rename identity and bootstrap
1944
+ marker as proof a fleet has already lived on this host whatever it is called now:
1945
+ it pages `no fleet identity to recover for agent <project>` instead of
1946
+ provisioning a second workspace beside the live orchestrator. Finish the rename
1947
+ and the next pass recovers normally.
1948
+
1949
+ A default tick sends one message (`customType` `omp-conductor.tick`, attributed
1950
+ to the user): the standing-loop prompt, the reporting-policy constraint re-read
1951
+ from conductor config on every tick, the delivery rule, and the mechanically
1952
+ computed operator-availability state. A configured `message` replaces the first
1953
+ three parts but not that clock state. The delivery rule is there
1954
+ because end-of-turn text reaches the operator's Telegram only on a turn that
1955
+ *began* as an inbound Telegram message: a tick is injected locally, so anything
1956
+ the session merely writes at the end of one is read by nobody, and a reportable
1957
+ event has to be delivered by an explicit `telegram_send` call the session
1958
+ watched succeed. The tick starts a turn if the session is idle; while a turn is
1959
+ streaming it is queued as a follow-up and consumed when that turn ends.
1960
+ It sends **nothing** when:
1961
+
1962
+ - `armedFile` is configured and missing;
1963
+ - `accessFile` is configured and the escalation channel is not verifiably up;
1964
+ - an earlier tick is still queued. Ticks coalesce rather than stack, so a slow
1965
+ turn cannot leave a backlog of heartbeats behind it — and two coalesced ticks
1966
+ in a row are the signal that the session is not slow but wedged, which is
1967
+ what the [stall marker](#a-wedged-session-and-the-marker-that-notices) is for.
1968
+
1969
+ ### One session per directory ticks, and it says which
1970
+
1971
+ Activation is a property of the *directory*, so before it arms anything the
1972
+ heartbeat asks whether this session is the orchestrator or merely a session
1973
+ standing in its directory. It has to: opening a second omp session in the fleet's
1974
+ cwd — a shell to read state, say — used to arm a second heartbeat that prompted
1975
+ *that* session with the standing loop, and with
1976
+ [`authority`](#configuration) delegated it would consider itself entitled to
1977
+ merge PRs and cut releases. Two brains, one queue, and nothing in the log to tell
1978
+ them apart.
1979
+
1980
+ **Under herdr** (`HERDR_ENV=1` with a `HERDR_PANE_ID`), the answer is the pane's
1981
+ registered agent name: the heartbeat asks `herdr agent list` for the entry whose
1982
+ `pane_id` is this pane's and ticks only when its `name` equals `agentName`. Fleetness
1983
+ is the *session* — every pane in it shares `HERDR_SESSION` and the cwd — and
1984
+ herdr's `agent` field is the *runtime*, `omp` for the orchestrator and for the
1985
+ shell beside it, so neither can tell them apart. The registered name can, it is
1986
+ what `herdr agent start fleet --kind omp --pane <id>` sets when the recovery plugin
1987
+ starts a fleet into an empty pane — and what `herdr agent rename <pane-id> fleet`
1988
+ sets on a pane whose agent herdr already detects, which is the only safe spelling
1989
+ while omp is running in it. It is the same identity the recovery plugin keys on. A
1990
+ pane with a different name, or no name at all, stays inert.
1991
+
1992
+ **Without herdr**, the session claims the directory in a sibling
1993
+ `.conductor-tick-owner.json` (pid, session file, claim time) and ticks only while
1994
+ it is the live claimant. Liveness is a **pid check, never a timestamp**: a crashed
1995
+ orchestrator's claim is reclaimed by the next session rather than wedging the
1996
+ fleet until somebody deletes a file, and a slow-but-running orchestrator never
1997
+ loses its claim to a lease that expired.
1998
+
1999
+ Declining is logged once, at session start, naming the holder — which is the whole
2000
+ point, because the original failure was that the second ticker was
2001
+ indistinguishable from the first:
2002
+
2003
+ ```text
2004
+ [omp-conductor] orchestrator tick inactive: pane w1:p1 (agent "fleet") owns the fleet tick here — this session will not tick
2005
+ [omp-conductor] orchestrator tick inactive: this pane is agent "scratch", not the fleet agent "fleet" — this session will not tick
2006
+ [omp-conductor] orchestrator tick inactive: pid 12345 (claimed 2026-01-02T03:04:05.000Z, session …/fleet.jsonl) owns the fleet tick in /home/conductor/.omp/conductor — this session will not tick
2007
+ ```
2008
+
2009
+ A `herdr agent list` that does not answer also declines, for the same reason the
2010
+ escalation channel fails closed: under herdr this session is one pane of several
2011
+ in that directory, and an unproven identity is exactly the case the check exists
2012
+ for. That includes `herdr` not being on the session's `PATH` — worth checking on a
2013
+ fleet host, where the orchestrator's environment comes from a unit file rather
2014
+ than a login shell — and `HERDR_BIN_PATH` names the binary when it is not, the
2015
+ same escape hatch the recovery plugin's `recover.sh` has. On a host with no herdr
2016
+ and no prior claimant — the ordinary single-session case — nothing changes.
2017
+
2018
+ ### The escalation channel is a gate, and it fails closed
2019
+
2020
+ Unattended dispatch is only defensible while a tier-2 escalation can reach a
2021
+ person. So `accessFile` is checked on **every** tick and never cached at session
2022
+ start: the bridge is reconfigured out-of-band, and a heartbeat that trusted a
2023
+ startup snapshot would keep dispatching for days after the channel went away. A
2024
+ stale arm marker must not outlive the channel that makes running unattended safe.
2025
+
2026
+ The check passes only when a bot token is resolvable — `TELEGRAM_BOT_TOKEN` in
2027
+ the environment, or in the `.env` beside `accessFile` — and the file parses to an
2028
+ object with `enabled: true` and exactly one `allowFrom` entry. Everything else
2029
+ stops the heartbeat: no token, so nothing outbound works at all; file missing,
2030
+ unreadable or truncated; not JSON, or JSON that is not an object; `enabled`
2031
+ absent or false; zero owners paired (nobody to page) or more than one (ambiguous:
2032
+ the conductor refuses to guess which human is on the hook). Failure modes are
2033
+ deliberately not distinguished in the decision: each one means a page lands
2034
+ nowhere. `omp-conductor status` is where they are told apart — its `telegram` row
2035
+ names the specific fault.
2036
+
2037
+ One caveat the file cannot express: omp-telegram binds its own copy of the token
2038
+ in `startBot()` at session start, and only when the bridge is switched on. It
2039
+ rebinds only on `/telegram token` or `/telegram on`. So writing a token into
2040
+ `.env` out-of-band — or flipping `enabled` to true by hand — restores tier-2
2041
+ paging immediately, because conductor sends those itself, while the bridge's own
2042
+ tools, `telegram_send` and `telegram_ask`, stay dead until you reload it. After
2043
+ either edit, run `/telegram on` in the orchestrator session. Until you do, ticks
2044
+ carry an explicit note that an amendment cannot be approved on this surface, and
2045
+ the conductor never assumes an answer it did not receive.
2046
+
2047
+ Leaving `accessFile` unset passes the gate, because an ordinary developer session
2048
+ that happens to have a `.conductor-tick.json` has no bridge to check. It is not an
2049
+ off switch for the check: **a fleet deploy always sets it.**
2050
+
2051
+ Every tick — sent or skipped — is logged with its reason (`not armed`,
2052
+ `escalation channel down`, `tick already pending`) to the omp log. `omp-conductor
2053
+ hold` is deliberately **not** one of the gates: hold stops the *dispatcher*
2054
+ claiming work, and the tick drives a different session — one whose duties
2055
+ (grooming the queue, draining escalations, reporting) are exactly what stays
2056
+ useful while dispatch is stopped. Its own off switch is the arm marker. Skips
2057
+ are deliberately silent in the UI: a disarmed fleet would otherwise raise a
2058
+ notification every interval, forever. The one exception is a malformed
2059
+ `.conductor-tick.json`,
2060
+ which notifies once at session start and leaves the heartbeat off; silent failure
2061
+ there is the failure mode the heartbeat exists to prevent. A conductor config that
2062
+ cannot supply a reporting scope logs `tick reporting scope: using material` once
2063
+ per session. The interval does not re-log it, because the file is unlikely to fix
2064
+ itself between two ticks.
2065
+
2066
+ ### A wedged session, and the marker that notices
2067
+
2068
+ Coalescing is also the only wedge detector this package has. On 2026-08-07 the
2069
+ dogfood fleet's orchestrator finished a turn, logged `ui.loop-blocked` right
2070
+ after an auto-compaction threshold decision, and never started another. The
2071
+ process stayed alive, so herdr's recovery — agent listed AND a non-shell
2072
+ foreground process — read healthy. The dispatch daemon is a separate process and
2073
+ kept working, so `/healthz` was green all night, while the one brain holding
2074
+ merge authority sat on a green PR it never merged. A tick injected two minutes
2075
+ into the wedge and an operator's Telegram message five minutes later both went
2076
+ unconsumed for 23 minutes, until a manual `SIGTERM`. The heartbeat logged `tick
2077
+ skipped: tick already pending` throughout, which is exactly what a merely slow
2078
+ turn looks like.
2079
+
2080
+ So the heartbeat counts them. Two consecutive coalesced ticks — a full hour at
2081
+ the reference 1800-second interval, generous by construction — mean the last
2082
+ prompt was never consumed, and the extension:
2083
+
2084
+ - writes `<session cwd>/.conductor-stalled`, one line of `<ISO timestamp>
2085
+ <diagnosis>`.
2086
+ - logs at **error** level: `orchestrator stalled: 2 ticks queued unconsumed —
2087
+ the agent loop is not draining; see .conductor-stalled`.
2088
+
2089
+ Both escapes deliberately leave the session, because a loop that cannot drain
2090
+ its queue cannot report on itself — that is the whole failure.
2091
+
2092
+ **The daemon reads it.** A marker nobody consumes is an artifact, not an alert,
2093
+ so the dispatch daemon checks it on its own five-minute tick — and *before* its
2094
+ pause check. The orchestrator is a different process and can be wedged while
2095
+ the fleet is deliberately paused, which is precisely the state the dogfood
2096
+ fleet was in when this happened. One tier-2 page per stall, keyed on the
2097
+ marker's own timestamp so a second wedge the same day is not swallowed as a
2098
+ repeat, re-armed when the marker clears, and latched only once the page is
2099
+ confirmed delivered — an escalation channel that fails on the one tick that
2100
+ noticed must not buy permanent silence.
2101
+
2102
+ It restarts nothing. A wedge lands mid-turn, and no other process can tell a
2103
+ half-applied edit from an idle loop; the operator attaches, looks, and decides.
2104
+
2105
+ **herdr-conductor deliberately does not read it**, though its liveness test
2106
+ (agent listed AND a non-shell foreground process) passes straight through a
2107
+ wedge. That plugin only runs on `startup`, `pane.exited` and
2108
+ `pane.agent_detected`, and a session that stays alive and stops working emits
2109
+ none of them — so the check could never fire during the wedge itself. What it
2110
+ *would* catch is the recovery afterwards: the marker survives a restart until
2111
+ the new session consumes a tick, so every operator SIGTERM-and-resume would
2112
+ page about the healthy session they just fixed. Telling those apart needs the
2113
+ process start time against the marker's, and herdr's `pane process-info`
2114
+ reports pids, not start times. The daemon gives up at most one tick of
2115
+ coverage and never cries wolf.
2116
+
2117
+ The first tick that actually sends clears the counter and deletes the marker,
2118
+ and it deletes one it did not write: recovery normally arrives as a fresh
2119
+ process resuming the same transcript, so the session doing the clearing is not
2120
+ the session that stalled. Nothing else removes the file. Neither the write nor
2121
+ the delete can take the heartbeat down — a filesystem error is logged and the
2122
+ tick carries on.
2123
+
2124
+ `omp-conductor status` reads the same marker from the **state directory** and
2125
+ prints one more line under the daemon block:
2126
+
2127
+ ```text
2128
+ orchestrator STALLED since 2026-08-07T06:27:55.123Z — 2 ticks queued unconsumed — the agent loop is not draining
2129
+ ```
2130
+
2131
+ That reading is the reference deploy's convention — the orchestrator session
2132
+ runs from `~/.omp/conductor`, which is the state directory — and it is
2133
+ one-directional: a line there proves a wedge, and its absence proves nothing,
2134
+ least of all on a fleet whose session lives somewhere else.
2135
+
2136
+ ## CLI reference
2137
+
2138
+ ```bash
2139
+ omp-conductor setup [area] [--no-ai] [--project NAME]
2140
+ omp-conductor setup host [NAME] [--project NAME]
2141
+ omp-conductor setup graph [--no-seed] [--print] [--project NAME]
2142
+ omp-conductor start [--port N] [--project NAME]
2143
+ omp-conductor --version
2144
+ omp-conductor stop
2145
+ omp-conductor restart [--now] [--timeout SECONDS] [--port N] [--project NAME]
2146
+ omp-conductor upgrade [--to VERSION] [--project NAME]
2147
+ omp-conductor upgrade-install --to VERSION [--project NAME]
2148
+ omp-conductor upgrade-rollback
2149
+ omp-conductor status [--project NAME]
2150
+ omp-conductor doctor [--project NAME] [--json] [--probe-telegram]
2151
+ omp-conductor ledger [--issue N] [--limit N] [--project NAME]
2152
+ omp-conductor board [--project NAME]
2153
+ omp-conductor hold [--keep-ticks] [--project NAME]
2154
+ omp-conductor stop [--pane] [--project NAME]
2155
+ omp-conductor arm [--project NAME]
2156
+ omp-conductor disarm [--project NAME]
2157
+ omp-conductor tail <issue> [--project NAME]
2158
+ omp-conductor extend <issue> --turns N [--project NAME]
2159
+ omp-conductor worker pause <issue> [--project NAME]
2160
+ omp-conductor worker resume <issue> [--project NAME]
2161
+ omp-conductor worker stop <issue> --reason TEXT [--project NAME]
2162
+ omp-conductor unblock <issue> [--force] [--no-requeue] [--project NAME]
2163
+ omp-conductor verb <conductor_*> [--project NAME] [--arg k=v ...]
2164
+ omp-conductor friction <escalation-digest|report-noise|report-surprise> --detail TEXT [--issue N] [--project NAME]
2165
+ omp-conductor event record --category NAME --summary TEXT --evidence REF [--occurred-at ISO] [--project NAME]
2166
+ omp-conductor report --text TEXT [--kind material|digest|tier2|decision-needed|fleet-stopped|confirmed-failure] [--events IDS] [--notices IDS] [--project NAME]
2167
+ omp-conductor message --text TEXT [--project NAME]
2168
+ omp-conductor decision open --question TEXT [--blocks TEXT] [--resolves-when COND] [--project NAME]
2169
+ omp-conductor decision resolve <id> --answer TEXT [--project NAME]
2170
+ omp-conductor decision withdraw <id> [--reason TEXT] [--project NAME]
2171
+ omp-conductor decision list [--project NAME]
2172
+ omp-conductor daemon [--once] [--port N] [--project NAME]
2173
+ omp-conductor resume [--project NAME]
2174
+ omp-conductor brief-upgrade [--migrate|--retrofit] [--apply] [--file PATH] [--project NAME]
2175
+ omp-conductor help
2176
+ ```
2177
+
2178
+ | Command | Scope | Behaviour |
2179
+ | --- | --- | --- |
2180
+ | `setup [area] [--no-ai] [--project NAME]` | project | The deterministic interview, in a plain terminal — the same prompts, the same one-writer apply sequence, and the same single consent gate as `omp-conductor setup`, which is now one dialog implementation of the shared surface rather than the only way in. Bare is a full first run, or — when the project already exists — a chooser of which area to amend. Naming an area positionally skips that chooser and amends only that area: `tracker`, `gates`, `caps`, `code-graph`, `authority`, `policy`, `escalation`, `reporting`, `brief`. `host` and `graph` are install subcommands rather than areas and are matched first; anything else exits `2` listing both vocabularies. Every prompt shows its current value as the default, and Enter accepts what you see; `Ctrl-C` at any prompt abandons the run and writes nothing. Setup also **reads your repos to propose answers**: the gates prompt is pre-filled from what CI actually runs, and the brief's `## Project context` and release procedure are drafted from every routing repo and shown for confirmation before anything is written. Each probe is a short session with **no shell, no editor and no verbs** in a throwaway shallow clone, and every answer is a proposal you edit or decline — a probe that cannot clone, cannot reach a model, or answers unusably costs you one warning and the shipped stub. `--no-ai` asks every question with the reading half removed. |
2181
+ | `setup host [--project NAME]` | host | Re-render and stage the systemd unit, then **run** the install: `install -m 0644` into `/etc/systemd/system`, `daemon-reload`, `enable`, `restart`. Stages the fleet recovery oneshot (`omp-conductor-recover.service`) and its playbook (`/usr/local/sbin/omp-conductor-recover`) alongside, and installs them **before** the fleet units: both fleet units carry `OnFailure=` to the recovery unit, so a crash-looped daemon or herdr session now collects evidence durably, attempts one bounded recovery, and pages tier-2 instead of dying silently (#485). Every command is shown with its exact argv, one confirm covers the batch, and `sudo` asks for your password once before the first step — or is skipped entirely on a fleet that genuinely runs as root. The first failure stops the rest and prints the un-run remainder verbatim so you can finish by hand. Refuses an *escalated* invocation (`sudo`, or `sudo -i`/`su -` detected by the invoking account disagreeing with the fleet's) before writing anything, naming both accounts, because staging derives the unit's `User=`/`HOME=` from whoever ran it. On a non-Linux host the files are still staged and only the `systemctl` steps are refused. |
2182
+ | `setup graph [--no-seed] [--print] [--project NAME]` | project | The code-graph install end to end, in one preview and one confirm: check the prerequisites read-only and stop before installing anything when `codebase-memory-mcp` is absent or no MCP entry mounts it (printing the entry to add); `git clone` each missing index-only checkout **as you, never through sudo**; install and enable `cbm-reindex.timer` as root; then seed one indexing run so the first fetch happens while you watch, and verify with the same probe `status` uses. A repo that does not verify is a failure with the remediation, not a success — staged-but-not-trusted is how you discover months later that no worker read an index. `--no-seed` enables the timer without the seeding run and says plainly the graph is unusable until it first fires; it never skips the prerequisite or clone steps. `--print` changes nothing. Exits `1` when no repo has [`graphProject`](#configuration). |
2183
+ | `start` | host | Start `herdr-fleet.service` when that optional unit is installed, clearing a previous pane-recovery pin, then start the dispatch daemon and wait until it answers `GET /healthz`. When `omp-conductor.service` is installed, systemd is the only start path: even `start --project NAME` restores the shared unit and uses the name only to verify that `/healthz` serves the requested project. A detached daemon is allowed only when the unit is proven absent. It never clears pause or arms ticks. Refuses if a daemon is already live, naming its pid; manager refusal or unprovable ownership is an error rather than a detached fallback. |
2184
+ | `stop` | fleet | Prefer `systemctl stop omp-conductor.service` when that unit's MainPID is the live daemon — systemd then owns the stop and will not schedule a restart for the exit it just requested. Otherwise `SIGTERM`, then `SIGKILL` after a 10-second grace period. Prints `not running` when there is nothing to stop, and tags the confirmation with `(via systemctl)` when the unit path was used. |
2185
+ | `restart [--now] [--timeout SECONDS] [--port N] [--project NAME]` | host | Drains the fleet by default: pause new claims, wait until live workers reach `0 / N` (bounded by `--timeout SECONDS`, default 1800 = 30 min), restart, then restore the prior dispatch state. A daemon serving multiple configured projects makes restart host-wide: `--project` is rejected because draining one queue and restarting the shared process would kill another project's workers. Prefer `systemctl restart` when the unit owns the live pid so the replacement stays supervised; only a host proven not to have the installed unit may fall back to the standalone stop/start path. `--now` skips the drain and restarts immediately, orphaning any live runs (old behaviour). A drain that hits `--timeout` restarts nothing and leaves dispatch paused — `omp-conductor resume` lifts it, or re-run `restart` to keep waiting. The new process **salvages dirty live worktrees before orphaning** those rows — see [Deploying a new package onto a busy fleet](#deploying-a-new-package-onto-a-busy-fleet). |
2186
+ | `upgrade [--to VERSION] [--project NAME]` | host | Deterministically update the Bun-global CLI, omp plugin, Herdr recovery plugin, and managed brief as one release. Resolves the npm version and exact `gitHead`, pauses only new claims, drains active workers, installs all surfaces, reloads Herdr and the daemon, waits for pane recovery, verifies identities and fleet health twice, then restores the original dispatch state. Host-wide by default: one daemon serves every configured project, so a bare run drains all of them and refreshes every brief. `--project` is rejected when the live daemon serves several projects — draining one queue and restarting the shared daemon would kill another's workers. A no-op when already current. Failure leaves dispatch paused. Must run outside a Herdr-managed session. The same transaction runs detached, without a human, as the fleet-installs-itself path: the orchestrator calls the `conductor_install` verb under the granted `install` shape, the daemon validates the version against npm and starts a transient systemd unit (`upgrade-install`) outside the pane and the daemon, and the first tick after the restart verifies version, `/healthz`, ticks, pane and `doctor` against the durable upgrade journal before restoring dispatch — rolling back and paging tier-2 on any gap. |
2187
+ | `status [--project NAME]` | fleet | Layered fleet report first: `dispatch` / `ticks` / next scheduled tick / `pane` / `recovery` / `herdr` / `telegram` / `brief` / `decisions` / optional `failure classes` and `code graph` / `daemon`, then the project body. The project body includes the latest completed dispatch timestamp, ready/routed/admitted counts, bounded hold groups, and the GitHub API budget (`graphql` / `core` remaining and reset, in the caps block); API failures are marked `DEGRADED` so queue starvation cannot look idle. Active-run lines overlay cooperative worker `paused`/`pausing` from `/healthz` without changing SQLite `running` state or the live worker count. The next tick comes from the live heartbeat process, not a guess from log timestamps. Telegram health uses `getMe` to prove API authentication without sending a message and reports inbound bridge configuration separately. Configured graphs report prerequisites, indexed repos, timer state, and refresh freshness without blocking dispatch. A `reports` block lists everything the outbox has not delivered, with its age, and prints `pending` (nobody has it) differently from `SENDING` (outcome unknown, it may already have arrived) — see [Report delivery](#report-delivery-the-outbox). The daemon block includes `rss` from `/healthz`; live workers add a busy-deploy warning. A `.conductor-stalled` marker adds an `orchestrator STALLED since …` line. |
2188
+ | `ledger [--issue N] [--limit N]` | project | The action audit: every [mediated-verb](#the-mediated-verbs-126) mutation and every next-attempt turn budget. Verb entries include the arguments, decision, named refusal, and resulting SHA. Turn-budget entries remain after an override is replaced or consumed. Reads (`conductor_pr_status`) are absent so polling cannot bury the signal. `--issue` narrows both histories; `--limit` defaults to 50. Recent verb refusals and pending turn overrides also appear in `status`. |
2189
+ | `board [--project NAME]` | fleet | Live keyboard-driven kanban over the same SQLite and `/healthz` truth as `status`, plus the tracker's current labels: Queue, Claimed, Running, Green, Blocked, Failed, Orphaned, the last 24 hours of Merged and Settled, and Parked (an issue the tracker has not confirmed closed — still open, or a label read that failed — so nothing dispatches it until a human labels it). Columns are mutually exclusive and describe current state, not the newest run row, so a requeued issue is queued rather than failed and a closed issue is neither. Refreshes run/spend/turn values every second, and health plus the label read every ten seconds. `Enter` follows the selected transcript in place; `u` invokes the existing unblock workflow on a Blocked, Failed, or Orphaned card; `i` / `p` open the issue / PR; `r` refreshes health; `?` shows all keys. Requires an interactive terminal of at least 50×20. |
2190
+ | `hold [--keep-ticks] [--project NAME]` | fleet | Soft stop: pause claiming **and** disarm ticks. Daemon and pane stay up. Prefer this when the intent is "stop the conductor" without killing processes. `--keep-ticks` pauses claiming but leaves the arm marker, so the heartbeat keeps reporting and `resume` alone restores the fleet — no fresh arm challenge. See [Stop the conductor](README.md#stop-the-conductor-hold--stop). |
2191
+ | `stop [--pane] [--project NAME]` | fleet | Stop the conductor: pause claiming, disarm ticks, then stop the dispatch daemon (systemctl-aware). Pane stays up unless `--pane` is passed. `stop --pane` also pins herdr-conductor recovery off for the conductor agent only — it does **not** stop `herdr-fleet.service` or any other herdr session. Fail-closed: exits nonzero unless the agent is proven gone. To bounce the daemon without stopping the fleet, use `restart`. |
2192
+ | `arm [--project NAME]` | fleet | Proof-gated: send a Telegram challenge and write this project's arm marker only after your reply appears as a user turn in the orchestrator transcript. The challenge names the project, so a host running two fleets is not ambiguous. Never auto-armed by `resume` / `hold`. |
2193
+ | `disarm [--project NAME]` | fleet | Remove this project's arm marker so its ticks skip; another project's ticks keep running. Also clears a pre-per-project shared `armed` marker while that marker is still what holds this fleet's gate open — otherwise the disarm would not disarm. Processes untouched. |
2194
+ | `tail <issue>` | project | Follow the newest run for that issue: the worker's assistant text as `assistant: …` and each tool it calls as `tool: <name>`, printed as they land. Workers are omp sessions inside the daemon rather than terminals, so this is the only way to watch one live — a herdr pane running it becomes an observation window. Starts from the top of the transcript, not the end, so attaching to a run that is already ten turns in shows those ten turns. Exits `1` with `no run recorded for #N` when the issue has never been dispatched, or `no transcript yet (state: …)` when the attempt has not opened one. Otherwise it runs until `Ctrl-C`, or until the run has finished and its transcript has been silent for five seconds, and prints `run ended: <state>`. |
2195
+ | `extend <issue> --turns N [--project NAME]` | project | Raise a live worker's effective turn ceiling through its owning daemon without restarting its session. If the latest run is failed, killed, orphaned, or blocked and has no live controller, store a one-shot ceiling for that issue's next claimed attempt instead. A next-attempt value must exceed the project base, every extension must stay at or below `workerMaxTurnsCeiling`, and live extensions remain monotonic. The pending value appears in `status`, is recorded in `ledger`, and is consumed atomically by one claim. |
2196
+ | `worker pause <issue>` / `worker resume <issue>` | project | Cooperatively park one live worker without changing its run state or lane. Pause aborts the active turn to harness idle and freezes the remaining wall-clock budget; resume continues the same session with a prompt to re-check its last action before repeating it. This is separate from fleet-level `hold`, which refuses new claims and work-starting mutations while allowing pre-pause completion work and releases. |
2197
+ | `worker stop <issue> --reason TEXT [--project NAME]` | project | Terminally end a running or cooperatively paused worker. The reason is required (1–500 characters) and persisted on the run. The command waits for settlement, records the distinct `stopped` state, salvages and publishes dirty work, removes `agent:in-progress` through the durable label outbox, and consumes neither failed-attempt nor continuation budget. If salvage fails, the tree holding the only copy stays in place and the command names it. Repeating stop is idempotent and reports the run's already-terminal state. |
2198
+ | `unblock <issue> [--force] [--no-requeue]` | project | Remove that issue's `blocked` and `failed` labels so an answered escalation can be claimed again, and restore the project queue label by default so the dispatcher actually sees it. `agent:in-progress` comes off too, but only when the newest recorded run is terminal — that row is the proof no worker still owns the issue, so a live run keeps the label (and the queue label stays off until that run settles), and so does an issue with no run row at all. Run history remains intact: blocks consume the independent continuation budget, not failed implementation attempts. The output reports both budgets and warns when either will make the next tick escalate instead of dispatch. The label changes go through the [label projection outbox](#how-one-tick-works): they are applied inline before the command returns, but **a tracker that refuses them (403, rate limit) no longer fails the verb** — it exits `0`, the intended label state is durable and the daemon retries it, and the output says `label sync queued (N pending) — the daemon retries` instead of claiming the labels were restored. Safety is preserved, but the issue is only claimable once the queue label itself lands: the queue read asks GitHub for issues carrying that label, so a refused queue-label add keeps the issue out of dispatch until projection succeeds. `--no-requeue` clears the state labels only, leaving the queue label untouched — the case where you are about to close the issue. **Refuses, clearing nothing and exiting `3`, when the newest attempt's work could not be committed and its worktree is the only copy** — re-claiming removes that tree. `--force` records the operator's acceptance on the run row and then clears; the salvage failure stays in history. Exits `2` when the issue number is missing or malformed. |
2199
+ | `verb <conductor_*> [--arg k=v ...]` | project | Run one [mediated verb](#the-mediated-verbs-126) as the orchestrator, from the CLI — the external-orchestrator half of the verb surface. Every argument goes in as a `--arg k=v` string; an orchestrator can merge (`conductor_pr_merge`), label (`conductor_label`), release (`conductor_release`), update a branch (`conductor_pr_update_branch`) or title/body (`conductor_pr_update`), or read PR state (`conductor_pr_status`). The daemon applies the same checks and writes the same ledger rows a session's call would; a missing `--arg` is refused exactly as a missing tool argument is, worker-only verbs (`conductor_push`, `conductor_pr_create`) are refused with `role-not-allowed`, and a refusal exits `3`. An unknown verb exits `2`. |
2200
+ | `friction <kind> --detail TEXT [--issue N]` | project | Record one bounded judgment the daemon cannot infer: an escalation belonged in a digest, or a tick report was noise/surprising. The detail is limited to 160 characters. One event never changes policy; three observations inside seven days make the aggregate eligible for one Learning-loop prompt, followed by a seven-day cooldown. |
2201
+ | `report --text TEXT [--kind material|digest|tier2|decision-needed|fleet-stopped|confirmed-failure]` | project | Hand a rendered report to the daemon's durable outbox. The command persists the text **before** anything can send and prints a durable handoff id. A material report submitted during quiet hours becomes a held notice until the window opens; otherwise it becomes a report whose delivery the daemon owns, retries with bounded backoff, and records. Delivery is [at-least-once](#report-delivery-the-outbox), so a crash mid-send is retried as a possible repeat and `delivered` never proves exactly one message. `--kind digest` is accepted at most once per local day, decided from the ledger; an unknown `--kind` exits `2`. The remaining kinds declare the report's interrupt category — the escalation handoff: the reporting policy decides between immediate delivery and a durable hold exactly as for a daemon escalation of that category, A repeated identical call exits `2` only while the earlier handoff is still queued undelivered; once it lands, the same text is admitted again (the handoff state decides, not a permanent ledger). Anything still owed appears in `status` with its age. |
2202
+ | `decision open --question TEXT [--blocks TEXT] [--resolves-when COND]` | project | Record a question the orchestrator has put to you, and print its id. A question that lives only in a session's context is lost at the next compaction — after which it is either asked twice or dropped silently. `--resolves-when` attaches a machine-checkable condition: `pr-merged:<https url>`, `pr-checks-green:<https url>`, `pr-mergeable:<https url>`, `issue-closed:<n>`, `npm-version:<pkg>@<version>`, or `rate-limit-reset:github`; anything else exits `2` listing the six forms. See [The decision ledger](#the-decision-ledger-136). |
2203
+ | `decision resolve <id> --answer TEXT` | project | Record what you decided. Exits `1` naming the id when it is unknown or no longer open, so a second answer cannot overwrite the first. |
2204
+ | `decision withdraw <id> [--reason TEXT]` | project | Close a question the session stopped needing, with why. Same guard as `resolve`. |
2205
+ | `decision list` | project | Open questions, oldest first: id, age, what each blocks, whether its condition is met, and the question. Prints `no open decisions` when there are none. |
2206
+ | `daemon` | host | Run the loop in the **foreground**, ticking every 5 minutes and serving `/healthz`. Admitted workers run in a tracked background pool, so settlement and capacity checks remain periodic while they work; shutdown drains the pool before closing the store. This is what `start` launches and what a systemd unit should call. |
2207
+ | `daemon --once` | host | Run a single tick, wait for workers admitted by that tick, and exit. No HTTP server or pidfile — a drill must not register itself as the daemon, or the next reader believes it and the real daemon's in-flight runs get reconciled as orphans. |
2208
+ | `--port N` | — | Accepted by `start`, `restart` and `daemon`. Both `--port 9000` and `--port=9000` work; missing or out of range exits `2` rather than falling back to the default, because probing the wrong endpoint is worse than a hard failure. |
2209
+ | `--project NAME` | — | Selects a project for project-scoped commands and foreground `daemon`. On an installed shared service, `start --project NAME` still starts the host-wide unit and uses the name only to verify `/healthz`; a draining `restart --project NAME` is rejected when that daemon serves multiple projects. A project-only daemon is available only through an explicit foreground `daemon --project NAME` or standalone start on a host proven not to have the unit. On a multi-project host, `setup host` refuses without a name and says which part needs it: the units it installs are host-global, and only the per-project tail (tick config, brief link) is per-project — pass `--project NAME` or the positional `setup host NAME`. |
2210
+ | `pause [--reason TEXT]` | fleet | Stop new claims and work-starting mutations only. The running daemon notices on its next tick; runs already in flight finish. The orchestrator may still merge, update, or label runs admitted before the pause, and may release when the release policy's own preconditions hold. The orchestrator heartbeat keeps ticking if armed — its gate is the arm marker, not this flag. Per-worker pause is separate. Prefer `hold` to silence both. `--reason TEXT` is recorded in the pause sentinel, which `status` shows as the pause provenance. |
2211
+ | `resume [--project NAME]` | fleet | Clear pause and any `stop --pane` recovery pin — does **not** re-arm. Run `arm` after an inbound Telegram proof to resume ticks. |
2212
+ | `--version`, `-V`, `version` | none | Print the installed `omp-conductor` package version and exit `0`. Works from the global binary and npm/plugin install because it reads the package metadata beside the shipped CLI. |
2213
+ | `brief-upgrade` | project | Inspect the package-floor + shared policy (when present) + `POLICY.md` overlay. Reports by default, naming which layers compose; see [Keeping a brief current](#keeping-a-brief-current). |
2214
+ | `--migrate` | — | Only for `brief-upgrade`. Lift a bannered `ORCHESTRATOR.md` owned half into `POLICY.md` and recompose. Dry-run unless `--apply`. |
2215
+ | `--retrofit` | — | Only for `brief-upgrade`. Propose (or with `--apply`, write) a `YOURS TO EDIT` banner before the first owned-topic heading on a hand-written brief. |
2216
+ | `--apply` | — | Only for `brief-upgrade`. Confirms `--migrate` / `--retrofit`. On its own it exits `2`: the legacy single-file merge was removed in 0.4.3. |
2217
+ | `--file PATH` | — | Only for `brief-upgrade`. Check a brief that is not where the wizard would have put it, on a host that may have no config at all. |
2218
+ | `help`, `--help`, `-h` | none | Print usage. An unknown or missing verb prints it too, and exits `2`. |
2219
+
2220
+ Pause is a sentinel under the state directory and survives a daemon restart.
2221
+ `hold --project NAME` writes `paused-<name>` for that project only; a bare
2222
+ `paused` file (legacy / all-projects) pauses every project. It refuses new claims
2223
+ and work-starting mutations, allows completion verbs only for runs admitted
2224
+ before the pause, and leaves `conductor_release` to its normal authority, grant,
2225
+ and precondition checks. Per-worker pause is independent. Hold also removes the
2226
+ arm marker the heartbeat reads, so both brains go quiet without killing processes.
2227
+
2228
+ Every one of these is a verb on the `omp-conductor` binary, and the Scope column
2229
+ is its project classification, declared once in `src/commands/context.ts`
2230
+ (`COMMAND_SCOPES`) and enforced by the tests there (#514):
2231
+
2232
+ - **project** — acts on exactly one configured project; a multi-project host
2233
+ demands `--project NAME` (or the documented positional form, e.g.
2234
+ `setup host NAME`).
2235
+ - **host** — acts on the one shared host (units, daemon, package surfaces);
2236
+ never resolves a project, so a multi-project config cannot make it ambiguous.
2237
+ `--project` is accepted only for the documented per-project tail that row
2238
+ names, and refused whenever it cannot be honoured (upgrade/restart reject
2239
+ narrowing the shared daemon; `upgrade-rollback` rejects it outright).
2240
+ - **fleet** — acts on every configured project by default (or `--all`), with
2241
+ `--project NAME` narrowing to one.
2242
+ - **—** (em dash) — the row is a flag rather than a command, or the command
2243
+ takes no project at all.
2244
+
2245
+ There is no in-session command: an omp session that wants any of them shells out
2246
+ to the binary, which is what keeps one implementation and one ledger entry per
2247
+ action.
2248
+
2249
+ ### Health endpoint
2250
+
2251
+ ```bash
2252
+ curl -s localhost:8787/healthz
2253
+ ```
2254
+
2255
+ ```json
2256
+ {
2257
+ "ok": true,
2258
+ "rssBytes": 123456789,
2259
+ "projects": [
2260
+ {
2261
+ "ok": true,
2262
+ "paused": false,
2263
+ "activeRuns": 1,
2264
+ "project": "demo",
2265
+ "dispatch": {
2266
+ "completedAt": 1786185678000,
2267
+ "ready": 8,
2268
+ "routed": 8,
2269
+ "admitted": 0,
2270
+ "degraded": true,
2271
+ "holds": [
2272
+ { "reason": "parent-lookup-error", "count": 8, "issues": [321, 320, 318] }
2273
+ ]
2274
+ },
2275
+ "codeGraph": {
2276
+ "configured": true,
2277
+ "status": "degraded",
2278
+ "checkedAt": "2026-08-08T13:00:00.000Z",
2279
+ "prerequisites": { "indexer": "present", "mcpMount": "missing" },
2280
+ "repos": [
2281
+ {
2282
+ "name": "api",
2283
+ "path": "/home/fleet/.cache/conductor-graph/acme/api",
2284
+ "clone": "present",
2285
+ "index": "present"
2286
+ }
2287
+ ],
2288
+ "timer": { "enabled": "enabled", "active": "active" },
2289
+ "refresh": {
2290
+ "result": "success",
2291
+ "fresh": true,
2292
+ "lastSuccessAt": "2026-08-08T12:50:00.000Z",
2293
+ "ageMs": 600000
2294
+ },
2295
+ "reasons": ["worker MCP configuration does not mount the indexer"]
2296
+ }
2297
+ }
2298
+ ]
2299
+ }
2300
+ ```
2301
+
2302
+ Any other path or method returns `404`. Top-level `ok` is process liveness across
2303
+ every served project; top-level `rssBytes` is the daemon's resident set.
2304
+ Per-project blocks keep `paused`, `activeRuns`, `dispatch`, `codeGraph`, and
2305
+ `workers`. Nonfatal admission errors and graph degradation keep `ok` `true` so a
2306
+ supervisor does not restart-loop. Inspect `dispatch.degraded` and its bounded
2307
+ reason groups for queue starvation; inspect `codeGraph` for configured graph
2308
+ health. `activeRuns` counts occupied issues — live workers plus green PRs
2309
+ awaiting merge.
2310
+
2311
+ ## What a worker may and may not do
2312
+
2313
+ Each worker gets one brief, one worktree, one branch, and no knowledge of the
2314
+ dispatcher. The brief is explicit about the boundary:
2315
+
2316
+ | It may | It must not |
2317
+ | --- | --- |
2318
+ | Read the issue and the repo's own guidance (`AGENTS.md`, `CLAUDE.md`, `CONTEXT.md`, relevant ADRs) before writing anything. | Touch any path outside its worktree, or switch branches. |
2319
+ | Query its repo's [code graph](#code-graph-discovery), when one is configured, by the project name whose `root_path` matches the clone its brief names. | Query that graph by its own cwd or worktree path — no index of a worktree exists — or treat what it returns as current. It is a snapshot of the clone's default branch; the real file in the worktree wins. |
2320
+ | Edit code inside its own worktree. | Weaken, skip, delete or loosen **any test it did not write** — that is a design question to escalate, and it is checked by diff review before the push. |
2321
+ | Add or update tests for behaviour it introduced. | Suppress a warning, delete an assertion, or special-case an input to make a check pass. |
2322
+ | Run the repo's configured cheap gates, each from its listed `cwd`, over the whole tree. | Run docker or image builds, production builds, browser/e2e suites, or the full test suite on the shared host — CI owns the heavy gates. |
2323
+ | Review its whole diff, then commit and publish once with `conductor_push`. One corrective push if CI is red. | Force-push, `git add -f`, or add AI/co-author attribution. There is no force path to reach: `conductor_push` publishes that run's branch fast-forward only and takes no other ref. Red twice means stop and report, not push a third time. |
2324
+ | Open a PR with `conductor_pr_create`, and poll CI to a verdict with `conductor_pr_status`. | Run `gh pr merge` — or reach `conductor_pr_merge`, which refuses a worker session mechanically. **A worker is never authorised to merge**, whoever else holds the authority, so PRs land one at a time with a freshness re-check; two workers merging concurrently is how agent PRs clobber each other. Who *may* merge is the [`authority`](#configuration) answer, and it is never the worker. The verb refusal is mechanical, and so is the channel: each one is bound to the pid the daemon spawned, so reaching for the orchestrator's socket is refused rather than honoured (see [The transport](#the-transport)). Shelling out to `gh` remains a prohibition, not an impossibility — a session shares the daemon's credentials. |
2325
+ | Escalate: ambiguity, a cross-repo contract, a needed credential, a product or data-migration decision, a blocking existing test, CI red twice, or most of the wall-clock budget burned. | Cut a release, push a tag, publish to npm, edit a deployment pin, deploy, or touch infrastructure or secrets. `conductor_release` and `conductor_label` refuse a worker whatever `releasePolicy` says, because the check compares the caller against the configured holder rather than ruling one value out. The in-session tripwire still blocks recognised release/deploy tool calls early and audits the attempt, but it is [defence in depth](#the-mediated-verbs-126), not the gate. |
2326
+
2327
+ The worker ends with a seven-line evidence report (issue, PR, observed head SHA,
2328
+ state, gates, changed, next). A textual `pushed-green` claim is not success: the
2329
+ daemon repeats the PR/head/check verification before it records that state.
2330
+
2331
+ ### Worker confinement and the integrity tripwire
2332
+
2333
+ A worker session is rooted at its worktree `cwd`. **Structured file tools are
2334
+ gated mechanically:** `runWorker` asks `createSession({ role: "worker" })`,
2335
+ which installs an inline harness extension that blocks `write` / `edit` /
2336
+ `read` / `grep` / `glob` when the tool's path resolves outside that worktree
2337
+ (symlink-aware). Target selection was already mechanical — only a repo in
2338
+ `routing.repos` is ever checked out — and the caps still bound *how much* work
2339
+ happens.
2340
+
2341
+ General shell access is not confined to the worktree. Its argument is an opaque
2342
+ program, so the brief still forbids path escape and the deploy-level answer is a
2343
+ least-privilege worker uid (below). The narrower release-policy tripwire does
2344
+ inspect explicit command shapes such as `git tag`, `npm publish`, and deploy
2345
+ verbs; it blocks those before execution when `releasePolicy` is `none`.
2346
+
2347
+ #### Integrity tripwire (package self-hash)
2348
+
2349
+ Separately, the conductor watches *itself*. At startup the daemon sha256s every
2350
+ `.ts` and `.md` file of its own installed `src/` — the dispatcher and the briefs
2351
+ both, since rewriting a brief buys more than rewriting the loop — and re-walks
2352
+ that tree on every tick (about 0.6 ms). Any difference at all, changed or added
2353
+ or removed, is read as the package having been modified underneath a running
2354
+ daemon: the tick claims nothing, the fleet is paused, and a tier-2 escalation
2355
+ naming the first few differing paths pages you **once**, not every five minutes.
2356
+
2357
+ **A normal deploy never trips it.** The baseline is recorded per daemon process,
2358
+ so installing a new build and restarting the unit re-records it from the new
2359
+ files; only a change that lands *while* a daemon is holding the package open can
2360
+ diverge from it. That also means `omp-conductor resume` on its own will not hold
2361
+ — the next tick re-walks, still differs, and pauses again. Put the files back, or
2362
+ restart onto the build you meant to be running.
2363
+
2364
+ This catches a worker (or human) that still managed to edit the live install —
2365
+ including via `bash` — after the fact. It is detection for the package boundary,
2366
+ not a substitute for the worktree gate or a dedicated uid.
2367
+
2368
+ #### Least-privilege worker uid (deploy)
2369
+
2370
+ The largest remaining win is OS-level: run the daemon (or at least worker
2371
+ sessions, when the harness supports a uid switch) as a user that can write only
2372
+ its worktrees and mirrors. A sketch that matches the reference single-host
2373
+ deploy:
2374
+
2375
+ 1. Create a system user, e.g. `conductor-worker`, with home under
2376
+ `/var/lib/conductor-worker` (or similar).
2377
+ 2. `chown` the project's `workspaceRoot` and `mirrorRoot` to that user; leave
2378
+ `~/.omp/conductor/config.json` readable only by the operator/daemon account
2379
+ (`0600` as shipped).
2380
+ 3. Do **not** put the worker uid in `docker` / `sudoers`, and do not give it the
2381
+ operator's `gh` auth if a narrower deploy token can open PRs in the routed
2382
+ repos alone.
2383
+ 4. Point the [example systemd unit](systemd/omp-conductor.service.example)
2384
+ `User=` / `Group=` at that account once the daemon itself should run
2385
+ unprivileged end-to-end.
2386
+
2387
+ Until that uid exists, a root-or-operator daemon still has a mechanical
2388
+ worktree gate on structured tools and an integrity tripwire on its own package —
2389
+ but `bash` plus host credentials remain a prompt-and-deploy problem.
2390
+
2391
+ ### The orchestrator is unconfined, deliberately
2392
+
2393
+ There is **no mechanical file gate on the orchestrator session**, and that is an
2394
+ operator decision rather than an omission (#143).
2395
+
2396
+ A previous release jailed it to an allowlist. That gate could only ever be
2397
+ installed by `createLocalSession`, so it existed exactly in the sessions this
2398
+ daemon spawns — and the supported shape for a heartbeat orchestrator is an
2399
+ `omp` session the operator starts themselves, which never had it. A boundary
2400
+ present in one deployment out of two is not a boundary, and the brief asserting
2401
+ it was absolute was the worse half of the bug: a session that believes it is
2402
+ gated stops checking itself.
2403
+
2404
+ What holds the orchestrator instead:
2405
+
2406
+ | | |
2407
+ | --- | --- |
2408
+ | **The brief** | `ORCHESTRATOR.md`'s hard boundaries — never read or edit a worker's checkout or the mirror cache; when you need a run's code, read its PR. |
2409
+ | **The action ledger** | Every `conductor_*` mutation and operator-selected next-attempt turn budget remains auditable. `omp-conductor ledger` shows both, including refused calls and consumed or replaced budget overrides. |
2410
+ | **The dispatcher** | Merge, label and release authority are checked in the daemon against the operator's grant, across a process boundary, never in the prompt. |
2411
+
2412
+ Unconfined means auditable, not licensed. `orchestratorReadPaths` is retired: it
2413
+ is still accepted in a config and ignored, so a fleet carrying it upgrades
2414
+ without editing anything.
2415
+
2416
+ ## The mediated verbs (#126)
2417
+
2418
+ A session can reach `gh`: it inherits the daemon's environment, credentials and
2419
+ all. It is told not to publish with it. These verbs are the sanctioned route
2420
+ instead, because the dispatcher owns the settlement record — a push or a PR the
2421
+ daemon did not perform is a run it cannot account for, and the checks that would
2422
+ have refused it never ran. The point of them is *where those checks run*: in the
2423
+ daemon, across a process boundary, not in a prompt the model can rewrite.
2424
+
2425
+ ### The verbs
2426
+
2427
+ | Verb | Allowed caller | What the daemon checks before acting |
2428
+ | --- | --- | --- |
2429
+ | `conductor_push` | the worker owning the run | The ref is exactly `refs/heads/<that run's branch>`. Fast-forward only; there is no force argument to reject because none is declared. |
2430
+ | `conductor_pr_create` | the worker owning the run | The run has no open PR (the same guard admission uses); head is the run branch; base is the repo's configured `defaultBranch`. |
2431
+ | `conductor_pr_status` | worker or orchestrator | Read-only — nothing to gate. A worker reads only its own run's PR; an orchestrator may name any syntactically valid PR URL, open, merged, or closed, and gets its live state and head (checks are reported when available; a merged or closed PR reports its state instead of an `expected OPEN` refusal). |
2432
+ | `conductor_pr_update_branch` | orchestrator, or the worker owning the run | The PR belongs to this project and is open. A worker may only name its own run's PR. |
2433
+ | `conductor_pr_merge` | **orchestrator only** | Ordinarily, `authority.merge` equals the caller. A hand-edited `recoveryMerges` entry may instead authorize one exact unrecorded PR/head/reason while held. In both paths, `headSha` equals the live head *at execution time*; checks are green at that same SHA; the project route and migration chain are valid; the project's single merge slot is free. |
2434
+ | `conductor_label` | **orchestrator only** | The label is in the project's own vocabulary. Lifecycle labels stay the daemon's. |
2435
+ | `conductor_release` | **orchestrator only** | `authority.release` equals the caller; the per-shape grant permits it; the artefact or environment was declared; the release preconditions hold; the `reason` is in the closed enum. `version-bump-pr` creates or re-validates one deterministic version-only PR and, on a later call, merges only its exact green head through the project's single merge slot. |
2436
+ | `conductor_install` | **orchestrator only** | Gated like a release act: the `install` shape defaults to `human` and a grant is what moves it. The daemon refuses a version npm does not expose with a full `gitHead`, refuses while another install is still in flight, and otherwise starts a detached transient unit that pauses, drains, installs the CLI/omp plugin/Herdr plugin and reloads — outside this session and the daemon. The unit never declares its own success; the first tick after the restart verifies and reports through the durable outbox. |
2437
+
2438
+ Standing merge and release authority use the same exact rule: **the caller's
2439
+ role must equal the configured holder.** `authority` has exactly two values, so
2440
+ a `!== "human"` test would have let a *worker* release. A worker is refused
2441
+ every release shape under the most permissive config there is. The exact
2442
+ operator-authored recovery tuple below is the sole authority exception inside
2443
+ `conductor_pr_merge`; a reviewed version bump is instead a `conductor_release`
2444
+ operation governed throughout by release authority.
2445
+
2446
+ `recoveryMerges` is deliberately narrower than standing merge authority. It is
2447
+ an operator-authored, one-PR escape hatch for a recovery branch that cannot have
2448
+ a run row—for example, a conflict repair created after the fleet was held. It
2449
+ does not admit new work, unpause the fleet, widen repository routing, bypass
2450
+ live-head or check validation, or make a general class of PRs mergeable.
2451
+ Authorizations are re-read from config on every call and every attempted merge
2452
+ is written to the ordinary verb ledger, including refusals.
2453
+
2454
+ For a repo with `release.versionFile`, call `conductor_release` with
2455
+ `shape=version-bump-pr` and the intended `v<semver>` tag. The requested version
2456
+ must be newer than the live semantic version. The first call creates
2457
+ `conductor/release-<version>` from the live default branch, changes only the
2458
+ declared JSON `version`, and opens a normal PR. Call it again after CI: the daemon
2459
+ re-reads that exact PR head, verifies the PR contains only the semantic version
2460
+ change, requires green checks, and merges with GitHub's exact-head guard. The
2461
+ ordinary action ledger records both calls. A raw source push is never delegated,
2462
+ and a worker makes no release decision.
2463
+
2464
+ For Git-backed releases, a repo that declares `release.versionFile` refuses both
2465
+ tag creation and a new tag push until the live default branch's version matches
2466
+ the requested tag. `git-tag` is idempotent when the named tag exists locally but
2467
+ has not been pushed: it re-points the tag to the verified live default-branch
2468
+ head. `git-push-tags` performs the same re-point immediately before pushing if
2469
+ the default branch moved between the two calls. A tag already published on
2470
+ origin is immutable: an identical tag is accepted as already complete, while a
2471
+ different published target is refused and must use a new tag name.
2472
+
2473
+ A `github-release` for the same repo likewise requires that reviewed tag to be
2474
+ present on origin and verifies the tag's version file before creating the
2475
+ release; it never lets GitHub synthesize the missing tag.
2476
+
2477
+ ### The transport
2478
+
2479
+ Identity is never an argument. `project`, `run`, `issue` and the caller's role
2480
+ come from **which socket the call arrived on**, and a request carrying any of
2481
+ those field names is refused outright, named. So a worker on run X cannot *ask*
2482
+ to merge run Y's PR — on its own channel that request is unexpressible.
2483
+
2484
+ **Each channel is bound to one process, because the modes cannot tell sessions
2485
+ apart.** Every session runs as the daemon's own uid, so it matches the *owner*
2486
+ class here: it can list this directory and connect to any socket in it, including
2487
+ the orchestrator's. Authorisation and the ledger both read the role from the
2488
+ channel, so a worker doing that would have been authorised as the orchestrator
2489
+ (under `authority.merge: "orchestrator"`) *and recorded as* the orchestrator. No
2490
+ file mode closes that — the owner bits belong to the uid the session already has.
2491
+
2492
+ So the daemon binds each channel to the **pid it spawned for that session**, and
2493
+ refuses a connection from anything else without answering it, logged the way an
2494
+ impersonation is. Until a channel is bound it refuses everything, because the
2495
+ socket necessarily exists before the child that connects to it. The kernel
2496
+ supplies the pid: `SO_PEERCRED` on Linux, `LOCAL_PEERPID` on macOS. A host where
2497
+ neither can be asked — no loadable libc, or the call refused — refuses every
2498
+ connection on a bound channel and says so at startup, rather than falling back to
2499
+ the uid, which under one shared uid is no check at all.
2500
+
2501
+ The residual is narrow, real, and worth stating: one uid can `ptrace` and signal
2502
+ its siblings, so a determined session can still interfere with the process that
2503
+ *is* bound. That is a far higher bar than connecting to a socket, and closing it
2504
+ needs separate OS principals.
2505
+
2506
+ ```
2507
+ <state dir>/verbs/ daemon-owned, mode 0711
2508
+ run-7-9a783d877d422b9e.sock 0600, bound for run 7
2509
+ run-9-1c40e2a5b6d3f018.sock 0600, bound for run 9
2510
+ orchestrator-4b1f...c2.sock 0600, the orchestrator's
2511
+ ```
2512
+
2513
+ **What these modes buy, and what they do not.** They keep every *other local
2514
+ account* out: `0711` on the parent is traversable but not listable, so no other
2515
+ user can enumerate the fleet's sockets, the suffixes are unguessable, and only
2516
+ the daemon's uid can connect to a `0600` socket at all.
2517
+
2518
+ They are **not** a boundary between runs. Sessions are child processes of the
2519
+ daemon running as its own uid, so a session matches the owner class on all of
2520
+ these: it could list the directory and connect to a sibling's socket. Each run is
2521
+ *handed* its own path and nothing else, which is a convention the run has no
2522
+ reason to break — not an enforcement. What makes breaking it visible is the
2523
+ [ledger](#the-ledger): every call is recorded with the channel it
2524
+ arrived on, so a worker calling on another run's socket is in the record.
2525
+
2526
+ Closing that properly needs the sessions to be different OS principals. A
2527
+ per-run credential boundary that did exactly this shipped and was removed in
2528
+ 0.5.0 — it worked, and the cost was that it also hid the operator's own model
2529
+ credential from every session, so nothing could start. It is not worth
2530
+ re-litigating without solving that first.
2531
+
2532
+ Before binding, the daemon verifies every component of the path is owned by
2533
+ itself (or root), free of symlinks, and unwritable by anyone else; a failed
2534
+ check **refuses dispatch** rather than degrading. Paths are unguessably
2535
+ suffixed, and only the daemon ever unlinks one.
2536
+
2537
+ Peer credentials are asserted server-side — `getpeereid` on macOS, `SO_PEERCRED`
2538
+ on Linux — and the daemon states at startup exactly what that buys rather than
2539
+ implying more. Sessions are child processes running under the daemon's own uid,
2540
+ so the peer check proves the caller is a local process on this host; it is the
2541
+ socket, not the uid, that says which run is calling. A connection whose peer
2542
+ cannot be read at all is closed with no reply and logged.
2543
+
2544
+ ```
2545
+ verb transport: verb sockets in ~/.omp/conductor/verbs (mode 711); each socket
2546
+ 0600 under the daemon's own uid; peer uid asserted with getpeereid
2547
+ ```
2548
+
2549
+ **No mutation route exists on the HTTP port**, and none may be added. That
2550
+ surface is unauthenticated loopback TCP reachable by any local user; a `PUT` or
2551
+ `POST` at any verb path answers 404, pinned by a test.
2552
+
2553
+ The child-side tool handler is a thin client only. It forwards arguments and
2554
+ renders the answer — no policy branch, no local fallback, no second route. With
2555
+ no socket it fails closed and says so, rather than reaching for `git push`.
2556
+
2557
+ ### The ledger
2558
+
2559
+ Every mutating verb call is recorded with its arguments, the decision, the
2560
+ named refusal reason and any resulting SHA. Reads are not: a status poll every
2561
+ thirty seconds would bury the refusals the record exists to surface.
2562
+
2563
+ Every `extend` that sets a next-attempt budget also appends an audit entry.
2564
+ Replacing or consuming the pending override does not erase that history.
2565
+
2566
+ ```console
2567
+ $ omp-conductor ledger --issue 7
2568
+ acme — 3 verb call(s), 1 refused (newest first)
2569
+ 2026-08-09 11:04:12 REFUSE conductor_pr_merge worker #7 [role-not-allowed]
2570
+ prUrl=https://github.com/acme/api/pull/7 headSha=9a783d8… reason=preconditions-met
2571
+ refused: merge authority is the orchestrator's, never a worker session's.
2572
+ 2026-08-09 10:58:03 ALLOW conductor_pr_create worker #7
2573
+ title=fix: settle the head check body=Closes acme/tracker#7
2574
+ opened https://github.com/acme/api/pull/7 (conductor/issue-7 → main).
2575
+ 2026-08-09 10:57:41 ALLOW conductor_push worker #7 9a783d877d42
2576
+ (no arguments)
2577
+ pushed refs/heads/conductor/issue-7 at 9a783d877d42….
2578
+ ```
2579
+
2580
+ The newest few also appear in `omp-conductor status`, because a refused merge is
2581
+ news: it means a session tried to do something the config does not permit.
2582
+
2583
+ `release-policy.ts` stays installed as defence in depth — it refuses early, in
2584
+ the session, with an explanation the model can act on in the same turn, and it
2585
+ leaves a durable record that something tried. It is no longer what *stops* a
2586
+ release. Treat a block there as evidence about a session's intentions; the
2587
+ daemon is what prevented it.
2588
+
2589
+
2590
+ ## Limitations
2591
+
2592
+ Known and deliberate in this version:
2593
+
2594
+ - **`gh` is shelled out to.** Every tracker operation spawns a process and does its
2595
+ own TLS handshake (roughly 200-400 ms each), and failures are classified by
2596
+ matching human-readable stderr rather than a status code. The upside is that no
2597
+ token is ever handled, stored or logged by the daemon.
2598
+ - **`listReady` fetches a single page of 100 issues.** A queue deeper than 100
2599
+ ready issues truncates silently. A backlog that size is a staffing problem before
2600
+ it is a paging one.
2601
+ - **Spend accounting depends on harness telemetry.** Cost arrives only when the
2602
+ harness run carries it; without it `spendUsd` reads `0`, `status` shows `$0.00`,
2603
+ and the daily-spend cap never fires. The turn and wall-clock ceilings are what
2604
+ actually bound a runaway in that case. Do not treat `$0.00` as proof that nothing
2605
+ was spent.
2606
+ - **GitHub is the only tracker.** The internal `Tracker` port is deliberately
2607
+ provider-neutral, but `tracker.kind` accepts only `"github"` today.
2608
+ - **One daemon serves every configured project by default.** `setup host` writes a
2609
+ unit without `--project`. Pass `--project NAME` only to filter a foreground or
2610
+ temporary daemon down to one project.
2611
+ - **Labels are matched exactly and case-sensitively.** `Ready-For-Agent` is not
2612
+ `ready-for-agent`, and the mismatch is silent: the issue is simply never picked
2613
+ up.
2614
+ - **No cross-process lock on the mirrors.** Two dispatch loops fetching the same
2615
+ repo at the same instant can collide on git's ref locks; the run fails and is
2616
+ retried rather than corrupted.
2617
+ - **Uniquely local mirror branches are retained.** Terminal runs are reaped
2618
+ automatically only after every commit exists on a remote ref. A failed salvage
2619
+ push deliberately leaves its branch and tree for an operator rather than
2620
+ trading disk hygiene for data loss.
2621
+ - **`stop` is a bounded best-effort drain.** A signal stops new ticks and the
2622
+ daemon waits for its active worker pool before closing the store. The CLI
2623
+ escalates to `SIGKILL` after 10 seconds, so a worker that needs longer is
2624
+ orphaned and salvaged on restart. Use `pause`, wait for `workers 0 / N`, then
2625
+ stop when a clean drain matters. The supervising unit keeps
2626
+ `SuccessExitStatus=0 143` and renders `Restart=always` (#546): systemd
2627
+ restarts the daemon on any exit — clean or signalled — **except** an explicit
2628
+ `systemctl stop`, which it records as intentional and never undoes. So an
2629
+ unattended SIGTERM (a worker's `bun test`) costs a few seconds of downtime
2630
+ instead of an outage that waits for a human, while the operator's own stop
2631
+ still stops and stays stopped. A genuine crash loop still trips the
2632
+ start-limit burst and reaches the recovery unit through `OnFailure=`, so the
2633
+ restart policy stays a visible detector rather than a silent spin. Prefer
2634
+ `omp-conductor stop` / `systemctl stop` over raw `kill`: a raw kill is not an
2635
+ intentional stop, and the daemon comes straight back.
2636
+ - **A failed orchestrator degrades quietly.** The daemon logs a warning and keeps
2637
+ running, but tier-1 escalations then land in issue comments — which is exactly the
2638
+ "nobody reads it until morning" path the orchestrator exists to avoid. The warning
2639
+ is in `daemon.log`; nothing pages you about it.
2640
+ - **Workers are not terminal panes, so you cannot watch them there.** Each
2641
+ worker is an omp session the daemon starts as a child process. The resident
2642
+ daemon tracks workers in a background pool so the five-minute loop keeps
2643
+ settling PRs and checking capacity; shutdown waits for that pool. Herdr still
2644
+ shows exactly one pane (the orchestrator's) regardless of concurrency.
2645
+
2646
+ The cap does work. The admission loop (`admitCandidates` in `src/daemon.ts`) computes
2647
+ `slots = maxConcurrentWorkers - live workers`, admits at most that many issues
2648
+ per tick, and dispatches them together. To see them, read `omp-conductor
2649
+ status`, which lists every occupied issue, or follow `daemon.log`.
2650
+ - **Report delivery is at-least-once, never exactly-once.** The Telegram Bot API
2651
+ takes no client-supplied idempotency key, so the window between "Telegram
2652
+ accepted it" and "SQLite recorded that" is irreducible. The daemon resolves it
2653
+ toward a duplicate — the report is retried and the retry says it may be a
2654
+ repeat — because a duplicate you can recognise by its report id is cheaper
2655
+ than a silently dropped page. `delivered` means Telegram accepted an attempt,
2656
+ not that exactly one message exists. See
2657
+ [Report delivery](#report-delivery-the-outbox).
2658
+ - **Workers stop at green PRs.** They are never authorised to merge, release or
2659
+ deploy: those actions default to a human, and while setup may grant either to
2660
+ the orchestrator, `authority` never grants them to a worker or the dispatch
2661
+ daemon. The verbs refuse a worker mechanically, and a worker reaching for
2662
+ another session's channel is refused too — each channel is bound to the pid the
2663
+ daemon spawned for it. What remains a prohibition rather than a gate is shelling
2664
+ out to `gh` directly: sessions inherit the daemon's credentials. That shows up as
2665
+ a mutation with no matching ledger entry, which is a mismatch an operator can
2666
+ find.
2667
+ - **The worker gate is partial, and the orchestrator has none.** A worker's
2668
+ structured `write` / `edit` / `read` / `grep` / `glob` calls are gated to its
2669
+ worktree by an inline harness extension; `bash` is not, so a shell one-liner
2670
+ can still leave the tree, and no claim in this README says otherwise. The
2671
+ orchestrator is [unconfined on purpose](#the-orchestrator-is-unconfined-deliberately)
2672
+ — its boundaries are its brief and the verb ledger. Prefer a
2673
+ [least-privilege worker uid](#least-privilege-worker-uid-deploy); the
2674
+ [integrity tripwire](#integrity-tripwire-package-self-hash) still pages if the
2675
+ installed package itself changes under a live daemon.
2676
+
2677
+
2678
+ ## License
2679
+
2680
+ MIT