omp-conductor 0.15.9 → 0.15.11

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (86) hide show
  1. package/README.md +273 -2543
  2. package/REFERENCE.md +2638 -0
  3. package/package.json +3 -2
  4. package/schema/config.schema.json +8 -23
  5. package/src/arm-challenge.ts +112 -0
  6. package/src/ask.ts +434 -0
  7. package/src/board.ts +81 -15
  8. package/src/brief-upgrade.ts +114 -8
  9. package/src/briefs/orchestrator.md +55 -29
  10. package/src/briefs/policy.md +14 -5
  11. package/src/briefs/worker.md +7 -1
  12. package/src/chain-check.ts +1 -1
  13. package/src/check-trailing-newlines.ts +82 -0
  14. package/src/cli.ts +190 -1391
  15. package/src/commands/arm.ts +21 -0
  16. package/src/commands/board.ts +23 -0
  17. package/src/commands/brief-upgrade.ts +186 -0
  18. package/src/commands/context.ts +49 -0
  19. package/src/commands/daemon.ts +71 -0
  20. package/src/commands/dashboard.ts +74 -0
  21. package/src/commands/decision.ts +103 -0
  22. package/src/commands/disarm.ts +21 -0
  23. package/src/commands/doctor.ts +98 -0
  24. package/src/commands/event.ts +62 -0
  25. package/src/commands/extend.ts +64 -0
  26. package/src/commands/friction.ts +56 -0
  27. package/src/commands/help.ts +9 -0
  28. package/src/commands/hold.ts +26 -0
  29. package/src/commands/intake.ts +134 -0
  30. package/src/commands/ledger.ts +69 -0
  31. package/src/commands/message.ts +48 -0
  32. package/src/commands/report.ts +170 -0
  33. package/src/commands/restart.ts +76 -0
  34. package/src/commands/resume.ts +58 -0
  35. package/src/commands/setup.ts +93 -0
  36. package/src/commands/start.ts +23 -0
  37. package/src/commands/stats.ts +131 -0
  38. package/src/commands/status.ts +48 -0
  39. package/src/commands/stop.ts +51 -0
  40. package/src/commands/tail.ts +109 -0
  41. package/src/commands/unblock.ts +39 -0
  42. package/src/commands/upgrade-install.ts +31 -0
  43. package/src/commands/upgrade-rollback.ts +23 -0
  44. package/src/commands/upgrade.ts +25 -0
  45. package/src/commands/verb.ts +83 -0
  46. package/src/commands/version.ts +30 -0
  47. package/src/commands/worker.ts +100 -0
  48. package/src/config-schema.ts +38 -1
  49. package/src/config.ts +10 -3
  50. package/src/daemon.ts +613 -94
  51. package/src/dashboard/app.js +120 -0
  52. package/src/dashboard/index.html +34 -0
  53. package/src/dashboard/server.ts +267 -0
  54. package/src/dashboard/style.css +180 -0
  55. package/src/decisions.ts +39 -14
  56. package/src/diff-flags.ts +131 -241
  57. package/src/doctor.ts +795 -0
  58. package/src/escalate.ts +60 -19
  59. package/src/failure-class.ts +29 -3
  60. package/src/fleet.ts +58 -1
  61. package/src/graph-health.ts +1 -1
  62. package/src/label-projection.ts +1 -1
  63. package/src/lifecycle.ts +198 -2
  64. package/src/notices.ts +9 -0
  65. package/src/omp.ts +2 -0
  66. package/src/orchestrator-tick.ts +315 -17
  67. package/src/release-policy.ts +135 -23
  68. package/src/reports.ts +19 -5
  69. package/src/setup-host.ts +420 -8
  70. package/src/setup-install.ts +69 -14
  71. package/src/setup-wizard.ts +199 -61
  72. package/src/setup.ts +131 -35
  73. package/src/stats.ts +331 -0
  74. package/src/store.ts +206 -21
  75. package/src/tracker/github.ts +27 -4
  76. package/src/types.ts +144 -31
  77. package/src/unblock.ts +55 -11
  78. package/src/upgrade-journal.ts +220 -0
  79. package/src/upgrade-verify.ts +506 -0
  80. package/src/upgrade.ts +295 -26
  81. package/src/verbs/actions.ts +73 -1
  82. package/src/verbs/protocol.ts +29 -4
  83. package/src/verbs/server.ts +183 -20
  84. package/systemd/omp-conductor-recover.sh +433 -0
  85. package/systemd/omp-conductor.service.example +7 -0
  86. package/systemd/recover-unit-test.sh +428 -0
package/REFERENCE.md ADDED
@@ -0,0 +1,2638 @@
1
+ # omp-conductor — Reference
2
+
3
+ The complete reference for `omp-conductor`, every surface and every internal. For
4
+ the 10-minute operator guide — what it is, install, quick start, operating the
5
+ fleet, and multi-project operation — read [`README.md`](README.md).
6
+
7
+ The sections below are the reference half of the documentation. Each guide
8
+ section links back here for the full detail it trims.
9
+
10
+ ## Install prerequisites
11
+
12
+ The guide's [Install](README.md#install) section links here. What the host needs
13
+ before anything runs:
14
+
15
+ `@oh-my-pi/pi-coding-agent` (`>=17.1.4`) is a **peer dependency** and must already
16
+ be present. If you run omp, it is.
17
+
18
+ Also required on the host:
19
+
20
+ - `bun`: the CLI and the daemon run on it (`Bun.serve` backs `/healthz`).
21
+ - **A model credential the *daemon's own account* can reach.** Sessions are child
22
+ processes of the daemon and inherit its environment and `$HOME` unmodified, so a
23
+ session authenticates with exactly what the daemon authenticates with — nothing
24
+ is injected and nothing is scrubbed. Any shape the harness itself understands
25
+ works, including the ordinary one:
26
+ - an **OAuth login** already recorded for that account (`omp` login state under
27
+ its `~/.omp/agent`). This is the common case and needs no configuration.
28
+ - a model **API key** in the daemon's environment (`ANTHROPIC_API_KEY`,
29
+ `OPENAI_API_KEY`, …). Note systemd starts the service with a clean
30
+ environment, so it has to be an `Environment=` line on the unit, not something
31
+ exported in your shell.
32
+ - the harness's **auth broker**, configured in that account's
33
+ `~/.omp/agent/config.yml`.
34
+
35
+ The account matters more than the shape: a login recorded under a *different*
36
+ account is invisible to the service. A unit running `User=fleet` cannot see
37
+ `root`'s login, and workers then die at turn 0 with `No model selected`. Such a
38
+ run is classified `env-start-failure` and charges **neither** the failure budget
39
+ nor a continuation — an environment fault is not a failed implementation, and a
40
+ run that recorded no turn, no commit and no error did not attempt anything — but
41
+ nothing dispatches successfully until the credential is reachable.
42
+ - `gh`, already authenticated: every tracker operation shells out to it, so the
43
+ daemon never handles a GitHub token itself.
44
+ - `git`: mirrors and worktrees.
45
+ - **[omp-telegram](https://www.npmjs.com/package/omp-telegram)**, for the
46
+ escalation channel. It is a separate package and is not vendored here.
47
+
48
+ Two different things depend on it, and they need different amounts of it:
49
+
50
+ - **Tier-2 paging** needs only its bot token. This package reads
51
+ `TELEGRAM_BOT_TOKEN` out of `$OMP_TELEGRAM_STATE_DIR/.env` (default
52
+ `~/.omp/agent/telegram/.env`) and posts to the chat id you configure. No
53
+ pairing required, and no token ever passes through this package's own config.
54
+ - **The interactive channel** — replying to an escalation, approving a brief
55
+ amendment from your phone — needs omp-telegram actually paired, which is what
56
+ writes `access.json`. The fleet heartbeat also reads that file and refuses to
57
+ tick unless exactly one owner is paired, on the grounds that unattended
58
+ dispatch is only defensible while a tier-2 page can reach a person.
59
+ - **Approving a Learning-loop amendment from a heartbeat tick** needs one more
60
+ setting than pairing: a notify destination. `/telegram notify` writes
61
+ `notifyMode` and `notifyChat` into the same `access.json`. omp-telegram
62
+ mounts its `telegram_ask` tool only for a turn that resolves a notify target,
63
+ and a locally injected tick resolves one only through that setting — so
64
+ without it the orchestrator can page you but cannot put a yes/no question in
65
+ front of you, which is the one thing the Learning loop's approval step
66
+ requires. `omp-conductor status` reports this on the `telegram` row, and a
67
+ tick that cannot ask says so in its own prompt and falls back to
68
+ `telegram_send`.
69
+
70
+ Set `notifyChat` even on a forum fleet. `/telegram topics` routes to a topic
71
+ this session claims at runtime, and that claim is not visible in
72
+ `access.json` — so a file carrying only `topicsChat` is reported as
73
+ unconfigured rather than guessed at, on the grounds that a health row which
74
+ reads green over a broken contract is worse than one that overstates a fault.
75
+
76
+ - **A fleet that answers instead of narrating** needs one key, set once:
77
+ `/telegram set profile daemon` (omp-telegram 0.11.0 or newer). Without it the
78
+ bridge behaves as it does on a laptop: it finalizes a real Telegram message
79
+ per assistant turn for as long as a conversation is active — so one answer
80
+ arrives as several messages, and a message that lands mid-tick keeps relaying
81
+ that tick's internal turns — and it posts every local run's closing text to
82
+ `notifyChat`, which on a host whose runs are heartbeat ticks means each tick's
83
+ working prose. The profile switches all of it off at the transport: text
84
+ reaches Telegram only through `telegram_send` / `telegram_ask`, the idle post
85
+ is suppressed, and `telegram_ask` stays mounted and aimed at the paired owner
86
+ on every turn — including a locally injected tick, so it also removes the need
87
+ for `notifyMode` above. Approval and blocked-input pings still fire; those
88
+ mean a human is needed, which is the point of the channel.
89
+
90
+ `omp-conductor status` reports an interactive profile on the `telegram` row,
91
+ and every tick composed on one carries a prompt line saying so. Note the two
92
+ settings pull against each other before the profile exists: setting
93
+ `notifyMode` to make `telegram_ask` mountable is exactly what arms the idle
94
+ post, so the correctly askable fleet was also the loud one.
95
+
96
+ With neither, tier 2 degrades to a comment on the issue. Nothing is broken in
97
+ that configuration: it is supported, just slower to reach you.
98
+
99
+
100
+ ## Updating
101
+
102
+ Run one command from a shell outside the target `herdr-fleet.service`:
103
+
104
+ ```bash
105
+ omp-conductor upgrade
106
+ ```
107
+
108
+ The fleet can also install its own fixes: with the `install` release shape
109
+ granted to the orchestrator (`omp-conductor setup authority`, gate **install**),
110
+ an orchestrator requests `conductor_install --arg version=X.Y.Z`. The daemon
111
+ refuses a version npm does not expose, then starts a transient systemd unit
112
+ (`omp-conductor-upgrade-install-<version>`) that runs the same transaction
113
+ outside the pane and the daemon, journaling every surface it touches to the
114
+ state-dir upgrade journal. The first tick after the restart reads that journal
115
+ and verifies independently — installed version, `/healthz`, ticks, pane,
116
+ `doctor` — then restores dispatch and reports through the durable outbox; any
117
+ gap triggers the detached rollback and a tier-2 page. The unit never declares
118
+ its own success, so a fleet is never left half-upgraded and quiet.
119
+
120
+ It resolves the latest published npm release, pauses new claims, drains active
121
+ workers, and pins that exact release across the Bun-global CLI, omp plugin, and
122
+ Herdr plugin. It also recomposes the conductor-owned brief floor, restarts Herdr
123
+ and the daemon, waits for pane recovery, verifies the installed identities and
124
+ layered fleet status twice, then restores the original dispatch state.
125
+
126
+ Use `omp-conductor upgrade --to X.Y.Z` for an explicit published version. The
127
+ command exits without changing anything when all three surfaces already use that
128
+ release, the Herdr plugin is pinned to its exact `gitHead`, and the brief is
129
+ current.
130
+
131
+ The upgrade is host-wide, because everything it replaces is: one daemon serves
132
+ every configured project, so its restart lands on all of them at once. A bare
133
+ `omp-conductor upgrade` therefore drains **every** project's workers and
134
+ refreshes **every** project's brief, and pauses them with the fleet-wide
135
+ sentinel — which leaves any per-project `hold` you set standing when it
136
+ restores dispatch.
137
+
138
+ `--project` narrows that only when it is truthful to do so. When a live daemon
139
+ recorded a single project, the command drains and restarts that project, and an
140
+ explicit `--project` naming a different one is rejected before pause or
141
+ installation. When the daemon serves every configured project and there is more
142
+ than one, `--project` is rejected too: draining one queue and then restarting
143
+ the shared daemon would kill another queue's workers without ever counting
144
+ them. Re-run without the flag.
145
+
146
+ Ticks remain in their existing armed or disarmed state, so an ordinary update
147
+ does not halt the exact pane or require another Telegram arm challenge. Progress
148
+ names the Bun-global CLI, omp plugin, Herdr plugin, brief, reloads, and both
149
+ verification passes separately.
150
+
151
+ An installation, brief, reload, or verification failure pauses dispatch and
152
+ attempts to restore the exact CLI/plugin identities that were present before the
153
+ command. If rollback also fails, the error names every failed restoration and
154
+ keeps dispatch paused; it never brings a known mixed fleet back into service.
155
+ The command never publishes npm, edits an install root, or delegates lifecycle
156
+ steps to an AI session. It refuses to run inside a Herdr-managed session because
157
+ an updater that restarts itself cannot verify the result.
158
+
159
+ After every upgrade — and after the initial install — run the mechanical host
160
+ verification:
161
+
162
+ ```bash
163
+ omp-conductor doctor
164
+ ```
165
+
166
+ It is read-only and exits 0 only when nothing failed. Each finding is one
167
+ deployment fault that has already cost a debugging session: expired `gh` auth
168
+ or scope gaps, case-mismatched routing/state labels (silent by design), the
169
+ installed systemd units drifting from the staged render, runtime-dir ownership,
170
+ config.json validation and backup freshness, sqlite `PRAGMA integrity_check`,
171
+ spend telemetry reading $0.00 on every completed run, invalid reporting
172
+ timezones, and Telegram delivery health. `--probe-telegram` is the one opt-in
173
+ side effect: one self-identified test message through the report transport.
174
+
175
+ ## Onboarding
176
+
177
+ `omp-conductor setup` is the whole of it. One command, in a plain terminal, doing
178
+ the two jobs onboarding has always had:
179
+
180
+ | Half | What it does |
181
+ | --- | --- |
182
+ | **The interview** | Asks the judgment no amount of repo reading produces, then writes it into `POLICY.md` as prose you own and can edit. |
183
+ | **The probes** | Reads your repos and *proposes* the rest — real CI gates, the project context, the release procedure — each a default you edit or a draft you confirm. |
184
+
185
+ The split matters because the two halves fail differently. A wrong config value is
186
+ a run that errors on the next tick; a wrong release boundary is a fleet that
187
+ publishes something at 03:00. The first is worth a validated prompt. The second is
188
+ worth being asked properly, which is why it is asked and never guessed.
189
+
190
+ ### What only you can answer
191
+
192
+ Always asked: **where the roadmap lives and what the current priority is.** A
193
+ tracker shows what is *open*, never what *matters*, and an orchestrator that cannot
194
+ rank work grooms by recency — which is how a stale issue outranks the thing you are
195
+ shipping this month.
196
+
197
+ Asked only when you grant the orchestrator a release shape, because a
198
+ humans-release fleet has no boundary to draw:
199
+
200
+ - **Where the orchestrator's leg ENDS**, in one sentence. If it cannot be said in
201
+ one sentence it is not a boundary, and a vague release mandate is what eventually
202
+ publishes something at 03:00.
203
+ - **What** may be released and from which branch; **when** — batched how, after
204
+ which *named* checks; **what proof** must be held first, results actually read
205
+ rather than an impression; **what must still be asked** every time; and **what
206
+ stays permanently forbidden**.
207
+ - **What makes a release worth cutting** — a sprint, an epic's children all closed,
208
+ N merged issues. Without it the orchestrator either releases per merge, a stream
209
+ of meaningless versions burning shared runners, or never releases at all.
210
+ - **Who owns the rollback.** Name a person and setup says so plainly: that person
211
+ already owns the release, so the honest configuration ends the agent's leg
212
+ *before* the irreversible step. It offers to move the boundary there; declining
213
+ is a choice, not a mistake.
214
+
215
+ Grant every release shape and setup pushes back once — credentials sitting in the
216
+ environment of a session that runs unattended for weeks, and a 03:00 rollback being
217
+ a judgement call under time pressure with partial information — then records what
218
+ you decide. It is your fleet.
219
+
220
+ ### What setup reads for you
221
+
222
+ Setup discovers factual defaults before it asks for them. `git remote get-url
223
+ origin` supplies the tracker repo and single-repo routing key; bounded `gh` calls
224
+ supply the default branch, existing queue/state labels, branch-protection checks,
225
+ environments, an unambiguous npm package name, and GitHub Projects/open milestones.
226
+ Each discovered value is shown with its evidence and remains editable at the same
227
+ prompt. Discovery only seeds a fresh interview: a re-run starts from the saved
228
+ project, so an operator-edited value is never guessed again.
229
+
230
+ Judgment and prose still belong to the existing confined model probes, and **every
231
+ answer is a proposal**:
232
+
233
+ - **Gates.** Reads each routing repo's CI workflows, `package.json` scripts and
234
+ `Makefile`/`justfile`, then pre-fills the [gates](#configuration) prompt with the
235
+ exact commands and the `cwd` each runs from, so your gates match what CI runs. It
236
+ reports the evidence it used, and an honest "this repo has no cheap local check"
237
+ is a real answer rather than an invented `npm test`.
238
+ - **Project context** and **the release procedure.** Drafted across *every* routing
239
+ repo — which repo owns which concern, which ship together, where the release
240
+ machinery actually lives — then shown to you in full and kept **only if you
241
+ confirm**. `POLICY.md` is re-read on every tick, so a paragraph you never read
242
+ would become an instruction the orchestrator follows all week.
243
+
244
+ A model probe has **no shell, no editor and no verbs**: it reads files and answers,
245
+ and a tool it was not given is refused rather than allowed. It is not a sandbox —
246
+ it runs as your own user and reads what you can read — which is why it is pointed
247
+ only at repos you configured yourself.
248
+
249
+ **Setup never fails because discovery or a probe did.** No `gh`, no auth or
250
+ network, a private repo, no omp peer, a clone failure, a timeout, or a malformed
251
+ reply each produces a warning and leaves the typed default in place. Every question
252
+ is still asked.
253
+
254
+ Skip the reading half entirely with `--no-ai`:
255
+
256
+ ```bash
257
+ omp-conductor setup --no-ai
258
+ ```
259
+
260
+ To fill in or revise just the brief later — the two `POLICY.md` sections above — run
261
+ the `brief` area, which re-asks the judgment questions and re-runs the probes:
262
+
263
+ ```bash
264
+ omp-conductor setup brief
265
+ ```
266
+
267
+ ### Changing one setting
268
+
269
+ `config.json` is wizard-written, so changing a value means running the wizard —
270
+ and a wizard that re-asks twenty questions to add one key is a wizard people edit
271
+ the file behind instead. So a re-run against a project that is already configured
272
+ opens with one question:
273
+
274
+ ```text
275
+ "platform" is already configured — what would you like to do?
276
+ > Change one area
277
+ asks one area's questions; every other answer is carried through from the saved config
278
+ Walk every question again
279
+ the full interview, every prompt pre-filled with what is configured now
280
+ Add another project
281
+ full interview for a new project; existing projects stay as they are
282
+ ```
283
+
284
+ Amending is the default. Pick it and the eight areas are listed with what each one
285
+ says right now, so the row you want is the row you can see:
286
+
287
+ ```text
288
+ Which area? Each row shows what it says now
289
+ tracker & repos — acme/platform, queue "ready-for-agent", "repo:" → platform, api, web, worker
290
+ gates — platform: bun run check; api: ruff check . @ backend; web: pnpm lint…
291
+ caps & worker model — 2 workers, 120 turns, 90m, $25/day, 2 attempts (all defaults) — harness default model
292
+ code graph — not configured — workers grep
293
+ authority — merge=orchestrator, release=orchestrator
294
+ escalation & triage — tier 2 pages Telegram 123456789, comments too, triage external
295
+ reporting scope — material — escalations, plus green PRs, second failures, and anything that stops the fleet
296
+ orchestrator brief — none at ~/.omp/conductor/worktrees/ORCHESTRATOR.md
297
+ ```
298
+
299
+ Only that area's questions are asked. Every other answer is read back out of
300
+ `config.json` and written again unchanged — the same answers, the same builder,
301
+ the same single confirm, so there is still exactly one thing in this package that
302
+ writes a config, and it still writes nothing before you agree. The consent screen
303
+ leads with the delta and then shows the whole project as it would be written:
304
+
305
+ ```text
306
+ amending code graph — project platform
307
+ was not configured — workers grep
308
+ now ~/.cache/conductor-graph/acme — 4 clone(s): platform, api, web, worker
309
+ carried over tracker & repos, gates, caps & worker model, authority, escalation & triage, reporting scope, orchestrator brief
310
+ read back from ~/.omp/conductor/config.json and rewritten unchanged
311
+ ```
312
+
313
+ A first run, or a project name this config has never seen, never sees either
314
+ question: there is nothing to amend, so it is the full interview exactly as
315
+ before. Choosing *Walk every question again* asks once to confirm the replace
316
+ (so a silent overwrite cannot happen from muscle-memory Enter) — every prompt
317
+ pre-filled with what is configured, Enter to keep it — with one wrinkle worth
318
+ knowing: the two authority confirms and the orchestrator-session confirm cannot
319
+ start on "yes", so Entering through the full interview **revokes** a delegation
320
+ rather than renewing it. Amending the `authority` area names the current grant in
321
+ the question, which is the safer way to leave one alone.
322
+
323
+ ### Adding another project
324
+
325
+ One daemon serves every configured project. To put a second (or third) fleet on
326
+ the same host without touching the first:
327
+
328
+ ```bash
329
+ omp-conductor setup --project second
330
+ # or pick "Add another project" from the re-run chooser
331
+ ```
332
+
333
+ That is a full interview for the new name only. Defaults land under
334
+ `~/.omp/conductor/projects/<name>/{worktrees,mirrors}` so two fleets never share
335
+ a cwd; the first project's existing flat `worktrees`/`mirrors` paths are never
336
+ migrated. A `workspaceRoot` that collides with another project is refused with
337
+ both names in the error. Re-using an existing name asks amend-or-replace before
338
+ anything is written.
339
+
340
+ After apply, setup provisions labels, brief, tick config (with `project` +
341
+ `agentName`), topic binding, smoke, and arm for the new project only, then prints:
342
+
343
+ - `omp-conductor restart --now` — the running daemon picks up the new project
344
+ only after reload (printed, not auto-run when workers are live)
345
+ - a copy-pasteable **herdr handoff**: `herdr --session conductor workspace create
346
+ --cwd <workspaceRoot> --label <project> --no-focus`, then `herdr --session
347
+ conductor agent start <project> --kind omp --pane <pane-id>` into that empty
348
+ pane (never into a live orchestrator), plus the `FLEET_CWDS` list for
349
+ multi-fleet recovery
350
+
351
+ `omp-conductor setup gates --project second` (and every other area) still amends
352
+ only that project.
353
+
354
+ ### Keeping a brief current
355
+
356
+ The standing prompt is three layers:
357
+
358
+ | Layer | File | Updates how? |
359
+ | --- | --- | --- |
360
+ | Package floor | `src/briefs/orchestrator.md` | Every tick recomposes it into `ORCHESTRATOR.md` from the installed package. Upgrade the package in this host's existing install root + restart is enough. |
361
+ | Shared policy | `$OMP_CONDUCTOR_HOME/SHARED_POLICY.md` (default `~/.omp/conductor/SHARED_POLICY.md`) | Optional, host-wide: one file applying to **every** project. Create it by hand beside `config.json`; it is re-read each tick like `POLICY.md`, so an edit binds the next heartbeat. The per-project `POLICY.md` overrides it where the two conflict. |
362
+ | Fleet policy | `POLICY.md` | Yours. Setup writes the scaffold once; the Learning loop edits only this file. |
363
+ | Composed view | `ORCHESTRATOR.md` | Regenerated from floor + `SHARED_POLICY.md` (when present) + `POLICY.md` on each tick (and at setup). Do not hand-amend it for durable policy. |
364
+ | Worker brief | `src/briefs/worker.md` | Read per run from the package. |
365
+
366
+ ```bash
367
+ omp-conductor brief-upgrade # report overlay / legacy state
368
+ omp-conductor brief-upgrade --migrate # dry-run: bannered ORCHESTRATOR.md → POLICY.md
369
+ omp-conductor brief-upgrade --migrate --apply
370
+ omp-conductor brief-upgrade --retrofit # #20: propose YOURS TO EDIT cut on a hand-written brief
371
+ omp-conductor brief-upgrade --retrofit --apply
372
+ ```
373
+
374
+ - **Overlay already active** (`POLICY.md` present): protocol updates need no brief-upgrade.
375
+ `brief-upgrade` reports which layers compose — floor, optional shared policy,
376
+ and `POLICY.md` — and the composed banner names all three when the shared one
377
+ exists, so a reader can tell which layer a paragraph came from.
378
+ - **Legacy bannered brief**: `--migrate` lifts the owned half into `POLICY.md`
379
+ and recomposes. Previous brief and policy versions go to
380
+ `$OMP_CONDUCTOR_HOME/backups/briefs/` (default
381
+ `~/.omp/conductor/backups/briefs/`), named with their source filename and
382
+ timestamp.
383
+ - **Hand-written brief** (no banner): `--retrofit` inserts the banner before the first Releases / Project context / Reporting / Amendments heading; then `--migrate`.
384
+ - **The legacy single-file merge is gone** (0.4.3). A bare `--apply` exits `2`
385
+ naming the two paths that remain, rather than rewriting a brief nobody asked
386
+ it to. `--migrate` is the cross-version ABI: an upgrade keeps calling it, and
387
+ an overlay fleet tolerates its absence because the floor recomposes each tick.
388
+ - **`status` names the layout**, so a fleet still on a legacy brief is visible
389
+ where an operator already looks: `brief overlay (package floor + POLICY.md)`,
390
+ or `brief legacy-bannered — run omp-conductor brief-upgrade`.
391
+ - **Existing sidecars**: a composed refresh relocates conductor-generated
392
+ `ORCHESTRATOR.md.bak-<timestamp>` and `POLICY.md.bak-<timestamp>` files into
393
+ that backup directory. Other `.bak` files stay untouched.
394
+
395
+ `--file PATH` checks a brief that is not where the wizard would have put it.
396
+
397
+ The **Learning loop** proposes diffs against `POLICY.md` for you to approve over
398
+ Telegram. It also learns from repeated operational friction. The daemon
399
+ automatically rolls up repairable admission holds; the orchestrator records
400
+ judgments code cannot make with:
401
+
402
+ ```bash
403
+ omp-conductor friction escalation-digest --detail "routine retry belonged in the digest" [--issue N]
404
+ omp-conductor friction report-noise --detail "green status repeated with no operator action"
405
+ omp-conductor friction report-surprise --detail "a material failure was missing from the report"
406
+ ```
407
+
408
+ Three observations within seven days make a bounded signal eligible for one
409
+ tick. After it is surfaced, that signal cools down for seven days. A signal is
410
+ evidence to investigate, never an automatic policy edit: the existing one-at-a-
411
+ time Telegram approval, `POLICY.md`-only edit, Hard-boundary prohibition, and
412
+ **Amendments** log still apply.
413
+
414
+ ## How one tick works
415
+
416
+ Per tick, for the daemon's project:
417
+
418
+ 1. **Verify and settle pushed PRs.** For every run in `pushed-pending`, repeat
419
+ the independent head/check verification; green → `pushed-green`, red →
420
+ `failed`, and still pending stays occupied. For every verified
421
+ `pushed-green` run, ask what became of its PR. Merged → `merged`; closed
422
+ without merging → `failed`. Unknown answers leave the row unchanged. Every row
423
+ that settles also loses its `agent:in-progress` label. This
424
+ maintenance runs even while dispatch is paused or workers are active, so
425
+ status converges on the five-minute tick cadence. It also runs above admission
426
+ so a row settled here frees its issue in the same tick. See
427
+ [what settles a green PR](#what-settles-a-green-pr).
428
+ 2. **Paused?** If the pause sentinel exists, the tick claims nothing and returns.
429
+ Settlement has already run, but no queue or admission work occurs. This makes
430
+ `omp-conductor hold` take effect without signalling the process.
431
+ 3. **List the queue.** Open issues in `tracker.repo` labelled `queueLabel`.
432
+ 4. **Filter and route.** An issue is eligible only if it carries the queue label
433
+ and none of the three state labels (`inProgress`, `blocked`, `failed`). Eligible
434
+ issues are partitioned into routable and unroutable.
435
+ 5. **Escalate the unroutable** at Tier 1, quoting the repo labels actually seen and
436
+ the configured repo names. These are never dispatched.
437
+ 6. **Check spend.** If spend since local midnight has reached `dailySpendUsd`, the
438
+ daemon **pauses itself**, pages at Tier 2, and returns.
439
+ 7. **Check capacity.** `maxConcurrentWorkers` minus *live* workers (runs in
440
+ `claimed` or `running`) gives the free slots. A pending or green PR occupies
441
+ its issue but not a slot: its worker is finished, and counting pushed PRs
442
+ would let two completed workers stop the fleet.
443
+ If no slot is free, the tick logs and returns.
444
+ 8. **Check the plan allowance.** If `caps.planUsage` names a window, the daemon
445
+ reads it (cached, see [Caps](#caps)) and holds *every* candidate under
446
+ `plan-usage-cap` when the window is at or over its threshold. Unlike the
447
+ spend cap this does **not** pause the daemon: the window resets on the
448
+ provider's clock, so dispatch resumes by itself.
449
+ 9. **Admit issues** up to the free slots, skipping any issue that already has
450
+ an active run — including a pending or green PR, so a second attempt cannot
451
+ land on a live PR. Repeated implementation failures consume
452
+ `maxAttemptsPerIssue`; cap kills, daemon orphans and answered blocks consume
453
+ the independent `maxContinuationsPerIssue`. Exhausting either escalates.
454
+ 10. **Ask the tracker whether the work already exists.** For each candidate that
455
+ survived step 9 — so at most one API call per free slot, never one per queued
456
+ issue — the daemon asks whether an **open** PR already closes it. An open PR
457
+ normally holds the issue. One narrow exception permits a routed continuation:
458
+ the latest run must be terminal, and the open PR must be that run's retained
459
+ work — either the PR URL it recorded or a PR opened on the branch it retained.
460
+ The branch half matters because a run can be cap-killed before its worker ever
461
+ opens a PR, leaving a retained branch and no recorded URL; a PR pushed to that
462
+ branch afterwards is still the continuation target. Drafts count because their
463
+ branch can hold the only copy of the work.
464
+ The tracker also finds work missing from a new, moved, restored, or cleared
465
+ store. If the check fails, the candidate is **held**, not admitted, and
466
+ retried next tick: the cost of holding is five minutes, the cost of admitting
467
+ on an unknown is a burned attempt and a duplicate PR. Only that candidate is
468
+ held, so a flaky API cannot stall the rest of the queue.
469
+ 11. **Record the pass.** Persist ready/routed/admitted counts and group every hold
470
+ under a stable reason code with at most five sample issue numbers. Tracker
471
+ failures mark the summary `DEGRADED`; capacity, sibling, open-PR and budget
472
+ holds remain normal policy state.
473
+ 12. **Dispatch** the admitted issues concurrently.
474
+
475
+ Then, per admitted issue:
476
+
477
+ 1. **Create the run row (`claimed`) — before any worktree or session exists.**
478
+ This ordering is the whole crash-safety story: the *row*, not a label, is the
479
+ guard against dispatching the same issue twice. It is local and written before
480
+ anything that can fail; if the process dies at any later point, the startup
481
+ orphan sweep marks the left-behind row `orphaned` and the orchestrator's drain
482
+ duty triages it (see below). The `agent:in-progress` label is a write-behind
483
+ projection of that row — enqueued in the same breath and applied to the tracker
484
+ by the projector with retry — so even a tracker that refuses the write cannot
485
+ recreate a duplicate PR while the guard is unavailable.
486
+ 2. Run the tick's post-admission projection pass, which applies freshly enqueued
487
+ label ops on the healthy path.
488
+ 3. Clear any stale tree for this issue, then add a fresh worktree at
489
+ `<workspaceRoot>/<issue>` cut from the bare mirror at `<mirrorRoot>/<repo>.git`,
490
+ on the run's branch off the repo's default branch.
491
+ 4. Allocate a session transcript under `<state dir>/sessions/`, one per attempt,
492
+ and move the run to `running`. The run record keeps the exact path and a
493
+ failure escalation quotes it, so you can read what the worker actually did.
494
+ 5. Run one omp session with the rendered brief, under the turn and wall-clock caps.
495
+ 6. Record the outcome:
496
+
497
+ | Outcome | Labels | Worktree | Escalation |
498
+ | --- | --- | --- | --- |
499
+ | `pushed-pending` | `agent:in-progress` stays while the daemon rechecks GitHub | removed | none |
500
+ | `pushed-green` | `agent:in-progress` stays while the PR is open | removed | none |
501
+ | `blocked` | swapped to `agent:blocked` | dirty tree committed to the branch, then removed | Tier 1 |
502
+ | `failed` / `killed` | swapped to `agent:failed` | dirty tree committed to the branch, then retained until the PR or issue is terminal | Tier 1 |
503
+ | unexpected error | swapped to `agent:failed` | same | Tier 1 |
504
+
505
+ `pushed-pending` and `pushed-green` are not the end of the row: later ticks
506
+ verify outstanding checks and settle the PR once it resolves, and a row that
507
+ settles gives up its `agent:in-progress` label. See
508
+ [what settles a green PR](#what-settles-a-green-pr).
509
+
510
+ Every label write the dispatcher makes goes through the label projection
511
+ outbox and is applied to the tracker by the projector with retry, strictly in
512
+ the order it was enqueued per issue — a later op for one issue never lands
513
+ before an earlier one that is still owed. A state-label swap on a dispatch
514
+ outcome enqueues the new label ahead of the old one's removal, so the issue
515
+ is never briefly bare (the shape eligibility reads as fresh work), while a
516
+ requeue (`swapToQueue`, `unblock`) enqueues its removals ahead of the queue
517
+ add for the same reason in reverse: the issue must not look claimable before
518
+ its stale state label is gone. A refused write is deferred with backoff
519
+ instead of dropped, and `status` shows any un-applied lag on a
520
+ `labels projection` row.
521
+
522
+ **Every continuable end salvages the tree first.** A turns-cap kill, a
523
+ wall-clock kill, a crash and a graceful block all leave a tree the next
524
+ attempt removes `--force` — only the run's *branch* is preserved across
525
+ attempts. So before the escalation is written, a dirty tree is committed to
526
+ the run's own branch as `wip(#<issue>): attempt <n> <ending> — auto-salvaged`
527
+ (everything, including files git has never seen) and pushed, and the
528
+ escalation says where it went: `WIP committed to <branch> @ <sha>`. A push
529
+ that is refused leaves the commit in this host's mirror and says so.
530
+
531
+ Blocking was excluded from this until #118, on the argument that a worker
532
+ which stops on purpose has turns left to commit for itself. It cost a
533
+ 34-file refactor: the worker blocked to ask whether a failing test was
534
+ obsolete — which is precisely a worker declining to commit a half-migrated
535
+ tree — and the daemon removed the tree seconds later, leaving the run branch
536
+ and `origin/main` on the same commit. A `pushed-green` or `pushed-pending`
537
+ run is now the only end that does not salvage: its deliverable is already on
538
+ a remote branch, and appending a WIP commit would turn the PR the daemon
539
+ just verified red.
540
+
541
+ **A salvage that fails keeps the tree and stops the issue.** If git refuses
542
+ the commit, the worktree is the only copy in existence, so it is retained
543
+ whatever the run's outcome was, the row records the failure, and the issue
544
+ is held out of dispatch with the `unsalvaged-wip` reason — because claiming
545
+ it is what would finally destroy the tree. `status` shows it under `wip` as
546
+ `UNSALVAGED`, and `omp-conductor unblock <n>` refuses. Recover the tree by
547
+ hand, then `unblock <n> --force` records that you accepted it and releases
548
+ the hold.
549
+
550
+ **A preserved tip is named to the next worker.** The sha is written to the
551
+ run row, shown by `status` and the board, and the continuation brief tells
552
+ the resuming worker the exact commit it is building on and that it is the
553
+ only copy.
554
+
555
+ Later ticks reap retained failure trees in bounded batches after the tracker
556
+ proves their PR merged/closed or their issue closed, provided no live run or
557
+ queued continuation owns the issue. Cleanup fetches remote refs first and
558
+ keeps any dirty tree or branch with uniquely local commits. Only then does it
559
+ remove the physical tree, prune registrations, and delete the obsolete local
560
+ mirror branch. Unknown tracker, network, repo, or git state is a no-op.
561
+
562
+ ### What a restart does to runs that were in flight
563
+
564
+ A `claimed` or `running` row is a promise that a worker process exists, and a
565
+ daemon that just started knows that promise is broken: its workers died with the
566
+ previous process. At startup — unless another daemon is alive, so a foreground
567
+ `daemon --once` cannot orphan a running daemon's real workers — every such row is
568
+ **salvaged first** (dirty tree → `wip(#N): attempt N killed by a daemon restart —
569
+ auto-salvaged` on the run's branch, same path as a turns-cap kill), then moved to
570
+ `orphaned`, with a log line naming the issue, the attempt and the worktree. That
571
+ frees the slots immediately; a fleet must never resume as deadlocked as it
572
+ crashed, and uncommitted edits must not wait for a human with `bun -e`.
573
+
574
+ Only the rows change after salvage. The issue keeps `agent:in-progress` — the
575
+ label is the crash guard against double-dispatch — and deciding what the dead
576
+ worker's remains are worth is the orchestrator's drain-duty judgement, spelled
577
+ out in its brief: an open green PR goes to the merge path, a salvaged sha is a
578
+ continuation hand-off, and a clean orphan has its label released so the next
579
+ tick re-claims it. Orphans consume `maxContinuationsPerIssue`, not failed
580
+ implementation attempts, so crashes cannot starve the retry needed for a real
581
+ code or CI failure — and a crash loop still escalates.
582
+
583
+ ### Deploying a new package onto a busy fleet
584
+
585
+ `systemctl restart` / `omp-conductor restart` is safe for **work product** once
586
+ this version is installed: startup salvage commits dirty trees before orphaning
587
+ rows, and salvage rewrites the mirror's managed `info/exclude` to the package's
588
+ current list before `git add` so a narrowed ignore cannot hide deliverables.
589
+
590
+ It is still disruptive for **in-flight sessions** because the worker process dies
591
+ and the attempt is spent. Update through the lifecycle command rather than
592
+ hand-installing or restarting individual surfaces:
593
+
594
+ ```bash
595
+ omp-conductor upgrade
596
+ ```
597
+
598
+ It pauses new claims, drains workers, installs one pinned release across all
599
+ surfaces, restarts, verifies twice, and resumes only if dispatch was initially
600
+ running. If an install, restart, or verification step fails, dispatch remains
601
+ paused and the command exits nonzero.
602
+
603
+ Do **not** edit files under the running install and expect the daemon to keep
604
+ dispatching — the integrity tripwire pauses and pages. Upgrade by whole release
605
+ so the new process records a fresh baseline.
606
+
607
+ ### What settles a green PR
608
+
609
+ The worker watches CI, then reports the PR URL and the exact remote head SHA it
610
+ observed. The daemon independently reads the PR again and requires it to be open,
611
+ non-draft, still at that head, and backed by a non-empty check rollup in which
612
+ every check succeeded or was skipped. Missing or nonterminal checks become
613
+ `pushed-pending` and are rechecked on later ticks; red or cancelled checks become
614
+ `failed` with a bounded job/log digest. Only verified evidence becomes
615
+ `pushed-green`.
616
+ If a failed report mentions PRs only in prose, the daemon retains the last URL
617
+ whose `owner/repo` matches the run's repository; links to other repositories are
618
+ ignored. This preserves the continuation target without trusting an unrelated
619
+ PR mentioned in the same report.
620
+
621
+ What happens after verification is a human's decision, taken minutes to days
622
+ later and never announced to the daemon — so every tick asks the tracker about
623
+ every pushed PR it is still holding:
624
+
625
+ | PR | Row becomes | Why |
626
+ | --- | --- | --- |
627
+ | merged | `merged` | The work landed. This is the state `merged` was reserved for. |
628
+ | closed without merging | `failed`, class `returned-for-revision`, with the PR in `lastError` | A human read the work and asked for another pass. Leaving it `pushed-green` strands the issue forever behind a PR nobody will merge, and calling it `merged` is a lie about code that is not on the base branch. `failed` is the honest row state; the class preserves the review decision and releases the issue so a re-queue can be attempted again. |
629
+ | still open | unchanged | The normal steady state. Its issue must stay occupied, or a second attempt lands on the live PR. |
630
+ | could not be determined | unchanged | A flaky network, a revoked token, a deleted PR. An unknown answer never settles a row; the next tick asks again for free. |
631
+
632
+ Run history is untouched. A PR closed without merging consumes one continuation,
633
+ not a failed implementation attempt; a merge spends neither budget. A settled
634
+ row also loses `agent:in-progress` from its issue: the row transition and label
635
+ removal are one fact, and a terminal answer about the PR proves no worker
636
+ process owns the issue,
637
+ so the duplicate-dispatch guard it exists for is spent. The removal is *enqueued
638
+ on the label projection outbox in the same breath as the terminal write* — a
639
+ durable local write that cannot fail on the tracker — so the row settles at once
640
+ and the projector applies the label with retry. A tracker that refuses the
641
+ removal can no longer strand it: the op stays owed, eligibility overlays the
642
+ pending removal so the issue is not held back by a label that is already decided
643
+ gone, and `status` shows the lag. Anything beyond that one release — a re-queue,
644
+ a `blocked` marker — is still the orchestrator's drain-duty judgement. One
645
+ unreachable PR costs its own row and nothing else; the rest of the sweep still
646
+ settles.
647
+
648
+ Until this existed, nothing ever revisited a `pushed-green` row: the startup
649
+ reconciler only settles rows that held a process, and `merged` went unwritten. On
650
+ 2026-08-07 the reference fleet reported three active runs whose PRs were all
651
+ merged and whose issues were all closed, through two daemon restarts — and because
652
+ the active set *is* the busy set, those three issues were permanently unclaimable.
653
+ A status page that has stopped being evidence is worse than no status page.
654
+
655
+ The label half of that outlived the row half by two days. On 2026-08-09 a merged
656
+ PR and a closed-unmerged one both settled their rows correctly and both left their
657
+ issues carrying `agent:in-progress`, which eligibility reads as "a worker owns
658
+ this" — with the brief forbidding the orchestrator from editing a state label and
659
+ `unblock` declining to clear that one, neither issue could ever be claimed again.
660
+
661
+ ### Base-branch health after merge
662
+
663
+ The daemon records two different facts after a merge:
664
+
665
+ - The **post-merge audit** attributes a regression to one merge. For up to 24
666
+ hours, it checks only push-triggered workflow runs for the exact merge SHA and
667
+ base branch. A new red result adds `base-branch-red` evidence and escalates.
668
+ - **Current health** drives `status` and release policy. On every sweep, the
669
+ daemon resolves the live head of each branch it merged into during the last
670
+ seven days, then reads only push-triggered runs for that head and branch. The
671
+ status row includes the head SHA and run count.
672
+
673
+ Current health is `green` only when every observed run completed successfully,
674
+ neutrally, or skipped. A failing conclusion is `red`; an in-progress or
675
+ unrecognised conclusion is `pending`; and a head with no push-triggered run is
676
+ `unknown`, never green. Pending and unknown heads are rechecked. Green and red
677
+ heads are read again when the branch moves, so an old verdict cannot describe a
678
+ new commit. If GitHub cannot return the head or its runs, the daemon keeps the
679
+ last honest row instead of replacing evidence with a network failure.
680
+
681
+ The `base-branch-green` release requirement reads this current live-head row for
682
+ the repository being released. Red, pending, unknown, and absent evidence all
683
+ refuse the release.
684
+
685
+ ### The settlement audit
686
+
687
+ Two lines of the worker report template were once taken on faith. `state:
688
+ pushed-green` stopped being believed in #85: a claim is now verified against
689
+ GitHub before a run settles. The `changed:` line stopped being *asked for* in
690
+ #488: the file list is derived at settlement from the pull request's own diff,
691
+ so a worker that writes no file list still settles with a correct one, and a
692
+ report whose list contradicts the diff is rewritten to say what the diff says.
693
+ The worker's narrative — what changed and why — is untouched; only the file list
694
+ stops being hand-authored.
695
+
696
+ So at settlement the daemon fetches the pull request's diff, derives the report's
697
+ `changed:` line from it, and checks the diff itself for weakened tests. What it
698
+ finds is a **settlement audit flag** — advisory, never a gate. A flagged run
699
+ settles exactly as an unflagged one does; nothing here can change a run's state,
700
+ hold a merge, or spend an attempt.
701
+
702
+ | Flag | Raised when |
703
+ | --- | --- |
704
+ | `changed-line-missing` | The settlement could not read the PR's diff, so no `changed:` file list was derived. The one disclosure fault left, and it is a tree-read failure, not a worker's. |
705
+ | `test-file-deleted` | A test file left the tree with no rename to account for it. |
706
+ | `test-disabled` | A `.skip` / `.only` / `xit` / `@pytest.mark.skip` / `t.Skip` marker appears on a line the PR added. |
707
+ | `assertions-removed` | An assertion was commented out, or a test file lost more assertions than it gained. |
708
+ | `test-timeout-raised` | A named timeout in a test file went up — compared against its own previous value, so a brand-new timeout is not a finding. |
709
+
710
+ A flag on a test file the dispatching issue never names is marked
711
+ `[unattributed]`: that is the "don't weaken tests you didn't write" case, and it
712
+ is the one worth reading first.
713
+
714
+ Flags reach you three ways: appended to the settlement report, stored on the run
715
+ row and shown under the run in `omp-conductor status` for as long as its PR is
716
+ open, and — once per flagged settlement — as a tier-1 escalation to the
717
+ orchestrator, whose brief says what judgement each flag invites.
718
+
719
+ **A clean, accurately reported PR produces nothing.** That is a design
720
+ constraint, not an aspiration: an audit that fires on honest work gets muted, and
721
+ a muted audit is worse than none because the fleet still believes it is being
722
+ checked. Every rule resolves ambiguity towards silence, and each accepts a named
723
+ blind spot to stay quiet — a renamed test file is not a deleted one (even when
724
+ git did not detect the rename, matched by basename), a `.skip` inside a string
725
+ literal or a recorded fixture is not a skip, and a rewritten test that keeps its
726
+ coverage is not a weakening.
727
+
728
+ The analyser is pure: it takes a parsed diff, the report and the issue text, and
729
+ returns flags. Only `Tracker.prDiff` touches the network, and a diff it cannot
730
+ read produces no flags *and says so* — silence about a diff nobody read is not a
731
+ clean bill.
732
+
733
+ ### Continuation runs
734
+
735
+ When a worktree is provisioned onto a branch that already exists in the mirror
736
+ (reattach after a prior attempt, orphan, or turns-cap auto-requeue), the worker
737
+ brief includes a **Continuation** section: read `git log` / `git diff` against
738
+ the default branch first, and do not recreate work already on the branch.
739
+
740
+ A **turns-cap kill with attempts remaining** salvages the tree, puts the queue
741
+ label back on, and skips the failed label so the next tick reclaims as a
742
+ continuation automatically.
743
+
744
+ ### Branch names
745
+
746
+ `<type>/<slug>`, where the type is `fix` when any label's last segment (after `:`
747
+ or `/`) is `bug`, and `feat` otherwise. The slug is the issue title folded to
748
+ `[a-z0-9-]`, and the whole ref is capped at 60 characters. It is computed from the
749
+ issue alone, so a retried run recomputes the same branch and finds its own work
750
+ instead of forking a second one.
751
+
752
+ ## Routing
753
+
754
+ An issue must carry **exactly one** `repo:<name>` label naming a repo in
755
+ `routing.repos`. The prefix is `routing.labelPrefix` and defaults to `repo:`.
756
+
757
+ Routing never guesses. An issue it cannot resolve to a single configured checkout
758
+ is handed back as unroutable:
759
+
760
+ | Reason | Condition |
761
+ | --- | --- |
762
+ | `no-repo-label` | The issue carries no label starting with the prefix. |
763
+ | `multiple-repo-labels` | It carries two or more distinct prefixed labels. A repeated identical label is deduplicated, not treated as an ambiguity. |
764
+ | `unknown-repo` | Its single prefixed label names a repo that is not in `routing.repos`. |
765
+
766
+ In all three cases the issue is **escalated at Tier 1 and never dispatched**. The
767
+ fix is always the same, and the escalation says so: put exactly one
768
+ `repo:<name>` label on the issue.
769
+
770
+ This is deliberate. A request that spans two repos, taken whole by one worker, is
771
+ the precise failure this guard exists to prevent: the worker cannot open a PR
772
+ against two checkouts, so it improvises — it vendors a copy, edits the wrong repo,
773
+ or produces a PR that cannot be merged without the other half. Splitting a
774
+ multi-repo request is a human decision about contracts; it is not something to
775
+ infer from a label. Sending the issue back costs a label edit; guessing costs a
776
+ bad merge.
777
+
778
+ ## Host sizing and memory
779
+
780
+ Workers are **child processes** of the daemon (plus one long-lived orchestrator
781
+ session), each its own pid, talking back over a unix socket. They are still
782
+ inside the service's cgroup, so systemd's Memory peak for
783
+ `omp-conductor.service` is daemon + every live worker + the orchestrator + any
784
+ MCP stdio children those sessions mount. `MemoryMax=` governs that whole total,
785
+ not one process.
786
+
787
+ The generated unit starts `omp-conductor daemon --port 8787` without a
788
+ `--project` filter, so one daemon serves every configured project. Its automatic
789
+ `MemoryMax=` tier uses the sum of resolved `maxConcurrentWorkers` values across
790
+ all projects: `3G` for a total of one worker, otherwise `5G`.
791
+
792
+ On the reference deploy that produced [issue #51](https://github.com/TerrifiedBug/conductor/issues/51):
793
+
794
+ | Shape | Observed |
795
+ | --- | --- |
796
+ | Idle / workers restarting | ~430 MB RSS for the daemon alone |
797
+ | Two workers + orchestrator, busy | **3.2–4.2 GB** Memory peak for the unit; up to ~800 MB swap |
798
+
799
+ That peak is **expected for concurrent SDK sessions**, not evidence of a
800
+ conductor-side leak: the SQLite store is disk-backed, admission state is
801
+ per-tick, and worker sessions are disposed when a run ends. What grows is the
802
+ session heap (conversation + tool output); a single graph-assisted run has been
803
+ measured in the hundreds of thousands of characters of tool output.
804
+
805
+ **Practical guidance**
806
+
807
+ - Prefer **≥16 GiB RAM** for the default `maxConcurrentWorkers: 2`, and do **not**
808
+ co-locate ClickHouse / other multi-GB services beside that fleet on an ≤8 GiB
809
+ box.
810
+ - On hosts under ~16 GiB, keep the **sum** of every project's
811
+ `maxConcurrentWorkers` at **1**. `omp-conductor setup` chooses that default
812
+ for a new project when it can read host RAM and warns before applying a
813
+ configuration whose combined capacity exceeds the host recommendation.
814
+ - Supervise the daemon with a unit that sets `SuccessExitStatus=0 143`; setup
815
+ renders `MemoryMax=3G` for one configured worker and `MemoryMax=5G` for two
816
+ or more.
817
+ - `omp-conductor status` prints daemon `rss` from `/healthz` when the process is
818
+ up, so you can see pressure without scraping journald.
819
+
820
+ ## Caps
821
+
822
+ Caps resolve per project: the global `defaults` block, then the project's own
823
+ `caps` layered on field by field, so a project that pins one cap still inherits the
824
+ rest. `0` is a real value (a hard stop), not "unset".
825
+
826
+ | Cap | Default | What it protects |
827
+ | --- | --- | --- |
828
+ | `maxConcurrentWorkers` | `2` (setup may write `1` on &lt;16 GiB hosts) | Parallel omp sessions, each a child process of the daemon and all inside its cgroup. Two, because **CI runner slots, not model tokens, are the usual throughput ceiling** — a third worker would starve its own PR checks on a small self-hosted runner pool. On hosts under ~16 GiB RAM, prefer `1` so the unit stays out of swap ([host sizing](#host-sizing-and-memory)). Raise it only if you actually have the runners *and* the RAM. |
829
+ | `maxConcurrentWorkersPerRepo` | `1` | Max live workers in the **same repo**. The mirror, branch-protection staleness and shared CI egress are all per-repo collision domains, so extra slots should land on other repos. Raise it only when a repo genuinely needs two workers at once. |
830
+ | `dailySpendUsd` | `25` | Rolling-day spend ceiling in USD, or `null` for no spend gate. `0` is a hard stop. Metered from assistant `usage.cost.total`. |
831
+ | `planUsage` | `null` (unmetered) | Subscription/plan allowance guard: `{ "windowId": "anthropic:7d", "maxUsedFraction": 0.85 }`, or `null` for no plan gate. Independent of `dailySpendUsd` — see [Plan allowance](#plan-allowance-planusage) below. |
832
+ | `workerMaxTurns` | `120` | Base ceiling for each new worker. Catches a session looping without converging; use `omp-conductor extend` to raise one live run or one issue's next attempt without changing this default. |
833
+ | `workerMaxTurnsCeiling` | `240` (twice the effective `workerMaxTurns` when omitted) | Upper bound for per-issue turn extensions. Prevents the loopback control from granting an unbounded worker budget. |
834
+ | `workerWallClockMs` | `5400000` (90 minutes) | Wall-clock ceiling for one worker. A session that is merely stuck spends no turns, so turns alone cannot detect it. |
835
+ | `maxAttemptsPerIssue` | `2` | Failed implementation or CI attempts before escalation. Operational stops do not consume this budget, so salvage can continue without stealing the retry needed for a real failure. |
836
+ | `maxContinuationsPerIssue` | `2` | Cap-kill, daemon-orphan and answered-block resumes before escalation. This independently bounds crash/resume loops. |
837
+
838
+ Days are counted from **local midnight**, matching how a human reads "today".
839
+
840
+ Set `dailySpendUsd` to `null` (wizard: blank) for no money gate — turns and wall-clock still apply. Hitting a numeric `dailySpendUsd` is not the same as hitting the other caps. A concurrency
841
+ limit simply defers work to a later tick. The spend cap **pauses the daemon and
842
+ pages at Tier 2**: a loop that is burning money has to halt itself, because
843
+ waiting for someone to notice tomorrow is how a runaway becomes expensive.
844
+ Work resumes only after `omp-conductor resume`.
845
+
846
+ `workerMaxTurns` and `workerWallClockMs` are enforced inside the session driver.
847
+ The daemon reads a live run's effective turn ceiling at every turn boundary. Use
848
+ `omp-conductor extend <issue> --turns N [--project NAME]` to raise it without
849
+ restarting or reconstructing the session. For a live worker, extension is
850
+ monotonic: equal or lower values are refused. If the latest run is failed,
851
+ killed, orphaned, or blocked, the command instead stores a one-shot ceiling for
852
+ that issue's next attempt. A next-attempt ceiling must exceed the effective
853
+ project base, and every extension must stay at or below
854
+ `workerMaxTurnsCeiling`. `status` shows both active ceilings and pending
855
+ next-attempt overrides. The store consumes an override atomically when it claims
856
+ the next run, so later attempts return to the project base. Config edits change
857
+ that base on the next tick but do not change workers already in flight. A cap
858
+ that fires aborts the run, records it as `killed`, and names the ceiling in the
859
+ escalation.
860
+
861
+ Pause one live worker cooperatively with
862
+ `omp-conductor worker pause <issue> [--project NAME]`. The daemon aborts the
863
+ active turn to an idle harness state, freezes the remaining wall-clock budget,
864
+ and keeps the run in the Running lane. `status` overlays `paused`/`pausing` from `/healthz` on that active-run line while the SQLite row stays `running`. `omp-conductor worker resume <issue>`
865
+ continues the same session with a prompt to re-check its last action before
866
+ proceeding. To end that run instead, use
867
+ `omp-conductor worker stop <issue> --reason TEXT [--project NAME]`. Stop works
868
+ from running or paused, salvages dirty work before removing the worktree, records
869
+ the distinct terminal `stopped` state, and removes `agent:in-progress` through
870
+ the label outbox. A salvage failure keeps the only copy in place and reports its
871
+ path. Stopped runs consume neither implementation-failure nor continuation
872
+ budget. Repeating stop reports the already-terminal state. These worker controls
873
+ are separate from fleet-level `pause`, which stops new claims.
874
+
875
+ ### Plan allowance (`planUsage`)
876
+
877
+ `dailySpendUsd` meters money, which is the only thing an API-billed account can
878
+ run out of. A fixed-price subscription cannot be expressed that way: the real
879
+ ceiling is a **provider allowance** — a weekly token window whose marginal
880
+ dollar cost is zero and whose exhaustion stops every session on the host.
881
+ Pricing that into the dollar meter would mean inventing a number.
882
+
883
+ `planUsage` is the second, independent guard. It reads `omp usage --json` — the
884
+ structured form of the harness `/usage` view — and holds new claims while the
885
+ named window is at or over its threshold. Running workers finish normally, and
886
+ **the daemon is not paused**: the window resets on the provider's clock, so
887
+ dispatch resumes by itself once a fresh reading is below the threshold. Nothing
888
+ estimates a plan quota from conductor's own transcript token counts.
889
+
890
+ ```json
891
+ "caps": {
892
+ "planUsage": { "windowId": "anthropic:7d", "maxUsedFraction": 0.85 }
893
+ }
894
+ ```
895
+
896
+ **Naming the window.** `limits` in the payload is a *list*, not a single
897
+ number: one Anthropic account reports `anthropic:5h`, `anthropic:7d` and the
898
+ tier-scoped `anthropic:7d:fable` at the same time, and other providers add
899
+ their own. So the cap names its window rather than taking whichever entry came
900
+ first. Run this on the fleet host and copy an `id`:
901
+
902
+ ```bash
903
+ omp usage --json | jq -r '.reports[].limits[] | "\(.id) \(.amount.usedFraction) \(.amount.unit)"'
904
+ ```
905
+
906
+ A bare window key (`"7d"`) also works, but **only** when exactly one reported
907
+ allowance carries it. On an Anthropic account `7d` matches two, and the guard
908
+ refuses to guess.
909
+
910
+ **`maxUsedFraction` is a fraction, not a percentage.** `0.85` holds at 85%.
911
+ A value outside `0`–`1` is rejected at config load, because `85` would mean
912
+ "hold at 8500% consumed" — a guard that reads as configured and can never fire.
913
+ Comparison always goes through the provider's `usedFraction`, never a raw
914
+ count: `unit` is `percent` for Anthropic and `unknown` with raw counts for
915
+ `xai-oauth`, so a threshold compared against `used` misreads any non-percent
916
+ provider by orders of magnitude.
917
+
918
+ **Availability policy.** The guard never displays a number it did not read, and
919
+ never shows a fabricated `0% used`. What each situation does:
920
+
921
+ | Situation | `status` / `board` | New claims |
922
+ | --- | --- | --- |
923
+ | `planUsage: null` | `unmetered` | admitted |
924
+ | Window below threshold | `5% / 85% of anthropic:7d used · resets in 6d 2h` | admitted |
925
+ | Window at or over threshold | same, plus `holding new claims` | **held** (`plan-usage-cap`, Tier 1) |
926
+ | No provider reports a readable allowance, or `omp usage --json` fails | `unavailable — <reason>` | admitted for up to 30 minutes, then **held** and paged at Tier 2 |
927
+ | `windowId` names a window the reading does not contain | `window "<id>" is not in this reading — Reported: …` | **held**, Tier 2 |
928
+ | `windowId` matches more than one allowance | `window "<id>" matches …` | **held**, Tier 2 |
929
+ | The window reports nothing a fraction can be derived from | `window "<id>" reports no comparable fraction …` | **held**, Tier 2 |
930
+
931
+ The split is deliberate. A *read error* is transient — a token refresh, a
932
+ provider 502, `omp` briefly absent mid-upgrade — and stalling a fleet on one
933
+ would cost more than admitting through it, since spend, turns, wall clock and
934
+ concurrency are all still enforced. Half an hour of continuous failure is not
935
+ an outage, it is a broken meter, and a plan-capped fleet running on a broken
936
+ meter is how the allowance gets spent to zero unnoticed. A *successful* read
937
+ that does not contain the configured window is not a read error at all: the
938
+ source answered, and it says the config names something that is not there. That
939
+ fails closed immediately, like every other config fault in this package, and
940
+ recovers by itself as soon as a reading contains the window again.
941
+
942
+ Readings are cached for 60 seconds (15 for a failure) so one tick costs one
943
+ provider call rather than one per candidate, and a cached reading is dropped
944
+ the moment its own `resetsAt` passes — that is what makes admission resume at
945
+ the rollover instead of a TTL later. `omp usage invalidate` clears omp's own
946
+ cache; conductor picks the change up at its next read.
947
+
948
+ Both controls are shown separately, never folded together — `status` prints a
949
+ `spend today` row and a `plan usage` row, and the board's admission line ends
950
+ with `spend $2.40/$25.00 | plan 5%/85%`.
951
+
952
+ ## Worker model
953
+
954
+ `workerModel` on a project pins the model its workers run on, as a pattern in
955
+ omp's own model/role syntax (whatever `/model` accepts). It sits beside `caps`
956
+ rather than inside them, because it is not a ceiling:
957
+
958
+ ```json
959
+ "workerModel": "smol"
960
+ ```
961
+
962
+ Omit it and the harness picks, which is the right answer until you have a reason.
963
+ The pattern is passed through unresolved: omp resolves it after its extensions
964
+ load, so a name this package has never heard of still works. If the harness cannot
965
+ honour the pattern it says so, and the daemon logs that per run:
966
+
967
+ ```text
968
+ #412 model fallback: <what the harness substituted>
969
+ ```
970
+
971
+ Worth reading the log for. A run that quietly used a weaker model than you chose
972
+ otherwise looks like a run that was merely unlucky.
973
+
974
+ ## Code-graph discovery
975
+
976
+ Optional, off unless you answer yes in the wizard, and worth answering yes to for
977
+ one measured reason: **workers spend most of a run finding code, not changing it.**
978
+ On the dogfood fleet a single run typically spends 30–62 `read` calls and 32–69
979
+ `bash` calls against 9–24 edits — 215–390k characters of tool output, roughly four
980
+ fifths of a 120-turn budget — and the runs that died at the turns cap died with
981
+ the work unfinished. A code graph answers "who calls this" and "where is this
982
+ defined" in one call instead of twenty greps.
983
+
984
+ ### Two things this package does not do for you
985
+
986
+ `omp-conductor` never installs, starts, imports, or depends on the indexer for
987
+ dispatch. With `graphProject` unset, nothing about dispatch, caps, escalation, or
988
+ status changes. A fresh host needs both of these before an index is worth
989
+ anything, and `setup graph` reports them as step 0:
990
+
991
+ 1. **`codebase-memory-mcp` on PATH** — a separate project,
992
+ [DeusData/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp).
993
+ 2. **Mounted as an MCP server** in `~/.omp/agent/mcp.json`, on the account the
994
+ daemon runs as. Miss this and the failure is silent: every index builds
995
+ correctly, no worker session can read any of them, so workers fall back to
996
+ grepping and the feature looks like a no-op. `setup graph` prints the entry.
997
+
998
+ Say yes and the wizard asks for one root, then derives one clone per routed repo
999
+ underneath it (default `~/.cache/conductor-graph/<org>/<repo>`) and writes it to
1000
+ each repo's [`graphProject`](#configuration). The only automatic interaction is
1001
+ a bounded, read-only health query; this package never clones, fetches, builds an
1002
+ index, or changes systemd. Dispatch, caps and escalation do not depend on graph
1003
+ health. On a fleet configured before this key existed, `omp-conductor setup` and
1004
+ the `code graph` area add it in two prompts — see
1005
+ [Changing one setting](#changing-one-setting).
1006
+
1007
+ ### Why the clone, and not your checkout or the worktree
1008
+
1009
+ This is the part that decides whether the feature helps or hurts, so it is worth
1010
+ being blunt about all three candidates.
1011
+
1012
+ | Directory | Why not |
1013
+ | --- | --- |
1014
+ | **The worker's worktree** | An index is keyed by the realpath of the directory it was built from, and has no git-worktree awareness. A run's `worktrees/<issue>` path is therefore *always* an empty project — a worker that queried its own cwd would get silence, conclude there is no graph, and spend the run grepping. This is why `graphProject` is an absolute path in the config and not something derived at run time. |
1015
+ | **Your own checkout** | Refreshing an index means resetting the clone to its default branch. In a directory you work in, that either destroys uncommitted work or — if it is made safe instead — indexes whatever feature branch you left checked out, so the fleet orients against your WIP. |
1016
+ | **A conductor mirror** | The daemon's mirrors are bare. There is no working tree to index. |
1017
+
1018
+ So `graphProject` names a fourth thing: a clone that exists only to be indexed,
1019
+ that nothing human ever edits, and that is therefore safe to `git reset --hard`
1020
+ every night. The worker brief names that path, tells the session to match it
1021
+ against `list_projects`' `root_path` and query by the `name` beside it, and says
1022
+ plainly that the graph is a snapshot which does **not** contain the worker's own
1023
+ edits — orient with it, then read the real file before changing it.
1024
+
1025
+ ### Creating and refreshing them
1026
+
1027
+ ```bash
1028
+ omp-conductor setup graph --print # print the plan: clones, index commands, units
1029
+ omp-conductor setup graph # run it: clone, install, enable, seed, verify
1030
+ ```
1031
+
1032
+ `setup graph --print` prints a `git clone` for every clone that does not exist yet, the
1033
+ one-shot index command per repo, and a `cbm-reindex.service` + `cbm-reindex.timer`
1034
+ pair built from the project's own repos and branches. `--write` stages all three
1035
+ in the state directory and prints the two `sudo` lines that install and enable
1036
+ them; it never runs `systemctl`.
1037
+
1038
+ **Run it as the account the fleet runs as, never under `sudo`** — it refuses if
1039
+ you try. Everything it derives resolves per-account: the config it loads, the
1040
+ state directory it stages into, and the `HOME`/`User=` it bakes into the unit.
1041
+ Under root you get a timer that goes green while writing indexes into
1042
+ `/root/.cache`, where no worker session looks — silent, and indistinguishable
1043
+ from the feature simply not helping. Only installing the units needs root, which
1044
+ is why that is two separate printed commands.
1045
+
1046
+ Two properties of the generated unit are deliberate:
1047
+
1048
+ - **It is a timer, not the server's own watcher.** That watcher lives inside a
1049
+ connected MCP session and dies with it, so an ephemeral worker session keeps
1050
+ nothing fresh. The refresh has to come from outside the fleet.
1051
+ - **It fails loudly.** The refresh is `set -euo pipefail`, then per repo
1052
+ `git fetch --prune origin` and `git reset --hard origin/<its own defaultBranch>`
1053
+ before indexing. Nothing is `|| true`-ed, so a fetch that has been broken for a
1054
+ week turns the unit red instead of quietly re-indexing a stale tree and exiting
1055
+ `0` — a green timer serving a month-old graph is worse than no graph at all.
1056
+
1057
+ The unit spells out `HOME` and an explicit `PATH`, because systemd supplies
1058
+ neither usefully: the indexer resolves its store from `HOME`, systemd's default
1059
+ `PATH` has no `~/.local/bin`, and the indexer shells out to `git`. Both are the
1060
+ user that ran `setup graph`; the unit sets no `User=`, so check them if that is
1061
+ not the account the timer runs as.
1062
+
1063
+ ### Seeing whether the graph is usable
1064
+
1065
+ When at least one routed repo has `graphProject`, `omp-conductor status` adds a
1066
+ `code graph` block. It proves the indexer is on `PATH`, the worker MCP config
1067
+ mounts it, every configured clone exists and exactly matches an indexed
1068
+ `root_path`, the refresh timer is enabled and active, and the last service run
1069
+ succeeded within 45 minutes. A running daemon refreshes this evidence every
1070
+ minute and publishes the cached result through `/healthz`; status probes the host
1071
+ directly when that cache is unavailable. Every command is read-only, runs with a
1072
+ one-second timeout, and graph degradation never changes `/healthz.ok` or blocks
1073
+ dispatch. Unconfigured projects omit the block entirely.
1074
+
1075
+ ## Escalation tiers
1076
+
1077
+ | Tier | Meaning | Raised by | Delivered to |
1078
+ | --- | --- | --- | --- |
1079
+ | 1 | "Not a human's problem yet" — the run is parked and safe. | Unroutable issue, blocked run, failed or killed run, dispatch error, attempts exhausted. | The orchestrator session, as an injected prompt. Falls back to an issue comment when no orchestrator is running, or when it will not accept the injection. |
1080
+ | 2 | "The fleet is stopped until you look." | Daily spend cap reached; the installed package changed under the running daemon. Either way the project is already paused. | Telegram, when `escalation.telegramChatId` is set and a bot token is readable; otherwise it falls back to the issue comment. |
1081
+
1082
+ **The orchestrator** is one persistent, file-backed session per daemon run, resumed
1083
+ across restarts so it remembers what it has already handled. Its `cwd` is the state
1084
+ directory, deliberately not a checkout. Delivery resolves when the harness *accepts*
1085
+ the prompt, not when the model answers it, so a tick never parks behind a model; an
1086
+ injection arriving mid-thought queues as a follow-up instead of interrupting the
1087
+ turn in flight. Its standing orders are explicit: re-brief the worker, file or
1088
+ comment on issues, or promote to tier 2, and never edit product code or push a
1089
+ branch. Merging is the one line worded from config — see [`authority`](#configuration).
1090
+ If it fails to start, the daemon logs a warning and runs on, with tier-1
1091
+ escalations degraded to issue comments.
1092
+
1093
+ **Or no orchestrator at all.** Set `escalation.orchestrator` to `"external"` when
1094
+ you already run your own supervising session — a visible TUI session in a pane,
1095
+ typically. The daemon then starts none of its own and every tier-1 escalation
1096
+ posts as an issue comment, which is what that session drains. One brain, and it
1097
+ is the one you can watch.
1098
+
1099
+ **Answering a tier 1 is only half of it.** A blocked or failed run leaves its state
1100
+ label on the issue, and eligibility reads any state label as disqualifying, so an
1101
+ answered issue that keeps one is never re-claimed and the answer is inert — nothing
1102
+ fails, the issue just stops existing as far as dispatch is concerned.
1103
+ [`omp-conductor unblock <issue>`](#cli-reference) is the way back: it clears the
1104
+ label through the same tracker the dispatcher writes with, including
1105
+ `agent:in-progress` when the newest recorded run is terminal, since a terminal row
1106
+ is proof the worker process is gone. The brief tells the
1107
+ orchestrator to run that verb rather than edit the label itself, and that is not a
1108
+ formality — orphan detection works by comparing `agent:in-progress` labels against
1109
+ live runs, and it is only trustworthy while every state label on the tracker was
1110
+ written by this package.
1111
+
1112
+ Tier 2 borrows the bot token that `omp-telegram` already owns, at
1113
+ `~/.omp/agent/telegram/.env` (or `$OMP_TELEGRAM_STATE_DIR/.env`). If you run that
1114
+ bot, Tier 2 needs no extra configuration beyond the chat id. If the token is
1115
+ absent, Tier 2 degrades to the issue comment instead of failing. The token is
1116
+ never logged, and it is redacted out of any error text that could reach a public
1117
+ issue comment.
1118
+
1119
+ **Escalations are deduplicated.** The dispatcher re-notices the same unroutable
1120
+ issue on every poll, so a ledger in the store — keyed by project, issue, tier and
1121
+ summary — makes a recurring condition page **once** and suppresses the five-minute
1122
+ repeats. The marker is recorded only on successful delivery, so a page that could
1123
+ not be delivered is retried on the next tick instead of being written off as sent.
1124
+ The spend-cap and integrity-tripwire summaries carry the date, so the same
1125
+ condition pages again tomorrow but only once per day.
1126
+
1127
+ If `fallbackToIssueComment` is off and no Telegram transport is configured,
1128
+ delivery throws instead of dropping silently. The failure is logged and retried,
1129
+ because a swallowed escalation looks exactly like a healthy fleet.
1130
+
1131
+ ## Report delivery (the outbox)
1132
+
1133
+ Escalations are the daemon's. **Reports** — the material events and the daily
1134
+ digest your [`reporting.scope`](README.md#your-workflow-vs-the-package) asks for — are
1135
+ written by the orchestrator, and until v0.3.26 they were also *delivered* by it:
1136
+ a report reached you only if the model remembered to call `telegram_send`. On
1137
+ 2026-08-06 a suite release and two tier-2 escalations were written that way and
1138
+ none of the three arrived, and nothing anywhere recorded that fact — an undelivered
1139
+ report and a quiet tick look identical.
1140
+
1141
+ ### Material events survive the session
1142
+
1143
+ A deferred digest does not use the session transcript as its source of truth.
1144
+ Record each ordinary outcome when it happens:
1145
+
1146
+ ```bash
1147
+ omp-conductor event record \
1148
+ --category merge \
1149
+ --summary "#42 merged" \
1150
+ --evidence "https://github.com/acme/api/pull/42"
1151
+ ```
1152
+
1153
+ `--category` is a short lowercase slug. `--summary` states the outcome, and
1154
+ `--evidence` names the issue, PR, release, run, commit, or URL that proves it.
1155
+ Use `--occurred-at <ISO timestamp>` when the event happened earlier; otherwise,
1156
+ the command uses the current time. The command writes one row to SQLite and
1157
+ sends nothing. The row survives later ticks, session compaction, session
1158
+ replacement, and daemon restarts.
1159
+
1160
+ When a digest is due, its tick prompt lists a bounded, oldest-first set of
1161
+ owed material events and deferred escalations. Each line includes its ledger id.
1162
+ The prompt gives the exact handoff shape:
1163
+
1164
+ ```bash
1165
+ omp-conductor report \
1166
+ --kind digest \
1167
+ --events EVENT_ID_1,EVENT_ID_2 \
1168
+ --notices NOTICE_ID_1,NOTICE_ID_2 \
1169
+ --text "<the whole digest>"
1170
+ ```
1171
+
1172
+ Remove the id of any row you did not use. Omitted rows stay owed. The report row
1173
+ and the named ledger rows are associated in one SQLite transaction. If the
1174
+ handoff fails, no row is consumed. If a daily report deduplicates against a
1175
+ daily report already queued that day, newly named rows also stay owed. If
1176
+ delivery exhausts its retry budget and the report becomes `failed`, its rows
1177
+ return to the owed backlog, where a replacement digest can claim them.
1178
+ `omp-conductor status` always shows the
1179
+ material-event and held-escalation backlog counts, including the age of the
1180
+ oldest row when one exists.
1181
+
1182
+ This accumulator does not poll GitHub and does not infer outcomes from tracker
1183
+ state. The orchestrator still decides what is material and records the evidence.
1184
+ The mechanism only makes that decision durable until a non-failed digest owns it.
1185
+
1186
+ Authorship still needs judgement the daemon does not have, so it stays with the
1187
+ model. Delivery does not, so it moved:
1188
+
1189
+ ```bash
1190
+ omp-conductor report --text "<the whole report>" # immediate report, when policy permits
1191
+ omp-conductor report \
1192
+ --text "<the whole digest>" --kind digest \
1193
+ --events EVENT_IDS --notices NOTICE_IDS # use row ids from its tick
1194
+ ```
1195
+
1196
+ The command persists the text before anything is sent and prints a durable
1197
+ handoff id. An immediate report admitted during quiet hours is stored as a held
1198
+ notice and prints that id; the daemon includes it in the next digest or in a
1199
+ catch-up report when the configured window opens. Otherwise it writes a
1200
+ `reports` row and prints its report id. Both survive the session being
1201
+ compacted, interrupted or restarted,
1202
+ and the daemon being restarted under it. The daemon delivers over the same bot
1203
+ token tier 2 uses, with bounded retries, and `omp-conductor status` lists
1204
+ anything it still owes. If availability closes after a material report was
1205
+ queued but before its first attempt, the outbox atomically converts that row to
1206
+ the same held-notice path instead of leaking the update through quiet hours.
1207
+
1208
+ ### Answering a person, in the thread they wrote in
1209
+
1210
+ A report is an update; an answer is a conversation, and it goes back where the
1211
+ question came from. `telegram_send` keeps the active forum topic **only while it
1212
+ names no chat** — `thread_id` defaults to the active topic when `chat_id` is
1213
+ omitted — so an orchestrator that helpfully supplied `chat_id` (and nothing
1214
+ else) answered three topic messages in the main chat instead (#366). The
1215
+ orchestrator floor now says to name **neither** target or **both**, and the
1216
+ dispatcher refuses `chat_id` without `thread_id` for a project that configured
1217
+ `escalation.telegramTopicId`. A flat-chat project is unaffected. When a pinned
1218
+ topic id has gone stale — the bridge re-claims pane topics across restarts — the
1219
+ live claim for the project's herdr space is used instead, falling back to a
1220
+ claim titled for the project, so a restart does not quietly move every page into
1221
+ the main chat (#407, #412).
1222
+
1223
+ A locally injected tick has no inbound message to inherit a topic from, so a
1224
+ bare `telegram_send` there has nothing to preserve. That turn addresses the
1225
+ operator from config instead:
1226
+
1227
+ ```bash
1228
+ omp-conductor message --text "<the message>" # this project's chat and topic
1229
+ omp-conductor message --text "QUESTION: cut 0.16.0 tonight?" # carries the decision category
1230
+ ```
1231
+
1232
+ It is not a bypass of the interrupt policy: the same availability decision an
1233
+ autonomous Telegram tool call gets is applied, so a message the policy defers is
1234
+ durably held for the digest or the working-hours catch-up and the command prints
1235
+ that held-notice id instead of claiming delivery. It is also not a report — it
1236
+ leaves no `reports` row, and nothing retries it.
1237
+
1238
+ ### Delivery is at-least-once, and the docs will not pretend otherwise
1239
+
1240
+ The Telegram Bot API accepts no client-supplied idempotency key and offers the
1241
+ bot no readable record of what it has already sent. There is nothing to replay a
1242
+ request against and nothing to reconcile with, so **exactly-once delivery cannot
1243
+ be built on this transport** and this package does not claim it. `delivered` is
1244
+ proof that Telegram accepted *an* attempt, never proof that exactly one message
1245
+ exists.
1246
+
1247
+ What it does instead is make the ambiguity explicit and always resolve it in the
1248
+ direction of the duplicate:
1249
+
1250
+ | State | Meaning | What you do |
1251
+ | --- | --- | --- |
1252
+ | `pending` | Nothing is in flight. Never attempted, or the last attempt failed **definitively** — see below. Nobody has this report. | Nothing. It retries on a bounded backoff (30s doubling to a 15-minute floor) and `status` shows the error. |
1253
+ | `sending` | A request left this host and its outcome was never learned: the daemon died, or the request was cut off after the bytes went out. Telegram may be holding the message. | Nothing, but expect a possible duplicate. Check the chat if you want to know now. |
1254
+ | `delivered` | Telegram answered `ok: true`. The row records the message id from the response **body**, not the HTTP status. | Nothing. |
1255
+ | `failed` | The retry budget ran out — six attempts, roughly half an hour. | Fix the transport. This state pages tier 2 in its own right — see below. |
1256
+
1257
+ The row is written `sending`, with the id of the attempt about to run, **before**
1258
+ the request is made. A crash in that window therefore leaves an explicitly
1259
+ ambiguous row rather than a silently lost one. The next daemon start sweeps every
1260
+ `sending` row, retries it, and the retried message carries the report id and a
1261
+ plain-English line saying it may already be in the chat. A duplicate you can spot
1262
+ by its report id is much the cheaper of the two mistakes; a silently dropped
1263
+ report is the entire reason this exists.
1264
+
1265
+ #### Two kinds of failure, and only one of them is quiet
1266
+
1267
+ A failed send is classified where the socket is watched, not by the caller, and
1268
+ the two classes are treated differently on purpose:
1269
+
1270
+ | Outcome | What happened | Row | Retry says |
1271
+ | --- | --- | --- | --- |
1272
+ | **Definitive** — nobody has it | Telegram answered and refused it (`{"ok":false}` under any status, or a non-2xx status), or the connection never opened at all (refused, DNS failure) so the request provably never left. | back to `pending`, backoff, attempt counted | nothing special — it *is* a first attempt |
1273
+ | **Outcome unknown** — Telegram might have it | The request was cut off after it left: timeout, abort, socket reset, `EPIPE`. Or the POST came back `200` and the **response body could not be read** — Telegram had already decided and the answer was lost coming back. | stays `sending`, flagged as a possible repeat | `POSSIBLE REPEAT`, with the report id to compare against |
1274
+
1275
+ Anything that cannot be classified confidently is treated as **outcome unknown**.
1276
+ That default is deliberate and is the safe direction: the worst case is a
1277
+ duplicate you were warned about, against a delivered report re-posted as though
1278
+ it were new, with nothing anywhere saying it might be a second copy.
1279
+
1280
+ The half of "never double-post" that *is* achievable is enforced: a report cannot
1281
+ be **concurrently** in flight twice. Claiming a report is a conditional update,
1282
+ so only one attempt can move a `pending` row, and every terminal transition names
1283
+ the attempt it is settling — a request that answers after its row was reclaimed
1284
+ is discarded rather than allowed to overwrite a newer attempt's outcome. That is
1285
+ what stops a retry storm.
1286
+
1287
+ ### A report nobody can deliver is itself news
1288
+
1289
+ A report that exhausts its retries is marked `failed` **and escalates as tier 2**.
1290
+ This rides the transport that just failed, which is deliberate and accepted: the
1291
+ common failure is a wrong chat id or a bot kicked from the chat, not a global
1292
+ Telegram outage, and in both of those the page reaches an operator who is
1293
+ otherwise being told nothing at all. If the whole channel is down the page
1294
+ degrades to a line in `daemon.log` and the `reports` block in `status`, which is
1295
+ then the only surface — a report has no tracker issue, so there is no issue
1296
+ comment to fall back to. The page goes through the ordinary escalation ledger and
1297
+ carries the report id, so one undeliverable report pages exactly once.
1298
+
1299
+ ### Daily digests are deduplicated from the ledger
1300
+
1301
+ With `digest.cadence: "daily"`, `--kind digest` is accepted at most once per
1302
+ **local** day, per project. The second hand-over on the same day is refused and
1303
+ told which report already holds the slot, including when that report has already
1304
+ been delivered. This is decided from the `reports` table, not from the model's
1305
+ memory of the last tick — a restarted or compacted session cannot send a second
1306
+ daily digest by forgetting the first. A `per-tick` digest carries no daily key,
1307
+ so later ticks can hand off newly accumulated rows. Material reports carry no
1308
+ dedupe key either: two events in a day are two events.
1309
+
1310
+ ### What `status` shows
1311
+
1312
+ ```text
1313
+ reports 1 pending · 1 sending · 0 failed (delivery is at-least-once — a retry may duplicate)
1314
+ 9f2c1ab0d3e4 pending material 12m old attempt 2/6, retry in 1m (telegram sendMessage rejected: {"ok":false,…)
1315
+ 4b7c1ad9e001 SENDING digest 3m old attempt 1, outcome unknown — the process that sent it never said; a daemon start retries it and the message will say it may be a repeat
1316
+ ```
1317
+
1318
+ `pending` and `sending` are printed differently because they ask different things
1319
+ of you, and every row carries its age — "1 report pending since 08:15Z" is the
1320
+ signal that was missing when the reports went nowhere. Delivered reports leave
1321
+ the block: it is a list of what you are still owed, not a log.
1322
+
1323
+ Delivery keeps running while the fleet is **paused**. Pause stops claiming, not
1324
+ your right to hear about work that already happened. It runs on its own
1325
+ thirty-second timer rather than the five-minute dispatch tick, so a report does
1326
+ not sit in the outbox for the length of a poll interval.
1327
+
1328
+ The tier-2 escalation ledger (`notifications`) is untouched by all of this. It is
1329
+ a bare dedupe key by design — its primary key *is* the key — which is exactly why
1330
+ reports needed a separate table rather than an extension of that one.
1331
+
1332
+ ## The decision ledger (#136)
1333
+
1334
+ The outbox above fixed reports the orchestrator sends. This fixes the ones it is
1335
+ **waiting on**. A question put to you — an amendment, a tier-2 decision, "do I
1336
+ ship this tonight?" — lived in exactly one place: the model's context. A
1337
+ compaction, a restart, or a tick that ran long lost the question *and* the fact
1338
+ that one was owed, after which the session either asked again (you answer twice)
1339
+ or dropped it silently (the decision never lands, and nothing anywhere says one
1340
+ is outstanding).
1341
+
1342
+ So questions are written down, and every tick's prompt carries what is still
1343
+ open — read from the store, never from what the session remembers asking:
1344
+
1345
+ ```bash
1346
+ omp-conductor decision open --question "ship 0.4.3 tonight?" \
1347
+ --blocks "the release" --resolves-when issue-closed:132
1348
+ omp-conductor decision list
1349
+ omp-conductor decision resolve <id> --answer "yes, after #132 lands"
1350
+ omp-conductor decision withdraw <id> --reason "the release slipped a week"
1351
+ ```
1352
+
1353
+ **`--resolves-when` is the part that makes a parked question wake up.** Six
1354
+ conditions, each one something this package can check without asking you:
1355
+
1356
+ | Condition | Met when |
1357
+ | --- | --- |
1358
+ | `pr-merged:<https url>` | `gh` reports that pull request merged. |
1359
+ | `pr-checks-green:<https url>` | Every check on that pull request has a green verdict (a non-empty list, all `success`/`neutral`); a failing or still-pending check is not met. |
1360
+ | `pr-mergeable:<https url>` | The pull request is mergeable (`clean`, not `unknown` or conflicting). |
1361
+ | `issue-closed:<number>` | That issue is closed on the tracker. |
1362
+ | `npm-version:<pkg>@<version>` | `npm view <pkg>@<version> version` succeeds — the version is published. |
1363
+ | `rate-limit-reset:github` | GraphQL quota on `github` has any remaining capacity again. |
1364
+
1365
+ The daemon evaluates them beside each tick, fire-and-forget: a hanging registry
1366
+ costs one unevaluated condition, never the tick. A row that transitions
1367
+ false→true also writes the same `.conductor-tick-requested` poke recover uses,
1368
+ so the orchestrator heartbeat fires promptly (mid-interval poll, still gated by
1369
+ arm/channel/pending single-flight) instead of waiting a full interval. The poke
1370
+ reason and the digest flag `[CONDITION MET — act on this now]` both surface the
1371
+ wake so the session acts when the answer becomes actionable. Repeated sweeps
1372
+ while the condition stays true do nothing further — the store marks the
1373
+ transition once. After a green-but-behind PR is updated through
1374
+ `conductor_pr_update_branch`, open a fresh `pr-checks-green` watch on the new
1375
+ head so the next green transition can wake merge review the same way; nothing
1376
+ here merges on its own.
1377
+
1378
+ Anything else exits `2` and lists the six forms. An unparseable condition on an
1379
+ existing row is *listed and never treated as met*: a grammar a future release
1380
+ adds must not make an old row unloadable, and a question must never be hidden by
1381
+ a condition nobody can check.
1382
+
1383
+ **Expiry is enforced, not remembered.** An unanswered question closes itself
1384
+ after seven days — the deadline the floor's parked-amendment protocol already
1385
+ promised — so the digest stays a list of live questions instead of a graveyard.
1386
+ Answering or withdrawing is explicit, and a second resolution of the same id is
1387
+ refused rather than overwriting the first answer.
1388
+
1389
+ `omp-conductor status` carries one row, `decisions`, reported whether or not
1390
+ anything is open: `decisions 2 open (oldest 26h)`, or `decisions none open`. A
1391
+ row that appeared only when something was outstanding would leave "it forgot to
1392
+ record the question" and "there genuinely is none" looking identical, which is
1393
+ the ambiguity this table exists to remove.
1394
+
1395
+ ## Failure classes and recovery by class (#132)
1396
+
1397
+ Every run that did not reach a merged PR used to end at a human. The
1398
+ orchestrator re-derived the same triage on each tick — read the row, read the
1399
+ PR's checks, decide whether to requeue, re-run, settle or escalate — and then
1400
+ threw the conclusion away. Measured on this fleet's own history: **half the
1401
+ spend produced no merged PR**, and a large share of it was not implementation
1402
+ failure at all but daemon restarts, cancelled runners and a base branch moving
1403
+ under a green PR.
1404
+
1405
+ So the daemon classifies each terminal non-success before the next dispatch,
1406
+ persists the class on the row, and performs the one recovery that class names.
1407
+
1408
+ | Class | Signals | Recovery | Budget |
1409
+ | --- | --- | --- | --- |
1410
+ | `env-start-failure` | turn 0 plus an explicit harness start error (`No model selected`, a rejected key) | escalate — the session never read the issue | none |
1411
+ | `settlement-stuck` | a row carrying a PR that has since merged | settle: release the label, mark the row merged | none |
1412
+ | `returned-for-revision` | a `pushed-green` or `pushed-pending` PR was closed without merging | none — preserve the review decision for a human re-queue | continuation |
1413
+ | `merge-conflict` | `pushed-green`, PR open, GitHub reports conflicting | requeue for a rebase continuation | continuation |
1414
+ | `question` | the worker stopped to ask something (`blocked`) | escalate, carrying the worker's own report as evidence | none |
1415
+ | `orphan-dirty` | orphaned with a failed salvage and no operator ack | hold — recorded only; the tree is the only copy | none |
1416
+ | `orphan-clean` | orphaned with nothing uncommitted | requeue | continuation |
1417
+ | `turn-cap-progress` | at the turn ceiling **with** a PR, head or salvage commit | continue from the branch | continuation |
1418
+ | `turn-cap-spinning` | at the ceiling with no PR and no commits | escalate with the last tool calls the transcript recorded — and the completion path deliberately does **not** requeue it | none |
1419
+ | `admin-kill` | killed *below* its own ceiling — a restart or a drain | requeue | none |
1420
+ | `ci-infra` | PR open, every unresolved check cancelled / timed out / stale | re-run the failed jobs | none |
1421
+ | `ci-deterministic` | PR open, a check genuinely reports `FAILURE` | escalate with the failing check names and links | failed attempt |
1422
+ | `dispatch-infra` | the conductor's own Git path failed before the worker's first turn | requeue, bounded by per-class strikes | none |
1423
+ | `provider-credit` | the provider refused the run for credit (HTTP 402, or its own out-of-credit text read off the transcript) | pause the fleet and require `omp-conductor resume` once the provider has credit | none |
1424
+ | `provider-transient` | the provider aborted a request stream before the run produced a verdict | requeue, bounded by per-class strikes | none |
1425
+ | `unknown` | anything unrecognised | escalate | as recorded |
1426
+
1427
+ **Unknown escalates; it never silently retries.** A shape this table does not
1428
+ recognise is a gap in the table, and a quiet requeue would spend a budget on a
1429
+ cause nobody has named — the behaviour this exists to end.
1430
+
1431
+ ### The budgets follow the cause
1432
+
1433
+ `failuresFor` (implementation attempts) excludes `ci-infra`, `settlement-stuck`,
1434
+ `env-start-failure`, `dispatch-infra`, `provider-credit`, `provider-transient`
1435
+ and `returned-for-revision`. `continuationsFor` excludes `admin-kill`,
1436
+ `settlement-stuck`, `env-start-failure`, `dispatch-infra`, `provider-credit`
1437
+ and `provider-transient`, but explicitly counts a failed
1438
+ `returned-for-revision` row. Environment, dispatch and provider faults charge
1439
+ neither budget because the issue did not receive a valid implementation
1440
+ attempt. A merge conflict and a returned review both charge a continuation:
1441
+ each asks for more work, but neither is a failed implementation attempt.
1442
+
1443
+ An **unclassified** row (every row written before 0.4.3) counts exactly as it
1444
+ did before classification existed. Upgrading therefore changes no existing
1445
+ budget: the columns are additive and nullable, and a pre-0.4.3 `conductor.db`
1446
+ opens unchanged.
1447
+
1448
+ ### Stale labels are reconciled
1449
+
1450
+ On 2026-08-09 four issues carried `agent:failed` while every one of them was
1451
+ already complete — residue of a turns-cap kill two days earlier that nothing in
1452
+ the loop ever revisited. The board counted four phantom failures while the
1453
+ genuinely stuck issues were invisible.
1454
+
1455
+ Each tick now reconciles the three state labels against the tracker:
1456
+
1457
+ - A **closed** issue never keeps an `agent:*` label.
1458
+ - An **open** issue carrying `failed` whose sub-issues have *all* closed loses
1459
+ the label and gets one comment naming them, deduplicated through the same
1460
+ notifications ledger escalations use.
1461
+
1462
+ Positive evidence only: a tracker that cannot list answers empty, and an empty
1463
+ answer removes nothing — the label is the interlock that keeps two workers off
1464
+ one issue.
1465
+
1466
+ ### Where you see it
1467
+
1468
+ - `omp-conductor status` grows a `failure classes (unrecovered)` block, counting
1469
+ only rows whose recovery has *not* run. Classes rather than row states,
1470
+ because a row state is not an issue state.
1471
+ - The board appends `[<class>]` to a card whose newest run carries one.
1472
+ - The tick prompt carries one line — `Auto-recovered since last tick: 3
1473
+ (merge-conflict #365, admin-kill #82, …) — already handled, do not re-triage
1474
+ these.` — so the orchestrator stops writing that paragraph by re-deriving it.
1475
+
1476
+ ## Configuration
1477
+
1478
+ The config lives at `$OMP_CONDUCTOR_HOME/config.json`, or
1479
+ `~/.omp/conductor/config.json` when that variable is unset. It is written with mode
1480
+ `0600` in a directory created `0700`, because it carries chat ids and clone URLs.
1481
+ That same directory holds the SQLite store (`conductor.db`), the `paused` sentinel,
1482
+ the `sessions/` worker transcripts, the `orchestrator/` session directory,
1483
+ `backups/briefs/` for timestamped brief and policy safety copies, and
1484
+ `release-policy-blocks.jsonl`, the append-only audit of mechanically rejected
1485
+ release/deploy calls.
1486
+
1487
+ Runtime state lives elsewhere, under `$OMP_CONDUCTOR_RUNTIME_DIR` (default
1488
+ `~/.omp/run/daemons/omp-conductor`): `daemon.json`, a mode-`0600` pidfile written
1489
+ atomically, and `daemon.log`, appended across every boot so the previous failure is
1490
+ still there when you go looking. It is kept apart from the config directory because
1491
+ it is meaningless after a reboot, and the pidfile's liveness is probed on every
1492
+ read — a stale one never blocks a `start`. Both `start` and a bare `daemon` write
1493
+ the pidfile, so a daemon run in the foreground under systemd is as visible to
1494
+ `status` as a backgrounded one; `daemon --once` writes nothing, because that drill
1495
+ is exactly what the orphan-reconciliation guard reads the pidfile to protect.
1496
+
1497
+ The file is validated on every read. A malformed config produces one readable error
1498
+ listing every fault, and the daemon refuses to start rather than running with half
1499
+ a project.
1500
+
1501
+ The same vocabulary the loader enforces ships as a JSON Schema at
1502
+ `schema/config.schema.json` in the installed package (draft 2020-12). Anything
1503
+ `saveConfig` writes carries a top-level `"$schema"` reference to that installed
1504
+ copy (resolved from the package's own location, so it points at a real file),
1505
+ which lets an editor that understands JSON Schema validate a hand-edited config
1506
+ as you type; a config without the key is just as valid. Regenerate the shipped
1507
+ schema from `ConfigSchema` (`src/config-schema.ts`) with:
1508
+
1509
+ ```sh
1510
+ bun run schema
1511
+ ```
1512
+
1513
+ and commit the resulting `schema/config.schema.json`. CI's `bun test` fails if the
1514
+ checked-in schema drifts from what the code renders, so you cannot forget the step.
1515
+
1516
+ `omp-conductor setup` is the only thing here that writes this file, and on a project
1517
+ it already knows it can rewrite one area of it without re-asking the rest — see
1518
+ [Changing one setting](#changing-one-setting).
1519
+
1520
+ `version` is `2`. A `version: 1` file still loads: caps it names that this build no
1521
+ longer enforces are dropped rather than treated as typos, and the next save writes
1522
+ it back as `2`. In a `version: 2` file an unrecognised cap key **is** an error,
1523
+ because there is nothing left to retire — a mistyped `dailySpendUSD` would
1524
+ otherwise read as configured while the real ceiling stayed the default.
1525
+
1526
+ A complete, valid config for one project with two target repos:
1527
+
1528
+ ```json
1529
+ {
1530
+ "version": 2,
1531
+ "defaults": {
1532
+ "maxConcurrentWorkers": 2,
1533
+ "dailySpendUsd": 25,
1534
+ "planUsage": { "windowId": "anthropic:7d", "maxUsedFraction": 0.85 },
1535
+ "workerMaxTurns": 120,
1536
+ "workerWallClockMs": 5400000,
1537
+ "maxAttemptsPerIssue": 2,
1538
+ "maxContinuationsPerIssue": 2
1539
+ },
1540
+ "projects": [
1541
+ {
1542
+ "name": "demo",
1543
+ "tracker": { "kind": "github", "repo": "acme/planning" },
1544
+ "queueLabel": "ready-for-agent",
1545
+ "stateLabels": {
1546
+ "inProgress": "agent:in-progress",
1547
+ "blocked": "agent:blocked",
1548
+ "failed": "agent:failed"
1549
+ },
1550
+ "routing": {
1551
+ "labelPrefix": "repo:",
1552
+ "repos": {
1553
+ "api": {
1554
+ "name": "api",
1555
+ "cloneUrl": "git@github.com:acme/api.git",
1556
+ "defaultBranch": "main",
1557
+ "gates": [
1558
+ { "cmd": "bun run lint", "cwd": "." },
1559
+ { "cmd": "bun test", "cwd": "." }
1560
+ ],
1561
+ "graphProject": "~/.cache/conductor-graph/acme/api",
1562
+ "migrations": { "dir": "backend/alembic/versions" },
1563
+ "release": { "versionFile": "omp/package.json" }
1564
+ },
1565
+ "worker": {
1566
+ "name": "worker",
1567
+ "cloneUrl": "git@github.com:acme/worker.git",
1568
+ "defaultBranch": "main",
1569
+ "gates": [
1570
+ { "cmd": "ruff check .", "cwd": "." },
1571
+ { "cmd": "pytest -q", "cwd": "backend" }
1572
+ ]
1573
+ }
1574
+ }
1575
+ },
1576
+ "caps": {
1577
+ "maxConcurrentWorkers": 1,
1578
+ "dailySpendUsd": 15
1579
+ },
1580
+ "workerModel": "smol",
1581
+ "escalation": {
1582
+ "telegramChatId": "123456789",
1583
+ "telegramTopicId": 8713,
1584
+ "fallbackToIssueComment": true,
1585
+ "orchestrator": "embedded"
1586
+ },
1587
+ "authority": {
1588
+ "merge": "human",
1589
+ "release": "human"
1590
+ },
1591
+ "releasePolicy": {
1592
+ "version-bump-pr": "human",
1593
+ "git-tag": "human",
1594
+ "git-push-tags": "human",
1595
+ "package-publish": "human",
1596
+ "github-release": "human",
1597
+ "deploy": "human"
1598
+ },
1599
+ "policy": {
1600
+ "merge": {
1601
+ "requiredChecks": ["build", "lint"],
1602
+ "baseFreshness": "up-to-date",
1603
+ "drafts": "block",
1604
+ "whenBehindBase": "update-branch"
1605
+ },
1606
+ "release": {
1607
+ "requires": ["runs-settled", "no-open-prs"],
1608
+ "requiredChecks": ["release"],
1609
+ "artefacts": ["@acme/sdk"],
1610
+ "environments": ["staging"]
1611
+ }
1612
+ },
1613
+ "recoveryMerges": [
1614
+ {
1615
+ "prUrl": "https://github.com/acme/api/pull/381",
1616
+ "headSha": "9a783d8f17071d63f2d5d764d43a29837c365920",
1617
+ "reason": "operator-instructed"
1618
+ }
1619
+ ],
1620
+ "reporting": {
1621
+ "scope": "material"
1622
+ },
1623
+ "workspaceRoot": "~/.omp/conductor/worktrees",
1624
+ "mirrorRoot": "~/.omp/conductor/mirrors"
1625
+ }
1626
+ ]
1627
+ }
1628
+ ```
1629
+
1630
+ Field notes:
1631
+
1632
+ | Field | Notes |
1633
+ | --- | --- |
1634
+ | `version` | Must be `2`. A `version: 1` file still loads, drops the caps this build no longer enforces, and is rewritten as `2` on the next save. Present from day one so a format change can be migrated instead of silently misread. |
1635
+ | `defaults` | Every `Caps` field. Anything omitted falls back to the built-in default. |
1636
+ | `tracker.repo` | `owner/repo`. `tracker.kind` may be omitted; `"github"` is the only accepted value. |
1637
+ | `queueLabel` | The one label meaning "a human has signed this off as agent-ready". Matched exactly, case-sensitively. |
1638
+ | `release.versionFile` | Optional, per repo: a repo-relative JSON file with a top-level string `version`, such as `omp/package.json`. Declares that tags must match the version already landed on the live default branch. A delegated `git-tag` for such a repo requires delegated `version-bump-pr` too; otherwise config loading fails with the missing preparation path instead of granting an impossible release. Absolute paths and `..` are refused. |
1639
+ | `groomBelow` | Optional; default `4`. Routable candidates below this count make the orchestrator's tick prompt say the queue is running low and to groom it (Duty 2). An integer ≥ 1; anything else degrades to the default. |
1640
+ | `stateLabels` | Optional; defaults to `agent:in-progress`, `agent:blocked`, `agent:failed`. |
1641
+ | `routing.labelPrefix` | Optional; defaults to `repo:`. |
1642
+ | `routing.repos` | At least one entry, or nothing can be routed. `name` defaults to the map key, `defaultBranch` to `main`. |
1643
+ | `gates` | The exact cheap commands CI also runs, each with the `cwd` it runs from (`cwd` defaults to `.`). Running the real gate locally is what makes an unattended push safe — a subset lets an error outside the source dir reach the runners. |
1644
+ | `graphProject` | Optional, per repo. Absolute path of the **index-only clone** whose code graph this repo's workers query — conductor's own disposable clone, pinned to the repo's default branch, never a checkout you work in and never a worker's worktree. Written by the wizard; `~` is expanded, and a relative path is an error rather than something resolved against whichever cwd happened to read the file. Absent means this repo has no graph and its briefs say nothing about one. See [Code-graph discovery](#code-graph-discovery). |
1645
+ | `migrations` | Optional, per repo: `{ "dir": "backend/alembic/versions" }`. Names the repo-relative directory of an Alembic-style ordered migration chain (`revision` / `down_revision` in `*.py`). When set, `conductor_pr_merge` **refuses** a merge that would corrupt the chain at the base tip: reusing a revision id another file already declares, deleting a published migration, or a merge that would leave the combined graph with more than one head (so a stale parent is refused, and a fork-repair merge migration that unifies the heads passes). Absent means the repo opts out of the chain check entirely. Repo-relative only: a leading `/` or `..` is an error. |
1646
+ | `caps` | Per-project overrides; omit it or pin only the fields you want to change. |
1647
+ | `escalation.fallbackToIssueComment` | Defaults to `true`. Absent means "yes, still tell me". |
1648
+ | `escalation.telegramTopicId` | Optional forum topic for everything conductor sends: tier-2 pages, reports, digests, arm challenges, `omp-conductor message`. Setup offers the topics omp-telegram has claimed, naming each one's herdr space. The bridge re-claims a pane's topic across restarts — including the restarts `upgrade` and `restart` perform — so a pinned id that is no longer claimed is replaced at send time by the live claim whose **herdr space** is this project, falling back to one titled for the project, logged without ids (#407, #412). The space is read first because the bridge titles a topic `ownAgentName ?? basename(cwd)`, and a multi-project host whose panes sit under one state directory gives every claim the same title. An identity two claims share is treated as no match at all rather than a guess. A pin that is still claimed always wins, so a deliberately separate topic is never hijacked. Absent keeps flat-chat behaviour. |
1649
+ | `escalation.orchestrator` | Optional; `"embedded"` (default) or `"external"`. `external` means an orchestrator session already runs elsewhere: the daemon starts none, and tier-1 escalations post as issue comments for that session to drain. Any other value is an error. |
1650
+ | `authority` | Optional; `{ "merge": …, "release": … }`, each `"human"` (default) or `"orchestrator"`. It grants nothing to the daemon — it words the orchestrator's standing orders and the Releases paragraph of the rendered brief, so the config and the prompt cannot disagree about who holds the merge button. Unknown keys and any other value are errors, never folded to the default. |
1651
+ | `releasePolicy` | Optional; a per-shape map whose values are `"human"` (default) or `"orchestrator"`. Shapes are `version-bump-pr`, `git-tag`, `git-push-tags`, `package-publish`, `github-release`, and `deploy`. The legacy `"none"` denies every shape; legacy `"operator-brief"` grants the artifact-producing shapes, including reviewed version preparation, but keeps deploy human-owned. The in-session tripwire blocks recognised raw release/deploy calls before execution. Every rejection is written to `release-policy-blocks.jsonl`; the heartbeat carries that day's count into the daily digest. This is the mechanical gate; `authority.release` still says who owns the decision. |
1652
+ | `recoveryMerges` | Optional, hand-edited recovery authority for a PR that has no conductor run record. Each entry is an exact `{ prUrl, headSha, reason: "operator-instructed" }` tuple. When all three values match, `conductor_pr_merge` may merge that one PR even while the fleet is held and even when standing merge authority is `"human"`. It still requires a routed project repo, an open PR at that exact live head, green checks, the migration-chain guard, and the single-flight lock—the same safety path as an ordinary merge. Duplicate PR URLs and malformed values make config loading fail closed. Setup preserves entries but never creates them. Remove an entry after the recovery is complete. |
1653
+ | `reporting` | Optional; a **legacy scope preset** (`reporting.scope` — `"material"` default, `"decisions"`, `"escalations"`) or the **explicit form** `{ "interruptOn": [...], "digest": { ... }, "availability": { ... } }`. The preset writes which categories may page the operator (`interruptOn`) and when the rollup happens (`digest.cadence`); the explicit form sets both directly and may add a weekly operator-availability window. The two forms are mutually exclusive in one config. See [Reporting policy](#reporting-policy-reporting). |
1654
+ | `orchestratorReadPaths` | **Retired in 0.4.3.** Still accepted in a config and ignored, so a fleet carrying it upgrades without an edit. It widened the orchestrator's file-tool allowlist; there is no allowlist any more — the orchestrator is [unconfined by design](#the-orchestrator-is-unconfined-deliberately). |
1655
+ | `policy` | Optional; the gating conditions a merge or a release must satisfy, in two sections — `policy.merge` and `policy.release`. Any member may be omitted and the loader fills it from the strict default; an unknown key in either section, or a value outside its vocabulary, is an error naming the field, never a silent downgrade. See [Merge and release preconditions](#merge-and-release-preconditions-policy). |
1656
+ | `workspaceRoot` / `mirrorRoot` | Optional; default to `worktrees/` and `mirrors/` under the state directory. `~` is expanded. |
1657
+
1658
+ Prefer an SSH `cloneUrl`, or an https URL backed by a credential helper. A clone URL
1659
+ with credentials embedded is persisted into the mirror's git config, exactly as it
1660
+ would be for a hand-run clone.
1661
+
1662
+ ### Reporting policy (`reporting`)
1663
+
1664
+ What may interrupt the operator's phone, and when the daily rollup happens. Two
1665
+ spellings, mutually exclusive in one config (the loader rejects a `scope` next to
1666
+ `interruptOn`/`digest`):
1667
+
1668
+ - **Preset** — `reporting.scope`, the three legacy values, mapped verbatim:
1669
+ - `material` (default) → `interruptOn: [tier2, decision-needed, fleet-stopped, confirmed-failure, material]`, digest `per-tick`.
1670
+ - `decisions` → `interruptOn: [tier2, decision-needed, fleet-stopped]`, digest `per-tick`.
1671
+ - `escalations` → `interruptOn: [tier2, fleet-stopped]`, digest `daily` (model-timed).
1672
+ - **Explicit** — `reporting: { "interruptOn": ["tier2", "fleet-stopped", ...], "digest": { "cadence": "none" | "per-tick" | "daily" } }`.
1673
+ `interruptOn` must be a non-empty array of known categories (`tier2`, `decision-needed`, `fleet-stopped`, `confirmed-failure`, `material`), each an escalation's tier-2 category. `daily` may add `at` (`HH:MM`, 24h) and `timezone` (a known IANA zone, defaulting to the host zone) — both only valid with `daily`.
1674
+
1675
+ The explicit form may add a weekly local-time window:
1676
+
1677
+ ```json
1678
+ {
1679
+ "reporting": {
1680
+ "interruptOn": ["tier2", "fleet-stopped"],
1681
+ "digest": { "cadence": "daily", "at": "17:00", "timezone": "Europe/London" },
1682
+ "availability": {
1683
+ "timezone": "Europe/London",
1684
+ "days": ["mon", "tue", "wed", "thu", "fri"],
1685
+ "start": "09:00",
1686
+ "end": "17:00",
1687
+ "bypass": ["fleet-stopped"]
1688
+ }
1689
+ }
1690
+ }
1691
+ ```
1692
+
1693
+ `timezone` must be a known IANA zone. For a daily digest, its timezone defaults
1694
+ to this value and must match it when both are set.
1695
+
1696
+ `days` is a non-empty set of `mon` through `sun`; `start` is inclusive and
1697
+ `end` is exclusive. A start later than the end defines an overnight window on
1698
+ the day it opens. `bypass` is an explicit list of known interrupt categories
1699
+ that may still page outside the window; it may be empty. For ordinary notices,
1700
+ a bypass has no effect on a category omitted from `interruptOn`. Urgent recovery
1701
+ notices may bypass category batching when the digest loop itself is unavailable,
1702
+ but they still require the configured availability bypass outside the window.
1703
+
1704
+ The setup wizard offers this as **Weekly availability window** and asks for the
1705
+ zone, days, start/end, bypass categories, and digest schedule: every tick,
1706
+ model-timed daily, disabled, or a fixed daily `HH:MM`. Re-running setup or
1707
+ amending reporting preselects and preserves the configured `none`, `per-tick`,
1708
+ or `daily` cadence. Choosing **Continuous (24-hour interrupts)** is the explicit
1709
+ opt-out and preserves the behavior of every existing config; an absent
1710
+ `availability` key also means continuous operation.
1711
+
1712
+ Outside the window, an otherwise interruptible escalation is stored durably
1713
+ instead of sent. A daily digest may consume it first. Otherwise the daemon
1714
+ atomically queues one working-hours catch-up report when the window opens,
1715
+ including after downtime; associating the held rows before delivery prevents a
1716
+ later tick from authoring a duplicate. Each heartbeat prompt names the
1717
+ mechanically computed current mode and next transition. `status` shows the same
1718
+ state plus the next digest opportunity (`due now`, every tick, disabled, or its
1719
+ next operator-local timestamp). Config, escalation routing, and report transport
1720
+ are re-read at tick or send time, so changing the window or Telegram target
1721
+ does not require a daemon restart.
1722
+
1723
+ Attachment-bearing autonomous Telegram sends cannot be replayed by the text
1724
+ digest, so they are blocked with an explicit “nothing sent or held” error rather
1725
+ than silently dropping their files.
1726
+
1727
+ A tier-2 escalation whose category is **not** in `interruptOn` is not dropped: it
1728
+ is held (`held_notices`) and the next accepted digest is its delivery authority.
1729
+ A `daily` digest is at-most-once per local day (`digest:<YYYY-MM-DD>` in the
1730
+ configured zone), which remains the delivery authority across restarts.
1731
+ `per-tick` digests are not daily-deduplicated, so a later tick can claim newly
1732
+ accumulated rows. A scheduled `daily` digest is only sent on a day it has not
1733
+ already run, once the local clock has passed `at`; a restart after `at` still
1734
+ sends today's (one catch-up), and a fully missed day is skipped, never sent late.
1735
+
1736
+ **What the scope does:** the [orchestrator heartbeat](#orchestrator-tick) appends
1737
+ the current constraint to every tick it sends, so the reporting contract arrives
1738
+ with the prompt instead of only in a brief the session read hours ago. Explicit
1739
+ policies name their actual interrupt categories and digest cadence. The policy is
1740
+ re-read from `~/.omp/conductor/config.json` on **every** tick, and escalation,
1741
+ direct Telegram, and durable report paths apply it mechanically. Turning the
1742
+ volume up or down — `omp-conductor setup` again, or an edit to the file — therefore
1743
+ binds the next tick without restarting the session.
1744
+
1745
+ No config, an unreadable or invalid config, or several unnamed projects fall
1746
+ back to the legacy `material` scope for the heartbeat and log the reason once.
1747
+ An invalid live availability policy blocks autonomous Telegram fail-closed; it
1748
+ does not guess that the operator is awake.
1749
+
1750
+ Changing the key later does not rewrite an `ORCHESTRATOR.md` you already have.
1751
+ The generated `POLICY.md` describes every scope without pinning the current
1752
+ choice; the tick constraint remains derived from live config. Keep any
1753
+ operator-owned reporting additions in `ORCHESTRATOR.md` consistent with it.
1754
+
1755
+ ### Merge and release preconditions (`policy`)
1756
+
1757
+ These used to be sentences in your `POLICY.md`: when a PR may be merged, what
1758
+ must be green, what a release requires. Prose cannot be checked, so every tick
1759
+ re-decided them by reading and interpreting them again. They are configuration
1760
+ now, `POLICY.md` keeps only judgement, and the rendered brief *describes* the
1761
+ policy instead of restating it — no threshold lives in two places.
1762
+
1763
+ `policy.merge`:
1764
+
1765
+ | Field | Values | Default | Means |
1766
+ | --- | --- | --- | --- |
1767
+ | `requiredChecks` | any check names | `[]` | Checks that must have concluded successfully. **Empty is the strict answer** — it means every check the PR reports, not "no checks". |
1768
+ | `baseFreshness` | `up-to-date`, `any` | `up-to-date` | Whether the head must be level with the base branch. `any` accepts a verdict produced against an older base. |
1769
+ | `drafts` | `block`, `allow` | `block` | Whether a draft PR can be merged at all. |
1770
+ | `whenBehindBase` | `update-branch`, `hold`, `escalate` | `update-branch` | What to do with a green PR that fell behind. `update-branch` runs `gh pr update-branch` and waits for the fresh run. Closing it and an admin bypass are not spellable. |
1771
+
1772
+ `policy.release`:
1773
+
1774
+ | Field | Values | Default | Means |
1775
+ | --- | --- | --- | --- |
1776
+ | `requires` | `runs-settled`, `no-open-prs`, `queue-drained`, `base-branch-green`, `epic-children-closed` | `["runs-settled"]` | What must already have landed. `runs-settled` reads each active run's PR fact at release time: a pushed run whose PR has merged counts as settled even when the settle sweep has not yet written the terminal row — so a hold-drained release does not wait an extra tick the operator reached the gate by holding. Live workers and unmerged/unknown PRs still refuse, and the message names which is which. `base-branch-green` requires the current live head's push-triggered workflow verdict for that routed repository to be green; pending, unknown, red, or no observation refuses release. Order and duplicates do not matter; the loader canonicalises. |
1777
+ | `requiredChecks` | any check names | `[]` | Checks that must be green on the branch being released. Empty means every check it reports. |
1778
+ | `artefacts` | any names | `[]` | The packages or images this project releases. **Empty denies**: nothing has been authorised to ship. |
1779
+ | `environments` | any names | `[]` | Deploy targets. **Empty denies** every environment. |
1780
+
1781
+ A project with no `policy` block loads as the whole default above, which is the
1782
+ strictest reading of the prose it replaced. `omp-conductor setup` asks for all of
1783
+ it under the **merge & release preconditions** area, so changing one condition
1784
+ costs eight prompts rather than a hand-edit — see
1785
+ [Changing one setting](#changing-one-setting).
1786
+
1787
+ This key grants nothing. Who *may* merge or release is
1788
+ [`authority`](#configuration), and which release tool calls are mechanically
1789
+ permitted is [`releasePolicy`](#configuration). `policy` says what must be true
1790
+ before the act, whoever is doing it.
1791
+
1792
+ #### Reasons are a closed vocabulary
1793
+
1794
+ Where an automated verb takes a `reason`, the argument is one value out of a
1795
+ fixed set, not free text — a reason a rule matches on is a reason that decides,
1796
+ and a decision made out of a model's own wording is one no two runs spell the
1797
+ same way. A reason outside its set is refused, and the refusal names every
1798
+ accepted value.
1799
+
1800
+ | Verb | Accepted reasons |
1801
+ | --- | --- |
1802
+ | merge | `preconditions-met`, `behind-base-refreshed`, `operator-instructed`, `release-blocking` |
1803
+ | release | `batch-complete`, `epic-closed`, `hotfix`, `operator-instructed` |
1804
+ | label change | `promoted-to-queue`, `re-briefed`, `needs-human`, `duplicate`, `superseded`, `out-of-scope` |
1805
+
1806
+ Free-form rationale still has a home: it rides alongside as a separate
1807
+ `rationale` field, is written into the audit trail verbatim, and is never
1808
+ parsed or matched by anything.
1809
+
1810
+ ## Orchestrator tick
1811
+
1812
+ The escalation path above assumes an orchestrator session that is actually
1813
+ running its loop. A 24/7 omp session with a standing brief and nobody typing into
1814
+ it never gets prompted, so it never runs anything. Installing
1815
+ `omp plugin install omp-conductor` also installs a heartbeat that prompts it.
1816
+
1817
+ The heartbeat is **inert unless the session cwd contains
1818
+ `.conductor-tick.json`**, so an ordinary session has no timer. `omp-conductor setup`
1819
+ writes this file for external orchestration. A manual configuration has this form:
1820
+
1821
+ ```json
1822
+ {
1823
+ "intervalSeconds": 900,
1824
+ "project": "fleet",
1825
+ "armedFile": "/home/fleet/.omp/conductor/armed-fleet",
1826
+ "accessFile": "/home/fleet/.omp/agent/telegram/access.json",
1827
+ "message": "Run your standing loop from ORCHESTRATOR.md now."
1828
+ }
1829
+ ```
1830
+
1831
+ | Key | Required | Default | Notes |
1832
+ | --- | --- | --- | --- |
1833
+ | `intervalSeconds` | yes | — | Whole seconds between ticks, minimum `60`. A tick costs a full turn of a frontier model, so a sub-minute period is refused rather than obeyed. |
1834
+ | `project` | no | the only configured project | Which conductor project this fleet session ticks for. `setup host` stamps it, one tick config per fleet cwd, and it is what lets a host with several configured projects resolve *this* fleet's brief, reporting policy, digest ledger and release grants. Omitting it is the pre-multi-project spelling: correct on a single-project host, and on a host with two or more it degrades every tick to the default reporting scope with no release grants — `status` and the tick log then name the one fix (`re-run omp-conductor setup host`). A name no configured project has degrades the same way. |
1835
+ | `budgetSeconds` | no | `600` | Seconds a turn may run before the tick guard refuses its remaining tool calls (#189), and before a queued operator message preempts them. An integer ≥ 60; anything else degrades to the default. |
1836
+ | `askTimeoutSeconds` | no | `300` | Seconds one `conductor_ask` call waits for the operator before its declared `on-timeout` outcome fires (#438). An integer 60–3600; anything else degrades to the default. The resolved ceiling is always capped at the turn budget, so an ask can never outlive the turn it runs in. |
1837
+ | `armedFile` | no | none — the gate passes | Path to the arm marker. A tick does nothing while the file is missing. **Re-read from disk on every tick**, so a `setup host` restamp onto `armed-<project>` binds on the next heartbeat instead of leaving a live pane watching the path it captured at session start. Relative paths resolve against the session cwd, so `state/armed` means `<cwd>/state/armed`. `setup host` writes `<state dir>/armed-<project>`, one marker per project, so `arm --project A` cannot arm B. A value it did not generate is left alone as your own choice. |
1838
+ | `accessFile` | no | none — the gate passes | Path to the Telegram bridge's `access.json`. Every tick re-reads it and requires `enabled: true` with exactly one entry in `allowFrom`. Relative paths resolve against the session cwd. **Configure this on any fleet deploy** — see below. |
1839
+ | `message` | no | `Tick <ISO timestamp>: re-read <workspaceRoot>/ORCHESTRATOR.md from disk, then run your standing loop from it.`, then the reporting-policy line, delivery rule, and mechanical availability state | When set, this text replaces the ordinary reporting-policy line and delivery rule, but the runtime-owned availability state is still appended: a custom prompt cannot infer whether the operator may be interrupted. Re-read from disk on **every** tick, so rewording it binds the next heartbeat instead of waiting for a session restart; a re-read that fails — caught mid-edit, removed, or invalid — keeps the value read at session start rather than stopping the heartbeat. `intervalSeconds` is *not* re-read: rescheduling a live timer still needs a restart. The default *orders* the session to re-read its brief, naming the path resolved from the project's `workspaceRoot`, because a standing prompt drifts out of a long-lived session's context while the file on disk does not. |
1840
+ | `agentName` | no | the project name, else `fleet` | The herdr agent name the orchestrator's pane is registered under. Under herdr this is the whole of the identity check below. `setup host` writes the project name, so two fleets in one herdr session are distinguishable; when no tick config names one, the fallback matches `AGENT_NAME=${AGENT_NAME:-fleet}` in the recovery plugin's `recover.sh`, so both halves key on one name. Rename the agent and set this to match. |
1841
+
1842
+ #### Bounded operator asks (#438)
1843
+
1844
+ On 2026-08-16 an unanswered `telegram_ask` blocked the orchestrator turn for
1845
+ six hours; every duty behind the question stopped with it. `telegram_ask` is
1846
+ omp-telegram's tool and it waits as long as the answer takes, so on a locally
1847
+ injected tick the extension refuses it and mounts its own bounded surface,
1848
+ `conductor_ask`:
1849
+
1850
+ - The call carries `question`, `on-timeout` (`auto-proceed` or `park`, required),
1851
+ and optionally `timeoutSeconds`, `blocks`, `recommended`, `options` and
1852
+ `category`.
1853
+ - The tool records a decision row first (durable, seven-day expiry), then
1854
+ delivers the question through the same path `omp-conductor message` uses —
1855
+ immediately when the reporting policy permits, durably held otherwise.
1856
+ - It waits at most the ceiling: `timeoutSeconds` if the ask names one, else the
1857
+ tick config's `askTimeoutSeconds`, else 300 seconds — always capped at the
1858
+ turn budget, so the ask can never outlive the turn it runs in. An ask issued
1859
+ without a timeout gets the default all the same.
1860
+ - On timeout, `auto-proceed` applies the recommended option and resolves the row
1861
+ with `"<option> (auto-applied on ask timeout)"` so the record never reads as a
1862
+ human choice; `park` leaves the row open and pending — re-surfaced in every
1863
+ tick prompt until answered or the seven-day expiry — and the caller takes the
1864
+ blocked work out of the claimable queue. A timeout is "nobody answered yet",
1865
+ never a cancelled/errored ask and never an operator no.
1866
+
1867
+ The refusal of the raw tool is mechanical (the tool-call gate), not a prompt
1868
+ reminder: a model cannot wait unbounded on a local tick even by omitting the
1869
+ timeout argument.
1870
+
1871
+ #### Upgrading from one shared arm marker
1872
+
1873
+ Before per-project markers every project was given the same `<state dir>/armed`,
1874
+ so arming one fleet armed all of them. `setup host` rewrites that value — and
1875
+ only that value — to `armed-<project>`. The old bare marker is honoured for one
1876
+ more cycle on a **single-project** host, so the upgrade never silently disarms a
1877
+ live fleet, and the next `arm` or `disarm` retires it. On a host with **two or
1878
+ more** projects it arms nothing: `status` reports `legacy global arm marker —
1879
+ re-run setup host, then arm per project`, and every tick stays disarmed until each
1880
+ project is armed on its own marker.
1881
+
1882
+ The same restamp renames the identity this pane ticks under: an `agentName` of
1883
+ `fleet` — the value every project used to be given — becomes the project name.
1884
+ **Under herdr that is an operator step, not a no-op.** Ownership is proved against
1885
+ the pane's registered herdr agent, so after re-running `setup host` the live fleet
1886
+ pane needs the new name.
1887
+
1888
+ **Rename the agent herdr already detects — do not `agent start`.** `herdr agent
1889
+ start` submits omp *into* the pane's existing shell and requires a pane sitting at
1890
+ a shell prompt with no agent on it; the live orchestrator pane is neither, so it
1891
+ is refused at best and starts a second omp in that pane at worst. `rename` touches
1892
+ no process and keeps the session as it is:
1893
+
1894
+ ```sh
1895
+ herdr --session <session> agent list # find the fleet's pane_id
1896
+ herdr --session <session> agent rename <pane-id> <project>
1897
+ ```
1898
+
1899
+ If the name cannot be reassigned in place, stop and resume rather than starting a
1900
+ second orchestrator — the same shape `recover.sh` uses, so the omp session is
1901
+ preserved rather than replaced:
1902
+
1903
+ ```sh
1904
+ herdr --session <session> agent get <pane-id> # note agent_session.value — the session ref
1905
+ # exit omp in that pane (/exit) so the pane is back at a shell prompt, then:
1906
+ herdr --session <session> agent start <project> --kind omp --pane <pane-id> -- --resume=<ref>
1907
+ ```
1908
+
1909
+ Until the pane carries the new name it declines to tick and logs which agent it
1910
+ actually is versus the one the tick config names, with the `rename` command in the
1911
+ line — the heartbeat fails closed and says so rather than letting two fleets both
1912
+ answer to `fleet`. Set `agentName` explicitly if you would rather keep the old
1913
+ name; a value that is not the shared default is never rewritten.
1914
+
1915
+ The marker itself needs no restart: `armedFile` is re-read from disk every tick, so
1916
+ the restamped path binds on the next heartbeat. Before 0.15.2 it was read once at
1917
+ session start, and a restamp under a live pane left that pane watching a path the
1918
+ restamp had just replaced — `status` reported `armed` from the file while the
1919
+ heartbeat skipped silently as "not armed", which writes no stall marker.
1920
+
1921
+ Recovery is fail-closed across that window. A restamped `agentName` moves the
1922
+ recovery plugin's own state files to per-agent paths that do not exist yet, and
1923
+ the live pane is still saved under `fleet`, so the snapshot offers no candidate
1924
+ for the new name. `herdr-conductor` treats the pre-rename identity and bootstrap
1925
+ marker as proof a fleet has already lived on this host whatever it is called now:
1926
+ it pages `no fleet identity to recover for agent <project>` instead of
1927
+ provisioning a second workspace beside the live orchestrator. Finish the rename
1928
+ and the next pass recovers normally.
1929
+
1930
+ A default tick sends one message (`customType` `omp-conductor.tick`, attributed
1931
+ to the user): the standing-loop prompt, the reporting-policy constraint re-read
1932
+ from conductor config on every tick, the delivery rule, and the mechanically
1933
+ computed operator-availability state. A configured `message` replaces the first
1934
+ three parts but not that clock state. The delivery rule is there
1935
+ because end-of-turn text reaches the operator's Telegram only on a turn that
1936
+ *began* as an inbound Telegram message: a tick is injected locally, so anything
1937
+ the session merely writes at the end of one is read by nobody, and a reportable
1938
+ event has to be delivered by an explicit `telegram_send` call the session
1939
+ watched succeed. The tick starts a turn if the session is idle; while a turn is
1940
+ streaming it is queued as a follow-up and consumed when that turn ends.
1941
+ It sends **nothing** when:
1942
+
1943
+ - `armedFile` is configured and missing;
1944
+ - `accessFile` is configured and the escalation channel is not verifiably up;
1945
+ - an earlier tick is still queued. Ticks coalesce rather than stack, so a slow
1946
+ turn cannot leave a backlog of heartbeats behind it — and two coalesced ticks
1947
+ in a row are the signal that the session is not slow but wedged, which is
1948
+ what the [stall marker](#a-wedged-session-and-the-marker-that-notices) is for.
1949
+
1950
+ ### One session per directory ticks, and it says which
1951
+
1952
+ Activation is a property of the *directory*, so before it arms anything the
1953
+ heartbeat asks whether this session is the orchestrator or merely a session
1954
+ standing in its directory. It has to: opening a second omp session in the fleet's
1955
+ cwd — a shell to read state, say — used to arm a second heartbeat that prompted
1956
+ *that* session with the standing loop, and with
1957
+ [`authority`](#configuration) delegated it would consider itself entitled to
1958
+ merge PRs and cut releases. Two brains, one queue, and nothing in the log to tell
1959
+ them apart.
1960
+
1961
+ **Under herdr** (`HERDR_ENV=1` with a `HERDR_PANE_ID`), the answer is the pane's
1962
+ registered agent name: the heartbeat asks `herdr agent list` for the entry whose
1963
+ `pane_id` is this pane's and ticks only when its `name` equals `agentName`. Fleetness
1964
+ is the *session* — every pane in it shares `HERDR_SESSION` and the cwd — and
1965
+ herdr's `agent` field is the *runtime*, `omp` for the orchestrator and for the
1966
+ shell beside it, so neither can tell them apart. The registered name can, it is
1967
+ what `herdr agent start fleet --kind omp --pane <id>` sets when the recovery plugin
1968
+ starts a fleet into an empty pane — and what `herdr agent rename <pane-id> fleet`
1969
+ sets on a pane whose agent herdr already detects, which is the only safe spelling
1970
+ while omp is running in it. It is the same identity the recovery plugin keys on. A
1971
+ pane with a different name, or no name at all, stays inert.
1972
+
1973
+ **Without herdr**, the session claims the directory in a sibling
1974
+ `.conductor-tick-owner.json` (pid, session file, claim time) and ticks only while
1975
+ it is the live claimant. Liveness is a **pid check, never a timestamp**: a crashed
1976
+ orchestrator's claim is reclaimed by the next session rather than wedging the
1977
+ fleet until somebody deletes a file, and a slow-but-running orchestrator never
1978
+ loses its claim to a lease that expired.
1979
+
1980
+ Declining is logged once, at session start, naming the holder — which is the whole
1981
+ point, because the original failure was that the second ticker was
1982
+ indistinguishable from the first:
1983
+
1984
+ ```text
1985
+ [omp-conductor] orchestrator tick inactive: pane w1:p1 (agent "fleet") owns the fleet tick here — this session will not tick
1986
+ [omp-conductor] orchestrator tick inactive: this pane is agent "scratch", not the fleet agent "fleet" — this session will not tick
1987
+ [omp-conductor] orchestrator tick inactive: pid 12345 (claimed 2026-01-02T03:04:05.000Z, session …/fleet.jsonl) owns the fleet tick in /home/conductor/.omp/conductor — this session will not tick
1988
+ ```
1989
+
1990
+ A `herdr agent list` that does not answer also declines, for the same reason the
1991
+ escalation channel fails closed: under herdr this session is one pane of several
1992
+ in that directory, and an unproven identity is exactly the case the check exists
1993
+ for. That includes `herdr` not being on the session's `PATH` — worth checking on a
1994
+ fleet host, where the orchestrator's environment comes from a unit file rather
1995
+ than a login shell — and `HERDR_BIN_PATH` names the binary when it is not, the
1996
+ same escape hatch the recovery plugin's `recover.sh` has. On a host with no herdr
1997
+ and no prior claimant — the ordinary single-session case — nothing changes.
1998
+
1999
+ ### The escalation channel is a gate, and it fails closed
2000
+
2001
+ Unattended dispatch is only defensible while a tier-2 escalation can reach a
2002
+ person. So `accessFile` is checked on **every** tick and never cached at session
2003
+ start: the bridge is reconfigured out-of-band, and a heartbeat that trusted a
2004
+ startup snapshot would keep dispatching for days after the channel went away. A
2005
+ stale arm marker must not outlive the channel that makes running unattended safe.
2006
+
2007
+ The check passes only when a bot token is resolvable — `TELEGRAM_BOT_TOKEN` in
2008
+ the environment, or in the `.env` beside `accessFile` — and the file parses to an
2009
+ object with `enabled: true` and exactly one `allowFrom` entry. Everything else
2010
+ stops the heartbeat: no token, so nothing outbound works at all; file missing,
2011
+ unreadable or truncated; not JSON, or JSON that is not an object; `enabled`
2012
+ absent or false; zero owners paired (nobody to page) or more than one (ambiguous:
2013
+ the conductor refuses to guess which human is on the hook). Failure modes are
2014
+ deliberately not distinguished in the decision: each one means a page lands
2015
+ nowhere. `omp-conductor status` is where they are told apart — its `telegram` row
2016
+ names the specific fault.
2017
+
2018
+ One caveat the file cannot express: omp-telegram binds its own copy of the token
2019
+ in `startBot()` at session start, and only when the bridge is switched on. It
2020
+ rebinds only on `/telegram token` or `/telegram on`. So writing a token into
2021
+ `.env` out-of-band — or flipping `enabled` to true by hand — restores tier-2
2022
+ paging immediately, because conductor sends those itself, while the bridge's own
2023
+ tools, `telegram_send` and `telegram_ask`, stay dead until you reload it. After
2024
+ either edit, run `/telegram on` in the orchestrator session. Until you do, ticks
2025
+ carry an explicit note that an amendment cannot be approved on this surface, and
2026
+ the conductor never assumes an answer it did not receive.
2027
+
2028
+ Leaving `accessFile` unset passes the gate, because an ordinary developer session
2029
+ that happens to have a `.conductor-tick.json` has no bridge to check. It is not an
2030
+ off switch for the check: **a fleet deploy always sets it.**
2031
+
2032
+ Every tick — sent or skipped — is logged with its reason (`not armed`,
2033
+ `escalation channel down`, `tick already pending`) to the omp log. `omp-conductor
2034
+ hold` is deliberately **not** one of the gates: hold stops the *dispatcher*
2035
+ claiming work, and the tick drives a different session — one whose duties
2036
+ (grooming the queue, draining escalations, reporting) are exactly what stays
2037
+ useful while dispatch is stopped. Its own off switch is the arm marker. Skips
2038
+ are deliberately silent in the UI: a disarmed fleet would otherwise raise a
2039
+ notification every interval, forever. The one exception is a malformed
2040
+ `.conductor-tick.json`,
2041
+ which notifies once at session start and leaves the heartbeat off; silent failure
2042
+ there is the failure mode the heartbeat exists to prevent. A conductor config that
2043
+ cannot supply a reporting scope logs `tick reporting scope: using material` once
2044
+ per session. The interval does not re-log it, because the file is unlikely to fix
2045
+ itself between two ticks.
2046
+
2047
+ ### A wedged session, and the marker that notices
2048
+
2049
+ Coalescing is also the only wedge detector this package has. On 2026-08-07 the
2050
+ dogfood fleet's orchestrator finished a turn, logged `ui.loop-blocked` right
2051
+ after an auto-compaction threshold decision, and never started another. The
2052
+ process stayed alive, so herdr's recovery — agent listed AND a non-shell
2053
+ foreground process — read healthy. The dispatch daemon is a separate process and
2054
+ kept working, so `/healthz` was green all night, while the one brain holding
2055
+ merge authority sat on a green PR it never merged. A tick injected two minutes
2056
+ into the wedge and an operator's Telegram message five minutes later both went
2057
+ unconsumed for 23 minutes, until a manual `SIGTERM`. The heartbeat logged `tick
2058
+ skipped: tick already pending` throughout, which is exactly what a merely slow
2059
+ turn looks like.
2060
+
2061
+ So the heartbeat counts them. Two consecutive coalesced ticks — a full hour at
2062
+ the reference 1800-second interval, generous by construction — mean the last
2063
+ prompt was never consumed, and the extension:
2064
+
2065
+ - writes `<session cwd>/.conductor-stalled`, one line of `<ISO timestamp>
2066
+ <diagnosis>`.
2067
+ - logs at **error** level: `orchestrator stalled: 2 ticks queued unconsumed —
2068
+ the agent loop is not draining; see .conductor-stalled`.
2069
+
2070
+ Both escapes deliberately leave the session, because a loop that cannot drain
2071
+ its queue cannot report on itself — that is the whole failure.
2072
+
2073
+ **The daemon reads it.** A marker nobody consumes is an artifact, not an alert,
2074
+ so the dispatch daemon checks it on its own five-minute tick — and *before* its
2075
+ pause check. The orchestrator is a different process and can be wedged while
2076
+ the fleet is deliberately paused, which is precisely the state the dogfood
2077
+ fleet was in when this happened. One tier-2 page per stall, keyed on the
2078
+ marker's own timestamp so a second wedge the same day is not swallowed as a
2079
+ repeat, re-armed when the marker clears, and latched only once the page is
2080
+ confirmed delivered — an escalation channel that fails on the one tick that
2081
+ noticed must not buy permanent silence.
2082
+
2083
+ It restarts nothing. A wedge lands mid-turn, and no other process can tell a
2084
+ half-applied edit from an idle loop; the operator attaches, looks, and decides.
2085
+
2086
+ **herdr-conductor deliberately does not read it**, though its liveness test
2087
+ (agent listed AND a non-shell foreground process) passes straight through a
2088
+ wedge. That plugin only runs on `startup`, `pane.exited` and
2089
+ `pane.agent_detected`, and a session that stays alive and stops working emits
2090
+ none of them — so the check could never fire during the wedge itself. What it
2091
+ *would* catch is the recovery afterwards: the marker survives a restart until
2092
+ the new session consumes a tick, so every operator SIGTERM-and-resume would
2093
+ page about the healthy session they just fixed. Telling those apart needs the
2094
+ process start time against the marker's, and herdr's `pane process-info`
2095
+ reports pids, not start times. The daemon gives up at most one tick of
2096
+ coverage and never cries wolf.
2097
+
2098
+ The first tick that actually sends clears the counter and deletes the marker,
2099
+ and it deletes one it did not write: recovery normally arrives as a fresh
2100
+ process resuming the same transcript, so the session doing the clearing is not
2101
+ the session that stalled. Nothing else removes the file. Neither the write nor
2102
+ the delete can take the heartbeat down — a filesystem error is logged and the
2103
+ tick carries on.
2104
+
2105
+ `omp-conductor status` reads the same marker from the **state directory** and
2106
+ prints one more line under the daemon block:
2107
+
2108
+ ```text
2109
+ orchestrator STALLED since 2026-08-07T06:27:55.123Z — 2 ticks queued unconsumed — the agent loop is not draining
2110
+ ```
2111
+
2112
+ That reading is the reference deploy's convention — the orchestrator session
2113
+ runs from `~/.omp/conductor`, which is the state directory — and it is
2114
+ one-directional: a line there proves a wedge, and its absence proves nothing,
2115
+ least of all on a fleet whose session lives somewhere else.
2116
+
2117
+ ## CLI reference
2118
+
2119
+ ```bash
2120
+ omp-conductor setup [area] [--no-ai] [--project NAME]
2121
+ omp-conductor setup host [--project NAME]
2122
+ omp-conductor setup graph [--no-seed] [--print] [--project NAME]
2123
+ omp-conductor start [--port N] [--project NAME]
2124
+ omp-conductor --version
2125
+ omp-conductor stop
2126
+ omp-conductor restart [--now] [--timeout SECONDS] [--port N] [--project NAME]
2127
+ omp-conductor upgrade [--to VERSION] [--project NAME]
2128
+ omp-conductor upgrade-install --to VERSION [--project NAME]
2129
+ omp-conductor upgrade-rollback [--project NAME]
2130
+ omp-conductor status [--project NAME]
2131
+ omp-conductor doctor [--project NAME] [--json] [--probe-telegram]
2132
+ omp-conductor ledger [--issue N] [--limit N] [--project NAME]
2133
+ omp-conductor board [--project NAME]
2134
+ omp-conductor hold [--keep-ticks] [--project NAME]
2135
+ omp-conductor stop [--pane] [--project NAME]
2136
+ omp-conductor arm [--project NAME]
2137
+ omp-conductor disarm [--project NAME]
2138
+ omp-conductor tail <issue> [--project NAME]
2139
+ omp-conductor extend <issue> --turns N [--project NAME]
2140
+ omp-conductor worker pause <issue> [--project NAME]
2141
+ omp-conductor worker resume <issue> [--project NAME]
2142
+ omp-conductor worker stop <issue> --reason TEXT [--project NAME]
2143
+ omp-conductor unblock <issue> [--force] [--no-requeue] [--project NAME]
2144
+ omp-conductor verb <conductor_*> [--project NAME] [--arg k=v ...]
2145
+ omp-conductor friction <escalation-digest|report-noise|report-surprise> --detail TEXT [--issue N] [--project NAME]
2146
+ omp-conductor event record --category NAME --summary TEXT --evidence REF [--occurred-at ISO] [--project NAME]
2147
+ omp-conductor report --text TEXT [--kind material|digest|tier2|decision-needed|fleet-stopped|confirmed-failure] [--events IDS] [--notices IDS] [--project NAME]
2148
+ omp-conductor message --text TEXT [--project NAME]
2149
+ omp-conductor decision open --question TEXT [--blocks TEXT] [--resolves-when COND] [--project NAME]
2150
+ omp-conductor decision resolve <id> --answer TEXT [--project NAME]
2151
+ omp-conductor decision withdraw <id> [--reason TEXT] [--project NAME]
2152
+ omp-conductor decision list [--project NAME]
2153
+ omp-conductor daemon [--once] [--port N] [--project NAME]
2154
+ omp-conductor resume [--project NAME]
2155
+ omp-conductor brief-upgrade [--migrate|--retrofit] [--apply] [--file PATH] [--project NAME]
2156
+ omp-conductor help
2157
+ ```
2158
+
2159
+ | Command | Behaviour |
2160
+ | --- | --- |
2161
+ | `setup [area] [--no-ai] [--project NAME]` | The deterministic interview, in a plain terminal — the same prompts, the same one-writer apply sequence, and the same single consent gate as `omp-conductor setup`, which is now one dialog implementation of the shared surface rather than the only way in. Bare is a full first run, or — when the project already exists — a chooser of which area to amend. Naming an area positionally skips that chooser and amends only that area: `tracker`, `gates`, `caps`, `code-graph`, `authority`, `policy`, `escalation`, `reporting`, `brief`. `host` and `graph` are install subcommands rather than areas and are matched first; anything else exits `2` listing both vocabularies. Every prompt shows its current value as the default, and Enter accepts what you see; `Ctrl-C` at any prompt abandons the run and writes nothing. Setup also **reads your repos to propose answers**: the gates prompt is pre-filled from what CI actually runs, and the brief's `## Project context` and release procedure are drafted from every routing repo and shown for confirmation before anything is written. Each probe is a short session with **no shell, no editor and no verbs** in a throwaway shallow clone, and every answer is a proposal you edit or decline — a probe that cannot clone, cannot reach a model, or answers unusably costs you one warning and the shipped stub. `--no-ai` asks every question with the reading half removed. |
2162
+ | `setup host [--project NAME]` | Re-render and stage the systemd unit, then **run** the install: `install -m 0644` into `/etc/systemd/system`, `daemon-reload`, `enable`, `restart`. Stages the fleet recovery oneshot (`omp-conductor-recover.service`) and its playbook (`/usr/local/sbin/omp-conductor-recover`) alongside, and installs them **before** the fleet units: both fleet units carry `OnFailure=` to the recovery unit, so a crash-looped daemon or herdr session now collects evidence durably, attempts one bounded recovery, and pages tier-2 instead of dying silently (#485). Every command is shown with its exact argv, one confirm covers the batch, and `sudo` asks for your password once before the first step — or is skipped entirely on a fleet that genuinely runs as root. The first failure stops the rest and prints the un-run remainder verbatim so you can finish by hand. Refuses an *escalated* invocation (`sudo`, or `sudo -i`/`su -` detected by the invoking account disagreeing with the fleet's) before writing anything, naming both accounts, because staging derives the unit's `User=`/`HOME=` from whoever ran it. On a non-Linux host the files are still staged and only the `systemctl` steps are refused. |
2163
+ | `setup graph [--no-seed] [--print] [--project NAME]` | The code-graph install end to end, in one preview and one confirm: check the prerequisites read-only and stop before installing anything when `codebase-memory-mcp` is absent or no MCP entry mounts it (printing the entry to add); `git clone` each missing index-only checkout **as you, never through sudo**; install and enable `cbm-reindex.timer` as root; then seed one indexing run so the first fetch happens while you watch, and verify with the same probe `status` uses. A repo that does not verify is a failure with the remediation, not a success — staged-but-not-trusted is how you discover months later that no worker read an index. `--no-seed` enables the timer without the seeding run and says plainly the graph is unusable until it first fires; it never skips the prerequisite or clone steps. `--print` changes nothing. Exits `1` when no repo has [`graphProject`](#configuration). |
2164
+ | `start` | Start `herdr-fleet.service` when that optional unit is installed, clearing a previous pane-recovery pin, then start the dispatch daemon and wait until it answers `GET /healthz`. When `omp-conductor.service` is installed, systemd is the only start path: even `start --project NAME` restores the shared unit and uses the name only to verify that `/healthz` serves the requested project. A detached daemon is allowed only when the unit is proven absent. It never clears pause or arms ticks. Refuses if a daemon is already live, naming its pid; manager refusal or unprovable ownership is an error rather than a detached fallback. |
2165
+ | `stop` | Prefer `systemctl stop omp-conductor.service` when that unit's MainPID is the live daemon — systemd then owns the stop and will not schedule a restart for the exit it just requested. Otherwise `SIGTERM`, then `SIGKILL` after a 10-second grace period. Prints `not running` when there is nothing to stop, and tags the confirmation with `(via systemctl)` when the unit path was used. |
2166
+ | `restart [--now] [--timeout SECONDS] [--port N] [--project NAME]` | Drains the fleet by default: pause new claims, wait until live workers reach `0 / N` (bounded by `--timeout SECONDS`, default 1800 = 30 min), restart, then restore the prior dispatch state. A daemon serving multiple configured projects makes restart host-wide: `--project` is rejected because draining one queue and restarting the shared process would kill another project's workers. Prefer `systemctl restart` when the unit owns the live pid so the replacement stays supervised; only a host proven not to have the installed unit may fall back to the standalone stop/start path. `--now` skips the drain and restarts immediately, orphaning any live runs (old behaviour). A drain that hits `--timeout` restarts nothing and leaves dispatch paused — `omp-conductor resume` lifts it, or re-run `restart` to keep waiting. The new process **salvages dirty live worktrees before orphaning** those rows — see [Deploying a new package onto a busy fleet](#deploying-a-new-package-onto-a-busy-fleet). |
2167
+ | `upgrade [--to VERSION] [--project NAME]` | Deterministically update the Bun-global CLI, omp plugin, Herdr recovery plugin, and managed brief as one release. Resolves the npm version and exact `gitHead`, pauses only new claims, drains active workers, installs all surfaces, reloads Herdr and the daemon, waits for pane recovery, verifies identities and fleet health twice, then restores the original dispatch state. Host-wide by default: one daemon serves every configured project, so a bare run drains all of them and refreshes every brief. `--project` is rejected when the live daemon serves several projects — draining one queue and restarting the shared daemon would kill another's workers. A no-op when already current. Failure leaves dispatch paused. Must run outside a Herdr-managed session. The same transaction runs detached, without a human, as the fleet-installs-itself path: the orchestrator calls the `conductor_install` verb under the granted `install` shape, the daemon validates the version against npm and starts a transient systemd unit (`upgrade-install`) outside the pane and the daemon, and the first tick after the restart verifies version, `/healthz`, ticks, pane and `doctor` against the durable upgrade journal before restoring dispatch — rolling back and paging tier-2 on any gap. |
2168
+ | `status [--project NAME]` | Layered fleet report first: `dispatch` / `ticks` / next scheduled tick / `pane` / `recovery` / `herdr` / `telegram` / `brief` / `decisions` / optional `failure classes` and `code graph` / `daemon`, then the project body. The project body includes the latest completed dispatch timestamp, ready/routed/admitted counts, bounded hold groups, and the GitHub API budget (`graphql` / `core` remaining and reset, in the caps block); API failures are marked `DEGRADED` so queue starvation cannot look idle. Active-run lines overlay cooperative worker `paused`/`pausing` from `/healthz` without changing SQLite `running` state or the live worker count. The next tick comes from the live heartbeat process, not a guess from log timestamps. Telegram health uses `getMe` to prove API authentication without sending a message and reports inbound bridge configuration separately. Configured graphs report prerequisites, indexed repos, timer state, and refresh freshness without blocking dispatch. A `reports` block lists everything the outbox has not delivered, with its age, and prints `pending` (nobody has it) differently from `SENDING` (outcome unknown, it may already have arrived) — see [Report delivery](#report-delivery-the-outbox). The daemon block includes `rss` from `/healthz`; live workers add a busy-deploy warning. A `.conductor-stalled` marker adds an `orchestrator STALLED since …` line. |
2169
+ | `ledger [--issue N] [--limit N]` | The action audit: every [mediated-verb](#the-mediated-verbs-126) mutation and every next-attempt turn budget. Verb entries include the arguments, decision, named refusal, and resulting SHA. Turn-budget entries remain after an override is replaced or consumed. Reads (`conductor_pr_status`) are absent so polling cannot bury the signal. `--issue` narrows both histories; `--limit` defaults to 50. Recent verb refusals and pending turn overrides also appear in `status`. |
2170
+ | `board [--project NAME]` | Live keyboard-driven kanban over the same SQLite and `/healthz` truth as `status`, plus the tracker's current labels: Queue, Claimed, Running, Green, Blocked, Failed, Orphaned, the last 24 hours of Merged and Settled, and Parked (an issue the tracker has not confirmed closed — still open, or a label read that failed — so nothing dispatches it until a human labels it). Columns are mutually exclusive and describe current state, not the newest run row, so a requeued issue is queued rather than failed and a closed issue is neither. Refreshes run/spend/turn values every second, and health plus the label read every ten seconds. `Enter` follows the selected transcript in place; `u` invokes the existing unblock workflow on a Blocked, Failed, or Orphaned card; `i` / `p` open the issue / PR; `r` refreshes health; `?` shows all keys. Requires an interactive terminal of at least 50×20. |
2171
+ | `hold [--keep-ticks] [--project NAME]` | Soft stop: pause claiming **and** disarm ticks. Daemon and pane stay up. Prefer this when the intent is "stop the conductor" without killing processes. `--keep-ticks` pauses claiming but leaves the arm marker, so the heartbeat keeps reporting and `resume` alone restores the fleet — no fresh arm challenge. See [Stop the conductor](README.md#stop-the-conductor-hold--stop). |
2172
+ | `stop [--pane] [--project NAME]` | Stop the conductor: pause claiming, disarm ticks, then stop the dispatch daemon (systemctl-aware). Pane stays up unless `--pane` is passed. `stop --pane` also pins herdr-conductor recovery off for the conductor agent only — it does **not** stop `herdr-fleet.service` or any other herdr session. Fail-closed: exits nonzero unless the agent is proven gone. To bounce the daemon without stopping the fleet, use `restart`. |
2173
+ | `arm [--project NAME]` | Proof-gated: send a Telegram challenge and write this project's arm marker only after your reply appears as a user turn in the orchestrator transcript. The challenge names the project, so a host running two fleets is not ambiguous. Never auto-armed by `resume` / `hold`. |
2174
+ | `disarm [--project NAME]` | Remove this project's arm marker so its ticks skip; another project's ticks keep running. Also clears a pre-per-project shared `armed` marker while that marker is still what holds this fleet's gate open — otherwise the disarm would not disarm. Processes untouched. |
2175
+ | `tail <issue>` | Follow the newest run for that issue: the worker's assistant text as `assistant: …` and each tool it calls as `tool: <name>`, printed as they land. Workers are omp sessions inside the daemon rather than terminals, so this is the only way to watch one live — a herdr pane running it becomes an observation window. Starts from the top of the transcript, not the end, so attaching to a run that is already ten turns in shows those ten turns. Exits `1` with `no run recorded for #N` when the issue has never been dispatched, or `no transcript yet (state: …)` when the attempt has not opened one. Otherwise it runs until `Ctrl-C`, or until the run has finished and its transcript has been silent for five seconds, and prints `run ended: <state>`. |
2176
+ | `extend <issue> --turns N [--project NAME]` | Raise a live worker's effective turn ceiling through its owning daemon without restarting its session. If the latest run is failed, killed, orphaned, or blocked and has no live controller, store a one-shot ceiling for that issue's next claimed attempt instead. A next-attempt value must exceed the project base, every extension must stay at or below `workerMaxTurnsCeiling`, and live extensions remain monotonic. The pending value appears in `status`, is recorded in `ledger`, and is consumed atomically by one claim. |
2177
+ | `worker pause <issue>` / `worker resume <issue>` | Cooperatively park one live worker without changing its run state or lane. Pause aborts the active turn to harness idle and freezes the remaining wall-clock budget; resume continues the same session with a prompt to re-check its last action before repeating it. This is separate from fleet-level `hold`, which refuses new claims and work-starting mutations while allowing pre-pause completion work and releases. |
2178
+ | `worker stop <issue> --reason TEXT [--project NAME]` | Terminally end a running or cooperatively paused worker. The reason is required (1–500 characters) and persisted on the run. The command waits for settlement, records the distinct `stopped` state, salvages and publishes dirty work, removes `agent:in-progress` through the durable label outbox, and consumes neither failed-attempt nor continuation budget. If salvage fails, the tree holding the only copy stays in place and the command names it. Repeating stop is idempotent and reports the run's already-terminal state. |
2179
+ | `unblock <issue> [--force] [--no-requeue]` | Remove that issue's `blocked` and `failed` labels so an answered escalation can be claimed again, and restore the project queue label by default so the dispatcher actually sees it. `agent:in-progress` comes off too, but only when the newest recorded run is terminal — that row is the proof no worker still owns the issue, so a live run keeps the label (and the queue label stays off until that run settles), and so does an issue with no run row at all. Run history remains intact: blocks consume the independent continuation budget, not failed implementation attempts. The output reports both budgets and warns when either will make the next tick escalate instead of dispatch. The label changes go through the [label projection outbox](#how-one-tick-works): they are applied inline before the command returns, but **a tracker that refuses them (403, rate limit) no longer fails the verb** — it exits `0`, the intended label state is durable and the daemon retries it, and the output says `label sync queued (N pending) — the daemon retries` instead of claiming the labels were restored. Safety is preserved, but the issue is only claimable once the queue label itself lands: the queue read asks GitHub for issues carrying that label, so a refused queue-label add keeps the issue out of dispatch until projection succeeds. `--no-requeue` clears the state labels only, leaving the queue label untouched — the case where you are about to close the issue. **Refuses, clearing nothing and exiting `3`, when the newest attempt's work could not be committed and its worktree is the only copy** — re-claiming removes that tree. `--force` records the operator's acceptance on the run row and then clears; the salvage failure stays in history. Exits `2` when the issue number is missing or malformed. |
2180
+ | `verb <conductor_*> [--arg k=v ...]` | Run one [mediated verb](#the-mediated-verbs-126) as the orchestrator, from the CLI — the external-orchestrator half of the verb surface. Every argument goes in as a `--arg k=v` string; an orchestrator can merge (`conductor_pr_merge`), label (`conductor_label`), release (`conductor_release`), update a branch (`conductor_pr_update_branch`) or title/body (`conductor_pr_update`), or read PR state (`conductor_pr_status`). The daemon applies the same checks and writes the same ledger rows a session's call would; a missing `--arg` is refused exactly as a missing tool argument is, worker-only verbs (`conductor_push`, `conductor_pr_create`) are refused with `role-not-allowed`, and a refusal exits `3`. An unknown verb exits `2`. |
2181
+ | `friction <kind> --detail TEXT [--issue N]` | Record one bounded judgment the daemon cannot infer: an escalation belonged in a digest, or a tick report was noise/surprising. The detail is limited to 160 characters. One event never changes policy; three observations inside seven days make the aggregate eligible for one Learning-loop prompt, followed by a seven-day cooldown. |
2182
+ | `report --text TEXT [--kind material|digest|tier2|decision-needed|fleet-stopped|confirmed-failure]` | Hand a rendered report to the daemon's durable outbox. The command persists the text **before** anything can send and prints a durable handoff id. A material report submitted during quiet hours becomes a held notice until the window opens; otherwise it becomes a report whose delivery the daemon owns, retries with bounded backoff, and records. Delivery is [at-least-once](#report-delivery-the-outbox), so a crash mid-send is retried as a possible repeat and `delivered` never proves exactly one message. `--kind digest` is accepted at most once per local day, decided from the ledger; an unknown `--kind` exits `2`. The remaining kinds declare the report's interrupt category — the escalation handoff: the reporting policy decides between immediate delivery and a durable hold exactly as for a daemon escalation of that category, A repeated identical call exits `2` only while the earlier handoff is still queued undelivered; once it lands, the same text is admitted again (the handoff state decides, not a permanent ledger). Anything still owed appears in `status` with its age. |
2183
+ | `decision open --question TEXT [--blocks TEXT] [--resolves-when COND]` | Record a question the orchestrator has put to you, and print its id. A question that lives only in a session's context is lost at the next compaction — after which it is either asked twice or dropped silently. `--resolves-when` attaches a machine-checkable condition: `pr-merged:<https url>`, `pr-checks-green:<https url>`, `pr-mergeable:<https url>`, `issue-closed:<n>`, `npm-version:<pkg>@<version>`, or `rate-limit-reset:github`; anything else exits `2` listing the six forms. See [The decision ledger](#the-decision-ledger-136). |
2184
+ | `decision resolve <id> --answer TEXT` | Record what you decided. Exits `1` naming the id when it is unknown or no longer open, so a second answer cannot overwrite the first. |
2185
+ | `decision withdraw <id> [--reason TEXT]` | Close a question the session stopped needing, with why. Same guard as `resolve`. |
2186
+ | `decision list` | Open questions, oldest first: id, age, what each blocks, whether its condition is met, and the question. Prints `no open decisions` when there are none. |
2187
+ | `daemon` | Run the loop in the **foreground**, ticking every 5 minutes and serving `/healthz`. Admitted workers run in a tracked background pool, so settlement and capacity checks remain periodic while they work; shutdown drains the pool before closing the store. This is what `start` launches and what a systemd unit should call. |
2188
+ | `daemon --once` | Run a single tick, wait for workers admitted by that tick, and exit. No HTTP server or pidfile — a drill must not register itself as the daemon, or the next reader believes it and the real daemon's in-flight runs get reconciled as orphans. |
2189
+ | `--port N` | Accepted by `start`, `restart` and `daemon`. Both `--port 9000` and `--port=9000` work; missing or out of range exits `2` rather than falling back to the default, because probing the wrong endpoint is worse than a hard failure. |
2190
+ | `--project NAME` | Selects a project for project-scoped commands and foreground `daemon`. On an installed shared service, `start --project NAME` still starts the host-wide unit and uses the name only to verify `/healthz`; a draining `restart --project NAME` is rejected when that daemon serves multiple projects. A project-only daemon is available only through an explicit foreground `daemon --project NAME` or standalone start on a host proven not to have the unit. |
2191
+ | `pause [--reason TEXT]` | Stop new claims and work-starting mutations only. The running daemon notices on its next tick; runs already in flight finish. The orchestrator may still merge, update, or label runs admitted before the pause, and may release when the release policy's own preconditions hold. The orchestrator heartbeat keeps ticking if armed — its gate is the arm marker, not this flag. Per-worker pause is separate. Prefer `hold` to silence both. `--reason TEXT` is recorded in the pause sentinel, which `status` shows as the pause provenance. |
2192
+ | `resume [--project NAME]` | Clear pause and any `stop --pane` recovery pin — does **not** re-arm. Run `arm` after an inbound Telegram proof to resume ticks. |
2193
+ | `--version`, `-V`, `version` | Print the installed `omp-conductor` package version and exit `0`. Works from the global binary and npm/plugin install because it reads the package metadata beside the shipped CLI. |
2194
+ | `brief-upgrade` | Inspect the package-floor + shared policy (when present) + `POLICY.md` overlay. Reports by default, naming which layers compose; see [Keeping a brief current](#keeping-a-brief-current). |
2195
+ | `--migrate` | Only for `brief-upgrade`. Lift a bannered `ORCHESTRATOR.md` owned half into `POLICY.md` and recompose. Dry-run unless `--apply`. |
2196
+ | `--retrofit` | Only for `brief-upgrade`. Propose (or with `--apply`, write) a `YOURS TO EDIT` banner before the first owned-topic heading on a hand-written brief. |
2197
+ | `--apply` | Only for `brief-upgrade`. Confirms `--migrate` / `--retrofit`. On its own it exits `2`: the legacy single-file merge was removed in 0.4.3. |
2198
+ | `--file PATH` | Only for `brief-upgrade`. Check a brief that is not where the wizard would have put it, on a host that may have no config at all. |
2199
+ | `help`, `--help`, `-h` | Print usage. An unknown or missing verb prints it too, and exits `2`. |
2200
+
2201
+ Pause is a sentinel under the state directory and survives a daemon restart.
2202
+ `hold --project NAME` writes `paused-<name>` for that project only; a bare
2203
+ `paused` file (legacy / all-projects) pauses every project. It refuses new claims
2204
+ and work-starting mutations, allows completion verbs only for runs admitted
2205
+ before the pause, and leaves `conductor_release` to its normal authority, grant,
2206
+ and precondition checks. Per-worker pause is independent. Hold also removes the
2207
+ arm marker the heartbeat reads, so both brains go quiet without killing processes.
2208
+
2209
+ Every one of these is a verb on the `omp-conductor` binary, each taking an optional
2210
+ `--project NAME`. There is no in-session command: an omp session that wants any of
2211
+ them shells out to the binary, which is what keeps one implementation and one ledger
2212
+ entry per action.
2213
+
2214
+ ### Health endpoint
2215
+
2216
+ ```bash
2217
+ curl -s localhost:8787/healthz
2218
+ ```
2219
+
2220
+ ```json
2221
+ {
2222
+ "ok": true,
2223
+ "rssBytes": 123456789,
2224
+ "projects": [
2225
+ {
2226
+ "ok": true,
2227
+ "paused": false,
2228
+ "activeRuns": 1,
2229
+ "project": "demo",
2230
+ "dispatch": {
2231
+ "completedAt": 1786185678000,
2232
+ "ready": 8,
2233
+ "routed": 8,
2234
+ "admitted": 0,
2235
+ "degraded": true,
2236
+ "holds": [
2237
+ { "reason": "parent-lookup-error", "count": 8, "issues": [321, 320, 318] }
2238
+ ]
2239
+ },
2240
+ "codeGraph": {
2241
+ "configured": true,
2242
+ "status": "degraded",
2243
+ "checkedAt": "2026-08-08T13:00:00.000Z",
2244
+ "prerequisites": { "indexer": "present", "mcpMount": "missing" },
2245
+ "repos": [
2246
+ {
2247
+ "name": "api",
2248
+ "path": "/home/fleet/.cache/conductor-graph/acme/api",
2249
+ "clone": "present",
2250
+ "index": "present"
2251
+ }
2252
+ ],
2253
+ "timer": { "enabled": "enabled", "active": "active" },
2254
+ "refresh": {
2255
+ "result": "success",
2256
+ "fresh": true,
2257
+ "lastSuccessAt": "2026-08-08T12:50:00.000Z",
2258
+ "ageMs": 600000
2259
+ },
2260
+ "reasons": ["worker MCP configuration does not mount the indexer"]
2261
+ }
2262
+ }
2263
+ ]
2264
+ }
2265
+ ```
2266
+
2267
+ Any other path or method returns `404`. Top-level `ok` is process liveness across
2268
+ every served project; top-level `rssBytes` is the daemon's resident set.
2269
+ Per-project blocks keep `paused`, `activeRuns`, `dispatch`, `codeGraph`, and
2270
+ `workers`. Nonfatal admission errors and graph degradation keep `ok` `true` so a
2271
+ supervisor does not restart-loop. Inspect `dispatch.degraded` and its bounded
2272
+ reason groups for queue starvation; inspect `codeGraph` for configured graph
2273
+ health. `activeRuns` counts occupied issues — live workers plus green PRs
2274
+ awaiting merge.
2275
+
2276
+ ## What a worker may and may not do
2277
+
2278
+ Each worker gets one brief, one worktree, one branch, and no knowledge of the
2279
+ dispatcher. The brief is explicit about the boundary:
2280
+
2281
+ | It may | It must not |
2282
+ | --- | --- |
2283
+ | Read the issue and the repo's own guidance (`AGENTS.md`, `CLAUDE.md`, `CONTEXT.md`, relevant ADRs) before writing anything. | Touch any path outside its worktree, or switch branches. |
2284
+ | Query its repo's [code graph](#code-graph-discovery), when one is configured, by the project name whose `root_path` matches the clone its brief names. | Query that graph by its own cwd or worktree path — no index of a worktree exists — or treat what it returns as current. It is a snapshot of the clone's default branch; the real file in the worktree wins. |
2285
+ | Edit code inside its own worktree. | Weaken, skip, delete or loosen **any test it did not write** — that is a design question to escalate, and it is checked by diff review before the push. |
2286
+ | Add or update tests for behaviour it introduced. | Suppress a warning, delete an assertion, or special-case an input to make a check pass. |
2287
+ | Run the repo's configured cheap gates, each from its listed `cwd`, over the whole tree. | Run docker or image builds, production builds, browser/e2e suites, or the full test suite on the shared host — CI owns the heavy gates. |
2288
+ | Review its whole diff, then commit and publish once with `conductor_push`. One corrective push if CI is red. | Force-push, `git add -f`, or add AI/co-author attribution. There is no force path to reach: `conductor_push` publishes that run's branch fast-forward only and takes no other ref. Red twice means stop and report, not push a third time. |
2289
+ | Open a PR with `conductor_pr_create`, and poll CI to a verdict with `conductor_pr_status`. | Run `gh pr merge` — or reach `conductor_pr_merge`, which refuses a worker session mechanically. **A worker is never authorised to merge**, whoever else holds the authority, so PRs land one at a time with a freshness re-check; two workers merging concurrently is how agent PRs clobber each other. Who *may* merge is the [`authority`](#configuration) answer, and it is never the worker. The verb refusal is mechanical, and so is the channel: each one is bound to the pid the daemon spawned, so reaching for the orchestrator's socket is refused rather than honoured (see [The transport](#the-transport)). Shelling out to `gh` remains a prohibition, not an impossibility — a session shares the daemon's credentials. |
2290
+ | Escalate: ambiguity, a cross-repo contract, a needed credential, a product or data-migration decision, a blocking existing test, CI red twice, or most of the wall-clock budget burned. | Cut a release, push a tag, publish to npm, edit a deployment pin, deploy, or touch infrastructure or secrets. `conductor_release` and `conductor_label` refuse a worker whatever `releasePolicy` says, because the check compares the caller against the configured holder rather than ruling one value out. The in-session tripwire still blocks recognised release/deploy tool calls early and audits the attempt, but it is [defence in depth](#the-mediated-verbs-126), not the gate. |
2291
+
2292
+ The worker ends with a seven-line evidence report (issue, PR, observed head SHA,
2293
+ state, gates, changed, next). A textual `pushed-green` claim is not success: the
2294
+ daemon repeats the PR/head/check verification before it records that state.
2295
+
2296
+ ### Worker confinement and the integrity tripwire
2297
+
2298
+ A worker session is rooted at its worktree `cwd`. **Structured file tools are
2299
+ gated mechanically:** `runWorker` asks `createSession({ role: "worker" })`,
2300
+ which installs an inline harness extension that blocks `write` / `edit` /
2301
+ `read` / `grep` / `glob` when the tool's path resolves outside that worktree
2302
+ (symlink-aware). Target selection was already mechanical — only a repo in
2303
+ `routing.repos` is ever checked out — and the caps still bound *how much* work
2304
+ happens.
2305
+
2306
+ General shell access is not confined to the worktree. Its argument is an opaque
2307
+ program, so the brief still forbids path escape and the deploy-level answer is a
2308
+ least-privilege worker uid (below). The narrower release-policy tripwire does
2309
+ inspect explicit command shapes such as `git tag`, `npm publish`, and deploy
2310
+ verbs; it blocks those before execution when `releasePolicy` is `none`.
2311
+
2312
+ #### Integrity tripwire (package self-hash)
2313
+
2314
+ Separately, the conductor watches *itself*. At startup the daemon sha256s every
2315
+ `.ts` and `.md` file of its own installed `src/` — the dispatcher and the briefs
2316
+ both, since rewriting a brief buys more than rewriting the loop — and re-walks
2317
+ that tree on every tick (about 0.6 ms). Any difference at all, changed or added
2318
+ or removed, is read as the package having been modified underneath a running
2319
+ daemon: the tick claims nothing, the fleet is paused, and a tier-2 escalation
2320
+ naming the first few differing paths pages you **once**, not every five minutes.
2321
+
2322
+ **A normal deploy never trips it.** The baseline is recorded per daemon process,
2323
+ so installing a new build and restarting the unit re-records it from the new
2324
+ files; only a change that lands *while* a daemon is holding the package open can
2325
+ diverge from it. That also means `omp-conductor resume` on its own will not hold
2326
+ — the next tick re-walks, still differs, and pauses again. Put the files back, or
2327
+ restart onto the build you meant to be running.
2328
+
2329
+ This catches a worker (or human) that still managed to edit the live install —
2330
+ including via `bash` — after the fact. It is detection for the package boundary,
2331
+ not a substitute for the worktree gate or a dedicated uid.
2332
+
2333
+ #### Least-privilege worker uid (deploy)
2334
+
2335
+ The largest remaining win is OS-level: run the daemon (or at least worker
2336
+ sessions, when the harness supports a uid switch) as a user that can write only
2337
+ its worktrees and mirrors. A sketch that matches the reference single-host
2338
+ deploy:
2339
+
2340
+ 1. Create a system user, e.g. `conductor-worker`, with home under
2341
+ `/var/lib/conductor-worker` (or similar).
2342
+ 2. `chown` the project's `workspaceRoot` and `mirrorRoot` to that user; leave
2343
+ `~/.omp/conductor/config.json` readable only by the operator/daemon account
2344
+ (`0600` as shipped).
2345
+ 3. Do **not** put the worker uid in `docker` / `sudoers`, and do not give it the
2346
+ operator's `gh` auth if a narrower deploy token can open PRs in the routed
2347
+ repos alone.
2348
+ 4. Point the [example systemd unit](systemd/omp-conductor.service.example)
2349
+ `User=` / `Group=` at that account once the daemon itself should run
2350
+ unprivileged end-to-end.
2351
+
2352
+ Until that uid exists, a root-or-operator daemon still has a mechanical
2353
+ worktree gate on structured tools and an integrity tripwire on its own package —
2354
+ but `bash` plus host credentials remain a prompt-and-deploy problem.
2355
+
2356
+ ### The orchestrator is unconfined, deliberately
2357
+
2358
+ There is **no mechanical file gate on the orchestrator session**, and that is an
2359
+ operator decision rather than an omission (#143).
2360
+
2361
+ A previous release jailed it to an allowlist. That gate could only ever be
2362
+ installed by `createLocalSession`, so it existed exactly in the sessions this
2363
+ daemon spawns — and the supported shape for a heartbeat orchestrator is an
2364
+ `omp` session the operator starts themselves, which never had it. A boundary
2365
+ present in one deployment out of two is not a boundary, and the brief asserting
2366
+ it was absolute was the worse half of the bug: a session that believes it is
2367
+ gated stops checking itself.
2368
+
2369
+ What holds the orchestrator instead:
2370
+
2371
+ | | |
2372
+ | --- | --- |
2373
+ | **The brief** | `ORCHESTRATOR.md`'s hard boundaries — never read or edit a worker's checkout or the mirror cache; when you need a run's code, read its PR. |
2374
+ | **The action ledger** | Every `conductor_*` mutation and operator-selected next-attempt turn budget remains auditable. `omp-conductor ledger` shows both, including refused calls and consumed or replaced budget overrides. |
2375
+ | **The dispatcher** | Merge, label and release authority are checked in the daemon against the operator's grant, across a process boundary, never in the prompt. |
2376
+
2377
+ Unconfined means auditable, not licensed. `orchestratorReadPaths` is retired: it
2378
+ is still accepted in a config and ignored, so a fleet carrying it upgrades
2379
+ without editing anything.
2380
+
2381
+ ## The mediated verbs (#126)
2382
+
2383
+ A session can reach `gh`: it inherits the daemon's environment, credentials and
2384
+ all. It is told not to publish with it. These verbs are the sanctioned route
2385
+ instead, because the dispatcher owns the settlement record — a push or a PR the
2386
+ daemon did not perform is a run it cannot account for, and the checks that would
2387
+ have refused it never ran. The point of them is *where those checks run*: in the
2388
+ daemon, across a process boundary, not in a prompt the model can rewrite.
2389
+
2390
+ ### The verbs
2391
+
2392
+ | Verb | Allowed caller | What the daemon checks before acting |
2393
+ | --- | --- | --- |
2394
+ | `conductor_push` | the worker owning the run | The ref is exactly `refs/heads/<that run's branch>`. Fast-forward only; there is no force argument to reject because none is declared. |
2395
+ | `conductor_pr_create` | the worker owning the run | The run has no open PR (the same guard admission uses); head is the run branch; base is the repo's configured `defaultBranch`. |
2396
+ | `conductor_pr_status` | worker or orchestrator | Read-only — nothing to gate. A worker reads only its own run's PR; an orchestrator may name any syntactically valid PR URL, open, merged, or closed, and gets its live state and head (checks are reported when available; a merged or closed PR reports its state instead of an `expected OPEN` refusal). |
2397
+ | `conductor_pr_update_branch` | orchestrator, or the worker owning the run | The PR belongs to this project and is open. A worker may only name its own run's PR. |
2398
+ | `conductor_pr_merge` | **orchestrator only** | Ordinarily, `authority.merge` equals the caller. A hand-edited `recoveryMerges` entry may instead authorize one exact unrecorded PR/head/reason while held. In both paths, `headSha` equals the live head *at execution time*; checks are green at that same SHA; the project route and migration chain are valid; the project's single merge slot is free. |
2399
+ | `conductor_label` | **orchestrator only** | The label is in the project's own vocabulary. Lifecycle labels stay the daemon's. |
2400
+ | `conductor_release` | **orchestrator only** | `authority.release` equals the caller; the per-shape grant permits it; the artefact or environment was declared; the release preconditions hold; the `reason` is in the closed enum. `version-bump-pr` creates or re-validates one deterministic version-only PR and, on a later call, merges only its exact green head through the project's single merge slot. |
2401
+ | `conductor_install` | **orchestrator only** | Gated like a release act: the `install` shape defaults to `human` and a grant is what moves it. The daemon refuses a version npm does not expose with a full `gitHead`, refuses while another install is still in flight, and otherwise starts a detached transient unit that pauses, drains, installs the CLI/omp plugin/Herdr plugin and reloads — outside this session and the daemon. The unit never declares its own success; the first tick after the restart verifies and reports through the durable outbox. |
2402
+
2403
+ Standing merge and release authority use the same exact rule: **the caller's
2404
+ role must equal the configured holder.** `authority` has exactly two values, so
2405
+ a `!== "human"` test would have let a *worker* release. A worker is refused
2406
+ every release shape under the most permissive config there is. The exact
2407
+ operator-authored recovery tuple below is the sole authority exception inside
2408
+ `conductor_pr_merge`; a reviewed version bump is instead a `conductor_release`
2409
+ operation governed throughout by release authority.
2410
+
2411
+ `recoveryMerges` is deliberately narrower than standing merge authority. It is
2412
+ an operator-authored, one-PR escape hatch for a recovery branch that cannot have
2413
+ a run row—for example, a conflict repair created after the fleet was held. It
2414
+ does not admit new work, unpause the fleet, widen repository routing, bypass
2415
+ live-head or check validation, or make a general class of PRs mergeable.
2416
+ Authorizations are re-read from config on every call and every attempted merge
2417
+ is written to the ordinary verb ledger, including refusals.
2418
+
2419
+ For a repo with `release.versionFile`, call `conductor_release` with
2420
+ `shape=version-bump-pr` and the intended `v<semver>` tag. The requested version
2421
+ must be newer than the live semantic version. The first call creates
2422
+ `conductor/release-<version>` from the live default branch, changes only the
2423
+ declared JSON `version`, and opens a normal PR. Call it again after CI: the daemon
2424
+ re-reads that exact PR head, verifies the PR contains only the semantic version
2425
+ change, requires green checks, and merges with GitHub's exact-head guard. The
2426
+ ordinary action ledger records both calls. A raw source push is never delegated,
2427
+ and a worker makes no release decision.
2428
+
2429
+ For Git-backed releases, a repo that declares `release.versionFile` refuses both
2430
+ tag creation and a new tag push until the live default branch's version matches
2431
+ the requested tag. `git-tag` is idempotent when the named tag exists locally but
2432
+ has not been pushed: it re-points the tag to the verified live default-branch
2433
+ head. `git-push-tags` performs the same re-point immediately before pushing if
2434
+ the default branch moved between the two calls. A tag already published on
2435
+ origin is immutable: an identical tag is accepted as already complete, while a
2436
+ different published target is refused and must use a new tag name.
2437
+
2438
+ A `github-release` for the same repo likewise requires that reviewed tag to be
2439
+ present on origin and verifies the tag's version file before creating the
2440
+ release; it never lets GitHub synthesize the missing tag.
2441
+
2442
+ ### The transport
2443
+
2444
+ Identity is never an argument. `project`, `run`, `issue` and the caller's role
2445
+ come from **which socket the call arrived on**, and a request carrying any of
2446
+ those field names is refused outright, named. So a worker on run X cannot *ask*
2447
+ to merge run Y's PR — on its own channel that request is unexpressible.
2448
+
2449
+ **Each channel is bound to one process, because the modes cannot tell sessions
2450
+ apart.** Every session runs as the daemon's own uid, so it matches the *owner*
2451
+ class here: it can list this directory and connect to any socket in it, including
2452
+ the orchestrator's. Authorisation and the ledger both read the role from the
2453
+ channel, so a worker doing that would have been authorised as the orchestrator
2454
+ (under `authority.merge: "orchestrator"`) *and recorded as* the orchestrator. No
2455
+ file mode closes that — the owner bits belong to the uid the session already has.
2456
+
2457
+ So the daemon binds each channel to the **pid it spawned for that session**, and
2458
+ refuses a connection from anything else without answering it, logged the way an
2459
+ impersonation is. Until a channel is bound it refuses everything, because the
2460
+ socket necessarily exists before the child that connects to it. The kernel
2461
+ supplies the pid: `SO_PEERCRED` on Linux, `LOCAL_PEERPID` on macOS. A host where
2462
+ neither can be asked — no loadable libc, or the call refused — refuses every
2463
+ connection on a bound channel and says so at startup, rather than falling back to
2464
+ the uid, which under one shared uid is no check at all.
2465
+
2466
+ The residual is narrow, real, and worth stating: one uid can `ptrace` and signal
2467
+ its siblings, so a determined session can still interfere with the process that
2468
+ *is* bound. That is a far higher bar than connecting to a socket, and closing it
2469
+ needs separate OS principals.
2470
+
2471
+ ```
2472
+ <state dir>/verbs/ daemon-owned, mode 0711
2473
+ run-7-9a783d877d422b9e.sock 0600, bound for run 7
2474
+ run-9-1c40e2a5b6d3f018.sock 0600, bound for run 9
2475
+ orchestrator-4b1f...c2.sock 0600, the orchestrator's
2476
+ ```
2477
+
2478
+ **What these modes buy, and what they do not.** They keep every *other local
2479
+ account* out: `0711` on the parent is traversable but not listable, so no other
2480
+ user can enumerate the fleet's sockets, the suffixes are unguessable, and only
2481
+ the daemon's uid can connect to a `0600` socket at all.
2482
+
2483
+ They are **not** a boundary between runs. Sessions are child processes of the
2484
+ daemon running as its own uid, so a session matches the owner class on all of
2485
+ these: it could list the directory and connect to a sibling's socket. Each run is
2486
+ *handed* its own path and nothing else, which is a convention the run has no
2487
+ reason to break — not an enforcement. What makes breaking it visible is the
2488
+ [ledger](#the-ledger): every call is recorded with the channel it
2489
+ arrived on, so a worker calling on another run's socket is in the record.
2490
+
2491
+ Closing that properly needs the sessions to be different OS principals. A
2492
+ per-run credential boundary that did exactly this shipped and was removed in
2493
+ 0.5.0 — it worked, and the cost was that it also hid the operator's own model
2494
+ credential from every session, so nothing could start. It is not worth
2495
+ re-litigating without solving that first.
2496
+
2497
+ Before binding, the daemon verifies every component of the path is owned by
2498
+ itself (or root), free of symlinks, and unwritable by anyone else; a failed
2499
+ check **refuses dispatch** rather than degrading. Paths are unguessably
2500
+ suffixed, and only the daemon ever unlinks one.
2501
+
2502
+ Peer credentials are asserted server-side — `getpeereid` on macOS, `SO_PEERCRED`
2503
+ on Linux — and the daemon states at startup exactly what that buys rather than
2504
+ implying more. Sessions are child processes running under the daemon's own uid,
2505
+ so the peer check proves the caller is a local process on this host; it is the
2506
+ socket, not the uid, that says which run is calling. A connection whose peer
2507
+ cannot be read at all is closed with no reply and logged.
2508
+
2509
+ ```
2510
+ verb transport: verb sockets in ~/.omp/conductor/verbs (mode 711); each socket
2511
+ 0600 under the daemon's own uid; peer uid asserted with getpeereid
2512
+ ```
2513
+
2514
+ **No mutation route exists on the HTTP port**, and none may be added. That
2515
+ surface is unauthenticated loopback TCP reachable by any local user; a `PUT` or
2516
+ `POST` at any verb path answers 404, pinned by a test.
2517
+
2518
+ The child-side tool handler is a thin client only. It forwards arguments and
2519
+ renders the answer — no policy branch, no local fallback, no second route. With
2520
+ no socket it fails closed and says so, rather than reaching for `git push`.
2521
+
2522
+ ### The ledger
2523
+
2524
+ Every mutating verb call is recorded with its arguments, the decision, the
2525
+ named refusal reason and any resulting SHA. Reads are not: a status poll every
2526
+ thirty seconds would bury the refusals the record exists to surface.
2527
+
2528
+ Every `extend` that sets a next-attempt budget also appends an audit entry.
2529
+ Replacing or consuming the pending override does not erase that history.
2530
+
2531
+ ```console
2532
+ $ omp-conductor ledger --issue 7
2533
+ acme — 3 verb call(s), 1 refused (newest first)
2534
+ 2026-08-09 11:04:12 REFUSE conductor_pr_merge worker #7 [role-not-allowed]
2535
+ prUrl=https://github.com/acme/api/pull/7 headSha=9a783d8… reason=preconditions-met
2536
+ refused: merge authority is the orchestrator's, never a worker session's.
2537
+ 2026-08-09 10:58:03 ALLOW conductor_pr_create worker #7
2538
+ title=fix: settle the head check body=Closes acme/tracker#7
2539
+ opened https://github.com/acme/api/pull/7 (conductor/issue-7 → main).
2540
+ 2026-08-09 10:57:41 ALLOW conductor_push worker #7 9a783d877d42
2541
+ (no arguments)
2542
+ pushed refs/heads/conductor/issue-7 at 9a783d877d42….
2543
+ ```
2544
+
2545
+ The newest few also appear in `omp-conductor status`, because a refused merge is
2546
+ news: it means a session tried to do something the config does not permit.
2547
+
2548
+ `release-policy.ts` stays installed as defence in depth — it refuses early, in
2549
+ the session, with an explanation the model can act on in the same turn, and it
2550
+ leaves a durable record that something tried. It is no longer what *stops* a
2551
+ release. Treat a block there as evidence about a session's intentions; the
2552
+ daemon is what prevented it.
2553
+
2554
+
2555
+ ## Limitations
2556
+
2557
+ Known and deliberate in this version:
2558
+
2559
+ - **`gh` is shelled out to.** Every tracker operation spawns a process and does its
2560
+ own TLS handshake (roughly 200-400 ms each), and failures are classified by
2561
+ matching human-readable stderr rather than a status code. The upside is that no
2562
+ token is ever handled, stored or logged by the daemon.
2563
+ - **`listReady` fetches a single page of 100 issues.** A queue deeper than 100
2564
+ ready issues truncates silently. A backlog that size is a staffing problem before
2565
+ it is a paging one.
2566
+ - **Spend accounting depends on harness telemetry.** Cost arrives only when the
2567
+ harness run carries it; without it `spendUsd` reads `0`, `status` shows `$0.00`,
2568
+ and the daily-spend cap never fires. The turn and wall-clock ceilings are what
2569
+ actually bound a runaway in that case. Do not treat `$0.00` as proof that nothing
2570
+ was spent.
2571
+ - **GitHub is the only tracker.** The internal `Tracker` port is deliberately
2572
+ provider-neutral, but `tracker.kind` accepts only `"github"` today.
2573
+ - **One daemon serves every configured project by default.** `setup host` writes a
2574
+ unit without `--project`. Pass `--project NAME` only to filter a foreground or
2575
+ temporary daemon down to one project.
2576
+ - **Labels are matched exactly and case-sensitively.** `Ready-For-Agent` is not
2577
+ `ready-for-agent`, and the mismatch is silent: the issue is simply never picked
2578
+ up.
2579
+ - **No cross-process lock on the mirrors.** Two dispatch loops fetching the same
2580
+ repo at the same instant can collide on git's ref locks; the run fails and is
2581
+ retried rather than corrupted.
2582
+ - **Uniquely local mirror branches are retained.** Terminal runs are reaped
2583
+ automatically only after every commit exists on a remote ref. A failed salvage
2584
+ push deliberately leaves its branch and tree for an operator rather than
2585
+ trading disk hygiene for data loss.
2586
+ - **`stop` is a bounded best-effort drain.** A signal stops new ticks and the
2587
+ daemon waits for its active worker pool before closing the store. The CLI
2588
+ escalates to `SIGKILL` after 10 seconds, so a worker that needs longer is
2589
+ orphaned and salvaged on restart. Use `pause`, wait for `workers 0 / N`, then
2590
+ stop when a clean drain matters. A supervising unit should set
2591
+ `SuccessExitStatus=0 143`, and operators should prefer `omp-conductor stop` /
2592
+ `systemctl stop` over raw `kill`, so `Restart=on-failure` cannot misread a
2593
+ deliberate stop as a crash.
2594
+ - **A failed orchestrator degrades quietly.** The daemon logs a warning and keeps
2595
+ running, but tier-1 escalations then land in issue comments — which is exactly the
2596
+ "nobody reads it until morning" path the orchestrator exists to avoid. The warning
2597
+ is in `daemon.log`; nothing pages you about it.
2598
+ - **Workers are not terminal panes, so you cannot watch them there.** Each
2599
+ worker is an omp session the daemon starts as a child process. The resident
2600
+ daemon tracks workers in a background pool so the five-minute loop keeps
2601
+ settling PRs and checking capacity; shutdown waits for that pool. Herdr still
2602
+ shows exactly one pane (the orchestrator's) regardless of concurrency.
2603
+
2604
+ The cap does work. The admission loop (`admitCandidates` in `src/daemon.ts`) computes
2605
+ `slots = maxConcurrentWorkers - live workers`, admits at most that many issues
2606
+ per tick, and dispatches them together. To see them, read `omp-conductor
2607
+ status`, which lists every occupied issue, or follow `daemon.log`.
2608
+ - **Report delivery is at-least-once, never exactly-once.** The Telegram Bot API
2609
+ takes no client-supplied idempotency key, so the window between "Telegram
2610
+ accepted it" and "SQLite recorded that" is irreducible. The daemon resolves it
2611
+ toward a duplicate — the report is retried and the retry says it may be a
2612
+ repeat — because a duplicate you can recognise by its report id is cheaper
2613
+ than a silently dropped page. `delivered` means Telegram accepted an attempt,
2614
+ not that exactly one message exists. See
2615
+ [Report delivery](#report-delivery-the-outbox).
2616
+ - **Workers stop at green PRs.** They are never authorised to merge, release or
2617
+ deploy: those actions default to a human, and while setup may grant either to
2618
+ the orchestrator, `authority` never grants them to a worker or the dispatch
2619
+ daemon. The verbs refuse a worker mechanically, and a worker reaching for
2620
+ another session's channel is refused too — each channel is bound to the pid the
2621
+ daemon spawned for it. What remains a prohibition rather than a gate is shelling
2622
+ out to `gh` directly: sessions inherit the daemon's credentials. That shows up as
2623
+ a mutation with no matching ledger entry, which is a mismatch an operator can
2624
+ find.
2625
+ - **The worker gate is partial, and the orchestrator has none.** A worker's
2626
+ structured `write` / `edit` / `read` / `grep` / `glob` calls are gated to its
2627
+ worktree by an inline harness extension; `bash` is not, so a shell one-liner
2628
+ can still leave the tree, and no claim in this README says otherwise. The
2629
+ orchestrator is [unconfined on purpose](#the-orchestrator-is-unconfined-deliberately)
2630
+ — its boundaries are its brief and the verb ledger. Prefer a
2631
+ [least-privilege worker uid](#least-privilege-worker-uid-deploy); the
2632
+ [integrity tripwire](#integrity-tripwire-package-self-hash) still pages if the
2633
+ installed package itself changes under a live daemon.
2634
+
2635
+
2636
+ ## License
2637
+
2638
+ MIT