omp-conductor 0.15.10 → 0.15.12

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (92) hide show
  1. package/README.md +273 -2544
  2. package/REFERENCE.md +2680 -0
  3. package/package.json +3 -2
  4. package/schema/config.schema.json +11 -23
  5. package/src/arm-challenge.ts +112 -0
  6. package/src/ask.ts +434 -0
  7. package/src/board.ts +81 -15
  8. package/src/brief-upgrade.ts +114 -8
  9. package/src/briefs/orchestrator.md +113 -34
  10. package/src/briefs/policy.md +33 -8
  11. package/src/briefs/worker.md +18 -9
  12. package/src/chain-check.ts +1 -1
  13. package/src/check-trailing-newlines.ts +82 -0
  14. package/src/cli.ts +225 -1406
  15. package/src/commands/arm.ts +21 -0
  16. package/src/commands/board.ts +23 -0
  17. package/src/commands/brief-upgrade.ts +186 -0
  18. package/src/commands/context.ts +150 -0
  19. package/src/commands/daemon.ts +71 -0
  20. package/src/commands/dashboard.ts +74 -0
  21. package/src/commands/decision.ts +103 -0
  22. package/src/commands/disarm.ts +21 -0
  23. package/src/commands/doctor.ts +100 -0
  24. package/src/commands/event.ts +62 -0
  25. package/src/commands/extend.ts +64 -0
  26. package/src/commands/friction.ts +56 -0
  27. package/src/commands/help.ts +9 -0
  28. package/src/commands/hold.ts +26 -0
  29. package/src/commands/intake.ts +155 -0
  30. package/src/commands/ledger.ts +69 -0
  31. package/src/commands/message.ts +96 -0
  32. package/src/commands/report.ts +206 -0
  33. package/src/commands/restart.ts +76 -0
  34. package/src/commands/resume.ts +58 -0
  35. package/src/commands/setup.ts +143 -0
  36. package/src/commands/start.ts +23 -0
  37. package/src/commands/stats.ts +131 -0
  38. package/src/commands/status.ts +48 -0
  39. package/src/commands/stop.ts +51 -0
  40. package/src/commands/tail.ts +109 -0
  41. package/src/commands/unblock.ts +39 -0
  42. package/src/commands/upgrade-install.ts +31 -0
  43. package/src/commands/upgrade-rollback.ts +32 -0
  44. package/src/commands/upgrade.ts +25 -0
  45. package/src/commands/verb.ts +83 -0
  46. package/src/commands/version.ts +30 -0
  47. package/src/commands/worker.ts +100 -0
  48. package/src/config-schema.ts +47 -1
  49. package/src/config.ts +45 -4
  50. package/src/daemon.ts +972 -101
  51. package/src/dashboard/app.js +459 -0
  52. package/src/dashboard/index.html +61 -0
  53. package/src/dashboard/server.ts +481 -0
  54. package/src/dashboard/style.css +348 -0
  55. package/src/decisions.ts +39 -14
  56. package/src/diff-flags.ts +131 -241
  57. package/src/doctor.ts +932 -0
  58. package/src/escalate.ts +2 -2
  59. package/src/failure-class.ts +66 -3
  60. package/src/fleet.ts +58 -1
  61. package/src/gitops.ts +157 -0
  62. package/src/graph-health.ts +1 -1
  63. package/src/label-projection.ts +1 -1
  64. package/src/lifecycle.ts +198 -2
  65. package/src/model-fallback.ts +177 -0
  66. package/src/notices.ts +9 -0
  67. package/src/omp.ts +93 -13
  68. package/src/orchestrator-tick.ts +414 -20
  69. package/src/orchestrator.ts +4 -4
  70. package/src/privileged.ts +10 -0
  71. package/src/release-policy.ts +342 -30
  72. package/src/reports.ts +19 -5
  73. package/src/session-host.ts +11 -5
  74. package/src/setup-host.ts +663 -25
  75. package/src/setup-install.ts +292 -28
  76. package/src/setup-wizard.ts +255 -74
  77. package/src/setup.ts +156 -35
  78. package/src/stats.ts +331 -0
  79. package/src/store.ts +219 -22
  80. package/src/tracker/github.ts +74 -4
  81. package/src/types.ts +215 -31
  82. package/src/unblock.ts +55 -11
  83. package/src/upgrade-journal.ts +220 -0
  84. package/src/upgrade-verify.ts +506 -0
  85. package/src/upgrade.ts +385 -58
  86. package/src/verbs/actions.ts +73 -1
  87. package/src/verbs/protocol.ts +29 -4
  88. package/src/verbs/server.ts +183 -20
  89. package/src/worker.ts +3 -3
  90. package/systemd/omp-conductor-recover.sh +433 -0
  91. package/systemd/omp-conductor.service.example +14 -3
  92. package/systemd/recover-unit-test.sh +428 -0
@@ -0,0 +1,433 @@
1
+ #!/usr/bin/env bash
2
+ #
3
+ # omp-conductor fleet recovery — the OnFailure= playbook for the two fleet
4
+ # units (omp-conductor.service, herdr-fleet.service).
5
+ #
6
+ # Why it exists. Both units crash-loop under systemd's default start limit and
7
+ # then sit in `failed` indefinitely, in silence: the daemon cannot report its
8
+ # own crash loop, and the orchestrator cannot report a session host that is not
9
+ # hosting it (#485). Each unit carries OnFailure=omp-conductor-recover.service,
10
+ # so the moment one of them is given up on, systemd starts this unit — outside
11
+ # the failed unit's cgroup and process tree, so the restarts below can never
12
+ # take the recovery down with them.
13
+ #
14
+ # The playbook is fixed and bounded, because it runs at the moment the system
15
+ # is least healthy with no orchestrator alive to review it:
16
+ #
17
+ # 1. Collect evidence durably, before anything else. journal tail, unit
18
+ # state, `omp-conductor doctor --json`, installed versions, and the
19
+ # pre-upgrade snapshot list land in $RECOVER_DIR so the record survives
20
+ # no matter what happens next.
21
+ # 2. Attempt at most ONE known-safe restorative action, chosen mechanically
22
+ # from that evidence:
23
+ # ROLLBACK — a fresh pre-upgrade config snapshot exists (an upgrade
24
+ # just landed and the daemon died after it): restore the
25
+ # durable bytes `upgrade` keeps, then start the daemon
26
+ # unit once.
27
+ # RECOVER_PANE — herdr-fleet.service failed: start it (reset-failed +
28
+ # start); its startup hook resumes the fleet pane by
29
+ # exact identity.
30
+ # RESTART_PEER — the daemon failed with no upgrade evidence: restart
31
+ # the peer unit (herdr-fleet.service) once. The fleet
32
+ # units share state; the peer is the only unit the
33
+ # recovery may touch that cannot re-loop the failure
34
+ # that already exhausted the failed unit's start limit.
35
+ # Nothing improvised: no diagnosis, no repo edits, no installed-tree
36
+ # writes, no blind restarts of the failed unit. If the evidence points at
37
+ # a code defect, the right output is the escalation — the normal fleet
38
+ # loop repairs it once the fleet is back.
39
+ # 3. Escalate tier-2 through the durable outbox (`omp-conductor report
40
+ # --kind tier2`), whether or not the action worked — a fleet that fell
41
+ # over and got itself up is still news.
42
+ # 4. Never loop. Consecutive failed attempts are capped at
43
+ # RECOVER_MAX_FAILED_ATTEMPTS via a durable counter (expired after
44
+ # RECOVER_FAIL_WINDOW_S so a genuinely new episode gets fresh attempts),
45
+ # the recovery unit itself carries StartLimitBurst=2 (rendered by
46
+ # setup-host.ts), and this unit never carries OnFailure= — there is no
47
+ # path where recovery re-triggers recovery.
48
+ #
49
+ # Runs as root (journalctl and systemctl need it). OMP_CONDUCTOR_HOME is set
50
+ # by the unit to the fleet's conductor home, so every omp-conductor call
51
+ # touches the fleet's own config and store. RECOVER_DRY_RUN=1 prints every
52
+ # decision and mutates nothing.
53
+
54
+ set -uo pipefail
55
+
56
+ PROG=omp-conductor-recover
57
+
58
+ # Every external command is overridable so tests can inject stubs; the unit
59
+ # renders an explicit PATH (this host's non-interactive PATH lacks `omp`,
60
+ # `herdr` and `omp-conductor`).
61
+ SYSTEMCTL=${SYSTEMCTL_BIN:-systemctl}
62
+ JOURNALCTL=${JOURNALCTL_BIN:-journalctl}
63
+ CONDUCTOR=${OMP_CONDUCTOR_BIN:-omp-conductor}
64
+
65
+ # The fleet's conductor home: where config.json, the sqlite store and the
66
+ # upgrade's durable snapshots live. The unit renders OMP_CONDUCTOR_HOME and
67
+ # RECOVER_STATE_DIR; a standalone run defaults to ~/.omp/conductor.
68
+ STATE_DIR=${RECOVER_STATE_DIR:-${OMP_CONDUCTOR_HOME:-$HOME/.omp/conductor}}
69
+ CONFIG_PATH=$STATE_DIR/config.json
70
+ BACKUP_ROOT=$STATE_DIR/backups/config
71
+ RECOVER_DIR=${RECOVER_DIR:-$STATE_DIR/recovery}
72
+
73
+ # The two fleet units, daemon first. Tests may narrow this list.
74
+ DAEMON_UNIT=${DAEMON_UNIT:-omp-conductor.service}
75
+ HERDR_UNIT=${HERDR_UNIT:-herdr-fleet.service}
76
+ FLEET_UNITS=${FLEET_UNITS:-$DAEMON_UNIT:$HERDR_UNIT}
77
+
78
+ # Bounds (see the header). The rendered unit's StartLimitIntervalSec should
79
+ # stay in agreement with RECOVER_FAIL_WINDOW_S.
80
+ MAX_FAILED_ATTEMPTS=${RECOVER_MAX_FAILED_ATTEMPTS:-2}
81
+ FAIL_WINDOW_S=${RECOVER_FAIL_WINDOW_S:-1800}
82
+ ROLLBACK_AGE_S=${RECOVER_ROLLBACK_AGE_S:-3600}
83
+ JOURNAL_LINES=${RECOVER_JOURNAL_LINES:-200}
84
+ VERIFY_POLL_S=${RECOVER_VERIFY_POLL_S:-2}
85
+ VERIFY_TRIES=${RECOVER_VERIFY_TRIES:-10}
86
+ PROJECT=${RECOVER_PROJECT:-}
87
+
88
+ # --------------------------------------------------------------------------
89
+ # output
90
+ # --------------------------------------------------------------------------
91
+
92
+ say() { printf '%s: %s\n' "$PROG" "$*"; }
93
+ log() { printf '%s: %s\n' "$PROG" "$*" >&2; }
94
+
95
+ dry_run() { [[ ${RECOVER_DRY_RUN:-0} == 1 ]]; }
96
+
97
+ stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
98
+ epoch_now() { date +%s; }
99
+
100
+ # --------------------------------------------------------------------------
101
+ # systemd access (the only live-host surface; tests inject stubs)
102
+ # --------------------------------------------------------------------------
103
+
104
+ # `is-failed` prints "failed" only for the failed state; an absent unit prints
105
+ # "inactive" and exits 3, which is not a failure.
106
+ unit_is_failed() { # <unit>
107
+ [[ $("$SYSTEMCTL" is-failed "$1" 2>/dev/null | tr -d '[:space:]') == failed ]]
108
+ }
109
+
110
+ unit_says_failed() { # <unit> — "is-failed" verdict + exit code, for logs
111
+ "$SYSTEMCTL" is-failed "$1" 2>/dev/null | tr -d '[:space:]'
112
+ }
113
+
114
+ # One bounded boot of a unit: clear the start limit, then start. Only ever
115
+ # used on the failed unit with fresh evidence behind it (the rollback leg) or
116
+ # on herdr-fleet (RECOVER_PANE), never as a blind daemon restart.
117
+ start_unit() { # <unit>
118
+ say "systemd: reset-failed $1"
119
+ "$SYSTEMCTL" reset-failed "$1" 2>/dev/null || true
120
+ say "systemd: start $1"
121
+ "$SYSTEMCTL" start "$1"
122
+ }
123
+
124
+ # Wait until the unit leaves the failed state, bounded by the poll window.
125
+ # Returns 0 when it did.
126
+ wait_no_longer_failed() { # <unit>
127
+ local tries=$VERIFY_TRIES
128
+ while (( tries > 0 )); do
129
+ if ! unit_is_failed "$1"; then
130
+ return 0
131
+ fi
132
+ sleep "$VERIFY_POLL_S"
133
+ tries=$((tries - 1))
134
+ done
135
+ return 1
136
+ }
137
+
138
+ # --------------------------------------------------------------------------
139
+ # durable evidence (written before anything else)
140
+ # --------------------------------------------------------------------------
141
+
142
+ # Write the full evidence bundle for one run and print its path. The file is
143
+ # the record a later human or the ordinary fleet loop reads (#378 wants it
144
+ # provenance-carrying): who, when, what was installed, what the units said,
145
+ # what the doctor said, and the journal tail of the failed unit.
146
+ write_evidence() { # <failed-unit or empty>
147
+ local unit=${1:-none} states="" snapshots="" doctor="" versions="" journal=""
148
+ local u
149
+ if ! mkdir -p "$RECOVER_DIR" 2>/dev/null; then
150
+ log "cannot create recovery dir $RECOVER_DIR"
151
+ return 2
152
+ fi
153
+ for u in ${FLEET_UNITS//:/ }; do
154
+ states="$states $u=$(unit_says_failed "$u")"
155
+ done
156
+ snapshots=$(ls -t "$BACKUP_ROOT"/config.json.pre-upgrade-* 2>/dev/null | head -n 5 || true)
157
+ doctor=$("$CONDUCTOR" doctor --json 2>&1 | head -c 4000 || true)
158
+ versions=$("$CONDUCTOR" --version 2>&1 | head -n 1 || true)
159
+ if [[ $unit != none ]]; then
160
+ journal=$("$JOURNALCTL" -u "$unit" -n "$JOURNAL_LINES" --no-pager 2>&1 || true)
161
+ fi
162
+ local file="$RECOVER_DIR/evidence-$(stamp)-$unit.txt"
163
+ {
164
+ printf '%s\n' "=== omp-conductor fleet recovery evidence (#485) ==="
165
+ printf 'failed unit: %s\n' "$unit"
166
+ printf 'triggered at: %s (epoch %s)\n' "$(stamp)" "$(epoch_now)"
167
+ printf 'recover script: %s\n' "$(readlink -f "$0" 2>/dev/null || printf '%s' "$0")"
168
+ printf 'units: %s\n' "$states"
169
+ printf 'installed: %s\n' "$versions"
170
+ printf '%s\n' 'pre-upgrade snapshots (newest first):'
171
+ printf '%s\n' "${snapshots:- (none)}"
172
+ printf '%s\n' '--- doctor ---'
173
+ printf '%s\n' "$doctor"
174
+ printf '%s\n' '--- journal tail ---'
175
+ printf '%s\n' "$journal"
176
+ printf '%s\n' '--- decision ---'
177
+ } >"$file" 2>/dev/null
178
+ if [[ ! -s $file ]]; then
179
+ log "could not write evidence to $file"
180
+ return 2
181
+ fi
182
+ printf '%s\n' "$file"
183
+ }
184
+
185
+ append_evidence() { # <evidence path> <text>
186
+ printf '%s\n' "$2" >>"$1" 2>/dev/null || true
187
+ }
188
+
189
+ # --------------------------------------------------------------------------
190
+ # escalation — always, through the durable outbox (#123, #288)
191
+ # --------------------------------------------------------------------------
192
+
193
+ # The single escalation path this repo has for "something must reach a human":
194
+ # `omp-conductor report --kind tier2` persists the row in the fleet's store
195
+ # (the durable outbox) before anything is sent, and the daemon owns delivery
196
+ # with bounded retries once it is back. This script only ever enqueues.
197
+ escalation_text() { # <failed-unit or empty> <action> <outcome> <evidence>
198
+ if [[ ${1:-none} == none ]]; then
199
+ printf '%s\n' \
200
+ "fleet recovery (#485): a recovery trigger fired but no fleet unit was failing when it ran — the fleet got itself up or was started externally; evidence=$4"
201
+ return
202
+ fi
203
+ printf '%s\n' \
204
+ "fleet recovery (#485): $1 entered failed state; recovery=${2:-none}; outcome=${3}; evidence=${4}; next: read the evidence file (and its doctor output) before touching the fleet"
205
+ }
206
+
207
+ escalate() { # <failed-unit or empty> <action-name> <outcome> <evidence-path>
208
+ local unit=${1:-none} text args
209
+ text=$(escalation_text "$unit" "$2" "$3" "$4")
210
+ if [[ -n $PROJECT ]]; then
211
+ args=("$CONDUCTOR" report --kind tier2 --project "$PROJECT" --text "$text")
212
+ else
213
+ args=("$CONDUCTOR" report --kind tier2 --text "$text")
214
+ fi
215
+ say "escalation: ${args[*]}"
216
+ if dry_run; then
217
+ return 0
218
+ fi
219
+ local out
220
+ out=$("${args[@]}" 2>&1)
221
+ local code=$?
222
+ if (( code != 0 )); then
223
+ log "escalation failed (exit $code): $out"
224
+ return 1
225
+ fi
226
+ printf '%s: escalation: %s\n' "$PROG" "$out" >&2
227
+ return 0
228
+ }
229
+
230
+ # --------------------------------------------------------------------------
231
+ # bounds — the durable counter ("a recovery that fails twice stops trying")
232
+ # --------------------------------------------------------------------------
233
+
234
+ MARKER="$RECOVER_DIR/attempts"
235
+
236
+ read_failed_attempts() {
237
+ local count ts now
238
+ [[ -f $MARKER ]] || { printf '0\n'; return; }
239
+ count=$(sed -n '1p' "$MARKER" 2>/dev/null || true)
240
+ ts=$(sed -n '2p' "$MARKER" 2>/dev/null || true)
241
+ now=$(epoch_now)
242
+ if [[ ! $count =~ ^[0-9]+$ || ! $ts =~ ^[0-9]+$ ]] || (( now - ts > FAIL_WINDOW_S )); then
243
+ printf '0\n'
244
+ return
245
+ fi
246
+ printf '%s\n' "$count"
247
+ }
248
+
249
+ bump_failed_attempts() {
250
+ local next
251
+ next=$(read_failed_attempts)
252
+ mkdir -p "$RECOVER_DIR" 2>/dev/null || true
253
+ printf '%s\n%s\n' "$((next + 1))" "$(epoch_now)" >"$MARKER"
254
+ }
255
+
256
+ clear_failed_attempts() {
257
+ rm -f "$MARKER"
258
+ }
259
+
260
+ # --------------------------------------------------------------------------
261
+ # decisions (pure; the test suite drives these directly)
262
+ # --------------------------------------------------------------------------
263
+
264
+ # The newest pre-upgrade config snapshot recent enough to implicate the current
265
+ # failure. `upgrade` persists these durably before it touches anything
266
+ # (configBackupDir(), config.json.pre-upgrade-*). A fresh one plus a dead
267
+ # daemon is the one mechanical "the upgrade did it" signal we are allowed to
268
+ # act on — fixing the new version's mistakes is a normal daemon-side repair,
269
+ # but reinstalling the previous package pinned to a durable version is not
270
+ # something this bounded playbook may improvise.
271
+ fresh_rollback_snapshot() {
272
+ local newest ts now
273
+ newest=$(ls -t "$BACKUP_ROOT"/config.json.pre-upgrade-* 2>/dev/null | head -n 1 || true)
274
+ [[ -n $newest ]] || return 1
275
+ now=$(epoch_now)
276
+ ts=$(stat -c %Y "$newest" 2>/dev/null || true)
277
+ [[ -n $ts && $((now - ts)) -le $ROLLBACK_AGE_S ]] || return 1
278
+ printf '%s\n' "$newest"
279
+ }
280
+
281
+ # The single action for one failed unit: `ROLLBACK|<snapshot>`,
282
+ # `RESTART_PEER|<herdr unit>`, `RECOVER_PANE|<herdr unit>` or `NONE|`.
283
+ # Mechanically selected, never improvised.
284
+ choose_action() { # <failed-unit>
285
+ local unit=$1 snap
286
+ if [[ $unit == "$DAEMON_UNIT" ]]; then
287
+ if snap=$(fresh_rollback_snapshot); then
288
+ printf 'ROLLBACK|%s\n' "$snap"
289
+ return
290
+ fi
291
+ printf 'RESTART_PEER|%s\n' "$HERDR_UNIT"
292
+ return
293
+ fi
294
+ if [[ $unit == "$HERDR_UNIT" ]]; then
295
+ printf 'RECOVER_PANE|%s\n' "$HERDR_UNIT"
296
+ return
297
+ fi
298
+ printf 'NONE|\n'
299
+ }
300
+
301
+ # --------------------------------------------------------------------------
302
+ # main
303
+ # --------------------------------------------------------------------------
304
+
305
+ run_recovery() {
306
+ local failed="" u action_name="" action_arg="" ev="" outcome="" ok=1
307
+ for u in ${FLEET_UNITS//:/ }; do
308
+ if unit_is_failed "$u"; then
309
+ failed=$u
310
+ break
311
+ fi
312
+ done
313
+
314
+ if dry_run; then
315
+ say "dry run: nothing will be written or executed"
316
+ [[ -z $failed ]] && say "plan: no unit failed — escalate only"
317
+ if [[ -n $failed ]]; then
318
+ IFS='|' read -r action_name action_arg <<<"$(choose_action "$failed")"
319
+ say "plan: failed unit $failed; action: $action_name"
320
+ fi
321
+ say "plan: escalation — $(escalation_text "$failed" "${action_name:-none}" "dry run" "DRY-RUN")"
322
+ return 0
323
+ fi
324
+
325
+ ev=$(write_evidence "$failed")
326
+ if [[ -z $ev ]]; then
327
+ log "could not write evidence under $RECOVER_DIR — treating the recovery as failed"
328
+ return 1
329
+ fi
330
+
331
+ local attempts
332
+ if [[ -z $failed ]]; then
333
+ # Nothing failed when recovery ran (the fleet got itself up, or an
334
+ # external start beat this activation): always escalate, and a previously
335
+ # exhausted episode counter gets a fresh start.
336
+ say "no unit failed when this recovery ran — the fleet got itself up (or was started externally); nothing to attempt"
337
+ append_evidence "$ev" "decision: NONE — no unit failed"
338
+ if escalate none none "no unit failed when the recovery ran" "$ev"; then
339
+ clear_failed_attempts
340
+ return 0
341
+ fi
342
+ return 1
343
+ fi
344
+
345
+ attempts=$(read_failed_attempts)
346
+ if [[ $attempts -ge $MAX_FAILED_ATTEMPTS ]]; then
347
+ say "attempts exhausted: $attempts consecutive failed recoveries inside the ${FAIL_WINDOW_S}s window; escalating only"
348
+ append_evidence "$ev" "decision: NONE (attempts exhausted, $attempts)"
349
+ if escalate "$failed" none "attempts exhausted ($attempts); no further recovery" "$ev"; then
350
+ return 0
351
+ fi
352
+ return 1
353
+ fi
354
+
355
+ say "failed unit: $failed"
356
+ IFS='|' read -r action_name action_arg <<<"$(choose_action "$failed")"
357
+ action_name=${action_name:-NONE}
358
+ append_evidence "$ev" "decision: $action_name"
359
+ if [[ $action_name == NONE ]]; then
360
+ say "no safe mechanical action for $failed; escalating"
361
+ if escalate "$failed" none "no known-safe action" "$ev"; then
362
+ return 0
363
+ fi
364
+ return 1
365
+ fi
366
+
367
+ # Exactly one action, no matter how many units are down.
368
+ case "$action_name" in
369
+ ROLLBACK)
370
+ if cp "$action_arg" "$CONFIG_PATH" && start_unit "$failed"; then
371
+ if wait_no_longer_failed "$failed"; then
372
+ ok=0
373
+ outcome="restored pre-upgrade snapshot ${action_arg##*/} and started $failed"
374
+ else
375
+ outcome="restored pre-upgrade snapshot and started $failed, but it is still failing"
376
+ fi
377
+ else
378
+ outcome="could not restore snapshot ${action_arg##*/} or start $failed"
379
+ fi
380
+ ;;
381
+ RECOVER_PANE)
382
+ if start_unit "$HERDR_UNIT"; then
383
+ if wait_no_longer_failed "$HERDR_UNIT"; then
384
+ ok=0
385
+ outcome="restarted session host $HERDR_UNIT; its startup hook recovers the fleet pane"
386
+ else
387
+ outcome="$HERDR_UNIT started but is still failing"
388
+ fi
389
+ else
390
+ outcome="could not start $HERDR_UNIT"
391
+ fi
392
+ ;;
393
+ RESTART_PEER)
394
+ if "$SYSTEMCTL" restart "$action_arg"; then
395
+ outcome="restarted peer unit $action_arg"
396
+ else
397
+ outcome="could not restart peer unit $action_arg"
398
+ fi
399
+ ;;
400
+ esac
401
+ say "outcome: $outcome"
402
+ append_evidence "$ev" "outcome: $outcome"
403
+
404
+ if (( ok == 0 )); then
405
+ append_evidence "$ev" "verify: $failed no longer failed"
406
+ say "recovery succeeded: $failed no longer failed"
407
+ clear_failed_attempts
408
+ else
409
+ append_evidence "$ev" "verify: $failed still failed"
410
+ say "recovery failed: $failed still failed; attempts remaining until exhausted"
411
+ bump_failed_attempts
412
+ fi
413
+
414
+ if escalate "$failed" "$action_name" "$outcome" "$ev"; then
415
+ if (( ok == 0 )); then
416
+ return 0
417
+ fi
418
+ return 1
419
+ fi
420
+ return 1
421
+ }
422
+
423
+ main() {
424
+ dry_run && say "dry run: printing decisions, mutating nothing"
425
+ if run_recovery; then
426
+ exit 0
427
+ fi
428
+ exit 1
429
+ }
430
+
431
+ if [[ ${BASH_SOURCE[0]} == "$0" ]]; then
432
+ main
433
+ fi
@@ -16,6 +16,12 @@
16
16
  # 4. Do not co-locate ClickHouse + other multi-GB services beside a 2-worker
17
17
  # fleet on a ≤8 GB box.
18
18
  #
19
+ # OnFailure (issue #485): this line arms the fleet recovery oneshot that #485
20
+ # ships via `setup host` — `omp-conductor-recover.service`, staged from
21
+ # omp-conductor-recover.sh. Do not remove it: a crash-looped unit used to sit in
22
+ # `failed` indefinitely with nobody watching it. Re-running `setup host`
23
+ # provisions the recovery unit and installs the playbook itself.
24
+ #
19
25
  # Install:
20
26
  # sudo install -m 0644 omp-conductor.service.example /etc/systemd/system/omp-conductor.service
21
27
  # sudo systemctl daemon-reload
@@ -27,6 +33,7 @@ Description=omp-conductor dispatch daemon
27
33
  Documentation=https://github.com/TerrifiedBug/conductor
28
34
  After=network-online.target
29
35
  Wants=network-online.target
36
+ OnFailure=omp-conductor-recover.service
30
37
 
31
38
  [Service]
32
39
  Type=simple
@@ -44,9 +51,13 @@ WorkingDirectory=/home/fleet
44
51
  # under a unit, call `daemon` directly.
45
52
  ExecStart=/home/fleet/.local/bin/omp-conductor daemon --port 8787
46
53
 
47
- Restart=on-failure
48
- # Handled SIGTERM exits 143; without this, Restart=on-failure brings the unit
49
- # straight back after `systemctl stop` / `omp-conductor stop`.
54
+ Restart=always
55
+ # Restart on any exit — clean, signalled or crashed — except an explicit
56
+ # `systemctl stop`, which systemd records as intentional and never undoes
57
+ # (issue #546). A stray SIGTERM (e.g. a worker's `bun test`) then costs
58
+ # seconds of downtime instead of leaving the fleet down until a human;
59
+ # a crash loop still trips the start-limit burst and reaches OnFailure above.
60
+ # SuccessExitStatus=0 143 keeps a handled SIGTERM shutdown logged as success.
50
61
  SuccessExitStatus=0 143
51
62
 
52
63
  MemoryAccounting=yes