omp-conductor 0.15.10 → 0.15.12
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +273 -2544
- package/REFERENCE.md +2680 -0
- package/package.json +3 -2
- package/schema/config.schema.json +11 -23
- package/src/arm-challenge.ts +112 -0
- package/src/ask.ts +434 -0
- package/src/board.ts +81 -15
- package/src/brief-upgrade.ts +114 -8
- package/src/briefs/orchestrator.md +113 -34
- package/src/briefs/policy.md +33 -8
- package/src/briefs/worker.md +18 -9
- package/src/chain-check.ts +1 -1
- package/src/check-trailing-newlines.ts +82 -0
- package/src/cli.ts +225 -1406
- package/src/commands/arm.ts +21 -0
- package/src/commands/board.ts +23 -0
- package/src/commands/brief-upgrade.ts +186 -0
- package/src/commands/context.ts +150 -0
- package/src/commands/daemon.ts +71 -0
- package/src/commands/dashboard.ts +74 -0
- package/src/commands/decision.ts +103 -0
- package/src/commands/disarm.ts +21 -0
- package/src/commands/doctor.ts +100 -0
- package/src/commands/event.ts +62 -0
- package/src/commands/extend.ts +64 -0
- package/src/commands/friction.ts +56 -0
- package/src/commands/help.ts +9 -0
- package/src/commands/hold.ts +26 -0
- package/src/commands/intake.ts +155 -0
- package/src/commands/ledger.ts +69 -0
- package/src/commands/message.ts +96 -0
- package/src/commands/report.ts +206 -0
- package/src/commands/restart.ts +76 -0
- package/src/commands/resume.ts +58 -0
- package/src/commands/setup.ts +143 -0
- package/src/commands/start.ts +23 -0
- package/src/commands/stats.ts +131 -0
- package/src/commands/status.ts +48 -0
- package/src/commands/stop.ts +51 -0
- package/src/commands/tail.ts +109 -0
- package/src/commands/unblock.ts +39 -0
- package/src/commands/upgrade-install.ts +31 -0
- package/src/commands/upgrade-rollback.ts +32 -0
- package/src/commands/upgrade.ts +25 -0
- package/src/commands/verb.ts +83 -0
- package/src/commands/version.ts +30 -0
- package/src/commands/worker.ts +100 -0
- package/src/config-schema.ts +47 -1
- package/src/config.ts +45 -4
- package/src/daemon.ts +972 -101
- package/src/dashboard/app.js +459 -0
- package/src/dashboard/index.html +61 -0
- package/src/dashboard/server.ts +481 -0
- package/src/dashboard/style.css +348 -0
- package/src/decisions.ts +39 -14
- package/src/diff-flags.ts +131 -241
- package/src/doctor.ts +932 -0
- package/src/escalate.ts +2 -2
- package/src/failure-class.ts +66 -3
- package/src/fleet.ts +58 -1
- package/src/gitops.ts +157 -0
- package/src/graph-health.ts +1 -1
- package/src/label-projection.ts +1 -1
- package/src/lifecycle.ts +198 -2
- package/src/model-fallback.ts +177 -0
- package/src/notices.ts +9 -0
- package/src/omp.ts +93 -13
- package/src/orchestrator-tick.ts +414 -20
- package/src/orchestrator.ts +4 -4
- package/src/privileged.ts +10 -0
- package/src/release-policy.ts +342 -30
- package/src/reports.ts +19 -5
- package/src/session-host.ts +11 -5
- package/src/setup-host.ts +663 -25
- package/src/setup-install.ts +292 -28
- package/src/setup-wizard.ts +255 -74
- package/src/setup.ts +156 -35
- package/src/stats.ts +331 -0
- package/src/store.ts +219 -22
- package/src/tracker/github.ts +74 -4
- package/src/types.ts +215 -31
- package/src/unblock.ts +55 -11
- package/src/upgrade-journal.ts +220 -0
- package/src/upgrade-verify.ts +506 -0
- package/src/upgrade.ts +385 -58
- package/src/verbs/actions.ts +73 -1
- package/src/verbs/protocol.ts +29 -4
- package/src/verbs/server.ts +183 -20
- package/src/worker.ts +3 -3
- package/systemd/omp-conductor-recover.sh +433 -0
- package/systemd/omp-conductor.service.example +14 -3
- package/systemd/recover-unit-test.sh +428 -0
|
@@ -0,0 +1,433 @@
|
|
|
1
|
+
#!/usr/bin/env bash
|
|
2
|
+
#
|
|
3
|
+
# omp-conductor fleet recovery — the OnFailure= playbook for the two fleet
|
|
4
|
+
# units (omp-conductor.service, herdr-fleet.service).
|
|
5
|
+
#
|
|
6
|
+
# Why it exists. Both units crash-loop under systemd's default start limit and
|
|
7
|
+
# then sit in `failed` indefinitely, in silence: the daemon cannot report its
|
|
8
|
+
# own crash loop, and the orchestrator cannot report a session host that is not
|
|
9
|
+
# hosting it (#485). Each unit carries OnFailure=omp-conductor-recover.service,
|
|
10
|
+
# so the moment one of them is given up on, systemd starts this unit — outside
|
|
11
|
+
# the failed unit's cgroup and process tree, so the restarts below can never
|
|
12
|
+
# take the recovery down with them.
|
|
13
|
+
#
|
|
14
|
+
# The playbook is fixed and bounded, because it runs at the moment the system
|
|
15
|
+
# is least healthy with no orchestrator alive to review it:
|
|
16
|
+
#
|
|
17
|
+
# 1. Collect evidence durably, before anything else. journal tail, unit
|
|
18
|
+
# state, `omp-conductor doctor --json`, installed versions, and the
|
|
19
|
+
# pre-upgrade snapshot list land in $RECOVER_DIR so the record survives
|
|
20
|
+
# no matter what happens next.
|
|
21
|
+
# 2. Attempt at most ONE known-safe restorative action, chosen mechanically
|
|
22
|
+
# from that evidence:
|
|
23
|
+
# ROLLBACK — a fresh pre-upgrade config snapshot exists (an upgrade
|
|
24
|
+
# just landed and the daemon died after it): restore the
|
|
25
|
+
# durable bytes `upgrade` keeps, then start the daemon
|
|
26
|
+
# unit once.
|
|
27
|
+
# RECOVER_PANE — herdr-fleet.service failed: start it (reset-failed +
|
|
28
|
+
# start); its startup hook resumes the fleet pane by
|
|
29
|
+
# exact identity.
|
|
30
|
+
# RESTART_PEER — the daemon failed with no upgrade evidence: restart
|
|
31
|
+
# the peer unit (herdr-fleet.service) once. The fleet
|
|
32
|
+
# units share state; the peer is the only unit the
|
|
33
|
+
# recovery may touch that cannot re-loop the failure
|
|
34
|
+
# that already exhausted the failed unit's start limit.
|
|
35
|
+
# Nothing improvised: no diagnosis, no repo edits, no installed-tree
|
|
36
|
+
# writes, no blind restarts of the failed unit. If the evidence points at
|
|
37
|
+
# a code defect, the right output is the escalation — the normal fleet
|
|
38
|
+
# loop repairs it once the fleet is back.
|
|
39
|
+
# 3. Escalate tier-2 through the durable outbox (`omp-conductor report
|
|
40
|
+
# --kind tier2`), whether or not the action worked — a fleet that fell
|
|
41
|
+
# over and got itself up is still news.
|
|
42
|
+
# 4. Never loop. Consecutive failed attempts are capped at
|
|
43
|
+
# RECOVER_MAX_FAILED_ATTEMPTS via a durable counter (expired after
|
|
44
|
+
# RECOVER_FAIL_WINDOW_S so a genuinely new episode gets fresh attempts),
|
|
45
|
+
# the recovery unit itself carries StartLimitBurst=2 (rendered by
|
|
46
|
+
# setup-host.ts), and this unit never carries OnFailure= — there is no
|
|
47
|
+
# path where recovery re-triggers recovery.
|
|
48
|
+
#
|
|
49
|
+
# Runs as root (journalctl and systemctl need it). OMP_CONDUCTOR_HOME is set
|
|
50
|
+
# by the unit to the fleet's conductor home, so every omp-conductor call
|
|
51
|
+
# touches the fleet's own config and store. RECOVER_DRY_RUN=1 prints every
|
|
52
|
+
# decision and mutates nothing.
|
|
53
|
+
|
|
54
|
+
set -uo pipefail
|
|
55
|
+
|
|
56
|
+
PROG=omp-conductor-recover
|
|
57
|
+
|
|
58
|
+
# Every external command is overridable so tests can inject stubs; the unit
|
|
59
|
+
# renders an explicit PATH (this host's non-interactive PATH lacks `omp`,
|
|
60
|
+
# `herdr` and `omp-conductor`).
|
|
61
|
+
SYSTEMCTL=${SYSTEMCTL_BIN:-systemctl}
|
|
62
|
+
JOURNALCTL=${JOURNALCTL_BIN:-journalctl}
|
|
63
|
+
CONDUCTOR=${OMP_CONDUCTOR_BIN:-omp-conductor}
|
|
64
|
+
|
|
65
|
+
# The fleet's conductor home: where config.json, the sqlite store and the
|
|
66
|
+
# upgrade's durable snapshots live. The unit renders OMP_CONDUCTOR_HOME and
|
|
67
|
+
# RECOVER_STATE_DIR; a standalone run defaults to ~/.omp/conductor.
|
|
68
|
+
STATE_DIR=${RECOVER_STATE_DIR:-${OMP_CONDUCTOR_HOME:-$HOME/.omp/conductor}}
|
|
69
|
+
CONFIG_PATH=$STATE_DIR/config.json
|
|
70
|
+
BACKUP_ROOT=$STATE_DIR/backups/config
|
|
71
|
+
RECOVER_DIR=${RECOVER_DIR:-$STATE_DIR/recovery}
|
|
72
|
+
|
|
73
|
+
# The two fleet units, daemon first. Tests may narrow this list.
|
|
74
|
+
DAEMON_UNIT=${DAEMON_UNIT:-omp-conductor.service}
|
|
75
|
+
HERDR_UNIT=${HERDR_UNIT:-herdr-fleet.service}
|
|
76
|
+
FLEET_UNITS=${FLEET_UNITS:-$DAEMON_UNIT:$HERDR_UNIT}
|
|
77
|
+
|
|
78
|
+
# Bounds (see the header). The rendered unit's StartLimitIntervalSec should
|
|
79
|
+
# stay in agreement with RECOVER_FAIL_WINDOW_S.
|
|
80
|
+
MAX_FAILED_ATTEMPTS=${RECOVER_MAX_FAILED_ATTEMPTS:-2}
|
|
81
|
+
FAIL_WINDOW_S=${RECOVER_FAIL_WINDOW_S:-1800}
|
|
82
|
+
ROLLBACK_AGE_S=${RECOVER_ROLLBACK_AGE_S:-3600}
|
|
83
|
+
JOURNAL_LINES=${RECOVER_JOURNAL_LINES:-200}
|
|
84
|
+
VERIFY_POLL_S=${RECOVER_VERIFY_POLL_S:-2}
|
|
85
|
+
VERIFY_TRIES=${RECOVER_VERIFY_TRIES:-10}
|
|
86
|
+
PROJECT=${RECOVER_PROJECT:-}
|
|
87
|
+
|
|
88
|
+
# --------------------------------------------------------------------------
|
|
89
|
+
# output
|
|
90
|
+
# --------------------------------------------------------------------------
|
|
91
|
+
|
|
92
|
+
say() { printf '%s: %s\n' "$PROG" "$*"; }
|
|
93
|
+
log() { printf '%s: %s\n' "$PROG" "$*" >&2; }
|
|
94
|
+
|
|
95
|
+
dry_run() { [[ ${RECOVER_DRY_RUN:-0} == 1 ]]; }
|
|
96
|
+
|
|
97
|
+
stamp() { date -u +%Y-%m-%dT%H:%M:%SZ; }
|
|
98
|
+
epoch_now() { date +%s; }
|
|
99
|
+
|
|
100
|
+
# --------------------------------------------------------------------------
|
|
101
|
+
# systemd access (the only live-host surface; tests inject stubs)
|
|
102
|
+
# --------------------------------------------------------------------------
|
|
103
|
+
|
|
104
|
+
# `is-failed` prints "failed" only for the failed state; an absent unit prints
|
|
105
|
+
# "inactive" and exits 3, which is not a failure.
|
|
106
|
+
unit_is_failed() { # <unit>
|
|
107
|
+
[[ $("$SYSTEMCTL" is-failed "$1" 2>/dev/null | tr -d '[:space:]') == failed ]]
|
|
108
|
+
}
|
|
109
|
+
|
|
110
|
+
unit_says_failed() { # <unit> — "is-failed" verdict + exit code, for logs
|
|
111
|
+
"$SYSTEMCTL" is-failed "$1" 2>/dev/null | tr -d '[:space:]'
|
|
112
|
+
}
|
|
113
|
+
|
|
114
|
+
# One bounded boot of a unit: clear the start limit, then start. Only ever
|
|
115
|
+
# used on the failed unit with fresh evidence behind it (the rollback leg) or
|
|
116
|
+
# on herdr-fleet (RECOVER_PANE), never as a blind daemon restart.
|
|
117
|
+
start_unit() { # <unit>
|
|
118
|
+
say "systemd: reset-failed $1"
|
|
119
|
+
"$SYSTEMCTL" reset-failed "$1" 2>/dev/null || true
|
|
120
|
+
say "systemd: start $1"
|
|
121
|
+
"$SYSTEMCTL" start "$1"
|
|
122
|
+
}
|
|
123
|
+
|
|
124
|
+
# Wait until the unit leaves the failed state, bounded by the poll window.
|
|
125
|
+
# Returns 0 when it did.
|
|
126
|
+
wait_no_longer_failed() { # <unit>
|
|
127
|
+
local tries=$VERIFY_TRIES
|
|
128
|
+
while (( tries > 0 )); do
|
|
129
|
+
if ! unit_is_failed "$1"; then
|
|
130
|
+
return 0
|
|
131
|
+
fi
|
|
132
|
+
sleep "$VERIFY_POLL_S"
|
|
133
|
+
tries=$((tries - 1))
|
|
134
|
+
done
|
|
135
|
+
return 1
|
|
136
|
+
}
|
|
137
|
+
|
|
138
|
+
# --------------------------------------------------------------------------
|
|
139
|
+
# durable evidence (written before anything else)
|
|
140
|
+
# --------------------------------------------------------------------------
|
|
141
|
+
|
|
142
|
+
# Write the full evidence bundle for one run and print its path. The file is
|
|
143
|
+
# the record a later human or the ordinary fleet loop reads (#378 wants it
|
|
144
|
+
# provenance-carrying): who, when, what was installed, what the units said,
|
|
145
|
+
# what the doctor said, and the journal tail of the failed unit.
|
|
146
|
+
write_evidence() { # <failed-unit or empty>
|
|
147
|
+
local unit=${1:-none} states="" snapshots="" doctor="" versions="" journal=""
|
|
148
|
+
local u
|
|
149
|
+
if ! mkdir -p "$RECOVER_DIR" 2>/dev/null; then
|
|
150
|
+
log "cannot create recovery dir $RECOVER_DIR"
|
|
151
|
+
return 2
|
|
152
|
+
fi
|
|
153
|
+
for u in ${FLEET_UNITS//:/ }; do
|
|
154
|
+
states="$states $u=$(unit_says_failed "$u")"
|
|
155
|
+
done
|
|
156
|
+
snapshots=$(ls -t "$BACKUP_ROOT"/config.json.pre-upgrade-* 2>/dev/null | head -n 5 || true)
|
|
157
|
+
doctor=$("$CONDUCTOR" doctor --json 2>&1 | head -c 4000 || true)
|
|
158
|
+
versions=$("$CONDUCTOR" --version 2>&1 | head -n 1 || true)
|
|
159
|
+
if [[ $unit != none ]]; then
|
|
160
|
+
journal=$("$JOURNALCTL" -u "$unit" -n "$JOURNAL_LINES" --no-pager 2>&1 || true)
|
|
161
|
+
fi
|
|
162
|
+
local file="$RECOVER_DIR/evidence-$(stamp)-$unit.txt"
|
|
163
|
+
{
|
|
164
|
+
printf '%s\n' "=== omp-conductor fleet recovery evidence (#485) ==="
|
|
165
|
+
printf 'failed unit: %s\n' "$unit"
|
|
166
|
+
printf 'triggered at: %s (epoch %s)\n' "$(stamp)" "$(epoch_now)"
|
|
167
|
+
printf 'recover script: %s\n' "$(readlink -f "$0" 2>/dev/null || printf '%s' "$0")"
|
|
168
|
+
printf 'units: %s\n' "$states"
|
|
169
|
+
printf 'installed: %s\n' "$versions"
|
|
170
|
+
printf '%s\n' 'pre-upgrade snapshots (newest first):'
|
|
171
|
+
printf '%s\n' "${snapshots:- (none)}"
|
|
172
|
+
printf '%s\n' '--- doctor ---'
|
|
173
|
+
printf '%s\n' "$doctor"
|
|
174
|
+
printf '%s\n' '--- journal tail ---'
|
|
175
|
+
printf '%s\n' "$journal"
|
|
176
|
+
printf '%s\n' '--- decision ---'
|
|
177
|
+
} >"$file" 2>/dev/null
|
|
178
|
+
if [[ ! -s $file ]]; then
|
|
179
|
+
log "could not write evidence to $file"
|
|
180
|
+
return 2
|
|
181
|
+
fi
|
|
182
|
+
printf '%s\n' "$file"
|
|
183
|
+
}
|
|
184
|
+
|
|
185
|
+
append_evidence() { # <evidence path> <text>
|
|
186
|
+
printf '%s\n' "$2" >>"$1" 2>/dev/null || true
|
|
187
|
+
}
|
|
188
|
+
|
|
189
|
+
# --------------------------------------------------------------------------
|
|
190
|
+
# escalation — always, through the durable outbox (#123, #288)
|
|
191
|
+
# --------------------------------------------------------------------------
|
|
192
|
+
|
|
193
|
+
# The single escalation path this repo has for "something must reach a human":
|
|
194
|
+
# `omp-conductor report --kind tier2` persists the row in the fleet's store
|
|
195
|
+
# (the durable outbox) before anything is sent, and the daemon owns delivery
|
|
196
|
+
# with bounded retries once it is back. This script only ever enqueues.
|
|
197
|
+
escalation_text() { # <failed-unit or empty> <action> <outcome> <evidence>
|
|
198
|
+
if [[ ${1:-none} == none ]]; then
|
|
199
|
+
printf '%s\n' \
|
|
200
|
+
"fleet recovery (#485): a recovery trigger fired but no fleet unit was failing when it ran — the fleet got itself up or was started externally; evidence=$4"
|
|
201
|
+
return
|
|
202
|
+
fi
|
|
203
|
+
printf '%s\n' \
|
|
204
|
+
"fleet recovery (#485): $1 entered failed state; recovery=${2:-none}; outcome=${3}; evidence=${4}; next: read the evidence file (and its doctor output) before touching the fleet"
|
|
205
|
+
}
|
|
206
|
+
|
|
207
|
+
escalate() { # <failed-unit or empty> <action-name> <outcome> <evidence-path>
|
|
208
|
+
local unit=${1:-none} text args
|
|
209
|
+
text=$(escalation_text "$unit" "$2" "$3" "$4")
|
|
210
|
+
if [[ -n $PROJECT ]]; then
|
|
211
|
+
args=("$CONDUCTOR" report --kind tier2 --project "$PROJECT" --text "$text")
|
|
212
|
+
else
|
|
213
|
+
args=("$CONDUCTOR" report --kind tier2 --text "$text")
|
|
214
|
+
fi
|
|
215
|
+
say "escalation: ${args[*]}"
|
|
216
|
+
if dry_run; then
|
|
217
|
+
return 0
|
|
218
|
+
fi
|
|
219
|
+
local out
|
|
220
|
+
out=$("${args[@]}" 2>&1)
|
|
221
|
+
local code=$?
|
|
222
|
+
if (( code != 0 )); then
|
|
223
|
+
log "escalation failed (exit $code): $out"
|
|
224
|
+
return 1
|
|
225
|
+
fi
|
|
226
|
+
printf '%s: escalation: %s\n' "$PROG" "$out" >&2
|
|
227
|
+
return 0
|
|
228
|
+
}
|
|
229
|
+
|
|
230
|
+
# --------------------------------------------------------------------------
|
|
231
|
+
# bounds — the durable counter ("a recovery that fails twice stops trying")
|
|
232
|
+
# --------------------------------------------------------------------------
|
|
233
|
+
|
|
234
|
+
MARKER="$RECOVER_DIR/attempts"
|
|
235
|
+
|
|
236
|
+
read_failed_attempts() {
|
|
237
|
+
local count ts now
|
|
238
|
+
[[ -f $MARKER ]] || { printf '0\n'; return; }
|
|
239
|
+
count=$(sed -n '1p' "$MARKER" 2>/dev/null || true)
|
|
240
|
+
ts=$(sed -n '2p' "$MARKER" 2>/dev/null || true)
|
|
241
|
+
now=$(epoch_now)
|
|
242
|
+
if [[ ! $count =~ ^[0-9]+$ || ! $ts =~ ^[0-9]+$ ]] || (( now - ts > FAIL_WINDOW_S )); then
|
|
243
|
+
printf '0\n'
|
|
244
|
+
return
|
|
245
|
+
fi
|
|
246
|
+
printf '%s\n' "$count"
|
|
247
|
+
}
|
|
248
|
+
|
|
249
|
+
bump_failed_attempts() {
|
|
250
|
+
local next
|
|
251
|
+
next=$(read_failed_attempts)
|
|
252
|
+
mkdir -p "$RECOVER_DIR" 2>/dev/null || true
|
|
253
|
+
printf '%s\n%s\n' "$((next + 1))" "$(epoch_now)" >"$MARKER"
|
|
254
|
+
}
|
|
255
|
+
|
|
256
|
+
clear_failed_attempts() {
|
|
257
|
+
rm -f "$MARKER"
|
|
258
|
+
}
|
|
259
|
+
|
|
260
|
+
# --------------------------------------------------------------------------
|
|
261
|
+
# decisions (pure; the test suite drives these directly)
|
|
262
|
+
# --------------------------------------------------------------------------
|
|
263
|
+
|
|
264
|
+
# The newest pre-upgrade config snapshot recent enough to implicate the current
|
|
265
|
+
# failure. `upgrade` persists these durably before it touches anything
|
|
266
|
+
# (configBackupDir(), config.json.pre-upgrade-*). A fresh one plus a dead
|
|
267
|
+
# daemon is the one mechanical "the upgrade did it" signal we are allowed to
|
|
268
|
+
# act on — fixing the new version's mistakes is a normal daemon-side repair,
|
|
269
|
+
# but reinstalling the previous package pinned to a durable version is not
|
|
270
|
+
# something this bounded playbook may improvise.
|
|
271
|
+
fresh_rollback_snapshot() {
|
|
272
|
+
local newest ts now
|
|
273
|
+
newest=$(ls -t "$BACKUP_ROOT"/config.json.pre-upgrade-* 2>/dev/null | head -n 1 || true)
|
|
274
|
+
[[ -n $newest ]] || return 1
|
|
275
|
+
now=$(epoch_now)
|
|
276
|
+
ts=$(stat -c %Y "$newest" 2>/dev/null || true)
|
|
277
|
+
[[ -n $ts && $((now - ts)) -le $ROLLBACK_AGE_S ]] || return 1
|
|
278
|
+
printf '%s\n' "$newest"
|
|
279
|
+
}
|
|
280
|
+
|
|
281
|
+
# The single action for one failed unit: `ROLLBACK|<snapshot>`,
|
|
282
|
+
# `RESTART_PEER|<herdr unit>`, `RECOVER_PANE|<herdr unit>` or `NONE|`.
|
|
283
|
+
# Mechanically selected, never improvised.
|
|
284
|
+
choose_action() { # <failed-unit>
|
|
285
|
+
local unit=$1 snap
|
|
286
|
+
if [[ $unit == "$DAEMON_UNIT" ]]; then
|
|
287
|
+
if snap=$(fresh_rollback_snapshot); then
|
|
288
|
+
printf 'ROLLBACK|%s\n' "$snap"
|
|
289
|
+
return
|
|
290
|
+
fi
|
|
291
|
+
printf 'RESTART_PEER|%s\n' "$HERDR_UNIT"
|
|
292
|
+
return
|
|
293
|
+
fi
|
|
294
|
+
if [[ $unit == "$HERDR_UNIT" ]]; then
|
|
295
|
+
printf 'RECOVER_PANE|%s\n' "$HERDR_UNIT"
|
|
296
|
+
return
|
|
297
|
+
fi
|
|
298
|
+
printf 'NONE|\n'
|
|
299
|
+
}
|
|
300
|
+
|
|
301
|
+
# --------------------------------------------------------------------------
|
|
302
|
+
# main
|
|
303
|
+
# --------------------------------------------------------------------------
|
|
304
|
+
|
|
305
|
+
run_recovery() {
|
|
306
|
+
local failed="" u action_name="" action_arg="" ev="" outcome="" ok=1
|
|
307
|
+
for u in ${FLEET_UNITS//:/ }; do
|
|
308
|
+
if unit_is_failed "$u"; then
|
|
309
|
+
failed=$u
|
|
310
|
+
break
|
|
311
|
+
fi
|
|
312
|
+
done
|
|
313
|
+
|
|
314
|
+
if dry_run; then
|
|
315
|
+
say "dry run: nothing will be written or executed"
|
|
316
|
+
[[ -z $failed ]] && say "plan: no unit failed — escalate only"
|
|
317
|
+
if [[ -n $failed ]]; then
|
|
318
|
+
IFS='|' read -r action_name action_arg <<<"$(choose_action "$failed")"
|
|
319
|
+
say "plan: failed unit $failed; action: $action_name"
|
|
320
|
+
fi
|
|
321
|
+
say "plan: escalation — $(escalation_text "$failed" "${action_name:-none}" "dry run" "DRY-RUN")"
|
|
322
|
+
return 0
|
|
323
|
+
fi
|
|
324
|
+
|
|
325
|
+
ev=$(write_evidence "$failed")
|
|
326
|
+
if [[ -z $ev ]]; then
|
|
327
|
+
log "could not write evidence under $RECOVER_DIR — treating the recovery as failed"
|
|
328
|
+
return 1
|
|
329
|
+
fi
|
|
330
|
+
|
|
331
|
+
local attempts
|
|
332
|
+
if [[ -z $failed ]]; then
|
|
333
|
+
# Nothing failed when recovery ran (the fleet got itself up, or an
|
|
334
|
+
# external start beat this activation): always escalate, and a previously
|
|
335
|
+
# exhausted episode counter gets a fresh start.
|
|
336
|
+
say "no unit failed when this recovery ran — the fleet got itself up (or was started externally); nothing to attempt"
|
|
337
|
+
append_evidence "$ev" "decision: NONE — no unit failed"
|
|
338
|
+
if escalate none none "no unit failed when the recovery ran" "$ev"; then
|
|
339
|
+
clear_failed_attempts
|
|
340
|
+
return 0
|
|
341
|
+
fi
|
|
342
|
+
return 1
|
|
343
|
+
fi
|
|
344
|
+
|
|
345
|
+
attempts=$(read_failed_attempts)
|
|
346
|
+
if [[ $attempts -ge $MAX_FAILED_ATTEMPTS ]]; then
|
|
347
|
+
say "attempts exhausted: $attempts consecutive failed recoveries inside the ${FAIL_WINDOW_S}s window; escalating only"
|
|
348
|
+
append_evidence "$ev" "decision: NONE (attempts exhausted, $attempts)"
|
|
349
|
+
if escalate "$failed" none "attempts exhausted ($attempts); no further recovery" "$ev"; then
|
|
350
|
+
return 0
|
|
351
|
+
fi
|
|
352
|
+
return 1
|
|
353
|
+
fi
|
|
354
|
+
|
|
355
|
+
say "failed unit: $failed"
|
|
356
|
+
IFS='|' read -r action_name action_arg <<<"$(choose_action "$failed")"
|
|
357
|
+
action_name=${action_name:-NONE}
|
|
358
|
+
append_evidence "$ev" "decision: $action_name"
|
|
359
|
+
if [[ $action_name == NONE ]]; then
|
|
360
|
+
say "no safe mechanical action for $failed; escalating"
|
|
361
|
+
if escalate "$failed" none "no known-safe action" "$ev"; then
|
|
362
|
+
return 0
|
|
363
|
+
fi
|
|
364
|
+
return 1
|
|
365
|
+
fi
|
|
366
|
+
|
|
367
|
+
# Exactly one action, no matter how many units are down.
|
|
368
|
+
case "$action_name" in
|
|
369
|
+
ROLLBACK)
|
|
370
|
+
if cp "$action_arg" "$CONFIG_PATH" && start_unit "$failed"; then
|
|
371
|
+
if wait_no_longer_failed "$failed"; then
|
|
372
|
+
ok=0
|
|
373
|
+
outcome="restored pre-upgrade snapshot ${action_arg##*/} and started $failed"
|
|
374
|
+
else
|
|
375
|
+
outcome="restored pre-upgrade snapshot and started $failed, but it is still failing"
|
|
376
|
+
fi
|
|
377
|
+
else
|
|
378
|
+
outcome="could not restore snapshot ${action_arg##*/} or start $failed"
|
|
379
|
+
fi
|
|
380
|
+
;;
|
|
381
|
+
RECOVER_PANE)
|
|
382
|
+
if start_unit "$HERDR_UNIT"; then
|
|
383
|
+
if wait_no_longer_failed "$HERDR_UNIT"; then
|
|
384
|
+
ok=0
|
|
385
|
+
outcome="restarted session host $HERDR_UNIT; its startup hook recovers the fleet pane"
|
|
386
|
+
else
|
|
387
|
+
outcome="$HERDR_UNIT started but is still failing"
|
|
388
|
+
fi
|
|
389
|
+
else
|
|
390
|
+
outcome="could not start $HERDR_UNIT"
|
|
391
|
+
fi
|
|
392
|
+
;;
|
|
393
|
+
RESTART_PEER)
|
|
394
|
+
if "$SYSTEMCTL" restart "$action_arg"; then
|
|
395
|
+
outcome="restarted peer unit $action_arg"
|
|
396
|
+
else
|
|
397
|
+
outcome="could not restart peer unit $action_arg"
|
|
398
|
+
fi
|
|
399
|
+
;;
|
|
400
|
+
esac
|
|
401
|
+
say "outcome: $outcome"
|
|
402
|
+
append_evidence "$ev" "outcome: $outcome"
|
|
403
|
+
|
|
404
|
+
if (( ok == 0 )); then
|
|
405
|
+
append_evidence "$ev" "verify: $failed no longer failed"
|
|
406
|
+
say "recovery succeeded: $failed no longer failed"
|
|
407
|
+
clear_failed_attempts
|
|
408
|
+
else
|
|
409
|
+
append_evidence "$ev" "verify: $failed still failed"
|
|
410
|
+
say "recovery failed: $failed still failed; attempts remaining until exhausted"
|
|
411
|
+
bump_failed_attempts
|
|
412
|
+
fi
|
|
413
|
+
|
|
414
|
+
if escalate "$failed" "$action_name" "$outcome" "$ev"; then
|
|
415
|
+
if (( ok == 0 )); then
|
|
416
|
+
return 0
|
|
417
|
+
fi
|
|
418
|
+
return 1
|
|
419
|
+
fi
|
|
420
|
+
return 1
|
|
421
|
+
}
|
|
422
|
+
|
|
423
|
+
main() {
|
|
424
|
+
dry_run && say "dry run: printing decisions, mutating nothing"
|
|
425
|
+
if run_recovery; then
|
|
426
|
+
exit 0
|
|
427
|
+
fi
|
|
428
|
+
exit 1
|
|
429
|
+
}
|
|
430
|
+
|
|
431
|
+
if [[ ${BASH_SOURCE[0]} == "$0" ]]; then
|
|
432
|
+
main
|
|
433
|
+
fi
|
|
@@ -16,6 +16,12 @@
|
|
|
16
16
|
# 4. Do not co-locate ClickHouse + other multi-GB services beside a 2-worker
|
|
17
17
|
# fleet on a ≤8 GB box.
|
|
18
18
|
#
|
|
19
|
+
# OnFailure (issue #485): this line arms the fleet recovery oneshot that #485
|
|
20
|
+
# ships via `setup host` — `omp-conductor-recover.service`, staged from
|
|
21
|
+
# omp-conductor-recover.sh. Do not remove it: a crash-looped unit used to sit in
|
|
22
|
+
# `failed` indefinitely with nobody watching it. Re-running `setup host`
|
|
23
|
+
# provisions the recovery unit and installs the playbook itself.
|
|
24
|
+
#
|
|
19
25
|
# Install:
|
|
20
26
|
# sudo install -m 0644 omp-conductor.service.example /etc/systemd/system/omp-conductor.service
|
|
21
27
|
# sudo systemctl daemon-reload
|
|
@@ -27,6 +33,7 @@ Description=omp-conductor dispatch daemon
|
|
|
27
33
|
Documentation=https://github.com/TerrifiedBug/conductor
|
|
28
34
|
After=network-online.target
|
|
29
35
|
Wants=network-online.target
|
|
36
|
+
OnFailure=omp-conductor-recover.service
|
|
30
37
|
|
|
31
38
|
[Service]
|
|
32
39
|
Type=simple
|
|
@@ -44,9 +51,13 @@ WorkingDirectory=/home/fleet
|
|
|
44
51
|
# under a unit, call `daemon` directly.
|
|
45
52
|
ExecStart=/home/fleet/.local/bin/omp-conductor daemon --port 8787
|
|
46
53
|
|
|
47
|
-
Restart=
|
|
48
|
-
#
|
|
49
|
-
#
|
|
54
|
+
Restart=always
|
|
55
|
+
# Restart on any exit — clean, signalled or crashed — except an explicit
|
|
56
|
+
# `systemctl stop`, which systemd records as intentional and never undoes
|
|
57
|
+
# (issue #546). A stray SIGTERM (e.g. a worker's `bun test`) then costs
|
|
58
|
+
# seconds of downtime instead of leaving the fleet down until a human;
|
|
59
|
+
# a crash loop still trips the start-limit burst and reaches OnFailure above.
|
|
60
|
+
# SuccessExitStatus=0 143 keeps a handled SIGTERM shutdown logged as success.
|
|
50
61
|
SuccessExitStatus=0 143
|
|
51
62
|
|
|
52
63
|
MemoryAccounting=yes
|