yamine 0.21.1 → 0.22.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
data/lib/yamine/cli.rb CHANGED
@@ -83,6 +83,7 @@ module Yamine
83
83
 
84
84
  Usage:
85
85
  yamine start One-setup-and-go: setup if needed, then boot -> https://<app>.localhost
86
+ yamine start --detach Same, in the background; returns once the app is healthy
86
87
  yamine setup One-shot workstation setup without booting (run once)
87
88
  yamine Bare form of `start` — boots every process in config/local.yml
88
89
  yamine get <name> Print URL for a service
@@ -103,10 +104,10 @@ module Yamine
103
104
  yamine hosts sync|clean Manage /etc/hosts entries
104
105
  yamine kamal <variant> Preview-deploy snippet for Kamal
105
106
  yamine stop Stop this app's backend + routes
106
- yamine restart Touch tmp/restart.txt
107
+ yamine restart Touch tmp/restart.txt (a supervised app is stopped)
107
108
  yamine log [-F] [n] Tail (or follow) log/development.log
108
109
 
109
- Flags: --variant, --tld, --force, --app-port, --wait (default), --no-wait, --json, --branch
110
+ Flags: --variant, --tld, --force, --app-port, --wait (default), --no-wait, --detach, --json, --branch
110
111
  Env: YAMINE_VARIANT/TLD/PORT/STATE_DIR/AGENT, YAMINE_BRANCH=1
111
112
  HELP
112
113
  end
@@ -0,0 +1,256 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Yamine
4
+ # Stopping a spawned process AND everything it started.
5
+ #
6
+ # Every run-mode process is spawned as ["sh", "-c", cmd] (see
7
+ # BootCommand.collect_spawns), so the pid yamine owns is the SHELL and
8
+ # the app is the shell's child. Signalling that pid alone is not
9
+ # stopping the app: on linux the shell's child is outside the shell's
10
+ # own signal scope, so TERM to the shell leaves the app running,
11
+ # reparented to init. `yamine stop` reports "Stopped <host>" and the
12
+ # app goes on serving, the route is gone, and nothing holds a handle
13
+ # to the process anymore. macOS forwards the signal, which is why a
14
+ # linux-only defect stayed invisible to the unit suite.
15
+ #
16
+ # Two halves, and the second is only safe with the first:
17
+ # * spawn each process as its own group leader (pgroup: true), so
18
+ # the whole tree shares a group that dies together;
19
+ # * signal the GROUP (-pid), which reaches the shell and every
20
+ # descendant in one syscall.
21
+ #
22
+ # A negative pid means "the process group whose id is that number",
23
+ # not "this process and its children" — so it is only correct for a
24
+ # pid that actually LEADS a group, and the two ways to know that are
25
+ # not equally good.
26
+ #
27
+ # The kernel can be asked, and it is the only option for a pid that
28
+ # came from a file (`yamine stop` and `yamine worktree remove` run in
29
+ # a different process than the spawn, with nothing but a pid from a
30
+ # route entry or a sidecar). But that is a reading, and it has two
31
+ # ways of being wrong:
32
+ #
33
+ # * ESRCH. The pid is gone — yet a process group outlives its
34
+ # leader, and the members behind it are exactly what needs
35
+ # stopping: a `sh -c` shell that died on its own leaves the app
36
+ # running in its group with nothing left to signal it by pid. A
37
+ # dead leader answers ESRCH, which reads exactly like "leads no
38
+ # group".
39
+ #
40
+ # * "Same group as me". A pid that never led a group answers this
41
+ # (a directly-spawned puma), and so does one of ours that has not
42
+ # run its setpgid yet — `pgroup: true` puts that in the child, so
43
+ # whether the parent can look first is a property of the spawn
44
+ # path, not something to bet a stop on.
45
+ #
46
+ # So a process spawned through `ProcessTree.spawn` is recorded as a
47
+ # leader, and the record is a fact about what we asked for rather than
48
+ # a reading of the moment. Everything else is the kernel's answer, and
49
+ # a pid number on its own is never evidence: 1234 may be a live
50
+ # process that inherited its group, and `kill(-1234)` would then hit
51
+ # whatever unrelated group wears that id.
52
+ module ProcessTree
53
+ # Pids we spawned as group leaders. A plain Hash, deliberately
54
+ # unlocked: every operation on it is a single call the GVL makes
55
+ # atomic, and yamine stops its processes from inside a trap handler
56
+ # (BootCommand.trap_cleanup), where Mutex#synchronize raises
57
+ # "can't be called from trap context" — a stop path that only works
58
+ # outside a trap is a stop that never happens on Ctrl-C.
59
+ PGROUP_LEADERS = {}
60
+ # Exit statuses of what we spawned, by pid, filled in by the reaper
61
+ # that `detach` starts. Same reasoning, same shape, and the same
62
+ # bounded cost as PGROUP_LEADERS above: an entry is a few dozen bytes
63
+ # per process this one ever spawned, dropped by `forget` whenever
64
+ # the pid is signalled.
65
+ STATUSES = {}
66
+
67
+ module_function
68
+
69
+ # Spawn a child and reap it without blocking us, keeping what it
70
+ # exited with. A drop-in for Process.detach — nobody waits on a live
71
+ # backend, a dead pid is all the boot needs — that also keeps the one
72
+ # thing only a reaper can know. Without it the answer dies with the
73
+ # child: `kill(0, pid)` says a process is gone and nothing says
74
+ # whether it exited 0 or was killed, so "web is down" is all a
75
+ # supervisor could ever report.
76
+ #
77
+ # Unlocked, like PGROUP_LEADERS, and a Hash read from a supervision
78
+ # thread: every operation is a single call the GVL makes atomic.
79
+ def detach(pid)
80
+ return pid unless pid.to_i.positive?
81
+
82
+ Thread.new do
83
+ _waited, status = ::Process.waitpid2(pid)
84
+ STATUSES[pid] = status
85
+ rescue SystemCallError
86
+ nil
87
+ end
88
+ pid
89
+ end
90
+
91
+ # Process::Status for a pid we spawned, or nil when there is nothing
92
+ # to report: not our child, or it has not been reaped yet.
93
+ #
94
+ # Never blocks, which is the whole point — the caller is a loop that
95
+ # has other routes to watch, and asking a live child how it is doing
96
+ # would hang it. So a status that has not arrived yet is nil, and the
97
+ # caller asks again on its next pass.
98
+ def status(pid)
99
+ recorded = STATUSES[pid]
100
+ return recorded if recorded
101
+
102
+ # Not one of ours to have reaped (a pid out of a route entry, a
103
+ # double in a test): ask the kernel. WNOHANG so a live pid is not
104
+ # waited on, and ECHILD — which is the answer for anything that is
105
+ # not a child of ours — is a nil, not a failure.
106
+ _waited, status = ::Process.waitpid2(pid, ::Process::WNOHANG)
107
+ status
108
+ rescue StandardError
109
+ nil
110
+ end
111
+
112
+ # Spawn a boot process as its own group leader and record that it is
113
+ # one. This is the only place a boot process is created, so "every
114
+ # process we boot leads a group" is one fact in one place rather
115
+ # than a convention spread across spawns.
116
+ def spawn(*args, **opts)
117
+ pid = ::Process.spawn(*args, pgroup: true, **opts)
118
+ note_group_leader(pid)
119
+ pid
120
+ end
121
+
122
+ # A pid we spawned with pgroup: true leads its own group — that is
123
+ # what we asked for, so the kernel is not consulted. Two readings it
124
+ # could give are both wrong here: "still in my group" while the
125
+ # child has not run its setpgid, and ESRCH once it is gone while
126
+ # the members it left behind are still running.
127
+ def note_group_leader(pid)
128
+ return unless pid.to_i.positive?
129
+
130
+ PGROUP_LEADERS[pid] = true
131
+ end
132
+
133
+ def known_group_leader?(pid)
134
+ PGROUP_LEADERS.key?(pid)
135
+ end
136
+
137
+ # Drop every record of a pid we have finished with. Both tables are
138
+ # about a spawn, not about a number that has to keep meaning one:
139
+ # leaving them behind is how a pid that later gets recycled gets
140
+ # signalled as a group it never led, or reported with an exit status
141
+ # that belongs to somebody else.
142
+ def forget(pid)
143
+ PGROUP_LEADERS.delete(pid)
144
+ STATUSES.delete(pid)
145
+ end
146
+
147
+ # Does this pid lead its own process group? Ours, if we spawned it.
148
+ # Otherwise the kernel — and its "no" is the answer we want: a dead
149
+ # pid and a process that inherited its group are exactly the pids
150
+ # that must be signalled on their own, never by group id.
151
+ def group_leader?(pid)
152
+ return true if known_group_leader?(pid)
153
+
154
+ Process.getpgid(pid) == pid
155
+ rescue SystemCallError
156
+ false
157
+ end
158
+
159
+ # TERM a spawned process and its whole tree. Returns true when the
160
+ # signal was delivered, false when there was nothing left to signal
161
+ # (already dead, or a pid we may not touch) — which is the end
162
+ # state every caller wants, not an error, so this never raises.
163
+ def term(pid, signal: "TERM")
164
+ return false unless pid.to_i.positive?
165
+
166
+ group_signal(pid, signal) || pid_signal(pid, signal)
167
+ ensure
168
+ # The record is about a spawn, not about a pid that must keep
169
+ # this meaning: dropping it here keeps a pid that later gets
170
+ # recycled from being signalled as a group it never led.
171
+ forget(pid) if pid.to_i.positive?
172
+ end
173
+
174
+ # The whole group, one syscall. Best-effort: returns false when
175
+ # there is no such group left, so the caller can fall back.
176
+ def group_signal(pid, signal)
177
+ return false unless group_leader?(pid)
178
+
179
+ Process.kill(signal, -pid)
180
+ true
181
+ rescue SystemCallError
182
+ # ESRCH: the group outlived neither the leader nor its members.
183
+ # EPERM: the group exists but is not ours to signal. Either way a
184
+ # single-pid signal is the last thing left to try.
185
+ false
186
+ end
187
+
188
+ def pid_signal(pid, signal)
189
+ Process.kill(signal, pid)
190
+ true
191
+ rescue SystemCallError
192
+ false
193
+ end
194
+
195
+ # Stop a tree and make sure it is GONE — for the callers that go on
196
+ # to delete the directory it runs in, drop the route, or report
197
+ # "Stopped <host>". `term` alone is not that promise: it hands the
198
+ # signal to the group and returns, and the kernel's sweep covers
199
+ # the members that exist at that instant. A `sh -c` shell that is
200
+ # still assembling its tree forks its app *after* the sweep, and
201
+ # that app never sees the signal at all.
202
+ #
203
+ # Measured on linux: TERM the group of a `sh -c "ruby <app>"` a
204
+ # few milliseconds after the spawn and the shell dies while the app
205
+ # comes up behind it, reparented to init and serving forever —
206
+ # a live process whose working directory is on the way out, which
207
+ # is the exact shape of the bug this module exists to kill. A stop
208
+ # that only asks cannot rule that out, so this asks and then
209
+ # insists: TERM, give the group a moment to leave on its own
210
+ # terms, KILL whatever is still in it.
211
+ #
212
+ # The liveness question is asked of the GROUP, not of the pid: a
213
+ # dead leader that is a zombie still answers `kill(0, pid)`, and a
214
+ # live tree behind one does not answer to its pid at all. The group
215
+ # is the only handle that tells the two apart.
216
+ #
217
+ # `term` is deliberately left alone: it is the primitive a signal
218
+ # trap handler uses (BootCommand.trap_cleanup), and a trap handler
219
+ # must not sit in a sleep loop.
220
+ def terminate(pid, grace: 5)
221
+ return false unless term(pid)
222
+
223
+ deadline = monotonic + grace
224
+ sleep 0.05 while group?(pid) && monotonic < deadline
225
+ return true unless group?(pid)
226
+
227
+ # A group that outlived its grace period is not draining, it is
228
+ # stuck or ignoring us, and the caller is about to delete what it
229
+ # runs in. No group_leader? check here: the leader is very likely
230
+ # already gone, and it is the members behind it that need the
231
+ # signal.
232
+ Process.kill("KILL", -pid)
233
+ true
234
+ rescue SystemCallError
235
+ # The group went away between the last look and this signal.
236
+ true
237
+ end
238
+
239
+ # Does this process group still have a member in it? Signal 0 asks
240
+ # the kernel about the group rather than about a pid, so it stays
241
+ # true while anything behind a dead leader is still running, and
242
+ # goes false only once the tree really is down.
243
+ def group?(pid)
244
+ return false unless pid.to_i.positive?
245
+
246
+ Process.kill(0, -pid)
247
+ true
248
+ rescue SystemCallError
249
+ false
250
+ end
251
+
252
+ def monotonic
253
+ Process.clock_gettime(Process::CLOCK_MONOTONIC)
254
+ end
255
+ end
256
+ end