yamine 0.21.1 → 0.21.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 8e920f5f6d29a92e0a6e81ca671bccf62e0a264ec453f9b5bc4d42d4e56980d3
4
- data.tar.gz: cc06342930ee7c882cf30fd2818f1a60495c8a59e44ec6779cc666f569493ac0
3
+ metadata.gz: 618e529eeea12aa43397e084455136a4062bd55a546e904a1cb4170e3fc1466f
4
+ data.tar.gz: aa598da83330345ca84a60b168539262aa516df7440d6d91611b5cc853c584ec
5
5
  SHA512:
6
- metadata.gz: 4164dabf9bd5537023ad59d426a958a5e284a90f0226cc83e0533cf635edf0c65acc16dd311d58a88df83d7abdcdfb92034703ecda06d89af08bb5ec96c73635
7
- data.tar.gz: 1e3af2b55e54885fe8aeed98856c4371eab280ec981cd2644ebb75aba0206752ae44307c0252f51188a778426efbb882ef2285c63f142def39f4e876f77302d6
6
+ metadata.gz: 848bb6e6cced84dc132e38dc2ce4828c3514b4fca2c4596e796e9e938b1b2f23d164aadea56e14aed9e8c13d7f732723ede5b377f443db068f74bfd7a9e88410
7
+ data.tar.gz: 453dd9432df8527eace777b7894abf8540b60f103243ce4917593cd021c38506a0aa76614ca0eb926caea5ebfca6f25177729148e5342f22e1b571f7571d104a
data/CHANGELOG.md CHANGED
@@ -2,6 +2,40 @@
2
2
 
3
3
  ## [Unreleased]
4
4
 
5
+ ## [0.21.2] — 2026-09-29
6
+
7
+ ### Fixed
8
+
9
+ - **Linux no longer orphans the app behind a boot's shell.** Every
10
+ process is spawned as `["sh", "-c", cmd]`, so the pid yamine tracks
11
+ is the *shell* — and on Linux a TERM to that shell does not reach the
12
+ app behind it, which reparents to init and keeps running. Every stop
13
+ path inherited the hole: `yamine stop` printed "Stopped <host>" and
14
+ left the app serving, Ctrl-C left the whole tree up, and
15
+ `yamine worktree remove` deleted the directory under a live process.
16
+ (macOS forwards the signal, which is why the unit suite never saw
17
+ it.) Each process now spawns as its own process-group leader and
18
+ stop signals the *group*, so one syscall reaches the shell and
19
+ everything below it. Only a pid that actually leads a group is
20
+ signalled by group id — the kernel is asked first — so the paths that
21
+ never had a shell in front of them (a directly-spawned puma, a pid
22
+ that is already gone) keep signalling exactly as before.
23
+ - **A stop that lands while the tree is still being built no longer
24
+ leaves a live process behind.** The group signal reaches the members
25
+ that exist at the instant it is sent, and a `sh -c` shell that has
26
+ not forked its app yet forks it *after* that sweep — so the app never
27
+ hears the TERM, reparents to init, and goes on serving with its
28
+ directory already on the way out. Measured on linux: TERM the group a
29
+ few milliseconds after the spawn and the shell dies while the app
30
+ comes up behind it and stays up. Asking was never the same as
31
+ stopping, and every caller that goes on to report "Stopped", drop the
32
+ route, or delete the worktree was exposed. Stops now ask and then
33
+ insist: TERM the group, let it leave on its own terms briefly, and
34
+ KILL whatever is still in it — liveness asked of the *group*, since a
35
+ dead leader that is a zombie still answers to its own pid while a live
36
+ tree behind it does not. The signal-trap path keeps the single-shot
37
+ signal, because a trap handler must not sleep.
38
+
5
39
  ## [0.21.1] — 2026-09-29
6
40
 
7
41
  ### Fixed
@@ -345,10 +345,19 @@ module Yamine
345
345
  stop_spawned_pid(app.pid)
346
346
  end
347
347
 
348
+ # The pid of a spawned process is its `sh -c` shell, so the whole
349
+ # group gets the signal: the app behind the shell is what must
350
+ # actually stop. Falls back to the single pid for a process that
351
+ # leads no group, and never raises (see ProcessTree).
352
+ #
353
+ # Escalated, because this is the "a failure leaves nothing running"
354
+ # contract and asking is not enough: a shell still assembling its
355
+ # tree forks the app after the signal sweep, and that app is then
356
+ # unreachable by the signal that was just sent to its group
357
+ # (ProcessTree.terminate). The signal-trap path stays on the
358
+ # single-shot `term` — a trap handler must not sleep.
348
359
  def stop_spawned_pid(pid)
349
- Process.kill("TERM", pid)
350
- rescue SystemCallError
351
- nil
360
+ ProcessTree.terminate(pid)
352
361
  end
353
362
 
354
363
  # A background process (proxy: false) is spawned, logged, and
@@ -853,9 +862,10 @@ module Yamine
853
862
  $stderr.puts " log: #{log}" if File.file?(log)
854
863
  reporter&.note("#{name} exited; cleaning up routes")
855
864
  cleanup_routes(ctx, hostnames)
856
- # Kill remaining children
865
+ # Kill remaining children — by group, so a `sh -c` backend's
866
+ # process dies with the shell we hold a pid for.
857
867
  named_pids.each_key do |other|
858
- Process.kill("TERM", other) rescue nil
868
+ ProcessTree.term(other)
859
869
  end
860
870
  exit 0
861
871
  end
@@ -893,7 +903,11 @@ module Yamine
893
903
  %w[INT TERM].each do |sig|
894
904
  trap(sig) do
895
905
  cleanup_routes(ctx, hostnames)
896
- pids.each { |pid| Process.kill("TERM", pid) rescue nil }
906
+ # The children are in their own process groups, so the
907
+ # terminal's signal reaches only us — this is the one place
908
+ # their trees get told to stop. By group, not by pid: a
909
+ # `sh -c` backend's real process is behind the pid we hold.
910
+ pids.each { |pid| ProcessTree.term(pid) }
897
911
  exit 0
898
912
  end
899
913
  end
@@ -905,7 +919,9 @@ module Yamine
905
919
  entry = ctx.store.find(h)
906
920
  pid = ctx.backend_pid_for(entry) if entry
907
921
  if pid && ProxyControl.pid_alive?(pid)
908
- Process.kill("TERM", pid) rescue nil
922
+ # Group, not pid: the sidecar names the `sh -c` shell, and
923
+ # the app behind it is the process that must stop.
924
+ ProcessTree.term(pid)
909
925
  end
910
926
  ctx.store.remove_route(h, owner_pid: Process.pid) rescue nil
911
927
  FileUtils.rm_f(File.join(ctx.store.dir, "backend-#{h}.pid"))
@@ -206,14 +206,21 @@ module Yamine
206
206
 
207
207
  backend_pid = ctx.backend_pid_for(entry)
208
208
  if backend_pid && ProxyControl.pid_alive?(backend_pid)
209
- begin
210
- Process.kill("TERM", backend_pid)
209
+ # The sidecar names the `sh -c` shell a boot-mode backend was
210
+ # spawned behind, so the whole group is signalled — otherwise
211
+ # "Stopped" is reported while the app keeps serving (linux).
212
+ # And it escalates rather than merely asking: a shell still
213
+ # assembling its tree forks the app after the signal sweep,
214
+ # so that app never hears the TERM and "Stopped" would be a
215
+ # lie (ProcessTree.terminate). A false return means there
216
+ # was nothing left to signal.
217
+ if ProcessTree.terminate(backend_pid)
211
218
  if ctx.wait_for_exit(backend_pid, timeout: 10)
212
219
  stopped << "#{hostname} (backend #{backend_pid})"
213
220
  else
214
221
  stopped << "#{hostname} (backend #{backend_pid} still draining)"
215
222
  end
216
- rescue SystemCallError
223
+ else
217
224
  gone << hostname
218
225
  end
219
226
  else
@@ -604,12 +604,20 @@ module Yamine
604
604
  hostname = entry["hostname"]
605
605
  backend_pid = ctx.backend_pid_for(entry)
606
606
  if backend_pid && ProxyControl.pid_alive?(backend_pid)
607
- begin
608
- Process.kill("TERM", backend_pid)
609
- ctx.wait_for_exit(backend_pid, timeout: 10)
610
- rescue SystemCallError
611
- nil
612
- end
607
+ # Group, not pid: a boot-mode backend's sidecar names the
608
+ # `sh -c` shell, and on linux the app behind that shell does
609
+ # not stop with it (a live process whose cwd is about to be
610
+ # deleted is exactly what must not survive here). A pid that
611
+ # leads no group is signalled on its own, as before.
612
+ #
613
+ # And not a bare TERM: the directory is about to be deleted
614
+ # out from under whatever is left, so this stop has to be
615
+ # able to report "stopped" and mean it. `terminate`
616
+ # escalates, because a shell that is still assembling its
617
+ # tree forks the app after the signal sweep, and that app
618
+ # never hears the TERM at all (ProcessTree.terminate).
619
+ ProcessTree.terminate(backend_pid)
620
+ ctx.wait_for_exit(backend_pid, timeout: 10)
613
621
  end
614
622
  ctx.store.remove_route(hostname)
615
623
  FileUtils.rm_f(File.join(ctx.store.dir, "backend-#{hostname}.pid"))
@@ -0,0 +1,201 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Yamine
4
+ # Stopping a spawned process AND everything it started.
5
+ #
6
+ # Every run-mode process is spawned as ["sh", "-c", cmd] (see
7
+ # BootCommand.collect_spawns), so the pid yamine owns is the SHELL and
8
+ # the app is the shell's child. Signalling that pid alone is not
9
+ # stopping the app: on linux the shell's child is outside the shell's
10
+ # own signal scope, so TERM to the shell leaves the app running,
11
+ # reparented to init. `yamine stop` reports "Stopped <host>" and the
12
+ # app goes on serving, the route is gone, and nothing holds a handle
13
+ # to the process anymore. macOS forwards the signal, which is why a
14
+ # linux-only defect stayed invisible to the unit suite.
15
+ #
16
+ # Two halves, and the second is only safe with the first:
17
+ # * spawn each process as its own group leader (pgroup: true), so
18
+ # the whole tree shares a group that dies together;
19
+ # * signal the GROUP (-pid), which reaches the shell and every
20
+ # descendant in one syscall.
21
+ #
22
+ # A negative pid means "the process group whose id is that number",
23
+ # not "this process and its children" — so it is only correct for a
24
+ # pid that actually LEADS a group, and the two ways to know that are
25
+ # not equally good.
26
+ #
27
+ # The kernel can be asked, and it is the only option for a pid that
28
+ # came from a file (`yamine stop` and `yamine worktree remove` run in
29
+ # a different process than the spawn, with nothing but a pid from a
30
+ # route entry or a sidecar). But that is a reading, and it has two
31
+ # ways of being wrong:
32
+ #
33
+ # * ESRCH. The pid is gone — yet a process group outlives its
34
+ # leader, and the members behind it are exactly what needs
35
+ # stopping: a `sh -c` shell that died on its own leaves the app
36
+ # running in its group with nothing left to signal it by pid. A
37
+ # dead leader answers ESRCH, which reads exactly like "leads no
38
+ # group".
39
+ #
40
+ # * "Same group as me". A pid that never led a group answers this
41
+ # (a directly-spawned puma), and so does one of ours that has not
42
+ # run its setpgid yet — `pgroup: true` puts that in the child, so
43
+ # whether the parent can look first is a property of the spawn
44
+ # path, not something to bet a stop on.
45
+ #
46
+ # So a process spawned through `ProcessTree.spawn` is recorded as a
47
+ # leader, and the record is a fact about what we asked for rather than
48
+ # a reading of the moment. Everything else is the kernel's answer, and
49
+ # a pid number on its own is never evidence: 1234 may be a live
50
+ # process that inherited its group, and `kill(-1234)` would then hit
51
+ # whatever unrelated group wears that id.
52
+ module ProcessTree
53
+ # Pids we spawned as group leaders. A plain Hash, deliberately
54
+ # unlocked: every operation on it is a single call the GVL makes
55
+ # atomic, and yamine stops its processes from inside a trap handler
56
+ # (BootCommand.trap_cleanup), where Mutex#synchronize raises
57
+ # "can't be called from trap context" — a stop path that only works
58
+ # outside a trap is a stop that never happens on Ctrl-C.
59
+ PGROUP_LEADERS = {}
60
+
61
+ module_function
62
+
63
+ # Spawn a boot process as its own group leader and record that it is
64
+ # one. This is the only place a boot process is created, so "every
65
+ # process we boot leads a group" is one fact in one place rather
66
+ # than a convention spread across spawns.
67
+ def spawn(*args, **opts)
68
+ pid = ::Process.spawn(*args, pgroup: true, **opts)
69
+ note_group_leader(pid)
70
+ pid
71
+ end
72
+
73
+ # A pid we spawned with pgroup: true leads its own group — that is
74
+ # what we asked for, so the kernel is not consulted. Two readings it
75
+ # could give are both wrong here: "still in my group" while the
76
+ # child has not run its setpgid, and ESRCH once it is gone while
77
+ # the members it left behind are still running.
78
+ def note_group_leader(pid)
79
+ return unless pid.to_i.positive?
80
+
81
+ PGROUP_LEADERS[pid] = true
82
+ end
83
+
84
+ def known_group_leader?(pid)
85
+ PGROUP_LEADERS.key?(pid)
86
+ end
87
+
88
+ def forget(pid)
89
+ PGROUP_LEADERS.delete(pid)
90
+ end
91
+
92
+ # Does this pid lead its own process group? Ours, if we spawned it.
93
+ # Otherwise the kernel — and its "no" is the answer we want: a dead
94
+ # pid and a process that inherited its group are exactly the pids
95
+ # that must be signalled on their own, never by group id.
96
+ def group_leader?(pid)
97
+ return true if known_group_leader?(pid)
98
+
99
+ Process.getpgid(pid) == pid
100
+ rescue SystemCallError
101
+ false
102
+ end
103
+
104
+ # TERM a spawned process and its whole tree. Returns true when the
105
+ # signal was delivered, false when there was nothing left to signal
106
+ # (already dead, or a pid we may not touch) — which is the end
107
+ # state every caller wants, not an error, so this never raises.
108
+ def term(pid, signal: "TERM")
109
+ return false unless pid.to_i.positive?
110
+
111
+ group_signal(pid, signal) || pid_signal(pid, signal)
112
+ ensure
113
+ # The record is about a spawn, not about a pid that must keep
114
+ # this meaning: dropping it here keeps a pid that later gets
115
+ # recycled from being signalled as a group it never led.
116
+ forget(pid) if pid.to_i.positive?
117
+ end
118
+
119
+ # The whole group, one syscall. Best-effort: returns false when
120
+ # there is no such group left, so the caller can fall back.
121
+ def group_signal(pid, signal)
122
+ return false unless group_leader?(pid)
123
+
124
+ Process.kill(signal, -pid)
125
+ true
126
+ rescue SystemCallError
127
+ # ESRCH: the group outlived neither the leader nor its members.
128
+ # EPERM: the group exists but is not ours to signal. Either way a
129
+ # single-pid signal is the last thing left to try.
130
+ false
131
+ end
132
+
133
+ def pid_signal(pid, signal)
134
+ Process.kill(signal, pid)
135
+ true
136
+ rescue SystemCallError
137
+ false
138
+ end
139
+
140
+ # Stop a tree and make sure it is GONE — for the callers that go on
141
+ # to delete the directory it runs in, drop the route, or report
142
+ # "Stopped <host>". `term` alone is not that promise: it hands the
143
+ # signal to the group and returns, and the kernel's sweep covers
144
+ # the members that exist at that instant. A `sh -c` shell that is
145
+ # still assembling its tree forks its app *after* the sweep, and
146
+ # that app never sees the signal at all.
147
+ #
148
+ # Measured on linux: TERM the group of a `sh -c "ruby <app>"` a
149
+ # few milliseconds after the spawn and the shell dies while the app
150
+ # comes up behind it, reparented to init and serving forever —
151
+ # a live process whose working directory is on the way out, which
152
+ # is the exact shape of the bug this module exists to kill. A stop
153
+ # that only asks cannot rule that out, so this asks and then
154
+ # insists: TERM, give the group a moment to leave on its own
155
+ # terms, KILL whatever is still in it.
156
+ #
157
+ # The liveness question is asked of the GROUP, not of the pid: a
158
+ # dead leader that is a zombie still answers `kill(0, pid)`, and a
159
+ # live tree behind one does not answer to its pid at all. The group
160
+ # is the only handle that tells the two apart.
161
+ #
162
+ # `term` is deliberately left alone: it is the primitive a signal
163
+ # trap handler uses (BootCommand.trap_cleanup), and a trap handler
164
+ # must not sit in a sleep loop.
165
+ def terminate(pid, grace: 5)
166
+ return false unless term(pid)
167
+
168
+ deadline = monotonic + grace
169
+ sleep 0.05 while group?(pid) && monotonic < deadline
170
+ return true unless group?(pid)
171
+
172
+ # A group that outlived its grace period is not draining, it is
173
+ # stuck or ignoring us, and the caller is about to delete what it
174
+ # runs in. No group_leader? check here: the leader is very likely
175
+ # already gone, and it is the members behind it that need the
176
+ # signal.
177
+ Process.kill("KILL", -pid)
178
+ true
179
+ rescue SystemCallError
180
+ # The group went away between the last look and this signal.
181
+ true
182
+ end
183
+
184
+ # Does this process group still have a member in it? Signal 0 asks
185
+ # the kernel about the group rather than about a pid, so it stays
186
+ # true while anything behind a dead leader is still running, and
187
+ # goes false only once the tree really is down.
188
+ def group?(pid)
189
+ return false unless pid.to_i.positive?
190
+
191
+ Process.kill(0, -pid)
192
+ true
193
+ rescue SystemCallError
194
+ false
195
+ end
196
+
197
+ def monotonic
198
+ Process.clock_gettime(Process::CLOCK_MONOTONIC)
199
+ end
200
+ end
201
+ end
data/lib/yamine/runner.rb CHANGED
@@ -96,6 +96,12 @@ module Yamine
96
96
  # view yamine's failure payloads use. Interleaving matches the
97
97
  # foreman/without-yamine experience.
98
98
  #
99
+ # Spawned through ProcessTree, not Kernel#spawn: the command is an
100
+ # argv that starts with `sh -c` (collect_spawns), so the pid we get
101
+ # back is the shell. Its own process group is what makes the process
102
+ # behind it stoppable — stop signals the group (ProcessTree.term),
103
+ # and the shell cannot leave its child outside the group it leads.
104
+ #
99
105
  # extra_env is the config's env.clear/env.secret for this process —
100
106
  # yamine's own keys (PORT, YAMINE_URL, DATABASE_URL, the TLS vars)
101
107
  # win over it, because those describe the boot rather than the app.
@@ -106,7 +112,7 @@ module Yamine
106
112
  env = child_env(dir, url: url, port: port, rails_dev_host: rails_dev_host,
107
113
  database_url: database_url, database_env: database_env, extra_env: extra_env)
108
114
  path = log_path(dir, name)
109
- pid = with_clean_env { spawn(env, *command, chdir: dir, out: path, err: [:child, :out]) }
115
+ pid = with_clean_env { ProcessTree.spawn(env, *command, chdir: dir, out: path, err: [:child, :out]) }
110
116
  Process.detach(pid)
111
117
  target = "127.0.0.1:#{port}"
112
118
  if register
@@ -130,12 +136,16 @@ module Yamine
130
136
  # whole tree up concurrently, then polls every route until healthy.
131
137
  # Returns the placeholder App (target known before bind). Fate of
132
138
  # the backend is decided by wait, not by spawn.
139
+ #
140
+ # Through ProcessTree, for the same reason as boot_run: a `sh -c`
141
+ # pid is not the process doing the work, so its group is the only
142
+ # handle that reaches it.
133
143
  def spawn_http(name:, hostname:, url:, dir:, command:, port:, rails_dev_host: nil,
134
144
  database_url: nil, database_env: nil, force: false, extra_env: nil)
135
145
  env = child_env(dir, url: url, port: port, rails_dev_host: rails_dev_host,
136
146
  database_url: database_url, database_env: database_env, extra_env: extra_env)
137
147
  path = log_path(dir, name)
138
- pid = with_clean_env { spawn(env, *command, chdir: dir, out: path, err: [:child, :out]) }
148
+ pid = with_clean_env { ProcessTree.spawn(env, *command, chdir: dir, out: path, err: [:child, :out]) }
139
149
  Process.detach(pid)
140
150
  App.new(name: name, hostname: hostname, url: url, pid: pid,
141
151
  target: "127.0.0.1:#{port}", kind: "tcp", command: command)
@@ -340,10 +350,14 @@ module Yamine
340
350
 
341
351
  private
342
352
 
353
+ # A backend we spawned is stopped by its whole group: the pid we
354
+ # hold for a run-mode process is the `sh -c` shell, and on linux the
355
+ # app behind it does not see a signal aimed at that shell alone.
356
+ # A pid that leads no group (the socket spawns, a pid already gone)
357
+ # falls back to the single-pid signal. Never raises: a stopped
358
+ # backend is the goal, and "it was already dead" is a success here.
343
359
  def stop_pid(pid)
344
- Process.kill("TERM", pid)
345
- rescue SystemCallError
346
- nil
360
+ ProcessTree.term(pid)
347
361
  end
348
362
  end
349
363
  end
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module Yamine
4
- VERSION = "0.21.1"
4
+ VERSION = "0.21.2"
5
5
  end
data/lib/yamine.rb CHANGED
@@ -19,6 +19,7 @@ require_relative "yamine/log"
19
19
  require_relative "yamine/route_store"
20
20
  require_relative "yamine/certs"
21
21
  require_relative "yamine/ports"
22
+ require_relative "yamine/process_tree"
22
23
  require_relative "yamine/hosts"
23
24
  require_relative "yamine/proxy"
24
25
  require_relative "yamine/proxy_control"
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: yamine
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.21.1
4
+ version: 0.21.2
5
5
  platform: ruby
6
6
  authors:
7
7
  - Kaka Ruto
@@ -105,6 +105,7 @@ files:
105
105
  - lib/yamine/ports.rb
106
106
  - lib/yamine/privileged_payload.rb
107
107
  - lib/yamine/probe.rb
108
+ - lib/yamine/process_tree.rb
108
109
  - lib/yamine/procfile.rb
109
110
  - lib/yamine/proxy.rb
110
111
  - lib/yamine/proxy_control.rb