yamine 0.21.0 → 0.21.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +55 -0
- data/README.md +4 -3
- data/lib/yamine/cli/boot.rb +23 -7
- data/lib/yamine/cli/routes.rb +10 -3
- data/lib/yamine/cli/worktree.rb +14 -6
- data/lib/yamine/database.rb +1 -0
- data/lib/yamine/process_tree.rb +201 -0
- data/lib/yamine/proxy.rb +28 -2
- data/lib/yamine/runner.rb +19 -5
- data/lib/yamine/version.rb +1 -1
- data/lib/yamine.rb +1 -0
- metadata +2 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 618e529eeea12aa43397e084455136a4062bd55a546e904a1cb4170e3fc1466f
|
|
4
|
+
data.tar.gz: aa598da83330345ca84a60b168539262aa516df7440d6d91611b5cc853c584ec
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 848bb6e6cced84dc132e38dc2ce4828c3514b4fca2c4596e796e9e938b1b2f23d164aadea56e14aed9e8c13d7f732723ede5b377f443db068f74bfd7a9e88410
|
|
7
|
+
data.tar.gz: 453dd9432df8527eace777b7894abf8540b60f103243ce4917593cd021c38506a0aa76614ca0eb926caea5ebfca6f25177729148e5342f22e1b571f7571d104a
|
data/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,61 @@
|
|
|
2
2
|
|
|
3
3
|
## [Unreleased]
|
|
4
4
|
|
|
5
|
+
## [0.21.2] — 2026-09-29
|
|
6
|
+
|
|
7
|
+
### Fixed
|
|
8
|
+
|
|
9
|
+
- **Linux no longer orphans the app behind a boot's shell.** Every
|
|
10
|
+
process is spawned as `["sh", "-c", cmd]`, so the pid yamine tracks
|
|
11
|
+
is the *shell* — and on Linux a TERM to that shell does not reach the
|
|
12
|
+
app behind it, which reparents to init and keeps running. Every stop
|
|
13
|
+
path inherited the hole: `yamine stop` printed "Stopped <host>" and
|
|
14
|
+
left the app serving, Ctrl-C left the whole tree up, and
|
|
15
|
+
`yamine worktree remove` deleted the directory under a live process.
|
|
16
|
+
(macOS forwards the signal, which is why the unit suite never saw
|
|
17
|
+
it.) Each process now spawns as its own process-group leader and
|
|
18
|
+
stop signals the *group*, so one syscall reaches the shell and
|
|
19
|
+
everything below it. Only a pid that actually leads a group is
|
|
20
|
+
signalled by group id — the kernel is asked first — so the paths that
|
|
21
|
+
never had a shell in front of them (a directly-spawned puma, a pid
|
|
22
|
+
that is already gone) keep signalling exactly as before.
|
|
23
|
+
- **A stop that lands while the tree is still being built no longer
|
|
24
|
+
leaves a live process behind.** The group signal reaches the members
|
|
25
|
+
that exist at the instant it is sent, and a `sh -c` shell that has
|
|
26
|
+
not forked its app yet forks it *after* that sweep — so the app never
|
|
27
|
+
hears the TERM, reparents to init, and goes on serving with its
|
|
28
|
+
directory already on the way out. Measured on linux: TERM the group a
|
|
29
|
+
few milliseconds after the spawn and the shell dies while the app
|
|
30
|
+
comes up behind it and stays up. Asking was never the same as
|
|
31
|
+
stopping, and every caller that goes on to report "Stopped", drop the
|
|
32
|
+
route, or delete the worktree was exposed. Stops now ask and then
|
|
33
|
+
insist: TERM the group, let it leave on its own terms briefly, and
|
|
34
|
+
KILL whatever is still in it — liveness asked of the *group*, since a
|
|
35
|
+
dead leader that is a zombie still answers to its own pid while a live
|
|
36
|
+
tree behind it does not. The signal-trap path keeps the single-shot
|
|
37
|
+
signal, because a trap handler must not sleep.
|
|
38
|
+
|
|
39
|
+
## [0.21.1] — 2026-09-29
|
|
40
|
+
|
|
41
|
+
### Fixed
|
|
42
|
+
|
|
43
|
+
- **Claiming a per-worktree database no longer raises `NoMethodError`
|
|
44
|
+
on Ruby 3.2/3.3.** The claim stamps the state file with
|
|
45
|
+
`Time#iso8601`, but `lib/yamine/database.rb` never required the
|
|
46
|
+
stdlib `time` library, and on 3.2/3.3 nothing else loads it
|
|
47
|
+
implicitly — so every claim raised (34 unit errors on those rubies
|
|
48
|
+
in CI). 0.21.0 is affected. One `require "time"`, alongside the
|
|
49
|
+
stdlibs that file already pulls in.
|
|
50
|
+
- **A closed WebSocket upgrade no longer leaks a thread on Linux.**
|
|
51
|
+
`pipe_both` tore each end down with `close` from the sibling
|
|
52
|
+
thread, and a close from another thread does not interrupt that
|
|
53
|
+
copy's blocking read on Linux (it does on macOS) — so every closed
|
|
54
|
+
upgrade parked a thread forever, `handle` never unwound, and they
|
|
55
|
+
accumulated in the root daemon. Teardown is now symmetric
|
|
56
|
+
`shutdown(2)`, which wakes the parked read with EOF; both sockets
|
|
57
|
+
still close exactly once, and no timeout was added to the
|
|
58
|
+
unbounded upgrade path.
|
|
59
|
+
|
|
5
60
|
## [0.21.0] — 2026-09-28
|
|
6
61
|
|
|
7
62
|
### Security
|
data/README.md
CHANGED
|
@@ -508,10 +508,11 @@ bundle exec rake test
|
|
|
508
508
|
self-signed, and there is no buffering or rate limiting.
|
|
509
509
|
|
|
510
510
|
|
|
511
|
-
The `
|
|
511
|
+
The `test/fixtures/apps/` fixture fleet exercises
|
|
512
512
|
detection, inference, and boot across Rails variants, Roda, Sinatra,
|
|
513
|
-
bare Rack, Jekyll, compound Procfiles, and a monorepo.
|
|
514
|
-
|
|
513
|
+
bare Rack, Jekyll, compound Procfiles, and a monorepo. It ships with
|
|
514
|
+
the repo and runs as part of the unit suite across the Ruby matrix —
|
|
515
|
+
no sibling checkout, no separate CI job.
|
|
515
516
|
|
|
516
517
|
## License
|
|
517
518
|
|
data/lib/yamine/cli/boot.rb
CHANGED
|
@@ -345,10 +345,19 @@ module Yamine
|
|
|
345
345
|
stop_spawned_pid(app.pid)
|
|
346
346
|
end
|
|
347
347
|
|
|
348
|
+
# The pid of a spawned process is its `sh -c` shell, so the whole
|
|
349
|
+
# group gets the signal: the app behind the shell is what must
|
|
350
|
+
# actually stop. Falls back to the single pid for a process that
|
|
351
|
+
# leads no group, and never raises (see ProcessTree).
|
|
352
|
+
#
|
|
353
|
+
# Escalated, because this is the "a failure leaves nothing running"
|
|
354
|
+
# contract and asking is not enough: a shell still assembling its
|
|
355
|
+
# tree forks the app after the signal sweep, and that app is then
|
|
356
|
+
# unreachable by the signal that was just sent to its group
|
|
357
|
+
# (ProcessTree.terminate). The signal-trap path stays on the
|
|
358
|
+
# single-shot `term` — a trap handler must not sleep.
|
|
348
359
|
def stop_spawned_pid(pid)
|
|
349
|
-
|
|
350
|
-
rescue SystemCallError
|
|
351
|
-
nil
|
|
360
|
+
ProcessTree.terminate(pid)
|
|
352
361
|
end
|
|
353
362
|
|
|
354
363
|
# A background process (proxy: false) is spawned, logged, and
|
|
@@ -853,9 +862,10 @@ module Yamine
|
|
|
853
862
|
$stderr.puts " log: #{log}" if File.file?(log)
|
|
854
863
|
reporter&.note("#{name} exited; cleaning up routes")
|
|
855
864
|
cleanup_routes(ctx, hostnames)
|
|
856
|
-
# Kill remaining children
|
|
865
|
+
# Kill remaining children — by group, so a `sh -c` backend's
|
|
866
|
+
# process dies with the shell we hold a pid for.
|
|
857
867
|
named_pids.each_key do |other|
|
|
858
|
-
|
|
868
|
+
ProcessTree.term(other)
|
|
859
869
|
end
|
|
860
870
|
exit 0
|
|
861
871
|
end
|
|
@@ -893,7 +903,11 @@ module Yamine
|
|
|
893
903
|
%w[INT TERM].each do |sig|
|
|
894
904
|
trap(sig) do
|
|
895
905
|
cleanup_routes(ctx, hostnames)
|
|
896
|
-
|
|
906
|
+
# The children are in their own process groups, so the
|
|
907
|
+
# terminal's signal reaches only us — this is the one place
|
|
908
|
+
# their trees get told to stop. By group, not by pid: a
|
|
909
|
+
# `sh -c` backend's real process is behind the pid we hold.
|
|
910
|
+
pids.each { |pid| ProcessTree.term(pid) }
|
|
897
911
|
exit 0
|
|
898
912
|
end
|
|
899
913
|
end
|
|
@@ -905,7 +919,9 @@ module Yamine
|
|
|
905
919
|
entry = ctx.store.find(h)
|
|
906
920
|
pid = ctx.backend_pid_for(entry) if entry
|
|
907
921
|
if pid && ProxyControl.pid_alive?(pid)
|
|
908
|
-
|
|
922
|
+
# Group, not pid: the sidecar names the `sh -c` shell, and
|
|
923
|
+
# the app behind it is the process that must stop.
|
|
924
|
+
ProcessTree.term(pid)
|
|
909
925
|
end
|
|
910
926
|
ctx.store.remove_route(h, owner_pid: Process.pid) rescue nil
|
|
911
927
|
FileUtils.rm_f(File.join(ctx.store.dir, "backend-#{h}.pid"))
|
data/lib/yamine/cli/routes.rb
CHANGED
|
@@ -206,14 +206,21 @@ module Yamine
|
|
|
206
206
|
|
|
207
207
|
backend_pid = ctx.backend_pid_for(entry)
|
|
208
208
|
if backend_pid && ProxyControl.pid_alive?(backend_pid)
|
|
209
|
-
|
|
210
|
-
|
|
209
|
+
# The sidecar names the `sh -c` shell a boot-mode backend was
|
|
210
|
+
# spawned behind, so the whole group is signalled — otherwise
|
|
211
|
+
# "Stopped" is reported while the app keeps serving (linux).
|
|
212
|
+
# And it escalates rather than merely asking: a shell still
|
|
213
|
+
# assembling its tree forks the app after the signal sweep,
|
|
214
|
+
# so that app never hears the TERM and "Stopped" would be a
|
|
215
|
+
# lie (ProcessTree.terminate). A false return means there
|
|
216
|
+
# was nothing left to signal.
|
|
217
|
+
if ProcessTree.terminate(backend_pid)
|
|
211
218
|
if ctx.wait_for_exit(backend_pid, timeout: 10)
|
|
212
219
|
stopped << "#{hostname} (backend #{backend_pid})"
|
|
213
220
|
else
|
|
214
221
|
stopped << "#{hostname} (backend #{backend_pid} still draining)"
|
|
215
222
|
end
|
|
216
|
-
|
|
223
|
+
else
|
|
217
224
|
gone << hostname
|
|
218
225
|
end
|
|
219
226
|
else
|
data/lib/yamine/cli/worktree.rb
CHANGED
|
@@ -604,12 +604,20 @@ module Yamine
|
|
|
604
604
|
hostname = entry["hostname"]
|
|
605
605
|
backend_pid = ctx.backend_pid_for(entry)
|
|
606
606
|
if backend_pid && ProxyControl.pid_alive?(backend_pid)
|
|
607
|
-
|
|
608
|
-
|
|
609
|
-
|
|
610
|
-
|
|
611
|
-
|
|
612
|
-
|
|
607
|
+
# Group, not pid: a boot-mode backend's sidecar names the
|
|
608
|
+
# `sh -c` shell, and on linux the app behind that shell does
|
|
609
|
+
# not stop with it (a live process whose cwd is about to be
|
|
610
|
+
# deleted is exactly what must not survive here). A pid that
|
|
611
|
+
# leads no group is signalled on its own, as before.
|
|
612
|
+
#
|
|
613
|
+
# And not a bare TERM: the directory is about to be deleted
|
|
614
|
+
# out from under whatever is left, so this stop has to be
|
|
615
|
+
# able to report "stopped" and mean it. `terminate`
|
|
616
|
+
# escalates, because a shell that is still assembling its
|
|
617
|
+
# tree forks the app after the signal sweep, and that app
|
|
618
|
+
# never hears the TERM at all (ProcessTree.terminate).
|
|
619
|
+
ProcessTree.terminate(backend_pid)
|
|
620
|
+
ctx.wait_for_exit(backend_pid, timeout: 10)
|
|
613
621
|
end
|
|
614
622
|
ctx.store.remove_route(hostname)
|
|
615
623
|
FileUtils.rm_f(File.join(ctx.store.dir, "backend-#{hostname}.pid"))
|
data/lib/yamine/database.rb
CHANGED
|
@@ -0,0 +1,201 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module Yamine
|
|
4
|
+
# Stopping a spawned process AND everything it started.
|
|
5
|
+
#
|
|
6
|
+
# Every run-mode process is spawned as ["sh", "-c", cmd] (see
|
|
7
|
+
# BootCommand.collect_spawns), so the pid yamine owns is the SHELL and
|
|
8
|
+
# the app is the shell's child. Signalling that pid alone is not
|
|
9
|
+
# stopping the app: on linux the shell's child is outside the shell's
|
|
10
|
+
# own signal scope, so TERM to the shell leaves the app running,
|
|
11
|
+
# reparented to init. `yamine stop` reports "Stopped <host>" and the
|
|
12
|
+
# app goes on serving, the route is gone, and nothing holds a handle
|
|
13
|
+
# to the process anymore. macOS forwards the signal, which is why a
|
|
14
|
+
# linux-only defect stayed invisible to the unit suite.
|
|
15
|
+
#
|
|
16
|
+
# Two halves, and the second is only safe with the first:
|
|
17
|
+
# * spawn each process as its own group leader (pgroup: true), so
|
|
18
|
+
# the whole tree shares a group that dies together;
|
|
19
|
+
# * signal the GROUP (-pid), which reaches the shell and every
|
|
20
|
+
# descendant in one syscall.
|
|
21
|
+
#
|
|
22
|
+
# A negative pid means "the process group whose id is that number",
|
|
23
|
+
# not "this process and its children" — so it is only correct for a
|
|
24
|
+
# pid that actually LEADS a group, and the two ways to know that are
|
|
25
|
+
# not equally good.
|
|
26
|
+
#
|
|
27
|
+
# The kernel can be asked, and it is the only option for a pid that
|
|
28
|
+
# came from a file (`yamine stop` and `yamine worktree remove` run in
|
|
29
|
+
# a different process than the spawn, with nothing but a pid from a
|
|
30
|
+
# route entry or a sidecar). But that is a reading, and it has two
|
|
31
|
+
# ways of being wrong:
|
|
32
|
+
#
|
|
33
|
+
# * ESRCH. The pid is gone — yet a process group outlives its
|
|
34
|
+
# leader, and the members behind it are exactly what needs
|
|
35
|
+
# stopping: a `sh -c` shell that died on its own leaves the app
|
|
36
|
+
# running in its group with nothing left to signal it by pid. A
|
|
37
|
+
# dead leader answers ESRCH, which reads exactly like "leads no
|
|
38
|
+
# group".
|
|
39
|
+
#
|
|
40
|
+
# * "Same group as me". A pid that never led a group answers this
|
|
41
|
+
# (a directly-spawned puma), and so does one of ours that has not
|
|
42
|
+
# run its setpgid yet — `pgroup: true` puts that in the child, so
|
|
43
|
+
# whether the parent can look first is a property of the spawn
|
|
44
|
+
# path, not something to bet a stop on.
|
|
45
|
+
#
|
|
46
|
+
# So a process spawned through `ProcessTree.spawn` is recorded as a
|
|
47
|
+
# leader, and the record is a fact about what we asked for rather than
|
|
48
|
+
# a reading of the moment. Everything else is the kernel's answer, and
|
|
49
|
+
# a pid number on its own is never evidence: 1234 may be a live
|
|
50
|
+
# process that inherited its group, and `kill(-1234)` would then hit
|
|
51
|
+
# whatever unrelated group wears that id.
|
|
52
|
+
module ProcessTree
|
|
53
|
+
# Pids we spawned as group leaders. A plain Hash, deliberately
|
|
54
|
+
# unlocked: every operation on it is a single call the GVL makes
|
|
55
|
+
# atomic, and yamine stops its processes from inside a trap handler
|
|
56
|
+
# (BootCommand.trap_cleanup), where Mutex#synchronize raises
|
|
57
|
+
# "can't be called from trap context" — a stop path that only works
|
|
58
|
+
# outside a trap is a stop that never happens on Ctrl-C.
|
|
59
|
+
PGROUP_LEADERS = {}
|
|
60
|
+
|
|
61
|
+
module_function
|
|
62
|
+
|
|
63
|
+
# Spawn a boot process as its own group leader and record that it is
|
|
64
|
+
# one. This is the only place a boot process is created, so "every
|
|
65
|
+
# process we boot leads a group" is one fact in one place rather
|
|
66
|
+
# than a convention spread across spawns.
|
|
67
|
+
def spawn(*args, **opts)
|
|
68
|
+
pid = ::Process.spawn(*args, pgroup: true, **opts)
|
|
69
|
+
note_group_leader(pid)
|
|
70
|
+
pid
|
|
71
|
+
end
|
|
72
|
+
|
|
73
|
+
# A pid we spawned with pgroup: true leads its own group — that is
|
|
74
|
+
# what we asked for, so the kernel is not consulted. Two readings it
|
|
75
|
+
# could give are both wrong here: "still in my group" while the
|
|
76
|
+
# child has not run its setpgid, and ESRCH once it is gone while
|
|
77
|
+
# the members it left behind are still running.
|
|
78
|
+
def note_group_leader(pid)
|
|
79
|
+
return unless pid.to_i.positive?
|
|
80
|
+
|
|
81
|
+
PGROUP_LEADERS[pid] = true
|
|
82
|
+
end
|
|
83
|
+
|
|
84
|
+
def known_group_leader?(pid)
|
|
85
|
+
PGROUP_LEADERS.key?(pid)
|
|
86
|
+
end
|
|
87
|
+
|
|
88
|
+
def forget(pid)
|
|
89
|
+
PGROUP_LEADERS.delete(pid)
|
|
90
|
+
end
|
|
91
|
+
|
|
92
|
+
# Does this pid lead its own process group? Ours, if we spawned it.
|
|
93
|
+
# Otherwise the kernel — and its "no" is the answer we want: a dead
|
|
94
|
+
# pid and a process that inherited its group are exactly the pids
|
|
95
|
+
# that must be signalled on their own, never by group id.
|
|
96
|
+
def group_leader?(pid)
|
|
97
|
+
return true if known_group_leader?(pid)
|
|
98
|
+
|
|
99
|
+
Process.getpgid(pid) == pid
|
|
100
|
+
rescue SystemCallError
|
|
101
|
+
false
|
|
102
|
+
end
|
|
103
|
+
|
|
104
|
+
# TERM a spawned process and its whole tree. Returns true when the
|
|
105
|
+
# signal was delivered, false when there was nothing left to signal
|
|
106
|
+
# (already dead, or a pid we may not touch) — which is the end
|
|
107
|
+
# state every caller wants, not an error, so this never raises.
|
|
108
|
+
def term(pid, signal: "TERM")
|
|
109
|
+
return false unless pid.to_i.positive?
|
|
110
|
+
|
|
111
|
+
group_signal(pid, signal) || pid_signal(pid, signal)
|
|
112
|
+
ensure
|
|
113
|
+
# The record is about a spawn, not about a pid that must keep
|
|
114
|
+
# this meaning: dropping it here keeps a pid that later gets
|
|
115
|
+
# recycled from being signalled as a group it never led.
|
|
116
|
+
forget(pid) if pid.to_i.positive?
|
|
117
|
+
end
|
|
118
|
+
|
|
119
|
+
# The whole group, one syscall. Best-effort: returns false when
|
|
120
|
+
# there is no such group left, so the caller can fall back.
|
|
121
|
+
def group_signal(pid, signal)
|
|
122
|
+
return false unless group_leader?(pid)
|
|
123
|
+
|
|
124
|
+
Process.kill(signal, -pid)
|
|
125
|
+
true
|
|
126
|
+
rescue SystemCallError
|
|
127
|
+
# ESRCH: the group outlived neither the leader nor its members.
|
|
128
|
+
# EPERM: the group exists but is not ours to signal. Either way a
|
|
129
|
+
# single-pid signal is the last thing left to try.
|
|
130
|
+
false
|
|
131
|
+
end
|
|
132
|
+
|
|
133
|
+
def pid_signal(pid, signal)
|
|
134
|
+
Process.kill(signal, pid)
|
|
135
|
+
true
|
|
136
|
+
rescue SystemCallError
|
|
137
|
+
false
|
|
138
|
+
end
|
|
139
|
+
|
|
140
|
+
# Stop a tree and make sure it is GONE — for the callers that go on
|
|
141
|
+
# to delete the directory it runs in, drop the route, or report
|
|
142
|
+
# "Stopped <host>". `term` alone is not that promise: it hands the
|
|
143
|
+
# signal to the group and returns, and the kernel's sweep covers
|
|
144
|
+
# the members that exist at that instant. A `sh -c` shell that is
|
|
145
|
+
# still assembling its tree forks its app *after* the sweep, and
|
|
146
|
+
# that app never sees the signal at all.
|
|
147
|
+
#
|
|
148
|
+
# Measured on linux: TERM the group of a `sh -c "ruby <app>"` a
|
|
149
|
+
# few milliseconds after the spawn and the shell dies while the app
|
|
150
|
+
# comes up behind it, reparented to init and serving forever —
|
|
151
|
+
# a live process whose working directory is on the way out, which
|
|
152
|
+
# is the exact shape of the bug this module exists to kill. A stop
|
|
153
|
+
# that only asks cannot rule that out, so this asks and then
|
|
154
|
+
# insists: TERM, give the group a moment to leave on its own
|
|
155
|
+
# terms, KILL whatever is still in it.
|
|
156
|
+
#
|
|
157
|
+
# The liveness question is asked of the GROUP, not of the pid: a
|
|
158
|
+
# dead leader that is a zombie still answers `kill(0, pid)`, and a
|
|
159
|
+
# live tree behind one does not answer to its pid at all. The group
|
|
160
|
+
# is the only handle that tells the two apart.
|
|
161
|
+
#
|
|
162
|
+
# `term` is deliberately left alone: it is the primitive a signal
|
|
163
|
+
# trap handler uses (BootCommand.trap_cleanup), and a trap handler
|
|
164
|
+
# must not sit in a sleep loop.
|
|
165
|
+
def terminate(pid, grace: 5)
|
|
166
|
+
return false unless term(pid)
|
|
167
|
+
|
|
168
|
+
deadline = monotonic + grace
|
|
169
|
+
sleep 0.05 while group?(pid) && monotonic < deadline
|
|
170
|
+
return true unless group?(pid)
|
|
171
|
+
|
|
172
|
+
# A group that outlived its grace period is not draining, it is
|
|
173
|
+
# stuck or ignoring us, and the caller is about to delete what it
|
|
174
|
+
# runs in. No group_leader? check here: the leader is very likely
|
|
175
|
+
# already gone, and it is the members behind it that need the
|
|
176
|
+
# signal.
|
|
177
|
+
Process.kill("KILL", -pid)
|
|
178
|
+
true
|
|
179
|
+
rescue SystemCallError
|
|
180
|
+
# The group went away between the last look and this signal.
|
|
181
|
+
true
|
|
182
|
+
end
|
|
183
|
+
|
|
184
|
+
# Does this process group still have a member in it? Signal 0 asks
|
|
185
|
+
# the kernel about the group rather than about a pid, so it stays
|
|
186
|
+
# true while anything behind a dead leader is still running, and
|
|
187
|
+
# goes false only once the tree really is down.
|
|
188
|
+
def group?(pid)
|
|
189
|
+
return false unless pid.to_i.positive?
|
|
190
|
+
|
|
191
|
+
Process.kill(0, -pid)
|
|
192
|
+
true
|
|
193
|
+
rescue SystemCallError
|
|
194
|
+
false
|
|
195
|
+
end
|
|
196
|
+
|
|
197
|
+
def monotonic
|
|
198
|
+
Process.clock_gettime(Process::CLOCK_MONOTONIC)
|
|
199
|
+
end
|
|
200
|
+
end
|
|
201
|
+
end
|
data/lib/yamine/proxy.rb
CHANGED
|
@@ -566,25 +566,51 @@ module Yamine
|
|
|
566
566
|
# the other. Deliberately outside the idle bound: an idle websocket
|
|
567
567
|
# is legitimate traffic (silence is the normal state), and framing
|
|
568
568
|
# no longer applies once the connection is hijacked.
|
|
569
|
+
#
|
|
570
|
+
# Whichever direction reaches EOF first tears down BOTH ends, and it
|
|
571
|
+
# has to do that with shutdown before close. The sibling copy is
|
|
572
|
+
# parked in a blocking read on the socket being torn down, and a
|
|
573
|
+
# close from another thread does not interrupt that read on linux (it
|
|
574
|
+
# does on macOS) — so close-only left the parked copy alive forever
|
|
575
|
+
# and `t1.join` never returned: one leaked thread per closed upgrade,
|
|
576
|
+
# plus a `handle` that never unwound, in the root daemon. shutdown(2)
|
|
577
|
+
# changes the socket's kernel state instead of dropping the fd, so
|
|
578
|
+
# the parked read comes back with EOF at once — deterministic, and
|
|
579
|
+
# with no timeout added to a path that must stay unbounded.
|
|
569
580
|
def pipe_both(client, backend)
|
|
570
581
|
t1 = Thread.new do
|
|
571
582
|
IO.copy_stream(client, backend)
|
|
572
583
|
rescue IOError, SystemCallError
|
|
573
584
|
nil
|
|
574
585
|
ensure
|
|
575
|
-
backend
|
|
586
|
+
teardown_half(backend)
|
|
576
587
|
end
|
|
577
588
|
t2 = Thread.new do
|
|
578
589
|
IO.copy_stream(backend, client)
|
|
579
590
|
rescue IOError, SystemCallError
|
|
580
591
|
nil
|
|
581
592
|
ensure
|
|
582
|
-
client
|
|
593
|
+
teardown_half(client)
|
|
583
594
|
end
|
|
584
595
|
t1.join
|
|
585
596
|
t2.join
|
|
586
597
|
end
|
|
587
598
|
|
|
599
|
+
# End one end of a hijacked connection: unblock the sibling copy
|
|
600
|
+
# first, then close. Runs from an ensure, so it swallows everything
|
|
601
|
+
# — an exception escaping here would unwind into the acceptor.
|
|
602
|
+
def teardown_half(sock)
|
|
603
|
+
# An SSLSocket (the TLS client's socket) has no #shutdown; its
|
|
604
|
+
# kernel socket is one level down, and that is the fd the parked
|
|
605
|
+
# read is sitting on.
|
|
606
|
+
io = sock.respond_to?(:shutdown) ? sock : sock.to_io
|
|
607
|
+
io.shutdown(Socket::SHUT_RDWR)
|
|
608
|
+
rescue StandardError
|
|
609
|
+
nil
|
|
610
|
+
ensure
|
|
611
|
+
sock.close rescue nil
|
|
612
|
+
end
|
|
613
|
+
|
|
588
614
|
# DNS-rebinding boundary: the proxy binds loopback, so any website
|
|
589
615
|
# that tricks a browser into requesting 127.0.0.1 reaches us. A
|
|
590
616
|
# foreign Host (outside our TLDs) gets a bare 404 that names nothing —
|
data/lib/yamine/runner.rb
CHANGED
|
@@ -96,6 +96,12 @@ module Yamine
|
|
|
96
96
|
# view yamine's failure payloads use. Interleaving matches the
|
|
97
97
|
# foreman/without-yamine experience.
|
|
98
98
|
#
|
|
99
|
+
# Spawned through ProcessTree, not Kernel#spawn: the command is an
|
|
100
|
+
# argv that starts with `sh -c` (collect_spawns), so the pid we get
|
|
101
|
+
# back is the shell. Its own process group is what makes the process
|
|
102
|
+
# behind it stoppable — stop signals the group (ProcessTree.term),
|
|
103
|
+
# and the shell cannot leave its child outside the group it leads.
|
|
104
|
+
#
|
|
99
105
|
# extra_env is the config's env.clear/env.secret for this process —
|
|
100
106
|
# yamine's own keys (PORT, YAMINE_URL, DATABASE_URL, the TLS vars)
|
|
101
107
|
# win over it, because those describe the boot rather than the app.
|
|
@@ -106,7 +112,7 @@ module Yamine
|
|
|
106
112
|
env = child_env(dir, url: url, port: port, rails_dev_host: rails_dev_host,
|
|
107
113
|
database_url: database_url, database_env: database_env, extra_env: extra_env)
|
|
108
114
|
path = log_path(dir, name)
|
|
109
|
-
pid = with_clean_env { spawn(env, *command, chdir: dir, out: path, err: [:child, :out]) }
|
|
115
|
+
pid = with_clean_env { ProcessTree.spawn(env, *command, chdir: dir, out: path, err: [:child, :out]) }
|
|
110
116
|
Process.detach(pid)
|
|
111
117
|
target = "127.0.0.1:#{port}"
|
|
112
118
|
if register
|
|
@@ -130,12 +136,16 @@ module Yamine
|
|
|
130
136
|
# whole tree up concurrently, then polls every route until healthy.
|
|
131
137
|
# Returns the placeholder App (target known before bind). Fate of
|
|
132
138
|
# the backend is decided by wait, not by spawn.
|
|
139
|
+
#
|
|
140
|
+
# Through ProcessTree, for the same reason as boot_run: a `sh -c`
|
|
141
|
+
# pid is not the process doing the work, so its group is the only
|
|
142
|
+
# handle that reaches it.
|
|
133
143
|
def spawn_http(name:, hostname:, url:, dir:, command:, port:, rails_dev_host: nil,
|
|
134
144
|
database_url: nil, database_env: nil, force: false, extra_env: nil)
|
|
135
145
|
env = child_env(dir, url: url, port: port, rails_dev_host: rails_dev_host,
|
|
136
146
|
database_url: database_url, database_env: database_env, extra_env: extra_env)
|
|
137
147
|
path = log_path(dir, name)
|
|
138
|
-
pid = with_clean_env { spawn(env, *command, chdir: dir, out: path, err: [:child, :out]) }
|
|
148
|
+
pid = with_clean_env { ProcessTree.spawn(env, *command, chdir: dir, out: path, err: [:child, :out]) }
|
|
139
149
|
Process.detach(pid)
|
|
140
150
|
App.new(name: name, hostname: hostname, url: url, pid: pid,
|
|
141
151
|
target: "127.0.0.1:#{port}", kind: "tcp", command: command)
|
|
@@ -340,10 +350,14 @@ module Yamine
|
|
|
340
350
|
|
|
341
351
|
private
|
|
342
352
|
|
|
353
|
+
# A backend we spawned is stopped by its whole group: the pid we
|
|
354
|
+
# hold for a run-mode process is the `sh -c` shell, and on linux the
|
|
355
|
+
# app behind it does not see a signal aimed at that shell alone.
|
|
356
|
+
# A pid that leads no group (the socket spawns, a pid already gone)
|
|
357
|
+
# falls back to the single-pid signal. Never raises: a stopped
|
|
358
|
+
# backend is the goal, and "it was already dead" is a success here.
|
|
343
359
|
def stop_pid(pid)
|
|
344
|
-
|
|
345
|
-
rescue SystemCallError
|
|
346
|
-
nil
|
|
360
|
+
ProcessTree.term(pid)
|
|
347
361
|
end
|
|
348
362
|
end
|
|
349
363
|
end
|
data/lib/yamine/version.rb
CHANGED
data/lib/yamine.rb
CHANGED
|
@@ -19,6 +19,7 @@ require_relative "yamine/log"
|
|
|
19
19
|
require_relative "yamine/route_store"
|
|
20
20
|
require_relative "yamine/certs"
|
|
21
21
|
require_relative "yamine/ports"
|
|
22
|
+
require_relative "yamine/process_tree"
|
|
22
23
|
require_relative "yamine/hosts"
|
|
23
24
|
require_relative "yamine/proxy"
|
|
24
25
|
require_relative "yamine/proxy_control"
|
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: yamine
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.21.
|
|
4
|
+
version: 0.21.2
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Kaka Ruto
|
|
@@ -105,6 +105,7 @@ files:
|
|
|
105
105
|
- lib/yamine/ports.rb
|
|
106
106
|
- lib/yamine/privileged_payload.rb
|
|
107
107
|
- lib/yamine/probe.rb
|
|
108
|
+
- lib/yamine/process_tree.rb
|
|
108
109
|
- lib/yamine/procfile.rb
|
|
109
110
|
- lib/yamine/proxy.rb
|
|
110
111
|
- lib/yamine/proxy_control.rb
|