cable_room 0.7.0.beta3 → 0.8.0.beta2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -4,9 +4,9 @@ module CableRoom
4
4
  class Host
5
5
  # The parent process behind `cable_room server --workers N`. It forks N children, runs one
6
6
  # block in each, replaces a child that dies, and passes SIGTERM and SIGINT on to every child
7
- # before exiting itself. It never hosts a room, never opens Redis, and never starts a Bus:
8
- # each child builds its own after the fork (`CableRoom.after_fork!`), so nothing with a
9
- # socket or a thread behind it is ever shared between processes.
7
+ # before exiting itself. It never hosts a room, never subscribes the Bus, and never starts a
8
+ # Host thread: each child builds its own after the fork (`CableRoom.after_fork!`), so nothing
9
+ # with a socket or a thread behind it is ever shared between processes.
10
10
  #
11
11
  # Supervisor.new(count: 2, logger: logger) { |index| run_a_host(index) }.run
12
12
  #
@@ -18,16 +18,21 @@ module CableRoom
18
18
  #
19
19
  # Restarts: a child that dies is replaced. One that ran for at least `stable_after` seconds
20
20
  # comes back at once; one that died sooner is treated as failing to boot and comes back
21
- # after a delay that doubles from `first_backoff` up to `max_backoff`, so a broken app forks
22
- # a few times a minute rather than thousands of times a second.
21
+ # after a delay that doubles from `first_backoff` up to `max_backoff`. The delay only resets
22
+ # once a replacement stays up `stable_after` seconds, so the fastest any slot can cycle is
23
+ # once per `stable_after` seconds, and a broken app settles at one fork per slot every
24
+ # `max_backoff` seconds rather than thousands a second.
23
25
  #
24
- # Signals: SIGTERM and SIGINT to the parent are relayed to every live child and the parent
25
- # then waits, with no time limit of its own, until they've all exited (a child migrating rooms
26
- # may take a while; the container's `timeout -k` is the outer bound). A signal sent to one
27
- # child touches only that child: it exits, and the parent replaces it. The handlers do
28
- # nothing but write a byte to a pipe; the main loop reads the pipe, so no real work runs in
29
- # signal context. SIGCHLD writes the same pipe to wake the loop when a child exits, and the
30
- # loop also polls every `POLL_INTERVAL` seconds in case a wake-up is ever missed.
26
+ # Signals: SIGTERM and SIGINT to the parent are relayed to every live child, and the parent
27
+ # waits up to `shutdown_timeout` seconds for them to exit, then SIGKILLs whatever is left.
28
+ # A second SIGTERM or SIGINT while waiting doesn't wait any longer: it SIGKILLs the children
29
+ # at once. Either way the parent exits 0 once they're all gone. A signal sent to one child
30
+ # touches only that child: it exits, and the parent replaces it.
31
+ #
32
+ # The handlers do nothing but write a line to a pipe; the main loop reads the pipe, so no
33
+ # real work runs in signal context. SIGCHLD writes the same pipe to wake the loop when a
34
+ # child exits, and the loop also polls every `POLL_INTERVAL` seconds in case a wake-up is
35
+ # ever missed.
31
36
  #
32
37
  # If the parent itself dies without relaying anything (SIGKILL, a crash), each child notices
33
38
  # through a second pipe (`watch_parent`) and stops as if it had been sent SIGTERM, so a dead
@@ -40,7 +45,7 @@ module CableRoom
40
45
 
41
46
  attr_reader :count, :logger
42
47
 
43
- def initialize(count:, logger:, stable_after: 5, first_backoff: 1, max_backoff: 30, &body)
48
+ def initialize(count:, logger:, stable_after: 5, first_backoff: 1, max_backoff: 30, shutdown_timeout: 25, &body)
44
49
  raise ArgumentError, "count must be at least 1 (got #{count.inspect})" unless count.is_a?(Integer) && count >= 1
45
50
  raise ArgumentError, "Supervisor needs a block to run in each worker" unless body
46
51
 
@@ -49,12 +54,14 @@ module CableRoom
49
54
  @stable_after = stable_after
50
55
  @first_backoff = first_backoff
51
56
  @max_backoff = max_backoff
57
+ @shutdown_timeout = shutdown_timeout
52
58
  @body = body
53
59
 
54
60
  @workers = {} # index => Worker, for every live child
55
61
  @restart_at = {} # index => monotonic time to fork a replacement
56
62
  @backoff = {} # index => the delay used for that slot's last boot failure
57
63
  @stopping = false
64
+ @kill_at = nil # monotonic time to SIGKILL children that haven't exited
58
65
  end
59
66
 
60
67
  # Fork the workers and supervise them until a stop signal has arrived and every child has
@@ -70,7 +77,12 @@ module CableRoom
70
77
  wait_for_wakeup
71
78
  reap_exited_workers
72
79
  break if @stopping && @workers.empty?
73
- start_due_restarts unless @stopping
80
+
81
+ if @stopping
82
+ kill_stragglers if now >= @kill_at
83
+ else
84
+ start_due_restarts
85
+ end
74
86
  end
75
87
 
76
88
  logger.info "all workers exited"
@@ -117,10 +129,10 @@ module CableRoom
117
129
  nil
118
130
  end
119
131
 
120
- # Sleep until a signal or child exit wakes us, a restart is due, or the poll interval
121
- # passes. Any stop signal on the pipe starts the shutdown.
132
+ # Sleep until a signal or child exit wakes us, a restart or the kill deadline is due, or
133
+ # the poll interval passes. Any stop signal on the pipe starts (or escalates) the shutdown.
122
134
  def wait_for_wakeup
123
- timeout = [POLL_INTERVAL, seconds_until_next_restart].compact.min
135
+ timeout = [POLL_INTERVAL, seconds_until_next_restart, seconds_until_kill].compact.min
124
136
  ready, = IO.select([@signal_reader], nil, nil, timeout)
125
137
  return unless ready
126
138
 
@@ -134,17 +146,27 @@ module CableRoom
134
146
 
135
147
  def begin_stopping(signal)
136
148
  if @stopping
137
- # A second Ctrl-C or TERM: pass it along again, in case a child missed the first
138
- relay(signal)
149
+ # A second Ctrl-C or TERM: whoever sent it doesn't want to wait for a graceful stop
150
+ logger.warn "got SIG#{signal} again, killing #{@workers.size} worker(s) now"
151
+ relay("KILL")
152
+ @kill_at = Float::INFINITY
139
153
  return
140
154
  end
141
155
 
142
156
  @stopping = true
143
157
  @restart_at.clear
144
- logger.info "got SIG#{signal}, relaying to #{@workers.size} worker(s) and waiting for them"
158
+ @kill_at = now + @shutdown_timeout
159
+ logger.info "got SIG#{signal}, relaying to #{@workers.size} worker(s) and waiting up to #{@shutdown_timeout}s for them"
145
160
  relay(signal)
146
161
  end
147
162
 
163
+ # Past the shutdown deadline: kill what's left, once (the next reap collects them)
164
+ def kill_stragglers
165
+ logger.warn "#{@workers.size} worker(s) still running after #{@shutdown_timeout}s, killing them"
166
+ relay("KILL")
167
+ @kill_at = Float::INFINITY
168
+ end
169
+
148
170
  def relay(signal)
149
171
  @workers.each_value do |worker|
150
172
  Process.kill(signal, worker.pid)
@@ -209,6 +231,10 @@ module CableRoom
209
231
  next_at && [next_at - now, 0].max
210
232
  end
211
233
 
234
+ def seconds_until_kill
235
+ @kill_at && [@kill_at - now, 0].max
236
+ end
237
+
212
238
  def start_worker(index)
213
239
  pid = fork_worker(index)
214
240
  @workers[index] = Worker.new(index, pid, now)
@@ -259,8 +285,8 @@ module CableRoom
259
285
 
260
286
  # Once every child has closed its copy, the parent holds the only write end of the lifeline
261
287
  # pipe, so a read here returns EOF exactly when the parent is gone: crashed, or SIGKILLed by
262
- # the container's `timeout -k`. A worker must not carry on hosting rooms with nobody
263
- # supervising it, so it stops itself the same way a relayed SIGTERM would stop it.
288
+ # the container. A worker must not carry on hosting rooms with nobody supervising it, so it
289
+ # stops itself the same way a relayed SIGTERM would stop it.
264
290
  def watch_parent(reader)
265
291
  thread = Thread.new do
266
292
  reader.read(1)
@@ -1,50 +1,37 @@
1
- require 'concurrent'
2
-
3
1
  module CableRoom
4
2
  class Host
5
- # The thread pool every room's work runs on. Deliberately *not* an ActionCable::Server::Worker
6
- # subclass: that would put Room work on the same shared `:work` callback chain as ordinary
7
- # ActionCable connections, and that chain is leaky by construction -- `ActiveSupport::Callbacks`
8
- # re-injects a callback added to the base class into every existing descendant regardless of
9
- # load order, so an app's own tenant-switching hook (PandaPal's, say, registered from a Rails
10
- # initializer well after this class is defined) ends up wrapping Room work too, whether we want
11
- # it to or not. Owning a distinct namespace here avoids that category of problem entirely:
12
- # nothing outside cable_room can attach to Room work by surprise. The one seam an app gets is
13
- # a Room's own `before_work`/`after_work`/`around_work` (see Room::Callbacks and
14
- # Host::Runner#with_room_context).
15
- class WorkerPool
16
- attr_reader :executor
17
-
18
- def initialize(max_size: 5)
19
- @executor = Concurrent::ThreadPoolExecutor.new(
20
- name: "CableRoom",
21
- min_threads: 1,
22
- max_threads: max_size,
23
- max_queue: 0,
24
- )
3
+ # The thread pool every room's work runs on. It's an ActionCable Worker so the `:work`
4
+ # callbacks Rails installs (the executor wrap, ActiveRecord log tagging) still apply to room
5
+ # work exactly as they did when rooms were channels.
6
+ #
7
+ # The "connection" passed around here is the room's Host::Runner. ActionCable's Worker was
8
+ # written for connections; rooms don't have one, but the runner fills the same role: it's
9
+ # the thing with a logger and an error reporter.
10
+ class WorkerPool < ActionCable::Server::Worker
11
+ set_callback :work, :around do |_, blk|
12
+ pconn = ActionCable::Server::Worker.connection
13
+ ActionCable::Server::Worker.connection = connection
14
+ blk.call
15
+ ensure
16
+ ActionCable::Server::Worker.connection = pconn
25
17
  end
26
18
 
27
- # Reduces every exception to a proper report instead of ActionCable's Worker#invoke, which
28
- # logs a line and calls a no-argument `handle_exception`, discarding the error itself. Rooms
29
- # run all of their work through here, so report it properly. `connection:` is always a
30
- # Host::Runner in practice (see Runner#post_work); the else branch is a defensive fallback.
19
+ # ActionCable's Worker#invoke reduces every exception to a log line and a no-argument
20
+ # `handle_exception` call, which discards the error itself. Rooms run all of their work
21
+ # through here, so report it properly instead.
31
22
  def invoke(receiver, method, *args, connection:, &block)
32
- receiver.send method, *args, &block
33
- rescue Exception => e
34
- if connection.respond_to?(:report_work_error)
35
- connection.report_work_error(e)
36
- else
37
- logger.error "There was an exception - #{e.class}(#{e.message})"
38
- logger.error Array(e.backtrace).join("\n")
39
- CableRoom.report_error(e, connection: connection)
23
+ work(connection) do
24
+ receiver.send method, *args, &block
25
+ rescue Exception => e
26
+ if connection.respond_to?(:report_work_error)
27
+ connection.report_work_error(e)
28
+ else
29
+ logger.error "There was an exception - #{e.class}(#{e.message})"
30
+ logger.error Array(e.backtrace).join("\n")
31
+ CableRoom.report_error(e, connection: connection)
32
+ end
40
33
  end
41
34
  end
42
-
43
- private
44
-
45
- def logger
46
- ActionCable.server.logger
47
- end
48
35
  end
49
36
  end
50
37
  end