yamine 0.10.2 → 0.10.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 06f020588bc63954dd44b5a99b598995ec3e6f562c03756b867d26ff44178af8
4
- data.tar.gz: 97f1613f7ad01d1362e4098651048addd3a4a738912c6bd83427e3ea9453ad8d
3
+ metadata.gz: 073eb6f5386c8500d62348d2093ebc11fa13dc291f6907d11410bcb1b81e09eb
4
+ data.tar.gz: 9750f59d2dd9cce11832d232d3d1cb11248b31e52545133ab2b998d21bdf9549
5
5
  SHA512:
6
- metadata.gz: 3f03809157db470e905542751e16d8a6cb6edd6b91a78637f20e85fb5670e6b30e15f717d1bbaed714c70b077a756fe97744bfb27521304606b2d661dea57eab
7
- data.tar.gz: 83bb42d5eb49f6ee4007d04fa3e1b13024a46044d1afc7756a126f43384f60cb9c6ca57d7440d77ffcbacd37b9d95b04dbf8eb933bbf1676a10440e4951bdba7
6
+ metadata.gz: a0cb84ba9b5fccda18721eb0de3bbb05667cffe97caf8882a3d3217d8dbee339319af2c72090f3b5bee41df198b01cc345e2dcb0e74b8ca9af202f2cd688e7ee
7
+ data.tar.gz: 89f533e131495a41e6e1613e09017b5e3b4d3a834bbaa5507bf2d13bdcb16054fb9320f36585d9b858d4f34c96734b3a5285fe04a0a8e01fac8227689573dfb4
data/CHANGELOG.md CHANGED
@@ -1,5 +1,77 @@
1
1
  # Changelog
2
2
 
3
+ ## [0.10.4] — 2026-09-11
4
+
5
+ ### Fixed
6
+
7
+ - **A stale `tmp/pids/server.pid` no longer kills the boot 2s in.** The
8
+ anywaye incident: an interrupted earlier run left puma alive, Rails
9
+ refused to start ("A server is already running (pid: N)"), `web` died
10
+ before its first healthcheck, and the failure line only said "process
11
+ exited" — the actual reason sat in the log file. Now:
12
+ - Pre-flight `Readiness.server_pid_conflict` checks the pidfile BEFORE
13
+ spawning. A live holder is a conflict, not a mystery crash; a dead
14
+ pid is fine (Rails overwrites it).
15
+ - When the holder is this app's own leftover puma (`puma ... [app]`),
16
+ yamine reaps it automatically. Anything else fails with the exact
17
+ fix (`kill N && rm tmp/pids/server.pid`), or `--force` to take over.
18
+ - `Readiness.fatal_line` recognizes the common killers (already
19
+ running, address in use, missing gems, missing database) in the
20
+ process's own log and puts the real reason in the failure line.
21
+ - **No more orphans on any failure path.** The incident left puma +
22
+ solid-queue + ask-app-server running for 24 minutes with no routes
23
+ after the CLI exited. Now `Runner#boot_run` reaps the backend when
24
+ registration is refused (quota/conflict), `boot_concurrent` tracks
25
+ every spawned HTTP backend in `children` before adopt, and `boot_all`
26
+ wraps spawn+boot so a raise mid-boot kills every child and removes
27
+ every route already registered. Tests pin all three (with `expects`
28
+ so they cannot pass vacuously).
29
+
30
+ ## [0.10.3] — 2026-09-11
31
+
32
+ ### Added
33
+
34
+ - **`yamine start` waits by default** (was opt-in `--wait`). Every HTTP
35
+ process is spawned concurrently, each `healthcheck: { path: /up,
36
+ timeout: 30 }` or TCP accept is polled, and routes are registered only
37
+ when all are healthy (`ready:` + `WaitPayload` with `--json`). On
38
+ failure everything spawned (including background children) is killed and
39
+ the CLI exits 1 with the failed process's own `log/yamine-<name>.log`
40
+ tail and no half-booted routes. `--wait` is kept as a no-op alias;
41
+ `--no-wait` restores the old sequential fire-and-forget path.
42
+ - `Yamine::Readiness` phased events (`phase`/`action`/`status`/
43
+ `duration_ms`/`detail`), `check_deps` pre-flight, `wait_healthy`
44
+ (healthcheck path or TCP accept), concurrent `wait_all` (one thread per
45
+ process, dead-PID short-circuit), per-phase timeouts (deps 30s, db 15s,
46
+ schema 90s, process 45s or per-healthcheck). `Yamine::WaitPayload`
47
+ success/failure payloads. `Runner#spawn_http`/`adopt`/`spawned?`, `boot_run`
48
+ output is logged and `Database.exists?` splits from `ensure_exists` for
49
+ provenance without a second probe.
50
+ - Top-level `db: false` opts out of per-worktree databases (validated
51
+ boolean-or-mapping, `x-` extensions still ignored), `db.schema_load`
52
+ moves to top-level `db`. Template discovery checks `ENV` then top-level
53
+ `env.clear` then per-process `env.clear`.
54
+
55
+ ### Fixed
56
+
57
+ - **`yamine --force` (and bare `yamine --wait`/`--no-wait`) no longer hits
58
+ the "no longer supported" error.** Boot-flag dispatch now routes bare
59
+ `yamine --flag` to `run_inferred` instead of the legacy single-name
60
+ path, and `--help`/`--version` are handled before boot. `run_named`
61
+ now guides `unknown flag \`--flag\`` toward `--help`.
62
+ - **`yamine` no longer wastes a 2s health wait before reporting the
63
+ ownership conflict.** The conflict you hit (`"anywaye.localhost" is
64
+ already registered by a running agent "kaka@...` + stale background
65
+ jobs already spawned) now fails before any spawn, `prune_stale` discards
66
+ crashed state, and the message distinguishes same-agent-different-dir
67
+ ("stop the other instance") from foreign agent ("work in your own
68
+ worktree"). Guidance now says `yamine start --force (or bare \`yamine
69
+ --force\`)`.
70
+ - **Credentials-only `DATABASE_URL`s no longer share silently.** When no
71
+ template exists but `database.yml`/`Gemfile` proves server-backed, boot
72
+ warns loudly with the `env.clear` + `local.secrets` fix. Injection is
73
+ still `DATABASE_URL`, no `ask-auth` dep, no Rails load at boot.
74
+
3
75
  ## [0.10.2] — 2026-09-10
4
76
 
5
77
  ### Added
@@ -46,8 +46,12 @@ module Yamine
46
46
  end
47
47
 
48
48
  def run_named(ctx, name, _args)
49
- $stderr.puts "Error: `yamine #{name}` is no longer supported."
50
- $stderr.puts " All processes come from config/local.yml. Run `yamine init` to create one."
49
+ if name.to_s.start_with?("-")
50
+ $stderr.puts "Error: unknown flag `#{name}`. Try `yamine --help` or `yamine start --help`."
51
+ else
52
+ $stderr.puts "Error: `yamine #{name}` is no longer supported."
53
+ $stderr.puts " All processes come from config/local.yml. Run `yamine init` to create one."
54
+ end
51
55
  exit 1
52
56
  end
53
57
 
@@ -86,6 +90,28 @@ module Yamine
86
90
  exit 1
87
91
  end
88
92
 
93
+ # Pre-flight: a live tmp/pids/server.pid makes Rails refuse to
94
+ # boot ("A server is already running") the moment web starts.
95
+ # When the holder is this app's own puma (`puma ... [app]`), take
96
+ # over automatically — that is one of ours. Anything else is not
97
+ # ours to kill: fail with the exact fix before spawning.
98
+ conflict_ok, conflict_detail, conflict_pid = Readiness.server_pid_conflict(Dir.pwd)
99
+ unless conflict_ok
100
+ if own_orphan_puma?(conflict_pid, resolved.app, Dir.pwd)
101
+ say opts, " reaping own orphaned server (pid #{conflict_pid})"
102
+ reap_stale_server(conflict_pid)
103
+ elsif opts[:force]
104
+ say opts, " --force: taking over pid #{conflict_pid}"
105
+ reap_stale_server(conflict_pid)
106
+ else
107
+ $stderr.puts "Error: #{conflict_detail}."
108
+ $stderr.puts " Rails will refuse to boot while it runs. Free it first:"
109
+ $stderr.puts " kill #{conflict_pid} && rm tmp/pids/server.pid"
110
+ $stderr.puts " Or re-run with --force to let yamine take over that pid."
111
+ exit 1
112
+ end
113
+ end
114
+
89
115
  runner = Runner.new(store: ctx.store)
90
116
  children = []
91
117
  routes_registered = []
@@ -107,12 +133,21 @@ module Yamine
107
133
  with_wait = !opts[:no_wait]
108
134
  spawn_plan = collect_spawns(ctx, runner, resolved, opts, db_url, children)
109
135
 
110
- if with_wait
111
- boot_concurrent(ctx, runner, resolved, opts, spawn_plan, children,
112
- routes_registered, db_url, db_name, db_created, events: events)
113
- else
114
- boot_sequential(ctx, runner, resolved, opts, spawn_plan, children,
115
- routes_registered, db_url, events: events)
136
+ begin
137
+ if with_wait
138
+ boot_concurrent(ctx, runner, resolved, opts, spawn_plan, children,
139
+ routes_registered, db_url, db_name, db_created, events: events)
140
+ else
141
+ boot_sequential(ctx, runner, resolved, opts, spawn_plan, children,
142
+ routes_registered, db_url, events: events)
143
+ end
144
+ rescue StandardError
145
+ # A raise mid-boot (route conflict, quota, unexpected error)
146
+ # must never leave half a tree running: reap every child and
147
+ # every route already registered, then surface the error.
148
+ children.each { |c| stop_spawned_pid(c[:pid]) }
149
+ cleanup_routes(ctx, routes_registered.flat_map { |r| r[:hostnames] })
150
+ raise
116
151
  end
117
152
 
118
153
  background = processes.select { |_, v| v["proxy"] == false }
@@ -133,6 +168,35 @@ module Yamine
133
168
  supervise_tree(ctx, all_hostnames, named_pids, reporter: events)
134
169
  end
135
170
 
171
+ # TERM the old server and make sure the pidfile no longer names
172
+ # a live process before we spawn (wait a moment; then remove the
173
+ # file if the process is still refusing to die, so Rails can boot
174
+ # over it — the socket it held is already ours to take).
175
+ def reap_stale_server(pid)
176
+ Process.kill("TERM", pid)
177
+ deadline = Process.clock_gettime(Process::CLOCK_MONOTONIC) + 5
178
+ while Process.clock_gettime(Process::CLOCK_MONOTONIC) < deadline
179
+ break unless ProxyControl.pid_alive?(pid)
180
+ sleep 0.2
181
+ end
182
+ FileUtils.rm_f(File.join(Dir.pwd, "tmp", "pids", "server.pid"))
183
+ rescue SystemCallError
184
+ FileUtils.rm_f(File.join(Dir.pwd, "tmp", "pids", "server.pid"))
185
+ end
186
+
187
+ # Is this pid a leftover backend of THIS app? `puma (tcp://...)
188
+ # [app]` names the app; we also accept a command containing the
189
+ # app dir. Only then may we reap it unasked.
190
+ def own_orphan_puma?(pid, app, dir)
191
+ _out, status = Open3.capture2("ps", "-p", pid.to_s, "-o", "command=")
192
+ return false unless status.success?
193
+
194
+ cmd = _out.to_s
195
+ cmd.include?("[#{app}]") || cmd.include?(File.expand_path(dir))
196
+ rescue SystemCallError
197
+ false
198
+ end
199
+
136
200
  # One pass over config: background processes spawn immediately
137
201
  # (no route to wait on), HTTP processes become a spawn plan the
138
202
  # boot mode executes. Compound-line refusal happens here so both
@@ -194,12 +258,18 @@ module Yamine
194
258
  apps = {}
195
259
  plan.each do |item|
196
260
  say opts, " [#{item[:name]}] #{item[:url]}"
261
+ app = runner.spawn_http(name: item[:name], hostname: item[:hostname],
262
+ url: item[:url], dir: Dir.pwd, command: item[:command],
263
+ port: item[:port], rails_dev_host: item[:hostname],
264
+ database_url: db_url, force: opts[:force])
265
+ # Tracked before registration so any later failure (a raise
266
+ # during adopt, a crash while another process waits) can
267
+ # always reap it. Duplicate names in named_pids are harmless.
268
+ children << { name: item[:name], pid: app.pid }
197
269
  apps[item[:name]] = {
198
270
  item: item,
199
- app: runner.spawn_http(name: item[:name], hostname: item[:hostname],
200
- url: item[:url], dir: Dir.pwd, command: item[:command],
201
- port: item[:port], rails_dev_host: item[:hostname],
202
- database_url: db_url, force: opts[:force])
271
+ log_path: File.join(Dir.pwd, "log", "yamine-#{item[:name]}.log"),
272
+ app: app
203
273
  }
204
274
  end
205
275
  wait_result = Readiness.wait_all(apps, sink: events)
@@ -298,20 +368,32 @@ module Yamine
298
368
  # Compares our worktree dir (spec.dir) against the owner's: same
299
369
  # dir + same agent is a restart, anything else names the owner.
300
370
  def check_worktree_ownership!(ctx, resolved, force:)
371
+ return if force
372
+
301
373
  mine = Agent.name
302
374
  here = File.expand_path(Dir.pwd)
375
+ # Stale pids (agent died, worktree deleted) must not block a restart.
376
+ ctx.store.prune_stale
303
377
  Resolver.hostnames(resolved).each do |hostname|
304
378
  entry = ctx.store.find(hostname)
305
379
  next unless entry
306
380
  next unless entry["pid"] != 0 && ProxyControl.pid_alive?(entry["pid"])
307
381
  next if entry["agent"] == mine && entry.dig("spec", "dir") == here
308
- next if force
309
382
 
310
383
  owner = entry["agent"] && !entry["agent"].empty? ? "agent #{entry["agent"].inspect}" : "PID #{entry["pid"]}"
311
384
  dir = entry.dig("spec", "dir")
312
- $stderr.puts "Error: #{hostname} is live and owned by #{owner}#{dir ? " (#{dir})" : ""}."
313
- $stderr.puts " Work in your own worktree (each branch gets its own URL), or take over explicitly:"
314
- $stderr.puts " yamine start --force"
385
+ same_agent = entry["agent"] == mine
386
+ same_dir = entry.dig("spec", "dir") == here
387
+ if same_agent && !same_dir
388
+ $stderr.puts "Error: #{hostname} was started from a different directory by the same agent #{mine.inspect}:"
389
+ $stderr.puts " running: #{dir || "(unknown)"}"
390
+ $stderr.puts " current: #{here}"
391
+ $stderr.puts " Stop the other instance first (`yamine stop` there) or take over explicitly:"
392
+ else
393
+ $stderr.puts "Error: #{hostname} is live and owned by #{owner}#{dir ? " (#{dir})" : ""}."
394
+ $stderr.puts " Work in your own worktree (each branch gets its own URL), or take over explicitly:"
395
+ end
396
+ $stderr.puts " yamine start --force (or bare `yamine --force`)"
315
397
  exit 1
316
398
  end
317
399
  end
data/lib/yamine/cli.rb CHANGED
@@ -25,9 +25,21 @@ module Yamine
25
25
 
26
26
  def run(argv)
27
27
  args = argv.dup
28
- if args.empty? || (!SUBCOMMANDS.include?(args.first) && !args.first.start_with?("-"))
28
+ if args.empty?
29
29
  return BootCommand.run_inferred(Context.new, args)
30
30
  end
31
+ # --help / --version are handled by `help` below, not by boot.
32
+ if args.first&.start_with?("-")
33
+ first = args.first.to_s
34
+ if %w[--help -h --version -v].include?(first)
35
+ return help if %w[--help -h].include?(first)
36
+ return (puts "yamine #{VERSION}"; 0) if %w[--version -v].include?(first)
37
+ end
38
+ return BootCommand.run_inferred(Context.new, args)
39
+ end
40
+ unless SUBCOMMANDS.include?(args.first)
41
+ return BootCommand.run_named(Context.new, args.first, args[1..])
42
+ end
31
43
 
32
44
  cmd = args.shift
33
45
  ctx = Context.new
@@ -168,8 +168,7 @@ module Yamine
168
168
  until Process.clock_gettime(Process::CLOCK_MONOTONIC) > deadline
169
169
  unless process_alive?(app.pid)
170
170
  ms = elapsed_ms(started)
171
- return fail_result(name, started, ms,
172
- "process exited before becoming healthy — see log/yamine-#{name}.log", sink)
171
+ return fail_result(name, started, ms, died_detail(name, slot), sink)
173
172
  end
174
173
  ok, detail = probe(item[:hostname], item[:port], path: path)
175
174
  if ok
@@ -182,13 +181,24 @@ module Yamine
182
181
  ms = elapsed_ms(started)
183
182
  status = process_alive?(app.pid) ? "timeout" : "fail"
184
183
  detail = process_alive?(app.pid) ? "no healthy response within #{timeout}s (#{last_error})" :
185
- "process exited before becoming healthy — see log/yamine-#{name}.log"
184
+ died_detail(name, slot)
186
185
  event = Event.new(phase: "process", action: name, status: status,
187
186
  duration_ms: ms, detail: detail)
188
187
  sink&.event(event)
189
188
  { name: name, status: status, phase: "process", detail: detail, duration_ms: ms }
190
189
  end
191
190
 
191
+ # Why a process that died before answering died: the recognized
192
+ # fatal line when there is one, else point at its log.
193
+ def died_detail(name, slot)
194
+ reason = fatal_line(slot[:log_path])
195
+ if reason
196
+ "process exited: #{reason} (log/yamine-#{name}.log)"
197
+ else
198
+ "process exited before becoming healthy — see log/yamine-#{name}.log"
199
+ end
200
+ end
201
+
192
202
  def ok_result(name, started, ms, path, sink)
193
203
  detail = path ? "healthcheck #{path} returned 2xx-3xx" : "port accepted connection"
194
204
  event = Event.new(phase: "process", action: name, status: "ok",
@@ -204,6 +214,53 @@ module Yamine
204
214
  { name: name, status: "fail", phase: "process", detail: detail, duration_ms: ms }
205
215
  end
206
216
 
217
+ # Rails refuses to boot when tmp/pids/server.pid names a live
218
+ # process ("A server is already running (pid: N)") — the classic
219
+ # interrupted-run leftover. Report the exact reason and the fix
220
+ # BEFORE spawning anything, instead of letting `web` die 2s later.
221
+ # Returns [ok, detail, pid]; ok is true when there is nothing in the
222
+ # way (no file, unreadable, or a dead pid — Rails overwrites those).
223
+ def server_pid_conflict(dir = Dir.pwd)
224
+ path = File.join(dir, "tmp", "pids", "server.pid")
225
+ return [true, nil, nil] unless File.file?(path)
226
+
227
+ pid = File.read(path).strip.to_i
228
+ return [true, nil, nil] unless pid.positive? && process_alive?(pid)
229
+
230
+ [false, "a Rails server is already running (pid #{pid}, from #{path})", pid]
231
+ rescue SystemCallError, ArgumentError
232
+ [true, nil, nil]
233
+ end
234
+
235
+ # Known fatal patterns in a dying backend's log. The log line that
236
+ # killed the process is more useful than "process exited" — agents
237
+ # should not have to open the file to learn why.
238
+ FATAL_PATTERNS = [
239
+ [/A server is already running \(pid:\s*(\d+)/i,
240
+ ->(m) { "another server is already running (pid #{m[1]}); stale tmp/pids/server.pid" }],
241
+ [/Address already in use/i,
242
+ ->(_m) { "address already in use; another process holds the port" }],
243
+ [/Bundler::GemNotFound|Could not find .* locally installed gems|run `?bundle install/i,
244
+ ->(_m) { "gems missing; run `bundle install`" }],
245
+ [/ActiveRecord::NoDatabaseError|PG::ConnectionBad|database .* does not exist/i,
246
+ ->(_m) { "database unreachable or missing; check DATABASE_URL / the database server" }]
247
+ ].freeze
248
+
249
+ def fatal_line(path)
250
+ return nil unless path && File.file?(path)
251
+
252
+ lines = File.readlines(path).last(80)
253
+ FATAL_PATTERNS.each do |pattern, describe|
254
+ match = lines.reverse.find { |l| l.match?(pattern) }
255
+ next unless match
256
+
257
+ return describe.call(match.match(pattern))
258
+ end
259
+ nil
260
+ rescue SystemCallError
261
+ nil
262
+ end
263
+
207
264
  def process_alive?(pid)
208
265
  Process.kill(0, pid)
209
266
  true
data/lib/yamine/runner.rb CHANGED
@@ -103,10 +103,17 @@ module Yamine
103
103
  Process.detach(pid)
104
104
  target = "127.0.0.1:#{port}"
105
105
  if register
106
- spec ||= { "dir" => File.expand_path(dir), "proc" => name }
107
- @store.add_route(hostname, target, Process.pid, kind: "tcp",
108
- force: force, spec: spec)
109
- write_backend_pid(hostname, pid)
106
+ begin
107
+ spec ||= { "dir" => File.expand_path(dir), "proc" => name }
108
+ @store.add_route(hostname, target, Process.pid, kind: "tcp",
109
+ force: force, spec: spec)
110
+ write_backend_pid(hostname, pid)
111
+ rescue StandardError
112
+ # Registration refused (quota, conflict): the backend is
113
+ # already spawned — never leave it running without a route.
114
+ stop_pid(pid)
115
+ raise
116
+ end
110
117
  end
111
118
  App.new(name: name, hostname: hostname, url: url, pid: pid,
112
119
  target: target, kind: "tcp", command: command)
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module Yamine
4
- VERSION = "0.10.2"
4
+ VERSION = "0.10.4"
5
5
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: yamine
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.10.2
4
+ version: 0.10.4
5
5
  platform: ruby
6
6
  authors:
7
7
  - Kaka Ruto