railwatch 0.1.0 → 0.1.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/AGENTS.md +1 -1
- data/CHANGELOG.md +9 -0
- data/README.md +44 -166
- data/docs/ai-and-mcp.md +9 -9
- data/docs/configuration.md +419 -392
- data/docs/getting-started.md +10 -8
- data/docs/records.md +304 -280
- data/docs/source-maps.md +1 -1
- data/docs/testing.md +17 -17
- data/docs/troubleshooting.md +90 -87
- data/lib/railwatch/secret_safety.rb +3 -1
- data/lib/railwatch/version.rb +1 -1
- data/lib/tasks/railwatch_tasks.rake +4 -4
- metadata +1 -1
data/docs/source-maps.md
CHANGED
|
@@ -54,7 +54,7 @@ them and cannot resolve minified locations accurately.
|
|
|
54
54
|
|
|
55
55
|
For a custom release uploader, POST the raw `.map` bytes to
|
|
56
56
|
`/ingest/sourcemaps` with `Content-Type: application/octet-stream`,
|
|
57
|
-
`Authorization: Bearer
|
|
57
|
+
`Authorization: Bearer rw_...`, `X-Railwatch-Deploy`, and
|
|
58
58
|
`X-Railwatch-Filename` (the generated JavaScript URL path without a leading
|
|
59
59
|
slash). A successful response is HTTP 201 with
|
|
60
60
|
`{"ok":true,"filename":"vite/assets/index-abc.js","bytes":1234}`.
|
data/docs/testing.md
CHANGED
|
@@ -2,13 +2,13 @@
|
|
|
2
2
|
|
|
3
3
|
Railwatch already watches every query, N+1, span, exception, and outgoing
|
|
4
4
|
request your app makes. The same instrumentation works in your test suite,
|
|
5
|
-
which means a spec can assert on them
|
|
5
|
+
which means a spec can assert on them. CI can fail a pull request that
|
|
6
6
|
adds an N+1 or doubles a page's query count.
|
|
7
7
|
|
|
8
8
|
## Set-up
|
|
9
9
|
|
|
10
|
-
RSpec
|
|
11
|
-
for you
|
|
10
|
+
RSpec: add one line to `spec/rails_helper.rb`. The install generator adds it
|
|
11
|
+
for you.
|
|
12
12
|
|
|
13
13
|
```ruby
|
|
14
14
|
require "rspec/rails"
|
|
@@ -18,7 +18,7 @@ require "railwatch/rspec"
|
|
|
18
18
|
That requires `railwatch/spec_helper`, includes `Railwatch::SpecHelper` into every
|
|
19
19
|
example group, and defines the matchers below.
|
|
20
20
|
|
|
21
|
-
Minitest
|
|
21
|
+
Minitest: the same thing in `test/test_helper.rb`.
|
|
22
22
|
|
|
23
23
|
```ruby
|
|
24
24
|
require "rails/test_help"
|
|
@@ -31,7 +31,7 @@ end
|
|
|
31
31
|
|
|
32
32
|
Railwatch must be *enabled* in the test environment or every block would look
|
|
33
33
|
empty. `config.enabled?` is true when `config.enabled` is set and a token is
|
|
34
|
-
present, so set any non-blank `RAILWATCH_TOKEN` for the test env
|
|
34
|
+
present, so set any non-blank `RAILWATCH_TOKEN` for the test env. Records go to
|
|
35
35
|
an in-memory transport, never over the network. If Railwatch is disabled, the
|
|
36
36
|
matchers raise `Railwatch::SpecHelper::Disabled` rather than quietly passing.
|
|
37
37
|
|
|
@@ -50,9 +50,9 @@ expect { user.reload }.to have_railwatch_queries(exactly: 1)
|
|
|
50
50
|
expect { Report.generate }.to have_railwatch_queries(at_least: 1)
|
|
51
51
|
```
|
|
52
52
|
|
|
53
|
-
|
|
54
|
-
raises `ArgumentError`. On failure the message lists every statement,
|
|
55
|
-
truncated to 120 characters, so CI output says what to go and fix:
|
|
53
|
+
Pass exactly one of `at_most:`, `exactly:`, `at_least:`. Passing two, or
|
|
54
|
+
none, raises `ArgumentError`. On failure the message lists every statement,
|
|
55
|
+
each truncated to 120 characters, so CI output says what to go and fix:
|
|
56
56
|
|
|
57
57
|
```
|
|
58
58
|
expected the block to run at most 1 database queries, but it ran 3:
|
|
@@ -61,7 +61,7 @@ expected the block to run at most 1 database queries, but it ran 3:
|
|
|
61
61
|
3. SELECT "gadgets".* FROM "gadgets" WHERE "gadgets"."id" = ?
|
|
62
62
|
```
|
|
63
63
|
|
|
64
|
-
Cached queries don't count
|
|
64
|
+
Cached queries don't count. They never become `query` records.
|
|
65
65
|
|
|
66
66
|
### `have_railwatch_n_plus_one`
|
|
67
67
|
|
|
@@ -88,7 +88,7 @@ expect { importer.run }.to record_railwatch_exception(ArgumentError)
|
|
|
88
88
|
expect { importer.run }.not_to record_railwatch_exceptions
|
|
89
89
|
```
|
|
90
90
|
|
|
91
|
-
These see anything that reaches `Rails.error
|
|
91
|
+
These see anything that reaches `Rails.error`: `Rails.error.handle`,
|
|
92
92
|
`Rails.error.report`, `Railwatch.report`, and unhandled exceptions a request
|
|
93
93
|
spec's middleware catches. A block that raises out of the matcher still
|
|
94
94
|
raises; nothing is swallowed.
|
|
@@ -116,8 +116,8 @@ produces the same statement listing on failure.
|
|
|
116
116
|
## Where the matchers work
|
|
117
117
|
|
|
118
118
|
Anywhere. A request spec's `get "/widgets"` opens and closes its own
|
|
119
|
-
execution
|
|
120
|
-
|
|
119
|
+
execution. So its whole tree is visible by the time the block returns:
|
|
120
|
+
queries, N+1s, outgoing HTTP.
|
|
121
121
|
|
|
122
122
|
```ruby
|
|
123
123
|
expect { get "/widgets" }.to have_railwatch_queries(at_most: 6)
|
|
@@ -129,7 +129,7 @@ execution for the duration of the assertion and closed afterwards. No parent
|
|
|
129
129
|
opened yourself has its records read straight off that execution's buffer.
|
|
130
130
|
|
|
131
131
|
Under the hood every matcher calls `Railwatch::SpecHelper#railwatch_capture`,
|
|
132
|
-
which is public
|
|
132
|
+
which is public. Use it directly for anything the matchers don't cover:
|
|
133
133
|
|
|
134
134
|
```ruby
|
|
135
135
|
records = railwatch_capture { get "/widgets" }
|
|
@@ -164,12 +164,12 @@ RSpec.describe "performance budgets", type: :request do
|
|
|
164
164
|
end
|
|
165
165
|
```
|
|
166
166
|
|
|
167
|
-
Two ways to run it
|
|
168
|
-
|
|
169
|
-
job list
|
|
167
|
+
Two ways to run it. Tag these examples and run them as their own CI step
|
|
168
|
+
with `bundle exec rspec --tag performance`, so a budget failure is obvious in
|
|
169
|
+
the job list. Or leave them in the main suite so any pull request that adds a
|
|
170
170
|
query fails immediately. Either way the failure message names the statements,
|
|
171
171
|
so the fix is usually an `includes` one line away.
|
|
172
172
|
|
|
173
173
|
Seed enough rows in `before` that an N+1 actually crosses
|
|
174
|
-
`config.n_plus_one_threshold
|
|
174
|
+
`config.n_plus_one_threshold`. With three records, a five-query threshold
|
|
175
175
|
never fires and the gate passes on code that would melt in production.
|
data/docs/troubleshooting.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Troubleshooting
|
|
2
2
|
|
|
3
3
|
Start with `bin/rails railwatch:doctor`. It checks every piece of the
|
|
4
|
-
install in one pass and prints a `✓`/`✗` line per piece
|
|
4
|
+
install in one pass and prints a `✓`/`✗` line per piece. The sections
|
|
5
5
|
below are keyed to those lines.
|
|
6
6
|
|
|
7
7
|
```sh
|
|
@@ -9,14 +9,14 @@ bin/rails railwatch:doctor
|
|
|
9
9
|
```
|
|
10
10
|
|
|
11
11
|
The task exits non-zero only when **token** or **ingest reachable**
|
|
12
|
-
fails. Everything else is informational
|
|
12
|
+
fails. Everything else is informational. A `✗` there means a feature
|
|
13
13
|
isn't wired, not that the install is broken.
|
|
14
14
|
|
|
15
15
|
| Doctor line | What a `✗` means |
|
|
16
16
|
|---|---|
|
|
17
17
|
| `token` | `RAILWATCH_TOKEN` is unset or empty. Fatal: nothing is recorded at all. |
|
|
18
18
|
| `ingest url` | `ingest_url` isn't a parseable HTTP(S) URL. |
|
|
19
|
-
| `ingest reachable` | `GET {ingest_url}/ingest/ping` didn't return success. Fatal. |
|
|
19
|
+
| `ingest reachable` | `GET {ingest_url}/ingest/ping` didn't return success. Fatal. The ping carries the token, so a missing or wrong token fails this line too; fix `token` first. |
|
|
20
20
|
| `request middleware` | `Railwatch::Middleware::Request` isn't in the stack, so requests aren't executions. |
|
|
21
21
|
| `engine mounted` | `mount Railwatch::Engine, at: "/railwatch"` is missing from `config/routes.rb`; the browser beacon has nowhere to post. |
|
|
22
22
|
| `deploy` | `config.deploy` is unset — records ship, charts get no deploy markers. |
|
|
@@ -33,37 +33,38 @@ isn't wired, not that the install is broken.
|
|
|
33
33
|
**Symptom.** The environment's pages stay empty however much traffic the
|
|
34
34
|
app takes.
|
|
35
35
|
|
|
36
|
-
Work down this list
|
|
37
|
-
different angles
|
|
36
|
+
Work down this list. The first five are the same root cause seen from
|
|
37
|
+
different angles: Railwatch decided not to record.
|
|
38
38
|
|
|
39
39
|
**The token is missing or blank.** `Railwatch.enabled?` is
|
|
40
40
|
`config.enabled && token.present?`. With no token the engine's
|
|
41
41
|
`railwatch.subscribe` initializer returns early, so no subscribers and no
|
|
42
|
-
patches are installed at all
|
|
43
|
-
development. Fix: set `RAILWATCH_TOKEN`, restart, re-run
|
|
44
|
-
|
|
45
|
-
enough to spot a truncated or quoted value
|
|
42
|
+
patches are installed at all. This is by design, so the gem is inert in
|
|
43
|
+
development. Fix: set `RAILWATCH_TOKEN`, restart, and re-run
|
|
44
|
+
`railwatch:doctor`. The `token` line prints the first 6 characters and
|
|
45
|
+
the length, which is enough to spot a truncated or quoted value.
|
|
46
46
|
|
|
47
47
|
**The token is wrong.** A 401 from the ingest marks the transport
|
|
48
48
|
permanently unauthorized: no further flush is attempted for the lifetime
|
|
49
|
-
of that process. Fixing the env var isn't enough
|
|
49
|
+
of that process. Fixing the env var isn't enough. Restart the process.
|
|
50
50
|
`railwatch:doctor`'s `ingest reachable` line catches this before you
|
|
51
51
|
deploy.
|
|
52
52
|
|
|
53
53
|
**`RAILWATCH_INGEST_URL` points somewhere else.** Records go where you sent
|
|
54
|
-
them. `railwatch:status` prints the URL it is actually using
|
|
54
|
+
them. `railwatch:status` prints the URL it is actually using. Compare it
|
|
55
55
|
against the platform you're looking at. Self-hosting: see
|
|
56
56
|
[`self-hosting.md`](self-hosting.md).
|
|
57
57
|
|
|
58
58
|
**`config.enabled` is false.** `RAILWATCH_ENABLED=0` (or `false`/`no`/`off`)
|
|
59
59
|
turns everything off with a valid token present.
|
|
60
60
|
|
|
61
|
-
**Sample rates are at zero.** `c.sample = { requests: 0.0 }`
|
|
62
|
-
per-route `railwatch_never_sample` macro on
|
|
63
|
-
|
|
64
|
-
values. Note that an unhandled exception still
|
|
65
|
-
execution,
|
|
66
|
-
signature of a low sample rate
|
|
61
|
+
**Sample rates are at zero.** `c.sample = { requests: 0.0 }` means no
|
|
62
|
+
request records. So does the per-route `railwatch_never_sample` macro on
|
|
63
|
+
a controller. The `sample rates` doctor line prints the effective
|
|
64
|
+
values. Note that an unhandled exception can still ship from a
|
|
65
|
+
sampled-out execution, subject to the `exceptions` rate. So "exceptions
|
|
66
|
+
arrive but nothing else does" is the signature of a low sample rate
|
|
67
|
+
rather than a broken install.
|
|
67
68
|
|
|
68
69
|
**The record type is ignored.** `c.ignore` drops a type before it is
|
|
69
70
|
built. The `ignored record types` doctor line prints the list. Ignoring
|
|
@@ -71,14 +72,14 @@ built. The `ignored record types` doctor line prints the list. Ignoring
|
|
|
71
72
|
|
|
72
73
|
**You're looking at the test environment.** Requiring `railwatch/rspec`
|
|
73
74
|
(or `railwatch/minitest`) swaps the reporter's transport for an in-memory
|
|
74
|
-
one
|
|
75
|
+
one. A suite records normally but never sends anything over the
|
|
75
76
|
network. Independently: the health sampler, the session flusher, and the
|
|
76
77
|
profiler all refuse to start when `Rails.env.test?`.
|
|
77
78
|
|
|
78
79
|
Still nothing? Set `RAILWATCH_DEBUG=1` and restart. Internal diagnostics go
|
|
79
|
-
to stderr prefixed `[railwatch]
|
|
80
|
-
become `log` records about themselves
|
|
81
|
-
... }` gets the same failures as a callback.
|
|
80
|
+
to stderr prefixed `[railwatch]`. They never go to `Rails.logger`, so they
|
|
81
|
+
can't become `log` records about themselves. `Railwatch.on_unrecoverable
|
|
82
|
+
{ |e| ... }` gets the same failures as a callback.
|
|
82
83
|
|
|
83
84
|
## Doubled scheduled_task records
|
|
84
85
|
|
|
@@ -86,16 +87,16 @@ become `log` records about themselves). `Railwatch.on_unrecoverable { |e|
|
|
|
86
87
|
page, at the same minute, usually with different `drift`.
|
|
87
88
|
|
|
88
89
|
**Cause.** Two Solid Queue supervisors are running against the same
|
|
89
|
-
queue database
|
|
90
|
-
`SOLID_QUEUE_IN_PUMA=true` while a dedicated job role
|
|
91
|
-
two containers of the job role. Each supervisor has
|
|
92
|
-
scheduler and its own workers, so the job really is
|
|
93
|
-
Railwatch doesn't dedupe: it records one
|
|
94
|
-
`perform.active_job`, in whichever process performed
|
|
95
|
-
rows are a true report of a doubled run.
|
|
90
|
+
queue database. That could be a stale `bin/jobs` left over from a
|
|
91
|
+
previous `bin/dev`, `SOLID_QUEUE_IN_PUMA=true` while a dedicated job role
|
|
92
|
+
is also booted, or two containers of the job role. Each supervisor has
|
|
93
|
+
its own recurring scheduler and its own workers, so the job really is
|
|
94
|
+
performed twice. Railwatch doesn't dedupe: it records one
|
|
95
|
+
`scheduled_task` per `perform.active_job`, in whichever process performed
|
|
96
|
+
it. The doubled rows are a true report of a doubled run.
|
|
96
97
|
|
|
97
98
|
**Fix.** Run one supervisor. `SolidQueue::Process.where(kind:
|
|
98
|
-
"Supervisor")` tells you how many think they're alive
|
|
99
|
+
"Supervisor")` tells you how many think they're alive. `bin/kamal app
|
|
99
100
|
logs -r job` tells you which containers are booting one. The same
|
|
100
101
|
duplication also doubles the work itself, so this is worth fixing
|
|
101
102
|
regardless of what the dashboard says.
|
|
@@ -111,9 +112,8 @@ fine.
|
|
|
111
112
|
the real class never runs.
|
|
112
113
|
|
|
113
114
|
**Fix.** Re-prepend the patch onto the replacement, once, before the
|
|
114
|
-
suite. This is exactly what the gem's own suite does
|
|
115
|
-
|
|
116
|
-
same:
|
|
115
|
+
suite. This is exactly what the gem's own suite does, in
|
|
116
|
+
`spec/spec_helper.rb`. App suites using WebMock should do the same:
|
|
117
117
|
|
|
118
118
|
```ruby
|
|
119
119
|
# WebMock replaces ::Net::HTTP with a subclass whose #request short-circuits
|
|
@@ -132,35 +132,35 @@ end
|
|
|
132
132
|
have workers.
|
|
133
133
|
|
|
134
134
|
**What actually happens.** Ruby routes `fork`, `Process.fork`, and
|
|
135
|
-
`Kernel#fork` through `Process._fork
|
|
136
|
-
with Rails' own `ActiveSupport::ForkTracker
|
|
137
|
-
uses to reset its connection pools
|
|
135
|
+
`Kernel#fork` through `Process._fork`. Railwatch registers one callback
|
|
136
|
+
with Rails' own `ActiveSupport::ForkTracker`, the same hook Active Record
|
|
137
|
+
uses to reset its connection pools. Before the child returns from
|
|
138
138
|
`fork`, it replaces the inherited reporter buffer, drop accounting,
|
|
139
139
|
transport policy state, mutexes, condition variables, dead threads, and the
|
|
140
|
-
profiler's process-global state
|
|
141
|
-
otherwise leave the child permanently unable to profile
|
|
140
|
+
profiler's process-global state. A parent's in-flight profile would
|
|
141
|
+
otherwise leave the child permanently unable to profile.
|
|
142
142
|
The parent's half-finished session map is discarded too. The child then
|
|
143
143
|
emits its own `process` record and starts fresh health/session threads for
|
|
144
|
-
its role. Parent records remain owned by and delivered from the parent
|
|
145
|
-
|
|
144
|
+
its role. Parent records remain owned by and delivered from the parent.
|
|
145
|
+
They can never be replayed by every child. **No `on_worker_boot`
|
|
146
146
|
configuration is needed**, in Puma cluster mode or in Solid Queue's
|
|
147
147
|
forked workers.
|
|
148
148
|
|
|
149
149
|
The synchronization objects are replaced without locking them. That is
|
|
150
|
-
deliberate
|
|
150
|
+
deliberate. If another parent thread owned a mutex at the instant of
|
|
151
151
|
`fork`, Ruby preserves the locked mutex in the child but not the thread
|
|
152
152
|
that could unlock it.
|
|
153
153
|
|
|
154
154
|
**When a process legitimately reports nothing.** `Health.start!` returns
|
|
155
155
|
early unless the process's role is `web` or `worker`, and
|
|
156
156
|
`Sessions.start!` only runs for `web`. Role detection is
|
|
157
|
-
`Railwatch::Subscribers::ProcessInfo.role`, in this order
|
|
158
|
-
Solid Queue is loaded and `$PROGRAM_NAME` includes `"jobs"` (or the
|
|
159
|
-
command starts with `solid_queue:`)
|
|
160
|
-
`$PROGRAM_NAME` ends in `rake
|
|
161
|
-
`process`. The worker check comes first deliberately
|
|
162
|
-
a job container too. A console or a rake task ships no health
|
|
163
|
-
design, and both modules also return early in the `test` env.
|
|
157
|
+
`Railwatch::Subscribers::ProcessInfo.role`, in this order. First `worker`,
|
|
158
|
+
when Solid Queue is loaded and `$PROGRAM_NAME` includes `"jobs"` (or the
|
|
159
|
+
command starts with `solid_queue:`). Then `console`. Then `command`, when
|
|
160
|
+
`$PROGRAM_NAME` ends in `rake`. Then `web`, when Puma is defined. Else
|
|
161
|
+
`process`. The worker check comes first deliberately, because Puma is
|
|
162
|
+
loaded in a job container too. A console or a rake task ships no health
|
|
163
|
+
records by design, and both modules also return early in the `test` env.
|
|
164
164
|
|
|
165
165
|
## Memory growth with tail sampling on
|
|
166
166
|
|
|
@@ -169,68 +169,71 @@ design, and both modules also return early in the `test` env.
|
|
|
169
169
|
|
|
170
170
|
**Cause.** That is the trade-off, not a leak. With head sampling only, a
|
|
171
171
|
sampled-out execution builds and buffers nothing. With tail sampling on,
|
|
172
|
-
*every* execution buffers its child records
|
|
173
|
-
logs, view renders
|
|
174
|
-
|
|
172
|
+
*every* execution buffers its child records for its whole lifetime:
|
|
173
|
+
queries, cache events, logs, view renders. The keep-or-discard decision
|
|
174
|
+
can't be made until it ends.
|
|
175
175
|
|
|
176
176
|
**What to check.**
|
|
177
177
|
|
|
178
178
|
- Per execution, the buffer is capped at `Execution::MAX_RECORDS`
|
|
179
|
-
(10,000). Past that, records are dropped and counted
|
|
179
|
+
(10,000). Past that, records are dropped and counted. The count is
|
|
180
180
|
added to the reporter's drop counter so the loss is visible on the
|
|
181
181
|
platform rather than silent.
|
|
182
182
|
- `c.buffer_size` (default 10,000, the same as `MAX_RECORDS`) caps the
|
|
183
183
|
process-wide queue between the app and the reporter thread.
|
|
184
|
-
Oldest-dropped-first, also counted. Do not set it below `MAX_RECORDS
|
|
185
|
-
|
|
186
|
-
a tree larger than the queue loses its own first records
|
|
187
|
-
the outgoing requests a long job made before it started
|
|
188
|
-
Keeping far more executions than before means far more
|
|
189
|
-
arriving at this queue
|
|
190
|
-
- `c.profile_slow_ms` compounds it
|
|
184
|
+
Oldest-dropped-first, also counted. Do not set it below `MAX_RECORDS`.
|
|
185
|
+
An execution's tree is written to the queue in one go when it ends, so
|
|
186
|
+
a tree larger than the queue loses its own first records. Typically
|
|
187
|
+
those are the outgoing requests a long job made before it started
|
|
188
|
+
writing. Keeping far more executions than before means far more
|
|
189
|
+
records arriving at this queue. Raise it, or lower what you keep.
|
|
190
|
+
- `c.profile_slow_ms` compounds it. It profiles every tail-buffering
|
|
191
191
|
execution from its first line and throws away the fast ones, so the
|
|
192
192
|
profiler's stack table is held alongside the record buffer.
|
|
193
193
|
- `c.failure_context` buffers sampled-out executions too, but a ring of
|
|
194
194
|
that many records each rather than all of them. If RSS climbed after
|
|
195
|
-
setting it, lower the count
|
|
195
|
+
setting it, lower the count. It is a per-execution bound, so the
|
|
196
196
|
process-wide cost is that many records times the executions running
|
|
197
197
|
concurrently.
|
|
198
198
|
|
|
199
|
-
**Fix.**
|
|
200
|
-
|
|
201
|
-
|
|
202
|
-
|
|
199
|
+
**Fix.** The threshold decides what is kept, not what is buffered:
|
|
200
|
+
with tail sampling on, every execution buffers until it ends, so
|
|
201
|
+
lowering `tail_sample_slow_ms` keeps more trees rather than fewer. To
|
|
202
|
+
buffer less, turn tail sampling off and use `Railwatch.keep!` on the
|
|
203
|
+
specific paths you care about, or drop the highest-volume child types
|
|
204
|
+
with `c.ignore` so they are never buffered.
|
|
203
205
|
|
|
204
206
|
## Profiles never appear
|
|
205
207
|
|
|
206
208
|
**Symptom.** `c.profile_sample` is set but no `profile` records ship and
|
|
207
209
|
no request is marked `profiled`.
|
|
208
210
|
|
|
209
|
-
**Cause.** No profiler backend is installed. Railwatch doesn't vendor one
|
|
211
|
+
**Cause.** No profiler backend is installed. Railwatch doesn't vendor one.
|
|
210
212
|
`Railwatch::Profiler.available?` is false unless `vernier` or `stackprof`
|
|
211
213
|
loads, and every profiling setting is inert while it is. The doctor's
|
|
212
214
|
`profiler backend` line reports this.
|
|
213
215
|
|
|
214
216
|
**Fix.** Add `gem "vernier"` (Ruby ≥ 3.2, preferred) or
|
|
215
217
|
`gem "stackprof"`. Two other reasons a profile can be absent even with a
|
|
216
|
-
backend
|
|
217
|
-
stays off rather than falling back
|
|
218
|
-
is skipped unless `profile_sample` is explicitly non-zero. Both
|
|
219
|
-
are process-global, so an execution that starts while another
|
|
220
|
-
being profiled is simply not profiled
|
|
218
|
+
backend. One is `c.profiler` pinned to a name that doesn't load; profiling
|
|
219
|
+
stays off rather than falling back. The other is the `test` env, where
|
|
220
|
+
profiling is skipped unless `profile_sample` is explicitly non-zero. Both
|
|
221
|
+
backends are process-global, so an execution that starts while another
|
|
222
|
+
one is being profiled is simply not profiled. That is expected, not a
|
|
223
|
+
bug.
|
|
221
224
|
|
|
222
225
|
## No deploy marker on the charts
|
|
223
226
|
|
|
224
227
|
**Symptom.** Charts have no vertical deploy lines; the Releases page
|
|
225
228
|
groups everything under one blank release.
|
|
226
229
|
|
|
227
|
-
**Cause.** `config.deploy` is unset. The doctor's `deploy` line says so
|
|
228
|
-
|
|
229
|
-
or initializer it came from.
|
|
230
|
+
**Cause.** `config.deploy` is unset. The doctor's `deploy` line says so.
|
|
231
|
+
When it is set, the line names the environment variable, `REVISION`, Git
|
|
232
|
+
checkout, or initializer it came from.
|
|
230
233
|
|
|
231
234
|
**Fix.** Set one of them; they are read in this order:
|
|
232
235
|
|
|
233
|
-
1. `RAILWATCH_DEPLOY
|
|
236
|
+
1. `RAILWATCH_DEPLOY`, the explicit override on any platform.
|
|
234
237
|
2. `KAMAL_VERSION`.
|
|
235
238
|
3. `GIT_REV`, `GIT_SHA`, `SOURCE_VERSION`, `HEROKU_SLUG_COMMIT`,
|
|
236
239
|
`RENDER_GIT_COMMIT`, the tag from `FLY_IMAGE_REF`,
|
|
@@ -243,8 +246,8 @@ Full 40-character SHAs are shortened to 12 characters. Or assign `deploy` in
|
|
|
243
246
|
the initializer. Set `detect_deploy`/`RAILWATCH_DETECT_DEPLOY` to false to ignore
|
|
244
247
|
steps 3–5. The value is stamped on every record, so a change only affects
|
|
245
248
|
records shipped after the restart. Note that
|
|
246
|
-
`config.deploy` and the deploy *marker* are two different things
|
|
247
|
-
marker
|
|
249
|
+
`config.deploy` and the deploy *marker* are two different things. The
|
|
250
|
+
marker, with its commit list, comes from `railwatch:deploy` or the Kamal
|
|
248
251
|
hook below.
|
|
249
252
|
|
|
250
253
|
## The Kamal hook doesn't fire
|
|
@@ -255,14 +258,14 @@ hook below.
|
|
|
255
258
|
|
|
256
259
|
- **`RAILWATCH_TOKEN` isn't exported to the hook.** The first thing
|
|
257
260
|
`.kamal/hooks/post-deploy` does is `[ -z "$RAILWATCH_TOKEN" ] && exit 0`.
|
|
258
|
-
The hook runs on the deployer machine, in your shell
|
|
259
|
-
container
|
|
260
|
-
*app* isn't necessarily in the deployer's environment. Export it there
|
|
261
|
-
|
|
261
|
+
The hook runs on the deployer machine, in your shell, not in a
|
|
262
|
+
container. So a token that only exists in `.kamal/secrets` for the
|
|
263
|
+
*app* isn't necessarily in the deployer's environment. Export it there,
|
|
264
|
+
or source the same secret store your CI uses.
|
|
262
265
|
- **`curl` or `ruby` isn't on the deployer, or `RAILWATCH_INGEST_URL`
|
|
263
266
|
isn't set.** Then the hook falls back to
|
|
264
267
|
`bin/kamal app exec --primary --reuse "bin/rails railwatch:deploy[$KAMAL_VERSION]"`,
|
|
265
|
-
which records the same deploy **minus the commit list
|
|
268
|
+
which records the same deploy **minus the commit list**. A container
|
|
266
269
|
has the code but not the git history. If your deploys show up without
|
|
267
270
|
commits, this is the path you're on.
|
|
268
271
|
- **The hook isn't there.** The install generator only writes it when
|
|
@@ -278,8 +281,8 @@ it exits 0 regardless.
|
|
|
278
281
|
fewer lines and highlights nothing.
|
|
279
282
|
|
|
280
283
|
**Cause.** Full-text search uses SQLite's FTS5 (`logs_fts`). The
|
|
281
|
-
platform checks for both
|
|
282
|
-
|
|
284
|
+
platform checks for both a SQLite adapter *and* the `logs_fts` table.
|
|
285
|
+
When either is absent it falls back to `message LIKE '%...%'`. That
|
|
283
286
|
fallback is a plain substring match: no phrase or negation syntax, and no
|
|
284
287
|
snippet highlighting.
|
|
285
288
|
|
|
@@ -294,11 +297,11 @@ self-hosted platform operator should rebuild the documented search index.
|
|
|
294
297
|
**Symptom.** Every other page has data; Tenants shows nothing.
|
|
295
298
|
|
|
296
299
|
**Cause.** `tenant` is a column on every telemetry row, filled from
|
|
297
|
-
`Railwatch::Context.current_tenant
|
|
300
|
+
`Railwatch::Context.current_tenant`. A tenant only exists as a GROUP BY
|
|
298
301
|
over those rows. If nothing ever sets it, every row has a null tenant and
|
|
299
302
|
there is nothing to group.
|
|
300
303
|
|
|
301
|
-
**Fix.** Apps on `activerecord-tenanted` get it free
|
|
304
|
+
**Fix.** Apps on `activerecord-tenanted` get it free. Railwatch reads
|
|
302
305
|
`ActiveRecord::Base.current_tenant` / `TenantRecord.current_tenant` with
|
|
303
306
|
no configuration. Everyone else sets it explicitly, as early in the
|
|
304
307
|
request as the tenant is known:
|
|
@@ -313,7 +316,7 @@ does not retroactively apply to it.
|
|
|
313
316
|
|
|
314
317
|
## See also
|
|
315
318
|
|
|
316
|
-
- [`configuration.md`](configuration.md)
|
|
317
|
-
- [`records.md`](records.md)
|
|
318
|
-
- [`faq.md`](faq.md)
|
|
319
|
+
- [`configuration.md`](configuration.md): every option and its default.
|
|
320
|
+
- [`records.md`](records.md): what each record type contains.
|
|
321
|
+
- [`faq.md`](faq.md): overhead, retention, PII, and what happens when
|
|
319
322
|
the platform is unreachable.
|
|
@@ -6,7 +6,9 @@ module Railwatch
|
|
|
6
6
|
# Read-only Git checks used by the installer and doctor. Tokens are never
|
|
7
7
|
# returned in diagnostics: callers get a path or a short prefix only.
|
|
8
8
|
module SecretSafety
|
|
9
|
-
|
|
9
|
+
# rw_ is the ingest token prefix; lt_ was the prefix before 0.1.1 and is
|
|
10
|
+
# still matched so a token minted earlier is still caught in a tracked file.
|
|
11
|
+
TOKEN_PATTERN = /\b(?:rw|lt)_[A-Za-z0-9_-]{6,}\b/
|
|
10
12
|
TOKEN_FILE_GLOBS = [ ".env", ".env.*", ".kamal/secrets", "config/deploy.yml",
|
|
11
13
|
"config/initializers/*.rb" ].freeze
|
|
12
14
|
|
data/lib/railwatch/version.rb
CHANGED
|
@@ -186,11 +186,11 @@ namespace :railwatch do
|
|
|
186
186
|
|
|
187
187
|
1. Sign in (or sign up) at #{base.url("/dashboard")}
|
|
188
188
|
2. New application, then New environment (production, staging, ...)
|
|
189
|
-
3. The environment's token (
|
|
189
|
+
3. The environment's token (rw_...) is shown once, right after it is created.
|
|
190
190
|
|
|
191
191
|
Then set it where this app reads its environment:
|
|
192
192
|
|
|
193
|
-
RAILWATCH_TOKEN=
|
|
193
|
+
RAILWATCH_TOKEN=rw_...#{"\n RAILWATCH_INGEST_URL=#{base.host_url}" if base.self_hosted?}
|
|
194
194
|
|
|
195
195
|
With Kamal: bin/rails generate railwatch:install --prompt-token --kamal-secrets
|
|
196
196
|
Verify: bin/rails railwatch:doctor
|
|
@@ -204,12 +204,12 @@ namespace :railwatch do
|
|
|
204
204
|
task mcp: :environment do
|
|
205
205
|
base = Railwatch::Endpoints.new(Railwatch.config)
|
|
206
206
|
mcp = base.url("/mcp")
|
|
207
|
-
token = "
|
|
207
|
+
token = "rwp_your_token_here"
|
|
208
208
|
puts <<~TEXT
|
|
209
209
|
Railwatch MCP server: #{mcp}
|
|
210
210
|
|
|
211
211
|
An MCP token is per person, not per app: Settings -> Profile -> "API & MCP
|
|
212
|
-
token" at #{base.url("/settings/profile")}. It starts with
|
|
212
|
+
token" at #{base.url("/settings/profile")}. It starts with rwp_ and is
|
|
213
213
|
shown once. Everything below is scoped to whatever accounts that user
|
|
214
214
|
belongs to.
|
|
215
215
|
|