dead_bro 0.2.31 → 0.2.33
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +13 -0
- data/FEATURES.md +333 -0
- data/README.md +2 -2
- data/lib/dead_bro/configuration.rb +23 -11
- data/lib/dead_bro/job_subscriber.rb +42 -113
- data/lib/dead_bro/sql_subscriber.rb +104 -2
- data/lib/dead_bro/sql_tracking_middleware.rb +1 -0
- data/lib/dead_bro/subscriber.rb +44 -33
- data/lib/dead_bro/version.rb +1 -1
- metadata +3 -2
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: f1569a9a3be524d19ba616313570f01cb251647a76384347993655a8ffbaef9e
|
|
4
|
+
data.tar.gz: 3a091b0d844b12cfb9b61cf7b0d2eec04514bb5ef988fba7203d45b4ab237e67
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: ebc44f95253369e7c8ab6f3873ed9143e0d72382c477bae39ca09491c22c93b236b2bc1dcc07599dd755797564f4a0f25f34104cf1f95246ea8d91fa9b4c1935
|
|
7
|
+
data.tar.gz: 6dd40034800398ceac4e82895d427e484ee7fd1b321a22882c1138c0e73514bfb4675a14437abd88cb6f4e5327c01ad85f2c7e75269956c4dd9d4e240a946714
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,18 @@
|
|
|
1
1
|
## [Unreleased]
|
|
2
2
|
|
|
3
|
+
## [0.2.32] - 2026-09-22
|
|
4
|
+
|
|
5
|
+
### Fixed
|
|
6
|
+
- **Errored requests now ship the SQL/cache/memory/etc data captured before the exception, not just the exception itself.** Previously, as soon as a request raised (e.g. an `ActiveRecord::LockWaitTimeout` mid-transaction), the subscriber discarded every already-captured `sql_queries`, `cache_events`, `ar_instantiation_count`, `memory_events`, `gc_stats`, etc. entry and sent only exception metadata — so error pages showed a backtrace with zero query context even when the gem had captured it. The error payload now carries the same detail fields a successful request would.
|
|
7
|
+
- **Background job failures were reported to the dashboard as successful completions, not dropped.** There is no `exception.active_job` event anywhere in Rails/ActiveJob — the handler that subscribed to it never fired. `perform.active_job` (the real event) fires whether or not the job raised, but its handler built `status: "completed"` unconditionally and never inspected `data[:exception_object]`, so a failing job wasn't invisible — it was recorded as if it had succeeded. This silently skewed job success-rate and duration numbers for every account on a version prior to this one; there's nothing to backfill, since no exception data was ever sent for those occurrences. The handler now branches on `data[:exception_object]`: a failing job reports `status: "failed"` with `error`, `fingerprint`, and `cause_chain`, and ships regardless of sampling.
|
|
8
|
+
- **Known limitation:** this only covers a job whose exception escapes `perform_now` uncaught. A job using `retry_on`, `discard_on`, or a custom `rescue_from` has the exception handled *inside* `perform_now`, before `perform.active_job`'s payload is ever built — so a retried or discarded job still reports `status: "completed"` for every handled attempt. Only the final attempt of a `retry_on` job that exhausts its retries with no block given re-raises far enough to be visible. Covering retried/discarded attempts would need separate instrumentation on `enqueue_retry.active_job` / `retry_stopped.active_job` / `discard.active_job` — not done here.
|
|
9
|
+
- **A request opening 5 or more transactions/savepoints could trip a false N+1 flag.** Every Rails adapter checked (MySQL/PostgreSQL/SQLite3, 7.1–8.1) logs transaction-control statements under the single name `"TRANSACTION"`, which the old skip list (`data[:name] == "BEGIN" || ...`) never matched — so `BEGIN`/`COMMIT`/etc. weren't ignored, they were tracked as ordinary queries: sanitized, counted toward N+1 detection, and folded into one aggregate per normalized statement. Five or more transactions in one request hit `N_PLUS_ONE_THRESHOLD` and flagged that aggregate `n_plus_one: true`, which the backend trusts directly for the aggregate payload format. Fixed as a consequence of the `TXN_CONTROL_NAMES` match below, not a separate change.
|
|
10
|
+
|
|
11
|
+
### Added
|
|
12
|
+
- Transaction-control statements (`BEGIN`/`COMMIT`/`ROLLBACK`/`SAVEPOINT`/`RELEASE`) are now captured as lightweight breadcrumbs (`transaction_events`, with offsets) instead of being tracked as ordinary queries — so a transaction that rolled back before an error shows up explicitly in the request trace, instead of as an anonymous, sanitized `sql_queries` entry.
|
|
13
|
+
|
|
14
|
+
## [0.2.31] - 2026-08-30
|
|
15
|
+
|
|
3
16
|
### Added
|
|
4
17
|
- Monitor thread now sends a synchronous heartbeat on startup before the first collection tick. This ensures remote settings — including `monitor_enabled` — are applied from the very first reporting cycle, so Sidekiq workers and other non-web processes that have not yet sent any metrics still receive the correct configuration immediately on boot rather than waiting up to 60 seconds for the first scheduled tick.
|
|
5
18
|
|
data/FEATURES.md
ADDED
|
@@ -0,0 +1,333 @@
|
|
|
1
|
+
# ApmBro Feature List
|
|
2
|
+
|
|
3
|
+
A comprehensive feature list for comparing ApmBro with other APM (Application Performance Monitoring) tools.
|
|
4
|
+
|
|
5
|
+
## Core Architecture
|
|
6
|
+
|
|
7
|
+
- **Rails Integration**: Automatic subscription to Rails events via ActiveSupport::Notifications
|
|
8
|
+
- **Zero-Configuration Setup**: Works out of the box with minimal configuration
|
|
9
|
+
- **Asynchronous Metrics Posting**: Non-blocking HTTP requests using background threads
|
|
10
|
+
- **Thread-Local Storage**: Per-request metric collection using thread-local variables
|
|
11
|
+
- **Circuit Breaker Pattern**: Built-in circuit breaker to prevent cascading failures when APM endpoint is down
|
|
12
|
+
- **Deploy Tracking**: Automatic deploy ID resolution from multiple sources (Rails settings, ENV vars, Heroku, Git)
|
|
13
|
+
|
|
14
|
+
## Request Tracking
|
|
15
|
+
|
|
16
|
+
### Controller Action Monitoring
|
|
17
|
+
- **Automatic Tracking**: Tracks all controller actions automatically
|
|
18
|
+
- **Request Duration**: Measures total request processing time
|
|
19
|
+
- **HTTP Method & Path**: Captures HTTP method and request path
|
|
20
|
+
- **Status Codes**: Tracks HTTP response status codes
|
|
21
|
+
- **View Runtime**: Separate tracking of view rendering time
|
|
22
|
+
- **Database Runtime**: Separate tracking of database query time
|
|
23
|
+
- **Request Parameters**: Captures request parameters (with sensitive data filtering)
|
|
24
|
+
- **User Agent**: Tracks user agent strings
|
|
25
|
+
- **User ID Extraction**: Extracts authenticated user ID (supports Warden)
|
|
26
|
+
- **Environment Context**: Tracks Rails environment (development, staging, production)
|
|
27
|
+
|
|
28
|
+
### Request Sampling
|
|
29
|
+
- **Configurable Sample Rate**: Percentage-based sampling (1-100%)
|
|
30
|
+
- **Random Sampling**: Each request has random chance of being tracked
|
|
31
|
+
- **Consistent Per-Request**: Sampling decision applies to all metrics for a request
|
|
32
|
+
- **Error Override**: Errors are always tracked regardless of sampling
|
|
33
|
+
- **Cost Optimization**: Reduces data volume and costs for high-traffic applications
|
|
34
|
+
|
|
35
|
+
### Exclusion Rules
|
|
36
|
+
- **Controller Exclusion**: Exclude entire controllers from tracking
|
|
37
|
+
- **Action Exclusion**: Exclude specific controller#action combinations
|
|
38
|
+
- **Wildcard Support**: Pattern matching with `*` wildcards (e.g., `Admin::*`, `Admin::*#*`)
|
|
39
|
+
- **Job Exclusion**: Exclude specific background jobs from tracking
|
|
40
|
+
- **Flexible Configuration**: Configure via initializer, Rails settings, or environment variables
|
|
41
|
+
|
|
42
|
+
## SQL Query Tracking
|
|
43
|
+
|
|
44
|
+
### Query Details
|
|
45
|
+
- **Full SQL Tracking**: Captures all SQL queries executed during requests and jobs
|
|
46
|
+
- **Query Sanitization**: Automatically sanitizes SQL to remove sensitive data
|
|
47
|
+
- **Query Name**: Tracks query names (e.g., "User Load", "User Update")
|
|
48
|
+
- **Duration Measurement**: Precise query execution time in milliseconds
|
|
49
|
+
- **Cache Detection**: Identifies cached queries
|
|
50
|
+
- **Connection ID**: Tracks database connection ID
|
|
51
|
+
- **Call Stack Traces**: Full backtrace showing where queries were executed
|
|
52
|
+
- **Object Allocations**: Optional tracking of object allocations per query
|
|
53
|
+
|
|
54
|
+
### Query Performance Analysis
|
|
55
|
+
- **Slow Query Detection**: Configurable threshold for identifying slow queries
|
|
56
|
+
- **EXPLAIN ANALYZE**: Automatic execution plan capture for slow queries
|
|
57
|
+
- **Background Execution**: EXPLAIN ANALYZE runs in separate thread (non-blocking)
|
|
58
|
+
- **Multi-Database Support**: Works with PostgreSQL, MySQL, SQLite, and others
|
|
59
|
+
- **Smart Filtering**: Automatically skips transaction queries (BEGIN, COMMIT, ROLLBACK)
|
|
60
|
+
- **Execution Plan Details**:
|
|
61
|
+
- PostgreSQL: Full EXPLAIN ANALYZE with buffer usage statistics
|
|
62
|
+
- MySQL: EXPLAIN ANALYZE with actual execution times
|
|
63
|
+
- SQLite: EXPLAIN QUERY PLAN output
|
|
64
|
+
- **Query Optimization Insights**: Helps identify missing indexes, full table scans, JOIN issues
|
|
65
|
+
|
|
66
|
+
## View Rendering Tracking
|
|
67
|
+
|
|
68
|
+
### View Performance
|
|
69
|
+
- **Template Rendering**: Tracks main template rendering
|
|
70
|
+
- **Partial Rendering**: Tracks partial template rendering with cache key information
|
|
71
|
+
- **Collection Rendering**: Tracks collection rendering (partials in loops)
|
|
72
|
+
- **Rendering Duration**: Precise timing for each view component
|
|
73
|
+
- **Virtual Path Tracking**: Tracks view virtual paths
|
|
74
|
+
- **Layout Information**: Captures layout usage
|
|
75
|
+
|
|
76
|
+
### View Analysis
|
|
77
|
+
- **Slow View Detection**: Identifies the slowest rendering views
|
|
78
|
+
- **Frequency Analysis**: Tracks most frequently rendered views
|
|
79
|
+
- **Cache Hit Rate**: Calculates cache hit rates for partials
|
|
80
|
+
- **Collection Cache Analysis**: Tracks cache hit rates for collection rendering
|
|
81
|
+
- **Performance Metrics**:
|
|
82
|
+
- Total views rendered per request
|
|
83
|
+
- Total view rendering duration
|
|
84
|
+
- Average view rendering duration
|
|
85
|
+
- Breakdown by view type (template, partial, collection)
|
|
86
|
+
|
|
87
|
+
## Memory Tracking & Leak Detection
|
|
88
|
+
|
|
89
|
+
### Lightweight Memory Tracking (Default)
|
|
90
|
+
- **Memory Usage Monitoring**: Tracks memory consumption per request using GC stats
|
|
91
|
+
- **Memory Growth Tracking**: Measures memory growth during request processing
|
|
92
|
+
- **GC Statistics**: Tracks garbage collection count and heap pages
|
|
93
|
+
- **Minimal Performance Impact**: ~0.1ms overhead per request
|
|
94
|
+
- **Memory Before/After**: Captures memory state at request start and end
|
|
95
|
+
|
|
96
|
+
### Detailed Allocation Tracking (Optional)
|
|
97
|
+
- **Object Allocation Tracking**: Detailed tracking of object allocations (disabled by default)
|
|
98
|
+
- **Allocation Sampling**: Configurable sampling rate for allocations
|
|
99
|
+
- **Large Object Detection**: Identifies objects larger than 1MB threshold
|
|
100
|
+
- **Memory Snapshots**: Periodic memory snapshots during request processing
|
|
101
|
+
- **Object Count Tracking**: Tracks object counts before and after requests
|
|
102
|
+
- **Performance Impact**: ~2-5ms overhead per request (only when enabled)
|
|
103
|
+
|
|
104
|
+
### Memory Leak Detection
|
|
105
|
+
- **Pattern Detection**: Detects growing memory patterns over time
|
|
106
|
+
- **GC Efficiency Analysis**: Monitors garbage collection effectiveness
|
|
107
|
+
- **Heap Page Tracking**: Tracks heap page growth
|
|
108
|
+
- **Request Correlation**: Correlates memory growth with specific controllers/actions
|
|
109
|
+
|
|
110
|
+
## Background Job Tracking
|
|
111
|
+
|
|
112
|
+
### Job Execution Monitoring
|
|
113
|
+
- **ActiveJob Integration**: Automatic tracking when ActiveJob is available
|
|
114
|
+
- **Job Class Tracking**: Tracks job class names
|
|
115
|
+
- **Job ID**: Captures unique job identifiers
|
|
116
|
+
- **Queue Name**: Tracks which queue processed the job
|
|
117
|
+
- **Job Arguments**: Captures job arguments (with sensitive data filtering)
|
|
118
|
+
- **Duration Measurement**: Precise job execution time in milliseconds
|
|
119
|
+
- **Status Tracking**: Tracks job status (completed or failed)
|
|
120
|
+
|
|
121
|
+
### Job Error Tracking
|
|
122
|
+
- **Exception Capture**: Captures exceptions from failed jobs
|
|
123
|
+
- **Exception Class**: Tracks exception class names
|
|
124
|
+
- **Exception Messages**: Captures exception messages (truncated to 1000 chars)
|
|
125
|
+
- **Backtraces**: Full exception backtraces (first 50 lines)
|
|
126
|
+
- **SQL Query Context**: Includes SQL queries executed during failed jobs
|
|
127
|
+
- **Memory Context**: Includes memory usage during job execution
|
|
128
|
+
|
|
129
|
+
### Job SQL Tracking
|
|
130
|
+
- **SQL Query Tracking**: Tracks all SQL queries executed during job processing
|
|
131
|
+
- **Query Details**: Same detailed SQL tracking as request tracking
|
|
132
|
+
- **Query Context**: Full context of database operations in background jobs
|
|
133
|
+
|
|
134
|
+
## Cache Tracking
|
|
135
|
+
|
|
136
|
+
### Cache Operations
|
|
137
|
+
- **Read Operations**: Tracks cache read operations
|
|
138
|
+
- **Write Operations**: Tracks cache write operations
|
|
139
|
+
- **Delete Operations**: Tracks cache delete operations
|
|
140
|
+
- **Existence Checks**: Tracks cache existence checks
|
|
141
|
+
- **Fetch Operations**: Tracks cache fetch with hit/miss detection
|
|
142
|
+
- **Multi-Read Operations**: Tracks cache read_multi operations
|
|
143
|
+
- **Multi-Write Operations**: Tracks cache write_multi operations
|
|
144
|
+
- **Cache Generation**: Tracks cache generation events
|
|
145
|
+
|
|
146
|
+
### Cache Analysis
|
|
147
|
+
- **Cache Hit Detection**: Identifies cache hits vs misses
|
|
148
|
+
- **Cache Key Tracking**: Tracks cache keys (truncated to 200 chars)
|
|
149
|
+
- **Store Information**: Identifies which cache store was used
|
|
150
|
+
- **Namespace Tracking**: Tracks cache namespaces
|
|
151
|
+
- **Duration Measurement**: Precise timing for each cache operation
|
|
152
|
+
- **Hit Rate Calculation**: Calculates cache hit rates per request
|
|
153
|
+
|
|
154
|
+
## Redis Tracking
|
|
155
|
+
|
|
156
|
+
### Redis Command Tracking
|
|
157
|
+
- **Command Monitoring**: Tracks all Redis commands executed
|
|
158
|
+
- **Command Name**: Captures Redis command names (GET, SET, etc.)
|
|
159
|
+
- **Key Tracking**: Tracks Redis keys (truncated to 200 chars)
|
|
160
|
+
- **Argument Count**: Tracks number of arguments per command
|
|
161
|
+
- **Database Selection**: Tracks which Redis database is used
|
|
162
|
+
- **Duration Measurement**: Precise timing for each Redis command
|
|
163
|
+
- **Error Tracking**: Captures Redis command errors
|
|
164
|
+
|
|
165
|
+
### Advanced Redis Features
|
|
166
|
+
- **Pipeline Support**: Tracks Redis pipeline operations with command counts
|
|
167
|
+
- **Multi/Transaction Support**: Tracks Redis MULTI/EXEC transactions
|
|
168
|
+
- **ActiveSupport Integration**: Subscribes to ActiveSupport::Notifications for Redis events
|
|
169
|
+
- **Client Instrumentation**: Direct instrumentation of Redis::Client for comprehensive coverage
|
|
170
|
+
|
|
171
|
+
## Error Tracking
|
|
172
|
+
|
|
173
|
+
### Exception Handling
|
|
174
|
+
- **Automatic Exception Capture**: Captures exceptions from controller actions
|
|
175
|
+
- **Exception Class**: Tracks exception class names
|
|
176
|
+
- **Exception Messages**: Captures exception messages (truncated to 1000 chars)
|
|
177
|
+
- **Full Backtraces**: Captures complete exception backtraces (first 50 lines)
|
|
178
|
+
- **Request Context**: Includes full request context with exceptions
|
|
179
|
+
- **Error Flagging**: Errors are marked and always tracked (even with sampling)
|
|
180
|
+
|
|
181
|
+
### Error Context
|
|
182
|
+
- **Controller/Action**: Identifies where the error occurred
|
|
183
|
+
- **Request Parameters**: Includes request parameters at time of error
|
|
184
|
+
- **User Information**: Includes user ID if available
|
|
185
|
+
- **SQL Queries**: Includes SQL queries executed before error
|
|
186
|
+
- **Memory State**: Includes memory usage at time of error
|
|
187
|
+
- **Log Messages**: Includes application logs captured during request
|
|
188
|
+
|
|
189
|
+
## HTTP Instrumentation
|
|
190
|
+
|
|
191
|
+
### Outgoing HTTP Tracking
|
|
192
|
+
- **HTTP Request Tracking**: Tracks outgoing HTTP requests (via middleware)
|
|
193
|
+
- **Request Context**: Captures HTTP request details
|
|
194
|
+
- **Response Context**: Captures HTTP response details
|
|
195
|
+
- **Duration Measurement**: Tracks HTTP request duration
|
|
196
|
+
|
|
197
|
+
## Configuration & Flexibility
|
|
198
|
+
|
|
199
|
+
### Configuration Options
|
|
200
|
+
- **API Key Management**: Multiple sources (config, Rails credentials, ENV)
|
|
201
|
+
- **Endpoint Configuration**: Configurable endpoint URL
|
|
202
|
+
- **Timeout Settings**: Configurable open and read timeouts
|
|
203
|
+
- **Enable/Disable Toggle**: Can be enabled/disabled via configuration
|
|
204
|
+
- **Environment Detection**: Automatic Rails environment detection
|
|
205
|
+
|
|
206
|
+
### Circuit Breaker Configuration
|
|
207
|
+
- **Failure Threshold**: Configurable failure threshold (default: 3)
|
|
208
|
+
- **Recovery Timeout**: Configurable recovery timeout (default: 60 seconds)
|
|
209
|
+
- **Retry Timeout**: Configurable retry timeout (default: 300 seconds)
|
|
210
|
+
- **Enable/Disable**: Can enable/disable circuit breaker
|
|
211
|
+
|
|
212
|
+
### Memory Tracking Configuration
|
|
213
|
+
- **Memory Tracking Toggle**: Enable/disable memory tracking
|
|
214
|
+
- **Allocation Tracking Toggle**: Enable/disable detailed allocation tracking
|
|
215
|
+
- **Sampling Configuration**: Configurable request sampling rate
|
|
216
|
+
|
|
217
|
+
### Query Analysis Configuration
|
|
218
|
+
- **Slow Query Threshold**: Configurable threshold in milliseconds (default: 500ms)
|
|
219
|
+
- **EXPLAIN ANALYZE Toggle**: Enable/disable automatic EXPLAIN ANALYZE
|
|
220
|
+
|
|
221
|
+
## Data Safety & Privacy
|
|
222
|
+
|
|
223
|
+
### Data Sanitization
|
|
224
|
+
- **SQL Sanitization**: Automatically sanitizes SQL queries
|
|
225
|
+
- **Parameter Filtering**: Filters sensitive parameters (password, token, secret, key)
|
|
226
|
+
- **Argument Truncation**: Limits and truncates job arguments
|
|
227
|
+
- **Key Truncation**: Truncates cache and Redis keys to 200 characters
|
|
228
|
+
- **Value Truncation**: Recursively truncates nested values to prevent huge payloads
|
|
229
|
+
- **String Limits**: Limits string values (e.g., user agent to 200 chars, messages to 1000 chars)
|
|
230
|
+
|
|
231
|
+
### Data Limits
|
|
232
|
+
- **Array Limits**: Limits array sizes (e.g., first 10 job arguments, first 5 array elements)
|
|
233
|
+
- **Hash Limits**: Limits hash key counts (e.g., first 20 hash keys, first 30 params)
|
|
234
|
+
- **Backtrace Limits**: Limits backtraces to first 50 lines
|
|
235
|
+
- **Allocation Limits**: Limits allocations tracked per request (max 1000)
|
|
236
|
+
|
|
237
|
+
## Performance & Reliability
|
|
238
|
+
|
|
239
|
+
### Performance Optimizations
|
|
240
|
+
- **Asynchronous Posting**: Non-blocking HTTP requests
|
|
241
|
+
- **Lightweight Default Mode**: Minimal overhead in default configuration
|
|
242
|
+
- **Sampling Support**: Reduces data volume for high-traffic applications
|
|
243
|
+
- **Thread-Local Storage**: Efficient per-request data collection
|
|
244
|
+
- **Background EXPLAIN**: EXPLAIN ANALYZE runs in background thread
|
|
245
|
+
|
|
246
|
+
### Reliability Features
|
|
247
|
+
- **Circuit Breaker**: Prevents cascading failures
|
|
248
|
+
- **Error Handling**: Comprehensive error handling to prevent instrumentation failures
|
|
249
|
+
- **Graceful Degradation**: Continues working even if some features fail
|
|
250
|
+
- **Timeout Protection**: Configurable timeouts prevent hanging requests
|
|
251
|
+
|
|
252
|
+
## Integration & Compatibility
|
|
253
|
+
|
|
254
|
+
### Framework Support
|
|
255
|
+
- **Rails Integration**: Full Rails integration via Railtie
|
|
256
|
+
- **ActiveSupport Notifications**: Uses ActiveSupport::Notifications for event subscription
|
|
257
|
+
- **ActiveRecord Integration**: Tracks ActiveRecord SQL queries
|
|
258
|
+
- **ActiveJob Integration**: Tracks ActiveJob background jobs
|
|
259
|
+
- **ActionView Integration**: Tracks ActionView rendering
|
|
260
|
+
|
|
261
|
+
### Database Support
|
|
262
|
+
- **PostgreSQL**: Full support with EXPLAIN ANALYZE
|
|
263
|
+
- **MySQL**: Full support with EXPLAIN ANALYZE
|
|
264
|
+
- **SQLite**: Full support with EXPLAIN QUERY PLAN
|
|
265
|
+
- **Other Databases**: Basic support with standard EXPLAIN
|
|
266
|
+
|
|
267
|
+
### Cache Store Support
|
|
268
|
+
- **All Cache Stores**: Works with any Rails cache store
|
|
269
|
+
- **Multi-Store Support**: Tracks cache operations across different stores
|
|
270
|
+
|
|
271
|
+
### Redis Support
|
|
272
|
+
- **Redis Gem**: Works with redis gem
|
|
273
|
+
- **Client Instrumentation**: Direct instrumentation of Redis::Client
|
|
274
|
+
- **Pipeline Support**: Tracks Redis pipelines
|
|
275
|
+
- **Transaction Support**: Tracks Redis MULTI/EXEC transactions
|
|
276
|
+
|
|
277
|
+
## Logging & Debugging
|
|
278
|
+
|
|
279
|
+
### Application Logging
|
|
280
|
+
- **Log Capture**: Captures application logs during request processing
|
|
281
|
+
- **Log Context**: Includes logs in metric payloads
|
|
282
|
+
- **Debug Logging**: Optional debug logging for skipped requests
|
|
283
|
+
|
|
284
|
+
## Deployment & Environment
|
|
285
|
+
|
|
286
|
+
### Deploy Tracking
|
|
287
|
+
- **Deploy ID Resolution**: Multiple sources for deploy identification (`Configuration#deploy_id=` wins when set, then ENV in `Configuration::DEPLOY_REVISION_ENV_KEYS` order—including `DEAD_BRO_DEPLOY_ID`, git/CI vars, `DD_VERSION`, etc.), otherwise a **per-process UUID** (fine for single dyno/process; unusable alone for fleets like ECS replicas)
|
|
288
|
+
- **Revision Tracking**: Includes deploy/revision ID in all metric payloads
|
|
289
|
+
|
|
290
|
+
### Environment Support
|
|
291
|
+
- **Rails Environment**: Automatic Rails environment detection
|
|
292
|
+
- **Rack Environment**: Fallback to RACK_ENV or RAILS_ENV
|
|
293
|
+
- **Environment Context**: Includes environment in all metric payloads
|
|
294
|
+
|
|
295
|
+
## Data Collection & Transmission
|
|
296
|
+
|
|
297
|
+
### Metric Payload Structure
|
|
298
|
+
- **Structured Data**: Well-structured JSON payloads
|
|
299
|
+
- **Event Names**: Descriptive event names for different metric types
|
|
300
|
+
- **Timestamp Tracking**: ISO8601 timestamps for all metrics
|
|
301
|
+
- **Metadata**: Rich metadata including environment, host, deploy ID
|
|
302
|
+
|
|
303
|
+
### HTTP Client
|
|
304
|
+
- **HTTPS Support**: Secure HTTPS communication
|
|
305
|
+
- **Bearer Token Auth**: API key authentication via Bearer tokens
|
|
306
|
+
- **JSON Encoding**: JSON-encoded payloads
|
|
307
|
+
- **Custom Headers**: Proper Content-Type and Authorization headers
|
|
308
|
+
|
|
309
|
+
## Comparison-Ready Features
|
|
310
|
+
|
|
311
|
+
### Unique Differentiators
|
|
312
|
+
1. **Automatic EXPLAIN ANALYZE**: Background execution plan capture for slow queries
|
|
313
|
+
2. **Lightweight Memory Tracking**: Low-overhead memory monitoring by default
|
|
314
|
+
3. **Comprehensive Cache Tracking**: Detailed cache operation tracking
|
|
315
|
+
4. **Redis Instrumentation**: Full Redis command tracking including pipelines
|
|
316
|
+
5. **View Rendering Analysis**: Detailed view performance analysis with cache hit rates
|
|
317
|
+
6. **Flexible Exclusion Rules**: Wildcard support for controller/job exclusion
|
|
318
|
+
7. **Request Sampling**: Configurable percentage-based sampling
|
|
319
|
+
8. **Circuit Breaker**: Built-in resilience for APM endpoint failures
|
|
320
|
+
9. **Multi-Source Configuration**: Flexible configuration from multiple sources
|
|
321
|
+
10. **Deploy Tracking**: Automatic deploy ID resolution from multiple sources
|
|
322
|
+
|
|
323
|
+
### Standard APM Features
|
|
324
|
+
- Request/response tracking
|
|
325
|
+
- SQL query tracking
|
|
326
|
+
- Error tracking
|
|
327
|
+
- Background job tracking
|
|
328
|
+
- Memory tracking
|
|
329
|
+
- Performance metrics
|
|
330
|
+
- Exception handling
|
|
331
|
+
- User context
|
|
332
|
+
- Environment tracking
|
|
333
|
+
|
data/README.md
CHANGED
|
@@ -94,11 +94,11 @@ DeadBro can automatically capture the query plan (`EXPLAIN`) of slow SELECT quer
|
|
|
94
94
|
|
|
95
95
|
### Configuration
|
|
96
96
|
|
|
97
|
-
- **`
|
|
97
|
+
- **`explain_enabled`** (default: `false`) — this is the one setting that must be turned on in Ruby; the dashboard toggle can only turn capture *off*, never on, as an extra safety rail:
|
|
98
98
|
|
|
99
99
|
```ruby
|
|
100
100
|
DeadBro.configure do |config|
|
|
101
|
-
config.
|
|
101
|
+
config.explain_enabled = true
|
|
102
102
|
end
|
|
103
103
|
```
|
|
104
104
|
|
|
@@ -12,8 +12,12 @@ module DeadBro
|
|
|
12
12
|
# Local-only opt-in for EXPLAIN plan capture on slow queries. The gem runs
|
|
13
13
|
# DB statements (plan-only EXPLAIN) when this is on, so it must be enabled
|
|
14
14
|
# in the app's own config — remote settings can only turn it OFF, never on
|
|
15
|
-
# (see apply_remote_settings). Effective state: #
|
|
16
|
-
attr_accessor :
|
|
15
|
+
# (see apply_remote_settings). Effective state: #explain_active?
|
|
16
|
+
attr_accessor :explain_enabled
|
|
17
|
+
|
|
18
|
+
# Compatibility for initializers written before the EXPLAIN option was renamed.
|
|
19
|
+
alias_method :explain_analyze_enabled, :explain_enabled
|
|
20
|
+
alias_method :explain_analyze_enabled=, :explain_enabled=
|
|
17
21
|
|
|
18
22
|
# Remote-managed settings (overwritten by backend JSON `settings` on successful API responses)
|
|
19
23
|
attr_accessor :memory_tracking_enabled, :allocation_tracking_enabled, :allocation_sample_rate,
|
|
@@ -66,7 +70,7 @@ module DeadBro
|
|
|
66
70
|
|
|
67
71
|
REMOTE_SETTING_KEYS = %w[
|
|
68
72
|
enabled sample_rate memory_tracking_enabled allocation_tracking_enabled allocation_sample_rate
|
|
69
|
-
|
|
73
|
+
explain_enabled slow_query_threshold_ms max_sql_queries_to_send max_logs_to_send
|
|
70
74
|
watch_enabled excluded_controllers excluded_jobs exclusive_controllers exclusive_jobs
|
|
71
75
|
monitor_enabled enable_db_stats enable_process_stats enable_system_stats
|
|
72
76
|
sample_rates_by_type
|
|
@@ -94,11 +98,11 @@ module DeadBro
|
|
|
94
98
|
# sampling, allocation-source tracing) runs on this % of requests so the
|
|
95
99
|
# ~2-5ms overhead can be capped without turning the feature fully off.
|
|
96
100
|
@allocation_sample_rate = 100
|
|
97
|
-
@
|
|
101
|
+
@explain_enabled = false
|
|
98
102
|
# Remote kill switch for EXPLAIN capture. Defaults to true so a local
|
|
99
103
|
# opt-in works against older backends that never send the key; any
|
|
100
|
-
# response that sends
|
|
101
|
-
@
|
|
104
|
+
# response that sends explain_enabled: false turns capture off.
|
|
105
|
+
@remote_explain_enabled = true
|
|
102
106
|
@slow_query_threshold_ms = 500
|
|
103
107
|
@max_sql_queries_to_send = 500
|
|
104
108
|
@max_logs_to_send = 100
|
|
@@ -166,6 +170,12 @@ module DeadBro
|
|
|
166
170
|
def apply_remote_settings(hash)
|
|
167
171
|
return unless hash.is_a?(Hash)
|
|
168
172
|
|
|
173
|
+
# Accept older servers; the canonical key wins when both are present.
|
|
174
|
+
hash = hash.transform_keys(&:to_s)
|
|
175
|
+
if hash.key?("explain_analyze_enabled") && !hash.key?("explain_enabled")
|
|
176
|
+
hash["explain_enabled"] = hash["explain_analyze_enabled"]
|
|
177
|
+
end
|
|
178
|
+
|
|
169
179
|
@settings_mutex.synchronize do
|
|
170
180
|
hash.each do |key, value|
|
|
171
181
|
k = key.to_s
|
|
@@ -174,11 +184,11 @@ module DeadBro
|
|
|
174
184
|
case k
|
|
175
185
|
when "sample_rate", "allocation_sample_rate", "slow_query_threshold_ms", "max_sql_queries_to_send", "max_logs_to_send"
|
|
176
186
|
send(:"#{k}=", value.to_i)
|
|
177
|
-
when "
|
|
187
|
+
when "explain_enabled"
|
|
178
188
|
# EXPLAIN runs statements against the customer DB, so the backend
|
|
179
189
|
# must never be able to switch it on — only off. The local opt-in
|
|
180
|
-
# (
|
|
181
|
-
@
|
|
190
|
+
# (explain_enabled) stays untouched; see #explain_active?
|
|
191
|
+
@remote_explain_enabled = !!value
|
|
182
192
|
when "enabled", "memory_tracking_enabled", "allocation_tracking_enabled", "watch_enabled",
|
|
183
193
|
"monitor_enabled", "enable_db_stats", "enable_process_stats", "enable_system_stats"
|
|
184
194
|
send(:"#{k}=", !!value)
|
|
@@ -193,10 +203,12 @@ module DeadBro
|
|
|
193
203
|
|
|
194
204
|
# EXPLAIN capture requires BOTH the local opt-in and the remote flag —
|
|
195
205
|
# locally off means off no matter what the backend sends.
|
|
196
|
-
def
|
|
197
|
-
!!(@
|
|
206
|
+
def explain_active?
|
|
207
|
+
!!(@explain_enabled && @remote_explain_enabled)
|
|
198
208
|
end
|
|
199
209
|
|
|
210
|
+
alias_method :explain_analyze_active?, :explain_active?
|
|
211
|
+
|
|
200
212
|
def heartbeat_due?
|
|
201
213
|
return false if api_key.nil?
|
|
202
214
|
last_heartbeat_attempt_at.nil? || (Time.now.utc - last_heartbeat_attempt_at) >= HEARTBEAT_INTERVAL
|
|
@@ -9,7 +9,6 @@ end
|
|
|
9
9
|
module DeadBro
|
|
10
10
|
class JobSubscriber
|
|
11
11
|
JOB_EVENT_NAME = "perform.active_job"
|
|
12
|
-
JOB_EXCEPTION_EVENT_NAME = "exception.active_job"
|
|
13
12
|
|
|
14
13
|
def self.subscribe!(client: Client.new)
|
|
15
14
|
# Snap GC state before the job runs so stop_request_tracking gets a valid diff
|
|
@@ -19,7 +18,28 @@ module DeadBro
|
|
|
19
18
|
rescue
|
|
20
19
|
end
|
|
21
20
|
|
|
22
|
-
# Track job execution
|
|
21
|
+
# Track job execution — success AND failure both land here. ActiveJob wraps
|
|
22
|
+
# perform with `instrument(:perform) { super }` (see
|
|
23
|
+
# ActiveJob::Instrumentation#instrument), and AS::Notifications.instrument
|
|
24
|
+
# still runs its "finish" listeners (with :exception / :exception_object set
|
|
25
|
+
# in the payload) when the block raises, then re-raises. There is no separate
|
|
26
|
+
# "exception.active_job" event anywhere in Rails/ActiveJob — a prior version
|
|
27
|
+
# of this file subscribed to one, so it never fired; the job wasn't dropped,
|
|
28
|
+
# it fired *this* event and built status: "completed" unconditionally,
|
|
29
|
+
# silently reporting every failing job as a success. Branching on
|
|
30
|
+
# data[:exception_object] here is the only place a job failure can actually
|
|
31
|
+
# be detected.
|
|
32
|
+
#
|
|
33
|
+
# Known gap: this only sees an exception that escapes perform_now uncaught.
|
|
34
|
+
# ActiveJob::Base#perform_now (Execution) rescues internally and hands off to
|
|
35
|
+
# rescue_with_handler — which is exactly what retry_on/discard_on/rescue_from
|
|
36
|
+
# are built on (ActiveJob::Exceptions) — before Instrumentation's `super`
|
|
37
|
+
# even returns. A handled retry or discard returns normally with no
|
|
38
|
+
# exception attached, so perform.active_job still reports status:
|
|
39
|
+
# "completed" for every handled attempt; only a retry_on job's final,
|
|
40
|
+
# unhandled raise (attempts exhausted, no block given) is visible here.
|
|
41
|
+
# Covering handled attempts would need separate instrumentation on
|
|
42
|
+
# enqueue_retry.active_job / retry_stopped.active_job / discard.active_job.
|
|
23
43
|
ActiveSupport::Notifications.subscribe(JOB_EVENT_NAME) do |name, started, finished, _unique_id, data|
|
|
24
44
|
begin
|
|
25
45
|
if DeadBro.configuration.skip_tracking?
|
|
@@ -40,12 +60,14 @@ module DeadBro
|
|
|
40
60
|
rescue
|
|
41
61
|
end
|
|
42
62
|
|
|
63
|
+
exception = data[:exception_object]
|
|
64
|
+
has_error = !exception.nil?
|
|
65
|
+
|
|
43
66
|
# Skip out via sampling before we build any payload — jobs can be chatty
|
|
44
67
|
# enough that even the "cheap" stop/analyze work matters under load.
|
|
45
|
-
#
|
|
46
|
-
# always sends errors with force: true.
|
|
68
|
+
# Errors always ship regardless of sampling, matching Subscriber's web path.
|
|
47
69
|
job_type_key = "#{job_class_name}#perform"
|
|
48
|
-
unless DeadBro.configuration.should_sample?(job_type_key)
|
|
70
|
+
unless has_error || DeadBro.configuration.should_sample?(job_type_key)
|
|
49
71
|
drain_job_tracking
|
|
50
72
|
next
|
|
51
73
|
end
|
|
@@ -71,6 +93,7 @@ module DeadBro
|
|
|
71
93
|
|
|
72
94
|
# Get SQL queries executed during this job
|
|
73
95
|
sql_queries = DeadBro::SqlSubscriber.stop_request_tracking
|
|
96
|
+
transaction_events = DeadBro::SqlSubscriber.last_transaction_events
|
|
74
97
|
dependency_events = job_dependency_payload
|
|
75
98
|
db_connection_stats = defined?(DeadBro::DbConnectionSubscriber) ? DeadBro::DbConnectionSubscriber.stop_request_tracking : {}
|
|
76
99
|
gc_pressure = defined?(DeadBro::GcTracker) ? DeadBro::GcTracker.stop_request_tracking : {}
|
|
@@ -117,8 +140,9 @@ module DeadBro
|
|
|
117
140
|
db_connection_checkouts: db_connection_stats[:checkouts],
|
|
118
141
|
gc_pressure: gc_pressure,
|
|
119
142
|
ar_instantiation_count: ar_instantiation_count,
|
|
120
|
-
status: "completed",
|
|
143
|
+
status: has_error ? "failed" : "completed",
|
|
121
144
|
sql_queries: sql_queries,
|
|
145
|
+
transaction_events: transaction_events,
|
|
122
146
|
rails_env: DeadBro.env,
|
|
123
147
|
host: DeadBro.safe_hostname,
|
|
124
148
|
process_kind: DeadBro.process_kind,
|
|
@@ -130,115 +154,19 @@ module DeadBro
|
|
|
130
154
|
logs: DeadBro.logger.logs
|
|
131
155
|
}.merge(dependency_events)
|
|
132
156
|
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
begin
|
|
141
|
-
if DeadBro.configuration.skip_tracking?
|
|
142
|
-
drain_job_tracking
|
|
143
|
-
next
|
|
144
|
-
end
|
|
145
|
-
|
|
146
|
-
job_class_name = data[:job].class.name
|
|
147
|
-
if DeadBro.configuration.excluded_job?(job_class_name)
|
|
148
|
-
next
|
|
149
|
-
end
|
|
150
|
-
# If exclusive_jobs is defined and not empty, only track matching jobs
|
|
151
|
-
unless DeadBro.configuration.exclusive_job?(job_class_name)
|
|
152
|
-
next
|
|
153
|
-
end
|
|
154
|
-
rescue
|
|
155
|
-
end
|
|
156
|
-
|
|
157
|
-
duration_ms = ((finished - started) * 1000.0).round(2)
|
|
158
|
-
exception = data[:exception_object]
|
|
159
|
-
queue_duration_ms = job_queue_duration_ms(data[:job], started)
|
|
160
|
-
|
|
161
|
-
# Ensure tracking was started (fallback if perform_start.active_job didn't fire)
|
|
162
|
-
unless DeadBro::SqlSubscriber.tracking_active?
|
|
163
|
-
DeadBro.logger.clear
|
|
164
|
-
Thread.current[DeadBro::TRACKING_START_TIME_KEY] = Time.now
|
|
165
|
-
DeadBro::SqlSubscriber.start_request_tracking
|
|
166
|
-
start_job_dependency_tracking
|
|
167
|
-
DeadBro::DbConnectionSubscriber.start_request_tracking if defined?(DeadBro::DbConnectionSubscriber)
|
|
168
|
-
DeadBro::WatchTracker.start_request_tracking if defined?(DeadBro::WatchTracker)
|
|
169
|
-
if DeadBro.configuration.allocation_tracking_enabled && defined?(DeadBro::MemoryTrackingSubscriber)
|
|
170
|
-
DeadBro::MemoryTrackingSubscriber.start_request_tracking
|
|
171
|
-
else
|
|
172
|
-
DeadBro::LightweightMemoryTracker.start_request_tracking if defined?(DeadBro::LightweightMemoryTracker)
|
|
173
|
-
end
|
|
174
|
-
end
|
|
175
|
-
|
|
176
|
-
# Get SQL queries executed during this job
|
|
177
|
-
sql_queries = DeadBro::SqlSubscriber.stop_request_tracking
|
|
178
|
-
dependency_events = job_dependency_payload
|
|
179
|
-
db_connection_stats = defined?(DeadBro::DbConnectionSubscriber) ? DeadBro::DbConnectionSubscriber.stop_request_tracking : {}
|
|
180
|
-
gc_pressure = defined?(DeadBro::GcTracker) ? DeadBro::GcTracker.stop_request_tracking : {}
|
|
181
|
-
ar_instantiation_count = defined?(DeadBro::ArObjectTracker) ? DeadBro::ArObjectTracker.stop_request_tracking : nil
|
|
182
|
-
watch_events = defined?(DeadBro::WatchTracker) ? DeadBro::WatchTracker.stop_request_tracking : []
|
|
183
|
-
|
|
184
|
-
# Stop memory tracking and get collected memory data
|
|
185
|
-
if DeadBro.configuration.allocation_tracking_enabled && defined?(DeadBro::MemoryTrackingSubscriber)
|
|
186
|
-
detailed_memory = DeadBro::MemoryTrackingSubscriber.stop_request_tracking
|
|
187
|
-
memory_performance = DeadBro::MemoryTrackingSubscriber.analyze_memory_performance(detailed_memory)
|
|
188
|
-
# Keep memory_events compact and user-friendly (no large raw arrays)
|
|
189
|
-
memory_events = {
|
|
190
|
-
memory_before: detailed_memory[:memory_before],
|
|
191
|
-
memory_after: detailed_memory[:memory_after],
|
|
192
|
-
duration_seconds: detailed_memory[:duration_seconds],
|
|
193
|
-
allocations_count: (detailed_memory[:allocations] || []).length,
|
|
194
|
-
memory_snapshots_count: (detailed_memory[:memory_snapshots] || []).length,
|
|
195
|
-
large_objects_count: (detailed_memory[:large_objects] || []).length
|
|
196
|
-
}
|
|
197
|
-
else
|
|
198
|
-
lightweight_memory = DeadBro::LightweightMemoryTracker.stop_request_tracking
|
|
199
|
-
# Separate raw readings from derived performance metrics to avoid duplicating data
|
|
200
|
-
memory_events = {
|
|
201
|
-
memory_before: lightweight_memory[:memory_before],
|
|
202
|
-
memory_after: lightweight_memory[:memory_after]
|
|
203
|
-
}
|
|
204
|
-
memory_performance = {
|
|
205
|
-
memory_growth_mb: lightweight_memory[:memory_growth_mb],
|
|
206
|
-
gc_count_increase: lightweight_memory[:gc_count_increase],
|
|
207
|
-
heap_pages_increase: lightweight_memory[:heap_pages_increase],
|
|
208
|
-
duration_seconds: lightweight_memory[:duration_seconds]
|
|
209
|
-
}
|
|
157
|
+
if has_error
|
|
158
|
+
payload[:exception_class] = exception.class.name
|
|
159
|
+
payload[:message] = exception.message.to_s[0, 1000]
|
|
160
|
+
payload[:backtrace] = Array(exception.backtrace).first(50)
|
|
161
|
+
payload[:fingerprint] = DeadBro::Subscriber.compute_error_fingerprint(exception)
|
|
162
|
+
payload[:cause_chain] = DeadBro::Subscriber.build_cause_chain(exception)
|
|
163
|
+
payload[:error] = true
|
|
210
164
|
end
|
|
211
165
|
|
|
212
|
-
|
|
213
|
-
|
|
214
|
-
|
|
215
|
-
|
|
216
|
-
arguments: safe_arguments(data[:job].arguments),
|
|
217
|
-
started_at: started.utc.iso8601(3),
|
|
218
|
-
duration_ms: duration_ms,
|
|
219
|
-
queue_duration_ms: queue_duration_ms,
|
|
220
|
-
db_connection_wait_ms: db_connection_stats[:wait_ms],
|
|
221
|
-
db_connection_checkouts: db_connection_stats[:checkouts],
|
|
222
|
-
gc_pressure: gc_pressure,
|
|
223
|
-
ar_instantiation_count: ar_instantiation_count,
|
|
224
|
-
status: "failed",
|
|
225
|
-
sql_queries: sql_queries,
|
|
226
|
-
exception_class: exception&.class&.name,
|
|
227
|
-
message: exception&.message&.to_s&.[](0, 1000),
|
|
228
|
-
backtrace: Array(exception&.backtrace).first(50),
|
|
229
|
-
rails_env: DeadBro.env,
|
|
230
|
-
host: DeadBro.safe_hostname,
|
|
231
|
-
process_kind: DeadBro.process_kind,
|
|
232
|
-
memory_usage: memory_usage_mb,
|
|
233
|
-
gc_stats: gc_stats,
|
|
234
|
-
memory_events: memory_events,
|
|
235
|
-
memory_performance: memory_performance,
|
|
236
|
-
watch_events: watch_events,
|
|
237
|
-
logs: DeadBro.logger.logs
|
|
238
|
-
}.merge(dependency_events)
|
|
239
|
-
|
|
240
|
-
event_name = exception&.class&.name || "ActiveJob::Exception"
|
|
241
|
-
client.post_metric(event_name: event_name, payload: payload, force: true)
|
|
166
|
+
# force: true — errors always ship, bypassing sampling by design; for
|
|
167
|
+
# completions the sampling decision above already accounted for any
|
|
168
|
+
# per-job-type override, so client#post_metric must not re-roll it globally.
|
|
169
|
+
client.post_metric(event_name: name, payload: payload, force: true)
|
|
242
170
|
end
|
|
243
171
|
rescue
|
|
244
172
|
# Never raise from instrumentation install
|
|
@@ -248,6 +176,7 @@ module DeadBro
|
|
|
248
176
|
# build a payload (excluded job / sampled out). Matches Subscriber.drain_request_tracking.
|
|
249
177
|
def self.drain_job_tracking
|
|
250
178
|
# wait_for_explains: false — result is discarded, don't block on pending plans.
|
|
179
|
+
# stop_request_tracking also pops the transaction-events stack (see its comment).
|
|
251
180
|
DeadBro::SqlSubscriber.stop_request_tracking(wait_for_explains: false) if defined?(DeadBro::SqlSubscriber)
|
|
252
181
|
Thread.current[:dead_bro_http_events] = nil
|
|
253
182
|
DeadBro::CacheSubscriber.stop_request_tracking if defined?(DeadBro::CacheSubscriber)
|
|
@@ -16,7 +16,11 @@ module DeadBro
|
|
|
16
16
|
THREAD_LOCAL_EXPLAIN_PENDING_KEY = :dead_bro_explain_pending
|
|
17
17
|
THREAD_LOCAL_CALL_COUNTS_KEY = :dead_bro_sql_call_counts
|
|
18
18
|
THREAD_LOCAL_AGGREGATES_KEY = :dead_bro_sql_aggregates
|
|
19
|
+
THREAD_LOCAL_TXN_EVENTS_KEY = :dead_bro_sql_txn_events
|
|
19
20
|
MAX_TRACKED_QUERIES = 1000
|
|
21
|
+
# Transaction control breadcrumbs (BEGIN/COMMIT/ROLLBACK/...) are cheap and rare
|
|
22
|
+
# compared to queries, but a pathological retry loop could still spam them.
|
|
23
|
+
MAX_TRACKED_TXN_EVENTS = 200
|
|
20
24
|
|
|
21
25
|
# Number of identical queries within one request that triggers N+1 detection.
|
|
22
26
|
N_PLUS_ONE_THRESHOLD = 5
|
|
@@ -37,6 +41,43 @@ module DeadBro
|
|
|
37
41
|
SANITIZE_SKIP_SENSITIVE_WHEN_NO_KEYWORDS = /password|token|secret|key|ssn|credit_card/i
|
|
38
42
|
SANITIZE_SKIP_WHERE_WHEN_NO_KEYWORD = /WHERE/i
|
|
39
43
|
|
|
44
|
+
# Rails' own adapters (MySQL, PostgreSQL, SQLite3, at least 7.1 through 8.1 —
|
|
45
|
+
# the versions checked directly in this repo) log every transaction-control
|
|
46
|
+
# statement under the single generic name "TRANSACTION"
|
|
47
|
+
# (`internal_execute("BEGIN", "TRANSACTION", ...)`, `internal_execute("COMMIT",
|
|
48
|
+
# "TRANSACTION", ...)`, etc. — see AbstractMysqlAdapter/PostgreSQL::DatabaseStatements/
|
|
49
|
+
# SQLite3::DatabaseStatements). "BEGIN"/"COMMIT"/"ROLLBACK"/"SAVEPOINT"/"RELEASE"
|
|
50
|
+
# as literal `name` values are kept here only as a defensive fallback for
|
|
51
|
+
# adapters/older Rails versions that might still emit them directly — the actual
|
|
52
|
+
# operation for a "TRANSACTION"-named event is derived from the SQL text itself
|
|
53
|
+
# (see transaction_operation_for below), since the name alone can't distinguish
|
|
54
|
+
# BEGIN from COMMIT from a SAVEPOINT.
|
|
55
|
+
TXN_CONTROL_NAMES = %w[BEGIN COMMIT ROLLBACK SAVEPOINT RELEASE TRANSACTION].freeze
|
|
56
|
+
|
|
57
|
+
TXN_OPERATION_FROM_SQL = [
|
|
58
|
+
[/\A\s*ROLLBACK\s+TO\s+SAVEPOINT/i, "ROLLBACK TO SAVEPOINT"],
|
|
59
|
+
[/\A\s*RELEASE\s+SAVEPOINT/i, "RELEASE SAVEPOINT"],
|
|
60
|
+
[/\A\s*SAVEPOINT/i, "SAVEPOINT"],
|
|
61
|
+
[/\A\s*BEGIN/i, "BEGIN"],
|
|
62
|
+
[/\A\s*COMMIT/i, "COMMIT"],
|
|
63
|
+
[/\A\s*ROLLBACK/i, "ROLLBACK"]
|
|
64
|
+
].freeze
|
|
65
|
+
|
|
66
|
+
# data[:name] is only ever "TRANSACTION" in current Rails — this derives the
|
|
67
|
+
# actual verb (BEGIN/COMMIT/ROLLBACK/SAVEPOINT/RELEASE SAVEPOINT/ROLLBACK TO
|
|
68
|
+
# SAVEPOINT) from the SQL text so breadcrumbs are still meaningful. A literal
|
|
69
|
+
# non-"TRANSACTION" name (older Rails/adapter) is passed through unchanged.
|
|
70
|
+
def self.transaction_operation_for(name, sql)
|
|
71
|
+
return name unless name == "TRANSACTION"
|
|
72
|
+
sql_str = sql.to_s
|
|
73
|
+
TXN_OPERATION_FROM_SQL.each do |re, op|
|
|
74
|
+
return op if sql_str.match?(re)
|
|
75
|
+
end
|
|
76
|
+
name
|
|
77
|
+
rescue
|
|
78
|
+
name
|
|
79
|
+
end
|
|
80
|
+
|
|
40
81
|
# True when there is at least one active tracking context (e.g. for nested jobs).
|
|
41
82
|
def self.tracking_active?
|
|
42
83
|
stack = Thread.current[THREAD_LOCAL_KEY]
|
|
@@ -50,6 +91,13 @@ module DeadBro
|
|
|
50
91
|
stack.last
|
|
51
92
|
end
|
|
52
93
|
|
|
94
|
+
# Current transaction-events array (top of stack); nil if no active tracking.
|
|
95
|
+
def self.current_txn_events_array
|
|
96
|
+
stack = Thread.current[THREAD_LOCAL_TXN_EVENTS_KEY]
|
|
97
|
+
return nil unless stack.is_a?(Array) && stack.any?
|
|
98
|
+
stack.last
|
|
99
|
+
end
|
|
100
|
+
|
|
53
101
|
# Sum of SQL counts and durations recorded so far in the active tracking
|
|
54
102
|
# context. Used by WatchTracker to attribute SQL to a DeadBro.watch block.
|
|
55
103
|
def self.current_sql_metrics
|
|
@@ -96,6 +144,27 @@ module DeadBro
|
|
|
96
144
|
sql.to_s.downcase
|
|
97
145
|
end
|
|
98
146
|
|
|
147
|
+
# Records a lightweight breadcrumb for a transaction-control statement
|
|
148
|
+
# (BEGIN/COMMIT/ROLLBACK/SAVEPOINT/RELEASE) — enough to show, e.g., that the
|
|
149
|
+
# queries preceding an error ran inside a transaction that was rolled back,
|
|
150
|
+
# without the cost of treating it like a tracked query (no backtrace/EXPLAIN).
|
|
151
|
+
def self.record_transaction_event(name, started, finished)
|
|
152
|
+
current = current_txn_events_array
|
|
153
|
+
return unless current
|
|
154
|
+
return if current.length >= MAX_TRACKED_TXN_EVENTS
|
|
155
|
+
|
|
156
|
+
tracking_start = Thread.current[DeadBro::TRACKING_START_TIME_KEY]
|
|
157
|
+
start_offset_ms = tracking_start ? ((started - tracking_start) * 1000.0).round(2) : nil
|
|
158
|
+
|
|
159
|
+
current << {
|
|
160
|
+
event: name,
|
|
161
|
+
duration_ms: ((finished - started) * 1000.0).round(2),
|
|
162
|
+
start_offset_ms: start_offset_ms
|
|
163
|
+
}
|
|
164
|
+
rescue
|
|
165
|
+
nil
|
|
166
|
+
end
|
|
167
|
+
|
|
99
168
|
def self.subscribe!
|
|
100
169
|
# Subscribe with a start/finish listener to measure allocations per query
|
|
101
170
|
if ActiveSupport::Notifications.notifier.respond_to?(:subscribe)
|
|
@@ -106,7 +175,13 @@ module DeadBro
|
|
|
106
175
|
end
|
|
107
176
|
|
|
108
177
|
ActiveSupport::Notifications.subscribe(SQL_EVENT_NAME) do |name, started, finished, _unique_id, data|
|
|
109
|
-
next if data[:name] == "SCHEMA" || data[:name] == "CACHE"
|
|
178
|
+
next if data[:name] == "SCHEMA" || data[:name] == "CACHE"
|
|
179
|
+
|
|
180
|
+
if TXN_CONTROL_NAMES.include?(data[:name])
|
|
181
|
+
record_transaction_event(transaction_operation_for(data[:name], data[:sql]), started, finished)
|
|
182
|
+
next
|
|
183
|
+
end
|
|
184
|
+
|
|
110
185
|
# Only track queries that are part of the current request (top of stack for nested jobs)
|
|
111
186
|
current = current_queries_array
|
|
112
187
|
next unless current
|
|
@@ -208,6 +283,7 @@ module DeadBro
|
|
|
208
283
|
Thread.current[THREAD_LOCAL_EXPLAIN_PENDING_KEY] = []
|
|
209
284
|
(Thread.current[THREAD_LOCAL_CALL_COUNTS_KEY] ||= []) << {}
|
|
210
285
|
(Thread.current[THREAD_LOCAL_AGGREGATES_KEY] ||= []) << {}
|
|
286
|
+
(Thread.current[THREAD_LOCAL_TXN_EVENTS_KEY] ||= []) << []
|
|
211
287
|
end
|
|
212
288
|
|
|
213
289
|
# wait_for_explains: false on drain paths (excluded / sampled-out requests)
|
|
@@ -223,6 +299,21 @@ module DeadBro
|
|
|
223
299
|
cc_stack = Thread.current[THREAD_LOCAL_CALL_COUNTS_KEY]
|
|
224
300
|
cc_stack.pop if cc_stack.is_a?(Array) && cc_stack.any?
|
|
225
301
|
|
|
302
|
+
# Transaction-control breadcrumbs are popped here, as part of this method's
|
|
303
|
+
# existing lifecycle, rather than through a separate stop_transaction_tracking
|
|
304
|
+
# call. start_request_tracking has many callers across this gem (and its own
|
|
305
|
+
# specs) that only know about query tracking; requiring every one of them to
|
|
306
|
+
# also remember a second, separately-paired stop call is exactly how a frame
|
|
307
|
+
# gets pushed and never popped — leaking across every later request/job on
|
|
308
|
+
# the same thread. Piggybacking on stop_request_tracking, which is already
|
|
309
|
+
# reliably paired with every start_request_tracking call, means there is
|
|
310
|
+
# nothing new for any caller to remember. A caller that wants the events
|
|
311
|
+
# reads last_transaction_events immediately afterward.
|
|
312
|
+
txn_stack = Thread.current[THREAD_LOCAL_TXN_EVENTS_KEY]
|
|
313
|
+
Thread.current[:dead_bro_last_transaction_events] =
|
|
314
|
+
(txn_stack.is_a?(Array) && txn_stack.any?) ? txn_stack.pop : []
|
|
315
|
+
Thread.current[THREAD_LOCAL_TXN_EVENTS_KEY] = nil if txn_stack.nil? || txn_stack.empty?
|
|
316
|
+
|
|
226
317
|
# Fold any completed EXPLAIN plans from raw queries into their aggregate entry
|
|
227
318
|
raw_queries.each do |q|
|
|
228
319
|
next unless q[:explain_plan]
|
|
@@ -244,6 +335,17 @@ module DeadBro
|
|
|
244
335
|
aggregates_h.values.sort_by { |a| -a[:total_duration_ms] }
|
|
245
336
|
end
|
|
246
337
|
|
|
338
|
+
# The transaction-control breadcrumbs captured during the tracking window that
|
|
339
|
+
# this thread's most recent stop_request_tracking call just ended. Consumes
|
|
340
|
+
# (clears) the thread-local on read, not just on write: otherwise it would sit
|
|
341
|
+
# pinned in Thread.current between requests (a Puma/Sidekiq thread is reused),
|
|
342
|
+
# and — worse — any caller that reads it without an immediately-preceding
|
|
343
|
+
# stop_request_tracking would silently get the *previous* request's
|
|
344
|
+
# breadcrumbs instead of an empty result.
|
|
345
|
+
def self.last_transaction_events
|
|
346
|
+
Thread.current[:dead_bro_last_transaction_events].tap { Thread.current[:dead_bro_last_transaction_events] = nil } || []
|
|
347
|
+
end
|
|
348
|
+
|
|
247
349
|
# Upper bound on pending EXPLAIN threads per request. Each thread checks
|
|
248
350
|
# out an AR pool connection, so this must stay well below common pool
|
|
249
351
|
# sizes (Rails default is 5) to avoid starving the app under a storm.
|
|
@@ -304,7 +406,7 @@ module DeadBro
|
|
|
304
406
|
EXPLAIN_DML_KEYWORD_RE = /\b(INSERT|UPDATE|DELETE|MERGE)\b/i.freeze
|
|
305
407
|
|
|
306
408
|
def self.should_explain_query?(duration_ms, sql)
|
|
307
|
-
return false unless DeadBro.configuration.
|
|
409
|
+
return false unless DeadBro.configuration.explain_active?
|
|
308
410
|
return false if duration_ms < DeadBro.configuration.slow_query_threshold_ms
|
|
309
411
|
return false unless sql.is_a?(String)
|
|
310
412
|
|
|
@@ -94,6 +94,7 @@ module DeadBro
|
|
|
94
94
|
if defined?(DeadBro::SqlSubscriber)
|
|
95
95
|
Thread.current[:dead_bro_sql_queries]
|
|
96
96
|
Thread.current[:dead_bro_sql_queries] = nil
|
|
97
|
+
Thread.current[DeadBro::SqlSubscriber::THREAD_LOCAL_TXN_EVENTS_KEY] = nil
|
|
97
98
|
end
|
|
98
99
|
|
|
99
100
|
if defined?(DeadBro::CacheSubscriber)
|
data/lib/dead_bro/subscriber.rb
CHANGED
|
@@ -58,6 +58,7 @@ module DeadBro
|
|
|
58
58
|
|
|
59
59
|
# Stop SQL tracking and get collected queries (this was started by the request)
|
|
60
60
|
sql_queries = DeadBro::SqlSubscriber.stop_request_tracking
|
|
61
|
+
transaction_events = DeadBro::SqlSubscriber.last_transaction_events
|
|
61
62
|
|
|
62
63
|
# Stop cache, redis, and elasticsearch tracking
|
|
63
64
|
cache_events = defined?(DeadBro::CacheSubscriber) ? DeadBro::CacheSubscriber.stop_request_tracking : []
|
|
@@ -147,6 +148,42 @@ module DeadBro
|
|
|
147
148
|
})
|
|
148
149
|
end
|
|
149
150
|
|
|
151
|
+
# Everything captured during the request regardless of outcome — shared between
|
|
152
|
+
# the error and success payloads below so an errored request (e.g. a lock wait
|
|
153
|
+
# timeout raised mid-transaction) still ships the SQL/cache/memory/etc data that
|
|
154
|
+
# led up to it, not just the exception. This used to be built only for the
|
|
155
|
+
# success path, so every errored request shipped with sql_count: 0 and no trace
|
|
156
|
+
# even though the gem had already captured it.
|
|
157
|
+
detail_fields = {
|
|
158
|
+
view_runtime_ms: data[:view_runtime],
|
|
159
|
+
db_runtime_ms: data[:db_runtime],
|
|
160
|
+
memory_usage: memory_usage_mb,
|
|
161
|
+
gc_stats: gc_stats,
|
|
162
|
+
sql_count: sql_count(data),
|
|
163
|
+
sql_queries: sql_queries,
|
|
164
|
+
transaction_events: transaction_events,
|
|
165
|
+
http_outgoing: Thread.current[:dead_bro_http_events] || [],
|
|
166
|
+
cache_events: cache_events,
|
|
167
|
+
redis_events: redis_events,
|
|
168
|
+
elasticsearch_events: elasticsearch_events,
|
|
169
|
+
cache_hits: cache_hits(data),
|
|
170
|
+
cache_misses: cache_misses(data),
|
|
171
|
+
view_events: view_events,
|
|
172
|
+
view_performance: view_performance,
|
|
173
|
+
memory_events: memory_events,
|
|
174
|
+
memory_performance: memory_performance,
|
|
175
|
+
allocation_phases: allocation_phases,
|
|
176
|
+
rack_duration_ms: rack_duration_ms,
|
|
177
|
+
queue_duration_ms: Thread.current[:dead_bro_queue_duration_ms],
|
|
178
|
+
db_connection_wait_ms: db_connection_stats[:wait_ms],
|
|
179
|
+
db_connection_checkouts: db_connection_stats[:checkouts],
|
|
180
|
+
gc_pressure: gc_pressure,
|
|
181
|
+
ar_instantiation_count: ar_instantiation_count,
|
|
182
|
+
cpu_time_ms: cpu_time_ms,
|
|
183
|
+
watch_events: watch_events,
|
|
184
|
+
logs: DeadBro.logger.logs
|
|
185
|
+
}
|
|
186
|
+
|
|
150
187
|
# Report exceptions attached to this action (e.g. controller/view errors)
|
|
151
188
|
if data[:exception] || data[:exception_object]
|
|
152
189
|
begin
|
|
@@ -154,7 +191,7 @@ module DeadBro
|
|
|
154
191
|
exception_obj = data[:exception_object]
|
|
155
192
|
backtrace = Array(exception_obj&.backtrace).first(50)
|
|
156
193
|
|
|
157
|
-
error_payload =
|
|
194
|
+
error_payload = detail_fields.merge(
|
|
158
195
|
controller: data[:controller],
|
|
159
196
|
action: data[:action],
|
|
160
197
|
format: data[:format],
|
|
@@ -174,9 +211,8 @@ module DeadBro
|
|
|
174
211
|
backtrace: backtrace,
|
|
175
212
|
fingerprint: compute_error_fingerprint(exception_obj),
|
|
176
213
|
cause_chain: build_cause_chain(exception_obj),
|
|
177
|
-
error: true
|
|
178
|
-
|
|
179
|
-
}
|
|
214
|
+
error: true
|
|
215
|
+
)
|
|
180
216
|
|
|
181
217
|
event_name = (exception_class || exception_obj&.class&.name || "exception").to_s
|
|
182
218
|
client.post_metric(event_name: event_name, payload: error_payload, force: true)
|
|
@@ -186,7 +222,7 @@ module DeadBro
|
|
|
186
222
|
end
|
|
187
223
|
end
|
|
188
224
|
|
|
189
|
-
payload =
|
|
225
|
+
payload = detail_fields.merge(
|
|
190
226
|
controller: data[:controller],
|
|
191
227
|
action: data[:action],
|
|
192
228
|
format: data[:format],
|
|
@@ -195,40 +231,14 @@ module DeadBro
|
|
|
195
231
|
status: data[:status],
|
|
196
232
|
started_at: started.utc.iso8601(3),
|
|
197
233
|
duration_ms: duration_ms,
|
|
198
|
-
view_runtime_ms: data[:view_runtime],
|
|
199
|
-
db_runtime_ms: data[:db_runtime],
|
|
200
234
|
host: DeadBro.safe_hostname,
|
|
201
235
|
request_host: safe_request_host(data),
|
|
202
236
|
rails_env: DeadBro.env,
|
|
203
237
|
process_kind: DeadBro.process_kind,
|
|
204
238
|
params: safe_params(data),
|
|
205
239
|
user_agent: safe_user_agent(data),
|
|
206
|
-
user_id: extract_user_id(data)
|
|
207
|
-
|
|
208
|
-
gc_stats: gc_stats,
|
|
209
|
-
sql_count: sql_count(data),
|
|
210
|
-
sql_queries: sql_queries,
|
|
211
|
-
http_outgoing: Thread.current[:dead_bro_http_events] || [],
|
|
212
|
-
cache_events: cache_events,
|
|
213
|
-
redis_events: redis_events,
|
|
214
|
-
elasticsearch_events: elasticsearch_events,
|
|
215
|
-
cache_hits: cache_hits(data),
|
|
216
|
-
cache_misses: cache_misses(data),
|
|
217
|
-
view_events: view_events,
|
|
218
|
-
view_performance: view_performance,
|
|
219
|
-
memory_events: memory_events,
|
|
220
|
-
memory_performance: memory_performance,
|
|
221
|
-
allocation_phases: allocation_phases,
|
|
222
|
-
rack_duration_ms: rack_duration_ms,
|
|
223
|
-
queue_duration_ms: Thread.current[:dead_bro_queue_duration_ms],
|
|
224
|
-
db_connection_wait_ms: db_connection_stats[:wait_ms],
|
|
225
|
-
db_connection_checkouts: db_connection_stats[:checkouts],
|
|
226
|
-
gc_pressure: gc_pressure,
|
|
227
|
-
ar_instantiation_count: ar_instantiation_count,
|
|
228
|
-
cpu_time_ms: cpu_time_ms,
|
|
229
|
-
watch_events: watch_events,
|
|
230
|
-
logs: DeadBro.logger.logs
|
|
231
|
-
}
|
|
240
|
+
user_id: extract_user_id(data)
|
|
241
|
+
)
|
|
232
242
|
# force: true — the sampling decision (global or per-request-type) was
|
|
233
243
|
# already made above; client#post_metric must not re-roll it with the
|
|
234
244
|
# global-only rate, which would silently override a per-type sample rate.
|
|
@@ -242,6 +252,7 @@ module DeadBro
|
|
|
242
252
|
def self.drain_request_tracking
|
|
243
253
|
# wait_for_explains: false — the result is discarded, so don't block this
|
|
244
254
|
# thread waiting on pending EXPLAIN plans.
|
|
255
|
+
# stop_request_tracking also pops the transaction-events stack (see its comment).
|
|
245
256
|
DeadBro::SqlSubscriber.stop_request_tracking(wait_for_explains: false) if defined?(DeadBro::SqlSubscriber)
|
|
246
257
|
DeadBro::CacheSubscriber.stop_request_tracking if defined?(DeadBro::CacheSubscriber)
|
|
247
258
|
DeadBro::RedisSubscriber.stop_request_tracking if defined?(DeadBro::RedisSubscriber)
|
data/lib/dead_bro/version.rb
CHANGED
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: dead_bro
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.2.
|
|
4
|
+
version: 0.2.33
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Emanuel Comsa
|
|
@@ -18,6 +18,7 @@ extensions: []
|
|
|
18
18
|
extra_rdoc_files: []
|
|
19
19
|
files:
|
|
20
20
|
- CHANGELOG.md
|
|
21
|
+
- FEATURES.md
|
|
21
22
|
- README.md
|
|
22
23
|
- lib/dead_bro.rb
|
|
23
24
|
- lib/dead_bro/allocation_source_sampler.rb
|
|
@@ -80,7 +81,7 @@ required_rubygems_version: !ruby/object:Gem::Requirement
|
|
|
80
81
|
- !ruby/object:Gem::Version
|
|
81
82
|
version: '0'
|
|
82
83
|
requirements: []
|
|
83
|
-
rubygems_version: 4.0.
|
|
84
|
+
rubygems_version: 4.0.16
|
|
84
85
|
specification_version: 4
|
|
85
86
|
summary: Minimal APM for Rails apps.
|
|
86
87
|
test_files: []
|