toga-ai 1.0.650 → 1.0.651
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
|
@@ -18,7 +18,7 @@
|
|
|
18
18
|
| [Compass VIP Support Importer (worker2)](features/compass-vip-support-importer.md) | A worker2 action that ingests Compass's quarterly VIP spreadsheet and assigns each VIP user's support technician by setting `Users.c_supportedByUserId` in `Clie | worker2/Worker/Client/Compass/VipSupport.php |
|
|
19
19
|
| [Creating Worker Actions](features/creating-worker-actions.md) | How to add a new callable Worker action — a PHP class whose `public static` methods are invoked as background jobs (via webhook, cron, or `_Worker::runTask()`). | worker2/Worker/, worker2/Controller/Index.php, _underscore/Worker.php |
|
|
20
20
|
| [Cross-account AWS access for worker2 crons (_Component_Aws_Workloads)](features/cross-account-aws-access.md) | `_Component_Aws_Workloads` is the **standard, and only sanctioned, way any new worker2 cron obtains AWS access** — for any account, any region, any AWS SDK clie | worker2/Component/Aws/Workloads/Workloads.php, worker2/Config/production.ini, worker2/Worker/Infrastructure/CloudWatch.php, worker2/Worker/Monitor/Fleet.php, library/app/worker.php, library/app/cloud.php |
|
|
21
|
-
| [Elastic Beanstalk health monitor → OneUptime (ElasticBeanstalkHealth)](features/elastic-beanstalk-health-monitor.md) | `_Worker_Infrastructure_CloudWatch::ElasticBeanstalkHealth()` is a worker2 cron that reads each Elastic Beanstalk (EB) environment's
|
|
21
|
+
| [Elastic Beanstalk health monitor → OneUptime (ElasticBeanstalkHealth)](features/elastic-beanstalk-health-monitor.md) | `_Worker_Infrastructure_CloudWatch::ElasticBeanstalkHealth()` is a worker2 cron that reads each Elastic Beanstalk (EB) environment's health and reports it to a | worker2/Worker/Infrastructure/CloudWatch.php |
|
|
22
22
|
| [Elite Freshservice Sync (worker2)](features/elite-freshservice-sync.md) | `_Worker_Elite` processes Freshservice webhook events and syncs them into TOGA 2. | worker2/Worker/Elite.php, worker2/Config/dev-kmaramreddy-laptop.ini |
|
|
23
23
|
| [Error Escalation Cron (Errors::Escalate → ClickUp / email)](features/error-escalation-cron.md) | `_Worker_Infrastructure_Errors::Escalate` (renamed from `SyncWithClickup`) is the sole owner of **escalation, de-escalation, ClickUp ticketing, reminders, busin | worker2/Worker/Infrastructure/Errors.php, worker2/Worker/Notification/Email.php, worker2/Worker/Notification/EmailTemplate.php, worker2/Worker/Client/True.php, worker2/Worker/Clickup/ErrorTask.php, worker2/Worker/Clickup.php, worker2/Controller/Index.php, worker2/Config/production.ini, _underscore/Model/Core/Logs/Issue.php, dbchanges2/Core/2026-07-30a - Error escalation cron job.sql, dbchanges2/Logs/2026-08-03a - Issue clickupPriority.sql, dbchanges2/Core/2026-08-04a - Error neglect digest cron job.sql |
|
|
24
24
|
| [Error-Issue Auto-Resolution & Reopen (frequency-decay lifecycle)](features/error-issue-auto-resolution.md) | The error system could escalate and de-escalate an Issue's *urgency* but had no concept of an Issue being **resolved**. | worker2/Worker/Infrastructure/Errors.php, worker2/Worker/Clickup/ErrorTask.php, _underscore/Model/Core/Logs/Issue.php, tools/mvc/errors/get.php, tools/mvc/errors/issue/get.php, dbchanges2/Logs/2026-08-05a - Issue status baseline and auto-resolution.sql |
|
|
@@ -6,7 +6,7 @@ project: Worker
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-08-
|
|
9
|
+
updated: 2026-08-25
|
|
10
10
|
owners: [jcardinal]
|
|
11
11
|
files:
|
|
12
12
|
- worker2/Worker/Infrastructure/CloudWatch.php
|
|
@@ -19,209 +19,253 @@ related:
|
|
|
19
19
|
## Summary
|
|
20
20
|
|
|
21
21
|
`_Worker_Infrastructure_CloudWatch::ElasticBeanstalkHealth()` is a worker2 cron that reads
|
|
22
|
-
each Elastic Beanstalk (EB) environment's
|
|
23
|
-
|
|
24
|
-
|
|
22
|
+
each Elastic Beanstalk (EB) environment's health and reports it to a OneUptime **Incoming
|
|
23
|
+
Request** monitor, which pages the team. It is a **"dumb reporter, smart monitor"**
|
|
24
|
+
heartbeat (the same Pattern-B contract as
|
|
25
25
|
[OneUptime push-metric monitors](./oneuptime-worker2-monitoring.md)): the worker decides
|
|
26
26
|
pass/fail and POSTs decided **string tokens** (`alarm=HIGH/OK`, `probe=DEGRADED/OK`) that
|
|
27
27
|
OneUptime string-matches — OneUptime cannot compare numbers on a pushed body.
|
|
28
28
|
|
|
29
|
-
The load-bearing design decision this doc records is the **paging (alarm) criteria
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
29
|
+
The load-bearing design decision this doc records is the **paging (alarm) criteria**, and
|
|
30
|
+
the key **2026-08-25** change: the 5xx page is now driven by a **CloudWatch windowed
|
|
31
|
+
request-rate gate**, fully **decoupled from EB's own `HealthStatus`**. A non-5xx
|
|
32
|
+
`Degraded`/`Severe` still pages immediately off EB health; but whether a *5xx* problem pages
|
|
33
|
+
is decided by the true per-environment 5xx **rate over a trailing window** (default 5 min),
|
|
34
|
+
not by EB's ~10-second `ApplicationMetrics` snapshot.
|
|
35
|
+
|
|
36
|
+
**Why the change:** the old gate read EB `describeEnvironmentHealth`'s `ApplicationMetrics`,
|
|
37
|
+
a ~10-**second** window. At low traffic a single 500 in a near-empty bucket reads as 100%
|
|
38
|
+
5xx and false-pages. This actually fired: worker-production paged "100% requests failing"
|
|
39
|
+
while the healthd log showed **16 500s / 516 req = 3.1%** over 17 minutes — the 100% was one
|
|
40
|
+
500 alone in a 1-request 10s bucket. The windowed CloudWatch gate averages over minutes and
|
|
41
|
+
across instances, so a lone 500 can no longer dominate.
|
|
35
42
|
|
|
36
43
|
**Critical behavior:** the method is **non-fatal and never throws** — it always returns a
|
|
37
44
|
string. worker2 has **no DLQ and a 3600s SQS visibility timeout**, so any uncaught 500
|
|
38
45
|
becomes a poison-message storm; an outer `catch(\Throwable)` is the backstop and the push
|
|
39
|
-
itself runs non-fatal.
|
|
46
|
+
itself runs non-fatal. CloudWatch read failures are caught **internally** and return `null`.
|
|
40
47
|
|
|
41
48
|
## Key files / entry points
|
|
42
49
|
|
|
43
50
|
- `worker2/Worker/Infrastructure/CloudWatch.php` — `_Worker_Infrastructure_CloudWatch`,
|
|
44
51
|
action `ElasticBeanstalkHealth`, dispatched via
|
|
45
52
|
`_Worker::runTask('Infrastructure/CloudWatch/ElasticBeanstalkHealth', {awsAccountId,
|
|
46
|
-
environments, oneuptimeUrl})`.
|
|
53
|
+
environments, oneuptimeUrl, …thresholds})`.
|
|
54
|
+
- New private `fetchWindowedRequestCounts()` reads the CloudWatch metrics (below).
|
|
47
55
|
- AWS access is obtained through `_Component_Aws_Workloads::assumeCredentials()` (STS-assume
|
|
48
56
|
a read-only role per account, one assume reused across regions) — see
|
|
49
57
|
[Cross-account AWS access](./cross-account-aws-access.md), which also documents the
|
|
50
|
-
`environments` region→env-names map parameter shape.
|
|
58
|
+
`environments` region→env-names map parameter shape. The **same assumed credentials** build
|
|
59
|
+
both the `ElasticBeanstalkClient` and the new `CloudWatchClient`.
|
|
51
60
|
- `oneuptimeUrl` arrives as a **cron parameter and is a push credential** — never log it and
|
|
52
61
|
never record its value in a doc (see gotchas).
|
|
53
62
|
|
|
54
63
|
## Alarm / paging criteria (the decision logic)
|
|
55
64
|
|
|
56
|
-
|
|
57
|
-
`Severe`. `Suspended` never pages.
|
|
65
|
+
Two independent paths now feed `alarm=HIGH`:
|
|
58
66
|
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
67
|
+
1. **EB-health path (non-5xx).** EB `HealthStatus` `Degraded` or `Severe` that is **not**
|
|
68
|
+
attributable solely to 5xx (latency, instances down, a failed deploy, a mixed cause, or
|
|
69
|
+
no stated cause) **pages immediately** — unchanged, still driven off EB health.
|
|
70
|
+
`Suspended` never pages. `causesAttributeSolelyTo5xx()` is what tells a solely-5xx event
|
|
71
|
+
apart from everything else.
|
|
72
|
+
2. **CloudWatch windowed 5xx-rate path.** For a solely-5xx `Degraded`/`Severe`, the page is
|
|
73
|
+
decided **not** by EB but by the true windowed rate from CloudWatch. Given the window's
|
|
74
|
+
`requests` (ApplicationRequestsTotal) and `fivexx` (ApplicationRequests5xx), page when:
|
|
75
|
+
- `fivexx >= min5xxAbsolute` **(absolute-volume floor)** — pages regardless of share, or
|
|
76
|
+
- `requests >= minRequests` **AND** `sharePercent >= min5xxPercent` **(share gate)** — the
|
|
77
|
+
`minRequests` floor keeps a near-idle window off the share gate so a lone 500 can't page.
|
|
62
78
|
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
- the **absolute** 5xx count in the window is at/above `HTTP_5XX_ALARM_ABSOLUTE_COUNT`
|
|
67
|
-
(default 20), **or**
|
|
68
|
-
- the **5xx share** of requests is at/above `HTTP_5XX_ALARM_RATIO` (default 0.80).
|
|
69
|
-
- **Zero requests in the window → no page** (nothing observed; also avoids divide-by-zero).
|
|
79
|
+
Because the 5xx decision reads CloudWatch directly, a real 5xx flood pages **even if EB
|
|
80
|
+
hasn't escalated yet**, and a benign low-traffic 5xx blip does not page even though EB flags
|
|
81
|
+
`Degraded`.
|
|
70
82
|
|
|
71
|
-
|
|
72
|
-
when the 5xx share is high, so real 5xx floods reach `Severe` *first* and never touch the
|
|
73
|
-
gated `Degraded` branch. When the gate applied only to `Degraded`, the 80% threshold was
|
|
74
|
-
effectively dead — a real 50%-5xx `Severe` event paged despite the setting. Gating both
|
|
75
|
-
tiers makes the threshold actually govern.
|
|
83
|
+
### Cron parameters (replaces the former class constants)
|
|
76
84
|
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
ratio gate only applies when **every** stated cause references 5xx (and there is at least
|
|
80
|
-
one); anything else pages.
|
|
85
|
+
The four thresholds are **optional cron parameters** spread as **named arguments** from
|
|
86
|
+
`CronJobs.parameters`, with method-signature defaults:
|
|
81
87
|
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
`
|
|
85
|
-
|
|
88
|
+
| Parameter | Default | Meaning |
|
|
89
|
+
|---|---|---|
|
|
90
|
+
| `windowMinutes` | `5` | Trailing window (minutes) over which the CloudWatch rate is summed |
|
|
91
|
+
| `minRequests` | `100` | Minimum windowed request count before the **share** gate applies |
|
|
92
|
+
| `min5xxPercent` | `20.0` | 5xx share (percent) at/above which a solely-5xx env pages, once `minRequests` is met |
|
|
93
|
+
| `min5xxAbsolute` | `20` | Absolute windowed 5xx count at/above which a solely-5xx env pages regardless of share |
|
|
86
94
|
|
|
87
|
-
|
|
95
|
+
Removed this iteration: `shouldPageForHealth()` and the old
|
|
96
|
+
`HTTP_5XX_ALARM_RATIO` / `HTTP_5XX_ALARM_ABSOLUTE_COUNT` constants (superseded by the
|
|
97
|
+
windowed gate and the four parameters).
|
|
88
98
|
|
|
89
|
-
|
|
90
|
-
the monitor sets **`alarm=HIGH`** for that environment (a blind read must never masquerade
|
|
91
|
-
as healthy) **and** sets **`probe=DEGRADED`** (coverage gap, distinct from a measured-bad
|
|
92
|
-
signal — same alarm-vs-probe split as the multi-client monitors in
|
|
93
|
-
[OneUptime push-metric monitors](./oneuptime-worker2-monitoring.md)).
|
|
99
|
+
### Per-monitor threshold tuning (one cron watches many environments)
|
|
94
100
|
|
|
95
|
-
|
|
101
|
+
A single `ElasticBeanstalkHealth` cron watches environments of **very different scale**, so
|
|
102
|
+
thresholds are tuned per monitor. Measured 3h baselines: api-production-1 ~922–1737 req/5min,
|
|
103
|
+
apiproxy ~400–1650, but eu-west-1 envs only ~40–90 req/5min; 5xx was 0 across all.
|
|
96
104
|
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
105
|
+
- **Worker + worker 1.0 monitors** keep the **method defaults** (`100 / 20% / 20`) — the low
|
|
106
|
+
absolute floor suits ~150 req/5min traffic.
|
|
107
|
+
- **API and API-Proxy monitors** (which each also include the low-traffic eu-west-1 env) are
|
|
108
|
+
overridden to **`windowMinutes=5, minRequests=30, min5xxPercent=20, min5xxAbsolute=100`** —
|
|
109
|
+
the lower request floor keeps the small env on the share gate, while the higher absolute
|
|
110
|
+
floor stops the busy envs paging on minor blips (20 scattered 500s on ~1300 req = <2%). The
|
|
111
|
+
override is applied by the dbchanges2 migration below.
|
|
112
|
+
|
|
113
|
+
### Fail-safe: an unmeasurable solely-5xx event still pages
|
|
114
|
+
|
|
115
|
+
The windowed metrics are **per-instance only** (see below), and some envs never publish them
|
|
116
|
+
(the 1.0 worker env `agilant-worker` publishes only `EnvironmentHealth`). For such an env the
|
|
117
|
+
gate has **no data** and returns **zeros** (not `null`). To avoid silently missing a real
|
|
118
|
+
flood, when EB reports a **solely-5xx `Degraded`/`Severe`** but the window is **empty
|
|
119
|
+
(unmeasurable)**, the monitor **pages** with gate token **`health-5xx-unmeasured`** rather
|
|
120
|
+
than trusting the empty read.
|
|
121
|
+
|
|
122
|
+
A CloudWatch read *failure* (as opposed to empty data) sets **`probe=DEGRADED`** (coverage
|
|
123
|
+
gap), distinct from a measured-bad `alarm=HIGH` — same alarm-vs-probe split as the
|
|
124
|
+
multi-client monitors in [OneUptime push-metric monitors](./oneuptime-worker2-monitoring.md).
|
|
125
|
+
|
|
126
|
+
## Reading the windowed 5xx rate from CloudWatch (durable AWS reference)
|
|
127
|
+
|
|
128
|
+
- **EB application-request metrics are PER-INSTANCE ONLY.** `ApplicationRequestsTotal`,
|
|
129
|
+
`ApplicationRequests5xx` (and `…2xx/3xx/4xx`) are published under dimensions
|
|
130
|
+
`{EnvironmentName, InstanceId}` in namespace `AWS/ElasticBeanstalk`. **There is no
|
|
131
|
+
environment-only rollup.** You must sum across instances yourself with metric-math over a
|
|
132
|
+
search expression:
|
|
133
|
+
|
|
134
|
+
```
|
|
135
|
+
SUM(SEARCH('{AWS/ElasticBeanstalk,EnvironmentName,InstanceId} MetricName="ApplicationRequestsTotal" EnvironmentName="<env>"','Sum',<periodSeconds>))
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
The `SEARCH` auto-discovers every instance's series and `SUM` folds them into one, so the
|
|
139
|
+
window total is correct as instances come and go. Query total and 5xx as two expressions
|
|
140
|
+
over `windowMinutes`.
|
|
141
|
+
- **Not every environment publishes these metrics.** worker-production, api-production-*, and
|
|
142
|
+
apiproxy-production-* all publish them; the **1.0 worker env `agilant-worker` publishes
|
|
143
|
+
ONLY `EnvironmentHealth`** (no app-request metrics) → the gate reads zeros there → the
|
|
144
|
+
`health-5xx-unmeasured` fail-safe above covers it.
|
|
145
|
+
- **IAM already sufficient.** The `WorkloadsRuntime` role grants `cloudwatch:*` in **both**
|
|
146
|
+
account `654654170868` (worker/api) and `502614707982` (legacy worker), so no IAM change
|
|
147
|
+
was needed for the CloudWatch reads.
|
|
148
|
+
|
|
149
|
+
### EB enhanced-health facts (still relevant — the EB-health path)
|
|
125
150
|
|
|
126
151
|
- **HealthStatus severity ladder:** `Ok → Info → Warning → Degraded → Severe` (plus the grey
|
|
127
|
-
states `Pending`/`Unknown`/`Suspended`/`NoData`). `Degraded` is
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
environment effectively not serving."
|
|
131
|
-
- **`ApplicationMetrics` returns RAW COUNTS, not percentages** — despite the API-reference
|
|
132
|
-
prose saying "percentage"/"per second". Fields: `Duration` (Integer seconds, usually 10),
|
|
133
|
-
`RequestCount` (Integer, **total** requests over the window),
|
|
134
|
-
`StatusCodes.{Status2xx,Status3xx,Status4xx,Status5xx}` (Integer counts). AWS's own
|
|
135
|
-
example: 2xx 3391 + 5xx 843 = RequestCount 4234. With no traffic, `RequestCount=0` and
|
|
136
|
-
`StatusCodes` may be **absent** → treat as "no data," not "zero failures." The 5xx-ratio
|
|
137
|
-
denominator is `RequestCount` (the authoritative total).
|
|
152
|
+
states `Pending`/`Unknown`/`Suspended`/`NoData`). `Degraded` is routinely tripped by benign
|
|
153
|
+
transients (Auto Scaling scale-up, a mid-deploy dip), which is why a solely-5xx `Degraded`
|
|
154
|
+
is now gated on the CloudWatch rate rather than paged outright.
|
|
138
155
|
- **Why an env is Degraded comes from `Causes`:** EB `Causes` strings literally contain e.g.
|
|
139
156
|
`"19.9 % of the requests are failing with HTTP 5xx."` A substring check reliably
|
|
140
|
-
**classifies** 5xx
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
- **A failed application DEPLOYMENT is NOT reliably reflected in HealthStatus** —
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
157
|
+
**classifies** solely-5xx vs. not; a wording change fails safe (treated as non-5xx →
|
|
158
|
+
pages). `describeEnvironmentHealth` is called with `AttributeNames` including
|
|
159
|
+
`HealthStatus`, `Status`, `Color`, `Causes`.
|
|
160
|
+
- **A failed application DEPLOYMENT is NOT reliably reflected in HealthStatus** — the
|
|
161
|
+
authoritative structured deploy signal is per-instance `describeInstancesHealth` →
|
|
162
|
+
`Deployment.Status`. (A deploy that fails without ever degrading health remains an accepted
|
|
163
|
+
residual gap.)
|
|
164
|
+
- **IAM:** `describeEnvironmentHealth` is covered by the managed `AWSElasticBeanstalkReadOnly`
|
|
165
|
+
policy and requires enhanced health enabled.
|
|
166
|
+
|
|
167
|
+
## Deploy order (mandatory) — code BEFORE the migration
|
|
168
|
+
|
|
169
|
+
worker2 code **must** deploy **before** the dbchanges2 threshold migration runs.
|
|
170
|
+
|
|
171
|
+
- new-code + old-params → **fine** (the params are optional with defaults).
|
|
172
|
+
- **old-code + new-params → FATAL.** The dispatcher spreads `CronJobs.parameters` as **named
|
|
173
|
+
arguments**; an old method signature that lacks `windowMinutes`/`minRequests`/… throws
|
|
174
|
+
**"Unknown named parameter"** and **every run fails** → the monitor goes dark → OneUptime
|
|
175
|
+
raises an offline incident.
|
|
176
|
+
|
|
177
|
+
Correct order: **1) deploy worker2 → 2) apply the SQL → 3) add the OneUptime criterion.**
|
|
178
|
+
|
|
179
|
+
## dbchanges2 migration (the API/API-Proxy override)
|
|
180
|
+
|
|
181
|
+
`dbchanges2/Core/2026-08-25a - ApiApiproxy Beanstalk Health 5xx Thresholds.sql` applies the
|
|
182
|
+
API/API-Proxy override with a JSON_SET that **adds only the four keys**, preserving
|
|
183
|
+
`awsAccountId`/`environments`/`oneuptimeUrl`:
|
|
184
|
+
|
|
185
|
+
```sql
|
|
186
|
+
UPDATE CronJobs
|
|
187
|
+
SET parameters = JSON_SET(parameters,
|
|
188
|
+
'$.windowMinutes', 5, '$.minRequests', 30,
|
|
189
|
+
'$.min5xxPercent', 20, '$.min5xxAbsolute', 100)
|
|
190
|
+
WHERE action = 'Infrastructure/CloudWatch/ElasticBeanstalkHealth'
|
|
191
|
+
AND name IN (<the two API monitor names>);
|
|
192
|
+
```
|
|
193
|
+
|
|
194
|
+
`Core.CronJobs.parameters` is confirmed a native `JSON` column, so `JSON_SET` merges rather
|
|
195
|
+
than overwrites.
|
|
151
196
|
|
|
152
197
|
## OneUptime monitor config (external, not code)
|
|
153
198
|
|
|
154
199
|
Paging depends on the OneUptime Incoming-Request monitor's **string-match** criteria on the
|
|
155
|
-
POSTed body — `Contains "alarm":"HIGH"`
|
|
156
|
-
This is external OneUptime configuration, not worker2 code.
|
|
157
|
-
|
|
158
|
-
## Design history — rejected direction (this session)
|
|
200
|
+
POSTed body — `Contains "alarm":"HIGH"` pages.
|
|
159
201
|
|
|
160
|
-
|
|
161
|
-
and
|
|
162
|
-
|
|
163
|
-
|
|
164
|
-
|
|
165
|
-
|
|
166
|
-
`
|
|
167
|
-
iteration:** a deploy that fails fast *without ever degrading health* is not caught.
|
|
202
|
+
**Known alerting gap (recommended fix):** the current monitors match `"status":"error"`,
|
|
203
|
+
`"alarm":"HIGH"`, and heartbeat-gap rules, but **nothing matches `"probe":"DEGRADED"`**. So a
|
|
204
|
+
CloudWatch-metrics **read failure while EB reads Ok** (body = `alarm:OK, probe:DEGRADED,
|
|
205
|
+
status:reporting`) is **invisible to alerting**. Add a criterion matching
|
|
206
|
+
`Contains "probe":"DEGRADED"` → a **Degraded-severity, auto-resolving** incident, placed
|
|
207
|
+
**before** the "online" recovery rule (which also matches `alarm:OK`). The code already emits
|
|
208
|
+
the `probe` token; this is a OneUptime config change the developer applies.
|
|
168
209
|
|
|
169
210
|
## Known gaps / follow-ups (not done)
|
|
170
211
|
|
|
171
|
-
- **
|
|
172
|
-
`
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
- **
|
|
176
|
-
`shouldPageForHealth()` and `causesAttributeSolelyTo5xx()` are pure and ideal to unit-test
|
|
177
|
-
(ratio at exactly 0.80; absolute count 19 vs 20; `Severe` over a solely-5xx cause; a
|
|
178
|
-
mixed 5xx + non-5xx cause; empty causes; null metrics; zero requests) — deferred pending a
|
|
179
|
-
harness.
|
|
212
|
+
- **No first-party test harness in worker2** (all tests are vendor/). The gate decision and
|
|
213
|
+
`causesAttributeSolelyTo5xx()` are pure and ideal to unit-test (share at exactly
|
|
214
|
+
`min5xxPercent`; absolute at `min5xxAbsolute-1` vs `min5xxAbsolute`; `requests` just below
|
|
215
|
+
`minRequests`; empty/zero window; mixed 5xx + non-5xx cause) — deferred pending a harness.
|
|
216
|
+
- **Deploy fails without degrading health** — still not caught (accepted residual).
|
|
180
217
|
|
|
181
218
|
## Gotchas / known issues
|
|
182
219
|
|
|
220
|
+
- **Deploy worker2 code BEFORE the threshold migration** — old-code + new-params is fatal
|
|
221
|
+
("Unknown named parameter" on every run → monitor dark). See Deploy order.
|
|
183
222
|
- **Keep the method non-fatal — never let it throw.** worker2 has no DLQ and a 3600s SQS
|
|
184
|
-
visibility timeout
|
|
185
|
-
|
|
186
|
-
- **
|
|
187
|
-
`
|
|
188
|
-
-
|
|
189
|
-
`
|
|
190
|
-
-
|
|
191
|
-
|
|
223
|
+
visibility timeout; an outer `catch(\Throwable)` is the backstop and CloudWatch failures are
|
|
224
|
+
caught internally (return `null`).
|
|
225
|
+
- **EB app-request metrics are per-instance only** — there is no env rollup; you must
|
|
226
|
+
`SUM(SEARCH(...))` across instances or you undercount.
|
|
227
|
+
- **An empty window ≠ zero failures.** An env with no published app-request metrics (e.g. 1.0
|
|
228
|
+
`agilant-worker`) reads zeros; a solely-5xx EB event over an empty window pages via the
|
|
229
|
+
`health-5xx-unmeasured` fail-safe rather than reading healthy.
|
|
230
|
+
- **`minRequests` floor guards the near-idle edge** the old code couldn't — the share gate is
|
|
231
|
+
suppressed until the window has real volume, so a lone 500 in a quiet window no longer pages.
|
|
232
|
+
- **OneUptime has no `probe:DEGRADED` criterion** — a metrics-read failure while EB is Ok is
|
|
233
|
+
currently unalerted until that criterion is added.
|
|
192
234
|
- **`oneuptimeUrl` is a push credential** — it arrives as a cron parameter; never log it or
|
|
193
235
|
record its value in a doc.
|
|
194
|
-
- **Near-idle ratio edge (residual, accepted).** With the tiny-sample guard removed, a
|
|
195
|
-
near-idle env whose window holds a tiny all-error sample (e.g. its only request being a
|
|
196
|
-
500 = 100%) still trips the ratio and pages. Optional future guard: require a minimum 5xx
|
|
197
|
-
count before the ratio applies (the absolute floor gates high volume, not this low-volume
|
|
198
|
-
edge).
|
|
199
236
|
|
|
200
237
|
## Change history
|
|
238
|
+
- 2026-08-25 — Replaced the EB `ApplicationMetrics` 5xx gate (a ~10s window that false-paged
|
|
239
|
+
"100% failing" on a lone 500 in a near-empty bucket — real worker-production incident: 16
|
|
240
|
+
500s / 516 req = 3.1% over 17 min) with a **CloudWatch windowed request-rate gate**. Reads
|
|
241
|
+
true per-env `ApplicationRequestsTotal`/`ApplicationRequests5xx` over a trailing window
|
|
242
|
+
(default 5 min) via `SUM(SEARCH('{AWS/ElasticBeanstalk,EnvironmentName,InstanceId}…','Sum',
|
|
243
|
+
period))` (these metrics are per-instance only — no env rollup). The 5xx page is now
|
|
244
|
+
**decoupled from EB `HealthStatus`**: pages when `5xx >= min5xxAbsolute` OR `requests >=
|
|
245
|
+
minRequests AND share >= min5xxPercent`; a non-5xx `Degraded`/`Severe` still pages
|
|
246
|
+
immediately. Added four optional cron params (`windowMinutes=5, minRequests=100,
|
|
247
|
+
min5xxPercent=20.0, min5xxAbsolute=20`) spread as named arguments; removed
|
|
248
|
+
`shouldPageForHealth()` and the `HTTP_5XX_ALARM_RATIO`/`HTTP_5XX_ALARM_ABSOLUTE_COUNT`
|
|
249
|
+
constants. Added a `health-5xx-unmeasured` fail-safe (solely-5xx over an empty/unpublished
|
|
250
|
+
window pages rather than reading healthy). Per-monitor tuning: worker + worker-1.0 keep
|
|
251
|
+
defaults; API/API-Proxy overridden to `30 / 20% / 100` (dbchanges2
|
|
252
|
+
`2026-08-25a - ApiApiproxy Beanstalk Health 5xx Thresholds.sql`, JSON_SET on the native
|
|
253
|
+
`CronJobs.parameters` JSON column). Recorded the mandatory deploy order (code before the
|
|
254
|
+
migration — old-code+new-params throws "Unknown named parameter") and the OneUptime gap
|
|
255
|
+
(no `probe:DEGRADED` criterion, so a metrics-read failure while EB reads Ok is unalerted).
|
|
256
|
+
WorkloadsRuntime already grants `cloudwatch:*` in accounts 654654170868 and 502614707982 —
|
|
257
|
+
no IAM change. (jcardinal)
|
|
201
258
|
- 2026-08-19 — Overhauled the 5xx alarm gating in `shouldPageForHealth()`. The 80% share
|
|
202
|
-
gate now
|
|
203
|
-
threshold dead because EB escalates real 5xx floods straight to `Severe`
|
|
204
|
-
`
|
|
205
|
-
|
|
206
|
-
|
|
207
|
-
|
|
208
|
-
cause references 5xx, ≥1) + a `causeIsFivexx()` helper, so a mixed cause (5xx trickle beside
|
|
209
|
-
a real non-5xx failure) always pages — a security-review gap the Severe change would have
|
|
210
|
-
widened. Added `HTTP_5XX_ALARM_ABSOLUTE_COUNT` (20): a solely-5xx env pages regardless of
|
|
211
|
-
share once the absolute 5xx count reaches the floor (closes the ratio's magnitude-blindness,
|
|
212
|
-
e.g. 60k of 100k). Removed `HTTP_5XX_MIN_REQUEST_COUNT` (was 20) and its tiny-sample
|
|
213
|
-
page-anyway guard — a small number of 5xx is now judged purely on the ratio; below the
|
|
214
|
-
absolute floor, page only if share ≥ 0.80, and zero requests → no page. Accepted residual:
|
|
215
|
-
a near-idle env with a tiny all-error sample still trips the ratio. (jcardinal)
|
|
259
|
+
gate now applied to **both** `Degraded` and `Severe` (was `Degraded`-only, which left the
|
|
260
|
+
threshold dead because EB escalates real 5xx floods straight to `Severe`). Narrowed 5xx
|
|
261
|
+
classification: replaced `causesAttributeTo5xx()` (any cause mentions 5xx) with
|
|
262
|
+
`causesAttributeSolelyTo5xx()` (every cause references 5xx, ≥1). Added
|
|
263
|
+
`HTTP_5XX_ALARM_ABSOLUTE_COUNT` (20). Removed `HTTP_5XX_MIN_REQUEST_COUNT` tiny-sample
|
|
264
|
+
guard. (jcardinal)
|
|
216
265
|
- 2026-08-17 — Created. Documented the `ElasticBeanstalkHealth` alarm/paging criteria (page
|
|
217
|
-
on `Degraded`/`Severe`; suppress a 5xx-driven `Degraded` unless
|
|
218
|
-
`
|
|
219
|
-
|
|
220
|
-
|
|
221
|
-
|
|
222
|
-
|
|
223
|
-
`ApplicationMetrics` sets the ratio; failed deploys aren't reliably in HealthStatus —
|
|
224
|
-
`describeInstancesHealth.Deployment.Status` is authoritative; `AWSElasticBeanstalkReadOnly`
|
|
225
|
-
covers the reads). Recorded the reverted `['Severe']`-only + deployment-failure-detector
|
|
226
|
-
direction and the accepted residual gap (a deploy that fails without degrading health).
|
|
227
|
-
(jcardinal)
|
|
266
|
+
on `Degraded`/`Severe`; suppress a 5xx-driven `Degraded` unless share ≥ ratio), the
|
|
267
|
+
blind-read fail-safe (`alarm=HIGH` + `probe=DEGRADED`), and durable EB enhanced-health
|
|
268
|
+
facts (severity ladder; `ApplicationMetrics` raw counts; `Causes` classifies; failed
|
|
269
|
+
deploys aren't reliably in HealthStatus). (jcardinal)
|
|
270
|
+
</content>
|
|
271
|
+
</invoke>
|
package/package.json
CHANGED