toga-ai 1.0.650 → 1.0.651

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -18,7 +18,7 @@
18
18
  | [Compass VIP Support Importer (worker2)](features/compass-vip-support-importer.md) | A worker2 action that ingests Compass's quarterly VIP spreadsheet and assigns each VIP user's support technician by setting `Users.c_supportedByUserId` in `Clie | worker2/Worker/Client/Compass/VipSupport.php |
19
19
  | [Creating Worker Actions](features/creating-worker-actions.md) | How to add a new callable Worker action — a PHP class whose `public static` methods are invoked as background jobs (via webhook, cron, or `_Worker::runTask()`). | worker2/Worker/, worker2/Controller/Index.php, _underscore/Worker.php |
20
20
  | [Cross-account AWS access for worker2 crons (_Component_Aws_Workloads)](features/cross-account-aws-access.md) | `_Component_Aws_Workloads` is the **standard, and only sanctioned, way any new worker2 cron obtains AWS access** — for any account, any region, any AWS SDK clie | worker2/Component/Aws/Workloads/Workloads.php, worker2/Config/production.ini, worker2/Worker/Infrastructure/CloudWatch.php, worker2/Worker/Monitor/Fleet.php, library/app/worker.php, library/app/cloud.php |
21
- | [Elastic Beanstalk health monitor → OneUptime (ElasticBeanstalkHealth)](features/elastic-beanstalk-health-monitor.md) | `_Worker_Infrastructure_CloudWatch::ElasticBeanstalkHealth()` is a worker2 cron that reads each Elastic Beanstalk (EB) environment's **enhanced-health** status | worker2/Worker/Infrastructure/CloudWatch.php |
21
+ | [Elastic Beanstalk health monitor → OneUptime (ElasticBeanstalkHealth)](features/elastic-beanstalk-health-monitor.md) | `_Worker_Infrastructure_CloudWatch::ElasticBeanstalkHealth()` is a worker2 cron that reads each Elastic Beanstalk (EB) environment's health and reports it to a | worker2/Worker/Infrastructure/CloudWatch.php |
22
22
  | [Elite Freshservice Sync (worker2)](features/elite-freshservice-sync.md) | `_Worker_Elite` processes Freshservice webhook events and syncs them into TOGA 2. | worker2/Worker/Elite.php, worker2/Config/dev-kmaramreddy-laptop.ini |
23
23
  | [Error Escalation Cron (Errors::Escalate → ClickUp / email)](features/error-escalation-cron.md) | `_Worker_Infrastructure_Errors::Escalate` (renamed from `SyncWithClickup`) is the sole owner of **escalation, de-escalation, ClickUp ticketing, reminders, busin | worker2/Worker/Infrastructure/Errors.php, worker2/Worker/Notification/Email.php, worker2/Worker/Notification/EmailTemplate.php, worker2/Worker/Client/True.php, worker2/Worker/Clickup/ErrorTask.php, worker2/Worker/Clickup.php, worker2/Controller/Index.php, worker2/Config/production.ini, _underscore/Model/Core/Logs/Issue.php, dbchanges2/Core/2026-07-30a - Error escalation cron job.sql, dbchanges2/Logs/2026-08-03a - Issue clickupPriority.sql, dbchanges2/Core/2026-08-04a - Error neglect digest cron job.sql |
24
24
  | [Error-Issue Auto-Resolution & Reopen (frequency-decay lifecycle)](features/error-issue-auto-resolution.md) | The error system could escalate and de-escalate an Issue's *urgency* but had no concept of an Issue being **resolved**. | worker2/Worker/Infrastructure/Errors.php, worker2/Worker/Clickup/ErrorTask.php, _underscore/Model/Core/Logs/Issue.php, tools/mvc/errors/get.php, tools/mvc/errors/issue/get.php, dbchanges2/Logs/2026-08-05a - Issue status baseline and auto-resolution.sql |
@@ -6,7 +6,7 @@ project: Worker
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-08-19
9
+ updated: 2026-08-25
10
10
  owners: [jcardinal]
11
11
  files:
12
12
  - worker2/Worker/Infrastructure/CloudWatch.php
@@ -19,209 +19,253 @@ related:
19
19
  ## Summary
20
20
 
21
21
  `_Worker_Infrastructure_CloudWatch::ElasticBeanstalkHealth()` is a worker2 cron that reads
22
- each Elastic Beanstalk (EB) environment's **enhanced-health** status and reports it to a
23
- OneUptime **Incoming Request** monitor, which pages the team. It is a **"dumb reporter,
24
- smart monitor"** heartbeat (the same Pattern-B contract as
22
+ each Elastic Beanstalk (EB) environment's health and reports it to a OneUptime **Incoming
23
+ Request** monitor, which pages the team. It is a **"dumb reporter, smart monitor"**
24
+ heartbeat (the same Pattern-B contract as
25
25
  [OneUptime push-metric monitors](./oneuptime-worker2-monitoring.md)): the worker decides
26
26
  pass/fail and POSTs decided **string tokens** (`alarm=HIGH/OK`, `probe=DEGRADED/OK`) that
27
27
  OneUptime string-matches — OneUptime cannot compare numbers on a pushed body.
28
28
 
29
- The load-bearing design decision this doc records is the **paging (alarm) criteria**: page
30
- on EB `Degraded`/`Severe`, but **suppress the common "an occasional HTTP 500 tripped
31
- Degraded/Severe" false alarm** unless 5xx errors actually dominate the request mix (by
32
- share) or reach a real absolute volume. The 5xx gate applies to **both** `Degraded` and
33
- `Severe`, and only when EB attributes the problem **solely** to 5xx — a mixed cause always
34
- pages.
29
+ The load-bearing design decision this doc records is the **paging (alarm) criteria**, and
30
+ the key **2026-08-25** change: the 5xx page is now driven by a **CloudWatch windowed
31
+ request-rate gate**, fully **decoupled from EB's own `HealthStatus`**. A non-5xx
32
+ `Degraded`/`Severe` still pages immediately off EB health; but whether a *5xx* problem pages
33
+ is decided by the true per-environment 5xx **rate over a trailing window** (default 5 min),
34
+ not by EB's ~10-second `ApplicationMetrics` snapshot.
35
+
36
+ **Why the change:** the old gate read EB `describeEnvironmentHealth`'s `ApplicationMetrics`,
37
+ a ~10-**second** window. At low traffic a single 500 in a near-empty bucket reads as 100%
38
+ 5xx and false-pages. This actually fired: worker-production paged "100% requests failing"
39
+ while the healthd log showed **16 500s / 516 req = 3.1%** over 17 minutes — the 100% was one
40
+ 500 alone in a 1-request 10s bucket. The windowed CloudWatch gate averages over minutes and
41
+ across instances, so a lone 500 can no longer dominate.
35
42
 
36
43
  **Critical behavior:** the method is **non-fatal and never throws** — it always returns a
37
44
  string. worker2 has **no DLQ and a 3600s SQS visibility timeout**, so any uncaught 500
38
45
  becomes a poison-message storm; an outer `catch(\Throwable)` is the backstop and the push
39
- itself runs non-fatal.
46
+ itself runs non-fatal. CloudWatch read failures are caught **internally** and return `null`.
40
47
 
41
48
  ## Key files / entry points
42
49
 
43
50
  - `worker2/Worker/Infrastructure/CloudWatch.php` — `_Worker_Infrastructure_CloudWatch`,
44
51
  action `ElasticBeanstalkHealth`, dispatched via
45
52
  `_Worker::runTask('Infrastructure/CloudWatch/ElasticBeanstalkHealth', {awsAccountId,
46
- environments, oneuptimeUrl})`.
53
+ environments, oneuptimeUrl, …thresholds})`.
54
+ - New private `fetchWindowedRequestCounts()` reads the CloudWatch metrics (below).
47
55
  - AWS access is obtained through `_Component_Aws_Workloads::assumeCredentials()` (STS-assume
48
56
  a read-only role per account, one assume reused across regions) — see
49
57
  [Cross-account AWS access](./cross-account-aws-access.md), which also documents the
50
- `environments` region→env-names map parameter shape.
58
+ `environments` region→env-names map parameter shape. The **same assumed credentials** build
59
+ both the `ElasticBeanstalkClient` and the new `CloudWatchClient`.
51
60
  - `oneuptimeUrl` arrives as a **cron parameter and is a push credential** — never log it and
52
61
  never record its value in a doc (see gotchas).
53
62
 
54
63
  ## Alarm / paging criteria (the decision logic)
55
64
 
56
- Per environment, `alarm=HIGH` (page) when the EB `HealthStatus` is `Degraded` **or**
57
- `Severe`. `Suspended` never pages.
65
+ Two independent paths now feed `alarm=HIGH`:
58
66
 
59
- The **5xx gate applies to both `Degraded` and `Severe`.** A non-5xx or mixed-cause
60
- `Degraded`/`Severe` — and any env with no stated cause — **always pages**; the gate only
61
- ever suppresses an env EB attributes **solely** to 5xx.
67
+ 1. **EB-health path (non-5xx).** EB `HealthStatus` `Degraded` or `Severe` that is **not**
68
+ attributable solely to 5xx (latency, instances down, a failed deploy, a mixed cause, or
69
+ no stated cause) **pages immediately** — unchanged, still driven off EB health.
70
+ `Suspended` never pages. `causesAttributeSolelyTo5xx()` is what tells a solely-5xx event
71
+ apart from everything else.
72
+ 2. **CloudWatch windowed 5xx-rate path.** For a solely-5xx `Degraded`/`Severe`, the page is
73
+ decided **not** by EB but by the true windowed rate from CloudWatch. Given the window's
74
+ `requests` (ApplicationRequestsTotal) and `fivexx` (ApplicationRequests5xx), page when:
75
+ - `fivexx >= min5xxAbsolute` **(absolute-volume floor)** — pages regardless of share, or
76
+ - `requests >= minRequests` **AND** `sharePercent >= min5xxPercent` **(share gate)** — the
77
+ `minRequests` floor keeps a near-idle window off the share gate so a lone 500 can't page.
62
78
 
63
- - **Non-5xx or mixed-cause `Degraded`/`Severe`** (latency, instances down, a failed deploy,
64
- or a 5xx trickle *alongside* a real non-5xx failure) **always pages.**
65
- - **Solely-5xx `Degraded`/`Severe` pages only on a genuine flood** — either:
66
- - the **absolute** 5xx count in the window is at/above `HTTP_5XX_ALARM_ABSOLUTE_COUNT`
67
- (default 20), **or**
68
- - the **5xx share** of requests is at/above `HTTP_5XX_ALARM_RATIO` (default 0.80).
69
- - **Zero requests in the window → no page** (nothing observed; also avoids divide-by-zero).
79
+ Because the 5xx decision reads CloudWatch directly, a real 5xx flood pages **even if EB
80
+ hasn't escalated yet**, and a benign low-traffic 5xx blip does not page even though EB flags
81
+ `Degraded`.
70
82
 
71
- **Why the gate now covers `Severe` too:** EB escalates an environment to `Severe` precisely
72
- when the 5xx share is high, so real 5xx floods reach `Severe` *first* and never touch the
73
- gated `Degraded` branch. When the gate applied only to `Degraded`, the 80% threshold was
74
- effectively dead — a real 50%-5xx `Severe` event paged despite the setting. Gating both
75
- tiers makes the threshold actually govern.
83
+ ### Cron parameters (replaces the former class constants)
76
84
 
77
- **Why "solely" 5xx, not "any" 5xx cause:** a mixed-cause env (a 5xx trickle plus a real
78
- non-5xx failure such as a failed deploy) must not be silenced by a sub-threshold ratio. The
79
- ratio gate only applies when **every** stated cause references 5xx (and there is at least
80
- one); anything else pages.
85
+ The four thresholds are **optional cron parameters** spread as **named arguments** from
86
+ `CronJobs.parameters`, with method-signature defaults:
81
87
 
82
- **Why an absolute floor in addition to the ratio:** the ratio alone is magnitude-blind —
83
- 60k failing out of 100k reads as 60% and would be suppressed, though it is plainly a flood.
84
- `HTTP_5XX_ALARM_ABSOLUTE_COUNT` pages such an env regardless of share; tune it to the busiest
85
- monitored environment's request volume.
88
+ | Parameter | Default | Meaning |
89
+ |---|---|---|
90
+ | `windowMinutes` | `5` | Trailing window (minutes) over which the CloudWatch rate is summed |
91
+ | `minRequests` | `100` | Minimum windowed request count before the **share** gate applies |
92
+ | `min5xxPercent` | `20.0` | 5xx share (percent) at/above which a solely-5xx env pages, once `minRequests` is met |
93
+ | `min5xxAbsolute` | `20` | Absolute windowed 5xx count at/above which a solely-5xx env pages regardless of share |
86
94
 
87
- ### Fail-safe: a blind read pages, never reads healthy
95
+ Removed this iteration: `shouldPageForHealth()` and the old
96
+ `HTTP_5XX_ALARM_RATIO` / `HTTP_5XX_ALARM_ABSOLUTE_COUNT` constants (superseded by the
97
+ windowed gate and the four parameters).
88
98
 
89
- If the health read (`describeEnvironmentHealth`) throws `AwsException` for an environment,
90
- the monitor sets **`alarm=HIGH`** for that environment (a blind read must never masquerade
91
- as healthy) **and** sets **`probe=DEGRADED`** (coverage gap, distinct from a measured-bad
92
- signal — same alarm-vs-probe split as the multi-client monitors in
93
- [OneUptime push-metric monitors](./oneuptime-worker2-monitoring.md)).
99
+ ### Per-monitor threshold tuning (one cron watches many environments)
94
100
 
95
- ### Configurable constants (top of the class)
101
+ A single `ElasticBeanstalkHealth` cron watches environments of **very different scale**, so
102
+ thresholds are tuned per monitor. Measured 3h baselines: api-production-1 ~922–1737 req/5min,
103
+ apiproxy ~400–1650, but eu-west-1 envs only ~40–90 req/5min; 5xx was 0 across all.
96
104
 
97
- | Constant | Default | Meaning |
98
- |---|---|---|
99
- | `HEALTH_ALARM_STATUSES` | `['Degraded','Severe']` | HealthStatus values that page |
100
- | `HTTP_5XX_ALARM_RATIO` | `0.80` | 5xx share (0.0–1.0) at/above which a solely-5xx `Degraded`/`Severe` pages |
101
- | `HTTP_5XX_ALARM_ABSOLUTE_COUNT` | `20` | Absolute 5xx count in the window at/above which a solely-5xx `Degraded`/`Severe` pages regardless of share |
102
- | `HTTP_5XX_CAUSE_MARKERS` | `['5xx','http 5']` | Case-insensitive substrings identifying a 5xx-attributed EB `Cause` |
103
-
104
- ### Implementation shape
105
-
106
- - `shouldPageForHealth(healthStatus, causes, applicationMetrics): [bool, ?float]` — returns
107
- the page decision and the observed 5xx ratio.
108
- - `causesAttributeSolelyTo5xx(causes): bool` — true only when **every** EB `Cause` references
109
- 5xx **and there is at least one** cause. Replaces the former `causesAttributeTo5xx()`, which
110
- returned true if *any* cause mentioned 5xx and so let a mixed-cause env be gated.
111
- - `causeIsFivexx(cause): bool` — helper that substring-matches `HTTP_5XX_CAUSE_MARKERS`
112
- against a single EB `Cause` string (case-insensitive); `causesAttributeSolelyTo5xx` requires
113
- it to hold for all causes.
114
- - The ratio is computed as `StatusCodes.Status5xx / RequestCount` from `ApplicationMetrics`
115
- (raw counts — see below), **not** by parsing the percentage out of the `Causes` text. The
116
- absolute-count floor reads `StatusCodes.Status5xx` directly.
117
- - `describeEnvironmentHealth` is called with
118
- `AttributeNames = ['HealthStatus','Status','Color','Causes','ApplicationMetrics']`.
119
- - Per-environment report + OneUptime payload now include the per-env `alarm` token and the
120
- observed `fivexxRatio`.
121
-
122
- ## How EB enhanced-health is read (durable AWS reference)
123
-
124
- Facts verified against current AWS docs; they inform why the logic above is shaped as it is.
105
+ - **Worker + worker 1.0 monitors** keep the **method defaults** (`100 / 20% / 20`) — the low
106
+ absolute floor suits ~150 req/5min traffic.
107
+ - **API and API-Proxy monitors** (which each also include the low-traffic eu-west-1 env) are
108
+ overridden to **`windowMinutes=5, minRequests=30, min5xxPercent=20, min5xxAbsolute=100`** —
109
+ the lower request floor keeps the small env on the share gate, while the higher absolute
110
+ floor stops the busy envs paging on minor blips (20 scattered 500s on ~1300 req = <2%). The
111
+ override is applied by the dbchanges2 migration below.
112
+
113
+ ### Fail-safe: an unmeasurable solely-5xx event still pages
114
+
115
+ The windowed metrics are **per-instance only** (see below), and some envs never publish them
116
+ (the 1.0 worker env `agilant-worker` publishes only `EnvironmentHealth`). For such an env the
117
+ gate has **no data** and returns **zeros** (not `null`). To avoid silently missing a real
118
+ flood, when EB reports a **solely-5xx `Degraded`/`Severe`** but the window is **empty
119
+ (unmeasurable)**, the monitor **pages** with gate token **`health-5xx-unmeasured`** rather
120
+ than trusting the empty read.
121
+
122
+ A CloudWatch read *failure* (as opposed to empty data) sets **`probe=DEGRADED`** (coverage
123
+ gap), distinct from a measured-bad `alarm=HIGH` — same alarm-vs-probe split as the
124
+ multi-client monitors in [OneUptime push-metric monitors](./oneuptime-worker2-monitoring.md).
125
+
126
+ ## Reading the windowed 5xx rate from CloudWatch (durable AWS reference)
127
+
128
+ - **EB application-request metrics are PER-INSTANCE ONLY.** `ApplicationRequestsTotal`,
129
+ `ApplicationRequests5xx` (and `…2xx/3xx/4xx`) are published under dimensions
130
+ `{EnvironmentName, InstanceId}` in namespace `AWS/ElasticBeanstalk`. **There is no
131
+ environment-only rollup.** You must sum across instances yourself with metric-math over a
132
+ search expression:
133
+
134
+ ```
135
+ SUM(SEARCH('{AWS/ElasticBeanstalk,EnvironmentName,InstanceId} MetricName="ApplicationRequestsTotal" EnvironmentName="<env>"','Sum',<periodSeconds>))
136
+ ```
137
+
138
+ The `SEARCH` auto-discovers every instance's series and `SUM` folds them into one, so the
139
+ window total is correct as instances come and go. Query total and 5xx as two expressions
140
+ over `windowMinutes`.
141
+ - **Not every environment publishes these metrics.** worker-production, api-production-*, and
142
+ apiproxy-production-* all publish them; the **1.0 worker env `agilant-worker` publishes
143
+ ONLY `EnvironmentHealth`** (no app-request metrics) → the gate reads zeros there → the
144
+ `health-5xx-unmeasured` fail-safe above covers it.
145
+ - **IAM already sufficient.** The `WorkloadsRuntime` role grants `cloudwatch:*` in **both**
146
+ account `654654170868` (worker/api) and `502614707982` (legacy worker), so no IAM change
147
+ was needed for the CloudWatch reads.
148
+
149
+ ### EB enhanced-health facts (still relevant — the EB-health path)
125
150
 
126
151
  - **HealthStatus severity ladder:** `Ok → Info → Warning → Degraded → Severe` (plus the grey
127
- states `Pending`/`Unknown`/`Suspended`/`NoData`). `Degraded` is the "high failure" tier
128
- and is **routinely tripped by benign transients** (Auto Scaling scale-up, a mid-deploy
129
- dip), which is exactly why `Degraded` alone is noisy. `Severe` = "very high failure /
130
- environment effectively not serving."
131
- - **`ApplicationMetrics` returns RAW COUNTS, not percentages** — despite the API-reference
132
- prose saying "percentage"/"per second". Fields: `Duration` (Integer seconds, usually 10),
133
- `RequestCount` (Integer, **total** requests over the window),
134
- `StatusCodes.{Status2xx,Status3xx,Status4xx,Status5xx}` (Integer counts). AWS's own
135
- example: 2xx 3391 + 5xx 843 = RequestCount 4234. With no traffic, `RequestCount=0` and
136
- `StatusCodes` may be **absent** → treat as "no data," not "zero failures." The 5xx-ratio
137
- denominator is `RequestCount` (the authoritative total).
152
+ states `Pending`/`Unknown`/`Suspended`/`NoData`). `Degraded` is routinely tripped by benign
153
+ transients (Auto Scaling scale-up, a mid-deploy dip), which is why a solely-5xx `Degraded`
154
+ is now gated on the CloudWatch rate rather than paged outright.
138
155
  - **Why an env is Degraded comes from `Causes`:** EB `Causes` strings literally contain e.g.
139
156
  `"19.9 % of the requests are failing with HTTP 5xx."` A substring check reliably
140
- **classifies** 5xx-driven vs. not (fail-safe: a wording change → treated as a normal
141
- `Degraded` → pages). The **ratio itself must come from `ApplicationMetrics`**, never by
142
- parsing the number out of `Causes`.
143
- - **A failed application DEPLOYMENT is NOT reliably reflected in HealthStatus** — it can read
144
- `Warning`/`Degraded`, or fail before any red request-failure signal. The authoritative
145
- structured deploy signal is per-instance `describeInstancesHealth` →
146
- `Deployment.Status ∈ {'In Progress','Deployed','Failed'}`. (This detector was evaluated and
147
- removed this iteration — see Design history.)
148
- - **IAM:** `describeEnvironmentHealth` / `describeInstancesHealth` / `describeEvents` are all
149
- covered by the managed `AWSElasticBeanstalkReadOnly` policy and require enhanced health
150
- enabled. No per-call charge.
157
+ **classifies** solely-5xx vs. not; a wording change fails safe (treated as non-5xx →
158
+ pages). `describeEnvironmentHealth` is called with `AttributeNames` including
159
+ `HealthStatus`, `Status`, `Color`, `Causes`.
160
+ - **A failed application DEPLOYMENT is NOT reliably reflected in HealthStatus** — the
161
+ authoritative structured deploy signal is per-instance `describeInstancesHealth` →
162
+ `Deployment.Status`. (A deploy that fails without ever degrading health remains an accepted
163
+ residual gap.)
164
+ - **IAM:** `describeEnvironmentHealth` is covered by the managed `AWSElasticBeanstalkReadOnly`
165
+ policy and requires enhanced health enabled.
166
+
167
+ ## Deploy order (mandatory) — code BEFORE the migration
168
+
169
+ worker2 code **must** deploy **before** the dbchanges2 threshold migration runs.
170
+
171
+ - new-code + old-params → **fine** (the params are optional with defaults).
172
+ - **old-code + new-params → FATAL.** The dispatcher spreads `CronJobs.parameters` as **named
173
+ arguments**; an old method signature that lacks `windowMinutes`/`minRequests`/… throws
174
+ **"Unknown named parameter"** and **every run fails** → the monitor goes dark → OneUptime
175
+ raises an offline incident.
176
+
177
+ Correct order: **1) deploy worker2 → 2) apply the SQL → 3) add the OneUptime criterion.**
178
+
179
+ ## dbchanges2 migration (the API/API-Proxy override)
180
+
181
+ `dbchanges2/Core/2026-08-25a - ApiApiproxy Beanstalk Health 5xx Thresholds.sql` applies the
182
+ API/API-Proxy override with a JSON_SET that **adds only the four keys**, preserving
183
+ `awsAccountId`/`environments`/`oneuptimeUrl`:
184
+
185
+ ```sql
186
+ UPDATE CronJobs
187
+ SET parameters = JSON_SET(parameters,
188
+ '$.windowMinutes', 5, '$.minRequests', 30,
189
+ '$.min5xxPercent', 20, '$.min5xxAbsolute', 100)
190
+ WHERE action = 'Infrastructure/CloudWatch/ElasticBeanstalkHealth'
191
+ AND name IN (<the two API monitor names>);
192
+ ```
193
+
194
+ `Core.CronJobs.parameters` is confirmed a native `JSON` column, so `JSON_SET` merges rather
195
+ than overwrites.
151
196
 
152
197
  ## OneUptime monitor config (external, not code)
153
198
 
154
199
  Paging depends on the OneUptime Incoming-Request monitor's **string-match** criteria on the
155
- POSTed body — `Contains "alarm":"HIGH"` (and, if desired, `Contains "probe":"DEGRADED"`).
156
- This is external OneUptime configuration, not worker2 code.
157
-
158
- ## Design history — rejected direction (this session)
200
+ POSTed body — `Contains "alarm":"HIGH"` pages.
159
201
 
160
- An earlier iteration set `HEALTH_ALARM_STATUSES = ['Severe']` (dropping `Degraded` entirely)
161
- and **added a separate deployment-failure detector** (`describeInstancesHealth` →
162
- `Deployment.Status = 'Failed'`) to still catch deploy failures. That was **reverted** after a
163
- scope change: the team decided they **do** want to page on `Degraded` (to catch a flood of
164
- 500s and non-500 degradations), with only the occasional-500 case suppressed via the 80%
165
- ratio. The deployment-failure detector was **removed** — deploy failures that surface as
166
- `Degraded` are now covered by the `Degraded` alarm. **Residual gap accepted for this
167
- iteration:** a deploy that fails fast *without ever degrading health* is not caught.
202
+ **Known alerting gap (recommended fix):** the current monitors match `"status":"error"`,
203
+ `"alarm":"HIGH"`, and heartbeat-gap rules, but **nothing matches `"probe":"DEGRADED"`**. So a
204
+ CloudWatch-metrics **read failure while EB reads Ok** (body = `alarm:OK, probe:DEGRADED,
205
+ status:reporting`) is **invisible to alerting**. Add a criterion matching
206
+ `Contains "probe":"DEGRADED"` → a **Degraded-severity, auto-resolving** incident, placed
207
+ **before** the "online" recovery rule (which also matches `alarm:OK`). The code already emits
208
+ the `probe` token; this is a OneUptime config change the developer applies.
168
209
 
169
210
  ## Known gaps / follow-ups (not done)
170
211
 
171
- - **Per-region client construction is not individually guarded.** Each
172
- `new ElasticBeanstalkClient` is not in its own try/catch, so a bad region/creds aborts the
173
- **whole run** via the outer backstop instead of paging just that region's envs as blind.
174
- Candidate follow-up now that blind reads page.
175
- - **No first-party test harness exists in worker2** (all tests are vendor/).
176
- `shouldPageForHealth()` and `causesAttributeSolelyTo5xx()` are pure and ideal to unit-test
177
- (ratio at exactly 0.80; absolute count 19 vs 20; `Severe` over a solely-5xx cause; a
178
- mixed 5xx + non-5xx cause; empty causes; null metrics; zero requests) — deferred pending a
179
- harness.
212
+ - **No first-party test harness in worker2** (all tests are vendor/). The gate decision and
213
+ `causesAttributeSolelyTo5xx()` are pure and ideal to unit-test (share at exactly
214
+ `min5xxPercent`; absolute at `min5xxAbsolute-1` vs `min5xxAbsolute`; `requests` just below
215
+ `minRequests`; empty/zero window; mixed 5xx + non-5xx cause) — deferred pending a harness.
216
+ - **Deploy fails without degrading health** — still not caught (accepted residual).
180
217
 
181
218
  ## Gotchas / known issues
182
219
 
220
+ - **Deploy worker2 code BEFORE the threshold migration** — old-code + new-params is fatal
221
+ ("Unknown named parameter" on every run → monitor dark). See Deploy order.
183
222
  - **Keep the method non-fatal — never let it throw.** worker2 has no DLQ and a 3600s SQS
184
- visibility timeout, so any uncaught 500 becomes a poison-message storm. The push runs
185
- non-fatal and an outer `catch(\Throwable)` is the backstop.
186
- - **A blind health read must page, not read healthy** — an `AwsException` on
187
- `describeEnvironmentHealth` sets `alarm=HIGH` + `probe=DEGRADED` for that env.
188
- - **`ApplicationMetrics` is raw counts, not percentages** — divide `Status5xx` by
189
- `RequestCount`; `RequestCount=0`/absent `StatusCodes` means "no data," not "zero failures."
190
- - **Classify 5xx-attribution from `Causes` text, but take the ratio from `ApplicationMetrics`
191
- — never parse the percentage out of `Causes`.**
223
+ visibility timeout; an outer `catch(\Throwable)` is the backstop and CloudWatch failures are
224
+ caught internally (return `null`).
225
+ - **EB app-request metrics are per-instance only** — there is no env rollup; you must
226
+ `SUM(SEARCH(...))` across instances or you undercount.
227
+ - **An empty window ≠ zero failures.** An env with no published app-request metrics (e.g. 1.0
228
+ `agilant-worker`) reads zeros; a solely-5xx EB event over an empty window pages via the
229
+ `health-5xx-unmeasured` fail-safe rather than reading healthy.
230
+ - **`minRequests` floor guards the near-idle edge** the old code couldn't — the share gate is
231
+ suppressed until the window has real volume, so a lone 500 in a quiet window no longer pages.
232
+ - **OneUptime has no `probe:DEGRADED` criterion** — a metrics-read failure while EB is Ok is
233
+ currently unalerted until that criterion is added.
192
234
  - **`oneuptimeUrl` is a push credential** — it arrives as a cron parameter; never log it or
193
235
  record its value in a doc.
194
- - **Near-idle ratio edge (residual, accepted).** With the tiny-sample guard removed, a
195
- near-idle env whose window holds a tiny all-error sample (e.g. its only request being a
196
- 500 = 100%) still trips the ratio and pages. Optional future guard: require a minimum 5xx
197
- count before the ratio applies (the absolute floor gates high volume, not this low-volume
198
- edge).
199
236
 
200
237
  ## Change history
238
+ - 2026-08-25 — Replaced the EB `ApplicationMetrics` 5xx gate (a ~10s window that false-paged
239
+ "100% failing" on a lone 500 in a near-empty bucket — real worker-production incident: 16
240
+ 500s / 516 req = 3.1% over 17 min) with a **CloudWatch windowed request-rate gate**. Reads
241
+ true per-env `ApplicationRequestsTotal`/`ApplicationRequests5xx` over a trailing window
242
+ (default 5 min) via `SUM(SEARCH('{AWS/ElasticBeanstalk,EnvironmentName,InstanceId}…','Sum',
243
+ period))` (these metrics are per-instance only — no env rollup). The 5xx page is now
244
+ **decoupled from EB `HealthStatus`**: pages when `5xx >= min5xxAbsolute` OR `requests >=
245
+ minRequests AND share >= min5xxPercent`; a non-5xx `Degraded`/`Severe` still pages
246
+ immediately. Added four optional cron params (`windowMinutes=5, minRequests=100,
247
+ min5xxPercent=20.0, min5xxAbsolute=20`) spread as named arguments; removed
248
+ `shouldPageForHealth()` and the `HTTP_5XX_ALARM_RATIO`/`HTTP_5XX_ALARM_ABSOLUTE_COUNT`
249
+ constants. Added a `health-5xx-unmeasured` fail-safe (solely-5xx over an empty/unpublished
250
+ window pages rather than reading healthy). Per-monitor tuning: worker + worker-1.0 keep
251
+ defaults; API/API-Proxy overridden to `30 / 20% / 100` (dbchanges2
252
+ `2026-08-25a - ApiApiproxy Beanstalk Health 5xx Thresholds.sql`, JSON_SET on the native
253
+ `CronJobs.parameters` JSON column). Recorded the mandatory deploy order (code before the
254
+ migration — old-code+new-params throws "Unknown named parameter") and the OneUptime gap
255
+ (no `probe:DEGRADED` criterion, so a metrics-read failure while EB reads Ok is unalerted).
256
+ WorkloadsRuntime already grants `cloudwatch:*` in accounts 654654170868 and 502614707982 —
257
+ no IAM change. (jcardinal)
201
258
  - 2026-08-19 — Overhauled the 5xx alarm gating in `shouldPageForHealth()`. The 80% share
202
- gate now applies to **both** `Degraded` and `Severe` (was `Degraded`-only, which left the
203
- threshold dead because EB escalates real 5xx floods straight to `Severe` — a 50%-5xx
204
- `Severe` had been paging despite the 0.80 setting). This Severe-gating trade-off was made
205
- deliberately with an independent architecture second opinion on record (cto returned
206
- DISAGREE-WITH-ALTERNATIVE; developer accepted). Narrowed 5xx classification: replaced
207
- `causesAttributeTo5xx()` (any cause mentions 5xx) with `causesAttributeSolelyTo5xx()` (every
208
- cause references 5xx, ≥1) + a `causeIsFivexx()` helper, so a mixed cause (5xx trickle beside
209
- a real non-5xx failure) always pages — a security-review gap the Severe change would have
210
- widened. Added `HTTP_5XX_ALARM_ABSOLUTE_COUNT` (20): a solely-5xx env pages regardless of
211
- share once the absolute 5xx count reaches the floor (closes the ratio's magnitude-blindness,
212
- e.g. 60k of 100k). Removed `HTTP_5XX_MIN_REQUEST_COUNT` (was 20) and its tiny-sample
213
- page-anyway guard — a small number of 5xx is now judged purely on the ratio; below the
214
- absolute floor, page only if share ≥ 0.80, and zero requests → no page. Accepted residual:
215
- a near-idle env with a tiny all-error sample still trips the ratio. (jcardinal)
259
+ gate now applied to **both** `Degraded` and `Severe` (was `Degraded`-only, which left the
260
+ threshold dead because EB escalates real 5xx floods straight to `Severe`). Narrowed 5xx
261
+ classification: replaced `causesAttributeTo5xx()` (any cause mentions 5xx) with
262
+ `causesAttributeSolelyTo5xx()` (every cause references 5xx, ≥1). Added
263
+ `HTTP_5XX_ALARM_ABSOLUTE_COUNT` (20). Removed `HTTP_5XX_MIN_REQUEST_COUNT` tiny-sample
264
+ guard. (jcardinal)
216
265
  - 2026-08-17 — Created. Documented the `ElasticBeanstalkHealth` alarm/paging criteria (page
217
- on `Degraded`/`Severe`; suppress a 5xx-driven `Degraded` unless the 5xx share ≥
218
- `HTTP_5XX_ALARM_RATIO` 0.80 with an `HTTP_5XX_MIN_REQUEST_COUNT` 20 tiny-sample guard;
219
- `Severe` and non-5xx `Degraded` always page), the blind-read fail-safe (`alarm=HIGH` +
220
- `probe=DEGRADED`), the four configurable constants, and the `shouldPageForHealth` /
221
- `causesAttributeTo5xx` helper shape. Captured durable EB enhanced-health facts (severity
222
- ladder; `ApplicationMetrics` returns raw counts not percentages; `Causes` classifies but
223
- `ApplicationMetrics` sets the ratio; failed deploys aren't reliably in HealthStatus —
224
- `describeInstancesHealth.Deployment.Status` is authoritative; `AWSElasticBeanstalkReadOnly`
225
- covers the reads). Recorded the reverted `['Severe']`-only + deployment-failure-detector
226
- direction and the accepted residual gap (a deploy that fails without degrading health).
227
- (jcardinal)
266
+ on `Degraded`/`Severe`; suppress a 5xx-driven `Degraded` unless share ≥ ratio), the
267
+ blind-read fail-safe (`alarm=HIGH` + `probe=DEGRADED`), and durable EB enhanced-health
268
+ facts (severity ladder; `ApplicationMetrics` raw counts; `Causes` classifies; failed
269
+ deploys aren't reliably in HealthStatus). (jcardinal)
270
+ </content>
271
+ </invoke>
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "toga-ai",
3
- "version": "1.0.650",
3
+ "version": "1.0.651",
4
4
  "description": "TOGA Technology Team Claude Knowledge System — shared AI coding harness with skills, knowledge base CLI, and project installer for Claude Code.",
5
5
  "keywords": [
6
6
  "claude",