toga-ai 1.0.478 → 1.0.479
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/knowledge/2.0/apps/_underscore/INDEX.md +1 -1
- package/knowledge/2.0/apps/_underscore/features/email-send-pipeline.md +1 -0
- package/knowledge/2.0/apps/_underscore/features/per-client-database-connections.md +36 -1
- package/knowledge/2.0/apps/worker2/INDEX.md +2 -1
- package/knowledge/2.0/apps/worker2/features/all-client-email-queue-monitor.md +170 -0
- package/knowledge/2.0/apps/worker2/features/oneuptime-worker2-monitoring.md +19 -1
- package/knowledge/INDEX.md +1 -1
- package/package.json +1 -1
|
@@ -24,7 +24,7 @@
|
|
|
24
24
|
| [_Model magic-field access (__get without __isset)](features/model-magic-field-access.md) | `_Model` exposes DB columns as "magic" properties via `__get()`, but it defines **no** `__isset()`. | _underscore/Model/Core/Model.php |
|
|
25
25
|
| [_Model::save() vs raw _Query — no atomic conditional update](features/model-save-vs-query-atomic-update.md) | `_Model::save()` is a plain load-then-write ORM primitive and **cannot express an atomic conditional update** (an optimistic-concurrency / row-claim guard such | _underscore/Model.php, _underscore/Query.php |
|
|
26
26
|
| [NetSuite REST Client (_Component_Api_Netsuite) — record writes & SuiteQL](features/netsuite-rest-client.md) | `_Component_Api_Netsuite` is the **2.0 `_underscore` NetSuite REST client** — the shared primitive every worker2/api2 NetSuite caller uses for record GETs, Suit | _underscore/Component/Api/Netsuite/Netsuite.php |
|
|
27
|
-
| [Per-Client Database Connections & the Local Logs Trap](features/per-client-database-connections.md) | When `_underscore` serves a request for a client it opens **three distinct per-client database connections**, not one. | _underscore/Database.php, _underscore/ApiRequest.php, _underscore/Model/Client/Logs/Api.php, api2/Controller/Index.php |
|
|
27
|
+
| [Per-Client Database Connections & the Local Logs Trap](features/per-client-database-connections.md) | When `_underscore` serves a request for a client it opens **three distinct per-client database connections**, not one. | _underscore/Database.php, _underscore/Query.php, _underscore/ApiRequest.php, _underscore/Model/Client/Logs/Api.php, api2/Controller/Index.php |
|
|
28
28
|
| [Persona Name Translation (PersonaTranslations sidecar)](features/persona-name-translation.md) | Serves Persona **names** in multiple languages by adding a per-language **sidecar** table `PersonaTranslations`, reusing the platform's existing metadata-driven | _underscore/Model/Client/PersonaTranslation.php, dbchanges2/Client/2026-07-22b - PersonaTranslations.sql, dbchanges2/Core/2026-07-22a - PersonaTranslationsRecord.sql, dbchanges2/Client/2026-07-22c - PersonaTranslationsAcl.sql, dbchanges2/Client_CompassCanada/2026-07-22 - PersonaTranslationsFrench.sql, toga2-commerce/src/pages/Account/view/MySettingsView.tsx |
|
|
29
29
|
| [Recursive Item Fulfillments (upstream mirroring)](features/recursive-item-fulfillments.md) | In a multi-tier supply chain a sales order (SO) spawns a purchase order (PO) that becomes another SO downstream, and so on. | _underscore/Model/Client/ItemFulfillment.php, _underscore/Model/Client/ItemFulfillmentItem.php, _underscore/Model/Client/ItemFulfillmentItemUnit.php, _underscore/Model/Client/ItemFulfillmentPackage.php, _underscore/Model/Compass/AdvanceShippingNotice.php, dbchanges2/Core/2026-02-13 - 75601 - RecursiveItemFulfillmentCreation.sql, dbchanges2/Core/2026-06-04 - RecursiveItemFulfillmentPut.sql, dbchanges2/Client_Compass/2026-07-02a - FixSA133377TrackingSerialAndDuplicateIF.sql |
|
|
30
30
|
| [Surface Resolver (_Model_Core_Surface::resolve — replaces Page::meta)](features/surface-resolver.md) | The runtime for the platform-wide **Surface** UI presentation layer: 9 `_underscore` models plus a cached resolver, `_Model_Core_Surface::resolve(&$api, string | _underscore/Model/Core/Surface.php, _underscore/Model/Client/AclRecordScript.php, _underscore/Model/Core/RecordScript.php, dbchanges2/Core/2026-06-30a - SurfaceMetaGroupAndSalesOrderSections.sql, dbchanges2/Client/2026-06-30a - SurfaceMetaGroupAcl.sql, dbchanges2/Client_Quad/2026-07-01a - GrantSurfacesMetaGroupScriptAcl.sql, dbchanges2/Client_CompassCanada/2026-07-01a - GrantSurfacesMetaGroupScriptAcl.sql, dbchanges2/Core/2026-06-29b - SurfaceMetaPublicReadAcl.sql, dbchanges2/Client/2026-06-29c - SurfaceRecordScriptAcl.sql, dbchanges2/Core/2026-06-29c - SurfaceDebugPhpMethodFix.sql, dbchanges2/Client_Compass/2026-07-15f - SalesOrderRecordActionsRemoveDeadConfigRuleOverrides.sql, dbchanges2/Core/2026-07-17h - Update - ClearApprovalsFilterButtonConfig.sql, dbchanges2/Client_Compass/2026-07-17a - SalesOrderApprovalActionsOverride.sql, dbchanges2/Client_Compass/2026-07-17b - SalesOrderApprovalsFilterButtonOverride.sql, dbchanges2/Client_CompassCanada/2026-07-17a - SalesOrderApprovalActionsOverride.sql, dbchanges2/Client_CompassCanada/2026-07-17b - SalesOrderApprovalsFilterButtonOverride.sql, dbchanges2/Client_Quad/2026-07-17a - SalesOrderApprovalActionsOverride.sql, dbchanges2/Client_Quad/2026-07-17b - SalesOrderApprovalsFilterButtonOverride.sql, _underscore/Model/Core/SurfaceElement.php, _underscore/Model/Core/Action.php, _underscore/Model/Core/Vocabulary.php, _underscore/Model/Core/VocabularyTerm.php, _underscore/Model/Core/Message.php, _underscore/Model/Client/SurfaceOverride.php, _underscore/Model/Client/MessageTranslation.php, _underscore/Model/Client/ThemeToken.php, _underscore/Model/Core/Page.php, api2/Component/Api/V2/V2.php |
|
|
@@ -6,10 +6,11 @@ project: _Underscore
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-07-
|
|
9
|
+
updated: 2026-07-30
|
|
10
10
|
owners: ["dfranks", "jcardinal", "mhammontree", "apeterson", "kyalamarthi"]
|
|
11
11
|
files:
|
|
12
12
|
- _underscore/Database.php
|
|
13
|
+
- _underscore/Query.php
|
|
13
14
|
- _underscore/ApiRequest.php
|
|
14
15
|
- _underscore/Model/Client/Logs/Api.php
|
|
15
16
|
- api2/Controller/Index.php
|
|
@@ -17,6 +18,7 @@ related:
|
|
|
17
18
|
- ../architecture.md
|
|
18
19
|
- ../workflows/local-db-refresh-from-beta.md
|
|
19
20
|
- ./database-alias-repointing.md
|
|
21
|
+
- ../../worker2/features/all-client-email-queue-monitor.md
|
|
20
22
|
---
|
|
21
23
|
|
|
22
24
|
## Summary
|
|
@@ -112,6 +114,31 @@ here — they live in `Config/*.ini`.)
|
|
|
112
114
|
`Core.Clients.clientDatabaseId` → `Core.Databases.name`. This mirrors the log-database resolution
|
|
113
115
|
already done in `_underscore`'s `Email.php`. A string-built name fails only for the clients whose
|
|
114
116
|
names happen to diverge, so it passes local testing and breaks in production.
|
|
117
|
+
- **The same rule applies to the LOG database: never build `'Logs_' . $clientIdentifier`.**
|
|
118
|
+
Resolve `Core.Clients.logDatabaseId` → `Core.Databases.name` (`Compass_Usa` → `Logs_Compass`).
|
|
119
|
+
A string-built log-DB name silently targets a schema that does not exist.
|
|
120
|
+
- **`_Database::registerClientDatabases()` is unsafe for a loop over ALL clients.** It INNER
|
|
121
|
+
JOINs the client database *and* archive database alongside the log database, so a client
|
|
122
|
+
missing either row drops out of the result set entirely — and the method then dereferences the
|
|
123
|
+
null row (`$database->clientId`) and throws. It also opens three connections per client when a
|
|
124
|
+
caller may need only one. For an all-client fan-out, run a **log-database-only** variant with
|
|
125
|
+
**LEFT OUTER JOINs** and report clients that resolve to nothing, rather than letting them
|
|
126
|
+
vanish. Worked example: [All-Client Email Queue Monitor](../../worker2/features/all-client-email-queue-monitor.md).
|
|
127
|
+
- **A per-client loop must NOT reuse the shared aliases.** `DB_CLIENT` / `DB_CLIENT_LOGS` /
|
|
128
|
+
`Archive` are single slots — each iteration clobbers the previous one, and all live connection
|
|
129
|
+
state is alias-keyed. Register each client's database **under its real schema name with no
|
|
130
|
+
alias** (guard with `isset(_Database::$_registers[$name])`) when you need many clients open at
|
|
131
|
+
once.
|
|
132
|
+
- **There is no `_Query::setTimeout()`.** Only `_ApiRequest::setTimeout()` exists. A 2.0 request
|
|
133
|
+
or worker therefore has **no per-query timeout**: a hung database cluster blocks until the
|
|
134
|
+
connection itself times out. This matters most in fan-out loops across many client clusters,
|
|
135
|
+
where one unreachable cluster stalls the whole pass.
|
|
136
|
+
- **Classify driver failures by MySQL error NUMBER, not message text.** `_Query` throws a plain
|
|
137
|
+
`Exception` whose message embeds the driver error on an `Error #: <number>` line
|
|
138
|
+
(`_underscore/Query.php:346`). Matching text like `"doesn't exist"` is locale- and
|
|
139
|
+
engine-dependent (MySQL and MariaDB word it differently). Parse `/Error #:\s*(\d+)/` and
|
|
140
|
+
compare to the code — e.g. `1146` (`ER_NO_SUCH_TABLE`) is how a worker tolerates a per-tenant
|
|
141
|
+
table that predates a migration.
|
|
115
142
|
- **Client DBs do not share a charset, so any hardcoded `COLLATE` is invalid somewhere.** Emitting
|
|
116
143
|
e.g. `COLLATE utf8mb4_unicode_ci` against a `latin1` column errors with `COLLATION ... is not
|
|
117
144
|
valid for CHARACTER SET 'latin1'`. For cross-tenant string comparison use `CAST(expr AS BINARY)`
|
|
@@ -152,6 +179,14 @@ here — they live in `Config/*.ini`.)
|
|
|
152
179
|
|
|
153
180
|
## Change history
|
|
154
181
|
|
|
182
|
+
- 2026-07-30 — Added four framework-level gotchas surfaced while building the all-client email
|
|
183
|
+
queue monitor: the log-DB name has the same never-string-build rule as the client DB;
|
|
184
|
+
`registerClientDatabases()` is unsafe for an all-client loop (INNER JOINs on client+archive
|
|
185
|
+
drop clients then throw on the null row, and it opens three connections); a per-client loop
|
|
186
|
+
must register under the **real schema name with no alias** because the aliases are single
|
|
187
|
+
slots; **`_Query::setTimeout()` does not exist**, so 2.0 has no per-query timeout; and driver
|
|
188
|
+
failures must be classified by MySQL **error number** parsed from `Error #: <n>` rather than
|
|
189
|
+
message text. (jcardinal)
|
|
155
190
|
- 2026-07-28 — **Corrected a wrong rule:** a client's schema name is *not* reliably
|
|
156
191
|
`'Client_' . clientIdentifier` (`Compass_Usa` → `Client_Compass`); resolve
|
|
157
192
|
`Clients.clientDatabaseId` → `Databases.name`. Recorded that client DBs do **not** share a
|
|
@@ -4,6 +4,7 @@
|
|
|
4
4
|
|-----|---------|-------|
|
|
5
5
|
| [Worker (worker2) Architecture](architecture.md) | Worker (repo `worker2`) is an AWS Elastic Beanstalk **Worker Tier** application that processes background jobs. | worker2/Controller/Index.php, worker2/Worker/, worker2/LambdaFunctions/, _underscore/Worker.php, worker2/composer.json |
|
|
6
6
|
| [Deploy-Time Auto-Registration to the Shared ALB Target Group (non-production)](features/alb-target-group-auto-registration.md) | TOGA does **not** pay for EB-managed load-balancer registration, so an EB instance is normally **not** added to its environment's ALB target group — a fresh or | worker2/.platform/hooks/postdeploy/060_register_instance_to_shared_application_load_balancer.sh, worker2/ebs/register_instance_to_shared_application_load_balancer.php, worker2/.platform/hooks/prebuild/_shared/040-write-instance-id.sh, worker2/.platform/hooks/prebuild/_shared/041-write-region.sh, worker2/.platform/hooks/postdeploy/015_install_composer.sh, api2/.platform/hooks/postdeploy/060_register_instance_to_shared_application_load_balancer.sh, api2/ebs/register_instance_to_shared_application_load_balancer.php |
|
|
7
|
+
| [All-Client Email Queue Monitor (Monitor/Operations/EmailQueue)](features/all-client-email-queue-monitor.md) | `_Worker_Monitor_Operations::EmailQueue()` is the cross-client health check for the 2.0 outbound email queue (`Logs_<Client>.Email` — see [2.0 Email Send Pipeli | worker2/Worker/Monitor/Operations.php, _underscore/Database.php, _underscore/Query.php, dbchanges2/Logs_Client/2026-05-21 - Email.sql |
|
|
7
8
|
| [Automated PR Merger — Concurrent Force-Push Clobber Race](features/automated-pr-merger-force-push-race.md) | The automated PR merger `_Worker_Team_GitHub::Merge` (`worker2` `Worker/Team/Github.php`) merges approved PRs to `_production` by **force-pushing from a clone t | Worker/Team/Github.php |
|
|
8
9
|
| [ClickUp Connectivity Watchdog](features/clickup-connectivity-watchdog.md) | A cron watchdog that emails when the ClickUp integration looks disconnected during business hours. | worker2/Worker/Clickup/Health.php, worker2/Database/ClickupHealthWatchdog.sql |
|
|
9
10
|
| [ClickUp Design Sprint Automation (Final Design Outcome)](features/clickup-design-sprint-automation.md) | `_Worker_Clickup_Design` is meant to drive the design-sprint workflow in ClickUp via the API, replacing a set of native ClickUp automations. | worker2/Worker/Clickup/Design.php, worker2/Worker/Clickup.php, worker2/Controller/ClickupDesignTest.php, _underscore/Component/Api/Clickup/Clickup.php |
|
|
@@ -24,7 +25,7 @@
|
|
|
24
25
|
| [NetSuite Supporting-Record Webhook Importer (the reusable recipe)](features/netsuite-supporting-record-webhook-importer.md) | A single **repeatable recipe** for porting a legacy daily-pull NetSuite *supporting-record* importer (the lookup/dimension tables behind Forecast2 — Employees, | worker2/Worker/Netsuite/Employee.php, worker2/Worker/Netsuite/Account.php, worker2/Worker/Netsuite/Classification.php, worker2/Worker/Netsuite/Customer.php, worker2/Worker/Netsuite/Item.php, worker2/Worker/Netsuite.php, _underscore/Model/Forecast/Employee.php, _underscore/Model/Forecast/Account.php, _underscore/Model/Forecast/Classification.php, _underscore/Component/Forecast/Db/Db.php, test/@dave/test_employee_lifecycle.php, test/@dave/test_account_lifecycle.php, test/@dave/test_classification_lifecycle.php, test/@dave/NetSuite/api-message-queue/ue_api_msg_queue_enqueue.js, worker/crons/toga2/forecast2/import_supporting_records.php |
|
|
25
26
|
| [Background Email-Template Worker (_Worker_Notification_EmailTemplate)](features/notification-email-template.md) | `_Worker_Notification_EmailTemplate::Send(...)` dispatches a **stored, client-defined `EmailTemplates` row off-thread** as a background WorkerJob. | worker2/Worker/Notification/EmailTemplate.php, worker2/Worker/Client/True.php, _underscore/Model/Client/EmailTemplate.php |
|
|
26
27
|
| [DB-Driven Notification (Internal) Email](features/notification-email.md) | Internal/notification emails (merge-conflict alerts, ops notices — anything system-generated, not client-facing transactional mail) are sent through one worker | worker2/Worker/Notification/Email.php, _underscore/Model/Client/EmailTemplate.php, dbchanges2/Client/2026-06-23a - EmailTemplateWrapper.sql, dbchanges2/Client_True/2026-06-23a - EmailTemplateWrapper.sql |
|
|
27
|
-
| [OneUptime push-metric monitors for 2.0 workers](features/oneuptime-worker2-monitoring.md) | A second, **OneUptime-reporting** monitoring pattern for the 2.0 worker2 tier, ported from the 1.0 `App_SystemMonitor_Compass` monitors. | worker2/Worker/Monitor/Compass.php, worker2/Worker/Client/Compass.php, worker2/composer.json, _underscore/Cloud.php |
|
|
28
|
+
| [OneUptime push-metric monitors for 2.0 workers](features/oneuptime-worker2-monitoring.md) | A second, **OneUptime-reporting** monitoring pattern for the 2.0 worker2 tier, ported from the 1.0 `App_SystemMonitor_Compass` monitors. | worker2/Worker/Monitor/Compass.php, worker2/Worker/Client/Compass.php, worker2/composer.json, _underscore/Cloud.php, worker2/Worker/Monitor/Operations.php |
|
|
28
29
|
| [Platform Cache Cleanup (_Worker_Platform_Cache — Clean + Truncate)](features/platform-cache-cleanup.md) | `_Worker_Platform_Cache` owns maintenance of the shared **Cache** cluster that backs api2's [multi-client data retrieval](../../api2/features/cross-client-data- | worker2/Worker/Platform/Cache.php, worker2/Controller/Index.php, worker2/_.php, dbchanges2/Core/2026-07-27a - PlatformCacheCleanCron.sql |
|
|
29
30
|
| [Startech Webhook Handler (worker2)](features/startech-webhook-handler.md) | Receives inbound webhook events from Startech (Easeedesk) and creates or updates the corresponding ticket in TOGA 2.0. | worker2/Worker/Startech.php |
|
|
30
31
|
| [Talos (TOGa IQ) Meeting-Notes Integration & Token Auto-Refresh (consumer)](features/talos-meeting-notes-integration.md) | How a **dev tool / agent consumes Talos (TOGa IQ)** to query the team meeting-notes corpus programmatically. | .claude/skills/plan-ticket/scripts/talos.js |
|
|
@@ -0,0 +1,170 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: All-Client Email Queue Monitor (Monitor/Operations/EmailQueue)
|
|
3
|
+
framework: "2.0"
|
|
4
|
+
repo: worker2
|
|
5
|
+
project: Worker
|
|
6
|
+
client: shared
|
|
7
|
+
type: feature
|
|
8
|
+
status: active
|
|
9
|
+
updated: 2026-07-30
|
|
10
|
+
owners: ["jcardinal"]
|
|
11
|
+
files:
|
|
12
|
+
- worker2/Worker/Monitor/Operations.php
|
|
13
|
+
- _underscore/Database.php
|
|
14
|
+
- _underscore/Query.php
|
|
15
|
+
- dbchanges2/Logs_Client/2026-05-21 - Email.sql
|
|
16
|
+
related:
|
|
17
|
+
- ./oneuptime-worker2-monitoring.md
|
|
18
|
+
- ./monitoring-framework.md
|
|
19
|
+
- ../../_underscore/features/email-send-pipeline.md
|
|
20
|
+
- ../../_underscore/features/per-client-database-connections.md
|
|
21
|
+
---
|
|
22
|
+
|
|
23
|
+
## Summary
|
|
24
|
+
|
|
25
|
+
`_Worker_Monitor_Operations::EmailQueue()` is the cross-client health check for the 2.0
|
|
26
|
+
outbound email queue (`Logs_<Client>.Email` — see
|
|
27
|
+
[2.0 Email Send Pipeline](../../_underscore/features/email-send-pipeline.md)). It resolves
|
|
28
|
+
**every** client from `Core.Clients`, connects to each client's own logs database, counts
|
|
29
|
+
`FAILED` / `PENDING` rows in `Email`, and pushes one aggregate payload to a single OneUptime
|
|
30
|
+
Incoming Request monitor with the **offending clients named** in the body.
|
|
31
|
+
|
|
32
|
+
It was previously hardcoded to `Logs_Compass`, so a failing email queue for any other client
|
|
33
|
+
was completely invisible. The action path and method signature are unchanged, so the existing
|
|
34
|
+
`Core.WorkerJobs` cron entry needed no change. (Out-of-code follow-up: the OneUptime monitor
|
|
35
|
+
was still named "Compass Email Queue" and needs renaming in the UI.)
|
|
36
|
+
|
|
37
|
+
This is **shared internal infrastructure, not a Compass feature** — Compass USA is simply the
|
|
38
|
+
client it used to be limited to.
|
|
39
|
+
|
|
40
|
+
## Key files / entry points
|
|
41
|
+
|
|
42
|
+
- `worker2/Worker/Monitor/Operations.php` — `abstract class _Worker_Monitor_Operations`.
|
|
43
|
+
- `public static function EmailQueue(): string` — the action, route
|
|
44
|
+
`Monitor/Operations/EmailQueue`.
|
|
45
|
+
- `private static function resolveClientLogDatabases(): array` — one `DB_CORE` query
|
|
46
|
+
returning every client + its log database name + this environment's hosts.
|
|
47
|
+
- `private static function isMissingTableError(Throwable $e): bool` — MySQL error-number
|
|
48
|
+
classification (see gotchas).
|
|
49
|
+
- `const MYSQL_ERROR_NO_SUCH_TABLE = 1146;`
|
|
50
|
+
- No other file in any repo references this class (verified by tree-wide grep) — the
|
|
51
|
+
autoloader and `_Worker::runTask()` reach it indirectly by action path.
|
|
52
|
+
|
|
53
|
+
## How it works
|
|
54
|
+
|
|
55
|
+
1. **Resolve every client's log database.** One query on `_underscore::DB_CORE` joins
|
|
56
|
+
`Clients` → `Databases` (via `Clients.logDatabaseId`) → `DatabaseHosts` →
|
|
57
|
+
`Environments` (slug) → `Regions` (code). Environment slug follows the
|
|
58
|
+
`dev-*` → `dev` collapsing used by `_Worker_Team_Transcripts::initialize()`; region comes
|
|
59
|
+
from `_Cloud::getRegion()`. Rows are ordered so the closest region for a client comes
|
|
60
|
+
first; later rows for the same client are skipped.
|
|
61
|
+
2. **Register each client's log DB under its REAL database name, with no alias.** A per-client
|
|
62
|
+
loop cannot reuse the shared `_underscore::DB_CLIENT_LOGS` alias — every iteration would
|
|
63
|
+
clobber it (all live connection state is alias-keyed; see
|
|
64
|
+
[Re-pointing a DB alias mid-request](../../_underscore/features/database-alias-repointing.md)).
|
|
65
|
+
Registration is guarded with `if (!isset(_Database::$_registers[$logDatabaseName]))`.
|
|
66
|
+
3. **Count per client**, index-friendly:
|
|
67
|
+
`SELECT status, COUNT(*) FROM Email WHERE status IN ('FAILED','PENDING') GROUP BY status`.
|
|
68
|
+
4. **Bucket each client** into `checked`, `skipped` (no `Email` table), `unreachable` (any
|
|
69
|
+
other query/connection failure) or `unresolved` (no log database or no host for this
|
|
70
|
+
environment).
|
|
71
|
+
5. **Decide the tokens and push once** to the same OneUptime monitor.
|
|
72
|
+
|
|
73
|
+
### Alarm rule (explicit product decision)
|
|
74
|
+
|
|
75
|
+
Alarm on **`FAILED > 0` only**. `PENDING` is reported as context and **never** alarms — a
|
|
76
|
+
queue with work in it is normal, and the sender cron runs every minute.
|
|
77
|
+
|
|
78
|
+
### Payload contract — `alarm` and `probe` are SEPARATE tokens
|
|
79
|
+
|
|
80
|
+
OneUptime can only string-match a request body (no numeric comparison, no JSON key
|
|
81
|
+
targeting), so the worker decides and emits tokens:
|
|
82
|
+
|
|
83
|
+
| Field | Meaning |
|
|
84
|
+
|---|---|
|
|
85
|
+
| `status` | `reporting`, or `error` reserved for the **Core lookup itself** failing (checker fully blind) |
|
|
86
|
+
| `alarm` | `HIGH` \| `OK` — emails are actually failing (`failedCount > 0`) |
|
|
87
|
+
| `probe` | `OK` \| `DEGRADED` — set `DEGRADED` when `unreachableClients` is non-empty |
|
|
88
|
+
| `failedCount`, `pendingCount`, `alarmFailedCountThreshold` | raw counts + the threshold used |
|
|
89
|
+
| `offendingClients` | comma-separated client **names** with per-client failed counts |
|
|
90
|
+
| `worstClient`, `worstFailedCount` | the single worst tenant |
|
|
91
|
+
| `clientsChecked`, `clientsSkipped` | coverage counters |
|
|
92
|
+
| `unreachableClients`, `unresolvedClients` | named coverage gaps |
|
|
93
|
+
| `checkedAtUtc` | `gmdate('c')` |
|
|
94
|
+
|
|
95
|
+
**Decision:** on a monitor that covers all clients, *"emails are failing"* and *"we could not
|
|
96
|
+
look at some clients"* need distinct OneUptime criteria. Collapsing them means a logs-cluster
|
|
97
|
+
outage reads as healthy email delivery. Hence two independent tokens rather than one.
|
|
98
|
+
|
|
99
|
+
## Data model
|
|
100
|
+
|
|
101
|
+
`Email` / `EmailAttachment` ship in the `Logs_Client` blank
|
|
102
|
+
(`dbchanges2/Logs_Client/2026-05-21 - Email.sql`), so every client logs database should have
|
|
103
|
+
them. `Email.status` is `enum('PENDING','SENT','FAILED')`, with `retryCount` and
|
|
104
|
+
`failureReason`. Index `idx_email_status` exists on `status`.
|
|
105
|
+
|
|
106
|
+
## Client variations
|
|
107
|
+
|
|
108
|
+
None — the check is uniform across every client in `Core.Clients`. Per-client differences that
|
|
109
|
+
matter are the **log database name** and whether the `Email` migration has been applied.
|
|
110
|
+
|
|
111
|
+
## Gotchas / known issues
|
|
112
|
+
|
|
113
|
+
- **A client's database name is NOT derivable from `clientIdentifier`.** `Compass_Usa` →
|
|
114
|
+
`Logs_Compass`. Always resolve `Clients.logDatabaseId` → `Databases.name`. Building
|
|
115
|
+
`'Logs_' . $clientIdentifier` silently targets a nonexistent database. See
|
|
116
|
+
[per-client database connections](../../_underscore/features/per-client-database-connections.md).
|
|
117
|
+
- **Do not use `_Database::registerClientDatabases()` for an all-client fan-out.** It INNER
|
|
118
|
+
JOINs the client and archive databases alongside the log database, so a client missing either
|
|
119
|
+
row drops out and the method then dereferences a null row and throws. It also opens three
|
|
120
|
+
connections when only the log one is needed. Run a log-database-only variant with LEFT OUTER
|
|
121
|
+
JOINs instead.
|
|
122
|
+
- **INNER JOINs in an all-client resolution query silently un-monitor clients.** A client with
|
|
123
|
+
a null `logDatabaseId`, or with no `DatabaseHosts` row for the current environment, vanishes
|
|
124
|
+
from the result set with no log line and no alert token — the exact blind spot the monitor
|
|
125
|
+
exists to eliminate. LEFT OUTER JOIN every level and report those clients under an explicit
|
|
126
|
+
`unresolvedClients` bucket.
|
|
127
|
+
- **Follow-on trap once left-joined:** a `DatabaseHosts` row belonging to a *different*
|
|
128
|
+
environment shows up as a populated host and looks resolved. Null the host columns unless the
|
|
129
|
+
environment join matched:
|
|
130
|
+
`IF(LogEnvironments.id IS NULL, NULL, LogDatabaseHosts.hostCluster)`.
|
|
131
|
+
- **Detect "table does not exist" by MySQL error NUMBER, not message text.** `_Query` throws a
|
|
132
|
+
plain `Exception` whose message embeds the driver error on an `Error #: <number>` line
|
|
133
|
+
(`_underscore/Query.php:346`). Message wording is locale- and engine-dependent (MySQL vs
|
|
134
|
+
MariaDB differ); parse `/Error #:\s*(\d+)/` and compare to `1146` (`ER_NO_SUCH_TABLE`). A
|
|
135
|
+
logs DB predating the `Email` migration is a **skip**, not an outage.
|
|
136
|
+
- **An unconditional `SUM(status = 'X')` cannot use an index.**
|
|
137
|
+
`SELECT SUM(status='FAILED') FROM Email` full-scans, reading every historical `SENT` row
|
|
138
|
+
despite `idx_email_status`. Use `WHERE status IN (...) GROUP BY status`. This runs per client
|
|
139
|
+
on a cron, so the difference compounds.
|
|
140
|
+
- **There is no `_Query::setTimeout()` in the 2.0 framework** (only `_ApiRequest::setTimeout()`).
|
|
141
|
+
A 2.0 worker therefore has **no per-query timeout**, so a hung database cluster blocks the
|
|
142
|
+
fan-out until the connection itself times out. Relevant to any loop over all clients.
|
|
143
|
+
- **A passing dev run does not prove cross-cluster correctness.** Non-prod is a single
|
|
144
|
+
all-in-one database endpoint; production splits Core / `Client_*` / `Logs_*` / `Archive_*` /
|
|
145
|
+
`Cache` across **separate clusters**.
|
|
146
|
+
- **Verification status of this capture:** `php -l` passes; php-reviewer and sql-reviewer both
|
|
147
|
+
returned 0 critical and every actionable finding was fixed. It has **not** been executed
|
|
148
|
+
against any environment and **no database was queried** — the `Core` join shape is verified
|
|
149
|
+
against framework source only, not against real rows.
|
|
150
|
+
|
|
151
|
+
## Change history
|
|
152
|
+
|
|
153
|
+
- 2026-07-30 — Refactored `EmailQueue()` from Compass-only (hardcoded `Logs_Compass` via
|
|
154
|
+
`initialize()` + `const DB_LOGS_COMPASS`, both removed) to an all-client fan-out
|
|
155
|
+
(121 → 309 lines): added `resolveClientLogDatabases()`, per-client registration under the real
|
|
156
|
+
database name with no alias, LEFT-OUTER-JOIN resolution with an `unresolvedClients` bucket,
|
|
157
|
+
error-number-based missing-table skip, index-friendly `GROUP BY status` counting, and the
|
|
158
|
+
separate `alarm` / `probe` tokens. Alarm on `FAILED > 0` only; `PENDING` never alarms. Action
|
|
159
|
+
path and signature unchanged, so no cron change. Uncommitted in the worker2 working tree at
|
|
160
|
+
capture time. (jcardinal)
|
|
161
|
+
|
|
162
|
+
## Related docs
|
|
163
|
+
|
|
164
|
+
- [OneUptime push-metric monitors for 2.0 workers](./oneuptime-worker2-monitoring.md) — the
|
|
165
|
+
push/token pattern and OneUptime criteria this monitor follows.
|
|
166
|
+
- [Monitoring Framework](./monitoring-framework.md) — the parallel DB-driven, email-alert pattern.
|
|
167
|
+
- [2.0 Email Send Pipeline](../../_underscore/features/email-send-pipeline.md) — what fills the
|
|
168
|
+
queue this monitor watches.
|
|
169
|
+
- [Per-Client Database Connections](../../_underscore/features/per-client-database-connections.md)
|
|
170
|
+
— log-database name resolution.
|
|
@@ -6,15 +6,17 @@ project: Worker
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-07-
|
|
9
|
+
updated: 2026-07-30
|
|
10
10
|
owners: ["jcardinal"]
|
|
11
11
|
files:
|
|
12
12
|
- worker2/Worker/Monitor/Compass.php
|
|
13
13
|
- worker2/Worker/Client/Compass.php
|
|
14
14
|
- worker2/composer.json
|
|
15
15
|
- _underscore/Cloud.php
|
|
16
|
+
- worker2/Worker/Monitor/Operations.php
|
|
16
17
|
related:
|
|
17
18
|
- ./monitoring-framework.md
|
|
19
|
+
- ./all-client-email-queue-monitor.md
|
|
18
20
|
- ../../_underscore/features/cloud-s3-helpers.md
|
|
19
21
|
- ../../../1.0/apps/worker/features/oneuptime-worker-uptime-monitoring.md
|
|
20
22
|
---
|
|
@@ -98,6 +100,17 @@ OneUptime criteria:
|
|
|
98
100
|
prevents flapping back to Operational while the alarm is still HIGH).
|
|
99
101
|
- **not received in 15 min** → Offline.
|
|
100
102
|
|
|
103
|
+
### Separate the ALARM signal from the PROBE-health signal on multi-client monitors
|
|
104
|
+
|
|
105
|
+
When one monitor covers **every** client (rather than one client's integration), "the metric is
|
|
106
|
+
bad" and "we could not look at some clients" need **distinct** OneUptime criteria. Collapsing
|
|
107
|
+
them means a logs-cluster outage reads as healthy delivery. Emit **two independent string
|
|
108
|
+
tokens**: `alarm` (`HIGH`/`OK`) for the measured condition and `probe` (`OK`/`DEGRADED`) for
|
|
109
|
+
coverage gaps, and name the affected clients in the body so the incident is actionable without a
|
|
110
|
+
log dive. `status: "error"` stays reserved for the checker being **fully** blind (its own
|
|
111
|
+
resolution query failed). Worked example:
|
|
112
|
+
[All-Client Email Queue Monitor](./all-client-email-queue-monitor.md).
|
|
113
|
+
|
|
101
114
|
### Heartbeat / cron cadence timing
|
|
102
115
|
|
|
103
116
|
The 10-min-Degraded / 15-min-Offline heartbeat thresholds pair with a **5-minute** cron
|
|
@@ -179,6 +192,11 @@ check for another client.
|
|
|
179
192
|
monitors must pass the region (see [Cloud S3 helpers](../../_underscore/features/cloud-s3-helpers.md)).
|
|
180
193
|
|
|
181
194
|
## Change history
|
|
195
|
+
- 2026-07-30 — Added the multi-client refinement of the payload contract: a **separate `probe`
|
|
196
|
+
token** (`OK`/`DEGRADED`) alongside `alarm`, so a partial logs-cluster outage cannot read as a
|
|
197
|
+
healthy metric, plus naming the offending clients in the body. Recorded that the previously
|
|
198
|
+
Compass-only `Monitor/Operations/EmailQueue` is now an all-client monitor — see
|
|
199
|
+
[All-Client Email Queue Monitor](./all-client-email-queue-monitor.md). (jcardinal)
|
|
182
200
|
- 2026-07-14 — Grew `_Worker_Monitor_Compass` to a 9-monitor suite (S3/DB/mailbox);
|
|
183
201
|
finalized the payload contract (`status`/`alarm`/`checkedAtUtc`) and the not-received
|
|
184
202
|
10-min-Degraded / 15-min-Offline + OK-gated recovery criteria; standardized cron on
|
package/knowledge/INDEX.md
CHANGED
|
@@ -19,7 +19,7 @@ _Auto-generated by `knowledge.js index`. Do not hand-edit._
|
|
|
19
19
|
## 2.0 framework
|
|
20
20
|
|
|
21
21
|
- **_underscore** (_Underscore) _(framework core)_ — 40 doc(s) → [2.0/apps/_underscore/INDEX.md](2.0/apps/_underscore/INDEX.md)
|
|
22
|
-
- **worker2** (Worker) —
|
|
22
|
+
- **worker2** (Worker) — 36 doc(s) → [2.0/apps/worker2/INDEX.md](2.0/apps/worker2/INDEX.md)
|
|
23
23
|
- **api2** (API) — 20 doc(s) → [2.0/apps/api2/INDEX.md](2.0/apps/api2/INDEX.md)
|
|
24
24
|
- **dbchanges2** (Database Changes) _(framework core)_ — 3 doc(s) → [2.0/apps/dbchanges2/INDEX.md](2.0/apps/dbchanges2/INDEX.md)
|
|
25
25
|
- **toga2-supply** (TOGa Supply) — 3 doc(s) → [2.0/apps/toga2-supply/INDEX.md](2.0/apps/toga2-supply/INDEX.md)
|
package/package.json
CHANGED