toga-ai 1.0.645 → 1.0.647

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (23) hide show
  1. package/knowledge/1.0/apps/library/features/cron-execution-monitoring.md +25 -3
  2. package/knowledge/1.0/apps/worker/INDEX.md +2 -1
  3. package/knowledge/1.0/apps/worker/architecture.md +35 -2
  4. package/knowledge/1.0/apps/worker/features/oneuptime-worker-uptime-monitoring.md +40 -3
  5. package/knowledge/1.0/apps/worker/workflows/tracing-a-worker-cron-run-in-production.md +124 -0
  6. package/knowledge/2.0/apps/_underscore/features/email-template-sending.md +17 -0
  7. package/knowledge/2.0/apps/_underscore/features/per-client-database-connections.md +17 -2
  8. package/knowledge/2.0/apps/api2/architecture.md +56 -8
  9. package/knowledge/2.0/apps/worker2/INDEX.md +2 -1
  10. package/knowledge/2.0/apps/worker2/features/cross-account-aws-access.md +42 -3
  11. package/knowledge/2.0/apps/worker2/features/oneuptime-worker2-monitoring.md +39 -0
  12. package/knowledge/2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md +175 -0
  13. package/knowledge/2.0/standards/backend-php.md +18 -2
  14. package/knowledge/INDEX.md +2 -2
  15. package/knowledge/clients/compass-usa/INDEX.md +1 -0
  16. package/knowledge/clients/compass-usa/features/odp-edi-850-item-resolution.md +179 -0
  17. package/knowledge/clients/compass-usa/features/odp-edi-855-acknowledgement-and-overquantity-guard.md +29 -2
  18. package/knowledge/clients/compass-usa/profile.md +10 -1
  19. package/knowledge/clients/compass-usa/workflows/odp-order-pipeline-to-netsuite.md +39 -2
  20. package/knowledge/clients/pcmaticb2b/INDEX.md +1 -1
  21. package/knowledge/clients/pcmaticb2b/features/startech-ticket-sync.md +5 -2
  22. package/knowledge/clients/pcmaticb2b/profile.md +4 -2
  23. package/package.json +1 -1
@@ -6,12 +6,13 @@ project: Library
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-07-06
10
- owners: [dfranks]
9
+ updated: 2026-08-25
10
+ owners: [dfranks, bala]
11
11
  files:
12
12
  - library/app/framework.php
13
13
  related:
14
14
  - ../../worker/architecture.md
15
+ - ../../worker/workflows/tracing-a-worker-cron-run-in-production.md
15
16
  ---
16
17
 
17
18
  ## Summary
@@ -38,13 +39,28 @@ row in `db_log` and of the per-job Sentry check-in monitors.
38
39
  cast silently discards the sub-second precision even when the schema column is `decimal(10,3)`.
39
40
  - **Schema:** the table is provisioned by the dbchanges migration
40
41
  `dbchanges/Common/DF/2026-5-7 Cron Checkin.sql` (table `CronJobExecutions`,
41
- `executionTimeSeconds decimal(10,3)`, no `note` column).
42
+ `executionTimeSeconds decimal(10,3)`). **Correction (2026-08-25): prod `Common.CronJobExecutions`
43
+ does have a `note` TEXT column** — verified by writing to it and reading it back. `App_Framework`
44
+ never populates it, which makes it the practical place to park **temporary** step tracing for a
45
+ cron you are diagnosing (worker crons have no readable stdout — see
46
+ [Tracing a worker cron run in production](../../worker/workflows/tracing-a-worker-cron-run-in-production.md)).
47
+ Remove the tracing when you are done, and never write payloads or credentials into it.
48
+ - **Don't confuse the two `note`s.** `cronInitialization()` writes `note = 'Started execution'` to
49
+ the **`Log`** table on `db_log` (the older `CRON`/`recordType` row), not to `CronJobExecutions`.
42
50
  - **Design decision:** 1.0 records executions by writing **directly to the shared DB**
43
51
  (`db_common`), *not* by POSTing from 1.0 to a 2.0 API endpoint. Per Jeff Cardinal, 1.0 code
44
52
  must not POST to 2.0 code; the direct shared-DB write is the sanctioned path (TRUE-78182).
45
53
 
46
54
  ## Gotchas
47
55
 
56
+ - **⚠ A run skipped by the overlap guard leaves NO ROW — absence is ambiguous.**
57
+ `cronInitialization()` calls `exitIfProcessRunning()` **before** the `CronJobExecutions` INSERT,
58
+ and the `cronFinished(true)` that guard then calls does nothing because `cronLogId` was never
59
+ set. So a blocked run is indistinguishable from a run that never launched: **a missing row for
60
+ an expected slot means "skipped", not "hung"**, and a hung run looks different (it has
61
+ `dtCheckIn` with no `dtCheckOut`). `isProcessRunning()` matches **any** `ps -ef` line containing
62
+ the script path (it only excludes `/bin/sh` and its own pid), so a stray `tail -f` or editor on
63
+ that path blocks the cron indefinitely.
48
64
  - **`db_log` has no `CronJobExecutions` table.** A `cronFinished()` UPDATE that runs against the
49
65
  `db_log` connection throws on **every** 1.0 cron job. Both the INSERT and the UPDATE must
50
66
  target `db_common`.
@@ -60,6 +76,12 @@ row in `db_log` and of the per-job Sentry check-in monitors.
60
76
  logic. (Fixed 2026-07-06: consolidated back to one `db_common` INSERT/UPDATE pair.)
61
77
 
62
78
  ## Change history
79
+ - 2026-08-25 — Corrected the schema note (prod `CronJobExecutions` **does** carry a `note` TEXT
80
+ column, unused by `App_Framework`, usable for temporary cron tracing) and separated it from the
81
+ `note = 'Started execution'` row `cronInitialization()` writes to `Log` on `db_log`. Recorded
82
+ that the overlap guard runs **before** the INSERT, so a skipped run writes no row at all and
83
+ cannot be told apart from "never launched" — and that `isProcessRunning()`'s `ps -ef` substring
84
+ match lets any process holding the script path block the cron. No code change. (bala)
63
85
  - 2026-07-06 — Fixed a merge (`origin/_production`) that reintroduced duplicate
64
86
  `CronJobExecutions` writes and pointed one `cronFinished()` UPDATE at `db_log` (where the
65
87
  table doesn't exist), which would have thrown on every 1.0 cron. Consolidated to one
@@ -11,9 +11,10 @@
11
11
  | [NetSuite Sales Order Sales Rep Sourcing (Staples & ODP EDI orders)](features/netsuite-sales-order-sales-rep-sourcing.md) | How the **sales rep** on a NetSuite Sales Order is determined for the two 1.0 `worker` EDI order-creation integrations (Staples cXML and Compass/ODP EDI). | worker/crons/toga2/compass/workflow/5_create_netsuite_sales_orders_from_office_depot_purchase_orders.php, worker/crons/sync/staples/sync_staples_cxml.php, test/@Mark/NetSuite/TRUE_80451_customer_salesrep_diag.php |
12
12
  | [NetSuite → TOGa Supply Per-Client Sync (thin wrappers)](features/netsuite-togasupply-per-client-sync.md) | Syncs NetSuite transactions (sales orders, purchase orders, invoices, item receipts, item fulfillments, inventory adjustments) into each TOGa Supply (2.0) clien | worker/crons/toga2/netsuite/common_sync_togasupply.php, worker/crons/toga2/netsuite/sync_togasupply_canon.php, worker/crons/toga2/netsuite/sync_togasupply_elite.php, worker/schedules/cron.worker.sync.json, dbchanges2/_modules/netsuite/2026-04-01 - Parameters.sql, library/app/api/toga2.php, library/app/api/netsuite/rest.php, library/app/framework.php, library/app/systemmonitor/netsuiteintegration.php, test/@srija/Elite Testing/Service Requests/test_sync_togasupply_elite_section.php, test/@srija/Elite Testing/Service Requests/test_diagnose_togasupply_elite.php |
13
13
  | [OneUptime Server monitor + disk/memory hygiene on the 1.0 worker EB host](features/oneuptime-server-monitor-host-hygiene.md) | The 1.0 `agilant-worker` EB environment runs on the **legacy Amazon Linux 1 PHP 7.2 platform** (Apache httpd/prefork, s3fs mounts, cron) and repeatedly went dow | worker/.ebextensions/040_disk_memory_hygiene.config, worker/.ebextensions/045_oneuptime_agent.config, worker/ebs/cron.worker.php, worker/ebs/mount-s3fs-folders.php, worker/ebs/apache_settings.php, worker/ebs/setup_phpini.php |
14
- | [OneUptime external uptime monitoring for 1.0 workers](features/oneuptime-worker-uptime-monitoring.md) | Every 1.0 worker box self-reports its liveness to an external OneUptime monitor once per minute by curl-POSTing to a per-worker "Incoming Request" heartbeat URL | library/app/worker.php, worker/crons/worker/worker_heartbeat.php |
14
+ | [OneUptime external uptime monitoring for 1.0 workers](features/oneuptime-worker-uptime-monitoring.md) | Every 1.0 worker box self-reports its liveness to an external OneUptime monitor once per minute by curl-POSTing to a per-worker "Incoming Request" heartbeat URL | library/app/worker.php, worker/crons/worker/worker_heartbeat.php, worker/ebs/cron.worker.php, worker/.ebextensions/045_oneuptime_agent.config |
15
15
  | [Prudential: Send Shipments for the Day report (daily cron)](features/send-shipments-for-the-day.md) | Daily cron (9:00 PM) that emails Prudential and Dell stakeholders an Excel report of all devices shipped that day, including tracking number, serial number, emp | worker/crons/notifications/reports/send_shipments_for_the_day.php |
16
16
  | [Staples cXML Order Import (SFTP → NetSuite)](features/staples-cxml-order-import.md) | `sync_staples_cxml.php` is an **hourly** cron (runs at **:45**) that imports Staples cXML purchase orders from SFTP into NetSuite as Sales Orders, then writes a | worker/crons/sync/staples/sync_staples_cxml.php |
17
17
  | [Diagnosing frozen 1.0 worker cron check-ins (Sentry "missed" flood)](workflows/diagnosing-frozen-cron-checkins.md) | When 1.0 worker cron timestamps freeze and Sentry project `worker1` fills with **`missed`** check-ins, the intuitive diagnosis — a wedged `App_Framework::isProc | worker/.ebextensions/cron.config, library/app/worker.php |
18
18
  | [isFulfillable Multi-Client Backfill (all togasupply clients)](workflows/isfulfillable-multi-client-backfill.md) | One-time backfill that catches up `Items.isFulfillable` on **existing** items across **all 17 togasupply clients** (AIG, Broward Sheriff, Canon, Endeavor Health | worker/crons/toga2/netsuite/backfill_isfulfillable_all_clients.php, library/app/api/toga2.php |
19
19
  | [Onboarding a Client to the NetSuite TOGa Supply Sync](workflows/onboarding-client-to-netsuite-togasupply-sync.md) | How to add a new TOGa 2 client to the per-client NetSuite → TOGa Supply importer (`worker/crons/toga2/netsuite/`). | worker/crons/toga2/netsuite/sync_togasupply.php, worker/crons/toga2/netsuite/common_sync_togasupply.php, worker/crons/toga2/netsuite/sync_togasupply_elite.php, worker/schedules/cron.worker.sync.json, dbchanges2/_modules/netsuite/2026-04-01 - Parameters.sql, dbchanges2/_modules/netsuite/2026-08-05 - CLEAN NETSUITE CLINET.SQL |
20
+ | [Tracing a 1.0 worker cron run in production (no stdout, silent skips)](workflows/tracing-a-worker-cron-run-in-production.md) | How to answer *"did this cron actually run, and what did it do?"* on the 1.0 `worker` tier, where **there is no usable stdout** and **a skipped run leaves no tr | worker/ebs/cron.worker.php, worker/schedules/cron.worker.sync.json, library/app/framework.php |
@@ -6,8 +6,8 @@ project: Worker
6
6
  client: shared
7
7
  type: architecture
8
8
  status: active
9
- updated: 2026-07-29
10
- owners: [jcardinal, sking]
9
+ updated: 2026-08-24
10
+ owners: [jcardinal, sking, bala]
11
11
  files:
12
12
  - worker/index.php
13
13
  - worker/_/app/framework.php
@@ -20,6 +20,8 @@ files:
20
20
  related:
21
21
  - ../library/architecture.md
22
22
  - ../dbchanges/workflows/authoring-and-shipping-sql-files.md
23
+ - ./features/oneuptime-worker-uptime-monitoring.md
24
+ - ../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md
23
25
  ---
24
26
 
25
27
  ## Summary
@@ -44,6 +46,7 @@ which **self-elects a distinct role** (`notification`, `database`, `infrastructu
44
46
  under `crons/`; the schedule (cron registration) is the source of truth for what runs — a script
45
47
  that isn't scheduled never executes. Client integrations live under `crons/toga2/<client>/`.
46
48
  Use prepared statements for all SQL; never interpolate input.
49
+ A worker that fails deploy-time role election runs NO crons while EB still reads healthy.
47
50
 
48
51
  ## How a job becomes a cron (the dispatch pipeline)
49
52
 
@@ -108,6 +111,32 @@ is the most important and least obvious part of the architecture.
108
111
  > Net effect: roles are a **claim-the-first-free-slot pool**, and a dead worker's role is
109
112
  > reclaimed by its replacement within minutes. There is no static instance→role mapping to edit.
110
113
 
114
+ ### The failure mode: NO NAME = NO CRONTAB (and it is silent)
115
+
116
+ The election is the **single point of failure for an entire instance**, and the architecture gives
117
+ it no second chance:
118
+
119
+ - The claim runs **only at DEPLOY time** — `ebs/cron.worker.php` is written as the EB `appdeploy`
120
+ enact hook. **It never retries.** An instance that fails to claim stays nameless until the next
121
+ deploy or until something terminates it.
122
+ - The claim writes `/etc/worker-role` and **only then** appends `cron.worker.<role>.json` to the
123
+ crontab. So a nameless instance has **no role crontab at all** — it runs *nothing*, not even
124
+ `worker_heartbeat.php`, while **EB reports the environment healthy**. The box is up; it just
125
+ does no work.
126
+ - **The claim's DB connect is `@mysqli_connect(...)` — error-suppressed.** A boot-time connect
127
+ failure to `Vision_Log` therefore leaves the instance nameless **with no log line anywhere**.
128
+ This is the prime suspect for the 2026-08-24 incident (2 of 7 workers roleless for ~3 hours),
129
+ and it violates the no-`@`-suppression rule in `rules/toga/common/coding-style.md`. Removing the
130
+ `@` and logging the failure is the recommended fix.
131
+ - `Workers` **only ever holds instances that SUCCEEDED**, so a failed claim leaves no row —
132
+ **absence is the only evidence**, and detecting it requires the AWS instance list.
133
+
134
+ Because every OneUptime monitor on this tier is keyed by **role** rather than by instance, none of
135
+ them can see this (see
136
+ [the blind spot](./features/oneuptime-worker-uptime-monitoring.md)). External coverage now comes
137
+ from the 2.0
138
+ [Worker Fleet Role-Assignment Monitor](../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md).
139
+
111
140
  ## Anatomy of a cron script
112
141
 
113
142
  Every script is self-contained and follows this boilerplate:
@@ -208,6 +237,10 @@ web face (`mvc/` GET routes for login/logout/404); the tier's real work is the c
208
237
 
209
238
  ## Conventions & gotchas
210
239
 
240
+ - **A worker that fails role election is silently dead, not degraded.** No name → no crontab →
241
+ the instance runs nothing while EB reads healthy, and the `@`-suppressed `mysqli_connect` in
242
+ `ebs/cron.worker.php` logs nothing. Never treat "EB is green" as evidence the fleet is working;
243
+ check `Vision_Log.Workers` against the actual EB instance list.
211
244
  - **`dbchanges` auto-apply is branch-gated, and `_production` is NOT auto-applied.**
212
245
  `crons/infrastructure/execute_dbchanges.php` runs every 2 minutes but only for branches matching
213
246
  `_%`, deriving the env by stripping the leading `_` and requiring `config.<env>.ini` in the
@@ -6,12 +6,17 @@ project: Worker
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-07-13
10
- owners: ["jcardinal"]
9
+ updated: 2026-08-24
10
+ owners: ["jcardinal", "bala"]
11
11
  files:
12
12
  - library/app/worker.php
13
13
  - worker/crons/worker/worker_heartbeat.php
14
- related: ["oneuptime-server-monitor-host-hygiene"]
14
+ - worker/ebs/cron.worker.php
15
+ - worker/.ebextensions/045_oneuptime_agent.config
16
+ related:
17
+ - ./oneuptime-server-monitor-host-hygiene.md
18
+ - ../architecture.md
19
+ - ../../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md
15
20
  ---
16
21
 
17
22
  ## Summary
@@ -57,6 +62,31 @@ Database, Infrastructure, TOGa, TOGa Desk, and Catalog, then importing each into
57
62
  manually. To add a new worker's monitor, clone an existing monitor export the same way and
58
63
  register the resulting heartbeat URL in `App_Worker::$oneUptimeEndpoints`.
59
64
 
65
+ ## THE BLIND SPOT — every monitor here is keyed by ROLE, so a ROLELESS instance is invisible
66
+
67
+ Load-bearing limitation, learned the hard way (2026-08-24: **2 of the 7 production workers ran
68
+ roleless for ~3 hours undetected**). This layer, and the Server monitors in
69
+ [host hygiene](./oneuptime-server-monitor-host-hygiene.md), are keyed by `workerName` — **14
70
+ monitors, none keyed by instance**:
71
+
72
+ - `worker_heartbeat.php` only pings `App_Worker::$oneUptimeEndpoints[$row['workerName']]` for
73
+ the row matching **its own** `instanceId`. No row → no ping target → it pushes nothing.
74
+ - The `045_oneuptime_agent.config` Server agent reads `/etc/worker-role` and **exits clean when
75
+ it is empty.**
76
+
77
+ So an instance that never claimed a role reports to nothing — and because every role is still
78
+ owned by *someone*, **no role-keyed monitor is missing a ping either.** The fleet reads 100%
79
+ healthy while a box does zero work (role assignment happens **only at deploy time**, and no role
80
+ means the role's `cron.worker.<role>.json` is never appended to the crontab, so the instance runs
81
+ **nothing**).
82
+
83
+ `Vision_Log.Workers` only ever holds instances that **succeeded** in claiming a role, so the
84
+ failure leaves no row and no log line — **absence is the only evidence**, and detecting an
85
+ absence requires the AWS instance list, which nothing on this tier consults. That gap is what
86
+ the 2.0
87
+ [Worker Fleet Role-Assignment Monitor](../../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md)
88
+ now closes from outside the fleet. **Do not try to close it with another role-keyed monitor.**
89
+
60
90
  ## Gotchas
61
91
 
62
92
  - Do **not** hardcode the heartbeat URLs/UUIDs anywhere but `App_Worker::$oneUptimeEndpoints`
@@ -67,6 +97,13 @@ register the resulting heartbeat URL in `App_Worker::$oneUptimeEndpoints`.
67
97
  pending cleanup — the live source of truth is the `App_Worker::$oneUptimeEndpoints` map.
68
98
 
69
99
  ## Change history
100
+ - 2026-08-24 — Documented **the blind spot**: all 14 monitors on this tier are keyed by
101
+ `workerName`, so an instance that claimed **no** role pings nothing and leaves every role-keyed
102
+ monitor still green — which is how 2 of 7 production workers ran roleless for ~3 hours
103
+ undetected. Recorded that `Vision_Log.Workers` only holds successful claims (absence is the
104
+ only evidence) and that the gap is now covered from outside the fleet by the 2.0
105
+ [Worker Fleet Role-Assignment Monitor](../../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md).
106
+ (bala)
70
107
  - 2026-07-13 — Added external OneUptime heartbeat push for all 1.0 workers on top of the
71
108
  existing internal DB-heartbeat mechanism; endpoints keyed by workerName in
72
109
  `App_Worker::$oneUptimeEndpoints`. (jcardinal)
@@ -0,0 +1,124 @@
1
+ ---
2
+ title: Tracing a 1.0 worker cron run in production (no stdout, silent skips)
3
+ framework: "1.0"
4
+ repo: worker
5
+ project: Worker
6
+ client: shared
7
+ type: workflow
8
+ status: active
9
+ updated: 2026-08-25
10
+ owners: ["bala"]
11
+ files:
12
+ - worker/ebs/cron.worker.php
13
+ - worker/schedules/cron.worker.sync.json
14
+ - library/app/framework.php
15
+ related:
16
+ - ../architecture.md
17
+ - ./diagnosing-frozen-cron-checkins.md
18
+ - ../../library/features/cron-execution-monitoring.md
19
+ - ../../../1.0/standards/backend-php.md
20
+ ---
21
+
22
+ ## Summary
23
+ How to answer *"did this cron actually run, and what did it do?"* on the 1.0 `worker` tier, where
24
+ **there is no usable stdout** and **a skipped run leaves no trace at all**. Written from a
25
+ production diagnostic session on
26
+ `worker/crons/toga2/compass/workflow/3a_import_office_depot_purchase_orders.php`; the mechanics
27
+ apply to every cron on the tier.
28
+
29
+ Use this before you conclude a cron "never fired" — three different situations look identical
30
+ from the outside.
31
+
32
+ ## Step 1 — get the real schedule from `schedules/*.json`, never from the file
33
+ The `.php` file's own header comment is not authoritative and is routinely stale. So are skills,
34
+ docs and tickets. `worker/schedules/cron.<env>[.<role>].json` is the only source of truth.
35
+
36
+ Worked example: the entry *"Download EDI From S3 Files & Create PO - EDI Office Depot"* in
37
+ `cron.worker.sync.json` is `*/5 * * * *` — **every 5 minutes**. The cron file's header said
38
+ "EVERY HOUR" (corrected 2026-08-25) and at least one skill still says hourly. A wrong assumed
39
+ frequency makes you read a normal gap as an outage.
40
+
41
+ Also check whether the job is registered in the env you are testing in at all: this one is **not**
42
+ in `cron.beta.json`, so beta never runs it on a schedule. Nothing on beta is evidence about it.
43
+
44
+ ## Step 2 — echo/print goes nowhere useful, so do not debug with it
45
+ The crontab line is assembled in `worker/ebs/cron.worker.php`, and the two blocks that build it
46
+ differ:
47
+
48
+ | Schedule file | Command written | Where output goes |
49
+ |---|---|---|
50
+ | base `cron.<env>.json` | `php /var/www/html/crons/<script>` | **nowhere** — the ` > /var/www/cache/WORKER_ERROR_$RANDOM` redirect is **commented out** |
51
+ | role `cron.worker.<role>.json` | `php /var/www/html/crons/<script> > /var/www/cache/WORKER_ERROR_$RANDOM` | a **randomly named** file under `/var/www/cache/` on whichever of the 7 instances ran it |
52
+
53
+ Either way you cannot go and read it: base-schedule output is discarded outright, and a
54
+ role-schedule run leaves an unpredictable filename on a box in an autoscaled fleet. **Treat
55
+ worker cron stdout as write-only.**
56
+
57
+ ## Step 3 — put temporary tracing somewhere you can query
58
+ What worked: write timestamped step markers into the **`note`** column of the run's
59
+ `Common.CronJobExecutions` row (the row `App_Framework::cronInitialization()` already inserted),
60
+ then read them back over the DB tooling from your desk. `dtCheckIn`, `dtCheckOut` and
61
+ `executionTimeSeconds` on the same row give you the shape of the run for free.
62
+
63
+ Rules: keep it to a few markers, and **remove the tracing once diagnosed** (it was removed in
64
+ this session). Never write payloads or credentials into `note`.
65
+
66
+ ## Step 4 — read `Common.CronJobExecutions` to tell "ran", "hung" and "skipped" apart
67
+ Query the legacy env's `Common.CronJobExecutions` filtering `job LIKE` the script path:
68
+
69
+ - **`dtCheckIn` + `dtCheckOut` + `executionTimeSeconds`** — it ran and finished; the duration
70
+ tells you whether it fits inside its schedule slot.
71
+ - **`dtCheckIn` with no `dtCheckOut`** — it started and never completed (crash, timeout, or the
72
+ instance went away).
73
+ - **no row for a slot you expected** — the run was **skipped by the overlap guard**, not "never
74
+ fired". A blocked run writes **nothing**: `cronInitialization()` calls
75
+ `exitIfProcessRunning()` **before** the INSERT, and the `cronFinished(true)` it then calls is a
76
+ no-op because `cronLogId` is unset. See
77
+ [Cron Execution Monitoring](../../library/features/cron-execution-monitoring.md).
78
+
79
+ ### The overlap guard is a `ps -ef` substring match — anything can block it
80
+ `App_Framework::isProcessRunning()` returns true for **any** `ps -ef` line containing the script
81
+ path (it only excludes `/bin/sh` lines and its own pid). So a developer's `tail -f`, an editor, or
82
+ any shell holding that path in its command line **blocks the cron indefinitely** while looking
83
+ like a legitimate overlap. Check `ps -ef` for what is actually holding the name before assuming a
84
+ long-running instance of the job itself.
85
+
86
+ ## Step 5 — if runs overlap, look at what the job does before its real work
87
+ A job whose own runtime exceeds its schedule interval silently loses most of its slots to the
88
+ guard. Measure the phases, not the total: in the ODP case an instrumented run showed the S3
89
+ listing alone returning **138,472 objects and taking 32 seconds** before any PO was touched, with
90
+ each PO then costing roughly 40 seconds — comfortably past a 5-minute schedule. Prefixes the job
91
+ skips while processing are still listed, so accumulated `SENT/` and `OUTBOX/` objects inflate
92
+ every run.
93
+
94
+ ## Deploy-time trap: a PHP 8-only syntax kills the whole cron with no visible error
95
+ This tier runs **PHP 7.2**, so a **PHP 8 named argument** (`someFunction(items: $x)`) is a
96
+ **parse error**, not a runtime warning:
97
+
98
+ ```
99
+ PHP Parse error: syntax error, unexpected ':', expecting ',' or ')'
100
+ ```
101
+
102
+ The script then does nothing at all on every tick — and per Steps 2 and 4 you will see no output
103
+ and no `CronJobExecutions` row, i.e. it looks exactly like "the cron never fired". Typed
104
+ parameters and return types (`: void`, `: array`, `: bool`, `?array`, `?object`, `object $x`) are
105
+ PHP 7.0/7.1 and **are** safe here; they are already used in
106
+ `worker/crons/toga2/compass/backfill_items_isfulfillable.php` and
107
+ `update_item_fulfillments_in_netsuite.php`.
108
+
109
+ `1.0/standards/backend-php.md` already forbids named arguments in 1.0 — but a workspace
110
+ `CLAUDE.md` / `.github/copilot-instructions.md` that says *"use PHP 8 Named Parameters"* applies
111
+ to **2.0 only**. The 1.0 standard wins in `library/`, `worker/` and every other 1.0 app.
112
+
113
+ **Do not read a PHP 8 function call in existing code as proof the tier is PHP 8.**
114
+ `library/app/systemmonitor/500error.php` uses `str_contains()`; that file is not a version
115
+ signal, it is a latent bug.
116
+
117
+ ## Change history
118
+ - 2026-08-25 — Documented from a prod diagnostic session on the Compass ODP 850 importer: the
119
+ schedule JSON (not the file header) is authoritative, worker cron stdout is unreadable (base
120
+ block's `WORKER_ERROR` redirect commented out, role block writes a random filename),
121
+ `Common.CronJobExecutions.note` is the practical place for temporary tracing, a run skipped by
122
+ the `ps -ef` overlap guard writes **no row at all** (guard runs before the INSERT), the guard
123
+ matches any process holding the script path, and a PHP 8 named argument parse-errors the whole
124
+ cron silently on this PHP 7.2 tier. (bala)
@@ -132,6 +132,16 @@ can keep using `sendEmail($api, ...)`.
132
132
  predated the commit. That is a **deploy gap, not a code bug** — a redeploy fixes it, and no code
133
133
  change should be made. Same known-issue as
134
134
  [api2 environment-variable-drives-underscore-branch](../../api2/features/environment-variable-drives-underscore-branch.md).
135
+ - **⚠ An explicit `null` recipient is a 400, not a 500 (fixed 2026-08-25).** `$to`/`$cc`/`$bcc` on
136
+ `sendEmail()`/`send()`/`dispatch()` are typed `string|array|null` (a scripted-API caller can pass an
137
+ argument that is present-but-`null` — the `[]` default only applies when the arg is **omitted**).
138
+ Before the fix they were non-nullable `string|array`, so an explicit `null` `$to` threw
139
+ `Argument #3 ($to) must be of type array|string, null given` → **HTTP 500** — this was the platform's
140
+ **largest 5xx contributor** (prod `Logs.Issue` ref `2F`, 828 occurrences). `dispatch()` now normalizes
141
+ each recipient at the top (`(array)($x ?? [])`); after merging the template's stored
142
+ `EmailTemplateOutgoingEmailAddress` rows, if `$to` is **still empty** it throws `_Exception_Validation`
143
+ (→ **HTTP 400**) rather than a TypeError-500 or sending a recipientless email. A missing recipient is
144
+ client input error, not a server fault.
135
145
  - **`sendEmail()`'s signature is load-bearing for scripted APIs** — the Record Script engine
136
146
  (`api2/Component/Api/V2/V2.php`, ~line 3594) calls the method with `api` as a named
137
147
  argument, so the first param must stay `&$api`. Do not "clean it up" by removing it.
@@ -198,6 +208,13 @@ worker method) in-process instead.
198
208
  restored it — and that `sendEmail` has been variadic since 2024-12-24, so the spread was always
199
209
  correct and May was the regression. Documented the contract on `sendEmail()`'s docblock plus a marker
200
210
  comment at each of the four Quad call sites. (apeterson)
211
+ - 2026-08-25 — Made `$to`/`$cc`/`$bcc` nullable (`string|array|null`) on `sendEmail()`/`send()`/
212
+ `dispatch()` and normalized each to an array at the top of `dispatch()` (`(array)($x ?? [])`),
213
+ fixing the platform's **largest 5xx contributor** (prod `Logs.Issue` ref `2F`, 828 occurrences): a
214
+ scripted-API caller passing an explicit `null` recipient hit `Argument #3 ($to) must be of type
215
+ array|string, null given` → HTTP 500. After merging the template's stored recipient addresses, an
216
+ empty `$to` now throws `_Exception_Validation` (HTTP 400) instead of a TypeError-500 or a
217
+ recipientless send. Added the corresponding gotcha. (jcardinal)
201
218
  - 2026-08-18 — Fixed the Quad order-approval/rejection emails' production 500
202
219
  (`str_replace(): Argument #2 must be string`, EO-1): the four `Model/Quad/ApprovalDecision.php`
203
220
  + `Model/Quad/SalesOrder.php` call sites passed the template-vars map as a bare positional
@@ -6,8 +6,8 @@ project: _Underscore
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-08-17
10
- owners: ["dfranks", "jcardinal", "mhammontree", "apeterson", "kyalamarthi"]
9
+ updated: 2026-08-24
10
+ owners: ["dfranks", "jcardinal", "mhammontree", "apeterson", "kyalamarthi", "bala"]
11
11
  files:
12
12
  - _underscore/Database.php
13
13
  - _underscore/Model.php
@@ -177,6 +177,16 @@ here — they live in `Config/*.ini`.)
177
177
  - Related 1.0 analogue: the legacy `App_` worker has the same hazard writing to `Logs.API`
178
178
  (`db_logs`) — the laptop trap there is documented separately in the worker NetSuite bootstrap
179
179
  notes.
180
+ - **`_Database::register()` is LAZY — it stores connection config and never opens a socket.**
181
+ The connect happens on first use. So adding a static alias registration to an app's `_.php`
182
+ bootstrap is **safe across every environment**, even one whose config group points at a
183
+ localhost that lacks the schema: nothing fails until something actually queries that alias.
184
+ This is what makes "register the alias globally, use it in one cron" a two-line change rather
185
+ than a per-environment config exercise (worked example: `DB_VISION_LOGS` in
186
+ [the Worker Fleet Role-Assignment Monitor](../../worker2/features/worker-fleet-role-assignment-monitor.md)).
187
+ The corollary is the trap the rest of this doc describes: because registration is free and
188
+ silent, a **bad** registration also stays silent until the request that needs it blows up with
189
+ `Unknown database`.
180
190
 
181
191
  ## WITHDRAWN — the "alias-keyed `$_modelCache` cross-tenant leak" hypothesis (2026-08-17)
182
192
 
@@ -210,6 +220,11 @@ also hits `_modules/<module>/`** — is recorded in
210
220
 
211
221
  ## Change history
212
222
 
223
+ - 2026-08-24 — Recorded that **`_Database::register()` is lazy** (stores config, opens no
224
+ socket; the connect happens on first use), so a static alias registration in an app's `_.php`
225
+ is safe in every environment including a localhost config that lacks the schema — and that
226
+ this is also why a *bad* registration stays silent until first query. Verified while adding
227
+ `DB_VISION_LOGS` to worker2. (bala)
213
228
  - 2026-08-17 (later pass) — **WITHDRAWN, supersedes the entry below.** The alias-keyed
214
229
  `$_modelCache` cross-tenant-leak hypothesis is **not supported** and is no longer a live security
215
230
  concern. The six `c_` columns are **AIG's own**: `_modules/netsuite/2026-07-10a -
@@ -6,7 +6,7 @@ project: API
6
6
  client: shared
7
7
  type: architecture
8
8
  status: active
9
- updated: 2026-08-20
9
+ updated: 2026-08-25
10
10
  owners: [jcardinal, bala, mhammontree, dfranks]
11
11
  files:
12
12
  - api2/Controller/Index.php
@@ -114,14 +114,22 @@ One ~2,000-line `execute()` then `processRoutePairs()`:
114
114
  5. **Transaction logging** — every request logged (to client/core Logs DB, or as a JSONL
115
115
  line shipped by CloudWatch when `[api] log_filepath` is set).
116
116
 
117
- > **⚠ Auto-generated `Api.transactionId` collides under concurrency → 1062 → HTTP 500.**
117
+ > **Auto-generated `Api.transactionId` used to collide under concurrency → 1062 → HTTP 500 (FIXED 2026-08-25).**
118
118
  > Separate from the *client-supplied* `transactionId` uniqueness check (EV-5, above): the inbound
119
- > request-logger inserts its `Api` log row with a UNIQUE `transactionId` set to a
120
- > millisecond-precision timestamp (`Y-m-d H:i:s.v`). Concurrent nested writes generated within the
121
- > same millisecond collide on that UNIQUE key → MySQL **1062** → **HTTP 500**. This breaks ingestion
122
- > for high-volume senders (seen on the Compass/Veyer ASN feed, `sourceIp 34.232.23.158`) and is
123
- > platform-wide. Fix direction: make the logged `transactionId` unique-enough (uuid) or
124
- > retry-on-1062. Until fixed, a burst of concurrent posts can intermittently 500 with no app-level cause.
119
+ > request-logger inserted its `Api` log row with a UNIQUE `transactionId` (`Logs.Api.transactionId`,
120
+ > `varchar(255) UNIQUE`, in both Core `Logs` and `Logs_<Client>`) set to a millisecond-precision
121
+ > timestamp (`Y-m-d H:i:s.v`). Concurrent inserts in the same millisecond collided on that UNIQUE key
122
+ > → MySQL **1062** → the logger threw and turned a *successful* request into **HTTP 500 (EO-1)**. This
123
+ > broke ingestion for high-volume senders (seen on the Compass/Veyer ASN feed, `sourceIp 34.232.23.158`;
124
+ > confirmed in prod `Logs.Issue` reference `1Z`, 304 occurrences, `Duplicate entry
125
+ > '<ms-timestamp>' for key 'Api.transactionId'`, stack `/v2/users/me → _Model_Client_User::me() →
126
+ > internalApiRequest() → _Model->save()`).
127
+ > **Fix:** the two DB-insert log sites now set the log-row `transactionId` to a random uuid
128
+ > (`_String::generateUuid()`) so it cannot collide — `V2.php` `internalApiRequest()` logger (~L2370)
129
+ > and `Controller/Index.php` uncaught-`Throwable` error-recovery logger (~L456). The request's own base
130
+ > `transactionId` was already a uuid when auto-generated (`V2.php` ~L2130); only the bare-ms log-row id
131
+ > was at fault. The **sibling CloudWatch JSONL file-append logger** (`Controller/Index.php` ~L433) is
132
+ > **intentionally left bare-ms** — a JSONL append has no UNIQUE key and cannot collide.
125
133
 
126
134
  ## CRUD engine — `processRoutePairs()`
127
135
 
@@ -216,6 +224,34 @@ IMDSv2 token on the `curl` command line (visible in `ps`/`/proc/<pid>/cmdline`,
216
224
  back silently to IMDSv1 — fix both alongside the key. *(Location + remediation only; no key material
217
225
  is recorded anywhere.)*
218
226
 
227
+ ## EB "Degraded" health alerts vs. the ALB's own view (monitoring — read before triaging)
228
+
229
+ An Elastic Beanstalk enhanced-health transition — *"Environment health transitioned Ok→Degraded …
230
+ One or more TargetGroups … in a reduced health state: awseb-AWSEB-<id> - Degraded"* — **does NOT
231
+ mean the ALB marked a target unhealthy.** On `api-production-1` (AWS account `654654170868`), during
232
+ such an event the ALB's own view was **fully healthy throughout**: `HealthyHostCount=2`,
233
+ `UnHealthyHostCount=0`, ELB `5xx=0`, `TargetConnectionErrorCount=0`.
234
+
235
+ The two subsystems measure different things and legitimately disagree during any transient:
236
+
237
+ - **EB enhanced health** rates the target group from its **own per-instance request-latency/status
238
+ sampling on a point-in-time snapshot** — a short traffic burst (observed `RequestCount ~10x`
239
+ baseline, peaking ~1,936/min, with a `TargetResponseTime` max outlier ~7.7 s) is enough to flip it
240
+ to *Degraded*.
241
+ - **The target-group console** shows **current ALB health**, which recovers within ~45 s (3 checks).
242
+ So when you open the console after the alert, it "looks perfectly healthy" — because it is, now.
243
+
244
+ api2's EB target group health-checks **`/health`** (interval 15 s, timeout 5 s, unhealthy threshold
245
+ 5 ≈ **75 s to trip**, healthy threshold 3 ≈ **45 s to recover**). A transient burst is **expected
246
+ behavior, not an ALB/target-group defect** — do not chase a phantom target failure.
247
+
248
+ **The exception:** some *Degraded* events are instead genuine **"X% HTTP 5xx"** — those are real
249
+ code bugs, not transient bursts. The 2026-08 batch traced to the `transactionId`-1062 collision
250
+ (gotcha above), the platform-wide EmailTemplate null-recipient TypeError, and the Pcmaticb2b ticket
251
+ interceptor's plain-`Exception` rejections — all now returning 4xx/fixed. Distinguish the two by
252
+ reading the ALB 5xx metric and `Logs.Issue`: zero 5xx + a traffic spike = transient EB snapshot;
253
+ a sustained 5xx rate = a code bug to fix.
254
+
219
255
  ## Known issues / accepted risks
220
256
 
221
257
  Open items a maintainer should know before changing this tier. None are "bugs to fix right now" —
@@ -289,6 +325,18 @@ they are the known sharp edges. Do not re-discover these from scratch.
289
325
  single-caller branches in `V2.php` as unverified until exercised directly.
290
326
 
291
327
  ## Change history
328
+ - 2026-08-25 — **Resolved the auto-generated `Api.transactionId` 1062 collision** (was the 2026-07-28
329
+ gotcha): the two DB-insert log sites now uuid the log-row `transactionId` (`_String::generateUuid()`)
330
+ — `V2.php` `internalApiRequest()` logger (~L2370) and `Controller/Index.php` error-recovery logger
331
+ (~L456); the CloudWatch JSONL append (~L433) is intentionally left bare-ms (no UNIQUE key). Confirmed
332
+ the production symptom (`Logs.Issue` ref `1Z`, 304 occurrences, `/v2/users/me` → `_Model->save()`).
333
+ Also added a **monitoring section**: an EB enhanced-health *Ok→Degraded* "TargetGroups … reduced
334
+ health state" transition does **not** mean the ALB marked a target unhealthy — on `api-production-1`
335
+ (acct `654654170868`) the ALB view stayed fully healthy (HealthyHostCount=2, 5xx=0) while a short
336
+ ~10x traffic burst (~1,936/min, TargetResponseTime max ~7.7 s) flipped EB's point-in-time snapshot;
337
+ the target-group console recovers in ~45 s so it "looks healthy" after the fact. Recorded api2's
338
+ `/health` check timing (15 s / 5 s / trip 75 s / recover 45 s) and how to tell a transient burst
339
+ (0 ALB 5xx) from a genuine "X% HTTP 5xx" code bug. (jcardinal)
292
340
  - 2026-07-28 — Added Known issue #10: cross-client page-number paging is inherently O(page) (deep pages materialize every preceding row into the Cache cluster), with caller-facing keyset cursors identified as the durable fix but left unbuilt/unscoped, plus the `DEEPEN_CHUNK_MAX_RECORDS = 500` tuning tradeoff against `getFullModelData()` expansion and `CURL_TIMEOUT_SECONDS = 30`. Added a change-guidance bullet that single-caller branches in the untested `V2.php` monolith get zero incidental coverage (the keyset `LIMIT` syntax error shipped invisibly for a full cycle). (jcardinal)
293
341
  - 2026-07-28 — Documented the previously unrecorded `060_register_instance_to_shared_application_load_balancer` postdeploy hook pair in the Deployment section (non-prod self-registration into the same-named ALB target group; production skipped), and recorded a **committed IAM access key** in `ebs/register_instance_to_shared_application_load_balancer.php` as a security note + Known issue #9 — location, line range, commit subject, and remediation only (rotate, audit CloudTrail, move to the instance profile; history rewrite is a separate sign-off). Flagged api2's copy as the unhardened original vs. the new worker2 reference implementation. (jcardinal)
294
342
  - 2026-07-28 — Added a consolidated **Known issues / accepted risks** section (8 items), absorbing the previously free-floating deferred raw-exception-disclosure follow-up as item 1, so the tier's sharp edges (unrotated committed secrets, pre-execute phase still outside the main guard, local Logs DB name mismatch, permissive CORS, unpinned `_underscore` build clone, untested `V2.php` monolith, JWT rotation overlap window) are in one place instead of scattered. Recorded that `DB_CACHE` is resolved by name (`Databases.name = 'Cache'`), never by a hardcoded id, which differs per Core instance. (jcardinal)
@@ -17,7 +17,7 @@
17
17
  | [ClickUp Work Type Automation (Committed / Conditional / Stretch)](features/clickup-work-type-automation.md) | The ClickUp webhook handler (`_Worker_Clickup`) automatically maintains each task's **Work Type** custom field — `Committed`, `Conditional`, or `Stretch` — base | worker2/Worker/Clickup.php, worker2/Tests/Worker/ClickupWorkTypeTest.php |
18
18
  | [Compass VIP Support Importer (worker2)](features/compass-vip-support-importer.md) | A worker2 action that ingests Compass's quarterly VIP spreadsheet and assigns each VIP user's support technician by setting `Users.c_supportedByUserId` in `Clie | worker2/Worker/Client/Compass/VipSupport.php |
19
19
  | [Creating Worker Actions](features/creating-worker-actions.md) | How to add a new callable Worker action — a PHP class whose `public static` methods are invoked as background jobs (via webhook, cron, or `_Worker::runTask()`). | worker2/Worker/, worker2/Controller/Index.php, _underscore/Worker.php |
20
- | [Cross-account AWS access for worker2 crons (_Component_Aws_Workloads)](features/cross-account-aws-access.md) | `_Component_Aws_Workloads` is the **standard, and only sanctioned, way any new worker2 cron obtains AWS access** — for any account, any region, any AWS SDK clie | worker2/Component/Aws/Workloads/Workloads.php, worker2/Config/production.ini, worker2/Worker/Infrastructure/CloudWatch.php |
20
+ | [Cross-account AWS access for worker2 crons (_Component_Aws_Workloads)](features/cross-account-aws-access.md) | `_Component_Aws_Workloads` is the **standard, and only sanctioned, way any new worker2 cron obtains AWS access** — for any account, any region, any AWS SDK clie | worker2/Component/Aws/Workloads/Workloads.php, worker2/Config/production.ini, worker2/Worker/Infrastructure/CloudWatch.php, worker2/Worker/Monitor/Fleet.php, library/app/worker.php, library/app/cloud.php |
21
21
  | [Elastic Beanstalk health monitor → OneUptime (ElasticBeanstalkHealth)](features/elastic-beanstalk-health-monitor.md) | `_Worker_Infrastructure_CloudWatch::ElasticBeanstalkHealth()` is a worker2 cron that reads each Elastic Beanstalk (EB) environment's **enhanced-health** status | worker2/Worker/Infrastructure/CloudWatch.php |
22
22
  | [Elite Freshservice Sync (worker2)](features/elite-freshservice-sync.md) | `_Worker_Elite` processes Freshservice webhook events and syncs them into TOGA 2. | worker2/Worker/Elite.php, worker2/Config/dev-kmaramreddy-laptop.ini |
23
23
  | [Error Escalation Cron (Errors::Escalate → ClickUp / email)](features/error-escalation-cron.md) | `_Worker_Infrastructure_Errors::Escalate` (renamed from `SyncWithClickup`) is the sole owner of **escalation, de-escalation, ClickUp ticketing, reminders, busin | worker2/Worker/Infrastructure/Errors.php, worker2/Worker/Notification/Email.php, worker2/Worker/Notification/EmailTemplate.php, worker2/Worker/Client/True.php, worker2/Worker/Clickup/ErrorTask.php, worker2/Worker/Clickup.php, worker2/Controller/Index.php, worker2/Config/production.ini, _underscore/Model/Core/Logs/Issue.php, dbchanges2/Core/2026-07-30a - Error escalation cron job.sql, dbchanges2/Logs/2026-08-03a - Issue clickupPriority.sql, dbchanges2/Core/2026-08-04a - Error neglect digest cron job.sql |
@@ -47,6 +47,7 @@
47
47
  | [Centralized Tracking-Status Refresh (worker2 Sync cron, all clients)](features/tracking-status-refresh.md) | A single platform-wide cron keeps `TrackingNumbers.status` current for **every provisioned client** until each shipment reaches a terminal state, by polling Fed | worker2/Worker/Sync/TrackingNumbers.php, _underscore/Component/Library/Carriers/Fedex/Fedex.php, _underscore/Component/Library/Carriers/Usps/Usps.php, _underscore/Component/Library/Carriers/Ups/Ups.php, _underscore/Model/Client/TrackingNumber.php, dbchanges2/Client/2026-07-30 - TrackingNumbers_RefreshColumns.sql, dbchanges2/Client/2026-07-30 - Acl_TrackingStatus_Grant.sql, dbchanges2/Client_Tdsynnex/2026-07-30 - Apis_TdsynnexKey.sql, dbchanges2/Core/2026-07-30 - TrackingRefresh_Cron.sql |
48
48
  | [VAPI Webhook Handler (worker2 — AI-BDR end-of-call processing)](features/vapi-webhook-handler.md) | `_Worker_Vapi` ([worker2/Worker/Vapi.php](worker2/Worker/Vapi.php)) is the **PHP side of the AI-BDR call loop** — the webhook that receives VAPI's end-of-call r | worker2/Worker/Vapi.php, worker2/Worker/Ai/Bdr/Vapi.php, worker2/Controller/Index.php |
49
49
  | [WJE Freshservice Sync (worker2)](features/wje-freshservice-sync.md) | WJE ("WJE IT", helpdesk `wje.freshservice.com`) is a **Freshservice**-based help-desk client whose tickets, contacts, assets, groups, categories, and canned res | worker2/Worker/Wje.php, _underscore/Component/Api/Wje/Wje.php, _underscore/Model/Wje/Ticket.php, _underscore/Model/Wje/TicketNote.php, _underscore/Model/Wje/Contact.php, _underscore/Model/Wje/Unit.php, _underscore/Model/Wje/TicketTeam.php, _underscore/Model/Wje/TicketCategory.php, _underscore/Model/Wje/AssetType.php, _underscore/Model/Wje/PredefinedReply.php, library/app/api/wje.php, worker/crons/toga2/wje/import_supporting_records.php, worker/crons/toga2/wje/sync_togasupply_wje.php, worker/crons/notifications/reports/wje/wje_common.php, library/app/systemmonitor/wje.php, dbchanges2/Client_Wje/2024-10-04 - WjeOnboarding.sql |
50
+ | [1.0 Worker Fleet Role-Assignment Monitor (Monitor/Fleet/RoleAssignment)](features/worker-fleet-role-assignment-monitor.md) | `_Worker_Monitor_Fleet::RoleAssignment()` is a **2.0 worker2 cron that watches the 1.0 `worker` fleet from the outside**. | worker2/Worker/Monitor/Fleet.php, worker2/_.php, dbchanges2/Core/2026-08-24 - Worker Role Assignment Monitor.sql, worker/ebs/cron.worker.php, worker/crons/worker/worker_heartbeat.php, library/app/worker.php |
50
51
  | [PHP Runtime Upgrade on Elastic Beanstalk (worker2 8.3 → 8.5 + PhpSpreadsheet 1.x → 3.x)](workflows/php-runtime-upgrade-dependency-audit.md) | The procedure used to move worker2 from **PHP 8.3 to PHP 8.5** on Elastic Beanstalk, and the dependency work that had to land first. | worker2/composer.json, worker2/composer.lock, worker2/Worker/Team/Sprint.php, worker2/Worker/Client/TowFoundation/ProcessReceipts.php, worker2/Worker/Forecast/Import.php |
51
52
  | [Running worker2 locally against real NetSuite (and what EB does instead)](workflows/running-worker2-locally.md) | How to boot **worker2** on a developer machine and run a real worker action against the **live NetSuite** account and a **local** `Forecast` database. | worker2/index.php, worker2/composer.json, worker2/.ebextensions/php_include_underscore.config, worker2/.ebextensions/git.json, worker2/.ebextensions/git.php, worker2/Config/production.ini, _underscore/Component/Api/Netsuite/Netsuite.php, dbchanges2/Logs_Client/2026-04-08_BLANK_CLIENT_LOGS_DATABASE.SQL |
52
53
  | [Ticket → ClickUp Pseudocode Planning (Talos-grounded)](workflows/ticket-to-pseudocode-planning.md) | A repeatable procedure for turning a ClickUp ticket into a reviewed, formatted implementation plan posted back to the ticket's `📝 Pseudocode` custom field. | test/@dave/clickup_md2delta.js |
@@ -6,18 +6,22 @@ project: Worker
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-08-11
10
- owners: [jcardinal]
9
+ updated: 2026-08-24
10
+ owners: [jcardinal, bala]
11
11
  files:
12
12
  - worker2/Component/Aws/Workloads/Workloads.php
13
13
  - worker2/Config/production.ini
14
14
  - worker2/Worker/Infrastructure/CloudWatch.php
15
+ - worker2/Worker/Monitor/Fleet.php
16
+ - library/app/worker.php
17
+ - library/app/cloud.php
15
18
  related:
16
19
  - ../../_underscore/features/component-model-namespace-registration.md
17
20
  - ../../_underscore/features/config-group-access.md
18
21
  - ../../_underscore/features/cloud-s3-helpers.md
19
22
  - ./oneuptime-worker2-monitoring.md
20
23
  - ./elastic-beanstalk-health-monitor.md
24
+ - ./worker-fleet-role-assignment-monitor.md
21
25
  ---
22
26
 
23
27
  ## Summary
@@ -79,7 +83,8 @@ region is only a runtime argument to `client()`, never a reason for a second rol
79
83
  - **654654170868** — production hub. Home region `us-east-1`; also used in `us-west-2` and
80
84
  `eu-west-1`. worker2 runs here on Elastic Beanstalk; its instance profile role
81
85
  `ElasticBeanstalk-EC2-Instance-Profile` is the caller identity.
82
- - **502614707982** — legacy account, `us-west-2`.
86
+ - **502614707982** — legacy account, `us-west-2`. **This is where the 1.0 `worker` fleet
87
+ actually runs** — see the account-confusion gotcha below.
83
88
  - **975050298201** — non-production account, `us-east-1`.
84
89
 
85
90
  Each `WorkloadsRuntime` role has broad service permissions (`elasticbeanstalk:*`, `ec2:*`,
@@ -144,8 +149,33 @@ code yet. New code must not adopt the static key.
144
149
  **Never record the access-key values in knowledge.** They exist in `production.ini`; that is
145
150
  their only home.
146
151
 
152
+ ### Which account the 1.0 worker fleet lives in (easy and expensive to get wrong)
153
+
154
+ The 1.0 `worker` EC2 instances are in **502614707982**, `us-west-2` (EB environment
155
+ `agilant-worker`, confirmed in the EC2 console).
156
+
157
+ **`App_Worker::AWS_WORKER_QUEUE_URL` in `library/app/worker.php` points at 654654170868 and
158
+ misleads** — that account holds only the **SQS queue the 1.0 workers read from**, not the
159
+ instances. Anything that needs to *see* the worker boxes (EB, EC2, health) must target
160
+ **502614707982**.
161
+
162
+ worker 1.0 authenticates there with a **long-lived hardcoded IAM access key in
163
+ `library/app/cloud.php`** — a committed credential that should be rotated and moved to a role
164
+ (value deliberately not recorded here; see [Legacy static key](#legacy-static-key--being-retired)).
165
+ worker2 does **not** use it: it assumes `WorkloadsRuntime` cross-account from its EB instance
166
+ role, which is exactly the migration this component exists to enable. First worker2 consumer in
167
+ that account is the
168
+ [Worker Fleet Role-Assignment Monitor](./worker-fleet-role-assignment-monitor.md).
169
+
170
+ **Scope EB reads by `EnvironmentName`, never account-wide.** 502614707982 also runs
171
+ `agilant-worker-alpha` and `agilant-worker-beta` as **separate EB environments that have no role
172
+ registry at all**. An account-level EC2 scan would sweep them in and misreport them.
173
+
147
174
  ## Gotchas / known issues
148
175
 
176
+ - **Do not assume the 1.0 worker fleet is in the 868 production hub.** `AWS_WORKER_QUEUE_URL`
177
+ (654654170868) is the SQS queue only; the instances are in the legacy account 502614707982,
178
+ `us-west-2`. See [the section above](#which-account-the-10-worker-fleet-lives-in-easy-and-expensive-to-get-wrong).
149
179
  - Do not construct `StsClient` or pass explicit credentials by hand in new crons — the whole
150
180
  point is that the instance role is the caller and the component owns the assume-role step.
151
181
  - The instance-role-as-caller only works **on the EB tier**. Off-tier (e.g. a local CLI run)
@@ -161,6 +191,15 @@ their only home.
161
191
  see [creating-worker-actions.md](./creating-worker-actions.md).
162
192
 
163
193
  ## Change history
194
+ - 2026-08-24 — Pinned down **which account the 1.0 `worker` fleet actually runs in**:
195
+ **502614707982** (`us-west-2`, EB env `agilant-worker`), **not** the 654654170868 production hub
196
+ that `App_Worker::AWS_WORKER_QUEUE_URL` implies — that account holds only the SQS queue the 1.0
197
+ workers read from. Recorded that worker 1.0 reaches it with the hardcoded static key in
198
+ `library/app/cloud.php` (rotate + move to a role) while worker2 assumes `WorkloadsRuntime` from
199
+ its instance role, and that EB reads there must be scoped by `EnvironmentName` because
200
+ `agilant-worker-alpha`/`-beta` are separate registry-less environments in the same account.
201
+ Found while building the
202
+ [Worker Fleet Role-Assignment Monitor](./worker-fleet-role-assignment-monitor.md). (bala)
164
203
  - 2026-08-11 — Noted that `ElasticBeanstalkHealth`'s `$environments` must be typed `object|array`
165
204
  (not `array`): the nested JSON object arrives as `stdClass` because the dispatcher casts only
166
205
  the top level of `CronJobs.parameters`, so a strict `array` throws a `TypeError` at dispatch.