toga-ai 1.0.645 → 1.0.647
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/knowledge/1.0/apps/library/features/cron-execution-monitoring.md +25 -3
- package/knowledge/1.0/apps/worker/INDEX.md +2 -1
- package/knowledge/1.0/apps/worker/architecture.md +35 -2
- package/knowledge/1.0/apps/worker/features/oneuptime-worker-uptime-monitoring.md +40 -3
- package/knowledge/1.0/apps/worker/workflows/tracing-a-worker-cron-run-in-production.md +124 -0
- package/knowledge/2.0/apps/_underscore/features/email-template-sending.md +17 -0
- package/knowledge/2.0/apps/_underscore/features/per-client-database-connections.md +17 -2
- package/knowledge/2.0/apps/api2/architecture.md +56 -8
- package/knowledge/2.0/apps/worker2/INDEX.md +2 -1
- package/knowledge/2.0/apps/worker2/features/cross-account-aws-access.md +42 -3
- package/knowledge/2.0/apps/worker2/features/oneuptime-worker2-monitoring.md +39 -0
- package/knowledge/2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md +175 -0
- package/knowledge/2.0/standards/backend-php.md +18 -2
- package/knowledge/INDEX.md +2 -2
- package/knowledge/clients/compass-usa/INDEX.md +1 -0
- package/knowledge/clients/compass-usa/features/odp-edi-850-item-resolution.md +179 -0
- package/knowledge/clients/compass-usa/features/odp-edi-855-acknowledgement-and-overquantity-guard.md +29 -2
- package/knowledge/clients/compass-usa/profile.md +10 -1
- package/knowledge/clients/compass-usa/workflows/odp-order-pipeline-to-netsuite.md +39 -2
- package/knowledge/clients/pcmaticb2b/INDEX.md +1 -1
- package/knowledge/clients/pcmaticb2b/features/startech-ticket-sync.md +5 -2
- package/knowledge/clients/pcmaticb2b/profile.md +4 -2
- package/package.json +1 -1
|
@@ -6,12 +6,13 @@ project: Library
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-
|
|
10
|
-
owners: [dfranks]
|
|
9
|
+
updated: 2026-08-25
|
|
10
|
+
owners: [dfranks, bala]
|
|
11
11
|
files:
|
|
12
12
|
- library/app/framework.php
|
|
13
13
|
related:
|
|
14
14
|
- ../../worker/architecture.md
|
|
15
|
+
- ../../worker/workflows/tracing-a-worker-cron-run-in-production.md
|
|
15
16
|
---
|
|
16
17
|
|
|
17
18
|
## Summary
|
|
@@ -38,13 +39,28 @@ row in `db_log` and of the per-job Sentry check-in monitors.
|
|
|
38
39
|
cast silently discards the sub-second precision even when the schema column is `decimal(10,3)`.
|
|
39
40
|
- **Schema:** the table is provisioned by the dbchanges migration
|
|
40
41
|
`dbchanges/Common/DF/2026-5-7 Cron Checkin.sql` (table `CronJobExecutions`,
|
|
41
|
-
`executionTimeSeconds decimal(10,3)
|
|
42
|
+
`executionTimeSeconds decimal(10,3)`). **Correction (2026-08-25): prod `Common.CronJobExecutions`
|
|
43
|
+
does have a `note` TEXT column** — verified by writing to it and reading it back. `App_Framework`
|
|
44
|
+
never populates it, which makes it the practical place to park **temporary** step tracing for a
|
|
45
|
+
cron you are diagnosing (worker crons have no readable stdout — see
|
|
46
|
+
[Tracing a worker cron run in production](../../worker/workflows/tracing-a-worker-cron-run-in-production.md)).
|
|
47
|
+
Remove the tracing when you are done, and never write payloads or credentials into it.
|
|
48
|
+
- **Don't confuse the two `note`s.** `cronInitialization()` writes `note = 'Started execution'` to
|
|
49
|
+
the **`Log`** table on `db_log` (the older `CRON`/`recordType` row), not to `CronJobExecutions`.
|
|
42
50
|
- **Design decision:** 1.0 records executions by writing **directly to the shared DB**
|
|
43
51
|
(`db_common`), *not* by POSTing from 1.0 to a 2.0 API endpoint. Per Jeff Cardinal, 1.0 code
|
|
44
52
|
must not POST to 2.0 code; the direct shared-DB write is the sanctioned path (TRUE-78182).
|
|
45
53
|
|
|
46
54
|
## Gotchas
|
|
47
55
|
|
|
56
|
+
- **⚠ A run skipped by the overlap guard leaves NO ROW — absence is ambiguous.**
|
|
57
|
+
`cronInitialization()` calls `exitIfProcessRunning()` **before** the `CronJobExecutions` INSERT,
|
|
58
|
+
and the `cronFinished(true)` that guard then calls does nothing because `cronLogId` was never
|
|
59
|
+
set. So a blocked run is indistinguishable from a run that never launched: **a missing row for
|
|
60
|
+
an expected slot means "skipped", not "hung"**, and a hung run looks different (it has
|
|
61
|
+
`dtCheckIn` with no `dtCheckOut`). `isProcessRunning()` matches **any** `ps -ef` line containing
|
|
62
|
+
the script path (it only excludes `/bin/sh` and its own pid), so a stray `tail -f` or editor on
|
|
63
|
+
that path blocks the cron indefinitely.
|
|
48
64
|
- **`db_log` has no `CronJobExecutions` table.** A `cronFinished()` UPDATE that runs against the
|
|
49
65
|
`db_log` connection throws on **every** 1.0 cron job. Both the INSERT and the UPDATE must
|
|
50
66
|
target `db_common`.
|
|
@@ -60,6 +76,12 @@ row in `db_log` and of the per-job Sentry check-in monitors.
|
|
|
60
76
|
logic. (Fixed 2026-07-06: consolidated back to one `db_common` INSERT/UPDATE pair.)
|
|
61
77
|
|
|
62
78
|
## Change history
|
|
79
|
+
- 2026-08-25 — Corrected the schema note (prod `CronJobExecutions` **does** carry a `note` TEXT
|
|
80
|
+
column, unused by `App_Framework`, usable for temporary cron tracing) and separated it from the
|
|
81
|
+
`note = 'Started execution'` row `cronInitialization()` writes to `Log` on `db_log`. Recorded
|
|
82
|
+
that the overlap guard runs **before** the INSERT, so a skipped run writes no row at all and
|
|
83
|
+
cannot be told apart from "never launched" — and that `isProcessRunning()`'s `ps -ef` substring
|
|
84
|
+
match lets any process holding the script path block the cron. No code change. (bala)
|
|
63
85
|
- 2026-07-06 — Fixed a merge (`origin/_production`) that reintroduced duplicate
|
|
64
86
|
`CronJobExecutions` writes and pointed one `cronFinished()` UPDATE at `db_log` (where the
|
|
65
87
|
table doesn't exist), which would have thrown on every 1.0 cron. Consolidated to one
|
|
@@ -11,9 +11,10 @@
|
|
|
11
11
|
| [NetSuite Sales Order Sales Rep Sourcing (Staples & ODP EDI orders)](features/netsuite-sales-order-sales-rep-sourcing.md) | How the **sales rep** on a NetSuite Sales Order is determined for the two 1.0 `worker` EDI order-creation integrations (Staples cXML and Compass/ODP EDI). | worker/crons/toga2/compass/workflow/5_create_netsuite_sales_orders_from_office_depot_purchase_orders.php, worker/crons/sync/staples/sync_staples_cxml.php, test/@Mark/NetSuite/TRUE_80451_customer_salesrep_diag.php |
|
|
12
12
|
| [NetSuite → TOGa Supply Per-Client Sync (thin wrappers)](features/netsuite-togasupply-per-client-sync.md) | Syncs NetSuite transactions (sales orders, purchase orders, invoices, item receipts, item fulfillments, inventory adjustments) into each TOGa Supply (2.0) clien | worker/crons/toga2/netsuite/common_sync_togasupply.php, worker/crons/toga2/netsuite/sync_togasupply_canon.php, worker/crons/toga2/netsuite/sync_togasupply_elite.php, worker/schedules/cron.worker.sync.json, dbchanges2/_modules/netsuite/2026-04-01 - Parameters.sql, library/app/api/toga2.php, library/app/api/netsuite/rest.php, library/app/framework.php, library/app/systemmonitor/netsuiteintegration.php, test/@srija/Elite Testing/Service Requests/test_sync_togasupply_elite_section.php, test/@srija/Elite Testing/Service Requests/test_diagnose_togasupply_elite.php |
|
|
13
13
|
| [OneUptime Server monitor + disk/memory hygiene on the 1.0 worker EB host](features/oneuptime-server-monitor-host-hygiene.md) | The 1.0 `agilant-worker` EB environment runs on the **legacy Amazon Linux 1 PHP 7.2 platform** (Apache httpd/prefork, s3fs mounts, cron) and repeatedly went dow | worker/.ebextensions/040_disk_memory_hygiene.config, worker/.ebextensions/045_oneuptime_agent.config, worker/ebs/cron.worker.php, worker/ebs/mount-s3fs-folders.php, worker/ebs/apache_settings.php, worker/ebs/setup_phpini.php |
|
|
14
|
-
| [OneUptime external uptime monitoring for 1.0 workers](features/oneuptime-worker-uptime-monitoring.md) | Every 1.0 worker box self-reports its liveness to an external OneUptime monitor once per minute by curl-POSTing to a per-worker "Incoming Request" heartbeat URL | library/app/worker.php, worker/crons/worker/worker_heartbeat.php |
|
|
14
|
+
| [OneUptime external uptime monitoring for 1.0 workers](features/oneuptime-worker-uptime-monitoring.md) | Every 1.0 worker box self-reports its liveness to an external OneUptime monitor once per minute by curl-POSTing to a per-worker "Incoming Request" heartbeat URL | library/app/worker.php, worker/crons/worker/worker_heartbeat.php, worker/ebs/cron.worker.php, worker/.ebextensions/045_oneuptime_agent.config |
|
|
15
15
|
| [Prudential: Send Shipments for the Day report (daily cron)](features/send-shipments-for-the-day.md) | Daily cron (9:00 PM) that emails Prudential and Dell stakeholders an Excel report of all devices shipped that day, including tracking number, serial number, emp | worker/crons/notifications/reports/send_shipments_for_the_day.php |
|
|
16
16
|
| [Staples cXML Order Import (SFTP → NetSuite)](features/staples-cxml-order-import.md) | `sync_staples_cxml.php` is an **hourly** cron (runs at **:45**) that imports Staples cXML purchase orders from SFTP into NetSuite as Sales Orders, then writes a | worker/crons/sync/staples/sync_staples_cxml.php |
|
|
17
17
|
| [Diagnosing frozen 1.0 worker cron check-ins (Sentry "missed" flood)](workflows/diagnosing-frozen-cron-checkins.md) | When 1.0 worker cron timestamps freeze and Sentry project `worker1` fills with **`missed`** check-ins, the intuitive diagnosis — a wedged `App_Framework::isProc | worker/.ebextensions/cron.config, library/app/worker.php |
|
|
18
18
|
| [isFulfillable Multi-Client Backfill (all togasupply clients)](workflows/isfulfillable-multi-client-backfill.md) | One-time backfill that catches up `Items.isFulfillable` on **existing** items across **all 17 togasupply clients** (AIG, Broward Sheriff, Canon, Endeavor Health | worker/crons/toga2/netsuite/backfill_isfulfillable_all_clients.php, library/app/api/toga2.php |
|
|
19
19
|
| [Onboarding a Client to the NetSuite TOGa Supply Sync](workflows/onboarding-client-to-netsuite-togasupply-sync.md) | How to add a new TOGa 2 client to the per-client NetSuite → TOGa Supply importer (`worker/crons/toga2/netsuite/`). | worker/crons/toga2/netsuite/sync_togasupply.php, worker/crons/toga2/netsuite/common_sync_togasupply.php, worker/crons/toga2/netsuite/sync_togasupply_elite.php, worker/schedules/cron.worker.sync.json, dbchanges2/_modules/netsuite/2026-04-01 - Parameters.sql, dbchanges2/_modules/netsuite/2026-08-05 - CLEAN NETSUITE CLINET.SQL |
|
|
20
|
+
| [Tracing a 1.0 worker cron run in production (no stdout, silent skips)](workflows/tracing-a-worker-cron-run-in-production.md) | How to answer *"did this cron actually run, and what did it do?"* on the 1.0 `worker` tier, where **there is no usable stdout** and **a skipped run leaves no tr | worker/ebs/cron.worker.php, worker/schedules/cron.worker.sync.json, library/app/framework.php |
|
|
@@ -6,8 +6,8 @@ project: Worker
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: architecture
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-
|
|
10
|
-
owners: [jcardinal, sking]
|
|
9
|
+
updated: 2026-08-24
|
|
10
|
+
owners: [jcardinal, sking, bala]
|
|
11
11
|
files:
|
|
12
12
|
- worker/index.php
|
|
13
13
|
- worker/_/app/framework.php
|
|
@@ -20,6 +20,8 @@ files:
|
|
|
20
20
|
related:
|
|
21
21
|
- ../library/architecture.md
|
|
22
22
|
- ../dbchanges/workflows/authoring-and-shipping-sql-files.md
|
|
23
|
+
- ./features/oneuptime-worker-uptime-monitoring.md
|
|
24
|
+
- ../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md
|
|
23
25
|
---
|
|
24
26
|
|
|
25
27
|
## Summary
|
|
@@ -44,6 +46,7 @@ which **self-elects a distinct role** (`notification`, `database`, `infrastructu
|
|
|
44
46
|
under `crons/`; the schedule (cron registration) is the source of truth for what runs — a script
|
|
45
47
|
that isn't scheduled never executes. Client integrations live under `crons/toga2/<client>/`.
|
|
46
48
|
Use prepared statements for all SQL; never interpolate input.
|
|
49
|
+
A worker that fails deploy-time role election runs NO crons while EB still reads healthy.
|
|
47
50
|
|
|
48
51
|
## How a job becomes a cron (the dispatch pipeline)
|
|
49
52
|
|
|
@@ -108,6 +111,32 @@ is the most important and least obvious part of the architecture.
|
|
|
108
111
|
> Net effect: roles are a **claim-the-first-free-slot pool**, and a dead worker's role is
|
|
109
112
|
> reclaimed by its replacement within minutes. There is no static instance→role mapping to edit.
|
|
110
113
|
|
|
114
|
+
### The failure mode: NO NAME = NO CRONTAB (and it is silent)
|
|
115
|
+
|
|
116
|
+
The election is the **single point of failure for an entire instance**, and the architecture gives
|
|
117
|
+
it no second chance:
|
|
118
|
+
|
|
119
|
+
- The claim runs **only at DEPLOY time** — `ebs/cron.worker.php` is written as the EB `appdeploy`
|
|
120
|
+
enact hook. **It never retries.** An instance that fails to claim stays nameless until the next
|
|
121
|
+
deploy or until something terminates it.
|
|
122
|
+
- The claim writes `/etc/worker-role` and **only then** appends `cron.worker.<role>.json` to the
|
|
123
|
+
crontab. So a nameless instance has **no role crontab at all** — it runs *nothing*, not even
|
|
124
|
+
`worker_heartbeat.php`, while **EB reports the environment healthy**. The box is up; it just
|
|
125
|
+
does no work.
|
|
126
|
+
- **The claim's DB connect is `@mysqli_connect(...)` — error-suppressed.** A boot-time connect
|
|
127
|
+
failure to `Vision_Log` therefore leaves the instance nameless **with no log line anywhere**.
|
|
128
|
+
This is the prime suspect for the 2026-08-24 incident (2 of 7 workers roleless for ~3 hours),
|
|
129
|
+
and it violates the no-`@`-suppression rule in `rules/toga/common/coding-style.md`. Removing the
|
|
130
|
+
`@` and logging the failure is the recommended fix.
|
|
131
|
+
- `Workers` **only ever holds instances that SUCCEEDED**, so a failed claim leaves no row —
|
|
132
|
+
**absence is the only evidence**, and detecting it requires the AWS instance list.
|
|
133
|
+
|
|
134
|
+
Because every OneUptime monitor on this tier is keyed by **role** rather than by instance, none of
|
|
135
|
+
them can see this (see
|
|
136
|
+
[the blind spot](./features/oneuptime-worker-uptime-monitoring.md)). External coverage now comes
|
|
137
|
+
from the 2.0
|
|
138
|
+
[Worker Fleet Role-Assignment Monitor](../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md).
|
|
139
|
+
|
|
111
140
|
## Anatomy of a cron script
|
|
112
141
|
|
|
113
142
|
Every script is self-contained and follows this boilerplate:
|
|
@@ -208,6 +237,10 @@ web face (`mvc/` GET routes for login/logout/404); the tier's real work is the c
|
|
|
208
237
|
|
|
209
238
|
## Conventions & gotchas
|
|
210
239
|
|
|
240
|
+
- **A worker that fails role election is silently dead, not degraded.** No name → no crontab →
|
|
241
|
+
the instance runs nothing while EB reads healthy, and the `@`-suppressed `mysqli_connect` in
|
|
242
|
+
`ebs/cron.worker.php` logs nothing. Never treat "EB is green" as evidence the fleet is working;
|
|
243
|
+
check `Vision_Log.Workers` against the actual EB instance list.
|
|
211
244
|
- **`dbchanges` auto-apply is branch-gated, and `_production` is NOT auto-applied.**
|
|
212
245
|
`crons/infrastructure/execute_dbchanges.php` runs every 2 minutes but only for branches matching
|
|
213
246
|
`_%`, deriving the env by stripping the leading `_` and requiring `config.<env>.ini` in the
|
|
@@ -6,12 +6,17 @@ project: Worker
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-
|
|
10
|
-
owners: ["jcardinal"]
|
|
9
|
+
updated: 2026-08-24
|
|
10
|
+
owners: ["jcardinal", "bala"]
|
|
11
11
|
files:
|
|
12
12
|
- library/app/worker.php
|
|
13
13
|
- worker/crons/worker/worker_heartbeat.php
|
|
14
|
-
|
|
14
|
+
- worker/ebs/cron.worker.php
|
|
15
|
+
- worker/.ebextensions/045_oneuptime_agent.config
|
|
16
|
+
related:
|
|
17
|
+
- ./oneuptime-server-monitor-host-hygiene.md
|
|
18
|
+
- ../architecture.md
|
|
19
|
+
- ../../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md
|
|
15
20
|
---
|
|
16
21
|
|
|
17
22
|
## Summary
|
|
@@ -57,6 +62,31 @@ Database, Infrastructure, TOGa, TOGa Desk, and Catalog, then importing each into
|
|
|
57
62
|
manually. To add a new worker's monitor, clone an existing monitor export the same way and
|
|
58
63
|
register the resulting heartbeat URL in `App_Worker::$oneUptimeEndpoints`.
|
|
59
64
|
|
|
65
|
+
## THE BLIND SPOT — every monitor here is keyed by ROLE, so a ROLELESS instance is invisible
|
|
66
|
+
|
|
67
|
+
Load-bearing limitation, learned the hard way (2026-08-24: **2 of the 7 production workers ran
|
|
68
|
+
roleless for ~3 hours undetected**). This layer, and the Server monitors in
|
|
69
|
+
[host hygiene](./oneuptime-server-monitor-host-hygiene.md), are keyed by `workerName` — **14
|
|
70
|
+
monitors, none keyed by instance**:
|
|
71
|
+
|
|
72
|
+
- `worker_heartbeat.php` only pings `App_Worker::$oneUptimeEndpoints[$row['workerName']]` for
|
|
73
|
+
the row matching **its own** `instanceId`. No row → no ping target → it pushes nothing.
|
|
74
|
+
- The `045_oneuptime_agent.config` Server agent reads `/etc/worker-role` and **exits clean when
|
|
75
|
+
it is empty.**
|
|
76
|
+
|
|
77
|
+
So an instance that never claimed a role reports to nothing — and because every role is still
|
|
78
|
+
owned by *someone*, **no role-keyed monitor is missing a ping either.** The fleet reads 100%
|
|
79
|
+
healthy while a box does zero work (role assignment happens **only at deploy time**, and no role
|
|
80
|
+
means the role's `cron.worker.<role>.json` is never appended to the crontab, so the instance runs
|
|
81
|
+
**nothing**).
|
|
82
|
+
|
|
83
|
+
`Vision_Log.Workers` only ever holds instances that **succeeded** in claiming a role, so the
|
|
84
|
+
failure leaves no row and no log line — **absence is the only evidence**, and detecting an
|
|
85
|
+
absence requires the AWS instance list, which nothing on this tier consults. That gap is what
|
|
86
|
+
the 2.0
|
|
87
|
+
[Worker Fleet Role-Assignment Monitor](../../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md)
|
|
88
|
+
now closes from outside the fleet. **Do not try to close it with another role-keyed monitor.**
|
|
89
|
+
|
|
60
90
|
## Gotchas
|
|
61
91
|
|
|
62
92
|
- Do **not** hardcode the heartbeat URLs/UUIDs anywhere but `App_Worker::$oneUptimeEndpoints`
|
|
@@ -67,6 +97,13 @@ register the resulting heartbeat URL in `App_Worker::$oneUptimeEndpoints`.
|
|
|
67
97
|
pending cleanup — the live source of truth is the `App_Worker::$oneUptimeEndpoints` map.
|
|
68
98
|
|
|
69
99
|
## Change history
|
|
100
|
+
- 2026-08-24 — Documented **the blind spot**: all 14 monitors on this tier are keyed by
|
|
101
|
+
`workerName`, so an instance that claimed **no** role pings nothing and leaves every role-keyed
|
|
102
|
+
monitor still green — which is how 2 of 7 production workers ran roleless for ~3 hours
|
|
103
|
+
undetected. Recorded that `Vision_Log.Workers` only holds successful claims (absence is the
|
|
104
|
+
only evidence) and that the gap is now covered from outside the fleet by the 2.0
|
|
105
|
+
[Worker Fleet Role-Assignment Monitor](../../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md).
|
|
106
|
+
(bala)
|
|
70
107
|
- 2026-07-13 — Added external OneUptime heartbeat push for all 1.0 workers on top of the
|
|
71
108
|
existing internal DB-heartbeat mechanism; endpoints keyed by workerName in
|
|
72
109
|
`App_Worker::$oneUptimeEndpoints`. (jcardinal)
|
|
@@ -0,0 +1,124 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Tracing a 1.0 worker cron run in production (no stdout, silent skips)
|
|
3
|
+
framework: "1.0"
|
|
4
|
+
repo: worker
|
|
5
|
+
project: Worker
|
|
6
|
+
client: shared
|
|
7
|
+
type: workflow
|
|
8
|
+
status: active
|
|
9
|
+
updated: 2026-08-25
|
|
10
|
+
owners: ["bala"]
|
|
11
|
+
files:
|
|
12
|
+
- worker/ebs/cron.worker.php
|
|
13
|
+
- worker/schedules/cron.worker.sync.json
|
|
14
|
+
- library/app/framework.php
|
|
15
|
+
related:
|
|
16
|
+
- ../architecture.md
|
|
17
|
+
- ./diagnosing-frozen-cron-checkins.md
|
|
18
|
+
- ../../library/features/cron-execution-monitoring.md
|
|
19
|
+
- ../../../1.0/standards/backend-php.md
|
|
20
|
+
---
|
|
21
|
+
|
|
22
|
+
## Summary
|
|
23
|
+
How to answer *"did this cron actually run, and what did it do?"* on the 1.0 `worker` tier, where
|
|
24
|
+
**there is no usable stdout** and **a skipped run leaves no trace at all**. Written from a
|
|
25
|
+
production diagnostic session on
|
|
26
|
+
`worker/crons/toga2/compass/workflow/3a_import_office_depot_purchase_orders.php`; the mechanics
|
|
27
|
+
apply to every cron on the tier.
|
|
28
|
+
|
|
29
|
+
Use this before you conclude a cron "never fired" — three different situations look identical
|
|
30
|
+
from the outside.
|
|
31
|
+
|
|
32
|
+
## Step 1 — get the real schedule from `schedules/*.json`, never from the file
|
|
33
|
+
The `.php` file's own header comment is not authoritative and is routinely stale. So are skills,
|
|
34
|
+
docs and tickets. `worker/schedules/cron.<env>[.<role>].json` is the only source of truth.
|
|
35
|
+
|
|
36
|
+
Worked example: the entry *"Download EDI From S3 Files & Create PO - EDI Office Depot"* in
|
|
37
|
+
`cron.worker.sync.json` is `*/5 * * * *` — **every 5 minutes**. The cron file's header said
|
|
38
|
+
"EVERY HOUR" (corrected 2026-08-25) and at least one skill still says hourly. A wrong assumed
|
|
39
|
+
frequency makes you read a normal gap as an outage.
|
|
40
|
+
|
|
41
|
+
Also check whether the job is registered in the env you are testing in at all: this one is **not**
|
|
42
|
+
in `cron.beta.json`, so beta never runs it on a schedule. Nothing on beta is evidence about it.
|
|
43
|
+
|
|
44
|
+
## Step 2 — echo/print goes nowhere useful, so do not debug with it
|
|
45
|
+
The crontab line is assembled in `worker/ebs/cron.worker.php`, and the two blocks that build it
|
|
46
|
+
differ:
|
|
47
|
+
|
|
48
|
+
| Schedule file | Command written | Where output goes |
|
|
49
|
+
|---|---|---|
|
|
50
|
+
| base `cron.<env>.json` | `php /var/www/html/crons/<script>` | **nowhere** — the ` > /var/www/cache/WORKER_ERROR_$RANDOM` redirect is **commented out** |
|
|
51
|
+
| role `cron.worker.<role>.json` | `php /var/www/html/crons/<script> > /var/www/cache/WORKER_ERROR_$RANDOM` | a **randomly named** file under `/var/www/cache/` on whichever of the 7 instances ran it |
|
|
52
|
+
|
|
53
|
+
Either way you cannot go and read it: base-schedule output is discarded outright, and a
|
|
54
|
+
role-schedule run leaves an unpredictable filename on a box in an autoscaled fleet. **Treat
|
|
55
|
+
worker cron stdout as write-only.**
|
|
56
|
+
|
|
57
|
+
## Step 3 — put temporary tracing somewhere you can query
|
|
58
|
+
What worked: write timestamped step markers into the **`note`** column of the run's
|
|
59
|
+
`Common.CronJobExecutions` row (the row `App_Framework::cronInitialization()` already inserted),
|
|
60
|
+
then read them back over the DB tooling from your desk. `dtCheckIn`, `dtCheckOut` and
|
|
61
|
+
`executionTimeSeconds` on the same row give you the shape of the run for free.
|
|
62
|
+
|
|
63
|
+
Rules: keep it to a few markers, and **remove the tracing once diagnosed** (it was removed in
|
|
64
|
+
this session). Never write payloads or credentials into `note`.
|
|
65
|
+
|
|
66
|
+
## Step 4 — read `Common.CronJobExecutions` to tell "ran", "hung" and "skipped" apart
|
|
67
|
+
Query the legacy env's `Common.CronJobExecutions` filtering `job LIKE` the script path:
|
|
68
|
+
|
|
69
|
+
- **`dtCheckIn` + `dtCheckOut` + `executionTimeSeconds`** — it ran and finished; the duration
|
|
70
|
+
tells you whether it fits inside its schedule slot.
|
|
71
|
+
- **`dtCheckIn` with no `dtCheckOut`** — it started and never completed (crash, timeout, or the
|
|
72
|
+
instance went away).
|
|
73
|
+
- **no row for a slot you expected** — the run was **skipped by the overlap guard**, not "never
|
|
74
|
+
fired". A blocked run writes **nothing**: `cronInitialization()` calls
|
|
75
|
+
`exitIfProcessRunning()` **before** the INSERT, and the `cronFinished(true)` it then calls is a
|
|
76
|
+
no-op because `cronLogId` is unset. See
|
|
77
|
+
[Cron Execution Monitoring](../../library/features/cron-execution-monitoring.md).
|
|
78
|
+
|
|
79
|
+
### The overlap guard is a `ps -ef` substring match — anything can block it
|
|
80
|
+
`App_Framework::isProcessRunning()` returns true for **any** `ps -ef` line containing the script
|
|
81
|
+
path (it only excludes `/bin/sh` lines and its own pid). So a developer's `tail -f`, an editor, or
|
|
82
|
+
any shell holding that path in its command line **blocks the cron indefinitely** while looking
|
|
83
|
+
like a legitimate overlap. Check `ps -ef` for what is actually holding the name before assuming a
|
|
84
|
+
long-running instance of the job itself.
|
|
85
|
+
|
|
86
|
+
## Step 5 — if runs overlap, look at what the job does before its real work
|
|
87
|
+
A job whose own runtime exceeds its schedule interval silently loses most of its slots to the
|
|
88
|
+
guard. Measure the phases, not the total: in the ODP case an instrumented run showed the S3
|
|
89
|
+
listing alone returning **138,472 objects and taking 32 seconds** before any PO was touched, with
|
|
90
|
+
each PO then costing roughly 40 seconds — comfortably past a 5-minute schedule. Prefixes the job
|
|
91
|
+
skips while processing are still listed, so accumulated `SENT/` and `OUTBOX/` objects inflate
|
|
92
|
+
every run.
|
|
93
|
+
|
|
94
|
+
## Deploy-time trap: a PHP 8-only syntax kills the whole cron with no visible error
|
|
95
|
+
This tier runs **PHP 7.2**, so a **PHP 8 named argument** (`someFunction(items: $x)`) is a
|
|
96
|
+
**parse error**, not a runtime warning:
|
|
97
|
+
|
|
98
|
+
```
|
|
99
|
+
PHP Parse error: syntax error, unexpected ':', expecting ',' or ')'
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
The script then does nothing at all on every tick — and per Steps 2 and 4 you will see no output
|
|
103
|
+
and no `CronJobExecutions` row, i.e. it looks exactly like "the cron never fired". Typed
|
|
104
|
+
parameters and return types (`: void`, `: array`, `: bool`, `?array`, `?object`, `object $x`) are
|
|
105
|
+
PHP 7.0/7.1 and **are** safe here; they are already used in
|
|
106
|
+
`worker/crons/toga2/compass/backfill_items_isfulfillable.php` and
|
|
107
|
+
`update_item_fulfillments_in_netsuite.php`.
|
|
108
|
+
|
|
109
|
+
`1.0/standards/backend-php.md` already forbids named arguments in 1.0 — but a workspace
|
|
110
|
+
`CLAUDE.md` / `.github/copilot-instructions.md` that says *"use PHP 8 Named Parameters"* applies
|
|
111
|
+
to **2.0 only**. The 1.0 standard wins in `library/`, `worker/` and every other 1.0 app.
|
|
112
|
+
|
|
113
|
+
**Do not read a PHP 8 function call in existing code as proof the tier is PHP 8.**
|
|
114
|
+
`library/app/systemmonitor/500error.php` uses `str_contains()`; that file is not a version
|
|
115
|
+
signal, it is a latent bug.
|
|
116
|
+
|
|
117
|
+
## Change history
|
|
118
|
+
- 2026-08-25 — Documented from a prod diagnostic session on the Compass ODP 850 importer: the
|
|
119
|
+
schedule JSON (not the file header) is authoritative, worker cron stdout is unreadable (base
|
|
120
|
+
block's `WORKER_ERROR` redirect commented out, role block writes a random filename),
|
|
121
|
+
`Common.CronJobExecutions.note` is the practical place for temporary tracing, a run skipped by
|
|
122
|
+
the `ps -ef` overlap guard writes **no row at all** (guard runs before the INSERT), the guard
|
|
123
|
+
matches any process holding the script path, and a PHP 8 named argument parse-errors the whole
|
|
124
|
+
cron silently on this PHP 7.2 tier. (bala)
|
|
@@ -132,6 +132,16 @@ can keep using `sendEmail($api, ...)`.
|
|
|
132
132
|
predated the commit. That is a **deploy gap, not a code bug** — a redeploy fixes it, and no code
|
|
133
133
|
change should be made. Same known-issue as
|
|
134
134
|
[api2 environment-variable-drives-underscore-branch](../../api2/features/environment-variable-drives-underscore-branch.md).
|
|
135
|
+
- **⚠ An explicit `null` recipient is a 400, not a 500 (fixed 2026-08-25).** `$to`/`$cc`/`$bcc` on
|
|
136
|
+
`sendEmail()`/`send()`/`dispatch()` are typed `string|array|null` (a scripted-API caller can pass an
|
|
137
|
+
argument that is present-but-`null` — the `[]` default only applies when the arg is **omitted**).
|
|
138
|
+
Before the fix they were non-nullable `string|array`, so an explicit `null` `$to` threw
|
|
139
|
+
`Argument #3 ($to) must be of type array|string, null given` → **HTTP 500** — this was the platform's
|
|
140
|
+
**largest 5xx contributor** (prod `Logs.Issue` ref `2F`, 828 occurrences). `dispatch()` now normalizes
|
|
141
|
+
each recipient at the top (`(array)($x ?? [])`); after merging the template's stored
|
|
142
|
+
`EmailTemplateOutgoingEmailAddress` rows, if `$to` is **still empty** it throws `_Exception_Validation`
|
|
143
|
+
(→ **HTTP 400**) rather than a TypeError-500 or sending a recipientless email. A missing recipient is
|
|
144
|
+
client input error, not a server fault.
|
|
135
145
|
- **`sendEmail()`'s signature is load-bearing for scripted APIs** — the Record Script engine
|
|
136
146
|
(`api2/Component/Api/V2/V2.php`, ~line 3594) calls the method with `api` as a named
|
|
137
147
|
argument, so the first param must stay `&$api`. Do not "clean it up" by removing it.
|
|
@@ -198,6 +208,13 @@ worker method) in-process instead.
|
|
|
198
208
|
restored it — and that `sendEmail` has been variadic since 2024-12-24, so the spread was always
|
|
199
209
|
correct and May was the regression. Documented the contract on `sendEmail()`'s docblock plus a marker
|
|
200
210
|
comment at each of the four Quad call sites. (apeterson)
|
|
211
|
+
- 2026-08-25 — Made `$to`/`$cc`/`$bcc` nullable (`string|array|null`) on `sendEmail()`/`send()`/
|
|
212
|
+
`dispatch()` and normalized each to an array at the top of `dispatch()` (`(array)($x ?? [])`),
|
|
213
|
+
fixing the platform's **largest 5xx contributor** (prod `Logs.Issue` ref `2F`, 828 occurrences): a
|
|
214
|
+
scripted-API caller passing an explicit `null` recipient hit `Argument #3 ($to) must be of type
|
|
215
|
+
array|string, null given` → HTTP 500. After merging the template's stored recipient addresses, an
|
|
216
|
+
empty `$to` now throws `_Exception_Validation` (HTTP 400) instead of a TypeError-500 or a
|
|
217
|
+
recipientless send. Added the corresponding gotcha. (jcardinal)
|
|
201
218
|
- 2026-08-18 — Fixed the Quad order-approval/rejection emails' production 500
|
|
202
219
|
(`str_replace(): Argument #2 must be string`, EO-1): the four `Model/Quad/ApprovalDecision.php`
|
|
203
220
|
+ `Model/Quad/SalesOrder.php` call sites passed the template-vars map as a bare positional
|
|
@@ -6,8 +6,8 @@ project: _Underscore
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-08-
|
|
10
|
-
owners: ["dfranks", "jcardinal", "mhammontree", "apeterson", "kyalamarthi"]
|
|
9
|
+
updated: 2026-08-24
|
|
10
|
+
owners: ["dfranks", "jcardinal", "mhammontree", "apeterson", "kyalamarthi", "bala"]
|
|
11
11
|
files:
|
|
12
12
|
- _underscore/Database.php
|
|
13
13
|
- _underscore/Model.php
|
|
@@ -177,6 +177,16 @@ here — they live in `Config/*.ini`.)
|
|
|
177
177
|
- Related 1.0 analogue: the legacy `App_` worker has the same hazard writing to `Logs.API`
|
|
178
178
|
(`db_logs`) — the laptop trap there is documented separately in the worker NetSuite bootstrap
|
|
179
179
|
notes.
|
|
180
|
+
- **`_Database::register()` is LAZY — it stores connection config and never opens a socket.**
|
|
181
|
+
The connect happens on first use. So adding a static alias registration to an app's `_.php`
|
|
182
|
+
bootstrap is **safe across every environment**, even one whose config group points at a
|
|
183
|
+
localhost that lacks the schema: nothing fails until something actually queries that alias.
|
|
184
|
+
This is what makes "register the alias globally, use it in one cron" a two-line change rather
|
|
185
|
+
than a per-environment config exercise (worked example: `DB_VISION_LOGS` in
|
|
186
|
+
[the Worker Fleet Role-Assignment Monitor](../../worker2/features/worker-fleet-role-assignment-monitor.md)).
|
|
187
|
+
The corollary is the trap the rest of this doc describes: because registration is free and
|
|
188
|
+
silent, a **bad** registration also stays silent until the request that needs it blows up with
|
|
189
|
+
`Unknown database`.
|
|
180
190
|
|
|
181
191
|
## WITHDRAWN — the "alias-keyed `$_modelCache` cross-tenant leak" hypothesis (2026-08-17)
|
|
182
192
|
|
|
@@ -210,6 +220,11 @@ also hits `_modules/<module>/`** — is recorded in
|
|
|
210
220
|
|
|
211
221
|
## Change history
|
|
212
222
|
|
|
223
|
+
- 2026-08-24 — Recorded that **`_Database::register()` is lazy** (stores config, opens no
|
|
224
|
+
socket; the connect happens on first use), so a static alias registration in an app's `_.php`
|
|
225
|
+
is safe in every environment including a localhost config that lacks the schema — and that
|
|
226
|
+
this is also why a *bad* registration stays silent until first query. Verified while adding
|
|
227
|
+
`DB_VISION_LOGS` to worker2. (bala)
|
|
213
228
|
- 2026-08-17 (later pass) — **WITHDRAWN, supersedes the entry below.** The alias-keyed
|
|
214
229
|
`$_modelCache` cross-tenant-leak hypothesis is **not supported** and is no longer a live security
|
|
215
230
|
concern. The six `c_` columns are **AIG's own**: `_modules/netsuite/2026-07-10a -
|
|
@@ -6,7 +6,7 @@ project: API
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: architecture
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-08-
|
|
9
|
+
updated: 2026-08-25
|
|
10
10
|
owners: [jcardinal, bala, mhammontree, dfranks]
|
|
11
11
|
files:
|
|
12
12
|
- api2/Controller/Index.php
|
|
@@ -114,14 +114,22 @@ One ~2,000-line `execute()` then `processRoutePairs()`:
|
|
|
114
114
|
5. **Transaction logging** — every request logged (to client/core Logs DB, or as a JSONL
|
|
115
115
|
line shipped by CloudWatch when `[api] log_filepath` is set).
|
|
116
116
|
|
|
117
|
-
>
|
|
117
|
+
> **Auto-generated `Api.transactionId` used to collide under concurrency → 1062 → HTTP 500 (FIXED 2026-08-25).**
|
|
118
118
|
> Separate from the *client-supplied* `transactionId` uniqueness check (EV-5, above): the inbound
|
|
119
|
-
> request-logger
|
|
120
|
-
>
|
|
121
|
-
>
|
|
122
|
-
>
|
|
123
|
-
>
|
|
124
|
-
>
|
|
119
|
+
> request-logger inserted its `Api` log row with a UNIQUE `transactionId` (`Logs.Api.transactionId`,
|
|
120
|
+
> `varchar(255) UNIQUE`, in both Core `Logs` and `Logs_<Client>`) set to a millisecond-precision
|
|
121
|
+
> timestamp (`Y-m-d H:i:s.v`). Concurrent inserts in the same millisecond collided on that UNIQUE key
|
|
122
|
+
> → MySQL **1062** → the logger threw and turned a *successful* request into **HTTP 500 (EO-1)**. This
|
|
123
|
+
> broke ingestion for high-volume senders (seen on the Compass/Veyer ASN feed, `sourceIp 34.232.23.158`;
|
|
124
|
+
> confirmed in prod `Logs.Issue` reference `1Z`, 304 occurrences, `Duplicate entry
|
|
125
|
+
> '<ms-timestamp>' for key 'Api.transactionId'`, stack `/v2/users/me → _Model_Client_User::me() →
|
|
126
|
+
> internalApiRequest() → _Model->save()`).
|
|
127
|
+
> **Fix:** the two DB-insert log sites now set the log-row `transactionId` to a random uuid
|
|
128
|
+
> (`_String::generateUuid()`) so it cannot collide — `V2.php` `internalApiRequest()` logger (~L2370)
|
|
129
|
+
> and `Controller/Index.php` uncaught-`Throwable` error-recovery logger (~L456). The request's own base
|
|
130
|
+
> `transactionId` was already a uuid when auto-generated (`V2.php` ~L2130); only the bare-ms log-row id
|
|
131
|
+
> was at fault. The **sibling CloudWatch JSONL file-append logger** (`Controller/Index.php` ~L433) is
|
|
132
|
+
> **intentionally left bare-ms** — a JSONL append has no UNIQUE key and cannot collide.
|
|
125
133
|
|
|
126
134
|
## CRUD engine — `processRoutePairs()`
|
|
127
135
|
|
|
@@ -216,6 +224,34 @@ IMDSv2 token on the `curl` command line (visible in `ps`/`/proc/<pid>/cmdline`,
|
|
|
216
224
|
back silently to IMDSv1 — fix both alongside the key. *(Location + remediation only; no key material
|
|
217
225
|
is recorded anywhere.)*
|
|
218
226
|
|
|
227
|
+
## EB "Degraded" health alerts vs. the ALB's own view (monitoring — read before triaging)
|
|
228
|
+
|
|
229
|
+
An Elastic Beanstalk enhanced-health transition — *"Environment health transitioned Ok→Degraded …
|
|
230
|
+
One or more TargetGroups … in a reduced health state: awseb-AWSEB-<id> - Degraded"* — **does NOT
|
|
231
|
+
mean the ALB marked a target unhealthy.** On `api-production-1` (AWS account `654654170868`), during
|
|
232
|
+
such an event the ALB's own view was **fully healthy throughout**: `HealthyHostCount=2`,
|
|
233
|
+
`UnHealthyHostCount=0`, ELB `5xx=0`, `TargetConnectionErrorCount=0`.
|
|
234
|
+
|
|
235
|
+
The two subsystems measure different things and legitimately disagree during any transient:
|
|
236
|
+
|
|
237
|
+
- **EB enhanced health** rates the target group from its **own per-instance request-latency/status
|
|
238
|
+
sampling on a point-in-time snapshot** — a short traffic burst (observed `RequestCount ~10x`
|
|
239
|
+
baseline, peaking ~1,936/min, with a `TargetResponseTime` max outlier ~7.7 s) is enough to flip it
|
|
240
|
+
to *Degraded*.
|
|
241
|
+
- **The target-group console** shows **current ALB health**, which recovers within ~45 s (3 checks).
|
|
242
|
+
So when you open the console after the alert, it "looks perfectly healthy" — because it is, now.
|
|
243
|
+
|
|
244
|
+
api2's EB target group health-checks **`/health`** (interval 15 s, timeout 5 s, unhealthy threshold
|
|
245
|
+
5 ≈ **75 s to trip**, healthy threshold 3 ≈ **45 s to recover**). A transient burst is **expected
|
|
246
|
+
behavior, not an ALB/target-group defect** — do not chase a phantom target failure.
|
|
247
|
+
|
|
248
|
+
**The exception:** some *Degraded* events are instead genuine **"X% HTTP 5xx"** — those are real
|
|
249
|
+
code bugs, not transient bursts. The 2026-08 batch traced to the `transactionId`-1062 collision
|
|
250
|
+
(gotcha above), the platform-wide EmailTemplate null-recipient TypeError, and the Pcmaticb2b ticket
|
|
251
|
+
interceptor's plain-`Exception` rejections — all now returning 4xx/fixed. Distinguish the two by
|
|
252
|
+
reading the ALB 5xx metric and `Logs.Issue`: zero 5xx + a traffic spike = transient EB snapshot;
|
|
253
|
+
a sustained 5xx rate = a code bug to fix.
|
|
254
|
+
|
|
219
255
|
## Known issues / accepted risks
|
|
220
256
|
|
|
221
257
|
Open items a maintainer should know before changing this tier. None are "bugs to fix right now" —
|
|
@@ -289,6 +325,18 @@ they are the known sharp edges. Do not re-discover these from scratch.
|
|
|
289
325
|
single-caller branches in `V2.php` as unverified until exercised directly.
|
|
290
326
|
|
|
291
327
|
## Change history
|
|
328
|
+
- 2026-08-25 — **Resolved the auto-generated `Api.transactionId` 1062 collision** (was the 2026-07-28
|
|
329
|
+
gotcha): the two DB-insert log sites now uuid the log-row `transactionId` (`_String::generateUuid()`)
|
|
330
|
+
— `V2.php` `internalApiRequest()` logger (~L2370) and `Controller/Index.php` error-recovery logger
|
|
331
|
+
(~L456); the CloudWatch JSONL append (~L433) is intentionally left bare-ms (no UNIQUE key). Confirmed
|
|
332
|
+
the production symptom (`Logs.Issue` ref `1Z`, 304 occurrences, `/v2/users/me` → `_Model->save()`).
|
|
333
|
+
Also added a **monitoring section**: an EB enhanced-health *Ok→Degraded* "TargetGroups … reduced
|
|
334
|
+
health state" transition does **not** mean the ALB marked a target unhealthy — on `api-production-1`
|
|
335
|
+
(acct `654654170868`) the ALB view stayed fully healthy (HealthyHostCount=2, 5xx=0) while a short
|
|
336
|
+
~10x traffic burst (~1,936/min, TargetResponseTime max ~7.7 s) flipped EB's point-in-time snapshot;
|
|
337
|
+
the target-group console recovers in ~45 s so it "looks healthy" after the fact. Recorded api2's
|
|
338
|
+
`/health` check timing (15 s / 5 s / trip 75 s / recover 45 s) and how to tell a transient burst
|
|
339
|
+
(0 ALB 5xx) from a genuine "X% HTTP 5xx" code bug. (jcardinal)
|
|
292
340
|
- 2026-07-28 — Added Known issue #10: cross-client page-number paging is inherently O(page) (deep pages materialize every preceding row into the Cache cluster), with caller-facing keyset cursors identified as the durable fix but left unbuilt/unscoped, plus the `DEEPEN_CHUNK_MAX_RECORDS = 500` tuning tradeoff against `getFullModelData()` expansion and `CURL_TIMEOUT_SECONDS = 30`. Added a change-guidance bullet that single-caller branches in the untested `V2.php` monolith get zero incidental coverage (the keyset `LIMIT` syntax error shipped invisibly for a full cycle). (jcardinal)
|
|
293
341
|
- 2026-07-28 — Documented the previously unrecorded `060_register_instance_to_shared_application_load_balancer` postdeploy hook pair in the Deployment section (non-prod self-registration into the same-named ALB target group; production skipped), and recorded a **committed IAM access key** in `ebs/register_instance_to_shared_application_load_balancer.php` as a security note + Known issue #9 — location, line range, commit subject, and remediation only (rotate, audit CloudTrail, move to the instance profile; history rewrite is a separate sign-off). Flagged api2's copy as the unhardened original vs. the new worker2 reference implementation. (jcardinal)
|
|
294
342
|
- 2026-07-28 — Added a consolidated **Known issues / accepted risks** section (8 items), absorbing the previously free-floating deferred raw-exception-disclosure follow-up as item 1, so the tier's sharp edges (unrotated committed secrets, pre-execute phase still outside the main guard, local Logs DB name mismatch, permissive CORS, unpinned `_underscore` build clone, untested `V2.php` monolith, JWT rotation overlap window) are in one place instead of scattered. Recorded that `DB_CACHE` is resolved by name (`Databases.name = 'Cache'`), never by a hardcoded id, which differs per Core instance. (jcardinal)
|
|
@@ -17,7 +17,7 @@
|
|
|
17
17
|
| [ClickUp Work Type Automation (Committed / Conditional / Stretch)](features/clickup-work-type-automation.md) | The ClickUp webhook handler (`_Worker_Clickup`) automatically maintains each task's **Work Type** custom field — `Committed`, `Conditional`, or `Stretch` — base | worker2/Worker/Clickup.php, worker2/Tests/Worker/ClickupWorkTypeTest.php |
|
|
18
18
|
| [Compass VIP Support Importer (worker2)](features/compass-vip-support-importer.md) | A worker2 action that ingests Compass's quarterly VIP spreadsheet and assigns each VIP user's support technician by setting `Users.c_supportedByUserId` in `Clie | worker2/Worker/Client/Compass/VipSupport.php |
|
|
19
19
|
| [Creating Worker Actions](features/creating-worker-actions.md) | How to add a new callable Worker action — a PHP class whose `public static` methods are invoked as background jobs (via webhook, cron, or `_Worker::runTask()`). | worker2/Worker/, worker2/Controller/Index.php, _underscore/Worker.php |
|
|
20
|
-
| [Cross-account AWS access for worker2 crons (_Component_Aws_Workloads)](features/cross-account-aws-access.md) | `_Component_Aws_Workloads` is the **standard, and only sanctioned, way any new worker2 cron obtains AWS access** — for any account, any region, any AWS SDK clie | worker2/Component/Aws/Workloads/Workloads.php, worker2/Config/production.ini, worker2/Worker/Infrastructure/CloudWatch.php |
|
|
20
|
+
| [Cross-account AWS access for worker2 crons (_Component_Aws_Workloads)](features/cross-account-aws-access.md) | `_Component_Aws_Workloads` is the **standard, and only sanctioned, way any new worker2 cron obtains AWS access** — for any account, any region, any AWS SDK clie | worker2/Component/Aws/Workloads/Workloads.php, worker2/Config/production.ini, worker2/Worker/Infrastructure/CloudWatch.php, worker2/Worker/Monitor/Fleet.php, library/app/worker.php, library/app/cloud.php |
|
|
21
21
|
| [Elastic Beanstalk health monitor → OneUptime (ElasticBeanstalkHealth)](features/elastic-beanstalk-health-monitor.md) | `_Worker_Infrastructure_CloudWatch::ElasticBeanstalkHealth()` is a worker2 cron that reads each Elastic Beanstalk (EB) environment's **enhanced-health** status | worker2/Worker/Infrastructure/CloudWatch.php |
|
|
22
22
|
| [Elite Freshservice Sync (worker2)](features/elite-freshservice-sync.md) | `_Worker_Elite` processes Freshservice webhook events and syncs them into TOGA 2. | worker2/Worker/Elite.php, worker2/Config/dev-kmaramreddy-laptop.ini |
|
|
23
23
|
| [Error Escalation Cron (Errors::Escalate → ClickUp / email)](features/error-escalation-cron.md) | `_Worker_Infrastructure_Errors::Escalate` (renamed from `SyncWithClickup`) is the sole owner of **escalation, de-escalation, ClickUp ticketing, reminders, busin | worker2/Worker/Infrastructure/Errors.php, worker2/Worker/Notification/Email.php, worker2/Worker/Notification/EmailTemplate.php, worker2/Worker/Client/True.php, worker2/Worker/Clickup/ErrorTask.php, worker2/Worker/Clickup.php, worker2/Controller/Index.php, worker2/Config/production.ini, _underscore/Model/Core/Logs/Issue.php, dbchanges2/Core/2026-07-30a - Error escalation cron job.sql, dbchanges2/Logs/2026-08-03a - Issue clickupPriority.sql, dbchanges2/Core/2026-08-04a - Error neglect digest cron job.sql |
|
|
@@ -47,6 +47,7 @@
|
|
|
47
47
|
| [Centralized Tracking-Status Refresh (worker2 Sync cron, all clients)](features/tracking-status-refresh.md) | A single platform-wide cron keeps `TrackingNumbers.status` current for **every provisioned client** until each shipment reaches a terminal state, by polling Fed | worker2/Worker/Sync/TrackingNumbers.php, _underscore/Component/Library/Carriers/Fedex/Fedex.php, _underscore/Component/Library/Carriers/Usps/Usps.php, _underscore/Component/Library/Carriers/Ups/Ups.php, _underscore/Model/Client/TrackingNumber.php, dbchanges2/Client/2026-07-30 - TrackingNumbers_RefreshColumns.sql, dbchanges2/Client/2026-07-30 - Acl_TrackingStatus_Grant.sql, dbchanges2/Client_Tdsynnex/2026-07-30 - Apis_TdsynnexKey.sql, dbchanges2/Core/2026-07-30 - TrackingRefresh_Cron.sql |
|
|
48
48
|
| [VAPI Webhook Handler (worker2 — AI-BDR end-of-call processing)](features/vapi-webhook-handler.md) | `_Worker_Vapi` ([worker2/Worker/Vapi.php](worker2/Worker/Vapi.php)) is the **PHP side of the AI-BDR call loop** — the webhook that receives VAPI's end-of-call r | worker2/Worker/Vapi.php, worker2/Worker/Ai/Bdr/Vapi.php, worker2/Controller/Index.php |
|
|
49
49
|
| [WJE Freshservice Sync (worker2)](features/wje-freshservice-sync.md) | WJE ("WJE IT", helpdesk `wje.freshservice.com`) is a **Freshservice**-based help-desk client whose tickets, contacts, assets, groups, categories, and canned res | worker2/Worker/Wje.php, _underscore/Component/Api/Wje/Wje.php, _underscore/Model/Wje/Ticket.php, _underscore/Model/Wje/TicketNote.php, _underscore/Model/Wje/Contact.php, _underscore/Model/Wje/Unit.php, _underscore/Model/Wje/TicketTeam.php, _underscore/Model/Wje/TicketCategory.php, _underscore/Model/Wje/AssetType.php, _underscore/Model/Wje/PredefinedReply.php, library/app/api/wje.php, worker/crons/toga2/wje/import_supporting_records.php, worker/crons/toga2/wje/sync_togasupply_wje.php, worker/crons/notifications/reports/wje/wje_common.php, library/app/systemmonitor/wje.php, dbchanges2/Client_Wje/2024-10-04 - WjeOnboarding.sql |
|
|
50
|
+
| [1.0 Worker Fleet Role-Assignment Monitor (Monitor/Fleet/RoleAssignment)](features/worker-fleet-role-assignment-monitor.md) | `_Worker_Monitor_Fleet::RoleAssignment()` is a **2.0 worker2 cron that watches the 1.0 `worker` fleet from the outside**. | worker2/Worker/Monitor/Fleet.php, worker2/_.php, dbchanges2/Core/2026-08-24 - Worker Role Assignment Monitor.sql, worker/ebs/cron.worker.php, worker/crons/worker/worker_heartbeat.php, library/app/worker.php |
|
|
50
51
|
| [PHP Runtime Upgrade on Elastic Beanstalk (worker2 8.3 → 8.5 + PhpSpreadsheet 1.x → 3.x)](workflows/php-runtime-upgrade-dependency-audit.md) | The procedure used to move worker2 from **PHP 8.3 to PHP 8.5** on Elastic Beanstalk, and the dependency work that had to land first. | worker2/composer.json, worker2/composer.lock, worker2/Worker/Team/Sprint.php, worker2/Worker/Client/TowFoundation/ProcessReceipts.php, worker2/Worker/Forecast/Import.php |
|
|
51
52
|
| [Running worker2 locally against real NetSuite (and what EB does instead)](workflows/running-worker2-locally.md) | How to boot **worker2** on a developer machine and run a real worker action against the **live NetSuite** account and a **local** `Forecast` database. | worker2/index.php, worker2/composer.json, worker2/.ebextensions/php_include_underscore.config, worker2/.ebextensions/git.json, worker2/.ebextensions/git.php, worker2/Config/production.ini, _underscore/Component/Api/Netsuite/Netsuite.php, dbchanges2/Logs_Client/2026-04-08_BLANK_CLIENT_LOGS_DATABASE.SQL |
|
|
52
53
|
| [Ticket → ClickUp Pseudocode Planning (Talos-grounded)](workflows/ticket-to-pseudocode-planning.md) | A repeatable procedure for turning a ClickUp ticket into a reviewed, formatted implementation plan posted back to the ticket's `📝 Pseudocode` custom field. | test/@dave/clickup_md2delta.js |
|
|
@@ -6,18 +6,22 @@ project: Worker
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-08-
|
|
10
|
-
owners: [jcardinal]
|
|
9
|
+
updated: 2026-08-24
|
|
10
|
+
owners: [jcardinal, bala]
|
|
11
11
|
files:
|
|
12
12
|
- worker2/Component/Aws/Workloads/Workloads.php
|
|
13
13
|
- worker2/Config/production.ini
|
|
14
14
|
- worker2/Worker/Infrastructure/CloudWatch.php
|
|
15
|
+
- worker2/Worker/Monitor/Fleet.php
|
|
16
|
+
- library/app/worker.php
|
|
17
|
+
- library/app/cloud.php
|
|
15
18
|
related:
|
|
16
19
|
- ../../_underscore/features/component-model-namespace-registration.md
|
|
17
20
|
- ../../_underscore/features/config-group-access.md
|
|
18
21
|
- ../../_underscore/features/cloud-s3-helpers.md
|
|
19
22
|
- ./oneuptime-worker2-monitoring.md
|
|
20
23
|
- ./elastic-beanstalk-health-monitor.md
|
|
24
|
+
- ./worker-fleet-role-assignment-monitor.md
|
|
21
25
|
---
|
|
22
26
|
|
|
23
27
|
## Summary
|
|
@@ -79,7 +83,8 @@ region is only a runtime argument to `client()`, never a reason for a second rol
|
|
|
79
83
|
- **654654170868** — production hub. Home region `us-east-1`; also used in `us-west-2` and
|
|
80
84
|
`eu-west-1`. worker2 runs here on Elastic Beanstalk; its instance profile role
|
|
81
85
|
`ElasticBeanstalk-EC2-Instance-Profile` is the caller identity.
|
|
82
|
-
- **502614707982** — legacy account, `us-west-2`.
|
|
86
|
+
- **502614707982** — legacy account, `us-west-2`. **This is where the 1.0 `worker` fleet
|
|
87
|
+
actually runs** — see the account-confusion gotcha below.
|
|
83
88
|
- **975050298201** — non-production account, `us-east-1`.
|
|
84
89
|
|
|
85
90
|
Each `WorkloadsRuntime` role has broad service permissions (`elasticbeanstalk:*`, `ec2:*`,
|
|
@@ -144,8 +149,33 @@ code yet. New code must not adopt the static key.
|
|
|
144
149
|
**Never record the access-key values in knowledge.** They exist in `production.ini`; that is
|
|
145
150
|
their only home.
|
|
146
151
|
|
|
152
|
+
### Which account the 1.0 worker fleet lives in (easy and expensive to get wrong)
|
|
153
|
+
|
|
154
|
+
The 1.0 `worker` EC2 instances are in **502614707982**, `us-west-2` (EB environment
|
|
155
|
+
`agilant-worker`, confirmed in the EC2 console).
|
|
156
|
+
|
|
157
|
+
**`App_Worker::AWS_WORKER_QUEUE_URL` in `library/app/worker.php` points at 654654170868 and
|
|
158
|
+
misleads** — that account holds only the **SQS queue the 1.0 workers read from**, not the
|
|
159
|
+
instances. Anything that needs to *see* the worker boxes (EB, EC2, health) must target
|
|
160
|
+
**502614707982**.
|
|
161
|
+
|
|
162
|
+
worker 1.0 authenticates there with a **long-lived hardcoded IAM access key in
|
|
163
|
+
`library/app/cloud.php`** — a committed credential that should be rotated and moved to a role
|
|
164
|
+
(value deliberately not recorded here; see [Legacy static key](#legacy-static-key--being-retired)).
|
|
165
|
+
worker2 does **not** use it: it assumes `WorkloadsRuntime` cross-account from its EB instance
|
|
166
|
+
role, which is exactly the migration this component exists to enable. First worker2 consumer in
|
|
167
|
+
that account is the
|
|
168
|
+
[Worker Fleet Role-Assignment Monitor](./worker-fleet-role-assignment-monitor.md).
|
|
169
|
+
|
|
170
|
+
**Scope EB reads by `EnvironmentName`, never account-wide.** 502614707982 also runs
|
|
171
|
+
`agilant-worker-alpha` and `agilant-worker-beta` as **separate EB environments that have no role
|
|
172
|
+
registry at all**. An account-level EC2 scan would sweep them in and misreport them.
|
|
173
|
+
|
|
147
174
|
## Gotchas / known issues
|
|
148
175
|
|
|
176
|
+
- **Do not assume the 1.0 worker fleet is in the 868 production hub.** `AWS_WORKER_QUEUE_URL`
|
|
177
|
+
(654654170868) is the SQS queue only; the instances are in the legacy account 502614707982,
|
|
178
|
+
`us-west-2`. See [the section above](#which-account-the-10-worker-fleet-lives-in-easy-and-expensive-to-get-wrong).
|
|
149
179
|
- Do not construct `StsClient` or pass explicit credentials by hand in new crons — the whole
|
|
150
180
|
point is that the instance role is the caller and the component owns the assume-role step.
|
|
151
181
|
- The instance-role-as-caller only works **on the EB tier**. Off-tier (e.g. a local CLI run)
|
|
@@ -161,6 +191,15 @@ their only home.
|
|
|
161
191
|
see [creating-worker-actions.md](./creating-worker-actions.md).
|
|
162
192
|
|
|
163
193
|
## Change history
|
|
194
|
+
- 2026-08-24 — Pinned down **which account the 1.0 `worker` fleet actually runs in**:
|
|
195
|
+
**502614707982** (`us-west-2`, EB env `agilant-worker`), **not** the 654654170868 production hub
|
|
196
|
+
that `App_Worker::AWS_WORKER_QUEUE_URL` implies — that account holds only the SQS queue the 1.0
|
|
197
|
+
workers read from. Recorded that worker 1.0 reaches it with the hardcoded static key in
|
|
198
|
+
`library/app/cloud.php` (rotate + move to a role) while worker2 assumes `WorkloadsRuntime` from
|
|
199
|
+
its instance role, and that EB reads there must be scoped by `EnvironmentName` because
|
|
200
|
+
`agilant-worker-alpha`/`-beta` are separate registry-less environments in the same account.
|
|
201
|
+
Found while building the
|
|
202
|
+
[Worker Fleet Role-Assignment Monitor](./worker-fleet-role-assignment-monitor.md). (bala)
|
|
164
203
|
- 2026-08-11 — Noted that `ElasticBeanstalkHealth`'s `$environments` must be typed `object|array`
|
|
165
204
|
(not `array`): the nested JSON object arrives as `stdClass` because the dispatcher casts only
|
|
166
205
|
the top level of `CronJobs.parameters`, so a strict `array` throws a `TypeError` at dispatch.
|