toga-ai 1.0.645 → 1.0.646

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -6,12 +6,13 @@ project: Library
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-07-06
10
- owners: [dfranks]
9
+ updated: 2026-08-25
10
+ owners: [dfranks, bala]
11
11
  files:
12
12
  - library/app/framework.php
13
13
  related:
14
14
  - ../../worker/architecture.md
15
+ - ../../worker/workflows/tracing-a-worker-cron-run-in-production.md
15
16
  ---
16
17
 
17
18
  ## Summary
@@ -38,13 +39,28 @@ row in `db_log` and of the per-job Sentry check-in monitors.
38
39
  cast silently discards the sub-second precision even when the schema column is `decimal(10,3)`.
39
40
  - **Schema:** the table is provisioned by the dbchanges migration
40
41
  `dbchanges/Common/DF/2026-5-7 Cron Checkin.sql` (table `CronJobExecutions`,
41
- `executionTimeSeconds decimal(10,3)`, no `note` column).
42
+ `executionTimeSeconds decimal(10,3)`). **Correction (2026-08-25): prod `Common.CronJobExecutions`
43
+ does have a `note` TEXT column** — verified by writing to it and reading it back. `App_Framework`
44
+ never populates it, which makes it the practical place to park **temporary** step tracing for a
45
+ cron you are diagnosing (worker crons have no readable stdout — see
46
+ [Tracing a worker cron run in production](../../worker/workflows/tracing-a-worker-cron-run-in-production.md)).
47
+ Remove the tracing when you are done, and never write payloads or credentials into it.
48
+ - **Don't confuse the two `note`s.** `cronInitialization()` writes `note = 'Started execution'` to
49
+ the **`Log`** table on `db_log` (the older `CRON`/`recordType` row), not to `CronJobExecutions`.
42
50
  - **Design decision:** 1.0 records executions by writing **directly to the shared DB**
43
51
  (`db_common`), *not* by POSTing from 1.0 to a 2.0 API endpoint. Per Jeff Cardinal, 1.0 code
44
52
  must not POST to 2.0 code; the direct shared-DB write is the sanctioned path (TRUE-78182).
45
53
 
46
54
  ## Gotchas
47
55
 
56
+ - **⚠ A run skipped by the overlap guard leaves NO ROW — absence is ambiguous.**
57
+ `cronInitialization()` calls `exitIfProcessRunning()` **before** the `CronJobExecutions` INSERT,
58
+ and the `cronFinished(true)` that guard then calls does nothing because `cronLogId` was never
59
+ set. So a blocked run is indistinguishable from a run that never launched: **a missing row for
60
+ an expected slot means "skipped", not "hung"**, and a hung run looks different (it has
61
+ `dtCheckIn` with no `dtCheckOut`). `isProcessRunning()` matches **any** `ps -ef` line containing
62
+ the script path (it only excludes `/bin/sh` and its own pid), so a stray `tail -f` or editor on
63
+ that path blocks the cron indefinitely.
48
64
  - **`db_log` has no `CronJobExecutions` table.** A `cronFinished()` UPDATE that runs against the
49
65
  `db_log` connection throws on **every** 1.0 cron job. Both the INSERT and the UPDATE must
50
66
  target `db_common`.
@@ -60,6 +76,12 @@ row in `db_log` and of the per-job Sentry check-in monitors.
60
76
  logic. (Fixed 2026-07-06: consolidated back to one `db_common` INSERT/UPDATE pair.)
61
77
 
62
78
  ## Change history
79
+ - 2026-08-25 — Corrected the schema note (prod `CronJobExecutions` **does** carry a `note` TEXT
80
+ column, unused by `App_Framework`, usable for temporary cron tracing) and separated it from the
81
+ `note = 'Started execution'` row `cronInitialization()` writes to `Log` on `db_log`. Recorded
82
+ that the overlap guard runs **before** the INSERT, so a skipped run writes no row at all and
83
+ cannot be told apart from "never launched" — and that `isProcessRunning()`'s `ps -ef` substring
84
+ match lets any process holding the script path block the cron. No code change. (bala)
63
85
  - 2026-07-06 — Fixed a merge (`origin/_production`) that reintroduced duplicate
64
86
  `CronJobExecutions` writes and pointed one `cronFinished()` UPDATE at `db_log` (where the
65
87
  table doesn't exist), which would have thrown on every 1.0 cron. Consolidated to one
@@ -11,9 +11,10 @@
11
11
  | [NetSuite Sales Order Sales Rep Sourcing (Staples & ODP EDI orders)](features/netsuite-sales-order-sales-rep-sourcing.md) | How the **sales rep** on a NetSuite Sales Order is determined for the two 1.0 `worker` EDI order-creation integrations (Staples cXML and Compass/ODP EDI). | worker/crons/toga2/compass/workflow/5_create_netsuite_sales_orders_from_office_depot_purchase_orders.php, worker/crons/sync/staples/sync_staples_cxml.php, test/@Mark/NetSuite/TRUE_80451_customer_salesrep_diag.php |
12
12
  | [NetSuite → TOGa Supply Per-Client Sync (thin wrappers)](features/netsuite-togasupply-per-client-sync.md) | Syncs NetSuite transactions (sales orders, purchase orders, invoices, item receipts, item fulfillments, inventory adjustments) into each TOGa Supply (2.0) clien | worker/crons/toga2/netsuite/common_sync_togasupply.php, worker/crons/toga2/netsuite/sync_togasupply_canon.php, worker/crons/toga2/netsuite/sync_togasupply_elite.php, worker/schedules/cron.worker.sync.json, dbchanges2/_modules/netsuite/2026-04-01 - Parameters.sql, library/app/api/toga2.php, library/app/api/netsuite/rest.php, library/app/framework.php, library/app/systemmonitor/netsuiteintegration.php, test/@srija/Elite Testing/Service Requests/test_sync_togasupply_elite_section.php, test/@srija/Elite Testing/Service Requests/test_diagnose_togasupply_elite.php |
13
13
  | [OneUptime Server monitor + disk/memory hygiene on the 1.0 worker EB host](features/oneuptime-server-monitor-host-hygiene.md) | The 1.0 `agilant-worker` EB environment runs on the **legacy Amazon Linux 1 PHP 7.2 platform** (Apache httpd/prefork, s3fs mounts, cron) and repeatedly went dow | worker/.ebextensions/040_disk_memory_hygiene.config, worker/.ebextensions/045_oneuptime_agent.config, worker/ebs/cron.worker.php, worker/ebs/mount-s3fs-folders.php, worker/ebs/apache_settings.php, worker/ebs/setup_phpini.php |
14
- | [OneUptime external uptime monitoring for 1.0 workers](features/oneuptime-worker-uptime-monitoring.md) | Every 1.0 worker box self-reports its liveness to an external OneUptime monitor once per minute by curl-POSTing to a per-worker "Incoming Request" heartbeat URL | library/app/worker.php, worker/crons/worker/worker_heartbeat.php |
14
+ | [OneUptime external uptime monitoring for 1.0 workers](features/oneuptime-worker-uptime-monitoring.md) | Every 1.0 worker box self-reports its liveness to an external OneUptime monitor once per minute by curl-POSTing to a per-worker "Incoming Request" heartbeat URL | library/app/worker.php, worker/crons/worker/worker_heartbeat.php, worker/ebs/cron.worker.php, worker/.ebextensions/045_oneuptime_agent.config |
15
15
  | [Prudential: Send Shipments for the Day report (daily cron)](features/send-shipments-for-the-day.md) | Daily cron (9:00 PM) that emails Prudential and Dell stakeholders an Excel report of all devices shipped that day, including tracking number, serial number, emp | worker/crons/notifications/reports/send_shipments_for_the_day.php |
16
16
  | [Staples cXML Order Import (SFTP → NetSuite)](features/staples-cxml-order-import.md) | `sync_staples_cxml.php` is an **hourly** cron (runs at **:45**) that imports Staples cXML purchase orders from SFTP into NetSuite as Sales Orders, then writes a | worker/crons/sync/staples/sync_staples_cxml.php |
17
17
  | [Diagnosing frozen 1.0 worker cron check-ins (Sentry "missed" flood)](workflows/diagnosing-frozen-cron-checkins.md) | When 1.0 worker cron timestamps freeze and Sentry project `worker1` fills with **`missed`** check-ins, the intuitive diagnosis — a wedged `App_Framework::isProc | worker/.ebextensions/cron.config, library/app/worker.php |
18
18
  | [isFulfillable Multi-Client Backfill (all togasupply clients)](workflows/isfulfillable-multi-client-backfill.md) | One-time backfill that catches up `Items.isFulfillable` on **existing** items across **all 17 togasupply clients** (AIG, Broward Sheriff, Canon, Endeavor Health | worker/crons/toga2/netsuite/backfill_isfulfillable_all_clients.php, library/app/api/toga2.php |
19
19
  | [Onboarding a Client to the NetSuite TOGa Supply Sync](workflows/onboarding-client-to-netsuite-togasupply-sync.md) | How to add a new TOGa 2 client to the per-client NetSuite → TOGa Supply importer (`worker/crons/toga2/netsuite/`). | worker/crons/toga2/netsuite/sync_togasupply.php, worker/crons/toga2/netsuite/common_sync_togasupply.php, worker/crons/toga2/netsuite/sync_togasupply_elite.php, worker/schedules/cron.worker.sync.json, dbchanges2/_modules/netsuite/2026-04-01 - Parameters.sql, dbchanges2/_modules/netsuite/2026-08-05 - CLEAN NETSUITE CLINET.SQL |
20
+ | [Tracing a 1.0 worker cron run in production (no stdout, silent skips)](workflows/tracing-a-worker-cron-run-in-production.md) | How to answer *"did this cron actually run, and what did it do?"* on the 1.0 `worker` tier, where **there is no usable stdout** and **a skipped run leaves no tr | worker/ebs/cron.worker.php, worker/schedules/cron.worker.sync.json, library/app/framework.php |
@@ -6,8 +6,8 @@ project: Worker
6
6
  client: shared
7
7
  type: architecture
8
8
  status: active
9
- updated: 2026-07-29
10
- owners: [jcardinal, sking]
9
+ updated: 2026-08-24
10
+ owners: [jcardinal, sking, bala]
11
11
  files:
12
12
  - worker/index.php
13
13
  - worker/_/app/framework.php
@@ -20,6 +20,8 @@ files:
20
20
  related:
21
21
  - ../library/architecture.md
22
22
  - ../dbchanges/workflows/authoring-and-shipping-sql-files.md
23
+ - ./features/oneuptime-worker-uptime-monitoring.md
24
+ - ../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md
23
25
  ---
24
26
 
25
27
  ## Summary
@@ -44,6 +46,7 @@ which **self-elects a distinct role** (`notification`, `database`, `infrastructu
44
46
  under `crons/`; the schedule (cron registration) is the source of truth for what runs — a script
45
47
  that isn't scheduled never executes. Client integrations live under `crons/toga2/<client>/`.
46
48
  Use prepared statements for all SQL; never interpolate input.
49
+ A worker that fails deploy-time role election runs NO crons while EB still reads healthy.
47
50
 
48
51
  ## How a job becomes a cron (the dispatch pipeline)
49
52
 
@@ -108,6 +111,32 @@ is the most important and least obvious part of the architecture.
108
111
  > Net effect: roles are a **claim-the-first-free-slot pool**, and a dead worker's role is
109
112
  > reclaimed by its replacement within minutes. There is no static instance→role mapping to edit.
110
113
 
114
+ ### The failure mode: NO NAME = NO CRONTAB (and it is silent)
115
+
116
+ The election is the **single point of failure for an entire instance**, and the architecture gives
117
+ it no second chance:
118
+
119
+ - The claim runs **only at DEPLOY time** — `ebs/cron.worker.php` is written as the EB `appdeploy`
120
+ enact hook. **It never retries.** An instance that fails to claim stays nameless until the next
121
+ deploy or until something terminates it.
122
+ - The claim writes `/etc/worker-role` and **only then** appends `cron.worker.<role>.json` to the
123
+ crontab. So a nameless instance has **no role crontab at all** — it runs *nothing*, not even
124
+ `worker_heartbeat.php`, while **EB reports the environment healthy**. The box is up; it just
125
+ does no work.
126
+ - **The claim's DB connect is `@mysqli_connect(...)` — error-suppressed.** A boot-time connect
127
+ failure to `Vision_Log` therefore leaves the instance nameless **with no log line anywhere**.
128
+ This is the prime suspect for the 2026-08-24 incident (2 of 7 workers roleless for ~3 hours),
129
+ and it violates the no-`@`-suppression rule in `rules/toga/common/coding-style.md`. Removing the
130
+ `@` and logging the failure is the recommended fix.
131
+ - `Workers` **only ever holds instances that SUCCEEDED**, so a failed claim leaves no row —
132
+ **absence is the only evidence**, and detecting it requires the AWS instance list.
133
+
134
+ Because every OneUptime monitor on this tier is keyed by **role** rather than by instance, none of
135
+ them can see this (see
136
+ [the blind spot](./features/oneuptime-worker-uptime-monitoring.md)). External coverage now comes
137
+ from the 2.0
138
+ [Worker Fleet Role-Assignment Monitor](../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md).
139
+
111
140
  ## Anatomy of a cron script
112
141
 
113
142
  Every script is self-contained and follows this boilerplate:
@@ -208,6 +237,10 @@ web face (`mvc/` GET routes for login/logout/404); the tier's real work is the c
208
237
 
209
238
  ## Conventions & gotchas
210
239
 
240
+ - **A worker that fails role election is silently dead, not degraded.** No name → no crontab →
241
+ the instance runs nothing while EB reads healthy, and the `@`-suppressed `mysqli_connect` in
242
+ `ebs/cron.worker.php` logs nothing. Never treat "EB is green" as evidence the fleet is working;
243
+ check `Vision_Log.Workers` against the actual EB instance list.
211
244
  - **`dbchanges` auto-apply is branch-gated, and `_production` is NOT auto-applied.**
212
245
  `crons/infrastructure/execute_dbchanges.php` runs every 2 minutes but only for branches matching
213
246
  `_%`, deriving the env by stripping the leading `_` and requiring `config.<env>.ini` in the
@@ -6,12 +6,17 @@ project: Worker
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-07-13
10
- owners: ["jcardinal"]
9
+ updated: 2026-08-24
10
+ owners: ["jcardinal", "bala"]
11
11
  files:
12
12
  - library/app/worker.php
13
13
  - worker/crons/worker/worker_heartbeat.php
14
- related: ["oneuptime-server-monitor-host-hygiene"]
14
+ - worker/ebs/cron.worker.php
15
+ - worker/.ebextensions/045_oneuptime_agent.config
16
+ related:
17
+ - ./oneuptime-server-monitor-host-hygiene.md
18
+ - ../architecture.md
19
+ - ../../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md
15
20
  ---
16
21
 
17
22
  ## Summary
@@ -57,6 +62,31 @@ Database, Infrastructure, TOGa, TOGa Desk, and Catalog, then importing each into
57
62
  manually. To add a new worker's monitor, clone an existing monitor export the same way and
58
63
  register the resulting heartbeat URL in `App_Worker::$oneUptimeEndpoints`.
59
64
 
65
+ ## THE BLIND SPOT — every monitor here is keyed by ROLE, so a ROLELESS instance is invisible
66
+
67
+ Load-bearing limitation, learned the hard way (2026-08-24: **2 of the 7 production workers ran
68
+ roleless for ~3 hours undetected**). This layer, and the Server monitors in
69
+ [host hygiene](./oneuptime-server-monitor-host-hygiene.md), are keyed by `workerName` — **14
70
+ monitors, none keyed by instance**:
71
+
72
+ - `worker_heartbeat.php` only pings `App_Worker::$oneUptimeEndpoints[$row['workerName']]` for
73
+ the row matching **its own** `instanceId`. No row → no ping target → it pushes nothing.
74
+ - The `045_oneuptime_agent.config` Server agent reads `/etc/worker-role` and **exits clean when
75
+ it is empty.**
76
+
77
+ So an instance that never claimed a role reports to nothing — and because every role is still
78
+ owned by *someone*, **no role-keyed monitor is missing a ping either.** The fleet reads 100%
79
+ healthy while a box does zero work (role assignment happens **only at deploy time**, and no role
80
+ means the role's `cron.worker.<role>.json` is never appended to the crontab, so the instance runs
81
+ **nothing**).
82
+
83
+ `Vision_Log.Workers` only ever holds instances that **succeeded** in claiming a role, so the
84
+ failure leaves no row and no log line — **absence is the only evidence**, and detecting an
85
+ absence requires the AWS instance list, which nothing on this tier consults. That gap is what
86
+ the 2.0
87
+ [Worker Fleet Role-Assignment Monitor](../../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md)
88
+ now closes from outside the fleet. **Do not try to close it with another role-keyed monitor.**
89
+
60
90
  ## Gotchas
61
91
 
62
92
  - Do **not** hardcode the heartbeat URLs/UUIDs anywhere but `App_Worker::$oneUptimeEndpoints`
@@ -67,6 +97,13 @@ register the resulting heartbeat URL in `App_Worker::$oneUptimeEndpoints`.
67
97
  pending cleanup — the live source of truth is the `App_Worker::$oneUptimeEndpoints` map.
68
98
 
69
99
  ## Change history
100
+ - 2026-08-24 — Documented **the blind spot**: all 14 monitors on this tier are keyed by
101
+ `workerName`, so an instance that claimed **no** role pings nothing and leaves every role-keyed
102
+ monitor still green — which is how 2 of 7 production workers ran roleless for ~3 hours
103
+ undetected. Recorded that `Vision_Log.Workers` only holds successful claims (absence is the
104
+ only evidence) and that the gap is now covered from outside the fleet by the 2.0
105
+ [Worker Fleet Role-Assignment Monitor](../../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md).
106
+ (bala)
70
107
  - 2026-07-13 — Added external OneUptime heartbeat push for all 1.0 workers on top of the
71
108
  existing internal DB-heartbeat mechanism; endpoints keyed by workerName in
72
109
  `App_Worker::$oneUptimeEndpoints`. (jcardinal)
@@ -0,0 +1,124 @@
1
+ ---
2
+ title: Tracing a 1.0 worker cron run in production (no stdout, silent skips)
3
+ framework: "1.0"
4
+ repo: worker
5
+ project: Worker
6
+ client: shared
7
+ type: workflow
8
+ status: active
9
+ updated: 2026-08-25
10
+ owners: ["bala"]
11
+ files:
12
+ - worker/ebs/cron.worker.php
13
+ - worker/schedules/cron.worker.sync.json
14
+ - library/app/framework.php
15
+ related:
16
+ - ../architecture.md
17
+ - ./diagnosing-frozen-cron-checkins.md
18
+ - ../../library/features/cron-execution-monitoring.md
19
+ - ../../../1.0/standards/backend-php.md
20
+ ---
21
+
22
+ ## Summary
23
+ How to answer *"did this cron actually run, and what did it do?"* on the 1.0 `worker` tier, where
24
+ **there is no usable stdout** and **a skipped run leaves no trace at all**. Written from a
25
+ production diagnostic session on
26
+ `worker/crons/toga2/compass/workflow/3a_import_office_depot_purchase_orders.php`; the mechanics
27
+ apply to every cron on the tier.
28
+
29
+ Use this before you conclude a cron "never fired" — three different situations look identical
30
+ from the outside.
31
+
32
+ ## Step 1 — get the real schedule from `schedules/*.json`, never from the file
33
+ The `.php` file's own header comment is not authoritative and is routinely stale. So are skills,
34
+ docs and tickets. `worker/schedules/cron.<env>[.<role>].json` is the only source of truth.
35
+
36
+ Worked example: the entry *"Download EDI From S3 Files & Create PO - EDI Office Depot"* in
37
+ `cron.worker.sync.json` is `*/5 * * * *` — **every 5 minutes**. The cron file's header said
38
+ "EVERY HOUR" (corrected 2026-08-25) and at least one skill still says hourly. A wrong assumed
39
+ frequency makes you read a normal gap as an outage.
40
+
41
+ Also check whether the job is registered in the env you are testing in at all: this one is **not**
42
+ in `cron.beta.json`, so beta never runs it on a schedule. Nothing on beta is evidence about it.
43
+
44
+ ## Step 2 — echo/print goes nowhere useful, so do not debug with it
45
+ The crontab line is assembled in `worker/ebs/cron.worker.php`, and the two blocks that build it
46
+ differ:
47
+
48
+ | Schedule file | Command written | Where output goes |
49
+ |---|---|---|
50
+ | base `cron.<env>.json` | `php /var/www/html/crons/<script>` | **nowhere** — the ` > /var/www/cache/WORKER_ERROR_$RANDOM` redirect is **commented out** |
51
+ | role `cron.worker.<role>.json` | `php /var/www/html/crons/<script> > /var/www/cache/WORKER_ERROR_$RANDOM` | a **randomly named** file under `/var/www/cache/` on whichever of the 7 instances ran it |
52
+
53
+ Either way you cannot go and read it: base-schedule output is discarded outright, and a
54
+ role-schedule run leaves an unpredictable filename on a box in an autoscaled fleet. **Treat
55
+ worker cron stdout as write-only.**
56
+
57
+ ## Step 3 — put temporary tracing somewhere you can query
58
+ What worked: write timestamped step markers into the **`note`** column of the run's
59
+ `Common.CronJobExecutions` row (the row `App_Framework::cronInitialization()` already inserted),
60
+ then read them back over the DB tooling from your desk. `dtCheckIn`, `dtCheckOut` and
61
+ `executionTimeSeconds` on the same row give you the shape of the run for free.
62
+
63
+ Rules: keep it to a few markers, and **remove the tracing once diagnosed** (it was removed in
64
+ this session). Never write payloads or credentials into `note`.
65
+
66
+ ## Step 4 — read `Common.CronJobExecutions` to tell "ran", "hung" and "skipped" apart
67
+ Query the legacy env's `Common.CronJobExecutions` filtering `job LIKE` the script path:
68
+
69
+ - **`dtCheckIn` + `dtCheckOut` + `executionTimeSeconds`** — it ran and finished; the duration
70
+ tells you whether it fits inside its schedule slot.
71
+ - **`dtCheckIn` with no `dtCheckOut`** — it started and never completed (crash, timeout, or the
72
+ instance went away).
73
+ - **no row for a slot you expected** — the run was **skipped by the overlap guard**, not "never
74
+ fired". A blocked run writes **nothing**: `cronInitialization()` calls
75
+ `exitIfProcessRunning()` **before** the INSERT, and the `cronFinished(true)` it then calls is a
76
+ no-op because `cronLogId` is unset. See
77
+ [Cron Execution Monitoring](../../library/features/cron-execution-monitoring.md).
78
+
79
+ ### The overlap guard is a `ps -ef` substring match — anything can block it
80
+ `App_Framework::isProcessRunning()` returns true for **any** `ps -ef` line containing the script
81
+ path (it only excludes `/bin/sh` lines and its own pid). So a developer's `tail -f`, an editor, or
82
+ any shell holding that path in its command line **blocks the cron indefinitely** while looking
83
+ like a legitimate overlap. Check `ps -ef` for what is actually holding the name before assuming a
84
+ long-running instance of the job itself.
85
+
86
+ ## Step 5 — if runs overlap, look at what the job does before its real work
87
+ A job whose own runtime exceeds its schedule interval silently loses most of its slots to the
88
+ guard. Measure the phases, not the total: in the ODP case an instrumented run showed the S3
89
+ listing alone returning **138,472 objects and taking 32 seconds** before any PO was touched, with
90
+ each PO then costing roughly 40 seconds — comfortably past a 5-minute schedule. Prefixes the job
91
+ skips while processing are still listed, so accumulated `SENT/` and `OUTBOX/` objects inflate
92
+ every run.
93
+
94
+ ## Deploy-time trap: a PHP 8-only syntax kills the whole cron with no visible error
95
+ This tier runs **PHP 7.2**, so a **PHP 8 named argument** (`someFunction(items: $x)`) is a
96
+ **parse error**, not a runtime warning:
97
+
98
+ ```
99
+ PHP Parse error: syntax error, unexpected ':', expecting ',' or ')'
100
+ ```
101
+
102
+ The script then does nothing at all on every tick — and per Steps 2 and 4 you will see no output
103
+ and no `CronJobExecutions` row, i.e. it looks exactly like "the cron never fired". Typed
104
+ parameters and return types (`: void`, `: array`, `: bool`, `?array`, `?object`, `object $x`) are
105
+ PHP 7.0/7.1 and **are** safe here; they are already used in
106
+ `worker/crons/toga2/compass/backfill_items_isfulfillable.php` and
107
+ `update_item_fulfillments_in_netsuite.php`.
108
+
109
+ `1.0/standards/backend-php.md` already forbids named arguments in 1.0 — but a workspace
110
+ `CLAUDE.md` / `.github/copilot-instructions.md` that says *"use PHP 8 Named Parameters"* applies
111
+ to **2.0 only**. The 1.0 standard wins in `library/`, `worker/` and every other 1.0 app.
112
+
113
+ **Do not read a PHP 8 function call in existing code as proof the tier is PHP 8.**
114
+ `library/app/systemmonitor/500error.php` uses `str_contains()`; that file is not a version
115
+ signal, it is a latent bug.
116
+
117
+ ## Change history
118
+ - 2026-08-25 — Documented from a prod diagnostic session on the Compass ODP 850 importer: the
119
+ schedule JSON (not the file header) is authoritative, worker cron stdout is unreadable (base
120
+ block's `WORKER_ERROR` redirect commented out, role block writes a random filename),
121
+ `Common.CronJobExecutions.note` is the practical place for temporary tracing, a run skipped by
122
+ the `ps -ef` overlap guard writes **no row at all** (guard runs before the INSERT), the guard
123
+ matches any process holding the script path, and a PHP 8 named argument parse-errors the whole
124
+ cron silently on this PHP 7.2 tier. (bala)
@@ -6,8 +6,8 @@ project: _Underscore
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-08-17
10
- owners: ["dfranks", "jcardinal", "mhammontree", "apeterson", "kyalamarthi"]
9
+ updated: 2026-08-24
10
+ owners: ["dfranks", "jcardinal", "mhammontree", "apeterson", "kyalamarthi", "bala"]
11
11
  files:
12
12
  - _underscore/Database.php
13
13
  - _underscore/Model.php
@@ -177,6 +177,16 @@ here — they live in `Config/*.ini`.)
177
177
  - Related 1.0 analogue: the legacy `App_` worker has the same hazard writing to `Logs.API`
178
178
  (`db_logs`) — the laptop trap there is documented separately in the worker NetSuite bootstrap
179
179
  notes.
180
+ - **`_Database::register()` is LAZY — it stores connection config and never opens a socket.**
181
+ The connect happens on first use. So adding a static alias registration to an app's `_.php`
182
+ bootstrap is **safe across every environment**, even one whose config group points at a
183
+ localhost that lacks the schema: nothing fails until something actually queries that alias.
184
+ This is what makes "register the alias globally, use it in one cron" a two-line change rather
185
+ than a per-environment config exercise (worked example: `DB_VISION_LOGS` in
186
+ [the Worker Fleet Role-Assignment Monitor](../../worker2/features/worker-fleet-role-assignment-monitor.md)).
187
+ The corollary is the trap the rest of this doc describes: because registration is free and
188
+ silent, a **bad** registration also stays silent until the request that needs it blows up with
189
+ `Unknown database`.
180
190
 
181
191
  ## WITHDRAWN — the "alias-keyed `$_modelCache` cross-tenant leak" hypothesis (2026-08-17)
182
192
 
@@ -210,6 +220,11 @@ also hits `_modules/<module>/`** — is recorded in
210
220
 
211
221
  ## Change history
212
222
 
223
+ - 2026-08-24 — Recorded that **`_Database::register()` is lazy** (stores config, opens no
224
+ socket; the connect happens on first use), so a static alias registration in an app's `_.php`
225
+ is safe in every environment including a localhost config that lacks the schema — and that
226
+ this is also why a *bad* registration stays silent until first query. Verified while adding
227
+ `DB_VISION_LOGS` to worker2. (bala)
213
228
  - 2026-08-17 (later pass) — **WITHDRAWN, supersedes the entry below.** The alias-keyed
214
229
  `$_modelCache` cross-tenant-leak hypothesis is **not supported** and is no longer a live security
215
230
  concern. The six `c_` columns are **AIG's own**: `_modules/netsuite/2026-07-10a -
@@ -17,7 +17,7 @@
17
17
  | [ClickUp Work Type Automation (Committed / Conditional / Stretch)](features/clickup-work-type-automation.md) | The ClickUp webhook handler (`_Worker_Clickup`) automatically maintains each task's **Work Type** custom field — `Committed`, `Conditional`, or `Stretch` — base | worker2/Worker/Clickup.php, worker2/Tests/Worker/ClickupWorkTypeTest.php |
18
18
  | [Compass VIP Support Importer (worker2)](features/compass-vip-support-importer.md) | A worker2 action that ingests Compass's quarterly VIP spreadsheet and assigns each VIP user's support technician by setting `Users.c_supportedByUserId` in `Clie | worker2/Worker/Client/Compass/VipSupport.php |
19
19
  | [Creating Worker Actions](features/creating-worker-actions.md) | How to add a new callable Worker action — a PHP class whose `public static` methods are invoked as background jobs (via webhook, cron, or `_Worker::runTask()`). | worker2/Worker/, worker2/Controller/Index.php, _underscore/Worker.php |
20
- | [Cross-account AWS access for worker2 crons (_Component_Aws_Workloads)](features/cross-account-aws-access.md) | `_Component_Aws_Workloads` is the **standard, and only sanctioned, way any new worker2 cron obtains AWS access** — for any account, any region, any AWS SDK clie | worker2/Component/Aws/Workloads/Workloads.php, worker2/Config/production.ini, worker2/Worker/Infrastructure/CloudWatch.php |
20
+ | [Cross-account AWS access for worker2 crons (_Component_Aws_Workloads)](features/cross-account-aws-access.md) | `_Component_Aws_Workloads` is the **standard, and only sanctioned, way any new worker2 cron obtains AWS access** — for any account, any region, any AWS SDK clie | worker2/Component/Aws/Workloads/Workloads.php, worker2/Config/production.ini, worker2/Worker/Infrastructure/CloudWatch.php, worker2/Worker/Monitor/Fleet.php, library/app/worker.php, library/app/cloud.php |
21
21
  | [Elastic Beanstalk health monitor → OneUptime (ElasticBeanstalkHealth)](features/elastic-beanstalk-health-monitor.md) | `_Worker_Infrastructure_CloudWatch::ElasticBeanstalkHealth()` is a worker2 cron that reads each Elastic Beanstalk (EB) environment's **enhanced-health** status | worker2/Worker/Infrastructure/CloudWatch.php |
22
22
  | [Elite Freshservice Sync (worker2)](features/elite-freshservice-sync.md) | `_Worker_Elite` processes Freshservice webhook events and syncs them into TOGA 2. | worker2/Worker/Elite.php, worker2/Config/dev-kmaramreddy-laptop.ini |
23
23
  | [Error Escalation Cron (Errors::Escalate → ClickUp / email)](features/error-escalation-cron.md) | `_Worker_Infrastructure_Errors::Escalate` (renamed from `SyncWithClickup`) is the sole owner of **escalation, de-escalation, ClickUp ticketing, reminders, busin | worker2/Worker/Infrastructure/Errors.php, worker2/Worker/Notification/Email.php, worker2/Worker/Notification/EmailTemplate.php, worker2/Worker/Client/True.php, worker2/Worker/Clickup/ErrorTask.php, worker2/Worker/Clickup.php, worker2/Controller/Index.php, worker2/Config/production.ini, _underscore/Model/Core/Logs/Issue.php, dbchanges2/Core/2026-07-30a - Error escalation cron job.sql, dbchanges2/Logs/2026-08-03a - Issue clickupPriority.sql, dbchanges2/Core/2026-08-04a - Error neglect digest cron job.sql |
@@ -47,6 +47,7 @@
47
47
  | [Centralized Tracking-Status Refresh (worker2 Sync cron, all clients)](features/tracking-status-refresh.md) | A single platform-wide cron keeps `TrackingNumbers.status` current for **every provisioned client** until each shipment reaches a terminal state, by polling Fed | worker2/Worker/Sync/TrackingNumbers.php, _underscore/Component/Library/Carriers/Fedex/Fedex.php, _underscore/Component/Library/Carriers/Usps/Usps.php, _underscore/Component/Library/Carriers/Ups/Ups.php, _underscore/Model/Client/TrackingNumber.php, dbchanges2/Client/2026-07-30 - TrackingNumbers_RefreshColumns.sql, dbchanges2/Client/2026-07-30 - Acl_TrackingStatus_Grant.sql, dbchanges2/Client_Tdsynnex/2026-07-30 - Apis_TdsynnexKey.sql, dbchanges2/Core/2026-07-30 - TrackingRefresh_Cron.sql |
48
48
  | [VAPI Webhook Handler (worker2 — AI-BDR end-of-call processing)](features/vapi-webhook-handler.md) | `_Worker_Vapi` ([worker2/Worker/Vapi.php](worker2/Worker/Vapi.php)) is the **PHP side of the AI-BDR call loop** — the webhook that receives VAPI's end-of-call r | worker2/Worker/Vapi.php, worker2/Worker/Ai/Bdr/Vapi.php, worker2/Controller/Index.php |
49
49
  | [WJE Freshservice Sync (worker2)](features/wje-freshservice-sync.md) | WJE ("WJE IT", helpdesk `wje.freshservice.com`) is a **Freshservice**-based help-desk client whose tickets, contacts, assets, groups, categories, and canned res | worker2/Worker/Wje.php, _underscore/Component/Api/Wje/Wje.php, _underscore/Model/Wje/Ticket.php, _underscore/Model/Wje/TicketNote.php, _underscore/Model/Wje/Contact.php, _underscore/Model/Wje/Unit.php, _underscore/Model/Wje/TicketTeam.php, _underscore/Model/Wje/TicketCategory.php, _underscore/Model/Wje/AssetType.php, _underscore/Model/Wje/PredefinedReply.php, library/app/api/wje.php, worker/crons/toga2/wje/import_supporting_records.php, worker/crons/toga2/wje/sync_togasupply_wje.php, worker/crons/notifications/reports/wje/wje_common.php, library/app/systemmonitor/wje.php, dbchanges2/Client_Wje/2024-10-04 - WjeOnboarding.sql |
50
+ | [1.0 Worker Fleet Role-Assignment Monitor (Monitor/Fleet/RoleAssignment)](features/worker-fleet-role-assignment-monitor.md) | `_Worker_Monitor_Fleet::RoleAssignment()` is a **2.0 worker2 cron that watches the 1.0 `worker` fleet from the outside**. | worker2/Worker/Monitor/Fleet.php, worker2/_.php, dbchanges2/Core/2026-08-24 - Worker Role Assignment Monitor.sql, worker/ebs/cron.worker.php, worker/crons/worker/worker_heartbeat.php, library/app/worker.php |
50
51
  | [PHP Runtime Upgrade on Elastic Beanstalk (worker2 8.3 → 8.5 + PhpSpreadsheet 1.x → 3.x)](workflows/php-runtime-upgrade-dependency-audit.md) | The procedure used to move worker2 from **PHP 8.3 to PHP 8.5** on Elastic Beanstalk, and the dependency work that had to land first. | worker2/composer.json, worker2/composer.lock, worker2/Worker/Team/Sprint.php, worker2/Worker/Client/TowFoundation/ProcessReceipts.php, worker2/Worker/Forecast/Import.php |
51
52
  | [Running worker2 locally against real NetSuite (and what EB does instead)](workflows/running-worker2-locally.md) | How to boot **worker2** on a developer machine and run a real worker action against the **live NetSuite** account and a **local** `Forecast` database. | worker2/index.php, worker2/composer.json, worker2/.ebextensions/php_include_underscore.config, worker2/.ebextensions/git.json, worker2/.ebextensions/git.php, worker2/Config/production.ini, _underscore/Component/Api/Netsuite/Netsuite.php, dbchanges2/Logs_Client/2026-04-08_BLANK_CLIENT_LOGS_DATABASE.SQL |
52
53
  | [Ticket → ClickUp Pseudocode Planning (Talos-grounded)](workflows/ticket-to-pseudocode-planning.md) | A repeatable procedure for turning a ClickUp ticket into a reviewed, formatted implementation plan posted back to the ticket's `📝 Pseudocode` custom field. | test/@dave/clickup_md2delta.js |
@@ -6,18 +6,22 @@ project: Worker
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-08-11
10
- owners: [jcardinal]
9
+ updated: 2026-08-24
10
+ owners: [jcardinal, bala]
11
11
  files:
12
12
  - worker2/Component/Aws/Workloads/Workloads.php
13
13
  - worker2/Config/production.ini
14
14
  - worker2/Worker/Infrastructure/CloudWatch.php
15
+ - worker2/Worker/Monitor/Fleet.php
16
+ - library/app/worker.php
17
+ - library/app/cloud.php
15
18
  related:
16
19
  - ../../_underscore/features/component-model-namespace-registration.md
17
20
  - ../../_underscore/features/config-group-access.md
18
21
  - ../../_underscore/features/cloud-s3-helpers.md
19
22
  - ./oneuptime-worker2-monitoring.md
20
23
  - ./elastic-beanstalk-health-monitor.md
24
+ - ./worker-fleet-role-assignment-monitor.md
21
25
  ---
22
26
 
23
27
  ## Summary
@@ -79,7 +83,8 @@ region is only a runtime argument to `client()`, never a reason for a second rol
79
83
  - **654654170868** — production hub. Home region `us-east-1`; also used in `us-west-2` and
80
84
  `eu-west-1`. worker2 runs here on Elastic Beanstalk; its instance profile role
81
85
  `ElasticBeanstalk-EC2-Instance-Profile` is the caller identity.
82
- - **502614707982** — legacy account, `us-west-2`.
86
+ - **502614707982** — legacy account, `us-west-2`. **This is where the 1.0 `worker` fleet
87
+ actually runs** — see the account-confusion gotcha below.
83
88
  - **975050298201** — non-production account, `us-east-1`.
84
89
 
85
90
  Each `WorkloadsRuntime` role has broad service permissions (`elasticbeanstalk:*`, `ec2:*`,
@@ -144,8 +149,33 @@ code yet. New code must not adopt the static key.
144
149
  **Never record the access-key values in knowledge.** They exist in `production.ini`; that is
145
150
  their only home.
146
151
 
152
+ ### Which account the 1.0 worker fleet lives in (easy and expensive to get wrong)
153
+
154
+ The 1.0 `worker` EC2 instances are in **502614707982**, `us-west-2` (EB environment
155
+ `agilant-worker`, confirmed in the EC2 console).
156
+
157
+ **`App_Worker::AWS_WORKER_QUEUE_URL` in `library/app/worker.php` points at 654654170868 and
158
+ misleads** — that account holds only the **SQS queue the 1.0 workers read from**, not the
159
+ instances. Anything that needs to *see* the worker boxes (EB, EC2, health) must target
160
+ **502614707982**.
161
+
162
+ worker 1.0 authenticates there with a **long-lived hardcoded IAM access key in
163
+ `library/app/cloud.php`** — a committed credential that should be rotated and moved to a role
164
+ (value deliberately not recorded here; see [Legacy static key](#legacy-static-key--being-retired)).
165
+ worker2 does **not** use it: it assumes `WorkloadsRuntime` cross-account from its EB instance
166
+ role, which is exactly the migration this component exists to enable. First worker2 consumer in
167
+ that account is the
168
+ [Worker Fleet Role-Assignment Monitor](./worker-fleet-role-assignment-monitor.md).
169
+
170
+ **Scope EB reads by `EnvironmentName`, never account-wide.** 502614707982 also runs
171
+ `agilant-worker-alpha` and `agilant-worker-beta` as **separate EB environments that have no role
172
+ registry at all**. An account-level EC2 scan would sweep them in and misreport them.
173
+
147
174
  ## Gotchas / known issues
148
175
 
176
+ - **Do not assume the 1.0 worker fleet is in the 868 production hub.** `AWS_WORKER_QUEUE_URL`
177
+ (654654170868) is the SQS queue only; the instances are in the legacy account 502614707982,
178
+ `us-west-2`. See [the section above](#which-account-the-10-worker-fleet-lives-in-easy-and-expensive-to-get-wrong).
149
179
  - Do not construct `StsClient` or pass explicit credentials by hand in new crons — the whole
150
180
  point is that the instance role is the caller and the component owns the assume-role step.
151
181
  - The instance-role-as-caller only works **on the EB tier**. Off-tier (e.g. a local CLI run)
@@ -161,6 +191,15 @@ their only home.
161
191
  see [creating-worker-actions.md](./creating-worker-actions.md).
162
192
 
163
193
  ## Change history
194
+ - 2026-08-24 — Pinned down **which account the 1.0 `worker` fleet actually runs in**:
195
+ **502614707982** (`us-west-2`, EB env `agilant-worker`), **not** the 654654170868 production hub
196
+ that `App_Worker::AWS_WORKER_QUEUE_URL` implies — that account holds only the SQS queue the 1.0
197
+ workers read from. Recorded that worker 1.0 reaches it with the hardcoded static key in
198
+ `library/app/cloud.php` (rotate + move to a role) while worker2 assumes `WorkloadsRuntime` from
199
+ its instance role, and that EB reads there must be scoped by `EnvironmentName` because
200
+ `agilant-worker-alpha`/`-beta` are separate registry-less environments in the same account.
201
+ Found while building the
202
+ [Worker Fleet Role-Assignment Monitor](./worker-fleet-role-assignment-monitor.md). (bala)
164
203
  - 2026-08-11 — Noted that `ElasticBeanstalkHealth`'s `$environments` must be typed `object|array`
165
204
  (not `array`): the nested JSON object arrives as `stdClass` because the dispatcher casts only
166
205
  the top level of `CronJobs.parameters`, so a strict `array` throws a `TypeError` at dispatch.
@@ -139,6 +139,35 @@ POSTed JSON body is addressed via the `requestBody` prefix — e.g.
139
139
  - Docs: https://oneuptime.com/docs/monitor/incident-alert-templating and
140
140
  https://oneuptime.com/docs/en/monitor/incoming-request-monitor
141
141
 
142
+ ### The monitor IMPORT JSON — exact shape (verified 2026-08-24)
143
+
144
+ Monitors are provisioned by importing a **resource-export JSON**. The importer is strict and
145
+ its error messages are the only documentation, so the verified shape:
146
+
147
+ ```
148
+ { fileType: "oneuptime-resource-export", schemaVersion: 1, resourceType: "Monitor", items: [ … ] }
149
+ ```
150
+
151
+ - **`monitorType` must be the UI DISPLAY name, with a space — `"Incoming Request"`, NOT
152
+ `"IncomingRequest"`.** The importer rejects the concatenated form and helpfully lists the
153
+ valid values in the error.
154
+ - Criteria nest deeply, each level carrying a `_type` tag:
155
+ `monitorSteps{_type:"MonitorSteps", value:{ monitorStepsInstanceArray:[ {_type:"MonitorStep",
156
+ value:{ monitorCriteria:{_type:"MonitorCriteria", value:{ monitorCriteriaInstanceArray:[ … ] }}}} ]}}`
157
+ - Each **criterion instance** carries `monitorStatusId`, `filterCondition` (`All`/`Any`),
158
+ `filters[]`, `incidents[]`, `isEnabled`, `name`, `description`.
159
+ - **Timing filters use `checkOn: "Incoming Request"`** with `filterType`
160
+ `"Not Recieved In Minutes"` / `"Recieved In Minutes"` (OneUptime's own misspelling) —
161
+ **not** `checkOn: "Request Body"`, which is only for the string-token matching above.
162
+ - `incidentSeverityId` is wrapped: `{_type: "ObjectID", value: "…"}`.
163
+ - There is **no `projectId`** in the export, and **status/severity ObjectIDs are
164
+ project-specific** — clone them out of an existing monitor's export rather than inventing them.
165
+
166
+ **Practical fix for the [activation-ordering trap](#gotchas--known-issues):** instead of
167
+ importing the monitor twice (body criteria first, absence criteria after the first push), import
168
+ the absence criteria **in the same file with `isEnabled: false`** and simply flip them on once a
169
+ push has landed.
170
+
142
171
  ### Heartbeat / cron cadence timing
143
172
 
144
173
  The 10-min-Degraded / 15-min-Offline heartbeat thresholds pair with a **5-minute** cron
@@ -327,6 +356,16 @@ check for another client.
327
356
  monitors must pass the region (see [Cloud S3 helpers](../../_underscore/features/cloud-s3-helpers.md)).
328
357
 
329
358
  ## Change history
359
+ - 2026-08-24 — Documented the **monitor import-JSON schema** the runbook previously described
360
+ only conceptually: the `oneuptime-resource-export` envelope, `monitorType` must be the
361
+ spaced display name `"Incoming Request"`, the `_type`-tagged
362
+ `monitorSteps → monitorStepsInstanceArray → monitorCriteria → monitorCriteriaInstanceArray`
363
+ nesting, per-criterion fields, timing filters on `checkOn: "Incoming Request"` (not
364
+ `"Request Body"`), the `ObjectID`-wrapped `incidentSeverityId`, and that status/severity
365
+ ObjectIDs are project-specific with no `projectId` in the export. Added the practical dodge for
366
+ the activation-ordering trap: import the absence criteria with `isEnabled: false` and flip them
367
+ on after the first push, instead of importing twice. Verified while provisioning the
368
+ [Worker Fleet Role-Assignment Monitor](./worker-fleet-role-assignment-monitor.md). (bala)
330
369
  - 2026-08-20 — Added the **SFTP monitor** shape (phpseclib3 `rawlist` + age-cutoff count +
331
370
  decided token; login/list failure → `status:error` connectivity check) and its two
332
371
  runtime gotchas: `rawlist` returns **arrays** not objects, and the SFTP type constants are
@@ -0,0 +1,175 @@
1
+ ---
2
+ title: 1.0 Worker Fleet Role-Assignment Monitor (Monitor/Fleet/RoleAssignment)
3
+ framework: "2.0"
4
+ repo: worker2
5
+ project: Worker
6
+ client: shared
7
+ type: feature
8
+ status: active
9
+ updated: 2026-08-24
10
+ owners: ["bala"]
11
+ files:
12
+ - worker2/Worker/Monitor/Fleet.php
13
+ - worker2/_.php
14
+ - dbchanges2/Core/2026-08-24 - Worker Role Assignment Monitor.sql
15
+ - worker/ebs/cron.worker.php
16
+ - worker/crons/worker/worker_heartbeat.php
17
+ - library/app/worker.php
18
+ related:
19
+ - ./oneuptime-worker2-monitoring.md
20
+ - ./cross-account-aws-access.md
21
+ - ./elastic-beanstalk-health-monitor.md
22
+ - ../../../1.0/apps/worker/architecture.md
23
+ - ../../../1.0/apps/worker/features/oneuptime-worker-uptime-monitoring.md
24
+ ---
25
+
26
+ ## Summary
27
+
28
+ `_Worker_Monitor_Fleet::RoleAssignment()` is a **2.0 worker2 cron that watches the 1.0
29
+ `worker` fleet from the outside**. Every 5 minutes it lists the EC2 instances Elastic
30
+ Beanstalk reports for the `agilant-worker` environment and **compares that list against the
31
+ `Vision_Log.Workers` role registry**, then pushes three findings to a OneUptime Incoming
32
+ Request monitor:
33
+
34
+ | Finding | Meaning |
35
+ |---|---|
36
+ | `namelessInstances` | An instance is running in AWS but owns **no** row in `Workers` — it claimed no role, so it is running **no crons at all**. |
37
+ | `unownedRoles` | A role in `App_Worker::$desiredWorkers` that **nobody** currently owns. |
38
+ | `staleRoles` | A role whose `dtHeartbeat` has stopped advancing. |
39
+
40
+ **Why it exists:** 2 of the 7 production workers ran **roleless for roughly 3 hours with
41
+ nobody noticing**. Every pre-existing monitor is keyed by *role*, so an instance that never
42
+ got a role is invisible to all of them **by design** (see
43
+ [the blind spot](#the-blind-spot-why-none-of-the-14-existing-monitors-could-see-this)).
44
+ Detection is now about 10 minutes.
45
+
46
+ It follows the standard team token contract from
47
+ [OneUptime push-metric monitors](./oneuptime-worker2-monitoring.md): the worker decides,
48
+ OneUptime string-matches. `status = reporting | error`, `alarm = HIGH | OK`.
49
+
50
+ ## Key files / entry points
51
+
52
+ - `worker2/Worker/Monitor/Fleet.php` — `_Worker_Monitor_Fleet`, action
53
+ **`Monitor/Fleet/RoleAssignment`** (routing verified: `Monitor/Fleet/RoleAssignment` maps to
54
+ `_Worker_Monitor_Fleet::RoleAssignment`).
55
+ - `worker2/_.php` — adds `const DB_VISION_LOGS = 'Vision_Log'` and its
56
+ `_Database::register()` call on the **`[database1]`** connection group.
57
+ - `dbchanges2/Core/2026-08-24 - Worker Role Assignment Monitor.sql` — the `Core.CronJobs`
58
+ row: `*/5 * * * *`, `parameters` NULL, **`isActive = 0`**.
59
+ - AWS access is obtained by assuming `WorkloadsRuntime` via
60
+ `_Component_Aws_Workloads::client(ElasticBeanstalkClient::class, ...)` — see
61
+ [Cross-account AWS access](./cross-account-aws-access.md). No static keys.
62
+
63
+ All tunables are **class constants with cron-parameter overrides**, so the schedule and the
64
+ thresholds live in the `CronJobs` row rather than in a deploy.
65
+
66
+ ## How it works
67
+
68
+ 1. Assume `WorkloadsRuntime` in the **legacy** account (`502614707982`, `us-west-2`) and call
69
+ **`describeInstancesHealth`** on the EB environment **`agilant-worker`**.
70
+ 2. Read `Vision_Log.Workers` (`workerInstanceId` UNIQUE, `workerName` UNIQUE, `dtHeartbeat`)
71
+ through the new `DB_VISION_LOGS` alias.
72
+ 3. Diff the two sets in both directions, plus a heartbeat-freshness pass, producing the three
73
+ findings above.
74
+ 4. Apply the **grace period** (default 10 min): an instance that is still booting has not had
75
+ a chance to claim a role yet and must not page.
76
+ 5. POST the decided tokens to the OneUptime monitor, non-fatal.
77
+
78
+ **Fail-safe:** a failed AWS call *or* a failed registry read pushes **`status = error`**, so a
79
+ blind checker is never mistaken for a healthy fleet. This is the same rule as
80
+ [the EB health monitor](./elastic-beanstalk-health-monitor.md) — a read that could not happen
81
+ is never reported as "fine."
82
+
83
+ ### Why `Vision_Log` needed no new plumbing
84
+
85
+ worker2's **`[database1]` group already points at the legacy production cluster**
86
+ (cluster id `clwbyqvxdm4q`, `us-west-2`) — it is what already backs the `TOGaDeskSupport`
87
+ alias. `Vision_Log` lives on that **same single all-in-one legacy cluster** (verified: `Core`,
88
+ `TOGaDeskSupport`, `Vision` and `Vision_Log` are all present on it). So the whole database
89
+ change was **two lines in `_.php`**: no new credentials, no security-group change, no
90
+ `Core.Databases` rows.
91
+
92
+ Two facts that make that safe, both verified this session:
93
+
94
+ - All **9** `worker2/Config/*.ini` files define `[database1]`, so the registration resolves in
95
+ every environment.
96
+ - **`_Database::register()` is lazy** — it stores connection config and never opens a socket —
97
+ so a static registration in `_.php` costs nothing and cannot break a dev config whose
98
+ `[database1]` points at localhost.
99
+
100
+ ## The blind spot (why none of the 14 existing monitors could see this)
101
+
102
+ This is the durable "why" the incident exposed, and the reason this monitor had to reach for
103
+ AWS rather than the database alone.
104
+
105
+ - 1.0 role assignment happens in `worker/ebs/cron.worker.php` and **only at DEPLOY time** (it
106
+ is written as the EB `appdeploy` enact hook). The script claims a free name, writes
107
+ `/etc/worker-role`, and **only then** appends that role's `cron.worker.<name>.json` entries.
108
+ - Therefore **no name means no crontab**. A roleless instance runs *nothing*, while EB happily
109
+ reports the environment healthy — the box is up, it just does no work.
110
+ - **All 14 existing OneUptime monitors are keyed by ROLE, not by instance.**
111
+ `crons/worker/worker_heartbeat.php` only pings the endpoint for the row matching **its own**
112
+ `instanceId`, and the `045_oneuptime_agent.config` infrastructure agent reads
113
+ `/etc/worker-role` and **exits clean when it is empty**. A roleless instance therefore
114
+ reports to nothing, and no role-keyed monitor is missing a ping.
115
+ - `Vision_Log.Workers` **only ever holds instances that SUCCEEDED** in claiming a role. The
116
+ failure leaves no row and no log line, so **absence is the only evidence** — and seeing an
117
+ absence requires the AWS instance list. That is precisely why this monitor exists in this
118
+ shape.
119
+
120
+ ## Design history — built in worker2, not in worker 1.0
121
+
122
+ It was **first built in 1.0** (`worker` owns the `Vision_Log` connection natively), then moved
123
+ to worker2 and **the 1.0 version was deleted**. Recorded so it is not resurrected:
124
+
125
+ - worker2 already reaches the same cluster (above), so the 1.0 home bought nothing.
126
+ - worker2 uses the **EB instance role**; 1.0 would have used the hardcoded static AWS key.
127
+ - Schedule and parameters live in a **`CronJobs` row**, not in a deploy.
128
+ - Most importantly, it watches the 1.0 fleet **from outside it** instead of from a member of it.
129
+
130
+ The obvious objection — "a monitor running on the fleet it watches has a blind spot" — is
131
+ covered either way by OneUptime's own **absence criteria** (10 min Degraded / 15 min Offline),
132
+ which catch the monitor itself dying. The other three reasons decided it.
133
+
134
+ ## Verification performed at build time
135
+
136
+ - **24/24 logic tests pass** against the real `Fleet.php` (private methods driven by
137
+ reflection): today's healthy fleet, a **replay of the actual incident**, the grace period,
138
+ stale heartbeats, a NULL-name spare row, a ghost row for a terminated instance, and an empty
139
+ registry.
140
+ - `Core.CronJobs` columns, uuid uniqueness, and the absence of a clashing `Monitor/Fleet%`
141
+ action all confirmed.
142
+ - **Unverified at capture time:** whether `WorkloadsRuntime` exists in `502614707982` with
143
+ `elasticbeanstalk:DescribeInstancesHealth`. If it does not, the monitor pushes
144
+ `status = error` — loud, not silent.
145
+
146
+ ## Gotchas / known issues
147
+
148
+ - **The `CronJobs` row ships `isActive = 0` and must be flipped on after worker2 deploys.**
149
+ Follow the activation ordering in
150
+ [OneUptime push-metric monitors](./oneuptime-worker2-monitoring.md) — confirm a push has
151
+ landed before enabling the absence criteria.
152
+ - **Never widen the AWS scope to an account-level EC2 scan.** The same account also runs
153
+ `agilant-worker-alpha` and `agilant-worker-beta` as **separate EB environments with no role
154
+ registry at all**. `describeInstancesHealth` is scoped by `EnvironmentName`, which correctly
155
+ excludes them; an account-wide scan would flag every alpha/beta box as roleless.
156
+ - **Known false-positive risk, deliberately NOT guarded yet.**
157
+ `worker_heartbeat.php` terminates a stale instance and **then** deletes its `Workers` row. If
158
+ EB still lists that instance in `describeInstancesHealth` for a minute or two after the row is
159
+ gone, this monitor sees an instance with no role and raises `alarm = HIGH` — a possible false
160
+ page during a deploy or a recycle. It was not pre-solved because **how long EB keeps a
161
+ terminating instance in the health list is unverified**. If it shows up in practice, the fix is
162
+ a one-line filter on the instance's health status.
163
+ - **The OneUptime push URL is a credential** — never log it, never record its value.
164
+ - The class lives under `Worker/Monitor/` (singular), matching the OneUptime push monitors —
165
+ not the `Worker/Monitors/` (plural) orchestrator framework.
166
+
167
+ ## Change history
168
+ - 2026-08-24 — Created. Added `_Worker_Monitor_Fleet::RoleAssignment()`
169
+ (`Monitor/Fleet/RoleAssignment`, `*/5 * * * *`, `isActive = 0` pending deploy) reporting
170
+ `namelessInstances` / `unownedRoles` / `staleRoles` to OneUptime, after 2 of 7 production
171
+ workers ran roleless for ~3 hours undetected. Registered `DB_VISION_LOGS` on the existing
172
+ `[database1]` legacy-cluster group (2 lines, no new credentials). Recorded the role-keyed
173
+ monitoring blind spot that made the incident invisible, the decision to build in worker2
174
+ rather than 1.0 (and delete the 1.0 version), and the unguarded terminating-instance
175
+ false-positive risk. worker2 `_production` e75690f, dbchanges2 `_main` 17732e2. (bala)
@@ -5,7 +5,7 @@ _Auto-generated by `knowledge.js index`. Do not hand-edit._
5
5
  ## 1.0 framework
6
6
 
7
7
  - **library** (Library) _(framework core)_ — 19 doc(s) → [1.0/apps/library/INDEX.md](1.0/apps/library/INDEX.md)
8
- - **worker** (Worker) — 26 doc(s) → [1.0/apps/worker/INDEX.md](1.0/apps/worker/INDEX.md)
8
+ - **worker** (Worker) — 28 doc(s) → [1.0/apps/worker/INDEX.md](1.0/apps/worker/INDEX.md)
9
9
  - **dbchanges** (Database Changes) _(framework core)_ — 1 doc(s) → [1.0/apps/dbchanges/INDEX.md](1.0/apps/dbchanges/INDEX.md)
10
10
  - **worker1.5** (Worker 1.5) — 0 doc(s) → [1.0/apps/worker1.5/INDEX.md](1.0/apps/worker1.5/INDEX.md)
11
11
  - **togadesk** (TOGa Desk) — 13 doc(s) → [1.0/apps/togadesk/INDEX.md](1.0/apps/togadesk/INDEX.md)
@@ -19,7 +19,7 @@ _Auto-generated by `knowledge.js index`. Do not hand-edit._
19
19
  ## 2.0 framework
20
20
 
21
21
  - **_underscore** (_Underscore) _(framework core)_ — 62 doc(s) → [2.0/apps/_underscore/INDEX.md](2.0/apps/_underscore/INDEX.md)
22
- - **worker2** (Worker) — 55 doc(s) → [2.0/apps/worker2/INDEX.md](2.0/apps/worker2/INDEX.md)
22
+ - **worker2** (Worker) — 56 doc(s) → [2.0/apps/worker2/INDEX.md](2.0/apps/worker2/INDEX.md)
23
23
  - **api2** (API) — 25 doc(s) → [2.0/apps/api2/INDEX.md](2.0/apps/api2/INDEX.md)
24
24
  - **dbchanges2** (Database Changes) _(framework core)_ — 8 doc(s) → [2.0/apps/dbchanges2/INDEX.md](2.0/apps/dbchanges2/INDEX.md)
25
25
  - **toga2-supply** (TOGa Supply) — 7 doc(s) → [2.0/apps/toga2-supply/INDEX.md](2.0/apps/toga2-supply/INDEX.md)
@@ -12,6 +12,7 @@
12
12
  | [Compass MITS PO rejection — zero line items (400 + Issue + PM alert email)](features/mits-po-zero-line-item-rejection.md) | 2.0 | MITS transmits Compass purchase orders that carry **`"purchaseOrderItems": []`** (real example: PO `50309196-1`, `mitsSalesOrder` `MR244630`). | _underscore/Model/Compass/PurchaseOrder.php, _underscore/Model/Compass/Usa/PurchaseOrder.php, _underscore/Model/Compass/Canada/PurchaseOrder.php, worker2/Worker/Client/Compass/reports/PurchaseOrderRejectionEmail.php |
13
13
  | [Compass MITS Sales-Order Transmission — Rejection Alerting (Issue/Event, not email-in-cron)](features/mits-sales-order-transmission-alerting.md) | 2.0 | When MITS **rejects** a Compass sales order transmitted by the 1.0 cron `1_transmit_compass_sales_orders_to_mits.php`, the alert is no longer an email built ins | worker/crons/toga2/compass/workflow/1_transmit_compass_sales_orders_to_mits.php, worker2/Worker/Infrastructure/Errors.php, _underscore/Model/Core/Logs/Issue.php, _underscore/Model/Core/Logs/IssueEmailAddress.php, dbchanges2/Logs/2026-08-04a - MITS rejection business recipients.sql |
14
14
  | [Compass MR/MA Order Auto-Approval & Status Gate](features/mr-ma-order-approval-and-status.md) | 2.0 | Compass **MR** and **MA** sales orders are system-generated from the MITS / Office Depot EDI pipeline (they do not originate as user-entered SA orders) and must | _underscore/Model/Compass/SalesOrder.php, _underscore/Model/Compass/PurchaseOrder.php, worker/crons/toga2/compass/workflow/3a_import_office_depot_purchase_orders.php |
15
+ | [Compass ODP EDI 850 Line-Item Resolution (VA part number to IN SKU fallback)](features/odp-edi-850-item-resolution.md) | 1.0 | How each **PO1 line** on an inbound Office Depot (ODP) **EDI 850** is resolved to a real `Client_Compass` catalog item before cron `3a_import_office_depot_purch | worker/crons/toga2/compass/workflow/3a_import_office_depot_purchase_orders.php, library/app/Edi.php |
15
16
  | [Compass Office Depot EDI 855 Acknowledgement + Over-Quantity PO Guard](features/odp-edi-855-acknowledgement-and-overquantity-guard.md) | 1.0 | How Compass acknowledges Office Depot (ODP) inbound **EDI 850** purchase orders with an **X12 855**, and the **over-quantity guard** that rejects a duplicate PO | worker/crons/toga2/compass/workflow/3a_import_office_depot_purchase_orders.php, worker/crons/toga2/compass/workflow/4_transmit_office_depot_po_acknowledgements.php, library/app/edi.php |
16
17
  | [Compass PEOPLE-File User Lifecycle (duplicate accounts, reactivation grace window, raw-SQL deactivation)](features/people-file-user-lifecycle.md) | 2.0 | The nightly **PEOPLE** file cron (`_Worker_Client_Compass_PeopleFile`, `worker2/Worker/Client/Compass/PeopleFile.php`) owns the whole `Users` row lifecycle for | worker2/Worker/Client/Compass/PeopleFile.php |
17
18
  | [Persona Model & Levy-Sector Gating (worker2 PEOPLE cron)](features/persona-model-and-levy-gating.md) | 2.0 | Compass USA catalogue visibility is driven by **personas** in `Client_Compass`. | worker2/Worker/Client/Compass/PeopleFile.php |
@@ -0,0 +1,179 @@
1
+ ---
2
+ title: Compass ODP EDI 850 Line-Item Resolution (VA part number to IN SKU fallback)
3
+ framework: "1.0"
4
+ repo: worker
5
+ project: Worker
6
+ client: compass-usa
7
+ type: client-feature
8
+ status: active
9
+ updated: 2026-08-25
10
+ owners: ["bala"]
11
+ files:
12
+ - worker/crons/toga2/compass/workflow/3a_import_office_depot_purchase_orders.php
13
+ - library/app/Edi.php
14
+ related:
15
+ - odp-edi-855-acknowledgement-and-overquantity-guard.md
16
+ - ../workflows/odp-order-pipeline-to-netsuite.md
17
+ - ../workflows/odp-duplicate-po-line-cleanup.md
18
+ - ../workflows/odp-edi-import-recovery.md
19
+ - ../profile.md
20
+ ---
21
+
22
+ ## Summary
23
+ How each **PO1 line** on an inbound Office Depot (ODP) **EDI 850** is resolved to a real
24
+ `Client_Compass` catalog item before cron `3a_import_office_depot_purchase_orders.php` writes
25
+ anything. ODP sends **two** identifiers per line and only one of them is reliable:
26
+
27
+ | PO1 element | Qualifier | Parsed into | Reliability |
28
+ |---|---|---|---|
29
+ | `elements[7]` | `VA` | `partNumber` | **free text that drifts** — a mistyped character arrives verbatim |
30
+ | `elements[9]` | `IN` | `vendorPartNumber` (ODP's own SKU) | **correct every time observed** |
31
+
32
+ (Mapping verified in `App_Edi::translatePurchaseOrder850()`, the `case 'PO1'` block. So on a
33
+ Compass PO the `VendorItems.vendorPartNumber` value **is** ODP's SKU, and 3a's later
34
+ PO-item-to-SO-item matching already keyed on it — only sales-order-item creation trusted the VA
35
+ text.)
36
+
37
+ Before 2026-08-25 the importer created the ODP sales-order item from the **VA text only**, so a
38
+ single wrong character killed the whole PO import. Resolution now runs **first**, for **every
39
+ line**, and nothing is written unless all lines resolve.
40
+
41
+ ## The resolution ladder (`resolveOfficeDepotEdiItemPartNumbers()`)
42
+ Runs on the parsed 850 **before** the duplicate-PO check, the over-quantity guard, and any
43
+ SalesOrder / PurchaseOrder / item creation.
44
+
45
+ 1. **VA part number matches the Compass catalog** — use it as-is
46
+ (`partNumberSource = 'ediPartNumber'`).
47
+ 2. **Fall back to ODP's own SKU** — `VendorItems.vendorPartNumber = <IN value>` with
48
+ `vendorId = App_Client_Compass::VENDOR_ID__OFFICE_DEPOT` gives `Items.partNumber`, which is
49
+ then **confirmed to be in the Compass catalog** (`partNumberSource = 'vendorSku'`).
50
+ Auto-resolves **only when exactly ONE catalog candidate remains.**
51
+ 3. **Otherwise the line is unresolved** — it is reported with a plain-English reason and
52
+ **nothing at all is imported for that PO**.
53
+
54
+ Two api2 `GET`s back it, both batched per file (all lines' values in one `in` clause):
55
+ `/items` joined to `Catalogs` filtered on `Catalogs.uuid = UUID_ITEM_CATALOG`
56
+ (`961dbcbc-1943-11ef-8e3b-0aae2980db55`, the "Compass" catalog the cron posts against), and
57
+ `/vendor-items` joined to `Items` filtered on `vendorPartNumber IN (...)` plus the ODP vendor id.
58
+ All matching is **case-insensitive** (keys upper-cased, original casing returned).
59
+
60
+ ### Why SKU resolution and not "clean up the string"
61
+ The real failures were not a single normalisable pattern:
62
+
63
+ | ODP PO | VA part number sent | Correct part number | ODP SKU (IN) |
64
+ |---|---|---|---|
65
+ | `41762948-1170` (2026-08-25) | `US-AND-CROTHALL-POM` (extra `M`) | `US-AND-CROTHALL-PO` | `4976158` |
66
+ | `41242800-1170`, `41243306-1170` (2026-06-17) | `ODTABKITTINGOPT3` (missing `P`) | `ODPTABKITTINGOPT3` | `4900941` |
67
+
68
+ A character-stripping / fuzzy-normalising rule fixes the first and **not** the June pair. The
69
+ SKU is the only identifier that was right in every case, so the SKU is the fallback key.
70
+
71
+ ## Never create an item to match a bad inbound string
72
+ Hand-patching these failures in the past created **duplicate junk items in the Agilant catalog**
73
+ — e.g. `ODTABKITTINGOPT3` (item id 2715) now sits alongside the real `ODPTABKITTINGOPT3`
74
+ (id 2738). Creating an item so a mangled part number "matches" is the wrong fix: it splits one
75
+ physical SKU across two `Items` rows, which is the same shape as the
76
+ [dual-catalog duplicate PO-line bug](../workflows/odp-duplicate-po-line-cleanup.md). Fix the
77
+ `VendorItems` mapping (or have ODP correct their part number) instead.
78
+
79
+ ## ODP VendorItems span two catalogs, and a few SKUs are ambiguous
80
+ Verified on prod `Client_Compass`, `VendorItems.vendorId = 1`: rows point at items in
81
+ **catalogId 1 ("Compass")**, **catalogId 2 ("Agilant")**, and **22 rows carry a null catalog**.
82
+ Three SKUs map to **more than one catalog-1 item**:
83
+
84
+ | ODP SKU | Catalog-1 candidates |
85
+ |---|---|
86
+ | `7128204` | `AW5M5UT` / `AW5M5UT-1` |
87
+ | `7423324` | `INC019BTBK` / `INC019BTBK-NS` |
88
+ | `8030904` | `C30705092` / `C30705092-DS` |
89
+
90
+ This is why step 2 requires **exactly one** candidate. An ambiguous SKU is **reported, never
91
+ guessed** — picking one silently would ship the wrong item. A SKU whose only match sits outside
92
+ catalog 1 is treated the same way (unresolved), which is what the extra catalog-confirmation
93
+ lookup is for.
94
+
95
+ ## Nothing is written until every line resolves
96
+ The gate is a single `$shouldProcessOrder = empty($itemResolution['unresolvedItems'])` check at
97
+ the top of the per-PO loop, ahead of all writes.
98
+
99
+ **What it prevents:** the importer used to create the ODP sales order **first** and add items one
100
+ at a time, so a bad line left a **half-built order in production** — with no PO and no 855.
101
+ Exactly what happened to ODP SO `480238825001` on 2026-08-20: lines 1-3 created, line 4 failed,
102
+ no PurchaseOrder, no acknowledgement. Now a failed line leaves the PO completely untouched, so
103
+ re-dropping the 850 is a clean retry (see
104
+ [import recovery](../workflows/odp-edi-import-recovery.md)).
105
+
106
+ The resolved part numbers are also what the downstream checks now see:
107
+ - the existing `SKIP_IF_CONTAINS_PART_NUMBERS` skip list, and
108
+ - the **over-quantity guard** (`officeDepotPurchaseOrderExceedsCompassDemand()`), which keys on
109
+ `partNumber` and **silently skips any part it cannot match** — so before this change a mangled
110
+ part number bypassed duplicate-quantity protection entirely. See
111
+ [855 Acknowledgement + Over-Quantity Guard](odp-edi-855-acknowledgement-and-overquantity-guard.md).
112
+
113
+ ## `ediPartNumber` — what ODP sent is kept alongside what we resolved
114
+ Every parsed item now carries **both**: `ediPartNumber` (ODP's original VA text) and `partNumber`
115
+ (the resolved catalog value), plus `partNumberSource` (`ediPartNumber` | `vendorSku` | `null`).
116
+ An EDI acknowledgement must reflect the partner's own values, so the **reject 855 is built from
117
+ `ediPartNumber`**.
118
+
119
+ > **The accept 855 still echoes OUR corrected part number.** Cron
120
+ > `4_transmit_office_depot_po_acknowledgements.php` reads `Items.partNumber` off the stored PO,
121
+ > so a SKU-corrected line acknowledges the **real Compass part number**, not the string ODP sent
122
+ > (PO109 still carries their SKU). Accepted deliberately on 2026-08-25. To change it,
123
+ > `ediPartNumber` has to be carried onto the PurchaseOrder record — it is not persisted today.
124
+
125
+ ## A failed lookup degrades to the OLD behaviour, it does not block
126
+ If either api2 lookup returns `isSuccess = false`, `buildItemsUnchangedAfterFailedLookup()`
127
+ returns every line **exactly as ODP sent it** (`ediPartNumber` populated, `partNumber`
128
+ untouched, no unresolved lines) and the import proceeds as it did before this feature existed.
129
+
130
+ **Why:** blocking on a lookup failure would turn a one-line problem into *no ODP PO imports at
131
+ all*. Worst case now equals the old behaviour, never worse. The failed lookup still emails.
132
+
133
+ ## The failure email (one sender, no template in the cron)
134
+ `sendErrorNotification(string $message, ?object $response = null)` is the **only** sender —
135
+ subject `Office Depot PO Import error: <PO numbers>`, the raw EDI attached as
136
+ `EDI_<PO numbers>.txt`. The API-response `print_r` dump is appended **only when a response is
137
+ passed**, so a business-readable message is no longer buried in a dump.
138
+
139
+ `buildItemResolutionErrorMessage()` builds that body: the Compass PO / Compass SO / ODP SO, then
140
+ per failed line the **ODP part number, the ODP SKU, the qty and the exact reason**, closing with
141
+ the fix (set the SKU up as an ODP vendor item against the correct Compass item, or have ODP
142
+ correct the part number, then re-drop the attached file into `OfficeDepot/` on `agilant-as2`).
143
+ All interpolated values pass through `htmlspecialchars()`.
144
+
145
+ **No email template belongs in the cron** — `App_Email_Agilant` already wraps the body via
146
+ `App_Email_Template::renderHtml()`. A second sender and a cron-local template were both
147
+ explicitly rejected.
148
+
149
+ ## Helper map (all in cron 3a)
150
+ | Function | Job |
151
+ |---|---|
152
+ | `resolveOfficeDepotEdiItemPartNumbers()` | the ladder; returns `['items' => ..., 'unresolvedItems' => ...]` |
153
+ | `fetchCatalogPartNumbersByPartNumber()` | `/items` + `Catalogs.uuid` to an upper-cased part-number lookup (`null` = call failed) |
154
+ | `fetchPartNumbersByVendorSku()` | `/vendor-items` + ODP vendor id to SKU then **every** candidate part number |
155
+ | `filterToCatalogPartNumbers()` | keeps only candidates confirmed in the Compass catalog |
156
+ | `buildItemsUnchangedAfterFailedLookup()` | the degrade-to-old-behaviour path |
157
+ | `buildItemResolutionFailureReason()` | one plain sentence per failed line |
158
+ | `buildItemResolutionErrorMessage()` | the full email body |
159
+ | `uniqueNonEmptyValues()` | trim / drop blanks / de-dupe before an `in` clause |
160
+
161
+ ## Verification (2026-08-25)
162
+ `php -l` clean. A standalone harness ran the real resolver functions against prod-shaped lookup
163
+ data — 20 checks covering the `41762948-1170` fix, the June `ODTABKITTINGOPT3` pair, an
164
+ ambiguous SKU, a SKU pointing outside the Compass catalog, a line with no SKU, and the
165
+ failed-lookup degradation. Prod after deploy: ODP `PurchaseOrders` id 112094 number
166
+ `41762948-1170` created 13:32:54 with all 4 items correct, `SalesOrderItems` line 4 id 613590
167
+ created as `US-AND-CROTHALL-PO`, lines 1-3 reused (not duplicated), `FileLog` 878902.
168
+
169
+ ## Change history
170
+ - 2026-08-25 — Built the VA-to-IN resolution ladder so a drifted ODP part number no longer kills
171
+ a PO import: exact catalog match first, then ODP's own SKU through `VendorItems` (vendorId 1)
172
+ confirmed in the Compass catalog, auto-resolving only on a single candidate and reporting
173
+ ambiguity otherwise. Moved resolution **ahead of all writes** (a bad line used to leave a
174
+ half-built ODP SO — `480238825001`, 2026-08-20), fed the resolved part numbers to the skip list
175
+ and the over-quantity guard (which silently skips unmatched parts, so a mangled part number had
176
+ been bypassing duplicate protection), added `ediPartNumber` so the reject 855 echoes ODP's own
177
+ value, and made a failed lookup degrade to the pre-change behaviour rather than block. One
178
+ error sender only, no template in the cron. Recorded the dual-catalog `VendorItems` spread and
179
+ the 3 ambiguous SKUs. Deployed to prod and verified (commit `fc23bb46`). (bala)
@@ -6,13 +6,14 @@ project: Worker
6
6
  client: compass-usa
7
7
  type: client-feature
8
8
  status: active
9
- updated: 2026-08-13
10
- owners: ["jcardinal"]
9
+ updated: 2026-08-25
10
+ owners: ["jcardinal", "bala"]
11
11
  files:
12
12
  - worker/crons/toga2/compass/workflow/3a_import_office_depot_purchase_orders.php
13
13
  - worker/crons/toga2/compass/workflow/4_transmit_office_depot_po_acknowledgements.php
14
14
  - library/app/edi.php
15
15
  related:
16
+ - odp-edi-850-item-resolution.md
16
17
  - ../workflows/odp-order-pipeline-to-netsuite.md
17
18
  - ../workflows/odp-duplicate-po-line-cleanup.md
18
19
  - mits-po-transmission-to-vendors.md
@@ -65,6 +66,25 @@ The demand relay it walks is the standard ODP chain — Compass SO (top, e.g. `M
65
66
  Compass PO → ODP SO → ODP PO (leaf) — via the `SalesOrders_PurchaseOrders` and
66
67
  `PurchaseOrders_SalesOrders` bridges (the same 6-table join the acknowledgement cron uses).
67
68
 
69
+ ### ⚠ The guard only works on RESOLVED part numbers (2026-08-25)
70
+ Because the guard keys on `partNumber` and **ignores any part it cannot match**, a part number
71
+ that ODP mistyped used to slip past duplicate-quantity protection entirely — the guard saw an
72
+ unknown part and passed the PO. Cron 3a now resolves every 850 line to a real Compass catalog part
73
+ number **before** the guard runs, and feeds the guard the resolved values. Anything that fails to
74
+ resolve stops the import outright, so an unmatched part can no longer reach the guard as a silent
75
+ pass. Ladder, ambiguity rules and the degrade-on-lookup-failure policy:
76
+ [ODP EDI 850 Line-Item Resolution](odp-edi-850-item-resolution.md).
77
+
78
+ ## Which part number each 855 echoes
79
+ An acknowledgement must reflect the partner's own values, so the two paths now differ:
80
+
81
+ - **Reject 855 (from 3a)** — built from **`ediPartNumber`**, the raw `VA` part number ODP sent,
82
+ which the resolver preserves on every parsed item alongside the corrected `partNumber`.
83
+ - **Accept 855 (cron 4)** — reads `Items.partNumber` off the **stored** PO, so a SKU-corrected
84
+ line echoes back **our** part number, not ODP's. `PO109` still carries their SKU, so the line is
85
+ still identifiable to them. Accepted deliberately; carrying `ediPartNumber` onto the
86
+ PurchaseOrder record is what it would take to change (it is not persisted today).
87
+
68
88
  ## Why the reject is sent inline, not by cron 4
69
89
  A rejected PO is **never persisted**, so cron 4 — which keys off `PurchaseOrders` rows with
70
90
  `dtAcknowledged IS NULL` — can never see it. The importer therefore **builds and sends the
@@ -127,6 +147,13 @@ A legacy, abandoned sketch (`compass/edi/3_send_edi_855(reject).php`) used **non
127
147
  codes `BAK01 = A1` / `BAK02 = RE` — that path was **deliberately NOT followed**.
128
148
 
129
149
  ## Change history
150
+ - 2026-08-25 — The over-quantity guard is now fed **resolved** part numbers (it ignores parts it
151
+ cannot match, so a mistyped ODP part number had been bypassing duplicate protection), the
152
+ **reject 855 echoes `ediPartNumber`** (what ODP sent) instead of our corrected value, and the
153
+ accept 855's use of our own `Items.partNumber` was recorded as a known, accepted difference.
154
+ Cron 3a's `sendErrorNotification()` became `(string $message, ?object $response = null)` so a
155
+ business-readable failure email is not buried in an API `print_r` dump. Details in the new
156
+ [ODP EDI 850 Line-Item Resolution](odp-edi-850-item-resolution.md). (bala)
130
157
  - 2026-08-13 — Rejections now raise a business alert: cron `3a` throws
131
158
  `App_Exception_Business` (`COMPASS_OFFICE_DEPOT_PO_IMPORT_REJECTED_OVER_QUANTITY`,
132
159
  `URGENCY_HIGH`) so every over-quantity/duplicate reject surfaces as a Logs.Issue instead of
@@ -18,7 +18,7 @@ project: _Underscore
18
18
  client: compass-usa
19
19
  type: profile
20
20
  status: active
21
- updated: 2026-08-20
21
+ updated: 2026-08-25
22
22
  owners: [jcardinal, bala, tcox, apeterson, dfranks]
23
23
  files: []
24
24
  related:
@@ -34,6 +34,7 @@ related:
34
34
  - features/stranded-approval-reassignment.md
35
35
  - workflows/cross-kit-bundle-corruption.md
36
36
  - workflows/odp-duplicate-po-line-cleanup.md
37
+ - features/odp-edi-850-item-resolution.md
37
38
  - features/odp-edi-855-acknowledgement-and-overquantity-guard.md
38
39
  - ../../2.0/apps/worker2/features/compass-vip-support-importer.md
39
40
  - ../../2.0/apps/toga2-commerce/features/expedited-shipping-gating.md
@@ -99,6 +100,14 @@ separate, related client (see its own profile).
99
100
  rejects the whole PO and sends a reject **EDI 855**; ODP POs are acknowledged with an 855
100
101
  (accept via cron 4, reject inline from 3a) built by a shared `App_Edi` method. See
101
102
  [ODP EDI 855 Acknowledgement + Over-Quantity Guard](features/odp-edi-855-acknowledgement-and-overquantity-guard.md).
103
+ - **ODP's part numbers on the 850 drift; their SKU does not (2026-08-25).** A `PO1` line carries a
104
+ free-text `VA` part number **and** ODP's own `IN` SKU. A single mistyped character in the `VA`
105
+ text used to kill the whole PO import (and half-build the ODP sales order). Cron 3a now resolves
106
+ every line — catalog match first, then the SKU through `VendorItems` (vendorId 1) — before it
107
+ writes anything, and never guesses on an ambiguous SKU:
108
+ [ODP EDI 850 Line-Item Resolution](features/odp-edi-850-item-resolution.md).
109
+ **Never create an `Items` row so a mangled inbound part number "matches"** — that is what put
110
+ duplicate junk items in the Agilant catalog (`ODTABKITTINGOPT3` vs `ODPTABKITTINGOPT3`).
102
111
  - ASN ingestion entry points: cXML to the V2 API (logged in `Logs_Compass.Api`) and
103
112
  `worker/crons/toga2/compass/workflow/3b_import_strategic_systems_advance_shipping_notices.php`.
104
113
 
@@ -6,7 +6,7 @@ project: Worker
6
6
  client: compass-usa
7
7
  type: workflow
8
8
  status: active
9
- updated: 2026-08-21
9
+ updated: 2026-08-25
10
10
  owners: ["rgirish", "bala", "dfranks", "jcardinal", "mhammontree"]
11
11
  files:
12
12
  - worker/crons/toga2/compass/workflow/1_transmit_compass_sales_orders_to_mits.php
@@ -19,7 +19,9 @@ files:
19
19
  - library/app/client/compass.php
20
20
  related:
21
21
  - 1.0/apps/worker/features/netsuite-sales-order-sales-rep-sourcing.md
22
+ - clients/compass-usa/features/odp-edi-850-item-resolution.md
22
23
  - clients/compass-usa/features/odp-edi-855-acknowledgement-and-overquantity-guard.md
24
+ - 1.0/apps/worker/workflows/tracing-a-worker-cron-run-in-production.md
23
25
  - clients/compass-usa/workflows/odp-edi-import-recovery.md
24
26
  - clients/compass-usa/workflows/order-lifecycle-and-data-integrity.md
25
27
  - clients/compass-usa/features/mits-po-transmission-to-vendors.md
@@ -97,6 +99,27 @@ duplicates and backfills whatever is missing**:
97
99
  `customerPurchaseOrder` stays `NULL` on a repaired order and must be set manually. See
98
100
  [Recovering a lost ODP EDI import](odp-edi-import-recovery.md).
99
101
 
102
+ ### Every 850 line must resolve to a catalog item before anything is written (2026-08-25)
103
+ 3a resolves **all** PO1 lines to real `Client_Compass` catalog part numbers as the **first** thing
104
+ it does per PO — ahead of the duplicate-PO check, the over-quantity guard and every write. If any
105
+ line fails, **nothing is created for that PO** and the failure is emailed with the EDI attached.
106
+
107
+ Before this, 3a created the ODP SalesOrder first and added items one at a time, so one bad line
108
+ left a **half-built order in prod** with no PurchaseOrder and no 855 (ODP SO `480238825001`,
109
+ 2026-08-20). Combined with the idempotency table above, a failed import is now always a clean
110
+ re-drop rather than a repair job. Ladder, ambiguity rules and the failed-lookup fallback:
111
+ [ODP EDI 850 Line-Item Resolution](../features/odp-edi-850-item-resolution.md).
112
+
113
+ ### ⚠ Known issue (NOT fixed): the `OfficeDepot/` listing makes 3a runs overlap
114
+ An instrumented prod run on 2026-08-25 showed the S3 listing returning **138,472 objects** and
115
+ taking **32 seconds** before any PO was processed, with each PO then costing roughly 40 seconds.
116
+ A file-bearing run therefore blows past the `*/5` schedule and the following runs are **silently
117
+ skipped** by `App_Framework::exitIfProcessRunning()` — leaving **no `CronJobExecutions` row at
118
+ all**, which reads as "the cron never fired". The object count is driven by `OfficeDepot/SENT/`
119
+ and `OfficeDepot/OUTBOX/` accumulating: 3a skips those prefixes when *processing* but still
120
+ *lists* them. Fix candidates (separate work): prune `SENT/`, or list with a tighter prefix. See
121
+ [Tracing a worker cron run in production](../../1.0/apps/worker/workflows/tracing-a-worker-cron-run-in-production.md).
122
+
100
123
  ### The EDI 850 → TOGa identifier mapping (not guessable — this is the key to tracing an ODP order)
101
124
 
102
125
  | 850 segment | Meaning | Where it lands |
@@ -105,7 +128,14 @@ duplicates and backfills whatever is missing**:
105
128
  | **`REF~QC`** | the **ODP SALES ORDER number** (e.g. `475390815001`) | `SalesOrders.number` with `customerId = 1` |
106
129
  | **`REF~EU` / `REF~PO`** | the **Compass MITS PO number** (e.g. `50310002-1`) | its `c_mitsSalesOrder` field holds the ODP SO number |
107
130
  | **`REF~LU`** | the **Compass sales order** (e.g. `SA135273`) | `SalesOrders` with `customerId = 2` |
108
- | **`PO1`** | the line | qty, UOM, unit price, vendor part, ODP item id |
131
+ | **`PO1`** | the line | qty, UOM, unit price, and **two** item identifiers (below) |
132
+
133
+ **`PO1` carries two item identifiers and only one of them is trustworthy:** `elements[7]` (the
134
+ `VA` qualifier) is parsed into **`partNumber`** and is **free text that drifts**; `elements[9]`
135
+ (the `IN` qualifier) is parsed into **`vendorPartNumber`** and is **ODP's own SKU**, correct in
136
+ every observed case. So on a Compass PO the `VendorItems.vendorPartNumber` value *is* ODP's SKU.
137
+ Cron 3a resolves each line through that SKU when the `VA` text does not match the catalog — see
138
+ [ODP EDI 850 Line-Item Resolution](../features/odp-edi-850-item-resolution.md).
109
139
 
110
140
  **⚠ `REF~QC` — not `BEG` — is the value to search on for the ODP SalesOrder.** Searching
111
141
  `PurchaseOrders` / `SalesOrders` for the `BEG` number finds nothing for the SO and wastes time.
@@ -242,6 +272,13 @@ wrong record makes every order look "stuck." Use the **join**, never a number ma
242
272
  transmission feature.)
243
273
 
244
274
  ## Change history
275
+ - 2026-08-25 — Cron 3a now **resolves every 850 line to a catalog item before it writes anything**
276
+ (a bad line used to leave a half-built ODP SO); documented the `PO1` two-identifier detail
277
+ (`elements[7]` = drifting `VA` text, `elements[9]` = ODP's reliable `IN` SKU) behind the new
278
+ [850 item-resolution feature](../features/odp-edi-850-item-resolution.md). Flagged the
279
+ **unfixed** `OfficeDepot/` listing problem: 138,472 objects / 32 s per run pushes 3a past its
280
+ `*/5` slot, so following runs are skipped by the overlap guard with **no `CronJobExecutions`
281
+ row** to show it. (bala)
245
282
  - 2026-08-21 — **Cron 5 sales rep** no longer hardcodes employee 34877: it reads the end
246
283
  customer's native NetSuite `salesRep` RecordRef (`App_NetSuite::getCustomer()`, cached per
247
284
  run, failures logged and treated as "no rep", left unset when absent), with the end-customer
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "toga-ai",
3
- "version": "1.0.645",
3
+ "version": "1.0.646",
4
4
  "description": "TOGA Technology Team Claude Knowledge System — shared AI coding harness with skills, knowledge base CLI, and project installer for Claude Code.",
5
5
  "keywords": [
6
6
  "claude",