toga-ai 1.0.645 → 1.0.646
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/knowledge/1.0/apps/library/features/cron-execution-monitoring.md +25 -3
- package/knowledge/1.0/apps/worker/INDEX.md +2 -1
- package/knowledge/1.0/apps/worker/architecture.md +35 -2
- package/knowledge/1.0/apps/worker/features/oneuptime-worker-uptime-monitoring.md +40 -3
- package/knowledge/1.0/apps/worker/workflows/tracing-a-worker-cron-run-in-production.md +124 -0
- package/knowledge/2.0/apps/_underscore/features/per-client-database-connections.md +17 -2
- package/knowledge/2.0/apps/worker2/INDEX.md +2 -1
- package/knowledge/2.0/apps/worker2/features/cross-account-aws-access.md +42 -3
- package/knowledge/2.0/apps/worker2/features/oneuptime-worker2-monitoring.md +39 -0
- package/knowledge/2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md +175 -0
- package/knowledge/INDEX.md +2 -2
- package/knowledge/clients/compass-usa/INDEX.md +1 -0
- package/knowledge/clients/compass-usa/features/odp-edi-850-item-resolution.md +179 -0
- package/knowledge/clients/compass-usa/features/odp-edi-855-acknowledgement-and-overquantity-guard.md +29 -2
- package/knowledge/clients/compass-usa/profile.md +10 -1
- package/knowledge/clients/compass-usa/workflows/odp-order-pipeline-to-netsuite.md +39 -2
- package/package.json +1 -1
|
@@ -6,12 +6,13 @@ project: Library
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-
|
|
10
|
-
owners: [dfranks]
|
|
9
|
+
updated: 2026-08-25
|
|
10
|
+
owners: [dfranks, bala]
|
|
11
11
|
files:
|
|
12
12
|
- library/app/framework.php
|
|
13
13
|
related:
|
|
14
14
|
- ../../worker/architecture.md
|
|
15
|
+
- ../../worker/workflows/tracing-a-worker-cron-run-in-production.md
|
|
15
16
|
---
|
|
16
17
|
|
|
17
18
|
## Summary
|
|
@@ -38,13 +39,28 @@ row in `db_log` and of the per-job Sentry check-in monitors.
|
|
|
38
39
|
cast silently discards the sub-second precision even when the schema column is `decimal(10,3)`.
|
|
39
40
|
- **Schema:** the table is provisioned by the dbchanges migration
|
|
40
41
|
`dbchanges/Common/DF/2026-5-7 Cron Checkin.sql` (table `CronJobExecutions`,
|
|
41
|
-
`executionTimeSeconds decimal(10,3)
|
|
42
|
+
`executionTimeSeconds decimal(10,3)`). **Correction (2026-08-25): prod `Common.CronJobExecutions`
|
|
43
|
+
does have a `note` TEXT column** — verified by writing to it and reading it back. `App_Framework`
|
|
44
|
+
never populates it, which makes it the practical place to park **temporary** step tracing for a
|
|
45
|
+
cron you are diagnosing (worker crons have no readable stdout — see
|
|
46
|
+
[Tracing a worker cron run in production](../../worker/workflows/tracing-a-worker-cron-run-in-production.md)).
|
|
47
|
+
Remove the tracing when you are done, and never write payloads or credentials into it.
|
|
48
|
+
- **Don't confuse the two `note`s.** `cronInitialization()` writes `note = 'Started execution'` to
|
|
49
|
+
the **`Log`** table on `db_log` (the older `CRON`/`recordType` row), not to `CronJobExecutions`.
|
|
42
50
|
- **Design decision:** 1.0 records executions by writing **directly to the shared DB**
|
|
43
51
|
(`db_common`), *not* by POSTing from 1.0 to a 2.0 API endpoint. Per Jeff Cardinal, 1.0 code
|
|
44
52
|
must not POST to 2.0 code; the direct shared-DB write is the sanctioned path (TRUE-78182).
|
|
45
53
|
|
|
46
54
|
## Gotchas
|
|
47
55
|
|
|
56
|
+
- **⚠ A run skipped by the overlap guard leaves NO ROW — absence is ambiguous.**
|
|
57
|
+
`cronInitialization()` calls `exitIfProcessRunning()` **before** the `CronJobExecutions` INSERT,
|
|
58
|
+
and the `cronFinished(true)` that guard then calls does nothing because `cronLogId` was never
|
|
59
|
+
set. So a blocked run is indistinguishable from a run that never launched: **a missing row for
|
|
60
|
+
an expected slot means "skipped", not "hung"**, and a hung run looks different (it has
|
|
61
|
+
`dtCheckIn` with no `dtCheckOut`). `isProcessRunning()` matches **any** `ps -ef` line containing
|
|
62
|
+
the script path (it only excludes `/bin/sh` and its own pid), so a stray `tail -f` or editor on
|
|
63
|
+
that path blocks the cron indefinitely.
|
|
48
64
|
- **`db_log` has no `CronJobExecutions` table.** A `cronFinished()` UPDATE that runs against the
|
|
49
65
|
`db_log` connection throws on **every** 1.0 cron job. Both the INSERT and the UPDATE must
|
|
50
66
|
target `db_common`.
|
|
@@ -60,6 +76,12 @@ row in `db_log` and of the per-job Sentry check-in monitors.
|
|
|
60
76
|
logic. (Fixed 2026-07-06: consolidated back to one `db_common` INSERT/UPDATE pair.)
|
|
61
77
|
|
|
62
78
|
## Change history
|
|
79
|
+
- 2026-08-25 — Corrected the schema note (prod `CronJobExecutions` **does** carry a `note` TEXT
|
|
80
|
+
column, unused by `App_Framework`, usable for temporary cron tracing) and separated it from the
|
|
81
|
+
`note = 'Started execution'` row `cronInitialization()` writes to `Log` on `db_log`. Recorded
|
|
82
|
+
that the overlap guard runs **before** the INSERT, so a skipped run writes no row at all and
|
|
83
|
+
cannot be told apart from "never launched" — and that `isProcessRunning()`'s `ps -ef` substring
|
|
84
|
+
match lets any process holding the script path block the cron. No code change. (bala)
|
|
63
85
|
- 2026-07-06 — Fixed a merge (`origin/_production`) that reintroduced duplicate
|
|
64
86
|
`CronJobExecutions` writes and pointed one `cronFinished()` UPDATE at `db_log` (where the
|
|
65
87
|
table doesn't exist), which would have thrown on every 1.0 cron. Consolidated to one
|
|
@@ -11,9 +11,10 @@
|
|
|
11
11
|
| [NetSuite Sales Order Sales Rep Sourcing (Staples & ODP EDI orders)](features/netsuite-sales-order-sales-rep-sourcing.md) | How the **sales rep** on a NetSuite Sales Order is determined for the two 1.0 `worker` EDI order-creation integrations (Staples cXML and Compass/ODP EDI). | worker/crons/toga2/compass/workflow/5_create_netsuite_sales_orders_from_office_depot_purchase_orders.php, worker/crons/sync/staples/sync_staples_cxml.php, test/@Mark/NetSuite/TRUE_80451_customer_salesrep_diag.php |
|
|
12
12
|
| [NetSuite → TOGa Supply Per-Client Sync (thin wrappers)](features/netsuite-togasupply-per-client-sync.md) | Syncs NetSuite transactions (sales orders, purchase orders, invoices, item receipts, item fulfillments, inventory adjustments) into each TOGa Supply (2.0) clien | worker/crons/toga2/netsuite/common_sync_togasupply.php, worker/crons/toga2/netsuite/sync_togasupply_canon.php, worker/crons/toga2/netsuite/sync_togasupply_elite.php, worker/schedules/cron.worker.sync.json, dbchanges2/_modules/netsuite/2026-04-01 - Parameters.sql, library/app/api/toga2.php, library/app/api/netsuite/rest.php, library/app/framework.php, library/app/systemmonitor/netsuiteintegration.php, test/@srija/Elite Testing/Service Requests/test_sync_togasupply_elite_section.php, test/@srija/Elite Testing/Service Requests/test_diagnose_togasupply_elite.php |
|
|
13
13
|
| [OneUptime Server monitor + disk/memory hygiene on the 1.0 worker EB host](features/oneuptime-server-monitor-host-hygiene.md) | The 1.0 `agilant-worker` EB environment runs on the **legacy Amazon Linux 1 PHP 7.2 platform** (Apache httpd/prefork, s3fs mounts, cron) and repeatedly went dow | worker/.ebextensions/040_disk_memory_hygiene.config, worker/.ebextensions/045_oneuptime_agent.config, worker/ebs/cron.worker.php, worker/ebs/mount-s3fs-folders.php, worker/ebs/apache_settings.php, worker/ebs/setup_phpini.php |
|
|
14
|
-
| [OneUptime external uptime monitoring for 1.0 workers](features/oneuptime-worker-uptime-monitoring.md) | Every 1.0 worker box self-reports its liveness to an external OneUptime monitor once per minute by curl-POSTing to a per-worker "Incoming Request" heartbeat URL | library/app/worker.php, worker/crons/worker/worker_heartbeat.php |
|
|
14
|
+
| [OneUptime external uptime monitoring for 1.0 workers](features/oneuptime-worker-uptime-monitoring.md) | Every 1.0 worker box self-reports its liveness to an external OneUptime monitor once per minute by curl-POSTing to a per-worker "Incoming Request" heartbeat URL | library/app/worker.php, worker/crons/worker/worker_heartbeat.php, worker/ebs/cron.worker.php, worker/.ebextensions/045_oneuptime_agent.config |
|
|
15
15
|
| [Prudential: Send Shipments for the Day report (daily cron)](features/send-shipments-for-the-day.md) | Daily cron (9:00 PM) that emails Prudential and Dell stakeholders an Excel report of all devices shipped that day, including tracking number, serial number, emp | worker/crons/notifications/reports/send_shipments_for_the_day.php |
|
|
16
16
|
| [Staples cXML Order Import (SFTP → NetSuite)](features/staples-cxml-order-import.md) | `sync_staples_cxml.php` is an **hourly** cron (runs at **:45**) that imports Staples cXML purchase orders from SFTP into NetSuite as Sales Orders, then writes a | worker/crons/sync/staples/sync_staples_cxml.php |
|
|
17
17
|
| [Diagnosing frozen 1.0 worker cron check-ins (Sentry "missed" flood)](workflows/diagnosing-frozen-cron-checkins.md) | When 1.0 worker cron timestamps freeze and Sentry project `worker1` fills with **`missed`** check-ins, the intuitive diagnosis — a wedged `App_Framework::isProc | worker/.ebextensions/cron.config, library/app/worker.php |
|
|
18
18
|
| [isFulfillable Multi-Client Backfill (all togasupply clients)](workflows/isfulfillable-multi-client-backfill.md) | One-time backfill that catches up `Items.isFulfillable` on **existing** items across **all 17 togasupply clients** (AIG, Broward Sheriff, Canon, Endeavor Health | worker/crons/toga2/netsuite/backfill_isfulfillable_all_clients.php, library/app/api/toga2.php |
|
|
19
19
|
| [Onboarding a Client to the NetSuite TOGa Supply Sync](workflows/onboarding-client-to-netsuite-togasupply-sync.md) | How to add a new TOGa 2 client to the per-client NetSuite → TOGa Supply importer (`worker/crons/toga2/netsuite/`). | worker/crons/toga2/netsuite/sync_togasupply.php, worker/crons/toga2/netsuite/common_sync_togasupply.php, worker/crons/toga2/netsuite/sync_togasupply_elite.php, worker/schedules/cron.worker.sync.json, dbchanges2/_modules/netsuite/2026-04-01 - Parameters.sql, dbchanges2/_modules/netsuite/2026-08-05 - CLEAN NETSUITE CLINET.SQL |
|
|
20
|
+
| [Tracing a 1.0 worker cron run in production (no stdout, silent skips)](workflows/tracing-a-worker-cron-run-in-production.md) | How to answer *"did this cron actually run, and what did it do?"* on the 1.0 `worker` tier, where **there is no usable stdout** and **a skipped run leaves no tr | worker/ebs/cron.worker.php, worker/schedules/cron.worker.sync.json, library/app/framework.php |
|
|
@@ -6,8 +6,8 @@ project: Worker
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: architecture
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-
|
|
10
|
-
owners: [jcardinal, sking]
|
|
9
|
+
updated: 2026-08-24
|
|
10
|
+
owners: [jcardinal, sking, bala]
|
|
11
11
|
files:
|
|
12
12
|
- worker/index.php
|
|
13
13
|
- worker/_/app/framework.php
|
|
@@ -20,6 +20,8 @@ files:
|
|
|
20
20
|
related:
|
|
21
21
|
- ../library/architecture.md
|
|
22
22
|
- ../dbchanges/workflows/authoring-and-shipping-sql-files.md
|
|
23
|
+
- ./features/oneuptime-worker-uptime-monitoring.md
|
|
24
|
+
- ../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md
|
|
23
25
|
---
|
|
24
26
|
|
|
25
27
|
## Summary
|
|
@@ -44,6 +46,7 @@ which **self-elects a distinct role** (`notification`, `database`, `infrastructu
|
|
|
44
46
|
under `crons/`; the schedule (cron registration) is the source of truth for what runs — a script
|
|
45
47
|
that isn't scheduled never executes. Client integrations live under `crons/toga2/<client>/`.
|
|
46
48
|
Use prepared statements for all SQL; never interpolate input.
|
|
49
|
+
A worker that fails deploy-time role election runs NO crons while EB still reads healthy.
|
|
47
50
|
|
|
48
51
|
## How a job becomes a cron (the dispatch pipeline)
|
|
49
52
|
|
|
@@ -108,6 +111,32 @@ is the most important and least obvious part of the architecture.
|
|
|
108
111
|
> Net effect: roles are a **claim-the-first-free-slot pool**, and a dead worker's role is
|
|
109
112
|
> reclaimed by its replacement within minutes. There is no static instance→role mapping to edit.
|
|
110
113
|
|
|
114
|
+
### The failure mode: NO NAME = NO CRONTAB (and it is silent)
|
|
115
|
+
|
|
116
|
+
The election is the **single point of failure for an entire instance**, and the architecture gives
|
|
117
|
+
it no second chance:
|
|
118
|
+
|
|
119
|
+
- The claim runs **only at DEPLOY time** — `ebs/cron.worker.php` is written as the EB `appdeploy`
|
|
120
|
+
enact hook. **It never retries.** An instance that fails to claim stays nameless until the next
|
|
121
|
+
deploy or until something terminates it.
|
|
122
|
+
- The claim writes `/etc/worker-role` and **only then** appends `cron.worker.<role>.json` to the
|
|
123
|
+
crontab. So a nameless instance has **no role crontab at all** — it runs *nothing*, not even
|
|
124
|
+
`worker_heartbeat.php`, while **EB reports the environment healthy**. The box is up; it just
|
|
125
|
+
does no work.
|
|
126
|
+
- **The claim's DB connect is `@mysqli_connect(...)` — error-suppressed.** A boot-time connect
|
|
127
|
+
failure to `Vision_Log` therefore leaves the instance nameless **with no log line anywhere**.
|
|
128
|
+
This is the prime suspect for the 2026-08-24 incident (2 of 7 workers roleless for ~3 hours),
|
|
129
|
+
and it violates the no-`@`-suppression rule in `rules/toga/common/coding-style.md`. Removing the
|
|
130
|
+
`@` and logging the failure is the recommended fix.
|
|
131
|
+
- `Workers` **only ever holds instances that SUCCEEDED**, so a failed claim leaves no row —
|
|
132
|
+
**absence is the only evidence**, and detecting it requires the AWS instance list.
|
|
133
|
+
|
|
134
|
+
Because every OneUptime monitor on this tier is keyed by **role** rather than by instance, none of
|
|
135
|
+
them can see this (see
|
|
136
|
+
[the blind spot](./features/oneuptime-worker-uptime-monitoring.md)). External coverage now comes
|
|
137
|
+
from the 2.0
|
|
138
|
+
[Worker Fleet Role-Assignment Monitor](../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md).
|
|
139
|
+
|
|
111
140
|
## Anatomy of a cron script
|
|
112
141
|
|
|
113
142
|
Every script is self-contained and follows this boilerplate:
|
|
@@ -208,6 +237,10 @@ web face (`mvc/` GET routes for login/logout/404); the tier's real work is the c
|
|
|
208
237
|
|
|
209
238
|
## Conventions & gotchas
|
|
210
239
|
|
|
240
|
+
- **A worker that fails role election is silently dead, not degraded.** No name → no crontab →
|
|
241
|
+
the instance runs nothing while EB reads healthy, and the `@`-suppressed `mysqli_connect` in
|
|
242
|
+
`ebs/cron.worker.php` logs nothing. Never treat "EB is green" as evidence the fleet is working;
|
|
243
|
+
check `Vision_Log.Workers` against the actual EB instance list.
|
|
211
244
|
- **`dbchanges` auto-apply is branch-gated, and `_production` is NOT auto-applied.**
|
|
212
245
|
`crons/infrastructure/execute_dbchanges.php` runs every 2 minutes but only for branches matching
|
|
213
246
|
`_%`, deriving the env by stripping the leading `_` and requiring `config.<env>.ini` in the
|
|
@@ -6,12 +6,17 @@ project: Worker
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-
|
|
10
|
-
owners: ["jcardinal"]
|
|
9
|
+
updated: 2026-08-24
|
|
10
|
+
owners: ["jcardinal", "bala"]
|
|
11
11
|
files:
|
|
12
12
|
- library/app/worker.php
|
|
13
13
|
- worker/crons/worker/worker_heartbeat.php
|
|
14
|
-
|
|
14
|
+
- worker/ebs/cron.worker.php
|
|
15
|
+
- worker/.ebextensions/045_oneuptime_agent.config
|
|
16
|
+
related:
|
|
17
|
+
- ./oneuptime-server-monitor-host-hygiene.md
|
|
18
|
+
- ../architecture.md
|
|
19
|
+
- ../../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md
|
|
15
20
|
---
|
|
16
21
|
|
|
17
22
|
## Summary
|
|
@@ -57,6 +62,31 @@ Database, Infrastructure, TOGa, TOGa Desk, and Catalog, then importing each into
|
|
|
57
62
|
manually. To add a new worker's monitor, clone an existing monitor export the same way and
|
|
58
63
|
register the resulting heartbeat URL in `App_Worker::$oneUptimeEndpoints`.
|
|
59
64
|
|
|
65
|
+
## THE BLIND SPOT — every monitor here is keyed by ROLE, so a ROLELESS instance is invisible
|
|
66
|
+
|
|
67
|
+
Load-bearing limitation, learned the hard way (2026-08-24: **2 of the 7 production workers ran
|
|
68
|
+
roleless for ~3 hours undetected**). This layer, and the Server monitors in
|
|
69
|
+
[host hygiene](./oneuptime-server-monitor-host-hygiene.md), are keyed by `workerName` — **14
|
|
70
|
+
monitors, none keyed by instance**:
|
|
71
|
+
|
|
72
|
+
- `worker_heartbeat.php` only pings `App_Worker::$oneUptimeEndpoints[$row['workerName']]` for
|
|
73
|
+
the row matching **its own** `instanceId`. No row → no ping target → it pushes nothing.
|
|
74
|
+
- The `045_oneuptime_agent.config` Server agent reads `/etc/worker-role` and **exits clean when
|
|
75
|
+
it is empty.**
|
|
76
|
+
|
|
77
|
+
So an instance that never claimed a role reports to nothing — and because every role is still
|
|
78
|
+
owned by *someone*, **no role-keyed monitor is missing a ping either.** The fleet reads 100%
|
|
79
|
+
healthy while a box does zero work (role assignment happens **only at deploy time**, and no role
|
|
80
|
+
means the role's `cron.worker.<role>.json` is never appended to the crontab, so the instance runs
|
|
81
|
+
**nothing**).
|
|
82
|
+
|
|
83
|
+
`Vision_Log.Workers` only ever holds instances that **succeeded** in claiming a role, so the
|
|
84
|
+
failure leaves no row and no log line — **absence is the only evidence**, and detecting an
|
|
85
|
+
absence requires the AWS instance list, which nothing on this tier consults. That gap is what
|
|
86
|
+
the 2.0
|
|
87
|
+
[Worker Fleet Role-Assignment Monitor](../../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md)
|
|
88
|
+
now closes from outside the fleet. **Do not try to close it with another role-keyed monitor.**
|
|
89
|
+
|
|
60
90
|
## Gotchas
|
|
61
91
|
|
|
62
92
|
- Do **not** hardcode the heartbeat URLs/UUIDs anywhere but `App_Worker::$oneUptimeEndpoints`
|
|
@@ -67,6 +97,13 @@ register the resulting heartbeat URL in `App_Worker::$oneUptimeEndpoints`.
|
|
|
67
97
|
pending cleanup — the live source of truth is the `App_Worker::$oneUptimeEndpoints` map.
|
|
68
98
|
|
|
69
99
|
## Change history
|
|
100
|
+
- 2026-08-24 — Documented **the blind spot**: all 14 monitors on this tier are keyed by
|
|
101
|
+
`workerName`, so an instance that claimed **no** role pings nothing and leaves every role-keyed
|
|
102
|
+
monitor still green — which is how 2 of 7 production workers ran roleless for ~3 hours
|
|
103
|
+
undetected. Recorded that `Vision_Log.Workers` only holds successful claims (absence is the
|
|
104
|
+
only evidence) and that the gap is now covered from outside the fleet by the 2.0
|
|
105
|
+
[Worker Fleet Role-Assignment Monitor](../../../2.0/apps/worker2/features/worker-fleet-role-assignment-monitor.md).
|
|
106
|
+
(bala)
|
|
70
107
|
- 2026-07-13 — Added external OneUptime heartbeat push for all 1.0 workers on top of the
|
|
71
108
|
existing internal DB-heartbeat mechanism; endpoints keyed by workerName in
|
|
72
109
|
`App_Worker::$oneUptimeEndpoints`. (jcardinal)
|
|
@@ -0,0 +1,124 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Tracing a 1.0 worker cron run in production (no stdout, silent skips)
|
|
3
|
+
framework: "1.0"
|
|
4
|
+
repo: worker
|
|
5
|
+
project: Worker
|
|
6
|
+
client: shared
|
|
7
|
+
type: workflow
|
|
8
|
+
status: active
|
|
9
|
+
updated: 2026-08-25
|
|
10
|
+
owners: ["bala"]
|
|
11
|
+
files:
|
|
12
|
+
- worker/ebs/cron.worker.php
|
|
13
|
+
- worker/schedules/cron.worker.sync.json
|
|
14
|
+
- library/app/framework.php
|
|
15
|
+
related:
|
|
16
|
+
- ../architecture.md
|
|
17
|
+
- ./diagnosing-frozen-cron-checkins.md
|
|
18
|
+
- ../../library/features/cron-execution-monitoring.md
|
|
19
|
+
- ../../../1.0/standards/backend-php.md
|
|
20
|
+
---
|
|
21
|
+
|
|
22
|
+
## Summary
|
|
23
|
+
How to answer *"did this cron actually run, and what did it do?"* on the 1.0 `worker` tier, where
|
|
24
|
+
**there is no usable stdout** and **a skipped run leaves no trace at all**. Written from a
|
|
25
|
+
production diagnostic session on
|
|
26
|
+
`worker/crons/toga2/compass/workflow/3a_import_office_depot_purchase_orders.php`; the mechanics
|
|
27
|
+
apply to every cron on the tier.
|
|
28
|
+
|
|
29
|
+
Use this before you conclude a cron "never fired" — three different situations look identical
|
|
30
|
+
from the outside.
|
|
31
|
+
|
|
32
|
+
## Step 1 — get the real schedule from `schedules/*.json`, never from the file
|
|
33
|
+
The `.php` file's own header comment is not authoritative and is routinely stale. So are skills,
|
|
34
|
+
docs and tickets. `worker/schedules/cron.<env>[.<role>].json` is the only source of truth.
|
|
35
|
+
|
|
36
|
+
Worked example: the entry *"Download EDI From S3 Files & Create PO - EDI Office Depot"* in
|
|
37
|
+
`cron.worker.sync.json` is `*/5 * * * *` — **every 5 minutes**. The cron file's header said
|
|
38
|
+
"EVERY HOUR" (corrected 2026-08-25) and at least one skill still says hourly. A wrong assumed
|
|
39
|
+
frequency makes you read a normal gap as an outage.
|
|
40
|
+
|
|
41
|
+
Also check whether the job is registered in the env you are testing in at all: this one is **not**
|
|
42
|
+
in `cron.beta.json`, so beta never runs it on a schedule. Nothing on beta is evidence about it.
|
|
43
|
+
|
|
44
|
+
## Step 2 — echo/print goes nowhere useful, so do not debug with it
|
|
45
|
+
The crontab line is assembled in `worker/ebs/cron.worker.php`, and the two blocks that build it
|
|
46
|
+
differ:
|
|
47
|
+
|
|
48
|
+
| Schedule file | Command written | Where output goes |
|
|
49
|
+
|---|---|---|
|
|
50
|
+
| base `cron.<env>.json` | `php /var/www/html/crons/<script>` | **nowhere** — the ` > /var/www/cache/WORKER_ERROR_$RANDOM` redirect is **commented out** |
|
|
51
|
+
| role `cron.worker.<role>.json` | `php /var/www/html/crons/<script> > /var/www/cache/WORKER_ERROR_$RANDOM` | a **randomly named** file under `/var/www/cache/` on whichever of the 7 instances ran it |
|
|
52
|
+
|
|
53
|
+
Either way you cannot go and read it: base-schedule output is discarded outright, and a
|
|
54
|
+
role-schedule run leaves an unpredictable filename on a box in an autoscaled fleet. **Treat
|
|
55
|
+
worker cron stdout as write-only.**
|
|
56
|
+
|
|
57
|
+
## Step 3 — put temporary tracing somewhere you can query
|
|
58
|
+
What worked: write timestamped step markers into the **`note`** column of the run's
|
|
59
|
+
`Common.CronJobExecutions` row (the row `App_Framework::cronInitialization()` already inserted),
|
|
60
|
+
then read them back over the DB tooling from your desk. `dtCheckIn`, `dtCheckOut` and
|
|
61
|
+
`executionTimeSeconds` on the same row give you the shape of the run for free.
|
|
62
|
+
|
|
63
|
+
Rules: keep it to a few markers, and **remove the tracing once diagnosed** (it was removed in
|
|
64
|
+
this session). Never write payloads or credentials into `note`.
|
|
65
|
+
|
|
66
|
+
## Step 4 — read `Common.CronJobExecutions` to tell "ran", "hung" and "skipped" apart
|
|
67
|
+
Query the legacy env's `Common.CronJobExecutions` filtering `job LIKE` the script path:
|
|
68
|
+
|
|
69
|
+
- **`dtCheckIn` + `dtCheckOut` + `executionTimeSeconds`** — it ran and finished; the duration
|
|
70
|
+
tells you whether it fits inside its schedule slot.
|
|
71
|
+
- **`dtCheckIn` with no `dtCheckOut`** — it started and never completed (crash, timeout, or the
|
|
72
|
+
instance went away).
|
|
73
|
+
- **no row for a slot you expected** — the run was **skipped by the overlap guard**, not "never
|
|
74
|
+
fired". A blocked run writes **nothing**: `cronInitialization()` calls
|
|
75
|
+
`exitIfProcessRunning()` **before** the INSERT, and the `cronFinished(true)` it then calls is a
|
|
76
|
+
no-op because `cronLogId` is unset. See
|
|
77
|
+
[Cron Execution Monitoring](../../library/features/cron-execution-monitoring.md).
|
|
78
|
+
|
|
79
|
+
### The overlap guard is a `ps -ef` substring match — anything can block it
|
|
80
|
+
`App_Framework::isProcessRunning()` returns true for **any** `ps -ef` line containing the script
|
|
81
|
+
path (it only excludes `/bin/sh` lines and its own pid). So a developer's `tail -f`, an editor, or
|
|
82
|
+
any shell holding that path in its command line **blocks the cron indefinitely** while looking
|
|
83
|
+
like a legitimate overlap. Check `ps -ef` for what is actually holding the name before assuming a
|
|
84
|
+
long-running instance of the job itself.
|
|
85
|
+
|
|
86
|
+
## Step 5 — if runs overlap, look at what the job does before its real work
|
|
87
|
+
A job whose own runtime exceeds its schedule interval silently loses most of its slots to the
|
|
88
|
+
guard. Measure the phases, not the total: in the ODP case an instrumented run showed the S3
|
|
89
|
+
listing alone returning **138,472 objects and taking 32 seconds** before any PO was touched, with
|
|
90
|
+
each PO then costing roughly 40 seconds — comfortably past a 5-minute schedule. Prefixes the job
|
|
91
|
+
skips while processing are still listed, so accumulated `SENT/` and `OUTBOX/` objects inflate
|
|
92
|
+
every run.
|
|
93
|
+
|
|
94
|
+
## Deploy-time trap: a PHP 8-only syntax kills the whole cron with no visible error
|
|
95
|
+
This tier runs **PHP 7.2**, so a **PHP 8 named argument** (`someFunction(items: $x)`) is a
|
|
96
|
+
**parse error**, not a runtime warning:
|
|
97
|
+
|
|
98
|
+
```
|
|
99
|
+
PHP Parse error: syntax error, unexpected ':', expecting ',' or ')'
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
The script then does nothing at all on every tick — and per Steps 2 and 4 you will see no output
|
|
103
|
+
and no `CronJobExecutions` row, i.e. it looks exactly like "the cron never fired". Typed
|
|
104
|
+
parameters and return types (`: void`, `: array`, `: bool`, `?array`, `?object`, `object $x`) are
|
|
105
|
+
PHP 7.0/7.1 and **are** safe here; they are already used in
|
|
106
|
+
`worker/crons/toga2/compass/backfill_items_isfulfillable.php` and
|
|
107
|
+
`update_item_fulfillments_in_netsuite.php`.
|
|
108
|
+
|
|
109
|
+
`1.0/standards/backend-php.md` already forbids named arguments in 1.0 — but a workspace
|
|
110
|
+
`CLAUDE.md` / `.github/copilot-instructions.md` that says *"use PHP 8 Named Parameters"* applies
|
|
111
|
+
to **2.0 only**. The 1.0 standard wins in `library/`, `worker/` and every other 1.0 app.
|
|
112
|
+
|
|
113
|
+
**Do not read a PHP 8 function call in existing code as proof the tier is PHP 8.**
|
|
114
|
+
`library/app/systemmonitor/500error.php` uses `str_contains()`; that file is not a version
|
|
115
|
+
signal, it is a latent bug.
|
|
116
|
+
|
|
117
|
+
## Change history
|
|
118
|
+
- 2026-08-25 — Documented from a prod diagnostic session on the Compass ODP 850 importer: the
|
|
119
|
+
schedule JSON (not the file header) is authoritative, worker cron stdout is unreadable (base
|
|
120
|
+
block's `WORKER_ERROR` redirect commented out, role block writes a random filename),
|
|
121
|
+
`Common.CronJobExecutions.note` is the practical place for temporary tracing, a run skipped by
|
|
122
|
+
the `ps -ef` overlap guard writes **no row at all** (guard runs before the INSERT), the guard
|
|
123
|
+
matches any process holding the script path, and a PHP 8 named argument parse-errors the whole
|
|
124
|
+
cron silently on this PHP 7.2 tier. (bala)
|
|
@@ -6,8 +6,8 @@ project: _Underscore
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-08-
|
|
10
|
-
owners: ["dfranks", "jcardinal", "mhammontree", "apeterson", "kyalamarthi"]
|
|
9
|
+
updated: 2026-08-24
|
|
10
|
+
owners: ["dfranks", "jcardinal", "mhammontree", "apeterson", "kyalamarthi", "bala"]
|
|
11
11
|
files:
|
|
12
12
|
- _underscore/Database.php
|
|
13
13
|
- _underscore/Model.php
|
|
@@ -177,6 +177,16 @@ here — they live in `Config/*.ini`.)
|
|
|
177
177
|
- Related 1.0 analogue: the legacy `App_` worker has the same hazard writing to `Logs.API`
|
|
178
178
|
(`db_logs`) — the laptop trap there is documented separately in the worker NetSuite bootstrap
|
|
179
179
|
notes.
|
|
180
|
+
- **`_Database::register()` is LAZY — it stores connection config and never opens a socket.**
|
|
181
|
+
The connect happens on first use. So adding a static alias registration to an app's `_.php`
|
|
182
|
+
bootstrap is **safe across every environment**, even one whose config group points at a
|
|
183
|
+
localhost that lacks the schema: nothing fails until something actually queries that alias.
|
|
184
|
+
This is what makes "register the alias globally, use it in one cron" a two-line change rather
|
|
185
|
+
than a per-environment config exercise (worked example: `DB_VISION_LOGS` in
|
|
186
|
+
[the Worker Fleet Role-Assignment Monitor](../../worker2/features/worker-fleet-role-assignment-monitor.md)).
|
|
187
|
+
The corollary is the trap the rest of this doc describes: because registration is free and
|
|
188
|
+
silent, a **bad** registration also stays silent until the request that needs it blows up with
|
|
189
|
+
`Unknown database`.
|
|
180
190
|
|
|
181
191
|
## WITHDRAWN — the "alias-keyed `$_modelCache` cross-tenant leak" hypothesis (2026-08-17)
|
|
182
192
|
|
|
@@ -210,6 +220,11 @@ also hits `_modules/<module>/`** — is recorded in
|
|
|
210
220
|
|
|
211
221
|
## Change history
|
|
212
222
|
|
|
223
|
+
- 2026-08-24 — Recorded that **`_Database::register()` is lazy** (stores config, opens no
|
|
224
|
+
socket; the connect happens on first use), so a static alias registration in an app's `_.php`
|
|
225
|
+
is safe in every environment including a localhost config that lacks the schema — and that
|
|
226
|
+
this is also why a *bad* registration stays silent until first query. Verified while adding
|
|
227
|
+
`DB_VISION_LOGS` to worker2. (bala)
|
|
213
228
|
- 2026-08-17 (later pass) — **WITHDRAWN, supersedes the entry below.** The alias-keyed
|
|
214
229
|
`$_modelCache` cross-tenant-leak hypothesis is **not supported** and is no longer a live security
|
|
215
230
|
concern. The six `c_` columns are **AIG's own**: `_modules/netsuite/2026-07-10a -
|
|
@@ -17,7 +17,7 @@
|
|
|
17
17
|
| [ClickUp Work Type Automation (Committed / Conditional / Stretch)](features/clickup-work-type-automation.md) | The ClickUp webhook handler (`_Worker_Clickup`) automatically maintains each task's **Work Type** custom field — `Committed`, `Conditional`, or `Stretch` — base | worker2/Worker/Clickup.php, worker2/Tests/Worker/ClickupWorkTypeTest.php |
|
|
18
18
|
| [Compass VIP Support Importer (worker2)](features/compass-vip-support-importer.md) | A worker2 action that ingests Compass's quarterly VIP spreadsheet and assigns each VIP user's support technician by setting `Users.c_supportedByUserId` in `Clie | worker2/Worker/Client/Compass/VipSupport.php |
|
|
19
19
|
| [Creating Worker Actions](features/creating-worker-actions.md) | How to add a new callable Worker action — a PHP class whose `public static` methods are invoked as background jobs (via webhook, cron, or `_Worker::runTask()`). | worker2/Worker/, worker2/Controller/Index.php, _underscore/Worker.php |
|
|
20
|
-
| [Cross-account AWS access for worker2 crons (_Component_Aws_Workloads)](features/cross-account-aws-access.md) | `_Component_Aws_Workloads` is the **standard, and only sanctioned, way any new worker2 cron obtains AWS access** — for any account, any region, any AWS SDK clie | worker2/Component/Aws/Workloads/Workloads.php, worker2/Config/production.ini, worker2/Worker/Infrastructure/CloudWatch.php |
|
|
20
|
+
| [Cross-account AWS access for worker2 crons (_Component_Aws_Workloads)](features/cross-account-aws-access.md) | `_Component_Aws_Workloads` is the **standard, and only sanctioned, way any new worker2 cron obtains AWS access** — for any account, any region, any AWS SDK clie | worker2/Component/Aws/Workloads/Workloads.php, worker2/Config/production.ini, worker2/Worker/Infrastructure/CloudWatch.php, worker2/Worker/Monitor/Fleet.php, library/app/worker.php, library/app/cloud.php |
|
|
21
21
|
| [Elastic Beanstalk health monitor → OneUptime (ElasticBeanstalkHealth)](features/elastic-beanstalk-health-monitor.md) | `_Worker_Infrastructure_CloudWatch::ElasticBeanstalkHealth()` is a worker2 cron that reads each Elastic Beanstalk (EB) environment's **enhanced-health** status | worker2/Worker/Infrastructure/CloudWatch.php |
|
|
22
22
|
| [Elite Freshservice Sync (worker2)](features/elite-freshservice-sync.md) | `_Worker_Elite` processes Freshservice webhook events and syncs them into TOGA 2. | worker2/Worker/Elite.php, worker2/Config/dev-kmaramreddy-laptop.ini |
|
|
23
23
|
| [Error Escalation Cron (Errors::Escalate → ClickUp / email)](features/error-escalation-cron.md) | `_Worker_Infrastructure_Errors::Escalate` (renamed from `SyncWithClickup`) is the sole owner of **escalation, de-escalation, ClickUp ticketing, reminders, busin | worker2/Worker/Infrastructure/Errors.php, worker2/Worker/Notification/Email.php, worker2/Worker/Notification/EmailTemplate.php, worker2/Worker/Client/True.php, worker2/Worker/Clickup/ErrorTask.php, worker2/Worker/Clickup.php, worker2/Controller/Index.php, worker2/Config/production.ini, _underscore/Model/Core/Logs/Issue.php, dbchanges2/Core/2026-07-30a - Error escalation cron job.sql, dbchanges2/Logs/2026-08-03a - Issue clickupPriority.sql, dbchanges2/Core/2026-08-04a - Error neglect digest cron job.sql |
|
|
@@ -47,6 +47,7 @@
|
|
|
47
47
|
| [Centralized Tracking-Status Refresh (worker2 Sync cron, all clients)](features/tracking-status-refresh.md) | A single platform-wide cron keeps `TrackingNumbers.status` current for **every provisioned client** until each shipment reaches a terminal state, by polling Fed | worker2/Worker/Sync/TrackingNumbers.php, _underscore/Component/Library/Carriers/Fedex/Fedex.php, _underscore/Component/Library/Carriers/Usps/Usps.php, _underscore/Component/Library/Carriers/Ups/Ups.php, _underscore/Model/Client/TrackingNumber.php, dbchanges2/Client/2026-07-30 - TrackingNumbers_RefreshColumns.sql, dbchanges2/Client/2026-07-30 - Acl_TrackingStatus_Grant.sql, dbchanges2/Client_Tdsynnex/2026-07-30 - Apis_TdsynnexKey.sql, dbchanges2/Core/2026-07-30 - TrackingRefresh_Cron.sql |
|
|
48
48
|
| [VAPI Webhook Handler (worker2 — AI-BDR end-of-call processing)](features/vapi-webhook-handler.md) | `_Worker_Vapi` ([worker2/Worker/Vapi.php](worker2/Worker/Vapi.php)) is the **PHP side of the AI-BDR call loop** — the webhook that receives VAPI's end-of-call r | worker2/Worker/Vapi.php, worker2/Worker/Ai/Bdr/Vapi.php, worker2/Controller/Index.php |
|
|
49
49
|
| [WJE Freshservice Sync (worker2)](features/wje-freshservice-sync.md) | WJE ("WJE IT", helpdesk `wje.freshservice.com`) is a **Freshservice**-based help-desk client whose tickets, contacts, assets, groups, categories, and canned res | worker2/Worker/Wje.php, _underscore/Component/Api/Wje/Wje.php, _underscore/Model/Wje/Ticket.php, _underscore/Model/Wje/TicketNote.php, _underscore/Model/Wje/Contact.php, _underscore/Model/Wje/Unit.php, _underscore/Model/Wje/TicketTeam.php, _underscore/Model/Wje/TicketCategory.php, _underscore/Model/Wje/AssetType.php, _underscore/Model/Wje/PredefinedReply.php, library/app/api/wje.php, worker/crons/toga2/wje/import_supporting_records.php, worker/crons/toga2/wje/sync_togasupply_wje.php, worker/crons/notifications/reports/wje/wje_common.php, library/app/systemmonitor/wje.php, dbchanges2/Client_Wje/2024-10-04 - WjeOnboarding.sql |
|
|
50
|
+
| [1.0 Worker Fleet Role-Assignment Monitor (Monitor/Fleet/RoleAssignment)](features/worker-fleet-role-assignment-monitor.md) | `_Worker_Monitor_Fleet::RoleAssignment()` is a **2.0 worker2 cron that watches the 1.0 `worker` fleet from the outside**. | worker2/Worker/Monitor/Fleet.php, worker2/_.php, dbchanges2/Core/2026-08-24 - Worker Role Assignment Monitor.sql, worker/ebs/cron.worker.php, worker/crons/worker/worker_heartbeat.php, library/app/worker.php |
|
|
50
51
|
| [PHP Runtime Upgrade on Elastic Beanstalk (worker2 8.3 → 8.5 + PhpSpreadsheet 1.x → 3.x)](workflows/php-runtime-upgrade-dependency-audit.md) | The procedure used to move worker2 from **PHP 8.3 to PHP 8.5** on Elastic Beanstalk, and the dependency work that had to land first. | worker2/composer.json, worker2/composer.lock, worker2/Worker/Team/Sprint.php, worker2/Worker/Client/TowFoundation/ProcessReceipts.php, worker2/Worker/Forecast/Import.php |
|
|
51
52
|
| [Running worker2 locally against real NetSuite (and what EB does instead)](workflows/running-worker2-locally.md) | How to boot **worker2** on a developer machine and run a real worker action against the **live NetSuite** account and a **local** `Forecast` database. | worker2/index.php, worker2/composer.json, worker2/.ebextensions/php_include_underscore.config, worker2/.ebextensions/git.json, worker2/.ebextensions/git.php, worker2/Config/production.ini, _underscore/Component/Api/Netsuite/Netsuite.php, dbchanges2/Logs_Client/2026-04-08_BLANK_CLIENT_LOGS_DATABASE.SQL |
|
|
52
53
|
| [Ticket → ClickUp Pseudocode Planning (Talos-grounded)](workflows/ticket-to-pseudocode-planning.md) | A repeatable procedure for turning a ClickUp ticket into a reviewed, formatted implementation plan posted back to the ticket's `📝 Pseudocode` custom field. | test/@dave/clickup_md2delta.js |
|
|
@@ -6,18 +6,22 @@ project: Worker
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-08-
|
|
10
|
-
owners: [jcardinal]
|
|
9
|
+
updated: 2026-08-24
|
|
10
|
+
owners: [jcardinal, bala]
|
|
11
11
|
files:
|
|
12
12
|
- worker2/Component/Aws/Workloads/Workloads.php
|
|
13
13
|
- worker2/Config/production.ini
|
|
14
14
|
- worker2/Worker/Infrastructure/CloudWatch.php
|
|
15
|
+
- worker2/Worker/Monitor/Fleet.php
|
|
16
|
+
- library/app/worker.php
|
|
17
|
+
- library/app/cloud.php
|
|
15
18
|
related:
|
|
16
19
|
- ../../_underscore/features/component-model-namespace-registration.md
|
|
17
20
|
- ../../_underscore/features/config-group-access.md
|
|
18
21
|
- ../../_underscore/features/cloud-s3-helpers.md
|
|
19
22
|
- ./oneuptime-worker2-monitoring.md
|
|
20
23
|
- ./elastic-beanstalk-health-monitor.md
|
|
24
|
+
- ./worker-fleet-role-assignment-monitor.md
|
|
21
25
|
---
|
|
22
26
|
|
|
23
27
|
## Summary
|
|
@@ -79,7 +83,8 @@ region is only a runtime argument to `client()`, never a reason for a second rol
|
|
|
79
83
|
- **654654170868** — production hub. Home region `us-east-1`; also used in `us-west-2` and
|
|
80
84
|
`eu-west-1`. worker2 runs here on Elastic Beanstalk; its instance profile role
|
|
81
85
|
`ElasticBeanstalk-EC2-Instance-Profile` is the caller identity.
|
|
82
|
-
- **502614707982** — legacy account, `us-west-2`.
|
|
86
|
+
- **502614707982** — legacy account, `us-west-2`. **This is where the 1.0 `worker` fleet
|
|
87
|
+
actually runs** — see the account-confusion gotcha below.
|
|
83
88
|
- **975050298201** — non-production account, `us-east-1`.
|
|
84
89
|
|
|
85
90
|
Each `WorkloadsRuntime` role has broad service permissions (`elasticbeanstalk:*`, `ec2:*`,
|
|
@@ -144,8 +149,33 @@ code yet. New code must not adopt the static key.
|
|
|
144
149
|
**Never record the access-key values in knowledge.** They exist in `production.ini`; that is
|
|
145
150
|
their only home.
|
|
146
151
|
|
|
152
|
+
### Which account the 1.0 worker fleet lives in (easy and expensive to get wrong)
|
|
153
|
+
|
|
154
|
+
The 1.0 `worker` EC2 instances are in **502614707982**, `us-west-2` (EB environment
|
|
155
|
+
`agilant-worker`, confirmed in the EC2 console).
|
|
156
|
+
|
|
157
|
+
**`App_Worker::AWS_WORKER_QUEUE_URL` in `library/app/worker.php` points at 654654170868 and
|
|
158
|
+
misleads** — that account holds only the **SQS queue the 1.0 workers read from**, not the
|
|
159
|
+
instances. Anything that needs to *see* the worker boxes (EB, EC2, health) must target
|
|
160
|
+
**502614707982**.
|
|
161
|
+
|
|
162
|
+
worker 1.0 authenticates there with a **long-lived hardcoded IAM access key in
|
|
163
|
+
`library/app/cloud.php`** — a committed credential that should be rotated and moved to a role
|
|
164
|
+
(value deliberately not recorded here; see [Legacy static key](#legacy-static-key--being-retired)).
|
|
165
|
+
worker2 does **not** use it: it assumes `WorkloadsRuntime` cross-account from its EB instance
|
|
166
|
+
role, which is exactly the migration this component exists to enable. First worker2 consumer in
|
|
167
|
+
that account is the
|
|
168
|
+
[Worker Fleet Role-Assignment Monitor](./worker-fleet-role-assignment-monitor.md).
|
|
169
|
+
|
|
170
|
+
**Scope EB reads by `EnvironmentName`, never account-wide.** 502614707982 also runs
|
|
171
|
+
`agilant-worker-alpha` and `agilant-worker-beta` as **separate EB environments that have no role
|
|
172
|
+
registry at all**. An account-level EC2 scan would sweep them in and misreport them.
|
|
173
|
+
|
|
147
174
|
## Gotchas / known issues
|
|
148
175
|
|
|
176
|
+
- **Do not assume the 1.0 worker fleet is in the 868 production hub.** `AWS_WORKER_QUEUE_URL`
|
|
177
|
+
(654654170868) is the SQS queue only; the instances are in the legacy account 502614707982,
|
|
178
|
+
`us-west-2`. See [the section above](#which-account-the-10-worker-fleet-lives-in-easy-and-expensive-to-get-wrong).
|
|
149
179
|
- Do not construct `StsClient` or pass explicit credentials by hand in new crons — the whole
|
|
150
180
|
point is that the instance role is the caller and the component owns the assume-role step.
|
|
151
181
|
- The instance-role-as-caller only works **on the EB tier**. Off-tier (e.g. a local CLI run)
|
|
@@ -161,6 +191,15 @@ their only home.
|
|
|
161
191
|
see [creating-worker-actions.md](./creating-worker-actions.md).
|
|
162
192
|
|
|
163
193
|
## Change history
|
|
194
|
+
- 2026-08-24 — Pinned down **which account the 1.0 `worker` fleet actually runs in**:
|
|
195
|
+
**502614707982** (`us-west-2`, EB env `agilant-worker`), **not** the 654654170868 production hub
|
|
196
|
+
that `App_Worker::AWS_WORKER_QUEUE_URL` implies — that account holds only the SQS queue the 1.0
|
|
197
|
+
workers read from. Recorded that worker 1.0 reaches it with the hardcoded static key in
|
|
198
|
+
`library/app/cloud.php` (rotate + move to a role) while worker2 assumes `WorkloadsRuntime` from
|
|
199
|
+
its instance role, and that EB reads there must be scoped by `EnvironmentName` because
|
|
200
|
+
`agilant-worker-alpha`/`-beta` are separate registry-less environments in the same account.
|
|
201
|
+
Found while building the
|
|
202
|
+
[Worker Fleet Role-Assignment Monitor](./worker-fleet-role-assignment-monitor.md). (bala)
|
|
164
203
|
- 2026-08-11 — Noted that `ElasticBeanstalkHealth`'s `$environments` must be typed `object|array`
|
|
165
204
|
(not `array`): the nested JSON object arrives as `stdClass` because the dispatcher casts only
|
|
166
205
|
the top level of `CronJobs.parameters`, so a strict `array` throws a `TypeError` at dispatch.
|
|
@@ -139,6 +139,35 @@ POSTed JSON body is addressed via the `requestBody` prefix — e.g.
|
|
|
139
139
|
- Docs: https://oneuptime.com/docs/monitor/incident-alert-templating and
|
|
140
140
|
https://oneuptime.com/docs/en/monitor/incoming-request-monitor
|
|
141
141
|
|
|
142
|
+
### The monitor IMPORT JSON — exact shape (verified 2026-08-24)
|
|
143
|
+
|
|
144
|
+
Monitors are provisioned by importing a **resource-export JSON**. The importer is strict and
|
|
145
|
+
its error messages are the only documentation, so the verified shape:
|
|
146
|
+
|
|
147
|
+
```
|
|
148
|
+
{ fileType: "oneuptime-resource-export", schemaVersion: 1, resourceType: "Monitor", items: [ … ] }
|
|
149
|
+
```
|
|
150
|
+
|
|
151
|
+
- **`monitorType` must be the UI DISPLAY name, with a space — `"Incoming Request"`, NOT
|
|
152
|
+
`"IncomingRequest"`.** The importer rejects the concatenated form and helpfully lists the
|
|
153
|
+
valid values in the error.
|
|
154
|
+
- Criteria nest deeply, each level carrying a `_type` tag:
|
|
155
|
+
`monitorSteps{_type:"MonitorSteps", value:{ monitorStepsInstanceArray:[ {_type:"MonitorStep",
|
|
156
|
+
value:{ monitorCriteria:{_type:"MonitorCriteria", value:{ monitorCriteriaInstanceArray:[ … ] }}}} ]}}`
|
|
157
|
+
- Each **criterion instance** carries `monitorStatusId`, `filterCondition` (`All`/`Any`),
|
|
158
|
+
`filters[]`, `incidents[]`, `isEnabled`, `name`, `description`.
|
|
159
|
+
- **Timing filters use `checkOn: "Incoming Request"`** with `filterType`
|
|
160
|
+
`"Not Recieved In Minutes"` / `"Recieved In Minutes"` (OneUptime's own misspelling) —
|
|
161
|
+
**not** `checkOn: "Request Body"`, which is only for the string-token matching above.
|
|
162
|
+
- `incidentSeverityId` is wrapped: `{_type: "ObjectID", value: "…"}`.
|
|
163
|
+
- There is **no `projectId`** in the export, and **status/severity ObjectIDs are
|
|
164
|
+
project-specific** — clone them out of an existing monitor's export rather than inventing them.
|
|
165
|
+
|
|
166
|
+
**Practical fix for the [activation-ordering trap](#gotchas--known-issues):** instead of
|
|
167
|
+
importing the monitor twice (body criteria first, absence criteria after the first push), import
|
|
168
|
+
the absence criteria **in the same file with `isEnabled: false`** and simply flip them on once a
|
|
169
|
+
push has landed.
|
|
170
|
+
|
|
142
171
|
### Heartbeat / cron cadence timing
|
|
143
172
|
|
|
144
173
|
The 10-min-Degraded / 15-min-Offline heartbeat thresholds pair with a **5-minute** cron
|
|
@@ -327,6 +356,16 @@ check for another client.
|
|
|
327
356
|
monitors must pass the region (see [Cloud S3 helpers](../../_underscore/features/cloud-s3-helpers.md)).
|
|
328
357
|
|
|
329
358
|
## Change history
|
|
359
|
+
- 2026-08-24 — Documented the **monitor import-JSON schema** the runbook previously described
|
|
360
|
+
only conceptually: the `oneuptime-resource-export` envelope, `monitorType` must be the
|
|
361
|
+
spaced display name `"Incoming Request"`, the `_type`-tagged
|
|
362
|
+
`monitorSteps → monitorStepsInstanceArray → monitorCriteria → monitorCriteriaInstanceArray`
|
|
363
|
+
nesting, per-criterion fields, timing filters on `checkOn: "Incoming Request"` (not
|
|
364
|
+
`"Request Body"`), the `ObjectID`-wrapped `incidentSeverityId`, and that status/severity
|
|
365
|
+
ObjectIDs are project-specific with no `projectId` in the export. Added the practical dodge for
|
|
366
|
+
the activation-ordering trap: import the absence criteria with `isEnabled: false` and flip them
|
|
367
|
+
on after the first push, instead of importing twice. Verified while provisioning the
|
|
368
|
+
[Worker Fleet Role-Assignment Monitor](./worker-fleet-role-assignment-monitor.md). (bala)
|
|
330
369
|
- 2026-08-20 — Added the **SFTP monitor** shape (phpseclib3 `rawlist` + age-cutoff count +
|
|
331
370
|
decided token; login/list failure → `status:error` connectivity check) and its two
|
|
332
371
|
runtime gotchas: `rawlist` returns **arrays** not objects, and the SFTP type constants are
|
|
@@ -0,0 +1,175 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: 1.0 Worker Fleet Role-Assignment Monitor (Monitor/Fleet/RoleAssignment)
|
|
3
|
+
framework: "2.0"
|
|
4
|
+
repo: worker2
|
|
5
|
+
project: Worker
|
|
6
|
+
client: shared
|
|
7
|
+
type: feature
|
|
8
|
+
status: active
|
|
9
|
+
updated: 2026-08-24
|
|
10
|
+
owners: ["bala"]
|
|
11
|
+
files:
|
|
12
|
+
- worker2/Worker/Monitor/Fleet.php
|
|
13
|
+
- worker2/_.php
|
|
14
|
+
- dbchanges2/Core/2026-08-24 - Worker Role Assignment Monitor.sql
|
|
15
|
+
- worker/ebs/cron.worker.php
|
|
16
|
+
- worker/crons/worker/worker_heartbeat.php
|
|
17
|
+
- library/app/worker.php
|
|
18
|
+
related:
|
|
19
|
+
- ./oneuptime-worker2-monitoring.md
|
|
20
|
+
- ./cross-account-aws-access.md
|
|
21
|
+
- ./elastic-beanstalk-health-monitor.md
|
|
22
|
+
- ../../../1.0/apps/worker/architecture.md
|
|
23
|
+
- ../../../1.0/apps/worker/features/oneuptime-worker-uptime-monitoring.md
|
|
24
|
+
---
|
|
25
|
+
|
|
26
|
+
## Summary
|
|
27
|
+
|
|
28
|
+
`_Worker_Monitor_Fleet::RoleAssignment()` is a **2.0 worker2 cron that watches the 1.0
|
|
29
|
+
`worker` fleet from the outside**. Every 5 minutes it lists the EC2 instances Elastic
|
|
30
|
+
Beanstalk reports for the `agilant-worker` environment and **compares that list against the
|
|
31
|
+
`Vision_Log.Workers` role registry**, then pushes three findings to a OneUptime Incoming
|
|
32
|
+
Request monitor:
|
|
33
|
+
|
|
34
|
+
| Finding | Meaning |
|
|
35
|
+
|---|---|
|
|
36
|
+
| `namelessInstances` | An instance is running in AWS but owns **no** row in `Workers` — it claimed no role, so it is running **no crons at all**. |
|
|
37
|
+
| `unownedRoles` | A role in `App_Worker::$desiredWorkers` that **nobody** currently owns. |
|
|
38
|
+
| `staleRoles` | A role whose `dtHeartbeat` has stopped advancing. |
|
|
39
|
+
|
|
40
|
+
**Why it exists:** 2 of the 7 production workers ran **roleless for roughly 3 hours with
|
|
41
|
+
nobody noticing**. Every pre-existing monitor is keyed by *role*, so an instance that never
|
|
42
|
+
got a role is invisible to all of them **by design** (see
|
|
43
|
+
[the blind spot](#the-blind-spot-why-none-of-the-14-existing-monitors-could-see-this)).
|
|
44
|
+
Detection is now about 10 minutes.
|
|
45
|
+
|
|
46
|
+
It follows the standard team token contract from
|
|
47
|
+
[OneUptime push-metric monitors](./oneuptime-worker2-monitoring.md): the worker decides,
|
|
48
|
+
OneUptime string-matches. `status = reporting | error`, `alarm = HIGH | OK`.
|
|
49
|
+
|
|
50
|
+
## Key files / entry points
|
|
51
|
+
|
|
52
|
+
- `worker2/Worker/Monitor/Fleet.php` — `_Worker_Monitor_Fleet`, action
|
|
53
|
+
**`Monitor/Fleet/RoleAssignment`** (routing verified: `Monitor/Fleet/RoleAssignment` maps to
|
|
54
|
+
`_Worker_Monitor_Fleet::RoleAssignment`).
|
|
55
|
+
- `worker2/_.php` — adds `const DB_VISION_LOGS = 'Vision_Log'` and its
|
|
56
|
+
`_Database::register()` call on the **`[database1]`** connection group.
|
|
57
|
+
- `dbchanges2/Core/2026-08-24 - Worker Role Assignment Monitor.sql` — the `Core.CronJobs`
|
|
58
|
+
row: `*/5 * * * *`, `parameters` NULL, **`isActive = 0`**.
|
|
59
|
+
- AWS access is obtained by assuming `WorkloadsRuntime` via
|
|
60
|
+
`_Component_Aws_Workloads::client(ElasticBeanstalkClient::class, ...)` — see
|
|
61
|
+
[Cross-account AWS access](./cross-account-aws-access.md). No static keys.
|
|
62
|
+
|
|
63
|
+
All tunables are **class constants with cron-parameter overrides**, so the schedule and the
|
|
64
|
+
thresholds live in the `CronJobs` row rather than in a deploy.
|
|
65
|
+
|
|
66
|
+
## How it works
|
|
67
|
+
|
|
68
|
+
1. Assume `WorkloadsRuntime` in the **legacy** account (`502614707982`, `us-west-2`) and call
|
|
69
|
+
**`describeInstancesHealth`** on the EB environment **`agilant-worker`**.
|
|
70
|
+
2. Read `Vision_Log.Workers` (`workerInstanceId` UNIQUE, `workerName` UNIQUE, `dtHeartbeat`)
|
|
71
|
+
through the new `DB_VISION_LOGS` alias.
|
|
72
|
+
3. Diff the two sets in both directions, plus a heartbeat-freshness pass, producing the three
|
|
73
|
+
findings above.
|
|
74
|
+
4. Apply the **grace period** (default 10 min): an instance that is still booting has not had
|
|
75
|
+
a chance to claim a role yet and must not page.
|
|
76
|
+
5. POST the decided tokens to the OneUptime monitor, non-fatal.
|
|
77
|
+
|
|
78
|
+
**Fail-safe:** a failed AWS call *or* a failed registry read pushes **`status = error`**, so a
|
|
79
|
+
blind checker is never mistaken for a healthy fleet. This is the same rule as
|
|
80
|
+
[the EB health monitor](./elastic-beanstalk-health-monitor.md) — a read that could not happen
|
|
81
|
+
is never reported as "fine."
|
|
82
|
+
|
|
83
|
+
### Why `Vision_Log` needed no new plumbing
|
|
84
|
+
|
|
85
|
+
worker2's **`[database1]` group already points at the legacy production cluster**
|
|
86
|
+
(cluster id `clwbyqvxdm4q`, `us-west-2`) — it is what already backs the `TOGaDeskSupport`
|
|
87
|
+
alias. `Vision_Log` lives on that **same single all-in-one legacy cluster** (verified: `Core`,
|
|
88
|
+
`TOGaDeskSupport`, `Vision` and `Vision_Log` are all present on it). So the whole database
|
|
89
|
+
change was **two lines in `_.php`**: no new credentials, no security-group change, no
|
|
90
|
+
`Core.Databases` rows.
|
|
91
|
+
|
|
92
|
+
Two facts that make that safe, both verified this session:
|
|
93
|
+
|
|
94
|
+
- All **9** `worker2/Config/*.ini` files define `[database1]`, so the registration resolves in
|
|
95
|
+
every environment.
|
|
96
|
+
- **`_Database::register()` is lazy** — it stores connection config and never opens a socket —
|
|
97
|
+
so a static registration in `_.php` costs nothing and cannot break a dev config whose
|
|
98
|
+
`[database1]` points at localhost.
|
|
99
|
+
|
|
100
|
+
## The blind spot (why none of the 14 existing monitors could see this)
|
|
101
|
+
|
|
102
|
+
This is the durable "why" the incident exposed, and the reason this monitor had to reach for
|
|
103
|
+
AWS rather than the database alone.
|
|
104
|
+
|
|
105
|
+
- 1.0 role assignment happens in `worker/ebs/cron.worker.php` and **only at DEPLOY time** (it
|
|
106
|
+
is written as the EB `appdeploy` enact hook). The script claims a free name, writes
|
|
107
|
+
`/etc/worker-role`, and **only then** appends that role's `cron.worker.<name>.json` entries.
|
|
108
|
+
- Therefore **no name means no crontab**. A roleless instance runs *nothing*, while EB happily
|
|
109
|
+
reports the environment healthy — the box is up, it just does no work.
|
|
110
|
+
- **All 14 existing OneUptime monitors are keyed by ROLE, not by instance.**
|
|
111
|
+
`crons/worker/worker_heartbeat.php` only pings the endpoint for the row matching **its own**
|
|
112
|
+
`instanceId`, and the `045_oneuptime_agent.config` infrastructure agent reads
|
|
113
|
+
`/etc/worker-role` and **exits clean when it is empty**. A roleless instance therefore
|
|
114
|
+
reports to nothing, and no role-keyed monitor is missing a ping.
|
|
115
|
+
- `Vision_Log.Workers` **only ever holds instances that SUCCEEDED** in claiming a role. The
|
|
116
|
+
failure leaves no row and no log line, so **absence is the only evidence** — and seeing an
|
|
117
|
+
absence requires the AWS instance list. That is precisely why this monitor exists in this
|
|
118
|
+
shape.
|
|
119
|
+
|
|
120
|
+
## Design history — built in worker2, not in worker 1.0
|
|
121
|
+
|
|
122
|
+
It was **first built in 1.0** (`worker` owns the `Vision_Log` connection natively), then moved
|
|
123
|
+
to worker2 and **the 1.0 version was deleted**. Recorded so it is not resurrected:
|
|
124
|
+
|
|
125
|
+
- worker2 already reaches the same cluster (above), so the 1.0 home bought nothing.
|
|
126
|
+
- worker2 uses the **EB instance role**; 1.0 would have used the hardcoded static AWS key.
|
|
127
|
+
- Schedule and parameters live in a **`CronJobs` row**, not in a deploy.
|
|
128
|
+
- Most importantly, it watches the 1.0 fleet **from outside it** instead of from a member of it.
|
|
129
|
+
|
|
130
|
+
The obvious objection — "a monitor running on the fleet it watches has a blind spot" — is
|
|
131
|
+
covered either way by OneUptime's own **absence criteria** (10 min Degraded / 15 min Offline),
|
|
132
|
+
which catch the monitor itself dying. The other three reasons decided it.
|
|
133
|
+
|
|
134
|
+
## Verification performed at build time
|
|
135
|
+
|
|
136
|
+
- **24/24 logic tests pass** against the real `Fleet.php` (private methods driven by
|
|
137
|
+
reflection): today's healthy fleet, a **replay of the actual incident**, the grace period,
|
|
138
|
+
stale heartbeats, a NULL-name spare row, a ghost row for a terminated instance, and an empty
|
|
139
|
+
registry.
|
|
140
|
+
- `Core.CronJobs` columns, uuid uniqueness, and the absence of a clashing `Monitor/Fleet%`
|
|
141
|
+
action all confirmed.
|
|
142
|
+
- **Unverified at capture time:** whether `WorkloadsRuntime` exists in `502614707982` with
|
|
143
|
+
`elasticbeanstalk:DescribeInstancesHealth`. If it does not, the monitor pushes
|
|
144
|
+
`status = error` — loud, not silent.
|
|
145
|
+
|
|
146
|
+
## Gotchas / known issues
|
|
147
|
+
|
|
148
|
+
- **The `CronJobs` row ships `isActive = 0` and must be flipped on after worker2 deploys.**
|
|
149
|
+
Follow the activation ordering in
|
|
150
|
+
[OneUptime push-metric monitors](./oneuptime-worker2-monitoring.md) — confirm a push has
|
|
151
|
+
landed before enabling the absence criteria.
|
|
152
|
+
- **Never widen the AWS scope to an account-level EC2 scan.** The same account also runs
|
|
153
|
+
`agilant-worker-alpha` and `agilant-worker-beta` as **separate EB environments with no role
|
|
154
|
+
registry at all**. `describeInstancesHealth` is scoped by `EnvironmentName`, which correctly
|
|
155
|
+
excludes them; an account-wide scan would flag every alpha/beta box as roleless.
|
|
156
|
+
- **Known false-positive risk, deliberately NOT guarded yet.**
|
|
157
|
+
`worker_heartbeat.php` terminates a stale instance and **then** deletes its `Workers` row. If
|
|
158
|
+
EB still lists that instance in `describeInstancesHealth` for a minute or two after the row is
|
|
159
|
+
gone, this monitor sees an instance with no role and raises `alarm = HIGH` — a possible false
|
|
160
|
+
page during a deploy or a recycle. It was not pre-solved because **how long EB keeps a
|
|
161
|
+
terminating instance in the health list is unverified**. If it shows up in practice, the fix is
|
|
162
|
+
a one-line filter on the instance's health status.
|
|
163
|
+
- **The OneUptime push URL is a credential** — never log it, never record its value.
|
|
164
|
+
- The class lives under `Worker/Monitor/` (singular), matching the OneUptime push monitors —
|
|
165
|
+
not the `Worker/Monitors/` (plural) orchestrator framework.
|
|
166
|
+
|
|
167
|
+
## Change history
|
|
168
|
+
- 2026-08-24 — Created. Added `_Worker_Monitor_Fleet::RoleAssignment()`
|
|
169
|
+
(`Monitor/Fleet/RoleAssignment`, `*/5 * * * *`, `isActive = 0` pending deploy) reporting
|
|
170
|
+
`namelessInstances` / `unownedRoles` / `staleRoles` to OneUptime, after 2 of 7 production
|
|
171
|
+
workers ran roleless for ~3 hours undetected. Registered `DB_VISION_LOGS` on the existing
|
|
172
|
+
`[database1]` legacy-cluster group (2 lines, no new credentials). Recorded the role-keyed
|
|
173
|
+
monitoring blind spot that made the incident invisible, the decision to build in worker2
|
|
174
|
+
rather than 1.0 (and delete the 1.0 version), and the unguarded terminating-instance
|
|
175
|
+
false-positive risk. worker2 `_production` e75690f, dbchanges2 `_main` 17732e2. (bala)
|
package/knowledge/INDEX.md
CHANGED
|
@@ -5,7 +5,7 @@ _Auto-generated by `knowledge.js index`. Do not hand-edit._
|
|
|
5
5
|
## 1.0 framework
|
|
6
6
|
|
|
7
7
|
- **library** (Library) _(framework core)_ — 19 doc(s) → [1.0/apps/library/INDEX.md](1.0/apps/library/INDEX.md)
|
|
8
|
-
- **worker** (Worker) —
|
|
8
|
+
- **worker** (Worker) — 28 doc(s) → [1.0/apps/worker/INDEX.md](1.0/apps/worker/INDEX.md)
|
|
9
9
|
- **dbchanges** (Database Changes) _(framework core)_ — 1 doc(s) → [1.0/apps/dbchanges/INDEX.md](1.0/apps/dbchanges/INDEX.md)
|
|
10
10
|
- **worker1.5** (Worker 1.5) — 0 doc(s) → [1.0/apps/worker1.5/INDEX.md](1.0/apps/worker1.5/INDEX.md)
|
|
11
11
|
- **togadesk** (TOGa Desk) — 13 doc(s) → [1.0/apps/togadesk/INDEX.md](1.0/apps/togadesk/INDEX.md)
|
|
@@ -19,7 +19,7 @@ _Auto-generated by `knowledge.js index`. Do not hand-edit._
|
|
|
19
19
|
## 2.0 framework
|
|
20
20
|
|
|
21
21
|
- **_underscore** (_Underscore) _(framework core)_ — 62 doc(s) → [2.0/apps/_underscore/INDEX.md](2.0/apps/_underscore/INDEX.md)
|
|
22
|
-
- **worker2** (Worker) —
|
|
22
|
+
- **worker2** (Worker) — 56 doc(s) → [2.0/apps/worker2/INDEX.md](2.0/apps/worker2/INDEX.md)
|
|
23
23
|
- **api2** (API) — 25 doc(s) → [2.0/apps/api2/INDEX.md](2.0/apps/api2/INDEX.md)
|
|
24
24
|
- **dbchanges2** (Database Changes) _(framework core)_ — 8 doc(s) → [2.0/apps/dbchanges2/INDEX.md](2.0/apps/dbchanges2/INDEX.md)
|
|
25
25
|
- **toga2-supply** (TOGa Supply) — 7 doc(s) → [2.0/apps/toga2-supply/INDEX.md](2.0/apps/toga2-supply/INDEX.md)
|
|
@@ -12,6 +12,7 @@
|
|
|
12
12
|
| [Compass MITS PO rejection — zero line items (400 + Issue + PM alert email)](features/mits-po-zero-line-item-rejection.md) | 2.0 | MITS transmits Compass purchase orders that carry **`"purchaseOrderItems": []`** (real example: PO `50309196-1`, `mitsSalesOrder` `MR244630`). | _underscore/Model/Compass/PurchaseOrder.php, _underscore/Model/Compass/Usa/PurchaseOrder.php, _underscore/Model/Compass/Canada/PurchaseOrder.php, worker2/Worker/Client/Compass/reports/PurchaseOrderRejectionEmail.php |
|
|
13
13
|
| [Compass MITS Sales-Order Transmission — Rejection Alerting (Issue/Event, not email-in-cron)](features/mits-sales-order-transmission-alerting.md) | 2.0 | When MITS **rejects** a Compass sales order transmitted by the 1.0 cron `1_transmit_compass_sales_orders_to_mits.php`, the alert is no longer an email built ins | worker/crons/toga2/compass/workflow/1_transmit_compass_sales_orders_to_mits.php, worker2/Worker/Infrastructure/Errors.php, _underscore/Model/Core/Logs/Issue.php, _underscore/Model/Core/Logs/IssueEmailAddress.php, dbchanges2/Logs/2026-08-04a - MITS rejection business recipients.sql |
|
|
14
14
|
| [Compass MR/MA Order Auto-Approval & Status Gate](features/mr-ma-order-approval-and-status.md) | 2.0 | Compass **MR** and **MA** sales orders are system-generated from the MITS / Office Depot EDI pipeline (they do not originate as user-entered SA orders) and must | _underscore/Model/Compass/SalesOrder.php, _underscore/Model/Compass/PurchaseOrder.php, worker/crons/toga2/compass/workflow/3a_import_office_depot_purchase_orders.php |
|
|
15
|
+
| [Compass ODP EDI 850 Line-Item Resolution (VA part number to IN SKU fallback)](features/odp-edi-850-item-resolution.md) | 1.0 | How each **PO1 line** on an inbound Office Depot (ODP) **EDI 850** is resolved to a real `Client_Compass` catalog item before cron `3a_import_office_depot_purch | worker/crons/toga2/compass/workflow/3a_import_office_depot_purchase_orders.php, library/app/Edi.php |
|
|
15
16
|
| [Compass Office Depot EDI 855 Acknowledgement + Over-Quantity PO Guard](features/odp-edi-855-acknowledgement-and-overquantity-guard.md) | 1.0 | How Compass acknowledges Office Depot (ODP) inbound **EDI 850** purchase orders with an **X12 855**, and the **over-quantity guard** that rejects a duplicate PO | worker/crons/toga2/compass/workflow/3a_import_office_depot_purchase_orders.php, worker/crons/toga2/compass/workflow/4_transmit_office_depot_po_acknowledgements.php, library/app/edi.php |
|
|
16
17
|
| [Compass PEOPLE-File User Lifecycle (duplicate accounts, reactivation grace window, raw-SQL deactivation)](features/people-file-user-lifecycle.md) | 2.0 | The nightly **PEOPLE** file cron (`_Worker_Client_Compass_PeopleFile`, `worker2/Worker/Client/Compass/PeopleFile.php`) owns the whole `Users` row lifecycle for | worker2/Worker/Client/Compass/PeopleFile.php |
|
|
17
18
|
| [Persona Model & Levy-Sector Gating (worker2 PEOPLE cron)](features/persona-model-and-levy-gating.md) | 2.0 | Compass USA catalogue visibility is driven by **personas** in `Client_Compass`. | worker2/Worker/Client/Compass/PeopleFile.php |
|
|
@@ -0,0 +1,179 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Compass ODP EDI 850 Line-Item Resolution (VA part number to IN SKU fallback)
|
|
3
|
+
framework: "1.0"
|
|
4
|
+
repo: worker
|
|
5
|
+
project: Worker
|
|
6
|
+
client: compass-usa
|
|
7
|
+
type: client-feature
|
|
8
|
+
status: active
|
|
9
|
+
updated: 2026-08-25
|
|
10
|
+
owners: ["bala"]
|
|
11
|
+
files:
|
|
12
|
+
- worker/crons/toga2/compass/workflow/3a_import_office_depot_purchase_orders.php
|
|
13
|
+
- library/app/Edi.php
|
|
14
|
+
related:
|
|
15
|
+
- odp-edi-855-acknowledgement-and-overquantity-guard.md
|
|
16
|
+
- ../workflows/odp-order-pipeline-to-netsuite.md
|
|
17
|
+
- ../workflows/odp-duplicate-po-line-cleanup.md
|
|
18
|
+
- ../workflows/odp-edi-import-recovery.md
|
|
19
|
+
- ../profile.md
|
|
20
|
+
---
|
|
21
|
+
|
|
22
|
+
## Summary
|
|
23
|
+
How each **PO1 line** on an inbound Office Depot (ODP) **EDI 850** is resolved to a real
|
|
24
|
+
`Client_Compass` catalog item before cron `3a_import_office_depot_purchase_orders.php` writes
|
|
25
|
+
anything. ODP sends **two** identifiers per line and only one of them is reliable:
|
|
26
|
+
|
|
27
|
+
| PO1 element | Qualifier | Parsed into | Reliability |
|
|
28
|
+
|---|---|---|---|
|
|
29
|
+
| `elements[7]` | `VA` | `partNumber` | **free text that drifts** — a mistyped character arrives verbatim |
|
|
30
|
+
| `elements[9]` | `IN` | `vendorPartNumber` (ODP's own SKU) | **correct every time observed** |
|
|
31
|
+
|
|
32
|
+
(Mapping verified in `App_Edi::translatePurchaseOrder850()`, the `case 'PO1'` block. So on a
|
|
33
|
+
Compass PO the `VendorItems.vendorPartNumber` value **is** ODP's SKU, and 3a's later
|
|
34
|
+
PO-item-to-SO-item matching already keyed on it — only sales-order-item creation trusted the VA
|
|
35
|
+
text.)
|
|
36
|
+
|
|
37
|
+
Before 2026-08-25 the importer created the ODP sales-order item from the **VA text only**, so a
|
|
38
|
+
single wrong character killed the whole PO import. Resolution now runs **first**, for **every
|
|
39
|
+
line**, and nothing is written unless all lines resolve.
|
|
40
|
+
|
|
41
|
+
## The resolution ladder (`resolveOfficeDepotEdiItemPartNumbers()`)
|
|
42
|
+
Runs on the parsed 850 **before** the duplicate-PO check, the over-quantity guard, and any
|
|
43
|
+
SalesOrder / PurchaseOrder / item creation.
|
|
44
|
+
|
|
45
|
+
1. **VA part number matches the Compass catalog** — use it as-is
|
|
46
|
+
(`partNumberSource = 'ediPartNumber'`).
|
|
47
|
+
2. **Fall back to ODP's own SKU** — `VendorItems.vendorPartNumber = <IN value>` with
|
|
48
|
+
`vendorId = App_Client_Compass::VENDOR_ID__OFFICE_DEPOT` gives `Items.partNumber`, which is
|
|
49
|
+
then **confirmed to be in the Compass catalog** (`partNumberSource = 'vendorSku'`).
|
|
50
|
+
Auto-resolves **only when exactly ONE catalog candidate remains.**
|
|
51
|
+
3. **Otherwise the line is unresolved** — it is reported with a plain-English reason and
|
|
52
|
+
**nothing at all is imported for that PO**.
|
|
53
|
+
|
|
54
|
+
Two api2 `GET`s back it, both batched per file (all lines' values in one `in` clause):
|
|
55
|
+
`/items` joined to `Catalogs` filtered on `Catalogs.uuid = UUID_ITEM_CATALOG`
|
|
56
|
+
(`961dbcbc-1943-11ef-8e3b-0aae2980db55`, the "Compass" catalog the cron posts against), and
|
|
57
|
+
`/vendor-items` joined to `Items` filtered on `vendorPartNumber IN (...)` plus the ODP vendor id.
|
|
58
|
+
All matching is **case-insensitive** (keys upper-cased, original casing returned).
|
|
59
|
+
|
|
60
|
+
### Why SKU resolution and not "clean up the string"
|
|
61
|
+
The real failures were not a single normalisable pattern:
|
|
62
|
+
|
|
63
|
+
| ODP PO | VA part number sent | Correct part number | ODP SKU (IN) |
|
|
64
|
+
|---|---|---|---|
|
|
65
|
+
| `41762948-1170` (2026-08-25) | `US-AND-CROTHALL-POM` (extra `M`) | `US-AND-CROTHALL-PO` | `4976158` |
|
|
66
|
+
| `41242800-1170`, `41243306-1170` (2026-06-17) | `ODTABKITTINGOPT3` (missing `P`) | `ODPTABKITTINGOPT3` | `4900941` |
|
|
67
|
+
|
|
68
|
+
A character-stripping / fuzzy-normalising rule fixes the first and **not** the June pair. The
|
|
69
|
+
SKU is the only identifier that was right in every case, so the SKU is the fallback key.
|
|
70
|
+
|
|
71
|
+
## Never create an item to match a bad inbound string
|
|
72
|
+
Hand-patching these failures in the past created **duplicate junk items in the Agilant catalog**
|
|
73
|
+
— e.g. `ODTABKITTINGOPT3` (item id 2715) now sits alongside the real `ODPTABKITTINGOPT3`
|
|
74
|
+
(id 2738). Creating an item so a mangled part number "matches" is the wrong fix: it splits one
|
|
75
|
+
physical SKU across two `Items` rows, which is the same shape as the
|
|
76
|
+
[dual-catalog duplicate PO-line bug](../workflows/odp-duplicate-po-line-cleanup.md). Fix the
|
|
77
|
+
`VendorItems` mapping (or have ODP correct their part number) instead.
|
|
78
|
+
|
|
79
|
+
## ODP VendorItems span two catalogs, and a few SKUs are ambiguous
|
|
80
|
+
Verified on prod `Client_Compass`, `VendorItems.vendorId = 1`: rows point at items in
|
|
81
|
+
**catalogId 1 ("Compass")**, **catalogId 2 ("Agilant")**, and **22 rows carry a null catalog**.
|
|
82
|
+
Three SKUs map to **more than one catalog-1 item**:
|
|
83
|
+
|
|
84
|
+
| ODP SKU | Catalog-1 candidates |
|
|
85
|
+
|---|---|
|
|
86
|
+
| `7128204` | `AW5M5UT` / `AW5M5UT-1` |
|
|
87
|
+
| `7423324` | `INC019BTBK` / `INC019BTBK-NS` |
|
|
88
|
+
| `8030904` | `C30705092` / `C30705092-DS` |
|
|
89
|
+
|
|
90
|
+
This is why step 2 requires **exactly one** candidate. An ambiguous SKU is **reported, never
|
|
91
|
+
guessed** — picking one silently would ship the wrong item. A SKU whose only match sits outside
|
|
92
|
+
catalog 1 is treated the same way (unresolved), which is what the extra catalog-confirmation
|
|
93
|
+
lookup is for.
|
|
94
|
+
|
|
95
|
+
## Nothing is written until every line resolves
|
|
96
|
+
The gate is a single `$shouldProcessOrder = empty($itemResolution['unresolvedItems'])` check at
|
|
97
|
+
the top of the per-PO loop, ahead of all writes.
|
|
98
|
+
|
|
99
|
+
**What it prevents:** the importer used to create the ODP sales order **first** and add items one
|
|
100
|
+
at a time, so a bad line left a **half-built order in production** — with no PO and no 855.
|
|
101
|
+
Exactly what happened to ODP SO `480238825001` on 2026-08-20: lines 1-3 created, line 4 failed,
|
|
102
|
+
no PurchaseOrder, no acknowledgement. Now a failed line leaves the PO completely untouched, so
|
|
103
|
+
re-dropping the 850 is a clean retry (see
|
|
104
|
+
[import recovery](../workflows/odp-edi-import-recovery.md)).
|
|
105
|
+
|
|
106
|
+
The resolved part numbers are also what the downstream checks now see:
|
|
107
|
+
- the existing `SKIP_IF_CONTAINS_PART_NUMBERS` skip list, and
|
|
108
|
+
- the **over-quantity guard** (`officeDepotPurchaseOrderExceedsCompassDemand()`), which keys on
|
|
109
|
+
`partNumber` and **silently skips any part it cannot match** — so before this change a mangled
|
|
110
|
+
part number bypassed duplicate-quantity protection entirely. See
|
|
111
|
+
[855 Acknowledgement + Over-Quantity Guard](odp-edi-855-acknowledgement-and-overquantity-guard.md).
|
|
112
|
+
|
|
113
|
+
## `ediPartNumber` — what ODP sent is kept alongside what we resolved
|
|
114
|
+
Every parsed item now carries **both**: `ediPartNumber` (ODP's original VA text) and `partNumber`
|
|
115
|
+
(the resolved catalog value), plus `partNumberSource` (`ediPartNumber` | `vendorSku` | `null`).
|
|
116
|
+
An EDI acknowledgement must reflect the partner's own values, so the **reject 855 is built from
|
|
117
|
+
`ediPartNumber`**.
|
|
118
|
+
|
|
119
|
+
> **The accept 855 still echoes OUR corrected part number.** Cron
|
|
120
|
+
> `4_transmit_office_depot_po_acknowledgements.php` reads `Items.partNumber` off the stored PO,
|
|
121
|
+
> so a SKU-corrected line acknowledges the **real Compass part number**, not the string ODP sent
|
|
122
|
+
> (PO109 still carries their SKU). Accepted deliberately on 2026-08-25. To change it,
|
|
123
|
+
> `ediPartNumber` has to be carried onto the PurchaseOrder record — it is not persisted today.
|
|
124
|
+
|
|
125
|
+
## A failed lookup degrades to the OLD behaviour, it does not block
|
|
126
|
+
If either api2 lookup returns `isSuccess = false`, `buildItemsUnchangedAfterFailedLookup()`
|
|
127
|
+
returns every line **exactly as ODP sent it** (`ediPartNumber` populated, `partNumber`
|
|
128
|
+
untouched, no unresolved lines) and the import proceeds as it did before this feature existed.
|
|
129
|
+
|
|
130
|
+
**Why:** blocking on a lookup failure would turn a one-line problem into *no ODP PO imports at
|
|
131
|
+
all*. Worst case now equals the old behaviour, never worse. The failed lookup still emails.
|
|
132
|
+
|
|
133
|
+
## The failure email (one sender, no template in the cron)
|
|
134
|
+
`sendErrorNotification(string $message, ?object $response = null)` is the **only** sender —
|
|
135
|
+
subject `Office Depot PO Import error: <PO numbers>`, the raw EDI attached as
|
|
136
|
+
`EDI_<PO numbers>.txt`. The API-response `print_r` dump is appended **only when a response is
|
|
137
|
+
passed**, so a business-readable message is no longer buried in a dump.
|
|
138
|
+
|
|
139
|
+
`buildItemResolutionErrorMessage()` builds that body: the Compass PO / Compass SO / ODP SO, then
|
|
140
|
+
per failed line the **ODP part number, the ODP SKU, the qty and the exact reason**, closing with
|
|
141
|
+
the fix (set the SKU up as an ODP vendor item against the correct Compass item, or have ODP
|
|
142
|
+
correct the part number, then re-drop the attached file into `OfficeDepot/` on `agilant-as2`).
|
|
143
|
+
All interpolated values pass through `htmlspecialchars()`.
|
|
144
|
+
|
|
145
|
+
**No email template belongs in the cron** — `App_Email_Agilant` already wraps the body via
|
|
146
|
+
`App_Email_Template::renderHtml()`. A second sender and a cron-local template were both
|
|
147
|
+
explicitly rejected.
|
|
148
|
+
|
|
149
|
+
## Helper map (all in cron 3a)
|
|
150
|
+
| Function | Job |
|
|
151
|
+
|---|---|
|
|
152
|
+
| `resolveOfficeDepotEdiItemPartNumbers()` | the ladder; returns `['items' => ..., 'unresolvedItems' => ...]` |
|
|
153
|
+
| `fetchCatalogPartNumbersByPartNumber()` | `/items` + `Catalogs.uuid` to an upper-cased part-number lookup (`null` = call failed) |
|
|
154
|
+
| `fetchPartNumbersByVendorSku()` | `/vendor-items` + ODP vendor id to SKU then **every** candidate part number |
|
|
155
|
+
| `filterToCatalogPartNumbers()` | keeps only candidates confirmed in the Compass catalog |
|
|
156
|
+
| `buildItemsUnchangedAfterFailedLookup()` | the degrade-to-old-behaviour path |
|
|
157
|
+
| `buildItemResolutionFailureReason()` | one plain sentence per failed line |
|
|
158
|
+
| `buildItemResolutionErrorMessage()` | the full email body |
|
|
159
|
+
| `uniqueNonEmptyValues()` | trim / drop blanks / de-dupe before an `in` clause |
|
|
160
|
+
|
|
161
|
+
## Verification (2026-08-25)
|
|
162
|
+
`php -l` clean. A standalone harness ran the real resolver functions against prod-shaped lookup
|
|
163
|
+
data — 20 checks covering the `41762948-1170` fix, the June `ODTABKITTINGOPT3` pair, an
|
|
164
|
+
ambiguous SKU, a SKU pointing outside the Compass catalog, a line with no SKU, and the
|
|
165
|
+
failed-lookup degradation. Prod after deploy: ODP `PurchaseOrders` id 112094 number
|
|
166
|
+
`41762948-1170` created 13:32:54 with all 4 items correct, `SalesOrderItems` line 4 id 613590
|
|
167
|
+
created as `US-AND-CROTHALL-PO`, lines 1-3 reused (not duplicated), `FileLog` 878902.
|
|
168
|
+
|
|
169
|
+
## Change history
|
|
170
|
+
- 2026-08-25 — Built the VA-to-IN resolution ladder so a drifted ODP part number no longer kills
|
|
171
|
+
a PO import: exact catalog match first, then ODP's own SKU through `VendorItems` (vendorId 1)
|
|
172
|
+
confirmed in the Compass catalog, auto-resolving only on a single candidate and reporting
|
|
173
|
+
ambiguity otherwise. Moved resolution **ahead of all writes** (a bad line used to leave a
|
|
174
|
+
half-built ODP SO — `480238825001`, 2026-08-20), fed the resolved part numbers to the skip list
|
|
175
|
+
and the over-quantity guard (which silently skips unmatched parts, so a mangled part number had
|
|
176
|
+
been bypassing duplicate protection), added `ediPartNumber` so the reject 855 echoes ODP's own
|
|
177
|
+
value, and made a failed lookup degrade to the pre-change behaviour rather than block. One
|
|
178
|
+
error sender only, no template in the cron. Recorded the dual-catalog `VendorItems` spread and
|
|
179
|
+
the 3 ambiguous SKUs. Deployed to prod and verified (commit `fc23bb46`). (bala)
|
package/knowledge/clients/compass-usa/features/odp-edi-855-acknowledgement-and-overquantity-guard.md
CHANGED
|
@@ -6,13 +6,14 @@ project: Worker
|
|
|
6
6
|
client: compass-usa
|
|
7
7
|
type: client-feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-08-
|
|
10
|
-
owners: ["jcardinal"]
|
|
9
|
+
updated: 2026-08-25
|
|
10
|
+
owners: ["jcardinal", "bala"]
|
|
11
11
|
files:
|
|
12
12
|
- worker/crons/toga2/compass/workflow/3a_import_office_depot_purchase_orders.php
|
|
13
13
|
- worker/crons/toga2/compass/workflow/4_transmit_office_depot_po_acknowledgements.php
|
|
14
14
|
- library/app/edi.php
|
|
15
15
|
related:
|
|
16
|
+
- odp-edi-850-item-resolution.md
|
|
16
17
|
- ../workflows/odp-order-pipeline-to-netsuite.md
|
|
17
18
|
- ../workflows/odp-duplicate-po-line-cleanup.md
|
|
18
19
|
- mits-po-transmission-to-vendors.md
|
|
@@ -65,6 +66,25 @@ The demand relay it walks is the standard ODP chain — Compass SO (top, e.g. `M
|
|
|
65
66
|
Compass PO → ODP SO → ODP PO (leaf) — via the `SalesOrders_PurchaseOrders` and
|
|
66
67
|
`PurchaseOrders_SalesOrders` bridges (the same 6-table join the acknowledgement cron uses).
|
|
67
68
|
|
|
69
|
+
### ⚠ The guard only works on RESOLVED part numbers (2026-08-25)
|
|
70
|
+
Because the guard keys on `partNumber` and **ignores any part it cannot match**, a part number
|
|
71
|
+
that ODP mistyped used to slip past duplicate-quantity protection entirely — the guard saw an
|
|
72
|
+
unknown part and passed the PO. Cron 3a now resolves every 850 line to a real Compass catalog part
|
|
73
|
+
number **before** the guard runs, and feeds the guard the resolved values. Anything that fails to
|
|
74
|
+
resolve stops the import outright, so an unmatched part can no longer reach the guard as a silent
|
|
75
|
+
pass. Ladder, ambiguity rules and the degrade-on-lookup-failure policy:
|
|
76
|
+
[ODP EDI 850 Line-Item Resolution](odp-edi-850-item-resolution.md).
|
|
77
|
+
|
|
78
|
+
## Which part number each 855 echoes
|
|
79
|
+
An acknowledgement must reflect the partner's own values, so the two paths now differ:
|
|
80
|
+
|
|
81
|
+
- **Reject 855 (from 3a)** — built from **`ediPartNumber`**, the raw `VA` part number ODP sent,
|
|
82
|
+
which the resolver preserves on every parsed item alongside the corrected `partNumber`.
|
|
83
|
+
- **Accept 855 (cron 4)** — reads `Items.partNumber` off the **stored** PO, so a SKU-corrected
|
|
84
|
+
line echoes back **our** part number, not ODP's. `PO109` still carries their SKU, so the line is
|
|
85
|
+
still identifiable to them. Accepted deliberately; carrying `ediPartNumber` onto the
|
|
86
|
+
PurchaseOrder record is what it would take to change (it is not persisted today).
|
|
87
|
+
|
|
68
88
|
## Why the reject is sent inline, not by cron 4
|
|
69
89
|
A rejected PO is **never persisted**, so cron 4 — which keys off `PurchaseOrders` rows with
|
|
70
90
|
`dtAcknowledged IS NULL` — can never see it. The importer therefore **builds and sends the
|
|
@@ -127,6 +147,13 @@ A legacy, abandoned sketch (`compass/edi/3_send_edi_855(reject).php`) used **non
|
|
|
127
147
|
codes `BAK01 = A1` / `BAK02 = RE` — that path was **deliberately NOT followed**.
|
|
128
148
|
|
|
129
149
|
## Change history
|
|
150
|
+
- 2026-08-25 — The over-quantity guard is now fed **resolved** part numbers (it ignores parts it
|
|
151
|
+
cannot match, so a mistyped ODP part number had been bypassing duplicate protection), the
|
|
152
|
+
**reject 855 echoes `ediPartNumber`** (what ODP sent) instead of our corrected value, and the
|
|
153
|
+
accept 855's use of our own `Items.partNumber` was recorded as a known, accepted difference.
|
|
154
|
+
Cron 3a's `sendErrorNotification()` became `(string $message, ?object $response = null)` so a
|
|
155
|
+
business-readable failure email is not buried in an API `print_r` dump. Details in the new
|
|
156
|
+
[ODP EDI 850 Line-Item Resolution](odp-edi-850-item-resolution.md). (bala)
|
|
130
157
|
- 2026-08-13 — Rejections now raise a business alert: cron `3a` throws
|
|
131
158
|
`App_Exception_Business` (`COMPASS_OFFICE_DEPOT_PO_IMPORT_REJECTED_OVER_QUANTITY`,
|
|
132
159
|
`URGENCY_HIGH`) so every over-quantity/duplicate reject surfaces as a Logs.Issue instead of
|
|
@@ -18,7 +18,7 @@ project: _Underscore
|
|
|
18
18
|
client: compass-usa
|
|
19
19
|
type: profile
|
|
20
20
|
status: active
|
|
21
|
-
updated: 2026-08-
|
|
21
|
+
updated: 2026-08-25
|
|
22
22
|
owners: [jcardinal, bala, tcox, apeterson, dfranks]
|
|
23
23
|
files: []
|
|
24
24
|
related:
|
|
@@ -34,6 +34,7 @@ related:
|
|
|
34
34
|
- features/stranded-approval-reassignment.md
|
|
35
35
|
- workflows/cross-kit-bundle-corruption.md
|
|
36
36
|
- workflows/odp-duplicate-po-line-cleanup.md
|
|
37
|
+
- features/odp-edi-850-item-resolution.md
|
|
37
38
|
- features/odp-edi-855-acknowledgement-and-overquantity-guard.md
|
|
38
39
|
- ../../2.0/apps/worker2/features/compass-vip-support-importer.md
|
|
39
40
|
- ../../2.0/apps/toga2-commerce/features/expedited-shipping-gating.md
|
|
@@ -99,6 +100,14 @@ separate, related client (see its own profile).
|
|
|
99
100
|
rejects the whole PO and sends a reject **EDI 855**; ODP POs are acknowledged with an 855
|
|
100
101
|
(accept via cron 4, reject inline from 3a) built by a shared `App_Edi` method. See
|
|
101
102
|
[ODP EDI 855 Acknowledgement + Over-Quantity Guard](features/odp-edi-855-acknowledgement-and-overquantity-guard.md).
|
|
103
|
+
- **ODP's part numbers on the 850 drift; their SKU does not (2026-08-25).** A `PO1` line carries a
|
|
104
|
+
free-text `VA` part number **and** ODP's own `IN` SKU. A single mistyped character in the `VA`
|
|
105
|
+
text used to kill the whole PO import (and half-build the ODP sales order). Cron 3a now resolves
|
|
106
|
+
every line — catalog match first, then the SKU through `VendorItems` (vendorId 1) — before it
|
|
107
|
+
writes anything, and never guesses on an ambiguous SKU:
|
|
108
|
+
[ODP EDI 850 Line-Item Resolution](features/odp-edi-850-item-resolution.md).
|
|
109
|
+
**Never create an `Items` row so a mangled inbound part number "matches"** — that is what put
|
|
110
|
+
duplicate junk items in the Agilant catalog (`ODTABKITTINGOPT3` vs `ODPTABKITTINGOPT3`).
|
|
102
111
|
- ASN ingestion entry points: cXML to the V2 API (logged in `Logs_Compass.Api`) and
|
|
103
112
|
`worker/crons/toga2/compass/workflow/3b_import_strategic_systems_advance_shipping_notices.php`.
|
|
104
113
|
|
|
@@ -6,7 +6,7 @@ project: Worker
|
|
|
6
6
|
client: compass-usa
|
|
7
7
|
type: workflow
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-08-
|
|
9
|
+
updated: 2026-08-25
|
|
10
10
|
owners: ["rgirish", "bala", "dfranks", "jcardinal", "mhammontree"]
|
|
11
11
|
files:
|
|
12
12
|
- worker/crons/toga2/compass/workflow/1_transmit_compass_sales_orders_to_mits.php
|
|
@@ -19,7 +19,9 @@ files:
|
|
|
19
19
|
- library/app/client/compass.php
|
|
20
20
|
related:
|
|
21
21
|
- 1.0/apps/worker/features/netsuite-sales-order-sales-rep-sourcing.md
|
|
22
|
+
- clients/compass-usa/features/odp-edi-850-item-resolution.md
|
|
22
23
|
- clients/compass-usa/features/odp-edi-855-acknowledgement-and-overquantity-guard.md
|
|
24
|
+
- 1.0/apps/worker/workflows/tracing-a-worker-cron-run-in-production.md
|
|
23
25
|
- clients/compass-usa/workflows/odp-edi-import-recovery.md
|
|
24
26
|
- clients/compass-usa/workflows/order-lifecycle-and-data-integrity.md
|
|
25
27
|
- clients/compass-usa/features/mits-po-transmission-to-vendors.md
|
|
@@ -97,6 +99,27 @@ duplicates and backfills whatever is missing**:
|
|
|
97
99
|
`customerPurchaseOrder` stays `NULL` on a repaired order and must be set manually. See
|
|
98
100
|
[Recovering a lost ODP EDI import](odp-edi-import-recovery.md).
|
|
99
101
|
|
|
102
|
+
### Every 850 line must resolve to a catalog item before anything is written (2026-08-25)
|
|
103
|
+
3a resolves **all** PO1 lines to real `Client_Compass` catalog part numbers as the **first** thing
|
|
104
|
+
it does per PO — ahead of the duplicate-PO check, the over-quantity guard and every write. If any
|
|
105
|
+
line fails, **nothing is created for that PO** and the failure is emailed with the EDI attached.
|
|
106
|
+
|
|
107
|
+
Before this, 3a created the ODP SalesOrder first and added items one at a time, so one bad line
|
|
108
|
+
left a **half-built order in prod** with no PurchaseOrder and no 855 (ODP SO `480238825001`,
|
|
109
|
+
2026-08-20). Combined with the idempotency table above, a failed import is now always a clean
|
|
110
|
+
re-drop rather than a repair job. Ladder, ambiguity rules and the failed-lookup fallback:
|
|
111
|
+
[ODP EDI 850 Line-Item Resolution](../features/odp-edi-850-item-resolution.md).
|
|
112
|
+
|
|
113
|
+
### ⚠ Known issue (NOT fixed): the `OfficeDepot/` listing makes 3a runs overlap
|
|
114
|
+
An instrumented prod run on 2026-08-25 showed the S3 listing returning **138,472 objects** and
|
|
115
|
+
taking **32 seconds** before any PO was processed, with each PO then costing roughly 40 seconds.
|
|
116
|
+
A file-bearing run therefore blows past the `*/5` schedule and the following runs are **silently
|
|
117
|
+
skipped** by `App_Framework::exitIfProcessRunning()` — leaving **no `CronJobExecutions` row at
|
|
118
|
+
all**, which reads as "the cron never fired". The object count is driven by `OfficeDepot/SENT/`
|
|
119
|
+
and `OfficeDepot/OUTBOX/` accumulating: 3a skips those prefixes when *processing* but still
|
|
120
|
+
*lists* them. Fix candidates (separate work): prune `SENT/`, or list with a tighter prefix. See
|
|
121
|
+
[Tracing a worker cron run in production](../../1.0/apps/worker/workflows/tracing-a-worker-cron-run-in-production.md).
|
|
122
|
+
|
|
100
123
|
### The EDI 850 → TOGa identifier mapping (not guessable — this is the key to tracing an ODP order)
|
|
101
124
|
|
|
102
125
|
| 850 segment | Meaning | Where it lands |
|
|
@@ -105,7 +128,14 @@ duplicates and backfills whatever is missing**:
|
|
|
105
128
|
| **`REF~QC`** | the **ODP SALES ORDER number** (e.g. `475390815001`) | `SalesOrders.number` with `customerId = 1` |
|
|
106
129
|
| **`REF~EU` / `REF~PO`** | the **Compass MITS PO number** (e.g. `50310002-1`) | its `c_mitsSalesOrder` field holds the ODP SO number |
|
|
107
130
|
| **`REF~LU`** | the **Compass sales order** (e.g. `SA135273`) | `SalesOrders` with `customerId = 2` |
|
|
108
|
-
| **`PO1`** | the line | qty, UOM, unit price,
|
|
131
|
+
| **`PO1`** | the line | qty, UOM, unit price, and **two** item identifiers (below) |
|
|
132
|
+
|
|
133
|
+
**`PO1` carries two item identifiers and only one of them is trustworthy:** `elements[7]` (the
|
|
134
|
+
`VA` qualifier) is parsed into **`partNumber`** and is **free text that drifts**; `elements[9]`
|
|
135
|
+
(the `IN` qualifier) is parsed into **`vendorPartNumber`** and is **ODP's own SKU**, correct in
|
|
136
|
+
every observed case. So on a Compass PO the `VendorItems.vendorPartNumber` value *is* ODP's SKU.
|
|
137
|
+
Cron 3a resolves each line through that SKU when the `VA` text does not match the catalog — see
|
|
138
|
+
[ODP EDI 850 Line-Item Resolution](../features/odp-edi-850-item-resolution.md).
|
|
109
139
|
|
|
110
140
|
**⚠ `REF~QC` — not `BEG` — is the value to search on for the ODP SalesOrder.** Searching
|
|
111
141
|
`PurchaseOrders` / `SalesOrders` for the `BEG` number finds nothing for the SO and wastes time.
|
|
@@ -242,6 +272,13 @@ wrong record makes every order look "stuck." Use the **join**, never a number ma
|
|
|
242
272
|
transmission feature.)
|
|
243
273
|
|
|
244
274
|
## Change history
|
|
275
|
+
- 2026-08-25 — Cron 3a now **resolves every 850 line to a catalog item before it writes anything**
|
|
276
|
+
(a bad line used to leave a half-built ODP SO); documented the `PO1` two-identifier detail
|
|
277
|
+
(`elements[7]` = drifting `VA` text, `elements[9]` = ODP's reliable `IN` SKU) behind the new
|
|
278
|
+
[850 item-resolution feature](../features/odp-edi-850-item-resolution.md). Flagged the
|
|
279
|
+
**unfixed** `OfficeDepot/` listing problem: 138,472 objects / 32 s per run pushes 3a past its
|
|
280
|
+
`*/5` slot, so following runs are skipped by the overlap guard with **no `CronJobExecutions`
|
|
281
|
+
row** to show it. (bala)
|
|
245
282
|
- 2026-08-21 — **Cron 5 sales rep** no longer hardcodes employee 34877: it reads the end
|
|
246
283
|
customer's native NetSuite `salesRep` RecordRef (`App_NetSuite::getCustomer()`, cached per
|
|
247
284
|
run, failures logged and treated as "no rep", left unset when absent), with the end-customer
|
package/package.json
CHANGED