toga-ai 1.0.508 → 1.0.510
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/knowledge/1.0/apps/library/INDEX.md +1 -1
- package/knowledge/1.0/apps/library/features/error-capture-1-0.md +47 -5
- package/knowledge/1.0/apps/library/features/toga2-api-client-and-bridge.md +19 -1
- package/knowledge/1.0/apps/tools/features/errors-curation-console.md +57 -0
- package/knowledge/2.0/apps/_underscore/features/error-reporting-issue-event.md +53 -4
- package/knowledge/2.0/apps/api2/features/api-payload-interceptors.md +64 -1
- package/knowledge/2.0/apps/api2/features/request-logging.md +31 -1
- package/knowledge/2.0/apps/worker2/INDEX.md +1 -1
- package/knowledge/2.0/apps/worker2/features/error-escalation-cron.md +85 -0
- package/knowledge/2.0/apps/worker2/features/notification-email.md +22 -3
- package/knowledge/INDEX.md +1 -1
- package/knowledge/clients/compass-usa/INDEX.md +2 -1
- package/knowledge/clients/compass-usa/workflows/odp-edi-import-recovery.md +111 -0
- package/knowledge/clients/compass-usa/workflows/odp-order-pipeline-to-netsuite.md +70 -9
- package/knowledge/clients/compass-usa/workflows/order-lifecycle-and-data-integrity.md +44 -2
- package/package.json +1 -1
|
@@ -8,7 +8,7 @@
|
|
|
8
8
|
| [Diagnostic Dialog — View Recommended Services Routing](features/diagnostic-dialog-view-recommended-services.md) | Two "View Recommended Services" buttons exist in the TOGa Refresh 2026 SR view: 1. | library/app/model/toga/diagnostic.php, library/app/model/servicerequest.php |
|
|
9
9
|
| [Elite Freshservice Sync (library)](features/elite-freshservice-sync.md) | `App_Api_Toga2` in `library/app/api/toga2.php` orchestrates bidirectional sync between TOGA 2 and TOGaDesk. | library/app/api/toga2.php |
|
|
10
10
|
| [Branded HTML Email Templates (App_Email_Template)](features/email-templates.md) | `App_Email_Template` (`app/email/template.php`) is the base class for branded HTML emails in the 1.0 (`App_`) framework. | library/app/email/template.php, library/app/email/agilant.php |
|
|
11
|
-
| [Error Capture in 1.0 (App_Error_Capture → shared 2.0 Logs DB)](features/error-capture-1-0.md) | The 1.0 side of the platform error-reporting pipeline (TRUE-78188). | library/app/error/capture.php, library/app/error.php, library/app/exception/business.php, library/app/cloud.php, worker/config.worker.ini, worker/crons/toga2/compass/workflow/1_transmit_compass_sales_orders_to_mits.php |
|
|
11
|
+
| [Error Capture in 1.0 (App_Error_Capture → shared 2.0 Logs DB)](features/error-capture-1-0.md) | The 1.0 side of the platform error-reporting pipeline (TRUE-78188). | library/app/error/capture.php, library/app/error.php, library/app/exception/business.php, library/app/api/toga2.php, library/app/cloud.php, worker/config.worker.ini, worker/crons/toga2/compass/workflow/1_transmit_compass_sales_orders_to_mits.php |
|
|
12
12
|
| [1.0 MVC Page Pattern & New-App Skeleton](features/mvc-page-pattern-and-app-skeleton.md) | This is the **reusable recipe for standing up a new 1.0 (`App_`) application** and for adding pages to one — the folder-based MVC routing, the page lifecycle, t | library/app/framework.php, library/app/frameworkindex.php, library/app/mvc.php, library/app/database.php, library/app/model.php, library/app/config.php |
|
|
13
13
|
| [isFulfillable from NetSuite during Item Sync (Phase 1)](features/netsuite-item-isfulfillable-sync.md) | This is the **1.0 (Phase 1)** half of the `isFulfillable` feature: reading the NetSuite `isfulfillable` flag during item sync and stamping it onto the **Agilant | library/app/netsuite.php, library/app/api/toga2.php, worker/crons/toga2/netsuite/common_sync_togasupply.php, worker/crons/toga2/netsuite/backfill_isfulfillable_jul5.php |
|
|
14
14
|
| [NetSuite SuiteQL/REST API Reference](features/netsuite-suiteql-api-reference.md) | General working reference for the Agilant NetSuite integration: how to authenticate, how SuiteQL behaves, and the confirmed schema of the tables/columns/codes w | library/app/api/netsuite/rest.php, library/ssl/netsuite_ec_key.pem, test/@dave/Junk Drawer/nsq.php |
|
|
@@ -12,11 +12,13 @@ files:
|
|
|
12
12
|
- library/app/error/capture.php
|
|
13
13
|
- library/app/error.php
|
|
14
14
|
- library/app/exception/business.php
|
|
15
|
+
- library/app/api/toga2.php
|
|
15
16
|
- library/app/cloud.php
|
|
16
17
|
- worker/config.worker.ini
|
|
17
18
|
- worker/crons/toga2/compass/workflow/1_transmit_compass_sales_orders_to_mits.php
|
|
18
19
|
related:
|
|
19
20
|
- ../architecture.md
|
|
21
|
+
- ./toga2-api-client-and-bridge.md
|
|
20
22
|
- ../../../2.0/apps/_underscore/features/error-reporting-issue-event.md
|
|
21
23
|
- ../../../2.0/apps/worker2/features/error-escalation-cron.md
|
|
22
24
|
- ../../tools/features/errors-curation-console.md
|
|
@@ -76,13 +78,42 @@ whichever occurrence happened to be seen first — i.e. one arbitrary order numb
|
|
|
76
78
|
|
|
77
79
|
Compass USA's sales-order → MITS transmit rejection was the first business-exception use case.
|
|
78
80
|
|
|
79
|
-
### Client attribution
|
|
81
|
+
### Client attribution — resolved from the API client uuid at `App_Api_Toga2::authenticate()`
|
|
80
82
|
|
|
81
83
|
`App_Error::setCurrentClientId()` exists in 1.0 for parity with 2.0's ambient current-client, but
|
|
82
|
-
**1.0 has no central hook** equivalent to `_Database::registerClientDatabases()`.
|
|
83
|
-
|
|
84
|
-
`
|
|
85
|
-
|
|
84
|
+
**1.0 has no central hook** equivalent to `_Database::registerClientDatabases()`. Because of that,
|
|
85
|
+
the explicit setter was only ever called by **one** hand-edited Compass cron, so
|
|
86
|
+
**every other 1.0 cron recorded `clientId` NULL**. Evidence: `Logs.Event` held **2,093 rows with
|
|
87
|
+
`clientId` NULL** against **107** with a client (all of them 2.0/api2 JWT traffic); issue `1Q`
|
|
88
|
+
traced back to the 1.0 cron `crons/toga2/quad/import_po_and_asn.php`.
|
|
89
|
+
|
|
90
|
+
**`App_Api_Toga2::authenticate()` is 1.0's true choke point.** `send()` authenticates before every
|
|
91
|
+
call, and every 1.0→2.0 cron arrives there holding its `UUID_CLIENT`. It now calls
|
|
92
|
+
`App_Error_Capture::setCurrentClientUuid($uuid)` — a **store only, no lookup**, because
|
|
93
|
+
`authenticate()` is a happy-path function and must not take on a query.
|
|
94
|
+
|
|
95
|
+
`App_Error_Capture` then resolves uuid → `Core.Clients.id` **lazily**, only when an error is
|
|
96
|
+
actually captured, cached per process (`resolveClientIdFromUuid()`, private). It walks
|
|
97
|
+
**`CORE_DATABASE_LINK_CANDIDATES = ['db_toga2core', 'db_prod_toga2core', 'db_core']`** because 1.0
|
|
98
|
+
apps register the same Core connection under different link names. **Every failure path returns
|
|
99
|
+
`null` and records no client rather than throwing.**
|
|
100
|
+
|
|
101
|
+
Precedence is unchanged: an explicit `clientId` argument wins, then an explicit
|
|
102
|
+
`setCurrentClientId()`, then the ambient uuid.
|
|
103
|
+
|
|
104
|
+
> **⚠ The resolution must run INSIDE the existing
|
|
105
|
+
> `App_Error::setThrowExceptionsEnabled(false)` guard.** An unconfigured DB link **warns** before
|
|
106
|
+
> it fails, and in 1.0 a warning outside that guard terminates the cron (see the first gotcha
|
|
107
|
+
> below). Do not hoist the lookup out of the guarded block.
|
|
108
|
+
|
|
109
|
+
### The reference encoder is a byte-identical mirror of 2.0's
|
|
110
|
+
|
|
111
|
+
`capture.php` carries its own copy of `encodeReference()` plus `REFERENCE_ALPHABET`
|
|
112
|
+
(`'0123456789ACDEFGHJKMNPQRSTUVWXYZ'`) and `REFERENCE_OFFSET` (`128`), matching
|
|
113
|
+
`_Model_Core_Logs_Issue` **byte for byte**. Both write the same **UNIQUE** `Issue.reference`
|
|
114
|
+
column in the same shared table, so a divergence between them is a failed INSERT and a **lost**
|
|
115
|
+
error. Change one, change both, deploy together. Rationale for the alphabet and the offset lives in
|
|
116
|
+
the [2.0 doc](../../../2.0/apps/_underscore/features/error-reporting-issue-event.md#the-quotable-reference--base-32-over-an-unambiguous-alphabet-reworked-2026-08-04).
|
|
86
117
|
|
|
87
118
|
### Fingerprint fallback (no usable stack)
|
|
88
119
|
|
|
@@ -151,6 +182,17 @@ exists in `api2`/`worker2`. So 1.0 reads the AL1 container config at
|
|
|
151
182
|
|
|
152
183
|
## Change history
|
|
153
184
|
|
|
185
|
+
- 2026-08-04 (later, **uncommitted/undeployed** at time of writing) — **Fixed: 1.0 errors recorded
|
|
186
|
+
`clientId` NULL on every cron but one** (2,093 NULL `Logs.Event` rows vs 107 with a client; issue
|
|
187
|
+
`1Q` from `crons/toga2/quad/import_po_and_asn.php`). `App_Api_Toga2::authenticate()` — 1.0's real
|
|
188
|
+
choke point, since `send()` authenticates before every call — now records the ambient client uuid
|
|
189
|
+
(store only, no lookup), and `App_Error_Capture` resolves uuid → `Core.Clients.id` lazily and
|
|
190
|
+
per-process on capture, walking `CORE_DATABASE_LINK_CANDIDATES`
|
|
191
|
+
(`db_toga2core`/`db_prod_toga2core`/`db_core`) because 1.0 apps name the Core link differently;
|
|
192
|
+
every failure path returns null rather than throwing, and the lookup deliberately runs inside the
|
|
193
|
+
`setThrowExceptionsEnabled(false)` guard because an unconfigured link *warns* before it fails.
|
|
194
|
+
Also mirrored 2.0's **base-32 reference encoder** (`REFERENCE_ALPHABET` + `REFERENCE_OFFSET = 128`)
|
|
195
|
+
byte-for-byte. (jcardinal)
|
|
154
196
|
- 2026-08-04 — Built the 1.0 side of TRUE-78188: `App_Error_Capture` writing occurrences into the
|
|
155
197
|
shared 2.0 `Logs` DB (opt-in via `[database_toga2logs]`, hand-written escaped SQL);
|
|
156
198
|
`App_Exception_Business` with a fingerprint string byte-identical to 2.0's so one `issueKey`
|
|
@@ -6,7 +6,7 @@ project: Library
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-08-
|
|
9
|
+
updated: 2026-08-04
|
|
10
10
|
owners: [jcardinal, mhammontree]
|
|
11
11
|
files:
|
|
12
12
|
- library/app/api/toga2.php
|
|
@@ -22,6 +22,7 @@ related:
|
|
|
22
22
|
- netsuite-suiteql-rest-shim.md
|
|
23
23
|
- ../../worker/features/netsuite-togasupply-per-client-sync.md
|
|
24
24
|
- ../architecture.md
|
|
25
|
+
- ./error-capture-1-0.md
|
|
25
26
|
---
|
|
26
27
|
|
|
27
28
|
## Summary
|
|
@@ -77,6 +78,18 @@ belong to the NetSuite importer — documented in the per-client-sync doc, not h
|
|
|
77
78
|
- Returns the decoded response object; callers read `->data->{resource}`, `->meta->nextPage`,
|
|
78
79
|
`->isSuccess`, `->status`, `->messages[].code`.
|
|
79
80
|
|
|
81
|
+
> **`authenticate()` is also 1.0's ambient-client choke point for error reporting** (added
|
|
82
|
+
> 2026-08-04). Every 1.0→2.0 cron passes through it holding its `UUID_CLIENT`, so it calls
|
|
83
|
+
> `App_Error_Capture::setCurrentClientUuid()` — **store only, no lookup**: this is a happy-path
|
|
84
|
+
> function and must not acquire a query. The uuid → `Core.Clients.id` resolution happens lazily
|
|
85
|
+
> inside error capture. Keep it that way; see
|
|
86
|
+
> [error-capture-1-0](./error-capture-1-0.md).
|
|
87
|
+
|
|
88
|
+
> **Latent, pre-existing, uninvestigated:** several `self::$togaClientUuid` references sit in
|
|
89
|
+
> methods of this class (roughly lines 2880–4250) with **no such property declared** — an
|
|
90
|
+
> undefined-property bug waiting to fire (and in 1.0, a warning inside an unguarded path
|
|
91
|
+
> terminates the cron). Not introduced by, and not fixed by, the 2026-08-04 error-reporting work.
|
|
92
|
+
|
|
80
93
|
### Options DSL (how 1.0 expresses 2.0 queries)
|
|
81
94
|
|
|
82
95
|
`$options` is an assoc array assembled into the 2.0 query string:
|
|
@@ -340,6 +353,11 @@ enable flags** and an optional `$monitorTogadeskDepartmentIds[]`:
|
|
|
340
353
|
|
|
341
354
|
## Change history
|
|
342
355
|
|
|
356
|
+
- 2026-08-04 (**uncommitted/undeployed** at time of writing) — `authenticate()` now records the
|
|
357
|
+
ambient client uuid for error reporting (`App_Error_Capture::setCurrentClientUuid()`, store-only)
|
|
358
|
+
because it is the one place every 1.0→2.0 cron passes through holding its `UUID_CLIENT` — this is
|
|
359
|
+
what fixed 1.0 errors recording `clientId` NULL. Also noted the pre-existing undeclared
|
|
360
|
+
`self::$togaClientUuid` references (~L2880–4250) as a latent undefined-property bug. (jcardinal)
|
|
343
361
|
- 2026-08-03 (later pass) — TRUE-79401 **static regression harness + one correction.** BUILT a
|
|
344
362
|
no-DB PHP harness (`test/@Mark/AIG/test_multi_email.php`, 12 cases, PHP 7.2.33, `php -l` clean)
|
|
345
363
|
that feeds mocked `/contacts` fragments through the email transform and asserts the delimited
|
|
@@ -72,6 +72,48 @@ curated text, affected clients) right, with an always-present Type chip.
|
|
|
72
72
|
time.
|
|
73
73
|
- Hourly-occurrence chart; fingerprints table re-sized for the narrow column with an icon-only
|
|
74
74
|
Move button.
|
|
75
|
+
- Header chip order is **Issue / Environment / Type / Urgency** — Type was moved ahead of Urgency so
|
|
76
|
+
urgency finishes the row at the far right.
|
|
77
|
+
|
|
78
|
+
### Table presentation rules (established 2026-08-04)
|
|
79
|
+
|
|
80
|
+
**Stacked datetimes in every table on both pages** — listing *Last seen*; detail *Events → When*,
|
|
81
|
+
*Affected clients → Last*, *ClickUp episodes → Opened/Resolved*, *Fingerprints → First seen*. A
|
|
82
|
+
shared `$datetime` closure emits
|
|
83
|
+
`<span class="dt"><span class="dt-date">…</span><span class="dt-time">…</span></span>`, each part
|
|
84
|
+
`nowrap`, split on the single space MySQL puts between date and time; a blank value renders an em
|
|
85
|
+
dash. Tabular figures keep digits aligned down the column.
|
|
86
|
+
|
|
87
|
+
> **Why:** letting a datetime wrap naturally let the browser pick its own break point, and the point
|
|
88
|
+
> it picked was **inside the day of the month** (`2026-08-` / `04`). This replaced the old
|
|
89
|
+
> `.wrap-date` / `.wrap-anywhere` approach, whose now-dead CSS rule was removed.
|
|
90
|
+
|
|
91
|
+
**The summary column has a 100px floor via table overflow, not a cell `min-width`.** The summary
|
|
92
|
+
`<col>` is deliberately left **unwidthed** so it absorbs leftover space — which is also why *every
|
|
93
|
+
other* column must declare a width, or `col-type` absorbs it instead. The floor comes from
|
|
94
|
+
`.issue-listing { min-width: 1026px }` inside a new `.table-scroll { overflow-x: auto }` wrapper.
|
|
95
|
+
|
|
96
|
+
> **⚠ Under `table-layout: fixed` a browser compresses ALL columns proportionally at narrow
|
|
97
|
+
> viewports and ignores `min-width` on a cell.** A cell-level `min-width` therefore cannot work —
|
|
98
|
+
> the table has to overflow and scroll. `col-when` narrowed 168px → 104px now that dates stack (the
|
|
99
|
+
> 64px went to the summary), and a duplicate `.col-type` rule was removed in favour of the
|
|
100
|
+
> pre-existing 104px one. **The pager sits outside the scroll wrapper.**
|
|
101
|
+
|
|
102
|
+
**Whole-row click replaced the two per-cell links** (Issue ID and Summary). The `<tr>` carries
|
|
103
|
+
`class="row-link"`, `data-href`, `tabindex="0"` and an `aria-label`; `.issue-id-link` was renamed
|
|
104
|
+
`.issue-id` (keeping the accent colour as the affordance) and `tr.row-link` got cursor/hover/
|
|
105
|
+
`focus-visible` rules.
|
|
106
|
+
|
|
107
|
+
This page was deliberately **JavaScript-free**, and this is the one thing that cannot be done
|
|
108
|
+
without it: no element but an anchor is clickable, and an anchor cannot wrap a `<tr>`. A small
|
|
109
|
+
**delegated** click+keydown handler was added. Ctrl/Cmd/Shift-click still opens a new tab;
|
|
110
|
+
Enter/Space work because removing the anchors removed the tab stop; clicks on `a`, `button`,
|
|
111
|
+
`input`, `select`, `textarea`, `summary` or `label` are ignored so future in-row controls keep their
|
|
112
|
+
own behaviour.
|
|
113
|
+
|
|
114
|
+
> **Known trade-off:** a screen reader now announces these as **labelled rows, not links**.
|
|
115
|
+
> `role="link"` was deliberately **not** applied, because it would stop the element being a table
|
|
116
|
+
> row.
|
|
75
117
|
|
|
76
118
|
**This page notifies nobody.** Escalation, ClickUp ticketing, and email are owned entirely by
|
|
77
119
|
`_Worker_Infrastructure_Errors::Escalate`. Adding a notify action here would create a second,
|
|
@@ -128,6 +170,21 @@ They were renamed from plural on 2026-08-01, after the pipeline was already in p
|
|
|
128
170
|
|
|
129
171
|
## Change history
|
|
130
172
|
|
|
173
|
+
- 2026-08-04 (later, **uncommitted/undeployed** at time of writing) — UI pass on both `/errors`
|
|
174
|
+
pages: **stacked datetimes** in every table via a shared `$datetime` closure emitting
|
|
175
|
+
`.dt`/`.dt-date`/`.dt-time` (a naturally wrapping datetime broke inside the day of the month), and
|
|
176
|
+
the dead `.wrap-date`/`.wrap-anywhere` rules removed; a **100px floor on the summary column**
|
|
177
|
+
delivered as `.issue-listing { min-width: 1026px }` inside a new `.table-scroll` overflow wrapper
|
|
178
|
+
(with the pager outside it), because `table-layout: fixed` ignores a cell `min-width` and
|
|
179
|
+
compresses all columns proportionally — `col-when` narrowed 168px → 104px and a duplicate
|
|
180
|
+
`.col-type` rule removed; the two per-cell links replaced by a **whole-row click**
|
|
181
|
+
(`tr.row-link` + `data-href` + `tabindex` + `aria-label` and one small delegated click/keydown
|
|
182
|
+
handler — the single unavoidable use of JS on this deliberately JS-free page, since an anchor
|
|
183
|
+
cannot wrap a `<tr>`), honouring Ctrl/Cmd/Shift-click and ignoring clicks on interactive
|
|
184
|
+
descendants, with the accepted trade-off that screen readers announce labelled rows rather than
|
|
185
|
+
links (`role="link"` would stop the element being a row); and issue-detail header order changed to
|
|
186
|
+
Issue / Environment / Type / Urgency. Note **`library` and `tools` must both deploy** for any of
|
|
187
|
+
this to be visible. (jcardinal)
|
|
131
188
|
- 2026-08-04 — Console build-out: the listing gained real pagination (25/page with a filtered
|
|
132
189
|
`COUNT`, replacing a silent `LIMIT 200`), an EB Environment column **and filter**, a
|
|
133
190
|
BUSINESS/TECHNICAL Type column derived from `issueKey` (no new DB column), and a 3-column filter
|
|
@@ -190,15 +190,47 @@ A uniform shape is what lets a viewer render context **generically** (see the To
|
|
|
190
190
|
card/JSON-tree renderer). Consumers must read **both** the old and new key positions during the
|
|
191
191
|
7-day `Event` retention overlap.
|
|
192
192
|
|
|
193
|
-
### The quotable reference
|
|
193
|
+
### The quotable reference — base-32 over an unambiguous alphabet (reworked 2026-08-04)
|
|
194
194
|
|
|
195
|
-
`
|
|
196
|
-
|
|
197
|
-
|
|
195
|
+
`Issue.reference` encodes `Issue.id` in **base 32** over
|
|
196
|
+
`REFERENCE_ALPHABET = '0123456789ACDEFGHJKMNPQRSTUVWXYZ'`, most-significant character first, no
|
|
197
|
+
padding (`_Model_Core_Logs_Issue::encodeReference()`). **All ten digits are kept; `O`, `I`, `L`
|
|
198
|
+
and `B` are removed** because in the contexts a reference actually gets used — read aloud on a
|
|
199
|
+
call, typed from a screenshot, pasted into a ClickUp title — they impersonate `0`, `1`, `1` and
|
|
200
|
+
`8`. A reference exists to be quoted, so ambiguity is the one defect it cannot have.
|
|
201
|
+
|
|
202
|
+
It replaced a decimal-counter + letter, base-26 scheme (`1A`, `2Q`, `3O`).
|
|
203
|
+
|
|
204
|
+
> **⚠ `REFERENCE_OFFSET = 128` exists to stop the new scheme from OVERWRITING the old one —
|
|
205
|
+
> do not remove it.** `Issue.reference` is **UNIQUE**, and every legacy value begins `1`–`3`.
|
|
206
|
+
> With no offset, new id 75 encodes to `2A`, which id 27 already holds — the INSERT fails and the
|
|
207
|
+
> error is **lost instead of recorded**, precisely the failure mode this pipeline exists to
|
|
208
|
+
> prevent. The offset burns the first 128 (= 4 × 32) code points so every new reference starts at
|
|
209
|
+
> `4` or later and both generations coexist forever. Verified by script over 200,000 ids: no
|
|
210
|
+
> collisions among themselves, none against the legacy values, no banned characters.
|
|
211
|
+
|
|
212
|
+
The column is `varchar(16)`, so the growth to 3 characters (≈ id 897) and 4 (≈ id 29,000) cannot
|
|
213
|
+
truncate. Because `AUTO_INCREMENT` never reuses ids, gaps are expected and permanent — a reference
|
|
214
|
+
can never be recycled onto a different problem.
|
|
215
|
+
|
|
216
|
+
**The 1.0 encoder in `library/app/error/capture.php` must stay byte-identical** — both write the
|
|
217
|
+
same UNIQUE column in the same shared table. Change one, change the other, deploy together.
|
|
198
218
|
|
|
199
219
|
The reference is written by PHP in a second statement (see MySQL constraints below), which is
|
|
200
220
|
why the column is nullable even though it is always populated.
|
|
201
221
|
|
|
222
|
+
#### Existing data is NOT being truncated (decided 2026-08-04)
|
|
223
|
+
|
|
224
|
+
Nothing needs deleting to adopt the new scheme. References are **stored, never re-derived from
|
|
225
|
+
the id at read time**: the listing search, the Tools Move-fingerprint prompt (which accepts a
|
|
226
|
+
reference *or* a numeric id) and every display path read the stored string, so both generations
|
|
227
|
+
work side by side. Accepted cost: 10 of the 68 existing issues keep ambiguous references (`1B`,
|
|
228
|
+
`1I`, `2B`, `2I`, `2L`, `2O`, `3B`, `3I`, `3L`, `3O`) until they age out through the normal GC.
|
|
229
|
+
|
|
230
|
+
> **If you ever do truncate `Logs.Issue`:** `TRUNCATE` **resets `AUTO_INCREMENT`**, and
|
|
231
|
+
> `IssueFingerprint`, `IssueEmailAddress`, `IssueClickupTask` and `Event` all reference
|
|
232
|
+
> `Issue.id` — they must be cleared in the **same** operation or they orphan onto reused ids.
|
|
233
|
+
|
|
202
234
|
### Shutdown handler
|
|
203
235
|
|
|
204
236
|
`register_shutdown_function` is registered with a **pre-allocated memory reserve**. Before this,
|
|
@@ -249,6 +281,13 @@ DB, so the pipeline is no longer 2.0-only; see
|
|
|
249
281
|
has to opt in by adding a `[database_toga2logs]` config section, so Sentry must remain for any 1.0
|
|
250
282
|
app that has not been wired up yet.
|
|
251
283
|
|
|
284
|
+
## Deploy order for a change that spans this pipeline
|
|
285
|
+
|
|
286
|
+
`_underscore` **first** (api2 and worker2 clone it at prebuild), then api2, worker2, `library`,
|
|
287
|
+
`worker`, `tools`. Anything touching the reference encoder or a shared `Logs` table name must land
|
|
288
|
+
in `_underscore` and `library` in the **same** release, and the `/errors` console work needs
|
|
289
|
+
**both** `library` and `tools` deployed before it is visible.
|
|
290
|
+
|
|
252
291
|
## Gotchas / known issues
|
|
253
292
|
|
|
254
293
|
- **Read-your-writes: reads go to the READ host and cannot see an uncommitted write.** A
|
|
@@ -379,6 +418,16 @@ clientUserId). **Neither was built.** As built instead:
|
|
|
379
418
|
|
|
380
419
|
## Change history
|
|
381
420
|
|
|
421
|
+
- 2026-08-04 (later, **uncommitted/undeployed** at time of writing) — **Reference encoding reworked
|
|
422
|
+
to base 32** over `'0123456789ACDEFGHJKMNPQRSTUVWXYZ'` (all digits kept; `O`/`I`/`L`/`B` dropped
|
|
423
|
+
as impersonators of `0`/`1`/`1`/`8`), replacing the base-26 decimal-counter+letter scheme, with
|
|
424
|
+
**`REFERENCE_OFFSET = 128`** so new values start at `4` and cannot collide with the legacy `1`–`3`
|
|
425
|
+
values on the UNIQUE column — a collision would fail the INSERT and *lose* the error. Verified
|
|
426
|
+
over 200,000 ids. The 1.0 encoder in `library/app/error/capture.php` mirrors it byte-for-byte.
|
|
427
|
+
**Decided: existing Issue/Event data is not truncated** — references are stored and never
|
|
428
|
+
re-derived, so both generations coexist; 10 of 68 issues keep ambiguous references until GC.
|
|
429
|
+
Recorded the TRUNCATE warning (it resets `AUTO_INCREMENT` and four child tables reference
|
|
430
|
+
`Issue.id`). Next production reference will be `68` (max id 72). (jcardinal)
|
|
382
431
|
- 2026-08-04 — **Decided: `_Exception_Business` extends `_Exception_Validation`** so api2 answers
|
|
383
432
|
4xx for a business condition while still recording it — with the ordering trap that every
|
|
384
433
|
`catch (_Exception_Validation …)` must now test business **first** (api2's controller and the
|
|
@@ -6,7 +6,7 @@ project: API
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-08-
|
|
9
|
+
updated: 2026-08-04
|
|
10
10
|
owners: ["mhammontree", "dfranks"]
|
|
11
11
|
files:
|
|
12
12
|
- api2/Component/Api/V2/V2.php
|
|
@@ -82,6 +82,50 @@ tenant's other APIs at all. See
|
|
|
82
82
|
> result above is authoritative; the code path is misleading. Anyone tempted to reason it out
|
|
83
83
|
> from V2.php will reach the wrong answer.
|
|
84
84
|
|
|
85
|
+
## The dispatch is UNGUARDED — a row naming a method that does not exist hard-fatals the endpoint
|
|
86
|
+
|
|
87
|
+
The derived name (step 2 above) is **pure convention with nothing validating it**, and there is
|
|
88
|
+
**no `phpMethod` column in production** to override it. All **four** call sites in
|
|
89
|
+
`api2/Component/Api/V2/V2.php` invoke the derived name **without a `method_exists()` check**:
|
|
90
|
+
|
|
91
|
+
| Site | Phase |
|
|
92
|
+
|------|-------|
|
|
93
|
+
| `~L8655` | builds the interceptor lookup |
|
|
94
|
+
| `~L3148` | pre-processing dispatch |
|
|
95
|
+
| `~L5694` | post-processing dispatch |
|
|
96
|
+
| `~L7411` / `~L7437` | alternate pre-processing path |
|
|
97
|
+
|
|
98
|
+
So a **single `ApiPayloadInterceptors` row whose (recordId, phase, method) resolves to a method the
|
|
99
|
+
model does not define takes the endpoint down** with a PHP fatal —
|
|
100
|
+
`Call to undefined method _Model_<Slug>_<Model>::postPost()` — not a handled 4xx/5xx. This is the
|
|
101
|
+
mirror image of the "missing row" failure mode above: a *missing* row is silent, an *extra/wrong*
|
|
102
|
+
row is fatal.
|
|
103
|
+
|
|
104
|
+
**Fix direction:** guard all four sites with `method_exists()` and raise a logged `Logs.Issue`
|
|
105
|
+
instead of letting the request fatal.
|
|
106
|
+
|
|
107
|
+
### The client-model resolution that decides which class must define the method
|
|
108
|
+
|
|
109
|
+
At `~L5681-5690` the engine derives the concrete class by `str_replace`-ing `_Model_Client_` →
|
|
110
|
+
`_Model_<jwt client slug>_` on `$record->model`, guarded by `class_exists`. For Compass USA that
|
|
111
|
+
yields `_Model_Compass_Usa_*`.
|
|
112
|
+
|
|
113
|
+
The chain is `_Model_Compass_Usa_X extends _Model_Compass_X extends _Model_Client_X`, so **a method
|
|
114
|
+
defined on the shared Compass base is inherited by both the Usa and Canada subclasses** — a
|
|
115
|
+
per-region subclass does *not* need its own copy. Do not add duplicate hooks per region.
|
|
116
|
+
|
|
117
|
+
### `Core.Records` facts for interceptor debugging
|
|
118
|
+
|
|
119
|
+
- **`Core.Records.model` is UNIQUE** — exactly one record row per model class, so an interceptor's
|
|
120
|
+
`recordId` maps to precisely one model. No ambiguity to resolve.
|
|
121
|
+
- The column is **`route`** (not `routePlural` / `routeSingular`) in the current schema. Example:
|
|
122
|
+
`recordId 15` = route `sales-order-items` = `_Model_Client_SalesOrderItem`.
|
|
123
|
+
- **Diagnostic:** join the client's active interceptor rows to `Core.Records` and check each derived
|
|
124
|
+
method actually exists in `_underscore/Model/<Slug>/`. Compass's legitimate rows (records **14**
|
|
125
|
+
sales-orders, **17** purchase-orders, **19** vendor-items, **21** items, **37** contacts, **55**
|
|
126
|
+
advance-shipping-notices, **179** approval-decisions) all map 1:1 to real methods — **the row whose
|
|
127
|
+
record has no corresponding method is the anomaly.**
|
|
128
|
+
|
|
85
129
|
## Worked example — the EV-10 that was not a code bug
|
|
86
130
|
|
|
87
131
|
`_Model_Client_ItemFulfillment::prePost()` defaults `itemFulfillmentStageId` to the shipped stage on
|
|
@@ -99,6 +143,14 @@ and the failing environment**. It is a small table, and the drift is usually exa
|
|
|
99
143
|
- **⚠ A hook that "isn't running" is a missing registration row before it is a code bug.** Check
|
|
100
144
|
`ApiPayloadInterceptors` for `(recordId, phase, method)` **first** — never start by re-reading the
|
|
101
145
|
PHP.
|
|
146
|
+
- **⚠ An interceptor row for a method that does not exist is a production outage, not a no-op.**
|
|
147
|
+
The dispatch is unguarded at all four call sites — the endpoint fatals with *"Call to undefined
|
|
148
|
+
method"*. If a write endpoint suddenly 500s for one client only, list that client's
|
|
149
|
+
`ApiPayloadInterceptors` rows and confirm every derived method exists.
|
|
150
|
+
- **⚠ A migration that enables an interceptor must insert `isActive = 0`.** Activation is a *data*
|
|
151
|
+
change that switches on a *code* path; if the PHP defining the method is not confirmed deployed to
|
|
152
|
+
that environment, the row takes the endpoint down. Insert inactive, verify the deploy, then flip
|
|
153
|
+
`isActive = 1` — and remember every environment activates independently.
|
|
102
154
|
- **The method name is derived, so a typo'd enum silently misses.** A row with
|
|
103
155
|
`prePostProcessing = 'PRE'`, `httpMethod = 'PUT'` resolves to `prePut`, not `prePost`; the engine
|
|
104
156
|
will simply find no method and move on.
|
|
@@ -114,6 +166,17 @@ and the failing environment**. It is a small table, and the drift is usually exa
|
|
|
114
166
|
|
|
115
167
|
## Change history
|
|
116
168
|
|
|
169
|
+
- 2026-08-04 — Documented the **inverse failure mode**: the derived-name dispatch is **unguarded at
|
|
170
|
+
all four `V2.php` call sites** (`~L8655` lookup, `~L3148` pre, `~L5694` post, `~L7411`/`~L7437`
|
|
171
|
+
alternate pre) and there is **no `phpMethod` column in production**, so one row naming a
|
|
172
|
+
nonexistent method hard-fatals the endpoint (*"Call to undefined method
|
|
173
|
+
`_Model_Compass_Usa_SalesOrderItem::postPost()`"*). Added the client-model resolution at
|
|
174
|
+
`~L5681-5690` (`_Model_Client_` → `_Model_<slug>_`, `class_exists`-guarded) and the
|
|
175
|
+
`Usa → Compass → Client` inheritance rule (define once on the Compass base; no per-region copy),
|
|
176
|
+
the `Core.Records` debugging facts (`model` is UNIQUE, column is `route`, record 15 =
|
|
177
|
+
`sales-order-items`, Compass's 1:1 legitimate row set), and the rule that a migration enabling an
|
|
178
|
+
interceptor must land `isActive = 0` first. Diagnosed from a production incident; no code change.
|
|
179
|
+
(dfranks)
|
|
117
180
|
- 2026-08-03 — Added two verified behaviours from a live dev probe: (1) interceptor registration
|
|
118
181
|
is **per-API** (`apiId` on the row, in both the Core and `Client_<X>` tables), so a guard in a
|
|
119
182
|
hook is scoped by **DB config rather than code** and a new row widens its blast radius; (2) the
|
|
@@ -6,7 +6,7 @@ project: API
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-08-
|
|
9
|
+
updated: 2026-08-04
|
|
10
10
|
owners: ["mhammontree", "dfranks"]
|
|
11
11
|
files:
|
|
12
12
|
- api2/Component/Api/V2/V2.php
|
|
@@ -15,7 +15,9 @@ files:
|
|
|
15
15
|
- _underscore/Model/Core/Logs/Api.php
|
|
16
16
|
related:
|
|
17
17
|
- ../architecture.md
|
|
18
|
+
- ./api-payload-interceptors.md
|
|
18
19
|
- ../../../../clients/aig/features/entitlement-intake.md
|
|
20
|
+
- ../../../../clients/compass-usa/workflows/order-lifecycle-and-data-integrity.md
|
|
19
21
|
---
|
|
20
22
|
|
|
21
23
|
## Summary
|
|
@@ -61,6 +63,28 @@ That is the deciding factor when choosing between `_Exception_Validation` (400,
|
|
|
61
63
|
request, no Issue) and `_Exception_Business` (500, Sentry + a routed Issue). Do **not** reach for
|
|
62
64
|
`_Exception_Business` merely to "make sure the rejection is recorded" — it already is.
|
|
63
65
|
|
|
66
|
+
### ⚠ But an uncaught PHP **fatal** can leave NO log row anywhere
|
|
67
|
+
|
|
68
|
+
The commit-the-logs / roll-back-the-data contract above holds for the **handled** paths — a
|
|
69
|
+
validation rejection or an exception the controller catches. It did **not** hold for an **uncaught
|
|
70
|
+
`Error`** observed in production on 2026-08-04 (a `Call to undefined method` raised from an
|
|
71
|
+
[API payload interceptor](api-payload-interceptors.md) dispatch): the fatal was reported to the
|
|
72
|
+
error tracker, but **`Logs_Compass.Api` contained zero 5xx rows for the entire day and the failing
|
|
73
|
+
request appears nowhere in it at all.** Suspected regression in the recent `Error.php` /
|
|
74
|
+
Issue-Event table normalization work.
|
|
75
|
+
|
|
76
|
+
Practical consequences:
|
|
77
|
+
|
|
78
|
+
- **Do not conclude "the request never happened" from an empty log table.** For a fatal, absence of
|
|
79
|
+
a log row is not evidence of absence of the request. Check the error tracker's per-event detail —
|
|
80
|
+
that may be the *only* record of the individual request.
|
|
81
|
+
- **Diagnose from the DATA, not the log.** The data transaction rolls back while a *previous*
|
|
82
|
+
successful request already committed its own rows, so a fatal on a multi-request write sequence
|
|
83
|
+
leaves a detectable **orphan / zero-child parent**. Sweep for those instead of grepping logs — see
|
|
84
|
+
[Compass Order Lifecycle & Data-Integrity Invariants](../../../clients/compass-usa/workflows/order-lifecycle-and-data-integrity.md#detecting-empty-shell-orders-left-by-a-fatal-mid-sequence).
|
|
85
|
+
- **Orphan sweeps only find PARTIAL casualties.** An attempt that rolled back completely leaves no
|
|
86
|
+
row in any table and, with no API log row, no trail at all.
|
|
87
|
+
|
|
64
88
|
## Blind spots when auditing a client's API traffic
|
|
65
89
|
|
|
66
90
|
Looking only at `Logs_<Client>.Api` misses four sources:
|
|
@@ -110,6 +134,12 @@ re-send.
|
|
|
110
134
|
|
|
111
135
|
## Change history
|
|
112
136
|
|
|
137
|
+
- 2026-08-04 — Added the limit of the commit-logs contract: an **uncaught PHP fatal** (a
|
|
138
|
+
`Call to undefined method` from an interceptor dispatch) produced **zero rows in
|
|
139
|
+
`Logs_Compass.Api`** — no 5xx row for the whole day and no row for the failing request — while
|
|
140
|
+
still reaching the error tracker. Suspected regression in the `Error.php` / Issue-Event
|
|
141
|
+
normalization work. So log-based diagnosis fails for this class of 500; sweep the **data** for
|
|
142
|
+
zero-child parents instead, and accept that fully-rolled-back attempts leave no trail. (dfranks)
|
|
113
143
|
- 2026-08-03 — Corrected a common wrong assumption: an **HTTP 400 leaves a full log row**.
|
|
114
144
|
`Controller/Index.php` L369-377 commits `DB_LOGS`/`DB_CLIENT_LOGS` and rolls back only the
|
|
115
145
|
business transaction, so rejections are queryable and attributable to the credential; what they
|
|
@@ -18,7 +18,7 @@
|
|
|
18
18
|
| [Compass VIP Support Importer (worker2)](features/compass-vip-support-importer.md) | A worker2 action that ingests Compass's quarterly VIP spreadsheet and assigns each VIP user's support technician by setting `Users.c_supportedByUserId` in `Clie | worker2/Worker/Client/Compass/VipSupport.php |
|
|
19
19
|
| [Creating Worker Actions](features/creating-worker-actions.md) | How to add a new callable Worker action — a PHP class whose `public static` methods are invoked as background jobs (via webhook, cron, or `_Worker::runTask()`). | worker2/Worker/, worker2/Controller/Index.php, _underscore/Worker.php |
|
|
20
20
|
| [Elite Freshservice Sync (worker2)](features/elite-freshservice-sync.md) | `_Worker_Elite` processes Freshservice webhook events and syncs them into TOGA 2. | worker2/Worker/Elite.php, worker2/Config/dev-kmaramreddy-laptop.ini |
|
|
21
|
-
| [Error Escalation Cron (Errors::Escalate → ClickUp / email)](features/error-escalation-cron.md) | `_Worker_Infrastructure_Errors::Escalate` (renamed from `SyncWithClickup`) is the sole owner of **escalation, de-escalation, ClickUp ticketing, reminders, busin | worker2/Worker/Infrastructure/Errors.php, worker2/Worker/Clickup/ErrorTask.php, worker2/Worker/Clickup.php, worker2/Controller/Index.php, worker2/Config/production.ini, _underscore/Model/Core/Logs/Issue.php, dbchanges2/Core/2026-07-30a - Error escalation cron job.sql, dbchanges2/Logs/2026-08-03a - Issue clickupPriority.sql |
|
|
21
|
+
| [Error Escalation Cron (Errors::Escalate → ClickUp / email)](features/error-escalation-cron.md) | `_Worker_Infrastructure_Errors::Escalate` (renamed from `SyncWithClickup`) is the sole owner of **escalation, de-escalation, ClickUp ticketing, reminders, busin | worker2/Worker/Infrastructure/Errors.php, worker2/Worker/Notification/Email.php, worker2/Worker/Notification/EmailTemplate.php, worker2/Worker/Client/True.php, worker2/Worker/Clickup/ErrorTask.php, worker2/Worker/Clickup.php, worker2/Controller/Index.php, worker2/Config/production.ini, _underscore/Model/Core/Logs/Issue.php, dbchanges2/Core/2026-07-30a - Error escalation cron job.sql, dbchanges2/Logs/2026-08-03a - Issue clickupPriority.sql |
|
|
22
22
|
| [Etilize Catalog Item Import & Refresh](features/etilize-catalog-item-import.md) | Client-generic catalog onboarding from an S3 CSV plus an Etilize re-pull. | worker2/Worker/Etilize/Items.php |
|
|
23
23
|
| [Etilize Item Translation Import](features/etilize-item-translation-import.md) | The abstract worker class `_Worker_Etilize_ItemTranslations` imports **non-English** item text from Etilize into the client's `ItemTranslations` table. | worker2/Worker/Etilize/ItemTranslations.php |
|
|
24
24
|
| [Monitoring Framework (Orchestrator + Child Monitors)](features/monitoring-framework.md) | A unified, DB-driven monitoring framework for business-critical data flows (Compass POs, Prudential asset imports, AIG closed claims, …). | worker2/Worker/Monitor.php, worker2/Worker/Monitors/, worker2/Worker/Monitors/RateEntitlement.php, worker2/Worker/Notification/Email.php, worker2/Worker/Rate.php, dbchanges2/Core/2026-05-21 - Monitors.sql, dbchanges2/Core/2026-06-29a - Rate Entitlement Contract Monitor.sql |
|
|
@@ -10,6 +10,9 @@ updated: 2026-08-04
|
|
|
10
10
|
owners: ["jcardinal"]
|
|
11
11
|
files:
|
|
12
12
|
- worker2/Worker/Infrastructure/Errors.php
|
|
13
|
+
- worker2/Worker/Notification/Email.php
|
|
14
|
+
- worker2/Worker/Notification/EmailTemplate.php
|
|
15
|
+
- worker2/Worker/Client/True.php
|
|
13
16
|
- worker2/Worker/Clickup/ErrorTask.php
|
|
14
17
|
- worker2/Worker/Clickup.php
|
|
15
18
|
- worker2/Controller/Index.php
|
|
@@ -23,6 +26,8 @@ related:
|
|
|
23
26
|
- ../../../1.0/apps/library/features/error-capture-1-0.md
|
|
24
27
|
- ./creating-worker-actions.md
|
|
25
28
|
- ./clickup-project-routing.md
|
|
29
|
+
- ./notification-email.md
|
|
30
|
+
- ./notification-email-template.md
|
|
26
31
|
---
|
|
27
32
|
|
|
28
33
|
## Summary
|
|
@@ -107,10 +112,64 @@ which is safe only because the automation never sets status itself.
|
|
|
107
112
|
> stamping `dtAcknowledged` ~1s after `dtCreated` and permanently disabling the neglect axis for
|
|
108
113
|
> that issue. Status webhooks arriving within 60s of creation are ignored.
|
|
109
114
|
|
|
115
|
+
### ⚠ Escalation email had never been sent once — two independent causes (fixed 2026-08-04)
|
|
116
|
+
|
|
117
|
+
Evidence: `Logs_True.Email` contained **zero** error-escalation rows *ever*, while unrelated
|
|
118
|
+
worker2 mail (merge-conflict notices, sprint reports) delivered within ~60s — so the mail pipeline
|
|
119
|
+
itself was healthy. Every `Infrastructure/Errors/Escalate` run reported `"emailsSent":0` alongside
|
|
120
|
+
`"failures":0` and therefore **read as healthy**. All 48 HIGH/URGENT issues sat at
|
|
121
|
+
`notificationCount` 4 — the last rung of the ladder — having sent nothing.
|
|
122
|
+
|
|
123
|
+
**Root cause A (structural — the important one).** `processTechnicalIssue()` returned early when
|
|
124
|
+
the ClickUp task would read identically (priority already correct **and** count unchanged), and
|
|
125
|
+
**that `return` sat before the `if ($shouldNotify)` block.** `persistIssueState()` still ran on the
|
|
126
|
+
way out with `didNotify: $shouldNotify`, stamping `dtLastNotified` and incrementing
|
|
127
|
+
`notificationCount`. So a run that sent nothing still **consumed a rung** of the
|
|
128
|
+
30/60/240/1440-minute `REMINDER_INTERVAL_MINUTES` ladder — and since a live issue normally sits at
|
|
129
|
+
a settled priority with an unchanged count, that was **nearly every run**.
|
|
130
|
+
|
|
131
|
+
Fix: the ClickUp write is now skipped on its own merits, inside
|
|
132
|
+
`if ($needsPriorityWrite || $hasCountChanged)`, and the notification decision is made
|
|
133
|
+
**independently, below it**.
|
|
134
|
+
|
|
135
|
+
> **Lesson worth generalizing: "does ClickUp need writing?" and "does anyone need telling?" are two
|
|
136
|
+
> separate questions.** Conflating them into one early return silenced the entire reminder ladder
|
|
137
|
+
> while every metric reported success.
|
|
138
|
+
|
|
139
|
+
**Root cause B.** Both send paths built a `_Email` and called `->send()` **inline inside the cron's
|
|
140
|
+
open `DB_LOGS` transaction**. `_Email::send()` resolves the client's log database, registers a
|
|
141
|
+
second connection and toggles the read host from in there; anything it threw was caught into
|
|
142
|
+
`error_log` only — invisible in the console *and* in the run summary. Both paths now **enqueue via
|
|
143
|
+
`_Worker::runTask`**, so a failure is a visible failed `WorkerJobs` row.
|
|
144
|
+
|
|
145
|
+
### `emailFailures` in the run summary
|
|
146
|
+
|
|
147
|
+
`Escalate()`'s JSON gained **`emailFailures`**, threaded through `processIssue()` /
|
|
148
|
+
`processTechnicalIssue()`'s outcome arrays and counting notifications that were *meant* to go out
|
|
149
|
+
and did not. It is reported **separately from `failures`** (which counts issues that threw). It
|
|
150
|
+
exists specifically so the failure mode above — `emailsSent:0` with `failures:0` reading as
|
|
151
|
+
healthy — cannot recur silently.
|
|
152
|
+
|
|
110
153
|
### Notification and task presentation
|
|
111
154
|
|
|
112
155
|
- The technical escalation email to **devteam@togatech.com** now **coexists** with the ClickUp
|
|
113
156
|
update. It was an `elseif`, so only one of the two ever fired.
|
|
157
|
+
- **The technical escalation email uses the TOGA Technology stored email template.**
|
|
158
|
+
`sendTechnicalEscalationEmail()` enqueues **`Notification/EmailTemplate/Send`** with
|
|
159
|
+
`clientIdentifier` `'True'` and
|
|
160
|
+
`_Worker_Client_True::EMAIL_TEMPLATE_UUID__TOGA_TECHNOLOGY`
|
|
161
|
+
(`232d4edb-c2fa-4a8b-b5b9-d5800c962e19`) — the same pattern `Team/Transcripts.php` uses. Verified
|
|
162
|
+
against `Client_True.EmailTemplates`: `isActive = 1`; `{subject}` appears in **both** the subject
|
|
163
|
+
line and the branded header banner; `{body}` sits inside a **600px-wide** content cell that
|
|
164
|
+
already sets Plus Jakarta Sans 16px; the row carries its own `sendFromEmailAddress`
|
|
165
|
+
(`noreply@togatech.com`) and a single priority column (`NORMAL`).
|
|
166
|
+
- **`buildTechnicalEmailBody()` is therefore a FRAGMENT, not a document.** Its wrapper was
|
|
167
|
+
changed from `max-width:680px` + its own `font-family` to `max-width:100%` and no font —
|
|
168
|
+
680px would overflow the template's 600px cell.
|
|
169
|
+
- **Per-send priority is not available on the template path**, so urgency rides in the subject
|
|
170
|
+
and body instead.
|
|
171
|
+
- **Business-issue emails moved to `Notification/Email/Send`**, which *does* keep per-client log
|
|
172
|
+
attribution and per-send priority.
|
|
114
173
|
- ClickUp titles carry the **reference only** — `[1E]`, not `[1E-1]`. Including the occurrence
|
|
115
174
|
number made the title churn on every escalation.
|
|
116
175
|
- **Occurrences** and **Last Occurrence** custom fields are synced, and task bodies open with a
|
|
@@ -172,6 +231,13 @@ this controller needs the same treatment.**
|
|
|
172
231
|
|
|
173
232
|
## Gotchas / known issues
|
|
174
233
|
|
|
234
|
+
- **Never send email inline inside this cron's open transaction — enqueue it.** `_Email::send()`
|
|
235
|
+
registers a second connection and toggles the read host, and a throw from in there was only ever
|
|
236
|
+
`error_log`'d. Use `_Worker::runTask` so a failure becomes a visible failed `WorkerJobs` row.
|
|
237
|
+
- **A "nothing changed, return early" shortcut must not sit above the notification block.** It
|
|
238
|
+
consumed reminder-ladder rungs without sending anything, for months, while reporting
|
|
239
|
+
`emailsSent:0 / failures:0` as success. Decide the outbound write and the notification
|
|
240
|
+
independently.
|
|
175
241
|
- **Persist state before, or independently of, any outbound call.** A regression chain took the
|
|
176
242
|
whole run down every minute: `syncCustomFields()` read `$issue->dtLastOccurred`, which
|
|
177
243
|
`loadWorkingSet()` never selected; `_Error::errorHandler()` promotes the undefined-property
|
|
@@ -207,6 +273,25 @@ this controller needs the same treatment.**
|
|
|
207
273
|
|
|
208
274
|
## Change history
|
|
209
275
|
|
|
276
|
+
- 2026-08-04 (later, **uncommitted/undeployed** at time of writing) — **Fixed: escalation email had
|
|
277
|
+
never been sent, ever**, for two independent reasons. (A) `processTechnicalIssue()` returned early
|
|
278
|
+
when the ClickUp task would read identically, and that return sat *above* the `if ($shouldNotify)`
|
|
279
|
+
block while `persistIssueState()` still stamped `dtLastNotified`/`notificationCount` on the way
|
|
280
|
+
out — so nearly every run silently consumed a rung of the 30/60/240/1440-minute ladder; all 48
|
|
281
|
+
HIGH/URGENT issues had reached rung 4 having sent zero mail. The ClickUp write is now gated on its
|
|
282
|
+
own merits (`$needsPriorityWrite || $hasCountChanged`) and the notification decision is made
|
|
283
|
+
independently below it. (B) Both send paths built `_Email` and called `->send()` **inline inside
|
|
284
|
+
the open `DB_LOGS` transaction**, where a throw was only `error_log`'d — both now enqueue via
|
|
285
|
+
`_Worker::runTask`. **Built:** the technical escalation email now goes through the stored TOGA
|
|
286
|
+
Technology template (`Notification/EmailTemplate/Send`, clientIdentifier `True`,
|
|
287
|
+
`_Worker_Client_True::EMAIL_TEMPLATE_UUID__TOGA_TECHNOLOGY`), which makes
|
|
288
|
+
`buildTechnicalEmailBody()` a fragment (`max-width:100%`, no font — 680px would overflow the
|
|
289
|
+
template's 600px cell) and removes per-send priority (urgency moves into the subject/body);
|
|
290
|
+
business emails moved to `Notification/Email/Send`, which keeps per-client log attribution and
|
|
291
|
+
priority. **Built:** an **`emailFailures`** counter in `Escalate()`'s run summary, separate from
|
|
292
|
+
`failures`, so `emailsSent:0 / failures:0` can never read as healthy again. Verified by reading
|
|
293
|
+
only — proof will be `emailsSent` going non-zero (or `emailFailures` naming the problem) on the
|
|
294
|
+
first Escalate run after the worker2 deploy. (jcardinal)
|
|
210
295
|
- 2026-08-04 — **Fixed: worker job failures were never captured at all** — worker2's front
|
|
211
296
|
controller catches `Throwable` and returns 200, so six catch sites in
|
|
212
297
|
`worker2/Controller/Index.php` now call `_Error::captureException()` after rollback and before the
|
|
@@ -6,8 +6,8 @@ project: Worker
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-
|
|
10
|
-
owners: ["mhammontree"]
|
|
9
|
+
updated: 2026-08-04
|
|
10
|
+
owners: ["mhammontree", "jcardinal"]
|
|
11
11
|
files:
|
|
12
12
|
- worker2/Worker/Notification/Email.php
|
|
13
13
|
- _underscore/Model/Client/EmailTemplate.php
|
|
@@ -39,7 +39,7 @@ migration backfilled first).
|
|
|
39
39
|
- `worker2/Worker/Notification/Email.php` — `abstract _Worker_Notification_Email`. One method:
|
|
40
40
|
`Send(string $clientIdentifier, string $subject, string $body, string|array $to=[],
|
|
41
41
|
string|array $cc=[], string|array $bcc=[], string $fromEmail='donotreply@togatech.com',
|
|
42
|
-
string $fromName='TOGA Technology')`. The entry point. Self-registers the client DB (no
|
|
42
|
+
string $fromName='TOGA Technology', ?int $priority=null)`. The entry point. Self-registers the client DB (no
|
|
43
43
|
`initialize()` — see below), wraps the body, and sends.
|
|
44
44
|
- `_underscore/Model/Client/EmailTemplate.php` — supplies `WRAPPER_UUID` and
|
|
45
45
|
`renderWrappedBody($subject, $body)`. Documented in
|
|
@@ -69,6 +69,20 @@ migration backfilled first).
|
|
|
69
69
|
5. **Send** via a plain `_Email`: `setClientIdentifier`, `addTo/Cc/Bcc`, `setFrom`,
|
|
70
70
|
`setSubject`, `setBody($wrapped ?? $raw)`, `send()`.
|
|
71
71
|
|
|
72
|
+
### Optional per-send priority (added 2026-08-04)
|
|
73
|
+
|
|
74
|
+
The trailing **`?int $priority = null`** maps to `_Email::setPriority()`, which emits the
|
|
75
|
+
`X-Priority` / `Importance` / `X-MSMail-Priority` headers. Omit it and nothing changes — the
|
|
76
|
+
dispatcher spreads **named** parameters, so a new trailing optional parameter is backward
|
|
77
|
+
compatible for every existing caller and queued job.
|
|
78
|
+
|
|
79
|
+
It was added so that moving the error pipeline's **business** escalation email off an inline
|
|
80
|
+
`_Email` and onto this worker did not silently drop the urgency headers (see
|
|
81
|
+
[error-escalation-cron](./error-escalation-cron.md)). Note the sibling
|
|
82
|
+
[`Notification/EmailTemplate/Send`](./notification-email-template.md) path has **no** per-send
|
|
83
|
+
priority — a stored template carries its own single priority column — so urgency there has to live
|
|
84
|
+
in the subject and body.
|
|
85
|
+
|
|
72
86
|
### Enqueuing it
|
|
73
87
|
|
|
74
88
|
`_Worker::runTask('Notification/Email/Send', ['clientIdentifier'=>…, 'subject'=>…, 'body'=>…, …])`
|
|
@@ -130,6 +144,11 @@ working recipe — verified in classic + new Outlook:
|
|
|
130
144
|
are replaced simultaneously, so a `{body}` literal inside the subject can't be re-expanded.
|
|
131
145
|
|
|
132
146
|
## Change history
|
|
147
|
+
- 2026-08-04 (**uncommitted/undeployed** at time of writing) — Added an optional trailing
|
|
148
|
+
`?int $priority = null` to `Send()`, mapping to `_Email::setPriority()`
|
|
149
|
+
(`X-Priority`/`Importance`/`X-MSMail-Priority`). Backward compatible because the dispatcher
|
|
150
|
+
spreads named parameters. Added so the error pipeline's business escalation email kept its urgency
|
|
151
|
+
headers when it moved off an inline `_Email` onto this worker. (jcardinal)
|
|
133
152
|
- 2026-06-24 — Built the DB-driven notification-email mechanism (TRUE-79240): new
|
|
134
153
|
`_Worker_Notification_Email::Send` entry point; branded shell moved out of code (deleted
|
|
135
154
|
`_underscore/Email/Template.php`, which also carried the 1.0 double-render bug) into a reserved
|
package/knowledge/INDEX.md
CHANGED
|
@@ -5,7 +5,7 @@ _Auto-generated by `knowledge.js index`. Do not hand-edit._
|
|
|
5
5
|
## 1.0 framework
|
|
6
6
|
|
|
7
7
|
- **library** (Library) _(framework core)_ — 15 doc(s) → [1.0/apps/library/INDEX.md](1.0/apps/library/INDEX.md)
|
|
8
|
-
- **worker** (Worker) —
|
|
8
|
+
- **worker** (Worker) — 20 doc(s) → [1.0/apps/worker/INDEX.md](1.0/apps/worker/INDEX.md)
|
|
9
9
|
- **dbchanges** (Database Changes) _(framework core)_ — 1 doc(s) → [1.0/apps/dbchanges/INDEX.md](1.0/apps/dbchanges/INDEX.md)
|
|
10
10
|
- **worker1.5** (Worker 1.5) — 0 doc(s) → [1.0/apps/worker1.5/INDEX.md](1.0/apps/worker1.5/INDEX.md)
|
|
11
11
|
- **togadesk** (TOGa Desk) — 11 doc(s) → [1.0/apps/togadesk/INDEX.md](1.0/apps/togadesk/INDEX.md)
|
|
@@ -12,5 +12,6 @@
|
|
|
12
12
|
| [Compass USA](profile.md) | 2.0 | Compass USA is a TOGA client running a multi-tier supply-chain commerce operation. | |
|
|
13
13
|
| [Compass Cross-Kit Bundle Corruption — Detection & Repair](workflows/cross-kit-bundle-corruption.md) | 2.0 | A frontend regression in `toga2-commerce`'s edit-order bundle submission mis-attributed bundle (kit) line items and **fees/warranties** to the **wrong kit**, pe | src/api/syncSalesOrderItemsFromLocalStorageCartToApi.ts |
|
|
14
14
|
| [Compass Office Depot Duplicate PO-Line Cleanup (dual-catalog SKU)](workflows/odp-duplicate-po-line-cleanup.md) | 1.0 | The NetSuite fulfillment-sales-order importer duplicated Office Depot purchase-order lines on Compass USA because it reconciled the fulfillment SO against the e | library/app/api/toga2.php, dbchanges2/Client_Compass/2026-07-22a - CleanupOfficeDepotDuplicatePurchaseOrderItems.sql |
|
|
15
|
-
| [Compass ODP
|
|
15
|
+
| [Recovering a Lost Compass ODP EDI 850 Import (re-drop from Logs.FileLog)](workflows/odp-edi-import-recovery.md) | 1.0 | How to recover a Compass **Office Depot EDI 850** import that failed partway — the case where cron **3a** created the ODP SalesOrder header, the follow-up item | worker/crons/toga2/compass/workflow/3a_import_office_depot_purchase_orders.php, worker/crons/toga2/compass/workflow/5_create_netsuite_sales_orders_from_office_depot_purchase_orders.php, worker/schedules/cron.worker.sync.json |
|
|
16
|
+
| [Compass ODP Order Pipeline to NetSuite (numbered worker crons)](workflows/odp-order-pipeline-to-netsuite.md) | 1.0 | The end-to-end **Compass Office Depot (ODP) order → NetSuite** pipeline as it actually runs through the 1.0 `worker` crons under `worker/crons/toga2/compass/`, | worker/crons/toga2/compass/workflow/1_transmit_compass_sales_orders_to_mits.php, worker/crons/toga2/compass/workflow/2_transmit_mits_purchase_orders_to_vendors.php, worker/crons/toga2/compass/edi/1_download_edi_s3_create_po_toga.php, worker/crons/toga2/compass/workflow/3a_import_office_depot_purchase_orders.php, worker/crons/toga2/compass/workflow/5_create_netsuite_sales_orders_from_office_depot_purchase_orders.php, worker/schedules/cron.worker.sync.json, library/app/client/compass.php |
|
|
16
17
|
| [Compass Order Lifecycle & Data-Integrity Invariants](workflows/order-lifecycle-and-data-integrity.md) | 2.0 | End-to-end map of how a Compass order flows through the `Client_Compass` (2.0) database and the **expected raw-data shape** at each link/ASN/IF level. | |
|
|
@@ -0,0 +1,111 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Recovering a Lost Compass ODP EDI 850 Import (re-drop from Logs.FileLog)
|
|
3
|
+
framework: "1.0"
|
|
4
|
+
repo: worker
|
|
5
|
+
project: Worker
|
|
6
|
+
client: compass-usa
|
|
7
|
+
type: workflow
|
|
8
|
+
status: active
|
|
9
|
+
updated: 2026-08-04
|
|
10
|
+
owners: ["dfranks"]
|
|
11
|
+
files:
|
|
12
|
+
- worker/crons/toga2/compass/workflow/3a_import_office_depot_purchase_orders.php
|
|
13
|
+
- worker/crons/toga2/compass/workflow/5_create_netsuite_sales_orders_from_office_depot_purchase_orders.php
|
|
14
|
+
- worker/schedules/cron.worker.sync.json
|
|
15
|
+
related:
|
|
16
|
+
- clients/compass-usa/workflows/odp-order-pipeline-to-netsuite.md
|
|
17
|
+
- clients/compass-usa/workflows/order-lifecycle-and-data-integrity.md
|
|
18
|
+
- ../../../2.0/apps/api2/features/api-payload-interceptors.md
|
|
19
|
+
- ../../../2.0/apps/api2/features/request-logging.md
|
|
20
|
+
- clients/compass-usa/profile.md
|
|
21
|
+
---
|
|
22
|
+
|
|
23
|
+
## Summary
|
|
24
|
+
|
|
25
|
+
How to recover a Compass **Office Depot EDI 850** import that failed partway — the case where cron
|
|
26
|
+
**3a** created the ODP SalesOrder header, the follow-up item write died (e.g. an api2 fatal), and 3a
|
|
27
|
+
then **deleted the S3 object**, leaving a zero-item order and nothing to retry.
|
|
28
|
+
|
|
29
|
+
The recovery works because the raw x12 is archived in `Logs.FileLog` and **3a is fully idempotent**:
|
|
30
|
+
re-uploading the file makes it repair the items, the ODP PO, and both bridges. Verified end-to-end
|
|
31
|
+
in **production** on 2026-08-04 for ODP PO `41608261-1135` / ODP SO `475390815001`.
|
|
32
|
+
|
|
33
|
+
## When to use this
|
|
34
|
+
|
|
35
|
+
Symptoms: a zero-item ODP SalesOrder (see the orphan sweeps in
|
|
36
|
+
[order lifecycle & data integrity](order-lifecycle-and-data-integrity.md)), and cron 5 re-attempting
|
|
37
|
+
the NetSuite SO every 10 minutes because it selects on `c_dtTransmittedToNetsuite IS NULL`.
|
|
38
|
+
|
|
39
|
+
## The runbook
|
|
40
|
+
|
|
41
|
+
### Step 1 — retrieve the archived x12
|
|
42
|
+
|
|
43
|
+
The raw EDI is on the **legacy-core cluster**, table **`Logs.FileLog`**:
|
|
44
|
+
|
|
45
|
+
- `job = 'ODP_EDI'`
|
|
46
|
+
- `fileName = '<odpPoNumber>.x12'` — note **`.x12`**, even though the inbound S3 filename is
|
|
47
|
+
`EDI_<odpPoNumber>.txt`.
|
|
48
|
+
|
|
49
|
+
**Query it bounded by `job` + `fileTimestamp`.** An unbounded `fileName LIKE` scan times out on this
|
|
50
|
+
table.
|
|
51
|
+
|
|
52
|
+
Extract with `mysql --raw --batch --skip-column-names`, then **strip the trailing newline the client
|
|
53
|
+
appends** so the byte count matches the stored `LENGTH(fileData)` exactly. A byte-count mismatch means
|
|
54
|
+
you have a corrupted payload — do not upload it.
|
|
55
|
+
|
|
56
|
+
### Step 2 — know that S3 has nothing left
|
|
57
|
+
|
|
58
|
+
3a **deletes the S3 object after processing** (`:769`), so a failed import leaves no file to retry.
|
|
59
|
+
Recovery is a **re-upload**, to:
|
|
60
|
+
|
|
61
|
+
```
|
|
62
|
+
s3://agilant-as2/OfficeDepot/EDI_<odpPoNumber>.txt
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
### Step 3 — clear the cause BEFORE re-uploading
|
|
66
|
+
|
|
67
|
+
**Confirm no `ApiPayloadInterceptors` row exists for the `sales-order-items` record** (`recordId 15`)
|
|
68
|
+
whose derived method the model does not define. If one is still active, the item write fatals again
|
|
69
|
+
**and the file is deleted a second time**. See
|
|
70
|
+
[API payload interceptors](../../../2.0/apps/api2/features/api-payload-interceptors.md).
|
|
71
|
+
|
|
72
|
+
### Step 4 — let 3a repair it (≤ 5 minutes)
|
|
73
|
+
|
|
74
|
+
3a runs `*/5`. Because every step is existence-guarded it **reuses** the existing SO and **backfills**
|
|
75
|
+
what is missing: unmatched items (matched case-insensitively by `partNumber`), the ODP PO if absent,
|
|
76
|
+
and both the `PurchaseOrders_SalesOrders` and `PurchaseOrderItems_SalesOrderItems` bridges. No
|
|
77
|
+
duplicates are created.
|
|
78
|
+
|
|
79
|
+
**Then set `customerPurchaseOrder` manually** — 3a never updates the SalesOrder *header* on an
|
|
80
|
+
existing order, so it stays `NULL`. Prefer direct SQL over a `PUT`; the reasoning and its two
|
|
81
|
+
tradeoffs are in the
|
|
82
|
+
[lifecycle doc](order-lifecycle-and-data-integrity.md#repairing-a-header-only-scalar-field-direct-sql-vs-put).
|
|
83
|
+
|
|
84
|
+
### Step 5 — let cron 5 transmit to NetSuite
|
|
85
|
+
|
|
86
|
+
Cron 5 (`*/10`) picks the now-complete ODP SO up and stamps `c_dtTransmittedToNetsuite` +
|
|
87
|
+
`c_netsuiteInternalSalesOrderId`. It excludes `number LIKE 'MA%'`. Its stamps landing is the
|
|
88
|
+
confirmation the recovery finished.
|
|
89
|
+
|
|
90
|
+
## Gotchas / known issues
|
|
91
|
+
|
|
92
|
+
- **⚠ Re-uploading before fixing the root cause destroys the file again** (3a deletes on the way
|
|
93
|
+
through). Step 3 is not optional.
|
|
94
|
+
- **⚠ The archive filename extension differs from the S3 one** — `.x12` in `Logs.FileLog`,
|
|
95
|
+
`EDI_<po>.txt` in S3. Searching for the wrong one finds nothing.
|
|
96
|
+
- **Don't search for the ODP SalesOrder by the `BEG` number.** `BEG` is the ODP *PO*; the SO number
|
|
97
|
+
is `REF~QC`. Full segment map in the
|
|
98
|
+
[ODP pipeline doc](odp-order-pipeline-to-netsuite.md).
|
|
99
|
+
- **`Logs.FileLog` has limited retention** (~30 days) — beyond that there is no archived x12 and the
|
|
100
|
+
recovery must come from ODP re-sending.
|
|
101
|
+
- A **fully** rolled-back import leaves no row anywhere and no API log row, so it will not show up in
|
|
102
|
+
an orphan sweep at all; only the error tracker's per-event detail can enumerate those.
|
|
103
|
+
|
|
104
|
+
## Change history
|
|
105
|
+
|
|
106
|
+
- 2026-08-04 — Documented the runbook after recovering a production ODP order whose item write was
|
|
107
|
+
killed by an api2 fatal: `Logs.FileLog` (`job = 'ODP_EDI'`, `<po>.x12`, bounded by `fileTimestamp`)
|
|
108
|
+
as the x12 archive, the trailing-newline/byte-count check, the re-upload path to
|
|
109
|
+
`s3://agilant-as2/OfficeDepot/`, the mandatory "clear the interceptor first" step, 3a's idempotent
|
|
110
|
+
repair plus the `customerPurchaseOrder` header gap, and cron 5's `*/10` stamps as the completion
|
|
111
|
+
signal. (dfranks)
|
|
@@ -6,16 +6,18 @@ project: Worker
|
|
|
6
6
|
client: compass-usa
|
|
7
7
|
type: workflow
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-08-
|
|
10
|
-
owners: ["rgirish", "bala"]
|
|
9
|
+
updated: 2026-08-04
|
|
10
|
+
owners: ["rgirish", "bala", "dfranks"]
|
|
11
11
|
files:
|
|
12
12
|
- worker/crons/toga2/compass/workflow/1_transmit_compass_sales_orders_to_mits.php
|
|
13
13
|
- worker/crons/toga2/compass/workflow/2_transmit_mits_purchase_orders_to_vendors.php
|
|
14
14
|
- worker/crons/toga2/compass/edi/1_download_edi_s3_create_po_toga.php
|
|
15
15
|
- worker/crons/toga2/compass/workflow/3a_import_office_depot_purchase_orders.php
|
|
16
16
|
- worker/crons/toga2/compass/workflow/5_create_netsuite_sales_orders_from_office_depot_purchase_orders.php
|
|
17
|
+
- worker/schedules/cron.worker.sync.json
|
|
17
18
|
- library/app/client/compass.php
|
|
18
19
|
related:
|
|
20
|
+
- clients/compass-usa/workflows/odp-edi-import-recovery.md
|
|
19
21
|
- clients/compass-usa/workflows/order-lifecycle-and-data-integrity.md
|
|
20
22
|
- clients/compass-usa/features/mits-po-transmission-to-vendors.md
|
|
21
23
|
- clients/compass-usa/features/mits-po-to-so-item-linking.md
|
|
@@ -40,7 +42,8 @@ Commerce → Compass SalesOrder (SalesOrders.customerId = 2)
|
|
|
40
42
|
MITS → PurchaseOrder (PurchaseOrders.vendorId = 1 = OFFICE DEPOT; POST /v2/purchase-orders; PO number e.g. 50305906-1)
|
|
41
43
|
[cron 2_transmit_mits_purchase_orders_to_vendors] → ODP (cXML; email fallback per integration) [sets PurchaseOrders.dtSubmitted]
|
|
42
44
|
ODP → 850 EDI dropped to S3 (AS2) (bucket agilant-as2, prefix OfficeDepot/)
|
|
43
|
-
[cron
|
|
45
|
+
[cron workflow/3a_import_office_depot_purchase_orders — NOT edi/1, see correction below]
|
|
46
|
+
→ creates ODP SalesOrder (customerId = 1); deletes the S3 file after processing
|
|
44
47
|
ODP SalesOrder (customerId = 1)
|
|
45
48
|
[cron 5_create_netsuite_sales_orders_from_office_depot_purchase_orders, hourly] → NetSuite SO
|
|
46
49
|
writes back c_dtTransmittedToNetsuite + c_netsuiteInternalSalesOrderId ONTO the ODP SO (customerId=1)
|
|
@@ -56,12 +59,57 @@ ODP SalesOrder (customerId = 1)
|
|
|
56
59
|
`PurchaseOrders.dtSubmitted` when transmitted (~2 min after PO creation). (See the dedicated
|
|
57
60
|
[MITS PO Transmission to Vendors](../features/mits-po-transmission-to-vendors.md) feature for
|
|
58
61
|
the item-less-PO gotcha and the CXML-vs-EMAIL routing detail.)
|
|
59
|
-
3. **`
|
|
60
|
-
`agilant-as2`, prefix `OfficeDepot/` (
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
62
|
+
3. **`workflow/3a_import_office_depot_purchase_orders.php`** — pulls the ODP **850** from S3 bucket
|
|
63
|
+
`agilant-as2`, prefix `OfficeDepot/` (excluding `OUTBOX/` and `SENT/`; `:46` S3 read,
|
|
64
|
+
`:54-58` prefix filters) and creates the downstream **ODP SalesOrder** (customerId=1 = Office
|
|
65
|
+
Depot). **Deletes the S3 object after processing** (`:769`) — so an absent S3 object is expected
|
|
66
|
+
once ingested, not evidence of a miss, *and* a failed import leaves nothing to retry.
|
|
67
|
+
|
|
68
|
+
> **⚠ Correction (2026-08-04): `edi/1_download_edi_s3_create_po_toga.php` is DEAD CODE.** It
|
|
69
|
+
> appears in **no** schedule file. The scheduled entry *named* "Download EDI From S3 Files &
|
|
70
|
+
> Create PO - EDI Office Depot" (`*/5 * * * *`, `worker/schedules/cron.worker.sync.json`
|
|
71
|
+
> `:254-259`) actually points at **`toga2/compass/workflow/3a_import_office_depot_purchase_orders.php`**.
|
|
72
|
+
> The two scripts have **different idempotency and different API call shapes** — `edi/1`'s create
|
|
73
|
+
> branch POSTs the SO with *nested* items and a `customer.name`, while **3a POSTs the header alone
|
|
74
|
+
> (uuids) and then POSTs the items in a separate request**. Reasoning from `edi/1` gives wrong
|
|
75
|
+
> answers about what production does.
|
|
76
|
+
>
|
|
77
|
+
> **General rule: never infer which script runs from its filename or folder.** Verify against
|
|
78
|
+
> `worker/schedules/cron.*.json` — and note the schedule entry's **name can be actively
|
|
79
|
+
> misleading**.
|
|
80
|
+
|
|
81
|
+
### 3a is idempotent — re-dropping an EDI file is safe
|
|
82
|
+
|
|
83
|
+
Every step of 3a is gated by a `totalRecordCount == 0` existence check, so a re-drop produces **no
|
|
84
|
+
duplicates and backfills whatever is missing**:
|
|
85
|
+
|
|
86
|
+
| Step | Guard | Behavior on re-run |
|
|
87
|
+
|------|-------|--------------------|
|
|
88
|
+
| ODP SalesOrder | `:162-227` existence check | reuses the existing SO |
|
|
89
|
+
| SO items | `:470-498` | matches existing items **case-insensitively by `partNumber`**; creates only unmatched ones |
|
|
90
|
+
| ODP PurchaseOrder | `:356-436` | creates only if missing |
|
|
91
|
+
| `PurchaseOrders_SalesOrders` + `PurchaseOrderItems_SalesOrderItems` bridges | `:632` | creates only if missing |
|
|
92
|
+
|
|
93
|
+
**KNOWN GAP: 3a never updates the SalesOrder HEADER on an existing order.** So
|
|
94
|
+
`customerPurchaseOrder` stays `NULL` on a repaired order and must be set manually. See
|
|
95
|
+
[Recovering a lost ODP EDI import](odp-edi-import-recovery.md).
|
|
96
|
+
|
|
97
|
+
### The EDI 850 → TOGa identifier mapping (not guessable — this is the key to tracing an ODP order)
|
|
98
|
+
|
|
99
|
+
| 850 segment | Meaning | Where it lands |
|
|
100
|
+
|-------------|---------|----------------|
|
|
101
|
+
| **`BEG`** | the **Office Depot PO number** (e.g. `41608261-1135`) | `PurchaseOrders.number`, vendor "Agilant Solutions Inc" |
|
|
102
|
+
| **`REF~QC`** | the **ODP SALES ORDER number** (e.g. `475390815001`) | `SalesOrders.number` with `customerId = 1` |
|
|
103
|
+
| **`REF~EU` / `REF~PO`** | the **Compass MITS PO number** (e.g. `50310002-1`) | its `c_mitsSalesOrder` field holds the ODP SO number |
|
|
104
|
+
| **`REF~LU`** | the **Compass sales order** (e.g. `SA135273`) | `SalesOrders` with `customerId = 2` |
|
|
105
|
+
| **`PO1`** | the line | qty, UOM, unit price, vendor part, ODP item id |
|
|
106
|
+
|
|
107
|
+
**⚠ `REF~QC` — not `BEG` — is the value to search on for the ODP SalesOrder.** Searching
|
|
108
|
+
`PurchaseOrders` / `SalesOrders` for the `BEG` number finds nothing for the SO and wastes time.
|
|
109
|
+
(This is the segment-level detail behind the "don't search by the MITS PO number" rule below.)
|
|
110
|
+
4. **`workflow/5_create_netsuite_sales_orders_from_office_depot_purchase_orders.php`** — scheduled
|
|
111
|
+
**every 10 minutes** (`*/10`, verified against `worker/schedules/` on 2026-08-04; this doc
|
|
112
|
+
previously said "hourly"); picks up ODP SalesOrders
|
|
65
113
|
(`OfficeDepotSalesOrders.customerId = 1 AND c_dtTransmittedToNetsuite IS NULL`), creates the
|
|
66
114
|
SO in NetSuite, and writes back `c_dtTransmittedToNetsuite` + `c_netsuiteInternalSalesOrderId`
|
|
67
115
|
**onto the ODP SalesOrder (customerId=1), never the Compass SO (customerId=2)**.
|
|
@@ -70,6 +118,10 @@ ODP SalesOrder (customerId = 1)
|
|
|
70
118
|
- **computer-kit orders** (Bundles 187–196) wait for a **2nd PO** before syncing.
|
|
71
119
|
- a **30-minute delay** after the first ODP PO is created.
|
|
72
120
|
|
|
121
|
+
Because the selection is only `c_dtTransmittedToNetsuite IS NULL` (plus `number NOT LIKE 'MA%'`),
|
|
122
|
+
a **zero-item ODP SalesOrder sits in this queue retrying every 10 minutes** until its lines are
|
|
123
|
+
restored — the retry loop is the symptom, not the cause.
|
|
124
|
+
|
|
73
125
|
## Cron 5 NetSuite SO creation — failure modes & the two ODP error emails
|
|
74
126
|
|
|
75
127
|
Two different Compass ODP error emails exist, from **two different pipeline stages**. They look
|
|
@@ -141,6 +193,15 @@ wrong record makes every order look "stuck." Use the **join**, never a number ma
|
|
|
141
193
|
transmission feature.)
|
|
142
194
|
|
|
143
195
|
## Change history
|
|
196
|
+
- 2026-08-04 — **Corrected the 850-import step**: `edi/1_download_edi_s3_create_po_toga.php` is
|
|
197
|
+
**dead code in no schedule**; the schedule entry named "Download EDI From S3 Files & Create PO -
|
|
198
|
+
EDI Office Depot" (`*/5`) actually runs `workflow/3a_import_office_depot_purchase_orders.php`,
|
|
199
|
+
which has different idempotency and posts the SO header and its items as **two separate API
|
|
200
|
+
requests**. Added 3a's full existence-check/idempotency map (re-dropping a file is safe; header is
|
|
201
|
+
never updated so `customerPurchaseOrder` stays NULL), the **EDI 850 → TOGa identifier mapping**
|
|
202
|
+
(`BEG` = ODP PO, `REF~QC` = ODP SO — the one to search on, `REF~EU`/`REF~PO` = MITS PO,
|
|
203
|
+
`REF~LU` = Compass SO), and corrected cron 5's schedule to `*/10`. Production incident
|
|
204
|
+
investigation; no code change. (dfranks)
|
|
144
205
|
- 2026-08-03 — Documented cron 5's "Please enter a value for Ext Price" root cause (`rate` set
|
|
145
206
|
only for NetSuite item groups, never plain items → NetSuite `add()` USER_ERROR, order re-emails
|
|
146
207
|
hourly while stuck), the two distinct ODP error emails (3a PO-import vs 5 SO-creation) and cron
|
|
@@ -5,8 +5,8 @@ project: _Underscore
|
|
|
5
5
|
client: compass-usa
|
|
6
6
|
type: workflow
|
|
7
7
|
status: active
|
|
8
|
-
updated: 2026-
|
|
9
|
-
owners: ["jcardinal", "bala"]
|
|
8
|
+
updated: 2026-08-04
|
|
9
|
+
owners: ["jcardinal", "bala", "dfranks"]
|
|
10
10
|
files: []
|
|
11
11
|
related:
|
|
12
12
|
- clients/compass-usa/profile.md
|
|
@@ -93,6 +93,41 @@ one-upstream for bundles.
|
|
|
93
93
|
5. **Mirror back-links:** every IF on a non-source SO (ODP/Agilant) should have
|
|
94
94
|
`upstreamItemFulfillmentId` set, and each of its items `upstreamItemFulfillmentItemId` set.
|
|
95
95
|
|
|
96
|
+
## Detecting "empty shell" orders left by a fatal mid-sequence
|
|
97
|
+
|
|
98
|
+
The ODP EDI import creates a SalesOrder **header** and its **line items** in **separate API
|
|
99
|
+
requests** (`POST /sales-orders`, then `POST /sales-order-items` ~4s later — see the
|
|
100
|
+
[ODP pipeline doc](odp-order-pipeline-to-netsuite.md)). If the *second* request dies, the header is
|
|
101
|
+
already committed and the items roll back, leaving a **zero-item sales order**. This is the
|
|
102
|
+
signature to sweep for whenever an api2 write fatals — and, because a fatal can leave **no API log
|
|
103
|
+
row at all** ([request logging](../../../2.0/apps/api2/features/request-logging.md)), the **data is
|
|
104
|
+
the only trail**.
|
|
105
|
+
|
|
106
|
+
Sweeps against `Client_Compass`:
|
|
107
|
+
|
|
108
|
+
1. **Zero-item sales orders** — `SalesOrders LEFT JOIN SalesOrderItems … GROUP BY so.id HAVING
|
|
109
|
+
COUNT(soi.id) = 0`.
|
|
110
|
+
2. **Zero-item purchase orders** — the same shape over `PurchaseOrders` / `PurchaseOrderItems`.
|
|
111
|
+
3. **ODP SOs with no upstream PO** — ODP SOs (`customerId = 1`) `LEFT JOIN
|
|
112
|
+
PurchaseOrders_SalesOrders WHERE` the bridge row `IS NULL`.
|
|
113
|
+
|
|
114
|
+
> **⚠ LIMITATION: this finds only PARTIAL casualties.** An attempt that rolled back *completely*
|
|
115
|
+
> leaves no row in any table, and with no API log row there is no trail anywhere — only the error
|
|
116
|
+
> tracker's **per-event detail** can enumerate those. Never report an orphan sweep as a complete
|
|
117
|
+
> casualty list.
|
|
118
|
+
|
|
119
|
+
### Repairing a header-only scalar field: direct SQL vs `PUT`
|
|
120
|
+
|
|
121
|
+
When only a scalar needs setting on an existing order (e.g. `customerPurchaseOrder`, which the ODP
|
|
122
|
+
importer never backfills), **direct SQL is the deliberate choice over `PUT /sales-orders/{uuid}`**: a
|
|
123
|
+
`PUT` fires the record-14 `POST`/`PUT` [payload
|
|
124
|
+
interceptor](../../../2.0/apps/api2/features/api-payload-interceptors.md) and drags the whole order
|
|
125
|
+
back through MITS/NetSuite logic for one field.
|
|
126
|
+
|
|
127
|
+
Accept the two tradeoffs: **no API log row** for the change, and **`dtUpdated` still moves** (the
|
|
128
|
+
column is `ON UPDATE CURRENT_TIMESTAMP`), which can look like a fresh modification to anything doing
|
|
129
|
+
`dtUpdated`-based change detection downstream.
|
|
130
|
+
|
|
96
131
|
## Verified schema corrections (older code/scripts assume these wrongly)
|
|
97
132
|
- **`ItemFulfillmentPackages` does NOT exist** — header IF tracking is `ItemFulfillments_TrackingNumbers`.
|
|
98
133
|
- `ItemFulfillmentItemUnits` has **no `trackingNumberId`** — unit tracking is the bridge
|
|
@@ -120,6 +155,13 @@ runtime workflow.
|
|
|
120
155
|
- High-multiplier over-fulfillment (5×–20×) does not fit the split-PO spurious-link pattern — separate cause.
|
|
121
156
|
|
|
122
157
|
## Change history
|
|
158
|
+
- 2026-08-04 — Added the **"empty shell" order signature** — the ODP importer writes the SO header
|
|
159
|
+
and its items in two separate API requests, so a fatal on the second leaves a committed,
|
|
160
|
+
zero-item order — plus the three orphan sweeps that detect it (zero-item SOs, zero-item POs, ODP
|
|
161
|
+
SOs with no `PurchaseOrders_SalesOrders` bridge) and the limitation that they only find *partial*
|
|
162
|
+
casualties. Also recorded why a header-only scalar repair uses **direct SQL over `PUT`** (a `PUT`
|
|
163
|
+
fires the record-14 interceptor and the whole MITS/NetSuite path) and its two tradeoffs (no API
|
|
164
|
+
log row; `dtUpdated` still moves via `ON UPDATE CURRENT_TIMESTAMP`). (dfranks)
|
|
123
165
|
- 2026-06-30 — Updated the IF-stage invariant: `itemFulfillmentStageId` is now NOT NULL and
|
|
124
166
|
picked/packed/shipped stages are all in use (resolved by slug), superseding "stage 3 = Shipped;
|
|
125
167
|
stages 1/2 unused". Compass order status remains shipped-only. (bala)
|
package/package.json
CHANGED