toga-ai 1.0.844 → 1.0.845

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -6,7 +6,7 @@ project: Worker
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-09-17
9
+ updated: 2026-09-18
10
10
  owners: ["dfranks", "bala", "jcardinal", "mhammontree", "snaredla", "rgirish"]
11
11
  files:
12
12
  - worker/crons/toga2/netsuite/common_sync_togasupply.php
@@ -106,8 +106,13 @@ The value is `<seconds>-<STATE>`. Three readings, and only the first is a real f
106
106
 
107
107
  - **`1-RUNNING`** (window collapsed ÷3 to 1s) = throwing on **every** run = real freeze. Go read the
108
108
  error (`Logs.Event` by clientId → file:line → `Logs_<Client>.Api` on the transactionId).
109
- - **A large-window `*-RUNNING`** (e.g. `864000-RUNNING`, `288000-RUNNING`) = healthy mid-run **or** a
110
- deliberate pause/backfill — **NOT** stuck. Do not "fix" it.
109
+ - **A large-window `*-RUNNING`** (e.g. `864000-RUNNING`, `288000-RUNNING`) = **usually** healthy
110
+ mid-run or a deliberate pause/backfill. **⚠ CORRECTED 2026-09-18 — this is NOT sufficient on its
111
+ own, and the old "NOT stuck, do not fix it" wording caused a real miss.** NYCHH showed
112
+ `864000-RUNNING` while its cursor sat **2.5 years** stale, and the session skipped it on that
113
+ reading. **Always read `NETSUITE_LAST_SYNC_CURSOR_<SECTION>` together with the mode.** A full
114
+ window with an ancient cursor is **broken, not busy** — see the future-cursor bug below, which
115
+ produces exactly this healthy-looking shape.
111
116
  - **`IDLE` + cursor frozen exactly at a reset value + NO error events = the section toggle is off**
112
117
  (`IS_ENABLED_INTEGRATION_<SECTION> = false` in the wrapper). Resetting the cursor does **nothing**
113
118
  while the toggle is off. Worked example: NYCHH SALES_ORDERS looked "stuck" after a cursor reset to
@@ -322,6 +327,70 @@ hit **every section and every client**, not just the by-reference ones.
322
327
  GET's UTC. `strtotime()` and `App_Date::convertToSQLDate()` both use the script tz (Chicago) and are
323
328
  the trap. Any manual cursor rewind value you set by hand must also be **Eastern**.
324
329
 
330
+ ### ⚠ …and the empty-window write is STILL WRONG — it stores a cursor ONE HOUR IN THE FUTURE (found 2026-09-18, Elite prod — NOT FIXED)
331
+
332
+ **This CORRECTS the section directly above.** The 2026-09-09 fix made the *read* and the
333
+ *per-record write* agree on Eastern, and its claim that "the **whole** cursor pipeline runs on one
334
+ clock" is **wrong**. The empty-window (`|0`) write it introduced has the opposite error, and it is
335
+ live in every section for every client.
336
+
337
+ - **The symptom is a section that looks perfectly healthy.** Elite's POs sat at `864000-IDLE` with a
338
+ recent cursor and synced **nothing**. No error, no shrinking window, no alert. This is the
339
+ hardest failure shape in the engine to spot, because every dashboard signal says fine.
340
+ - **Root cause — a double Eastern conversion.** `common_sync_togasupply.php:945` (POs) writes
341
+ `App_Date::convertToTimezone($timeNow, 'Y-m-d H:i:s', 'America/New_York')`. `$timeNow` is an
342
+ **epoch**, and `convertToTimezone` renders it in Eastern. The server clock is
343
+ **America/Chicago**, so the stored string is **local time + 1 hour**. The next run reads that
344
+ string back as the window **start**, converts it to Eastern **again** for the NetSuite query, and
345
+ asks for records modified after a moment that has not happened yet.
346
+ - **It is self-sustaining.** Nothing matches → the window is empty → the `|0` branch writes another
347
+ future cursor. Forever.
348
+ - **Prod proof (Elite `NETSUITE_LAST_SYNC_CURSOR_PURCHASE_ORDERS`), every write exactly +60 min and
349
+ always ending `|0`:**
350
+
351
+ | Written at (server, Chicago) | Cursor stored |
352
+ |---|---|
353
+ | 10:36:09 | 11:36:08 |
354
+ | 10:26:10 | 11:26:08 |
355
+ | 10:22:41 | 11:22:40 |
356
+
357
+ - **`|0` is the tell.** The `|0` suffix marks the empty-window branch
358
+ (`$countProcessedRecords === 0`). A cursor that is both `|0` **and** in the future is this bug.
359
+ - **Scope: every section, every client.** The same line pattern is at **838** (sales orders),
360
+ **945** (purchase orders), **1031** (invoices), **1200** (item receipts) and the fulfillments
361
+ section. It only bites once a section **catches up** and starts hitting empty windows — which is
362
+ exactly what happened when Elite's PO backfill finished. Likely also explains **Endeavor Health**
363
+ (window growing, cursor frozen, no errors).
364
+ - **⚠ OPEN — NOT FIXED.** `common_sync_togasupply.php` is shared by all 22 netsuite clients and the
365
+ developer has not given the go-ahead. Read-only investigation as of 2026-09-18.
366
+ - **The nuance the 2026-09-17 "ruled out" note got right, and what it got wrong.** The earlier
367
+ entry that recorded "a cursor timezone theory — **wrong**" was right about the **per-record**
368
+ write (the `common` line ~821 path, which correctly uses the NetSuite `lastModifiedDate`). It was
369
+ wrong as a blanket statement. **Both are true: the per-record write is correct; the empty-window
370
+ write is broken.** Do not use one to dismiss the other.
371
+
372
+ #### Reading `NETSUITE_EXECUTION_MODE_*`: the mode alone is NOT enough — always read the cursor with it
373
+
374
+ How to survey the fleet for stuck clients (client list:
375
+ `test/team/2.0 deployment/NetSuite_Clients_Db.txt`, 22 netsuite clients):
376
+
377
+ | Value | Reading |
378
+ |---|---|
379
+ | `864000-IDLE` | healthy |
380
+ | a **shrinking** number with `RUNNING` | a section that is throwing — go read `Logs.Issue` |
381
+ | `<max>-RUNNING` | usually just mid-run — **but check the cursor before you believe it** |
382
+
383
+ **⚠ CORRECTION (2026-09-18): `864000-RUNNING` was read as "healthy, in progress" and skipped. That
384
+ was wrong.** NYCHH showed `864000-RUNNING` while its cursor was **2.5 years stale**. A full-window
385
+ `RUNNING` with an ancient cursor is **broken, not busy**. Always pair the mode with
386
+ `NETSUITE_LAST_SYNC_CURSOR_<section>` — neither value is a health signal on its own.
387
+
388
+ **Fleet sweep result, 2026-09-18** (all 22 clients): Endeavor Health `INVOICES 1-RUNNING`, AIG
389
+ `SALES_ORDERS 8-RUNNING`, Compass `INVOICES 96000-RUNNING`, NYCHH both sections 2.5 years behind,
390
+ Elite known. The other 17 were clean. After resetting the modes to `864000-IDLE`, **AIG and Compass
391
+ recovered on their own**; **Endeavor Health's window grew but its cursor stayed frozen** — the
392
+ future-cursor signature above.
393
+
325
394
  ### ⚠ ITEM_RECEIPTS froze at RUNNING on a case-mismatched Units serial → duplicate-unit INSERT (EV-10 / 1062) (found + fixed 2026-09-11, NYCHH prod)
326
395
 
327
396
  `syncItemReceiptFromNetsuite()` (`library/app/api/toga2.php` ~L4022) builds a PHP array of the
@@ -732,7 +801,7 @@ nested count:
732
801
  line the fulfillment referenced, the match failed, and the section fail-loud-froze. The flat list
733
802
  replaces the truncated nested one before match/reconstruct.
734
803
 
735
- ### Retiring the stale SalesOrder a reclassified TransferOrder leaves behind (2026-09-17, NOT YET DEPLOYED)
804
+ ### Retiring the stale SalesOrder a reclassified TransferOrder leaves behind (DEPLOYED + RUNNING 2026-09-18)
736
805
 
737
806
  **Turning transfer-order detection on for an existing client strands every order it already
738
807
  imported as a SalesOrder.** The same NetSuite order now syncs to a `TransferOrder`, but the old
@@ -741,15 +810,25 @@ and `Items._qtyOnHand` goes negative. Measured on Elite: **218** SalesOrders dup
741
810
  TransferOrder on `c_netsuiteInternalSalesOrderId`, and **217 of 217** matched pairs are identical on
742
811
  order number, line count *and* total quantity.
743
812
 
744
- `App_Api_Toga2::retireStaleSalesOrderForTransferOrder()` (new, `private static`) is called at the
813
+ `App_Api_Toga2::retireStaleSalesOrderForTransferOrder()` (`private static`) is called at the
745
814
  **END** of `syncTransferOrderFromNetsuite()` so the TO and its lines exist first. Per SO line:
746
815
 
747
816
  1. move `ItemFulfillmentItems`: `salesOrderItemId` → `transferOrderItemId`,
748
- 2. move `PurchaseOrderItems_SalesOrderItems` → `PurchaseOrderItems_TransferOrderItems`,
749
- 3. `DELETE` the `SalesOrderItem`;
817
+ 2. move `PurchaseOrderItems_SalesOrderItems` → `PurchaseOrderItems_TransferOrderItems` (upstream),
818
+ 3. **(b2)** move `SalesOrderItems_PurchaseOrderItems` → `PurchaseOrderItems_TransferOrderItems`
819
+ (downstream — added 2026-09-18),
820
+ 4. `DELETE` the `SalesOrderItem`;
821
+
822
+ then at header level:
823
+
824
+ 5. **(d2)** collapse **both** `SalesOrders_PurchaseOrders` (downstream) **and**
825
+ `PurchaseOrders_SalesOrders` (upstream) onto `PurchaseOrders_TransferOrders`
826
+ — upstream added 2026-09-18,
827
+ 6. **(e)** repoint `ItemFulfillments.salesOrderId` → `transferOrderId` (the fulfillment **header**,
828
+ added 2026-09-18),
829
+ 7. `DELETE` the `SalesOrder`.
750
830
 
751
- then move `SalesOrders_PurchaseOrders` → `PurchaseOrders_TransferOrders` and `DELETE` the
752
- `SalesOrder`. **Move before delete, always** — the bridge FKs are `RESTRICT`.
831
+ **Move before delete, always** — the bridge FKs are `RESTRICT`.
753
832
 
754
833
  **The two fulfillment-parent columns are mutually exclusive, which is what makes the move safe.**
755
834
  Prod check: 391 rows TO-side, 293 SO-side, **0 with both, 0 with neither**.
@@ -781,6 +860,68 @@ fleet-wide grant is `dbchanges2/_modules/netsuite/2026-09-17a - StaleSalesOrderR
781
860
  load-bearing for this method's safety. Skipped for now because Elite TOs are 1–5 lines; flagged by
782
861
  php-reviewer.
783
862
 
863
+ #### ⚠ It shipped missing FOUR more FK dependencies — and the fix for that is a QUERY, not more guessing (2026-09-18)
864
+
865
+ The method deployed, ran, and then failed on **one blocker after another**, each revealed only by
866
+ fixing the one before it. Found and fixed in the order they surfaced:
867
+
868
+ | # | Missing dependency | Rows on the 218 stale orders | Resolution |
869
+ |---|---|---|---|
870
+ | a | `ItemFulfillments.salesOrderId` — the fulfillment **HEADER** | **all 218** | repoint to the TransferOrder (step e) |
871
+ | b | `PurchaseOrders_SalesOrders` — the **upstream** header bridge | **218** (vs 123 downstream) | step d2 |
872
+ | c | `SalesOrderItems_PurchaseOrderItems` — the **downstream** line bridge | 7 (vs 669 upstream) | step b2 |
873
+ | d | `Invoices.salesOrderId` | 1 (SalesOrder 240734) | **cannot** be moved — `Invoices` has no `transferOrderId`; correctly throws |
874
+
875
+ **(a) would have failed every single order.** `ItemFulfillments` carries its own
876
+ `salesOrderId`/`transferOrderId` pair — the same either-or as the line table (prod: 334 rows, all
877
+ SO-side, **0 with both, 0 with neither**). The original method only moved the fulfillment **items**.
878
+ The fulfillment itself is real shipping history and is **never deleted** — only repointed.
879
+
880
+ **The pattern behind (b) and (c): every SO↔PO bridge is TWO tables, and it is easy to fix one and
881
+ miss its sibling.**
882
+
883
+ | Handled first | Missed sibling |
884
+ |---|---|
885
+ | `SalesOrders_PurchaseOrders` | `PurchaseOrders_SalesOrders` |
886
+ | `PurchaseOrderItems_SalesOrderItems` | `SalesOrderItems_PurchaseOrderItems` |
887
+
888
+ Note the volumes cut **both ways** — upstream dominates at header level (218 vs 123) and downstream
889
+ is the rare one at line level (7 vs 669). Neither direction is "the main one"; you cannot skip one
890
+ because it looked small on another client. Direction semantics are on the
891
+ [SO↔PO bridge direction doc](../../../2.0/apps/_underscore/features/sales-order-purchase-order-bridge-direction.md).
892
+
893
+ **⚠ THE METHOD THAT ENDED THE WHACK-A-MOLE — make this step 1 of any "delete a parent record" work,
894
+ not step 5.** Stop guessing which children exist. Ask the schema, then count:
895
+
896
+ ```sql
897
+ -- 1. every FK pointing at the parent (and its line table)
898
+ SELECT TABLE_NAME, COLUMN_NAME, REFERENCED_TABLE_NAME
899
+ FROM information_schema.KEY_COLUMN_USAGE
900
+ WHERE TABLE_SCHEMA = 'Client_Elite'
901
+ AND REFERENCED_TABLE_NAME IN ('SalesOrders', 'SalesOrderItems');
902
+
903
+ -- 2. then COUNT actual rows in each, restricted to the affected orders
904
+ ```
905
+
906
+ That pair of queries found **12** tables referencing `SalesOrders` and **9** referencing
907
+ `SalesOrderItems`. Only the **4** above had any rows. The other **15 were empty** and needed no
908
+ code at all: `Entitlements_SalesOrders`, `SalesOrderEmailAddresses`, `SalesOrderNotes`,
909
+ `SalesOrders_Payments`, `SalesOrders_SalesSupportUsers`, `SalesOrders_TransferOrders`,
910
+ `TransferOrders_SalesOrders`, `SalesOrderItems_CommittedUnits`, `SalesOrderItems_Tags`, child lines
911
+ via `parentSalesOrderItemId`, `SalesOrderItems_TransferOrderItems`,
912
+ `TransferOrderItems_SalesOrderItems`. **Counting first is what tells you which of the 21 you
913
+ actually have to write code for** — the schema alone would have sent us after all 21.
914
+
915
+ **Verified working in Elite prod, 2026-09-18:**
916
+
917
+ | Measure | Before → after |
918
+ |---|---|
919
+ | stale duplicate SalesOrders | 218 → **187** |
920
+ | `SalesOrders` total | 340 → **309** |
921
+ | fulfillments on the TransferOrder side | 1 → **37** (moved, **not** deleted) |
922
+ | `PurchaseOrders_TransferOrders` | 17 → **104** |
923
+ | `PurchaseOrderItems_TransferOrderItems` | 28 → **163** |
924
+
784
925
  ### ⚠ TWO code paths create `PurchaseOrderItems` — only one stamped `fulfillmentType` (fixed 2026-09-17)
785
926
 
786
927
  `fulfillmentType` was stamped by `syncPurchaseOrderFromNetsuite()` (from the order-level
@@ -822,13 +963,39 @@ Fix, applied at **both** item-level delete sites: **re-read the item's tracking
822
963
  before deleting, instead of trusting the run-start snapshot — the same re-fetch pattern this file
823
964
  already uses for the single-tracking-number fan-out. A `totalRecordCount` guard was added to both.
824
965
 
825
- **Scope was confirmed, not assumed:** every FK-1451 failure that day was on
826
- `/v2/item-fulfillment-items/` — **zero** on units or tracking rows — so the unit-level lists were
827
- deliberately left alone.
966
+ **⚠ "Scope was confirmed, not assumed" was WRONG — there were THREE layers, not one
967
+ (corrected 2026-09-18).** Clearing `ItemFulfillmentItems_TrackingNumbers` only exposed the next
968
+ layer down: `ItemFulfillmentItemUnits.itemFulfillmentItemId`. Same root cause each time — a
969
+ run-start snapshot used to delete children while a later pass **in the same run** creates more of
970
+ them. Confirmed on the blocking row (item `b96515b3`): **trackingRows 1, unitRows 1,
971
+ unitTrackingRows 1** — each layer hidden behind the one above it. The "zero failures on units"
972
+ observation was true *at that moment* only because the tracking layer failed first and never let
973
+ the run reach the units.
974
+
975
+ **Second fix:** the stale-item delete block now re-reads units **live** via the existing paginated
976
+ helper `getItemFulfillmentItemUnitsByItemFulfillmentUuid()`, then deletes, in order, each unit's
977
+ tracking bridge rows → the unit → the item.
978
+
979
+ **Lesson: with nested children, fixing the FK you can see just reveals the next one. Enumerate the
980
+ whole child tree up front** — the same `information_schema.KEY_COLUMN_USAGE` + row-count method
981
+ written up under the stale-SalesOrder retirement section above.
828
982
 
829
983
  The exception was **not** swallowed. `send()` throws by default; the throw aborted the run and left
830
984
  the mode `RUNNING`. That is the designed fail-loud behaviour working correctly.
831
985
 
986
+ **⚠ OPEN as of 2026-09-18 — the fix is committed but the running process is NOT executing it.**
987
+ The unit-layer fix is committed (`2247a70b`, working tree clean) and Elite still fails in prod.
988
+
989
+ **The diagnostic that proves "committed ≠ running": compare the ACTUAL API call sequence in
990
+ `Logs_<client>.Api` against what the committed code would emit.** The committed code performs a
991
+ `GET /v2/item-fulfillment-item-tracking-numbers` immediately before the delete. That GET **does not
992
+ appear** in the log. At 10:27 and again at 11:07 the sequence was still two plain
993
+ `GET /v2/item-fulfillment-items` followed by the failing `DELETE` — i.e. the **old** code path.
994
+
995
+ This is a deploy / opcache / checkout-path problem, **unresolved**. Use this technique before
996
+ re-debugging any "fix that didn't work": the API log is a faithful trace of which code is actually
997
+ running, and it settles the question in one query.
998
+
832
999
  **General rule for this engine: any list used to delete children must be re-read immediately before
833
1000
  the delete if ANY later pass in the same run can create more of them.** A run-start snapshot is only
834
1001
  safe for read-only use.
@@ -841,7 +1008,11 @@ order — record these so nobody re-walks them:
841
1008
  1. backfill lag — no,
842
1009
  2. the `property_exists` gate — no,
843
1010
  3. ACL on the field — grants exist (`recordFieldId` **2594/2595**, roleId 3, `isWritable = 1`),
844
- 4. a cursor timezone theory — **wrong**; the code correctly uses `America/New_York`.
1011
+ 4. a cursor timezone theory — **wrong *for this symptom*, and only for the PER-RECORD write**;
1012
+ that path correctly uses `America/New_York`. **⚠ Do not read this as "cursor timezones are
1013
+ fine" — corrected 2026-09-18:** the **empty-window (`|0`) write** genuinely is broken and
1014
+ stores a cursor one hour in the future. See the future-cursor section above. Both facts are
1015
+ true at once.
845
1016
 
846
1017
  **Actual reason:** those orders are now classified as **Transfer Orders** (`$0` + `holdInvoice`), so
847
1018
  the sync routes them to `syncTransferOrderFromNetsuite()` and they never reach the sales-order path
@@ -1824,6 +1995,27 @@ library (or vice versa) crashes GroWrk and Adyen on their next sync run.
1824
1995
  [NetSuite Sync Alert Monitor](../../library/features/netsuite-sync-alert-monitor.md).
1825
1996
 
1826
1997
  ## Change history
1998
+ - 2026-09-18 — **Three corrections and one deployed fix.** (1) The stale-SalesOrder retirement
1999
+ method shipped **missing four FK dependencies** and failed on each in turn: `ItemFulfillments`
2000
+ **header** `salesOrderId` (all 218 orders), the **upstream** `PurchaseOrders_SalesOrders` header
2001
+ bridge (218 rows vs 123 downstream), the **downstream** `SalesOrderItems_PurchaseOrderItems` line
2002
+ bridge (7 vs 669 upstream), and `Invoices.salesOrderId` (1 row, unmovable, throws by design). Two
2003
+ of the four were the **opposite direction of a bridge already handled** — every SO↔PO bridge is
2004
+ two tables. Recorded the method that ended the guessing: query
2005
+ `information_schema.KEY_COLUMN_USAGE` for every FK on the parent (12 + 9 tables) and **count rows
2006
+ on the affected orders** — only 4 of 21 had any. Deployed and verified on Elite prod: stale
2007
+ duplicates 218→187, fulfillments moved (not deleted) 1→37, `PurchaseOrders_TransferOrders` 17→104.
2008
+ (2) **Corrected the "scope was confirmed" claim on the FK-1451 fix** — it had **three** layers,
2009
+ not one: tracking rows → `ItemFulfillmentItemUnits` → the item. Fixed by re-reading units live.
2010
+ **Still failing in prod: committed (`2247a70b`) ≠ running**, proven by comparing the actual call
2011
+ sequence in `Logs_<client>.Api` against what the committed code emits. (3) **Found the
2012
+ empty-window cursor write stores a time ONE HOUR IN THE FUTURE**, for every section and every
2013
+ client — `convertToTimezone()` renders an epoch in Eastern while the server runs Chicago, so an
2014
+ IDLE-looking section syncs nothing forever. **This corrects the 2026-09-09 "whole pipeline is on
2015
+ one clock" claim and the 2026-09-17 "cursor timezone theory ruled out" note** (that note was
2016
+ right about the per-record write only). **NOT FIXED** — shared file, 22 clients. (4) Added the
2017
+ fleet-survey rule: **`864000-RUNNING` is not a health signal — always read the cursor with the
2018
+ mode**; NYCHH was 2.5 years stale while showing a full window. (rgirish)
1827
2019
  - 2026-09-17 — **Cross-client sync-freeze debugging (Compass / Prudential / NYCHH / Elite).**
1828
2020
  (1) **Compass INVOICES** unfrozen (`96000-RUNNING`, `Logs.Issue` 792): `syncInvoiceFromNetsuite`'s
1829
2021
  billed-SO lookup keyed on `c_netsuiteInternalSalesOrderId` with a `==1` guard, but Compass keeps two
@@ -6,7 +6,7 @@ project: _Underscore
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-09-17
9
+ updated: 2026-09-18
10
10
  owners: ["jcardinal", "mhammontree", "tcox", "bala", "apeterson", "rgirish"]
11
11
  files:
12
12
  - api2/Component/Api/V2/V2.php
@@ -182,6 +182,41 @@ recovered to `864000-IDLE` and the cursor moved 2026-02-03 → 2026-03-28.
182
182
  permission with no field permissions is not a partial grant — it is a grant that reads as a
183
183
  successful, empty, count-less response.
184
184
 
185
+ #### ⚠ The 2026-09-17a migration only covered the PO-FIRST siblings — 326/330 have ZERO grants everywhere (found 2026-09-18, NYCHH)
186
+
187
+ `2026-09-17a` granted records **325** (`transfer-orders` PO-first header bridge) and **329**
188
+ (PO-first line bridge). Their **TO-first** siblings were never in the migration at all:
189
+
190
+ | `Core.Records` id | route | state |
191
+ |---|---|---|
192
+ | **326** | `transfer-orders-purchase-orders` | **zero** `AclRecordPermissions` rows |
193
+ | **330** | `transfer-order-items-purchase-order-items` | **zero** `AclRecordPermissions` rows |
194
+
195
+ **Symptom is a HARD denial, unlike Elite's silent one.**
196
+ `GET /v2/transfer-order-items-purchase-order-items` → **403 `EZ-1`**. Contrast with the
197
+ `200 + WZ-1 + no totalRecordCount` shape above: `EZ-1` means the **record** layer is missing
198
+ entirely, `WZ-1` means the record layer is fine and the **field** layer is empty. The two error
199
+ codes tell you which layer to go fix — read the code before you start adding rows.
200
+
201
+ **Found on NYCHH, 2.5 years behind:** `NETSUITE_LAST_SYNC_CURSOR_PURCHASE_ORDERS` at
202
+ **2024-02-16**, `..._INVOICES` at **2024-02-22**, while the execution mode read a healthy-looking
203
+ `864000-RUNNING`.
204
+
205
+ **Red herring worth recording: the error envelope says `api: "Agilant"`.** That looks like a
206
+ *different* API needing *different* grants. It is not — NYCHH's Agilant API carries `roleId` 1 and
207
+ 3, the same as everywhere else, so **roleId 3 is still the right target**. Check the API's roles
208
+ before you chase a second role.
209
+
210
+ **Client-scoped SQL handed to the developer** — 5 statements, the same four-layer chain and
211
+ `NOT EXISTS` guards as `2026-09-17a`: permissions for 326/330, `'all'` expressions, logic groups,
212
+ logic-group expressions, and **6 field grants** — **2214 / 2215 / 2216** for record 326 and
213
+ **2230 / 2231 / 2232** for record 330. `id` (**2213 / 2229**) is **deliberately excluded**, matching
214
+ the 325/329 pattern above. **⚠ NOT YET RUN as of 2026-09-18.**
215
+
216
+ **⚠ This is probably not NYCHH-only.** 326/330 were never in the module migration, so **any** client
217
+ whose sync touches the TO-first bridges will hit the same 403. Treat it as a fleet gap owed a
218
+ follow-up module migration, not a one-tenant patch.
219
+
185
220
  ### ⚠ A multi-client ACL migration must SELF-HEAL — every tenant is broken differently
186
221
 
187
222
  Do not write a fleet-wide ACL migration as "insert the rows the reference client has." Surveying
@@ -858,6 +893,17 @@ hardcoded `Core.RecordFields` id literals instead of a subselect.
858
893
  and every repo is on the **same branch** so the generated model matches the DB.
859
894
 
860
895
  ## Change history
896
+ - 2026-09-18 — **The `2026-09-17a` fleet migration has a gap: it granted only the PO-first bridges
897
+ 325/329, leaving the TO-first siblings 326 (`transfer-orders-purchase-orders`) and 330
898
+ (`transfer-order-items-purchase-order-items`) with ZERO `AclRecordPermissions` rows.** Found on
899
+ NYCHH, which was **2.5 years behind** (PO cursor 2024-02-16, invoices 2024-02-22) while its
900
+ execution mode read a healthy `864000-RUNNING`. Symptom is **403 `EZ-1`** — a hard denial —
901
+ which distinguishes a missing **record** layer from Elite's `200 + WZ-1` missing **field** layer;
902
+ use the error code to pick the layer. Recorded the client-scoped fix (5 guarded statements, field
903
+ grants **2214/2215/2216** for 326 and **2230/2231/2232** for 330, `id` 2213/2229 deliberately
904
+ excluded) and the red herring that the envelope's `api: "Agilant"` does **not** mean a different
905
+ role — that API carries roleId 1 and 3 like everywhere else. **SQL not yet run**, and the gap is
906
+ likely fleet-wide, not NYCHH-only. (rgirish)
861
907
  - 2026-09-17 — Added the **field-layer silent-200** failure shape: record access with **zero**
862
908
  `AclFieldPermissions` rows returns HTTP **200** carrying `WZ-1` and **omits**
863
909
  `meta.totalRecordCount`, so fail-loud code that guards on that key throws while blaming record
@@ -6,8 +6,8 @@ project: _Underscore
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-09-16
10
- owners: [apeterson, bala, jcardinal]
9
+ updated: 2026-09-18
10
+ owners: [apeterson, bala, jcardinal, rgirish]
11
11
  files:
12
12
  - _underscore/Model/Client/PurchaseOrders/SalesOrder.php
13
13
  - _underscore/Model/Client/SalesOrders/PurchaseOrder.php
@@ -81,6 +81,7 @@ All **35** prod client schemas contain all four tables (`SalesOrders`, `Purchase
81
81
  - **Downstream-only assumptions are baked into shared code.** Any tenant whose POs are purely upstream (verified: NYCHH) silently gets `null`/empty from the shared downstream-reading helpers rather than an error.
82
82
  - **⚠ Compass Canada is NOT upstream-only, and it DOES have model overrides** — corrected 2026-08-26. The earlier "no `Model/Compasscanada/` directory" reading looked in the wrong place: the folder is **`_underscore/Model/Compass/Canada/`**, and `_Model_Compass_Canada_SalesOrder extends _Model_Compass_SalesOrder`, so Compass Canada inherits **Compass's** `_purchaseOrders` override, not the base one. Its **upstream bridge is completely empty (0 rows)** while 735 of 788 orders resolve a PO downstream. When checking whether a tenant has overrides, search for the class name (`grep -rn "class _Model_.*_SalesOrder extends"`) rather than guessing a directory — Compass's tenants nest one level deeper than everyone else's.
83
83
  - **⚠ Do NOT paper over the direction with an `ojoin` on a shared single-record fetch.** Tempting because it works in one place: the TOGa Supply NYCHH orders **list** joins the upstream bridge and is correct — but only because NYCHH's upstream bridge happens to be **1:1** (823 rows / 823 distinct sales orders, max 1 PO per SO). The same join in the shared `fetchOrdersDetails` (`toga2-supply/src/pages/Orders/api/OrdersApi.ts:397`) breaks three tenants: **Compass** has up to **6,440** downstream POs on one sales order (row multiplication on a single-record fetch), **Compass Canada's upstream bridge is empty** (735 of 788 would drop to **zero**), and **Prudential** has only **1,868 of 26,843** orders upstream (92% populated would become 7%). A scalar `GROUP_CONCAT` calculated field is the right mechanism precisely because it collapses many POs without multiplying rows.
84
+ - **⚠ When you DELETE a sales order, you must move BOTH directions of BOTH bridges — fixing one and missing its sibling is the default failure (2026-09-18).** Elite's stale-SalesOrder retirement shipped handling only `SalesOrders_PurchaseOrders` (downstream header) and `PurchaseOrderItems_SalesOrderItems` (upstream line), and failed on prod against the two siblings it skipped — `PurchaseOrders_SalesOrders` (upstream header) and `SalesOrderItems_PurchaseOrderItems` (downstream line). **Neither direction is "the main one":** on the same 218 orders, upstream dominated at header level (**218 vs 123**) while downstream was the rare one at line level (**7 vs 669**), so you cannot skip a direction because it looked empty on another tenant. The bridge FKs are `RESTRICT`, so a missed sibling is a hard FK error rather than silent — but you only find it one blocker at a time. **Enumerate instead:** query `information_schema.KEY_COLUMN_USAGE` for every FK referencing the parent, then count rows on the affected records, before writing any delete code. Details on the [per-client sync engine doc](../../../1.0/apps/worker/features/netsuite-togasupply-per-client-sync.md).
84
85
 
85
86
  ## Related
86
87
  - [Sales-order PO Number sourcing](./sales-order-po-number-sourcing.md)
@@ -6,7 +6,7 @@ project: Worker
6
6
  client: elite
7
7
  type: client-feature
8
8
  status: active
9
- updated: 2026-09-17
9
+ updated: 2026-09-18
10
10
  owners: ["snaredla", "jcardinal", "rgirish"]
11
11
  files:
12
12
  - worker/crons/toga2/netsuite/sync_togasupply_elite.php
@@ -339,22 +339,70 @@ ships to **all 22 clients** carrying `netsuite` in `_modules.txt` and is written
339
339
  [ACL permission chain](../../../2.0/apps/_underscore/features/acl-permission-chain.md). **Run so far
340
340
  on `Client_Elite` only.**
341
341
 
342
- ### Elite's stale SalesOrder backfill is queued, not done
342
+ ### Elite's stale SalesOrder backfill — RUNNING, 218 → 187 (2026-09-18)
343
343
 
344
344
  Enabling transfer-order detection left **218** Elite `SalesOrders` that duplicate a `TransferOrder`
345
345
  (**217 of 217** matched pairs identical on order number, line count and total quantity). They still
346
346
  carry fulfillments, which is what drove `_qtyOnHand` negative — see
347
- [inventory quantities](./inventory-quantities-drop-ship.md). The retirement method that cleans them
348
- up is built but **NOT yet deployed**; mechanics on the
347
+ [inventory quantities](./inventory-quantities-drop-ship.md).
348
+
349
+ `retireStaleSalesOrderForTransferOrder()` is **deployed and working**. Verified on prod
350
+ 2026-09-18: stale duplicates **218 → 187**, `SalesOrders` **340 → 309**, fulfillments moved onto
351
+ the TransferOrder side **1 → 37** (moved, never deleted), `PurchaseOrders_TransferOrders`
352
+ **17 → 104**, `PurchaseOrderItems_TransferOrderItems` **28 → 163**.
353
+
354
+ Getting there took **four more FK dependencies** the first version missed — including two that
355
+ were the *opposite direction* of a bridge already handled. Full table, the row counts, and the
356
+ `information_schema` method that found them all at once are on the
357
+ [engine doc](../../../1.0/apps/worker/features/netsuite-togasupply-per-client-sync.md).
358
+
359
+ **Blockers:**
360
+
361
+ 1. ~~SalesOrder **240734** has 4 `InvoiceItems` and will throw~~ — **CLEARED 2026-09-18**, see
362
+ below.
363
+ 2. **11 of the 218** have no matching TransferOrder (4 are TOGa-created `SA1000xx` with no NetSuite
364
+ id). **Still open — manual review.**
365
+
366
+ #### Decision: the blocking invoice on SalesOrder 240734 was DELETED, not kept (2026-09-18)
367
+
368
+ Invoice **43** / number **277274** / NetSuite invoice id **6215693**, 4 `InvoiceItems`, 1
369
+ `InvoiceTrackingNumbers` row, created 2026-09-15.
370
+
371
+ **Why not keep it** — the answer is a documented team decision in the code at
372
+ `library/app/api/toga2.php:3350-3362`:
373
+
374
+ > *"Per team decision we do NOT import invoices billed against a transfer order at all."*
375
+
376
+ `Invoices` has **no `transferOrderId` column**, so such an invoice would import with a null order
377
+ link; `syncInvoiceFromNetsuite()` returns early for them. This invoice was a **leftover from before
378
+ Elite's transfer-order detection was enabled**, when 240734 was still a plain SalesOrder. Under
379
+ today's rule the sync would simply skip it. Deleting it puts the data where the current rule says
380
+ it belongs, and **NetSuite keeps the source record** — nothing is lost.
381
+
382
+ **Delete order matters — children first:** `InvoiceTrackingNumbers` → `InvoiceItems` → `Invoices`.
383
+ Executed and verified in prod: **0 / 0 / 0**. Order 240734 now has 0 invoice items and is clear to
384
+ retire.
385
+
386
+ ### ⚠ OPEN — the item-fulfillment FK fix is committed but NOT running in prod (2026-09-18)
387
+
388
+ `NETSUITE_EXECUTION_MODE_ITEM_FULFILLMENTS` is still failing on item `b96515b3`. The bug turned out
389
+ to have **three** layers (tracking rows → `ItemFulfillmentItemUnits` → the item), and the fix for
390
+ the deepest one is committed (`2247a70b`, working tree clean) — but the running process is
391
+ executing the **old** code. Proven from `Logs_Elite.Api`: the committed code does a
392
+ `GET /v2/item-fulfillment-item-tracking-numbers` immediately before the delete and that GET never
393
+ appears; at 10:27 and 11:07 the sequence was still two plain `GET /v2/item-fulfillment-items` then
394
+ the failing `DELETE`. **Deploy / opcache / checkout-path issue, unresolved.** Mechanics and the
395
+ general diagnostic are on the
349
396
  [engine doc](../../../1.0/apps/worker/features/netsuite-togasupply-per-client-sync.md).
350
397
 
351
- Two blockers to clear by hand first:
398
+ ### ⚠ OPEN — Elite's PO cursor is written ONE HOUR IN THE FUTURE (2026-09-18)
352
399
 
353
- 1. **SalesOrder 240734 has 4 `InvoiceItems` and WILL throw** — `InvoiceItems` has no
354
- `transferOrderItemId`, so those rows cannot be moved. Deliberate fail-loud; it blocks the
355
- backfill until resolved.
356
- 2. **11 of the 218 have no matching TransferOrder** (4 are TOGa-created `SA1000xx` with no NetSuite
357
- id). Manual review.
400
+ Elite's POs sat at a healthy-looking `864000-IDLE` with a current cursor and synced **nothing**.
401
+ Every empty-window write was server time **+60 minutes**, always ending `|0` (10:36:09 → cursor
402
+ 11:36:08). This is an engine-wide bug in the empty-window (`|0`) branch, not an Elite
403
+ configuration problem, and it only appears **after a section catches up** — which is exactly what
404
+ Elite's PO backfill just did. **NOT FIXED** (shared file, 22 clients). Full root cause on the
405
+ [engine doc](../../../1.0/apps/worker/features/netsuite-togasupply-per-client-sync.md).
358
406
 
359
407
  ## Gotchas / known issues
360
408
 
@@ -397,6 +445,18 @@ Two blockers to clear by hand first:
397
445
  rewinding `NETSUITE_LAST_SYNC_DATETIME_SALES_ORDERS`.
398
446
 
399
447
  ## Change history
448
+ - 2026-09-18 — **Stale-SalesOrder backfill is live: 218 → 187.** The retirement method needed
449
+ four more FK dependencies before it would run (fulfillment **header** `salesOrderId`, the
450
+ upstream `PurchaseOrders_SalesOrders`, the downstream `SalesOrderItems_PurchaseOrderItems`, and
451
+ `Invoices.salesOrderId`); fulfillments were **moved** to the TransferOrder side, never deleted
452
+ (1 → 37). **Cleared blocker 1 by deleting invoice 43 / 277274 / NetSuite 6215693** — justified
453
+ by the team decision at `toga2.php:3350-3362` that invoices billed against a transfer order are
454
+ not imported at all (`Invoices` has no `transferOrderId`), so it was a pre-detection leftover;
455
+ deleted children-first and verified 0/0/0. 11 unmatched orders remain. **Two things left OPEN:**
456
+ the three-layer item-fulfillment FK fix is committed (`2247a70b`) but the **running process is
457
+ still on the old code** (proven from the `Logs_Elite.Api` call sequence), and Elite's PO cursor
458
+ is being written **one hour in the future** by the engine's empty-window branch — an IDLE-looking
459
+ section that syncs nothing. (rgirish)
400
460
  - 2026-09-17 — Corrected the stale "transfer orders pre-set but not in use / 0 rows" note in
401
461
  *Elite-only wrapper deviations*: `IS_ENABLED_INTEGRATION_TRANSFER_ORDERS` is now `true`
402
462
  (ZERO_DOLLAR_HOLD) and prod `Client_Elite.TransferOrders` holds 218 rows. (jcardinal)
@@ -8,8 +8,8 @@ project: Worker
8
8
  client: endeavor-health
9
9
  type: profile
10
10
  status: active
11
- updated: 2026-09-16
12
- owners: ["jcardinal"]
11
+ updated: 2026-09-18
12
+ owners: ["jcardinal", "rgirish"]
13
13
  files: []
14
14
  related:
15
15
  - ../../1.0/apps/worker/features/netsuite-togasupply-per-client-sync.md
@@ -23,6 +23,19 @@ Endeavor Health TOGa-Supply NetSuite integration client on the shared 1.0 worker
23
23
  - NetSuite transactions sync into the 2.0 `Client_<Id>` tenant via the TOGa2 API through a thin `worker/crons/toga2/netsuite/sync_togasupply_*.php` wrapper.
24
24
  - Profile seeded 2026-08-29 fixing Endeavor's stuck `ITEM_RECEIPTS` section; item-receipt path migrated SOAP → REST (`fetchItemReceiptById`). Mechanics + diagnosis playbook: [per-client sync](../../1.0/apps/worker/features/netsuite-togasupply-per-client-sync.md).
25
25
 
26
+ ## ⚠ OPEN (2026-09-18) — INVOICES frozen, and it is probably the future-cursor bug
27
+
28
+ The 2026-09-18 fleet sweep found `NETSUITE_EXECUTION_MODE_INVOICES` at **`1-RUNNING`**. After a
29
+ manual reset to `864000-IDLE`, the **window grew back but the cursor stayed frozen** and no errors
30
+ appeared.
31
+
32
+ That is the exact signature of the engine's **empty-window cursor write storing a time one hour in
33
+ the future** — an IDLE-looking section that syncs nothing forever. Root cause, the `|0` tell, and
34
+ the fact that it is **not yet fixed** (shared file, all 22 clients) are on the
35
+ [per-client sync engine doc](../../1.0/apps/worker/features/netsuite-togasupply-per-client-sync.md).
36
+
37
+ **Still frozen as of 2026-09-18** — do not re-diagnose this as an Endeavor-specific problem.
38
+
26
39
  ## Integration touchpoints
27
40
 
28
41
  - **`worker` / `library` (1.0)** — shared togasupply sync engine + `App_Api_Toga2` bridge.
@@ -6,7 +6,7 @@ project: Library
6
6
  client: nychh
7
7
  type: client-feature
8
8
  status: active
9
- updated: 2026-09-16
9
+ updated: 2026-09-18
10
10
  owners: [jcardinal, bala, rgirish]
11
11
  files:
12
12
  - library/app/api/toga2.php
@@ -182,6 +182,22 @@ Two NYCHH-specific gates, both of which bit during this work:
182
182
 
183
183
  WARNING: NYCHH's `NETSUITE_LAST_SYNC_CURSOR_ITEM_RECEIPTS` is stale at **2024-08-06** - re-enabling item receipts replays about two years unless it is rewound forward first.
184
184
 
185
+ ## ⚠ OPEN (2026-09-18) — NYCHH sync is 2.5 YEARS behind on a 403 EZ-1 from records 326/330
186
+
187
+ Found by a fleet sweep of `NETSUITE_EXECUTION_MODE_*` across all 22 netsuite clients. NYCHH was the worst: `NETSUITE_LAST_SYNC_CURSOR_PURCHASE_ORDERS` at **2024-02-16** and `NETSUITE_LAST_SYNC_CURSOR_INVOICES` at **2024-02-22**.
188
+
189
+ **⚠ The mode read `864000-RUNNING` the whole time and was taken as "healthy, in progress".** That reading was wrong — a full window with a 2.5-year-old cursor is **broken, not busy**. Always pair the execution mode with the cursor; neither is a health signal alone.
190
+
191
+ **Error:** `GET /v2/transfer-order-items-purchase-order-items` → **403 `EZ-1`**.
192
+
193
+ **Root cause — a different ACL gap from Elite's.** Records **326** (`transfer-orders-purchase-orders`) and **330** (`transfer-order-items-purchase-order-items`) — the **TO-first** direction — have **zero** `AclRecordPermissions` rows. The `2026-09-17a` module migration only covered the PO-first siblings **325/329**. The error shape differs from Elite's silent `200 + WZ-1`: `EZ-1` means the **record** layer is missing, `WZ-1` means the record layer is fine and the **field** layer is empty — use the code to pick which layer to fix.
194
+
195
+ **Red herring:** the envelope reports `api: "Agilant"`, which looks like a different API needing different grants. It is not — NYCHH's Agilant API carries `roleId` **1 and 3**, so roleId 3 is still correct.
196
+
197
+ **Fix prepared, ⚠ NOT YET RUN:** client-scoped SQL, 5 statements, same four-layer chain and `NOT EXISTS` guards as `2026-09-17a` — permissions for 326/330, `'all'` expressions, logic groups, logic-group expressions, and 6 field grants (**2214/2215/2216** for 326; **2230/2231/2232** for 330; `id` **2213/2229** deliberately excluded, matching 325/329). Chain mechanics on the [ACL permission chain](../../../2.0/apps/_underscore/features/acl-permission-chain.md).
198
+
199
+ **Not NYCHH-only.** 326/330 were never in the module migration, so any client whose sync touches the TO-first bridges will hit the same 403.
200
+
185
201
  ## Gotchas / known issues
186
202
 
187
203
  - **🚩 The import sends almost everything to catch-all destination location 2 "NYC Health + Hospitals" — 1,832 of 1,939 transfer orders**, next highest Coney Island at 27, and **every one has `createdByUserId` NULL** (277 on 2026-08-29, 1,528 on 2026-08-30, then 1–3/day). Whether that location is a real receiving place or the import should resolve the actual hospital is **open with the PM**; its address is currently Jacobi's. Do not assume the destination is correct: [location shipping addresses](./location-shipping-addresses.md).
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "toga-ai",
3
- "version": "1.0.844",
3
+ "version": "1.0.845",
4
4
  "description": "TOGA Technology Team Claude Knowledge System — shared AI coding harness with skills, knowledge base CLI, and project installer for Claude Code.",
5
5
  "keywords": [
6
6
  "claude",