toga-ai 1.0.844 → 1.0.846

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -6,7 +6,7 @@ project: Library
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-09-16
9
+ updated: 2026-09-18
10
10
  owners: ["jcardinal", "mhammontree"]
11
11
  files:
12
12
  - library/app/error/capture.php
@@ -116,6 +116,7 @@ Two things are specifically 1.0:
116
116
 
117
117
  ## Gotchas
118
118
 
119
+ - **Looking for a 1.0 cron fatal? Query `Logs.Issue` on environment `prod`, ordered by `dtLastOccurred`.** The legacy `Logs.Errors` and `Vision_Log.ErrorLog` tables do **not** carry 1.0 cron fatals — searching them wastes time (2026-09-18, DOE outbound-cron outage). Uncaught exceptions reach here via `App_Error::handleException` → `captureException()`, which returns the short quotable reference. **A `catch` BYPASSES capture** — if you catch per record, call `App_Error_Capture::captureException($e)` yourself, and put only the returned reference (never `$e->getMessage()`, which can carry a raw vendor response body) into an email or log.
119
120
  - **A `*/` inside PHP docblock prose silently closes the comment early.** Writing something like `list*/fetch*` in a `/** … */` block ends the docblock at the `*/`, and `php -l` then reports a confusing `syntax error, unexpected token` on a later line. Hit while documenting the NetSuite throw sites. Avoid the literal `*/` in docblock text.
120
121
  - **⚠ A grouped `Logs.Issue` row can show a STALE error id — read the live api2 transaction log for the current error.** Two facts combine to mislead: (1) `Event.errorMessage` is **`varchar(255)`**, so a long API response is **truncated** — often before the trailing error id; and (2) `Logs.Issue.subject` / `Issue.errorMessage` hold the **first-occurrence** text, so a long-lived issue keeps showing the error id from whenever it was first seen even after the per-call error has changed. Observed 2026-08-17: issue `3H` still displayed `EO-1` / `3G-6` (the original undefined-method fatal) while the current per-call failure was actually `EV-11` (an FK 1451 on a DELETE). **To see what is failing *now*, read the live api2 transaction log (`Logs_<client>.Api.responsePayload`) for the specific call — not the grouped `Issue` row.**
121
122
  - **Any PHP warning inside capture code would kill the request.** 1.0's `App_Error::handleError()` routes **every** PHP warning into `handleException()`, which renders the red box and calls `exit()` — so a single warning raised while capturing (an unreachable IMDS endpoint is the everyday case: `file_get_contents` warns before returning `false`) terminates the request or cron mid-run while reporting an unrelated error. 2.0's equivalent handler *throws*, which a `catch` absorbs. `captureException()` therefore sets `App_Error::setThrowExceptionsEnabled(false)` for its own duration and restores it in `finally`. This disables the **handler**, not real throws — `App_Query` raises a plain `throw new Exception` on DB errors, so genuine failures still reach the catch. **This is a general 1.0 trap for any code that must not escalate**, not just error capture.
@@ -6,7 +6,7 @@ project: Worker
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-09-17
9
+ updated: 2026-09-18
10
10
  owners: ["dfranks", "bala", "jcardinal", "mhammontree", "snaredla", "rgirish"]
11
11
  files:
12
12
  - worker/crons/toga2/netsuite/common_sync_togasupply.php
@@ -106,8 +106,13 @@ The value is `<seconds>-<STATE>`. Three readings, and only the first is a real f
106
106
 
107
107
  - **`1-RUNNING`** (window collapsed ÷3 to 1s) = throwing on **every** run = real freeze. Go read the
108
108
  error (`Logs.Event` by clientId → file:line → `Logs_<Client>.Api` on the transactionId).
109
- - **A large-window `*-RUNNING`** (e.g. `864000-RUNNING`, `288000-RUNNING`) = healthy mid-run **or** a
110
- deliberate pause/backfill — **NOT** stuck. Do not "fix" it.
109
+ - **A large-window `*-RUNNING`** (e.g. `864000-RUNNING`, `288000-RUNNING`) = **usually** healthy
110
+ mid-run or a deliberate pause/backfill. **⚠ CORRECTED 2026-09-18 — this is NOT sufficient on its
111
+ own, and the old "NOT stuck, do not fix it" wording caused a real miss.** NYCHH showed
112
+ `864000-RUNNING` while its cursor sat **2.5 years** stale, and the session skipped it on that
113
+ reading. **Always read `NETSUITE_LAST_SYNC_CURSOR_<SECTION>` together with the mode.** A full
114
+ window with an ancient cursor is **broken, not busy** — see the future-cursor bug below, which
115
+ produces exactly this healthy-looking shape.
111
116
  - **`IDLE` + cursor frozen exactly at a reset value + NO error events = the section toggle is off**
112
117
  (`IS_ENABLED_INTEGRATION_<SECTION> = false` in the wrapper). Resetting the cursor does **nothing**
113
118
  while the toggle is off. Worked example: NYCHH SALES_ORDERS looked "stuck" after a cursor reset to
@@ -322,6 +327,70 @@ hit **every section and every client**, not just the by-reference ones.
322
327
  GET's UTC. `strtotime()` and `App_Date::convertToSQLDate()` both use the script tz (Chicago) and are
323
328
  the trap. Any manual cursor rewind value you set by hand must also be **Eastern**.
324
329
 
330
+ ### ⚠ …and the empty-window write is STILL WRONG — it stores a cursor ONE HOUR IN THE FUTURE (found 2026-09-18, Elite prod — NOT FIXED)
331
+
332
+ **This CORRECTS the section directly above.** The 2026-09-09 fix made the *read* and the
333
+ *per-record write* agree on Eastern, and its claim that "the **whole** cursor pipeline runs on one
334
+ clock" is **wrong**. The empty-window (`|0`) write it introduced has the opposite error, and it is
335
+ live in every section for every client.
336
+
337
+ - **The symptom is a section that looks perfectly healthy.** Elite's POs sat at `864000-IDLE` with a
338
+ recent cursor and synced **nothing**. No error, no shrinking window, no alert. This is the
339
+ hardest failure shape in the engine to spot, because every dashboard signal says fine.
340
+ - **Root cause — a double Eastern conversion.** `common_sync_togasupply.php:945` (POs) writes
341
+ `App_Date::convertToTimezone($timeNow, 'Y-m-d H:i:s', 'America/New_York')`. `$timeNow` is an
342
+ **epoch**, and `convertToTimezone` renders it in Eastern. The server clock is
343
+ **America/Chicago**, so the stored string is **local time + 1 hour**. The next run reads that
344
+ string back as the window **start**, converts it to Eastern **again** for the NetSuite query, and
345
+ asks for records modified after a moment that has not happened yet.
346
+ - **It is self-sustaining.** Nothing matches → the window is empty → the `|0` branch writes another
347
+ future cursor. Forever.
348
+ - **Prod proof (Elite `NETSUITE_LAST_SYNC_CURSOR_PURCHASE_ORDERS`), every write exactly +60 min and
349
+ always ending `|0`:**
350
+
351
+ | Written at (server, Chicago) | Cursor stored |
352
+ |---|---|
353
+ | 10:36:09 | 11:36:08 |
354
+ | 10:26:10 | 11:26:08 |
355
+ | 10:22:41 | 11:22:40 |
356
+
357
+ - **`|0` is the tell.** The `|0` suffix marks the empty-window branch
358
+ (`$countProcessedRecords === 0`). A cursor that is both `|0` **and** in the future is this bug.
359
+ - **Scope: every section, every client.** The same line pattern is at **838** (sales orders),
360
+ **945** (purchase orders), **1031** (invoices), **1200** (item receipts) and the fulfillments
361
+ section. It only bites once a section **catches up** and starts hitting empty windows — which is
362
+ exactly what happened when Elite's PO backfill finished. Likely also explains **Endeavor Health**
363
+ (window growing, cursor frozen, no errors).
364
+ - **⚠ OPEN — NOT FIXED.** `common_sync_togasupply.php` is shared by all 22 netsuite clients and the
365
+ developer has not given the go-ahead. Read-only investigation as of 2026-09-18.
366
+ - **The nuance the 2026-09-17 "ruled out" note got right, and what it got wrong.** The earlier
367
+ entry that recorded "a cursor timezone theory — **wrong**" was right about the **per-record**
368
+ write (the `common` line ~821 path, which correctly uses the NetSuite `lastModifiedDate`). It was
369
+ wrong as a blanket statement. **Both are true: the per-record write is correct; the empty-window
370
+ write is broken.** Do not use one to dismiss the other.
371
+
372
+ #### Reading `NETSUITE_EXECUTION_MODE_*`: the mode alone is NOT enough — always read the cursor with it
373
+
374
+ How to survey the fleet for stuck clients (client list:
375
+ `test/team/2.0 deployment/NetSuite_Clients_Db.txt`, 22 netsuite clients):
376
+
377
+ | Value | Reading |
378
+ |---|---|
379
+ | `864000-IDLE` | healthy |
380
+ | a **shrinking** number with `RUNNING` | a section that is throwing — go read `Logs.Issue` |
381
+ | `<max>-RUNNING` | usually just mid-run — **but check the cursor before you believe it** |
382
+
383
+ **⚠ CORRECTION (2026-09-18): `864000-RUNNING` was read as "healthy, in progress" and skipped. That
384
+ was wrong.** NYCHH showed `864000-RUNNING` while its cursor was **2.5 years stale**. A full-window
385
+ `RUNNING` with an ancient cursor is **broken, not busy**. Always pair the mode with
386
+ `NETSUITE_LAST_SYNC_CURSOR_<section>` — neither value is a health signal on its own.
387
+
388
+ **Fleet sweep result, 2026-09-18** (all 22 clients): Endeavor Health `INVOICES 1-RUNNING`, AIG
389
+ `SALES_ORDERS 8-RUNNING`, Compass `INVOICES 96000-RUNNING`, NYCHH both sections 2.5 years behind,
390
+ Elite known. The other 17 were clean. After resetting the modes to `864000-IDLE`, **AIG and Compass
391
+ recovered on their own**; **Endeavor Health's window grew but its cursor stayed frozen** — the
392
+ future-cursor signature above.
393
+
325
394
  ### ⚠ ITEM_RECEIPTS froze at RUNNING on a case-mismatched Units serial → duplicate-unit INSERT (EV-10 / 1062) (found + fixed 2026-09-11, NYCHH prod)
326
395
 
327
396
  `syncItemReceiptFromNetsuite()` (`library/app/api/toga2.php` ~L4022) builds a PHP array of the
@@ -732,7 +801,7 @@ nested count:
732
801
  line the fulfillment referenced, the match failed, and the section fail-loud-froze. The flat list
733
802
  replaces the truncated nested one before match/reconstruct.
734
803
 
735
- ### Retiring the stale SalesOrder a reclassified TransferOrder leaves behind (2026-09-17, NOT YET DEPLOYED)
804
+ ### Retiring the stale SalesOrder a reclassified TransferOrder leaves behind (DEPLOYED + RUNNING 2026-09-18)
736
805
 
737
806
  **Turning transfer-order detection on for an existing client strands every order it already
738
807
  imported as a SalesOrder.** The same NetSuite order now syncs to a `TransferOrder`, but the old
@@ -741,15 +810,25 @@ and `Items._qtyOnHand` goes negative. Measured on Elite: **218** SalesOrders dup
741
810
  TransferOrder on `c_netsuiteInternalSalesOrderId`, and **217 of 217** matched pairs are identical on
742
811
  order number, line count *and* total quantity.
743
812
 
744
- `App_Api_Toga2::retireStaleSalesOrderForTransferOrder()` (new, `private static`) is called at the
813
+ `App_Api_Toga2::retireStaleSalesOrderForTransferOrder()` (`private static`) is called at the
745
814
  **END** of `syncTransferOrderFromNetsuite()` so the TO and its lines exist first. Per SO line:
746
815
 
747
816
  1. move `ItemFulfillmentItems`: `salesOrderItemId` → `transferOrderItemId`,
748
- 2. move `PurchaseOrderItems_SalesOrderItems` → `PurchaseOrderItems_TransferOrderItems`,
749
- 3. `DELETE` the `SalesOrderItem`;
817
+ 2. move `PurchaseOrderItems_SalesOrderItems` → `PurchaseOrderItems_TransferOrderItems` (upstream),
818
+ 3. **(b2)** move `SalesOrderItems_PurchaseOrderItems` → `PurchaseOrderItems_TransferOrderItems`
819
+ (downstream — added 2026-09-18),
820
+ 4. `DELETE` the `SalesOrderItem`;
821
+
822
+ then at header level:
823
+
824
+ 5. **(d2)** collapse **both** `SalesOrders_PurchaseOrders` (downstream) **and**
825
+ `PurchaseOrders_SalesOrders` (upstream) onto `PurchaseOrders_TransferOrders`
826
+ — upstream added 2026-09-18,
827
+ 6. **(e)** repoint `ItemFulfillments.salesOrderId` → `transferOrderId` (the fulfillment **header**,
828
+ added 2026-09-18),
829
+ 7. `DELETE` the `SalesOrder`.
750
830
 
751
- then move `SalesOrders_PurchaseOrders` → `PurchaseOrders_TransferOrders` and `DELETE` the
752
- `SalesOrder`. **Move before delete, always** — the bridge FKs are `RESTRICT`.
831
+ **Move before delete, always** — the bridge FKs are `RESTRICT`.
753
832
 
754
833
  **The two fulfillment-parent columns are mutually exclusive, which is what makes the move safe.**
755
834
  Prod check: 391 rows TO-side, 293 SO-side, **0 with both, 0 with neither**.
@@ -781,6 +860,68 @@ fleet-wide grant is `dbchanges2/_modules/netsuite/2026-09-17a - StaleSalesOrderR
781
860
  load-bearing for this method's safety. Skipped for now because Elite TOs are 1–5 lines; flagged by
782
861
  php-reviewer.
783
862
 
863
+ #### ⚠ It shipped missing FOUR more FK dependencies — and the fix for that is a QUERY, not more guessing (2026-09-18)
864
+
865
+ The method deployed, ran, and then failed on **one blocker after another**, each revealed only by
866
+ fixing the one before it. Found and fixed in the order they surfaced:
867
+
868
+ | # | Missing dependency | Rows on the 218 stale orders | Resolution |
869
+ |---|---|---|---|
870
+ | a | `ItemFulfillments.salesOrderId` — the fulfillment **HEADER** | **all 218** | repoint to the TransferOrder (step e) |
871
+ | b | `PurchaseOrders_SalesOrders` — the **upstream** header bridge | **218** (vs 123 downstream) | step d2 |
872
+ | c | `SalesOrderItems_PurchaseOrderItems` — the **downstream** line bridge | 7 (vs 669 upstream) | step b2 |
873
+ | d | `Invoices.salesOrderId` | 1 (SalesOrder 240734) | **cannot** be moved — `Invoices` has no `transferOrderId`; correctly throws |
874
+
875
+ **(a) would have failed every single order.** `ItemFulfillments` carries its own
876
+ `salesOrderId`/`transferOrderId` pair — the same either-or as the line table (prod: 334 rows, all
877
+ SO-side, **0 with both, 0 with neither**). The original method only moved the fulfillment **items**.
878
+ The fulfillment itself is real shipping history and is **never deleted** — only repointed.
879
+
880
+ **The pattern behind (b) and (c): every SO↔PO bridge is TWO tables, and it is easy to fix one and
881
+ miss its sibling.**
882
+
883
+ | Handled first | Missed sibling |
884
+ |---|---|
885
+ | `SalesOrders_PurchaseOrders` | `PurchaseOrders_SalesOrders` |
886
+ | `PurchaseOrderItems_SalesOrderItems` | `SalesOrderItems_PurchaseOrderItems` |
887
+
888
+ Note the volumes cut **both ways** — upstream dominates at header level (218 vs 123) and downstream
889
+ is the rare one at line level (7 vs 669). Neither direction is "the main one"; you cannot skip one
890
+ because it looked small on another client. Direction semantics are on the
891
+ [SO↔PO bridge direction doc](../../../2.0/apps/_underscore/features/sales-order-purchase-order-bridge-direction.md).
892
+
893
+ **⚠ THE METHOD THAT ENDED THE WHACK-A-MOLE — make this step 1 of any "delete a parent record" work,
894
+ not step 5.** Stop guessing which children exist. Ask the schema, then count:
895
+
896
+ ```sql
897
+ -- 1. every FK pointing at the parent (and its line table)
898
+ SELECT TABLE_NAME, COLUMN_NAME, REFERENCED_TABLE_NAME
899
+ FROM information_schema.KEY_COLUMN_USAGE
900
+ WHERE TABLE_SCHEMA = 'Client_Elite'
901
+ AND REFERENCED_TABLE_NAME IN ('SalesOrders', 'SalesOrderItems');
902
+
903
+ -- 2. then COUNT actual rows in each, restricted to the affected orders
904
+ ```
905
+
906
+ That pair of queries found **12** tables referencing `SalesOrders` and **9** referencing
907
+ `SalesOrderItems`. Only the **4** above had any rows. The other **15 were empty** and needed no
908
+ code at all: `Entitlements_SalesOrders`, `SalesOrderEmailAddresses`, `SalesOrderNotes`,
909
+ `SalesOrders_Payments`, `SalesOrders_SalesSupportUsers`, `SalesOrders_TransferOrders`,
910
+ `TransferOrders_SalesOrders`, `SalesOrderItems_CommittedUnits`, `SalesOrderItems_Tags`, child lines
911
+ via `parentSalesOrderItemId`, `SalesOrderItems_TransferOrderItems`,
912
+ `TransferOrderItems_SalesOrderItems`. **Counting first is what tells you which of the 21 you
913
+ actually have to write code for** — the schema alone would have sent us after all 21.
914
+
915
+ **Verified working in Elite prod, 2026-09-18:**
916
+
917
+ | Measure | Before → after |
918
+ |---|---|
919
+ | stale duplicate SalesOrders | 218 → **187** |
920
+ | `SalesOrders` total | 340 → **309** |
921
+ | fulfillments on the TransferOrder side | 1 → **37** (moved, **not** deleted) |
922
+ | `PurchaseOrders_TransferOrders` | 17 → **104** |
923
+ | `PurchaseOrderItems_TransferOrderItems` | 28 → **163** |
924
+
784
925
  ### ⚠ TWO code paths create `PurchaseOrderItems` — only one stamped `fulfillmentType` (fixed 2026-09-17)
785
926
 
786
927
  `fulfillmentType` was stamped by `syncPurchaseOrderFromNetsuite()` (from the order-level
@@ -822,13 +963,39 @@ Fix, applied at **both** item-level delete sites: **re-read the item's tracking
822
963
  before deleting, instead of trusting the run-start snapshot — the same re-fetch pattern this file
823
964
  already uses for the single-tracking-number fan-out. A `totalRecordCount` guard was added to both.
824
965
 
825
- **Scope was confirmed, not assumed:** every FK-1451 failure that day was on
826
- `/v2/item-fulfillment-items/` — **zero** on units or tracking rows — so the unit-level lists were
827
- deliberately left alone.
966
+ **⚠ "Scope was confirmed, not assumed" was WRONG — there were THREE layers, not one
967
+ (corrected 2026-09-18).** Clearing `ItemFulfillmentItems_TrackingNumbers` only exposed the next
968
+ layer down: `ItemFulfillmentItemUnits.itemFulfillmentItemId`. Same root cause each time — a
969
+ run-start snapshot used to delete children while a later pass **in the same run** creates more of
970
+ them. Confirmed on the blocking row (item `b96515b3`): **trackingRows 1, unitRows 1,
971
+ unitTrackingRows 1** — each layer hidden behind the one above it. The "zero failures on units"
972
+ observation was true *at that moment* only because the tracking layer failed first and never let
973
+ the run reach the units.
974
+
975
+ **Second fix:** the stale-item delete block now re-reads units **live** via the existing paginated
976
+ helper `getItemFulfillmentItemUnitsByItemFulfillmentUuid()`, then deletes, in order, each unit's
977
+ tracking bridge rows → the unit → the item.
978
+
979
+ **Lesson: with nested children, fixing the FK you can see just reveals the next one. Enumerate the
980
+ whole child tree up front** — the same `information_schema.KEY_COLUMN_USAGE` + row-count method
981
+ written up under the stale-SalesOrder retirement section above.
828
982
 
829
983
  The exception was **not** swallowed. `send()` throws by default; the throw aborted the run and left
830
984
  the mode `RUNNING`. That is the designed fail-loud behaviour working correctly.
831
985
 
986
+ **⚠ OPEN as of 2026-09-18 — the fix is committed but the running process is NOT executing it.**
987
+ The unit-layer fix is committed (`2247a70b`, working tree clean) and Elite still fails in prod.
988
+
989
+ **The diagnostic that proves "committed ≠ running": compare the ACTUAL API call sequence in
990
+ `Logs_<client>.Api` against what the committed code would emit.** The committed code performs a
991
+ `GET /v2/item-fulfillment-item-tracking-numbers` immediately before the delete. That GET **does not
992
+ appear** in the log. At 10:27 and again at 11:07 the sequence was still two plain
993
+ `GET /v2/item-fulfillment-items` followed by the failing `DELETE` — i.e. the **old** code path.
994
+
995
+ This is a deploy / opcache / checkout-path problem, **unresolved**. Use this technique before
996
+ re-debugging any "fix that didn't work": the API log is a faithful trace of which code is actually
997
+ running, and it settles the question in one query.
998
+
832
999
  **General rule for this engine: any list used to delete children must be re-read immediately before
833
1000
  the delete if ANY later pass in the same run can create more of them.** A run-start snapshot is only
834
1001
  safe for read-only use.
@@ -841,7 +1008,11 @@ order — record these so nobody re-walks them:
841
1008
  1. backfill lag — no,
842
1009
  2. the `property_exists` gate — no,
843
1010
  3. ACL on the field — grants exist (`recordFieldId` **2594/2595**, roleId 3, `isWritable = 1`),
844
- 4. a cursor timezone theory — **wrong**; the code correctly uses `America/New_York`.
1011
+ 4. a cursor timezone theory — **wrong *for this symptom*, and only for the PER-RECORD write**;
1012
+ that path correctly uses `America/New_York`. **⚠ Do not read this as "cursor timezones are
1013
+ fine" — corrected 2026-09-18:** the **empty-window (`|0`) write** genuinely is broken and
1014
+ stores a cursor one hour in the future. See the future-cursor section above. Both facts are
1015
+ true at once.
845
1016
 
846
1017
  **Actual reason:** those orders are now classified as **Transfer Orders** (`$0` + `holdInvoice`), so
847
1018
  the sync routes them to `syncTransferOrderFromNetsuite()` and they never reach the sales-order path
@@ -1824,6 +1995,27 @@ library (or vice versa) crashes GroWrk and Adyen on their next sync run.
1824
1995
  [NetSuite Sync Alert Monitor](../../library/features/netsuite-sync-alert-monitor.md).
1825
1996
 
1826
1997
  ## Change history
1998
+ - 2026-09-18 — **Three corrections and one deployed fix.** (1) The stale-SalesOrder retirement
1999
+ method shipped **missing four FK dependencies** and failed on each in turn: `ItemFulfillments`
2000
+ **header** `salesOrderId` (all 218 orders), the **upstream** `PurchaseOrders_SalesOrders` header
2001
+ bridge (218 rows vs 123 downstream), the **downstream** `SalesOrderItems_PurchaseOrderItems` line
2002
+ bridge (7 vs 669 upstream), and `Invoices.salesOrderId` (1 row, unmovable, throws by design). Two
2003
+ of the four were the **opposite direction of a bridge already handled** — every SO↔PO bridge is
2004
+ two tables. Recorded the method that ended the guessing: query
2005
+ `information_schema.KEY_COLUMN_USAGE` for every FK on the parent (12 + 9 tables) and **count rows
2006
+ on the affected orders** — only 4 of 21 had any. Deployed and verified on Elite prod: stale
2007
+ duplicates 218→187, fulfillments moved (not deleted) 1→37, `PurchaseOrders_TransferOrders` 17→104.
2008
+ (2) **Corrected the "scope was confirmed" claim on the FK-1451 fix** — it had **three** layers,
2009
+ not one: tracking rows → `ItemFulfillmentItemUnits` → the item. Fixed by re-reading units live.
2010
+ **Still failing in prod: committed (`2247a70b`) ≠ running**, proven by comparing the actual call
2011
+ sequence in `Logs_<client>.Api` against what the committed code emits. (3) **Found the
2012
+ empty-window cursor write stores a time ONE HOUR IN THE FUTURE**, for every section and every
2013
+ client — `convertToTimezone()` renders an epoch in Eastern while the server runs Chicago, so an
2014
+ IDLE-looking section syncs nothing forever. **This corrects the 2026-09-09 "whole pipeline is on
2015
+ one clock" claim and the 2026-09-17 "cursor timezone theory ruled out" note** (that note was
2016
+ right about the per-record write only). **NOT FIXED** — shared file, 22 clients. (4) Added the
2017
+ fleet-survey rule: **`864000-RUNNING` is not a health signal — always read the cursor with the
2018
+ mode**; NYCHH was 2.5 years stale while showing a full window. (rgirish)
1827
2019
  - 2026-09-17 — **Cross-client sync-freeze debugging (Compass / Prudential / NYCHH / Elite).**
1828
2020
  (1) **Compass INVOICES** unfrozen (`96000-RUNNING`, `Logs.Issue` 792): `syncInvoiceFromNetsuite`'s
1829
2021
  billed-SO lookup keyed on `c_netsuiteInternalSalesOrderId` with a `==1` guard, but Compass keeps two
@@ -6,10 +6,12 @@ project: Worker
6
6
  client: shared
7
7
  type: workflow
8
8
  status: active
9
- updated: 2026-09-16
10
- owners: ["bala", "kyalamarthi", "rgirish"]
9
+ updated: 2026-09-18
10
+ owners: ["bala", "kyalamarthi", "rgirish", "mhammontree"]
11
11
  files:
12
12
  - worker/ebs/cron.worker.php
13
+ - library/app/error.php
14
+ - library/app/error/capture.php
13
15
  - worker/schedules/cron.worker.sync.json
14
16
  - worker/schedules/cron.worker.infrastructure.json
15
17
  - worker/.ebextensions/009_setup_phpini.config
@@ -78,6 +80,16 @@ Query the legacy env's `Common.CronJobExecutions` filtering `job LIKE` the scrip
78
80
 
79
81
  **The overlap guard is a `ps -ef` substring match — anything can block it.** `App_Framework::isProcessRunning()` returns true for **any** `ps -ef` line containing the script path (it only excludes `/bin/sh` lines and its own pid). So a developer's `tail -f`, an editor, or any shell holding that path in its command line **blocks the cron indefinitely** while looking like a legitimate overlap. Check `ps -ef` for what is actually holding the name before assuming a long-running instance of the job itself.
80
82
 
83
+ ### Step 4b — a cron that DIED mid-run: the fatal is in the shared 2.0 `Logs.Issue`, not the legacy log tables
84
+ A `dtCheckIn` with **no `dtCheckOut`** (Step 4) means it crashed. Find out why here:
85
+
86
+ - 1.0 uncaught exceptions route through `App_Error::handleException` → `App_Error_Capture::captureException()`, which records the error in the **shared 2.0 `Logs.Issue`** table with a short quotable reference (e.g. `X8`) plus `errorMessage` and trace.
87
+ - Query `Logs.Issue` on **environment `prod`** (**not** `legacy`), ordered by `dtLastOccurred`.
88
+ - **The legacy `Logs.Errors` and `Vision_Log.ErrorLog` tables do NOT carry 1.0 cron fatals** — searching them wastes time. Worked example: a 2026-09-18 DOE outbound-cron outage was found this way only after both legacy tables came back empty.
89
+ - Mechanics, opt-in config and gotchas: [Error capture in 1.0](../../library/features/error-capture-1-0.md).
90
+
91
+ > Remember the notice trap: any PHP notice/warning **terminates** a 1.0 cron (see the 1.0 back-end standard), so "crashed mid-run" often means a notice, not a thrown exception.
92
+
81
93
  ### Step 5 — if runs overlap, look at what the job does before its real work
82
94
  A job whose own runtime exceeds its schedule interval silently loses most of its slots to the guard. Measure the phases, not the total: in the ODP case an instrumented run showed the S3 listing alone returning **138,472 objects and taking 32 seconds** before any PO was touched, with each PO then costing roughly 40 seconds — comfortably past a 5-minute schedule. Prefixes the job skips while processing are still listed, so accumulated `SENT/` and `OUTBOX/` objects inflate every run.
83
95
 
@@ -6,7 +6,7 @@ project: _Underscore
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-09-17
9
+ updated: 2026-09-18
10
10
  owners: ["jcardinal", "mhammontree", "tcox", "bala", "apeterson", "rgirish"]
11
11
  files:
12
12
  - api2/Component/Api/V2/V2.php
@@ -182,6 +182,41 @@ recovered to `864000-IDLE` and the cursor moved 2026-02-03 → 2026-03-28.
182
182
  permission with no field permissions is not a partial grant — it is a grant that reads as a
183
183
  successful, empty, count-less response.
184
184
 
185
+ #### ⚠ The 2026-09-17a migration only covered the PO-FIRST siblings — 326/330 have ZERO grants everywhere (found 2026-09-18, NYCHH)
186
+
187
+ `2026-09-17a` granted records **325** (`transfer-orders` PO-first header bridge) and **329**
188
+ (PO-first line bridge). Their **TO-first** siblings were never in the migration at all:
189
+
190
+ | `Core.Records` id | route | state |
191
+ |---|---|---|
192
+ | **326** | `transfer-orders-purchase-orders` | **zero** `AclRecordPermissions` rows |
193
+ | **330** | `transfer-order-items-purchase-order-items` | **zero** `AclRecordPermissions` rows |
194
+
195
+ **Symptom is a HARD denial, unlike Elite's silent one.**
196
+ `GET /v2/transfer-order-items-purchase-order-items` → **403 `EZ-1`**. Contrast with the
197
+ `200 + WZ-1 + no totalRecordCount` shape above: `EZ-1` means the **record** layer is missing
198
+ entirely, `WZ-1` means the record layer is fine and the **field** layer is empty. The two error
199
+ codes tell you which layer to go fix — read the code before you start adding rows.
200
+
201
+ **Found on NYCHH, 2.5 years behind:** `NETSUITE_LAST_SYNC_CURSOR_PURCHASE_ORDERS` at
202
+ **2024-02-16**, `..._INVOICES` at **2024-02-22**, while the execution mode read a healthy-looking
203
+ `864000-RUNNING`.
204
+
205
+ **Red herring worth recording: the error envelope says `api: "Agilant"`.** That looks like a
206
+ *different* API needing *different* grants. It is not — NYCHH's Agilant API carries `roleId` 1 and
207
+ 3, the same as everywhere else, so **roleId 3 is still the right target**. Check the API's roles
208
+ before you chase a second role.
209
+
210
+ **Client-scoped SQL handed to the developer** — 5 statements, the same four-layer chain and
211
+ `NOT EXISTS` guards as `2026-09-17a`: permissions for 326/330, `'all'` expressions, logic groups,
212
+ logic-group expressions, and **6 field grants** — **2214 / 2215 / 2216** for record 326 and
213
+ **2230 / 2231 / 2232** for record 330. `id` (**2213 / 2229**) is **deliberately excluded**, matching
214
+ the 325/329 pattern above. **⚠ NOT YET RUN as of 2026-09-18.**
215
+
216
+ **⚠ This is probably not NYCHH-only.** 326/330 were never in the module migration, so **any** client
217
+ whose sync touches the TO-first bridges will hit the same 403. Treat it as a fleet gap owed a
218
+ follow-up module migration, not a one-tenant patch.
219
+
185
220
  ### ⚠ A multi-client ACL migration must SELF-HEAL — every tenant is broken differently
186
221
 
187
222
  Do not write a fleet-wide ACL migration as "insert the rows the reference client has." Surveying
@@ -858,6 +893,17 @@ hardcoded `Core.RecordFields` id literals instead of a subselect.
858
893
  and every repo is on the **same branch** so the generated model matches the DB.
859
894
 
860
895
  ## Change history
896
+ - 2026-09-18 — **The `2026-09-17a` fleet migration has a gap: it granted only the PO-first bridges
897
+ 325/329, leaving the TO-first siblings 326 (`transfer-orders-purchase-orders`) and 330
898
+ (`transfer-order-items-purchase-order-items`) with ZERO `AclRecordPermissions` rows.** Found on
899
+ NYCHH, which was **2.5 years behind** (PO cursor 2024-02-16, invoices 2024-02-22) while its
900
+ execution mode read a healthy `864000-RUNNING`. Symptom is **403 `EZ-1`** — a hard denial —
901
+ which distinguishes a missing **record** layer from Elite's `200 + WZ-1` missing **field** layer;
902
+ use the error code to pick the layer. Recorded the client-scoped fix (5 guarded statements, field
903
+ grants **2214/2215/2216** for 326 and **2230/2231/2232** for 330, `id` 2213/2229 deliberately
904
+ excluded) and the red herring that the envelope's `api: "Agilant"` does **not** mean a different
905
+ role — that API carries roleId 1 and 3 like everywhere else. **SQL not yet run**, and the gap is
906
+ likely fleet-wide, not NYCHH-only. (rgirish)
861
907
  - 2026-09-17 — Added the **field-layer silent-200** failure shape: record access with **zero**
862
908
  `AclFieldPermissions` rows returns HTTP **200** carrying `WZ-1` and **omits**
863
909
  `meta.totalRecordCount`, so fail-loud code that guards on that key throws while blaming record
@@ -6,8 +6,8 @@ project: _Underscore
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-09-16
10
- owners: [apeterson, bala, jcardinal]
9
+ updated: 2026-09-18
10
+ owners: [apeterson, bala, jcardinal, rgirish]
11
11
  files:
12
12
  - _underscore/Model/Client/PurchaseOrders/SalesOrder.php
13
13
  - _underscore/Model/Client/SalesOrders/PurchaseOrder.php
@@ -81,6 +81,7 @@ All **35** prod client schemas contain all four tables (`SalesOrders`, `Purchase
81
81
  - **Downstream-only assumptions are baked into shared code.** Any tenant whose POs are purely upstream (verified: NYCHH) silently gets `null`/empty from the shared downstream-reading helpers rather than an error.
82
82
  - **⚠ Compass Canada is NOT upstream-only, and it DOES have model overrides** — corrected 2026-08-26. The earlier "no `Model/Compasscanada/` directory" reading looked in the wrong place: the folder is **`_underscore/Model/Compass/Canada/`**, and `_Model_Compass_Canada_SalesOrder extends _Model_Compass_SalesOrder`, so Compass Canada inherits **Compass's** `_purchaseOrders` override, not the base one. Its **upstream bridge is completely empty (0 rows)** while 735 of 788 orders resolve a PO downstream. When checking whether a tenant has overrides, search for the class name (`grep -rn "class _Model_.*_SalesOrder extends"`) rather than guessing a directory — Compass's tenants nest one level deeper than everyone else's.
83
83
  - **⚠ Do NOT paper over the direction with an `ojoin` on a shared single-record fetch.** Tempting because it works in one place: the TOGa Supply NYCHH orders **list** joins the upstream bridge and is correct — but only because NYCHH's upstream bridge happens to be **1:1** (823 rows / 823 distinct sales orders, max 1 PO per SO). The same join in the shared `fetchOrdersDetails` (`toga2-supply/src/pages/Orders/api/OrdersApi.ts:397`) breaks three tenants: **Compass** has up to **6,440** downstream POs on one sales order (row multiplication on a single-record fetch), **Compass Canada's upstream bridge is empty** (735 of 788 would drop to **zero**), and **Prudential** has only **1,868 of 26,843** orders upstream (92% populated would become 7%). A scalar `GROUP_CONCAT` calculated field is the right mechanism precisely because it collapses many POs without multiplying rows.
84
+ - **⚠ When you DELETE a sales order, you must move BOTH directions of BOTH bridges — fixing one and missing its sibling is the default failure (2026-09-18).** Elite's stale-SalesOrder retirement shipped handling only `SalesOrders_PurchaseOrders` (downstream header) and `PurchaseOrderItems_SalesOrderItems` (upstream line), and failed on prod against the two siblings it skipped — `PurchaseOrders_SalesOrders` (upstream header) and `SalesOrderItems_PurchaseOrderItems` (downstream line). **Neither direction is "the main one":** on the same 218 orders, upstream dominated at header level (**218 vs 123**) while downstream was the rare one at line level (**7 vs 669**), so you cannot skip a direction because it looked empty on another tenant. The bridge FKs are `RESTRICT`, so a missed sibling is a hard FK error rather than silent — but you only find it one blocker at a time. **Enumerate instead:** query `information_schema.KEY_COLUMN_USAGE` for every FK referencing the parent, then count rows on the affected records, before writing any delete code. Details on the [per-client sync engine doc](../../../1.0/apps/worker/features/netsuite-togasupply-per-client-sync.md).
84
85
 
85
86
  ## Related
86
87
  - [Sales-order PO Number sourcing](./sales-order-po-number-sourcing.md)
@@ -6,11 +6,12 @@ project: TOGa View Frontend
6
6
  client: rate
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-09-16
10
- owners: ["bala"]
9
+ updated: 2026-09-18
10
+ owners: ["bala", "mhammontree"]
11
11
  files:
12
12
  - toga2-view/src/components/MobileNav/MobileNav.tsx
13
13
  - toga2-view/public/assets/Rate_Logo.svg
14
+ - toga2-view/src/context/AuthContext.tsx
14
15
  related:
15
16
  - clients/rate/profile.md
16
17
  - 2.0/apps/toga2-view/features/service-card.md
@@ -47,6 +48,8 @@ Files:
47
48
 
48
49
  - **Use `Rate_Logo.svg`, not `rateDark1.png` in the header.** The PNG has a 1.625:1 aspect ratio; at `h-[24px]` it renders only 39 px wide. The Figma Logos component is 58.5 × 24 px — only the SVG export matches.
49
50
  - **`EnvironmentBadge` is intentionally kept** in `App.tsx` — a dev/QA tool, must not be removed.
51
+ - **⚠ OPEN DEFECT — Logout does nothing, and never has (found 2026-09-18, TRUE-82118, NOT fixed).** `MobileNav.tsx:29` sets the Logout nav item's action to `() => console.log("Logging out...")` — a stub. The button (~line 162) calls `item.action?.()` then closes the drawer, so it logs one line and leaves the user signed in. **Two working implementations exist and NEITHER is called anywhere:** `AuthContext.tsx:84` `logout()` (clears accessToken/refreshToken/user, resets state) and `src/utils/auth.ts` `logout()` (clears tokens, redirects to `/login`) — that whole file is imported by nothing.
52
+ - **Testers: clearing local tokens ends OUR session, not the SAML IdP's.** The next load can silently sign the same user straight back in. To come back as a different test user, use a private window or a separate browser profile.
50
53
  - **Page content top offset:** pages behind the fixed header need `mt-[72px]` to avoid overlap. The `allowedPaths` list in `MobileNavToggle` controls which routes get the nav shell.
51
54
 
52
55
  ## Related
@@ -6,7 +6,7 @@ project: Worker
6
6
  client: elite
7
7
  type: client-feature
8
8
  status: active
9
- updated: 2026-09-17
9
+ updated: 2026-09-18
10
10
  owners: ["snaredla", "jcardinal", "rgirish"]
11
11
  files:
12
12
  - worker/crons/toga2/netsuite/sync_togasupply_elite.php
@@ -339,22 +339,70 @@ ships to **all 22 clients** carrying `netsuite` in `_modules.txt` and is written
339
339
  [ACL permission chain](../../../2.0/apps/_underscore/features/acl-permission-chain.md). **Run so far
340
340
  on `Client_Elite` only.**
341
341
 
342
- ### Elite's stale SalesOrder backfill is queued, not done
342
+ ### Elite's stale SalesOrder backfill — RUNNING, 218 → 187 (2026-09-18)
343
343
 
344
344
  Enabling transfer-order detection left **218** Elite `SalesOrders` that duplicate a `TransferOrder`
345
345
  (**217 of 217** matched pairs identical on order number, line count and total quantity). They still
346
346
  carry fulfillments, which is what drove `_qtyOnHand` negative — see
347
- [inventory quantities](./inventory-quantities-drop-ship.md). The retirement method that cleans them
348
- up is built but **NOT yet deployed**; mechanics on the
347
+ [inventory quantities](./inventory-quantities-drop-ship.md).
348
+
349
+ `retireStaleSalesOrderForTransferOrder()` is **deployed and working**. Verified on prod
350
+ 2026-09-18: stale duplicates **218 → 187**, `SalesOrders` **340 → 309**, fulfillments moved onto
351
+ the TransferOrder side **1 → 37** (moved, never deleted), `PurchaseOrders_TransferOrders`
352
+ **17 → 104**, `PurchaseOrderItems_TransferOrderItems` **28 → 163**.
353
+
354
+ Getting there took **four more FK dependencies** the first version missed — including two that
355
+ were the *opposite direction* of a bridge already handled. Full table, the row counts, and the
356
+ `information_schema` method that found them all at once are on the
357
+ [engine doc](../../../1.0/apps/worker/features/netsuite-togasupply-per-client-sync.md).
358
+
359
+ **Blockers:**
360
+
361
+ 1. ~~SalesOrder **240734** has 4 `InvoiceItems` and will throw~~ — **CLEARED 2026-09-18**, see
362
+ below.
363
+ 2. **11 of the 218** have no matching TransferOrder (4 are TOGa-created `SA1000xx` with no NetSuite
364
+ id). **Still open — manual review.**
365
+
366
+ #### Decision: the blocking invoice on SalesOrder 240734 was DELETED, not kept (2026-09-18)
367
+
368
+ Invoice **43** / number **277274** / NetSuite invoice id **6215693**, 4 `InvoiceItems`, 1
369
+ `InvoiceTrackingNumbers` row, created 2026-09-15.
370
+
371
+ **Why not keep it** — the answer is a documented team decision in the code at
372
+ `library/app/api/toga2.php:3350-3362`:
373
+
374
+ > *"Per team decision we do NOT import invoices billed against a transfer order at all."*
375
+
376
+ `Invoices` has **no `transferOrderId` column**, so such an invoice would import with a null order
377
+ link; `syncInvoiceFromNetsuite()` returns early for them. This invoice was a **leftover from before
378
+ Elite's transfer-order detection was enabled**, when 240734 was still a plain SalesOrder. Under
379
+ today's rule the sync would simply skip it. Deleting it puts the data where the current rule says
380
+ it belongs, and **NetSuite keeps the source record** — nothing is lost.
381
+
382
+ **Delete order matters — children first:** `InvoiceTrackingNumbers` → `InvoiceItems` → `Invoices`.
383
+ Executed and verified in prod: **0 / 0 / 0**. Order 240734 now has 0 invoice items and is clear to
384
+ retire.
385
+
386
+ ### ⚠ OPEN — the item-fulfillment FK fix is committed but NOT running in prod (2026-09-18)
387
+
388
+ `NETSUITE_EXECUTION_MODE_ITEM_FULFILLMENTS` is still failing on item `b96515b3`. The bug turned out
389
+ to have **three** layers (tracking rows → `ItemFulfillmentItemUnits` → the item), and the fix for
390
+ the deepest one is committed (`2247a70b`, working tree clean) — but the running process is
391
+ executing the **old** code. Proven from `Logs_Elite.Api`: the committed code does a
392
+ `GET /v2/item-fulfillment-item-tracking-numbers` immediately before the delete and that GET never
393
+ appears; at 10:27 and 11:07 the sequence was still two plain `GET /v2/item-fulfillment-items` then
394
+ the failing `DELETE`. **Deploy / opcache / checkout-path issue, unresolved.** Mechanics and the
395
+ general diagnostic are on the
349
396
  [engine doc](../../../1.0/apps/worker/features/netsuite-togasupply-per-client-sync.md).
350
397
 
351
- Two blockers to clear by hand first:
398
+ ### ⚠ OPEN — Elite's PO cursor is written ONE HOUR IN THE FUTURE (2026-09-18)
352
399
 
353
- 1. **SalesOrder 240734 has 4 `InvoiceItems` and WILL throw** — `InvoiceItems` has no
354
- `transferOrderItemId`, so those rows cannot be moved. Deliberate fail-loud; it blocks the
355
- backfill until resolved.
356
- 2. **11 of the 218 have no matching TransferOrder** (4 are TOGa-created `SA1000xx` with no NetSuite
357
- id). Manual review.
400
+ Elite's POs sat at a healthy-looking `864000-IDLE` with a current cursor and synced **nothing**.
401
+ Every empty-window write was server time **+60 minutes**, always ending `|0` (10:36:09 → cursor
402
+ 11:36:08). This is an engine-wide bug in the empty-window (`|0`) branch, not an Elite
403
+ configuration problem, and it only appears **after a section catches up** — which is exactly what
404
+ Elite's PO backfill just did. **NOT FIXED** (shared file, 22 clients). Full root cause on the
405
+ [engine doc](../../../1.0/apps/worker/features/netsuite-togasupply-per-client-sync.md).
358
406
 
359
407
  ## Gotchas / known issues
360
408
 
@@ -397,6 +445,18 @@ Two blockers to clear by hand first:
397
445
  rewinding `NETSUITE_LAST_SYNC_DATETIME_SALES_ORDERS`.
398
446
 
399
447
  ## Change history
448
+ - 2026-09-18 — **Stale-SalesOrder backfill is live: 218 → 187.** The retirement method needed
449
+ four more FK dependencies before it would run (fulfillment **header** `salesOrderId`, the
450
+ upstream `PurchaseOrders_SalesOrders`, the downstream `SalesOrderItems_PurchaseOrderItems`, and
451
+ `Invoices.salesOrderId`); fulfillments were **moved** to the TransferOrder side, never deleted
452
+ (1 → 37). **Cleared blocker 1 by deleting invoice 43 / 277274 / NetSuite 6215693** — justified
453
+ by the team decision at `toga2.php:3350-3362` that invoices billed against a transfer order are
454
+ not imported at all (`Invoices` has no `transferOrderId`), so it was a pre-detection leftover;
455
+ deleted children-first and verified 0/0/0. 11 unmatched orders remain. **Two things left OPEN:**
456
+ the three-layer item-fulfillment FK fix is committed (`2247a70b`) but the **running process is
457
+ still on the old code** (proven from the `Logs_Elite.Api` call sequence), and Elite's PO cursor
458
+ is being written **one hour in the future** by the engine's empty-window branch — an IDLE-looking
459
+ section that syncs nothing. (rgirish)
400
460
  - 2026-09-17 — Corrected the stale "transfer orders pre-set but not in use / 0 rows" note in
401
461
  *Elite-only wrapper deviations*: `IS_ENABLED_INTEGRATION_TRANSFER_ORDERS` is now `true`
402
462
  (ZERO_DOLLAR_HOLD) and prod `Client_Elite.TransferOrders` holds 218 rows. (jcardinal)
@@ -8,8 +8,8 @@ project: Worker
8
8
  client: endeavor-health
9
9
  type: profile
10
10
  status: active
11
- updated: 2026-09-16
12
- owners: ["jcardinal"]
11
+ updated: 2026-09-18
12
+ owners: ["jcardinal", "rgirish"]
13
13
  files: []
14
14
  related:
15
15
  - ../../1.0/apps/worker/features/netsuite-togasupply-per-client-sync.md
@@ -23,6 +23,19 @@ Endeavor Health TOGa-Supply NetSuite integration client on the shared 1.0 worker
23
23
  - NetSuite transactions sync into the 2.0 `Client_<Id>` tenant via the TOGa2 API through a thin `worker/crons/toga2/netsuite/sync_togasupply_*.php` wrapper.
24
24
  - Profile seeded 2026-08-29 fixing Endeavor's stuck `ITEM_RECEIPTS` section; item-receipt path migrated SOAP → REST (`fetchItemReceiptById`). Mechanics + diagnosis playbook: [per-client sync](../../1.0/apps/worker/features/netsuite-togasupply-per-client-sync.md).
25
25
 
26
+ ## ⚠ OPEN (2026-09-18) — INVOICES frozen, and it is probably the future-cursor bug
27
+
28
+ The 2026-09-18 fleet sweep found `NETSUITE_EXECUTION_MODE_INVOICES` at **`1-RUNNING`**. After a
29
+ manual reset to `864000-IDLE`, the **window grew back but the cursor stayed frozen** and no errors
30
+ appeared.
31
+
32
+ That is the exact signature of the engine's **empty-window cursor write storing a time one hour in
33
+ the future** — an IDLE-looking section that syncs nothing forever. Root cause, the `|0` tell, and
34
+ the fact that it is **not yet fixed** (shared file, all 22 clients) are on the
35
+ [per-client sync engine doc](../../1.0/apps/worker/features/netsuite-togasupply-per-client-sync.md).
36
+
37
+ **Still frozen as of 2026-09-18** — do not re-diagnose this as an Endeavor-specific problem.
38
+
26
39
  ## Integration touchpoints
27
40
 
28
41
  - **`worker` / `library` (1.0)** — shared togasupply sync engine + `App_Api_Toga2` bridge.
@@ -90,6 +90,30 @@ Current selection query:
90
90
  - The standalone assignment-group block was **deleted** — the main payload already diffs `assignment_group`, so it was a duplicate GET, a second code path, and the source of the lockout stamp.
91
91
  - Debug pin already in the file: the commented `#AND NYCDOETickets.ticketNumber = '...'` line scopes a pass to one ticket.
92
92
 
93
+ ### Batch loop, run caps & per-ticket isolation (TRUE-82129 current shape)
94
+ Before this the cron ran the selection query **once per ticket under `LIMIT 1`** with a `sleep(1)` after every ticket — one ticket per 6-minute run, ~10 days for a full pass over the queue.
95
+ - **One `LIMIT 250` (`BATCH_SIZE`) read into a PHP array, consumed with `array_shift`, refilled only when empty.** The query is `type=ALL`, ~35,021 rows, `Using temporary; Using filesort` (it drives off `managed_service_orders`, so `Common.NYCDOETickets.idx_dtLastChecked` is unusable at this join order). A filesort already sorts every surviving row before returning the top N, **so a bigger LIMIT costs nothing extra**.
96
+ - **`BATCH_SIZE` throttles NOTHING** — it caps rows per SELECT only, because the loop refills. Lowering it is **not** a rate-limit control; a reviewer proposed exactly that and it would not have worked. The rate cap is **`MAX_TICKETS_PER_RUN` (300)**.
97
+ - **The per-ticket `sleep(1)` was removed** — it burned ~70% of the 340s run window against a 0.35s API call. **The `sleep(1)` after the Assigned→In Progress PATCH was KEPT** — it waits for ServiceNow to apply the change before the re-read. Do not delete that one.
98
+ - **Per-ticket `try/catch (Throwable)`** so one bad ticket cannot strand the rest, plus **`MAX_CONSECUTIVE_FAILURES` (10)**: an unbroken failure run stops the loop and **throws after the alert emails**, so a systemic break still ends as a visibly failed run.
99
+ - Measured in production 2026-09-18: **300 tickets in 171s (0.57s/ticket), zero non-200**; eligible queue 2,406 → 1,845 over two runs.
100
+
101
+ > **DECISION (Mark, 2026-09-18) — narrow, reasoned exception to `1.0/standards/backend-php.md` § *Fail-loud is the default for import/sync/cron code*.** That rule exists because a watermark can advance PAST a lost record. Here `dtSynced` advances **only after a successful PATCH**, so a thrown ticket advances nothing, loses nothing, and is retried next run; `MAX_CONSECUTIVE_FAILURES` preserves the fail-loud guarantee. **Do not re-propose removing the per-ticket catch without also removing this reasoning.**
102
+
103
+ > **FOLLOW-UP (proposed, NOT done — own ticket):** rewrite the selection to drive off `Common.NYCDOETickets` (subquery on that table alone, then join and re-filter) so `idx_dtLastChecked` is usable. Needs an over-fetch factor sized from real data. `sql-reviewer` recommended keeping it out of TRUE-82129.
104
+
105
+ ### Alert-email safety (TRUE-82129, from cso review)
106
+ - **Never put `$e->getMessage()` into an email or a log on this cron.** Exceptions from `App_Api_NYCDOEV2` / `App_ApiTransaction` embed the **raw ServiceNow response body** — DOE requester PII, and on an auth failure the body of a request carrying the OAuth secret.
107
+ - The catch calls **`App_Error_Capture::captureException($e)` explicitly** (catching normally bypasses that capture) and the email carries **only the issue reference**. See [Error capture in 1.0](../../../1.0/apps/library/features/error-capture-1-0.md).
108
+ - `htmlspecialchars` on every interpolated value in both alert emails; failure list capped at `MAX_FAILURES_LISTED` (25) with the exact count in the subject.
109
+ - The unmapped-assignment-group email now groups **by assignment group** with an exact count and up to `MAX_TICKETS_LISTED_PER_GROUP` (10) ticket numbers each.
110
+
111
+ ### No output from this cron (team rule)
112
+ All three `error_log()` calls were removed (the per-ticket technician sys_id message, TRUE-82027's "alerted on N failures" line, and one added in TRUE-82129). **1.0 worker crons must produce no output** — they run unattended under CRON and `error_log()` goes to stderr on CLI. Failures surface only in `Logs.Issue` and the throttled alert email. General 1.0 worker rule, not DOE-specific (Mark, 2026-09-18).
113
+
114
+ ### Missing technician `referenceId` is the NORMAL state, not an incident
115
+ **265 of 389 (68%) TOGaDesk technicians have no `people.referenceId`** (ServiceNow sys_id), so `u_technician` is left **unchanged** on those pushes — deliberate: blanking it would wipe a technician ServiceNow may legitimately hold. State, hold, ETA and assignment_group still sync. This is why the per-ticket log line was removed. Reconcile with `worker/crons/sync/nycdoe/report_technician_identities.php` / `backfill_technician_identities.php` (needs its own ticket).
116
+
93
117
  ### Failure alerting instead of a throw (TRUE-82027)
94
118
  **DECISION (Mark, 2026-09-17): no circuit breaker** — 1.0 is transitioning to 2.0 under a separate task, so the cheap path was chosen.
95
119
  - A failed ServiceNow PATCH **no longer throws**; it records the incident and the run continues, then all failures are emailed **once per run** to sking@togatech.com / mhammontree@togatech.com.
@@ -155,12 +179,20 @@ The SNOW round-trip note reconciliation (delete-and-reinsert in `process_tickets
155
179
  - **A merged PR is not a deployed change — TRUE-81044 is merged to `_production` but NOT deployed** (merge commit worker `2f4dcfe9e`). Production payloads on 2026-09-15 still write `u_status_task` and still never send `u_other_reason_on_hold` (**0 occurrences across 9,914 PATCHes since 2026-08-20**). Read the live payload in `Logs.API` before concluding a fix is live.
156
180
  - Holds are **never auto-released** — if a SNOW user manually changes status, the outbound cron correctly re-asserts On Hold; that is intended, not a bug. Clear the hold in TOGaDesk.
157
181
  - **Diagnosing which writer dropped a hold:** `repair_order_history` (`repairid, userid, dtStamp, note`) records *user-driven* status changes. A hold reverting with **no** history row between the hold and the revert = a raw-`UPDATE` writer (`updateStatus`/`qqStatus`), **not** the worker inbound cron (`process_tickets.php` contains no reference to `ORDER_ASSIGNED_AWAITING_SCHEDULING`). Note that **since TRUE-80060** a *status-changing* `class.repair.php::addNotes` now **calls** `updateStatus()` (so it too can move the cache via that raw `UPDATE`); a **pure comment** still writes nothing to `repair_orders.status`. This history-gap technique is how TRUE-80060 was pinned to `qqStatus()` rather than the SNOW round-trip.
158
- - **KNOWN REMAINING GAP (needs a follow-up ticket) — the invalid-assignment-group email reports only the LAST ticket of a run.** `$invalidAssignmentGroups` is declared **inside** the outer `do/while`, so it resets every ticket; the email body now says so explicitly. Agreed fix: declare it before the loop (as `$failedPatches` already is) **and** group output by assignment group with a capped ticket list — one unmapped group can cover 60+ tickets. Rare in practice; the default group is TOGA's own, historically named "ASI Systems".
182
+ - **⚠ OUTAGE (TRUE-82129, 2026-09-18) — `send_ticket_updates.php` died on its FIRST ticket of EVERY run, for a whole day.** `App_Database::fetchOne()` (like `fetchRow()` / `numRows()`) takes its result **by reference** (`&$res`, `library/app/database.php` L336/L352/L362). TRUE-82027 (commit `b63fb0136`) passed `App_Database::query(...)` into it **inline**, which raises the notice *"Only variables should be passed by reference"* — and 1.0's error handler **terminates the run on a notice**. Live from the 2026-09-18 10:18 deploy until fixed. **Fix:** assign the `query()` result to a variable first. Symptom set worth recognizing: zero ServiceNow calls after the deploy, `Common.CronJobExecutions.dtCheckOut` NULL on every run, and exactly **one** `NYCDOETickets.dtLastChecked` stamp per 6-minute run. General 1.0 trap — see the 1.0 back-end standard.
183
+ - **⚠ "The outbound cron re-asserts On Hold every run" (business rule 2) was only true in theory between 2026-09-18 10:18 and the TRUE-82129 fix** — the cron was dying on its first ticket. Check the cron actually completes before trusting that guarantee.
184
+ - **OPEN QUESTION — `dtSynced` is not a reliable outbound-only signal.** The 12 reported incidents (INC2298009, INC2298110, INC2297667, INC2299269-77) had `Common.NYCDOETickets.dtSynced` advance from ~2026-09-17/18 05:00–09:46 to ~08:57–09:00 on 2026-09-18 **while the outbound cron was not running**. Suspected `process_tickets.php` (inbound also writes `dtSynced`) but **not confirmed**. Confirm before reading `dtSynced` as proof of an outbound push.
185
+ - **⚠ SECURITY — hardcoded live ServiceNow credentials in `library/app/api/nycdoev2.php` / `nycdoe.php`.** Found 2026-09-18, not fixed; detail in [ServiceNow / ASN Integration](servicenow-integration.md).
159
186
  - **KNOWN REMAINING GAP (needs a follow-up ticket):** both `qqStatus()` hold gates live inside the `type IN (ONSITE_REPAIR, ONSITE_SERVICE)` branch. The `DEPOT_REPAIR` branch (`repairorder.php` from ~L799) has **no hold gate at all**, so a `DEPOT_REPAIR` order set to any hold status is still silently recomputed out of hold by `updateStatus()` — same class of bug, untouched by TRUE-80060.
160
187
  - **KNOWN REMAINING GAP (needs a follow-up ticket) — inbound note reconciliation is paged too small.** `process_tickets.php`'s inbound comment reconciliation fetches only `page_size=50, page_number=1` of the SNOW journal, then **deletes any local note not in that set** (the remainder-delete ~L1326–1330). An order with **>50** comments/work_notes can have local notes silently deleted. It is also inconsistent with `addNotes`, which fetches `page_size=100`.
161
188
  - **KNOWN REMAINING GAP — `buildDoubleFieldArray` collapses unsynced notes.** In `process_tickets`, `buildDoubleFieldArray` keys on `IFNULL(referenceId,0)`, so **all** NULL-`referenceId` notes collapse to key `'0'`, making the remainder-delete handle multiple unsynced notes inconsistently.
162
189
  - `qqStatus()` collapses all recognized hold sub-statuses to base `HOLD` in `repair_orders.status`; the specific hold **sub-reason persists only in `repair_order_notes.newstatus`**. Any status badge / report / filter reading `repair_orders.status` directly cannot distinguish the six hold variants after a recompute.
163
190
 
191
+ ## Change history
192
+ - 2026-09-18 — TRUE-82129: fixed the by-reference `fetchOne()` notice that killed every run; batch loop + run caps replaced LIMIT 1 + per-ticket sleep; alert emails de-PII'd and escaped; unmapped-group email now grouped per run. (mhammontree)
193
+ - 2026-09-17 — TRUE-82027: outbound push runs once per order per pass; `dtLastChecked` split from `dtSynced`; failure alerting instead of a throw. (mhammontree)
194
+ - 2026-08-20 — TRUE-81044: `state` + `hold_reason` + `u_eta` + `u_other_reason_on_hold` pushed together; stopped writing `u_status_task`. (mhammontree)
195
+
164
196
  ## Related
165
197
  - [NYCDOE ServiceNow / ASN Integration](servicenow-integration.md) — full integration map, crons, schedules, and `App_Api_NYCDOEV2` plumbing.
166
198
  - [NYC DOE client profile](../profile.md)
@@ -156,6 +156,8 @@ This probe is the client's **only 2.0 footprint**. DOE is **1.0 by default** —
156
156
 
157
157
  ## Gotchas
158
158
 
159
+ - **⚠ SECURITY (cso review 2026-09-18 — NOT fixed, needs its own ticket): live production ServiceNow credentials are hardcoded in committed PHP.** `library/app/api/nycdoev2.php` and `library/app/api/nycdoe.php` hold the ServiceNow **`clientSecret` and `refreshToken` as literals** (stage values too). Treat as compromised: move them to config/env and **rotate with DOE**. Separately, `Vision_Log.ErrorLog` rows dump `$_SERVER` including live AWS access key id / secret in plain text — its own ticket, same rotation treatment. **No credential values are recorded here or anywhere in the KB.**
160
+
159
161
  - **⚠ OPEN BUG / NAMED PATTERN — "Lenovo Off-Layout ASN File" (silent column shift + unit collapse).** Investigated 2026-08-11 (read-only; **no code fix exists**). Lenovo sends NYCDOE ASNs in **two different column layouts**, and `library/app/asnprocessor/lenovo.php` picks its column mapping by **column COUNT with no header-name validation** (`:110`, `if ($colCount >= 30`), so the off-layout file is parsed with the wrong mapping and silently mis-read.
160
162
  - **The two layouts.** The normal daily feed `DOE-Lenovo-ASN-MM-DD-YYYY.CSV` (~2 MB) has **30** columns and matches the code's mapping. A second file, `DOE-Open-Lenovo-MM-DD-YYYY.csv` (**lowercase** `.csv`, ~45 KB), has **31** columns in a **different order**: `Contact Email` moves from column 29 to column **6**, and a trailing `Full Shipment` column is appended. Everything from `Street Address` onward is therefore read **one column too early**. The "extra comma" repair heuristic at `:119-127` does not catch this (unchanged since March 2026 — commits `3442ddcf`, `6956dfbf`, `864cf073`, so not a regression).
161
163
  - **The four wrong values** (`lenovo.php:129-156`): `serialNumber` ← `$cols[22]` lands on **System Quantity**, a constant `"1"` on every row (the real serial ends up in `assetTag`); `orderQuantity` ← `$cols[21]` lands on a blank `Quantity Ordered`; `partNumber` ← `$cols[17]` lands on the long **Product Description** instead of the OEM SKU.
@@ -6,7 +6,7 @@ project: Library
6
6
  client: nychh
7
7
  type: client-feature
8
8
  status: active
9
- updated: 2026-09-16
9
+ updated: 2026-09-18
10
10
  owners: [jcardinal, bala, rgirish]
11
11
  files:
12
12
  - library/app/api/toga2.php
@@ -182,6 +182,22 @@ Two NYCHH-specific gates, both of which bit during this work:
182
182
 
183
183
  WARNING: NYCHH's `NETSUITE_LAST_SYNC_CURSOR_ITEM_RECEIPTS` is stale at **2024-08-06** - re-enabling item receipts replays about two years unless it is rewound forward first.
184
184
 
185
+ ## ⚠ OPEN (2026-09-18) — NYCHH sync is 2.5 YEARS behind on a 403 EZ-1 from records 326/330
186
+
187
+ Found by a fleet sweep of `NETSUITE_EXECUTION_MODE_*` across all 22 netsuite clients. NYCHH was the worst: `NETSUITE_LAST_SYNC_CURSOR_PURCHASE_ORDERS` at **2024-02-16** and `NETSUITE_LAST_SYNC_CURSOR_INVOICES` at **2024-02-22**.
188
+
189
+ **⚠ The mode read `864000-RUNNING` the whole time and was taken as "healthy, in progress".** That reading was wrong — a full window with a 2.5-year-old cursor is **broken, not busy**. Always pair the execution mode with the cursor; neither is a health signal alone.
190
+
191
+ **Error:** `GET /v2/transfer-order-items-purchase-order-items` → **403 `EZ-1`**.
192
+
193
+ **Root cause — a different ACL gap from Elite's.** Records **326** (`transfer-orders-purchase-orders`) and **330** (`transfer-order-items-purchase-order-items`) — the **TO-first** direction — have **zero** `AclRecordPermissions` rows. The `2026-09-17a` module migration only covered the PO-first siblings **325/329**. The error shape differs from Elite's silent `200 + WZ-1`: `EZ-1` means the **record** layer is missing, `WZ-1` means the record layer is fine and the **field** layer is empty — use the code to pick which layer to fix.
194
+
195
+ **Red herring:** the envelope reports `api: "Agilant"`, which looks like a different API needing different grants. It is not — NYCHH's Agilant API carries `roleId` **1 and 3**, so roleId 3 is still correct.
196
+
197
+ **Fix prepared, ⚠ NOT YET RUN:** client-scoped SQL, 5 statements, same four-layer chain and `NOT EXISTS` guards as `2026-09-17a` — permissions for 326/330, `'all'` expressions, logic groups, logic-group expressions, and 6 field grants (**2214/2215/2216** for 326; **2230/2231/2232** for 330; `id` **2213/2229** deliberately excluded, matching 325/329). Chain mechanics on the [ACL permission chain](../../../2.0/apps/_underscore/features/acl-permission-chain.md).
198
+
199
+ **Not NYCHH-only.** 326/330 were never in the module migration, so any client whose sync touches the TO-first bridges will hit the same 403.
200
+
185
201
  ## Gotchas / known issues
186
202
 
187
203
  - **🚩 The import sends almost everything to catch-all destination location 2 "NYC Health + Hospitals" — 1,832 of 1,939 transfer orders**, next highest Coney Island at 27, and **every one has `createdByUserId` NULL** (277 on 2026-08-29, 1,528 on 2026-08-30, then 1–3/day). Whether that location is a real receiving place or the import should resolve the actual hospital is **open with the PM**; its address is currently Jacobi's. Do not assume the destination is correct: [location shipping addresses](./location-shipping-addresses.md).
@@ -7,7 +7,7 @@
7
7
  | [Rate home-warranty Terms & Conditions content — verbatim transcript policy](features/home-warranty-terms-content.md) | 2.0 | Why `TERMS_HW.ts` is a byte-verbatim transcript of Rate's legal form and must never be hand-edited; open before touching the home-warranty T&C content or regene |
8
8
  | [Rate Monthly Reconciliation Report](features/monthly-reconciliation-report.md) | 1.0 | Monthly cron that emails an Excel reconciliation of Rate subscription orders + PayPal payments; open when changing the report columns, filters, or recipients. |
9
9
  | [Rate SalesOrder → NetSuite CashSale Export (postPost)](features/netsuite-cashsale-export.md) | 2.0 | How a Rate SalesOrder postPost writes a NetSuite CashSale (ship-to address + line-item internalId resolution); open when touching Rate's NetSuite export. |
10
- | [Rate PayPal Subscription Purchase & Webhook Pipeline](features/paypal-subscription-purchase-webhook.md) | 2.0 | How Rate PayPal subscription purchases are created (browser-only) and how the worker2 webhook handles lifecycle events; open before touching Rate purchase/webho |
10
+ | [Rate PayPal Subscription Purchase & Webhook Pipeline](features/paypal-subscription-purchase-webhook.md) | 2.0 | How Rate PayPal subscription purchases are created (browser primary, webhook fallback) and how the worker2 webhook handles lifecycle events; open before touchin |
11
11
  | [Rate SAML SSO](features/saml-sso.md) | 2.0 | How Rate users sign in via Azure AD SAML through the TOGa gateway, plus assertion attributes and the loan-officer upsert; open when debugging Rate login or user |
12
12
  | [Service Card Entitlement Display](features/service-card-entitlements.md) | 2.0 | How Rate's home/services pages render one service card per purchased entitlement (per-entitlement, own pinned address for warranty); open when touching service |
13
13
  | [Rate Service-Purchase Confirmation Emails (Tech / Warranty)](features/service-purchase-emails.md) | 2.0 | How Rate purchase-confirmation emails (tech vs warranty) are triggered from the entitlement postPost and sent off-thread from DB-stored templates; open when edi |
@@ -6,7 +6,7 @@ project: _Underscore
6
6
  client: rate
7
7
  type: client-feature
8
8
  status: active
9
- updated: 2026-09-16
9
+ updated: 2026-09-18
10
10
  owners: [mhammontree, tcox, rgirish]
11
11
  files:
12
12
  - _underscore/Model/Rate/Entitlement.php
@@ -150,6 +150,11 @@ Because the interceptor swallows failures, detection rides the **committed outbo
150
150
  - **Ordering: TRUE-81049 must land before/with TRUE-80575.** TRUE-80575's worker2 handler re-reads the entitlement and alerts when `c_aigContractId` is empty — until identifiers persist, that check false-alarms on **every** purchase. TRUE-80575 (1 `_underscore` commit + 3 `worker2` commits) sat un-PR'd from 2026-08-11, never deployed; PRs raised 2026-08-18.
151
151
  - **No persistent failure flag** — the only durable evidence of a failed contract creation is the outbound API log row, not the entitlement.
152
152
  - **Address-gated** — `postPost` does nothing when the payload lacks a primary contact address, so such an entitlement never attempts contract creation (and never logs a `/contract` call).
153
+ - **⚠ OPEN DEFECT (found 2026-09-18, TRUE-82118, NOT fixed) — a TECH purchase attempts an AIG WARRANTY contract.** `_underscore/Model/Rate/Entitlement.php` ~line 491 gates the whole contract-creation block on *"does the payload have an address"* (`!empty($payload->contact->primaryContactAddress->address)`), **not** on whether the product is a warranty. `$isWarranty` is only used for reporting. So any tech purchase by a contact that happens to have an address tries to create an AIG contract. Fix direction: gate that block on `$isWarranty` **and** the address. Deferred because `_underscore` deploys were blocked until 2026-09-24 (Rohan) and TRUE-82118 had to ship 2026-09-18.
154
+ - **⚠ The address comes from the RESPONSE payload, not the request.** api2 builds it from the contact's saved address, so **this cannot be prevented from worker2 or toga2-view — only in `_underscore`.** Proof: dev-sandbox `Logs_Rate.Api` id 62300 — `requestPayload` has no `primaryContactAddress`, `responsePayload` does.
155
+ - Observed twice on dev-sandbox (`Logs_Rate.Api` 62254, 62299): OUT `POST /wishservices/api/authentication/login`, `responseCode` 0 — no egress, so it stopped at auth.
156
+ - **Exposure: only contacts WITH an address** — 11 of 666 production Rate contacts, all existing Whole Home Warranty holders.
157
+ - Most likely outcome is AIG **rejecting** it (we send `saleItemId` as the item **title**, e.g. `"1 Year Unlimited Tech Support - Annual"`, which is not in AIG's catalog), swallowed silently — but it **would** fire the `_Worker_Monitors_RateEntitlement` alert to devteam@togatech.com and look like a broken warranty purchase. Worst case, unproven: AIG accepts, a real warranty contract is created for a tech sale, and `persistAigContractIdentifiers()` writes the identifiers onto the **tech** entitlement — it runs on the success path regardless of product.
153
158
  - **Auth-induced failures are blind to the monitor** — see the detection limitation.
154
159
  - **The AIG carrier cancel (`/contract/cancel`) is chained off the SAME `BILLING.SUBSCRIPTION.CANCELLED` event** worker2 uses to confirm a portal cancellation — so the two flows share a trigger and a failure mode: while `Rate/Webhook` jobs time out, neither the AIG cancel nor the `dateCancelled` stamp happens. See [subscription cancellation](subscription-cancellation.md). **"AIG" here is the *carrier*, not the `Client_Aig` tenant** — do not resolve it as a client slug.
155
160
  - **The AIG-success branch REASSIGNS `$payload->uuid`** (`$payload->uuid = _String::generateUuid()`, repurposed as the AIG contract number) in the same `postPost` that pins the WH service address. On the pre-redesign build the pin step ran **after** this reassignment and keyed off `$payload->uuid`, so **every successful AIG contract silently broke the WH address pin** (the pin lookup matched no row). Preserve the ordering fix — snapshot the entitlement uuid **before** this branch. See the [WH per-address purchase guard](whole-home-warranty-purchase-guard.md) "AIG-success clobbers the pin" gotcha (`_underscore` `_beta` `9c2e4b1c`).
@@ -6,7 +6,7 @@ project: Worker
6
6
  client: rate
7
7
  type: client-feature
8
8
  status: active
9
- updated: 2026-09-16
9
+ updated: 2026-09-18
10
10
  owners: ["mhammontree"]
11
11
  files:
12
12
  - worker2/Worker/Rate.php
@@ -18,6 +18,8 @@ files:
18
18
  - toga2-view/src/pages/CheckOut/viewModel/useCheckoutPageViewModel.ts
19
19
  - toga2-view/src/pages/CheckOut/api/checkoutApi.ts
20
20
  - toga2-view/src/pages/Activation/view/Activation.tsx
21
+ - toga2-view/src/constants/storageKeys.ts
22
+ - test/@Mark/Rate/replay_paypal_webhook.php
21
23
  related:
22
24
  - profile.md
23
25
  - service-card-entitlements.md
@@ -26,30 +28,38 @@ related:
26
28
  - aig-contract-creation.md
27
29
  - subscription-cancellation.md
28
30
  ---
29
- How Rate PayPal subscription purchases are created (browser-only) and how the worker2 webhook handles lifecycle events; open before touching Rate purchase/webhook, backfilling a missed purchase, or the watchdog-timeout defect.
31
+ How Rate PayPal subscription purchases are created (browser primary, webhook fallback) and how the worker2 webhook handles lifecycle events; open before touching Rate purchase/webhook, replaying a webhook, or backfilling a missed purchase.
30
32
 
31
33
  ## Summary
32
- Rate customers buy home-tech-support and home-warranty subscriptions through PayPal on `toga2-view`. The purchase is created **entirely client-side**; `_Worker_Rate::Webhook` (worker2) handles only post-purchase lifecycle events (renewal sales orders and deactivations) against an entitlement that **already exists**.
34
+ Rate customers buy home-tech-support and home-warranty subscriptions through PayPal on `toga2-view`. The browser is the **primary** creation path; `_Worker_Rate::Webhook` (worker2) is the **server-side fallback** (`BILLING.SUBSCRIPTION.ACTIVATED`) and handles post-purchase lifecycle events (renewal sales orders, deactivations).
33
35
 
34
- **⚠ STATUS (2026-08-10, TRUE-80575) — A FIXED (uncommitted), B DECIDED (not built):**
35
- - **A. Every `Rate/Webhook` worker job fails** (watchdog kill at ~322s vs a 300s watchdog; prod `Core.WorkerJobs` 560657, 563520, 563544, 563545, 574010, 2026-08-01 → 2026-08-03). **Fix written, NOT committed** — it sits in the `worker2` working tree on branch `TRUE-80575`. See *Idempotency (rebuilt)*.
36
- - **B. New purchases have no server-side creation path at all.** Only the customer's browser can create an entitlement. If the browser never completes the call, PayPal keeps the money and TOGa records nothing. **Approach decided** — add a `BILLING.SUBSCRIPTION.ACTIVATED` handler to `_Worker_Rate::Webhook` as a *fallback net*; the browser `onApprove` POST stays the PRIMARY synchronous path. **Not yet built**, and for warranty it is **blocked** by the `toga2-view` work to persist the covered property pre-payment.
37
- - **A and B are separate.** Fixing the watchdog does **not** close the money-taken-nothing-recorded hole.
36
+ **STATUS (2026-09-18, TRUE-82118):** TRUE-80575's two defects are **fixed and deployed** — the watchdog timeout (A) and the missing server-side creation path (B, the `BILLING.SUBSCRIPTION.ACTIVATED` fallback handler). TRUE-82118 then fixed a defect **in** that fallback: it rejected every Tech Support purchase. See *ACTIVATED fallback*. Committed on branch `TRUE-82118` (`worker2` 28d6569, `toga2-view` 6e4b01c), **not yet pushed**.
38
37
 
39
38
  ## How it works
40
39
 
41
- ### New purchase (browser-only)
40
+ ### New purchase (browser — primary path)
42
41
  1. `usePayPalSubscription.ts` resolves a plan id from `PAYPAL_PLAN_IDS` (`paypalService.ts`) and opens the PayPal subscription flow.
43
42
  2. On PayPal approval, the SDK's `onApprove` callback (`usePayPalSubscription.ts:151-177`) fires **in the customer's browser**.
44
43
  3. `handleActivateService` (`useCheckoutPageViewModel.ts:402-489`) builds a single nested `POST /entitlements` payload — Entitlement + Subscription + SalesOrder + SalesOrderItem + Payment — and posts it via `checkoutApi.ts:20`.
45
- 4. That browser call is the **only** thing that creates the records.
44
+ 4. That browser call normally creates the records; the ACTIVATED webhook covers it when it does not.
45
+ 5. **Safety net (TRUE-82118, `6e4b01c`).** `onApprove` writes the PayPal subscription id, user uuid, plan and amount to `sessionStorage` (keys in `src/constants/storageKeys.ts`) **before** the API call; checkout clears it only once the entitlement exists. The buyer is charged the moment `onApprove` fires, so a failed POST / closed tab / dropped connection still leaves a trace support can follow. All storage access is `try/catch`-wrapped.
46
+ 6. **`return_url` / `cancel_url` removed** from `application_context` in `createSubscription`. **This is NOT a confirmed root cause** — PayPal's App Switch docs say client-side callbacks "will continue to work". They were removed because nothing in the app reads either URL, and pointing PayPal's redirect at `/activation?success=true` is unsafe on its own: that page reports success unconditionally (10s countdown → `/home`, never reads `subscription_id`), so any flow landing there without `onApprove` having run tells a paying customer it worked while no entitlement exists.
46
47
 
47
48
  ### Lifecycle (worker2)
48
- PayPal → Lambda → SQS → `_Worker_Rate::Webhook`, which handles `PAYMENT.SALE.COMPLETED` (renewal sales order on an existing entitlement) and six deactivation events. There is **no** `BILLING.SUBSCRIPTION.ACTIVATED` handler. For a new purchase the worker dead-ends early: `Rate.php:112` throws *"Subscription not found for token"*, `Rate.php:394` throws *"Entitlement not found"*.
49
+ PayPal → Lambda → SQS → `_Worker_Rate::Webhook`, which handles `BILLING.SUBSCRIPTION.ACTIVATED` (server-side fallback creation), `PAYMENT.SALE.COMPLETED` (renewal sales order on an existing entitlement) and six deactivation events.
49
50
 
50
- ### Idempotency (rebuilt 2026-08-10 — code in the `TRUE-80575` worker2 working tree, uncommitted)
51
+ ### ACTIVATED fallback — `handleSubscriptionActivated` (`worker2/Worker/Rate.php`)
52
+ The server-side net for a purchase the browser never posted. **It is a fallback, not the primary path** — the browser `onApprove` POST still creates the records in the normal case.
53
+
54
+ **`plan_requires_aig` is a REQUIRED per-plan config entry** (`worker2/Config/production.ini`). `resolveProductForPlan()` **refuses** a plan that has no entry rather than defaulting to `false`, and tests with `array_key_exists` so an explicit `"0"` stays valid. The resulting `$requiresAig` now decides whether the AIG-only guards run at all, so a new warranty plan added without that line would sell a warranty with no address, no phone, and no AIG contract check. Both Tech Support plans declare `"0"`.
55
+
56
+ **Guards are AIG-gated, not universal (TRUE-82118 fix, `28d6569`).** Step 5's *"contact must have an address"* check was written for the AIG-backed Whole Home Warranty but ran for **every** plan, so **every Tech Support purchase was blocked** — Tech Support is not AIG-backed and has no covered property anywhere in its flow. Step 5 is now gated on `$requiresAig` (matching how step 5b already gated the phone check), and `contact.primaryContactAddress` is included in the create payload **only** when the contact has one.
57
+
58
+ **Still UNVERIFIED end to end.** The 2026-09-18 sandbox proof (below) exercised the **frontend** path only. To exercise the fallback, block the `POST /v2/entitlements` request in devtools after PayPal approval so only the webhook runs.
59
+
60
+ ### Idempotency (rebuilt 2026-08-10, TRUE-80575 — shipped)
51
61
  **Guard on the business record, not on a log or a job row.** `worker2/Worker/Rate.php`:
52
- - `isEventProcessed()` **deleted** entirely, along with its call site and the `Logs_Rate` lookup.
62
+ - `isEventProcessed()` was **deleted** entirely, along with its call site and the `Logs_Rate` lookup.
53
63
  - New `findPaymentByTransactionId()` — `GET /payments` filtered on `transactionIdentifier`, mirroring the existing `findSubscriptionByToken()`.
54
64
  - `handlePaymentCompleted()` now opens with a guard that returns **success without writing** if that PayPal sale is already recorded.
55
65
  - **`DB_LOGS_RATE` registration is deliberately KEPT** — still required by the outbound `_ApiRequest::setLogging(true)` calls. Do not remove it as "dead" when tidying.
@@ -77,22 +87,52 @@ Across 7 real events on two days (`Core.WorkerJobs` 563520/563544/563545/332143/
77
87
 
78
88
  ### Investigating this
79
89
  - **Defect A cannot be reproduced locally via a sandbox webhook.** Non-production worker2 is not SQS-driven (manually HTTP-invoked), so a PayPal sandbox webhook never traverses the Lambda→SQS→worker path and the watchdog never fires. Reproduce by timing the `isEventProcessed()` query directly against a large `Logs_Rate.Api`.
80
- - **Defect B** must be exercised via a real sandbox browser checkout.
90
+ - **The fallback handler** must be exercised via a real sandbox browser checkout with `POST /v2/entitlements` blocked in devtools.
81
91
  - **Always bound `Logs_Rate.Api` queries by `dtStamp` plus a `method`/route filter.** Unfiltered windows return hundreds of unrelated bulk contact-sync `PUT` rows and truncate.
82
92
 
93
+ ### Replaying a webhook worker2 already received (`test/@Mark/Rate/replay_paypal_webhook.php`)
94
+ Re-enqueues a stored PayPal webhook through the same `Rate/Webhook` action, replaying the exact payload from the `Core.WorkerJobs.parameters` column. Needed because when a handler refuses or defers an event the customer **has already paid**, and once PayPal's retry window closes that `WorkerJobs` row is the only surviving copy of the payload — PayPal will not re-deliver.
95
+
96
+ - Dry-run by default; `APPLY=1` to enqueue; `CONFIRM_PRODUCTION=1` on top for production. One job per run. Refuses any action that is not `Rate/Webhook`.
97
+ - **Replay passes signature checks only because `[paypal] verify_signature` is OFF** (the EB worker tier cannot reach `api-m.paypal.com`). If that is ever turned on, replay breaks.
98
+ - **Safe to re-run:** `handleSubscriptionActivated` returns early when `Subscriptions.token` exists, and `handlePaymentCompleted` returns early when the sale id is already on a payment row.
99
+ - **Order matters: ACTIVATED first, then `PAYMENT.SALE.COMPLETED`.**
100
+ - Not committed — the `test` repo is not branched.
101
+
102
+ ### Evidence (TRUE-82118) — the 2026-09-17 Rate incident
103
+ A Rate customer paid **$9.97 for Tech Support** on 2026-09-17 (PayPal subscription `I-3S346AJE68C5`) and got **no entitlement and no confirmation email**. Chain, all read-only against production:
104
+ - `Client_Rate.Contacts` 919 exists (created 2026-09-13 at SSO login), **no entitlement rows**.
105
+ - **Zero** `POST /v2/entitlements` from her browser in `Logs_Rate.Api` on either date — only GETs. So the browser path never fired and the fallback was the only remaining chance.
106
+ - `Core.WorkerJobs` 1628022 (`SUBSCRIPTION.CREATED`), 1628789 (`PAYMENT.SALE.COMPLETED`), 1628790 (`SUBSCRIPTION.ACTIVATED`) all `isSuccess=1`.
107
+ - The ACTIVATED handler **blocked on the address guard**: prod `Logs.Issue` id 798, reference `WX`, issueKey `RATE_PAYPAL_ACTIVATION_NO_SERVICE_ADDRESS`, urgency HIGH, still OPEN.
108
+ - `PAYMENT.SALE.COMPLETED` returned `deferred` because no subscription row existed.
109
+ - **No Tech Support purchase has EVER succeeded in production.** The only tech entitlements (`ET100007-9`) are from test user "andy america" in 2025; every real purchase since is warranty. That is why the universal address guard went unnoticed.
110
+
111
+ ### Verification (dev-sandbox, 2026-09-18) — frontend path only
112
+ Two browser purchases by contact 8, both correct:
113
+ - `ET100079` (id 112), Tech Support **Monthly**, $9.97, subscription 93, `dateStart` 2026-09-18, `dateEnd` 2026-10-18.
114
+ - `ET100080` (id 113), Tech Support **Annual**, $99.00, subscription 94, `dateStart` 2026-09-18, `dateEnd` **2027**-09-18 (yearly correctly added a year, not a month).
115
+ - Both: `serviceAddressId` NULL and both AIG columns NULL — **correct for tech**.
116
+ - Email: `Logs_Rate.Email` id 1, status SENT 16s after queueing, `retryCount` 0, correct TECH template, every placeholder substituted. Recipient rewritten to `devteam@goagilant.com` by the sandbox debug redirect — expected.
117
+ - **LIMIT: this proves the FRONTEND path only.** The entitlement came from the browser (`POST /v2/entitlements` 201 from a residential IP) and no `Rate/Webhook` job ran, so the worker2 fallback that was actually fixed is **still unverified**.
118
+
83
119
  ### Evidence (TRUE-80575)
84
120
  Zero `POST /v2/entitlements` rows in `Logs_Rate.Api` for all of 2026-08-01, while `Client_Rate` User 755 / Contact 773 / Customer 752 (Kwadwo "Drew" Moore) were created at 16:25:16 pre-payment and PayPal charged $89.97 at 16:32:47 (txn `0DY909616V1516622`, subscription `I-411CRRVH81U8`).
85
121
 
86
122
  ### Decisions
87
- - **Defect B approach (decided 2026-08-10):** add a `BILLING.SUBSCRIPTION.ACTIVATED` handler to `_Worker_Rate::Webhook`. The browser `onApprove` POST remains the **PRIMARY synchronous path**; the webhook is the **fallback net, not a replacement**. Not yet built.
123
+ - **`plan_requires_aig` is REQUIRED, not defaulted (2026-09-18).** A missing entry is a hard refusal. Defaulting to `false` would silently sell a warranty with none of its guards.
124
+ - **Repair procedure for the 2026-09-17 customer (decided, NOT yet executed — blocked on the worker2 deploy).** After the fix deploys, replay `Core.WorkerJobs` **1628790 (ACTIVATED) first, then 1628789 (`PAYMENT.SALE.COMPLETED`)**. PayPal will not re-deliver a 2026-09-17 webhook. The `PAYMENT.SALE.COMPLETED` replay also extends `Subscriptions.dateEnd` one billing period to 2026-10-17, matching PayPal's own `next_billing_time` in the stored payload. Resolve prod `Logs.Issue` reference `WX` once she is whole.
125
+ - **The ACTIVATED handler is the fallback net, not a replacement** for the browser `onApprove` POST.
88
126
  - **Rejected:** framing TRUE-80575 as a single "webhook broken" bug.
89
127
  - **Rejected (with evidence):** `Core.WorkerJobs` as an event ledger.
90
128
  - **Deferred to the lead developers:** a durable idempotency table. Do not write a migration.
91
129
  - **Resolved:** the "Rosetta White double charge" (recorded 2026-08-04) is closed. Both sales (`04T839079C8838430` and `10T52264HU684770N`, $89.97 each, 2026-07-10) have been **refunded**. No backfill was performed and none is needed — her covered-property address predates `validateAddress` logging and was never recoverable anyway.
92
130
 
93
131
  ### Still open (carry forward)
94
- - **`toga2-view`:** persist the covered property **pre-payment** — this **BLOCKS** the ACTIVATED handler for warranty; and add the missing `yearly_warranty` key to `PAYPAL_PLAN_IDS`.
95
- - The `BILLING.SUBSCRIPTION.ACTIVATED` handler itself.
132
+ - **Execute the 2026-09-17 repair replay** once the TRUE-82118 worker2 commit is pushed and deployed.
133
+ - **Verify the ACTIVATED fallback itself** — never exercised end to end (see *Verification*).
134
+ - **`_underscore` `Entitlement.php::postPost` attempts an AIG warranty contract for a TECH purchase** when the contact happens to have an address. Open, deferred. See [AIG contract creation](aig-contract-creation.md).
135
+ - **`toga2-view`:** persist the covered property **pre-payment**; and add the missing `yearly_warranty` key to `PAYPAL_PLAN_IDS`.
96
136
  - The webhook still does **not return HTTP 200 on failure paths**, so PayPal retries indefinitely.
97
137
  - Webhook **signature verification still disabled** — the endpoint is forgeable.
98
138
  - The PayPal `client_secret` is **still committed in plaintext** in `worker2/` and `api2/` `Config/production.ini` — rotate and move it out of the repos. (Location only; no value here.)
@@ -105,8 +145,9 @@ Zero `POST /v2/entitlements` rows in `Logs_Rate.Api` for all of 2026-08-01, whil
105
145
  - **`isSuccess = 0` on a `Rate/Webhook` job does NOT mean "no side effects."** A watchdog-killed job's PHP process can keep running and complete its API writes after the row is marked failed. Job 560657 started 07:53:03, killed 07:58:18, yet its `POST /v2/sales-orders` landed 08:00:10 and `POST /v2/entitlement-sales-orders` at 08:00:41 (worker IP 34.232.23.158). **Re-running or manually remediating a failed job risks duplicates** — always check `Logs_Rate.Api` first. This race also explains why some purchases appear to have worked.
106
146
  - **Yearly Home Warranty cannot be purchased at all (separate live bug, needs its own ticket).** `PAYPAL_PLAN_IDS` (`paypalService.ts:31-41`) defines `monthly_tech`, `yearly_tech`, `monthly_warranty` — but **no `yearly_warranty`**. `getPlanId()` builds `` `${plan}_${serviceType}` `` → `yearly_warranty` → `undefined` → `onError("Invalid plan configuration")`.
107
147
  - **Security: the webhook endpoint is unauthenticated and forgeable.** With `[paypal] verify_signature` disabled (documented in-code as a temporary tradeoff for the blocked egress), anyone who knows the endpoint URL can create a paid SalesOrder. Separately, the live PayPal `client_secret` is committed in plaintext in the `[paypal]` block of **both** `worker2/Config/production.ini` and `api2/Config/production.ini` — rotate and move it out of the repo. (Location only; value not recorded in this KB.)
108
- - **Defect A now also blocks portal cancellations, not just renewals.** As of 2026-08-04 the portal's self-service cancel flow relies on this worker consuming `BILLING.SUBSCRIPTION.CANCELLED` to stamp `Subscriptions.dateCancelled` (its reconciliation channel) — with every `Rate/Webhook` job dying, that path is dead. See [Rate tech-support subscription cancellation](subscription-cancellation.md).
109
- - **Hypothesis, NOT confirmed — the `return_url` lands on a page that creates nothing.** `usePayPalSubscription.ts:126` sets `return_url` to `${origin}/activation?success=true`. If the SDK falls back to a full-page redirect (popup blocked, mobile in-app browser), PayPal redirects there and `onApprove` never runs; `Activation.tsx` is a static "success, redirecting in 10s" page that reads no subscription id and calls no API. The DB can only prove the *absence* of the call, not why the browser didn't make it. Treat as a plausible mechanism, not a cause.
148
+ - **This worker is also the portal-cancellation channel.** The self-service cancel flow relies on it consuming `BILLING.SUBSCRIPTION.CANCELLED` to stamp `Subscriptions.dateCancelled` — any outage here kills cancellations too, not just renewals. See [subscription cancellation](subscription-cancellation.md).
149
+ - **⚠ A guard written for the warranty will silently kill the tech product.** Both products share one `handleSubscriptionActivated`, and tech has **no** covered property and **no** AIG contract. Any new check on address / phone / AIG must be gated on `$requiresAig`. This is exactly how every Tech Support purchase was blocked for months without anyone noticing — no tech purchase had ever succeeded in production, so there was no "it used to work" signal.
150
+ - **`/activation?success=true` reports success unconditionally.** `Activation.tsx` never reads `subscription_id` and calls no API — it counts down 10s and goes to `/home`. Any flow that lands there without `onApprove` having run tells a paying customer it worked while no entitlement exists. Do not point a redirect at it. (The `return_url`/`cancel_url` were removed for this reason; PayPal suppressing `onApprove` remains **unproven** — their App Switch docs say client-side callbacks keep working.)
110
151
 
111
152
  ## Related
112
153
  - [Rate profile](profile.md)
@@ -6,7 +6,7 @@ project: _Underscore
6
6
  client: rate
7
7
  type: client-feature
8
8
  status: active
9
- updated: 2026-09-16
9
+ updated: 2026-09-18
10
10
  owners: [mhammontree]
11
11
  files:
12
12
  - _underscore/Model/Rate/Entitlement.php
@@ -76,6 +76,8 @@ Remaining items are **optional / separate, not go-live blockers**:
76
76
  - **Shared interceptor with TRUE-79533 (WH per-address guard) — same `prePost`/`postPost`.** This workflow's `postPost` trigger + helpers live in the **same** `_Model_Rate_Entitlement` interceptors that the [Whole Home Warranty per-address guard](whole-home-warranty-purchase-guard.md) edits, and TRUE-79251 also adds the passThrough `const PASSTHROUGH_KEY_FULFILLMENT_PREFERENCE` + the contractFulfillmentPreference (EV-8) fix inside `prePost`. A bad `_production`→`_beta` merge already mis-resolved this once and fatally broke every Rate purchase save. When merging either ticket across branches, hand-merge into ONE `prePost()` keeping both features' logic — see the **shared-interceptor merge hazard** gotcha on that doc for the correct resolution and why `php -l` is not a sufficient gate.
77
77
  - **`border-radius` on a `<td>` renders SQUARE in many clients (Outlook).** Use an `<img>` for circular badges/icons (e.g. the green circle-check) instead of a CSS-rounded table cell.
78
78
  - **ASCII-only subjects (mojibake workaround).** `_Email::send()` does not set PHPMailer `CharSet=UTF-8`, so non-ASCII subjects/body (em dash, curly quotes, accents) mojibake. This workflow uses ASCII-only subjects; the proper root-cause fix is tracked on [`email-template-sending.md`](../../../2.0/apps/_underscore/features/email-template-sending.md).
79
+ - **⚠ A non-prod tenant missing its `EmailTemplates` seed looks like a code error, not a data gap.** It surfaces as a FAILED `Notification/EmailTemplate/Send` worker job naming the template **uuid**: *"Could not load `_Model_Client_EmailTemplate` object based upon these criteria: `uuid` = '2c4f8a1e-…'"* (dev-sandbox `Core.WorkerJobs` 141941, 2026-09-18). dev-sandbox `Client_Rate.EmailTemplates` held only 2 rows (Forgot Password, Talos); prod has both Rate templates as ids 3 and 4. Fix: run the existing `dbchanges2/Client_Rate/2026-06-30a - Rate purchase email templates.sql` on that environment. Check the tenant's `EmailTemplates` rows before debugging the sender.
80
+ - **⚠ OPEN — both templates render badly in email dark mode (from TRUE-79251; needs a design decision, not a code fix).** (1) The CTA button becomes an **empty dark bar** — background `#21292d` disappears into the client's dark theme and the white label is invisible. (2) The TOGA footer wordmark shows inside a **dark block**, because the image was rasterized onto the exact footer colour `#21292d`, which the client lightens in dark mode. Light mode is fine, which is why neither was caught — the *"flatten the asset onto the exact region background"* recipe above is precisely what breaks under dark mode.
79
81
  - **Worker DB registration** — see [`notification-email-template.md`](../../../2.0/apps/worker2/features/notification-email-template.md): worker actions do not auto-register `DB_CLIENT`.
80
82
 
81
83
  ## Related
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "toga-ai",
3
- "version": "1.0.844",
3
+ "version": "1.0.846",
4
4
  "description": "TOGA Technology Team Claude Knowledge System — shared AI coding harness with skills, knowledge base CLI, and project installer for Claude Code.",
5
5
  "keywords": [
6
6
  "claude",