toga-ai 1.0.678 → 1.0.680
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/knowledge/1.0/apps/worker/INDEX.md +3 -2
- package/knowledge/1.0/apps/worker/features/forecast2-netsuite-reconciliation.md +16 -1
- package/knowledge/1.0/apps/worker/features/forecast2-supporting-records-import.md +210 -0
- package/knowledge/1.0/apps/worker/workflows/tracing-a-worker-cron-run-in-production.md +39 -2
- package/knowledge/2.0/apps/talos/INDEX.md +1 -0
- package/knowledge/2.0/apps/talos/architecture.md +23 -0
- package/knowledge/2.0/apps/talos/features/chat-frontend.md +172 -0
- package/knowledge/2.0/apps/toga-blox/features/talos-assistant.md +52 -0
- package/knowledge/2.0/apps/toga25-supply/features/talos-integration.md +103 -2
- package/knowledge/2.0/apps/worker2/features/netsuite-supporting-record-backfill-worker.md +46 -8
- package/knowledge/2.0/apps/worker2/features/netsuite-supporting-record-webhook-importer.md +43 -12
- package/knowledge/INDEX.md +3 -2
- package/knowledge/clients/true/profile.md +3 -1
- package/knowledge/registry.json +8 -2
- package/knowledge/sessions/2026-08-28-openorderitems-location-fix-kyalamarthi.md +107 -0
- package/package.json +1 -1
|
@@ -7,7 +7,8 @@
|
|
|
7
7
|
| [Compass Manager Approval Reminder Emails (1.0 worker crons)](features/compass-manager-approval-reminder-emails.md) | Two 1.0 worker crons nag approvers about sales orders still waiting on a decision — one per Compass tenant. | worker/crons/toga2/compass/compass_email_reminders.php, worker/crons/toga2/compasscanada/compass_email_reminders.php, worker/schedules/cron.worker.sync.json, worker1.5/crons/toga2/compass/compass_email_reminders.php, worker1.5/schedules/cron.worker.json |
|
|
8
8
|
| [Compass Partial In-Transit & Delivered Emails (per package)](features/compass-partial-in-transit-delivered-emails.md) | Compass USA and Compass Canada send a **per-package** in-transit email (and a matching delivered email) instead of one email listing the whole order. | worker/crons/toga2/compass/update_salesorder_status_from_odp.php, worker/crons/toga2/compass/send_delivered_email.php, worker/crons/toga2/compasscanada/update_salesorder_status_from_odp.php, worker/crons/toga2/compasscanada/workflow/3_update_salesorder_status_from_grand_and_toy.php, worker/crons/toga2/compasscanada/send_delivered_email.php, worker/crons/toga2/compasscanada/workflow/test_partial_in_transit_email.php, worker/crons/toga2/compasscanada/workflow/test_partial_delivered_email.php, library/app/client/compasscanada.php |
|
|
9
9
|
| [Elite TOGA 2.0 → TOGaDeskSupport Standalone Attachment Sync](features/elite-togadesk-attachment-sync.md) | `sync_togadesk_elite_attachments.php` is a standalone cron (every 5 minutes) that syncs file attachments from TOGA 2.0 into TOGaDeskSupport for Elite. | worker/crons/toga2/elite/sync_togadesk_elite_attachments.php, worker/crons/toga2/elite/test_sync_togadesk_elite_attachments.php |
|
|
10
|
-
| [Forecast2 ↔ NetSuite Reconciliation & Trueup Tooling](features/forecast2-netsuite-reconciliation.md) | CLI tools to **audit** and **repair** drift between the production `Forecast` DB (core2) and NetSuite. | test/@dave/checker.php, worker2/Component/Forecast/SaleImport/SaleImport.php, test/@dave/looper.php, test/@dave/reconcile_netsuite_totals.php, test/@dave/fixer.php, tools/bin/forecast/fixer.php, test/@dave/analyze_netsuite_forecast_diff.php, test/@dave/trueup_sales.php, test/@dave/reconcile_drift_2023plus.php, test/@dave/probe_invoice_gap_2026.php, test/@dave/probe_creditmemo_gap_detail.php, test/@dave/trueup_open_orders.php, test/@dave/loop_trueup_open_orders.php, test/@dave/trueup_opportunities.php, test/@dave/probe_sales_gap_direct.php, test/@dave/probe_missing_oo_timing.php, test/@dave/probe_missing_oo_createdby.php, test/@dave/probe_drift_so_dates.php, test/@dave/probe_profit_invoices.php, test/@dave/probe_profit_gap.php, worker/crons/toga2/forecast2/common_import_sales_from_netsuite.php, worker/crons/toga2/forecast2/periodic_forecast_discrepancy_fix_open_orders.php, worker/crons/toga2/forecast2/import_open_orders.php, worker/schedules/cron.worker.infrastructure.json |
|
|
10
|
+
| [Forecast2 ↔ NetSuite Reconciliation & Trueup Tooling](features/forecast2-netsuite-reconciliation.md) | CLI tools to **audit** and **repair** drift between the production `Forecast` DB (core2) and NetSuite. | test/@dave/checker.php, worker2/Component/Forecast/SaleImport/SaleImport.php, test/@dave/looper.php, test/@dave/reconcile_netsuite_totals.php, test/@dave/fixer.php, tools/bin/forecast/fixer.php, tools/bin/forecast/checker.php, test/@dave/analyze_netsuite_forecast_diff.php, test/@dave/trueup_sales.php, test/@dave/reconcile_drift_2023plus.php, test/@dave/probe_invoice_gap_2026.php, test/@dave/probe_creditmemo_gap_detail.php, test/@dave/trueup_open_orders.php, test/@dave/loop_trueup_open_orders.php, test/@dave/trueup_opportunities.php, test/@dave/probe_sales_gap_direct.php, test/@dave/probe_missing_oo_timing.php, test/@dave/probe_missing_oo_createdby.php, test/@dave/probe_drift_so_dates.php, test/@dave/probe_profit_invoices.php, test/@dave/probe_profit_gap.php, worker/crons/toga2/forecast2/common_import_sales_from_netsuite.php, worker/crons/toga2/forecast2/periodic_forecast_discrepancy_fix_open_orders.php, worker/crons/toga2/forecast2/import_open_orders.php, worker/schedules/cron.worker.infrastructure.json |
|
|
11
|
+
| [Forecast2 Supporting-Records Nightly Import (1.0 cron) — Customers & ConsolidatedCustomers](features/forecast2-supporting-records-import.md) | The nightly **1.0** pull that keeps the Forecast2 lookup/dimension tables (`Forecast.Accounts`, `Classifications`, `Customers`, `ConsolidatedCustomers`, `Employ | worker/crons/toga2/forecast2/import_supporting_records.php, worker/crons/toga2/forecast2/common_import_sales_from_netsuite.php, worker/schedules/cron.worker.infrastructure.json, worker/.ebextensions/009_setup_phpini.config, worker/crons/sync/netsuite/netsuite_customers.php, library/app/api/netsuite/rest.php, library/app/model/forecast2/customer.php, library/app/model/forecast2/consolidatedcustomer.php |
|
|
11
12
|
| [NetSuite Sales Order Sales Rep Sourcing (Staples & ODP EDI orders)](features/netsuite-sales-order-sales-rep-sourcing.md) | How the **sales rep** on a NetSuite Sales Order is determined for the two 1.0 `worker` EDI order-creation integrations (Staples cXML and Compass/ODP EDI). | worker/crons/toga2/compass/workflow/5_create_netsuite_sales_orders_from_office_depot_purchase_orders.php, worker/crons/sync/staples/sync_staples_cxml.php, test/@Mark/NetSuite/TRUE_80451_customer_salesrep_diag.php |
|
|
12
13
|
| [NetSuite → TOGa Supply Per-Client Sync (thin wrappers)](features/netsuite-togasupply-per-client-sync.md) | Syncs NetSuite transactions (sales orders, purchase orders, invoices, item receipts, item fulfillments, inventory adjustments) into each TOGa Supply (2.0) clien | worker/crons/toga2/netsuite/common_sync_togasupply.php, worker/crons/toga2/netsuite/sync_togasupply_canon.php, worker/crons/toga2/netsuite/sync_togasupply_elite.php, worker/schedules/cron.worker.sync.json, dbchanges2/_modules/netsuite/2026-04-01 - Parameters.sql, library/app/api/toga2.php, library/app/api/netsuite/rest.php, library/app/framework.php, library/app/systemmonitor/netsuiteintegration.php, test/@srija/Elite Testing/Service Requests/test_sync_togasupply_elite_section.php, test/@srija/Elite Testing/Service Requests/test_diagnose_togasupply_elite.php |
|
|
13
14
|
| [OneUptime Server monitor + disk/memory hygiene on the 1.0 worker EB host](features/oneuptime-server-monitor-host-hygiene.md) | The 1.0 `agilant-worker` EB environment runs on the **legacy Amazon Linux 1 PHP 7.2 platform** (Apache httpd/prefork, s3fs mounts, cron) and repeatedly went dow | worker/.ebextensions/040_disk_memory_hygiene.config, worker/.ebextensions/045_oneuptime_agent.config, worker/ebs/cron.worker.php, worker/ebs/mount-s3fs-folders.php, worker/ebs/apache_settings.php, worker/ebs/setup_phpini.php |
|
|
@@ -17,4 +18,4 @@
|
|
|
17
18
|
| [Diagnosing frozen 1.0 worker cron check-ins (Sentry "missed" flood)](workflows/diagnosing-frozen-cron-checkins.md) | When 1.0 worker cron timestamps freeze and Sentry project `worker1` fills with **`missed`** check-ins, the intuitive diagnosis — a wedged `App_Framework::isProc | worker/.ebextensions/cron.config, library/app/worker.php |
|
|
18
19
|
| [isFulfillable Multi-Client Backfill (all togasupply clients)](workflows/isfulfillable-multi-client-backfill.md) | One-time backfill that catches up `Items.isFulfillable` on **existing** items across **all 17 togasupply clients** (AIG, Broward Sheriff, Canon, Endeavor Health | worker/crons/toga2/netsuite/backfill_isfulfillable_all_clients.php, library/app/api/toga2.php |
|
|
19
20
|
| [Onboarding a Client to the NetSuite TOGa Supply Sync](workflows/onboarding-client-to-netsuite-togasupply-sync.md) | How to add a new TOGa 2 client to the per-client NetSuite → TOGa Supply importer (`worker/crons/toga2/netsuite/`). | worker/crons/toga2/netsuite/sync_togasupply.php, worker/crons/toga2/netsuite/common_sync_togasupply.php, worker/crons/toga2/netsuite/sync_togasupply_elite.php, worker/schedules/cron.worker.sync.json, dbchanges2/_modules/netsuite/2026-04-01 - Parameters.sql, dbchanges2/_modules/netsuite/2026-08-05 - CLEAN NETSUITE CLINET.SQL |
|
|
20
|
-
| [Tracing a 1.0 worker cron run in production (no stdout, silent skips)](workflows/tracing-a-worker-cron-run-in-production.md) | How to answer *"did this cron actually run, and what did it do?"* on the 1.0 `worker` tier, where **there is no usable stdout** and **a skipped run leaves no tr | worker/ebs/cron.worker.php, worker/schedules/cron.worker.sync.json, library/app/framework.php |
|
|
21
|
+
| [Tracing a 1.0 worker cron run in production (no stdout, silent skips)](workflows/tracing-a-worker-cron-run-in-production.md) | How to answer *"did this cron actually run, and what did it do?"* on the 1.0 `worker` tier, where **there is no usable stdout** and **a skipped run leaves no tr | worker/ebs/cron.worker.php, worker/schedules/cron.worker.sync.json, worker/schedules/cron.worker.infrastructure.json, worker/.ebextensions/009_setup_phpini.config, library/app/framework.php |
|
|
@@ -6,7 +6,7 @@ project: Worker
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-08-
|
|
9
|
+
updated: 2026-08-28
|
|
10
10
|
owners: [dfranks, jcardinal, kyalamarthi]
|
|
11
11
|
files:
|
|
12
12
|
- test/@dave/checker.php
|
|
@@ -15,6 +15,7 @@ files:
|
|
|
15
15
|
- test/@dave/reconcile_netsuite_totals.php
|
|
16
16
|
- test/@dave/fixer.php
|
|
17
17
|
- tools/bin/forecast/fixer.php
|
|
18
|
+
- tools/bin/forecast/checker.php
|
|
18
19
|
- test/@dave/analyze_netsuite_forecast_diff.php
|
|
19
20
|
- test/@dave/trueup_sales.php
|
|
20
21
|
- test/@dave/reconcile_drift_2023plus.php
|
|
@@ -643,9 +644,23 @@ None — Forecast2 is a single shared dataset.
|
|
|
643
644
|
enqueuer execution log on the NetSuite side).
|
|
644
645
|
- These tools live in `test/@dave/` (developer tooling), but `trueup_open_orders` has been run
|
|
645
646
|
against production. The `Defaults`/checkpoint mechanics of the scheduled sync are separate.
|
|
647
|
+
- **⚠ SECURITY — the `tools` repo copies of these scripts hardcode PRODUCTION credentials inline.**
|
|
648
|
+
Found 2026-08-28: `tools/bin/forecast/fixer.php:1424` and `tools/bin/forecast/checker.php:678`
|
|
649
|
+
carry live production **Forecast read-replica** database credentials **inline in source**, not via
|
|
650
|
+
`worker/config.worker.ini`. Per the team security rule these are **committed secrets and must be
|
|
651
|
+
treated as compromised and rotated**; deleting the lines is not sufficient because they are in git
|
|
652
|
+
history. Values are deliberately **not** reproduced here, and must never be pasted into a knowledge
|
|
653
|
+
doc, a session file, or a ticket. Not yet actioned — **needs its own ticket**, and a senior should
|
|
654
|
+
sequence the rotation against whatever else reads that replica.
|
|
646
655
|
|
|
647
656
|
## Change history
|
|
648
657
|
|
|
658
|
+
- 2026-08-28 — **⚠ Security finding (not fixed):** the `tools` repo copies of these scripts —
|
|
659
|
+
`tools/bin/forecast/fixer.php:1424` and `tools/bin/forecast/checker.php:678` — hardcode **live
|
|
660
|
+
production Forecast read-replica credentials inline in source**, rather than reading
|
|
661
|
+
`worker/config.worker.ini`. Treat as compromised, rotate, and remember the values are in git history
|
|
662
|
+
so removing the lines is not enough. No values recorded here. Surfaced incidentally while
|
|
663
|
+
investigating TRUE-81213. (kyalamarthi)
|
|
649
664
|
- 2026-08-27 — **Taught the nightly open-orders discrepancy-fix cron about `locationId` — and turned it
|
|
650
665
|
into its own backfill (TRUE-79162).** The cron had **zero** occurrences of "location" in 775 lines
|
|
651
666
|
(line SuiteQL, both SQL builders, and `$compare` all omitted it) and it delete+reinserts rows, so it
|
|
@@ -0,0 +1,210 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Forecast2 Supporting-Records Nightly Import (1.0 cron) — Customers & ConsolidatedCustomers
|
|
3
|
+
framework: "1.0"
|
|
4
|
+
repo: worker
|
|
5
|
+
project: Worker
|
|
6
|
+
client: shared
|
|
7
|
+
type: feature
|
|
8
|
+
status: active
|
|
9
|
+
updated: 2026-08-28
|
|
10
|
+
owners: ["kyalamarthi"]
|
|
11
|
+
files:
|
|
12
|
+
- worker/crons/toga2/forecast2/import_supporting_records.php
|
|
13
|
+
- worker/crons/toga2/forecast2/common_import_sales_from_netsuite.php
|
|
14
|
+
- worker/schedules/cron.worker.infrastructure.json
|
|
15
|
+
- worker/.ebextensions/009_setup_phpini.config
|
|
16
|
+
- worker/crons/sync/netsuite/netsuite_customers.php
|
|
17
|
+
- library/app/api/netsuite/rest.php
|
|
18
|
+
- library/app/model/forecast2/customer.php
|
|
19
|
+
- library/app/model/forecast2/consolidatedcustomer.php
|
|
20
|
+
related:
|
|
21
|
+
- ./forecast2-netsuite-reconciliation.md
|
|
22
|
+
- ../workflows/tracing-a-worker-cron-run-in-production.md
|
|
23
|
+
- ../architecture.md
|
|
24
|
+
- ../../../../2.0/apps/worker2/features/netsuite-supporting-record-webhook-importer.md
|
|
25
|
+
- ../../../../2.0/apps/worker2/features/netsuite-supporting-record-backfill-worker.md
|
|
26
|
+
---
|
|
27
|
+
|
|
28
|
+
## Summary
|
|
29
|
+
|
|
30
|
+
The nightly **1.0** pull that keeps the Forecast2 lookup/dimension tables (`Forecast.Accounts`,
|
|
31
|
+
`Classifications`, `Customers`, `ConsolidatedCustomers`, `Employees`) in step with NetSuite. It is the
|
|
32
|
+
older sibling — and the standing backstop — of the
|
|
33
|
+
[2.0 webhook importer](../../../../2.0/apps/worker2/features/netsuite-supporting-record-webhook-importer.md),
|
|
34
|
+
and it is what ultimately feeds **Power BI**'s customer dimension, including the
|
|
35
|
+
**Consolidated Business Name** column.
|
|
36
|
+
|
|
37
|
+
This doc goes deep on the **Customers** section, because that is where the interesting behavior
|
|
38
|
+
lives: it is the only section that writes a **parent** table (`ConsolidatedCustomers`) as a
|
|
39
|
+
side-effect of walking **children** (`Customers`), and that asymmetry is what produced TRUE-81213.
|
|
40
|
+
|
|
41
|
+
## Key files / entry points
|
|
42
|
+
|
|
43
|
+
- **`worker/crons/toga2/forecast2/import_supporting_records.php`** — the scheduled wrapper. Sets the
|
|
44
|
+
`SHOULD_SYNC_*` consts (`ACCOUNTS`, `CLASSIFICATIONS`, `CUSTOMERS`, `EMPLOYEES` = true; jobs/items/
|
|
45
|
+
sales/open-orders/opportunities = false) and then `require`s the shared body. Wraps the run in
|
|
46
|
+
`App_Worker::checkIn('forecast-import-supporting-records')` / `checkOut()`.
|
|
47
|
+
- **`worker/crons/toga2/forecast2/common_import_sales_from_netsuite.php`** — the shared body, one
|
|
48
|
+
`if (SHOULD_SYNC_X …)` block per record type, executed **in file order**:
|
|
49
|
+
accounts (`:85`) → classifications (`:126`) → **customers (`:193`)** → employees (`:290`) → items,
|
|
50
|
+
sales, opportunities, open orders (all off for this wrapper).
|
|
51
|
+
- **`App_Api_Netsuite_Rest::listCustomers()`** (`library/app/api/netsuite/rest.php:294`) — the source
|
|
52
|
+
query. **No date/`lastmodifieddate` filter**: every pass pulls *every* NetSuite customer.
|
|
53
|
+
- **`worker/schedules/cron.worker.infrastructure.json`** — schedule entry *"Forecast 2.0 - Import
|
|
54
|
+
supporting records (2:00 AM)."*, `0 2 * * *`, `active: 1`.
|
|
55
|
+
|
|
56
|
+
## Where and when it actually runs
|
|
57
|
+
|
|
58
|
+
- The entry lives in the **`infrastructure`** role schedule, so it runs on **exactly one** of the 7
|
|
59
|
+
production `worker` instances — whichever currently holds that role. Role assignment is dynamic;
|
|
60
|
+
the live roster is `Vision_Log.Workers` (`workerName`, `workerInstanceId`, `dtHeartbeat`) and the
|
|
61
|
+
instances are also EC2-tagged `Worker=<ROLE>-WORKER`. Instance ids are ephemeral — a stale heartbeat
|
|
62
|
+
gets the instance terminated and the role re-claimed by its replacement. See
|
|
63
|
+
[Worker architecture](../architecture.md) and
|
|
64
|
+
[tracing a cron run](../workflows/tracing-a-worker-cron-run-in-production.md).
|
|
65
|
+
- **`0 2 * * *` is 2:00 AM CENTRAL, not UTC.** `worker/.ebextensions/009_setup_phpini.config` symlinks
|
|
66
|
+
`/etc/localtime` to `America/Chicago`, so `crond` — and the production MySQL `NOW()` you compare
|
|
67
|
+
against — are both on Central (07:00 UTC while CDT is in effect).
|
|
68
|
+
`periodic_forecast_discrepancy_fix.php` is scheduled at the **same instant on the same instance**.
|
|
69
|
+
- **Manual trigger:** `php /var/www/html/crons/toga2/forecast2/import_supporting_records.php` — the
|
|
70
|
+
path is exactly the `cron` value from the schedule JSON, rooted at `/var/www/html/crons/`.
|
|
71
|
+
|
|
72
|
+
## How it works — the Customers section (`:193` onward)
|
|
73
|
+
|
|
74
|
+
Two writes happen per NetSuite customer, and **they are guarded differently**. This is the single
|
|
75
|
+
most important thing to understand about this cron.
|
|
76
|
+
|
|
77
|
+
1. **Block A — the parent (`ConsolidatedCustomers`), lines ~214-253, UNGUARDED.**
|
|
78
|
+
The custom field `custentity10` (`NETSUITE_CUSTOM_FIELD_ID__CONSOLIDATED_CUSTOMER`) is read off the
|
|
79
|
+
customer, and its referenced group is upserted into `ConsolidatedCustomers` keyed on
|
|
80
|
+
`netsuiteInternalId` (create if unseen, rename if the name differs), then cached in
|
|
81
|
+
`$lookupConsolidatedCustomers`. This block has always worked.
|
|
82
|
+
2. **Block B — the child link (`Customers.consolidatedCustomerId`), lines ~256-259 onward, GUARDED.**
|
|
83
|
+
The row is written only when it is new **or** its change guard trips.
|
|
84
|
+
|
|
85
|
+
**`ConsolidatedCustomers` is not a synced dimension table.** Nothing anywhere enumerates NetSuite's
|
|
86
|
+
consolidated-customer list; parents exist **only as a by-product of walking customers**. Consequences
|
|
87
|
+
worth internalizing before you diagnose anything here:
|
|
88
|
+
|
|
89
|
+
- The parent table is **complete and then some** — prod holds **2,742** rows with **0** NULL/empty
|
|
90
|
+
names against **2,729** distinct `custentity10` values actually in use. It carries a *surplus*
|
|
91
|
+
because **nothing ever deletes from it**.
|
|
92
|
+
- Therefore *"a Consolidated Business Name is blank"* is almost never a **missing group**; it is a
|
|
93
|
+
**missing link**. The two failures have different causes and different fixes.
|
|
94
|
+
|
|
95
|
+
## The TRUE-81213 fix (`&&` → `||`), commit `b846e97c`
|
|
96
|
+
|
|
97
|
+
The Block B guard read:
|
|
98
|
+
|
|
99
|
+
```php
|
|
100
|
+
name != entityId && consolidatedCustomerId != $consolidatedCustomerId
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
so an existing customer was rewritten **only when BOTH changed**. Names almost never change, which
|
|
104
|
+
made `Forecast.Customers.consolidatedCustomerId` effectively **WRITE-ONCE** — set at first insert and
|
|
105
|
+
never repaired afterwards. A Consolidated Customer assigned in NetSuite *after* the customer row
|
|
106
|
+
already existed was never linked, and Power BI reported a blank Consolidated Business Name for it
|
|
107
|
+
forever. The operator is now `||` (either field differing triggers the update), with an inline
|
|
108
|
+
comment recording why.
|
|
109
|
+
|
|
110
|
+
**Blast radius was ~19× the ticket.** The ticket described 33 strategic accounts; the real population
|
|
111
|
+
was ~788. Production, before → after the first corrected run:
|
|
112
|
+
|
|
113
|
+
| Measure | Before | After |
|
|
114
|
+
|---|---|---|
|
|
115
|
+
| Customers with a `consolidatedCustomerId` | 4,062 | **4,855** (+793) |
|
|
116
|
+
| Strategic accounts with a blank Consolidated Business Name | 33 | **1** |
|
|
117
|
+
|
|
118
|
+
The "33" was only the visible tip: those are the strategic accounts that *also* had `accountType`
|
|
119
|
+
populated, and `accountType` exists for just **253 of 9,213** customers. The number that actually
|
|
120
|
+
predicted the outcome was NetSuite's **4,850 customers carrying a `custentity10`** — which the +793
|
|
121
|
+
matched almost exactly.
|
|
122
|
+
|
|
123
|
+
## Diagnosing a blank Consolidated Business Name
|
|
124
|
+
|
|
125
|
+
1. **Count orphan groups before anything else.** `ConsolidatedCustomers` rows with **zero** member
|
|
126
|
+
customers were the bug's signature (**192** of them pre-fix) — parents dutifully created by Block A
|
|
127
|
+
while Block B refused to link the child. A high orphan count means *broken links*; genuinely
|
|
128
|
+
*missing groups* would instead show up as customers whose `custentity10` has no row at all. Do not
|
|
129
|
+
conflate the two.
|
|
130
|
+
2. **Do not use `accountType` as a proxy for coverage.** It is set on 2.7% of customers, so any
|
|
131
|
+
"blank" count filtered by it under-reports the real population by more than an order of magnitude.
|
|
132
|
+
3. **Compare against NetSuite's `custentity10` population**, not against the ticket's row count.
|
|
133
|
+
|
|
134
|
+
## Causes that look right and are not (both ruled out in production)
|
|
135
|
+
|
|
136
|
+
- **The SuiteQL alias-casing bug in the 2.0 handler is NOT this.**
|
|
137
|
+
`worker2/Worker/Netsuite/Customer.php` selects `BUILTIN.DF(custentity10) AS consolidatedName` and
|
|
138
|
+
reads `$record->consolidatedName`; SuiteQL lower-cases aliases, so that read is always `NULL`. It has
|
|
139
|
+
caused **zero** data damage: **0 NULL names across all 2,742 rows**. `custentity10` is *also* selected
|
|
140
|
+
**unaliased**, so the FK link still resolves; only a brand-new parent's display name could be lost,
|
|
141
|
+
and this nightly cron creates parents first, with correct names, via its own lowercase alias. Still a
|
|
142
|
+
real latent bug — just not this one.
|
|
143
|
+
- **`accountType` being populated does NOT mean the webhook ran.** Prod `Core.WorkerJobs` has only
|
|
144
|
+
`Netsuite/Customer/put` **30** + `post` **5** = **35** jobs *ever* (all since 2026-08-19), whereas
|
|
145
|
+
`Netsuite/CustomerAccountTypeBackfill/Backfill` ran **twice** on 2026-08-19. The backfill issues a raw
|
|
146
|
+
`UPDATE … SET accountType` and never touches `consolidatedCustomerId` — which is exactly why 253
|
|
147
|
+
customers carry an `accountType` next to a stale consolidated link.
|
|
148
|
+
|
|
149
|
+
## Why no backfill migration was needed — and why one would have been wrong
|
|
150
|
+
|
|
151
|
+
`listCustomers()` has **no date filter**, so the corrected cron re-links the whole population on its
|
|
152
|
+
next pass. That is not theory: one run wrote 793 rows. **The fix is self-healing; ship the code and
|
|
153
|
+
wait one night.**
|
|
154
|
+
|
|
155
|
+
A `dbchanges` SQL backfill would have been the **wrong tool**, not merely redundant: the
|
|
156
|
+
customer→group mapping exists *only* in NetSuite's `custentity10`. The database cannot derive it, and
|
|
157
|
+
name-matching would have silently assigned customers to the **wrong** group. Contrast
|
|
158
|
+
[`accountType`](../../../../2.0/apps/worker2/features/netsuite-supporting-record-backfill-worker.md),
|
|
159
|
+
which genuinely did need a backfill worker — precisely because the 1.0 model does not declare that
|
|
160
|
+
column, so no nightly writer exists for it at all.
|
|
161
|
+
|
|
162
|
+
**Rule of thumb:** if the nightly full pull owns the column, fix the pull. Write a backfill only for a
|
|
163
|
+
column the pull cannot or will not write.
|
|
164
|
+
|
|
165
|
+
## Gotchas / known issues
|
|
166
|
+
|
|
167
|
+
- **A manual run silently no-ops while another run is in flight.** `import_supporting_records.php`
|
|
168
|
+
opens with `if (App_Framework::isProcessRunning()) exit;` — no output, no
|
|
169
|
+
`Common.CronJobExecutions` row. Check `ps -ef` for the script path before concluding your manual run
|
|
170
|
+
did nothing. (The guard is a substring match, so an editor or `tail -f` holding that path blocks it
|
|
171
|
+
too — see [tracing a cron run](../workflows/tracing-a-worker-cron-run-in-production.md).)
|
|
172
|
+
- **Customers is the THIRD section.** A PHP warning anywhere in accounts or classifications routes to
|
|
173
|
+
`App_Error::handleError` → `exit` on this tier, so the process dies **before customers is ever
|
|
174
|
+
reached**. "Customers didn't sync" is frequently "an earlier section blew up".
|
|
175
|
+
- **The 1.0 `worker` app cannot be booted on a laptop** — it dies on a missing `vendor/autoload.php`.
|
|
176
|
+
There is no local path to running this cron; probe NetSuite from a 2.0 harness instead (see
|
|
177
|
+
[running worker2 locally](../../../../2.0/apps/worker2/workflows/running-worker2-locally.md)).
|
|
178
|
+
- **⚠ Do not confuse this pipeline with `crons/sync/netsuite/netsuite_customers.php`** — a *different*
|
|
179
|
+
customer sync writing `db_netsuite.NetsuiteCustomers` (which feeds `PowerBISales`), and it is
|
|
180
|
+
**dead**. Line 20 calls `App_NetSuite::listAllCustomers()`, which **does not exist** anywhere in
|
|
181
|
+
`library` (only `listCustomers`, `listAllCustomersForSalesRep`, `listAllCustomersForParentCustomer`).
|
|
182
|
+
It is `active: 1` every 4 hours in `cron.worker.sync.json`, so it fatals on **every** run and has been
|
|
183
|
+
starving its table. The same file is also insert-only (no `else` branch) and its unguarded
|
|
184
|
+
`foreach ($customer->customFieldList->customField …)` would abort the whole run on a customer with no
|
|
185
|
+
custom fields. **Not fixed — needs its own ticket.**
|
|
186
|
+
|
|
187
|
+
## Change history
|
|
188
|
+
|
|
189
|
+
- 2026-08-28 — **TRUE-81213 fixed (`b846e97c`, verified in production).** The Block B change guard used
|
|
190
|
+
`&&` where it meant `||`, making `Customers.consolidatedCustomerId` write-once and blanking Power BI's
|
|
191
|
+
Consolidated Business Name for every customer whose group was assigned after its row existed. First
|
|
192
|
+
corrected run: consolidated links **4,062 → 4,855 (+793)**, blank strategic accounts **33 → 1**.
|
|
193
|
+
Created this doc to record what the investigation established beyond the fix: `ConsolidatedCustomers`
|
|
194
|
+
is a **by-product** of the unguarded Block A (complete, 2,742 rows, surplus, never deleted from), so
|
|
195
|
+
the bug's signature is **192 orphan groups**, not missing groups; `accountType` (253/9,213) is a
|
|
196
|
+
badly-biased proxy for coverage; the SuiteQL `consolidatedName` casing bug and "the webhook ran" were
|
|
197
|
+
both **ruled out**; and the fix is **self-healing** because `listCustomers()` has no date filter, so a
|
|
198
|
+
SQL backfill was unnecessary *and* would have mis-assigned groups. Also recorded where/when the cron
|
|
199
|
+
runs (infrastructure role, 2:00 AM **Central**) and the dead
|
|
200
|
+
`sync/netsuite/netsuite_customers.php` sibling. (kyalamarthi)
|
|
201
|
+
|
|
202
|
+
## Related docs
|
|
203
|
+
|
|
204
|
+
- [NetSuite Supporting-Record Webhook Importer](../../../../2.0/apps/worker2/features/netsuite-supporting-record-webhook-importer.md)
|
|
205
|
+
— the 2.0 real-time replacement this cron backstops.
|
|
206
|
+
- [NetSuite Supporting-Record Backfill Worker](../../../../2.0/apps/worker2/features/netsuite-supporting-record-backfill-worker.md)
|
|
207
|
+
— the bulk reconciliation path, and when it is the wrong tool.
|
|
208
|
+
- [Forecast2 ↔ NetSuite Reconciliation & Trueup Tooling](./forecast2-netsuite-reconciliation.md).
|
|
209
|
+
- [Tracing a 1.0 worker cron run in production](../workflows/tracing-a-worker-cron-run-in-production.md).
|
|
210
|
+
- [Worker (1.0 Framework) Architecture](../architecture.md) — the 7-instance role election.
|
|
@@ -6,15 +6,18 @@ project: Worker
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: workflow
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-08-
|
|
10
|
-
owners: ["bala"]
|
|
9
|
+
updated: 2026-08-28
|
|
10
|
+
owners: ["bala", "kyalamarthi"]
|
|
11
11
|
files:
|
|
12
12
|
- worker/ebs/cron.worker.php
|
|
13
13
|
- worker/schedules/cron.worker.sync.json
|
|
14
|
+
- worker/schedules/cron.worker.infrastructure.json
|
|
15
|
+
- worker/.ebextensions/009_setup_phpini.config
|
|
14
16
|
- library/app/framework.php
|
|
15
17
|
related:
|
|
16
18
|
- ../architecture.md
|
|
17
19
|
- ./diagnosing-frozen-cron-checkins.md
|
|
20
|
+
- ../features/forecast2-supporting-records-import.md
|
|
18
21
|
- ../../library/features/cron-execution-monitoring.md
|
|
19
22
|
- ../../../1.0/standards/backend-php.md
|
|
20
23
|
---
|
|
@@ -41,6 +44,31 @@ frequency makes you read a normal gap as an outage.
|
|
|
41
44
|
Also check whether the job is registered in the env you are testing in at all: this one is **not**
|
|
42
45
|
in `cron.beta.json`, so beta never runs it on a schedule. Nothing on beta is evidence about it.
|
|
43
46
|
|
|
47
|
+
### Which of the 7 boxes runs it — and in which timezone
|
|
48
|
+
|
|
49
|
+
Two facts you need before you go looking for evidence, both easy to get wrong:
|
|
50
|
+
|
|
51
|
+
- **The role file names the instance.** A job in `cron.worker.<role>.json` runs on **exactly one**
|
|
52
|
+
instance: whichever currently holds that role. There is **no static instance→role mapping** —
|
|
53
|
+
assignment is dynamic and claim-the-first-free-slot (see
|
|
54
|
+
[Worker architecture](../architecture.md)). To find today's box: read
|
|
55
|
+
**`Vision_Log.Workers`** (`workerName`, `workerInstanceId`, `dtHeartbeat`), or look for the EC2 tag
|
|
56
|
+
**`Worker=<ROLE>-WORKER`**. Instance ids are **ephemeral** — a stale heartbeat gets the instance
|
|
57
|
+
terminated and its role re-claimed by the replacement, so an instance id you noted last week is
|
|
58
|
+
probably meaningless today. Worked example: `toga2/forecast2/import_supporting_records.php` is in
|
|
59
|
+
`cron.worker.infrastructure.json`, so only the current **`infrastructure`** worker ever runs it.
|
|
60
|
+
- **⚠ The crontab runs on CENTRAL time, not UTC.** `.ebextensions/009_setup_phpini.config` symlinks
|
|
61
|
+
`/etc/localtime` to `America/Chicago`, so `0 2 * * *` means **2:00 AM Central** (07:00 UTC under
|
|
62
|
+
CDT). The production MySQL also returns Central from `NOW()`, so timestamps you pull back agree
|
|
63
|
+
with the schedule — but anything you compare against a UTC source (CloudWatch, an AWS console, a
|
|
64
|
+
NetSuite timestamp) is off by 5–6 hours. Most "the cron didn't run at the time I expected"
|
|
65
|
+
reports are this.
|
|
66
|
+
|
|
67
|
+
**To run it by hand:** `php /var/www/html/crons/<path exactly as it appears in the schedule JSON>` —
|
|
68
|
+
e.g. `php /var/www/html/crons/toga2/forecast2/import_supporting_records.php`. Do this on the
|
|
69
|
+
role-holding instance, and read Step 4's overlap-guard note first: a manual run **silently exits**
|
|
70
|
+
if a scheduled run is already in flight.
|
|
71
|
+
|
|
44
72
|
## Step 2 — echo/print goes nowhere useful, so do not debug with it
|
|
45
73
|
The crontab line is assembled in `worker/ebs/cron.worker.php`, and the two blocks that build it
|
|
46
74
|
differ:
|
|
@@ -115,6 +143,15 @@ to **2.0 only**. The 1.0 standard wins in `library/`, `worker/` and every other
|
|
|
115
143
|
signal, it is a latent bug.
|
|
116
144
|
|
|
117
145
|
## Change history
|
|
146
|
+
- 2026-08-28 — Added **"which of the 7 boxes runs it — and in which timezone"** to Step 1, after a
|
|
147
|
+
production trace of the Forecast2 supporting-records cron (TRUE-81213). A `cron.worker.<role>.json`
|
|
148
|
+
job runs on exactly one dynamically-assigned instance: find it via `Vision_Log.Workers` or the EC2
|
|
149
|
+
tag `Worker=<ROLE>-WORKER`, never a remembered instance id. And **the crontab runs on Central**
|
|
150
|
+
(`.ebextensions/009_setup_phpini.config` symlinks `/etc/localtime` to `America/Chicago`), as does
|
|
151
|
+
production MySQL `NOW()` — so `0 2 * * *` is 2:00 AM Central, and any comparison against a UTC
|
|
152
|
+
source is off by 5–6 hours. Recorded the manual-trigger form
|
|
153
|
+
`php /var/www/html/crons/<schedule-json path>` and the fact that it silently exits under the overlap
|
|
154
|
+
guard. (kyalamarthi)
|
|
118
155
|
- 2026-08-25 — Documented from a prod diagnostic session on the Compass ODP 850 importer: the
|
|
119
156
|
schedule JSON (not the file header) is authoritative, worker cron stdout is unreadable (base
|
|
120
157
|
block's `WORKER_ERROR` redirect commented out, role block writes a random filename),
|
|
@@ -4,6 +4,7 @@
|
|
|
4
4
|
|-----|---------|-------|
|
|
5
5
|
| [TOGa IQ (talos) Architecture](architecture.md) | **TOGa IQ** is TOGA Technology's AI agent platform. | talos/libs/aegra-api/src/aegra_api/main.py, talos/libs/aegra-api/src/aegra_api/settings.py, talos/libs/aegra-api/src/aegra_api/config.py, talos/libs/aegra-api/src/aegra_api/core/tenant_router.py, talos/libs/aegra-api/src/aegra_api/core/control_plane_db.py, talos/libs/aegra-api/src/aegra_api/core/auth_middleware.py, talos/libs/aegra-api/src/aegra_api/services/run_executor.py, talos/libs/aegra-api/src/aegra_api/services/langgraph_service.py, talos/libs/aegra-api/src/aegra_api/services/graph_factory.py, talos/libs/aegra-api/src/aegra_api/services/streaming_service.py, talos/agents/talos_agent/graph.py, talos/docker-compose.yml, talos/deployments/docker/Dockerfile |
|
|
6
6
|
| [aegra-api — Agent Protocol HTTP + Execution Pipeline](features/aegra-api.md) | `aegra-api` is the **FastAPI Agent Protocol server** at the heart of TOGa IQ. | talos/libs/aegra-api/src/aegra_api/main.py, talos/libs/aegra-api/src/aegra_api/settings.py, talos/libs/aegra-api/src/aegra_api/config.py, talos/libs/aegra-api/src/aegra_api/api/assistants.py, talos/libs/aegra-api/src/aegra_api/api/threads.py, talos/libs/aegra-api/src/aegra_api/api/runs.py, talos/libs/aegra-api/src/aegra_api/api/store.py, talos/libs/aegra-api/src/aegra_api/api/mcp.py, talos/libs/aegra-api/src/aegra_api/api/knowledge_bases.py, talos/libs/aegra-api/src/aegra_api/core/auth_middleware.py, talos/libs/aegra-api/src/aegra_api/core/auth_deps.py, talos/libs/aegra-api/src/aegra_api/core/tenant_router.py, talos/libs/aegra-api/src/aegra_api/core/control_plane_db.py, talos/libs/aegra-api/src/aegra_api/core/redis_manager.py, talos/libs/aegra-api/src/aegra_api/core/encryption.py, talos/libs/aegra-api/src/aegra_api/middleware/content_type_fix.py, talos/libs/aegra-api/src/aegra_api/middleware/rate_limiter.py, talos/libs/aegra-api/src/aegra_api/middleware/logger_middleware.py, talos/libs/aegra-api/src/aegra_api/services/broker.py, talos/libs/aegra-api/src/aegra_api/services/redis_broker.py, talos/libs/aegra-api/src/aegra_api/services/executor.py, talos/libs/aegra-api/src/aegra_api/services/local_executor.py, talos/libs/aegra-api/src/aegra_api/services/worker_executor.py, talos/libs/aegra-api/src/aegra_api/services/run_executor.py, talos/libs/aegra-api/src/aegra_api/services/langgraph_service.py, talos/libs/aegra-api/src/aegra_api/services/graph_factory.py, talos/libs/aegra-api/src/aegra_api/services/graph_streaming.py, talos/libs/aegra-api/src/aegra_api/services/streaming_service.py, talos/libs/aegra-api/src/aegra_api/services/event_store.py, talos/libs/aegra-api/alembic/env.py |
|
|
7
|
+
| [TOGa IQ Chat Frontend (Next.js) — streaming stack & Aegra wiring](features/chat-frontend.md) | The repo **`agilantsolutions/talos`** is the TOGa IQ **chat front-end** — the browser client that talks to the Aegra Agent-Protocol backend. | talos/package.json, talos/next.config.mjs, talos/src/providers/Stream.tsx, talos/src/modules/chat/viewmodel/use-chat-viewmodel.tsx, talos/src/lib/api-url.ts, talos/src/lib/env.ts, talos/src/lib/api/client.ts, talos/src/app/api/[..._path]/route.ts, talos/.env.development, talos/.env.beta, talos/.env.gamma, talos/.env.production |
|
|
7
8
|
| [Deployment — Docker, Compose, Entrypoint, External PG/Redis](features/deployment.md) | TOGa IQ ships as a **single container** (`aegra` service) wrapping the `aegra-api` FastAPI server. | talos/docker-compose.yml, talos/deployments/docker/Dockerfile, talos/deployments/docker/entrypoint.sh |
|
|
8
9
|
| [MCP Servers — clickup-mcp and toga-db-mcp](features/mcp-servers.md) | Two internal **FastMCP** servers exposed over **HTTP** with API-key auth and PM2 process management: - **`clickup-mcp`** — ClickUp workspace surface (spaces / f | talos/mcp-servers/clickup-mcp/src, talos/mcp-servers/clickup-mcp/ecosystem.config.js, talos/mcp-servers/clickup-mcp/ecosystem.dev.config.js, talos/mcp-servers/clickup-mcp/pyproject.toml, talos/mcp-servers/clickup-mcp/.env.example, talos/mcp-servers/toga-db-mcp/src, talos/mcp-servers/toga-db-mcp/clusters.yaml, talos/mcp-servers/toga-db-mcp/ecosystem.config.js, talos/mcp-servers/toga-db-mcp/pyproject.toml, talos/mcp-servers/toga-db-mcp/.env.example |
|
|
9
10
|
| [Observability — Langfuse, OTEL, Prometheus, OneUptime](features/observability.md) | TOGa IQ uses **two complementary tracing planes** plus optional Prometheus metrics and external uptime monitoring: - **Langfuse (native v3 SDK)** — LLM-shaped t | talos/libs/aegra-api/src/aegra_api/observability/__init__.py, talos/libs/aegra-api/src/aegra_api/observability/setup.py, talos/libs/aegra-api/src/aegra_api/observability/base.py, talos/libs/aegra-api/src/aegra_api/observability/langfuse_provider.py, talos/libs/aegra-api/src/aegra_api/observability/langfuse_client.py, talos/libs/aegra-api/src/aegra_api/observability/otel.py, talos/libs/aegra-api/src/aegra_api/observability/metrics.py, talos/libs/aegra-api/src/aegra_api/observability/span_enrichment.py, talos/libs/aegra-api/src/aegra_api/observability/targets |
|
|
@@ -45,6 +45,22 @@ dependency** — it is a self-contained Python product that consumes TOGa data
|
|
|
45
45
|
read-only via the `toga-db-mcp` server. It lives under `2.0/apps/talos/` because
|
|
46
46
|
its consumers are 2.0-era TOGa apps (TOGa Hub, TOGa View) and TogaHub auth.
|
|
47
47
|
|
|
48
|
+
> **Repo scope — this doc covers the BACKEND ONLY, which is not in the `talos` repo.**
|
|
49
|
+
> Verified 2026-08-28 (`git ls-tree` across all 16 `origin/*` refs): no branch of
|
|
50
|
+
> `agilantsolutions/talos` contains `libs/aegra-api`, `agents/`, `mcp-servers/`, or
|
|
51
|
+
> `docker-compose.yml` — every branch is the chat **front-end**, documented in
|
|
52
|
+
> [features/chat-frontend.md](features/chat-frontend.md). The Aegra backend described below lives
|
|
53
|
+
> in a separate, currently unregistered repo, so the `files:` paths above do not resolve.
|
|
54
|
+
|
|
55
|
+
**Critical rules:** Clients speak the **Agent Protocol** — never invent a custom wire shape; SDK and
|
|
56
|
+
LangGraph-Studio compatibility depend on it. **Tenant provisioning is explicit with no auto-create**,
|
|
57
|
+
so an unregistered `org_id` gets a clean **403 on every call even with a valid token** — provision
|
|
58
|
+
the tenant before debugging the caller. `AUTH_TYPE` ∈ {`togahub`, `entra`, `noop`} and **`noop` must
|
|
59
|
+
never ship to prod**. `toga-db-mcp` is the **only** sanctioned path into TOGa MySQL (read-only,
|
|
60
|
+
`LIMIT` 1–1000, enforced inside the MCP rather than the agent), and the Fernet `MCP_ENCRYPTION_KEY`
|
|
61
|
+
is a single **global** key shared by every tenant — not per-tenant — so treat rotating it as a
|
|
62
|
+
cross-tenant event.
|
|
63
|
+
|
|
48
64
|
## Top-level layout
|
|
49
65
|
|
|
50
66
|
```
|
|
@@ -194,4 +210,11 @@ Always use Langfuse's cache-corrected cost (it undercounts cache badly).
|
|
|
194
210
|
inside the MCP, not the agent.
|
|
195
211
|
|
|
196
212
|
## Change history
|
|
213
|
+
- 2026-08-28 — Scoped this doc to the Aegra backend and recorded that it is **not** the `talos` repo
|
|
214
|
+
(verified across all 16 origin refs); the `talos` repo is the chat front-end, now documented in
|
|
215
|
+
`features/chat-frontend.md`. Dropped `"language": "python"` from the `talos` `registry.json` entry
|
|
216
|
+
so the front-end repo stops loading `2.0/standards/python.md`. **Deferred half — once the Aegra
|
|
217
|
+
backend repo's name is known, register it and set `"language": "python"` on THAT entry**; until
|
|
218
|
+
then no repo loads the Python standard. Added a backend-scoped `Critical rules:` line to the
|
|
219
|
+
Summary (the pre-existing one under Talos Pricing Platform covers pricing only). (apeterson)
|
|
197
220
|
- 2026-06-16 — Initial architecture doc for talos / TOGa IQ under 2.0/apps. (akhokhani)
|
|
@@ -0,0 +1,172 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: TOGa IQ Chat Frontend (Next.js) — streaming stack & Aegra wiring
|
|
3
|
+
framework: "2.0"
|
|
4
|
+
repo: talos
|
|
5
|
+
project: TOGa IQ
|
|
6
|
+
client: shared
|
|
7
|
+
type: feature
|
|
8
|
+
status: active
|
|
9
|
+
updated: 2026-08-28
|
|
10
|
+
owners: [apeterson]
|
|
11
|
+
files:
|
|
12
|
+
- talos/package.json
|
|
13
|
+
- talos/next.config.mjs
|
|
14
|
+
- talos/src/providers/Stream.tsx
|
|
15
|
+
- talos/src/modules/chat/viewmodel/use-chat-viewmodel.tsx
|
|
16
|
+
- talos/src/lib/api-url.ts
|
|
17
|
+
- talos/src/lib/env.ts
|
|
18
|
+
- talos/src/lib/api/client.ts
|
|
19
|
+
- talos/src/app/api/[..._path]/route.ts
|
|
20
|
+
- talos/.env.development
|
|
21
|
+
- talos/.env.beta
|
|
22
|
+
- talos/.env.gamma
|
|
23
|
+
- talos/.env.production
|
|
24
|
+
related:
|
|
25
|
+
- ../architecture.md
|
|
26
|
+
- aegra-api.md
|
|
27
|
+
- ../../toga-blox/features/talos-assistant.md
|
|
28
|
+
- ../../toga25-supply/features/talos-integration.md
|
|
29
|
+
---
|
|
30
|
+
|
|
31
|
+
## Summary
|
|
32
|
+
|
|
33
|
+
The repo **`agilantsolutions/talos`** is the TOGa IQ **chat front-end** — the browser client that
|
|
34
|
+
talks to the Aegra Agent-Protocol backend. It is the reference implementation of **token-streaming
|
|
35
|
+
chat** in-house, so it is what any other app (toga25-supply, toga2-desk, toga2-commerce) should copy
|
|
36
|
+
when wiring a real Talos backend.
|
|
37
|
+
|
|
38
|
+
This doc exists because the front-end half of TOGa IQ was previously undocumented: anyone told
|
|
39
|
+
*"look at how talos does streaming"* landed on
|
|
40
|
+
[`../architecture.md`](../architecture.md), which describes only the Python/FastAPI backend.
|
|
41
|
+
|
|
42
|
+
## Repo identity — read this before trusting `architecture.md`
|
|
43
|
+
|
|
44
|
+
**Verified 2026-08-28 (`git ls-tree` across all 16 `origin/*` refs): no branch of
|
|
45
|
+
`agilantsolutions/talos` contains `libs/aegra-api`, `agents/`, `mcp-servers/`, or
|
|
46
|
+
`docker-compose.yml`.** Every branch is a front-end app. The Python/Aegra backend that
|
|
47
|
+
`2.0/apps/talos/architecture.md` documents (and lists under `files: talos/libs/aegra-api/…`) lives
|
|
48
|
+
in a **different repo** that is not yet in `registry.json`.
|
|
49
|
+
|
|
50
|
+
Two live consequences until that is reconciled:
|
|
51
|
+
|
|
52
|
+
- `registry.json` marks `talos` as `"language": "python"`, so `/kickoff` loads
|
|
53
|
+
`2.0/standards/python.md` for a **TypeScript/Next.js** repo — and does *not* load
|
|
54
|
+
`2.0/standards/frontend.md`. Pull the front-end standard in manually when working here.
|
|
55
|
+
- The `files:` list on the architecture doc does not resolve against this repo, so file-based
|
|
56
|
+
knowledge search (`search --file=`) cannot bridge from Aegra code to that doc.
|
|
57
|
+
|
|
58
|
+
## Two generations of the app live side by side
|
|
59
|
+
|
|
60
|
+
| Branches | Stack | Build / deploy |
|
|
61
|
+
|---|---|---|
|
|
62
|
+
| `_production` (and older: `ai-model-cache-fix`, `previous-version`, `talos-golive`, `version2-api`) | **Vite** SPA (`vite.config.js`, `index.html`) | `buildspec.yml` + `.platform/` → Elastic Beanstalk |
|
|
63
|
+
| `_beta`, `_next`, `TRUE-77539`, `TRUE-77871` | **Next.js 15** App Router, React 19 | `amplify.yml` → Amplify |
|
|
64
|
+
|
|
65
|
+
So the Next.js rewrite described below is **not yet in `_production`** — production still serves the
|
|
66
|
+
older Vite SPA. Check which branch you are reading before drawing conclusions.
|
|
67
|
+
|
|
68
|
+
## The streaming stack
|
|
69
|
+
|
|
70
|
+
- **`useStream` from `@langchain/langgraph-sdk/react`** — *not* `@langchain/react`. This is the
|
|
71
|
+
LangGraph SDK's own React binding.
|
|
72
|
+
- Wrapped in an **app-owned** context (`StreamProvider` → `StreamSession` →
|
|
73
|
+
`StreamContext.Provider`, consumed via `useStreamContext()`), not the library's provider. That is
|
|
74
|
+
what lets app-specific auth/tenant/thread values be injected.
|
|
75
|
+
- Generative-UI messages come from a second subpath, `@langchain/langgraph-sdk/react-ui`
|
|
76
|
+
(`uiMessageReducer`, `isUIMessage`, `isRemoveUIMessage`), folded into stream state via
|
|
77
|
+
`onCustomEvent`.
|
|
78
|
+
|
|
79
|
+
Versions (`package.json`): `@langchain/core ^1.0.2`, `@langchain/langgraph ^1.0.1`,
|
|
80
|
+
`@langchain/langgraph-sdk ^1.0.0`, `langgraph-nextjs-api-passthrough ^0.0.4`, `next ^15.4.10`,
|
|
81
|
+
`react ^19.0.0`, `nuqs ^2.4.1`.
|
|
82
|
+
|
|
83
|
+
The `useStream` config actually used (`src/providers/Stream.tsx`):
|
|
84
|
+
|
|
85
|
+
```ts
|
|
86
|
+
useTypedStream({
|
|
87
|
+
apiUrl, // absolute Aegra host — see below
|
|
88
|
+
apiKey: apiKey ?? undefined, // X-Api-Key (LangSmith-style), separate from the bearer
|
|
89
|
+
assistantId, // URL query state, defaults to env.assistantId ("agent")
|
|
90
|
+
threadId: threadId ?? null, // URL query state via nuqs
|
|
91
|
+
fetchStateHistory: true,
|
|
92
|
+
defaultHeaders: bearerToken ? { Authorization: `Bearer ${bearerToken}` } : undefined,
|
|
93
|
+
onCustomEvent: …, // reduces UIMessage events into state.ui
|
|
94
|
+
onThreadId: …, // writes the new threadId to the URL, then refetches the thread list
|
|
95
|
+
});
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
Two independent credentials ride along: an **API key** (`apiKey`) and a **user bearer token**
|
|
99
|
+
(`defaultHeaders`). The bearer is an **api2** JWT — `NEXT_PUBLIC_AUTH_API_URL` is
|
|
100
|
+
`https://api.togahub.com/v2` (prod/gamma) or `https://api.beta.togahub.com/v2` (dev/beta) — issued
|
|
101
|
+
for the **TogaHub** client, matching Aegra's `AUTH_TYPE=togahub`.
|
|
102
|
+
|
|
103
|
+
**Readiness probe:** `checkGraphStatus()` `fetch`es `${apiUrl}/info` on mount and toasts on failure.
|
|
104
|
+
A cheap, copyable way to derive a real "connected / not connected" state.
|
|
105
|
+
|
|
106
|
+
**`threadId` and `assistantId` live in the URL** via `useQueryState` from **`nuqs`**, which is
|
|
107
|
+
Next-only. A Vite/react-router app must use `useSearchParams` instead — do not lift this part
|
|
108
|
+
verbatim.
|
|
109
|
+
|
|
110
|
+
## The browser calls Aegra DIRECTLY; the Next.js proxy route is vestigial
|
|
111
|
+
|
|
112
|
+
This is the load-bearing finding for anyone reusing the pattern.
|
|
113
|
+
|
|
114
|
+
`getApiUrl()` returns `env.apiUrl` = `NEXT_PUBLIC_API_URL` = an **absolute external host**
|
|
115
|
+
(`https://api.beta.togaiq.com`), and that same absolute URL is handed to `useStream({ apiUrl })`.
|
|
116
|
+
Nothing routes through `src/app/api/[..._path]/route.ts`. The passthrough route is dead code:
|
|
117
|
+
|
|
118
|
+
- It reads a **server-only** `LANGGRAPH_API_URL`, which **`.env.beta` does not define at all** — the
|
|
119
|
+
route would 500 there.
|
|
120
|
+
- `env.ts`'s own doc comment on `apiUrl` still says *"proxied via /api"*. **That comment is stale**
|
|
121
|
+
and contradicts the value it annotates.
|
|
122
|
+
|
|
123
|
+
**Consequence 1 (enabling):** Aegra already serves CORS to a browser on a different origin.
|
|
124
|
+
A **static SPA with no server route handler** — e.g. toga25-supply on Amplify — *can* call Aegra
|
|
125
|
+
directly. An earlier assumption that a server-side proxy was structurally required is **wrong**.
|
|
126
|
+
|
|
127
|
+
**Consequence 2 (constraining):** the dead proxy pins the request origin:
|
|
128
|
+
|
|
129
|
+
```ts
|
|
130
|
+
// The backend validates the requesting domain against its Domains table.
|
|
131
|
+
// Always send Origin: http://talos so the backend recognizes us.
|
|
132
|
+
headers["Origin"] = "http://talos";
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
`Origin` is a **forbidden header** — a browser cannot set it from JS. So if that `Domains`-table
|
|
136
|
+
gate is real, a new consumer origin needs **its own `Domains` row**; it cannot spoof its way in the
|
|
137
|
+
way this server-side route did.
|
|
138
|
+
|
|
139
|
+
## Environment config — every environment points at BETA
|
|
140
|
+
|
|
141
|
+
Verified 2026-08-28:
|
|
142
|
+
|
|
143
|
+
| File | `NEXT_PUBLIC_API_URL` | `LANGGRAPH_API_URL` (server-only, unused) |
|
|
144
|
+
|---|---|---|
|
|
145
|
+
| `.env.development` | `api.beta.togaiq.com` | `api.beta.togaiq.com` |
|
|
146
|
+
| `.env.beta` | `api.beta.togaiq.com` | **not defined** |
|
|
147
|
+
| `.env.gamma` | `api.beta.togaiq.com` | not defined |
|
|
148
|
+
| `.env.production` | `api.beta.togaiq.com` | `api.beta.togaiq.com` — commented *"Currently not live with new version of ACL"* |
|
|
149
|
+
|
|
150
|
+
**No environment in this repo is configured against a production Aegra host.** Treat "is there a
|
|
151
|
+
live prod Aegra?" as an open question to confirm with the backend owners (`akhokhani`, `jcardinal`)
|
|
152
|
+
rather than an assumption.
|
|
153
|
+
|
|
154
|
+
## Gotchas
|
|
155
|
+
|
|
156
|
+
- **`@langchain/react` is a different package** from what this app uses. It is a newer wrapper
|
|
157
|
+
(v1.0.33) over the same `@langchain/langgraph-sdk`, and it adds a `@langchain/core ^1.1.48` peer
|
|
158
|
+
plus media hooks/players. Do not assume talos validates it — talos uses the SDK's own
|
|
159
|
+
`/react` subpath.
|
|
160
|
+
- **Do not copy `nuqs`** into a non-Next app (see above).
|
|
161
|
+
- **`env.ts` is the only sanctioned env read** ("never access `process.env` directly elsewhere") and
|
|
162
|
+
`validateEnv()` throws at startup on missing vars. Keep new env reads inside it — but note its
|
|
163
|
+
comments have already drifted from its values once.
|
|
164
|
+
|
|
165
|
+
## Change history
|
|
166
|
+
- 2026-08-28 — Documented the previously-undocumented TOGa IQ chat front-end: `useStream` from
|
|
167
|
+
`@langchain/langgraph-sdk/react` wrapped in an app-owned context, the browser's **direct
|
|
168
|
+
cross-origin call to Aegra** (proving a static SPA needs no server proxy) and the vestigial
|
|
169
|
+
Next.js passthrough route with its unspoofable pinned `Origin`, the Vite `_production` vs Next.js
|
|
170
|
+
`_beta`/`_next` split, all-environments-point-at-beta env config, and the finding that no `talos`
|
|
171
|
+
branch contains the Python backend that `architecture.md` describes. Research only — no code
|
|
172
|
+
changed. (apeterson)
|
|
@@ -25,6 +25,7 @@ related:
|
|
|
25
25
|
- ../workflows/local-link-into-a-consumer-app.md
|
|
26
26
|
- ../../toga25-supply/features/talos-integration.md
|
|
27
27
|
- ../../ai-bdr/features/landing-chat-drawer.md
|
|
28
|
+
- ../../talos/features/chat-frontend.md
|
|
28
29
|
---
|
|
29
30
|
|
|
30
31
|
## Summary
|
|
@@ -140,8 +141,52 @@ Open questions were handed to the UX/UI team.
|
|
|
140
141
|
same 40 fail with the Talos changes stashed. Do not read them as a regression from this work.
|
|
141
142
|
- The drawer variant itself was **uncommitted working tree** at capture (2026-08-27).
|
|
142
143
|
|
|
144
|
+
## The contract cannot express streaming — a real backend REQUIRES a contract change
|
|
145
|
+
|
|
146
|
+
Reviewed 2026-08-28 against `TRUE-80692`. The 08-27 drawer-variant work widened `TalosResponse`
|
|
147
|
+
(added `message`, `suggestions`, `cta`; made `paras` optional), but **every one of those is still a
|
|
148
|
+
field on a single resolved value** — the contract remains **one-shot**:
|
|
149
|
+
|
|
150
|
+
```ts
|
|
151
|
+
onSend: (text: string) => Promise<TalosResponse> // resolves ONCE, with the finished answer
|
|
152
|
+
```
|
|
153
|
+
|
|
154
|
+
`TalosMessageModel` / `TalosConversation` are documented "in-memory only". There is **no token
|
|
155
|
+
delta, no `threadId`, no tool-call surface, and no interrupt representation** anywhere in
|
|
156
|
+
`types.ts`.
|
|
157
|
+
|
|
158
|
+
**So token streaming is NOT an adapter-implementation detail.** No adapter, however clever, can
|
|
159
|
+
stream through a `Promise<TalosResponse>` — the promise resolves once. Wiring the live
|
|
160
|
+
Aegra/LangGraph backend (which streams tokens, emits tool calls, and surfaces HITL interrupts —
|
|
161
|
+
see [TOGa IQ chat frontend](../../talos/features/chat-frontend.md)) means changing this component's
|
|
162
|
+
props. That was accepted deliberately rather than shipping a non-streaming v1.
|
|
163
|
+
|
|
164
|
+
**Agreed direction (DECIDED 2026-08-28, NOT yet implemented):** make `TalosPanel` **controlled** —
|
|
165
|
+
the host passes `messages` / `isStreaming` / `onSubmit` / `onStop` — plus an optional
|
|
166
|
+
`connection?: "connecting" | "ready" | "error"` defaulting to `"ready"`.
|
|
167
|
+
|
|
168
|
+
Two properties make this safe for a published package:
|
|
169
|
+
|
|
170
|
+
- **Non-breaking.** `connection` defaults to `"ready"`, so existing consumers and
|
|
171
|
+
`createTalosStubAdapter()` keep working untouched.
|
|
172
|
+
- **Consistent with what is already here.** `open` / `expanded` / `width` are *already* controlled
|
|
173
|
+
props owned by the host; moving `messages` to the host is the same pattern, not a new one.
|
|
174
|
+
|
|
175
|
+
The panel also needs real **"connecting" / "failed to connect"** states: today a thrown `onSend` is
|
|
176
|
+
its *only* failure path, which cannot represent "we never got a token to begin with."
|
|
177
|
+
|
|
178
|
+
**The stream itself does NOT belong in blox.** Per
|
|
179
|
+
[frontend.md §21](../../../standards/frontend.md) (promote only after a 2nd consumer validates the
|
|
180
|
+
contract — Talos streaming is n=1) and **§13(b)** (blox peers do not hoist, so a `@langchain/core`
|
|
181
|
+
peer would force all five consumer apps to install it, and a barrel import can surface that as a
|
|
182
|
+
Rollup *"failed to resolve import"* for a component they never referenced). blox keeps
|
|
183
|
+
`TalosLauncher` / `TalosPanel` as pure UI; the LangChain client stays app-local. blox's unbundled
|
|
184
|
+
`tsc` + `fix-esm-imports` build is also fragile against a dep with subpath exports.
|
|
185
|
+
|
|
143
186
|
## Gotchas
|
|
144
187
|
|
|
188
|
+
- **`onSend` cannot stream** — see the section above. A live streaming backend needs the
|
|
189
|
+
`TRUE-80692` contract change, not just an adapter.
|
|
145
190
|
- The real chat backend is a later blox ticket — `onSend` is stubbed here; do not assume live
|
|
146
191
|
responses from blox alone. **A working implementation exists outside blox and is now wired to
|
|
147
192
|
this component:** `bdr/src/lib/talosChat.ts` implements the same
|
|
@@ -167,6 +212,13 @@ Open questions were handed to the UX/UI team.
|
|
|
167
212
|
currently consumes it via a local symlink.
|
|
168
213
|
|
|
169
214
|
## Change history
|
|
215
|
+
- 2026-08-28 — Review only (no blox code change): established that the `onSend(text) =>
|
|
216
|
+
Promise<TalosResponse>` contract **cannot express token streaming** — the 08-27 `message`/
|
|
217
|
+
`suggestions`/`cta` additions widen the *payload* but it still resolves once — so a live
|
|
218
|
+
streaming backend requires a `TRUE-80692` props change. Decided on a **controlled** `TalosPanel`
|
|
219
|
+
(`messages`/`isStreaming`/`onSubmit`/`onStop`) plus an optional `connection` prop defaulting to
|
|
220
|
+
`"ready"` so the change stays non-breaking. Also decided the LangChain client stays **app-local,
|
|
221
|
+
never a blox peer** (frontend.md §13b/§21). (apeterson)
|
|
170
222
|
- 2026-08-27 — BUILT the **`drawer` variant** on `TRUE-80692` (uncommitted at capture): scrim +
|
|
171
223
|
slide-in aside that stays mounted while closed, title bar, greeting-as-first-bubble
|
|
172
224
|
(`welcome`), quick-reply chips, single-line composer, host `footer` slot, and a `features`
|
|
@@ -6,7 +6,7 @@ project: TOGa 2.5 Supply
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-08-
|
|
9
|
+
updated: 2026-08-28
|
|
10
10
|
owners: [apeterson]
|
|
11
11
|
files:
|
|
12
12
|
- toga25-supply/src/layout/AppLayout/AppLayout.tsx
|
|
@@ -16,6 +16,8 @@ files:
|
|
|
16
16
|
- toga25-supply/src/assets/talos-owl.png
|
|
17
17
|
related:
|
|
18
18
|
- ../../toga-blox/features/talos-assistant.md
|
|
19
|
+
- ../../talos/features/chat-frontend.md
|
|
20
|
+
- ../../../standards/frontend.md
|
|
19
21
|
---
|
|
20
22
|
|
|
21
23
|
## Summary
|
|
@@ -29,7 +31,8 @@ state and injects the adapter; the launcher renders in the header, the panel at
|
|
|
29
31
|
|
|
30
32
|
- **AppLayout owns the state** — `open` / `expanded` / `width` / `dockWidth` — and injects the
|
|
31
33
|
adapter:
|
|
32
|
-
- `onSend`: stub
|
|
34
|
+
- `onSend`: **still the stub.** The live streaming backend is designed but unbuilt — see
|
|
35
|
+
*Real streaming backend* below.
|
|
33
36
|
- `onOpenRecord(id)` → navigate `/sales-orders?sales-orders=<id>`.
|
|
34
37
|
- `onAction("approvals")` → navigate `/sales-orders`.
|
|
35
38
|
- `onDockWidthChange(px)` → reserve layout width for the docked panel.
|
|
@@ -48,5 +51,103 @@ header's right edge sitting **under** the panel, hiding the launcher + avatar. *
|
|
|
48
51
|
the reserved dock width (`marginRight` = the `onDockWidthChange` value) to the **outer app
|
|
49
52
|
wrapper** so the header shifts along with the content.
|
|
50
53
|
|
|
54
|
+
## Real streaming backend — design (DECIDED 2026-08-28, NOT yet implemented)
|
|
55
|
+
|
|
56
|
+
`onSend` is still the stub. The plan below replaces it with the live Aegra/LangGraph backend,
|
|
57
|
+
copying the proven approach in [TOGa IQ's chat front-end](../../talos/features/chat-frontend.md).
|
|
58
|
+
No code has been written yet.
|
|
59
|
+
|
|
60
|
+
### Package choice: `@langchain/langgraph-sdk/react`, app-local
|
|
61
|
+
|
|
62
|
+
- **Use `useStream` from `@langchain/langgraph-sdk/react`** — the same import talos uses — **not
|
|
63
|
+
`@langchain/react`.** Reasons: one fewer peer (`@langchain/react` additionally needs
|
|
64
|
+
`@langchain/core ^1.1.48`); the path is already proven in-house so
|
|
65
|
+
`talos/src/providers/Stream.tsx` can be lifted nearly verbatim; and `@langchain/react`'s extra
|
|
66
|
+
surface is largely **media hooks/players** that blox has no components to render.
|
|
67
|
+
- **Take only `useStream` for v1.** Add `useToolCalls` when tool activity is actually shown, and the
|
|
68
|
+
headless interrupt helpers when doing human-in-the-loop — *that* is where `@langchain/react` earns
|
|
69
|
+
its keep, so revisit it then, not now.
|
|
70
|
+
- **Write our own context provider** (as talos does) rather than the library's `StreamProvider` —
|
|
71
|
+
supply-specific token / tenant / `threadId` / Aegra URL must be injected.
|
|
72
|
+
- **It stays in supply, never in blox.** Per [frontend.md §21](../../../standards/frontend.md) this
|
|
73
|
+
is n=1, and per §13(b) a blox peer would force `@langchain/core` on all five consumer apps. blox
|
|
74
|
+
keeps `TalosPanel` as pure UI.
|
|
75
|
+
|
|
76
|
+
### Why a provider at all — three jobs
|
|
77
|
+
|
|
78
|
+
1. **Singleton.** `useStream` is stateful; two call sites = two independent conversations.
|
|
79
|
+
2. **Survive route changes.** The panel is docked in `AppLayout` while the user navigates. If the
|
|
80
|
+
stream lived inside the panel, any unmount would kill an in-flight response and lose history — so
|
|
81
|
+
it must sit **above the router outlet**.
|
|
82
|
+
3. **Be the app↔blox boundary.** Supply's token/tenant/`threadId`/Aegra URL enter here and LangGraph
|
|
83
|
+
message shapes are translated here, keeping `TalosPanel` app-agnostic and reusable by desk and
|
|
84
|
+
commerce later.
|
|
85
|
+
|
|
86
|
+
### Planned shape
|
|
87
|
+
|
|
88
|
+
```
|
|
89
|
+
src/api/talos.ts — sole VITE_TALOS_API read chokepoint + token exchange
|
|
90
|
+
src/providers/TalosStreamProvider.tsx — the single useStream call, exposed via context
|
|
91
|
+
src/layout/AppLayout/viewModel/useTalosViewModel.ts — flat shape for the panel
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
- **The mapper is where the estimate lives.** LangGraph `Message[]` → blox `TalosMessageModel[]` do
|
|
95
|
+
**not** line up: `type: "ai" | "human"` vs `role: "talos" | "user"`; streamed content blocks vs a
|
|
96
|
+
finished `paras: string[]`; and tool calls and interrupts have **no representation at all** in
|
|
97
|
+
blox's model. This is the same gap that forces the `TRUE-80692` contract change.
|
|
98
|
+
- **`threadId` in URL search params via react-router 7 `useSearchParams`** — **not `nuqs`**, which is
|
|
99
|
+
Next-only. Do not copy that part of talos's provider (§5, URL as source of truth).
|
|
100
|
+
- **Token held in React Query** per [frontend.md §3](../../../standards/frontend.md) — fetched server
|
|
101
|
+
state, never Zustand.
|
|
102
|
+
- **Tear down token + `threadId` + conversation on logout and tenant switch** per §18.
|
|
103
|
+
|
|
104
|
+
### Async auth gating — and the anti-pattern
|
|
105
|
+
|
|
106
|
+
`useStream` takes the token as a **value** (`defaultHeaders`), not a promise, so you **cannot await
|
|
107
|
+
it inside the hook**. That single constraint decides the design:
|
|
108
|
+
|
|
109
|
+
- Awaiting the token inside `onSend` (via `queryClient.ensureQueryData`, reusing the panel's existing
|
|
110
|
+
thinking/error states, lazily fetching on first send) works **only** for the non-streaming one-shot
|
|
111
|
+
contract.
|
|
112
|
+
- For **streaming**, the gate must sit **above** the hook — do not mount the stream until the token
|
|
113
|
+
query resolves. This is exactly why the panel's `connection` prop stops being optional polish and
|
|
114
|
+
becomes required.
|
|
115
|
+
|
|
116
|
+
**Anti-pattern — do not use `useSuspenseQuery`/Suspense for this gate.** It suspends the subtree, so
|
|
117
|
+
the panel **unmounts and remounts** — flashing and dropping whatever the user typed in the composer —
|
|
118
|
+
and in a shared library the thrown promise surfaces in host apps that never opted into a boundary.
|
|
119
|
+
Keep it a plain query.
|
|
120
|
+
|
|
121
|
+
**Open product decision:** mid-stream 401 behavior — silent refresh and resume, vs. surface an error
|
|
122
|
+
and let the user resend. Decide deliberately rather than discovering it in QA.
|
|
123
|
+
|
|
124
|
+
## Aegra-side blockers — backend config, zero front-end code
|
|
125
|
+
|
|
126
|
+
Each of these masquerades as a front-end bug. **Confirm them with the Aegra owners
|
|
127
|
+
(`akhokhani`, `jcardinal`) BEFORE building the provider** — answer (c) especially decides whether the
|
|
128
|
+
auth/loading layer gets built at all.
|
|
129
|
+
|
|
130
|
+
- **(a) CORS allowlist** must include supply's **per-client hostnames** (nychh, compass, quad, …).
|
|
131
|
+
The symptom is a request that never leaves the browser — no response to debug. Note the browser
|
|
132
|
+
calls Aegra **directly** cross-origin (proven by talos), and a static Amplify SPA has no server
|
|
133
|
+
route handler to proxy through, so this is unavoidable. Related: Aegra also validates the
|
|
134
|
+
requesting domain against a `Domains` table, and `Origin` is a forbidden header a browser cannot
|
|
135
|
+
set — so a new origin needs its **own `Domains` row**.
|
|
136
|
+
- **(b) Tenant provisioning is explicit, with no auto-create.** An unregistered org gets a clean
|
|
137
|
+
**403 on every call even with a valid token** (control-plane DB / `tenant_router` LRU — see
|
|
138
|
+
[talos architecture](../../talos/architecture.md)).
|
|
139
|
+
- **(c) Token audience.** Aegra `AUTH_TYPE=togahub` validates api2 JWTs, but supply's token is issued
|
|
140
|
+
for the **supply** client, not TogaHub. Either Aegra accepts it (trivial) or supply needs a
|
|
141
|
+
handoff — api2 already exposes `/auth/delegator`, `/auth/encrypted`, and
|
|
142
|
+
`/auth/encrypted-user-uuid` for exactly this cross-app identity pass.
|
|
143
|
+
|
|
51
144
|
## Change history
|
|
145
|
+
- 2026-08-28 — Design/research only, no code: chose `useStream` from
|
|
146
|
+
`@langchain/langgraph-sdk/react` (not `@langchain/react`), kept **app-local** rather than in blox
|
|
147
|
+
(frontend.md §13b/§21), and designed a route-surviving singleton `TalosStreamProvider` above the
|
|
148
|
+
router outlet plus a LangGraph→blox message mapper. Recorded that the token cannot be awaited
|
|
149
|
+
inside `useStream`, so streaming must gate **above** the hook (and **not** via
|
|
150
|
+
`useSuspenseQuery`, which remounts the panel and drops composer text), and logged three
|
|
151
|
+
Aegra-side blockers (per-client CORS hostnames, explicit tenant provisioning → 403, supply-vs-
|
|
152
|
+
TogaHub token audience) to confirm with the backend owners first. (apeterson)
|
|
52
153
|
- 2026-08-20 — Wired the shared blox Talos assistant into supply for all clients: AppLayout hosts state + injects the adapter (stub `onSend`; `onOpenRecord`/`onAction` navigate to sales orders; `onDockWidthChange` reserves layout width), launcher in a new header `talosLauncher` slot, panel at app root, owl avatar + Plus Jakarta Sans loaded locally. Fixed the docked panel hiding the header launcher by reserving dock width on the outer wrapper (apeterson).
|
|
@@ -6,7 +6,7 @@ project: Worker
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-08-
|
|
9
|
+
updated: 2026-08-28
|
|
10
10
|
owners: ["kyalamarthi"]
|
|
11
11
|
files:
|
|
12
12
|
- worker2/Worker/Netsuite/CustomerAccountTypeBackfill.php
|
|
@@ -17,6 +17,7 @@ files:
|
|
|
17
17
|
related:
|
|
18
18
|
- ./netsuite-supporting-record-webhook-importer.md
|
|
19
19
|
- ../workflows/running-worker2-locally.md
|
|
20
|
+
- ../../../../1.0/apps/worker/features/forecast2-supporting-records-import.md
|
|
20
21
|
- ../../../../1.0/apps/library/features/netsuite-suiteql-api-reference.md
|
|
21
22
|
---
|
|
22
23
|
|
|
@@ -79,6 +80,31 @@ only ongoing writer** — and:
|
|
|
79
80
|
- **Therefore: after any bulk reclassification in NetSuite, re-run this backfill.** It is the only
|
|
80
81
|
thing that will bring `Forecast` back into agreement.
|
|
81
82
|
|
|
83
|
+
## When a backfill is the WRONG tool
|
|
84
|
+
|
|
85
|
+
The test is **"does an existing writer already own this column, over the whole population?"** — not
|
|
86
|
+
"is the data currently wrong?".
|
|
87
|
+
|
|
88
|
+
- **A full, unfiltered nightly pull makes the fix self-healing — ship the code and wait one night.**
|
|
89
|
+
TRUE-81213 (blank Consolidated Business Names) looked exactly like backfill work, and was not:
|
|
90
|
+
`App_Api_Netsuite_Rest::listCustomers()` pulls **every** NetSuite customer with **no date filter**, so
|
|
91
|
+
the corrected cron re-linked **793 rows on its very next pass**. No migration, no worker.
|
|
92
|
+
- **Worse than redundant — a SQL backfill would have been actively wrong there.** The customer→group
|
|
93
|
+
mapping exists **only** in NetSuite's `custentity10`; the database cannot derive it, so the only
|
|
94
|
+
available SQL heuristic (name matching) would have silently assigned customers to the **wrong**
|
|
95
|
+
consolidated group.
|
|
96
|
+
- **`accountType` is the contrasting case that genuinely needs this worker:** the 1.0 model does not
|
|
97
|
+
declare the column, so **no** nightly writer exists — there is nothing to self-heal.
|
|
98
|
+
|
|
99
|
+
### Corollary: a populated column does NOT prove the webhook ran
|
|
100
|
+
|
|
101
|
+
Do not infer webhook coverage from data being present. Prod `Core.WorkerJobs` holds only
|
|
102
|
+
`Netsuite/Customer/put` **30** + `post` **5** = **35** customer jobs *ever* (all since 2026-08-19),
|
|
103
|
+
while `Netsuite/CustomerAccountTypeBackfill/Backfill` ran **twice** on 2026-08-19 — and **this backfill
|
|
104
|
+
issues a raw `UPDATE … SET accountType`, touching nothing else**. That is precisely why 253 customers
|
|
105
|
+
carried a correct `accountType` alongside a stale `consolidatedCustomerId`, which is the misreading
|
|
106
|
+
that sent TRUE-81213's investigation down the wrong path. **Count the jobs; do not infer them.**
|
|
107
|
+
|
|
82
108
|
## Production results (2026-08-19)
|
|
83
109
|
|
|
84
110
|
- **248 of 9,183** NetSuite customers have an Account Type set — **2.7%**. The other 97% are legitimately
|
|
@@ -87,13 +113,14 @@ only ongoing writer** — and:
|
|
|
87
113
|
|
|
88
114
|
## Gotchas / known issues
|
|
89
115
|
|
|
90
|
-
-
|
|
91
|
-
`worker/crons/toga2/forecast2/common_import_sales_from_netsuite.php:256-259`
|
|
92
|
-
`name != … && consolidatedCustomerId != …` where it
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
116
|
+
- **✅ The 1.0 nightly cron was effectively INSERT-ONLY for customers — FIXED 2026-08-28.**
|
|
117
|
+
`worker/crons/toga2/forecast2/common_import_sales_from_netsuite.php:256-259` guarded the update with
|
|
118
|
+
`name != … && consolidatedCustomerId != …` where it meant `||`, so an existing row was rewritten only
|
|
119
|
+
when **both** values changed. TRUE-81213 / commit `b846e97c` changed it to `||`; verified in
|
|
120
|
+
production. The cron now **does** heal a stale `name` or `consolidatedCustomerId`
|
|
121
|
+
(**not** `accountType` — see above, the 1.0 model does not declare it). Historical evidence of the
|
|
122
|
+
bug (prod, 2026-08-19): Forecast held **9,206** customers vs **9,183** in NetSuite; the cron still
|
|
123
|
+
never prunes deleted customers.
|
|
97
124
|
- **`dryRun` is `true` by default** — a run that "did nothing" almost certainly just needed
|
|
98
125
|
`dryRun=false`. Read the returned summary before concluding the data was already correct.
|
|
99
126
|
- **SuiteQL lower-cases every column alias**, so a camelCase alias read back as camelCase is silently
|
|
@@ -106,6 +133,15 @@ only ongoing writer** — and:
|
|
|
106
133
|
|
|
107
134
|
## Change history
|
|
108
135
|
|
|
136
|
+
- 2026-08-28 — Added **"When a backfill is the WRONG tool"** after TRUE-81213 proved the point: a
|
|
137
|
+
blank-Consolidated-Business-Name defect that looked like backfill work was **self-healing**, because
|
|
138
|
+
`listCustomers()` pulls the full population with no date filter (the corrected cron re-linked 793 rows
|
|
139
|
+
in one pass) — and a SQL backfill would have been *actively wrong*, since the customer→group mapping
|
|
140
|
+
lives only in NetSuite's `custentity10` and name-matching would mis-assign groups. Also recorded the
|
|
141
|
+
corollary that **a populated column does not prove the webhook ran** (35 customer webhook jobs ever,
|
|
142
|
+
vs. two runs of this backfill's raw `UPDATE … SET accountType`), and marked the 1.0 cron's
|
|
143
|
+
`&&`-should-be-`||` guard **fixed** (`b846e97c`) — it is now a real backstop for `name` and
|
|
144
|
+
`consolidatedCustomerId`. (kyalamarthi)
|
|
109
145
|
- 2026-08-19 — Created after building `_Worker_Netsuite_CustomerAccountTypeBackfill` (TRUE-80206) on
|
|
110
146
|
the `ItemFulfillableBackfill` pattern: recorded the batching recipe (page 1,000 / SuiteQL `IN` 500 /
|
|
111
147
|
UPDATE 500), the `dryRun`-defaults-to-true and countable-summary-on-every-path contract, and the
|
|
@@ -123,4 +159,6 @@ only ongoing writer** — and:
|
|
|
123
159
|
- [Running worker2 locally against real NetSuite](../workflows/running-worker2-locally.md) — how this was
|
|
124
160
|
tested before it touched production.
|
|
125
161
|
- [NetSuite SuiteQL/REST API Reference](../../../../1.0/apps/library/features/netsuite-suiteql-api-reference.md)
|
|
162
|
+
- [Forecast2 Supporting-Records Nightly Import (1.0 cron)](../../../../1.0/apps/worker/features/forecast2-supporting-records-import.md)
|
|
163
|
+
— the full pull that made TRUE-81213 self-healing, and the orphan-group diagnostic.
|
|
126
164
|
— SuiteQL mechanics, `BUILTIN.DF`, the lower-cased-alias rule, and the custom-field map.
|
|
@@ -6,7 +6,7 @@ project: Worker
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-08-
|
|
9
|
+
updated: 2026-08-28
|
|
10
10
|
owners: ["dfranks", "bala", "kyalamarthi"]
|
|
11
11
|
files:
|
|
12
12
|
- worker2/Worker/Netsuite/Employee.php
|
|
@@ -36,6 +36,7 @@ related:
|
|
|
36
36
|
- ../../_underscore/features/forecast-sale-import.md
|
|
37
37
|
- ../../_underscore/features/netsuite-rest-client.md
|
|
38
38
|
- ../architecture.md
|
|
39
|
+
- ../../../../1.0/apps/worker/features/forecast2-supporting-records-import.md
|
|
39
40
|
---
|
|
40
41
|
|
|
41
42
|
## Summary
|
|
@@ -95,8 +96,9 @@ recipe and its two recurring variants so a new one is a fill-in-the-blanks job,
|
|
|
95
96
|
records, the parent/supervisor backfill). It shares `import_supporting_records.php` with the other
|
|
96
97
|
sections, so **never disable the wrapper.**
|
|
97
98
|
⚠ **But do not assume the backstop covers your new column** — see *the backstop is narrower than it
|
|
98
|
-
looks* under Gotchas. A column the 1.0 model does not declare has **no** nightly writer at all
|
|
99
|
-
|
|
99
|
+
looks* under Gotchas. A column the 1.0 model does not declare has **no** nightly writer at all.
|
|
100
|
+
(Customers' cron *was* effectively insert-only until TRUE-81213 — fixed 2026-08-28.)
|
|
101
|
+
When there is no backstop, ship a
|
|
100
102
|
[backfill worker](./netsuite-supporting-record-backfill-worker.md) with the feature.
|
|
101
103
|
|
|
102
104
|
## Recurring variants
|
|
@@ -119,7 +121,8 @@ recipe and its two recurring variants so a new one is a fill-in-the-blanks job,
|
|
|
119
121
|
`custentity_account_type` (NetSuite Entity Field "Account Type", internal id **10719**) maps
|
|
120
122
|
**1 = Strategic Emerging · 2 = Not Strategic · 3 = Strategic Growth · 4 = Strategic Maintain ·
|
|
121
123
|
5 = Strategic Draining**. Power BI reads `Forecast.Customers.accountType` directly, so a raw id
|
|
122
|
-
renders as `3` on the report. Only **248 of 9,183** customers (2.7%) have the field set at all
|
|
124
|
+
renders as `3` on the report. Only **248 of 9,183** customers (2.7%) have the field set at all
|
|
125
|
+
(**253 of 9,213** as of 2026-08-28) —
|
|
123
126
|
mostly-blank is correct, not a symptom. If you are unsure of a custom field's type, let the
|
|
124
127
|
[backfill worker's dry run probe it](./netsuite-supporting-record-backfill-worker.md).
|
|
125
128
|
- **A column with TWO writers must agree byte-for-byte on format** (Customers `name`). SOAP returned a
|
|
@@ -141,10 +144,17 @@ recipe and its two recurring variants so a new one is a fill-in-the-blanks job,
|
|
|
141
144
|
aliases, or read the lower-cased key.** This has now caused two separate bugs:
|
|
142
145
|
- it broke the first cut of the account-type backfill — every customer was misreported as "blank in
|
|
143
146
|
NetSuite" **while the job reported success**; only testing caught it;
|
|
144
|
-
- **still open
|
|
145
|
-
`
|
|
146
|
-
|
|
147
|
-
|
|
147
|
+
- **still open in `_production`, but measured harmless (2026-08-28):**
|
|
148
|
+
`worker2/Worker/Netsuite/Customer.php` (~line 147) selects
|
|
149
|
+
`BUILTIN.DF(custentity10) AS consolidatedName` and reads `$record->consolidatedName`, so the
|
|
150
|
+
webhook never sets a consolidated-customer **name**. Production damage is **zero — 0 NULL/empty
|
|
151
|
+
names across all 2,742 `Forecast.ConsolidatedCustomers` rows** — because `custentity10` is *also*
|
|
152
|
+
selected **unaliased**, so the FK link still resolves, and the nightly 1.0 cron creates every parent
|
|
153
|
+
first with a correct name via its own lowercase `consol_name` alias. Only a brand-new parent seen by
|
|
154
|
+
the webhook *before* the cron could lose its display name. So: still worth fixing, but it is **not**
|
|
155
|
+
a live data-loss path and it was **not** the cause of TRUE-81213 (see below).
|
|
156
|
+
⚠ The fix currently exists **only as an uncommitted working-tree edit** on branch `TRUE-80742`,
|
|
157
|
+
which has no commits of its own — it is committed to no branch and will be lost by a checkout.
|
|
148
158
|
- **SuiteQL over the REST client has NO server-side placeholder binding.** `_Component_Api_Netsuite::send`
|
|
149
159
|
passes the query string as-is, so the **`(int)` cast on an already-int id is the injection control** —
|
|
150
160
|
the same codebase-wide idiom used by `lookupId()` and the Item/Opportunity/Employee handlers. A
|
|
@@ -163,14 +173,22 @@ recipe and its two recurring variants so a new one is a fill-in-the-blanks job,
|
|
|
163
173
|
covers columns the **1.0** model declares. `App_Model_Forecast2_Customer`
|
|
164
174
|
(`library/app/model/forecast2/customer.php`) does not declare `accountType`, so the 2 AM
|
|
165
175
|
"Forecast 2.0 — Import supporting records" cron can neither write **nor clobber** it — the webhook is
|
|
166
|
-
the **only** ongoing writer.
|
|
167
|
-
guard (`common_import_sales_from_netsuite.php:256-259`) uses `name != … && consolidatedCustomerId != …`
|
|
168
|
-
where it means `||`, so an existing row is rewritten only when *both* changed and renames never
|
|
169
|
-
propagate (prod: 9,206 Forecast customers vs 9,183 in NetSuite). **Latent bug, not fixed.** And because
|
|
176
|
+
the **only** ongoing writer. And because
|
|
170
177
|
a NetSuite **Mass Update / CSV import without "Run Server SuiteScript"** fires no User Event, a bulk
|
|
171
178
|
reclassification reaches neither writer. Ship a
|
|
172
179
|
[backfill worker](./netsuite-supporting-record-backfill-worker.md) for any column in this position and
|
|
173
180
|
re-run it after every bulk change.
|
|
181
|
+
- **✅ RESOLVED 2026-08-28 — the customer cron is no longer insert-only.** This doc previously recorded
|
|
182
|
+
the `&&`-should-be-`||` change guard at
|
|
183
|
+
`worker/crons/toga2/forecast2/common_import_sales_from_netsuite.php:256-259` as a **latent bug, not
|
|
184
|
+
fixed**. It is **fixed** — TRUE-81213, commit **`b846e97c`**, deployed and **verified in production**.
|
|
185
|
+
The guard rewrote an existing customer only when `name` **AND** `consolidatedCustomerId` had *both*
|
|
186
|
+
changed, so `consolidatedCustomerId` was effectively **write-once** and Power BI showed a blank
|
|
187
|
+
Consolidated Business Name for every customer whose group was assigned after its row existed. First
|
|
188
|
+
corrected run: consolidated links **4,062 → 4,855 (+793)**, blank strategic accounts **33 → 1**.
|
|
189
|
+
The nightly cron therefore now *is* a genuine backstop for `name` and `consolidatedCustomerId`
|
|
190
|
+
(`accountType` remains unbacked — that part of the bullet above still stands). Full write-up:
|
|
191
|
+
[Forecast2 supporting-records nightly import](../../../../1.0/apps/worker/features/forecast2-supporting-records-import.md).
|
|
174
192
|
- **⚠ Unrelated observation, production, updated 2026-08-19: the item job types are still at 0% success,
|
|
175
193
|
the cause is now known, and the volume is growing.** `Core.WorkerJobs` in prod shows
|
|
176
194
|
`Netsuite/InventoryItem/post|put` at **~1,163 failures / 0 successes**, all with
|
|
@@ -182,6 +200,17 @@ recipe and its two recurring variants so a new one is a fill-in-the-blanks job,
|
|
|
182
200
|
|
|
183
201
|
## Change history
|
|
184
202
|
|
|
203
|
+
- 2026-08-28 — **The legacy customer cron's `&&`-should-be-`||` change guard is FIXED** (TRUE-81213,
|
|
204
|
+
commit `b846e97c`, verified in production) — this doc had it recorded as "latent bug, not fixed".
|
|
205
|
+
It had made `Forecast.Customers.consolidatedCustomerId` **write-once**, blanking Power BI's
|
|
206
|
+
Consolidated Business Name; the corrected cron self-healed **+793** links (4,062 → 4,855) on its first
|
|
207
|
+
run and took blank strategic accounts from **33 → 1**. So the nightly cron *is* now a real backstop for
|
|
208
|
+
`name` + `consolidatedCustomerId` (still not for `accountType`). Also **measured** the still-open
|
|
209
|
+
`consolidatedName` alias-casing bug: **zero** production damage (0 NULL names in 2,742 rows) because
|
|
210
|
+
`custentity10` is selected unaliased and the cron creates parents first — and noted that its fix is
|
|
211
|
+
sitting uncommitted on branch `TRUE-80742`. Deep dive lives in the new
|
|
212
|
+
[1.0 supporting-records import doc](../../../../1.0/apps/worker/features/forecast2-supporting-records-import.md).
|
|
213
|
+
(kyalamarthi)
|
|
185
214
|
- 2026-08-19 — **Customers went live for the first time (TRUE-80206)** and produced four additions.
|
|
186
215
|
(1) **Root cause of the whole ticket:** the `customscript_ue_amq_enqueue` deployment for Customer
|
|
187
216
|
(deployment internal id 9, `customdeploy9`) sat at `Status = TESTING` / `isdeployed = F` while all 15
|
|
@@ -213,6 +242,8 @@ recipe and its two recurring variants so a new one is a fill-in-the-blanks job,
|
|
|
213
242
|
|
|
214
243
|
- [NetSuite Supporting-Record Backfill Worker](./netsuite-supporting-record-backfill-worker.md) — the bulk
|
|
215
244
|
reconciliation counterpart; ship one whenever the legacy cron does not back your column.
|
|
245
|
+
- [Forecast2 Supporting-Records Nightly Import (1.0 cron)](../../../../1.0/apps/worker/features/forecast2-supporting-records-import.md)
|
|
246
|
+
— the legacy backstop this recipe replaces, and the `ConsolidatedCustomers` linkage in detail.
|
|
216
247
|
- [NetSuite → TOGA Opportunity Sync](./netsuite-opportunity-sync.md) — header+children transaction sibling.
|
|
217
248
|
- [NetSuite → Forecast Open-Orders Sync](./netsuite-salesorder-open-orders-sync.md).
|
|
218
249
|
- [Forecast.Sales NetSuite import engine](../../_underscore/features/forecast-sale-import.md) — the
|
package/knowledge/INDEX.md
CHANGED
|
@@ -5,7 +5,7 @@ _Auto-generated by `knowledge.js index`. Do not hand-edit._
|
|
|
5
5
|
## 1.0 framework
|
|
6
6
|
|
|
7
7
|
- **library** (Library) _(framework core)_ — 20 doc(s) → [1.0/apps/library/INDEX.md](1.0/apps/library/INDEX.md)
|
|
8
|
-
- **worker** (Worker) —
|
|
8
|
+
- **worker** (Worker) — 29 doc(s) → [1.0/apps/worker/INDEX.md](1.0/apps/worker/INDEX.md)
|
|
9
9
|
- **dbchanges** (Database Changes) _(framework core)_ — 1 doc(s) → [1.0/apps/dbchanges/INDEX.md](1.0/apps/dbchanges/INDEX.md)
|
|
10
10
|
- **worker1.5** (Worker 1.5) — 0 doc(s) → [1.0/apps/worker1.5/INDEX.md](1.0/apps/worker1.5/INDEX.md)
|
|
11
11
|
- **togadesk** (TOGa Desk) — 13 doc(s) → [1.0/apps/togadesk/INDEX.md](1.0/apps/togadesk/INDEX.md)
|
|
@@ -26,7 +26,7 @@ _Auto-generated by `knowledge.js index`. Do not hand-edit._
|
|
|
26
26
|
- **saml** (SAML SSO Gateway) — 5 doc(s) → [2.0/apps/saml/INDEX.md](2.0/apps/saml/INDEX.md)
|
|
27
27
|
- **toga2-view** (TOGa View Frontend) — 12 doc(s) → [2.0/apps/toga2-view/INDEX.md](2.0/apps/toga2-view/INDEX.md)
|
|
28
28
|
- **toga2-hub** (TOGa Hub) — 2 doc(s) → [2.0/apps/toga2-hub/INDEX.md](2.0/apps/toga2-hub/INDEX.md)
|
|
29
|
-
- **talos** (TOGa IQ) —
|
|
29
|
+
- **talos** (TOGa IQ) — 8 doc(s) → [2.0/apps/talos/INDEX.md](2.0/apps/talos/INDEX.md)
|
|
30
30
|
- **voice-to-voice** (TOGa Voice) — 4 doc(s) → [2.0/apps/voice-to-voice/INDEX.md](2.0/apps/voice-to-voice/INDEX.md)
|
|
31
31
|
- **ai-bdr** (AI-BDR) — 13 doc(s) → [2.0/apps/ai-bdr/INDEX.md](2.0/apps/ai-bdr/INDEX.md)
|
|
32
32
|
- **toga2-commerce** (TOGa Commerce) — 20 doc(s) → [2.0/apps/toga2-commerce/INDEX.md](2.0/apps/toga2-commerce/INDEX.md)
|
|
@@ -40,6 +40,7 @@ _Auto-generated by `knowledge.js index`. Do not hand-edit._
|
|
|
40
40
|
- **websocket** (WebSocket Server) — 2 doc(s) → [standalone/apps/websocket/INDEX.md](standalone/apps/websocket/INDEX.md)
|
|
41
41
|
- **forward** (Forwarder) — 3 doc(s) → [standalone/apps/forward/INDEX.md](standalone/apps/forward/INDEX.md)
|
|
42
42
|
- **claude** (Claude Harness) — 8 doc(s) → [standalone/apps/claude/INDEX.md](standalone/apps/claude/INDEX.md)
|
|
43
|
+
- **powerbi** (Power BI) — 0 doc(s) → [standalone/apps/powerbi/INDEX.md](standalone/apps/powerbi/INDEX.md)
|
|
43
44
|
|
|
44
45
|
## Clients
|
|
45
46
|
|
|
@@ -8,11 +8,13 @@ apps:
|
|
|
8
8
|
- dbchanges2
|
|
9
9
|
- api2
|
|
10
10
|
- togatech
|
|
11
|
+
- worker
|
|
12
|
+
- library
|
|
11
13
|
project: _Underscore
|
|
12
14
|
client: true
|
|
13
15
|
type: profile
|
|
14
16
|
status: active
|
|
15
|
-
updated: 2026-08-
|
|
17
|
+
updated: 2026-08-28
|
|
16
18
|
owners: [jcardinal, kyalamarthi, tcox]
|
|
17
19
|
files: []
|
|
18
20
|
related:
|
package/knowledge/registry.json
CHANGED
|
@@ -145,8 +145,7 @@
|
|
|
145
145
|
"project": "TOGa IQ",
|
|
146
146
|
"framework": "2.0",
|
|
147
147
|
"role": "app",
|
|
148
|
-
"dependsOn": []
|
|
149
|
-
"language": "python"
|
|
148
|
+
"dependsOn": []
|
|
150
149
|
},
|
|
151
150
|
{
|
|
152
151
|
"repo": "test",
|
|
@@ -236,5 +235,12 @@
|
|
|
236
235
|
"framework": "standalone",
|
|
237
236
|
"role": "app",
|
|
238
237
|
"dependsOn": []
|
|
238
|
+
},
|
|
239
|
+
{
|
|
240
|
+
"repo": "powerbi",
|
|
241
|
+
"project": "Power BI",
|
|
242
|
+
"framework": "standalone",
|
|
243
|
+
"role": "app",
|
|
244
|
+
"dependsOn": []
|
|
239
245
|
}
|
|
240
246
|
]
|
|
@@ -0,0 +1,107 @@
|
|
|
1
|
+
---
|
|
2
|
+
type: session
|
|
3
|
+
slug: openorderitems-location-fix
|
|
4
|
+
title: Fix Forecast.OpenOrderItems.locationId NULL on 67% of rows
|
|
5
|
+
author: kyalamarthi
|
|
6
|
+
repos: [worker, worker2, dbchanges2, library, test]
|
|
7
|
+
framework: "both"
|
|
8
|
+
client: shared
|
|
9
|
+
status: active
|
|
10
|
+
created: 2026-08-28
|
|
11
|
+
updated: 2026-08-28
|
|
12
|
+
---
|
|
13
|
+
|
|
14
|
+
# Session: openorderitems-location-fix
|
|
15
|
+
**Date:** 2026-08-28
|
|
16
|
+
**Project/Repo:** worker2 + worker + dbchanges2 (2.0, with 1.0 cron/library findings)
|
|
17
|
+
**Task:** Populate `Forecast.OpenOrderItems.locationId`, which was NULL on 67% of rows, starving Power BI of the warehouse dimension on open orders (TRUE-79162; TRUE-79078 for the reverted `tl.location` select).
|
|
18
|
+
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
## What WORKED
|
|
22
|
+
|
|
23
|
+
- **Measured the gap definitively** via the `toga-db` MCP against prod `Forecast`: **2,883 of 4,308 rows NULL (67%)** across **1,366 of 1,827 orders**, and critically **ZERO orders have a mix** of NULL and populated rows — proof it is writer-attribution, not per-line data. The table is live: the NULL count drifted 2,883 -> 2,844 within an hour as webhooks fired.
|
|
24
|
+
- **Disproved the documented root cause.** The KB said `Forecast.Locations` is empty. It is healthy: 184 rows, 172 with `netsuiteParentLocationId`, 12 roots, **0 rows with a missing parent**. The leaf->root rollup has everything it needs, and `SyncAll` did NOT need to run first.
|
|
25
|
+
- **Pinned the real cause with git.** `locationId` is written only as a side effect of a `salesOrder` webhook re-import, which shipped **2026-06-29** (`worker2` ec66962, `_underscore` b9a4472c) with no backfill. The NULL-rate-by-month gradient matches that cutover exactly: 2026-08 1.2%, 2026-07 16%, 2026-06 77%, 2026-05 and earlier ~90-100%.
|
|
26
|
+
- **Found the second cause:** the live 3 AM cron `periodic_forecast_discrepancy_fix_open_orders.php` (`active: 1`) had **zero occurrences of "location" in 775 lines** and delete+reinserts rows, so anything it recreated came back NULL.
|
|
27
|
+
- **The window split that reshaped the whole fix:** 1,711 rows / 750 orders fall INSIDE the cron's `[today-2y, today+10y]` window; **1,133 rows / 608 orders fall before it** (back to 2019-04-24) and are scanned by nothing.
|
|
28
|
+
- **Built and ran a read-only probe** `test/@dave/probe_ooi_location_gap.php` (verified zero write statements) against live NetSuite: 4,387 importable lines, **100% resolvable, 0 legitimately-NULL, 0 dimension gaps**. Expected values are 6 roots (71 New York 2,506 lines; 2: 1,262; 137: 297; 139: 263; 67: 36; 138: 23).
|
|
29
|
+
- **Verified the line join key at scale:** `OpenOrderItems.lineNumber` == `transactionline.id` — 99 of 99 randomly sampled prod rows plus all 7 multi-root orders matched, **zero join misses**. `linesequencenumber` diverges on 2,061 of 4,387 lines (~47%).
|
|
30
|
+
- **Verified order 235272 end to end by hand** (SalesOrd internalId 6087746, status B): 6 lines all NULL locally, all on NetSuite leaf `4` = "New York", which is itself a root -> `Locations.id 71`. Note `tranid 235272` is reused across three NetSuite records (Invoice 5397885, ItemShip 5418115, SalesOrd 6087746).
|
|
31
|
+
- **Phase 1 (cron fix) implemented:** 5 edits — `location AS locationid` in the line SQL; `Locations` lookup + `$hasLocationCol` / `netsuiteParentLocationId` guards via the existing `forecastColumnExists()`; `resolveOoiLocationId()` copied from `tools/bin/forecast/fixer.php:1328-1342`; `locationId` threaded through both SQL builders; a null-safe `locationId` entry in `$compare`. Lint clean, no PHP 7.4+ tokens introduced.
|
|
32
|
+
- **A stub harness that renders the generated SQL caught a real bug** before any run: the first ODKU clause produced `classificationId = VALUES(classificationId),,` plus a missing separator before `itemId`, which would have broken **every insert the cron makes**. `php -l` did not catch it.
|
|
33
|
+
- **Phase 2 (backfill worker) verified end to end** — `_Worker_Netsuite_OpenOrderLocationBackfill`, against live NetSuite with a local `Forecast` fixture, on order 6087746: dry run reported 6 rows leaf 4 -> 71 and **wrote nothing**; live run updated 6; DB showed all 6 = 71; immediate re-run reported *"nothing to do, 0 rows"*.
|
|
34
|
+
- **Phase 3:** added `ORDER BY id` to `Location.php:89` (offset-paged with no stable sort — latent drop/duplicate bug), and wrote the dbchanges2 `Core/2026-08-27a` migration scheduling `Netsuite/Location/SyncAll` daily 2:30 AM. **Tested the migration twice** against a throwaway local `CronJobs` table: inserts once, guard holds on re-run.
|
|
35
|
+
- **Unblocked local worker2 execution** (see failures below for the traps): app-root-first `include_path`, blanked `[database]` password, created local `ClientLogs` + `Forecast` schemas.
|
|
36
|
+
- **Built `test/@dave/nsq2.php`**, a working read-only SuiteQL runner, and used it to confirm the restored `listSalesOrders()` query shape returns `location: "4"` / `location_name: "New York"` from live NetSuite.
|
|
37
|
+
- **PRs opened** on branch `TRUE-80326` for `worker`, `worker2`, `dbchanges2` — all three committed and pushed.
|
|
38
|
+
- **`/capture` published 7 docs**, commit `59f2cd8` to `agilantsolutions/claude` `_main`, `validate: OK (30 repos)`.
|
|
39
|
+
|
|
40
|
+
## What did NOT work — DO NOT RETRY THESE
|
|
41
|
+
|
|
42
|
+
- **`test/@dave/Junk Drawer/NetSuite/nsq.php` cannot start at all.** Its bootstrap does `realpath(__DIR__ . '/../../../worker')`, which from its current folder resolves to `test/worker` — a path that does not exist (the file was moved deeper without fixing it). Use `test/@dave/nsq2.php` instead.
|
|
43
|
+
- **Booting worker2 locally with only `_underscore` prepended to `include_path` fails.** Exact error: `Exception: Could not determine application root!` thrown from `C:\www\library\_.php:17`. Cause: `class _underscore` lives in `worker2/_.php`, and `C:\www\library\_.php` (the unrelated 1.0 framework file, already on the path) **shadows** it. The **app root must come FIRST**: `set_include_path(<worker2> . PATH_SEPARATOR . <_underscore> . PATH_SEPARATOR . get_include_path())`.
|
|
44
|
+
- **The framework's own error handler crashes while reporting that failure**, masking the real exception: `Error: Class "_underscore" not found in C:\www\_underscore\Error.php:373`. Wrap `require 'index.php'` in try/catch to see the true error.
|
|
45
|
+
- **NetSuite auth from a laptop fails on the database, not on credentials.** First `Access denied for user 'root'@'localhost' (using password: YES)` at `_underscore/Database.php:495`, then `Unknown database 'clientlogs'`. Root cause: **`send()` does NOT call `setLogging(false)`** — there are zero `setLogging` occurrences in `_underscore/Component/Api/Netsuite/Netsuite.php` and all three `_ApiRequest` constructions pass only 4 args, so logging defaults ON and every call writes to `DB_CLIENT_LOGS`. The KB previously claimed the opposite (now corrected). A laptop with no `ClientLogs` schema cannot make even a read-only SuiteQL call.
|
|
46
|
+
- **`GROUP_CONCAT` for bulk row export truncates silently** at 1,024 chars (~100 of 700 rows) — `group_concat_max_len` default. Fix: the optimizer hint `SELECT /*+ SET_VAR(group_concat_max_len=4194304) */ …`, which works through the `toga-db` MCP and returned all 4,308 rows in two blobs.
|
|
47
|
+
- **Bash heredocs above roughly 900 whitespace-separated tokens fail** with `unexpected EOF while looking for matching quote` — a command-length limit, not a quoting error. The write is rejected atomically (no partial file). Split into ~500-token chunks.
|
|
48
|
+
- **`php -l` cannot catch malformed SQL string assembly.** The ODKU double-comma bug lint-passed cleanly. Only rendering the built SQL through a stub harness exposed it.
|
|
49
|
+
- **Wrote the dbchanges2 migration into the WRONG checkout.** `C:\www\2.0\dbchanges2` is a **stale duplicate** (branch TRUE-80742, newest commit 2026-08-19, 144 `Core/` files). The live repo is **`C:\www\dbchanges2`** (`_main`, committed 2026-08-27, 151 files). The developer could not see the file. Same trap already documented for `worker2` and `_underscore`; the `repo-path-dbchanges2` memory has been corrected. `worker` and `library` have only one checkout each.
|
|
50
|
+
- **`git checkout -- <file>` is refused by the GateGuard destructive-command hook** even after presenting the required facts — unlike the write gate, it does not mark-and-allow on retry. Reverse the edit with the Edit tool instead (or set `ECC_GATEGUARD=off`).
|
|
51
|
+
- **I twice asserted `import_open_orders.php` is `active: 0`. It is `active: 1`.** This came from an early exploration report and was repeated without re-verification. It matters a great deal — see Blockers.
|
|
52
|
+
- **My MySQL error-1093 explanation was wrong.** I attributed the clean run to the `FROM (SELECT 1) AS placeholder` form. bala pushed a tested correction (MySQL 8.0.30): **1093 is UPDATE/DELETE-only**, and an unwrapped self-referencing `INSERT … SELECT … WHERE NOT EXISTS` runs clean. The migration is fine; the reasoning was not.
|
|
53
|
+
- **`knowledge.js kickoff-preflight`'s feature matcher has a dead zone.** `--q="forecast locations netsuite backfill import"` matched **0** feature docs; `--q=netsuite` matched **21** (~74k tokens, `estimate.heavy`). Neither is usable — hand-pick the read list.
|
|
54
|
+
- **`knowledge.js publish` was deliberately NOT used for the capture.** A second capture agent was writing the same repo concurrently, and `publish` does `git add knowledge/`, which would have swept its unapproved drafts into the commit. Only the 7 intended files were staged, so the auto-generated `INDEX.md` files are briefly stale on the remote.
|
|
55
|
+
|
|
56
|
+
## Not tried yet (candidates for next session)
|
|
57
|
+
|
|
58
|
+
- **THE LIVE THREAD — three writers to `Forecast.Locations`, unverified.** `common_import_sales_from_netsuite.php:1636-1934` **auto-INSERTs and RENAMEs** `Locations` rows and writes `locationId` itself. It is gated behind `SHOULD_SYNC_OPEN_ORDERS` (`:1585`), which `import_open_orders.php` sets — and that cron is **active: 1**. It has been dormant since June only because `$nsItem->location` was always null; `tl.location` came back in `library` on **2026-08-28** (jcardinal's NYCHH transfer-order work, `rest.php:448`), so **this code just woke up**. It now coexists with `_Worker_Netsuite_Location::SyncAll` (newly scheduled) and the fixed 3 AM cron. Nobody has checked whether they agree on `Locations.name`.
|
|
59
|
+
- **The backfill has never been run against prod** — only the local fixture. Sequence: `dryRun: true` scoped with `dateOrderTo: "2024-08-26"`, read the "order no longer open" list, then live.
|
|
60
|
+
- **PHP 7.2 lint never performed.** `worker` and `library` run 7.2 in prod; only `C:\xampp\php` (8.2) exists here.
|
|
61
|
+
- **Open/closed status of the 608 pre-window orders never checked in prod.** Some of those 1,133 rows are likely stale and should be removed rather than given a location — the backfill reports them deliberately instead of filling them.
|
|
62
|
+
- **Non-prod parity migration not written.** dev-sandbox has no `Locations` table and no `OpenOrderItems.locationId`, so none of this is testable outside prod. Complicated by MySQL 8 having no `ADD COLUMN IF NOT EXISTS`.
|
|
63
|
+
- **Leaf-level location as a new column** — deferred, if reporting ever wants sub-location granularity.
|
|
64
|
+
- **Post-deploy verification:** confirm the NULL census drops by ~1,711 after the next 3 AM run, and does **not** climb the following night.
|
|
65
|
+
- **Security ticket + rotation:** `tools/bin/forecast/fixer.php:1424` and `checker.php:678` carry inline production database credentials (surfaced by a concurrent capture agent; values deliberately not recorded).
|
|
66
|
+
|
|
67
|
+
## Current file state
|
|
68
|
+
|
|
69
|
+
| File | Status | Notes |
|
|
70
|
+
|------|--------|-------|
|
|
71
|
+
| `worker/crons/toga2/forecast2/periodic_forecast_discrepancy_fix_open_orders.php` | Modified, **committed** (PR, branch `TRUE-80326`) | 5 edits: `location AS locationid` in line SQL; `Locations` lookup + `$hasLocationCol` guards; `resolveOoiLocationId()`; `locationId` in both builders (converted to array params); `locationId` in `$compare`. Lint clean on 8.2 only. |
|
|
72
|
+
| `worker2/Worker/Netsuite/OpenOrderLocationBackfill.php` | **New**, committed (PR) | `dryRun` default true, reuses `_Component_Forecast_Db::resolveLocationId`, `AND locationId IS NULL` guard, reports non-open orders instead of filling them. Verified on the local fixture, never run in prod. |
|
|
73
|
+
| `worker2/Worker/Netsuite/Location.php` | Modified, committed (PR) | Added `ORDER BY id` to the offset-paged `SELECT id, name, parent FROM location`. |
|
|
74
|
+
| `dbchanges2/Core/2026-08-27a - Insert - Netsuite Location SyncAll CronJob.sql` | **New**, committed (PR) | In the LIVE repo `C:\www\dbchanges2`. Daily 2:30 AM Central, `isActive=1`, `NOT EXISTS`-guarded, pre-generated v4 UUID literal. |
|
|
75
|
+
| `library/app/api/netsuite/rest.php` | **Reverted — no local change** | Phase 4 was written, verified against live NetSuite, then reverted at the developer's request. Subsequently found **unnecessary**: `tl.location` is back at `:448` upstream as of `27ae6a68` (2026-08-28). Patch + full-file backup remain in the session scratchpad but are now redundant. |
|
|
76
|
+
| `test/@dave/probe_ooi_location_gap.php` | **New**, uncommitted (untracked) | Read-only auditor; zero write statements verified by grep. Takes `sos.txt` + `locations.txt` as argv. |
|
|
77
|
+
| `test/@dave/nsq2.php` | **New**, uncommitted (untracked) | Working read-only SuiteQL runner replacing the broken `nsq.php`. |
|
|
78
|
+
| `test/@dave/ooi_location_gap.csv` | **New**, uncommitted | 4,387 rows of per-line results. Advised NOT to commit (production order identifiers). |
|
|
79
|
+
| `test/@dave/ooi_expected_location_by_order.csv` | **New**, uncommitted | 1,366 order -> expected `locationId` rows. Advised NOT to commit. |
|
|
80
|
+
| `worker2/Config/dev-kyalamarthi-laptop.ini` | Modified, untracked by git | `[database] password` blanked to match local MySQL (root, no password). Backup in scratchpad. **Local only — never commit.** |
|
|
81
|
+
| Local MySQL `ClientLogs` schema | **Created** | 6 tables from `dbchanges2/Logs_Client/2026-08-14 - BLANK CLIENT LOGS DATABASE.sql`. Required because every NetSuite call logs. |
|
|
82
|
+
| Local MySQL `Forecast` schema | **Created** | Minimal fixture: `Locations` (184 rows) + `OpenOrderItems` (6 canary rows for order 235272). |
|
|
83
|
+
| `knowledge/` (team repo) | **Pushed**, commit `59f2cd8` | 7 docs updated by `/capture`. |
|
|
84
|
+
|
|
85
|
+
## Decisions made
|
|
86
|
+
|
|
87
|
+
- **`locationId` stores the ROOT warehouse, not the line's leaf sub-location.** Rationale: all 1,425 already-populated rows hold one of 6 roots; `resolveLocationId` rolls up deliberately because reporting wants the top level; storing leaves would give backfilled rows a different meaning from existing ones. Data: 25+ distinct leaves collapse to 6 roots, and within an order the leaf almost never varies (15 of 1,366 orders; only 7 resolve to >1 root, each with exactly 1 NULL row). **Rejected:** leaf storage. **Deferred:** leaf as a NEW column.
|
|
88
|
+
- **Fix the cron FIRST, backfill second.** Rationale: the cron's delete+reinsert erases the column, so a backfill-first order regrows the gap the same night. Adding `locationId` to `$compare` also makes the cron repair all 1,711 in-window rows by itself, shrinking the backfill from 2,883 rows to 1,133.
|
|
89
|
+
- **A targeted backfill worker, not a re-import.** **Rejected:** enqueuing `Netsuite/SalesOrder/put` for 1,366 orders — the KB documents a REST-sublist read-lag race that can WIPE an order's rows, and re-import needlessly recomputes revenue/profit. **Rejected:** a one-off `dbchanges2` UPDATE — not repeatable, and that repo has never contained an `UPDATE OpenOrderItems`.
|
|
90
|
+
- **The backfill reuses `_Component_Forecast_Db::resolveLocationId`** rather than adding a third copy of the rollup, so it cannot disagree with the live importer.
|
|
91
|
+
- **The probe deliberately does NOT use `resolveLocationId`** for classification, because that function returns NULL both when NetSuite has no location and when the leaf is unknown to `Locations` — collapsing exactly the two buckets an audit must separate.
|
|
92
|
+
- **Converted the cron's two SQL builders to a single array parameter** (team rule: 4+ params must use an associative array) rather than adding a 13th positional argument.
|
|
93
|
+
- **Guarded the `Core.CronJobs` insert with `NOT EXISTS`** even though every sibling registration in that folder is unguarded — dbchanges2 rule #7 now requires it.
|
|
94
|
+
- **`library/rest.php` split into its own PR** (TRUE-79078) rather than bundled, because it is an independent regression with tier-wide blast radius. Later moot — it landed upstream.
|
|
95
|
+
|
|
96
|
+
## Blockers
|
|
97
|
+
|
|
98
|
+
- **UNRESOLVED AND TIME-SENSITIVE: three writers now target `Forecast.Locations`.** The 1.0 rollup at `common_import_sales_from_netsuite.php:1636-1934` auto-INSERTs and RENAMEs rows there; it is reached by `import_open_orders.php`, which is **`active: 1`**, and it went live again when `tl.location` returned to `library` on 2026-08-28. It has not been reconciled against `_Worker_Netsuite_Location::SyncAll` (newly scheduled 2:30 AM) or the fixed 3 AM cron. Both run tonight.
|
|
99
|
+
- **No PHP 7.2 binary on this machine** (only `C:\xampp\php`, 8.2), so 7.2 compatibility of the `worker` change is unproven. A single 7.4+ token in the 1.0 tier is a production outage; the PR should not be marked lint-verified.
|
|
100
|
+
- **Nothing in this work is testable outside production** — dev-sandbox lacks the `Locations` table and the `locationId` column entirely.
|
|
101
|
+
|
|
102
|
+
## Exact next step
|
|
103
|
+
|
|
104
|
+
> Before tonight's crons run, confirm the newly re-animated 1.0 rollup cannot fight `SyncAll` over `Forecast.Locations.name`. Run `git -C C:\www\worker grep -n "INSERT INTO Locations" -A 6 crons/toga2/forecast2/common_import_sales_from_netsuite.php` and compare the `rootName` it writes (from `App_Api_Netsuite_Rest::listLocations()`) against the `name` that `_Worker_Netsuite_Location::SyncAll` stores (NetSuite `location.name`) for the same 12 root ids. If they can differ, gate one writer before 2:30 AM.
|
|
105
|
+
|
|
106
|
+
---
|
|
107
|
+
_Saved by /session-save on 2026-08-28_
|
package/package.json
CHANGED