toga-ai 1.0.673 → 1.0.675

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -5,6 +5,7 @@
5
5
  | [Worker (worker2) Architecture](architecture.md) | Worker (repo `worker2`) is an AWS Elastic Beanstalk **Worker Tier** application that processes background jobs. | worker2/Controller/Index.php, worker2/Worker/, worker2/LambdaFunctions/, _underscore/Worker.php, worker2/composer.json |
6
6
  | [Deploy-Time Auto-Registration to the Shared ALB Target Group (non-production)](features/alb-target-group-auto-registration.md) | TOGA does **not** pay for EB-managed load-balancer registration, so an EB instance is normally **not** added to its environment's ALB target group — a fresh or | worker2/.platform/hooks/postdeploy/060_register_instance_to_shared_application_load_balancer.sh, worker2/ebs/register_instance_to_shared_application_load_balancer.php, worker2/.platform/hooks/_shared/040-write-instance-id.sh, worker2/.platform/hooks/_shared/041-write-region.sh, worker2/.platform/hooks/_shared/042-write-eb-environment.sh, worker2/.platform/hooks/postdeploy/015_install_composer.sh, saml/.platform/hooks/_shared/040-write-instance-id.sh, saml/.platform/hooks/_shared/041-write-region.sh, saml/.platform/hooks/_shared/042-write-eb-environment.sh, saml/.platform/hooks/prestart/040-write-instance-id.sh, saml/.platform/hooks/prestart/041-write-region.sh, saml/.platform/hooks/postdeploy/040-write-instance-id.sh, saml/.platform/hooks/postdeploy/041-write-region.sh, saml/.platform/hooks/postdeploy/042-write-eb-environment.sh, saml/.platform/hooks/postdeploy/060_register_instance_to_shared_application_load_balancer.sh, saml/ebs/register_instance_to_shared_application_load_balancer.php, saml/composer.json, api2/.platform/hooks/prebuild/_shared/042-write-eb-environment.sh, api2/.platform/hooks/postdeploy/060_register_instance_to_shared_application_load_balancer.sh, api2/ebs/register_instance_to_shared_application_load_balancer.php |
7
7
  | [All-Client Email Queue Monitor (Monitor/Operations/EmailQueue)](features/all-client-email-queue-monitor.md) | `_Worker_Monitor_Operations::EmailQueue()` is the cross-client health check for the 2.0 outbound email queue (`Logs_<Client>.Email` — see [2.0 Email Send Pipeli | worker2/Worker/Monitor/Operations.php, _underscore/Database.php, _underscore/Query.php, dbchanges2/Logs_Client/2026-05-21 - Email.sql |
8
+ | [API 2.0 Stability Monitor (Monitor/Operations/ApiStability)](features/api-stability-monitor.md) | `_Worker_Monitor_Operations::ApiStability()` is a cross-client/operational **end-to-end round-trip probe** of the production 2.0 API. | worker2/Worker/Monitor/Operations.php, dbchanges2/Core/2026-08-28a - API 2.0 Stability Monitor Cron.sql |
8
9
  | [Automated PR Merger — Concurrent Force-Push Clobber Race](features/automated-pr-merger-force-push-race.md) | The automated PR merger `_Worker_Team_GitHub::Merge` (`worker2` `Worker/Team/Github.php`) merges approved PRs to `_production` by **force-pushing from a clone t | Worker/Team/Github.php |
9
10
  | [Callback Scheduling ("call me back") — worker2 AI-BDR](features/callback-scheduling.md) | The **AI-BDR callback path**: what happens between a prospect saying *"call me back later"* on a Vapi call and the dialer actually placing that second call. | worker2/Worker/Vapi.php, worker2/Worker/Ai/Bdr/Vapi.php, worker2/Controller/Index.php |
10
11
  | [ClickUp Connectivity Watchdog](features/clickup-connectivity-watchdog.md) | A cron watchdog that emails when the ClickUp integration looks disconnected during business hours. | worker2/Worker/Clickup/Health.php, worker2/Database/ClickupHealthWatchdog.sql |
@@ -0,0 +1,140 @@
1
+ ---
2
+ title: API 2.0 Stability Monitor (Monitor/Operations/ApiStability)
3
+ framework: "2.0"
4
+ repo: worker2
5
+ project: Worker
6
+ client: shared
7
+ type: feature
8
+ status: active
9
+ updated: 2026-08-28
10
+ owners: ["jcardinal"]
11
+ files:
12
+ - worker2/Worker/Monitor/Operations.php
13
+ - dbchanges2/Core/2026-08-28a - API 2.0 Stability Monitor Cron.sql
14
+ related:
15
+ - ./oneuptime-worker2-monitoring.md
16
+ - ./netsuite-integrations-monitor.md
17
+ - ./all-client-email-queue-monitor.md
18
+ - ./elastic-beanstalk-health-monitor.md
19
+ ---
20
+
21
+ ## Summary
22
+
23
+ `_Worker_Monitor_Operations::ApiStability()` is a cross-client/operational **end-to-end
24
+ round-trip probe** of the production 2.0 API. Unlike a log-scraper (blind to a dead
25
+ endpoint) or a cheap health-endpoint liveness probe, it performs a real authenticated
26
+ call chain against `api.togahub.com/v2` and only reports OK when the API round-trips a
27
+ representative read correctly. It follows the established "dumb reporter, smart monitor"
28
+ OneUptime push pattern — same shape as the sibling
29
+ [NetsuiteIntegrations](./netsuite-integrations-monitor.md) and
30
+ [EmailQueue](./all-client-email-queue-monitor.md) monitors on
31
+ `_Worker_Monitor_Operations`. Shared internal operational infrastructure, **not** a client
32
+ feature.
33
+
34
+ **Motivating incident.** A `Core.Records` model failed to deploy to production. Because the
35
+ 2.0 API requires a model for every `Core.Records` entry, any call touching the missing
36
+ model returned HTTP 500 — while the Elastic Beanstalk instances stayed "healthy" and the
37
+ `/v2/health` short-circuit route kept returning 200. A monitor that actually round-trips
38
+ the API would have gone red immediately; a health-endpoint probe and an EB health check
39
+ both stayed green. This probe closes that gap.
40
+
41
+ ## Key files / entry points
42
+
43
+ - `worker2/Worker/Monitor/Operations.php` — `abstract class _Worker_Monitor_Operations`;
44
+ `public static function ApiStability(): string`, route
45
+ `Monitor/Operations/ApiStability`. All config is `ALL_CAPS` in-method locals (no
46
+ arguments): `$ONEUPTIME_URL` (push credential), `$API_BASE`, timeouts, expected statuses,
47
+ `$EXPECTED_CLIENT_AUTH_TYPE = 'SSO'`, `$DOMAINS_QUERY`, and `$ORIGIN`.
48
+ - `dbchanges2/Core/2026-08-28a - API 2.0 Stability Monitor Cron.sql` — the `Core.CronJobs`
49
+ seed row (see below).
50
+
51
+ ## How it works
52
+
53
+ Two sequential sub-checks against **production** `api.togahub.com/v2` (the `$API_BASE` is
54
+ hardcoded — the monitor watches prod regardless of which environment runs it). HTTP goes
55
+ through `_ApiRequest` with `ENCODE__JSON` (so `execute()` returns the decoded object) and a
56
+ fresh `transactionId` per request via `_String::generateUuid()`:
57
+
58
+ 1. **Auth** — `POST /v2/auth/public?transactionId=<uuid>` with an `Origin` header (see the
59
+ gotcha below). Expect **HTTP 201**; extract the bearer token from `data.tokens.access`.
60
+ 2. **Representative read** — `GET /v2/domains?<fields/join/where>&transactionId=<uuid>` with
61
+ `Authorization: Bearer <token>` and the same `Origin`. The query is a representative
62
+ joined read across `Domains` + `Environments` + `ClientAuthentications` + `Apps`, scoped
63
+ to the `compass.togasupply.com` PRODUCTION domain. Expect **HTTP 200 AND**
64
+ `data.domains[0].ClientAuthentications.type === 'SSO'` (`$EXPECTED_CLIENT_AUTH_TYPE`).
65
+ This is the exact call that failed during the missing-model incident.
66
+
67
+ Any failed sub-check sets `alarm = 'DEGRADED'` (else `OK`). The method pushes
68
+ `{status:'reporting', alarm, authStatus, domainsStatus, failureCount, failures,
69
+ checkedAtUtc}` to OneUptime, then returns a structured summary string recorded in
70
+ `WorkerJobs.output`.
71
+
72
+ Compass USA's production supply domain is used only as a **representative** live call — this
73
+ is shared platform/liveness knowledge, not client-specific behavior.
74
+
75
+ ### Never-throw / non-fatal push
76
+
77
+ Logging is off, `throwExceptionsOnFailure` is off, and the push is wrapped so a
78
+ monitoring-side failure never fails the worker job — the standard reporter contract from
79
+ [the OneUptime push-metric pattern](./oneuptime-worker2-monitoring.md).
80
+
81
+ ## Cron seed (dbchanges2)
82
+
83
+ `dbchanges2/Core/2026-08-28a - API 2.0 Stability Monitor Cron.sql` — one `Core.CronJobs`
84
+ row: action `Monitor/Operations/ApiStability`, schedule `*/5 * * * *` (every 5 min, all
85
+ hours), `maxExecutionTime` 120, `parameters` NULL (no arguments), **`isActive = 0`**
86
+ (ships inert per the deploy-inert convention — flip to 1 after the OneUptime monitor exists
87
+ and its push URL is set). Core-database-only references (cluster-isolation satisfied); no
88
+ `Core.Records`/`RecordFields` involved.
89
+
90
+ ## Gotchas / known issues
91
+
92
+ - **`POST /v2/auth/public` returns HTTP 401 without an `Origin` request header.** The
93
+ public-auth endpoint carries **no** client credential; the API resolves which client/app
94
+ the public token is minted for from the request's `Origin` header. A server-side call must
95
+ set it explicitly (`Origin: https://compass.togasupply.com`) — the browser supplies Origin
96
+ automatically, which is why the front end's public-token flow
97
+ (`toga-blox-npm/src/api/auth.ts` POSTs to `/auth/public` with a null body + JSON
98
+ content-type) never sees this. Send the same `Origin` on the `/domains` call too, mirroring
99
+ the browser. This is the non-obvious fact for anyone scripting a server-side call to the
100
+ public-auth endpoint. (Verified 2026-08-28: with Origin, auth → 201 and domains → 200 +
101
+ `SSO`; without it, auth → 401.)
102
+ - **The push URL is a credential** — it lives only as the in-method `$ONEUPTIME_URL` local,
103
+ never logged and never written into a doc or into dbchanges2.
104
+ - **Watches PRODUCTION regardless of runner.** `$API_BASE` is hardcoded to prod, so a
105
+ non-prod worker still probes prod — intended, but do not "fix" it to the running env.
106
+
107
+ ## OneUptime import artifact
108
+
109
+ Provisioned by importing an Incoming-Request monitor export (a local staging artifact,
110
+ `C:\STAGE\monitor-export-api-2.0-stability-2026-08-28.json` — **not** a repo file) with three
111
+ criteria: body Contains `"alarm":"DEGRADED"` → Degraded + incident (templated with
112
+ `{{requestBody.failures}}` etc.); body Contains `"alarm":"OK"` → Operational;
113
+ not-received-25-min → Offline + incident. Per the
114
+ [activation-ordering trap](./oneuptime-worker2-monitoring.md#gotchas--known-issues), the
115
+ absence criterion ships `isEnabled: false` — enable it after the first push lands.
116
+ Project-specific status/severity ObjectIDs were cloned from an existing monitor export.
117
+
118
+ ## Change history
119
+
120
+ - 2026-08-28 — Built `ApiStability()` on `_Worker_Monitor_Operations` — a two-step
121
+ end-to-end round-trip probe (auth/public → 201 → bearer, then a joined `/v2/domains` read
122
+ → 200 + `SSO`) against production `api.togahub.com/v2`, plus the `Core.CronJobs` seed
123
+ (`*/5 * * * *`, `isActive=0`) and an Incoming-Request OneUptime import. Motivated by a
124
+ missing `Core.Records` model that 500'd real calls while EB/health-endpoint stayed green.
125
+ Recorded the durable gotcha that `POST /v2/auth/public` **401s without an `Origin`
126
+ header** (the API resolves the client from Origin). Verified working end-to-end after the
127
+ Origin fix. (jcardinal)
128
+
129
+ ## Related docs
130
+
131
+ - [OneUptime push-metric monitors for 2.0 workers](./oneuptime-worker2-monitoring.md) — the
132
+ push/token pattern, OneUptime criteria, import-JSON schema, deploy-inert and
133
+ activation-ordering conventions, and the liveness-probe-vs-log-scraper distinction.
134
+ - [NetSuite Integrations Monitor](./netsuite-integrations-monitor.md) and
135
+ [All-Client Email Queue Monitor](./all-client-email-queue-monitor.md) — sibling
136
+ `_Worker_Monitor_Operations` monitors modeled on the same reporter contract.
137
+ - [Elastic Beanstalk Health Monitor](./elastic-beanstalk-health-monitor.md) — EB-level
138
+ health this probe deliberately reaches past (EB stayed healthy during the incident).
139
+ </content>
140
+ </invoke>
@@ -6,7 +6,7 @@ project: Worker
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-08-20
9
+ updated: 2026-08-28
10
10
  owners: ["jcardinal", "mhammontree", "bala"]
11
11
  files:
12
12
  - worker2/Worker/Monitor/Compass.php
@@ -195,6 +195,21 @@ touches it. Rules for the probe:
195
195
  duplicate credentials into a second repo and race the shared access token. Accept and record the
196
196
  coverage gap rather than forking the client.
197
197
 
198
+ #### A round-trip probe is stronger than a health-endpoint probe
199
+
200
+ A health endpoint that short-circuits before any real work (`/v2/health` → 200) proves the
201
+ process is up, but **not** that the API can actually serve a request. A missing `Core.Records`
202
+ model once 500'd every real call while `/v2/health` and EB health both stayed green. Where an
203
+ integration matters, add an **end-to-end round-trip probe** that authenticates and reads a
204
+ representative record, asserting on the returned data — not just the HTTP status. First
205
+ instance: [API 2.0 Stability Monitor](./api-stability-monitor.md), which mints a public token
206
+ then reads `/v2/domains` and asserts `ClientAuthentications.type === 'SSO'`.
207
+
208
+ **Gotcha for a server-side public-auth call:** `POST /v2/auth/public` returns **HTTP 401
209
+ without an `Origin` request header** — the endpoint carries no client credential and resolves
210
+ the client/app from `Origin`. The browser supplies it automatically; a scripted call must set
211
+ it explicitly. See the API-stability doc for details.
212
+
198
213
  ### Tunable thresholds live in `Core.CronJobs.parameters`, everything else stays a local
199
214
 
200
215
  The alarm threshold is the **one** setting that goes in the DB, as
@@ -356,6 +371,12 @@ check for another client.
356
371
  monitors must pass the region (see [Cloud S3 helpers](../../_underscore/features/cloud-s3-helpers.md)).
357
372
 
358
373
  ## Change history
374
+ - 2026-08-28 — Added the **round-trip probe** refinement (an authenticated end-to-end read
375
+ that asserts on returned data is stronger than a `/v2/health` short-circuit or EB health
376
+ check — a missing `Core.Records` model 500'd real calls while both stayed green) and the
377
+ durable gotcha that `POST /v2/auth/public` **401s without an `Origin` header** (the API
378
+ resolves the client from Origin). First instance:
379
+ [API 2.0 Stability Monitor](./api-stability-monitor.md). (jcardinal)
359
380
  - 2026-08-24 — Documented the **monitor import-JSON schema** the runbook previously described
360
381
  only conceptually: the `oneuptime-resource-export` envelope, `monitorType` must be the
361
382
  spaced display name `"Incoming Request"`, the `_type`-tagged
@@ -410,6 +431,7 @@ check for another client.
410
431
  5-min cadence → 10/15-min heartbeat timing. (jcardinal)
411
432
 
412
433
  ## Related docs
434
+ - [API 2.0 Stability Monitor](./api-stability-monitor.md) — the first end-to-end round-trip probe instance (auth/public + /domains SSO assertion)
413
435
  - [Monitoring Framework](./monitoring-framework.md) — the parallel DB-driven, email-alert monitoring pattern
414
436
  - [Cloud S3 helpers](../../_underscore/features/cloud-s3-helpers.md) — `_Cloud::getS3Objects()` used to list the EDI bucket
415
437
  - [OneUptime 1.0 worker uptime monitoring](../../../1.0/apps/worker/features/oneuptime-worker-uptime-monitoring.md) — the 1.0 push-heartbeat predecessor
@@ -19,7 +19,7 @@ _Auto-generated by `knowledge.js index`. Do not hand-edit._
19
19
  ## 2.0 framework
20
20
 
21
21
  - **_underscore** (_Underscore) _(framework core)_ — 70 doc(s) → [2.0/apps/_underscore/INDEX.md](2.0/apps/_underscore/INDEX.md)
22
- - **worker2** (Worker) — 57 doc(s) → [2.0/apps/worker2/INDEX.md](2.0/apps/worker2/INDEX.md)
22
+ - **worker2** (Worker) — 58 doc(s) → [2.0/apps/worker2/INDEX.md](2.0/apps/worker2/INDEX.md)
23
23
  - **api2** (API) — 25 doc(s) → [2.0/apps/api2/INDEX.md](2.0/apps/api2/INDEX.md)
24
24
  - **dbchanges2** (Database Changes) _(framework core)_ — 13 doc(s) → [2.0/apps/dbchanges2/INDEX.md](2.0/apps/dbchanges2/INDEX.md)
25
25
  - **toga2-supply** (TOGa Supply) — 9 doc(s) → [2.0/apps/toga2-supply/INDEX.md](2.0/apps/toga2-supply/INDEX.md)
@@ -0,0 +1,244 @@
1
+ ---
2
+ type: session
3
+ slug: nycdoe-2-0-integration
4
+ title: NYCDOE 2.0 ServiceNow/ASN/NetSuite Integration — Design & Pseudocode Prep (TRUE-80519)
5
+ author: mhammontree
6
+ repos: [worker, library, togadesk, dbchanges, _underscore, worker2, dbchanges2, api2]
7
+ framework: "both"
8
+ client: nycdoe
9
+ status: active
10
+ created: 2026-08-28
11
+ updated: 2026-08-28
12
+ ---
13
+
14
+ # Session: nycdoe-2-0-integration
15
+ **Date:** 2026-08-28
16
+ **Project/Repo:** worker / library / togadesk / dbchanges (1.0) + _underscore / worker2 / dbchanges2 / api2 (2.0)
17
+ **Task:** Multi-day research/design session (TRUE-80519) producing the rundown and 2.0 redesign of
18
+ NYCDOE's ServiceNow/ASN/NetSuite integration, for Mark to author his own pseudocode proposal for
19
+ Jeff (dev team director) to approve. Explicitly a planning ticket — no code changes. Mid-session,
20
+ Northwell entered the picture as a second client needing the same SFTP→NetSuite pipeline (installs
21
+ only, no ServiceNow/RITM until ~next year), which reframed the work toward a client-agnostic,
22
+ multi-tenant-scalable design. Session ends with Jeff formally scoping **Phase 1 = the SFTP
23
+ ingestion portion**, with an explicit requirement to scale to **1000+ clients**.
24
+
25
+ ---
26
+
27
+ ## What WORKED
28
+
29
+ - **ASN-first design principle, held throughout without revision**: nothing is ever created in the
30
+ 2.0 system except from a confirmed ingested ASN. ServiceNow and NetSuite are sync *targets*,
31
+ never trusted sources. This survived every subsequent design pressure-test this session.
32
+ - **Case → ServiceRequest is the correct DOE entity model — confirmed directly by Jeff/Mark,
33
+ correcting an earlier 3-tier assumption**: `Case` = the 1.0 `managed_service_order` equivalent,
34
+ `ServiceRequest` = the 1.0 `repair_order` equivalent. **DOE will NOT use `Ticket` at all** — an
35
+ earlier design (this session) had assumed `Case → ServiceRequest → Ticket`, which was wrong for
36
+ DOE specifically.
37
+ - **`worker2/Worker/Netsuite/SalesOrder.php` confirmed real, generic, multi-client, already in
38
+ production** (`Netsuite/SalesOrder::Sync`) — verified by direct code read, not assumed. Includes
39
+ freeze/idempotency semantics already. Reuse this; do not write a new SO push.
40
+ - **`_Trait_Netsuite_PurchaseOrder::syncNetsuitePurchaseOrder()` and
41
+ `_Trait_Netsuite_ItemReceipt::syncNetsuiteItemReceipt()` confirmed real, working, inbound sync
42
+ methods** (NetSuite → Toga) via direct code read — not previously known to exist before this
43
+ session's code review.
44
+ - **Confirmed gap, verified by code search (not assumed)**: no outbound serial/inventory-assignment
45
+ push to NetSuite exists anywhere in `_underscore`/`worker2`. `c_netsuiteInternalInventoryAssignmentId`
46
+ is declared on `Trait_Netsuite_Unit`/`ItemFulfillmentItemUnit` but has zero sync logic behind it.
47
+ This is genuinely new work (the 2.0 equivalent of 1.0's `2_send_serials_to_netsuite.php`).
48
+ - **`Client_Nycdoe` confirmed already provisioned with the full standard 2.0 platform schema**
49
+ (`PurchaseOrders`, `SalesOrders`, `PurchaseOrders_SalesOrders`, `PurchaseOrderItems_SalesOrderItems`,
50
+ `ItemReceipts`, `Units`, `Cases`, `ServiceRequests`, `Tickets`, `Locations`) — verified via live
51
+ prod schema query, 0 rows in every business table. No new client database needed; this is
52
+ additive migrations only.
53
+ - **`Locations` table reuse confirmed workable for DOE school-site data** — real columns checked
54
+ (`locationTypeId`, `parentLocationId`, `primaryLocationAddressId`, `allowReceiving`/`allowSending`).
55
+ Only new work: 2–3 `c_` fields (LDA/location code, building code, SNOW `sys_id`) + one
56
+ `LocationTypes` seed row. No new table.
57
+ - **`Files` table reuse confirmed workable for the SFTP import ledger** — real columns checked
58
+ (`checksum`, `sourceUrl`, `size`, `name`, `type`). `checksum` gives content-based dedupe for free.
59
+ Workflow-state fields (`c_status`, `c_vendorId`) go as custom fields directly on `Files` — NOT a
60
+ separate bridge/companion table, since Files↔ledger-status is a strict 1:1 relationship (bridge
61
+ tables are for genuine one-to-many/many-to-many only — this is the actual platform convention,
62
+ confirmed by every real example found this session, not written down as an explicit rule anywhere).
63
+ - **`Core.Partners` / `Core.PartnerApiIdentities` confirmed real via direct code read, with a
64
+ complete, working, production reference implementation** in `worker2/Worker/Startech.php`:
65
+ resolve `(partnerName, partnerIdentifier)` → `PartnerApiIdentity` → `clientId` → register that
66
+ client's DB → `clientApiId` → that client's own `Apis` credentials (uuid + secret). This is the
67
+ mechanism for "which clients have manufacturer X's SFTP feature enabled" and "what credentials do
68
+ we use to act as that client" — no need to invent a parallel mechanism.
69
+ - **Live production evidence (real ASN CSV + real NetSuite screenshots) root-caused a genuine 1.0
70
+ bug**: one customer PO (`WR270026786`), one ASN file, containing both a serialized Lexmark
71
+ printer and a non-serialized Belkin cable, produced **two separate NetSuite Sales Orders a day
72
+ apart**. Root cause (best-supported theory, not directly log-confirmed): 1.0 has two
73
+ uncoordinated, independently-scheduled ingestion pipelines (`import_asn.php` polling ServiceNow's
74
+ API for serialized items vs. `legacy_import_asn.php` reading the SFTP file for non-serialized
75
+ items) that can race each other on the same shipment. A single ASN-file-only pipeline (this
76
+ session's design) structurally removes the race.
77
+ - **NetSuite PO vs. SO scoping clarified and used to settle a genuine design flip-flop**: a NetSuite
78
+ PO is vendor-scoped (Lexmark and Belkin can never share one PO); a Sales Order is customer-scoped
79
+ and the platform's `PurchaseOrders_SalesOrders` bridge already supports one SO fulfilled by
80
+ multiple vendor POs. This means the original "SO/PO stamp at the Case/header level" design likely
81
+ still holds — the observed 2-SO behavior is attributed to 1.0's pipeline race, not a genuine
82
+ NetSuite/business requirement. (Not fully confirmed — see Not Tried Yet.)
83
+ - **Two independent, opt-in `dbchanges2` modules designed**: `_modules/servicenow/` (RITM/ticketing
84
+ linkage, generic) and `_modules/sftp-import/` (file ledger, generic, no DOE-specific fields).
85
+ Deliberately separate because Northwell needs SFTP/NetSuite now but not ServiceNow until ~next
86
+ year — a client can adopt either independently.
87
+ - **Non-serialized quantity-progress display problem solved architecturally, not with a display
88
+ hack**: Skyler wants "10/10" not "1/10" for a 10-cable line. Resolution: track
89
+ `receivedQty`/`orderedQty` directly on `ServiceRequest`, sourced from NetSuite's own `RECEIVED`
90
+ field on that line's PO — no per-`Unit` counting needed at all, for serialized or non-serialized
91
+ items, once SO/PO lives at the right grain.
92
+ - **Technician identity gap confirmed via direct schema check**: no synced roster exists in 1.0 —
93
+ identity resolution is live, per-call (`getUserByEmail`/`getUserSysIdByName` against SNOW's
94
+ `sys_user`), stored on two columns of the generic `TOGaDeskSupport.people` table
95
+ (`referenceId`=sys_id, `serviceNowEmail`), reconciled only by two **manual, unscheduled** scripts.
96
+ 2.0 target: a synced roster via the `servicenow` module; `Tickets_Users` (already exists in
97
+ `Client_Nycdoe`) is the direct 2.0 analog of `repair_order_technicians` — no new bridge table.
98
+ - **Explicit, deliberate decision NOT to poll ServiceNow's Task API as a trigger**, even though
99
+ ServiceNow's own PCS Vendor Integration Guide documents that as its default vendor-integration
100
+ pattern (confirmed to be the same `/api/nycd3/v2/data/tasks` endpoint 1.0's `listIncItems()`
101
+ already calls). Kept the three SNOW docs reviewed (Site Search API, PCS Vendor Guide, API
102
+ Migration Guide) purely as API-shape reference material.
103
+ - **`cmn_location` confirmed (via SNOW's own API doc) to be refreshed only nightly from LCGMS** —
104
+ decisively validates syncing locations locally over live-per-line SNOW calls (zero freshness lost).
105
+ - Produced a client-agnostic architecture doc (`2.0-generic-sftp-netsuite-integration-pattern.md`)
106
+ separating reusable platform pattern from NYCDOE-specific decisions, once Northwell made that
107
+ separation necessary. Confirmed as the right instinct by the user.
108
+ - Jeff confirmed independently (without prompting) that chained `postPost` interceptors are hard to
109
+ troubleshoot — validating this session's own prior recommendation to keep interceptor usage to
110
+ single, well-scoped, already-proven hops (SO creation → NetSuite dispatch; Ticket/ServiceRequest
111
+ state change → SNOW sync) rather than chaining the whole SFTP-to-NetSuite pipeline through them.
112
+
113
+ ## What did NOT work — DO NOT RETRY THESE
114
+
115
+ - **Assuming S3 alone (not a temp SFTP subfolder) should bridge 1.0 and 2.0 during parallel
116
+ testing** — wrong. The SFTP server is externally managed by someone else on the team and is
117
+ permanent infrastructure, not going away. The actual plan: 1.0's script copies the file to a temp
118
+ subfolder **on the same SFTP server** before its existing delete step; 2.0's worker points at that
119
+ temp folder pre-cutover and the real manufacturer folder post-cutover — **only a path config value
120
+ changes at cutover, not the mechanism**. This means 2.0 needs real SFTP connectivity
121
+ (phpseclib, credentials) built from day one — my initial "defer SFTP, read from S3 instead" idea
122
+ was based on the wrong assumption about where the bridge lived.
123
+ - **Recommending a separate "thin companion ledger table" referencing `Files.id` for workflow
124
+ status** — wrong per platform convention. Files↔ledger-status is a strict 1:1 relationship, not
125
+ one-to-many, so the idiomatic answer is custom fields (`c_status`, `c_vendorId`) directly on
126
+ `Files`, not a joined table. Self-corrected only after the user directly asked "what's the
127
+ platform's actual preference here" — should have applied the shape-driven test (1:1 → custom
128
+ field; 1:many/many:many → bridge table) the first time.
129
+ - **Claiming "using the ASN as sole source of truth" would fix the `WR270026786` dual-SO bug** —
130
+ imprecise on first pass. ServiceNow was never involved in producing that specific data (both
131
+ items came from the same ASN file); the actual fix is the Case-level SO grouping design, a
132
+ different decision than the ASN-first principle. Corrected when the user pushed back directly.
133
+ - **Concluding (from the `WR270029067`/`WR270026786` live examples) that the NetSuite SO/PO stamp
134
+ should move from Case-header level down to ServiceRequest/item level** — this flip-flopped twice
135
+ in one session (header → item-level → back to header) before landing on: header-level SO,
136
+ vendor-scoped POs underneath via the existing bridge table, with the observed 1.0 dual-SO
137
+ behavior attributed to the two-pipeline race rather than a genuine requirement. Should have
138
+ applied the PO-is-vendor-scoped / SO-is-customer-scoped distinction *before* jumping to revise a
139
+ locked-in design from one live data point.
140
+ - **MCP DB tool (`TOGa Database Integration`) disconnected mid-session** and the `/mcp` reconnect
141
+ notice was misleading — tools showed as "reconnected" in the system reminder but were not
142
+ actually reachable via `ToolSearch` or direct call for a period. Required a full VS Code restart
143
+ to actually restore connectivity. If this recurs: don't trust the reconnect notice alone — verify
144
+ with an actual tool call before relying on it.
145
+
146
+ ## Not tried yet (candidates for next session)
147
+
148
+ - Confirm whether `syncNetsuitePurchaseOrder`/`syncNetsuiteItemReceipt` are actually wired to a live
149
+ trigger (a Record Script row?) in production, or currently dormant — no caller found in
150
+ `api2`/`worker2` source; could be wired via DB-driven Record Script, which wouldn't show up in a
151
+ code grep.
152
+ - Confirm with whoever owns NetSuite/accounting whether one-SO-multiple-vendor-PO (e.g. Lexmark +
153
+ Belkin under one Case) is actually the desired business behavior, or whether always-separate SOs
154
+ per vendor is intentional — this is the one still-open piece of the SO/PO grain flip-flop above.
155
+ - Confirm per-vendor SFTP delete/move permissions (Mark manually verified his own SFTP credentials
156
+ can create/write/delete; still need to confirm 1.0's actual service-account credentials have the
157
+ same access, and separately confirm whether a server-side `rename()` is supported vs.
158
+ full download+reupload).
159
+ - Confirm actual Lenovo/Apple/Lexmark ASN file sizes before finalizing S3-staging vs.
160
+ re-fetch-by-name for fanned-out per-file worker jobs.
161
+ - Confirm whether ServiceNow's data model supports multiple RITMs per customer PO cleanly (relevant
162
+ to the item-grain RITM linkage decision).
163
+ - §4.1 (PO-suffix format variance defeating dedupe — `-RPL`/`-REPL`/etc.) is still fully open, not
164
+ addressed by any of this session's redesign work.
165
+ - **Incidents (INCs) — entirely deferred**, was scheduled for a "Wednesday" follow-up session; not
166
+ yet covered in any depth (only that "installs = RITM, not incidents" boundary is established).
167
+ - Design the exact schema for the per-client FTP-credentials table and the per-client
168
+ manufacturer-to-import table that Jeff's latest meeting notes call for (`manufacturerId` +
169
+ `ftpCredentialId`, "FTP ID... use for dynamic connections to FTPs").
170
+ - Design the concrete discovery-cron-per-manufacturer / fan-out-per-client mechanism for the
171
+ confirmed 1000+ client scale target — including whether `Core.PartnerApiIdentities` querying at
172
+ that volume needs any indexing/pagination consideration.
173
+ - Open question from Jeff's notes, unresolved: **do all manufacturers use the same ASN file format
174
+ across different clients** (e.g. Lenovo-to-Northwell vs. Lenovo-to-NYCDOE), or does format vary
175
+ per client/contract? Affects whether per-manufacturer parsers are safely shareable across clients
176
+ or need per-(client × manufacturer) overrides.
177
+ - Northwell-specific plan file deliberately not yet started (Northwell will get its own ClickUp
178
+ ticket; this session's scope stays DOE-only per the actual ticket title/user story).
179
+ - Design the exact trigger mechanism (interceptor vs. explicit call) for the outbound NetSuite
180
+ serial/inventory push once it's built.
181
+ - Design the unified generic "sync `Ticket`/`ServiceRequest` state to RITM" worker action discussed
182
+ across the hold-status-sync and proof-of-delivery conversations (replacing what are two separate,
183
+ polling-based scripts in 1.0 with one event-driven mechanism, multiple trigger points).
184
+
185
+ ## Current file state
186
+
187
+ | File | Status | Notes |
188
+ |------|--------|-------|
189
+ | `C:\Users\mhammontree\.claude\plans\true-80519-this-is-a-sunny-riddle.md` | Needs update | Main NYCDOE plan file, §1–11 (1.0 rundown, §7 redesign principles, §8 SNOW API wishlist, §10 concrete architecture outline). **Not yet updated** with: Case→ServiceRequest-only (no Ticket) model, the `Core.Partners`/`PartnerApiIdentities` credential mechanism, or Jeff's Phase-1-SFTP-scope + 1000+ client requirement. |
190
+ | `C:\Users\mhammontree\.claude\plans\true-80519-file-mapping.md` | Needs update | 1:1 mapping of every 1.0 file/table to its 2.0 status. Not yet revised for the no-Ticket model change. |
191
+ | `C:\Users\mhammontree\.claude\plans\2.0-generic-sftp-netsuite-integration-pattern.md` | Needs update | Client-agnostic pattern doc. Still references the earlier speculative "Integrations table" concept — needs correcting to the confirmed `Core.Partners`/`PartnerApiIdentities` mechanism, the no-Ticket model, and the 1000+ client scale requirement. |
192
+ | No code files | Unchanged | Entire session was research/design/planning; no implementation work performed. |
193
+
194
+ ## Decisions made
195
+
196
+ - **ASN-first (never trust SNOW/NetSuite to create/change our records)** — root cause of nearly all
197
+ 1.0 bugs is treating these external systems as reliable/always-in-sync. Rejected: SNOW's own
198
+ documented default vendor pattern (poll the Task API) — deliberately not adopted.
199
+ - **`Case → ServiceRequest`, no `Ticket`, for DOE** — matches 1.0's actual `MSO → repair_order`
200
+ two-tier shape exactly. Rejected: the original 3-tier `Case → ServiceRequest → Ticket` design from
201
+ earlier in this same session, abandoned once Jeff/Mark clarified DOE never had a Ticket-equivalent.
202
+ - **NetSuite SO/PO stamp stays at `Case` (header) level**, supporting multiple vendor-scoped POs per
203
+ SO via the existing bridge table. Rejected: moving the stamp to `ServiceRequest`/item level
204
+ (seriously considered and reverted within this session after the PO-is-vendor-scoped /
205
+ SO-is-customer-scoped distinction was raised).
206
+ - **Interceptor usage narrowed to single, proven hops** (SO creation → NetSuite dispatch;
207
+ state-change → SNOW sync). Rejected: chaining the entire SFTP-to-NetSuite pipeline through nested
208
+ `postPost` interceptors — Jeff independently flagged troubleshooting difficulty as the reason.
209
+ - **File ledger via custom fields on the existing `Files` table.** Rejected: a new dedicated ledger
210
+ table, and a thin companion table referencing `Files.id` — both unnecessary given the real
211
+ `checksum`/`sourceUrl` columns and the 1:1 relationship shape.
212
+ - **SFTP credential/client resolution via `Core.Partners`/`Core.PartnerApiIdentities`**, reusing the
213
+ exact pattern already proven in `worker2/Worker/Startech.php`. Rejected: inventing a new generic
214
+ "Integrations" table from scratch (an earlier speculative proposal this session).
215
+ - **Two independent, opt-in `dbchanges2` modules** (`servicenow`, `sftp-import`) rather than one
216
+ bundled module — Northwell needs one now, the other later. Rejected: a single combined module.
217
+ - **A separate, client-agnostic architecture document**, apart from the NYCDOE-specific plan file —
218
+ needed once Northwell entered the picture. Rejected: keeping everything in one NYCDOE-titled file.
219
+ - **Discovery cron scheduled per-manufacturer, not per-client**, fanning out per client via
220
+ `_Worker::runTask()` — required to hit the confirmed 1000+ client scale target without creating
221
+ 1000+ `Core.CronJobs` rows.
222
+ - **Mark authors his own pseudocode for Jeff's approval; this session's output (plan files) is raw
223
+ material only, never submitted directly** — standing instruction, repeated and honored throughout
224
+ (Mark is not permitted to submit Claude's own output as his proposal per his team's process).
225
+ - **This ticket's (TRUE-80519) formal scope stays DOE-only; Northwell gets its own future ClickUp
226
+ ticket** — the ticket's literal title/user story is DOE-specific; avoids risking a "Rework" cycle
227
+ from the ClickBot-enforced pseudocode-review gate on this ticket.
228
+
229
+ ## Blockers
230
+
231
+ None currently. The DB MCP tool had a transient disconnect mid-session (see "did NOT work" above)
232
+ but is confirmed working again as of the last successful query this session.
233
+
234
+ ## Exact next step
235
+
236
+ > Update `true-80519-this-is-a-sunny-riddle.md` and `2.0-generic-sftp-netsuite-integration-pattern.md`
237
+ > to reflect: (a) the `Case → ServiceRequest`-only model for DOE (no `Ticket`), (b) the
238
+ > `Core.Partners`/`Core.PartnerApiIdentities` credential-resolution mechanism (replacing the
239
+ > speculative "Integrations table" idea), and (c) Jeff's Phase 1 = SFTP scope decision with the
240
+ > explicit 1000+ client scalability requirement. Then design the concrete schema for the per-client
241
+ > FTP-credentials table and manufacturer-import table referenced in Jeff's latest meeting notes.
242
+
243
+ ---
244
+ _Saved by /session-save on 2026-08-28_
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "toga-ai",
3
- "version": "1.0.673",
3
+ "version": "1.0.675",
4
4
  "description": "TOGA Technology Team Claude Knowledge System — shared AI coding harness with skills, knowledge base CLI, and project installer for Claude Code.",
5
5
  "keywords": [
6
6
  "claude",