toga-ai 1.0.649 → 1.0.651

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -37,6 +37,8 @@
37
37
  | [Persona Name Translation (PersonaTranslations sidecar)](features/persona-name-translation.md) | Serves Persona **names** in multiple languages by adding a per-language **sidecar** table `PersonaTranslations`, reusing the platform's existing metadata-driven | _underscore/Model/Client/PersonaTranslation.php, dbchanges2/Client/2026-07-22b - PersonaTranslations.sql, dbchanges2/Core/2026-07-22a - PersonaTranslationsRecord.sql, dbchanges2/Client/2026-07-22c - PersonaTranslationsAcl.sql, dbchanges2/Client_CompassCanada/2026-07-22 - PersonaTranslationsFrench.sql, toga2-commerce/src/pages/Account/view/MySettingsView.tsx |
38
38
  | [Record Change Audit Log (Logs_<Client>.Record / RecordField) — reading a field's history](features/record-change-audit-log.md) | Every 2.0 client schema has a sibling **logs** schema `Logs_<Tenant>` (e.g. | _underscore/Model/Client/Logs/Record.php, _underscore/Model/Client/Logs/RecordField.php, _underscore/Model/Client/Logs/CustomRecordField.php |
39
39
  | [Recursive Item Fulfillments (upstream mirroring)](features/recursive-item-fulfillments.md) | In a multi-tier supply chain a sales order (SO) spawns a purchase order (PO) that becomes another SO downstream, and so on. | _underscore/Model/Client/ItemFulfillment.php, _underscore/Model/Client/ItemFulfillmentItem.php, _underscore/Model/Client/ItemFulfillmentItemUnit.php, _underscore/Model/Client/ItemFulfillmentPackage.php, _underscore/Model/Compass/AdvanceShippingNotice.php, dbchanges2/Core/2026-02-13 - 75601 - RecursiveItemFulfillmentCreation.sql, dbchanges2/Core/2026-06-04 - RecursiveItemFulfillmentPut.sql, dbchanges2/Client_Compass/2026-07-02a - FixSA133377TrackingSerialAndDuplicateIF.sql, dbchanges2/Client_Compass/2026-08-18 - FixSA135471HeroItemHalfQuantityFulfillment.sql |
40
+ | [Sales-order \"PO Number\" sourcing — one surface element, `_purchaseOrders`, tenant SQL underneath](features/sales-order-po-number-sourcing.md) | The sales-order modal's **PO Number** field is **not** tenant-split at the display layer. | _underscore/Model/Client/SalesOrder.php, _underscore/Model/Compass/SalesOrder.php, toga25-supply/src/pages/SalesOrders/helpers/surfaceBundleToTenantFields.ts, toga25-supply/src/pages/SalesOrders/view/SalesOrderRecordModalLayout/hooks/usePurchaseOrderDetails.ts, dbchanges2/Core/2026-08-25a - PoNumberDetailFieldPurchaseOrdersValueKey.sql, dbchanges2/Core/2026-07-23a - PoNumberDetailFieldValueKey.sql, dbchanges2/Client_Quad/2026-08-19a - PoDetailsButtonMostRecentPoSelection.sql |
41
+ | [SO↔PO bridge tables are TWO tables in OPPOSITE directions (upstream vs downstream)](features/sales-order-purchase-order-bridge-direction.md) | There are **two** bridge tables linking sales orders and purchase orders, and they mean **opposite things**. | _underscore/Model/Client/PurchaseOrders/SalesOrder.php, _underscore/Model/Client/SalesOrders/PurchaseOrder.php, _underscore/Model/Client/ItemFulfillment.php, _underscore/Model/Client/SalesOrder.php |
40
42
  | [Per-client sales-order status filter (Surface FILTER_SET → table meta `filterOptions`)](features/sales-order-status-filter-surface.md) | The status-filter dropdown on the sales-orders table is **per client**, driven by a Surface `FILTER_SET` rather than by the raw contents of the client's `SalesO | _underscore/Model/Client/TableView.php, _underscore/Model/Core/Surface.php, _underscore/Model/Client/SalesOrder.php, _underscore/Model/Quad/SalesOrder.php, _underscore/Model/Prudential/SalesOrder.php, toga-blox/src/components/Table/hooks/useFetchTablePageMeta.ts, toga-blox/src/api/types.ts, dbchanges2/Core/2026-08-24b - SalesOrderStatusFilterSurfaceSeed.sql, dbchanges2/Client_Compass/2026-08-24b - SalesOrderStatusFilterHides.sql, dbchanges2/Client_CompassCanada/2026-08-24b - SalesOrderStatusFilterHides.sql, dbchanges2/Client_Quad/2026-08-24b - SalesOrderStatusFilterHides.sql, dbchanges2/Client_Nychh/2026-08-24b - SalesOrderStatusFilterHides.sql |
41
43
  | [_String helpers — ASCII-safe HTML entity encoding (and the parseBetween trap)](features/string-html-entity-helpers.md) | `_String` is the 2.0 framework's static string utility class. | _underscore/String.php |
42
44
  | [Surface Resolver (_Model_Core_Surface::resolve — replaces Page::meta)](features/surface-resolver.md) | The runtime for the platform-wide **Surface** UI presentation layer: 9 `_underscore` models plus a cached resolver, `_Model_Core_Surface::resolve(&$api, string | _underscore/Model/Client/Language.php, dbchanges2/Core/2026-08-21 - SalesOrderDecisionSurfacesReseed.sql, dbchanges2/Core/2026-08-24 - RestoreApproveDenyRowActionsVisibility.sql, dbchanges2/Client_Quad/2026-08-24 - ProdPortApprovePoNumberEnabledRule.sql, dbchanges2/Client_Compass/2026-08-24 - ProdPortApprovalsGateAndApproveStepTwoRule.sql, dbchanges2/Client_CompassCanada/2026-08-24 - ProdPortApprovalsGateAndApproveStepTwoRule.sql, dbchanges2/Client_Nychh/2026-08-24 - FixInventoryGroupingsUnitsTopologyOverrideIds.sql, toga25-supply/src/App.tsx, toga25-supply/src/surface/useFetchSurfaceMeta.ts, toga25-supply/src/contexts/AuthContext.tsx, dbchanges2/Core/2026-07-21a - SalesOrderDecisionSummarySurfaceSeed.sql, dbchanges2/Core/2026-07-21b - SalesOrderDecisionActionSurfaceSeed.sql, dbchanges2/Core/2026-08-21a - SalesOrderDecisionSurfacesReseed.sql, dbchanges2/Client_Compass/2026-07-21c - SalesOrderDecisionSummaryTotalConcat.sql, dbchanges2/Client_Quad/2026-07-21c - SalesOrderDecisionSummaryOverride.sql, dbchanges2/Client_CompassCanada/2026-07-21c - SalesOrderDecisionSummaryTotalConcat.sql, dbchanges2/Client_Compass/2026-08-04a - ApprovalDetailsAssignedManagerPreferredStage.sql, _underscore/Model/Core/Surface.php, _underscore/Model/Client/AclRecordScript.php, _underscore/Model/Core/RecordScript.php, dbchanges2/Core/2026-06-30a - SurfaceMetaGroupAndSalesOrderSections.sql, dbchanges2/Client/2026-06-30a - SurfaceMetaGroupAcl.sql, dbchanges2/Client_Quad/2026-07-01a - GrantSurfacesMetaGroupScriptAcl.sql, dbchanges2/Client_CompassCanada/2026-07-01a - GrantSurfacesMetaGroupScriptAcl.sql, dbchanges2/Core/2026-06-29b - SurfaceMetaPublicReadAcl.sql, dbchanges2/Client/2026-06-29c - SurfaceRecordScriptAcl.sql, dbchanges2/Core/2026-06-29c - SurfaceDebugPhpMethodFix.sql, dbchanges2/Client_Compass/2026-07-15f - SalesOrderRecordActionsRemoveDeadConfigRuleOverrides.sql, dbchanges2/Core/2026-07-17h - Update - ClearApprovalsFilterButtonConfig.sql, dbchanges2/Client_Compass/2026-07-17a - SalesOrderApprovalActionsOverride.sql, dbchanges2/Client_Compass/2026-07-17b - SalesOrderApprovalsFilterButtonOverride.sql, dbchanges2/Client_CompassCanada/2026-07-17a - SalesOrderApprovalActionsOverride.sql, dbchanges2/Client_CompassCanada/2026-07-17b - SalesOrderApprovalsFilterButtonOverride.sql, dbchanges2/Client_Quad/2026-07-17a - SalesOrderApprovalActionsOverride.sql, dbchanges2/Client_Quad/2026-07-17b - SalesOrderApprovalsFilterButtonOverride.sql, _underscore/Model/Core/SurfaceElement.php, _underscore/Model/Core/Action.php, _underscore/Model/Core/Vocabulary.php, _underscore/Model/Core/VocabularyTerm.php, _underscore/Model/Core/Message.php, _underscore/Model/Client/SurfaceOverride.php, _underscore/Model/Client/MessageTranslation.php, _underscore/Model/Client/ThemeToken.php, _underscore/Model/Core/Page.php, api2/Component/Api/V2/V2.php |
@@ -0,0 +1,88 @@
1
+ ---
2
+ title: "Sales-order \"PO Number\" sourcing — one surface element, `_purchaseOrders`, tenant SQL underneath"
3
+ framework: "2.0"
4
+ repo: _underscore
5
+ project: _Underscore
6
+ client: shared
7
+ type: feature
8
+ status: active
9
+ updated: 2026-08-25
10
+ owners: [apeterson]
11
+ files:
12
+ - _underscore/Model/Client/SalesOrder.php
13
+ - _underscore/Model/Compass/SalesOrder.php
14
+ - toga25-supply/src/pages/SalesOrders/helpers/surfaceBundleToTenantFields.ts
15
+ - toga25-supply/src/pages/SalesOrders/view/SalesOrderRecordModalLayout/hooks/usePurchaseOrderDetails.ts
16
+ - dbchanges2/Core/2026-08-25a - PoNumberDetailFieldPurchaseOrdersValueKey.sql
17
+ - dbchanges2/Core/2026-07-23a - PoNumberDetailFieldValueKey.sql
18
+ - dbchanges2/Client_Quad/2026-08-19a - PoDetailsButtonMostRecentPoSelection.sql
19
+ related:
20
+ - ./sales-order-purchase-order-bridge-direction.md
21
+ - ./calculated-sql-fields.md
22
+ - ./surface-resolver.md
23
+ - ../../toga25-supply/features/surface-frontend.md
24
+ - ../../dbchanges2/features/surface-layer-schema.md
25
+ ---
26
+
27
+ ## Summary
28
+
29
+ The sales-order modal's **PO Number** field is **not** tenant-split at the display layer. As of
30
+ 2026-08-25 every tenant renders it the same way:
31
+
32
+ ```
33
+ Core.SurfaceElements id 7 (Order Details "PO Number", messages.label = salesOrder.field.poNumber)
34
+ → config.valueKey = "_purchaseOrders"
35
+ → FE adapter toga25-supply/src/pages/SalesOrders/helpers/surfaceBundleToTenantFields.ts
36
+ → renders the order's `_purchaseOrders` calculated field (comma-joined string)
37
+ ```
38
+
39
+ **All tenant variation lives underneath, in the SQL behind `_purchaseOrders`** — not in the
40
+ surface, not in the front end. This supersedes the earlier (TRUE-81190) understanding that the
41
+ *display path* itself was tenant-split.
42
+
43
+ ## How it works
44
+
45
+ ### The display binding
46
+ `SurfaceElements` id 7 carries `config = {"valueKey": "_purchaseOrders"}`. `isPersonaValue` is
47
+ deliberately **not** set: the value is a scalar comma-joined string, not a persona array.
48
+
49
+ `surfaceBundleToTenantFields.ts` exposes `PO_NUMBER_VALUE_KEY` and resolves it against the order
50
+ via `resolvePathInsensitive(valueKey)`, normalising an empty string to an em-dash.
51
+
52
+ ### The tenant layer (`_purchaseOrders`)
53
+ | Tenant | Implementation | Direction read |
54
+ |---|---|---|
55
+ | **Compass USA** | overrides in `_underscore/Model/Compass/SalesOrder.php:238` — a `UNION ALL` of its own vendor POs plus Office Depot's POs traversed SO→PO→SO→PO across the multi-tier chain | downstream **+ upstream** (only tenant reading upstream) |
56
+ | **Quad Graphics** | no override → base implementation | downstream only |
57
+ | **Compass Canada** | no `Model/Compasscanada/` directory exists at all → base implementation | downstream only |
58
+ | base (`_Model_Client_SalesOrder:348`) | the shared calculated field | **downstream only** (`SalesOrders_PurchaseOrders`) |
59
+
60
+ ### `poSelection` is a different thing
61
+ The `poSelection` config on the `sales-order-record-actions` surface picks **which** linked PO the
62
+ *editable* flow treats as primary (Enter-PO modal prefill, `approveButtonRule`, completion state).
63
+ Core default is `"first"`; Quad is overridden to `"mostRecent"` by
64
+ `Client_Quad/2026-08-19a`. It reads the **downstream join**, not `_purchaseOrders`. Do not conflate
65
+ the two — changing one does not move the other.
66
+
67
+ ## Gotchas
68
+
69
+ - **A surface-config change is only real if the FE actually honours the binding.** Before
70
+ 2026-08-25 the FE constant *matched* the surface's `valueKey` and then **discarded** it,
71
+ substituting a hardcoded field path. The config was read and thrown away — so shipping the
72
+ migration alone would have silently done nothing. Whenever you repoint a `valueKey`, grep the FE
73
+ for the old literal and confirm it resolves *through* the config, not around it.
74
+ - **Surface seed `id`s are deterministic; surface `uuid`s are not.** Seeds insert with `UUID()`, so
75
+ the element uuid differs per environment. A uuid quoted in an older migration will **not** match
76
+ the uuid in a live payload. Key surface updates on the literal `id` (here, `id = 7`).
77
+ - **Downstream-only base implementation is a trap for upstream tenants.** On an upstream-only
78
+ tenant `_purchaseOrders` resolves cleanly (no `EV-8`) and returns `null` forever. See
79
+ [SO↔PO bridge direction](./sales-order-purchase-order-bridge-direction.md) and
80
+ [NYCHH PO Number](../../../clients/nychh/features/po-number-upstream-direction.md).
81
+ - Dropping `isPersonaValue` is intentional; re-adding it will make the FE try to treat the string
82
+ as a persona array.
83
+
84
+ ## Change history
85
+ - 2026-08-25 — Repointed `SurfaceElements` id 7 to `_purchaseOrders` for all clients
86
+ (`Core/2026-08-25a`, reverting `Core/2026-07-23a`) and made the FE adapter honour the surface
87
+ `valueKey` instead of a hardcoded path; recorded the corrected per-tenant map, which supersedes
88
+ the tenant-split display understanding (apeterson)
@@ -0,0 +1,72 @@
1
+ ---
2
+ title: "SO↔PO bridge tables are TWO tables in OPPOSITE directions (upstream vs downstream)"
3
+ framework: "2.0"
4
+ repo: _underscore
5
+ project: _Underscore
6
+ client: shared
7
+ type: feature
8
+ status: active
9
+ updated: 2026-08-25
10
+ owners: [apeterson]
11
+ files:
12
+ - _underscore/Model/Client/PurchaseOrders/SalesOrder.php
13
+ - _underscore/Model/Client/SalesOrders/PurchaseOrder.php
14
+ - _underscore/Model/Client/ItemFulfillment.php
15
+ - _underscore/Model/Client/SalesOrder.php
16
+ related:
17
+ - ./sales-order-po-number-sourcing.md
18
+ - ./recursive-item-fulfillments.md
19
+ - ./fulfillable-item-propagation.md
20
+ - ../../api2/features/v2-rest-query-contract.md
21
+ - ../../../clients/nychh/features/po-number-upstream-direction.md
22
+ ---
23
+
24
+ ## Summary
25
+
26
+ There are **two** bridge tables linking sales orders and purchase orders, and they mean
27
+ **opposite things**. Both names are legal under the bridge-table naming convention (sibling
28
+ tables may be named in either order), and the two API routes differ only in word order — so
29
+ **nothing at the API surface tells you which direction you are querying**. Querying the wrong
30
+ one returns `200` with an empty array, which reads like a permission or data bug and is not.
31
+
32
+ | Table | Direction | Meaning | Model | API route | `Core.Records` id |
33
+ |---|---|---|---|---|---|
34
+ | `PurchaseOrders_SalesOrders` | **UPSTREAM** | a sales order created **FROM** a purchase order — the *customer's* PO | `_Model_Client_PurchaseOrders_SalesOrder` | `purchase-order-sales-orders` | 276 |
35
+ | `SalesOrders_PurchaseOrders` | **DOWNSTREAM** | a purchase order created **FROM** a sales order — a *vendor* PO | `_Model_Client_SalesOrders_PurchaseOrder` | `sales-order-purchase-orders` | 275 |
36
+
37
+ Read the name as **"`<source>`\_`<thing created from it>`"**. The *first* table name is the
38
+ origin; the second is what was generated from it.
39
+
40
+ ## How it works
41
+
42
+ - Both `Records` rows are `aclDatabase = CLIENT`, `childPolicy = MATCH_UPSERT`, so both are
43
+ per-tenant data and both accept nested writes that match-or-create the link row.
44
+ - Which direction a tenant actually populates depends on how their orders enter the platform:
45
+ - Customer sends TOGa a PO which becomes a sales order → **upstream** rows.
46
+ - TOGa raises vendor POs to fulfil a sales order → **downstream** rows.
47
+ - A multi-tier tenant (Compass USA) has **both**, chained SO→PO→SO→PO.
48
+ - The only in-code statement of the distinction today is a docblock at
49
+ `_underscore/Model/Client/ItemFulfillment.php:136-140`.
50
+ - The base `_purchaseOrders` calculated field on `_Model_Client_SalesOrder` (line ~348) reads the
51
+ **downstream** table only. See
52
+ [sales-order PO Number sourcing](./sales-order-po-number-sourcing.md).
53
+
54
+ ## Gotchas
55
+
56
+ - **An empty array is the symptom.** `200` / `totalRecordCount 0` with **no** `EZ-*` or `EV-*`
57
+ message is the signature of *"the join matched nothing"* — i.e. wrong direction — not ACL and
58
+ not a missing migration. A permission failure would be `EZ-1`; an unknown field `EV-8`/`EV-9`.
59
+ - **Check both directions before debugging anything else.** Count rows in
60
+ `PurchaseOrders_SalesOrders` and `SalesOrders_PurchaseOrders` for the order. If upstream = 1 and
61
+ downstream = 0, the caller is on the wrong route. This diagnosis is one query and saves hours.
62
+ - **The routes are near-homographs.** `sales-order-purchase-orders` vs
63
+ `purchase-order-sales-orders`. Read them aloud before wiring a front-end call.
64
+ - **Downstream-only assumptions are baked into shared code.** Any tenant whose POs are purely
65
+ upstream (verified: NYCHH; suspected: Compass Canada, which has no `Model/Compasscanada/`
66
+ overrides at all) silently gets `null`/empty from the shared downstream-reading helpers rather
67
+ than an error. Upstream-only tenants need an explicit upstream counterpart field.
68
+
69
+ ## Change history
70
+ - 2026-08-25 — Documented the two-direction bridge after an NYCHH "empty array" debug that was
71
+ purely a wrong-direction query; recorded route/model/record-id mapping and the empty-vs-EZ
72
+ diagnostic signature (apeterson)
@@ -8,3 +8,4 @@
8
8
  | [Auditing a client DB that drifted from its models (partially applied module migration)](workflows/client-schema-drift-audit.md) | A recurring 2.0 failure mode: **one client's database drifts from what the PHP models declare**, usually because a `_modules/<module>/` migration was applied to | dbchanges2/Client_Growrk/2026-05-28.sql, dbchanges2/Client_Growrk/2026-08-10c - GrowrkServiceRequestCustomFieldsCatchUp.sql, dbchanges2/Client_Growrk/2026-08-10d - GrowrkServiceRequestTypeAndDispositionSeeds.sql, dbchanges2/_modules/netsuite/2026-07-10a - UnitInventoryFields.sql, dbchanges2/Client_Growrk/2026-08-10 - GrowrkUnitInventoryFieldsCatchUp.sql, dbchanges2/Client_Growrk/2026-08-10b - GrowrkUnitItemDescriptionAcl.sql, dbchanges2/Client_Growrk/_modules.txt |
9
9
  | [Local vs prod MySQL config parity — why “it passed locally” is not evidence](workflows/local-vs-prod-mysql-config-parity.md) | Several migration failures that look like "prod-only bugs" are actually **per-machine MySQL server-configuration differences**. | |
10
10
  | [Repairing non-prod metadata drift (works in prod, broken in beta/dev-sandbox)](workflows/nonprod-metadata-drift-repair.md) | Almost all 2.0 platform behavior is **metadata** — `Core.Records`/`RecordFields`, `Core.RecordScripts`, `Core.ApiPayloadInterceptors`, and per-client `Acl*` row | dbchanges2/Core/2026-06-30a - ItemFulfillmentStageDefaultInterceptor.sql, dbchanges2/Client_Compass/2026-08-06 - RemoveBrokenSalesOrderItemPostPostInterceptor.sql, dbchanges2/Core/2026-07-16a - TrackingNumberSignatureTypeRecordField.sql, dbchanges2/_modules/netsuite/2026-07-10a - UnitInventoryFields.sql, dbchanges2/Client_Aig/_modules.txt, dbchanges2/Client/2026-07-22c - TrackingNumberMeasureIdsFieldPermission.sql, api2/Config/beta.ini, api2/Config/sandbox-dev.ini, dbchanges2/Client/2026-07-22a - EntitlementServiceAddressId.sql, dbchanges2/Client/2026-07-23a - EntitlementServiceAddressIdFieldPermission.sql |
11
+ | [Verifying whether a dbchanges2 migration actually ran in an environment](workflows/verifying-a-migration-ran.md) | dbchanges2 has no execution ledger you can query — "did this file run here?" has to be answered from the **rows the file would have produced**. | Client/2026-08-11b - SalesOrderPurchaseOrdersField.sql, Client/2026-08-13 - UnitsForItemsPO_SalesOrderItemsJoin.sql |
@@ -0,0 +1,58 @@
1
+ ---
2
+ title: "Verifying whether a dbchanges2 migration actually ran in an environment"
3
+ framework: "2.0"
4
+ repo: dbchanges2
5
+ project: Database Changes
6
+ client: shared
7
+ type: workflow
8
+ status: active
9
+ updated: 2026-08-25
10
+ owners: [apeterson]
11
+ files:
12
+ - Client/2026-08-11b - SalesOrderPurchaseOrdersField.sql
13
+ - Client/2026-08-13 - UnitsForItemsPO_SalesOrderItemsJoin.sql
14
+ related:
15
+ - ../architecture.md
16
+ - ../features/surface-layer-schema.md
17
+ ---
18
+
19
+ ## Summary
20
+
21
+ dbchanges2 has no execution ledger you can query — "did this file run here?" has to be answered
22
+ from the **rows the file would have produced**. This procedure gets a reliable answer, and names
23
+ the two ways the naive answer is wrong.
24
+
25
+ ## Steps
26
+
27
+ 1. **Preferred — match the hardcoded uuid.** If the migration inserts a literal v4 uuid, select on
28
+ it. Present ⇒ this file ran. This is the only check that identifies the *file* rather than the
29
+ *effect*.
30
+ 2. **Fallback — match the structural signature.** When the uuid column is unavailable or the row
31
+ was seeded with `UUID()`, select on the **exact column-value combination** the file writes.
32
+ Worked example — `Client/2026-08-13 - UnitsForItemsPO_SalesOrderItemsJoin.sql` was confirmed on
33
+ `Client_Nychh` by a `TableViewJoins` row with `sortOrder = 23`, `joinRecordId = 15`,
34
+ `joinOnRecordFieldId = 60`, `parentRecordFieldId = 368`, `type = OUTER`, on the tableview slug
35
+ `units-for-items-for-purchase-orders`.
36
+ 3. **For a guarded file, find its distinguishing side-effect.** Pick something in the file that a
37
+ pre-existing row would *not* already have — typically the ACL grant pair it adds. Worked example
38
+ — `Client/2026-08-11b - SalesOrderPurchaseOrdersField.sql` was confirmed on `Client_Nychh` by the
39
+ grant pair **exactly Base + API, `isWritable = 0`**; Compass already had `_purchaseOrders` as
40
+ `CustomRecordFields` id 77 under a *different* uuid, so the field row alone proved nothing.
41
+ 4. **State the conclusion as "the effect is present in env X"**, not "the file ran", unless step 1
42
+ succeeded.
43
+
44
+ ## Gotchas
45
+
46
+ - **A guarded `INSERT ... WHERE NOT EXISTS` is a NO-OP where the row already exists under a
47
+ different uuid.** Therefore: *row absent* ≠ "the file did not run", and *row present* ≠ "this
48
+ file created it". Both directions of the naive inference are unsound.
49
+ - Seed rows created with `UUID()` differ per environment, so cross-environment uuid comparison is
50
+ meaningless. Compare **ids and structure**, not uuids. (Same reason surface migrations key on the
51
+ literal `id` — see [Surface Layer Schema](../features/surface-layer-schema.md).)
52
+ - Fan-out folders (`Client/`) run per tenant, so a file can be present in one client schema and
53
+ genuinely absent in another. Always verify against the specific schema, not "the environment".
54
+
55
+ ## Change history
56
+ - 2026-08-25 — Captured the uuid-first / structural-signature-fallback procedure and the
57
+ guarded-insert unsoundness, after confirming `Client/2026-08-11b` and `Client/2026-08-13` on
58
+ `Client_Nychh` in sandbox-client (apeterson)
@@ -18,7 +18,7 @@
18
18
  | [Compass VIP Support Importer (worker2)](features/compass-vip-support-importer.md) | A worker2 action that ingests Compass's quarterly VIP spreadsheet and assigns each VIP user's support technician by setting `Users.c_supportedByUserId` in `Clie | worker2/Worker/Client/Compass/VipSupport.php |
19
19
  | [Creating Worker Actions](features/creating-worker-actions.md) | How to add a new callable Worker action — a PHP class whose `public static` methods are invoked as background jobs (via webhook, cron, or `_Worker::runTask()`). | worker2/Worker/, worker2/Controller/Index.php, _underscore/Worker.php |
20
20
  | [Cross-account AWS access for worker2 crons (_Component_Aws_Workloads)](features/cross-account-aws-access.md) | `_Component_Aws_Workloads` is the **standard, and only sanctioned, way any new worker2 cron obtains AWS access** — for any account, any region, any AWS SDK clie | worker2/Component/Aws/Workloads/Workloads.php, worker2/Config/production.ini, worker2/Worker/Infrastructure/CloudWatch.php, worker2/Worker/Monitor/Fleet.php, library/app/worker.php, library/app/cloud.php |
21
- | [Elastic Beanstalk health monitor → OneUptime (ElasticBeanstalkHealth)](features/elastic-beanstalk-health-monitor.md) | `_Worker_Infrastructure_CloudWatch::ElasticBeanstalkHealth()` is a worker2 cron that reads each Elastic Beanstalk (EB) environment's **enhanced-health** status | worker2/Worker/Infrastructure/CloudWatch.php |
21
+ | [Elastic Beanstalk health monitor → OneUptime (ElasticBeanstalkHealth)](features/elastic-beanstalk-health-monitor.md) | `_Worker_Infrastructure_CloudWatch::ElasticBeanstalkHealth()` is a worker2 cron that reads each Elastic Beanstalk (EB) environment's health and reports it to a | worker2/Worker/Infrastructure/CloudWatch.php |
22
22
  | [Elite Freshservice Sync (worker2)](features/elite-freshservice-sync.md) | `_Worker_Elite` processes Freshservice webhook events and syncs them into TOGA 2. | worker2/Worker/Elite.php, worker2/Config/dev-kmaramreddy-laptop.ini |
23
23
  | [Error Escalation Cron (Errors::Escalate → ClickUp / email)](features/error-escalation-cron.md) | `_Worker_Infrastructure_Errors::Escalate` (renamed from `SyncWithClickup`) is the sole owner of **escalation, de-escalation, ClickUp ticketing, reminders, busin | worker2/Worker/Infrastructure/Errors.php, worker2/Worker/Notification/Email.php, worker2/Worker/Notification/EmailTemplate.php, worker2/Worker/Client/True.php, worker2/Worker/Clickup/ErrorTask.php, worker2/Worker/Clickup.php, worker2/Controller/Index.php, worker2/Config/production.ini, _underscore/Model/Core/Logs/Issue.php, dbchanges2/Core/2026-07-30a - Error escalation cron job.sql, dbchanges2/Logs/2026-08-03a - Issue clickupPriority.sql, dbchanges2/Core/2026-08-04a - Error neglect digest cron job.sql |
24
24
  | [Error-Issue Auto-Resolution & Reopen (frequency-decay lifecycle)](features/error-issue-auto-resolution.md) | The error system could escalate and de-escalate an Issue's *urgency* but had no concept of an Issue being **resolved**. | worker2/Worker/Infrastructure/Errors.php, worker2/Worker/Clickup/ErrorTask.php, _underscore/Model/Core/Logs/Issue.php, tools/mvc/errors/get.php, tools/mvc/errors/issue/get.php, dbchanges2/Logs/2026-08-05a - Issue status baseline and auto-resolution.sql |
@@ -6,7 +6,7 @@ project: Worker
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-08-19
9
+ updated: 2026-08-25
10
10
  owners: [jcardinal]
11
11
  files:
12
12
  - worker2/Worker/Infrastructure/CloudWatch.php
@@ -19,209 +19,253 @@ related:
19
19
  ## Summary
20
20
 
21
21
  `_Worker_Infrastructure_CloudWatch::ElasticBeanstalkHealth()` is a worker2 cron that reads
22
- each Elastic Beanstalk (EB) environment's **enhanced-health** status and reports it to a
23
- OneUptime **Incoming Request** monitor, which pages the team. It is a **"dumb reporter,
24
- smart monitor"** heartbeat (the same Pattern-B contract as
22
+ each Elastic Beanstalk (EB) environment's health and reports it to a OneUptime **Incoming
23
+ Request** monitor, which pages the team. It is a **"dumb reporter, smart monitor"**
24
+ heartbeat (the same Pattern-B contract as
25
25
  [OneUptime push-metric monitors](./oneuptime-worker2-monitoring.md)): the worker decides
26
26
  pass/fail and POSTs decided **string tokens** (`alarm=HIGH/OK`, `probe=DEGRADED/OK`) that
27
27
  OneUptime string-matches — OneUptime cannot compare numbers on a pushed body.
28
28
 
29
- The load-bearing design decision this doc records is the **paging (alarm) criteria**: page
30
- on EB `Degraded`/`Severe`, but **suppress the common "an occasional HTTP 500 tripped
31
- Degraded/Severe" false alarm** unless 5xx errors actually dominate the request mix (by
32
- share) or reach a real absolute volume. The 5xx gate applies to **both** `Degraded` and
33
- `Severe`, and only when EB attributes the problem **solely** to 5xx — a mixed cause always
34
- pages.
29
+ The load-bearing design decision this doc records is the **paging (alarm) criteria**, and
30
+ the key **2026-08-25** change: the 5xx page is now driven by a **CloudWatch windowed
31
+ request-rate gate**, fully **decoupled from EB's own `HealthStatus`**. A non-5xx
32
+ `Degraded`/`Severe` still pages immediately off EB health; but whether a *5xx* problem pages
33
+ is decided by the true per-environment 5xx **rate over a trailing window** (default 5 min),
34
+ not by EB's ~10-second `ApplicationMetrics` snapshot.
35
+
36
+ **Why the change:** the old gate read EB `describeEnvironmentHealth`'s `ApplicationMetrics`,
37
+ a ~10-**second** window. At low traffic a single 500 in a near-empty bucket reads as 100%
38
+ 5xx and false-pages. This actually fired: worker-production paged "100% requests failing"
39
+ while the healthd log showed **16 500s / 516 req = 3.1%** over 17 minutes — the 100% was one
40
+ 500 alone in a 1-request 10s bucket. The windowed CloudWatch gate averages over minutes and
41
+ across instances, so a lone 500 can no longer dominate.
35
42
 
36
43
  **Critical behavior:** the method is **non-fatal and never throws** — it always returns a
37
44
  string. worker2 has **no DLQ and a 3600s SQS visibility timeout**, so any uncaught 500
38
45
  becomes a poison-message storm; an outer `catch(\Throwable)` is the backstop and the push
39
- itself runs non-fatal.
46
+ itself runs non-fatal. CloudWatch read failures are caught **internally** and return `null`.
40
47
 
41
48
  ## Key files / entry points
42
49
 
43
50
  - `worker2/Worker/Infrastructure/CloudWatch.php` — `_Worker_Infrastructure_CloudWatch`,
44
51
  action `ElasticBeanstalkHealth`, dispatched via
45
52
  `_Worker::runTask('Infrastructure/CloudWatch/ElasticBeanstalkHealth', {awsAccountId,
46
- environments, oneuptimeUrl})`.
53
+ environments, oneuptimeUrl, …thresholds})`.
54
+ - New private `fetchWindowedRequestCounts()` reads the CloudWatch metrics (below).
47
55
  - AWS access is obtained through `_Component_Aws_Workloads::assumeCredentials()` (STS-assume
48
56
  a read-only role per account, one assume reused across regions) — see
49
57
  [Cross-account AWS access](./cross-account-aws-access.md), which also documents the
50
- `environments` region→env-names map parameter shape.
58
+ `environments` region→env-names map parameter shape. The **same assumed credentials** build
59
+ both the `ElasticBeanstalkClient` and the new `CloudWatchClient`.
51
60
  - `oneuptimeUrl` arrives as a **cron parameter and is a push credential** — never log it and
52
61
  never record its value in a doc (see gotchas).
53
62
 
54
63
  ## Alarm / paging criteria (the decision logic)
55
64
 
56
- Per environment, `alarm=HIGH` (page) when the EB `HealthStatus` is `Degraded` **or**
57
- `Severe`. `Suspended` never pages.
65
+ Two independent paths now feed `alarm=HIGH`:
58
66
 
59
- The **5xx gate applies to both `Degraded` and `Severe`.** A non-5xx or mixed-cause
60
- `Degraded`/`Severe` — and any env with no stated cause — **always pages**; the gate only
61
- ever suppresses an env EB attributes **solely** to 5xx.
67
+ 1. **EB-health path (non-5xx).** EB `HealthStatus` `Degraded` or `Severe` that is **not**
68
+ attributable solely to 5xx (latency, instances down, a failed deploy, a mixed cause, or
69
+ no stated cause) **pages immediately** — unchanged, still driven off EB health.
70
+ `Suspended` never pages. `causesAttributeSolelyTo5xx()` is what tells a solely-5xx event
71
+ apart from everything else.
72
+ 2. **CloudWatch windowed 5xx-rate path.** For a solely-5xx `Degraded`/`Severe`, the page is
73
+ decided **not** by EB but by the true windowed rate from CloudWatch. Given the window's
74
+ `requests` (ApplicationRequestsTotal) and `fivexx` (ApplicationRequests5xx), page when:
75
+ - `fivexx >= min5xxAbsolute` **(absolute-volume floor)** — pages regardless of share, or
76
+ - `requests >= minRequests` **AND** `sharePercent >= min5xxPercent` **(share gate)** — the
77
+ `minRequests` floor keeps a near-idle window off the share gate so a lone 500 can't page.
62
78
 
63
- - **Non-5xx or mixed-cause `Degraded`/`Severe`** (latency, instances down, a failed deploy,
64
- or a 5xx trickle *alongside* a real non-5xx failure) **always pages.**
65
- - **Solely-5xx `Degraded`/`Severe` pages only on a genuine flood** — either:
66
- - the **absolute** 5xx count in the window is at/above `HTTP_5XX_ALARM_ABSOLUTE_COUNT`
67
- (default 20), **or**
68
- - the **5xx share** of requests is at/above `HTTP_5XX_ALARM_RATIO` (default 0.80).
69
- - **Zero requests in the window → no page** (nothing observed; also avoids divide-by-zero).
79
+ Because the 5xx decision reads CloudWatch directly, a real 5xx flood pages **even if EB
80
+ hasn't escalated yet**, and a benign low-traffic 5xx blip does not page even though EB flags
81
+ `Degraded`.
70
82
 
71
- **Why the gate now covers `Severe` too:** EB escalates an environment to `Severe` precisely
72
- when the 5xx share is high, so real 5xx floods reach `Severe` *first* and never touch the
73
- gated `Degraded` branch. When the gate applied only to `Degraded`, the 80% threshold was
74
- effectively dead — a real 50%-5xx `Severe` event paged despite the setting. Gating both
75
- tiers makes the threshold actually govern.
83
+ ### Cron parameters (replaces the former class constants)
76
84
 
77
- **Why "solely" 5xx, not "any" 5xx cause:** a mixed-cause env (a 5xx trickle plus a real
78
- non-5xx failure such as a failed deploy) must not be silenced by a sub-threshold ratio. The
79
- ratio gate only applies when **every** stated cause references 5xx (and there is at least
80
- one); anything else pages.
85
+ The four thresholds are **optional cron parameters** spread as **named arguments** from
86
+ `CronJobs.parameters`, with method-signature defaults:
81
87
 
82
- **Why an absolute floor in addition to the ratio:** the ratio alone is magnitude-blind —
83
- 60k failing out of 100k reads as 60% and would be suppressed, though it is plainly a flood.
84
- `HTTP_5XX_ALARM_ABSOLUTE_COUNT` pages such an env regardless of share; tune it to the busiest
85
- monitored environment's request volume.
88
+ | Parameter | Default | Meaning |
89
+ |---|---|---|
90
+ | `windowMinutes` | `5` | Trailing window (minutes) over which the CloudWatch rate is summed |
91
+ | `minRequests` | `100` | Minimum windowed request count before the **share** gate applies |
92
+ | `min5xxPercent` | `20.0` | 5xx share (percent) at/above which a solely-5xx env pages, once `minRequests` is met |
93
+ | `min5xxAbsolute` | `20` | Absolute windowed 5xx count at/above which a solely-5xx env pages regardless of share |
86
94
 
87
- ### Fail-safe: a blind read pages, never reads healthy
95
+ Removed this iteration: `shouldPageForHealth()` and the old
96
+ `HTTP_5XX_ALARM_RATIO` / `HTTP_5XX_ALARM_ABSOLUTE_COUNT` constants (superseded by the
97
+ windowed gate and the four parameters).
88
98
 
89
- If the health read (`describeEnvironmentHealth`) throws `AwsException` for an environment,
90
- the monitor sets **`alarm=HIGH`** for that environment (a blind read must never masquerade
91
- as healthy) **and** sets **`probe=DEGRADED`** (coverage gap, distinct from a measured-bad
92
- signal — same alarm-vs-probe split as the multi-client monitors in
93
- [OneUptime push-metric monitors](./oneuptime-worker2-monitoring.md)).
99
+ ### Per-monitor threshold tuning (one cron watches many environments)
94
100
 
95
- ### Configurable constants (top of the class)
101
+ A single `ElasticBeanstalkHealth` cron watches environments of **very different scale**, so
102
+ thresholds are tuned per monitor. Measured 3h baselines: api-production-1 ~922–1737 req/5min,
103
+ apiproxy ~400–1650, but eu-west-1 envs only ~40–90 req/5min; 5xx was 0 across all.
96
104
 
97
- | Constant | Default | Meaning |
98
- |---|---|---|
99
- | `HEALTH_ALARM_STATUSES` | `['Degraded','Severe']` | HealthStatus values that page |
100
- | `HTTP_5XX_ALARM_RATIO` | `0.80` | 5xx share (0.0–1.0) at/above which a solely-5xx `Degraded`/`Severe` pages |
101
- | `HTTP_5XX_ALARM_ABSOLUTE_COUNT` | `20` | Absolute 5xx count in the window at/above which a solely-5xx `Degraded`/`Severe` pages regardless of share |
102
- | `HTTP_5XX_CAUSE_MARKERS` | `['5xx','http 5']` | Case-insensitive substrings identifying a 5xx-attributed EB `Cause` |
103
-
104
- ### Implementation shape
105
-
106
- - `shouldPageForHealth(healthStatus, causes, applicationMetrics): [bool, ?float]` — returns
107
- the page decision and the observed 5xx ratio.
108
- - `causesAttributeSolelyTo5xx(causes): bool` — true only when **every** EB `Cause` references
109
- 5xx **and there is at least one** cause. Replaces the former `causesAttributeTo5xx()`, which
110
- returned true if *any* cause mentioned 5xx and so let a mixed-cause env be gated.
111
- - `causeIsFivexx(cause): bool` — helper that substring-matches `HTTP_5XX_CAUSE_MARKERS`
112
- against a single EB `Cause` string (case-insensitive); `causesAttributeSolelyTo5xx` requires
113
- it to hold for all causes.
114
- - The ratio is computed as `StatusCodes.Status5xx / RequestCount` from `ApplicationMetrics`
115
- (raw counts — see below), **not** by parsing the percentage out of the `Causes` text. The
116
- absolute-count floor reads `StatusCodes.Status5xx` directly.
117
- - `describeEnvironmentHealth` is called with
118
- `AttributeNames = ['HealthStatus','Status','Color','Causes','ApplicationMetrics']`.
119
- - Per-environment report + OneUptime payload now include the per-env `alarm` token and the
120
- observed `fivexxRatio`.
121
-
122
- ## How EB enhanced-health is read (durable AWS reference)
123
-
124
- Facts verified against current AWS docs; they inform why the logic above is shaped as it is.
105
+ - **Worker + worker 1.0 monitors** keep the **method defaults** (`100 / 20% / 20`) — the low
106
+ absolute floor suits ~150 req/5min traffic.
107
+ - **API and API-Proxy monitors** (which each also include the low-traffic eu-west-1 env) are
108
+ overridden to **`windowMinutes=5, minRequests=30, min5xxPercent=20, min5xxAbsolute=100`** —
109
+ the lower request floor keeps the small env on the share gate, while the higher absolute
110
+ floor stops the busy envs paging on minor blips (20 scattered 500s on ~1300 req = <2%). The
111
+ override is applied by the dbchanges2 migration below.
112
+
113
+ ### Fail-safe: an unmeasurable solely-5xx event still pages
114
+
115
+ The windowed metrics are **per-instance only** (see below), and some envs never publish them
116
+ (the 1.0 worker env `agilant-worker` publishes only `EnvironmentHealth`). For such an env the
117
+ gate has **no data** and returns **zeros** (not `null`). To avoid silently missing a real
118
+ flood, when EB reports a **solely-5xx `Degraded`/`Severe`** but the window is **empty
119
+ (unmeasurable)**, the monitor **pages** with gate token **`health-5xx-unmeasured`** rather
120
+ than trusting the empty read.
121
+
122
+ A CloudWatch read *failure* (as opposed to empty data) sets **`probe=DEGRADED`** (coverage
123
+ gap), distinct from a measured-bad `alarm=HIGH` — same alarm-vs-probe split as the
124
+ multi-client monitors in [OneUptime push-metric monitors](./oneuptime-worker2-monitoring.md).
125
+
126
+ ## Reading the windowed 5xx rate from CloudWatch (durable AWS reference)
127
+
128
+ - **EB application-request metrics are PER-INSTANCE ONLY.** `ApplicationRequestsTotal`,
129
+ `ApplicationRequests5xx` (and `…2xx/3xx/4xx`) are published under dimensions
130
+ `{EnvironmentName, InstanceId}` in namespace `AWS/ElasticBeanstalk`. **There is no
131
+ environment-only rollup.** You must sum across instances yourself with metric-math over a
132
+ search expression:
133
+
134
+ ```
135
+ SUM(SEARCH('{AWS/ElasticBeanstalk,EnvironmentName,InstanceId} MetricName="ApplicationRequestsTotal" EnvironmentName="<env>"','Sum',<periodSeconds>))
136
+ ```
137
+
138
+ The `SEARCH` auto-discovers every instance's series and `SUM` folds them into one, so the
139
+ window total is correct as instances come and go. Query total and 5xx as two expressions
140
+ over `windowMinutes`.
141
+ - **Not every environment publishes these metrics.** worker-production, api-production-*, and
142
+ apiproxy-production-* all publish them; the **1.0 worker env `agilant-worker` publishes
143
+ ONLY `EnvironmentHealth`** (no app-request metrics) → the gate reads zeros there → the
144
+ `health-5xx-unmeasured` fail-safe above covers it.
145
+ - **IAM already sufficient.** The `WorkloadsRuntime` role grants `cloudwatch:*` in **both**
146
+ account `654654170868` (worker/api) and `502614707982` (legacy worker), so no IAM change
147
+ was needed for the CloudWatch reads.
148
+
149
+ ### EB enhanced-health facts (still relevant — the EB-health path)
125
150
 
126
151
  - **HealthStatus severity ladder:** `Ok → Info → Warning → Degraded → Severe` (plus the grey
127
- states `Pending`/`Unknown`/`Suspended`/`NoData`). `Degraded` is the "high failure" tier
128
- and is **routinely tripped by benign transients** (Auto Scaling scale-up, a mid-deploy
129
- dip), which is exactly why `Degraded` alone is noisy. `Severe` = "very high failure /
130
- environment effectively not serving."
131
- - **`ApplicationMetrics` returns RAW COUNTS, not percentages** — despite the API-reference
132
- prose saying "percentage"/"per second". Fields: `Duration` (Integer seconds, usually 10),
133
- `RequestCount` (Integer, **total** requests over the window),
134
- `StatusCodes.{Status2xx,Status3xx,Status4xx,Status5xx}` (Integer counts). AWS's own
135
- example: 2xx 3391 + 5xx 843 = RequestCount 4234. With no traffic, `RequestCount=0` and
136
- `StatusCodes` may be **absent** → treat as "no data," not "zero failures." The 5xx-ratio
137
- denominator is `RequestCount` (the authoritative total).
152
+ states `Pending`/`Unknown`/`Suspended`/`NoData`). `Degraded` is routinely tripped by benign
153
+ transients (Auto Scaling scale-up, a mid-deploy dip), which is why a solely-5xx `Degraded`
154
+ is now gated on the CloudWatch rate rather than paged outright.
138
155
  - **Why an env is Degraded comes from `Causes`:** EB `Causes` strings literally contain e.g.
139
156
  `"19.9 % of the requests are failing with HTTP 5xx."` A substring check reliably
140
- **classifies** 5xx-driven vs. not (fail-safe: a wording change → treated as a normal
141
- `Degraded` → pages). The **ratio itself must come from `ApplicationMetrics`**, never by
142
- parsing the number out of `Causes`.
143
- - **A failed application DEPLOYMENT is NOT reliably reflected in HealthStatus** — it can read
144
- `Warning`/`Degraded`, or fail before any red request-failure signal. The authoritative
145
- structured deploy signal is per-instance `describeInstancesHealth` →
146
- `Deployment.Status ∈ {'In Progress','Deployed','Failed'}`. (This detector was evaluated and
147
- removed this iteration — see Design history.)
148
- - **IAM:** `describeEnvironmentHealth` / `describeInstancesHealth` / `describeEvents` are all
149
- covered by the managed `AWSElasticBeanstalkReadOnly` policy and require enhanced health
150
- enabled. No per-call charge.
157
+ **classifies** solely-5xx vs. not; a wording change fails safe (treated as non-5xx →
158
+ pages). `describeEnvironmentHealth` is called with `AttributeNames` including
159
+ `HealthStatus`, `Status`, `Color`, `Causes`.
160
+ - **A failed application DEPLOYMENT is NOT reliably reflected in HealthStatus** — the
161
+ authoritative structured deploy signal is per-instance `describeInstancesHealth` →
162
+ `Deployment.Status`. (A deploy that fails without ever degrading health remains an accepted
163
+ residual gap.)
164
+ - **IAM:** `describeEnvironmentHealth` is covered by the managed `AWSElasticBeanstalkReadOnly`
165
+ policy and requires enhanced health enabled.
166
+
167
+ ## Deploy order (mandatory) — code BEFORE the migration
168
+
169
+ worker2 code **must** deploy **before** the dbchanges2 threshold migration runs.
170
+
171
+ - new-code + old-params → **fine** (the params are optional with defaults).
172
+ - **old-code + new-params → FATAL.** The dispatcher spreads `CronJobs.parameters` as **named
173
+ arguments**; an old method signature that lacks `windowMinutes`/`minRequests`/… throws
174
+ **"Unknown named parameter"** and **every run fails** → the monitor goes dark → OneUptime
175
+ raises an offline incident.
176
+
177
+ Correct order: **1) deploy worker2 → 2) apply the SQL → 3) add the OneUptime criterion.**
178
+
179
+ ## dbchanges2 migration (the API/API-Proxy override)
180
+
181
+ `dbchanges2/Core/2026-08-25a - ApiApiproxy Beanstalk Health 5xx Thresholds.sql` applies the
182
+ API/API-Proxy override with a JSON_SET that **adds only the four keys**, preserving
183
+ `awsAccountId`/`environments`/`oneuptimeUrl`:
184
+
185
+ ```sql
186
+ UPDATE CronJobs
187
+ SET parameters = JSON_SET(parameters,
188
+ '$.windowMinutes', 5, '$.minRequests', 30,
189
+ '$.min5xxPercent', 20, '$.min5xxAbsolute', 100)
190
+ WHERE action = 'Infrastructure/CloudWatch/ElasticBeanstalkHealth'
191
+ AND name IN (<the two API monitor names>);
192
+ ```
193
+
194
+ `Core.CronJobs.parameters` is confirmed a native `JSON` column, so `JSON_SET` merges rather
195
+ than overwrites.
151
196
 
152
197
  ## OneUptime monitor config (external, not code)
153
198
 
154
199
  Paging depends on the OneUptime Incoming-Request monitor's **string-match** criteria on the
155
- POSTed body — `Contains "alarm":"HIGH"` (and, if desired, `Contains "probe":"DEGRADED"`).
156
- This is external OneUptime configuration, not worker2 code.
157
-
158
- ## Design history — rejected direction (this session)
200
+ POSTed body — `Contains "alarm":"HIGH"` pages.
159
201
 
160
- An earlier iteration set `HEALTH_ALARM_STATUSES = ['Severe']` (dropping `Degraded` entirely)
161
- and **added a separate deployment-failure detector** (`describeInstancesHealth` →
162
- `Deployment.Status = 'Failed'`) to still catch deploy failures. That was **reverted** after a
163
- scope change: the team decided they **do** want to page on `Degraded` (to catch a flood of
164
- 500s and non-500 degradations), with only the occasional-500 case suppressed via the 80%
165
- ratio. The deployment-failure detector was **removed** — deploy failures that surface as
166
- `Degraded` are now covered by the `Degraded` alarm. **Residual gap accepted for this
167
- iteration:** a deploy that fails fast *without ever degrading health* is not caught.
202
+ **Known alerting gap (recommended fix):** the current monitors match `"status":"error"`,
203
+ `"alarm":"HIGH"`, and heartbeat-gap rules, but **nothing matches `"probe":"DEGRADED"`**. So a
204
+ CloudWatch-metrics **read failure while EB reads Ok** (body = `alarm:OK, probe:DEGRADED,
205
+ status:reporting`) is **invisible to alerting**. Add a criterion matching
206
+ `Contains "probe":"DEGRADED"` → a **Degraded-severity, auto-resolving** incident, placed
207
+ **before** the "online" recovery rule (which also matches `alarm:OK`). The code already emits
208
+ the `probe` token; this is a OneUptime config change the developer applies.
168
209
 
169
210
  ## Known gaps / follow-ups (not done)
170
211
 
171
- - **Per-region client construction is not individually guarded.** Each
172
- `new ElasticBeanstalkClient` is not in its own try/catch, so a bad region/creds aborts the
173
- **whole run** via the outer backstop instead of paging just that region's envs as blind.
174
- Candidate follow-up now that blind reads page.
175
- - **No first-party test harness exists in worker2** (all tests are vendor/).
176
- `shouldPageForHealth()` and `causesAttributeSolelyTo5xx()` are pure and ideal to unit-test
177
- (ratio at exactly 0.80; absolute count 19 vs 20; `Severe` over a solely-5xx cause; a
178
- mixed 5xx + non-5xx cause; empty causes; null metrics; zero requests) — deferred pending a
179
- harness.
212
+ - **No first-party test harness in worker2** (all tests are vendor/). The gate decision and
213
+ `causesAttributeSolelyTo5xx()` are pure and ideal to unit-test (share at exactly
214
+ `min5xxPercent`; absolute at `min5xxAbsolute-1` vs `min5xxAbsolute`; `requests` just below
215
+ `minRequests`; empty/zero window; mixed 5xx + non-5xx cause) — deferred pending a harness.
216
+ - **Deploy fails without degrading health** — still not caught (accepted residual).
180
217
 
181
218
  ## Gotchas / known issues
182
219
 
220
+ - **Deploy worker2 code BEFORE the threshold migration** — old-code + new-params is fatal
221
+ ("Unknown named parameter" on every run → monitor dark). See Deploy order.
183
222
  - **Keep the method non-fatal — never let it throw.** worker2 has no DLQ and a 3600s SQS
184
- visibility timeout, so any uncaught 500 becomes a poison-message storm. The push runs
185
- non-fatal and an outer `catch(\Throwable)` is the backstop.
186
- - **A blind health read must page, not read healthy** — an `AwsException` on
187
- `describeEnvironmentHealth` sets `alarm=HIGH` + `probe=DEGRADED` for that env.
188
- - **`ApplicationMetrics` is raw counts, not percentages** — divide `Status5xx` by
189
- `RequestCount`; `RequestCount=0`/absent `StatusCodes` means "no data," not "zero failures."
190
- - **Classify 5xx-attribution from `Causes` text, but take the ratio from `ApplicationMetrics`
191
- — never parse the percentage out of `Causes`.**
223
+ visibility timeout; an outer `catch(\Throwable)` is the backstop and CloudWatch failures are
224
+ caught internally (return `null`).
225
+ - **EB app-request metrics are per-instance only** — there is no env rollup; you must
226
+ `SUM(SEARCH(...))` across instances or you undercount.
227
+ - **An empty window ≠ zero failures.** An env with no published app-request metrics (e.g. 1.0
228
+ `agilant-worker`) reads zeros; a solely-5xx EB event over an empty window pages via the
229
+ `health-5xx-unmeasured` fail-safe rather than reading healthy.
230
+ - **`minRequests` floor guards the near-idle edge** the old code couldn't — the share gate is
231
+ suppressed until the window has real volume, so a lone 500 in a quiet window no longer pages.
232
+ - **OneUptime has no `probe:DEGRADED` criterion** — a metrics-read failure while EB is Ok is
233
+ currently unalerted until that criterion is added.
192
234
  - **`oneuptimeUrl` is a push credential** — it arrives as a cron parameter; never log it or
193
235
  record its value in a doc.
194
- - **Near-idle ratio edge (residual, accepted).** With the tiny-sample guard removed, a
195
- near-idle env whose window holds a tiny all-error sample (e.g. its only request being a
196
- 500 = 100%) still trips the ratio and pages. Optional future guard: require a minimum 5xx
197
- count before the ratio applies (the absolute floor gates high volume, not this low-volume
198
- edge).
199
236
 
200
237
  ## Change history
238
+ - 2026-08-25 — Replaced the EB `ApplicationMetrics` 5xx gate (a ~10s window that false-paged
239
+ "100% failing" on a lone 500 in a near-empty bucket — real worker-production incident: 16
240
+ 500s / 516 req = 3.1% over 17 min) with a **CloudWatch windowed request-rate gate**. Reads
241
+ true per-env `ApplicationRequestsTotal`/`ApplicationRequests5xx` over a trailing window
242
+ (default 5 min) via `SUM(SEARCH('{AWS/ElasticBeanstalk,EnvironmentName,InstanceId}…','Sum',
243
+ period))` (these metrics are per-instance only — no env rollup). The 5xx page is now
244
+ **decoupled from EB `HealthStatus`**: pages when `5xx >= min5xxAbsolute` OR `requests >=
245
+ minRequests AND share >= min5xxPercent`; a non-5xx `Degraded`/`Severe` still pages
246
+ immediately. Added four optional cron params (`windowMinutes=5, minRequests=100,
247
+ min5xxPercent=20.0, min5xxAbsolute=20`) spread as named arguments; removed
248
+ `shouldPageForHealth()` and the `HTTP_5XX_ALARM_RATIO`/`HTTP_5XX_ALARM_ABSOLUTE_COUNT`
249
+ constants. Added a `health-5xx-unmeasured` fail-safe (solely-5xx over an empty/unpublished
250
+ window pages rather than reading healthy). Per-monitor tuning: worker + worker-1.0 keep
251
+ defaults; API/API-Proxy overridden to `30 / 20% / 100` (dbchanges2
252
+ `2026-08-25a - ApiApiproxy Beanstalk Health 5xx Thresholds.sql`, JSON_SET on the native
253
+ `CronJobs.parameters` JSON column). Recorded the mandatory deploy order (code before the
254
+ migration — old-code+new-params throws "Unknown named parameter") and the OneUptime gap
255
+ (no `probe:DEGRADED` criterion, so a metrics-read failure while EB reads Ok is unalerted).
256
+ WorkloadsRuntime already grants `cloudwatch:*` in accounts 654654170868 and 502614707982 —
257
+ no IAM change. (jcardinal)
201
258
  - 2026-08-19 — Overhauled the 5xx alarm gating in `shouldPageForHealth()`. The 80% share
202
- gate now applies to **both** `Degraded` and `Severe` (was `Degraded`-only, which left the
203
- threshold dead because EB escalates real 5xx floods straight to `Severe` — a 50%-5xx
204
- `Severe` had been paging despite the 0.80 setting). This Severe-gating trade-off was made
205
- deliberately with an independent architecture second opinion on record (cto returned
206
- DISAGREE-WITH-ALTERNATIVE; developer accepted). Narrowed 5xx classification: replaced
207
- `causesAttributeTo5xx()` (any cause mentions 5xx) with `causesAttributeSolelyTo5xx()` (every
208
- cause references 5xx, ≥1) + a `causeIsFivexx()` helper, so a mixed cause (5xx trickle beside
209
- a real non-5xx failure) always pages — a security-review gap the Severe change would have
210
- widened. Added `HTTP_5XX_ALARM_ABSOLUTE_COUNT` (20): a solely-5xx env pages regardless of
211
- share once the absolute 5xx count reaches the floor (closes the ratio's magnitude-blindness,
212
- e.g. 60k of 100k). Removed `HTTP_5XX_MIN_REQUEST_COUNT` (was 20) and its tiny-sample
213
- page-anyway guard — a small number of 5xx is now judged purely on the ratio; below the
214
- absolute floor, page only if share ≥ 0.80, and zero requests → no page. Accepted residual:
215
- a near-idle env with a tiny all-error sample still trips the ratio. (jcardinal)
259
+ gate now applied to **both** `Degraded` and `Severe` (was `Degraded`-only, which left the
260
+ threshold dead because EB escalates real 5xx floods straight to `Severe`). Narrowed 5xx
261
+ classification: replaced `causesAttributeTo5xx()` (any cause mentions 5xx) with
262
+ `causesAttributeSolelyTo5xx()` (every cause references 5xx, ≥1). Added
263
+ `HTTP_5XX_ALARM_ABSOLUTE_COUNT` (20). Removed `HTTP_5XX_MIN_REQUEST_COUNT` tiny-sample
264
+ guard. (jcardinal)
216
265
  - 2026-08-17 — Created. Documented the `ElasticBeanstalkHealth` alarm/paging criteria (page
217
- on `Degraded`/`Severe`; suppress a 5xx-driven `Degraded` unless the 5xx share ≥
218
- `HTTP_5XX_ALARM_RATIO` 0.80 with an `HTTP_5XX_MIN_REQUEST_COUNT` 20 tiny-sample guard;
219
- `Severe` and non-5xx `Degraded` always page), the blind-read fail-safe (`alarm=HIGH` +
220
- `probe=DEGRADED`), the four configurable constants, and the `shouldPageForHealth` /
221
- `causesAttributeTo5xx` helper shape. Captured durable EB enhanced-health facts (severity
222
- ladder; `ApplicationMetrics` returns raw counts not percentages; `Causes` classifies but
223
- `ApplicationMetrics` sets the ratio; failed deploys aren't reliably in HealthStatus —
224
- `describeInstancesHealth.Deployment.Status` is authoritative; `AWSElasticBeanstalkReadOnly`
225
- covers the reads). Recorded the reverted `['Severe']`-only + deployment-failure-detector
226
- direction and the accepted residual gap (a deploy that fails without degrading health).
227
- (jcardinal)
266
+ on `Degraded`/`Severe`; suppress a 5xx-driven `Degraded` unless share ≥ ratio), the
267
+ blind-read fail-safe (`alarm=HIGH` + `probe=DEGRADED`), and durable EB enhanced-health
268
+ facts (severity ladder; `ApplicationMetrics` raw counts; `Causes` classifies; failed
269
+ deploys aren't reliably in HealthStatus). (jcardinal)
270
+ </content>
271
+ </invoke>
@@ -18,10 +18,10 @@ _Auto-generated by `knowledge.js index`. Do not hand-edit._
18
18
 
19
19
  ## 2.0 framework
20
20
 
21
- - **_underscore** (_Underscore) _(framework core)_ — 62 doc(s) → [2.0/apps/_underscore/INDEX.md](2.0/apps/_underscore/INDEX.md)
21
+ - **_underscore** (_Underscore) _(framework core)_ — 66 doc(s) → [2.0/apps/_underscore/INDEX.md](2.0/apps/_underscore/INDEX.md)
22
22
  - **worker2** (Worker) — 56 doc(s) → [2.0/apps/worker2/INDEX.md](2.0/apps/worker2/INDEX.md)
23
23
  - **api2** (API) — 25 doc(s) → [2.0/apps/api2/INDEX.md](2.0/apps/api2/INDEX.md)
24
- - **dbchanges2** (Database Changes) _(framework core)_ — 8 doc(s) → [2.0/apps/dbchanges2/INDEX.md](2.0/apps/dbchanges2/INDEX.md)
24
+ - **dbchanges2** (Database Changes) _(framework core)_ — 9 doc(s) → [2.0/apps/dbchanges2/INDEX.md](2.0/apps/dbchanges2/INDEX.md)
25
25
  - **toga2-supply** (TOGa Supply) — 7 doc(s) → [2.0/apps/toga2-supply/INDEX.md](2.0/apps/toga2-supply/INDEX.md)
26
26
  - **saml** (SAML SSO Gateway) — 4 doc(s) → [2.0/apps/saml/INDEX.md](2.0/apps/saml/INDEX.md)
27
27
  - **toga2-view** (TOGa View Frontend) — 12 doc(s) → [2.0/apps/toga2-view/INDEX.md](2.0/apps/toga2-view/INDEX.md)
@@ -22,6 +22,7 @@ updated: 2026-08-25
22
22
  owners: [jcardinal, bala, tcox, apeterson, dfranks]
23
23
  files: []
24
24
  related:
25
+ - ../../2.0/apps/_underscore/features/sales-order-po-number-sourcing.md
25
26
  - features/persona-model-and-levy-gating.md
26
27
  - features/people-file-user-lifecycle.md
27
28
  - workflows/persona-refactor-migration.md
@@ -2,4 +2,5 @@
2
2
 
3
3
  | Doc | Framework | Summary | Files |
4
4
  |-----|-----------|---------|-------|
5
+ | [NYCHH PO links are UPSTREAM — the downstream SO→PO route returns empty](features/po-number-upstream-direction.md) | 2.0 | NYCHH's sales orders are created **from the customer's purchase order**, so their SO↔PO links live in the **upstream** table `PurchaseOrders_SalesOrders` (route | _underscore/Model/Client/SalesOrder.php, toga25-supply/src/pages/SalesOrders/view/SalesOrderRecordModalLayout/hooks/usePurchaseOrderDetails.ts, dbchanges2/Client/2026-08-11b - SalesOrderPurchaseOrdersField.sql |
5
6
  | [NYC Health & Hospitals](profile.md) | 2.0 | NYC Health & Hospitals (NYCHH) is a TOGA 2.0 client on the `_underscore` platform, prod schema `Client_Nychh`. | dbchanges2/Client_Nychh/2026-08-18 - InventoryUnitsItemColumns.sql, dbchanges2/Client_Nychh/2026-08-24 - FixInventoryGroupingsUnitsTopologyOverrideIds.sql |
@@ -0,0 +1,59 @@
1
+ ---
2
+ title: "NYCHH PO links are UPSTREAM — the downstream SO→PO route returns empty"
3
+ framework: "2.0"
4
+ repo: _underscore
5
+ project: _Underscore
6
+ client: nychh
7
+ type: client-feature
8
+ status: active
9
+ updated: 2026-08-25
10
+ owners: [apeterson]
11
+ files:
12
+ - _underscore/Model/Client/SalesOrder.php
13
+ - toga25-supply/src/pages/SalesOrders/view/SalesOrderRecordModalLayout/hooks/usePurchaseOrderDetails.ts
14
+ - dbchanges2/Client/2026-08-11b - SalesOrderPurchaseOrdersField.sql
15
+ related:
16
+ - ../../../2.0/apps/_underscore/features/sales-order-purchase-order-bridge-direction.md
17
+ - ../../../2.0/apps/_underscore/features/sales-order-po-number-sourcing.md
18
+ - ../profile.md
19
+ ---
20
+
21
+ ## Summary
22
+
23
+ NYCHH's sales orders are created **from the customer's purchase order**, so their SO↔PO links live
24
+ in the **upstream** table `PurchaseOrders_SalesOrders` (route `purchase-order-sales-orders`,
25
+ record 276). They have **no** downstream `SalesOrders_PurchaseOrders` rows.
26
+
27
+ This produced a symptom that looked like a bug and was not: the sales-order modal's PO Number came
28
+ back as an empty array while the sales-orders **list** column showed a PO number for the same
29
+ order. The front end was calling the **downstream** route; the list column's table-view join uses
30
+ the **upstream** table.
31
+
32
+ ## How it works
33
+
34
+ - Verified on `Client_Nychh` (sandbox-client) for the affected order: **upstream = 1 row,
35
+ downstream = 0 rows**.
36
+ - The API returned `200` with `totalRecordCount 0` and **no** `EZ-*`/`EV-*` messages — the
37
+ signature of a join matching nothing, not of a permission or field-resolution failure.
38
+ - **ACL was clean and was ruled out:** `AclRecordPermissions` on records **275 and 276** both grant
39
+ `roleId` 1 and 3 full CRUD in `Client_Nychh`; the JWT carried client roles `[1,2]`. This is *not*
40
+ the Compass Canada-style role mis-seed.
41
+ - Both relevant migrations were confirmed to have run on `Client_Nychh`
42
+ (`Client/2026-08-11b`, `Client/2026-08-13`) — see
43
+ [verifying a migration ran](../../../2.0/apps/dbchanges2/workflows/verifying-a-migration-ran.md).
44
+
45
+ ## Gotchas
46
+
47
+ - 🚨 **Open consequence.** The shared `_purchaseOrders` calculated field reads the **downstream**
48
+ table (`_underscore/Model/Client/SalesOrder.php:348`). Now that the modal's PO Number binds to
49
+ `_purchaseOrders`, NYCHH resolves it without an `EV-8` but gets **`null` forever**. NYCHH needs an
50
+ **upstream counterpart** — either an upstream-aware override of `_purchaseOrders` for this tenant
51
+ or an upstream variant field. **No tenant precedent exists for upstream-only:** Compass USA reads
52
+ both, Quad and Compass Canada are downstream-only.
53
+ - Do not "fix" this by adding downstream rows. The direction is correct for NYCHH's business model;
54
+ the shared field is what is incomplete.
55
+
56
+ ## Change history
57
+ - 2026-08-25 — Diagnosed the empty-array PO Number as a wrong-direction query (upstream data,
58
+ downstream route); ruled out ACL and migrations; recorded the still-open `_purchaseOrders`
59
+ downstream-only gap for this tenant (apeterson)
@@ -23,6 +23,8 @@ files:
23
23
  related:
24
24
  - ../../2.0/apps/_underscore/features/tracking-number-bridges.md
25
25
  - ../../2.0/apps/_underscore/features/tableview-joins.md
26
+ - ./features/po-number-upstream-direction.md
27
+ - ../../2.0/apps/_underscore/features/sales-order-purchase-order-bridge-direction.md
26
28
  ---
27
29
 
28
30
  ## Summary
@@ -52,6 +54,11 @@ table views. Client-specific DB change-sets live in `dbchanges2/Client_Nychh/`.
52
54
  tag added after the ItemShip never bumps the fulfillment sync. See
53
55
  [NYCHH Asset-Tag Backfill](../../2.0/apps/worker2/features/nychh-asset-tag-backfill.md).
54
56
 
57
+ - **PO links are UPSTREAM (`PurchaseOrders_SalesOrders`), not downstream.** The downstream
58
+ `sales-order-purchase-orders` route returns an empty array for NYCHH by design, and the shared
59
+ `_purchaseOrders` field (downstream-only) resolves to `null` for this tenant — open gap. See
60
+ [NYCHH PO Number direction](./features/po-number-upstream-direction.md).
61
+
55
62
  ## Notes
56
63
  - Tracking data: record 318 (item-level) is currently empty for this client; their tracking
57
64
  populates the IF/shipment level (record 317). The rebuilt views use 318 (per the Compass
@@ -2,5 +2,6 @@
2
2
 
3
3
  | Doc | Framework | Summary | Files |
4
4
  |-----|-----------|---------|-------|
5
+ | [Quad Enter PO Details — single-vendor-by-country selection and the EV-12 `vendorItem` failure](features/enter-po-details-vendor-selection-ev12.md) | 2.0 | Adding a PO number for Quad through **Enter PO Details** can `400` with **`EV-12` "unable to find a unique match"** on field **`vendorItem`**. | _underscore/Model/Quad/PurchaseOrder.php |
5
6
  | [Quad: PO + ASN email importer (import_po_and_asn.php)](features/po-asn-email-import.md) | 1.0 | Quad Graphics' supplier emails Purchase Order and Advance Shipping Notice CSVs into a monitored Microsoft 365 mailbox. | worker/crons/toga2/quad/import_po_and_asn.php |
6
7
  | [Quad Graphics](profile.md) | 2.0 | Quad Graphics is a TOGA 2.0 client on the `_underscore` platform, prod schema `Client_Quad`. | |
@@ -0,0 +1,62 @@
1
+ ---
2
+ title: "Quad Enter PO Details — single-vendor-by-country selection and the EV-12 `vendorItem` failure"
3
+ framework: "2.0"
4
+ repo: _underscore
5
+ project: _Underscore
6
+ client: quad
7
+ type: client-feature
8
+ status: active
9
+ updated: 2026-08-25
10
+ owners: [apeterson]
11
+ files:
12
+ - _underscore/Model/Quad/PurchaseOrder.php
13
+ related:
14
+ - ../../../2.0/apps/api2/features/v2-api-error-codes.md
15
+ - ../../../2.0/apps/_underscore/features/sales-order-po-number-sourcing.md
16
+ - ../../../2.0/standards/backend-php.md
17
+ - ../profile.md
18
+ ---
19
+
20
+ ## Summary
21
+
22
+ Adding a PO number for Quad through **Enter PO Details** can `400` with
23
+ **`EV-12` "unable to find a unique match"** on field **`vendorItem`**. The failure is a **code
24
+ defect plus a country-scoped catalog coverage gap** — not a bad data reset and not ACL. Quad's
25
+ `Client_Quad` data was verified healthy (592 `VendorItems`, 18 active vendors, 15 with a complete
26
+ `Vendors → Addresses → States → Countries` chain).
27
+
28
+ ## How it works
29
+
30
+ 1. The FE payload contains **no items**. Quad's `prePost` interceptor
31
+ (`_underscore/Model/Quad/PurchaseOrder.php:9` → `populatePurchaseOrderFromSalesOrder`) builds
32
+ them from the sales order's lines.
33
+ 2. `getVendorByCountry` (line ~186) picks **one vendor for the whole PO**: first the vendor covering
34
+ the **most** of the order's items *in the ship-to country*, else **any** vendor in that country.
35
+ 3. Line ~79 then restricts the `VendorItems` join to that **single** vendor.
36
+ 4. When the chosen vendor has no `VendorItems` row for a line, lines ~127-139 fall back to sending a
37
+ **descriptor** (`vendor` + `item` + `vendorPartNumber`).
38
+
39
+ ## Gotchas
40
+
41
+ - 🚨 **The descriptor fallback cannot work — its comment is wrong.** The comment claims the API can
42
+ "create/match" the descriptor. The field's child policy is **`MATCH`**, which only *looks up* and
43
+ **never creates**. Zero matches (or 2+) ⇒ `EV-12`. This is the primary defect.
44
+ - **Real-world instance:** order ships to Poland; picked vendor **Bechtle Direct Polska** (PL,
45
+ correct); item **MGDN4HN/A** is mapped **only** to General Technologies (IN, India). No Polish
46
+ supplier for the part ⇒ `EV-12`. Reproduced on a second order with a different vendor/item pair —
47
+ it is **systemic**, not one-off.
48
+ - 🚨 **Secondary defect — nondeterministic vendor choice.** `getVendorByCountry`'s ranking query is
49
+ `ORDER BY vendorItemCount DESC LIMIT 1` with **no tie-break column**. On tied counts the winner is
50
+ arbitrary and can **change after a database reload** (row order/ids move). This is exactly the
51
+ hazard in [backend-php standards](../../../2.0/standards/backend-php.md) ("a `LIMIT 1` on a
52
+ non-unique key MUST have a deterministic `ORDER BY`"). Fix is to append `, v.id`.
53
+ - **A vendor with a broken address link is invisible to vendor selection.** The queries require the
54
+ full `Vendors → Addresses → States → Countries` chain; a vendor missing any hop is silently
55
+ skipped and the PO is silently handed to a different vendor.
56
+ - **Recommended remediation (not yet applied):** raise a useful error naming the chosen vendor and
57
+ the unmapped item instead of an opaque `EV-12`, and add the deterministic tie-break.
58
+
59
+ ## Change history
60
+ - 2026-08-25 — Diagnosed EV-12 on `vendorItem` as the `MATCH`-policy descriptor fallback that can
61
+ never create, driven by single-vendor-by-country selection against a country-scoped catalog gap;
62
+ also recorded the missing `LIMIT 1` tie-break in `getVendorByCountry` (apeterson)
@@ -20,6 +20,8 @@ owners: ["jcardinal", "bala", "apeterson", "ajean"]
20
20
  files: []
21
21
  related:
22
22
  - ./features/po-asn-email-import.md
23
+ - ./features/enter-po-details-vendor-selection-ev12.md
24
+ - ../../2.0/apps/_underscore/features/sales-order-po-number-sourcing.md
23
25
  - ../../2.0/apps/_underscore/features/tracking-number-bridges.md
24
26
  - ../../2.0/apps/_underscore/features/item-fulfillment-stage-lifecycle-and-order-status.md
25
27
  - ../../2.0/apps/_underscore/features/surface-resolver.md
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "toga-ai",
3
- "version": "1.0.649",
3
+ "version": "1.0.651",
4
4
  "description": "TOGA Technology Team Claude Knowledge System — shared AI coding harness with skills, knowledge base CLI, and project installer for Claude Code.",
5
5
  "keywords": [
6
6
  "claude",