toga-ai 1.0.495 → 1.0.496
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/knowledge/1.0/apps/worker/INDEX.md +1 -0
- package/knowledge/1.0/apps/worker/workflows/diagnosing-frozen-cron-checkins.md +103 -0
- package/knowledge/2.0/apps/_underscore/features/email-send-pipeline.md +37 -1
- package/knowledge/2.0/apps/api2/INDEX.md +1 -1
- package/knowledge/2.0/apps/api2/features/request-logging.md +2 -1
- package/knowledge/INDEX.md +1 -1
- package/package.json +1 -1
|
@@ -10,5 +10,6 @@
|
|
|
10
10
|
| [NetSuite → TOGa Supply Per-Client Sync (thin wrappers)](features/netsuite-togasupply-per-client-sync.md) | Syncs NetSuite transactions (sales orders, purchase orders, invoices, item receipts, item fulfillments, inventory adjustments) into each TOGa Supply (2.0) clien | worker/crons/toga2/netsuite/common_sync_togasupply.php, worker/crons/toga2/netsuite/sync_togasupply_canon.php, worker/schedules/cron.worker.sync.json, dbchanges2/_modules/netsuite/2026-04-01 - Parameters.sql, library/app/api/netsuite/rest.php, library/app/systemmonitor/netsuiteintegration.php |
|
|
11
11
|
| [OneUptime external uptime monitoring for 1.0 workers](features/oneuptime-worker-uptime-monitoring.md) | Every 1.0 worker box self-reports its liveness to an external OneUptime monitor once per minute by curl-POSTing to a per-worker "Incoming Request" heartbeat URL | library/app/worker.php, worker/crons/worker/worker_heartbeat.php |
|
|
12
12
|
| [Prudential: Send Shipments for the Day report (daily cron)](features/send-shipments-for-the-day.md) | Daily cron (9:00 PM) that emails Prudential and Dell stakeholders an Excel report of all devices shipped that day, including tracking number, serial number, emp | worker/crons/notifications/reports/send_shipments_for_the_day.php |
|
|
13
|
+
| [Diagnosing frozen 1.0 worker cron check-ins (Sentry "missed" flood)](workflows/diagnosing-frozen-cron-checkins.md) | When 1.0 worker cron timestamps freeze and Sentry project `worker1` fills with **`missed`** check-ins, the intuitive diagnosis — a wedged `App_Framework::isProc | worker/.ebextensions/cron.config, library/app/worker.php |
|
|
13
14
|
| [isFulfillable Multi-Client Backfill (all togasupply clients)](workflows/isfulfillable-multi-client-backfill.md) | One-time backfill that catches up `Items.isFulfillable` on **existing** items across **all 17 togasupply clients** (AIG, Broward Sheriff, Canon, Endeavor Health | worker/crons/toga2/netsuite/backfill_isfulfillable_all_clients.php, library/app/api/toga2.php |
|
|
14
15
|
| [Onboarding a Client to the NetSuite TOGa Supply Sync](workflows/onboarding-client-to-netsuite-togasupply-sync.md) | How to add a new TOGa 2 client to the per-client NetSuite → TOGa Supply importer (`worker/crons/toga2/netsuite/`). | worker/crons/toga2/netsuite/sync_togasupply.php, worker/crons/toga2/netsuite/common_sync_togasupply.php, worker/schedules/cron.worker.sync.json, dbchanges2/_modules/netsuite/2026-04-01 - Parameters.sql |
|
|
@@ -0,0 +1,103 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Diagnosing frozen 1.0 worker cron check-ins (Sentry "missed" flood)
|
|
3
|
+
framework: "1.0"
|
|
4
|
+
repo: worker
|
|
5
|
+
project: Worker
|
|
6
|
+
client: shared
|
|
7
|
+
type: workflow
|
|
8
|
+
status: active
|
|
9
|
+
updated: 2026-08-03
|
|
10
|
+
owners: ["dfranks"]
|
|
11
|
+
files:
|
|
12
|
+
- worker/.ebextensions/cron.config
|
|
13
|
+
- library/app/worker.php
|
|
14
|
+
related:
|
|
15
|
+
- ../architecture.md
|
|
16
|
+
- ../features/oneuptime-worker-uptime-monitoring.md
|
|
17
|
+
- ../../library/features/cron-execution-monitoring.md
|
|
18
|
+
---
|
|
19
|
+
|
|
20
|
+
## Summary
|
|
21
|
+
|
|
22
|
+
When 1.0 worker cron timestamps freeze and Sentry project `worker1` fills with **`missed`**
|
|
23
|
+
check-ins, the intuitive diagnosis — a wedged `App_Framework::isProcessRunning()` lock — is
|
|
24
|
+
usually **wrong**. The common cause is **Elastic Beanstalk deployment churn replacing the EC2
|
|
25
|
+
instances underneath the crons**. This workflow is the ordered path that distinguishes the two
|
|
26
|
+
and the interpretation traps along the way.
|
|
27
|
+
|
|
28
|
+
## Where to look (name mapping)
|
|
29
|
+
|
|
30
|
+
Sentry project **`worker1`** is the 1.0 worker tier running in the Elastic Beanstalk
|
|
31
|
+
environment **`agilant-worker`** (application "Agilant", `us-west-2`, account
|
|
32
|
+
`502614707982`, reachable with the `default` SSO profile). It is a multi-instance autoscaled
|
|
33
|
+
fleet. **There is no EB environment literally named `worker1`** — the Sentry project name and
|
|
34
|
+
the EB environment name do not match, and looking for the wrong one costs real time.
|
|
35
|
+
|
|
36
|
+
## The failure mode
|
|
37
|
+
|
|
38
|
+
An EB deploy that aborts — events read like `Failed to deploy application`,
|
|
39
|
+
`Unsuccessful command execution on instance id(s) ...`,
|
|
40
|
+
`Cannot complete command execution ... no longer running` — causes EB to **replace
|
|
41
|
+
instances**. In-flight cron processes die with their host, and each Sentry monitor's
|
|
42
|
+
`lastCheckIn` freezes at whatever its now-terminated instance last reported. Nothing is
|
|
43
|
+
stuck; nothing needs clearing.
|
|
44
|
+
|
|
45
|
+
Two consequences worth knowing:
|
|
46
|
+
|
|
47
|
+
- **Check-ins resume on their own**, typically within ~1 minute of the deploy completing.
|
|
48
|
+
No manual intervention, no lock to clear, no instance to restart.
|
|
49
|
+
- EB may warn that an aborted deploy left instances on **mixed application versions**. That
|
|
50
|
+
does need a re-deploy of the intended version — the recovery of the crons does not imply
|
|
51
|
+
the fleet is on the version you meant to ship.
|
|
52
|
+
|
|
53
|
+
## Diagnostic path
|
|
54
|
+
|
|
55
|
+
1. **List the project's monitors and find the stall boundary.**
|
|
56
|
+
`GET https://sentry.io/api/0/organizations/<org>/monitors/?project=worker1` — read each
|
|
57
|
+
monitor's `environments[].lastCheckIn`. The cluster of timestamps that all stop at roughly
|
|
58
|
+
the same moment is the boundary.
|
|
59
|
+
2. **Pull per-monitor check-in history.**
|
|
60
|
+
`GET .../monitors/<slug>/checkins/` — a `missed` entry **with no `duration`** means the run
|
|
61
|
+
never started or never reported. A run that started and then hung looks different (it has a
|
|
62
|
+
start and no clean check-out). This is the first signal separating "process never launched"
|
|
63
|
+
from "process wedged".
|
|
64
|
+
3. **KEY DISCRIMINATOR — are non-monitored crons on the same host still throwing ordinary
|
|
65
|
+
exceptions in the same Sentry project's issue stream?** If other crons are still erroring
|
|
66
|
+
*after* the stall boundary, then PHP runs and Sentry connectivity both work and the host is
|
|
67
|
+
alive. That rules out "host down / Sentry unreachable" and points at instance replacement
|
|
68
|
+
or a per-cron blockage.
|
|
69
|
+
4. **Confirm against Elastic Beanstalk.**
|
|
70
|
+
```
|
|
71
|
+
aws elasticbeanstalk describe-events --environment-name agilant-worker --start-time <ISO>
|
|
72
|
+
aws elasticbeanstalk describe-environment-health --environment-name agilant-worker --attribute-names All
|
|
73
|
+
```
|
|
74
|
+
Correlate deploy / abort / instance-add / instance-remove timestamps against the stall
|
|
75
|
+
boundary. A deploy window that brackets the boundary is the answer.
|
|
76
|
+
|
|
77
|
+
## Interpretation caveats
|
|
78
|
+
|
|
79
|
+
- **Timezones.** Sentry returns **UTC**; the developer reporting the symptom is usually
|
|
80
|
+
quoting local **CDT (UTC-5)**. Convert before concluding the windows "don't line up" — they
|
|
81
|
+
usually do.
|
|
82
|
+
- **Establish each monitor's baseline before calling a stale timestamp new.** The `worker1`
|
|
83
|
+
monitor list chronically contains monitors that have been silently dead for weeks or months
|
|
84
|
+
for reasons unrelated to any current incident. During an incident these are noise and must
|
|
85
|
+
not be read as "failed to recover." (Deliberately not listing which ones — that goes stale.
|
|
86
|
+
Check each monitor's own history.)
|
|
87
|
+
- **Separate high-frequency from low-frequency monitors.** A 5-minute monitor should recover
|
|
88
|
+
within minutes; an hourly or daily one will not move until its next scheduled slot, and its
|
|
89
|
+
silence proves nothing.
|
|
90
|
+
- **Alerting-hygiene gap.** Chronically dead cron monitors mean nobody is being alerted on
|
|
91
|
+
those jobs at all. Worth a separate follow-up whenever you notice one.
|
|
92
|
+
|
|
93
|
+
## After recovery
|
|
94
|
+
|
|
95
|
+
A multi-minute cron outage on this tier means missed windows for order processing, email
|
|
96
|
+
sending, and inbound order imports. These are catch-up-on-next-run designs and generally
|
|
97
|
+
self-drain — but **eyeball queue depth / backlog after recovery rather than assuming clean**.
|
|
98
|
+
|
|
99
|
+
## Change history
|
|
100
|
+
|
|
101
|
+
- 2026-08-03 — Documented from a production diagnostic session: frozen cron timestamps +
|
|
102
|
+
`missed` check-in flood on `worker1` traced to EB deployment churn on `agilant-worker`, not
|
|
103
|
+
a wedged `isProcessRunning()` lock. (dfranks)
|
|
@@ -7,7 +7,7 @@ client: shared
|
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
9
|
updated: 2026-07-29
|
|
10
|
-
owners: ["bala", tcox]
|
|
10
|
+
owners: ["bala", "tcox", "dfranks"]
|
|
11
11
|
files:
|
|
12
12
|
- _underscore/Email.php
|
|
13
13
|
- worker2/Worker/Infrastructure/Email/Send.php
|
|
@@ -71,11 +71,47 @@ SMTP** using PHPMailer, then flips each row to `SENT` or `FAILED`. So a 2.0 emai
|
|
|
71
71
|
`DevTeam@goagilant.com` is a known-good verified sender identity in worker2 config. See
|
|
72
72
|
[ai-bdr web-funnel-app.md](../../ai-bdr/features/web-funnel-app.md).
|
|
73
73
|
|
|
74
|
+
- **⚠ READ-AFTER-WRITE ON THE READER REPLICA LOSES `send()`'S OWN LOG ROW (deterministic).**
|
|
75
|
+
`_Email::send()` INSERTs the `PENDING` `Logs_[Client].Email` row, then reads `$emailLog->id`
|
|
76
|
+
back. On a `_Model`, `save()` sets the PK in memory from `getInsertId()` but does **not** clear
|
|
77
|
+
`_model_needsToBeInitialized`, so the next `->id` access triggers `__get` → `initialize()`,
|
|
78
|
+
which issues a fresh `SELECT ... WHERE id=<new id>`. `DB_CLIENT_LOGS` is registered with a
|
|
79
|
+
**separate `readHost`**, so that SELECT is routed to the **reader replica** — but the INSERT is
|
|
80
|
+
still inside an **uncommitted** lazy transaction, so it has produced no binlog event and the
|
|
81
|
+
replica **can never** see it during this request. The read returns 0 rows and `initialize()`
|
|
82
|
+
throws `"exactly 1 row expected, 0"`. This is **100% deterministic, not a lag race** — an
|
|
83
|
+
uncommitted write cannot replicate. A connection *can* see its own uncommitted writes
|
|
84
|
+
(read-your-own-writes), so forcing the read onto the writer is the fix.
|
|
85
|
+
- **Why only some emails hit it:** approval/order-placed emails send from inside the **api2
|
|
86
|
+
CRUD POST path** (and worker2), which call `_Database::setIsReadHostEnabled(false)`
|
|
87
|
+
(`V2.php` create loop; worker2 `Controller/Index.php`) — their identical read-back hits the
|
|
88
|
+
**writer** and succeeds. The scripted-API path (`/email-templates/sendEmail`) never suppresses
|
|
89
|
+
the read host, so it throws every time. This was the root cause of the Compass In-Transit
|
|
90
|
+
email outage — see
|
|
91
|
+
[`compass-partial-in-transit-delivered-emails.md`](../../../1.0/apps/worker/features/compass-partial-in-transit-delivered-emails.md).
|
|
92
|
+
- **Fix direction (smallest blast radius):** wrap the log-row `save()` + id read (+ attachment
|
|
93
|
+
loop) in `_Email::send()` with `_Database::setIsReadHostEnabled(false)`, restoring the prior
|
|
94
|
+
value in a `finally` (get/set at `Database.php`). Fixes all scripted-API emails and hardens the
|
|
95
|
+
interceptor path.
|
|
96
|
+
- **Rejected alternatives:** (a) `commit` right after `save()` before the read — recreates the
|
|
97
|
+
failure as an *intermittent* replication-lag race AND breaks the request's one-transaction
|
|
98
|
+
boundary (orphans the row on a later throw, prematurely commits everything else buffered on the
|
|
99
|
+
connection, strands attachment/API-log writes on a closed lazy transaction). (b) Clear
|
|
100
|
+
`_model_needsToBeInitialized` after INSERT in `_Model::save()` so `->id` reads from memory
|
|
101
|
+
(zero queries) — conceptually the *correct* fix but changes every model in the ORM, too broad
|
|
102
|
+
for a hotfix.
|
|
103
|
+
|
|
74
104
|
## Change history
|
|
75
105
|
- 2026-07-29 — Noted that the SES SMTP credentials/settings now have a **non-PHP consumer**: the
|
|
76
106
|
BDR funnel sends via nodemailer against the same SES host (us-west-2, port 587, TLS) using
|
|
77
107
|
`TOGA_SMTP_*` env vars, so a rotation must be coordinated with BDR's `.env.local` + Amplify
|
|
78
108
|
vars or BDR's summary email goes dark. (tcox)
|
|
109
|
+
- 2026-07-29 — Documented the **deterministic read-after-write failure** in `_Email::send()`:
|
|
110
|
+
reading `$emailLog->id` after INSERT triggers `_Model::__get`→`initialize()`→ a reader-replica
|
|
111
|
+
SELECT for an uncommitted row → `"exactly 1 row expected, 0"` throw. Only the scripted-API send
|
|
112
|
+
path fails (CRUD/worker paths suppress the read host via `setIsReadHostEnabled(false)`); this was
|
|
113
|
+
the root cause of the Compass In-Transit outage. Fix direction = suppress the read host around the
|
|
114
|
+
log-row save+read in `send()`. (dfranks)
|
|
79
115
|
- 2026-07-28 — Created: documented that `_Email::send()` **queues** (writes a `PENDING`
|
|
80
116
|
`Logs_[Client].Email` row + attachment BLOBs) rather than transmitting, and that the worker2
|
|
81
117
|
`Infrastructure/Email/Send` cron (`Core.CronJobs` `* * * * *`) does the actual SES SMTP send
|
|
@@ -11,7 +11,7 @@
|
|
|
11
11
|
| [Nested FK object embedding is gated by the CHILD record's own ACL](features/nested-fk-acl-embedding.md) | When the V2 JSON engine serializes a foreign-key field into a **nested object** (in `getFullModelData()`, ~V2.php L6016-6060), it re-checks the **child** record | api2/Component/Api/V2/V2.php, dbchanges2/Client_Compass/2026-07-23b - PurchaseOrdersRecordReadAcl.sql |
|
|
12
12
|
| [Nested-relationship writes & child matching (link vs. create)](features/nested-relationship-writes.md) | When a 2.0 API write payload (`POST`/`PUT`) contains a **nested related object** (e.g. | api2/Component/Api/V2/V2.php, _underscore/Model/Client/ContactEmailAddress.php |
|
|
13
13
|
| [Record Scripts (computed/aggregate /v2 endpoints — the authoring contract)](features/record-scripts.md) | In api2 you almost never write a controller. | api2/Component/Api/V2/V2.php, _underscore/Model/Team/Sprint.php, _underscore/Model/Client/ItemFulfillment.php, _underscore/Query.php |
|
|
14
|
-
| [V2 Request Logging & Where Requests Land (client vs core log DB)](features/request-logging.md) | The V2 engine logs **every inbound request** — success *and* failure, with response code and payload — and routes each log entry to the **client** log or the ** | api2/Component/Api/V2/V2.php, _underscore/Model/Client/Logs/Api.php, _underscore/Model/Core/Logs/Api.php |
|
|
14
|
+
| [V2 Request Logging & Where Requests Land (client vs core log DB)](features/request-logging.md) | The V2 engine logs **every inbound request** — success *and* failure, with response code and payload — and routes each log entry to the **client** log or the ** | api2/Component/Api/V2/V2.php, api2/Controller/Index.php, _underscore/Model/Client/Logs/Api.php, _underscore/Model/Core/Logs/Api.php |
|
|
15
15
|
| [POST + JSON-body args for scripted APIs](features/scripted-api-post-body-args.md) | The V2 engine can run a Record Script (scripted API) for a **POST** request, and a scripted API can receive its arguments from the **JSON request body** instead | api2/Component/Api/V2/V2.php |
|
|
16
16
|
| [TOGa IQ Sprint Dashboard API (Record Scripts)](features/sprint-dashboard-api.md) | The internal **TOGa IQ sprint dashboard** is served in production by **six api2 Record Scripts** on `_Model_Team_Sprint` (`_underscore/Model/Team/Sprint.php`), | _underscore/Model/Team/Sprint.php, api2/Component/Api/V2/V2.php, dbchanges2/Core/2026-07-24a - SprintDashboardRecordScripts.sql, dbchanges2/Client_True/2026-07-24b - SprintDashboardScriptAcl.sql |
|
|
17
17
|
| [Surface action-state via the surface=<slug> request option (M2M-safe)](features/surface-meta-option.md) | An opt-in V2 engine request option, `surface=<slug>`, that attaches per-record UI action state (`isVisible`/`isEnabled`) to a GET response **under `meta.surface | api2/Component/Api/V2/V2.php, _underscore/Model/Core/Surface.php |
|
|
@@ -7,9 +7,10 @@ client: shared
|
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
9
|
updated: 2026-07-29
|
|
10
|
-
owners: ["mhammontree"]
|
|
10
|
+
owners: ["mhammontree", "dfranks"]
|
|
11
11
|
files:
|
|
12
12
|
- api2/Component/Api/V2/V2.php
|
|
13
|
+
- api2/Controller/Index.php
|
|
13
14
|
- _underscore/Model/Client/Logs/Api.php
|
|
14
15
|
- _underscore/Model/Core/Logs/Api.php
|
|
15
16
|
related:
|
package/knowledge/INDEX.md
CHANGED
|
@@ -5,7 +5,7 @@ _Auto-generated by `knowledge.js index`. Do not hand-edit._
|
|
|
5
5
|
## 1.0 framework
|
|
6
6
|
|
|
7
7
|
- **library** (Library) _(framework core)_ — 14 doc(s) → [1.0/apps/library/INDEX.md](1.0/apps/library/INDEX.md)
|
|
8
|
-
- **worker** (Worker) —
|
|
8
|
+
- **worker** (Worker) — 18 doc(s) → [1.0/apps/worker/INDEX.md](1.0/apps/worker/INDEX.md)
|
|
9
9
|
- **dbchanges** (Database Changes) _(framework core)_ — 1 doc(s) → [1.0/apps/dbchanges/INDEX.md](1.0/apps/dbchanges/INDEX.md)
|
|
10
10
|
- **worker1.5** (Worker 1.5) — 0 doc(s) → [1.0/apps/worker1.5/INDEX.md](1.0/apps/worker1.5/INDEX.md)
|
|
11
11
|
- **togadesk** (TOGa Desk) — 10 doc(s) → [1.0/apps/togadesk/INDEX.md](1.0/apps/togadesk/INDEX.md)
|
package/package.json
CHANGED