toga-ai 1.0.504 → 1.0.506
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/knowledge/1.0/apps/library/INDEX.md +1 -0
- package/knowledge/1.0/apps/library/features/error-capture-1-0.md +165 -0
- package/knowledge/1.0/apps/tools/INDEX.md +1 -1
- package/knowledge/1.0/apps/tools/features/errors-curation-console.md +54 -1
- package/knowledge/1.0/apps/tools/features/talos-kb-documents-admin.md +13 -1
- package/knowledge/1.0/apps/worker/INDEX.md +1 -0
- package/knowledge/1.0/apps/worker/features/oneuptime-server-monitor-host-hygiene.md +204 -0
- package/knowledge/1.0/apps/worker/features/oneuptime-worker-uptime-monitoring.md +1 -1
- package/knowledge/1.0/standards/backend-php.md +45 -1
- package/knowledge/2.0/apps/_underscore/INDEX.md +1 -1
- package/knowledge/2.0/apps/_underscore/features/error-reporting-issue-event.md +116 -10
- package/knowledge/2.0/apps/api2/features/v2-api-error-codes.md +24 -0
- package/knowledge/2.0/apps/worker2/INDEX.md +2 -2
- package/knowledge/2.0/apps/worker2/features/alb-target-group-auto-registration.md +25 -1
- package/knowledge/2.0/apps/worker2/features/error-escalation-cron.md +86 -4
- package/knowledge/INDEX.md +2 -2
- package/knowledge/sessions/2026-08-04-new-error-handler-jcardinal.md +129 -0
- package/package.json +1 -1
|
@@ -8,6 +8,7 @@
|
|
|
8
8
|
| [Diagnostic Dialog — View Recommended Services Routing](features/diagnostic-dialog-view-recommended-services.md) | Two "View Recommended Services" buttons exist in the TOGa Refresh 2026 SR view: 1. | library/app/model/toga/diagnostic.php, library/app/model/servicerequest.php |
|
|
9
9
|
| [Elite Freshservice Sync (library)](features/elite-freshservice-sync.md) | `App_Api_Toga2` in `library/app/api/toga2.php` orchestrates bidirectional sync between TOGA 2 and TOGaDesk. | library/app/api/toga2.php |
|
|
10
10
|
| [Branded HTML Email Templates (App_Email_Template)](features/email-templates.md) | `App_Email_Template` (`app/email/template.php`) is the base class for branded HTML emails in the 1.0 (`App_`) framework. | library/app/email/template.php, library/app/email/agilant.php |
|
|
11
|
+
| [Error Capture in 1.0 (App_Error_Capture → shared 2.0 Logs DB)](features/error-capture-1-0.md) | The 1.0 side of the platform error-reporting pipeline (TRUE-78188). | library/app/error/capture.php, library/app/error.php, library/app/exception/business.php, library/app/cloud.php, worker/config.worker.ini, worker/crons/toga2/compass/workflow/1_transmit_compass_sales_orders_to_mits.php |
|
|
11
12
|
| [1.0 MVC Page Pattern & New-App Skeleton](features/mvc-page-pattern-and-app-skeleton.md) | This is the **reusable recipe for standing up a new 1.0 (`App_`) application** and for adding pages to one — the folder-based MVC routing, the page lifecycle, t | library/app/framework.php, library/app/frameworkindex.php, library/app/mvc.php, library/app/database.php, library/app/model.php, library/app/config.php |
|
|
12
13
|
| [isFulfillable from NetSuite during Item Sync (Phase 1)](features/netsuite-item-isfulfillable-sync.md) | This is the **1.0 (Phase 1)** half of the `isFulfillable` feature: reading the NetSuite `isfulfillable` flag during item sync and stamping it onto the **Agilant | library/app/netsuite.php, library/app/api/toga2.php, worker/crons/toga2/netsuite/common_sync_togasupply.php, worker/crons/toga2/netsuite/backfill_isfulfillable_jul5.php |
|
|
13
14
|
| [NetSuite SuiteQL/REST API Reference](features/netsuite-suiteql-api-reference.md) | General working reference for the Agilant NetSuite integration: how to authenticate, how SuiteQL behaves, and the confirmed schema of the tables/columns/codes w | library/app/api/netsuite/rest.php, library/ssl/netsuite_ec_key.pem, test/@dave/Junk Drawer/nsq.php |
|
|
@@ -0,0 +1,165 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: "Error Capture in 1.0 (App_Error_Capture → shared 2.0 Logs DB)"
|
|
3
|
+
framework: "1.0"
|
|
4
|
+
repo: library
|
|
5
|
+
project: Library
|
|
6
|
+
client: shared
|
|
7
|
+
type: feature
|
|
8
|
+
status: active
|
|
9
|
+
updated: 2026-08-04
|
|
10
|
+
owners: ["jcardinal"]
|
|
11
|
+
files:
|
|
12
|
+
- library/app/error/capture.php
|
|
13
|
+
- library/app/error.php
|
|
14
|
+
- library/app/exception/business.php
|
|
15
|
+
- library/app/cloud.php
|
|
16
|
+
- worker/config.worker.ini
|
|
17
|
+
- worker/crons/toga2/compass/workflow/1_transmit_compass_sales_orders_to_mits.php
|
|
18
|
+
related:
|
|
19
|
+
- ../architecture.md
|
|
20
|
+
- ../../../2.0/apps/_underscore/features/error-reporting-issue-event.md
|
|
21
|
+
- ../../../2.0/apps/worker2/features/error-escalation-cron.md
|
|
22
|
+
- ../../tools/features/errors-curation-console.md
|
|
23
|
+
- ../../standards/backend-php.md
|
|
24
|
+
---
|
|
25
|
+
|
|
26
|
+
## Summary
|
|
27
|
+
|
|
28
|
+
The 1.0 side of the platform error-reporting pipeline (TRUE-78188). `App_Error_Capture`
|
|
29
|
+
(`library/app/error/capture.php`) records 1.0 errors as occurrences in the **same shared 2.0 Logs
|
|
30
|
+
database** (`Logs.Issue` / `IssueFingerprint` / `Event`) that 2.0 writes to, so 1.0 and 2.0 feed
|
|
31
|
+
**one issue list, one Tools console, one escalation cron**. Before this, 1.0 errors existed only
|
|
32
|
+
in Sentry and were invisible to the issue system.
|
|
33
|
+
|
|
34
|
+
`App_Error_Capture::captureException()` is the single 1.0 capture entry point and returns the
|
|
35
|
+
quotable `Issue.reference`, mirroring `_Error::captureException()` in 2.0. Add new 1.0 capture
|
|
36
|
+
paths by calling it — never by re-implementing persistence.
|
|
37
|
+
|
|
38
|
+
Read the 2.0 doc
|
|
39
|
+
([error-reporting-issue-event](../../../2.0/apps/_underscore/features/error-reporting-issue-event.md))
|
|
40
|
+
for the data model, fingerprint scheme, reference encoding, and the MySQL constraints. This doc
|
|
41
|
+
covers only what is **different or additional in 1.0**.
|
|
42
|
+
|
|
43
|
+
## Opt-in by config section
|
|
44
|
+
|
|
45
|
+
Capture requires a **`[database_toga2logs]`** section in the app's config (added to
|
|
46
|
+
`worker/config.worker.ini`). **An app with no such section silently does not capture.** That is
|
|
47
|
+
deliberate: `library` is shared by every 1.0 app, and a shared library must not fatal or spam an
|
|
48
|
+
app that has not been wired to the Logs DB yet. Wiring a new 1.0 app into error reporting is
|
|
49
|
+
therefore purely a config change.
|
|
50
|
+
|
|
51
|
+
> `dbname` is **not** necessarily `Logs` — 2.0 resolves the Core Logs database name at runtime
|
|
52
|
+
> from `Core.Databases` id 11. Read the real name per environment (same caveat as the Tools
|
|
53
|
+
> console).
|
|
54
|
+
|
|
55
|
+
## How it works
|
|
56
|
+
|
|
57
|
+
### Hand-written, fully escaped SQL
|
|
58
|
+
|
|
59
|
+
1.0 has **no `_Model`**, so every statement in `capture.php` is hand-written against
|
|
60
|
+
`App_Database`, with every interpolated value escaped (`App_Database::sqlEscape()`) or cast. There
|
|
61
|
+
is no model layer to absorb a schema change here: a rename in the shared Logs DB must be applied
|
|
62
|
+
to this file in the same release as the migration (the same exposure the Tools console has).
|
|
63
|
+
|
|
64
|
+
### `App_Exception_Business` — one issue across both frameworks
|
|
65
|
+
|
|
66
|
+
`library/app/exception/business.php` mirrors 2.0's `_Exception_Business`: a stable `issueKey` plus
|
|
67
|
+
a `minimumUrgency` floor. **The fingerprint text is byte-identical to 2.0's** (`'issueKey: ' .
|
|
68
|
+
KEY`), so the *same key* thrown from either framework resolves to **one** Issue with one
|
|
69
|
+
occurrence count and one set of recipients. Do not "tidy" that string on either side — the
|
|
70
|
+
cross-framework identity is the whole point.
|
|
71
|
+
|
|
72
|
+
Business issues are titled by their key: `Logs.Issue.subject` holds the `issueKey` (e.g.
|
|
73
|
+
`COMPASS_SALES_ORDER_MITS_TRANSMIT_REJECTED`) while `Issue.errorMessage` and each
|
|
74
|
+
`Event.errorMessage` keep the per-occurrence text. Otherwise the issue ends up named after
|
|
75
|
+
whichever occurrence happened to be seen first — i.e. one arbitrary order number.
|
|
76
|
+
|
|
77
|
+
Compass USA's sales-order → MITS transmit rejection was the first business-exception use case.
|
|
78
|
+
|
|
79
|
+
### Client attribution
|
|
80
|
+
|
|
81
|
+
`App_Error::setCurrentClientId()` exists in 1.0 for parity with 2.0's ambient current-client, but
|
|
82
|
+
**1.0 has no central hook** equivalent to `_Database::registerClientDatabases()`. Client-scoped
|
|
83
|
+
1.0 crons therefore **set it explicitly** (see
|
|
84
|
+
`worker/crons/toga2/compass/workflow/1_transmit_compass_sales_orders_to_mits.php`). Capture
|
|
85
|
+
prefers an explicit `clientId` argument and falls back to the ambient value.
|
|
86
|
+
|
|
87
|
+
### Fingerprint fallback (no usable stack)
|
|
88
|
+
|
|
89
|
+
Frames are keyed on **file + function**, not `file:line`, so inserting lines above a throw site
|
|
90
|
+
does not split an issue. When there is no usable application stack (fatals, OOM), identity is
|
|
91
|
+
**exception class + entry point (script basename) + message with digits masked** — never the
|
|
92
|
+
exception's own file/line, which for an OOM is wherever the allocation happened to cross the
|
|
93
|
+
limit. Same scheme as 2.0; keep the two implementations in lockstep.
|
|
94
|
+
|
|
95
|
+
### The 1.0 "red box" error page
|
|
96
|
+
|
|
97
|
+
The rendered error page now headlines the **issue reference** (e.g. `1A-1`). Sentry's event id is
|
|
98
|
+
a fallback shown only when our own capture produced nothing. The `sleep(13)` that existed purely
|
|
99
|
+
to let Sentry ingest before the link was clicked is gone. Business exceptions get a different
|
|
100
|
+
headline and **do** render their message — safe for that class only, because a business message is
|
|
101
|
+
written for a reader and carries no file, line, or trace.
|
|
102
|
+
|
|
103
|
+
### `Logs.Event.context` shape
|
|
104
|
+
|
|
105
|
+
Identical in both frameworks: **every top-level key holds an object**, and scalars are grouped
|
|
106
|
+
under a `Reference` object. Top-level keys are the display names — `Reference`, `GET`, `POST`,
|
|
107
|
+
`Runtime`, `AWS`, `SERVER`. `GET`/`POST` are cast to objects so an empty one serializes as `{}`,
|
|
108
|
+
not `[]`. `$_POST` is captured and redacted on the same key-name basis as `$_GET` (note PHP only
|
|
109
|
+
populates `$_POST` for form encodings, so JSON API bodies deliberately do not appear). The
|
|
110
|
+
over-length shed list was renamed in lockstep and its marker is an object too.
|
|
111
|
+
|
|
112
|
+
### Elastic Beanstalk environment name on AL1
|
|
113
|
+
|
|
114
|
+
1.0 runs **"64bit Amazon Linux/2.8.8" (AL1)**, which predates
|
|
115
|
+
`/opt/elasticbeanstalk/bin/get-config`, and the `042-write-eb-environment.sh` deploy hook only
|
|
116
|
+
exists in `api2`/`worker2`. So 1.0 reads the AL1 container config at
|
|
117
|
+
`/opt/elasticbeanstalk/deploy/configuration/containerconfiguration`, searching **recursively** for
|
|
118
|
+
`environment_name` (its position moved between platform versions), then falls back to
|
|
119
|
+
`aws ec2 describe-tags`.
|
|
120
|
+
|
|
121
|
+
## Gotchas / known issues
|
|
122
|
+
|
|
123
|
+
- **Any PHP warning inside capture code would kill the request.** 1.0's
|
|
124
|
+
`App_Error::handleError()` routes **every** PHP warning into `handleException()`, which renders
|
|
125
|
+
the red box and calls `exit()` — so a single warning raised while capturing (an unreachable IMDS
|
|
126
|
+
endpoint is the everyday case: `file_get_contents` warns before returning `false`) terminates
|
|
127
|
+
the request or cron mid-run while reporting an unrelated error. 2.0's equivalent handler
|
|
128
|
+
*throws*, which a `catch` absorbs. `captureException()` therefore sets
|
|
129
|
+
`App_Error::setThrowExceptionsEnabled(false)` for its own duration and restores it in
|
|
130
|
+
`finally`. This disables the **handler**, not real throws — `App_Query` raises a plain
|
|
131
|
+
`throw new Exception` on DB errors, so genuine failures still reach the catch. **This is a
|
|
132
|
+
general 1.0 trap for any code that must not escalate**, not just error capture.
|
|
133
|
+
- **No typed properties anywhere in `library/`.** 1.0 runs **PHP 7.2**; a deploy failed with
|
|
134
|
+
`syntax error, unexpected '?'` on `private static ?string $memoryReserve`. Nine typed properties
|
|
135
|
+
across `error.php`, `error/capture.php`, and `exception/business.php` were converted to untyped
|
|
136
|
+
with `@var` docblocks. Parameter and return types are fine (7.0/7.1).
|
|
137
|
+
- **An early `return` can make a fallback unreachable.** Two bugs in the first EB-environment
|
|
138
|
+
attempt: a `return` when `get-config` was not executable skipped the `describe-tags` fallback
|
|
139
|
+
entirely — i.e. the fallback was dead on the only platform that needed it — and the CLI call
|
|
140
|
+
omitted `App_Cloud::$preCommands`, so it did not authenticate the way every other 1.0 AWS call
|
|
141
|
+
does. **Unverified:** whether that IAM identity actually holds `ec2:DescribeTags`.
|
|
142
|
+
- **Fingerprint scheme changes re-group existing issues once.** Moving to file+function frames and
|
|
143
|
+
the class+entryPoint fallback changes hashes, so open issues re-fingerprint on first occurrence
|
|
144
|
+
after deploy. Deliberate; merge old hashes onto the curated Issue with the Tools
|
|
145
|
+
merge-fingerprint box.
|
|
146
|
+
- **Pre-existing committed secrets (location only — not fixed, out of scope).**
|
|
147
|
+
`library/app/cloud.php` holds a hardcoded **AWS access key id and secret access key** as
|
|
148
|
+
class constants/statics (~lines 23–24), used by `App_Cloud::$preCommands` for every 1.0 AWS CLI
|
|
149
|
+
call. Unrelated to this work. Treat as compromised: move to SSM like the other parameters and
|
|
150
|
+
rotate.
|
|
151
|
+
|
|
152
|
+
## Change history
|
|
153
|
+
|
|
154
|
+
- 2026-08-04 — Built the 1.0 side of TRUE-78188: `App_Error_Capture` writing occurrences into the
|
|
155
|
+
shared 2.0 `Logs` DB (opt-in via `[database_toga2logs]`, hand-written escaped SQL);
|
|
156
|
+
`App_Exception_Business` with a fingerprint string byte-identical to 2.0's so one `issueKey`
|
|
157
|
+
spans both frameworks; business issues titled by `issueKey`; the red-box page headlining the
|
|
158
|
+
issue reference and dropping the Sentry `sleep(13)`; the uniform object-per-top-level-key
|
|
159
|
+
`Event.context` shape with `$_POST` capture; AL1 container-config + `describe-tags` resolution of
|
|
160
|
+
the EB environment name; explicit `setCurrentClientId()` for client-scoped crons. Recorded two
|
|
161
|
+
framework-level traps: PHP 7.2 forbids typed properties in `library/`, and
|
|
162
|
+
`App_Error::handleError` escalates any warning to a rendered page + `exit()`, which capture
|
|
163
|
+
suppresses with `setThrowExceptionsEnabled(false)`/`finally`. Also fixed a latent `$this->errStr`
|
|
164
|
+
in a static branch of `error.php` and noted the pre-existing hardcoded AWS keys in
|
|
165
|
+
`app/cloud.php`. (jcardinal)
|
|
@@ -7,7 +7,7 @@
|
|
|
7
7
|
| [CloudFront Client Setup](features/cloudfront-client-setup.md) | An SSO-gated admin tool at **`/devops/cloudfront-clients`** in the Tools 1.0 app that onboards a client onto **CloudFront + Route 53 across multiple AWS account | tools/_/app/devops/cloudfront.php, tools/mvc/devops/cloudfront-clients/get.php, tools/mvc/devops/cloudfront-clients/post.php, tools/assets/js/cloudfront-clients.js, tools/assets/css/cloudfront-clients.css, tools/_/app/nav.php, tools/_/app/frameworkindex.php, tools/config.production.ini |
|
|
8
8
|
| [Design Demo Admin](features/design-demo-admin.md) | A self-serve admin UI at **`/design`** in the SSO-protected **Tools** app that lets the design team publish self-contained "Claude Design" HTML exports as **ver | tools/_/app/design/github.php, tools/mvc/design/get.php, tools/mvc/design/post.php, tools/assets/css/design.css, tools/assets/js/design.js, tools/_/app/frameworkindex.php, tools/_/app/nav.php, tools/composer.json |
|
|
9
9
|
| [Tools — Developers Folder (UUID & Password Generators)](features/developer-tools.md) | The first two tools shipped in the Tools app, both under the **Developers** folder and gated to personas **Development Team** / **TOGa Technology**. | tools/mvc/developers/uuid/get.php, tools/mvc/developers/password/get.php |
|
|
10
|
-
| [/errors Curation Console (Tools → shared Core Logs DB)](features/errors-curation-console.md) | Internal-only triage/curation screen for the 2.0 Issue/Event error-reporting pipeline, built as a 1.0 Tools MVC page reading the **shared Core Logs DB** through | tools/mvc/errors/get.php, tools/mvc/errors/post.php, tools/_/app/nav.php, tools/config.production.ini |
|
|
10
|
+
| [/errors Curation Console (Tools → shared Core Logs DB)](features/errors-curation-console.md) | Internal-only triage/curation screen for the 2.0 Issue/Event error-reporting pipeline, built as a 1.0 Tools MVC page reading the **shared Core Logs DB** through | tools/mvc/errors/get.php, tools/mvc/errors/post.php, tools/mvc/errors/issue/get.php, tools/assets/css/style.css, tools/_/app/nav.php, tools/config.production.ini |
|
|
11
11
|
| [Legacy Email Notifier (tools /email-migration/notify)](features/legacy-email-notifier.md) | An SSO-gated admin tool at **`/email-migration/notify`** (nav group **Email Migration** > **Legacy Notifier**, personas `['TOGa Technology','Development Team']` | tools/mvc/email-migration/notify/get.php, tools/mvc/email-migration/notify/post.php, tools/_/app/nav.php |
|
|
12
12
|
| [Tools MVC — Routing, CSRF & App_Database Access Patterns](features/mvc-data-access-patterns.md) | The load-bearing 1.0 (`App_`) framework conventions a developer needs when adding a page to the Tools app — URL routing, CSRF, and DB access through `App_Databa | tools/_/app/nav.php, tools/mvc/get.php, library/app/error.php |
|
|
13
13
|
| [Tools Persona-Gated Navigation (App_Nav)](features/persona-gated-navigation.md) | `App_Nav` is the Tools app's two-level, **persona-gated** navigation. | tools/_/app/nav.php, tools/mvc/get.php |
|
|
@@ -6,11 +6,13 @@ project: Tools
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-08-
|
|
9
|
+
updated: 2026-08-04
|
|
10
10
|
owners: ["jcardinal"]
|
|
11
11
|
files:
|
|
12
12
|
- tools/mvc/errors/get.php
|
|
13
13
|
- tools/mvc/errors/post.php
|
|
14
|
+
- tools/mvc/errors/issue/get.php
|
|
15
|
+
- tools/assets/css/style.css
|
|
14
16
|
- tools/_/app/nav.php
|
|
15
17
|
- tools/config.production.ini
|
|
16
18
|
related:
|
|
@@ -18,6 +20,7 @@ related:
|
|
|
18
20
|
- ./mvc-data-access-patterns.md
|
|
19
21
|
- ../../../2.0/apps/_underscore/features/error-reporting-issue-event.md
|
|
20
22
|
- ../../../2.0/apps/worker2/features/error-escalation-cron.md
|
|
23
|
+
- ../../library/features/error-capture-1-0.md
|
|
21
24
|
---
|
|
22
25
|
|
|
23
26
|
## Summary
|
|
@@ -41,6 +44,35 @@ happening and were deliberately never ticketed. Nobody had that view before.
|
|
|
41
44
|
- **Any save sets `isManaged`**, which does double duty: it protects the curated text from being
|
|
42
45
|
overwritten by later occurrences, **and** it exempts the Issue from the cron's GC sweep.
|
|
43
46
|
|
|
47
|
+
### Routing and pages
|
|
48
|
+
|
|
49
|
+
Filesystem routing (1.0 Tools MVC): `mvc/errors/` = **`/errors`** (the listing) and
|
|
50
|
+
`mvc/errors/issue/` = **`/errors/issue`** (the detail page).
|
|
51
|
+
|
|
52
|
+
**Listing** — 25 per page with a filtered `COUNT` (it was previously a silent `LIMIT 200`, so
|
|
53
|
+
older issues simply did not exist as far as a curator could tell); bold monospace Issue ID; an
|
|
54
|
+
**EB Environment** column read from the newest event's context; a **Type** column showing
|
|
55
|
+
`BUSINESS` or `TECHNICAL`; and search + three dropdown filters (including EB Environment) laid out
|
|
56
|
+
as a **3-column grid** — the original wrapping flex row broke to one control per line.
|
|
57
|
+
|
|
58
|
+
**`Type` needs no new column.** The presence of `Issue.issueKey` *is* the business/technical
|
|
59
|
+
signal. Do not add a `type` column for this (see the 2.0 doc's "proposed and not built" note).
|
|
60
|
+
|
|
61
|
+
**Issue detail** — a two-column split: technical evidence left, business curation (recipients,
|
|
62
|
+
curated text, affected clients) right, with an always-present Type chip.
|
|
63
|
+
|
|
64
|
+
- **Stack trace and context come from the MOST RECENT event**, not the Issue's first-seen
|
|
65
|
+
snapshot. The first occurrence is rarely the one you are debugging.
|
|
66
|
+
- **Context renders as cards**, one per top-level context group, ordered largest-first by recursive
|
|
67
|
+
node count, each a collapsible JSON tree built on native `<details>`/`<summary>` — **no
|
|
68
|
+
JavaScript** — with typed value colouring and `[REDACTED]` called out. This only works because
|
|
69
|
+
capture now emits a uniform object-per-top-level-key shape.
|
|
70
|
+
- A **view-layer** key humanizer (`elasticBeanstalkEnvironment` → "Elastic Beanstalk Environment";
|
|
71
|
+
ALL_CAPS `$_SERVER` keys left verbatim). It is presentation only — never rename keys at capture
|
|
72
|
+
time.
|
|
73
|
+
- Hourly-occurrence chart; fingerprints table re-sized for the narrow column with an icon-only
|
|
74
|
+
Move button.
|
|
75
|
+
|
|
44
76
|
**This page notifies nobody.** Escalation, ClickUp ticketing, and email are owned entirely by
|
|
45
77
|
`_Worker_Infrastructure_Errors::Escalate`. Adding a notify action here would create a second,
|
|
46
78
|
untracked alerting path.
|
|
@@ -80,11 +112,32 @@ They were renamed from plural on 2026-08-01, after the pipeline was already in p
|
|
|
80
112
|
- **Cross-tenant surface.** The Core Logs DB is shared by every client, so this console can show
|
|
81
113
|
one client's captured context to a viewer looking at another client's problem. It is
|
|
82
114
|
internal-only and persona-gated for that reason — do not expose any part of it to a client.
|
|
115
|
+
- **⚠ Never `JSON_EXTRACT` `Event.context`.** The column is `TEXT` and a context truncated at
|
|
116
|
+
`CONTEXT_MAX_LENGTH` is **invalid JSON** — `JSON_EXTRACT` would raise and take the whole page
|
|
117
|
+
down. The EB Environment filter therefore matches with **`SUBSTRING_INDEX`**, uses `EXISTS` (has
|
|
118
|
+
this issue *ever* occurred on that deployment), and validates the submitted value against the
|
|
119
|
+
discovered list — an allowlist, which also rules out `LIKE` wildcards.
|
|
120
|
+
- **The distinct-EB-environment query is a full scan of `Event`.** Acceptable only because that
|
|
121
|
+
table is bounded by the 7-day purge. At volume the proper fix is denormalising the EB environment
|
|
122
|
+
onto its own `Event` column **at capture time**.
|
|
123
|
+
- **Read both the old and new `Event.context` key positions.** The context shape was restructured
|
|
124
|
+
on 2026-08-04 (one object per top-level key, scalars under `Reference`); with 7-day retention,
|
|
125
|
+
both shapes coexist for a week after any such change.
|
|
83
126
|
- **Curated text is authoritative.** Once `isManaged` is set, later occurrences no longer update
|
|
84
127
|
subject/description, so a stale curated subject stays stale until someone edits it again.
|
|
85
128
|
|
|
86
129
|
## Change history
|
|
87
130
|
|
|
131
|
+
- 2026-08-04 — Console build-out: the listing gained real pagination (25/page with a filtered
|
|
132
|
+
`COUNT`, replacing a silent `LIMIT 200`), an EB Environment column **and filter**, a
|
|
133
|
+
BUSINESS/TECHNICAL Type column derived from `issueKey` (no new DB column), and a 3-column filter
|
|
134
|
+
grid. New `/errors/issue` detail page split technical/curation into two columns, showing the
|
|
135
|
+
**most recent** event's trace and context, with context rendered as largest-first cards over a
|
|
136
|
+
JavaScript-free `<details>` JSON tree, a view-layer key humanizer, a Type chip, and an
|
|
137
|
+
hourly-occurrence chart. Recorded the load-bearing gotcha that `Event.context` must be matched
|
|
138
|
+
with `SUBSTRING_INDEX` and never `JSON_EXTRACT` (a truncated context is invalid JSON and would
|
|
139
|
+
take the page down), that the filter value is allowlisted against the discovered list, and that
|
|
140
|
+
the distinct-values scan is only safe because `Event` is bounded by the 7-day purge. (jcardinal)
|
|
88
141
|
- 2026-08-01 — Updated every hardcoded table name in `tools/mvc/errors/get.php` (25+ refs) and
|
|
89
142
|
`post.php` (15+ refs) for the Core Logs singular rename (`Issues`→`Issue`, `Events`→`Event`,
|
|
90
143
|
`IssueFingerprints`→`IssueFingerprint`, `IssueClickupTasks`→`IssueClickupTask`,
|
|
@@ -6,7 +6,7 @@ project: Tools
|
|
|
6
6
|
client: shared
|
|
7
7
|
type: feature
|
|
8
8
|
status: active
|
|
9
|
-
updated: 2026-
|
|
9
|
+
updated: 2026-08-04
|
|
10
10
|
owners: [jcardinal, ajean]
|
|
11
11
|
files:
|
|
12
12
|
- tools/mvc/talos/kb-documents/get.php
|
|
@@ -303,6 +303,13 @@ Refactor from a single hard-coded `development-team` KB to per-AI-model, data-dr
|
|
|
303
303
|
idempotent, **fails the deploy loudly** if `pgsql` won't load) and
|
|
304
304
|
`hooks/postdeploy/01-restart-php.sh` (restarts php-fpm/httpd). **EB runs `.platform/hooks`
|
|
305
305
|
regardless of the git exec bit** — worker2's equivalents are committed `100644`.
|
|
306
|
+
> **Rewritten 2026-08-04 (in `tools` and `worker2`) — it failed on PHP 8.5.** The old hook tried
|
|
307
|
+
> the **generic** `php-pgsql` first, which resolves against the default 8.1 stream and can
|
|
308
|
+
> conflict; it also discarded `dnf` output and only checked that `pgsql` loaded, not
|
|
309
|
+
> **`pdo_pgsql`** (the extension `App_Pg` actually needs). It now tries the **versioned**
|
|
310
|
+
> candidate first, captures and prints `dnf` output so a failure is diagnosable from
|
|
311
|
+
> `eb-hooks.log`, checks **both** extensions, and honours a `PGSQL_HOOK_OPTIONAL=1` escape hatch
|
|
312
|
+
> for environments that must deploy without Postgres.
|
|
306
313
|
- Apply the dbchanges2 `Team/2026-06-30a..e` + `Core/2026-06-30a` migrations in order.
|
|
307
314
|
- Verify the Bedrock KB region / data-source name against the live account (TODOs in code).
|
|
308
315
|
|
|
@@ -397,6 +404,11 @@ Refactor from a single hard-coded `development-team` KB to per-AI-model, data-dr
|
|
|
397
404
|
|
|
398
405
|
## Change history
|
|
399
406
|
|
|
407
|
+
- 2026-08-04 — Rewrote the `prebuild/01-install-php-pgsql.sh` hook (here and in `worker2`), which
|
|
408
|
+
failed on PHP 8.5: it tried the generic `php-pgsql` (resolving against the default 8.1 stream),
|
|
409
|
+
swallowed `dnf` output, and checked only `pgsql` rather than `pdo_pgsql`. Now tries the versioned
|
|
410
|
+
candidate first, prints `dnf` output, verifies both extensions, and supports
|
|
411
|
+
`PGSQL_HOOK_OPTIONAL=1`. (jcardinal)
|
|
400
412
|
- 2026-07-29 — **Fixed the empty Vocabulary page: the per-model scoping shipped with no data
|
|
401
413
|
backfill.** `vocabulary/get.php` reads strictly `WHERE c_trueAiModelId = <model>` (no NULL
|
|
402
414
|
fallback), but `Team/2026-07-27a` added the column `DEFAULT NULL` and nothing populated it —
|
|
@@ -8,6 +8,7 @@
|
|
|
8
8
|
| [Elite TOGA 2.0 → TOGaDeskSupport Standalone Attachment Sync](features/elite-togadesk-attachment-sync.md) | `sync_togadesk_elite_attachments.php` is a standalone cron (every 5 minutes) that syncs file attachments from TOGA 2.0 into TOGaDeskSupport for Elite. | worker/crons/toga2/elite/sync_togadesk_elite_attachments.php, worker/crons/toga2/elite/test_sync_togadesk_elite_attachments.php |
|
|
9
9
|
| [Forecast2 ↔ NetSuite Reconciliation & Trueup Tooling](features/forecast2-netsuite-reconciliation.md) | CLI tools to **audit** and **repair** drift between the production `Forecast` DB (core2) and NetSuite. | test/@dave/checker.php, worker2/Component/Forecast/SaleImport/SaleImport.php, test/@dave/looper.php, test/@dave/reconcile_netsuite_totals.php, test/@dave/fixer.php, test/@dave/analyze_netsuite_forecast_diff.php, test/@dave/trueup_sales.php, test/@dave/reconcile_drift_2023plus.php, test/@dave/probe_invoice_gap_2026.php, test/@dave/probe_creditmemo_gap_detail.php, test/@dave/trueup_open_orders.php, test/@dave/loop_trueup_open_orders.php, test/@dave/trueup_opportunities.php, test/@dave/probe_sales_gap_direct.php, test/@dave/probe_missing_oo_timing.php, test/@dave/probe_missing_oo_createdby.php, test/@dave/probe_drift_so_dates.php, test/@dave/probe_profit_invoices.php, test/@dave/probe_profit_gap.php, worker/crons/toga2/forecast2/common_import_sales_from_netsuite.php, worker/crons/toga2/forecast2/periodic_forecast_discrepancy_fix_open_orders.php, worker/crons/toga2/forecast2/import_open_orders.php, worker/schedules/cron.worker.infrastructure.json |
|
|
10
10
|
| [NetSuite → TOGa Supply Per-Client Sync (thin wrappers)](features/netsuite-togasupply-per-client-sync.md) | Syncs NetSuite transactions (sales orders, purchase orders, invoices, item receipts, item fulfillments, inventory adjustments) into each TOGa Supply (2.0) clien | worker/crons/toga2/netsuite/common_sync_togasupply.php, worker/crons/toga2/netsuite/sync_togasupply_canon.php, worker/schedules/cron.worker.sync.json, dbchanges2/_modules/netsuite/2026-04-01 - Parameters.sql, library/app/api/netsuite/rest.php, library/app/systemmonitor/netsuiteintegration.php |
|
|
11
|
+
| [OneUptime Server monitor + disk/memory hygiene on the 1.0 worker EB host](features/oneuptime-server-monitor-host-hygiene.md) | The 1.0 `agilant-worker` EB environment runs on the **legacy Amazon Linux 1 PHP 7.2 platform** (Apache httpd/prefork, s3fs mounts, cron) and repeatedly went dow | worker/.ebextensions/040_disk_memory_hygiene.config, worker/.ebextensions/045_oneuptime_agent.config, worker/ebs/cron.worker.php, worker/ebs/mount-s3fs-folders.php, worker/ebs/apache_settings.php, worker/ebs/setup_phpini.php |
|
|
11
12
|
| [OneUptime external uptime monitoring for 1.0 workers](features/oneuptime-worker-uptime-monitoring.md) | Every 1.0 worker box self-reports its liveness to an external OneUptime monitor once per minute by curl-POSTing to a per-worker "Incoming Request" heartbeat URL | library/app/worker.php, worker/crons/worker/worker_heartbeat.php |
|
|
12
13
|
| [Prudential: Send Shipments for the Day report (daily cron)](features/send-shipments-for-the-day.md) | Daily cron (9:00 PM) that emails Prudential and Dell stakeholders an Excel report of all devices shipped that day, including tracking number, serial number, emp | worker/crons/notifications/reports/send_shipments_for_the_day.php |
|
|
13
14
|
| [Diagnosing frozen 1.0 worker cron check-ins (Sentry "missed" flood)](workflows/diagnosing-frozen-cron-checkins.md) | When 1.0 worker cron timestamps freeze and Sentry project `worker1` fills with **`missed`** check-ins, the intuitive diagnosis — a wedged `App_Framework::isProc | worker/.ebextensions/cron.config, library/app/worker.php |
|
|
@@ -0,0 +1,204 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: OneUptime Server monitor + disk/memory hygiene on the 1.0 worker EB host
|
|
3
|
+
framework: "1.0"
|
|
4
|
+
repo: worker
|
|
5
|
+
project: Worker
|
|
6
|
+
client: shared
|
|
7
|
+
type: feature
|
|
8
|
+
status: active
|
|
9
|
+
updated: 2026-08-03
|
|
10
|
+
owners: ["jcardinal"]
|
|
11
|
+
files:
|
|
12
|
+
- worker/.ebextensions/040_disk_memory_hygiene.config
|
|
13
|
+
- worker/.ebextensions/045_oneuptime_agent.config
|
|
14
|
+
- worker/ebs/cron.worker.php
|
|
15
|
+
- worker/ebs/mount-s3fs-folders.php
|
|
16
|
+
- worker/ebs/apache_settings.php
|
|
17
|
+
- worker/ebs/setup_phpini.php
|
|
18
|
+
related: ["oneuptime-worker-uptime-monitoring"]
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
## Summary
|
|
22
|
+
|
|
23
|
+
The 1.0 `agilant-worker` EB environment runs on the **legacy Amazon Linux 1 PHP 7.2
|
|
24
|
+
platform** (Apache httpd/prefork, s3fs mounts, cron) and repeatedly went down from the EC2
|
|
25
|
+
instance running **out of memory and out of disk**. Two things address it:
|
|
26
|
+
|
|
27
|
+
1. **Host-level observability** — a OneUptime **Server monitor** per worker role, fed by the
|
|
28
|
+
OneUptime **Infrastructure Agent** running on the box (CPU / memory / disk / network /
|
|
29
|
+
process metrics every 30s). This is *not* a heartbeat — see
|
|
30
|
+
[[oneuptime-worker-uptime-monitoring]] for the separate Incoming-Request heartbeat layer.
|
|
31
|
+
2. **Disk/memory hygiene** — a `.ebextensions` config that reaps stale files and rotates
|
|
32
|
+
logs (safe/additive), plus a set of **behavior-changing fixes that are documented but
|
|
33
|
+
NOT applied** (Apache prefork sizing, PHP `memory_limit`, s3fs `ensure_diskfree`,
|
|
34
|
+
cron `flock`).
|
|
35
|
+
|
|
36
|
+
## Key files / entry points
|
|
37
|
+
|
|
38
|
+
| File | Purpose | Risk |
|
|
39
|
+
|---|---|---|
|
|
40
|
+
| `.ebextensions/040_disk_memory_hygiene.config` | Deletes stale junk (s3fs `/tmp` cache, `/var/www/cache` staging, orphaned `WORKER_ERROR_*`, deploy litter) + rotates logs | Safe / additive — deletes only stale files |
|
|
41
|
+
| `.ebextensions/045_oneuptime_agent.config` | Installs + supervises the OneUptime infrastructure-agent (static Go binary, v11.7.4) | Safe / additive — no app behavior change |
|
|
42
|
+
|
|
43
|
+
## How the Server monitor works
|
|
44
|
+
|
|
45
|
+
- The agent POSTs metrics every **30s** to
|
|
46
|
+
`{oneuptime_url}/server-monitor/response/ingest/{secret_key}`; OneUptime evaluates the
|
|
47
|
+
monitor's **Criteria** against them to decide Operational / Degraded / Offline.
|
|
48
|
+
- **AL1 has no systemd**, so `045_…` deliberately skips the agent's `configure`/`start`
|
|
49
|
+
subcommands (they register a systemd/SysV service and fail here). It writes
|
|
50
|
+
`/etc/oneuptime-infrastructure-agent/config.json` directly and runs the foreground **`run`**
|
|
51
|
+
subcommand under a **1-minute cron watchdog** — the same systemd-bypass the ODP Voice
|
|
52
|
+
container monitor uses.
|
|
53
|
+
- Metrics reported: CPU (aggregate + per-core + user/system/idle/iowait/steal breakdown),
|
|
54
|
+
load average 1/5/15, memory + swap, **disk per mount** (total/used/free/%, device, fstype,
|
|
55
|
+
I/O counters — criteria target a specific path), network (aggregate + per-interface bytes/
|
|
56
|
+
packets/errors/drops, TCP connection counts; cumulative since boot, OneUptime derives
|
|
57
|
+
throughput from deltas), host info (platform/kernel/uptime/virtualization), and
|
|
58
|
+
**per-process** pid/name/command/cpu%/mem%/status/threads/user (enables
|
|
59
|
+
"process is executing" criteria).
|
|
60
|
+
|
|
61
|
+
## Provisioning — already done, per-role model
|
|
62
|
+
|
|
63
|
+
`agilant-worker` is one EB Worker tier whose instances each claim ONE stable role
|
|
64
|
+
(`ebs/cron.worker.php`), so there is **one Server monitor per role (7)**, grouped under the
|
|
65
|
+
OneUptime label **`Worker 1.0`**. Each monitor already carries the criteria below, and its
|
|
66
|
+
secret is **hardcoded in `045_oneuptime_agent.config`**, selected at runtime by the role in
|
|
67
|
+
`/etc/worker-role`. Nothing to create — you only need to deploy.
|
|
68
|
+
|
|
69
|
+
| Role | Monitor name |
|
|
70
|
+
|------|--------------|
|
|
71
|
+
| notification | Notification Worker 1.0 - System Health |
|
|
72
|
+
| database | Database Worker 1.0 - System Health |
|
|
73
|
+
| infrastructure | Infrastructure Worker 1.0 - System Health |
|
|
74
|
+
| toga | TOGa Worker 1.0 - System Health |
|
|
75
|
+
| togadesk | TOGa Desk Worker 1.0 - System Health |
|
|
76
|
+
| sync | Sync Worker 1.0 - System Health |
|
|
77
|
+
| catalog | Catalog Worker 1.0 - System Health |
|
|
78
|
+
|
|
79
|
+
**Key rotation:** regenerate that monitor's secret in OneUptime and update the matching
|
|
80
|
+
`case` arm in `045_…`. Instance-ids churn but roles are stable, so a recycled box
|
|
81
|
+
re-reports to the same role monitor automatically.
|
|
82
|
+
|
|
83
|
+
**Creating a monitor from scratch** (only if a new role is added): Monitors → Create Monitor
|
|
84
|
+
→ type **Server / VM** → the secret key appears on the monitor's Settings tab after save.
|
|
85
|
+
Via REST it is `POST /api/monitor` with `{"data":{"projectId":…,"name":…,"monitorType":"Server"}}`
|
|
86
|
+
and an `ApiKey` header — but the secret is server-generated and **not** returned in the
|
|
87
|
+
create body; read it from Settings.
|
|
88
|
+
|
|
89
|
+
**Verify after deploy**, on the instance:
|
|
90
|
+
|
|
91
|
+
```bash
|
|
92
|
+
pgrep -fa oneuptime-infrastructure-agent # process running
|
|
93
|
+
tail -n 40 /var/log/oneuptime-agent.log # agent log
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
Last Ping updates within ~1 min.
|
|
97
|
+
|
|
98
|
+
## Criteria (order matters — first match wins, catch-all Operational last)
|
|
99
|
+
|
|
100
|
+
| Order | Result state | Check | Condition | Threshold |
|
|
101
|
+
|---|---|---|---|---|
|
|
102
|
+
| 1 | Offline | CPU Usage % | Greater Than | 90 |
|
|
103
|
+
| 2 | Offline | Memory Usage % | Greater Than | 90 |
|
|
104
|
+
| 3 | Degraded | Disk Usage % (path `/`) | Greater Than | 90 |
|
|
105
|
+
| 4 | Degraded | CPU Usage % | Greater Than | 75 |
|
|
106
|
+
| 5 | Degraded | Memory Usage % | Greater Than | 75 |
|
|
107
|
+
| 6 | Operational (fallback) | CPU Usage % | Less Than Or Equal To | 100 |
|
|
108
|
+
|
|
109
|
+
Each criterion's filter condition is **All**. Recommended additions for this worker
|
|
110
|
+
specifically: Degraded if disk `/` > **80%** (earlier warning — disk is this box's actual
|
|
111
|
+
failure mode), Degraded if swap % used > 50% (early OOM signal once swap exists), and
|
|
112
|
+
Offline if the Apache process is not executing (Server Process → Command contains `httpd`
|
|
113
|
+
→ Is Not Executing).
|
|
114
|
+
|
|
115
|
+
Via API, criteria live in `monitorSteps.monitorCriteria.criteriaInstances[]`; each instance
|
|
116
|
+
has `filters[]` (`checkOn`, operator, `value`, plus `serverMonitorOptions.diskPath` for disk
|
|
117
|
+
checks), a `filterCondition` (`All`/`Any`), and a `monitorStatusId`.
|
|
118
|
+
|
|
119
|
+
## Root cause of the outages
|
|
120
|
+
|
|
121
|
+
### Disk
|
|
122
|
+
|
|
123
|
+
| Cause | Evidence | Status |
|
|
124
|
+
|---|---|---|
|
|
125
|
+
| s3fs `use_cache=/tmp` **unbounded**, never evicted | `ebs/mount-s3fs-folders.php:53`, `ebs/s3fs.json` (3 buckets) | **Not fixed** — needs `ensure_diskfree` (below) |
|
|
126
|
+
| `/var/www/cache` 0777 with no reaper; DB dumps up to **12 GB/part**, deleted only on successful upload | `ebs/setup_export_cache_folders.php:11-13`, `crons/database/database_backup.php:75,314,337` | Reaper cron shipped in `040_…` |
|
|
127
|
+
| `WORKER_ERROR_*` reaper never scheduled | `crons/worker/check_worker_errors.php` — present in **no** `schedules/*.json` | Reaper cron shipped in `040_…` |
|
|
128
|
+
| Deploy litter (`/tmp/phpredis*`, s3fs source) | `010_setup_redis.config:3-6`, `006_mount-s3fs.config:16-29` | One-off cleanup shipped in `040_…` |
|
|
129
|
+
|
|
130
|
+
### Memory
|
|
131
|
+
|
|
132
|
+
| Cause | Evidence | Status |
|
|
133
|
+
|---|---|---|
|
|
134
|
+
| Apache prefork massively over-provisioned: `MaxRequestWorkers = ceil(3200*3/37) = 260` while PHP `memory_limit = 2G` | `ebs/apache_settings.php:3,8` + `ebs/setup_phpini.php:22` | **Not fixed** |
|
|
135
|
+
| Long CLI cron jobs overlap (minute-ly schedules, no `flock`), each entitled to 2 GB | `ebs/cron.worker.php`, `schedules/cron.worker*.json` (26–103 jobs) | **Not fixed** |
|
|
136
|
+
| Standing band-aid: Apache `graceful` every 20 min "to close db connections" | `schedules/cron.worker.json:3-7` | **Not fixed** — would be retired by `MaxConnectionsPerChild` |
|
|
137
|
+
|
|
138
|
+
## Fixes documented but NOT applied — review before applying
|
|
139
|
+
|
|
140
|
+
These change app behavior. Apply deliberately.
|
|
141
|
+
|
|
142
|
+
**Bound the s3fs cache** — `ebs/mount-s3fs-folders.php`, after the `use_cache` option is
|
|
143
|
+
appended (line 53):
|
|
144
|
+
|
|
145
|
+
```php
|
|
146
|
+
$mount_command .= ' -o ' . escapeshellarg('ensure_diskfree=2000'); // MB; s3fs evicts its own cache to keep >=2GB free
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
**Right-size Apache** — `ebs/apache_settings.php`: measure the real child size (the command
|
|
150
|
+
is in the file at line 4 — with `memory_limit=2G`, heavy endpoints are far bigger than the
|
|
151
|
+
assumed 37 MB), then replace the `× 3` overcommit on line 8 with
|
|
152
|
+
`MaxRequestWorkers = floor(availableRAM_MB * 0.7 / measuredChildMB)` (≈20–25 on a 4 GB box,
|
|
153
|
+
not 260) and lower `StartServers`/`MinSpareServers` to 4–8. Add `MaxConnectionsPerChild 500`
|
|
154
|
+
inside the `<IfModule prefork.c>` block so children recycle and return leaked memory to the
|
|
155
|
+
OS — this is what retires the 20-minute graceful-restart hack.
|
|
156
|
+
|
|
157
|
+
**Lower cron memory pressure** — `ebs/setup_phpini.php`: drop the global
|
|
158
|
+
`memory_limit = 2G` (line 22) to `512M` and raise it per-script via `ini_set()` only in the
|
|
159
|
+
few CLI batch jobs that need it; set `display_errors = off` (line 20) in production, logging
|
|
160
|
+
to a rotated file instead.
|
|
161
|
+
|
|
162
|
+
**Stop cron overlap** — where `ebs/cron.worker.php` builds each cron command, prefix
|
|
163
|
+
minute-ly/2-minute jobs with a non-blocking lock:
|
|
164
|
+
`flock -n /tmp/lock.<jobname> php /var/www/html/crons/<jobname>`.
|
|
165
|
+
|
|
166
|
+
**Swap (optional stopgap)** — only *after* the disk reapers are in place; this box is
|
|
167
|
+
disk-constrained and a 2 GB swapfile on a near-full root volume makes disk worse:
|
|
168
|
+
|
|
169
|
+
```yaml
|
|
170
|
+
commands:
|
|
171
|
+
02_swap:
|
|
172
|
+
test: "[ ! -f /var/swapfile ]"
|
|
173
|
+
command: "dd if=/dev/zero of=/var/swapfile bs=1M count=2048 && chmod 600 /var/swapfile && mkswap /var/swapfile && swapon /var/swapfile"
|
|
174
|
+
```
|
|
175
|
+
|
|
176
|
+
## Gotchas
|
|
177
|
+
|
|
178
|
+
- **Disk path `/` works here** — this is a real EC2 host, unlike the ODP Voice container
|
|
179
|
+
case where "Disk metrics for path / were unavailable" appears because the container root
|
|
180
|
+
isn't in the reported disk list. s3fs (fuse) mounts may **not** be reported as normal
|
|
181
|
+
partitions, so prefer alerting on `/` rather than the s3fs mount.
|
|
182
|
+
- The agent is kept alive by a **cron watchdog, not a service** — if it stops reporting,
|
|
183
|
+
check the watchdog and `/var/log/oneuptime-agent.log`, not `service`/`systemctl`.
|
|
184
|
+
- Monitor secrets are hardcoded per-role in `045_…`; treat that file as a credential
|
|
185
|
+
location. Do not copy secrets into docs, logs, or other files.
|
|
186
|
+
- **Security follow-up (urgent, out of scope of the monitoring work):** hardcoded live AWS
|
|
187
|
+
credentials and DB passwords are committed in plaintext in `ebs/cron.worker.php:14-15`,
|
|
188
|
+
`crons/worker/worker_heartbeat.php:12`, `crons/database/database_backup*.php`, and
|
|
189
|
+
`config.worker.ini`. Rotate and move to EB environment properties / an IAM instance role.
|
|
190
|
+
|
|
191
|
+
## References
|
|
192
|
+
|
|
193
|
+
- Agent + subcommands — https://github.com/OneUptime/oneuptime/blob/master/InfrastructureAgent/README.md
|
|
194
|
+
- Server monitor — https://oneuptime.com/docs/monitor/server-monitor
|
|
195
|
+
- Criteria instance API — https://oneuptime.com/reference/monitor-criteria-instance
|
|
196
|
+
- Monitor REST API — https://oneuptime.com/reference/monitor
|
|
197
|
+
- Agent releases (v11.7.4) — https://github.com/OneUptime/oneuptime/releases
|
|
198
|
+
|
|
199
|
+
## Change history
|
|
200
|
+
- 2026-08-03 — Captured from `worker/docs/oneuptime-server-monitor.md` (that in-repo doc is
|
|
201
|
+
being deleted; this is now the source of truth). Documents the per-role OneUptime Server
|
|
202
|
+
monitors + infrastructure agent (`045_…`), the shipped disk/memory hygiene reapers
|
|
203
|
+
(`040_…`), the diagnosed disk/memory root causes, and the reviewed-but-unapplied
|
|
204
|
+
behavior-changing fixes. (jcardinal)
|
|
@@ -5,11 +5,12 @@ project: Library
|
|
|
5
5
|
client: shared
|
|
6
6
|
type: standard
|
|
7
7
|
status: active
|
|
8
|
-
updated: 2026-
|
|
8
|
+
updated: 2026-08-04
|
|
9
9
|
owners: [jcardinal, rgirish, mhammontree]
|
|
10
10
|
files: []
|
|
11
11
|
related:
|
|
12
12
|
- ../apps/library/architecture.md
|
|
13
|
+
- ../apps/library/features/error-capture-1-0.md
|
|
13
14
|
- ../../2.0/standards/backend-php.md
|
|
14
15
|
---
|
|
15
16
|
|
|
@@ -259,6 +260,49 @@ if ($condition) {
|
|
|
259
260
|
$piece = substr($value, 0, 8);
|
|
260
261
|
```
|
|
261
262
|
|
|
263
|
+
### No typed properties (PHP 7.2)
|
|
264
|
+
|
|
265
|
+
* 1.0 runs **PHP 7.2**, so **typed properties (7.4+) are forbidden anywhere in `library/` or a 1.0
|
|
266
|
+
app.** Declare class attributes untyped with an `@var` docblock:
|
|
267
|
+
|
|
268
|
+
```php
|
|
269
|
+
// WRONG in 1.0 — deploy fails with "syntax error, unexpected '?'"
|
|
270
|
+
private static ?string $memoryReserve = null;
|
|
271
|
+
|
|
272
|
+
// CORRECT in 1.0
|
|
273
|
+
/** @var string|null */
|
|
274
|
+
private static $memoryReserve = null;
|
|
275
|
+
```
|
|
276
|
+
|
|
277
|
+
* **Parameter and return types ARE fine** (PHP 7.0/7.1) — this rule is only about *properties*.
|
|
278
|
+
* This is the opposite of the 2.0 standard, which requires typed properties. The universal
|
|
279
|
+
`rules/toga/common/coding-style.md` rule is already **scoped to 2.0** and cross-references this
|
|
280
|
+
standard for the 1.0 `@var` docblock form — so there is no conflict to resolve, just two different
|
|
281
|
+
rules for two different runtimes.
|
|
282
|
+
* Real failure: a deploy of `library/app/error/capture.php` broke on
|
|
283
|
+
`private static ?string $memoryReserve`; nine typed properties across three files had to be
|
|
284
|
+
converted.
|
|
285
|
+
|
|
286
|
+
### Any PHP warning terminates the request — how to opt out locally
|
|
287
|
+
|
|
288
|
+
`App_Error::handleError()` routes **every** PHP warning into `handleException()`, which renders the
|
|
289
|
+
error page and calls **`exit()`**. (2.0's equivalent *throws*, which a `catch` can absorb — 1.0 does
|
|
290
|
+
not give you that.) So a warning raised inside code that must not escalate — error capture, a
|
|
291
|
+
monitoring ping, an IMDS probe where `file_get_contents` warns before returning `false`, a retry
|
|
292
|
+
loop — **kills the request or cron mid-run while reporting an unrelated error.**
|
|
293
|
+
|
|
294
|
+
Two sanctioned patterns, in order of preference:
|
|
295
|
+
|
|
296
|
+
1. **A local handler around the risky call** — `set_error_handler()` throwing an `ErrorException`,
|
|
297
|
+
then `restore_error_handler()`. Use this when you want to *catch and retry* (the NetSuite
|
|
298
|
+
SSL-drop retry loop does this).
|
|
299
|
+
2. **`App_Error::setThrowExceptionsEnabled(false)` with a `finally` restore** — use this when the
|
|
300
|
+
code must never escalate at all (`App_Error_Capture::captureException()` does this for its whole
|
|
301
|
+
duration). This disables the **handler**, not real `throw`s: `App_Query` raises a plain
|
|
302
|
+
`throw new Exception` on a DB error, so genuine failures still reach your `catch`.
|
|
303
|
+
|
|
304
|
+
Never leave either one in effect beyond the block that needs it.
|
|
305
|
+
|
|
262
306
|
### Documentation
|
|
263
307
|
|
|
264
308
|
* **DocBlocks:** framework (`library`) classes/methods are well documented with a description and `@param`/`@return`/`@throws`. Application code is less consistent — document new and modified methods to the framework standard.
|
|
@@ -17,7 +17,7 @@
|
|
|
17
17
|
| [Re-pointing a DB alias mid-request (_Database::register park/restore)](features/database-alias-repointing.md) | `_Database` keys **all live per-database runtime state by the connection ALIAS** (`Client` / `_underscore::DB_CLIENT`, `ClientLogs`, `Archive`), **not** by the | _underscore/Database.php, _underscore/Query.php, api2/Component/Api/V2/V2.php, api2/Component/Api/CrossClient/CrossClient.php |
|
|
18
18
|
| [2.0 Email Send Pipeline (queue + Send worker)](features/email-send-pipeline.md) | In 2.0, `_Email::send()` **does not transmit** — it queues the message. | _underscore/Email.php, worker2/Worker/Infrastructure/Email/Send.php |
|
|
19
19
|
| [Client Email Template Sending](features/email-template-sending.md) | `_Model_Client_EmailTemplate` sends a stored, client-defined email template by UUID. | _underscore/Model/Client/EmailTemplate.php, _underscore/Model/Client/EmailTemplateOutgoingEmailAddress.php, _underscore/Email.php |
|
|
20
|
-
| [Error Reporting — Issue/Event Capture, Fingerprinting & Aggregation](features/error-reporting-issue-event.md) | Platform-wide error reporting for TOGA 2.0, built on an **Issue / Event** aggregation model in the **shared Core Logs DB**. | _underscore/Error.php, _underscore/Exception/Business.php, _underscore/Model/Core/Logs/Issue.php, _underscore/Model/Core/Logs/Event.php, _underscore/Model/Core/Logs/IssueFingerprint.php, _underscore/Model/Core/Logs/IssueClickupTask.php, _underscore/Model/Core/Logs/IssueEmailAddress.php, _underscore/Model/Core/Logs/IssueAreaOwner.php, dbchanges2/Logs/2026-07-30a - Error reporting Issues and Events.sql, dbchanges2/Logs/2026-08-01a - Rename tables to singular.sql, dbchanges2/Logs_Client/2026-08-01a - Drop Error table.sql, dbchanges2/Core/2026-07-30a - Error escalation cron job.sql |
|
|
20
|
+
| [Error Reporting — Issue/Event Capture, Fingerprinting & Aggregation](features/error-reporting-issue-event.md) | Platform-wide error reporting for TOGA 2.0, built on an **Issue / Event** aggregation model in the **shared Core Logs DB**. | _underscore/Error.php, _underscore/Database.php, _underscore/Exception/Business.php, api2/Controller/Index.php, worker2/Controller/Index.php, _underscore/Model/Core/Logs/Issue.php, _underscore/Model/Core/Logs/Event.php, _underscore/Model/Core/Logs/IssueFingerprint.php, _underscore/Model/Core/Logs/IssueClickupTask.php, _underscore/Model/Core/Logs/IssueEmailAddress.php, _underscore/Model/Core/Logs/IssueAreaOwner.php, dbchanges2/Logs/2026-07-30a - Error reporting Issues and Events.sql, dbchanges2/Logs/2026-08-01a - Rename tables to singular.sql, dbchanges2/Logs_Client/2026-08-01a - Drop Error table.sql, dbchanges2/Core/2026-07-30a - Error escalation cron job.sql |
|
|
21
21
|
| [Record-Changed Event Publishing (_Event::publish to SQS)](features/event-publish-sqs.md) | `_Event::publish()` (in `_underscore/Event.php`) is the PHP side of the real-time event pipeline. | _underscore/Event.php |
|
|
22
22
|
| [Forecast.Sales NetSuite import engine (real-time webhook)](features/forecast-sale-import.md) | Real-time importer that takes a NetSuite **sale** record and writes its lines into `Forecast.Sales` (the Forecast2 revenue table). | worker2/Component/Forecast/SaleImport/SaleImport.php, worker2/Component/Forecast/Db/Db.php, _underscore/Component/Api/Netsuite/Netsuite.php, worker2/Worker/Netsuite/Invoice.php, worker2/Worker/Netsuite/CashSale.php, worker2/Worker/Netsuite/CreditMemo.php, worker2/Worker/Netsuite/CashRefund.php, worker2/Worker/Netsuite/JournalEntry.php, worker2/Worker/Netsuite/Opportunity.php, worker2/Worker/Netsuite/SalesOrder.php, dbchanges2/Forecast/2026-06-26a - Add journalEntry to Sales transaction type enum.sql, test/@dave/test_invoice_lifecycle.php, test/@dave/test_je_lifecycle.php, test/@dave/test_creditmemo_lifecycle.php, test/@dave/test_cashsale_lifecycle.php, test/@dave/test_cashrefund_lifecycle.php, test/@dave/test_fetchrecord_routes.php, test/@dave/verify_je_classification.php, test/@dave/probe_je_accounts.php, test/@dave/probe_je_shape.php, test/@dave/fixer.php, test/@dave/Junk Drawer/NetSuite/api-message-queue/ue_api_msg_queue_enqueue.js, test/@dave/Junk Drawer/NetSuite/api-message-queue/dev_ue_api_msg_queue_enqueue.js |
|
|
23
23
|
| [isFulfillable Propagation Up the SO↔PO Chain](features/fulfillable-item-propagation.md) | `Items.isFulfillable` is a boolean that gates whether a storefront line's **Qty Fulfilled** cell is actionable. | _underscore/Model/Client/Item.php, _underscore/Model/Compass/Item.php, dbchanges2/Core/2026-07-17 - Items isFulfillable RecordField.sql, dbchanges2/Core/2026-07-17 - RegisterItemIsFulfillableInterceptors.sql, dbchanges2/Client/2026-07-17 - ItemsisFulfillable.sql |
|