toga-ai 1.0.672 → 1.0.674

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -6,7 +6,7 @@ project: SAML SSO Gateway
6
6
  client: shared
7
7
  type: architecture
8
8
  status: active
9
- updated: 2026-06-11
9
+ updated: 2026-08-27
10
10
  owners: ["rgirish", "jcardinal"]
11
11
  files:
12
12
  - saml/index.php
@@ -71,6 +71,14 @@ To onboard a new SSO client, add the class in `_underscore` — not in this repo
71
71
  - **`_underscore` pulled at deploy time** via `.ebextensions/git.php` + prebuild hook — clones branch `_production`. Deploy `saml` after `_underscore` merges to pick up framework changes.
72
72
  - **CodePipeline:** Source (GitHub `_production`) → ManualApproval → Deploy (EB).
73
73
  - **Composer deps** installed at postdeploy via `015_install_composer.sh`.
74
+ - **Non-prod shared-ALB self-registration (2026-08-27).** Postdeploy hook
75
+ `060_register_instance_to_shared_application_load_balancer.sh` + `ebs/register_instance_to_shared_application_load_balancer.php`
76
+ register a non-production `saml` instance into the target group named exactly its EB env name;
77
+ **skips `production`** (so `saml-production` is unaffected) and always `exit 0`. Requires
78
+ `aws/aws-sdk-php` (now in `composer.json`, lock/vendor to be regenerated) and two IAM actions on
79
+ the instance profile (`elasticloadbalancing:DescribeTargetGroups` region-scoped +
80
+ `RegisterTargets` scoped to the non-prod target-group ARN). See
81
+ [ALB target-group auto-registration](../worker2/features/alb-target-group-auto-registration.md).
74
82
 
75
83
  ## Dependencies
76
84
 
@@ -88,3 +96,4 @@ To onboard a new SSO client, add the class in `_underscore` — not in this repo
88
96
  ## Change history
89
97
  - 2026-06-11 — Added `transactionCommit()` before redirect; wrapped `getAuthenticatedSsoUser()` in try/catch with Sentry; echo+exit on auth failure (rgirish)
90
98
  - 2026-06-25 — Sharpened signature-verification gotcha with explicit threat model and flagged it as the repo's top security priority; documented plaintext secrets in `production.ini` + hardcoded Sentry DSN in `Controller/Index.php` with SSM Parameter Store recommendation (jcardinal)
99
+ - 2026-08-27 — Documented non-prod shared-ALB self-registration postdeploy hook (`060_...sh` + `ebs/register_instance_to_shared_application_load_balancer.php`); skips `production`, requires `aws/aws-sdk-php` + two IAM actions; cross-linked worker2 ALB auto-registration feature (jcardinal)
@@ -3,8 +3,9 @@
3
3
  | Doc | Summary | Files |
4
4
  |-----|---------|-------|
5
5
  | [Worker (worker2) Architecture](architecture.md) | Worker (repo `worker2`) is an AWS Elastic Beanstalk **Worker Tier** application that processes background jobs. | worker2/Controller/Index.php, worker2/Worker/, worker2/LambdaFunctions/, _underscore/Worker.php, worker2/composer.json |
6
- | [Deploy-Time Auto-Registration to the Shared ALB Target Group (non-production)](features/alb-target-group-auto-registration.md) | TOGA does **not** pay for EB-managed load-balancer registration, so an EB instance is normally **not** added to its environment's ALB target group — a fresh or | worker2/.platform/hooks/postdeploy/060_register_instance_to_shared_application_load_balancer.sh, worker2/ebs/register_instance_to_shared_application_load_balancer.php, worker2/.platform/hooks/prebuild/_shared/040-write-instance-id.sh, worker2/.platform/hooks/prebuild/_shared/041-write-region.sh, worker2/.platform/hooks/prebuild/_shared/042-write-eb-environment.sh, api2/.platform/hooks/prebuild/_shared/042-write-eb-environment.sh, worker2/.platform/hooks/postdeploy/015_install_composer.sh, api2/.platform/hooks/postdeploy/060_register_instance_to_shared_application_load_balancer.sh, api2/ebs/register_instance_to_shared_application_load_balancer.php |
6
+ | [Deploy-Time Auto-Registration to the Shared ALB Target Group (non-production)](features/alb-target-group-auto-registration.md) | TOGA does **not** pay for EB-managed load-balancer registration, so an EB instance is normally **not** added to its environment's ALB target group — a fresh or | worker2/.platform/hooks/postdeploy/060_register_instance_to_shared_application_load_balancer.sh, worker2/ebs/register_instance_to_shared_application_load_balancer.php, worker2/.platform/hooks/_shared/040-write-instance-id.sh, worker2/.platform/hooks/_shared/041-write-region.sh, worker2/.platform/hooks/_shared/042-write-eb-environment.sh, worker2/.platform/hooks/postdeploy/015_install_composer.sh, saml/.platform/hooks/_shared/040-write-instance-id.sh, saml/.platform/hooks/_shared/041-write-region.sh, saml/.platform/hooks/_shared/042-write-eb-environment.sh, saml/.platform/hooks/prestart/040-write-instance-id.sh, saml/.platform/hooks/prestart/041-write-region.sh, saml/.platform/hooks/postdeploy/040-write-instance-id.sh, saml/.platform/hooks/postdeploy/041-write-region.sh, saml/.platform/hooks/postdeploy/042-write-eb-environment.sh, saml/.platform/hooks/postdeploy/060_register_instance_to_shared_application_load_balancer.sh, saml/ebs/register_instance_to_shared_application_load_balancer.php, saml/composer.json, api2/.platform/hooks/prebuild/_shared/042-write-eb-environment.sh, api2/.platform/hooks/postdeploy/060_register_instance_to_shared_application_load_balancer.sh, api2/ebs/register_instance_to_shared_application_load_balancer.php |
7
7
  | [All-Client Email Queue Monitor (Monitor/Operations/EmailQueue)](features/all-client-email-queue-monitor.md) | `_Worker_Monitor_Operations::EmailQueue()` is the cross-client health check for the 2.0 outbound email queue (`Logs_<Client>.Email` — see [2.0 Email Send Pipeli | worker2/Worker/Monitor/Operations.php, _underscore/Database.php, _underscore/Query.php, dbchanges2/Logs_Client/2026-05-21 - Email.sql |
8
+ | [API 2.0 Stability Monitor (Monitor/Operations/ApiStability)](features/api-stability-monitor.md) | `_Worker_Monitor_Operations::ApiStability()` is a cross-client/operational **end-to-end round-trip probe** of the production 2.0 API. | worker2/Worker/Monitor/Operations.php, dbchanges2/Core/2026-08-28a - API 2.0 Stability Monitor Cron.sql |
8
9
  | [Automated PR Merger — Concurrent Force-Push Clobber Race](features/automated-pr-merger-force-push-race.md) | The automated PR merger `_Worker_Team_GitHub::Merge` (`worker2` `Worker/Team/Github.php`) merges approved PRs to `_production` by **force-pushing from a clone t | Worker/Team/Github.php |
9
10
  | [Callback Scheduling ("call me back") — worker2 AI-BDR](features/callback-scheduling.md) | The **AI-BDR callback path**: what happens between a prospect saying *"call me back later"* on a Vapi call and the dialer actually placing that second call. | worker2/Worker/Vapi.php, worker2/Worker/Ai/Bdr/Vapi.php, worker2/Controller/Index.php |
10
11
  | [ClickUp Connectivity Watchdog](features/clickup-connectivity-watchdog.md) | A cron watchdog that emails when the ClickUp integration looks disconnected during business hours. | worker2/Worker/Clickup/Health.php, worker2/Database/ClickupHealthWatchdog.sql |
@@ -6,20 +6,32 @@ project: Worker
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-08-04
9
+ updated: 2026-08-27
10
10
  owners: [jcardinal]
11
11
  files:
12
12
  - worker2/.platform/hooks/postdeploy/060_register_instance_to_shared_application_load_balancer.sh
13
13
  - worker2/ebs/register_instance_to_shared_application_load_balancer.php
14
- - worker2/.platform/hooks/prebuild/_shared/040-write-instance-id.sh
15
- - worker2/.platform/hooks/prebuild/_shared/041-write-region.sh
16
- - worker2/.platform/hooks/prebuild/_shared/042-write-eb-environment.sh
17
- - api2/.platform/hooks/prebuild/_shared/042-write-eb-environment.sh
14
+ - worker2/.platform/hooks/_shared/040-write-instance-id.sh
15
+ - worker2/.platform/hooks/_shared/041-write-region.sh
16
+ - worker2/.platform/hooks/_shared/042-write-eb-environment.sh
18
17
  - worker2/.platform/hooks/postdeploy/015_install_composer.sh
18
+ - saml/.platform/hooks/_shared/040-write-instance-id.sh
19
+ - saml/.platform/hooks/_shared/041-write-region.sh
20
+ - saml/.platform/hooks/_shared/042-write-eb-environment.sh
21
+ - saml/.platform/hooks/prestart/040-write-instance-id.sh
22
+ - saml/.platform/hooks/prestart/041-write-region.sh
23
+ - saml/.platform/hooks/postdeploy/040-write-instance-id.sh
24
+ - saml/.platform/hooks/postdeploy/041-write-region.sh
25
+ - saml/.platform/hooks/postdeploy/042-write-eb-environment.sh
26
+ - saml/.platform/hooks/postdeploy/060_register_instance_to_shared_application_load_balancer.sh
27
+ - saml/ebs/register_instance_to_shared_application_load_balancer.php
28
+ - saml/composer.json
29
+ - api2/.platform/hooks/prebuild/_shared/042-write-eb-environment.sh
19
30
  - api2/.platform/hooks/postdeploy/060_register_instance_to_shared_application_load_balancer.sh
20
31
  - api2/ebs/register_instance_to_shared_application_load_balancer.php
21
32
  related:
22
33
  - ../architecture.md
34
+ - ../../saml/architecture.md
23
35
  - ../../api2/workflows/codepipeline-codeconnections-deploy.md
24
36
  - ../../api2/architecture.md
25
37
  ---
@@ -34,8 +46,10 @@ pair automates that step for **non-production** environments: on every deploy th
34
46
  registers **itself** into the target group whose **name exactly equals the EB environment name**.
35
47
 
36
48
  Originally built in `api2`; ported to `worker2` on 2026-07-28 with three security hardenings
37
- (instance-profile credentials, native-PHP IMDSv2, fail-closed token handling). **`worker2` is now
38
- the reference implementation — `api2`'s copy has known defects, see below.**
49
+ (instance-profile credentials, native-PHP IMDSv2, fail-closed token handling), and to `saml`
50
+ (SAML SSO Gateway) on 2026-08-27 as a byte-faithful copy of the hardened `worker2` variant.
51
+ **`worker2` is now the reference implementation — port from it, not from `api2`, whose copy has
52
+ known defects (see below).**
39
53
 
40
54
  Non-prod `worker2` environments need this because they are reached **over HTTP for manual job
41
55
  invocation**, not through SQS.
@@ -50,8 +64,8 @@ invocation**, not through SQS.
50
64
  2. **`ebs/register_instance_to_shared_application_load_balancer.php`**
51
65
  - Reads the instance id and region from
52
66
  `/var/app/current/storage/instance-id.txt` and `storage/region.txt` (written earlier by the
53
- `_shared/040-write-instance-id.sh` / `041-write-region.sh` prebuild hooks), falling back to a
54
- direct **IMDSv2** lookup.
67
+ `_shared/040-write-instance-id.sh` / `041-write-region.sh` hooks — see the layout note
68
+ below), falling back to a direct **IMDSv2** lookup.
55
69
  - Reads the EB environment name via `get-config container -k environment_name`.
56
70
  - Uses the AWS SDK (`aws/aws-sdk-php`) ELBv2 client to `DescribeTargetGroups` and selects the
57
71
  target group whose **`TargetGroupName` is an exact string match** for the environment name —
@@ -64,13 +78,23 @@ invocation**, not through SQS.
64
78
 
65
79
  | Hook | Provides |
66
80
  |---|---|
67
- | `prebuild/_shared/040-write-instance-id.sh` | `storage/instance-id.txt` |
68
- | `prebuild/_shared/041-write-region.sh` | `storage/region.txt` |
69
- | `prebuild/_shared/042-write-eb-environment.sh` | `storage/eb-environment.txt` |
81
+ | `_shared/040-write-instance-id.sh` (run via a phase wrapper) | `storage/instance-id.txt` |
82
+ | `_shared/041-write-region.sh` (run via a phase wrapper) | `storage/region.txt` |
83
+ | `_shared/042-write-eb-environment.sh` (run via a phase wrapper) | `storage/eb-environment.txt` |
70
84
  | `postdeploy/015_install_composer.sh` | `vendor/` (the AWS SDK) |
71
85
 
72
86
  Renumber it below `015` and the SDK autoloader does not exist yet.
73
87
 
88
+ ### Real on-disk hook layout — `_shared` + thin phase wrappers (corrected 2026-08-27)
89
+
90
+ The canonical `040`/`041`/`042` scripts live at **`.platform/hooks/_shared/0xx.sh`** — **not**
91
+ under `prebuild/`, as earlier revisions of this doc stated. Each deploy phase that needs one
92
+ carries a **thin wrapper** in its own hook dir (`.platform/hooks/prestart/0xx.sh`,
93
+ `.platform/hooks/postdeploy/0xx.sh`) that just runs `bash ../_shared/0xx.sh`. In the `worker2`
94
+ and `saml` trees, `040`/`041` have both a `prestart` and a `postdeploy` wrapper, while `042` has
95
+ a `postdeploy` wrapper only (no `prestart` copy). When porting, copy the canonical `_shared`
96
+ script **and** the phase wrappers the tier needs — the wrappers are what actually execute.
97
+
74
98
  ### `042-write-eb-environment.sh` — the EB environment name, cached to disk (2026-08-04)
75
99
 
76
100
  Added to the same `_shared` hook family (plus a postdeploy hook) in **both `api2` and `worker2`**,
@@ -95,6 +119,21 @@ an identical `015_install_composer.sh`; and the same `get-config environment -k
95
119
  convention already used by `.ebextensions/git.php`. Check these four before porting the pair to
96
120
  any other 2.0 tier.
97
121
 
122
+ **`saml` port (2026-08-27) — the two prerequisites a bare tier is missing.** `saml` had *none*
123
+ of the hook family (prebuild was `git.sh` only, postdeploy was `015_install_composer.sh` only,
124
+ no `ebs/` dir), so the port copied the canonical `_shared/040/041/042` scripts, their phase
125
+ wrappers, `060_…​.sh`, and the hardened `ebs/register_instance_to_shared_application_load_balancer.php`.
126
+ The two prerequisites it did **not** already meet:
127
+
128
+ - **`aws/aws-sdk-php` was absent** from `composer.json` — added `aws/aws-sdk-php: ^3.337`
129
+ (the SDK's ELBv2 client is what the script uses). **`composer.lock` + `vendor/` must be
130
+ regenerated** with `composer require aws/aws-sdk-php:^3.337` before deploy (a developer step;
131
+ not done in the porting session).
132
+ - **Stale `config.platform.php` pin.** `saml` pinned `7.2.33`, but the EB box runs PHP 8.5 (AL2023
133
+ / platform 4.13.1) and `aws-sdk-php ^3.337` needs PHP ≥ 8.1, so the stale pin would block
134
+ dependency resolution. Corrected to `8.1.0`. **Check the platform pin on any tier before
135
+ porting** — a low pin silently blocks the SDK install.
136
+
98
137
  ## Credentials — instance profile only (decision, 2026-07-28)
99
138
 
100
139
  **Deploy-time AWS credentials come from the EC2 instance profile. Never from source, never from
@@ -141,6 +180,12 @@ implementation instead:
141
180
  group name first when a non-prod env is unreachable after a deploy.
142
181
  - **Exit 0 hides failures.** Registration problems will not show in deploy status — read
143
182
  `/var/log/eb-hooks.log` on the instance.
183
+ - **`saml` does nothing until it has a non-prod env + matching target group.** The team KB lists
184
+ `saml`'s only environment as `saml-production`, which this hook **self-skips**. The `saml` port
185
+ registers a target only on a **non-prod** `saml` env, and only if (a) that env's EC2 instance
186
+ profile carries the two IAM actions below (Describe scoped by `aws:RequestedRegion`, Register
187
+ scoped to the **non-prod** target-group ARN — never `targetgroup/*/*`), and (b) a target group
188
+ exists named **exactly** the non-prod env name. Naming equality is the whole contract.
144
189
  - **The hook turns on reachability; it does not secure it.** The target group and listener already
145
190
  exist by convention, but because this hook is what actually puts non-prod instances behind the
146
191
  shared ALB, the **listener rules and security groups must restrict non-prod to internal/VPN
@@ -172,6 +217,13 @@ which lists the other hooks but not `060` — pending an elevated-doc update.
172
217
 
173
218
  ## Change history
174
219
 
220
+ - 2026-08-27 — Ported the hardened `worker2` hook pair to **`saml`** (SAML SSO Gateway): copied
221
+ `_shared/040/041/042`, their prestart/postdeploy wrappers, `060_…​.sh`, and the instance-profile
222
+ `ebs/register_instance_to_shared_application_load_balancer.php`. Added `aws/aws-sdk-php: ^3.337`
223
+ to `saml/composer.json` and corrected its stale `config.platform.php` pin `7.2.33 → 8.1.0`
224
+ (lock/vendor still to be regenerated by a developer). Also **corrected this doc's file paths**:
225
+ the canonical `040/041/042` scripts live at `.platform/hooks/_shared/`, not `prebuild/_shared/`,
226
+ and run via thin per-phase wrappers. (jcardinal)
175
227
  - 2026-08-04 — Added `_shared/042-write-eb-environment.sh` (+ a postdeploy hook) to **api2 and
176
228
  worker2**, caching the EB environment name to disk from `get-config` → `$EB_ENVIRONMENT_NAME` →
177
229
  the IMDS tag. Read from a file rather than shelled out to at use time because its first consumer
@@ -0,0 +1,140 @@
1
+ ---
2
+ title: API 2.0 Stability Monitor (Monitor/Operations/ApiStability)
3
+ framework: "2.0"
4
+ repo: worker2
5
+ project: Worker
6
+ client: shared
7
+ type: feature
8
+ status: active
9
+ updated: 2026-08-28
10
+ owners: ["jcardinal"]
11
+ files:
12
+ - worker2/Worker/Monitor/Operations.php
13
+ - dbchanges2/Core/2026-08-28a - API 2.0 Stability Monitor Cron.sql
14
+ related:
15
+ - ./oneuptime-worker2-monitoring.md
16
+ - ./netsuite-integrations-monitor.md
17
+ - ./all-client-email-queue-monitor.md
18
+ - ./elastic-beanstalk-health-monitor.md
19
+ ---
20
+
21
+ ## Summary
22
+
23
+ `_Worker_Monitor_Operations::ApiStability()` is a cross-client/operational **end-to-end
24
+ round-trip probe** of the production 2.0 API. Unlike a log-scraper (blind to a dead
25
+ endpoint) or a cheap health-endpoint liveness probe, it performs a real authenticated
26
+ call chain against `api.togahub.com/v2` and only reports OK when the API round-trips a
27
+ representative read correctly. It follows the established "dumb reporter, smart monitor"
28
+ OneUptime push pattern — same shape as the sibling
29
+ [NetsuiteIntegrations](./netsuite-integrations-monitor.md) and
30
+ [EmailQueue](./all-client-email-queue-monitor.md) monitors on
31
+ `_Worker_Monitor_Operations`. Shared internal operational infrastructure, **not** a client
32
+ feature.
33
+
34
+ **Motivating incident.** A `Core.Records` model failed to deploy to production. Because the
35
+ 2.0 API requires a model for every `Core.Records` entry, any call touching the missing
36
+ model returned HTTP 500 — while the Elastic Beanstalk instances stayed "healthy" and the
37
+ `/v2/health` short-circuit route kept returning 200. A monitor that actually round-trips
38
+ the API would have gone red immediately; a health-endpoint probe and an EB health check
39
+ both stayed green. This probe closes that gap.
40
+
41
+ ## Key files / entry points
42
+
43
+ - `worker2/Worker/Monitor/Operations.php` — `abstract class _Worker_Monitor_Operations`;
44
+ `public static function ApiStability(): string`, route
45
+ `Monitor/Operations/ApiStability`. All config is `ALL_CAPS` in-method locals (no
46
+ arguments): `$ONEUPTIME_URL` (push credential), `$API_BASE`, timeouts, expected statuses,
47
+ `$EXPECTED_CLIENT_AUTH_TYPE = 'SSO'`, `$DOMAINS_QUERY`, and `$ORIGIN`.
48
+ - `dbchanges2/Core/2026-08-28a - API 2.0 Stability Monitor Cron.sql` — the `Core.CronJobs`
49
+ seed row (see below).
50
+
51
+ ## How it works
52
+
53
+ Two sequential sub-checks against **production** `api.togahub.com/v2` (the `$API_BASE` is
54
+ hardcoded — the monitor watches prod regardless of which environment runs it). HTTP goes
55
+ through `_ApiRequest` with `ENCODE__JSON` (so `execute()` returns the decoded object) and a
56
+ fresh `transactionId` per request via `_String::generateUuid()`:
57
+
58
+ 1. **Auth** — `POST /v2/auth/public?transactionId=<uuid>` with an `Origin` header (see the
59
+ gotcha below). Expect **HTTP 201**; extract the bearer token from `data.tokens.access`.
60
+ 2. **Representative read** — `GET /v2/domains?<fields/join/where>&transactionId=<uuid>` with
61
+ `Authorization: Bearer <token>` and the same `Origin`. The query is a representative
62
+ joined read across `Domains` + `Environments` + `ClientAuthentications` + `Apps`, scoped
63
+ to the `compass.togasupply.com` PRODUCTION domain. Expect **HTTP 200 AND**
64
+ `data.domains[0].ClientAuthentications.type === 'SSO'` (`$EXPECTED_CLIENT_AUTH_TYPE`).
65
+ This is the exact call that failed during the missing-model incident.
66
+
67
+ Any failed sub-check sets `alarm = 'DEGRADED'` (else `OK`). The method pushes
68
+ `{status:'reporting', alarm, authStatus, domainsStatus, failureCount, failures,
69
+ checkedAtUtc}` to OneUptime, then returns a structured summary string recorded in
70
+ `WorkerJobs.output`.
71
+
72
+ Compass USA's production supply domain is used only as a **representative** live call — this
73
+ is shared platform/liveness knowledge, not client-specific behavior.
74
+
75
+ ### Never-throw / non-fatal push
76
+
77
+ Logging is off, `throwExceptionsOnFailure` is off, and the push is wrapped so a
78
+ monitoring-side failure never fails the worker job — the standard reporter contract from
79
+ [the OneUptime push-metric pattern](./oneuptime-worker2-monitoring.md).
80
+
81
+ ## Cron seed (dbchanges2)
82
+
83
+ `dbchanges2/Core/2026-08-28a - API 2.0 Stability Monitor Cron.sql` — one `Core.CronJobs`
84
+ row: action `Monitor/Operations/ApiStability`, schedule `*/5 * * * *` (every 5 min, all
85
+ hours), `maxExecutionTime` 120, `parameters` NULL (no arguments), **`isActive = 0`**
86
+ (ships inert per the deploy-inert convention — flip to 1 after the OneUptime monitor exists
87
+ and its push URL is set). Core-database-only references (cluster-isolation satisfied); no
88
+ `Core.Records`/`RecordFields` involved.
89
+
90
+ ## Gotchas / known issues
91
+
92
+ - **`POST /v2/auth/public` returns HTTP 401 without an `Origin` request header.** The
93
+ public-auth endpoint carries **no** client credential; the API resolves which client/app
94
+ the public token is minted for from the request's `Origin` header. A server-side call must
95
+ set it explicitly (`Origin: https://compass.togasupply.com`) — the browser supplies Origin
96
+ automatically, which is why the front end's public-token flow
97
+ (`toga-blox-npm/src/api/auth.ts` POSTs to `/auth/public` with a null body + JSON
98
+ content-type) never sees this. Send the same `Origin` on the `/domains` call too, mirroring
99
+ the browser. This is the non-obvious fact for anyone scripting a server-side call to the
100
+ public-auth endpoint. (Verified 2026-08-28: with Origin, auth → 201 and domains → 200 +
101
+ `SSO`; without it, auth → 401.)
102
+ - **The push URL is a credential** — it lives only as the in-method `$ONEUPTIME_URL` local,
103
+ never logged and never written into a doc or into dbchanges2.
104
+ - **Watches PRODUCTION regardless of runner.** `$API_BASE` is hardcoded to prod, so a
105
+ non-prod worker still probes prod — intended, but do not "fix" it to the running env.
106
+
107
+ ## OneUptime import artifact
108
+
109
+ Provisioned by importing an Incoming-Request monitor export (a local staging artifact,
110
+ `C:\STAGE\monitor-export-api-2.0-stability-2026-08-28.json` — **not** a repo file) with three
111
+ criteria: body Contains `"alarm":"DEGRADED"` → Degraded + incident (templated with
112
+ `{{requestBody.failures}}` etc.); body Contains `"alarm":"OK"` → Operational;
113
+ not-received-25-min → Offline + incident. Per the
114
+ [activation-ordering trap](./oneuptime-worker2-monitoring.md#gotchas--known-issues), the
115
+ absence criterion ships `isEnabled: false` — enable it after the first push lands.
116
+ Project-specific status/severity ObjectIDs were cloned from an existing monitor export.
117
+
118
+ ## Change history
119
+
120
+ - 2026-08-28 — Built `ApiStability()` on `_Worker_Monitor_Operations` — a two-step
121
+ end-to-end round-trip probe (auth/public → 201 → bearer, then a joined `/v2/domains` read
122
+ → 200 + `SSO`) against production `api.togahub.com/v2`, plus the `Core.CronJobs` seed
123
+ (`*/5 * * * *`, `isActive=0`) and an Incoming-Request OneUptime import. Motivated by a
124
+ missing `Core.Records` model that 500'd real calls while EB/health-endpoint stayed green.
125
+ Recorded the durable gotcha that `POST /v2/auth/public` **401s without an `Origin`
126
+ header** (the API resolves the client from Origin). Verified working end-to-end after the
127
+ Origin fix. (jcardinal)
128
+
129
+ ## Related docs
130
+
131
+ - [OneUptime push-metric monitors for 2.0 workers](./oneuptime-worker2-monitoring.md) — the
132
+ push/token pattern, OneUptime criteria, import-JSON schema, deploy-inert and
133
+ activation-ordering conventions, and the liveness-probe-vs-log-scraper distinction.
134
+ - [NetSuite Integrations Monitor](./netsuite-integrations-monitor.md) and
135
+ [All-Client Email Queue Monitor](./all-client-email-queue-monitor.md) — sibling
136
+ `_Worker_Monitor_Operations` monitors modeled on the same reporter contract.
137
+ - [Elastic Beanstalk Health Monitor](./elastic-beanstalk-health-monitor.md) — EB-level
138
+ health this probe deliberately reaches past (EB stayed healthy during the incident).
139
+ </content>
140
+ </invoke>
@@ -6,7 +6,7 @@ project: Worker
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-08-20
9
+ updated: 2026-08-28
10
10
  owners: ["jcardinal", "mhammontree", "bala"]
11
11
  files:
12
12
  - worker2/Worker/Monitor/Compass.php
@@ -195,6 +195,21 @@ touches it. Rules for the probe:
195
195
  duplicate credentials into a second repo and race the shared access token. Accept and record the
196
196
  coverage gap rather than forking the client.
197
197
 
198
+ #### A round-trip probe is stronger than a health-endpoint probe
199
+
200
+ A health endpoint that short-circuits before any real work (`/v2/health` → 200) proves the
201
+ process is up, but **not** that the API can actually serve a request. A missing `Core.Records`
202
+ model once 500'd every real call while `/v2/health` and EB health both stayed green. Where an
203
+ integration matters, add an **end-to-end round-trip probe** that authenticates and reads a
204
+ representative record, asserting on the returned data — not just the HTTP status. First
205
+ instance: [API 2.0 Stability Monitor](./api-stability-monitor.md), which mints a public token
206
+ then reads `/v2/domains` and asserts `ClientAuthentications.type === 'SSO'`.
207
+
208
+ **Gotcha for a server-side public-auth call:** `POST /v2/auth/public` returns **HTTP 401
209
+ without an `Origin` request header** — the endpoint carries no client credential and resolves
210
+ the client/app from `Origin`. The browser supplies it automatically; a scripted call must set
211
+ it explicitly. See the API-stability doc for details.
212
+
198
213
  ### Tunable thresholds live in `Core.CronJobs.parameters`, everything else stays a local
199
214
 
200
215
  The alarm threshold is the **one** setting that goes in the DB, as
@@ -356,6 +371,12 @@ check for another client.
356
371
  monitors must pass the region (see [Cloud S3 helpers](../../_underscore/features/cloud-s3-helpers.md)).
357
372
 
358
373
  ## Change history
374
+ - 2026-08-28 — Added the **round-trip probe** refinement (an authenticated end-to-end read
375
+ that asserts on returned data is stronger than a `/v2/health` short-circuit or EB health
376
+ check — a missing `Core.Records` model 500'd real calls while both stayed green) and the
377
+ durable gotcha that `POST /v2/auth/public` **401s without an `Origin` header** (the API
378
+ resolves the client from Origin). First instance:
379
+ [API 2.0 Stability Monitor](./api-stability-monitor.md). (jcardinal)
359
380
  - 2026-08-24 — Documented the **monitor import-JSON schema** the runbook previously described
360
381
  only conceptually: the `oneuptime-resource-export` envelope, `monitorType` must be the
361
382
  spaced display name `"Incoming Request"`, the `_type`-tagged
@@ -410,6 +431,7 @@ check for another client.
410
431
  5-min cadence → 10/15-min heartbeat timing. (jcardinal)
411
432
 
412
433
  ## Related docs
434
+ - [API 2.0 Stability Monitor](./api-stability-monitor.md) — the first end-to-end round-trip probe instance (auth/public + /domains SSO assertion)
413
435
  - [Monitoring Framework](./monitoring-framework.md) — the parallel DB-driven, email-alert monitoring pattern
414
436
  - [Cloud S3 helpers](../../_underscore/features/cloud-s3-helpers.md) — `_Cloud::getS3Objects()` used to list the EDI bucket
415
437
  - [OneUptime 1.0 worker uptime monitoring](../../../1.0/apps/worker/features/oneuptime-worker-uptime-monitoring.md) — the 1.0 push-heartbeat predecessor
@@ -19,7 +19,7 @@ _Auto-generated by `knowledge.js index`. Do not hand-edit._
19
19
  ## 2.0 framework
20
20
 
21
21
  - **_underscore** (_Underscore) _(framework core)_ — 70 doc(s) → [2.0/apps/_underscore/INDEX.md](2.0/apps/_underscore/INDEX.md)
22
- - **worker2** (Worker) — 57 doc(s) → [2.0/apps/worker2/INDEX.md](2.0/apps/worker2/INDEX.md)
22
+ - **worker2** (Worker) — 58 doc(s) → [2.0/apps/worker2/INDEX.md](2.0/apps/worker2/INDEX.md)
23
23
  - **api2** (API) — 25 doc(s) → [2.0/apps/api2/INDEX.md](2.0/apps/api2/INDEX.md)
24
24
  - **dbchanges2** (Database Changes) _(framework core)_ — 13 doc(s) → [2.0/apps/dbchanges2/INDEX.md](2.0/apps/dbchanges2/INDEX.md)
25
25
  - **toga2-supply** (TOGa Supply) — 9 doc(s) → [2.0/apps/toga2-supply/INDEX.md](2.0/apps/toga2-supply/INDEX.md)
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "toga-ai",
3
- "version": "1.0.672",
3
+ "version": "1.0.674",
4
4
  "description": "TOGA Technology Team Claude Knowledge System — shared AI coding harness with skills, knowledge base CLI, and project installer for Claude Code.",
5
5
  "keywords": [
6
6
  "claude",