toga-ai 1.0.171 → 1.0.173

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -3,6 +3,7 @@
3
3
  | Doc | Summary | Files |
4
4
  |-----|---------|-------|
5
5
  | [Worker (worker2) Architecture](architecture.md) | Worker (repo `worker2`) is an AWS Elastic Beanstalk **Worker Tier** application that processes background jobs. | worker2/Controller/Index.php, worker2/Worker/, worker2/LambdaFunctions/, _underscore/Worker.php |
6
+ | [ClickUp Connectivity Watchdog](features/clickup-connectivity-watchdog.md) | A cron watchdog that emails when the ClickUp integration looks disconnected during business hours. | worker2/Worker/Clickup/Health.php, worker2/Database/ClickupHealthWatchdog.sql |
6
7
  | [ClickUp Project & Opportunity Multi-List Routing](features/clickup-project-routing.md) | Routes ClickUp tasks into the correct **secondary multi-list memberships** based on their custom-field values, via the `clickup` webhook. | worker2/Worker/Clickup/Project.php, worker2/Worker/Clickup.php |
7
8
  | [ClickUp Work Type Automation (Committed / Conditional / Stretch)](features/clickup-work-type-automation.md) | The ClickUp webhook handler (`_Worker_Clickup`) automatically maintains each task's **Work Type** custom field — `Committed`, `Conditional`, or `Stretch` — base | worker2/Worker/Clickup.php, worker2/Tests/Worker/ClickupWorkTypeTest.php |
8
9
  | [Creating Worker Actions](features/creating-worker-actions.md) | How to add a new callable Worker action — a PHP class whose `public static` methods are invoked as background jobs (via webhook, cron, or `_Worker::runTask()`). | worker2/Worker/, worker2/Controller/Index.php, _underscore/Worker.php |
@@ -0,0 +1,81 @@
1
+ ---
2
+ title: ClickUp Connectivity Watchdog
3
+ framework: "2.0"
4
+ repo: worker2
5
+ project: Worker
6
+ client: shared
7
+ type: feature
8
+ status: active
9
+ updated: 2026-06-23
10
+ owners: [ajean]
11
+ files:
12
+ - worker2/Worker/Clickup/Health.php
13
+ - worker2/Database/ClickupHealthWatchdog.sql
14
+ related:
15
+ - ./clickup-project-routing.md
16
+ - ../architecture.md
17
+ ---
18
+
19
+ ## Summary
20
+
21
+ A cron watchdog that emails when the ClickUp integration looks disconnected during business
22
+ hours. Implemented in `_Worker_Clickup_Health` (action `Clickup/Health/Check`). It exists
23
+ because the ClickUp webhook subscriptions can silently die (a dropped subscription, a ClickUp
24
+ outage, or our endpoint failing), and nobody notices until reporting/automations drift.
25
+
26
+ ## Key files / entry points
27
+
28
+ - `worker2/Worker/Clickup/Health.php` — `_Worker_Clickup_Health::Check(): string` (the cron
29
+ action) + `initialize()` (registers the `Team` DB, same pattern as `_Worker_Team_Transcripts`).
30
+ - `worker2/Database/ClickupHealthWatchdog.sql` — run-once reference: the `Team.ClickUpHealthState`
31
+ table + the `Core.CronJobs` row. Canonical migrations live in `dbchanges2`
32
+ (`Team/2026-06-23a …`, `Core/2026-06-23a …`).
33
+
34
+ ## How it works
35
+
36
+ Scheduled every 15 min on weekdays 09:00–16:45 Central (cron `0,15,30,45 9-16 * * 1-5`).
37
+ `Check()` evaluates two signals and emails `ajean@togatech.com` if either fires:
38
+
39
+ 1. **Inbound silence** — `MAX(dtCreated)` of `Core.WorkerJobs WHERE action = 'Clickup/Webhook'`
40
+ (exact match, so the watchdog's own `Clickup/Health/Check` rows don't count). Elapsed minutes
41
+ are computed by **MySQL `TIMESTAMPDIFF`** (no PHP/DB timezone mismatch). Alerts when ≥ 60 min
42
+ silent, but only from the **second business hour onward** (a one-hour warm-up so the expected
43
+ overnight gap isn't flagged at 09:00).
44
+ 2. **Webhook health** — `GET /team/{id}/webhook`; any subscription whose endpoint contains
45
+ `togahub.com` with `health->status == 'failing'`. The probe is wrapped in try/catch — if the
46
+ probe call itself throws, that is treated as a connectivity signal too.
47
+
48
+ De-duplication: `Team.ClickUpHealthState` (single row) records `dtLastAlerted`; an ongoing
49
+ outage re-emails at most once per 60 min (`shouldAlert`), and the flag is cleared when ClickUp
50
+ is healthy again (`clearAlert`) so each fresh outage notifies once.
51
+
52
+ ## Data model
53
+
54
+ - **`Team.ClickUpHealthState`** — single-row alert state: `id`, `uuid`, `dtCreated`, `dtUpdated`,
55
+ `dtLastAlerted` (NULL when healthy), `lastReason`. The row is created lazily on first alert.
56
+ - **`Core.CronJobs`** — one row, action `Clickup/Health/Check`, `maxExecutionTime` 120.
57
+
58
+ ## Client variations
59
+
60
+ None — internal monitoring for Agilant's ClickUp workspace.
61
+
62
+ ## Gotchas / known issues
63
+
64
+ - **Detects ClickUp-side problems only.** It runs *inside* the worker/cron pipeline, so it
65
+ assumes that pipeline is healthy. It will **not** fire if the EB worker tier, the CronScheduler
66
+ Lambda, or the database is itself down — full worker-down detection needs an external uptime
67
+ monitor (not built).
68
+ - **The inbound query must match `Clickup/Webhook` exactly.** A `LIKE 'Clickup/%'` would also
69
+ match the watchdog's own `Clickup/Health/Check` rows and the timer would never expire.
70
+ - **Email identifier is `'True'`** (the internal/team `clientIdentifier`, consistent with other
71
+ team alerts via `_Worker_Notification_Email::Send`).
72
+ - **Business window is enforced in code too** (Mon–Fri, 09:00–17:00 Central), not just in the
73
+ cron schedule, so a manual run outside hours is a safe no-op.
74
+
75
+ ## Change history
76
+ - 2026-06-23 — Created. Watchdog cron emails on 60-min inbound webhook silence or a failing subscription during business hours; deduped via `Team.ClickUpHealthState`. (ajean)
77
+
78
+ ## Related docs
79
+
80
+ - [ClickUp Project & Opportunity Multi-List Routing](./clickup-project-routing.md) — the integration this watches.
81
+ - [Worker (worker2) Architecture](../architecture.md) — CronJobs, WorkerJobs, always-HTTP-200.
@@ -6,21 +6,23 @@ project: Worker
6
6
  client: shared
7
7
  type: feature
8
8
  status: active
9
- updated: 2026-06-09
10
- owners: [jcardinal]
9
+ updated: 2026-06-23
10
+ owners: [jcardinal, ajean]
11
11
  files:
12
12
  - worker2/Worker/Clickup/Project.php
13
13
  - worker2/Worker/Clickup.php
14
14
  related:
15
15
  - ../architecture.md
16
16
  - ./creating-worker-actions.md
17
+ - ./clickup-connectivity-watchdog.md
17
18
  ---
18
19
 
19
20
  ## Summary
20
21
 
21
22
  Routes ClickUp tasks into the correct **secondary multi-list memberships** based on their
22
23
  custom-field values, via the `clickup` webhook. Implemented in `_Worker_Clickup_Project`
23
- (an abstract worker class) with three independent entry points covering three ClickUp spaces.
24
+ (an abstract worker class) with three independent routing entry points covering three ClickUp
25
+ spaces, plus a **hub-placement guard** that keeps Epics/Opportunities anchored to their hub.
24
26
 
25
27
  **Hybrid design (the central constraint):** ClickUp's public **v2 API forbids changing or
26
28
  removing a task's home list** (`400 TASK_035 — Task home list cannot be altered`). So:
@@ -35,8 +37,9 @@ removing a task's home list** (`400 TASK_035 — Task home list cannot be altere
35
37
  - `worker2/Worker/Clickup.php` — delegates from `_Worker_Clickup::Webhook()`:
36
38
  - `taskStatusUpdated` (~L643): calls `handleEpicUpdate($taskId)` unconditionally.
37
39
  - `taskUpdated` (~L826–828): behind a **self-trigger guard**, calls all three handlers.
40
+ - `taskCreated` and `taskMoved`: call `enforceHubPlacement($taskId)` (the guard, below).
38
41
 
39
- Three public entry points (all `public static function …(string $taskId): void`):
42
+ Three routing entry points (all `public static function …(string $taskId): void`):
40
43
 
41
44
  | Handler | Space (guard) | Driver fields | Effect |
42
45
  |---|---|---|---|
@@ -44,13 +47,39 @@ Three public entry points (all `public static function …(string $taskId): void
44
47
  | `handleOpportunityUpdate` | Opportunity/RFP Hub `90113928591` | Business Unit, Project Category, Opportunity Stage | Adds to the LS folder's **Opportunities** list when BU = "Lifecycle Services", category maps to an LS folder, and stage is not `Closed Lost` / not Won/Signed; removes otherwise. |
45
48
  | `handleAmOpportunityUpdate` | Opportunity/RFP Hub `90113928591` | Business Unit, Project Category, Opportunity Stage, Project Phase | Two passes: (1) space-level A&M Opportunities list; (2) A&M folder list driven by Opportunity Stage, with Execution→Completion driven by Project Phase = "Complete". |
46
49
 
50
+ ### Hub-placement guard — `enforceHubPlacement(string $taskId): void` (added 2026-06-23)
51
+
52
+ Keeps **Epics in Project Hub** and **Opportunities in Opportunity/RFP Hub**. Called from the
53
+ `taskCreated` and `taskMoved` cases in `Worker/Clickup.php`. Unlike the three routing handlers
54
+ (which identify tasks structurally), it identifies the type by **custom Task Type**
55
+ (`custom_item_id`, configured via the `TASK_TYPE_EPIC_ID` / `TASK_TYPE_OPPORTUNITY_ID`
56
+ constants). A task is **misplaced** when its **home** `space->id` is not its hub — i.e. it was
57
+ created in, or moved into, Lifecycle Services / Advisory & Modernization (or any other space).
58
+
59
+ For a misplaced top-level task it:
60
+ 1. **Backfills metadata** (best-effort, LS/A&M only) — Business Unit from the home space
61
+ (`getBusinessUnitForSpaceId()`), Project Category from the home folder
62
+ (`getProjectCategoryForFolderId()`, the **inverse** of `getLsFolderKey`/`getAmFolderKey` —
63
+ keep them in sync). Written via `setDropdownFieldByName()`, which handles both `drop_down`
64
+ and `labels` field shapes and writes only when the value differs.
65
+ 2. **Re-adds the correct hub intake list** as a secondary membership (Epic →
66
+ `PROJECT_HUB_INTAKE_LIST_ID`, Opportunity → `OPPORTUNITY_HUB_INTAKE_LIST_ID`). The home list
67
+ itself cannot be moved by the API (TASK_035) — a human must do that.
68
+ 3. **Posts a one-time comment** (only when it actually just added the membership) telling the
69
+ user to move the home list back manually.
70
+
71
+ **Ships inert**: with `TASK_TYPE_*_ID` both `null` it is a no-op, so it can deploy before the
72
+ ids are discovered. Configure via the `DiscoverTaskTypes()` / `DiscoverIds()` actions, then
73
+ `RegisterHubGuardWebhooks()`.
74
+
47
75
  ## How it works
48
76
 
49
77
  1. The `clickup` webhook → `_Worker_Clickup::Webhook()` → delegation in `Worker/Clickup.php`.
50
78
  2. Each handler fetches task details via `_Worker_Clickup::getTaskDetails()` (a per-invocation
51
79
  `static` cache shared across all handlers in one job), then guards on **space id** and
52
- **top-level** (`parent` empty). Epics/tasks are identified structurally — ClickUp Task
53
- Types is not enabled, so `task_type` is unreliable.
80
+ **top-level** (`parent` empty). The three routing handlers identify Epics/tasks
81
+ **structurally** (space + top-level). Custom **Task Types are now configured**, and the
82
+ newer `enforceHubPlacement` guard keys on `custom_item_id` to recognize Epics/Opportunities.
54
83
  3. Custom fields are read by name via `extractCustomFields()` (one pass, last-non-null-wins,
55
84
  trimmed) → mapped to a target list id → reconciled by `applyMultiListPlacement()`:
56
85
  POST the target if absent, DELETE every other in-scope list (DELETE failures are caught
@@ -80,8 +109,9 @@ None in MySQL — all state lives in ClickUp. Each invocation is recorded as a n
80
109
 
81
110
  ClickUp structure encoded as constants in `Project.php`: 4 LS folders
82
111
  (Onsite, Factory, Audio-Visual, Managed Services), 4 A&M folders (Retainer, Security,
83
- AI & Data Center, Cloud), plus space-level lists. Regenerate with the `DiscoverIds()` /
84
- `DiscoverAmIds()` diagnostic actions if the ClickUp structure changes.
112
+ AI & Data Center, Cloud), plus space-level lists, the hub intake lists, and the Epic/Opportunity
113
+ custom Task Type ids. Regenerate with the `DiscoverIds()` / `DiscoverAmIds()` /
114
+ `DiscoverTaskTypes()` diagnostic actions if the ClickUp structure changes.
85
115
 
86
116
  ## Client variations
87
117
 
@@ -89,38 +119,52 @@ None — uniform across all clients (shared internal automation for Agilant's Cl
89
119
 
90
120
  ## Gotchas / known issues
91
121
 
122
+ - **Hub-guard correctness rests on the HOME space:** `enforceHubPlacement` filters on
123
+ `taskDetails->space->id`, which always follows the task's **home** list. A hub Epic that the
124
+ routing handlers mirror into LS/A&M as a *secondary* membership still reports its home hub as
125
+ the space, so the guard correctly ignores it. Filtering on home space is what stops the guard
126
+ from fighting the secondary-membership routing.
127
+ - **`getProjectCategoryForFolderId()` must stay in sync** with `getLsFolderKey`/`getAmFolderKey`
128
+ and the `*_FOLDER_*` constants — it is their inverse (folder id → Project Category value).
92
129
  - **List-name slash spacing differs by space, intentionally:** LS uses
93
130
  `"Client Scoping / Discovery"` (space before slash); A&M uses `"Client Scoping/ Discovery"`
94
131
  (no space). These match the real ClickUp list names — do not "fix" one to match the other.
95
132
  - **String-typed list IDs:** `collectCurrentMemberships()` returns IDs as **strings** and
96
133
  avoids array-key dedup, because PHP coerces numeric-string keys to int and would break the
97
- strict `in_array` comparisons against the string constants.
98
- - **Shared `getTaskDetails()` cache is stale after writes:** all three handlers share one
99
- cached task fetch per job. After one handler mutates memberships, the cached `locations`
100
- are stale for the others. Currently safe only because the space guards are mutually
101
- exclusive (a task lives in one space → exactly one handler does work). Wiring a second
102
- handler onto `taskStatusUpdated`, or overlapping list scopes, would break this.
103
- - **`RATE_LIMIT_DELAY_US` (120ms) is defined but unused** — pacing relies entirely on
104
- `_Component_Api_Clickup::send()`'s internal `sleep(1)`. `handleAmOpportunityUpdate` can
105
- issue several placement passes (multiple POST/DELETE) per event.
134
+ strict `in_array` comparisons against the string constants. (`getProjectCategoryForFolderId`
135
+ relies on that same numeric-string→int coercion being *consistent* between map keys and the
136
+ lookup, so it is correct; a hidden/leading-zero folder id falls through to `null`.)
137
+ - **Shared `getTaskDetails()` cache is stale after writes:** all handlers share one cached task
138
+ fetch per job. After one handler mutates memberships, the cached `locations` are stale for the
139
+ others. Safe today because the space guards are mutually exclusive (a task lives in one space →
140
+ exactly one handler does work).
141
+ - **`RATE_LIMIT_DELAY_US` (120ms):** now used by `enforceHubPlacement` / `setDropdownFieldByName`
142
+ (`usleep`); the three routing handlers still rely on `_Component_Api_Clickup::send()`'s internal
143
+ `sleep(1)`.
106
144
  - **POST failures are fatal, DELETE failures are tolerated** — a failed ADD leaves the task
107
145
  out of its correct list (surfaced as job failure); a failed DELETE (most likely TASK_035,
108
146
  which shouldn't occur on secondary lists) is logged with `[ClickupProject]` and skipped.
109
- - **Webhook subscriptions** must cover each space. Three exist, all → `webhook.togahub.com/clickup`:
147
+ - **Webhook subscriptions** must cover each space, all → `webhook.togahub.com/clickup`. Routing:
110
148
  legacy sprint space `90020178491`, Project Hub `90113939341`, Lifecycle Services
111
- `90114156087`. Opportunity Hub coverage is registered via `RegisterOpportunityWebhook()`.
149
+ `90114156087`; Opportunity Hub via `RegisterOpportunityWebhook()`. The hub guard additionally
150
+ needs **`taskCreated` on LS + A&M** and a **workspace-wide `taskMoved`**, registered via
151
+ `RegisterHubGuardWebhooks()` (which skips any already covered to avoid double-delivery).
112
152
 
113
153
  ## Diagnostics
114
154
 
115
- Read-only actions (no writes), invoked via `curl -X POST https://worker.togahub.com/` with
155
+ Read-only / setup actions, invoked via `curl -X POST https://worker.togahub.com/` with
116
156
  `{"action":"Clickup/Project/<Method>","parameters":{...}}`:
117
- `DiagnoseOpportunity`, `DiagnoseAmTask` (routing verdicts), `DiscoverIds`, `DiscoverAmIds`
118
- (print constant declarations), `RegisterOpportunityWebhook` (one-time setup).
157
+ `DiagnoseOpportunity`, `DiagnoseAmTask` (routing verdicts), `DiscoverIds`, `DiscoverAmIds`,
158
+ `DiscoverTaskTypes` (print constant declarations / custom-task-type ids),
159
+ `TestHubGuardMappings` (pure self-test of the folder→category & space→BU maps; expect
160
+ `RESULT: PASS`), `RegisterOpportunityWebhook` and `RegisterHubGuardWebhooks` (one-time setup).
119
161
 
120
162
  ## Change history
163
+ - 2026-06-23 — Added the hub-placement guard (`enforceHubPlacement` on `taskCreated`/`taskMoved`): re-anchors misplaced Epics/Opportunities to their hub by custom Task Type + home space, backfills Business Unit/Project Category, and notifies. Added `DiscoverTaskTypes`, `RegisterHubGuardWebhooks`, `TestHubGuardMappings`. Corrected the earlier "Task Types not enabled" note. (ajean)
121
164
  - 2026-06-09 — Documented ClickUp secondary multi-list routing (hybrid native-automation + worker design, three space handlers, self-trigger guard). (jcardinal)
122
165
 
123
166
  ## Related docs
124
167
 
125
168
  - [Worker (worker2) Architecture](../architecture.md) — always-HTTP-200, commit-before-SQS.
126
169
  - [Creating Worker Actions](./creating-worker-actions.md) — the worker-action contract.
170
+ - [ClickUp Connectivity Watchdog](./clickup-connectivity-watchdog.md) — the email-alert health check.
@@ -16,7 +16,7 @@ _Auto-generated by `knowledge.js index`. Do not hand-edit._
16
16
  ## 2.0 framework
17
17
 
18
18
  - **_underscore** (_Underscore) _(framework core)_ — 11 doc(s) → [2.0/apps/_underscore/INDEX.md](2.0/apps/_underscore/INDEX.md)
19
- - **worker2** (Worker) — 10 doc(s) → [2.0/apps/worker2/INDEX.md](2.0/apps/worker2/INDEX.md)
19
+ - **worker2** (Worker) — 11 doc(s) → [2.0/apps/worker2/INDEX.md](2.0/apps/worker2/INDEX.md)
20
20
  - **api2** (API) — 4 doc(s) → [2.0/apps/api2/INDEX.md](2.0/apps/api2/INDEX.md)
21
21
  - **dbchanges2** (Database Changes) _(framework core)_ — 1 doc(s) → [2.0/apps/dbchanges2/INDEX.md](2.0/apps/dbchanges2/INDEX.md)
22
22
  - **toga2-supply** (TOGa Supply) — 3 doc(s) → [2.0/apps/toga2-supply/INDEX.md](2.0/apps/toga2-supply/INDEX.md)
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "toga-ai",
3
- "version": "1.0.171",
3
+ "version": "1.0.173",
4
4
  "description": "TOGA Technology Team Claude Knowledge System — shared AI coding harness with skills, knowledge base CLI, and project installer for Claude Code.",
5
5
  "keywords": [
6
6
  "claude",
@@ -47,7 +47,12 @@ to change the threshold. If it prints "No team sessions saved yet.", say so and
47
47
  No team sessions saved yet. Save one at the end of your work: /session-save [name]
48
48
  ```
49
49
 
50
- **`list` mode:** render this table and stop here.
50
+ Whenever you render this table to the developer (in `list` mode **and** in the
51
+ no-argument discovery flow), explicitly disclose that delete is available — e.g. a
52
+ trailing line: "To retire a stale session, run `/session-resume delete <id>`." A
53
+ developer must never have to already know the command to discover it.
54
+
55
+ **`list` mode:** render this table (with the delete disclosure) and stop here.
51
56
 
52
57
  **`delete <id>` mode:** confirm with the developer, then run
53
58
  `node "<TEAM_REPO>/knowledge.js" session-delete <id> --msg="session: retire <id>"`,
@@ -61,7 +66,8 @@ report the `SESSION: DELETED|NOT_FOUND|CONFLICT` result, and stop.
61
66
  - `<slug-or-id>` → the row whose id equals it or contains the slug. If several match
62
67
  (same slug, different authors/dates), show those candidates (date + author) and ask which.
63
68
  - If none match: "No session found matching '<x>'. Available: [list ids]."
64
- - No argument → ask: "Which session should I load? (id or slug)"
69
+ - No argument → ask: "Which session should I load? (id or slug — or `delete <id>` to
70
+ retire a stale one)"
65
71
 
66
72
  ---
67
73