toga-ai 1.0.512 → 1.0.513
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
|
@@ -49,6 +49,26 @@ happening and were deliberately never ticketed. Nobody had that view before.
|
|
|
49
49
|
Filesystem routing (1.0 Tools MVC): `mvc/errors/` = **`/errors`** (the listing) and
|
|
50
50
|
`mvc/errors/issue/` = **`/errors/issue`** (the detail page).
|
|
51
51
|
|
|
52
|
+
**`/errors/issue` takes two parameters, one meaning each** (2026-08-04):
|
|
53
|
+
|
|
54
|
+
| Parameter | Means | Use |
|
|
55
|
+
| --- | --- | --- |
|
|
56
|
+
| `?issue=4KP` | `Issue.reference` — the identifier everything else quotes (ClickUp title, escalation email, neglect digest, production error page) | **all new links**, including the listing rows and worker2's `buildToolsIssueUrl()` |
|
|
57
|
+
| `?issueId=9` | `Issue.id` | the original form, kept working **indefinitely** |
|
|
58
|
+
|
|
59
|
+
`?issueId=` is never removed: links built from it are already sitting in ClickUp tasks and emails
|
|
60
|
+
sent months ago that cannot be rewritten.
|
|
61
|
+
|
|
62
|
+
> **Why two parameters instead of one that accepts either.** References are **base 32 over an
|
|
63
|
+
> alphabet that includes digits**, so an ordinary reference — `'68'`, `'137'`, `'400'` — is all
|
|
64
|
+
> digits and **cannot be distinguished from a primary key by inspection**. A combined parameter
|
|
65
|
+
> would have to guess; naming the parameter after the column removes the ambiguity. An earlier
|
|
66
|
+
> combined `?issue=` with reference-first/id-fallback precedence was written and then replaced.
|
|
67
|
+
|
|
68
|
+
`?issue=` also accepts the **event** form (`?issue=4KP-12` strips the event half), because that is
|
|
69
|
+
what people copy out of an error page or ClickUp comment. An unknown reference leaves the id at 0
|
|
70
|
+
and renders the existing "no such issue" view.
|
|
71
|
+
|
|
52
72
|
**Listing** — 25 per page with a filtered `COUNT` (it was previously a silent `LIMIT 200`, so
|
|
53
73
|
older issues simply did not exist as far as a curator could tell); bold monospace Issue ID; an
|
|
54
74
|
**EB Environment** column read from the newest event's context; a **Type** column showing
|
|
@@ -75,6 +95,22 @@ curated text, affected clients) right, with an always-present Type chip.
|
|
|
75
95
|
- Header chip order is **Issue / Environment / Type / Urgency** — Type was moved ahead of Urgency so
|
|
76
96
|
urgency finishes the row at the far right.
|
|
77
97
|
|
|
98
|
+
### Neglect indicator (2026-08-04)
|
|
99
|
+
|
|
100
|
+
Both pages now surface the **neglect** contribution to urgency: the listing computes
|
|
101
|
+
`openTaskAgeHours` as a correlated subquery and renders a `$describeNeglect` amber note
|
|
102
|
+
(`.neglect-note`, `#8a5a00`) **under the urgency badge**; the detail page's *Acknowledged* fact row
|
|
103
|
+
reads "not yet — waiting 3 days" alongside a NORMAL/HIGH neglect chip.
|
|
104
|
+
|
|
105
|
+
**Why:** the urgency badge alone is genuinely misleading. A **HIGH beside "Lifetime 1"** is a
|
|
106
|
+
statement about *response time*, not about the error, and nothing on the page let a reader tell
|
|
107
|
+
those apart. **Nothing renders while the clock is running but under 24h** — below that the clock
|
|
108
|
+
contributes nothing to urgency, so saying "neglect" would overstate it.
|
|
109
|
+
|
|
110
|
+
**Deliberately scoped to the neglect contribution only.** The console never claims *which* axis is
|
|
111
|
+
winning: computing the full driving axis in Tools would mean copying the volume, sustained and
|
|
112
|
+
release-band thresholds too, and that much duplicated policy would rot.
|
|
113
|
+
|
|
78
114
|
### Table presentation rules (established 2026-08-04)
|
|
79
115
|
|
|
80
116
|
**Stacked datetimes in every table on both pages** — listing *Last seen*; detail *Events → When*,
|
|
@@ -143,6 +179,17 @@ They were renamed from plural on 2026-08-01, after the pipeline was already in p
|
|
|
143
179
|
|
|
144
180
|
## Gotchas / known issues
|
|
145
181
|
|
|
182
|
+
- **⚠ OPEN BUG (found 2026-08-04, deliberately left unfixed pending the developer's go-ahead):
|
|
183
|
+
`tools/mvc/errors/post.php` ~line 181, the merge-fingerprint action, will merge into the WRONG
|
|
184
|
+
issue.** It takes a reference **or** a numeric id in **one** field and disambiguates with
|
|
185
|
+
`ctype_digit()`. That was safe under the old **base-26** scheme, where every reference contained a
|
|
186
|
+
letter. Under the new **base-32** alphabet an all-digit reference like `'68'` is treated as **id
|
|
187
|
+
68** and silently repoints the fingerprint at the wrong Issue. Fix is either reference-first-then-id
|
|
188
|
+
precedence or two separate inputs (as `/errors/issue` now does).
|
|
189
|
+
- **The 24/48 neglect thresholds are duplicated as plain numbers in both Tools views.** Tools is
|
|
190
|
+
1.0 and cannot read worker2's `NEGLECT_NORMAL_HOURS` / `NEGLECT_HIGH_HOURS`; both sites carry a
|
|
191
|
+
comment naming those constants as the source of truth. **Changing the policy means changing three
|
|
192
|
+
places** (worker2 + both Tools views).
|
|
146
193
|
- **A 2.0 table rename breaks this page silently until it is run.** This console has **no model
|
|
147
194
|
layer** — `get.php` and `post.php` interpolate table names into hand-written `App_Database`
|
|
148
195
|
SQL (40+ literal occurrences between them). Nothing here fails at deploy time; it fails when
|
|
@@ -170,6 +217,22 @@ They were renamed from plural on 2026-08-01, after the pipeline was already in p
|
|
|
170
217
|
|
|
171
218
|
## Change history
|
|
172
219
|
|
|
220
|
+
- 2026-08-04 (latest, **uncommitted/undeployed** at time of writing) — **Built:** a **neglect
|
|
221
|
+
indicator** on both pages — listing gained an `openTaskAgeHours` correlated subquery and a
|
|
222
|
+
`$describeNeglect` amber note (`.neglect-note`, `#8a5a00`) under the urgency badge; the detail
|
|
223
|
+
page's *Acknowledged* row now reads "not yet — waiting 3 days" with a NORMAL/HIGH neglect chip.
|
|
224
|
+
The urgency badge alone was misleading (a HIGH beside "Lifetime 1" describes response time, not
|
|
225
|
+
the error) and nothing distinguished the two; nothing renders under 24h since the clock is
|
|
226
|
+
contributing nothing yet. Scoped to the neglect contribution only — computing the winning axis
|
|
227
|
+
would mean duplicating the volume/sustained/release-band thresholds as well. **Built:**
|
|
228
|
+
`/errors/issue` now accepts **`?issue=<reference>`** (also tolerating the `4KP-12` event form) and
|
|
229
|
+
the listing links by reference; `?issueId=<id>` is kept working **indefinitely** for links already
|
|
230
|
+
sitting in old ClickUp tasks and emails. Two parameters rather than one, because base-32
|
|
231
|
+
references can be all digits and are indistinguishable from a primary key. **Recorded two
|
|
232
|
+
maintenance hazards:** the 24/48 thresholds are now duplicated as literals in both Tools views
|
|
233
|
+
(Tools is 1.0 and cannot read worker2's constants — three places to change), and an **open bug** in
|
|
234
|
+
`post.php`'s merge-fingerprint `ctype_digit()` disambiguation, which under base-32 will merge into
|
|
235
|
+
the wrong Issue (flagged, left unfixed). (jcardinal)
|
|
173
236
|
- 2026-08-04 (later, **uncommitted/undeployed** at time of writing) — UI pass on both `/errors`
|
|
174
237
|
pages: **stacked datetimes** in every table via a shared `$datetime` closure emitting
|
|
175
238
|
`.dt`/`.dt-date`/`.dt-time` (a naturally wrapping datetime broke inside the day of the month), and
|
|
@@ -18,7 +18,7 @@
|
|
|
18
18
|
| [Compass VIP Support Importer (worker2)](features/compass-vip-support-importer.md) | A worker2 action that ingests Compass's quarterly VIP spreadsheet and assigns each VIP user's support technician by setting `Users.c_supportedByUserId` in `Clie | worker2/Worker/Client/Compass/VipSupport.php |
|
|
19
19
|
| [Creating Worker Actions](features/creating-worker-actions.md) | How to add a new callable Worker action — a PHP class whose `public static` methods are invoked as background jobs (via webhook, cron, or `_Worker::runTask()`). | worker2/Worker/, worker2/Controller/Index.php, _underscore/Worker.php |
|
|
20
20
|
| [Elite Freshservice Sync (worker2)](features/elite-freshservice-sync.md) | `_Worker_Elite` processes Freshservice webhook events and syncs them into TOGA 2. | worker2/Worker/Elite.php, worker2/Config/dev-kmaramreddy-laptop.ini |
|
|
21
|
-
| [Error Escalation Cron (Errors::Escalate → ClickUp / email)](features/error-escalation-cron.md) | `_Worker_Infrastructure_Errors::Escalate` (renamed from `SyncWithClickup`) is the sole owner of **escalation, de-escalation, ClickUp ticketing, reminders, busin | worker2/Worker/Infrastructure/Errors.php, worker2/Worker/Notification/Email.php, worker2/Worker/Notification/EmailTemplate.php, worker2/Worker/Client/True.php, worker2/Worker/Clickup/ErrorTask.php, worker2/Worker/Clickup.php, worker2/Controller/Index.php, worker2/Config/production.ini, _underscore/Model/Core/Logs/Issue.php, dbchanges2/Core/2026-07-30a - Error escalation cron job.sql, dbchanges2/Logs/2026-08-03a - Issue clickupPriority.sql |
|
|
21
|
+
| [Error Escalation Cron (Errors::Escalate → ClickUp / email)](features/error-escalation-cron.md) | `_Worker_Infrastructure_Errors::Escalate` (renamed from `SyncWithClickup`) is the sole owner of **escalation, de-escalation, ClickUp ticketing, reminders, busin | worker2/Worker/Infrastructure/Errors.php, worker2/Worker/Notification/Email.php, worker2/Worker/Notification/EmailTemplate.php, worker2/Worker/Client/True.php, worker2/Worker/Clickup/ErrorTask.php, worker2/Worker/Clickup.php, worker2/Controller/Index.php, worker2/Config/production.ini, _underscore/Model/Core/Logs/Issue.php, dbchanges2/Core/2026-07-30a - Error escalation cron job.sql, dbchanges2/Logs/2026-08-03a - Issue clickupPriority.sql, dbchanges2/Core/2026-08-04a - Error neglect digest cron job.sql |
|
|
22
22
|
| [Etilize Catalog Item Import & Refresh](features/etilize-catalog-item-import.md) | Client-generic catalog onboarding from an S3 CSV plus an Etilize re-pull. | worker2/Worker/Etilize/Items.php |
|
|
23
23
|
| [Etilize Item Translation Import](features/etilize-item-translation-import.md) | The abstract worker class `_Worker_Etilize_ItemTranslations` imports **non-English** item text from Etilize into the client's `ItemTranslations` table. | worker2/Worker/Etilize/ItemTranslations.php |
|
|
24
24
|
| [Monitoring Framework (Orchestrator + Child Monitors)](features/monitoring-framework.md) | A unified, DB-driven monitoring framework for business-critical data flows (Compass POs, Prudential asset imports, AIG closed claims, …). | worker2/Worker/Monitor.php, worker2/Worker/Monitors/, worker2/Worker/Monitors/RateEntitlement.php, worker2/Worker/Notification/Email.php, worker2/Worker/Rate.php, dbchanges2/Core/2026-05-21 - Monitors.sql, dbchanges2/Core/2026-06-29a - Rate Entitlement Contract Monitor.sql |
|
|
@@ -20,6 +20,7 @@ files:
|
|
|
20
20
|
- _underscore/Model/Core/Logs/Issue.php
|
|
21
21
|
- dbchanges2/Core/2026-07-30a - Error escalation cron job.sql
|
|
22
22
|
- dbchanges2/Logs/2026-08-03a - Issue clickupPriority.sql
|
|
23
|
+
- dbchanges2/Core/2026-08-04a - Error neglect digest cron job.sql
|
|
23
24
|
related:
|
|
24
25
|
- ../../_underscore/features/error-reporting-issue-event.md
|
|
25
26
|
- ../../../1.0/apps/tools/features/errors-curation-console.md
|
|
@@ -47,13 +48,93 @@ through this one action, so there is exactly one place to mute or debug alerting
|
|
|
47
48
|
- **Volume** — occurrences in a **1-hour** window: **NORMAL 5 / HIGH 10 / URGENT 20** (retuned
|
|
48
49
|
2026-08-04 from 20/50/100 over 10 minutes — a 10-minute window never fired for the
|
|
49
50
|
low-and-steady failures that actually matter).
|
|
50
|
-
- **Neglect** — hours an *unacknowledged* ClickUp task has sat: **
|
|
51
|
+
- **Neglect** — hours an *unacknowledged* ClickUp task has sat: **NORMAL 24 / HIGH 48. Capped at
|
|
52
|
+
HIGH — there is no URGENT rung.**
|
|
51
53
|
- **Sustained-24h** — an issue still occurring a day later escalates on duration alone.
|
|
52
54
|
- **Floor** — the Issue's declared `minimumUrgency` applies under *all* axes.
|
|
53
55
|
|
|
54
56
|
The neglect axis exists because volume alone cannot see a rare-but-serious error, nor a serious
|
|
55
57
|
error that everyone is ignoring.
|
|
56
58
|
|
|
59
|
+
#### ⚠ Neglect is capped at HIGH and slowed to 24h/48h (decided + fixed 2026-08-04)
|
|
60
|
+
|
|
61
|
+
`NEGLECT_NORMAL_HOURS` 1 → **24**, `NEGLECT_HIGH_HOURS` 4 → **48**, and
|
|
62
|
+
**`NEGLECT_URGENT_HOURS` and its branch were removed** from `resolveNeglectUrgency()`. This
|
|
63
|
+
**supersedes** the earlier 1/4/24 ladder with an URGENT rung.
|
|
64
|
+
|
|
65
|
+
**Why:** neglect answers *"is anyone on this"* — a question about the **team's response**, never
|
|
66
|
+
about how much damage the error is doing. Only the volume axes measure damage, so only they may
|
|
67
|
+
claim URGENT. Left uncapped on the old clock, any error that fired once and went unacknowledged
|
|
68
|
+
for an afternoon became HIGH, then URGENT the next day. That is how the queue filled up:
|
|
69
|
+
**48 of 68 issues sat at HIGH**, and issue **3H (id 60) was HIGH off a single lifetime
|
|
70
|
+
occurrence** — volume 0 in the 60-minute window, sustained 1 against a threshold of 50,
|
|
71
|
+
`minimumUrgency` LOW, with the entire HIGH coming from its ClickUp task sitting unacknowledged
|
|
72
|
+
for 4.93 hours.
|
|
73
|
+
|
|
74
|
+
**The neglect axis is a FLOOR, not a ceiling.** `resolveNeglectUrgency()` is one of four inputs
|
|
75
|
+
to `highestUrgency()`, which takes the **max** — so a neglected issue that starts firing still
|
|
76
|
+
escalates on volume: two days unacknowledged contributes HIGH, and 20 occurrences in the window
|
|
77
|
+
take it straight to URGENT regardless. Capping the axis lowers the floor *neglect alone* can
|
|
78
|
+
raise; it takes nothing away from the occurrence count. Upward moves apply the same run (only
|
|
79
|
+
**downward** moves go through the `REQUIRED_QUIET_RUNS` dwell), so volume takeover is same-run.
|
|
80
|
+
|
|
81
|
+
**Intended side effect, accepted:** `SPRINT_PROMOTION_URGENCY` is `'URGENT'`, so with neglect
|
|
82
|
+
capped at HIGH, **neglect alone can no longer promote an issue into the sprint** — only a genuine
|
|
83
|
+
occurrence rate can.
|
|
84
|
+
|
|
85
|
+
**Deploy effect.** All 50 then-HIGH issues had open unacknowledged tasks younger than 24h, so
|
|
86
|
+
their neglect contribution becomes LOW and they recompute from volume/sustained/floor within ~3
|
|
87
|
+
runs of the worker2 deploy. Expect a **burst of ClickUp priority PUTs** on the first runs as
|
|
88
|
+
`clickupPriority` re-syncs.
|
|
89
|
+
|
|
90
|
+
### Daily neglect digest email (`Infrastructure/Errors/NeglectDigest`)
|
|
91
|
+
|
|
92
|
+
A second worker action — `public NeglectDigest()`, with `loadNeglectedIssues()` and
|
|
93
|
+
`buildNeglectDigestBody()` — registered in `Core.CronJobs` as action
|
|
94
|
+
`Infrastructure/Errors/NeglectDigest` at **`0 8 * * *`** (08:00 CT: after the overnight cleanup
|
|
95
|
+
jobs, before the working day), `maxExecutionTime` 300. Migration:
|
|
96
|
+
`dbchanges2/Core/2026-08-04a - Error neglect digest cron job.sql`.
|
|
97
|
+
|
|
98
|
+
**Why a digest exists at all — the governing principle:** neglect means *"nobody is looking"*, so
|
|
99
|
+
**every signal that requires somebody to look is self-defeating.** It is also the necessary
|
|
100
|
+
counterpart to capping the axis: with neglect slowed and capped, an ignored issue no longer
|
|
101
|
+
announces itself by climbing the urgency scale, so something has to come and tell you.
|
|
102
|
+
|
|
103
|
+
Sends via `_Worker::runTask('Notification/EmailTemplate/Send')` with the TOGA Technology stored
|
|
104
|
+
template — the **same path as the technical escalation email**. Recipient is `[_underscore]
|
|
105
|
+
error_escalation_email`, gated by the same `error_escalation_enabled` key, so non-production mails
|
|
106
|
+
nobody. The body is a **fragment** for the template's `{body}` (600px cell; the template owns font
|
|
107
|
+
and width), exactly like `buildTechnicalEmailBody()`.
|
|
108
|
+
|
|
109
|
+
Three design decisions, each load-bearing:
|
|
110
|
+
|
|
111
|
+
- **It sends even when the list is empty.** An empty digest states *"nothing is neglected"*;
|
|
112
|
+
silence cannot be distinguished from a broken notifier — which is precisely the failure that hid
|
|
113
|
+
the escalation email being dead for days while every run reported success. The **subject
|
|
114
|
+
reflects the state** ("No neglected error issues" vs "2 error issues nobody has picked up") so
|
|
115
|
+
it is triageable from the inbox list without opening it.
|
|
116
|
+
- **One email covering every issue, never one per issue.** A per-issue alert for a condition that
|
|
117
|
+
by nature affects many issues at once is how a channel becomes wallpaper — and the reminder
|
|
118
|
+
ladder already carries the per-issue nagging.
|
|
119
|
+
- **Floored at `NEGLECT_NORMAL_HOURS`**, not "everything open". That constant is where the cron
|
|
120
|
+
itself starts treating an issue as neglected, so the digest reports only things that have
|
|
121
|
+
crossed a line. A task opened this morning is not neglected, it is new. Ordered **oldest
|
|
122
|
+
first** so the worst case is the first line read.
|
|
123
|
+
|
|
124
|
+
Verified against production: the query returns **2 rows** today (1B and 1I, both 24h
|
|
125
|
+
unacknowledged).
|
|
126
|
+
|
|
127
|
+
**Rejected alternatives:**
|
|
128
|
+
|
|
129
|
+
- *Mutating the ClickUp task **title** to say "neglected"* — the title is deliberately just the
|
|
130
|
+
reference and is never rewritten; rewriting churns the activity feed and breaks the
|
|
131
|
+
quote-the-reference contract, and it only helps someone already looking at the task.
|
|
132
|
+
- *A new `Issue` column (`dtNeglectedSince`, `neglectLevel`)* — fully derivable from
|
|
133
|
+
`IssueClickupTask.dtCreated` + `Issue.dtAcknowledged`, so storing it is denormalised state that
|
|
134
|
+
can drift, and it still needs something to read it out.
|
|
135
|
+
- *A ClickUp custom field* — **deferred, not rejected**: it writes every run and duplicates what
|
|
136
|
+
the digest says. Revisit only if the digest proves insufficient.
|
|
137
|
+
|
|
57
138
|
### Fast up, slow down
|
|
58
139
|
|
|
59
140
|
Escalation applies **the same run** a band is crossed. De-escalation requires the count to sit
|
|
@@ -108,6 +189,15 @@ ClickUp is the acknowledge/resolve system — developers already live there, and
|
|
|
108
189
|
Tools would just be ignored. **Any** status change on an error task counts as acknowledgement,
|
|
109
190
|
which is safe only because the automation never sets status itself.
|
|
110
191
|
|
|
192
|
+
> **`dtAcknowledged` is set by a status change OR by an assignee being added** — both, and this was
|
|
193
|
+
> already implemented (confirmed 2026-08-04, no change needed). `worker2/Worker/Clickup.php:856`
|
|
194
|
+
> routes the ClickUp `taskUpdated` webhook into
|
|
195
|
+
> `_Worker_Clickup_ErrorTask::handleAssigneeChange()`, which calls `markAcknowledged()` (and also
|
|
196
|
+
> feeds the per-area ownership learner). Both routes are subject to the 60-second creation grace
|
|
197
|
+
> window below. Production: 8 of 71 issues acknowledged. **Worth stating explicitly, because the
|
|
198
|
+
> acknowledgement definition is what both the neglect axis and the daily digest key off**, and it
|
|
199
|
+
> is not visible from the escalation-cron side.
|
|
200
|
+
|
|
111
201
|
> **60-second creation grace window (fix, 2026-08-04).** ClickUp fires a status webhook for a
|
|
112
202
|
> task's *initial* status, so the automation was acknowledging its **own** newly created tasks —
|
|
113
203
|
> stamping `dtAcknowledged` ~1s after `dtCreated` and permanently disabling the neglect axis for
|
|
@@ -176,6 +266,32 @@ healthy — cannot recur silently.
|
|
|
176
266
|
- **Occurrences** and **Last Occurrence** custom fields are synced, and task bodies open with a
|
|
177
267
|
**"Debug in Tools"** deep link.
|
|
178
268
|
|
|
269
|
+
#### Tools deep links address the issue by REFERENCE, not id
|
|
270
|
+
|
|
271
|
+
`TOOLS_ISSUE_PATH` changed from `'/errors/issue?issueId='` to **`'/errors/issue?issue='`**, and
|
|
272
|
+
`buildToolsIssueUrl()` now emits the **reference**, `rawurlencode`d, falling back to the id only
|
|
273
|
+
when a row has no reference. **Two parameters, one meaning each** — deliberately not one parameter
|
|
274
|
+
that accepts either:
|
|
275
|
+
|
|
276
|
+
| Parameter | Means | Use |
|
|
277
|
+
| --- | --- | --- |
|
|
278
|
+
| `?issue=4KP` | `Issue.reference` — what everything else quotes (ClickUp title, escalation email, digest, production error page) | **all new links** |
|
|
279
|
+
| `?issueId=9` | `Issue.id` | the original form, kept working **indefinitely** |
|
|
280
|
+
|
|
281
|
+
`?issueId=` stays supported forever because links built from it are already sitting in ClickUp
|
|
282
|
+
tasks and emails sent months ago that cannot be rewritten.
|
|
283
|
+
|
|
284
|
+
> **Why separate parameters rather than one that accepts either.** References are **base 32 over
|
|
285
|
+
> an alphabet that includes digits**, so an ordinary reference — `'68'`, `'137'`, `'400'` — is all
|
|
286
|
+
> digits and **cannot be told apart from a primary key by inspection**. A single combined parameter
|
|
287
|
+
> would have to guess; naming the parameter after the column removes the ambiguity entirely. (An
|
|
288
|
+
> earlier combined `?issue=` with reference-first / id-fallback precedence was written and then
|
|
289
|
+
> replaced by this.)
|
|
290
|
+
|
|
291
|
+
`?issue=` also accepts the **event** form — `?issue=4KP-12` strips the event half and lands on the
|
|
292
|
+
issue — because that is what people copy out of an error page or a ClickUp comment. An unknown
|
|
293
|
+
reference leaves the id at 0 and renders the existing "no such issue" view.
|
|
294
|
+
|
|
179
295
|
### Area-level owner learning (observation only)
|
|
180
296
|
|
|
181
297
|
`IssueAreaOwner` learns an owner after **5 in a row** and ships **observation-only** — a hint in
|
|
@@ -232,6 +348,12 @@ this controller needs the same treatment.**
|
|
|
232
348
|
|
|
233
349
|
## Gotchas / known issues
|
|
234
350
|
|
|
351
|
+
- **An issue held at HIGH purely by neglect never de-escalates on its own.**
|
|
352
|
+
`applyDeEscalationDwell()` needs the computed target to sit **below** the stored urgency for
|
|
353
|
+
`REQUIRED_QUIET_RUNS` (3) consecutive runs, but the neglect axis keeps recomputing the target
|
|
354
|
+
**at** the stored level — so the dwell counter never starts. That is why issue 3H showed
|
|
355
|
+
`consecutiveQuietWindows = 0` after five hours of total silence. **Acknowledging the ClickUp task
|
|
356
|
+
is what releases it**, and it then drops within 3 cron runs (~3 min).
|
|
235
357
|
- **Never send email inline inside this cron's open transaction — enqueue it.** `_Email::send()`
|
|
236
358
|
registers a second connection and toggles the read host, and a throw from in there was only ever
|
|
237
359
|
`error_log`'d. Use `_Worker::runTask` so a failure becomes a visible failed `WorkerJobs` row.
|
|
@@ -299,6 +421,29 @@ this controller needs the same treatment.**
|
|
|
299
421
|
**non-retroactive** — the first email lands on the next live occurrence. Also noted that prod
|
|
300
422
|
`Logs.IssueEmailAddress` was still empty, making Compass's MITS rejection the platform's first
|
|
301
423
|
business-routed issue. No code change. (dfranks)
|
|
424
|
+
- 2026-08-04 (latest, **uncommitted/undeployed** at time of writing) — **Decided + fixed: the
|
|
425
|
+
neglect urgency axis is capped at HIGH and slowed to 24h/48h**, superseding the 1/4/24 ladder
|
|
426
|
+
recorded earlier the same day (`NEGLECT_NORMAL_HOURS` 1→24, `NEGLECT_HIGH_HOURS` 4→48,
|
|
427
|
+
`NEGLECT_URGENT_HOURS` and its branch removed). Neglect measures the **team's response**, not the
|
|
428
|
+
damage the error is doing, so only the volume axes may claim URGENT; uncapped on the old clock a
|
|
429
|
+
single-occurrence error became HIGH in an afternoon and URGENT the next day — 48 of 68 issues were
|
|
430
|
+
at HIGH and issue 3H (id 60) was HIGH off one lifetime occurrence. The axis remains a **floor, not
|
|
431
|
+
a ceiling** (`highestUrgency()` takes the max, so volume still escalates immediately). Accepted
|
|
432
|
+
side effect: `SPRINT_PROMOTION_URGENCY` is URGENT, so neglect alone can no longer promote into the
|
|
433
|
+
sprint. **Built:** a **daily neglect digest** (`Infrastructure/Errors/NeglectDigest`,
|
|
434
|
+
`Core.CronJobs` at `0 8 * * *`, migration `dbchanges2/Core/2026-08-04a`) — the counterpart to the
|
|
435
|
+
cap, since a capped axis no longer announces an ignored issue; it **sends even when empty** (a
|
|
436
|
+
broken notifier is indistinguishable from silence), is **one email for all issues**, and floors at
|
|
437
|
+
`NEGLECT_NORMAL_HOURS`, oldest first. Rejected: rewriting the ClickUp title, a derived
|
|
438
|
+
`dtNeglectedSince`/`neglectLevel` column; deferred: a ClickUp custom field. **Built:**
|
|
439
|
+
`TOOLS_ISSUE_PATH` now emits `?issue=<reference>` (rawurlencoded, id fallback) with `?issueId=`
|
|
440
|
+
kept working indefinitely — separate parameters because base-32 references can be all digits and
|
|
441
|
+
are indistinguishable from a primary key. **Discovered:** assignee-based acknowledgement was
|
|
442
|
+
already implemented (`Clickup.php:856` → `handleAssigneeChange()` → `markAcknowledged()`), and
|
|
443
|
+
that a neglect-held HIGH never de-escalates until the task is acknowledged. Verification:
|
|
444
|
+
`php -l` clean on every changed file; the digest SQL run against production (2 rows); neglect
|
|
445
|
+
impact quantified across all 68 issues; `0 8 * * *` checked against existing daily job
|
|
446
|
+
conventions. (jcardinal)
|
|
302
447
|
- 2026-08-04 (later, **uncommitted/undeployed** at time of writing) — **Fixed: escalation email had
|
|
303
448
|
never been sent, ever**, for two independent reasons. (A) `processTechnicalIssue()` returned early
|
|
304
449
|
when the ClickUp task would read identically, and that return sat *above* the `if ($shouldNotify)`
|
package/package.json
CHANGED