toga-ai 1.0.532 → 1.0.533
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/knowledge/1.0/apps/tools/features/errors-curation-console.md +22 -0
- package/knowledge/2.0/apps/_underscore/features/error-reporting-issue-event.md +27 -0
- package/knowledge/2.0/apps/worker2/INDEX.md +1 -0
- package/knowledge/2.0/apps/worker2/features/error-escalation-cron.md +10 -0
- package/knowledge/2.0/apps/worker2/features/error-issue-auto-resolution.md +209 -0
- package/knowledge/INDEX.md +1 -1
- package/package.json +1 -1
|
@@ -20,6 +20,7 @@ related:
|
|
|
20
20
|
- ./mvc-data-access-patterns.md
|
|
21
21
|
- ../../../2.0/apps/_underscore/features/error-reporting-issue-event.md
|
|
22
22
|
- ../../../2.0/apps/worker2/features/error-escalation-cron.md
|
|
23
|
+
- ../../../2.0/apps/worker2/features/error-issue-auto-resolution.md
|
|
23
24
|
- ../../library/features/error-capture-1-0.md
|
|
24
25
|
---
|
|
25
26
|
|
|
@@ -183,6 +184,20 @@ second line — the two halves then read as unrelated fragments. It is now `disp
|
|
|
183
184
|
winning: computing the full driving axis in Tools would mean copying the volume, sustained and
|
|
184
185
|
release-band thresholds too, and that much duplicated policy would rot.
|
|
185
186
|
|
|
187
|
+
### Issue status: Resolved vs. Closed (2026-08-05)
|
|
188
|
+
|
|
189
|
+
Both pages surface the new Issue-level **OPEN → RESOLVED** lifecycle (owned by the worker2
|
|
190
|
+
frequency-decay cron — see
|
|
191
|
+
[error-issue-auto-resolution](../../../2.0/apps/worker2/features/error-issue-auto-resolution.md)):
|
|
192
|
+
|
|
193
|
+
- An **Open / Resolved / All** status filter (**default Open**; the submitted value is
|
|
194
|
+
**allowlisted** before it reaches the query).
|
|
195
|
+
- A **Resolved badge** on the listing rows and the detail header.
|
|
196
|
+
- A **`dtAutoResolved`** fact row on the detail page.
|
|
197
|
+
- **⚠ The ClickUp-episodes table column "Resolved" was renamed to "Closed"** to break the naming
|
|
198
|
+
collision. In this console an **episode** is *Closed* (`IssueClickupTask.dtResolved`); an
|
|
199
|
+
**Issue** is *Resolved* (`Issue.status` / `dtAutoResolved`). Keep the two labels distinct.
|
|
200
|
+
|
|
186
201
|
### Table presentation rules (established 2026-08-04)
|
|
187
202
|
|
|
188
203
|
**Stacked datetimes in every table on both pages** — listing *Last seen*; detail *Events → When*,
|
|
@@ -297,6 +312,13 @@ They were renamed from plural on 2026-08-01, after the pipeline was already in p
|
|
|
297
312
|
|
|
298
313
|
## Change history
|
|
299
314
|
|
|
315
|
+
- 2026-08-05 — **Built:** the `/errors` console surfaces the new Issue-level RESOLVED lifecycle —
|
|
316
|
+
an **Open/Resolved/All** status filter (default Open, allowlisted value), a **Resolved badge**
|
|
317
|
+
(listing + detail header), a **`dtAutoResolved`** fact row, and the ClickUp-episodes column
|
|
318
|
+
header renamed **"Resolved" → "Closed"** to break the collision with the new Issue-level status
|
|
319
|
+
(an *episode* is Closed; an *Issue* is Resolved). Lifecycle behaviour is owned by the worker2 cron
|
|
320
|
+
([error-issue-auto-resolution](../../../2.0/apps/worker2/features/error-issue-auto-resolution.md)).
|
|
321
|
+
(jcardinal)
|
|
300
322
|
- 2026-08-04 (latest, **uncommitted/undeployed**; CSS/markup/nav only — no PHP logic, schema or
|
|
301
323
|
worker change) — **Fixed:** the context cards were too narrow and too short. `.context-card-body`
|
|
302
324
|
`max-height` 340px → 510px, and the grid's `repeat(auto-fit, minmax(320px, 1fr))` (five columns on
|
|
@@ -29,6 +29,7 @@ related:
|
|
|
29
29
|
- ./per-client-database-connections.md
|
|
30
30
|
- ./email-send-pipeline.md
|
|
31
31
|
- ../../worker2/features/error-escalation-cron.md
|
|
32
|
+
- ../../worker2/features/error-issue-auto-resolution.md
|
|
32
33
|
- ../../../1.0/apps/tools/features/errors-curation-console.md
|
|
33
34
|
- ../../../1.0/apps/library/features/error-capture-1-0.md
|
|
34
35
|
- ../../api2/features/v2-api-error-codes.md
|
|
@@ -77,6 +78,24 @@ Issue itself. Anything that reads like "this client's issue" is a modelling erro
|
|
|
77
78
|
|
|
78
79
|
`IssueEmailAddresses.clientId = 0` means **all clients**.
|
|
79
80
|
|
|
81
|
+
### Issue-level lifecycle columns (added 2026-08-05)
|
|
82
|
+
|
|
83
|
+
`Issue` carries an **OPEN → RESOLVED** lifecycle written only by the worker2 frequency-decay cron:
|
|
84
|
+
`status enum('OPEN','RESOLVED')`, `dtAutoResolved`, and a durable firing baseline
|
|
85
|
+
(`baselineGapSeconds`, `baselineSampleCount`, `dtBaselineAnchor`, `baselineAnchorOccurrences`) —
|
|
86
|
+
migration `dbchanges2/Logs/2026-08-05a`. Full behaviour lives in
|
|
87
|
+
[error-issue-auto-resolution](../../worker2/features/error-issue-auto-resolution.md). Two identity
|
|
88
|
+
rules matter from the data-model side:
|
|
89
|
+
|
|
90
|
+
> **⚠ `Issue.dtAutoResolved` is deliberately NOT named `dtResolved`.** `dtResolved` already means
|
|
91
|
+
> the **episode**-level `IssueClickupTask.dtResolved`; a same-named `Issue` column would invite a
|
|
92
|
+
> wrong join between the Issue lifecycle and an episode close. Keep the two levels distinct — an
|
|
93
|
+
> **episode** closes on every ClickUp completion; an **Issue** resolves only when firing stops.
|
|
94
|
+
>
|
|
95
|
+
> **⚠ Never maintain the baseline scalars from the capture path** (`_underscore/Error.php` or the
|
|
96
|
+
> mirrored `library/app/error/capture.php`) — they are owned by the cron, because `Event` purges at
|
|
97
|
+
> 7 days but a low-frequency Issue's baseline must persist for months.
|
|
98
|
+
|
|
80
99
|
## How it works
|
|
81
100
|
|
|
82
101
|
### Fingerprinting (identity)
|
|
@@ -520,6 +539,14 @@ clientUserId). **Neither was built.** As built instead:
|
|
|
520
539
|
|
|
521
540
|
## Change history
|
|
522
541
|
|
|
542
|
+
- 2026-08-05 — **Data model:** `Issue` gained an Issue-level **OPEN → RESOLVED** lifecycle
|
|
543
|
+
(`status`, `dtAutoResolved`) plus a durable firing baseline (`baselineGapSeconds`,
|
|
544
|
+
`baselineSampleCount`, `dtBaselineAnchor`, `baselineAnchorOccurrences`), migration
|
|
545
|
+
`dbchanges2/Logs/2026-08-05a`. Recorded the two identity rules: `dtAutoResolved` is **not**
|
|
546
|
+
`dtResolved` (that token is the episode-level `IssueClickupTask.dtResolved`), and the baseline
|
|
547
|
+
scalars are cron-owned — **never** written from the capture path, since `Event` purges at 7 days
|
|
548
|
+
while the baseline must persist for months. Behaviour lives in
|
|
549
|
+
[error-issue-auto-resolution](../../worker2/features/error-issue-auto-resolution.md). (jcardinal)
|
|
523
550
|
- 2026-08-04 (latest, **uncommitted/undeployed** at time of writing) — **DECISION REVERSED:
|
|
524
551
|
`$GLOBALS` IS now captured**, as a new top-level `GLOBALS` group, in both frameworks. The earlier
|
|
525
552
|
entry recording that `$GLOBALS` would deliberately never be captured is **superseded and has been
|
|
@@ -19,6 +19,7 @@
|
|
|
19
19
|
| [Creating Worker Actions](features/creating-worker-actions.md) | How to add a new callable Worker action — a PHP class whose `public static` methods are invoked as background jobs (via webhook, cron, or `_Worker::runTask()`). | worker2/Worker/, worker2/Controller/Index.php, _underscore/Worker.php |
|
|
20
20
|
| [Elite Freshservice Sync (worker2)](features/elite-freshservice-sync.md) | `_Worker_Elite` processes Freshservice webhook events and syncs them into TOGA 2. | worker2/Worker/Elite.php, worker2/Config/dev-kmaramreddy-laptop.ini |
|
|
21
21
|
| [Error Escalation Cron (Errors::Escalate → ClickUp / email)](features/error-escalation-cron.md) | `_Worker_Infrastructure_Errors::Escalate` (renamed from `SyncWithClickup`) is the sole owner of **escalation, de-escalation, ClickUp ticketing, reminders, busin | worker2/Worker/Infrastructure/Errors.php, worker2/Worker/Notification/Email.php, worker2/Worker/Notification/EmailTemplate.php, worker2/Worker/Client/True.php, worker2/Worker/Clickup/ErrorTask.php, worker2/Worker/Clickup.php, worker2/Controller/Index.php, worker2/Config/production.ini, _underscore/Model/Core/Logs/Issue.php, dbchanges2/Core/2026-07-30a - Error escalation cron job.sql, dbchanges2/Logs/2026-08-03a - Issue clickupPriority.sql, dbchanges2/Core/2026-08-04a - Error neglect digest cron job.sql |
|
|
22
|
+
| [Error-Issue Auto-Resolution & Reopen (frequency-decay lifecycle)](features/error-issue-auto-resolution.md) | The error system could escalate and de-escalate an Issue's *urgency* but had no concept of an Issue being **resolved**. | worker2/Worker/Infrastructure/Errors.php, worker2/Worker/Clickup/ErrorTask.php, _underscore/Model/Core/Logs/Issue.php, tools/mvc/errors/get.php, tools/mvc/errors/issue/get.php, dbchanges2/Logs/2026-08-05a - Issue status baseline and auto-resolution.sql |
|
|
22
23
|
| [Etilize Catalog Item Import & Refresh](features/etilize-catalog-item-import.md) | Client-generic catalog onboarding from an S3 CSV plus an Etilize re-pull. | worker2/Worker/Etilize/Items.php |
|
|
23
24
|
| [Etilize Item Translation Import](features/etilize-item-translation-import.md) | The abstract worker class `_Worker_Etilize_ItemTranslations` imports **non-English** item text from Etilize into the client's `ItemTranslations` table. | worker2/Worker/Etilize/ItemTranslations.php |
|
|
24
25
|
| [Monitoring Framework (Orchestrator + Child Monitors)](features/monitoring-framework.md) | A unified, DB-driven monitoring framework for business-critical data flows (Compass POs, Prudential asset imports, AIG closed claims, …). | worker2/Worker/Monitor.php, worker2/Worker/Monitors/, worker2/Worker/Monitors/RateEntitlement.php, worker2/Worker/Notification/Email.php, worker2/Worker/Rate.php, dbchanges2/Core/2026-05-21 - Monitors.sql, dbchanges2/Core/2026-06-29a - Rate Entitlement Contract Monitor.sql |
|
|
@@ -25,6 +25,7 @@ related:
|
|
|
25
25
|
- ../../_underscore/features/error-reporting-issue-event.md
|
|
26
26
|
- ../../../1.0/apps/tools/features/errors-curation-console.md
|
|
27
27
|
- ../../../1.0/apps/library/features/error-capture-1-0.md
|
|
28
|
+
- ./error-issue-auto-resolution.md
|
|
28
29
|
- ./creating-worker-actions.md
|
|
29
30
|
- ./clickup-project-routing.md
|
|
30
31
|
- ./notification-email.md
|
|
@@ -420,6 +421,15 @@ this controller needs the same treatment.**
|
|
|
420
421
|
|
|
421
422
|
## Change history
|
|
422
423
|
|
|
424
|
+
- 2026-08-05 — **Built (separate doc):** the cron now owns an Issue-level **OPEN → RESOLVED**
|
|
425
|
+
lifecycle — it auto-resolves an Issue once its events decay to zero for a duration proportional to
|
|
426
|
+
the Issue's own firing cadence (durable EWMA baseline on `Issue`), closes the ClickUp ticket
|
|
427
|
+
DB-first, and auto-reopens under the same id on recurrence. Resolution keys only on the baseline +
|
|
428
|
+
quiet duration, never on `urgency`/`minimumUrgency`. The `ErrorTask.php` webhook still closes
|
|
429
|
+
episodes/resets counters but **never** sets `Issue.status`; `recurrenceCount` still increments
|
|
430
|
+
exactly once per episode close. Full mechanics, constants, the two resolution paths
|
|
431
|
+
(`autoResolveIssue` + the ticketless `sweepResolveQuietIssues` sweep) and the sweep's index live in
|
|
432
|
+
[error-issue-auto-resolution](./error-issue-auto-resolution.md). (jcardinal)
|
|
423
433
|
- 2026-08-04 (verified against `origin/_production` + the prod Logs cluster) — Recorded the
|
|
424
434
|
**recipient-routing semantics** a developer needs before configuring a business alert: routing is
|
|
425
435
|
**both** business email and dev-team ClickUp, evaluated independently (adding a recipient is
|
|
@@ -0,0 +1,209 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Error-Issue Auto-Resolution & Reopen (frequency-decay lifecycle)
|
|
3
|
+
framework: "2.0"
|
|
4
|
+
repo: worker2
|
|
5
|
+
project: Worker
|
|
6
|
+
client: shared
|
|
7
|
+
type: feature
|
|
8
|
+
status: active
|
|
9
|
+
updated: 2026-08-05
|
|
10
|
+
owners: ["jcardinal"]
|
|
11
|
+
files:
|
|
12
|
+
- worker2/Worker/Infrastructure/Errors.php
|
|
13
|
+
- worker2/Worker/Clickup/ErrorTask.php
|
|
14
|
+
- _underscore/Model/Core/Logs/Issue.php
|
|
15
|
+
- tools/mvc/errors/get.php
|
|
16
|
+
- tools/mvc/errors/issue/get.php
|
|
17
|
+
- dbchanges2/Logs/2026-08-05a - Issue status baseline and auto-resolution.sql
|
|
18
|
+
related:
|
|
19
|
+
- ./error-escalation-cron.md
|
|
20
|
+
- ../../_underscore/features/error-reporting-issue-event.md
|
|
21
|
+
- ../../../1.0/apps/tools/features/errors-curation-console.md
|
|
22
|
+
---
|
|
23
|
+
|
|
24
|
+
## Summary
|
|
25
|
+
|
|
26
|
+
The error system could escalate and de-escalate an Issue's *urgency* but had no concept of an
|
|
27
|
+
Issue being **resolved**. This feature adds an Issue-level **OPEN → RESOLVED** lifecycle to the
|
|
28
|
+
every-minute escalation cron (`_Worker_Infrastructure_Errors`): once an Issue's events **decay to
|
|
29
|
+
zero for a duration proportional to how often that Issue normally fires**, the cron auto-resolves
|
|
30
|
+
it (and closes its ClickUp ticket); if the same Issue recurs, the cron auto-reopens it **under the
|
|
31
|
+
same internal id**. Resolution keys **only** on a trusted firing baseline plus a quiet duration —
|
|
32
|
+
never on human severity (`urgency`/`minimumUrgency`).
|
|
33
|
+
|
|
34
|
+
This complements urgency escalation (see
|
|
35
|
+
[error-escalation-cron](./error-escalation-cron.md)) and the Issue/Event capture model (see
|
|
36
|
+
[error-reporting-issue-event](../../_underscore/features/error-reporting-issue-event.md)). It does
|
|
37
|
+
**not** replace the existing episode/recurrence mechanics — episodes and ClickUp task linking were
|
|
38
|
+
already built; this adds the Issue-level status flip on top.
|
|
39
|
+
|
|
40
|
+
## Data model additions (`Logs.Issue`)
|
|
41
|
+
|
|
42
|
+
Migration: `dbchanges2/Logs/2026-08-05a - Issue status baseline and auto-resolution.sql` (run on
|
|
43
|
+
the **Logs** cluster).
|
|
44
|
+
|
|
45
|
+
- **`status enum('OPEN','RESOLVED')`** + **`dtAutoResolved`** — the Issue-level lifecycle.
|
|
46
|
+
- **Durable firing baseline** (all scalars on `Issue`, cron-maintained):
|
|
47
|
+
`baselineGapSeconds` (EWMA of inter-arrival gap), `baselineSampleCount`, `dtBaselineAnchor`,
|
|
48
|
+
`baselineAnchorOccurrences`.
|
|
49
|
+
- New index **`Issue_status_dtLastOccurred (status, dtLastOccurred)`** for the ticketless sweep
|
|
50
|
+
(see Performance).
|
|
51
|
+
|
|
52
|
+
> **⚠ The resolution column is `dtAutoResolved`, NOT `dtResolved` — deliberately.** The token
|
|
53
|
+
> `dtResolved` already means the **episode**-level `IssueClickupTask.dtResolved` (~6 usages). A
|
|
54
|
+
> same-named column on `Issue` would invite a wrong join between an Issue's lifecycle and an
|
|
55
|
+
> episode's close. The two levels are distinct: an episode closes every time the ClickUp ticket is
|
|
56
|
+
> completed; the Issue resolves only when firing actually stops.
|
|
57
|
+
|
|
58
|
+
## How it works
|
|
59
|
+
|
|
60
|
+
### Durable frequency baseline (why not the Event table)
|
|
61
|
+
|
|
62
|
+
`Events` **purge after 7 days**, but a low-frequency Issue may need to stay quiet for **months**
|
|
63
|
+
before it is safe to call resolved. So "how often does this Issue fire" cannot be recomputed from
|
|
64
|
+
`Event` — it lives on durable `Issue` scalars instead.
|
|
65
|
+
|
|
66
|
+
- `computeBaselineUpdate` maintains an **EWMA of the inter-arrival gap** (`baselineGapSeconds`),
|
|
67
|
+
folded into `persistIssueState` on the worker2 cron. `BASELINE_EWMA_ALPHA_PERCENT = 30`.
|
|
68
|
+
- **Bursts** that fire many times between two 1-minute runs are handled by dividing elapsed time
|
|
69
|
+
by the **`totalOccurrences` delta**, so a burst does not read as one giant gap.
|
|
70
|
+
- The baseline is **trusted** (Issue eligible to resolve) only once
|
|
71
|
+
`baselineSampleCount >= BASELINE_MIN_SAMPLES (3)`.
|
|
72
|
+
|
|
73
|
+
> **⚠ The baseline is maintained ONLY on the worker2 cron path, never on the capture path.** Do
|
|
74
|
+
> **not** update these scalars from `_underscore/Error.php` or `library/app/error/capture.php` (the
|
|
75
|
+
> byte-mirrored capture path). Baseline maintenance belongs to `computeBaselineUpdate` /
|
|
76
|
+
> `persistIssueState` in the cron.
|
|
77
|
+
|
|
78
|
+
### The quiet-window formula
|
|
79
|
+
|
|
80
|
+
```
|
|
81
|
+
requiredQuiet = clamp(RESOLVE_QUIET_MULTIPLE * baselineGapSeconds,
|
|
82
|
+
RESOLVE_MIN_QUIET_HOURS, RESOLVE_MAX_QUIET_DAYS)
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
Constants in `Errors.php`: `RESOLVE_QUIET_MULTIPLE = 3`, `RESOLVE_MIN_QUIET_HOURS = 24`,
|
|
86
|
+
`RESOLVE_MAX_QUIET_DAYS = 90`. A high-frequency Issue therefore needs only the **24h floor** of
|
|
87
|
+
silence; a monthly Issue needs roughly **90 days**. An Issue with events decayed to zero for
|
|
88
|
+
`requiredQuiet` (and a trusted baseline) is resolved.
|
|
89
|
+
|
|
90
|
+
> **Resolution keys ONLY on trusted baseline + quiet duration — never on `urgency` or
|
|
91
|
+
> `minimumUrgency`.** Product decision: an Issue that has genuinely stopped firing is resolved
|
|
92
|
+
> regardless of the human severity floor. Severity governs *how loudly* an active Issue is
|
|
93
|
+
> escalated, not *whether a silent one is closed*.
|
|
94
|
+
|
|
95
|
+
### Two resolution paths (a working-set coverage gap)
|
|
96
|
+
|
|
97
|
+
The every-minute working set does not select every Issue — a quiet LOW Issue with no open ClickUp
|
|
98
|
+
task is never picked up. Resolution is therefore split by division of labour:
|
|
99
|
+
|
|
100
|
+
1. **`autoResolveIssue`** (inside `processIssue`) — for Issues **with an open ClickUp ticket**,
|
|
101
|
+
which are always in the working set via the open-task branch. It also PUTs ClickUp
|
|
102
|
+
`status = 'complete'`.
|
|
103
|
+
2. **`sweepResolveQuietIssues`** — a **set-based sweep** for **ticketless** Issues that have
|
|
104
|
+
fallen out of the every-minute working set. This is what closes the coverage gap.
|
|
105
|
+
|
|
106
|
+
**`autoResolveIssue` is DB-first, then best-effort ClickUp:** close the episode → set
|
|
107
|
+
`status = RESOLVED`, `urgency = LOW`, `occurrences = 0`, `consecutiveQuietWindows = 0`, commit →
|
|
108
|
+
**then** PUT ClickUp `status = 'complete'` (logged and swallowed on failure). DB-first ordering
|
|
109
|
+
makes the ClickUp webhook echo an **idempotent no-op** (episode already closed → 0 rows). The
|
|
110
|
+
ClickUp terminal string is **`'complete'`** (`_Worker_Clickup`, `COMPLETE_STATUS_TYPES =
|
|
111
|
+
['done','closed']`).
|
|
112
|
+
|
|
113
|
+
**No team email is sent on resolve** — guaranteed by an **unconditional early return in
|
|
114
|
+
`processIssue`** before `shouldNotify` / email / sprint logic.
|
|
115
|
+
|
|
116
|
+
### Reopen (same id, never the old ticket)
|
|
117
|
+
|
|
118
|
+
- **`sweepReopenResolvedIssues`** + **`reopenIssue`** flip `status` back to **OPEN** when
|
|
119
|
+
`dtLastOccurred > dtAutoResolved`.
|
|
120
|
+
- On reopen, the baseline anchor is **re-seeded**: `dtBaselineAnchor = dtLastOccurred` and
|
|
121
|
+
`baselineAnchorOccurrences = totalOccurrences`, so the long quiet gap is **not** folded into the
|
|
122
|
+
EWMA (see the reopen bug below).
|
|
123
|
+
- The recurrence ticketing itself (a **new** ClickUp ticket linked to prior episodes via
|
|
124
|
+
`createClickupTask` + `loadPreviousTaskIdentifiers` + `POST /task/{new}/link/{prev}`) was
|
|
125
|
+
**already built** — this feature only adds the Issue-level status flip. The **same internal Issue
|
|
126
|
+
id** is reused; the old ticket is never reopened.
|
|
127
|
+
|
|
128
|
+
### Webhook safety — the ClickUp side must NEVER set `Issue.status`
|
|
129
|
+
|
|
130
|
+
`ErrorTask.php` `resolveEpisode` closes the episode and resets counters via the shared
|
|
131
|
+
`_Worker_Infrastructure_Errors::resetIssueForNextEpisode`, but it **never** writes
|
|
132
|
+
`Issue.status`. A developer completing a ClickUp ticket is treated as a *hypothesis* that the cron
|
|
133
|
+
then verifies by watching whether events actually stop — only the frequency-decay cron may write
|
|
134
|
+
`RESOLVED`.
|
|
135
|
+
|
|
136
|
+
`recurrenceCount` increments **exactly once per episode close** (a single atomic
|
|
137
|
+
`dtResolved IS NULL` guard), whether the closer is the webhook or the cron's `closeOpenEpisode`, so
|
|
138
|
+
the DB-first resolve → webhook echo can never double-count.
|
|
139
|
+
|
|
140
|
+
## Performance
|
|
141
|
+
|
|
142
|
+
`sweepResolveQuietIssues` puts a constant `dtLastOccurred < NOW() - INTERVAL 86400 SECOND`
|
|
143
|
+
pre-filter **before** the per-row `LEAST`/`GREATEST` clamp, so the new
|
|
144
|
+
`Issue_status_dtLastOccurred (status, dtLastOccurred)` index can **range-scan** rather than
|
|
145
|
+
degrading to a status-only lookup. The correlated `NOT EXISTS` (ticketless test) uses
|
|
146
|
+
`IssueClickupTask (issueId, dtResolved)`.
|
|
147
|
+
|
|
148
|
+
## Tools `/errors` UI
|
|
149
|
+
|
|
150
|
+
(See [errors-curation-console](../../../1.0/apps/tools/features/errors-curation-console.md) for the
|
|
151
|
+
console's own conventions.)
|
|
152
|
+
|
|
153
|
+
- **Open / Resolved / All** status filter (default **Open**; the request value is allowlisted).
|
|
154
|
+
- **Resolved badge** on both the listing and the detail header.
|
|
155
|
+
- **`dtAutoResolved`** rendered as a fact row on the detail page.
|
|
156
|
+
- The episode-table column header **"Resolved" → "Closed"** — deliberately renamed to break the
|
|
157
|
+
naming collision with the new **Issue-level** `status`. In the UI, an **episode** is now *Closed*;
|
|
158
|
+
an **Issue** is *Resolved*.
|
|
159
|
+
|
|
160
|
+
## Deploy order
|
|
161
|
+
|
|
162
|
+
**Run the SQL migration first, then deploy the code:**
|
|
163
|
+
|
|
164
|
+
1. `dbchanges2/Logs/2026-08-05a - Issue status baseline and auto-resolution.sql` on the **Logs**
|
|
165
|
+
database.
|
|
166
|
+
2. Then deploy `_underscore` + `worker2` + `tools`.
|
|
167
|
+
|
|
168
|
+
> **⚠ Order-critical.** Code before its migration crashes `loadWorkingSet` on the missing
|
|
169
|
+
> `status` / baseline columns (a missing column in a working-set SELECT is a **runtime exception**
|
|
170
|
+
> in 2.0, not a null — see the escalation-cron doc).
|
|
171
|
+
|
|
172
|
+
## Gotchas / known issues
|
|
173
|
+
|
|
174
|
+
- **Only the frequency-decay cron may write `Issue.status = RESOLVED`.** The ClickUp webhook
|
|
175
|
+
(`ErrorTask.php`) closes episodes and resets counters but must never touch `Issue.status`.
|
|
176
|
+
- **`autoResolveIssue` must stay DB-first.** Closing the episode and committing `RESOLVED` before
|
|
177
|
+
the ClickUp PUT is what makes the webhook echo a 0-row no-op; reversing the order reintroduces a
|
|
178
|
+
double-close race.
|
|
179
|
+
- **A RESOLVED Issue with no new events must short-circuit the pipeline.** (Bug caught in review.)
|
|
180
|
+
For **floored** Issues — whose urgency never reaches LOW — a RESOLVED Issue stayed in the working
|
|
181
|
+
set and **spuriously created a new ClickUp ticket**. Fixed with an explicit **RESOLVED
|
|
182
|
+
early-return guard** plus dropping `urgency = LOW` / `occurrences = 0` on resolve.
|
|
183
|
+
- **Reopen must re-seed the baseline anchor.** (Bug caught in review.) Without re-seeding
|
|
184
|
+
`dtBaselineAnchor` / `baselineAnchorOccurrences` on reopen, the **entire quiet period** was folded
|
|
185
|
+
into the EWMA, inflating the quiet window roughly **1.6× per resolve→reopen cycle**. Always
|
|
186
|
+
re-anchor on reopen.
|
|
187
|
+
- **`baselineSampleCount < 3` → not eligible to resolve.** An Issue without a trusted baseline is
|
|
188
|
+
never auto-resolved, so a brand-new Issue cannot resolve itself out of existence on the strength
|
|
189
|
+
of one quiet hour.
|
|
190
|
+
|
|
191
|
+
## Change history
|
|
192
|
+
|
|
193
|
+
- 2026-08-05 — **Built:** Issue-level auto-resolution and reopen by frequency decay. New
|
|
194
|
+
`Issue.status enum('OPEN','RESOLVED')` + `dtAutoResolved`, a durable cron-maintained firing
|
|
195
|
+
baseline (`baselineGapSeconds` EWMA + `baselineSampleCount`/`dtBaselineAnchor`/
|
|
196
|
+
`baselineAnchorOccurrences`, trusted at sample count ≥ 3), and a quiet-window
|
|
197
|
+
`clamp(3 × baselineGapSeconds, 24h, 90d)`. Two resolution paths — per-issue `autoResolveIssue`
|
|
198
|
+
(DB-first, then ClickUp `complete`) for ticketed Issues and a set-based `sweepResolveQuietIssues`
|
|
199
|
+
for ticketless ones (closing the working-set coverage gap) — plus `sweepReopenResolvedIssues` /
|
|
200
|
+
`reopenIssue` that flip status back to OPEN and re-anchor the baseline. Resolution keys only on a
|
|
201
|
+
trusted baseline + quiet duration, never on urgency; no team email on resolve; the ClickUp webhook
|
|
202
|
+
never sets `Issue.status`. Tools `/errors` gained an Open/Resolved/All filter, a Resolved badge,
|
|
203
|
+
a `dtAutoResolved` row, and the episode column renamed "Resolved" → "Closed". Two review bugs
|
|
204
|
+
fixed: a RESOLVED-with-no-events fall-through that spuriously re-ticketed floored issues, and a
|
|
205
|
+
reopen that folded the quiet period into the EWMA (~1.6× window inflation) until the anchor was
|
|
206
|
+
re-seeded. Migration `dbchanges2/Logs/2026-08-05a` runs before the `_underscore`+worker2+tools
|
|
207
|
+
deploy. (jcardinal)
|
|
208
|
+
</content>
|
|
209
|
+
</invoke>
|
package/knowledge/INDEX.md
CHANGED
|
@@ -19,7 +19,7 @@ _Auto-generated by `knowledge.js index`. Do not hand-edit._
|
|
|
19
19
|
## 2.0 framework
|
|
20
20
|
|
|
21
21
|
- **_underscore** (_Underscore) _(framework core)_ — 46 doc(s) → [2.0/apps/_underscore/INDEX.md](2.0/apps/_underscore/INDEX.md)
|
|
22
|
-
- **worker2** (Worker) —
|
|
22
|
+
- **worker2** (Worker) — 41 doc(s) → [2.0/apps/worker2/INDEX.md](2.0/apps/worker2/INDEX.md)
|
|
23
23
|
- **api2** (API) — 22 doc(s) → [2.0/apps/api2/INDEX.md](2.0/apps/api2/INDEX.md)
|
|
24
24
|
- **dbchanges2** (Database Changes) _(framework core)_ — 5 doc(s) → [2.0/apps/dbchanges2/INDEX.md](2.0/apps/dbchanges2/INDEX.md)
|
|
25
25
|
- **toga2-supply** (TOGa Supply) — 5 doc(s) → [2.0/apps/toga2-supply/INDEX.md](2.0/apps/toga2-supply/INDEX.md)
|
package/package.json
CHANGED