toga-ai 1.0.464 → 1.0.465

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -239,6 +239,18 @@ they are the known sharp edges. Do not re-discover these from scratch.
239
239
  [worker2 does](../worker2/features/alb-target-group-auto-registration.md). The same script's
240
240
  IMDSv2 handling (token on the `curl` command line, silent IMDSv1 fallback) must be fixed with it.
241
241
  See the Security note above.
242
+ 10. **Cross-client page-number paging is O(page), and the deepen chunk size is an untuned
243
+ performance dial.** Serving a deep page of a
244
+ [cross-client listing](features/cross-client-data-retrieval.md) requires materializing every
245
+ preceding row into the Cache cluster — page 4430 means 110,747 rows. Adaptive chunking
246
+ (`DEEPEN_CHUNK_MAX_RECORDS = 500`) makes it cheaper, not cheap. The durable fix is exposing
247
+ **keyset cursors to the caller** so cost is independent of depth; V2 already supports
248
+ `seekValues` + `seekUuid` correctly, so the mechanism exists but the caller-facing contract is
249
+ **unbuilt and not yet scoped**. Related dial: chunk size trades round trips against
250
+ sub-response size — a narrowed `fields` request stays small at any chunk, but an unrestricted
251
+ listing expands every row through `getFullModelData()` at the requested depth, which is the case
252
+ where 500 may need lowering. Watch it against `CURL_TIMEOUT_SECONDS = 30`. Do not raise the
253
+ chunk constant without checking sub-response size for wide, deep listings.
242
254
 
243
255
  ## When making changes here
244
256
 
@@ -254,8 +266,13 @@ they are the known sharp edges. Do not re-discover these from scratch.
254
266
  the duplicate check.
255
267
  - Preserve the commit-logs / rollback-data-on-failure invariant when editing the controller
256
268
  or `execute()`.
269
+ - **A code path with exactly one caller gets zero incidental coverage.** V2's keyset-pagination
270
+ branch emitted syntactically invalid SQL for an entire development cycle because only the
271
+ cross-client deepen path ever sets `$isKeysetMode`. In a ~2,000-line untested monolith, treat
272
+ single-caller branches in `V2.php` as unverified until exercised directly.
257
273
 
258
274
  ## Change history
275
+ - 2026-07-28 — Added Known issue #10: cross-client page-number paging is inherently O(page) (deep pages materialize every preceding row into the Cache cluster), with caller-facing keyset cursors identified as the durable fix but left unbuilt/unscoped, plus the `DEEPEN_CHUNK_MAX_RECORDS = 500` tuning tradeoff against `getFullModelData()` expansion and `CURL_TIMEOUT_SECONDS = 30`. Added a change-guidance bullet that single-caller branches in the untested `V2.php` monolith get zero incidental coverage (the keyset `LIMIT` syntax error shipped invisibly for a full cycle). (jcardinal)
259
276
  - 2026-07-28 — Documented the previously unrecorded `060_register_instance_to_shared_application_load_balancer` postdeploy hook pair in the Deployment section (non-prod self-registration into the same-named ALB target group; production skipped), and recorded a **committed IAM access key** in `ebs/register_instance_to_shared_application_load_balancer.php` as a security note + Known issue #9 — location, line range, commit subject, and remediation only (rotate, audit CloudTrail, move to the instance profile; history rewrite is a separate sign-off). Flagged api2's copy as the unhardened original vs. the new worker2 reference implementation. (jcardinal)
260
277
  - 2026-07-28 — Added a consolidated **Known issues / accepted risks** section (8 items), absorbing the previously free-floating deferred raw-exception-disclosure follow-up as item 1, so the tier's sharp edges (unrotated committed secrets, pre-execute phase still outside the main guard, local Logs DB name mismatch, permissive CORS, unpinned `_underscore` build clone, untested `V2.php` monolith, JWT rotation overlap window) are in one place instead of scattered. Recorded that `DB_CACHE` is resolved by name (`Databases.name = 'Cache'`), never by a hardcoded id, which differs per Core instance. (jcardinal)
261
278
  - 2026-07-28 — Added gotcha: the request-logger's auto-generated `Api.transactionId` (millisecond timestamp `Y-m-d H:i:s.v`, UNIQUE) collides under concurrent same-millisecond nested writes → MySQL 1062 → HTTP 500; platform-wide, observed on the Compass/Veyer ASN feed (`sourceIp 34.232.23.158`). Distinct from the client-supplied `transactionId`/EV-5 uniqueness contract. Fix direction: uuid the logged id or retry-on-1062. (bala)
@@ -42,10 +42,11 @@ is a **scatter-gather** engine: it fans out to each entitled client's own V2 API
42
42
  streams with a watermark, and caches the merged result per-user/per-query on a dedicated Cache
43
43
  cluster.
44
44
 
45
- > **Status: working end to end on local dev** (multi-client fan-out, ACL-filtered `client` object,
46
- > typed multi-column sort, keyset paging, narrowed `fields`, per-query cache isolation).
47
- > **Not yet deployed** — next stop is external testing. See
48
- > [api2 architecture](../architecture.md) for the accepted risks carried into that test.
45
+ > **Status: working end to end on local dev AND on a real deployed environment** — first
46
+ > deployed test 2026-07-28 on the **dev sandbox** (`api.beta.togahub.com`, dev-sandbox cluster).
47
+ > That test exposed a page-2-and-beyond failure whose root cause and four follow-on defects are
48
+ > now fixed and verified (see *Deep paging*, *Gotchas* and the change history). See
49
+ > [api2 architecture](../architecture.md) for the accepted risks carried alongside it.
49
50
 
50
51
  ## How it works
51
52
 
@@ -107,6 +108,73 @@ deepening and returned empty pages.
107
108
  the watermark cannot make any row safe, so querying it is a wasted round trip. At 1000 clients that
108
109
  is ~1 sub-request per pass instead of 1000 — the single biggest efficiency property of the design.
109
110
 
111
+ **Worked example (the canonical shape to reason about).** Sorting `-dateOrder` DESC across Compass
112
+ + NYCHH: Compass has 170 rows sharing `dateOrder = 2026-07-27` while NYCHH's *newest* row is
113
+ `2026-06-18`. Compass therefore holds the watermark for many consecutive pages, and NYCHH must
114
+ **not** be re-queried until Compass's cursor passes `2026-06-18`. Verified against live data: after
115
+ the 25-row page-1 cursor, 145 of the 170 same-date rows remain (25 + 145 = 170), i.e. the
116
+ multi-column seek predicate is correct across a large tie block.
117
+
118
+ ### The merge never serves past the watermark (`isDepthUnproven`)
119
+
120
+ `servePage()` used to do a blind `OFFSET`/`LIMIT` over the result table. When deepening **could not**
121
+ cover the requested page — a client failed, a cursor stalled, or `MAX_DEEPEN_ITERATIONS` ran out —
122
+ it still served a full page of rows that a lagging client may legitimately displace, i.e. it
123
+ presented a mis-ordered page as authoritative. One failed sub-request was enough to return 25 wrong
124
+ rows.
125
+
126
+ The page is now **clamped to `Tables.safeRecordCount`**, and any shortfall raises the new public
127
+ `$isDepthUnproven`, which V2 turns into a **WARNING** message telling the caller to retry to
128
+ continue deepening. Two new meta fields: `meta.crossClient.safeRecordCount` and
129
+ `meta.crossClient.isDepthUnproven`. `needsDeepening()` and the flag are deduped onto one new
130
+ `hasUnexhaustedClients()` helper.
131
+
132
+ **Subtlety — the flag is only raised while `hasUnexhaustedClients()` is true.** Once every client is
133
+ exhausted, `safeRecordCount` *is* the true end of the result set, so a short final page is correct
134
+ and must **not** warn.
135
+
136
+ ### Adaptive deepen chunking (sub-request page size ≠ caller's page size)
137
+
138
+ Because only the watermark client is deepened per pass, and each pass used to fetch the **caller's**
139
+ `recordsPerPage`, reaching page N cost ~N sequential sub-requests — and `MAX_DEEPEN_ITERATIONS = 50`
140
+ capped one HTTP request at ~1250 new rows. Requesting page 254 returned an *empty* page plus the
141
+ new warning.
142
+
143
+ The internal fetch size is now decoupled from the caller's:
144
+
145
+ ```
146
+ chunkSize = min(DEEPEN_CHUNK_MAX_RECORDS /* 500 */,
147
+ max(callerRecordsPerPage, target - safeRowCount()))
148
+ ```
149
+
150
+ Page 1 of 25 still fetches exactly 25; page 254 fetches 500 per pass — ~13 round trips instead of
151
+ 254. `fanOut()` takes `chunkSize` instead of `recordsPerPage`.
152
+
153
+ The caller's `recordsPerPage` still governs `servePage()`'s `LIMIT`, `needsDeepening()`'s target,
154
+ `reportedCounts()`, and the **cache identity** (`Tables.recordsPerPage` + query hash). Only the
155
+ internal fetch size changed — **no schema change**.
156
+
157
+ **Chunk-size tuning.** Chunk size trades round trips against sub-response size. A narrowed `fields`
158
+ request stays small at any chunk; an *unrestricted* listing expands every row through
159
+ `getFullModelData()` at the requested depth, and that is the case where 500 may need lowering.
160
+ Watch it against `CURL_TIMEOUT_SECONDS = 30`.
161
+
162
+ ### Only a cursor-less sub-response may set a client's total
163
+
164
+ Once the seek predicate is in the `WHERE` clause, V2's found-rows count is the **remainder after the
165
+ cursor**, not the client's total. `collectRows()` therefore tracks an `$isFirstFetch` map and only a
166
+ usable, cursor-less (offset-mode) sub-response establishes `totalRecordCount`; later deepen passes
167
+ leave the recorded total standing. Without this guard, fixing the keyset `LIMIT` bug would have made
168
+ reported totals silently **shrink** on every deepen pass.
169
+
170
+ ### Diagnosing a fan-out failure from the client logs
171
+
172
+ In `Logs_<Client>.Api`, a sub-request that logged its `POST /v2/auth/encrypted-user-uuid` (201) but
173
+ has **no corresponding GET row** proves the data sub-request died **before** the logging step (a
174
+ fatal) — which is what distinguishes it from a clean 0-row response. That single observation
175
+ isolated the keyset `LIMIT` root cause. Cross-check the *other* client's log too: the **absence** of
176
+ any row there confirms watermark-only deepening queried only the intended client.
177
+
110
178
  ### Per-row `client` object (ACL-filtered)
111
179
 
112
180
  Every returned row carries a `client` object filtered by the caller's **CORE** roles
@@ -157,12 +225,45 @@ cron.
157
225
  in client queries — because the PHP-side watermark merge compares bytes. Any linguistic collation
158
226
  would order differently and the merge could certify rows complete that MySQL then re-orders.
159
227
  - **Never use the raw stored row count** to decide paging depth; only `safeRecordCount`.
228
+ - **Never serve rows past `safeRecordCount`.** A short page + `isDepthUnproven` + a WARNING is
229
+ correct; a full page of possibly-displaceable rows is not.
230
+ - **A failed client is UNKNOWN, never exhausted.** Only a client that provably ran out of rows may
231
+ be marked `exhausted` — the watermark math treats exhausted clients as having no remainder.
232
+ - **Only a cursor-less sub-response may set a client's `totalRecordCount`** — a seek-filtered
233
+ response counts the remainder, not the total.
234
+ - **Anything compared against a sub-request's page size must compare against `chunkSize`**, not the
235
+ caller's `recordsPerPage`, now that the two differ.
160
236
  - **Every outbound sub-request needs its own globally-unique `transactionId`.**
161
237
  - **Forward the caller's original query string verbatim** — parsed `$httpOptions` do not round-trip
162
238
  (see the url-decode gotcha below). `MAX_CONCURRENT_CLIENT_FETCHES = 25` bounds fan-out.
163
239
 
164
240
  ## Gotchas
165
241
 
242
+ - **V2's keyset `LIMIT` had no leading newline — every seek sub-request was a SQL syntax error.**
243
+ In `processRoutePairs()` (`V2.php`, keyset branch, ~line 4113) the keyset path did
244
+ `$sql .= ('LIMIT ' . $recordsPerPage)` while the OFFSET path had a leading newline. The `ORDER BY`
245
+ immediately above ends in a bare `ASC`, so the generated SQL read
246
+ `` CAST(`uuid` AS BINARY) ASCLIMIT 25 ``. It hid for a whole development cycle because **only the
247
+ cross-client deepen path ever sets `$isKeysetMode`** — page 1 uses OFFSET and nothing else in api2
248
+ exercises keyset mode. This was the root cause of the entire "page 2+ returns wrong results" report.
249
+ *Lesson: a code path with exactly one caller gets zero incidental coverage.*
250
+ - **`emptyClientResult()` returning `exhausted => true` turned a transient error into a permanently
251
+ wrong cached answer.** `safeRowCount()` skips exhausted clients when computing the watermark, so a
252
+ failed client dropped out of the watermark → every already-cached row looked provably safe →
253
+ `safeRecordCount` was committed too high → `needsDeepening()` returned false for the **life of the
254
+ cache entry**, permanently excluding that client from every future deepen pass. It is now `false`
255
+ (unknown remainder). No runaway risk: the existing `$newRowCount === 0` stall guard still breaks
256
+ the loop.
257
+ - **`persistCursors()` erased the reported totals of every client it did not fetch this pass.** It
258
+ `DELETE`s all `Tables_Clients` rows then re-`INSERT`s with `$counts[$clientId] ?? null`, but a
259
+ deepen pass only populates `$counts` for the single watermark-holding client. Every other client's
260
+ `totalRecordCount`/`prohibitedRecordCount` became NULL → `meta.totalRecordCount` and
261
+ `totalPageCount` went to 0 → and because `nextPage` derives from `totalPageCount`, `nextPage`
262
+ became `null`, dead-ending any caller that pages by following `nextPage`. Fix: `SELECT` the
263
+ existing counts before the `DELETE` and carry them forward **per field**.
264
+ - **The exhaustion test had to move to `chunkSize` in the same commit as adaptive chunking.**
265
+ Comparing a full 500-row block against a caller page size of 25 declares the client "exhausted" and
266
+ silently reintroduces the `emptyClientResult` class of bug above.
166
267
  - **V2 NEVER url-decodes its query string.** It splits `QUERY_STRING` on `&` and `=` and consumes
167
268
  the raw values. Anything containing `,`, `[`, `"`, `%`, `+` or a space arrives mangled — this bit
168
269
  the feature twice (the JSON `seekValues` array, and a comma-separated `fields` list corrupted into
@@ -200,7 +301,35 @@ cron.
200
301
  - **`curl_multi` busy-spin guard** — both multi loops `usleep(100)` when
201
302
  `curl_multi_select() === -1`.
202
303
 
304
+ ## Known limitation — deep page-number paging is inherently O(page)
305
+
306
+ Serving page 4430 means materializing all 110,747 preceding rows into the cache. Adaptive chunking
307
+ makes that cheaper, not cheap: cost still grows with depth, because a page **number** cannot be
308
+ resolved without knowing everything before it.
309
+
310
+ The durable answer is **exposing keyset cursors to the caller** so cost is independent of depth — the
311
+ mechanism already exists (V2 now supports `seekValues` + `seekUuid` correctly). This was raised with
312
+ the developer on 2026-07-28 and **deliberately left unbuilt / not yet scoped**: record it as a known
313
+ limitation and a candidate direction, not as pending work.
314
+
203
315
  ## Change history
316
+ - 2026-07-28 — **First test on a real deployed environment** (dev sandbox,
317
+ `api.beta.togahub.com`), which reported wrong results on page 2+. Root cause: V2's keyset `LIMIT`
318
+ was concatenated without a leading newline, so **every** seek sub-request was a SQL syntax error —
319
+ invisible until now because only the cross-client deepen path sets `$isKeysetMode`. Fixed, plus the
320
+ four defects it masked: `emptyClientResult()` marking a *failed* client `exhausted` (which
321
+ permanently over-committed `safeRecordCount` and excluded that client from all future deepening);
322
+ `persistCursors()` NULL-ing the totals of every client not fetched in the pass (killing
323
+ `totalRecordCount`/`totalPageCount`/`nextPage`); totals being taken from cursor-filtered
324
+ sub-responses (now only a cursor-less first fetch sets a client's total); and `servePage()` serving
325
+ blind `OFFSET`/`LIMIT` past the watermark (now clamped to `safeRecordCount`, with a new public
326
+ `$isDepthUnproven` → V2 WARNING, `hasUnexhaustedClients()` helper, and new
327
+ `meta.crossClient.safeRecordCount` / `isDepthUnproven`). Added **adaptive deepen chunking**
328
+ (`DEEPEN_CHUNK_MAX_RECORDS = 500`; per-pass `chunkSize`, `fanOut()` and the exhaustion test both
329
+ switched to it) so deep pages cost ~page/500 round trips instead of ~page — no schema change, cache
330
+ identity unchanged. Recorded the Compass/NYCHH tie-block worked example, the log-based fan-out
331
+ failure diagnostic, and deep page-number paging being O(page) as a known limitation (caller-facing
332
+ keyset cursors left unbuilt). Verified working on dev sandbox. (jcardinal)
204
333
  - 2026-07-28 — Took the feature from never-executed to working end to end. Row identity moved from
205
334
  the numeric id to `rowUuid`/`seekUuid` (a V2 LIST never returns `id`); per-query
206
335
  `TableResults_<id>` tables with typed, direction-indexed sort columns; multi-column `seekValues`
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "toga-ai",
3
- "version": "1.0.464",
3
+ "version": "1.0.465",
4
4
  "description": "TOGA Technology Team Claude Knowledge System — shared AI coding harness with skills, knowledge base CLI, and project installer for Claude Code.",
5
5
  "keywords": [
6
6
  "claude",