hikoutei 0.4.2 → 0.4.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (140) hide show
  1. package/dist/application/sync/outbound/SheetsEffectDispatcher.d.ts +64 -0
  2. package/dist/application/sync/outbound/SheetsEffectDispatcher.d.ts.map +1 -0
  3. package/dist/application/sync/outbound/SheetsEffectDispatcher.js +440 -0
  4. package/dist/application/sync/outbound/SheetsEffectDispatcher.js.map +1 -0
  5. package/dist/application/sync/service/SyncServiceBootstrap.d.ts +3 -4
  6. package/dist/application/sync/service/SyncServiceBootstrap.d.ts.map +1 -1
  7. package/dist/application/sync/service/SyncServiceBootstrap.js +7 -4
  8. package/dist/application/sync/service/SyncServiceBootstrap.js.map +1 -1
  9. package/dist/application/sync/telemetry/syncTiming.d.ts +20 -33
  10. package/dist/application/sync/telemetry/syncTiming.d.ts.map +1 -1
  11. package/dist/application/sync/telemetry/syncTiming.js +14 -12
  12. package/dist/application/sync/telemetry/syncTiming.js.map +1 -1
  13. package/dist/infrastructure/storage/errors.d.ts +10 -5
  14. package/dist/infrastructure/storage/errors.d.ts.map +1 -1
  15. package/dist/infrastructure/storage/errors.js +10 -7
  16. package/dist/infrastructure/storage/errors.js.map +1 -1
  17. package/dist/infrastructure/storage/index.d.ts +6 -4
  18. package/dist/infrastructure/storage/index.d.ts.map +1 -1
  19. package/dist/infrastructure/storage/index.js +3 -2
  20. package/dist/infrastructure/storage/index.js.map +1 -1
  21. package/dist/infrastructure/storage/sqlite/schema.d.ts +12 -11
  22. package/dist/infrastructure/storage/sqlite/schema.d.ts.map +1 -1
  23. package/dist/infrastructure/storage/sqlite/schema.js +17 -109
  24. package/dist/infrastructure/storage/sqlite/schema.js.map +1 -1
  25. package/dist/infrastructure/storage/state/canonical/canonicalCommit.d.ts +2 -2
  26. package/dist/infrastructure/storage/state/canonical/canonicalCommit.d.ts.map +1 -1
  27. package/dist/infrastructure/storage/state/canonical/canonicalCommit.js +2 -2
  28. package/dist/infrastructure/storage/state/canonical/canonicalCommit.js.map +1 -1
  29. package/dist/infrastructure/storage/state/mapped/mappedPersistenceContext.d.ts +1 -2
  30. package/dist/infrastructure/storage/state/mapped/mappedPersistenceContext.d.ts.map +1 -1
  31. package/dist/infrastructure/storage/state/mapped/mappedPersistenceContext.js +1 -2
  32. package/dist/infrastructure/storage/state/mapped/mappedPersistenceContext.js.map +1 -1
  33. package/dist/infrastructure/storage/state/observation/observationCanonical.d.ts +1 -1
  34. package/dist/infrastructure/storage/state/observation/observationCanonical.d.ts.map +1 -1
  35. package/dist/infrastructure/storage/state/observation/observationQuarantine.d.ts +1 -1
  36. package/dist/infrastructure/storage/state/observation/observationQuarantine.d.ts.map +1 -1
  37. package/dist/infrastructure/storage/state/observation/observationQuarantine.js +1 -1
  38. package/dist/infrastructure/storage/state/observation/observationQuarantine.js.map +1 -1
  39. package/dist/infrastructure/storage/state/observation/observationTypes.d.ts +1 -1
  40. package/dist/infrastructure/storage/state/observation/observationTypes.d.ts.map +1 -1
  41. package/dist/infrastructure/storage/state/observation/observationValidation.d.ts +1 -1
  42. package/dist/infrastructure/storage/state/observation/observationValidation.d.ts.map +1 -1
  43. package/dist/infrastructure/storage/state/observation/observationWriter.d.ts +1 -1
  44. package/dist/infrastructure/storage/state/observation/observationWriter.d.ts.map +1 -1
  45. package/dist/infrastructure/storage/state/observation/observationWriter.js +1 -2
  46. package/dist/infrastructure/storage/state/observation/observationWriter.js.map +1 -1
  47. package/dist/infrastructure/storage/state/resolution/resolutionWriter.d.ts +1 -2
  48. package/dist/infrastructure/storage/state/resolution/resolutionWriter.d.ts.map +1 -1
  49. package/dist/infrastructure/storage/state/resolution/resolutionWriter.js +1 -2
  50. package/dist/infrastructure/storage/state/resolution/resolutionWriter.js.map +1 -1
  51. package/dist/infrastructure/storage/state/resolution/resolutionWriterContracts.d.ts +1 -1
  52. package/dist/infrastructure/storage/state/resolution/resolutionWriterContracts.d.ts.map +1 -1
  53. package/dist/infrastructure/storage/state/resolution/resolutionWriterHelpers.d.ts +1 -1
  54. package/dist/infrastructure/storage/state/resolution/resolutionWriterHelpers.d.ts.map +1 -1
  55. package/dist/infrastructure/storage/state/resolution/resolutionWriterHelpers.js +1 -1
  56. package/dist/infrastructure/storage/state/resolution/resolutionWriterHelpers.js.map +1 -1
  57. package/dist/infrastructure/storage/state/resolution/resolutionWriterSql.d.ts +1 -1
  58. package/dist/infrastructure/storage/state/resolution/resolutionWriterSql.d.ts.map +1 -1
  59. package/dist/infrastructure/storage/state/resolution/resolutionWriterSql.js +2 -2
  60. package/dist/infrastructure/storage/state/resolution/resolutionWriterSql.js.map +1 -1
  61. package/dist/infrastructure/storage/sync/shared/spreadsheetAuthority.d.ts +1 -1
  62. package/dist/infrastructure/storage/sync/shared/spreadsheetAuthority.d.ts.map +1 -1
  63. package/dist/infrastructure/storage/sync/shared/spreadsheetAuthority.js +1 -1
  64. package/dist/infrastructure/storage/sync/shared/spreadsheetAuthority.js.map +1 -1
  65. package/dist/infrastructure/storage/sync/shared/syncRegistry.d.ts +1 -1
  66. package/dist/infrastructure/storage/sync/shared/syncRegistry.js +1 -1
  67. package/package.json +13 -7
  68. package/dist/application/sync/outbound/effects/AdaptiveEffectBatchController.d.ts +0 -66
  69. package/dist/application/sync/outbound/effects/AdaptiveEffectBatchController.d.ts.map +0 -1
  70. package/dist/application/sync/outbound/effects/AdaptiveEffectBatchController.js +0 -123
  71. package/dist/application/sync/outbound/effects/AdaptiveEffectBatchController.js.map +0 -1
  72. package/dist/application/sync/outbound/effects/SyncEffectSupervisor.d.ts +0 -111
  73. package/dist/application/sync/outbound/effects/SyncEffectSupervisor.d.ts.map +0 -1
  74. package/dist/application/sync/outbound/effects/SyncEffectSupervisor.js +0 -369
  75. package/dist/application/sync/outbound/effects/SyncEffectSupervisor.js.map +0 -1
  76. package/dist/application/sync/outbound/effects/SyncEffectWorker.d.ts +0 -127
  77. package/dist/application/sync/outbound/effects/SyncEffectWorker.d.ts.map +0 -1
  78. package/dist/application/sync/outbound/effects/SyncEffectWorker.js +0 -552
  79. package/dist/application/sync/outbound/effects/SyncEffectWorker.js.map +0 -1
  80. package/dist/application/sync/outbound/effects/SyncEffectWorkerConstants.d.ts +0 -85
  81. package/dist/application/sync/outbound/effects/SyncEffectWorkerConstants.d.ts.map +0 -1
  82. package/dist/application/sync/outbound/effects/SyncEffectWorkerConstants.js +0 -74
  83. package/dist/application/sync/outbound/effects/SyncEffectWorkerConstants.js.map +0 -1
  84. package/dist/application/sync/outbound/effects/SyncEffectWorkerDispatch.d.ts +0 -31
  85. package/dist/application/sync/outbound/effects/SyncEffectWorkerDispatch.d.ts.map +0 -1
  86. package/dist/application/sync/outbound/effects/SyncEffectWorkerDispatch.js +0 -222
  87. package/dist/application/sync/outbound/effects/SyncEffectWorkerDispatch.js.map +0 -1
  88. package/dist/application/sync/outbound/effects/SyncEffectWorkerHelpers.d.ts +0 -14
  89. package/dist/application/sync/outbound/effects/SyncEffectWorkerHelpers.d.ts.map +0 -1
  90. package/dist/application/sync/outbound/effects/SyncEffectWorkerHelpers.js +0 -25
  91. package/dist/application/sync/outbound/effects/SyncEffectWorkerHelpers.js.map +0 -1
  92. package/dist/application/sync/outbound/effects/SyncEffectWorkerRouting.d.ts +0 -61
  93. package/dist/application/sync/outbound/effects/SyncEffectWorkerRouting.d.ts.map +0 -1
  94. package/dist/application/sync/outbound/effects/SyncEffectWorkerRouting.js +0 -296
  95. package/dist/application/sync/outbound/effects/SyncEffectWorkerRouting.js.map +0 -1
  96. package/dist/application/sync/outbound/effects/SyncEffectWorkerTiming.d.ts +0 -17
  97. package/dist/application/sync/outbound/effects/SyncEffectWorkerTiming.d.ts.map +0 -1
  98. package/dist/application/sync/outbound/effects/SyncEffectWorkerTiming.js +0 -80
  99. package/dist/application/sync/outbound/effects/SyncEffectWorkerTiming.js.map +0 -1
  100. package/dist/application/sync/outbound/effects/SyncEffectWorkerTransitions.d.ts +0 -13
  101. package/dist/application/sync/outbound/effects/SyncEffectWorkerTransitions.d.ts.map +0 -1
  102. package/dist/application/sync/outbound/effects/SyncEffectWorkerTransitions.js +0 -248
  103. package/dist/application/sync/outbound/effects/SyncEffectWorkerTransitions.js.map +0 -1
  104. package/dist/infrastructure/storage/sync/outbound/effectOutbox.d.ts +0 -141
  105. package/dist/infrastructure/storage/sync/outbound/effectOutbox.d.ts.map +0 -1
  106. package/dist/infrastructure/storage/sync/outbound/effectOutbox.js +0 -318
  107. package/dist/infrastructure/storage/sync/outbound/effectOutbox.js.map +0 -1
  108. package/dist/infrastructure/storage/sync/outbound/effectOutboxContracts.d.ts +0 -143
  109. package/dist/infrastructure/storage/sync/outbound/effectOutboxContracts.d.ts.map +0 -1
  110. package/dist/infrastructure/storage/sync/outbound/effectOutboxContracts.js +0 -16
  111. package/dist/infrastructure/storage/sync/outbound/effectOutboxContracts.js.map +0 -1
  112. package/dist/infrastructure/storage/sync/outbound/effectOutboxSql.d.ts +0 -33
  113. package/dist/infrastructure/storage/sync/outbound/effectOutboxSql.d.ts.map +0 -1
  114. package/dist/infrastructure/storage/sync/outbound/effectOutboxSql.js +0 -259
  115. package/dist/infrastructure/storage/sync/outbound/effectOutboxSql.js.map +0 -1
  116. package/dist/infrastructure/storage/sync/outbound/effectOutboxSupport.d.ts +0 -27
  117. package/dist/infrastructure/storage/sync/outbound/effectOutboxSupport.d.ts.map +0 -1
  118. package/dist/infrastructure/storage/sync/outbound/effectOutboxSupport.js +0 -316
  119. package/dist/infrastructure/storage/sync/outbound/effectOutboxSupport.js.map +0 -1
  120. package/dist/infrastructure/storage/sync/shared/writerLease.d.ts +0 -76
  121. package/dist/infrastructure/storage/sync/shared/writerLease.d.ts.map +0 -1
  122. package/dist/infrastructure/storage/sync/shared/writerLease.js +0 -189
  123. package/dist/infrastructure/storage/sync/shared/writerLease.js.map +0 -1
  124. package/docs/advanced-sheets-gateway-concurrency-problem.md +0 -434
  125. package/docs/architecture.md +0 -224
  126. package/docs/ci.md +0 -293
  127. package/docs/code-guidelines.md +0 -248
  128. package/docs/development.md +0 -74
  129. package/docs/gateway-removal-inventory.md +0 -147
  130. package/docs/git-workflow.md +0 -224
  131. package/docs/google-sheets-sync-scaling-strategy.md +0 -459
  132. package/docs/mikro-orm-adapter-spike.md +0 -89
  133. package/docs/quick-start.md +0 -137
  134. package/docs/sql-layer-plan.md +0 -59
  135. package/docs/sync-bulk-write-benchmark.md +0 -2117
  136. package/docs/sync-observability.md +0 -100
  137. package/docs/task-queue-write-model.md +0 -640
  138. package/docs/typed-sheets-mvp-scope-2026-06-29.md +0 -490
  139. package/docs/typed-sheets-plan.md +0 -417
  140. package/docs/write-and-synchronization-flow.md +0 -179
@@ -1,2117 +0,0 @@
1
- # Sync bulk-write benchmark
2
-
3
- ## Black-box server and Locust workload
4
-
5
- The test-only server harness treats Hikoutei as a server-side library rather
6
- than calling the sync worker directly from a benchmark script:
7
-
8
- ```sh
9
- node --env-file=.env .local/hikoutei-load-server.mjs
10
- ```
11
-
12
- It binds to `127.0.0.1:8787` by default and reports a run-specific persistent
13
- SQLite path, System_State/User_Input tab names, and JSONL log path from
14
- `GET /health`. It does not delete the database or Sheet data on shutdown.
15
-
16
- Run the mixed workload from another terminal:
17
-
18
- ```sh
19
- locust -f .local/locustfile.py \
20
- --host http://127.0.0.1:8787 \
21
- --users 20 \
22
- --spawn-rate 5 \
23
- --run-time 5m \
24
- --headless
25
- ```
26
-
27
- The Locust user weights are read 40%, create 20%, application update 25%, and
28
- User_Input simulation 15%. The User_Input task calls the test-only
29
- `/__test/user-input` endpoint, which changes the real Sheet and leaves the
30
- normal polling service to apply the value to SQLite. The server records HTTP,
31
- flush, Gateway, polling, error, and over-30-second events in the JSONL log;
32
- the over-30-second observer never stops the server or workload. Inspect
33
- `GET /metrics` after Locust stops, then stop the server with `Ctrl-C`; the
34
- persistent run data remains for inspection.
35
-
36
- ## 2026-08-03 — black-box Locust run and configuration diagnosis
37
-
38
- - Server run: `load-1785732937019-23536468`
39
- - Log: `.local/hikoutei-load-load-1785732937019-23536468.jsonl`
40
- - Workload: `.local/locustfile.py` mixed CRUD/User_Input tasks
41
- - Final HTTP log: **1,508** requests; 1,201 successful responses and 307
42
- failures, including 264 creates, 582 reads, 375 updates, and 209 User_Input
43
- requests
44
- - Gateway log: **444** requests; 246 successes, 190 operation failures, and 8
45
- timeouts; 4 requests exceeded 30 seconds
46
- - Final local state snapshot: **264** SQLite entity rows, 332 applied effects,
47
- 7 blocked candidates, 601 pending effects, and 234 processing effects
48
-
49
- The test did create visible tabs, but the server read a different
50
- `TYPED_SHEETS_GATEWAY_SHEET_ID` from the local `.env` than the spreadsheet ID
51
- provided for the intended target. On the spreadsheet actually used by the
52
- server, the following tabs were visible: `TS_Load_load-1785732937019-23536468_System`
53
- (265 rows including the header) and `TS_Load_load-1785732937019-23536468_Input`
54
- (26 rows including the header). The apparent missing table was therefore a
55
- configuration-target mismatch, not a provisioning omission. The intended sheet
56
- ID must be placed in `.env` before the next run; the shared secret must also be
57
- valid for that deployed gateway/sheet configuration.
58
-
59
- The run was not a clean functional pass. Concurrent User_Input simulation and
60
- polling contended on the Apps Script observation lock, producing repeated
61
- `Could not acquire the sync observation gateway lock` errors; the resulting
62
- failed effects then blocked later ORM writes with `user_input projection is
63
- blocked: latest effect is failed`. The persistent SQLite file and actual test
64
- Sheet data were retained for diagnosis.
65
-
66
- ## 2026-08-03 — lock-refactor Locust run
67
-
68
- - Branch: `perf/adaptive-sync-performance`
69
- - Server run: `load-1785737483461-793bddd2`
70
- - Exact server command: `node --env-file=.env .local/hikoutei-load-server.mjs`
71
- - Exact Locust command: `/usr/local/bin/locust -f .local/locustfile.py --host http://127.0.0.1:8787`
72
- (user count and duration were controlled from the Locust UI)
73
- - Log: `.local/hikoutei-load-load-1785737483461-793bddd2.jsonl`
74
- - Database: `.local/hikoutei-load-load-1785737483461-793bddd2.sqlite`
75
- - Backend: local Node `v24.3.0`, persistent SQLite/MikroORM, deployed Apps Script
76
- gateway, and the real target Sheet
77
- - Scenario: mixed read/create/update/User_Input workload; HTTP activity ran from
78
- `2026-08-03T06:12:25Z` through `2026-08-03T06:17:44Z` (about 5m 19s)
79
- - Setup: one provisioning Gateway call, 4,013 ms, excluded from steady-state
80
- counts below
81
-
82
- | Metric | No-setup / steady-state result | Full recorded run |
83
- | --- | ---: | ---: |
84
- | HTTP requests | 4,883 workload requests | 4,907 through workload stop, including 21 health, 2 metrics, and 1 unknown request |
85
- | HTTP success/failure | 3,119 / 1,764 workload responses | 3,142 / 1,765 responses |
86
- | Gateway requests | 763 sync calls; 718 success / 45 failure | 764 including setup |
87
- | HTTP latency | p50 5.16 ms; p95 2,820.89 ms | max 60,004.28 ms |
88
- | Gateway latency | p50 2,312 ms; p95 30,314 ms | max 60,004 ms; 35 over 30 s |
89
-
90
- Workload response breakdown: create 505 successful and 522 failed; read 1,957
91
- successful; update 582 successful and 603 failed; User_Input 75 returned 202,
92
- 598 returned 404 before the row was projected, and 41 returned 500. Gateway
93
- failures during the workload were 22 invalid responses, 15 `method_not_allowed`
94
- responses, 6 timeouts, and 2 remote operation failures. No
95
- `Could not acquire the sync observation gateway lock` message occurred during
96
- the workload window. One such lock error appeared later while background
97
- polling continued after Locust stopped.
98
-
99
- The dominant workload failure was projection poisoning: 1,124 HTTP operation
100
- errors reported `user_input projection is blocked: latest effect is
101
- blocked_candidate`. The workload also recorded 7 polling errors for invalid
102
- observed projection evidence. A post-load metrics snapshot showed 505 entity
103
- rows, 656 applied outbox effects, 25 blocked candidates, 61 failed effects,
104
- 1,264 pending effects, and 177 processing effects; this snapshot includes
105
- background synchronization after the Locust workload stopped. The Locust UI
106
- process remained open without new workload requests, and the server remained
107
- alive for diagnosis. All Sheet, SQLite, and JSONL data were retained.
108
-
109
- ## 2026-08-03 — pending User_Input retry Locust run
110
-
111
- - Branch: `perf/adaptive-sync-performance`
112
- - Server run: `load-1785738645774-79e64130`
113
- - Exact server command: `node --env-file=.env .local/hikoutei-load-server.mjs`
114
- - Exact Locust command: `/usr/local/bin/locust -f .local/locustfile.py --host http://127.0.0.1:8787`
115
- (user count and duration were controlled from the Locust UI)
116
- - Log: `.local/hikoutei-load-load-1785738645774-79e64130.jsonl`
117
- - Database: `.local/hikoutei-load-load-1785738645774-79e64130.sqlite`
118
- - Backend: local Node `v24.3.0`, persistent SQLite/MikroORM, deployed Apps Script
119
- gateway, and the real target Sheet
120
- - Scenario: mixed read/create/update/User_Input workload with bounded User_Input
121
- retry/backoff; workload ran from `2026-08-03T06:31:37Z` through
122
- `2026-08-03T06:36:58Z` (about 5m 21s)
123
- - Setup: one provisioning call and pre-load worker warm-up excluded from the
124
- steady-state Gateway counts
125
-
126
- | Metric | No-setup / steady-state result | Locust result |
127
- | --- | ---: | ---: |
128
- | Requests | 7,754 CRUD/User_Input requests | 7,774 total requests |
129
- | Success/failure | 4,519 / 3,235 workload responses | fail ratio **41.61%** |
130
- | Gateway | 1,173 calls; 1,165 success / 8 failure | p50 2,132 ms; p95 6,624 ms |
131
- | HTTP latency | aggregate median 10 ms; p95 2,300 ms | max 52,409 ms |
132
- | Slow Gateway calls | 12 over 30 seconds | max 52,407 ms |
133
-
134
- Locust's stopped report had 3,116 successful reads, 1,590 creates with 1,386
135
- failures, 1,924 updates with 1,698 failures, and 1,124 User_Input attempts
136
- with 151 failures. Of the User_Input failures, 148 reached the bounded retry
137
- deadline and 3 were HTTP 500 responses. The harness therefore stopped counting
138
- most projection-lag 404 responses as immediate Locust failures, but those rows
139
- still did not become successful Sheet edits: the server log recorded 943 404s
140
- and 188 successful 202 edits.
141
-
142
- The main failure remained projection poisoning. The server recorded 3,084
143
- `user_input projection is blocked: latest effect is blocked_candidate` errors,
144
- matching the create/update failures. During the workload, the Gateway had 3
145
- remote operation failures, 4 invalid responses, and 1 `method_not_allowed`
146
- response; polling also recorded 2 observation-lock acquisition failures and 7
147
- invalid observed-evidence errors. The retry queue improved classification of
148
- asynchronous row lag, but it cannot repair a projection whose effect is already
149
- blocked. A post-load metrics snapshot showed 204 entity rows, 360 applied
150
- outbox effects, 25 blocked candidates, 2 failed effects, 397 pending effects,
151
- and 93 processing effects. The Locust state was stopped with zero users; the
152
- server and all generated data remain available for diagnosis.
153
-
154
- Compared with the previous lock-refactor run, the retry queue reduced User_Input
155
- failures reported by Locust from immediate row-not-found failures to 148
156
- retry-deadline failures, but the underlying blocked-candidate cascade remains
157
- the next bottleneck. The two observation-lock failures also confirm that the
158
- remaining contention is in the retained full-observation path, not the
159
- lock-free values-only read.
160
-
161
- ## 2026-08-03 — Sync_Conflicts and automatic system-wins Locust run
162
-
163
- - Branch: `perf/adaptive-sync-performance`
164
- - Server command: `node --env-file=.env .local/hikoutei-load-server.mjs`
165
- - Locust command: `/usr/local/bin/locust -f .local/locustfile.py --host http://127.0.0.1:8787` (user count and duration were controlled from the Locust UI)
166
- - Server run: `load-1785748381328-41f5654c`
167
- - Log: `.local/hikoutei-load-load-1785748381328-41f5654c.jsonl`
168
- - Database: `.local/hikoutei-load-load-1785748381328-41f5654c.sqlite`
169
- - Backend: local Node `v24.3.0`, persistent SQLite/MikroORM, deployed Apps Script gateway, and the real target Sheet
170
- - Scenario: mixed read/create/update/User_Input workload with mandatory
171
- `System_State`, `User_Input`, and `Sync_Conflicts` projections
172
- - Workload window: `2026-08-03T09:13:59.580Z`–`2026-08-03T09:19:17.434Z`
173
- (about 5m 18s); setup provisioning took 5,436 ms and is excluded below
174
- - Dataset/result: 456 SQLite entity rows and 127 captured conflicts
175
-
176
- | Metric | No-setup / steady-state result | Full server record |
177
- | --- | ---: | ---: |
178
- | Application HTTP requests | 2,648 workload requests | 2,669 including 21 health requests before the workload endpoint closed |
179
- | HTTP outcome | 2,068 2xx / 246 projection-lag 404 / 334 500 | 2,089 2xx / 246 404 / 334 500 |
180
- | Locust-effective failure ratio | **334 / 2,648 = 12.61%** | 404 responses were marked success by the bounded retry queue |
181
- | Gateway requests | 428 sync calls; 308 success / 120 failure | 429 including setup |
182
- | Gateway latency | p50 4,344 ms; p95 34,246 ms | max 60,011 ms; 92 over 30 s; 4 over 60 s |
183
-
184
- Application breakdown was 456 successful creates and 93 failed creates, 1,044
185
- successful reads, 526 successful updates and 130 failed updates, and 42
186
- successful User_Input changes, 246 projection-lag 404s, and 111 failed
187
- User_Input requests. The main errors were 223 requests blocked by a previously
188
- failed `user_input` effect, 90 invalid Gateway JSON responses, 17
189
- `Use a signed POST request` responses, and 4 Gateway timeouts. During the
190
- workload window there were 2 polling evidence errors and 1 effect-worker
191
- postcondition/claim mismatch. The 404s are not counted as Locust failures by
192
- `locustfile.py`; they remain bounded pending projection retries.
193
-
194
- The new conflict path itself behaved correctly in SQLite: all **127/127**
195
- `sync_conflict` rows ended `RESOLVED`, all **127/127**
196
- `acknowledge_system` commands ended `applied` with role `sync_operator`, and
197
- there were **0** active candidate pointers. No `latest effect is
198
- blocked_candidate` cascade appeared in the server errors; the remaining
199
- application failures were caused by failed Gateway/materialization effects.
200
- The durable outbox retained the remote work for recovery. At the post-load
201
- metrics snapshot (after background synchronization continued), the outbox was
202
- `applied=390`, `pending=1,267`, `processing=60`, `failed=375`,
203
- `blocked_candidate=4`, and `superseded=127` at the
204
- `2026-08-03T09:26:09Z` metrics snapshot. The 127 conflict-audit effects
205
- were `30 applied`, `77 pending`, and `20 failed`, so SQLite resolution was
206
- committed even though the remote `Sync_Conflicts` projection had not fully
207
- converged.
208
-
209
- This is a functional improvement over the preceding retry-queue run: the
210
- unresolved conflict/candidate cascade was eliminated and the observed failure
211
- ratio was lower, but it is **not a clean remote-delivery benchmark**. The
212
- remaining bottleneck is Apps Script instability/latency (`invalid JSON`,
213
- `method_not_allowed`, operation failures, and timeouts), which poisoned
214
- System_State/User_Input effects and left a large durable backlog. The server
215
- and all Sheet, SQLite, and JSONL artifacts were retained; the server continued
216
- running after Locust stopped, so later background metrics are explicitly not
217
- part of the steady-state workload numbers.
218
-
219
- A later background-only inspection at `2026-08-03T09:50:42Z` found that the
220
- server was still alive but the remote projection had not converged: SQLite now
221
- contained **333** resolved conflicts (up from 127 at the workload snapshot),
222
- with zero active candidate pointers. This increase represents the same
223
- User_Input edits being observed again while their canonical reconcile effects
224
- remain failed/pending; it is not evidence that SQLite left conflicts open, but
225
- it does show that a continuously running worker cannot drain a permanently
226
- unavailable remote projection. The outbox then stood at `applied=794`,
227
- `pending=887`, `processing=64`, `failed=553`, `blocked_candidate=4`, and
228
- `superseded=333`; the server was still handling a long Gateway append when the
229
- 20-second before/after check was taken. This follow-up is diagnostic and is not
230
- part of the steady-state workload result above.
231
-
232
- ## 2026-08-02 — performance branch baseline
233
-
234
- - Branch: `perf/adaptive-sync-performance`
235
- - Base: `origin/develop` at `63ebddc`, including the merged contract PRs #155 and #156
236
- - Environment: macOS arm64, Node `v24.3.0`, npm `11.4.2`
237
-
238
- ### Local implementation verification (synthetic, no live Sheets)
239
-
240
- These commands exercise the in-process SQLite authority plus fake/in-memory
241
- Apps Script gateways. They are correctness and contract checks, not latency
242
- measurements: no real spreadsheet, credentials, or Apps Script quota is
243
- involved, so timings are not comparable to the live baselines below.
244
-
245
- - `npm test` — 30 files, 197 tests passed
246
- - `npm run typecheck` — passed
247
- - `npm run typecheck:test` — passed
248
- - `npm run build` — passed
249
- - `npm pack --dry-run` — passed
250
- - `git diff --check` — passed
251
-
252
- The regression suite now covers the adaptive/outbound work directly with
253
- synthetic fixtures: the values-only preflight escalation rules, formula/merged/
254
- error deferral to the periodic safety full scan, safety-scan scheduling and
255
- coalescing, backlog convergence across worker passes, no full read on unchanged
256
- adaptive passes, response-loss backoff, and the new internal polling timing
257
- phases emitted through the diagnostics sink (canonical state read, values-only
258
- read, fast comparison, full metadata observation, persistence, overall, and safety-scan lag).
259
-
260
- ### 2026-08-02 — local synthetic latency sample
261
-
262
- - Branch: `perf/adaptive-sync-performance`
263
- - Exact command: `./node_modules/.bin/vite-node scripts/bench-local-sync-performance.ts`
264
- - Backend: in-process SQLite/MikroORM plus `FakeSyncSheetGateway`; no network or Apps Script quota
265
- - Dataset: 66 unchanged User_Input rows; 7 steady-state samples per mode
266
- - Setup: entity seeding, effect delivery, and first warm-up pass excluded
267
-
268
- | Scenario | No-setup / steady-state result | Gateway calls during run |
269
- | --- | ---: | --- |
270
- | Adaptive values-only polling | mean **1.604 ms** (min 1.271 / max 2.582) | 7 values-only reads, 0 full snapshots |
271
- | Full metadata polling | mean **5.642 ms** (min 5.228 / max 6.656) | 0 values-only reads, 7 full snapshots |
272
- | 66-row outbound update | SQLite flush **68.988 ms**; delivery **216.973 ms**; 66 applied | 8 `applyEffects` calls |
273
-
274
- The local adaptive sample is about **3.5× faster** than the local full path in
275
- this no-network fixture, but it is not comparable to the historical live
276
- 27.7-second versus 2.2-second measurements below. The useful correctness signal
277
- is that unchanged adaptive samples used no full metadata reads. The live
278
- measurement below confirms the same behavior through the deployed gateway.
279
-
280
- ### 2026-08-02 — live Apps Script polling smoke benchmark
281
-
282
- - Branch: `perf/adaptive-sync-performance`
283
- - Exact command: `LIVE_POLL_ONLY=1 node --env-file=.env .local/user-input-polling-live.mjs`
284
- - Slow-request guard: synchronization Gateway requests with `durationMs > 30,000` fail the benchmark; provisioning and cleanup requests are recorded but excluded from this failure gate
285
- - Backend: deployed Apps Script gateway and real Google Sheet
286
- - Dataset: 1 entity row, System_State + User_Input projections
287
- - Setup: temporary-tab provisioning and first safety scan are reported separately;
288
- steady-state polling excludes provisioning
289
-
290
- | Scenario | No-setup / steady-state result | Evidence |
291
- | --- | ---: | --- |
292
- | Initial unchanged adaptive poll | **2,504 ms** | `mode=adaptive`, `fullMetadataTables=0`, 1 values-only read, 0 full snapshots |
293
- | Remote edit request | **1,663.4 ms** | Apps Script `setValue()` + `flush()` |
294
- | Changed adaptive poll | **3,859 ms** | 1,580 ms values-only read + 2,262 ms full metadata read; SQLite status became `approved` |
295
- | First ORM insert flush | **12.39 ms** | SQLite/outbox commit |
296
- | Insert remote delivery | **5,185.8 ms** | 2 effects applied |
297
-
298
- The service's automatic initial safety scan took **2,567 ms** and used one full
299
- metadata request. The latency guard inspected **8 synchronization Gateway
300
- requests** and found **0** over the 30-second threshold. Temporary Sheet tabs
301
- were cleaned up after the run. This live run proves the current harness reaches
302
- the deployed gateway and that unchanged polling avoids the full metadata request
303
- while changed polling escalates and persists correctly. The JSON result includes
304
- `gatewayLatencyGuard` with the threshold and guarded-request summary. A
305
- functional failure (for example, the SQLite value is not `approved`) is reported
306
- separately from a `slow_sync_gateway_operation` latency failure.
307
-
308
- #### Post-guard verification rerun
309
-
310
- - Commands:
311
- - `LIVE_POLL_ONLY=1 node --env-file=.env .local/user-input-polling-live.mjs`
312
- - `LIVE_CRUD_ONLY=1 node --env-file=.env .local/user-input-polling-live.mjs`
313
- - `node --env-file=.env .local/user-input-polling-live.mjs`
314
- - Dataset: 1 row per isolated run; provisioning excluded from steady-state values
315
- - Functional result: all runs completed; User_Input status became `approved`; all
316
- CRUD deliveries applied 2 effects per operation; no conflicts
317
-
318
- | Run | No-setup / steady-state result | Guard result |
319
- | --- | --- | --- |
320
- | `LIVE_POLL_ONLY=1` | Initial poll **2,304 ms**; changed poll **5,345 ms**; insert delivery **6,194.66 ms** | 8 guarded requests, 0 slow |
321
- | `LIVE_CRUD_ONLY=1` | Initial poll **2,153 ms**; update delivery **8,515.16 ms**; delete delivery **6,372.73 ms** | 9 guarded requests, 0 slow |
322
- | Full harness | Initial poll **2,728 ms**; steady polls **3,141.38 / 2,489.48 / 2,227.51 ms**; update delivery **7,593.92 ms**; delete delivery **7,249.12 ms** | 15 guarded requests, 0 slow |
323
-
324
- #### Quota observation for this verification batch
325
-
326
- The three live runs above made **38 external Apps Script Web App POSTs** in total:
327
- 32 synchronization requests, 3 provisioning requests, and 3 cleanup requests. Every observed gateway request returned HTTP 200; no 403/429,
328
- `RESOURCE_EXHAUSTED`, or quota-related gateway error occurred. The 30-second
329
- gateway guard also found 0 slow synchronization requests.
330
-
331
- This is an observed sample, not a remaining-quota counter. The current gateway
332
- uses Apps Script `SpreadsheetApp`, not the Google Sheets REST API, so the Sheets
333
- REST request quota cannot be inferred from these POST counts. Apps Script
334
- service/runtime quotas are account- and project-dependent, and Google does not
335
- return their remaining amount through this harness. The exact remaining quota
336
- must be checked in the linked Google Cloud/API or Apps Script execution/usage
337
- console.
338
-
339
- ### 2026-08-02 — progressive live bidirectional verification
340
-
341
- - Branch: `perf/adaptive-sync-performance`
342
- - Exact command: `node --env-file=.env .local/progressive-bidirectional-live.mjs`
343
- - Backend: deployed Apps Script gateway and real Google Sheet
344
- - Matrix: cumulative **1 → 10 → 25 → 66 → 100 rows** in one isolated service
345
- - Each stage performed both directions: app flush/effect delivery, bulk User_Input
346
- edit, changed polling, SQLite verification, and one unchanged polling pass
347
- - Policy: 30-second Gateway requests were recorded but never stopped the matrix;
348
- all stages completed before the final result was reported
349
-
350
- | Rows | Added | App flush | Outbound delivery | Changed poll | Unchanged poll | Max Gateway request | Result |
351
- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | --- |
352
- | 1 | 1 | 12.25 ms | 6,304.94 ms | 4,003.09 ms | 1,768.76 ms | 3,510 ms | passed |
353
- | 10 | 9 | 13.18 ms | 13,100.62 ms | 10,130.85 ms | 2,356.88 ms | 8,121 ms | passed |
354
- | 25 | 15 | 16.71 ms | 16,375.68 ms | 7,866.09 ms | 2,037.32 ms | 8,462 ms | passed |
355
- | 66 | 41 | 34.96 ms | 78,406.70 ms | 16,210.92 ms | 1,845.36 ms | 19,087 ms | passed |
356
- | 100 | 34 | 23.89 ms | 120,235.71 ms | 27,529.99 ms | 1,881.89 ms | 26,694 ms | passed |
357
-
358
- All stages persisted the expected row count, changed polling applied every
359
- expected row with one full-metadata escalation and zero conflicts, and unchanged
360
- polling stayed on the values-only path. The matrix made **47 synchronization
361
- Gateway requests** (49 including one provisioning and one cleanup request),
362
- found **0** over 30 seconds, and had no functional or cleanup failure. The 66- and 100-row worker/stage totals exceeded 30 seconds because they
363
- contained multiple Gateway requests; no individual synchronization Gateway
364
- request exceeded the threshold.
365
-
366
- This is a one-row smoke benchmark, not a 66-row throughput result. Network
367
- latency, Apps Script execution variance, Sheet lock contention, and quota remain
368
- caveats. The progressive result above is still a single 100-row cumulative run,
369
- not a repeated production-load benchmark. Historical live comparison baselines
370
- from earlier branches remain:
371
-
372
- | Scenario | Setup | No-setup / steady-state result | Source |
373
- | --- | ---: | ---: | --- |
374
- | Full snapshot polling, 66 rows | excluded | 27,652 ms steady state | 2026-07-27 phase trace below |
375
- | Values-only polling, 66 rows | excluded | 2,240 ms steady state | 2026-07-27 lightweight polling below |
376
- | Outbound 20-order stage | excluded | 28,182 ms / 5.0 rows/s | 2026-07-27 timing run below |
377
- | Outbound 370-order stage | excluded | 36,865 ms / 30.1 rows/s | 2026-07-27 progressive run below |
378
-
379
- SQLite remains the application authority and Sheets the asynchronous projection:
380
- full metadata fidelity, response-loss recovery, and the periodic safety full scan
381
- stay required while the values-only fast path avoids remote full reads on
382
- unchanged passes.
383
-
384
- ## 2026-07-24 — raw Apps Script write
385
-
386
- - Branch: `benchmark/apps-script-bulk-write`
387
- - Script: `apps-script/gateway/BulkWriteBenchmark.gs`
388
- - Entry point: `runBulkWriteBenchmark100()`
389
- - Backend: Apps Script `SpreadsheetApp`
390
- - Target: `__typed_sheets_bulk_write_benchmark`
391
- - Dataset: 100 rows × 6 columns = 600 cells
392
- - Write mode: one contiguous `setValues()` followed by one `flush()`
393
- - Setup: the benchmark appended at row 22; one-time sheet setup was excluded
394
- from the measured write interval
395
-
396
- | Measurement | Setup | `setValues()` | `flush()` | Steady-state total |
397
- | --- | ---: | ---: | ---: | ---: |
398
- | 100 rows / 600 cells | 713 ms | 96 ms | 278 ms | 374 ms |
399
-
400
- Throughput was 267.38 rows/second and 1,604.28 cells/second.
401
-
402
- ## Comparison: full sync request
403
-
404
- The previous clean-DB test sent 20 customer effects through the complete
405
- Gateway path with `POST /sync/once?maxEffects=20`. It took 75,409 ms and all 20
406
- effects were applied. That path included snapshot reads, row metadata, CAS
407
- checks, postcondition reads, receipts, and request tracing.
408
-
409
- The raw write benchmark is therefore roughly 1,000 times faster per row than
410
- the complete 20-effect sync request. This is not a correctness comparison: the
411
- benchmark intentionally bypasses concurrency checks, user-edit detection,
412
- receipts, and postcondition verification.
413
-
414
- ## 2026-07-24 — isolated stage benchmark
415
-
416
- - Branch: `benchmark/apps-script-bulk-write`
417
- - Script: `apps-script/gateway/BulkWriteBenchmark.gs`
418
- - Entry point: `runBulkWriteStageBenchmark20()`
419
- - Backend: Apps Script `SpreadsheetApp`
420
- - Dataset: 20 rows × 6 columns
421
- - Setup time is excluded from each stage measurement.
422
-
423
- | Stage | Setup | Operation | Flush | Measured total |
424
- | --- | ---: | ---: | ---: | ---: |
425
- | `metadata_read` | 35,117 ms | 3,609 ms | — | 3,609 ms |
426
- | `metadata_write` | 33,273 ms | 31,637 ms | 59 ms | 31,696 ms |
427
- | `snapshot_read` | 19,540 ms | 4,303 ms | — | 4,303 ms |
428
- | `cas_compare` | 0 ms | 2 ms | — | 2 ms |
429
- | `postcondition_read` | 18,280 ms | 2,646 ms | — | 2,646 ms |
430
- | `receipt_read` | 1,200 ms | 119 ms | — | 119 ms |
431
- | `receipt_write` | 1,597 ms | 430 ms | 352 ms | 782 ms |
432
-
433
- The dominant measured stage is `metadata_write`: rewriting three row metadata
434
- entries for 20 rows took about 31.7 seconds, while the `flush()` itself took
435
- only 59 ms. This identifies the per-row Developer Metadata API calls as the
436
- primary bottleneck. The in-memory CAS comparison and receipt operations are
437
- not significant for this test size.
438
-
439
- ## 2026-07-24 — visible-state metadata migration
440
-
441
- The Gateway batch path now treats SQLite as the authority for
442
- `visibleRevision` (가시적 버전) and `visibleHash` (가시적 상태 해시):
443
-
444
- - Sheet Developer Metadata keeps only the projection-local row anchor.
445
- - Batch CAS compares the bulk-read cell hash with the effect's
446
- `expectedVisibleHash` from SQLite; it no longer reads or rewrites a
447
- revision/hash metadata pair for every changed row.
448
- - Successful results and receipts derive the next revision as
449
- `expectedVisibleRevision + 1`.
450
- - Batch postconditions verify the raw cell hash in one range read.
451
- - The receipt sheet remains because it is durable response-loss evidence, not
452
- per-row Developer Metadata.
453
- - Existing visible revision/hash metadata is not deleted; it is left inert so
454
- this migration does not perform a destructive cleanup of user sheets.
455
-
456
- This change is source-level and requires deploying the updated
457
- `apps-script/gateway/Code.gs` before a live benchmark. The next comparison
458
- should rerun `runBulkWriteStageBenchmark20()` and a 20-effect sync against a
459
- fresh test sheet, then verify that `metadata_write` no longer appears in the
460
- production batch path and that SQLite confirmations still advance normally.
461
-
462
- ## 2026-07-24 — new Gateway progressive load test
463
-
464
- - Branch: `benchmark/apps-script-bulk-write`
465
- - Harness: `.local/lib-test-beta-0.2.0-beta1`
466
- - Backend: MikroORM + SQLite locally, deployed Apps Script Gateway remotely
467
- - Database: fresh `typed-sheets-load-test-new-gateway.sqlite`
468
- - Setup: provisioned five projection tabs and seeded 20 customers plus 20
469
- products; the 40-effect seed sync took 46.7 seconds and applied all effects
470
- with no failures. Setup is excluded from the stage comparison below.
471
- - Gateway behavior: visible revision/hash row metadata disabled; anchor
472
- metadata and receipt sheet retained.
473
- - Stage runner: the local Node fetch runner created orders in groups of
474
- `1/5/10/20` and called `/sync/once?maxEffects=1/5/10/20` once after each
475
- group. Each order produced four effects in this scenario.
476
-
477
- | Orders added | `maxEffects` | Order writes | Sync request | Applied | Failed | Pending after stage |
478
- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
479
- | 1 | 1 | 6.03 s | 5.97 s | 1 | 0 | 3 |
480
- | 5 | 5 | 11.07 s | 11.03 s | 5 | 0 | 18 |
481
- | 10 | 10 | 16.22 s | 16.15 s | 10 | 0 | 48 |
482
- | 20 | 20 | 23.24 s | 23.10 s | 20 | 0 | 108 |
483
-
484
- The stage sync cost was approximately 1.15 seconds per effect at the 20-effect
485
- batch, compared with 5.97 seconds per effect at the single-effect batch. The
486
- cost is still substantial, but the 20-effect request is about 3.3× faster than
487
- the earlier 75.4-second 20-effect run that rewrote visible row metadata.
488
-
489
- The accumulated 108-effect backlog was then drained with
490
- `/sync/once?maxEffects=100`:
491
-
492
- | Drain pass | Duration | Selected | Applied | Deferred | Failed | Pending after pass |
493
- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
494
- | 1 | 59.85 s | 92 | 60 | 32 | 0 | 48 |
495
- | 2 | 69.49 s | 48 | 48 | 0 | 0 | 0 |
496
-
497
- Final SQLite counts were 36 orders, 36 order items, and 36 payments. The final
498
- outbox contained 184 applied effects and no failed effects. This confirms that
499
- the new Gateway path can drain this clean backlog, but it does not yet prove
500
- that arbitrary response loss or a manually edited Sheet will always converge.
501
-
502
- ## 2026-07-24 — write-batch read reduction
503
-
504
- - Branch: `benchmark/apps-script-bulk-write`
505
- - Source: `apps-script/gateway/Code.gs`
506
- - Scope: the `applyEffects` write path only; the standalone `readSnapshot` path
507
- remains unchanged
508
-
509
- The write path no longer uses the reader-oriented full snapshot at both sides
510
- of a batch:
511
-
512
- - `createSyncBatchContext_` now reads one raw-value range and row anchors using
513
- `readSyncBatchState_`; it does not read formulas, display values, or merged
514
- ranges, and it does not repeat the raw-value range read.
515
- - The postcondition check now reads only the changed rows, grouped into
516
- contiguous ranges, instead of reading every registered row from row 2 to the
517
- sheet's last row.
518
- - The final full snapshot read was removed from the write response. SQLite is
519
- authoritative for visible revision/hash state, while each changed effect is
520
- still verified from raw Sheet cells before its receipt is written. The
521
- optional `snapshotHash` in an apply response is therefore `null`.
522
- - Stable row anchors and the receipt sheet remain unchanged because they are
523
- still required for row identity and response-loss recovery.
524
-
525
- Static verification passed with `node --check`, the focused Apps Script source
526
- and sync-client tests (9 tests), `npm run typecheck`, and `npm run build`. A live
527
- comparison requires deploying this `Code.gs` first. The next benchmark should
528
- compare the `prepare_batch_context`, `batch_flush`, and postcondition phase
529
- durations against the progressive run above, especially on a sheet with many
530
- existing rows.
531
-
532
- ## 2026-07-24 — fresh-sheet progressive load test after a new deployment
533
-
534
- - Branch: `benchmark/apps-script-bulk-write`
535
- - Harness: `.local/lib-test-beta-0.2.0-beta1`
536
- - Backend: MikroORM + SQLite locally, newly deployed Apps Script Gateway
537
- - Database: fresh `data/typed-sheets-load-test-progressive-20260724.sqlite`
538
- - Dataset: 20 customers and 20 products seeded into a new spreadsheet; seed
539
- synchronization was 40 effects and is excluded from the progressive stage
540
- comparison
541
- - Setup commands: `POST /sync/provision`, `POST /load-test/seed` with
542
- `customers=20`, `products=20`, then `POST /sync/once?maxEffects=100`
543
- - Progressive script: create order groups of `1/5/10/20`; after each group,
544
- call `POST /sync/once?maxEffects=<group size>` and record `/metrics` before
545
- and after the call
546
- - Background worker: disabled; synchronization was explicitly invoked by the
547
- test runner
548
-
549
- The 40-effect seed took 49.87 seconds and applied all effects without failure.
550
-
551
- | Orders added | `maxEffects` | Order writes | Sync request | Applied | Failed | Pending after stage |
552
- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
553
- | 1 | 1 | 21.98 ms | 6.40 s | 1 | 0 | 3 |
554
- | 5 | 5 | 42.29 ms | 9.13 s | 5 | 0 | 18 |
555
- | 10 | 10 | 77.42 ms | 14.72 s | 10 | 0 | 48 |
556
- | 20 | 20 | 139.47 ms | 27.39 s | 20 | 0 | 108 |
557
-
558
- The first drain pass selected 92 effects, applied 60, deferred 32 because of
559
- the Gateway prefix limit, and took 56.62 seconds. The following 16-effect
560
- `load-payments-state` request exceeded the 120-second client timeout. Its
561
- response-loss recovery then issued one full `readSnapshot` request per unknown
562
- effect, with individual reads taking roughly 7–16 seconds. The observed state
563
- after that recovery was 162 applied, 16 failed, and 6 pending; every failed
564
- effect had `postcondition_unavailable`.
565
-
566
- A follow-up drain selected 21 effects (the 16 failed effects were eligible for
567
- retry plus 5 pending effects), took 184.39 seconds, applied only the 5 pending
568
- effects, and left the same 16 failures. Final metrics were:
569
-
570
- | SQLite table | Count |
571
- | --- | ---: |
572
- | customers | 20 |
573
- | products | 20 |
574
- | orders | 36 |
575
- | orderItems | 36 |
576
- | payments | 36 |
577
-
578
- The final outbox contained 167 applied, 16 failed, and 1 pending effect. This
579
- test confirms the dominant failure path is not raw `setValues()` throughput:
580
- an oversized or slow payment batch reaches the 120-second transport timeout,
581
- then failed-effect retry multiplies the cost by performing full snapshot reads
582
- per effect. The next fix should make postcondition recovery batch-aware and
583
- avoid retrying all failed effects in the normal drain pass without a bounded
584
- backoff or explicit recovery mode.
585
-
586
- ## 2026-07-24 — fresh-sheet progressive load test with batch postcondition recovery
587
-
588
- - Branch: `benchmark/apps-script-bulk-write`
589
- - Harness: `.local/lib-test-beta-0.2.0-beta1`
590
- - Backend: MikroORM + SQLite locally, newly deployed Apps Script Gateway
591
- - Database: fresh `data/typed-sheets-load-test-progressive-20260724-new-sheet.sqlite`
592
- - Dataset: 20 customers and 20 products seeded into a fresh spreadsheet; seed
593
- synchronization applied 40 effects without failure and is excluded from the
594
- progressive stage comparison
595
- - Background worker: disabled; synchronization was explicitly invoked with
596
- `POST /sync/once`
597
- - Recovery path: `readEffectPostconditions` batch API; no per-effect full
598
- `readSnapshot` recovery call
599
-
600
- | Orders added | `maxEffects` | Order writes | Sync request | Selected | Applied | Requeued | Failed | Pending after stage |
601
- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
602
- | 1 | 1 | 57.16 ms | 5.96 s | 1 | 1 | 0 | 0 | 3 |
603
- | 5 | 5 | 44.42 ms | 125.51 s | 5 | 0 | 5 | 0 | 23 |
604
- | 10 | 10 | 77.86 ms | 15.38 s | 10 | 10 | 0 | 0 | 53 |
605
- | 20 | 20 | 558.76 ms | 18.54 s | 20 | 20 | 0 | 0 | 113 |
606
-
607
- The five-effect stage exceeded the 120-second client timeout, but the new
608
- batch postcondition recovery classified all five effects as `unapplied` and
609
- returned them to `pending`; it did not create `postcondition_unavailable`
610
- failures or issue one full snapshot read per effect. The other stages applied
611
- normally.
612
-
613
- The backlog was drained in multiple passes because the Apps Script Gateway
614
- still limits each apply prefix to 20 effects, and product stock effects for the
615
- same target must respect predecessor ordering:
616
-
617
- | Drain pass | Selected | Applied | Deferred | Pending after pass |
618
- | ---: | ---: | ---: | ---: | ---: |
619
- | 1 | 78 | 46 | 32 | 67 |
620
- | 2 | 33 | 33 | 0 | 34 |
621
- | 3 | 1 | 1 | 0 | 33 |
622
- | Ordered-chain passes | 25 | 25 | 0 | 0 |
623
-
624
- Final metrics were 184 applied effects, zero failed effects, and zero pending
625
- effects. SQLite contained 36 orders, 36 order items, and 36 payments. This
626
- reproduces the variable Gateway latency, but confirms that batch postcondition
627
- recovery prevents the previous per-effect snapshot amplification and allows
628
- the backlog to converge to zero.
629
-
630
- ## 2026-07-25 — published beta progressive API load test
631
-
632
- - Branch: `refactor/thin-sync-gateway`
633
- - Harness: `.local/lib-test-beta-0.2.0-beta1`
634
- - Package: published `typed-sheets@0.2.0-beta-1` (not a local workspace link)
635
- - Backend: MikroORM + SQLite locally, deployed Apps Script Gateway remotely
636
- - Gateway configuration: the new untracked `.env` values supplied for this
637
- test; the secret is intentionally not recorded here
638
- - Background worker: disabled; the load runner invoked `/sync/once` itself
639
- - Existing state: the test DB already contained one customer, one product, and
640
- two applied seed effects; the spreadsheet was also reused, so this is not a
641
- clean-sheet comparison
642
- - Commands used:
643
-
644
- ```sh
645
- LOAD_TEST_DURATION_MS=5000 LOAD_TEST_CONCURRENCY=1 \
646
- LOAD_TEST_WRITE_RATIO=0.7 LOAD_TEST_SEED_CUSTOMERS=5 \
647
- LOAD_TEST_SEED_PRODUCTS=5 LOAD_TEST_SYNC=1 LOAD_TEST_DRAIN=0 \
648
- node --env-file-if-exists=.env load-test.mjs
649
-
650
- LOAD_TEST_DURATION_MS=10000 LOAD_TEST_CONCURRENCY=2 \
651
- LOAD_TEST_WRITE_RATIO=0.7 LOAD_TEST_SEED_CUSTOMERS=5 \
652
- LOAD_TEST_SEED_PRODUCTS=5 LOAD_TEST_SYNC=1 LOAD_TEST_DRAIN=0 \
653
- node --env-file-if-exists=.env load-test.mjs
654
- ```
655
-
656
- | Stage | Duration | Concurrency | Generated orders | Total requests | Requests/s | Sync calls recorded | Applied | Pending | Processing | Failed |
657
- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
658
- | Low | 5 s | 1 | 643 | 918 | 181.30 | 0 | 2 | 3,558 | 10 | 0 |
659
- | Medium | 10 s | 2 | 876 | 1,283 | 127.24 | 0 | 18 | 8,429 | 45 | 0 |
660
-
661
- The `/orders` route combines order writes and order-list reads, so its latency
662
- is not a write-only measurement. The low-stage mixed route latency was p50
663
- 6.18 ms / p95 9.74 ms / p99 13.93 ms; the medium-stage latency was p50
664
- 14.67 ms / p95 28.77 ms / p99 39.45 ms. All HTTP requests succeeded.
665
-
666
- The load runner did not receive a completed `maxEffects=50` sync response before
667
- each short stage ended, so those calls were aborted from the runner's point of
668
- view. A standalone `maxEffects=1` request after the low stage selected and
669
- applied one effect successfully in roughly 3.3 seconds. This gives a direct
670
- steady-state reference for the consumer path, excluding the load runner's
671
- health and seed setup.
672
-
673
- The medium stage already saturated the projection path: local SQLite produced
674
- orders much faster than the Gateway consumed effects, growing the backlog from
675
- 3,558 to 8,429 pending effects while failed effects remained at zero. The high
676
- stage was intentionally not run because it would add backlog and external
677
- Gateway traffic without changing this conclusion. A clean-sheet drain test is
678
- still required separately.
679
-
680
- ## 2026-07-25 — new Gateway fast-append and management-loop smoke test
681
-
682
- - Branch: `refactor/thin-sync-gateway`
683
- - Harness: `.local/lib-test-beta-0.2.0-beta1`
684
- - Backend: MikroORM + SQLite locally, newly deployed Apps Script Gateway
685
- - Database: fresh `data/typed-sheets-load-test-20260725-new-gateway.sqlite`
686
- - Gateway configuration: new untracked `.env` values supplied for this test;
687
- the secret is intentionally not recorded here
688
- - Background worker: enabled; reconciliation(불일치 보정) management loop was
689
- enabled with its default 60-second interval
690
- - Command:
691
-
692
- ```sh
693
- env LOAD_TEST_PROVISION=1 LOAD_TEST_DURATION_MS=3000 \
694
- LOAD_TEST_CONCURRENCY=1 LOAD_TEST_WRITE_RATIO=1 \
695
- LOAD_TEST_SEED_CUSTOMERS=20 LOAD_TEST_SEED_PRODUCTS=20 \
696
- LOAD_TEST_DRAIN=1 LOAD_TEST_SYNC_DRAIN_TIMEOUT_MS=120000 \
697
- npm run load-test
698
- ```
699
-
700
- | Stage | Setup / provision | Traffic duration | Generated orders | HTTP errors | Applied | Pending | Processing | Failed |
701
- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
702
- | 20/20 seed | 13.24 s | 3 s | 435 | 0 | 300 | 2,178 | 100 | 0 |
703
-
704
- The local API successfully created 435 orders, 834 order items, and 435
705
- payments; all 562 HTTP requests succeeded. The Gateway accepted
706
- `fastAppendRows` batches of 20, 80, and 100 rows, and each returned `applied`,
707
- but the remote requests still took approximately 16–37 seconds each. The
708
- SQLite producer therefore outpaced the Gateway consumer and the drain did not
709
- reach zero within 120 seconds.
710
-
711
- The management loop exposed a compatibility defect after provisioning:
712
- `readSnapshot` returned a protocol version that the Node client rejected as
713
- unsupported. Consequently reconciliation did not run, so this result validated
714
- fast append only and did not validate automatic drift repair. The protocol was
715
- aligned in the follow-up run below.
716
-
717
- ## 2026-07-26 — protocol-aligned Gateway load test
718
-
719
- - Branch: `refactor/thin-sync-gateway`
720
- - Harness: `.local/lib-test-beta-0.2.0-beta1`
721
- - Backend: MikroORM + SQLite locally, deployed Apps Script Gateway remotely
722
- - Database: fresh `/private/tmp/typed-sheets-protocol-fixed-20260726.sqlite`
723
- - Gateway configuration: current untracked `.env` values supplied for this
724
- test; the secret is intentionally not recorded here
725
- - Background worker: enabled; reconciliation(불일치 보정) management loop was
726
- enabled
727
- - The SQLite database was fresh, but the remote spreadsheet was reused and was
728
- not an empty-sheet comparison. The first reconciliation scan observed
729
- existing remote rows.
730
- - Server command:
731
-
732
- ```sh
733
- TYPED_SHEETS_DB_PATH=/private/tmp/typed-sheets-protocol-fixed-20260726.sqlite \
734
- TYPED_SHEETS_API_PORT=3200 TYPED_SHEETS_BACKGROUND_WORKER=1 npm start
735
- ```
736
-
737
- - Load command:
738
-
739
- ```sh
740
- LOAD_TEST_PROVISION=1 LOAD_TEST_DURATION_MS=3000 \
741
- LOAD_TEST_CONCURRENCY=1 LOAD_TEST_WRITE_RATIO=1 \
742
- LOAD_TEST_SEED_CUSTOMERS=20 LOAD_TEST_SEED_PRODUCTS=20 \
743
- LOAD_TEST_DRAIN=1 LOAD_TEST_SYNC_DRAIN_TIMEOUT_MS=120000 \
744
- npm run load-test
745
- ```
746
-
747
- | Metric | Result |
748
- | --- | ---: |
749
- | Provision setup | 7.49 s |
750
- | Traffic duration | 3 s |
751
- | Generated orders | 465 |
752
- | Total HTTP requests | 591 |
753
- | HTTP errors | 0 |
754
- | Steady-state `/orders` requests | 465 successful; p50 6.21 ms / p95 8.93 ms / p99 12.53 ms |
755
- | Total elapsed including drain wait | 130.99 s |
756
- | Final applied effects | 20 |
757
- | Final pending effects | 3,605 |
758
- | Final failed effects | 0 |
759
-
760
- The protocol error disappeared: every observed `readSnapshot` request returned
761
- `ok: true`, including reconciliation reads. However, the first reconciliation
762
- scan on the reused sheet found 1,350 desired rows and 500 scanned remote rows,
763
- then enqueued 885 correction effects. Those corrections competed with the
764
- load-test writes. The Gateway applied the first 20-row `fastAppendRows` batch,
765
- but returned `operation_failed` for later 80-row and 100-row batches; the
766
- worker requeued those effects and no final `failed` rows remained. This run
767
- therefore confirms protocol compatibility, but not successful convergence.
768
-
769
- The result also shows that starting reconciliation immediately against a
770
- non-empty reused sheet contaminates a producer-throughput test. The next
771
- benchmark should use a clean spreadsheet, disable reconciliation for the raw
772
- fast-append measurement, then run reconciliation separately with a controlled
773
- drift case.
774
-
775
- ## 2026-07-26 — clean Sheet protocol-aligned load test
776
-
777
- - Branch: `refactor/thin-sync-gateway`
778
- - Harness: `.local/lib-test-beta-0.2.0-beta1`
779
- - Backend: MikroORM + SQLite locally, newly deployed Apps Script Gateway
780
- - Database: fresh `/private/tmp/typed-sheets-new-sheet-20260726.sqlite`
781
- - Gateway configuration: new untracked `.env` values supplied for this test;
782
- the secret is intentionally not recorded here
783
- - Background worker: enabled; reconciliation(불일치 보정) management loop was
784
- enabled
785
- - The remote Sheet was new and passed provisioning successfully.
786
- - Command:
787
-
788
- ```sh
789
- LOAD_TEST_PROVISION=1 LOAD_TEST_DURATION_MS=3000 \
790
- LOAD_TEST_CONCURRENCY=1 LOAD_TEST_WRITE_RATIO=1 \
791
- LOAD_TEST_SEED_CUSTOMERS=20 LOAD_TEST_SEED_PRODUCTS=20 \
792
- LOAD_TEST_DRAIN=1 LOAD_TEST_SYNC_DRAIN_TIMEOUT_MS=120000 \
793
- npm run load-test
794
- ```
795
-
796
- | Metric | Result |
797
- | --- | ---: |
798
- | Provision setup | 16.89 s |
799
- | Traffic duration | 3 s |
800
- | Generated orders | 393 |
801
- | Generated order items | 759 |
802
- | Generated payments | 393 |
803
- | Total HTTP requests | 519 |
804
- | HTTP errors | 0 |
805
- | Steady-state `/orders` requests | 393 successful; p50 7.29 ms / p95 11.26 ms / p99 15.95 ms |
806
- | Total elapsed including drain wait | 140.32 s |
807
- | Final applied effects | 100 |
808
- | Final pending effects | 2,164 |
809
- | Final processing effects | 100 |
810
- | Final failed effects | 0 |
811
-
812
- The clean Sheet removed the previous reused-data problem for the initial
813
- fast-append calls: 20-row and 80-row `fastAppendRows` batches were both
814
- applied successfully, and no protocol or `operation_failed` error was observed
815
- during those batches. However, the management loop later saw the canonical
816
- SQLite state growing faster than the Sheet snapshot and enqueued additional
817
- reconciliation work. The worker could not drain the resulting queue within
818
- 120 seconds; 2,164 effects remained pending and 100 remained processing when
819
- the load runner stopped. A separate 20-effect correction request observed an
820
- `operation_failed` response while the server was being shut down, so that
821
- late response should not be treated as a clean fast-append result.
822
-
823
- This isolates the current conclusion more clearly: a clean Sheet can accept
824
- the tested 20- and 80-row fast-append batches, but the combined producer plus
825
- reconciliation workload still exceeds the worker/Gateway drain rate. A raw
826
- fast-append benchmark with reconciliation disabled is still needed to measure
827
- the Gateway ceiling without management-loop work competing for the same worker.
828
-
829
- ## 2026-07-26 — new Gateway progressive load test
830
-
831
- - Branch: `refactor/thin-sync-gateway`
832
- - Harness: `.local/lib-test-beta-0.2.0-beta1`
833
- - Backend: MikroORM + SQLite locally, newly deployed Apps Script Gateway
834
- - Database: fresh `/private/tmp/typed-sheets-progressive-20260726-01.sqlite`
835
- - Gateway configuration: new untracked values supplied for this test; the
836
- secret is intentionally not recorded here
837
- - Provisioning: five projection sheets created and initialized successfully
838
- - Background worker: enabled; reconciliation management loop enabled
839
- - Traffic: write-only order workload (`LOAD_TEST_WRITE_RATIO=1`), 5 seconds per
840
- stage, 20 customers and 20 products seeded at each stage, no drain wait
841
-
842
- The server was started with the Gateway environment variables supplied for this
843
- run and `TYPED_SHEETS_WORKER_MAX_EFFECTS=100`. Each stage reused the same local
844
- database and newly provisioned spreadsheet so the table and outbox totals show
845
- the cumulative effect of the stages.
846
-
847
- | Stage | Concurrency | Orders | `/orders` req/s | p50 / p95 / p99 (ms) | Final outbox |
848
- | --- | ---: | ---: | ---: | ---: | --- |
849
- | 1 | 1 | 683 | 135.11 | 6.97 / 11.33 / 15.80 | applied 340, pending 4,084 |
850
- | 2 | 2 | 582 | 112.71 | 15.53 / 27.20 / 36.95 | applied 340, failed 100, pending 8,431 |
851
- | 3 | 4 | 481 | 93.66 | 42.46 / 67.13 / 141.94 | applied 340, failed 66, pending 12,462, processing 80 |
852
-
853
- All API requests succeeded with zero HTTP errors. The SQLite totals after the
854
- third stage were 1,746 orders, 3,385 order items, and 1,746 payments. The
855
- write path therefore accepted traffic, but the projection path did not keep up:
856
- the backlog increased by roughly 4,000 effects per stage while the applied
857
- count remained at 340.
858
-
859
- The server logs identified two concrete causes:
860
-
861
- 1. The reconciliation scan observed 340 Sheet rows but 4,434 desired SQLite
862
- rows and enqueued 1,197 correction effects while the original outbox was
863
- still being drained. Reconciliation is therefore amplifying the backlog
864
- when it runs before the initial append queue has caught up.
865
- 2. The correction effects include the `_deleted` tombstone field, but the
866
- registered ranges for the test sheets do not declare `_deleted`. Gateway
867
- postcondition recovery repeatedly failed with:
868
-
869
- ```text
870
- postcondition_read_failed: Effect field is not a registered header: _deleted
871
- ```
872
-
873
- There was also a direct Gateway cost signal: a `readSnapshot` for a 320-row
874
- sheet took 65.553 seconds. This is outside the fast append write itself, but it
875
- shows that running full reconciliation snapshots during a high-write backlog
876
- can monopolize the same worker and make convergence worse.
877
-
878
- This run is not a successful convergence test. It demonstrates that the raw
879
- SQLite API path remains responsive, while the current combination of pending
880
- effect drain, immediate reconciliation, and an invalid tombstone field cannot
881
- reach outbox zero.
882
-
883
- ## 2026-07-26 — pure fast-append Gateway throughput
884
-
885
- - Branch: `refactor/thin-sync-gateway`
886
- - Harness: temporary `.local/fast-append-load-test.mjs`
887
- - Backend: direct `AppsScriptOperationClient` + `createFastAppendRowsOperation`
888
- calls; no SQLite
889
- outbox, effect worker, reconciliation, snapshot, postcondition, receipt, or
890
- delete operation was involved
891
- - Database: `/private/tmp/typed-sheets-fast-append-20260726-01.sqlite` was used
892
- only to provision the five projection registry entries; it was not part of
893
- the measured write path
894
- - Spreadsheet: newly provisioned `LoadTest_Customers` projection
895
- - Gateway mode: the newly deployed fast-append-only Code.gs; a verification
896
- `readSnapshot` request was rejected with `Fast append benchmark accepts only
897
- fastAppendRows.`
898
- - Request limit: 100 rows per Gateway call; the 200-row stage used two calls
899
-
900
- | Batch rows | Gateway calls | Applied rows | Elapsed | Rows/s | Cells/s |
901
- | ---: | ---: | ---: | ---: | ---: | ---: |
902
- | 20 | 1 | 20 | 3,309.04 ms | 6.04 | 36.26 |
903
- | 50 | 1 | 50 | 2,766.91 ms | 18.07 | 108.42 |
904
- | 100 | 1 | 100 | 2,506.87 ms | 39.89 | 239.34 |
905
- | 200 | 2 | 200 | 5,459.14 ms | 36.64 | 219.82 |
906
-
907
- All 370 requested rows were acknowledged as `applied`; no row was left for a
908
- worker or reconciliation pass. This is the first clean measurement of the
909
- Gateway's current fast-append path. It confirms that the pure path can process
910
- the complete tested load, but its steady-state ceiling is approximately 36–40
911
- rows/s for six-column rows, with roughly 2.5–2.8 seconds of latency per 50–100
912
- row request. The earlier backlog and `_deleted` failures were therefore not
913
- part of this measurement.
914
-
915
- ## 2026-07-26 — continuous drain follow-up
916
-
917
- - Branch: `refactor/thin-sync-gateway`
918
- - Harness: `.local/lib-test-beta-0.2.0-beta1`
919
- - Backend: the same local SQLite database and new Sheet from the preceding
920
- clean-Sheet test
921
- - Server command:
922
-
923
- ```sh
924
- TYPED_SHEETS_DB_PATH=/private/tmp/typed-sheets-new-sheet-20260726.sqlite \
925
- TYPED_SHEETS_API_PORT=3200 TYPED_SHEETS_BACKGROUND_WORKER=1 npm start
926
- ```
927
-
928
- - The server was restarted with the previous outbox intact. A local metrics
929
- poll ran every 15 seconds for 180 seconds; no new application traffic was
930
- generated.
931
-
932
- | Metric | Result |
933
- | --- | ---: |
934
- | Initial recovered state | applied 180 / pending 2,084 / processing 80 / failed 20 |
935
- | Applied during observation | +240 |
936
- | Final applied effects | 420 |
937
- | Final pending effects | 2,337 |
938
- | Final processing effects | 0 |
939
- | Final failed effects | 20 |
940
- | Observation duration | 180 s |
941
-
942
- The worker did continue consuming effects after the load runner had stopped;
943
- this disproves the idea that the 120-second test timeout closes the Gateway or
944
- stops synchronization. However, the queue did not converge. Reconciliation
945
- continued to enqueue correction effects while the worker was consuming the
946
- existing queue, so pending temporarily increased from 2,104 to 2,417 before
947
- ending at 2,337.
948
-
949
- The 20 failed effects were repeatedly rejected during batched postcondition
950
- recovery with:
951
-
952
- ```text
953
- postcondition_read_failed: Effect field is not a registered header: _deleted
954
- ```
955
-
956
- Those effects reached four attempts and remained failed. This is a concrete
957
- blocking defect independent of the 120-second drain window: the worker can
958
- consume some work indefinitely, but these effects cannot complete until the
959
- registered Sheet headers and the postcondition payload agree. The continuous
960
- test therefore confirms ongoing consumption, but not eventual convergence.
961
-
962
- ## 2026-07-27 — lib-test-beta thin-interface fast-append benchmark
963
-
964
- - Branch: `refactor/thin-sync-gateway`
965
- - Historical harness: `.local/lib-test-beta-0.2.0-beta1/fast-append-load-test.mjs`
966
- (removed after this benchmark)
967
- - Server: `.local/lib-test-beta-0.2.0-beta1/server.mjs`, using the current
968
- `AppsScriptOperationClient` and `createFastAppendRowsOperation` library APIs
969
- - Database: fresh `/private/tmp/typed-sheets-lib-test-fast-append-20260727.sqlite`
970
- used only to initialize the lib-test runtime; SQLite outbox/effect processing
971
- was not part of the measured path
972
- - Gateway: newly supplied deployment; URL, shared secret, and spreadsheet ID
973
- are intentionally not recorded
974
- - Background effect worker: disabled
975
- - Reconciliation: disabled
976
- - Setup: one unmeasured operation created/initialized `LoadTest_Customers` and
977
- wrote its six-column header
978
- - Measured scenario: three sequential requests, one contiguous `setValues()`
979
- operation per stage; no metadata, snapshot, CAS, receipt, postcondition, or
980
- delete work
981
-
982
- Command used for the measured runner:
983
-
984
- ```sh
985
- LOAD_TEST_BASE_URL=http://127.0.0.1:3205 \
986
- node .local/lib-test-beta-0.2.0-beta1/fast-append-load-test.mjs
987
- ```
988
-
989
- | Batch rows | Gateway calls | Applied rows | Elapsed | Rows/s | Cells/s |
990
- | ---: | ---: | ---: | ---: | ---: | ---: |
991
- | 20 | 1 | 20 | 2,275 ms | 8.79 | 52.75 |
992
- | 100 | 1 | 100 | 2,729 ms | 36.64 | 219.86 |
993
- | 370 | 1 | 370 | 3,792 ms | 97.57 | 585.44 |
994
-
995
- All 490 rows across the three stages returned `applied`, with HTTP 200 and no
996
- client or remote error. The 20-row stage includes the fixed HTTP/Apps Script
997
- startup cost; larger contiguous writes amortized that cost substantially. The
998
- 370-row one-request result is materially faster than the earlier 200-row
999
- two-request measurement (36.64 rows/s), but it is not a complete production
1000
- sync test because it bypasses SQLite outbox consumption and reconciliation.
1001
-
1002
- ## 2026-07-27 — operational User/Order/OrderItem end-to-end test
1003
-
1004
- - Branch: `refactor/thin-sync-gateway`
1005
- - Harness: `.local/typed-sheets-e2e.mjs`
1006
- - Command: `npm run build`, then `node .local/typed-sheets-e2e.mjs` with the
1007
- supplied Gateway environment variables; secret and Sheet ID are intentionally
1008
- not recorded
1009
- - Package boundary: the server imported the built package entry
1010
- (`dist/index.js`), `typed-sheets/orm`, and `typed-sheets/mikro-orm`; the test
1011
- runner rejected a `src` import and did not call the Gateway directly
1012
- - Domain tables: `User`, `Order`, and `OrderItem`; the library's internal
1013
- registry and outbox tables were also created by the MikroORM SQLite adapter
1014
- - Gateway: the supplied thin `Code.gs` deployment; provisioning and all remote
1015
- writes went through `AppsScriptOperationClient` and
1016
- `AppsScriptOperationSyncGateway`
1017
- - Spreadsheet setup: each run used three unique tabs derived from its run ID,
1018
- so existing Sheet data was neither deleted nor reused
1019
- - Worker: started automatically by the server after successful provisioning;
1020
- maximum 100 effects per pass
1021
- - Reconciliation: disabled for this isolated SQLite-to-Sheets create path
1022
- - API concurrency: 4 HTTP clients
1023
-
1024
- The functional smoke flow created one User, one Order, and two OrderItems, then
1025
- read the User-Order-OrderItem graph through the HTTP API. Each load stage used
1026
- the same API and EntityManager path, generated durable SQLite outbox effects,
1027
- and waited for the server-owned worker to drain them. The stage measurements
1028
- exclude one-time package build, SQLite migration, Sheet provisioning, and the
1029
- functional smoke setup; they are the no-setup/steady-state measurements.
1030
-
1031
- | Stage | New Orders | New OrderItems | Setup excluded | Stage time | SQLite rows = Sheet rows | Failed effects |
1032
- | ---: | ---: | ---: | :---: | ---: | --- | ---: |
1033
- | Smoke | 1 | 2 | — | functional check | 1 User / 1 Order / 2 Items | 0 |
1034
- | 20 | 20 | 40 | yes | 12,673 ms | 1 / 21 / 42 | 0 |
1035
- | 100 | 100 | 200 | yes | 16,918 ms | 1 / 121 / 242 | 0 |
1036
- | 370 | 370 | 740 | yes | 40,196 ms | 1 / 491 / 982 | 0 |
1037
-
1038
- The final cumulative worker count was 1,474 applied effects, with zero failed
1039
- effects and zero ready outbox effects. Every dedicated Sheet tab matched the
1040
- SQLite domain count. The server was stopped and restarted against the same
1041
- SQLite file; the original Order and its two OrderItems were read successfully
1042
- after restart.
1043
-
1044
- ## 2026-07-27 — operational test with idle-gated reconciliation
1045
-
1046
- - Branch: `refactor/thin-sync-gateway`
1047
- - Harness: `.local/typed-sheets-e2e.mjs`
1048
- - Backend: built `typed-sheets` package, MikroORM, and local SQLite
1049
- - Gateway: the previously supplied Apps Script deployment; credentials and
1050
- spreadsheet ID are intentionally not recorded
1051
- - Scenario command (Gateway environment variables were supplied separately):
1052
-
1053
- ```sh
1054
- TYPED_SHEETS_OPERATIONAL_LOAD_STAGES=1 \
1055
- TYPED_SHEETS_OPERATIONAL_RECONCILIATION_INTERVAL_MS=1000 \
1056
- TYPED_SHEETS_OPERATIONAL_DRAIN_TIMEOUT_MS=120000 \
1057
- node .local/typed-sheets-e2e.mjs
1058
- ```
1059
-
1060
- - Reconciliation policy: run only after the worker pass is idle and SQLite has
1061
- no `pending` or `processing` outbox effects
1062
- - Dataset: one functional smoke flow, followed by one additional Order and two
1063
- OrderItems through the HTTP API
1064
- - Setup: package build, SQLite migration, Sheet provisioning, and functional
1065
- smoke setup are excluded from the stage timing
1066
-
1067
- | Stage | New Orders | New OrderItems | Stage time | SQLite rows = Sheet rows | Failed effects |
1068
- | ---: | ---: | ---: | ---: | --- | ---: |
1069
- | 1 | 1 | 2 | 38,724 ms | 1 User / 2 Orders / 4 Items | 0 |
1070
-
1071
- The operational E2E completed successfully. It imported the built package,
1072
- created all three domain tables through the library, wrote through the SQLite
1073
- outbox and server-owned worker, read the entity graph through the HTTP API, and
1074
- verified the same row counts in Sheets. The server restart check also passed.
1075
- The final outbox contained 21 applied effects and zero pending or processing
1076
- effects.
1077
-
1078
- The test also exposed a remaining system-level issue. Even after the worker
1079
- became idle, reconciliation observed a stale or incomplete Sheet snapshot and
1080
- created correction effects:
1081
-
1082
- | Reconciliation observation | Desired rows | Drifted rows | Missing rows | Correction effects |
1083
- | --- | ---: | ---: | ---: | ---: |
1084
- | First scan | 4 | 0 | 4 | 4 |
1085
- | Later scan | 7 | 4 | 3 | 7 |
1086
-
1087
- Fourteen reconciliation effects were eventually applied, which is why the
1088
- final row counts converged. This means the idle gate prevents reconciliation
1089
- from competing with an actively pending outbox, but it does not yet prove that
1090
- an `applied` worker result and the subsequent Sheet snapshot are immediately
1091
- consistent. The normal write result, Gateway snapshot behavior, and
1092
- reconciliation identity matching need separate investigation before claiming
1093
- that the correction path is only a rare safety net.
1094
-
1095
- Compared with the earlier pure fast-append benchmark (370 synthetic rows in
1096
- 3,792 ms), this run is not a raw `setValues()` comparison: it includes 1,110
1097
- entity effects, SQLite flush/outbox work, worker leasing, multiple Gateway
1098
- requests, and Sheet row-count verification. It confirms that the current
1099
- library-to-worker-to-thin-Gateway create path converges for this workload, not
1100
- that update/delete or reconciliation paths are complete.
1101
-
1102
- ## 2026-07-27 — batched anchor lookup polling benchmark
1103
-
1104
- - Branch: `refactor/thin-sync-gateway`
1105
- - Harness: `.local/typed-sheets-e2e.mjs`
1106
- - Backend: built `typed-sheets` package, MikroORM, and local SQLite
1107
- - Gateway: an operational `Code.gs` deployment; URL, secret, and spreadsheet ID
1108
- are intentionally not recorded
1109
- - Reconciliation: disabled so it could not add correction effects during the
1110
- polling measurement
1111
- - Scenario: create one smoke User plus 20 Orders and 40 OrderItems through the
1112
- HTTP API, wait for the server-owned effect worker to drain, then run the
1113
- polling path twice
1114
- - Command shape (Gateway environment variables were supplied separately):
1115
-
1116
- ```sh
1117
- npm run build
1118
- TYPED_SHEETS_OPERATIONAL_RECONCILIATION=0 \
1119
- TYPED_SHEETS_OPERATIONAL_POLLING_BENCHMARK=1 \
1120
- TYPED_SHEETS_OPERATIONAL_POLLING_RUNS=2 \
1121
- TYPED_SHEETS_OPERATIONAL_LOAD_STAGES=20 \
1122
- node .local/typed-sheets-e2e.mjs
1123
- ```
1124
-
1125
- The observation operation was changed from one `getDeveloperMetadata()` call
1126
- per row to one Sheet-scoped `DeveloperMetadataFinder` search per operation.
1127
- The returned row locations are indexed in memory and reused by both anchor
1128
- assignment and snapshot construction. Existing duplicate-anchor and
1129
- unanchored-row behavior remains unchanged.
1130
-
1131
- | Poll run | Rows scanned | Anchors assigned | Users | Orders | OrderItems | Poll time |
1132
- | --- | ---: | ---: | ---: | ---: | ---: | ---: |
1133
- | First (includes anchor setup) | 64 | 64 | 1 | 21 | 42 | 56,159 ms |
1134
- | Second (steady state) | 64 | 0 | 1 | 21 | 42 | 55,401 ms |
1135
-
1136
- The load stage itself took 12,483 ms and produced matching SQLite/Sheet row
1137
- counts with zero failed effects. The first poll assigned one anchor per row;
1138
- this write remains intentionally per-row because `fastAppend` does not create
1139
- Developer Metadata. The second poll had no anchor writes, but its duration was
1140
- almost unchanged. Therefore the batch lookup removes the explicit row-by-row
1141
- metadata-read pattern, but it is not yet a material end-to-end polling speedup.
1142
-
1143
- The remaining polling cost is now likely distributed across the full range
1144
- reads (`values`, formulas, display values, and merged ranges), per-cell
1145
- normalization and SHA-256 hashing, and conversion of the returned metadata
1146
- locations. A phase-level benchmark is required before claiming that any one of
1147
- those is the next dominant stage. This result also means the earlier estimate
1148
- of a few minutes for 10,000 rows cannot be accepted as a measured expectation;
1149
- the current steady-state result is still too slow to extrapolate safely.
1150
-
1151
- ## 2026-07-27 — operational timing run on the newly deployed Gateway
1152
-
1153
- - Branch: `refactor/thin-sync-gateway`
1154
- - Harness: `.local/typed-sheets-e2e.mjs`
1155
- - Backend: built `typed-sheets` package, MikroORM, and local SQLite
1156
- - Package boundary: the server imported `dist/index.js`; the runner did not
1157
- import `src/**` or call SQLite/Gateway operations directly
1158
- - Gateway: the newly supplied Apps Script deployment; URL, secret, and
1159
- spreadsheet ID are intentionally not recorded
1160
- - Reconciliation: disabled for this isolated write-path measurement
1161
- - API concurrency: 4; load stage: 20 new Orders with 40 new OrderItems
1162
- - Sheet setup: three unique projection tabs were provisioned for this run;
1163
- setup and provisioning are excluded from the stage timer
1164
- - Command shape (Gateway environment variables were supplied separately):
1165
-
1166
- ```sh
1167
- TYPED_SHEETS_OPERATIONAL_RECONCILIATION=0 \
1168
- TYPED_SHEETS_OPERATIONAL_LOAD_STAGES=20 \
1169
- TYPED_SHEETS_OPERATIONAL_DRAIN_TIMEOUT_MS=240000 \
1170
- node .local/typed-sheets-e2e.mjs
1171
- ```
1172
-
1173
- The functional smoke path created a User, an Order, and two OrderItems,
1174
- read the entity graph through the HTTP API, updated an Order through the
1175
- EntityManager, and removed the Order and its two related test items. Delete is
1176
- represented in the System_State projection by the library's
1177
- `__typed_sheets_deleted` tombstone, so the verification compares active rows,
1178
- not physical Sheet row count.
1179
-
1180
- | Stage | New Orders | New OrderItems | Stage time | SQLite active rows = Sheet active rows | Failed effects |
1181
- | ---: | ---: | ---: | ---: | --- | ---: |
1182
- | Smoke | 1 | 2 | functional check | 1 User / 1 Order / 2 Items | 0 |
1183
- | 20 | 20 | 40 | 28,182 ms | 1 / 21 / 42 | 0 |
1184
-
1185
- The stage finished with 69 applied effects, zero failed effects, and no ready
1186
- or active effects. The server was stopped and restarted against the same
1187
- SQLite file; the original Order and its two OrderItems were readable after
1188
- restart. All remote requests returned HTTP 200.
1189
-
1190
- The timing sink separated local ORM work, worker orchestration, and the remote
1191
- Gateway. The largest steady-state totals were:
1192
-
1193
- | Scope | Operation | Phase | Calls | Total | Max |
1194
- | --- | --- | --- | ---: | ---: | ---: |
1195
- | Worker | append | `append_gateway_dispatch` | 9 | 21,698 ms | 4,301 ms |
1196
- | Worker | update-like | `regular_gateway_dispatch` | 3 | 12,512 ms | 5,217 ms |
1197
- | Gateway | update-like | `dispatcher_eval` | 3 | 6,553 ms | 2,920 ms |
1198
- | Gateway | update-like | `script_lock` | 3 | 6,318 ms | 2,830 ms |
1199
- | Gateway | append | `dispatcher_flush` | 9 | 2,091 ms | 293 ms |
1200
- | Gateway | append | `append_range_lookup` | 9 | 896 ms | 248 ms |
1201
- | Gateway | append | `set_values` | 9 | 92 ms | 17 ms |
1202
- | ORM flush | append | `flush_total` | 23 | 55 ms | 4 ms |
1203
- | ORM flush | delete | `flush_total` | 1 | 2 ms | 2 ms |
1204
-
1205
- This run confirms that local SQLite persistence and outbox creation are not the
1206
- current bottleneck: ORM flushes were measured in milliseconds. The dominant
1207
- cost is the Gateway round trip and Apps Script dispatcher/lock overhead. The
1208
- raw append `setValues()` work itself was only 92 ms across nine requests; the
1209
- remaining append time is request dispatch, range lookup, and remote execution
1210
- overhead. At the ORM layer `em.remove()` is correctly classified as `delete`,
1211
- but its System_State materialization is a tombstone write and therefore appears
1212
- in the Gateway's regular update-like timing path rather than as a physical row
1213
- delete.
1214
-
1215
- This is a successful functional and timing run for the 20-order operational
1216
- scenario. It does not establish the sustained 100/370-order drain ceiling, nor
1217
- does it test reconciliation or user-edit conflict handling; those require
1218
- separate runs so their reads and corrective effects do not contaminate this
1219
- write-path measurement.
1220
-
1221
- ## 2026-07-27 — progressive operational throughput run
1222
-
1223
- - Branch: `refactor/thin-sync-gateway`
1224
- - Harness: `.local/typed-sheets-e2e.mjs`
1225
- - Backend: built `typed-sheets` package, MikroORM, and local SQLite
1226
- - Package boundary: the server imported `dist/index.js`; the runner did not
1227
- import `src/**` or call SQLite/Gateway operations directly
1228
- - Gateway: the newly supplied Apps Script deployment; credentials are not
1229
- recorded
1230
- - Reconciliation: disabled for this isolated throughput measurement
1231
- - Stage Sheet verification: disabled so the result measures API ingestion plus
1232
- worker drain, without an additional full snapshot read after every stage
1233
- - Load stages: 20, 100, and 370 new Orders; each Order created two OrderItems
1234
- - API concurrency: 4; Gateway timeout: 120 seconds; drain timeout: 240 seconds
1235
- - Command shape (Gateway environment variables were supplied separately):
1236
-
1237
- ```sh
1238
- TYPED_SHEETS_OPERATIONAL_RECONCILIATION=0 \
1239
- TYPED_SHEETS_OPERATIONAL_LOAD_STAGES='20,100,370' \
1240
- TYPED_SHEETS_OPERATIONAL_VERIFY_STAGE_SHEETS=0 \
1241
- TYPED_SHEETS_OPERATIONAL_DRAIN_TIMEOUT_MS=240000 \
1242
- node .local/typed-sheets-e2e.mjs
1243
- ```
1244
-
1245
- | Stage | New Orders | New OrderItems | Materialized rows | Stage time | Approx. rows/s | Cumulative applied | Failed |
1246
- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
1247
- | 20 | 20 | 40 | 60 | 11,987 ms | 5.0 | 69 | 0 |
1248
- | 100 | 100 | 200 | 300 | 14,960 ms | 20.1 | 369 | 0 |
1249
- | 370 | 370 | 740 | 1,110 | 36,865 ms | 30.1 | 1,479 | 0 |
1250
-
1251
- All three stages completed through the real HTTP server and the deployed
1252
- Gateway. The worker reached the expected applied count at every stage, no
1253
- effect failed, the server restarted successfully against the same SQLite file,
1254
- and the persisted Order/OrderItem graph remained readable after restart.
1255
-
1256
- The dominant measured cost was still the remote worker-to-Gateway path:
1257
-
1258
- | Scope | Phase | Calls | Total | Max |
1259
- | --- | --- | ---: | ---: | ---: |
1260
- | Worker | `append_gateway_dispatch` | 27 | 67,825 ms | 3,964 ms |
1261
- | Worker | `worker_total` (append) | 20 | 71,592 ms | 6,068 ms |
1262
- | Gateway | `append_range_lookup` | 27 | 3,941 ms | 938 ms |
1263
- | Gateway | `set_values` | 27 | 370 ms | 52 ms |
1264
- | ORM flush | `flush_total` (append) | 493 | 1,118 ms | 7 ms |
1265
-
1266
- This separates the bottleneck more clearly: local SQLite/ORM flush work is
1267
- roughly millisecond-scale, and the raw Gateway `setValues()` work is small.
1268
- Most elapsed time is HTTP/Apps Script dispatch and Gateway range lookup. The
1269
- stage timer includes API entity creation and worker drain, so it is an
1270
- end-to-end operational throughput result, not a raw `setValues()` benchmark.
1271
-
1272
- The polling probe returned zero rows and zero tables on both steady-state runs.
1273
- That is **not** a measured user-input polling speed: this operational fixture
1274
- registers only `system_state` projections, so there is no `user_input`
1275
- projection for the optimized polling path to scan. User-input polling needs a
1276
- separate fixture with a valid user-input registration and data.
1277
-
1278
- ## 2026-07-27 — corrected full-table polling run
1279
-
1280
- The previous section's polling probe was invalid for this operational scenario:
1281
- it filtered by the projection label `user_input`, even though the test's User,
1282
- Order, and OrderItem tabs are all domain tables represented by `system_state`.
1283
- The polling helper now batches all registered projections; `user_input` remains
1284
- an ownership/edit-mode classification, not a domain table name.
1285
-
1286
- - Branch: `refactor/thin-sync-gateway`
1287
- - Harness: `.local/typed-sheets-e2e.mjs`
1288
- - Backend and package boundary: same as the progressive throughput run above
1289
- - Reconciliation: disabled
1290
- - Load stage: 20 Orders and 40 OrderItems, plus the functional smoke rows
1291
- - Polling: one combined Gateway observation request for User, Order, and
1292
- OrderItem, followed by two polling passes
1293
- - Command shape (Gateway environment variables were supplied separately):
1294
-
1295
- ```sh
1296
- TYPED_SHEETS_OPERATIONAL_RECONCILIATION=0 \
1297
- TYPED_SHEETS_OPERATIONAL_LOAD_STAGES=20 \
1298
- TYPED_SHEETS_OPERATIONAL_POLLING_BENCHMARK=1 \
1299
- TYPED_SHEETS_OPERATIONAL_POLLING_RUNS=2 \
1300
- TYPED_SHEETS_OPERATIONAL_VERIFY_STAGE_SHEETS=0 \
1301
- node .local/typed-sheets-e2e.mjs
1302
- ```
1303
-
1304
- | Poll pass | Rows scanned | Anchors assigned | User rows | Order rows | OrderItem rows | Elapsed |
1305
- | --- | ---: | ---: | ---: | ---: | ---: | ---: |
1306
- | First (anchor setup) | 66 | 60 | 1 | 22 | 43 | 34,919 ms |
1307
- | Second (steady state) | 66 | 0 | 1 | 22 | 43 | 75,150 ms |
1308
-
1309
- Both passes completed without failed effects. The first pass assigned anchors
1310
- to rows created by fast-append; the second pass reused all 66 anchors. The
1311
- steady-state pass was slower than the first despite doing no metadata writes,
1312
- so the dominant polling cost is not only anchor assignment. The full snapshot
1313
- read, Apps Script lock/dispatch latency, range reads, normalization, hashing,
1314
- and SQLite observation persistence all remain in the path.
1315
-
1316
- A separate 20/100/370 run reached the 370 stage but could not start its full
1317
- table polling pass: the Gateway rejected the Order observation while adding
1318
- row Developer Metadata because the test Spreadsheet exceeded its allowed
1319
- metadata storage. This is an independent scalability limit in the current
1320
- anchor-based polling design; it prevents claiming a 370-row polling result.
1321
-
1322
- ## 2026-07-27 — polling phase trace
1323
-
1324
- The observation Gateway and local ingestion path were instrumented to separate
1325
- remote Apps Script work from SQLite comparison and observation persistence. The
1326
- same 20-Order/40-OrderItem operational fixture scanned 66 rows across the three
1327
- domain tabs.
1328
-
1329
- | Poll pass | User Gateway | Order Gateway | OrderItem Gateway | Local SQLite/observation | Total polling |
1330
- | --- | ---: | ---: | ---: | ---: | ---: |
1331
- | First, 60 anchors assigned | 994 ms | 8,917 ms | 18,464 ms | 8 ms | 31,371 ms |
1332
- | Second, steady state | 635 ms | 6,748 ms | 17,661 ms | 6 ms | 27,652 ms |
1333
-
1334
- The first pass spent 14,544 ms assigning 60 row anchors: 4,808 ms for Order
1335
- and 9,735 ms for OrderItem. The second pass assigned no anchors, but still
1336
- spent 10,402 ms reading Developer Metadata, so removing anchor writes alone is
1337
- not enough.
1338
-
1339
- The largest steady-state remote phases were:
1340
-
1341
- | Phase | User | Order | OrderItem | Combined |
1342
- | --- | ---: | ---: | ---: | ---: |
1343
- | `anchor_metadata_read` | 228 ms | 2,794 ms | 7,380 ms | 10,402 ms |
1344
- | `row_normalization` | 30 ms | 786 ms | 5,705 ms | 6,521 ms |
1345
- | `snapshot_hash` | 111 ms | 2,502 ms | 4,204 ms | 6,817 ms |
1346
-
1347
- These phases are nested inside `snapshot_build`, so their combined values must
1348
- not be added again to the per-table Gateway totals. The raw `values_read`
1349
- phase was only 12 ms in steady state, and local SQLite shadow comparison plus
1350
- observation persistence took 6 ms. Therefore the current polling bottleneck is
1351
- not SQLite and not the basic Sheet range read; it is the Developer Metadata
1352
- finder plus repeated Apps Script hashing/normalization work. OrderItem alone
1353
- accounted for 17,661 ms of the 27,652 ms steady-state pass.
1354
-
1355
- This confirms that the current full-snapshot polling path is not suitable as a
1356
- frequent normal synchronization mechanism at this size. It is currently more
1357
- appropriate as an infrequent safety scan, while `onEdit` or a redesigned
1358
- lightweight identity/change path handles the normal user-edit signal. The
1359
- 370-row polling pass remains unmeasured because the current anchor metadata
1360
- quota is reached before the scan completes.
1361
-
1362
- ## 2026-07-27 — lightweight values-only polling
1363
-
1364
- - Branch: `refactor/thin-sync-gateway`
1365
- - Harness: `.local/typed-sheets-e2e.mjs`
1366
- - Backend: built `typed-sheets` package, MikroORM, and local SQLite
1367
- - Package boundary: the server imported `dist/index.js`; the runner did not
1368
- import `src/**` or call SQLite/Gateway operations directly
1369
- - Gateway: the supplied Apps Script deployment; URL, secret, and spreadsheet
1370
- ID are intentionally not recorded
1371
- - Reconciliation: disabled
1372
- - Dataset: one smoke User plus 20 Orders and 40 OrderItems (66 Sheet rows)
1373
- - Scenario: the server-owned worker drained the outbox, then `/poll` issued one
1374
- batched values-only read for User, Order, and OrderItem twice
1375
- - Command shape (Gateway environment variables were supplied separately):
1376
-
1377
- ```sh
1378
- npm run build
1379
- TYPED_SHEETS_OPERATIONAL_RECONCILIATION=0 \
1380
- TYPED_SHEETS_OPERATIONAL_LOAD_STAGES=20 \
1381
- TYPED_SHEETS_OPERATIONAL_POLLING_BENCHMARK=1 \
1382
- TYPED_SHEETS_OPERATIONAL_POLLING_RUNS=2 \
1383
- TYPED_SHEETS_OPERATIONAL_VERIFY_STAGE_SHEETS=0 \
1384
- TYPED_SHEETS_OPERATIONAL_SKIP_TIMING_SUMMARY=1 \
1385
- node .local/typed-sheets-e2e.mjs
1386
- ```
1387
-
1388
- | Poll pass | Rows scanned | Unchanged | Changed | Unknown/invalid | Elapsed | Remote read total |
1389
- | --- | ---: | ---: | ---: | ---: | ---: | ---: |
1390
- | First | 66 | 66 | 0 | 0 / 0 | 2,109 ms | 573 ms |
1391
- | Second, steady state | 66 | 66 | 0 | 0 / 0 | 2,240 ms | 530 ms |
1392
-
1393
- Each pass used one signed Apps Script request containing three independent
1394
- `getValues()` operations. No `LockService`, Developer Metadata, snapshot hash,
1395
- receipt, or observation persistence was used. The remote phases per table were
1396
- approximately 210–220 ms for User, 155–210 ms for Order, and 156–163 ms for
1397
- OrderItem; the remainder was HTTP dispatch and local canonical comparison.
1398
-
1399
- The test passed through the real HTTP server, package boundary, SQLite outbox,
1400
- background worker, and deployed Gateway. The final rows were classified as
1401
- unchanged, including active and tombstoned System_State rows, and the server
1402
- restart check passed. Compared with the previous steady-state full-snapshot
1403
- poll (27,652 ms for the same 66-row shape), the lightweight path was roughly
1404
- 12 times faster. This is a read/compare benchmark: changed rows are returned
1405
- to the caller, but this pass does not yet turn them into evaluated user-edit
1406
- events or canonical writes.
1407
-
1408
- ## Caveats
1409
-
1410
- - The raw benchmark measures a throughput upper bound, not production sync
1411
- behavior.
1412
- - The benchmark uses an isolated sheet and synthetic string values.
1413
- - Network, HTTP, lock acquisition, retry, and response-loss behavior are not
1414
- included.
1415
- - The result strongly indicates that the current Gateway validation and
1416
- metadata path, rather than raw `setValues()` throughput, is the dominant
1417
- bottleneck.
1418
-
1419
- ## 2026-08-03 — clean Locust smoke after adaptive-sync implementation
1420
-
1421
- - Branch: `perf/adaptive-sync-performance`
1422
- - Command:
1423
- `locust -f .local/locustfile.py --host http://127.0.0.1:8787 --headless
1424
- --users 10 --spawn-rate 2 --run-time 60s --csv
1425
- .local/locust-20260803-231300-clean2 --html
1426
- .local/locust-20260803-231300-clean2.html --only-summary`
1427
- - Server: `.local/hikoutei-load-server.mjs`, Node.js 24.3, local SQLite,
1428
- deployed Apps Script Gateway, fresh run ID `load-20260803-231003-clean2`
1429
- - Dataset/scenario: fresh SQLite and fresh `System_State`, `User_Input`, and
1430
- `Sync_Conflicts` projection tabs; 10 mixed Locust users; 2 users/second
1431
- ramp; 60-second workload window.
1432
- - No-setup/steady-state scope: server provisioning and startup were excluded;
1433
- the table below starts after `/health` became ready and includes only the
1434
- Locust workload window. The separate drain snapshot includes background
1435
- worker/polling traffic after Locust stopped.
1436
-
1437
- | Workload | Requests | Failures | p50 | p95 | Maximum | Throughput |
1438
- | --- | ---: | ---: | ---: | ---: | ---: | ---: |
1439
- | `GET /users/:id` | 134 | 0 (0%) | 2 ms | 4 ms | 6 ms | 2.29/s |
1440
- | `POST /users` | 81 | 0 (0%) | 5 ms | 6 ms | 19 ms | 1.38/s |
1441
- | `PATCH /users/:id` | 76 | 0 (0%) | 5 ms | 8 ms | 22 ms | 1.30/s |
1442
- | `POST /__test/user-input` | 38 | 8 (21.05%) | 2.2 s | 31.0 s | 31.9 s | 0.65/s |
1443
- | **Aggregated** | **339** | **8 (2.36%)** | **4 ms** | **2.3 s** | **31.9 s** | **5.78/s** |
1444
-
1445
- The server-side Gateway snapshot after the workload and approximately two
1446
- minutes of background drain contained 100 Gateway requests, 34 failures, p50
1447
- 3.71 s, p95 33.47 s, and maximum 60.00 s. The failures were dominated by
1448
- intermittent non-JSON/timeout Gateway responses. The SQLite outbox had 306
1449
- `pending`, 8 `processing`, 2 `applied`, and 2 `superseded` effects at the
1450
- snapshot; it had not converged, so this is not a successful drain benchmark.
1451
-
1452
- Compared with the earlier 2,648-request run (12.61% HTTP failures, Gateway
1453
- p50 4.34 s, p95 34.25 s, maximum 60.01 s), the clean smoke had a lower HTTP
1454
- failure rate for entity create/read/update traffic, but its User_Input path
1455
- still failed and the outbox did not drain. The runs are not a production
1456
- performance comparison until the Gateway deployment is stable and the
1457
- background outbox converges under the same workload.
1458
-
1459
- Artifacts:
1460
-
1461
- - `.local/locust-20260803-231300-clean2_stats.csv`
1462
- - `.local/locust-20260803-231300-clean2_failures.csv`
1463
- - `.local/locust-20260803-231300-clean2.html`
1464
- - `.local/hikoutei-load-load-20260803-231003-clean2.jsonl`
1465
-
1466
- Caveats: the Gateway returned intermittent HTTP 404/non-JSON responses and
1467
- 60-second client timeouts during the run; Python 3.9 emitted Locust's
1468
- end-of-life warning; and the remote deployment was not independently
1469
- re-deployed during this measurement.
1470
-
1471
- ## 2026-08-03 — Gateway baseline transport gate (stopped)
1472
-
1473
- - Branch: `perf/adaptive-sync-performance`
1474
- - Command:
1475
- `node --env-file=.env --input-type=module <gateway-baseline-probe>`
1476
- - Exact probe: 100 sequential (`concurrency=1`) signed `applyOperations` calls
1477
- using a read-only function that returned the bound spreadsheet ID and an
1478
- integer nonce; no Sheet rows or tabs were created or changed.
1479
- - Environment: current built `dist/`, configured Apps Script Gateway URL,
1480
- configured shared secret, configured spreadsheet ID; secrets and IDs are not
1481
- recorded here.
1482
- - No-setup/steady-state scope: every request was measured after the client was
1483
- constructed; there was no application or Sheet setup work in the request
1484
- loop.
1485
-
1486
- | Requests | Success | Failure | Failure rate | p50/p95/maximum | Duration |
1487
- | ---: | ---: | ---: | ---: | --- | ---: |
1488
- | 100 | 0 | 100 | **100%** | not applicable (all HTTP 405) | 425,373 ms |
1489
-
1490
- Every failure decoded as `invalid_sync_gateway_response` with HTTP 405 and the
1491
- message `Code.gs response was not valid JSON`. This exceeds the 1% stop
1492
- criterion, so no further live sync/load refactor or User_Input migration should
1493
- be accepted until the deployed Gateway URL, redirect/method preservation,
1494
- permissions, deployment version, and quota/runtime path are isolated. This
1495
- probe is diagnostic only and does not establish Sheet convergence.
1496
-
1497
- Artifact: `.local/gateway-baseline-100-20260804.json`.
1498
-
1499
- ## 2026-08-04 — Gateway baseline transport gate (redeployed, passed)
1500
-
1501
- - Branch: `perf/adaptive-sync-performance`
1502
- - Command: ephemeral signed no-op probe using the built
1503
- `AppsScriptOperationClient`; Gateway credentials were supplied only through
1504
- process environment and are not recorded.
1505
- - Exact probe: 100 sequential (`concurrency=1`) signed `applyOperations` calls
1506
- using the same read-only function as the stopped baseline. No Sheet rows or
1507
- tabs were created or changed.
1508
- - Environment: redeployed Apps Script Web App `/exec`, current built `dist/`,
1509
- 60-second client timeout. Setup was excluded from the measured loop.
1510
-
1511
- | Requests | Success | Failure | Failure rate | p50 | p95 | Maximum | Duration |
1512
- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
1513
- | 100 | 100 | 0 | **0%** | 1,956 ms | 4,511 ms | 20,994 ms | 254,786 ms |
1514
-
1515
- All responses were valid HTTP 200 JSON envelopes. Every request followed the
1516
- expected `POST /exec` 302 redirect to the Apps Script `macros/echo` endpoint,
1517
- which returned HTTP 200. Compared with the previous 100% failure baseline, the
1518
- redeployment fixed the transport gate. This probe proves transport and signed
1519
- no-op execution only; it does not establish Sheet write convergence or remove
1520
- the need for the subsequent single-append, replay, and concurrency checks.
1521
-
1522
- Artifact: `.local/gateway-baseline-100-20260804-fixed.json`.
1523
-
1524
- ## 2026-08-04 — Code.gs-only write correctness gate (passed)
1525
-
1526
- - Branch: `perf/adaptive-sync-performance`
1527
- - Command: `node .local/gateway-write-verification-live.mjs` with Gateway
1528
- credentials supplied only through ephemeral process environment; credentials
1529
- are not recorded.
1530
- - Backend: redeployed Apps Script Web App `/exec`, built `dist/`, no Advanced
1531
- Sheets Service or manifest activation.
1532
- - Scenario: create a uniquely named temporary tab, append one row, replay the
1533
- same effect, submit a same-effect payload mismatch, submit a duplicate
1534
- identity, simulate response loss after the remote append response, replay the
1535
- uncertain effect, inspect row/receipt counts, and delete the temporary tab and
1536
- newly created receipt tab. Setup and cleanup were outside the correctness
1537
- assertions; no existing user tab was retained or modified.
1538
-
1539
- | Check | Result |
1540
- | --- | --- |
1541
- | Single append | `applied`, receipt-backed visible hash/revision |
1542
- | Exact replay | `applied`, target rows 2 including header, one receipt per effect |
1543
- | Payload mismatch | rejected with `operation_failed` |
1544
- | Duplicate identity | rejected with `operation_failed` |
1545
- | Simulated response loss | classified as network uncertainty |
1546
- | Response-loss replay | `applied`, no duplicate target/receipt row |
1547
- | Cleanup | temporary target and receipt tabs deleted |
1548
-
1549
- This is the first live write-path verification of the built-in
1550
- `SpreadsheetApp` implementation. It proves the default single-`Code.gs` path can
1551
- append, replay, reject unsafe reuse, and recover a lost response without the
1552
- Advanced Sheets Service. It does not yet establish multi-user throughput or
1553
- eliminate the documented two-flush crash window between target and receipt
1554
- writes.
1555
-
1556
- Artifact: `.local/gateway-write-verification-20260804.json`.
1557
-
1558
- ## 2026-08-04 — Direct Sheets API service-account gate (blocked at auth)
1559
-
1560
- - Branch: `feature/service-account-sheets`
1561
- - Command: `node --env-file=.env scripts/bench/direct-sheets-api-service-account.mjs`
1562
- - Backend: Node direct Google Sheets API client using `GoogleAuth` and a local
1563
- service-account JSON path; Apps Script, `Code.gs`, Advanced Sheets Service, and
1564
- Drive API were not used. Credential values and spreadsheet identifiers are not
1565
- recorded.
1566
- - Setup: Stage 0 attempted one spreadsheet metadata read. No benchmark tab was
1567
- created because authentication/authorization failed before setup.
1568
-
1569
- | Stage | Result |
1570
- | --- | --- |
1571
- | Environment/credential shape | local file readable and service-account JSON shape valid |
1572
- | Stage 0 metadata/auth smoke | blocked: classified `permission` / HTTP `403` |
1573
- | Stage 1 raw append | not run |
1574
- | Stage 2 key reads | not run |
1575
- | Stage 3 postcondition | not run |
1576
- | Stage 4 response-loss replay | not run |
1577
- | Cleanup | `ok`, 0 tabs created, 0 deleted, 0 remaining |
1578
-
1579
- This run confirms that the benchmark reaches the Google Sheets API but does not
1580
- prove Direct API performance or correctness. The external setup still requires
1581
- the Sheets API to be enabled in the Service Account's Cloud project and the
1582
- target test Spreadsheet to be shared with that Service Account as an Editor.
1583
- The 403 result is preserved as an environment/permission gate, not converted to
1584
- a benchmark success or a production conclusion.
1585
-
1586
- Artifact: `.local/direct-sheets-api-2026-08-04T12-48-35-851Z-098c074b.json`.
1587
-
1588
- A retry after the external access check produced the same classified result:
1589
- Stage 0 returned `permission` / HTTP `403` after about 1.1 seconds, no temporary
1590
- tab was created, and cleanup remained `ok` with 0 created, 0 deleted, and 0
1591
- remaining tabs. The retry artifact is
1592
- `.local/direct-sheets-api-2026-08-04T13-40-14-546Z-0dfd69b2.json`.
1593
-
1594
- A redacted one-request diagnostic confirmed the API detail as
1595
- `PERMISSION_DENIED`, reason `forbidden`, message `The caller does not have
1596
- permission`; this is an access mismatch rather than a timeout, quota, or
1597
- benchmark failure.
1598
-
1599
- ## 2026-08-04 — Direct Sheets API service-account benchmark (access restored)
1600
-
1601
- - Branch: `feature/service-account-sheets`
1602
- - Command: `node --env-file=.env scripts/bench/direct-sheets-api-service-account.mjs`
1603
- - Backend: Node `@googleapis/sheets` client with `GoogleAuth` and a local
1604
- service-account JSON. The credential path was changed to the ignored local
1605
- file under `.local/credentials/`; no credential value or spreadsheet ID is
1606
- recorded here. Apps Script, `Code.gs`, Advanced Sheets Service, and Drive API
1607
- were not used.
1608
- - Dataset: 4,000 requests total across 20 cells: 1/10/100/500 rows per
1609
- request × concurrency 1/2/4/10/20, 20 requests per cell. Setup and cleanup
1610
- are reported separately from the steady-state append measurements.
1611
-
1612
- | Setup/check | Result |
1613
- | --- | --- |
1614
- | Stage 0 metadata smoke | passed in 1,033 ms; one existing tab observed |
1615
- | Temporary tabs/header setup | 21/21 tabs, 21 header appends, 12,450 ms |
1616
- | Cleanup | passed; 21/21 deleted, 0 generated tabs remaining |
1617
- | Total run | 123,057 ms; artifact exit code 0 |
1618
-
1619
- ### Stage 1 — raw append
1620
-
1621
- The table reports `successful requests / requests`, accepted rows, and the
1622
- cell-local rows/s value. A failed request was classified as `rate_limited`
1623
- (HTTP `429`); no append response anomaly was recorded.
1624
-
1625
- | Rows/request | Concurrency | Append success | Rows appended | Rows/s | Error |
1626
- | ---: | ---: | ---: | ---: | ---: | --- |
1627
- | 1 | 1 | 20/20 | 20 | 1.9 | — |
1628
- | 1 | 2 | 20/20 | 20 | 3.7 | — |
1629
- | 1 | 4 | 10/20 | 10 | 3.3 | 10 × 429 |
1630
- | 1 | 10 | 0/20 | 0 | 0.0 | 20 × 429 |
1631
- | 1 | 20 | 0/20 | 0 | 0.0 | 20 × 429 |
1632
- | 10 | 1 | 0/20 | 0 | 0.0 | 20 × 429 |
1633
- | 10 | 2 | 0/20 | 0 | 0.0 | 20 × 429 |
1634
- | 10 | 4 | 0/20 | 0 | 0.0 | 20 × 429 |
1635
- | 10 | 10 | 0/20 | 0 | 0.0 | 20 × 429 |
1636
- | 10 | 20 | 0/20 | 0 | 0.0 | 20 × 429 |
1637
- | 100 | 1 | 0/20 | 0 | 0.0 | 20 × 429 |
1638
- | 100 | 2 | 0/20 | 0 | 0.0 | 20 × 429 |
1639
- | 100 | 4 | 0/20 | 0 | 0.0 | 20 × 429 |
1640
- | 100 | 10 | 0/20 | 0 | 0.0 | 20 × 429 |
1641
- | 100 | 20 | 16/20 | 1,600 | 689.4 | 4 × 429 |
1642
- | 500 | 1 | 20/20 | 10,000 | 856.6 | — |
1643
- | 500 | 2 | 20/20 | 10,000 | 1,702.9 | — |
1644
- | 500 | 4 | 8/20 | 4,000 | 1,284.6 | 12 × 429 |
1645
- | 500 | 10 | 0/20 | 0 | 0.0 | 20 × 429 |
1646
- | 500 | 20 | 0/20 | 0 | 0.0 | 20 × 429 |
1647
-
1648
- Aggregate Stage 1: `114/400` requests succeeded, `25,650` rows were
1649
- accepted, and `286` requests were classified as `429`. Across all requests,
1650
- p50/p95/p99/max latency was `495/986/1,841/2,320 ms`; the measured aggregate
1651
- steady-state rate was `341.6 rows/s`. This is direct Sheets API quota evidence,
1652
- not evidence that concurrency 20 is production-capable.
1653
-
1654
- ### Stage 2 — key read and batch comparison
1655
-
1656
- Stage 2 issued one leading-window `values.get` and one four-range
1657
- `values.batchGet` per cell. All 20 single-range reads returned HTTP success;
1658
- seven batch reads returned classified `bad_request` / HTTP `400`. The remaining
1659
- batch reads compared keys order-independently. Key-read latency was
1660
- p50/p95/max `521/532/545 ms`; batch-read latency was `524/673/835 ms`.
1661
-
1662
- | Cell | Key window observed/missing | Batch result |
1663
- | --- | ---: | --- |
1664
- | r1_c1 | 20/0 | 20 matched, 0 missing |
1665
- | r1_c2 | 20/0 | 20 matched, 0 missing |
1666
- | r1_c4 | 10/10 | 10 matched, 10 missing |
1667
- | r1_c10 | 0/20 | 0 matched, 20 missing |
1668
- | r1_c20 | 0/20 | 0 matched, 20 missing |
1669
- | r10_c1 | 0/100 | 0 matched, 200 missing |
1670
- | r10_c2 | 0/100 | 0 matched, 200 missing |
1671
- | r10_c4 | 0/100 | 0 matched, 200 missing |
1672
- | r10_c10 | 0/100 | 0 matched, 200 missing |
1673
- | r10_c20 | 0/100 | 0 matched, 200 missing |
1674
- | r100_c1 | 0/100 | HTTP 400 |
1675
- | r100_c2 | 0/100 | HTTP 400 |
1676
- | r100_c4 | 0/100 | HTTP 400 |
1677
- | r100_c10 | 0/100 | HTTP 400 |
1678
- | r100_c20 | 100/100 | 1,600 matched, 400 missing |
1679
- | r500_c1 | 100/0 | 10,000 matched, 0 missing |
1680
- | r500_c2 | 100/0 | 10,000 matched, 0 missing |
1681
- | r500_c4 | 100/100 | HTTP 400 |
1682
- | r500_c10 | 0/100 | HTTP 400 |
1683
- | r500_c20 | 0/100 | HTTP 400 |
1684
-
1685
- The missing rows reflect the preceding Stage 1 quota failures and concurrent
1686
- partial writes; they are not treated as successful row delivery.
1687
-
1688
- ### Stage 3 — batch postcondition
1689
-
1690
- One `values.get` postcondition read was issued per cell. `4/20` cells passed
1691
- both row-count and deterministic-key checks. Postcondition latency was
1692
- p50/p95/max `524/681/697 ms`. The other cells had missing rows consistent with
1693
- Stage 1's `429` results.
1694
-
1695
- ### Stage 4 — response-loss boundary
1696
-
1697
- The full sweep's two replay appends were both unsuccessful after quota
1698
- pressure, so that sweep recorded `unexpected_count` rather than claiming
1699
- duplicate evidence. A separate five-row minimal probe was then run with the
1700
- same Direct API client and fresh temporary tab:
1701
-
1702
- | Check | Result |
1703
- | --- | --- |
1704
- | First append; response discarded | succeeded, 403 ms |
1705
- | Deterministic replay | succeeded, 390 ms |
1706
- | Readback | 10 rows, 5 unique keys, 5 duplicates |
1707
- | Verdict | `duplicate_replay` |
1708
- | Cleanup | passed; 1/1 tab deleted, 0 remaining |
1709
-
1710
- This confirms that raw `values.append` has no idempotency/receipt protection;
1711
- replaying after response loss duplicated all five rows. It is evidence for
1712
- keeping SQLite outbox, receipt, and recovery semantics outside this raw
1713
- transport, not a production-safe provider result.
1714
-
1715
- Artifacts:
1716
-
1717
- - `.local/direct-sheets-api-2026-08-04T13-53-31-116Z-33b98b28.json`
1718
- - `.local/direct-sheets-api-response-loss-2026-08-04T13-58-53-371Z-2a65266b.json`
1719
-
1720
- ## 2026-08-04 — Direct Sheets API batch-write comparison: UpdateCells vs values.batchUpdate
1721
-
1722
- - Branch: `feature/service-account-sheets`
1723
- - Command: `node --env-file=.env scripts/bench/direct-sheets-api-batch-update.mjs`
1724
- - Backend: Node `v24.3.0`, `@googleapis/sheets` client with `GoogleAuth` and a
1725
- local service-account JSON. Apps Script, `Code.gs`, Advanced Sheets Service,
1726
- and Drive API were not used. Credential values and the spreadsheet ID are
1727
- not recorded.
1728
- - Dataset: total **10,000 records × 3 string cells** (`bench_key`, `seq`,
1729
- `payload`) per scenario, split across 1/2/4/10/20 temporary tabs
1730
- (10,000/5,000/2,500/1,000/500 rows per tab). Each scenario is exactly one
1731
- write request, always sequential (request concurrency 1; no `Promise.all` or
1732
- parallel workers). Warm-up 1 (unmeasured) and measured repetitions 3 per
1733
- scenario.
1734
- - Two raw write paths on the same spreadsheet:
1735
- 1. `spreadsheets.batchUpdate` + one `UpdateCellsRequest` per tab
1736
- (`userEnteredValue.stringValue`, `fields: "userEnteredValue"`), grid
1737
- pre-sized to rows+1 × 3 columns.
1738
- 2. `spreadsheets.values.batchUpdate` + one `ValueRange` per tab
1739
- (`valueInputOption: "RAW"`), same grid.
1740
- - Setup (addSheet with pre-sized grid, one batched header write per scenario),
1741
- write round-trip latency, verification, and cleanup are recorded
1742
- separately; verification time is never part of write latency. Measured
1743
- writes are never auto-retried; `429`/timeouts/4xx/5xx would be recorded as
1744
- failures.
1745
- - **Every attempt writes its own unique deterministic key set**: the attempt
1746
- marker (`w0` warm-up, `r0`/`r1`/`r2` measured repetitions) is folded into
1747
- every key, so each attempt writes DIFFERENT keys into the same fixed
1748
- ranges. Stale data from an earlier attempt can never pass a later
1749
- attempt's verification (leftover keys surface as `missing` + `extra`); a
1750
- no-op or partial write cannot be silently masked by prior data.
1751
- - **Every measured repetition is verified individually, always**: after each
1752
- of the 30 measured writes — including after timeout/4xx/5xx/response-format
1753
- failures, because a lost-response write may still have applied — a separate
1754
- `values.batchGet` re-reads every tab and compares exact per-tab row counts
1755
- plus the deterministic key set (missing/duplicate/unexpected keys) against
1756
- THAT attempt's expected rows. Write-response outcome and verification
1757
- outcome are stored independently per repetition (`responseOk`, `verified`,
1758
- `outcome`); a repetition is a successful benchmark write only when BOTH
1759
- pass. A failed response whose data was nevertheless verified is preserved
1760
- as `write_response_failed_but_data_verified` evidence and never counted as
1761
- a success. Verification latency is recorded per repetition and never
1762
- enters write latency.
1763
- - **Scenarios are isolated and the matrix stops on contamination**: each
1764
- scenario's temporary tabs are deleted right after that scenario (before
1765
- the next one starts). If a per-scenario cleanup fails, the matrix is
1766
- aborted with an explicit `matrixAborted` reason — no later scenario ever
1767
- runs on possibly contaminated state — and the `finally` recovery cleanup
1768
- retries the leftovers. The recovery cleanup is idempotent: it resolves
1769
- remaining tabs from a metadata read (single source of truth); a generated
1770
- title the metadata does NOT list is already gone (e.g. a delete whose
1771
- response was lost after applying) and is forgotten instead of being
1772
- re-deleted, so a clean zero-remaining state never becomes a false cleanup
1773
- failure. Map ids are used as a fallback only when metadata is unavailable.
1774
- - **Scenario order is a deterministic seeded shuffle** (seed
1775
- `DIRECT_BATCH_SEED`, default `20260804`; mulberry32 + Fisher–Yates) of all
1776
- (tabCount, api) pairs — no fixed-order bias (updateCells always first, tab
1777
- counts always ascending), no uncontrolled randomness. The seed and the
1778
- exact order are recorded in the artifact. Request concurrency stays 1.
1779
- - **Signals**: SIGINT/SIGTERM request an orderly stop — the current request
1780
- settles, cleanup runs, the artifact is written, and the process exits
1781
- nonzero; cleanup is never skipped to exit faster.
1782
- - **Configuration safety**: numeric env values (`DIRECT_BATCH_TOTAL_ROWS`,
1783
- `DIRECT_BATCH_WARMUP`, `DIRECT_BATCH_REPETITIONS`, each
1784
- `DIRECT_BATCH_TAB_COUNTS` entry, `DIRECT_BATCH_SEED`) must be safe
1785
- integers; values above `Number.MAX_SAFE_INTEGER` are rejected with an
1786
- `unsafe_integer` classification before any tab is created. An unexpected
1787
- fatal error preserves the whole recorded artifact (stages, completed
1788
- scenarios, cleanup) and appends a redacted fatal classification plus a
1789
- nonzero overall verdict instead of replacing it.
1790
- - Exit code 0 only for a complete matrix: every measured write response-ok
1791
- AND post-write verified, clean setup, per-scenario cleanups and the final
1792
- recovery cleanup all complete, no interrupt. Any partial outcome (failed or
1793
- unverified write, setup/verification/cleanup problem, matrix abort, signal
1794
- interrupt) exits nonzero with a classified reason.
1795
-
1796
- ### Result
1797
-
1798
- Rerun with the corrected semantics (per-attempt unique key sets, always-run
1799
- verification, seeded order). All 30 measured writes (10 scenarios × 3
1800
- repetitions) returned HTTP 200 with matching response signals, and **every
1801
- repetition's own post-write verification passed**: per-tab row counts matched
1802
- exactly and each attempt's key set was complete with zero
1803
- missing/duplicate/unexpected keys (all repetitions `outcome: success`;
1804
- `write_response_failed_but_data_verified` never occurred). Per-repetition
1805
- verification reads took ~0.5–1.1 s each. Each scenario's tabs were deleted
1806
- before the next scenario (per-scenario cleanup 0.8–1.8 s each); the final
1807
- recovery cleanup found 0 remaining generated tabs. The run exited 0
1808
- (complete matrix: 30/30 measured writes verified). Stage 0 metadata/auth
1809
- smoke took 1,154 ms; the whole run took 125,731 ms.
1810
-
1811
- Actual scenario order (seed 20260804): `valuesBatchUpdate@20` →
1812
- `updateCells@20` → `updateCells@4` → `valuesBatchUpdate@4` →
1813
- `valuesBatchUpdate@1` → `valuesBatchUpdate@10` → `updateCells@10` →
1814
- `updateCells@1` → `updateCells@2` → `valuesBatchUpdate@2`.
1815
-
1816
- | API | Tabs | Rows/tab | Payload bytes | Rep latencies (ms) | p50 / p95 / max (ms) | Rows/s | Cells/s |
1817
- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
1818
- | `values.batchUpdate` | 20 | 500 | 846,765 | 2,090 / 2,220 / 2,272 | 2,220 / 2,272 / 2,272 | 4,558 | 13,673 |
1819
- | `updateCells` | 20 | 500 | 2,067,308 | 2,816 / 3,419 / 3,361 | 3,361 / 3,419 / 3,419 | 3,126 | 9,379 |
1820
- | `updateCells` | 4 | 2,500 | 2,060,473 | 1,968 / 2,491 / 3,241 | 2,491 / 3,241 / 3,241 | 3,896 | 11,689 |
1821
- | `values.batchUpdate` | 4 | 2,500 | 840,379 | 1,640 / 2,301 / 1,785 | 1,785 / 2,301 / 2,301 | 5,238 | 15,715 |
1822
- | `values.batchUpdate` | 1 | 10,000 | 840,122 | 2,783 / 1,507 / 2,527 | 2,527 / 2,783 / 2,783 | 4,401 | 13,204 |
1823
- | `values.batchUpdate` | 10 | 1,000 | 840,905 | 1,312 / 1,538 / 1,317 | 1,317 / 1,538 / 1,538 | 7,200 | 21,601 |
1824
- | `updateCells` | 10 | 1,000 | 2,061,157 | 2,071 / 2,248 / 1,902 | 2,071 / 2,248 / 2,248 | 4,822 | 14,467 |
1825
- | `updateCells` | 1 | 10,000 | 2,060,128 | 1,899 / 2,141 / 1,850 | 1,899 / 2,141 / 2,141 | 5,093 | 15,280 |
1826
- | `updateCells` | 2 | 5,000 | 2,060,243 | 1,665 / 1,602 / 1,699 | 1,665 / 1,699 / 1,699 | 6,041 | 18,124 |
1827
- | `values.batchUpdate` | 2 | 5,000 | 840,207 | 1,454 / 1,545 / 1,761 | 1,545 / 1,761 / 1,761 | 6,303 | 18,908 |
1828
-
1829
- Per-scenario setup (tabs + headers) took 824–1,610 ms; per-repetition
1830
- verification reads took ~0.5–1.1 s each and all 30 passed (per-scenario
1831
- verification totals: 1.7–2.4 s for its 3 repetitions). Per-scenario cleanup
1832
- (delete + remaining-tab check) took 0.8–1.8 s each. Warm-up latencies
1833
- followed the same pattern as the measured repetitions. Rows/s is
1834
- `(successful repetitions × 10,000) / sum of successful repetition durations`
1835
- — total proven rows over the summed round-trip time of those repetitions;
1836
- cells/s is rows/s × 3.
1837
-
1838
- ### Reading the result
1839
-
1840
- - **`values.batchUpdate` is the faster and smaller path again**: its payload
1841
- is about 840 KB versus 2.06 MB for `updateCells` (per-cell JSON envelope
1842
- cost), and it was faster in 7 of 10 scenario cells (1.31–2.78 s vs
1843
- 1.60–3.42 s round-trip; 4,401–7,200 rows/s vs 3,126–6,041 rows/s).
1844
- - **Tab-count effects remain modest and noisy**: this run's fastest cell was
1845
- `values.batchUpdate` with 10 tabs (7,200 rows/s) and its slowest was
1846
- `updateCells` with 20 tabs (3,126 rows/s), while the single-tab cells sat
1847
- mid-pack — a different relative pattern from the previous fixed-order run
1848
- (where single-tab was slowest for both APIs). The comparison is dominated
1849
- by network/backend variance between scenario cells, not by a hard tab
1850
- count rule; three repetitions per cell cannot prove one.
1851
- - This is a raw-transport comparison on one spreadsheet, one run, three
1852
- repetitions per cell, each repetition individually verified against its
1853
- own unique key set. It measures the Google Sheets REST path only and says
1854
- nothing about the Hikoutei worker/coordinator drain rate, outbox, receipts,
1855
- CAS, or Apps Script gateway behavior. The 10,000-row request is the
1856
- largest tested here; larger sizes were not attempted.
1857
-
1858
- ### Quota observation
1859
-
1860
- A first run of this benchmark (before the setup batching below) hit HTTP `429`
1861
- on 4 of 8 cleanup `deleteSheet` chunks after roughly 120 requests, stranding
1862
- 34 tabs; the stranded tabs were deleted by a paced follow-up. The fixed rerun
1863
- reduced setup writes from ~2 per tab to O(1) per scenario (grid pre-sized in
1864
- `addSheet`, headers written with one `values.batchUpdate`). The corrected run
1865
- above added per-repetition verification reads (30 `batchGet` calls) and
1866
- per-scenario cleanup (10 delete rounds plus remaining-tab checks) and still
1867
- completed with zero `429`s, zero verification failures, and zero stranded
1868
- tabs. Per-scenario cleanup also bounds the damage of an interrupt: at most the
1869
- in-progress scenario's tabs can remain, and the signal handler runs the
1870
- recovery cleanup before exiting nonzero. One intermediate rerun (16:54, see
1871
- artifacts) hit the always-run verification gate with a verification
1872
- normalization bug (raw `ValueRange` objects compared instead of their
1873
- `values` arrays) and correctly exited 1 with `0 of 30 measured writes
1874
- verified` rather than reporting a false success. This is observed quota
1875
- pressure evidence, not a remaining-quota counter; Sheets REST write quota is
1876
- per-minute per project and must be checked in the Cloud console.
1877
-
1878
- ### How to run
1879
-
1880
- ```sh
1881
- node --env-file=.env scripts/bench/direct-sheets-api-batch-update.mjs
1882
- ```
1883
-
1884
- Required environment: `GOOGLE_APPLICATION_CREDENTIALS` (service-account JSON
1885
- readable locally) and `GOOGLE_SHEETS_TEST_SPREADSHEET_ID` (the service account
1886
- must be shared on that spreadsheet as an Editor). Optional knobs:
1887
-
1888
- - `DIRECT_BATCH_TAB_COUNTS=1,2,4,10,20` — tab-count matrix (each count must
1889
- divide the total rows evenly)
1890
- - `DIRECT_BATCH_TOTAL_ROWS=10000` — records per scenario
1891
- - `DIRECT_BATCH_WARMUP=1` — unmeasured warm-up requests per scenario
1892
- - `DIRECT_BATCH_REPETITIONS=3` — measured repetitions per scenario
1893
- - `DIRECT_BATCH_SEED=20260804` — deterministic scenario-order shuffle seed
1894
- (unsigned 32-bit integer; the seed and actual order are recorded in the
1895
- artifact)
1896
-
1897
- Invalid env values (non-integers, values above `Number.MAX_SAFE_INTEGER`,
1898
- non-positive counts, uneven tab splits, out-of-range seeds) are classified
1899
- and abort before any tab is created. If authentication or sharing is blocked,
1900
- Stage 0 fails with a classified `auth`/`permission` result and cleanup stays
1901
- trivially clean; the run is never reported as a benchmark success in that
1902
- case. The exit code is 0 only when the complete matrix ran: every measured
1903
- write response-ok AND post-write verified, setup clean, per-scenario cleanups
1904
- and final recovery cleanup complete, no signal interrupt; anything less exits
1905
- 1 with a classified reason in the artifact.
1906
-
1907
- Artifacts (raw JSON, in the ignored `.local/` directory):
1908
-
1909
- - `.local/direct-sheets-api-batch-update-2026-08-04T16-57-19-098Z-65c1e7a5.json` (clean run above: per-attempt unique keys, always-run verification, seeded order)
1910
- - `.local/direct-sheets-api-batch-update-2026-08-04T16-54-50-919Z-0c3ffca3.json` (superseded: verification normalization bug — every write correctly failed the always-run verification gate, exit 1, no false success)
1911
- - `.local/direct-sheets-api-batch-update-2026-08-04T16-21-51-910Z-7013fc02.json` (superseded: pre-fix semantics — fixed order, shared key set across repetitions, verification only after response-ok writes)
1912
- - `.local/direct-sheets-api-batch-update-2026-08-04T15-59-24-945Z-b7134a7a.json` (superseded: pre-fix semantics — single aggregate verification, rows/s under-reported 3×)
1913
- - `.local/direct-sheets-api-batch-update-2026-08-04T15-49-31-777Z-ba3d0297.json` (first run: pre-fix `values.batchUpdate` range bug + cleanup 429 evidence)
1914
-
1915
- ## 2026-08-05 — Integrated Apps Script `appendTestBatch` load
1916
-
1917
- - Branch: `feature/service-account-sheets`
1918
- - Command: `node --env-file=.env --env-file=.local/gateway-test-override.env .local/run-append-load.mjs`
1919
- - Backend: deployed Apps Script Gateway with the Advanced Sheets v4 service;
1920
- credentials and spreadsheet identifiers are intentionally not recorded.
1921
- - Dataset: **10,000 entity inserts** into one temporary `System_State` tab,
1922
- one SQLite outbox/runtime, sequential worker dispatch, 1,000 append rows per
1923
- request, and a 1,100 ms minimum interval between append request starts.
1924
- Updates/deletes were not included.
1925
- - Setup, SQLite seed/flush, worker upload, verification, and cleanup were
1926
- measured separately. The worker made 10 successful append requests plus
1927
- provisioning, verification, and cleanup requests; no gateway request failed.
1928
-
1929
- | Phase | Duration |
1930
- | --- | ---: |
1931
- | Runtime setup/provisioning | 4.68 s |
1932
- | SQLite seed + outbox flush | 4.52 s |
1933
- | Worker append drain | 798.10 s (13m 18.10s) |
1934
- | Remote verification | 3.56 s |
1935
- | Remote cleanup | 2.66 s |
1936
-
1937
- Append request durations were 51.43–119.89 s (p50 75.90 s, p95/max
1938
- 119.89 s). All 10,000 rows were applied and remotely verified; cleanup left no
1939
- load tabs. The measured integrated drain rate was approximately **12.53
1940
- rows/s**, despite the 1,000-row API batches. This is materially slower than the
1941
- raw service-account `values.batchUpdate` result above because the integrated
1942
- path also pays Apps Script execution, receipt handling, identity/postcondition
1943
- scans, gateway transport, and worker lease overhead. This run confirms
1944
- correctness and batching, but does **not** show that the end-to-end upload
1945
- bottleneck has been eliminated.
1946
-
1947
- Artifact:
1948
-
1949
- - `.local/append-test-load-msfndo05-ecd2cf98.json`
1950
-
1951
- Caveat: this is one sequential run on one spreadsheet; Apps Script cold starts,
1952
- quota pressure, network latency, and the current receipt/postcondition work can
1953
- vary substantially. The authority fence was reset on the dedicated test
1954
- spreadsheet before the run; no production spreadsheet should be reset this
1955
- way.
1956
-
1957
- ## 2026-08-05 — Direct Google Sheets API outbound worker load
1958
-
1959
- - Branch: `feature/service-account-sheets`
1960
- - Command: `node --env-file=.env .local/run-append-load-direct.mjs`
1961
- - Backend: direct Google Sheets REST API outbound worker (service account via
1962
- Application Default Credentials, `GOOGLE_APPLICATION_CREDENTIALS`) against
1963
- the shared test spreadsheet; SQLite outbox/worker state machine unchanged;
1964
- Apps Script gateway not involved in the write path (provisioning of the
1965
- temporary tab was done by the service account itself). Credentials and
1966
- spreadsheet identifiers are intentionally not recorded.
1967
- - Dataset: **10,000 entity inserts** into one temporary `System_State` tab
1968
- (`id`, `status`, `__typed_sheets_deleted`), one SQLite outbox/runtime,
1969
- sequential worker dispatch, 1,000 append rows per `batchUpdate`, the worker
1970
- append throttle (1,100 ms between append request starts) and the provider's
1971
- read/write request-start limiters (1,100 ms per class) all active.
1972
- Updates/deletes were not included.
1973
- - Setup, SQLite seed/flush, worker upload, verification, and cleanup were
1974
- measured separately. The worker made 10 successful append batches plus
1975
- per-batch preflight reads, verification, and cleanup; no provider request
1976
- failed (20/20 requests ok).
1977
-
1978
- | Phase | Duration |
1979
- | --- | ---: |
1980
- | Runtime setup/provisioning | 3.08 s |
1981
- | SQLite seed + outbox flush | 4.44 s |
1982
- | Worker append drain | 39.47 s |
1983
- | Remote verification | 0.50 s |
1984
- | Remote cleanup | 1.06 s |
1985
-
1986
- Append `batchUpdate` durations were 0.99–1.59 s (p50 1.45 s, max 1.59 s);
1987
- per-batch preflight reads (sheet enumeration + grid data, two calls per batch
1988
- spaced by the 1,100 ms read limiter) were 1.21–2.03 s. All 10,000 rows were
1989
- applied (10 batches × 1,000), zero failures/deferred/requeues, and the read-back
1990
- verification confirmed 10,000 rows with the expected first/last keys; cleanup
1991
- left no load tabs and removed the receipt tab.
1992
-
1993
- ### Comparison with the previous integrated run (2026-08-05, same spreadsheet
1994
- class, same 10,000 rows, same 1,000-row batch and 1,100 ms interval):
1995
-
1996
- | Metric | Apps Script gateway | Direct Sheets API | Delta |
1997
- | --- | ---: | ---: | ---: |
1998
- | Worker append drain | 798.10 s | 39.47 s | **20.2× faster** |
1999
- | Integrated throughput | 12.53 rows/s | 253.4 rows/s | +20.2× |
2000
- | Per-batch latency p50 | 75.90 s | 1.45 s | 52× lower |
2001
- | Per-batch latency max | 119.89 s | 1.59 s | 75× lower |
2002
- | Failed/uncertain effects | 0 | 0 | same |
2003
-
2004
- Steady-state (drain-only) result: **10,000 rows in 39.47 s ≈ 253 rows/s**;
2005
- excluding the 1,100 ms request-start pacing of the worker throttle and the
2006
- provider read/write limiters, the raw `batchUpdate` time alone was ~13 s for
2007
- all 10 batches (1.3 s/batch). The drain time is dominated by the deliberate
2008
- quota pacing, not by Sheets latency.
2009
-
2010
- Artifact:
2011
-
2012
- - `.local/append-test-load-direct-msfuq2yb-3a862d7a.json`
2013
-
2014
- Caveats: one sequential run on one spreadsheet; a provider bug found during
2015
- this live test (receipt-tab discovery: `spreadsheets.get` with `ranges`
2016
- returns only intersecting sheets, so the second batch could not see the
2017
- receipt tab created by the first and attempted to recreate it, failing the
2018
- batch with 400 INVALID_ARGUMENT) was fixed in the provider before this run
2019
- (2-call preflight: sheet enumeration without ranges, then target+receipt data
2020
- ranges) and is covered by regression tests. Results are not a claim about
2021
- quota headroom at higher concurrency, update/delete throughput, or
2022
- multi-writer behavior; the direct provider additionally paces reads at
2023
- 1,100 ms, so observation-heavy workloads were not measured here.
2024
-
2025
- ## 2026-08-05 — Direct outbound update/delete correctness gate (passed)
2026
-
2027
- - Branch: `feature/service-account-sheets`
2028
- - Command: `node --env-file=.env .local/run-update-delete-direct.mjs`
2029
- - Backend: direct Google Sheets API outbound worker (service account via ADC)
2030
- on the shared test spreadsheet; two projections per entity (`system_state`
2031
- tab A:C and `user_input` tab A:B), one SQLite outbox/runtime, no Apps
2032
- Script involvement. Credentials and spreadsheet identifiers are not
2033
- recorded.
2034
- - Dataset: 100 entities (id/status), seeded through the mapped ORM, then 50
2035
- bulk updates, two simulated human-edit CAS scenarios, and 25 entity
2036
- deletions, each step drained by the real effect worker through the direct
2037
- provider and verified by service-account read-back.
2038
-
2039
- | Step | Result |
2040
- | --- | --- |
2041
- | Seed + append (100 system + 100 input rows) | 200 effects applied; both tabs 100 rows |
2042
- | Bulk update 50 → "paid" | 100 effects applied; both tabs exactly 50 "paid" |
2043
- | Human edit on user_input tab + local update | input mirror effect parked as `blocked_candidate`/`candidate_guard_mismatch`; sheet keeps human value; system tab applied "paid-v2" |
2044
- | Human edit on system tab + local update | system effect recorded as durable `conflict`/`visible_guard_mismatch`; sheet keeps "human-sys-edit"; input mirror applied |
2045
- | Delete 25 entities | 25 system rows flagged `__typed_sheets_deleted=true`, 25 user_input rows physically removed; 75 rows remain; human-edited rows preserved |
2046
- | Cleanup | passed; load tabs + receipt tab removed |
2047
-
2048
- This gate confirms the direct provider's guarded update (visible-hash CAS
2049
- with the same conflict semantics as the gateway path: `blocked_candidate` for
2050
- candidate-protected user_input effects, `conflict` for system projection
2051
- effects), physical delete with full-row guard, receipt-backed evidence, and
2052
- per-step worker/outbox bookkeeping against the real Sheets API. It is a
2053
- correctness gate, not a throughput measurement; delete/update volume here is
2054
- small (≤100 effects per step) and paced by the same 1,100 ms request-start
2055
- limiters.
2056
-
2057
- Artifact:
2058
-
2059
- - `.local/update-delete-direct-msfv3rd0-72dc5b6e.json`
2060
-
2061
- Caveats: single run on one spreadsheet; the guard scenarios mutate rows
2062
- outside the worker and rely on the worker's existing conflict semantics, so
2063
- their expected statuses are `blocked_candidate`/`conflict`, not failures.
2064
-
2065
- ## 2026-08-05 — Direct full-provider live correctness gate (passed)
2066
-
2067
- - Branch: `feature/service-account-sheets`
2068
- - Command: `node --env-file=.env scripts/ci/run-api-scenario.mjs --backend=live --outbound=direct`
2069
- - Backend: the full `GoogleSheetsApiSyncProvider` (service account via ADC)
2070
- on the shared test spreadsheet, provisioning from an EMPTY spreadsheet state
2071
- (no tabs) through the tracked scenario. No Apps Script gateway, no
2072
- `TYPED_SHEETS_GATEWAY_*` secrets. Credentials and spreadsheet identifiers
2073
- are never recorded; the report keeps only `sheetMatched: true`.
2074
-
2075
- | Step | Result |
2076
- | --- | --- |
2077
- | Provisioning from empty spreadsheet | 3 tabs created + headers initialized (one atomic batch), `createdSheets`/`initializedHeaders` asserted |
2078
- | ORM create → worker append | system fast append + user_input create; receipt tab evidence (2 receipt rows, no duplicates) |
2079
- | Guarded update + read-back | ORM update → drain → sheet value verified |
2080
- | Mapped polling of SA-simulated human edit | full observation → SQLite entity reflects `edited-by-user` (appliedRows 1) |
2081
- | Stale input edit | input mirror effect parked `blocked_candidate`/`candidate_guard_mismatch`; sheet keeps `human-2` |
2082
- | Stale system edit | system effect recorded as durable `conflict`/`visible_guard_mismatch`; sheet keeps `human-sys` |
2083
- | Delete | system row tombstoned (`__typed_sheets_deleted=true`), user_input row physically removed; 3 system rows / 2 input rows remain |
2084
- | Anchor evidence | user_input rows anchored (polling observation pass), system rows identity-located, no duplicate materialization (unique ids) |
2085
- | Cleanup | passed; fixture tabs + receipt tab removed |
2086
-
2087
- Totals: 19 steps, 35 assertions, all passed. Setup 1.61 s; steady-state
2088
- drain/observation/verification 39.41 s; total 42.06 s (paced by the same
2089
- 1,100 ms per-request-class limiters). Artifact:
2090
-
2091
- - `.local/live-direct-parity-2026-08-05T20-33-29.json`
2092
-
2093
- This gate verifies the full-provider parity checklist (R9 of the promotion
2094
- plan) against the real Sheets API: provisioning from empty state, append /
2095
- update / delete with receipts, mapped User_Input polling into SQLite, both
2096
- CAS guard paths, anchor evidence, and cleanup — with zero Apps Script
2097
- involvement. It is a correctness gate, not a throughput measurement.
2098
-
2099
- Caveats: single sequential run on one spreadsheet; the tracked workflow
2100
- (`.github/workflows/live-integration.yml`) runs the same scenario
2101
- service-account-only via `GOOGLE_SERVICE_ACCOUNT_CREDENTIALS_JSON` + `GOOGLE_SHEETS_TEST_SPREADSHEET_ID`.
2102
-
2103
- ## Direct full-provider parity (tracked live scenario)
2104
-
2105
- No new throughput measurement was run for the full-provider promotion; this
2106
- document records only evidence that was actually measured. The full direct
2107
- provider (`googleSheetsApi`) is exercised live by the tracked scenario
2108
- (`scripts/ci/run-api-scenario.mjs --backend live --outbound direct`), which is
2109
- the correctness gate for provisioning from an empty spreadsheet, append /
2110
- update / delete delivery, receipt evidence, mapped User_Input polling,
2111
- stale-edit CAS guards (`blocked_candidate` / `visible_guard_mismatch`), anchor
2112
- evidence, and cleanup — all through the service account with no Apps Script.
2113
- The unguarded raw-transport throughput numbers above do not include receipts
2114
- or compare-and-set, so they are not provider throughput claims. Live artifacts
2115
- from the update/delete correctness gate live under `.local/`
2116
- (`.local/update-delete-direct-*.json`), and the workflow report is uploaded as
2117
- the `hikoutei-live-google-sheets-direct` artifact.