jskelet 0.1.1 → 0.1.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +5 -0
- package/CHANGELOG.md +129 -2
- package/README.md +21 -7
- package/bin/jskelet.mjs +6 -6
- package/docs/03-routing.md +48 -9
- package/docs/04-render-ve-sablonlar.md +2 -2
- package/docs/05-islands.md +59 -6
- package/docs/06-cache.md +240 -26
- package/docs/07-yapilandirma.md +108 -7
- package/docs/08-build.md +4 -4
- package/docs/09-dev-araclari.md +5 -0
- package/docs/12-panel-ve-oturum.md +384 -0
- package/docs/README.md +25 -2
- package/docs/en/01-getting-started.md +292 -0
- package/docs/en/02-architecture.md +305 -0
- package/docs/en/03-routing.md +493 -0
- package/docs/en/04-rendering.md +504 -0
- package/docs/en/05-islands.md +492 -0
- package/docs/en/06-caching.md +640 -0
- package/docs/en/07-configuration.md +789 -0
- package/docs/en/08-build.md +383 -0
- package/docs/en/09-dev-tools.md +314 -0
- package/docs/en/10-deployment.md +332 -0
- package/docs/en/11-migration.md +360 -0
- package/docs/en/12-dashboards-and-sessions.md +392 -0
- package/docs/en/README.md +112 -0
- package/package.json +4 -2
- package/src/build/build.mjs +1 -1
- package/src/build/tasks/client.mjs +2 -2
- package/src/build/tasks/fonts.mjs +3 -3
- package/src/build/tasks/icons.mjs +1 -1
- package/src/build/tasks/images.mjs +2 -2
- package/src/client/devtools/overlay.js +196 -164
- package/src/client/devtools/report.js +96 -96
- package/src/client/form.js +192 -0
- package/src/client/index.js +10 -1
- package/src/client/registry.js +78 -4
- package/src/client/swap.js +188 -0
- package/src/config/defaults.js +83 -0
- package/src/config/index.js +129 -18
- package/src/config/pattern.js +1 -1
- package/src/dev-server.mjs +1 -1
- package/src/http/control-flow.js +16 -1
- package/src/http/cookies.js +257 -0
- package/src/http/request-context.js +162 -0
- package/src/index.js +26 -2
- package/src/init.mjs +32 -31
- package/src/log.mjs +8 -2
- package/src/logo.png +0 -0
- package/src/runtime/alias-hooks.mjs +1 -1
- package/src/server/assets.js +1 -1
- package/src/server/create-app.js +12 -4
- package/src/server/data-cache.js +244 -0
- package/src/server/dev/devtools.js +6 -2
- package/src/server/dev/report.js +8 -1
- package/src/server/dev/version-check.mjs +139 -0
- package/src/server/head-hints.js +1 -1
- package/src/server/html-cache.js +32 -6
- package/src/server/middleware/csrf.js +134 -0
- package/src/server/prewarm.js +164 -19
- package/src/server/render.js +256 -20
- package/src/server/router.js +14 -7
- package/src/server/status-page.js +1 -1
- package/src/version.mjs +9 -4
- package/src/views/components/loader.js +1 -1
- package/src/views/helpers/tags.js +53 -1
|
@@ -0,0 +1,640 @@
|
|
|
1
|
+
# 06 — Caching and prewarm
|
|
2
|
+
|
|
3
|
+
This document explains JSkelet's ISR substitute in full detail: the HTML TTL
|
|
4
|
+
cache and its stale-while-revalidate behaviour, where `revalidate` comes from,
|
|
5
|
+
how the cache key is built, the values of the `X-JSkelet-Cache` header, why the
|
|
6
|
+
compressed body is kept in the cache, per-request memoization
|
|
7
|
+
(`withRequestCache` / `cache()`), the data cache (`withDataCache`), how upstream
|
|
8
|
+
failures affect the cache (`reportUpstreamFailure`) and the prewarm round at
|
|
9
|
+
server startup. The
|
|
10
|
+
measurement rationale behind the decisions is in
|
|
11
|
+
[02-architecture.md](./02-architecture.md), and the full reference of config
|
|
12
|
+
fields is in [07-configuration.md](./07-configuration.md).
|
|
13
|
+
|
|
14
|
+
## The big picture
|
|
15
|
+
|
|
16
|
+
```
|
|
17
|
+
route(controller, { revalidate })
|
|
18
|
+
└─ withHtmlCache(key, ttl, producer) ← TTL + stale-while-revalidate
|
|
19
|
+
└─ withUpstreamTracking(...) ← missing data detection
|
|
20
|
+
└─ withRequestCache(...) ← per-request memoization
|
|
21
|
+
└─ produce() → controller + renderPage
|
|
22
|
+
└─ withDataCache(...) ← upstream data cache
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
The order matters: the **per-request cache must be innermost** so that two
|
|
26
|
+
calls in the same render collapse into a single upstream request; **upstream
|
|
27
|
+
tracking must be inside the HTML cache** so that output produced with missing
|
|
28
|
+
data is not written to the cache.
|
|
29
|
+
|
|
30
|
+
How the two caches divide the work:
|
|
31
|
+
|
|
32
|
+
| | HTML cache | Data cache |
|
|
33
|
+
| --- | --- | --- |
|
|
34
|
+
| What it holds | The whole page (+ its compressed body) | The JSON that came from upstream |
|
|
35
|
+
| Entry size | ~100-200 kB | ~1-20 kB |
|
|
36
|
+
| Entry limit | 500 (`cache().maxEntries`) | 10,000 (`cache().data.maxEntries`) |
|
|
37
|
+
| Who benefits | Pages with traffic: not even rendered | The long tail: rendered, but without going to the API |
|
|
38
|
+
|
|
39
|
+
In practice this distinction means: on a site with tens of thousands of paths it
|
|
40
|
+
is impossible to keep every page hot as HTML — a warm-up that goes past 500
|
|
41
|
+
entries deletes what it just warmed. For the long tail the goal is not "have the
|
|
42
|
+
HTML ready" but **"have the data that produces the page available without going
|
|
43
|
+
to the API"**. Then a page that was never warmed is also produced within
|
|
44
|
+
milliseconds on the first visit, and spends no quota.
|
|
45
|
+
|
|
46
|
+
## Public versus per-visitor
|
|
47
|
+
|
|
48
|
+
Everything in this document applies to HTML that **can go to everyone
|
|
49
|
+
unchanged**. There is no identity in the cache key (only path + query), so a
|
|
50
|
+
page in the cache is the answer for that path, not the answer for whoever asked
|
|
51
|
+
for it first.
|
|
52
|
+
|
|
53
|
+
A page that depends on the user therefore takes a separate path:
|
|
54
|
+
|
|
55
|
+
```js
|
|
56
|
+
app.get("/dashboard", route(async ({ req }) => { … }, { private: true }));
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
`private: true` does three things at once: the cache is disabled, a
|
|
60
|
+
`cache.html` pattern **cannot** override that decision, and the response is
|
|
61
|
+
sent with `private, no-store` and `Vary: Cookie`, without an ETag. The details
|
|
62
|
+
and the session/CSRF side are in
|
|
63
|
+
[12-dashboards-and-sessions.md](./12-dashboards-and-sessions.md).
|
|
64
|
+
|
|
65
|
+
If you forget the flag, the framework does not stay quiet: as soon as the
|
|
66
|
+
controller reads `Cookie`, `Authorization` or `req.session`/`req.user`, the
|
|
67
|
+
render is marked and **not written** to the cache. In development the request
|
|
68
|
+
fails with an explanation, in production it is served with `no-store` and
|
|
69
|
+
logged. The guard is a last line of defence, not an excuse — the right place is
|
|
70
|
+
`private: true`.
|
|
71
|
+
|
|
72
|
+
## `revalidate` — where the TTL comes from
|
|
73
|
+
|
|
74
|
+
A route's TTL can come from two sources, and **the config wins**:
|
|
75
|
+
|
|
76
|
+
1. `route(controller, { revalidate: 60 })` — the route's own duration.
|
|
77
|
+
2. The matching pattern inside `jskelet.config.mjs` → `cache().html`. If it
|
|
78
|
+
exists it overrides the route's value.
|
|
79
|
+
|
|
80
|
+
The one exception is `private: true`: a matching pattern is ignored. The lock is
|
|
81
|
+
one-way, because a mistake in the other direction means a silent data leak.
|
|
82
|
+
|
|
83
|
+
```js
|
|
84
|
+
// jskelet.config.mjs
|
|
85
|
+
export default {
|
|
86
|
+
async cache() {
|
|
87
|
+
return {
|
|
88
|
+
html: {
|
|
89
|
+
"/": 60,
|
|
90
|
+
"/news/:slug": 300,
|
|
91
|
+
"/tag/:slug": 120,
|
|
92
|
+
},
|
|
93
|
+
};
|
|
94
|
+
},
|
|
95
|
+
};
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
Overriding from the config makes it possible to tune the freshness profile of
|
|
99
|
+
the whole site from a single file; you do not have to walk through the route
|
|
100
|
+
files.
|
|
101
|
+
|
|
102
|
+
The resolution result is **remembered per path**, so a pattern scan is not done
|
|
103
|
+
on every request. If there is no `cache().html` rule at all, the route's own
|
|
104
|
+
value is used directly.
|
|
105
|
+
|
|
106
|
+
If `revalidate` is not given, or is 0, the page is **not cached at all**: every
|
|
107
|
+
request is rendered and the response is sent with
|
|
108
|
+
`Cache-Control: private, no-store` and no ETag. No `X-JSkelet-Cache` header is
|
|
109
|
+
written either — the cache path never ran, so `MISS` would be misleading.
|
|
110
|
+
|
|
111
|
+
Sending `no-store` on a dynamic page is deliberate. HTTP treats a response with
|
|
112
|
+
no directives as "heuristically cacheable", so an intermediate proxy or the
|
|
113
|
+
browser's back button could store HTML produced for a single visitor.
|
|
114
|
+
|
|
115
|
+
The cache also only kicks in for `GET` requests.
|
|
116
|
+
|
|
117
|
+
## The cache key
|
|
118
|
+
|
|
119
|
+
```
|
|
120
|
+
`${req.path}?${new URLSearchParams(query).toString()}`
|
|
121
|
+
```
|
|
122
|
+
|
|
123
|
+
So the path **and all query parameters** are part of the key. `/list?page=2`
|
|
124
|
+
and `/list?page=3` are separate entries.
|
|
125
|
+
|
|
126
|
+
The practical consequence: a page that does not depend on the query string
|
|
127
|
+
produces a separate entry for every combination when it is called with
|
|
128
|
+
different campaign parameters (`?utm_source=…`). Stripping such parameters at
|
|
129
|
+
the reverse proxy layer, or turning off the cache (by not supplying
|
|
130
|
+
`revalidate`), is a reasonable precaution; by default the store holds at most
|
|
131
|
+
500 entries and evicts the oldest with LRU.
|
|
132
|
+
|
|
133
|
+
## Stale-while-revalidate
|
|
134
|
+
|
|
135
|
+
The entry structure:
|
|
136
|
+
|
|
137
|
+
```
|
|
138
|
+
expiresAt = now + ttl
|
|
139
|
+
staleUntil = now + ttl * 2 (STALE_FACTOR = 1)
|
|
140
|
+
```
|
|
141
|
+
|
|
142
|
+
Read behaviour:
|
|
143
|
+
|
|
144
|
+
| State | Response | Background |
|
|
145
|
+
| --- | --- | --- |
|
|
146
|
+
| `now < expiresAt` | The cached HTML, `HIT` | — |
|
|
147
|
+
| `expiresAt ≤ now < staleUntil` | The cached HTML **immediately**, `STALE` | A refresh is started |
|
|
148
|
+
| `now ≥ staleUntil` | The entry is deleted, fresh render, `MISS` | — |
|
|
149
|
+
|
|
150
|
+
A failure of the refresh inside the stale window does not affect the request:
|
|
151
|
+
the old HTML stays valid for the whole window and the error is only logged
|
|
152
|
+
(`[html-cache] background refresh failed: …`).
|
|
153
|
+
|
|
154
|
+
Concurrent refreshes for the same key are collapsed into a single run (the
|
|
155
|
+
`inflight` map): a hundred concurrent requests fall to one render.
|
|
156
|
+
|
|
157
|
+
The gain: after the first warm-up no request waits for a render. The price: the
|
|
158
|
+
data in the HTML can be at most `revalidate + one refresh round` behind. That
|
|
159
|
+
price is acceptable, because live fields such as prices are updated on the
|
|
160
|
+
client over WebSocket.
|
|
161
|
+
|
|
162
|
+
The store is an LRU: an accessed entry is moved to the end, and once the limit
|
|
163
|
+
(`cache().maxEntries`, 500 by default) is exceeded the oldest is evicted.
|
|
164
|
+
|
|
165
|
+
## What gets written to the cache
|
|
166
|
+
|
|
167
|
+
Only output that satisfies **both** of these two conditions is stored:
|
|
168
|
+
|
|
169
|
+
1. `status === 200`
|
|
170
|
+
2. `degraded !== true` — no transient upstream failure was reported during the
|
|
171
|
+
render.
|
|
172
|
+
|
|
173
|
+
So 404 pages, redirects and HTML produced with missing data do not enter the
|
|
174
|
+
cache.
|
|
175
|
+
|
|
176
|
+
## Response headers
|
|
177
|
+
|
|
178
|
+
`route()` writes `X-JSkelet-Cache` on every response (the header name can be
|
|
179
|
+
changed with `brand.cacheHeader`):
|
|
180
|
+
|
|
181
|
+
| Value | Meaning |
|
|
182
|
+
| --- | --- |
|
|
183
|
+
| `HIT` | From the cache, fresh |
|
|
184
|
+
| `STALE` | From the cache, expired; being refreshed in the background |
|
|
185
|
+
| `MISS` | Rendered on this request (or the cache is off) |
|
|
186
|
+
|
|
187
|
+
On cacheable responses, additionally:
|
|
188
|
+
|
|
189
|
+
```
|
|
190
|
+
Cache-Control: public, max-age=0, s-maxage=<revalidate>, stale-while-revalidate=60
|
|
191
|
+
```
|
|
192
|
+
|
|
193
|
+
`max-age=0` turns off storage in the browser, `s-maxage` announces the duration
|
|
194
|
+
to intermediate layers (CDN, reverse proxy). This way, when a CDN sits in
|
|
195
|
+
front, the same freshness model works across both layers together.
|
|
196
|
+
|
|
197
|
+
## Storing the compressed body
|
|
198
|
+
|
|
199
|
+
Every cached entry carries an `encoded` map and shares the same lifetime as the
|
|
200
|
+
HTML. The first time a page is requested with brotli or gzip the output is
|
|
201
|
+
computed and put in the map; on subsequent requests the same buffer is sent.
|
|
202
|
+
The same page is not re-brotli'd on every request.
|
|
203
|
+
|
|
204
|
+
On this path `Content-Encoding`, `Vary` and `Content-Length` are written
|
|
205
|
+
directly by `route()`; the compression middleware does not kick in because it
|
|
206
|
+
sees `Content-Encoding`.
|
|
207
|
+
|
|
208
|
+
`HEAD` requests are not compressed (there is no body). If the client accepts
|
|
209
|
+
neither brotli nor gzip, plain HTML is sent.
|
|
210
|
+
|
|
211
|
+
## Per-request memoization: `cache()`
|
|
212
|
+
|
|
213
|
+
The equivalent of React's `cache()` function: calls made with the same
|
|
214
|
+
arguments within the same request run only once.
|
|
215
|
+
|
|
216
|
+
```js
|
|
217
|
+
// lib/api/articles.js
|
|
218
|
+
import { cache } from "jskelet";
|
|
219
|
+
|
|
220
|
+
export const getArticle = cache(async (slug) => {
|
|
221
|
+
const response = await fetch(`${process.env.API_ORIGIN}/articles/${slug}`);
|
|
222
|
+
return response.json();
|
|
223
|
+
});
|
|
224
|
+
```
|
|
225
|
+
|
|
226
|
+
Now if both the controller and `hooks.layoutContext()` ask for the same article
|
|
227
|
+
in the same render, a single upstream request is made.
|
|
228
|
+
|
|
229
|
+
Details:
|
|
230
|
+
|
|
231
|
+
- The context is carried with `AsyncLocalStorage` and is set up by
|
|
232
|
+
`withRequestCache()` inside `route()`.
|
|
233
|
+
- **Without a context, memoization is disabled** and the function is called
|
|
234
|
+
directly. Calling it from a script or from inside another process is safe.
|
|
235
|
+
- The key is `JSON.stringify(args)`; argument-less calls share the `""` key. Do
|
|
236
|
+
not use it with arguments that cannot be serialised (functions, `Symbol`,
|
|
237
|
+
circular objects).
|
|
238
|
+
- What is stored is the function's **return value**, that is, the Promise
|
|
239
|
+
itself for `async` functions. Because the same Promise is shared, concurrent
|
|
240
|
+
calls collapse too.
|
|
241
|
+
- `withRequestCache(run)` is exported; it can be used to set up the same scope
|
|
242
|
+
outside `route()` (for example in an Express handler you wrote yourself).
|
|
243
|
+
|
|
244
|
+
## Cross-request data cache: `withDataCache`
|
|
245
|
+
|
|
246
|
+
`cache()` only lives for the duration of **a single request**. What it takes to
|
|
247
|
+
protect the long tail from the API quota is a data layer that lives across
|
|
248
|
+
requests, has a TTL and refreshes itself:
|
|
249
|
+
|
|
250
|
+
```js
|
|
251
|
+
// lib/api/articles.js
|
|
252
|
+
import { withDataCache, reportUpstreamFailure } from "jskelet";
|
|
253
|
+
|
|
254
|
+
export async function getArticle(slug) {
|
|
255
|
+
return withDataCache(`news:${slug}`, 600, async () => {
|
|
256
|
+
const response = await fetch(`${process.env.API_ORIGIN}/articles/${slug}`);
|
|
257
|
+
|
|
258
|
+
if (!response.ok) {
|
|
259
|
+
reportUpstreamFailure({ status: response.status, path: `/articles/${slug}` });
|
|
260
|
+
return null;
|
|
261
|
+
}
|
|
262
|
+
|
|
263
|
+
return response.json();
|
|
264
|
+
});
|
|
265
|
+
}
|
|
266
|
+
```
|
|
267
|
+
|
|
268
|
+
The wrapper form of the same pattern — the key is derived from the arguments:
|
|
269
|
+
|
|
270
|
+
```js
|
|
271
|
+
import { dataCache } from "jskelet";
|
|
272
|
+
|
|
273
|
+
export const getArticle = dataCache(
|
|
274
|
+
async (slug) => apiGet(`/articles/${slug}`),
|
|
275
|
+
{ key: "news", revalidate: 600 },
|
|
276
|
+
);
|
|
277
|
+
```
|
|
278
|
+
|
|
279
|
+
Behaviour:
|
|
280
|
+
|
|
281
|
+
| State | Result |
|
|
282
|
+
| --- | --- |
|
|
283
|
+
| Fresh entry | Returns immediately, the `producer` does not run |
|
|
284
|
+
| TTL expired, still inside the stale window | The stale value returns **immediately**, the refresh runs in the background |
|
|
285
|
+
| No entry | The `producer` is awaited |
|
|
286
|
+
| The `producer` failed, a stale entry exists | The stale value returns, warning: `[data-cache] producer failed, serving stale value: …` |
|
|
287
|
+
| The `producer` failed, there is no entry | The error goes to the caller |
|
|
288
|
+
|
|
289
|
+
Details:
|
|
290
|
+
|
|
291
|
+
- **Concurrent calls for the same key collapse into one upstream request.** This
|
|
292
|
+
is the behaviour that saves the most quota during warm-up rounds: if 50 pages
|
|
293
|
+
want the same index data, the API is called once.
|
|
294
|
+
- **`null` and `undefined` are not stored.** An application's HTTP client
|
|
295
|
+
usually returns `null` on failure; storing that would freeze a transient 429
|
|
296
|
+
into "no data" for the whole TTL. Pass `{ storeEmpty: true }` if you want the
|
|
297
|
+
empty answer stored deliberately.
|
|
298
|
+
- **The stale window is longer than the HTML one**: `staleFactor` defaults to 10,
|
|
299
|
+
so an entry stays as an emergency fallback for 11 times its TTL. Stale data is
|
|
300
|
+
better than an incomplete page. It can be turned off per key with
|
|
301
|
+
`{ staleFactor: 0 }`.
|
|
302
|
+
- The key belongs entirely to the application: distinctions such as language,
|
|
303
|
+
version or page number go into the key (`news:en:v2:${slug}`).
|
|
304
|
+
- When the TTL is `0` the cache is disabled and the `producer` runs on every
|
|
305
|
+
call — enough to switch a setting off temporarily.
|
|
306
|
+
|
|
307
|
+
The management surface:
|
|
308
|
+
|
|
309
|
+
| Function | What it does |
|
|
310
|
+
| --- | --- |
|
|
311
|
+
| `withDataCache(key, ttlSeconds, producer, options?)` | The main entry point |
|
|
312
|
+
| `dataCache(fn, { key, revalidate, … })` | The function wrapper |
|
|
313
|
+
| `clearDataCache(prefix?)` | Drops the entries matching the prefix (or all of them), returns how many were removed |
|
|
314
|
+
| `getDataCacheSize()` | The number of entries |
|
|
315
|
+
| `getDataCacheEntries()` | A dump: `{ key, stale, expiresIn }`. The value itself is not returned. |
|
|
316
|
+
|
|
317
|
+
`clearDataCache("news:")` is the counterpart of a "this content was updated"
|
|
318
|
+
webhook: it drops one section's data so the next HTML refresh picks up the new
|
|
319
|
+
content.
|
|
320
|
+
|
|
321
|
+
## Degraded render: `reportUpstreamFailure`
|
|
322
|
+
|
|
323
|
+
If upstream went down during the render, the output contains missing data.
|
|
324
|
+
Rather than serving such HTML for the whole TTL, the right behaviour is to
|
|
325
|
+
**never write it** to the cache: the next request tries again.
|
|
326
|
+
|
|
327
|
+
The dependency direction is deliberately inverted: the framework does not know
|
|
328
|
+
about the data layer, the data layer notifies the framework. If nobody ever
|
|
329
|
+
calls it, the cost is an empty array.
|
|
330
|
+
|
|
331
|
+
```js
|
|
332
|
+
// lib/api/client.js
|
|
333
|
+
import { reportUpstreamFailure } from "jskelet";
|
|
334
|
+
|
|
335
|
+
export async function apiGet(path) {
|
|
336
|
+
try {
|
|
337
|
+
const response = await fetch(`${process.env.API_ORIGIN}${path}`);
|
|
338
|
+
|
|
339
|
+
if (!response.ok) {
|
|
340
|
+
reportUpstreamFailure({ status: response.status, path });
|
|
341
|
+
return null;
|
|
342
|
+
}
|
|
343
|
+
|
|
344
|
+
return response.json();
|
|
345
|
+
} catch (error) {
|
|
346
|
+
// No response at all: status 0 means a network error.
|
|
347
|
+
reportUpstreamFailure({ status: 0, path });
|
|
348
|
+
return null;
|
|
349
|
+
}
|
|
350
|
+
}
|
|
351
|
+
```
|
|
352
|
+
|
|
353
|
+
### Distinguishing transient and permanent failures
|
|
354
|
+
|
|
355
|
+
| State | Counts as | Result |
|
|
356
|
+
| --- | --- | --- |
|
|
357
|
+
| `0` (network error), `408`, `425`, `429`, `>= 500` | **Transient** | The page is not written to the cache, warning: `[render] <path> was produced with missing data, not caching it (…)` |
|
|
358
|
+
| Others (`400`, `403`, `404`, …) | **Permanent** | Only a warning: `[render] <path> was produced with missing data, upstream is failing permanently (…)`. The cache is not blocked. |
|
|
359
|
+
|
|
360
|
+
Permanent failures not blocking the cache is deliberate: deterministic answers
|
|
361
|
+
do not get better by retrying. Turning the cache off because of them would mean
|
|
362
|
+
rendering the page from scratch on every visit — the content comes back just as
|
|
363
|
+
incomplete, and the visitor only pays the render time.
|
|
364
|
+
|
|
365
|
+
Output produced with missing data is **not offered to shared caches** either: a
|
|
366
|
+
`degraded` response gets `private, no-store` instead of `public, s-maxage=…`.
|
|
367
|
+
Taking back the "do not store" decision at the CDN would repeat the same mistake
|
|
368
|
+
one layer up. The diagnostic header (`X-JSkelet-Cache: MISS`) is still written.
|
|
369
|
+
|
|
370
|
+
### When `notFound()` coincides with a transient failure
|
|
371
|
+
|
|
372
|
+
A controller that calls `notFound()` because no data arrived can turn the whole
|
|
373
|
+
site into 404s when upstream is rate limited — and because those 404s enter the
|
|
374
|
+
cache, a temporary quota problem becomes a "this page does not exist" answer for
|
|
375
|
+
the whole TTL. For a search engine that is a permanent loss.
|
|
376
|
+
|
|
377
|
+
The framework separates the two cases: if a **transient** upstream failure was
|
|
378
|
+
reported during the render, `notFound()` is not served as a 404.
|
|
379
|
+
|
|
380
|
+
| During the render | Result of `notFound()` |
|
|
381
|
+
| --- | --- |
|
|
382
|
+
| A transient failure exists (`429`, `5xx`, network error) | `503`, `Retry-After: 30`, `no-store` — **not** written to the cache, the next request produces the real content |
|
|
383
|
+
| A permanent failure (`404`, `403`…) or no failure | A normal `404` |
|
|
384
|
+
|
|
385
|
+
The log line:
|
|
386
|
+
`[render] /news/x returned notFound() while upstream is failing (429 /api/...), serving an uncached 503 instead`
|
|
387
|
+
|
|
388
|
+
So when upstream runs out of quota the page is produced dynamically, without
|
|
389
|
+
being written to the cache; nothing is frozen as "missing".
|
|
390
|
+
|
|
391
|
+
## Managing the cache
|
|
392
|
+
|
|
393
|
+
`jskelet` exports these functions:
|
|
394
|
+
|
|
395
|
+
| Function | What it does |
|
|
396
|
+
| --- | --- |
|
|
397
|
+
| `withHtmlCache(key, ttlSeconds, producer)` | For using the cache directly. If `ttlSeconds` is 0 the producer always runs. |
|
|
398
|
+
| `clearHtmlCache()` | Empties the store completely. |
|
|
399
|
+
| `getHtmlCacheSize()` | The number of entries. |
|
|
400
|
+
| `getHtmlCacheEntries()` | A dump: `{ key, bytes, status, stale, expiresIn, encodings }`. The HTML body is not returned, only its size. |
|
|
401
|
+
|
|
402
|
+
To write an admin endpoint:
|
|
403
|
+
|
|
404
|
+
```js
|
|
405
|
+
import { clearHtmlCache, getHtmlCacheEntries } from "jskelet";
|
|
406
|
+
|
|
407
|
+
export default function register(app) {
|
|
408
|
+
app.post("/_admin/cache/clear", (req, res) => {
|
|
409
|
+
if (req.headers["x-admin-token"] !== process.env.ADMIN_TOKEN) {
|
|
410
|
+
res.status(404).end();
|
|
411
|
+
return;
|
|
412
|
+
}
|
|
413
|
+
clearHtmlCache();
|
|
414
|
+
res.json({ ok: true });
|
|
415
|
+
});
|
|
416
|
+
|
|
417
|
+
app.get("/_admin/cache", (req, res) => {
|
|
418
|
+
res.json(getHtmlCacheEntries());
|
|
419
|
+
});
|
|
420
|
+
}
|
|
421
|
+
```
|
|
422
|
+
|
|
423
|
+
The dev server also clears the cache by itself whenever the manifest changes:
|
|
424
|
+
the stored HTML would be carrying asset URLs with old hashes, and if it were
|
|
425
|
+
not cleared the page would keep requesting a deleted file
|
|
426
|
+
([09-dev-tools.md](./09-dev-tools.md)).
|
|
427
|
+
|
|
428
|
+
Because the cache lives in process memory, if you run more than one
|
|
429
|
+
process/replica each one has its own cache; `clearHtmlCache()` only affects the
|
|
430
|
+
process it is called in.
|
|
431
|
+
|
|
432
|
+
## Prewarm — warming up at startup
|
|
433
|
+
|
|
434
|
+
The equivalent of Next's build-time prerender, except the output is not written
|
|
435
|
+
to disk: since the cache lives in process memory, the warm-up also happens when
|
|
436
|
+
the process comes up. The gain is the same — the first visitor does not wait
|
|
437
|
+
for a cold render — but the data is not frozen; every entry ages with the
|
|
438
|
+
route's `revalidate` and is refreshed in the background with
|
|
439
|
+
stale-while-revalidate.
|
|
440
|
+
|
|
441
|
+
The warm-up is done with **real HTTP requests**
|
|
442
|
+
(`http://127.0.0.1:<port>`), so that the cache key, the compression and the
|
|
443
|
+
middleware chain are exactly the same as with normal traffic.
|
|
444
|
+
|
|
445
|
+
### `hooks.prewarmPaths()`
|
|
446
|
+
|
|
447
|
+
The application declares which paths get warmed; usually it is the very same
|
|
448
|
+
function that produces the sitemap.
|
|
449
|
+
|
|
450
|
+
```js
|
|
451
|
+
// jskelet.config.mjs
|
|
452
|
+
export default {
|
|
453
|
+
hooks: {
|
|
454
|
+
async prewarmPaths() {
|
|
455
|
+
const slugs = await getAllArticleSlugs();
|
|
456
|
+
return ["/", "/markets", ...slugs.map((slug) => `/news/${slug}`)];
|
|
457
|
+
},
|
|
458
|
+
},
|
|
459
|
+
};
|
|
460
|
+
```
|
|
461
|
+
|
|
462
|
+
Rules:
|
|
463
|
+
|
|
464
|
+
- If it does not return an array a warning is printed and no warm-up happens.
|
|
465
|
+
- Only strings starting with `/` are taken.
|
|
466
|
+
- Ones starting with one of the `prewarmSkip` prefixes are skipped. The default
|
|
467
|
+
list: `/api/`, `/_fragment/`, `/__jskelet/`. Session-dependent pages should
|
|
468
|
+
not be warmed.
|
|
469
|
+
- Deduplication **preserves order**: when no `priority` is given, the order the
|
|
470
|
+
application provides is meaningful — put the most important pages first.
|
|
471
|
+
- If this hook is not defined the warm-up is never set up; not even the timer
|
|
472
|
+
is started.
|
|
473
|
+
|
|
474
|
+
### Round logic
|
|
475
|
+
|
|
476
|
+
1. The list is collected. If it is longer than `max` (400 by default) a slice is
|
|
477
|
+
selected: the paths matching `priority` are taken first **on every round**,
|
|
478
|
+
and the remaining slots are filled from the queue.
|
|
479
|
+
2. `concurrency` workers send requests in parallel (4 in prod, 2 in dev). Less
|
|
480
|
+
parallelism in dev: so the scan does not compete for CPU with the render of
|
|
481
|
+
the page you currently have open in the browser.
|
|
482
|
+
3. If `rps` is given, the round never goes above that rate — no matter the
|
|
483
|
+
parallelism.
|
|
484
|
+
4. **A single serial retry round** is performed for the failed paths after
|
|
485
|
+
waiting `retryDelayMs` (`concurrency: 1`). The wait is deliberate: rate limit
|
|
486
|
+
windows are on the order of seconds, so retrying immediately just earns the
|
|
487
|
+
same 429.
|
|
488
|
+
5. A summary is logged:
|
|
489
|
+
`[prewarm] warmed 128/130 pages, 2 failed, 5 recovered on the retry pass (12.4s)`
|
|
490
|
+
|
|
491
|
+
### Warm-up order: `priority`
|
|
492
|
+
|
|
493
|
+
```js
|
|
494
|
+
// jskelet.config.mjs
|
|
495
|
+
cache: () => ({
|
|
496
|
+
prewarm: {
|
|
497
|
+
priority: [
|
|
498
|
+
"/",
|
|
499
|
+
"/markets/:path*",
|
|
500
|
+
/-comments$/,
|
|
501
|
+
],
|
|
502
|
+
},
|
|
503
|
+
}),
|
|
504
|
+
```
|
|
505
|
+
|
|
506
|
+
The pattern syntax (`/news/:slug`) and a plain `RegExp` can be used together;
|
|
507
|
+
the latter is for rules the pattern syntax does not cover, such as "everything
|
|
508
|
+
ending in `-comments`". Whatever is written first is warmed first; paths that
|
|
509
|
+
match nothing go to the queue and keep their relative order.
|
|
510
|
+
|
|
511
|
+
### Drip warm-up: `rotate` + `rps` + `intervalSeconds`
|
|
512
|
+
|
|
513
|
+
On a site with 10,000 paths, warming everything in a single round is neither
|
|
514
|
+
possible (the HTML cache holds 500 entries) nor right (the API quota runs out).
|
|
515
|
+
The correct behaviour is to spread the list over time:
|
|
516
|
+
|
|
517
|
+
```js
|
|
518
|
+
prewarm: {
|
|
519
|
+
max: 300, // 300 pages per round
|
|
520
|
+
rps: 4, // at most 4 requests per second
|
|
521
|
+
intervalSeconds: 300, // a round every 5 minutes
|
|
522
|
+
rotate: true, // the queue continues where it left off
|
|
523
|
+
priority: ["/", "/markets/:path*"],
|
|
524
|
+
}
|
|
525
|
+
```
|
|
526
|
+
|
|
527
|
+
In this setup the priority pages are refreshed on every round, the rest of the
|
|
528
|
+
queue is walked end to end across rounds, and upstream never sees more than four
|
|
529
|
+
requests per second. Used together with the data cache, the warm-up barely
|
|
530
|
+
reaches the API after the second round: it reads from the data layer.
|
|
531
|
+
|
|
532
|
+
With rotation on, the paths left outside the limit are not lost, they are left
|
|
533
|
+
for the next round; the log distinguishes this:
|
|
534
|
+
`… , 700 deferred to the next pass`. With `rotate: false` you get the classic
|
|
535
|
+
behaviour — every round warms the same first slice of the list and the rest is
|
|
536
|
+
never warmed (`… , 700 over the limit`).
|
|
537
|
+
|
|
538
|
+
If a round takes longer than `intervalSeconds`, a new round is not started;
|
|
539
|
+
overlapping rounds would put twice the load on upstream.
|
|
540
|
+
|
|
541
|
+
The requests go out with the headers `user-agent: jskelet-prewarm`
|
|
542
|
+
(`brand.prewarmUserAgent`) and `accept-encoding: br, gzip`; the second one so
|
|
543
|
+
that the compressed body enters the cache too.
|
|
544
|
+
|
|
545
|
+
If `DEV_TOKEN` is set, the warm-up carries the token as a cookie; otherwise the
|
|
546
|
+
dev gate returns 404 for all pages and the cache never fills.
|
|
547
|
+
|
|
548
|
+
The request list in the dev panel and the terminal filter out requests carrying
|
|
549
|
+
`prewarmUserAgent`: so that hundreds of warm-up requests do not flood the view.
|
|
550
|
+
Progress shows up in the badge next to the bubble.
|
|
551
|
+
|
|
552
|
+
### Timing
|
|
553
|
+
|
|
554
|
+
- The warm-up starts at boot **with a delay**: so it does not compete with the
|
|
555
|
+
first real requests. The default delay is 500 ms in prod and 3000 ms in dev.
|
|
556
|
+
Longer in dev, because a file save restarts the process and the timer dies
|
|
557
|
+
with it; it only warms up once the server stays quiet for a while.
|
|
558
|
+
- If `PREWARM_INTERVAL_SECONDS` / `cache().prewarm.intervalSeconds` > 0 the
|
|
559
|
+
round is repeated periodically. Because entries age with `revalidate` and the
|
|
560
|
+
visitor does not wait thanks to stale-while-revalidate, this is **optional**;
|
|
561
|
+
it is for setups that also want to keep pages that are never visited warm.
|
|
562
|
+
- All timers are `unref()`ed: they do not delay process shutdown.
|
|
563
|
+
- No warm-up failure takes the process down.
|
|
564
|
+
|
|
565
|
+
### Settings
|
|
566
|
+
|
|
567
|
+
Order of precedence: **environment variable → config → code default.** Env
|
|
568
|
+
comes first so that one-off experiments can be done without editing the config.
|
|
569
|
+
|
|
570
|
+
| Setting | Env | `cache().prewarm` | Default |
|
|
571
|
+
| --- | --- | --- | --- |
|
|
572
|
+
| On/off | `PREWARM=0` disables it, `PREWARM=1` overrides the config and enables it | `enabled` | `true` |
|
|
573
|
+
| Maximum paths per round | `PREWARM_MAX` | `max` | `400` |
|
|
574
|
+
| Parallelism | `PREWARM_CONCURRENCY` | `concurrency` | prod 4, dev 2 |
|
|
575
|
+
| Requests per second | `PREWARM_RPS` | `rps` | `0` (unlimited) |
|
|
576
|
+
| Startup delay (ms) | `PREWARM_DELAY_MS` | `delayMs` | prod 500, dev 3000 |
|
|
577
|
+
| Retry round delay (ms) | `PREWARM_RETRY_DELAY_MS` | `retryDelayMs` | `2000` |
|
|
578
|
+
| Period (seconds) | `PREWARM_INTERVAL_SECONDS` | `intervalSeconds` | `0` (off) |
|
|
579
|
+
| Queue rotation | — | `rotate` | `true` |
|
|
580
|
+
| Warm-up order | — | `priority` | `[]` |
|
|
581
|
+
|
|
582
|
+
Numeric settings only accept **positive and finite** values; an invalid value
|
|
583
|
+
silently falls through to the next layer.
|
|
584
|
+
|
|
585
|
+
### Triggering by hand
|
|
586
|
+
|
|
587
|
+
```js
|
|
588
|
+
import { prewarm, prewarmProgress } from "jskelet";
|
|
589
|
+
|
|
590
|
+
await prewarm({ origin: "http://127.0.0.1:3000" }); // paths from the hook
|
|
591
|
+
await prewarm({ origin, paths: ["/", "/markets"] }); // only these paths
|
|
592
|
+
await prewarm({ origin, quiet: true }); // without printing a summary
|
|
593
|
+
```
|
|
594
|
+
|
|
595
|
+
If `paths` is given the hook is never called. The return value is
|
|
596
|
+
`{ ok, failed, total, elapsed }`.
|
|
597
|
+
|
|
598
|
+
`prewarmProgress` holds the live state and the dev panel reads it:
|
|
599
|
+
|
|
600
|
+
```js
|
|
601
|
+
{
|
|
602
|
+
active, done, total, ok, failed, startedAt, finishedAt,
|
|
603
|
+
entries: [{ path, status, ms, bytes, cache, error }],
|
|
604
|
+
}
|
|
605
|
+
```
|
|
606
|
+
|
|
607
|
+
The `cache` field inside `entries` is that path's `X-JSkelet-Cache` response;
|
|
608
|
+
from there you can see whether the warm-up round really returned `MISS` and
|
|
609
|
+
filled the cache.
|
|
610
|
+
|
|
611
|
+
## Diagnosis: common situations
|
|
612
|
+
|
|
613
|
+
- **Every request returns `MISS`.** The route was not given a `revalidate`, or
|
|
614
|
+
the pattern inside `cache().html` gives 0 seconds. Or the page returns a code
|
|
615
|
+
other than `status: 200`.
|
|
616
|
+
- **The page returns `MISS` but upstream is healthy.** A transient upstream
|
|
617
|
+
failure may have been reported; look for the line `was produced with missing
|
|
618
|
+
data, not caching it` in the log.
|
|
619
|
+
- **Stale data all the time.** `revalidate` is too high; remember that the real
|
|
620
|
+
lag is at most `revalidate` + one refresh round.
|
|
621
|
+
- **The cache is bloating.** Because query parameters go into the key, campaign
|
|
622
|
+
parameters may be multiplying entries.
|
|
623
|
+
- **The warm-up never runs.** `hooks.prewarmPaths` is not defined, `PREWARM=0`
|
|
624
|
+
is set, or `cache().prewarm.enabled === false`.
|
|
625
|
+
- **The warm-up round pushes the API into 429.** No `rps` was given. Lowering
|
|
626
|
+
`concurrency` is not enough; the setting that protects the quota is the total
|
|
627
|
+
rate. The lasting fix is the data cache: after the second round the warm-up
|
|
628
|
+
does not reach upstream.
|
|
629
|
+
- **The warm-up list is longer than `max` and its tail never warms.** `rotate`
|
|
630
|
+
may be `false`; the `over the limit` phrase in the log shows this.
|
|
631
|
+
- **A whole section returns 404.** Upstream may be down. In that case a 503 that
|
|
632
|
+
does not enter the cache is now returned instead of a 404; look for the
|
|
633
|
+
`returned notFound() while upstream is failing` line in the log.
|
|
634
|
+
|
|
635
|
+
## What's next
|
|
636
|
+
|
|
637
|
+
- The full reference of config fields and the env table:
|
|
638
|
+
[07-configuration.md](./07-configuration.md)
|
|
639
|
+
- Watching the cache from the dev panel: [09-dev-tools.md](./09-dev-tools.md)
|
|
640
|
+
- Using it together with a CDN/reverse proxy: [10-deployment.md](./10-deployment.md)
|