pmtiles-swarm 0.55.2 → 0.56.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -7,6 +7,65 @@
7
7
  ### 🐞 Bug fixes
8
8
  - _...Add new stuff here..._
9
9
 
10
+ ## 0.56.1
11
+ ### ✨ Features and improvements
12
+ - **Comments trimmed back to house style, and the reasoning moved into the docs where it belongs.**
13
+ Recent work had been leaving 15–25 line explanatory blocks in the source; measured across
14
+ `catalog.js` that was 85 new comment lines against 76 new code lines. The argument for a decision
15
+ drifts out of date faster in a comment than in prose, and buries the code while it does it.
16
+
17
+ Two new sections in [docs/internals.md](docs/internals.md) give the displaced reasoning a home:
18
+ **Re-reading a summary an older prober wrote**, and **A validator for a URL that stays put** —
19
+ which is where the ETag story now lives, including why the infohash was the wrong choice and what
20
+ a PMTiles reader needs to see for `If-Range` and cross-origin `ETag` to work.
21
+
22
+ No behaviour change.
23
+
24
+ ## 0.56.0
25
+ ### ✨ Features and improvements
26
+ - **MapLibre Tiles are recognised.** PMTiles tile type `6` is MLT, and the tile-type table stopped at
27
+ `5` — so an MLT archive probed as `unknown` and was refused a tile endpoint it could have served,
28
+ even though the extension map had known about `.mlt` for a while.
29
+
30
+ No new TileJSON key was needed for it. MapLibre spells the tile encoding of a vector source
31
+ `encoding`, exactly as it does the elevation packing of a raster-dem one, and applies TileJSON
32
+ members to both after the source is constructed — so `encoding: "mlt"` reaches a vector source the
33
+ same way `encoding: "terrarium"` reaches a raster one. One key, two meanings, told apart by the
34
+ source type. The MLT value is read from the header rather than the metadata, because an archive
35
+ whose tile type says MapLibre Tiles is MLT-encoded and nothing needs to say so twice.
36
+
37
+ - **A resume-data shortfall is now reported.** The sidecar has been answering with how many torrents
38
+ were asked to write and how many managed it before the deadline, and both numbers were discarded —
39
+ by the engine wrapper, and again by the timer that called it. A torrent that does not write is one
40
+ that gets re-hashed on the next start, which for a 700 GiB archive is the difference between
41
+ seeding in seconds and seeding in half an hour. That is what "why is everything at 0%" looks like
42
+ from outside, and the silence here is part of why it was hard to see.
43
+
44
+ ### 🐞 Bug fixes
45
+ - **The two pages disagreed about how large an archive was.** The console rounded to whole units
46
+ above ten and the public page always kept a decimal, so the same archive read as `81 GiB` on one
47
+ and `80.6 GiB` on the other. Both now keep a decimal from KiB up — these are mostly archive sizes,
48
+ and half a gigabyte is worth seeing — and a test holds the two helpers character-for-character
49
+ identical, since nothing about a duplicated function announces when it stops being a copy.
50
+
51
+ ## 0.55.3
52
+ ### 🐞 Bug fixes
53
+ - **A `/latest/` document could never be updated once a client had cached it.** The `ETag` was the
54
+ infohash, which says which _build_ a category resolved to — and these documents carry more than
55
+ that. A TileJSON also carries the archive's summary; a magnet also carries its web seeds and
56
+ trackers; the feed carries both. So enriching a summary (0.55.1 reading an `encoding` an older
57
+ prober had missed) or adding a web seed changed the body while the infohash stayed put: every cache
58
+ in the path revalidated, was told `304`, and went on serving the old document indefinitely. That is
59
+ not a cache being stale for a minute; it is a document that can never change again.
60
+
61
+ These are now tagged over what is actually sent. `/latest/<category>/archive.pmtiles` and the
62
+ `.torrent` redirect keep the infohash, because for those the infohash really is the whole content.
63
+
64
+ - **A successful metadata re-read was logged as a failure.** The line counting vector layers ran for
65
+ raster archives too, where there are none, and the `TypeError` went to the catch beside it — which
66
+ reported a backfill that had already written its result as "no vector layers yet". A misleading log
67
+ is worse than none when it is what somebody is reading to find out whether the thing works.
68
+
10
69
  ## 0.55.2
11
70
  ### 🐞 Bug fixes
12
71
  - **The stale-summary re-read never ran for a category URL.** 0.55.1 added it to
package/docs/internals.md CHANGED
@@ -18,6 +18,8 @@ Operator-facing documentation is elsewhere — see [publishing](publishing.md),
18
18
  - [Serving an MBTiles archive](#serving-an-mbtiles-archive)
19
19
  - [Answering for a tile that is not there](#answering-for-a-tile-that-is-not-there)
20
20
  - [Reading an archive that is still arriving](#reading-an-archive-that-is-still-arriving)
21
+ - [Re-reading a summary an older prober wrote](#re-reading-a-summary-an-older-prober-wrote)
22
+ - [A validator for a URL that stays put](#a-validator-for-a-url-that-stays-put)
21
23
  - [The health endpoint](#the-health-endpoint)
22
24
  - [The externally visible base URL](#the-externally-visible-base-url)
23
25
  - [Scheduled sources](#scheduled-sources)
@@ -310,6 +312,86 @@ nothing to read yet and the same archive at 100% will, so a permanent
310
312
  would put a swarm read behind each one. The limiter is in memory on purpose: a
311
313
  restart is a reasonable moment to try again.
312
314
 
315
+ ## Re-reading a summary an older prober wrote
316
+
317
+ An archive's summary — format, zoom range, bounds, vector layers, encoding — is
318
+ read once, written into the catalog, and never questioned again. That is right
319
+ as far as it goes: re-reading a header out of the swarm is not free, and for a
320
+ given infohash the answer cannot change.
321
+
322
+ It goes wrong the moment the prober learns to read something new. Every archive
323
+ probed before that keeps a summary with a hole in it, permanently, and the only
324
+ way out is to remove and re-add the archive by hand.
325
+
326
+ `encoding` was the case that made this plain. The key had been sitting in the
327
+ metadata of archives a node had been serving for months; teaching the prober to
328
+ read it changed nothing at all, because nothing ever asked again.
329
+
330
+ So `summarize()` stamps a `summaryVersion`, and a summary older than the current
331
+ prober is re-read once in the background — on a request for the archive's
332
+ TileJSON, rate-limited to once a minute per archive. Raise `SUMMARY_VERSION`
333
+ whenever a field is added and every archive in the catalog picks it up on its
334
+ own, with nothing to do on upgrade.
335
+
336
+ Two things this needs to get right, both learned the hard way:
337
+
338
+ - **Every route that serves a TileJSON has to trigger it**, not only
339
+ `/archives/<infohash>/tiles.json`. `/latest/<category>/tiles.json` is the URL
340
+ a style points at, so leaving it out missed exactly the archives that were
341
+ being consumed the documented way.
342
+ - **The write-back must not be conditioned on one field.** It began as a
343
+ vector-layer backfill and returned early unless layers turned up, which would
344
+ have discarded the encoding it had just gone to fetch.
345
+
346
+ ## A validator for a URL that stays put
347
+
348
+ Everything under `/archives/` is addressed by infohash: the URL changes when the
349
+ content does, so it can be cached for a year and marked `immutable`. A `/latest/`
350
+ URL is the opposite — stable on purpose, with the content moving underneath it —
351
+ so it needs a validator, and a short TTL alone is a guess.
352
+
353
+ The obvious validator is the infohash of the archive the category resolved to.
354
+ It is wrong, and it fails in a way that has no bottom: these documents carry more
355
+ than which build they name. A TileJSON also carries the archive's summary; a
356
+ magnet also carries its web seeds and trackers; the feed carries both. Enrich a
357
+ summary or add a web seed and the body changes while the infohash does not — so
358
+ every cache in the path revalidates, is told `304`, and goes on serving the old
359
+ document. Not stale for a minute: unable to be updated again for as long as the
360
+ archive exists.
361
+
362
+ So these are tagged over the body they are sending. Two nodes answering
363
+ identically still produce identical tags, which is the property the infohash was
364
+ chosen for; two nodes answering differently no longer claim otherwise.
365
+
366
+ `/latest/<category>/archive.pmtiles` and the `.torrent` redirect keep the
367
+ infohash, because for those the infohash really is the whole content.
368
+
369
+ ### Where the ranges have to be careful
370
+
371
+ A PMTiles reader does not fetch a file. It fetches a header, then a root
372
+ directory, then leaf directories, then tiles, over minutes or hours. If a
373
+ rebuild lands partway through, the offsets it read from the old build address
374
+ bytes in the new one — which does not fail loudly. It decodes as the wrong tile,
375
+ or as nothing, with no error anywhere naming the cause.
376
+
377
+ `If-Range` is therefore honoured on the category range endpoint: a range
378
+ conditioned on a build that is no longer current is refused _as a range_ and
379
+ answered in full, which a reader survives. `Last-Modified` is suppressed there so
380
+ a client cannot condition on a date instead — a build restored from a backup can
381
+ be newer while looking older.
382
+
383
+ The official reader closes the loop from its side, comparing the `ETag` of every
384
+ response against the one it saw first and re-reading the header when they differ.
385
+ That only works if it can see the tag: `ETag` is not exposed to cross-origin
386
+ JavaScript by default, so these routes send `Access-Control-Expose-Headers`.
387
+ Unexposed, the reader compares against `null`, the comparison never fires, and it
388
+ splices two builds in silence.
389
+
390
+ A weak validator is no better. The reader discards any tag beginning with `W/`,
391
+ which is exactly what Express derives from a file's size and mtime — and that
392
+ tag also differs per node, so two nodes behind a load balancer would hand a
393
+ reader two tags for byte-identical archives.
394
+
313
395
  ## The health endpoint
314
396
 
315
397
  For a load balancer, which needs three things: no credential, a cheap answer, and
@@ -157,14 +157,23 @@ answers 404 for them exactly as the feeds do.
157
157
 
158
158
  ### How a client knows the build moved
159
159
 
160
- Every one of these carries an `ETag`, and the tag is the infohash of the archive
161
- it resolved to:
160
+ Every one of these carries an `ETag` over the document it is sending:
162
161
 
163
162
  ```
164
- ETag: "913d671f3a28c5b8d605e28cf6bf01e293d36e86"
163
+ ETag: "6b8f1c2d…"
165
164
  Cache-Control: public, max-age=60, must-revalidate
166
165
  ```
167
166
 
167
+ **Over the document, not over the infohash.** The infohash is the obvious
168
+ choice and it is wrong here: it says which _build_ a category resolved to, and
169
+ these documents carry more than that. A TileJSON also carries the archive's
170
+ summary; a magnet also carries its web seeds and trackers. Enrich a summary or
171
+ add a web seed and the body changes while the infohash does not — so every cache
172
+ in the path revalidates, is told `304`, and goes on serving the old document.
173
+ Not stale for a minute: unable to be updated at all. Only
174
+ `/latest/{category}/archive.pmtiles` and the `.torrent` redirect are tagged by
175
+ infohash, because for those the infohash really is the whole content.
176
+
168
177
  A short TTL on its own is a guess. At five minutes, every client and every proxy
169
178
  in front of one serves the previous build for up to five minutes after a rollover
170
179
  and not one of them can tell it is doing so. The infohash is the honest answer:
package/docs/tilejson.md CHANGED
@@ -172,7 +172,22 @@ MapLibre applies every TileJSON member to the source after the source is
172
172
  constructed, so an `encoding` here **overrides** one written in the style. That
173
173
  is the intended direction: the archive is the thing that knows.
174
174
 
175
- Anything other than the three values the style specification defines is dropped
175
+ ### `mlt`
176
+
177
+ The same key carries one more value, for a different kind of source. MapLibre
178
+ spells the tile encoding of a vector source `encoding` too, and `mlt` there means
179
+ the tiles are [MapLibre Tiles](https://maplibre.org/maplibre-tile-spec/) rather
180
+ than MVT — so there was no new key to invent for it, only a value to allow
181
+ through.
182
+
183
+ This one comes from the header rather than the metadata. PMTiles has a tile type
184
+ for MLT (`6`), and an archive whose tile type says MapLibre Tiles _is_
185
+ MLT-encoded; nothing needs to be written in the metadata to make it so.
186
+ Elevation packing is the opposite case — the header knows the tile is WebP and
187
+ nothing about what its three channels mean — which is why the two are read from
188
+ different places despite sharing a key.
189
+
190
+ Anything other than the values the style specification defines is dropped
176
191
  rather than passed on. A client handed an encoding it does not recognise is
177
192
  worse off than one handed nothing, because nothing at least leaves it free to
178
193
  use its own default.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pmtiles-swarm",
3
- "version": "0.55.2",
3
+ "version": "0.56.1",
4
4
  "description": "BitTorrent distribution for PMTiles map archives: create torrents, watch folders, publish and subscribe to RSS feeds, and seed through qBittorrent or an embedded client",
5
5
  "type": "module",
6
6
  "main": "src/index.js",
package/src/api.js CHANGED
@@ -366,13 +366,8 @@ export function createApp({
366
366
  const metadataRetries = new Map();
367
367
  const needsReread = (summary, infoHash) => {
368
368
  if (!summary) return false;
369
- // Two reasons to go back to the archive. The first is the one this began
370
- // as: a vector archive whose layers have not been read yet. The second is
371
- // the general case it turned out to be — a summary written by an older
372
- // prober, missing whatever that prober did not know to look for. Without
373
- // it, adding a field to the summary reaches new archives only, and every
374
- // archive already in the catalog keeps its hole until somebody removes and
375
- // re-adds it by hand.
369
+ // Either the layers were never read, or an older prober wrote this and did
370
+ // not know to look for what the summary now carries.
376
371
  const stale = summary.summaryVersion !== SUMMARY_VERSION;
377
372
  const missingLayers = summary.format === 'pbf' && !summary.vectorLayers;
378
373
  if (!stale && !missingLayers) return false;
@@ -404,19 +399,18 @@ export function createApp({
404
399
  tiles
405
400
  .summarize(entry.infoHash, { timeoutMs })
406
401
  .then(async (summary) => {
407
- // Written back whenever the read produced something newer than what
408
- // was stored. Returning early unless vector layers turned up was right
409
- // when layers were the only thing this looked for; it would now throw
410
- // away the encoding it went to fetch.
411
402
  const fresh = summary.summaryVersion !== entry.pmtiles?.summaryVersion;
412
403
  if (!summary.vectorLayers && !fresh) return;
413
404
  await catalog.put({
414
405
  infoHash: entry.infoHash,
415
406
  pmtiles: { ...entry.pmtiles, ...summary },
416
407
  });
408
+ // Raster archives have no layers to count.
417
409
  console.log(
418
- `[tiles] read ${summary.vectorLayers.length} vector layers for ` +
419
- `${entry.name} out of the swarm`,
410
+ summary.vectorLayers
411
+ ? `[tiles] read ${summary.vectorLayers.length} vector layers for ` +
412
+ `${entry.name} out of the swarm`
413
+ : `[tiles] re-read the metadata for ${entry.name}`,
420
414
  );
421
415
  })
422
416
  .catch((error) => {
@@ -1041,17 +1035,16 @@ export function createApp({
1041
1035
  // `location.origin` there names the one port that is not for the
1042
1036
  // public — and a web seed built from it is handed to every peer in
1043
1037
  // the swarm.
1044
- // Whatever was actually published for this archive, where one was:
1045
- // the record is the truth about what peers hold, and it need not be
1046
- // what this node would build today.
1038
+ // What was actually published, which need not be what this node
1039
+ // would build today.
1047
1040
  url:
1048
1041
  entry.selfWebSeedUrl ??
1049
1042
  `${publishingBase({
1050
1043
  config,
1051
1044
  requestBase: baseUrl(req),
1052
1045
  })}/archives/${entry.infoHash}/archive.pmtiles`,
1053
- // The base the switch would use, so the console can offer it for
1054
- // editing before anything permanent is written.
1046
+ // Worked out here: the console runs on the admin listener, whose port
1047
+ // no peer can reach.
1055
1048
  base: publishingBase({ config, requestBase: baseUrl(req) }),
1056
1049
  },
1057
1050
  });
@@ -1177,14 +1170,8 @@ export function createApp({
1177
1170
  const body = req.body ?? {};
1178
1171
  try {
1179
1172
  const result = await library.setPublishing(req.params.infoHash, body, {
1180
- // Typed into the field beside the switch, where it beats
1181
- // everything else: this is the one URL a person gets to decide
1182
- // deliberately, because it is the one that cannot be taken back.
1173
+ // The typed field wins; otherwise the request, on the public port.
1183
1174
  publishingUrl: body.publishingUrl,
1184
- // Otherwise the request it is being asked on is the best evidence
1185
- // available, the same reasoning the TileJSON already uses — and
1186
- // with the public port rather than the admin one, since a seed URL
1187
- // on a listener bound to localhost reaches no peer at all.
1188
1175
  baseUrl: baseUrl(req),
1189
1176
  });
1190
1177
  res.json(result);
@@ -2052,9 +2039,7 @@ export function createApp({
2052
2039
  magnet: entry.magnet,
2053
2040
  torrent: `${baseUrl(req)}/archives/${entry.infoHash}/archive.torrent`,
2054
2041
  webSeeds: entry.webSeeds ?? [],
2055
- // Present only when this archive is offered as a download. Absent
2056
- // otherwise rather than null: a public document should say what is
2057
- // on offer, not enumerate what is being withheld.
2042
+ // Absent rather than null where it is not on offer.
2058
2043
  ...(publishingFor(entry, config).publicDownload
2059
2044
  ? {
2060
2045
  archive: `${baseUrl(req)}/archives/${entry.infoHash}/archive.pmtiles`,
@@ -2103,34 +2088,22 @@ export function createApp({
2103
2088
  };
2104
2089
 
2105
2090
  /**
2106
- * Tags a response as "whichever build of this category is current".
2091
+ * Tags a `/latest/` response so a cache can tell when the build moves.
2107
2092
  *
2108
- * Everything under `/archives/` is addressed by infohash and may be cached
2109
- * for a year, because the URL changes when the content does. A `/latest/`
2110
- * URL is the exact opposite: it is stable on purpose, so the content
2111
- * underneath it moves and the URL alone gives a cache no way to notice. A
2112
- * short TTL was the only thing these endpoints had, and a TTL is a guess —
2113
- * at five minutes, every client and every proxy in front of one serves the
2114
- * previous build for up to five minutes after a rollover, and not one of
2115
- * them can tell whether it is doing so.
2116
- *
2117
- * The infohash is the honest validator. It changes exactly when the archive
2118
- * a category resolves to changes, never otherwise, and it is the same value
2119
- * on every node in the swarm — so two nodes behind a load balancer agree
2120
- * about what is current, rather than each inventing a tag of its own from a
2121
- * body hash or an mtime and defeating the cache whenever a client is sent
2122
- * to the other one.
2093
+ * See docs/serving-tiles.md — "How a client knows the build moved".
2123
2094
  * @param {import('express').Response} res - The response to tag.
2124
- * @param {object} entry - The archive this URL resolved to.
2095
+ * @param {object|string} content - What is being sent.
2125
2096
  */
2126
- const tagAsLatest = (res, entry) => {
2127
- // Quoted, because that is what an entity-tag is. Unquoted, a validator
2128
- // never matches an If-None-Match and every revalidation misses silently.
2129
- res.setHeader('etag', `"${entry.infoHash}"`);
2130
- // must-revalidate because serving this one stale is not merely "slightly
2131
- // old": a reader holding part of one build and asking for the rest after
2132
- // a rollover assembles a file that never existed. A minute of cache with
2133
- // a mandatory check at the end of it is the trade.
2097
+ const tagAsLatest = (res, content) => {
2098
+ // Over the body, not the infohash: these documents carry more than which
2099
+ // build they resolved to, so the infohash can stay put while they change.
2100
+ const text =
2101
+ typeof content === 'string' ? content : JSON.stringify(content);
2102
+ // Quoted, or `fresh` never matches it and every revalidation misses.
2103
+ res.setHeader(
2104
+ 'etag',
2105
+ `"${crypto.createHash('sha1').update(text).digest('hex')}"`,
2106
+ );
2134
2107
  res.setHeader('cache-control', 'public, max-age=60, must-revalidate');
2135
2108
  };
2136
2109
 
@@ -2238,12 +2211,8 @@ export function createApp({
2238
2211
  route(async (req, res) => {
2239
2212
  res.setHeader('access-control-allow-origin', '*');
2240
2213
  const categories = describeCategories(req);
2241
- // Tagged over the categories alone, deliberately. Express would happily
2242
- // hash the whole body for us, but `generatedAt` is a fresh timestamp on
2243
- // every request — so that tag would differ every time and no client
2244
- // could ever revalidate. What a consumer polls this for is whether the
2245
- // set of categories or the build each resolves to has moved, and that
2246
- // is exactly what this covers.
2214
+ // Over the categories alone: `generatedAt` changes every request, so a
2215
+ // tag covering it could never match.
2247
2216
  res.setHeader(
2248
2217
  'etag',
2249
2218
  `"${crypto.createHash('sha1').update(JSON.stringify(categories)).digest('hex')}"`,
@@ -2274,10 +2243,7 @@ export function createApp({
2274
2243
  });
2275
2244
  }
2276
2245
 
2277
- // The same background re-read the per-archive route does. Left out
2278
- // here, it missed exactly the archives that matter most: /latest/ is the
2279
- // URL a style points at, so a category anyone actually consumes through
2280
- // the documented path was the one place a stale summary never healed.
2246
+ // /latest/ is the URL a style points at, so it has to heal too.
2281
2247
  startMetadataBackfill(entry);
2282
2248
 
2283
2249
  res.setHeader('access-control-allow-origin', '*');
@@ -2287,8 +2253,7 @@ export function createApp({
2287
2253
  // makes that re-read cheap — a client already holding the current build
2288
2254
  // gets a 304 and no body, which Express does on its own once an ETag is
2289
2255
  // set before the response goes out.
2290
- tagAsLatest(res, entry);
2291
- res.json({
2256
+ const doc = {
2292
2257
  ...buildTileJson(entry, baseUrl(req)),
2293
2258
  // Names what it resolved to, so a consumer can tell one build from the
2294
2259
  // next without diffing the tile URLs.
@@ -2298,7 +2263,9 @@ export function createApp({
2298
2263
  name: entry.name,
2299
2264
  createdAt: entry.createdAt,
2300
2265
  },
2301
- });
2266
+ };
2267
+ tagAsLatest(res, doc);
2268
+ res.json(doc);
2302
2269
  }),
2303
2270
  );
2304
2271
 
@@ -2319,34 +2286,23 @@ export function createApp({
2319
2286
  // this one says so, and names what it resolved to while it is at it. The
2320
2287
  // tag is the infohash being redirected to, which is precisely the thing
2321
2288
  // that changes when this redirect starts pointing somewhere else.
2322
- tagAsLatest(res, entry);
2323
- res.redirect(
2324
- 302,
2325
- `${baseUrl(req)}/archives/${entry.infoHash}/archive.torrent`,
2326
- );
2289
+ const target = `${baseUrl(req)}/archives/${entry.infoHash}/archive.torrent`;
2290
+ tagAsLatest(res, target);
2291
+ res.redirect(302, target);
2327
2292
  });
2328
2293
 
2329
2294
  app.get('/latest/:category/magnet', (req, res) => {
2330
2295
  const entry = newestIn(req.params.category, req);
2331
2296
  if (!entry) return res.status(404).json({ error: 'no such category' });
2332
- tagAsLatest(res, entry);
2333
- res.type('text/plain').send(entry.magnet ?? '');
2297
+ const magnet = entry.magnet ?? '';
2298
+ tagAsLatest(res, magnet);
2299
+ res.type('text/plain').send(magnet);
2334
2300
  });
2335
2301
 
2336
2302
  /**
2337
- * The headers a browser must be told it may read from a range response.
2303
+ * Headers a cross-origin PMTiles reader needs and would not otherwise see.
2338
2304
  *
2339
- * Only a handful of response headers reach JavaScript on a cross-origin
2340
- * fetch, and not one of the headers a PMTiles reader depends on is among
2341
- * them. The official reader records the ETag of the first response and
2342
- * compares every response after it against that value — which is how it
2343
- * notices that the archive moved underneath a read in progress. Cross-origin
2344
- * and unexposed, it reads `null` instead, a comparison against null never
2345
- * fires, and it would splice two builds together in silence: the exact
2346
- * failure the tag exists to prevent. Content-Range earns its place through a
2347
- * narrower case — an archive smaller than the reader's first 16 KiB request
2348
- * is answered 416, and it parses the real length out of that header before
2349
- * retrying.
2305
+ * See docs/serving-tiles.md — "How a client knows the build moved".
2350
2306
  */
2351
2307
  const RANGE_CORS_HEADERS = {
2352
2308
  'access-control-allow-origin': '*',
@@ -2357,24 +2313,8 @@ export function createApp({
2357
2313
  /**
2358
2314
  * Serves a byte range for an archive this node does not hold whole.
2359
2315
  *
2360
- * **Experimental.** It closes the loop the cache mode was built for: the
2361
- * node holds no bytes, a reader asks for some, and the pieces covering them
2362
- * are pulled out of the swarm on demand — the same path the tile endpoint
2363
- * has always taken internally, one HTTP layer further out and sharing the
2364
- * same piece cache and open handle.
2365
- *
2366
- * It is off by default and should stay off for anything public, for reasons
2367
- * that are properties of the arrangement rather than of this code:
2368
- *
2369
- * - Every byte is somebody else's upload. A node in cache mode is not an
2370
- * origin; putting one behind a URL that looks like an origin turns each
2371
- * request into swarm traffic that this node neither paid for nor holds.
2372
- * - A piece read takes as long as the swarm takes. That is fine for a tile,
2373
- * which the reader asked for and will wait on, and poor for an HTTP client
2374
- * with its own patience.
2375
- * - There is no way to answer a request for the whole file that is not
2376
- * "download 700 GiB through BitTorrent and stream it out". So a range is
2377
- * required, and a large one is refused.
2316
+ * Experimental, off by default, and bounded: a Range is required and a large
2317
+ * one refused. See docs/configuration.md — `serveArchiveFromSwarm`.
2378
2318
  * @param {import('express').Request} req - The request, for its Range.
2379
2319
  * @param {import('express').Response} res - The response.
2380
2320
  * @param {object} entry - The archive.
@@ -2396,17 +2336,12 @@ export function createApp({
2396
2336
  });
2397
2337
  }
2398
2338
 
2399
- // A range, and only a range. Without one the answer is the whole archive,
2400
- // and the whole archive is not on this disk — serving it would mean
2401
- // pulling every piece through the swarm to stream it back out, which is
2402
- // both enormous and somebody else's bandwidth.
2339
+ // Without a Range the answer is the whole archive, pulled piece by piece
2340
+ // through the swarm to stream back out.
2403
2341
  const asked = req.headers.range;
2404
2342
  if (!asked) {
2405
2343
  res.setHeader('accept-ranges', 'bytes');
2406
- // 409 rather than 411. The request is well formed -- 411 is about a
2407
- // missing Content-Length on the way in -- and what is wrong is the state
2408
- // of the thing being asked for: the whole file is not here, and fetching
2409
- // it through the swarm to stream back out is not an answer.
2344
+ // 409, not 411: the request is well formed, the resource is not here.
2410
2345
  return res.status(409).json({
2411
2346
  error:
2412
2347
  'this node does not hold this archive, so it can only answer a ' +
@@ -2415,8 +2350,7 @@ export function createApp({
2415
2350
  }
2416
2351
 
2417
2352
  const ranges = parseRange(size, asked, { combine: true });
2418
- // -1 is unsatisfiable, -2 is malformed, and more than one range would mean
2419
- // a multipart response that nothing reading PMTiles has ever asked for.
2353
+ // -1 unsatisfiable, -2 malformed; multipart is not worth supporting here.
2420
2354
  if (ranges === -1) {
2421
2355
  res.setHeader('content-range', `bytes */${size}`);
2422
2356
  return res.status(416).json({ error: 'range not satisfiable' });
@@ -2437,9 +2371,8 @@ export function createApp({
2437
2371
  });
2438
2372
  }
2439
2373
 
2440
- // A deadline of its own, shorter than the piece timeout underneath. An
2441
- // HTTP client gives up long before libtorrent does, and a request left
2442
- // hanging on a piece nobody is seeding holds a connection open for nothing.
2374
+ // Shorter than the piece timeout underneath: an HTTP client gives up long
2375
+ // before libtorrent does.
2443
2376
  const controller = new AbortController();
2444
2377
  const deadline = setTimeout(
2445
2378
  () => controller.abort(),
@@ -2456,8 +2389,7 @@ export function createApp({
2456
2389
  for (const [name, value] of Object.entries(RANGE_CORS_HEADERS)) {
2457
2390
  res.setHeader(name, value);
2458
2391
  }
2459
- // The same tag a complete copy would answer with, because it is the same
2460
- // content: the infohash names these bytes wherever they were read from.
2392
+ // The same tag a complete copy answers with: same bytes, same name.
2461
2393
  res.setHeader('etag', `"${entry.infoHash}"`);
2462
2394
  res.setHeader('cache-control', 'public, max-age=31536000, immutable');
2463
2395
  res.send(body);
@@ -2476,30 +2408,13 @@ export function createApp({
2476
2408
  /**
2477
2409
  * The current build of a category, as a file, by byte range.
2478
2410
  *
2479
- * The same thing `/archives/<infohash>/archive.pmtiles` is, addressed the
2480
- * way a style or a long-lived configuration wants to address it: by what it
2481
- * is rather than by which build it happens to be. Point any PMTiles reader
2482
- * at this and it keeps working across rebuilds, with no infohash to chase.
2483
- *
2484
- * The hazard is the one the immutable route never has to think about. A
2485
- * PMTiles reader does not fetch a file; it fetches a header, then a root
2486
- * directory, then leaf directories, then tiles, over minutes or hours. If a
2487
- * rebuild lands partway through, the offsets it read from the old build
2488
- * address bytes in the new one. That does not fail loudly — it decodes as
2489
- * the wrong tile, or as nothing, with no error anywhere that names the
2490
- * cause. So the validator is the infohash and `If-Range` is honoured: a
2491
- * range conditioned on a build that is no longer current is refused as a
2492
- * range and answered in full, which a reader survives, instead of being
2493
- * spliced, which it does not.
2411
+ * `If-Range` is honoured so a rebuild landing mid-read cannot splice two
2412
+ * builds. See docs/serving-tiles.md — "How a client knows the build moved".
2494
2413
  */
2495
2414
  app.get('/latest/:category/archive.pmtiles', (req, res) => {
2496
2415
  const entry = newestIn(req.params.category, req);
2497
2416
  if (!entry) return res.status(404).json({ error: 'no such category' });
2498
2417
 
2499
- // Whether this node hands out whole archives at all is the operator's
2500
- // call, not a consequence of having the file. Everything else published
2501
- // here is either small or metered by the request; this is up to 700 GiB to
2502
- // anyone who knows an infohash, so it is asked for rather than assumed.
2503
2418
  if (!publishingFor(entry, config).serveArchive) {
2504
2419
  return res.status(403).json({
2505
2420
  error:
@@ -2508,10 +2423,8 @@ export function createApp({
2508
2423
  });
2509
2424
  }
2510
2425
 
2511
- // Not on this disk. Either the node can read it out of the swarm, which is
2512
- // what cache mode is for, or it says so — but it never answers a range
2513
- // from a sparse file, where the unwritten space reads as zeroes and looks
2514
- // exactly like data.
2426
+ // Never answered from the sparse file, where unwritten space reads as
2427
+ // zeroes and looks exactly like data.
2515
2428
  if (entry.complete === false) return serveFromSwarm(req, res, entry);
2516
2429
  const file = entry.savePath ? path.join(entry.savePath, entry.name) : null;
2517
2430
  if (!file) return res.status(404).json({ error: 'no file for it here' });
@@ -2566,16 +2479,15 @@ export function createApp({
2566
2479
  // validator: a tag meaning "still empty" is indistinguishable from one
2567
2480
  // meaning "still this build", and a reader would cache the emptiness past
2568
2481
  // the arrival of the thing it was waiting for.
2569
- if (entry) tagAsLatest(res, entry);
2570
- res.type('application/rss+xml').send(
2571
- renderFeed(entry ? [entry] : [], {
2572
- title: `${config.feedTitle ?? 'PMTiles archives'} — ${category}, latest`,
2573
- baseUrl: baseUrl(req),
2574
- copyright: config.feedCopyright,
2575
- category,
2576
- maxItems: 1,
2577
- }),
2578
- );
2482
+ const body = renderFeed(entry ? [entry] : [], {
2483
+ title: `${config.feedTitle ?? 'PMTiles archives'} — ${category}, latest`,
2484
+ baseUrl: baseUrl(req),
2485
+ copyright: config.feedCopyright,
2486
+ category,
2487
+ maxItems: 1,
2488
+ });
2489
+ if (entry) tagAsLatest(res, body);
2490
+ res.type('application/rss+xml').send(body);
2579
2491
  });
2580
2492
 
2581
2493
  app.get('/feed.xml', (req, res) => {
@@ -2839,10 +2751,6 @@ export function createApp({
2839
2751
  const entry = catalog.get(req.params.infoHash);
2840
2752
  if (!entry) return res.status(404).json({ error: 'not found' });
2841
2753
 
2842
- // Whether this node hands out whole archives at all is the operator's
2843
- // call, not a consequence of having the file. Everything else published
2844
- // here is either small or metered by the request; this is up to 700 GiB to
2845
- // anyone who knows an infohash, so it is asked for rather than assumed.
2846
2754
  if (!publishingFor(entry, config).serveArchive) {
2847
2755
  return res.status(403).json({
2848
2756
  error:
package/src/catalog.js CHANGED
@@ -58,42 +58,18 @@ export function normalizeCategories(source) {
58
58
  }
59
59
 
60
60
  /**
61
- * What this node offers of an archive's own bytes, over plain HTTP.
61
+ * What this node offers of an archive's own bytes over HTTP.
62
62
  *
63
- * Three separate decisions, because they are three separate exposures and
64
- * conflating them takes the choice away from whoever runs the node:
65
- *
66
- * - `serveArchive` — whether `/archives/<hash>/archive.pmtiles` answers at
67
- * all. This is the one that decides whether a stranger who knows an infohash
68
- * can pull 700 GiB off the box.
69
- * - `selfWebSeed` — whether this node's own URL is written into the torrent's
70
- * `url-list`, so every peer in the swarm fetches from it.
71
- * - `publicDownload` — whether the public catalogue page offers it as a
72
- * download. Serving a file to a reader that already knows the URL and
73
- * advertising it on a page are not the same act.
74
- *
75
- * The last two are ANDed with the first rather than merely defaulting from it.
76
- * A node cannot be a web seed for a file it will not serve, and a download link
77
- * that 403s is worse than no link — so a catalog edited by hand into that state
78
- * is read as the safe thing rather than obeyed into an incoherent one.
63
+ * Three separate exposures, resolved per archive then per node. See
64
+ * docs/configuration.md — "Offering the archive file itself".
79
65
  * @param {object} [entry] - A catalog entry, whose fields win where set.
80
66
  * @param {object} [config] - The node's defaults.
81
67
  * @returns {{serveArchive: boolean, selfWebSeed: boolean, publicDownload: boolean}} - Resolved.
82
68
  */
83
69
  export function publishingFor(entry, config) {
84
70
  const serveArchive = entry?.serveArchive ?? config?.serveArchive ?? false;
85
- // Whether the whole file is here. `serveArchive` does not depend on it —
86
- // `serveArchiveFromSwarm` exists precisely so a node holding nothing can
87
- // still answer a bounded range — but a download link does. A link labelled
88
- // "download" on a public page has to hand over the file, and a cache-mode
89
- // node cannot: it answers 409, or refuses a request with no Range, which
90
- // reads as a broken node rather than as a deliberate limit.
91
- //
92
- // `selfWebSeed` is guarded separately, in setPublishing, because it has a
93
- // lifecycle rather than a state: the URL is written into a .torrent and has
94
- // to be added and withdrawn at the right moments, not merely reported as
95
- // off. Resolving it to false here would withdraw a published seed the first
96
- // time an archive was rechecked.
71
+ // A download link has to hand over the file; a bounded range need not.
72
+ // selfWebSeed is guarded in setPublishing instead — it has a lifecycle.
97
73
  const held = entry?.complete !== false;
98
74
  return {
99
75
  serveArchive,
@@ -275,20 +251,10 @@ export class Catalog {
275
251
  }
276
252
  }
277
253
  /**
278
- * Whether a URL is one other people could plausibly fetch.
279
- *
280
- * A web seed is not a setting; it is written into the `.torrent` and the
281
- * magnet, and served to everyone who asks for either. Nothing rewrites it
282
- * afterwards — the file is handed out byte for byte — so a URL that names
283
- * this machine's own loopback interface is not a mistake that gets corrected
284
- * later. It is distributed, followed, and retried by every peer in the swarm
285
- * for as long as the torrent exists.
254
+ * Whether a URL is one other peers could plausibly fetch.
286
255
  *
287
- * Loopback is the only case refused outright, because it cannot be right for
288
- * anybody: `127.0.0.1` means the peer's own machine, not this one. A private
289
- * address is a different matter — a node syncing to its own peers across a LAN
290
- * is a real arrangement, and this is not the place to overrule it — so that is
291
- * reported rather than blocked.
256
+ * Loopback is refused; a private address is reported and allowed. See
257
+ * docs/configuration.md — `selfWebSeed`.
292
258
  * @param {string} url - The candidate web seed.
293
259
  * @returns {{ok: boolean, why?: string, warning?: string}} - Whether to
294
260
  * publish it, and what to say about it either way.
@@ -343,24 +309,10 @@ export function reachability(url) {
343
309
  }
344
310
 
345
311
  /**
346
- * The base a URL gets when it is going to outlive the request that made it.
347
- *
348
- * Almost every URL this node emits is worked out per request, on purpose: a
349
- * node answering on several domains should name itself as whichever one was
350
- * asked, and `publicUrl` is deliberately left unset to allow that. That is the
351
- * right answer for a TileJSON, a `.torrent` link or a style URL — read once,
352
- * by whoever asked, and correct for them.
353
- *
354
- * A web seed is not that. It is written into the `.torrent` and the magnet and
355
- * served byte for byte to everyone who asks for either, so it has to be one
356
- * address rather than whichever the last request happened to arrive on. This
357
- * is where that address comes from, in order of how deliberate it is:
358
- *
359
- * 1. Given outright — the field beside the switch in the console.
360
- * 2. `publishingUrl`, the node's answer for exactly this question.
361
- * 3. `publicUrl`, if a node has overridden everything anyway.
362
- * 4. The request, which is a guess, but the same guess every other URL makes.
312
+ * The base for a URL that will outlive the request that made it.
363
313
  *
314
+ * Given outright, then `publishingUrl`, then `publicUrl`, then the request.
315
+ * See docs/configuration.md — `publishingUrl`.
364
316
  * @param {object} [options] - `explicit`, `config` and `requestBase`.
365
317
  * @returns {string} - A base with no trailing slash, or an empty string.
366
318
  */
@@ -372,9 +324,7 @@ export function publishingBase({ explicit, config, requestBase } = {}) {
372
324
  requestBase,
373
325
  ];
374
326
  for (const candidate of candidates) {
375
- // An empty or whitespace value means "not set", not "use an empty base" —
376
- // the same reading `publicUrl` gets, and for the same reason: clearing a
377
- // key by emptying it is what an operator naturally does to a JSON file.
327
+ // Empty means "not set", not "use an empty base".
378
328
  const value = String(candidate ?? '').trim();
379
329
  if (value) return value.replace(/\/+$/, '');
380
330
  }
package/src/config.js CHANGED
@@ -166,82 +166,29 @@ const DEFAULTS = {
166
166
  */
167
167
  md5: false,
168
168
  /**
169
- * Answer `/archives/<infohash>/archive.pmtiles` — the archive itself, by
170
- * byte range, which is what every PMTiles reader actually wants.
171
- *
172
- * Off by default, and deliberately so: this is the whole file, and an
173
- * archive here can be 700 GiB. Everything else this node publishes is either
174
- * small (TileJSON, a .torrent, a feed) or metered by the request (one tile),
175
- * so turning a node on has never meant offering its disk to anyone who knows
176
- * an infohash. This would, so it is asked for rather than assumed.
177
- *
178
- * The node's answer, not the only one: a watched folder, a scheduled source
179
- * or an individual archive may carry its own, and is obeyed either way.
169
+ * Answer `/archives/<infohash>/archive.pmtiles`. Off by default: it is the
170
+ * whole file. See docs/configuration.md — `serveArchive`.
180
171
  */
181
172
  serveArchive: false,
182
173
  /**
183
- * Write this node's own archive URL into the torrent's `url-list`, so every
184
- * peer in the swarm can fetch from it over HTTP.
185
- *
186
- * A web seed is the difference between a cold tile taking tens of seconds
187
- * and taking under one, and it is what makes an archive usable before it has
188
- * any peers at all. It is also an open invitation: a seed URL is followed by
189
- * everyone who holds the torrent, not only by people who visit this node.
190
- *
191
- * Means nothing without `serveArchive`, and is read as off without it — a
192
- * web seed URL that refuses the request is worse than none, because a client
193
- * spends its retries on it.
174
+ * Publish this node as a web seed for the archives it holds. Read as off
175
+ * without `serveArchive`. See docs/configuration.md — `selfWebSeed`.
194
176
  */
195
177
  selfWebSeed: false,
196
178
  /**
197
- * Offer the archive as a download on the public catalogue page.
198
- *
199
- * Separate from `serveArchive` on purpose. Serving a file to a reader that
200
- * was given the URL and advertising it on a page are different acts, and a
201
- * node can reasonably want the first without the second: the endpoint exists
202
- * for a style or a peer, and putting a 700 GiB link in front of every casual
203
- * visitor is a different decision. Read as off without `serveArchive`, since
204
- * a link that answers 403 is worse than no link.
179
+ * Offer the archive as a download on the public catalogue page. Needs
180
+ * `serveArchive` and a complete copy. See docs/configuration.md.
205
181
  */
206
182
  publicDownload: false,
207
183
  /**
208
- * Answer a byte range for an archive this node does not hold, by pulling the
209
- * covering pieces out of the swarm on demand.
210
- *
211
- * **Experimental, and not recommended for anything public.** It closes the
212
- * loop cache mode was built for — the node holds no bytes, a reader asks for
213
- * some, and the pieces arrive from peers — over the same path the tile
214
- * endpoint has always used internally, sharing its piece cache and its open
215
- * handle. As a way to point an ordinary PMTiles reader at an archive nothing
216
- * here has a copy of, it works.
217
- *
218
- * The reservations are properties of the arrangement rather than of the
219
- * implementation:
220
- *
221
- * - Every byte is somebody else's upload. A cache-mode node is not an origin,
222
- * and putting one behind a URL that looks like one turns each request into
223
- * swarm traffic it neither paid for nor holds.
224
- * - A piece read takes as long as the swarm takes. Acceptable for a tile,
225
- * which a reader asked for and will wait on; poor for an HTTP client with
226
- * its own patience and its own timeout.
227
- * - There is no honest answer to a request for the whole file, so a Range
228
- * header is required and a large one is refused.
184
+ * Answer a byte range for an archive this node does not hold, from the
185
+ * swarm. Experimental, and not for anything public — see
186
+ * docs/configuration.md — `serveArchiveFromSwarm`.
229
187
  */
230
188
  serveArchiveFromSwarm: false,
231
- /**
232
- * The largest range `serveArchiveFromSwarm` will fetch at once, in bytes.
233
- *
234
- * Sized for what a PMTiles reader actually asks for — a 16 KiB header, a
235
- * directory, a tile — rather than for bulk transfer, which is the thing this
236
- * must not quietly become.
237
- */
189
+ /** The largest range `serveArchiveFromSwarm` will fetch at once, in bytes. */
238
190
  swarmRangeLimitBytes: 8 * 1024 * 1024,
239
- /**
240
- * How long to wait for the swarm before giving up on a range, in
241
- * milliseconds. Shorter than the piece timeout underneath on purpose: an
242
- * HTTP client gives up long before libtorrent does, and a request left
243
- * hanging on a piece nobody is seeding holds a connection open for nothing.
244
- */
191
+ /** How long to wait for the swarm before giving up on a range, in ms. */
245
192
  swarmRangeTimeoutMs: 30000,
246
193
  /**
247
194
  * How long an unfinished download is kept before startup treats it as
@@ -274,23 +221,9 @@ const DEFAULTS = {
274
221
  /** Public base URL, used to build absolute links in the RSS feed and TileJSON. */
275
222
  publicUrl: undefined,
276
223
  /**
277
- * The one address to use for URLs that outlive the request that made them.
278
- *
279
- * Almost every URL this node emits is worked out per request, deliberately: a
280
- * node answering on several domains should name itself as whichever one was
281
- * asked for, and leaving `publicUrl` unset is what allows that. Read once, by
282
- * whoever asked, a per-request answer is the correct answer.
283
- *
284
- * A web seed is not read once. It is written into the `.torrent` and the
285
- * magnet, served byte for byte to everyone who fetches either, and never
286
- * rewritten — so it has to be one address rather than whichever the last
287
- * request happened to arrive on. Set this to that address.
288
- *
289
- * Narrower than `publicUrl` on purpose. `publicUrl` overrides every URL the
290
- * node emits and so gives up the multi-domain behaviour; this overrides only
291
- * the ones that have to be permanent. Unset falls back to `publicUrl`, then
292
- * to the request — and the console lets the address be typed at the moment
293
- * the switch is turned on, which beats both.
224
+ * The address for URLs that outlive the request that made them — today, web
225
+ * seeds. Narrower than `publicUrl`, which overrides every URL and so gives
226
+ * up answering on several domains. See docs/configuration.md.
294
227
  */
295
228
  publishingUrl: undefined,
296
229
  /**
@@ -430,20 +430,30 @@ export class CompositeEngine {
430
430
  * therefore re-hashed its whole store on every start, which for 800 GB is
431
431
  * half an hour of disk before it serves anything.
432
432
  * @param {string} [infoHash] - One archive, or all of them when omitted.
433
- * @returns {Promise<void>} - Resolves once every engine has been asked.
433
+ * @returns {Promise<{written: number, asked: number}>} - Totals across every
434
+ * engine that keeps resume data.
434
435
  */
435
436
  async saveResume(infoHash) {
437
+ // Summed rather than dropped, so a caller can tell the difference between
438
+ // "every torrent wrote" and "half of them will be re-hashed on the next
439
+ // start". An engine that fails outright contributes nothing to either
440
+ // total, which is right: it was never asked in a way that counted.
441
+ let written = 0;
442
+ let asked = 0;
436
443
  for (const engine of [this.#primary, ...this.#secondaries]) {
437
444
  // WebTorrent keeps none, and says so by not offering the method.
438
445
  if (!engine.saveResume) continue;
439
446
  try {
440
- await engine.saveResume(infoHash);
447
+ const result = (await engine.saveResume(infoHash)) ?? {};
448
+ written += Number(result.written) || 0;
449
+ asked += Number(result.asked) || 0;
441
450
  } catch (error) {
442
451
  console.warn(
443
452
  `[composite] ${engine.name} could not save resume data: ${error.message}`,
444
453
  );
445
454
  }
446
455
  }
456
+ return { written, asked };
447
457
  }
448
458
 
449
459
  /**
@@ -759,10 +759,17 @@ export class LibtorrentEngine {
759
759
  /**
760
760
  * Persists resume data, so the next start skips re-hashing the store.
761
761
  * @param {string} [infoHash] - One torrent, or all when omitted.
762
- * @returns {Promise<void>} - Resolves once saved.
762
+ * @returns {Promise<{written: number, asked: number}>} - How many torrents
763
+ * were told to write resume data, and how many actually did before the
764
+ * deadline.
763
765
  */
764
766
  async saveResume(infoHash) {
765
- await this.#call('save_resume', { infoHash });
767
+ // Returned rather than discarded. The sidecar reports both numbers, and
768
+ // the gap between them is the thing worth knowing: a torrent that did not
769
+ // write is one that gets re-hashed on the next start, which for a 700 GiB
770
+ // archive is the difference between seeding in seconds and seeding in half
771
+ // an hour. That answer was being thrown away here.
772
+ return this.#call('save_resume', { infoHash });
766
773
  }
767
774
 
768
775
  /**
package/src/index.js CHANGED
@@ -529,6 +529,21 @@ PMTILES_SWARM_PUBLIC_URL
529
529
  resumeTimer = setInterval(() => {
530
530
  engine
531
531
  .saveResume()
532
+ .then((result) => {
533
+ // Said out loud, because the shortfall is the thing that costs.
534
+ // A torrent that did not write inside the deadline is one that gets
535
+ // re-hashed on the next start, and for a 700 GiB archive that is the
536
+ // difference between seeding in seconds and seeding in half an hour
537
+ // -- which is exactly what "why is everything at 0%" looks like from
538
+ // the outside. Silence here is what made that hard to see.
539
+ const { written = 0, asked = 0 } = result ?? {};
540
+ if (asked > 0 && written < asked) {
541
+ console.warn(
542
+ `[resume] ${written} of ${asked} torrents wrote resume data ` +
543
+ 'in time; the rest will be re-hashed on the next start',
544
+ );
545
+ }
546
+ })
532
547
  .catch((error) =>
533
548
  console.warn(`[resume] could not save: ${error.message}`),
534
549
  );
package/src/library.js CHANGED
@@ -528,17 +528,8 @@ export class Library {
528
528
  async finalize(infoHash) {
529
529
  const settled = await this.#finalizeOnce(infoHash);
530
530
 
531
- // Only now, and never at import. A web seed URL for an archive that is
532
- // still arriving answers 409, and a peer handed a URL that refuses spends
533
- // its retries on it — worse than no web seed at all, and unfixable
534
- // afterwards, because the URL is in the .torrent every peer holds. So a
535
- // subscription records the intention when it joins and it is acted on
536
- // here, at the first moment this node actually holds the whole file.
537
- //
538
- // Safe to reach on an archive that was already complete: setPublishing
539
- // works from what is on record rather than from a transition, so this does
540
- // nothing the second time — and does the right thing the first time for an
541
- // archive that finished before the setting existed.
531
+ // Only now: a web seed for an archive still arriving answers 409, and the
532
+ // URL cannot be taken back out of the torrents peers already hold.
542
533
  if (settled && publishingFor(settled, this.#config).selfWebSeed) {
543
534
  try {
544
535
  await this.setPublishing(infoHash, {});
@@ -2713,13 +2704,9 @@ export class Library {
2713
2704
  /**
2714
2705
  * Rewrites an archive's web seed list, in the .torrent and everywhere else.
2715
2706
  *
2716
- * This is safe on a torrent already in circulation. BEP 19's `url-list` is a
2717
- * top-level key in the metainfo and the infohash covers only the `info`
2718
- * dictionary, so changing one leaves the infohash untouched — every magnet,
2719
- * peer and published reference stays valid. The check below asserts that
2720
- * rather than trusting it: if a rewrite ever did change the infohash, the
2721
- * result would be a different torrent wearing the old one's name, which is
2722
- * worth refusing loudly.
2707
+ * Safe on a torrent in circulation: `url-list` sits outside the info
2708
+ * dictionary, so the infohash is untouched. Asserted below rather than
2709
+ * trusted.
2723
2710
  * @param {object} entry - The catalog entry to rewrite.
2724
2711
  * @param {string[]} wanted - The complete new list.
2725
2712
  * @returns {Promise<{webSeeds: string[], parsed: object}>} - The new list and
@@ -2821,13 +2808,10 @@ export class Library {
2821
2808
  }
2822
2809
 
2823
2810
  /**
2824
- * Drops web seeds from an archive, leaving everything else about it alone.
2811
+ * Drops web seeds from an archive, leaving everything else alone.
2825
2812
  *
2826
- * The running engine is not told. libtorrent has no "forget this url seed"
2827
- * that every version answers to, and the consequence of it keeping one is
2828
- * bounded — it retries a URL that refuses and gives up on it. What matters
2829
- * is that the .torrent and the magnet stop handing the URL to anybody new,
2830
- * which is what this does.
2813
+ * The running engine is not told: there is no portable way to retract a url
2814
+ * seed, and what matters is that nobody new is handed it.
2831
2815
  * @param {string} infoHash - The archive.
2832
2816
  * @param {string[]} urls - Web seed URLs to drop. Unknown ones are ignored.
2833
2817
  * @returns {Promise<{webSeeds: string[], dropped: string[]}>} - What is left,
@@ -2857,17 +2841,9 @@ export class Library {
2857
2841
  /**
2858
2842
  * Sets what this node offers of one archive's own bytes over HTTP.
2859
2843
  *
2860
- * Three separate switches — see `publishingFor` in catalog.js for why they
2861
- * are separate — and one side effect: turning `selfWebSeed` on writes this
2862
- * node's own archive URL into the torrent's `url-list`, and turning it off
2863
- * takes that URL back out. The URL used is remembered, because the base can
2864
- * change underneath a node and removing "whatever we would build today"
2865
- * would leave yesterday's URL in the torrent for ever.
2866
- *
2867
- * Turning `serveArchive` off takes the other two with it, and that is not a
2868
- * quiet tidy-up: it means a URL already handed to every peer holding the
2869
- * torrent stops answering. The caller is told what was withdrawn so it can
2870
- * say so.
2844
+ * `selfWebSeed` has a side effect: it writes this node's URL into the
2845
+ * torrent, and takes it out again. See docs/configuration.md — "Offering the
2846
+ * archive file itself".
2871
2847
  * @param {string} infoHash - The archive.
2872
2848
  * @param {object} changes - Any of the three, as booleans. Anything else is
2873
2849
  * ignored, so a caller may send only what it is changing.
@@ -2878,20 +2854,14 @@ export class Library {
2878
2854
  const entry = this.#catalog.get(infoHash);
2879
2855
  if (!entry) throw new Error('unknown archive');
2880
2856
 
2881
- // Only what was actually asked for is recorded. An archive that says
2882
- // nothing about a setting goes on deferring to the node, which is what
2883
- // makes changing the node's answer reach the archives that never had one
2884
- // of their own.
2857
+ // Only what was asked for: an archive that says nothing defers to the node.
2885
2858
  const wanted = { ...entry };
2886
2859
  for (const key of ['serveArchive', 'selfWebSeed', 'publicDownload']) {
2887
2860
  if (typeof changes[key] === 'boolean') wanted[key] = changes[key];
2888
2861
  }
2889
2862
 
2890
- // With serving off the other two cannot be true, and writing that down
2891
- // matters more than it looks: left as a latent `true`, either would spring
2892
- // back the moment serving was turned on again — re-publishing this node as
2893
- // a web seed, or re-listing a 700 GiB download, as a side effect of a
2894
- // decision about something else.
2863
+ // Written down rather than left latent, or either would spring back the
2864
+ // next time serving was turned on.
2895
2865
  if (!publishingFor(wanted, this.#config).serveArchive) {
2896
2866
  wanted.selfWebSeed = false;
2897
2867
  wanted.publicDownload = false;
@@ -2899,21 +2869,14 @@ export class Library {
2899
2869
  const after = publishingFor(wanted, this.#config);
2900
2870
  await this.#catalog.put(wanted);
2901
2871
 
2902
- // Driven by what is on record rather than by the transition, so calling
2903
- // this twice does nothing the second time and calling it on an archive
2904
- // that was created with the setting already on still writes the seed. A
2905
- // transition test looked equivalent and was not: an import that inherits
2906
- // `selfWebSeed: true` from the node has no "before" in which it was off,
2907
- // so nothing would ever have published it.
2872
+ // Driven by what is on record, not by a transition: an import that
2873
+ // inherits the setting has no "before" in which it was off.
2908
2874
  const published = entry.selfWebSeedUrl ?? null;
2909
2875
  let webSeed = published;
2910
2876
  let warning = null;
2911
2877
 
2912
- // A node cannot be a web seed for bytes it does not have. It might be able
2913
- // to *answer* for them, where serveArchiveFromSwarm is on -- but answering
2914
- // by fetching from the swarm and then advertising that to the swarm is a
2915
- // loop with an amplifier in it: every peer that takes the seed makes this
2916
- // node download the piece again to serve it.
2878
+ // Answering from the swarm and advertising that to the swarm is a loop
2879
+ // with an amplifier in it.
2917
2880
  if (after.selfWebSeed && !published && entry.complete === false) {
2918
2881
  throw new Error(
2919
2882
  'this node does not hold a complete copy of this archive, so it ' +
@@ -2937,10 +2900,7 @@ export class Library {
2937
2900
  );
2938
2901
  }
2939
2902
  webSeed = `${base}/archives/${infoHash}/archive.pmtiles`;
2940
- // Checked before it goes anywhere. Nothing rewrites a web seed once it
2941
- // is in a .torrent -- the file is served byte for byte to everyone who
2942
- // asks -- so a URL that cannot work is not a mistake that gets corrected
2943
- // on the next request. It is distributed and then retried for ever.
2903
+ // Nothing rewrites a web seed once it is in a .torrent.
2944
2904
  const reach = reachability(webSeed);
2945
2905
  if (!reach.ok) throw new Error(reach.why);
2946
2906
  if (reach.warning) console.warn(`[web seed] ${reach.warning}`);
@@ -2952,10 +2912,7 @@ export class Library {
2952
2912
  // and then vanish in the same call.
2953
2913
  await this.#catalog.put({ infoHash, selfWebSeedUrl: webSeed });
2954
2914
  } else if (!after.selfWebSeed && published) {
2955
- // Whatever was actually published, which is not necessarily the URL this
2956
- // node would build for itself today: the base can change underneath a
2957
- // node, and removing "whatever we would say now" would leave yesterday's
2958
- // URL in the torrent for ever.
2915
+ // What was published, not what this node would build today.
2959
2916
  await this.removeWebSeeds(infoHash, [published]);
2960
2917
  await this.#catalog.put({ infoHash, selfWebSeedUrl: null });
2961
2918
  webSeed = null;
@@ -2964,12 +2921,8 @@ export class Library {
2964
2921
  return {
2965
2922
  ...after,
2966
2923
  webSeed,
2967
- // Published anyway, and said out loud. A node syncing to its own peers
2968
- // across a LAN is a real arrangement and not this code's to overrule.
2969
2924
  warning,
2970
- // Named so a caller can warn about the one change that is not merely a
2971
- // setting moving: this URL is already in the hands of every peer that
2972
- // holds the torrent, and they will go on trying it for a while.
2925
+ // Peers holding the torrent keep trying the URL until they refresh.
2973
2926
  withdrewWebSeed: Boolean(published) && !after.selfWebSeed,
2974
2927
  };
2975
2928
  }
@@ -3165,12 +3118,8 @@ export class Library {
3165
3118
  stale: false,
3166
3119
  });
3167
3120
 
3168
- // Publishing this node as a web seed is the one of the three with an
3169
- // effect beyond a stored boolean: it writes a URL into the .torrent and
3170
- // the magnet. Done here so a watched folder and an RSS subscription get it
3171
- // without each having to know, and warned rather than thrown because an
3172
- // import that produced a good archive should not be failed by a seed URL
3173
- // that could not be built.
3121
+ // Here rather than in each caller, and warned rather than thrown: a good
3122
+ // archive should not be failed by a seed URL that could not be built.
3174
3123
  if (publishingFor(entry, this.#config).selfWebSeed) {
3175
3124
  try {
3176
3125
  await this.setPublishing(entry.infoHash, {});
@@ -18,6 +18,8 @@ const TILE_TYPES = {
18
18
  3: { format: 'jpeg', contentType: 'image/jpeg' },
19
19
  4: { format: 'webp', contentType: 'image/webp' },
20
20
  5: { format: 'avif', contentType: 'image/avif' },
21
+ // No MIME is registered for MLT, and nothing selects a tile by content type.
22
+ 6: { format: 'mlt', contentType: 'application/octet-stream' },
21
23
  };
22
24
 
23
25
  /**
@@ -59,34 +61,23 @@ export function metadataFlag(value) {
59
61
  }
60
62
 
61
63
  /**
62
- * What a `raster-dem` archive says about how its elevations are packed.
64
+ * The tile encoding an archive declares, if it is one MapLibre understands.
63
65
  *
64
- * Nothing in a PMTiles header says this — the header knows the tile is WebP,
65
- * not what the three channels mean — so the only place it can come from is the
66
- * archive's own metadata. Without it a consumer falls back to a default, and
67
- * the default is wrong for exactly the archives that most need to say
68
- * something: a terrarium-packed DEM read as `mapbox` decodes every mountain
69
- * into noise, silently, with a plausible-looking map on screen.
70
- *
71
- * The values are the ones the style specification defines, and anything else
72
- * is dropped rather than passed on. A tile client handed an encoding it does
73
- * not recognise is worse off than one handed nothing, because nothing at least
74
- * leaves it free to use its own default.
66
+ * See docs/tilejson.md — `encoding`.
75
67
  * @param {unknown} value - Whatever the metadata held.
76
68
  * @returns {string | undefined} - A known encoding, or undefined.
77
69
  */
78
70
  export function metadataEncoding(value) {
79
71
  if (typeof value !== 'string') return undefined;
80
72
  const text = value.trim().toLowerCase();
81
- return ['terrarium', 'mapbox', 'custom'].includes(text) ? text : undefined;
73
+ // An unrecognised value is dropped: it would cost a client its own default.
74
+ return ['terrarium', 'mapbox', 'custom', 'mlt'].includes(text)
75
+ ? text
76
+ : undefined;
82
77
  }
83
78
 
84
79
  /**
85
- * The numbers a `custom` encoding is meaningless without.
86
- *
87
- * `encoding: "custom"` says "the channels mean what these four factors say",
88
- * so carrying the word and not the factors publishes an archive nobody can
89
- * read. They travel together or not at all.
80
+ * The four factors a `custom` encoding is unreadable without. All or nothing.
90
81
  * @param {object} metadata - The archive's metadata.
91
82
  * @returns {object | undefined} - The four factors, or undefined.
92
83
  */
@@ -111,22 +102,11 @@ export function customEncodingFactors(metadata) {
111
102
  * @returns {Promise<PMTilesSummary>} - The summary.
112
103
  */
113
104
  /**
114
- * What this prober reads, as a number that goes up when that changes.
115
- *
116
- * A summary is stored in the catalog and never read again, which is right --
117
- * re-reading a header out of the swarm is not free, and the answer does not
118
- * change for a given infohash. It goes wrong the moment the prober learns to
119
- * read something new: every archive probed before that keeps a summary with a
120
- * hole in it, for ever, and the only way out was to remove and re-add it.
121
- *
122
- * `encoding` was the case that made this obvious. It sat in the metadata of
123
- * archives this node had been serving for months, and adding the code to read
124
- * it changed nothing at all, because nothing ever asked again.
125
- *
126
- * So: stamp what the prober knew at the time, and re-read once when that is
127
- * behind. Raise this whenever a field is added to the summary below.
105
+ * What this prober reads. Raise it whenever a field is added to the summary,
106
+ * and archives probed by an older build are re-read once. See
107
+ * docs/internals.md — "Re-reading a summary an older prober wrote".
128
108
  */
129
- export const SUMMARY_VERSION = 2;
109
+ export const SUMMARY_VERSION = 3;
130
110
 
131
111
  export function summarize(header, metadata = {}) {
132
112
  const type = TILE_TYPES[header.tileType] ?? TILE_TYPES[0];
@@ -165,12 +145,9 @@ export function summarize(header, metadata = {}) {
165
145
  // the same key, so an archive built to be served there carries the answer
166
146
  // with it and does not have to be configured again here.
167
147
  sparse: metadataFlag(metadata.sparse),
168
- // How a raster-dem archive packs elevation into pixels: terrarium, mapbox
169
- // or custom. Read from the archive rather than configured here, for the
170
- // same reason `sparse` is — the archive knows, and a mirror of it should
171
- // not have to be told again.
172
- encoding: metadataEncoding(metadata.encoding),
173
- // Only where the encoding is custom, since they mean nothing otherwise.
148
+ // The header settles MLT; only the metadata can settle elevation packing.
149
+ encoding:
150
+ type.format === 'mlt' ? 'mlt' : metadataEncoding(metadata.encoding),
174
151
  encodingFactors:
175
152
  metadataEncoding(metadata.encoding) === 'custom'
176
153
  ? customEncodingFactors(metadata)
@@ -889,17 +889,30 @@
889
889
  <script type="module">
890
890
  const $ = (id) => document.getElementById(id);
891
891
 
892
- /** Formats a byte count for people rather than machines. */
892
+ /** What this page shows in place of a size it has not got. */
893
+ const EMPTY = '—';
894
+ /**
895
+ * Formats a byte count for people rather than machines.
896
+ *
897
+ * Kept character-for-character identical to the same helper on the other
898
+ * page, and held there by a test. They drifted once: this one rounded to
899
+ * whole units above ten and the other always kept a decimal, so the same
900
+ * archive read as 81 GiB on one page and 80.6 GiB on the other. Two
901
+ * numbers for one fact is worse than either number.
902
+ *
903
+ * One decimal from KiB up, because these are mostly archive sizes and
904
+ * half a gigabyte of difference is worth seeing.
905
+ */
893
906
  const bytes = (n) => {
894
- if (!Number.isFinite(n) || n <= 0) return '—';
907
+ if (!Number.isFinite(n) || n <= 0) return EMPTY;
895
908
  const units = ['B', 'KiB', 'MiB', 'GiB', 'TiB'];
896
909
  let value = n;
897
910
  let unit = 0;
898
911
  while (value >= 1024 && unit < units.length - 1) {
899
912
  value /= 1024;
900
- unit++;
913
+ unit += 1;
901
914
  }
902
- return `${value.toFixed(value < 10 && unit > 0 ? 1 : 0)} ${units[unit]}`;
915
+ return `${value.toFixed(unit === 0 ? 0 : 1)} ${units[unit]}`;
903
916
  };
904
917
 
905
918
  const rate = (n) => (n > 0 ? `${bytes(n)}/s` : '—');
@@ -262,8 +262,22 @@
262
262
  return node;
263
263
  };
264
264
 
265
+ /** What this page shows in place of a size it has not got. */
266
+ const EMPTY = null;
267
+ /**
268
+ * Formats a byte count for people rather than machines.
269
+ *
270
+ * Kept character-for-character identical to the same helper on the other
271
+ * page, and held there by a test. They drifted once: this one rounded to
272
+ * whole units above ten and the other always kept a decimal, so the same
273
+ * archive read as 81 GiB on one page and 80.6 GiB on the other. Two
274
+ * numbers for one fact is worse than either number.
275
+ *
276
+ * One decimal from KiB up, because these are mostly archive sizes and
277
+ * half a gigabyte of difference is worth seeing.
278
+ */
265
279
  const bytes = (n) => {
266
- if (!Number.isFinite(n) || n <= 0) return null;
280
+ if (!Number.isFinite(n) || n <= 0) return EMPTY;
267
281
  const units = ['B', 'KiB', 'MiB', 'GiB', 'TiB'];
268
282
  let value = n;
269
283
  let unit = 0;