pmtiles-swarm 0.94.0 → 0.96.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -7,6 +7,75 @@
7
7
  ### 🐞 Bug fixes
8
8
  - _...Add new stuff here..._
9
9
 
10
+ ## 0.96.0
11
+ ### ✨ Features and improvements
12
+ - **The swarm-read limit is memory now, not a count.** `tiles.maxOpenSwarmArchives: 16` becomes
13
+ `tiles.swarmCacheBytes: 1 GiB`, because a count was the wrong unit for it. What is expensive about
14
+ a cache-mode reader is its piece cache, which is **RAM** — a map of whole pieces in the node's own
15
+ heap — and how big it is depends on the torrent: `max(64 MiB, 8 × pieceLength)`. Sixteen readers
16
+ is a gigabyte against the 4 MiB pieces this project creates and two against the 16 MiB pieces a
17
+ planet torrent usually has, which is not a limit anybody chose.
18
+
19
+ Each reader is asked what its cache costs rather than assumed, so the count now follows from the
20
+ memory: halve `pieceCacheBytes` and twice as many archives stay open for the same gigabyte. The
21
+ last reader is never closed, whatever the budget says — the archive just opened is the one being
22
+ read. `maxOpenSwarmArchives` remains for a node that wants a hard count as well; unset by default.
23
+
24
+ Complete archives are unaffected: they hold a file descriptor and no piece cache, and are bounded
25
+ by `maxOpenArchives`.
26
+ - **`init --systemd` regenerates the unit without touching the configuration.** The unit is derived
27
+ from the configuration and from how many archives the library holds, and both move — a watched
28
+ folder added later belongs in `ReadWritePaths`, and a grown library needs longer to write its
29
+ resume data than the unit allows. Running it again used to be refused, which sent people to
30
+ `--force`, which replaces `swarm.config.json` outright: tokens, stacks, feeds and all, to
31
+ regenerate a file beside it. It now reads the configuration, writes only the unit, and says so —
32
+ along with the `diff` to run before replacing what is installed, since `ReadWritePaths` is derived
33
+ and a path added to the installed unit by hand will not be in the new one.
34
+
35
+ ### 🐞 Bug fixes
36
+ - **The unit's `ReadWritePaths` left out where an export writes.** `stackExports[].savePath` and
37
+ `.publishDir`, and a watched folder's `publishDir`, were not among the directories derived from
38
+ the configuration — so a nightly bake to a directory named nowhere else ran for an hour and was
39
+ refused the write at the end, with that directory's permissions perfect. Which is the exact
40
+ failure the derived list exists to prevent. Re-run `init --systemd` to pick them up.
41
+
42
+ ## 0.95.0
43
+ ### ✨ Features and improvements
44
+ - **The defaults now assume a library of hundreds, not a handful.** `tiles.maxOpenArchives` goes
45
+ from 16 to 128 and `tiles.directoryCacheEntries` from 200 to 2000. Sixteen was set for the most
46
+ expensive kind of open archive and then applied to every kind, which is wrong for the node this is
47
+ increasingly used to build: a stack assembled from a provider's file index names several hundred
48
+ sources, and a bake walks every one of them. At sixteen such a run spent most of its time
49
+ reopening archives it had just closed, and each reopen re-reads a header and a directory.
50
+
51
+ So the limit is now three limits, because the handles cost different things. A complete archive is
52
+ a file descriptor and the unit allows 65535 of them. A **cache-mode** archive carries a piece cache
53
+ sized from the torrent's piece length — at 16 MiB pieces a hundred of them is gigabytes — so
54
+ `tiles.maxOpenSwarmArchives` keeps the old ceiling of 16 and is counted separately. A **URL or
55
+ bucket** archive holds an HTTP reader and nothing else, and costs a network round trip to reopen,
56
+ so `tiles.maxOpenRemoteArchives` is 64. A node serving only cache-mode archives behaves as before.
57
+
58
+ - **The unit's stop timeout is derived from the library it was written for.** Stopping writes resume
59
+ data for every archive and the node allows two seconds apiece, so a fixed five minutes covered a
60
+ library of 145 and no more — silently, because outgrowing it produces no error, just a library
61
+ that comes back at 0% and re-hashes for hours. `init --systemd` now counts the catalog and sizes
62
+ `TimeoutStopSec` to fit, and a node whose library has outgrown a default unit says so at startup.
63
+
64
+
65
+ ### 🐞 Bug fixes
66
+ - **A sidecar left running by a previous start is now killed rather than fought.** This is the
67
+ restart that has to be done two or three times before it takes. The sidecar exits when its pipe
68
+ closes, which covers an orderly stop — but not one in the middle of hashing, since libtorrent does
69
+ not hand control back until the check finishes, and not a node killed outright. What is left holds
70
+ the listen port and the resume directory, the next start fails against it, and since a failed
71
+ start leaves its own sidecar they accumulate.
72
+
73
+ Stopping now insists: the shutdown request goes first so resume data is still saved, and a sidecar
74
+ that has not gone within a few seconds is killed rather than left behind. And a start reaps what a
75
+ previous run left before spawning anything — by pid, but only where `/proc` confirms that pid is
76
+ really a sidecar, since pids are reused and killing whatever inherited one would be worse than the
77
+ problem.
78
+
10
79
  ## 0.94.0
11
80
  ### 🐞 Bug fixes
12
81
  - **Saving settings rewrote a proxy list nobody had touched.** A trusted-proxy list may be stored as
@@ -724,20 +724,70 @@ through the swarm, pulling only the pieces a requested tile lives in — which i
724
724
  what lets a machine with 10 GiB free serve a 700 GiB planet. See
725
725
  [serving-tiles.md](serving-tiles.md).
726
726
 
727
- | setting | default | |
728
- | ----------------------------- | -------- | ---------------------------------------------------------------- |
729
- | `tiles.maxOpenArchives` | `16` | each holds a descriptor or a torrent reader plus its piece cache |
730
- | `tiles.directoryCacheEntries` | `200` | header and directory cache, shared across archives |
731
- | `tiles.pieceCacheBytes` | unset | sized from the torrent's piece length when unset |
732
- | `tiles.hydrateIdleMs` | unset | idle time before background hydration resumes |
733
- | `tiles.pieceTimeoutMs` | `120000` | how long to wait for one piece |
734
- | `tiles.readyTimeoutMs` | `60000` | how long to wait for torrent metadata |
735
- | `tiles.metadataTimeoutMs` | `120000` | how long a background metadata read may take |
736
- | `tiles.headerTimeoutMs` | `12000` | how long a TileJSON request waits for a header |
737
- | `tiles.sparse` | unset | `true` for 404 on a missing tile, `false` for 204 |
738
-
739
- Leave `pieceCacheBytes` unset unless you have a reason. A fixed budget is a trap
740
- with 16 MiB pieces, since 64 MiB holds only four.
727
+ | setting | default | |
728
+ | ----------------------------- | -------- | ------------------------------------------------------------ |
729
+ | `tiles.maxOpenArchives` | `128` | complete archives kept open; each is a file descriptor |
730
+ | `tiles.swarmCacheBytes` | `1 GiB` | **memory** the piece caches of swarm reads share |
731
+ | `tiles.maxOpenSwarmArchives` | unset | a hard count of swarm readers as well, if you want one |
732
+ | `tiles.maxOpenRemoteArchives` | `64` | readers for a URL or a bucket; cheap to hold, dear to reopen |
733
+ | `tiles.directoryCacheEntries` | `2000` | header and directory cache, shared across archives |
734
+ | `tiles.pieceCacheBytes` | unset | sized from the torrent's piece length when unset |
735
+ | `tiles.hydrateIdleMs` | unset | idle time before background hydration resumes |
736
+ | `tiles.pieceTimeoutMs` | `120000` | how long to wait for one piece |
737
+ | `tiles.readyTimeoutMs` | `60000` | how long to wait for torrent metadata |
738
+ | `tiles.metadataTimeoutMs` | `120000` | how long a background metadata read may take |
739
+ | `tiles.headerTimeoutMs` | `12000` | how long a TileJSON request waits for a header |
740
+ | `tiles.sparse` | unset | `true` for 404 on a missing tile, `false` for 204 |
741
+
742
+ ### Why a swarm read caches pieces in memory at all
743
+
744
+ The swarm's unit is a **piece** — 4 MiB for archives this project creates,
745
+ commonly 16 MiB for a planet torrent. A tile is a few kilobytes inside one. So
746
+ fetching a piece to serve one tile and throwing it away means the next tile
747
+ pays for the whole piece again, and PMTiles guarantees there will be a next
748
+ tile in the same piece: its tiles are laid out in Hilbert order, so tiles that
749
+ are neighbours on the map are neighbours in the file. A tile lookup also reads
750
+ the header and a directory before the tile itself, and those live in pieces
751
+ that every single request would otherwise re-fetch.
752
+
753
+ Not on disk, because a disk copy of the pieces is what cache mode exists to
754
+ avoid: the point is a machine with 10 GiB free serving a 700 GiB planet. The
755
+ engine keeps whatever it keeps; this is the layer above it, and evicting from
756
+ it costs nothing.
757
+
758
+ **Trading cache size for open archives.** The two settings multiply. Halving
759
+ `pieceCacheBytes` doubles how many archives stay open within the same
760
+ `swarmCacheBytes`, at the cost of a lower hit rate on each. `16 MiB × 64
761
+ archives` and `64 MiB × 16 archives` are the same gigabyte spent differently:
762
+ the first suits a node reading a tile here and a tile there across many
763
+ archives, the second a node serving a region of a few.
764
+
765
+ Leave `pieceCacheBytes` unset unless you have such a reason. A fixed budget is
766
+ a trap with 16 MiB pieces, since 64 MiB holds only four.
767
+
768
+ ### Three limits, because the handles cost different things
769
+
770
+ They were one limit, set at sixteen for the most expensive kind and applied to
771
+ every kind. That is wrong for the node this is increasingly used to build: a
772
+ stack assembled from a provider's file index names several hundred sources, and
773
+ a bake walks every one of them. At sixteen such a run spends most of its time
774
+ reopening archives it has just closed, and each reopen re-reads a header and a
775
+ directory.
776
+
777
+ - **A complete archive** is a file descriptor and its share of the directory
778
+ cache. The unit allows 65535 descriptors, so hundreds of these are nothing.
779
+ - **A cache-mode archive** carries a piece cache, and it is **memory** — a map
780
+ of whole pieces in the node's own heap, not a directory on disk. Bounded by
781
+ `swarmCacheBytes` rather than by a count, because a count is the wrong unit:
782
+ a reader allows itself `max(64 MiB, 8 × pieceLength)`, so sixteen of them is
783
+ 1 GiB against the 4 MiB pieces this project creates and 2 GiB against the
784
+ 16 MiB pieces a planet torrent usually has. Stated as memory, the number
785
+ means what an operator cares about and the count follows from it.
786
+ - **A URL or bucket archive** holds an HTTP reader and the summary it has
787
+ already read. Cheap to keep, and expensive to reopen: a reopen costs a header
788
+ and a directory fetch over the network rather than off a disk.
789
+
790
+ A node serving nothing but cache-mode archives behaves exactly as it did.
741
791
 
742
792
  `headerTimeoutMs` is shorter than the others on purpose: somebody is waiting on
743
793
  it. A cache-mode archive with no web seed and no peers has nothing to read a
package/docs/internals.md CHANGED
@@ -25,6 +25,7 @@ Operator-facing documentation is elsewhere — see [publishing](publishing.md),
25
25
  - [Scheduled sources](#scheduled-sources)
26
26
  - [Running two engines at once](#running-two-engines-at-once)
27
27
  - [Why libtorrent runs as a sidecar](#why-libtorrent-runs-as-a-sidecar)
28
+ - [A sidecar that outlives the node](#a-sidecar-that-outlives-the-node)
28
29
  - [Making a torrent out of a map](#making-a-torrent-out-of-a-map)
29
30
  - [Resuming a partial download](#resuming-a-partial-download)
30
31
  - [Retiring and pruning a subscription](#retiring-and-pruning-a-subscription)
@@ -579,6 +580,40 @@ under its final name until it finishes. Do not point a web server at a libtorren
579
580
  save path that is also serving web seeds. Completion is still recorded correctly
580
581
  — the watcher finds no marked file and notes that the archive is whole.
581
582
 
583
+ ### A sidecar that outlives the node
584
+
585
+ The one failure mode of running libtorrent in another process, and the reason a
586
+ service can need starting two or three times before it takes.
587
+
588
+ The sidecar exits when its pipe closes, which covers an orderly stop and most
589
+ crashes. It does not cover a sidecar in the middle of **hashing**: libtorrent
590
+ does not hand control back until a check finishes, so the process notices
591
+ neither the closed pipe nor a signal until it does — minutes on a large
592
+ archive, hours on a planet. Nor does it cover a node killed outright, which
593
+ takes no part in stopping anything.
594
+
595
+ What is left behind still holds the listen port and the resume directory. The
596
+ next start fails against it, and since a failed start leaves its own sidecar,
597
+ they accumulate. In the journal it reads as:
598
+
599
+ ```
600
+ pmtiles-swarm.service: Unit process 102573 (python) remains running after unit stopped.
601
+ pmtiles-swarm.service: Failed to kill control group ...: Invalid argument
602
+ ```
603
+
604
+ Two things prevent it. Stopping **insists**: the shutdown request goes first,
605
+ so resume data is saved either way, and a sidecar that has not gone within a
606
+ few seconds is killed rather than left. And a start **reaps** what a previous
607
+ run left, before it spawns anything — the sidecar's pid is recorded beside its
608
+ resume data, and a recorded pid is killed only when `/proc` says that process
609
+ is really a sidecar. Pids are reused, and killing whatever inherited one would
610
+ be worse than the problem.
611
+
612
+ Neither is a substitute for `KillMode=mixed` in the unit. That is what stops
613
+ systemd signalling the sidecar at the same moment as the node, which kills it
614
+ before it can write anything down — see
615
+ [running-as-a-service.md](running-as-a-service.md).
616
+
582
617
  ## Making a torrent out of a map
583
618
 
584
619
  Two choices here are specific to distributing maps rather than generic files.
@@ -189,6 +189,50 @@ The unit below is what it generates, with your paths in it. Worth reading
189
189
  either way — a generated file you do not understand is a hand-written one you
190
190
  have not written yet.
191
191
 
192
+ ## Regenerating the unit on a node that already runs
193
+
194
+ The unit is derived from two things that move: the configuration, and how many
195
+ archives the library holds. A watched folder added six months ago belongs in
196
+ `ReadWritePaths`, and a library that has grown needs longer to write its resume
197
+ data than the unit allows. So running this again is ordinary:
198
+
199
+ ```
200
+ sudo pmtiles-swarm init --systemd --config /etc/pmtiles-swarm/swarm.config.json
201
+ ```
202
+
203
+ **It reads the configuration and writes only the unit**, next to that
204
+ configuration — never into `/etc/systemd/system`, and never over
205
+ `swarm.config.json`. Installing it is your step, and worth a look first:
206
+
207
+ ```
208
+ diff /etc/systemd/system/pmtiles-swarm.service \
209
+ /etc/pmtiles-swarm/pmtiles-swarm.service
210
+ sudo cp /etc/pmtiles-swarm/pmtiles-swarm.service /etc/systemd/system/
211
+ sudo systemctl daemon-reload
212
+ sudo systemctl restart pmtiles-swarm
213
+ ```
214
+
215
+ **`--force` is a different command.** It replaces `swarm.config.json` with a
216
+ fresh one — tokens, stacks, feeds and all — and it is for starting over, not
217
+ for regenerating a unit. Nothing about writing the unit needs it.
218
+
219
+ **Run it as root, not as the service account.** It writes into `/etc`, which
220
+ the service account cannot do and should not be able to do. What decides the
221
+ `User=` line is `--user`, which defaults to `pmtiles-swarm`; pass it if the
222
+ account is called something else:
223
+
224
+ ```
225
+ sudo pmtiles-swarm init --systemd \
226
+ --config /etc/pmtiles-swarm/swarm.config.json \
227
+ --user maps
228
+ ```
229
+
230
+ `ReadWritePaths` is derived from the configuration, so **a path added to the
231
+ installed unit by hand is not in the generated one**. That is what the diff is
232
+ for. The durable fix is to move such a path into the configuration — as a save
233
+ location, a watched folder, or a cache directory — after which it is derived
234
+ like the rest.
235
+
192
236
  ## The unit
193
237
 
194
238
  Annotated, because every directive here has cost somebody a diagnosis.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pmtiles-swarm",
3
- "version": "0.94.0",
3
+ "version": "0.96.0",
4
4
  "description": "BitTorrent distribution for PMTiles map archives: create torrents, watch folders, publish and subscribe to RSS feeds, and seed through qBittorrent or an embedded client",
5
5
  "type": "module",
6
6
  "main": "src/index.js",
package/src/config.js CHANGED
@@ -364,12 +364,62 @@ const DEFAULTS = {
364
364
  */
365
365
  tiles: {
366
366
  /**
367
- * Open archives kept alive at once. Each holds a file descriptor or a
368
- * torrent reader plus its piece cache, so this bounds both.
367
+ * Open archives kept alive at once.
368
+ *
369
+ * Sized for a library of hundreds, because that is what a node that
370
+ * merges layers holds: a stack built from a provider's file index names
371
+ * four hundred sources and a bake walks every one of them. At sixteen,
372
+ * such a run spent most of its time reopening archives it had just
373
+ * closed, and each reopen re-reads a header and a directory.
374
+ *
375
+ * A complete archive open is a file descriptor and its share of the
376
+ * directory cache, which is why this can be large -- the unit allows
377
+ * 65535 of them. An archive read through the swarm is not: it carries a
378
+ * piece cache, and those are bounded separately below.
379
+ */
380
+ maxOpenArchives: 128,
381
+ /**
382
+ * Memory the piece caches of swarm-read archives may hold between them.
383
+ *
384
+ * A count would be the wrong unit. What is expensive about a cache-mode
385
+ * reader is its piece cache, and how big that is depends on the torrent:
386
+ * `max(64 MiB, 8 x pieceLength)`, so 64 MiB for the 4 MiB pieces this
387
+ * project creates and 128 MiB for the 16 MiB pieces a planet torrent
388
+ * usually has. Sixteen readers is a gigabyte in one case and two in the
389
+ * other, which is not a limit anybody chose.
390
+ *
391
+ * Stated as memory, the number means what an operator actually cares
392
+ * about, and the count follows from it: lower `pieceCacheBytes` and more
393
+ * archives stay open, at the same ceiling. A budget rather than an
394
+ * allocation -- an archive that has served three tiles holds three
395
+ * pieces.
396
+ */
397
+ swarmCacheBytes: 1024 * 1024 * 1024,
398
+ /**
399
+ * A hard count of swarm-read archives, for a node that wants one.
400
+ *
401
+ * Unset by default: the budget above is the better limit, because it is
402
+ * the thing that runs out. Set this as well and the smaller wins.
403
+ */
404
+ maxOpenSwarmArchives: undefined,
405
+ /**
406
+ * Open archives read straight from a URL or a bucket.
407
+ *
408
+ * Cheap in memory -- an HTTP reader and the summary it has already read
409
+ * -- and expensive to reopen, since a reopen costs a header and a
410
+ * directory fetch over the network. So this sits well above the swarm
411
+ * limit and below the local one.
412
+ */
413
+ maxOpenRemoteArchives: 64,
414
+ /**
415
+ * Header and directory cache entries, shared across every archive.
416
+ *
417
+ * One archive contributes several: a header, a root directory, and a leaf
418
+ * per region being read. Two hundred was a handful of archives' worth, so
419
+ * a stack over hundreds of sources evicted its own directories between
420
+ * one tile and the next.
369
421
  */
370
- maxOpenArchives: 16,
371
- /** Header and directory cache entries, shared across every archive. */
372
- directoryCacheEntries: 200,
422
+ directoryCacheEntries: 2000,
373
423
  /**
374
424
  * Byte budget for the piece cache of one swarm-read archive. Unset sizes
375
425
  * it from the torrent's piece length, which is safer than a fixed budget.
@@ -729,8 +779,16 @@ export function writablePaths(config, configPath) {
729
779
  // directory holding it is written to as surely as any of the above.
730
780
  configPath ? path.dirname(path.resolve(configPath)) : undefined,
731
781
  ...(config?.watch ?? []).map((entry) => entry?.path),
782
+ ...(config?.watch ?? []).map((entry) => entry?.publishDir),
732
783
  ...(config?.locations ?? []).map((entry) => entry?.path),
733
784
  ...(config?.subscriptions ?? []).map((entry) => entry?.savePath),
785
+ // Where a scheduled export puts the archive it just built, and where it
786
+ // publishes a copy of it. Missed for as long as this list has existed:
787
+ // a nightly bake to a directory named nowhere else ran for an hour and
788
+ // was refused the write at the end, with the directory's permissions
789
+ // perfect -- which is the exact failure this list exists to prevent.
790
+ ...(config?.stackExports ?? []).map((entry) => entry?.savePath),
791
+ ...(config?.stackExports ?? []).map((entry) => entry?.publishDir),
734
792
  ].filter((value) => typeof value === 'string' && value);
735
793
 
736
794
  // Deduplicated but deliberately not collapsed into common ancestors. A
@@ -1,4 +1,5 @@
1
1
  import { spawn } from 'node:child_process';
2
+ import fs from 'node:fs/promises';
2
3
  import { createRequire } from 'node:module';
3
4
  import path from 'node:path';
4
5
 
@@ -29,6 +30,122 @@ function resolveSidecar() {
29
30
  }
30
31
  }
31
32
 
33
+ /**
34
+ * Ends a child process, politely and then not.
35
+ * @param {object} child - The process.
36
+ * @param {number} graceMs - How long the polite request gets.
37
+ * @returns {Promise<void>} - When it has gone.
38
+ */
39
+ export async function stopChild(child, graceMs) {
40
+ if (!child || child.exitCode !== null || child.signalCode !== null) return;
41
+ const gone = new Promise((resolve) => child.once('exit', resolve));
42
+ child.kill();
43
+ const settled = await Promise.race([
44
+ gone.then(() => true),
45
+ new Promise((resolve) => {
46
+ const timer = setTimeout(() => resolve(false), graceMs);
47
+ timer.unref?.();
48
+ }),
49
+ ]);
50
+ if (settled) return;
51
+ console.warn(
52
+ '[libtorrent] the sidecar did not stop when asked -- it is usually ' +
53
+ 'hashing, which cannot be interrupted -- so it is being killed. Its ' +
54
+ 'resume data was already saved.',
55
+ );
56
+ child.kill('SIGKILL');
57
+ await gone;
58
+ }
59
+
60
+ /**
61
+ * Where the pid of the running sidecar is written.
62
+ * @param {string} [resumeDir] - Where resume data is kept.
63
+ * @returns {string|null} - The path, or null when there is nowhere to put it.
64
+ */
65
+ function sidecarPidPath(resumeDir) {
66
+ return resumeDir ? path.join(resumeDir, 'sidecar.pid') : null;
67
+ }
68
+
69
+ /**
70
+ * Records which process is the sidecar, so a later start can recognise it.
71
+ * @param {string} [resumeDir] - Where resume data is kept.
72
+ * @param {number} pid - The sidecar.
73
+ * @returns {Promise<void>} - When written.
74
+ */
75
+ async function rememberSidecarPid(resumeDir, pid) {
76
+ const file = sidecarPidPath(resumeDir);
77
+ if (!file) return;
78
+ await fs.writeFile(file, String(pid)).catch(() => {});
79
+ }
80
+
81
+ /**
82
+ * Forgets the recorded pid.
83
+ * @param {string} [resumeDir] - Where resume data is kept.
84
+ * @returns {Promise<void>} - When removed.
85
+ */
86
+ async function forgetSidecarPid(resumeDir) {
87
+ const file = sidecarPidPath(resumeDir);
88
+ if (!file) return;
89
+ await fs.rm(file, { force: true }).catch(() => {});
90
+ }
91
+
92
+ /**
93
+ * Kills a sidecar left behind by a previous run.
94
+ *
95
+ * A node killed outright -- SIGKILL, an OOM, a power cut -- takes no part in
96
+ * stopping its sidecar, and a sidecar mid-hash does not notice its pipe close.
97
+ * It goes on holding the listen port and the resume directory, and the next
98
+ * start fails against it. Reaped here rather than lived with, because the
99
+ * alternative is the operator restarting the service until it takes.
100
+ *
101
+ * Identified by its command line, not by the pid alone: pids are reused, and
102
+ * killing whatever inherited one would be far worse than the problem.
103
+ * @param {string} [resumeDir] - Where resume data is kept.
104
+ * @returns {Promise<boolean>} - Whether one was killed.
105
+ */
106
+ export async function reapStaleSidecar(resumeDir) {
107
+ const file = sidecarPidPath(resumeDir);
108
+ if (!file) return false;
109
+ let pid;
110
+ try {
111
+ pid = Number(await fs.readFile(file, 'utf8'));
112
+ } catch {
113
+ return false;
114
+ }
115
+ if (!Number.isInteger(pid) || pid <= 1 || pid === process.pid) {
116
+ await forgetSidecarPid(resumeDir);
117
+ return false;
118
+ }
119
+
120
+ // Only where the process table can be read as files, which is where this
121
+ // runs as a service. Elsewhere a stale pid is left alone: guessing is worse.
122
+ let cmdline;
123
+ try {
124
+ cmdline = await fs.readFile(`/proc/${pid}/cmdline`, 'utf8');
125
+ } catch {
126
+ await forgetSidecarPid(resumeDir);
127
+ return false;
128
+ }
129
+ if (!cmdline.includes('libtorrent_sidecar')) {
130
+ await forgetSidecarPid(resumeDir);
131
+ return false;
132
+ }
133
+
134
+ console.warn(
135
+ `[libtorrent] a sidecar from a previous run is still running (pid ${pid}). ` +
136
+ 'It holds the listen port and the resume directory this one needs, so ' +
137
+ 'it is being killed. This is what a start that has to be repeated ' +
138
+ 'two or three times looks like.',
139
+ );
140
+ try {
141
+ process.kill(pid, 'SIGKILL');
142
+ } catch {
143
+ // Gone between reading and killing, which is the good outcome.
144
+ }
145
+ await forgetSidecarPid(resumeDir);
146
+ return true;
147
+ }
148
+
32
149
  /**
33
150
  * A SeedEngine backed by libtorrent, through a sidecar process.
34
151
  *
@@ -117,6 +234,9 @@ export class LibtorrentEngine {
117
234
 
118
235
  this.#ready = new Promise((resolve, reject) => {
119
236
  const script = this.#options.script ?? resolveSidecar();
237
+ // Before the spawn, not after: a sidecar from a previous run holds the
238
+ // port this one is about to ask for.
239
+ const reaped = reapStaleSidecar(this.#options.resumeDir);
120
240
 
121
241
  const child = spawn(this.#options.python, [script], {
122
242
  stdio: ['pipe', 'pipe', 'pipe'],
@@ -140,6 +260,7 @@ export class LibtorrentEngine {
140
260
  },
141
261
  });
142
262
  this.#child = child;
263
+ reaped.then(() => rememberSidecarPid(this.#options.resumeDir, child.pid));
143
264
 
144
265
  // Nothing of the last sidecar's is carried into this one. Being killed
145
266
  // does not wait for a newline, so a sidecar that died partway through a
@@ -802,9 +923,19 @@ export class LibtorrentEngine {
802
923
  await this.#call('shutdown', {}, options.timeoutMs ?? 15000).catch(
803
924
  () => {},
804
925
  );
805
- this.#child?.kill();
926
+
927
+ // Asked, then insisted. A sidecar in the middle of hashing is not reading
928
+ // its pipe and does not act on a signal until libtorrent hands control
929
+ // back, which on a large archive is minutes -- so a plain SIGTERM left it
930
+ // running after the node had gone. systemd then reports a unit process
931
+ // that remains after the unit stopped, the next start finds the old one
932
+ // still holding the listen port, and the library comes back holding
933
+ // nothing. That is the restart that has to be done two or three times.
934
+ const child = this.#child;
806
935
  this.#child = null;
807
936
  this.#ready = null;
937
+ await stopChild(child, options.killGraceMs ?? 5000);
938
+ await forgetSidecarPid(this.#options.resumeDir);
808
939
  }
809
940
 
810
941
  /**
package/src/index.js CHANGED
@@ -28,6 +28,7 @@ import {
28
28
  installSignalHandlers,
29
29
  runStoppers,
30
30
  } from './shutdown.js';
31
+ import { TIMEOUT_STOP_SECONDS } from './systemd.js';
31
32
  import { ScheduledSourceManager } from './sources.js';
32
33
  import { StackExportScheduler } from './stack-exports.js';
33
34
  import { StackFeedSubscriber } from './stack-feed.js';
@@ -338,6 +339,23 @@ PMTILES_SWARM_PUBLIC_URL
338
339
  );
339
340
 
340
341
  const catalogued = catalog.list().length;
342
+ // Said once at startup, because the way this fails is silent. Stopping
343
+ // writes resume data for every archive and the engine is given two seconds
344
+ // apiece; a unit that allows less than that is killed partway through, and
345
+ // what comes back re-hashes every archive whose resume data never landed.
346
+ // There is no error at that moment -- only a library at 0% and hours of
347
+ // disk. The generated unit derives its allowance from the library it was
348
+ // written for, so this is really "your library has grown since then".
349
+ const stopNeeds = Math.ceil(engineStopMs(catalogued) / 1000);
350
+ if (stopNeeds > TIMEOUT_STOP_SECONDS) {
351
+ console.warn(
352
+ `[shutdown] stopping this library needs about ${stopNeeds}s to write ` +
353
+ `resume data, which is more than the ${TIMEOUT_STOP_SECONDS}s ` +
354
+ 'a default systemd unit allows. Re-run `pmtiles-swarm init --systemd` ' +
355
+ 'to regenerate the unit, or raise TimeoutStopSec by hand -- a stop cut ' +
356
+ 'short re-hashes every archive it had not saved.',
357
+ );
358
+ }
341
359
  if (catalogued > 0) {
342
360
  // Reported, never fatal. Restore already tolerates a failure per archive;
343
361
  // what this catches is the whole call coming apart -- and the node it
@@ -3,7 +3,7 @@ import fs from 'node:fs/promises';
3
3
  import os from 'node:os';
4
4
  import path from 'node:path';
5
5
  import { hashPassword } from './auth.js';
6
- import { writablePaths } from './config.js';
6
+ import { loadConfig, writablePaths } from './config.js';
7
7
  import { directoryCommands, unitFor } from './systemd.js';
8
8
 
9
9
  /**
@@ -78,6 +78,25 @@ function firstConfig({ dataDir, savePath, passwordHash }) {
78
78
  };
79
79
  }
80
80
 
81
+ /**
82
+ * How many archives the catalog holds, if there is one to read.
83
+ *
84
+ * Best effort by design: this runs before the node does, on a machine that
85
+ * may have no data directory at all, and a missing or unreadable catalog is
86
+ * an install rather than an error.
87
+ * @param {object} config - The resolved configuration.
88
+ * @returns {Promise<number>} - The count, or zero.
89
+ */
90
+ async function countCatalog(config) {
91
+ try {
92
+ const file = path.join(config.dataDir, 'catalog.json');
93
+ const held = JSON.parse(await fs.readFile(file, 'utf8'));
94
+ return Array.isArray(held.entries) ? held.entries.length : 0;
95
+ } catch {
96
+ return 0;
97
+ }
98
+ }
99
+
81
100
  /**
82
101
  * Writes a first configuration file.
83
102
  *
@@ -128,10 +147,49 @@ export async function runInit(options = {}, write = console.log) {
128
147
  .access(configPath)
129
148
  .then(() => true)
130
149
  .catch(() => false);
150
+
151
+ // A node that already has a configuration and wants the unit is the
152
+ // ordinary reason to run this a second time: the unit is derived from the
153
+ // configuration and from how many archives the library holds, and both move.
154
+ // Refusing sent people to --force, which replaces the configuration --
155
+ // tokens, stacks, feeds and all -- to regenerate a file beside it.
156
+ if (exists && options.systemd && !options.force) {
157
+ const held = await loadConfig(configPath);
158
+ const unitPath = path.join(
159
+ path.dirname(configPath),
160
+ 'pmtiles-swarm.service',
161
+ );
162
+ await fs.writeFile(
163
+ unitPath,
164
+ unitFor({
165
+ config: held,
166
+ configPath,
167
+ user: options.user,
168
+ execStart: options.execStart,
169
+ archives: await countCatalog(held),
170
+ }),
171
+ );
172
+ write(`Wrote ${unitPath}`);
173
+ write('');
174
+ write(` ${configPath} was read, not written.`);
175
+ write('');
176
+ write(' Compare it with the one that is installed before replacing it:');
177
+ write('');
178
+ write(` diff /etc/systemd/system/pmtiles-swarm.service ${unitPath}`);
179
+ write('');
180
+ write(' ReadWritePaths is derived from the configuration, so a path');
181
+ write(' added to the installed unit by hand is not in this one. Move it');
182
+ write(' into the configuration — as a save location, a watched folder or');
183
+ write(' a cache path — and it will be derived from now on.');
184
+ return 0;
185
+ }
186
+
131
187
  if (exists && !options.force) {
132
188
  write(
133
189
  `${configPath} already exists. Pass --force to replace it — and take a ` +
134
- 'copy first, since the tokens in it are not recoverable.',
190
+ 'copy first, since the tokens in it are not recoverable. To write ' +
191
+ 'just the unit from the configuration that is there, pass --systemd ' +
192
+ 'without --force.',
135
193
  );
136
194
  return 1;
137
195
  }
@@ -164,6 +222,10 @@ export async function runInit(options = {}, write = console.log) {
164
222
  configPath,
165
223
  user: options.user,
166
224
  execStart: options.execStart,
225
+ // So the stop timeout fits the library this node already holds. On a
226
+ // fresh install there is none and the default stands; re-running this
227
+ // after the library has grown is what keeps the two together.
228
+ archives: await countCatalog(config),
167
229
  }),
168
230
  );
169
231
  write(`Wrote ${unitPath}`);
package/src/systemd.js CHANGED
@@ -1,5 +1,6 @@
1
1
  import os from 'node:os';
2
2
  import path from 'node:path';
3
+ import { engineStopMs } from './shutdown.js';
3
4
  import { writablePaths } from './config.js';
4
5
 
5
6
  /**
@@ -26,7 +27,29 @@ import { writablePaths } from './config.js';
26
27
  */
27
28
 
28
29
  /** What the resume save can want, before the library has grown into it. */
29
- const TIMEOUT_STOP_SECONDS = 300;
30
+ export const TIMEOUT_STOP_SECONDS = 300;
31
+
32
+ /**
33
+ * How long systemd should allow a stop, for a library of this size.
34
+ *
35
+ * The node gives its engine two seconds an archive to write resume data, and
36
+ * systemd's allowance has to outlast that or the save is cut off partway --
37
+ * which is exactly the library that returns at 0% and re-hashes for hours.
38
+ * A fixed five minutes covered a library of 145 and no more, silently, so
39
+ * this is derived the same way `ReadWritePaths` and the thread pool are.
40
+ *
41
+ * Rounded up to whole minutes, with a wide margin over the node's own budget:
42
+ * being generous here costs nothing, since systemd stops waiting the moment
43
+ * the process is gone.
44
+ * @param {object} config - The resolved configuration.
45
+ * @param {number} [archives] - How many the catalog holds.
46
+ * @returns {number} - Seconds for `TimeoutStopSec`.
47
+ */
48
+ export function stopTimeoutFor(config, archives = 0) {
49
+ const budget = engineStopMs(archives) / 1000;
50
+ const wanted = Math.ceil((budget * 1.5 + 60) / 60) * 60;
51
+ return Math.max(TIMEOUT_STOP_SECONDS, wanted);
52
+ }
30
53
 
31
54
  /** libuv's own, which is what a unit that says nothing gets. */
32
55
  const DEFAULT_THREADPOOL = 4;
@@ -68,11 +91,13 @@ export function unitFor({
68
91
  user = 'pmtiles-swarm',
69
92
  execStart,
70
93
  workingDirectory,
94
+ archives = 0,
71
95
  }) {
72
96
  const home = workingDirectory ?? `/var/lib/${user}`;
73
97
  const binary = execStart ?? `${home}/node_modules/.bin/pmtiles-swarm`;
74
98
  const paths = writablePaths(config, configPath);
75
99
  const threadPool = threadPoolFor(config);
100
+ const stopSeconds = stopTimeoutFor(config, archives);
76
101
 
77
102
  // Wrapped the way systemd's own examples are, because this list grows with
78
103
  // every watched folder and a single line of them is unreadable in a diff.
@@ -128,9 +153,14 @@ RestartSec=5
128
153
  # Stopping writes resume data for every archive, and the node allows its engine
129
154
  # two seconds per torrent to do it. This has to outlast that: killed mid-save,
130
155
  # every archive that had not been written re-hashes its whole store on the way
131
- # back up, which for a 700 GiB archive is hours. Raise it as the library grows
132
- # — tools/resume-doctor.py prints the arithmetic for yours.
133
- TimeoutStopSec=${TIMEOUT_STOP_SECONDS}
156
+ # back up, which for a 700 GiB archive is hours.
157
+ #
158
+ # Derived from the ${archives} archive(s) this node holds, with room to grow.
159
+ # A fixed five minutes covered a library of 145 and no more, and covered it
160
+ # silently — the symptom of outgrowing it is not an error but a library that
161
+ # comes back at 0%. Re-running init --systemd recomputes this;
162
+ # tools/resume-doctor.py prints the arithmetic for yours.
163
+ TimeoutStopSec=${stopSeconds}
134
164
 
135
165
  # The node stops the sidecar itself, and needs it alive to do so. The default,
136
166
  # control-group, signals both at once and the sidecar dies before it can write
package/src/tiles.js CHANGED
@@ -49,6 +49,42 @@ export class TileReadError extends Error {
49
49
  }
50
50
  }
51
51
 
52
+ /** What a reader's piece cache holds before it has seen the torrent. */
53
+ const STARTING_CACHE_BYTES = 64 * 1024 * 1024;
54
+
55
+ /**
56
+ * Which swarm-read archives to close for their piece caches to fit a budget.
57
+ *
58
+ * Each reader is asked what its cache costs rather than assumed, because only
59
+ * it knows: the budget it sets for itself is `max(64 MiB, 8 x pieceLength)`,
60
+ * and the piece length is the torrent's. So a node reading 4 MiB-piece
61
+ * archives keeps four times as many open as one reading 16 MiB-piece
62
+ * archives, for the same memory -- which is the arithmetic an operator would
63
+ * otherwise have to do by hand with every piece length in front of them.
64
+ *
65
+ * Least recently used first, and never the last one: the archive just opened
66
+ * is the one being read, and closing it to meet a budget it alone exceeds
67
+ * would serve no tile at all.
68
+ * @param {Array} held - `[key, handle]` pairs, least recently used first.
69
+ * @param {number} budget - Bytes the piece caches may hold between them.
70
+ * @returns {Array} - The pairs to close, in the order to close them.
71
+ */
72
+ export function overSwarmBudget(held, budget) {
73
+ if (!Number.isFinite(budget) || budget <= 0) return [];
74
+ const swarm = held.filter(([, one]) => one?.mode === 'swarm');
75
+ const cost = (one) => one.source?.stats?.cacheBudget ?? STARTING_CACHE_BYTES;
76
+
77
+ let total = swarm.reduce((sum, [, one]) => sum + cost(one), 0);
78
+ const closing = [];
79
+ for (const pair of swarm) {
80
+ if (total <= budget) break;
81
+ if (swarm.length - closing.length <= 1) break;
82
+ closing.push(pair);
83
+ total -= cost(pair[1]);
84
+ }
85
+ return closing;
86
+ }
87
+
52
88
  /**
53
89
  * Opens archives on demand and reads tiles out of them.
54
90
  */
@@ -76,7 +112,7 @@ export class TileStore {
76
112
  this.#engine = engine;
77
113
  this.#config = config;
78
114
  this.#directoryCache = new SharedPromiseCache(
79
- config.tiles?.directoryCacheEntries ?? 200,
115
+ config.tiles?.directoryCacheEntries ?? 2000,
80
116
  );
81
117
  }
82
118
 
@@ -240,7 +276,7 @@ export class TileStore {
240
276
  };
241
277
  this.#remote.set(url, handle);
242
278
 
243
- const limit = this.#config.tiles?.maxOpenRemoteArchives ?? 16;
279
+ const limit = this.#config.tiles?.maxOpenRemoteArchives ?? 64;
244
280
  while (this.#remote.size > limit) {
245
281
  const [oldest] = this.#remote.keys();
246
282
  this.#remote.delete(oldest);
@@ -431,15 +467,56 @@ export class TileStore {
431
467
  const handle = await this.#openArchive(entry);
432
468
  this.#open.set(entry.infoHash, handle);
433
469
 
434
- const limit = this.#config.tiles?.maxOpenArchives ?? 16;
435
- while (this.#open.size > limit) {
436
- const [oldest, victim] = this.#open.entries().next().value;
437
- this.#open.delete(oldest);
438
- await this.#release(victim);
470
+ // Two budgets, because the two kinds of handle cost quite different
471
+ // things. A complete archive is a file descriptor; one read through the
472
+ // swarm carries a piece cache sized from the torrent's piece length,
473
+ // which with 16 MiB pieces is tens of megabytes each. A single limit had
474
+ // to be set for the expensive kind, which left a library of hundreds of
475
+ // local archives reopening files it had just closed.
476
+ const tiles = this.#config.tiles ?? {};
477
+ await this.#evict(() => true, tiles.maxOpenArchives ?? 128);
478
+ if (tiles.maxOpenSwarmArchives !== undefined) {
479
+ await this.#evict(
480
+ (one) => one.mode === 'swarm',
481
+ tiles.maxOpenSwarmArchives,
482
+ );
439
483
  }
484
+ await this.#evictSwarmMemory(tiles.swarmCacheBytes ?? 1024 * 1024 * 1024);
440
485
  return handle;
441
486
  }
442
487
 
488
+ /**
489
+ * Closes the least recently used handles until a budget is met.
490
+ *
491
+ * Least recently used first, which a Map gives for nothing: reading one
492
+ * moves it to the end, so iteration order is oldest first.
493
+ * @param {Function} counts - Whether a handle counts against this budget.
494
+ * @param {number} limit - How many of them to keep.
495
+ * @returns {Promise<void>} - When the closing is done.
496
+ */
497
+ async #evict(counts, limit) {
498
+ if (!Number.isFinite(limit) || limit < 1) return;
499
+ const held = [...this.#open].filter(([, one]) => counts(one));
500
+ while (held.length > limit) {
501
+ const [key, victim] = held.shift();
502
+ this.#open.delete(key);
503
+ await this.#release(victim);
504
+ }
505
+ }
506
+
507
+ /**
508
+ * Closes swarm-read archives until their piece caches fit a byte budget.
509
+ * @param {number} budget - Bytes the piece caches may hold between them.
510
+ * @returns {Promise<void>} - When the closing is done.
511
+ */
512
+ async #evictSwarmMemory(budget) {
513
+ for (const [key] of overSwarmBudget([...this.#open], budget)) {
514
+ const victim = this.#open.get(key);
515
+ this.#open.delete(key);
516
+ await this.#release(victim);
517
+ }
518
+ }
519
+
443
520
  /**
444
521
  * Opens an archive, choosing between the local file and the swarm.
445
522
  * @param {object} entry - Catalog entry.
@@ -5279,17 +5279,47 @@ Every piece is hashed against the ` +
5279
5279
  key: 'tiles.maxOpenArchives',
5280
5280
  label: 'Archives kept open',
5281
5281
  type: 'number',
5282
- placeholder: '16',
5282
+ placeholder: '128',
5283
5283
  help:
5284
- 'Each holds a file descriptor, and in cache mode a piece ' +
5285
- 'cache as well. Past the number of archives this node serves ' +
5286
- 'it stops doing anything.',
5284
+ 'A complete archive open is a file descriptor and its share ' +
5285
+ 'of the directory cache, so this can be large — the unit ' +
5286
+ 'allows 65535 of them. Sized for a node that merges layers, ' +
5287
+ 'where one stack names hundreds of sources and a bake walks ' +
5288
+ 'every one. Past the number of archives this node serves it ' +
5289
+ 'stops doing anything.',
5290
+ },
5291
+ {
5292
+ key: 'tiles.swarmCacheBytes',
5293
+ label: 'Piece cache memory, all swarm reads',
5294
+ type: 'number',
5295
+ placeholder: '1073741824',
5296
+ help:
5297
+ 'RAM, not disk. An archive read through the swarm keeps whole ' +
5298
+ 'pieces in memory, because the swarm hands out pieces and a ' +
5299
+ 'tile is a few kilobytes of one — without it, every tile pays ' +
5300
+ 'for a 4 or 16 MiB fetch. Each reader allows itself ' +
5301
+ '<code>max(64 MiB, 8 × piece length)</code>, so this is the ' +
5302
+ 'number that decides how many stay open: lower ' +
5303
+ '<b>Piece cache bytes</b> below and more of them fit. A ' +
5304
+ 'budget, not an allocation — a reader that has served three ' +
5305
+ 'tiles holds three pieces. Complete archives are unaffected: ' +
5306
+ 'they hold a file descriptor and no cache.',
5307
+ },
5308
+ {
5309
+ key: 'tiles.maxOpenRemoteArchives',
5310
+ label: 'URL and bucket archives kept open',
5311
+ type: 'number',
5312
+ placeholder: '64',
5313
+ help:
5314
+ 'Cheap to hold — an HTTP reader and the summary it has ' +
5315
+ 'already read — and expensive to reopen, since a reopen costs ' +
5316
+ 'a header and a directory fetch over the network.',
5287
5317
  },
5288
5318
  {
5289
5319
  key: 'tiles.directoryCacheEntries',
5290
5320
  label: 'Directory cache entries',
5291
5321
  type: 'number',
5292
- placeholder: '200',
5322
+ placeholder: '2000',
5293
5323
  restart: true,
5294
5324
  help:
5295
5325
  'PMTiles finds a tile through a root directory and then a ' +