pmtiles-swarm 0.56.1 → 0.59.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +144 -0
- package/docs/configuration.md +25 -3
- package/docs/running-as-a-service.md +198 -4
- package/package.json +2 -2
- package/src/config.js +73 -0
- package/src/engines/composite.js +7 -5
- package/src/engines/libtorrent.js +19 -4
- package/src/index.js +53 -3
- package/src/init-command.js +231 -0
- package/src/library.js +83 -6
- package/src/shutdown.js +53 -5
- package/src/sources.js +3 -0
- package/src/subscriptions.js +4 -0
- package/src/systemd.js +158 -0
- package/src/web/index.html +108 -31
- package/swarm.config.json.sample +10 -54
package/CHANGELOG.md
CHANGED
|
@@ -7,6 +7,150 @@
|
|
|
7
7
|
### 🐞 Bug fixes
|
|
8
8
|
- _...Add new stuff here..._
|
|
9
9
|
|
|
10
|
+
## 0.59.0
|
|
11
|
+
### ✨ Features and improvements
|
|
12
|
+
- **Requires pmtiles-torrent 0.10.2**, which is what actually ends the re-checking: a
|
|
13
|
+
`seedOnly` add now discards resume data that would cancel the claim, and the periodic save
|
|
14
|
+
leaves a hashing torrent alone. Everything below only stops the node making more of it.
|
|
15
|
+
- **`pmtiles-swarm init` writes a first configuration, and optionally the unit to run it.**
|
|
16
|
+
Every path it writes is absolute, which is the one mistake it exists to make impossible: a
|
|
17
|
+
relative path resolves against the config file, the documented layout puts that file in
|
|
18
|
+
`/etc`, and so `./data` there means a catalog, a resume directory and potentially a 700 GiB
|
|
19
|
+
archive on the configuration partition. State under `/etc` is refused rather than warned
|
|
20
|
+
about — at the moment a config is written there is nothing to migrate, and the same mistake
|
|
21
|
+
found later costs a stopped service and a careful move.
|
|
22
|
+
|
|
23
|
+
`--systemd` also writes `pmtiles-swarm.service` beside it, with `ReadWritePaths` **derived
|
|
24
|
+
from the configuration it just wrote**. That derivation is the point. A unit and a config
|
|
25
|
+
that disagree is every systemd failure this project has diagnosed, and none of them look
|
|
26
|
+
like what they are: a `savePath` missing from that line is refused inside the unit's
|
|
27
|
+
namespace before any permission bit is read, so the directory's owner and mode are both
|
|
28
|
+
perfect and the write still fails. Two files generated from one source cannot drift.
|
|
29
|
+
|
|
30
|
+
It installs nothing. The unit is written next to the configuration and the `cp` into
|
|
31
|
+
`/etc/systemd/system` is one of the commands it prints, alongside an `install -d` for every
|
|
32
|
+
directory involved with the right owner and mode.
|
|
33
|
+
|
|
34
|
+
`--password` is hashed before it is written, and left out entirely when none is given.
|
|
35
|
+
There is deliberately no placeholder: `auth.password` accepts plaintext, so `REPLACE-ME` in
|
|
36
|
+
that field is a working password until somebody notices — a credential that looks set and
|
|
37
|
+
is not.
|
|
38
|
+
|
|
39
|
+
- **Two scripts for diagnosing a library that re-checks on every start**, in `tools/`.
|
|
40
|
+
`resume-doctor.py` reads a node's real configuration, catalog, stored `.torrent` files and
|
|
41
|
+
resume directory and says, for each archive, what libtorrent will do on the next start and
|
|
42
|
+
why — applying libtorrent's own rules and citing the file and line each came from. It also
|
|
43
|
+
reads the unit through `systemctl show`, does the arithmetic on every deadline that can cut
|
|
44
|
+
a resume save short, and hashes pieces rather than trusting the catalog. `resume-experiment.py`
|
|
45
|
+
proves the five behaviours involved on whatever libtorrent is actually installed, because
|
|
46
|
+
1.2, 2.0 and 2.1 differ enough that a claim verified on one is not a claim about the other.
|
|
47
|
+
|
|
48
|
+
### 🐞 Bug fixes
|
|
49
|
+
- **A stop no longer abandons the sidecar mid-write, which is where resume data was going.**
|
|
50
|
+
Three separate bounds decided how long the engine step had, and the smallest won: eight
|
|
51
|
+
seconds for the step, fifteen for the whole shutdown, fifteen for the shutdown RPC. The
|
|
52
|
+
sidecar allows each torrent two seconds of its resume-save budget, so past four archives the
|
|
53
|
+
node gave up first and every torrent it had not persisted re-hashed its whole store on the
|
|
54
|
+
way back up. That is the state a library gets stuck in: checking, on every start, for hours.
|
|
55
|
+
|
|
56
|
+
The engine step is now worked out from the catalog — two seconds a torrent, over a floor —
|
|
57
|
+
and the shutdown watchdog is derived from the steps it is meant to contain rather than being
|
|
58
|
+
a second deadline kept in agreement with them by hand. It was not in agreement.
|
|
59
|
+
|
|
60
|
+
The same fixed 60s applied to the periodic save, so past thirty archives the call gave up
|
|
61
|
+
before the sidecar finished, and the written/asked counts never came back — which is why the
|
|
62
|
+
shortfall this reports could not be seen from outside. **`TimeoutStopSec` in the documented
|
|
63
|
+
unit rises from 45 to 300 seconds**, and existing installs need it raised by hand.
|
|
64
|
+
|
|
65
|
+
- **A `.torrent` that moved with `dataDir` is found again instead of silently becoming a
|
|
66
|
+
magnet.** `torrentPath` is recorded absolute, so moving state out of `/etc` — which
|
|
67
|
+
`docs/running-as-a-service.md` tells you to do — left every catalog entry naming a directory
|
|
68
|
+
that no longer existed. Nothing repointed them and nothing complained, because an unreadable
|
|
69
|
+
`.torrent` was treated as "use the magnet instead".
|
|
70
|
+
|
|
71
|
+
That fallback is the damage rather than a graceful degradation. A magnet carries no metadata
|
|
72
|
+
and neither does resume data, so the archive waits on BEP 9 for a file list that only a peer
|
|
73
|
+
can supply — and for an archive this node originated there is nobody to ask. Seen in the
|
|
74
|
+
field: twenty archives at 0% in `downloading_metadata`, indefinitely, after one documented
|
|
75
|
+
migration. The current `dataDir` is now tried second, the recorded path is corrected in
|
|
76
|
+
place so the warning is printed once rather than for ever, and falling back to a magnet at
|
|
77
|
+
all now says so.
|
|
78
|
+
|
|
79
|
+
## 0.58.1
|
|
80
|
+
### 🐞 Bug fixes
|
|
81
|
+
- **The sample configuration put every piece of state under the config file.** `"dataDir": "./data"`,
|
|
82
|
+
`"savePath": "./data/torrents-data"` and `"resumeDir": "./data/resume"` all resolve against the
|
|
83
|
+
config file — and the service guide puts that file in `/etc`. Anyone following both documents ended
|
|
84
|
+
up with a catalog, a resume directory and their archives on the partition meant for configuration,
|
|
85
|
+
having done nothing wrong. The sample now uses absolute paths, a test enforces it, and the node
|
|
86
|
+
warns at startup if state resolves under `/etc` anyway.
|
|
87
|
+
|
|
88
|
+
It is also shorter. A first config should get a node running, not demonstrate the whole surface —
|
|
89
|
+
`docs/configuration.md` is where the rest lives.
|
|
90
|
+
|
|
91
|
+
- **Moving state out of `/etc` had a trap in the instructions.** `mv OLD/data NEW/data` nests when the
|
|
92
|
+
destination exists, which it does after the documented setup — so the real directory ends up one
|
|
93
|
+
level too deep, the node writes a fresh empty catalog beside it, and an intact library reads as
|
|
94
|
+
lost. `docs/running-as-a-service.md` now gives a form that cannot nest, says to repoint
|
|
95
|
+
`libtorrent.resumeDir` as well as `dataDir`, and has you count catalog entries before and after.
|
|
96
|
+
|
|
97
|
+
## 0.58.0
|
|
98
|
+
### ✨ Features and improvements
|
|
99
|
+
- **A stable name for every kind of import, not just watched folders.** `latestLink` and
|
|
100
|
+
`latestLinkType` are now offered on watched web locations, RSS feeds and remote nodes as well — one
|
|
101
|
+
path a consumer can hold while the build behind it changes.
|
|
102
|
+
|
|
103
|
+
A scheduled source had the feature all along and no way to ask for it: the console never showed the
|
|
104
|
+
column. It was also ignoring `latestLinkType`, so a source asking for a hard link quietly got a
|
|
105
|
+
symlink. A subscription could not ask at all.
|
|
106
|
+
|
|
107
|
+
For a subscription the name is pointed at the archive **when the download finishes**. Until then
|
|
108
|
+
there is a marker file, or a sparse one still filling in, and a name resolving to either is worse
|
|
109
|
+
than no name: whatever opens it reads zeroes rather than failing.
|
|
110
|
+
|
|
111
|
+
- **The node says something when its state has landed in `/etc`.** Nobody chooses that — the
|
|
112
|
+
documented service layout puts the config file there, every path resolves relative to that file,
|
|
113
|
+
and the sample reads `"dataDir": "./data"`. So the catalog and the resume directory end up on the
|
|
114
|
+
partition meant for configuration. Warned rather than corrected: it is a real path that works, and
|
|
115
|
+
moving a running node's data would be worse than saying so.
|
|
116
|
+
|
|
117
|
+
### 🐞 Bug fixes
|
|
118
|
+
|
|
119
|
+
## 0.57.0
|
|
120
|
+
### ✨ Features and improvements
|
|
121
|
+
- **The three publishing switches are one **Local file** column on the import tables.** Three
|
|
122
|
+
dropdowns per row cost more width than the choice is worth, on tables that already scroll
|
|
123
|
+
sideways — and as an import default these are almost always decided together:
|
|
124
|
+
|
|
125
|
+
| option | serves the file | web seed | on the public page |
|
|
126
|
+
| --- | --- | --- | --- |
|
|
127
|
+
| `node` | — | — | — |
|
|
128
|
+
| `off` | no | — | — |
|
|
129
|
+
| `http` | yes | no | no |
|
|
130
|
+
| `http + catalog` | yes | no | yes |
|
|
131
|
+
| `http + web seed` | yes | yes | no |
|
|
132
|
+
| `http + web seed + catalog` | yes | yes | yes |
|
|
133
|
+
|
|
134
|
+
`http + catalog` is there rather than left out to keep a tidier ladder: a dropdown that could not
|
|
135
|
+
say it would round that state up to the nearest option it had, turning a web seed on — and
|
|
136
|
+
publishing the node to the swarm — because somebody re-saved an unrelated row. Every combination
|
|
137
|
+
survives being written and read back, which is what the round-trip test checks.
|
|
138
|
+
|
|
139
|
+
The archive details panel still offers the three separately. A row here sets a policy for what
|
|
140
|
+
arrives; the panel is where one archive gets picked over.
|
|
141
|
+
|
|
142
|
+
## 0.56.2
|
|
143
|
+
### 🐞 Bug fixes
|
|
144
|
+
- **Adding a save location in the console did nothing.** The row was read correctly and then thrown
|
|
145
|
+
away: the settings pane below renders every config key it does not explicitly skip as a raw-JSON
|
|
146
|
+
textarea, the skip list named `watch`, `sources` and `subscriptions` but not `locations`, and in
|
|
147
|
+
`saveSettings` the textarea loop runs after the row editors — so a copy of the list as it was when
|
|
148
|
+
the pane was drawn overwrote the one with the new row in it.
|
|
149
|
+
|
|
150
|
+
The skip is now derived from the registered row editors rather than listed by hand, so this cannot
|
|
151
|
+
happen again to the next editor somebody adds. The list was the bug; keeping a list and adding one
|
|
152
|
+
more name to it would have been the same bug waiting.
|
|
153
|
+
|
|
10
154
|
## 0.56.1
|
|
11
155
|
### ✨ Features and improvements
|
|
12
156
|
- **Comments trimmed back to house style, and the reasoning moved into the docs where it belongs.**
|
package/docs/configuration.md
CHANGED
|
@@ -488,9 +488,19 @@ moment a download finishes.
|
|
|
488
488
|
|
|
489
489
|
The node's setting is the default. A [watched folder](#watched-folders), a
|
|
490
490
|
[scheduled source](#scheduled-sources), an [RSS feed](#subscriptions) and a
|
|
491
|
-
remote node may each carry their own — the **
|
|
492
|
-
|
|
493
|
-
|
|
491
|
+
remote node may each carry their own — the **Local file** column on those
|
|
492
|
+
tables, which offers the three as one choice:
|
|
493
|
+
|
|
494
|
+
| option | `serveArchive` | `selfWebSeed` | `publicDownload` |
|
|
495
|
+
| --------------------------- | -------------- | ------------- | ---------------- |
|
|
496
|
+
| `node` | unset | unset | unset |
|
|
497
|
+
| `off` | `false` | — | — |
|
|
498
|
+
| `http` | `true` | — | — |
|
|
499
|
+
| `http + catalog` | `true` | — | `true` |
|
|
500
|
+
| `http + web seed` | `true` | `true` | — |
|
|
501
|
+
| `http + web seed + catalog` | `true` | `true` | `true` |
|
|
502
|
+
|
|
503
|
+
`node` means "no opinion" rather than "off". And any individual archive can be switched in the console, under **HTTP
|
|
494
504
|
sources** in its details.
|
|
495
505
|
|
|
496
506
|
An archive that says nothing goes on following the node, so changing the node's
|
|
@@ -586,6 +596,17 @@ it, and what turns a cold tile read from tens of seconds into well under one.
|
|
|
586
596
|
`webSeedBase` on its own assumes the watched folder is already the web root, since
|
|
587
597
|
nothing is moved.
|
|
588
598
|
|
|
599
|
+
**Every table that imports an archive offers this**, not just watched folders:
|
|
600
|
+
a [scheduled source](#scheduled-sources), an [RSS feed](#subscriptions) and a
|
|
601
|
+
remote node each take `latestLink` and `latestLinkType` too. The point is the
|
|
602
|
+
same in all four — a consumer holds one path and never learns the name of any
|
|
603
|
+
particular build.
|
|
604
|
+
|
|
605
|
+
For a subscription the name is pointed at the archive **when the download
|
|
606
|
+
finishes**, not when it is joined: until then there is a marker file or a sparse
|
|
607
|
+
one still filling in, and a name resolving to either is worse than no name at
|
|
608
|
+
all, because whatever opens it reads zeroes rather than failing.
|
|
609
|
+
|
|
589
610
|
`latestLinkType: 'hard'` is for a name something reads the archive _through_. A
|
|
590
611
|
hard link still resolves after the build it names is retired, where a symlink is
|
|
591
612
|
left pointing at nothing. The other kind stays the fallback in both directions,
|
|
@@ -625,6 +646,7 @@ entry gives either a `url` template or an `index` directory:
|
|
|
625
646
|
| `everyHours` | an interval instead |
|
|
626
647
|
| `md5` | overrides the node's [`md5`](#md5) for this source alone |
|
|
627
648
|
| `serveArchive`, `selfWebSeed`, `publicDownload` | override what this node offers of the archives this source fetches |
|
|
649
|
+
| `latestLink`, `latestLinkType` | a stable second name for the newest build, as on a watched folder |
|
|
628
650
|
|
|
629
651
|
Prefer a template where the naming is predictable: it asks a direct question,
|
|
630
652
|
gets a direct answer, and needs the upstream to publish no listing at all.
|
|
@@ -151,8 +151,50 @@ sudo apt-get install -y python3-libtorrent
|
|
|
151
151
|
sudo -u pmtiles-swarm python3 -c "import libtorrent; print(libtorrent.__version__)"
|
|
152
152
|
```
|
|
153
153
|
|
|
154
|
+
## The configuration, and a unit for it
|
|
155
|
+
|
|
156
|
+
`init` writes both, and the unit's `ReadWritePaths` is derived from the
|
|
157
|
+
configuration it just wrote. That derivation is the reason to use it: a unit and
|
|
158
|
+
a configuration that disagree is the failure this whole document is arranged
|
|
159
|
+
around, and two files generated from one source cannot.
|
|
160
|
+
|
|
161
|
+
```sh
|
|
162
|
+
sudo -u pmtiles-swarm -H /var/lib/pmtiles-swarm/node_modules/.bin/pmtiles-swarm init --config /etc/pmtiles-swarm/swarm.config.json --data-dir /var/lib/pmtiles-swarm/data --save-path /mnt/store/torrent-data --systemd --password 'the console password'
|
|
163
|
+
```
|
|
164
|
+
|
|
165
|
+
That writes the configuration, generates `pmtiles-swarm.service` beside it, and
|
|
166
|
+
prints a command for every directory involved — each with the right owner and
|
|
167
|
+
mode. Nothing privileged happens and nothing is installed: the unit is written
|
|
168
|
+
next to the configuration, and copying it into `/etc/systemd/system` is one of
|
|
169
|
+
the lines it hands you.
|
|
170
|
+
|
|
171
|
+
Two refusals worth knowing before you type it. State under `/etc` is rejected
|
|
172
|
+
rather than warned about, because at the moment a configuration is written there
|
|
173
|
+
is nothing to migrate — the same mistake found six months later costs a stopped
|
|
174
|
+
service and a careful move. And a configuration that already exists is left
|
|
175
|
+
alone unless `--force` says otherwise, since the credentials in it are not
|
|
176
|
+
recoverable.
|
|
177
|
+
|
|
178
|
+
`--password` is hashed before it is written. Leave it off and there is no
|
|
179
|
+
console password at all: the generated API key is the way in until you set one.
|
|
180
|
+
There is deliberately no placeholder, because `auth.password` accepts plaintext
|
|
181
|
+
and a placeholder in that field is a working password until somebody notices.
|
|
182
|
+
|
|
183
|
+
**Re-run it after adding a watched folder**, or a subscription with a save path
|
|
184
|
+
of its own. Both are directories the node writes to, and neither reaches
|
|
185
|
+
`ReadWritePaths` by itself. `tools/resume-doctor.py` checks a running node's
|
|
186
|
+
unit against its live configuration and says which directories are missing.
|
|
187
|
+
|
|
188
|
+
The unit below is what it generates, with your paths in it. Worth reading
|
|
189
|
+
either way — a generated file you do not understand is a hand-written one you
|
|
190
|
+
have not written yet.
|
|
191
|
+
|
|
154
192
|
## The unit
|
|
155
193
|
|
|
194
|
+
Annotated, because every directive here has cost somebody a diagnosis.
|
|
195
|
+
`init --systemd` writes this file with your paths already in it; what
|
|
196
|
+
follows is what those lines are for.
|
|
197
|
+
|
|
156
198
|
```ini
|
|
157
199
|
[Unit]
|
|
158
200
|
Description=pmtiles-swarm
|
|
@@ -177,8 +219,12 @@ Restart=always
|
|
|
177
219
|
RestartSec=5
|
|
178
220
|
|
|
179
221
|
# Stopping announces "stopped" to every tracker, releases the data directory
|
|
180
|
-
# lock
|
|
181
|
-
|
|
222
|
+
# lock, cancels downloads in flight and writes resume data for every archive.
|
|
223
|
+
# That last part is what scales: the node allows its engine two seconds per
|
|
224
|
+
# torrent, so a library of fifty wants well over a minute and systemd has to
|
|
225
|
+
# outlast it. Too low here and the process is killed mid-save, which costs a
|
|
226
|
+
# re-hash of everything unwritten on the way back up.
|
|
227
|
+
TimeoutStopSec=300
|
|
182
228
|
|
|
183
229
|
# The node stops the sidecar itself, and needs it alive to do so. The default
|
|
184
230
|
# signals both at once, which kills the sidecar before it can write its resume
|
|
@@ -233,6 +279,17 @@ itself, in order, and waits for the resume data to be written. Nothing is left
|
|
|
233
279
|
running: anything still alive when `TimeoutStopSec` expires is killed, sidecar
|
|
234
280
|
included.
|
|
235
281
|
|
|
282
|
+
**`TimeoutStopSec`, and why it is not a round number you pick once.** Stopping
|
|
283
|
+
writes resume data for every archive, and the node allows its engine two seconds
|
|
284
|
+
per torrent to do it — so the figure that matters grows with the library. Three
|
|
285
|
+
hundred seconds covers a hundred and forty archives; a larger node needs more.
|
|
286
|
+
Set it too low and systemd kills the process mid-save, and every archive that
|
|
287
|
+
had not been written re-hashes its whole store on the way back up, which for a
|
|
288
|
+
700 GiB archive is hours.
|
|
289
|
+
|
|
290
|
+
`tools/resume-doctor.py` prints the arithmetic for your library and says whether
|
|
291
|
+
the unit's value covers it.
|
|
292
|
+
|
|
236
293
|
**No `ExecStop=`.** systemd already sends `SIGTERM`, and the node handles it
|
|
237
294
|
from the moment it starts. `ExecStop=/bin/kill -15 $MAINPID` is redundant, and
|
|
238
295
|
becomes wrong under `KillMode=process` — that one leaves the rest of the unit
|
|
@@ -256,8 +313,10 @@ directory:
|
|
|
256
313
|
-> /etc/pmtiles-swarm/data
|
|
257
314
|
```
|
|
258
315
|
|
|
259
|
-
That is rarely what you want for a service
|
|
260
|
-
|
|
316
|
+
That is rarely what you want for a service, and the node says so on startup —
|
|
317
|
+
`/etc` is for configuration, and a catalog, a resume directory and possibly an
|
|
318
|
+
archive do not belong on that partition. The shipped sample uses absolute paths
|
|
319
|
+
for exactly this reason. Use them for anything holding data:
|
|
261
320
|
|
|
262
321
|
```json
|
|
263
322
|
{
|
|
@@ -267,6 +326,33 @@ holding data:
|
|
|
267
326
|
}
|
|
268
327
|
```
|
|
269
328
|
|
|
329
|
+
#### Moving state that already landed in the wrong place
|
|
330
|
+
|
|
331
|
+
Repoint **every** path that lived under the directory being moved, not just
|
|
332
|
+
`dataDir`: `libtorrent.resumeDir` is the one that is easy to miss, because it is
|
|
333
|
+
absolute and sits inside it.
|
|
334
|
+
|
|
335
|
+
And move with a destination that cannot nest. `mv /etc/pmtiles-swarm/data
|
|
336
|
+
/var/lib/pmtiles-swarm/data` puts the source _inside_ the destination when the
|
|
337
|
+
destination already exists — which it does, since the setup above creates it —
|
|
338
|
+
leaving `/var/lib/pmtiles-swarm/data/data`. The node then finds no catalog where
|
|
339
|
+
it was told to look, writes a fresh empty one, and every archive appears to have
|
|
340
|
+
been lost.
|
|
341
|
+
|
|
342
|
+
```bash
|
|
343
|
+
sudo systemctl stop pmtiles-swarm
|
|
344
|
+
sudo python3 -c "import json;print(len(json.load(open('OLD/catalog.json'))['entries']))"
|
|
345
|
+
|
|
346
|
+
sudo mv /etc/pmtiles-swarm/data /var/lib/pmtiles-swarm/ # note: no second 'data'
|
|
347
|
+
# repoint dataDir *and* libtorrent.resumeDir, then
|
|
348
|
+
|
|
349
|
+
sudo systemctl start pmtiles-swarm
|
|
350
|
+
sudo python3 -c "import json;print(len(json.load(open('NEW/catalog.json'))['entries']))"
|
|
351
|
+
```
|
|
352
|
+
|
|
353
|
+
The entry count before and after is the check that matters: an empty new catalog
|
|
354
|
+
is indistinguishable from a lost library until you count.
|
|
355
|
+
|
|
270
356
|
### 2. Every one of them in `ReadWritePaths`
|
|
271
357
|
|
|
272
358
|
`ProtectSystem=strict` presents the whole filesystem as read-only inside the
|
|
@@ -750,3 +836,111 @@ once rather than diagnosing later:
|
|
|
750
836
|
`journalctl -u pmtiles-swarm | grep -i onComplete` — and a hook redirecting its
|
|
751
837
|
own output to a file will have nothing for the journal to show, which is not
|
|
752
838
|
the same as not having run.
|
|
839
|
+
|
|
840
|
+
## When archives re-check on every start
|
|
841
|
+
|
|
842
|
+
The symptom is always the same shape — a restart, and then a library sitting at
|
|
843
|
+
0% or grinding through a re-hash of archives that are whole on the disk. The
|
|
844
|
+
causes are not the same shape at all, so guessing between them is expensive:
|
|
845
|
+
re-hashing a 700 GiB archive is half an hour during which it serves nobody.
|
|
846
|
+
|
|
847
|
+
Two scripts in `tools/` answer it without guessing. Neither starts anything,
|
|
848
|
+
neither writes anything, and both are safe against a running node.
|
|
849
|
+
|
|
850
|
+
```sh
|
|
851
|
+
sudo -u pmtiles-swarm python3 tools/resume-doctor.py \
|
|
852
|
+
-c /etc/pmtiles-swarm/swarm.config.json
|
|
853
|
+
```
|
|
854
|
+
|
|
855
|
+
**`resume-doctor.py`** reads your configuration, catalog, stored `.torrent`
|
|
856
|
+
files and resume directory, and says for each archive what libtorrent will do on
|
|
857
|
+
the next start and why — seed at once, stat its files, or re-hash the store. It
|
|
858
|
+
applies libtorrent's own rules rather than a guess at them, and cites the file
|
|
859
|
+
and line each one came from. It also reads the unit through `systemctl show`,
|
|
860
|
+
does the arithmetic on every deadline that can cut a resume save short, and
|
|
861
|
+
scans the journal for the failures that leave a trace.
|
|
862
|
+
|
|
863
|
+
```sh
|
|
864
|
+
sudo -u pmtiles-swarm python3 tools/resume-experiment.py
|
|
865
|
+
```
|
|
866
|
+
|
|
867
|
+
**`resume-experiment.py`** proves the behaviour on the libtorrent you actually
|
|
868
|
+
have, rather than the one the documentation was written against. It builds a
|
|
869
|
+
real torrent over a temporary file, saves real resume data, adds it again, and
|
|
870
|
+
reports what happened. 2.0 and 2.1 differ enough to be worth the thirty seconds.
|
|
871
|
+
|
|
872
|
+
Three things they exist to tell apart, because the first two look identical from
|
|
873
|
+
outside and want opposite fixes:
|
|
874
|
+
|
|
875
|
+
- **No resume data is not the problem.** An archive recorded complete is added
|
|
876
|
+
with `seed_mode`, and that claim stands on its own: libtorrent stats the files
|
|
877
|
+
and seeds, without hashing a byte. A missing resume file costs nothing.
|
|
878
|
+
- **Stale resume data is the problem, and it is worse than none.** A resume file
|
|
879
|
+
holding a _partial_ bitfield cancels the `seed_mode` claim — libtorrent drops
|
|
880
|
+
it if a single piece is unset. The archive then comes back as a downloader
|
|
881
|
+
with every byte already on disk, and fetches again what it has. Deleting that
|
|
882
|
+
one file is the repair; with none, it seeds at once.
|
|
883
|
+
- **A short or missing file is a genuine re-check.** `mismatching_file_size` is
|
|
884
|
+
the only thing here that really does mean re-hashing, and it is about the data
|
|
885
|
+
on disk rather than about resume data at all. libtorrent 2.x records no
|
|
886
|
+
mtimes, so nothing in this depends on a timestamp.
|
|
887
|
+
|
|
888
|
+
Which archives carry stale resume data is not random, and the pattern is worth
|
|
889
|
+
knowing before it looks like a second bug. An archive from a watched folder was
|
|
890
|
+
read end to end before it was ever registered: it is complete from its first
|
|
891
|
+
moment and no partial bitfield is ever written for it. An archive joined from a
|
|
892
|
+
feed, a magnet or a peer is genuinely partial for hours, with partial resume
|
|
893
|
+
data written for it every `resumeSaveIntervalSeconds` throughout. Only the
|
|
894
|
+
second kind has anything to leave behind, so trouble concentrating there is this
|
|
895
|
+
failure's ordinary shape rather than something the import path introduced.
|
|
896
|
+
|
|
897
|
+
Two more things the doctor looks at, because they are the ones a snapshot alone
|
|
898
|
+
gets wrong.
|
|
899
|
+
|
|
900
|
+
**Whether the catalog still points at its own `.torrent` files.** Each entry
|
|
901
|
+
records `torrentPath` as an absolute path, and restore reads that rather than
|
|
902
|
+
recomputing it — so moving `dataDir` out of `/etc`, which this guide tells you
|
|
903
|
+
to do, leaves every recorded path behind. Nothing warns, because a missing
|
|
904
|
+
`.torrent` is not treated as an error: restore quietly falls back to the magnet.
|
|
905
|
+
That fallback is the damage. A magnet carries no metadata, resume data does not
|
|
906
|
+
carry it either, and so the archive waits on BEP 9 for a file list that only a
|
|
907
|
+
peer can supply — from a swarm where this node is the origin. It sits at 0% in
|
|
908
|
+
`downloading_metadata` indefinitely, which no amount of re-checking will fix.
|
|
909
|
+
|
|
910
|
+
**Whether the periodic save gets through the whole library.** Every torrent is
|
|
911
|
+
asked to save in the same cycle, so a healthy `resumeDir` has all its files
|
|
912
|
+
within one interval of each other. Ages spread wider than that mean some cycles
|
|
913
|
+
are finishing early, and the ages alone cannot say whether the same archives
|
|
914
|
+
lose every time. `--watch` measures it rather than inferring it:
|
|
915
|
+
|
|
916
|
+
```sh
|
|
917
|
+
sudo -u pmtiles-swarm python3 tools/resume-doctor.py \
|
|
918
|
+
-c /etc/pmtiles-swarm/swarm.config.json --watch 660
|
|
919
|
+
```
|
|
920
|
+
|
|
921
|
+
It reports each save cycle as it happens, how many of them covered every
|
|
922
|
+
archive, and which archives were never written at all.
|
|
923
|
+
|
|
924
|
+
**A poisoned archive may heal itself, so run this twice.** The commonest way a
|
|
925
|
+
whole archive acquires a partial bitfield is a save taken while it was
|
|
926
|
+
re-checking: `write_resume_data` truncates `have_pieces` to the pieces checked
|
|
927
|
+
so far, and the five-minute timer lands inside that window every time a large
|
|
928
|
+
archive checks. That shortened bitfield is what cancels `seed_mode` on the next
|
|
929
|
+
start — so the archive comes back partial, checks again, and writes another one.
|
|
930
|
+
The loop needs no restart to keep itself going.
|
|
931
|
+
|
|
932
|
+
The doctor tells the two apart by length: a bitfield shorter than the torrent
|
|
933
|
+
was written mid-check, and one at full length is what a finished check
|
|
934
|
+
concluded. A check that completes rewrites its own resume file correctly, so an
|
|
935
|
+
archive poisoned in one run can read `seeds at once` in the next with nothing
|
|
936
|
+
done to it. **Only the archives that stay poisoned across two runs, several
|
|
937
|
+
minutes apart, are worth deleting a file for.**
|
|
938
|
+
|
|
939
|
+
Two things it deliberately does not treat as faults. An archive whose data sits
|
|
940
|
+
outside the `savePathLayout` shape is normal — one adopted from a file this node
|
|
941
|
+
already holds keeps that file, and a node may keep archives beside whatever
|
|
942
|
+
produced them. What matters is whether the data is where the entry says, which
|
|
943
|
+
is checked directly. And every directory an archive actually lives in has to
|
|
944
|
+
appear in `ReadWritePaths`, not just the `savePath` in the configuration; the
|
|
945
|
+
doctor checks each one in use, because a deliberately-placed archive is exactly
|
|
946
|
+
the case that gets left out of the unit.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pmtiles-swarm",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.59.0",
|
|
4
4
|
"description": "BitTorrent distribution for PMTiles map archives: create torrents, watch folders, publish and subscribe to RSS feeds, and seed through qBittorrent or an embedded client",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "src/index.js",
|
|
@@ -46,7 +46,7 @@
|
|
|
46
46
|
"maplibre-gl": "^6.2.0",
|
|
47
47
|
"parse-torrent": "^11.0.24",
|
|
48
48
|
"pmtiles": "^4.4.1",
|
|
49
|
-
"pmtiles-torrent": "^0.
|
|
49
|
+
"pmtiles-torrent": "^0.10.2",
|
|
50
50
|
"webtorrent": "^3.0.21"
|
|
51
51
|
},
|
|
52
52
|
"engines": {
|
package/src/config.js
CHANGED
|
@@ -507,6 +507,76 @@ function clone(value) {
|
|
|
507
507
|
return value;
|
|
508
508
|
}
|
|
509
509
|
|
|
510
|
+
/**
|
|
511
|
+
* Every directory this configuration means to write to.
|
|
512
|
+
*
|
|
513
|
+
* Which is exactly what `ReadWritePaths` has to name. Under
|
|
514
|
+
* `ProtectSystem=strict` a directory missing from that line is refused inside
|
|
515
|
+
* the unit's namespace, before any permission bit is consulted — so it fails
|
|
516
|
+
* with ownership and mode both perfect, which is a hard thing to go looking
|
|
517
|
+
* for. Deriving the list from the config rather than writing it out by hand is
|
|
518
|
+
* the only way the two cannot drift.
|
|
519
|
+
*
|
|
520
|
+
* See docs/running-as-a-service.md — "Where it writes".
|
|
521
|
+
* @param {object} config - A config whose paths have been resolved.
|
|
522
|
+
* @param {string} [configPath] - The config file, whose directory is written to.
|
|
523
|
+
* @returns {string[]} - Absolute directories, deduplicated and shortest-first.
|
|
524
|
+
*/
|
|
525
|
+
export function writablePaths(config, configPath) {
|
|
526
|
+
const found = [
|
|
527
|
+
config?.dataDir,
|
|
528
|
+
config?.savePath,
|
|
529
|
+
config?.cacheSavePath,
|
|
530
|
+
config?.libtorrent?.resumeDir,
|
|
531
|
+
config?.torrentDropDir,
|
|
532
|
+
// The console rewrites the configuration when a token is minted, so the
|
|
533
|
+
// directory holding it is written to as surely as any of the above.
|
|
534
|
+
configPath ? path.dirname(path.resolve(configPath)) : undefined,
|
|
535
|
+
...(config?.watch ?? []).map((entry) => entry?.path),
|
|
536
|
+
...(config?.locations ?? []).map((entry) => entry?.path),
|
|
537
|
+
...(config?.subscriptions ?? []).map((entry) => entry?.savePath),
|
|
538
|
+
].filter((value) => typeof value === 'string' && value);
|
|
539
|
+
|
|
540
|
+
// Deduplicated but deliberately not collapsed into common ancestors. A
|
|
541
|
+
// shorter list would grant the same access — `ReadWritePaths` covers a
|
|
542
|
+
// directory and everything under it — but the two callers want different
|
|
543
|
+
// things from it, and only one of them would be served. Creating the
|
|
544
|
+
// directories needs each of them named; and collapsing quietly widens the
|
|
545
|
+
// grant to whatever ancestor happens to be shared, which for a config beside
|
|
546
|
+
// its data is the whole tree above both.
|
|
547
|
+
return [...new Set(found.map((value) => path.resolve(value)))].sort();
|
|
548
|
+
}
|
|
549
|
+
|
|
550
|
+
/**
|
|
551
|
+
* Complains about state that has landed in /etc.
|
|
552
|
+
*
|
|
553
|
+
* Nobody chooses this. The documented service layout puts the config file in
|
|
554
|
+
* /etc, every path resolves relative to that file, and the sample read
|
|
555
|
+
* "./data" — so the catalog and the resume directory ended up on the partition
|
|
556
|
+
* meant for configuration. Warned rather than corrected: it is a real path
|
|
557
|
+
* that works, and moving a running node's data would be worse than saying so.
|
|
558
|
+
*
|
|
559
|
+
* See docs/running-as-a-service.md — "The paths in the configuration".
|
|
560
|
+
* @param {object} config - A config whose paths have been resolved.
|
|
561
|
+
* @returns {string[]} - What to say, empty when there is nothing to say.
|
|
562
|
+
*/
|
|
563
|
+
export function stateUnderEtc(config) {
|
|
564
|
+
const named = [
|
|
565
|
+
['dataDir', config?.dataDir],
|
|
566
|
+
['savePath', config?.savePath],
|
|
567
|
+
];
|
|
568
|
+
return named
|
|
569
|
+
.filter(
|
|
570
|
+
([, value]) => typeof value === 'string' && value.startsWith('/etc/'),
|
|
571
|
+
)
|
|
572
|
+
.map(
|
|
573
|
+
([name, value]) =>
|
|
574
|
+
`[config] ${name} is ${value}. /etc is for configuration; state ` +
|
|
575
|
+
'belongs under /var/lib. Give it an absolute path, or let systemd ' +
|
|
576
|
+
'StateDirectory= make one.',
|
|
577
|
+
);
|
|
578
|
+
}
|
|
579
|
+
|
|
510
580
|
/**
|
|
511
581
|
* Loads configuration from a JSON file, applying defaults and environment
|
|
512
582
|
* overrides. A missing file is fine — the defaults run.
|
|
@@ -573,6 +643,9 @@ export async function loadConfig(configPath) {
|
|
|
573
643
|
const { savePath, conflict } = resolveSavePath(config);
|
|
574
644
|
config.savePath = path.resolve(base, savePath);
|
|
575
645
|
config.savePathConflict = conflict;
|
|
646
|
+
|
|
647
|
+
for (const message of stateUnderEtc(config)) console.warn(message);
|
|
648
|
+
|
|
576
649
|
// Kept in step, because older code and older configs both reach for it.
|
|
577
650
|
config.webtorrent = { ...config.webtorrent, savePath: config.savePath };
|
|
578
651
|
if (config.libtorrent) config.libtorrent.savePath = config.savePath;
|
package/src/engines/composite.js
CHANGED
|
@@ -433,7 +433,7 @@ export class CompositeEngine {
|
|
|
433
433
|
* @returns {Promise<{written: number, asked: number}>} - Totals across every
|
|
434
434
|
* engine that keeps resume data.
|
|
435
435
|
*/
|
|
436
|
-
async saveResume(infoHash) {
|
|
436
|
+
async saveResume(infoHash, options = {}) {
|
|
437
437
|
// Summed rather than dropped, so a caller can tell the difference between
|
|
438
438
|
// "every torrent wrote" and "half of them will be re-hashed on the next
|
|
439
439
|
// start". An engine that fails outright contributes nothing to either
|
|
@@ -444,7 +444,7 @@ export class CompositeEngine {
|
|
|
444
444
|
// WebTorrent keeps none, and says so by not offering the method.
|
|
445
445
|
if (!engine.saveResume) continue;
|
|
446
446
|
try {
|
|
447
|
-
const result = (await engine.saveResume(infoHash)) ?? {};
|
|
447
|
+
const result = (await engine.saveResume(infoHash, options)) ?? {};
|
|
448
448
|
written += Number(result.written) || 0;
|
|
449
449
|
asked += Number(result.asked) || 0;
|
|
450
450
|
} catch (error) {
|
|
@@ -653,19 +653,21 @@ export class CompositeEngine {
|
|
|
653
653
|
* Shuts every engine down.
|
|
654
654
|
* @returns {Promise<void>} - Resolves once all are stopped.
|
|
655
655
|
*/
|
|
656
|
-
async destroy() {
|
|
656
|
+
async destroy(options = {}) {
|
|
657
657
|
this.#stopping = true;
|
|
658
658
|
if (this.#timer) clearInterval(this.#timer);
|
|
659
659
|
this.#timer = undefined;
|
|
660
660
|
|
|
661
661
|
for (const engine of this.#secondaries) {
|
|
662
662
|
await engine
|
|
663
|
-
.destroy()
|
|
663
|
+
.destroy(options)
|
|
664
664
|
.catch((error) =>
|
|
665
665
|
console.error(`[engine] ${engine.name}: ${error.message}`),
|
|
666
666
|
);
|
|
667
667
|
}
|
|
668
|
-
|
|
668
|
+
// Forwarded, because the primary is the one holding resume data and the
|
|
669
|
+
// budget was worked out from how much of it there is to write.
|
|
670
|
+
await this.#primary.destroy(options);
|
|
669
671
|
}
|
|
670
672
|
}
|
|
671
673
|
|
|
@@ -763,25 +763,40 @@ export class LibtorrentEngine {
|
|
|
763
763
|
* were told to write resume data, and how many actually did before the
|
|
764
764
|
* deadline.
|
|
765
765
|
*/
|
|
766
|
-
async saveResume(infoHash) {
|
|
766
|
+
async saveResume(infoHash, options = {}) {
|
|
767
767
|
// Returned rather than discarded. The sidecar reports both numbers, and
|
|
768
768
|
// the gap between them is the thing worth knowing: a torrent that did not
|
|
769
769
|
// write is one that gets re-hashed on the next start, which for a 700 GiB
|
|
770
770
|
// archive is the difference between seeding in seconds and seeding in half
|
|
771
771
|
// an hour. That answer was being thrown away here.
|
|
772
|
-
|
|
772
|
+
//
|
|
773
|
+
// The default #call timeout is 60s and the sidecar's own budget is two
|
|
774
|
+
// seconds per torrent, so past thirty archives the call gave up first —
|
|
775
|
+
// and then the counts never came back, which is why the shortfall this
|
|
776
|
+
// reports could not be seen from outside.
|
|
777
|
+
return this.#call('save_resume', { infoHash }, options.timeoutMs);
|
|
773
778
|
}
|
|
774
779
|
|
|
775
780
|
/**
|
|
776
781
|
* Saves resume data and stops the sidecar.
|
|
782
|
+
*
|
|
783
|
+
* The timeout is the caller's to set, because only the caller knows how many
|
|
784
|
+
* torrents are about to be written down and the sidecar spends two seconds
|
|
785
|
+
* per torrent. Fifteen seconds was the fixed value, which meant every
|
|
786
|
+
* library past seven archives had its resume save cut off — and each torrent
|
|
787
|
+
* that missed was re-hashed in full on the way back up.
|
|
788
|
+
* @param {object} [options] - Overrides.
|
|
789
|
+
* @param {number} [options.timeoutMs] - How long the sidecar gets to finish.
|
|
777
790
|
* @returns {Promise<void>} - Resolves once stopped.
|
|
778
791
|
*/
|
|
779
|
-
async destroy() {
|
|
792
|
+
async destroy(options = {}) {
|
|
780
793
|
// Set before anything else, so the exit this is about to cause is
|
|
781
794
|
// recognised as intended by the handler that sees it.
|
|
782
795
|
this.#stopping = true;
|
|
783
796
|
if (!this.#child) return;
|
|
784
|
-
await this.#call('shutdown', {}, 15000).catch(
|
|
797
|
+
await this.#call('shutdown', {}, options.timeoutMs ?? 15000).catch(
|
|
798
|
+
() => {},
|
|
799
|
+
);
|
|
785
800
|
this.#child?.kill();
|
|
786
801
|
this.#child = null;
|
|
787
802
|
this.#ready = null;
|