pmtiles-swarm 0.58.1 → 0.60.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +137 -0
- package/docs/haproxy.md +74 -3
- package/docs/running-as-a-service.md +167 -2
- package/package.json +2 -2
- package/src/api.js +198 -102
- package/src/config.js +42 -2
- package/src/engines/composite.js +7 -5
- package/src/engines/libtorrent.js +19 -4
- package/src/incomplete.js +9 -0
- package/src/index.js +53 -3
- package/src/init-command.js +231 -0
- package/src/library.js +121 -6
- package/src/shutdown.js +53 -5
- package/src/systemd.js +158 -0
- package/src/tilejson.js +27 -3
- package/src/web/index.html +1 -0
package/CHANGELOG.md
CHANGED
|
@@ -7,6 +7,143 @@
|
|
|
7
7
|
### 🐞 Bug fixes
|
|
8
8
|
- _...Add new stuff here..._
|
|
9
9
|
|
|
10
|
+
## 0.60.0
|
|
11
|
+
### ✨ Features and improvements
|
|
12
|
+
- **A tile URL that survives a rebuild.** Every archive is addressed by infohash, which is
|
|
13
|
+
what makes a tile cacheable for a year and what makes it useless in an application: the URL
|
|
14
|
+
changes with every build, so anything that wrote one down is pinned to a build that
|
|
15
|
+
eventually stops existing. A category is the only stable handle this system has, and now it
|
|
16
|
+
has a tile endpoint of its own — `/latest/<category>/{z}/{x}/{y}.<ext>` — which resolves to
|
|
17
|
+
whichever build is current on every request.
|
|
18
|
+
|
|
19
|
+
The category TileJSON advertises it as `tiles`, so a style written once keeps working across
|
|
20
|
+
a rebuild without being re-fetched for the URLs alone. The immutable template is still
|
|
21
|
+
published beside it as `latest.tiles`, because it is still the better URL for anything that
|
|
22
|
+
can re-read the document: it is content-addressed, so it caches for a year and never
|
|
23
|
+
revalidates. Offering only one of the two would be choosing for the consumer, and the right
|
|
24
|
+
answer differs by consumer.
|
|
25
|
+
|
|
26
|
+
The console offers the template as a copyable **XYZ** button next to the TileJSON one, which
|
|
27
|
+
is the form a Leaflet layer, an OpenLayers source or a GIS client actually wants.
|
|
28
|
+
|
|
29
|
+
Deliberately not a redirect to the immutable URL, though every other `/latest/` route is one.
|
|
30
|
+
A redirect costs a round trip and a map asks for hundreds of tiles: what is a negligible
|
|
31
|
+
indirection for a `.torrent` is the difference between a map that feels immediate and one
|
|
32
|
+
that does not. It is cached as the moving target it is — `max-age=300, must-revalidate`,
|
|
33
|
+
tagged with the build it resolved to, so a revalidation is a 304 while that build stands and
|
|
34
|
+
a miss the moment it moves.
|
|
35
|
+
|
|
36
|
+
- **A benchmark for "why does this feel slower than the other one", in `tools/tile-bench.mjs`.**
|
|
37
|
+
It reads two TileJSON documents, picks tiles inside the zoom range and bounds both can serve,
|
|
38
|
+
and requests the same set from each — one server at a time, because run together they compete
|
|
39
|
+
for the same link and each measures the other's load as its own latency.
|
|
40
|
+
|
|
41
|
+
It reports percentiles rather than an average, since what makes a map feel slow is the tail
|
|
42
|
+
and a mean built from nineteen fast requests does not move for the twentieth. Time to first
|
|
43
|
+
byte is separated from the total, which is the difference between a slow lookup and a slow
|
|
44
|
+
link — they want opposite fixes.
|
|
45
|
+
|
|
46
|
+
It also detects a pool of unequal nodes, which is a common cause and an invisible one: half
|
|
47
|
+
the tiles arrive quickly and half do not, and balanced evenly the mean looks tolerable
|
|
48
|
+
throughout. Two distinct groups are reported as two, with a histogram, and `--header` tallies
|
|
49
|
+
a response header naming which backend answered. `--a-origin`/`--b-origin` send the tile
|
|
50
|
+
requests somewhere other than the document was read from, which is the only way to measure
|
|
51
|
+
one node directly: a node with `publicUrl` set answers with that name however it was asked.
|
|
52
|
+
|
|
53
|
+
- **`docs/haproxy.md` covers a pool whose nodes are not the same speed.** Round robin assumes
|
|
54
|
+
the pool is interchangeable, and a tile server on an NVMe disk and one on a spinning disk are
|
|
55
|
+
not — an archive read is a seek into a large file, which is what a spinning disk is worst at.
|
|
56
|
+
`backup`, `weight` and least-connections are compared, along with what each does and does not
|
|
57
|
+
fix.
|
|
58
|
+
|
|
59
|
+
### 🐞 Bug fixes
|
|
60
|
+
- **An archive at 100% and seeding could serve no tiles until the node was restarted.** Which
|
|
61
|
+
source an archive is read through is decided once, when a reader opens it, and every other
|
|
62
|
+
thing that can change that answer already invalidates the reader: a pause, a resume, a mode
|
|
63
|
+
change, a move, a finished download. A finished *check* did not — and it is the easiest of
|
|
64
|
+
them to reach, because during `checking_files` libtorrent reports `progress` as the fraction
|
|
65
|
+
hashed so far, which is indistinguishable from a download sitting at the same figure.
|
|
66
|
+
|
|
67
|
+
So a tile read arriving mid-check opened against the swarm, correctly, and kept that handle
|
|
68
|
+
afterwards. The archive then read from a swarm whose only member is this node, while the
|
|
69
|
+
whole file sat on the disk beside it, and nothing evicted the handle short of a restart. The
|
|
70
|
+
completion sweep looked straight past it: an entry already recorded complete never reaches
|
|
71
|
+
the code that would have noticed.
|
|
72
|
+
|
|
73
|
+
The sweep now drops a reader that is going to the swarm for an archive the disk says is
|
|
74
|
+
whole. The disk is checked before the handle is dropped rather than after, or a reader that
|
|
75
|
+
would only re-open against the swarm anyway would be invalidated on every sweep, for ever.
|
|
76
|
+
Cache mode is left alone: it reads from the swarm because that is what it is for.
|
|
77
|
+
|
|
78
|
+
## 0.59.0
|
|
79
|
+
### ✨ Features and improvements
|
|
80
|
+
- **Requires pmtiles-torrent 0.10.2**, which is what actually ends the re-checking: a
|
|
81
|
+
`seedOnly` add now discards resume data that would cancel the claim, and the periodic save
|
|
82
|
+
leaves a hashing torrent alone. Everything below only stops the node making more of it.
|
|
83
|
+
- **`pmtiles-swarm init` writes a first configuration, and optionally the unit to run it.**
|
|
84
|
+
Every path it writes is absolute, which is the one mistake it exists to make impossible: a
|
|
85
|
+
relative path resolves against the config file, the documented layout puts that file in
|
|
86
|
+
`/etc`, and so `./data` there means a catalog, a resume directory and potentially a 700 GiB
|
|
87
|
+
archive on the configuration partition. State under `/etc` is refused rather than warned
|
|
88
|
+
about — at the moment a config is written there is nothing to migrate, and the same mistake
|
|
89
|
+
found later costs a stopped service and a careful move.
|
|
90
|
+
|
|
91
|
+
`--systemd` also writes `pmtiles-swarm.service` beside it, with `ReadWritePaths` **derived
|
|
92
|
+
from the configuration it just wrote**. That derivation is the point. A unit and a config
|
|
93
|
+
that disagree is every systemd failure this project has diagnosed, and none of them look
|
|
94
|
+
like what they are: a `savePath` missing from that line is refused inside the unit's
|
|
95
|
+
namespace before any permission bit is read, so the directory's owner and mode are both
|
|
96
|
+
perfect and the write still fails. Two files generated from one source cannot drift.
|
|
97
|
+
|
|
98
|
+
It installs nothing. The unit is written next to the configuration and the `cp` into
|
|
99
|
+
`/etc/systemd/system` is one of the commands it prints, alongside an `install -d` for every
|
|
100
|
+
directory involved with the right owner and mode.
|
|
101
|
+
|
|
102
|
+
`--password` is hashed before it is written, and left out entirely when none is given.
|
|
103
|
+
There is deliberately no placeholder: `auth.password` accepts plaintext, so `REPLACE-ME` in
|
|
104
|
+
that field is a working password until somebody notices — a credential that looks set and
|
|
105
|
+
is not.
|
|
106
|
+
|
|
107
|
+
- **Two scripts for diagnosing a library that re-checks on every start**, in `tools/`.
|
|
108
|
+
`resume-doctor.py` reads a node's real configuration, catalog, stored `.torrent` files and
|
|
109
|
+
resume directory and says, for each archive, what libtorrent will do on the next start and
|
|
110
|
+
why — applying libtorrent's own rules and citing the file and line each came from. It also
|
|
111
|
+
reads the unit through `systemctl show`, does the arithmetic on every deadline that can cut
|
|
112
|
+
a resume save short, and hashes pieces rather than trusting the catalog. `resume-experiment.py`
|
|
113
|
+
proves the five behaviours involved on whatever libtorrent is actually installed, because
|
|
114
|
+
1.2, 2.0 and 2.1 differ enough that a claim verified on one is not a claim about the other.
|
|
115
|
+
|
|
116
|
+
### 🐞 Bug fixes
|
|
117
|
+
- **A stop no longer abandons the sidecar mid-write, which is where resume data was going.**
|
|
118
|
+
Three separate bounds decided how long the engine step had, and the smallest won: eight
|
|
119
|
+
seconds for the step, fifteen for the whole shutdown, fifteen for the shutdown RPC. The
|
|
120
|
+
sidecar allows each torrent two seconds of its resume-save budget, so past four archives the
|
|
121
|
+
node gave up first and every torrent it had not persisted re-hashed its whole store on the
|
|
122
|
+
way back up. That is the state a library gets stuck in: checking, on every start, for hours.
|
|
123
|
+
|
|
124
|
+
The engine step is now worked out from the catalog — two seconds a torrent, over a floor —
|
|
125
|
+
and the shutdown watchdog is derived from the steps it is meant to contain rather than being
|
|
126
|
+
a second deadline kept in agreement with them by hand. It was not in agreement.
|
|
127
|
+
|
|
128
|
+
The same fixed 60s applied to the periodic save, so past thirty archives the call gave up
|
|
129
|
+
before the sidecar finished, and the written/asked counts never came back — which is why the
|
|
130
|
+
shortfall this reports could not be seen from outside. **`TimeoutStopSec` in the documented
|
|
131
|
+
unit rises from 45 to 300 seconds**, and existing installs need it raised by hand.
|
|
132
|
+
|
|
133
|
+
- **A `.torrent` that moved with `dataDir` is found again instead of silently becoming a
|
|
134
|
+
magnet.** `torrentPath` is recorded absolute, so moving state out of `/etc` — which
|
|
135
|
+
`docs/running-as-a-service.md` tells you to do — left every catalog entry naming a directory
|
|
136
|
+
that no longer existed. Nothing repointed them and nothing complained, because an unreadable
|
|
137
|
+
`.torrent` was treated as "use the magnet instead".
|
|
138
|
+
|
|
139
|
+
That fallback is the damage rather than a graceful degradation. A magnet carries no metadata
|
|
140
|
+
and neither does resume data, so the archive waits on BEP 9 for a file list that only a peer
|
|
141
|
+
can supply — and for an archive this node originated there is nobody to ask. Seen in the
|
|
142
|
+
field: twenty archives at 0% in `downloading_metadata`, indefinitely, after one documented
|
|
143
|
+
migration. The current `dataDir` is now tried second, the recorded path is corrected in
|
|
144
|
+
place so the warning is printed once rather than for ever, and falling back to a magnet at
|
|
145
|
+
all now says so.
|
|
146
|
+
|
|
10
147
|
## 0.58.1
|
|
11
148
|
### 🐞 Bug fixes
|
|
12
149
|
- **The sample configuration put every piece of state under the config file.** `"dataDir": "./data"`,
|
package/docs/haproxy.md
CHANGED
|
@@ -114,6 +114,70 @@ gives archive affinity, at the price of concentrating one archive on one node.
|
|
|
114
114
|
For nodes holding complete copies there is no cold read to avoid, so this is a
|
|
115
115
|
cost with no benefit.
|
|
116
116
|
|
|
117
|
+
### When the nodes are not the same speed
|
|
118
|
+
|
|
119
|
+
Round robin assumes the pool is interchangeable, and a tile server on an NVMe
|
|
120
|
+
disk and one on a spinning disk are not. Every archive read is a seek into a
|
|
121
|
+
large file — the PMTiles directory, then the leaf, then the tile — and that is
|
|
122
|
+
precisely the access pattern a spinning disk is worst at. The two can easily
|
|
123
|
+
differ by an order of magnitude.
|
|
124
|
+
|
|
125
|
+
What that does to a pool is worse than the average suggests. Balanced evenly,
|
|
126
|
+
half the tiles arrive quickly and half do not, and a map is judged by the slow
|
|
127
|
+
half: panning stutters wherever the slow node answered. The mean looks
|
|
128
|
+
tolerable throughout, which is why this is easy to miss and easy to blame on
|
|
129
|
+
the software.
|
|
130
|
+
|
|
131
|
+
Three ways out, and they answer different questions.
|
|
132
|
+
|
|
133
|
+
**`backup`** keeps the slow node out of service entirely until the fast one
|
|
134
|
+
fails:
|
|
135
|
+
|
|
136
|
+
```
|
|
137
|
+
backend pmtiles-swarm
|
|
138
|
+
option httpchk GET /health HTTP/1.1
|
|
139
|
+
http-check expect status 200
|
|
140
|
+
server node1 172.16.1.49:8090 check inter 5s
|
|
141
|
+
server node2 172.16.1.41:8090 check inter 5s backup
|
|
142
|
+
```
|
|
143
|
+
|
|
144
|
+
Every request goes to `node1` while it is healthy, and the pool falls to
|
|
145
|
+
`node2` when it is not. Latency is then the fast node's latency, and the slow
|
|
146
|
+
one is redundancy rather than capacity. This is the right answer when the fast
|
|
147
|
+
node can carry the load alone, which for tile serving it usually can.
|
|
148
|
+
|
|
149
|
+
**`weight`** keeps both in service and sends proportionally less to the slow
|
|
150
|
+
one:
|
|
151
|
+
|
|
152
|
+
```
|
|
153
|
+
server node1 172.16.1.49:8090 check inter 5s weight 200
|
|
154
|
+
server node2 172.16.1.41:8090 check inter 5s weight 20
|
|
155
|
+
```
|
|
156
|
+
|
|
157
|
+
Worth it only when you actually need the second node's throughput. It does not
|
|
158
|
+
remove the slow answers, it makes them rarer — a tenth of requests at ten times
|
|
159
|
+
the latency is still a tenth of requests a person notices.
|
|
160
|
+
|
|
161
|
+
**Least Connections** is the better of the two ways to keep both in service,
|
|
162
|
+
and it needs no numbers chosen by hand. A slow node holds each connection
|
|
163
|
+
longer, so it accumulates open ones and stops being picked until it catches up:
|
|
164
|
+
the pool tunes itself to what the disks are doing today rather than to a weight
|
|
165
|
+
somebody guessed last year. It shares `weight`'s limitation — the slow answers
|
|
166
|
+
become proportionally rarer, not absent — and it pairs with `backup` without
|
|
167
|
+
conflict, since a backup server is held out of the pool whatever the algorithm.
|
|
168
|
+
|
|
169
|
+
`tools/tile-bench.mjs` shows which situation you are in. A pool of unequal
|
|
170
|
+
nodes answers in two distinct groups rather than one spread, and the tool says
|
|
171
|
+
so explicitly; `--header` tallies a response header naming the backend, if
|
|
172
|
+
HAProxy is set to add one:
|
|
173
|
+
|
|
174
|
+
```
|
|
175
|
+
http-response set-header X-Served-By %s
|
|
176
|
+
```
|
|
177
|
+
|
|
178
|
+
Run it before and after the change. What should move is the tail: `total p90`
|
|
179
|
+
and `total p99` collapse towards `p50` once one node is answering everything.
|
|
180
|
+
|
|
117
181
|
### HTTP/2
|
|
118
182
|
|
|
119
183
|
Enable it on the frontend and leave _HTTP/2 without TLS_ unchecked: the client
|
|
@@ -219,9 +283,16 @@ says so, which is worse than having no check at all.
|
|
|
219
283
|
|
|
220
284
|
**Cache tiles by infohash aggressively.** `/archives/<infohash>/…` is immutable
|
|
221
285
|
by construction — an infohash names those bytes and no others — and is served
|
|
222
|
-
with `max-age=31536000, immutable`.
|
|
223
|
-
|
|
224
|
-
|
|
286
|
+
with `max-age=31536000, immutable`.
|
|
287
|
+
|
|
288
|
+
Everything under `/latest/<category>/` is the opposite, because it moves on
|
|
289
|
+
every build. The documents — `tiles.json`, the feeds — are served with
|
|
290
|
+
`max-age=60, must-revalidate`; the tiles with `max-age=300`. Both carry an ETag
|
|
291
|
+
naming the build they resolved to, so a revalidation is a 304 while the build
|
|
292
|
+
stands and a miss the moment it moves. Let Cloudflare revalidate rather than
|
|
293
|
+
overriding either with a page rule: the whole point of the category tile URL is
|
|
294
|
+
that it is safe to write into an application, and it is only safe while the
|
|
295
|
+
edge notices a rebuild.
|
|
225
296
|
|
|
226
297
|
**Do not strip or rewrite the ETag on either.** It is the infohash, which is
|
|
227
298
|
how a PMTiles reader notices that the archive moved underneath a read already
|
|
@@ -151,8 +151,50 @@ sudo apt-get install -y python3-libtorrent
|
|
|
151
151
|
sudo -u pmtiles-swarm python3 -c "import libtorrent; print(libtorrent.__version__)"
|
|
152
152
|
```
|
|
153
153
|
|
|
154
|
+
## The configuration, and a unit for it
|
|
155
|
+
|
|
156
|
+
`init` writes both, and the unit's `ReadWritePaths` is derived from the
|
|
157
|
+
configuration it just wrote. That derivation is the reason to use it: a unit and
|
|
158
|
+
a configuration that disagree is the failure this whole document is arranged
|
|
159
|
+
around, and two files generated from one source cannot.
|
|
160
|
+
|
|
161
|
+
```sh
|
|
162
|
+
sudo -u pmtiles-swarm -H /var/lib/pmtiles-swarm/node_modules/.bin/pmtiles-swarm init --config /etc/pmtiles-swarm/swarm.config.json --data-dir /var/lib/pmtiles-swarm/data --save-path /mnt/store/torrent-data --systemd --password 'the console password'
|
|
163
|
+
```
|
|
164
|
+
|
|
165
|
+
That writes the configuration, generates `pmtiles-swarm.service` beside it, and
|
|
166
|
+
prints a command for every directory involved — each with the right owner and
|
|
167
|
+
mode. Nothing privileged happens and nothing is installed: the unit is written
|
|
168
|
+
next to the configuration, and copying it into `/etc/systemd/system` is one of
|
|
169
|
+
the lines it hands you.
|
|
170
|
+
|
|
171
|
+
Two refusals worth knowing before you type it. State under `/etc` is rejected
|
|
172
|
+
rather than warned about, because at the moment a configuration is written there
|
|
173
|
+
is nothing to migrate — the same mistake found six months later costs a stopped
|
|
174
|
+
service and a careful move. And a configuration that already exists is left
|
|
175
|
+
alone unless `--force` says otherwise, since the credentials in it are not
|
|
176
|
+
recoverable.
|
|
177
|
+
|
|
178
|
+
`--password` is hashed before it is written. Leave it off and there is no
|
|
179
|
+
console password at all: the generated API key is the way in until you set one.
|
|
180
|
+
There is deliberately no placeholder, because `auth.password` accepts plaintext
|
|
181
|
+
and a placeholder in that field is a working password until somebody notices.
|
|
182
|
+
|
|
183
|
+
**Re-run it after adding a watched folder**, or a subscription with a save path
|
|
184
|
+
of its own. Both are directories the node writes to, and neither reaches
|
|
185
|
+
`ReadWritePaths` by itself. `tools/resume-doctor.py` checks a running node's
|
|
186
|
+
unit against its live configuration and says which directories are missing.
|
|
187
|
+
|
|
188
|
+
The unit below is what it generates, with your paths in it. Worth reading
|
|
189
|
+
either way — a generated file you do not understand is a hand-written one you
|
|
190
|
+
have not written yet.
|
|
191
|
+
|
|
154
192
|
## The unit
|
|
155
193
|
|
|
194
|
+
Annotated, because every directive here has cost somebody a diagnosis.
|
|
195
|
+
`init --systemd` writes this file with your paths already in it; what
|
|
196
|
+
follows is what those lines are for.
|
|
197
|
+
|
|
156
198
|
```ini
|
|
157
199
|
[Unit]
|
|
158
200
|
Description=pmtiles-swarm
|
|
@@ -177,8 +219,12 @@ Restart=always
|
|
|
177
219
|
RestartSec=5
|
|
178
220
|
|
|
179
221
|
# Stopping announces "stopped" to every tracker, releases the data directory
|
|
180
|
-
# lock
|
|
181
|
-
|
|
222
|
+
# lock, cancels downloads in flight and writes resume data for every archive.
|
|
223
|
+
# That last part is what scales: the node allows its engine two seconds per
|
|
224
|
+
# torrent, so a library of fifty wants well over a minute and systemd has to
|
|
225
|
+
# outlast it. Too low here and the process is killed mid-save, which costs a
|
|
226
|
+
# re-hash of everything unwritten on the way back up.
|
|
227
|
+
TimeoutStopSec=300
|
|
182
228
|
|
|
183
229
|
# The node stops the sidecar itself, and needs it alive to do so. The default
|
|
184
230
|
# signals both at once, which kills the sidecar before it can write its resume
|
|
@@ -233,6 +279,17 @@ itself, in order, and waits for the resume data to be written. Nothing is left
|
|
|
233
279
|
running: anything still alive when `TimeoutStopSec` expires is killed, sidecar
|
|
234
280
|
included.
|
|
235
281
|
|
|
282
|
+
**`TimeoutStopSec`, and why it is not a round number you pick once.** Stopping
|
|
283
|
+
writes resume data for every archive, and the node allows its engine two seconds
|
|
284
|
+
per torrent to do it — so the figure that matters grows with the library. Three
|
|
285
|
+
hundred seconds covers a hundred and forty archives; a larger node needs more.
|
|
286
|
+
Set it too low and systemd kills the process mid-save, and every archive that
|
|
287
|
+
had not been written re-hashes its whole store on the way back up, which for a
|
|
288
|
+
700 GiB archive is hours.
|
|
289
|
+
|
|
290
|
+
`tools/resume-doctor.py` prints the arithmetic for your library and says whether
|
|
291
|
+
the unit's value covers it.
|
|
292
|
+
|
|
236
293
|
**No `ExecStop=`.** systemd already sends `SIGTERM`, and the node handles it
|
|
237
294
|
from the moment it starts. `ExecStop=/bin/kill -15 $MAINPID` is redundant, and
|
|
238
295
|
becomes wrong under `KillMode=process` — that one leaves the rest of the unit
|
|
@@ -779,3 +836,111 @@ once rather than diagnosing later:
|
|
|
779
836
|
`journalctl -u pmtiles-swarm | grep -i onComplete` — and a hook redirecting its
|
|
780
837
|
own output to a file will have nothing for the journal to show, which is not
|
|
781
838
|
the same as not having run.
|
|
839
|
+
|
|
840
|
+
## When archives re-check on every start
|
|
841
|
+
|
|
842
|
+
The symptom is always the same shape — a restart, and then a library sitting at
|
|
843
|
+
0% or grinding through a re-hash of archives that are whole on the disk. The
|
|
844
|
+
causes are not the same shape at all, so guessing between them is expensive:
|
|
845
|
+
re-hashing a 700 GiB archive is half an hour during which it serves nobody.
|
|
846
|
+
|
|
847
|
+
Two scripts in `tools/` answer it without guessing. Neither starts anything,
|
|
848
|
+
neither writes anything, and both are safe against a running node.
|
|
849
|
+
|
|
850
|
+
```sh
|
|
851
|
+
sudo -u pmtiles-swarm python3 tools/resume-doctor.py \
|
|
852
|
+
-c /etc/pmtiles-swarm/swarm.config.json
|
|
853
|
+
```
|
|
854
|
+
|
|
855
|
+
**`resume-doctor.py`** reads your configuration, catalog, stored `.torrent`
|
|
856
|
+
files and resume directory, and says for each archive what libtorrent will do on
|
|
857
|
+
the next start and why — seed at once, stat its files, or re-hash the store. It
|
|
858
|
+
applies libtorrent's own rules rather than a guess at them, and cites the file
|
|
859
|
+
and line each one came from. It also reads the unit through `systemctl show`,
|
|
860
|
+
does the arithmetic on every deadline that can cut a resume save short, and
|
|
861
|
+
scans the journal for the failures that leave a trace.
|
|
862
|
+
|
|
863
|
+
```sh
|
|
864
|
+
sudo -u pmtiles-swarm python3 tools/resume-experiment.py
|
|
865
|
+
```
|
|
866
|
+
|
|
867
|
+
**`resume-experiment.py`** proves the behaviour on the libtorrent you actually
|
|
868
|
+
have, rather than the one the documentation was written against. It builds a
|
|
869
|
+
real torrent over a temporary file, saves real resume data, adds it again, and
|
|
870
|
+
reports what happened. 2.0 and 2.1 differ enough to be worth the thirty seconds.
|
|
871
|
+
|
|
872
|
+
Three things they exist to tell apart, because the first two look identical from
|
|
873
|
+
outside and want opposite fixes:
|
|
874
|
+
|
|
875
|
+
- **No resume data is not the problem.** An archive recorded complete is added
|
|
876
|
+
with `seed_mode`, and that claim stands on its own: libtorrent stats the files
|
|
877
|
+
and seeds, without hashing a byte. A missing resume file costs nothing.
|
|
878
|
+
- **Stale resume data is the problem, and it is worse than none.** A resume file
|
|
879
|
+
holding a _partial_ bitfield cancels the `seed_mode` claim — libtorrent drops
|
|
880
|
+
it if a single piece is unset. The archive then comes back as a downloader
|
|
881
|
+
with every byte already on disk, and fetches again what it has. Deleting that
|
|
882
|
+
one file is the repair; with none, it seeds at once.
|
|
883
|
+
- **A short or missing file is a genuine re-check.** `mismatching_file_size` is
|
|
884
|
+
the only thing here that really does mean re-hashing, and it is about the data
|
|
885
|
+
on disk rather than about resume data at all. libtorrent 2.x records no
|
|
886
|
+
mtimes, so nothing in this depends on a timestamp.
|
|
887
|
+
|
|
888
|
+
Which archives carry stale resume data is not random, and the pattern is worth
|
|
889
|
+
knowing before it looks like a second bug. An archive from a watched folder was
|
|
890
|
+
read end to end before it was ever registered: it is complete from its first
|
|
891
|
+
moment and no partial bitfield is ever written for it. An archive joined from a
|
|
892
|
+
feed, a magnet or a peer is genuinely partial for hours, with partial resume
|
|
893
|
+
data written for it every `resumeSaveIntervalSeconds` throughout. Only the
|
|
894
|
+
second kind has anything to leave behind, so trouble concentrating there is this
|
|
895
|
+
failure's ordinary shape rather than something the import path introduced.
|
|
896
|
+
|
|
897
|
+
Two more things the doctor looks at, because they are the ones a snapshot alone
|
|
898
|
+
gets wrong.
|
|
899
|
+
|
|
900
|
+
**Whether the catalog still points at its own `.torrent` files.** Each entry
|
|
901
|
+
records `torrentPath` as an absolute path, and restore reads that rather than
|
|
902
|
+
recomputing it — so moving `dataDir` out of `/etc`, which this guide tells you
|
|
903
|
+
to do, leaves every recorded path behind. Nothing warns, because a missing
|
|
904
|
+
`.torrent` is not treated as an error: restore quietly falls back to the magnet.
|
|
905
|
+
That fallback is the damage. A magnet carries no metadata, resume data does not
|
|
906
|
+
carry it either, and so the archive waits on BEP 9 for a file list that only a
|
|
907
|
+
peer can supply — from a swarm where this node is the origin. It sits at 0% in
|
|
908
|
+
`downloading_metadata` indefinitely, which no amount of re-checking will fix.
|
|
909
|
+
|
|
910
|
+
**Whether the periodic save gets through the whole library.** Every torrent is
|
|
911
|
+
asked to save in the same cycle, so a healthy `resumeDir` has all its files
|
|
912
|
+
within one interval of each other. Ages spread wider than that mean some cycles
|
|
913
|
+
are finishing early, and the ages alone cannot say whether the same archives
|
|
914
|
+
lose every time. `--watch` measures it rather than inferring it:
|
|
915
|
+
|
|
916
|
+
```sh
|
|
917
|
+
sudo -u pmtiles-swarm python3 tools/resume-doctor.py \
|
|
918
|
+
-c /etc/pmtiles-swarm/swarm.config.json --watch 660
|
|
919
|
+
```
|
|
920
|
+
|
|
921
|
+
It reports each save cycle as it happens, how many of them covered every
|
|
922
|
+
archive, and which archives were never written at all.
|
|
923
|
+
|
|
924
|
+
**A poisoned archive may heal itself, so run this twice.** The commonest way a
|
|
925
|
+
whole archive acquires a partial bitfield is a save taken while it was
|
|
926
|
+
re-checking: `write_resume_data` truncates `have_pieces` to the pieces checked
|
|
927
|
+
so far, and the five-minute timer lands inside that window every time a large
|
|
928
|
+
archive checks. That shortened bitfield is what cancels `seed_mode` on the next
|
|
929
|
+
start — so the archive comes back partial, checks again, and writes another one.
|
|
930
|
+
The loop needs no restart to keep itself going.
|
|
931
|
+
|
|
932
|
+
The doctor tells the two apart by length: a bitfield shorter than the torrent
|
|
933
|
+
was written mid-check, and one at full length is what a finished check
|
|
934
|
+
concluded. A check that completes rewrites its own resume file correctly, so an
|
|
935
|
+
archive poisoned in one run can read `seeds at once` in the next with nothing
|
|
936
|
+
done to it. **Only the archives that stay poisoned across two runs, several
|
|
937
|
+
minutes apart, are worth deleting a file for.**
|
|
938
|
+
|
|
939
|
+
Two things it deliberately does not treat as faults. An archive whose data sits
|
|
940
|
+
outside the `savePathLayout` shape is normal — one adopted from a file this node
|
|
941
|
+
already holds keeps that file, and a node may keep archives beside whatever
|
|
942
|
+
produced them. What matters is whether the data is where the entry says, which
|
|
943
|
+
is checked directly. And every directory an archive actually lives in has to
|
|
944
|
+
appear in `ReadWritePaths`, not just the `savePath` in the configuration; the
|
|
945
|
+
doctor checks each one in use, because a deliberately-placed archive is exactly
|
|
946
|
+
the case that gets left out of the unit.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pmtiles-swarm",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.60.0",
|
|
4
4
|
"description": "BitTorrent distribution for PMTiles map archives: create torrents, watch folders, publish and subscribe to RSS feeds, and seed through qBittorrent or an embedded client",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "src/index.js",
|
|
@@ -46,7 +46,7 @@
|
|
|
46
46
|
"maplibre-gl": "^6.2.0",
|
|
47
47
|
"parse-torrent": "^11.0.24",
|
|
48
48
|
"pmtiles": "^4.4.1",
|
|
49
|
-
"pmtiles-torrent": "^0.
|
|
49
|
+
"pmtiles-torrent": "^0.10.2",
|
|
50
50
|
"webtorrent": "^3.0.21"
|
|
51
51
|
},
|
|
52
52
|
"engines": {
|