pmtiles-swarm 0.7.1 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -7,6 +7,75 @@
7
7
  ### 🐞 Bug fixes
8
8
  - _...Add new stuff here..._
9
9
 
10
+ ## 0.9.0
11
+ ### ✨ Features and improvements
12
+ - **`pmtiles-swarm status`**, which asks a running node what it is doing and reads the answer out
13
+ loud. It takes the same config file the node runs with, so the address, the port and the
14
+ credential come from one place rather than being remembered and retyped. That is the whole
15
+ point of it: the API is on `adminPort` rather than the public port, the node binds where `host`
16
+ says and that is usually not loopback, and it accepts `authorization: Bearer` and not
17
+ `x-api-key`. Get any one of those wrong by hand and the answer is a refused connection or a 401,
18
+ both of which read as a broken node rather than as a mistyped command — which is exactly how
19
+ they were read while diagnosing the archive fixed below.
20
+
21
+ It names the case that is otherwise silent: an archive the catalog holds and the engine does
22
+ not, which through `curl` is a row of empty columns and looks like a corrupt archive. Just after
23
+ a start it is normal and passes; persisting, the engine refused it and the log says why. Exits
24
+ non-zero when the node does not answer or its engine is down, so it can be the last step of a
25
+ deployment script, and `--json` hands back the raw replies for anything that would rather parse.
26
+
27
+ Also warns when `--config` names a file that is not there. Startup ignores that on purpose, so
28
+ a first run can write one — but for a question about a running node the silence is
29
+ misleading, since the answer then describes the default address and looks entirely real.
30
+
31
+ - **[docs/haproxy.md](docs/haproxy.md) now covers the backend pool**: why round robin rather than
32
+ the plugin's default of Source-IP Hash, which fails quietly behind a CDN by pinning nearly all
33
+ traffic to one node while the rest sit idle and healthy; when least-connections or URI hash are
34
+ worth having instead; and what HTTP/2 on the frontend does and does not change about balancing.
35
+
36
+ ### 🐞 Bug fixes
37
+ - **An archive built from a watched folder no longer sits at 0%, seeding nobody, for a quarter of
38
+ an hour.** The libtorrent engine dropped `seedOnly` on its way to the sidecar, so libtorrent
39
+ re-hashed an 81 GiB archive that had been read end to end moments earlier to produce its
40
+ torrent. Everything else already handled it — the library sets it in five places, the
41
+ composite engine checks it against what the primary reports, qBittorrent has its own flag for
42
+ it — and this one engine silently did not pass it on. Needs pmtiles-torrent 0.4.1, which the
43
+ existing dependency range picks up on a fresh install.
44
+ - **`docs/running-as-a-service.md` no longer suggests checking a node with `curl localhost:8091`.**
45
+ It names loopback and sends no credential, so on a node bound to its LAN address with a key
46
+ configured it fails twice over, in the two ways that look most like a broken node. It now uses
47
+ the status command.
48
+
49
+ ## 0.8.0
50
+ ### ✨ Features and improvements
51
+ - **`GET /health`**, for a load balancer: 200 when this node can serve, 503 when its engine
52
+ cannot, no credential and nothing to parse. It asks the engine rather than itself, which is the
53
+ distinction that makes it worth having — a feed is built from the catalogue and never touches
54
+ the swarm, so a balancer checking `/feed.xml`, the nearest thing that existed, gets 200 from a
55
+ node whose engine is dead and keeps sending it traffic. The answer is cached for two seconds,
56
+ because a balancer asks often and each check is an inter-process round trip, and it is sent
57
+ `no-store`: a stale health check keeps a dead node in rotation for as long as whatever cached
58
+ it says so.
59
+ - **`GET /archives/<infohash>/ready`**, which answers a different question — whether a *particular*
60
+ archive has become servable on this node. 200 once its header and, for vector, its layers have
61
+ been read; 503 with which half is missing; **415** for an archive that can never be served,
62
+ since MBTiles is distributed here but cannot be read a byte range at a time and polling it would
63
+ be polling for ever; 404 when it is not here. It reports rather than acts, starting no read and
64
+ waiting for nothing — a probe that does work on demand is a probe that can be used to make a
65
+ node do work on demand.
66
+ - **[docs/haproxy.md](docs/haproxy.md)**, written against the OPNsense plugin: the health monitor
67
+ field by field, how the check interval trades against failover time, timeouts long enough for a
68
+ web seed to finish, `X-Forwarded-Proto`, gating a deployment on `/ready` — and a table of what a
69
+ reverse proxy in front of a BitTorrent node simply cannot carry.
70
+
71
+ ### 🐞 Bug fixes
72
+ - **Head-warming no longer tries to read a PMTiles header out of a `.osm.pbf`.** The guard was
73
+ `entry.kind && entry.kind !== 'pmtiles'`, and `guessKind` answers `undefined` for anything it
74
+ does not recognise — so it never fired for exactly the archives it existed to exclude. Every
75
+ planet dump being mirrored, and every MBTiles archive, was read as though it had a header,
76
+ failed, and came back on the backoff for ever. The test is positive now: the kind has to *be*
77
+ PMTiles, taken from the entry where it is known and from the file name where it is not.
78
+
10
79
  ## 0.7.1
11
80
  ### 🐞 Bug fixes
12
81
  - **A feed no longer deletes its only complete copy.** Retention was written for a watched folder
package/README.md CHANGED
@@ -33,6 +33,9 @@ node src/index.js --config swarm.config.json
33
33
  address.
34
34
  - **[docs/running-as-a-service.md](docs/running-as-a-service.md)** — a systemd unit, why
35
35
  `Restart=always` is required rather than optional, and which ports want a firewall rule.
36
+ - **[docs/haproxy.md](docs/haproxy.md)** — health monitors, what a reverse proxy carries and
37
+ what it cannot, timeouts long enough for a web seed, and gating a deployment on
38
+ `/archives/<infohash>/ready`.
36
39
  - **[docs/architecture-diagram.md](docs/architecture-diagram.md)** — how a publishing node, a
37
40
  serving tier, the swarm and both kinds of client fit together.
38
41
 
@@ -586,6 +589,38 @@ network equipment: peers request 16 KiB blocks whatever the piece size. The sett
586
589
  matters there is `maxConnections`, since every peer holds a NAT table entry. See
587
590
  [docs/publishing.md](docs/publishing.md).
588
591
 
592
+ ## Asking a running node what it is doing
593
+
594
+ ```sh
595
+ node src/index.js status --config /etc/pmtiles-swarm/swarm.config.json
596
+ ```
597
+
598
+ ```
599
+ engine libtorrent ready
600
+ version 0.8.0
601
+ 17 archives, 1 the engine does not know about
602
+
603
+ NAME SIZE STATE PROGRESS
604
+ planetiler-openmaptiles-260803.pmtiles 81 GiB seeding 100%
605
+ planet-260803.osm.pbf 94 GiB downloading 37%
606
+ planetiler-openmaptiles-260810.pmtiles 83 GiB — —
607
+ ```
608
+
609
+ It reads the same config file the node runs with, so the address, the port and the credential come
610
+ from one place rather than being remembered and retyped. That matters more than it sounds: the API
611
+ is on `adminPort`, not the public port; the node binds where `host` says, which is usually not
612
+ loopback; and it accepts `authorization: Bearer`, not `x-api-key`. Get any one of those wrong with
613
+ `curl` and the answer is a refused connection or a 401 — both of which read as a broken node rather
614
+ than a mistyped command.
615
+
616
+ An archive with a state of `—` is one the catalog holds and the engine is not. Directly after a
617
+ restart that is normal and passes. Persisting, it means the engine refused it, and the log says
618
+ why.
619
+
620
+ Exit status is 0 when the node answered and its engine is up, 1 when it did not or is not — so it
621
+ works in a deployment script. `--json` gives the raw `/api/status` and `/api/torrents` replies for
622
+ anything that wants to parse rather than read.
623
+
589
624
  ## API
590
625
 
591
626
  | Method | Path | Purpose |
@@ -626,6 +661,8 @@ matters there is `maxConnections`, since every peer holds a NAT table entry. See
626
661
  | `GET` | `/api/catalog` | The whole catalogue, for a peer keeping itself in step |
627
662
  | `GET` | `/archives/:infoHash/tiles.json` | TileJSON — **public** |
628
663
  | `GET` | `/archives/:infoHash/:z/:x/:y.:ext` | One tile — **public** |
664
+ | `GET` | `/health` | 200 when this node can serve, 503 when its engine cannot — **public**, no credential, for a load balancer |
665
+ | `GET` | `/archives/:infoHash/ready` | Whether this node can serve *this* archive yet: 200 ready, 503 not yet, 415 never — **public** |
629
666
  | `GET` | `/archives/:infoHash/archive.torrent` | The `.torrent` a torrent-aware client joins with — **public** |
630
667
  | `GET` | `/archives/:infoHash/preview` | Map preview for one archive — admin, not public |
631
668
  | `GET` | `/latest/:category/tiles.json`, `/:name.torrent`, `/magnet` | The newest build in a category — **public**. The torrent name is yours to choose, so a link can read `planetiler-openmaptiles-latest.torrent`; it redirects to the immutable URL, which names the download after the build it actually is |
@@ -0,0 +1,235 @@
1
+ # Behind HAProxy
2
+
3
+ Written against the **OPNsense HAProxy plugin**, which is configured through a
4
+ form rather than a file. The field names below are that plugin's; the generated
5
+ configuration is shown alongside so the same thing can be written by hand.
6
+
7
+ ## What HAProxy carries, and what it cannot
8
+
9
+ This is the first thing to get straight, because a reverse proxy in front of a
10
+ BitTorrent node carries less than it looks like it does.
11
+
12
+ | | Through HAProxy? |
13
+ | --- | --- |
14
+ | Tiles, TileJSON, feeds, `.torrent` files, health checks | **Yes** — ordinary HTTP on 8090 |
15
+ | The console and API on 8091 | **No.** Do not publish it |
16
+ | BitTorrent peers on 6881 | **No.** Not HTTP; needs its own forward |
17
+ | WebRTC to browser peers | **No.** Outbound only, and needs nothing |
18
+
19
+ A node reachable *only* through HAProxy has no incoming peers at all. It will
20
+ still seed to peers it connects to, and still serve tiles, but it is half a
21
+ member of the swarm. 6881 wants a firewall rule of its own, straight to the
22
+ node, TCP **and** UDP.
23
+
24
+ ## The health monitor
25
+
26
+ **Services → HAProxy → Settings → Health Monitors → Add**
27
+
28
+ | Field | Value |
29
+ | --- | --- |
30
+ | Name | `pmtiles-swarm` |
31
+ | Check type | `HTTP` |
32
+ | Check interval | `2s` or `5s` — see below |
33
+ | Port to check | *empty*, so it uses the real server's port |
34
+ | HTTP method | `GET` |
35
+ | Request URI | `/health` |
36
+ | HTTP version | `HTTP/1.1` |
37
+ | HTTP host | anything; `localhost` is fine |
38
+ | Custom HTTP check → Expression | exact string match for the HTTP status code |
39
+ | Custom HTTP check → Value | `200` |
40
+
41
+ `HTTP/1.1` requires a `Host` header, which is why the field is there. The node
42
+ does not route by it, so its value only matters if something in front of it
43
+ does.
44
+
45
+ Which produces:
46
+
47
+ ```
48
+ backend pmtiles-swarm
49
+ option httpchk GET /health HTTP/1.1
50
+ http-check expect status 200
51
+ server node1 172.16.1.49:8090 check inter 5s
52
+ server node2 172.16.1.41:8090 check inter 5s
53
+ ```
54
+
55
+ **Do not point the monitor at `/` or at `/feed.xml`.** Both answer 200 from a
56
+ node whose engine is dead — a feed is built from the catalogue and never
57
+ touches the swarm — so the balancer would keep sending traffic to a node that
58
+ cannot serve a tile. `/health` asks the engine and answers 503 when it does not
59
+ reply.
60
+
61
+ ### How often to check
62
+
63
+ `/health` caches its answer for two seconds, so anything faster than that asks
64
+ the same question twice and gets the same answer. At two seconds, near enough
65
+ every check is a real round trip to the engine — which enumerates every torrent
66
+ the node holds, so the cost grows with the catalogue.
67
+
68
+ What the interval buys is failover speed, and it is the interval multiplied by
69
+ HAProxy's `fall` count, which defaults to 3:
70
+
71
+ | Check interval | Marked down after |
72
+ | --- | --- |
73
+ | `2s` | about 6 seconds |
74
+ | `5s` | about 15 seconds |
75
+
76
+ Neither is wrong. Six seconds is worth having if a node dying mid-request
77
+ matters; fifteen is worth having if it does not, and costs the node less. What
78
+ is not worth having is a sub-second interval, which only asks a cached answer
79
+ more often.
80
+
81
+ Note also that the node comes *back* on `rise` successful checks — 2 by
82
+ default — so a flapping node re-enters rotation quickly whichever you choose.
83
+
84
+ ## The backend pool
85
+
86
+ **Settings → Backend Pools → Add**, mode **HTTP (Layer 7)**.
87
+
88
+ ### Balancing algorithm
89
+
90
+ **Round Robin.** Tile requests are stateless, numerous and roughly the same
91
+ size, which is the case round robin is for.
92
+
93
+ **Not Source-IP Hash, even though it is the default here.** It exists for
94
+ sticky sessions, and tiles have no session to be sticky about. Worse, it fails
95
+ quietly in exactly the setup this document assumes: behind a CDN every request
96
+ arrives from a handful of edge addresses, so hashing on the source pins almost
97
+ all traffic to one node while the others sit idle and healthy. A NAT'd office
98
+ or a mobile carrier does the same thing on a smaller scale.
99
+
100
+ Two others are worth knowing about, for arrangements this is not:
101
+
102
+ **Least Connections**, when long transfers share the backend with short ones.
103
+ A web-seed range request can run for minutes while a tile takes milliseconds,
104
+ and round robin will happily queue tiles behind a transfer. If web seeds are
105
+ served from a different host — an ordinary web server in front of the published
106
+ directory, which is the usual arrangement — that variance is not here and round
107
+ robin is simpler.
108
+
109
+ **URI Hash**, for a tier of **cache-mode** nodes. Such a node holds only the
110
+ pieces it has read, so sending the same region to different nodes makes each of
111
+ them pay the cold read separately — the one real cost of scaling a cache-mode
112
+ tier horizontally. `balance uri depth 2` hashes on `/archives/<infohash>` and
113
+ gives archive affinity, at the price of concentrating one archive on one node.
114
+ For nodes holding complete copies there is no cold read to avoid, so this is a
115
+ cost with no benefit.
116
+
117
+ ### HTTP/2
118
+
119
+ Enable it on the frontend and leave *HTTP/2 without TLS* unchecked: the client
120
+ gets HTTP/2, the node is spoken to over HTTP/1.1, which is what it speaks.
121
+ Balancing in HTTP mode is per request rather than per connection, so a client
122
+ multiplexing a hundred tile requests over one HTTP/2 connection still has them
123
+ spread across the pool.
124
+
125
+ ## Real servers
126
+
127
+ **Settings → Real Servers → Add**, one per node, port **8090**.
128
+
129
+ The admin port has no place in a public backend. It carries the console and the
130
+ API, and while both are behind authentication, an internet-facing sign-in page
131
+ is a thing to decide on deliberately rather than to acquire by copying a
132
+ backend.
133
+
134
+ ## Tell the node it is behind a proxy
135
+
136
+ HAProxy terminates TLS, so without being told, the node believes every request
137
+ arrived over plain HTTP and advertises `http://` tile URLs. A browser that
138
+ loaded the map over HTTPS then blocks every one of them as mixed content, which
139
+ looks like an empty map rather than a configuration mistake.
140
+
141
+ Two ways, and they are alternatives rather than a pair:
142
+
143
+ ```json
144
+ { "publicUrl": "https://swarm.example.org" }
145
+ ```
146
+
147
+ One canonical URL whatever the request said. Simple, and right when the node is
148
+ only ever reached one way.
149
+
150
+ ```json
151
+ { "trustProxy": "172.16.1.0/24" }
152
+ ```
153
+
154
+ Derive it per request from `X-Forwarded-Proto` and the `Host`, so the same node
155
+ answers correctly on its LAN address and through the proxy. **Name the proxy
156
+ rather than saying `true`**: the node also listens on its LAN address directly,
157
+ and `true` would let anything that can reach that port claim any protocol and
158
+ host it likes, which decides what URLs your feed hands out.
159
+
160
+ For that to work HAProxy has to send the header. In the plugin it is
161
+ **Public Service → Advanced → Option pass-through**, or in configuration:
162
+
163
+ ```
164
+ frontend https-in
165
+ http-request set-header X-Forwarded-Proto https
166
+ ```
167
+
168
+ `Host` is passed through by default.
169
+
170
+ ## Timeouts
171
+
172
+ The defaults are wrong for this. A web seed is an HTTP range request for part of
173
+ an archive, and a peer pulling a large one holds the connection for as long as
174
+ the transfer takes.
175
+
176
+ ```
177
+ defaults
178
+ timeout connect 5s
179
+ timeout client 1h
180
+ timeout server 1h
181
+ ```
182
+
183
+ An hour is not generous here; a client fetching tens of gigabytes over a slow
184
+ link exceeds anything shorter, and the failure looks like a corrupt download
185
+ rather than a timeout.
186
+
187
+ ## Deploying a new build
188
+
189
+ `/health` says whether a node should be sent traffic. It says nothing about
190
+ whether a *particular* archive has become servable there, which is the question
191
+ worth asking after a build lands.
192
+
193
+ ```sh
194
+ INFOHASH=5e1c143c400d15aaacfb1c748d4ab6d1b46c5df5
195
+ for node in 172.16.1.49 172.16.1.41; do
196
+ until curl -fsS "http://$node:8090/archives/$INFOHASH/ready" >/dev/null; do
197
+ sleep 10
198
+ done
199
+ echo "$node is ready"
200
+ done
201
+ ```
202
+
203
+ Ask each node **directly**, not through the balancer — through it you learn that
204
+ *some* node is ready, which is not the same thing and is exactly the wrong
205
+ answer when deciding whether to move traffic.
206
+
207
+ `/ready` answers 415 for an archive that can never be served, so a loop like the
208
+ one above would never end if an MBTiles archive reached it. Treat 415 as a
209
+ separate case if that is possible.
210
+
211
+ ## Cloudflare in front
212
+
213
+ Two rules worth setting deliberately.
214
+
215
+ **Never cache `/health` or `/archives/*/ready`.** Both send `no-store` and
216
+ Cloudflare honours it, but a page rule that caches everything can override that.
217
+ A cached health check keeps a dead node in rotation for as long as the cache
218
+ says so, which is worse than having no check at all.
219
+
220
+ **Cache tiles by infohash aggressively.** `/archives/<infohash>/…` is immutable
221
+ by construction — an infohash names those bytes and no others — and is served
222
+ with `max-age=31536000, immutable`. `/latest/<category>/…` is the opposite: it
223
+ moves on every build and is served with `max-age=300`.
224
+
225
+ WebRTC does not pass through Cloudflare, and neither does BitTorrent. Browser
226
+ peers reach the node over ICE, and a `wss://` tracker is the only part of that
227
+ which is HTTP at all.
228
+
229
+ ## When it looks like the swarm is empty
230
+
231
+ Almost always the peer port rather than the proxy. HAProxy carries none of it,
232
+ so a node behind one with no forward for 6881 shows connected peers only where
233
+ it made the connection itself. See
234
+ [docs/engines.md](engines.md#ports-and-reachability) for which ports want
235
+ forwarding and which want nothing.
@@ -441,9 +441,23 @@ Then check it is actually serving:
441
441
 
442
442
  ```sh
443
443
  curl -fsS localhost:8090/feed.xml >/dev/null && echo "public surface ok"
444
- curl -fsS localhost:8091/api/status | head -c 200
444
+ node /opt/pmtiles-swarm/src/index.js status \
445
+ --config /etc/pmtiles-swarm/swarm.config.json
445
446
  ```
446
447
 
448
+ Ask through the status command rather than with `curl`. Reaching the API by hand
449
+ means getting the bind address, the admin port and the credential right in one
450
+ go, and each of them fails in a way that looks like a broken node: a node bound
451
+ to its LAN address refuses a request to `localhost`, and the header it accepts
452
+ is `authorization: Bearer`, so anything else is a 401. The status command reads
453
+ all three out of the config file the service is running with. It exits non-zero
454
+ when the node does not answer or its engine is down, so it also works as the
455
+ last step of a deployment script.
456
+
457
+ An archive listed with a state of `—` is one the catalog holds and the engine is
458
+ not. Just after a start that is normal and passes within a minute or so. If it
459
+ persists, the engine refused it, and the journal says why.
460
+
447
461
  And that it can write where it is supposed to, which nothing above proves:
448
462
 
449
463
  ```sh
@@ -200,6 +200,85 @@ For an archive published as a mutable torrent, the block also carries
200
200
  rather than pinning to the version the document was generated from. See
201
201
  [publishing](publishing.md).
202
202
 
203
+ ## Health checks
204
+
205
+ ```
206
+ GET /health
207
+ ```
208
+
209
+ 200 when this node can serve, 503 when it cannot, no credential and no body
210
+ worth parsing — which is what a load balancer needs. It is on the public
211
+ surface, so it answers on the same port the tiles do.
212
+
213
+ **It asks the engine, not just itself.** A reply means the engine answered a
214
+ round trip, which is the difference worth reporting: a feed is built from the
215
+ catalog and never touches the swarm, so a balancer checking `/feed.xml` gets
216
+ 200 from a node whose engine is dead and keeps sending it traffic.
217
+
218
+ The answer is cached for a couple of seconds, because a balancer asks often and
219
+ each check costs an inter-process round trip. A node that has just died leaves
220
+ rotation one check later than it otherwise would.
221
+
222
+ It sends `cache-control: no-store`. A stale health check is worse than none —
223
+ it keeps a dead node in rotation for as long as whatever cached it says so.
224
+
225
+ ```
226
+ backend tiles
227
+ option httpchk GET /health
228
+ http-check expect status 200
229
+ server node1 10.0.0.11:8090 check inter 5s
230
+ ```
231
+
232
+ Configured through a form rather than a file, and with the rest of what a proxy
233
+ in front of this needs — timeouts, `X-Forwarded-Proto`, and the ports it cannot
234
+ carry — in [docs/haproxy.md](haproxy.md).
235
+
236
+ ### Whether one archive is servable yet
237
+
238
+ ```
239
+ GET /archives/<infohash>/ready
240
+ ```
241
+
242
+ A different question, and worth keeping apart from the one above. `/health`
243
+ decides whether a node should be sent traffic at all; this says whether a
244
+ newly published archive has become servable *here* — which is what you want
245
+ after a build lands and before pointing anything at it.
246
+
247
+ | | |
248
+ | --- | --- |
249
+ | **200** | Ready. Its header has been read, and a vector archive has its layers |
250
+ | **503** | Not yet — ask again. The body says which half is missing |
251
+ | **415** | Never. MBTiles is distributed here but cannot be read a byte range at a time, so waiting would be waiting for ever |
252
+ | **404** | Not on this node |
253
+
254
+ The codes differ because the responses differ: one is "poll me", one is "stop
255
+ polling", and a script that treats them alike either gives up too early or
256
+ waits for something that will never happen.
257
+
258
+ It **reports rather than acts** — it starts no read and waits for nothing. A
259
+ probe that does work on demand is a probe that can be used to make a node do
260
+ work on demand. Reading the head is the head warmer's job; this only says
261
+ whether it has happened.
262
+
263
+ So the shape of a deployment check across a serving tier is: publish, then poll
264
+ every node until each answers 200, then move `latest`.
265
+
266
+ ```sh
267
+ for node in 10.0.0.11 10.0.0.12; do
268
+ until curl -fsS "http://$node:8090/archives/$INFOHASH/ready" >/dev/null; do
269
+ sleep 10
270
+ done
271
+ done
272
+ ```
273
+
274
+ `curl -f` fails on 503 and on 415 alike, so treat 415 separately if an MBTiles
275
+ archive could ever reach that loop — otherwise it never ends.
276
+
277
+ An archive this node holds *completely* can still be read from disk with the
278
+ engine down, so 503 is a statement about the node rather than about every
279
+ request it could answer. That is the right way round for a balancer with
280
+ somewhere else to send the traffic.
281
+
203
282
  ## Caching
204
283
 
205
284
  Tiles are served `Cache-Control: public, max-age=31536000, immutable`.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pmtiles-swarm",
3
- "version": "0.7.1",
3
+ "version": "0.9.0",
4
4
  "description": "BitTorrent distribution for PMTiles map archives: create torrents, watch folders, publish and subscribe to RSS feeds, and seed through qBittorrent or an embedded client",
5
5
  "type": "module",
6
6
  "main": "src/index.js",
package/src/api.js CHANGED
@@ -13,6 +13,7 @@ import {
13
13
  isPublicSurface,
14
14
  } from './auth.js';
15
15
  import { normalizeCategories } from './catalog.js';
16
+ import { guessKind } from './library.js';
16
17
  import { QBittorrentEngine } from './engines/qbittorrent.js';
17
18
  import { RESTART_REQUIRED, redactConfig, saveConfig } from './config.js';
18
19
  import { freeSpace, listLocations } from './locations.js';
@@ -397,6 +398,59 @@ export function createApp({
397
398
  }),
398
399
  );
399
400
 
401
+ /**
402
+ * Whether this node should be sent traffic.
403
+ *
404
+ * For a load balancer, which needs three things this did not have: no
405
+ * credential, a cheap answer, and a status code rather than a body to parse.
406
+ * A balancer checking `/feed.xml` instead — the nearest thing that existed —
407
+ * gets 200 from a node whose engine is dead, because a feed is built from
408
+ * the catalog and never touches the swarm.
409
+ *
410
+ * Readiness rather than liveness. `engine.list()` is a round trip to the
411
+ * engine, so a reply means the sidecar is answering and not merely that Node
412
+ * is; that is the difference worth reporting, since everything that makes
413
+ * this node useful to a swarm goes through it.
414
+ *
415
+ * Cached, because a balancer asks every couple of seconds and an IPC round
416
+ * trip per check is a cost with nothing to show for it. The window is short
417
+ * enough that a node which has just died is out of rotation within one more
418
+ * check than it would have been.
419
+ */
420
+ let healthChecked = 0;
421
+ let healthOk = true;
422
+ let healthError;
423
+ const HEALTH_TTL_MS = 2000;
424
+
425
+ app.get(
426
+ '/health',
427
+ route(async (_req, res) => {
428
+ const now = Date.now();
429
+ if (now - healthChecked >= HEALTH_TTL_MS) {
430
+ healthChecked = now;
431
+ try {
432
+ await engine.list();
433
+ healthOk = true;
434
+ healthError = undefined;
435
+ } catch (error) {
436
+ healthOk = false;
437
+ healthError = error.message;
438
+ }
439
+ }
440
+
441
+ // Never cached anywhere. A stale health check is worse than none: it
442
+ // keeps a dead node in rotation for as long as whatever cached it says.
443
+ res.setHeader('cache-control', 'no-store');
444
+ res.setHeader('access-control-allow-origin', '*');
445
+ res.status(healthOk ? 200 : 503).json({
446
+ status: healthOk ? 'ok' : 'unavailable',
447
+ engine: engine.name,
448
+ version: VERSION,
449
+ ...(healthOk ? {} : { error: healthError }),
450
+ });
451
+ }),
452
+ );
453
+
400
454
  app.get(
401
455
  '/api/status',
402
456
  route(async (_req, res) => {
@@ -1828,6 +1882,95 @@ export function createApp({
1828
1882
  // with authentication configured this answered 401 — to the very callers it
1829
1883
  // exists for, since this is the URL the TileJSON torrent block advertises and
1830
1884
  // the one a syncing peer follows.
1885
+ /**
1886
+ * Whether this node can serve *this* archive yet.
1887
+ *
1888
+ * A different question from `/health`, and answered separately because
1889
+ * nobody asks it per request. `/health` decides whether a node should be
1890
+ * sent traffic at all; this decides whether a newly published archive has
1891
+ * become servable here — which is what you want to know after a build lands
1892
+ * and before pointing anything at it.
1893
+ *
1894
+ * Reports rather than acts. It starts no read and waits for nothing: a probe
1895
+ * that does work on demand is a probe that can be used to make a node do
1896
+ * work on demand. The head warmer is what makes an archive ready; this only
1897
+ * says whether it has.
1898
+ *
1899
+ * The three answers are deliberately different codes, because they call for
1900
+ * different responses from whoever asked. 503 is "not yet, ask again". 415
1901
+ * is "never" — an MBTiles archive is distributed here and cannot be read a
1902
+ * byte range at a time, so waiting for it would be waiting for ever. 404 is
1903
+ * "not here at all".
1904
+ */
1905
+ app.get(
1906
+ '/archives/:infoHash/ready',
1907
+ route(async (req, res) => {
1908
+ const entry = catalog.get(req.params.infoHash);
1909
+ // Never cached. The whole value of this is that it changes.
1910
+ res.setHeader('cache-control', 'no-store');
1911
+ res.setHeader('access-control-allow-origin', '*');
1912
+
1913
+ if (!entry) {
1914
+ return res.status(404).json({ ready: false, reason: 'unknown archive' });
1915
+ }
1916
+
1917
+ const kind = entry.kind ?? guessKind(entry.name ?? '');
1918
+ const shape = {
1919
+ infoHash: entry.infoHash,
1920
+ name: entry.name,
1921
+ kind: kind ?? 'unknown',
1922
+ complete: entry.complete === true,
1923
+ };
1924
+
1925
+ if (kind !== 'pmtiles') {
1926
+ return res.status(415).json({
1927
+ ...shape,
1928
+ ready: false,
1929
+ reason:
1930
+ `this is ${kind ? `a ${kind}` : 'not a PMTiles'} archive, and only ` +
1931
+ 'PMTiles can be read a byte range at a time — it will not become ' +
1932
+ 'servable by waiting',
1933
+ });
1934
+ }
1935
+
1936
+ const summary = entry.pmtiles;
1937
+ // A summary that names a format is one a header was actually read for.
1938
+ // Anything else is a partial left by a read that did not finish.
1939
+ if (!summary?.format) {
1940
+ return res.status(503).json({
1941
+ ...shape,
1942
+ ready: false,
1943
+ reason: 'its header has not been read yet',
1944
+ });
1945
+ }
1946
+
1947
+ // Vector tiles without their layer list can be served, but nothing can
1948
+ // be styled from them: the TileJSON a map asks for would carry no
1949
+ // vector_layers. The metadata sits wherever the writer put it, which for
1950
+ // planetiler is after every tile, so it routinely arrives long after the
1951
+ // header.
1952
+ if (summary.format === 'pbf' && !summary.vectorLayers) {
1953
+ return res.status(503).json({
1954
+ ...shape,
1955
+ ready: false,
1956
+ format: summary.format,
1957
+ reason: 'its metadata has not been read yet, so it carries no vector layers',
1958
+ });
1959
+ }
1960
+
1961
+ res.json({
1962
+ ...shape,
1963
+ ready: true,
1964
+ format: summary.format,
1965
+ minZoom: summary.minZoom,
1966
+ maxZoom: summary.maxZoom,
1967
+ ...(summary.vectorLayers
1968
+ ? { vectorLayers: summary.vectorLayers.length }
1969
+ : {}),
1970
+ });
1971
+ }),
1972
+ );
1973
+
1831
1974
  app.get(
1832
1975
  '/archives/:infoHash/archive.torrent',
1833
1976
  route(async (req, res) => {
package/src/auth.js CHANGED
@@ -144,6 +144,8 @@ export function isPublicSurface(path) {
144
144
  if (/^\/archives\/[^/]+\/preview\/?$/.test(path)) return false;
145
145
 
146
146
  return (
147
+ // A load balancer checks this, and it checks the public port.
148
+ path === '/health' ||
147
149
  path === '/api/catalog' ||
148
150
  path === '/api/catalog/' ||
149
151
  path === '/feed.xml' ||
@@ -254,6 +254,12 @@ export class LibtorrentEngine {
254
254
  savePath: request.savePath ?? this.#options.savePath,
255
255
  mode: request.mode,
256
256
  paused: request.paused,
257
+ // The caller's claim that the data is already on disk, which for an
258
+ // archive created here it is: the file was read end to end a moment ago
259
+ // to produce the torrent. Without passing it on, libtorrent hashes the
260
+ // whole archive again before seeding a byte — a quarter of an hour for
261
+ // an 81 GiB build, during which it reads as 0% and serves nobody.
262
+ seedOnly: request.seedOnly,
257
263
  });
258
264
  return result.infoHash;
259
265
  }
package/src/index.js CHANGED
@@ -92,20 +92,26 @@ function createOneEngine(name, config) {
92
92
  * @returns {Promise<void>} - Resolves once listening.
93
93
  */
94
94
  async function main() {
95
- const { values } = parseArgs({
95
+ const { values, positionals } = parseArgs({
96
96
  options: {
97
97
  config: { type: 'string', short: 'c' },
98
98
  port: { type: 'string', short: 'p' },
99
99
  help: { type: 'boolean', short: 'h' },
100
+ json: { type: 'boolean' },
100
101
  },
101
- allowPositionals: false,
102
+ allowPositionals: true,
102
103
  });
103
104
 
104
105
  if (values.help) {
105
106
  console.log(`pmtiles-swarm — BitTorrent distribution for PMTiles archives
106
107
 
108
+ Usage:
109
+ pmtiles-swarm [--config FILE] start the node
110
+ pmtiles-swarm status [--config FILE] ask a running node what it is doing
111
+
107
112
  --config, -c path to a JSON config file
108
113
  --port, -p override the listen port
114
+ --json machine-readable output, for the status command
109
115
  --help, -h this message
110
116
 
111
117
  Environment: PMTILES_SWARM_PORT, PMTILES_SWARM_DATA_DIR, PMTILES_SWARM_ENGINE,
@@ -118,6 +124,23 @@ PMTILES_SWARM_PUBLIC_URL
118
124
  const config = await loadConfig(values.config);
119
125
  if (values.port) config.port = Number(values.port);
120
126
 
127
+ // Asking rather than starting. Everything it needs — which address the admin
128
+ // listener is on, which port, and the credential — comes from the same
129
+ // configuration the node runs with, so there is nothing to pass and nothing
130
+ // to get wrong.
131
+ if (positionals[0] === 'status') {
132
+ const { runStatus } = await import('./status-command.js');
133
+ process.exitCode = await runStatus(config, { json: values.json });
134
+ return;
135
+ }
136
+
137
+ if (positionals.length > 0) {
138
+ console.error(`unknown command: ${positionals[0]}`);
139
+ console.error('try: pmtiles-swarm status');
140
+ process.exitCode = 2;
141
+ return;
142
+ }
143
+
121
144
  // Everything that has to be stopped, in the order it should be stopped,
122
145
  // filled in as startup proceeds.
123
146
  //
package/src/prewarm.js CHANGED
@@ -21,6 +21,8 @@
21
21
  * the first few seconds if the right few kilobytes are asked for first.
22
22
  */
23
23
 
24
+ import { guessKind } from './library.js';
25
+
24
26
  /** The first wait after an attempt that did not finish the job. */
25
27
  const DEFAULT_BACKOFF_SECONDS = 15;
26
28
 
@@ -91,9 +93,20 @@ export class HeadWarmer {
91
93
  * @returns {boolean} - True to attempt a read.
92
94
  */
93
95
  due(entry) {
94
- // Only PMTiles has a head worth reading. A .osm.pbf from a feed is not an
95
- // archive this can say anything about.
96
- if (entry.kind && entry.kind !== 'pmtiles') return false;
96
+ // Only PMTiles has a head worth reading, and this has to be a positive
97
+ // test rather than the absence of a negative one.
98
+ //
99
+ // `guessKind` answers `undefined` for anything it does not recognise — a
100
+ // .osm.pbf from a feed, for instance — so `entry.kind && entry.kind !==
101
+ // 'pmtiles'` never fired for exactly the archives it was meant to exclude.
102
+ // Every planet dump being mirrored was read as though it had a PMTiles
103
+ // header, failed, and came back on the backoff for ever.
104
+ //
105
+ // Taken from the entry where it is known and from the name where it is
106
+ // not, since an archive joined by magnet has no kind until its metadata
107
+ // arrives.
108
+ const kind = entry.kind ?? guessKind(entry.name ?? '');
109
+ if (kind !== 'pmtiles') return false;
97
110
 
98
111
  // A summary that names a format is one a header was actually read for.
99
112
  // Anything else — an empty object, or one left behind by a read that raced
@@ -0,0 +1,233 @@
1
+ /**
2
+ * `pmtiles-swarm status` — asking a running node what it is doing.
3
+ *
4
+ * This exists because interrogating a node meant getting four separate things
5
+ * right at once: which address the admin listener is bound to, which port,
6
+ * which header carries the credential, and where in the JSON the answer lives.
7
+ * Getting any one wrong produces something that looks like a broken archive —
8
+ * a refused connection, a 401, or a row of nulls — rather than like a mistyped
9
+ * command. Every one of those is derivable from the configuration file the
10
+ * node is already running with, so nothing here needs to be passed or
11
+ * remembered.
12
+ */
13
+
14
+ import { access } from 'node:fs/promises';
15
+
16
+ /**
17
+ * Whether a path can be read.
18
+ * @param {string} path - The path.
19
+ * @returns {Promise<boolean>} - True when it is there.
20
+ */
21
+ async function readable(path) {
22
+ try {
23
+ await access(path);
24
+ return true;
25
+ } catch {
26
+ return false;
27
+ }
28
+ }
29
+
30
+ /** Columns, and how wide the name column may grow before it is cut. */
31
+ const NAME_WIDTH = 44;
32
+
33
+ /**
34
+ * A size in bytes, as a person would write it.
35
+ * @param {number} value - Bytes.
36
+ * @returns {string} - e.g. "81 GiB".
37
+ */
38
+ export function bytes(value) {
39
+ if (!Number.isFinite(value) || value <= 0) return '—';
40
+ const units = ['B', 'KiB', 'MiB', 'GiB', 'TiB'];
41
+ let size = value;
42
+ let unit = 0;
43
+ while (size >= 1024 && unit < units.length - 1) {
44
+ size /= 1024;
45
+ unit += 1;
46
+ }
47
+ return `${size < 10 && unit > 0 ? size.toFixed(1) : Math.round(size)} ${units[unit]}`;
48
+ }
49
+
50
+ /**
51
+ * Where this node's API is, according to its own configuration.
52
+ *
53
+ * The admin listener where there is one, since that is where the API lives on
54
+ * a node that separates them. A wildcard bind is reported as loopback: `::` is
55
+ * what the node listens on, not an address anything can connect to.
56
+ * @param {object} config - Resolved configuration.
57
+ * @returns {string} - An origin, e.g. "http://172.16.1.49:8091".
58
+ */
59
+ export function adminUrl(config) {
60
+ const host = config.adminHost ?? config.host ?? '127.0.0.1';
61
+ const port = config.adminPort ?? config.port ?? 8090;
62
+ const reachable =
63
+ host === '0.0.0.0' || host === '::' || host === '' ? '127.0.0.1' : host;
64
+ // A bare IPv6 address needs brackets before it is a URL.
65
+ const bracketed =
66
+ reachable.includes(':') && !reachable.startsWith('[')
67
+ ? `[${reachable}]`
68
+ : reachable;
69
+ return `http://${bracketed}:${port}`;
70
+ }
71
+
72
+ /**
73
+ * The header a request to this node's API needs, if any.
74
+ *
75
+ * `authorization: Bearer`, which is the only form the node accepts — not
76
+ * `x-api-key`, whatever the convention elsewhere.
77
+ * @param {object} config - Resolved configuration.
78
+ * @returns {object} - Headers to send.
79
+ */
80
+ export function authHeaders(config) {
81
+ const key = config.auth?.apiKey;
82
+ return key ? { authorization: `Bearer ${key}` } : {};
83
+ }
84
+
85
+ /**
86
+ * One line per archive, plus what the engine says about the node.
87
+ * @param {object} answer - `{ status, torrents }` as the API returned them.
88
+ * @returns {string} - The report.
89
+ */
90
+ export function formatStatus({ status, torrents }) {
91
+ const lines = [];
92
+ const engine = status?.engine;
93
+ lines.push(
94
+ `engine ${engine?.name ?? 'unknown'}` +
95
+ (engine?.ok === false ? ` UNAVAILABLE — ${engine.error ?? ''}` : ' ready'),
96
+ );
97
+ if (status?.version) lines.push(`version ${status.version}`);
98
+
99
+ const rows = torrents ?? [];
100
+ // An archive the engine has never heard of is the case worth naming. It is
101
+ // in the catalog, it has a size, and every live column is empty — which
102
+ // reads as a broken archive and is usually a node that has not finished
103
+ // starting, or one that could not add it.
104
+ const unknown = rows.filter((row) => !row.status).length;
105
+ lines.push(
106
+ `${rows.length} archive${rows.length === 1 ? '' : 's'}` +
107
+ (unknown > 0 ? `, ${unknown} the engine does not know about` : ''),
108
+ );
109
+ lines.push('');
110
+
111
+ if (rows.length === 0) return `${lines.join('\n')}\n`;
112
+
113
+ const head =
114
+ 'NAME'.padEnd(NAME_WIDTH) +
115
+ 'SIZE'.padStart(9) +
116
+ ' ' +
117
+ 'STATE'.padEnd(12) +
118
+ 'PROGRESS'.padStart(8);
119
+ lines.push(head);
120
+
121
+ for (const row of rows) {
122
+ const name =
123
+ row.name.length > NAME_WIDTH - 1
124
+ ? `${row.name.slice(0, NAME_WIDTH - 2)}…`
125
+ : row.name;
126
+ const state = row.status?.state ?? (row.paused ? 'paused' : '—');
127
+ const progress =
128
+ typeof row.status?.progress === 'number'
129
+ ? `${Math.round(row.status.progress * 100)}%`
130
+ : '—';
131
+ lines.push(
132
+ name.padEnd(NAME_WIDTH) +
133
+ bytes(row.size).padStart(9) +
134
+ ' ' +
135
+ String(state).padEnd(12) +
136
+ progress.padStart(8),
137
+ );
138
+ }
139
+
140
+ if (unknown > 0) {
141
+ lines.push('');
142
+ lines.push(
143
+ 'An archive with no state is one the engine is not holding. If the node',
144
+ );
145
+ lines.push(
146
+ 'has just started it may still be handing them back; if it persists, the',
147
+ );
148
+ lines.push('log will say why it could not be added.');
149
+ }
150
+
151
+ return `${lines.join('\n')}\n`;
152
+ }
153
+
154
+ /**
155
+ * Asks a running node for its status and prints it.
156
+ * @param {object} config - Resolved configuration.
157
+ * @param {object} [options] - Injectable fetch and output, for testing.
158
+ * @returns {Promise<number>} - The exit code.
159
+ */
160
+ export async function runStatus(config, options = {}) {
161
+ const {
162
+ fetch: get = globalThis.fetch,
163
+ out = (text) => process.stdout.write(text),
164
+ err = (text) => process.stderr.write(text),
165
+ json = false,
166
+ } = options;
167
+
168
+ // A named config file that is not there is silently ignored on startup, so
169
+ // that a first run can write one. Here that silence is misleading: a typo in
170
+ // --config means this reports on the default address with no key, which is a
171
+ // different node from the one that was asked about, and the answer looks
172
+ // real. Say it, rather than letting it be discovered later.
173
+ if (config.configPath && !(await readable(config.configPath))) {
174
+ err(
175
+ `no config file at ${config.configPath} — using defaults, ` +
176
+ 'which is probably not the node you meant.\n',
177
+ );
178
+ }
179
+
180
+ const base = adminUrl(config);
181
+ const headers = authHeaders(config);
182
+
183
+ let status;
184
+ let torrents;
185
+ try {
186
+ const [statusReply, torrentsReply] = await Promise.all([
187
+ get(`${base}/api/status`, { headers }),
188
+ get(`${base}/api/torrents`, { headers }),
189
+ ]);
190
+
191
+ // Said plainly, because a 401 here means the key in this configuration is
192
+ // not the key the node is running with — which is a different problem from
193
+ // the node being down, and looks identical without being told.
194
+ if (statusReply.status === 401 || statusReply.status === 403) {
195
+ err(
196
+ `${base} refused the credential in this configuration file.\n` +
197
+ 'The node is running, but with a different auth.apiKey.\n',
198
+ );
199
+ return 1;
200
+ }
201
+ if (!statusReply.ok) {
202
+ err(`${base}/api/status answered ${statusReply.status}\n`);
203
+ return 1;
204
+ }
205
+
206
+ status = await statusReply.json();
207
+ torrents = torrentsReply.ok ? await torrentsReply.json() : [];
208
+ } catch (error) {
209
+ // A refused connection is the commonest failure and the least obvious: the
210
+ // node binds where the configuration says, which is often not loopback.
211
+ // Node's fetch reports every one of them as "fetch failed" and puts the
212
+ // part worth reading — refused, timed out, no such host — in `cause`.
213
+ const reason = error.cause?.code
214
+ ? `${error.message} (${error.cause.code})`
215
+ : error.message;
216
+ err(
217
+ `could not reach ${base}: ${reason}\n` +
218
+ 'That address comes from adminHost and adminPort in this configuration ' +
219
+ 'file.\nIs the node running, and bound where this says?\n',
220
+ );
221
+ return 1;
222
+ }
223
+
224
+ if (json) {
225
+ out(`${JSON.stringify({ status, torrents }, null, 2)}\n`);
226
+ } else {
227
+ out(formatStatus({ status, torrents }));
228
+ }
229
+
230
+ // Usable from a script: the engine being unreachable is the thing worth
231
+ // failing on, and it is what /health reports to a load balancer.
232
+ return status?.engine?.ok === false ? 1 : 0;
233
+ }