@pitlane/crawler 0.2.0 → 0.2.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (3) hide show
  1. package/CHANGELOG.md +22 -27
  2. package/README.md +30 -24
  3. package/package.json +3 -3
package/CHANGELOG.md CHANGED
@@ -1,40 +1,35 @@
1
1
  # @pitlane/crawler
2
2
 
3
+ ## 0.2.2
4
+
5
+ Documentation only. No crawler code changed.
6
+
7
+ - The npm description is now "In-memory route crawling and static path discovery for Remix.", short enough to read in a registry listing.
8
+ - Three things the README described loosely now match what the code does. Under `spider`, a redirect queues a relative `Location` rather than any same-origin target. The `Crawl failed` message drops `statusText` when the response carries none. An absolute or protocol-relative href is skipped because requests are dispatched under a placeholder origin, which is a firmer reason than "belongs to another origin".
9
+ - `staticPaths` leaves out a route pinned to a protocol or a hostname, since its href is not a path. That exclusion was never written down.
10
+ - The options table says `none` where a default does not exist, and a Documentation section links the crawling guide, the prerendering guide, and the API reference.
11
+
12
+ ## 0.2.1
13
+
14
+ Target Remix `3.0.0-rc.2`.
15
+
16
+ - No crawler code changed. `crawl()` still dispatches into a router's `fetch` and yields the same `{ pathname, filepath, response }` records.
17
+ - The `remix` peer stays at `^3.0.0-rc.1`, which already admits rc.2. Nothing needs to move to install this alongside either prerelease.
18
+ - Tested against `remix@3.0.0-rc.2`.
19
+
3
20
  ## 0.2.0
4
21
 
5
22
  Target Remix `3.0.0-rc.1`.
6
23
 
7
- - Raised the `remix` peer dependency to `^3.0.0-rc.1` (from
8
- `^3.0.0-beta.10`). The crawler's own API is unchanged — `crawl()` still
9
- dispatches into a router's `fetch` and yields the same
10
- `{ pathname, filepath, response }` records.
11
- - rc.1 moved the framework's DOM attributes into the `data-rmx-*` namespace,
12
- which is visible here because the crawler decides what to follow: the opt-out
13
- an app writes on an anchor is now `data-rmx-document`, not `rmx-document`.
14
- Following itself reads `href`, `rel`, and `<meta name="robots">`, so no
15
- crawler code changed.
24
+ - Raised the `remix` peer dependency to `^3.0.0-rc.1` (from `^3.0.0-beta.10`). The crawler's own API is unchanged — `crawl()` still dispatches into a router's `fetch` and yields the same `{ pathname, filepath, response }` records.
25
+ - rc.1 moved the framework's DOM attributes into the `data-rmx-*` namespace, which is visible here because the crawler decides what to follow: the opt-out an app writes on an anchor is now `data-rmx-document`, not `rmx-document`. Following itself reads `href`, `rel`, and `<meta name="robots">`, so no crawler code changed.
16
26
  - Tested against `remix@3.0.0-rc.1`.
17
27
 
18
28
  ## 0.1.0
19
29
 
20
30
  Initial release.
21
31
 
22
- - `crawl(router, options)` — walks an app by dispatching requests into its
23
- router's `fetch`, yielding `{ pathname, filepath, response }` per path.
24
- Follows `<a href>` and `<link rel="alternate">`, queues the assets a page
25
- references, honours `rel="nofollow"` and `<meta name="robots">`, skips
26
- cross-origin and non-navigable hrefs, and visits each path once. A redirect
27
- yields nothing and reports through `onRedirect`, since there is no document
28
- to write and the app still answers the path at runtime; under `spider` the
29
- same-origin target is queued instead. Any other non-2xx response aborts the
30
- crawl. `paths`, `spider`, `assets`, `concurrency`, `ignorePageNofollow`, and
31
- `onRedirect` configure it.
32
- - `staticPaths(routes)` — the paths a Remix 3 route map can serve with no
33
- params, deduplicated and sorted. `GET` and method-agnostic routes whose
34
- patterns declare no variables or wildcards.
35
- - The API comes from [remix-run/remix#11150](https://github.com/remix-run/remix/pull/11150),
36
- which was closed with the implementation kept beside the Remix docs site.
37
- Two deliberate differences: `assets` is a new option, because a bundler has
38
- usually emitted those files already, and the first error wins over the last
39
- when several paths fail under concurrency.
32
+ - `crawl(router, options)` — walks an app by dispatching requests into its router's `fetch`, yielding `{ pathname, filepath, response }` per path. Follows `<a href>` and `<link rel="alternate">`, queues the assets a page references, honours `rel="nofollow"` and `<meta name="robots">`, skips cross-origin and non-navigable hrefs, and visits each path once. A redirect yields nothing and reports through `onRedirect`, since there is no document to write and the app still answers the path at runtime; under `spider` the same-origin target is queued instead. Any other non-2xx response aborts the crawl. `paths`, `spider`, `assets`, `concurrency`, `ignorePageNofollow`, and `onRedirect` configure it.
33
+ - `staticPaths(routes)` — the paths a Remix 3 route map can serve with no params, deduplicated and sorted. `GET` and method-agnostic routes whose patterns declare no variables or wildcards.
34
+ - The API comes from [remix-run/remix#11150](https://github.com/remix-run/remix/pull/11150), which was closed with the implementation kept beside the Remix docs site. Two deliberate differences: `assets` is a new option, because a bundler has usually emitted those files already, and the first error wins over the last when several paths fail under concurrency.
40
35
  - Tested against `remix@3.0.0-beta.10`.
package/README.md CHANGED
@@ -1,10 +1,10 @@
1
1
  # @pitlane/crawler
2
2
 
3
- Spider a [Remix 3](https://remix.run) fetch router in memory.
3
+ Spider a [Remix](https://remix.run) fetch router in memory.
4
4
 
5
- `crawl(router)` dispatches requests straight into `router.fetch` and yields every response it gets back, following the links each page contains. No socket, no server, no browser, no HTTP: the router is the whole transport, so an app can be walked wherever the app itself runs.
5
+ `crawl(router)` follows links in rendered pages and yields successful responses. Requests go directly to the router's fetch handler, without an HTTP server.
6
6
 
7
- That makes prerendering a `for await` loop.
7
+ Prerendering is then a `for await` loop that writes each response to disk:
8
8
 
9
9
  ```ts
10
10
  import { crawl } from "@pitlane/crawler";
@@ -31,46 +31,46 @@ vp add @pitlane/crawler
31
31
 
32
32
  Requires `remix@^3.0.0-rc.1` as a peer.
33
33
 
34
- Using the [`remix()` Vite plugin](https://pitlane.tools/package/dev/)? You do not need this package directly — `remix({ prerender })` runs it for you. See the [prerendering guide](https://pitlane.tools/guides/prerendering).
34
+ The [`remix()` Vite plugin](https://pitlane.tools/package/dev/) already depends on this package and runs it for `remix({ prerender })`, so a prerendered Vite app needs no direct install. The [prerendering guide](https://pitlane.tools/guides/prerendering) covers that path.
35
35
 
36
- For everything else a walk is good for — static exports, sitemaps, link checks, render smoke tests — see the [crawling guide](https://pitlane.tools/guides/crawler).
36
+ Install it directly for the other uses of a walk: static exports, sitemaps, link checks, render smoke tests. Those are in the [crawling guide](https://pitlane.tools/guides/crawler).
37
37
 
38
38
  ## `crawl(router, options?)`
39
39
 
40
40
  Returns an async iterator of `{ pathname, filepath, response }`, one per fetched path.
41
41
 
42
- - `pathname` — the path that was requested.
43
- - `filepath` — where the response belongs on disk. HTML gets `<pathname>/index.html` so a static host serves it back for the original path; everything else keeps its own path.
44
- - `response` — the router's response, body unread.
42
+ - `pathname` is the path that was requested.
43
+ - `filepath` is where the response belongs on disk. HTML gets `<pathname>/index.html` so a static host serves it back for the original path; everything else keeps its own path.
44
+ - `response` is the router's response, body unread.
45
45
 
46
46
  Results arrive in completion order, and every path is fetched at most once.
47
47
 
48
- A redirect yields nothing: there is no document to write, and the app still answers the path at runtime. It reports through `onRedirect` instead, and when `spider` is on the same-origin target is queued, so a crawl seeded at a `/` that points elsewhere still finds the site. Any other non-2xx response aborts the crawl with `Crawl failed: <status> <statusText> (<pathname>)`.
48
+ A redirect yields nothing: there is no document to write, and the app still answers the path at runtime. It reports through `onRedirect` instead, and with `spider` on, a relative `Location` is queued, so a crawl seeded at a `/` that points elsewhere still finds the site. Any other non-2xx response aborts the crawl with `Crawl failed: <status> <statusText> (<pathname>)`, dropping `statusText` when the response carries none.
49
49
 
50
- | Option | Type | Default | Purpose |
51
- | -------------------- | ------------------------------- | ------- | -------------------------------------------------------------------------------------------------------------------------------------- |
52
- | `paths` | `string[]` | `["/"]` | Where to start. |
53
- | `spider` | `boolean` | `true` | Follow `<a href>` and `<link rel="alternate">` to find more paths. |
54
- | `assets` | `boolean` | `true` | Queue the `<link href>`, `<script src>`, and `<img src>` each page references. Turn it off when a bundler already emitted those files. |
55
- | `concurrency` | `number` | `1` | How many paths to fetch at once. |
56
- | `ignorePageNofollow` | `(pathname: string) => boolean` | — | Crawl a page's links even though the page asked robots not to follow them. |
57
- | `onRedirect` | `(pathname, location) => void` | — | Called for a path that redirected instead of returning a document. `location` is `null` when the redirect named none. |
50
+ | Option | Type | Default | Purpose |
51
+ | --- | --- | --- | --- |
52
+ | `paths` | `string[]` | `["/"]` | Where to start. |
53
+ | `spider` | `boolean` | `true` | Follow `<a href>` and `<link rel="alternate">` to find more paths. |
54
+ | `assets` | `boolean` | `true` | Queue the `<link href>`, `<script src>`, and `<img src>` each page references. Turn it off when a bundler already emitted those files. |
55
+ | `concurrency` | `number` | `1` | How many paths to fetch at once. |
56
+ | `ignorePageNofollow` | `(pathname: string) => boolean` | none | Crawl a page's links even though the page asked robots not to follow them. |
57
+ | `onRedirect` | `(pathname, location) => void` | none | Called for a path that redirected instead of returning a document. `location` is `null` when the redirect named none. |
58
58
 
59
59
  The first argument is anything with a `fetch(request: Request)` method: a `createRouter()` router, a built server bundle's default export, a worker-style `{ fetch }` object.
60
60
 
61
- ### What spidering respects
61
+ ### What spidering skips
62
62
 
63
- Crawling stops where a crawler should stop, so a run over a real site does not wander:
63
+ The spider leaves these alone:
64
64
 
65
65
  - `rel="nofollow"` on a link, and `<meta name="robots" content="nofollow">` (or `googlebot`) on a page.
66
- - Absolute and protocol-relative URLs, which belong to another origin.
66
+ - Absolute and protocol-relative URLs. Requests are dispatched under a placeholder origin, so an href that names a host is out of the crawl's reach.
67
67
  - `#fragment`, `mailto:`, `tel:`, `javascript:`, and `data:` hrefs.
68
68
 
69
- `ignorePageNofollow` is the escape hatch for the case where a page's `nofollow` is aimed at search engines rather than at you — a versioned docs tree that should not be indexed but does need to be built.
69
+ Use `ignorePageNofollow` when a page's `nofollow` is aimed at search engines rather than at the build, such as a versioned docs tree that should not be indexed but does need to be written out.
70
70
 
71
71
  ## `staticPaths(routes)`
72
72
 
73
- The question that comes before a crawl: which paths can this app serve with no params?
73
+ `staticPaths` reads a route map and returns the paths the app can serve with no params, which is what a prerender pass needs before it knows any dynamic values:
74
74
 
75
75
  ```ts
76
76
  import { staticPaths } from "@pitlane/crawler";
@@ -85,7 +85,7 @@ let routes = route({
85
85
  staticPaths(routes); // ["/", "/blog"]
86
86
  ```
87
87
 
88
- A route qualifies when it answers `GET` (or any method) and its pattern declares no variables or wildcards. `/blog/:slug` is left out, because its values live outside the route map. Results are deduplicated and sorted, so a build that renders them lists its output the same way every time.
88
+ A route qualifies when it answers `GET` (or any method) and its pattern declares no variables or wildcards. `/blog/:slug` is left out, because its values live outside the route map, and so is a route pinned to a protocol or hostname, because its href is not a path. Results are deduplicated and sorted, so a build that renders them lists its output the same way every time.
89
89
 
90
90
  ## Provenance
91
91
 
@@ -94,7 +94,13 @@ The `crawl` API comes from [remix-run/remix#11150](https://github.com/remix-run/
94
94
  - `assets` is new. Upstream always queues a page's assets, which is right for a site with no bundler and wrong for one where Vite already emitted them.
95
95
  - The first error wins when several paths fail at once, rather than the last.
96
96
 
97
- `staticPaths` has no upstream counterpart. It is the Remix 3 answer to React Router's `getStaticPaths`: a Remix router exposes no route table, but the route map an app builds it from is an ordinary object, and that is the thing worth reading.
97
+ `staticPaths` has no upstream counterpart. It covers the ground React Router's `getStaticPaths` covers: a Remix router exposes no route table, but the route map an app builds it from is an ordinary object, so `staticPaths` walks that instead.
98
+
99
+ ## Documentation
100
+
101
+ - [Crawling guide](https://pitlane.tools/guides/crawler)
102
+ - [Prerendering guide](https://pitlane.tools/guides/prerendering)
103
+ - [API reference](https://pitlane.tools/package/crawler/)
98
104
 
99
105
  ## License
100
106
 
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "@pitlane/crawler",
3
- "version": "0.2.0",
4
- "description": "crawl() — spider a Remix 3 fetch router in memory to prerender it, plus static-path discovery from a route map.",
3
+ "version": "0.2.2",
4
+ "description": "In-memory route crawling and static path discovery for Remix.",
5
5
  "keywords": [
6
6
  "crawl",
7
7
  "fetch-router",
@@ -35,7 +35,7 @@
35
35
  },
36
36
  "devDependencies": {
37
37
  "@types/node": "^25.5.0",
38
- "remix": "3.0.0-rc.1",
38
+ "remix": "3.0.0-rc.2",
39
39
  "typescript": "^7.0.2",
40
40
  "vite-plus": "^0.2.6"
41
41
  },