pagetrace 0.9.1 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -6,6 +6,55 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
6
6
  and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
7
7
  While the version is below 1.0.0, breaking changes ship in a minor release.
8
8
 
9
+ ## [0.10.0] - 2026-09-08
10
+
11
+ ### Added
12
+
13
+ - Internal link integrity. `link.broken` (error) for a link pointing at no page,
14
+ and `link.broken.added` for one this build broke — the diff stays quiet about
15
+ breakage it did not introduce, so CI fails on the regression rather than on
16
+ the backlog. Only broken links are stored, keeping a site's navigation out of
17
+ the lockfile and its diff readable. A `--dir` crawl is authoritative; an
18
+ origin crawl confirms each candidate with a real request, because a sitemap
19
+ routinely omits pages that are live. External links are not checked.
20
+
21
+ - Redirect drift. A crawl now records where each route actually landed, taken
22
+ from the response itself at no extra request. `redirect.added` and
23
+ `redirect.changed` (warn) and `redirect.removed` (info) report a route that
24
+ starts, moves or stops redirecting — a whole class of change that was
25
+ previously invisible, since a redirected route still answers. A redirect that
26
+ only adds or drops a trailing slash is server configuration, not drift, and is
27
+ not reported.
28
+ - `canonical.redirects` (warn): a canonical naming a URL that redirects is one
29
+ the engine may ignore. Only raised when the target was actually crawled.
30
+ - Sitemap health. `sitemap.dead` (error) for a sitemap entry answering 404, and
31
+ `sitemap.redirect` (warn) for one that redirects. Both read data the crawl
32
+ already gathered.
33
+ - `--format sarif` on `check`, for `github/codeql-action/upload-sarif`. Findings
34
+ become pull request annotations and Security tab entries.
35
+ - `pagetrace init`: writes a config file and takes the first snapshot. An
36
+ existing config is never overwritten.
37
+ - Thirteen more Schema.org types in the rich-results table, taking it from 20 to
38
+ 33: Book, Dataset, QAPage, Question, Answer, ClaimReview, Movie, ProfilePage,
39
+ DiscussionForumPosting, ImageObject, SpecialAnnouncement,
40
+ EmployerAggregateRating and PodcastSeries.
41
+
42
+ ### Changed
43
+
44
+ - An origin crawl now honours `robots.txt` `Disallow`, with longest-match-wins
45
+ and Allow winning ties, so `Disallow: /` plus `Allow: /blog` crawls the blog.
46
+ **A staging origin that serves `Disallow: /` now yields no pages** — pass
47
+ `--ignore-robots`, or set `"ignoreRobots": true`, to crawl a site you own.
48
+ - **`sitemap.dead` is an error**, so a sitemap that already lists a dead URL
49
+ fails `check` on the first run after upgrading. The problem predates the rule.
50
+
51
+ ### Fixed
52
+
53
+ - A field absent from an older lockfile read as a transition rather than as
54
+ missing information, because the scalar diff compared `undefined` against
55
+ `null`. Every optional field added from here on would have reported a phantom
56
+ change on every page of the first run.
57
+
9
58
  ## [0.9.1] - 2026-09-07
10
59
 
11
60
  ### Changed
@@ -281,6 +330,7 @@ Initial release. `snapshot`, `check` and `audit` commands; filesystem and HTTP
281
330
  crawling; diff classified by transition; absolute, cross-page and hreflang audit
282
331
  rules; pretty, JSON, markdown, GitHub and HTML reporters.
283
332
 
333
+ [0.10.0]: https://github.com/shyamexe/pagetrace/compare/v0.9.1...v0.10.0
284
334
  [0.9.1]: https://github.com/shyamexe/pagetrace/compare/v0.9.0...v0.9.1
285
335
  [0.9.0]: https://github.com/shyamexe/pagetrace/compare/v0.8.1...v0.9.0
286
336
  [0.8.1]: https://github.com/shyamexe/pagetrace/compare/v0.8.0...v0.8.1
package/README.md CHANGED
@@ -21,13 +21,21 @@ Requires Node 20.19 or newer. No native modules, three small dependencies.
21
21
 
22
22
  ## Use
23
23
 
24
- Record a baseline from your build output:
24
+ Set up a config file and the first baseline in one step:
25
+
26
+ ```bash
27
+ npx pagetrace init --dir ./out
28
+ ```
29
+
30
+ Or record the baseline on its own:
25
31
 
26
32
  ```bash
27
33
  npx pagetrace snapshot --dir ./out
28
34
  git add pagetrace.lock.json
29
35
  ```
30
36
 
37
+ `init` never overwrites an existing config; it refreshes the lockfile and leaves your edits alone.
38
+
31
39
  Check every build against it:
32
40
 
33
41
  ```bash
@@ -113,6 +121,8 @@ npx pagetrace snapshot --url https://example.com --limit 200
113
121
 
114
122
  Routes are discovered from `robots.txt` sitemap declarations, falling back to `/sitemap.xml`. Sitemap indexes are followed one level, up to 50 children, and expansion stops once `--limit` is satisfied. Gzipped children are recognised but not read.
115
123
 
124
+ Paths that `robots.txt` disallows are skipped, with the longest matching rule winning so an `Allow` exception still gets crawled. A staging origin that serves `Disallow: /` would therefore yield nothing — pass `--ignore-robots` (or set `"ignoreRobots": true` in the config) to crawl a site you own anyway.
125
+
116
126
  URLs pointing at another host are skipped. A page that cannot be fetched stops the run with an error rather than being dropped from the snapshot — a page silently missing from a crawl is indistinguishable from a page you deleted, and reporting a transient outage as a site-wide deletion is worse than failing.
117
127
 
118
128
  ### Version and updates
@@ -146,10 +156,20 @@ pagetrace update # install it globally
146
156
  | `aeo.crawler.newly_blocked` | error | `robots.txt` started blocking an AI crawler |
147
157
  | `aeo.llmstxt.removed` | error | `/llms.txt` disappeared |
148
158
  | `page.removed` | warn | A route in the lockfile is no longer there |
159
+ | `redirect.added` | warn | A route that used to answer directly now redirects |
160
+ | `redirect.changed` | warn | A route redirects somewhere new |
161
+ | `canonical.redirects` | warn | A canonical points at a URL that redirects |
162
+ | `link.broken.added` | error | A page started linking to a URL that does not exist |
163
+ | `sitemap.dead` | error | The sitemap lists a URL that answers 404 |
164
+ | `sitemap.redirect` | warn | The sitemap lists a URL that redirects |
149
165
  | `og.removed` / `hreflang.removed` | warn | Social or i18n tags dropped |
150
166
  | `title.changed` | info | Ordinary copy edit |
151
167
 
152
- Alongside the diff, `check` runs absolute rules: missing title, canonical, `h1`, or description; JSON-LD required and recommended properties for the twenty Schema.org types Google supports as rich results; thin content; images without `alt`; and AEO signals like whether the page opens with something an answer engine can quote. Disable with `--no-audit`.
168
+ Internal links are checked too. Only the broken ones are stored, so a site's navigation never lands in the lockfile: `link.broken` for a link that is already dead, `link.broken.added` for one this build broke. A `--dir` crawl is authoritative — the build directory is the whole site — while a crawl confirms each candidate with a real request first, because a sitemap routinely omits pages that are live. External links are deliberately not checked: a Cloudflare 403 and a rate limit both look like a dead page, and that is where link checkers earn their reputation for false positives.
169
+
170
+ Redirects are recorded from the response itself, so they cost no extra requests. A redirect that only adds or drops a trailing slash is server configuration rather than drift and is not reported. `canonical.redirects` is only raised when the canonical's target was actually crawled, so a `--limit` run cannot invent it.
171
+
172
+ Alongside the diff, `check` runs absolute rules: missing title, canonical, `h1`, or description; JSON-LD required and recommended properties for the thirty-three Schema.org types Google supports as rich results; thin content; images without `alt`; and AEO signals like whether the page opens with something an answer engine can quote. Disable with `--no-audit`.
153
173
 
154
174
  ## Config
155
175
 
@@ -200,7 +220,18 @@ Without the Action:
200
220
  - run: npx pagetrace check --dir ./out --baseline-branch origin/main --format github
201
221
  ```
202
222
 
203
- `--format` accepts `pretty`, `json`, `markdown` (sized for a PR comment), and `github` (workflow annotations).
223
+ `--format` accepts `pretty`, `json`, `markdown` (sized for a PR comment), `github` (workflow annotations) and `sarif`.
224
+
225
+ SARIF puts the findings in the Security tab and on the pull request itself, which survives longer than a comment:
226
+
227
+ ```yaml
228
+ - run: npx pagetrace check --dir ./out --baseline-branch origin/main --format sarif > pagetrace.sarif
229
+ - uses: github/codeql-action/upload-sarif@v3
230
+ with:
231
+ sarif_file: pagetrace.sarif
232
+ ```
233
+
234
+ Needs `security-events: write`. Routes are not source files, so GitHub lists each finding without anchoring it to a line in the diff.
204
235
 
205
236
  ## Programmatic API
206
237