pagesight 0.9.0 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -8,93 +8,40 @@ See your site the way search engines and AI see it.
8
8
  npm install pagesight
9
9
  ```
10
10
 
11
- Most SEO tools flag "title over 60 characters" and "only one H1 allowed." [Google's own engineers say those rules don't exist.](#why-not-other-seo-tools) Pagesight skips the myths and goes to the sources.
11
+ Google Search Console + PageSpeed Insights + CrUX + 139 AI crawlers. One package.
12
12
 
13
13
  ## Tools
14
14
 
15
- ### `inspect`
16
-
17
- Ask Google: is this page indexed? What canonical did you choose? Any crawl errors? Structured data issues?
18
-
19
- Returns index status, canonical (yours vs Google's), crawl status, rich results validation, sitemaps, and referring URLs.
20
-
21
- ### `pagespeed`
22
-
23
- Run Google Lighthouse on any URL:
24
-
25
- - **Scores**: performance, accessibility, best-practices, seo
26
- - **Core Web Vitals (lab)**: FCP, LCP, TBT, CLS, Speed Index, TTI
27
- - **CrUX field data**: real Chrome user metrics (page + origin)
28
- - **Opportunities**: ranked by severity with potential savings
29
- - **Strategy**: `mobile` or `desktop`
30
-
31
- ### `crux`
32
-
33
- Real-world Core Web Vitals from Chrome users (28-day rolling window):
34
-
35
- - **Metrics**: LCP, FCP, INP, CLS, TTFB, RTT, navigation types, form factors
36
- - **Granularity**: by URL or origin, by device (DESKTOP, PHONE, TABLET)
37
- - **Data**: p75 values + histogram distributions
38
-
39
- ### `crux_history`
40
-
41
- Core Web Vitals trends over time — up to 40 weekly data points (~10 months):
42
-
43
- - Trend detection (improved/stable/worse) with percentage change
44
- - Recent data points table for core metrics
45
- - Custom period count (1-40)
46
-
47
- ### `performance`
48
-
49
- Google Search Console search analytics:
50
-
51
- - **Dimensions**: `query`, `page`, `country`, `device`, `date`, `searchAppearance`, `hour`
52
- - **Search types**: `web`, `image`, `video`, `news`, `discover`, `googleNews`
53
- - **Filters**: `equals`, `contains`, `notEquals`, `notContains`, `includingRegex`, `excludingRegex`
54
- - **Pagination**: up to 25,000 rows
55
-
56
- ### `robots`
57
-
58
- Analyze any site's robots.txt:
59
-
60
- - **Syntax validation** per [RFC 9309](https://www.rfc-editor.org/rfc/rfc9309)
61
- - **AI crawler audit** — 139+ bots from the [ai-robots-txt](https://github.com/ai-robots-txt/ai.robots.txt) community registry
62
- - **Bot categories**: training scrapers, AI search crawlers, AI assistants, AI agents
63
- - **Per-bot status**: blocked or allowed, with the matched rule
15
+ | Tool | What it does |
16
+ |------|-------------|
17
+ | `audit` | One-call site audit. Runs pagespeed + metatags + robots + sitemaps + inspect in parallel. Returns prioritized findings. |
18
+ | `pagespeed` | Lighthouse scores, Core Web Vitals, opportunities with resource URLs, failing audits with selectors and fix links. |
19
+ | `metatags` | OG, Twitter Card, canonical, JSON-LD with schema validation, redirect chain, image validation. |
20
+ | `inspect` | Google index status, canonical choice, crawl state, rich results. |
21
+ | `sample_inspect` | Sample URLs from a sitemap and batch-inspect via GSC. Diagnoses indexing patterns. |
22
+ | `performance` | Search analytics — clicks, impressions, CTR, position. `compare: true` for period-over-period deltas. |
23
+ | `crux` | Real-world Core Web Vitals from Chrome users (p75, histograms). |
24
+ | `crux_history` | CWV trends over time — up to 40 weekly data points. |
25
+ | `robots` | robots.txt validation (RFC 9309) + AI crawler audit (139+ bots). |
26
+ | `sitemaps` | GSC properties and sitemaps with submitted/indexed counts. |
27
+ | `setup` | Auth status and OAuth setup. |
64
28
 
65
29
  ```
66
- === robots.txt: https://www.nytimes.com ===
67
- AI Crawlers: 35 blocked, 104 allowed (of 139 known)
68
-
69
- BLOCKED GPTBot (OpenAI) — GPT model training
70
- BLOCKED ClaudeBot (Anthropic) — Claude model training
71
- ALLOWED Claude-User (Anthropic) — User-initiated fetching
30
+ === Site Audit: https://fipe.chat ===
31
+
32
+ HIGH Missing canonical URL
33
+ HIGH 7,772 sitemap URLs submitted, 0 indexed
34
+ MEDIUM Missing og:image — no social preview image
35
+ MEDIUM Accessibility score: 89/100
36
+ LOW Missing Twitter Card tags
37
+ LOW No structured data (JSON-LD) found
72
38
  ```
73
39
 
74
- ### `sitemaps`
75
-
76
- Search Console properties and sitemaps (read-only):
77
-
78
- - `list_sites` — all properties with permission level
79
- - `get_site` — details for a specific property
80
- - `list_sitemaps` — sitemaps with error/warning counts
81
- - `get_sitemap` — full details for a specific sitemap
82
-
83
- ### `setup`
84
-
85
- Check auth status or walk through OAuth interactively.
86
-
87
40
  ## Setup
88
41
 
89
- ### 1. Google Cloud project
90
-
91
- 1. Go to [Google Cloud Console](https://console.cloud.google.com/)
92
- 2. Create a project (or use existing)
93
- 3. Enable: **Search Console API**, **PageSpeed Insights API**, **Chrome UX Report API**
94
- 4. Create **OAuth client ID** (Desktop app) — for Search Console
95
- 5. Create **API key** — for PageSpeed and CrUX
96
-
97
- ### 2. Configure
42
+ 1. [Google Cloud Console](https://console.cloud.google.com/) — enable Search Console API, PageSpeed Insights API, Chrome UX Report API
43
+ 2. Create OAuth client ID (Desktop app) + API key
44
+ 3. Configure:
98
45
 
99
46
  ```env
100
47
  GSC_CLIENT_ID=your-client-id.apps.googleusercontent.com
@@ -103,16 +50,16 @@ GSC_REFRESH_TOKEN=your-refresh-token
103
50
  GOOGLE_API_KEY=your-api-key
104
51
  ```
105
52
 
106
- The `robots` tool works without any credentials.
53
+ `robots`, `metatags`, and `pagespeed` work without credentials.
107
54
 
108
- ### 3. Use with your AI assistant
55
+ ### MCP config
109
56
 
110
57
  ```json
111
58
  {
112
59
  "mcpServers": {
113
60
  "pagesight": {
114
- "command": "bun",
115
- "args": ["run", "/path/to/pagesight/src/index.ts"],
61
+ "command": "npx",
62
+ "args": ["pagesight"],
116
63
  "env": {
117
64
  "GSC_CLIENT_ID": "your-client-id",
118
65
  "GSC_CLIENT_SECRET": "your-secret",
@@ -124,36 +71,15 @@ The `robots` tool works without any credentials.
124
71
  }
125
72
  ```
126
73
 
127
- Then just ask:
128
-
129
- ```
130
- "Is https://mysite.com indexed?"
131
- "Run pagespeed on my homepage"
132
- "Which AI crawlers can access my site?"
133
- "How have my Core Web Vitals changed?"
134
- "Which queries bring traffic to this page?"
135
- ```
136
-
137
74
  ## Why not other SEO tools?
138
75
 
139
76
  We checked every common SEO "rule" against official Google documentation:
140
77
 
141
- - **"Title must be under 60 characters"** — Google: "there's no limit." Gary Illyes: "an externally made-up metric."
142
- - **"Meta description must be 155 characters"** — Google: "there's no limit on how long a meta description can be."
143
- - **"Only one H1 per page"** — John Mueller: "You can use H1 tags as often as you want. There's no limit."
144
- - **"Minimum 300 words per page"** — Mueller: "the number of words on a page is not a quality factor."
145
- - **"Text-to-HTML ratio matters"** — Mueller: "it makes absolutely no sense at all for SEO."
78
+ - **"Title must be under 60 characters"** — Gary Illyes: "an externally made-up metric."
79
+ - **"Only one H1 per page"** — John Mueller: "You can use H1 tags as often as you want."
80
+ - **"Minimum 300 words per page"** — Mueller: "not a quality factor."
146
81
 
147
- Tools that flag these are reporting their opinions. Pagesight only reports what the sources actually return.
148
-
149
- ## Development
150
-
151
- ```bash
152
- bun install
153
- bun run start # start server
154
- bun run lint # biome check
155
- bun run format # biome format
156
- ```
82
+ Pagesight only reports what the sources actually return.
157
83
 
158
84
  ## License
159
85
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pagesight",
3
- "version": "0.9.0",
3
+ "version": "0.10.0",
4
4
  "description": "See your site the way search engines and AI see it.",
5
5
  "keywords": [
6
6
  "seo",
@@ -91,7 +91,7 @@ function formatOpportunities(audits: Record<string, PsiAudit>): string[] {
91
91
  }
92
92
 
93
93
  function formatDiagnostics(audits: Record<string, PsiAudit>): string[] {
94
- const failing: Array<{ title: string; displayValue: string }> = [];
94
+ const failing: PsiAudit[] = [];
95
95
 
96
96
  for (const audit of Object.values(audits)) {
97
97
  if (
@@ -100,15 +100,31 @@ function formatDiagnostics(audits: Record<string, PsiAudit>): string[] {
100
100
  (audit.scoreDisplayMode === "numeric" || audit.scoreDisplayMode === "metricSavings") &&
101
101
  audit.displayValue
102
102
  ) {
103
- failing.push({ title: audit.title, displayValue: audit.displayValue });
103
+ failing.push(audit);
104
104
  }
105
105
  }
106
106
 
107
107
  if (failing.length === 0) return [];
108
108
 
109
+ failing.sort((a, b) => (a.score ?? 0) - (b.score ?? 0));
110
+
109
111
  const lines: string[] = ["--- Diagnostics ---", ""];
110
- for (const item of failing.slice(0, 10)) {
111
- lines.push(`${item.title}: ${item.displayValue}`);
112
+ for (const audit of failing.slice(0, 10)) {
113
+ lines.push(`${audit.title}: ${audit.displayValue}`);
114
+
115
+ const linkMatch = audit.description?.match(/\[.*?\]\((https?:\/\/[^)]+)\)/);
116
+ if (linkMatch) lines.push(` Learn more: ${linkMatch[1]}`);
117
+
118
+ const items = audit.details?.items;
119
+ if (items && items.length > 0) {
120
+ for (const item of items.slice(0, 3)) {
121
+ lines.push(...formatDetailItem(item));
122
+ }
123
+ if (items.length > 3) {
124
+ lines.push(` ... and ${items.length - 3} more`);
125
+ }
126
+ }
127
+ lines.push("");
112
128
  }
113
129
 
114
130
  return lines;