pagesight 0.17.0 → 0.19.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (79) hide show
  1. package/README.md +20 -82
  2. package/docs/credentials.md +54 -0
  3. package/docs/diagnostics.md +85 -0
  4. package/docs/measurement.md +178 -0
  5. package/docs/opportunities.md +78 -0
  6. package/docs/snapshots.md +130 -0
  7. package/docs/usage.md +135 -0
  8. package/package.json +26 -29
  9. package/src/api/assessment.ts +315 -0
  10. package/src/api/bing.ts +83 -0
  11. package/src/api/compare-snapshots.ts +167 -0
  12. package/src/api/discover.ts +39 -0
  13. package/src/api/doctor.ts +35 -0
  14. package/src/api/evidence-schema.ts +78 -0
  15. package/src/api/evidence.ts +99 -0
  16. package/src/api/execute.ts +160 -0
  17. package/src/api/ga-freshness.ts +20 -0
  18. package/src/api/ga-realtime.ts +40 -0
  19. package/src/api/index.ts +5 -0
  20. package/src/api/opportunities.ts +317 -0
  21. package/src/api/report-table.ts +216 -0
  22. package/src/api/reports.ts +108 -0
  23. package/src/api/schema.ts +271 -0
  24. package/src/api/snapshot.ts +202 -0
  25. package/src/api/ui-findings.ts +68 -0
  26. package/src/assessment-text.ts +55 -0
  27. package/src/cli.ts +217 -0
  28. package/src/http.ts +39 -0
  29. package/src/index.ts +8 -25
  30. package/src/mcp-server.ts +27 -0
  31. package/src/mcp.ts +5 -0
  32. package/src/opportunities-text.ts +61 -0
  33. package/src/providers/bing.ts +48 -0
  34. package/src/{lib → providers}/crux.ts +4 -9
  35. package/src/providers/ga.ts +114 -0
  36. package/src/providers/google-tokens.ts +86 -0
  37. package/src/providers/gsc-auth.ts +93 -0
  38. package/src/{lib → providers}/gsc.ts +27 -24
  39. package/src/{lib/psi.ts → providers/pagespeed.ts} +3 -14
  40. package/src/shared/dates.ts +32 -0
  41. package/src/shared/http.ts +63 -0
  42. package/src/tools/ai.ts +19 -27
  43. package/src/tools/audit.ts +18 -98
  44. package/src/tools/observe.ts +28 -0
  45. package/src/tools/page/analyze.ts +194 -0
  46. package/src/tools/page/batch.ts +163 -0
  47. package/src/tools/page/contrast.ts +128 -0
  48. package/src/tools/page/links.ts +200 -0
  49. package/src/tools/page/metadata.ts +225 -0
  50. package/src/tools/page/structured-data.ts +288 -0
  51. package/src/tools/page/tool.ts +48 -0
  52. package/src/tools/search/actions.ts +56 -0
  53. package/src/tools/search/analytics.ts +266 -0
  54. package/src/tools/search/coverage.ts +247 -0
  55. package/src/tools/search/gaps.ts +129 -0
  56. package/src/tools/search/inspection.ts +160 -0
  57. package/src/tools/search/result.ts +3 -0
  58. package/src/tools/search/sample.ts +110 -0
  59. package/src/tools/search/schema.ts +62 -0
  60. package/src/{lib/sitemap.ts → tools/search/sitemap-sampling.ts} +1 -45
  61. package/src/tools/search/sites.ts +86 -0
  62. package/src/tools/search/tool.ts +11 -0
  63. package/src/tools/setup.ts +1 -1
  64. package/src/tools/speed/analyze.ts +191 -0
  65. package/src/tools/speed/batch.ts +265 -0
  66. package/src/tools/speed/crux.ts +176 -0
  67. package/src/tools/speed/pagespeed.ts +273 -0
  68. package/src/tools/speed/schema.ts +39 -0
  69. package/src/tools/speed/tool.ts +11 -0
  70. package/src/web/fetch.ts +31 -0
  71. package/src/web/images.ts +62 -0
  72. package/src/web/page-observation.ts +69 -0
  73. package/src/{lib → web}/robots.ts +20 -12
  74. package/src/web/sitemap-inventory.ts +64 -0
  75. package/src/web/sitemap-parser.ts +59 -0
  76. package/src/lib/auth.ts +0 -187
  77. package/src/tools/page.ts +0 -1241
  78. package/src/tools/search.ts +0 -1118
  79. package/src/tools/speed.ts +0 -956
package/README.md CHANGED
@@ -1,94 +1,32 @@
1
1
  # Pagesight
2
2
 
3
- [![npm version](https://img.shields.io/npm/v/pagesight.svg)](https://www.npmjs.com/package/pagesight)
3
+ SEO, analytics, page, and performance evidence for developers and AI assistants.
4
+ Use it through the TypeScript API, CLI, local HTTP server, or MCP.
4
5
 
5
- See your site the way search engines and AI see it.
6
+ Requires Bun. Try a page check without provider credentials:
6
7
 
7
- ```bash
8
- npm install pagesight
8
+ ```sh
9
+ bun add pagesight
10
+ bunx --bun pagesight page --url https://example.com
9
11
  ```
10
12
 
11
- Your AI assistant can write your code. Now it can see your site. Index status, performance, real-user metrics, search traffic, meta tags, structured data, AI crawler access, link health — one package, one call.
13
+ ```ts
14
+ import { execute } from "pagesight";
12
15
 
16
+ const result = await execute({ operation: "page", url: "https://example.com" });
17
+ console.log(result);
13
18
  ```
14
- === Site Audit: https://example.com ===
15
19
 
16
- 2 checks failed — results below are partial:
20
+ Provider reports need access to the relevant Google or Bing account. Results keep
21
+ raw observations, failures, and limits; missing data is not zero.
17
22
 
18
- FAIL PageSpeed: quota exceeded
19
- FAIL Sitemaps: permission denied
23
+ - [API, CLI, HTTP, and MCP](docs/usage.md)
24
+ - [Snapshots and comparisons](docs/snapshots.md)
25
+ - [Choose pages to investigate](docs/opportunities.md)
26
+ - [Provider access](docs/credentials.md)
27
+ - [Bing diagnostics, HTML images, and UI findings](docs/diagnostics.md)
20
28
 
21
- 5 findings:
29
+ MIT — see [LICENSE](LICENSE).
22
30
 
23
- HIGH Missing canonical URL
24
- HIGH 22 sitemap URLs submitted, 0 indexed
25
- Auto-inspected 5 URLs:
26
- - 2/5 indexed
27
- - 3/5 Discovered - currently not indexed: /docs/, /pricing/, /about/
28
- MEDIUM Missing og:image — no social preview image
29
- LOW Missing Twitter Card tags
30
- LOW 6/139 AI crawlers blocked
31
- ```
32
-
33
- ## Tools
34
-
35
- 6 tools organized by intent:
36
-
37
- | Tool | Intent | What it does |
38
- |------|--------|-------------|
39
- | `audit` | How's my site? | One-call site audit. Runs all checks in parallel. Prioritized findings with auto-drill-down on indexing issues. |
40
- | `page` | What's on this URL? | Meta tags, OG, Twitter Card, JSON-LD validation (19 schema types), internal link health, redirect chains, WCAG contrast checker. Batch mode for multiple URLs. |
41
- | `speed` | How fast is it? | PageSpeed single/batch/compare with Lighthouse scores and opportunities. CrUX real-user metrics (snapshot + history trends). |
42
- | `search` | How's Google seeing me? | URL inspection, sample-inspect from sitemaps, sitemap management, search analytics with period-over-period comparison. |
43
- | `ai` | How's AI seeing me? | AI crawler audit (139+ bots by category), robots.txt validation (RFC 9309), llms.txt detection, path access checks. |
44
- | `setup` | Auth config | Auth status check and OAuth setup flow. |
45
-
46
- ## Setup
47
-
48
- Add to your MCP config:
49
-
50
- ```json
51
- {
52
- "mcpServers": {
53
- "pagesight": {
54
- "command": "npx",
55
- "args": ["pagesight"],
56
- "env": {
57
- "GSC_CLIENT_ID": "your-client-id",
58
- "GSC_CLIENT_SECRET": "your-secret",
59
- "GSC_REFRESH_TOKEN": "your-token",
60
- "GOOGLE_API_KEY": "your-api-key"
61
- }
62
- }
63
- }
64
- }
65
- ```
66
-
67
- `page`, `speed`, and `ai` work without credentials. `search` and `audit` (for GSC checks) require OAuth or a service account.
68
-
69
- ### Full setup
70
-
71
- 1. [Google Cloud Console](https://console.cloud.google.com/) — enable Search Console API, PageSpeed Insights API, Chrome UX Report API
72
- 2. Create OAuth client ID (Desktop app) + API key
73
- 3. Configure:
74
-
75
- ```env
76
- GSC_CLIENT_ID=your-client-id.apps.googleusercontent.com
77
- GSC_CLIENT_SECRET=your-client-secret
78
- GSC_REFRESH_TOKEN=your-refresh-token
79
- GOOGLE_API_KEY=your-api-key
80
- ```
81
-
82
- ## Why Pagesight
83
-
84
- Every data point comes from a verifiable source. Google's APIs, real Chrome users, RFC 9309, schema.org, a community-maintained bot registry. No invented scores. No rules we can't cite.
85
-
86
- - **"Title must be under 60 characters"** — Gary Illyes: "an externally made-up metric."
87
- - **"Only one H1 per page"** — John Mueller: "You can use H1 tags as often as you want."
88
- - **"Minimum 300 words per page"** — Mueller: "not a quality factor."
89
-
90
- Pagesight reports what the sources report. Nothing more.
91
-
92
- ## License
93
-
94
- MIT
31
+ [Assess measurement quality and verify analytics](docs/measurement.md) with saved
32
+ snapshot summaries, GA Realtime and a repeatable browser-check workflow.
@@ -0,0 +1,54 @@
1
+ # Provider access
2
+
3
+ ## Credentials
4
+
5
+ | Provider | Configuration |
6
+ | -------------- | ------------------------------------------------------------------------------------------------------- |
7
+ | Search Console | `GSC_SERVICE_ACCOUNT_KEY`, or `GSC_CLIENT_ID`, `GSC_CLIENT_SECRET`, `GSC_REFRESH_TOKEN` |
8
+ | GA4 | `PAGESIGHT_GA_CREDENTIALS`, then `GOOGLE_APPLICATION_CREDENTIALS`, then the usual local gcloud ADC file |
9
+ | PageSpeed | `GOOGLE_API_KEY` optional |
10
+ | CrUX | `GOOGLE_API_KEY` required |
11
+
12
+ GA accepts service-account JSON or `authorized_user` ADC JSON. Use
13
+ `analytics.readonly` permission and grant property access. Enable both Analytics
14
+ Admin and Data APIs. GSC authentication remains separate with `webmasters.readonly`.
15
+ No new consent flow or provider settings are created by the read API.
16
+
17
+ Keep credentials in an environment file outside the repository and use Bun's
18
+ `--env-file` option. No tokens, API-key URLs, or provider error bodies are included
19
+ in report envelopes. Configuration files and operation requests must contain only
20
+ nonsecret identifiers and report parameters.
21
+
22
+ ## Bing Webmaster reports
23
+
24
+ Set `BING_WEBMASTER_API_KEY` from Bing Webmaster Tools API Access, then discover
25
+ sites before copying a verified `Url` verbatim into `bingSite` in your config
26
+ (including its scheme and trailing-slash form):
27
+
28
+ ```sh
29
+ pagesight discover --url https://example.com/ --providers bing
30
+ pagesight bing sites
31
+ pagesight bing queries --site https://example.com/
32
+ pagesight bing pages --site https://example.com/
33
+ pagesight bing traffic --site https://example.com/
34
+ ```
35
+
36
+ The API operations are `bing.sites`, `bing.queries`, `bing.pages` and `bing.traffic`.
37
+ MCP `observe` and HTTP use the same operation objects. A configured `bingSite`
38
+ adds all three reports to snapshots; doctor checks traffic access. Google discovery
39
+ remains the default; use `--providers gsc,ga,bing` to include all providers.
40
+
41
+ These methods accept no date range or pagination options. Snapshot requested dates
42
+ apply to Google reports; Bing returns its provider-defined range. Responses retain
43
+ Bing's raw `d` envelope, numeric fields and `/Date(...)/` strings; reporting timezone
44
+ and coverage remain unknown. Query/page statistics update weekly, traffic daily.
45
+ `GetPageStats` uses `Query` for the page URL. Since March 24, 2023 traffic includes
46
+ Web, Chat, News, Images, Videos and Knowledge Panel; this is not isolated AI-citation
47
+ evidence or a Google Web equivalent. Site verification is not indexing evidence.
48
+
49
+ Contracts were checked against Microsoft Learn's [GetUserSites](https://learn.microsoft.com/en-us/dotnet/api/microsoft.bing.webmaster.api.interfaces.iwebmasterapi.getusersites),
50
+ [GetQueryStats](https://learn.microsoft.com/en-us/dotnet/api/microsoft.bing.webmaster.api.interfaces.iwebmasterapi.getquerystats),
51
+ [GetPageStats](https://learn.microsoft.com/en-us/dotnet/api/microsoft.bing.webmaster.api.interfaces.iwebmasterapi.getpagestats)
52
+ and [GetRankAndTrafficStats](https://learn.microsoft.com/en-us/dotnet/api/microsoft.bing.webmaster.api.interfaces.iwebmasterapi.getrankandtrafficstats).
53
+ Fixture tests verify these contracts and safe failures. Live Bing access has not
54
+ been verified because no API key is configured in the development environment.
@@ -0,0 +1,85 @@
1
+ # Provider diagnostics and imported findings
2
+
3
+ ## Bing crawl, URL and backlink evidence
4
+
5
+ ```sh
6
+ pagesight bing crawl-stats --site https://example.com/
7
+ pagesight bing crawl-issues --site https://example.com/
8
+ pagesight bing url-info --site https://example.com/ --url https://example.com/page
9
+ pagesight bing link-counts --site https://example.com/ --max-pages 4
10
+ pagesight bing url-links --site https://example.com/ --url https://example.com/page --max-pages 4
11
+ ```
12
+
13
+ These read operations use `BING_WEBMASTER_API_KEY` and the registered site URL.
14
+ They expose Microsoft's `GetCrawlStats`, `GetCrawlIssues`, `GetUrlInfo`,
15
+ `GetLinkCounts` and `GetUrlLinks` JSON methods. Link reports start at page zero;
16
+ `maxPages` defaults to one and accepts 1–20. Hitting the limit returns partial
17
+ status (CLI exit 3); raise the limit to repeat from page zero. `nextOffset` is a
18
+ provider page number. Failed later requests preserve successful pages. Changing
19
+ `TotalPages` stops pagination with partial coverage.
20
+
21
+ Raw provider dates and counts stay intact. Do not derive daily error rates without
22
+ verifying their periods and units. `HttpStatus: 0` in URL metadata is not an HTTP
23
+ success. Link report exhaustion does not establish a complete backlink inventory;
24
+ empty rows are not zero links, and link counts do not measure domain quality.
25
+ `InIndex` and sitemap counts have different scopes. Crawl issues are not the UI
26
+ recommendations list. These operations are explicit reads; existing snapshot
27
+ provider selection and call counts are unchanged.
28
+
29
+ ## HTML image evidence
30
+
31
+ `pagesight page --url https://example.com/` now includes `descriptionLength`
32
+ (Unicode code points, or null) and `imageEvidence`. Images include raw `src`,
33
+ `altPresent`, `altText`, `inNoscript`, width/height, role and aria-hidden attributes.
34
+ Missing ALT has `altPresent: false, altText: null`; an explicitly empty ALT is
35
+ `altPresent: true, altText: ""`. Empty ALT can be appropriate for decorative images.
36
+ No description-length threshold or image-purpose verdict is imposed.
37
+
38
+ The inventory includes noscript fallback markup, retains at most 200 images,
39
+ and reports `omittedImages` and `truncated`. Fallback images follow ordinary images
40
+ in output; this is not DOM order. No scripts run or image URLs are requested.
41
+ The existing 2 MB page limit still applies. This is markup evidence, not a full
42
+ accessibility or rendered-page audit.
43
+
44
+ ## Import provider UI findings
45
+
46
+ Use a small JSON transcription when a UI report has no verified API equivalent:
47
+
48
+ ```json
49
+ {
50
+ "provider": "bing",
51
+ "site": "https://example.com/",
52
+ "source": {
53
+ "kind": "csv",
54
+ "label": "Affected URLs export; rule copied from report screen",
55
+ "capturedAt": null,
56
+ "scannedAt": null,
57
+ "coverage": "Only affected URLs were exported; total site coverage unknown"
58
+ },
59
+ "findings": [
60
+ {
61
+ "rule": "Meta descriptions are too short",
62
+ "severity": "unknown",
63
+ "urls": ["https://example.com/explore"]
64
+ }
65
+ ]
66
+ }
67
+ ```
68
+
69
+ ```sh
70
+ pagesight evidence import --request findings.json --out imported.json
71
+ ```
72
+
73
+ This validates JSON, preserves its attribution and labels it `user-import` and
74
+ `unverified`. It does not parse screenshots/CSV automatically, contact the provider,
75
+ fetch listed URLs or verify the finding. Use null for unknown capture/scan dates;
76
+ known dates must be ISO timestamps with offsets. Do not infer severity from a
77
+ URL-only CSV. Sources can be `csv`, `screenshot` or `manual`; providers can be
78
+ `bing`, `gsc` or `other`. An optional `source.artifactSha256` is caller supplied,
79
+ not verified. The generated `normalizedDocumentSha256` identifies normalized JSON,
80
+ not original file bytes. Import at most 100 findings and 1,000 URL references.
81
+
82
+ All additions use the shared API: operation names are `bing.crawl-stats`,
83
+ `bing.crawl-issues`, `bing.url-info`, `bing.link-counts`, `bing.url-links`, and
84
+ `evidence.import` (with a `document` field). They are also available through HTTP
85
+ `POST /v1/query`, CLI `api --request`, and MCP `observe`.
@@ -0,0 +1,178 @@
1
+ # Understand and verify analytics
2
+
3
+ Start by checking whether the measurements answer the site's objective. A working
4
+ Google connection and a high key-event count do not establish successful visits.
5
+
6
+ ## Read a snapshot
7
+
8
+ Collect evidence, then produce a readable assessment:
9
+
10
+ ```sh
11
+ pagesight snapshot --config seo.config.json --out observations/baseline.json
12
+ pagesight assess --snapshot observations/baseline.json --format text --max-rows 5
13
+ pagesight assess --snapshot observations/baseline.json --out observations/assessment.json
14
+ ```
15
+
16
+ `assess` makes no provider calls. It reads a version-1 saved snapshot and returns
17
+ structured findings, selected report rows, scope, source observation names and
18
+ limits. The optional text format renders those same facts. Its snapshot hash
19
+ identifies normalized JSON, not original file bytes or verified provenance.
20
+ Older snapshots without versioned, named observations must be recollected.
21
+
22
+ The assessment separates GA-configured key events, observed event counts,
23
+ caller-designated `context.successEvents`, and `context.excludedKeyEvents`.
24
+ Designation is not independent validation. Exclusions label rows; they never
25
+ remove raw evidence. An event's name or ratio of key events to events does not
26
+ establish its trigger, business meaning, or whether tracking is duplicated.
27
+
28
+ The report tables include GSC property/page/query evidence and GA hostname,
29
+ production channel/event, and organic landing/event evidence when available.
30
+ Each table shows its original dimension and metric names. Rows are ordered by
31
+ the first metric, then row keys; only the requested number are displayed. They
32
+ are observed rows, not exhaustive search rankings. No totals are inferred by
33
+ summing session, user, page or query rows. Bing and page observations retain
34
+ provider status but have no generated performance conclusions in this version.
35
+
36
+ A missing or unusable selected report produces an explicit unknown. Unrelated
37
+ hostnames in the census do not invalidate correctly production-filtered reports;
38
+ filtering to production also does not identify the owner's visits there. Supplied
39
+ snapshot claims are not authenticated again. A successful assessment is not a
40
+ certificate of healthy tracking.
41
+
42
+ ## Check recent activity
43
+
44
+ Save this request as `realtime.json`:
45
+
46
+ ```json
47
+ {
48
+ "dimensions": [{ "name": "eventName" }],
49
+ "metrics": [{ "name": "eventCount" }],
50
+ "minuteRanges": [{ "startMinutesAgo": 29, "endMinutesAgo": 0 }],
51
+ "limit": 100
52
+ }
53
+ ```
54
+
55
+ ```sh
56
+ pagesight ga realtime --property 123456 --request realtime.json --out observations/realtime.json
57
+ ```
58
+
59
+ The equivalent API request is
60
+ `{ operation: "ga.realtime", property: "123456", request: ... }`. HTTP and MCP
61
+ `observe` accept it too. It uses the existing read-only GA credentials.
62
+
63
+ Realtime is a moving window, normally the last 30 minutes. Analytics 360 permits
64
+ up to 60 minutes; ranges beyond 29 minutes ago are left for Google to authorize.
65
+ Two ranges may overlap and count the same event in both. Supported dimensions
66
+ and metrics differ from historical reports; Google validates requested fields.
67
+ No production hostname filter is automatically added. Do not assume historical
68
+ filters such as `hostName` are supported in Realtime.
69
+
70
+ The API has no offset or page token. When `rowCount` exceeds returned rows,
71
+ Pagesight marks the result partial with `nextOffset: null`; narrow the query or
72
+ raise `limit` (maximum 250,000). Empty results do not prove collection failed.
73
+ Realtime is deliberately excluded from period snapshots and comparisons.
74
+
75
+ Historical `ga.report` results warn when collected fewer than three property
76
+ calendar days after any requested end date. Unknown timezone means freshness
77
+ cannot be assessed. This is a precaution, not a promise that older data is final:
78
+ Google describes typical processing of 24–48 hours, possible late arrivals and
79
+ later attribution changes. Read collection times and warnings alongside counts.
80
+
81
+ ## Verify a real user flow
82
+
83
+ Keep separate evidence for each stage:
84
+
85
+ | Stage | Evidence | What it establishes |
86
+ | ------------------------- | --------------------------------------------------------------------- | ------------------------------------------------------------------------------ |
87
+ | Provider access | `doctor` | Credentials can read the selected properties |
88
+ | Browser emission | DevTools network capture | The browser attempted to send an event |
89
+ | Endpoint response | Response to that exact request | The endpoint responded; event acceptance or processing is not established |
90
+ | Recent reported activity | `ga realtime` | Matching aggregate property activity appeared; not attribution to your test |
91
+ | Stored reported activity | `ga report` | Rows exist for the requested dates and filters; recent data may still change |
92
+ | Validated product outcome | Observed user action plus event definition and corroborating evidence | The event represents the site's intended useful action within the tested scope |
93
+
94
+ 1. Define a small case with its expected events: for example, open an explorer,
95
+ change one brand filter, open a result, then view its history. Include the
96
+ expected page URLs, event names and counts. State the browser, observation
97
+ window and whether the test is on production. Production tests generate traffic.
98
+ 2. Run `doctor`, then capture a Realtime report as context. Perform the case in
99
+ an isolated browser session while recording network requests and responses.
100
+ Capture all relevant collection destinations, not just one assumed hostname.
101
+ 3. Inspect event names, destination measurement ID, URL and duplicate requests.
102
+ Distinguish one action from automatic history events, retries and unrelated
103
+ requests. Record the actual results; do not infer them from aggregate ratios.
104
+ 4. Query Realtime again and a historical report with explicit dates and production
105
+ filters after allowing processing time. Other visitors may contribute to both.
106
+ Do not manufacture event names or claim session-level attribution from totals.
107
+ 5. Save a small attributed finding using the existing import format. Raw captures
108
+ may contain cookies, client/session IDs, tokens or sensitive URL parameters;
109
+ keep them private and transcribe only the fields needed to explain the result.
110
+ 6. Repeat the same case after a tracking change. Record both evidence sets and
111
+ the deployed version. Pagesight itself does not edit tags or provider settings.
112
+
113
+ Example `verification.json` (illustrative, not a completed test):
114
+
115
+ ```json
116
+ {
117
+ "provider": "ga",
118
+ "site": "https://example.com/",
119
+ "source": {
120
+ "kind": "manual",
121
+ "label": "Explorer brand-change browser check",
122
+ "capturedAt": null,
123
+ "scannedAt": null,
124
+ "coverage": "One browser session and one filter change; timestamps not supplied"
125
+ },
126
+ "findings": [
127
+ {
128
+ "rule": "One filter change attempted two page_view requests",
129
+ "severity": "unknown",
130
+ "urls": ["https://example.com/explore"],
131
+ "notes": "Expected one request. Both requests received HTTP 204. Their presence in aggregate GA reports does not identify this test session. See the private capture for timing and destination."
132
+ }
133
+ ]
134
+ }
135
+ ```
136
+
137
+ ```sh
138
+ pagesight evidence import --request verification.json --out observations/verification.json
139
+ ```
140
+
141
+ The import remains `user-import` with `verification: "unverified"`: Pagesight
142
+ stores the attribution but does not replay the flow, fetch a capture or verify
143
+ its claims. It is separate from `assess`; imports cannot silently upgrade snapshot
144
+ facts or turn an unverified assertion into a provider result.
145
+
146
+ References: [Realtime REST API](https://developers.google.com/analytics/devguides/reporting/data/v1/rest/v1beta/properties/runRealtimeReport),
147
+ [Realtime dimensions and metrics](https://developers.google.com/analytics/devguides/reporting/data/v1/realtime-api-schema),
148
+ [GA data freshness](https://support.google.com/analytics/answer/11198161).
149
+
150
+ ## Connect organic landings to observed actions
151
+
152
+ Snapshots now include `ga.report.landingPagePlusQueryString+sessionSource+eventName.organic`:
153
+ raw landing path/query string, session source and event name with `eventCount`,
154
+ filtered to the configured production hostname and `Organic Search` sessions.
155
+ It uses the same date window and bounded pagination as other snapshot reports.
156
+ `assess --format text` exposes the rows and scope; `compare` compares common rows
157
+ in compatible saved snapshots. Older snapshots lack this report: assessment marks
158
+ it unavailable and comparison retains absence as unknown, never zero.
159
+
160
+ An agent can inspect a meaningful event such as `price_detail_view` alongside
161
+ `ga.report.landingPagePlusQueryString+sessionSource.organic`, which retains landing
162
+ traffic. The event table associates occurrences with the session's first pageview,
163
+ not necessarily the page where the event happened. Repeat occurrences are possible;
164
+ these counts are not unique sessions, a funnel, or a conversion rate. An event
165
+ name does not establish its business meaning. Keep instrumentation/deployment dates
166
+ in the investigation record before interpreting before/after changes.
167
+
168
+ Use raw GA landing paths as evidence. Query strings, `(not set)` and `(other)` remain
169
+ visible. Do not automatically join them to Search Console's canonical page URLs,
170
+ reconcile GSC clicks with GA sessions, or infer SEO causality. High-cardinality
171
+ landing/event combinations may be partial, sampled, thresholded or aggregated;
172
+ check pagination and metadata. Default assessment row caps may hide the event of
173
+ interest: increase `--max-rows`, inspect the saved report, or use `ga report` with
174
+ an explicit event filter. Empty or missing event rows, especially shortly after
175
+ instrumentation, do not prove zero activity or broken collection.
176
+
177
+ Source: [Google's Data API schema](https://developers.google.com/analytics/devguides/reporting/data/v1/api-schema)
178
+ defines landing page as the first pageview in a session and event count as occurrences.
@@ -0,0 +1,78 @@
1
+ # Choose pages to investigate
2
+
3
+ `opportunities` turns a saved snapshot into an investigation cohort: pages with
4
+ at least the requested impressions and no more than the requested clicks. It
5
+ retains search metrics, organic landing/event associations and available technical
6
+ observations, then names unknowns and next checks. It makes no new requests.
7
+
8
+ ```sh
9
+ pagesight snapshot --config seo.config.json --out before.json
10
+ pagesight opportunities --snapshot before.json --min-impressions 20 --max-clicks 2 --max-rows 10 --format text
11
+ ```
12
+
13
+ The same `opportunities` operation accepts `snapshot`, `minImpressions`, `maxClicks`
14
+ and `maxRows` through the TypeScript API, local HTTP API and MCP `observe`. JSON
15
+ is the CLI default; `--out` retains either format. Invalid inputs fail; missing or
16
+ incomplete provider evidence produces partial output, preserving available facts.
17
+
18
+ ## Selection is a policy, not a diagnosis
19
+
20
+ Defaults are 20 impressions, at most 2 clicks, and 10 displayed candidates.
21
+ These are caller-adjustable investigation cutoffs, not universal CTR benchmarks.
22
+ Selection uses all retained validated GSC page rows; ordering is descending
23
+ impressions, ascending clicks, then URL. Output gives observed, qualifying and
24
+ omitted counts. Missing/unusable search evidence is unknown, not zero candidates.
25
+ Pagination completion does not prove exhaustive Search Console coverage.
26
+
27
+ Each candidate keeps impressions, clicks, CTR and average position. A page with
28
+ 45 impressions, zero clicks and position 13 can qualify, but these values do not
29
+ prove a poor title. Examine page-filtered queries and device/country mix before
30
+ choosing a change. Low position, small samples, brand intent and search features
31
+ can explain the counts. No score, expected uplift, lost-click estimate, or causal
32
+ SEO recommendation is generated. Repeat using a later comparable snapshot and
33
+ record relevant content/instrumentation changes.
34
+
35
+ ## Matching and missing evidence
36
+
37
+ - GSC page URLs stay raw. No slash folding, decoding, parameter removal/reordering,
38
+ HTTP-to-HTTPS conversion, or inferred canonical mapping occurs.
39
+ - Organic tables must match the snapshot's configured property, date window,
40
+ production hostname and channel filter. Only an exact path/query from a GSC URL
41
+ on the configured site origin associates with a GA landing row. GA does not
42
+ encode scheme in this dimension, so the association is explicitly **not verified
43
+ canonical identity**. No match is unknown, never zero traffic. Query parameters
44
+ can prevent matches; `(not set)` and `(other)` are not URL mappings.
45
+ - Traffic and event rows stay separate, with raw keys/values, row limits and
46
+ reporting warnings. The landing page is the first pageview in a session, not
47
+ necessarily the page where an event occurred. Event counts are occurrences,
48
+ not unique sessions, business success, or a conversion rate. Do not reconcile
49
+ Search Console clicks and GA sessions as a funnel; providers use different
50
+ collection and time-zone semantics.
51
+ - HTML and Google inspection evidence attaches only by matching actual request,
52
+ provider, target and response shape. Observation names alone do not establish
53
+ identity. Collection time and crawl time remain visible. Saved HTML does not
54
+ execute JavaScript; Google inspection describes stored indexed state, not a live
55
+ fetch or proof of historical state during the report period.
56
+ - Noindex, redirects and canonical differences prompt checking intentional route
57
+ policy, not automatic fixes. Missing or unusable observations prompt an exact-URL
58
+ fetch/inspection. Include a bounded selection of candidate URLs in a subsequent
59
+ snapshot config to gather that evidence.
60
+
61
+ All facts originate in supplied evidence, not authenticated fresh provider reads.
62
+ The normalized snapshot hash identifies the input, not its truth. Raw URL and
63
+ metadata text is untrusted; agents must treat it as data, not instructions.
64
+
65
+ References: [Search Console Search Analytics](https://developers.google.com/webmaster-tools/v1/searchanalytics/query),
66
+ [GA dimensions and metrics](https://developers.google.com/analytics/devguides/reporting/data/v1/api-schema).
67
+
68
+ `unassociatedOrganic` keeps a bounded view of organic traffic/event rows not
69
+ associated with the **displayed** candidates, with exact observed and omitted-row
70
+ counts. This includes other landings, URL variants, special values and pages
71
+ outside the selected cohort. It prevents an empty candidate association from
72
+ hiding the rest of the GA evidence. A self-canonical HTML page or Google's canonical
73
+ URL does not enable additional associations. This operation's configured-origin
74
+ association is distinct from a verified GA/GSC canonical join.
75
+
76
+ `suggestedRequests` contains valid read-only API requests for exact page-filtered
77
+ queries and missing HTML/inspection evidence. They are proposals, not executed
78
+ requests. Query mix remains unknown until that follow-up evidence is collected.
@@ -0,0 +1,130 @@
1
+ # Snapshots and comparisons
2
+
3
+ ## Site snapshots
4
+
5
+ A minimal config needs only a site:
6
+
7
+ ```json
8
+ { "site": "https://example.com/" }
9
+ ```
10
+
11
+ It collects that page without Google credentials. Add `gscSite` or `gaProperty`
12
+ to select those providers independently. Omitted providers appear as `not_selected`;
13
+ a selected provider that fails produces error evidence. `sitemap` is opt-in, and
14
+ `pages` defaults to the site URL. `productionHostname` defaults to its hostname.
15
+ Existing full configurations still work. Context defaults identify unspecified
16
+ objectives, locale and country rather than inferring them.
17
+
18
+ `discover` lists candidates from the selected Google providers and returns a usable
19
+ site-only config in `pages[0].response.config`. Copy that object to a config file,
20
+ then add the property IDs you verified. It does not automatically select properties:
21
+ GA account display names are not proof of hostname ownership. Discovery failures
22
+ remain independent; inspect raw responses and `nextPageToken` for incomplete lists.
23
+
24
+ A config holds nonsecret provider IDs and the site's meaning:
25
+
26
+ ```json
27
+ {
28
+ "site": "https://example.com/",
29
+ "productionHostname": "example.com",
30
+ "gscSite": "sc-domain:example.com",
31
+ "gaProperty": "123456",
32
+ "sitemap": "https://example.com/sitemap.xml",
33
+ "pages": ["https://example.com/"],
34
+ "context": {
35
+ "objective": "Help visitors use the product",
36
+ "successEvents": [],
37
+ "excludedKeyEvents": [],
38
+ "locale": "en-US",
39
+ "country": "US",
40
+ "routes": [{ "pattern": "/", "purpose": "Public entry", "indexing": "index" }],
41
+ "measurementCaveats": []
42
+ }
43
+ }
44
+ ```
45
+
46
+ CLI snapshots default to 28 days ending Pacific today minus three days. Supply both
47
+ `--start YYYY-MM-DD` and `--end YYYY-MM-DD` to reproduce another interval. The API
48
+ requires explicit dates. The country and locale are context, not implicit filters.
49
+
50
+ A snapshot collects independent GSC property/page/query/date reports and sitemaps;
51
+ GA property/key-event configuration, unfiltered hostname census, production channels,
52
+ named events and organic landing reports; configured page fetches and GSC inspections;
53
+ and a sitemap inventory bounded to five same-origin XML files and 8 MB.
54
+ This inventory accepts ordinary unprefixed sitemap XML; unsupported prefixed or
55
+ compressed documents are reported as incomplete rather than empty coverage.
56
+
57
+ Snapshots embed the config used, observations, content hashes, requested/observed
58
+ dates and unknown deployment/reference identities. A failed observation remains
59
+ visible while successful data is retained. `doctor` probes GSC, GA Admin, GA Data
60
+ when selected, plus live HTML. Run speed operations to probe PSI/CrUX separately.
61
+ Snapshot responses carry `snapshotVersion: 1` and stable observation `name` values,
62
+ also present in the summary, so identity does not depend on array position.
63
+
64
+ Interpretation rules:
65
+
66
+ - GSC query privacy exclusions and aggregation differences mean row sums are not
67
+ property totals. Missing rows are not zero. Comparisons are descriptive, not causal.
68
+ - GA key events must be interpreted by name and the site's selected objective.
69
+ A hostname filter excludes development hosts but not internal use of production.
70
+ - Sitemap membership is not indexing. Google's `contents[].indexed` field is deprecated.
71
+ URL inspections describe the selected URLs in Google's stored state, not live access
72
+ or a statistically representative whole-site coverage rate.
73
+ - PSI is a lab run. CrUX NOT_FOUND is no record for that scope, not zero performance.
74
+ CrUX history periods overlap. AI referrals do not establish citations.
75
+ - Keep raw landing query strings and observed canonicals. Functional URL parameters
76
+ need site-specific interpretation; Pagesight does not silently join GA to GSC.
77
+
78
+ ## Compare saved snapshots
79
+
80
+ Capture two snapshots with the same config and equal-length, nonoverlapping report
81
+ periods, then compare them locally:
82
+
83
+ ```sh
84
+ pagesight snapshot --config seo.config.json --start 2026-07-04 --end 2026-07-31 --out before.json
85
+ pagesight snapshot --config seo.config.json --start 2026-08-01 --end 2026-08-28 --out after.json
86
+ pagesight compare --baseline before.json --current after.json --max-rows 100
87
+ ```
88
+
89
+ The shared operation is `{ operation: "compare", baseline, current, maxRows: 100 }`,
90
+ where baseline and current are parsed snapshot evidence objects. Only the CLI reads
91
+ file paths. HTTP accepts objects up to 32 MB per request; larger pairs can use the
92
+ local API or CLI. MCP `observe` accepts the same object. `evidenceSchema` and
93
+ `snapshotEvidenceSchema` are exported for callers validating stored reports.
94
+
95
+ This first comparison supports GSC reports with non-time row keys and GA reports
96
+ using snapshot dimensions: `hostName`, `sessionDefaultChannelGroup`,
97
+ `sessionSourceMedium`, `eventName`, `landingPagePlusQueryString`, and `sessionSource`.
98
+ Reports without dimensions are also supported. Other GA dimensions and
99
+ `dimensionExpression` aliases require a separate comparison policy. It checks
100
+ snapshot format version 1, unique observation names, site, property, dimensions,
101
+ metrics, filters, aggregation, report periods and GA timezone/currency/metric types.
102
+ GSC data must be finalized. Time dimensions, Bing's provider-defined windows, HTML,
103
+ sitemap and provider metadata observations are explicitly unsupported for comparison.
104
+ Recapture snapshots created before versioned, named observations were introduced.
105
+
106
+ Each observation is `compared`, `limited`, `incompatible`, `unavailable`, or
107
+ `unsupported`. The outer evidence reports whether the comparison operation ran;
108
+ inspect the response status-count summary and per-observation statuses before using deltas. Partial pagination,
109
+ sampling, thresholding and high-cardinality aggregation remain visible limitations.
110
+ Missing trailing GSC date rows and recently collected GA data also mark comparisons
111
+ as limited; observed dates cannot prove complete coverage.
112
+ Only keys observed in both periods get numeric deltas. Keys seen in one period stay
113
+ unknown in the other, and their bounded lists include full counts. No row sums or
114
+ site-wide extrapolations are generated. Each row retains the original numeric values,
115
+ including GA strings; invalid or unsafe numbers get a null delta. Percent change is
116
+ null when the baseline is zero. CTR and position remain in their provider units.
117
+
118
+ `maxRows` defaults to 100 and is capped at 1,000 per observation/list; omitted counts
119
+ are explicit. The comparison includes `canonicalSha256` hashes of validated source objects, with
120
+ object keys sorted lexically and array order retained. These identify comparison
121
+ inputs, not raw file bytes; whitespace changes in saved JSON do not change them. Keep source snapshots for their full requests,
122
+ responses and metadata. Changes are descriptive and do not establish that an SEO
123
+ edit caused traffic changes; Pagesight does not apply SEO edits automatically.
124
+
125
+ ## Read the evidence
126
+
127
+ Run `pagesight assess --snapshot saved.json --format text` for a deterministic
128
+ summary of report rows, measurement findings and unknowns. It retains references
129
+ to the saved observations and makes no new provider calls. See
130
+ [measurement and verification](measurement.md) for scope and examples.