mcp-scraper 0.38.2 → 0.40.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (111) hide show
  1. package/README.md +5 -2
  2. package/package.json +5 -6
  3. package/dist/bin/api-server.cjs +0 -58752
  4. package/dist/bin/api-server.cjs.map +0 -1
  5. package/dist/bin/api-server.d.cts +0 -1
  6. package/dist/bin/api-server.d.ts +0 -1
  7. package/dist/bin/api-server.js +0 -38
  8. package/dist/bin/api-server.js.map +0 -1
  9. package/dist/bin/mcp-scraper-cli.cjs +0 -2671
  10. package/dist/bin/mcp-scraper-cli.cjs.map +0 -1
  11. package/dist/bin/mcp-scraper-cli.d.cts +0 -1
  12. package/dist/bin/mcp-scraper-cli.d.ts +0 -1
  13. package/dist/bin/mcp-scraper-cli.js +0 -742
  14. package/dist/bin/mcp-scraper-cli.js.map +0 -1
  15. package/dist/bin/mcp-scraper-install.cjs +0 -129
  16. package/dist/bin/mcp-scraper-install.cjs.map +0 -1
  17. package/dist/bin/mcp-scraper-install.d.cts +0 -1
  18. package/dist/bin/mcp-scraper-install.d.ts +0 -1
  19. package/dist/bin/mcp-scraper-install.js +0 -27
  20. package/dist/bin/mcp-scraper-install.js.map +0 -1
  21. package/dist/bin/mcp-stdio-server.cjs +0 -12264
  22. package/dist/bin/mcp-stdio-server.cjs.map +0 -1
  23. package/dist/bin/mcp-stdio-server.d.cts +0 -1
  24. package/dist/bin/mcp-stdio-server.d.ts +0 -1
  25. package/dist/bin/mcp-stdio-server.js +0 -135
  26. package/dist/bin/mcp-stdio-server.js.map +0 -1
  27. package/dist/bin/paa-harvest.cjs +0 -3808
  28. package/dist/bin/paa-harvest.cjs.map +0 -1
  29. package/dist/bin/paa-harvest.d.cts +0 -1
  30. package/dist/bin/paa-harvest.d.ts +0 -1
  31. package/dist/bin/paa-harvest.js +0 -44
  32. package/dist/bin/paa-harvest.js.map +0 -1
  33. package/dist/chunk-345BQXZH.js +0 -712
  34. package/dist/chunk-345BQXZH.js.map +0 -1
  35. package/dist/chunk-44HZLHDV.js +0 -52
  36. package/dist/chunk-44HZLHDV.js.map +0 -1
  37. package/dist/chunk-AZRPG43B.js +0 -617
  38. package/dist/chunk-AZRPG43B.js.map +0 -1
  39. package/dist/chunk-CB5C3BPB.js +0 -135
  40. package/dist/chunk-CB5C3BPB.js.map +0 -1
  41. package/dist/chunk-EGJKUB4Q.js +0 -276
  42. package/dist/chunk-EGJKUB4Q.js.map +0 -1
  43. package/dist/chunk-FQI5PFE7.js +0 -1866
  44. package/dist/chunk-FQI5PFE7.js.map +0 -1
  45. package/dist/chunk-FRYT3ID4.js +0 -684
  46. package/dist/chunk-FRYT3ID4.js.map +0 -1
  47. package/dist/chunk-G3P3ZDB4.js +0 -69
  48. package/dist/chunk-G3P3ZDB4.js.map +0 -1
  49. package/dist/chunk-K443GQY5.js +0 -24
  50. package/dist/chunk-K443GQY5.js.map +0 -1
  51. package/dist/chunk-N7KUTTCC.js +0 -3007
  52. package/dist/chunk-N7KUTTCC.js.map +0 -1
  53. package/dist/chunk-NGM237OO.js +0 -3410
  54. package/dist/chunk-NGM237OO.js.map +0 -1
  55. package/dist/chunk-NKCCGADE.js +0 -11285
  56. package/dist/chunk-NKCCGADE.js.map +0 -1
  57. package/dist/chunk-NNW3O6ZD.js +0 -108
  58. package/dist/chunk-NNW3O6ZD.js.map +0 -1
  59. package/dist/chunk-QZXKQB7Y.js +0 -414
  60. package/dist/chunk-QZXKQB7Y.js.map +0 -1
  61. package/dist/chunk-SFRMFGQ6.js +0 -158
  62. package/dist/chunk-SFRMFGQ6.js.map +0 -1
  63. package/dist/chunk-YCI2PNCS.js +0 -499
  64. package/dist/chunk-YCI2PNCS.js.map +0 -1
  65. package/dist/chunk-YODBNTTN.js +0 -7
  66. package/dist/chunk-YODBNTTN.js.map +0 -1
  67. package/dist/db-C5KVCOYT.js +0 -239
  68. package/dist/db-C5KVCOYT.js.map +0 -1
  69. package/dist/extract-bundle-KUBX6N6Z.js +0 -568
  70. package/dist/extract-bundle-KUBX6N6Z.js.map +0 -1
  71. package/dist/index.cjs +0 -4160
  72. package/dist/index.cjs.map +0 -1
  73. package/dist/index.d.cts +0 -413
  74. package/dist/index.d.ts +0 -413
  75. package/dist/index.js +0 -338
  76. package/dist/index.js.map +0 -1
  77. package/dist/location-data-repository-TTWF3OTM.js +0 -35
  78. package/dist/location-data-repository-TTWF3OTM.js.map +0 -1
  79. package/dist/server-5EX6XBIA.js +0 -33596
  80. package/dist/server-5EX6XBIA.js.map +0 -1
  81. package/dist/site-extract-repository-XSPJTCIL.js +0 -62
  82. package/dist/site-extract-repository-XSPJTCIL.js.map +0 -1
  83. package/dist/worker-XCPU4YSN.js +0 -142
  84. package/dist/worker-XCPU4YSN.js.map +0 -1
  85. package/docs/adr/0001-in-page-graphql-interception-for-anti-bot-scraping.md +0 -58
  86. package/docs/adr/0002-hybrid-smart-rag-vault-retrieval.md +0 -62
  87. package/docs/adr/0003-waive-unrecoverable-scheduled-model-cost.md +0 -22
  88. package/docs/adr/README.md +0 -13
  89. package/docs/final-tooling-spec.md +0 -206
  90. package/docs/hosted-location-data.md +0 -108
  91. package/docs/kernel-proxy-future-enhancements.md +0 -80
  92. package/docs/mcp-tool-craft-lint.generated.md +0 -183
  93. package/docs/mcp-tool-design-guide.md +0 -225
  94. package/docs/mcp-tool-manifest.generated.json +0 -22871
  95. package/docs/mcp-tool-quality-spec.md +0 -240
  96. package/docs/oauth-legal-review.md +0 -38
  97. package/docs/seo-crawl-report-spec.md +0 -287
  98. package/docs/specs/api-forge-spec.md +0 -234
  99. package/docs/specs/connected-services-control-plane-decoupling-spec.md +0 -1044
  100. package/docs/specs/deferred-work-spec.md +0 -86
  101. package/docs/specs/google-drive-bulk-access-and-mcp-schema-passthrough-spec.md +0 -1689
  102. package/docs/specs/kernel-stealth-captcha-test-matrix.md +0 -278
  103. package/docs/specs/main-mcp-integration-ownership-spec.md +0 -1164
  104. package/docs/specs/mcp-tool-definition-quality-audit-spec.md +0 -1602
  105. package/docs/specs/meta-ad-creative-media-resolution-spec.md +0 -31
  106. package/docs/specs/multimodal-image-memory-architecture-spec.md +0 -1022
  107. package/docs/specs/oauth-mcp-spec.md +0 -213
  108. package/docs/specs/query-fanout-transport-contract-fix.md +0 -45
  109. package/docs/specs/relationship-workspace-ai-behavior-plan.md +0 -26
  110. package/docs/specs/unified-credit-and-scheduled-execution-billing-spec.md +0 -995
  111. package/docs/tool-catalog-spec.md +0 -388
@@ -1,240 +0,0 @@
1
- # MCP Tool Quality Spec
2
-
3
- This spec defines the shipping bar for MCP Scraper tools. It exists because MCP behavior is model-facing: a tool can be technically callable and still fail if the AI cannot infer when to use it, how to fill inputs, or how to chain the result.
4
-
5
- ## What Actually Steers The AI
6
-
7
- The model is primarily affected by what the MCP client receives from `tools/list` and `tools/call`.
8
-
9
- 1. Tool name
10
- 2. Tool title and annotations
11
- 3. Tool description
12
- 4. Input schema field names, descriptions, defaults, enums, and limits
13
- 5. Output schema and `structuredContent`
14
- 6. Tool result text and next-step guidance
15
- 7. Error shape and retry guidance
16
-
17
- README files, website copy, and public skill markdown help humans and skill loaders, but they do not reliably affect runtime AI behavior unless the client explicitly injects those files into context.
18
-
19
- ## Tool Boundary Rules
20
-
21
- Each tool must have one primary job.
22
-
23
- - Split tools when the user intent, cost model, result shape, or follow-up workflow differs.
24
- - Do not overload one tool with unrelated modes if a model could choose the wrong path.
25
- - Descriptions for adjacent tools must explicitly say when not to use each tool.
26
- - If a workflow has a natural sequence, encode it in the description and result guidance.
27
-
28
- Example:
29
-
30
- - `maps_search`: find multiple Google Maps candidates for a category, market, lead list, or "more than the 3-pack".
31
- - `maps_place_intel`: hydrate one known business or one selected candidate with full profile details and optional reviews.
32
-
33
- ## Tool Naming
34
-
35
- Names must be short, stable, and action-oriented.
36
-
37
- - Use domain + action: `maps_search`, `maps_place_intel`, `facebook_ad_search`.
38
- - Avoid generic names such as `search`, `lookup`, `get_data`, or `run`.
39
- - Avoid names that imply broader capability than the tool has.
40
- - Do not rename a public tool without a compatibility plan.
41
-
42
- ## Tool Descriptions
43
-
44
- Descriptions are model instructions. They must be concise but operational.
45
-
46
- Every description must include:
47
-
48
- - What the tool does.
49
- - When to use it.
50
- - When not to use it if there is an adjacent tool.
51
- - Important defaults and hard caps.
52
- - Cost-sensitive behavior when relevant.
53
- - The expected next tool when chaining is common.
54
- - Whether reports are saved locally.
55
-
56
- Bad:
57
-
58
- ```text
59
- Search Google Maps.
60
- ```
61
-
62
- Good:
63
-
64
- ```text
65
- Search Google Maps for multiple businesses/profiles by category, niche, keyword, or local market. Use this when the user asks for several Google Business Profiles, GMBs, GBPs, leads, prospects, competitors, or "more than the 3-pack." Returns up to 50 candidates. Default maxResults is 10; maximum is 50. Use maps_place_intel afterward only when a selected business needs full details and reviews.
66
- ```
67
-
68
- ## Input Schema
69
-
70
- Each input field must have a clear description.
71
-
72
- Required:
73
-
74
- - Required fields must be actually required in the schema.
75
- - Defaults must be encoded in the schema, not only in prose.
76
- - Hard caps must be encoded in the schema.
77
- - Fields that are often confused must say what not to put there.
78
- - Location and query fields must tell the model to split location from the topic when possible.
79
- - Enum fields must explain when to choose each option.
80
-
81
- For numeric limits:
82
-
83
- - Normal/default value belongs in `.default(...)`.
84
- - Maximum value belongs in `.max(...)`.
85
- - Description must say when to use higher values.
86
-
87
- ## Output Schema And Structured Content
88
-
89
- Any tool whose output may be consumed by another tool should have `outputSchema` and return `structuredContent`.
90
-
91
- Required for chaining tools:
92
-
93
- - Return arrays/objects in `structuredContent`; do not force the model to parse Markdown.
94
- - Keep `content` as a human-readable report.
95
- - Ensure `structuredContent` validates against `outputSchema`.
96
- - Include IDs, URLs, names, positions, counts, and any fields needed for the next tool.
97
-
98
- Example:
99
-
100
- `maps_search` must return `structuredContent.results[]` with at least:
101
-
102
- - `position`
103
- - `name`
104
- - `placeUrl`
105
- - `cid`
106
- - `cidDecimal`
107
- - `rating`
108
- - `reviewCount`
109
- - `category`
110
- - `address`
111
- - `websiteUrl`
112
- - `directionsUrl`
113
- - `metadata`
114
-
115
- `maps_search` and `directory_workflow` must also return sanitized `attempts[]` records for Maps rotation visibility. Include attempt number, max attempts, outcome, retry flag, proxy mode/source/suffix, proxy target level/location/ZIP, browser session suffix, observed IP city/region, error, result count, and duration. Do not expose full browser session IDs, full proxy IDs, or API keys.
116
-
117
- ## Tool Annotations
118
-
119
- Every public tool should define annotations.
120
-
121
- Use:
122
-
123
- - `readOnlyHint: true` for research, scrape, search, transcript, and inspect tools.
124
- - `destructiveHint: false` unless the tool mutates or deletes user data.
125
- - `idempotentHint: false` for live web searches because results, billing, and anti-bot state can change.
126
- - `openWorldHint: true` for tools that access public web or external live systems.
127
- - A human-readable `title`.
128
-
129
- Annotations are hints, not a replacement for descriptions.
130
-
131
- ## Result Text
132
-
133
- Human-readable text still matters. It is what users see and what some clients preserve in context.
134
-
135
- Every successful result should include:
136
-
137
- - Clear title.
138
- - Returned count versus requested count when relevant.
139
- - Key result table or summary.
140
- - Saved report path when report saving is enabled.
141
- - Next-step guidance for common follow-up actions.
142
-
143
- For chained tools, the result text should name the next tool explicitly.
144
-
145
- ## Error Format
146
-
147
- Errors must help the model choose the next action.
148
-
149
- Every API error should include:
150
-
151
- - `error` or `error_code`
152
- - Human-readable message
153
- - Whether retry is reasonable when known
154
- - Enough context to avoid repeating the same bad call
155
-
156
- Common cases:
157
-
158
- - Auth failure: tell user API key is invalid or missing. Do not retry.
159
- - Insufficient balance: return balance, required credits, and top-up URL.
160
- - CAPTCHA/block: for browser-service Google SERP tools, detect it immediately and treat the session as failed; say it is temporary and retryable with a fresh browser session. Do not wait for a solver.
161
- - Validation error: identify the bad or missing field.
162
- - Timeout/cancel: say whether the server attempted cleanup.
163
-
164
- ## Cost And Concurrency
165
-
166
- Tools that cost credits or hold jobs must expose that in metadata or result text.
167
-
168
- Required:
169
-
170
- - Add cost entry to `CREDIT_COST_CATALOG`.
171
- - Add ledger operation for billable work.
172
- - Include refund path on failure where applicable.
173
- - Respect the account concurrency model.
174
- - Tool descriptions should warn when a tool is expensive or long-running.
175
-
176
- ## Docs Surfaces
177
-
178
- For a new or changed public tool, update all applicable surfaces:
179
-
180
- - `src/mcp/paa-mcp-server.ts`
181
- - `src/mcp/mcp-tool-schemas.ts`
182
- - `src/mcp/mcp-response-formatter.ts`
183
- - API route and API schema
184
- - `README.md`
185
- - `public/skill.md`
186
- - `public/codex-skill.md`
187
- - `public/skills/mcp-scraper/skill.md`
188
- - Dashboard UI when users can trigger the workflow there
189
- - Live protocol tool list tests
190
-
191
- ## Packaging And Deployment
192
-
193
- MCP changes often have two release surfaces.
194
-
195
- NPX package:
196
-
197
- - Bump `package.json` and `package-lock.json` when publishing npm.
198
- - Rebuild `dist` with `npx tsup`.
199
- - Verify `npm pack --dry-run` includes the rebuilt stdio binary.
200
- - Verify built output contains the new tool name, description, schema, and formatter behavior.
201
-
202
- Hosted API:
203
-
204
- - Deploy API routes before or with the MCP package.
205
- - A new MCP tool that calls a new hosted endpoint is broken until production has that endpoint.
206
- - Answer "was this prod?" by separating API deployment from npm package publication.
207
-
208
- ## Required Tests
209
-
210
- Every new public MCP tool needs tests at the right level for risk.
211
-
212
- Minimum:
213
-
214
- - Schema default and hard cap tests.
215
- - Formatter test when result text or `structuredContent` matters.
216
- - Tool list/protocol test updated with the new tool.
217
- - Billing/ledger count tests updated for new billable operations.
218
- - Typecheck.
219
- - Unit/contract suite.
220
-
221
- For live-web tools:
222
-
223
- - One live smoke test or saved live evidence for the core workflow.
224
- - If anti-bot behavior is likely, capture failure mode and retry guidance.
225
-
226
- ## Definition Of Done
227
-
228
- A public MCP tool change is not done until all are true:
229
-
230
- - Tool name and boundary are clear.
231
- - Tool description tells the model when to use it and when not to.
232
- - Input schema encodes defaults, limits, and field-level instructions.
233
- - Output is structured when the result will be chained.
234
- - Result text is useful to humans and names the next tool when appropriate.
235
- - Errors are actionable.
236
- - Billing and refunds are correct.
237
- - Dashboard, docs, skill text, and README are updated where relevant.
238
- - `dist` is rebuilt for NPX package changes.
239
- - Production API deployment is accounted for separately from npm publication.
240
- - Tests and smoke evidence prove the workflow.
@@ -1,38 +0,0 @@
1
- # OAuth legal-page review notes
2
-
3
- > Counsel review is required before treating these drafts as final legal advice or launching connected-provider access broadly.
4
-
5
- Prepared July 11, 2026 for `public/privacy.html` and `public/terms.html`.
6
-
7
- ## Official policy sources
8
-
9
- - Google API Services User Data Policy / Limited Use: <https://developers.google.com/terms/api-services-user-data-policy>
10
- - Google APIs Terms: <https://developers.google.com/terms/>
11
- - YouTube API Services Terms: <https://developers.google.com/youtube/terms/api-services-terms-of-service>
12
- - YouTube API Services Developer Policies: <https://developers.google.com/youtube/terms/developer-policies>
13
- - YouTube Terms of Service: <https://www.youtube.com/t/terms>
14
- - Google Privacy Policy: <https://policies.google.com/privacy>
15
- - Meta Platform Terms: <https://developers.facebook.com/terms/>
16
- - Meta Privacy Policy: <https://www.facebook.com/privacy/policy/>
17
- - LinkedIn API Terms of Use: <https://www.linkedin.com/legal/l/api-terms-of-use>
18
- - X Developer Agreement: <https://docs.x.com/developer-terms/agreement>
19
- - X Developer Policy: <https://docs.x.com/developer-terms/policy>
20
- - Nango security and credential storage: <https://nango.dev/docs/guides/platform/security>
21
- - Nango connection deletion API: <https://nango.dev/docs/reference/backend/http-api/connections/delete>
22
-
23
- ## Repo-grounded implementation facts
24
-
25
- - `src/api/nango-control.ts` sends the signed-in identity to the separate Nango control service and retrieves only connections associated with that identity.
26
- - The Integrations UI states that connections are private per login and that OAuth tokens remain server-side.
27
- - `src/api/server.ts` has account deletion/deactivation, Stripe billing, and Resend transactional-email paths.
28
- - `src/api/db.ts` uses Turso/libSQL in production; `vercel.json` defines the hosted application surface.
29
- - Scheduled actions bind an exact connection ID to an allowlist of tools.
30
-
31
- ## Items counsel and the operator must confirm
32
-
33
- 1. Legal entity name, mailing address, governing law, and dispute forum are not present in the repo and should be added if required.
34
- 2. Confirm actual deletion orchestration removes Nango connections and connected-provider data when `/account/delete` runs; until then, the public policy correctly instructs users to email support for connection deletion confirmation.
35
- 3. Confirm exact operational retention periods, backup lifetime, subprocessor list, international-transfer mechanism, and DPA availability.
36
- 4. Confirm the Service age threshold and jurisdiction-specific consumer/privacy disclosures.
37
- 5. Confirm every provider action has the promised consent gate. In particular, YouTube and X require express action consent, and LinkedIn's general API terms restrict automated posting.
38
- 6. Re-run policy review before adding scopes, providers, advertising uses, model training, or new data recipients.
@@ -1,287 +0,0 @@
1
- # SEO Crawl Report — Screaming Frog parity spec (scrape-report scope)
2
-
3
- Goal: from our `extract_site` crawl HTML, reproduce the most valuable ~50% of Screaming Frog's
4
- per-URL crawl reports — **titles, metas, headings, indexability, canonicals, content, structured
5
- data, and internal-link analysis** — and emit them as structured files in the bulk-scrape folder
6
- plus an audit report.
7
-
8
- Scope boundary (user directive): **scrape reports only** — per-URL crawl data and the link graph we
9
- can derive from it. Out of scope here: JS-rendering diffs, log-file analysis, GA/GSC API joins,
10
- spell-check, PageSpeed/Lighthouse, AMP validation, crawl scheduling.
11
-
12
- ---
13
-
14
- ## 0. Current state (verified)
15
-
16
- - `PageData` (`src/api/site-extractor.ts:11`): `url, status, via, title, metaDescription, h1,
17
- headings[{level,text}], wordCount, schemaTypes[], canonicalUrl, internalLinks (count),
18
- externalLinks (count), bodyMarkdown, schema[]`.
19
- - `parsePageData(url, html, status, via)` (`src/api/site-extractor.ts:45`) — HTML only, **no response headers**.
20
- - Fetch layers that DO have headers/timing:
21
- - plain: `fetchAndParse` → `res: Response` (`res.headers`).
22
- - rotating: `rotating-proxy-crawl.ts` `fetchBatch` → playwright `resp` (`resp.status()`, `resp.headers()`).
23
- - `RotatingFetchResult` (`rotating-proxy-crawl.ts:5`) = `{ url, html, status, via }` — must be extended.
24
- - Bulk output: `saveBulkSite(siteUrl, BulkPage[])` (`mcp-response-formatter.ts:68`); `BulkPage` only
25
- carries `url,title,bodyMarkdown,metaDescription,schemaTypes`. Writes `index.md` + `pages/*.md`.
26
- - `formatExtractSite` (`mcp-response-formatter.ts`) maps `pages` → `BulkPage` for the folder.
27
-
28
- ---
29
-
30
- ## 1. Screaming Frog feature inventory → coverage decision
31
-
32
- Legend: ✅ already captured · ◑ partial · ➕ ADD (in this spec) · 🔗 link-graph (Section 3) · ⛔ out of scope
33
-
34
- ### 1.1 Response / crawl
35
- | SF data | Decision | Source |
36
- |---|---|---|
37
- | Address (URL) | ✅ | `PageData.url` |
38
- | Status code + status text | ✅ | `PageData.status` |
39
- | Content-Type | ➕ | response header |
40
- | Response time (ms) | ➕ | fetch timing |
41
- | Size (bytes, HTML transfer) | ➕ | `Content-Length` / `html.length` |
42
- | Redirect URL + redirect type | ➕ | `Location` header / 3xx |
43
- | Last-Modified | ➕ | response header |
44
- | Crawl depth (clicks from start) | 🔗 | computed from link graph |
45
- | Indexability + reason | ➕ | meta-robots + X-Robots + canonical + status |
46
-
47
- ### 1.2 On-page elements
48
- | SF data | Decision | Source |
49
- |---|---|---|
50
- | Title 1 + length | ✅/➕ | have title; ADD length + pixel-width estimate |
51
- | Title — pixel width | ➕ | char-width table estimate (flag approximate) |
52
- | Meta description + length + pixel width | ✅/➕ | have text; ADD lengths |
53
- | Meta keywords | ➕ | regex (low value, cheap) |
54
- | H1 (1st + 2nd) + length | ✅/➕ | have headings[]; derive H1-1/H1-2 + length |
55
- | H2 (1st/2nd) + count | ✅ | from headings[] |
56
- | Word count | ✅ | `PageData.wordCount` |
57
- | Text ratio (text/HTML) | ➕ | `bodyText.length / html.length` |
58
- | Meta robots | ➕ | regex |
59
- | X-Robots-Tag | ➕ | response header |
60
- | Canonical link element | ✅ | `PageData.canonicalUrl` |
61
- | rel=next / rel=prev | ➕ | regex |
62
- | hreflang entries | ➕ | regex (list of {lang,href}) |
63
- | Open Graph tags | ➕ | regex (og:title/description/image/type) |
64
- | Twitter card tags | ➕ | regex |
65
- | Mobile/AMP alternate | ➕ | regex `<link rel="amphtml">`, `alternate` |
66
-
67
- ### 1.3 Images
68
- | SF data | Decision | Source |
69
- |---|---|---|
70
- | Image count | ➕ | count `<img>` |
71
- | Images missing alt | ➕ | `<img>` without non-empty `alt` |
72
- | Alt text over N chars | ➕ | alt length check |
73
- | Image >100KB | ⛔ (needs per-asset fetch) | skip in v1 |
74
-
75
- ### 1.4 Structured data
76
- | SF data | Decision | Source |
77
- |---|---|---|
78
- | Schema types present | ✅ | `PageData.schemaTypes` |
79
- | Raw JSON-LD | ✅ | `PageData.schema` |
80
- | Microdata / RDFa | ◑ | JSON-LD only in v1 |
81
- | Validation errors | ◑ | shape checks only (not full Google validator) |
82
-
83
- ### 1.5 Content
84
- | SF data | Decision | Source |
85
- |---|---|---|
86
- | Low content / thin pages | ➕ | wordCount threshold |
87
- | Exact duplicates | ➕ | hash of normalized body |
88
- | Near-duplicates | ◑ | simhash (v2) — exact-hash in v1 |
89
-
90
- ### 1.6 Links — **the high-value gap (user-flagged)**
91
- | SF data | Decision | Source |
92
- |---|---|---|
93
- | Outlinks (count) | ✅ | `internalLinks`/`externalLinks` counts |
94
- | Outlinks edge list (target, anchor, rel, position) | 🔗➕ | replace counts with edge capture |
95
- | Inlinks + unique inlinks per URL | 🔗 | computed (invert outlinks) |
96
- | Anchor text per inlink | 🔗 | from edges |
97
- | Internal nofollow | 🔗➕ | `rel="nofollow"` on edge |
98
- | Crawl depth | 🔗 | BFS from start over internal edges |
99
- | Orphan URLs | 🔗 | in sitemap/crawl set but zero internal inlinks |
100
- | Broken internal/external links | 🔗 | edge target status ∈ 4xx/5xx |
101
- | Redirecting links | 🔗 | edge target status ∈ 3xx |
102
-
103
- ### 1.7 Issue filters (the actionable "reports" — computed cross-page)
104
- Titles: missing, duplicate, >60 char/>561px, <30 char, multiple, same-as-H1.
105
- Meta desc: missing, duplicate, >155/>985px, <70, multiple.
106
- H1: missing, duplicate, multiple, >70 char.
107
- H2: missing, multiple.
108
- Canonical: missing, canonicalised (non-self), multiple, non-indexable canonical target.
109
- Directives: noindex, nofollow.
110
- Response: 3xx (incl. chains/loops via edges), 4xx broken, 5xx, blocked-by-robots.
111
- URL: >115 char, uppercase, underscores, params, non-ASCII, duplicate.
112
- Content: thin (<X words), exact duplicate (hash collision).
113
- Images: missing alt.
114
- Structured data: present-but-invalid (shape), missing on key templates.
115
- Links: broken internal/external, redirected internal, orphan pages.
116
-
117
- ---
118
-
119
- ## 2. Suggested data model (THE deliverable)
120
-
121
- ### 2.1 Extend `PageData` (`src/api/site-extractor.ts:11`)
122
- Add these fields (keep existing):
123
-
124
- ```ts
125
- // response/header-derived (require fetch-layer plumbing, Section 3.2)
126
- contentType: string | null
127
- responseTimeMs: number | null
128
- sizeBytes: number | null
129
- redirectUrl: string | null // Location on 3xx
130
- lastModified: string | null
131
- xRobotsTag: string | null
132
-
133
- // on-page additions
134
- titleLength: number | null
135
- titlePixels: number | null // estimate, approximate
136
- metaDescLength: number | null
137
- metaKeywords: string | null
138
- h1_2: string | null // second H1 if present
139
- h2Count: number
140
- metaRobots: string | null
141
- relNext: string | null
142
- relPrev: string | null
143
- hreflang: Array<{ lang: string; href: string }>
144
- og: { title?: string; description?: string; image?: string; type?: string } | null
145
- twitter: { card?: string; title?: string; description?: string } | null
146
- ampHref: string | null
147
-
148
- // images
149
- imageCount: number
150
- imagesMissingAlt: number
151
-
152
- // content
153
- textRatio: number // bodyText.length / html.length
154
- contentHash: string // sha1 of normalized visible text
155
-
156
- // indexability (derived, see 3.3)
157
- indexable: boolean
158
- indexabilityReason: string | null // 'noindex' | 'canonicalised' | 'non-200' | 'x-robots-noindex' | null
159
-
160
- // LINKS — replace the two counts with the edge list
161
- outlinks: Array<{ href: string; anchor: string; rel: string | null; internal: boolean }>
162
- // keep internalLinks/externalLinks as derived counts for back-compat
163
- ```
164
-
165
- ### 2.2 Link edge + graph model (post-crawl, computed)
166
- ```ts
167
- interface LinkEdge { from: string; to: string; anchor: string; rel: string | null; internal: boolean }
168
-
169
- interface PageLinkMetrics {
170
- url: string
171
- inlinks: number
172
- uniqueInlinks: number
173
- outlinksInternal: number
174
- outlinksExternal: number
175
- crawlDepth: number | null // BFS hops from startUrl over internal 200 edges; null = orphan/unreachable
176
- orphan: boolean // crawled/in-sitemap but 0 internal inlinks
177
- topAnchors: string[] // most common inbound anchor texts
178
- }
179
- ```
180
-
181
- ### 2.3 Structured outputs written into the bulk folder
182
- Alongside `index.md` + `pages/*.md`, add:
183
- - `pages.jsonl` — one JSON line per page = full extended `PageData` minus `bodyMarkdown`/`schema`
184
- (those stay in `pages/*.md`). This is the "Internal" tab.
185
- - `links.jsonl` — one `LinkEdge` per line (internal + external). The "All Outlinks" export.
186
- - `link-metrics.jsonl` — one `PageLinkMetrics` per URL. The inlinks/depth/orphan view.
187
- - `issues.json` — the computed issue filters (Section 1.7) as `{ issueKey: { count, urls[] } }`.
188
- - `report.md` — human summary (counts per issue, top offenders) = the audit deliverable.
189
-
190
- ---
191
-
192
- ## 3. Implementation blueprint (atomic)
193
-
194
- ### 3.1 Capture: extend `parsePageData` — `src/api/site-extractor.ts:45`
195
- Signature change:
196
- ```ts
197
- function parsePageData(
198
- url: string, html: string, status: number, via: 'fetch'|'browser',
199
- resp?: { headers?: Record<string,string>; responseTimeMs?: number; redirectUrl?: string|null }
200
- ): PageData
201
- ```
202
- Add, inside the function, regex/derivations:
203
- - `titleLength = title?.length`; `titlePixels = estimatePixels(title)` (new helper, char-width table).
204
- - `metaDescLength`, `metaKeywords` via `<meta name="keywords">`.
205
- - `h1_2`/`h2Count` from existing `headings` array.
206
- - `metaRobots` via `<meta name="robots" content="...">`.
207
- - `relNext`/`relPrev` via `<link rel="next|prev">`.
208
- - `hreflang[]` via `<link rel="alternate" hreflang="..">`.
209
- - `og`/`twitter` via `<meta property="og:.."|name="twitter:..">`.
210
- - `ampHref` via `<link rel="amphtml">`.
211
- - `imageCount` = matches of `<img`; `imagesMissingAlt` = `<img>` lacking non-empty `alt`.
212
- - `textRatio = bodyText.length / Math.max(1, html.length)`.
213
- - `contentHash = sha1(bodyText.replace(/\s+/g,' ').trim())` (`node:crypto`).
214
- - header fields from `resp`: `contentType, xRobotsTag, lastModified, sizeBytes, responseTimeMs, redirectUrl`.
215
- - `outlinks[]`: replace the count-only loop (`site-extractor.ts:102-110`) — for each `<a href>` also
216
- capture anchor text and `rel`; classify `internal` by origin; keep `internalLinks`/`externalLinks`
217
- as `.filter().length` for back-compat.
218
- - `indexable`/`indexabilityReason`: `false` if status!=200, or metaRobots/xRobots contains `noindex`,
219
- or canonical present and != self → reason set accordingly.
220
-
221
- New helper `estimatePixels(s)` — sum per-char widths from a static map (approx Arial 'M'≈14, 'i'≈4…);
222
- mark all pixel fields "approximate" in docs. ~25 lines.
223
-
224
- ### 3.2 Plumb headers/timing into both fetch paths
225
- - `RotatingFetchResult` (`rotating-proxy-crawl.ts:5`): add
226
- `headers?: Record<string,string>; responseTimeMs?: number; redirectUrl?: string|null`.
227
- In `fetchBatch` (`:60-98`): time the `goto`, `resp.headers()`, capture `Location` when 3xx.
228
- - `fetchAndParse` (`site-extractor.ts:115`): build the same `resp` object from `res.headers` + timing,
229
- pass as 5th arg to `parsePageData`.
230
- - In `extractSite` rotating branch (`site-extractor.ts:~185`): pass `r.headers/responseTimeMs/redirectUrl`
231
- into `parsePageData(r.url, r.html, r.status, 'browser', {...})`.
232
-
233
- ### 3.3 Post-crawl link graph — new file `src/api/seo-link-graph.ts`
234
- ```ts
235
- export function buildLinkGraph(pages: PageData[], startUrl: string):
236
- { edges: LinkEdge[]; metrics: Map<string, PageLinkMetrics> }
237
- ```
238
- Logic:
239
- - Flatten `pages[].outlinks` → `edges` (from = page.url).
240
- - Build `inlinks` map by inverting internal edges; `uniqueInlinks` = distinct `from`.
241
- - `crawlDepth`: BFS from `startUrl` over internal edges whose target status==200; unreached → null.
242
- - `orphan`: page in set with 0 internal inlinks and url != startUrl.
243
- - `topAnchors`: top-3 inbound anchors by frequency.
244
- - Status join: map target→status from `pages` to flag broken (4xx/5xx) / redirect (3xx) edges.
245
-
246
- ### 3.4 Issue computation — new file `src/api/seo-issues.ts`
247
- ```ts
248
- export function computeIssues(pages: PageData[], metrics: Map<string,PageLinkMetrics>):
249
- Record<string, { count: number; urls: string[] }>
250
- ```
251
- Thresholds (SF defaults): title >60/<30 char & >561px; meta >155/<70 & >985px; H1 >70; URL >115;
252
- thin <200 words (configurable). Duplicates: group by exact `title`/`metaDescription`/`contentHash`,
253
- flag groups size>1. Broken/redirect/orphan from `metrics`/edges. Indexability from `PageData`.
254
-
255
- ### 3.5 Write structured outputs — extend `saveBulkSite` (`mcp-response-formatter.ts:68`)
256
- - Widen `BulkPage` → accept full extended `PageData` (or pass `PageData[]` directly).
257
- - After writing `pages/*.md`, also `writeFileSync`:
258
- - `pages.jsonl` (PageData minus bodyMarkdown/schema),
259
- - `links.jsonl`, `link-metrics.jsonl`, `issues.json`, `report.md`.
260
- - `formatExtractSite` already has full `pages`; pass them through (today it maps to the slim BulkPage —
261
- change to pass the full objects, call `buildLinkGraph` + `computeIssues` before `saveBulkSite`).
262
- - Return extra paths in the bulk summary + `structuredContent` (e.g. `pagesJsonl`, `issuesFile`).
263
-
264
- ### 3.6 The skill — `seo-crawl-audit`
265
- Orchestrator (no new server code):
266
- 1. Call `extract_site(url, rotateProxies:true)` → folder path from response.
267
- 2. Read `pages.jsonl` + `issues.json` + `link-metrics.jsonl` from the folder.
268
- 3. Emit a prioritized SEO report: response-code breakdown, title/meta/H1 issues with offender lists,
269
- indexability summary, thin/duplicate content, structured-data coverage, and the internal-link
270
- section (orphans, deepest pages, most-linked, broken internal links).
271
- 4. Hand the full link graph to the existing `site-architecture-auditor` skill for equity/architecture
272
- scoring (it already builds this graph from SF exports — we now feed it ours).
273
-
274
- ---
275
-
276
- ## 4. Build order (phased, each independently shippable)
277
- 1. **P1 — page fields (no headers):** 3.1 minus header fields + 3.5 `pages.jsonl`. Unlocks titles/metas/
278
- H1/canonical/indexability(meta)/thin/dup/schema reports immediately.
279
- 2. **P2 — link graph:** 3.1 outlinks edge capture + 3.3 + `links.jsonl`/`link-metrics.jsonl`. Unlocks
280
- inlinks/depth/orphans/broken-links — the user's headline ask.
281
- 3. **P3 — headers/timing:** 3.2 (both fetch paths). Adds content-type/size/response-time/X-Robots/redirects.
282
- 4. **P4 — issues + report:** 3.4 + `issues.json`/`report.md`.
283
- 5. **P5 — skill:** 3.6 + handoff to `site-architecture-auditor`.
284
-
285
- ## 5. Explicit non-goals (v1)
286
- Per-asset image weight, microdata/RDFa, full structured-data validation, near-duplicate simhash,
287
- JS-render diffing, pixel-width exactness (estimate only), redirect-chain hop-by-hop beyond one hop.