mcp-scraper 0.38.2 → 0.40.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +5 -2
- package/package.json +5 -6
- package/dist/bin/api-server.cjs +0 -58752
- package/dist/bin/api-server.cjs.map +0 -1
- package/dist/bin/api-server.d.cts +0 -1
- package/dist/bin/api-server.d.ts +0 -1
- package/dist/bin/api-server.js +0 -38
- package/dist/bin/api-server.js.map +0 -1
- package/dist/bin/mcp-scraper-cli.cjs +0 -2671
- package/dist/bin/mcp-scraper-cli.cjs.map +0 -1
- package/dist/bin/mcp-scraper-cli.d.cts +0 -1
- package/dist/bin/mcp-scraper-cli.d.ts +0 -1
- package/dist/bin/mcp-scraper-cli.js +0 -742
- package/dist/bin/mcp-scraper-cli.js.map +0 -1
- package/dist/bin/mcp-scraper-install.cjs +0 -129
- package/dist/bin/mcp-scraper-install.cjs.map +0 -1
- package/dist/bin/mcp-scraper-install.d.cts +0 -1
- package/dist/bin/mcp-scraper-install.d.ts +0 -1
- package/dist/bin/mcp-scraper-install.js +0 -27
- package/dist/bin/mcp-scraper-install.js.map +0 -1
- package/dist/bin/mcp-stdio-server.cjs +0 -12264
- package/dist/bin/mcp-stdio-server.cjs.map +0 -1
- package/dist/bin/mcp-stdio-server.d.cts +0 -1
- package/dist/bin/mcp-stdio-server.d.ts +0 -1
- package/dist/bin/mcp-stdio-server.js +0 -135
- package/dist/bin/mcp-stdio-server.js.map +0 -1
- package/dist/bin/paa-harvest.cjs +0 -3808
- package/dist/bin/paa-harvest.cjs.map +0 -1
- package/dist/bin/paa-harvest.d.cts +0 -1
- package/dist/bin/paa-harvest.d.ts +0 -1
- package/dist/bin/paa-harvest.js +0 -44
- package/dist/bin/paa-harvest.js.map +0 -1
- package/dist/chunk-345BQXZH.js +0 -712
- package/dist/chunk-345BQXZH.js.map +0 -1
- package/dist/chunk-44HZLHDV.js +0 -52
- package/dist/chunk-44HZLHDV.js.map +0 -1
- package/dist/chunk-AZRPG43B.js +0 -617
- package/dist/chunk-AZRPG43B.js.map +0 -1
- package/dist/chunk-CB5C3BPB.js +0 -135
- package/dist/chunk-CB5C3BPB.js.map +0 -1
- package/dist/chunk-EGJKUB4Q.js +0 -276
- package/dist/chunk-EGJKUB4Q.js.map +0 -1
- package/dist/chunk-FQI5PFE7.js +0 -1866
- package/dist/chunk-FQI5PFE7.js.map +0 -1
- package/dist/chunk-FRYT3ID4.js +0 -684
- package/dist/chunk-FRYT3ID4.js.map +0 -1
- package/dist/chunk-G3P3ZDB4.js +0 -69
- package/dist/chunk-G3P3ZDB4.js.map +0 -1
- package/dist/chunk-K443GQY5.js +0 -24
- package/dist/chunk-K443GQY5.js.map +0 -1
- package/dist/chunk-N7KUTTCC.js +0 -3007
- package/dist/chunk-N7KUTTCC.js.map +0 -1
- package/dist/chunk-NGM237OO.js +0 -3410
- package/dist/chunk-NGM237OO.js.map +0 -1
- package/dist/chunk-NKCCGADE.js +0 -11285
- package/dist/chunk-NKCCGADE.js.map +0 -1
- package/dist/chunk-NNW3O6ZD.js +0 -108
- package/dist/chunk-NNW3O6ZD.js.map +0 -1
- package/dist/chunk-QZXKQB7Y.js +0 -414
- package/dist/chunk-QZXKQB7Y.js.map +0 -1
- package/dist/chunk-SFRMFGQ6.js +0 -158
- package/dist/chunk-SFRMFGQ6.js.map +0 -1
- package/dist/chunk-YCI2PNCS.js +0 -499
- package/dist/chunk-YCI2PNCS.js.map +0 -1
- package/dist/chunk-YODBNTTN.js +0 -7
- package/dist/chunk-YODBNTTN.js.map +0 -1
- package/dist/db-C5KVCOYT.js +0 -239
- package/dist/db-C5KVCOYT.js.map +0 -1
- package/dist/extract-bundle-KUBX6N6Z.js +0 -568
- package/dist/extract-bundle-KUBX6N6Z.js.map +0 -1
- package/dist/index.cjs +0 -4160
- package/dist/index.cjs.map +0 -1
- package/dist/index.d.cts +0 -413
- package/dist/index.d.ts +0 -413
- package/dist/index.js +0 -338
- package/dist/index.js.map +0 -1
- package/dist/location-data-repository-TTWF3OTM.js +0 -35
- package/dist/location-data-repository-TTWF3OTM.js.map +0 -1
- package/dist/server-5EX6XBIA.js +0 -33596
- package/dist/server-5EX6XBIA.js.map +0 -1
- package/dist/site-extract-repository-XSPJTCIL.js +0 -62
- package/dist/site-extract-repository-XSPJTCIL.js.map +0 -1
- package/dist/worker-XCPU4YSN.js +0 -142
- package/dist/worker-XCPU4YSN.js.map +0 -1
- package/docs/adr/0001-in-page-graphql-interception-for-anti-bot-scraping.md +0 -58
- package/docs/adr/0002-hybrid-smart-rag-vault-retrieval.md +0 -62
- package/docs/adr/0003-waive-unrecoverable-scheduled-model-cost.md +0 -22
- package/docs/adr/README.md +0 -13
- package/docs/final-tooling-spec.md +0 -206
- package/docs/hosted-location-data.md +0 -108
- package/docs/kernel-proxy-future-enhancements.md +0 -80
- package/docs/mcp-tool-craft-lint.generated.md +0 -183
- package/docs/mcp-tool-design-guide.md +0 -225
- package/docs/mcp-tool-manifest.generated.json +0 -22871
- package/docs/mcp-tool-quality-spec.md +0 -240
- package/docs/oauth-legal-review.md +0 -38
- package/docs/seo-crawl-report-spec.md +0 -287
- package/docs/specs/api-forge-spec.md +0 -234
- package/docs/specs/connected-services-control-plane-decoupling-spec.md +0 -1044
- package/docs/specs/deferred-work-spec.md +0 -86
- package/docs/specs/google-drive-bulk-access-and-mcp-schema-passthrough-spec.md +0 -1689
- package/docs/specs/kernel-stealth-captcha-test-matrix.md +0 -278
- package/docs/specs/main-mcp-integration-ownership-spec.md +0 -1164
- package/docs/specs/mcp-tool-definition-quality-audit-spec.md +0 -1602
- package/docs/specs/meta-ad-creative-media-resolution-spec.md +0 -31
- package/docs/specs/multimodal-image-memory-architecture-spec.md +0 -1022
- package/docs/specs/oauth-mcp-spec.md +0 -213
- package/docs/specs/query-fanout-transport-contract-fix.md +0 -45
- package/docs/specs/relationship-workspace-ai-behavior-plan.md +0 -26
- package/docs/specs/unified-credit-and-scheduled-execution-billing-spec.md +0 -995
- package/docs/tool-catalog-spec.md +0 -388
|
@@ -1,240 +0,0 @@
|
|
|
1
|
-
# MCP Tool Quality Spec
|
|
2
|
-
|
|
3
|
-
This spec defines the shipping bar for MCP Scraper tools. It exists because MCP behavior is model-facing: a tool can be technically callable and still fail if the AI cannot infer when to use it, how to fill inputs, or how to chain the result.
|
|
4
|
-
|
|
5
|
-
## What Actually Steers The AI
|
|
6
|
-
|
|
7
|
-
The model is primarily affected by what the MCP client receives from `tools/list` and `tools/call`.
|
|
8
|
-
|
|
9
|
-
1. Tool name
|
|
10
|
-
2. Tool title and annotations
|
|
11
|
-
3. Tool description
|
|
12
|
-
4. Input schema field names, descriptions, defaults, enums, and limits
|
|
13
|
-
5. Output schema and `structuredContent`
|
|
14
|
-
6. Tool result text and next-step guidance
|
|
15
|
-
7. Error shape and retry guidance
|
|
16
|
-
|
|
17
|
-
README files, website copy, and public skill markdown help humans and skill loaders, but they do not reliably affect runtime AI behavior unless the client explicitly injects those files into context.
|
|
18
|
-
|
|
19
|
-
## Tool Boundary Rules
|
|
20
|
-
|
|
21
|
-
Each tool must have one primary job.
|
|
22
|
-
|
|
23
|
-
- Split tools when the user intent, cost model, result shape, or follow-up workflow differs.
|
|
24
|
-
- Do not overload one tool with unrelated modes if a model could choose the wrong path.
|
|
25
|
-
- Descriptions for adjacent tools must explicitly say when not to use each tool.
|
|
26
|
-
- If a workflow has a natural sequence, encode it in the description and result guidance.
|
|
27
|
-
|
|
28
|
-
Example:
|
|
29
|
-
|
|
30
|
-
- `maps_search`: find multiple Google Maps candidates for a category, market, lead list, or "more than the 3-pack".
|
|
31
|
-
- `maps_place_intel`: hydrate one known business or one selected candidate with full profile details and optional reviews.
|
|
32
|
-
|
|
33
|
-
## Tool Naming
|
|
34
|
-
|
|
35
|
-
Names must be short, stable, and action-oriented.
|
|
36
|
-
|
|
37
|
-
- Use domain + action: `maps_search`, `maps_place_intel`, `facebook_ad_search`.
|
|
38
|
-
- Avoid generic names such as `search`, `lookup`, `get_data`, or `run`.
|
|
39
|
-
- Avoid names that imply broader capability than the tool has.
|
|
40
|
-
- Do not rename a public tool without a compatibility plan.
|
|
41
|
-
|
|
42
|
-
## Tool Descriptions
|
|
43
|
-
|
|
44
|
-
Descriptions are model instructions. They must be concise but operational.
|
|
45
|
-
|
|
46
|
-
Every description must include:
|
|
47
|
-
|
|
48
|
-
- What the tool does.
|
|
49
|
-
- When to use it.
|
|
50
|
-
- When not to use it if there is an adjacent tool.
|
|
51
|
-
- Important defaults and hard caps.
|
|
52
|
-
- Cost-sensitive behavior when relevant.
|
|
53
|
-
- The expected next tool when chaining is common.
|
|
54
|
-
- Whether reports are saved locally.
|
|
55
|
-
|
|
56
|
-
Bad:
|
|
57
|
-
|
|
58
|
-
```text
|
|
59
|
-
Search Google Maps.
|
|
60
|
-
```
|
|
61
|
-
|
|
62
|
-
Good:
|
|
63
|
-
|
|
64
|
-
```text
|
|
65
|
-
Search Google Maps for multiple businesses/profiles by category, niche, keyword, or local market. Use this when the user asks for several Google Business Profiles, GMBs, GBPs, leads, prospects, competitors, or "more than the 3-pack." Returns up to 50 candidates. Default maxResults is 10; maximum is 50. Use maps_place_intel afterward only when a selected business needs full details and reviews.
|
|
66
|
-
```
|
|
67
|
-
|
|
68
|
-
## Input Schema
|
|
69
|
-
|
|
70
|
-
Each input field must have a clear description.
|
|
71
|
-
|
|
72
|
-
Required:
|
|
73
|
-
|
|
74
|
-
- Required fields must be actually required in the schema.
|
|
75
|
-
- Defaults must be encoded in the schema, not only in prose.
|
|
76
|
-
- Hard caps must be encoded in the schema.
|
|
77
|
-
- Fields that are often confused must say what not to put there.
|
|
78
|
-
- Location and query fields must tell the model to split location from the topic when possible.
|
|
79
|
-
- Enum fields must explain when to choose each option.
|
|
80
|
-
|
|
81
|
-
For numeric limits:
|
|
82
|
-
|
|
83
|
-
- Normal/default value belongs in `.default(...)`.
|
|
84
|
-
- Maximum value belongs in `.max(...)`.
|
|
85
|
-
- Description must say when to use higher values.
|
|
86
|
-
|
|
87
|
-
## Output Schema And Structured Content
|
|
88
|
-
|
|
89
|
-
Any tool whose output may be consumed by another tool should have `outputSchema` and return `structuredContent`.
|
|
90
|
-
|
|
91
|
-
Required for chaining tools:
|
|
92
|
-
|
|
93
|
-
- Return arrays/objects in `structuredContent`; do not force the model to parse Markdown.
|
|
94
|
-
- Keep `content` as a human-readable report.
|
|
95
|
-
- Ensure `structuredContent` validates against `outputSchema`.
|
|
96
|
-
- Include IDs, URLs, names, positions, counts, and any fields needed for the next tool.
|
|
97
|
-
|
|
98
|
-
Example:
|
|
99
|
-
|
|
100
|
-
`maps_search` must return `structuredContent.results[]` with at least:
|
|
101
|
-
|
|
102
|
-
- `position`
|
|
103
|
-
- `name`
|
|
104
|
-
- `placeUrl`
|
|
105
|
-
- `cid`
|
|
106
|
-
- `cidDecimal`
|
|
107
|
-
- `rating`
|
|
108
|
-
- `reviewCount`
|
|
109
|
-
- `category`
|
|
110
|
-
- `address`
|
|
111
|
-
- `websiteUrl`
|
|
112
|
-
- `directionsUrl`
|
|
113
|
-
- `metadata`
|
|
114
|
-
|
|
115
|
-
`maps_search` and `directory_workflow` must also return sanitized `attempts[]` records for Maps rotation visibility. Include attempt number, max attempts, outcome, retry flag, proxy mode/source/suffix, proxy target level/location/ZIP, browser session suffix, observed IP city/region, error, result count, and duration. Do not expose full browser session IDs, full proxy IDs, or API keys.
|
|
116
|
-
|
|
117
|
-
## Tool Annotations
|
|
118
|
-
|
|
119
|
-
Every public tool should define annotations.
|
|
120
|
-
|
|
121
|
-
Use:
|
|
122
|
-
|
|
123
|
-
- `readOnlyHint: true` for research, scrape, search, transcript, and inspect tools.
|
|
124
|
-
- `destructiveHint: false` unless the tool mutates or deletes user data.
|
|
125
|
-
- `idempotentHint: false` for live web searches because results, billing, and anti-bot state can change.
|
|
126
|
-
- `openWorldHint: true` for tools that access public web or external live systems.
|
|
127
|
-
- A human-readable `title`.
|
|
128
|
-
|
|
129
|
-
Annotations are hints, not a replacement for descriptions.
|
|
130
|
-
|
|
131
|
-
## Result Text
|
|
132
|
-
|
|
133
|
-
Human-readable text still matters. It is what users see and what some clients preserve in context.
|
|
134
|
-
|
|
135
|
-
Every successful result should include:
|
|
136
|
-
|
|
137
|
-
- Clear title.
|
|
138
|
-
- Returned count versus requested count when relevant.
|
|
139
|
-
- Key result table or summary.
|
|
140
|
-
- Saved report path when report saving is enabled.
|
|
141
|
-
- Next-step guidance for common follow-up actions.
|
|
142
|
-
|
|
143
|
-
For chained tools, the result text should name the next tool explicitly.
|
|
144
|
-
|
|
145
|
-
## Error Format
|
|
146
|
-
|
|
147
|
-
Errors must help the model choose the next action.
|
|
148
|
-
|
|
149
|
-
Every API error should include:
|
|
150
|
-
|
|
151
|
-
- `error` or `error_code`
|
|
152
|
-
- Human-readable message
|
|
153
|
-
- Whether retry is reasonable when known
|
|
154
|
-
- Enough context to avoid repeating the same bad call
|
|
155
|
-
|
|
156
|
-
Common cases:
|
|
157
|
-
|
|
158
|
-
- Auth failure: tell user API key is invalid or missing. Do not retry.
|
|
159
|
-
- Insufficient balance: return balance, required credits, and top-up URL.
|
|
160
|
-
- CAPTCHA/block: for browser-service Google SERP tools, detect it immediately and treat the session as failed; say it is temporary and retryable with a fresh browser session. Do not wait for a solver.
|
|
161
|
-
- Validation error: identify the bad or missing field.
|
|
162
|
-
- Timeout/cancel: say whether the server attempted cleanup.
|
|
163
|
-
|
|
164
|
-
## Cost And Concurrency
|
|
165
|
-
|
|
166
|
-
Tools that cost credits or hold jobs must expose that in metadata or result text.
|
|
167
|
-
|
|
168
|
-
Required:
|
|
169
|
-
|
|
170
|
-
- Add cost entry to `CREDIT_COST_CATALOG`.
|
|
171
|
-
- Add ledger operation for billable work.
|
|
172
|
-
- Include refund path on failure where applicable.
|
|
173
|
-
- Respect the account concurrency model.
|
|
174
|
-
- Tool descriptions should warn when a tool is expensive or long-running.
|
|
175
|
-
|
|
176
|
-
## Docs Surfaces
|
|
177
|
-
|
|
178
|
-
For a new or changed public tool, update all applicable surfaces:
|
|
179
|
-
|
|
180
|
-
- `src/mcp/paa-mcp-server.ts`
|
|
181
|
-
- `src/mcp/mcp-tool-schemas.ts`
|
|
182
|
-
- `src/mcp/mcp-response-formatter.ts`
|
|
183
|
-
- API route and API schema
|
|
184
|
-
- `README.md`
|
|
185
|
-
- `public/skill.md`
|
|
186
|
-
- `public/codex-skill.md`
|
|
187
|
-
- `public/skills/mcp-scraper/skill.md`
|
|
188
|
-
- Dashboard UI when users can trigger the workflow there
|
|
189
|
-
- Live protocol tool list tests
|
|
190
|
-
|
|
191
|
-
## Packaging And Deployment
|
|
192
|
-
|
|
193
|
-
MCP changes often have two release surfaces.
|
|
194
|
-
|
|
195
|
-
NPX package:
|
|
196
|
-
|
|
197
|
-
- Bump `package.json` and `package-lock.json` when publishing npm.
|
|
198
|
-
- Rebuild `dist` with `npx tsup`.
|
|
199
|
-
- Verify `npm pack --dry-run` includes the rebuilt stdio binary.
|
|
200
|
-
- Verify built output contains the new tool name, description, schema, and formatter behavior.
|
|
201
|
-
|
|
202
|
-
Hosted API:
|
|
203
|
-
|
|
204
|
-
- Deploy API routes before or with the MCP package.
|
|
205
|
-
- A new MCP tool that calls a new hosted endpoint is broken until production has that endpoint.
|
|
206
|
-
- Answer "was this prod?" by separating API deployment from npm package publication.
|
|
207
|
-
|
|
208
|
-
## Required Tests
|
|
209
|
-
|
|
210
|
-
Every new public MCP tool needs tests at the right level for risk.
|
|
211
|
-
|
|
212
|
-
Minimum:
|
|
213
|
-
|
|
214
|
-
- Schema default and hard cap tests.
|
|
215
|
-
- Formatter test when result text or `structuredContent` matters.
|
|
216
|
-
- Tool list/protocol test updated with the new tool.
|
|
217
|
-
- Billing/ledger count tests updated for new billable operations.
|
|
218
|
-
- Typecheck.
|
|
219
|
-
- Unit/contract suite.
|
|
220
|
-
|
|
221
|
-
For live-web tools:
|
|
222
|
-
|
|
223
|
-
- One live smoke test or saved live evidence for the core workflow.
|
|
224
|
-
- If anti-bot behavior is likely, capture failure mode and retry guidance.
|
|
225
|
-
|
|
226
|
-
## Definition Of Done
|
|
227
|
-
|
|
228
|
-
A public MCP tool change is not done until all are true:
|
|
229
|
-
|
|
230
|
-
- Tool name and boundary are clear.
|
|
231
|
-
- Tool description tells the model when to use it and when not to.
|
|
232
|
-
- Input schema encodes defaults, limits, and field-level instructions.
|
|
233
|
-
- Output is structured when the result will be chained.
|
|
234
|
-
- Result text is useful to humans and names the next tool when appropriate.
|
|
235
|
-
- Errors are actionable.
|
|
236
|
-
- Billing and refunds are correct.
|
|
237
|
-
- Dashboard, docs, skill text, and README are updated where relevant.
|
|
238
|
-
- `dist` is rebuilt for NPX package changes.
|
|
239
|
-
- Production API deployment is accounted for separately from npm publication.
|
|
240
|
-
- Tests and smoke evidence prove the workflow.
|
|
@@ -1,38 +0,0 @@
|
|
|
1
|
-
# OAuth legal-page review notes
|
|
2
|
-
|
|
3
|
-
> Counsel review is required before treating these drafts as final legal advice or launching connected-provider access broadly.
|
|
4
|
-
|
|
5
|
-
Prepared July 11, 2026 for `public/privacy.html` and `public/terms.html`.
|
|
6
|
-
|
|
7
|
-
## Official policy sources
|
|
8
|
-
|
|
9
|
-
- Google API Services User Data Policy / Limited Use: <https://developers.google.com/terms/api-services-user-data-policy>
|
|
10
|
-
- Google APIs Terms: <https://developers.google.com/terms/>
|
|
11
|
-
- YouTube API Services Terms: <https://developers.google.com/youtube/terms/api-services-terms-of-service>
|
|
12
|
-
- YouTube API Services Developer Policies: <https://developers.google.com/youtube/terms/developer-policies>
|
|
13
|
-
- YouTube Terms of Service: <https://www.youtube.com/t/terms>
|
|
14
|
-
- Google Privacy Policy: <https://policies.google.com/privacy>
|
|
15
|
-
- Meta Platform Terms: <https://developers.facebook.com/terms/>
|
|
16
|
-
- Meta Privacy Policy: <https://www.facebook.com/privacy/policy/>
|
|
17
|
-
- LinkedIn API Terms of Use: <https://www.linkedin.com/legal/l/api-terms-of-use>
|
|
18
|
-
- X Developer Agreement: <https://docs.x.com/developer-terms/agreement>
|
|
19
|
-
- X Developer Policy: <https://docs.x.com/developer-terms/policy>
|
|
20
|
-
- Nango security and credential storage: <https://nango.dev/docs/guides/platform/security>
|
|
21
|
-
- Nango connection deletion API: <https://nango.dev/docs/reference/backend/http-api/connections/delete>
|
|
22
|
-
|
|
23
|
-
## Repo-grounded implementation facts
|
|
24
|
-
|
|
25
|
-
- `src/api/nango-control.ts` sends the signed-in identity to the separate Nango control service and retrieves only connections associated with that identity.
|
|
26
|
-
- The Integrations UI states that connections are private per login and that OAuth tokens remain server-side.
|
|
27
|
-
- `src/api/server.ts` has account deletion/deactivation, Stripe billing, and Resend transactional-email paths.
|
|
28
|
-
- `src/api/db.ts` uses Turso/libSQL in production; `vercel.json` defines the hosted application surface.
|
|
29
|
-
- Scheduled actions bind an exact connection ID to an allowlist of tools.
|
|
30
|
-
|
|
31
|
-
## Items counsel and the operator must confirm
|
|
32
|
-
|
|
33
|
-
1. Legal entity name, mailing address, governing law, and dispute forum are not present in the repo and should be added if required.
|
|
34
|
-
2. Confirm actual deletion orchestration removes Nango connections and connected-provider data when `/account/delete` runs; until then, the public policy correctly instructs users to email support for connection deletion confirmation.
|
|
35
|
-
3. Confirm exact operational retention periods, backup lifetime, subprocessor list, international-transfer mechanism, and DPA availability.
|
|
36
|
-
4. Confirm the Service age threshold and jurisdiction-specific consumer/privacy disclosures.
|
|
37
|
-
5. Confirm every provider action has the promised consent gate. In particular, YouTube and X require express action consent, and LinkedIn's general API terms restrict automated posting.
|
|
38
|
-
6. Re-run policy review before adding scopes, providers, advertising uses, model training, or new data recipients.
|
|
@@ -1,287 +0,0 @@
|
|
|
1
|
-
# SEO Crawl Report — Screaming Frog parity spec (scrape-report scope)
|
|
2
|
-
|
|
3
|
-
Goal: from our `extract_site` crawl HTML, reproduce the most valuable ~50% of Screaming Frog's
|
|
4
|
-
per-URL crawl reports — **titles, metas, headings, indexability, canonicals, content, structured
|
|
5
|
-
data, and internal-link analysis** — and emit them as structured files in the bulk-scrape folder
|
|
6
|
-
plus an audit report.
|
|
7
|
-
|
|
8
|
-
Scope boundary (user directive): **scrape reports only** — per-URL crawl data and the link graph we
|
|
9
|
-
can derive from it. Out of scope here: JS-rendering diffs, log-file analysis, GA/GSC API joins,
|
|
10
|
-
spell-check, PageSpeed/Lighthouse, AMP validation, crawl scheduling.
|
|
11
|
-
|
|
12
|
-
---
|
|
13
|
-
|
|
14
|
-
## 0. Current state (verified)
|
|
15
|
-
|
|
16
|
-
- `PageData` (`src/api/site-extractor.ts:11`): `url, status, via, title, metaDescription, h1,
|
|
17
|
-
headings[{level,text}], wordCount, schemaTypes[], canonicalUrl, internalLinks (count),
|
|
18
|
-
externalLinks (count), bodyMarkdown, schema[]`.
|
|
19
|
-
- `parsePageData(url, html, status, via)` (`src/api/site-extractor.ts:45`) — HTML only, **no response headers**.
|
|
20
|
-
- Fetch layers that DO have headers/timing:
|
|
21
|
-
- plain: `fetchAndParse` → `res: Response` (`res.headers`).
|
|
22
|
-
- rotating: `rotating-proxy-crawl.ts` `fetchBatch` → playwright `resp` (`resp.status()`, `resp.headers()`).
|
|
23
|
-
- `RotatingFetchResult` (`rotating-proxy-crawl.ts:5`) = `{ url, html, status, via }` — must be extended.
|
|
24
|
-
- Bulk output: `saveBulkSite(siteUrl, BulkPage[])` (`mcp-response-formatter.ts:68`); `BulkPage` only
|
|
25
|
-
carries `url,title,bodyMarkdown,metaDescription,schemaTypes`. Writes `index.md` + `pages/*.md`.
|
|
26
|
-
- `formatExtractSite` (`mcp-response-formatter.ts`) maps `pages` → `BulkPage` for the folder.
|
|
27
|
-
|
|
28
|
-
---
|
|
29
|
-
|
|
30
|
-
## 1. Screaming Frog feature inventory → coverage decision
|
|
31
|
-
|
|
32
|
-
Legend: ✅ already captured · ◑ partial · ➕ ADD (in this spec) · 🔗 link-graph (Section 3) · ⛔ out of scope
|
|
33
|
-
|
|
34
|
-
### 1.1 Response / crawl
|
|
35
|
-
| SF data | Decision | Source |
|
|
36
|
-
|---|---|---|
|
|
37
|
-
| Address (URL) | ✅ | `PageData.url` |
|
|
38
|
-
| Status code + status text | ✅ | `PageData.status` |
|
|
39
|
-
| Content-Type | ➕ | response header |
|
|
40
|
-
| Response time (ms) | ➕ | fetch timing |
|
|
41
|
-
| Size (bytes, HTML transfer) | ➕ | `Content-Length` / `html.length` |
|
|
42
|
-
| Redirect URL + redirect type | ➕ | `Location` header / 3xx |
|
|
43
|
-
| Last-Modified | ➕ | response header |
|
|
44
|
-
| Crawl depth (clicks from start) | 🔗 | computed from link graph |
|
|
45
|
-
| Indexability + reason | ➕ | meta-robots + X-Robots + canonical + status |
|
|
46
|
-
|
|
47
|
-
### 1.2 On-page elements
|
|
48
|
-
| SF data | Decision | Source |
|
|
49
|
-
|---|---|---|
|
|
50
|
-
| Title 1 + length | ✅/➕ | have title; ADD length + pixel-width estimate |
|
|
51
|
-
| Title — pixel width | ➕ | char-width table estimate (flag approximate) |
|
|
52
|
-
| Meta description + length + pixel width | ✅/➕ | have text; ADD lengths |
|
|
53
|
-
| Meta keywords | ➕ | regex (low value, cheap) |
|
|
54
|
-
| H1 (1st + 2nd) + length | ✅/➕ | have headings[]; derive H1-1/H1-2 + length |
|
|
55
|
-
| H2 (1st/2nd) + count | ✅ | from headings[] |
|
|
56
|
-
| Word count | ✅ | `PageData.wordCount` |
|
|
57
|
-
| Text ratio (text/HTML) | ➕ | `bodyText.length / html.length` |
|
|
58
|
-
| Meta robots | ➕ | regex |
|
|
59
|
-
| X-Robots-Tag | ➕ | response header |
|
|
60
|
-
| Canonical link element | ✅ | `PageData.canonicalUrl` |
|
|
61
|
-
| rel=next / rel=prev | ➕ | regex |
|
|
62
|
-
| hreflang entries | ➕ | regex (list of {lang,href}) |
|
|
63
|
-
| Open Graph tags | ➕ | regex (og:title/description/image/type) |
|
|
64
|
-
| Twitter card tags | ➕ | regex |
|
|
65
|
-
| Mobile/AMP alternate | ➕ | regex `<link rel="amphtml">`, `alternate` |
|
|
66
|
-
|
|
67
|
-
### 1.3 Images
|
|
68
|
-
| SF data | Decision | Source |
|
|
69
|
-
|---|---|---|
|
|
70
|
-
| Image count | ➕ | count `<img>` |
|
|
71
|
-
| Images missing alt | ➕ | `<img>` without non-empty `alt` |
|
|
72
|
-
| Alt text over N chars | ➕ | alt length check |
|
|
73
|
-
| Image >100KB | ⛔ (needs per-asset fetch) | skip in v1 |
|
|
74
|
-
|
|
75
|
-
### 1.4 Structured data
|
|
76
|
-
| SF data | Decision | Source |
|
|
77
|
-
|---|---|---|
|
|
78
|
-
| Schema types present | ✅ | `PageData.schemaTypes` |
|
|
79
|
-
| Raw JSON-LD | ✅ | `PageData.schema` |
|
|
80
|
-
| Microdata / RDFa | ◑ | JSON-LD only in v1 |
|
|
81
|
-
| Validation errors | ◑ | shape checks only (not full Google validator) |
|
|
82
|
-
|
|
83
|
-
### 1.5 Content
|
|
84
|
-
| SF data | Decision | Source |
|
|
85
|
-
|---|---|---|
|
|
86
|
-
| Low content / thin pages | ➕ | wordCount threshold |
|
|
87
|
-
| Exact duplicates | ➕ | hash of normalized body |
|
|
88
|
-
| Near-duplicates | ◑ | simhash (v2) — exact-hash in v1 |
|
|
89
|
-
|
|
90
|
-
### 1.6 Links — **the high-value gap (user-flagged)**
|
|
91
|
-
| SF data | Decision | Source |
|
|
92
|
-
|---|---|---|
|
|
93
|
-
| Outlinks (count) | ✅ | `internalLinks`/`externalLinks` counts |
|
|
94
|
-
| Outlinks edge list (target, anchor, rel, position) | 🔗➕ | replace counts with edge capture |
|
|
95
|
-
| Inlinks + unique inlinks per URL | 🔗 | computed (invert outlinks) |
|
|
96
|
-
| Anchor text per inlink | 🔗 | from edges |
|
|
97
|
-
| Internal nofollow | 🔗➕ | `rel="nofollow"` on edge |
|
|
98
|
-
| Crawl depth | 🔗 | BFS from start over internal edges |
|
|
99
|
-
| Orphan URLs | 🔗 | in sitemap/crawl set but zero internal inlinks |
|
|
100
|
-
| Broken internal/external links | 🔗 | edge target status ∈ 4xx/5xx |
|
|
101
|
-
| Redirecting links | 🔗 | edge target status ∈ 3xx |
|
|
102
|
-
|
|
103
|
-
### 1.7 Issue filters (the actionable "reports" — computed cross-page)
|
|
104
|
-
Titles: missing, duplicate, >60 char/>561px, <30 char, multiple, same-as-H1.
|
|
105
|
-
Meta desc: missing, duplicate, >155/>985px, <70, multiple.
|
|
106
|
-
H1: missing, duplicate, multiple, >70 char.
|
|
107
|
-
H2: missing, multiple.
|
|
108
|
-
Canonical: missing, canonicalised (non-self), multiple, non-indexable canonical target.
|
|
109
|
-
Directives: noindex, nofollow.
|
|
110
|
-
Response: 3xx (incl. chains/loops via edges), 4xx broken, 5xx, blocked-by-robots.
|
|
111
|
-
URL: >115 char, uppercase, underscores, params, non-ASCII, duplicate.
|
|
112
|
-
Content: thin (<X words), exact duplicate (hash collision).
|
|
113
|
-
Images: missing alt.
|
|
114
|
-
Structured data: present-but-invalid (shape), missing on key templates.
|
|
115
|
-
Links: broken internal/external, redirected internal, orphan pages.
|
|
116
|
-
|
|
117
|
-
---
|
|
118
|
-
|
|
119
|
-
## 2. Suggested data model (THE deliverable)
|
|
120
|
-
|
|
121
|
-
### 2.1 Extend `PageData` (`src/api/site-extractor.ts:11`)
|
|
122
|
-
Add these fields (keep existing):
|
|
123
|
-
|
|
124
|
-
```ts
|
|
125
|
-
// response/header-derived (require fetch-layer plumbing, Section 3.2)
|
|
126
|
-
contentType: string | null
|
|
127
|
-
responseTimeMs: number | null
|
|
128
|
-
sizeBytes: number | null
|
|
129
|
-
redirectUrl: string | null // Location on 3xx
|
|
130
|
-
lastModified: string | null
|
|
131
|
-
xRobotsTag: string | null
|
|
132
|
-
|
|
133
|
-
// on-page additions
|
|
134
|
-
titleLength: number | null
|
|
135
|
-
titlePixels: number | null // estimate, approximate
|
|
136
|
-
metaDescLength: number | null
|
|
137
|
-
metaKeywords: string | null
|
|
138
|
-
h1_2: string | null // second H1 if present
|
|
139
|
-
h2Count: number
|
|
140
|
-
metaRobots: string | null
|
|
141
|
-
relNext: string | null
|
|
142
|
-
relPrev: string | null
|
|
143
|
-
hreflang: Array<{ lang: string; href: string }>
|
|
144
|
-
og: { title?: string; description?: string; image?: string; type?: string } | null
|
|
145
|
-
twitter: { card?: string; title?: string; description?: string } | null
|
|
146
|
-
ampHref: string | null
|
|
147
|
-
|
|
148
|
-
// images
|
|
149
|
-
imageCount: number
|
|
150
|
-
imagesMissingAlt: number
|
|
151
|
-
|
|
152
|
-
// content
|
|
153
|
-
textRatio: number // bodyText.length / html.length
|
|
154
|
-
contentHash: string // sha1 of normalized visible text
|
|
155
|
-
|
|
156
|
-
// indexability (derived, see 3.3)
|
|
157
|
-
indexable: boolean
|
|
158
|
-
indexabilityReason: string | null // 'noindex' | 'canonicalised' | 'non-200' | 'x-robots-noindex' | null
|
|
159
|
-
|
|
160
|
-
// LINKS — replace the two counts with the edge list
|
|
161
|
-
outlinks: Array<{ href: string; anchor: string; rel: string | null; internal: boolean }>
|
|
162
|
-
// keep internalLinks/externalLinks as derived counts for back-compat
|
|
163
|
-
```
|
|
164
|
-
|
|
165
|
-
### 2.2 Link edge + graph model (post-crawl, computed)
|
|
166
|
-
```ts
|
|
167
|
-
interface LinkEdge { from: string; to: string; anchor: string; rel: string | null; internal: boolean }
|
|
168
|
-
|
|
169
|
-
interface PageLinkMetrics {
|
|
170
|
-
url: string
|
|
171
|
-
inlinks: number
|
|
172
|
-
uniqueInlinks: number
|
|
173
|
-
outlinksInternal: number
|
|
174
|
-
outlinksExternal: number
|
|
175
|
-
crawlDepth: number | null // BFS hops from startUrl over internal 200 edges; null = orphan/unreachable
|
|
176
|
-
orphan: boolean // crawled/in-sitemap but 0 internal inlinks
|
|
177
|
-
topAnchors: string[] // most common inbound anchor texts
|
|
178
|
-
}
|
|
179
|
-
```
|
|
180
|
-
|
|
181
|
-
### 2.3 Structured outputs written into the bulk folder
|
|
182
|
-
Alongside `index.md` + `pages/*.md`, add:
|
|
183
|
-
- `pages.jsonl` — one JSON line per page = full extended `PageData` minus `bodyMarkdown`/`schema`
|
|
184
|
-
(those stay in `pages/*.md`). This is the "Internal" tab.
|
|
185
|
-
- `links.jsonl` — one `LinkEdge` per line (internal + external). The "All Outlinks" export.
|
|
186
|
-
- `link-metrics.jsonl` — one `PageLinkMetrics` per URL. The inlinks/depth/orphan view.
|
|
187
|
-
- `issues.json` — the computed issue filters (Section 1.7) as `{ issueKey: { count, urls[] } }`.
|
|
188
|
-
- `report.md` — human summary (counts per issue, top offenders) = the audit deliverable.
|
|
189
|
-
|
|
190
|
-
---
|
|
191
|
-
|
|
192
|
-
## 3. Implementation blueprint (atomic)
|
|
193
|
-
|
|
194
|
-
### 3.1 Capture: extend `parsePageData` — `src/api/site-extractor.ts:45`
|
|
195
|
-
Signature change:
|
|
196
|
-
```ts
|
|
197
|
-
function parsePageData(
|
|
198
|
-
url: string, html: string, status: number, via: 'fetch'|'browser',
|
|
199
|
-
resp?: { headers?: Record<string,string>; responseTimeMs?: number; redirectUrl?: string|null }
|
|
200
|
-
): PageData
|
|
201
|
-
```
|
|
202
|
-
Add, inside the function, regex/derivations:
|
|
203
|
-
- `titleLength = title?.length`; `titlePixels = estimatePixels(title)` (new helper, char-width table).
|
|
204
|
-
- `metaDescLength`, `metaKeywords` via `<meta name="keywords">`.
|
|
205
|
-
- `h1_2`/`h2Count` from existing `headings` array.
|
|
206
|
-
- `metaRobots` via `<meta name="robots" content="...">`.
|
|
207
|
-
- `relNext`/`relPrev` via `<link rel="next|prev">`.
|
|
208
|
-
- `hreflang[]` via `<link rel="alternate" hreflang="..">`.
|
|
209
|
-
- `og`/`twitter` via `<meta property="og:.."|name="twitter:..">`.
|
|
210
|
-
- `ampHref` via `<link rel="amphtml">`.
|
|
211
|
-
- `imageCount` = matches of `<img`; `imagesMissingAlt` = `<img>` lacking non-empty `alt`.
|
|
212
|
-
- `textRatio = bodyText.length / Math.max(1, html.length)`.
|
|
213
|
-
- `contentHash = sha1(bodyText.replace(/\s+/g,' ').trim())` (`node:crypto`).
|
|
214
|
-
- header fields from `resp`: `contentType, xRobotsTag, lastModified, sizeBytes, responseTimeMs, redirectUrl`.
|
|
215
|
-
- `outlinks[]`: replace the count-only loop (`site-extractor.ts:102-110`) — for each `<a href>` also
|
|
216
|
-
capture anchor text and `rel`; classify `internal` by origin; keep `internalLinks`/`externalLinks`
|
|
217
|
-
as `.filter().length` for back-compat.
|
|
218
|
-
- `indexable`/`indexabilityReason`: `false` if status!=200, or metaRobots/xRobots contains `noindex`,
|
|
219
|
-
or canonical present and != self → reason set accordingly.
|
|
220
|
-
|
|
221
|
-
New helper `estimatePixels(s)` — sum per-char widths from a static map (approx Arial 'M'≈14, 'i'≈4…);
|
|
222
|
-
mark all pixel fields "approximate" in docs. ~25 lines.
|
|
223
|
-
|
|
224
|
-
### 3.2 Plumb headers/timing into both fetch paths
|
|
225
|
-
- `RotatingFetchResult` (`rotating-proxy-crawl.ts:5`): add
|
|
226
|
-
`headers?: Record<string,string>; responseTimeMs?: number; redirectUrl?: string|null`.
|
|
227
|
-
In `fetchBatch` (`:60-98`): time the `goto`, `resp.headers()`, capture `Location` when 3xx.
|
|
228
|
-
- `fetchAndParse` (`site-extractor.ts:115`): build the same `resp` object from `res.headers` + timing,
|
|
229
|
-
pass as 5th arg to `parsePageData`.
|
|
230
|
-
- In `extractSite` rotating branch (`site-extractor.ts:~185`): pass `r.headers/responseTimeMs/redirectUrl`
|
|
231
|
-
into `parsePageData(r.url, r.html, r.status, 'browser', {...})`.
|
|
232
|
-
|
|
233
|
-
### 3.3 Post-crawl link graph — new file `src/api/seo-link-graph.ts`
|
|
234
|
-
```ts
|
|
235
|
-
export function buildLinkGraph(pages: PageData[], startUrl: string):
|
|
236
|
-
{ edges: LinkEdge[]; metrics: Map<string, PageLinkMetrics> }
|
|
237
|
-
```
|
|
238
|
-
Logic:
|
|
239
|
-
- Flatten `pages[].outlinks` → `edges` (from = page.url).
|
|
240
|
-
- Build `inlinks` map by inverting internal edges; `uniqueInlinks` = distinct `from`.
|
|
241
|
-
- `crawlDepth`: BFS from `startUrl` over internal edges whose target status==200; unreached → null.
|
|
242
|
-
- `orphan`: page in set with 0 internal inlinks and url != startUrl.
|
|
243
|
-
- `topAnchors`: top-3 inbound anchors by frequency.
|
|
244
|
-
- Status join: map target→status from `pages` to flag broken (4xx/5xx) / redirect (3xx) edges.
|
|
245
|
-
|
|
246
|
-
### 3.4 Issue computation — new file `src/api/seo-issues.ts`
|
|
247
|
-
```ts
|
|
248
|
-
export function computeIssues(pages: PageData[], metrics: Map<string,PageLinkMetrics>):
|
|
249
|
-
Record<string, { count: number; urls: string[] }>
|
|
250
|
-
```
|
|
251
|
-
Thresholds (SF defaults): title >60/<30 char & >561px; meta >155/<70 & >985px; H1 >70; URL >115;
|
|
252
|
-
thin <200 words (configurable). Duplicates: group by exact `title`/`metaDescription`/`contentHash`,
|
|
253
|
-
flag groups size>1. Broken/redirect/orphan from `metrics`/edges. Indexability from `PageData`.
|
|
254
|
-
|
|
255
|
-
### 3.5 Write structured outputs — extend `saveBulkSite` (`mcp-response-formatter.ts:68`)
|
|
256
|
-
- Widen `BulkPage` → accept full extended `PageData` (or pass `PageData[]` directly).
|
|
257
|
-
- After writing `pages/*.md`, also `writeFileSync`:
|
|
258
|
-
- `pages.jsonl` (PageData minus bodyMarkdown/schema),
|
|
259
|
-
- `links.jsonl`, `link-metrics.jsonl`, `issues.json`, `report.md`.
|
|
260
|
-
- `formatExtractSite` already has full `pages`; pass them through (today it maps to the slim BulkPage —
|
|
261
|
-
change to pass the full objects, call `buildLinkGraph` + `computeIssues` before `saveBulkSite`).
|
|
262
|
-
- Return extra paths in the bulk summary + `structuredContent` (e.g. `pagesJsonl`, `issuesFile`).
|
|
263
|
-
|
|
264
|
-
### 3.6 The skill — `seo-crawl-audit`
|
|
265
|
-
Orchestrator (no new server code):
|
|
266
|
-
1. Call `extract_site(url, rotateProxies:true)` → folder path from response.
|
|
267
|
-
2. Read `pages.jsonl` + `issues.json` + `link-metrics.jsonl` from the folder.
|
|
268
|
-
3. Emit a prioritized SEO report: response-code breakdown, title/meta/H1 issues with offender lists,
|
|
269
|
-
indexability summary, thin/duplicate content, structured-data coverage, and the internal-link
|
|
270
|
-
section (orphans, deepest pages, most-linked, broken internal links).
|
|
271
|
-
4. Hand the full link graph to the existing `site-architecture-auditor` skill for equity/architecture
|
|
272
|
-
scoring (it already builds this graph from SF exports — we now feed it ours).
|
|
273
|
-
|
|
274
|
-
---
|
|
275
|
-
|
|
276
|
-
## 4. Build order (phased, each independently shippable)
|
|
277
|
-
1. **P1 — page fields (no headers):** 3.1 minus header fields + 3.5 `pages.jsonl`. Unlocks titles/metas/
|
|
278
|
-
H1/canonical/indexability(meta)/thin/dup/schema reports immediately.
|
|
279
|
-
2. **P2 — link graph:** 3.1 outlinks edge capture + 3.3 + `links.jsonl`/`link-metrics.jsonl`. Unlocks
|
|
280
|
-
inlinks/depth/orphans/broken-links — the user's headline ask.
|
|
281
|
-
3. **P3 — headers/timing:** 3.2 (both fetch paths). Adds content-type/size/response-time/X-Robots/redirects.
|
|
282
|
-
4. **P4 — issues + report:** 3.4 + `issues.json`/`report.md`.
|
|
283
|
-
5. **P5 — skill:** 3.6 + handoff to `site-architecture-auditor`.
|
|
284
|
-
|
|
285
|
-
## 5. Explicit non-goals (v1)
|
|
286
|
-
Per-asset image weight, microdata/RDFa, full structured-data validation, near-duplicate simhash,
|
|
287
|
-
JS-render diffing, pixel-width exactness (estimate only), redirect-chain hop-by-hop beyond one hop.
|