mcp-scraper 0.35.1 → 0.36.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +9 -9
- package/dist/bin/api-server.cjs +26406 -21015
- package/dist/bin/api-server.cjs.map +1 -1
- package/dist/bin/api-server.js +3 -3
- package/dist/bin/mcp-scraper-cli.cjs +51 -7
- package/dist/bin/mcp-scraper-cli.cjs.map +1 -1
- package/dist/bin/mcp-scraper-cli.js +48 -5
- package/dist/bin/mcp-scraper-cli.js.map +1 -1
- package/dist/bin/mcp-scraper-install.cjs +2 -2
- package/dist/bin/mcp-scraper-install.cjs.map +1 -1
- package/dist/bin/mcp-scraper-install.js +2 -2
- package/dist/bin/mcp-stdio-server.cjs +944 -215
- package/dist/bin/mcp-stdio-server.cjs.map +1 -1
- package/dist/bin/mcp-stdio-server.js +8 -8
- package/dist/bin/paa-harvest.cjs +125 -70
- package/dist/bin/paa-harvest.cjs.map +1 -1
- package/dist/bin/paa-harvest.js +4 -4
- package/dist/chunk-345BQXZH.js +712 -0
- package/dist/chunk-345BQXZH.js.map +1 -0
- package/dist/{chunk-M2S27J6Z.js → chunk-44HZLHDV.js} +10 -1
- package/dist/chunk-44HZLHDV.js.map +1 -0
- package/dist/chunk-4HO66323.js +7 -0
- package/dist/chunk-4HO66323.js.map +1 -0
- package/dist/{chunk-BWXLTWF7.js → chunk-4ZIJ3BKZ.js} +6 -4
- package/dist/chunk-4ZIJ3BKZ.js.map +1 -0
- package/dist/{chunk-NPMW5HUS.js → chunk-5RULXBJ7.js} +879 -237
- package/dist/chunk-5RULXBJ7.js.map +1 -0
- package/dist/chunk-AN3VQARU.js +684 -0
- package/dist/chunk-AN3VQARU.js.map +1 -0
- package/dist/{chunk-XVVNKASZ.js → chunk-ANCGXUQJ.js} +118 -73
- package/dist/chunk-ANCGXUQJ.js.map +1 -0
- package/dist/{chunk-3HBPKR5G.js → chunk-D7LM5QZN.js} +3 -3
- package/dist/{chunk-4ZB3X6BQ.js → chunk-E5UEELA7.js} +16 -2
- package/dist/{chunk-4ZB3X6BQ.js.map → chunk-E5UEELA7.js.map} +1 -1
- package/dist/{chunk-ZID3WQID.js → chunk-FQI5PFE7.js} +9 -71
- package/dist/chunk-FQI5PFE7.js.map +1 -0
- package/dist/chunk-G3P3ZDB4.js +69 -0
- package/dist/chunk-G3P3ZDB4.js.map +1 -0
- package/dist/{chunk-YRGSEY5L.js → chunk-G7KAVJ3F.js} +2 -2
- package/dist/{chunk-YRGSEY5L.js.map → chunk-G7KAVJ3F.js.map} +1 -1
- package/dist/{chunk-62DQAWPF.js → chunk-IFYER7O4.js} +367 -42
- package/dist/chunk-IFYER7O4.js.map +1 -0
- package/dist/chunk-O2MCWFXQ.js +499 -0
- package/dist/chunk-O2MCWFXQ.js.map +1 -0
- package/dist/chunk-QZXKQB7Y.js +414 -0
- package/dist/chunk-QZXKQB7Y.js.map +1 -0
- package/dist/{db-YAI5AQOI.js → db-N6MPVMEF.js} +8 -2
- package/dist/{extract-bundle-ONWZVV55.js → extract-bundle-M4SDJG3V.js} +284 -98
- package/dist/extract-bundle-M4SDJG3V.js.map +1 -0
- package/dist/index.cjs +129 -70
- package/dist/index.cjs.map +1 -1
- package/dist/index.d.cts +11 -0
- package/dist/index.d.ts +11 -0
- package/dist/index.js +4 -4
- package/dist/location-data-repository-Z4NQOU5Y.js +35 -0
- package/dist/{server-GKUTC73B.js → server-ZIGAFKOL.js} +10791 -7629
- package/dist/server-ZIGAFKOL.js.map +1 -0
- package/dist/site-extract-repository-PGQZNW6V.js +62 -0
- package/dist/site-extract-repository-PGQZNW6V.js.map +1 -0
- package/dist/{worker-645BZPEK.js → worker-EBB6CTGW.js} +7 -7
- package/docs/hosted-location-data.md +108 -0
- package/docs/mcp-tool-craft-lint.generated.md +6 -3
- package/docs/mcp-tool-manifest.generated.json +1253 -222
- package/docs/mcp-tool-quality-spec.md +1 -1
- package/docs/specs/connected-services-control-plane-decoupling-spec.md +1044 -0
- package/docs/specs/kernel-stealth-captcha-test-matrix.md +278 -0
- package/docs/specs/multimodal-image-memory-architecture-spec.md +1022 -0
- package/docs/specs/unified-credit-and-scheduled-execution-billing-spec.md +36 -27
- package/package.json +6 -5
- package/dist/chunk-62DQAWPF.js.map +0 -1
- package/dist/chunk-BWXLTWF7.js.map +0 -1
- package/dist/chunk-M2S27J6Z.js.map +0 -1
- package/dist/chunk-NPMW5HUS.js.map +0 -1
- package/dist/chunk-R7EETU7Z.js +0 -419
- package/dist/chunk-R7EETU7Z.js.map +0 -1
- package/dist/chunk-U44TPRST.js +0 -130
- package/dist/chunk-U44TPRST.js.map +0 -1
- package/dist/chunk-XVVNKASZ.js.map +0 -1
- package/dist/chunk-YR4LJ6AQ.js +0 -7
- package/dist/chunk-YR4LJ6AQ.js.map +0 -1
- package/dist/chunk-YV2FUEBX.js +0 -851
- package/dist/chunk-YV2FUEBX.js.map +0 -1
- package/dist/chunk-ZID3WQID.js.map +0 -1
- package/dist/extract-bundle-ONWZVV55.js.map +0 -1
- package/dist/server-GKUTC73B.js.map +0 -1
- package/dist/site-extract-repository-L6BHWVDU.js +0 -30
- /package/dist/{chunk-3HBPKR5G.js.map → chunk-D7LM5QZN.js.map} +0 -0
- /package/dist/{db-YAI5AQOI.js.map → db-N6MPVMEF.js.map} +0 -0
- /package/dist/{site-extract-repository-L6BHWVDU.js.map → location-data-repository-Z4NQOU5Y.js.map} +0 -0
- /package/dist/{worker-645BZPEK.js.map → worker-EBB6CTGW.js.map} +0 -0
package/README.md
CHANGED
|
@@ -39,7 +39,9 @@ npx -y -p mcp-scraper@latest mcp-scraper-cli agent install codex
|
|
|
39
39
|
npx -y -p mcp-scraper@latest mcp-scraper-cli agent prompt agent-packet
|
|
40
40
|
```
|
|
41
41
|
|
|
42
|
-
`agent install claude --apply` upserts the Claude Code user-scope `mcp-scraper` entry to `npx -y
|
|
42
|
+
`agent install claude --apply` upserts the Claude Code user-scope `mcp-scraper` entry to `npx -y --package mcp-scraper@latest mcp-scraper`. Fully exit Claude Code and open a new Claude terminal after applying; MCP servers are attached when Claude starts.
|
|
43
|
+
|
|
44
|
+
The registered command uses the long `--package` flag deliberately. Claude Code's `mcp add` leaks short flags that appear after `--` back into its own option parsing, so a registered `-p` makes it reject its own `--scope`/`-s` argument with a misleading `unknown option` error. If the registration ever fails, the previous entry is captured beforehand and restored automatically.
|
|
43
45
|
|
|
44
46
|
Check usage and upgrade concurrency from a normal terminal:
|
|
45
47
|
|
|
@@ -88,7 +90,7 @@ Build the branded one-click bundle:
|
|
|
88
90
|
npm run build:mcpb
|
|
89
91
|
```
|
|
90
92
|
|
|
91
|
-
The generated bundle is written to `build/mcpb/mcp-scraper-<version>.mcpb` and copied to `public/downloads/` for the hosted download. The current public bundle is `https://mcpscraper.dev/downloads/mcp-scraper.mcpb` (`0.
|
|
93
|
+
The generated bundle is written to `build/mcpb/mcp-scraper-<version>.mcpb` and copied to `public/downloads/` for the hosted download. The current public bundle is `https://mcpscraper.dev/downloads/mcp-scraper.mcpb` (`0.34.1`, SHA-256 `b217a7273ae7ee406003619de7f4bf95b4c8ad241c7f723244187c1bb557d947`). Install it by opening or dragging it into Claude Desktop. Claude displays the `MCP Scraper` install card, icon, and API-key configuration field from the bundle manifest.
|
|
92
94
|
|
|
93
95
|
The MCPB install exposes every tool — web-intelligence plus all `browser_*` tools — through the one `mcp-scraper` server.
|
|
94
96
|
|
|
@@ -152,7 +154,7 @@ env = { MCP_SCRAPER_API_KEY = "sk_live_your_key" }
|
|
|
152
154
|
|
|
153
155
|
- `harvest_paa`
|
|
154
156
|
- `search_serp`
|
|
155
|
-
- `extract_url`
|
|
157
|
+
- `extract_url` — extract normal or Wayback-replayed page copy; Wayback results omit playback chrome and can include a timestamp-matched featured image.
|
|
156
158
|
- `map_site_urls`
|
|
157
159
|
- `map_wayback_snapshots` — count and inventory Wayback captures across an inclusive date range without downloading page bodies. Supports exact pages, prefixes, hosts, domains, or selected URLs; reports exact versus lower-bound counts, unique URLs/content digests, monthly coverage, missing months, and optional timestamp rows.
|
|
158
160
|
- `extract_site` — crawl a live site, batch one archived site snapshot from a Wayback replay URL, or pass a `wayback` plan for whole-site, single-page, or selected-page timelines across explicit months or a `from`/`to` range. Timeline ZIPs include month folders and a capture matrix.
|
|
@@ -166,7 +168,7 @@ env = { MCP_SCRAPER_API_KEY = "sk_live_your_key" }
|
|
|
166
168
|
- `instagram_media_download` — extract and download one Instagram post/reel/tv URL, optionally through a saved hosted browser `profile` for authenticated access. Returns text/caption, image URL/downloads, selected video/audio MP4 tracks, optional muxed MP4 when `ffmpeg` is available, optional transcript, and browser details.
|
|
167
169
|
- `maps_search` — search Google's localized local-results list for multiple business/profile candidates. Use for GMB/GBP prospect lists, competitors, categories, and anything needing more than the Google 3-pack. It opens the rendered business card, reads the profile dialog, then closes it before continuing to the next ranked card. Set `includeServices: true` to return services and areas served without collecting review cards. `maxResults` defaults to 10 and is capped at 50.
|
|
168
170
|
- `maps_place_intel` — hydrate one known/named Google Maps business with profile details and optional reviews. Use after `maps_search` when a selected candidate needs full details.
|
|
169
|
-
- `directory_workflow` — build city-by-city directory/prospecting datasets from Census place selection plus localized Google business searches. Use it for requests like "all cities over 100k population in Tennessee, then get 20 roofers from Maps."
|
|
171
|
+
- `directory_workflow` — build city-by-city directory/prospecting datasets from Census place selection plus localized Google business searches. Use it for requests like "all cities over 100k population in Tennessee, then get 20 roofers from Maps." Supply the business category, state, and market limits; MCP Scraper manages search transport and retry behavior internally. The saved CSV includes `source_location`, `result_position`, `business_name`, `review_stars`, `review_count`, `category`, `address`, `phone`, `hours_status`, `website_url`, `directions_url`, `place_url`, `cid`, `cid_decimal`, Census population, and ZIP groups.
|
|
170
172
|
- `workflow_list` — list higher-level workflow IDs plus AI-facing recipes for market analysis, ICP research, forum/review acquisition, brand design briefings, CRO audits, positioning briefs, content gaps, and AI search visibility audits.
|
|
171
173
|
- `workflow_suggest` — route a high-level business goal to the right workflow/tool chain before spending credits.
|
|
172
174
|
- `workflow_run` — run hosted workflows such as `agent-packet`, `local-competitive-audit`, `map-comparison`, `serp-comparison`, `paa-expansion-brief`, and `ai-overview-language`; returns run metadata, summary, and artifact IDs.
|
|
@@ -179,7 +181,7 @@ env = { MCP_SCRAPER_API_KEY = "sk_live_your_key" }
|
|
|
179
181
|
|
|
180
182
|
- `list_service_connections` — list this caller's tenant-owned Nango OAuth and official remote MCP connections, including verified provider-side account email/name when exposed, exact live reads, gated actions, permanently blocked administrative tools, credential transport, and schema-discovery metadata. Provider identity is distinct from the MCP Scraper login, and connections are never shared between customers.
|
|
181
183
|
- `describe_service_connection_tool` — fetch the sanitized live MCP Tool definition for one tool listed on one tenant-owned connection, including its current callability, input schema, optional output schema, safe annotations, and schema hash. Use this before constructing provider-native arguments; provider functions stay behind the generic bridges instead of becoming dozens of permanent top-level tools.
|
|
182
|
-
- `export_connected_service_data` — fetch a fresh bounded Gmail, Google Calendar, Google Search Console, Zoom, Resend, or Meta time range in one MCP call. Search Console's `search_console_performance` dataset walks accessible properties and bounded live Search Analytics pages with signed continuation. Small exports return inline; larger exports become private JSONL retained for seven days with a 15-minute signed URL. For relationship work, gather source evidence first: inspect existing People records, resolve the exact provider account, use RFC3339 `from`/`to` for Gmail ranges longer than 90 days, preserve provider provenance when writing a linked Communication, and never treat an export as permission to mutate the source account.
|
|
184
|
+
- `export_connected_service_data` — fetch a fresh bounded Gmail, Google Calendar, Google Search Console, Zoom, Resend, or Meta time range in one MCP call. Zoom's `zoom_transcripts` dataset resolves VTT files from recording metadata and downloads them through the authenticated connection without looping the separately rate-limited `get-meeting-transcript` function. Search Console's `search_console_performance` dataset walks accessible properties and bounded live Search Analytics pages with signed continuation. Small exports return inline; larger exports become private JSONL retained for seven days with a 15-minute signed URL. For relationship work, gather source evidence first: inspect existing People records, resolve the exact provider account, use RFC3339 `from`/`to` for Gmail ranges longer than 90 days, preserve provider provenance when writing a linked Communication, and never treat an export as permission to mutate the source account.
|
|
183
185
|
- `export_search_console_table_data` — filter up to 50,000 Search Console rows already persisted by a scheduled `connection_sync` and create a private renewable JSONL artifact without calling Google again. Get the typed `gsc_performance_*` table name from `list_service_connections`, inspect it with `table-describe`, and use the same filters with `table-query` for interactive analysis.
|
|
184
186
|
- `renew_connected_data_download` — issue a fresh 15-minute signed URL for an unexpired private export artifact without pulling the provider again.
|
|
185
187
|
- `read_service_connection` — run one small live read by exact allowlisted name across Nango OAuth or official remote MCP connections, including bounded Google Drive inventory, change, Doc, Sheet, and text-file tools. Do not loop it over a time range when `export_connected_service_data` supports that provider's collection.
|
|
@@ -220,15 +222,13 @@ Google Search Console exposes eight bounded reads and eight gated property and s
|
|
|
220
222
|
|
|
221
223
|
For accurate annotated videos, do not guess annotation times from a script. Start the replay, navigate until each target is visible and stable, call `browser_replay_mark` for each callout, then stop the replay and pass the returned annotations to `browser_replay_annotate` with the returned `source_width` and `source_height`.
|
|
222
224
|
|
|
223
|
-
For
|
|
224
|
-
|
|
225
|
-
For Google Maps tools (`maps_search` and `directory_workflow`), leave `proxyMode` unset for the default direct Google route. Localization is carried by the city in the query, UULE, `gl`, and `hl`; do not treat proxy geolocation as the source of truth. Use `proxyMode: "location"` only when an explicit residential-proxy experiment or compatibility path needs it. Retryable failures open a new browser session; location mode can additionally rotate its disposable proxy. Successful structured responses include sanitized attempt telemetry without exposing full proxy or browser IDs.
|
|
225
|
+
For Google SERP and Maps tools, callers provide the query, two-letter country code (`gl`), language (`hl`), device when relevant, and an optional city or region. MCP Scraper owns transport selection, anti-bot handling, and bounded retries internally; those implementation controls and receipts are not part of the public tool contract.
|
|
226
226
|
|
|
227
227
|
The `mcp-scraper` server (and the MCPB bundle, which runs it) exposes both sections through one MCP server.
|
|
228
228
|
|
|
229
229
|
All MCP tools expose output schemas and return `structuredContent` with the IDs, URLs, CSV paths, transcripts, browser session handles, replay paths, artifacts, recipe fields, or blueprint fields needed by the next step. Browser Agent tools keep a JSON text block for older clients, but structured data is the primary contract. All tools carry MCP annotations; file-writing tools such as replay downloads and annotations state their filesystem side effects.
|
|
230
230
|
|
|
231
|
-
The canonical tool inventory is generated at `docs/mcp-tool-manifest.generated.json`.
|
|
231
|
+
The canonical tool inventory is generated at `docs/mcp-tool-manifest.generated.json`. The current local unified server exposes 170 tools: 81 scraper, browser, workflow, billing, and connected-service tools plus 89 durable-memory tools. Release verification compares the exact local and hosted tool-name sets, not only the count.
|
|
232
232
|
|
|
233
233
|
For contract parity, stdio and MCPB memory calls invoke the matching public tool on the hosted MCP Scraper `/mcp` endpoint. The hosted aggregate runtime owns MCP Scraper-specific billing, scheduling, credential, and in-process cutover policy; its internal `/memory/mcp-call` bridge is a fallback to the standalone memory service, not the public stdio execution path. Direct `mcp-memory` OAuth and stdio clients continue to use `memory.mcpscraper.dev` and must be verified as a separate dependent release surface.
|
|
234
234
|
|