mcp-scraper 0.51.2 → 0.52.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +3 -3
- package/dist/bin/api-server.cjs +2836 -882
- package/dist/bin/api-server.cjs.map +1 -1
- package/dist/bin/api-server.js +2 -2
- package/dist/bin/mcp-scraper-cli.cjs +5 -1
- package/dist/bin/mcp-scraper-cli.cjs.map +1 -1
- package/dist/bin/mcp-scraper-cli.js +3 -3
- package/dist/bin/mcp-scraper-install.cjs +1 -1
- package/dist/bin/mcp-scraper-install.cjs.map +1 -1
- package/dist/bin/mcp-scraper-install.js +1 -1
- package/dist/bin/mcp-stdio-server.cjs +372 -41
- package/dist/bin/mcp-stdio-server.cjs.map +1 -1
- package/dist/bin/mcp-stdio-server.js +2 -2
- package/dist/bin/paa-harvest.cjs +4 -0
- package/dist/bin/paa-harvest.cjs.map +1 -1
- package/dist/bin/paa-harvest.js +2 -2
- package/dist/chunk-EUBO6E43.js +7 -0
- package/dist/chunk-EUBO6E43.js.map +1 -0
- package/dist/{chunk-ZSJHRZY5.js → chunk-J7TH5KU7.js} +2 -2
- package/dist/{chunk-FCZ3QCZP.js → chunk-R6NOADCR.js} +373 -42
- package/dist/chunk-R6NOADCR.js.map +1 -0
- package/dist/chunk-SG2PEHR3.js +562 -0
- package/dist/chunk-SG2PEHR3.js.map +1 -0
- package/dist/{chunk-GOZIG6HD.js → chunk-TVZD37SE.js} +2 -2
- package/dist/{chunk-3ZUBQQPQ.js → chunk-V7CEOUBF.js} +5 -1
- package/dist/chunk-V7CEOUBF.js.map +1 -0
- package/dist/{extract-bundle-RCTNANCH.js → extract-bundle-GUEUCTAE.js} +2 -2
- package/dist/index.cjs +4 -0
- package/dist/index.cjs.map +1 -1
- package/dist/index.js +2 -2
- package/dist/{server-KAEHVDY7.js → server-HHHZD6Q6.js} +1857 -401
- package/dist/server-HHHZD6Q6.js.map +1 -0
- package/dist/{worker-KQN673JF.js → worker-JL4TG6IF.js} +3 -3
- package/package.json +2 -2
- package/dist/chunk-27FMOD6S.js +0 -430
- package/dist/chunk-27FMOD6S.js.map +0 -1
- package/dist/chunk-3ZUBQQPQ.js.map +0 -1
- package/dist/chunk-FCZ3QCZP.js.map +0 -1
- package/dist/chunk-JNRSR5ZJ.js +0 -7
- package/dist/chunk-JNRSR5ZJ.js.map +0 -1
- package/dist/server-KAEHVDY7.js.map +0 -1
- /package/dist/{chunk-ZSJHRZY5.js.map → chunk-J7TH5KU7.js.map} +0 -0
- /package/dist/{chunk-GOZIG6HD.js.map → chunk-TVZD37SE.js.map} +0 -0
- /package/dist/{extract-bundle-RCTNANCH.js.map → extract-bundle-GUEUCTAE.js.map} +0 -0
- /package/dist/{worker-KQN673JF.js.map → worker-JL4TG6IF.js.map} +0 -0
package/README.md
CHANGED
|
@@ -90,7 +90,7 @@ Build the branded one-click bundle:
|
|
|
90
90
|
npm run build:mcpb
|
|
91
91
|
```
|
|
92
92
|
|
|
93
|
-
The generated bundle is written to `build/mcpb/mcp-scraper-<version>.mcpb` and copied to `public/downloads/` for the hosted download. The current public bundle is `https://mcpscraper.dev/downloads/mcp-scraper.mcpb` (`0.
|
|
93
|
+
The generated bundle is written to `build/mcpb/mcp-scraper-<version>.mcpb` and copied to `public/downloads/` for the hosted download. The current public bundle is `https://mcpscraper.dev/downloads/mcp-scraper.mcpb` (`0.52.1`, SHA-256 `02db181ab00dbc3629a30fbb054313019464cede7d959a6902e7637d09ae40ca`). Install it by opening or dragging it into Claude Desktop. Claude displays the `MCP Scraper` install card, icon, and API-key configuration field from the bundle manifest.
|
|
94
94
|
|
|
95
95
|
The MCPB install exposes every tool — web-intelligence plus all `browser_*` tools — through the one `mcp-scraper` server.
|
|
96
96
|
|
|
@@ -154,7 +154,7 @@ env = { MCP_SCRAPER_API_KEY = "sk_live_your_key" }
|
|
|
154
154
|
|
|
155
155
|
- `harvest_paa`
|
|
156
156
|
- `search_serp`
|
|
157
|
-
- `extract_url` — extract normal or Wayback-replayed page copy; Wayback results omit playback chrome and can include a timestamp-matched featured image.
|
|
157
|
+
- `extract_url` — extract normal or Wayback-replayed page copy; Wayback results omit playback chrome and can include a timestamp-matched featured image. Set `preserveMedia:true` to union static and rendered/lazy media, collapse responsive variants, attach up to `maxInlineImages` AI-readable images, and receive an owner-scoped ZIP manifest readable with `archive_read`. Branding output ranks the site logo separately from evidence-bounded proof images such as certifications, awards, memberships, partner/customer marks, and press mentions.
|
|
158
158
|
- `map_site_urls`
|
|
159
159
|
- `map_wayback_snapshots` — count and inventory Wayback captures across an inclusive date range without downloading page bodies. Supports exact pages, prefixes, hosts, domains, or selected URLs; reports exact versus lower-bound counts, unique URLs/content digests, monthly coverage, missing months, and optional timestamp rows.
|
|
160
160
|
- `extract_site` — crawl a live site, batch one archived site snapshot from a Wayback replay URL, or pass a `wayback` plan for whole-site, single-page, or selected-page timelines across explicit months or a `from`/`to` range. Timeline ZIPs include month folders and a capture matrix.
|
|
@@ -168,7 +168,7 @@ env = { MCP_SCRAPER_API_KEY = "sk_live_your_key" }
|
|
|
168
168
|
- `instagram_profile_content` — discover Instagram profile grid content links for a handle or profile URL, optionally through a saved hosted browser `profile` for authenticated access. Returns collected post/reel/tv URLs, profile counts, type counts, shortcodes, browser details, pagination attempts, stop reason, and limitations.
|
|
169
169
|
- `instagram_media_download` — extract and download one Instagram post/reel/tv URL, optionally through a saved hosted browser `profile` for authenticated access. Returns text/caption, image URL/downloads, selected video/audio MP4 tracks, optional muxed MP4 when `ffmpeg` is available, optional transcript, and browser details.
|
|
170
170
|
- `maps_search` — search Google's localized local-results list for multiple business/profile candidates. Use for GMB/GBP prospect lists, competitors, categories, and anything needing more than the Google 3-pack. It opens the rendered business card, reads the profile dialog, then closes it before continuing to the next ranked card. Set `includeServices: true` to return services and areas served without collecting review cards. `maxResults` defaults to 10 and is capped at 50.
|
|
171
|
-
- `maps_place_intel` — hydrate one known/named Google Maps business with profile details and optional
|
|
171
|
+
- `maps_place_intel` — hydrate one known/named Google Maps business with profile details, entity IDs/CID, services, service areas, review aggregates/cards, and optional photos. Set `includeImages:true`, choose `imageScope:"owner"` or `"all"`, and tune `maxImages`; results include ownership evidence, completion state, bounded AI image blocks, and an owner-scoped ZIP manifest readable with `archive_read`.
|
|
172
172
|
- `directory_workflow` — build city-by-city directory/prospecting datasets from Census place selection plus localized Google business searches. Use it for requests like "all cities over 100k population in Tennessee, then get 20 roofers from Maps." Supply the business category, state, and market limits; MCP Scraper manages search transport and retry behavior internally. The saved CSV includes `source_location`, `result_position`, `business_name`, `review_stars`, `review_count`, `category`, `address`, `phone`, `hours_status`, `website_url`, `directions_url`, `place_url`, `cid`, `cid_decimal`, Census population, and ZIP groups.
|
|
173
173
|
- `workflow_list` — list higher-level workflow IDs plus AI-facing recipes for market analysis, ICP research, forum/review acquisition, brand design briefings, CRO audits, positioning briefs, content gaps, and AI search visibility audits.
|
|
174
174
|
- `workflow_suggest` — route a high-level business goal to the right workflow/tool chain before spending credits.
|