mcp-scraper 0.51.1 → 0.52.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. package/README.md +4 -4
  2. package/dist/bin/api-server.cjs +2912 -904
  3. package/dist/bin/api-server.cjs.map +1 -1
  4. package/dist/bin/api-server.js +2 -2
  5. package/dist/bin/mcp-scraper-cli.cjs +5 -1
  6. package/dist/bin/mcp-scraper-cli.cjs.map +1 -1
  7. package/dist/bin/mcp-scraper-cli.js +3 -3
  8. package/dist/bin/mcp-scraper-install.cjs +1 -1
  9. package/dist/bin/mcp-scraper-install.cjs.map +1 -1
  10. package/dist/bin/mcp-scraper-install.js +1 -1
  11. package/dist/bin/mcp-stdio-server.cjs +384 -43
  12. package/dist/bin/mcp-stdio-server.cjs.map +1 -1
  13. package/dist/bin/mcp-stdio-server.js +2 -2
  14. package/dist/bin/paa-harvest.cjs +4 -0
  15. package/dist/bin/paa-harvest.cjs.map +1 -1
  16. package/dist/bin/paa-harvest.js +2 -2
  17. package/dist/{chunk-ZSJHRZY5.js → chunk-J7TH5KU7.js} +2 -2
  18. package/dist/{chunk-2K4RCBYY.js → chunk-OT2AF7SH.js} +385 -44
  19. package/dist/chunk-OT2AF7SH.js.map +1 -0
  20. package/dist/chunk-SG2PEHR3.js +562 -0
  21. package/dist/chunk-SG2PEHR3.js.map +1 -0
  22. package/dist/{chunk-GOZIG6HD.js → chunk-TVZD37SE.js} +2 -2
  23. package/dist/{chunk-3ZUBQQPQ.js → chunk-V7CEOUBF.js} +5 -1
  24. package/dist/chunk-V7CEOUBF.js.map +1 -0
  25. package/dist/chunk-VIAXQTZK.js +7 -0
  26. package/dist/chunk-VIAXQTZK.js.map +1 -0
  27. package/dist/{extract-bundle-RCTNANCH.js → extract-bundle-GUEUCTAE.js} +2 -2
  28. package/dist/index.cjs +4 -0
  29. package/dist/index.cjs.map +1 -1
  30. package/dist/index.js +2 -2
  31. package/dist/{server-JR5XPTXS.js → server-7EXEDKAU.js} +1920 -420
  32. package/dist/server-7EXEDKAU.js.map +1 -0
  33. package/dist/{worker-KQN673JF.js → worker-JL4TG6IF.js} +3 -3
  34. package/package.json +1 -1
  35. package/dist/chunk-27FMOD6S.js +0 -430
  36. package/dist/chunk-27FMOD6S.js.map +0 -1
  37. package/dist/chunk-2K4RCBYY.js.map +0 -1
  38. package/dist/chunk-3ZUBQQPQ.js.map +0 -1
  39. package/dist/chunk-XLUITZUM.js +0 -7
  40. package/dist/chunk-XLUITZUM.js.map +0 -1
  41. package/dist/server-JR5XPTXS.js.map +0 -1
  42. /package/dist/{chunk-ZSJHRZY5.js.map → chunk-J7TH5KU7.js.map} +0 -0
  43. /package/dist/{chunk-GOZIG6HD.js.map → chunk-TVZD37SE.js.map} +0 -0
  44. /package/dist/{extract-bundle-RCTNANCH.js.map → extract-bundle-GUEUCTAE.js.map} +0 -0
  45. /package/dist/{worker-KQN673JF.js.map → worker-JL4TG6IF.js.map} +0 -0
package/README.md CHANGED
@@ -90,7 +90,7 @@ Build the branded one-click bundle:
90
90
  npm run build:mcpb
91
91
  ```
92
92
 
93
- The generated bundle is written to `build/mcpb/mcp-scraper-<version>.mcpb` and copied to `public/downloads/` for the hosted download. The current public bundle is `https://mcpscraper.dev/downloads/mcp-scraper.mcpb` (`0.51.1`, SHA-256 `c460ad09a5fbe65088078feeede7c7f49a3af83b3ed6b5228be4f618e0479425`). Install it by opening or dragging it into Claude Desktop. Claude displays the `MCP Scraper` install card, icon, and API-key configuration field from the bundle manifest.
93
+ The generated bundle is written to `build/mcpb/mcp-scraper-<version>.mcpb` and copied to `public/downloads/` for the hosted download. The current public bundle is `https://mcpscraper.dev/downloads/mcp-scraper.mcpb` (`0.52.0`, SHA-256 `d50498579a60ce32fc12effa56fb70aa30ba873a970055d65a61aa3bdb4f4886`). Install it by opening or dragging it into Claude Desktop. Claude displays the `MCP Scraper` install card, icon, and API-key configuration field from the bundle manifest.
94
94
 
95
95
  The MCPB install exposes every tool — web-intelligence plus all `browser_*` tools — through the one `mcp-scraper` server.
96
96
 
@@ -154,7 +154,7 @@ env = { MCP_SCRAPER_API_KEY = "sk_live_your_key" }
154
154
 
155
155
  - `harvest_paa`
156
156
  - `search_serp`
157
- - `extract_url` — extract normal or Wayback-replayed page copy; Wayback results omit playback chrome and can include a timestamp-matched featured image.
157
+ - `extract_url` — extract normal or Wayback-replayed page copy; Wayback results omit playback chrome and can include a timestamp-matched featured image. Set `preserveMedia:true` to union static and rendered/lazy media, collapse responsive variants, attach up to `maxInlineImages` AI-readable images, and receive an owner-scoped ZIP manifest readable with `archive_read`. Branding output ranks the site logo separately from evidence-bounded proof images such as certifications, awards, memberships, partner/customer marks, and press mentions.
158
158
  - `map_site_urls`
159
159
  - `map_wayback_snapshots` — count and inventory Wayback captures across an inclusive date range without downloading page bodies. Supports exact pages, prefixes, hosts, domains, or selected URLs; reports exact versus lower-bound counts, unique URLs/content digests, monthly coverage, missing months, and optional timestamp rows.
160
160
  - `extract_site` — crawl a live site, batch one archived site snapshot from a Wayback replay URL, or pass a `wayback` plan for whole-site, single-page, or selected-page timelines across explicit months or a `from`/`to` range. Timeline ZIPs include month folders and a capture matrix.
@@ -168,7 +168,7 @@ env = { MCP_SCRAPER_API_KEY = "sk_live_your_key" }
168
168
  - `instagram_profile_content` — discover Instagram profile grid content links for a handle or profile URL, optionally through a saved hosted browser `profile` for authenticated access. Returns collected post/reel/tv URLs, profile counts, type counts, shortcodes, browser details, pagination attempts, stop reason, and limitations.
169
169
  - `instagram_media_download` — extract and download one Instagram post/reel/tv URL, optionally through a saved hosted browser `profile` for authenticated access. Returns text/caption, image URL/downloads, selected video/audio MP4 tracks, optional muxed MP4 when `ffmpeg` is available, optional transcript, and browser details.
170
170
  - `maps_search` — search Google's localized local-results list for multiple business/profile candidates. Use for GMB/GBP prospect lists, competitors, categories, and anything needing more than the Google 3-pack. It opens the rendered business card, reads the profile dialog, then closes it before continuing to the next ranked card. Set `includeServices: true` to return services and areas served without collecting review cards. `maxResults` defaults to 10 and is capped at 50.
171
- - `maps_place_intel` — hydrate one known/named Google Maps business with profile details and optional reviews. Use after `maps_search` when a selected candidate needs full details.
171
+ - `maps_place_intel` — hydrate one known/named Google Maps business with profile details, entity IDs/CID, services, service areas, review aggregates/cards, and optional photos. Set `includeImages:true`, choose `imageScope:"owner"` or `"all"`, and tune `maxImages`; results include ownership evidence, completion state, bounded AI image blocks, and an owner-scoped ZIP manifest readable with `archive_read`.
172
172
  - `directory_workflow` — build city-by-city directory/prospecting datasets from Census place selection plus localized Google business searches. Use it for requests like "all cities over 100k population in Tennessee, then get 20 roofers from Maps." Supply the business category, state, and market limits; MCP Scraper manages search transport and retry behavior internally. The saved CSV includes `source_location`, `result_position`, `business_name`, `review_stars`, `review_count`, `category`, `address`, `phone`, `hours_status`, `website_url`, `directions_url`, `place_url`, `cid`, `cid_decimal`, Census population, and ZIP groups.
173
173
  - `workflow_list` — list higher-level workflow IDs plus AI-facing recipes for market analysis, ICP research, forum/review acquisition, brand design briefings, CRO audits, positioning briefs, content gaps, and AI search visibility audits.
174
174
  - `workflow_suggest` — route a high-level business goal to the right workflow/tool chain before spending credits.
@@ -210,7 +210,7 @@ Google Search Console exposes eight bounded reads and eight gated property and s
210
210
 
211
211
  - Start unknown-location recall with `memory-search`. It uses Hybrid Smart RAG and returns ranked excerpts, not complete notes.
212
212
  - Read strong candidates with `memory-get` before relying on, summarizing, linking, or editing them.
213
- - Use `memory-list` for an exhaustive vault inventory. It returns every note's metadata plus a sorted list of every represented folder and nested folder.
213
+ - Use `memory-list` for an exhaustive inventory. Pass `allVaults:true` for every note and folder across the whole entitled Memory account; omit it for one logical vault. Account-wide results include aggregate totals and complete per-vault note/folder groups, so a default-vault count is never mistaken for the account total.
214
214
  - Use `list-memory-tags` for the complete account-wide canonical tag vocabulary, aliases, usage counts, and per-vault distribution.
215
215
  - `memory-put` replaces the complete note body. For edits, read the note, merge the requested change into the full body, write with `baseRevision`, then read back and verify it.
216
216