mcp-scraper 0.40.2 → 0.41.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (62) hide show
  1. package/README.md +2 -2
  2. package/dist/bin/api-server.cjs +2264 -337
  3. package/dist/bin/api-server.cjs.map +1 -1
  4. package/dist/bin/api-server.js +3 -3
  5. package/dist/bin/mcp-scraper-cli.cjs +1 -1
  6. package/dist/bin/mcp-scraper-cli.cjs.map +1 -1
  7. package/dist/bin/mcp-scraper-cli.js +1 -1
  8. package/dist/bin/mcp-scraper-install.cjs +1 -1
  9. package/dist/bin/mcp-scraper-install.cjs.map +1 -1
  10. package/dist/bin/mcp-scraper-install.js +1 -1
  11. package/dist/bin/mcp-stdio-server.cjs +180 -92
  12. package/dist/bin/mcp-stdio-server.cjs.map +1 -1
  13. package/dist/bin/mcp-stdio-server.js +6 -5
  14. package/dist/bin/mcp-stdio-server.js.map +1 -1
  15. package/dist/bin/paa-harvest.cjs +31 -9
  16. package/dist/bin/paa-harvest.cjs.map +1 -1
  17. package/dist/bin/paa-harvest.js +4 -4
  18. package/dist/{chunk-XORPNO3Z.js → chunk-27DUCRAZ.js} +2 -2
  19. package/dist/{chunk-44HZLHDV.js → chunk-5UN33CGU.js} +38 -1
  20. package/dist/chunk-5UN33CGU.js.map +1 -0
  21. package/dist/{chunk-X2LKCX6H.js → chunk-6J57U6HA.js} +3 -3
  22. package/dist/chunk-BBI7RGOT.js +99 -0
  23. package/dist/chunk-BBI7RGOT.js.map +1 -0
  24. package/dist/{chunk-NVXNEOUQ.js → chunk-CS4HE6QY.js} +47 -11
  25. package/dist/chunk-CS4HE6QY.js.map +1 -0
  26. package/dist/{chunk-3PIWJS6Y.js → chunk-EHES33KB.js} +2 -2
  27. package/dist/{chunk-GUVKHCKE.js → chunk-IQTT7CAB.js} +94 -3
  28. package/dist/chunk-IQTT7CAB.js.map +1 -0
  29. package/dist/{chunk-7XBBFBYY.js → chunk-JVIW4GK2.js} +44 -12
  30. package/dist/chunk-JVIW4GK2.js.map +1 -0
  31. package/dist/chunk-OOB35KFT.js +7 -0
  32. package/dist/chunk-OOB35KFT.js.map +1 -0
  33. package/dist/{chunk-LP6E462I.js → chunk-SIE5LZ2V.js} +74 -106
  34. package/dist/chunk-SIE5LZ2V.js.map +1 -0
  35. package/dist/{db-W3CP562I.js → db-LVVU6NOK.js} +16 -2
  36. package/dist/{extract-bundle-K4PG3RZJ.js → extract-bundle-SKXEDG4Z.js} +4 -4
  37. package/dist/index.cjs +31 -9
  38. package/dist/index.cjs.map +1 -1
  39. package/dist/index.js +4 -4
  40. package/dist/{location-data-repository-RLQX6SNM.js → location-data-repository-NBTCOA7B.js} +3 -3
  41. package/dist/{server-KM3CFWCF.js → server-2JJPCZH4.js} +1812 -162
  42. package/dist/server-2JJPCZH4.js.map +1 -0
  43. package/dist/{site-extract-repository-2SMMFKKL.js → site-extract-repository-6GX72U4N.js} +4 -4
  44. package/dist/{worker-FXAGFYOE.js → worker-5RWLFSAJ.js} +27 -28
  45. package/dist/worker-5RWLFSAJ.js.map +1 -0
  46. package/package.json +1 -1
  47. package/dist/chunk-44HZLHDV.js.map +0 -1
  48. package/dist/chunk-7XBBFBYY.js.map +0 -1
  49. package/dist/chunk-GUVKHCKE.js.map +0 -1
  50. package/dist/chunk-LP6E462I.js.map +0 -1
  51. package/dist/chunk-NVXNEOUQ.js.map +0 -1
  52. package/dist/chunk-ZUJLSICT.js +0 -7
  53. package/dist/chunk-ZUJLSICT.js.map +0 -1
  54. package/dist/server-KM3CFWCF.js.map +0 -1
  55. package/dist/worker-FXAGFYOE.js.map +0 -1
  56. /package/dist/{chunk-XORPNO3Z.js.map → chunk-27DUCRAZ.js.map} +0 -0
  57. /package/dist/{chunk-X2LKCX6H.js.map → chunk-6J57U6HA.js.map} +0 -0
  58. /package/dist/{chunk-3PIWJS6Y.js.map → chunk-EHES33KB.js.map} +0 -0
  59. /package/dist/{db-W3CP562I.js.map → db-LVVU6NOK.js.map} +0 -0
  60. /package/dist/{extract-bundle-K4PG3RZJ.js.map → extract-bundle-SKXEDG4Z.js.map} +0 -0
  61. /package/dist/{location-data-repository-RLQX6SNM.js.map → location-data-repository-NBTCOA7B.js.map} +0 -0
  62. /package/dist/{site-extract-repository-2SMMFKKL.js.map → site-extract-repository-6GX72U4N.js.map} +0 -0
package/README.md CHANGED
@@ -90,7 +90,7 @@ Build the branded one-click bundle:
90
90
  npm run build:mcpb
91
91
  ```
92
92
 
93
- The generated bundle is written to `build/mcpb/mcp-scraper-<version>.mcpb` and copied to `public/downloads/` for the hosted download. The current public bundle is `https://mcpscraper.dev/downloads/mcp-scraper.mcpb` (`0.40.2`, SHA-256 `a1047098851d810debdfed2dd1c1dc0334a930367746667fb3370ecc2a762b04`). Install it by opening or dragging it into Claude Desktop. Claude displays the `MCP Scraper` install card, icon, and API-key configuration field from the bundle manifest.
93
+ The generated bundle is written to `build/mcpb/mcp-scraper-<version>.mcpb` and copied to `public/downloads/` for the hosted download. The current public bundle is `https://mcpscraper.dev/downloads/mcp-scraper.mcpb` (`0.41.0`, SHA-256 `7f32f079e82a49acf936bd9b91d1c0b481e1949e0f52a27a5e20f312d7265228`). Install it by opening or dragging it into Claude Desktop. Claude displays the `MCP Scraper` install card, icon, and API-key configuration field from the bundle manifest.
94
94
 
95
95
  The MCPB install exposes every tool — web-intelligence plus all `browser_*` tools — through the one `mcp-scraper` server.
96
96
 
@@ -185,7 +185,7 @@ env = { MCP_SCRAPER_API_KEY = "sk_live_your_key" }
185
185
 
186
186
  - `list_service_connections` — list this caller's tenant-owned Nango OAuth and official remote MCP connections, including verified provider-side account email/name when exposed, exact live reads, gated actions, permanently blocked administrative tools, credential transport, and schema-discovery metadata. Provider identity is distinct from the MCP Scraper login, and connections are never shared between customers.
187
187
  - `describe_service_connection_tool` — fetch the sanitized live MCP Tool definition for one tool listed on one tenant-owned connection, including its current callability, input schema, optional output schema, safe annotations, and schema hash. Use this before constructing provider-native arguments; provider functions stay behind the generic bridges instead of becoming dozens of permanent top-level tools.
188
- - `export_connected_service_data` — fetch a fresh bounded Gmail, Google Calendar, Google Search Console, Zoom, Resend, or Meta time range in one MCP call. Zoom's `zoom_transcripts` dataset resolves VTT files from recording metadata and downloads them through the authenticated connection without looping the separately rate-limited `get-meeting-transcript` function. Search Console's `search_console_performance` dataset walks accessible properties and bounded live Search Analytics pages with signed continuation. Small exports return inline; larger exports become private JSONL retained for seven days with a 15-minute signed URL. For relationship work, gather source evidence first: inspect existing People records, resolve the exact provider account, use RFC3339 `from`/`to` for Gmail ranges longer than 90 days, preserve provider provenance when writing a linked Communication, and never treat an export as permission to mutate the source account.
188
+ - `export_connected_service_data` — fetch a fresh Gmail, Google Calendar, Google Search Console, Zoom, Slack, Resend, or Meta dataset in one MCP call. Slack's `slack_channel_messages` dataset accepts a `channelId`, paginates top-level history, fetches threaded replies in bounded parallel batches, preserves file metadata, honors retry delays, and supports `allTime:true`; it never joins or changes the channel. Zoom's `zoom_transcripts` dataset resolves VTT files from recording metadata and downloads them through the authenticated connection without looping the separately rate-limited `get-meeting-transcript` function. Search Console's `search_console_performance` dataset walks accessible properties and bounded live Search Analytics pages with continuation. Small exports return inline; larger exports become private JSONL retained for seven days with a 15-minute signed URL. For relationship work, gather source evidence first: inspect existing People records, resolve the exact provider account, preserve provider provenance when writing a linked Communication, and never treat an export as permission to mutate the source account.
189
189
  - `export_search_console_table_data` — filter up to 50,000 Search Console rows already persisted by a scheduled `connection_sync` and create a private renewable JSONL artifact without calling Google again. Get the typed `gsc_performance_*` table name from `list_service_connections`, inspect it with `table-describe`, and use the same filters with `table-query` for interactive analysis.
190
190
  - `renew_connected_data_download` — issue a fresh 15-minute signed URL for an unexpired private export artifact without pulling the provider again.
191
191
  - `read_service_connection` — run one small live read by exact allowlisted name across Nango OAuth or official remote MCP connections, including bounded Google Drive inventory, change, Doc, Sheet, and text-file tools. Do not loop it over a time range when `export_connected_service_data` supports that provider's collection.