mcp-scraper 0.43.5 → 0.44.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (32) hide show
  1. package/README.md +4 -3
  2. package/dist/bin/api-server.cjs +14020 -9827
  3. package/dist/bin/api-server.cjs.map +1 -1
  4. package/dist/bin/api-server.js +1 -1
  5. package/dist/bin/mcp-scraper-cli.cjs +1 -1
  6. package/dist/bin/mcp-scraper-cli.cjs.map +1 -1
  7. package/dist/bin/mcp-scraper-cli.js +1 -1
  8. package/dist/bin/mcp-scraper-install.cjs +2 -2
  9. package/dist/bin/mcp-scraper-install.cjs.map +1 -1
  10. package/dist/bin/mcp-scraper-install.js +2 -2
  11. package/dist/bin/mcp-stdio-server.cjs +592 -12
  12. package/dist/bin/mcp-stdio-server.cjs.map +1 -1
  13. package/dist/bin/mcp-stdio-server.js +5 -3
  14. package/dist/bin/mcp-stdio-server.js.map +1 -1
  15. package/dist/{chunk-6WLNXYRG.js → chunk-27FMOD6S.js} +1 -2
  16. package/dist/chunk-27FMOD6S.js.map +1 -0
  17. package/dist/{chunk-TC4GDB2Q.js → chunk-GRODXRUY.js} +2 -2
  18. package/dist/{chunk-TC4GDB2Q.js.map → chunk-GRODXRUY.js.map} +1 -1
  19. package/dist/{chunk-RSGUS5V6.js → chunk-HH5CE5LP.js} +594 -13
  20. package/dist/chunk-HH5CE5LP.js.map +1 -0
  21. package/dist/chunk-RUDXFPKE.js +7 -0
  22. package/dist/chunk-RUDXFPKE.js.map +1 -0
  23. package/dist/{extract-bundle-ITPFWRNI.js → extract-bundle-U3MNYDKC.js} +2 -2
  24. package/dist/{server-BTDXZA5U.js → server-75YGJWSI.js} +5955 -2438
  25. package/dist/server-75YGJWSI.js.map +1 -0
  26. package/package.json +1 -1
  27. package/dist/chunk-6WLNXYRG.js.map +0 -1
  28. package/dist/chunk-H4RPUFRR.js +0 -7
  29. package/dist/chunk-H4RPUFRR.js.map +0 -1
  30. package/dist/chunk-RSGUS5V6.js.map +0 -1
  31. package/dist/server-BTDXZA5U.js.map +0 -1
  32. /package/dist/{extract-bundle-ITPFWRNI.js.map → extract-bundle-U3MNYDKC.js.map} +0 -0
package/README.md CHANGED
@@ -90,7 +90,7 @@ Build the branded one-click bundle:
90
90
  npm run build:mcpb
91
91
  ```
92
92
 
93
- The generated bundle is written to `build/mcpb/mcp-scraper-<version>.mcpb` and copied to `public/downloads/` for the hosted download. The current public bundle is `https://mcpscraper.dev/downloads/mcp-scraper.mcpb` (`0.43.5`, SHA-256 `2c70cc4490f44d78f6803ab2e2b2e2fcdde72949532be053192744815bebe2aa`). Install it by opening or dragging it into Claude Desktop. Claude displays the `MCP Scraper` install card, icon, and API-key configuration field from the bundle manifest.
93
+ The generated bundle is written to `build/mcpb/mcp-scraper-<version>.mcpb` and copied to `public/downloads/` for the hosted download. The current public bundle is `https://mcpscraper.dev/downloads/mcp-scraper.mcpb` (`0.44.0`, SHA-256 `83ce32972445fc1b3b308d7e79d6ad7248fb1485d4a13635249b85ea47ac6a61`). Install it by opening or dragging it into Claude Desktop. Claude displays the `MCP Scraper` install card, icon, and API-key configuration field from the bundle manifest.
94
94
 
95
95
  The MCPB install exposes every tool — web-intelligence plus all `browser_*` tools — through the one `mcp-scraper` server.
96
96
 
@@ -178,6 +178,7 @@ env = { MCP_SCRAPER_API_KEY = "sk_live_your_key" }
178
178
  - `editorial_reading_room_guide` — load the reusable editorial workflow, content contract, or compact example before turning dense supplied material into a reading surface.
179
179
  - `create_editorial_reading_room` — render fully authored, source-grounded articles into one self-contained mobile-first HTML reading room with contents, hamburger navigation, search, jump links, progress, text sizing, evening mode, and visible provenance. Hosted clients receive a private seven-day artifact; local stdio clients receive an openable file under the MCP Scraper output directory.
180
180
  - `renew_editorial_reading_room_download` — issue a fresh signed URL for an unexpired private reading-room artifact.
181
+ - `report_artifact_read` — read owner-scoped text and JSONL artifacts through the authenticated MCP connection when a model sandbox cannot open the optional signed download URL. Continue with `nextOffset` until it is null; ZIP archives use `archive_read`.
181
182
  - `rank_tracker_workflow` — generate a database schema, cron/heartbeat plan, ingestion workflow, metrics list, and implementation prompt for building rank trackers. It has modes for Maps rankings via `directory_workflow`/`maps_search`, organic rankings via `search_serp`, AI Overview citation tracking, and PAA source presence tracking. This planning tool does not spend credits.
182
183
  - `credits_info`
183
184
 
@@ -185,7 +186,7 @@ env = { MCP_SCRAPER_API_KEY = "sk_live_your_key" }
185
186
 
186
187
  - `list_service_connections` — list this caller's tenant-owned Nango OAuth and official remote MCP connections, including verified provider-side account email/name when exposed, exact live reads, gated actions, permanently blocked administrative tools, credential transport, and schema-discovery metadata. Provider identity is distinct from the MCP Scraper login, and connections are never shared between customers.
187
188
  - `describe_service_connection_tool` — fetch the sanitized live MCP Tool definition for one tool listed on one tenant-owned connection, including its current callability, input schema, optional output schema, safe annotations, and schema hash. Use this before constructing provider-native arguments; provider functions stay behind the generic bridges instead of becoming dozens of permanent top-level tools.
188
- - `export_connected_service_data` — fetch a fresh Gmail, Google Calendar, Google Search Console, Zoom, Slack, Resend, or Meta dataset in one MCP call. Slack's `slack_channel_messages` dataset accepts a `channelId`, paginates top-level history, fetches threaded replies in bounded parallel batches, preserves file metadata, honors retry delays, and supports `allTime:true`; it never joins or changes the channel. Zoom's `zoom_transcripts` dataset resolves VTT files from recording metadata and downloads them through the authenticated connection without looping the separately rate-limited `get-meeting-transcript` function. Search Console's `search_console_performance` dataset walks accessible properties and bounded live Search Analytics pages with continuation. Small exports return inline; larger exports become private JSONL retained for seven days with a 15-minute signed URL. For relationship work, gather source evidence first: inspect existing People records, resolve the exact provider account, preserve provider provenance when writing a linked Communication, and never treat an export as permission to mutate the source account.
189
+ - `export_connected_service_data` — fetch a fresh Gmail, Google Calendar, Google Search Console, Zoom, Slack, Resend, or Meta dataset in one MCP call. Slack's `slack_channel_messages` dataset accepts a `channelId`, paginates top-level history, fetches threaded replies in bounded parallel batches, preserves file metadata, honors retry delays, and supports `allTime:true`; it never joins or changes the channel. Zoom's `zoom_transcripts` dataset resolves VTT files from recording metadata and downloads them through the authenticated connection without looping the separately rate-limited `get-meeting-transcript` function. Search Console's `search_console_performance` dataset walks accessible properties and bounded live Search Analytics pages with continuation. Small exports return inline; larger exports become private JSONL retained for seven days with exact `report_artifact_read` arguments plus an optional 15-minute human download URL. For relationship work, gather source evidence first: inspect existing People records, resolve the exact provider account, preserve provider provenance when writing a linked Communication, and never treat an export as permission to mutate the source account.
189
190
  - `export_search_console_table_data` — filter up to 50,000 Search Console rows already persisted by a scheduled `connection_sync` and create a private renewable JSONL artifact without calling Google again. Get the typed `gsc_performance_*` table name from `list_service_connections`, inspect it with `table-describe`, and use the same filters with `table-query` for interactive analysis.
190
191
  - `renew_connected_data_download` — issue a fresh 15-minute signed URL for an unexpired private export artifact without pulling the provider again.
191
192
  - `read_service_connection` — run one small live read by exact allowlisted name across Nango OAuth or official remote MCP connections, including bounded Google Drive inventory, change, Doc, Sheet, and text-file tools. Do not loop it over a time range when `export_connected_service_data` supports that provider's collection.
@@ -232,7 +233,7 @@ The `mcp-scraper` server (and the MCPB bundle, which runs it) exposes both secti
232
233
 
233
234
  All MCP tools expose output schemas and return `structuredContent` with the IDs, URLs, CSV paths, transcripts, browser session handles, replay paths, artifacts, recipe fields, or blueprint fields needed by the next step. Browser Agent tools keep a JSON text block for older clients, but structured data is the primary contract. All tools carry MCP annotations; file-writing tools such as replay downloads and annotations state their filesystem side effects.
234
235
 
235
- The canonical tool inventory is generated at `docs/mcp-tool-manifest.generated.json`. The unified server exposes 198 tools: 97 scraper, browser, workflow, billing, and connected-service tools plus 101 durable-memory tools. Release verification compares the exact local and hosted tool-name sets, not only the count.
236
+ The canonical tool inventory is generated at `docs/mcp-tool-manifest.generated.json`. The unified server exposes 215 tools: 114 scraper, browser, workflow, billing, and connected-service tools plus 101 durable-memory tools. The scraper inventory includes governed Local Sourcebook tools that follow a Memory-style contract/tag/prepare/validate/capture sequence before paid acquisition, plus Transparent Commons tools for public entity search, planning, validation, governed contribution, ledgers, needs-link discovery, and saved filters. Local Sourcebook publication remains administrator-only. Release verification compares the exact local and hosted tool-name sets, not only the count.
236
237
 
237
238
  For contract parity, stdio and MCPB memory calls invoke the matching public tool on the hosted MCP Scraper `/mcp` endpoint. The hosted aggregate runtime owns MCP Scraper-specific billing, scheduling, credential, and in-process cutover policy; its internal `/memory/mcp-call` bridge is a fallback to the standalone memory service, not the public stdio execution path. Direct `mcp-memory` OAuth and stdio clients continue to use `memory.mcpscraper.dev` and must be verified as a separate dependent release surface.
238
239