mcp-scraper 0.34.0 → 0.35.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +6 -5
- package/dist/bin/api-server.cjs +3788 -2742
- package/dist/bin/api-server.cjs.map +1 -1
- package/dist/bin/api-server.js +2 -2
- package/dist/bin/mcp-scraper-cli.cjs +1 -1
- package/dist/bin/mcp-scraper-cli.cjs.map +1 -1
- package/dist/bin/mcp-scraper-cli.js +1 -1
- package/dist/bin/mcp-scraper-install.cjs +2 -2
- package/dist/bin/mcp-scraper-install.cjs.map +1 -1
- package/dist/bin/mcp-scraper-install.js +2 -2
- package/dist/bin/mcp-stdio-server.cjs +2129 -1932
- package/dist/bin/mcp-stdio-server.cjs.map +1 -1
- package/dist/bin/mcp-stdio-server.js +5 -5
- package/dist/bin/paa-harvest.js +2 -2
- package/dist/{chunk-V36LS5YV.js → chunk-4ZB3X6BQ.js} +4 -3
- package/dist/chunk-4ZB3X6BQ.js.map +1 -0
- package/dist/{chunk-JBSGBSGT.js → chunk-BWXLTWF7.js} +21 -16
- package/dist/chunk-BWXLTWF7.js.map +1 -0
- package/dist/{chunk-J32XSMJI.js → chunk-NPMW5HUS.js} +2350 -2151
- package/dist/chunk-NPMW5HUS.js.map +1 -0
- package/dist/{chunk-SUOHUXQS.js → chunk-U44TPRST.js} +3 -3
- package/dist/chunk-U44TPRST.js.map +1 -0
- package/dist/{chunk-4KQVIYIH.js → chunk-XVVNKASZ.js} +2 -2
- package/dist/chunk-YR4LJ6AQ.js +7 -0
- package/dist/chunk-YR4LJ6AQ.js.map +1 -0
- package/dist/{chunk-2EDFOQD7.js → chunk-YRGSEY5L.js} +2 -2
- package/dist/{chunk-2EDFOQD7.js.map → chunk-YRGSEY5L.js.map} +1 -1
- package/dist/chunk-YV2FUEBX.js +851 -0
- package/dist/chunk-YV2FUEBX.js.map +1 -0
- package/dist/{extract-bundle-GPZRLBVY.js → extract-bundle-ONWZVV55.js} +59 -8
- package/dist/extract-bundle-ONWZVV55.js.map +1 -0
- package/dist/index.js +2 -2
- package/dist/{server-LKVPEQFE.js → server-GKUTC73B.js} +278 -52
- package/dist/server-GKUTC73B.js.map +1 -0
- package/dist/{site-extract-repository-JHMVHENZ.js → site-extract-repository-L6BHWVDU.js} +3 -3
- package/dist/{worker-ZZHSYI3Y.js → worker-645BZPEK.js} +4 -4
- package/docs/mcp-tool-craft-lint.generated.md +171 -49
- package/docs/mcp-tool-manifest.generated.json +508 -10
- package/docs/specs/query-fanout-transport-contract-fix.md +45 -0
- package/package.json +1 -1
- package/dist/chunk-J32XSMJI.js.map +0 -1
- package/dist/chunk-JBSGBSGT.js.map +0 -1
- package/dist/chunk-JWIE5NCR.js +0 -284
- package/dist/chunk-JWIE5NCR.js.map +0 -1
- package/dist/chunk-NFMIHEYE.js +0 -7
- package/dist/chunk-NFMIHEYE.js.map +0 -1
- package/dist/chunk-SUOHUXQS.js.map +0 -1
- package/dist/chunk-V36LS5YV.js.map +0 -1
- package/dist/extract-bundle-GPZRLBVY.js.map +0 -1
- package/dist/server-LKVPEQFE.js.map +0 -1
- /package/dist/{chunk-4KQVIYIH.js.map → chunk-XVVNKASZ.js.map} +0 -0
- /package/dist/{site-extract-repository-JHMVHENZ.js.map → site-extract-repository-L6BHWVDU.js.map} +0 -0
- /package/dist/{worker-ZZHSYI3Y.js.map → worker-645BZPEK.js.map} +0 -0
package/README.md
CHANGED
|
@@ -88,7 +88,7 @@ Build the branded one-click bundle:
|
|
|
88
88
|
npm run build:mcpb
|
|
89
89
|
```
|
|
90
90
|
|
|
91
|
-
The generated bundle is written to `build/mcpb/mcp-scraper-<version>.mcpb` and copied to `public/downloads/` for the hosted download. The current public bundle is `https://mcpscraper.dev/downloads/mcp-scraper.mcpb` (`0.
|
|
91
|
+
The generated bundle is written to `build/mcpb/mcp-scraper-<version>.mcpb` and copied to `public/downloads/` for the hosted download. The current public bundle is `https://mcpscraper.dev/downloads/mcp-scraper.mcpb` (`0.35.1`, SHA-256 `95edd500a9879b1506b622bb2fc3c2bff3bcde419b927889512c19c1aec90ae8`). Install it by opening or dragging it into Claude Desktop. Claude displays the `MCP Scraper` install card, icon, and API-key configuration field from the bundle manifest.
|
|
92
92
|
|
|
93
93
|
The MCPB install exposes every tool — web-intelligence plus all `browser_*` tools — through the one `mcp-scraper` server.
|
|
94
94
|
|
|
@@ -154,7 +154,8 @@ env = { MCP_SCRAPER_API_KEY = "sk_live_your_key" }
|
|
|
154
154
|
- `search_serp`
|
|
155
155
|
- `extract_url`
|
|
156
156
|
- `map_site_urls`
|
|
157
|
-
- `
|
|
157
|
+
- `map_wayback_snapshots` — count and inventory Wayback captures across an inclusive date range without downloading page bodies. Supports exact pages, prefixes, hosts, domains, or selected URLs; reports exact versus lower-bound counts, unique URLs/content digests, monthly coverage, missing months, and optional timestamp rows.
|
|
158
|
+
- `extract_site` — crawl a live site, batch one archived site snapshot from a Wayback replay URL, or pass a `wayback` plan for whole-site, single-page, or selected-page timelines across explicit months or a `from`/`to` range. Timeline ZIPs include month folders and a capture matrix.
|
|
158
159
|
- `youtube_harvest`
|
|
159
160
|
- `youtube_transcribe`
|
|
160
161
|
- `facebook_ad_search`
|
|
@@ -213,7 +214,7 @@ Google Search Console exposes eight bounded reads and eight gated property and s
|
|
|
213
214
|
- `browser_replay_download` — download and save the replay MP4 under `MCP_SCRAPER_OUTPUT_DIR/browser-replays`.
|
|
214
215
|
- `browser_replay_mark` — while recording, locate a DOM target and return a replay-timed annotation object.
|
|
215
216
|
- `browser_replay_annotate` — download a replay MP4, render timed boxes, circles, underlines, arrows, and labels using annotation objects from `browser_replay_mark` or exact bounds from `browser_locate`, and save a new annotated MP4 under `MCP_SCRAPER_OUTPUT_DIR/browser-replays`.
|
|
216
|
-
- `browser_capture_fanout` — capture ChatGPT/Claude AI-search fan-out from an open logged-in hosted session.
|
|
217
|
+
- `browser_capture_fanout` — capture ChatGPT/Claude AI-search fan-out from an open logged-in hosted session. Every client receives the complete structured capture inline. Installed stdio/MCPB clients can use `export=true` for durable `fanout.json`, query/source/citation/domain/snippet CSVs, TSV, and `report.html` under `MCP_SCRAPER_OUTPUT_DIR/fanout`; hosted OAuth clients receive `exports: null` and use the inline data.
|
|
217
218
|
- `browser_close`
|
|
218
219
|
- `browser_list_sessions`
|
|
219
220
|
|
|
@@ -227,7 +228,7 @@ The `mcp-scraper` server (and the MCPB bundle, which runs it) exposes both secti
|
|
|
227
228
|
|
|
228
229
|
All MCP tools expose output schemas and return `structuredContent` with the IDs, URLs, CSV paths, transcripts, browser session handles, replay paths, artifacts, recipe fields, or blueprint fields needed by the next step. Browser Agent tools keep a JSON text block for older clients, but structured data is the primary contract. All tools carry MCP annotations; file-writing tools such as replay downloads and annotations state their filesystem side effects.
|
|
229
230
|
|
|
230
|
-
The canonical tool inventory is generated at `docs/mcp-tool-manifest.generated.json`. Both the `mcp-scraper` stdio server and the hosted endpoint at `https://mcpscraper.dev/mcp` expose the same
|
|
231
|
+
The canonical tool inventory is generated at `docs/mcp-tool-manifest.generated.json`. Both the `mcp-scraper` stdio server and the hosted endpoint at `https://mcpscraper.dev/mcp` expose the same 168 tools: 79 scraper, browser, workflow, billing, and connected-service tools plus 89 durable-memory tools. Release verification compares the exact local and remote tool-name sets, not only the count.
|
|
231
232
|
|
|
232
233
|
For contract parity, stdio and MCPB memory calls invoke the matching public tool on the hosted MCP Scraper `/mcp` endpoint. The hosted aggregate runtime owns MCP Scraper-specific billing, scheduling, credential, and in-process cutover policy; its internal `/memory/mcp-call` bridge is a fallback to the standalone memory service, not the public stdio execution path. Direct `mcp-memory` OAuth and stdio clients continue to use `memory.mcpscraper.dev` and must be verified as a separate dependent release surface.
|
|
233
234
|
|
|
@@ -245,7 +246,7 @@ The `mcp-scraper` NPX stdio server also exposes saved reports as MCP resources:
|
|
|
245
246
|
- `BROWSER_AGENT_PROFILE_NAME` is optional and sets the default saved hosted browser profile for `mcp-scraper` stdio sessions. Aliases: `BROWSER_SERVICE_PROFILE_NAME`, `KERNEL_BROWSER_PROFILE_NAME`, `KERNEL_PROFILE_NAME`.
|
|
246
247
|
- `BROWSER_AGENT_PROFILE_SAVE_CHANGES=true` is optional hosted setup behavior. It persists cookies and storage back to the named profile when `browser_close` deletes the hosted browser session. Aliases: `BROWSER_SERVICE_PROFILE_SAVE_CHANGES`, `KERNEL_BROWSER_PROFILE_SAVE_CHANGES`, `KERNEL_PROFILE_SAVE_CHANGES`.
|
|
247
248
|
|
|
248
|
-
Every web intelligence tool call made through `mcp-scraper` saves a full Markdown report to disk by default and returns the file path in the MCP response. The hosted `/mcp` endpoint returns reports inline only and never writes files. Browser replay downloads are saved by `browser_replay_download` under `MCP_SCRAPER_OUTPUT_DIR/browser-replays`. AI fan-out
|
|
249
|
+
Every web intelligence tool call made through `mcp-scraper` saves a full Markdown report to disk by default and returns the file path in the MCP response. The hosted `/mcp` endpoint returns reports inline only and never writes files. Browser replay downloads are saved by `browser_replay_download` under `MCP_SCRAPER_OUTPUT_DIR/browser-replays`. AI fan-out captures are always returned inline; only installed stdio/MCPB clients write optional `export=true` files under `MCP_SCRAPER_OUTPUT_DIR/fanout`, returning relative paths. Hosted clients always receive `exports: null`.
|
|
249
250
|
|
|
250
251
|
## Updating Existing Installs
|
|
251
252
|
|