mcp-scraper 0.93.0 → 0.93.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (44) hide show
  1. package/CHANGELOG.md +37 -25
  2. package/README.md +3 -3
  3. package/dist/bin/api-server.js +1 -1
  4. package/dist/bin/mcp-scraper-cli.js +1 -1
  5. package/dist/bin/mcp-scraper-core.js +1 -1
  6. package/dist/bin/mcp-scraper-install.js +1 -1
  7. package/dist/bin/mcp-stdio-server.js +1 -1
  8. package/dist/bin/paa-harvest.js +1 -1
  9. package/dist/chunk-2SP57VCG.js +15 -15
  10. package/dist/{chunk-PMOPJOSJ.js → chunk-5JKBFNYF.js} +1 -1
  11. package/dist/{chunk-35QGIWR6.js → chunk-5VB3I7UX.js} +4 -4
  12. package/dist/chunk-73LAMUS5.js +6 -0
  13. package/dist/{chunk-LBMRUR4E.js → chunk-AFFMCO7R.js} +2 -2
  14. package/dist/{chunk-4YLZR4OR.js → chunk-CG2FHIBA.js} +1 -1
  15. package/dist/chunk-DT2FYN6N.js +1 -1
  16. package/dist/{chunk-XG6GCEUE.js → chunk-EWR7BPPD.js} +1 -1
  17. package/dist/{chunk-H3SJLJYM.js → chunk-IJF4VXJ7.js} +1 -1
  18. package/dist/{chunk-IK5BG7MO.js → chunk-KPDNM7E7.js} +1 -1
  19. package/dist/{chunk-M5VZVMPC.js → chunk-MJ2OICVY.js} +3 -3
  20. package/dist/{chunk-5KQO27NH.js → chunk-NLKA6SHC.js} +1 -1
  21. package/dist/chunk-ORB4RHCK.js +4 -4
  22. package/dist/{chunk-VRVO4OQ3.js → chunk-QWSWHUOK.js} +150 -150
  23. package/dist/chunk-SH5KB4P7.js +1 -1
  24. package/dist/chunk-UN6FVDJQ.js +1 -0
  25. package/dist/{chunk-M5QHXNFZ.js → chunk-VVG7LFFF.js} +3 -3
  26. package/dist/chunk-WJ4XFLS4.js +15 -15
  27. package/dist/chunk-XJ6PJWT5.js +1 -0
  28. package/dist/chunk-Y6MKMSOC.js +1 -1
  29. package/dist/{extract-bundle-M2XM4BEW.js → extract-bundle-MQOAKQDV.js} +1 -1
  30. package/dist/{gmail-service-EGATZIBE.js → gmail-service-6V5MFBYH.js} +1 -1
  31. package/dist/index.cjs +21 -21
  32. package/dist/index.d.cts +15 -15
  33. package/dist/index.d.ts +15 -15
  34. package/dist/index.js +1 -1
  35. package/dist/{server-JLR2YJ7V.js → server-GRB6VNT6.js} +43 -43
  36. package/dist/{site-extract-repository-3NFFFPXQ.js → site-extract-repository-2LFDP6FZ.js} +1 -1
  37. package/dist/{stripe-event-worker-PUYXT5DS.js → stripe-event-worker-IJXA6T37.js} +1 -1
  38. package/dist/worker-US6CTFYG.js +1 -0
  39. package/package.json +136 -17
  40. package/THIRD_PARTY_NOTICES.html +0 -203
  41. package/dist/chunk-DUJ56SHU.js +0 -1
  42. package/dist/chunk-RZEZJAID.js +0 -1
  43. package/dist/chunk-XASDACRP.js +0 -6
  44. package/dist/worker-UG4OXZWR.js +0 -1
package/CHANGELOG.md CHANGED
@@ -4,6 +4,18 @@ All notable changes to MCP Scraper are documented here. The format is based on [
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [0.93.2] - 2026-09-24
8
+
9
+ ### Fixed
10
+
11
+ - Show Google Search and other provider costs in the CEO report with separate calculated, actual-or-measured, and unresolved receipt coverage. Failed queries now stop the report, missing Bright Data SERP rates leave an explicit delivered-but-unpriced receipt, and historical cost rows no longer imply a margin from subscription MRR.
12
+
13
+ ## [0.93.1] - 2026-09-24
14
+
15
+ ### Fixed
16
+
17
+ - Keep explicitly requested two-page Google searches on the provider that supports them. Two-page full requests return organic listings with rich feature status marked unsupported and cost 40 Credits; if page 2 cannot be delivered, an available page 1 is returned as partial and billed for one page.
18
+
7
19
  ## [0.93.0] - 2026-09-24
8
20
 
9
21
  ### Added
@@ -95,7 +107,7 @@ All notable changes to MCP Scraper are documented here. The format is based on [
95
107
 
96
108
  - Retire the inactive Personal Assistant from server startup, HTTP and MCP routing, web navigation, Scheduler transitions, generated contracts, and Memory tool registration while preserving its source history, persisted data, secrets, and provider resources behind the verified archive branch.
97
109
  - Make two-page SERP capture perform a distinct page-two request, preserve page provenance and query/location intent, and report local-pack evidence as present, absent, incomplete, or unknown instead of silently claiming completeness.
98
- - Make Maps retries follow one immutable egress plan, retain location intent, close every browser/proxy attempt, and treat an already-gone Option 1 session as successful cleanup.
110
+ - Make Maps retries follow one immutable egress plan, retain location intent, close every browser/proxy attempt, and treat an already-gone Kernel session as successful cleanup.
99
111
 
100
112
  ### Fixed
101
113
 
@@ -109,13 +121,13 @@ All notable changes to MCP Scraper are documented here. The format is based on [
109
121
 
110
122
  - Add the fixture-tested operation, attempt, provider-receipt, reconciliation, and billing-link foundation needed to trace failed and retried SERP, PAA, Maps, extraction, and transcription work without treating missing provider cost as zero.
111
123
  - Add a generated MCP cost-coverage manifest and release gate that tracks every hosted input flag plus each named execution method's retry, timeout, billing, and cost-accounting owner.
112
- - Trace direct Option 2 SERP requests and Option 2/Option 1 browser attempts into the shared operation timeline, including provider identities, retry or fallback causality, immediate estimates, hourly due-gated Browser API reconciliation, and customer billing links.
113
- - Add a protected four-cell PAA benchmark runner for Option 1 and Option 2 at 20 and 40 complete questions, with a dry-run contract gate, one paid attempt per cell, sanitized provider evidence, and no automatic replacement run.
124
+ - Trace direct Bright Data SERP requests and Bright Data/Kernel browser attempts into the shared operation timeline, including provider identities, retry or fallback causality, immediate estimates, hourly due-gated Browser API reconciliation, and customer billing links.
125
+ - Add a protected four-cell PAA benchmark runner for Kernel and Bright Data at 20 and 40 complete questions, with a dry-run contract gate, one paid attempt per cell, sanitized provider evidence, and no automatic replacement run.
114
126
  - Add a once-daily due-gated purge of raw provider identifiers and verbose diagnostic errors after 30 days while preserving normalized cost facts and hashed correlation keys.
115
127
 
116
128
  ### Fixed
117
129
 
118
- - Skip Option 1 proxy resolution when the active attempt uses Option 2 Browser, removing avoidable Option 1 API traffic from Option 2-first SERP and PAA work.
130
+ - Skip Kernel proxy resolution when the active attempt uses Bright Data Browser, removing avoidable Kernel API traffic from Bright Data-first SERP and PAA work.
119
131
 
120
132
  ## [0.90.4] - 2026-09-22
121
133
 
@@ -127,21 +139,21 @@ All notable changes to MCP Scraper are documented here. The format is based on [
127
139
 
128
140
  ### Fixed
129
141
 
130
- - Route ordinary Option 2 SERP searches through the zone's native parsed proxy, returning Google results within the existing bounded deadline while preserving the REST endpoint as a configuration fallback.
142
+ - Route ordinary Bright Data SERP searches through the zone's native parsed proxy, returning Google results within the existing bounded deadline while preserving the REST endpoint as a configuration fallback.
131
143
  - Keep Browser API credentials and interactive PAA behavior independent from the native SERP transport, with one provider request and no hidden retry or browser fallback.
132
144
 
133
145
  ## [0.90.2] - 2026-09-21
134
146
 
135
147
  ### Fixed
136
148
 
137
- - Unwrap Option 2's object-valued REST response body before mapping parsed SERP fields, so successful direct searches return their organic results instead of an empty collection.
149
+ - Unwrap Bright Data's object-valued REST response body before mapping parsed SERP fields, so successful direct searches return their organic results instead of an empty collection.
138
150
 
139
151
  ## [0.90.1] - 2026-09-21
140
152
 
141
153
  ### Changed
142
154
 
143
- - Route ordinary `search_serp` calls through one parsed Option 2 SERP API request when the dedicated production zone is configured, avoiding browser startup and cleanup while leaving interactive PAA on Browser API.
144
- - Keep saved SERP identities on their existing Option 1-backed path and retain the browser path as the configuration fallback for local and unconfigured environments.
155
+ - Route ordinary `search_serp` calls through one parsed Bright Data SERP API request when the dedicated production zone is configured, avoiding browser startup and cleanup while leaving interactive PAA on Browser API.
156
+ - Keep saved SERP identities on their existing Kernel-backed path and retain the browser path as the configuration fallback for local and unconfigured environments.
145
157
 
146
158
  ### Fixed
147
159
 
@@ -184,7 +196,7 @@ All notable changes to MCP Scraper are documented here. The format is based on [
184
196
 
185
197
  ### Changed
186
198
 
187
- - Cap every Option 1 browser session at a ten-minute absolute lifetime and move expired-session reconciliation from the minute root cron to a dedicated hourly audit.
199
+ - Cap every Kernel browser session at a ten-minute absolute lifetime and move expired-session reconciliation from the minute root cron to a dedicated hourly audit.
188
200
  - Disable the inactive Personal Assistant reminder, reconciliation, and inbound cron schedules while preserving their routes and implementation.
189
201
  - Disable the inactive Personal Memory heartbeat and weekly rollup registrations in the production scheduler while preserving their implementation.
190
202
  - Pause twice-daily automatic memory optimization while preserving the workflow for deliberate use, preventing all-vault fan-out from consuming background execution capacity.
@@ -193,10 +205,10 @@ All notable changes to MCP Scraper are documented here. The format is based on [
193
205
 
194
206
  ### Fixed
195
207
 
196
- - Treat Option 1's not-found response during legacy session deletion as successful cleanup, preventing already-closed sessions from retrying forever.
208
+ - Treat Kernel's not-found response during legacy session deletion as successful cleanup, preventing already-closed sessions from retrying forever.
197
209
  - Show browser sessions awaiting provider cleanup separately in the CTO report instead of hiding them behind a non-null close timestamp.
198
210
  - Cast the analytics pruning clock before PostgreSQL interval arithmetic so the root cron no longer fails every minute while pruning scheduled occurrences.
199
- - Explicitly delete Option 1 screenshot sessions after capture, including when closing the browser connection fails.
211
+ - Explicitly delete Kernel screenshot sessions after capture, including when closing the browser connection fails.
200
212
 
201
213
  ## [0.89.7] - 2026-09-18
202
214
 
@@ -315,7 +327,7 @@ All notable changes to MCP Scraper are documented here. The format is based on [
315
327
 
316
328
  ### Fixed
317
329
 
318
- - Give each `reddit_thread` retrieval a 300-second end-to-end deadline, with two 60-second Option 1 attempts and two 60-second managed-browser backup attempts, instead of exhausting the full retry ladder in about 50 seconds. The MCP client now waits long enough to receive the endpoint's structured terminal result.
330
+ - Give each `reddit_thread` retrieval a 300-second end-to-end deadline, with two 60-second Kernel attempts and two 60-second managed-browser backup attempts, instead of exhausting the full retry ladder in about 50 seconds. The MCP client now waits long enough to receive the endpoint's structured terminal result.
319
331
 
320
332
  ## [0.88.2] - 2026-09-02
321
333
 
@@ -352,19 +364,19 @@ All notable changes to MCP Scraper are documented here. The format is based on [
352
364
 
353
365
  ### Fixed
354
366
 
355
- - Kept Option 2 telemetry lookup off the Reddit response critical path and reallocated the saved time to 17-second backup attempts, so all four provider attempts can finish before production ends the request.
367
+ - Kept Bright Data telemetry lookup off the Reddit response critical path and reallocated the saved time to 17-second backup attempts, so all four provider attempts can finish before production ends the request.
356
368
 
357
369
  ## [0.86.4] - 2026-09-02
358
370
 
359
371
  ### Fixed
360
372
 
361
- - Kept the complete two-primary, two-backup Reddit retry ladder inside the production request window by limiting Option 1 attempts to 8 seconds, Option 2 attempts to 14 seconds, and browser cleanup to 1 second.
373
+ - Kept the complete two-primary, two-backup Reddit retry ladder inside the production request window by limiting Kernel attempts to 8 seconds, Bright Data attempts to 14 seconds, and browser cleanup to 1 second.
362
374
 
363
375
  ## [0.86.3] - 2026-09-02
364
376
 
365
377
  ### Fixed
366
378
 
367
- - Applied 45-second Option 1 and 35-second Option 2 deadlines to the complete Reddit browser-attempt lifecycle, and made known-thread primary attempts find and click the target through DuckDuckGo before the residential landing.
379
+ - Applied 45-second Kernel and 35-second Bright Data deadlines to the complete Reddit browser-attempt lifecycle, and made known-thread primary attempts find and click the target through DuckDuckGo before the residential landing.
368
380
 
369
381
  ## [0.86.2] - 2026-09-02
370
382
 
@@ -513,13 +525,13 @@ All notable changes to MCP Scraper are documented here. The format is based on [
513
525
 
514
526
  ### Added
515
527
 
516
- - Added a Option 1-only Reddit workflow that searches DuckDuckGo with a `site:reddit.com` query, switches the same browser to a residential proxy before clicking the selected result, and reads modern Reddit posts plus bounded rendered-comment expansion through dedicated search, thread, and combined REST endpoints.
517
- - Added a bounded managed-browser backup for Reddit thread hydration after the primary Option 1 attempt fails or returns fewer than the semantic target, capped at 25 comments with measured bandwidth, duration, CAPTCHA, closure, and provider-cost telemetry.
528
+ - Added a Kernel-only Reddit workflow that searches DuckDuckGo with a `site:reddit.com` query, switches the same browser to a residential proxy before clicking the selected result, and reads modern Reddit posts plus bounded rendered-comment expansion through dedicated search, thread, and combined REST endpoints.
529
+ - Added a bounded managed-browser backup for Reddit thread hydration after the primary Kernel attempt fails or returns fewer than the semantic target, capped at 25 comments with measured bandwidth, duration, CAPTCHA, closure, and provider-cost telemetry.
518
530
 
519
531
  ### Changed
520
532
 
521
- - Routed the production `reddit_thread` and `reddit_trending` MCP tools through modern Reddit on Option 1 residential sessions, with DuckDuckGo site search for trend discovery; removed Google and old Reddit from their active execution path while preserving tool names, billing rates, bounded partial results, and refunds.
522
- - Cost probes now include Reddit Option 1 sessions and any managed-browser fallback bytes and cost in the same request receipt, and identify when the backup contributed to total cost.
533
+ - Routed the production `reddit_thread` and `reddit_trending` MCP tools through modern Reddit on Kernel residential sessions, with DuckDuckGo site search for trend discovery; removed Google and old Reddit from their active execution path while preserving tool names, billing rates, bounded partial results, and refunds.
534
+ - Cost probes now include Reddit Kernel sessions and any managed-browser fallback bytes and cost in the same request receipt, and identify when the backup contributed to total cost.
523
535
 
524
536
  ### Fixed
525
537
 
@@ -564,7 +576,7 @@ All notable changes to MCP Scraper are documented here. The format is based on [
564
576
 
565
577
  ### Fixed
566
578
 
567
- - Persisted per-control PAA dispatch and 0.7/1.0/1.4-second confirmation telemetry in durable checkpoints, exposed recent interaction and attempt correlation through MCP status, attached Option 2 session IDs immediately after browser launch, and finalized dangling attempt rows during lease recovery without blocking customer settlement.
579
+ - Persisted per-control PAA dispatch and 0.7/1.0/1.4-second confirmation telemetry in durable checkpoints, exposed recent interaction and attempt correlation through MCP status, attached Bright Data session IDs immediately after browser launch, and finalized dangling attempt rows during lease recovery without blocking customer settlement.
568
580
  - Prevented inline style, script, and hidden DOM text inside Google answer containers from falsely confirming that PAA answer material loaded.
569
581
  - Routed canonical `/assistant` page loads to the web app and the redacted private Assistant readiness endpoint to the main API function, preventing production 404s after the 0.79.1 launch.
570
582
 
@@ -588,7 +600,7 @@ All notable changes to MCP Scraper are documented here. The format is based on [
588
600
  - Added Scheduling as the canonical Personal Assistant setup surface, with connection readiness for Gmail, Calendar, Zoom, browser profiles, Memory, SMS, and email; exact schedule confirmation; approval and spend review; run history; and explicit watch/takeover states.
589
601
  - Added owner-scoped browser profiles that can hold multiple independently verified login bindings, while every browser schedule grant selects one exact profile, login, domain, and action set.
590
602
  - Added immutable schedule revisions, readiness receipts, append-only activation records, additive legacy schedule projection, and single-owner occurrence transition receipts so migration cannot silently infer browser authority or double-dispatch work.
591
- - Added Option 1 and private-Mac browser runtime boundaries with collision-resistant tenant namespaces, per-owner concurrency ceilings, bounded sessions, explicit and timeout cleanup, owner-qualified account deletion, and provider deletion readback.
603
+ - Added Kernel and private-Mac browser runtime boundaries with collision-resistant tenant namespaces, per-owner concurrency ceilings, bounded sessions, explicit and timeout cleanup, owner-qualified account deletion, and provider deletion readback.
592
604
  - Added an owner-controlled Personal Assistant that brings SMS/MMS, Gmail, Google Calendar, Zoom, browser work, reminders, and Memory context packets into one governed workflow with immutable plans, approval checkpoints, spend limits, and durable receipts.
593
605
  - Added Twilio number discovery, owned-number attachment, purchase and registration previews, Messaging Service readiness, signed inbound and delivery webhooks, safe MMS ingestion, deterministic opt-out handling, single and reviewed bulk messaging, and reconciliation for unknown provider outcomes.
594
606
  - Added immutable, revisioned Memory context packets with source and attachment provenance, Gmail full-message imports, MMS media metadata, lifecycle controls, and readback verification against the selected vault.
@@ -625,7 +637,7 @@ All notable changes to MCP Scraper are documented here. The format is based on [
625
637
  - Made `maxQuestions` an explicit target count rather than a traversal-depth control, with separate discovery and material-completeness diagnostics.
626
638
  - Preserved complete People Also Ask, AI Overview, and organic-result link provenance in JSON, structured MCP output, and CSV while classifying plain links and Google redirect links explicitly.
627
639
  - Resolved opaque Google `/goto` targets through bounded concurrent manual-redirect requests with active-browser interception as a fallback, without following publisher destinations and without dropping unresolved material.
628
- - Aligned the bounded PAA production-provider canary with the public `maxQuestions` contract and made Option 2 the default test provider.
640
+ - Aligned the bounded PAA production-provider canary with the public `maxQuestions` contract and made Bright Data the default test provider.
629
641
 
630
642
  ### Fixed
631
643
 
@@ -833,7 +845,7 @@ All notable changes to MCP Scraper are documented here. The format is based on [
833
845
 
834
846
  - Added portable `harvest_paa_start` and `harvest_paa_status` tools for durable long-running PAA research, with stable idempotency recovery, progress, attempt provenance, completeness, billing state, and bounded provider telemetry.
835
847
  - Added progressive PAA checkpoints that preserve and merge the best unique rows across browser retries and stale-job recovery instead of losing already captured questions when a provider session or caller is interrupted.
836
- - Added exact Option 2 browser-session identity, sanitized Session Logs enrichment, disconnect attribution, bandwidth usage telemetry, and retryable reconciliation without making provider telemetry a prerequisite for result delivery.
848
+ - Added exact Bright Data browser-session identity, sanitized Session Logs enrichment, disconnect attribution, bandwidth usage telemetry, and retryable reconciliation without making provider telemetry a prerequisite for result delivery.
837
849
 
838
850
  ### Changed
839
851
 
@@ -1566,7 +1578,7 @@ All notable changes to MCP Scraper are documented here. The format is based on [
1566
1578
 
1567
1579
  ### Changed
1568
1580
 
1569
- - PAA browser work now uses Option 1's co-located Playwright execution with stealth mode's default managed proxy and native browser metadata. Location is expressed only through Google UULE, CAPTCHA solver waiting is capped at 60 seconds, and a fresh session is allowed once only when no useful data was captured.
1581
+ - PAA browser work now uses Kernel's co-located Playwright execution with stealth mode's default managed proxy and native browser metadata. Location is expressed only through Google UULE, CAPTCHA solver waiting is capped at 60 seconds, and a fresh session is allowed once only when no useful data was captured.
1570
1582
  - PAA invocations stop browser work at 250 seconds inside the 280-second application budget, reserving 30 seconds for persistence, cleanup, and settlement. The legacy cron worker no longer claims Inngest-owned PAA jobs.
1571
1583
 
1572
1584
  ### Fixed
@@ -1802,7 +1814,7 @@ All notable changes to MCP Scraper are documented here. The format is based on [
1802
1814
 
1803
1815
  ### Changed
1804
1816
 
1805
- - `maps_search` now applies a transport ladder across its retry attempts so it can recover from Google soft-blocks instead of only retrying the same way. The first attempt is unchanged (Option 1's default stealth ISP proxy, direct navigation). Subsequent retries switch to direct egress and arrive at Google through a cross-site redirect (the combination that measurably clears blocks a cold navigation triggers); the final escalation attempt uses direct egress without the redirect and accepts any egress country. This only affects the `proxyMode: 'none'` default path and only its retries — a first-attempt success behaves exactly as before.
1817
+ - `maps_search` now applies a transport ladder across its retry attempts so it can recover from Google soft-blocks instead of only retrying the same way. The first attempt is unchanged (Kernel's default stealth ISP proxy, direct navigation). Subsequent retries switch to direct egress and arrive at Google through a cross-site redirect (the combination that measurably clears blocks a cold navigation triggers); the final escalation attempt uses direct egress without the redirect and accepts any egress country. This only affects the `proxyMode: 'none'` default path and only its retries — a first-attempt success behaves exactly as before.
1806
1818
 
1807
1819
  ## [0.32.1] - 2026-07-22
1808
1820
 
package/README.md CHANGED
@@ -175,7 +175,7 @@ Build the branded one-click bundle:
175
175
  npm run build:mcpb
176
176
  ```
177
177
 
178
- The generated bundle is written to `build/mcpb/mcp-scraper-<version>.mcpb` and copied to `public/downloads/` for the hosted download. The current public bundle is `https://mcpscraper.dev/downloads/mcp-scraper.mcpb` (`0.93.0`, SHA-256 `418a8520d7a552ec2c3bdbe0716e2aacc0b58255f397459a31ddb2f14ac97c07`). Install it by opening or dragging it into Claude Desktop. Claude displays the `MCP Scraper` install card, icon, API-key configuration field, and manually curated current-release message from the bundle manifest.
178
+ The generated bundle is written to `build/mcpb/mcp-scraper-<version>.mcpb` and copied to `public/downloads/` for the hosted download. The current public bundle is `https://mcpscraper.dev/downloads/mcp-scraper.mcpb` (`0.93.2`, SHA-256 `1c52ef996152bed5e17ba731f90a883cb4e1ad840826764d62d2e193a91d0a38`). Install it by opening or dragging it into Claude Desktop. Claude displays the `MCP Scraper` install card, icon, API-key configuration field, and manually curated current-release message from the bundle manifest.
179
179
 
180
180
  The MCPB install exposes every tool — web-intelligence plus all `browser_*` tools — through the one `mcp-scraper` server.
181
181
 
@@ -366,8 +366,8 @@ The `mcp-scraper` NPX stdio server also exposes saved reports as MCP resources:
366
366
  - `MCP_SCRAPER_OUTPUT_DIR` is optional and defaults to `~/Downloads/mcp-scraper`.
367
367
  - `MCP_SCRAPER_SAVE_REPORTS=false` disables automatic Markdown report files.
368
368
  - `MCP_SCRAPER_KEY_PATH` is optional. When no API key env var is set, the server also reads `~/.mcp-scraper-key` for compatibility with older installs.
369
- - `BROWSER_AGENT_PROFILE_NAME` is optional and sets the default saved hosted browser profile for `mcp-scraper` stdio sessions. Aliases: `BROWSER_SERVICE_PROFILE_NAME`, `Option 1_BROWSER_PROFILE_NAME`, `Option 1_PROFILE_NAME`.
370
- - `BROWSER_AGENT_PROFILE_SAVE_CHANGES=true` is optional hosted setup behavior. It persists cookies and storage back to the named profile when `browser_close` deletes the hosted browser session. Aliases: `BROWSER_SERVICE_PROFILE_SAVE_CHANGES`, `Option 1_BROWSER_PROFILE_SAVE_CHANGES`, `Option 1_PROFILE_SAVE_CHANGES`.
369
+ - `BROWSER_AGENT_PROFILE_NAME` is optional and sets the default saved hosted browser profile for `mcp-scraper` stdio sessions. Aliases: `BROWSER_SERVICE_PROFILE_NAME`, `KERNEL_BROWSER_PROFILE_NAME`, `KERNEL_PROFILE_NAME`.
370
+ - `BROWSER_AGENT_PROFILE_SAVE_CHANGES=true` is optional hosted setup behavior. It persists cookies and storage back to the named profile when `browser_close` deletes the hosted browser session. Aliases: `BROWSER_SERVICE_PROFILE_SAVE_CHANGES`, `KERNEL_BROWSER_PROFILE_SAVE_CHANGES`, `KERNEL_PROFILE_SAVE_CHANGES`.
371
371
 
372
372
  Hosted operators can isolate authorization state in a dedicated Turso/libSQL database without changing the public MCP tool catalog or API-key authentication. The secured store uses atomic authorization-code exchange and refresh rotation, keyed secret lookup, bounded encrypted replay receipts, authority epochs, and fail-closed maintenance behavior. Production migration and rollback are controlled data moves, not ordinary mode flips; see [MCP OAuth operations](docs/operations/mcp-oauth-runbook.md). Existing client setup and reconnect behavior are unchanged in Phase 1.
373
373
 
@@ -1,3 +1,3 @@
1
1
  #!/usr/bin/env node
2
2
  import{readFileSync as s}from"fs";function c(){try{for(let r of s(".env","utf8").split(`
3
- `)){let o=r.indexOf("=");if(o<1||r.trimStart().startsWith("#"))continue;let e=r.slice(0,o).trim();process.env[e]||(process.env[e]=r.slice(o+1).trim())}}catch{}}c();async function a(){let[{serve:r},{app:o},{startWorker:e},{migrate:i}]=await Promise.all([import("@hono/node-server"),import("../server-JLR2YJ7V.js"),import("../worker-UG4OXZWR.js"),import("../db-B5XJTOGN.js")]),n=parseInt(process.env.PORT??"3001");try{if(await i(),process.env.ANALYTICS_DATABASE_URL){let{migrateAnalytics:t}=await import("../analytics-repository-IHOFBSUV.js");await t()}e(),r({fetch:o.fetch,port:n},t=>{console.log(`[server] http://localhost:${t.port}`),console.log(`[server] admin auth: ${process.env.ADMIN_KEY?"configured":"not configured"}`)})}catch(t){console.error("[startup] server preflight failed",t instanceof Error?t.name:"unknown_error"),process.exit(1)}}a();
3
+ `)){let o=r.indexOf("=");if(o<1||r.trimStart().startsWith("#"))continue;let e=r.slice(0,o).trim();process.env[e]||(process.env[e]=r.slice(o+1).trim())}}catch{}}c();async function a(){let[{serve:r},{app:o},{startWorker:e},{migrate:i}]=await Promise.all([import("@hono/node-server"),import("../server-GRB6VNT6.js"),import("../worker-US6CTFYG.js"),import("../db-B5XJTOGN.js")]),n=parseInt(process.env.PORT??"3001");try{if(await i(),process.env.ANALYTICS_DATABASE_URL){let{migrateAnalytics:t}=await import("../analytics-repository-IHOFBSUV.js");await t()}e(),r({fetch:o.fetch,port:n},t=>{console.log(`[server] http://localhost:${t.port}`),console.log(`[server] admin auth: ${process.env.ADMIN_KEY?"configured":"not configured"}`)})}catch(t){console.error("[startup] server preflight failed",t instanceof Error?t.name:"unknown_error"),process.exit(1)}}a();
@@ -1,5 +1,5 @@
1
1
  #!/usr/bin/env node
2
- import{b as E,c as N,d as M,e as L,f as T,j as U}from"../chunk-W2T4NTCE.js";import"../chunk-KJQXUZ4Y.js";import"../chunk-Y6MKMSOC.js";import{g as v,j as b,k as D,l as K,m as H}from"../chunk-WO3N5FH2.js";import"../chunk-HE45FFBU.js";import{a as P}from"../chunk-DUJ56SHU.js";import{Command as he}from"commander";import{spawn as ne}from"child_process";import{mkdir as ke,writeFile as Pe}from"fs/promises";import{basename as Ce,join as Z}from"path";function se(e){return e.apiKey?.trim()||"sk_live_your_key"}function ce(e){return e.packageSpec?.trim()||"mcp-scraper@latest"}function A(e={}){return["-y","--package",ce(e),"mcp-scraper"]}function pe(e){let n={MCP_SCRAPER_API_KEY:se(e)},c=e.browserProfileName?.trim();return c&&(n.BROWSER_AGENT_PROFILE_NAME=c),e.browserProfileSaveChanges===!0&&(n.BROWSER_AGENT_PROFILE_SAVE_CHANGES="true"),n}function q(){return["mcp","remove","mcp-scraper","-s","user"]}function J(){return["mcp","get","mcp-scraper"]}function B(e){let n=e.match(/^\s*Command:\s*(.+?)\s*$/m)?.[1];if(!n)return null;let c=e.match(/^\s*Args:\s*(.*?)\s*$/m)?.[1]??"",i=c.length?c.split(/\s+/):[],p={},u=e.split(/^\s*Environment:\s*$/m)[1];if(u)for(let a of u.split(`
2
+ import{b as E,c as N,d as M,e as L,f as T,j as U}from"../chunk-W2T4NTCE.js";import"../chunk-KJQXUZ4Y.js";import"../chunk-Y6MKMSOC.js";import{g as v,j as b,k as D,l as K,m as H}from"../chunk-WO3N5FH2.js";import"../chunk-HE45FFBU.js";import{a as P}from"../chunk-XJ6PJWT5.js";import{Command as he}from"commander";import{spawn as ne}from"child_process";import{mkdir as ke,writeFile as Pe}from"fs/promises";import{basename as Ce,join as Z}from"path";function se(e){return e.apiKey?.trim()||"sk_live_your_key"}function ce(e){return e.packageSpec?.trim()||"mcp-scraper@latest"}function A(e={}){return["-y","--package",ce(e),"mcp-scraper"]}function pe(e){let n={MCP_SCRAPER_API_KEY:se(e)},c=e.browserProfileName?.trim();return c&&(n.BROWSER_AGENT_PROFILE_NAME=c),e.browserProfileSaveChanges===!0&&(n.BROWSER_AGENT_PROFILE_SAVE_CHANGES="true"),n}function q(){return["mcp","remove","mcp-scraper","-s","user"]}function J(){return["mcp","get","mcp-scraper"]}function B(e){let n=e.match(/^\s*Command:\s*(.+?)\s*$/m)?.[1];if(!n)return null;let c=e.match(/^\s*Args:\s*(.*?)\s*$/m)?.[1]??"",i=c.length?c.split(/\s+/):[],p={},u=e.split(/^\s*Environment:\s*$/m)[1];if(u)for(let a of u.split(`
3
3
  `)){let l=a.match(/^\s{2,}([A-Za-z_][A-Za-z0-9_]*)=(.*)$/);if(!l){if(a.trim().length&&!/^\s{2,}/.test(a))break;continue}p[l[1]]=l[2]}return{command:n,args:i,env:p}}function j(e){let n=["mcp","add","mcp-scraper","--scope","user"];for(let[c,i]of Object.entries(e.env))n.push("--env",`${c}=${i}`);return n.push("--",e.command,...e.args),n}function G(e={}){let n=["mcp","add","mcp-scraper","--scope","user"];for(let[c,i]of Object.entries(pe(e)))n.push("--env",`${c}=${i}`);return n.push("--","npx",...A(e)),n}function O(e){if(e==="claude-code")return"claude";if(e==="claude"||D.hosts.some(n=>n.id===e))return e;throw new Error('Unknown host "'+e+'". Use: codex, claude, claude-code, claude-desktop, cursor, windsurf, cline, or user-action-only')}function ue(e){return K(e==="claude"?"claude-code":e)}function W(e,n={}){let c=O(e),i=ue(c),p="Restart the MCP client so it starts a fresh npx process.",u='MCP_SCRAPER_API_KEY="$MCP_SCRAPER_API_KEY" npx -y -p mcp-scraper@latest mcp-scraper-cli agent install claude --apply',a=`X-Ray install protocol: ${v} (${b})`;return c==="codex"?["# Codex MCP config",a,i.exactConfig,"",`Continuation: ${i.continuation}`,`Rollback: ${i.rollback}`,"",p].join(`
4
4
  `):c==="claude"?["# Claude Code command",a,i.exactConfig,"","# One-command Claude Code setup",u,"",`Continuation: ${i.continuation}`,`Rollback: ${i.rollback}`,"",p].join(`
5
5
  `):c==="claude-desktop"?["# Claude Desktop config",a,i.exactConfig,"","Desktop Extension: https://mcpscraper.dev/downloads/mcp-scraper.mcpb",`Continuation: ${i.continuation}`,`Rollback: ${i.rollback}`,p].join(`
@@ -1,2 +1,2 @@
1
1
  #!/usr/bin/env node
2
- import{a as e}from"../chunk-4YLZR4OR.js";import"../chunk-VRVO4OQ3.js";import"../chunk-DT2FYN6N.js";import"../chunk-W2BVJ7S2.js";import"../chunk-RK2VCTZI.js";import"../chunk-ORB4RHCK.js";import"../chunk-TMB56NCA.js";import"../chunk-HUV2WTRW.js";import"../chunk-YGBTTW5D.js";import"../chunk-RZEZJAID.js";import"../chunk-SH5KB4P7.js";import"../chunk-2SP57VCG.js";import"../chunk-6DTXIZY2.js";import"../chunk-35QGIWR6.js";import"../chunk-WO3N5FH2.js";import"../chunk-HE45FFBU.js";import"../chunk-DUJ56SHU.js";import"../chunk-WJ4XFLS4.js";var _=["harvest_paa","search_serp","extract_url","diff_page","map_site_urls","map_wayback_snapshots","extract_site","analyze_site_similarity","audit_site","check_site_export","site_export_read","site_export_image","archive_read","youtube_harvest","youtube_transcribe","facebook_page_intel","facebook_ad_search","reddit_thread","reddit_trending","video_frame_analysis","video_frame_analysis_status","facebook_ad_transcribe","google_ads_search","google_ads_page_intel","google_ads_transcribe","facebook_video_transcribe","instagram_profile_content","instagram_media_download","maps_place_intel","maps_search","trustpilot_reviews","g2_reviews","capture_serp_snapshot","capture_serp_page_snapshots"];e({toolsets:new Set(["paa","serp"]),allowedToolNames:_});
2
+ import{a as e}from"../chunk-CG2FHIBA.js";import"../chunk-QWSWHUOK.js";import"../chunk-DT2FYN6N.js";import"../chunk-W2BVJ7S2.js";import"../chunk-RK2VCTZI.js";import"../chunk-ORB4RHCK.js";import"../chunk-TMB56NCA.js";import"../chunk-HUV2WTRW.js";import"../chunk-YGBTTW5D.js";import"../chunk-UN6FVDJQ.js";import"../chunk-SH5KB4P7.js";import"../chunk-2SP57VCG.js";import"../chunk-6DTXIZY2.js";import"../chunk-5VB3I7UX.js";import"../chunk-WO3N5FH2.js";import"../chunk-HE45FFBU.js";import"../chunk-XJ6PJWT5.js";import"../chunk-WJ4XFLS4.js";var _=["harvest_paa","search_serp","extract_url","diff_page","map_site_urls","map_wayback_snapshots","extract_site","analyze_site_similarity","audit_site","check_site_export","site_export_read","site_export_image","archive_read","youtube_harvest","youtube_transcribe","facebook_page_intel","facebook_ad_search","reddit_thread","reddit_trending","video_frame_analysis","video_frame_analysis_status","facebook_ad_transcribe","google_ads_search","google_ads_page_intel","google_ads_transcribe","facebook_video_transcribe","instagram_profile_content","instagram_media_download","maps_place_intel","maps_search","trustpilot_reviews","g2_reviews","capture_serp_snapshot","capture_serp_page_snapshots"];e({toolsets:new Set(["paa","serp"]),allowedToolNames:_});
@@ -1,3 +1,3 @@
1
1
  #!/usr/bin/env node
2
- import{a as s}from"../chunk-35QGIWR6.js";import{a as e}from"../chunk-DUJ56SHU.js";var r=process.argv.includes("--no-color")||process.env.NO_COLOR!==void 0||process.env.FORCE_COLOR==="0"||!process.stdout.isTTY,n=process.argv.includes("--help")||process.argv.includes("-h");n&&(process.stdout.write(["Usage: mcp-scraper-install [--no-color]","","Prints the branded MCP Scraper terminal install card and copyable install commands.","mcp-scraper prints the same card in a human terminal and runs as the MCP stdio server in clients.",""].join(`
2
+ import{a as s}from"../chunk-5VB3I7UX.js";import{a as e}from"../chunk-XJ6PJWT5.js";var r=process.argv.includes("--no-color")||process.env.NO_COLOR!==void 0||process.env.FORCE_COLOR==="0"||!process.stdout.isTTY,n=process.argv.includes("--help")||process.argv.includes("-h");n&&(process.stdout.write(["Usage: mcp-scraper-install [--no-color]","","Prints the branded MCP Scraper terminal install card and copyable install commands.","mcp-scraper prints the same card in a human terminal and runs as the MCP stdio server in clients.",""].join(`
3
3
  `)),process.exit(0));process.stdout.write(s({version:e,color:!r,apiKeyConfigured:!!process.env.MCP_SCRAPER_API_KEY?.trim()}));
@@ -1,2 +1,2 @@
1
1
  #!/usr/bin/env node
2
- import{a as r}from"../chunk-4YLZR4OR.js";import"../chunk-VRVO4OQ3.js";import"../chunk-DT2FYN6N.js";import"../chunk-W2BVJ7S2.js";import"../chunk-RK2VCTZI.js";import"../chunk-ORB4RHCK.js";import"../chunk-TMB56NCA.js";import"../chunk-HUV2WTRW.js";import"../chunk-YGBTTW5D.js";import"../chunk-RZEZJAID.js";import"../chunk-SH5KB4P7.js";import"../chunk-2SP57VCG.js";import"../chunk-6DTXIZY2.js";import"../chunk-35QGIWR6.js";import"../chunk-WO3N5FH2.js";import"../chunk-HE45FFBU.js";import"../chunk-DUJ56SHU.js";import"../chunk-WJ4XFLS4.js";r();
2
+ import{a as r}from"../chunk-CG2FHIBA.js";import"../chunk-QWSWHUOK.js";import"../chunk-DT2FYN6N.js";import"../chunk-W2BVJ7S2.js";import"../chunk-RK2VCTZI.js";import"../chunk-ORB4RHCK.js";import"../chunk-TMB56NCA.js";import"../chunk-HUV2WTRW.js";import"../chunk-YGBTTW5D.js";import"../chunk-UN6FVDJQ.js";import"../chunk-SH5KB4P7.js";import"../chunk-2SP57VCG.js";import"../chunk-6DTXIZY2.js";import"../chunk-5VB3I7UX.js";import"../chunk-WO3N5FH2.js";import"../chunk-HE45FFBU.js";import"../chunk-XJ6PJWT5.js";import"../chunk-WJ4XFLS4.js";r();
@@ -1,2 +1,2 @@
1
1
  #!/usr/bin/env node
2
- import{w as t}from"../chunk-M5VZVMPC.js";import"../chunk-M5QHXNFZ.js";import{b as r}from"../chunk-SH5KB4P7.js";import"../chunk-2SP57VCG.js";import"../chunk-6DTXIZY2.js";import"../chunk-Y6MKMSOC.js";import"../chunk-WJ4XFLS4.js";import{Command as s,Option as a}from"commander";var i=new s;i.name("paa-harvest").description("Recursively extract Google People Also Ask questions").requiredOption("-q, --query <query>","Seed query").option("-l, --location <location>",'Location name (e.g. "austin" or "Austin,Texas,United States")').option("--gl <gl>","Google country code","us").option("--hl <hl>","Google language code","en").option("-d, --depth <depth>","BFS depth (1-30)","3").option("-m, --max-questions <n>","Max questions to harvest","100").option("-o, --output <dir>","Output directory","./paa-output").option("-f, --format <format>","Output format: json, csv, or both","both").option("--headless","Run browser in headless mode",!1).option("--profile <dir>","Persistent browser profile directory").option("--proxy <url>","Proxy server URL").option("--browser-api-key <key>","Browser service API key (or set BROWSER_SERVICE_API_KEY env var)").addOption(new a("--\u006b\u0065\u0072\u006e\u0065\u006c-api-key <key>").hideHelp()).action(async e=>{try{let o=await t({query:e.query,location:e.location,gl:e.gl,hl:e.hl,depth:parseInt(e.depth,10),maxQuestions:parseInt(e.maxQuestions,10),outputDir:e.output,format:e.format,headless:e.headless,profileDir:e.profile,proxy:e.proxy,\u006b\u0065\u0072\u006e\u0065\u006cApiKey:e.browserApiKey??e.\u006b\u0065\u0072\u006e\u0065\u006cApiKey??r()});console.log(JSON.stringify({totalQuestions:o.totalQuestions,outputDir:o.stats.seed}))}catch(o){console.error(o instanceof Error?o.message:String(o)),process.exit(1)}});async function n(){await i.parseAsync()}n();
2
+ import{w as t}from"../chunk-MJ2OICVY.js";import"../chunk-VVG7LFFF.js";import{b as r}from"../chunk-SH5KB4P7.js";import"../chunk-2SP57VCG.js";import"../chunk-6DTXIZY2.js";import"../chunk-Y6MKMSOC.js";import"../chunk-WJ4XFLS4.js";import{Command as s,Option as a}from"commander";var i=new s;i.name("paa-harvest").description("Recursively extract Google People Also Ask questions").requiredOption("-q, --query <query>","Seed query").option("-l, --location <location>",'Location name (e.g. "austin" or "Austin,Texas,United States")').option("--gl <gl>","Google country code","us").option("--hl <hl>","Google language code","en").option("-d, --depth <depth>","BFS depth (1-30)","3").option("-m, --max-questions <n>","Max questions to harvest","100").option("-o, --output <dir>","Output directory","./paa-output").option("-f, --format <format>","Output format: json, csv, or both","both").option("--headless","Run browser in headless mode",!1).option("--profile <dir>","Persistent browser profile directory").option("--proxy <url>","Proxy server URL").option("--browser-api-key <key>","Browser service API key (or set BROWSER_SERVICE_API_KEY env var)").addOption(new a("--kernel-api-key <key>").hideHelp()).action(async e=>{try{let o=await t({query:e.query,location:e.location,gl:e.gl,hl:e.hl,depth:parseInt(e.depth,10),maxQuestions:parseInt(e.maxQuestions,10),outputDir:e.output,format:e.format,headless:e.headless,profileDir:e.profile,proxy:e.proxy,kernelApiKey:e.browserApiKey??e.kernelApiKey??r()});console.log(JSON.stringify({totalQuestions:o.totalQuestions,outputDir:o.stats.seed}))}catch(o){console.error(o instanceof Error?o.message:String(o)),process.exit(1)}});async function n(){await i.parseAsync()}n();
@@ -1,15 +1,15 @@
1
- import{c as l}from"./chunk-6DTXIZY2.js";import{t as _}from"./chunk-WJ4XFLS4.js";var c=166667e-10,N=.0001333336,I=new Set(["serp","fb_search","fb_ad","instagram"]),p=.00111,L=4,R=4e-4;function o(e,r){let t=process.env[e]?.trim();if(!t)return r;let s=Number(t);return Number.isFinite(s)&&s>=0?s:r}var A=o("NANGO_USD_PER_CONNECTION_MONTH",1),U=o("NANGO_USD_PER_FUNCTION_RUN",1e-4),m=o("NANGO_USD_PER_PROXY_REQUEST",1e-4),g=o("NANGO_USD_PER_COMPUTE_SEC",2e-4),O=o("\u0042\u0052\u0049\u0047\u0048\u0054\u0044\u0041\u0054\u0041_BROWSER_USD_PER_GB",8),D=o("\u0042\u0052\u0049\u0047\u0048\u0054\u0044\u0041\u0054\u0041_SERP_USD_PER_REQUEST",.002836168);function u(e,r){return Math.max(0,e)/1e3*(r?N:c)}function d(e,r){return e==="fal_wizper"?Math.max(0,r)/L*p:e==="deepinfra_qwen"?Math.max(0,r)/1e3*R:e==="openrouter"||e==="mcp_memory_video"||e==="mcp_memory_ai"?Math.max(0,r):e==="nango_connection"?Math.max(0,r)*A:e==="nango_function_run"?Math.max(0,r)*U:e==="nango_proxy_request"?Math.max(0,r)*m:e==="nango_compute"?Math.max(0,r)*g:e==="\u0062\u0072\u0069\u0067\u0068\u0074\u0064\u0061\u0074\u0061_browser_api"?Math.max(0,r)/1e9*O:0}import{randomUUID as i}from"crypto";var E=!1,n=null;async function T(){if(!E)return n||(n=S().finally(()=>{n=null}),n)}async function S(){let e=_(),r=await e.execute(`
1
+ import{c as l}from"./chunk-6DTXIZY2.js";import{t as _}from"./chunk-WJ4XFLS4.js";var c=166667e-10,N=.0001333336,I=new Set(["serp","fb_search","fb_ad","instagram"]),p=.00111,L=4,R=4e-4;function o(e,r){let t=process.env[e]?.trim();if(!t)return r;let s=Number(t);return Number.isFinite(s)&&s>=0?s:r}var A=o("NANGO_USD_PER_CONNECTION_MONTH",1),U=o("NANGO_USD_PER_FUNCTION_RUN",1e-4),m=o("NANGO_USD_PER_PROXY_REQUEST",1e-4),g=o("NANGO_USD_PER_COMPUTE_SEC",2e-4),O=o("BRIGHTDATA_BROWSER_USD_PER_GB",8),D=o("BRIGHTDATA_SERP_USD_PER_REQUEST",.002836168);function u(e,r){return Math.max(0,e)/1e3*(r?N:c)}function d(e,r){return e==="fal_wizper"?Math.max(0,r)/L*p:e==="deepinfra_qwen"?Math.max(0,r)/1e3*R:e==="openrouter"||e==="mcp_memory_video"||e==="mcp_memory_ai"?Math.max(0,r):e==="nango_connection"?Math.max(0,r)*A:e==="nango_function_run"?Math.max(0,r)*U:e==="nango_proxy_request"?Math.max(0,r)*m:e==="nango_compute"?Math.max(0,r)*g:e==="brightdata_browser_api"?Math.max(0,r)/1e9*O:0}import{randomUUID as i}from"crypto";var E=!1,n=null;async function T(){if(!E)return n||(n=S().finally(()=>{n=null}),n)}async function S(){let e=_(),r=await e.execute(`
2
2
  SELECT
3
- (SELECT COUNT(*) FROM sqlite_master WHERE type = 'table' AND name IN ('\u006b\u0065\u0072\u006e\u0065\u006c_session_log', 'vendor_usage_log', 'cost_probe_runs')) = 3
4
- AND (SELECT COUNT(*) FROM pragma_table_info('\u006b\u0065\u0072\u006e\u0065\u006c_session_log') WHERE name IN ('proxy_source', 'proxy_type', 'method')) = 3
3
+ (SELECT COUNT(*) FROM sqlite_master WHERE type = 'table' AND name IN ('kernel_session_log', 'vendor_usage_log', 'cost_probe_runs')) = 3
4
+ AND (SELECT COUNT(*) FROM pragma_table_info('kernel_session_log') WHERE name IN ('proxy_source', 'proxy_type', 'method')) = 3
5
5
  AND (SELECT COUNT(*) FROM pragma_table_info('vendor_usage_log') WHERE name IN ('method', 'source_key', 'provider_duration_ms', 'provider_captcha', 'provider_status')) = 5
6
6
  AND (SELECT COUNT(*) FROM sqlite_master WHERE type = 'index' AND name = 'vendor_usage_log_vendor_source_key') = 1
7
7
  AND (SELECT COUNT(*) FROM pragma_table_info('cost_probe_runs') WHERE name IN ('units', 'unit_type', 'mode')) = 3
8
8
  AS ready
9
9
  `);if(Number(r.rows[0]?.ready??0)===1){E=!0;return}await e.execute(`
10
- CREATE TABLE IF NOT EXISTS \u006b\u0065\u0072\u006e\u0065\u006c_session_log (
10
+ CREATE TABLE IF NOT EXISTS kernel_session_log (
11
11
  id TEXT PRIMARY KEY,
12
- \u006b\u0065\u0072\u006e\u0065\u006c_session_id TEXT,
12
+ kernel_session_id TEXT,
13
13
  op TEXT,
14
14
  source TEXT NOT NULL,
15
15
  probe_run_id TEXT,
@@ -26,7 +26,7 @@ import{c as l}from"./chunk-6DTXIZY2.js";import{t as _}from"./chunk-WJ4XFLS4.js";
26
26
  error TEXT,
27
27
  created_at TEXT NOT NULL DEFAULT (datetime('now'))
28
28
  )
29
- `),await e.execute("CREATE INDEX IF NOT EXISTS \u006b\u0065\u0072\u006e\u0065\u006c_session_log_op ON \u006b\u0065\u0072\u006e\u0065\u006c_session_log(op)"),await e.execute("CREATE INDEX IF NOT EXISTS \u006b\u0065\u0072\u006e\u0065\u006c_session_log_probe ON \u006b\u0065\u0072\u006e\u0065\u006c_session_log(probe_run_id)"),await e.execute("CREATE INDEX IF NOT EXISTS \u006b\u0065\u0072\u006e\u0065\u006c_session_log_created ON \u006b\u0065\u0072\u006e\u0065\u006c_session_log(created_at)");try{await e.execute("ALTER TABLE \u006b\u0065\u0072\u006e\u0065\u006c_session_log ADD COLUMN proxy_source TEXT")}catch{}try{await e.execute("ALTER TABLE \u006b\u0065\u0072\u006e\u0065\u006c_session_log ADD COLUMN proxy_type TEXT")}catch{}try{await e.execute("ALTER TABLE \u006b\u0065\u0072\u006e\u0065\u006c_session_log ADD COLUMN method TEXT")}catch{}await e.execute(`
29
+ `),await e.execute("CREATE INDEX IF NOT EXISTS kernel_session_log_op ON kernel_session_log(op)"),await e.execute("CREATE INDEX IF NOT EXISTS kernel_session_log_probe ON kernel_session_log(probe_run_id)"),await e.execute("CREATE INDEX IF NOT EXISTS kernel_session_log_created ON kernel_session_log(created_at)");try{await e.execute("ALTER TABLE kernel_session_log ADD COLUMN proxy_source TEXT")}catch{}try{await e.execute("ALTER TABLE kernel_session_log ADD COLUMN proxy_type TEXT")}catch{}try{await e.execute("ALTER TABLE kernel_session_log ADD COLUMN method TEXT")}catch{}await e.execute(`
30
30
  CREATE TABLE IF NOT EXISTS vendor_usage_log (
31
31
  id TEXT PRIMARY KEY,
32
32
  op TEXT,
@@ -54,20 +54,20 @@ import{c as l}from"./chunk-6DTXIZY2.js";import{t as _}from"./chunk-WJ4XFLS4.js";
54
54
  success INTEGER NOT NULL DEFAULT 0,
55
55
  error TEXT,
56
56
  http_only INTEGER NOT NULL DEFAULT 0,
57
- \u006b\u0065\u0072\u006e\u0065\u006c_used INTEGER NOT NULL DEFAULT 0,
58
- \u006b\u0065\u0072\u006e\u0065\u006c_sessions INTEGER NOT NULL DEFAULT 0,
59
- \u006b\u0065\u0072\u006e\u0065\u006c_seconds_total REAL NOT NULL DEFAULT 0,
60
- \u006b\u0065\u0072\u006e\u0065\u006c_tier_observed TEXT,
61
- est_\u006b\u0065\u0072\u006e\u0065\u006c_cost_usd_headless REAL NOT NULL DEFAULT 0,
62
- est_\u006b\u0065\u0072\u006e\u0065\u006c_cost_usd_headful REAL NOT NULL DEFAULT 0,
57
+ kernel_used INTEGER NOT NULL DEFAULT 0,
58
+ kernel_sessions INTEGER NOT NULL DEFAULT 0,
59
+ kernel_seconds_total REAL NOT NULL DEFAULT 0,
60
+ kernel_tier_observed TEXT,
61
+ est_kernel_cost_usd_headless REAL NOT NULL DEFAULT 0,
62
+ est_kernel_cost_usd_headful REAL NOT NULL DEFAULT 0,
63
63
  vendor_cost_usd REAL NOT NULL DEFAULT 0,
64
64
  charged_credits REAL NOT NULL DEFAULT 0,
65
65
  margin_usd_headless_starter REAL NOT NULL DEFAULT 0,
66
66
  margin_usd_headful_starter REAL NOT NULL DEFAULT 0,
67
67
  notes TEXT
68
68
  )
69
- `),await e.execute("CREATE INDEX IF NOT EXISTS cost_probe_runs_tool ON cost_probe_runs(tool)");try{await e.execute("ALTER TABLE cost_probe_runs ADD COLUMN units REAL")}catch{}try{await e.execute("ALTER TABLE cost_probe_runs ADD COLUMN unit_type TEXT")}catch{}try{await e.execute("ALTER TABLE cost_probe_runs ADD COLUMN mode TEXT")}catch{}E=!0}async function f(e){try{await T();let r=l(),t=Math.max(0,e.closedAtMs-e.openedAtMs);await _().execute({sql:`INSERT INTO \u006b\u0065\u0072\u006e\u0065\u006c_session_log
70
- (id, \u006b\u0065\u0072\u006e\u0065\u006c_session_id, op, source, probe_run_id, user_id, stealth, headless_sent, proxy_used, proxy_source, proxy_type, fallback, opened_at, closed_at, duration_ms, est_cost_usd_headless, est_cost_usd_headful, error, method)
71
- VALUES (?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?)`,args:[i(),e.\u006b\u0065\u0072\u006e\u0065\u006cSessionId??null,r?.op??null,e.source,r?.probeRunId??null,r?.userId??null,a(e.stealth),a(e.headlessSent),a(e.proxyUsed),e.proxySource??null,e.proxyType??null,e.fallback?1:0,new Date(e.openedAtMs).toISOString(),new Date(e.closedAtMs).toISOString(),t,u(t,!1),u(t,!0),e.error??null,r?.subOp??null]})}catch(r){console.warn("[cost-telemetry] record\u004b\u0065\u0072\u006e\u0065\u006cSession failed:",r instanceof Error?r.message:String(r))}}async function v(e){try{await T();let r=l();return(await _().execute({sql:`INSERT OR IGNORE INTO vendor_usage_log
69
+ `),await e.execute("CREATE INDEX IF NOT EXISTS cost_probe_runs_tool ON cost_probe_runs(tool)");try{await e.execute("ALTER TABLE cost_probe_runs ADD COLUMN units REAL")}catch{}try{await e.execute("ALTER TABLE cost_probe_runs ADD COLUMN unit_type TEXT")}catch{}try{await e.execute("ALTER TABLE cost_probe_runs ADD COLUMN mode TEXT")}catch{}E=!0}async function f(e){try{await T();let r=l(),t=Math.max(0,e.closedAtMs-e.openedAtMs);await _().execute({sql:`INSERT INTO kernel_session_log
70
+ (id, kernel_session_id, op, source, probe_run_id, user_id, stealth, headless_sent, proxy_used, proxy_source, proxy_type, fallback, opened_at, closed_at, duration_ms, est_cost_usd_headless, est_cost_usd_headful, error, method)
71
+ VALUES (?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?,?)`,args:[i(),e.kernelSessionId??null,r?.op??null,e.source,r?.probeRunId??null,r?.userId??null,a(e.stealth),a(e.headlessSent),a(e.proxyUsed),e.proxySource??null,e.proxyType??null,e.fallback?1:0,new Date(e.openedAtMs).toISOString(),new Date(e.closedAtMs).toISOString(),t,u(t,!1),u(t,!0),e.error??null,r?.subOp??null]})}catch(r){console.warn("[cost-telemetry] recordKernelSession failed:",r instanceof Error?r.message:String(r))}}async function v(e){try{await T();let r=l();return(await _().execute({sql:`INSERT OR IGNORE INTO vendor_usage_log
72
72
  (id, op, probe_run_id, user_id, vendor, model, units, unit_type, est_cost_usd, error, method, source_key, provider_duration_ms, provider_captcha, provider_status)
73
73
  VALUES (?,?,?,?,?,?,?,?,?,?,?,?,?,?,?)`,args:[i(),e.op??r?.op??null,e.probeRunId??r?.probeRunId??null,e.userId??r?.userId??null,e.vendor,e.model??null,e.units,e.unitType,d(e.vendor,e.units),e.error??null,e.method??r?.subOp??null,e.sourceKey??null,e.providerDurationMs??null,a(e.providerCaptcha),e.providerStatus??null]})).rowsAffected===1}catch(r){return console.warn("[cost-telemetry] recordVendorUsage failed:",r instanceof Error?r.message:String(r)),!1}}function a(e){return e==null?null:e?1:0}export{c as a,N as b,I as c,u as d,d as e,T as f,f as g,v as h};
@@ -1,4 +1,4 @@
1
- import{a as j,b as T,c as P,d as C,f as S}from"./chunk-TMB56NCA.js";import{ba as w}from"./chunk-RZEZJAID.js";import{Kb as N,Tb as D,d as m,e as I,i as u,t as s}from"./chunk-WJ4XFLS4.js";import{createHash as W}from"crypto";import{gunzipSync as $,gzipSync as q}from"zlib";var y="site-extracts/",L=10080*60*1e3,H=900*1e3;function U(){return process.env.SITE_EXTRACT_ARTIFACT_READ_WRITE_TOKEN?.trim()||process.env.PRIVATE_ARTIFACT_READ_WRITE_TOKEN?.trim()||process.env.CONNECTED_DATA_READ_WRITE_TOKEN?.trim()||process.env.CONNECTED_DATA_BLOB_READ_WRITE_TOKEN?.trim()||null}function v(){return process.env.VERCEL==="1"||process.env.NODE_ENV==="production"}function p(){return{prefix:y,artifactTtlMs:L,downloadTtlMs:H,token:U()}}function h(t){return j(t,y)}async function V(t){let e=await P({policy:p(),ownerId:t.ownerId,artifactKey:`${t.jobId}.zip`,createdAt:t.createdAt,filename:`${t.jobId}-site-export.zip`,contentType:"application/zip",content:t.content});return{key:e.artifactId,url:e.downloadUrl??"",bytes:e.bytes,contentType:e.contentType,filename:e.filename,sha256:e.sha256,expiresAt:e.expiresAt,downloadUrlExpiresAt:e.downloadUrlExpiresAt,kind:"bundle"}}async function G(t){let e=await T({policy:p(),ownerId:t.ownerId,scopeSegments:[t.jobId,"images"],artifactKey:`${t.imageId}.bin`,createdAt:new Date,filename:t.filename,contentType:t.contentType,content:t.content});return{key:e.artifactId,url:e.downloadUrl??"",bytes:e.bytes,contentType:e.contentType,filename:e.filename,sha256:e.sha256,expiresAt:e.expiresAt,downloadUrlExpiresAt:e.downloadUrlExpiresAt,kind:"image",imageId:t.imageId,sourceUrl:t.sourceUrl,sourcePage:t.sourcePage}}async function O(t){return T({policy:p(),ownerId:t.ownerId,scopeSegments:[t.jobId,"content"],artifactKey:`${t.chunkKey}.json.gz`,createdAt:t.createdAt,filename:`${t.chunkKey}.json.gz`,contentType:"application/gzip",content:t.content})}async function Q(t){return C({policy:p(),artifactId:t.artifactId,ownerId:t.ownerId})}async function J(t){return S({policy:p(),artifactId:t,maxBytes:50*1024*1024})}async function Z(t){return h(t.artifactId)!==t.ownerId?null:J(t.artifactId)}async function k(t){return h(t.artifactId)!==t.ownerId?null:S({policy:p(),artifactId:t.artifactId,maxBytes:25*1024*1024})}async function tt(t){return h(t.artifactId)!==t.ownerId?null:S({policy:p(),artifactId:t.artifactId,maxBytes:10*1024*1024})}async function et(t={}){let e=t.now??new Date,n=U();if(!n)return{deleted:0,store:v()?"none":"local"};if(!t.force&&!(e.getUTCHours()===3&&e.getUTCMinutes()===21))return{deleted:0,store:"private-vercel-blob",skipped:!0};let r=e.getTime()-L,{list:i,del:a}=await import("@vercel/blob"),o,c=0;for(let d=0;d<20;d+=1){let l=await i({prefix:y,token:n,limit:1e3,cursor:o}),f=l.blobs.filter(A=>new Date(A.uploadedAt).getTime()<=r);if(f.length>0&&(await a(f.map(A=>A.pathname),{token:n}),c+=f.length),!l.hasMore||!l.cursor)break;o=l.cursor}return{deleted:c,store:"private-vercel-blob"}}var F=10,b=20*1024*1024;function E(t){return W("sha256").update(t).digest("hex")}function B(t){return!t.acquiredHtml||t.extractionStatus==="failed"?null:{pageId:t.pageId??E(t.archivedUrl??t.url),url:t.archivedUrl??t.url,html:t.acquiredHtml,mainHtml:t.mainHtml??"",markdown:t.bodyMarkdown,htmlSha256:t.htmlSha256??E(t.acquiredHtml),markdownSha256:t.markdownSha256??E(t.bodyMarkdown)}}function R(t){return Buffer.from(JSON.stringify({version:"site-extract-content.v1",pages:t}))}async function M(t){let e=t.pages.flatMap(o=>{let c=B(o);return c?[c]:[]}),n=new Map,r=[],i=0,a=async()=>{if(!r.length)return;let o=R(r);if(o.length>b)throw new Error(`site export content chunk exceeds ${b} bytes`);let c=E(r.map(l=>l.pageId).join("\0")).slice(0,16),d=await O({ownerId:t.ownerId,jobId:t.jobId,createdAt:t.createdAt,chunkKey:`content-${String(i).padStart(4,"0")}-${c}`,content:q(o,{level:6})});for(let l of r)n.set(l.url,{artifactId:d.artifactId,artifactSha256:d.sha256,pageId:l.pageId,uncompressedBytes:o.length});i++,r=[]};for(let o of e){let c=[...r,o];if(r.length>0&&(c.length>F||R(c).length>b)&&await a(),r.push(o),R(r).length>b)throw new Error(`page ${o.pageId} exceeds the durable content chunk limit`)}return await a(),n}function it(t){let e=new Map;return async n=>{let r=e.get(n.artifactId);if(!r){let a=await k({artifactId:n.artifactId,ownerId:t});if(!a)throw new Error("site export content chunk is missing or unauthorized");if(E(a)!==n.artifactSha256)throw new Error("site export content chunk checksum mismatch");let o=$(a);if(o.length>b)throw new Error("site export content chunk exceeds its read limit");if(r=JSON.parse(o.toString("utf8")),r.version!=="site-extract-content.v1"||!Array.isArray(r.pages))throw new Error("site export content chunk has an unsupported contract");for(e.set(n.artifactId,r);e.size>2;)e.delete(e.keys().next().value)}let i=r.pages.find(a=>a.pageId===n.pageId);if(!i)throw new Error("site export page is missing from its content chunk");if(E(i.html)!==i.htmlSha256||E(i.markdown)!==i.markdownSha256)throw new Error("site export page checksum mismatch");return i}}var g="xray_entitlement";function pt(t){let e=Math.max(1,Number(t.options.effectiveMaxPages??t.options.maxPages??1)),n=Math.max(e,Number(t.options.requestedMaxPages??e)),r=n>e;return{requestedMaxPages:n,effectiveMaxPages:e,creditLimited:r,creditTruncated:r&&t.totalUrls>=e}}function Et(t,e=!1){return t.successfulUrls===0?"failed":t.failedUrls>0||t.remainingUrls>0||e?"partial":"complete"}function _(t){let e=Number(t.total_urls??0),n=t.attempted_urls==null,r=Number(n?t.done_urls??0:t.attempted_urls),i=Number(n?t.done_urls??0:t.successful_urls??0),a=Number(t.failed_urls??Math.max(0,r-i));return{id:String(t.id),userId:t.user_id!=null?Number(t.user_id):null,idempotencyKey:t.idempotency_key!=null?String(t.idempotency_key):null,requestFingerprint:t.request_fingerprint!=null?String(t.request_fingerprint):null,status:t.status!=null?String(t.status):"pending",startUrl:String(t.start_url??""),options:t.options?JSON.parse(String(t.options)):{},totalUrls:e,doneUrls:r,attemptedUrls:r,successfulUrls:i,failedUrls:a,remainingUrls:Math.max(0,e-r),artifacts:t.artifacts?JSON.parse(String(t.artifacts)):null,error:t.error!=null?String(t.error):null,publicError:I(t.public_error_json),billedMc:t.billed_mc!=null?Number(t.billed_mc):null,createdAt:String(t.created_at??""),updatedAt:String(t.updated_at??"")}}async function mt(t,e,n,r){await s().execute({sql:`INSERT INTO site_extract_jobs (id, user_id, status, start_url, options, created_at, updated_at)
1
+ import{a as j,b as T,c as P,d as C,f as S}from"./chunk-TMB56NCA.js";import{ba as w}from"./chunk-UN6FVDJQ.js";import{Kb as N,Tb as D,d as m,e as I,i as u,t as s}from"./chunk-WJ4XFLS4.js";import{createHash as W}from"crypto";import{gunzipSync as $,gzipSync as q}from"zlib";var y="site-extracts/",L=10080*60*1e3,H=900*1e3;function U(){return process.env.SITE_EXTRACT_ARTIFACT_READ_WRITE_TOKEN?.trim()||process.env.PRIVATE_ARTIFACT_READ_WRITE_TOKEN?.trim()||process.env.CONNECTED_DATA_READ_WRITE_TOKEN?.trim()||process.env.CONNECTED_DATA_BLOB_READ_WRITE_TOKEN?.trim()||null}function v(){return process.env.VERCEL==="1"||process.env.NODE_ENV==="production"}function p(){return{prefix:y,artifactTtlMs:L,downloadTtlMs:H,token:U()}}function h(t){return j(t,y)}async function V(t){let e=await P({policy:p(),ownerId:t.ownerId,artifactKey:`${t.jobId}.zip`,createdAt:t.createdAt,filename:`${t.jobId}-site-export.zip`,contentType:"application/zip",content:t.content});return{key:e.artifactId,url:e.downloadUrl??"",bytes:e.bytes,contentType:e.contentType,filename:e.filename,sha256:e.sha256,expiresAt:e.expiresAt,downloadUrlExpiresAt:e.downloadUrlExpiresAt,kind:"bundle"}}async function G(t){let e=await T({policy:p(),ownerId:t.ownerId,scopeSegments:[t.jobId,"images"],artifactKey:`${t.imageId}.bin`,createdAt:new Date,filename:t.filename,contentType:t.contentType,content:t.content});return{key:e.artifactId,url:e.downloadUrl??"",bytes:e.bytes,contentType:e.contentType,filename:e.filename,sha256:e.sha256,expiresAt:e.expiresAt,downloadUrlExpiresAt:e.downloadUrlExpiresAt,kind:"image",imageId:t.imageId,sourceUrl:t.sourceUrl,sourcePage:t.sourcePage}}async function O(t){return T({policy:p(),ownerId:t.ownerId,scopeSegments:[t.jobId,"content"],artifactKey:`${t.chunkKey}.json.gz`,createdAt:t.createdAt,filename:`${t.chunkKey}.json.gz`,contentType:"application/gzip",content:t.content})}async function Q(t){return C({policy:p(),artifactId:t.artifactId,ownerId:t.ownerId})}async function J(t){return S({policy:p(),artifactId:t,maxBytes:50*1024*1024})}async function Z(t){return h(t.artifactId)!==t.ownerId?null:J(t.artifactId)}async function k(t){return h(t.artifactId)!==t.ownerId?null:S({policy:p(),artifactId:t.artifactId,maxBytes:25*1024*1024})}async function tt(t){return h(t.artifactId)!==t.ownerId?null:S({policy:p(),artifactId:t.artifactId,maxBytes:10*1024*1024})}async function et(t={}){let e=t.now??new Date,n=U();if(!n)return{deleted:0,store:v()?"none":"local"};if(!t.force&&!(e.getUTCHours()===3&&e.getUTCMinutes()===21))return{deleted:0,store:"private-vercel-blob",skipped:!0};let r=e.getTime()-L,{list:i,del:a}=await import("@vercel/blob"),o,c=0;for(let d=0;d<20;d+=1){let l=await i({prefix:y,token:n,limit:1e3,cursor:o}),f=l.blobs.filter(A=>new Date(A.uploadedAt).getTime()<=r);if(f.length>0&&(await a(f.map(A=>A.pathname),{token:n}),c+=f.length),!l.hasMore||!l.cursor)break;o=l.cursor}return{deleted:c,store:"private-vercel-blob"}}var F=10,b=20*1024*1024;function E(t){return W("sha256").update(t).digest("hex")}function B(t){return!t.acquiredHtml||t.extractionStatus==="failed"?null:{pageId:t.pageId??E(t.archivedUrl??t.url),url:t.archivedUrl??t.url,html:t.acquiredHtml,mainHtml:t.mainHtml??"",markdown:t.bodyMarkdown,htmlSha256:t.htmlSha256??E(t.acquiredHtml),markdownSha256:t.markdownSha256??E(t.bodyMarkdown)}}function R(t){return Buffer.from(JSON.stringify({version:"site-extract-content.v1",pages:t}))}async function M(t){let e=t.pages.flatMap(o=>{let c=B(o);return c?[c]:[]}),n=new Map,r=[],i=0,a=async()=>{if(!r.length)return;let o=R(r);if(o.length>b)throw new Error(`site export content chunk exceeds ${b} bytes`);let c=E(r.map(l=>l.pageId).join("\0")).slice(0,16),d=await O({ownerId:t.ownerId,jobId:t.jobId,createdAt:t.createdAt,chunkKey:`content-${String(i).padStart(4,"0")}-${c}`,content:q(o,{level:6})});for(let l of r)n.set(l.url,{artifactId:d.artifactId,artifactSha256:d.sha256,pageId:l.pageId,uncompressedBytes:o.length});i++,r=[]};for(let o of e){let c=[...r,o];if(r.length>0&&(c.length>F||R(c).length>b)&&await a(),r.push(o),R(r).length>b)throw new Error(`page ${o.pageId} exceeds the durable content chunk limit`)}return await a(),n}function it(t){let e=new Map;return async n=>{let r=e.get(n.artifactId);if(!r){let a=await k({artifactId:n.artifactId,ownerId:t});if(!a)throw new Error("site export content chunk is missing or unauthorized");if(E(a)!==n.artifactSha256)throw new Error("site export content chunk checksum mismatch");let o=$(a);if(o.length>b)throw new Error("site export content chunk exceeds its read limit");if(r=JSON.parse(o.toString("utf8")),r.version!=="site-extract-content.v1"||!Array.isArray(r.pages))throw new Error("site export content chunk has an unsupported contract");for(e.set(n.artifactId,r);e.size>2;)e.delete(e.keys().next().value)}let i=r.pages.find(a=>a.pageId===n.pageId);if(!i)throw new Error("site export page is missing from its content chunk");if(E(i.html)!==i.htmlSha256||E(i.markdown)!==i.markdownSha256)throw new Error("site export page checksum mismatch");return i}}var g="xray_entitlement";function pt(t){let e=Math.max(1,Number(t.options.effectiveMaxPages??t.options.maxPages??1)),n=Math.max(e,Number(t.options.requestedMaxPages??e)),r=n>e;return{requestedMaxPages:n,effectiveMaxPages:e,creditLimited:r,creditTruncated:r&&t.totalUrls>=e}}function Et(t,e=!1){return t.successfulUrls===0?"failed":t.failedUrls>0||t.remainingUrls>0||e?"partial":"complete"}function _(t){let e=Number(t.total_urls??0),n=t.attempted_urls==null,r=Number(n?t.done_urls??0:t.attempted_urls),i=Number(n?t.done_urls??0:t.successful_urls??0),a=Number(t.failed_urls??Math.max(0,r-i));return{id:String(t.id),userId:t.user_id!=null?Number(t.user_id):null,idempotencyKey:t.idempotency_key!=null?String(t.idempotency_key):null,requestFingerprint:t.request_fingerprint!=null?String(t.request_fingerprint):null,status:t.status!=null?String(t.status):"pending",startUrl:String(t.start_url??""),options:t.options?JSON.parse(String(t.options)):{},totalUrls:e,doneUrls:r,attemptedUrls:r,successfulUrls:i,failedUrls:a,remainingUrls:Math.max(0,e-r),artifacts:t.artifacts?JSON.parse(String(t.artifacts)):null,error:t.error!=null?String(t.error):null,publicError:I(t.public_error_json),billedMc:t.billed_mc!=null?Number(t.billed_mc):null,createdAt:String(t.created_at??""),updatedAt:String(t.updated_at??"")}}async function mt(t,e,n,r){await s().execute({sql:`INSERT INTO site_extract_jobs (id, user_id, status, start_url, options, created_at, updated_at)
2
2
  VALUES (?, ?, 'pending', ?, ?, datetime('now'), datetime('now'))`,args:[t,e,n,JSON.stringify(r)]})}async function K(t,e){let r=await s().execute({sql:"SELECT * FROM site_extract_jobs WHERE user_id = ? AND idempotency_key = ? LIMIT 1",args:[t,e]});return r.rows[0]?_(r.rows[0]):null}async function gt(t){let n=await s().execute({sql:`INSERT OR IGNORE INTO site_extract_jobs
3
3
  (id, user_id, status, start_url, options, idempotency_key, request_fingerprint, created_at, updated_at)
4
4
  VALUES (?, ?, 'pending', ?, ?, ?, ?, datetime('now'), datetime('now'))`,args:[t.jobId,t.userId,t.startUrl,JSON.stringify(t.options),t.idempotencyKey,t.requestFingerprint]}),r=await K(t.userId,t.idempotencyKey);if(!r)throw new Error("idempotent extract job was not persisted");return{job:r,created:n.rowsAffected===1,conflict:r.requestFingerprint!==t.requestFingerprint}}async function x(t){let n=await s().execute({sql:"SELECT * FROM site_extract_jobs WHERE id = ?",args:[t]});return n.rows[0]?_(n.rows[0]):null}async function _t(t){return(await s().execute({sql:"SELECT * FROM site_extract_jobs WHERE user_id = ? ORDER BY created_at DESC LIMIT 50",args:[t]})).rows.map(r=>_(r))}async function xt(t=25){return(await s().execute({sql:`SELECT * FROM site_extract_jobs
@@ -1,4 +1,4 @@
1
- var a={message:"search_serp light mode returns organic Google positions, URLs, titles, and descriptions. Full mode adds available same-page SERP features. Both default to one page; request pages:2 explicitly for two. Light costs 20 Credits per delivered page, or 35 Credits when the backup supplies it. Full costs 35 Credits per delivered page."};var _={reset:"\x1B[0m",cyan:"\x1B[36m",lime:"\x1B[32m",amber:"\x1B[33m",red:"\x1B[31m",muted:"\x1B[90m",bold:"\x1B[1m"};function r(t,e,n){return n?`${_[e]}${t}${_.reset}`:t}function o(t,e,n){let s=t.padEnd(9," ");return` ${r(s,"muted",n)} ${e.join(r(" . ","muted",n))}`}function p(t){let e=t.color??!0,n=t.apiKeyConfigured?"$MCP_SCRAPER_API_KEY":"sk_live_your_key",s=String.raw`
1
+ var s={message:"Google Search now preserves delivery evidence for provider-cost reporting even when an upstream billing rate is unavailable. Search inputs, results, and customer prices are unchanged."};var _={reset:"\x1B[0m",cyan:"\x1B[36m",lime:"\x1B[32m",amber:"\x1B[33m",red:"\x1B[31m",muted:"\x1B[90m",bold:"\x1B[1m"};function r(t,e,n){return n?`${_[e]}${t}${_.reset}`:t}function o(t,e,n){let a=t.padEnd(9," ");return` ${r(a,"muted",n)} ${e.join(r(" . ","muted",n))}`}function p(t){let e=t.color??!0,n=t.apiKeyConfigured?"$MCP_SCRAPER_API_KEY":"sk_live_your_key",a=String.raw`
2
2
  __ __ ____ ____
3
3
  | \/ |/ ___| _ \
4
4
  | |\/| | | | |_) |
@@ -10,7 +10,7 @@ var a={message:"search_serp light mode returns organic Google positions, URLs, t
10
10
  \___ \| | | |_) | / _ \ | |_) | _| | |_) |
11
11
  ___) | |___| _ < / ___ \| __/| |___| _ <
12
12
  |____/ \____|_| \_\/_/ \_\_| |_____|_| \_\
13
- `,i=[`MCP_SCRAPER_API_KEY=${n} npx -y -p mcp-scraper@latest \\`," mcp-scraper-cli agent install claude --apply"].join(`
14
- `),c=["[mcp_servers.mcp-scraper]",'command = "npx"','args = ["-y", "-p", "mcp-scraper@latest", "mcp-scraper"]',`env = { MCP_SCRAPER_API_KEY = "${n}" }`].join(`
15
- `);return[r(`mcp-scraper v${t.version}`,"bold",e),r("> mcp-scraper-install","muted",e),r(s,"amber",e),`${r("MCP Scraper Agent","cyan",e)} . v${t.version} . mcpscraper.dev`,"1/1 install surfaces ready",r(`Newest in v${t.version}: ${a.message}`,"lime",e),"",`${r("Tools","cyan",e)} ${r("(362 MCP tools)","muted",e)}`,o("search",["harvest_paa","search_serp","maps_search","maps_place_intel"],e),o("extract",["extract_url","map_site_urls","extract_site","audit_site","directory_workflow"],e),o("build",["create_editorial_reading_room","rank_tracker_workflow","portable HTML"],e),o("media",["youtube_harvest","youtube_transcribe","facebook_reels_inventory","facebook_ad_search","facebook_page_intel","facebook_ad_transcribe","facebook_video_transcribe","tiktok_video_transcribe","instagram_profile_content","instagram_media_download","reddit_thread"],e),o("browser",["serp_identity_create","serp_identity_list","browser_open","browser_profile_connect","browser_profile_list","browser_close","browser_screenshot","browser_read","browser_locate","browser_replay_mark","browser_replay_annotate"],e),o("connect",["list_service_connections","describe_service_connection_tool","import_service_connection_to_memory","export_connected_service_data","renew_connected_data_download","read_service_connection","call_service_connection_action"],e),o("commons",["commons_search_entities","commons_get_entity_linkset","commons_prepare_entity","commons_submit_entity","commons_prepare_publication","commons_claim_publication","commons_publish_editorial","commons_get_publication"],e),o("account",["credits_info","reports","MCP resources"],e),o("memory",["memory-put","memory-get","memory-search","list-vaults","record-fact","list-scheduled-actions"],e),`${r("Workflows","cyan",e)} ${r("(MCP + CLI + API)","muted",e)}`,o("route",["workflow_list","workflow_suggest","workflow_run","workflow_step","workflow_status","workflow_artifact_read"],e),o("seo",["directory","agent-packet","competitive audit","map/serp comparison","PAA/AIO briefs","scheduled runs"],e),"",r("Usage tips:","amber",e),"Run mcp-scraper-install for this visible card. Run mcp-scraper-cli for setup utilities and subcommands.","Run mcp-scraper in a human terminal to print this card; MCP clients get the same command as a silent stdio server.","Explicit card command: npx -y -p mcp-scraper@latest mcp-scraper-install","Hosted browser sessions use direct/no-proxy egress by default.","Customer auth setup: run browser_profile_connect, send the watch_url, let the user sign in, then call browser_profile_list until AUTHENTICATED.","Connected account ranges: call export_connected_service_data once. It handles Gmail, Calendar, Zoom, and Resend pagination; do not loop read_service_connection over individual records.","Connected account RAG: call import_service_connection_to_memory for one bounded approved read. It writes a redacted, untrusted snapshot to a stable Memory path and embeds it for search.","Stack logins / reconnect: run browser_profile_connect again with the same profile name and another domain to add accounts or refresh a login.","Start with workflow_suggest for broad jobs like market analysis, ICP research, CRO audits, brand briefs, content gaps, and AI visibility.","For MCP clients, use mcp-scraper so one install can mix SERP, Maps, browser, reports, and saved MCP resources.","If you hit the concurrency limit, add 2 browsers for $5/month with mcp-scraper-cli billing concurrency checkout.","",`${r("Ready.","lime",e)} Install the combined MCP server with one command:`,"",r("Setup doctor","amber",e),"npx -y -p mcp-scraper@latest mcp-scraper-cli doctor","",r("Hosted profile setup","amber",e),'In your MCP client, call browser_profile_connect with email="seo@example.com" and domain="chatgpt.com".',"Give the returned watch_url to the user. After they sign in, call browser_profile_list, then browser_open with the returned profile. Add more logins by calling browser_profile_connect again with the same profile and a new domain.","",r("Claude Code one-command setup","amber",e),i,"Then fully exit Claude Code and open a new Claude terminal. Check with: claude mcp list","",r("Codex config","amber",e),c,"",r("Claude Desktop Extension","amber",e),"Download: https://mcpscraper.dev/downloads/mcp-scraper.mcpb","",r("Safety note:","muted",e),"mcp-scraper prints this card only when stdin/stdout are an interactive TTY. In MCP clients it writes only JSON-RPC to stdout.","Use --stdio or MCP_SCRAPER_FORCE_STDIO=1 to force server mode from a terminal.",""].join(`
13
+ `,c=[`MCP_SCRAPER_API_KEY=${n} npx -y -p mcp-scraper@latest \\`," mcp-scraper-cli agent install claude --apply"].join(`
14
+ `),i=["[mcp_servers.mcp-scraper]",'command = "npx"','args = ["-y", "-p", "mcp-scraper@latest", "mcp-scraper"]',`env = { MCP_SCRAPER_API_KEY = "${n}" }`].join(`
15
+ `);return[r(`mcp-scraper v${t.version}`,"bold",e),r("> mcp-scraper-install","muted",e),r(a,"amber",e),`${r("MCP Scraper Agent","cyan",e)} . v${t.version} . mcpscraper.dev`,"1/1 install surfaces ready",r(`Newest in v${t.version}: ${s.message}`,"lime",e),"",`${r("Tools","cyan",e)} ${r("(362 MCP tools)","muted",e)}`,o("search",["harvest_paa","search_serp","maps_search","maps_place_intel"],e),o("extract",["extract_url","map_site_urls","extract_site","audit_site","directory_workflow"],e),o("build",["create_editorial_reading_room","rank_tracker_workflow","portable HTML"],e),o("media",["youtube_harvest","youtube_transcribe","facebook_reels_inventory","facebook_ad_search","facebook_page_intel","facebook_ad_transcribe","facebook_video_transcribe","tiktok_video_transcribe","instagram_profile_content","instagram_media_download","reddit_thread"],e),o("browser",["serp_identity_create","serp_identity_list","browser_open","browser_profile_connect","browser_profile_list","browser_close","browser_screenshot","browser_read","browser_locate","browser_replay_mark","browser_replay_annotate"],e),o("connect",["list_service_connections","describe_service_connection_tool","import_service_connection_to_memory","export_connected_service_data","renew_connected_data_download","read_service_connection","call_service_connection_action"],e),o("commons",["commons_search_entities","commons_get_entity_linkset","commons_prepare_entity","commons_submit_entity","commons_prepare_publication","commons_claim_publication","commons_publish_editorial","commons_get_publication"],e),o("account",["credits_info","reports","MCP resources"],e),o("memory",["memory-put","memory-get","memory-search","list-vaults","record-fact","list-scheduled-actions"],e),`${r("Workflows","cyan",e)} ${r("(MCP + CLI + API)","muted",e)}`,o("route",["workflow_list","workflow_suggest","workflow_run","workflow_step","workflow_status","workflow_artifact_read"],e),o("seo",["directory","agent-packet","competitive audit","map/serp comparison","PAA/AIO briefs","scheduled runs"],e),"",r("Usage tips:","amber",e),"Run mcp-scraper-install for this visible card. Run mcp-scraper-cli for setup utilities and subcommands.","Run mcp-scraper in a human terminal to print this card; MCP clients get the same command as a silent stdio server.","Explicit card command: npx -y -p mcp-scraper@latest mcp-scraper-install","Hosted browser sessions use direct/no-proxy egress by default.","Customer auth setup: run browser_profile_connect, send the watch_url, let the user sign in, then call browser_profile_list until AUTHENTICATED.","Connected account ranges: call export_connected_service_data once. It handles Gmail, Calendar, Zoom, and Resend pagination; do not loop read_service_connection over individual records.","Connected account RAG: call import_service_connection_to_memory for one bounded approved read. It writes a redacted, untrusted snapshot to a stable Memory path and embeds it for search.","Stack logins / reconnect: run browser_profile_connect again with the same profile name and another domain to add accounts or refresh a login.","Start with workflow_suggest for broad jobs like market analysis, ICP research, CRO audits, brand briefs, content gaps, and AI visibility.","For MCP clients, use mcp-scraper so one install can mix SERP, Maps, browser, reports, and saved MCP resources.","If you hit the concurrency limit, add 2 browsers for $5/month with mcp-scraper-cli billing concurrency checkout.","",`${r("Ready.","lime",e)} Install the combined MCP server with one command:`,"",r("Setup doctor","amber",e),"npx -y -p mcp-scraper@latest mcp-scraper-cli doctor","",r("Hosted profile setup","amber",e),'In your MCP client, call browser_profile_connect with email="seo@example.com" and domain="chatgpt.com".',"Give the returned watch_url to the user. After they sign in, call browser_profile_list, then browser_open with the returned profile. Add more logins by calling browser_profile_connect again with the same profile and a new domain.","",r("Claude Code one-command setup","amber",e),c,"Then fully exit Claude Code and open a new Claude terminal. Check with: claude mcp list","",r("Codex config","amber",e),i,"",r("Claude Desktop Extension","amber",e),"Download: https://mcpscraper.dev/downloads/mcp-scraper.mcpb","",r("Safety note:","muted",e),"mcp-scraper prints this card only when stdin/stdout are an interactive TTY. In MCP clients it writes only JSON-RPC to stdout.","Use --stdio or MCP_SCRAPER_FORCE_STDIO=1 to force server mode from a terminal.",""].join(`
16
16
  `)}export{p as a};