scavio 0.14.0 → 0.15.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,6 +1,20 @@
1
1
  # Scavio
2
2
 
3
- TypeScript SDK for the [Scavio Search API](https://scavio.dev) — real-time Google, Amazon, Walmart, YouTube, Reddit, TikTok, TikTok Shop, Instagram, X, and LinkedIn data.
3
+ TypeScript SDK for the [Scavio API](https://scavio.dev) — real-time web scraping
4
+ and data extraction across 31 platforms on one API key, plus `extract()` to read
5
+ any URL as clean Markdown.
6
+
7
+ Structured JSON in, structured JSON out. No proxies, no headless browsers, no
8
+ per-site parsers to maintain.
9
+
10
+ - **Search and SERP** — Google (organic, news, maps, shopping, flights, hotels, trends, AI Mode)
11
+ - **E-commerce** — Amazon, Walmart, eBay, Target, Home Depot, TikTok Shop
12
+ - **Real estate and travel** — Zillow, Redfin, Booking, Airbnb, Tripadvisor
13
+ - **Reviews and local** — Yelp, G2, Capterra, Glassdoor, App Store, Google Play
14
+ - **Jobs and companies** — Indeed, Glassdoor, SEC EDGAR, Companies House
15
+ - **Ads transparency** — Google Ads Transparency Center, Meta Ad Library
16
+ - **Social and video** — YouTube, TikTok, Instagram, Threads, X, LinkedIn, Reddit, Kuaishou
17
+ - **Any other page** — `extract()` turns a URL into Markdown, plain text or raw HTML
4
18
 
5
19
  ## Install
6
20
 
@@ -21,6 +35,9 @@ const results = await client.search({ query: "web scraping api" });
21
35
  // Amazon product lookup
22
36
  const product = await client.amazon.product({ asin: "B09V3KXJPB" });
23
37
 
38
+ // Read any page as Markdown
39
+ const page = await client.extract({ url: "https://example.com/pricing" });
40
+
24
41
  // Check usage
25
42
  const usage = await client.getUsage();
26
43
  ```
@@ -41,7 +58,7 @@ const client = new Scavio({
41
58
  only to transient failures — HTTP 429, 500, 502, 503, 504 and network or timeout
42
59
  errors. Backoff is exponential with full jitter, capped at 8s, and a
43
60
  `Retry-After` header is honored when the API sends one. Non-transient errors
44
- (400, 401, 402, 404) are never retried.
61
+ (400, 401, 402, 404, 422) are never retried.
45
62
 
46
63
  `maxRequestsPerSecond` throttles the client so it never sends more than N
47
64
  requests in any one-second window. Your plan also has a server-side concurrency
@@ -50,6 +67,38 @@ Project, 3 on Bootstrap, 5 on Startup, 10 on Growth.
50
67
 
51
68
  ## API Reference
52
69
 
70
+ ### Extract (any URL)
71
+
72
+ `extract()` is a core endpoint, not a platform, so it lives on the client itself.
73
+ It reads any page and hands it back as readability Markdown, plain text or raw
74
+ HTML — the read-a-page primitive an agent or a RAG ingest needs.
75
+
76
+ ```typescript
77
+ const page = await client.extract({
78
+ url: "https://example.com/pricing",
79
+ format: "markdown", // "html" | "markdown" | "text", default "markdown"
80
+ mode: "normal", // "normal" | "advanced" | "ultra", default "normal"
81
+ });
82
+
83
+ page.content; // the page body
84
+ page.content_length; // characters returned
85
+ ```
86
+
87
+ Credits depend on `mode`, not on a flat per-call price: `normal` costs 1,
88
+ `advanced` costs 1, `ultra` costs 2. Billing happens only on a successful
89
+ extraction — a dead link, bot wall or timeout costs nothing.
90
+
91
+ - `normal` is a plain datacenter fetch. Start here.
92
+ - `advanced` renders the page in a headless browser. Use it when the content is
93
+ built client-side.
94
+ - `ultra` goes through residential IPs. Use it only when a bot wall blocks the
95
+ other two.
96
+
97
+ `html` is the raw page. `markdown` is a readability extraction with the
98
+ boilerplate stripped. `text` is that markdown flattened to plain text. URLs are
99
+ http(s) only, a bare host is upgraded to https, and loopback, private,
100
+ link-local and cloud-metadata hosts are rejected with a 400.
101
+
53
102
  ### Google
54
103
 
55
104
  Each method hits its own `/api/v2/google*` endpoint and returns Google's full
@@ -128,20 +177,47 @@ instead of the previous raw provider payload.
128
177
 
129
178
  ### Walmart
130
179
 
180
+ Seven endpoints. `search` and `product` changed shape in 0.15.0 and the other
181
+ five are new.
182
+
131
183
  ```typescript
132
- // Search products
133
184
  await client.walmart.search({
134
185
  query: "tv",
135
- min_price: 100, // optional
136
- max_price: 500, // optional
186
+ domain: "com", // "com" | "ca" | "com.mx"
187
+ page: 2, // 1-indexed
188
+ sort_by: "rating_high",
189
+ min_price: 100,
190
+ max_price: 500,
137
191
  });
138
192
 
139
- // Get product by ID
140
- await client.walmart.product({
141
- product_id: "123456",
142
- });
193
+ await client.walmart.product({ product_id: "13544111159" });
194
+ await client.walmart.reviews({ product_id: "13544111159", page: 2 });
195
+ await client.walmart.category({ category_id: "3944_133251_1095191" });
196
+ await client.walmart.offers({ product_id: "13544111159" });
197
+ await client.walmart.seller({ seller_id: "101138578" });
198
+ await client.walmart.sellerProducts({ seller_id: "101138578" });
143
199
  ```
144
200
 
201
+ Credits are a function of the body on `search` and `category`: `domain` "com" and
202
+ "ca" cost 1 credit, "com.mx" costs 2. The other five endpoints always cost 1.
203
+
204
+ #### Walmart changed in 0.15.0 (breaking)
205
+
206
+ - `device`, `delivery_zip` and `store_id` are retired. Sending one still returns
207
+ 200, with a top-level `warnings` array naming what was ignored.
208
+ - `domain` is **not** retired — it is the price-bearing param, and it is accepted
209
+ on `search` and `category` only. Walmart.ca product pages cannot be fetched, so
210
+ every product-keyed endpoint is walmart.com only.
211
+ - `page` (1-indexed) is the paging param; `start_page` remains a deprecated alias.
212
+ - `sort_by` gained `rating_high` and `new`.
213
+ - `fulfillment_speed` is `today` or `tomorrow` only. There is deliberately no
214
+ `2_days` (it leaks 3-4 day items) and no `anytime` — omit the param instead.
215
+ - `offers` returns the **buy-box seller only**, not the full offer list.
216
+ - `seller_id` must be the numeric catalog id (`seller_catalog_id`, returned by
217
+ `product` and `offers`). The GUID form of the id returns 404.
218
+ - `sellerProducts` has **no pagination** — roughly the first 40 server-rendered
219
+ items. `total_count` reports the seller's real catalog size.
220
+
145
221
  ### YouTube
146
222
 
147
223
  Credit cost varies by endpoint: `transcript` costs 8; `streams` costs 3;
@@ -431,6 +507,389 @@ await client.instagram.userFollowers({ username: "instagram", count: 50 });
431
507
  await client.instagram.userFollowings({ username: "instagram" });
432
508
  ```
433
509
 
510
+ ### Threads
511
+
512
+ Six endpoints, and the credit cost is a function of the body: **2 credits when
513
+ you address a user by `user_id`, 4 when you address them by `username`.** The
514
+ upstream handle lookup is dead, so a handle buys a second call. Only `profile`,
515
+ `userPosts` and `userReplies` are username-keyed; `post`, `postComments` and
516
+ `searchUsers` always cost 2.
517
+
518
+ ```typescript
519
+ // Resolve the handle once, then stay on the cheap path
520
+ const found = await client.threads.searchUsers({ query: "zuck" });
521
+
522
+ await client.threads.profile({ user_id: "63625256886" });
523
+ await client.threads.userPosts({ user_id: "63625256886", cursor });
524
+ await client.threads.userReplies({ user_id: "63625256886" });
525
+ await client.threads.post({ url: "https://www.threads.net/@zuck/post/..." });
526
+ await client.threads.postComments({ post_id: "3141..." });
527
+ ```
528
+
529
+ There is no Threads content search — `searchUsers` is people search and it is
530
+ the only search Threads exposes. Missing or conflicting identifiers return 422,
531
+ not 400; no match returns 404.
532
+
533
+ ### Kuaishou
534
+
535
+ Fourteen endpoints, priced **per endpoint** rather than flat: `videosBatch`
536
+ costs 40, `profile` and the four `search*` methods cost 10, `video` costs 2, and
537
+ everything else costs 1.
538
+
539
+ ```typescript
540
+ await client.kuaishou.userResolve({ share_link: "https://v.kuaishou.com/..." }); // 1
541
+ await client.kuaishou.userPosts({ user_id: "3xabc..." }); // 1
542
+ await client.kuaishou.tagFeed({ tag: "美食" }); // 1
543
+ await client.kuaishou.video({ photo_id: "3xdef..." }); // 2
544
+ await client.kuaishou.searchVideos({ keyword: "coffee" }); // 10
545
+ await client.kuaishou.videosBatch({ photo_ids: ["3xdef...", "3xghi..."] }); // 40
546
+ ```
547
+
548
+ `videosBatch` costs 40 whether you send 1 id or 20, so fill the batch. If all
549
+ you have is a share link, `userResolve()` turns it into a user id for 1 credit
550
+ rather than paying 10 for `profile()`. Missing identifiers return 422.
551
+
552
+ ### eBay
553
+
554
+ ```typescript
555
+ await client.ebay.search({ query: "airpods pro", sold: true, per_page: 120 });
556
+ await client.ebay.search({ seller: "musicmagpie" }); // no keyword needed
557
+ await client.ebay.product({ item_id: "126543210987" });
558
+ await client.ebay.seller({ seller: "musicmagpie" });
559
+ ```
560
+
561
+ 1 credit per call. `sold: true` searches completed listings that actually sold —
562
+ the price-research view; eBay publishes no headline count there, so
563
+ `total_results` comes back null. `per_page` accepts only 60, 120 or 240. `seller`
564
+ is a profile endpoint and cannot enumerate a catalogue — page a seller's
565
+ inventory through `search({ seller })` instead.
566
+
567
+ ### Target
568
+
569
+ ```typescript
570
+ await client.target.search({ keyword: "office chair", store_id: "1234" });
571
+ await client.target.category({ category_id: "5xtg6" });
572
+ await client.target.product({ tcin: "82291396" });
573
+ await client.target.reviews({ tcin: "82291396" });
574
+ ```
575
+
576
+ 1 credit per call, but these are the slowest endpoints in the SDK: product about
577
+ 4s, search about 9s, category about 37s, reviews about 40s. Raise `timeout`
578
+ accordingly. `reviews` returns at most 8 review bodies whatever `review_count`
579
+ says, and `limit` only trims — there is no paging. `seller_*` is null on
580
+ first-party stock, which means "sold by Target"; only Target Plus marketplace
581
+ listings name a vendor.
582
+
583
+ ### Home Depot
584
+
585
+ ```typescript
586
+ await client.homeDepot.search({ query: "cordless drill", page: 2 });
587
+ await client.homeDepot.product({ item_id: "313159056" });
588
+ await client.homeDepot.reviews({ item_id: "313159056", page: 2 });
589
+ ```
590
+
591
+ 2 credits per call. Search page size is fixed at 12, so paging is the only way
592
+ to read further; reviews are 30 per page and a page past `total_pages` is a 404.
593
+ `product` carries only a 10-review preview — `reviews` is the paginated surface.
594
+
595
+ ### Zillow
596
+
597
+ ```typescript
598
+ await client.zillow.search({
599
+ location: "Austin, TX",
600
+ listing_status: "for_rent",
601
+ min_price: 1500, // MONTHLY RENT on for_rent
602
+ max_price: 3000,
603
+ });
604
+ await client.zillow.property({ zpid: "29433327" });
605
+ await client.zillow.agentReviews({ screen_name: "jane-doe" });
606
+ ```
607
+
608
+ 1 credit per call. A bare ZIP works on its own but cannot be combined with a
609
+ filter or a sort — Zillow then geolocates the request and answers about another
610
+ city, so pass the city name whenever you filter. On `listing_status: "for_rent"`,
611
+ `min_price`/`max_price` mean monthly rent. `agentReviews` addresses an agent
612
+ profile, not a property, and returns the five reviews Zillow server-renders
613
+ (`total_review_count` is the real total). An unresolvable region is a 404.
614
+
615
+ ### Redfin
616
+
617
+ ```typescript
618
+ await client.redfin.search({ location: "https://www.redfin.com/city/30749/TX/Austin" });
619
+ await client.redfin.search({ region_id: 30749, region_type: 6 }); // 6 = city
620
+ await client.redfin.property({ property_id: "185301234" });
621
+ await client.redfin.market({ region_id: 30749, region_type: 6 });
622
+ ```
623
+
624
+ 1 credit per call. **City names are not accepted** on `location` — pass a
625
+ redfin.com region URL (`/city/`, `/neighborhood/`, `/county/`, `/zipcode/`) or
626
+ `region_id` plus `region_type`, which must be sent together. `region_id` is not
627
+ a ZIP code. `sold_within_days` is only valid with `listing_status: "sold"`.
628
+ `days_on_market` is always null in the response — do not build on it.
629
+
630
+ ### Booking
631
+
632
+ ```typescript
633
+ await client.booking.search({
634
+ destination: "Lisbon",
635
+ checkin: "2026-09-10",
636
+ checkout: "2026-09-13", // send both or neither
637
+ adults: 2,
638
+ currency: "USD",
639
+ });
640
+ await client.booking.hotel({ hotel: "the-independente" });
641
+ await client.booking.reviews({ hotel: "the-independente" });
642
+ ```
643
+
644
+ 1 credit per call. `checkin` and `checkout` must be sent together — Booking
645
+ ignores a lone check-in and prices its own date range. `hotel` and `reviews`
646
+ take dates for the same reason: Booking prices a stay, and the response echoes
647
+ whichever dates were used. `currency` defaults to USD; without it Booking prices
648
+ off the proxy exit and two identical requests disagree. A search with neither
649
+ `destination` nor `dest_id` returns Booking's homepage and still costs a credit.
650
+
651
+ ### Airbnb
652
+
653
+ ```typescript
654
+ await client.airbnb.search({
655
+ location: "Barcelona",
656
+ check_in: "2026-09-10",
657
+ check_out: "2026-09-15", // send both or neither
658
+ currency: "USD",
659
+ });
660
+ await client.airbnb.listing({ listing_id: "1234567890" });
661
+ await client.airbnb.reviews({ listing_id: "1234567890", limit: 50, offset: 50 });
662
+ ```
663
+
664
+ 1 credit per call. **Prices are search-only** — the listing page carries no
665
+ nightly rate under any parameters. A dateless search defaults to +30d for 5
666
+ nights and Airbnb A/Bs both the window and the prices, so the response flags it
667
+ as `dates_are_defaulted`; send real dates for anything you intend to compare.
668
+ The rating breakdown and review tags live on `listing`, while `reviews` returns
669
+ the review bodies with `limit`/`offset` paging.
670
+
671
+ ### Tripadvisor
672
+
673
+ **Start with `locations()`.** Every other endpoint is keyed by ids that exist
674
+ only inside TripAdvisor's own URLs.
675
+
676
+ ```typescript
677
+ const places = await client.tripadvisor.locations({ query: "Le Bernardin" });
678
+ await client.tripadvisor.search({ geo_id: "60763", category: "restaurants" });
679
+ await client.tripadvisor.location({ location_id: "426986", geo_id: "60763" });
680
+ await client.tripadvisor.reviews({ location_id: "426986", geo_id: "60763", page: 2 });
681
+ ```
682
+
683
+ 2 credits per call. A geo row from `locations()` gives the `geo_id` that
684
+ `search()` takes; a business row gives the `geo_id` + `location_id` pair that
685
+ `location()` and `reviews()` take. Page 1 of the reviews already ships inside
686
+ `location()`, so use `reviews()` to page past it. Review page size differs by
687
+ family (15 restaurants, 10 hotels and attractions), consecutive pages can repeat
688
+ one review at the boundary (de-duplicate on `review_id`), and a page past the
689
+ last is a 404.
690
+
691
+ ### Yelp
692
+
693
+ ```typescript
694
+ await client.yelp.search({ term: "ramen", location: "Seattle, WA" });
695
+ await client.yelp.business({ business_id: "..." });
696
+ await client.yelp.reviews({ business_id: "...", page: 2 }); // page 2, not 1
697
+ ```
698
+
699
+ 2 credits per call. `location` is effectively required — without it Yelp
700
+ geolocates off the proxy exit and the same request answers about a different
701
+ metro run to run. **Reviews page 1 is redundant**: it re-fetches the document
702
+ `business()` already returned and costs another 2 credits, so start at page 2.
703
+ Page size is fixed at 10 and a page past the last is a 404.
704
+
705
+ ### Indeed
706
+
707
+ ```typescript
708
+ await client.indeed.search({
709
+ query: "data engineer",
710
+ location: "Remote",
711
+ radius: 25, // 0, 5, 10, 15, 25, 35, 50 or 100 only
712
+ max_age_days: 7, // 1, 3, 7 or 14 only
713
+ });
714
+ await client.indeed.job({ job_id: "a1b2c3d4e5f6" });
715
+ await client.indeed.company({ company: "Stripe" });
716
+ await client.indeed.companyReviews({ company: "Stripe", page: 2 });
717
+ ```
718
+
719
+ 2 credits per call. `radius` and `max_age_days` are closed sets — Indeed ignores
720
+ anything else and silently returns the unfiltered set, so an unsupported radius
721
+ bills you for a search covering fifty miles. `min_salary` filters on Indeed's own
722
+ estimate for the role, not on a posted figure, so postings that publish no salary
723
+ still match. A location-only search (no `query`) is valid. Search is 10 postings
724
+ per page, company reviews 20.
725
+
726
+ ### Glassdoor
727
+
728
+ **Start with `companies()`.** The other three are keyed by an `employer_id` that
729
+ only exists inside Glassdoor's `/Overview/` URLs.
730
+
731
+ ```typescript
732
+ const hits = await client.glassdoor.companies({ query: "Stripe" });
733
+ const company = await client.glassdoor.company({ employer_id: "671932" });
734
+
735
+ // Pass the URLs the company response returns - halves the upstream work
736
+ await client.glassdoor.reviews({ url: company.reviews_url as string });
737
+ await client.glassdoor.salaries({ url: company.salaries_url as string });
738
+ ```
739
+
740
+ 1 credit per call. Glassdoor's login wall caps `reviews` at **three reviews per
741
+ response** — there is deliberately no `page` param; move the window with
742
+ `category` and `employment_status` and read `filtered_review_count` to see how
743
+ many match. Addressing `reviews` or `salaries` by `employer_id` costs two
744
+ upstream fetches because the slugs are case-sensitive and must be read off the
745
+ profile, so prefer the `reviews_url` / `salaries_url` the company response hands
746
+ back. These endpoints are slow and flaky by nature (company about 3-47s, reviews
747
+ about 75s, salaries about 41s) — raise `timeout` and keep retries on.
748
+
749
+ ### App Store
750
+
751
+ ```typescript
752
+ await client.appStore.search({ term: "meditation", limit: 100, country: "us" });
753
+ await client.appStore.app({ app_id: "com.apple.Pages" }); // or the numeric id
754
+ await client.appStore.reviews({ app_id: "361309726", page: 2, sort: "most_helpful" });
755
+ ```
756
+
757
+ 1 credit per call. **Search has no pagination** — `limit` (1-200) is the only
758
+ lever; every offset spelling is silently ignored. `app` accepts both a numeric
759
+ App Store id and a bundle id; `reviews` is numeric-only. Reviews hard-stop at
760
+ page 10 (50 per page), which is Apple's anonymous ceiling — reach further by
761
+ asking a different `country`. Reviews cannot 404: an unknown id and an app with
762
+ zero reviews return the same empty feed.
763
+
764
+ ### Google Play
765
+
766
+ ```typescript
767
+ await client.googlePlay.search({ query: "meditation", hl: "en", gl: "us" });
768
+ await client.googlePlay.app({ app_id: "com.spotify.music" });
769
+ await client.googlePlay.reviews({ app_id: "com.spotify.music", sort: "newest", count: 200 });
770
+ ```
771
+
772
+ 2 credits per call. **Search does not paginate** — one shelf of about 30 apps.
773
+ `hl` changes the storefront, not just the strings: title, description, install
774
+ formatting and content rating all move with it. The reviews `cursor` is opaque,
775
+ single-use, and encodes the sort as well as the position, so send it back with
776
+ the same `sort` it came from; a cursor past the last review is a 404. `app`
777
+ already returns the 20 reviews Play server-renders, plus the real install count
778
+ Play publishes but never displays.
779
+
780
+ ### G2
781
+
782
+ ```typescript
783
+ await client.g2.search({ query: "crm", rating: 4 });
784
+ await client.g2.product({ product_id: "notion" });
785
+ await client.g2.reviews({ product_id: "notion", page: 2, company_size: "enterprise" });
786
+ ```
787
+
788
+ **5 credits per call — the most expensive platform in the SDK**, because g2.com
789
+ bills 25 upstream credits per fetch. Retries are deliberately conservative for
790
+ that reason, and a bot wall arrives as a billed 200 rather than an error. G2
791
+ loads review text in a separate frame, so `product` carries no reviews — call
792
+ `reviews`, which is also the only place with exact per-star counts, pros/cons by
793
+ theme, and company-size / role / industry / region facets.
794
+
795
+ ### Capterra
796
+
797
+ ```typescript
798
+ const hits = await client.capterra.search({ query: "project management" });
799
+ await client.capterra.product({ product_id: "186596", slug: "Notion" });
800
+ await client.capterra.reviews({ product_id: "186596", slug: "Notion", page: 2 });
801
+ ```
802
+
803
+ 2 credits per call. **Search does not paginate** — Capterra fixes the result set
804
+ at 20, so there is no `page` param. `slug` is cosmetic on `product` but
805
+ load-bearing and case-sensitive on `reviews`: a wrong one silently serves page 1
806
+ under a billed 200, so pass back the `slug` or `reviews_url` that `search` or
807
+ `product` returned. Reviews are 25 per page, capped at page 100, and page 1
808
+ already ships inside `product`. `vendor` is null on the product profile —
809
+ Capterra does not publish it there.
810
+
811
+ ### SEC EDGAR
812
+
813
+ **Start with `lookup()`.** Callers hold a ticker; EDGAR is keyed by CIK.
814
+
815
+ ```typescript
816
+ const match = await client.sec.lookup({ query: "AAPL" });
817
+ await client.sec.company({ ticker: "AAPL" });
818
+ await client.sec.filings({ ticker: "AAPL", form: "10-K", include_history: true });
819
+ await client.sec.facts({ ticker: "AAPL", query: "revenue" });
820
+ await client.sec.concept({ ticker: "AAPL", concept: "NetIncomeLoss" });
821
+ await client.sec.search({ query: "climate risk", form: "10-K" });
822
+ ```
823
+
824
+ 1 credit per call, including `include_history: true`, which is the one call that
825
+ can buy up to 10 upstream fetches. Both `cik` and `ticker` accept either
826
+ spelling. XBRL concept tags are **case-sensitive** — `netincomeloss` is a 404
827
+ upstream, so use `facts()` to see what a filer actually reports. EDGAR's
828
+ "recent" block is not a fixed window: a decade for a quiet filer, about a year
829
+ for a prolific one.
830
+
831
+ ### Companies House
832
+
833
+ ```typescript
834
+ const hits = await client.companiesHouse.search({ query: "Monzo" });
835
+ await client.companiesHouse.company({ company_number: "09446231" });
836
+ await client.companiesHouse.officers({ company_number: "09446231", page: 2 });
837
+ await client.companiesHouse.filingHistory({ company_number: "SC090312" });
838
+ ```
839
+
840
+ 1 credit per call. `company_number` is deliberately loose — the register 404s on
841
+ numbers that lost their leading zeros or arrived lower-cased, so the transport
842
+ pads and upper-cases before asking. SC, NI, OC, SO, NC, FC, BR and CE prefixes
843
+ are all supported. Search is 20 per page and capped at page 50: the register
844
+ serves a 1000-result window per term whatever hit count it advertises, and page
845
+ 51 is an HTTP 416. Officers and filing history have no upper page bound — past
846
+ the last page you get an ordinary empty list.
847
+
848
+ ### Google Ads Transparency
849
+
850
+ ```typescript
851
+ const advertisers = await client.googleAds.advertisers({ query: "nike" });
852
+ const page1 = await client.googleAds.search({ advertiser_id: "AR123...", region: "DE" });
853
+ const page2 = await client.googleAds.search({
854
+ advertiser_id: "AR123...",
855
+ region: "DE",
856
+ cursor: page1.next_cursor as string, // re-send the SAME filters
857
+ });
858
+ await client.googleAds.creative({ advertiser_id: "AR123...", creative_id: "CR456..." });
859
+ ```
860
+
861
+ 1 credit per call. `search` paginates by `cursor` / `next_cursor` at 100 per
862
+ page — re-send the same filters alongside the cursor. `limit` is capped at 100 by
863
+ Google itself, which answers a larger request with zero rows rather than an
864
+ error. `advertisers` and `creative` do not paginate. **Impressions and reach are
865
+ EEA-only**: US creatives return null for `impressions_min`, `impressions_max` and
866
+ `first_shown` because Google only publishes reach where the DSA compels it. The
867
+ text, image and video format sets are disjoint — an advertiser's creatives never
868
+ overlap between them.
869
+
870
+ ### Meta Ad Library
871
+
872
+ ```typescript
873
+ const page1 = await client.metaAds.search({ query: "protein powder", country: "US" });
874
+ const page2 = await client.metaAds.search({
875
+ query: "protein powder",
876
+ cursor: page1.next_cursor as string,
877
+ });
878
+ await client.metaAds.advertiser({ page_id: "10150125871..." });
879
+ await client.metaAds.ad({ ad_archive_id: "1234567890123456" });
880
+ ```
881
+
882
+ 1 credit per call. `search` and `advertiser` paginate all the way through: page 1
883
+ returns 30 ads, then 10 per page via `next_cursor` — walk `has_next_page` to pull
884
+ a whole query or advertiser. The cursor is a self-contained opaque blob, so
885
+ paging is stateless and **the other filters are ignored when a cursor is
886
+ present**, so re-sending `query` alongside it is harmless and satisfies the
887
+ schema — the cursor already carries the filters. `total_results` caps at 50000 with
888
+ `total_is_capped: true`, because Meta only reports "more than 50,000". Spend,
889
+ reach, impressions and the paid-for-by disclosure exist on political and issue
890
+ ads only — set `ad_type: "political_and_issue_ads"` to expose them; commercial
891
+ ads leave those fields null.
892
+
434
893
  ### Usage
435
894
 
436
895
  ```typescript
@@ -460,13 +919,18 @@ All error classes:
460
919
  | `MissingAPIKeyError` | — | No API key provided |
461
920
  | `ScavioConnectionError` | — | Request never reached the API (DNS, reset, TLS) |
462
921
  | `ScavioTimeoutError` | — | Request exceeded the configured `timeout` |
463
- | `BadRequestError` | 400 | Invalid request parameters |
922
+ | `BadRequestError` | 400, 422 | Invalid request parameters |
464
923
  | `InvalidAPIKeyError` | 401 | Invalid API key |
465
924
  | `InsufficientCreditsError` | 402 | No credits remaining |
466
925
  | `NotFoundError` | 404 | No data upstream for that id (see TikTok Shop above) |
467
926
  | `RateLimitError` | 429 | Rate limit exceeded |
468
927
  | `ScavioAPIError` | other | Catch-all (has `.statusCode`) |
469
928
 
929
+ Threads and Kuaishou answer a missing or conflicting identifier with **422**, not
930
+ 400 — those routes have no 400 at all. Both map to `BadRequestError`, so one
931
+ `catch (e) { if (e instanceof BadRequestError) }` covers validation failures on
932
+ every platform. `e.statusCode` still reports whichever status the API sent.
933
+
470
934
  Every class extends `ScavioError`, so `catch (e) { if (e instanceof ScavioError) }`
471
935
  matches all of them. All except `MissingAPIKeyError`, `ScavioConnectionError` and
472
936
  `ScavioTimeoutError` carry `.statusCode` and `.responseBody`.
@@ -486,14 +950,24 @@ MIT
486
950
 
487
951
  ## About Scavio
488
952
 
489
- [Scavio](https://scavio.dev) is a unified [search API for AI agents](https://scavio.dev/search-api-for-ai-agents) — one API key, structured JSON, no scraping or proxies. A real-time [Tavily alternative](https://scavio.dev/alternatives/tavily) and [SerpAPI alternative](https://scavio.dev/alternatives/serpapi) with data from:
953
+ [Scavio](https://scavio.dev) is a unified web data and
954
+ [search API for AI agents](https://scavio.dev/search-api-for-ai-agents) — one API
955
+ key, structured JSON, no proxies or browser farms to run. It is a real-time
956
+ [Tavily alternative](https://scavio.dev/alternatives/tavily) and
957
+ [SerpAPI alternative](https://scavio.dev/alternatives/serpapi), and with
958
+ `extract()` it also covers the read-any-URL job people reach for Firecrawl to do.
959
+
960
+ What teams build on it:
490
961
 
491
- - [Google Search API](https://scavio.dev/google-search-api) — SERP results, news, images, maps, and knowledge graph
492
- - [Amazon Product API](https://scavio.dev/amazon-product-api) and [Walmart Product API](https://scavio.dev/walmart-product-api) — product search and details
493
- - [YouTube API](https://scavio.dev/youtube-transcript-api), [TikTok API](https://scavio.dev/tiktok-api), and [Instagram API](https://scavio.dev/instagram-api) — video and social media data
494
- - [Reddit API](https://scavio.dev/reddit-api) — posts and threaded comments
495
- - [X API](https://scavio.dev/docs/x-search) and [LinkedIn API](https://scavio.dev/docs/linkedin-person) — tweets, profiles, companies, and jobs
962
+ - **SERP and answer engines** — [Google Search API](https://scavio.dev/google-search-api) for organic results, news, images, maps and the knowledge graph
963
+ - **Price and catalog monitoring** — [Amazon Product API](https://scavio.dev/amazon-product-api), [Walmart Product API](https://scavio.dev/walmart-product-api), eBay, Target and Home Depot product, review and seller data
964
+ - **Real estate and travel pipelines** — Zillow and Redfin listings and market stats, Booking, Airbnb and Tripadvisor rates and reviews
965
+ - **Review mining and competitive research** — Yelp, G2, Capterra, Glassdoor, App Store and Google Play reviews on one shape
966
+ - **Recruiting and company intelligence** — Indeed jobs, Glassdoor salaries, SEC EDGAR filings and XBRL facts, UK Companies House officers and filing history
967
+ - **Ad and creative intelligence** — Google Ads Transparency Center and Meta Ad Library creatives
968
+ - **Social listening** — [YouTube API](https://scavio.dev/youtube-transcript-api), [TikTok API](https://scavio.dev/tiktok-api), [Instagram API](https://scavio.dev/instagram-api), [Reddit API](https://scavio.dev/reddit-api), [X API](https://scavio.dev/docs/x-search), [LinkedIn API](https://scavio.dev/docs/linkedin-person), Threads and Kuaishou
969
+ - **RAG ingestion** — `extract()` reads any URL and hands back readability Markdown ready to chunk and embed
496
970
 
497
971
  Teams choosing between providers can [compare Scavio vs alternatives](https://scavio.dev/compare) side by side.
498
972
 
499
- Get a free [API key](https://dashboard.scavio.dev) and explore the [documentation](https://scavio.dev/docs/introduction).
973
+ Get a free [API key](https://dashboard.scavio.dev/sign-up) and explore the [documentation](https://scavio.dev/docs/introduction).