scavio 0.14.0 → 0.16.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +495 -20
- package/dist/index.cjs +1624 -69
- package/dist/index.cjs.map +1 -1
- package/dist/index.d.cts +2923 -79
- package/dist/index.d.ts +2923 -79
- package/dist/index.js +1590 -68
- package/dist/index.js.map +1 -1
- package/package.json +36 -17
package/README.md
CHANGED
|
@@ -1,6 +1,20 @@
|
|
|
1
1
|
# Scavio
|
|
2
2
|
|
|
3
|
-
TypeScript SDK for the [Scavio
|
|
3
|
+
TypeScript SDK for the [Scavio API](https://scavio.dev) — real-time web scraping
|
|
4
|
+
and data extraction across 31 platforms on one API key, plus `extract()` to read
|
|
5
|
+
any URL as clean Markdown.
|
|
6
|
+
|
|
7
|
+
Structured JSON in, structured JSON out. No proxies, no headless browsers, no
|
|
8
|
+
per-site parsers to maintain.
|
|
9
|
+
|
|
10
|
+
- **Search and SERP** — Google (organic, news, maps, shopping, flights, hotels, trends, AI Mode)
|
|
11
|
+
- **E-commerce** — Amazon, Walmart, eBay, Target, Home Depot, TikTok Shop
|
|
12
|
+
- **Real estate and travel** — Zillow, Redfin, Booking, Airbnb, Tripadvisor
|
|
13
|
+
- **Reviews and local** — Yelp, G2, Capterra, Glassdoor, App Store, Google Play
|
|
14
|
+
- **Jobs and companies** — Indeed, Glassdoor, SEC EDGAR, Companies House
|
|
15
|
+
- **Ads transparency** — Google Ads Transparency Center, Meta Ad Library
|
|
16
|
+
- **Social and video** — YouTube, TikTok, Instagram, Threads, X, LinkedIn, Reddit, Kuaishou
|
|
17
|
+
- **Any other page** — `extract()` turns a URL into Markdown, plain text or raw HTML
|
|
4
18
|
|
|
5
19
|
## Install
|
|
6
20
|
|
|
@@ -21,6 +35,9 @@ const results = await client.search({ query: "web scraping api" });
|
|
|
21
35
|
// Amazon product lookup
|
|
22
36
|
const product = await client.amazon.product({ asin: "B09V3KXJPB" });
|
|
23
37
|
|
|
38
|
+
// Read any page as Markdown
|
|
39
|
+
const page = await client.extract({ url: "https://example.com/pricing" });
|
|
40
|
+
|
|
24
41
|
// Check usage
|
|
25
42
|
const usage = await client.getUsage();
|
|
26
43
|
```
|
|
@@ -32,7 +49,7 @@ const client = new Scavio({
|
|
|
32
49
|
apiKey: "sk_...", // or set SCAVIO_API_KEY env var
|
|
33
50
|
baseUrl: "https://api.scavio.dev", // default
|
|
34
51
|
timeout: 30_000, // ms, default
|
|
35
|
-
maxRequestsPerSecond: 1, // 1-
|
|
52
|
+
maxRequestsPerSecond: 1, // 1-50, default 1
|
|
36
53
|
maxRetries: 2, // default 2, set 0 to disable
|
|
37
54
|
});
|
|
38
55
|
```
|
|
@@ -41,15 +58,48 @@ const client = new Scavio({
|
|
|
41
58
|
only to transient failures — HTTP 429, 500, 502, 503, 504 and network or timeout
|
|
42
59
|
errors. Backoff is exponential with full jitter, capped at 8s, and a
|
|
43
60
|
`Retry-After` header is honored when the API sends one. Non-transient errors
|
|
44
|
-
(400, 401, 402, 404) are never retried.
|
|
61
|
+
(400, 401, 402, 404, 422) are never retried.
|
|
45
62
|
|
|
46
63
|
`maxRequestsPerSecond` throttles the client so it never sends more than N
|
|
47
64
|
requests in any one-second window. Your plan also has a server-side concurrency
|
|
48
|
-
limit on simultaneous in-flight requests: 1 on free and pay-as-you-go,
|
|
49
|
-
Project,
|
|
65
|
+
limit on simultaneous in-flight requests: 1 on free and pay-as-you-go, 5 on
|
|
66
|
+
Project, 10 on Bootstrap, 15 on Startup, 50 on Growth, and unlimited on
|
|
67
|
+
Enterprise.
|
|
50
68
|
|
|
51
69
|
## API Reference
|
|
52
70
|
|
|
71
|
+
### Extract (any URL)
|
|
72
|
+
|
|
73
|
+
`extract()` is a core endpoint, not a platform, so it lives on the client itself.
|
|
74
|
+
It reads any page and hands it back as readability Markdown, plain text or raw
|
|
75
|
+
HTML — the read-a-page primitive an agent or a RAG ingest needs.
|
|
76
|
+
|
|
77
|
+
```typescript
|
|
78
|
+
const page = await client.extract({
|
|
79
|
+
url: "https://example.com/pricing",
|
|
80
|
+
format: "markdown", // "html" | "markdown" | "text", default "markdown"
|
|
81
|
+
mode: "normal", // "normal" | "advanced" | "ultra", default "normal"
|
|
82
|
+
});
|
|
83
|
+
|
|
84
|
+
page.content; // the page body
|
|
85
|
+
page.content_length; // characters returned
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
Credits depend on `mode`, not on a flat per-call price: `normal` costs 1,
|
|
89
|
+
`advanced` costs 1, `ultra` costs 2. Billing happens only on a successful
|
|
90
|
+
extraction — a dead link, bot wall or timeout costs nothing.
|
|
91
|
+
|
|
92
|
+
- `normal` is a plain datacenter fetch. Start here.
|
|
93
|
+
- `advanced` renders the page in a headless browser. Use it when the content is
|
|
94
|
+
built client-side.
|
|
95
|
+
- `ultra` goes through residential IPs. Use it only when a bot wall blocks the
|
|
96
|
+
other two.
|
|
97
|
+
|
|
98
|
+
`html` is the raw page. `markdown` is a readability extraction with the
|
|
99
|
+
boilerplate stripped. `text` is that markdown flattened to plain text. URLs are
|
|
100
|
+
http(s) only, a bare host is upgraded to https, and loopback, private,
|
|
101
|
+
link-local and cloud-metadata hosts are rejected with a 400.
|
|
102
|
+
|
|
53
103
|
### Google
|
|
54
104
|
|
|
55
105
|
Each method hits its own `/api/v2/google*` endpoint and returns Google's full
|
|
@@ -128,20 +178,47 @@ instead of the previous raw provider payload.
|
|
|
128
178
|
|
|
129
179
|
### Walmart
|
|
130
180
|
|
|
181
|
+
Seven endpoints. `search` and `product` changed shape in 0.15.0 and the other
|
|
182
|
+
five are new.
|
|
183
|
+
|
|
131
184
|
```typescript
|
|
132
|
-
// Search products
|
|
133
185
|
await client.walmart.search({
|
|
134
186
|
query: "tv",
|
|
135
|
-
|
|
136
|
-
|
|
187
|
+
domain: "com", // "com" | "ca" | "com.mx"
|
|
188
|
+
page: 2, // 1-indexed
|
|
189
|
+
sort_by: "rating_high",
|
|
190
|
+
min_price: 100,
|
|
191
|
+
max_price: 500,
|
|
137
192
|
});
|
|
138
193
|
|
|
139
|
-
|
|
140
|
-
await client.walmart.
|
|
141
|
-
|
|
142
|
-
});
|
|
194
|
+
await client.walmart.product({ product_id: "13544111159" });
|
|
195
|
+
await client.walmart.reviews({ product_id: "13544111159", page: 2 });
|
|
196
|
+
await client.walmart.category({ category_id: "3944_133251_1095191" });
|
|
197
|
+
await client.walmart.offers({ product_id: "13544111159" });
|
|
198
|
+
await client.walmart.seller({ seller_id: "101138578" });
|
|
199
|
+
await client.walmart.sellerProducts({ seller_id: "101138578" });
|
|
143
200
|
```
|
|
144
201
|
|
|
202
|
+
Credits are a function of the body on `search` and `category`: `domain` "com" and
|
|
203
|
+
"ca" cost 1 credit, "com.mx" costs 2. The other five endpoints always cost 1.
|
|
204
|
+
|
|
205
|
+
#### Walmart changed in 0.15.0 (breaking)
|
|
206
|
+
|
|
207
|
+
- `device`, `delivery_zip` and `store_id` are retired. Sending one still returns
|
|
208
|
+
200, with a top-level `warnings` array naming what was ignored.
|
|
209
|
+
- `domain` is **not** retired — it is the price-bearing param, and it is accepted
|
|
210
|
+
on `search` and `category` only. Walmart.ca product pages cannot be fetched, so
|
|
211
|
+
every product-keyed endpoint is walmart.com only.
|
|
212
|
+
- `page` (1-indexed) is the paging param; `start_page` remains a deprecated alias.
|
|
213
|
+
- `sort_by` gained `rating_high` and `new`.
|
|
214
|
+
- `fulfillment_speed` is `today` or `tomorrow` only. There is deliberately no
|
|
215
|
+
`2_days` (it leaks 3-4 day items) and no `anytime` — omit the param instead.
|
|
216
|
+
- `offers` returns the **buy-box seller only**, not the full offer list.
|
|
217
|
+
- `seller_id` must be the numeric catalog id (`seller_catalog_id`, returned by
|
|
218
|
+
`product` and `offers`). The GUID form of the id returns 404.
|
|
219
|
+
- `sellerProducts` has **no pagination** — roughly the first 40 server-rendered
|
|
220
|
+
items. `total_count` reports the seller's real catalog size.
|
|
221
|
+
|
|
145
222
|
### YouTube
|
|
146
223
|
|
|
147
224
|
Credit cost varies by endpoint: `transcript` costs 8; `streams` costs 3;
|
|
@@ -431,6 +508,389 @@ await client.instagram.userFollowers({ username: "instagram", count: 50 });
|
|
|
431
508
|
await client.instagram.userFollowings({ username: "instagram" });
|
|
432
509
|
```
|
|
433
510
|
|
|
511
|
+
### Threads
|
|
512
|
+
|
|
513
|
+
Six endpoints, and the credit cost is a function of the body: **2 credits when
|
|
514
|
+
you address a user by `user_id`, 4 when you address them by `username`.** The
|
|
515
|
+
upstream handle lookup is dead, so a handle buys a second call. Only `profile`,
|
|
516
|
+
`userPosts` and `userReplies` are username-keyed; `post`, `postComments` and
|
|
517
|
+
`searchUsers` always cost 2.
|
|
518
|
+
|
|
519
|
+
```typescript
|
|
520
|
+
// Resolve the handle once, then stay on the cheap path
|
|
521
|
+
const found = await client.threads.searchUsers({ query: "zuck" });
|
|
522
|
+
|
|
523
|
+
await client.threads.profile({ user_id: "63625256886" });
|
|
524
|
+
await client.threads.userPosts({ user_id: "63625256886", cursor });
|
|
525
|
+
await client.threads.userReplies({ user_id: "63625256886" });
|
|
526
|
+
await client.threads.post({ url: "https://www.threads.net/@zuck/post/..." });
|
|
527
|
+
await client.threads.postComments({ post_id: "3141..." });
|
|
528
|
+
```
|
|
529
|
+
|
|
530
|
+
There is no Threads content search — `searchUsers` is people search and it is
|
|
531
|
+
the only search Threads exposes. Missing or conflicting identifiers return 422,
|
|
532
|
+
not 400; no match returns 404.
|
|
533
|
+
|
|
534
|
+
### Kuaishou
|
|
535
|
+
|
|
536
|
+
Fourteen endpoints, priced **per endpoint** rather than flat: `videosBatch`
|
|
537
|
+
costs 40, `profile` and the four `search*` methods cost 10, `video` costs 2, and
|
|
538
|
+
everything else costs 1.
|
|
539
|
+
|
|
540
|
+
```typescript
|
|
541
|
+
await client.kuaishou.userResolve({ share_link: "https://v.kuaishou.com/..." }); // 1
|
|
542
|
+
await client.kuaishou.userPosts({ user_id: "3xabc..." }); // 1
|
|
543
|
+
await client.kuaishou.tagFeed({ tag: "美食" }); // 1
|
|
544
|
+
await client.kuaishou.video({ photo_id: "3xdef..." }); // 2
|
|
545
|
+
await client.kuaishou.searchVideos({ keyword: "coffee" }); // 10
|
|
546
|
+
await client.kuaishou.videosBatch({ photo_ids: ["3xdef...", "3xghi..."] }); // 40
|
|
547
|
+
```
|
|
548
|
+
|
|
549
|
+
`videosBatch` costs 40 whether you send 1 id or 20, so fill the batch. If all
|
|
550
|
+
you have is a share link, `userResolve()` turns it into a user id for 1 credit
|
|
551
|
+
rather than paying 10 for `profile()`. Missing identifiers return 422.
|
|
552
|
+
|
|
553
|
+
### eBay
|
|
554
|
+
|
|
555
|
+
```typescript
|
|
556
|
+
await client.ebay.search({ query: "airpods pro", sold: true, per_page: 120 });
|
|
557
|
+
await client.ebay.search({ seller: "musicmagpie" }); // no keyword needed
|
|
558
|
+
await client.ebay.product({ item_id: "126543210987" });
|
|
559
|
+
await client.ebay.seller({ seller: "musicmagpie" });
|
|
560
|
+
```
|
|
561
|
+
|
|
562
|
+
1 credit per call. `sold: true` searches completed listings that actually sold —
|
|
563
|
+
the price-research view; eBay publishes no headline count there, so
|
|
564
|
+
`total_results` comes back null. `per_page` accepts only 60, 120 or 240. `seller`
|
|
565
|
+
is a profile endpoint and cannot enumerate a catalogue — page a seller's
|
|
566
|
+
inventory through `search({ seller })` instead.
|
|
567
|
+
|
|
568
|
+
### Target
|
|
569
|
+
|
|
570
|
+
```typescript
|
|
571
|
+
await client.target.search({ keyword: "office chair", store_id: "1234" });
|
|
572
|
+
await client.target.category({ category_id: "5xtg6" });
|
|
573
|
+
await client.target.product({ tcin: "82291396" });
|
|
574
|
+
await client.target.reviews({ tcin: "82291396" });
|
|
575
|
+
```
|
|
576
|
+
|
|
577
|
+
1 credit per call, but these are the slowest endpoints in the SDK: product about
|
|
578
|
+
4s, search about 9s, category about 37s, reviews about 40s. Raise `timeout`
|
|
579
|
+
accordingly. `reviews` returns at most 8 review bodies whatever `review_count`
|
|
580
|
+
says, and `limit` only trims — there is no paging. `seller_*` is null on
|
|
581
|
+
first-party stock, which means "sold by Target"; only Target Plus marketplace
|
|
582
|
+
listings name a vendor.
|
|
583
|
+
|
|
584
|
+
### Home Depot
|
|
585
|
+
|
|
586
|
+
```typescript
|
|
587
|
+
await client.homeDepot.search({ query: "cordless drill", page: 2 });
|
|
588
|
+
await client.homeDepot.product({ item_id: "313159056" });
|
|
589
|
+
await client.homeDepot.reviews({ item_id: "313159056", page: 2 });
|
|
590
|
+
```
|
|
591
|
+
|
|
592
|
+
2 credits per call. Search page size is fixed at 12, so paging is the only way
|
|
593
|
+
to read further; reviews are 30 per page and a page past `total_pages` is a 404.
|
|
594
|
+
`product` carries only a 10-review preview — `reviews` is the paginated surface.
|
|
595
|
+
|
|
596
|
+
### Zillow
|
|
597
|
+
|
|
598
|
+
```typescript
|
|
599
|
+
await client.zillow.search({
|
|
600
|
+
location: "Austin, TX",
|
|
601
|
+
listing_status: "for_rent",
|
|
602
|
+
min_price: 1500, // MONTHLY RENT on for_rent
|
|
603
|
+
max_price: 3000,
|
|
604
|
+
});
|
|
605
|
+
await client.zillow.property({ zpid: "29433327" });
|
|
606
|
+
await client.zillow.agentReviews({ screen_name: "jane-doe" });
|
|
607
|
+
```
|
|
608
|
+
|
|
609
|
+
1 credit per call. A bare ZIP works on its own but cannot be combined with a
|
|
610
|
+
filter or a sort — Zillow then geolocates the request and answers about another
|
|
611
|
+
city, so pass the city name whenever you filter. On `listing_status: "for_rent"`,
|
|
612
|
+
`min_price`/`max_price` mean monthly rent. `agentReviews` addresses an agent
|
|
613
|
+
profile, not a property, and returns the five reviews Zillow server-renders
|
|
614
|
+
(`total_review_count` is the real total). An unresolvable region is a 404.
|
|
615
|
+
|
|
616
|
+
### Redfin
|
|
617
|
+
|
|
618
|
+
```typescript
|
|
619
|
+
await client.redfin.search({ location: "https://www.redfin.com/city/30749/TX/Austin" });
|
|
620
|
+
await client.redfin.search({ region_id: 30749, region_type: 6 }); // 6 = city
|
|
621
|
+
await client.redfin.property({ property_id: "185301234" });
|
|
622
|
+
await client.redfin.market({ region_id: 30749, region_type: 6 });
|
|
623
|
+
```
|
|
624
|
+
|
|
625
|
+
1 credit per call. **City names are not accepted** on `location` — pass a
|
|
626
|
+
redfin.com region URL (`/city/`, `/neighborhood/`, `/county/`, `/zipcode/`) or
|
|
627
|
+
`region_id` plus `region_type`, which must be sent together. `region_id` is not
|
|
628
|
+
a ZIP code. `sold_within_days` is only valid with `listing_status: "sold"`.
|
|
629
|
+
`days_on_market` is always null in the response — do not build on it.
|
|
630
|
+
|
|
631
|
+
### Booking
|
|
632
|
+
|
|
633
|
+
```typescript
|
|
634
|
+
await client.booking.search({
|
|
635
|
+
destination: "Lisbon",
|
|
636
|
+
checkin: "2026-09-10",
|
|
637
|
+
checkout: "2026-09-13", // send both or neither
|
|
638
|
+
adults: 2,
|
|
639
|
+
currency: "USD",
|
|
640
|
+
});
|
|
641
|
+
await client.booking.hotel({ hotel: "the-independente" });
|
|
642
|
+
await client.booking.reviews({ hotel: "the-independente" });
|
|
643
|
+
```
|
|
644
|
+
|
|
645
|
+
1 credit per call. `checkin` and `checkout` must be sent together — Booking
|
|
646
|
+
ignores a lone check-in and prices its own date range. `hotel` and `reviews`
|
|
647
|
+
take dates for the same reason: Booking prices a stay, and the response echoes
|
|
648
|
+
whichever dates were used. `currency` defaults to USD; without it Booking prices
|
|
649
|
+
off the proxy exit and two identical requests disagree. A search with neither
|
|
650
|
+
`destination` nor `dest_id` returns Booking's homepage and still costs a credit.
|
|
651
|
+
|
|
652
|
+
### Airbnb
|
|
653
|
+
|
|
654
|
+
```typescript
|
|
655
|
+
await client.airbnb.search({
|
|
656
|
+
location: "Barcelona",
|
|
657
|
+
check_in: "2026-09-10",
|
|
658
|
+
check_out: "2026-09-15", // send both or neither
|
|
659
|
+
currency: "USD",
|
|
660
|
+
});
|
|
661
|
+
await client.airbnb.listing({ listing_id: "1234567890" });
|
|
662
|
+
await client.airbnb.reviews({ listing_id: "1234567890", limit: 50, offset: 50 });
|
|
663
|
+
```
|
|
664
|
+
|
|
665
|
+
1 credit per call. **Prices are search-only** — the listing page carries no
|
|
666
|
+
nightly rate under any parameters. A dateless search defaults to +30d for 5
|
|
667
|
+
nights and Airbnb A/Bs both the window and the prices, so the response flags it
|
|
668
|
+
as `dates_are_defaulted`; send real dates for anything you intend to compare.
|
|
669
|
+
The rating breakdown and review tags live on `listing`, while `reviews` returns
|
|
670
|
+
the review bodies with `limit`/`offset` paging.
|
|
671
|
+
|
|
672
|
+
### Tripadvisor
|
|
673
|
+
|
|
674
|
+
**Start with `locations()`.** Every other endpoint is keyed by ids that exist
|
|
675
|
+
only inside TripAdvisor's own URLs.
|
|
676
|
+
|
|
677
|
+
```typescript
|
|
678
|
+
const places = await client.tripadvisor.locations({ query: "Le Bernardin" });
|
|
679
|
+
await client.tripadvisor.search({ geo_id: "60763", category: "restaurants" });
|
|
680
|
+
await client.tripadvisor.location({ location_id: "426986", geo_id: "60763" });
|
|
681
|
+
await client.tripadvisor.reviews({ location_id: "426986", geo_id: "60763", page: 2 });
|
|
682
|
+
```
|
|
683
|
+
|
|
684
|
+
2 credits per call. A geo row from `locations()` gives the `geo_id` that
|
|
685
|
+
`search()` takes; a business row gives the `geo_id` + `location_id` pair that
|
|
686
|
+
`location()` and `reviews()` take. Page 1 of the reviews already ships inside
|
|
687
|
+
`location()`, so use `reviews()` to page past it. Review page size differs by
|
|
688
|
+
family (15 restaurants, 10 hotels and attractions), consecutive pages can repeat
|
|
689
|
+
one review at the boundary (de-duplicate on `review_id`), and a page past the
|
|
690
|
+
last is a 404.
|
|
691
|
+
|
|
692
|
+
### Yelp
|
|
693
|
+
|
|
694
|
+
```typescript
|
|
695
|
+
await client.yelp.search({ term: "ramen", location: "Seattle, WA" });
|
|
696
|
+
await client.yelp.business({ business_id: "..." });
|
|
697
|
+
await client.yelp.reviews({ business_id: "...", page: 2 }); // page 2, not 1
|
|
698
|
+
```
|
|
699
|
+
|
|
700
|
+
2 credits per call. `location` is effectively required — without it Yelp
|
|
701
|
+
geolocates off the proxy exit and the same request answers about a different
|
|
702
|
+
metro run to run. **Reviews page 1 is redundant**: it re-fetches the document
|
|
703
|
+
`business()` already returned and costs another 2 credits, so start at page 2.
|
|
704
|
+
Page size is fixed at 10 and a page past the last is a 404.
|
|
705
|
+
|
|
706
|
+
### Indeed
|
|
707
|
+
|
|
708
|
+
```typescript
|
|
709
|
+
await client.indeed.search({
|
|
710
|
+
query: "data engineer",
|
|
711
|
+
location: "Remote",
|
|
712
|
+
radius: 25, // 0, 5, 10, 15, 25, 35, 50 or 100 only
|
|
713
|
+
max_age_days: 7, // 1, 3, 7 or 14 only
|
|
714
|
+
});
|
|
715
|
+
await client.indeed.job({ job_id: "a1b2c3d4e5f6" });
|
|
716
|
+
await client.indeed.company({ company: "Stripe" });
|
|
717
|
+
await client.indeed.companyReviews({ company: "Stripe", page: 2 });
|
|
718
|
+
```
|
|
719
|
+
|
|
720
|
+
2 credits per call. `radius` and `max_age_days` are closed sets — Indeed ignores
|
|
721
|
+
anything else and silently returns the unfiltered set, so an unsupported radius
|
|
722
|
+
bills you for a search covering fifty miles. `min_salary` filters on Indeed's own
|
|
723
|
+
estimate for the role, not on a posted figure, so postings that publish no salary
|
|
724
|
+
still match. A location-only search (no `query`) is valid. Search is 10 postings
|
|
725
|
+
per page, company reviews 20.
|
|
726
|
+
|
|
727
|
+
### Glassdoor
|
|
728
|
+
|
|
729
|
+
**Start with `companies()`.** The other three are keyed by an `employer_id` that
|
|
730
|
+
only exists inside Glassdoor's `/Overview/` URLs.
|
|
731
|
+
|
|
732
|
+
```typescript
|
|
733
|
+
const hits = await client.glassdoor.companies({ query: "Stripe" });
|
|
734
|
+
const company = await client.glassdoor.company({ employer_id: "671932" });
|
|
735
|
+
|
|
736
|
+
// Pass the URLs the company response returns - halves the upstream work
|
|
737
|
+
await client.glassdoor.reviews({ url: company.reviews_url as string });
|
|
738
|
+
await client.glassdoor.salaries({ url: company.salaries_url as string });
|
|
739
|
+
```
|
|
740
|
+
|
|
741
|
+
1 credit per call. Glassdoor's login wall caps `reviews` at **three reviews per
|
|
742
|
+
response** — there is deliberately no `page` param; move the window with
|
|
743
|
+
`category` and `employment_status` and read `filtered_review_count` to see how
|
|
744
|
+
many match. Addressing `reviews` or `salaries` by `employer_id` costs two
|
|
745
|
+
upstream fetches because the slugs are case-sensitive and must be read off the
|
|
746
|
+
profile, so prefer the `reviews_url` / `salaries_url` the company response hands
|
|
747
|
+
back. These endpoints are slow and flaky by nature (company about 3-47s, reviews
|
|
748
|
+
about 75s, salaries about 41s) — raise `timeout` and keep retries on.
|
|
749
|
+
|
|
750
|
+
### App Store
|
|
751
|
+
|
|
752
|
+
```typescript
|
|
753
|
+
await client.appStore.search({ term: "meditation", limit: 100, country: "us" });
|
|
754
|
+
await client.appStore.app({ app_id: "com.apple.Pages" }); // or the numeric id
|
|
755
|
+
await client.appStore.reviews({ app_id: "361309726", page: 2, sort: "most_helpful" });
|
|
756
|
+
```
|
|
757
|
+
|
|
758
|
+
1 credit per call. **Search has no pagination** — `limit` (1-200) is the only
|
|
759
|
+
lever; every offset spelling is silently ignored. `app` accepts both a numeric
|
|
760
|
+
App Store id and a bundle id; `reviews` is numeric-only. Reviews hard-stop at
|
|
761
|
+
page 10 (50 per page), which is Apple's anonymous ceiling — reach further by
|
|
762
|
+
asking a different `country`. Reviews cannot 404: an unknown id and an app with
|
|
763
|
+
zero reviews return the same empty feed.
|
|
764
|
+
|
|
765
|
+
### Google Play
|
|
766
|
+
|
|
767
|
+
```typescript
|
|
768
|
+
await client.googlePlay.search({ query: "meditation", hl: "en", gl: "us" });
|
|
769
|
+
await client.googlePlay.app({ app_id: "com.spotify.music" });
|
|
770
|
+
await client.googlePlay.reviews({ app_id: "com.spotify.music", sort: "newest", count: 200 });
|
|
771
|
+
```
|
|
772
|
+
|
|
773
|
+
2 credits per call. **Search does not paginate** — one shelf of about 30 apps.
|
|
774
|
+
`hl` changes the storefront, not just the strings: title, description, install
|
|
775
|
+
formatting and content rating all move with it. The reviews `cursor` is opaque,
|
|
776
|
+
single-use, and encodes the sort as well as the position, so send it back with
|
|
777
|
+
the same `sort` it came from; a cursor past the last review is a 404. `app`
|
|
778
|
+
already returns the 20 reviews Play server-renders, plus the real install count
|
|
779
|
+
Play publishes but never displays.
|
|
780
|
+
|
|
781
|
+
### G2
|
|
782
|
+
|
|
783
|
+
```typescript
|
|
784
|
+
await client.g2.search({ query: "crm", rating: 4 });
|
|
785
|
+
await client.g2.product({ product_id: "notion" });
|
|
786
|
+
await client.g2.reviews({ product_id: "notion", page: 2, company_size: "enterprise" });
|
|
787
|
+
```
|
|
788
|
+
|
|
789
|
+
**5 credits per call — the most expensive platform in the SDK**, because g2.com
|
|
790
|
+
bills 25 upstream credits per fetch. Retries are deliberately conservative for
|
|
791
|
+
that reason, and a bot wall arrives as a billed 200 rather than an error. G2
|
|
792
|
+
loads review text in a separate frame, so `product` carries no reviews — call
|
|
793
|
+
`reviews`, which is also the only place with exact per-star counts, pros/cons by
|
|
794
|
+
theme, and company-size / role / industry / region facets.
|
|
795
|
+
|
|
796
|
+
### Capterra
|
|
797
|
+
|
|
798
|
+
```typescript
|
|
799
|
+
const hits = await client.capterra.search({ query: "project management" });
|
|
800
|
+
await client.capterra.product({ product_id: "186596", slug: "Notion" });
|
|
801
|
+
await client.capterra.reviews({ product_id: "186596", slug: "Notion", page: 2 });
|
|
802
|
+
```
|
|
803
|
+
|
|
804
|
+
2 credits per call. **Search does not paginate** — Capterra fixes the result set
|
|
805
|
+
at 20, so there is no `page` param. `slug` is cosmetic on `product` but
|
|
806
|
+
load-bearing and case-sensitive on `reviews`: a wrong one silently serves page 1
|
|
807
|
+
under a billed 200, so pass back the `slug` or `reviews_url` that `search` or
|
|
808
|
+
`product` returned. Reviews are 25 per page, capped at page 100, and page 1
|
|
809
|
+
already ships inside `product`. `vendor` is null on the product profile —
|
|
810
|
+
Capterra does not publish it there.
|
|
811
|
+
|
|
812
|
+
### SEC EDGAR
|
|
813
|
+
|
|
814
|
+
**Start with `lookup()`.** Callers hold a ticker; EDGAR is keyed by CIK.
|
|
815
|
+
|
|
816
|
+
```typescript
|
|
817
|
+
const match = await client.sec.lookup({ query: "AAPL" });
|
|
818
|
+
await client.sec.company({ ticker: "AAPL" });
|
|
819
|
+
await client.sec.filings({ ticker: "AAPL", form: "10-K", include_history: true });
|
|
820
|
+
await client.sec.facts({ ticker: "AAPL", query: "revenue" });
|
|
821
|
+
await client.sec.concept({ ticker: "AAPL", concept: "NetIncomeLoss" });
|
|
822
|
+
await client.sec.search({ query: "climate risk", form: "10-K" });
|
|
823
|
+
```
|
|
824
|
+
|
|
825
|
+
1 credit per call, including `include_history: true`, which is the one call that
|
|
826
|
+
can buy up to 10 upstream fetches. Both `cik` and `ticker` accept either
|
|
827
|
+
spelling. XBRL concept tags are **case-sensitive** — `netincomeloss` is a 404
|
|
828
|
+
upstream, so use `facts()` to see what a filer actually reports. EDGAR's
|
|
829
|
+
"recent" block is not a fixed window: a decade for a quiet filer, about a year
|
|
830
|
+
for a prolific one.
|
|
831
|
+
|
|
832
|
+
### Companies House
|
|
833
|
+
|
|
834
|
+
```typescript
|
|
835
|
+
const hits = await client.companiesHouse.search({ query: "Monzo" });
|
|
836
|
+
await client.companiesHouse.company({ company_number: "09446231" });
|
|
837
|
+
await client.companiesHouse.officers({ company_number: "09446231", page: 2 });
|
|
838
|
+
await client.companiesHouse.filingHistory({ company_number: "SC090312" });
|
|
839
|
+
```
|
|
840
|
+
|
|
841
|
+
1 credit per call. `company_number` is deliberately loose — the register 404s on
|
|
842
|
+
numbers that lost their leading zeros or arrived lower-cased, so the transport
|
|
843
|
+
pads and upper-cases before asking. SC, NI, OC, SO, NC, FC, BR and CE prefixes
|
|
844
|
+
are all supported. Search is 20 per page and capped at page 50: the register
|
|
845
|
+
serves a 1000-result window per term whatever hit count it advertises, and page
|
|
846
|
+
51 is an HTTP 416. Officers and filing history have no upper page bound — past
|
|
847
|
+
the last page you get an ordinary empty list.
|
|
848
|
+
|
|
849
|
+
### Google Ads Transparency
|
|
850
|
+
|
|
851
|
+
```typescript
|
|
852
|
+
const advertisers = await client.googleAds.advertisers({ query: "nike" });
|
|
853
|
+
const page1 = await client.googleAds.search({ advertiser_id: "AR123...", region: "DE" });
|
|
854
|
+
const page2 = await client.googleAds.search({
|
|
855
|
+
advertiser_id: "AR123...",
|
|
856
|
+
region: "DE",
|
|
857
|
+
cursor: page1.next_cursor as string, // re-send the SAME filters
|
|
858
|
+
});
|
|
859
|
+
await client.googleAds.creative({ advertiser_id: "AR123...", creative_id: "CR456..." });
|
|
860
|
+
```
|
|
861
|
+
|
|
862
|
+
1 credit per call. `search` paginates by `cursor` / `next_cursor` at 100 per
|
|
863
|
+
page — re-send the same filters alongside the cursor. `limit` is capped at 100 by
|
|
864
|
+
Google itself, which answers a larger request with zero rows rather than an
|
|
865
|
+
error. `advertisers` and `creative` do not paginate. **Impressions and reach are
|
|
866
|
+
EEA-only**: US creatives return null for `impressions_min`, `impressions_max` and
|
|
867
|
+
`first_shown` because Google only publishes reach where the DSA compels it. The
|
|
868
|
+
text, image and video format sets are disjoint — an advertiser's creatives never
|
|
869
|
+
overlap between them.
|
|
870
|
+
|
|
871
|
+
### Meta Ad Library
|
|
872
|
+
|
|
873
|
+
```typescript
|
|
874
|
+
const page1 = await client.metaAds.search({ query: "protein powder", country: "US" });
|
|
875
|
+
const page2 = await client.metaAds.search({
|
|
876
|
+
query: "protein powder",
|
|
877
|
+
cursor: page1.next_cursor as string,
|
|
878
|
+
});
|
|
879
|
+
await client.metaAds.advertiser({ page_id: "10150125871..." });
|
|
880
|
+
await client.metaAds.ad({ ad_archive_id: "1234567890123456" });
|
|
881
|
+
```
|
|
882
|
+
|
|
883
|
+
1 credit per call. `search` and `advertiser` paginate all the way through: page 1
|
|
884
|
+
returns 30 ads, then 10 per page via `next_cursor` — walk `has_next_page` to pull
|
|
885
|
+
a whole query or advertiser. The cursor is a self-contained opaque blob, so
|
|
886
|
+
paging is stateless and **the other filters are ignored when a cursor is
|
|
887
|
+
present**, so re-sending `query` alongside it is harmless and satisfies the
|
|
888
|
+
schema — the cursor already carries the filters. `total_results` caps at 50000 with
|
|
889
|
+
`total_is_capped: true`, because Meta only reports "more than 50,000". Spend,
|
|
890
|
+
reach, impressions and the paid-for-by disclosure exist on political and issue
|
|
891
|
+
ads only — set `ad_type: "political_and_issue_ads"` to expose them; commercial
|
|
892
|
+
ads leave those fields null.
|
|
893
|
+
|
|
434
894
|
### Usage
|
|
435
895
|
|
|
436
896
|
```typescript
|
|
@@ -460,13 +920,18 @@ All error classes:
|
|
|
460
920
|
| `MissingAPIKeyError` | — | No API key provided |
|
|
461
921
|
| `ScavioConnectionError` | — | Request never reached the API (DNS, reset, TLS) |
|
|
462
922
|
| `ScavioTimeoutError` | — | Request exceeded the configured `timeout` |
|
|
463
|
-
| `BadRequestError` | 400 | Invalid request parameters |
|
|
923
|
+
| `BadRequestError` | 400, 422 | Invalid request parameters |
|
|
464
924
|
| `InvalidAPIKeyError` | 401 | Invalid API key |
|
|
465
925
|
| `InsufficientCreditsError` | 402 | No credits remaining |
|
|
466
926
|
| `NotFoundError` | 404 | No data upstream for that id (see TikTok Shop above) |
|
|
467
927
|
| `RateLimitError` | 429 | Rate limit exceeded |
|
|
468
928
|
| `ScavioAPIError` | other | Catch-all (has `.statusCode`) |
|
|
469
929
|
|
|
930
|
+
Threads and Kuaishou answer a missing or conflicting identifier with **422**, not
|
|
931
|
+
400 — those routes have no 400 at all. Both map to `BadRequestError`, so one
|
|
932
|
+
`catch (e) { if (e instanceof BadRequestError) }` covers validation failures on
|
|
933
|
+
every platform. `e.statusCode` still reports whichever status the API sent.
|
|
934
|
+
|
|
470
935
|
Every class extends `ScavioError`, so `catch (e) { if (e instanceof ScavioError) }`
|
|
471
936
|
matches all of them. All except `MissingAPIKeyError`, `ScavioConnectionError` and
|
|
472
937
|
`ScavioTimeoutError` carry `.statusCode` and `.responseBody`.
|
|
@@ -486,14 +951,24 @@ MIT
|
|
|
486
951
|
|
|
487
952
|
## About Scavio
|
|
488
953
|
|
|
489
|
-
[Scavio](https://scavio.dev) is a unified
|
|
954
|
+
[Scavio](https://scavio.dev) is a unified web data and
|
|
955
|
+
[search API for AI agents](https://scavio.dev/search-api-for-ai-agents) — one API
|
|
956
|
+
key, structured JSON, no proxies or browser farms to run. It is a real-time
|
|
957
|
+
[Tavily alternative](https://scavio.dev/alternatives/tavily) and
|
|
958
|
+
[SerpAPI alternative](https://scavio.dev/alternatives/serpapi), and with
|
|
959
|
+
`extract()` it also covers the read-any-URL job people reach for Firecrawl to do.
|
|
960
|
+
|
|
961
|
+
What teams build on it:
|
|
490
962
|
|
|
491
|
-
- [Google Search API](https://scavio.dev/google-search-api)
|
|
492
|
-
- [Amazon Product API](https://scavio.dev/amazon-product-api)
|
|
493
|
-
-
|
|
494
|
-
-
|
|
495
|
-
-
|
|
963
|
+
- **SERP and answer engines** — [Google Search API](https://scavio.dev/google-search-api) for organic results, news, images, maps and the knowledge graph
|
|
964
|
+
- **Price and catalog monitoring** — [Amazon Product API](https://scavio.dev/amazon-product-api), [Walmart Product API](https://scavio.dev/walmart-product-api), eBay, Target and Home Depot product, review and seller data
|
|
965
|
+
- **Real estate and travel pipelines** — Zillow and Redfin listings and market stats, Booking, Airbnb and Tripadvisor rates and reviews
|
|
966
|
+
- **Review mining and competitive research** — Yelp, G2, Capterra, Glassdoor, App Store and Google Play reviews on one shape
|
|
967
|
+
- **Recruiting and company intelligence** — Indeed jobs, Glassdoor salaries, SEC EDGAR filings and XBRL facts, UK Companies House officers and filing history
|
|
968
|
+
- **Ad and creative intelligence** — Google Ads Transparency Center and Meta Ad Library creatives
|
|
969
|
+
- **Social listening** — [YouTube API](https://scavio.dev/youtube-transcript-api), [TikTok API](https://scavio.dev/tiktok-api), [Instagram API](https://scavio.dev/instagram-api), [Reddit API](https://scavio.dev/reddit-api), [X API](https://scavio.dev/docs/x-search), [LinkedIn API](https://scavio.dev/docs/linkedin-person), Threads and Kuaishou
|
|
970
|
+
- **RAG ingestion** — `extract()` reads any URL and hands back readability Markdown ready to chunk and embed
|
|
496
971
|
|
|
497
972
|
Teams choosing between providers can [compare Scavio vs alternatives](https://scavio.dev/compare) side by side.
|
|
498
973
|
|
|
499
|
-
Get a free [API key](https://dashboard.scavio.dev) and explore the [documentation](https://scavio.dev/docs/introduction).
|
|
974
|
+
Get a free [API key](https://dashboard.scavio.dev/sign-up) and explore the [documentation](https://scavio.dev/docs/introduction).
|