@datafuel/sdk 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Datafuel
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md ADDED
@@ -0,0 +1,192 @@
1
+ # @datafuel/sdk
2
+
3
+ TypeScript client for the [DataFuel](https://datafuel.ai) scraping API. No runtime dependencies, Node 20.3+.
4
+
5
+ ```bash
6
+ npm install @datafuel/sdk
7
+ ```
8
+
9
+ ```ts
10
+ import { DataFuel } from "@datafuel/sdk";
11
+
12
+ const df = new DataFuel(); // or new DataFuel("df_key_..."); reads DATAFUEL_API_KEY
13
+
14
+ const markdown = await df.markdown("https://example.com");
15
+ ```
16
+
17
+ ## Pick the call
18
+
19
+ | You have | Call | Waits? |
20
+ | ---------------------------- | ----------------------------------------------------- | ------------- |
21
+ | One URL | `scrape` / `markdown` | yes |
22
+ | A site, need its URL list | `map` | yes |
23
+ | A start URL, need many pages | `crawl`, or `startCrawl` + `waitCrawl` + `crawlPages` | `crawl` does |
24
+ | A list of known URLs | `runJob`, or `createJob` + `waitJob` + `jobResults` | `runJob` does |
25
+ | A question for an AI engine | `ask` | yes |
26
+
27
+ Start with plain `scrape`. Turn on `jsRendering` only when the page comes back empty: it is slower and costs five times the credits on a Basic proxy. `map` a section before you `crawl` it, it costs one credit and tells you how big it is.
28
+
29
+ ## Scrape
30
+
31
+ ```ts
32
+ import { DataFuel, Blocked } from "@datafuel/sdk";
33
+
34
+ const df = new DataFuel();
35
+
36
+ try {
37
+ const res = await df.scrape("https://shop.example.com/p/42", {
38
+ proxy: { type: "Premium", country: "US" },
39
+ jsRendering: true,
40
+ waitFor: "#price",
41
+ extract: { title: "h1", price: "#price" },
42
+ });
43
+ console.log(res.data); // { title: [...], price: [...] }
44
+ } catch (error) {
45
+ if (error instanceof Blocked) {
46
+ // 403/429/503 or an anti-bot wall. Refunded. error.protection names the vendor.
47
+ }
48
+ throw error;
49
+ }
50
+ ```
51
+
52
+ `Result` carries the metadata next to the content. Read it first:
53
+
54
+ | Field | Meaning |
55
+ | ------------------------ | ----------------------------------------------------------------- |
56
+ | `statusCode` | What the target answered. A 404 still completes and bills. |
57
+ | `finalUrl`, `redirected` | Where the page really came from. |
58
+ | `blocked`, `protection` | The target refused the request. The task failed and was refunded. |
59
+ | `creditsUsed` | Charged for this task, 0 when it failed. |
60
+
61
+ `res.text` returns html or markdown, `res.data` structured output, `res.image` screenshot bytes.
62
+
63
+ Page options, shared by `scrape`, jobs and crawls: `format` (`html`, `markdown`, `json`, `png`, `jpeg`), `jsRendering`, `waitFor`, `waitForTimeoutMs`, `jsInstructions`, `blockResource`, `mainContentOnly`, `includeImages`, `extract`, `extractRegex`, `template`, `method`, `body`, `contentType`, `headers`, `headerOrder`, `cookies`, `userAgent`, `userAgentType`, `ai`.
64
+
65
+ ## AI after the scrape
66
+
67
+ Point the page at an LLM with your own key and get structured data back:
68
+
69
+ ```ts
70
+ const res = await df.scrape(url, {
71
+ ai: {
72
+ prompt: "extract the product name and its price",
73
+ format: { name: "string", price: "number" },
74
+ provider: "openai", // openai, anthropic, google
75
+ model: "gpt-4o-mini",
76
+ apiKey: process.env.OPENAI_API_KEY,
77
+ },
78
+ });
79
+ res.data; // { name: "...", price: ... }
80
+ ```
81
+
82
+ Works on `scrape` and on URL jobs. Crawls reject it; the SDK says so before sending. `provider`, `model` and `apiKey` must be given together — a half-filled key would be spent on a scrape whose AI step then fails.
83
+
84
+ ## Crawl
85
+
86
+ ```ts
87
+ const id = await df.startCrawl("https://example.com/docs", {
88
+ maxPages: 200,
89
+ excludePaths: ["\\.pdf$"],
90
+ format: "markdown",
91
+ });
92
+ const status = await df.waitCrawl(id); // status.stop_reason: max_pages, max_depth_exhausted, ...
93
+
94
+ for await (const page of df.crawlPages(id)) {
95
+ if (page.pending || !page.ok) continue; // still running, or a failed page (refunded)
96
+ console.log(page.url, page.depth, page.text.length);
97
+ }
98
+ ```
99
+
100
+ `df.crawl(url, options)` does all three in one call:
101
+
102
+ ```ts
103
+ const crawl = await df.crawl("https://example.com/docs", { maxPages: 200 });
104
+ console.log(crawl.status.stop_reason, crawl.status.total_cost, crawl.pages.length);
105
+ ```
106
+
107
+ Check `stop_reason`: `insufficient_credits` means the crawl ended early. If the wait times out, `WaitTimeout.id` still holds the crawl id — it keeps running and billing server side, so pick it up again with `waitCrawl` and `crawlPages`.
108
+
109
+ Unset limits use the API defaults: 100 pages, depth 3, 5 pages in flight.
110
+
111
+ ## Jobs
112
+
113
+ ```ts
114
+ const results = await df.runJob(urls, { format: "markdown" });
115
+ for (const task of results.tasks) {
116
+ if (!task.ok) continue; // this URL failed, the others did not
117
+ console.log(task.text);
118
+ }
119
+ ```
120
+
121
+ `sequential: true` runs the URLs one after the other; the default runs them concurrently, bounded by your account's concurrency limit. `runAskJob(prompts, { engine })` does the same for a batch of prompts.
122
+
123
+ ## Map
124
+
125
+ ```ts
126
+ const site = await df.map("https://example.com", { search: "blog", limit: 500 });
127
+ for (const link of site.links) console.log(link.url);
128
+ ```
129
+
130
+ An empty `site.links` comes with a `site.reason`. `no_links_on_page` usually means the navigation is rendered client-side.
131
+
132
+ ## Errors
133
+
134
+ ```ts
135
+ import * as datafuel from "@datafuel/sdk";
136
+
137
+ try {
138
+ await df.scrape(url);
139
+ } catch (error) {
140
+ if (error instanceof datafuel.Blocked) {
141
+ /* subclass of TaskFailed */
142
+ }
143
+ if (error instanceof datafuel.InsufficientCredits) {
144
+ /* top up */
145
+ }
146
+ throw error;
147
+ }
148
+ ```
149
+
150
+ - `NoApiKey`: no key was passed and `DATAFUEL_API_KEY` is empty. Thrown before any request.
151
+ - `APIError`: the API refused the request. Subclasses: `Unauthorized`, `InsufficientCredits`, `RateLimited`, `NotFound`, `InvalidAttributes`, `IdempotencyKeyReused`.
152
+ - `ModuleUnavailable`, `EngineUnavailable`: an operator switched a task type or LLM engine off, e.g. during a provider outage. The reason is in the message, nothing is charged, and the SDK does not retry. `df.capabilities()` lists what is on.
153
+ - `TaskFailed`, and `Blocked` when the target refused: the API accepted the task but the page could not be scraped. The error carries `.result`, so the envelope is still readable. Failed tasks are refunded.
154
+ - `WaitTimeout`: a wait ran out of time. `.id` picks the work back up.
155
+
156
+ Tasks inside a job or a crawl never throw on their own — check `task.ok`, or call `task.raiseForStatus()`.
157
+
158
+ ## Retries and idempotency
159
+
160
+ Every write carries an `Idempotency-Key`, generated per request unless you pass `idempotencyKey`. That makes retries safe: network errors and 429/502/503/504 answers are retried with backoff, reusing the key, so a retry attaches to the task already running instead of charging twice. Other 4xx errors are never retried, and neither is an aborted request.
161
+
162
+ The same key is how a synchronous task survives a slow page: when the API answers "still processing", the SDK re-sends the identical request under the identical key until the result is ready or the call's timeout is reached.
163
+
164
+ A key replays its stored result, failures included. To run a failed scrape again, send a new request rather than the same key.
165
+
166
+ ## Options
167
+
168
+ ```ts
169
+ const df = new DataFuel({
170
+ apiKey,
171
+ baseUrl: "https://scraping-api.staging.datafuel.ai/api/v1",
172
+ timeoutMs: 300_000, // null disables the bound
173
+ maxRetries: 4,
174
+ pollIntervalMs: 5_000, // waitJob / waitCrawl
175
+ userAgent: "my-app/1.2",
176
+ fetch: myFetch,
177
+ });
178
+ ```
179
+
180
+ Keep `timeoutMs` generous: `scrape` waits until the page is ready, which can take minutes with `jsRendering`. Every call also takes its own `timeoutMs` and an `AbortSignal` as `signal`.
181
+
182
+ If you pass your own `fetch`, leave redirects off. fetch keeps custom headers across a redirect, so a redirect to another host would carry your `X-API-Key` to it. This SDK sends `redirect: "manual"`.
183
+
184
+ Full API reference: https://scraping-api.datafuel.ai/docs
185
+
186
+ ## MCP
187
+
188
+ To use DataFuel from Claude Code, Cursor, VS Code and other MCP clients instead of from code, run `npx -y @datafuel/mcp init`. See [@datafuel/mcp](https://www.npmjs.com/package/@datafuel/mcp).
189
+
190
+ ## License
191
+
192
+ MIT, see [LICENSE](LICENSE).