@datafuel/sdk 0.2.0 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -26,6 +26,8 @@ const markdown = await df.markdown("https://example.com");
26
26
  | A Google search | `search` | yes |
27
27
  | Many prompts or searches | `runAskJob` / `runSearchJob` | yes |
28
28
  | Earlier jobs, tasks, usage | `listJobs` / `listTasks` / `analytics` / `transactions` | yes |
29
+ | What a request may contain | `jsInstructions` / `aiProviders` / `proxyLocations` | yes |
30
+ | The wall in front of a URL | `checkProtection` | yes |
29
31
 
30
32
  Start with plain `scrape`. Turn on `jsRendering` only when the page comes back empty: it is slower and costs five times the credits on a Basic proxy. `map` a section before you `crawl` it, it costs one credit and tells you how big it is.
31
33
 
@@ -63,7 +65,22 @@ try {
63
65
 
64
66
  `res.text` returns html or markdown, `res.data` structured output, `res.image` screenshot bytes.
65
67
 
66
- Page options, shared by `scrape`, jobs and crawls: `format` (`html`, `markdown`, `json`, `png`, `jpeg`), `jsRendering`, `waitFor`, `waitForTimeoutMs`, `jsInstructions` (an object keyed by action, e.g. `{ click: "#more" }`; `df.jsInstructions()` lists the actions), `blockResource`, `mainContentOnly`, `includeImages`, `extract`, `extractRegex`, `template`, `method`, `body`, `contentType`, `headers`, `headerOrder`, `cookies`, `userAgent`, `userAgentType`, `ai`.
68
+ Page options, shared by `scrape`, jobs and crawls: `format` (`html`, `markdown`, `json`, `png`, `jpeg`), `jsRendering`, `waitFor`, `waitForTimeoutMs`, `jsInstructions`, `blockResource` (one resource type or an array, e.g. `["Image", "Font"]`), `mainContentOnly`, `includeImages`, `extract`, `extractRegex`, `template`, `method`, `body`, `contentType`, `headers`, `headerOrder`, `cookies`, `userAgent`, `userAgentType`, `ai`.
69
+
70
+ `jsInstructions` is an array of single-action objects. They run in the order you list them and an action can repeat:
71
+
72
+ ```ts
73
+ await df.scrape(url, {
74
+ jsRendering: true,
75
+ jsInstructions: [
76
+ { fill: ["input[name=q]", "laptops"] },
77
+ { click: "button[type=submit]" },
78
+ { wait_ms: 1000 },
79
+ ],
80
+ });
81
+ ```
82
+
83
+ The older form, one object keyed by action (`{ click: "#more" }`), still works, but its order is not guaranteed and an action cannot repeat. `df.jsInstructions()` lists the actions.
67
84
 
68
85
  `proxy` takes `type`, `country`, `city`, `state`, `asn`, and a sticky `sessionId` with `ttl` (seconds) on `scrape`, `map`, URL jobs and crawls. `df.proxyLocations()` and `df.proxyAsns(country)` list what a proxy type can exit from.
69
86
 
@@ -76,15 +93,14 @@ const res = await df.scrape(url, {
76
93
  ai: {
77
94
  prompt: "extract the product name and its price",
78
95
  format: { name: "string", price: "number" },
79
- provider: "openai", // openai, anthropic, google
80
- model: "gpt-4o-mini",
96
+ provider: "openai", // required; df.aiProviders() lists the providers and their models
81
97
  apiKey: process.env.OPENAI_API_KEY,
82
98
  },
83
99
  });
84
100
  res.data; // { name: "...", price: ... }
85
101
  ```
86
102
 
87
- Works on `scrape` and on URL jobs. Crawls reject it; the SDK says so before sending. `provider`, `model` and `apiKey` must be given together — a half-filled key would be spent on a scrape whose AI step then fails.
103
+ Works on `scrape` and on URL jobs. Crawls reject it; the SDK says so before sending. `provider` is required. `model` is optional: leave it out for the provider's default, or pick one of the models `df.aiProviders()` lists; any other value is rejected with `InvalidAttributes`.
88
104
 
89
105
  ## Crawl
90
106
 
@@ -141,7 +157,7 @@ const serp = await df.search("best crm", { country: "us", language: "en", page:
141
157
  console.log(serp.data); // parsed results page; format "html" or "markdown" for the raw page
142
158
  ```
143
159
 
144
- `search` also takes `location` (or `uule`, or `lat`/`lon` with `radius`), `googleDomain`, `tbs`, `safe`, `cr`, `lr`, `nfpr`, `filter` and `proxyCountry`.
160
+ `search` also takes `location` (or `uule`, or `lat`/`lon` with `radius`), `googleDomain`, `tbs`, `safe`, `cr`, `lr`, `nfpr` and `filter`. Searches leave through DataFuel's own pool: pick the market with `country` and `language`; `proxyCountry` is accepted but not used yet.
145
161
 
146
162
  ## Map
147
163
 
@@ -165,6 +181,23 @@ const history = await df.transactions({ operation: "refund", limit: 50 });
165
181
 
166
182
  `listJobs` and `listTasks` run newest first; pass `nextCursor` back as `cursor` until it is absent. Task items carry no result: call `getTask(id)` for it. Dates are `YYYY-MM-DD` in UTC (a `Date` is sent as its UTC day) and `endDate` is inclusive. `transactions` pages with `page` and `limit`, and its `sums` total each operation over the whole range. In `analytics`, `status_code` 0 means the target never answered (timeout, DNS).
167
183
 
184
+ ```ts
185
+ const credits = await df.balance();
186
+ const { plan_balance, payg_balance } = await df.balanceSplit();
187
+ ```
188
+
189
+ `balance` is what you can spend. It is made of plan credits and pay-as-you-go credits. Plan credits are spent first; unused ones roll over when the plan renews and expire if it is not renewed. Pay-as-you-go credits come from one-time credit packs (a `purchase` transaction), are spent after plan credits and never expire. Each transaction's `plan_amount` is the part of `amount` that moved plan credits.
190
+
191
+ ## Config and health
192
+
193
+ ```ts
194
+ await df.capabilities(); // which task types and LLM engines are on
195
+ await df.jsInstructions(); // the browser actions jsInstructions accepts
196
+ await df.aiProviders(); // the LLM providers and models ai accepts
197
+ await df.checkProtection(["https://shop.example.com/"]); // anti-bot vendor per URL, nothing scraped
198
+ const health = await df.health({ deep: true }); // health.ok is false while a dependency is down
199
+ ```
200
+
168
201
  ## Errors
169
202
 
170
203
  ```ts
@@ -184,7 +217,7 @@ try {
184
217
  ```
185
218
 
186
219
  - `NoApiKey`: no key was passed and `DATAFUEL_API_KEY` is empty. Thrown before any request.
187
- - `APIError`: the API refused the request. `.code` holds the API's error code (typed as `ErrorCode`). Subclasses: `Unauthorized`, `Forbidden`, `InsufficientCredits`, `RateLimited`, `NotFound`, `InvalidAttributes`, `IdempotencyKeyReused`, `JobNotCancellable`.
220
+ - `APIError`: the API refused the request. `.code` holds the API's error code (typed as `ErrorCode`). Subclasses: `Unauthorized`, `Forbidden`, `InsufficientCredits`, `RateLimited`, `NotFound`, `InvalidAttributes`, `IdempotencyKeyReused`, `JobNotCancellable`, `AlreadyExists` (a create collided with an existing task or job; nothing was charged, send it again).
188
221
  - `ModuleUnavailable`, `EngineUnavailable`: an operator switched a task type or LLM engine off, e.g. during a provider outage. The reason is in the message, nothing is charged, and the SDK does not retry. `df.capabilities()` lists what is on.
189
222
  - `TaskFailed`, and `Blocked` when the target refused: the API accepted the task but the page could not be scraped. The error carries `.result`, so the envelope is still readable. Failed tasks are refunded.
190
223
  - `WaitTimeout`: a wait ran out of time. `.id` picks the work back up.
@@ -217,7 +250,7 @@ Keep `timeoutMs` generous: `scrape` waits until the page is ready, which can tak
217
250
 
218
251
  If you pass your own `fetch`, leave redirects off. fetch keeps custom headers across a redirect, so a redirect to another host would carry your `X-API-Key` to it. This SDK sends `redirect: "manual"`.
219
252
 
220
- Full API reference: https://scraping-api.datafuel.ai/docs
253
+ Guides and the full API reference: https://docs.datafuel.ai
221
254
 
222
255
  ## MCP
223
256
 
package/dist/index.cjs CHANGED
@@ -21,6 +21,7 @@ var __toCommonJS = (mod) => __copyProps(__defProp({}, "__esModule", { value: tru
21
21
  var index_exports = {};
22
22
  __export(index_exports, {
23
23
  APIError: () => APIError,
24
+ AlreadyExists: () => AlreadyExists,
24
25
  Blocked: () => Blocked,
25
26
  Capabilities: () => Capabilities,
26
27
  CrawlPage: () => CrawlPage,
@@ -87,6 +88,8 @@ var IdempotencyKeyReused = class extends APIError {
87
88
  };
88
89
  var JobNotCancellable = class extends APIError {
89
90
  };
91
+ var AlreadyExists = class extends APIError {
92
+ };
90
93
  var Unavailable = class extends APIError {
91
94
  };
92
95
  var ModuleUnavailable = class extends Unavailable {
@@ -124,14 +127,15 @@ var BY_CODE = {
124
127
  IDEMPOTENCY_KEY_REUSED: IdempotencyKeyReused,
125
128
  INVALID_API_KEY: Unauthorized,
126
129
  FORBIDDEN: Forbidden,
127
- JOB_NOT_CANCELLABLE: JobNotCancellable
130
+ JOB_NOT_CANCELLABLE: JobNotCancellable,
131
+ TASK_ALREADY_EXISTS: AlreadyExists,
132
+ JOB_ALREADY_EXISTS: AlreadyExists
128
133
  };
129
134
  var BY_STATUS = {
130
135
  401: Unauthorized,
131
136
  402: InsufficientCredits,
132
137
  403: Forbidden,
133
138
  404: NotFound,
134
- 409: JobNotCancellable,
135
139
  422: IdempotencyKeyReused,
136
140
  429: RateLimited,
137
141
  503: Unavailable
@@ -152,19 +156,20 @@ function apiError(status, body, retryAfter = 0) {
152
156
 
153
157
  // src/core.ts
154
158
  var DEFAULT_BASE_URL = "https://scraping-api.datafuel.ai/api/v1";
155
- var VERSION = "0.2.0";
159
+ var VERSION = "0.4.0";
156
160
  var DEFAULT_TIMEOUT_MS = 18e4;
157
161
  var STILL_PROCESSING_DELAY_MS = 2e3;
158
162
  var STILL_PROCESSING = "TASK_STILL_PROCESSING";
159
163
  var MAX_IDEMPOTENCY_KEY = 255;
160
164
  var Request = class {
161
- constructor(method, path, params, body, idempotencyKey, auth = true) {
165
+ constructor(method, path, params, body, idempotencyKey, auth = true, degradedOk = false) {
162
166
  this.method = method;
163
167
  this.path = path;
164
168
  this.params = params;
165
169
  this.body = body;
166
170
  this.idempotencyKey = idempotencyKey;
167
171
  this.auth = auth;
172
+ this.degradedOk = degradedOk;
168
173
  }
169
174
  method;
170
175
  path;
@@ -172,6 +177,7 @@ var Request = class {
172
177
  body;
173
178
  idempotencyKey;
174
179
  auth;
180
+ degradedOk;
175
181
  /** GETs are safe by nature, writes because they carry an idempotency key. */
176
182
  get retryable() {
177
183
  return this.method === "GET" || this.idempotencyKey !== void 0;
@@ -219,15 +225,8 @@ function aiAttributes(ai) {
219
225
  return out;
220
226
  }
221
227
  function validateAI(ai) {
222
- const given = {
223
- provider: ai.provider !== void 0,
224
- model: ai.model !== void 0,
225
- apiKey: ai.apiKey !== void 0
226
- };
227
- const present = Object.values(given).filter(Boolean).length;
228
- if (present > 0 && present < 3) {
229
- const missing = Object.entries(given).filter(([, ok]) => !ok).map(([name]) => name).join(", ");
230
- throw new TypeError(`ai needs provider, model and apiKey together; missing: ${missing}`);
228
+ if (!ai.provider) {
229
+ throw new TypeError("ai needs a provider; aiProviders() lists the ones the API supports");
231
230
  }
232
231
  }
233
232
  function scrapeAttributes(options = {}) {
@@ -237,7 +236,9 @@ function scrapeAttributes(options = {}) {
237
236
  if (options.waitFor) attrs.wait_for_selector = options.waitFor;
238
237
  if (options.waitForTimeoutMs) attrs.wait_for_selector_timeout_ms = options.waitForTimeoutMs;
239
238
  if (options.jsInstructions) attrs.js_instructions = options.jsInstructions;
240
- if (options.blockResource) attrs.block_resource = options.blockResource;
239
+ if (options.blockResource?.length) {
240
+ attrs.block_resource = typeof options.blockResource === "string" ? options.blockResource : [...options.blockResource];
241
+ }
241
242
  if (options.mainContentOnly) attrs.main_content_only = true;
242
243
  if (options.includeImages !== void 0) attrs.include_images = options.includeImages;
243
244
  if (options.extract) attrs.extract_selector = selector(options.extract);
@@ -419,6 +420,17 @@ function crawlResultsRequest(crawlId, cursor, limit) {
419
420
  Object.keys(params).length > 0 ? params : void 0
420
421
  );
421
422
  }
423
+ function healthRequest(deep) {
424
+ return new Request(
425
+ "GET",
426
+ "/healthz",
427
+ deep ? { deep: "1" } : void 0,
428
+ void 0,
429
+ void 0,
430
+ false,
431
+ true
432
+ );
433
+ }
422
434
  function day(value) {
423
435
  return typeof value === "string" ? value : value.toISOString().slice(0, 10);
424
436
  }
@@ -651,7 +663,7 @@ var DataFuel = class {
651
663
  const signal = this.signalFor(opts);
652
664
  const send = this.fetchImpl;
653
665
  const response = await send(url, signal ? { ...init, signal } : init);
654
- return await parse(response);
666
+ return await parse(response, request.degradedOk);
655
667
  } catch (caught) {
656
668
  if (isAbort(caught)) {
657
669
  throw new TransportError("the request was aborted or timed out", { cause: caught });
@@ -1013,6 +1025,37 @@ var DataFuel = class {
1013
1025
  const body = record(await this.send(request, options));
1014
1026
  return Array.isArray(body.instructions) ? body.instructions : [];
1015
1027
  }
1028
+ /** The LLM providers and models `ai` accepts. Needs no key. */
1029
+ async aiProviders(options = {}) {
1030
+ const request = new Request(
1031
+ "GET",
1032
+ "/config/ai-providers",
1033
+ void 0,
1034
+ void 0,
1035
+ void 0,
1036
+ false
1037
+ );
1038
+ const body = record(await this.send(request, options));
1039
+ return Array.isArray(body.providers) ? body.providers : [];
1040
+ }
1041
+ /**
1042
+ * Whether the API is up. `deep` also checks the dependencies it needs to
1043
+ * serve scrapes. A degraded report is returned, not thrown: read `ok`.
1044
+ * Needs no key.
1045
+ */
1046
+ async health(options = {}) {
1047
+ const body = record(await this.send(healthRequest(options.deep ?? false), options));
1048
+ return { ...body, ok: body.status === "ok" };
1049
+ }
1050
+ /**
1051
+ * Which anti-bot protection sits in front of each URL. Nothing is scraped
1052
+ * and nothing is charged; invalid URLs are skipped.
1053
+ */
1054
+ async checkProtection(urls, options = {}) {
1055
+ return list(
1056
+ await this.send(new Request("POST", "/filter/check", void 0, [...urls]), options)
1057
+ );
1058
+ }
1016
1059
  /** Countries, regions and cities a proxy type can exit from. */
1017
1060
  async proxyLocations(options = {}) {
1018
1061
  const params = options.proxyType ? { proxy_type: options.proxyType } : void 0;
@@ -1028,13 +1071,22 @@ var DataFuel = class {
1028
1071
  await this.send(new Request("GET", "/config/proxy/asn", params), options)
1029
1072
  );
1030
1073
  }
1031
- /** Remaining credits. */
1074
+ /** Remaining credits: plan and pay-as-you-go together. */
1032
1075
  async balance(options = {}) {
1033
1076
  const body = await this.send(new Request("GET", "/users/@me/balance"), options);
1034
1077
  return intField(body, "balance");
1035
1078
  }
1079
+ /** Remaining credits by pool: plan credits (spent first) and pay-as-you-go credits. */
1080
+ async balanceSplit(options = {}) {
1081
+ const body = await this.send(new Request("GET", "/users/@me/balance"), options);
1082
+ return {
1083
+ balance: intField(body, "balance"),
1084
+ plan_balance: intField(body, "plan_balance"),
1085
+ payg_balance: intField(body, "payg_balance")
1086
+ };
1087
+ }
1036
1088
  /**
1037
- * Credit movements, newest first: purchases, usage, refunds, expiry.
1089
+ * Credit movements, newest first: plan assignments, credit pack purchases, usage, refunds, expiry.
1038
1090
  * `sums` totals each operation over the whole range, not just this page.
1039
1091
  */
1040
1092
  async transactions(options = {}) {
@@ -1075,7 +1127,7 @@ var DataFuel = class {
1075
1127
  function envApiKey() {
1076
1128
  return typeof process !== "undefined" ? process.env?.DATAFUEL_API_KEY : void 0;
1077
1129
  }
1078
- async function parse(response) {
1130
+ async function parse(response, degradedOk = false) {
1079
1131
  const text = await response.text();
1080
1132
  let body;
1081
1133
  if (text.length > 0) {
@@ -1085,6 +1137,7 @@ async function parse(response) {
1085
1137
  body = text;
1086
1138
  }
1087
1139
  }
1140
+ if (degradedOk && response.status === 503 && isHealthReport(body)) return body;
1088
1141
  if (!response.ok) {
1089
1142
  const header = response.headers.get("Retry-After");
1090
1143
  const retryAfter = header !== null && !Number.isNaN(Number(header)) ? Number(header) : 0;
@@ -1092,6 +1145,9 @@ async function parse(response) {
1092
1145
  }
1093
1146
  return body;
1094
1147
  }
1148
+ function isHealthReport(body) {
1149
+ return body !== null && typeof body === "object" && "status" in body;
1150
+ }
1095
1151
  function isAbort(error) {
1096
1152
  return error instanceof Error && (error.name === "AbortError" || error.name === "TimeoutError");
1097
1153
  }
@@ -1149,6 +1205,7 @@ function intField(body, name) {
1149
1205
  // Annotate the CommonJS export names for ESM import in node:
1150
1206
  0 && (module.exports = {
1151
1207
  APIError,
1208
+ AlreadyExists,
1152
1209
  Blocked,
1153
1210
  Capabilities,
1154
1211
  CrawlPage,