webcrawlerapi-langchain 0.1.1__tar.gz → 0.2.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: webcrawlerapi-langchain
3
- Version: 0.1.1
3
+ Version: 0.2.0
4
4
  Summary: LangChain integration for WebCrawlerAPI
5
5
  Home-page: https://github.com/webcrawlerapi/webcrawlerapi-langchain
6
6
  Author: WebCrawlerAPI
@@ -8,10 +8,10 @@ Author-email: support@webcrawlerapi.com
8
8
  Classifier: Programming Language :: Python :: 3
9
9
  Classifier: License :: OSI Approved :: MIT License
10
10
  Classifier: Operating System :: OS Independent
11
- Requires-Python: >=3.8
11
+ Requires-Python: >=3.10
12
12
  Description-Content-Type: text/markdown
13
- Requires-Dist: webcrawlerapi>=1.0.8
14
- Requires-Dist: langchain-core>=0.1.0
13
+ Requires-Dist: webcrawlerapi>=2.1.5
14
+ Requires-Dist: langchain-core>=0.3.0
15
15
  Dynamic: author
16
16
  Dynamic: author-email
17
17
  Dynamic: classifier
@@ -26,7 +26,7 @@ Dynamic: summary
26
26
 
27
27
  [WebcrawlerAPI](https://webcrawlerapi.com/) - is a website to LLM data API. It allows to convert websites and webpages markdown or cleaned content.
28
28
 
29
- **No subscription required**.
29
+ **Pay as you go (No subscription required).**
30
30
 
31
31
  This package provides LangChain integration for [WebCrawlerAPI](https://webcrawlerapi.com/), allowing you to easily use web crawling capabilities with [LangChain](https://www.langchain.com/) document processing pipeline.
32
32
 
@@ -40,6 +40,8 @@ pip install webcrawlerapi-langchain
40
40
 
41
41
  ## Usage
42
42
 
43
+ Requires Python 3.10 or newer. `api_key` and `base_url` fall back to the `WEBCRAWLERAPI_API_KEY` and `WEBCRAWLERAPI_BASE_URL` environment variables when not passed.
44
+
43
45
  ### Basic Loading
44
46
  ```python
45
47
  from webcrawlerapi_langchain import WebCrawlerAPILoader
@@ -48,7 +50,7 @@ from webcrawlerapi_langchain import WebCrawlerAPILoader
48
50
  loader = WebCrawlerAPILoader(
49
51
  url="https://example.com",
50
52
  api_key="your-api-key",
51
- scrape_type="markdown",
53
+ output_format="markdown",
52
54
  items_limit=10
53
55
  )
54
56
 
@@ -68,6 +70,7 @@ documents = await loader.aload()
68
70
  ```
69
71
 
70
72
  ### Lazy Loading
73
+ Yields each page as soon as it finishes crawling:
71
74
  ```python
72
75
  # Lazy loading
73
76
  for doc in loader.lazy_load():
@@ -81,16 +84,37 @@ async for doc in loader.alazy_load():
81
84
  print(doc.page_content[:100])
82
85
  ```
83
86
 
87
+ ### Error Handling
88
+ Crawling and API failures raise `WebCrawlerAPILoaderError`, which is importable from `webcrawlerapi_langchain`. Invalid parameters raise `ValueError`.
89
+ ```python
90
+ from webcrawlerapi_langchain import WebCrawlerAPILoaderError
91
+
92
+ try:
93
+ documents = loader.load()
94
+ except WebCrawlerAPILoaderError as e:
95
+ print(f"Crawl failed: {e}")
96
+ ```
97
+
84
98
  ## Configuration
85
99
 
86
100
  The loader accepts the following parameters:
87
101
 
88
- - `url`: The URL to crawl
89
- - `api_key`: Your WebCrawlerAPI API key
90
- - `scrape_type`: Type of scraping (html, cleaned, markdown)
91
- - `items_limit`: Maximum number of pages to crawl
102
+ - `url`: The URL to crawl (keyword arguments follow it)
103
+ - `api_key`: Your WebCrawlerAPI API key. Defaults to `WEBCRAWLERAPI_API_KEY`
104
+ - `base_url`: API base URL. Defaults to `WEBCRAWLERAPI_BASE_URL`
105
+ - `output_format`: Content format of each Document: `markdown` (default), `cleaned`, or `html`
106
+ - `scrape_type`: Deprecated, use `output_format` instead
107
+ - `items_limit`: Maximum number of pages to crawl (default `10`)
92
108
  - `whitelist_regexp`: Regex pattern for URL whitelist
93
109
  - `blacklist_regexp`: Regex pattern for URL blacklist
110
+ - `main_content_only`: Extract only the main content of each page (default `False`)
111
+ - `max_depth`: Maximum crawl depth. `0` is the seed URL only, `1` is the seed URL plus direct links
112
+ - `max_age`: Max age in seconds for cached content. `0` always fetches fresh content
113
+ - `respect_robots_txt`: Whether to respect robots.txt (default `False`)
114
+ - `keep_query_params`: Keep URL query params when deduplicating links
115
+ - `max_polls`: Maximum number of job status checks before giving up (default `100`)
116
+
117
+ Document metadata keys: `url`, `title`, `status_code`, `created_at`, `referred_url`, `depth`, `cost`, `job_id`, `item_id`.
94
118
 
95
119
  ### Links
96
120
  - [WebCrawlerAPI Python SDK](https://github.com/WebCrawlerAPI/webcrawlerapi-python-sdk)
@@ -2,7 +2,7 @@
2
2
 
3
3
  [WebcrawlerAPI](https://webcrawlerapi.com/) - is a website to LLM data API. It allows to convert websites and webpages markdown or cleaned content.
4
4
 
5
- **No subscription required**.
5
+ **Pay as you go (No subscription required).**
6
6
 
7
7
  This package provides LangChain integration for [WebCrawlerAPI](https://webcrawlerapi.com/), allowing you to easily use web crawling capabilities with [LangChain](https://www.langchain.com/) document processing pipeline.
8
8
 
@@ -16,6 +16,8 @@ pip install webcrawlerapi-langchain
16
16
 
17
17
  ## Usage
18
18
 
19
+ Requires Python 3.10 or newer. `api_key` and `base_url` fall back to the `WEBCRAWLERAPI_API_KEY` and `WEBCRAWLERAPI_BASE_URL` environment variables when not passed.
20
+
19
21
  ### Basic Loading
20
22
  ```python
21
23
  from webcrawlerapi_langchain import WebCrawlerAPILoader
@@ -24,7 +26,7 @@ from webcrawlerapi_langchain import WebCrawlerAPILoader
24
26
  loader = WebCrawlerAPILoader(
25
27
  url="https://example.com",
26
28
  api_key="your-api-key",
27
- scrape_type="markdown",
29
+ output_format="markdown",
28
30
  items_limit=10
29
31
  )
30
32
 
@@ -44,6 +46,7 @@ documents = await loader.aload()
44
46
  ```
45
47
 
46
48
  ### Lazy Loading
49
+ Yields each page as soon as it finishes crawling:
47
50
  ```python
48
51
  # Lazy loading
49
52
  for doc in loader.lazy_load():
@@ -57,16 +60,37 @@ async for doc in loader.alazy_load():
57
60
  print(doc.page_content[:100])
58
61
  ```
59
62
 
63
+ ### Error Handling
64
+ Crawling and API failures raise `WebCrawlerAPILoaderError`, which is importable from `webcrawlerapi_langchain`. Invalid parameters raise `ValueError`.
65
+ ```python
66
+ from webcrawlerapi_langchain import WebCrawlerAPILoaderError
67
+
68
+ try:
69
+ documents = loader.load()
70
+ except WebCrawlerAPILoaderError as e:
71
+ print(f"Crawl failed: {e}")
72
+ ```
73
+
60
74
  ## Configuration
61
75
 
62
76
  The loader accepts the following parameters:
63
77
 
64
- - `url`: The URL to crawl
65
- - `api_key`: Your WebCrawlerAPI API key
66
- - `scrape_type`: Type of scraping (html, cleaned, markdown)
67
- - `items_limit`: Maximum number of pages to crawl
78
+ - `url`: The URL to crawl (keyword arguments follow it)
79
+ - `api_key`: Your WebCrawlerAPI API key. Defaults to `WEBCRAWLERAPI_API_KEY`
80
+ - `base_url`: API base URL. Defaults to `WEBCRAWLERAPI_BASE_URL`
81
+ - `output_format`: Content format of each Document: `markdown` (default), `cleaned`, or `html`
82
+ - `scrape_type`: Deprecated, use `output_format` instead
83
+ - `items_limit`: Maximum number of pages to crawl (default `10`)
68
84
  - `whitelist_regexp`: Regex pattern for URL whitelist
69
85
  - `blacklist_regexp`: Regex pattern for URL blacklist
86
+ - `main_content_only`: Extract only the main content of each page (default `False`)
87
+ - `max_depth`: Maximum crawl depth. `0` is the seed URL only, `1` is the seed URL plus direct links
88
+ - `max_age`: Max age in seconds for cached content. `0` always fetches fresh content
89
+ - `respect_robots_txt`: Whether to respect robots.txt (default `False`)
90
+ - `keep_query_params`: Keep URL query params when deduplicating links
91
+ - `max_polls`: Maximum number of job status checks before giving up (default `100`)
92
+
93
+ Document metadata keys: `url`, `title`, `status_code`, `created_at`, `referred_url`, `depth`, `cost`, `job_id`, `item_id`.
70
94
 
71
95
  ### Links
72
96
  - [WebCrawlerAPI Python SDK](https://github.com/WebCrawlerAPI/webcrawlerapi-python-sdk)
@@ -2,11 +2,11 @@ from setuptools import setup, find_packages
2
2
 
3
3
  setup(
4
4
  name="webcrawlerapi-langchain",
5
- version="0.1.1",
5
+ version="0.2.0",
6
6
  packages=find_packages(),
7
7
  install_requires=[
8
- "webcrawlerapi>=1.0.8",
9
- "langchain-core>=0.1.0",
8
+ "webcrawlerapi>=2.1.5",
9
+ "langchain-core>=0.3.0",
10
10
  ],
11
11
  author="WebCrawlerAPI",
12
12
  author_email="support@webcrawlerapi.com",
@@ -19,5 +19,5 @@ setup(
19
19
  "License :: OSI Approved :: MIT License",
20
20
  "Operating System :: OS Independent",
21
21
  ],
22
- python_requires=">=3.8",
22
+ python_requires=">=3.10",
23
23
  )
@@ -0,0 +1,4 @@
1
+ from .loader import WebCrawlerAPILoader, WebCrawlerAPILoaderError
2
+
3
+ __version__ = "0.2.0"
4
+ __all__ = ["WebCrawlerAPILoader", "WebCrawlerAPILoaderError"]
@@ -1,7 +1,9 @@
1
- from typing import Dict, Iterator, List, Optional, AsyncIterator, Any, Literal
2
- import os
1
+ from typing import Iterator, List, Optional, AsyncIterator, Any, Literal
2
+ import asyncio
3
3
  import json
4
4
  import logging
5
+ import time
6
+ import warnings
5
7
  from langchain_core.documents import Document
6
8
  from langchain_core.utils import get_from_env
7
9
  from langchain_core.document_loaders import BaseLoader
@@ -10,10 +12,15 @@ from webcrawlerapi import WebCrawlerAPI
10
12
  # Configure logging
11
13
  logger = logging.getLogger(__name__)
12
14
 
15
+ OutputFormat = Literal["markdown", "cleaned", "html"]
16
+ ALLOWED_OUTPUT_FORMATS = ("markdown", "cleaned", "html")
17
+
18
+
13
19
  class WebCrawlerAPILoaderError(Exception):
14
20
  """Custom exception class for WebCrawlerAPILoader errors."""
15
21
  pass
16
22
 
23
+
17
24
  class WebCrawlerAPILoader(BaseLoader):
18
25
  """WebCrawlerAPI document loader integration.
19
26
 
@@ -31,12 +38,16 @@ class WebCrawlerAPILoader(BaseLoader):
31
38
  *,
32
39
  api_key: Optional[str] = None,
33
40
  base_url: Optional[str] = None,
34
- version: str = "v1",
35
- scrape_type: Literal["html", "cleaned", "markdown"] = "markdown",
41
+ output_format: Optional[OutputFormat] = None,
42
+ scrape_type: Optional[OutputFormat] = None,
36
43
  items_limit: int = 10,
37
- allow_subdomains: bool = False,
38
44
  whitelist_regexp: Optional[str] = None,
39
45
  blacklist_regexp: Optional[str] = None,
46
+ main_content_only: bool = False,
47
+ max_depth: Optional[int] = None,
48
+ max_age: Optional[int] = None,
49
+ respect_robots_txt: bool = False,
50
+ keep_query_params: Optional[bool] = None,
40
51
  max_polls: int = 100
41
52
  ):
42
53
  """Initialize the WebCrawlerAPI document loader.
@@ -45,13 +56,17 @@ class WebCrawlerAPILoader(BaseLoader):
45
56
  url: The URL to crawl
46
57
  api_key: Your WebCrawlerAPI API key. If not provided, will try to get from WEBCRAWLERAPI_API_KEY env var
47
58
  base_url: The base URL of the API. If not provided, will try to get from WEBCRAWLERAPI_BASE_URL env var
48
- version: API version to use (optional)
49
- scrape_type: Type of scraping (html, cleaned, markdown)
59
+ output_format: Content format of each Document (markdown, cleaned, html). Defaults to markdown
60
+ scrape_type: Deprecated. Use output_format instead
50
61
  items_limit: Maximum number of pages to crawl
51
- allow_subdomains: Whether to crawl subdomains
52
62
  whitelist_regexp: Regex pattern for URL whitelist
53
63
  blacklist_regexp: Regex pattern for URL blacklist
54
- max_polls: Maximum number of status checks before returning
64
+ main_content_only: Extract only the main content of each page
65
+ max_depth: Maximum crawl depth (0 = seed only, 1 = seed + direct links)
66
+ max_age: Max age in seconds for cached content. 0 = always fresh
67
+ respect_robots_txt: Whether to respect robots.txt
68
+ keep_query_params: Keep URL query params when deduplicating links
69
+ max_polls: Maximum number of status checks before giving up
55
70
  """
56
71
  if not url:
57
72
  raise ValueError("URL must be provided")
@@ -65,10 +80,17 @@ class WebCrawlerAPILoader(BaseLoader):
65
80
  if max_polls < 1:
66
81
  raise ValueError("max_polls must be greater than 0")
67
82
 
68
- if scrape_type not in ("html", "cleaned", "markdown"):
83
+ if scrape_type is not None:
84
+ warnings.warn(
85
+ "scrape_type is deprecated, use output_format instead",
86
+ DeprecationWarning,
87
+ stacklevel=2,
88
+ )
89
+ output_format = output_format or scrape_type or "markdown"
90
+ if output_format not in ALLOWED_OUTPUT_FORMATS:
69
91
  raise ValueError(
70
- f"Invalid scrape_type '{scrape_type}'. "
71
- "Allowed: 'html', 'cleaned', 'markdown'."
92
+ f"Invalid output_format '{output_format}'. "
93
+ "Allowed: 'markdown', 'cleaned', 'html'."
72
94
  )
73
95
 
74
96
  # Get API key and base URL from env vars if not provided
@@ -77,25 +99,46 @@ class WebCrawlerAPILoader(BaseLoader):
77
99
  raise ValueError("API key must be provided either through api_key parameter or WEBCRAWLERAPI_API_KEY environment variable")
78
100
 
79
101
  self.base_url = base_url or get_from_env(
80
- "base_url",
81
- "WEBCRAWLERAPI_BASE_URL",
102
+ "base_url",
103
+ "WEBCRAWLERAPI_BASE_URL",
82
104
  default="https://api.webcrawlerapi.com"
83
105
  )
84
106
 
85
107
  self.url = url
86
- self.version = version
87
- self.scrape_type = scrape_type
108
+ self.output_format = output_format
88
109
  self.items_limit = items_limit
89
- self.allow_subdomains = allow_subdomains
90
110
  self.whitelist_regexp = whitelist_regexp
91
111
  self.blacklist_regexp = blacklist_regexp
112
+ self.main_content_only = main_content_only
113
+ self.max_depth = max_depth
114
+ self.max_age = max_age
115
+ self.respect_robots_txt = respect_robots_txt
116
+ self.keep_query_params = keep_query_params
92
117
  self.max_polls = max_polls
93
118
 
94
- self.client = WebCrawlerAPI(
95
- api_key=self.api_key,
96
- base_url=self.base_url,
97
- version=version
98
- )
119
+ self.client = WebCrawlerAPI(api_key=self.api_key, base_url=self.base_url)
120
+
121
+ def _crawl_params(self) -> dict:
122
+ return {
123
+ "url": self.url,
124
+ "output_formats": [self.output_format],
125
+ "items_limit": self.items_limit,
126
+ "whitelist_regexp": self.whitelist_regexp,
127
+ "blacklist_regexp": self.blacklist_regexp,
128
+ "main_content_only": self.main_content_only,
129
+ "max_depth": self.max_depth,
130
+ "max_age": self.max_age,
131
+ "respect_robots_txt": self.respect_robots_txt,
132
+ "keep_query_params": self.keep_query_params,
133
+ }
134
+
135
+ def _fetch_item_content(self, item: Any) -> Optional[str]:
136
+ """Fetch item content in the configured output format."""
137
+ if self.output_format == "markdown":
138
+ return item.get_markdown()
139
+ if self.output_format == "cleaned":
140
+ return item.get_cleaned()
141
+ return item.get_html()
99
142
 
100
143
  def _create_document(self, item: Any) -> Optional[Document]:
101
144
  """Create a Document from a job item if it's valid.
@@ -104,27 +147,46 @@ class WebCrawlerAPILoader(BaseLoader):
104
147
  item: Job item from the API response
105
148
 
106
149
  Returns:
107
- Document if item is valid and has content, None otherwise
150
+ Document if item is done and has content, None otherwise
108
151
  """
109
152
  try:
110
- if not (item.status == "done" and item.content):
153
+ if item.status != "done":
154
+ return None
155
+
156
+ content = self._fetch_item_content(item)
157
+ if not content:
111
158
  return None
112
159
 
113
- doc = Document(
114
- page_content=item.content,
160
+ return Document(
161
+ page_content=content,
115
162
  metadata={
116
163
  "url": item.original_url,
117
164
  "title": item.title,
118
165
  "status_code": item.page_status_code,
119
166
  "created_at": item.created_at,
120
167
  "referred_url": item.referred_url,
121
- "cost": item.cost
168
+ "depth": item.depth,
169
+ "cost": item.cost,
170
+ "job_id": item.job_id,
171
+ "item_id": item.id,
122
172
  }
123
173
  )
124
- return doc
125
174
  except AttributeError as e:
126
175
  logger.error(f"Failed to create document from item: {e}")
127
176
  raise WebCrawlerAPILoaderError(f"Invalid job item format: {str(e)}") from e
177
+ except Exception as e:
178
+ logger.error(f"Failed to fetch content for {getattr(item, 'original_url', '?')}: {e}")
179
+ raise WebCrawlerAPILoaderError(f"Failed to fetch item content: {str(e)}") from e
180
+
181
+ @staticmethod
182
+ def _job_error_message(job: Any) -> str:
183
+ """Build an error message from failed job items, as the job itself carries no error field."""
184
+ errors = [
185
+ f"{item.original_url}: {item.error_code or 'error'} {item.last_error or ''}".strip()
186
+ for item in job.job_items
187
+ if item.status == "error"
188
+ ]
189
+ return "; ".join(errors) if errors else "Unknown error"
128
190
 
129
191
  def load(self) -> List[Document]:
130
192
  """Load data into Document objects.
@@ -134,21 +196,10 @@ class WebCrawlerAPILoader(BaseLoader):
134
196
 
135
197
  Raises:
136
198
  WebCrawlerAPILoaderError: If there are issues with the crawling job or API response
137
- ValueError: If the response format is invalid
138
- json.JSONDecodeError: If the API response contains invalid JSON
139
199
  """
140
200
  logger.info(f"Starting crawl for URL: {self.url}")
141
201
  try:
142
- job = self.client.crawl(
143
- url=self.url,
144
- scrape_type=self.scrape_type,
145
- items_limit=self.items_limit,
146
- allow_subdomains=self.allow_subdomains,
147
- whitelist_regexp=self.whitelist_regexp,
148
- blacklist_regexp=self.blacklist_regexp,
149
- max_polls=self.max_polls
150
- )
151
-
202
+ job = self.client.crawl(**self._crawl_params(), max_polls=self.max_polls)
152
203
  except json.JSONDecodeError as e:
153
204
  logger.error(f"JSON decode error in crawl response: {e}")
154
205
  raise WebCrawlerAPILoaderError(
@@ -159,22 +210,21 @@ class WebCrawlerAPILoader(BaseLoader):
159
210
  raise WebCrawlerAPILoaderError(f"API request failed: {str(e)}") from e
160
211
 
161
212
  if job.status == "error":
162
- error_msg = getattr(job, 'error', 'Unknown error')
213
+ error_msg = self._job_error_message(job)
163
214
  logger.error(f"Crawl job failed with error: {error_msg}")
164
215
  raise WebCrawlerAPILoaderError(f"Crawling job failed: {error_msg}")
165
216
 
166
- documents = []
217
+ if not job.is_terminal:
218
+ raise WebCrawlerAPILoaderError(
219
+ f"Maximum number of polls ({self.max_polls}) reached without job completion"
220
+ )
221
+
167
222
  logger.info(f"Processing {len(job.job_items)} job items")
223
+ documents = []
168
224
  for item in job.job_items:
169
- try:
170
- doc = self._create_document(item)
171
- if doc:
172
- documents.append(doc)
173
- except AttributeError as e:
174
- logger.error(f"Failed to process job item: {e}")
175
- raise WebCrawlerAPILoaderError(
176
- f"Invalid job item format - missing required field: {str(e)}"
177
- ) from e
225
+ doc = self._create_document(item)
226
+ if doc:
227
+ documents.append(doc)
178
228
 
179
229
  logger.info(f"Successfully created {len(documents)} documents")
180
230
  return documents
@@ -183,24 +233,14 @@ class WebCrawlerAPILoader(BaseLoader):
183
233
  """A lazy loader for Documents.
184
234
 
185
235
  Yields:
186
- Document objects one at a time as they are crawled.
236
+ Document objects one at a time as pages finish crawling.
187
237
 
188
238
  Raises:
189
239
  WebCrawlerAPILoaderError: If there are issues with the crawling job or API response
190
- ValueError: If the response format is invalid
191
- json.JSONDecodeError: If the API response contains invalid JSON
192
240
  """
193
241
  logger.info(f"Starting async crawl for URL: {self.url}")
194
242
  try:
195
- response = self.client.crawl_async(
196
- url=self.url,
197
- scrape_type=self.scrape_type,
198
- items_limit=self.items_limit,
199
- allow_subdomains=self.allow_subdomains,
200
- whitelist_regexp=self.whitelist_regexp,
201
- blacklist_regexp=self.blacklist_regexp
202
- )
203
-
243
+ response = self.client.crawl_async(**self._crawl_params())
204
244
  except json.JSONDecodeError as e:
205
245
  logger.error(f"JSON decode error in async crawl response: {e}")
206
246
  raise WebCrawlerAPILoaderError(
@@ -211,14 +251,12 @@ class WebCrawlerAPILoader(BaseLoader):
211
251
  raise WebCrawlerAPILoaderError(f"API request failed: {str(e)}") from e
212
252
 
213
253
  job_id = response.id
214
- polls = 0
215
254
  processed_items = set()
216
255
  logger.info(f"Starting to poll job {job_id}")
217
256
 
218
- while polls < self.max_polls:
257
+ for _ in range(self.max_polls):
219
258
  try:
220
259
  job = self.client.get_job(job_id)
221
-
222
260
  except json.JSONDecodeError as e:
223
261
  logger.error(f"JSON decode error in job status response: {e}")
224
262
  raise WebCrawlerAPILoaderError(
@@ -229,42 +267,34 @@ class WebCrawlerAPILoader(BaseLoader):
229
267
  raise WebCrawlerAPILoaderError(f"Failed to fetch job status: {str(e)}") from e
230
268
 
231
269
  if job.status == "error":
232
- error_msg = getattr(job, 'error', 'Unknown error')
270
+ error_msg = self._job_error_message(job)
233
271
  logger.error(f"Job failed with error: {error_msg}")
234
272
  raise WebCrawlerAPILoaderError(f"Crawling job failed: {error_msg}")
235
273
 
274
+ # Yield pages as soon as each item is done, without waiting for the whole job
236
275
  for item in job.job_items:
237
- if item.id not in processed_items:
238
- try:
239
- doc = self._create_document(item)
240
- if doc:
241
- processed_items.add(item.id)
242
- yield doc
243
- except AttributeError as e:
244
- logger.error(f"Failed to process job item: {e}")
245
- raise WebCrawlerAPILoaderError(
246
- f"Invalid job item format - missing required field: {str(e)}"
247
- ) from e
276
+ if item.id in processed_items or item.status not in ("done", "error"):
277
+ continue
278
+ processed_items.add(item.id)
279
+ doc = self._create_document(item)
280
+ if doc:
281
+ yield doc
248
282
 
249
283
  if job.is_terminal:
250
- logger.info("Job completed successfully")
251
- break
284
+ logger.info("Job completed")
285
+ return
252
286
 
253
287
  delay_seconds = (
254
288
  job.recommended_pull_delay_ms / 1000
255
289
  if job.recommended_pull_delay_ms
256
290
  else self.client.DEFAULT_POLL_DELAY_SECONDS
257
291
  )
258
-
259
- import time
260
292
  time.sleep(delay_seconds)
261
- polls += 1
262
293
 
263
- if not job.is_terminal and polls >= self.max_polls:
264
- logger.error(f"Job timed out after {self.max_polls} polls")
265
- raise WebCrawlerAPILoaderError(
266
- f"Maximum number of polls ({self.max_polls}) reached without job completion"
267
- )
294
+ logger.error(f"Job timed out after {self.max_polls} polls")
295
+ raise WebCrawlerAPILoaderError(
296
+ f"Maximum number of polls ({self.max_polls}) reached without job completion"
297
+ )
268
298
 
269
299
  async def aload(self) -> List[Document]:
270
300
  """Asynchronously load data into Document objects.
@@ -273,21 +303,23 @@ class WebCrawlerAPILoader(BaseLoader):
273
303
  List of Document objects, one for each crawled page.
274
304
 
275
305
  Raises:
276
- RuntimeError: If the crawling job fails
306
+ WebCrawlerAPILoaderError: If the crawling job fails
277
307
  """
278
- import asyncio
279
308
  return await asyncio.to_thread(self.load)
280
309
 
281
310
  async def alazy_load(self) -> AsyncIterator[Document]:
282
311
  """An async lazy loader for Documents.
283
312
 
284
313
  Yields:
285
- Document objects one at a time as they are crawled.
314
+ Document objects one at a time as pages finish crawling.
286
315
 
287
316
  Raises:
288
- RuntimeError: If the crawling job fails
317
+ WebCrawlerAPILoaderError: If the crawling job fails
289
318
  """
290
- import asyncio
291
- for doc in self.lazy_load():
319
+ iterator = self.lazy_load()
320
+ sentinel = object()
321
+ while True:
322
+ doc = await asyncio.to_thread(next, iterator, sentinel)
323
+ if doc is sentinel:
324
+ break
292
325
  yield doc
293
- await asyncio.sleep(0)
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: webcrawlerapi-langchain
3
- Version: 0.1.1
3
+ Version: 0.2.0
4
4
  Summary: LangChain integration for WebCrawlerAPI
5
5
  Home-page: https://github.com/webcrawlerapi/webcrawlerapi-langchain
6
6
  Author: WebCrawlerAPI
@@ -8,10 +8,10 @@ Author-email: support@webcrawlerapi.com
8
8
  Classifier: Programming Language :: Python :: 3
9
9
  Classifier: License :: OSI Approved :: MIT License
10
10
  Classifier: Operating System :: OS Independent
11
- Requires-Python: >=3.8
11
+ Requires-Python: >=3.10
12
12
  Description-Content-Type: text/markdown
13
- Requires-Dist: webcrawlerapi>=1.0.8
14
- Requires-Dist: langchain-core>=0.1.0
13
+ Requires-Dist: webcrawlerapi>=2.1.5
14
+ Requires-Dist: langchain-core>=0.3.0
15
15
  Dynamic: author
16
16
  Dynamic: author-email
17
17
  Dynamic: classifier
@@ -26,7 +26,7 @@ Dynamic: summary
26
26
 
27
27
  [WebcrawlerAPI](https://webcrawlerapi.com/) - is a website to LLM data API. It allows to convert websites and webpages markdown or cleaned content.
28
28
 
29
- **No subscription required**.
29
+ **Pay as you go (No subscription required).**
30
30
 
31
31
  This package provides LangChain integration for [WebCrawlerAPI](https://webcrawlerapi.com/), allowing you to easily use web crawling capabilities with [LangChain](https://www.langchain.com/) document processing pipeline.
32
32
 
@@ -40,6 +40,8 @@ pip install webcrawlerapi-langchain
40
40
 
41
41
  ## Usage
42
42
 
43
+ Requires Python 3.10 or newer. `api_key` and `base_url` fall back to the `WEBCRAWLERAPI_API_KEY` and `WEBCRAWLERAPI_BASE_URL` environment variables when not passed.
44
+
43
45
  ### Basic Loading
44
46
  ```python
45
47
  from webcrawlerapi_langchain import WebCrawlerAPILoader
@@ -48,7 +50,7 @@ from webcrawlerapi_langchain import WebCrawlerAPILoader
48
50
  loader = WebCrawlerAPILoader(
49
51
  url="https://example.com",
50
52
  api_key="your-api-key",
51
- scrape_type="markdown",
53
+ output_format="markdown",
52
54
  items_limit=10
53
55
  )
54
56
 
@@ -68,6 +70,7 @@ documents = await loader.aload()
68
70
  ```
69
71
 
70
72
  ### Lazy Loading
73
+ Yields each page as soon as it finishes crawling:
71
74
  ```python
72
75
  # Lazy loading
73
76
  for doc in loader.lazy_load():
@@ -81,16 +84,37 @@ async for doc in loader.alazy_load():
81
84
  print(doc.page_content[:100])
82
85
  ```
83
86
 
87
+ ### Error Handling
88
+ Crawling and API failures raise `WebCrawlerAPILoaderError`, which is importable from `webcrawlerapi_langchain`. Invalid parameters raise `ValueError`.
89
+ ```python
90
+ from webcrawlerapi_langchain import WebCrawlerAPILoaderError
91
+
92
+ try:
93
+ documents = loader.load()
94
+ except WebCrawlerAPILoaderError as e:
95
+ print(f"Crawl failed: {e}")
96
+ ```
97
+
84
98
  ## Configuration
85
99
 
86
100
  The loader accepts the following parameters:
87
101
 
88
- - `url`: The URL to crawl
89
- - `api_key`: Your WebCrawlerAPI API key
90
- - `scrape_type`: Type of scraping (html, cleaned, markdown)
91
- - `items_limit`: Maximum number of pages to crawl
102
+ - `url`: The URL to crawl (keyword arguments follow it)
103
+ - `api_key`: Your WebCrawlerAPI API key. Defaults to `WEBCRAWLERAPI_API_KEY`
104
+ - `base_url`: API base URL. Defaults to `WEBCRAWLERAPI_BASE_URL`
105
+ - `output_format`: Content format of each Document: `markdown` (default), `cleaned`, or `html`
106
+ - `scrape_type`: Deprecated, use `output_format` instead
107
+ - `items_limit`: Maximum number of pages to crawl (default `10`)
92
108
  - `whitelist_regexp`: Regex pattern for URL whitelist
93
109
  - `blacklist_regexp`: Regex pattern for URL blacklist
110
+ - `main_content_only`: Extract only the main content of each page (default `False`)
111
+ - `max_depth`: Maximum crawl depth. `0` is the seed URL only, `1` is the seed URL plus direct links
112
+ - `max_age`: Max age in seconds for cached content. `0` always fetches fresh content
113
+ - `respect_robots_txt`: Whether to respect robots.txt (default `False`)
114
+ - `keep_query_params`: Keep URL query params when deduplicating links
115
+ - `max_polls`: Maximum number of job status checks before giving up (default `100`)
116
+
117
+ Document metadata keys: `url`, `title`, `status_code`, `created_at`, `referred_url`, `depth`, `cost`, `job_id`, `item_id`.
94
118
 
95
119
  ### Links
96
120
  - [WebCrawlerAPI Python SDK](https://github.com/WebCrawlerAPI/webcrawlerapi-python-sdk)
@@ -0,0 +1,2 @@
1
+ webcrawlerapi>=2.1.5
2
+ langchain-core>=0.3.0
@@ -1,4 +0,0 @@
1
- from .loader import WebCrawlerAPILoader
2
-
3
- __version__ = "0.1.0"
4
- __all__ = ["WebCrawlerAPILoader"]
@@ -1,2 +0,0 @@
1
- webcrawlerapi>=1.0.8
2
- langchain-core>=0.1.0