webcrawlerapi-langchain 0.1.0__tar.gz → 0.2.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: webcrawlerapi-langchain
3
- Version: 0.1.0
3
+ Version: 0.2.0
4
4
  Summary: LangChain integration for WebCrawlerAPI
5
5
  Home-page: https://github.com/webcrawlerapi/webcrawlerapi-langchain
6
6
  Author: WebCrawlerAPI
@@ -8,10 +8,10 @@ Author-email: support@webcrawlerapi.com
8
8
  Classifier: Programming Language :: Python :: 3
9
9
  Classifier: License :: OSI Approved :: MIT License
10
10
  Classifier: Operating System :: OS Independent
11
- Requires-Python: >=3.8
11
+ Requires-Python: >=3.10
12
12
  Description-Content-Type: text/markdown
13
- Requires-Dist: webcrawlerapi==1.0.6
14
- Requires-Dist: langchain-core>=0.1.0
13
+ Requires-Dist: webcrawlerapi>=2.1.5
14
+ Requires-Dist: langchain-core>=0.3.0
15
15
  Dynamic: author
16
16
  Dynamic: author-email
17
17
  Dynamic: classifier
@@ -26,7 +26,7 @@ Dynamic: summary
26
26
 
27
27
  [WebcrawlerAPI](https://webcrawlerapi.com/) - is a website to LLM data API. It allows to convert websites and webpages markdown or cleaned content.
28
28
 
29
- **No subscription required**.
29
+ **Pay as you go (No subscription required).**
30
30
 
31
31
  This package provides LangChain integration for [WebCrawlerAPI](https://webcrawlerapi.com/), allowing you to easily use web crawling capabilities with [LangChain](https://www.langchain.com/) document processing pipeline.
32
32
 
@@ -40,6 +40,8 @@ pip install webcrawlerapi-langchain
40
40
 
41
41
  ## Usage
42
42
 
43
+ Requires Python 3.10 or newer. `api_key` and `base_url` fall back to the `WEBCRAWLERAPI_API_KEY` and `WEBCRAWLERAPI_BASE_URL` environment variables when not passed.
44
+
43
45
  ### Basic Loading
44
46
  ```python
45
47
  from webcrawlerapi_langchain import WebCrawlerAPILoader
@@ -48,7 +50,7 @@ from webcrawlerapi_langchain import WebCrawlerAPILoader
48
50
  loader = WebCrawlerAPILoader(
49
51
  url="https://example.com",
50
52
  api_key="your-api-key",
51
- scrape_type="markdown",
53
+ output_format="markdown",
52
54
  items_limit=10
53
55
  )
54
56
 
@@ -68,6 +70,7 @@ documents = await loader.aload()
68
70
  ```
69
71
 
70
72
  ### Lazy Loading
73
+ Yields each page as soon as it finishes crawling:
71
74
  ```python
72
75
  # Lazy loading
73
76
  for doc in loader.lazy_load():
@@ -81,16 +84,37 @@ async for doc in loader.alazy_load():
81
84
  print(doc.page_content[:100])
82
85
  ```
83
86
 
87
+ ### Error Handling
88
+ Crawling and API failures raise `WebCrawlerAPILoaderError`, which is importable from `webcrawlerapi_langchain`. Invalid parameters raise `ValueError`.
89
+ ```python
90
+ from webcrawlerapi_langchain import WebCrawlerAPILoaderError
91
+
92
+ try:
93
+ documents = loader.load()
94
+ except WebCrawlerAPILoaderError as e:
95
+ print(f"Crawl failed: {e}")
96
+ ```
97
+
84
98
  ## Configuration
85
99
 
86
100
  The loader accepts the following parameters:
87
101
 
88
- - `url`: The URL to crawl
89
- - `api_key`: Your WebCrawlerAPI API key
90
- - `scrape_type`: Type of scraping (html, cleaned, markdown)
91
- - `items_limit`: Maximum number of pages to crawl
102
+ - `url`: The URL to crawl (keyword arguments follow it)
103
+ - `api_key`: Your WebCrawlerAPI API key. Defaults to `WEBCRAWLERAPI_API_KEY`
104
+ - `base_url`: API base URL. Defaults to `WEBCRAWLERAPI_BASE_URL`
105
+ - `output_format`: Content format of each Document: `markdown` (default), `cleaned`, or `html`
106
+ - `scrape_type`: Deprecated, use `output_format` instead
107
+ - `items_limit`: Maximum number of pages to crawl (default `10`)
92
108
  - `whitelist_regexp`: Regex pattern for URL whitelist
93
109
  - `blacklist_regexp`: Regex pattern for URL blacklist
110
+ - `main_content_only`: Extract only the main content of each page (default `False`)
111
+ - `max_depth`: Maximum crawl depth. `0` is the seed URL only, `1` is the seed URL plus direct links
112
+ - `max_age`: Max age in seconds for cached content. `0` always fetches fresh content
113
+ - `respect_robots_txt`: Whether to respect robots.txt (default `False`)
114
+ - `keep_query_params`: Keep URL query params when deduplicating links
115
+ - `max_polls`: Maximum number of job status checks before giving up (default `100`)
116
+
117
+ Document metadata keys: `url`, `title`, `status_code`, `created_at`, `referred_url`, `depth`, `cost`, `job_id`, `item_id`.
94
118
 
95
119
  ### Links
96
120
  - [WebCrawlerAPI Python SDK](https://github.com/WebCrawlerAPI/webcrawlerapi-python-sdk)
@@ -2,7 +2,7 @@
2
2
 
3
3
  [WebcrawlerAPI](https://webcrawlerapi.com/) - is a website to LLM data API. It allows to convert websites and webpages markdown or cleaned content.
4
4
 
5
- **No subscription required**.
5
+ **Pay as you go (No subscription required).**
6
6
 
7
7
  This package provides LangChain integration for [WebCrawlerAPI](https://webcrawlerapi.com/), allowing you to easily use web crawling capabilities with [LangChain](https://www.langchain.com/) document processing pipeline.
8
8
 
@@ -16,6 +16,8 @@ pip install webcrawlerapi-langchain
16
16
 
17
17
  ## Usage
18
18
 
19
+ Requires Python 3.10 or newer. `api_key` and `base_url` fall back to the `WEBCRAWLERAPI_API_KEY` and `WEBCRAWLERAPI_BASE_URL` environment variables when not passed.
20
+
19
21
  ### Basic Loading
20
22
  ```python
21
23
  from webcrawlerapi_langchain import WebCrawlerAPILoader
@@ -24,7 +26,7 @@ from webcrawlerapi_langchain import WebCrawlerAPILoader
24
26
  loader = WebCrawlerAPILoader(
25
27
  url="https://example.com",
26
28
  api_key="your-api-key",
27
- scrape_type="markdown",
29
+ output_format="markdown",
28
30
  items_limit=10
29
31
  )
30
32
 
@@ -44,6 +46,7 @@ documents = await loader.aload()
44
46
  ```
45
47
 
46
48
  ### Lazy Loading
49
+ Yields each page as soon as it finishes crawling:
47
50
  ```python
48
51
  # Lazy loading
49
52
  for doc in loader.lazy_load():
@@ -57,16 +60,37 @@ async for doc in loader.alazy_load():
57
60
  print(doc.page_content[:100])
58
61
  ```
59
62
 
63
+ ### Error Handling
64
+ Crawling and API failures raise `WebCrawlerAPILoaderError`, which is importable from `webcrawlerapi_langchain`. Invalid parameters raise `ValueError`.
65
+ ```python
66
+ from webcrawlerapi_langchain import WebCrawlerAPILoaderError
67
+
68
+ try:
69
+ documents = loader.load()
70
+ except WebCrawlerAPILoaderError as e:
71
+ print(f"Crawl failed: {e}")
72
+ ```
73
+
60
74
  ## Configuration
61
75
 
62
76
  The loader accepts the following parameters:
63
77
 
64
- - `url`: The URL to crawl
65
- - `api_key`: Your WebCrawlerAPI API key
66
- - `scrape_type`: Type of scraping (html, cleaned, markdown)
67
- - `items_limit`: Maximum number of pages to crawl
78
+ - `url`: The URL to crawl (keyword arguments follow it)
79
+ - `api_key`: Your WebCrawlerAPI API key. Defaults to `WEBCRAWLERAPI_API_KEY`
80
+ - `base_url`: API base URL. Defaults to `WEBCRAWLERAPI_BASE_URL`
81
+ - `output_format`: Content format of each Document: `markdown` (default), `cleaned`, or `html`
82
+ - `scrape_type`: Deprecated, use `output_format` instead
83
+ - `items_limit`: Maximum number of pages to crawl (default `10`)
68
84
  - `whitelist_regexp`: Regex pattern for URL whitelist
69
85
  - `blacklist_regexp`: Regex pattern for URL blacklist
86
+ - `main_content_only`: Extract only the main content of each page (default `False`)
87
+ - `max_depth`: Maximum crawl depth. `0` is the seed URL only, `1` is the seed URL plus direct links
88
+ - `max_age`: Max age in seconds for cached content. `0` always fetches fresh content
89
+ - `respect_robots_txt`: Whether to respect robots.txt (default `False`)
90
+ - `keep_query_params`: Keep URL query params when deduplicating links
91
+ - `max_polls`: Maximum number of job status checks before giving up (default `100`)
92
+
93
+ Document metadata keys: `url`, `title`, `status_code`, `created_at`, `referred_url`, `depth`, `cost`, `job_id`, `item_id`.
70
94
 
71
95
  ### Links
72
96
  - [WebCrawlerAPI Python SDK](https://github.com/WebCrawlerAPI/webcrawlerapi-python-sdk)
@@ -2,11 +2,11 @@ from setuptools import setup, find_packages
2
2
 
3
3
  setup(
4
4
  name="webcrawlerapi-langchain",
5
- version="0.1.0",
5
+ version="0.2.0",
6
6
  packages=find_packages(),
7
7
  install_requires=[
8
- "webcrawlerapi==1.0.6",
9
- "langchain-core>=0.1.0",
8
+ "webcrawlerapi>=2.1.5",
9
+ "langchain-core>=0.3.0",
10
10
  ],
11
11
  author="WebCrawlerAPI",
12
12
  author_email="support@webcrawlerapi.com",
@@ -19,5 +19,5 @@ setup(
19
19
  "License :: OSI Approved :: MIT License",
20
20
  "Operating System :: OS Independent",
21
21
  ],
22
- python_requires=">=3.8",
22
+ python_requires=">=3.10",
23
23
  )
@@ -0,0 +1,4 @@
1
+ from .loader import WebCrawlerAPILoader, WebCrawlerAPILoaderError
2
+
3
+ __version__ = "0.2.0"
4
+ __all__ = ["WebCrawlerAPILoader", "WebCrawlerAPILoaderError"]
@@ -1,7 +1,9 @@
1
- from typing import Dict, Iterator, List, Optional, AsyncIterator, Any, Literal
2
- import os
1
+ from typing import Iterator, List, Optional, AsyncIterator, Any, Literal
2
+ import asyncio
3
3
  import json
4
4
  import logging
5
+ import time
6
+ import warnings
5
7
  from langchain_core.documents import Document
6
8
  from langchain_core.utils import get_from_env
7
9
  from langchain_core.document_loaders import BaseLoader
@@ -9,12 +11,16 @@ from webcrawlerapi import WebCrawlerAPI
9
11
 
10
12
  # Configure logging
11
13
  logger = logging.getLogger(__name__)
12
- logger.setLevel(logging.DEBUG)
14
+
15
+ OutputFormat = Literal["markdown", "cleaned", "html"]
16
+ ALLOWED_OUTPUT_FORMATS = ("markdown", "cleaned", "html")
17
+
13
18
 
14
19
  class WebCrawlerAPILoaderError(Exception):
15
20
  """Custom exception class for WebCrawlerAPILoader errors."""
16
21
  pass
17
22
 
23
+
18
24
  class WebCrawlerAPILoader(BaseLoader):
19
25
  """WebCrawlerAPI document loader integration.
20
26
 
@@ -32,12 +38,16 @@ class WebCrawlerAPILoader(BaseLoader):
32
38
  *,
33
39
  api_key: Optional[str] = None,
34
40
  base_url: Optional[str] = None,
35
- version: str = "v1",
36
- scrape_type: Literal["html", "cleaned", "markdown"] = "markdown",
41
+ output_format: Optional[OutputFormat] = None,
42
+ scrape_type: Optional[OutputFormat] = None,
37
43
  items_limit: int = 10,
38
- allow_subdomains: bool = False,
39
44
  whitelist_regexp: Optional[str] = None,
40
45
  blacklist_regexp: Optional[str] = None,
46
+ main_content_only: bool = False,
47
+ max_depth: Optional[int] = None,
48
+ max_age: Optional[int] = None,
49
+ respect_robots_txt: bool = False,
50
+ keep_query_params: Optional[bool] = None,
41
51
  max_polls: int = 100
42
52
  ):
43
53
  """Initialize the WebCrawlerAPI document loader.
@@ -46,13 +56,17 @@ class WebCrawlerAPILoader(BaseLoader):
46
56
  url: The URL to crawl
47
57
  api_key: Your WebCrawlerAPI API key. If not provided, will try to get from WEBCRAWLERAPI_API_KEY env var
48
58
  base_url: The base URL of the API. If not provided, will try to get from WEBCRAWLERAPI_BASE_URL env var
49
- version: API version to use (optional)
50
- scrape_type: Type of scraping (html, cleaned, markdown)
59
+ output_format: Content format of each Document (markdown, cleaned, html). Defaults to markdown
60
+ scrape_type: Deprecated. Use output_format instead
51
61
  items_limit: Maximum number of pages to crawl
52
- allow_subdomains: Whether to crawl subdomains
53
62
  whitelist_regexp: Regex pattern for URL whitelist
54
63
  blacklist_regexp: Regex pattern for URL blacklist
55
- max_polls: Maximum number of status checks before returning
64
+ main_content_only: Extract only the main content of each page
65
+ max_depth: Maximum crawl depth (0 = seed only, 1 = seed + direct links)
66
+ max_age: Max age in seconds for cached content. 0 = always fresh
67
+ respect_robots_txt: Whether to respect robots.txt
68
+ keep_query_params: Keep URL query params when deduplicating links
69
+ max_polls: Maximum number of status checks before giving up
56
70
  """
57
71
  if not url:
58
72
  raise ValueError("URL must be provided")
@@ -66,10 +80,17 @@ class WebCrawlerAPILoader(BaseLoader):
66
80
  if max_polls < 1:
67
81
  raise ValueError("max_polls must be greater than 0")
68
82
 
69
- if scrape_type not in ("html", "cleaned", "markdown"):
83
+ if scrape_type is not None:
84
+ warnings.warn(
85
+ "scrape_type is deprecated, use output_format instead",
86
+ DeprecationWarning,
87
+ stacklevel=2,
88
+ )
89
+ output_format = output_format or scrape_type or "markdown"
90
+ if output_format not in ALLOWED_OUTPUT_FORMATS:
70
91
  raise ValueError(
71
- f"Invalid scrape_type '{scrape_type}'. "
72
- "Allowed: 'html', 'cleaned', 'markdown'."
92
+ f"Invalid output_format '{output_format}'. "
93
+ "Allowed: 'markdown', 'cleaned', 'html'."
73
94
  )
74
95
 
75
96
  # Get API key and base URL from env vars if not provided
@@ -78,26 +99,46 @@ class WebCrawlerAPILoader(BaseLoader):
78
99
  raise ValueError("API key must be provided either through api_key parameter or WEBCRAWLERAPI_API_KEY environment variable")
79
100
 
80
101
  self.base_url = base_url or get_from_env(
81
- "base_url",
82
- "WEBCRAWLERAPI_BASE_URL",
102
+ "base_url",
103
+ "WEBCRAWLERAPI_BASE_URL",
83
104
  default="https://api.webcrawlerapi.com"
84
105
  )
85
106
 
86
107
  self.url = url
87
- self.version = version
88
- self.scrape_type = scrape_type
108
+ self.output_format = output_format
89
109
  self.items_limit = items_limit
90
- self.allow_subdomains = allow_subdomains
91
110
  self.whitelist_regexp = whitelist_regexp
92
111
  self.blacklist_regexp = blacklist_regexp
112
+ self.main_content_only = main_content_only
113
+ self.max_depth = max_depth
114
+ self.max_age = max_age
115
+ self.respect_robots_txt = respect_robots_txt
116
+ self.keep_query_params = keep_query_params
93
117
  self.max_polls = max_polls
94
118
 
95
- logger.debug(f"Initializing WebCrawlerAPILoader with URL: {url}, base_url: {self.base_url}")
96
- self.client = WebCrawlerAPI(
97
- api_key=self.api_key,
98
- base_url=self.base_url,
99
- version=version
100
- )
119
+ self.client = WebCrawlerAPI(api_key=self.api_key, base_url=self.base_url)
120
+
121
+ def _crawl_params(self) -> dict:
122
+ return {
123
+ "url": self.url,
124
+ "output_formats": [self.output_format],
125
+ "items_limit": self.items_limit,
126
+ "whitelist_regexp": self.whitelist_regexp,
127
+ "blacklist_regexp": self.blacklist_regexp,
128
+ "main_content_only": self.main_content_only,
129
+ "max_depth": self.max_depth,
130
+ "max_age": self.max_age,
131
+ "respect_robots_txt": self.respect_robots_txt,
132
+ "keep_query_params": self.keep_query_params,
133
+ }
134
+
135
+ def _fetch_item_content(self, item: Any) -> Optional[str]:
136
+ """Fetch item content in the configured output format."""
137
+ if self.output_format == "markdown":
138
+ return item.get_markdown()
139
+ if self.output_format == "cleaned":
140
+ return item.get_cleaned()
141
+ return item.get_html()
101
142
 
102
143
  def _create_document(self, item: Any) -> Optional[Document]:
103
144
  """Create a Document from a job item if it's valid.
@@ -106,30 +147,46 @@ class WebCrawlerAPILoader(BaseLoader):
106
147
  item: Job item from the API response
107
148
 
108
149
  Returns:
109
- Document if item is valid and has content, None otherwise
150
+ Document if item is done and has content, None otherwise
110
151
  """
111
152
  try:
112
- logger.debug(f"Processing job item: {item}")
113
- if not (item.status == "done" and item.content):
114
- logger.debug(f"Skipping item - status: {item.status}, has_content: {bool(item.content)}")
153
+ if item.status != "done":
154
+ return None
155
+
156
+ content = self._fetch_item_content(item)
157
+ if not content:
115
158
  return None
116
159
 
117
- doc = Document(
118
- page_content=item.content,
160
+ return Document(
161
+ page_content=content,
119
162
  metadata={
120
163
  "url": item.original_url,
121
164
  "title": item.title,
122
165
  "status_code": item.page_status_code,
123
166
  "created_at": item.created_at,
124
167
  "referred_url": item.referred_url,
125
- "cost": item.cost
168
+ "depth": item.depth,
169
+ "cost": item.cost,
170
+ "job_id": item.job_id,
171
+ "item_id": item.id,
126
172
  }
127
173
  )
128
- logger.debug(f"Created document from item: {item.original_url}")
129
- return doc
130
174
  except AttributeError as e:
131
175
  logger.error(f"Failed to create document from item: {e}")
132
176
  raise WebCrawlerAPILoaderError(f"Invalid job item format: {str(e)}") from e
177
+ except Exception as e:
178
+ logger.error(f"Failed to fetch content for {getattr(item, 'original_url', '?')}: {e}")
179
+ raise WebCrawlerAPILoaderError(f"Failed to fetch item content: {str(e)}") from e
180
+
181
+ @staticmethod
182
+ def _job_error_message(job: Any) -> str:
183
+ """Build an error message from failed job items, as the job itself carries no error field."""
184
+ errors = [
185
+ f"{item.original_url}: {item.error_code or 'error'} {item.last_error or ''}".strip()
186
+ for item in job.job_items
187
+ if item.status == "error"
188
+ ]
189
+ return "; ".join(errors) if errors else "Unknown error"
133
190
 
134
191
  def load(self) -> List[Document]:
135
192
  """Load data into Document objects.
@@ -139,27 +196,10 @@ class WebCrawlerAPILoader(BaseLoader):
139
196
 
140
197
  Raises:
141
198
  WebCrawlerAPILoaderError: If there are issues with the crawling job or API response
142
- ValueError: If the response format is invalid
143
- json.JSONDecodeError: If the API response contains invalid JSON
144
199
  """
145
200
  logger.info(f"Starting crawl for URL: {self.url}")
146
201
  try:
147
- logger.debug("Making crawl request with params: "
148
- f"scrape_type={self.scrape_type}, "
149
- f"items_limit={self.items_limit}, "
150
- f"allow_subdomains={self.allow_subdomains}")
151
-
152
- job = self.client.crawl(
153
- url=self.url,
154
- scrape_type=self.scrape_type,
155
- items_limit=self.items_limit,
156
- allow_subdomains=self.allow_subdomains,
157
- whitelist_regexp=self.whitelist_regexp,
158
- blacklist_regexp=self.blacklist_regexp,
159
- max_polls=self.max_polls
160
- )
161
- logger.debug(f"Received crawl response: {job}")
162
-
202
+ job = self.client.crawl(**self._crawl_params(), max_polls=self.max_polls)
163
203
  except json.JSONDecodeError as e:
164
204
  logger.error(f"JSON decode error in crawl response: {e}")
165
205
  raise WebCrawlerAPILoaderError(
@@ -170,22 +210,21 @@ class WebCrawlerAPILoader(BaseLoader):
170
210
  raise WebCrawlerAPILoaderError(f"API request failed: {str(e)}") from e
171
211
 
172
212
  if job.status == "error":
173
- error_msg = getattr(job, 'error', 'Unknown error')
213
+ error_msg = self._job_error_message(job)
174
214
  logger.error(f"Crawl job failed with error: {error_msg}")
175
215
  raise WebCrawlerAPILoaderError(f"Crawling job failed: {error_msg}")
176
216
 
177
- documents = []
217
+ if not job.is_terminal:
218
+ raise WebCrawlerAPILoaderError(
219
+ f"Maximum number of polls ({self.max_polls}) reached without job completion"
220
+ )
221
+
178
222
  logger.info(f"Processing {len(job.job_items)} job items")
223
+ documents = []
179
224
  for item in job.job_items:
180
- try:
181
- doc = self._create_document(item)
182
- if doc:
183
- documents.append(doc)
184
- except AttributeError as e:
185
- logger.error(f"Failed to process job item: {e}")
186
- raise WebCrawlerAPILoaderError(
187
- f"Invalid job item format - missing required field: {str(e)}"
188
- ) from e
225
+ doc = self._create_document(item)
226
+ if doc:
227
+ documents.append(doc)
189
228
 
190
229
  logger.info(f"Successfully created {len(documents)} documents")
191
230
  return documents
@@ -194,26 +233,14 @@ class WebCrawlerAPILoader(BaseLoader):
194
233
  """A lazy loader for Documents.
195
234
 
196
235
  Yields:
197
- Document objects one at a time as they are crawled.
236
+ Document objects one at a time as pages finish crawling.
198
237
 
199
238
  Raises:
200
239
  WebCrawlerAPILoaderError: If there are issues with the crawling job or API response
201
- ValueError: If the response format is invalid
202
- json.JSONDecodeError: If the API response contains invalid JSON
203
240
  """
204
241
  logger.info(f"Starting async crawl for URL: {self.url}")
205
242
  try:
206
- logger.debug("Making async crawl request")
207
- response = self.client.crawl_async(
208
- url=self.url,
209
- scrape_type=self.scrape_type,
210
- items_limit=self.items_limit,
211
- allow_subdomains=self.allow_subdomains,
212
- whitelist_regexp=self.whitelist_regexp,
213
- blacklist_regexp=self.blacklist_regexp
214
- )
215
- logger.debug(f"Received async crawl response: {response}")
216
-
243
+ response = self.client.crawl_async(**self._crawl_params())
217
244
  except json.JSONDecodeError as e:
218
245
  logger.error(f"JSON decode error in async crawl response: {e}")
219
246
  raise WebCrawlerAPILoaderError(
@@ -224,16 +251,12 @@ class WebCrawlerAPILoader(BaseLoader):
224
251
  raise WebCrawlerAPILoaderError(f"API request failed: {str(e)}") from e
225
252
 
226
253
  job_id = response.id
227
- polls = 0
228
254
  processed_items = set()
229
255
  logger.info(f"Starting to poll job {job_id}")
230
256
 
231
- while polls < self.max_polls:
257
+ for _ in range(self.max_polls):
232
258
  try:
233
- logger.debug(f"Polling job {job_id} (attempt {polls + 1}/{self.max_polls})")
234
259
  job = self.client.get_job(job_id)
235
- logger.debug(f"Received job status: {job.status}")
236
-
237
260
  except json.JSONDecodeError as e:
238
261
  logger.error(f"JSON decode error in job status response: {e}")
239
262
  raise WebCrawlerAPILoaderError(
@@ -244,44 +267,34 @@ class WebCrawlerAPILoader(BaseLoader):
244
267
  raise WebCrawlerAPILoaderError(f"Failed to fetch job status: {str(e)}") from e
245
268
 
246
269
  if job.status == "error":
247
- error_msg = getattr(job, 'error', 'Unknown error')
270
+ error_msg = self._job_error_message(job)
248
271
  logger.error(f"Job failed with error: {error_msg}")
249
272
  raise WebCrawlerAPILoaderError(f"Crawling job failed: {error_msg}")
250
273
 
274
+ # Yield pages as soon as each item is done, without waiting for the whole job
251
275
  for item in job.job_items:
252
- if item.id not in processed_items:
253
- try:
254
- doc = self._create_document(item)
255
- if doc:
256
- processed_items.add(item.id)
257
- logger.debug(f"Yielding document for URL: {item.original_url}")
258
- yield doc
259
- except AttributeError as e:
260
- logger.error(f"Failed to process job item: {e}")
261
- raise WebCrawlerAPILoaderError(
262
- f"Invalid job item format - missing required field: {str(e)}"
263
- ) from e
276
+ if item.id in processed_items or item.status not in ("done", "error"):
277
+ continue
278
+ processed_items.add(item.id)
279
+ doc = self._create_document(item)
280
+ if doc:
281
+ yield doc
264
282
 
265
283
  if job.is_terminal:
266
- logger.info("Job completed successfully")
267
- break
284
+ logger.info("Job completed")
285
+ return
268
286
 
269
287
  delay_seconds = (
270
288
  job.recommended_pull_delay_ms / 1000
271
289
  if job.recommended_pull_delay_ms
272
290
  else self.client.DEFAULT_POLL_DELAY_SECONDS
273
291
  )
274
-
275
- import time
276
- logger.debug(f"Waiting {delay_seconds} seconds before next poll")
277
292
  time.sleep(delay_seconds)
278
- polls += 1
279
293
 
280
- if not job.is_terminal and polls >= self.max_polls:
281
- logger.error(f"Job timed out after {self.max_polls} polls")
282
- raise WebCrawlerAPILoaderError(
283
- f"Maximum number of polls ({self.max_polls}) reached without job completion"
284
- )
294
+ logger.error(f"Job timed out after {self.max_polls} polls")
295
+ raise WebCrawlerAPILoaderError(
296
+ f"Maximum number of polls ({self.max_polls}) reached without job completion"
297
+ )
285
298
 
286
299
  async def aload(self) -> List[Document]:
287
300
  """Asynchronously load data into Document objects.
@@ -290,21 +303,23 @@ class WebCrawlerAPILoader(BaseLoader):
290
303
  List of Document objects, one for each crawled page.
291
304
 
292
305
  Raises:
293
- RuntimeError: If the crawling job fails
306
+ WebCrawlerAPILoaderError: If the crawling job fails
294
307
  """
295
- import asyncio
296
308
  return await asyncio.to_thread(self.load)
297
309
 
298
310
  async def alazy_load(self) -> AsyncIterator[Document]:
299
311
  """An async lazy loader for Documents.
300
312
 
301
313
  Yields:
302
- Document objects one at a time as they are crawled.
314
+ Document objects one at a time as pages finish crawling.
303
315
 
304
316
  Raises:
305
- RuntimeError: If the crawling job fails
317
+ WebCrawlerAPILoaderError: If the crawling job fails
306
318
  """
307
- import asyncio
308
- for doc in self.lazy_load():
319
+ iterator = self.lazy_load()
320
+ sentinel = object()
321
+ while True:
322
+ doc = await asyncio.to_thread(next, iterator, sentinel)
323
+ if doc is sentinel:
324
+ break
309
325
  yield doc
310
- await asyncio.sleep(0)
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: webcrawlerapi-langchain
3
- Version: 0.1.0
3
+ Version: 0.2.0
4
4
  Summary: LangChain integration for WebCrawlerAPI
5
5
  Home-page: https://github.com/webcrawlerapi/webcrawlerapi-langchain
6
6
  Author: WebCrawlerAPI
@@ -8,10 +8,10 @@ Author-email: support@webcrawlerapi.com
8
8
  Classifier: Programming Language :: Python :: 3
9
9
  Classifier: License :: OSI Approved :: MIT License
10
10
  Classifier: Operating System :: OS Independent
11
- Requires-Python: >=3.8
11
+ Requires-Python: >=3.10
12
12
  Description-Content-Type: text/markdown
13
- Requires-Dist: webcrawlerapi==1.0.6
14
- Requires-Dist: langchain-core>=0.1.0
13
+ Requires-Dist: webcrawlerapi>=2.1.5
14
+ Requires-Dist: langchain-core>=0.3.0
15
15
  Dynamic: author
16
16
  Dynamic: author-email
17
17
  Dynamic: classifier
@@ -26,7 +26,7 @@ Dynamic: summary
26
26
 
27
27
  [WebcrawlerAPI](https://webcrawlerapi.com/) - is a website to LLM data API. It allows to convert websites and webpages markdown or cleaned content.
28
28
 
29
- **No subscription required**.
29
+ **Pay as you go (No subscription required).**
30
30
 
31
31
  This package provides LangChain integration for [WebCrawlerAPI](https://webcrawlerapi.com/), allowing you to easily use web crawling capabilities with [LangChain](https://www.langchain.com/) document processing pipeline.
32
32
 
@@ -40,6 +40,8 @@ pip install webcrawlerapi-langchain
40
40
 
41
41
  ## Usage
42
42
 
43
+ Requires Python 3.10 or newer. `api_key` and `base_url` fall back to the `WEBCRAWLERAPI_API_KEY` and `WEBCRAWLERAPI_BASE_URL` environment variables when not passed.
44
+
43
45
  ### Basic Loading
44
46
  ```python
45
47
  from webcrawlerapi_langchain import WebCrawlerAPILoader
@@ -48,7 +50,7 @@ from webcrawlerapi_langchain import WebCrawlerAPILoader
48
50
  loader = WebCrawlerAPILoader(
49
51
  url="https://example.com",
50
52
  api_key="your-api-key",
51
- scrape_type="markdown",
53
+ output_format="markdown",
52
54
  items_limit=10
53
55
  )
54
56
 
@@ -68,6 +70,7 @@ documents = await loader.aload()
68
70
  ```
69
71
 
70
72
  ### Lazy Loading
73
+ Yields each page as soon as it finishes crawling:
71
74
  ```python
72
75
  # Lazy loading
73
76
  for doc in loader.lazy_load():
@@ -81,16 +84,37 @@ async for doc in loader.alazy_load():
81
84
  print(doc.page_content[:100])
82
85
  ```
83
86
 
87
+ ### Error Handling
88
+ Crawling and API failures raise `WebCrawlerAPILoaderError`, which is importable from `webcrawlerapi_langchain`. Invalid parameters raise `ValueError`.
89
+ ```python
90
+ from webcrawlerapi_langchain import WebCrawlerAPILoaderError
91
+
92
+ try:
93
+ documents = loader.load()
94
+ except WebCrawlerAPILoaderError as e:
95
+ print(f"Crawl failed: {e}")
96
+ ```
97
+
84
98
  ## Configuration
85
99
 
86
100
  The loader accepts the following parameters:
87
101
 
88
- - `url`: The URL to crawl
89
- - `api_key`: Your WebCrawlerAPI API key
90
- - `scrape_type`: Type of scraping (html, cleaned, markdown)
91
- - `items_limit`: Maximum number of pages to crawl
102
+ - `url`: The URL to crawl (keyword arguments follow it)
103
+ - `api_key`: Your WebCrawlerAPI API key. Defaults to `WEBCRAWLERAPI_API_KEY`
104
+ - `base_url`: API base URL. Defaults to `WEBCRAWLERAPI_BASE_URL`
105
+ - `output_format`: Content format of each Document: `markdown` (default), `cleaned`, or `html`
106
+ - `scrape_type`: Deprecated, use `output_format` instead
107
+ - `items_limit`: Maximum number of pages to crawl (default `10`)
92
108
  - `whitelist_regexp`: Regex pattern for URL whitelist
93
109
  - `blacklist_regexp`: Regex pattern for URL blacklist
110
+ - `main_content_only`: Extract only the main content of each page (default `False`)
111
+ - `max_depth`: Maximum crawl depth. `0` is the seed URL only, `1` is the seed URL plus direct links
112
+ - `max_age`: Max age in seconds for cached content. `0` always fetches fresh content
113
+ - `respect_robots_txt`: Whether to respect robots.txt (default `False`)
114
+ - `keep_query_params`: Keep URL query params when deduplicating links
115
+ - `max_polls`: Maximum number of job status checks before giving up (default `100`)
116
+
117
+ Document metadata keys: `url`, `title`, `status_code`, `created_at`, `referred_url`, `depth`, `cost`, `job_id`, `item_id`.
94
118
 
95
119
  ### Links
96
120
  - [WebCrawlerAPI Python SDK](https://github.com/WebCrawlerAPI/webcrawlerapi-python-sdk)
@@ -0,0 +1,2 @@
1
+ webcrawlerapi>=2.1.5
2
+ langchain-core>=0.3.0
@@ -1,4 +0,0 @@
1
- from .loader import WebCrawlerAPILoader
2
-
3
- __version__ = "0.1.0"
4
- __all__ = ["WebCrawlerAPILoader"]
@@ -1,2 +0,0 @@
1
- webcrawlerapi==1.0.6
2
- langchain-core>=0.1.0