webcrawlerapi-langchain 0.1.1__tar.gz → 0.2.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {webcrawlerapi_langchain-0.1.1 → webcrawlerapi_langchain-0.2.0}/PKG-INFO +34 -10
- {webcrawlerapi_langchain-0.1.1 → webcrawlerapi_langchain-0.2.0}/README.md +30 -6
- {webcrawlerapi_langchain-0.1.1 → webcrawlerapi_langchain-0.2.0}/setup.py +4 -4
- webcrawlerapi_langchain-0.2.0/webcrawlerapi_langchain/__init__.py +4 -0
- {webcrawlerapi_langchain-0.1.1 → webcrawlerapi_langchain-0.2.0}/webcrawlerapi_langchain/loader.py +127 -95
- {webcrawlerapi_langchain-0.1.1 → webcrawlerapi_langchain-0.2.0}/webcrawlerapi_langchain.egg-info/PKG-INFO +34 -10
- webcrawlerapi_langchain-0.2.0/webcrawlerapi_langchain.egg-info/requires.txt +2 -0
- webcrawlerapi_langchain-0.1.1/webcrawlerapi_langchain/__init__.py +0 -4
- webcrawlerapi_langchain-0.1.1/webcrawlerapi_langchain.egg-info/requires.txt +0 -2
- {webcrawlerapi_langchain-0.1.1 → webcrawlerapi_langchain-0.2.0}/setup.cfg +0 -0
- {webcrawlerapi_langchain-0.1.1 → webcrawlerapi_langchain-0.2.0}/webcrawlerapi_langchain.egg-info/SOURCES.txt +0 -0
- {webcrawlerapi_langchain-0.1.1 → webcrawlerapi_langchain-0.2.0}/webcrawlerapi_langchain.egg-info/dependency_links.txt +0 -0
- {webcrawlerapi_langchain-0.1.1 → webcrawlerapi_langchain-0.2.0}/webcrawlerapi_langchain.egg-info/top_level.txt +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: webcrawlerapi-langchain
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.2.0
|
|
4
4
|
Summary: LangChain integration for WebCrawlerAPI
|
|
5
5
|
Home-page: https://github.com/webcrawlerapi/webcrawlerapi-langchain
|
|
6
6
|
Author: WebCrawlerAPI
|
|
@@ -8,10 +8,10 @@ Author-email: support@webcrawlerapi.com
|
|
|
8
8
|
Classifier: Programming Language :: Python :: 3
|
|
9
9
|
Classifier: License :: OSI Approved :: MIT License
|
|
10
10
|
Classifier: Operating System :: OS Independent
|
|
11
|
-
Requires-Python: >=3.
|
|
11
|
+
Requires-Python: >=3.10
|
|
12
12
|
Description-Content-Type: text/markdown
|
|
13
|
-
Requires-Dist: webcrawlerapi>=1.
|
|
14
|
-
Requires-Dist: langchain-core>=0.
|
|
13
|
+
Requires-Dist: webcrawlerapi>=2.1.5
|
|
14
|
+
Requires-Dist: langchain-core>=0.3.0
|
|
15
15
|
Dynamic: author
|
|
16
16
|
Dynamic: author-email
|
|
17
17
|
Dynamic: classifier
|
|
@@ -26,7 +26,7 @@ Dynamic: summary
|
|
|
26
26
|
|
|
27
27
|
[WebcrawlerAPI](https://webcrawlerapi.com/) - is a website to LLM data API. It allows to convert websites and webpages markdown or cleaned content.
|
|
28
28
|
|
|
29
|
-
**No subscription required
|
|
29
|
+
**Pay as you go (No subscription required).**
|
|
30
30
|
|
|
31
31
|
This package provides LangChain integration for [WebCrawlerAPI](https://webcrawlerapi.com/), allowing you to easily use web crawling capabilities with [LangChain](https://www.langchain.com/) document processing pipeline.
|
|
32
32
|
|
|
@@ -40,6 +40,8 @@ pip install webcrawlerapi-langchain
|
|
|
40
40
|
|
|
41
41
|
## Usage
|
|
42
42
|
|
|
43
|
+
Requires Python 3.10 or newer. `api_key` and `base_url` fall back to the `WEBCRAWLERAPI_API_KEY` and `WEBCRAWLERAPI_BASE_URL` environment variables when not passed.
|
|
44
|
+
|
|
43
45
|
### Basic Loading
|
|
44
46
|
```python
|
|
45
47
|
from webcrawlerapi_langchain import WebCrawlerAPILoader
|
|
@@ -48,7 +50,7 @@ from webcrawlerapi_langchain import WebCrawlerAPILoader
|
|
|
48
50
|
loader = WebCrawlerAPILoader(
|
|
49
51
|
url="https://example.com",
|
|
50
52
|
api_key="your-api-key",
|
|
51
|
-
|
|
53
|
+
output_format="markdown",
|
|
52
54
|
items_limit=10
|
|
53
55
|
)
|
|
54
56
|
|
|
@@ -68,6 +70,7 @@ documents = await loader.aload()
|
|
|
68
70
|
```
|
|
69
71
|
|
|
70
72
|
### Lazy Loading
|
|
73
|
+
Yields each page as soon as it finishes crawling:
|
|
71
74
|
```python
|
|
72
75
|
# Lazy loading
|
|
73
76
|
for doc in loader.lazy_load():
|
|
@@ -81,16 +84,37 @@ async for doc in loader.alazy_load():
|
|
|
81
84
|
print(doc.page_content[:100])
|
|
82
85
|
```
|
|
83
86
|
|
|
87
|
+
### Error Handling
|
|
88
|
+
Crawling and API failures raise `WebCrawlerAPILoaderError`, which is importable from `webcrawlerapi_langchain`. Invalid parameters raise `ValueError`.
|
|
89
|
+
```python
|
|
90
|
+
from webcrawlerapi_langchain import WebCrawlerAPILoaderError
|
|
91
|
+
|
|
92
|
+
try:
|
|
93
|
+
documents = loader.load()
|
|
94
|
+
except WebCrawlerAPILoaderError as e:
|
|
95
|
+
print(f"Crawl failed: {e}")
|
|
96
|
+
```
|
|
97
|
+
|
|
84
98
|
## Configuration
|
|
85
99
|
|
|
86
100
|
The loader accepts the following parameters:
|
|
87
101
|
|
|
88
|
-
- `url`: The URL to crawl
|
|
89
|
-
- `api_key`: Your WebCrawlerAPI API key
|
|
90
|
-
- `
|
|
91
|
-
- `
|
|
102
|
+
- `url`: The URL to crawl (keyword arguments follow it)
|
|
103
|
+
- `api_key`: Your WebCrawlerAPI API key. Defaults to `WEBCRAWLERAPI_API_KEY`
|
|
104
|
+
- `base_url`: API base URL. Defaults to `WEBCRAWLERAPI_BASE_URL`
|
|
105
|
+
- `output_format`: Content format of each Document: `markdown` (default), `cleaned`, or `html`
|
|
106
|
+
- `scrape_type`: Deprecated, use `output_format` instead
|
|
107
|
+
- `items_limit`: Maximum number of pages to crawl (default `10`)
|
|
92
108
|
- `whitelist_regexp`: Regex pattern for URL whitelist
|
|
93
109
|
- `blacklist_regexp`: Regex pattern for URL blacklist
|
|
110
|
+
- `main_content_only`: Extract only the main content of each page (default `False`)
|
|
111
|
+
- `max_depth`: Maximum crawl depth. `0` is the seed URL only, `1` is the seed URL plus direct links
|
|
112
|
+
- `max_age`: Max age in seconds for cached content. `0` always fetches fresh content
|
|
113
|
+
- `respect_robots_txt`: Whether to respect robots.txt (default `False`)
|
|
114
|
+
- `keep_query_params`: Keep URL query params when deduplicating links
|
|
115
|
+
- `max_polls`: Maximum number of job status checks before giving up (default `100`)
|
|
116
|
+
|
|
117
|
+
Document metadata keys: `url`, `title`, `status_code`, `created_at`, `referred_url`, `depth`, `cost`, `job_id`, `item_id`.
|
|
94
118
|
|
|
95
119
|
### Links
|
|
96
120
|
- [WebCrawlerAPI Python SDK](https://github.com/WebCrawlerAPI/webcrawlerapi-python-sdk)
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
[WebcrawlerAPI](https://webcrawlerapi.com/) - is a website to LLM data API. It allows to convert websites and webpages markdown or cleaned content.
|
|
4
4
|
|
|
5
|
-
**No subscription required
|
|
5
|
+
**Pay as you go (No subscription required).**
|
|
6
6
|
|
|
7
7
|
This package provides LangChain integration for [WebCrawlerAPI](https://webcrawlerapi.com/), allowing you to easily use web crawling capabilities with [LangChain](https://www.langchain.com/) document processing pipeline.
|
|
8
8
|
|
|
@@ -16,6 +16,8 @@ pip install webcrawlerapi-langchain
|
|
|
16
16
|
|
|
17
17
|
## Usage
|
|
18
18
|
|
|
19
|
+
Requires Python 3.10 or newer. `api_key` and `base_url` fall back to the `WEBCRAWLERAPI_API_KEY` and `WEBCRAWLERAPI_BASE_URL` environment variables when not passed.
|
|
20
|
+
|
|
19
21
|
### Basic Loading
|
|
20
22
|
```python
|
|
21
23
|
from webcrawlerapi_langchain import WebCrawlerAPILoader
|
|
@@ -24,7 +26,7 @@ from webcrawlerapi_langchain import WebCrawlerAPILoader
|
|
|
24
26
|
loader = WebCrawlerAPILoader(
|
|
25
27
|
url="https://example.com",
|
|
26
28
|
api_key="your-api-key",
|
|
27
|
-
|
|
29
|
+
output_format="markdown",
|
|
28
30
|
items_limit=10
|
|
29
31
|
)
|
|
30
32
|
|
|
@@ -44,6 +46,7 @@ documents = await loader.aload()
|
|
|
44
46
|
```
|
|
45
47
|
|
|
46
48
|
### Lazy Loading
|
|
49
|
+
Yields each page as soon as it finishes crawling:
|
|
47
50
|
```python
|
|
48
51
|
# Lazy loading
|
|
49
52
|
for doc in loader.lazy_load():
|
|
@@ -57,16 +60,37 @@ async for doc in loader.alazy_load():
|
|
|
57
60
|
print(doc.page_content[:100])
|
|
58
61
|
```
|
|
59
62
|
|
|
63
|
+
### Error Handling
|
|
64
|
+
Crawling and API failures raise `WebCrawlerAPILoaderError`, which is importable from `webcrawlerapi_langchain`. Invalid parameters raise `ValueError`.
|
|
65
|
+
```python
|
|
66
|
+
from webcrawlerapi_langchain import WebCrawlerAPILoaderError
|
|
67
|
+
|
|
68
|
+
try:
|
|
69
|
+
documents = loader.load()
|
|
70
|
+
except WebCrawlerAPILoaderError as e:
|
|
71
|
+
print(f"Crawl failed: {e}")
|
|
72
|
+
```
|
|
73
|
+
|
|
60
74
|
## Configuration
|
|
61
75
|
|
|
62
76
|
The loader accepts the following parameters:
|
|
63
77
|
|
|
64
|
-
- `url`: The URL to crawl
|
|
65
|
-
- `api_key`: Your WebCrawlerAPI API key
|
|
66
|
-
- `
|
|
67
|
-
- `
|
|
78
|
+
- `url`: The URL to crawl (keyword arguments follow it)
|
|
79
|
+
- `api_key`: Your WebCrawlerAPI API key. Defaults to `WEBCRAWLERAPI_API_KEY`
|
|
80
|
+
- `base_url`: API base URL. Defaults to `WEBCRAWLERAPI_BASE_URL`
|
|
81
|
+
- `output_format`: Content format of each Document: `markdown` (default), `cleaned`, or `html`
|
|
82
|
+
- `scrape_type`: Deprecated, use `output_format` instead
|
|
83
|
+
- `items_limit`: Maximum number of pages to crawl (default `10`)
|
|
68
84
|
- `whitelist_regexp`: Regex pattern for URL whitelist
|
|
69
85
|
- `blacklist_regexp`: Regex pattern for URL blacklist
|
|
86
|
+
- `main_content_only`: Extract only the main content of each page (default `False`)
|
|
87
|
+
- `max_depth`: Maximum crawl depth. `0` is the seed URL only, `1` is the seed URL plus direct links
|
|
88
|
+
- `max_age`: Max age in seconds for cached content. `0` always fetches fresh content
|
|
89
|
+
- `respect_robots_txt`: Whether to respect robots.txt (default `False`)
|
|
90
|
+
- `keep_query_params`: Keep URL query params when deduplicating links
|
|
91
|
+
- `max_polls`: Maximum number of job status checks before giving up (default `100`)
|
|
92
|
+
|
|
93
|
+
Document metadata keys: `url`, `title`, `status_code`, `created_at`, `referred_url`, `depth`, `cost`, `job_id`, `item_id`.
|
|
70
94
|
|
|
71
95
|
### Links
|
|
72
96
|
- [WebCrawlerAPI Python SDK](https://github.com/WebCrawlerAPI/webcrawlerapi-python-sdk)
|
|
@@ -2,11 +2,11 @@ from setuptools import setup, find_packages
|
|
|
2
2
|
|
|
3
3
|
setup(
|
|
4
4
|
name="webcrawlerapi-langchain",
|
|
5
|
-
version="0.
|
|
5
|
+
version="0.2.0",
|
|
6
6
|
packages=find_packages(),
|
|
7
7
|
install_requires=[
|
|
8
|
-
"webcrawlerapi>=1.
|
|
9
|
-
"langchain-core>=0.
|
|
8
|
+
"webcrawlerapi>=2.1.5",
|
|
9
|
+
"langchain-core>=0.3.0",
|
|
10
10
|
],
|
|
11
11
|
author="WebCrawlerAPI",
|
|
12
12
|
author_email="support@webcrawlerapi.com",
|
|
@@ -19,5 +19,5 @@ setup(
|
|
|
19
19
|
"License :: OSI Approved :: MIT License",
|
|
20
20
|
"Operating System :: OS Independent",
|
|
21
21
|
],
|
|
22
|
-
python_requires=">=3.
|
|
22
|
+
python_requires=">=3.10",
|
|
23
23
|
)
|
{webcrawlerapi_langchain-0.1.1 → webcrawlerapi_langchain-0.2.0}/webcrawlerapi_langchain/loader.py
RENAMED
|
@@ -1,7 +1,9 @@
|
|
|
1
|
-
from typing import
|
|
2
|
-
import
|
|
1
|
+
from typing import Iterator, List, Optional, AsyncIterator, Any, Literal
|
|
2
|
+
import asyncio
|
|
3
3
|
import json
|
|
4
4
|
import logging
|
|
5
|
+
import time
|
|
6
|
+
import warnings
|
|
5
7
|
from langchain_core.documents import Document
|
|
6
8
|
from langchain_core.utils import get_from_env
|
|
7
9
|
from langchain_core.document_loaders import BaseLoader
|
|
@@ -10,10 +12,15 @@ from webcrawlerapi import WebCrawlerAPI
|
|
|
10
12
|
# Configure logging
|
|
11
13
|
logger = logging.getLogger(__name__)
|
|
12
14
|
|
|
15
|
+
OutputFormat = Literal["markdown", "cleaned", "html"]
|
|
16
|
+
ALLOWED_OUTPUT_FORMATS = ("markdown", "cleaned", "html")
|
|
17
|
+
|
|
18
|
+
|
|
13
19
|
class WebCrawlerAPILoaderError(Exception):
|
|
14
20
|
"""Custom exception class for WebCrawlerAPILoader errors."""
|
|
15
21
|
pass
|
|
16
22
|
|
|
23
|
+
|
|
17
24
|
class WebCrawlerAPILoader(BaseLoader):
|
|
18
25
|
"""WebCrawlerAPI document loader integration.
|
|
19
26
|
|
|
@@ -31,12 +38,16 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
31
38
|
*,
|
|
32
39
|
api_key: Optional[str] = None,
|
|
33
40
|
base_url: Optional[str] = None,
|
|
34
|
-
|
|
35
|
-
scrape_type:
|
|
41
|
+
output_format: Optional[OutputFormat] = None,
|
|
42
|
+
scrape_type: Optional[OutputFormat] = None,
|
|
36
43
|
items_limit: int = 10,
|
|
37
|
-
allow_subdomains: bool = False,
|
|
38
44
|
whitelist_regexp: Optional[str] = None,
|
|
39
45
|
blacklist_regexp: Optional[str] = None,
|
|
46
|
+
main_content_only: bool = False,
|
|
47
|
+
max_depth: Optional[int] = None,
|
|
48
|
+
max_age: Optional[int] = None,
|
|
49
|
+
respect_robots_txt: bool = False,
|
|
50
|
+
keep_query_params: Optional[bool] = None,
|
|
40
51
|
max_polls: int = 100
|
|
41
52
|
):
|
|
42
53
|
"""Initialize the WebCrawlerAPI document loader.
|
|
@@ -45,13 +56,17 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
45
56
|
url: The URL to crawl
|
|
46
57
|
api_key: Your WebCrawlerAPI API key. If not provided, will try to get from WEBCRAWLERAPI_API_KEY env var
|
|
47
58
|
base_url: The base URL of the API. If not provided, will try to get from WEBCRAWLERAPI_BASE_URL env var
|
|
48
|
-
|
|
49
|
-
scrape_type:
|
|
59
|
+
output_format: Content format of each Document (markdown, cleaned, html). Defaults to markdown
|
|
60
|
+
scrape_type: Deprecated. Use output_format instead
|
|
50
61
|
items_limit: Maximum number of pages to crawl
|
|
51
|
-
allow_subdomains: Whether to crawl subdomains
|
|
52
62
|
whitelist_regexp: Regex pattern for URL whitelist
|
|
53
63
|
blacklist_regexp: Regex pattern for URL blacklist
|
|
54
|
-
|
|
64
|
+
main_content_only: Extract only the main content of each page
|
|
65
|
+
max_depth: Maximum crawl depth (0 = seed only, 1 = seed + direct links)
|
|
66
|
+
max_age: Max age in seconds for cached content. 0 = always fresh
|
|
67
|
+
respect_robots_txt: Whether to respect robots.txt
|
|
68
|
+
keep_query_params: Keep URL query params when deduplicating links
|
|
69
|
+
max_polls: Maximum number of status checks before giving up
|
|
55
70
|
"""
|
|
56
71
|
if not url:
|
|
57
72
|
raise ValueError("URL must be provided")
|
|
@@ -65,10 +80,17 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
65
80
|
if max_polls < 1:
|
|
66
81
|
raise ValueError("max_polls must be greater than 0")
|
|
67
82
|
|
|
68
|
-
if scrape_type not
|
|
83
|
+
if scrape_type is not None:
|
|
84
|
+
warnings.warn(
|
|
85
|
+
"scrape_type is deprecated, use output_format instead",
|
|
86
|
+
DeprecationWarning,
|
|
87
|
+
stacklevel=2,
|
|
88
|
+
)
|
|
89
|
+
output_format = output_format or scrape_type or "markdown"
|
|
90
|
+
if output_format not in ALLOWED_OUTPUT_FORMATS:
|
|
69
91
|
raise ValueError(
|
|
70
|
-
f"Invalid
|
|
71
|
-
"Allowed: '
|
|
92
|
+
f"Invalid output_format '{output_format}'. "
|
|
93
|
+
"Allowed: 'markdown', 'cleaned', 'html'."
|
|
72
94
|
)
|
|
73
95
|
|
|
74
96
|
# Get API key and base URL from env vars if not provided
|
|
@@ -77,25 +99,46 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
77
99
|
raise ValueError("API key must be provided either through api_key parameter or WEBCRAWLERAPI_API_KEY environment variable")
|
|
78
100
|
|
|
79
101
|
self.base_url = base_url or get_from_env(
|
|
80
|
-
"base_url",
|
|
81
|
-
"WEBCRAWLERAPI_BASE_URL",
|
|
102
|
+
"base_url",
|
|
103
|
+
"WEBCRAWLERAPI_BASE_URL",
|
|
82
104
|
default="https://api.webcrawlerapi.com"
|
|
83
105
|
)
|
|
84
106
|
|
|
85
107
|
self.url = url
|
|
86
|
-
self.
|
|
87
|
-
self.scrape_type = scrape_type
|
|
108
|
+
self.output_format = output_format
|
|
88
109
|
self.items_limit = items_limit
|
|
89
|
-
self.allow_subdomains = allow_subdomains
|
|
90
110
|
self.whitelist_regexp = whitelist_regexp
|
|
91
111
|
self.blacklist_regexp = blacklist_regexp
|
|
112
|
+
self.main_content_only = main_content_only
|
|
113
|
+
self.max_depth = max_depth
|
|
114
|
+
self.max_age = max_age
|
|
115
|
+
self.respect_robots_txt = respect_robots_txt
|
|
116
|
+
self.keep_query_params = keep_query_params
|
|
92
117
|
self.max_polls = max_polls
|
|
93
118
|
|
|
94
|
-
self.client = WebCrawlerAPI(
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
119
|
+
self.client = WebCrawlerAPI(api_key=self.api_key, base_url=self.base_url)
|
|
120
|
+
|
|
121
|
+
def _crawl_params(self) -> dict:
|
|
122
|
+
return {
|
|
123
|
+
"url": self.url,
|
|
124
|
+
"output_formats": [self.output_format],
|
|
125
|
+
"items_limit": self.items_limit,
|
|
126
|
+
"whitelist_regexp": self.whitelist_regexp,
|
|
127
|
+
"blacklist_regexp": self.blacklist_regexp,
|
|
128
|
+
"main_content_only": self.main_content_only,
|
|
129
|
+
"max_depth": self.max_depth,
|
|
130
|
+
"max_age": self.max_age,
|
|
131
|
+
"respect_robots_txt": self.respect_robots_txt,
|
|
132
|
+
"keep_query_params": self.keep_query_params,
|
|
133
|
+
}
|
|
134
|
+
|
|
135
|
+
def _fetch_item_content(self, item: Any) -> Optional[str]:
|
|
136
|
+
"""Fetch item content in the configured output format."""
|
|
137
|
+
if self.output_format == "markdown":
|
|
138
|
+
return item.get_markdown()
|
|
139
|
+
if self.output_format == "cleaned":
|
|
140
|
+
return item.get_cleaned()
|
|
141
|
+
return item.get_html()
|
|
99
142
|
|
|
100
143
|
def _create_document(self, item: Any) -> Optional[Document]:
|
|
101
144
|
"""Create a Document from a job item if it's valid.
|
|
@@ -104,27 +147,46 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
104
147
|
item: Job item from the API response
|
|
105
148
|
|
|
106
149
|
Returns:
|
|
107
|
-
Document if item is
|
|
150
|
+
Document if item is done and has content, None otherwise
|
|
108
151
|
"""
|
|
109
152
|
try:
|
|
110
|
-
if
|
|
153
|
+
if item.status != "done":
|
|
154
|
+
return None
|
|
155
|
+
|
|
156
|
+
content = self._fetch_item_content(item)
|
|
157
|
+
if not content:
|
|
111
158
|
return None
|
|
112
159
|
|
|
113
|
-
|
|
114
|
-
page_content=
|
|
160
|
+
return Document(
|
|
161
|
+
page_content=content,
|
|
115
162
|
metadata={
|
|
116
163
|
"url": item.original_url,
|
|
117
164
|
"title": item.title,
|
|
118
165
|
"status_code": item.page_status_code,
|
|
119
166
|
"created_at": item.created_at,
|
|
120
167
|
"referred_url": item.referred_url,
|
|
121
|
-
"
|
|
168
|
+
"depth": item.depth,
|
|
169
|
+
"cost": item.cost,
|
|
170
|
+
"job_id": item.job_id,
|
|
171
|
+
"item_id": item.id,
|
|
122
172
|
}
|
|
123
173
|
)
|
|
124
|
-
return doc
|
|
125
174
|
except AttributeError as e:
|
|
126
175
|
logger.error(f"Failed to create document from item: {e}")
|
|
127
176
|
raise WebCrawlerAPILoaderError(f"Invalid job item format: {str(e)}") from e
|
|
177
|
+
except Exception as e:
|
|
178
|
+
logger.error(f"Failed to fetch content for {getattr(item, 'original_url', '?')}: {e}")
|
|
179
|
+
raise WebCrawlerAPILoaderError(f"Failed to fetch item content: {str(e)}") from e
|
|
180
|
+
|
|
181
|
+
@staticmethod
|
|
182
|
+
def _job_error_message(job: Any) -> str:
|
|
183
|
+
"""Build an error message from failed job items, as the job itself carries no error field."""
|
|
184
|
+
errors = [
|
|
185
|
+
f"{item.original_url}: {item.error_code or 'error'} {item.last_error or ''}".strip()
|
|
186
|
+
for item in job.job_items
|
|
187
|
+
if item.status == "error"
|
|
188
|
+
]
|
|
189
|
+
return "; ".join(errors) if errors else "Unknown error"
|
|
128
190
|
|
|
129
191
|
def load(self) -> List[Document]:
|
|
130
192
|
"""Load data into Document objects.
|
|
@@ -134,21 +196,10 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
134
196
|
|
|
135
197
|
Raises:
|
|
136
198
|
WebCrawlerAPILoaderError: If there are issues with the crawling job or API response
|
|
137
|
-
ValueError: If the response format is invalid
|
|
138
|
-
json.JSONDecodeError: If the API response contains invalid JSON
|
|
139
199
|
"""
|
|
140
200
|
logger.info(f"Starting crawl for URL: {self.url}")
|
|
141
201
|
try:
|
|
142
|
-
job = self.client.crawl(
|
|
143
|
-
url=self.url,
|
|
144
|
-
scrape_type=self.scrape_type,
|
|
145
|
-
items_limit=self.items_limit,
|
|
146
|
-
allow_subdomains=self.allow_subdomains,
|
|
147
|
-
whitelist_regexp=self.whitelist_regexp,
|
|
148
|
-
blacklist_regexp=self.blacklist_regexp,
|
|
149
|
-
max_polls=self.max_polls
|
|
150
|
-
)
|
|
151
|
-
|
|
202
|
+
job = self.client.crawl(**self._crawl_params(), max_polls=self.max_polls)
|
|
152
203
|
except json.JSONDecodeError as e:
|
|
153
204
|
logger.error(f"JSON decode error in crawl response: {e}")
|
|
154
205
|
raise WebCrawlerAPILoaderError(
|
|
@@ -159,22 +210,21 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
159
210
|
raise WebCrawlerAPILoaderError(f"API request failed: {str(e)}") from e
|
|
160
211
|
|
|
161
212
|
if job.status == "error":
|
|
162
|
-
error_msg =
|
|
213
|
+
error_msg = self._job_error_message(job)
|
|
163
214
|
logger.error(f"Crawl job failed with error: {error_msg}")
|
|
164
215
|
raise WebCrawlerAPILoaderError(f"Crawling job failed: {error_msg}")
|
|
165
216
|
|
|
166
|
-
|
|
217
|
+
if not job.is_terminal:
|
|
218
|
+
raise WebCrawlerAPILoaderError(
|
|
219
|
+
f"Maximum number of polls ({self.max_polls}) reached without job completion"
|
|
220
|
+
)
|
|
221
|
+
|
|
167
222
|
logger.info(f"Processing {len(job.job_items)} job items")
|
|
223
|
+
documents = []
|
|
168
224
|
for item in job.job_items:
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
|
|
172
|
-
documents.append(doc)
|
|
173
|
-
except AttributeError as e:
|
|
174
|
-
logger.error(f"Failed to process job item: {e}")
|
|
175
|
-
raise WebCrawlerAPILoaderError(
|
|
176
|
-
f"Invalid job item format - missing required field: {str(e)}"
|
|
177
|
-
) from e
|
|
225
|
+
doc = self._create_document(item)
|
|
226
|
+
if doc:
|
|
227
|
+
documents.append(doc)
|
|
178
228
|
|
|
179
229
|
logger.info(f"Successfully created {len(documents)} documents")
|
|
180
230
|
return documents
|
|
@@ -183,24 +233,14 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
183
233
|
"""A lazy loader for Documents.
|
|
184
234
|
|
|
185
235
|
Yields:
|
|
186
|
-
Document objects one at a time as
|
|
236
|
+
Document objects one at a time as pages finish crawling.
|
|
187
237
|
|
|
188
238
|
Raises:
|
|
189
239
|
WebCrawlerAPILoaderError: If there are issues with the crawling job or API response
|
|
190
|
-
ValueError: If the response format is invalid
|
|
191
|
-
json.JSONDecodeError: If the API response contains invalid JSON
|
|
192
240
|
"""
|
|
193
241
|
logger.info(f"Starting async crawl for URL: {self.url}")
|
|
194
242
|
try:
|
|
195
|
-
response = self.client.crawl_async(
|
|
196
|
-
url=self.url,
|
|
197
|
-
scrape_type=self.scrape_type,
|
|
198
|
-
items_limit=self.items_limit,
|
|
199
|
-
allow_subdomains=self.allow_subdomains,
|
|
200
|
-
whitelist_regexp=self.whitelist_regexp,
|
|
201
|
-
blacklist_regexp=self.blacklist_regexp
|
|
202
|
-
)
|
|
203
|
-
|
|
243
|
+
response = self.client.crawl_async(**self._crawl_params())
|
|
204
244
|
except json.JSONDecodeError as e:
|
|
205
245
|
logger.error(f"JSON decode error in async crawl response: {e}")
|
|
206
246
|
raise WebCrawlerAPILoaderError(
|
|
@@ -211,14 +251,12 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
211
251
|
raise WebCrawlerAPILoaderError(f"API request failed: {str(e)}") from e
|
|
212
252
|
|
|
213
253
|
job_id = response.id
|
|
214
|
-
polls = 0
|
|
215
254
|
processed_items = set()
|
|
216
255
|
logger.info(f"Starting to poll job {job_id}")
|
|
217
256
|
|
|
218
|
-
|
|
257
|
+
for _ in range(self.max_polls):
|
|
219
258
|
try:
|
|
220
259
|
job = self.client.get_job(job_id)
|
|
221
|
-
|
|
222
260
|
except json.JSONDecodeError as e:
|
|
223
261
|
logger.error(f"JSON decode error in job status response: {e}")
|
|
224
262
|
raise WebCrawlerAPILoaderError(
|
|
@@ -229,42 +267,34 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
229
267
|
raise WebCrawlerAPILoaderError(f"Failed to fetch job status: {str(e)}") from e
|
|
230
268
|
|
|
231
269
|
if job.status == "error":
|
|
232
|
-
error_msg =
|
|
270
|
+
error_msg = self._job_error_message(job)
|
|
233
271
|
logger.error(f"Job failed with error: {error_msg}")
|
|
234
272
|
raise WebCrawlerAPILoaderError(f"Crawling job failed: {error_msg}")
|
|
235
273
|
|
|
274
|
+
# Yield pages as soon as each item is done, without waiting for the whole job
|
|
236
275
|
for item in job.job_items:
|
|
237
|
-
if item.id not in
|
|
238
|
-
|
|
239
|
-
|
|
240
|
-
|
|
241
|
-
|
|
242
|
-
|
|
243
|
-
except AttributeError as e:
|
|
244
|
-
logger.error(f"Failed to process job item: {e}")
|
|
245
|
-
raise WebCrawlerAPILoaderError(
|
|
246
|
-
f"Invalid job item format - missing required field: {str(e)}"
|
|
247
|
-
) from e
|
|
276
|
+
if item.id in processed_items or item.status not in ("done", "error"):
|
|
277
|
+
continue
|
|
278
|
+
processed_items.add(item.id)
|
|
279
|
+
doc = self._create_document(item)
|
|
280
|
+
if doc:
|
|
281
|
+
yield doc
|
|
248
282
|
|
|
249
283
|
if job.is_terminal:
|
|
250
|
-
logger.info("Job completed
|
|
251
|
-
|
|
284
|
+
logger.info("Job completed")
|
|
285
|
+
return
|
|
252
286
|
|
|
253
287
|
delay_seconds = (
|
|
254
288
|
job.recommended_pull_delay_ms / 1000
|
|
255
289
|
if job.recommended_pull_delay_ms
|
|
256
290
|
else self.client.DEFAULT_POLL_DELAY_SECONDS
|
|
257
291
|
)
|
|
258
|
-
|
|
259
|
-
import time
|
|
260
292
|
time.sleep(delay_seconds)
|
|
261
|
-
polls += 1
|
|
262
293
|
|
|
263
|
-
|
|
264
|
-
|
|
265
|
-
|
|
266
|
-
|
|
267
|
-
)
|
|
294
|
+
logger.error(f"Job timed out after {self.max_polls} polls")
|
|
295
|
+
raise WebCrawlerAPILoaderError(
|
|
296
|
+
f"Maximum number of polls ({self.max_polls}) reached without job completion"
|
|
297
|
+
)
|
|
268
298
|
|
|
269
299
|
async def aload(self) -> List[Document]:
|
|
270
300
|
"""Asynchronously load data into Document objects.
|
|
@@ -273,21 +303,23 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
273
303
|
List of Document objects, one for each crawled page.
|
|
274
304
|
|
|
275
305
|
Raises:
|
|
276
|
-
|
|
306
|
+
WebCrawlerAPILoaderError: If the crawling job fails
|
|
277
307
|
"""
|
|
278
|
-
import asyncio
|
|
279
308
|
return await asyncio.to_thread(self.load)
|
|
280
309
|
|
|
281
310
|
async def alazy_load(self) -> AsyncIterator[Document]:
|
|
282
311
|
"""An async lazy loader for Documents.
|
|
283
312
|
|
|
284
313
|
Yields:
|
|
285
|
-
Document objects one at a time as
|
|
314
|
+
Document objects one at a time as pages finish crawling.
|
|
286
315
|
|
|
287
316
|
Raises:
|
|
288
|
-
|
|
317
|
+
WebCrawlerAPILoaderError: If the crawling job fails
|
|
289
318
|
"""
|
|
290
|
-
|
|
291
|
-
|
|
319
|
+
iterator = self.lazy_load()
|
|
320
|
+
sentinel = object()
|
|
321
|
+
while True:
|
|
322
|
+
doc = await asyncio.to_thread(next, iterator, sentinel)
|
|
323
|
+
if doc is sentinel:
|
|
324
|
+
break
|
|
292
325
|
yield doc
|
|
293
|
-
await asyncio.sleep(0)
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: webcrawlerapi-langchain
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.2.0
|
|
4
4
|
Summary: LangChain integration for WebCrawlerAPI
|
|
5
5
|
Home-page: https://github.com/webcrawlerapi/webcrawlerapi-langchain
|
|
6
6
|
Author: WebCrawlerAPI
|
|
@@ -8,10 +8,10 @@ Author-email: support@webcrawlerapi.com
|
|
|
8
8
|
Classifier: Programming Language :: Python :: 3
|
|
9
9
|
Classifier: License :: OSI Approved :: MIT License
|
|
10
10
|
Classifier: Operating System :: OS Independent
|
|
11
|
-
Requires-Python: >=3.
|
|
11
|
+
Requires-Python: >=3.10
|
|
12
12
|
Description-Content-Type: text/markdown
|
|
13
|
-
Requires-Dist: webcrawlerapi>=1.
|
|
14
|
-
Requires-Dist: langchain-core>=0.
|
|
13
|
+
Requires-Dist: webcrawlerapi>=2.1.5
|
|
14
|
+
Requires-Dist: langchain-core>=0.3.0
|
|
15
15
|
Dynamic: author
|
|
16
16
|
Dynamic: author-email
|
|
17
17
|
Dynamic: classifier
|
|
@@ -26,7 +26,7 @@ Dynamic: summary
|
|
|
26
26
|
|
|
27
27
|
[WebcrawlerAPI](https://webcrawlerapi.com/) - is a website to LLM data API. It allows to convert websites and webpages markdown or cleaned content.
|
|
28
28
|
|
|
29
|
-
**No subscription required
|
|
29
|
+
**Pay as you go (No subscription required).**
|
|
30
30
|
|
|
31
31
|
This package provides LangChain integration for [WebCrawlerAPI](https://webcrawlerapi.com/), allowing you to easily use web crawling capabilities with [LangChain](https://www.langchain.com/) document processing pipeline.
|
|
32
32
|
|
|
@@ -40,6 +40,8 @@ pip install webcrawlerapi-langchain
|
|
|
40
40
|
|
|
41
41
|
## Usage
|
|
42
42
|
|
|
43
|
+
Requires Python 3.10 or newer. `api_key` and `base_url` fall back to the `WEBCRAWLERAPI_API_KEY` and `WEBCRAWLERAPI_BASE_URL` environment variables when not passed.
|
|
44
|
+
|
|
43
45
|
### Basic Loading
|
|
44
46
|
```python
|
|
45
47
|
from webcrawlerapi_langchain import WebCrawlerAPILoader
|
|
@@ -48,7 +50,7 @@ from webcrawlerapi_langchain import WebCrawlerAPILoader
|
|
|
48
50
|
loader = WebCrawlerAPILoader(
|
|
49
51
|
url="https://example.com",
|
|
50
52
|
api_key="your-api-key",
|
|
51
|
-
|
|
53
|
+
output_format="markdown",
|
|
52
54
|
items_limit=10
|
|
53
55
|
)
|
|
54
56
|
|
|
@@ -68,6 +70,7 @@ documents = await loader.aload()
|
|
|
68
70
|
```
|
|
69
71
|
|
|
70
72
|
### Lazy Loading
|
|
73
|
+
Yields each page as soon as it finishes crawling:
|
|
71
74
|
```python
|
|
72
75
|
# Lazy loading
|
|
73
76
|
for doc in loader.lazy_load():
|
|
@@ -81,16 +84,37 @@ async for doc in loader.alazy_load():
|
|
|
81
84
|
print(doc.page_content[:100])
|
|
82
85
|
```
|
|
83
86
|
|
|
87
|
+
### Error Handling
|
|
88
|
+
Crawling and API failures raise `WebCrawlerAPILoaderError`, which is importable from `webcrawlerapi_langchain`. Invalid parameters raise `ValueError`.
|
|
89
|
+
```python
|
|
90
|
+
from webcrawlerapi_langchain import WebCrawlerAPILoaderError
|
|
91
|
+
|
|
92
|
+
try:
|
|
93
|
+
documents = loader.load()
|
|
94
|
+
except WebCrawlerAPILoaderError as e:
|
|
95
|
+
print(f"Crawl failed: {e}")
|
|
96
|
+
```
|
|
97
|
+
|
|
84
98
|
## Configuration
|
|
85
99
|
|
|
86
100
|
The loader accepts the following parameters:
|
|
87
101
|
|
|
88
|
-
- `url`: The URL to crawl
|
|
89
|
-
- `api_key`: Your WebCrawlerAPI API key
|
|
90
|
-
- `
|
|
91
|
-
- `
|
|
102
|
+
- `url`: The URL to crawl (keyword arguments follow it)
|
|
103
|
+
- `api_key`: Your WebCrawlerAPI API key. Defaults to `WEBCRAWLERAPI_API_KEY`
|
|
104
|
+
- `base_url`: API base URL. Defaults to `WEBCRAWLERAPI_BASE_URL`
|
|
105
|
+
- `output_format`: Content format of each Document: `markdown` (default), `cleaned`, or `html`
|
|
106
|
+
- `scrape_type`: Deprecated, use `output_format` instead
|
|
107
|
+
- `items_limit`: Maximum number of pages to crawl (default `10`)
|
|
92
108
|
- `whitelist_regexp`: Regex pattern for URL whitelist
|
|
93
109
|
- `blacklist_regexp`: Regex pattern for URL blacklist
|
|
110
|
+
- `main_content_only`: Extract only the main content of each page (default `False`)
|
|
111
|
+
- `max_depth`: Maximum crawl depth. `0` is the seed URL only, `1` is the seed URL plus direct links
|
|
112
|
+
- `max_age`: Max age in seconds for cached content. `0` always fetches fresh content
|
|
113
|
+
- `respect_robots_txt`: Whether to respect robots.txt (default `False`)
|
|
114
|
+
- `keep_query_params`: Keep URL query params when deduplicating links
|
|
115
|
+
- `max_polls`: Maximum number of job status checks before giving up (default `100`)
|
|
116
|
+
|
|
117
|
+
Document metadata keys: `url`, `title`, `status_code`, `created_at`, `referred_url`, `depth`, `cost`, `job_id`, `item_id`.
|
|
94
118
|
|
|
95
119
|
### Links
|
|
96
120
|
- [WebCrawlerAPI Python SDK](https://github.com/WebCrawlerAPI/webcrawlerapi-python-sdk)
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|