webcrawlerapi-langchain 0.1.0__tar.gz → 0.2.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {webcrawlerapi_langchain-0.1.0 → webcrawlerapi_langchain-0.2.0}/PKG-INFO +34 -10
- {webcrawlerapi_langchain-0.1.0 → webcrawlerapi_langchain-0.2.0}/README.md +30 -6
- {webcrawlerapi_langchain-0.1.0 → webcrawlerapi_langchain-0.2.0}/setup.py +4 -4
- webcrawlerapi_langchain-0.2.0/webcrawlerapi_langchain/__init__.py +4 -0
- {webcrawlerapi_langchain-0.1.0 → webcrawlerapi_langchain-0.2.0}/webcrawlerapi_langchain/loader.py +127 -112
- {webcrawlerapi_langchain-0.1.0 → webcrawlerapi_langchain-0.2.0}/webcrawlerapi_langchain.egg-info/PKG-INFO +34 -10
- webcrawlerapi_langchain-0.2.0/webcrawlerapi_langchain.egg-info/requires.txt +2 -0
- webcrawlerapi_langchain-0.1.0/webcrawlerapi_langchain/__init__.py +0 -4
- webcrawlerapi_langchain-0.1.0/webcrawlerapi_langchain.egg-info/requires.txt +0 -2
- {webcrawlerapi_langchain-0.1.0 → webcrawlerapi_langchain-0.2.0}/setup.cfg +0 -0
- {webcrawlerapi_langchain-0.1.0 → webcrawlerapi_langchain-0.2.0}/webcrawlerapi_langchain.egg-info/SOURCES.txt +0 -0
- {webcrawlerapi_langchain-0.1.0 → webcrawlerapi_langchain-0.2.0}/webcrawlerapi_langchain.egg-info/dependency_links.txt +0 -0
- {webcrawlerapi_langchain-0.1.0 → webcrawlerapi_langchain-0.2.0}/webcrawlerapi_langchain.egg-info/top_level.txt +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: webcrawlerapi-langchain
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.2.0
|
|
4
4
|
Summary: LangChain integration for WebCrawlerAPI
|
|
5
5
|
Home-page: https://github.com/webcrawlerapi/webcrawlerapi-langchain
|
|
6
6
|
Author: WebCrawlerAPI
|
|
@@ -8,10 +8,10 @@ Author-email: support@webcrawlerapi.com
|
|
|
8
8
|
Classifier: Programming Language :: Python :: 3
|
|
9
9
|
Classifier: License :: OSI Approved :: MIT License
|
|
10
10
|
Classifier: Operating System :: OS Independent
|
|
11
|
-
Requires-Python: >=3.
|
|
11
|
+
Requires-Python: >=3.10
|
|
12
12
|
Description-Content-Type: text/markdown
|
|
13
|
-
Requires-Dist: webcrawlerapi
|
|
14
|
-
Requires-Dist: langchain-core>=0.
|
|
13
|
+
Requires-Dist: webcrawlerapi>=2.1.5
|
|
14
|
+
Requires-Dist: langchain-core>=0.3.0
|
|
15
15
|
Dynamic: author
|
|
16
16
|
Dynamic: author-email
|
|
17
17
|
Dynamic: classifier
|
|
@@ -26,7 +26,7 @@ Dynamic: summary
|
|
|
26
26
|
|
|
27
27
|
[WebcrawlerAPI](https://webcrawlerapi.com/) - is a website to LLM data API. It allows to convert websites and webpages markdown or cleaned content.
|
|
28
28
|
|
|
29
|
-
**No subscription required
|
|
29
|
+
**Pay as you go (No subscription required).**
|
|
30
30
|
|
|
31
31
|
This package provides LangChain integration for [WebCrawlerAPI](https://webcrawlerapi.com/), allowing you to easily use web crawling capabilities with [LangChain](https://www.langchain.com/) document processing pipeline.
|
|
32
32
|
|
|
@@ -40,6 +40,8 @@ pip install webcrawlerapi-langchain
|
|
|
40
40
|
|
|
41
41
|
## Usage
|
|
42
42
|
|
|
43
|
+
Requires Python 3.10 or newer. `api_key` and `base_url` fall back to the `WEBCRAWLERAPI_API_KEY` and `WEBCRAWLERAPI_BASE_URL` environment variables when not passed.
|
|
44
|
+
|
|
43
45
|
### Basic Loading
|
|
44
46
|
```python
|
|
45
47
|
from webcrawlerapi_langchain import WebCrawlerAPILoader
|
|
@@ -48,7 +50,7 @@ from webcrawlerapi_langchain import WebCrawlerAPILoader
|
|
|
48
50
|
loader = WebCrawlerAPILoader(
|
|
49
51
|
url="https://example.com",
|
|
50
52
|
api_key="your-api-key",
|
|
51
|
-
|
|
53
|
+
output_format="markdown",
|
|
52
54
|
items_limit=10
|
|
53
55
|
)
|
|
54
56
|
|
|
@@ -68,6 +70,7 @@ documents = await loader.aload()
|
|
|
68
70
|
```
|
|
69
71
|
|
|
70
72
|
### Lazy Loading
|
|
73
|
+
Yields each page as soon as it finishes crawling:
|
|
71
74
|
```python
|
|
72
75
|
# Lazy loading
|
|
73
76
|
for doc in loader.lazy_load():
|
|
@@ -81,16 +84,37 @@ async for doc in loader.alazy_load():
|
|
|
81
84
|
print(doc.page_content[:100])
|
|
82
85
|
```
|
|
83
86
|
|
|
87
|
+
### Error Handling
|
|
88
|
+
Crawling and API failures raise `WebCrawlerAPILoaderError`, which is importable from `webcrawlerapi_langchain`. Invalid parameters raise `ValueError`.
|
|
89
|
+
```python
|
|
90
|
+
from webcrawlerapi_langchain import WebCrawlerAPILoaderError
|
|
91
|
+
|
|
92
|
+
try:
|
|
93
|
+
documents = loader.load()
|
|
94
|
+
except WebCrawlerAPILoaderError as e:
|
|
95
|
+
print(f"Crawl failed: {e}")
|
|
96
|
+
```
|
|
97
|
+
|
|
84
98
|
## Configuration
|
|
85
99
|
|
|
86
100
|
The loader accepts the following parameters:
|
|
87
101
|
|
|
88
|
-
- `url`: The URL to crawl
|
|
89
|
-
- `api_key`: Your WebCrawlerAPI API key
|
|
90
|
-
- `
|
|
91
|
-
- `
|
|
102
|
+
- `url`: The URL to crawl (keyword arguments follow it)
|
|
103
|
+
- `api_key`: Your WebCrawlerAPI API key. Defaults to `WEBCRAWLERAPI_API_KEY`
|
|
104
|
+
- `base_url`: API base URL. Defaults to `WEBCRAWLERAPI_BASE_URL`
|
|
105
|
+
- `output_format`: Content format of each Document: `markdown` (default), `cleaned`, or `html`
|
|
106
|
+
- `scrape_type`: Deprecated, use `output_format` instead
|
|
107
|
+
- `items_limit`: Maximum number of pages to crawl (default `10`)
|
|
92
108
|
- `whitelist_regexp`: Regex pattern for URL whitelist
|
|
93
109
|
- `blacklist_regexp`: Regex pattern for URL blacklist
|
|
110
|
+
- `main_content_only`: Extract only the main content of each page (default `False`)
|
|
111
|
+
- `max_depth`: Maximum crawl depth. `0` is the seed URL only, `1` is the seed URL plus direct links
|
|
112
|
+
- `max_age`: Max age in seconds for cached content. `0` always fetches fresh content
|
|
113
|
+
- `respect_robots_txt`: Whether to respect robots.txt (default `False`)
|
|
114
|
+
- `keep_query_params`: Keep URL query params when deduplicating links
|
|
115
|
+
- `max_polls`: Maximum number of job status checks before giving up (default `100`)
|
|
116
|
+
|
|
117
|
+
Document metadata keys: `url`, `title`, `status_code`, `created_at`, `referred_url`, `depth`, `cost`, `job_id`, `item_id`.
|
|
94
118
|
|
|
95
119
|
### Links
|
|
96
120
|
- [WebCrawlerAPI Python SDK](https://github.com/WebCrawlerAPI/webcrawlerapi-python-sdk)
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
[WebcrawlerAPI](https://webcrawlerapi.com/) - is a website to LLM data API. It allows to convert websites and webpages markdown or cleaned content.
|
|
4
4
|
|
|
5
|
-
**No subscription required
|
|
5
|
+
**Pay as you go (No subscription required).**
|
|
6
6
|
|
|
7
7
|
This package provides LangChain integration for [WebCrawlerAPI](https://webcrawlerapi.com/), allowing you to easily use web crawling capabilities with [LangChain](https://www.langchain.com/) document processing pipeline.
|
|
8
8
|
|
|
@@ -16,6 +16,8 @@ pip install webcrawlerapi-langchain
|
|
|
16
16
|
|
|
17
17
|
## Usage
|
|
18
18
|
|
|
19
|
+
Requires Python 3.10 or newer. `api_key` and `base_url` fall back to the `WEBCRAWLERAPI_API_KEY` and `WEBCRAWLERAPI_BASE_URL` environment variables when not passed.
|
|
20
|
+
|
|
19
21
|
### Basic Loading
|
|
20
22
|
```python
|
|
21
23
|
from webcrawlerapi_langchain import WebCrawlerAPILoader
|
|
@@ -24,7 +26,7 @@ from webcrawlerapi_langchain import WebCrawlerAPILoader
|
|
|
24
26
|
loader = WebCrawlerAPILoader(
|
|
25
27
|
url="https://example.com",
|
|
26
28
|
api_key="your-api-key",
|
|
27
|
-
|
|
29
|
+
output_format="markdown",
|
|
28
30
|
items_limit=10
|
|
29
31
|
)
|
|
30
32
|
|
|
@@ -44,6 +46,7 @@ documents = await loader.aload()
|
|
|
44
46
|
```
|
|
45
47
|
|
|
46
48
|
### Lazy Loading
|
|
49
|
+
Yields each page as soon as it finishes crawling:
|
|
47
50
|
```python
|
|
48
51
|
# Lazy loading
|
|
49
52
|
for doc in loader.lazy_load():
|
|
@@ -57,16 +60,37 @@ async for doc in loader.alazy_load():
|
|
|
57
60
|
print(doc.page_content[:100])
|
|
58
61
|
```
|
|
59
62
|
|
|
63
|
+
### Error Handling
|
|
64
|
+
Crawling and API failures raise `WebCrawlerAPILoaderError`, which is importable from `webcrawlerapi_langchain`. Invalid parameters raise `ValueError`.
|
|
65
|
+
```python
|
|
66
|
+
from webcrawlerapi_langchain import WebCrawlerAPILoaderError
|
|
67
|
+
|
|
68
|
+
try:
|
|
69
|
+
documents = loader.load()
|
|
70
|
+
except WebCrawlerAPILoaderError as e:
|
|
71
|
+
print(f"Crawl failed: {e}")
|
|
72
|
+
```
|
|
73
|
+
|
|
60
74
|
## Configuration
|
|
61
75
|
|
|
62
76
|
The loader accepts the following parameters:
|
|
63
77
|
|
|
64
|
-
- `url`: The URL to crawl
|
|
65
|
-
- `api_key`: Your WebCrawlerAPI API key
|
|
66
|
-
- `
|
|
67
|
-
- `
|
|
78
|
+
- `url`: The URL to crawl (keyword arguments follow it)
|
|
79
|
+
- `api_key`: Your WebCrawlerAPI API key. Defaults to `WEBCRAWLERAPI_API_KEY`
|
|
80
|
+
- `base_url`: API base URL. Defaults to `WEBCRAWLERAPI_BASE_URL`
|
|
81
|
+
- `output_format`: Content format of each Document: `markdown` (default), `cleaned`, or `html`
|
|
82
|
+
- `scrape_type`: Deprecated, use `output_format` instead
|
|
83
|
+
- `items_limit`: Maximum number of pages to crawl (default `10`)
|
|
68
84
|
- `whitelist_regexp`: Regex pattern for URL whitelist
|
|
69
85
|
- `blacklist_regexp`: Regex pattern for URL blacklist
|
|
86
|
+
- `main_content_only`: Extract only the main content of each page (default `False`)
|
|
87
|
+
- `max_depth`: Maximum crawl depth. `0` is the seed URL only, `1` is the seed URL plus direct links
|
|
88
|
+
- `max_age`: Max age in seconds for cached content. `0` always fetches fresh content
|
|
89
|
+
- `respect_robots_txt`: Whether to respect robots.txt (default `False`)
|
|
90
|
+
- `keep_query_params`: Keep URL query params when deduplicating links
|
|
91
|
+
- `max_polls`: Maximum number of job status checks before giving up (default `100`)
|
|
92
|
+
|
|
93
|
+
Document metadata keys: `url`, `title`, `status_code`, `created_at`, `referred_url`, `depth`, `cost`, `job_id`, `item_id`.
|
|
70
94
|
|
|
71
95
|
### Links
|
|
72
96
|
- [WebCrawlerAPI Python SDK](https://github.com/WebCrawlerAPI/webcrawlerapi-python-sdk)
|
|
@@ -2,11 +2,11 @@ from setuptools import setup, find_packages
|
|
|
2
2
|
|
|
3
3
|
setup(
|
|
4
4
|
name="webcrawlerapi-langchain",
|
|
5
|
-
version="0.
|
|
5
|
+
version="0.2.0",
|
|
6
6
|
packages=find_packages(),
|
|
7
7
|
install_requires=[
|
|
8
|
-
"webcrawlerapi
|
|
9
|
-
"langchain-core>=0.
|
|
8
|
+
"webcrawlerapi>=2.1.5",
|
|
9
|
+
"langchain-core>=0.3.0",
|
|
10
10
|
],
|
|
11
11
|
author="WebCrawlerAPI",
|
|
12
12
|
author_email="support@webcrawlerapi.com",
|
|
@@ -19,5 +19,5 @@ setup(
|
|
|
19
19
|
"License :: OSI Approved :: MIT License",
|
|
20
20
|
"Operating System :: OS Independent",
|
|
21
21
|
],
|
|
22
|
-
python_requires=">=3.
|
|
22
|
+
python_requires=">=3.10",
|
|
23
23
|
)
|
{webcrawlerapi_langchain-0.1.0 → webcrawlerapi_langchain-0.2.0}/webcrawlerapi_langchain/loader.py
RENAMED
|
@@ -1,7 +1,9 @@
|
|
|
1
|
-
from typing import
|
|
2
|
-
import
|
|
1
|
+
from typing import Iterator, List, Optional, AsyncIterator, Any, Literal
|
|
2
|
+
import asyncio
|
|
3
3
|
import json
|
|
4
4
|
import logging
|
|
5
|
+
import time
|
|
6
|
+
import warnings
|
|
5
7
|
from langchain_core.documents import Document
|
|
6
8
|
from langchain_core.utils import get_from_env
|
|
7
9
|
from langchain_core.document_loaders import BaseLoader
|
|
@@ -9,12 +11,16 @@ from webcrawlerapi import WebCrawlerAPI
|
|
|
9
11
|
|
|
10
12
|
# Configure logging
|
|
11
13
|
logger = logging.getLogger(__name__)
|
|
12
|
-
|
|
14
|
+
|
|
15
|
+
OutputFormat = Literal["markdown", "cleaned", "html"]
|
|
16
|
+
ALLOWED_OUTPUT_FORMATS = ("markdown", "cleaned", "html")
|
|
17
|
+
|
|
13
18
|
|
|
14
19
|
class WebCrawlerAPILoaderError(Exception):
|
|
15
20
|
"""Custom exception class for WebCrawlerAPILoader errors."""
|
|
16
21
|
pass
|
|
17
22
|
|
|
23
|
+
|
|
18
24
|
class WebCrawlerAPILoader(BaseLoader):
|
|
19
25
|
"""WebCrawlerAPI document loader integration.
|
|
20
26
|
|
|
@@ -32,12 +38,16 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
32
38
|
*,
|
|
33
39
|
api_key: Optional[str] = None,
|
|
34
40
|
base_url: Optional[str] = None,
|
|
35
|
-
|
|
36
|
-
scrape_type:
|
|
41
|
+
output_format: Optional[OutputFormat] = None,
|
|
42
|
+
scrape_type: Optional[OutputFormat] = None,
|
|
37
43
|
items_limit: int = 10,
|
|
38
|
-
allow_subdomains: bool = False,
|
|
39
44
|
whitelist_regexp: Optional[str] = None,
|
|
40
45
|
blacklist_regexp: Optional[str] = None,
|
|
46
|
+
main_content_only: bool = False,
|
|
47
|
+
max_depth: Optional[int] = None,
|
|
48
|
+
max_age: Optional[int] = None,
|
|
49
|
+
respect_robots_txt: bool = False,
|
|
50
|
+
keep_query_params: Optional[bool] = None,
|
|
41
51
|
max_polls: int = 100
|
|
42
52
|
):
|
|
43
53
|
"""Initialize the WebCrawlerAPI document loader.
|
|
@@ -46,13 +56,17 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
46
56
|
url: The URL to crawl
|
|
47
57
|
api_key: Your WebCrawlerAPI API key. If not provided, will try to get from WEBCRAWLERAPI_API_KEY env var
|
|
48
58
|
base_url: The base URL of the API. If not provided, will try to get from WEBCRAWLERAPI_BASE_URL env var
|
|
49
|
-
|
|
50
|
-
scrape_type:
|
|
59
|
+
output_format: Content format of each Document (markdown, cleaned, html). Defaults to markdown
|
|
60
|
+
scrape_type: Deprecated. Use output_format instead
|
|
51
61
|
items_limit: Maximum number of pages to crawl
|
|
52
|
-
allow_subdomains: Whether to crawl subdomains
|
|
53
62
|
whitelist_regexp: Regex pattern for URL whitelist
|
|
54
63
|
blacklist_regexp: Regex pattern for URL blacklist
|
|
55
|
-
|
|
64
|
+
main_content_only: Extract only the main content of each page
|
|
65
|
+
max_depth: Maximum crawl depth (0 = seed only, 1 = seed + direct links)
|
|
66
|
+
max_age: Max age in seconds for cached content. 0 = always fresh
|
|
67
|
+
respect_robots_txt: Whether to respect robots.txt
|
|
68
|
+
keep_query_params: Keep URL query params when deduplicating links
|
|
69
|
+
max_polls: Maximum number of status checks before giving up
|
|
56
70
|
"""
|
|
57
71
|
if not url:
|
|
58
72
|
raise ValueError("URL must be provided")
|
|
@@ -66,10 +80,17 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
66
80
|
if max_polls < 1:
|
|
67
81
|
raise ValueError("max_polls must be greater than 0")
|
|
68
82
|
|
|
69
|
-
if scrape_type not
|
|
83
|
+
if scrape_type is not None:
|
|
84
|
+
warnings.warn(
|
|
85
|
+
"scrape_type is deprecated, use output_format instead",
|
|
86
|
+
DeprecationWarning,
|
|
87
|
+
stacklevel=2,
|
|
88
|
+
)
|
|
89
|
+
output_format = output_format or scrape_type or "markdown"
|
|
90
|
+
if output_format not in ALLOWED_OUTPUT_FORMATS:
|
|
70
91
|
raise ValueError(
|
|
71
|
-
f"Invalid
|
|
72
|
-
"Allowed: '
|
|
92
|
+
f"Invalid output_format '{output_format}'. "
|
|
93
|
+
"Allowed: 'markdown', 'cleaned', 'html'."
|
|
73
94
|
)
|
|
74
95
|
|
|
75
96
|
# Get API key and base URL from env vars if not provided
|
|
@@ -78,26 +99,46 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
78
99
|
raise ValueError("API key must be provided either through api_key parameter or WEBCRAWLERAPI_API_KEY environment variable")
|
|
79
100
|
|
|
80
101
|
self.base_url = base_url or get_from_env(
|
|
81
|
-
"base_url",
|
|
82
|
-
"WEBCRAWLERAPI_BASE_URL",
|
|
102
|
+
"base_url",
|
|
103
|
+
"WEBCRAWLERAPI_BASE_URL",
|
|
83
104
|
default="https://api.webcrawlerapi.com"
|
|
84
105
|
)
|
|
85
106
|
|
|
86
107
|
self.url = url
|
|
87
|
-
self.
|
|
88
|
-
self.scrape_type = scrape_type
|
|
108
|
+
self.output_format = output_format
|
|
89
109
|
self.items_limit = items_limit
|
|
90
|
-
self.allow_subdomains = allow_subdomains
|
|
91
110
|
self.whitelist_regexp = whitelist_regexp
|
|
92
111
|
self.blacklist_regexp = blacklist_regexp
|
|
112
|
+
self.main_content_only = main_content_only
|
|
113
|
+
self.max_depth = max_depth
|
|
114
|
+
self.max_age = max_age
|
|
115
|
+
self.respect_robots_txt = respect_robots_txt
|
|
116
|
+
self.keep_query_params = keep_query_params
|
|
93
117
|
self.max_polls = max_polls
|
|
94
118
|
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
119
|
+
self.client = WebCrawlerAPI(api_key=self.api_key, base_url=self.base_url)
|
|
120
|
+
|
|
121
|
+
def _crawl_params(self) -> dict:
|
|
122
|
+
return {
|
|
123
|
+
"url": self.url,
|
|
124
|
+
"output_formats": [self.output_format],
|
|
125
|
+
"items_limit": self.items_limit,
|
|
126
|
+
"whitelist_regexp": self.whitelist_regexp,
|
|
127
|
+
"blacklist_regexp": self.blacklist_regexp,
|
|
128
|
+
"main_content_only": self.main_content_only,
|
|
129
|
+
"max_depth": self.max_depth,
|
|
130
|
+
"max_age": self.max_age,
|
|
131
|
+
"respect_robots_txt": self.respect_robots_txt,
|
|
132
|
+
"keep_query_params": self.keep_query_params,
|
|
133
|
+
}
|
|
134
|
+
|
|
135
|
+
def _fetch_item_content(self, item: Any) -> Optional[str]:
|
|
136
|
+
"""Fetch item content in the configured output format."""
|
|
137
|
+
if self.output_format == "markdown":
|
|
138
|
+
return item.get_markdown()
|
|
139
|
+
if self.output_format == "cleaned":
|
|
140
|
+
return item.get_cleaned()
|
|
141
|
+
return item.get_html()
|
|
101
142
|
|
|
102
143
|
def _create_document(self, item: Any) -> Optional[Document]:
|
|
103
144
|
"""Create a Document from a job item if it's valid.
|
|
@@ -106,30 +147,46 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
106
147
|
item: Job item from the API response
|
|
107
148
|
|
|
108
149
|
Returns:
|
|
109
|
-
Document if item is
|
|
150
|
+
Document if item is done and has content, None otherwise
|
|
110
151
|
"""
|
|
111
152
|
try:
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
153
|
+
if item.status != "done":
|
|
154
|
+
return None
|
|
155
|
+
|
|
156
|
+
content = self._fetch_item_content(item)
|
|
157
|
+
if not content:
|
|
115
158
|
return None
|
|
116
159
|
|
|
117
|
-
|
|
118
|
-
page_content=
|
|
160
|
+
return Document(
|
|
161
|
+
page_content=content,
|
|
119
162
|
metadata={
|
|
120
163
|
"url": item.original_url,
|
|
121
164
|
"title": item.title,
|
|
122
165
|
"status_code": item.page_status_code,
|
|
123
166
|
"created_at": item.created_at,
|
|
124
167
|
"referred_url": item.referred_url,
|
|
125
|
-
"
|
|
168
|
+
"depth": item.depth,
|
|
169
|
+
"cost": item.cost,
|
|
170
|
+
"job_id": item.job_id,
|
|
171
|
+
"item_id": item.id,
|
|
126
172
|
}
|
|
127
173
|
)
|
|
128
|
-
logger.debug(f"Created document from item: {item.original_url}")
|
|
129
|
-
return doc
|
|
130
174
|
except AttributeError as e:
|
|
131
175
|
logger.error(f"Failed to create document from item: {e}")
|
|
132
176
|
raise WebCrawlerAPILoaderError(f"Invalid job item format: {str(e)}") from e
|
|
177
|
+
except Exception as e:
|
|
178
|
+
logger.error(f"Failed to fetch content for {getattr(item, 'original_url', '?')}: {e}")
|
|
179
|
+
raise WebCrawlerAPILoaderError(f"Failed to fetch item content: {str(e)}") from e
|
|
180
|
+
|
|
181
|
+
@staticmethod
|
|
182
|
+
def _job_error_message(job: Any) -> str:
|
|
183
|
+
"""Build an error message from failed job items, as the job itself carries no error field."""
|
|
184
|
+
errors = [
|
|
185
|
+
f"{item.original_url}: {item.error_code or 'error'} {item.last_error or ''}".strip()
|
|
186
|
+
for item in job.job_items
|
|
187
|
+
if item.status == "error"
|
|
188
|
+
]
|
|
189
|
+
return "; ".join(errors) if errors else "Unknown error"
|
|
133
190
|
|
|
134
191
|
def load(self) -> List[Document]:
|
|
135
192
|
"""Load data into Document objects.
|
|
@@ -139,27 +196,10 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
139
196
|
|
|
140
197
|
Raises:
|
|
141
198
|
WebCrawlerAPILoaderError: If there are issues with the crawling job or API response
|
|
142
|
-
ValueError: If the response format is invalid
|
|
143
|
-
json.JSONDecodeError: If the API response contains invalid JSON
|
|
144
199
|
"""
|
|
145
200
|
logger.info(f"Starting crawl for URL: {self.url}")
|
|
146
201
|
try:
|
|
147
|
-
|
|
148
|
-
f"scrape_type={self.scrape_type}, "
|
|
149
|
-
f"items_limit={self.items_limit}, "
|
|
150
|
-
f"allow_subdomains={self.allow_subdomains}")
|
|
151
|
-
|
|
152
|
-
job = self.client.crawl(
|
|
153
|
-
url=self.url,
|
|
154
|
-
scrape_type=self.scrape_type,
|
|
155
|
-
items_limit=self.items_limit,
|
|
156
|
-
allow_subdomains=self.allow_subdomains,
|
|
157
|
-
whitelist_regexp=self.whitelist_regexp,
|
|
158
|
-
blacklist_regexp=self.blacklist_regexp,
|
|
159
|
-
max_polls=self.max_polls
|
|
160
|
-
)
|
|
161
|
-
logger.debug(f"Received crawl response: {job}")
|
|
162
|
-
|
|
202
|
+
job = self.client.crawl(**self._crawl_params(), max_polls=self.max_polls)
|
|
163
203
|
except json.JSONDecodeError as e:
|
|
164
204
|
logger.error(f"JSON decode error in crawl response: {e}")
|
|
165
205
|
raise WebCrawlerAPILoaderError(
|
|
@@ -170,22 +210,21 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
170
210
|
raise WebCrawlerAPILoaderError(f"API request failed: {str(e)}") from e
|
|
171
211
|
|
|
172
212
|
if job.status == "error":
|
|
173
|
-
error_msg =
|
|
213
|
+
error_msg = self._job_error_message(job)
|
|
174
214
|
logger.error(f"Crawl job failed with error: {error_msg}")
|
|
175
215
|
raise WebCrawlerAPILoaderError(f"Crawling job failed: {error_msg}")
|
|
176
216
|
|
|
177
|
-
|
|
217
|
+
if not job.is_terminal:
|
|
218
|
+
raise WebCrawlerAPILoaderError(
|
|
219
|
+
f"Maximum number of polls ({self.max_polls}) reached without job completion"
|
|
220
|
+
)
|
|
221
|
+
|
|
178
222
|
logger.info(f"Processing {len(job.job_items)} job items")
|
|
223
|
+
documents = []
|
|
179
224
|
for item in job.job_items:
|
|
180
|
-
|
|
181
|
-
|
|
182
|
-
|
|
183
|
-
documents.append(doc)
|
|
184
|
-
except AttributeError as e:
|
|
185
|
-
logger.error(f"Failed to process job item: {e}")
|
|
186
|
-
raise WebCrawlerAPILoaderError(
|
|
187
|
-
f"Invalid job item format - missing required field: {str(e)}"
|
|
188
|
-
) from e
|
|
225
|
+
doc = self._create_document(item)
|
|
226
|
+
if doc:
|
|
227
|
+
documents.append(doc)
|
|
189
228
|
|
|
190
229
|
logger.info(f"Successfully created {len(documents)} documents")
|
|
191
230
|
return documents
|
|
@@ -194,26 +233,14 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
194
233
|
"""A lazy loader for Documents.
|
|
195
234
|
|
|
196
235
|
Yields:
|
|
197
|
-
Document objects one at a time as
|
|
236
|
+
Document objects one at a time as pages finish crawling.
|
|
198
237
|
|
|
199
238
|
Raises:
|
|
200
239
|
WebCrawlerAPILoaderError: If there are issues with the crawling job or API response
|
|
201
|
-
ValueError: If the response format is invalid
|
|
202
|
-
json.JSONDecodeError: If the API response contains invalid JSON
|
|
203
240
|
"""
|
|
204
241
|
logger.info(f"Starting async crawl for URL: {self.url}")
|
|
205
242
|
try:
|
|
206
|
-
|
|
207
|
-
response = self.client.crawl_async(
|
|
208
|
-
url=self.url,
|
|
209
|
-
scrape_type=self.scrape_type,
|
|
210
|
-
items_limit=self.items_limit,
|
|
211
|
-
allow_subdomains=self.allow_subdomains,
|
|
212
|
-
whitelist_regexp=self.whitelist_regexp,
|
|
213
|
-
blacklist_regexp=self.blacklist_regexp
|
|
214
|
-
)
|
|
215
|
-
logger.debug(f"Received async crawl response: {response}")
|
|
216
|
-
|
|
243
|
+
response = self.client.crawl_async(**self._crawl_params())
|
|
217
244
|
except json.JSONDecodeError as e:
|
|
218
245
|
logger.error(f"JSON decode error in async crawl response: {e}")
|
|
219
246
|
raise WebCrawlerAPILoaderError(
|
|
@@ -224,16 +251,12 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
224
251
|
raise WebCrawlerAPILoaderError(f"API request failed: {str(e)}") from e
|
|
225
252
|
|
|
226
253
|
job_id = response.id
|
|
227
|
-
polls = 0
|
|
228
254
|
processed_items = set()
|
|
229
255
|
logger.info(f"Starting to poll job {job_id}")
|
|
230
256
|
|
|
231
|
-
|
|
257
|
+
for _ in range(self.max_polls):
|
|
232
258
|
try:
|
|
233
|
-
logger.debug(f"Polling job {job_id} (attempt {polls + 1}/{self.max_polls})")
|
|
234
259
|
job = self.client.get_job(job_id)
|
|
235
|
-
logger.debug(f"Received job status: {job.status}")
|
|
236
|
-
|
|
237
260
|
except json.JSONDecodeError as e:
|
|
238
261
|
logger.error(f"JSON decode error in job status response: {e}")
|
|
239
262
|
raise WebCrawlerAPILoaderError(
|
|
@@ -244,44 +267,34 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
244
267
|
raise WebCrawlerAPILoaderError(f"Failed to fetch job status: {str(e)}") from e
|
|
245
268
|
|
|
246
269
|
if job.status == "error":
|
|
247
|
-
error_msg =
|
|
270
|
+
error_msg = self._job_error_message(job)
|
|
248
271
|
logger.error(f"Job failed with error: {error_msg}")
|
|
249
272
|
raise WebCrawlerAPILoaderError(f"Crawling job failed: {error_msg}")
|
|
250
273
|
|
|
274
|
+
# Yield pages as soon as each item is done, without waiting for the whole job
|
|
251
275
|
for item in job.job_items:
|
|
252
|
-
if item.id not in
|
|
253
|
-
|
|
254
|
-
|
|
255
|
-
|
|
256
|
-
|
|
257
|
-
|
|
258
|
-
yield doc
|
|
259
|
-
except AttributeError as e:
|
|
260
|
-
logger.error(f"Failed to process job item: {e}")
|
|
261
|
-
raise WebCrawlerAPILoaderError(
|
|
262
|
-
f"Invalid job item format - missing required field: {str(e)}"
|
|
263
|
-
) from e
|
|
276
|
+
if item.id in processed_items or item.status not in ("done", "error"):
|
|
277
|
+
continue
|
|
278
|
+
processed_items.add(item.id)
|
|
279
|
+
doc = self._create_document(item)
|
|
280
|
+
if doc:
|
|
281
|
+
yield doc
|
|
264
282
|
|
|
265
283
|
if job.is_terminal:
|
|
266
|
-
logger.info("Job completed
|
|
267
|
-
|
|
284
|
+
logger.info("Job completed")
|
|
285
|
+
return
|
|
268
286
|
|
|
269
287
|
delay_seconds = (
|
|
270
288
|
job.recommended_pull_delay_ms / 1000
|
|
271
289
|
if job.recommended_pull_delay_ms
|
|
272
290
|
else self.client.DEFAULT_POLL_DELAY_SECONDS
|
|
273
291
|
)
|
|
274
|
-
|
|
275
|
-
import time
|
|
276
|
-
logger.debug(f"Waiting {delay_seconds} seconds before next poll")
|
|
277
292
|
time.sleep(delay_seconds)
|
|
278
|
-
polls += 1
|
|
279
293
|
|
|
280
|
-
|
|
281
|
-
|
|
282
|
-
|
|
283
|
-
|
|
284
|
-
)
|
|
294
|
+
logger.error(f"Job timed out after {self.max_polls} polls")
|
|
295
|
+
raise WebCrawlerAPILoaderError(
|
|
296
|
+
f"Maximum number of polls ({self.max_polls}) reached without job completion"
|
|
297
|
+
)
|
|
285
298
|
|
|
286
299
|
async def aload(self) -> List[Document]:
|
|
287
300
|
"""Asynchronously load data into Document objects.
|
|
@@ -290,21 +303,23 @@ class WebCrawlerAPILoader(BaseLoader):
|
|
|
290
303
|
List of Document objects, one for each crawled page.
|
|
291
304
|
|
|
292
305
|
Raises:
|
|
293
|
-
|
|
306
|
+
WebCrawlerAPILoaderError: If the crawling job fails
|
|
294
307
|
"""
|
|
295
|
-
import asyncio
|
|
296
308
|
return await asyncio.to_thread(self.load)
|
|
297
309
|
|
|
298
310
|
async def alazy_load(self) -> AsyncIterator[Document]:
|
|
299
311
|
"""An async lazy loader for Documents.
|
|
300
312
|
|
|
301
313
|
Yields:
|
|
302
|
-
Document objects one at a time as
|
|
314
|
+
Document objects one at a time as pages finish crawling.
|
|
303
315
|
|
|
304
316
|
Raises:
|
|
305
|
-
|
|
317
|
+
WebCrawlerAPILoaderError: If the crawling job fails
|
|
306
318
|
"""
|
|
307
|
-
|
|
308
|
-
|
|
319
|
+
iterator = self.lazy_load()
|
|
320
|
+
sentinel = object()
|
|
321
|
+
while True:
|
|
322
|
+
doc = await asyncio.to_thread(next, iterator, sentinel)
|
|
323
|
+
if doc is sentinel:
|
|
324
|
+
break
|
|
309
325
|
yield doc
|
|
310
|
-
await asyncio.sleep(0)
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: webcrawlerapi-langchain
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.2.0
|
|
4
4
|
Summary: LangChain integration for WebCrawlerAPI
|
|
5
5
|
Home-page: https://github.com/webcrawlerapi/webcrawlerapi-langchain
|
|
6
6
|
Author: WebCrawlerAPI
|
|
@@ -8,10 +8,10 @@ Author-email: support@webcrawlerapi.com
|
|
|
8
8
|
Classifier: Programming Language :: Python :: 3
|
|
9
9
|
Classifier: License :: OSI Approved :: MIT License
|
|
10
10
|
Classifier: Operating System :: OS Independent
|
|
11
|
-
Requires-Python: >=3.
|
|
11
|
+
Requires-Python: >=3.10
|
|
12
12
|
Description-Content-Type: text/markdown
|
|
13
|
-
Requires-Dist: webcrawlerapi
|
|
14
|
-
Requires-Dist: langchain-core>=0.
|
|
13
|
+
Requires-Dist: webcrawlerapi>=2.1.5
|
|
14
|
+
Requires-Dist: langchain-core>=0.3.0
|
|
15
15
|
Dynamic: author
|
|
16
16
|
Dynamic: author-email
|
|
17
17
|
Dynamic: classifier
|
|
@@ -26,7 +26,7 @@ Dynamic: summary
|
|
|
26
26
|
|
|
27
27
|
[WebcrawlerAPI](https://webcrawlerapi.com/) - is a website to LLM data API. It allows to convert websites and webpages markdown or cleaned content.
|
|
28
28
|
|
|
29
|
-
**No subscription required
|
|
29
|
+
**Pay as you go (No subscription required).**
|
|
30
30
|
|
|
31
31
|
This package provides LangChain integration for [WebCrawlerAPI](https://webcrawlerapi.com/), allowing you to easily use web crawling capabilities with [LangChain](https://www.langchain.com/) document processing pipeline.
|
|
32
32
|
|
|
@@ -40,6 +40,8 @@ pip install webcrawlerapi-langchain
|
|
|
40
40
|
|
|
41
41
|
## Usage
|
|
42
42
|
|
|
43
|
+
Requires Python 3.10 or newer. `api_key` and `base_url` fall back to the `WEBCRAWLERAPI_API_KEY` and `WEBCRAWLERAPI_BASE_URL` environment variables when not passed.
|
|
44
|
+
|
|
43
45
|
### Basic Loading
|
|
44
46
|
```python
|
|
45
47
|
from webcrawlerapi_langchain import WebCrawlerAPILoader
|
|
@@ -48,7 +50,7 @@ from webcrawlerapi_langchain import WebCrawlerAPILoader
|
|
|
48
50
|
loader = WebCrawlerAPILoader(
|
|
49
51
|
url="https://example.com",
|
|
50
52
|
api_key="your-api-key",
|
|
51
|
-
|
|
53
|
+
output_format="markdown",
|
|
52
54
|
items_limit=10
|
|
53
55
|
)
|
|
54
56
|
|
|
@@ -68,6 +70,7 @@ documents = await loader.aload()
|
|
|
68
70
|
```
|
|
69
71
|
|
|
70
72
|
### Lazy Loading
|
|
73
|
+
Yields each page as soon as it finishes crawling:
|
|
71
74
|
```python
|
|
72
75
|
# Lazy loading
|
|
73
76
|
for doc in loader.lazy_load():
|
|
@@ -81,16 +84,37 @@ async for doc in loader.alazy_load():
|
|
|
81
84
|
print(doc.page_content[:100])
|
|
82
85
|
```
|
|
83
86
|
|
|
87
|
+
### Error Handling
|
|
88
|
+
Crawling and API failures raise `WebCrawlerAPILoaderError`, which is importable from `webcrawlerapi_langchain`. Invalid parameters raise `ValueError`.
|
|
89
|
+
```python
|
|
90
|
+
from webcrawlerapi_langchain import WebCrawlerAPILoaderError
|
|
91
|
+
|
|
92
|
+
try:
|
|
93
|
+
documents = loader.load()
|
|
94
|
+
except WebCrawlerAPILoaderError as e:
|
|
95
|
+
print(f"Crawl failed: {e}")
|
|
96
|
+
```
|
|
97
|
+
|
|
84
98
|
## Configuration
|
|
85
99
|
|
|
86
100
|
The loader accepts the following parameters:
|
|
87
101
|
|
|
88
|
-
- `url`: The URL to crawl
|
|
89
|
-
- `api_key`: Your WebCrawlerAPI API key
|
|
90
|
-
- `
|
|
91
|
-
- `
|
|
102
|
+
- `url`: The URL to crawl (keyword arguments follow it)
|
|
103
|
+
- `api_key`: Your WebCrawlerAPI API key. Defaults to `WEBCRAWLERAPI_API_KEY`
|
|
104
|
+
- `base_url`: API base URL. Defaults to `WEBCRAWLERAPI_BASE_URL`
|
|
105
|
+
- `output_format`: Content format of each Document: `markdown` (default), `cleaned`, or `html`
|
|
106
|
+
- `scrape_type`: Deprecated, use `output_format` instead
|
|
107
|
+
- `items_limit`: Maximum number of pages to crawl (default `10`)
|
|
92
108
|
- `whitelist_regexp`: Regex pattern for URL whitelist
|
|
93
109
|
- `blacklist_regexp`: Regex pattern for URL blacklist
|
|
110
|
+
- `main_content_only`: Extract only the main content of each page (default `False`)
|
|
111
|
+
- `max_depth`: Maximum crawl depth. `0` is the seed URL only, `1` is the seed URL plus direct links
|
|
112
|
+
- `max_age`: Max age in seconds for cached content. `0` always fetches fresh content
|
|
113
|
+
- `respect_robots_txt`: Whether to respect robots.txt (default `False`)
|
|
114
|
+
- `keep_query_params`: Keep URL query params when deduplicating links
|
|
115
|
+
- `max_polls`: Maximum number of job status checks before giving up (default `100`)
|
|
116
|
+
|
|
117
|
+
Document metadata keys: `url`, `title`, `status_code`, `created_at`, `referred_url`, `depth`, `cost`, `job_id`, `item_id`.
|
|
94
118
|
|
|
95
119
|
### Links
|
|
96
120
|
- [WebCrawlerAPI Python SDK](https://github.com/WebCrawlerAPI/webcrawlerapi-python-sdk)
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|