npm - @crawlee/core - Versions diffs - 3.3.4-beta.16 → 3.3.4-beta.17 - Mend

@crawlee/core 3.3.4-beta.16 → 3.3.4-beta.17

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.

Files changed (2) hide show

package/README.md +108 -27
package/package.json +19 -12

package/README.md CHANGED Viewed

@@ -1,46 +1,127 @@
-# `@crawlee/core`
+<h1 align="center">
+    <a href="https://crawlee.dev">
+        <picture>
+          <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/apify/crawlee/master/website/static/img/crawlee-dark.svg?sanitize=true">
+          <img alt="Crawlee" src="https://raw.githubusercontent.com/apify/crawlee/master/website/static/img/crawlee-light.svg?sanitize=true" width="500">
+        </picture>
+    </a>
+    <br>
+    <small>A web scraping and browser automation library</small>
+</h1>
-Core set of classes required for Crawlee.
+<p align=center>
+    <a href="https://www.npmjs.com/package/@crawlee/core" rel="nofollow"><img src="https://img.shields.io/npm/v/@crawlee/core.svg" alt="NPM latest version" data-canonical-src="https://img.shields.io/npm/v/@crawlee/core/next.svg" style="max-width: 100%;"></a>
+    <a href="https://www.npmjs.com/package/@crawlee/core" rel="nofollow"><img src="https://img.shields.io/npm/dm/@crawlee/core.svg" alt="Downloads" data-canonical-src="https://img.shields.io/npm/dm/@crawlee/core.svg" style="max-width: 100%;"></a>
+    <a href="https://discord.gg/jyEM2PRvMU" rel="nofollow"><img src="https://img.shields.io/discord/801163717915574323?label=discord" alt="Chat on discord" data-canonical-src="https://img.shields.io/discord/801163717915574323?label=discord" style="max-width: 100%;"></a>
+    <a href="https://github.com/apify/crawlee/actions/workflows/test-and-release.yml"><img src="https://github.com/apify/crawlee/actions/workflows/test-and-release.yml/badge.svg?branch=master" alt="Build Status" style="max-width: 100%;"></a>
+</p>
-The [`crawlee`](https://www.npmjs.com/package/crawlee) package consists of several smaller packages, released separately under `@crawlee` namespace:
+> ℹ️ Crawlee is the successor to [Apify SDK](https://sdk.apify.com). 🎉 Fully rewritten in **TypeScript** for a better developer experience, and with even more powerful anti-blocking features. The interface is almost the same as Apify SDK so upgrading is a breeze. Read [the upgrading guide](https://crawlee.dev/docs/upgrading/upgrading-to-v3) to learn about the changes. ℹ️
-- [`@crawlee/core`](https://crawlee.dev/api/core): the base for all the crawler implementations, also contains things like `Request`, `RequestQueue`, `RequestList` or `Dataset` classes
-- [`@crawlee/cheerio`](https://crawlee.dev/api/cheerio-crawler): exports `CheerioCrawler`
-- [`@crawlee/playwright`](https://crawlee.dev/api/playwright-crawler): exports `PlaywrightCrawler`
-- [`@crawlee/puppeteer`](https://crawlee.dev/api/puppeteer-crawler): exports `PuppeteerCrawler`
-- [`@crawlee/linkedom`](https://crawlee.dev/api/jsdom-crawler): exports `LinkeDOMCrawler`
-- [`@crawlee/jsdom`](https://crawlee.dev/api/jsdom-crawler): exports `JSDOMCrawler`
-- [`@crawlee/basic`](https://crawlee.dev/api/basic-crawler): exports `BasicCrawler`
-- [`@crawlee/http`](https://crawlee.dev/api/http-crawler): exports `HttpCrawler` (which is used for creating [`@crawlee/jsdom`](https://crawlee.dev/api/jsdom-crawler) and [`@crawlee/cheerio`](https://crawlee.dev/api/cheerio-crawler))
-- [`@crawlee/browser`](https://crawlee.dev/api/browser-crawler): exports `BrowserCrawler` (which is used for creating [`@crawlee/playwright`](https://crawlee.dev/api/playwright-crawler) and [`@crawlee/puppeteer`](https://crawlee.dev/api/puppeteer-crawler))
-- [`@crawlee/memory-storage`](https://crawlee.dev/api/memory-storage): [`@apify/storage-local`](https://npmjs.com/package/@apify/storage-local) alternative
-- [`@crawlee/browser-pool`](https://crawlee.dev/api/browser-pool): previously [`browser-pool`](https://npmjs.com/package/browser-pool) package
-- [`@crawlee/utils`](https://crawlee.dev/api/utils): utility methods
-- [`@crawlee/types`](https://crawlee.dev/api/types): holds TS interfaces mainly about the [`StorageClient`](https://crawlee.dev/api/core/interface/StorageClient)
+Crawlee covers your crawling and scraping end-to-end and **helps you build reliable scrapers. Fast.**
-## Installing Crawlee
+Your crawlers will appear human-like and fly under the radar of modern bot protections even with the default configuration. Crawlee gives you the tools to crawl the web for links, scrape data, and store it to disk or cloud while staying configurable to suit your project's needs.
-Most of the Crawlee packages are extending and reexporting each other, so it's enough to install just the one you plan on using, e.g. `@crawlee/playwright` if you plan on using `playwright` - it already contains everything from the `@crawlee/browser` package, which includes everything from `@crawlee/basic`, which includes everything from `@crawlee/core`.
+Crawlee is available as the [`crawlee`](https://www.npmjs.com/package/crawlee) NPM package.
-If we don't care much about additional code being pulled in, we can just use the `crawlee` meta-package, which contains (re-exports) most of the `@crawlee/*` packages, and therefore contains all the crawler classes.
+> 👉 **View full documentation, guides and examples on the [Crawlee project website](https://crawlee.dev)** 👈
+## Installation
+We recommend visiting the [Introduction tutorial](https://crawlee.dev/docs/introduction) in Crawlee documentation for more information.
+> Crawlee requires **Node.js 16 or higher**.
+### With Crawlee CLI
+The fastest way to try Crawlee out is to use the **Crawlee CLI** and choose the **Getting started example**. The CLI will install all the necessary dependencies and add boilerplate code for you to play with.
 ```bash
-npm install crawlee
+npx crawlee create my-crawler
 ```
-Or if all we need is cheerio support, we can install only `@crawlee/cheerio`.
 ```bash
-npm install @crawlee/cheerio
+cd my-crawler
+npm start
 ```
-When using `playwright` or `puppeteer`, we still need to install those dependencies explicitly - this allows the users to be in control of which version will be used.
+### Manual installation
+If you prefer adding Crawlee **into your own project**, try the example below. Because it uses `PlaywrightCrawler` we also need to install [Playwright](https://playwright.dev). It's not bundled with Crawlee to reduce install size.
 ```bash
 npm install crawlee playwright
-# or npm install @crawlee/playwright playwright
 ```
-Alternatively we can also use the `crawlee` meta-package which contains (re-exports) most of the `@crawlee/*` packages, and therefore contains all the crawler classes.
+```js
+import { PlaywrightCrawler, Dataset } from 'crawlee';
+// PlaywrightCrawler crawls the web using a headless
+// browser controlled by the Playwright library.
+const crawler = new PlaywrightCrawler({
+    // Use the requestHandler to process each of the crawled pages.
+    async requestHandler({ request, page, enqueueLinks, log }) {
+        const title = await page.title();
+        log.info(`Title of ${request.loadedUrl} is '${title}'`);
+        // Save results as JSON to ./storage/datasets/default
+        await Dataset.pushData({ title, url: request.loadedUrl });
+        // Extract links from the current page
+        // and add them to the crawling queue.
+        await enqueueLinks();
+    },
+    // Uncomment this option to see the browser window.
+    // headless: false,
+});
+// Add first URL to the queue and start the crawl.
+await crawler.run(['https://crawlee.dev']);
+```
+By default, Crawlee stores data to `./storage` in the current working directory. You can override this directory via Crawlee configuration. For details, see [Configuration guide](https://crawlee.dev/docs/guides/configuration), [Request storage](https://crawlee.dev/docs/guides/request-storage) and [Result storage](https://crawlee.dev/docs/guides/result-storage).
+## 🛠 Features
+- Single interface for **HTTP and headless browser** crawling
+- Persistent **queue** for URLs to crawl (breadth & depth first)
+- Pluggable **storage** of both tabular data and files
+- Automatic **scaling** with available system resources
+- Integrated **proxy rotation** and session management
+- Lifecycles customizable with **hooks**
+- **CLI** to bootstrap your projects
+- Configurable **routing**, **error handling** and **retries**
+- **Dockerfiles** ready to deploy
+- Written in **TypeScript** with generics
+### 👾 HTTP crawling
+- Zero config **HTTP2 support**, even for proxies
+- Automatic generation of **browser-like headers**
+- Replication of browser **TLS fingerprints**
+- Integrated fast **HTML parsers**. Cheerio and JSDOM
+- Yes, you can scrape **JSON APIs** as well
+### 💻 Real browser crawling
+- JavaScript **rendering** and **screenshots**
+- **Headless** and **headful** support
+- Zero-config generation of **human-like fingerprints**
+- Automatic **browser management**
+- Use **Playwright** and **Puppeteer** with the same interface
+- **Chrome**, **Firefox**, **Webkit** and many others
+## Usage on the Apify platform
+Crawlee is open-source and runs anywhere, but since it's developed by [Apify](https://apify.com), it's easy to set up on the Apify platform and run in the cloud. Visit the [Apify SDK website](https://sdk.apify.com) to learn more about deploying Crawlee to the Apify platform.
+## Support
+If you find any bug or issue with Crawlee, please [submit an issue on GitHub](https://github.com/apify/crawlee/issues). For questions, you can ask on [Stack Overflow](https://stackoverflow.com/questions/tagged/apify), in GitHub Discussions or you can join our [Discord server](https://discord.com/invite/jyEM2PRvMU).
+## Contributing
+Your code contributions are welcome, and you'll be praised to eternity! If you have any ideas for improvements, either submit an issue or create a pull request. For contribution guidelines and the code of conduct, see [CONTRIBUTING.md](https://github.com/apify/crawlee/blob/master/CONTRIBUTING.md).
+## License
-> Sometimes you might want to use some utility methods from `@crawlee/utils`, so you might want to install that as well. This package contains some utilities that were previously available under `Apify.utils`. Browser related utilities can be also found in the crawler packages (e.g. `@crawlee/playwright`).
+This project is licensed under the Apache License 2.0 - see the [LICENSE.md](https://github.com/apify/crawlee/blob/master/LICENSE.md) file for details.

package/package.json CHANGED Viewed

@@ -1,18 +1,18 @@
 {
     "name": "@crawlee/core",
-    "version": "3.3.4-beta.16",
+    "version": "3.3.4-beta.17",
     "description": "The scalable web crawling and scraping library for JavaScript/Node.js. Enables development of data extraction and web automation jobs (not only) with headless Chrome and Puppeteer.",
     "engines": {
         "node": ">=16.0.0"
     },
-    "main": "./dist/index.js",
-    "module": "./dist/index.mjs",
-    "types": "./dist/index.d.ts",
+    "main": "./index.js",
+    "module": "./index.mjs",
+    "types": "./index.d.ts",
     "exports": {
         ".": {
-            "import": "./dist/index.mjs",
-            "require": "./dist/index.js",
-            "types": "./dist/index.d.ts"
+            "import": "./index.mjs",
+            "require": "./index.js",
+            "types": "./index.d.ts"
         },
         "./package.json": "./package.json"
     },
@@ -46,7 +46,7 @@
     "scripts": {
         "build": "yarn clean && yarn compile && yarn copy",
         "clean": "rimraf ./dist",
-        "compile": "tsc -p tsconfig.build.json && gen-esm-wrapper ./dist/index.js ./dist/index.mjs",
+        "compile": "tsc -p tsconfig.build.json && gen-esm-wrapper ./index.js ./index.mjs",
         "copy": "ts-node -T ../../scripts/copy.ts"
     },
     "publishConfig": {
@@ -59,9 +59,9 @@
         "@apify/pseudo_url": "^2.0.14",
         "@apify/timeout": "^0.3.0",
         "@apify/utilities": "^2.3.3",
-        "@crawlee/memory-storage": "^3.3.4-beta.16",
-        "@crawlee/types": "^3.3.4-beta.16",
-        "@crawlee/utils": "^3.3.4-beta.16",
+        "@crawlee/memory-storage": "^3.3.4-beta.17",
+        "@crawlee/types": "^3.3.4-beta.17",
+        "@crawlee/utils": "^3.3.4-beta.17",
         "@sapphire/async-queue": "^1.5.0",
         "@types/tough-cookie": "^4.0.2",
         "@vladfrangu/async_event_emitter": "^2.0.0",
@@ -77,5 +77,12 @@
         "tslib": "^2.4.0",
         "type-fest": "^3.0.0"
     },
-    "gitHead": "67a826c104e2e29a65ff89ccabaecf8b078f2c98"
+    "lerna": {
+        "command": {
+            "publish": {
+                "assets": []
+            }
+        }
+    },
+    "gitHead": "8bbc596f4dba0aa755df204d836c97c05fc94fca"
 }