@octocrawl/cli 0.0.0-stage → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +661 -0
- package/README.md +20 -2
- package/dist/cli.js +29184 -0
- package/package.json +37 -4
package/README.md
CHANGED
|
@@ -1,3 +1,21 @@
|
|
|
1
|
-
#
|
|
1
|
+
# @octocrawl/cli
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
Octocrawl on the command line: scrape, crawl, batch and map web pages into Markdown, tables and JSON, with an **Evidence Record** on every page (final URL, fetch time, HTTP status, robots.txt decision, hashes of what was read and delivered). It runs the Octocrawl engine in its own process; nothing is sent to an Octocrawl server. The same CLI is published as `octocrawl`, so `npx octocrawl` works too.
|
|
4
|
+
|
|
5
|
+
```bash
|
|
6
|
+
npx octocrawl scrape https://books.toscrape.com/ --markdown
|
|
7
|
+
npx octocrawl scrape https://en.wikipedia.org/wiki/List_of_countries_by_GDP_(nominal) --formats markdown,tables --out gdp/
|
|
8
|
+
npx octocrawl batch --urls-file urls.txt --formats markdown,tables --out results/
|
|
9
|
+
npx octocrawl map https://www.sitemaps.org/ --limit 50
|
|
10
|
+
npx octocrawl serve --port 8787
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
Every option of the REST API is a flag under its kebab-case name (`maxAge` is `--max-age`, `onlyMainContent: false` is `--no-only-main-content`); `octocrawl <command> --help` lists them. `--out <dir>` writes `results.jsonl`, `results.csv` (one row of evidence per page, failed pages included), each page's Markdown and each table as CSV. The task root (tasks, saved files, the page cache) is `--task-root`, else `W2L_TASK_ROOT`, else `.w2l/cli`.
|
|
14
|
+
|
|
15
|
+
- Node.js 22.13 or later. The browser lane uses Playwright's Chromium: run `npx playwright install chromium` once; without it, pages are read over HTTP only.
|
|
16
|
+
- `better-sqlite3` builds or downloads its native binding when installed. If your npm holds install scripts back, allow it (`npm install-scripts approve better-sqlite3`, or `allowScripts` in your package.json).
|
|
17
|
+
- robots.txt is obeyed. By default pages are requested with the user agent and client hints of the Chrome the browser lane runs; `--mode research` names itself as a research crawler instead, with your contact from `W2L_CONTACT`. There is no stealth mode.
|
|
18
|
+
- Paid browser services (Browserbase, Steel) are used only when you name them in `W2L_VENDORS` (for example `W2L_VENDORS=browserbase`) and their key is set; a key alone does nothing.
|
|
19
|
+
- `octocrawl serve` listens on 127.0.0.1. On any other address it needs a token (`--token` or `W2L_API_TOKEN`), since other machines could otherwise use it to reach your localhost and network.
|
|
20
|
+
|
|
21
|
+
Licence: AGPL-3.0-only. Source, documentation and the API reference: https://github.com/77777R7/Octocrawl
|