pi-unsloth-webtools 0.9.0 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -38,13 +38,15 @@ Both tools display their target in the TUI tool row: `web_search "query"` and `w
38
38
 
39
39
  Mirrors Unsloth Studio's `web_search` tool:
40
40
 
41
- - Searches exactly like Studio's pinned `ddgs==9.14.4` `DDGS.text()`: the same seven engines
42
- (duckduckgo, brave, google, mojeek, yahoo, yandex, wikipedia; bing is disabled upstream),
43
- the same provider deduplication, href-dedupe aggregator with frequency ordering (hrefs are
41
+ - Searches like Studio's pinned `ddgs==9.14.4` `DDGS.text()`: the Studio engines that still work
42
+ here (duckduckgo, yandex; the others are behind bot walls, see
43
+ [Known differences from Studio](#known-differences-from-studio)) plus startpage,
44
+ the same provider deduplication, and an href-dedupe aggregator (hrefs are
44
45
  canonicalized first — `utm_*`/tracking parameters and fragments are dropped and the URL is
45
46
  re-serialized, collapsing host-case, default-port, and trailing-slash variants — so the
46
- same page found via different tracking links collapses), and the same `SimpleFilterRanker`
47
- re-ranking. Formats results identically: `Title:` / `URL:` /
47
+ same page found via different tracking links collapses). Results are fused with weighted
48
+ reciprocal-rank fusion: each engine's own ordering, not a keyword guess, decides relevance,
49
+ and the fused list is capped per registrable domain. Formats results identically:
48
50
  `Snippet:` blocks separated by `---`, ending with the hint to call `web_fetch` to
49
51
  read a full page.
50
52
  - Rate-limit, timeout, and empty-result messages mirror Studio's `_search_failure_message`.
@@ -53,6 +55,11 @@ Mirrors Unsloth Studio's `web_search` tool:
53
55
  skipped); timeouts and cancellations are never retried.
54
56
  - Sweeps stop as soon as enough results are gathered: engines still in flight are aborted
55
57
  instead of being allowed to run to their timeout.
58
+ - Engine requests go through the browser-fingerprint transport first (`webFetch.transport`,
59
+ default `tls-first`) and fall back to the plain Node transport; a refusal on one transport is
60
+ retried on the other. Redirects are followed manually, so every hop is re-checked against the
61
+ website policy and private-address literals are refused. When a SOCKS5 proxy is configured the
62
+ sweep stays on the plain transport, keeping agent and Tor routing intact.
56
63
 
57
64
  ### web_fetch
58
65
 
@@ -172,9 +179,10 @@ third-party rendering service.
172
179
  still gets its chance.
173
180
  - Lightpanda identifies itself honestly and refuses to impersonate a browser user agent, so hard
174
181
  anti-bot walls are returned as failures (the direct error, or the incomplete-content note).
175
- - Binary resolution: `lightpanda` on `PATH`, then `PI_LIGHTPANDA_BIN`, then
176
- `webRender.lightpandaPath`. Prebuilt binaries exist for Linux (glibc; musl needs a source
177
- build) and macOS, plus Docker images; Windows needs WSL2.
182
+ - Binary resolution: `webRender.lightpandaPath`, then `PI_LIGHTPANDA_BIN`, then the launcher
183
+ installed by `scripts/install-lightpanda.sh`, then `lightpanda` on `PATH`. Prebuilt binaries
184
+ exist for Linux (glibc; musl needs a source build) and macOS, plus Docker images; Windows
185
+ needs WSL2.
178
186
  - Version matters: 1.0.0 renders JavaScript-heavy pages that the 0.2.x line cannot — measured on
179
187
  the same machine, IMDb went from a 76-byte empty document to 21k characters, dribbble from 32
180
188
  characters to 18k, and a Medium article from a challenge page to real text. Linux builds from
@@ -186,7 +194,9 @@ third-party rendering service.
186
194
  - `bash scripts/install-lightpanda.sh` installs the newest stable release. When the system glibc
187
195
  predates what the binary needs, it downloads Debian's `libc6` for the current stable suite,
188
196
  extracts it into the install directory, and writes a launcher shim — so 1.0.0 runs on a
189
- glibc 2.36 host without touching the system libraries. The script prints the path to use.
197
+ glibc 2.36 host without touching the system libraries. It also writes
198
+ `webRender.lightpandaPath` into the global settings, so the extension picks the binary up with
199
+ no further setup; `--no-configure` skips that and `--print-path` prints the launcher path for CI.
190
200
  - Disable the tier with `webRender.lightpandaEnabled: false`; `web_fetch` then stops after the
191
201
  network attempts and reports the original error.
192
202
 
@@ -215,11 +225,31 @@ third-party rendering service.
215
225
  any numeric-only line at a fixed edge position where page numbers appear on at least
216
226
  half the pages (so a one-off number sharing that position is dropped too, while fused
217
227
  labels like `Page 3 of 12` survive). Studio and pymupdf4llm return them verbatim.
218
- - Search engines: Node's `fetch` TLS fingerprint differs from ddgs's `primp`
219
- impersonation, so Google/Brave/Yahoo/Yandex may block or serve consent pages more
220
- aggressively (a blocked engine simply contributes no results). User agents are a
221
- fixed browser set plus ddgs's Android Google UA generator, not `fake_useragent`'s
222
- database.
228
+ - Search engines: the sweep speaks a browser's TLS/HTTP2 shape through `wreq-js` first (the same
229
+ transport `web_fetch` uses), so engine fingerprints match Chrome rather than Node's `fetch`;
230
+ the plain Node transport is the fallback. User agents on that fallback are a fixed browser set,
231
+ not `fake_useragent`'s database.
232
+ - Startpage: an extra engine beyond Studio's set, matching newer ddgs (its Google-backed index
233
+ means it fills the `google` provider slot). It answers a plain `GET /sp/search?query=` and
234
+ serves an Anubis proof-of-work challenge to clients that do not look like a browser, so it works
235
+ through the browser-fingerprint transport and the fallback's browser headers, but not a bare
236
+ Node `fetch`. Startpage's POST endpoint and its safesearch parameter are challenge-gated and are
237
+ not used.
238
+ - Unused engines: mojeek, yahoo, google, brave and wikipedia are not part of the sweep. A
239
+ non-JavaScript client cannot get past mojeek (JavaScript challenge) or yahoo (`_bv` bot beacon),
240
+ google and brave answer rate limits and bot interstitials instead of results (HTTP 429 in
241
+ testing), and wikipedia is an API for a source the other engines already return — so the sweep
242
+ no longer spends a slot on it, and `wikipedia.org` results are ranked like any other instead of
243
+ being forced to the top. The remaining three engines can be narrowed further per machine with
244
+ `webSearch.engines` (see Configuration).
245
+ - Ranking: Studio's `SimpleFilterRanker` (keyword buckets with `wikipedia.org` pinned first) is
246
+ replaced by weighted reciprocal-rank fusion over each engine's own ordering. On the
247
+ `npm run engine:eval` query set that lifts mean precision@5 from 0.28 (keyword buckets) to 0.35 —
248
+ the fusion keeps the ranking signal the engines already computed instead of guessing from query
249
+ substrings. Weights default to uniform because the three engines overlap so little (mean Jaccard
250
+ 0.07–0.22) that weighting mostly decides which engine dominates the list rather than which result
251
+ is better. A per-domain cap exists (`webSearch.maxPerHost`) but is off by default: it measurably
252
+ costs precision, and the uncapped fused top-5 already spans 4.3 of 5 distinct domains.
223
253
  - Empty sweeps: ddgs 9.14.4 raises the last engine exception; this port reports a
224
254
  timeout whenever any engine timed out, so the timeout message is not masked by later
225
255
  generic engine failures. The timeout budget bounds the entire sweep: per-engine
@@ -228,16 +258,16 @@ third-party rendering service.
228
258
  - Proxies: Studio routes through environment proxies; this port resolves and pins the target IP and
229
259
  tunnels that connection through `HTTPS_PROXY` / `HTTP_PROXY` / `ALL_PROXY` when the proxy is a
230
260
  SOCKS5 proxy (`NO_PROXY` exclusions respected; DNS stays local for the guard). Other proxy
231
- schemes fall back to a direct connection. The search path uses the process-wide `fetch`, so an
232
- agent-level proxy dispatcher applies there too — see
233
- [Companion: rotating exit IPs](#companion-rotating-exit-ips).
261
+ schemes fall back to a direct connection. The search sweep stays on the process-wide `fetch`
262
+ whenever a SOCKS5 proxy is configured, so an agent-level proxy dispatcher still applies there —
263
+ see [Companion: rotating exit IPs](#companion-rotating-exit-ips).
234
264
  - Dedup and titles: the aggregator keys on canonicalized hrefs (`utm_*`/tracking parameters
235
265
  and fragments stripped, then the URL re-serialized); fetched HTML pages are prefixed with
236
266
  the document `<title>`. Studio keys on raw hrefs and returns the converted body alone.
237
267
  - Upstream drift: current ddgs ships ten backends (adding bing, startpage, grokipedia),
238
268
  requires a `vqd` token for DuckDuckGo, and exposes an `extract()` mode. This port
239
- deliberately pins the Studio snapshot — seven engines, bing disabled upstream, no vqd,
240
- no pagination — so engine behavior matches Studio rather than ddgs head.
269
+ uses duckduckgo and yandex from the Studio snapshot plus startpage — bing stays
270
+ disabled, no vqd, no pagination — so engine behavior matches Studio rather than ddgs head.
241
271
 
242
272
  ## When to use alternatives
243
273
 
@@ -263,8 +293,9 @@ choose the best tool per URL. No need to fork this package to add those features
263
293
 
264
294
  [`pi-tor-proxy`](https://github.com/YuGiMob/pi-tor-proxy) routes pi's in-process `fetch` traffic
265
295
  through Tor (it downloads and manages its own Tor binary) and gives each pi instance its own
266
- circuit and exit IP. The search sweep uses the process-wide `fetch`, so it leaves through the
267
- current Tor exit, and many search engines rate-limit or challenge per outgoing IP —
296
+ circuit and exit IP. With a SOCKS5 proxy configured, the search sweep stays on the process-wide
297
+ `fetch`, so it leaves through the current Tor exit, and many search engines rate-limit or
298
+ challenge per outgoing IP —
268
299
  `/tor-cycle` swaps the exit those limits are counted against, while `/tor-country` and
269
300
  `/tor-exclude` constrain which exits are used.
270
301
 
@@ -307,9 +338,12 @@ Optional settings in `~/.pi/agent/settings.json` or `.pi/settings.json` (project
307
338
  | `websitePolicy` | none | Not read from settings. Tools run unrestricted by default; `websitePolicy` is a programmatic option the host passes to `webSearch` / `fetchPageText` |
308
339
  | `unslothWebTools.allowPrivateAddresses` / `webFetch.allowPrivateAddresses` | `true` | Opt out to restore the resolved-IP SSRF guard: private/loopback/link-local hosts (localhost, LAN IPs) are refused again. Non-canonical numeric IP encodings stay blocked either way |
309
340
  | `unslothWebTools.allowLocalFiles` / `webFetch.allowLocalFiles` | `true` | Opt out to refuse local files in `web_fetch` (`file://` URLs, absolute, `~/`, or `./` paths); when enabled, PDFs are extracted and HTML converted |
310
- | `webFetch.transport` / `unslothWebTools.transport` | `tls-first` | Fetch transport order: `tls-first` (default), `direct-first`, or `off` to disable the browser-fingerprint transport entirely |
341
+ | `webFetch.transport` / `unslothWebTools.transport` | `tls-first` | Transport order for `web_fetch` and `web_search` engine requests: `tls-first` (default), `direct-first`, or `off` to disable the browser-fingerprint transport entirely |
342
+ | `webSearch.engines` / `unslothWebTools.engines` | all three (duckduckgo, yandex, startpage) | Restrict `web_search` to a subset of engine names, e.g. `["duckduckgo", "yandex"]`; unknown names are ignored, and a list that matches nothing falls back to every engine |
343
+ | `webSearch.engineWeights` / `unslothWebTools.engineWeights` | all `1` | Per-engine fusion weight, e.g. `{"startpage": 2, "yandex": 0.5}`; only positive numbers are read |
344
+ | `webSearch.maxPerHost` / `unslothWebTools.maxPerHost` | `0` (no cap) | Most results one registrable domain may contribute; `0` disables the cap |
311
345
  | `webRender.lightpandaEnabled` / `unslothWebTools.lightpandaEnabled` | `true` | Opt out to disable local Lightpanda rendering |
312
- | `webRender.lightpandaPath` / `unslothWebTools.lightpandaPath` | `lightpanda` on `PATH` (`PI_LIGHTPANDA_BIN` fallback) | Path to the Lightpanda binary used for local rendering |
346
+ | `webRender.lightpandaPath` / `unslothWebTools.lightpandaPath` | launcher installed by `scripts/install-lightpanda.sh`, else `lightpanda` on `PATH` | Path to the Lightpanda binary used for local rendering |
313
347
  | `webRender.lightpandaCommand` / `unslothWebTools.lightpandaCommand` | none | Command prefix that launches the renderer, for WSL (`["wsl.exe","-e","<path>"]`) or containers; overrides `lightpandaPath`. The fetch flags are appended to it |
314
348
 
315
349
  Environment overrides: `PI_UNSLOTH_CACHE_DIR` changes the fetch cache directory, `PI_UNSLOTH_WEBTOOLS_STATS` opts into append-only sweep stats JSONL, `PI_CODING_AGENT_DIR` / `PI_AGENT_DIR` change the global settings directory, and `PI_LIGHTPANDA_BIN` points at the local renderer binary. Cache entries live 1 hour and stale copies are served only after a network failure. SOCKS5 proxies named by `HTTPS_PROXY`, `HTTP_PROXY`, or `ALL_PROXY` are honored on every fetch (`NO_PROXY` exclusions apply).
@@ -354,36 +388,21 @@ npm test
354
388
  npm run test:unit
355
389
  npm run test:smoke
356
390
  bash scripts/install-lightpanda.sh
357
- npm run compare:fetch
358
- npm run compare:browsers
359
391
  npm run camoufox:warmup
360
- npm run stealth:matrix
361
- /tmp/pyenv/bin/python scripts/stealth-python.py
392
+ npm run engine:eval
362
393
  ```
363
394
 
364
- `npm run compare:fetch` runs a live head-to-head of the direct fetch, the TLS-impersonation retry,
365
- local Lightpanda rendering over a target list. It accepts URLs as arguments and `--no-lightpanda`
366
- to drop the render tier. `npm run compare:browsers` adds a Camoufox column;
367
- install it separately (`npm i camoufox-js playwright-core && npx camoufox-js fetch`, plus GTK3
368
- libraries on Linux) and use `--seconds=N` to bound how long it waits out a JS challenge.
369
- `npm run stealth:matrix` compares stealth-browser options against one walled page (plus
370
- `bot.sannysoft.com` detection rows and a plain-page sanity check): raw CDP to a system Chromium
371
- (`PI_CHROMIUM_BIN` to point at it), Playwright with its bundled Chromium, Patchright, and Camoufox.
372
- Every browser dependency is loaded through a guarded dynamic import, so nothing is added to
373
- `package.json`; install whichever rows you want to measure. `--attempts=N` and `--seconds=N` bound
374
- the walled-page attempts, and passing row names runs a subset (`raw-cdp`, `playwright`, `patchright`,
375
- `camoufox`).
376
-
377
- `scripts/stealth-python.py` is the same idea for the Python-side options (nodriver, CloakBrowser,
378
- DrissionPage, cloudscraper, curl_cffi) against the same walled page; it needs a venv with those
379
- packages installed and is not wired into any npm script. Measured findings are in its module docstring.
380
-
381
395
  `npm run camoufox:warmup` measures what a warm Camoufox costs and buys: launch time, idle CPU and
382
396
  RSS, per-fetch latency with the browser already running, and whether a persistent profile
383
397
  (`user_data_dir`, pinned fingerprint) lets a Cloudflare clearance survive a restart. Flags:
384
398
  `--virtual` for a virtual display, `--pin=0` to let Camoufox rotate fingerprints, `--seconds=N`,
385
399
  `--idle=N`.
386
400
 
401
+ `npm run engine:eval` measures each search engine against a small hand-labelled developer query
402
+ set: yield, latency and failures, pairwise URL overlap, unique contribution, and offline fusion
403
+ simulations with candidate weight vectors and per-host caps. `--cache=FILE` reuses a previous run
404
+ without touching the network; `--delay`, `--timeout`, `--max` and `--transport` tune the sweep.
405
+
387
406
  ## Tests
388
407
 
389
408
  `npm test` runs the full suite. `npm run test:unit` skips the live-network smoke tests,
@@ -404,7 +423,7 @@ The suite ports Unsloth Studio's own tests for these tools:
404
423
  - `test/fetch-flow.test.ts`: GitHub README rewrite, deadline/cancellation, HTML sniffing
405
424
  (from `test_web_fetch_extraction.py`; the fetch client is injected via seams)
406
425
  - `test/engines.test.ts`: the ddgs engine port, normalizers, the XPath subset, the
407
- aggregator, the ranker, and the Wikipedia engine with a stubbed fetch
426
+ reciprocal-rank aggregator, the per-host cap, and the Startpage engine with a stubbed fetch
408
427
  - `test/pdf-parity.test.ts`: MuPDF engine capabilities, PDF 1.5 object streams,
409
428
  ASCII85Decode, font `/Differences` encodings, pymupdf4llm-style headings/links/tables
410
429
  - `test/entities.test.ts`: `decodeHtmlEntities` parity with CPython `html.unescape`,