pi-unsloth-webtools 0.2.1 → 0.2.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -26,7 +26,7 @@ Mirrors Unsloth Studio's `web_search` tool:
26
26
 
27
27
  - Searches exactly like Studio's pinned `ddgs==9.14.4` `DDGS.text()`: the same seven engines
28
28
  (duckduckgo, brave, google, mojeek, yahoo, yandex, wikipedia; bing is disabled upstream),
29
- the same provider-deduplication, href-dedupe aggregator with frequency ordering, and the
29
+ the same provider deduplication, href-dedupe aggregator with frequency ordering, and the
30
30
  same `SimpleFilterRanker` re-ranking. Formats results identically: `Title:` / `URL:` /
31
31
  `Snippet:` blocks separated by `---`, ending with the hint to pass `{"url": "<URL>"}` to
32
32
  read a full page.
@@ -45,7 +45,7 @@ Port of Studio's `_fetch_page_text` / `_fetch_url_raw` pipeline:
45
45
  validation and fetch.
46
46
  - GitHub repo root pages are rewritten to the unauthenticated README API
47
47
  (`Accept: application/vnd.github.raw+json`), falling back to the HTML page on failure.
48
- - Up to 5 redirect hops, each re-validated and re-resolved against the same rules.
48
+ - Up to 4 redirect hops, each re-validated and re-resolved against the same rules.
49
49
  - 512 KiB download cap (10 MiB for PDFs), overall deadline + per-hop socket timeouts, abort-aware
50
50
  (`signal` cancels mid-flight).
51
51
  - PDF text extraction via the official MuPDF.js engine (the same C library pymupdf wraps):
@@ -53,8 +53,8 @@ Port of Studio's `_fetch_page_text` / `_fetch_url_raw` pipeline:
53
53
  pymupdf4llm-style markdown layer (headings, bold/italic, code fences, links, tables)
54
54
  with Studio's corrupted/incomplete fallback to plain text.
55
55
  - Content sniffing: MIME allow/deny, binary magic signatures, PDF magic detection, and charset
56
- decoding (declared charset, BOM sniffing for UTF-8/16/32, cp1252 rescue for mislabeled
57
- single-byte pages).
56
+ decoding (declared charset, BOM sniffing for UTF-8/16/32, `<meta charset>` sniffing for
57
+ CJK and Windows/ISO encodings, cp1252 rescue for mislabeled single-byte pages).
58
58
  - HTML → Markdown conversion ported from Studio's dependency-free `_html_to_md.py`: headings,
59
59
  links, emphasis, lists, tables, blockquotes, code fences, entity decoding; hidden-element
60
60
  stripping (`hidden`, `aria-hidden`, inline styles); `<article>`/`<main>` main-content scoping
@@ -65,8 +65,22 @@ Port of Studio's `_fetch_page_text` / `_fetch_url_raw` pipeline:
65
65
  - HTML entity decoding replicates CPython's `html.unescape` (full 2,231-entry HTML5 table,
66
66
  longest-prefix rule, Windows-1252 numeric mappings), matching Studio byte-for-byte.
67
67
 
68
- See `ROADMAP.md` for the remaining gaps (per-line styling, table detection, TLS
69
- fingerprinting, proxies).
68
+ ## Known differences from Studio
69
+
70
+ - PDF styling: MuPDF.js exposes one font per line, so mixed-style lines style the
71
+ whole line instead of per-span; superscript, subscript, underline, strikeout, and
72
+ highlight markers are not emitted. Tables use a conservative text-grid detector:
73
+ aligned text tables are detected, drawn-rule-only tables are not.
74
+ - Search engines: Node's `fetch` TLS fingerprint differs from ddgs's `primp`
75
+ impersonation, so Google/Brave/Yahoo/Yandex may block or serve consent pages more
76
+ aggressively (a blocked engine simply contributes no results). User agents are a
77
+ fixed browser set plus ddgs's Android Google UA generator, not `fake_useragent`'s
78
+ database.
79
+ - Empty sweeps: ddgs 9.14.4 raises the last engine exception; this port reports a
80
+ timeout whenever any engine timed out, so the timeout message is not masked by later
81
+ generic engine failures.
82
+ - Proxies: Studio routes through environment proxies; this port always connects
83
+ directly with DNS pinning (deliberately out of scope).
70
84
 
71
85
  ## Development
72
86
 
@@ -74,27 +88,34 @@ fingerprinting, proxies).
74
88
  npm install
75
89
  npm run typecheck
76
90
  npm test
91
+ npm run test:unit
92
+ npm run test:smoke
77
93
  ```
78
94
 
79
95
  ## Tests
80
96
 
97
+ `npm test` runs the full suite. `npm run test:unit` skips the live-network smoke tests,
98
+ and `npm run test:smoke` runs only those.
99
+
81
100
  The suite ports Unsloth Studio's own tests for these tools:
82
101
 
83
- - `test/html-to-md.test.ts` — hidden-element stripping and main-content scoping (from
102
+ - `test/html-to-md.test.ts`: hidden-element stripping and main-content scoping (from
84
103
  `test_web_fetch_extraction.py`)
85
- - `test/header-strip.test.ts` — the header link-density suite plus article-vs-main selection
104
+ - `test/header-strip.test.ts`: the header link-density suite plus article-vs-main selection
86
105
  and boilerplate cases (from `test_web_fetch_extraction.py`)
87
- - `test/binary-guard.test.ts` — the MIME/magic/charset/PDF matrix (from
106
+ - `test/binary-guard.test.ts`: the MIME/magic/charset/PDF matrix (from
88
107
  `test_web_fetch_binary_guard.py`)
89
- - `test/web-search-policy.test.ts` — policy filtering, overfetch, and failure messages
108
+ - `test/web-search-policy.test.ts`: policy filtering, overfetch, and failure messages
90
109
  (from `test_web_access_policy.py`)
91
- - `test/fetch-flow.test.ts` — GitHub README rewrite, deadline/cancellation, HTML sniffing
110
+ - `test/fetch-flow.test.ts`: GitHub README rewrite, deadline/cancellation, HTML sniffing
92
111
  (from `test_web_fetch_extraction.py`; the fetch client is injected via seams)
93
- - `test/engines.test.ts` — the ddgs engine port: normalizers, the XPath subset, the
112
+ - `test/engines.test.ts`: the ddgs engine port, normalizers, the XPath subset, the
94
113
  aggregator, the ranker, and the Wikipedia engine with a stubbed fetch
95
- - `test/pdf-parity.test.ts` — MuPDF engine capabilities: PDF 1.5 object streams,
114
+ - `test/pdf-parity.test.ts`: MuPDF engine capabilities, PDF 1.5 object streams,
96
115
  ASCII85Decode, font `/Differences` encodings, pymupdf4llm-style headings/links/tables
97
- - `test/smoke.test.ts` — live network checks against real hosts
116
+ - `test/entities.test.ts`: `decodeHtmlEntities` parity with CPython `html.unescape`,
117
+ legacy refs, longest-prefix rule, Windows-1252 numeric mappings, invalid codepoints
118
+ - `test/smoke.test.ts`: live network checks against real hosts
98
119
 
99
120
  The seams (`seams.resolve` / `seams.request` / `rawFetch`) replace the network stack
100
121
  with fakes, mirroring how the Studio suite monkeypatches `_validate_and_resolve_host`
@@ -104,4 +125,4 @@ and `build_opener`.
104
125
 
105
126
  The ported logic derives from Unsloth Studio
106
127
  ([AGPL-3.0-only](https://github.com/unslothai/unsloth/blob/main/studio/LICENSE.AGPL-3.0)), so this
107
- package is released under the same **AGPL-3.0-only** license.
128
+ package is released under the same AGPL-3.0-only license.
package/engines.ts CHANGED
@@ -1,6 +1,7 @@
1
1
  import { randomBytes } from "node:crypto";
2
2
  import { decodeHtmlEntities, feedHtml } from "./html-to-md.ts";
3
3
  import type { AttrDict } from "./html-to-md.ts";
4
+ import { randomUserAgent } from "./user-agents.ts";
4
5
  export class EmptySweepError extends Error {
5
6
  constructor() {
6
7
  super("No results found");
@@ -110,6 +111,32 @@ interface XStep {
110
111
  preds: Pred[];
111
112
  terminal?: "text" | string;
112
113
  }
114
+ function parsePredicateBlocks(input: string, start: number): { preds: Pred[]; next: number } {
115
+ const preds: Pred[] = [];
116
+ let pos = start;
117
+ while (pos < input.length && input[pos] === "[") {
118
+ const innerStart = pos + 1;
119
+ let depth = 1;
120
+ let quote: string | null = null;
121
+ let j = innerStart;
122
+ while (j < input.length && depth) {
123
+ const c = input[j];
124
+ if (quote !== null) {
125
+ if (c === quote) quote = null;
126
+ } else if (c === "'" || c === '"') {
127
+ quote = c;
128
+ } else if (c === "[") {
129
+ depth++;
130
+ } else if (c === "]") {
131
+ depth--;
132
+ }
133
+ j++;
134
+ }
135
+ preds.push(parsePredExpr(input.slice(innerStart, j - 1)));
136
+ pos = j;
137
+ }
138
+ return { preds, next: pos };
139
+ }
113
140
 
114
141
  function parsePredExpr(input: string): Pred {
115
142
  let pos = 0;
@@ -173,35 +200,10 @@ function parsePredExpr(input: string): Pred {
173
200
  return { op: "desc", tag: name };
174
201
  }
175
202
  const name = word();
176
- const preds = parsePredBlocks();
203
+ const { preds, next } = parsePredicateBlocks(input, pos);
204
+ pos = next;
177
205
  return { op: "child", tag: name, preds };
178
206
  };
179
- const parsePredBlocks = (): Pred[] => {
180
- const preds: Pred[] = [];
181
- while (pos < input.length && input[pos] === "[") {
182
- const start = pos + 1;
183
- let depth = 1;
184
- let quote: string | null = null;
185
- let i = start;
186
- while (i < input.length && depth) {
187
- const c = input[i];
188
- if (quote !== null) {
189
- if (c === quote) quote = null;
190
- } else if (c === "'" || c === '"') {
191
- quote = c;
192
- } else if (c === "[") {
193
- depth++;
194
- } else if (c === "]") {
195
- depth--;
196
- }
197
- i++;
198
- }
199
- const inner = input.slice(start, i - 1);
200
- preds.push(parsePredExpr(inner));
201
- pos = i;
202
- }
203
- return preds;
204
- };
205
207
  const parseAnd = (): Pred => {
206
208
  let left = atom();
207
209
  while (true) {
@@ -274,28 +276,8 @@ function parsePath(expr: string): XStep[] {
274
276
  if (!m) break;
275
277
  const name = m[0];
276
278
  i += m[0].length;
277
- const preds: Pred[] = [];
278
- while (i < expr.length && expr[i] === "[") {
279
- const start = i + 1;
280
- let depth = 1;
281
- let quote: string | null = null;
282
- let j = start;
283
- while (j < expr.length && depth) {
284
- const c = expr[j];
285
- if (quote !== null) {
286
- if (c === quote) quote = null;
287
- } else if (c === "'" || c === '"') {
288
- quote = c;
289
- } else if (c === "[") {
290
- depth++;
291
- } else if (c === "]") {
292
- depth--;
293
- }
294
- j++;
295
- }
296
- preds.push(parsePredExpr(expr.slice(start, j - 1)));
297
- i = j;
298
- }
279
+ const { preds, next } = parsePredicateBlocks(expr, i);
280
+ i = next;
299
281
  steps.push({ axis, name, preds });
300
282
  }
301
283
  return steps;
@@ -409,19 +391,6 @@ export function extractResults(
409
391
  return results;
410
392
  }
411
393
 
412
- const USER_AGENTS = [
413
- "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36",
414
- "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36",
415
- "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36",
416
- "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:133.0) Gecko/20100101 Firefox/133.0",
417
- "Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:133.0) Gecko/20100101 Firefox/133.0",
418
- "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/18.2 Safari/605.1.15",
419
- ];
420
-
421
- function randomUserAgent(): string {
422
- return USER_AGENTS[Math.floor(Math.random() * USER_AGENTS.length)];
423
- }
424
-
425
394
  function googleUserAgent(): string {
426
395
  const devices: [string, string, number, number][] = [
427
396
  ["5.0", "SM-G900P Build/LRX21T", 39, 60],
@@ -822,7 +791,7 @@ export async function autoTextSearch(
822
791
  const seenProviders = new Set<string>();
823
792
  const aggregator = new ResultsAggregator();
824
793
  const ctx: EngineContext = { region: "us-en", safesearch: "moderate" };
825
- let err: unknown = null;
794
+ let timedOut = false;
826
795
  const uniqueProviders = new Set(engines.map((e) => e.provider)).size;
827
796
  const maxWorkers = Math.min(uniqueProviders, Math.ceil(maxResults / 10) + 1);
828
797
  let i = 0;
@@ -835,7 +804,7 @@ export async function autoTextSearch(
835
804
  seenProviders.add(engine.provider);
836
805
  }
837
806
  } catch (e) {
838
- err = e;
807
+ if (e instanceof Error && e.message.includes("timed out")) timedOut = true;
839
808
  }
840
809
  };
841
810
  while (i < engines.length) {
@@ -843,13 +812,14 @@ export async function autoTextSearch(
843
812
  const engine = engines[i++];
844
813
  if (seenProviders.has(engine.provider)) continue;
845
814
  pending.push(run(engine));
846
- if (pending.length >= maxWorkers || i >= maxWorkers) {
815
+ if (pending.length >= maxWorkers) {
847
816
  await Promise.allSettled(pending);
848
817
  pending = [];
849
818
  }
850
819
  }
820
+ await Promise.allSettled(pending);
851
821
  const results = rankResults(aggregator.extractDicts(), query);
852
822
  if (results.length) return results.slice(0, maxResults);
853
- if (err instanceof Error && err.message.includes("timed out")) throw new SearchTimeoutError();
823
+ if (timedOut) throw new SearchTimeoutError();
854
824
  throw new EmptySweepError();
855
825
  }
package/html-to-md.ts CHANGED
@@ -197,7 +197,7 @@ export function decodeHtmlEntities(text: string): string {
197
197
  return text.replace(CHARREF_RE, (whole, s: string) => {
198
198
  if (s[0] === "#") {
199
199
  const hex = s[1] === "x" || s[1] === "X";
200
- const num = parseInt(s.slice(2).replace(/;+$/, ""), hex ? 16 : 10);
200
+ const num = parseInt(s.slice(hex ? 2 : 1).replace(/;+$/, ""), hex ? 16 : 10);
201
201
  const mapped = INVALID_CHARREFS[num];
202
202
  if (mapped !== undefined) return mapped;
203
203
  if ((num >= 0xd800 && num <= 0xdfff) || num > 0x10ffff) return "\ufffd";
package/index.ts CHANGED
@@ -62,12 +62,12 @@ export default function (pi: ExtensionAPI) {
62
62
  name: "web_fetch",
63
63
  label: "Web Fetch",
64
64
  description:
65
- "Fetch a URL and return readable text content. HTML responses are converted to Markdown " +
66
- "with a main-content heuristic (article/main scoping, hidden-element and boilerplate " +
67
- "stripping); non-HTML text responses are returned as-is. GitHub repo root pages are " +
68
- "rewritten to the README API so the model reads the README instead of the repo page's UI " +
69
- "chrome. Blocks private/loopback/link-local targets (SSRF protection) and caps the " +
70
- "download size.",
65
+ "Fetch a URL and return its readable text. HTML pages are converted to Markdown using a " +
66
+ "main-content heuristic: article/main scoping plus hidden-element and boilerplate " +
67
+ "stripping. Non-HTML text is returned as-is. GitHub repo root pages are rewritten to the " +
68
+ "README API, so the README is returned instead of the repo page's UI chrome. " +
69
+ "Private/loopback/link-local targets are blocked (SSRF protection), and the download size " +
70
+ "is capped.",
71
71
  promptSnippet: "Fetch a web page and return readable text content",
72
72
  parameters: WebFetchParams,
73
73
  async execute(_toolCallId, params, signal, _onUpdate, _ctx) {
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-unsloth-webtools",
3
- "version": "0.2.1",
3
+ "version": "0.2.3",
4
4
  "type": "module",
5
5
  "description": "Pi extension: web_search and web_fetch tools ported from the Unsloth Studio codebase (DuckDuckGo search, SSRF-safe direct fetching, HTML-to-Markdown extraction)",
6
6
  "main": "index.ts",
@@ -30,8 +30,8 @@
30
30
  "engines.ts",
31
31
  "entities.ts",
32
32
  "pdf.ts",
33
+ "user-agents.ts",
33
34
  "README.md",
34
- "ROADMAP.md",
35
35
  "LICENSE"
36
36
  ],
37
37
  "pi": {
@@ -48,8 +48,10 @@
48
48
  },
49
49
  "scripts": {
50
50
  "test": "vitest run",
51
+ "test:unit": "vitest run --exclude **/smoke.test.ts",
52
+ "test:smoke": "vitest run test/smoke.test.ts",
51
53
  "typecheck": "tsc --noEmit",
52
- "prepublishOnly": "npm run typecheck"
54
+ "prepublishOnly": "npm run typecheck && npm run test:unit"
53
55
  },
54
56
  "devDependencies": {
55
57
  "@earendil-works/pi-coding-agent": "^0.84.0",
package/pdf.ts CHANGED
@@ -527,7 +527,7 @@ function assemblePages(
527
527
  return "";
528
528
  }
529
529
  if (pageLimitReached) {
530
- text += `\n\n... (PDF extraction page processing capped at ${MAX_WEB_PDF_PAGES} pages)`;
530
+ text += `\n\n... (PDF extraction is capped at ${MAX_WEB_PDF_PAGES} pages)`;
531
531
  }
532
532
  return text;
533
533
  }
package/user-agents.ts ADDED
@@ -0,0 +1,12 @@
1
+ const USER_AGENTS = [
2
+ "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36",
3
+ "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36",
4
+ "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36",
5
+ "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:133.0) Gecko/20100101 Firefox/133.0",
6
+ "Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:133.0) Gecko/20100101 Firefox/133.0",
7
+ "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/18.2 Safari/605.1.15",
8
+ ];
9
+
10
+ export function randomUserAgent(): string {
11
+ return USER_AGENTS[Math.floor(Math.random() * USER_AGENTS.length)];
12
+ }
package/web-access.ts CHANGED
@@ -173,46 +173,46 @@ export function checkUrlAccess(
173
173
  policy: WebsitePolicy | null,
174
174
  ): [boolean, string, string] {
175
175
  if (typeof url !== "string" || !url.trim()) {
176
- return [false, "Blocked: URL is empty.", ""];
176
+ return [false, "Blocked: the URL is empty.", ""];
177
177
  }
178
178
  const candidate = url.trim();
179
179
  if (
180
180
  Array.from(candidate).some((char) => /\s/.test(char) || char.charCodeAt(0) < 32) ||
181
181
  candidate.includes("\\")
182
182
  ) {
183
- return [false, "Blocked: URL contains invalid characters.", ""];
183
+ return [false, "Blocked: the URL contains invalid characters.", ""];
184
184
  }
185
185
  let parsed: URL;
186
186
  try {
187
187
  parsed = new URL(candidate);
188
188
  } catch {
189
- return [false, "Blocked: URL has an invalid hostname or port.", ""];
189
+ return [false, "Blocked: the URL has an invalid hostname or port.", ""];
190
190
  }
191
191
  const scheme = parsed.protocol.replace(/:$/, "").toLowerCase();
192
192
  if (scheme !== "http" && scheme !== "https") {
193
193
  return [false, "Blocked: only http/https URLs are allowed.", ""];
194
194
  }
195
195
  if (parsed.username || parsed.password || parsed.hostname.includes("%")) {
196
- return [false, "Blocked: URL credentials or encoded hostnames are not allowed.", ""];
196
+ return [false, "Blocked: URLs with credentials or encoded hostnames are not allowed.", ""];
197
197
  }
198
198
  if (!parsed.hostname) {
199
- return [false, "Blocked: URL has an invalid hostname or port.", ""];
199
+ return [false, "Blocked: the URL has an invalid hostname or port.", ""];
200
200
  }
201
201
  try {
202
202
  if (parsed.port && !(PORT_RE.test(parsed.port) && Number(parsed.port) >= 1 && Number(parsed.port) <= 65535)) {
203
- return [false, "Blocked: URL has an invalid hostname or port.", ""];
203
+ return [false, "Blocked: the URL has an invalid hostname or port.", ""];
204
204
  }
205
205
  } catch {
206
- return [false, "Blocked: URL has an invalid hostname or port.", ""];
206
+ return [false, "Blocked: the URL has an invalid hostname or port.", ""];
207
207
  }
208
208
  let hostname: string;
209
209
  try {
210
210
  hostname = normalizeDomain(parsed.hostname);
211
211
  } catch {
212
- return [false, "Blocked: URL has an invalid hostname or port.", ""];
212
+ return [false, "Blocked: the URL has an invalid hostname or port.", ""];
213
213
  }
214
214
  if (!hostnameAllowed(hostname, policy)) {
215
- return [false, `Blocked: website access policy disallows ${hostname}.`, hostname];
215
+ return [false, `Blocked: the website access policy disallows ${hostname}.`, hostname];
216
216
  }
217
217
  return [true, "", hostname];
218
218
  }
@@ -368,6 +368,8 @@ export function isPublicIp(ip: string): boolean {
368
368
  if (lower.startsWith("2001:db8")) return false;
369
369
  if (lower.startsWith("64:ff9b:")) return false;
370
370
  if (lower.startsWith("2001:10:")) return false;
371
+ if (lower.startsWith("2002:")) return false;
372
+ if (lower.startsWith("2001:0:") || lower.startsWith("2001::")) return false;
371
373
  const mapped = /^::ffff:(\d+\.\d+\.\d+\.\d+)$/.exec(lower);
372
374
  if (mapped) return isPublicIp(mapped[1]);
373
375
  if (lower.startsWith("::ffff:")) return false;
package/web-fetch.ts CHANGED
@@ -11,20 +11,11 @@ import {
11
11
  } from "./web-access.ts";
12
12
  import { htmlToMarkdown } from "./html-to-md.ts";
13
13
  import { extractPdfText, PdfParseError } from "./pdf.ts";
14
+ import { randomUserAgent } from "./user-agents.ts";
14
15
 
15
- const MIN_PAGE_CHARS = 2000;
16
16
  const MAX_FETCH_BYTES = 512 * 1024;
17
17
  const MAX_PDF_FETCH_BYTES = 10 * 1024 * 1024;
18
- const MAX_REDIRECTS = 5;
19
-
20
- const USER_AGENTS = [
21
- "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36",
22
- "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36",
23
- "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36",
24
- "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:133.0) Gecko/20100101 Firefox/133.0",
25
- "Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:133.0) Gecko/20100101 Firefox/133.0",
26
- "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/18.2 Safari/605.1.15",
27
- ];
18
+ const MAX_REQUESTS = 5;
28
19
 
29
20
  const UTF32_LE_BOM = Buffer.from([0xff, 0xfe, 0x00, 0x00]);
30
21
  const UTF32_BE_BOM = Buffer.from([0x00, 0x00, 0xfe, 0xff]);
@@ -237,6 +228,34 @@ function hasSingleByteTextEvidence(data: Buffer): boolean {
237
228
  return ascii / data.length >= MIN_SINGLE_BYTE_ASCII_RATIO;
238
229
  }
239
230
 
231
+ const CHARSET_ALIASES: Record<string, string> = {
232
+ gbk: "gbk",
233
+ gb2312: "gbk",
234
+ "gb-2312": "gbk",
235
+ cp936: "gbk",
236
+ "x-gbk": "gbk",
237
+ gb18030: "gb18030",
238
+ big5: "big5",
239
+ "big5-hkscs": "big5",
240
+ sjis: "shift_jis",
241
+ "x-sjis": "shift_jis",
242
+ cp932: "shift_jis",
243
+ "shift-jis": "shift_jis",
244
+ "euc-jp": "euc-jp",
245
+ "euc-kr": "euc-kr",
246
+ ksc5601: "euc-kr",
247
+ "ks_c_5601-1987": "euc-kr",
248
+ "ks_c_5601-1989": "euc-kr",
249
+ "iso-2022-jp": "iso-2022-jp",
250
+ "koi8-r": "koi8-r",
251
+ "koi8-u": "koi8-u",
252
+ cp866: "cp866",
253
+ "x-mac-cyrillic": "x-mac-cyrillic",
254
+ "windows-874": "windows-874",
255
+ cp874: "windows-874",
256
+ "tis-620": "tis-620",
257
+ };
258
+
240
259
  function normalizeCharset(name: string): string | null {
241
260
  const n = name.trim().replace(/["']/g, "").toLowerCase();
242
261
  switch (n) {
@@ -270,8 +289,28 @@ function normalizeCharset(name: string): string | null {
270
289
  case "utf-32be":
271
290
  return "utf-32be";
272
291
  default:
273
- return null;
292
+ break;
293
+ }
294
+ const alias = CHARSET_ALIASES[n];
295
+ if (alias !== undefined) return alias;
296
+ if (/^windows-125[0-8]$/.test(n) || /^iso-8859-(?:[2-9]|1[0-6])$/.test(n)) return n;
297
+ return null;
298
+ }
299
+
300
+ function sniffMetaCharset(bytes: Buffer): string | null {
301
+ const head = bytes.subarray(0, 1024).toString("latin1").toLowerCase();
302
+ const match =
303
+ /<meta\b[^>]*\bcharset\s*=\s*["']?\s*([a-z0-9_.\-]+)/.exec(head) ??
304
+ /<meta\b[^>]*\bhttp-equiv\s*=\s*["']?content-type["']?[^>]*\bcharset\s*=\s*["']?\s*([a-z0-9_.\-]+)/.exec(head);
305
+ return match ? normalizeCharset(match[1]) : null;
306
+ }
307
+
308
+ function sniffMetaCharsetForHtml(bytes: Buffer, contentType: string): string | null {
309
+ if (!contentType.includes("html")) {
310
+ const probe = bytes.subarray(0, 256).toString("latin1");
311
+ if (!looksLikeHtmlDocument(probe)) return null;
274
312
  }
313
+ return sniffMetaCharset(bytes);
275
314
  }
276
315
 
277
316
  function decodeUtf32(bytes: Buffer, littleEndian: boolean): string {
@@ -328,11 +367,38 @@ function decodeWithCodec(bytes: Buffer, codec: string | null): string {
328
367
  case "cp1252":
329
368
  return decodeSingleByte(bytes, true);
330
369
  case "iso8859-1":
331
- default:
332
370
  return decodeSingleByte(bytes, false);
371
+ default:
372
+ return decodeWithLabel(bytes, codec);
333
373
  }
334
374
  }
335
375
 
376
+ function decodeWithLabel(bytes: Buffer, label: string | null): string {
377
+ if (label === null) return decodeSingleByte(bytes, false);
378
+ if (label === "tis-620") return decodeTis620(bytes);
379
+ try {
380
+ return new TextDecoder(label, { fatal: false }).decode(bytes);
381
+ } catch {
382
+ return decodeSingleByte(bytes, false);
383
+ }
384
+ }
385
+
386
+ function decodeTis620(bytes: Buffer): string {
387
+ let out = "";
388
+ for (const byte of bytes) {
389
+ if (byte < 0x80) {
390
+ out += String.fromCharCode(byte);
391
+ } else if (byte >= 0xa1 && byte <= 0xfb) {
392
+ out += String.fromCodePoint(0x0e01 + byte - 0xa1);
393
+ } else if (byte === 0xa0) {
394
+ out += "\u00a0";
395
+ } else {
396
+ out += "\ufffd";
397
+ }
398
+ }
399
+ return out;
400
+ }
401
+
336
402
 
337
403
  async function resolveAndValidate(hostname: string, signal?: AbortSignal): Promise<ResolvedHost> {
338
404
  let addresses: { address: string; family: number }[];
@@ -346,7 +412,7 @@ async function resolveAndValidate(hostname: string, signal?: AbortSignal): Promi
346
412
  }
347
413
  for (const entry of addresses) {
348
414
  if (!isPublicIp(entry.address)) {
349
- return { ok: false, reason: `Blocked: refusing to fetch non-public address ${entry.address}.`, ip: "", family: 0 };
415
+ return { ok: false, reason: `Blocked: refusing to fetch the non-public address ${entry.address}.`, ip: "", family: 0 };
350
416
  }
351
417
  }
352
418
  const first = addresses[0];
@@ -386,20 +452,19 @@ function requestHop(opts: HopOptions): Promise<HopResponse> {
386
452
  const declaredPdf = String(res.headers["content-type"] ?? "").toLowerCase().includes("pdf");
387
453
  let limit = declaredPdf ? opts.maxPdfBytes : opts.maxBytes;
388
454
  let extendedForPdf = false;
389
- let done = false;
390
455
  const finish = (err: string | null, body: Buffer) => {
391
- if (done) return;
392
- done = true;
393
- if (err) reject(new Error(err));
394
- else
395
- resolve({
396
- status: res.statusCode ?? 0,
397
- headers: res.headers as Record<string, string | string[] | undefined>,
398
- body,
399
- });
456
+ settle(() => {
457
+ if (err) reject(new Error(err));
458
+ else
459
+ resolve({
460
+ status: res.statusCode ?? 0,
461
+ headers: res.headers as Record<string, string | string[] | undefined>,
462
+ body,
463
+ });
464
+ });
400
465
  };
401
466
  res.on("data", (chunk: Buffer) => {
402
- if (done) return;
467
+ if (settled) return;
403
468
  if (!declaredPdf && !extendedForPdf && total + chunk.length > opts.maxBytes) {
404
469
  if (hasPdfMagic(Buffer.concat(chunks))) {
405
470
  limit = opts.maxPdfBytes;
@@ -423,10 +488,17 @@ function requestHop(opts: HopOptions): Promise<HopResponse> {
423
488
  res.on("end", () => finish(null, Buffer.concat(chunks)));
424
489
  res.on("error", (err) => finish(err.message, Buffer.concat(chunks)));
425
490
  });
426
- request.on("timeout", () => request.destroy(new Error("timed out")));
427
- request.on("error", (err) => reject(err));
428
491
  const onAbort = () => request.destroy(new Error("cancelled"));
429
492
  opts.signal?.addEventListener("abort", onAbort, { once: true });
493
+ let settled = false;
494
+ const settle = (action: () => void) => {
495
+ if (settled) return;
496
+ settled = true;
497
+ opts.signal?.removeEventListener("abort", onAbort);
498
+ action();
499
+ };
500
+ request.on("timeout", () => request.destroy(new Error("timed out")));
501
+ request.on("error", (err) => settle(() => reject(err)));
430
502
  request.end();
431
503
  });
432
504
  }
@@ -460,9 +532,9 @@ export async function fetchUrlRaw(
460
532
  let currentUrl = url;
461
533
  let pinnedIp = resolved.ip;
462
534
  let pinnedFamily = resolved.family;
463
- const userAgent = USER_AGENTS[Math.floor(Math.random() * USER_AGENTS.length)];
535
+ const userAgent = randomUserAgent();
464
536
 
465
- for (let hop = 0; hop < MAX_REDIRECTS; hop++) {
537
+ for (let hop = 0; hop < MAX_REQUESTS; hop++) {
466
538
  budgetError = fetchBudgetExceeded(deadline, signal, now);
467
539
  if (budgetError !== null) return { error: budgetError, body: "", contentType: "" };
468
540
  const parsed = new URL(currentUrl);
@@ -471,6 +543,7 @@ export async function fetchUrlRaw(
471
543
  "User-Agent": userAgent,
472
544
  Host: hostHeader,
473
545
  Accept: "text/html,application/xhtml+xml,text/plain;q=0.9,*/*;q=0.5",
546
+ "Accept-Encoding": "identity",
474
547
  };
475
548
  if (options.extraHeaders) Object.assign(headers, options.extraHeaders);
476
549
  const inactivity = Math.max(1, deadline - now());
@@ -497,8 +570,9 @@ export async function fetchUrlRaw(
497
570
 
498
571
  if (response.status >= 300 && response.status < 400) {
499
572
  if (![301, 302, 303, 307, 308].includes(response.status)) {
573
+ const reason = http.STATUS_CODES[response.status] ?? "";
500
574
  return {
501
- error: `Failed to fetch URL: HTTP ${response.status} ${statusReason(response.status)}`,
575
+ error: `Failed to fetch URL: HTTP ${response.status}${reason ? ` ${reason}` : ""}`,
502
576
  body: "",
503
577
  contentType: "",
504
578
  };
@@ -507,12 +581,16 @@ export async function fetchUrlRaw(
507
581
  const location = Array.isArray(rawLocation) ? rawLocation[0] : rawLocation;
508
582
  if (!location) {
509
583
  return {
510
- error: "Failed to fetch URL: redirect missing Location header.",
584
+ error: "Failed to fetch URL: the redirect is missing a Location header.",
511
585
  body: "",
512
586
  contentType: "",
513
587
  };
514
588
  }
515
- currentUrl = new URL(location, currentUrl).toString();
589
+ try {
590
+ currentUrl = new URL(location, currentUrl).toString();
591
+ } catch {
592
+ return { error: "Failed to fetch URL: the redirect has an invalid Location.", body: "", contentType: "" };
593
+ }
516
594
  const [redirectAllowed, redirectReason, redirectHost] = checkUrlAccess(
517
595
  currentUrl,
518
596
  policy,
@@ -577,7 +655,10 @@ export async function fetchUrlRaw(
577
655
 
578
656
  const declaredCodec = declaredCharset ? normalizeCharset(declaredCharset) : null;
579
657
  const bomCodec = bomCodecFor(response.body);
580
- const rawHtml = decodeWithCodec(response.body, declaredCodec ?? bomCodec ?? "utf-8");
658
+ const rawHtml = decodeWithCodec(
659
+ response.body,
660
+ declaredCodec ?? bomCodec ?? sniffMetaCharsetForHtml(response.body, contentType) ?? "utf-8",
661
+ );
581
662
 
582
663
  if (looksBinary(rawHtml)) {
583
664
  let alt: string | null = null;
@@ -612,13 +693,6 @@ function bomCodecFor(bytes: Buffer): string | null {
612
693
  return null;
613
694
  }
614
695
 
615
- function statusReason(status: number): string {
616
- const reasons: Record<number, string> = {
617
- 301: "Moved Permanently", 302: "Found", 303: "See Other", 307: "Temporary Redirect", 308: "Permanent Redirect",
618
- };
619
- return reasons[status] ?? "";
620
- }
621
-
622
696
  export function truncatePageText(text: string, maxChars?: number): string {
623
697
  if (!text) return "(page returned no readable text)";
624
698
  if (typeof maxChars === "number" && maxChars > 0 && text.length > maxChars) {
package/web-search.ts CHANGED
@@ -64,23 +64,17 @@ export async function webSearch(
64
64
  const results = await client(effectiveQuery, wanted, signal);
65
65
  if (signal?.aborted) return "Search cancelled.";
66
66
  if (!results.length) return EMPTY_SEARCH_RESULTS[0];
67
- const parts: string[] = [];
67
+ const allowed: SearchResult[] = [];
68
68
  for (const result of results) {
69
- if (parts.length >= maxResults) break;
69
+ if (allowed.length >= maxResults) break;
70
70
  const href = String(result.href ?? "").trim();
71
71
  if (href && !checkUrlAccess(href, policy)[0]) continue;
72
- const title = String(result.title ?? "").replace(/\s+/g, " ");
73
- const snippet = String(result.body ?? "").replace(/\s+/g, " ");
74
- parts.push(`Title: ${title}\nURL: ${href}\nSnippet: ${snippet}`);
72
+ allowed.push(result);
75
73
  }
76
- if (!parts.length) return EMPTY_SEARCH_RESULTS[1];
77
- const text = parts.join("\n\n---\n\n");
78
- return (
79
- text +
80
- "\n\n---\n\nIMPORTANT: These are only short snippets. " +
81
- 'To get the full page content, call web_search with the url parameter (e.g. {"url": "<URL>"}).'
82
- );
74
+ if (!allowed.length) return EMPTY_SEARCH_RESULTS[1];
75
+ return formatSearchResults(allowed);
83
76
  } catch (err) {
77
+ if (signal?.aborted) return "Search cancelled.";
84
78
  return searchFailureMessage(err, timeoutMs);
85
79
  }
86
80
  }
@@ -88,7 +82,7 @@ export async function webSearch(
88
82
  export function searchFailureMessage(exc: unknown, timeoutMs = SEARCH_TIMEOUT_MS): string {
89
83
  if (exc instanceof SearchCancelled) return "Search cancelled.";
90
84
  if (exc instanceof SearchTimeoutError) {
91
- return `Search failed: the search engines did not respond within ${Math.round(timeoutMs / 1000)}s.`;
85
+ return `Search failed: the search engines did not respond within ${Math.round(timeoutMs / 1000)} seconds.`;
92
86
  }
93
87
  if (exc instanceof EmptySweepError || (exc instanceof Error && exc.message.includes("No results found"))) {
94
88
  return EMPTY_SEARCH_RESULTS[0];
@@ -98,15 +92,15 @@ export function searchFailureMessage(exc: unknown, timeoutMs = SEARCH_TIMEOUT_MS
98
92
 
99
93
  export function formatSearchResults(results: SearchResult[]): string {
100
94
  const parts = results.map((result) => {
101
- const title = result.title.replace(/\s+/g, " ");
102
- const href = result.href.trim();
103
- const snippet = result.body.replace(/\s+/g, " ");
95
+ const title = String(result.title ?? "").replace(/\s+/g, " ");
96
+ const href = String(result.href ?? "").trim();
97
+ const snippet = String(result.body ?? "").replace(/\s+/g, " ");
104
98
  return `Title: ${title}\nURL: ${href}\nSnippet: ${snippet}`;
105
99
  });
106
100
  const text = parts.join("\n\n---\n\n");
107
101
  return (
108
102
  text +
109
- "\n\n---\n\nIMPORTANT: These are only short snippets. " +
103
+ "\n\n---\n\nThese are only short snippets. " +
110
104
  'To get the full page content, call web_search with the url parameter (e.g. {"url": "<URL>"}).'
111
105
  );
112
106
  }
package/ROADMAP.md DELETED
@@ -1,77 +0,0 @@
1
- # Roadmap
2
-
3
- Tracking the remaining gaps between this port and Unsloth Studio's web tools, and the
4
- planned work to close them.
5
-
6
- ## Implemented
7
-
8
- ### PDF text extraction (MuPDF engine)
9
-
10
- `pdf.ts` uses the official **MuPDF.js** (`mupdf` npm package) — the same C engine that
11
- PyMuPDF wraps — replacing the earlier minimal built-in extractor:
12
-
13
- - Full xref handling: tables, cross-reference streams, and PDF 1.5+ object streams
14
- - All standard stream filters (FlateDecode, ASCII85Decode, LZW, RunLength, DCT, JPX, ...)
15
- - Font encodings and ToUnicode mapping (non-Latin text extracts correctly)
16
- - Encryption detection via `needsPassword()` (reported as unreadable, matching pymupdf's
17
- no-password behavior)
18
- - The markdown layer replicates pymupdf4llm's algorithm: `IdentifyHeaders` font-size
19
- heading detection, `get_raw_lines` line reconstruction (tolerance 3, 10% span-join
20
- delta), `write_text` styling (bold/italic/mono, code fences, bullets, link
21
- resolution with `%0x`-escaped URIs), and Studio's corrupted/incomplete fallback to
22
- plain MuPDF text with the exact thresholds from `backend/core/rag/parsers.py`
23
- - Table detection and pipe-markdown rendering matching pymupdf's `Table.to_markdown`
24
- shape (`|header|`, `|---|`, detail rows, `Col{i}` fill for empty headers)
25
-
26
- Known deltas vs pymupdf4llm:
27
-
28
- - Span-level styling: MuPDF.js's structured-text JSON exposes one font per line, so
29
- mixed-style lines (one bold word inside a body line) style the whole line instead of
30
- per-span. Line-level styling matches for homogeneous lines.
31
- - Superscript/subscript/underline/strikeout/highlight markers are not emitted (the
32
- JSON does not expose char-level flags).
33
- - Table detection is a conservative text-grid detector (column-start clustering with
34
- a 5 pt tolerance, contiguous multi-row bands) instead of PyMuPDF's
35
- `find_tables()` vector-graphics analysis. Aligned text tables are detected; tables
36
- defined only by drawn rules without aligned text are not.
37
-
38
- The minimal extractor remains as an automatic fallback when the `mupdf` package cannot
39
- be loaded (for example a stripped install).
40
-
41
- ## Not planned (explicit decisions)
42
-
43
- ### Proxy support
44
-
45
- Studio routes requests through urllib's environment proxies and honors
46
- `UNSLOTH_STUDIO_DISABLE_DNS_PINNING` for enterprise proxies. This port always connects
47
- directly with DNS pinning. Deliberately out of scope.
48
-
49
- ### Page size budgets
50
-
51
- Studio's `_page_char_budget()` sizes fetched pages to the serving model's context window.
52
- This port deliberately has **no character budget at all**: fetched pages and PDFs are
53
- returned in full (the PDF page-count cap of 50 pages remains, matching Studio). The raw
54
- download caps (512 KiB text / 10 MiB PDF) still bound what is fetched. A `maxChars`
55
- parameter remains available on the tool for callers that want to truncate.
56
-
57
- ## Known behavioral differences
58
-
59
- ### Search engine set
60
-
61
- The port implements ddgs 9.14.4's `DDGS.text()` exactly: the same seven engines
62
- (duckduckgo, brave, google, mojeek, yahoo, yandex, wikipedia — bing is `disabled` in
63
- ddgs upstream), the same provider-deduplication, href-dedupe aggregator with
64
- frequency ordering, and the same `SimpleFilterRanker` re-ranking. Remaining deltas:
65
-
66
- - **TLS fingerprinting**: ddgs uses `primp` with browser TLS impersonation. Node's
67
- `fetch` has a different fingerprint, so Google/Brave/Yahoo/Yandex may block or serve
68
- consent pages more aggressively. When an engine is blocked it simply contributes no
69
- results, exactly as when ddgs is blocked.
70
- - **User agents**: ddgs uses `fake_useragent`'s database; the port uses a fixed set of
71
- browser UAs plus ddgs's own Android Google UA generator.
72
-
73
- ### Entity decoding
74
-
75
- `decodeHtmlEntities` replicates CPython's `html.unescape` exactly (full HTML5 table,
76
- longest-prefix rule, Windows-1252 numeric mappings, invalid-codepoint handling), so the
77
- HTML-to-Markdown converter and search-result normalization match Studio's output.