pi-unsloth-webtools 0.2.1 → 0.2.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +36 -15
- package/engines.ts +36 -66
- package/html-to-md.ts +1 -1
- package/index.ts +6 -6
- package/package.json +5 -3
- package/pdf.ts +1 -1
- package/user-agents.ts +12 -0
- package/web-access.ts +11 -9
- package/web-fetch.ts +114 -40
- package/web-search.ts +11 -17
- package/ROADMAP.md +0 -77
package/README.md
CHANGED
|
@@ -26,7 +26,7 @@ Mirrors Unsloth Studio's `web_search` tool:
|
|
|
26
26
|
|
|
27
27
|
- Searches exactly like Studio's pinned `ddgs==9.14.4` `DDGS.text()`: the same seven engines
|
|
28
28
|
(duckduckgo, brave, google, mojeek, yahoo, yandex, wikipedia; bing is disabled upstream),
|
|
29
|
-
the same provider
|
|
29
|
+
the same provider deduplication, href-dedupe aggregator with frequency ordering, and the
|
|
30
30
|
same `SimpleFilterRanker` re-ranking. Formats results identically: `Title:` / `URL:` /
|
|
31
31
|
`Snippet:` blocks separated by `---`, ending with the hint to pass `{"url": "<URL>"}` to
|
|
32
32
|
read a full page.
|
|
@@ -45,7 +45,7 @@ Port of Studio's `_fetch_page_text` / `_fetch_url_raw` pipeline:
|
|
|
45
45
|
validation and fetch.
|
|
46
46
|
- GitHub repo root pages are rewritten to the unauthenticated README API
|
|
47
47
|
(`Accept: application/vnd.github.raw+json`), falling back to the HTML page on failure.
|
|
48
|
-
- Up to
|
|
48
|
+
- Up to 4 redirect hops, each re-validated and re-resolved against the same rules.
|
|
49
49
|
- 512 KiB download cap (10 MiB for PDFs), overall deadline + per-hop socket timeouts, abort-aware
|
|
50
50
|
(`signal` cancels mid-flight).
|
|
51
51
|
- PDF text extraction via the official MuPDF.js engine (the same C library pymupdf wraps):
|
|
@@ -53,8 +53,8 @@ Port of Studio's `_fetch_page_text` / `_fetch_url_raw` pipeline:
|
|
|
53
53
|
pymupdf4llm-style markdown layer (headings, bold/italic, code fences, links, tables)
|
|
54
54
|
with Studio's corrupted/incomplete fallback to plain text.
|
|
55
55
|
- Content sniffing: MIME allow/deny, binary magic signatures, PDF magic detection, and charset
|
|
56
|
-
decoding (declared charset, BOM sniffing for UTF-8/16/32,
|
|
57
|
-
single-byte pages).
|
|
56
|
+
decoding (declared charset, BOM sniffing for UTF-8/16/32, `<meta charset>` sniffing for
|
|
57
|
+
CJK and Windows/ISO encodings, cp1252 rescue for mislabeled single-byte pages).
|
|
58
58
|
- HTML → Markdown conversion ported from Studio's dependency-free `_html_to_md.py`: headings,
|
|
59
59
|
links, emphasis, lists, tables, blockquotes, code fences, entity decoding; hidden-element
|
|
60
60
|
stripping (`hidden`, `aria-hidden`, inline styles); `<article>`/`<main>` main-content scoping
|
|
@@ -65,8 +65,22 @@ Port of Studio's `_fetch_page_text` / `_fetch_url_raw` pipeline:
|
|
|
65
65
|
- HTML entity decoding replicates CPython's `html.unescape` (full 2,231-entry HTML5 table,
|
|
66
66
|
longest-prefix rule, Windows-1252 numeric mappings), matching Studio byte-for-byte.
|
|
67
67
|
|
|
68
|
-
|
|
69
|
-
|
|
68
|
+
## Known differences from Studio
|
|
69
|
+
|
|
70
|
+
- PDF styling: MuPDF.js exposes one font per line, so mixed-style lines style the
|
|
71
|
+
whole line instead of per-span; superscript, subscript, underline, strikeout, and
|
|
72
|
+
highlight markers are not emitted. Tables use a conservative text-grid detector:
|
|
73
|
+
aligned text tables are detected, drawn-rule-only tables are not.
|
|
74
|
+
- Search engines: Node's `fetch` TLS fingerprint differs from ddgs's `primp`
|
|
75
|
+
impersonation, so Google/Brave/Yahoo/Yandex may block or serve consent pages more
|
|
76
|
+
aggressively (a blocked engine simply contributes no results). User agents are a
|
|
77
|
+
fixed browser set plus ddgs's Android Google UA generator, not `fake_useragent`'s
|
|
78
|
+
database.
|
|
79
|
+
- Empty sweeps: ddgs 9.14.4 raises the last engine exception; this port reports a
|
|
80
|
+
timeout whenever any engine timed out, so the timeout message is not masked by later
|
|
81
|
+
generic engine failures.
|
|
82
|
+
- Proxies: Studio routes through environment proxies; this port always connects
|
|
83
|
+
directly with DNS pinning (deliberately out of scope).
|
|
70
84
|
|
|
71
85
|
## Development
|
|
72
86
|
|
|
@@ -74,27 +88,34 @@ fingerprinting, proxies).
|
|
|
74
88
|
npm install
|
|
75
89
|
npm run typecheck
|
|
76
90
|
npm test
|
|
91
|
+
npm run test:unit
|
|
92
|
+
npm run test:smoke
|
|
77
93
|
```
|
|
78
94
|
|
|
79
95
|
## Tests
|
|
80
96
|
|
|
97
|
+
`npm test` runs the full suite. `npm run test:unit` skips the live-network smoke tests,
|
|
98
|
+
and `npm run test:smoke` runs only those.
|
|
99
|
+
|
|
81
100
|
The suite ports Unsloth Studio's own tests for these tools:
|
|
82
101
|
|
|
83
|
-
- `test/html-to-md.test.ts
|
|
102
|
+
- `test/html-to-md.test.ts`: hidden-element stripping and main-content scoping (from
|
|
84
103
|
`test_web_fetch_extraction.py`)
|
|
85
|
-
- `test/header-strip.test.ts
|
|
104
|
+
- `test/header-strip.test.ts`: the header link-density suite plus article-vs-main selection
|
|
86
105
|
and boilerplate cases (from `test_web_fetch_extraction.py`)
|
|
87
|
-
- `test/binary-guard.test.ts
|
|
106
|
+
- `test/binary-guard.test.ts`: the MIME/magic/charset/PDF matrix (from
|
|
88
107
|
`test_web_fetch_binary_guard.py`)
|
|
89
|
-
- `test/web-search-policy.test.ts
|
|
108
|
+
- `test/web-search-policy.test.ts`: policy filtering, overfetch, and failure messages
|
|
90
109
|
(from `test_web_access_policy.py`)
|
|
91
|
-
- `test/fetch-flow.test.ts
|
|
110
|
+
- `test/fetch-flow.test.ts`: GitHub README rewrite, deadline/cancellation, HTML sniffing
|
|
92
111
|
(from `test_web_fetch_extraction.py`; the fetch client is injected via seams)
|
|
93
|
-
- `test/engines.test.ts
|
|
112
|
+
- `test/engines.test.ts`: the ddgs engine port, normalizers, the XPath subset, the
|
|
94
113
|
aggregator, the ranker, and the Wikipedia engine with a stubbed fetch
|
|
95
|
-
- `test/pdf-parity.test.ts
|
|
114
|
+
- `test/pdf-parity.test.ts`: MuPDF engine capabilities, PDF 1.5 object streams,
|
|
96
115
|
ASCII85Decode, font `/Differences` encodings, pymupdf4llm-style headings/links/tables
|
|
97
|
-
- `test/
|
|
116
|
+
- `test/entities.test.ts`: `decodeHtmlEntities` parity with CPython `html.unescape`,
|
|
117
|
+
legacy refs, longest-prefix rule, Windows-1252 numeric mappings, invalid codepoints
|
|
118
|
+
- `test/smoke.test.ts`: live network checks against real hosts
|
|
98
119
|
|
|
99
120
|
The seams (`seams.resolve` / `seams.request` / `rawFetch`) replace the network stack
|
|
100
121
|
with fakes, mirroring how the Studio suite monkeypatches `_validate_and_resolve_host`
|
|
@@ -104,4 +125,4 @@ and `build_opener`.
|
|
|
104
125
|
|
|
105
126
|
The ported logic derives from Unsloth Studio
|
|
106
127
|
([AGPL-3.0-only](https://github.com/unslothai/unsloth/blob/main/studio/LICENSE.AGPL-3.0)), so this
|
|
107
|
-
package is released under the same
|
|
128
|
+
package is released under the same AGPL-3.0-only license.
|
package/engines.ts
CHANGED
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
import { randomBytes } from "node:crypto";
|
|
2
2
|
import { decodeHtmlEntities, feedHtml } from "./html-to-md.ts";
|
|
3
3
|
import type { AttrDict } from "./html-to-md.ts";
|
|
4
|
+
import { randomUserAgent } from "./user-agents.ts";
|
|
4
5
|
export class EmptySweepError extends Error {
|
|
5
6
|
constructor() {
|
|
6
7
|
super("No results found");
|
|
@@ -110,6 +111,32 @@ interface XStep {
|
|
|
110
111
|
preds: Pred[];
|
|
111
112
|
terminal?: "text" | string;
|
|
112
113
|
}
|
|
114
|
+
function parsePredicateBlocks(input: string, start: number): { preds: Pred[]; next: number } {
|
|
115
|
+
const preds: Pred[] = [];
|
|
116
|
+
let pos = start;
|
|
117
|
+
while (pos < input.length && input[pos] === "[") {
|
|
118
|
+
const innerStart = pos + 1;
|
|
119
|
+
let depth = 1;
|
|
120
|
+
let quote: string | null = null;
|
|
121
|
+
let j = innerStart;
|
|
122
|
+
while (j < input.length && depth) {
|
|
123
|
+
const c = input[j];
|
|
124
|
+
if (quote !== null) {
|
|
125
|
+
if (c === quote) quote = null;
|
|
126
|
+
} else if (c === "'" || c === '"') {
|
|
127
|
+
quote = c;
|
|
128
|
+
} else if (c === "[") {
|
|
129
|
+
depth++;
|
|
130
|
+
} else if (c === "]") {
|
|
131
|
+
depth--;
|
|
132
|
+
}
|
|
133
|
+
j++;
|
|
134
|
+
}
|
|
135
|
+
preds.push(parsePredExpr(input.slice(innerStart, j - 1)));
|
|
136
|
+
pos = j;
|
|
137
|
+
}
|
|
138
|
+
return { preds, next: pos };
|
|
139
|
+
}
|
|
113
140
|
|
|
114
141
|
function parsePredExpr(input: string): Pred {
|
|
115
142
|
let pos = 0;
|
|
@@ -173,35 +200,10 @@ function parsePredExpr(input: string): Pred {
|
|
|
173
200
|
return { op: "desc", tag: name };
|
|
174
201
|
}
|
|
175
202
|
const name = word();
|
|
176
|
-
const preds =
|
|
203
|
+
const { preds, next } = parsePredicateBlocks(input, pos);
|
|
204
|
+
pos = next;
|
|
177
205
|
return { op: "child", tag: name, preds };
|
|
178
206
|
};
|
|
179
|
-
const parsePredBlocks = (): Pred[] => {
|
|
180
|
-
const preds: Pred[] = [];
|
|
181
|
-
while (pos < input.length && input[pos] === "[") {
|
|
182
|
-
const start = pos + 1;
|
|
183
|
-
let depth = 1;
|
|
184
|
-
let quote: string | null = null;
|
|
185
|
-
let i = start;
|
|
186
|
-
while (i < input.length && depth) {
|
|
187
|
-
const c = input[i];
|
|
188
|
-
if (quote !== null) {
|
|
189
|
-
if (c === quote) quote = null;
|
|
190
|
-
} else if (c === "'" || c === '"') {
|
|
191
|
-
quote = c;
|
|
192
|
-
} else if (c === "[") {
|
|
193
|
-
depth++;
|
|
194
|
-
} else if (c === "]") {
|
|
195
|
-
depth--;
|
|
196
|
-
}
|
|
197
|
-
i++;
|
|
198
|
-
}
|
|
199
|
-
const inner = input.slice(start, i - 1);
|
|
200
|
-
preds.push(parsePredExpr(inner));
|
|
201
|
-
pos = i;
|
|
202
|
-
}
|
|
203
|
-
return preds;
|
|
204
|
-
};
|
|
205
207
|
const parseAnd = (): Pred => {
|
|
206
208
|
let left = atom();
|
|
207
209
|
while (true) {
|
|
@@ -274,28 +276,8 @@ function parsePath(expr: string): XStep[] {
|
|
|
274
276
|
if (!m) break;
|
|
275
277
|
const name = m[0];
|
|
276
278
|
i += m[0].length;
|
|
277
|
-
const preds
|
|
278
|
-
|
|
279
|
-
const start = i + 1;
|
|
280
|
-
let depth = 1;
|
|
281
|
-
let quote: string | null = null;
|
|
282
|
-
let j = start;
|
|
283
|
-
while (j < expr.length && depth) {
|
|
284
|
-
const c = expr[j];
|
|
285
|
-
if (quote !== null) {
|
|
286
|
-
if (c === quote) quote = null;
|
|
287
|
-
} else if (c === "'" || c === '"') {
|
|
288
|
-
quote = c;
|
|
289
|
-
} else if (c === "[") {
|
|
290
|
-
depth++;
|
|
291
|
-
} else if (c === "]") {
|
|
292
|
-
depth--;
|
|
293
|
-
}
|
|
294
|
-
j++;
|
|
295
|
-
}
|
|
296
|
-
preds.push(parsePredExpr(expr.slice(start, j - 1)));
|
|
297
|
-
i = j;
|
|
298
|
-
}
|
|
279
|
+
const { preds, next } = parsePredicateBlocks(expr, i);
|
|
280
|
+
i = next;
|
|
299
281
|
steps.push({ axis, name, preds });
|
|
300
282
|
}
|
|
301
283
|
return steps;
|
|
@@ -409,19 +391,6 @@ export function extractResults(
|
|
|
409
391
|
return results;
|
|
410
392
|
}
|
|
411
393
|
|
|
412
|
-
const USER_AGENTS = [
|
|
413
|
-
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36",
|
|
414
|
-
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36",
|
|
415
|
-
"Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36",
|
|
416
|
-
"Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:133.0) Gecko/20100101 Firefox/133.0",
|
|
417
|
-
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:133.0) Gecko/20100101 Firefox/133.0",
|
|
418
|
-
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/18.2 Safari/605.1.15",
|
|
419
|
-
];
|
|
420
|
-
|
|
421
|
-
function randomUserAgent(): string {
|
|
422
|
-
return USER_AGENTS[Math.floor(Math.random() * USER_AGENTS.length)];
|
|
423
|
-
}
|
|
424
|
-
|
|
425
394
|
function googleUserAgent(): string {
|
|
426
395
|
const devices: [string, string, number, number][] = [
|
|
427
396
|
["5.0", "SM-G900P Build/LRX21T", 39, 60],
|
|
@@ -822,7 +791,7 @@ export async function autoTextSearch(
|
|
|
822
791
|
const seenProviders = new Set<string>();
|
|
823
792
|
const aggregator = new ResultsAggregator();
|
|
824
793
|
const ctx: EngineContext = { region: "us-en", safesearch: "moderate" };
|
|
825
|
-
let
|
|
794
|
+
let timedOut = false;
|
|
826
795
|
const uniqueProviders = new Set(engines.map((e) => e.provider)).size;
|
|
827
796
|
const maxWorkers = Math.min(uniqueProviders, Math.ceil(maxResults / 10) + 1);
|
|
828
797
|
let i = 0;
|
|
@@ -835,7 +804,7 @@ export async function autoTextSearch(
|
|
|
835
804
|
seenProviders.add(engine.provider);
|
|
836
805
|
}
|
|
837
806
|
} catch (e) {
|
|
838
|
-
|
|
807
|
+
if (e instanceof Error && e.message.includes("timed out")) timedOut = true;
|
|
839
808
|
}
|
|
840
809
|
};
|
|
841
810
|
while (i < engines.length) {
|
|
@@ -843,13 +812,14 @@ export async function autoTextSearch(
|
|
|
843
812
|
const engine = engines[i++];
|
|
844
813
|
if (seenProviders.has(engine.provider)) continue;
|
|
845
814
|
pending.push(run(engine));
|
|
846
|
-
if (pending.length >= maxWorkers
|
|
815
|
+
if (pending.length >= maxWorkers) {
|
|
847
816
|
await Promise.allSettled(pending);
|
|
848
817
|
pending = [];
|
|
849
818
|
}
|
|
850
819
|
}
|
|
820
|
+
await Promise.allSettled(pending);
|
|
851
821
|
const results = rankResults(aggregator.extractDicts(), query);
|
|
852
822
|
if (results.length) return results.slice(0, maxResults);
|
|
853
|
-
if (
|
|
823
|
+
if (timedOut) throw new SearchTimeoutError();
|
|
854
824
|
throw new EmptySweepError();
|
|
855
825
|
}
|
package/html-to-md.ts
CHANGED
|
@@ -197,7 +197,7 @@ export function decodeHtmlEntities(text: string): string {
|
|
|
197
197
|
return text.replace(CHARREF_RE, (whole, s: string) => {
|
|
198
198
|
if (s[0] === "#") {
|
|
199
199
|
const hex = s[1] === "x" || s[1] === "X";
|
|
200
|
-
const num = parseInt(s.slice(2).replace(/;+$/, ""), hex ? 16 : 10);
|
|
200
|
+
const num = parseInt(s.slice(hex ? 2 : 1).replace(/;+$/, ""), hex ? 16 : 10);
|
|
201
201
|
const mapped = INVALID_CHARREFS[num];
|
|
202
202
|
if (mapped !== undefined) return mapped;
|
|
203
203
|
if ((num >= 0xd800 && num <= 0xdfff) || num > 0x10ffff) return "\ufffd";
|
package/index.ts
CHANGED
|
@@ -62,12 +62,12 @@ export default function (pi: ExtensionAPI) {
|
|
|
62
62
|
name: "web_fetch",
|
|
63
63
|
label: "Web Fetch",
|
|
64
64
|
description:
|
|
65
|
-
"Fetch a URL and return readable text
|
|
66
|
-
"
|
|
67
|
-
"stripping
|
|
68
|
-
"
|
|
69
|
-
"
|
|
70
|
-
"
|
|
65
|
+
"Fetch a URL and return its readable text. HTML pages are converted to Markdown using a " +
|
|
66
|
+
"main-content heuristic: article/main scoping plus hidden-element and boilerplate " +
|
|
67
|
+
"stripping. Non-HTML text is returned as-is. GitHub repo root pages are rewritten to the " +
|
|
68
|
+
"README API, so the README is returned instead of the repo page's UI chrome. " +
|
|
69
|
+
"Private/loopback/link-local targets are blocked (SSRF protection), and the download size " +
|
|
70
|
+
"is capped.",
|
|
71
71
|
promptSnippet: "Fetch a web page and return readable text content",
|
|
72
72
|
parameters: WebFetchParams,
|
|
73
73
|
async execute(_toolCallId, params, signal, _onUpdate, _ctx) {
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-unsloth-webtools",
|
|
3
|
-
"version": "0.2.
|
|
3
|
+
"version": "0.2.3",
|
|
4
4
|
"type": "module",
|
|
5
5
|
"description": "Pi extension: web_search and web_fetch tools ported from the Unsloth Studio codebase (DuckDuckGo search, SSRF-safe direct fetching, HTML-to-Markdown extraction)",
|
|
6
6
|
"main": "index.ts",
|
|
@@ -30,8 +30,8 @@
|
|
|
30
30
|
"engines.ts",
|
|
31
31
|
"entities.ts",
|
|
32
32
|
"pdf.ts",
|
|
33
|
+
"user-agents.ts",
|
|
33
34
|
"README.md",
|
|
34
|
-
"ROADMAP.md",
|
|
35
35
|
"LICENSE"
|
|
36
36
|
],
|
|
37
37
|
"pi": {
|
|
@@ -48,8 +48,10 @@
|
|
|
48
48
|
},
|
|
49
49
|
"scripts": {
|
|
50
50
|
"test": "vitest run",
|
|
51
|
+
"test:unit": "vitest run --exclude **/smoke.test.ts",
|
|
52
|
+
"test:smoke": "vitest run test/smoke.test.ts",
|
|
51
53
|
"typecheck": "tsc --noEmit",
|
|
52
|
-
"prepublishOnly": "npm run typecheck"
|
|
54
|
+
"prepublishOnly": "npm run typecheck && npm run test:unit"
|
|
53
55
|
},
|
|
54
56
|
"devDependencies": {
|
|
55
57
|
"@earendil-works/pi-coding-agent": "^0.84.0",
|
package/pdf.ts
CHANGED
|
@@ -527,7 +527,7 @@ function assemblePages(
|
|
|
527
527
|
return "";
|
|
528
528
|
}
|
|
529
529
|
if (pageLimitReached) {
|
|
530
|
-
text += `\n\n... (PDF extraction
|
|
530
|
+
text += `\n\n... (PDF extraction is capped at ${MAX_WEB_PDF_PAGES} pages)`;
|
|
531
531
|
}
|
|
532
532
|
return text;
|
|
533
533
|
}
|
package/user-agents.ts
ADDED
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
const USER_AGENTS = [
|
|
2
|
+
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36",
|
|
3
|
+
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36",
|
|
4
|
+
"Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36",
|
|
5
|
+
"Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:133.0) Gecko/20100101 Firefox/133.0",
|
|
6
|
+
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:133.0) Gecko/20100101 Firefox/133.0",
|
|
7
|
+
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/18.2 Safari/605.1.15",
|
|
8
|
+
];
|
|
9
|
+
|
|
10
|
+
export function randomUserAgent(): string {
|
|
11
|
+
return USER_AGENTS[Math.floor(Math.random() * USER_AGENTS.length)];
|
|
12
|
+
}
|
package/web-access.ts
CHANGED
|
@@ -173,46 +173,46 @@ export function checkUrlAccess(
|
|
|
173
173
|
policy: WebsitePolicy | null,
|
|
174
174
|
): [boolean, string, string] {
|
|
175
175
|
if (typeof url !== "string" || !url.trim()) {
|
|
176
|
-
return [false, "Blocked: URL is empty.", ""];
|
|
176
|
+
return [false, "Blocked: the URL is empty.", ""];
|
|
177
177
|
}
|
|
178
178
|
const candidate = url.trim();
|
|
179
179
|
if (
|
|
180
180
|
Array.from(candidate).some((char) => /\s/.test(char) || char.charCodeAt(0) < 32) ||
|
|
181
181
|
candidate.includes("\\")
|
|
182
182
|
) {
|
|
183
|
-
return [false, "Blocked: URL contains invalid characters.", ""];
|
|
183
|
+
return [false, "Blocked: the URL contains invalid characters.", ""];
|
|
184
184
|
}
|
|
185
185
|
let parsed: URL;
|
|
186
186
|
try {
|
|
187
187
|
parsed = new URL(candidate);
|
|
188
188
|
} catch {
|
|
189
|
-
return [false, "Blocked: URL has an invalid hostname or port.", ""];
|
|
189
|
+
return [false, "Blocked: the URL has an invalid hostname or port.", ""];
|
|
190
190
|
}
|
|
191
191
|
const scheme = parsed.protocol.replace(/:$/, "").toLowerCase();
|
|
192
192
|
if (scheme !== "http" && scheme !== "https") {
|
|
193
193
|
return [false, "Blocked: only http/https URLs are allowed.", ""];
|
|
194
194
|
}
|
|
195
195
|
if (parsed.username || parsed.password || parsed.hostname.includes("%")) {
|
|
196
|
-
return [false, "Blocked:
|
|
196
|
+
return [false, "Blocked: URLs with credentials or encoded hostnames are not allowed.", ""];
|
|
197
197
|
}
|
|
198
198
|
if (!parsed.hostname) {
|
|
199
|
-
return [false, "Blocked: URL has an invalid hostname or port.", ""];
|
|
199
|
+
return [false, "Blocked: the URL has an invalid hostname or port.", ""];
|
|
200
200
|
}
|
|
201
201
|
try {
|
|
202
202
|
if (parsed.port && !(PORT_RE.test(parsed.port) && Number(parsed.port) >= 1 && Number(parsed.port) <= 65535)) {
|
|
203
|
-
return [false, "Blocked: URL has an invalid hostname or port.", ""];
|
|
203
|
+
return [false, "Blocked: the URL has an invalid hostname or port.", ""];
|
|
204
204
|
}
|
|
205
205
|
} catch {
|
|
206
|
-
return [false, "Blocked: URL has an invalid hostname or port.", ""];
|
|
206
|
+
return [false, "Blocked: the URL has an invalid hostname or port.", ""];
|
|
207
207
|
}
|
|
208
208
|
let hostname: string;
|
|
209
209
|
try {
|
|
210
210
|
hostname = normalizeDomain(parsed.hostname);
|
|
211
211
|
} catch {
|
|
212
|
-
return [false, "Blocked: URL has an invalid hostname or port.", ""];
|
|
212
|
+
return [false, "Blocked: the URL has an invalid hostname or port.", ""];
|
|
213
213
|
}
|
|
214
214
|
if (!hostnameAllowed(hostname, policy)) {
|
|
215
|
-
return [false, `Blocked: website access policy disallows ${hostname}.`, hostname];
|
|
215
|
+
return [false, `Blocked: the website access policy disallows ${hostname}.`, hostname];
|
|
216
216
|
}
|
|
217
217
|
return [true, "", hostname];
|
|
218
218
|
}
|
|
@@ -368,6 +368,8 @@ export function isPublicIp(ip: string): boolean {
|
|
|
368
368
|
if (lower.startsWith("2001:db8")) return false;
|
|
369
369
|
if (lower.startsWith("64:ff9b:")) return false;
|
|
370
370
|
if (lower.startsWith("2001:10:")) return false;
|
|
371
|
+
if (lower.startsWith("2002:")) return false;
|
|
372
|
+
if (lower.startsWith("2001:0:") || lower.startsWith("2001::")) return false;
|
|
371
373
|
const mapped = /^::ffff:(\d+\.\d+\.\d+\.\d+)$/.exec(lower);
|
|
372
374
|
if (mapped) return isPublicIp(mapped[1]);
|
|
373
375
|
if (lower.startsWith("::ffff:")) return false;
|
package/web-fetch.ts
CHANGED
|
@@ -11,20 +11,11 @@ import {
|
|
|
11
11
|
} from "./web-access.ts";
|
|
12
12
|
import { htmlToMarkdown } from "./html-to-md.ts";
|
|
13
13
|
import { extractPdfText, PdfParseError } from "./pdf.ts";
|
|
14
|
+
import { randomUserAgent } from "./user-agents.ts";
|
|
14
15
|
|
|
15
|
-
const MIN_PAGE_CHARS = 2000;
|
|
16
16
|
const MAX_FETCH_BYTES = 512 * 1024;
|
|
17
17
|
const MAX_PDF_FETCH_BYTES = 10 * 1024 * 1024;
|
|
18
|
-
const
|
|
19
|
-
|
|
20
|
-
const USER_AGENTS = [
|
|
21
|
-
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36",
|
|
22
|
-
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36",
|
|
23
|
-
"Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36",
|
|
24
|
-
"Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:133.0) Gecko/20100101 Firefox/133.0",
|
|
25
|
-
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:133.0) Gecko/20100101 Firefox/133.0",
|
|
26
|
-
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/18.2 Safari/605.1.15",
|
|
27
|
-
];
|
|
18
|
+
const MAX_REQUESTS = 5;
|
|
28
19
|
|
|
29
20
|
const UTF32_LE_BOM = Buffer.from([0xff, 0xfe, 0x00, 0x00]);
|
|
30
21
|
const UTF32_BE_BOM = Buffer.from([0x00, 0x00, 0xfe, 0xff]);
|
|
@@ -237,6 +228,34 @@ function hasSingleByteTextEvidence(data: Buffer): boolean {
|
|
|
237
228
|
return ascii / data.length >= MIN_SINGLE_BYTE_ASCII_RATIO;
|
|
238
229
|
}
|
|
239
230
|
|
|
231
|
+
const CHARSET_ALIASES: Record<string, string> = {
|
|
232
|
+
gbk: "gbk",
|
|
233
|
+
gb2312: "gbk",
|
|
234
|
+
"gb-2312": "gbk",
|
|
235
|
+
cp936: "gbk",
|
|
236
|
+
"x-gbk": "gbk",
|
|
237
|
+
gb18030: "gb18030",
|
|
238
|
+
big5: "big5",
|
|
239
|
+
"big5-hkscs": "big5",
|
|
240
|
+
sjis: "shift_jis",
|
|
241
|
+
"x-sjis": "shift_jis",
|
|
242
|
+
cp932: "shift_jis",
|
|
243
|
+
"shift-jis": "shift_jis",
|
|
244
|
+
"euc-jp": "euc-jp",
|
|
245
|
+
"euc-kr": "euc-kr",
|
|
246
|
+
ksc5601: "euc-kr",
|
|
247
|
+
"ks_c_5601-1987": "euc-kr",
|
|
248
|
+
"ks_c_5601-1989": "euc-kr",
|
|
249
|
+
"iso-2022-jp": "iso-2022-jp",
|
|
250
|
+
"koi8-r": "koi8-r",
|
|
251
|
+
"koi8-u": "koi8-u",
|
|
252
|
+
cp866: "cp866",
|
|
253
|
+
"x-mac-cyrillic": "x-mac-cyrillic",
|
|
254
|
+
"windows-874": "windows-874",
|
|
255
|
+
cp874: "windows-874",
|
|
256
|
+
"tis-620": "tis-620",
|
|
257
|
+
};
|
|
258
|
+
|
|
240
259
|
function normalizeCharset(name: string): string | null {
|
|
241
260
|
const n = name.trim().replace(/["']/g, "").toLowerCase();
|
|
242
261
|
switch (n) {
|
|
@@ -270,8 +289,28 @@ function normalizeCharset(name: string): string | null {
|
|
|
270
289
|
case "utf-32be":
|
|
271
290
|
return "utf-32be";
|
|
272
291
|
default:
|
|
273
|
-
|
|
292
|
+
break;
|
|
293
|
+
}
|
|
294
|
+
const alias = CHARSET_ALIASES[n];
|
|
295
|
+
if (alias !== undefined) return alias;
|
|
296
|
+
if (/^windows-125[0-8]$/.test(n) || /^iso-8859-(?:[2-9]|1[0-6])$/.test(n)) return n;
|
|
297
|
+
return null;
|
|
298
|
+
}
|
|
299
|
+
|
|
300
|
+
function sniffMetaCharset(bytes: Buffer): string | null {
|
|
301
|
+
const head = bytes.subarray(0, 1024).toString("latin1").toLowerCase();
|
|
302
|
+
const match =
|
|
303
|
+
/<meta\b[^>]*\bcharset\s*=\s*["']?\s*([a-z0-9_.\-]+)/.exec(head) ??
|
|
304
|
+
/<meta\b[^>]*\bhttp-equiv\s*=\s*["']?content-type["']?[^>]*\bcharset\s*=\s*["']?\s*([a-z0-9_.\-]+)/.exec(head);
|
|
305
|
+
return match ? normalizeCharset(match[1]) : null;
|
|
306
|
+
}
|
|
307
|
+
|
|
308
|
+
function sniffMetaCharsetForHtml(bytes: Buffer, contentType: string): string | null {
|
|
309
|
+
if (!contentType.includes("html")) {
|
|
310
|
+
const probe = bytes.subarray(0, 256).toString("latin1");
|
|
311
|
+
if (!looksLikeHtmlDocument(probe)) return null;
|
|
274
312
|
}
|
|
313
|
+
return sniffMetaCharset(bytes);
|
|
275
314
|
}
|
|
276
315
|
|
|
277
316
|
function decodeUtf32(bytes: Buffer, littleEndian: boolean): string {
|
|
@@ -328,11 +367,38 @@ function decodeWithCodec(bytes: Buffer, codec: string | null): string {
|
|
|
328
367
|
case "cp1252":
|
|
329
368
|
return decodeSingleByte(bytes, true);
|
|
330
369
|
case "iso8859-1":
|
|
331
|
-
default:
|
|
332
370
|
return decodeSingleByte(bytes, false);
|
|
371
|
+
default:
|
|
372
|
+
return decodeWithLabel(bytes, codec);
|
|
333
373
|
}
|
|
334
374
|
}
|
|
335
375
|
|
|
376
|
+
function decodeWithLabel(bytes: Buffer, label: string | null): string {
|
|
377
|
+
if (label === null) return decodeSingleByte(bytes, false);
|
|
378
|
+
if (label === "tis-620") return decodeTis620(bytes);
|
|
379
|
+
try {
|
|
380
|
+
return new TextDecoder(label, { fatal: false }).decode(bytes);
|
|
381
|
+
} catch {
|
|
382
|
+
return decodeSingleByte(bytes, false);
|
|
383
|
+
}
|
|
384
|
+
}
|
|
385
|
+
|
|
386
|
+
function decodeTis620(bytes: Buffer): string {
|
|
387
|
+
let out = "";
|
|
388
|
+
for (const byte of bytes) {
|
|
389
|
+
if (byte < 0x80) {
|
|
390
|
+
out += String.fromCharCode(byte);
|
|
391
|
+
} else if (byte >= 0xa1 && byte <= 0xfb) {
|
|
392
|
+
out += String.fromCodePoint(0x0e01 + byte - 0xa1);
|
|
393
|
+
} else if (byte === 0xa0) {
|
|
394
|
+
out += "\u00a0";
|
|
395
|
+
} else {
|
|
396
|
+
out += "\ufffd";
|
|
397
|
+
}
|
|
398
|
+
}
|
|
399
|
+
return out;
|
|
400
|
+
}
|
|
401
|
+
|
|
336
402
|
|
|
337
403
|
async function resolveAndValidate(hostname: string, signal?: AbortSignal): Promise<ResolvedHost> {
|
|
338
404
|
let addresses: { address: string; family: number }[];
|
|
@@ -346,7 +412,7 @@ async function resolveAndValidate(hostname: string, signal?: AbortSignal): Promi
|
|
|
346
412
|
}
|
|
347
413
|
for (const entry of addresses) {
|
|
348
414
|
if (!isPublicIp(entry.address)) {
|
|
349
|
-
return { ok: false, reason: `Blocked: refusing to fetch non-public address ${entry.address}.`, ip: "", family: 0 };
|
|
415
|
+
return { ok: false, reason: `Blocked: refusing to fetch the non-public address ${entry.address}.`, ip: "", family: 0 };
|
|
350
416
|
}
|
|
351
417
|
}
|
|
352
418
|
const first = addresses[0];
|
|
@@ -386,20 +452,19 @@ function requestHop(opts: HopOptions): Promise<HopResponse> {
|
|
|
386
452
|
const declaredPdf = String(res.headers["content-type"] ?? "").toLowerCase().includes("pdf");
|
|
387
453
|
let limit = declaredPdf ? opts.maxPdfBytes : opts.maxBytes;
|
|
388
454
|
let extendedForPdf = false;
|
|
389
|
-
let done = false;
|
|
390
455
|
const finish = (err: string | null, body: Buffer) => {
|
|
391
|
-
|
|
392
|
-
|
|
393
|
-
|
|
394
|
-
|
|
395
|
-
|
|
396
|
-
|
|
397
|
-
|
|
398
|
-
|
|
399
|
-
|
|
456
|
+
settle(() => {
|
|
457
|
+
if (err) reject(new Error(err));
|
|
458
|
+
else
|
|
459
|
+
resolve({
|
|
460
|
+
status: res.statusCode ?? 0,
|
|
461
|
+
headers: res.headers as Record<string, string | string[] | undefined>,
|
|
462
|
+
body,
|
|
463
|
+
});
|
|
464
|
+
});
|
|
400
465
|
};
|
|
401
466
|
res.on("data", (chunk: Buffer) => {
|
|
402
|
-
if (
|
|
467
|
+
if (settled) return;
|
|
403
468
|
if (!declaredPdf && !extendedForPdf && total + chunk.length > opts.maxBytes) {
|
|
404
469
|
if (hasPdfMagic(Buffer.concat(chunks))) {
|
|
405
470
|
limit = opts.maxPdfBytes;
|
|
@@ -423,10 +488,17 @@ function requestHop(opts: HopOptions): Promise<HopResponse> {
|
|
|
423
488
|
res.on("end", () => finish(null, Buffer.concat(chunks)));
|
|
424
489
|
res.on("error", (err) => finish(err.message, Buffer.concat(chunks)));
|
|
425
490
|
});
|
|
426
|
-
request.on("timeout", () => request.destroy(new Error("timed out")));
|
|
427
|
-
request.on("error", (err) => reject(err));
|
|
428
491
|
const onAbort = () => request.destroy(new Error("cancelled"));
|
|
429
492
|
opts.signal?.addEventListener("abort", onAbort, { once: true });
|
|
493
|
+
let settled = false;
|
|
494
|
+
const settle = (action: () => void) => {
|
|
495
|
+
if (settled) return;
|
|
496
|
+
settled = true;
|
|
497
|
+
opts.signal?.removeEventListener("abort", onAbort);
|
|
498
|
+
action();
|
|
499
|
+
};
|
|
500
|
+
request.on("timeout", () => request.destroy(new Error("timed out")));
|
|
501
|
+
request.on("error", (err) => settle(() => reject(err)));
|
|
430
502
|
request.end();
|
|
431
503
|
});
|
|
432
504
|
}
|
|
@@ -460,9 +532,9 @@ export async function fetchUrlRaw(
|
|
|
460
532
|
let currentUrl = url;
|
|
461
533
|
let pinnedIp = resolved.ip;
|
|
462
534
|
let pinnedFamily = resolved.family;
|
|
463
|
-
const userAgent =
|
|
535
|
+
const userAgent = randomUserAgent();
|
|
464
536
|
|
|
465
|
-
for (let hop = 0; hop <
|
|
537
|
+
for (let hop = 0; hop < MAX_REQUESTS; hop++) {
|
|
466
538
|
budgetError = fetchBudgetExceeded(deadline, signal, now);
|
|
467
539
|
if (budgetError !== null) return { error: budgetError, body: "", contentType: "" };
|
|
468
540
|
const parsed = new URL(currentUrl);
|
|
@@ -471,6 +543,7 @@ export async function fetchUrlRaw(
|
|
|
471
543
|
"User-Agent": userAgent,
|
|
472
544
|
Host: hostHeader,
|
|
473
545
|
Accept: "text/html,application/xhtml+xml,text/plain;q=0.9,*/*;q=0.5",
|
|
546
|
+
"Accept-Encoding": "identity",
|
|
474
547
|
};
|
|
475
548
|
if (options.extraHeaders) Object.assign(headers, options.extraHeaders);
|
|
476
549
|
const inactivity = Math.max(1, deadline - now());
|
|
@@ -497,8 +570,9 @@ export async function fetchUrlRaw(
|
|
|
497
570
|
|
|
498
571
|
if (response.status >= 300 && response.status < 400) {
|
|
499
572
|
if (![301, 302, 303, 307, 308].includes(response.status)) {
|
|
573
|
+
const reason = http.STATUS_CODES[response.status] ?? "";
|
|
500
574
|
return {
|
|
501
|
-
error: `Failed to fetch URL: HTTP ${response.status} ${
|
|
575
|
+
error: `Failed to fetch URL: HTTP ${response.status}${reason ? ` ${reason}` : ""}`,
|
|
502
576
|
body: "",
|
|
503
577
|
contentType: "",
|
|
504
578
|
};
|
|
@@ -507,12 +581,16 @@ export async function fetchUrlRaw(
|
|
|
507
581
|
const location = Array.isArray(rawLocation) ? rawLocation[0] : rawLocation;
|
|
508
582
|
if (!location) {
|
|
509
583
|
return {
|
|
510
|
-
error: "Failed to fetch URL: redirect missing Location header.",
|
|
584
|
+
error: "Failed to fetch URL: the redirect is missing a Location header.",
|
|
511
585
|
body: "",
|
|
512
586
|
contentType: "",
|
|
513
587
|
};
|
|
514
588
|
}
|
|
515
|
-
|
|
589
|
+
try {
|
|
590
|
+
currentUrl = new URL(location, currentUrl).toString();
|
|
591
|
+
} catch {
|
|
592
|
+
return { error: "Failed to fetch URL: the redirect has an invalid Location.", body: "", contentType: "" };
|
|
593
|
+
}
|
|
516
594
|
const [redirectAllowed, redirectReason, redirectHost] = checkUrlAccess(
|
|
517
595
|
currentUrl,
|
|
518
596
|
policy,
|
|
@@ -577,7 +655,10 @@ export async function fetchUrlRaw(
|
|
|
577
655
|
|
|
578
656
|
const declaredCodec = declaredCharset ? normalizeCharset(declaredCharset) : null;
|
|
579
657
|
const bomCodec = bomCodecFor(response.body);
|
|
580
|
-
const rawHtml = decodeWithCodec(
|
|
658
|
+
const rawHtml = decodeWithCodec(
|
|
659
|
+
response.body,
|
|
660
|
+
declaredCodec ?? bomCodec ?? sniffMetaCharsetForHtml(response.body, contentType) ?? "utf-8",
|
|
661
|
+
);
|
|
581
662
|
|
|
582
663
|
if (looksBinary(rawHtml)) {
|
|
583
664
|
let alt: string | null = null;
|
|
@@ -612,13 +693,6 @@ function bomCodecFor(bytes: Buffer): string | null {
|
|
|
612
693
|
return null;
|
|
613
694
|
}
|
|
614
695
|
|
|
615
|
-
function statusReason(status: number): string {
|
|
616
|
-
const reasons: Record<number, string> = {
|
|
617
|
-
301: "Moved Permanently", 302: "Found", 303: "See Other", 307: "Temporary Redirect", 308: "Permanent Redirect",
|
|
618
|
-
};
|
|
619
|
-
return reasons[status] ?? "";
|
|
620
|
-
}
|
|
621
|
-
|
|
622
696
|
export function truncatePageText(text: string, maxChars?: number): string {
|
|
623
697
|
if (!text) return "(page returned no readable text)";
|
|
624
698
|
if (typeof maxChars === "number" && maxChars > 0 && text.length > maxChars) {
|
package/web-search.ts
CHANGED
|
@@ -64,23 +64,17 @@ export async function webSearch(
|
|
|
64
64
|
const results = await client(effectiveQuery, wanted, signal);
|
|
65
65
|
if (signal?.aborted) return "Search cancelled.";
|
|
66
66
|
if (!results.length) return EMPTY_SEARCH_RESULTS[0];
|
|
67
|
-
const
|
|
67
|
+
const allowed: SearchResult[] = [];
|
|
68
68
|
for (const result of results) {
|
|
69
|
-
if (
|
|
69
|
+
if (allowed.length >= maxResults) break;
|
|
70
70
|
const href = String(result.href ?? "").trim();
|
|
71
71
|
if (href && !checkUrlAccess(href, policy)[0]) continue;
|
|
72
|
-
|
|
73
|
-
const snippet = String(result.body ?? "").replace(/\s+/g, " ");
|
|
74
|
-
parts.push(`Title: ${title}\nURL: ${href}\nSnippet: ${snippet}`);
|
|
72
|
+
allowed.push(result);
|
|
75
73
|
}
|
|
76
|
-
if (!
|
|
77
|
-
|
|
78
|
-
return (
|
|
79
|
-
text +
|
|
80
|
-
"\n\n---\n\nIMPORTANT: These are only short snippets. " +
|
|
81
|
-
'To get the full page content, call web_search with the url parameter (e.g. {"url": "<URL>"}).'
|
|
82
|
-
);
|
|
74
|
+
if (!allowed.length) return EMPTY_SEARCH_RESULTS[1];
|
|
75
|
+
return formatSearchResults(allowed);
|
|
83
76
|
} catch (err) {
|
|
77
|
+
if (signal?.aborted) return "Search cancelled.";
|
|
84
78
|
return searchFailureMessage(err, timeoutMs);
|
|
85
79
|
}
|
|
86
80
|
}
|
|
@@ -88,7 +82,7 @@ export async function webSearch(
|
|
|
88
82
|
export function searchFailureMessage(exc: unknown, timeoutMs = SEARCH_TIMEOUT_MS): string {
|
|
89
83
|
if (exc instanceof SearchCancelled) return "Search cancelled.";
|
|
90
84
|
if (exc instanceof SearchTimeoutError) {
|
|
91
|
-
return `Search failed: the search engines did not respond within ${Math.round(timeoutMs / 1000)}
|
|
85
|
+
return `Search failed: the search engines did not respond within ${Math.round(timeoutMs / 1000)} seconds.`;
|
|
92
86
|
}
|
|
93
87
|
if (exc instanceof EmptySweepError || (exc instanceof Error && exc.message.includes("No results found"))) {
|
|
94
88
|
return EMPTY_SEARCH_RESULTS[0];
|
|
@@ -98,15 +92,15 @@ export function searchFailureMessage(exc: unknown, timeoutMs = SEARCH_TIMEOUT_MS
|
|
|
98
92
|
|
|
99
93
|
export function formatSearchResults(results: SearchResult[]): string {
|
|
100
94
|
const parts = results.map((result) => {
|
|
101
|
-
const title = result.title.replace(/\s+/g, " ");
|
|
102
|
-
const href = result.href.trim();
|
|
103
|
-
const snippet = result.body.replace(/\s+/g, " ");
|
|
95
|
+
const title = String(result.title ?? "").replace(/\s+/g, " ");
|
|
96
|
+
const href = String(result.href ?? "").trim();
|
|
97
|
+
const snippet = String(result.body ?? "").replace(/\s+/g, " ");
|
|
104
98
|
return `Title: ${title}\nURL: ${href}\nSnippet: ${snippet}`;
|
|
105
99
|
});
|
|
106
100
|
const text = parts.join("\n\n---\n\n");
|
|
107
101
|
return (
|
|
108
102
|
text +
|
|
109
|
-
"\n\n---\n\
|
|
103
|
+
"\n\n---\n\nThese are only short snippets. " +
|
|
110
104
|
'To get the full page content, call web_search with the url parameter (e.g. {"url": "<URL>"}).'
|
|
111
105
|
);
|
|
112
106
|
}
|
package/ROADMAP.md
DELETED
|
@@ -1,77 +0,0 @@
|
|
|
1
|
-
# Roadmap
|
|
2
|
-
|
|
3
|
-
Tracking the remaining gaps between this port and Unsloth Studio's web tools, and the
|
|
4
|
-
planned work to close them.
|
|
5
|
-
|
|
6
|
-
## Implemented
|
|
7
|
-
|
|
8
|
-
### PDF text extraction (MuPDF engine)
|
|
9
|
-
|
|
10
|
-
`pdf.ts` uses the official **MuPDF.js** (`mupdf` npm package) — the same C engine that
|
|
11
|
-
PyMuPDF wraps — replacing the earlier minimal built-in extractor:
|
|
12
|
-
|
|
13
|
-
- Full xref handling: tables, cross-reference streams, and PDF 1.5+ object streams
|
|
14
|
-
- All standard stream filters (FlateDecode, ASCII85Decode, LZW, RunLength, DCT, JPX, ...)
|
|
15
|
-
- Font encodings and ToUnicode mapping (non-Latin text extracts correctly)
|
|
16
|
-
- Encryption detection via `needsPassword()` (reported as unreadable, matching pymupdf's
|
|
17
|
-
no-password behavior)
|
|
18
|
-
- The markdown layer replicates pymupdf4llm's algorithm: `IdentifyHeaders` font-size
|
|
19
|
-
heading detection, `get_raw_lines` line reconstruction (tolerance 3, 10% span-join
|
|
20
|
-
delta), `write_text` styling (bold/italic/mono, code fences, bullets, link
|
|
21
|
-
resolution with `%0x`-escaped URIs), and Studio's corrupted/incomplete fallback to
|
|
22
|
-
plain MuPDF text with the exact thresholds from `backend/core/rag/parsers.py`
|
|
23
|
-
- Table detection and pipe-markdown rendering matching pymupdf's `Table.to_markdown`
|
|
24
|
-
shape (`|header|`, `|---|`, detail rows, `Col{i}` fill for empty headers)
|
|
25
|
-
|
|
26
|
-
Known deltas vs pymupdf4llm:
|
|
27
|
-
|
|
28
|
-
- Span-level styling: MuPDF.js's structured-text JSON exposes one font per line, so
|
|
29
|
-
mixed-style lines (one bold word inside a body line) style the whole line instead of
|
|
30
|
-
per-span. Line-level styling matches for homogeneous lines.
|
|
31
|
-
- Superscript/subscript/underline/strikeout/highlight markers are not emitted (the
|
|
32
|
-
JSON does not expose char-level flags).
|
|
33
|
-
- Table detection is a conservative text-grid detector (column-start clustering with
|
|
34
|
-
a 5 pt tolerance, contiguous multi-row bands) instead of PyMuPDF's
|
|
35
|
-
`find_tables()` vector-graphics analysis. Aligned text tables are detected; tables
|
|
36
|
-
defined only by drawn rules without aligned text are not.
|
|
37
|
-
|
|
38
|
-
The minimal extractor remains as an automatic fallback when the `mupdf` package cannot
|
|
39
|
-
be loaded (for example a stripped install).
|
|
40
|
-
|
|
41
|
-
## Not planned (explicit decisions)
|
|
42
|
-
|
|
43
|
-
### Proxy support
|
|
44
|
-
|
|
45
|
-
Studio routes requests through urllib's environment proxies and honors
|
|
46
|
-
`UNSLOTH_STUDIO_DISABLE_DNS_PINNING` for enterprise proxies. This port always connects
|
|
47
|
-
directly with DNS pinning. Deliberately out of scope.
|
|
48
|
-
|
|
49
|
-
### Page size budgets
|
|
50
|
-
|
|
51
|
-
Studio's `_page_char_budget()` sizes fetched pages to the serving model's context window.
|
|
52
|
-
This port deliberately has **no character budget at all**: fetched pages and PDFs are
|
|
53
|
-
returned in full (the PDF page-count cap of 50 pages remains, matching Studio). The raw
|
|
54
|
-
download caps (512 KiB text / 10 MiB PDF) still bound what is fetched. A `maxChars`
|
|
55
|
-
parameter remains available on the tool for callers that want to truncate.
|
|
56
|
-
|
|
57
|
-
## Known behavioral differences
|
|
58
|
-
|
|
59
|
-
### Search engine set
|
|
60
|
-
|
|
61
|
-
The port implements ddgs 9.14.4's `DDGS.text()` exactly: the same seven engines
|
|
62
|
-
(duckduckgo, brave, google, mojeek, yahoo, yandex, wikipedia — bing is `disabled` in
|
|
63
|
-
ddgs upstream), the same provider-deduplication, href-dedupe aggregator with
|
|
64
|
-
frequency ordering, and the same `SimpleFilterRanker` re-ranking. Remaining deltas:
|
|
65
|
-
|
|
66
|
-
- **TLS fingerprinting**: ddgs uses `primp` with browser TLS impersonation. Node's
|
|
67
|
-
`fetch` has a different fingerprint, so Google/Brave/Yahoo/Yandex may block or serve
|
|
68
|
-
consent pages more aggressively. When an engine is blocked it simply contributes no
|
|
69
|
-
results, exactly as when ddgs is blocked.
|
|
70
|
-
- **User agents**: ddgs uses `fake_useragent`'s database; the port uses a fixed set of
|
|
71
|
-
browser UAs plus ddgs's own Android Google UA generator.
|
|
72
|
-
|
|
73
|
-
### Entity decoding
|
|
74
|
-
|
|
75
|
-
`decodeHtmlEntities` replicates CPython's `html.unescape` exactly (full HTML5 table,
|
|
76
|
-
longest-prefix rule, Windows-1252 numeric mappings, invalid-codepoint handling), so the
|
|
77
|
-
HTML-to-Markdown converter and search-result normalization match Studio's output.
|