@citation-media/html-parser 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Citation Media GmbH
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md ADDED
@@ -0,0 +1,103 @@
1
+ # @citation-media/html-parser
2
+
3
+ Extracts links, images, resources, metadata, and Markdown from HTML in Cloudflare Workers, in one streaming pass over the page. It builds on the runtime's [HTMLRewriter](https://developers.cloudflare.com/workers/runtime-apis/html-rewriter/), so it never holds a DOM: a page of several megabytes costs about as much memory as a small one.
4
+
5
+ ## Install
6
+
7
+ ```bash
8
+ npm install @citation-media/html-parser
9
+ ```
10
+
11
+ The package runs wherever `HTMLRewriter` is a global: Cloudflare Workers, `wrangler dev`, and `@cloudflare/vitest-pool-workers`. It has no dependencies.
12
+
13
+ ## Parse A Page
14
+
15
+ `parseHtml` returns only what you ask for. The input is the page as text or a `Response`, such as the result of `fetch`.
16
+
17
+ ```ts
18
+ import { parseHtml } from "@citation-media/html-parser";
19
+
20
+ const response = await fetch("https://example.com/");
21
+ const { links, images, meta, markdown } = await parseHtml(response, {
22
+ url: response.url,
23
+ links: { kinds: ["internal"], resources: false },
24
+ images: true,
25
+ meta: true,
26
+ markdown: true,
27
+ });
28
+ ```
29
+
30
+ | Option | What it collects |
31
+ | --- | --- |
32
+ | `url` | Required. Resolves relative URLs and tells internal links from external ones. |
33
+ | `links` | `true` or a filter: `kinds` (`internal`, `external`, `special`, `anchor`) and `resources: false` to drop downloads such as PDFs. |
34
+ | `images` | Images with `alt` (`null` when missing, `""` when decorative), `width`, `height`, `loading`, and whether they have a `srcset`. |
35
+ | `resources` | Scripts with their `src`, `type`, and attribute names, frame URLs, and `<link>` elements with their `rel`. |
36
+ | `meta` | Title, description, language, canonical URL, robots, viewport, generator, and Open Graph image. |
37
+ | `headings` | The first 200 headings with their level and text, and `headingCounts` per level without a limit. |
38
+ | `classPrefixes` | Counts of elements whose class starts with a prefix, such as `["elementor-", "brxe-", "wp-block-"]`. |
39
+ | `markdown` | `true` or options: `links` (default true), `images` (default false), `maxLength` (default 100 000). |
40
+ | `context` | A CSS selector such as `main` or `article`: only content inside matching elements counts. When nothing matches, the whole page counts, if the input is text. |
41
+ | `skipAriaHidden` | Leaves `aria-hidden` content out of the Markdown. Default true; set it to false for DOMs taken from a browser while a consent dialog is open, which hides the page that way. |
42
+
43
+ Links are deduplicated by URL, counted, and sorted by count. Each carries its `text`, `kind`, `resource`, `count`, and, where present, `rel`, `target`, and `aria-label`. URLs are absolute; paths keep their case.
44
+
45
+ The Markdown leaves out navigation, footer, forms, scripts, consent dialogs, visually hidden helpers, and binary or EXIF text, so it holds what a reader reads.
46
+
47
+ ## Clean A Page
48
+
49
+ `cleanHtml` returns HTML without what you remove. EXIF and XMP blocks and binary-looking text are always removed.
50
+
51
+ ```ts
52
+ import { cleanHtml } from "@citation-media/html-parser";
53
+
54
+ const html = await cleanHtml(page, {
55
+ url: "https://example.com/",
56
+ remove: {
57
+ scripts: true,
58
+ styles: true,
59
+ comments: true,
60
+ cookieConsent: true,
61
+ classes: true,
62
+ },
63
+ links: { kinds: ["internal"] },
64
+ context: "main",
65
+ });
66
+ ```
67
+
68
+ `remove` accepts `scripts`, `styles` (with style attributes and inline SVG), `comments`, `cookieConsent`, `ariaHidden`, `images`, `classes`, `ids`, `meta`, `linkTags`, and `hyphenation`. Links that `links` does not keep become their text.
69
+
70
+ ## Markdown By Workers AI
71
+
72
+ `markdownWithAi` cleans the page and converts it with Workers AI's document conversion. It handles tables and nested formatting better than the streaming Markdown, at the cost of a service call.
73
+
74
+ ```ts
75
+ import { markdownWithAi } from "@citation-media/html-parser";
76
+
77
+ const markdown = await markdownWithAi(env.AI, page, { url, images: true });
78
+ ```
79
+
80
+ ## Limits
81
+
82
+ The streaming Markdown follows HTMLRewriter's model: tables become rows without alignment, and bold or italic formatting is dropped. Use `markdownWithAi` where that matters. Elements that are opened but never closed, such as `<span/>`, keep their state to the end of the page.
83
+
84
+ ## Develop
85
+
86
+ Development needs Node.js 22.18 or later, which loads the TypeScript configuration of Oxlint and Oxfmt; `.nvmrc` pins Node.js 24.
87
+
88
+ ```bash
89
+ npm ci
90
+ npm test
91
+ npm run check
92
+ npm run build
93
+ ```
94
+
95
+ Tests run in workerd through `@cloudflare/vitest-pool-workers`, the runtime the package targets.
96
+
97
+ ## Release
98
+
99
+ Bump `version` in `package.json` and add a section to `CHANGELOG.md` on `main`, then push a matching tag such as `v0.2.0`. The release workflow publishes to npm with provenance and creates the GitHub release.
100
+
101
+ ## License
102
+
103
+ MIT
@@ -0,0 +1,21 @@
1
+ import type { CleanOptions } from "./clean.ts";
2
+ /**
3
+ * Markdown by Workers AI's document conversion, the second way besides the streaming Markdown of
4
+ * `parseHtml`. It handles tables and nested formatting better, at the cost of a service call. The
5
+ * page is cleaned first: scripts, styles, consent banners, and hidden helpers only cost tokens.
6
+ */
7
+ /** The part of the Workers AI binding this needs, so callers can pass `env.AI` as it is. */
8
+ export interface MarkdownConverter {
9
+ toMarkdown: (files: {
10
+ name: string;
11
+ blob: Blob;
12
+ }[]) => Promise<{
13
+ format: string;
14
+ data?: string;
15
+ }[]>;
16
+ }
17
+ export interface AiMarkdownOptions extends Pick<CleanOptions, "url" | "links" | "context"> {
18
+ /** Keeps images in the Markdown. Default false. */
19
+ images?: boolean;
20
+ }
21
+ export declare const markdownWithAi: (ai: MarkdownConverter, input: string | Response, { images, ...options }: AiMarkdownOptions) => Promise<string>;
@@ -0,0 +1,23 @@
1
+ import { cleanHtml } from "./clean.js";
2
+ export const markdownWithAi = async (ai, input, { images = false, ...options }) => {
3
+ const html = await cleanHtml(input, {
4
+ ...options,
5
+ remove: {
6
+ ariaHidden: true,
7
+ classes: true,
8
+ comments: true,
9
+ cookieConsent: true,
10
+ hyphenation: true,
11
+ ids: true,
12
+ images: !images,
13
+ linkTags: true,
14
+ meta: true,
15
+ scripts: true,
16
+ styles: true,
17
+ },
18
+ });
19
+ const [converted] = await ai.toMarkdown([
20
+ { blob: new Blob([html], { type: "text/html" }), name: "index.html" },
21
+ ]);
22
+ return converted?.format === "markdown" ? (converted.data ?? "") : "";
23
+ };
@@ -0,0 +1,32 @@
1
+ import type { LinkFilter } from "./parse.ts";
2
+ /**
3
+ * Cleans a page in one streaming pass, for callers that need HTML rather than extracted facts,
4
+ * such as a Markdown conversion by Workers AI. EXIF and XMP blobs and binary-looking text are
5
+ * always removed; everything else is removed only when asked.
6
+ */
7
+ export interface CleanOptions {
8
+ /** The page's address, needed to filter links by kind. */
9
+ url: string | URL;
10
+ remove?: {
11
+ scripts?: boolean;
12
+ /** Style elements, style attributes, and inline SVG. */
13
+ styles?: boolean;
14
+ comments?: boolean;
15
+ cookieConsent?: boolean;
16
+ ariaHidden?: boolean;
17
+ /** img, picture, and figure elements. */
18
+ images?: boolean;
19
+ classes?: boolean;
20
+ ids?: boolean;
21
+ meta?: boolean;
22
+ linkTags?: boolean;
23
+ /** Soft hyphens and `<wbr>`. */
24
+ hyphenation?: boolean;
25
+ };
26
+ /** Keeps only links that match; others become their text. */
27
+ links?: LinkFilter;
28
+ /** Keeps only the content of elements matching this selector. */
29
+ context?: string;
30
+ }
31
+ /** The cleaned page as HTML. */
32
+ export declare const cleanHtml: (input: string | Response, options: CleanOptions) => Promise<string>;
package/dist/clean.js ADDED
@@ -0,0 +1,143 @@
1
+ import { looksBinary } from "./text.js";
2
+ import { absoluteUrl, isResource, linkKind } from "./urls.js";
3
+ const cookieConsent = [
4
+ "[id^='brlbs-cmpnt']",
5
+ "[class^='brlbs-cmpnt']",
6
+ "[consent-skip-blocker]",
7
+ "#usercentrics-root",
8
+ "div[id*='consent-modal']",
9
+ "div[class*='consent-modal']",
10
+ "div[id*='consent-wrapper']",
11
+ "div[class*='consent-wrapper']",
12
+ "div[id*='consent-banner']",
13
+ "div[class*='consent-banner']",
14
+ ];
15
+ const remove = {
16
+ element(element) {
17
+ element.remove();
18
+ },
19
+ };
20
+ /** The cleaned page as HTML. */
21
+ export const cleanHtml = (input, options) => {
22
+ const base = new URL(options.url);
23
+ const removed = options.remove ?? {};
24
+ let rewriter = new HTMLRewriter();
25
+ if (options.context) {
26
+ // Everything outside the context is unwrapped or dropped; the context's subtrees stay whole.
27
+ let inContext = 0;
28
+ rewriter = rewriter
29
+ .on(options.context, {
30
+ element(element) {
31
+ try {
32
+ element.onEndTag(() => {
33
+ inContext -= 1;
34
+ });
35
+ inContext += 1;
36
+ }
37
+ catch {
38
+ // A void element such as <img> has no content to keep.
39
+ }
40
+ },
41
+ })
42
+ .on("*", {
43
+ element(element) {
44
+ // The context handler above runs first, so a matching element already counts as inside.
45
+ if (inContext === 0) {
46
+ element.removeAndKeepContent();
47
+ }
48
+ },
49
+ })
50
+ .onDocument({
51
+ text(chunk) {
52
+ if (inContext === 0) {
53
+ chunk.remove();
54
+ }
55
+ },
56
+ });
57
+ }
58
+ rewriter = rewriter
59
+ .on("*", {
60
+ element(element) {
61
+ if (element.tagName.includes(":")) {
62
+ element.remove();
63
+ return;
64
+ }
65
+ for (const [name, value] of element.attributes) {
66
+ if (name && value && looksBinary(value)) {
67
+ element.removeAttribute(name);
68
+ }
69
+ }
70
+ if (removed.classes) {
71
+ element.removeAttribute("class");
72
+ }
73
+ if (removed.ids) {
74
+ element.removeAttribute("id");
75
+ }
76
+ if (removed.styles) {
77
+ element.removeAttribute("style");
78
+ }
79
+ },
80
+ })
81
+ .onDocument({
82
+ comments(comment) {
83
+ if (removed.comments ||
84
+ /xpacket|xmpmeta|rdf:RDF|Adobe XMP Core|Exif|8BIM/iu.test(comment.text)) {
85
+ comment.remove();
86
+ }
87
+ },
88
+ text(chunk) {
89
+ if (looksBinary(chunk.text)) {
90
+ chunk.remove();
91
+ }
92
+ else if (removed.hyphenation && /&shy;|­/u.test(chunk.text)) {
93
+ chunk.replace(chunk.text.replaceAll(/&shy;|­/gu, ""), { html: true });
94
+ }
95
+ },
96
+ });
97
+ if (removed.scripts) {
98
+ rewriter = rewriter.on("script", remove);
99
+ }
100
+ if (removed.styles) {
101
+ rewriter = rewriter.on("style, svg", remove);
102
+ }
103
+ if (removed.cookieConsent) {
104
+ rewriter = rewriter.on(cookieConsent.join(", "), remove);
105
+ }
106
+ if (removed.ariaHidden) {
107
+ rewriter = rewriter.on("[aria-hidden='true'], .screen-reader-text, .sr-only, .visually-hidden, .skip-link", remove);
108
+ }
109
+ if (removed.images) {
110
+ rewriter = rewriter.on("img, picture, figure", remove);
111
+ }
112
+ if (removed.meta) {
113
+ rewriter = rewriter.on("meta", remove);
114
+ }
115
+ if (removed.linkTags) {
116
+ rewriter = rewriter.on("link", remove);
117
+ }
118
+ if (removed.hyphenation) {
119
+ rewriter = rewriter.on("wbr", remove);
120
+ }
121
+ const { links } = options;
122
+ if (links) {
123
+ rewriter = rewriter.on("a", {
124
+ element(element) {
125
+ const href = element.getAttribute("href");
126
+ if (!href) {
127
+ return;
128
+ }
129
+ const url = absoluteUrl(href, base);
130
+ const kind = linkKind(href, url, base);
131
+ const kept = (!links.kinds || links.kinds.includes(kind)) &&
132
+ (links.resources !== false || !isResource(url));
133
+ if (!kept) {
134
+ element.removeAndKeepContent();
135
+ }
136
+ },
137
+ });
138
+ }
139
+ const response = input instanceof Response
140
+ ? input
141
+ : new Response(input, { headers: { "content-type": "text/html" } });
142
+ return rewriter.transform(response).text();
143
+ };
@@ -0,0 +1,8 @@
1
+ export { markdownWithAi } from "./ai-markdown.ts";
2
+ export type { AiMarkdownOptions, MarkdownConverter } from "./ai-markdown.ts";
3
+ export { cleanHtml } from "./clean.ts";
4
+ export type { CleanOptions } from "./clean.ts";
5
+ export { parseHtml } from "./parse.ts";
6
+ export type { Heading, Image, Link, LinkFilter, MarkdownOptions, Meta, ParseOptions, ParseResult, Resources, Script, } from "./parse.ts";
7
+ export { absoluteUrl, isResource, linkKind, siteOf, specialProtocol, } from "./urls.ts";
8
+ export type { LinkKind } from "./urls.ts";
package/dist/index.js ADDED
@@ -0,0 +1,4 @@
1
+ export { markdownWithAi } from "./ai-markdown.js";
2
+ export { cleanHtml } from "./clean.js";
3
+ export { parseHtml } from "./parse.js";
4
+ export { absoluteUrl, isResource, linkKind, siteOf, specialProtocol, } from "./urls.js";
@@ -0,0 +1,20 @@
1
+ /**
2
+ * Builds Markdown from parser events while the page streams through: blocks for headings,
3
+ * paragraphs, list items, quotes, and table rows, inline links and images, and preformatted text
4
+ * that keeps its whitespace. It follows HTMLRewriter's limits: no aligned tables, and formatting such
5
+ * as bold or italic is dropped, because the content, not its styling, is what the output is for.
6
+ */
7
+ export declare class MarkdownWriter {
8
+ #private;
9
+ constructor(maxLength: number);
10
+ get full(): boolean;
11
+ text(chunk: string): void;
12
+ /** Starts a block such as `h2`, `p`, `li`, `blockquote`, `tr`, or `pre`. */
13
+ open(tag: string): void;
14
+ /** Ends a block; blocks without an end tag end when the next one opens. */
15
+ close(tag: string): void;
16
+ openLink(href: string): void;
17
+ closeLink(): void;
18
+ image(alt: string, src: string): void;
19
+ toString(): string;
20
+ }
@@ -0,0 +1,147 @@
1
+ import { collapseWhitespace, decodeEntities, looksBinary, removeHyphenation, } from "./text.js";
2
+ /**
3
+ * Builds Markdown from parser events while the page streams through: blocks for headings,
4
+ * paragraphs, list items, quotes, and table rows, inline links and images, and preformatted text
5
+ * that keeps its whitespace. It follows HTMLRewriter's limits: no aligned tables, and formatting such
6
+ * as bold or italic is dropped, because the content, not its styling, is what the output is for.
7
+ */
8
+ export class MarkdownWriter {
9
+ #parts = [];
10
+ #text = [];
11
+ #length = 0;
12
+ #maxLength;
13
+ #links = [];
14
+ #lists = [];
15
+ #preformatted = 0;
16
+ constructor(maxLength) {
17
+ this.#maxLength = maxLength;
18
+ }
19
+ get full() {
20
+ return this.#length >= this.#maxLength;
21
+ }
22
+ text(chunk) {
23
+ if (!this.full) {
24
+ this.#text.push(chunk);
25
+ }
26
+ }
27
+ /** Starts a block such as `h2`, `p`, `li`, `blockquote`, `tr`, or `pre`. */
28
+ open(tag) {
29
+ this.#flush();
30
+ const level = /^h(?<level>[1-6])$/u.exec(tag)?.groups?.level;
31
+ if (level) {
32
+ this.#push(`\n\n${"#".repeat(Number(level))} `);
33
+ }
34
+ else if (tag === "li") {
35
+ const list = this.#lists.at(-1);
36
+ if (list?.ordered) {
37
+ list.count += 1;
38
+ }
39
+ const indent = " ".repeat(Math.max(0, this.#lists.length - 1));
40
+ this.#push(`\n${indent}${list?.ordered ? `${list.count}.` : "-"} `);
41
+ }
42
+ else if (tag === "ul" || tag === "ol") {
43
+ this.#lists.push({ count: 0, ordered: tag === "ol" });
44
+ this.#push("\n");
45
+ }
46
+ else if (tag === "blockquote") {
47
+ this.#push("\n\n> ");
48
+ }
49
+ else if (tag === "tr") {
50
+ this.#push("\n|");
51
+ }
52
+ else if (tag === "td" || tag === "th") {
53
+ this.#push(" ");
54
+ }
55
+ else if (tag === "pre") {
56
+ this.#preformatted += 1;
57
+ this.#push("\n\n```\n");
58
+ }
59
+ else if (tag === "br") {
60
+ this.#push("\n");
61
+ }
62
+ else if (tag === "hr") {
63
+ this.#push("\n\n---\n\n");
64
+ }
65
+ else {
66
+ this.#push("\n\n");
67
+ }
68
+ }
69
+ /** Ends a block; blocks without an end tag end when the next one opens. */
70
+ close(tag) {
71
+ this.#flush();
72
+ if (tag === "ul" || tag === "ol") {
73
+ this.#lists.pop();
74
+ this.#push("\n");
75
+ }
76
+ else if (tag === "td" || tag === "th") {
77
+ this.#push(" |");
78
+ }
79
+ else if (tag === "pre") {
80
+ this.#preformatted = Math.max(0, this.#preformatted - 1);
81
+ this.#push("\n```\n\n");
82
+ }
83
+ else if (tag !== "li" && tag !== "tr") {
84
+ // List items and table rows end where the next one starts, so lists stay tight.
85
+ this.#push("\n\n");
86
+ }
87
+ }
88
+ openLink(href) {
89
+ this.#flush();
90
+ this.#links.push({ href, index: this.#parts.length });
91
+ this.#push("[");
92
+ }
93
+ closeLink() {
94
+ this.#flush();
95
+ const link = this.#links.pop();
96
+ if (!link) {
97
+ return;
98
+ }
99
+ // A link without text, such as an icon, keeps nothing: an empty [](url) helps no reader.
100
+ const label = this.#parts
101
+ .slice(link.index + 1)
102
+ .join("")
103
+ .trim();
104
+ if (!label) {
105
+ this.#parts.splice(link.index);
106
+ return;
107
+ }
108
+ this.#push(`](${link.href})`);
109
+ }
110
+ image(alt, src) {
111
+ this.#flush();
112
+ this.#push(`![${collapseWhitespace(alt).trim()}](${src})`);
113
+ }
114
+ toString() {
115
+ this.#flush();
116
+ // Code blocks keep their whitespace; everything between them is tidied.
117
+ return this.#parts
118
+ .join("")
119
+ .split(/(?<code>```[\s\S]*?```)/u)
120
+ .map((segment) => segment.startsWith("```")
121
+ ? segment
122
+ : segment
123
+ .replaceAll(/[ \t]+\n/gu, "\n")
124
+ .replaceAll(/\n[ \t]+(?=[^-\d> ])/gu, "\n"))
125
+ .join("")
126
+ .replaceAll(/\n{3,}/gu, "\n\n")
127
+ .trim()
128
+ .slice(0, this.#maxLength);
129
+ }
130
+ #flush() {
131
+ if (this.#text.length === 0) {
132
+ return;
133
+ }
134
+ const raw = removeHyphenation(decodeEntities(this.#text.join("")));
135
+ this.#text = [];
136
+ if (looksBinary(raw)) {
137
+ return;
138
+ }
139
+ this.#push(this.#preformatted > 0 ? raw : collapseWhitespace(raw));
140
+ }
141
+ #push(part) {
142
+ if (!this.full) {
143
+ this.#parts.push(part);
144
+ this.#length += part.length;
145
+ }
146
+ }
147
+ }
@@ -0,0 +1,108 @@
1
+ import type { LinkKind } from "./urls.ts";
2
+ /**
3
+ * One streaming pass over a page with Cloudflare's HTMLRewriter, which never builds a DOM: it
4
+ * collects links, images, resources, metadata, headings, and class counts, and writes Markdown at
5
+ * the same time. Memory stays flat even for pages of several megabytes.
6
+ */
7
+ export interface Link {
8
+ url: string;
9
+ /** Visible text of the first link to this URL. */
10
+ text: string;
11
+ kind: LinkKind;
12
+ /** Whether the URL points to a file download such as a PDF. */
13
+ resource: boolean;
14
+ /** How often the page links to this URL. */
15
+ count: number;
16
+ rel?: string;
17
+ target?: string;
18
+ /** `aria-label` of the first link, which names links whose text does not. */
19
+ label?: string;
20
+ }
21
+ export interface Image {
22
+ url: string;
23
+ /** `null` when the attribute is missing, `""` when the image is marked as decorative. */
24
+ alt: string | null;
25
+ width?: string;
26
+ height?: string;
27
+ loading?: string;
28
+ srcset: boolean;
29
+ }
30
+ export interface Script {
31
+ src?: string;
32
+ type?: string;
33
+ /** Names of the element's attributes, for consent markers such as `data-cookieconsent`. */
34
+ attributes: string[];
35
+ }
36
+ export interface Resources {
37
+ scripts: Script[];
38
+ frames: string[];
39
+ /** `<link>` elements with their `rel`, such as stylesheets and preconnects. */
40
+ links: {
41
+ href: string;
42
+ rel: string;
43
+ }[];
44
+ }
45
+ export interface Meta {
46
+ title?: string;
47
+ description?: string;
48
+ language?: string;
49
+ canonical?: string;
50
+ robots?: string;
51
+ viewport?: string;
52
+ generator?: string;
53
+ image?: string;
54
+ }
55
+ export interface Heading {
56
+ level: number;
57
+ text: string;
58
+ }
59
+ export interface LinkFilter {
60
+ /** Kinds to keep; all when not set. */
61
+ kinds?: LinkKind[];
62
+ /** Whether to keep links to downloads such as PDFs. Default true. */
63
+ resources?: boolean;
64
+ }
65
+ export interface MarkdownOptions {
66
+ /** Writes `[text](url)` for links. Default true. */
67
+ links?: boolean;
68
+ /** Writes `![alt](url)` for images. Default false. */
69
+ images?: boolean;
70
+ /** Longest output in characters. Default 100 000. */
71
+ maxLength?: number;
72
+ }
73
+ export interface ParseOptions {
74
+ /** The page's address, to resolve relative URLs and tell internal links from external ones. */
75
+ url: string | URL;
76
+ links?: boolean | LinkFilter;
77
+ images?: boolean;
78
+ resources?: boolean;
79
+ meta?: boolean;
80
+ headings?: boolean;
81
+ /** Counts elements whose class starts with a prefix, such as page builders' `elementor-`. */
82
+ classPrefixes?: string[];
83
+ markdown?: boolean | MarkdownOptions;
84
+ /**
85
+ * Collects only inside elements matching this selector, such as `main` or `article`. When
86
+ * nothing matches, the whole page counts; for a `Response` input, nothing is collected then.
87
+ */
88
+ context?: string;
89
+ /**
90
+ * Leaves `aria-hidden` content out of the Markdown. Default true. Set it to false for DOMs taken
91
+ * from a browser while a dialog is open: consent tools hide the whole page behind it that way.
92
+ */
93
+ skipAriaHidden?: boolean;
94
+ }
95
+ export interface ParseResult {
96
+ links?: Link[];
97
+ images?: Image[];
98
+ resources?: Resources;
99
+ meta?: Meta;
100
+ /** The first headings with their text, see `maxHeadings`. */
101
+ headings?: Heading[];
102
+ /** How many headings of each level the page has, uncapped: `headingCounts[0]` counts `h1`. */
103
+ headingCounts?: number[];
104
+ classCounts?: Record<string, number>;
105
+ markdown?: string;
106
+ }
107
+ /** Parses a page, as text or as a `Response` such as a `fetch` result, for what `options` ask. */
108
+ export declare const parseHtml: (input: string | Response, options: ParseOptions) => Promise<ParseResult>;
package/dist/parse.js ADDED
@@ -0,0 +1,380 @@
1
+ import { MarkdownWriter } from "./markdown.js";
2
+ import { collapseWhitespace, decodeEntities, removeHyphenation, } from "./text.js";
3
+ import { absoluteUrl, isResource, linkKind } from "./urls.js";
4
+ const blockTags = new Set([
5
+ "br",
6
+ "hr",
7
+ "h1",
8
+ "h2",
9
+ "h3",
10
+ "h4",
11
+ "h5",
12
+ "h6",
13
+ "p",
14
+ "li",
15
+ "ul",
16
+ "ol",
17
+ "blockquote",
18
+ "pre",
19
+ "tr",
20
+ "td",
21
+ "th",
22
+ "dt",
23
+ "dd",
24
+ "figcaption",
25
+ "section",
26
+ "article",
27
+ "div",
28
+ ]);
29
+ // Content that is not what a reader reads: page chrome, code, and consent dialogs.
30
+ const hiddenFromMarkdown = [
31
+ "script",
32
+ "style",
33
+ "noscript",
34
+ "template",
35
+ "svg",
36
+ "iframe",
37
+ "nav",
38
+ "footer",
39
+ "form",
40
+ "select",
41
+ "[hidden]",
42
+ "[role='dialog']",
43
+ "[aria-modal='true']",
44
+ "[id*='cookie']",
45
+ "[class*='cookie']",
46
+ "[id*='consent']",
47
+ "[class*='consent']",
48
+ "[id^='brlbs-cmpnt']",
49
+ "[class^='brlbs-cmpnt']",
50
+ "#BorlabsCookieBox",
51
+ "#usercentrics-root",
52
+ "#CybotCookiebotDialog",
53
+ "#cc-main",
54
+ ".screen-reader-text",
55
+ ".sr-only",
56
+ ".visually-hidden",
57
+ ];
58
+ const maxHeadings = 200;
59
+ const maxImages = 500;
60
+ const maxResources = 500;
61
+ // Follows an element to its end tag. Void elements such as <img> have none, and HTMLRewriter throws
62
+ // for them; elements closed implicitly, such as <li> without </li>, end with their parent.
63
+ const enter = (element, onEnd) => {
64
+ try {
65
+ element.onEndTag(onEnd);
66
+ return true;
67
+ }
68
+ catch {
69
+ return false;
70
+ }
71
+ };
72
+ const clean = (text) => collapseWhitespace(removeHyphenation(decodeEntities(text))).trim();
73
+ const addLink = (links, found, base) => {
74
+ // Cloudflare's e-mail obfuscation rewrites addresses to a script-decoded link.
75
+ if (found.url.includes("/cdn-cgi/l/email-protection")) {
76
+ return;
77
+ }
78
+ const existing = links.get(found.url);
79
+ if (existing) {
80
+ existing.count += 1;
81
+ return;
82
+ }
83
+ const { href, ...link } = found;
84
+ links.set(found.url, {
85
+ ...link,
86
+ count: 1,
87
+ kind: linkKind(href, found.url, base),
88
+ resource: isResource(found.url),
89
+ });
90
+ };
91
+ const filterLinks = (links, filter) => links
92
+ .filter((link) => (!filter.kinds || filter.kinds.includes(link.kind)) &&
93
+ (filter.resources !== false || !link.resource))
94
+ .toSorted((first, second) => second.count - first.count);
95
+ const collecting = (pass) => pass.inContext > 0;
96
+ const writing = (pass) => collecting(pass) && pass.hidden === 0 && pass.markdown !== undefined;
97
+ const hide = (pass) => ({
98
+ element(element) {
99
+ if (enter(element, () => (pass.hidden -= 1))) {
100
+ pass.hidden += 1;
101
+ }
102
+ },
103
+ });
104
+ const onStructure = (rewriter, pass) => {
105
+ const { context, classPrefixes = [] } = pass.options;
106
+ let next = rewriter;
107
+ if (context) {
108
+ next = next.on(context, {
109
+ element(element) {
110
+ pass.contextFound = true;
111
+ if (enter(element, () => (pass.inContext -= 1))) {
112
+ pass.inContext += 1;
113
+ }
114
+ },
115
+ });
116
+ }
117
+ next = next.on(hiddenFromMarkdown.join(", "), hide(pass));
118
+ if (pass.options.skipAriaHidden ?? true) {
119
+ next = next.on("[aria-hidden='true']", hide(pass));
120
+ }
121
+ return next.on("*", {
122
+ element(element) {
123
+ // Namespaced tags such as <x:xmpmeta> carry image metadata, not content.
124
+ if (element.tagName.includes(":")) {
125
+ hide(pass).element(element);
126
+ }
127
+ if (!collecting(pass)) {
128
+ return;
129
+ }
130
+ const classes = element.getAttribute("class")?.split(/\s+/u) ?? [];
131
+ for (const prefix of classPrefixes) {
132
+ if (classes.some((name) => name.startsWith(prefix))) {
133
+ pass.classCounts[prefix] = (pass.classCounts[prefix] ?? 0) + 1;
134
+ }
135
+ }
136
+ const tag = element.tagName.toLowerCase();
137
+ if (!writing(pass) || !blockTags.has(tag)) {
138
+ return;
139
+ }
140
+ pass.markdown?.open(tag);
141
+ enter(element, () => {
142
+ if (writing(pass)) {
143
+ pass.markdown?.close(tag);
144
+ }
145
+ });
146
+ },
147
+ });
148
+ };
149
+ const onLinks = (rewriter, pass) => rewriter.on("a[href]", {
150
+ element(element) {
151
+ const href = element.getAttribute("href") ?? "";
152
+ if (!collecting(pass) || !href) {
153
+ return;
154
+ }
155
+ const url = absoluteUrl(href, pass.base);
156
+ const text = [];
157
+ pass.link = text;
158
+ // Attributes are readable only while this handler runs, not in the end-tag handler.
159
+ const label = element.getAttribute("aria-label") ?? undefined;
160
+ const rel = element.getAttribute("rel") ?? undefined;
161
+ const target = element.getAttribute("target") ?? undefined;
162
+ const written = writing(pass) &&
163
+ pass.markdownOptions?.links !== false &&
164
+ !href.startsWith("#");
165
+ if (written) {
166
+ pass.markdown?.openLink(url);
167
+ }
168
+ enter(element, () => {
169
+ if (written) {
170
+ pass.markdown?.closeLink();
171
+ }
172
+ if (pass.linkFilter) {
173
+ addLink(pass.links, { href, label, rel, target, text: clean(text.join("")), url }, pass.base);
174
+ }
175
+ pass.link = undefined;
176
+ });
177
+ },
178
+ });
179
+ const onImages = (rewriter, pass) => rewriter.on("img", {
180
+ element(element) {
181
+ // Lazy-loading scripts often keep the real address in data-src until the image scrolls in.
182
+ const src = element.getAttribute("src") ?? element.getAttribute("data-src");
183
+ if (!collecting(pass) || !src || src.startsWith("data:")) {
184
+ return;
185
+ }
186
+ const url = absoluteUrl(src, pass.base);
187
+ const alt = element.getAttribute("alt");
188
+ if (pass.options.images &&
189
+ pass.images.size < maxImages &&
190
+ !pass.images.has(url)) {
191
+ pass.images.set(url, {
192
+ alt: alt === null ? null : clean(alt),
193
+ height: element.getAttribute("height") ?? undefined,
194
+ loading: element.getAttribute("loading") ?? undefined,
195
+ srcset: element.hasAttribute("srcset"),
196
+ url,
197
+ width: element.getAttribute("width") ?? undefined,
198
+ });
199
+ }
200
+ if (writing(pass) && pass.markdownOptions?.images) {
201
+ pass.markdown?.image(alt ?? "", url);
202
+ }
203
+ },
204
+ });
205
+ const onResources = (rewriter, pass) => {
206
+ const { resources, base } = pass;
207
+ return rewriter
208
+ .on("script", {
209
+ element(element) {
210
+ if (resources.scripts.length < maxResources) {
211
+ const src = element.getAttribute("src");
212
+ resources.scripts.push({
213
+ attributes: [...element.attributes].flatMap(([name]) => name ? [name] : []),
214
+ src: src ? absoluteUrl(src, base) : undefined,
215
+ type: element.getAttribute("type") ?? undefined,
216
+ });
217
+ }
218
+ },
219
+ })
220
+ .on("iframe[src]", {
221
+ element(element) {
222
+ if (resources.frames.length < maxResources) {
223
+ resources.frames.push(absoluteUrl(element.getAttribute("src") ?? "", base));
224
+ }
225
+ },
226
+ })
227
+ .on("link[href]", {
228
+ element(element) {
229
+ if (resources.links.length < maxResources) {
230
+ resources.links.push({
231
+ href: absoluteUrl(element.getAttribute("href") ?? "", base),
232
+ rel: element.getAttribute("rel") ?? "",
233
+ });
234
+ }
235
+ },
236
+ });
237
+ };
238
+ const metaFields = new Map([
239
+ ["description", "description"],
240
+ ["generator", "generator"],
241
+ ["og:image", "image"],
242
+ ["robots", "robots"],
243
+ ["viewport", "viewport"],
244
+ ]);
245
+ const onMeta = (rewriter, pass) => rewriter
246
+ .on("html", {
247
+ element(element) {
248
+ pass.meta.language = element.getAttribute("lang") ?? undefined;
249
+ },
250
+ })
251
+ .on("title", {
252
+ element(element) {
253
+ pass.title = [];
254
+ enter(element, () => {
255
+ pass.meta.title = clean((pass.title ?? []).join(""));
256
+ pass.title = undefined;
257
+ });
258
+ },
259
+ })
260
+ .on("meta", {
261
+ element(element) {
262
+ const name = element.getAttribute("name") ??
263
+ element.getAttribute("property") ??
264
+ "";
265
+ const field = metaFields.get(name.toLowerCase());
266
+ if (field) {
267
+ pass.meta[field] = element.getAttribute("content") ?? undefined;
268
+ }
269
+ },
270
+ })
271
+ .on("link[rel='canonical']", {
272
+ element(element) {
273
+ const href = element.getAttribute("href");
274
+ pass.meta.canonical = href ? absoluteUrl(href, pass.base) : undefined;
275
+ },
276
+ });
277
+ const onHeadings = (rewriter, pass) => rewriter.on("h1, h2, h3, h4, h5, h6", {
278
+ element(element) {
279
+ if (!collecting(pass)) {
280
+ return;
281
+ }
282
+ const level = Number(element.tagName.slice(1));
283
+ pass.headingCounts[level - 1] = (pass.headingCounts[level - 1] ?? 0) + 1;
284
+ pass.heading = { level, text: [] };
285
+ enter(element, () => {
286
+ const text = clean(pass.heading?.text.join("") ?? "");
287
+ if (pass.heading && text && pass.headings.length < maxHeadings) {
288
+ pass.headings.push({ level: pass.heading.level, text });
289
+ }
290
+ pass.heading = undefined;
291
+ });
292
+ },
293
+ });
294
+ const onText = (rewriter, pass) => rewriter.onDocument({
295
+ text(chunk) {
296
+ pass.title?.push(chunk.text);
297
+ if (!collecting(pass)) {
298
+ return;
299
+ }
300
+ pass.link?.push(chunk.text);
301
+ pass.heading?.text.push(chunk.text);
302
+ if (writing(pass)) {
303
+ pass.markdown?.text(chunk.text);
304
+ }
305
+ },
306
+ });
307
+ const resultOf = (pass) => {
308
+ const { options } = pass;
309
+ const result = {};
310
+ if (pass.linkFilter) {
311
+ result.links = filterLinks([...pass.links.values()], pass.linkFilter);
312
+ }
313
+ if (options.images) {
314
+ result.images = [...pass.images.values()];
315
+ }
316
+ if (options.resources) {
317
+ result.resources = pass.resources;
318
+ }
319
+ if (options.meta) {
320
+ result.meta = pass.meta;
321
+ }
322
+ if (options.headings) {
323
+ result.headings = pass.headings;
324
+ result.headingCounts = pass.headingCounts;
325
+ }
326
+ if (options.classPrefixes) {
327
+ result.classCounts = pass.classCounts;
328
+ }
329
+ if (pass.markdown) {
330
+ result.markdown = pass.markdown.toString();
331
+ }
332
+ return result;
333
+ };
334
+ /** Parses a page, as text or as a `Response` such as a `fetch` result, for what `options` ask. */
335
+ export const parseHtml = async (input, options) => {
336
+ const markdownOptions = options.markdown === true ? {} : options.markdown || undefined;
337
+ const pass = {
338
+ base: new URL(options.url),
339
+ classCounts: Object.fromEntries((options.classPrefixes ?? []).map((prefix) => [prefix, 0])),
340
+ contextFound: false,
341
+ heading: undefined,
342
+ headingCounts: [0, 0, 0, 0, 0, 0],
343
+ headings: [],
344
+ hidden: 0,
345
+ images: new Map(),
346
+ inContext: options.context ? 0 : 1,
347
+ link: undefined,
348
+ linkFilter: options.links === true ? {} : options.links || undefined,
349
+ links: new Map(),
350
+ markdown: markdownOptions
351
+ ? new MarkdownWriter(markdownOptions.maxLength ?? 100_000)
352
+ : undefined,
353
+ markdownOptions,
354
+ meta: {},
355
+ options,
356
+ resources: { frames: [], links: [], scripts: [] },
357
+ title: undefined,
358
+ };
359
+ let rewriter = onImages(onLinks(onStructure(new HTMLRewriter(), pass), pass), pass);
360
+ if (options.resources) {
361
+ rewriter = onResources(rewriter, pass);
362
+ }
363
+ if (options.meta) {
364
+ rewriter = onMeta(rewriter, pass);
365
+ }
366
+ if (options.headings) {
367
+ rewriter = onHeadings(rewriter, pass);
368
+ }
369
+ const response = input instanceof Response
370
+ ? input
371
+ : new Response(input, { headers: { "content-type": "text/html" } });
372
+ // The rewritten page is discarded; the handlers are what this pass is for.
373
+ await onText(rewriter, pass).transform(response).arrayBuffer();
374
+ // Like a selector that matches nothing in a DOM library, a missing context means the whole page.
375
+ // A response can be read only once, so this second pass needs the page as text.
376
+ if (options.context && !pass.contextFound && !(input instanceof Response)) {
377
+ return parseHtml(input, { ...options, context: undefined });
378
+ }
379
+ return resultOf(pass);
380
+ };
package/dist/text.d.ts ADDED
@@ -0,0 +1,16 @@
1
+ /**
2
+ * Text as HTMLRewriter delivers it: raw source text in chunks, with entities still encoded. The
3
+ * parser collects chunks per run of text and decodes them here, so an entity split across two
4
+ * chunks is still decoded correctly.
5
+ */
6
+ /** Decodes numeric and common named entities; unknown names stay as written. */
7
+ export declare const decodeEntities: (text: string) => string;
8
+ /** Soft hyphens and zero-width characters that split words for layout only. */
9
+ export declare const removeHyphenation: (text: string) => string;
10
+ /** Whitespace collapsed to single spaces, as a browser renders normal text. */
11
+ export declare const collapseWhitespace: (text: string) => string;
12
+ /**
13
+ * Whether text looks like binary data or an EXIF/XMP blob rather than words. The thresholds stay
14
+ * conservative so umlauts and other normal characters never trigger it.
15
+ */
16
+ export declare const looksBinary: (text: string) => boolean;
package/dist/text.js ADDED
@@ -0,0 +1,77 @@
1
+ /**
2
+ * Text as HTMLRewriter delivers it: raw source text in chunks, with entities still encoded. The
3
+ * parser collects chunks per run of text and decodes them here, so an entity split across two
4
+ * chunks is still decoded correctly.
5
+ */
6
+ const named = new Map(Object.entries({
7
+ Auml: "Ä",
8
+ Ouml: "Ö",
9
+ Uuml: "Ü",
10
+ amp: "&",
11
+ apos: "'",
12
+ auml: "ä",
13
+ bdquo: "„",
14
+ bull: "•",
15
+ copy: "©",
16
+ euro: "€",
17
+ gt: ">",
18
+ hellip: "…",
19
+ laquo: "«",
20
+ ldquo: "“",
21
+ lsquo: "‘",
22
+ lt: "<",
23
+ mdash: "—",
24
+ middot: "·",
25
+ nbsp: " ",
26
+ ndash: "–",
27
+ ouml: "ö",
28
+ quot: '"',
29
+ raquo: "»",
30
+ rdquo: "”",
31
+ reg: "®",
32
+ rsquo: "’",
33
+ sbquo: "‚",
34
+ shy: "",
35
+ szlig: "ß",
36
+ trade: "™",
37
+ uuml: "ü",
38
+ }));
39
+ const fromCodePoint = (value) => {
40
+ try {
41
+ return String.fromCodePoint(value);
42
+ }
43
+ catch {
44
+ return "";
45
+ }
46
+ };
47
+ /** Decodes numeric and common named entities; unknown names stay as written. */
48
+ export const decodeEntities = (text) => text.replaceAll(/&(?:#(?<decimal>\d+)|#x(?<hex>[\da-f]+)|(?<name>[a-z]+\d?));/giu, (entity, decimal, hex, name) => {
49
+ if (decimal) {
50
+ return fromCodePoint(Number(decimal));
51
+ }
52
+ if (hex) {
53
+ return fromCodePoint(Number.parseInt(hex, 16));
54
+ }
55
+ return named.get(name ?? "") ?? entity;
56
+ });
57
+ /** Soft hyphens and zero-width characters that split words for layout only. */
58
+ export const removeHyphenation = (text) => text.replaceAll(/[­​]/gu, "");
59
+ /** Whitespace collapsed to single spaces, as a browser renders normal text. */
60
+ export const collapseWhitespace = (text) => text.replaceAll(/\s+/gu, " ");
61
+ // Control characters other than tab and line breaks, which words never contain.
62
+ // oxlint-disable-next-line no-control-regex -- control characters are what this detects.
63
+ const controlCharacters = /[\u0000-\u0008\u000B\u000C\u000E-\u001F]/gu;
64
+ /**
65
+ * Whether text looks like binary data or an EXIF/XMP blob rather than words. The thresholds stay
66
+ * conservative so umlauts and other normal characters never trigger it.
67
+ */
68
+ export const looksBinary = (text) => {
69
+ if (text.length < 16) {
70
+ return false;
71
+ }
72
+ if (/8BIM|Exif|Photoshop 3\.0|Adobe XMP Core|xmpmeta|rdf:RDF|xpacket/iu.test(text)) {
73
+ return true;
74
+ }
75
+ const control = text.match(controlCharacters)?.length ?? 0;
76
+ return (text.includes("\u0000") || control > 20 || control / text.length > 0.05);
77
+ };
package/dist/urls.d.ts ADDED
@@ -0,0 +1,19 @@
1
+ /**
2
+ * URL rules shared by link and image extraction: resolving against the page, telling the site's
3
+ * own links from others, and recognising special protocols and downloadable resources.
4
+ */
5
+ export type LinkKind = "internal" | "external" | "special" | "anchor";
6
+ /** The special protocol of a URL, such as `mailto:`, or undefined for web and relative URLs. */
7
+ export declare const specialProtocol: (url: string) => string | undefined;
8
+ /** The registrable part of a host, good enough to tell a site's own hosts from third parties. */
9
+ export declare const siteOf: (host: string) => string;
10
+ /**
11
+ * Resolves `url` against the page. Scheme and host are normalised by `URL`; paths keep their case,
12
+ * because servers treat them as case-sensitive. `mail:` becomes `mailto:` and addresses are
13
+ * lowercased. Special URLs other than mail stay as written.
14
+ */
15
+ export declare const absoluteUrl: (url: string, base: URL) => string;
16
+ /** Whether the URL points to a file download such as a PDF, rather than a page. */
17
+ export declare const isResource: (url: string) => boolean;
18
+ /** What a link is, judged from its written `href` and its resolved URL. */
19
+ export declare const linkKind: (href: string, resolved: string, base: URL) => LinkKind;
package/dist/urls.js ADDED
@@ -0,0 +1,92 @@
1
+ /**
2
+ * URL rules shared by link and image extraction: resolving against the page, telling the site's
3
+ * own links from others, and recognising special protocols and downloadable resources.
4
+ */
5
+ const specialProtocols = [
6
+ "mailto:",
7
+ "mail:",
8
+ "tel:",
9
+ "sms:",
10
+ "ftp:",
11
+ "ftps:",
12
+ "ssh:",
13
+ "git:",
14
+ "svn:",
15
+ "news:",
16
+ "irc:",
17
+ "xmpp:",
18
+ "webcal:",
19
+ "skype:",
20
+ "slack:",
21
+ "spotify:",
22
+ "steam:",
23
+ "discord:",
24
+ // oxlint-disable-next-line eslint/no-script-url -- recognised so such links are not followed.
25
+ "javascript:",
26
+ "data:",
27
+ "blob:",
28
+ ];
29
+ const resourceExtensions = /\.(?:pdf|docx?|xlsx?|pptx?|zip|rar|tar|gz|mp3|mp4|avi|mov|jpe?g|png|gif|svg|webp|avif|csv)$/iu;
30
+ // Second-level labels under which registrable domains sit one level deeper, such as example.co.uk.
31
+ const secondLevels = /^(?:co|com|org|net|gov|gv|ac|edu)$/u;
32
+ /** The special protocol of a URL, such as `mailto:`, or undefined for web and relative URLs. */
33
+ export const specialProtocol = (url) => {
34
+ const lower = url.trim().toLowerCase();
35
+ return specialProtocols.find((protocol) => lower.startsWith(protocol));
36
+ };
37
+ /** The registrable part of a host, good enough to tell a site's own hosts from third parties. */
38
+ export const siteOf = (host) => {
39
+ const parts = host
40
+ .toLowerCase()
41
+ .replace(/^www\./u, "")
42
+ .split(".");
43
+ return parts.slice(secondLevels.test(parts.at(-2) ?? "") ? -3 : -2).join(".");
44
+ };
45
+ /**
46
+ * Resolves `url` against the page. Scheme and host are normalised by `URL`; paths keep their case,
47
+ * because servers treat them as case-sensitive. `mail:` becomes `mailto:` and addresses are
48
+ * lowercased. Special URLs other than mail stay as written.
49
+ */
50
+ export const absoluteUrl = (url, base) => {
51
+ const trimmed = url.trim();
52
+ const protocol = specialProtocol(trimmed);
53
+ if (protocol === "mail:" || protocol === "mailto:") {
54
+ return `mailto:${trimmed.slice(protocol.length).toLowerCase()}`;
55
+ }
56
+ if (protocol) {
57
+ return trimmed;
58
+ }
59
+ try {
60
+ return new URL(trimmed, base).href;
61
+ }
62
+ catch {
63
+ return trimmed;
64
+ }
65
+ };
66
+ /** Whether the URL points to a file download such as a PDF, rather than a page. */
67
+ export const isResource = (url) => {
68
+ try {
69
+ return resourceExtensions.test(new URL(url).pathname);
70
+ }
71
+ catch {
72
+ return resourceExtensions.test(url);
73
+ }
74
+ };
75
+ /** What a link is, judged from its written `href` and its resolved URL. */
76
+ export const linkKind = (href, resolved, base) => {
77
+ const written = href.trim();
78
+ if (specialProtocol(written)) {
79
+ return "special";
80
+ }
81
+ if (written.startsWith("#") || written.startsWith("/#")) {
82
+ return "anchor";
83
+ }
84
+ try {
85
+ return siteOf(new URL(resolved).hostname) === siteOf(base.hostname)
86
+ ? "internal"
87
+ : "external";
88
+ }
89
+ catch {
90
+ return "external";
91
+ }
92
+ };
package/package.json ADDED
@@ -0,0 +1,59 @@
1
+ {
2
+ "name": "@citation-media/html-parser",
3
+ "version": "0.1.0",
4
+ "description": "Streaming HTML parser for Cloudflare Workers: links, images, resources, and Markdown without a DOM, built on HTMLRewriter.",
5
+ "keywords": [
6
+ "cloudflare-workers",
7
+ "html",
8
+ "htmlrewriter",
9
+ "markdown",
10
+ "parser",
11
+ "scraping"
12
+ ],
13
+ "homepage": "https://github.com/Citation-Media/html-parser#readme",
14
+ "bugs": "https://github.com/Citation-Media/html-parser/issues",
15
+ "license": "MIT",
16
+ "repository": {
17
+ "type": "git",
18
+ "url": "git+https://github.com/Citation-Media/html-parser.git"
19
+ },
20
+ "files": [
21
+ "dist",
22
+ "README.md",
23
+ "LICENSE"
24
+ ],
25
+ "type": "module",
26
+ "sideEffects": false,
27
+ "main": "./dist/index.js",
28
+ "types": "./dist/index.d.ts",
29
+ "exports": {
30
+ ".": {
31
+ "types": "./dist/index.d.ts",
32
+ "default": "./dist/index.js"
33
+ }
34
+ },
35
+ "publishConfig": {
36
+ "access": "public"
37
+ },
38
+ "scripts": {
39
+ "build": "tsc -p tsconfig.build.json",
40
+ "check": "ultracite check && tsc --noEmit",
41
+ "fix": "ultracite fix",
42
+ "prepare": "npm run build",
43
+ "prepublishOnly": "npm run build && npm run check && npm test",
44
+ "test": "vitest run",
45
+ "typecheck": "tsc --noEmit"
46
+ },
47
+ "devDependencies": {
48
+ "@cloudflare/vitest-pool-workers": "0.22.0",
49
+ "@cloudflare/workers-types": "5.20260929.1",
50
+ "oxfmt": "0.68.0",
51
+ "oxlint": "1.83.0",
52
+ "typescript": "6.0.3",
53
+ "ultracite": "7.12.0",
54
+ "vitest": "4.1.11"
55
+ },
56
+ "engines": {
57
+ "node": ">=22.18.0"
58
+ }
59
+ }