decant-core 1.10.1 → 1.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -4,7 +4,7 @@
4
4
 
5
5
  [![npm version](https://img.shields.io/npm/v/decant-core?logo=npm&logoColor=white&label=npm&color=cb3837)](https://www.npmjs.com/package/decant-core)
6
6
  [![License: AGPL-3.0](https://img.shields.io/badge/License-AGPL--3.0-red.svg)](LICENSE)
7
- [![GitHub](https://img.shields.io/github/stars/Covai-Labs/decant-core?logo=github&logoColor=white&color=yellow&label=Stars)](https://github.com/Covai-Labs/decant-core/stargazers)
7
+ [![GitHub](https://img.shields.io/github/stars/Covai-Labs/decant?logo=github&logoColor=white&color=yellow&label=Stars)](https://github.com/Covai-Labs/decant/stargazers)
8
8
 
9
9
  Every AI chat exporter ends up solving the same problem: extracting conversations from ChatGPT, Claude, Gemini, Perplexity, DeepSeek and other constantly changing AI interfaces.
10
10
 
@@ -41,7 +41,7 @@ AI platforms don't expose stable public APIs for reading conversation history. N
41
41
 
42
42
  Maintaining that per-platform logic in every exporter is wasteful and fragile. `decant-core` centralizes it:
43
43
 
44
- - ✅ **18 AI chat platform parsers** with normalized output — you get structured messages, models, metadata and Markdown, not DOM soup.
44
+ - ✅ **19 AI chat platform parsers** with normalized output — you get structured messages, models, metadata and Markdown, not DOM soup.
45
45
  - ✅ **Web article extraction** — Mozilla Readability, Defuddle, and Article-Extractor run in parallel and arbitrate by content-quality scoring.
46
46
  - ✅ **Detection utilities** — tell an "AI chat page" apart from a "regular web page" before you decide which parser to run.
47
47
  - ✅ **Math & Markdown handling** — LaTeX normalization plus GFM tables/code fencing that survive round-trips into Obsidian, Logseq and Notion.
@@ -117,34 +117,35 @@ import { normalizeLatexMath } from "decant-core";
117
117
 
118
118
  ## Supported Platforms
119
119
 
120
- 18 AI chat platform parsers plus generic web article extraction:
121
-
122
- | Platform | Parser | Extraction strategy |
123
- | :---------------------------------- | :------------------------ | :----------------------------------------- |
124
- | **ChatGPT** | `ChatGPTParser` | DOM + internal API |
125
- | **Claude** | `ClaudeParser` | DOM + internal API + React fiber |
126
- | **Google Gemini** | `GeminiParser` | DOM + batchexecute RPC |
127
- | **Microsoft Copilot** | `CopilotParser` | DOM (multi-domain) |
128
- | **Perplexity** | `PerplexityParser` | Internal API + DOM fallback |
129
- | **DeepSeek** | `DeepSeekParser` | DOM + internal API (`fragments[]`) |
130
- | **Qwen** | `QwenParser` | DOM |
131
- | **Meta AI** | `MetaParser` | Internal API (GraphQL) + DOM fallback |
132
- | **Mistral / Le Chat** | `MistralParser` | DOM |
133
- | **Proton Lumo** | `LumoParser` | DOM (API is E2E-encrypted, not readable) |
134
- | **Z.ai** | `ZAiParser` | Internal API (chat + batch) + DOM fallback |
135
- | **Grok** | `GrokParser` | Internal API (response-node + load) + DOM |
136
- | **Google AI Studio** | `GoogleAIStudioParser` | DOM |
137
- | **NotebookLM** | `NotebookLMParser` | DOM |
138
- | **Google Search AI (AI Overviews)** | `GoogleSearchAIParser` | DOM |
139
- | **Gemini Cloud Assist** | `GeminiCloudAssistParser` | DOM |
140
- | **Joyland** | `JoylandParser` | DOM |
141
- | **Chub** | `ChubParser` | DOM |
142
- | **Generic Web Article** | `ArticleParser` | Readability + Defuddle + Article-Extractor |
120
+ 19 AI chat platform parsers plus generic web article extraction:
121
+
122
+ | Platform | Parser | Extraction strategy |
123
+ | :---------------------------------- | :------------------------ | :---------------------------------------------------- |
124
+ | **ChatGPT** | `ChatGPTParser` | DOM + internal API |
125
+ | **Claude** | `ClaudeParser` | DOM + internal API + React fiber |
126
+ | **Google Gemini** | `GeminiParser` | DOM + batchexecute RPC |
127
+ | **Microsoft Copilot** | `CopilotParser` | DOM (multi-domain) |
128
+ | **Perplexity** | `PerplexityParser` | Internal API + DOM fallback |
129
+ | **DeepSeek** | `DeepSeekParser` | DOM + internal API (`fragments[]`) |
130
+ | **Qwen** | `QwenParser` | DOM |
131
+ | **Meta AI** | `MetaParser` | Internal API (GraphQL) + DOM fallback |
132
+ | **Mistral / Le Chat** | `MistralParser` | DOM |
133
+ | **Proton Lumo** | `LumoParser` | DOM (API is E2E-encrypted, not readable) |
134
+ | **Z.ai** | `ZAiParser` | Internal API (chat + batch) + DOM fallback |
135
+ | **Grok** | `GrokParser` | Internal API (response-node + load) + DOM |
136
+ | **Google AI Studio** | `GoogleAIStudioParser` | DOM |
137
+ | **NotebookLM** | `NotebookLMParser` | DOM |
138
+ | **Google Search AI (AI Overviews)** | `GoogleSearchAIParser` | DOM |
139
+ | **Gemini Cloud Assist** | `GeminiCloudAssistParser` | DOM |
140
+ | **Joyland** | `JoylandParser` | DOM |
141
+ | **Chub** | `ChubParser` | DOM |
142
+ | **Duck.ai (DuckDuckGo AI)** | `DuckAIParser` | DOM (client-side privacy, no server chat history API) |
143
+ | **Generic Web Article** | `ArticleParser` | Readability + Defuddle + Article-Extractor |
143
144
 
144
145
  All parsers extend the base [`ChatParser`](ai/base.js) interface — a consistent `isAvailable(url)` +
145
146
  normalized `parse()` contract. For the full extraction-strategy breakdown and maintenance model, see
146
147
  [SUPPORTED_PLATFORMS.md](SUPPORTED_PLATFORMS.md), also published as the
147
- [platform matrix](https://covai-labs.github.io/decant-core/platforms/) on the developer docs site.
148
+ [platform matrix](https://covai-labs.github.io/decant/platforms/) on the developer docs site.
148
149
 
149
150
  ---
150
151
 
@@ -163,7 +164,7 @@ Sample test fixtures located in [`tests/fixtures/`](tests/fixtures/) consist of
163
164
  ## Used by
164
165
 
165
166
  - [AI Chat Exporter](https://github.com/Covai-Labs/ai-chat-exporter) — export, archive and transfer AI conversations between platforms.
166
- - [Decant](https://github.com/Covai-Labs/decant) — the distraction-free web clipper and research batcher.
167
+ - [Decant](https://github.com/Covai-Labs/decant-browser-extension) — the distraction-free web clipper and research batcher.
167
168
 
168
169
  These products are demonstrations of the library, not its purpose. Yours can be next — see [CONTRIBUTING.md](CONTRIBUTING.md).
169
170
 
package/ai/duck_ai.js ADDED
@@ -0,0 +1,209 @@
1
+ import { ChatParser } from "./base.js";
2
+ import { convertToMarkdown } from "../utils/html-to-markdown.js";
3
+
4
+ export function isDuckAiUrl(url) {
5
+ if (!url || typeof url !== "string") return false;
6
+ try {
7
+ const parsed = new URL(url);
8
+ const host = parsed.hostname.toLowerCase();
9
+ if (host === "duck.ai" || host.endsWith(".duck.ai")) {
10
+ return true;
11
+ }
12
+ if (host === "duckduckgo.com" || host.endsWith(".duckduckgo.com")) {
13
+ if (parsed.pathname === "/chat" || parsed.pathname.startsWith("/chat/")) {
14
+ return true;
15
+ }
16
+ return parsed.searchParams.get("ia") === "chat";
17
+ }
18
+ } catch {
19
+ return false;
20
+ }
21
+ return false;
22
+ }
23
+
24
+ export class DuckAIParser extends ChatParser {
25
+ name = "Duck.ai";
26
+
27
+ isAvailable(url) {
28
+ return isDuckAiUrl(url);
29
+ }
30
+
31
+ async parse() {
32
+ // 1. Extract Title
33
+ let title = "";
34
+ if (typeof document !== "undefined" && document.title) {
35
+ title = document.title
36
+ .replace(/\s*-\s*DuckDuckGo.*$/i, "")
37
+ .replace(/\s*-\s*Duck\.ai.*$/i, "")
38
+ .trim();
39
+ if (/^(?:duckduckgo\s*ai\s*chat|duck\.ai)$/i.test(title)) {
40
+ title = "";
41
+ }
42
+ }
43
+
44
+ if (!title && typeof document !== "undefined") {
45
+ const activeChat = document.querySelector(
46
+ '[data-testid="ChatsList"] [aria-current="page"], [data-testid="ChatsList"] button, [data-testid="ChatsList"] a',
47
+ );
48
+ if (activeChat && activeChat.textContent) {
49
+ title = activeChat.textContent.trim();
50
+ }
51
+ }
52
+
53
+ // 2. Extract Messages
54
+ const messages = [];
55
+ let latestModel = "";
56
+
57
+ const userElements =
58
+ typeof document !== "undefined"
59
+ ? Array.from(document.querySelectorAll('[data-testid="user-message"]'))
60
+ : [];
61
+
62
+ for (let i = 0; i < userElements.length; i++) {
63
+ const userEl = userElements[i];
64
+
65
+ // Process User Message
66
+ const userClone = userEl.cloneNode(true);
67
+ // Remove inline favicon images
68
+ userClone
69
+ .querySelectorAll(
70
+ 'img[src*="duckduckgo.com/ip3/"], img[src*="favicons"]',
71
+ )
72
+ .forEach((img) => img.remove());
73
+ const userText = convertToMarkdown(userClone);
74
+ if (userText.trim()) {
75
+ messages.push({
76
+ role: "User",
77
+ content: userText.trim(),
78
+ });
79
+ }
80
+
81
+ // Process Assistant Message
82
+ let assistantEl = userEl.nextElementSibling;
83
+ if (
84
+ !assistantEl ||
85
+ assistantEl.getAttribute("data-testid") === "user-message"
86
+ ) {
87
+ // Fallback: look within parent turn container for assistant message element
88
+ assistantEl =
89
+ userEl.parentElement?.querySelector('[id*="-assistant-message-"]') ||
90
+ null;
91
+ }
92
+
93
+ if (assistantEl) {
94
+ // Extract Model Badge if present
95
+ let turnModel = "";
96
+ const heading = assistantEl.querySelector(
97
+ '[id^="heading-"][id*="-assistant-message-"], [id*="-assistant-message-"] h2, [id*="-assistant-message-"] h3',
98
+ );
99
+ if (heading && heading.textContent) {
100
+ turnModel = heading.textContent.replace(/^[^a-zA-Z0-9]+/, "").trim();
101
+ if (turnModel) {
102
+ latestModel = turnModel;
103
+ }
104
+ }
105
+
106
+ const astClone = assistantEl.cloneNode(true);
107
+
108
+ // Remove header / heading
109
+ astClone
110
+ .querySelectorAll(
111
+ '[id^="heading-"], .OFnY5LgQty8A4PIj3XdY, button, [data-testid="feedback-prompt"]',
112
+ )
113
+ .forEach((el) => el.remove());
114
+
115
+ // Remove transient web search query box
116
+ astClone
117
+ .querySelectorAll('.f_6cBYM9KdwpDEKf3fGW, [class*="search"]')
118
+ .forEach((el) => el.remove());
119
+
120
+ // Remove actions / footer container
121
+ astClone
122
+ .querySelectorAll("._4aCfUrBe8vy05lxcMBX")
123
+ .forEach((el) => el.remove());
124
+
125
+ // Preprocess Streamdown code blocks into standard <pre><code class="language-xyz">...</code></pre>
126
+ astClone
127
+ .querySelectorAll('[data-streamdown="code-block"]')
128
+ .forEach((cb) => {
129
+ const header = cb.querySelector(
130
+ '[data-streamdown="code-block-header"]',
131
+ );
132
+ const lang =
133
+ header?.getAttribute("data-language") ||
134
+ header?.textContent?.trim().toLowerCase() ||
135
+ "";
136
+ const codeEl =
137
+ cb.querySelector('[data-streamdown="code-block-body"] code') ||
138
+ cb.querySelector("code");
139
+ const codeText = codeEl?.textContent || "";
140
+
141
+ const pre = astClone.ownerDocument.createElement("pre");
142
+ const code = astClone.ownerDocument.createElement("code");
143
+ if (lang) {
144
+ code.className = `language-${lang}`;
145
+ }
146
+ code.textContent = codeText;
147
+ pre.appendChild(code);
148
+ if (cb.parentNode) {
149
+ cb.parentNode.replaceChild(pre, cb);
150
+ }
151
+ });
152
+
153
+ // Clean up favicon images
154
+ astClone
155
+ .querySelectorAll(
156
+ 'img[src*="duckduckgo.com/ip3/"], img[src*="favicons"]',
157
+ )
158
+ .forEach((img) => img.remove());
159
+
160
+ // Format Read More citation links: ensure space before domain name in citation text
161
+ astClone
162
+ .querySelectorAll(
163
+ "section h4 + div li a span > span, section ul li a span > span",
164
+ )
165
+ .forEach((domainSpan) => {
166
+ const domainText = domainSpan.textContent.trim();
167
+ if (domainText) {
168
+ domainSpan.textContent = ` (${domainText})`;
169
+ }
170
+ });
171
+
172
+ const astText = convertToMarkdown(astClone);
173
+ if (astText.trim()) {
174
+ const msg = {
175
+ role: "Duck.ai",
176
+ content: astText.trim(),
177
+ };
178
+ if (turnModel) {
179
+ msg.model = turnModel;
180
+ }
181
+ messages.push(msg);
182
+ }
183
+ }
184
+ }
185
+
186
+ if (!title && messages.length > 0) {
187
+ const firstUser = messages.find((m) => m.role === "User");
188
+ if (firstUser && firstUser.content) {
189
+ title = firstUser.content.slice(0, 60).trim();
190
+ }
191
+ }
192
+ title = title || "Duck.ai Conversation";
193
+
194
+ const currentUrl =
195
+ typeof window !== "undefined" && window.location
196
+ ? window.location.href || ""
197
+ : "";
198
+
199
+ const metadata = {
200
+ Source: "Duck.ai",
201
+ Date: new Date().toLocaleString(),
202
+ Link: currentUrl,
203
+ Method: "DOM",
204
+ ...(latestModel ? { Model: latestModel } : {}),
205
+ };
206
+
207
+ return { title, messages, url: currentUrl, metadata };
208
+ }
209
+ }
package/ai/index.js CHANGED
@@ -28,6 +28,7 @@ export { GeminiCloudAssistParser } from "./gemini_cloud_assist.js";
28
28
  export { JoylandParser } from "./joyland.js";
29
29
  export { ChubParser } from "./chub.js";
30
30
  export { GrokParser } from "./grok.js";
31
+ export { DuckAIParser } from "./duck_ai.js";
31
32
 
32
33
  // Utilities
33
34
  export { convertToMarkdown, cleanMarkdown } from "../utils/html-to-markdown.js";
@@ -17,6 +17,7 @@ import { GeminiCloudAssistParser } from "../ai/gemini_cloud_assist.js";
17
17
  import { JoylandParser } from "../ai/joyland.js";
18
18
  import { ChubParser } from "../ai/chub.js";
19
19
  import { GrokParser } from "../ai/grok.js";
20
+ import { DuckAIParser } from "../ai/duck_ai.js";
20
21
 
21
22
  /**
22
23
  * Ordered list of parsers. First match wins.
@@ -32,6 +33,7 @@ export const parsers = [
32
33
  new QwenParser(),
33
34
  new MetaParser(),
34
35
  new MistralParser(),
36
+ new DuckAIParser(),
35
37
  new LumoParser(),
36
38
  new ZAiParser(),
37
39
  new GoogleAIStudioParser(),
@@ -1,3 +1,5 @@
1
+ import { isDuckAiUrl } from "../ai/duck_ai.js";
2
+
1
3
  /**
2
4
  * AI chat platform domain definitions.
3
5
  * Centralised so detection logic and UI (e.g. Decant's tip banner) share the same list.
@@ -23,6 +25,7 @@ export const AI_CHAT_DOMAINS = [
23
25
  "chub.ai",
24
26
  "characterhub.org",
25
27
  "grok.com",
28
+ "duck.ai",
26
29
  ];
27
30
 
28
31
  /**
@@ -51,4 +54,8 @@ export const URL_PATTERNS = [
51
54
  test: (url) => url.includes("console.cloud.google.com/gemini"),
52
55
  platform: "gemini-cloud-assist",
53
56
  },
57
+ {
58
+ test: isDuckAiUrl,
59
+ platform: "duck-ai",
60
+ },
54
61
  ];
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "decant-core",
3
- "version": "1.10.1",
3
+ "version": "1.11.0",
4
4
  "description": "A shared web extraction layer for AI conversations and regular web pages.",
5
5
  "type": "module",
6
6
  "engines": {
@@ -26,12 +26,12 @@
26
26
  ],
27
27
  "repository": {
28
28
  "type": "git",
29
- "url": "git+https://github.com/Covai-Labs/decant-core.git"
29
+ "url": "git+https://github.com/Covai-Labs/decant.git"
30
30
  },
31
31
  "bugs": {
32
- "url": "https://github.com/Covai-Labs/decant-core/issues"
32
+ "url": "https://github.com/Covai-Labs/decant/issues"
33
33
  },
34
- "homepage": "https://github.com/Covai-Labs/decant-core#readme",
34
+ "homepage": "https://github.com/Covai-Labs/decant#readme",
35
35
  "scripts": {
36
36
  "test": "node --test tests/*.test.js",
37
37
  "lint": "eslint .",