decant-core 1.10.1 → 1.12.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +14 -13
- package/ai/chatgpt.js +19 -6
- package/ai/claude.js +29 -3
- package/ai/deepseek.js +79 -4
- package/ai/duck_ai.js +209 -0
- package/ai/index.js +1 -0
- package/ai/z_ai.js +4 -5
- package/data/platforms.json +162 -0
- package/detection/detect-platform.js +2 -0
- package/detection/domains.js +7 -0
- package/package.json +10 -6
package/README.md
CHANGED
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
[](https://www.npmjs.com/package/decant-core)
|
|
6
6
|
[](LICENSE)
|
|
7
|
-
[](https://github.com/Covai-Labs/decant/stargazers)
|
|
8
8
|
|
|
9
9
|
Every AI chat exporter ends up solving the same problem: extracting conversations from ChatGPT, Claude, Gemini, Perplexity, DeepSeek and other constantly changing AI interfaces.
|
|
10
10
|
|
|
@@ -16,7 +16,7 @@ And every time one of those platforms changes its UI, seriously re-renders a mes
|
|
|
16
16
|
|
|
17
17
|
Existing chat exporters often suffer from two major flaws: they break whenever platform DOMs update, and many route user conversations through third-party servers.
|
|
18
18
|
|
|
19
|
-
`decant-core` was created to solve both at the foundation. Originally built to power local-first extensions like [AI Chat Exporter](https://
|
|
19
|
+
`decant-core` was created to solve both at the foundation. Originally built to power local-first extensions like [AI Chat Exporter](https://ace.covai.org/) and [Decant](https://decant.covai.org/), it decouples fragile platform parsing from presentation. By sharing this engine under AGPL-3.0, any browser extension, web clipper, archiver, or research tool can rely on a maintained, local-first extraction layer instead of reverse-engineering AI platforms in isolation.
|
|
20
20
|
|
|
21
21
|
```bash
|
|
22
22
|
npm install decant-core
|
|
@@ -41,7 +41,7 @@ AI platforms don't expose stable public APIs for reading conversation history. N
|
|
|
41
41
|
|
|
42
42
|
Maintaining that per-platform logic in every exporter is wasteful and fragile. `decant-core` centralizes it:
|
|
43
43
|
|
|
44
|
-
- ✅ **
|
|
44
|
+
- ✅ **19 AI chat platform parsers** with normalized output — you get structured messages, models, metadata and Markdown, not DOM soup.
|
|
45
45
|
- ✅ **Web article extraction** — Mozilla Readability, Defuddle, and Article-Extractor run in parallel and arbitrate by content-quality scoring.
|
|
46
46
|
- ✅ **Detection utilities** — tell an "AI chat page" apart from a "regular web page" before you decide which parser to run.
|
|
47
47
|
- ✅ **Math & Markdown handling** — LaTeX normalization plus GFM tables/code fencing that survive round-trips into Obsidian, Logseq and Notion.
|
|
@@ -117,34 +117,35 @@ import { normalizeLatexMath } from "decant-core";
|
|
|
117
117
|
|
|
118
118
|
## Supported Platforms
|
|
119
119
|
|
|
120
|
-
|
|
120
|
+
19 AI chat platform parsers plus generic web article extraction:
|
|
121
121
|
|
|
122
122
|
| Platform | Parser | Extraction strategy |
|
|
123
123
|
| :---------------------------------- | :------------------------ | :----------------------------------------- |
|
|
124
124
|
| **ChatGPT** | `ChatGPTParser` | DOM + internal API |
|
|
125
125
|
| **Claude** | `ClaudeParser` | DOM + internal API + React fiber |
|
|
126
126
|
| **Google Gemini** | `GeminiParser` | DOM + batchexecute RPC |
|
|
127
|
-
| **Microsoft Copilot** | `CopilotParser` | DOM
|
|
128
|
-
| **Perplexity** | `PerplexityParser` | Internal API + DOM
|
|
129
|
-
| **DeepSeek** | `DeepSeekParser` | DOM + internal API
|
|
127
|
+
| **Microsoft Copilot** | `CopilotParser` | DOM |
|
|
128
|
+
| **Perplexity** | `PerplexityParser` | Internal API + DOM |
|
|
129
|
+
| **DeepSeek** | `DeepSeekParser` | DOM + internal API |
|
|
130
130
|
| **Qwen** | `QwenParser` | DOM |
|
|
131
|
-
| **Meta AI** | `MetaParser` | Internal API
|
|
131
|
+
| **Meta AI** | `MetaParser` | Internal API + DOM |
|
|
132
132
|
| **Mistral / Le Chat** | `MistralParser` | DOM |
|
|
133
|
-
| **Proton Lumo** | `LumoParser` | DOM
|
|
134
|
-
| **Z.ai** | `ZAiParser` | Internal API
|
|
135
|
-
| **Grok** | `GrokParser` | Internal API
|
|
133
|
+
| **Proton Lumo** | `LumoParser` | DOM |
|
|
134
|
+
| **Z.ai** | `ZAiParser` | Internal API + DOM |
|
|
135
|
+
| **Grok** | `GrokParser` | Internal API + DOM |
|
|
136
136
|
| **Google AI Studio** | `GoogleAIStudioParser` | DOM |
|
|
137
137
|
| **NotebookLM** | `NotebookLMParser` | DOM |
|
|
138
138
|
| **Google Search AI (AI Overviews)** | `GoogleSearchAIParser` | DOM |
|
|
139
139
|
| **Gemini Cloud Assist** | `GeminiCloudAssistParser` | DOM |
|
|
140
140
|
| **Joyland** | `JoylandParser` | DOM |
|
|
141
141
|
| **Chub** | `ChubParser` | DOM |
|
|
142
|
+
| **Duck.ai (DuckDuckGo AI)** | `DuckAIParser` | DOM |
|
|
142
143
|
| **Generic Web Article** | `ArticleParser` | Readability + Defuddle + Article-Extractor |
|
|
143
144
|
|
|
144
145
|
All parsers extend the base [`ChatParser`](ai/base.js) interface — a consistent `isAvailable(url)` +
|
|
145
146
|
normalized `parse()` contract. For the full extraction-strategy breakdown and maintenance model, see
|
|
146
147
|
[SUPPORTED_PLATFORMS.md](SUPPORTED_PLATFORMS.md), also published as the
|
|
147
|
-
[platform matrix](https://covai-labs.github.io/decant
|
|
148
|
+
[platform matrix](https://covai-labs.github.io/decant/platforms/) on the developer docs site.
|
|
148
149
|
|
|
149
150
|
---
|
|
150
151
|
|
|
@@ -163,7 +164,7 @@ Sample test fixtures located in [`tests/fixtures/`](tests/fixtures/) consist of
|
|
|
163
164
|
## Used by
|
|
164
165
|
|
|
165
166
|
- [AI Chat Exporter](https://github.com/Covai-Labs/ai-chat-exporter) — export, archive and transfer AI conversations between platforms.
|
|
166
|
-
- [Decant](https://github.com/Covai-Labs/decant) — the distraction-free web clipper and research batcher.
|
|
167
|
+
- [Decant](https://github.com/Covai-Labs/decant-browser-extension) — the distraction-free web clipper and research batcher.
|
|
167
168
|
|
|
168
169
|
These products are demonstrations of the library, not its purpose. Yours can be next — see [CONTRIBUTING.md](CONTRIBUTING.md).
|
|
169
170
|
|
package/ai/chatgpt.js
CHANGED
|
@@ -831,6 +831,7 @@ export class ChatGPTParser extends ChatParser {
|
|
|
831
831
|
const messages = [];
|
|
832
832
|
for (const msg of apiMessages) {
|
|
833
833
|
let content = "";
|
|
834
|
+
let thinking = "";
|
|
834
835
|
for (const seg of msg.segments) {
|
|
835
836
|
if (seg.type === "text") {
|
|
836
837
|
content +=
|
|
@@ -843,7 +844,7 @@ export class ChatGPTParser extends ChatParser {
|
|
|
843
844
|
msg.imageGroupMap,
|
|
844
845
|
);
|
|
845
846
|
if (thoughtText) {
|
|
846
|
-
|
|
847
|
+
thinking += (thinking ? "\n\n" : "") + thoughtText;
|
|
847
848
|
}
|
|
848
849
|
} else if (seg.type === "image") {
|
|
849
850
|
const src = images[seg.fileId];
|
|
@@ -852,12 +853,22 @@ export class ChatGPTParser extends ChatParser {
|
|
|
852
853
|
}
|
|
853
854
|
}
|
|
854
855
|
}
|
|
855
|
-
|
|
856
|
-
if (
|
|
856
|
+
let fullContent = "";
|
|
857
|
+
if (thinking) {
|
|
858
|
+
fullContent += `<think>\n${thinking}\n</think>\n\n`;
|
|
859
|
+
}
|
|
860
|
+
if (content.trim()) {
|
|
861
|
+
fullContent += content.trim();
|
|
862
|
+
}
|
|
863
|
+
fullContent = fullContent.trim();
|
|
864
|
+
if (fullContent) {
|
|
857
865
|
const msgObj = {
|
|
858
866
|
role: msg.role,
|
|
859
|
-
content:
|
|
867
|
+
content: fullContent,
|
|
860
868
|
};
|
|
869
|
+
if (thinking) {
|
|
870
|
+
msgObj.thinking = thinking;
|
|
871
|
+
}
|
|
861
872
|
if (msg.timestamp) {
|
|
862
873
|
msgObj.timestamp = msg.timestamp;
|
|
863
874
|
}
|
|
@@ -876,8 +887,10 @@ export class ChatGPTParser extends ChatParser {
|
|
|
876
887
|
Link: currentUrl,
|
|
877
888
|
Model:
|
|
878
889
|
convoData?.model_slug ||
|
|
879
|
-
document.querySelector
|
|
880
|
-
|
|
890
|
+
(typeof document !== "undefined" && document.querySelector
|
|
891
|
+
? document.querySelector('[data-testid="model-selector-dropdown"]')
|
|
892
|
+
?.innerText
|
|
893
|
+
: null) ||
|
|
881
894
|
"ChatGPT",
|
|
882
895
|
Method: method,
|
|
883
896
|
};
|
package/ai/claude.js
CHANGED
|
@@ -402,12 +402,30 @@ export class ClaudeParser extends ChatParser {
|
|
|
402
402
|
const role = message.sender === "human" ? "User" : "Claude";
|
|
403
403
|
|
|
404
404
|
let contentStr = "";
|
|
405
|
+
let thinkingStr = "";
|
|
405
406
|
|
|
406
407
|
// Construct content
|
|
407
408
|
if (message.content && Array.isArray(message.content)) {
|
|
408
409
|
for (const block of message.content) {
|
|
409
|
-
if (block.type === "thinking"
|
|
410
|
-
|
|
410
|
+
if (block.type === "thinking") {
|
|
411
|
+
let thoughtText = "";
|
|
412
|
+
if (
|
|
413
|
+
typeof block.thinking === "string" &&
|
|
414
|
+
block.thinking.trim()
|
|
415
|
+
) {
|
|
416
|
+
thoughtText = block.thinking.trim();
|
|
417
|
+
} else if (Array.isArray(block.summaries)) {
|
|
418
|
+
thoughtText = block.summaries
|
|
419
|
+
.map((s) =>
|
|
420
|
+
typeof s === "string" ? s : s?.summary || "",
|
|
421
|
+
)
|
|
422
|
+
.map((s) => s.trim())
|
|
423
|
+
.filter(Boolean)
|
|
424
|
+
.join("\n");
|
|
425
|
+
}
|
|
426
|
+
if (thoughtText) {
|
|
427
|
+
thinkingStr += (thinkingStr ? "\n\n" : "") + thoughtText;
|
|
428
|
+
}
|
|
411
429
|
} else if (block.type === "text" && block.text) {
|
|
412
430
|
const cleanText = block.text
|
|
413
431
|
.replace(/<antArtifact[^>]*>[\s\S]*?<\/antArtifact>/g, "")
|
|
@@ -516,9 +534,17 @@ export class ClaudeParser extends ChatParser {
|
|
|
516
534
|
}
|
|
517
535
|
}
|
|
518
536
|
|
|
537
|
+
if (thinkingStr) {
|
|
538
|
+
contentStr = `<think>\n${thinkingStr}\n</think>\n\n` + contentStr;
|
|
539
|
+
}
|
|
540
|
+
|
|
519
541
|
contentStr = contentStr.trim();
|
|
520
542
|
if (contentStr) {
|
|
521
|
-
|
|
543
|
+
const msgObj = { role, content: contentStr };
|
|
544
|
+
if (thinkingStr) {
|
|
545
|
+
msgObj.thinking = thinkingStr;
|
|
546
|
+
}
|
|
547
|
+
messages.push(msgObj);
|
|
522
548
|
}
|
|
523
549
|
|
|
524
550
|
// Extract and push artifacts
|
package/ai/deepseek.js
CHANGED
|
@@ -98,7 +98,29 @@ async function fetchDeepSeekConversation(sessionId, token) {
|
|
|
98
98
|
const isUser = msgNode.role === "USER" || msgNode.role === "user";
|
|
99
99
|
const role = isUser ? "User" : "DeepSeek";
|
|
100
100
|
const content = extractDeepSeekMessageContent(msgNode);
|
|
101
|
-
|
|
101
|
+
let thinking = "";
|
|
102
|
+
if (!isUser && Array.isArray(msgNode.fragments)) {
|
|
103
|
+
const thinkFragments = msgNode.fragments.filter(
|
|
104
|
+
(f) => f && f.type === "THINK" && typeof f.content === "string",
|
|
105
|
+
);
|
|
106
|
+
thinking = thinkFragments
|
|
107
|
+
.map((f) => f.content.trim())
|
|
108
|
+
.filter(Boolean)
|
|
109
|
+
.join("\n\n");
|
|
110
|
+
}
|
|
111
|
+
let fullContent = "";
|
|
112
|
+
if (thinking) {
|
|
113
|
+
fullContent += `<think>\n${thinking}\n</think>\n\n`;
|
|
114
|
+
}
|
|
115
|
+
if (content.trim()) {
|
|
116
|
+
fullContent += content.trim();
|
|
117
|
+
}
|
|
118
|
+
fullContent = fullContent.trim();
|
|
119
|
+
const msg = { role, content: fullContent };
|
|
120
|
+
if (thinking) {
|
|
121
|
+
msg.thinking = thinking;
|
|
122
|
+
}
|
|
123
|
+
return msg;
|
|
102
124
|
})
|
|
103
125
|
.filter((msg) => msg.content.length > 0);
|
|
104
126
|
}
|
|
@@ -182,15 +204,40 @@ export class DeepSeekParser extends ChatParser {
|
|
|
182
204
|
|
|
183
205
|
outerElements.forEach((el) => {
|
|
184
206
|
let role = "Unknown";
|
|
207
|
+
let thinking = "";
|
|
185
208
|
if (el.matches(userSelector)) {
|
|
186
209
|
role = "User";
|
|
187
210
|
} else if (el.matches(assistantSelector)) {
|
|
188
211
|
role = "DeepSeek";
|
|
212
|
+
const messageContainer =
|
|
213
|
+
(el.closest && el.closest(".ds-message")) || el.parentElement;
|
|
214
|
+
if (messageContainer) {
|
|
215
|
+
const thinkContainers =
|
|
216
|
+
messageContainer.querySelectorAll(".ds-think-content");
|
|
217
|
+
if (thinkContainers.length > 0) {
|
|
218
|
+
thinking = Array.from(thinkContainers)
|
|
219
|
+
.map((tc) => convertToMarkdown(tc).trim())
|
|
220
|
+
.filter(Boolean)
|
|
221
|
+
.join("\n\n");
|
|
222
|
+
}
|
|
223
|
+
}
|
|
189
224
|
}
|
|
190
225
|
|
|
191
226
|
const text = convertToMarkdown(el);
|
|
227
|
+
let fullContent = "";
|
|
228
|
+
if (thinking) {
|
|
229
|
+
fullContent += `<think>\n${thinking}\n</think>\n\n`;
|
|
230
|
+
}
|
|
192
231
|
if (text.trim()) {
|
|
193
|
-
|
|
232
|
+
fullContent += text.trim();
|
|
233
|
+
}
|
|
234
|
+
fullContent = fullContent.trim();
|
|
235
|
+
if (fullContent) {
|
|
236
|
+
const msg = { role, content: fullContent };
|
|
237
|
+
if (thinking) {
|
|
238
|
+
msg.thinking = thinking;
|
|
239
|
+
}
|
|
240
|
+
messages.push(msg);
|
|
194
241
|
}
|
|
195
242
|
});
|
|
196
243
|
|
|
@@ -202,9 +249,37 @@ export class DeepSeekParser extends ChatParser {
|
|
|
202
249
|
messageRows.forEach((row) => {
|
|
203
250
|
const isUser = row.classList.contains("ds-user-message");
|
|
204
251
|
const role = isUser ? "User" : "DeepSeek";
|
|
205
|
-
const
|
|
252
|
+
const rowClone = row.cloneNode(true);
|
|
253
|
+
let thinking = "";
|
|
254
|
+
if (!isUser) {
|
|
255
|
+
const thinkContainers =
|
|
256
|
+
rowClone.querySelectorAll(".ds-think-content");
|
|
257
|
+
if (thinkContainers.length > 0) {
|
|
258
|
+
thinking = Array.from(thinkContainers)
|
|
259
|
+
.map((tc) => {
|
|
260
|
+
const md = convertToMarkdown(tc).trim();
|
|
261
|
+
tc.remove();
|
|
262
|
+
return md;
|
|
263
|
+
})
|
|
264
|
+
.filter(Boolean)
|
|
265
|
+
.join("\n\n");
|
|
266
|
+
}
|
|
267
|
+
}
|
|
268
|
+
const text = convertToMarkdown(rowClone);
|
|
269
|
+
let fullContent = "";
|
|
270
|
+
if (thinking) {
|
|
271
|
+
fullContent += `<think>\n${thinking}\n</think>\n\n`;
|
|
272
|
+
}
|
|
206
273
|
if (text.trim()) {
|
|
207
|
-
|
|
274
|
+
fullContent += text.trim();
|
|
275
|
+
}
|
|
276
|
+
fullContent = fullContent.trim();
|
|
277
|
+
if (fullContent) {
|
|
278
|
+
const msg = { role, content: fullContent };
|
|
279
|
+
if (thinking) {
|
|
280
|
+
msg.thinking = thinking;
|
|
281
|
+
}
|
|
282
|
+
messages.push(msg);
|
|
208
283
|
}
|
|
209
284
|
});
|
|
210
285
|
}
|
package/ai/duck_ai.js
ADDED
|
@@ -0,0 +1,209 @@
|
|
|
1
|
+
import { ChatParser } from "./base.js";
|
|
2
|
+
import { convertToMarkdown } from "../utils/html-to-markdown.js";
|
|
3
|
+
|
|
4
|
+
export function isDuckAiUrl(url) {
|
|
5
|
+
if (!url || typeof url !== "string") return false;
|
|
6
|
+
try {
|
|
7
|
+
const parsed = new URL(url);
|
|
8
|
+
const host = parsed.hostname.toLowerCase();
|
|
9
|
+
if (host === "duck.ai" || host.endsWith(".duck.ai")) {
|
|
10
|
+
return true;
|
|
11
|
+
}
|
|
12
|
+
if (host === "duckduckgo.com" || host.endsWith(".duckduckgo.com")) {
|
|
13
|
+
if (parsed.pathname === "/chat" || parsed.pathname.startsWith("/chat/")) {
|
|
14
|
+
return true;
|
|
15
|
+
}
|
|
16
|
+
return parsed.searchParams.get("ia") === "chat";
|
|
17
|
+
}
|
|
18
|
+
} catch {
|
|
19
|
+
return false;
|
|
20
|
+
}
|
|
21
|
+
return false;
|
|
22
|
+
}
|
|
23
|
+
|
|
24
|
+
export class DuckAIParser extends ChatParser {
|
|
25
|
+
name = "Duck.ai";
|
|
26
|
+
|
|
27
|
+
isAvailable(url) {
|
|
28
|
+
return isDuckAiUrl(url);
|
|
29
|
+
}
|
|
30
|
+
|
|
31
|
+
async parse() {
|
|
32
|
+
// 1. Extract Title
|
|
33
|
+
let title = "";
|
|
34
|
+
if (typeof document !== "undefined" && document.title) {
|
|
35
|
+
title = document.title
|
|
36
|
+
.replace(/\s*-\s*DuckDuckGo.*$/i, "")
|
|
37
|
+
.replace(/\s*-\s*Duck\.ai.*$/i, "")
|
|
38
|
+
.trim();
|
|
39
|
+
if (/^(?:duckduckgo\s*ai\s*chat|duck\.ai)$/i.test(title)) {
|
|
40
|
+
title = "";
|
|
41
|
+
}
|
|
42
|
+
}
|
|
43
|
+
|
|
44
|
+
if (!title && typeof document !== "undefined") {
|
|
45
|
+
const activeChat = document.querySelector(
|
|
46
|
+
'[data-testid="ChatsList"] [aria-current="page"], [data-testid="ChatsList"] button, [data-testid="ChatsList"] a',
|
|
47
|
+
);
|
|
48
|
+
if (activeChat && activeChat.textContent) {
|
|
49
|
+
title = activeChat.textContent.trim();
|
|
50
|
+
}
|
|
51
|
+
}
|
|
52
|
+
|
|
53
|
+
// 2. Extract Messages
|
|
54
|
+
const messages = [];
|
|
55
|
+
let latestModel = "";
|
|
56
|
+
|
|
57
|
+
const userElements =
|
|
58
|
+
typeof document !== "undefined"
|
|
59
|
+
? Array.from(document.querySelectorAll('[data-testid="user-message"]'))
|
|
60
|
+
: [];
|
|
61
|
+
|
|
62
|
+
for (let i = 0; i < userElements.length; i++) {
|
|
63
|
+
const userEl = userElements[i];
|
|
64
|
+
|
|
65
|
+
// Process User Message
|
|
66
|
+
const userClone = userEl.cloneNode(true);
|
|
67
|
+
// Remove inline favicon images
|
|
68
|
+
userClone
|
|
69
|
+
.querySelectorAll(
|
|
70
|
+
'img[src*="duckduckgo.com/ip3/"], img[src*="favicons"]',
|
|
71
|
+
)
|
|
72
|
+
.forEach((img) => img.remove());
|
|
73
|
+
const userText = convertToMarkdown(userClone);
|
|
74
|
+
if (userText.trim()) {
|
|
75
|
+
messages.push({
|
|
76
|
+
role: "User",
|
|
77
|
+
content: userText.trim(),
|
|
78
|
+
});
|
|
79
|
+
}
|
|
80
|
+
|
|
81
|
+
// Process Assistant Message
|
|
82
|
+
let assistantEl = userEl.nextElementSibling;
|
|
83
|
+
if (
|
|
84
|
+
!assistantEl ||
|
|
85
|
+
assistantEl.getAttribute("data-testid") === "user-message"
|
|
86
|
+
) {
|
|
87
|
+
// Fallback: look within parent turn container for assistant message element
|
|
88
|
+
assistantEl =
|
|
89
|
+
userEl.parentElement?.querySelector('[id*="-assistant-message-"]') ||
|
|
90
|
+
null;
|
|
91
|
+
}
|
|
92
|
+
|
|
93
|
+
if (assistantEl) {
|
|
94
|
+
// Extract Model Badge if present
|
|
95
|
+
let turnModel = "";
|
|
96
|
+
const heading = assistantEl.querySelector(
|
|
97
|
+
'[id^="heading-"][id*="-assistant-message-"], [id*="-assistant-message-"] h2, [id*="-assistant-message-"] h3',
|
|
98
|
+
);
|
|
99
|
+
if (heading && heading.textContent) {
|
|
100
|
+
turnModel = heading.textContent.replace(/^[^a-zA-Z0-9]+/, "").trim();
|
|
101
|
+
if (turnModel) {
|
|
102
|
+
latestModel = turnModel;
|
|
103
|
+
}
|
|
104
|
+
}
|
|
105
|
+
|
|
106
|
+
const astClone = assistantEl.cloneNode(true);
|
|
107
|
+
|
|
108
|
+
// Remove header / heading
|
|
109
|
+
astClone
|
|
110
|
+
.querySelectorAll(
|
|
111
|
+
'[id^="heading-"], .OFnY5LgQty8A4PIj3XdY, button, [data-testid="feedback-prompt"]',
|
|
112
|
+
)
|
|
113
|
+
.forEach((el) => el.remove());
|
|
114
|
+
|
|
115
|
+
// Remove transient web search query box
|
|
116
|
+
astClone
|
|
117
|
+
.querySelectorAll('.f_6cBYM9KdwpDEKf3fGW, [class*="search"]')
|
|
118
|
+
.forEach((el) => el.remove());
|
|
119
|
+
|
|
120
|
+
// Remove actions / footer container
|
|
121
|
+
astClone
|
|
122
|
+
.querySelectorAll("._4aCfUrBe8vy05lxcMBX")
|
|
123
|
+
.forEach((el) => el.remove());
|
|
124
|
+
|
|
125
|
+
// Preprocess Streamdown code blocks into standard <pre><code class="language-xyz">...</code></pre>
|
|
126
|
+
astClone
|
|
127
|
+
.querySelectorAll('[data-streamdown="code-block"]')
|
|
128
|
+
.forEach((cb) => {
|
|
129
|
+
const header = cb.querySelector(
|
|
130
|
+
'[data-streamdown="code-block-header"]',
|
|
131
|
+
);
|
|
132
|
+
const lang =
|
|
133
|
+
header?.getAttribute("data-language") ||
|
|
134
|
+
header?.textContent?.trim().toLowerCase() ||
|
|
135
|
+
"";
|
|
136
|
+
const codeEl =
|
|
137
|
+
cb.querySelector('[data-streamdown="code-block-body"] code') ||
|
|
138
|
+
cb.querySelector("code");
|
|
139
|
+
const codeText = codeEl?.textContent || "";
|
|
140
|
+
|
|
141
|
+
const pre = astClone.ownerDocument.createElement("pre");
|
|
142
|
+
const code = astClone.ownerDocument.createElement("code");
|
|
143
|
+
if (lang) {
|
|
144
|
+
code.className = `language-${lang}`;
|
|
145
|
+
}
|
|
146
|
+
code.textContent = codeText;
|
|
147
|
+
pre.appendChild(code);
|
|
148
|
+
if (cb.parentNode) {
|
|
149
|
+
cb.parentNode.replaceChild(pre, cb);
|
|
150
|
+
}
|
|
151
|
+
});
|
|
152
|
+
|
|
153
|
+
// Clean up favicon images
|
|
154
|
+
astClone
|
|
155
|
+
.querySelectorAll(
|
|
156
|
+
'img[src*="duckduckgo.com/ip3/"], img[src*="favicons"]',
|
|
157
|
+
)
|
|
158
|
+
.forEach((img) => img.remove());
|
|
159
|
+
|
|
160
|
+
// Format Read More citation links: ensure space before domain name in citation text
|
|
161
|
+
astClone
|
|
162
|
+
.querySelectorAll(
|
|
163
|
+
"section h4 + div li a span > span, section ul li a span > span",
|
|
164
|
+
)
|
|
165
|
+
.forEach((domainSpan) => {
|
|
166
|
+
const domainText = domainSpan.textContent.trim();
|
|
167
|
+
if (domainText) {
|
|
168
|
+
domainSpan.textContent = ` (${domainText})`;
|
|
169
|
+
}
|
|
170
|
+
});
|
|
171
|
+
|
|
172
|
+
const astText = convertToMarkdown(astClone);
|
|
173
|
+
if (astText.trim()) {
|
|
174
|
+
const msg = {
|
|
175
|
+
role: "Duck.ai",
|
|
176
|
+
content: astText.trim(),
|
|
177
|
+
};
|
|
178
|
+
if (turnModel) {
|
|
179
|
+
msg.model = turnModel;
|
|
180
|
+
}
|
|
181
|
+
messages.push(msg);
|
|
182
|
+
}
|
|
183
|
+
}
|
|
184
|
+
}
|
|
185
|
+
|
|
186
|
+
if (!title && messages.length > 0) {
|
|
187
|
+
const firstUser = messages.find((m) => m.role === "User");
|
|
188
|
+
if (firstUser && firstUser.content) {
|
|
189
|
+
title = firstUser.content.slice(0, 60).trim();
|
|
190
|
+
}
|
|
191
|
+
}
|
|
192
|
+
title = title || "Duck.ai Conversation";
|
|
193
|
+
|
|
194
|
+
const currentUrl =
|
|
195
|
+
typeof window !== "undefined" && window.location
|
|
196
|
+
? window.location.href || ""
|
|
197
|
+
: "";
|
|
198
|
+
|
|
199
|
+
const metadata = {
|
|
200
|
+
Source: "Duck.ai",
|
|
201
|
+
Date: new Date().toLocaleString(),
|
|
202
|
+
Link: currentUrl,
|
|
203
|
+
Method: "DOM",
|
|
204
|
+
...(latestModel ? { Model: latestModel } : {}),
|
|
205
|
+
};
|
|
206
|
+
|
|
207
|
+
return { title, messages, url: currentUrl, metadata };
|
|
208
|
+
}
|
|
209
|
+
}
|
package/ai/index.js
CHANGED
|
@@ -28,6 +28,7 @@ export { GeminiCloudAssistParser } from "./gemini_cloud_assist.js";
|
|
|
28
28
|
export { JoylandParser } from "./joyland.js";
|
|
29
29
|
export { ChubParser } from "./chub.js";
|
|
30
30
|
export { GrokParser } from "./grok.js";
|
|
31
|
+
export { DuckAIParser } from "./duck_ai.js";
|
|
31
32
|
|
|
32
33
|
// Utilities
|
|
33
34
|
export { convertToMarkdown, cleanMarkdown } from "../utils/html-to-markdown.js";
|
package/ai/z_ai.js
CHANGED
|
@@ -60,13 +60,12 @@ export function formatZaiMessage(entry) {
|
|
|
60
60
|
.filter(Boolean)
|
|
61
61
|
.join("\n\n");
|
|
62
62
|
if (reasoning) {
|
|
63
|
-
|
|
64
|
-
.split("\n")
|
|
65
|
-
.map((line) => `> ${line}`)
|
|
66
|
-
.join("\n");
|
|
67
|
-
content += `\n\n> 🧠 Thinking\n${quoted}`;
|
|
63
|
+
content = `<think>\n${reasoning}\n</think>\n\n${content}`;
|
|
68
64
|
}
|
|
69
65
|
const msg = { role, content };
|
|
66
|
+
if (reasoning) {
|
|
67
|
+
msg.thinking = reasoning;
|
|
68
|
+
}
|
|
70
69
|
if (entry.timestamp) {
|
|
71
70
|
try {
|
|
72
71
|
msg.timestamp = new Date(entry.timestamp * 1000).toISOString();
|
|
@@ -0,0 +1,162 @@
|
|
|
1
|
+
[
|
|
2
|
+
{
|
|
3
|
+
"id": "chatgpt",
|
|
4
|
+
"platform": "ChatGPT",
|
|
5
|
+
"parser": "ChatGPTParser",
|
|
6
|
+
"module": "decant-core/ai/chatgpt",
|
|
7
|
+
"strategy": "DOM + internal API",
|
|
8
|
+
"notes": "Fallback and RPC-assisted extraction; scroll/dedup helper for long threads."
|
|
9
|
+
},
|
|
10
|
+
{
|
|
11
|
+
"id": "claude",
|
|
12
|
+
"platform": "Claude",
|
|
13
|
+
"parser": "ClaudeParser",
|
|
14
|
+
"module": "decant-core/ai/claude",
|
|
15
|
+
"strategy": "DOM + internal API + React fiber",
|
|
16
|
+
"notes": "Internal API first with DOM fallback; reads the React tree for artifacts and structured blocks."
|
|
17
|
+
},
|
|
18
|
+
{
|
|
19
|
+
"id": "gemini",
|
|
20
|
+
"platform": "Google Gemini",
|
|
21
|
+
"parser": "GeminiParser",
|
|
22
|
+
"module": "decant-core/ai/gemini",
|
|
23
|
+
"strategy": "DOM + batchexecute RPC",
|
|
24
|
+
"notes": "Uses batchexecute RPC pagination with resilient DOM fallback."
|
|
25
|
+
},
|
|
26
|
+
{
|
|
27
|
+
"id": "copilot",
|
|
28
|
+
"platform": "Microsoft Copilot",
|
|
29
|
+
"parser": "CopilotParser",
|
|
30
|
+
"module": "decant-core/ai/copilot",
|
|
31
|
+
"strategy": "DOM",
|
|
32
|
+
"notes": "Multi-domain (bing + copilot) DOM extraction."
|
|
33
|
+
},
|
|
34
|
+
{
|
|
35
|
+
"id": "perplexity",
|
|
36
|
+
"platform": "Perplexity",
|
|
37
|
+
"parser": "PerplexityParser",
|
|
38
|
+
"module": "decant-core/ai/perplexity",
|
|
39
|
+
"strategy": "Internal API + DOM",
|
|
40
|
+
"notes": "Internal API first, DOM fallback for source citations and answers."
|
|
41
|
+
},
|
|
42
|
+
{
|
|
43
|
+
"id": "deepseek",
|
|
44
|
+
"platform": "DeepSeek",
|
|
45
|
+
"parser": "DeepSeekParser",
|
|
46
|
+
"module": "decant-core/ai/deepseek",
|
|
47
|
+
"strategy": "DOM + internal API",
|
|
48
|
+
"notes": "API-assisted parsing (fragments[]) with DOM fallback."
|
|
49
|
+
},
|
|
50
|
+
{
|
|
51
|
+
"id": "qwen",
|
|
52
|
+
"platform": "Qwen",
|
|
53
|
+
"parser": "QwenParser",
|
|
54
|
+
"module": "decant-core/ai/qwen",
|
|
55
|
+
"strategy": "DOM",
|
|
56
|
+
"notes": "Structured DOM extraction incl. file attachments."
|
|
57
|
+
},
|
|
58
|
+
{
|
|
59
|
+
"id": "meta",
|
|
60
|
+
"platform": "Meta AI",
|
|
61
|
+
"parser": "MetaParser",
|
|
62
|
+
"module": "decant-core/ai/meta",
|
|
63
|
+
"strategy": "Internal API + DOM",
|
|
64
|
+
"notes": "Internal GraphQL API (pinned + auto-resolved doc_ids) with DOM fallback."
|
|
65
|
+
},
|
|
66
|
+
{
|
|
67
|
+
"id": "mistral",
|
|
68
|
+
"platform": "Mistral / Le Chat",
|
|
69
|
+
"parser": "MistralParser",
|
|
70
|
+
"module": "decant-core/ai/mistral",
|
|
71
|
+
"strategy": "DOM",
|
|
72
|
+
"notes": "DOM extraction."
|
|
73
|
+
},
|
|
74
|
+
{
|
|
75
|
+
"id": "lumo",
|
|
76
|
+
"platform": "Proton Lumo",
|
|
77
|
+
"parser": "LumoParser",
|
|
78
|
+
"module": "decant-core/ai/lumo",
|
|
79
|
+
"strategy": "DOM",
|
|
80
|
+
"notes": "DOM only — API responses are E2E-encrypted and not readable."
|
|
81
|
+
},
|
|
82
|
+
{
|
|
83
|
+
"id": "z-ai",
|
|
84
|
+
"platform": "Z.ai",
|
|
85
|
+
"parser": "ZAiParser",
|
|
86
|
+
"module": "decant-core/ai/z_ai",
|
|
87
|
+
"strategy": "Internal API + DOM",
|
|
88
|
+
"notes": "Internal API (chat skeleton + batched bodies) with DOM fallback."
|
|
89
|
+
},
|
|
90
|
+
{
|
|
91
|
+
"id": "grok",
|
|
92
|
+
"platform": "Grok",
|
|
93
|
+
"parser": "GrokParser",
|
|
94
|
+
"module": "decant-core/ai/grok",
|
|
95
|
+
"strategy": "Internal API + DOM",
|
|
96
|
+
"notes": "Internal API (response-node ordering + load-responses bodies) with DOM fallback."
|
|
97
|
+
},
|
|
98
|
+
{
|
|
99
|
+
"id": "google-ai-studio",
|
|
100
|
+
"platform": "Google AI Studio",
|
|
101
|
+
"parser": "GoogleAIStudioParser",
|
|
102
|
+
"module": "decant-core/ai/google_ai_studio",
|
|
103
|
+
"strategy": "DOM",
|
|
104
|
+
"notes": "DOM extraction for aistudio.google.com sessions."
|
|
105
|
+
},
|
|
106
|
+
{
|
|
107
|
+
"id": "notebooklm",
|
|
108
|
+
"platform": "NotebookLM",
|
|
109
|
+
"parser": "NotebookLMParser",
|
|
110
|
+
"module": "decant-core/ai/notebooklm",
|
|
111
|
+
"strategy": "DOM",
|
|
112
|
+
"notes": "DOM extraction incl. notes and citations."
|
|
113
|
+
},
|
|
114
|
+
{
|
|
115
|
+
"id": "google-search-ai",
|
|
116
|
+
"platform": "Google Search AI (AI Overviews)",
|
|
117
|
+
"parser": "GoogleSearchAIParser",
|
|
118
|
+
"module": "decant-core/ai/google_search_ai",
|
|
119
|
+
"strategy": "DOM",
|
|
120
|
+
"notes": "DOM extraction for SGE overviews."
|
|
121
|
+
},
|
|
122
|
+
{
|
|
123
|
+
"id": "gemini-cloud-assist",
|
|
124
|
+
"platform": "Gemini Cloud Assist",
|
|
125
|
+
"parser": "GeminiCloudAssistParser",
|
|
126
|
+
"module": "decant-core/ai/gemini_cloud_assist",
|
|
127
|
+
"strategy": "DOM",
|
|
128
|
+
"notes": "DOM extraction for console.cloud.google.com assistants."
|
|
129
|
+
},
|
|
130
|
+
{
|
|
131
|
+
"id": "joyland",
|
|
132
|
+
"platform": "Joyland",
|
|
133
|
+
"parser": "JoylandParser",
|
|
134
|
+
"module": "decant-core/ai/joyland",
|
|
135
|
+
"strategy": "DOM",
|
|
136
|
+
"notes": "DOM extraction for character chat platforms."
|
|
137
|
+
},
|
|
138
|
+
{
|
|
139
|
+
"id": "chub",
|
|
140
|
+
"platform": "Chub",
|
|
141
|
+
"parser": "ChubParser",
|
|
142
|
+
"module": "decant-core/ai/chub",
|
|
143
|
+
"strategy": "DOM",
|
|
144
|
+
"notes": "DOM extraction."
|
|
145
|
+
},
|
|
146
|
+
{
|
|
147
|
+
"id": "duck-ai",
|
|
148
|
+
"platform": "Duck.ai (DuckDuckGo AI)",
|
|
149
|
+
"parser": "DuckAIParser",
|
|
150
|
+
"module": "decant-core/ai/duck_ai",
|
|
151
|
+
"strategy": "DOM",
|
|
152
|
+
"notes": "DOM (client-side privacy, no server chat history API)."
|
|
153
|
+
},
|
|
154
|
+
{
|
|
155
|
+
"id": "article",
|
|
156
|
+
"platform": "Generic Web Article",
|
|
157
|
+
"parser": "ArticleParser",
|
|
158
|
+
"module": "decant-core/article",
|
|
159
|
+
"strategy": "Readability + Defuddle + Article-Extractor",
|
|
160
|
+
"notes": "Runs three extractors concurrently and arbitrates by content-quality scoring."
|
|
161
|
+
}
|
|
162
|
+
]
|
|
@@ -17,6 +17,7 @@ import { GeminiCloudAssistParser } from "../ai/gemini_cloud_assist.js";
|
|
|
17
17
|
import { JoylandParser } from "../ai/joyland.js";
|
|
18
18
|
import { ChubParser } from "../ai/chub.js";
|
|
19
19
|
import { GrokParser } from "../ai/grok.js";
|
|
20
|
+
import { DuckAIParser } from "../ai/duck_ai.js";
|
|
20
21
|
|
|
21
22
|
/**
|
|
22
23
|
* Ordered list of parsers. First match wins.
|
|
@@ -32,6 +33,7 @@ export const parsers = [
|
|
|
32
33
|
new QwenParser(),
|
|
33
34
|
new MetaParser(),
|
|
34
35
|
new MistralParser(),
|
|
36
|
+
new DuckAIParser(),
|
|
35
37
|
new LumoParser(),
|
|
36
38
|
new ZAiParser(),
|
|
37
39
|
new GoogleAIStudioParser(),
|
package/detection/domains.js
CHANGED
|
@@ -1,3 +1,5 @@
|
|
|
1
|
+
import { isDuckAiUrl } from "../ai/duck_ai.js";
|
|
2
|
+
|
|
1
3
|
/**
|
|
2
4
|
* AI chat platform domain definitions.
|
|
3
5
|
* Centralised so detection logic and UI (e.g. Decant's tip banner) share the same list.
|
|
@@ -23,6 +25,7 @@ export const AI_CHAT_DOMAINS = [
|
|
|
23
25
|
"chub.ai",
|
|
24
26
|
"characterhub.org",
|
|
25
27
|
"grok.com",
|
|
28
|
+
"duck.ai",
|
|
26
29
|
];
|
|
27
30
|
|
|
28
31
|
/**
|
|
@@ -51,4 +54,8 @@ export const URL_PATTERNS = [
|
|
|
51
54
|
test: (url) => url.includes("console.cloud.google.com/gemini"),
|
|
52
55
|
platform: "gemini-cloud-assist",
|
|
53
56
|
},
|
|
57
|
+
{
|
|
58
|
+
test: isDuckAiUrl,
|
|
59
|
+
platform: "duck-ai",
|
|
60
|
+
},
|
|
54
61
|
];
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "decant-core",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.12.0",
|
|
4
4
|
"description": "A shared web extraction layer for AI conversations and regular web pages.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"engines": {
|
|
@@ -13,7 +13,8 @@
|
|
|
13
13
|
"./web": "./web/index.js",
|
|
14
14
|
"./web/*": "./web/*.js",
|
|
15
15
|
"./article": "./web/article.js",
|
|
16
|
-
"./detection/*": "./detection/*.js"
|
|
16
|
+
"./detection/*": "./detection/*.js",
|
|
17
|
+
"./platforms": "./data/platforms.json"
|
|
17
18
|
},
|
|
18
19
|
"files": [
|
|
19
20
|
"ai",
|
|
@@ -21,22 +22,25 @@
|
|
|
21
22
|
"detection",
|
|
22
23
|
"lib",
|
|
23
24
|
"utils",
|
|
25
|
+
"data",
|
|
24
26
|
"LICENSE",
|
|
25
27
|
"README.md"
|
|
26
28
|
],
|
|
27
29
|
"repository": {
|
|
28
30
|
"type": "git",
|
|
29
|
-
"url": "git+https://github.com/Covai-Labs/decant
|
|
31
|
+
"url": "git+https://github.com/Covai-Labs/decant.git"
|
|
30
32
|
},
|
|
31
33
|
"bugs": {
|
|
32
|
-
"url": "https://github.com/Covai-Labs/decant
|
|
34
|
+
"url": "https://github.com/Covai-Labs/decant/issues"
|
|
33
35
|
},
|
|
34
|
-
"homepage": "https://github.com/Covai-Labs/decant
|
|
36
|
+
"homepage": "https://github.com/Covai-Labs/decant#readme",
|
|
35
37
|
"scripts": {
|
|
36
38
|
"test": "node --test tests/*.test.js",
|
|
37
39
|
"lint": "eslint .",
|
|
38
40
|
"format": "prettier --write .",
|
|
39
|
-
"format:check": "prettier --check ."
|
|
41
|
+
"format:check": "prettier --check .",
|
|
42
|
+
"sync:platforms": "node scripts/sync-platforms.js",
|
|
43
|
+
"sync:check": "node scripts/sync-platforms.js --check"
|
|
40
44
|
},
|
|
41
45
|
"keywords": [
|
|
42
46
|
"ai",
|