decant-core 1.11.0 → 1.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -16,7 +16,7 @@ And every time one of those platforms changes its UI, seriously re-renders a mes
16
16
 
17
17
  Existing chat exporters often suffer from two major flaws: they break whenever platform DOMs update, and many route user conversations through third-party servers.
18
18
 
19
- `decant-core` was created to solve both at the foundation. Originally built to power local-first extensions like [AI Chat Exporter](https://ai-chat-exporter.covai.org/) and [Decant](https://decant.covai.org/), it decouples fragile platform parsing from presentation. By sharing this engine under AGPL-3.0, any browser extension, web clipper, archiver, or research tool can rely on a maintained, local-first extraction layer instead of reverse-engineering AI platforms in isolation.
19
+ `decant-core` was created to solve both at the foundation. Originally built to power local-first extensions like [AI Chat Exporter](https://ace.covai.org/) and [Decant](https://decant.covai.org/), it decouples fragile platform parsing from presentation. By sharing this engine under AGPL-3.0, any browser extension, web clipper, archiver, or research tool can rely on a maintained, local-first extraction layer instead of reverse-engineering AI platforms in isolation.
20
20
 
21
21
  ```bash
22
22
  npm install decant-core
@@ -119,28 +119,28 @@ import { normalizeLatexMath } from "decant-core";
119
119
 
120
120
  19 AI chat platform parsers plus generic web article extraction:
121
121
 
122
- | Platform | Parser | Extraction strategy |
123
- | :---------------------------------- | :------------------------ | :---------------------------------------------------- |
124
- | **ChatGPT** | `ChatGPTParser` | DOM + internal API |
125
- | **Claude** | `ClaudeParser` | DOM + internal API + React fiber |
126
- | **Google Gemini** | `GeminiParser` | DOM + batchexecute RPC |
127
- | **Microsoft Copilot** | `CopilotParser` | DOM (multi-domain) |
128
- | **Perplexity** | `PerplexityParser` | Internal API + DOM fallback |
129
- | **DeepSeek** | `DeepSeekParser` | DOM + internal API (`fragments[]`) |
130
- | **Qwen** | `QwenParser` | DOM |
131
- | **Meta AI** | `MetaParser` | Internal API (GraphQL) + DOM fallback |
132
- | **Mistral / Le Chat** | `MistralParser` | DOM |
133
- | **Proton Lumo** | `LumoParser` | DOM (API is E2E-encrypted, not readable) |
134
- | **Z.ai** | `ZAiParser` | Internal API (chat + batch) + DOM fallback |
135
- | **Grok** | `GrokParser` | Internal API (response-node + load) + DOM |
136
- | **Google AI Studio** | `GoogleAIStudioParser` | DOM |
137
- | **NotebookLM** | `NotebookLMParser` | DOM |
138
- | **Google Search AI (AI Overviews)** | `GoogleSearchAIParser` | DOM |
139
- | **Gemini Cloud Assist** | `GeminiCloudAssistParser` | DOM |
140
- | **Joyland** | `JoylandParser` | DOM |
141
- | **Chub** | `ChubParser` | DOM |
142
- | **Duck.ai (DuckDuckGo AI)** | `DuckAIParser` | DOM (client-side privacy, no server chat history API) |
143
- | **Generic Web Article** | `ArticleParser` | Readability + Defuddle + Article-Extractor |
122
+ | Platform | Parser | Extraction strategy |
123
+ | :---------------------------------- | :------------------------ | :----------------------------------------- |
124
+ | **ChatGPT** | `ChatGPTParser` | DOM + internal API |
125
+ | **Claude** | `ClaudeParser` | DOM + internal API + React fiber |
126
+ | **Google Gemini** | `GeminiParser` | DOM + batchexecute RPC |
127
+ | **Microsoft Copilot** | `CopilotParser` | DOM |
128
+ | **Perplexity** | `PerplexityParser` | Internal API + DOM |
129
+ | **DeepSeek** | `DeepSeekParser` | DOM + internal API |
130
+ | **Qwen** | `QwenParser` | DOM |
131
+ | **Meta AI** | `MetaParser` | Internal API + DOM |
132
+ | **Mistral / Le Chat** | `MistralParser` | DOM |
133
+ | **Proton Lumo** | `LumoParser` | DOM |
134
+ | **Z.ai** | `ZAiParser` | Internal API + DOM |
135
+ | **Grok** | `GrokParser` | Internal API + DOM |
136
+ | **Google AI Studio** | `GoogleAIStudioParser` | DOM |
137
+ | **NotebookLM** | `NotebookLMParser` | DOM |
138
+ | **Google Search AI (AI Overviews)** | `GoogleSearchAIParser` | DOM |
139
+ | **Gemini Cloud Assist** | `GeminiCloudAssistParser` | DOM |
140
+ | **Joyland** | `JoylandParser` | DOM |
141
+ | **Chub** | `ChubParser` | DOM |
142
+ | **Duck.ai (DuckDuckGo AI)** | `DuckAIParser` | DOM |
143
+ | **Generic Web Article** | `ArticleParser` | Readability + Defuddle + Article-Extractor |
144
144
 
145
145
  All parsers extend the base [`ChatParser`](ai/base.js) interface — a consistent `isAvailable(url)` +
146
146
  normalized `parse()` contract. For the full extraction-strategy breakdown and maintenance model, see
package/ai/chatgpt.js CHANGED
@@ -831,6 +831,7 @@ export class ChatGPTParser extends ChatParser {
831
831
  const messages = [];
832
832
  for (const msg of apiMessages) {
833
833
  let content = "";
834
+ let thinking = "";
834
835
  for (const seg of msg.segments) {
835
836
  if (seg.type === "text") {
836
837
  content +=
@@ -843,7 +844,7 @@ export class ChatGPTParser extends ChatParser {
843
844
  msg.imageGroupMap,
844
845
  );
845
846
  if (thoughtText) {
846
- content += `<details><summary>Thought Process</summary>\n\n${thoughtText}\n\n</details>\n\n`;
847
+ thinking += (thinking ? "\n\n" : "") + thoughtText;
847
848
  }
848
849
  } else if (seg.type === "image") {
849
850
  const src = images[seg.fileId];
@@ -852,12 +853,22 @@ export class ChatGPTParser extends ChatParser {
852
853
  }
853
854
  }
854
855
  }
855
- content = content.trim();
856
- if (content) {
856
+ let fullContent = "";
857
+ if (thinking) {
858
+ fullContent += `<think>\n${thinking}\n</think>\n\n`;
859
+ }
860
+ if (content.trim()) {
861
+ fullContent += content.trim();
862
+ }
863
+ fullContent = fullContent.trim();
864
+ if (fullContent) {
857
865
  const msgObj = {
858
866
  role: msg.role,
859
- content: content,
867
+ content: fullContent,
860
868
  };
869
+ if (thinking) {
870
+ msgObj.thinking = thinking;
871
+ }
861
872
  if (msg.timestamp) {
862
873
  msgObj.timestamp = msg.timestamp;
863
874
  }
@@ -876,8 +887,10 @@ export class ChatGPTParser extends ChatParser {
876
887
  Link: currentUrl,
877
888
  Model:
878
889
  convoData?.model_slug ||
879
- document.querySelector('[data-testid="model-selector-dropdown"]')
880
- ?.innerText ||
890
+ (typeof document !== "undefined" && document.querySelector
891
+ ? document.querySelector('[data-testid="model-selector-dropdown"]')
892
+ ?.innerText
893
+ : null) ||
881
894
  "ChatGPT",
882
895
  Method: method,
883
896
  };
package/ai/claude.js CHANGED
@@ -402,12 +402,30 @@ export class ClaudeParser extends ChatParser {
402
402
  const role = message.sender === "human" ? "User" : "Claude";
403
403
 
404
404
  let contentStr = "";
405
+ let thinkingStr = "";
405
406
 
406
407
  // Construct content
407
408
  if (message.content && Array.isArray(message.content)) {
408
409
  for (const block of message.content) {
409
- if (block.type === "thinking" && block.thinking) {
410
- contentStr += `> **Thinking Process:**\n> \n> ${block.thinking.replace(/\n/g, "\n> ")}\n\n`;
410
+ if (block.type === "thinking") {
411
+ let thoughtText = "";
412
+ if (
413
+ typeof block.thinking === "string" &&
414
+ block.thinking.trim()
415
+ ) {
416
+ thoughtText = block.thinking.trim();
417
+ } else if (Array.isArray(block.summaries)) {
418
+ thoughtText = block.summaries
419
+ .map((s) =>
420
+ typeof s === "string" ? s : s?.summary || "",
421
+ )
422
+ .map((s) => s.trim())
423
+ .filter(Boolean)
424
+ .join("\n");
425
+ }
426
+ if (thoughtText) {
427
+ thinkingStr += (thinkingStr ? "\n\n" : "") + thoughtText;
428
+ }
411
429
  } else if (block.type === "text" && block.text) {
412
430
  const cleanText = block.text
413
431
  .replace(/<antArtifact[^>]*>[\s\S]*?<\/antArtifact>/g, "")
@@ -516,9 +534,17 @@ export class ClaudeParser extends ChatParser {
516
534
  }
517
535
  }
518
536
 
537
+ if (thinkingStr) {
538
+ contentStr = `<think>\n${thinkingStr}\n</think>\n\n` + contentStr;
539
+ }
540
+
519
541
  contentStr = contentStr.trim();
520
542
  if (contentStr) {
521
- messages.push({ role, content: contentStr });
543
+ const msgObj = { role, content: contentStr };
544
+ if (thinkingStr) {
545
+ msgObj.thinking = thinkingStr;
546
+ }
547
+ messages.push(msgObj);
522
548
  }
523
549
 
524
550
  // Extract and push artifacts
package/ai/deepseek.js CHANGED
@@ -98,7 +98,29 @@ async function fetchDeepSeekConversation(sessionId, token) {
98
98
  const isUser = msgNode.role === "USER" || msgNode.role === "user";
99
99
  const role = isUser ? "User" : "DeepSeek";
100
100
  const content = extractDeepSeekMessageContent(msgNode);
101
- return { role, content: content.trim() };
101
+ let thinking = "";
102
+ if (!isUser && Array.isArray(msgNode.fragments)) {
103
+ const thinkFragments = msgNode.fragments.filter(
104
+ (f) => f && f.type === "THINK" && typeof f.content === "string",
105
+ );
106
+ thinking = thinkFragments
107
+ .map((f) => f.content.trim())
108
+ .filter(Boolean)
109
+ .join("\n\n");
110
+ }
111
+ let fullContent = "";
112
+ if (thinking) {
113
+ fullContent += `<think>\n${thinking}\n</think>\n\n`;
114
+ }
115
+ if (content.trim()) {
116
+ fullContent += content.trim();
117
+ }
118
+ fullContent = fullContent.trim();
119
+ const msg = { role, content: fullContent };
120
+ if (thinking) {
121
+ msg.thinking = thinking;
122
+ }
123
+ return msg;
102
124
  })
103
125
  .filter((msg) => msg.content.length > 0);
104
126
  }
@@ -182,15 +204,40 @@ export class DeepSeekParser extends ChatParser {
182
204
 
183
205
  outerElements.forEach((el) => {
184
206
  let role = "Unknown";
207
+ let thinking = "";
185
208
  if (el.matches(userSelector)) {
186
209
  role = "User";
187
210
  } else if (el.matches(assistantSelector)) {
188
211
  role = "DeepSeek";
212
+ const messageContainer =
213
+ (el.closest && el.closest(".ds-message")) || el.parentElement;
214
+ if (messageContainer) {
215
+ const thinkContainers =
216
+ messageContainer.querySelectorAll(".ds-think-content");
217
+ if (thinkContainers.length > 0) {
218
+ thinking = Array.from(thinkContainers)
219
+ .map((tc) => convertToMarkdown(tc).trim())
220
+ .filter(Boolean)
221
+ .join("\n\n");
222
+ }
223
+ }
189
224
  }
190
225
 
191
226
  const text = convertToMarkdown(el);
227
+ let fullContent = "";
228
+ if (thinking) {
229
+ fullContent += `<think>\n${thinking}\n</think>\n\n`;
230
+ }
192
231
  if (text.trim()) {
193
- messages.push({ role, content: text.trim() });
232
+ fullContent += text.trim();
233
+ }
234
+ fullContent = fullContent.trim();
235
+ if (fullContent) {
236
+ const msg = { role, content: fullContent };
237
+ if (thinking) {
238
+ msg.thinking = thinking;
239
+ }
240
+ messages.push(msg);
194
241
  }
195
242
  });
196
243
 
@@ -202,9 +249,37 @@ export class DeepSeekParser extends ChatParser {
202
249
  messageRows.forEach((row) => {
203
250
  const isUser = row.classList.contains("ds-user-message");
204
251
  const role = isUser ? "User" : "DeepSeek";
205
- const text = convertToMarkdown(row);
252
+ const rowClone = row.cloneNode(true);
253
+ let thinking = "";
254
+ if (!isUser) {
255
+ const thinkContainers =
256
+ rowClone.querySelectorAll(".ds-think-content");
257
+ if (thinkContainers.length > 0) {
258
+ thinking = Array.from(thinkContainers)
259
+ .map((tc) => {
260
+ const md = convertToMarkdown(tc).trim();
261
+ tc.remove();
262
+ return md;
263
+ })
264
+ .filter(Boolean)
265
+ .join("\n\n");
266
+ }
267
+ }
268
+ const text = convertToMarkdown(rowClone);
269
+ let fullContent = "";
270
+ if (thinking) {
271
+ fullContent += `<think>\n${thinking}\n</think>\n\n`;
272
+ }
206
273
  if (text.trim()) {
207
- messages.push({ role, content: text.trim() });
274
+ fullContent += text.trim();
275
+ }
276
+ fullContent = fullContent.trim();
277
+ if (fullContent) {
278
+ const msg = { role, content: fullContent };
279
+ if (thinking) {
280
+ msg.thinking = thinking;
281
+ }
282
+ messages.push(msg);
208
283
  }
209
284
  });
210
285
  }
package/ai/z_ai.js CHANGED
@@ -60,13 +60,12 @@ export function formatZaiMessage(entry) {
60
60
  .filter(Boolean)
61
61
  .join("\n\n");
62
62
  if (reasoning) {
63
- const quoted = reasoning
64
- .split("\n")
65
- .map((line) => `> ${line}`)
66
- .join("\n");
67
- content += `\n\n> 🧠 Thinking\n${quoted}`;
63
+ content = `<think>\n${reasoning}\n</think>\n\n${content}`;
68
64
  }
69
65
  const msg = { role, content };
66
+ if (reasoning) {
67
+ msg.thinking = reasoning;
68
+ }
70
69
  if (entry.timestamp) {
71
70
  try {
72
71
  msg.timestamp = new Date(entry.timestamp * 1000).toISOString();
@@ -0,0 +1,162 @@
1
+ [
2
+ {
3
+ "id": "chatgpt",
4
+ "platform": "ChatGPT",
5
+ "parser": "ChatGPTParser",
6
+ "module": "decant-core/ai/chatgpt",
7
+ "strategy": "DOM + internal API",
8
+ "notes": "Fallback and RPC-assisted extraction; scroll/dedup helper for long threads."
9
+ },
10
+ {
11
+ "id": "claude",
12
+ "platform": "Claude",
13
+ "parser": "ClaudeParser",
14
+ "module": "decant-core/ai/claude",
15
+ "strategy": "DOM + internal API + React fiber",
16
+ "notes": "Internal API first with DOM fallback; reads the React tree for artifacts and structured blocks."
17
+ },
18
+ {
19
+ "id": "gemini",
20
+ "platform": "Google Gemini",
21
+ "parser": "GeminiParser",
22
+ "module": "decant-core/ai/gemini",
23
+ "strategy": "DOM + batchexecute RPC",
24
+ "notes": "Uses batchexecute RPC pagination with resilient DOM fallback."
25
+ },
26
+ {
27
+ "id": "copilot",
28
+ "platform": "Microsoft Copilot",
29
+ "parser": "CopilotParser",
30
+ "module": "decant-core/ai/copilot",
31
+ "strategy": "DOM",
32
+ "notes": "Multi-domain (bing + copilot) DOM extraction."
33
+ },
34
+ {
35
+ "id": "perplexity",
36
+ "platform": "Perplexity",
37
+ "parser": "PerplexityParser",
38
+ "module": "decant-core/ai/perplexity",
39
+ "strategy": "Internal API + DOM",
40
+ "notes": "Internal API first, DOM fallback for source citations and answers."
41
+ },
42
+ {
43
+ "id": "deepseek",
44
+ "platform": "DeepSeek",
45
+ "parser": "DeepSeekParser",
46
+ "module": "decant-core/ai/deepseek",
47
+ "strategy": "DOM + internal API",
48
+ "notes": "API-assisted parsing (fragments[]) with DOM fallback."
49
+ },
50
+ {
51
+ "id": "qwen",
52
+ "platform": "Qwen",
53
+ "parser": "QwenParser",
54
+ "module": "decant-core/ai/qwen",
55
+ "strategy": "DOM",
56
+ "notes": "Structured DOM extraction incl. file attachments."
57
+ },
58
+ {
59
+ "id": "meta",
60
+ "platform": "Meta AI",
61
+ "parser": "MetaParser",
62
+ "module": "decant-core/ai/meta",
63
+ "strategy": "Internal API + DOM",
64
+ "notes": "Internal GraphQL API (pinned + auto-resolved doc_ids) with DOM fallback."
65
+ },
66
+ {
67
+ "id": "mistral",
68
+ "platform": "Mistral / Le Chat",
69
+ "parser": "MistralParser",
70
+ "module": "decant-core/ai/mistral",
71
+ "strategy": "DOM",
72
+ "notes": "DOM extraction."
73
+ },
74
+ {
75
+ "id": "lumo",
76
+ "platform": "Proton Lumo",
77
+ "parser": "LumoParser",
78
+ "module": "decant-core/ai/lumo",
79
+ "strategy": "DOM",
80
+ "notes": "DOM only — API responses are E2E-encrypted and not readable."
81
+ },
82
+ {
83
+ "id": "z-ai",
84
+ "platform": "Z.ai",
85
+ "parser": "ZAiParser",
86
+ "module": "decant-core/ai/z_ai",
87
+ "strategy": "Internal API + DOM",
88
+ "notes": "Internal API (chat skeleton + batched bodies) with DOM fallback."
89
+ },
90
+ {
91
+ "id": "grok",
92
+ "platform": "Grok",
93
+ "parser": "GrokParser",
94
+ "module": "decant-core/ai/grok",
95
+ "strategy": "Internal API + DOM",
96
+ "notes": "Internal API (response-node ordering + load-responses bodies) with DOM fallback."
97
+ },
98
+ {
99
+ "id": "google-ai-studio",
100
+ "platform": "Google AI Studio",
101
+ "parser": "GoogleAIStudioParser",
102
+ "module": "decant-core/ai/google_ai_studio",
103
+ "strategy": "DOM",
104
+ "notes": "DOM extraction for aistudio.google.com sessions."
105
+ },
106
+ {
107
+ "id": "notebooklm",
108
+ "platform": "NotebookLM",
109
+ "parser": "NotebookLMParser",
110
+ "module": "decant-core/ai/notebooklm",
111
+ "strategy": "DOM",
112
+ "notes": "DOM extraction incl. notes and citations."
113
+ },
114
+ {
115
+ "id": "google-search-ai",
116
+ "platform": "Google Search AI (AI Overviews)",
117
+ "parser": "GoogleSearchAIParser",
118
+ "module": "decant-core/ai/google_search_ai",
119
+ "strategy": "DOM",
120
+ "notes": "DOM extraction for SGE overviews."
121
+ },
122
+ {
123
+ "id": "gemini-cloud-assist",
124
+ "platform": "Gemini Cloud Assist",
125
+ "parser": "GeminiCloudAssistParser",
126
+ "module": "decant-core/ai/gemini_cloud_assist",
127
+ "strategy": "DOM",
128
+ "notes": "DOM extraction for console.cloud.google.com assistants."
129
+ },
130
+ {
131
+ "id": "joyland",
132
+ "platform": "Joyland",
133
+ "parser": "JoylandParser",
134
+ "module": "decant-core/ai/joyland",
135
+ "strategy": "DOM",
136
+ "notes": "DOM extraction for character chat platforms."
137
+ },
138
+ {
139
+ "id": "chub",
140
+ "platform": "Chub",
141
+ "parser": "ChubParser",
142
+ "module": "decant-core/ai/chub",
143
+ "strategy": "DOM",
144
+ "notes": "DOM extraction."
145
+ },
146
+ {
147
+ "id": "duck-ai",
148
+ "platform": "Duck.ai (DuckDuckGo AI)",
149
+ "parser": "DuckAIParser",
150
+ "module": "decant-core/ai/duck_ai",
151
+ "strategy": "DOM",
152
+ "notes": "DOM (client-side privacy, no server chat history API)."
153
+ },
154
+ {
155
+ "id": "article",
156
+ "platform": "Generic Web Article",
157
+ "parser": "ArticleParser",
158
+ "module": "decant-core/article",
159
+ "strategy": "Readability + Defuddle + Article-Extractor",
160
+ "notes": "Runs three extractors concurrently and arbitrates by content-quality scoring."
161
+ }
162
+ ]
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "decant-core",
3
- "version": "1.11.0",
3
+ "version": "1.12.0",
4
4
  "description": "A shared web extraction layer for AI conversations and regular web pages.",
5
5
  "type": "module",
6
6
  "engines": {
@@ -13,7 +13,8 @@
13
13
  "./web": "./web/index.js",
14
14
  "./web/*": "./web/*.js",
15
15
  "./article": "./web/article.js",
16
- "./detection/*": "./detection/*.js"
16
+ "./detection/*": "./detection/*.js",
17
+ "./platforms": "./data/platforms.json"
17
18
  },
18
19
  "files": [
19
20
  "ai",
@@ -21,6 +22,7 @@
21
22
  "detection",
22
23
  "lib",
23
24
  "utils",
25
+ "data",
24
26
  "LICENSE",
25
27
  "README.md"
26
28
  ],
@@ -36,7 +38,9 @@
36
38
  "test": "node --test tests/*.test.js",
37
39
  "lint": "eslint .",
38
40
  "format": "prettier --write .",
39
- "format:check": "prettier --check ."
41
+ "format:check": "prettier --check .",
42
+ "sync:platforms": "node scripts/sync-platforms.js",
43
+ "sync:check": "node scripts/sync-platforms.js --check"
40
44
  },
41
45
  "keywords": [
42
46
  "ai",