decant-core 1.11.0 → 1.13.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE CHANGED
@@ -1,3 +1,19 @@
1
+ ADDITIONAL PERMISSION UNDER GNU AGPL VERSION 3 SECTION 7:
2
+ FOSS LINKING EXCEPTION
3
+
4
+ Permission is granted to link, import, or bundle decant-core into projects
5
+ distributed under any OSI-approved open source license (including MPL-2.0, MIT,
6
+ Apache-2.0, and BSD) and distribute the resulting work under that project's
7
+ license, without requiring the enclosing project to be licensed under AGPLv3.
8
+ Any modifications directly made to decant-core source files remain subject to AGPLv3.
9
+
10
+ COMMERCIAL USE
11
+ If you wish to use decant-core in closed-source, proprietary, or commercial
12
+ software that cannot comply with the AGPLv3, a commercial license is available.
13
+ Please contact office@covai.org for licensing terms.
14
+
15
+ ==============================================================================
16
+
1
17
  GNU AFFERO GENERAL PUBLIC LICENSE
2
18
  Version 3, 19 November 2007
3
19
 
package/README.md CHANGED
@@ -16,7 +16,7 @@ And every time one of those platforms changes its UI, seriously re-renders a mes
16
16
 
17
17
  Existing chat exporters often suffer from two major flaws: they break whenever platform DOMs update, and many route user conversations through third-party servers.
18
18
 
19
- `decant-core` was created to solve both at the foundation. Originally built to power local-first extensions like [AI Chat Exporter](https://ai-chat-exporter.covai.org/) and [Decant](https://decant.covai.org/), it decouples fragile platform parsing from presentation. By sharing this engine under AGPL-3.0, any browser extension, web clipper, archiver, or research tool can rely on a maintained, local-first extraction layer instead of reverse-engineering AI platforms in isolation.
19
+ `decant-core` was created to solve both at the foundation. Originally built to power local-first extensions like [AI Chat Exporter](https://ace.covai.org/) and [Decant](https://decant.covai.org/), it decouples fragile platform parsing from presentation. By sharing this engine under AGPL-3.0, any browser extension, web clipper, archiver, or research tool can rely on a maintained, local-first extraction layer instead of reverse-engineering AI platforms in isolation.
20
20
 
21
21
  ```bash
22
22
  npm install decant-core
@@ -119,28 +119,28 @@ import { normalizeLatexMath } from "decant-core";
119
119
 
120
120
  19 AI chat platform parsers plus generic web article extraction:
121
121
 
122
- | Platform | Parser | Extraction strategy |
123
- | :---------------------------------- | :------------------------ | :---------------------------------------------------- |
124
- | **ChatGPT** | `ChatGPTParser` | DOM + internal API |
125
- | **Claude** | `ClaudeParser` | DOM + internal API + React fiber |
126
- | **Google Gemini** | `GeminiParser` | DOM + batchexecute RPC |
127
- | **Microsoft Copilot** | `CopilotParser` | DOM (multi-domain) |
128
- | **Perplexity** | `PerplexityParser` | Internal API + DOM fallback |
129
- | **DeepSeek** | `DeepSeekParser` | DOM + internal API (`fragments[]`) |
130
- | **Qwen** | `QwenParser` | DOM |
131
- | **Meta AI** | `MetaParser` | Internal API (GraphQL) + DOM fallback |
132
- | **Mistral / Le Chat** | `MistralParser` | DOM |
133
- | **Proton Lumo** | `LumoParser` | DOM (API is E2E-encrypted, not readable) |
134
- | **Z.ai** | `ZAiParser` | Internal API (chat + batch) + DOM fallback |
135
- | **Grok** | `GrokParser` | Internal API (response-node + load) + DOM |
136
- | **Google AI Studio** | `GoogleAIStudioParser` | DOM |
137
- | **NotebookLM** | `NotebookLMParser` | DOM |
138
- | **Google Search AI (AI Overviews)** | `GoogleSearchAIParser` | DOM |
139
- | **Gemini Cloud Assist** | `GeminiCloudAssistParser` | DOM |
140
- | **Joyland** | `JoylandParser` | DOM |
141
- | **Chub** | `ChubParser` | DOM |
142
- | **Duck.ai (DuckDuckGo AI)** | `DuckAIParser` | DOM (client-side privacy, no server chat history API) |
143
- | **Generic Web Article** | `ArticleParser` | Readability + Defuddle + Article-Extractor |
122
+ | Platform | Parser | Extraction strategy |
123
+ | :---------------------------------- | :------------------------ | :----------------------------------------- |
124
+ | **ChatGPT** | `ChatGPTParser` | DOM + internal API |
125
+ | **Claude** | `ClaudeParser` | DOM + internal API + React fiber |
126
+ | **Google Gemini** | `GeminiParser` | DOM + batchexecute RPC |
127
+ | **Microsoft Copilot** | `CopilotParser` | DOM |
128
+ | **Perplexity** | `PerplexityParser` | Internal API + DOM |
129
+ | **DeepSeek** | `DeepSeekParser` | DOM + internal API |
130
+ | **Qwen** | `QwenParser` | DOM |
131
+ | **Meta AI** | `MetaParser` | Internal API + DOM |
132
+ | **Mistral / Le Chat** | `MistralParser` | DOM |
133
+ | **Proton Lumo** | `LumoParser` | DOM |
134
+ | **Z.ai** | `ZAiParser` | Internal API + DOM |
135
+ | **Grok** | `GrokParser` | Internal API + DOM |
136
+ | **Google AI Studio** | `GoogleAIStudioParser` | DOM |
137
+ | **NotebookLM** | `NotebookLMParser` | DOM |
138
+ | **Google Search AI (AI Overviews)** | `GoogleSearchAIParser` | DOM |
139
+ | **Gemini Cloud Assist** | `GeminiCloudAssistParser` | DOM |
140
+ | **Joyland** | `JoylandParser` | DOM |
141
+ | **Chub** | `ChubParser` | DOM |
142
+ | **Duck.ai (DuckDuckGo AI)** | `DuckAIParser` | DOM |
143
+ | **Generic Web Article** | `ArticleParser` | Readability + Defuddle + Article-Extractor |
144
144
 
145
145
  All parsers extend the base [`ChatParser`](ai/base.js) interface — a consistent `isAvailable(url)` +
146
146
  normalized `parse()` contract. For the full extraction-strategy breakdown and maintenance model, see
@@ -151,9 +151,15 @@ normalized `parse()` contract. For the full extraction-strategy breakdown and ma
151
151
 
152
152
  ## License
153
153
 
154
- `decant-core` is licensed under the **GNU Affero General Public License v3.0 (AGPL-3.0-only)**.
154
+ `decant-core` is licensed under the **GNU Affero General Public License v3.0 (AGPL-3.0-only)** with a **FOSS Linking Exception**, alongside a **Commercial License** option.
155
155
 
156
- That choice is deliberate. AI platforms change constantly, and parser fixes belong in a shared commons so the whole ecosystem benefits — not siloed in a proprietary fork. If you use `decant-core`, network-based deployments that serve modified versions must also offer the corresponding source. Please review [`LICENSE`](LICENSE) before incorporating it into your project.
156
+ ### FOSS Linking Exception (Open Source)
157
+
158
+ Permission is granted to link, import, or bundle `decant-core` into projects distributed under any OSI-approved open source license (including **MPL-2.0**, **MIT**, **Apache-2.0**, and **BSD**) and distribute the resulting work under that project's license, without requiring the enclosing project to be licensed under AGPLv3. Any modifications directly made to `decant-core` source files remain subject to AGPLv3.
159
+
160
+ ### Commercial License
161
+
162
+ If you wish to use `decant-core` in closed-source, proprietary, or commercial software that cannot comply with the AGPLv3, a commercial license is available. Please contact `office@covai.org` for licensing terms.
157
163
 
158
164
  ### Third-Party Test Fixtures Notice
159
165
 
package/ai/chatgpt.js CHANGED
@@ -196,6 +196,221 @@ export function extractSharedConversationFromDom(
196
196
  return null;
197
197
  }
198
198
 
199
+ function cleanApiPartText(partText) {
200
+ return partText
201
+ .replace(/\u{E0000}[\u{E0000}-\u{E007F}]*/gu, "")
202
+ .replace(/citeturn\d+\w*/g, "")
203
+ .trim();
204
+ }
205
+
206
+ // Linearize the newer backend-api conversation shape, which returns a
207
+ // `messages` array instead of a `mapping` tree. Output segments match the
208
+ // mapping-based linearize() format ({ type: "text" | "thought" | "image" })
209
+ // so both flow through the same formatApiResult().
210
+ export function linearizeMessagesArray(apiMessages, includeImages) {
211
+ const messages = [];
212
+ if (!Array.isArray(apiMessages)) return messages;
213
+
214
+ const pushOrMerge = (entry, isThoughtMsg) => {
215
+ if (
216
+ entry.role === "ChatGPT" &&
217
+ messages.length > 0 &&
218
+ messages[messages.length - 1].role === "ChatGPT"
219
+ ) {
220
+ const prevMsg = messages[messages.length - 1];
221
+ if (isThoughtMsg) {
222
+ prevMsg.segments.unshift(...entry.segments);
223
+ } else {
224
+ prevMsg.segments.push(...entry.segments);
225
+ }
226
+ Object.assign(prevMsg.citeMap, entry.citeMap);
227
+ Object.assign(prevMsg.imageGroupMap, entry.imageGroupMap);
228
+ if (entry.timestamp && !prevMsg.timestamp) {
229
+ prevMsg.timestamp = entry.timestamp;
230
+ }
231
+ } else {
232
+ messages.push(entry);
233
+ }
234
+ };
235
+
236
+ for (const msg of apiMessages) {
237
+ if (!msg) continue;
238
+ const role = msg?.author?.role;
239
+ if (role !== "user" && role !== "assistant" && role !== "tool") continue;
240
+ if (msg.metadata?.is_visually_hidden_from_conversation === true) continue;
241
+
242
+ const content = msg.content || {};
243
+ const contentType = content.content_type;
244
+ const segments = [];
245
+ // Mirror the mapping path's thought detection (content types plus
246
+ // author/recipient markers).
247
+ const isThoughtMsg =
248
+ msg?.author?.name === "thought" ||
249
+ msg?.recipient === "thought" ||
250
+ contentType === "thought" ||
251
+ contentType === "thoughts" ||
252
+ contentType === "reasoning_recap" ||
253
+ msg?.metadata?.reasoning_status === "is_reasoning";
254
+
255
+ // Reasoning summaries: thoughts = [{ summary, content }]
256
+ if (Array.isArray(content.thoughts) && content.thoughts.length > 0) {
257
+ const thoughtParts = content.thoughts
258
+ .map((t) =>
259
+ t.summary ? `**${t.summary}**\n${t.content || ""}` : t.content || "",
260
+ )
261
+ .map((t) => t.trim())
262
+ .filter(Boolean);
263
+ if (thoughtParts.length > 0) {
264
+ segments.push({ type: "thought", content: thoughtParts.join("\n\n") });
265
+ }
266
+ }
267
+
268
+ // Reasoning recap (e.g. "Worked for 11s")
269
+ if (
270
+ contentType === "reasoning_recap" &&
271
+ typeof content.content === "string" &&
272
+ content.content.trim()
273
+ ) {
274
+ segments.push({ type: "thought", content: content.content.trim() });
275
+ }
276
+
277
+ // Tool invocations (content_type "code", e.g. Deep Research args JSON)
278
+ // and tool-role messages are not user-visible prose — skip them.
279
+ if (contentType !== "code" && role !== "tool") {
280
+ const parts = Array.isArray(content.parts) ? content.parts : [];
281
+ for (const part of parts) {
282
+ let partText = "";
283
+ let isThoughtPart = isThoughtMsg;
284
+ if (typeof part === "string") {
285
+ partText = part;
286
+ } else if (part && typeof part === "object") {
287
+ if (part.content_type === "text" && typeof part.text === "string") {
288
+ partText = part.text;
289
+ } else if (
290
+ part.content_type === "thought" &&
291
+ typeof part.text === "string"
292
+ ) {
293
+ partText = part.text;
294
+ isThoughtPart = true;
295
+ } else if (
296
+ part.content_type === "audio_transcription" &&
297
+ typeof part.text === "string"
298
+ ) {
299
+ partText = part.text;
300
+ } else if (
301
+ includeImages &&
302
+ part?.content_type === "image_asset_pointer" &&
303
+ part?.asset_pointer
304
+ ) {
305
+ segments.push({
306
+ type: "image",
307
+ fileId: part.asset_pointer.split("://")[1],
308
+ });
309
+ continue;
310
+ }
311
+ }
312
+ const text = partText ? cleanApiPartText(partText) : "";
313
+ if (text) {
314
+ segments.push({
315
+ type: isThoughtPart ? "thought" : "text",
316
+ content: text,
317
+ });
318
+ }
319
+ }
320
+
321
+ // Standalone content.text without parts (plain text / execution
322
+ // output), mirroring the mapping path.
323
+ if (
324
+ parts.length === 0 &&
325
+ typeof content.text === "string" &&
326
+ content.text.trim()
327
+ ) {
328
+ segments.push({ type: "text", content: content.text.trim() });
329
+ }
330
+ }
331
+
332
+ // Deep Research reports (widget_state), attachments, and Canvas
333
+ // documents from message metadata, mirroring the mapping path.
334
+ if (role !== "tool") {
335
+ const widgetRaw =
336
+ msg.metadata?.chatgpt_sdk?.widget_state ||
337
+ msg.metadata?.tool_response_metadata?.venus_widget_state;
338
+ if (widgetRaw) {
339
+ try {
340
+ const widget =
341
+ typeof widgetRaw === "string" ? JSON.parse(widgetRaw) : widgetRaw;
342
+ const reportText =
343
+ widget.report_message?.content?.parts?.[0] || widget.markdown;
344
+ const steering = widget.steering_acknowledgement;
345
+ let researchContent = "";
346
+ if (steering) researchContent += `${steering}\n\n`;
347
+ if (reportText) researchContent += reportText;
348
+ if (researchContent.trim()) {
349
+ segments.push({ type: "text", content: researchContent.trim() });
350
+ }
351
+ } catch {
352
+ // Ignore widget state JSON parse errors
353
+ }
354
+ }
355
+
356
+ if (
357
+ Array.isArray(msg.metadata?.attachments) &&
358
+ msg.metadata.attachments.length > 0
359
+ ) {
360
+ const fileNames = msg.metadata.attachments
361
+ .map((att) => att.name)
362
+ .filter(Boolean);
363
+ if (fileNames.length > 0) {
364
+ segments.push({
365
+ type: "text",
366
+ content: `[Attached: ${fileNames.join(", ")}]`,
367
+ });
368
+ }
369
+ }
370
+
371
+ if (msg.metadata?.canvas?.title) {
372
+ segments.push({
373
+ type: "text",
374
+ content: `[Canvas: ${msg.metadata.canvas.title}]`,
375
+ });
376
+ }
377
+ }
378
+
379
+ if (segments.length === 0) continue;
380
+
381
+ const displayRole = role === "user" ? "User" : "ChatGPT";
382
+ const timestamp = msg?.create_time
383
+ ? new Date(msg.create_time * 1000).toLocaleString()
384
+ : null;
385
+ // Citation / image-group references, mirroring the mapping path.
386
+ const citeMap = {};
387
+ const imageGroupMap = {};
388
+ for (const ref of msg?.metadata?.content_references ?? []) {
389
+ if (ref.matched_text) {
390
+ if (ref.items?.length) citeMap[ref.matched_text] = ref.items;
391
+ if (
392
+ ref.type === "image_group" ||
393
+ ref.matched_text.includes("image_group")
394
+ ) {
395
+ imageGroupMap[ref.matched_text] = ref;
396
+ }
397
+ }
398
+ }
399
+ pushOrMerge(
400
+ {
401
+ role: displayRole,
402
+ segments,
403
+ citeMap,
404
+ imageGroupMap,
405
+ timestamp,
406
+ },
407
+ isThoughtMsg,
408
+ );
409
+ }
410
+
411
+ return messages;
412
+ }
413
+
199
414
  export function linearize(mapping, includeImages, currentNodeId) {
200
415
  let path = [];
201
416
  const leafId = resolveActiveLeafNode(mapping, currentNodeId);
@@ -268,8 +483,20 @@ export function linearize(mapping, includeImages, currentNodeId) {
268
483
  ) {
269
484
  const segments = [];
270
485
  const parts = msg?.content?.parts ?? [];
486
+ // Tool-invocation payloads (content_type "code") are not user-visible
487
+ // prose, whether carried as standalone text or inside parts.
488
+ const isToolInvocation = msg?.content?.content_type === "code";
271
489
 
272
490
  for (const part of parts) {
491
+ if (isToolInvocation) {
492
+ if (!(
493
+ includeImages &&
494
+ part?.content_type === "image_asset_pointer" &&
495
+ part?.asset_pointer
496
+ )) {
497
+ continue;
498
+ }
499
+ }
273
500
  let partText = "";
274
501
  let isThoughtPart = isThoughtMsg;
275
502
 
@@ -316,11 +543,15 @@ export function linearize(mapping, includeImages, currentNodeId) {
316
543
  }
317
544
  }
318
545
 
319
- // Handle standalone content.text (e.g. execution_output or plain text)
546
+ // Handle standalone content.text (e.g. execution_output or plain text).
547
+ // Tool-invocation payloads (content_type "code", e.g. Deep Research
548
+ // "/Deep Research App/start" args JSON) are not user-visible prose —
549
+ // the site renders a status card instead — so never dump them as text.
320
550
  if (
321
551
  typeof msg.content?.text === "string" &&
322
552
  msg.content.text.trim() &&
323
- parts.length === 0
553
+ parts.length === 0 &&
554
+ msg.content?.content_type !== "code"
324
555
  ) {
325
556
  segments.push({ type: "text", content: msg.content.text.trim() });
326
557
  }
@@ -831,6 +1062,7 @@ export class ChatGPTParser extends ChatParser {
831
1062
  const messages = [];
832
1063
  for (const msg of apiMessages) {
833
1064
  let content = "";
1065
+ let thinking = "";
834
1066
  for (const seg of msg.segments) {
835
1067
  if (seg.type === "text") {
836
1068
  content +=
@@ -843,7 +1075,7 @@ export class ChatGPTParser extends ChatParser {
843
1075
  msg.imageGroupMap,
844
1076
  );
845
1077
  if (thoughtText) {
846
- content += `<details><summary>Thought Process</summary>\n\n${thoughtText}\n\n</details>\n\n`;
1078
+ thinking += (thinking ? "\n\n" : "") + thoughtText;
847
1079
  }
848
1080
  } else if (seg.type === "image") {
849
1081
  const src = images[seg.fileId];
@@ -852,12 +1084,22 @@ export class ChatGPTParser extends ChatParser {
852
1084
  }
853
1085
  }
854
1086
  }
855
- content = content.trim();
856
- if (content) {
1087
+ let fullContent = "";
1088
+ if (thinking) {
1089
+ fullContent += `<think>\n${thinking}\n</think>\n\n`;
1090
+ }
1091
+ if (content.trim()) {
1092
+ fullContent += content.trim();
1093
+ }
1094
+ fullContent = fullContent.trim();
1095
+ if (fullContent) {
857
1096
  const msgObj = {
858
1097
  role: msg.role,
859
- content: content,
1098
+ content: fullContent,
860
1099
  };
1100
+ if (thinking) {
1101
+ msgObj.thinking = thinking;
1102
+ }
861
1103
  if (msg.timestamp) {
862
1104
  msgObj.timestamp = msg.timestamp;
863
1105
  }
@@ -876,8 +1118,11 @@ export class ChatGPTParser extends ChatParser {
876
1118
  Link: currentUrl,
877
1119
  Model:
878
1120
  convoData?.model_slug ||
879
- document.querySelector('[data-testid="model-selector-dropdown"]')
880
- ?.innerText ||
1121
+ convoData?.default_model_slug ||
1122
+ (typeof document !== "undefined" && document.querySelector
1123
+ ? document.querySelector('[data-testid="model-selector-dropdown"]')
1124
+ ?.innerText
1125
+ : null) ||
881
1126
  "ChatGPT",
882
1127
  Method: method,
883
1128
  };
@@ -948,11 +1193,22 @@ export class ChatGPTParser extends ChatParser {
948
1193
  };
949
1194
  }
950
1195
 
951
- const apiMessages = linearize(
952
- result.data.mapping,
953
- includeImages,
954
- result.data.current_node,
955
- );
1196
+ // The backend returns either the legacy `mapping` tree or the newer
1197
+ // `messages` array shape — support both.
1198
+ let apiMessages = [];
1199
+ if (result.data.mapping) {
1200
+ apiMessages = linearize(
1201
+ result.data.mapping,
1202
+ includeImages,
1203
+ result.data.current_node,
1204
+ );
1205
+ }
1206
+ if (apiMessages.length === 0 && Array.isArray(result.data.messages)) {
1207
+ apiMessages = linearizeMessagesArray(
1208
+ result.data.messages,
1209
+ includeImages,
1210
+ );
1211
+ }
956
1212
  if (apiMessages.length > 0) {
957
1213
  return this.formatApiResult(
958
1214
  result.data,
@@ -84,16 +84,30 @@ if (!window.__chatgptHelperInjected) {
84
84
  let images = {};
85
85
  if (includeImages) {
86
86
  const fileIds = new Set();
87
- for (const node of Object.values(data.mapping)) {
87
+ const collectImagePointer = (part) => {
88
+ if (
89
+ part &&
90
+ part.content_type === "image_asset_pointer" &&
91
+ part.asset_pointer
92
+ ) {
93
+ fileIds.add(part.asset_pointer.split("://")[1]);
94
+ }
95
+ };
96
+ // Legacy `mapping` tree shape…
97
+ for (const node of Object.values(data.mapping || {})) {
88
98
  const msg = node.message;
89
99
  if (msg && msg.content && Array.isArray(msg.content.parts)) {
90
100
  for (const part of msg.content.parts) {
91
- if (
92
- part &&
93
- part.content_type === "image_asset_pointer" &&
94
- part.asset_pointer
95
- ) {
96
- fileIds.add(part.asset_pointer.split("://")[1]);
101
+ collectImagePointer(part);
102
+ }
103
+ }
104
+ }
105
+ // …and the newer `messages` array shape.
106
+ if (Array.isArray(data.messages)) {
107
+ for (const msg of data.messages) {
108
+ if (msg && msg.content && Array.isArray(msg.content.parts)) {
109
+ for (const part of msg.content.parts) {
110
+ collectImagePointer(part);
97
111
  }
98
112
  }
99
113
  }