pi-quiver 6.4.0 → 6.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +5 -0
- package/README.md +7 -7
- package/dist/bin/pi-quiver.js +184 -102
- package/extensions/doc_to_md.ts +1 -1
- package/lib/doc-to-md-core.ts +137 -87
- package/lib/doc-to-md-handle.ts +29 -11
- package/lib/doc-to-md-options.ts +3 -3
- package/package.json +1 -1
- package/scripts/doc_to_md.py +288 -6
package/CHANGELOG.md
CHANGED
|
@@ -8,6 +8,11 @@ Published to npm as `pi-quiver` (`pi install npm:pi-quiver`). Pushing a
|
|
|
8
8
|
via OIDC trusted publishing. The release helper at
|
|
9
9
|
`.agents/skills/release/scripts/release.sh` cuts the tag; CI publishes.
|
|
10
10
|
|
|
11
|
+
## v6.5.0 - 2026-09-29
|
|
12
|
+
|
|
13
|
+
- `doc_to_md` converts DOCX directly in the Python child (mammoth -> markdownify, python-docx text fallback) instead of LibreOffice -> PDF: heading styles survive as `#` headings, hyperlinks and footnotes are kept, pictures land under `images/` (#24). DOCX page semantics change: only author-inserted page breaks become `--- end of page.page_number=N ---` markers, `Page-Count` is suffixed `(explicit page breaks, not printed pages)` / `(no explicit page breaks)`, `pages` selects those segments and is rejected on a break-less file, and DOCX that falls back to LibreOffice (no DOCX-capable Python, or both DOCX engines failed) is marked degraded with `(LibreOffice pagination)` and rejects `pages`. LibreOffice is now optional for DOCX and still required for PPTX. Backend probe gains a `DOCX` line and pins `mammoth==1.13.0`, `markdownify==1.2.3`, `python-docx==1.2.0`; the managed venv moves to `doc-to-md-venv-v3` and the older venvs are removed after the first successful build.
|
|
14
|
+
- `doc_to_md` handle: the `Outline` gains a `p<N>` page column (the page whose marker closes each heading's segment) for every format that emits page markers, and prints `Outline: none` when the Markdown has no headings. DOCX `info` returns core properties and a heading TOC with segment pages, or `TOC: none (no heading styles found)`.
|
|
15
|
+
|
|
11
16
|
## v6.4.0 - 2026-09-28
|
|
12
17
|
|
|
13
18
|
- provider-stall-watchdog: per-model threshold overrides via a `models` map on `quiver.providerStallWatchdog` - glob keys like `lmstudio/*` override `firstEventMs`/`warningMs`/`recoveryMs` for matching models, so slow local servers get patience without raising the global defaults (#18).
|
package/README.md
CHANGED
|
@@ -65,7 +65,7 @@ A 300 KB changelog page never touches your context window - you get a preview an
|
|
|
65
65
|
| Extension | Tool | What it does |
|
|
66
66
|
| --- | --- | --- |
|
|
67
67
|
| `extensions/fetch.ts` | `fetch` | Retrieve URLs over HTTP(S). HTML -> Markdown (Readability extraction, Turndown conversion). Binary saved untouched to a temp file. GitHub issue/PR/repo/actions-run/actions-job URLs auto-route through `gh` (falls back to HTTP); failed runs/jobs include failed-step logs (best-effort, summary-only otherwise). Same size gate as `fetch`. Behavior lives in `lib/fetch-core.ts`; also exposed as the `pi-quiver fetch` CLI (see [Claude Code support](#claude-code-support)). |
|
|
68
|
-
| `extensions/doc_to_md.ts` | `doc_to_md` | Convert a local PDF/DOCX/PPTX/XLSX/XLS to a Markdown bundle on disk (`<stem>.md` + `images/` and spreadsheet `sheets/`) and return a handle (paths, page count, outline, diagnostics) - never inline Markdown. `info` mode inspects first; `pages` selects 1-based pages; every page ends with `--- end of page.page_number=N ---`. Tiers: pymupdf4llm -> PyMuPDF text (degraded) -> unpdf worker (no Python only). Excel -> sheet inventory, full CSVs, bounded previews, and optional rendered views. Settings under `quiver.docToMd`. |
|
|
68
|
+
| `extensions/doc_to_md.ts` | `doc_to_md` | Convert a local PDF/DOCX/PPTX/XLSX/XLS to a Markdown bundle on disk (`<stem>.md` + `images/` and spreadsheet `sheets/`) and return a handle (paths, page count, outline, diagnostics) - never inline Markdown. `info` mode inspects first; `pages` selects 1-based pages (DOCX: explicit-page-break segments); every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Tiers: pymupdf4llm -> PyMuPDF text (degraded) -> unpdf worker (no Python only). Excel -> sheet inventory, full CSVs, bounded previews, and optional rendered views. DOCX converts directly (mammoth, python-docx fallback); LibreOffice pagination marks the degraded route. The Outline lists `L<line>` and `p<page>` per heading. Settings under `quiver.docToMd`. |
|
|
69
69
|
| `extensions/session-name.ts` | `/session-name` | Manual + opt-in automatic session naming, naming rules and deny list, long-session revisits, and Ghostty/Herdr tab rename. OFF by default. |
|
|
70
70
|
| `extensions/sword-header.ts` | `/builtin-header` | Themed ASCII startup header replacing pi's default logo. OFF by default. |
|
|
71
71
|
| `extensions/fast-mode.ts` | `/fast` | Inject Anthropic fast-mode (`speed: "fast"` + `anthropic-beta: fast-mode-2026-02-01`) into every Claude Opus 4.8 / Opus 5 request, any thinking level. `--fast` flag + `/fast [on\|off\|status]`. OFF by default. |
|
|
@@ -133,8 +133,8 @@ The npm package's bundled JS deps install automatically on `pi install`. A few *
|
|
|
133
133
|
| Prerequisite | Needed by | If absent |
|
|
134
134
|
| --- | --- | --- |
|
|
135
135
|
| `gh` (GitHub CLI, installed + `gh auth login`) | `fetch` GitHub issue/PR/repo/actions-run/actions-job routing | Falls back to an HTTP fetch of the rendered page (private repos hit a login wall). |
|
|
136
|
-
| `uv` (+ managed Python 3.14, fetched on first use) | `doc_to_md` high-fidelity PDF and Excel conversion (preferred route), with `pymupdf4llm`, `openpyxl`, `xlrd`, and `
|
|
137
|
-
| LibreOffice (`soffice` on `PATH`) | `doc_to_md`
|
|
136
|
+
| `uv` (+ managed Python 3.14, fetched on first use) | `doc_to_md` high-fidelity PDF and Excel conversion and DOCX (preferred route), with `pymupdf4llm`, `openpyxl`, `xlrd`, `pillow`, `mammoth`, `markdownify`, and `python-docx` | Falls back to a system Python >= 3.12 or one-time managed venv; PDF degrades to `unpdf` only when no capable Python exists. Excel requires the Python backend (no JS fallback). |
|
|
137
|
+
| LibreOffice (`soffice` on `PATH`) | `doc_to_md` PPTX conversion, the DOCX fallback route, and Excel rendered views | PPTX errors with a remedy; DOCX converts directly via the Python backend (LibreOffice fills in when that backend lacks the DOCX packages or its DOCX child exits 1); Excel omits rendered views. |
|
|
138
138
|
|
|
139
139
|
None is a hard install-time dependency of the package; they are tools you provide in the environment where pi runs.
|
|
140
140
|
|
|
@@ -294,9 +294,9 @@ Each setting can also be overridden per-process via `PI_QUIVER_SLACK_ENABLED`, `
|
|
|
294
294
|
|
|
295
295
|
| Key | Default | Meaning |
|
|
296
296
|
|---|---|---|
|
|
297
|
-
| `primaryTimeoutMs` | `60000` | pymupdf4llm tier and unpdf tier deadline. |
|
|
298
|
-
| `fallbackTimeoutMs` | `30000` | PyMuPDF text tier, PDF info, and Excel rendered-view rasterization deadline. |
|
|
299
|
-
| `sofficeTimeoutMs` | `120000` |
|
|
297
|
+
| `primaryTimeoutMs` | `60000` | pymupdf4llm tier, DOCX child (`docx` mode), and unpdf tier deadline. |
|
|
298
|
+
| `fallbackTimeoutMs` | `30000` | PyMuPDF text tier (including the DOCX LibreOffice fallback), PDF and DOCX info, and Excel rendered-view rasterization deadline. |
|
|
299
|
+
| `sofficeTimeoutMs` | `120000` | PPTX -> PDF, the DOCX LibreOffice fallback, and the Excel rendered-view export deadline. |
|
|
300
300
|
| `excelTimeoutMs` | `60000` | Excel child and Excel info deadline. |
|
|
301
301
|
| `warmTimeoutMs` | `120000` | Absolute first-call backend discovery/bootstrap deadline. |
|
|
302
302
|
| `pymupdfVersion` | `1.27.2.3` | pymupdf4llm pin, minimum `1.27.0`. |
|
|
@@ -309,7 +309,7 @@ A bundle is `<outputDir>/<stem>.md` plus `<outputDir>/images/` and, for spreadsh
|
|
|
309
309
|
|
|
310
310
|
Excel needs a Python backend with openpyxl, xlrd and pillow - otherwise the call fails with `Remedy: install uv, or pip install openpyxl xlrd pillow`. The Markdown opens with a `## Sheets` table listing every sheet in workbook order (0-based `#`, `worksheet`/`chartsheet`, size, hidden, chart and image counts, rendered view, CSV link for non-empty worksheets), then one section per sheet: a `Data:` line linking the full-content CSV under `sheets/` for non-empty worksheets, chart metadata from the workbook model, embedded images, an optional rendered view, a preview of at most the first 100 rows x 50 columns, and - only when the preview is truncated - a `Columns:` profile (type, non-empty count, min/max, distinct up to 50). Sizes are the extent of non-empty cells (the `info` handle reports the raw worksheet dimensions instead, which may be larger). Rendered views (`images/<stem>-s<idx>.<fmt>`) are produced for sheets carrying charts or images when LibreOffice is on `PATH`: the workbook is exported one PDF page per sheet and rasterized under a 16 Mpx budget. Any LibreOffice or rasterization failure degrades to `Rendered view: unavailable (<reason>)` and a handle note; it never fails the conversion. Workbooks whose chartsheet drawings carry a zero-size anchor (openpyxl-authored files; Excel-authored files are unaffected) render as a degenerate page and are reported as such. `.xls` gets the inventory, CSVs and previews but no visual detection or rendering.
|
|
311
311
|
|
|
312
|
-
Worst-case wall time
|
|
312
|
+
Worst-case wall time: PDF `warmTimeoutMs (first call) + primaryTimeoutMs + fallbackTimeoutMs`; PPTX adds `sofficeTimeoutMs`; DOCX on the Python path `warmTimeoutMs + primaryTimeoutMs` (success or a terminal child failure), DOCX child exit 1 then LibreOffice `warmTimeoutMs + primaryTimeoutMs + sofficeTimeoutMs + primaryTimeoutMs + fallbackTimeoutMs`, DOCX without a DOCX-capable backend `warmTimeoutMs + sofficeTimeoutMs + primaryTimeoutMs + fallbackTimeoutMs`; Excel `warmTimeoutMs + excelTimeoutMs + sofficeTimeoutMs + fallbackTimeoutMs`. Add `KILL_GRACE_MS` (2000 ms) per kill. There is no cap on image count, image bytes, cell count or workbook memory - deliberately; the per-tier timeouts, the rendered-view pixel budget and `maxOutputBytes` are the bounds.
|
|
313
313
|
|
|
314
314
|
### Migrating from flat keys
|
|
315
315
|
|
package/dist/bin/pi-quiver.js
CHANGED
|
@@ -517,7 +517,7 @@ async function fetchUrl(opts) {
|
|
|
517
517
|
}
|
|
518
518
|
|
|
519
519
|
// lib/doc-to-md-core.ts
|
|
520
|
-
import { existsSync as existsSync2, mkdtempSync, renameSync as renameSync2, rmSync as rmSync2, statSync as statSync2 } from "node:fs";
|
|
520
|
+
import { existsSync as existsSync2, mkdtempSync, readdirSync as readdirSync2, renameSync as renameSync2, rmSync as rmSync2, statSync as statSync2 } from "node:fs";
|
|
521
521
|
import { spawn } from "node:child_process";
|
|
522
522
|
import { homedir, tmpdir as tmpdir3 } from "node:os";
|
|
523
523
|
import { basename, dirname, extname as extname3, join as join3, resolve as resolve2 } from "node:path";
|
|
@@ -697,11 +697,19 @@ function compactRanges(nums, maxEntries = 20) {
|
|
|
697
697
|
if (parts.length <= maxEntries) return parts.join(", ");
|
|
698
698
|
return `${parts.slice(0, maxEntries).join(", ")} (+${parts.length - maxEntries} more)`;
|
|
699
699
|
}
|
|
700
|
+
var PAGE_MARKER_RE = /^--- end of page\.page_number=(\d+) ---$/;
|
|
700
701
|
function scanOutline(md, max) {
|
|
701
702
|
const entries = [];
|
|
703
|
+
let pending = [];
|
|
702
704
|
let total = 0, inFence = false;
|
|
703
705
|
md.split("\n").forEach((raw, i) => {
|
|
704
706
|
const line = raw.replace(/\r$/, "");
|
|
707
|
+
const marker = line.match(PAGE_MARKER_RE);
|
|
708
|
+
if (marker) {
|
|
709
|
+
for (const e of pending) e.page = Number(marker[1]);
|
|
710
|
+
pending = [];
|
|
711
|
+
return;
|
|
712
|
+
}
|
|
705
713
|
if (/^\s*(```|~~~)/.test(line)) {
|
|
706
714
|
inFence = !inFence;
|
|
707
715
|
return;
|
|
@@ -710,23 +718,36 @@ function scanOutline(md, max) {
|
|
|
710
718
|
const m = line.match(/^(#{1,6}) (.*)$/);
|
|
711
719
|
if (!m) return;
|
|
712
720
|
total++;
|
|
713
|
-
if (entries.length < max)
|
|
721
|
+
if (entries.length < max) {
|
|
722
|
+
const e = { line: i + 1, level: m[1].length, title: trunc(m[2].trim(), TITLE_MAX), page: null };
|
|
723
|
+
entries.push(e);
|
|
724
|
+
pending.push(e);
|
|
725
|
+
}
|
|
714
726
|
});
|
|
715
727
|
return { entries, total };
|
|
716
728
|
}
|
|
717
729
|
function outlineLines(entries, total) {
|
|
730
|
+
if (total === 0) return ["Outline: none"];
|
|
718
731
|
if (entries.length === 0) return [];
|
|
719
|
-
const
|
|
720
|
-
const
|
|
732
|
+
const lineWidth = Math.max(3, ...entries.map((e) => `L${e.line}`.length)) + 2;
|
|
733
|
+
const paged = entries.filter((e) => e.page !== null);
|
|
734
|
+
const pageWidth = paged.length ? Math.max(...paged.map((e) => `p${e.page}`.length)) + 2 : 0;
|
|
735
|
+
const out = ["Outline:", ...entries.map((e) => ` ${`L${e.line}`.padEnd(lineWidth)}${pageWidth ? (e.page === null ? "" : `p${e.page}`).padEnd(pageWidth) : ""}${"#".repeat(e.level)} ${trunc(e.title, TITLE_MAX)}`)];
|
|
721
736
|
if (total > entries.length) out.push(` (+${total - entries.length} more)`);
|
|
722
737
|
return out;
|
|
723
738
|
}
|
|
739
|
+
function pageCountLabel(h) {
|
|
740
|
+
const n = h.pageCount ?? "?";
|
|
741
|
+
if (h.type !== "docx") return String(n);
|
|
742
|
+
if (h.tier === "docx") return `${n} (${(h.explicitBreaks ?? 0) > 0 ? "explicit page breaks, not printed pages" : "no explicit page breaks"})`;
|
|
743
|
+
return `${n} (LibreOffice pagination)`;
|
|
744
|
+
}
|
|
724
745
|
function formatHandle(h) {
|
|
725
746
|
const lines = [`Saved-To: ${h.savedTo}`];
|
|
726
747
|
if (h.imagesDir && h.imageCount > 0) lines.push(`Images-Dir: ${h.imagesDir}`);
|
|
727
748
|
if (h.sheetsDir) lines.push(`Sheets-Dir: ${h.sheetsDir}`);
|
|
728
749
|
lines.push(`Type: ${h.type} Engine: ${h.engine} Tier: ${h.tier}`);
|
|
729
|
-
lines.push(`Page-Count: ${h
|
|
750
|
+
lines.push(`Page-Count: ${pageCountLabel(h)} Pages: ${h.pages ? compactRanges(h.pages) : "all"} Images: ${h.imageCount} Size: ${formatSize2(h.bytes)} / ${h.lines} lines`);
|
|
730
751
|
if (h.degraded) lines.push(`Degraded: ${h.degraded}`);
|
|
731
752
|
if (h.fallbackReason) lines.push(`Fallback-Reason: ${h.fallbackReason}`);
|
|
732
753
|
const fe = [];
|
|
@@ -752,9 +773,9 @@ function formatInfoHandle(i, max) {
|
|
|
752
773
|
const meta = Object.entries(i.metadata).filter(([, v]) => v).map(([k, v]) => `${k[0].toUpperCase()}${k.slice(1)}: ${trunc(v, META_MAX_CHARS)}`);
|
|
753
774
|
if (meta.length) lines.push(meta.join(" "));
|
|
754
775
|
if (i.toc.length) {
|
|
755
|
-
lines.push("TOC:", ...i.toc.slice(0, max).map((t) => ` L${t.level} ${trunc(t.title, TITLE_MAX)} (p${t.page})`));
|
|
776
|
+
lines.push("TOC:", ...i.toc.slice(0, max).map((t) => ` L${t.level} ${trunc(t.title, TITLE_MAX)} (p${t.page ?? "?"})`));
|
|
756
777
|
if (i.tocTotal > max) lines.push(` (+${i.tocTotal - max} more)`);
|
|
757
|
-
}
|
|
778
|
+
} else if (i.type === "docx") lines.push("TOC: none (no heading styles found)");
|
|
758
779
|
return lines.join("\n");
|
|
759
780
|
}
|
|
760
781
|
|
|
@@ -767,11 +788,11 @@ var VERSION_RE = /^\d+(\.\d+)*$/;
|
|
|
767
788
|
var DOC_TO_MD_OPTIONS = [
|
|
768
789
|
{ key: "path", flag: null, type: "string", default: null, settable: false, help: "Local .pdf .docx .pptx .xlsx .xls file" },
|
|
769
790
|
{ key: "info", flag: "--info", type: "bool", default: false, settable: false, help: "Inspect only (page count, metadata, TOC or sheet inventory); no bundle" },
|
|
770
|
-
{ key: "pages", flag: "--pages", type: "pages", default: null, settable: false, help: 'Inclusive 1-based pages, e.g. "12-15" or "3,7,10-12" (PDF/DOCX/PPTX only); default all' },
|
|
791
|
+
{ key: "pages", flag: "--pages", type: "pages", default: null, settable: false, help: 'Inclusive 1-based pages, e.g. "12-15" or "3,7,10-12" (PDF/DOCX/PPTX only); default all. DOCX: selects explicit-page-break segments; rejected when the file has none' },
|
|
771
792
|
{ key: "outputDir", flag: "--output-dir", type: "string", default: null, settable: false, help: "Bundle root for <stem>.md + images/; default a per-call temp dir" },
|
|
772
793
|
{ key: "overwrite", flag: "--overwrite", type: "bool", default: false, settable: false, help: "Replace an existing completed <stem>.md bundle" },
|
|
773
|
-
{ key: "primaryTimeoutMs", flag: "--primary-timeout", type: "int", default: 6e4, settable: true, env: "PI_DOC_TO_MD_CONVERT_TIMEOUT_MS", help: "pymupdf4llm tier; also the unpdf tier" },
|
|
774
|
-
{ key: "fallbackTimeoutMs", flag: "--fallback-timeout", type: "int", default: 3e4, settable: true, help: "PyMuPDF get_text tier; also
|
|
794
|
+
{ key: "primaryTimeoutMs", flag: "--primary-timeout", type: "int", default: 6e4, settable: true, env: "PI_DOC_TO_MD_CONVERT_TIMEOUT_MS", help: "pymupdf4llm tier and DOCX child (docx mode); also the unpdf tier" },
|
|
795
|
+
{ key: "fallbackTimeoutMs", flag: "--fallback-timeout", type: "int", default: 3e4, settable: true, help: "PyMuPDF get_text tier (including DOCX LibreOffice fallback); also PDF and DOCX info and Excel rendered views" },
|
|
775
796
|
{ key: "sofficeTimeoutMs", flag: "--soffice-timeout", type: "int", default: 12e4, settable: true, env: "PI_DOC_TO_MD_SOFFICE_TIMEOUT_MS", help: "DOCX/PPTX -> PDF via LibreOffice; also Excel rendered views" },
|
|
776
797
|
{ key: "excelTimeoutMs", flag: "--excel-timeout", type: "int", default: 6e4, settable: true, help: "Excel child (both openpyxl loads); also info on Excel" },
|
|
777
798
|
{ key: "warmTimeoutMs", flag: "--warm-timeout", type: "int", default: 12e4, settable: true, env: "PI_DOC_TO_MD_WARM_TIMEOUT_MS", help: "Absolute backend discovery/bootstrap deadline (first call per process)" },
|
|
@@ -892,21 +913,22 @@ function renderHelp() {
|
|
|
892
913
|
}
|
|
893
914
|
|
|
894
915
|
// lib/doc-to-md-core.ts
|
|
895
|
-
var PACKAGE_PINS = { pymupdf4llm: TUNABLE_DEFAULTS.pymupdfVersion, openpyxl: "3.1.5", xlrd: "2.0.2", pillow: "12.3.0" };
|
|
916
|
+
var PACKAGE_PINS = { pymupdf4llm: TUNABLE_DEFAULTS.pymupdfVersion, openpyxl: "3.1.5", xlrd: "2.0.2", pillow: "12.3.0", mammoth: "1.13.0", markdownify: "1.2.3", "python-docx": "1.2.0" };
|
|
896
917
|
var KILL_GRACE_MS = 2e3;
|
|
897
|
-
var VENV_DIR_NAME = "doc-to-md-venv-
|
|
898
|
-
var
|
|
918
|
+
var VENV_DIR_NAME = "doc-to-md-venv-v3";
|
|
919
|
+
var LEGACY_VENV_DIR_NAMES = ["pymupdf-venv", "doc-to-md-venv-v2"];
|
|
899
920
|
var STDERR_CAP = 1e6;
|
|
900
921
|
var OUTPUT_MAX_BYTES = 2e7;
|
|
901
922
|
var EXCEL_PDF_FILTER = 'pdf:calc_pdf_Export:{"SinglePageSheets":{"type":"boolean","value":"true"}}';
|
|
923
|
+
var pinSpecs = (cfg) => [`pymupdf4llm==${cfg.pymupdfVersion}`, `openpyxl==${PACKAGE_PINS.openpyxl}`, `xlrd==${PACKAGE_PINS.xlrd}`, `pillow==${PACKAGE_PINS.pillow}`, `mammoth==${PACKAGE_PINS.mammoth}`, `markdownify==${PACKAGE_PINS.markdownify}`, `python-docx==${PACKAGE_PINS["python-docx"]}`];
|
|
902
924
|
function withArgs(cfg) {
|
|
903
|
-
return
|
|
925
|
+
return pinSpecs(cfg).flatMap((spec) => ["--with", spec]);
|
|
904
926
|
}
|
|
905
927
|
function pipInstallArgs(cfg) {
|
|
906
|
-
return ["-m", "pip", "install",
|
|
928
|
+
return ["-m", "pip", "install", ...pinSpecs(cfg)];
|
|
907
929
|
}
|
|
908
930
|
function warmArgs(cfg) {
|
|
909
|
-
return ["run", ...withArgs(cfg), "--python", "3.14", "python", "-c", "import pymupdf4llm, openpyxl, xlrd, PIL"];
|
|
931
|
+
return ["run", ...withArgs(cfg), "--python", "3.14", "python", "-c", "import pymupdf4llm, openpyxl, xlrd, PIL, mammoth, markdownify, docx"];
|
|
910
932
|
}
|
|
911
933
|
function uvChildArgs(cfg, script, mode) {
|
|
912
934
|
return ["run", ...withArgs(cfg), "--python", "3.14", "python", script, mode];
|
|
@@ -1058,6 +1080,11 @@ try:
|
|
|
1058
1080
|
print("XLSX", "yes")
|
|
1059
1081
|
except Exception:
|
|
1060
1082
|
print("XLSX", "no")
|
|
1083
|
+
try:
|
|
1084
|
+
import mammoth, markdownify, docx
|
|
1085
|
+
print("DOCX", "yes")
|
|
1086
|
+
except Exception:
|
|
1087
|
+
print("DOCX", "no")
|
|
1061
1088
|
`;
|
|
1062
1089
|
var PROBE_TIMEOUT_MS = 5e3;
|
|
1063
1090
|
function probeArgs() {
|
|
@@ -1065,8 +1092,8 @@ function probeArgs() {
|
|
|
1065
1092
|
}
|
|
1066
1093
|
var PYTHON_CANDIDATES = ["python3", "python"];
|
|
1067
1094
|
function parseProbeOutput(stdout) {
|
|
1068
|
-
const m = stdout.match(/^PY (\d+) (\d+)\r?\nPDF (yes|no)\r?\nXLSX (yes|no)\s*$/);
|
|
1069
|
-
return m ? { major: Number(m[1]), minor: Number(m[2]), pdf: m[3] === "yes", xlsx: m[4] === "yes" } : null;
|
|
1095
|
+
const m = stdout.match(/^PY (\d+) (\d+)\r?\nPDF (yes|no)\r?\nXLSX (yes|no)\r?\nDOCX (yes|no)\s*$/);
|
|
1096
|
+
return m ? { major: Number(m[1]), minor: Number(m[2]), pdf: m[3] === "yes", xlsx: m[4] === "yes", docx: m[5] === "yes" } : null;
|
|
1070
1097
|
}
|
|
1071
1098
|
function meetsFloor(p) {
|
|
1072
1099
|
return p.major > 3 || p.major === 3 && p.minor >= 12;
|
|
@@ -1091,7 +1118,7 @@ async function resolveBackend(cfg, deps, signal) {
|
|
|
1091
1118
|
};
|
|
1092
1119
|
const warm = await deps.run("uv", warmArgs(cfg), { timeoutMs: left(), capBytes: OUTPUT_MAX_BYTES, env: deps.env, signal });
|
|
1093
1120
|
if (signal?.aborted) throw new Error("aborted");
|
|
1094
|
-
if (warm.code === 0 && !warm.timedOut) return { kind: "uv", pdf: true, xlsx: true };
|
|
1121
|
+
if (warm.code === 0 && !warm.timedOut) return { kind: "uv", pdf: true, xlsx: true, docx: true };
|
|
1095
1122
|
const uvAbsent = warm.code === null && !warm.timedOut;
|
|
1096
1123
|
const isDeadline = (result) => result !== null && "kind" in result;
|
|
1097
1124
|
const probe = async (exe, tmp) => {
|
|
@@ -1107,18 +1134,19 @@ async function resolveBackend(cfg, deps, signal) {
|
|
|
1107
1134
|
const p = await probe(exe);
|
|
1108
1135
|
if (isDeadline(p)) return p;
|
|
1109
1136
|
if (!p || !meetsFloor(p)) continue;
|
|
1110
|
-
if (p.pdf) return { kind: "python", exe, pdf: true, xlsx: p.xlsx };
|
|
1137
|
+
if (p.pdf) return { kind: "python", exe, pdf: true, xlsx: p.xlsx, docx: p.docx };
|
|
1111
1138
|
eligible ??= { exe, version: `${p.major}.${p.minor}` };
|
|
1112
1139
|
}
|
|
1140
|
+
const healthy = (p) => meetsFloor(p) && p.pdf && p.xlsx && p.docx;
|
|
1113
1141
|
const venvDir = join3(deps.cacheRoot, VENV_DIR_NAME);
|
|
1114
1142
|
const venvExe = venvPython(venvDir, deps.platform);
|
|
1115
1143
|
const cached = await probe(venvExe);
|
|
1116
1144
|
if (isDeadline(cached)) return cached;
|
|
1117
|
-
if (cached &&
|
|
1145
|
+
if (cached && healthy(cached)) return { kind: "venv", exe: venvExe, pdf: true, xlsx: true, docx: true };
|
|
1118
1146
|
if (eligible) {
|
|
1119
1147
|
const recheck = await probe(venvExe);
|
|
1120
1148
|
if (isDeadline(recheck)) return recheck;
|
|
1121
|
-
if (recheck &&
|
|
1149
|
+
if (recheck && healthy(recheck)) return { kind: "venv", exe: venvExe, pdf: true, xlsx: true, docx: true };
|
|
1122
1150
|
const tmp = `${venvDir}.tmp-${deps.pid}`;
|
|
1123
1151
|
const bootFail = (stderr) => {
|
|
1124
1152
|
deps.rmrf(tmp);
|
|
@@ -1143,19 +1171,19 @@ async function resolveBackend(cfg, deps, signal) {
|
|
|
1143
1171
|
}
|
|
1144
1172
|
};
|
|
1145
1173
|
if (publish()) {
|
|
1146
|
-
deps.rmrf(join3(deps.cacheRoot,
|
|
1147
|
-
return { kind: "venv", exe: venvExe, pdf: true, xlsx: true };
|
|
1174
|
+
for (const legacy of LEGACY_VENV_DIR_NAMES) deps.rmrf(join3(deps.cacheRoot, legacy));
|
|
1175
|
+
return { kind: "venv", exe: venvExe, pdf: true, xlsx: true, docx: true };
|
|
1148
1176
|
}
|
|
1149
1177
|
const winner = await probe(venvExe, tmp);
|
|
1150
1178
|
if (isDeadline(winner)) return winner;
|
|
1151
|
-
if (winner &&
|
|
1179
|
+
if (winner && healthy(winner)) {
|
|
1152
1180
|
deps.rmrf(tmp);
|
|
1153
|
-
return { kind: "venv", exe: venvExe, pdf: true, xlsx: true };
|
|
1181
|
+
return { kind: "venv", exe: venvExe, pdf: true, xlsx: true, docx: true };
|
|
1154
1182
|
}
|
|
1155
1183
|
deps.rmrf(venvDir);
|
|
1156
1184
|
if (publish()) {
|
|
1157
|
-
deps.rmrf(join3(deps.cacheRoot,
|
|
1158
|
-
return { kind: "venv", exe: venvExe, pdf: true, xlsx: true };
|
|
1185
|
+
for (const legacy of LEGACY_VENV_DIR_NAMES) deps.rmrf(join3(deps.cacheRoot, legacy));
|
|
1186
|
+
return { kind: "venv", exe: venvExe, pdf: true, xlsx: true, docx: true };
|
|
1159
1187
|
}
|
|
1160
1188
|
return bootFail("rename after competing bootstrap");
|
|
1161
1189
|
}
|
|
@@ -1213,15 +1241,23 @@ async function tryConvertOffice(sofficeTimeoutMs, src, signal, run = runCapped,
|
|
|
1213
1241
|
throw e;
|
|
1214
1242
|
}
|
|
1215
1243
|
}
|
|
1216
|
-
|
|
1217
|
-
|
|
1218
|
-
if (r.
|
|
1219
|
-
|
|
1220
|
-
if (r.kind === "no-pdf") throw new Error("LibreOffice (soffice) ran but produced no usable PDF for this file. Ensure LibreOffice can open the document, or convert it to PDF manually first.");
|
|
1221
|
-
throw new Error(`soffice failed (code=${r.code} timedOut=${r.timedOut}): ${r.stderr}`);
|
|
1244
|
+
function officeFailure(r) {
|
|
1245
|
+
if (r.kind === "missing") return new Error("LibreOffice (soffice) is required to convert .docx/.pptx but was not found on PATH. Install LibreOffice or convert the file to PDF first.");
|
|
1246
|
+
if (r.kind === "no-pdf") return new Error("LibreOffice (soffice) ran but produced no usable PDF for this file. Ensure LibreOffice can open the document, or convert it to PDF manually first.");
|
|
1247
|
+
return new Error(`soffice failed (code=${r.code} timedOut=${r.timedOut}): ${r.stderr}`);
|
|
1222
1248
|
}
|
|
1223
1249
|
var DEGRADED_TEXT = "PyMuPDF text extraction - layout/tables not preserved";
|
|
1224
1250
|
var DEGRADED_UNPDF = "unpdf text extraction - structure not preserved";
|
|
1251
|
+
var DEGRADED_DOCX_TEXT = "python-docx text extraction - footnotes, hyperlinks, images not preserved";
|
|
1252
|
+
var DEGRADED_DOCX_OFFICE = "LibreOffice PDF route - heading styles and explicit page breaks not preserved; page numbers are LibreOffice pagination";
|
|
1253
|
+
var DOCX_PIP = "pip install mammoth markdownify python-docx";
|
|
1254
|
+
var DOCX_PAGES_OFFICE = `--pages on a DOCX needs the Python DOCX backend (explicit page-break segments); the LibreOffice route has none. Remedy: install uv, or ${DOCX_PIP}`;
|
|
1255
|
+
var docxRemedy = `Remedy: install uv, or ${DOCX_PIP} into a Python that already has pymupdf4llm`;
|
|
1256
|
+
var backendState = (b, missing) => b.kind === "none" ? b.reason : missing;
|
|
1257
|
+
var lacksDocx = (b) => b.kind === "none" || !b.docx;
|
|
1258
|
+
var clearStaging = (b) => {
|
|
1259
|
+
for (const f of readdirSync2(b.stagingDir)) rmSync2(join3(b.stagingDir, f), { recursive: true, force: true });
|
|
1260
|
+
};
|
|
1225
1261
|
var EXCEL_REMEDY = "Remedy: install uv, or pip install openpyxl xlrd pillow";
|
|
1226
1262
|
function resolveUnpdfWorker(candidates, exists) {
|
|
1227
1263
|
for (const candidate of candidates) if (exists(candidate)) return candidate;
|
|
@@ -1291,78 +1327,120 @@ async function convertDocument(o, signal, seams) {
|
|
|
1291
1327
|
let office = null;
|
|
1292
1328
|
try {
|
|
1293
1329
|
let pdfPath = inputPath;
|
|
1294
|
-
|
|
1295
|
-
office = await convertOffice(o.sofficeTimeoutMs, inputPath, signal);
|
|
1296
|
-
pdfPath = office.pdfPath;
|
|
1297
|
-
}
|
|
1298
|
-
const base = { path: pdfPath, pages: o.pages, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion };
|
|
1330
|
+
const base = { path: inputPath, pages: o.pages, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion };
|
|
1299
1331
|
let tier, engine, json, degraded = null, fallbackReason = null;
|
|
1332
|
+
let explicitBreaks = null;
|
|
1300
1333
|
let notes = [];
|
|
1301
|
-
|
|
1302
|
-
|
|
1303
|
-
|
|
1304
|
-
|
|
1305
|
-
|
|
1306
|
-
|
|
1334
|
+
let officeRoute = null;
|
|
1335
|
+
if (type === "docx" && !lacksDocx(backend)) {
|
|
1336
|
+
const d = await s.runTier("docx", base, b, signal, o.primaryTimeoutMs, backend);
|
|
1337
|
+
if (d.ok) {
|
|
1338
|
+
publishStaged(b);
|
|
1339
|
+
tier = "docx";
|
|
1340
|
+
engine = d.json.engine === "python-docx" ? "python-docx" : "mammoth";
|
|
1341
|
+
json = d.json;
|
|
1342
|
+
explicitBreaks = d.json.explicitBreaks ?? 0;
|
|
1343
|
+
if (d.json.degraded) {
|
|
1344
|
+
degraded = DEGRADED_DOCX_TEXT;
|
|
1345
|
+
fallbackReason = d.json.fallbackReason ?? null;
|
|
1346
|
+
}
|
|
1347
|
+
} else if ("userError" in d) throw new Error(d.userError);
|
|
1348
|
+
else if (d.reason === "exit 1") {
|
|
1349
|
+
officeRoute = `docx ${d.reason}${detailSuffix(d)}`;
|
|
1350
|
+
clearStaging(b);
|
|
1351
|
+
} else throw new Error(`Conversion failed: docx ${d.reason}${detailSuffix(d)}`);
|
|
1352
|
+
} else if (type === "docx") officeRoute = backend.kind === "none" ? backend.reason : "python backend lacks DOCX packages";
|
|
1353
|
+
if (type === "docx" && officeRoute !== null) {
|
|
1354
|
+
if (o.pages) throw new Error(officeRoute.startsWith("docx exit") ? `${DOCX_PAGES_OFFICE} (${officeRoute})` : DOCX_PAGES_OFFICE);
|
|
1355
|
+
const r = await s.office(o.sofficeTimeoutMs, inputPath, signal);
|
|
1356
|
+
if (!r.ok && r.kind === "missing") {
|
|
1357
|
+
if (officeRoute.startsWith("docx exit")) throw new Error(`Conversion failed: ${officeRoute}; LibreOffice (soffice) not found on PATH`);
|
|
1358
|
+
throw new Error(`DOCX conversion needs the Python DOCX packages or LibreOffice. Python backend: ${backendState(backend, "found without mammoth/markdownify/python-docx")}. ${docxRemedy}, or install LibreOffice (soffice)`);
|
|
1307
1359
|
}
|
|
1308
|
-
|
|
1309
|
-
|
|
1310
|
-
|
|
1311
|
-
|
|
1312
|
-
|
|
1313
|
-
|
|
1314
|
-
|
|
1315
|
-
|
|
1316
|
-
|
|
1317
|
-
|
|
1318
|
-
|
|
1319
|
-
|
|
1320
|
-
|
|
1321
|
-
|
|
1322
|
-
|
|
1323
|
-
|
|
1324
|
-
|
|
1325
|
-
|
|
1326
|
-
|
|
1327
|
-
|
|
1328
|
-
|
|
1360
|
+
if (!r.ok) throw officeRoute.startsWith("docx exit") ? new Error(`Conversion failed: ${officeRoute}; ${officeFailure(r).message}`) : officeFailure(r);
|
|
1361
|
+
office = r;
|
|
1362
|
+
pdfPath = r.pdfPath;
|
|
1363
|
+
}
|
|
1364
|
+
if (type === "pptx") {
|
|
1365
|
+
const r = await s.office(o.sofficeTimeoutMs, inputPath, signal);
|
|
1366
|
+
if (!r.ok && r.kind === "missing") throw new Error(`PPTX conversion needs LibreOffice (soffice); direct conversion is not available. Python backend: ${backendState(backend, "available")}. Remedy: install LibreOffice`);
|
|
1367
|
+
if (!r.ok) throw officeFailure(r);
|
|
1368
|
+
office = r;
|
|
1369
|
+
pdfPath = r.pdfPath;
|
|
1370
|
+
}
|
|
1371
|
+
const pdfBase = { ...base, path: pdfPath };
|
|
1372
|
+
if (json === void 0) {
|
|
1373
|
+
if (isExcel) {
|
|
1374
|
+
const r = await s.runTier("xlsx", base, b, signal, o.excelTimeoutMs, backend);
|
|
1375
|
+
if (!r.ok) {
|
|
1376
|
+
if ("userError" in r) throw new Error(r.userError);
|
|
1377
|
+
const remedy = r.reason.startsWith("timeout after") || r.reason === "output exceeded maxOutputBytes" ? ". Remedy: raise excelTimeoutMs" : "";
|
|
1378
|
+
throw new Error(`Excel conversion failed: ${r.reason}${detailSuffix(r)}${remedy}`);
|
|
1379
|
+
}
|
|
1380
|
+
publishSheetImages(b);
|
|
1381
|
+
publishSheetCsvs(b);
|
|
1382
|
+
tier = "excel";
|
|
1383
|
+
engine = type === "xls" ? "xlrd" : "openpyxl";
|
|
1384
|
+
json = r.json;
|
|
1385
|
+
notes = [...json.notes ?? []];
|
|
1386
|
+
const renderPages = json.renderPages ?? [];
|
|
1387
|
+
let skip = null;
|
|
1388
|
+
const perSheet = /* @__PURE__ */ new Map();
|
|
1389
|
+
if (renderPages.length) {
|
|
1390
|
+
const off = await s.office(o.sofficeTimeoutMs, inputPath, signal, runCapped, EXCEL_PDF_FILTER);
|
|
1391
|
+
if (!off.ok) skip = off.kind === "missing" ? "LibreOffice not found" : off.kind === "timeout" ? `soffice failed: timeout after ${o.sofficeTimeoutMs}ms` : off.kind === "exit" ? `soffice failed: exit ${off.code}` : "soffice produced no PDF";
|
|
1392
|
+
else {
|
|
1393
|
+
try {
|
|
1394
|
+
const rp = await s.runTier("render-pages", { path: off.pdfPath, sheetIndices: renderPages, expectedPages: json.sheetCount, imageDpi: o.imageDpi, imageFormat: o.imageFormat, stagingDir: b.stagingDir, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion }, b, signal, o.fallbackTimeoutMs, backend);
|
|
1395
|
+
if (!rp.ok) skip = `render failed: ${"userError" in rp ? rp.userError : rp.reason}`;
|
|
1396
|
+
else if (rp.json.ok === false) skip = rp.json.reason ?? "render failed";
|
|
1397
|
+
else {
|
|
1398
|
+
publishSheetImages(b);
|
|
1399
|
+
for (const f of rp.json.failed ?? []) perSheet.set(f.idx, f.reason);
|
|
1400
|
+
for (const d of rp.json.rendered ?? []) if (d.dpi < o.imageDpi) notes.push(`Rendered view s${d.idx}: rendered at ${d.dpi} dpi`);
|
|
1401
|
+
}
|
|
1402
|
+
} finally {
|
|
1403
|
+
off.cleanup();
|
|
1329
1404
|
}
|
|
1330
|
-
} finally {
|
|
1331
|
-
off.cleanup();
|
|
1332
1405
|
}
|
|
1406
|
+
if (skip) notes.push(`Rendered views skipped: ${skip}`);
|
|
1407
|
+
else if (perSheet.size) notes.push(`Rendered views: ${perSheet.size} of ${renderPages.length} unavailable`);
|
|
1408
|
+
}
|
|
1409
|
+
json = { ...json, markdown: reconcileRenderMarkers(json.markdown ?? "", renderPages, o.imageFormat, b.sourceMap, (idx) => perSheet.get(idx) ?? skip ?? "render failed") };
|
|
1410
|
+
} else if (backend.kind === "none") {
|
|
1411
|
+
const r = await s.runTier("pdf-text", pdfBase, b, signal, o.primaryTimeoutMs, backend);
|
|
1412
|
+
if (!r.ok) throw new Error("userError" in r ? r.userError : `Conversion failed: unpdf ${r.reason}${detailSuffix(r)}`);
|
|
1413
|
+
tier = "unpdf";
|
|
1414
|
+
engine = "unpdf";
|
|
1415
|
+
json = r.json;
|
|
1416
|
+
degraded = DEGRADED_UNPDF;
|
|
1417
|
+
} else {
|
|
1418
|
+
const p = await s.runTier("pdf-primary", pdfBase, b, signal, o.primaryTimeoutMs, backend);
|
|
1419
|
+
const kept = publishStaged(b);
|
|
1420
|
+
if (p.ok) {
|
|
1421
|
+
tier = "primary";
|
|
1422
|
+
engine = "pymupdf4llm";
|
|
1423
|
+
json = p.json;
|
|
1424
|
+
} else if ("userError" in p) throw new Error(p.userError);
|
|
1425
|
+
else {
|
|
1426
|
+
if (signal?.aborted) throw new Error("aborted");
|
|
1427
|
+
const keepPages = Object.fromEntries([...kept.entries()].map(([k, v]) => [String(k), v]));
|
|
1428
|
+
const f = await s.runTier("pdf-fallback", { ...pdfBase, keepPages }, b, signal, o.fallbackTimeoutMs, backend);
|
|
1429
|
+
publishStaged(b);
|
|
1430
|
+
if (!f.ok) throw new Error("userError" in f ? f.userError : `Conversion failed: primary ${p.reason}; fallback ${f.reason}${detailSuffix(f)}`);
|
|
1431
|
+
tier = "fallback";
|
|
1432
|
+
engine = "pymupdf-text";
|
|
1433
|
+
json = f.json;
|
|
1434
|
+
degraded = DEGRADED_TEXT;
|
|
1435
|
+
fallbackReason = `primary ${p.reason}`;
|
|
1333
1436
|
}
|
|
1334
|
-
if (skip) notes.push(`Rendered views skipped: ${skip}`);
|
|
1335
|
-
else if (perSheet.size) notes.push(`Rendered views: ${perSheet.size} of ${renderPages.length} unavailable`);
|
|
1336
|
-
}
|
|
1337
|
-
json = { ...json, markdown: reconcileRenderMarkers(json.markdown ?? "", renderPages, o.imageFormat, b.sourceMap, (idx) => perSheet.get(idx) ?? skip ?? "render failed") };
|
|
1338
|
-
} else if (backend.kind === "none") {
|
|
1339
|
-
const r = await s.runTier("pdf-text", base, b, signal, o.primaryTimeoutMs, backend);
|
|
1340
|
-
if (!r.ok) throw new Error("userError" in r ? r.userError : `Conversion failed: unpdf ${r.reason}${detailSuffix(r)}`);
|
|
1341
|
-
tier = "unpdf";
|
|
1342
|
-
engine = "unpdf";
|
|
1343
|
-
json = r.json;
|
|
1344
|
-
degraded = DEGRADED_UNPDF;
|
|
1345
|
-
} else {
|
|
1346
|
-
const p = await s.runTier("pdf-primary", base, b, signal, o.primaryTimeoutMs, backend);
|
|
1347
|
-
const kept = publishStaged(b);
|
|
1348
|
-
if (p.ok) {
|
|
1349
|
-
tier = "primary";
|
|
1350
|
-
engine = "pymupdf4llm";
|
|
1351
|
-
json = p.json;
|
|
1352
|
-
} else if ("userError" in p) throw new Error(p.userError);
|
|
1353
|
-
else {
|
|
1354
|
-
if (signal?.aborted) throw new Error("aborted");
|
|
1355
|
-
const keepPages = Object.fromEntries([...kept.entries()].map(([k, v]) => [String(k), v]));
|
|
1356
|
-
const f = await s.runTier("pdf-fallback", { ...base, keepPages }, b, signal, o.fallbackTimeoutMs, backend);
|
|
1357
|
-
publishStaged(b);
|
|
1358
|
-
if (!f.ok) throw new Error("userError" in f ? f.userError : `Conversion failed: primary ${p.reason}; fallback ${f.reason}${detailSuffix(f)}`);
|
|
1359
|
-
tier = "fallback";
|
|
1360
|
-
engine = "pymupdf-text";
|
|
1361
|
-
json = f.json;
|
|
1362
|
-
degraded = DEGRADED_TEXT;
|
|
1363
|
-
fallbackReason = `primary ${p.reason}`;
|
|
1364
1437
|
}
|
|
1365
1438
|
}
|
|
1439
|
+
if (officeRoute !== null) {
|
|
1440
|
+
degraded = DEGRADED_DOCX_OFFICE;
|
|
1441
|
+
fallbackReason = fallbackReason ? `${officeRoute}; ${fallbackReason}` : officeRoute;
|
|
1442
|
+
}
|
|
1443
|
+
if (tier === void 0 || engine === void 0 || json === void 0) throw new Error("internal: no tier produced output");
|
|
1366
1444
|
if (!isExcel) notes = json.notes ?? [];
|
|
1367
1445
|
const body = rewriteLinks(json.markdown ?? "", b.sourceMap);
|
|
1368
1446
|
validateImageLinks(body, b.manifest, b.csvManifest);
|
|
@@ -1377,7 +1455,7 @@ async function convertDocument(o, signal, seams) {
|
|
|
1377
1455
|
` : "") + body;
|
|
1378
1456
|
commitBundle(b, markdown);
|
|
1379
1457
|
const outline = scanOutline(markdown, o.outlineMaxEntries);
|
|
1380
|
-
const details = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total };
|
|
1458
|
+
const details = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total };
|
|
1381
1459
|
return { output: formatHandle(details), details };
|
|
1382
1460
|
} catch (e) {
|
|
1383
1461
|
abortBundle(b);
|
|
@@ -1398,9 +1476,13 @@ async function inspectDocument(o, signal, seams) {
|
|
|
1398
1476
|
let office = null;
|
|
1399
1477
|
try {
|
|
1400
1478
|
let path = inputPath;
|
|
1401
|
-
if (type === "docx"
|
|
1402
|
-
|
|
1403
|
-
|
|
1479
|
+
if (type === "docx") {
|
|
1480
|
+
if (lacksDocx(backend)) throw new Error(`DOCX inspection needs the Python DOCX packages. Python backend: ${backendState(backend, "found without mammoth/markdownify/python-docx")}. ${docxRemedy}`);
|
|
1481
|
+
} else if (type === "pptx") {
|
|
1482
|
+
const r2 = await s.office(o.sofficeTimeoutMs, inputPath, signal);
|
|
1483
|
+
if (!r2.ok) throw officeFailure(r2);
|
|
1484
|
+
office = r2;
|
|
1485
|
+
path = r2.pdfPath;
|
|
1404
1486
|
}
|
|
1405
1487
|
const r = await s.runTier("info", { path, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion }, { stagingDir: "" }, signal, isExcel ? o.excelTimeoutMs : o.fallbackTimeoutMs, backend);
|
|
1406
1488
|
if (!r.ok) {
|
package/extensions/doc_to_md.ts
CHANGED
|
@@ -35,7 +35,7 @@ export default function docToMdExtension(pi: ExtensionAPI) {
|
|
|
35
35
|
name: "doc_to_md",
|
|
36
36
|
label: "Convert doc to Markdown bundle",
|
|
37
37
|
description:
|
|
38
|
-
"Convert a local PDF/DOCX/PPTX/XLSX/XLS to a Markdown bundle on disk and return a handle (Saved-To, Images-Dir, Page-Count, Outline, diagnostics) - the Markdown itself is never inlined; read the Saved-To file (offset/limit) for content. `info: true` returns page count, metadata and TOC (or the sheet inventory) without converting - use it to pick `pages`. `pages` selects inclusive 1-based pages (PDF/
|
|
38
|
+
"Convert a local PDF/DOCX/PPTX/XLSX/XLS to a Markdown bundle on disk and return a handle (Saved-To, Images-Dir, Page-Count, Outline, diagnostics) - the Markdown itself is never inlined; read the Saved-To file (offset/limit) for content. `info: true` returns page count, metadata and TOC (or the sheet inventory) without converting - use it to pick `pages`. `pages` selects inclusive 1-based pages (PDF/PPTX) or explicit-page-break segments (DOCX; rejected when the file has none); every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Images are always extracted into images/ with relative links. Primary engine pymupdf4llm, fallback PyMuPDF text (degraded, marked), pure-JS unpdf only when no Python backend exists. Excel yields a sheet inventory (worksheets + chartsheets), a full CSV per non-empty worksheet under sheets/, a bounded preview with formulas and cached values, merged/hidden disclosure, and rendered chart views when LibreOffice is available. DOCX converts directly (mammoth; python-docx text fallback) and keeps heading styles, so the handle Outline lists `L<line>` and `p<segment>` per heading; LibreOffice (soffice) is optional for DOCX (fallback route, degraded) and Excel rendered views, and required for PPTX. Use `outputDir` for a durable bundle; without it the bundle lands in a per-call temp dir. Input must be a local file path (use fetch first for URLs).",
|
|
39
39
|
promptSnippet: "Convert a local PDF/DOCX/PPTX/XLSX to a Markdown bundle (handle returned; read Saved-To)",
|
|
40
40
|
parameters: buildSchema(),
|
|
41
41
|
|