pi-quiver 6.4.0 → 6.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -8,6 +8,11 @@ Published to npm as `pi-quiver` (`pi install npm:pi-quiver`). Pushing a
8
8
  via OIDC trusted publishing. The release helper at
9
9
  `.agents/skills/release/scripts/release.sh` cuts the tag; CI publishes.
10
10
 
11
+ ## v6.5.0 - 2026-09-29
12
+
13
+ - `doc_to_md` converts DOCX directly in the Python child (mammoth -> markdownify, python-docx text fallback) instead of LibreOffice -> PDF: heading styles survive as `#` headings, hyperlinks and footnotes are kept, pictures land under `images/` (#24). DOCX page semantics change: only author-inserted page breaks become `--- end of page.page_number=N ---` markers, `Page-Count` is suffixed `(explicit page breaks, not printed pages)` / `(no explicit page breaks)`, `pages` selects those segments and is rejected on a break-less file, and DOCX that falls back to LibreOffice (no DOCX-capable Python, or both DOCX engines failed) is marked degraded with `(LibreOffice pagination)` and rejects `pages`. LibreOffice is now optional for DOCX and still required for PPTX. Backend probe gains a `DOCX` line and pins `mammoth==1.13.0`, `markdownify==1.2.3`, `python-docx==1.2.0`; the managed venv moves to `doc-to-md-venv-v3` and the older venvs are removed after the first successful build.
14
+ - `doc_to_md` handle: the `Outline` gains a `p<N>` page column (the page whose marker closes each heading's segment) for every format that emits page markers, and prints `Outline: none` when the Markdown has no headings. DOCX `info` returns core properties and a heading TOC with segment pages, or `TOC: none (no heading styles found)`.
15
+
11
16
  ## v6.4.0 - 2026-09-28
12
17
 
13
18
  - provider-stall-watchdog: per-model threshold overrides via a `models` map on `quiver.providerStallWatchdog` - glob keys like `lmstudio/*` override `firstEventMs`/`warningMs`/`recoveryMs` for matching models, so slow local servers get patience without raising the global defaults (#18).
package/README.md CHANGED
@@ -65,7 +65,7 @@ A 300 KB changelog page never touches your context window - you get a preview an
65
65
  | Extension | Tool | What it does |
66
66
  | --- | --- | --- |
67
67
  | `extensions/fetch.ts` | `fetch` | Retrieve URLs over HTTP(S). HTML -> Markdown (Readability extraction, Turndown conversion). Binary saved untouched to a temp file. GitHub issue/PR/repo/actions-run/actions-job URLs auto-route through `gh` (falls back to HTTP); failed runs/jobs include failed-step logs (best-effort, summary-only otherwise). Same size gate as `fetch`. Behavior lives in `lib/fetch-core.ts`; also exposed as the `pi-quiver fetch` CLI (see [Claude Code support](#claude-code-support)). |
68
- | `extensions/doc_to_md.ts` | `doc_to_md` | Convert a local PDF/DOCX/PPTX/XLSX/XLS to a Markdown bundle on disk (`<stem>.md` + `images/` and spreadsheet `sheets/`) and return a handle (paths, page count, outline, diagnostics) - never inline Markdown. `info` mode inspects first; `pages` selects 1-based pages; every page ends with `--- end of page.page_number=N ---`. Tiers: pymupdf4llm -> PyMuPDF text (degraded) -> unpdf worker (no Python only). Excel -> sheet inventory, full CSVs, bounded previews, and optional rendered views. Settings under `quiver.docToMd`. |
68
+ | `extensions/doc_to_md.ts` | `doc_to_md` | Convert a local PDF/DOCX/PPTX/XLSX/XLS to a Markdown bundle on disk (`<stem>.md` + `images/` and spreadsheet `sheets/`) and return a handle (paths, page count, outline, diagnostics) - never inline Markdown. `info` mode inspects first; `pages` selects 1-based pages (DOCX: explicit-page-break segments); every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Tiers: pymupdf4llm -> PyMuPDF text (degraded) -> unpdf worker (no Python only). Excel -> sheet inventory, full CSVs, bounded previews, and optional rendered views. DOCX converts directly (mammoth, python-docx fallback); LibreOffice pagination marks the degraded route. The Outline lists `L<line>` and `p<page>` per heading. Settings under `quiver.docToMd`. |
69
69
  | `extensions/session-name.ts` | `/session-name` | Manual + opt-in automatic session naming, naming rules and deny list, long-session revisits, and Ghostty/Herdr tab rename. OFF by default. |
70
70
  | `extensions/sword-header.ts` | `/builtin-header` | Themed ASCII startup header replacing pi's default logo. OFF by default. |
71
71
  | `extensions/fast-mode.ts` | `/fast` | Inject Anthropic fast-mode (`speed: "fast"` + `anthropic-beta: fast-mode-2026-02-01`) into every Claude Opus 4.8 / Opus 5 request, any thinking level. `--fast` flag + `/fast [on\|off\|status]`. OFF by default. |
@@ -133,8 +133,8 @@ The npm package's bundled JS deps install automatically on `pi install`. A few *
133
133
  | Prerequisite | Needed by | If absent |
134
134
  | --- | --- | --- |
135
135
  | `gh` (GitHub CLI, installed + `gh auth login`) | `fetch` GitHub issue/PR/repo/actions-run/actions-job routing | Falls back to an HTTP fetch of the rendered page (private repos hit a login wall). |
136
- | `uv` (+ managed Python 3.14, fetched on first use) | `doc_to_md` high-fidelity PDF and Excel conversion (preferred route), with `pymupdf4llm`, `openpyxl`, `xlrd`, and `pillow` | Falls back to a system Python >= 3.12 or one-time managed venv; PDF degrades to `unpdf` only when no capable Python exists. Excel requires the Python backend (no JS fallback). |
137
- | LibreOffice (`soffice` on `PATH`) | `doc_to_md` DOCX/PPTX conversion | Office inputs error (no JS fallback for office->PDF); PDFs unaffected. |
136
+ | `uv` (+ managed Python 3.14, fetched on first use) | `doc_to_md` high-fidelity PDF and Excel conversion and DOCX (preferred route), with `pymupdf4llm`, `openpyxl`, `xlrd`, `pillow`, `mammoth`, `markdownify`, and `python-docx` | Falls back to a system Python >= 3.12 or one-time managed venv; PDF degrades to `unpdf` only when no capable Python exists. Excel requires the Python backend (no JS fallback). |
137
+ | LibreOffice (`soffice` on `PATH`) | `doc_to_md` PPTX conversion, the DOCX fallback route, and Excel rendered views | PPTX errors with a remedy; DOCX converts directly via the Python backend (LibreOffice fills in when that backend lacks the DOCX packages or its DOCX child exits 1); Excel omits rendered views. |
138
138
 
139
139
  None is a hard install-time dependency of the package; they are tools you provide in the environment where pi runs.
140
140
 
@@ -294,9 +294,9 @@ Each setting can also be overridden per-process via `PI_QUIVER_SLACK_ENABLED`, `
294
294
 
295
295
  | Key | Default | Meaning |
296
296
  |---|---|---|
297
- | `primaryTimeoutMs` | `60000` | pymupdf4llm tier and unpdf tier deadline. |
298
- | `fallbackTimeoutMs` | `30000` | PyMuPDF text tier, PDF info, and Excel rendered-view rasterization deadline. |
299
- | `sofficeTimeoutMs` | `120000` | DOCX/PPTX -> PDF deadline; also the Excel rendered-view export. |
297
+ | `primaryTimeoutMs` | `60000` | pymupdf4llm tier, DOCX child (`docx` mode), and unpdf tier deadline. |
298
+ | `fallbackTimeoutMs` | `30000` | PyMuPDF text tier (including the DOCX LibreOffice fallback), PDF and DOCX info, and Excel rendered-view rasterization deadline. |
299
+ | `sofficeTimeoutMs` | `120000` | PPTX -> PDF, the DOCX LibreOffice fallback, and the Excel rendered-view export deadline. |
300
300
  | `excelTimeoutMs` | `60000` | Excel child and Excel info deadline. |
301
301
  | `warmTimeoutMs` | `120000` | Absolute first-call backend discovery/bootstrap deadline. |
302
302
  | `pymupdfVersion` | `1.27.2.3` | pymupdf4llm pin, minimum `1.27.0`. |
@@ -309,7 +309,7 @@ A bundle is `<outputDir>/<stem>.md` plus `<outputDir>/images/` and, for spreadsh
309
309
 
310
310
  Excel needs a Python backend with openpyxl, xlrd and pillow - otherwise the call fails with `Remedy: install uv, or pip install openpyxl xlrd pillow`. The Markdown opens with a `## Sheets` table listing every sheet in workbook order (0-based `#`, `worksheet`/`chartsheet`, size, hidden, chart and image counts, rendered view, CSV link for non-empty worksheets), then one section per sheet: a `Data:` line linking the full-content CSV under `sheets/` for non-empty worksheets, chart metadata from the workbook model, embedded images, an optional rendered view, a preview of at most the first 100 rows x 50 columns, and - only when the preview is truncated - a `Columns:` profile (type, non-empty count, min/max, distinct up to 50). Sizes are the extent of non-empty cells (the `info` handle reports the raw worksheet dimensions instead, which may be larger). Rendered views (`images/<stem>-s<idx>.<fmt>`) are produced for sheets carrying charts or images when LibreOffice is on `PATH`: the workbook is exported one PDF page per sheet and rasterized under a 16 Mpx budget. Any LibreOffice or rasterization failure degrades to `Rendered view: unavailable (<reason>)` and a handle note; it never fails the conversion. Workbooks whose chartsheet drawings carry a zero-size anchor (openpyxl-authored files; Excel-authored files are unaffected) render as a degenerate page and are reported as such. `.xls` gets the inventory, CSVs and previews but no visual detection or rendering.
311
311
 
312
- Worst-case wall time is `warmTimeoutMs (first call) + sofficeTimeoutMs (Office only) + primaryTimeoutMs + fallbackTimeoutMs + KILL_GRACE_MS x kills` (Excel: `warmTimeoutMs + excelTimeoutMs + sofficeTimeoutMs + fallbackTimeoutMs + 2 * KILL_GRACE_MS`). There is no cap on image count, image bytes, cell count or workbook memory - deliberately; the per-tier timeouts, the rendered-view pixel budget and `maxOutputBytes` are the bounds.
312
+ Worst-case wall time: PDF `warmTimeoutMs (first call) + primaryTimeoutMs + fallbackTimeoutMs`; PPTX adds `sofficeTimeoutMs`; DOCX on the Python path `warmTimeoutMs + primaryTimeoutMs` (success or a terminal child failure), DOCX child exit 1 then LibreOffice `warmTimeoutMs + primaryTimeoutMs + sofficeTimeoutMs + primaryTimeoutMs + fallbackTimeoutMs`, DOCX without a DOCX-capable backend `warmTimeoutMs + sofficeTimeoutMs + primaryTimeoutMs + fallbackTimeoutMs`; Excel `warmTimeoutMs + excelTimeoutMs + sofficeTimeoutMs + fallbackTimeoutMs`. Add `KILL_GRACE_MS` (2000 ms) per kill. There is no cap on image count, image bytes, cell count or workbook memory - deliberately; the per-tier timeouts, the rendered-view pixel budget and `maxOutputBytes` are the bounds.
313
313
 
314
314
  ### Migrating from flat keys
315
315
 
@@ -517,7 +517,7 @@ async function fetchUrl(opts) {
517
517
  }
518
518
 
519
519
  // lib/doc-to-md-core.ts
520
- import { existsSync as existsSync2, mkdtempSync, renameSync as renameSync2, rmSync as rmSync2, statSync as statSync2 } from "node:fs";
520
+ import { existsSync as existsSync2, mkdtempSync, readdirSync as readdirSync2, renameSync as renameSync2, rmSync as rmSync2, statSync as statSync2 } from "node:fs";
521
521
  import { spawn } from "node:child_process";
522
522
  import { homedir, tmpdir as tmpdir3 } from "node:os";
523
523
  import { basename, dirname, extname as extname3, join as join3, resolve as resolve2 } from "node:path";
@@ -697,11 +697,19 @@ function compactRanges(nums, maxEntries = 20) {
697
697
  if (parts.length <= maxEntries) return parts.join(", ");
698
698
  return `${parts.slice(0, maxEntries).join(", ")} (+${parts.length - maxEntries} more)`;
699
699
  }
700
+ var PAGE_MARKER_RE = /^--- end of page\.page_number=(\d+) ---$/;
700
701
  function scanOutline(md, max) {
701
702
  const entries = [];
703
+ let pending = [];
702
704
  let total = 0, inFence = false;
703
705
  md.split("\n").forEach((raw, i) => {
704
706
  const line = raw.replace(/\r$/, "");
707
+ const marker = line.match(PAGE_MARKER_RE);
708
+ if (marker) {
709
+ for (const e of pending) e.page = Number(marker[1]);
710
+ pending = [];
711
+ return;
712
+ }
705
713
  if (/^\s*(```|~~~)/.test(line)) {
706
714
  inFence = !inFence;
707
715
  return;
@@ -710,23 +718,36 @@ function scanOutline(md, max) {
710
718
  const m = line.match(/^(#{1,6}) (.*)$/);
711
719
  if (!m) return;
712
720
  total++;
713
- if (entries.length < max) entries.push({ line: i + 1, level: m[1].length, title: trunc(m[2].trim(), TITLE_MAX) });
721
+ if (entries.length < max) {
722
+ const e = { line: i + 1, level: m[1].length, title: trunc(m[2].trim(), TITLE_MAX), page: null };
723
+ entries.push(e);
724
+ pending.push(e);
725
+ }
714
726
  });
715
727
  return { entries, total };
716
728
  }
717
729
  function outlineLines(entries, total) {
730
+ if (total === 0) return ["Outline: none"];
718
731
  if (entries.length === 0) return [];
719
- const width = Math.max(5, Math.max(...entries.map((e) => `L${e.line}`.length)) + 2);
720
- const out = ["Outline:", ...entries.map((e) => ` ${`L${e.line}`.padEnd(width)}${"#".repeat(e.level)} ${trunc(e.title, TITLE_MAX)}`)];
732
+ const lineWidth = Math.max(3, ...entries.map((e) => `L${e.line}`.length)) + 2;
733
+ const paged = entries.filter((e) => e.page !== null);
734
+ const pageWidth = paged.length ? Math.max(...paged.map((e) => `p${e.page}`.length)) + 2 : 0;
735
+ const out = ["Outline:", ...entries.map((e) => ` ${`L${e.line}`.padEnd(lineWidth)}${pageWidth ? (e.page === null ? "" : `p${e.page}`).padEnd(pageWidth) : ""}${"#".repeat(e.level)} ${trunc(e.title, TITLE_MAX)}`)];
721
736
  if (total > entries.length) out.push(` (+${total - entries.length} more)`);
722
737
  return out;
723
738
  }
739
+ function pageCountLabel(h) {
740
+ const n = h.pageCount ?? "?";
741
+ if (h.type !== "docx") return String(n);
742
+ if (h.tier === "docx") return `${n} (${(h.explicitBreaks ?? 0) > 0 ? "explicit page breaks, not printed pages" : "no explicit page breaks"})`;
743
+ return `${n} (LibreOffice pagination)`;
744
+ }
724
745
  function formatHandle(h) {
725
746
  const lines = [`Saved-To: ${h.savedTo}`];
726
747
  if (h.imagesDir && h.imageCount > 0) lines.push(`Images-Dir: ${h.imagesDir}`);
727
748
  if (h.sheetsDir) lines.push(`Sheets-Dir: ${h.sheetsDir}`);
728
749
  lines.push(`Type: ${h.type} Engine: ${h.engine} Tier: ${h.tier}`);
729
- lines.push(`Page-Count: ${h.pageCount ?? "?"} Pages: ${h.pages ? compactRanges(h.pages) : "all"} Images: ${h.imageCount} Size: ${formatSize2(h.bytes)} / ${h.lines} lines`);
750
+ lines.push(`Page-Count: ${pageCountLabel(h)} Pages: ${h.pages ? compactRanges(h.pages) : "all"} Images: ${h.imageCount} Size: ${formatSize2(h.bytes)} / ${h.lines} lines`);
730
751
  if (h.degraded) lines.push(`Degraded: ${h.degraded}`);
731
752
  if (h.fallbackReason) lines.push(`Fallback-Reason: ${h.fallbackReason}`);
732
753
  const fe = [];
@@ -752,9 +773,9 @@ function formatInfoHandle(i, max) {
752
773
  const meta = Object.entries(i.metadata).filter(([, v]) => v).map(([k, v]) => `${k[0].toUpperCase()}${k.slice(1)}: ${trunc(v, META_MAX_CHARS)}`);
753
774
  if (meta.length) lines.push(meta.join(" "));
754
775
  if (i.toc.length) {
755
- lines.push("TOC:", ...i.toc.slice(0, max).map((t) => ` L${t.level} ${trunc(t.title, TITLE_MAX)} (p${t.page})`));
776
+ lines.push("TOC:", ...i.toc.slice(0, max).map((t) => ` L${t.level} ${trunc(t.title, TITLE_MAX)} (p${t.page ?? "?"})`));
756
777
  if (i.tocTotal > max) lines.push(` (+${i.tocTotal - max} more)`);
757
- }
778
+ } else if (i.type === "docx") lines.push("TOC: none (no heading styles found)");
758
779
  return lines.join("\n");
759
780
  }
760
781
 
@@ -767,11 +788,11 @@ var VERSION_RE = /^\d+(\.\d+)*$/;
767
788
  var DOC_TO_MD_OPTIONS = [
768
789
  { key: "path", flag: null, type: "string", default: null, settable: false, help: "Local .pdf .docx .pptx .xlsx .xls file" },
769
790
  { key: "info", flag: "--info", type: "bool", default: false, settable: false, help: "Inspect only (page count, metadata, TOC or sheet inventory); no bundle" },
770
- { key: "pages", flag: "--pages", type: "pages", default: null, settable: false, help: 'Inclusive 1-based pages, e.g. "12-15" or "3,7,10-12" (PDF/DOCX/PPTX only); default all' },
791
+ { key: "pages", flag: "--pages", type: "pages", default: null, settable: false, help: 'Inclusive 1-based pages, e.g. "12-15" or "3,7,10-12" (PDF/DOCX/PPTX only); default all. DOCX: selects explicit-page-break segments; rejected when the file has none' },
771
792
  { key: "outputDir", flag: "--output-dir", type: "string", default: null, settable: false, help: "Bundle root for <stem>.md + images/; default a per-call temp dir" },
772
793
  { key: "overwrite", flag: "--overwrite", type: "bool", default: false, settable: false, help: "Replace an existing completed <stem>.md bundle" },
773
- { key: "primaryTimeoutMs", flag: "--primary-timeout", type: "int", default: 6e4, settable: true, env: "PI_DOC_TO_MD_CONVERT_TIMEOUT_MS", help: "pymupdf4llm tier; also the unpdf tier" },
774
- { key: "fallbackTimeoutMs", flag: "--fallback-timeout", type: "int", default: 3e4, settable: true, help: "PyMuPDF get_text tier; also info on PDF and Excel rendered views" },
794
+ { key: "primaryTimeoutMs", flag: "--primary-timeout", type: "int", default: 6e4, settable: true, env: "PI_DOC_TO_MD_CONVERT_TIMEOUT_MS", help: "pymupdf4llm tier and DOCX child (docx mode); also the unpdf tier" },
795
+ { key: "fallbackTimeoutMs", flag: "--fallback-timeout", type: "int", default: 3e4, settable: true, help: "PyMuPDF get_text tier (including DOCX LibreOffice fallback); also PDF and DOCX info and Excel rendered views" },
775
796
  { key: "sofficeTimeoutMs", flag: "--soffice-timeout", type: "int", default: 12e4, settable: true, env: "PI_DOC_TO_MD_SOFFICE_TIMEOUT_MS", help: "DOCX/PPTX -> PDF via LibreOffice; also Excel rendered views" },
776
797
  { key: "excelTimeoutMs", flag: "--excel-timeout", type: "int", default: 6e4, settable: true, help: "Excel child (both openpyxl loads); also info on Excel" },
777
798
  { key: "warmTimeoutMs", flag: "--warm-timeout", type: "int", default: 12e4, settable: true, env: "PI_DOC_TO_MD_WARM_TIMEOUT_MS", help: "Absolute backend discovery/bootstrap deadline (first call per process)" },
@@ -892,21 +913,22 @@ function renderHelp() {
892
913
  }
893
914
 
894
915
  // lib/doc-to-md-core.ts
895
- var PACKAGE_PINS = { pymupdf4llm: TUNABLE_DEFAULTS.pymupdfVersion, openpyxl: "3.1.5", xlrd: "2.0.2", pillow: "12.3.0" };
916
+ var PACKAGE_PINS = { pymupdf4llm: TUNABLE_DEFAULTS.pymupdfVersion, openpyxl: "3.1.5", xlrd: "2.0.2", pillow: "12.3.0", mammoth: "1.13.0", markdownify: "1.2.3", "python-docx": "1.2.0" };
896
917
  var KILL_GRACE_MS = 2e3;
897
- var VENV_DIR_NAME = "doc-to-md-venv-v2";
898
- var LEGACY_VENV_DIR_NAME = "pymupdf-venv";
918
+ var VENV_DIR_NAME = "doc-to-md-venv-v3";
919
+ var LEGACY_VENV_DIR_NAMES = ["pymupdf-venv", "doc-to-md-venv-v2"];
899
920
  var STDERR_CAP = 1e6;
900
921
  var OUTPUT_MAX_BYTES = 2e7;
901
922
  var EXCEL_PDF_FILTER = 'pdf:calc_pdf_Export:{"SinglePageSheets":{"type":"boolean","value":"true"}}';
923
+ var pinSpecs = (cfg) => [`pymupdf4llm==${cfg.pymupdfVersion}`, `openpyxl==${PACKAGE_PINS.openpyxl}`, `xlrd==${PACKAGE_PINS.xlrd}`, `pillow==${PACKAGE_PINS.pillow}`, `mammoth==${PACKAGE_PINS.mammoth}`, `markdownify==${PACKAGE_PINS.markdownify}`, `python-docx==${PACKAGE_PINS["python-docx"]}`];
902
924
  function withArgs(cfg) {
903
- return ["--with", `pymupdf4llm==${cfg.pymupdfVersion}`, "--with", `openpyxl==${PACKAGE_PINS.openpyxl}`, "--with", `xlrd==${PACKAGE_PINS.xlrd}`, "--with", `pillow==${PACKAGE_PINS.pillow}`];
925
+ return pinSpecs(cfg).flatMap((spec) => ["--with", spec]);
904
926
  }
905
927
  function pipInstallArgs(cfg) {
906
- return ["-m", "pip", "install", `pymupdf4llm==${cfg.pymupdfVersion}`, `openpyxl==${PACKAGE_PINS.openpyxl}`, `xlrd==${PACKAGE_PINS.xlrd}`, `pillow==${PACKAGE_PINS.pillow}`];
928
+ return ["-m", "pip", "install", ...pinSpecs(cfg)];
907
929
  }
908
930
  function warmArgs(cfg) {
909
- return ["run", ...withArgs(cfg), "--python", "3.14", "python", "-c", "import pymupdf4llm, openpyxl, xlrd, PIL"];
931
+ return ["run", ...withArgs(cfg), "--python", "3.14", "python", "-c", "import pymupdf4llm, openpyxl, xlrd, PIL, mammoth, markdownify, docx"];
910
932
  }
911
933
  function uvChildArgs(cfg, script, mode) {
912
934
  return ["run", ...withArgs(cfg), "--python", "3.14", "python", script, mode];
@@ -1058,6 +1080,11 @@ try:
1058
1080
  print("XLSX", "yes")
1059
1081
  except Exception:
1060
1082
  print("XLSX", "no")
1083
+ try:
1084
+ import mammoth, markdownify, docx
1085
+ print("DOCX", "yes")
1086
+ except Exception:
1087
+ print("DOCX", "no")
1061
1088
  `;
1062
1089
  var PROBE_TIMEOUT_MS = 5e3;
1063
1090
  function probeArgs() {
@@ -1065,8 +1092,8 @@ function probeArgs() {
1065
1092
  }
1066
1093
  var PYTHON_CANDIDATES = ["python3", "python"];
1067
1094
  function parseProbeOutput(stdout) {
1068
- const m = stdout.match(/^PY (\d+) (\d+)\r?\nPDF (yes|no)\r?\nXLSX (yes|no)\s*$/);
1069
- return m ? { major: Number(m[1]), minor: Number(m[2]), pdf: m[3] === "yes", xlsx: m[4] === "yes" } : null;
1095
+ const m = stdout.match(/^PY (\d+) (\d+)\r?\nPDF (yes|no)\r?\nXLSX (yes|no)\r?\nDOCX (yes|no)\s*$/);
1096
+ return m ? { major: Number(m[1]), minor: Number(m[2]), pdf: m[3] === "yes", xlsx: m[4] === "yes", docx: m[5] === "yes" } : null;
1070
1097
  }
1071
1098
  function meetsFloor(p) {
1072
1099
  return p.major > 3 || p.major === 3 && p.minor >= 12;
@@ -1091,7 +1118,7 @@ async function resolveBackend(cfg, deps, signal) {
1091
1118
  };
1092
1119
  const warm = await deps.run("uv", warmArgs(cfg), { timeoutMs: left(), capBytes: OUTPUT_MAX_BYTES, env: deps.env, signal });
1093
1120
  if (signal?.aborted) throw new Error("aborted");
1094
- if (warm.code === 0 && !warm.timedOut) return { kind: "uv", pdf: true, xlsx: true };
1121
+ if (warm.code === 0 && !warm.timedOut) return { kind: "uv", pdf: true, xlsx: true, docx: true };
1095
1122
  const uvAbsent = warm.code === null && !warm.timedOut;
1096
1123
  const isDeadline = (result) => result !== null && "kind" in result;
1097
1124
  const probe = async (exe, tmp) => {
@@ -1107,18 +1134,19 @@ async function resolveBackend(cfg, deps, signal) {
1107
1134
  const p = await probe(exe);
1108
1135
  if (isDeadline(p)) return p;
1109
1136
  if (!p || !meetsFloor(p)) continue;
1110
- if (p.pdf) return { kind: "python", exe, pdf: true, xlsx: p.xlsx };
1137
+ if (p.pdf) return { kind: "python", exe, pdf: true, xlsx: p.xlsx, docx: p.docx };
1111
1138
  eligible ??= { exe, version: `${p.major}.${p.minor}` };
1112
1139
  }
1140
+ const healthy = (p) => meetsFloor(p) && p.pdf && p.xlsx && p.docx;
1113
1141
  const venvDir = join3(deps.cacheRoot, VENV_DIR_NAME);
1114
1142
  const venvExe = venvPython(venvDir, deps.platform);
1115
1143
  const cached = await probe(venvExe);
1116
1144
  if (isDeadline(cached)) return cached;
1117
- if (cached && meetsFloor(cached) && cached.pdf && cached.xlsx) return { kind: "venv", exe: venvExe, pdf: true, xlsx: true };
1145
+ if (cached && healthy(cached)) return { kind: "venv", exe: venvExe, pdf: true, xlsx: true, docx: true };
1118
1146
  if (eligible) {
1119
1147
  const recheck = await probe(venvExe);
1120
1148
  if (isDeadline(recheck)) return recheck;
1121
- if (recheck && meetsFloor(recheck) && recheck.pdf && recheck.xlsx) return { kind: "venv", exe: venvExe, pdf: true, xlsx: true };
1149
+ if (recheck && healthy(recheck)) return { kind: "venv", exe: venvExe, pdf: true, xlsx: true, docx: true };
1122
1150
  const tmp = `${venvDir}.tmp-${deps.pid}`;
1123
1151
  const bootFail = (stderr) => {
1124
1152
  deps.rmrf(tmp);
@@ -1143,19 +1171,19 @@ async function resolveBackend(cfg, deps, signal) {
1143
1171
  }
1144
1172
  };
1145
1173
  if (publish()) {
1146
- deps.rmrf(join3(deps.cacheRoot, LEGACY_VENV_DIR_NAME));
1147
- return { kind: "venv", exe: venvExe, pdf: true, xlsx: true };
1174
+ for (const legacy of LEGACY_VENV_DIR_NAMES) deps.rmrf(join3(deps.cacheRoot, legacy));
1175
+ return { kind: "venv", exe: venvExe, pdf: true, xlsx: true, docx: true };
1148
1176
  }
1149
1177
  const winner = await probe(venvExe, tmp);
1150
1178
  if (isDeadline(winner)) return winner;
1151
- if (winner && meetsFloor(winner) && winner.pdf && winner.xlsx) {
1179
+ if (winner && healthy(winner)) {
1152
1180
  deps.rmrf(tmp);
1153
- return { kind: "venv", exe: venvExe, pdf: true, xlsx: true };
1181
+ return { kind: "venv", exe: venvExe, pdf: true, xlsx: true, docx: true };
1154
1182
  }
1155
1183
  deps.rmrf(venvDir);
1156
1184
  if (publish()) {
1157
- deps.rmrf(join3(deps.cacheRoot, LEGACY_VENV_DIR_NAME));
1158
- return { kind: "venv", exe: venvExe, pdf: true, xlsx: true };
1185
+ for (const legacy of LEGACY_VENV_DIR_NAMES) deps.rmrf(join3(deps.cacheRoot, legacy));
1186
+ return { kind: "venv", exe: venvExe, pdf: true, xlsx: true, docx: true };
1159
1187
  }
1160
1188
  return bootFail("rename after competing bootstrap");
1161
1189
  }
@@ -1213,15 +1241,23 @@ async function tryConvertOffice(sofficeTimeoutMs, src, signal, run = runCapped,
1213
1241
  throw e;
1214
1242
  }
1215
1243
  }
1216
- async function convertOffice(sofficeTimeoutMs, src, signal, run = runCapped) {
1217
- const r = await tryConvertOffice(sofficeTimeoutMs, src, signal, run);
1218
- if (r.ok) return r;
1219
- if (r.kind === "missing") throw new Error("LibreOffice (soffice) is required to convert .docx/.pptx but was not found on PATH. Install LibreOffice or convert the file to PDF first.");
1220
- if (r.kind === "no-pdf") throw new Error("LibreOffice (soffice) ran but produced no usable PDF for this file. Ensure LibreOffice can open the document, or convert it to PDF manually first.");
1221
- throw new Error(`soffice failed (code=${r.code} timedOut=${r.timedOut}): ${r.stderr}`);
1244
+ function officeFailure(r) {
1245
+ if (r.kind === "missing") return new Error("LibreOffice (soffice) is required to convert .docx/.pptx but was not found on PATH. Install LibreOffice or convert the file to PDF first.");
1246
+ if (r.kind === "no-pdf") return new Error("LibreOffice (soffice) ran but produced no usable PDF for this file. Ensure LibreOffice can open the document, or convert it to PDF manually first.");
1247
+ return new Error(`soffice failed (code=${r.code} timedOut=${r.timedOut}): ${r.stderr}`);
1222
1248
  }
1223
1249
  var DEGRADED_TEXT = "PyMuPDF text extraction - layout/tables not preserved";
1224
1250
  var DEGRADED_UNPDF = "unpdf text extraction - structure not preserved";
1251
+ var DEGRADED_DOCX_TEXT = "python-docx text extraction - footnotes, hyperlinks, images not preserved";
1252
+ var DEGRADED_DOCX_OFFICE = "LibreOffice PDF route - heading styles and explicit page breaks not preserved; page numbers are LibreOffice pagination";
1253
+ var DOCX_PIP = "pip install mammoth markdownify python-docx";
1254
+ var DOCX_PAGES_OFFICE = `--pages on a DOCX needs the Python DOCX backend (explicit page-break segments); the LibreOffice route has none. Remedy: install uv, or ${DOCX_PIP}`;
1255
+ var docxRemedy = `Remedy: install uv, or ${DOCX_PIP} into a Python that already has pymupdf4llm`;
1256
+ var backendState = (b, missing) => b.kind === "none" ? b.reason : missing;
1257
+ var lacksDocx = (b) => b.kind === "none" || !b.docx;
1258
+ var clearStaging = (b) => {
1259
+ for (const f of readdirSync2(b.stagingDir)) rmSync2(join3(b.stagingDir, f), { recursive: true, force: true });
1260
+ };
1225
1261
  var EXCEL_REMEDY = "Remedy: install uv, or pip install openpyxl xlrd pillow";
1226
1262
  function resolveUnpdfWorker(candidates, exists) {
1227
1263
  for (const candidate of candidates) if (exists(candidate)) return candidate;
@@ -1291,78 +1327,120 @@ async function convertDocument(o, signal, seams) {
1291
1327
  let office = null;
1292
1328
  try {
1293
1329
  let pdfPath = inputPath;
1294
- if (type === "docx" || type === "pptx") {
1295
- office = await convertOffice(o.sofficeTimeoutMs, inputPath, signal);
1296
- pdfPath = office.pdfPath;
1297
- }
1298
- const base = { path: pdfPath, pages: o.pages, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion };
1330
+ const base = { path: inputPath, pages: o.pages, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion };
1299
1331
  let tier, engine, json, degraded = null, fallbackReason = null;
1332
+ let explicitBreaks = null;
1300
1333
  let notes = [];
1301
- if (isExcel) {
1302
- const r = await s.runTier("xlsx", base, b, signal, o.excelTimeoutMs, backend);
1303
- if (!r.ok) {
1304
- if ("userError" in r) throw new Error(r.userError);
1305
- const remedy = r.reason.startsWith("timeout after") || r.reason === "output exceeded maxOutputBytes" ? ". Remedy: raise excelTimeoutMs" : "";
1306
- throw new Error(`Excel conversion failed: ${r.reason}${detailSuffix(r)}${remedy}`);
1334
+ let officeRoute = null;
1335
+ if (type === "docx" && !lacksDocx(backend)) {
1336
+ const d = await s.runTier("docx", base, b, signal, o.primaryTimeoutMs, backend);
1337
+ if (d.ok) {
1338
+ publishStaged(b);
1339
+ tier = "docx";
1340
+ engine = d.json.engine === "python-docx" ? "python-docx" : "mammoth";
1341
+ json = d.json;
1342
+ explicitBreaks = d.json.explicitBreaks ?? 0;
1343
+ if (d.json.degraded) {
1344
+ degraded = DEGRADED_DOCX_TEXT;
1345
+ fallbackReason = d.json.fallbackReason ?? null;
1346
+ }
1347
+ } else if ("userError" in d) throw new Error(d.userError);
1348
+ else if (d.reason === "exit 1") {
1349
+ officeRoute = `docx ${d.reason}${detailSuffix(d)}`;
1350
+ clearStaging(b);
1351
+ } else throw new Error(`Conversion failed: docx ${d.reason}${detailSuffix(d)}`);
1352
+ } else if (type === "docx") officeRoute = backend.kind === "none" ? backend.reason : "python backend lacks DOCX packages";
1353
+ if (type === "docx" && officeRoute !== null) {
1354
+ if (o.pages) throw new Error(officeRoute.startsWith("docx exit") ? `${DOCX_PAGES_OFFICE} (${officeRoute})` : DOCX_PAGES_OFFICE);
1355
+ const r = await s.office(o.sofficeTimeoutMs, inputPath, signal);
1356
+ if (!r.ok && r.kind === "missing") {
1357
+ if (officeRoute.startsWith("docx exit")) throw new Error(`Conversion failed: ${officeRoute}; LibreOffice (soffice) not found on PATH`);
1358
+ throw new Error(`DOCX conversion needs the Python DOCX packages or LibreOffice. Python backend: ${backendState(backend, "found without mammoth/markdownify/python-docx")}. ${docxRemedy}, or install LibreOffice (soffice)`);
1307
1359
  }
1308
- publishSheetImages(b);
1309
- publishSheetCsvs(b);
1310
- tier = "excel";
1311
- engine = type === "xls" ? "xlrd" : "openpyxl";
1312
- json = r.json;
1313
- notes = [...json.notes ?? []];
1314
- const renderPages = json.renderPages ?? [];
1315
- let skip = null;
1316
- const perSheet = /* @__PURE__ */ new Map();
1317
- if (renderPages.length) {
1318
- const off = await s.office(o.sofficeTimeoutMs, inputPath, signal, runCapped, EXCEL_PDF_FILTER);
1319
- if (!off.ok) skip = off.kind === "missing" ? "LibreOffice not found" : off.kind === "timeout" ? `soffice failed: timeout after ${o.sofficeTimeoutMs}ms` : off.kind === "exit" ? `soffice failed: exit ${off.code}` : "soffice produced no PDF";
1320
- else {
1321
- try {
1322
- const rp = await s.runTier("render-pages", { path: off.pdfPath, sheetIndices: renderPages, expectedPages: json.sheetCount, imageDpi: o.imageDpi, imageFormat: o.imageFormat, stagingDir: b.stagingDir, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion }, b, signal, o.fallbackTimeoutMs, backend);
1323
- if (!rp.ok) skip = `render failed: ${"userError" in rp ? rp.userError : rp.reason}`;
1324
- else if (rp.json.ok === false) skip = rp.json.reason ?? "render failed";
1325
- else {
1326
- publishSheetImages(b);
1327
- for (const f of rp.json.failed ?? []) perSheet.set(f.idx, f.reason);
1328
- for (const d of rp.json.rendered ?? []) if (d.dpi < o.imageDpi) notes.push(`Rendered view s${d.idx}: rendered at ${d.dpi} dpi`);
1360
+ if (!r.ok) throw officeRoute.startsWith("docx exit") ? new Error(`Conversion failed: ${officeRoute}; ${officeFailure(r).message}`) : officeFailure(r);
1361
+ office = r;
1362
+ pdfPath = r.pdfPath;
1363
+ }
1364
+ if (type === "pptx") {
1365
+ const r = await s.office(o.sofficeTimeoutMs, inputPath, signal);
1366
+ if (!r.ok && r.kind === "missing") throw new Error(`PPTX conversion needs LibreOffice (soffice); direct conversion is not available. Python backend: ${backendState(backend, "available")}. Remedy: install LibreOffice`);
1367
+ if (!r.ok) throw officeFailure(r);
1368
+ office = r;
1369
+ pdfPath = r.pdfPath;
1370
+ }
1371
+ const pdfBase = { ...base, path: pdfPath };
1372
+ if (json === void 0) {
1373
+ if (isExcel) {
1374
+ const r = await s.runTier("xlsx", base, b, signal, o.excelTimeoutMs, backend);
1375
+ if (!r.ok) {
1376
+ if ("userError" in r) throw new Error(r.userError);
1377
+ const remedy = r.reason.startsWith("timeout after") || r.reason === "output exceeded maxOutputBytes" ? ". Remedy: raise excelTimeoutMs" : "";
1378
+ throw new Error(`Excel conversion failed: ${r.reason}${detailSuffix(r)}${remedy}`);
1379
+ }
1380
+ publishSheetImages(b);
1381
+ publishSheetCsvs(b);
1382
+ tier = "excel";
1383
+ engine = type === "xls" ? "xlrd" : "openpyxl";
1384
+ json = r.json;
1385
+ notes = [...json.notes ?? []];
1386
+ const renderPages = json.renderPages ?? [];
1387
+ let skip = null;
1388
+ const perSheet = /* @__PURE__ */ new Map();
1389
+ if (renderPages.length) {
1390
+ const off = await s.office(o.sofficeTimeoutMs, inputPath, signal, runCapped, EXCEL_PDF_FILTER);
1391
+ if (!off.ok) skip = off.kind === "missing" ? "LibreOffice not found" : off.kind === "timeout" ? `soffice failed: timeout after ${o.sofficeTimeoutMs}ms` : off.kind === "exit" ? `soffice failed: exit ${off.code}` : "soffice produced no PDF";
1392
+ else {
1393
+ try {
1394
+ const rp = await s.runTier("render-pages", { path: off.pdfPath, sheetIndices: renderPages, expectedPages: json.sheetCount, imageDpi: o.imageDpi, imageFormat: o.imageFormat, stagingDir: b.stagingDir, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion }, b, signal, o.fallbackTimeoutMs, backend);
1395
+ if (!rp.ok) skip = `render failed: ${"userError" in rp ? rp.userError : rp.reason}`;
1396
+ else if (rp.json.ok === false) skip = rp.json.reason ?? "render failed";
1397
+ else {
1398
+ publishSheetImages(b);
1399
+ for (const f of rp.json.failed ?? []) perSheet.set(f.idx, f.reason);
1400
+ for (const d of rp.json.rendered ?? []) if (d.dpi < o.imageDpi) notes.push(`Rendered view s${d.idx}: rendered at ${d.dpi} dpi`);
1401
+ }
1402
+ } finally {
1403
+ off.cleanup();
1329
1404
  }
1330
- } finally {
1331
- off.cleanup();
1332
1405
  }
1406
+ if (skip) notes.push(`Rendered views skipped: ${skip}`);
1407
+ else if (perSheet.size) notes.push(`Rendered views: ${perSheet.size} of ${renderPages.length} unavailable`);
1408
+ }
1409
+ json = { ...json, markdown: reconcileRenderMarkers(json.markdown ?? "", renderPages, o.imageFormat, b.sourceMap, (idx) => perSheet.get(idx) ?? skip ?? "render failed") };
1410
+ } else if (backend.kind === "none") {
1411
+ const r = await s.runTier("pdf-text", pdfBase, b, signal, o.primaryTimeoutMs, backend);
1412
+ if (!r.ok) throw new Error("userError" in r ? r.userError : `Conversion failed: unpdf ${r.reason}${detailSuffix(r)}`);
1413
+ tier = "unpdf";
1414
+ engine = "unpdf";
1415
+ json = r.json;
1416
+ degraded = DEGRADED_UNPDF;
1417
+ } else {
1418
+ const p = await s.runTier("pdf-primary", pdfBase, b, signal, o.primaryTimeoutMs, backend);
1419
+ const kept = publishStaged(b);
1420
+ if (p.ok) {
1421
+ tier = "primary";
1422
+ engine = "pymupdf4llm";
1423
+ json = p.json;
1424
+ } else if ("userError" in p) throw new Error(p.userError);
1425
+ else {
1426
+ if (signal?.aborted) throw new Error("aborted");
1427
+ const keepPages = Object.fromEntries([...kept.entries()].map(([k, v]) => [String(k), v]));
1428
+ const f = await s.runTier("pdf-fallback", { ...pdfBase, keepPages }, b, signal, o.fallbackTimeoutMs, backend);
1429
+ publishStaged(b);
1430
+ if (!f.ok) throw new Error("userError" in f ? f.userError : `Conversion failed: primary ${p.reason}; fallback ${f.reason}${detailSuffix(f)}`);
1431
+ tier = "fallback";
1432
+ engine = "pymupdf-text";
1433
+ json = f.json;
1434
+ degraded = DEGRADED_TEXT;
1435
+ fallbackReason = `primary ${p.reason}`;
1333
1436
  }
1334
- if (skip) notes.push(`Rendered views skipped: ${skip}`);
1335
- else if (perSheet.size) notes.push(`Rendered views: ${perSheet.size} of ${renderPages.length} unavailable`);
1336
- }
1337
- json = { ...json, markdown: reconcileRenderMarkers(json.markdown ?? "", renderPages, o.imageFormat, b.sourceMap, (idx) => perSheet.get(idx) ?? skip ?? "render failed") };
1338
- } else if (backend.kind === "none") {
1339
- const r = await s.runTier("pdf-text", base, b, signal, o.primaryTimeoutMs, backend);
1340
- if (!r.ok) throw new Error("userError" in r ? r.userError : `Conversion failed: unpdf ${r.reason}${detailSuffix(r)}`);
1341
- tier = "unpdf";
1342
- engine = "unpdf";
1343
- json = r.json;
1344
- degraded = DEGRADED_UNPDF;
1345
- } else {
1346
- const p = await s.runTier("pdf-primary", base, b, signal, o.primaryTimeoutMs, backend);
1347
- const kept = publishStaged(b);
1348
- if (p.ok) {
1349
- tier = "primary";
1350
- engine = "pymupdf4llm";
1351
- json = p.json;
1352
- } else if ("userError" in p) throw new Error(p.userError);
1353
- else {
1354
- if (signal?.aborted) throw new Error("aborted");
1355
- const keepPages = Object.fromEntries([...kept.entries()].map(([k, v]) => [String(k), v]));
1356
- const f = await s.runTier("pdf-fallback", { ...base, keepPages }, b, signal, o.fallbackTimeoutMs, backend);
1357
- publishStaged(b);
1358
- if (!f.ok) throw new Error("userError" in f ? f.userError : `Conversion failed: primary ${p.reason}; fallback ${f.reason}${detailSuffix(f)}`);
1359
- tier = "fallback";
1360
- engine = "pymupdf-text";
1361
- json = f.json;
1362
- degraded = DEGRADED_TEXT;
1363
- fallbackReason = `primary ${p.reason}`;
1364
1437
  }
1365
1438
  }
1439
+ if (officeRoute !== null) {
1440
+ degraded = DEGRADED_DOCX_OFFICE;
1441
+ fallbackReason = fallbackReason ? `${officeRoute}; ${fallbackReason}` : officeRoute;
1442
+ }
1443
+ if (tier === void 0 || engine === void 0 || json === void 0) throw new Error("internal: no tier produced output");
1366
1444
  if (!isExcel) notes = json.notes ?? [];
1367
1445
  const body = rewriteLinks(json.markdown ?? "", b.sourceMap);
1368
1446
  validateImageLinks(body, b.manifest, b.csvManifest);
@@ -1377,7 +1455,7 @@ async function convertDocument(o, signal, seams) {
1377
1455
  ` : "") + body;
1378
1456
  commitBundle(b, markdown);
1379
1457
  const outline = scanOutline(markdown, o.outlineMaxEntries);
1380
- const details = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total };
1458
+ const details = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total };
1381
1459
  return { output: formatHandle(details), details };
1382
1460
  } catch (e) {
1383
1461
  abortBundle(b);
@@ -1398,9 +1476,13 @@ async function inspectDocument(o, signal, seams) {
1398
1476
  let office = null;
1399
1477
  try {
1400
1478
  let path = inputPath;
1401
- if (type === "docx" || type === "pptx") {
1402
- office = await convertOffice(o.sofficeTimeoutMs, inputPath, signal);
1403
- path = office.pdfPath;
1479
+ if (type === "docx") {
1480
+ if (lacksDocx(backend)) throw new Error(`DOCX inspection needs the Python DOCX packages. Python backend: ${backendState(backend, "found without mammoth/markdownify/python-docx")}. ${docxRemedy}`);
1481
+ } else if (type === "pptx") {
1482
+ const r2 = await s.office(o.sofficeTimeoutMs, inputPath, signal);
1483
+ if (!r2.ok) throw officeFailure(r2);
1484
+ office = r2;
1485
+ path = r2.pdfPath;
1404
1486
  }
1405
1487
  const r = await s.runTier("info", { path, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion }, { stagingDir: "" }, signal, isExcel ? o.excelTimeoutMs : o.fallbackTimeoutMs, backend);
1406
1488
  if (!r.ok) {
@@ -35,7 +35,7 @@ export default function docToMdExtension(pi: ExtensionAPI) {
35
35
  name: "doc_to_md",
36
36
  label: "Convert doc to Markdown bundle",
37
37
  description:
38
- "Convert a local PDF/DOCX/PPTX/XLSX/XLS to a Markdown bundle on disk and return a handle (Saved-To, Images-Dir, Page-Count, Outline, diagnostics) - the Markdown itself is never inlined; read the Saved-To file (offset/limit) for content. `info: true` returns page count, metadata and TOC (or the sheet inventory) without converting - use it to pick `pages`. `pages` selects inclusive 1-based pages (PDF/DOCX/PPTX); every page ends with `--- end of page.page_number=N ---`. Images are always extracted into images/ with relative links. Primary engine pymupdf4llm, fallback PyMuPDF text (degraded, marked), pure-JS unpdf only when no Python backend exists. Excel yields a sheet inventory (worksheets + chartsheets), a full CSV per non-empty worksheet under sheets/, a bounded preview with formulas and cached values, merged/hidden disclosure, and rendered chart views when LibreOffice is available. DOCX/PPTX need LibreOffice (soffice). Use `outputDir` for a durable bundle; without it the bundle lands in a per-call temp dir. Input must be a local file path (use fetch first for URLs).",
38
+ "Convert a local PDF/DOCX/PPTX/XLSX/XLS to a Markdown bundle on disk and return a handle (Saved-To, Images-Dir, Page-Count, Outline, diagnostics) - the Markdown itself is never inlined; read the Saved-To file (offset/limit) for content. `info: true` returns page count, metadata and TOC (or the sheet inventory) without converting - use it to pick `pages`. `pages` selects inclusive 1-based pages (PDF/PPTX) or explicit-page-break segments (DOCX; rejected when the file has none); every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Images are always extracted into images/ with relative links. Primary engine pymupdf4llm, fallback PyMuPDF text (degraded, marked), pure-JS unpdf only when no Python backend exists. Excel yields a sheet inventory (worksheets + chartsheets), a full CSV per non-empty worksheet under sheets/, a bounded preview with formulas and cached values, merged/hidden disclosure, and rendered chart views when LibreOffice is available. DOCX converts directly (mammoth; python-docx text fallback) and keeps heading styles, so the handle Outline lists `L<line>` and `p<segment>` per heading; LibreOffice (soffice) is optional for DOCX (fallback route, degraded) and Excel rendered views, and required for PPTX. Use `outputDir` for a durable bundle; without it the bundle lands in a per-call temp dir. Input must be a local file path (use fetch first for URLs).",
39
39
  promptSnippet: "Convert a local PDF/DOCX/PPTX/XLSX to a Markdown bundle (handle returned; read Saved-To)",
40
40
  parameters: buildSchema(),
41
41