pi-quiver 6.8.0 → 6.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -8,6 +8,17 @@ Published to npm as `pi-quiver` (`pi install npm:pi-quiver`). Pushing a
8
8
  via OIDC trusted publishing. The release helper at
9
9
  `.agents/skills/release/scripts/release.sh` cuts the tag; CI publishes.
10
10
 
11
+ ## v6.10.0 - 2026-10-02
12
+
13
+ - `doc_to_md`: new per-call `words` option (`--words`; not settable) writes `<stem>.words.json` - per selected PDF/image page, every text-layer word and every word inline OCR recognized in the same run, with display-space bbox (points; source pixels for images) and `source` `text`/`ocr`. Under `--ocr-mode all` each sidecar gets `ocr/<stem>-pNNN.words.json`. Never triggers OCR; Markdown, page stats and OCR outcome are unchanged. Handle gains `Words:`, `--json` gains `wordsPath`, `wordsReason`, `wordsErrors`, `ocr.wordSidecars`; `--info --words` is a usage error (#28).
14
+ - `doc_to_md`: `--help` and the generated skill list every bundle artifact under `Bundle layout` (#28).
15
+
16
+ ## v6.9.0 - 2026-10-01
17
+
18
+ - `doc_to_md`: every conversion that goes through the Python PDF tiers (PDF, PPTX, `.doc`, DOCX via LibreOffice) writes `<stem>.pages.json` (per-page `chars`, `images`, `imageCoverage`) and the handle prints `Page-Stats:`; CLI `--json` gains `pageStats`, `pageStatsPath`, `ocrDir`. The unpdf tier writes no stats (#26).
19
+ - `doc_to_md`: new per-call `ocrMode` (`--ocr-mode textless|all`). `all` forces OCR on an explicit `pages` selection in a second, kill-capped Python child and writes `ocr/<stem>-pNNN.md` sidecars; the Markdown stays byte-identical to the same selected-page call without OCR. Refused without `--ocr`, without explicit pages, or on non-PDF/PPTX/DOC inputs (exit 2); missing Tesseract data is a hard error. The `OCR:` line names written, no-text, failed, budget-stopped, killed and not-attempted pages, with the exact `--pages` to re-run for budget-stopped and not-attempted pages (#26).
20
+ - `doc_to_md`: `--help` and the generated `skills/doc-to-md/SKILL.md` end with the two-pass OCR usage block (#26).
21
+
11
22
  ## v6.8.0 - 2026-10-01
12
23
 
13
24
  - `session-name`: opt-in automatic naming starts in the background after three completed model/tool rounds. Shorter runs start a non-blocking attempt at run end; initial generation is best-effort with a 30-second local deadline and no automatic retry per session activation.
package/README.md CHANGED
@@ -65,7 +65,7 @@ A 300 KB changelog page never touches your context window - you get a preview an
65
65
  | Extension | Tool | What it does |
66
66
  | --- | --- | --- |
67
67
  | `extensions/fetch.ts` | `fetch` | Retrieve URLs over HTTP(S). HTML -> Markdown (Readability extraction, Turndown conversion). Binary saved untouched to a temp file. GitHub issue/PR/repo/actions-run/actions-job URLs auto-route through `gh` (falls back to HTTP); failed runs/jobs include failed-step logs (best-effort, summary-only otherwise). Same size gate as `fetch`. Behavior lives in `lib/fetch-core.ts`; also exposed as the `pi-quiver fetch` CLI (see [Claude Code support](#claude-code-support)). |
68
- | `extensions/doc_to_md.ts` | `doc_to_md` | Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, MSG/EML, HTML, or image file to a Markdown bundle on disk (`<stem>.md` + `images/`, spreadsheet `sheets/`, optional `pageImages` in `pages/`, and email `attachments/`) and return a handle (paths, page count, outline, diagnostics) - never inline Markdown. `info` inspects PDF/Office inputs; `pages` selects 1-based PDF/Office pages (DOCX: explicit-page-break segments); HTML, image, and email inputs reject both; every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Tiers: pymupdf4llm -> PyMuPDF text (degraded) -> unpdf worker (no Python only). Excel -> sheet inventory, full CSVs, bounded previews, and optional rendered views. HTML keeps local and `data:` images in the bundle and remote images as links; image inputs keep the original image. Scanned PDF pages keep a page picture on both Python tiers; opt-in OCR needs Tesseract language data. DOCX converts directly (mammoth, python-docx fallback); LibreOffice pagination marks the degraded route. The Outline lists `L<line>` and `p<page>` per heading. Settings under `quiver.docToMd`. |
68
+ | `extensions/doc_to_md.ts` | `doc_to_md` | Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, MSG/EML, HTML, or image file to a Markdown bundle on disk (`<stem>.md` + `images/`, spreadsheet `sheets/`, optional `pageImages` in `pages/`, and email `attachments/`) and return a handle (paths, page count, outline, diagnostics) - never inline Markdown. `info` inspects PDF/Office inputs; `pages` selects 1-based PDF/Office pages (DOCX: explicit-page-break segments); HTML, image, and email inputs reject both; every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Tiers: pymupdf4llm -> PyMuPDF text (degraded) -> unpdf worker (no Python only). Excel -> sheet inventory, full CSVs, bounded previews, and optional rendered views. HTML keeps local and `data:` images in the bundle and remote images as links; image inputs keep the original image. Scanned PDF pages keep a page picture on both Python tiers; opt-in OCR needs Tesseract language data. DOCX converts directly (mammoth, python-docx fallback); LibreOffice pagination marks the degraded route. The Outline lists `L<line>` and `p<page>` per heading. Settings under `quiver.docToMd`. Every conversion that goes through the Python PDF tiers (PDF, PPTX, `.doc`, DOCX via LibreOffice) writes `<stem>.pages.json` (per-page chars, image count, image coverage; handle line `Page-Stats:`); unpdf writes no stats. `ocrMode: "all"` with `ocr: true` and an explicit `pages` forces OCR on those pages into `ocr/<stem>-pNNN.md` sidecars, leaving the Markdown byte-identical to the same selected-page call without OCR. `words: true` (CLI `--words`, per-call only - no settings key) writes `<stem>.words.json` with every word's bbox per selected PDF/image page (`source` `text` or `ocr`), plus `ocr/<stem>-pNNN.words.json` beside each forced-OCR sidecar; it never triggers OCR. |
69
69
  | `extensions/session-name.ts` | `/session-name` | Manual + opt-in automatic session naming, naming rules and deny list, long-session revisits, and Ghostty/Herdr tab rename. OFF by default. |
70
70
  | `extensions/sword-header.ts` | `/builtin-header` | Themed ASCII startup header replacing pi's default logo. OFF by default. |
71
71
  | `extensions/fast-mode.ts` | `/fast` | Inject Anthropic fast-mode (`speed: "fast"` + `anthropic-beta: fast-mode-2026-02-01`) into every Claude Opus 4.8 / Opus 5 request, any thinking level. `--fast` flag + `/fast [on\|off\|status]`. OFF by default. |
@@ -310,7 +310,9 @@ Each setting can also be overridden per-process via `PI_QUIVER_SLACK_ENABLED`, `
310
310
  | `ocr` | `false` | Opt-in OCR for scanned pages and images when Tesseract language data is installed. |
311
311
  | `ocrLanguage` | `eng` | Plain `+`-joined Tesseract language codes for OCR. |
312
312
 
313
- A bundle is `<outputDir>/<stem>.md` plus `<outputDir>/images/` and, for spreadsheets with data, `<outputDir>/sheets/`; without `outputDir`, the tool creates a per-call temp root. A conversion owns `<stem>.md.lock` until it atomically publishes the Markdown. An existing `<stem>.md` fails the call unless `overwrite` is set; `overwrite` replaces that Markdown and the files it owns (`images/<stem>-p<N>-<n>.*`, `images/<stem>-s<idx>[-<n>].*`, `sheets/<stem>-s<idx>-<slug>.csv`), nothing else. Temp bundles are caller-owned - the tool never deletes a bundle it produced.
313
+ `ocrMode` (`textless` | `all`) is a per-call parameter and `--ocr-mode` flag only; a `quiver.docToMd.ocrMode` key is reported as unknown and ignored, so persisted settings cannot enable forced OCR or bypass its explicit-pages guard.
314
+
315
+ A bundle is `<outputDir>/<stem>.md` plus `<outputDir>/<stem>.pages.json` on Python PDF tiers, optional OCR sidecars under `<outputDir>/ocr/`, and `<outputDir>/images/` and, for spreadsheets with data, `<outputDir>/sheets/`; without `outputDir`, the tool creates a per-call temp root. A conversion owns `<stem>.md.lock` until it atomically publishes the Markdown. An existing `<stem>.md` fails the call unless `overwrite` is set; `overwrite` replaces that Markdown and the files it owns (`images/<stem>-p<N>-<n>.*`, `images/<stem>-s<idx>[-<n>].*`, `sheets/<stem>-s<idx>-<slug>.csv`, `<stem>.pages.json`, `ocr/<stem>-pNNN.md`), nothing else. Temp bundles are caller-owned - the tool never deletes a bundle it produced.
314
316
 
315
317
  Excel needs a Python backend with openpyxl, xlrd and pillow - otherwise the call fails with `Remedy: install uv, or pip install openpyxl xlrd pillow`. The Markdown opens with a `## Sheets` table listing every sheet in workbook order (0-based `#`, `worksheet`/`chartsheet`, size, hidden, chart and image counts, rendered view, CSV link for non-empty worksheets), then one section per sheet: a `Data:` line linking the full-content CSV under `sheets/` for non-empty worksheets, chart metadata from the workbook model, embedded images, an optional rendered view, a preview of at most the first 100 rows x 50 columns, and - only when the preview is truncated - a `Columns:` profile (type, non-empty count, min/max, distinct up to 50). Sizes are the extent of non-empty cells (the `info` handle reports the raw worksheet dimensions instead, which may be larger). Rendered views (`images/<stem>-s<idx>.<fmt>`) are produced for sheets carrying charts or images when LibreOffice is on `PATH`: the workbook is exported one PDF page per sheet and rasterized under a 16 Mpx budget. Any LibreOffice or rasterization failure degrades to `Rendered view: unavailable (<reason>)` and a handle note; it never fails the conversion. Workbooks whose chartsheet drawings carry a zero-size anchor (openpyxl-authored files; Excel-authored files are unaffected) render as a degenerate page and are reported as such. `.xls` gets the inventory, CSVs and previews but no visual detection or rendering.
316
318
 
@@ -581,6 +581,9 @@ function ownedCsvPattern(stem) {
581
581
  function ownedPagePattern(stem) {
582
582
  return new RegExp(`^${escRe(stem)}-p\\d+\\.[a-z0-9]+$`);
583
583
  }
584
+ function ownedOcrPattern(stem) {
585
+ return new RegExp(`^${escRe(stem)}-p\\d+(?:\\.words\\.json|\\.md)$`);
586
+ }
584
587
  var FILE_LINK_RE = /\[[^\]]*\]\(\s*((?:sheets|attachments)\/[^)\s]+)\s*\)/g;
585
588
  var IMG_LINK_RE = /!\[[^\]]*\]\(\s*(?:<([^>]*)>|([^)]*?))\s*\)|<img\b[^>]*\bsrc\s*=\s*(?:"([^"]*)"|'([^']*)'|([^\s>"']+))/gi;
586
589
  function imageTarget(m) {
@@ -627,6 +630,7 @@ function openBundle(root, requested, overwrite) {
627
630
  const sheetsDir = join2(root, "sheets");
628
631
  const pagesDir = join2(root, "pages");
629
632
  const attachmentsDir = join2(root, "attachments");
633
+ const ocrDir = join2(root, "ocr");
630
634
  try {
631
635
  if (existsSync(mdPath) && overwrite) {
632
636
  const owned = ownedPattern(stem), ownedCsv = ownedCsvPattern(stem), ownedPage = ownedPagePattern(stem);
@@ -644,6 +648,12 @@ function openBundle(root, requested, overwrite) {
644
648
  if (existsSync(attachmentsDir)) {
645
649
  for (const f of readdirSync(attachmentsDir)) if (attachments.has(f)) rmSync(join2(attachmentsDir, f), { force: true });
646
650
  }
651
+ rmSync(join2(root, `${stem}.pages.json`), { force: true });
652
+ rmSync(join2(root, `${stem}.words.json`), { force: true });
653
+ const ownedOcr = ownedOcrPattern(stem);
654
+ if (existsSync(ocrDir)) {
655
+ for (const f of readdirSync(ocrDir)) if (ownedOcr.test(f)) rmSync(join2(ocrDir, f), { force: true });
656
+ }
647
657
  }
648
658
  const lockId = randomBytes(6).toString("hex");
649
659
  const stagingDir = join2(imagesDir, `.stage-${lockId}`);
@@ -651,7 +661,8 @@ function openBundle(root, requested, overwrite) {
651
661
  mkdirSync2(stagingDir, { recursive: true });
652
662
  const pagesStagingDir = join2(pagesDir, `.stage-${lockId}`);
653
663
  const attachmentsStagingDir = join2(attachmentsDir, `.stage-${lockId}`);
654
- return { root, stem, renamedFrom, renameReason, mdPath, lockPath, imagesDir, stagingDir, lockId, sheetsDir, sheetsStagingDir, pagesDir, pagesStagingDir, attachmentsDir, attachmentsStagingDir, manifest: /* @__PURE__ */ new Set(), csvManifest: /* @__PURE__ */ new Set(), pageManifest: /* @__PURE__ */ new Set(), attachmentManifest: /* @__PURE__ */ new Set(), sourceMap: /* @__PURE__ */ new Map() };
664
+ const ocrStagingDir = join2(ocrDir, `.stage-${lockId}`);
665
+ return { root, stem, renamedFrom, renameReason, mdPath, lockPath, imagesDir, stagingDir, lockId, sheetsDir, sheetsStagingDir, pagesDir, pagesStagingDir, attachmentsDir, attachmentsStagingDir, ocrDir, ocrStagingDir, pageStatsPath: join2(root, `${stem}.pages.json`), wordsPath: join2(root, `${stem}.words.json`), ocrManifest: /* @__PURE__ */ new Set(), manifest: /* @__PURE__ */ new Set(), csvManifest: /* @__PURE__ */ new Set(), pageManifest: /* @__PURE__ */ new Set(), attachmentManifest: /* @__PURE__ */ new Set(), sourceMap: /* @__PURE__ */ new Map() };
655
666
  } catch (e) {
656
667
  rmSync(lockPath, { force: true });
657
668
  throw e;
@@ -715,6 +726,45 @@ function publishPageImages(b, pageCount) {
715
726
  }
716
727
  rmSync(b.pagesStagingDir, { recursive: true, force: true });
717
728
  }
729
+ function writePageStats(b, stats) {
730
+ writeFileSync2(b.pageStatsPath, `${JSON.stringify(stats, null, 2)}
731
+ `, "utf8");
732
+ }
733
+ function publishWords(b) {
734
+ const staged = join2(b.stagingDir, "words.json");
735
+ try {
736
+ if (!existsSync(staged)) throw new Error("child staged no words.json");
737
+ renameSync(staged, b.wordsPath);
738
+ return null;
739
+ } catch (e) {
740
+ rmSync(b.wordsPath, { force: true });
741
+ return `write failed - ${e.message}`;
742
+ }
743
+ }
744
+ function publishSidecars(b) {
745
+ const sidecars = /* @__PURE__ */ new Map(), wordSidecars = /* @__PURE__ */ new Map();
746
+ if (!existsSync(b.ocrStagingDir)) return { sidecars, wordSidecars };
747
+ for (const dir of readdirSync(b.ocrStagingDir).sort()) {
748
+ const m = dir.match(/^p(\d+)$/);
749
+ if (!m) continue;
750
+ const pageDir = join2(b.ocrStagingDir, dir);
751
+ if (!existsSync(join2(pageDir, ".done"))) continue;
752
+ const file = `${b.stem}-${dir}.md`;
753
+ if (!existsSync(join2(pageDir, file))) continue;
754
+ mkdirSync2(b.ocrDir, { recursive: true });
755
+ renameSync(join2(pageDir, file), join2(b.ocrDir, file));
756
+ b.ocrManifest.add(file);
757
+ sidecars.set(Number(m[1]), join2(b.ocrDir, file));
758
+ const wfile = `${b.stem}-${dir}.words.json`;
759
+ if (existsSync(join2(pageDir, wfile))) {
760
+ renameSync(join2(pageDir, wfile), join2(b.ocrDir, wfile));
761
+ b.ocrManifest.add(wfile);
762
+ wordSidecars.set(Number(m[1]), join2(b.ocrDir, wfile));
763
+ }
764
+ }
765
+ rmSync(b.ocrStagingDir, { recursive: true, force: true });
766
+ return { sidecars, wordSidecars };
767
+ }
718
768
  function publishAttachments(b) {
719
769
  if (!existsSync(b.attachmentsStagingDir)) return;
720
770
  for (const f of readdirSync(b.attachmentsStagingDir).sort()) {
@@ -772,6 +822,10 @@ function commitBundle(b, markdown) {
772
822
  fs.rmSync(b.attachmentsStagingDir, { recursive: true, force: true });
773
823
  } catch {
774
824
  }
825
+ try {
826
+ fs.rmSync(b.ocrStagingDir, { recursive: true, force: true });
827
+ } catch {
828
+ }
775
829
  try {
776
830
  fs.rmSync(b.lockPath, { force: true });
777
831
  } catch {
@@ -782,6 +836,10 @@ function abortBundle(b) {
782
836
  for (const f of b.csvManifest) rmSync(join2(b.sheetsDir, f), { force: true });
783
837
  for (const f of b.pageManifest) rmSync(join2(b.pagesDir, f), { force: true });
784
838
  for (const f of b.attachmentManifest) rmSync(join2(b.attachmentsDir, f), { force: true });
839
+ for (const f of b.ocrManifest) rmSync(join2(b.ocrDir, f), { force: true });
840
+ rmSync(b.ocrStagingDir, { recursive: true, force: true });
841
+ rmSync(b.pageStatsPath, { force: true });
842
+ rmSync(b.wordsPath, { force: true });
785
843
  rmSync(`${b.mdPath}.tmp`, { force: true });
786
844
  rmSync(b.stagingDir, { recursive: true, force: true });
787
845
  rmSync(b.sheetsStagingDir, { recursive: true, force: true });
@@ -820,7 +878,21 @@ function compactRanges(nums, maxEntries = 20) {
820
878
  }
821
879
  var INSTALL_HINT = "install Tesseract (see doc/doc-to-md.md), then rerun with ocr=true";
822
880
  var BARE_REASONS = ["fallback tier", "no Python backend"];
881
+ var pageList = (pages) => `page${pages.length === 1 ? "" : "s"} ${pages.join(", ")}`;
882
+ function forcedOcrLine(ocr) {
883
+ const clauses = [];
884
+ if (ocr.pages.length) clauses.push(`sidecars for ${pageList(ocr.pages)}`);
885
+ if (ocr.noText.length) clauses.push(`no text on ${pageList(ocr.noText)}`);
886
+ for (const page of ocr.ocrFailed) clauses.push(`failed on page ${page} (${ocr.ocrErrors[page] ?? "unknown error"})`);
887
+ if (ocr.budgetStopped.length) clauses.push(`budget-stopped ${pageList(ocr.budgetStopped)}`);
888
+ if (ocr.killed !== null) clauses.push(`page ${ocr.killed} killed the OCR child (timeout or crash - likely a compression bomb)`);
889
+ if (ocr.childError !== null) clauses.push(`OCR child failed before processing pages: ${ocr.childError}`);
890
+ if (ocr.notAttempted.length) clauses.push(`${pageList(ocr.notAttempted)} not attempted`);
891
+ const rerun = [...ocr.budgetStopped, ...ocr.notAttempted].sort((a, b) => a - b);
892
+ return `OCR: forced (${ocr.lang}) - ${clauses.join("; ")}${rerun.length ? ` - re-run with --pages ${rerun.join(",")}` : ""}`;
893
+ }
823
894
  function ocrLine(ocr, type) {
895
+ if (ocr.mode === "all") return forcedOcrLine(ocr);
824
896
  const image = type === "image";
825
897
  const rerun = ocr.tesseract ? "rerun with ocr=true" : INSTALL_HINT;
826
898
  switch (ocr.status) {
@@ -896,6 +968,12 @@ function formatHandle(h) {
896
968
  if (h.sheetsDir) lines.push(`Sheets-Dir: ${h.sheetsDir}`);
897
969
  if (h.pagesDir && h.pageImageCount > 0) lines.push(`Pages-Dir: ${h.pagesDir} (${h.pageImageCount} pages)`);
898
970
  else if (h.pageImagesReason) lines.push(`Pages-Dir: none - ${h.pageImagesReason}`);
971
+ if (h.pageStatsPath) lines.push(`Page-Stats: ${h.pageStatsPath}`);
972
+ if (h.wordsPath) {
973
+ const bad = Object.keys(h.wordsErrors).map(Number).sort((a, b) => a - b);
974
+ lines.push(`Words: ${h.wordsPath}${bad.length ? ` (extraction failed for pages ${compactRanges(bad)}: ${h.wordsErrors[bad[0]]})` : ""}`);
975
+ } else if (h.wordsReason) lines.push(`Words: ${h.wordsReason}`);
976
+ if (h.ocrDir) lines.push(`OCR-Dir: ${h.ocrDir}`);
899
977
  lines.push(`Type: ${h.type} Engine: ${h.engine} Tier: ${h.tier}`);
900
978
  lines.push(`Page-Count: ${pageCountLabel(h)} Pages: ${h.pages ? compactRanges(h.pages) : "all"} Images: ${h.imageCount} Size: ${formatSize2(h.bytes)} / ${h.lines} lines`);
901
979
  if (h.degraded) lines.push(`Degraded: ${h.degraded}`);
@@ -945,6 +1023,8 @@ var DOC_TO_MD_OPTIONS = [
945
1023
  { key: "outputDir", flag: "--output-dir", type: "string", default: null, settable: false, help: "Bundle root for <stem>.md + images/; default a per-call temp dir. <stem> = basename without extension with [^A-Za-z0-9._-]+ -> _ (empty -> document); a second call on the same stem writes <stem>-2.md" },
946
1024
  { key: "overwrite", flag: "--overwrite", type: "bool", default: false, settable: false, help: "Replace an existing completed <stem>.md bundle" },
947
1025
  { key: "pageImages", flag: "--page-images", type: "bool", default: false, settable: false, help: "Also render every selected page to pages/<stem>-pNNN.<imageFormat> at imageDpi (PDF, PPTX, .doc, DOCX via LibreOffice); off by default" },
1026
+ { key: "words", flag: "--words", type: "bool", default: false, settable: false, help: 'Write word positions: <stem>.words.json beside the Markdown lists every text-layer word of each selected page with its bbox (PDF points, top-left origin, display orientation; image inputs in source pixels) and the words inline OCR recognized, tagged source "text" or "ocr"; under --ocr-mode all the OCR words go to ocr/<stem>-pNNN.words.json beside each sidecar. Never triggers OCR. PDF and image inputs only.' },
1027
+ { key: "ocrMode", flag: "--ocr-mode", type: "enum", default: "textless", settable: false, enumValues: ["textless", "all"], help: "OCR policy: textless (default) OCRs only pages with an empty text layer, inline; all OCRs every selected page and writes the recognized text to ocr/<stem>-pNNN.md sidecars, leaving the Markdown untouched. all requires --ocr and an explicit --pages selection (PDF, PPTX, DOC)." },
948
1028
  { key: "primaryTimeoutMs", flag: "--primary-timeout", type: "int", default: 6e4, settable: true, env: "PI_DOC_TO_MD_CONVERT_TIMEOUT_MS", help: "pymupdf4llm tier and DOCX child (docx mode); also the unpdf tier" },
949
1029
  { key: "fallbackTimeoutMs", flag: "--fallback-timeout", type: "int", default: 3e4, settable: true, help: "PyMuPDF get_text tier (including DOCX LibreOffice fallback); also PDF and DOCX info and Excel rendered views" },
950
1030
  { key: "sofficeTimeoutMs", flag: "--soffice-timeout", type: "int", default: 12e4, settable: true, env: "PI_DOC_TO_MD_SOFFICE_TIMEOUT_MS", help: "DOCX/PPTX -> PDF via LibreOffice; also Excel rendered views" },
@@ -1051,11 +1131,40 @@ function resolveOptions(perCall, settings, env) {
1051
1131
  out[d.key] = value;
1052
1132
  }
1053
1133
  const o = out;
1054
- if (o.info && (o.pages !== null || o.outputDir !== null || o.overwrite || o.pageImages)) {
1055
- throw new UsageError("--info cannot be combined with --pages, --output-dir, --overwrite or --page-images");
1134
+ if (o.info && (o.pages !== null || o.outputDir !== null || o.overwrite || o.pageImages || o.words || o.ocrMode === "all")) {
1135
+ throw new UsageError("--info cannot be combined with --pages, --output-dir, --overwrite, --page-images, --words or --ocr-mode all");
1056
1136
  }
1057
1137
  return o;
1058
1138
  }
1139
+ var USAGE_PATTERNS = [
1140
+ "Two-pass OCR (PDF, PPTX, DOC):",
1141
+ " 1. <cmd> report.pdf --output-dir out --json",
1142
+ ' -> "pageStatsPath" points at out/report.pages.json; pages with few',
1143
+ ' chars and high imageCoverage are scans. "savedTo" is the Markdown.',
1144
+ " 2. <cmd> report.pdf --output-dir out --ocr --ocr-mode all --pages 2,7 --json",
1145
+ ' -> "ocr"."sidecars" maps 2 and 7 to out/ocr/report-2-p002.md and',
1146
+ " ...-p007.md (a second run in the same dir gets stem report-2);",
1147
+ " the Markdown of this run holds pages 2 and 7 only and equals what",
1148
+ " --pages 2,7 without --ocr would produce. Read sidecars and Markdown",
1149
+ " by the returned paths, never by guessing names.",
1150
+ " 3. --ocr-mode all refuses to run without --ocr and an explicit --pages.",
1151
+ ' The "ocr" object (OCR: line) names failed, budget-stopped, killed and',
1152
+ " not-attempted pages and the exact --pages to re-run.",
1153
+ " Details: doc/doc-to-md.md (bundle contract, failure buckets)."
1154
+ ].join("\n");
1155
+ var usagePatterns = (cmd) => USAGE_PATTERNS.replaceAll("<cmd>", cmd);
1156
+ var BUNDLE_LAYOUT = [
1157
+ { artifact: "<stem>.md", trigger: "always", content: "the Markdown", namedBy: "Saved-To: / savedTo" },
1158
+ { artifact: "images/", trigger: "embedded or extracted figures", content: "image files linked from the Markdown", namedBy: "Images-Dir: / imagesDir" },
1159
+ { artifact: "pages/<stem>-pNNN.<fmt>", trigger: "--page-images", content: "page renders at --image-dpi", namedBy: "Pages-Dir: / pagesDir" },
1160
+ { artifact: "sheets/", trigger: "Excel input", content: "one CSV per non-empty worksheet", namedBy: "Sheets-Dir: / sheetsDir" },
1161
+ { artifact: "attachments/", trigger: "email input", content: "saved attachments", namedBy: "Markdown attachment list" },
1162
+ { artifact: "<stem>.pages.json", trigger: "Python PDF tiers (PDF, PPTX, DOC, DOCX via LibreOffice; not unpdf)", content: "per-page chars, image count, image coverage", namedBy: "Page-Stats: / pageStatsPath" },
1163
+ { artifact: "<stem>.words.json", trigger: "--words", content: "per-page word boxes, source text/ocr", namedBy: "Words: / wordsPath" },
1164
+ { artifact: "ocr/<stem>-pNNN.md", trigger: "--ocr --ocr-mode all", content: "recognized text of a forced page", namedBy: "OCR-Dir: / ocr.sidecars" },
1165
+ { artifact: "ocr/<stem>-pNNN.words.json", trigger: "--ocr --ocr-mode all --words", content: "word boxes of that OCR", namedBy: "ocr.wordSidecars" }
1166
+ ];
1167
+ var bundleLayoutText = () => BUNDLE_LAYOUT.map((r) => ` ${r.artifact.padEnd(28)} ${r.trigger}; ${r.content}; named by ${r.namedBy}`).join("\n");
1059
1168
  function renderHelp() {
1060
1169
  const row = (d) => ` ${(d.flag ?? "<path>").padEnd(26)} ${d.help}${d.default !== null && d.key !== "info" && d.key !== "overwrite" ? ` (default ${d.default})` : ""}`;
1061
1170
  return [
@@ -1067,8 +1176,13 @@ function renderHelp() {
1067
1176
  "Tunables (also settable under quiver.docToMd in settings.json; per-call > settings > default):",
1068
1177
  ...DOC_TO_MD_OPTIONS.filter((d) => d.settable).map(row),
1069
1178
  "",
1070
- "Result: a handle (Saved-To, Images-Dir, Page-Count, Outline ...). Read the Saved-To file for the Markdown.",
1071
- "Exit codes: 0 success, 1 runtime error, 2 usage error."
1179
+ "Result: a handle (Saved-To, Images-Dir, Page-Stats, Page-Count, Outline ...). Read the Saved-To file for the Markdown.",
1180
+ "Exit codes: 0 success, 1 runtime error, 2 usage error.",
1181
+ "",
1182
+ "Bundle layout:",
1183
+ bundleLayoutText(),
1184
+ "",
1185
+ usagePatterns("pi-quiver doc-to-md")
1072
1186
  ].join("\n");
1073
1187
  }
1074
1188
 
@@ -1547,7 +1661,38 @@ function reconcileRenderMarkers(md, renderPages, fmt, sourceMap, reason) {
1547
1661
  if (/<!--rvs?:\d+-->/.test(md)) throw new Error("internal: unresolved render marker");
1548
1662
  return md;
1549
1663
  }
1550
- var emptyOcr = (lang) => ({ status: "off", lang, textless: [], pages: [], noText: [], ocrFailed: [], budgetStopped: [], reason: null, tesseract: null });
1664
+ var emptyOcr = (lang) => ({ status: "off", lang, textless: [], pages: [], noText: [], ocrFailed: [], budgetStopped: [], reason: null, tesseract: null, mode: "textless", sidecars: {}, wordSidecars: {}, ocrErrors: {}, killed: null, notAttempted: [], childError: null });
1665
+ var emptyOutcome = () => ({ written: [], noText: [], ocrFailed: [], ocrErrors: {}, budgetStopped: [], killed: null, notAttempted: [], childError: null });
1666
+ var ocrPagesTag = (n) => `p${String(n).padStart(3, "0")}`;
1667
+ var SIDE_MARKER_RE = /^--- end of page\.page_number=\d+ ---$/;
1668
+ function sidecarHasText(dir) {
1669
+ const file = readdirSync2(dir).find((f) => f.endsWith(".md"));
1670
+ if (!file) return false;
1671
+ return readFileSync2(join3(dir, file), "utf8").split("\n").slice(1).some((l) => l.trim() && !SIDE_MARKER_RE.test(l));
1672
+ }
1673
+ function recoverOcrPages(stagingDir, pages, detail) {
1674
+ const out = emptyOutcome();
1675
+ const activePath = join3(stagingDir, "active");
1676
+ const active = existsSync2(activePath) ? Number(readFileSync2(activePath, "utf8").trim()) || null : null;
1677
+ let sawPage = false;
1678
+ for (const n of pages) {
1679
+ const dir = join3(stagingDir, ocrPagesTag(n));
1680
+ if (existsSync2(join3(dir, ".done"))) {
1681
+ sawPage = true;
1682
+ (sidecarHasText(dir) ? out.written : out.noText).push(n);
1683
+ } else if (existsSync2(join3(dir, ".failed"))) {
1684
+ sawPage = true;
1685
+ out.ocrFailed.push(n);
1686
+ out.ocrErrors[n] = readFileSync2(join3(dir, ".failed"), "utf8").trim() || "unknown error";
1687
+ } else if (n === active) out.killed = n;
1688
+ else out.notAttempted.push(n);
1689
+ }
1690
+ if (active === null && !sawPage) out.childError = detail;
1691
+ return out;
1692
+ }
1693
+ function outcomeFromChild(j) {
1694
+ return { ...emptyOutcome(), written: j.written ?? [], noText: j.noText ?? [], ocrFailed: j.ocrFailed ?? [], ocrErrors: j.ocrErrors ?? {}, budgetStopped: j.budgetStopped ?? [] };
1695
+ }
1551
1696
  var OCR_SENTINEL_RE = /\x00OCR ([^\x00]*)\x00/g;
1552
1697
  function resolveOcrLabels(md, sourceMap) {
1553
1698
  return md.replace(OCR_SENTINEL_RE, (_, key) => {
@@ -1559,7 +1704,7 @@ function handleOcr(tier, type, o, json) {
1559
1704
  if (tier === "unpdf") return o.ocr ? { ...emptyOcr(o.ocrLanguage), status: "unavailable", reason: "no Python backend" } : null;
1560
1705
  const x = json.ocr;
1561
1706
  if (!x || type !== "image" && !x.textless.length && !x.pages.length && !x.ocrFailed.length) return null;
1562
- return x;
1707
+ return { ...emptyOcr(o.ocrLanguage), ...x };
1563
1708
  }
1564
1709
  async function convertDocument(o, signal, seams) {
1565
1710
  const s = { backend: (c) => getBackend(c, void 0, signal), runTier: runTierReal, office: tryConvertOffice, ...seams };
@@ -1568,10 +1713,17 @@ async function convertDocument(o, signal, seams) {
1568
1713
  if (!st || !st.isFile()) throw new Error(`Not a readable file: ${o.path}`);
1569
1714
  if (st.size === 0) throw new Error(`empty file: ${o.path}`);
1570
1715
  const type = classifyInput(inputPath);
1716
+ const forced = o.ocrMode === "all";
1717
+ if (forced) {
1718
+ if (!o.ocr) throw new UsageError("--ocr-mode all requires --ocr");
1719
+ if (o.pages === null) throw new UsageError('--ocr-mode all requires an explicit --pages selection (e.g. --pages 2,7); omitted pages and --pages "" mean all pages and are refused to keep OCR cost bounded');
1720
+ if (type !== "pdf" && type !== "pptx" && type !== "doc") throw new UsageError("--ocr-mode all applies to PDF, PPTX and DOC inputs only (DOCX pages are page-break segments, not PDF pages; convert the DOCX to PDF first)");
1721
+ }
1571
1722
  if (o.pages && (type === "html" || type === "image" || type === "email")) throw new Error(`--pages does not apply to ${type === "html" ? "HTML files" : type === "image" ? "images" : "email"}`);
1572
1723
  const isExcel = type === "xlsx" || type === "xlsm" || type === "xls";
1573
1724
  if (isExcel && o.pages) throw new Error("--pages does not apply to spreadsheets: worksheets have no stable page numbering");
1574
1725
  const backend = await s.backend({ pymupdfVersion: o.pymupdfVersion, warmTimeoutMs: o.warmTimeoutMs });
1726
+ if (forced && backend.kind === "none") throw new Error(`--ocr-mode all cannot run: no Python backend (${backend.reason})`);
1575
1727
  if (isExcel && (backend.kind === "none" || !backend.xlsx)) throw new Error(`Excel conversion needs a Python backend with openpyxl, xlrd and pillow (${backend.kind === "none" ? backend.reason : `${backend.kind} lacks the Excel packages`}). ${EXCEL_REMEDY}`);
1576
1728
  if (type === "email" && extname3(inputPath).toLowerCase() === ".msg" && (backend.kind === "none" || !backend.email)) throw new Error(`MSG conversion needs the extract-msg package. Python backend: ${backendState(backend, "found without extract-msg")}. Remedy: install uv, or pip install extract-msg markdownify into that Python`);
1577
1729
  if (type === "email" && extname3(inputPath).toLowerCase() !== ".msg" && lacksDocx(backend)) throw new Error(`EML conversion needs the Python DOCX/HTML packages (mammoth, markdownify, python-docx). Python backend: ${backendState(backend, "found without mammoth/markdownify/python-docx")}. ${docxRemedy}`);
@@ -1580,11 +1732,12 @@ async function convertDocument(o, signal, seams) {
1580
1732
  let office = null;
1581
1733
  try {
1582
1734
  let pdfPath = inputPath;
1583
- const base = { path: inputPath, pages: o.pages, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
1735
+ const base = { path: inputPath, pages: o.pages, ...o.words && (type === "pdf" || type === "image") ? { words: true } : {}, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr && !forced, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
1584
1736
  let tier, engine, json, degraded = null, fallbackReason = null;
1585
1737
  let explicitBreaks = null;
1586
1738
  let notes = [];
1587
1739
  let officeRoute = null;
1740
+ let copyReason = null;
1588
1741
  if (type === "html") {
1589
1742
  const prepared = await prepareHtml(inputPath, b.stagingDir);
1590
1743
  if (signal?.aborted) throw new Error("aborted");
@@ -1624,6 +1777,7 @@ async function convertDocument(o, signal, seams) {
1624
1777
  }
1625
1778
  if (signal?.aborted) throw new Error("aborted");
1626
1779
  if (json === void 0) {
1780
+ copyReason = reason;
1627
1781
  clearStaging(b);
1628
1782
  const file = `original${extname3(inputPath).toLowerCase()}`;
1629
1783
  const dir = join3(b.stagingDir, "p1");
@@ -1757,6 +1911,36 @@ async function convertDocument(o, signal, seams) {
1757
1911
  if (type === "doc" || officeRoute !== null) degraded = DEGRADED_DOCX_OFFICE;
1758
1912
  if (officeRoute !== null) fallbackReason = fallbackReason ? `${officeRoute}; ${fallbackReason}` : officeRoute;
1759
1913
  if (tier === void 0 || engine === void 0 || json === void 0) throw new Error("internal: no tier produced output");
1914
+ const pageStats = json.pageStats ?? null;
1915
+ if (pageStats) writePageStats(b, pageStats);
1916
+ let wordsPath = null, wordsReason = null;
1917
+ const wordsErrors = {};
1918
+ const takeWordsErrors = (j) => {
1919
+ for (const [k, v] of Object.entries(j?.wordsErrors ?? {})) {
1920
+ const page = Number(k);
1921
+ if (Number.isFinite(page)) wordsErrors[page] = v;
1922
+ }
1923
+ };
1924
+ if (o.words) {
1925
+ if (type !== "pdf" && type !== "image") wordsReason = `none - word positions apply to PDF and image inputs only (${type})`;
1926
+ else if (tier === "unpdf") wordsReason = "none - unpdf tier has no page geometry";
1927
+ else if (engine === "copy") wordsReason = `none - image copied without conversion (${copyReason})`;
1928
+ else if (json?.words === true) {
1929
+ wordsReason = publishWords(b);
1930
+ if (wordsReason === null) wordsPath = b.wordsPath;
1931
+ } else wordsReason = `write failed - ${json?.wordsErrors?.file ?? "child reported no words document"}`;
1932
+ takeWordsErrors(json);
1933
+ }
1934
+ let ocr = handleOcr(tier, type, o, json);
1935
+ if (forced) {
1936
+ const r = await s.runTier("ocr-pages", { path: pdfPath, pages: o.pages, ...o.words && type === "pdf" ? { words: true } : {}, stem: b.stem, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs, stagingDir: b.ocrStagingDir, dpi: o.imageDpi, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion }, b, signal, o.primaryTimeoutMs, backend);
1937
+ if (signal?.aborted || !r.ok && "reason" in r && r.reason === "aborted") throw new Error("aborted");
1938
+ if (r.ok && r.json.status === "unavailable") throw new Error(`OCR unavailable: ${r.json.reason} (install Tesseract; see doc/doc-to-md.md)`);
1939
+ const outcome = r.ok && r.json.status === "ran" ? outcomeFromChild(r.json) : recoverOcrPages(b.ocrStagingDir, o.pages, !r.ok ? "userError" in r ? r.userError : `${r.reason}${detailSuffix(r)}` : "malformed child output");
1940
+ const { sidecars, wordSidecars } = publishSidecars(b);
1941
+ ocr = { ...emptyOcr(o.ocrLanguage), ...json.ocr, status: "ran", reason: null, mode: "all", pages: outcome.written, noText: outcome.noText, ocrFailed: outcome.ocrFailed, ocrErrors: outcome.ocrErrors, budgetStopped: outcome.budgetStopped, killed: outcome.killed, notAttempted: outcome.notAttempted, childError: outcome.childError, sidecars: Object.fromEntries(sidecars), wordSidecars: Object.fromEntries(wordSidecars) };
1942
+ if (o.words && r.ok) takeWordsErrors(r.json);
1943
+ }
1760
1944
  if (!isExcel) notes = [...notes, ...json.notes ?? []];
1761
1945
  if (b.renamedFrom) notes.splice(notes[0]?.startsWith("preview truncated:") ? 1 : 0, 0, `renamed to ${b.stem} (${b.renameReason})`);
1762
1946
  const pageImagesReason = !o.pageImages || b.pageManifest.size || tier === "primary" || tier === "fallback" ? null : tier === "unpdf" ? "page images need the Python backend" : `${type} has no page geometry`;
@@ -1773,7 +1957,7 @@ async function convertDocument(o, signal, seams) {
1773
1957
  ` : "") + body;
1774
1958
  commitBundle(b, markdown);
1775
1959
  const outline = scanOutline(markdown, o.outlineMaxEntries);
1776
- const details = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr: handleOcr(tier, type, o, json) };
1960
+ const details = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr, pageStats, pageStatsPath: pageStats ? b.pageStatsPath : null, ocrDir: b.ocrManifest.size ? b.ocrDir : null, wordsPath, wordsReason, wordsErrors };
1777
1961
  return { output: formatHandle(details), details };
1778
1962
  } catch (e) {
1779
1963
  abortBundle(b);
@@ -1819,7 +2003,7 @@ async function inspectDocument(o, signal, seams) {
1819
2003
  }
1820
2004
 
1821
2005
  // bin/pi-quiver.ts
1822
- var USAGE = 'Usage: pi-quiver fetch <url> [--method GET|HEAD|POST] [--header "K: V"]... [--body <str>] [--raw] [--timeout-ms <n>]\n pi-quiver doc-to-md [--json] [--info] [--page-images] [--pages <spec>] [--output-dir <dir>] [--overwrite] [tunable flags] <path> (--help for all flags)';
2006
+ var USAGE = 'Usage: pi-quiver fetch <url> [--method GET|HEAD|POST] [--header "K: V"]... [--body <str>] [--raw] [--timeout-ms <n>]\n pi-quiver doc-to-md [--json] [--info] [--page-images] [--words] [--pages <spec>] [--output-dir <dir>] [--overwrite] [--ocr] [--ocr-mode textless|all] [tunable flags] <path> (--help for all flags)';
1823
2007
  function parseDocToMd(rest) {
1824
2008
  if (rest.includes("--help") || rest.includes("-h")) return { ok: true, cmd: "doc-to-md-help" };
1825
2009
  const json = rest.includes("--json");
@@ -1974,6 +2158,12 @@ ${USAGE}
1974
2158
  `);
1975
2159
  return 0;
1976
2160
  } catch (err) {
2161
+ if (err instanceof UsageError) {
2162
+ process.stderr.write(`${err.message}
2163
+ ${USAGE}
2164
+ `);
2165
+ return 2;
2166
+ }
1977
2167
  process.stderr.write(`doc-to-md failed: ${err instanceof Error ? err.message : String(err)}
1978
2168
  `);
1979
2169
  return 1;
@@ -35,7 +35,7 @@ export default function docToMdExtension(pi: ExtensionAPI) {
35
35
  name: "doc_to_md",
36
36
  label: "Convert doc to Markdown bundle",
37
37
  description:
38
- "Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, email (.msg/.eml), HTML (.html/.htm) or image (.png .jpg .jpeg .tif .tiff .bmp .gif) to a Markdown bundle on disk and return a handle (Saved-To, Images-Dir, Page-Count, Outline, diagnostics) - the Markdown itself is never inlined; read the Saved-To file (offset/limit) for content. `info: true` returns page count, metadata and TOC (or the sheet inventory) without converting - use it to pick `pages`. `pages` selects inclusive 1-based pages (PDF/PPTX) or explicit-page-break segments (DOCX; rejected when the file has none); every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Images are always extracted into images/ with relative links. Primary engine pymupdf4llm, fallback PyMuPDF text (degraded, marked), pure-JS unpdf only when no Python backend exists. Excel yields a sheet inventory (worksheets + chartsheets), a full CSV per non-empty worksheet under sheets/, a bounded preview with formulas and cached values, merged/hidden disclosure, and rendered chart views when LibreOffice is available. DOCX converts directly (mammoth; python-docx text fallback) and keeps heading styles, so the handle Outline lists `L<line>` and `p<segment>` per heading; LibreOffice (soffice) is optional for DOCX (fallback route, degraded) and Excel rendered views, and required for PPTX. Use `outputDir` for a durable bundle; without it the bundle lands in a per-call temp dir. Input must be a local file path (use fetch first for URLs). HTML converts without Readability (markdownify; Turndown fallback): local and data: images are copied into images/, remote images stay links. Pages without a text layer keep their picture in images/, or pages/ with `pageImages: true`; image inputs keep their picture in images/. `pageImages: true` writes page renders to pages/<stem>-pNNN, and email attachments are saved under attachments/. OCR is off by default: `ocr: true` adds recognized text (labeled as OCR) when Tesseract language data is installed; the handle's OCR: line says what ran.",
38
+ "Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, email (.msg/.eml), HTML (.html/.htm) or image (.png .jpg .jpeg .tif .tiff .bmp .gif) to a Markdown bundle on disk and return a handle (Saved-To, Images-Dir, Page-Count, Outline, diagnostics) - the Markdown itself is never inlined; read the Saved-To file (offset/limit) for content. `info: true` returns page count, metadata and TOC (or the sheet inventory) without converting - use it to pick `pages`. `pages` selects inclusive 1-based pages (PDF/PPTX) or explicit-page-break segments (DOCX; rejected when the file has none); every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Images are always extracted into images/ with relative links. Primary engine pymupdf4llm, fallback PyMuPDF text (degraded, marked), pure-JS unpdf only when no Python backend exists. Excel yields a sheet inventory (worksheets + chartsheets), a full CSV per non-empty worksheet under sheets/, a bounded preview with formulas and cached values, merged/hidden disclosure, and rendered chart views when LibreOffice is available. DOCX converts directly (mammoth; python-docx text fallback) and keeps heading styles, so the handle Outline lists `L<line>` and `p<segment>` per heading; LibreOffice (soffice) is optional for DOCX (fallback route, degraded) and Excel rendered views, and required for PPTX. Use `outputDir` for a durable bundle; without it the bundle lands in a per-call temp dir. Input must be a local file path (use fetch first for URLs). HTML converts without Readability (markdownify; Turndown fallback): local and data: images are copied into images/, remote images stay links. Pages without a text layer keep their picture in images/, or pages/ with `pageImages: true`; image inputs keep their picture in images/. `pageImages: true` writes page renders to pages/<stem>-pNNN, and email attachments are saved under attachments/. OCR is off by default: `ocr: true` adds recognized text (labeled as OCR) when Tesseract language data is installed; the handle's OCR: line says what ran. Two-pass OCR: read Page-Stats first, then re-run with ocr: true, ocrMode: \"all\" and an explicit pages selection to get ocr/ sidecars for the pages you name.",
39
39
  promptSnippet: "Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS/MSG/EML/HTML/image to a Markdown bundle (handle returned; read Saved-To)",
40
40
  parameters: buildSchema(),
41
41
 
@@ -1,23 +1,26 @@
1
1
  /**
2
2
  * Bundle protocol: a call owns `<stem>` for its whole duration via `<stem>.md.lock`;
3
- * children stage assets under `images/`, `sheets/`, `pages/`, and `attachments/` staging dirs;
4
- * Node publishes them to stem-prefixed files in those four dirs and records every file it
3
+ * children stage assets under `images/`, `sheets/`, `pages/`, `attachments/`, and `ocr/` staging dirs;
4
+ * Node publishes them to stem-prefixed files in those asset dirs and records every file it
5
5
  * wrote in a manifest, and commits `<stem>.md` atomically (tmp + rename).
6
6
  */
7
7
  import fs, { closeSync, existsSync, mkdirSync, openSync, readdirSync, readFileSync, renameSync, rmSync, statSync, unlinkSync, writeFileSync } from "node:fs";
8
8
  import { tmpdir } from "node:os";
9
9
  import { extname, join, resolve } from "node:path";
10
10
  import { randomBytes } from "node:crypto";
11
+ import type { PageStat } from "./doc-to-md-handle.ts";
11
12
 
12
13
  export interface Bundle {
13
14
  root: string; stem: string; renamedFrom: string | null; renameReason: string | null; mdPath: string; lockPath: string; imagesDir: string; stagingDir: string; lockId: string;
14
15
  sheetsDir: string; sheetsStagingDir: string;
15
16
  pagesDir: string; pagesStagingDir: string;
16
17
  attachmentsDir: string; attachmentsStagingDir: string;
18
+ ocrDir: string; ocrStagingDir: string; pageStatsPath: string; wordsPath: string;
17
19
  manifest: Set<string>;
18
20
  csvManifest: Set<string>;
19
21
  pageManifest: Set<string>;
20
22
  attachmentManifest: Set<string>;
23
+ ocrManifest: Set<string>;
21
24
  sourceMap: Map<string, string>;
22
25
  }
23
26
 
@@ -33,6 +36,8 @@ export function ownedCsvPattern(stem: string): RegExp {
33
36
 
34
37
  export function ownedPagePattern(stem: string): RegExp { return new RegExp(`^${escRe(stem)}-p\\d+\\.[a-z0-9]+$`); }
35
38
 
39
+ export function ownedOcrPattern(stem: string): RegExp { return new RegExp(`^${escRe(stem)}-p\\d+(?:\\.words\\.json|\\.md)$`); }
40
+
36
41
  const FILE_LINK_RE = /\[[^\]]*\]\(\s*((?:sheets|attachments)\/[^)\s]+)\s*\)/g;
37
42
  const IMG_LINK_RE = /!\[[^\]]*\]\(\s*(?:<([^>]*)>|([^)]*?))\s*\)|<img\b[^>]*\bsrc\s*=\s*(?:"([^"]*)"|'([^']*)'|([^\s>"']+))/gi;
38
43
 
@@ -75,6 +80,7 @@ export function openBundle(root: string, requested: string, overwrite: boolean):
75
80
  const sheetsDir = join(root, "sheets");
76
81
  const pagesDir = join(root, "pages");
77
82
  const attachmentsDir = join(root, "attachments");
83
+ const ocrDir = join(root, "ocr");
78
84
  try {
79
85
  if (existsSync(mdPath) && overwrite) {
80
86
  const owned = ownedPattern(stem), ownedCsv = ownedCsvPattern(stem), ownedPage = ownedPagePattern(stem);
@@ -87,6 +93,10 @@ export function openBundle(root: string, requested: string, overwrite: boolean):
87
93
  if (existsSync(sheetsDir)) for (const f of readdirSync(sheetsDir)) if (ownedCsv.test(f)) rmSync(join(sheetsDir, f), { force: true });
88
94
  if (existsSync(pagesDir)) for (const f of readdirSync(pagesDir)) if (ownedPage.test(f)) rmSync(join(pagesDir, f), { force: true });
89
95
  if (existsSync(attachmentsDir)) for (const f of readdirSync(attachmentsDir)) if (attachments.has(f)) rmSync(join(attachmentsDir, f), { force: true });
96
+ rmSync(join(root, `${stem}.pages.json`), { force: true });
97
+ rmSync(join(root, `${stem}.words.json`), { force: true });
98
+ const ownedOcr = ownedOcrPattern(stem);
99
+ if (existsSync(ocrDir)) for (const f of readdirSync(ocrDir)) if (ownedOcr.test(f)) rmSync(join(ocrDir, f), { force: true });
90
100
  }
91
101
  const lockId = randomBytes(6).toString("hex");
92
102
  const stagingDir = join(imagesDir, `.stage-${lockId}`);
@@ -94,7 +104,8 @@ export function openBundle(root: string, requested: string, overwrite: boolean):
94
104
  mkdirSync(stagingDir, { recursive: true });
95
105
  const pagesStagingDir = join(pagesDir, `.stage-${lockId}`);
96
106
  const attachmentsStagingDir = join(attachmentsDir, `.stage-${lockId}`);
97
- return { root, stem, renamedFrom, renameReason, mdPath, lockPath, imagesDir, stagingDir, lockId, sheetsDir, sheetsStagingDir, pagesDir, pagesStagingDir, attachmentsDir, attachmentsStagingDir, manifest: new Set(), csvManifest: new Set(), pageManifest: new Set(), attachmentManifest: new Set(), sourceMap: new Map() };
107
+ const ocrStagingDir = join(ocrDir, `.stage-${lockId}`);
108
+ return { root, stem, renamedFrom, renameReason, mdPath, lockPath, imagesDir, stagingDir, lockId, sheetsDir, sheetsStagingDir, pagesDir, pagesStagingDir, attachmentsDir, attachmentsStagingDir, ocrDir, ocrStagingDir, pageStatsPath: join(root, `${stem}.pages.json`), wordsPath: join(root, `${stem}.words.json`), ocrManifest: new Set(), manifest: new Set(), csvManifest: new Set(), pageManifest: new Set(), attachmentManifest: new Set(), sourceMap: new Map() };
98
109
  } catch (e) { rmSync(lockPath, { force: true }); throw e; }
99
110
  }
100
111
 
@@ -160,6 +171,49 @@ export function publishPageImages(b: Bundle, pageCount: number): void {
160
171
  rmSync(b.pagesStagingDir, { recursive: true, force: true });
161
172
  }
162
173
 
174
+ export function writePageStats(b: Bundle, stats: PageStat[]): void {
175
+ writeFileSync(b.pageStatsPath, `${JSON.stringify(stats, null, 2)}\n`, "utf8");
176
+ }
177
+
178
+ /** Rename within the bundle root keeps the complete words document atomic. */
179
+ export function publishWords(b: Bundle): string | null {
180
+ const staged = join(b.stagingDir, "words.json");
181
+ try {
182
+ if (!existsSync(staged)) throw new Error("child staged no words.json");
183
+ renameSync(staged, b.wordsPath);
184
+ return null;
185
+ } catch (e) {
186
+ rmSync(b.wordsPath, { force: true });
187
+ return `write failed - ${(e as Error).message}`;
188
+ }
189
+ }
190
+
191
+ /** Move `.done`-gated Markdown and words sidecars into `ocr/`; drop partial dirs and the checkpoint. Returns page -> absolute paths for each kind. */
192
+ export function publishSidecars(b: Bundle): { sidecars: Map<number, string>; wordSidecars: Map<number, string> } {
193
+ const sidecars = new Map<number, string>(), wordSidecars = new Map<number, string>();
194
+ if (!existsSync(b.ocrStagingDir)) return { sidecars, wordSidecars };
195
+ for (const dir of readdirSync(b.ocrStagingDir).sort()) {
196
+ const m = dir.match(/^p(\d+)$/);
197
+ if (!m) continue;
198
+ const pageDir = join(b.ocrStagingDir, dir);
199
+ if (!existsSync(join(pageDir, ".done"))) continue;
200
+ const file = `${b.stem}-${dir}.md`;
201
+ if (!existsSync(join(pageDir, file))) continue;
202
+ mkdirSync(b.ocrDir, { recursive: true });
203
+ renameSync(join(pageDir, file), join(b.ocrDir, file));
204
+ b.ocrManifest.add(file);
205
+ sidecars.set(Number(m[1]), join(b.ocrDir, file));
206
+ const wfile = `${b.stem}-${dir}.words.json`;
207
+ if (existsSync(join(pageDir, wfile))) {
208
+ renameSync(join(pageDir, wfile), join(b.ocrDir, wfile));
209
+ b.ocrManifest.add(wfile);
210
+ wordSidecars.set(Number(m[1]), join(b.ocrDir, wfile));
211
+ }
212
+ }
213
+ rmSync(b.ocrStagingDir, { recursive: true, force: true });
214
+ return { sidecars, wordSidecars };
215
+ }
216
+
163
217
  export function publishAttachments(b: Bundle): void {
164
218
  if (!existsSync(b.attachmentsStagingDir)) return;
165
219
  for (const f of readdirSync(b.attachmentsStagingDir).sort()) {
@@ -208,6 +262,7 @@ export function commitBundle(b: Bundle, markdown: string): void {
208
262
  try { fs.rmSync(b.sheetsStagingDir, { recursive: true, force: true }); } catch { /* best-effort */ }
209
263
  try { fs.rmSync(b.pagesStagingDir, { recursive: true, force: true }); } catch { /* best-effort */ }
210
264
  try { fs.rmSync(b.attachmentsStagingDir, { recursive: true, force: true }); } catch { /* best-effort */ }
265
+ try { fs.rmSync(b.ocrStagingDir, { recursive: true, force: true }); } catch { /* best-effort */ }
211
266
  try { fs.rmSync(b.lockPath, { force: true }); } catch { /* Markdown is published; cleanup is best-effort. */ }
212
267
  }
213
268
 
@@ -216,6 +271,10 @@ export function abortBundle(b: Bundle): void {
216
271
  for (const f of b.csvManifest) rmSync(join(b.sheetsDir, f), { force: true });
217
272
  for (const f of b.pageManifest) rmSync(join(b.pagesDir, f), { force: true });
218
273
  for (const f of b.attachmentManifest) rmSync(join(b.attachmentsDir, f), { force: true });
274
+ for (const f of b.ocrManifest) rmSync(join(b.ocrDir, f), { force: true });
275
+ rmSync(b.ocrStagingDir, { recursive: true, force: true });
276
+ rmSync(b.pageStatsPath, { force: true });
277
+ rmSync(b.wordsPath, { force: true });
219
278
  rmSync(`${b.mdPath}.tmp`, { force: true });
220
279
  rmSync(b.stagingDir, { recursive: true, force: true });
221
280
  rmSync(b.sheetsStagingDir, { recursive: true, force: true });
@@ -12,9 +12,9 @@ import { type ChildProcess, spawn } from "node:child_process";
12
12
  import { homedir, tmpdir } from "node:os";
13
13
  import { basename, dirname, extname, isAbsolute, join, resolve } from "node:path";
14
14
  import { fileURLToPath, pathToFileURL } from "node:url";
15
- import { type Bundle, abortBundle, commitBundle, openBundle, publishAttachments, publishPageImages, publishSheetCsvs, publishSheetImages, publishStaged, rewriteLinks, tempBundleRoot, validateImageLinks } from "./doc-to-md-bundle.ts";
16
- import { type Engine, type HandleData, type InfoData, type OcrInfo, type Tier, formatHandle, formatInfoHandle, scanOutline, type SheetInfo, type TocEntry } from "./doc-to-md-handle.ts";
17
- import { type DocToMdOptions, type InputType, IMAGE_EXTS, TUNABLE_DEFAULTS, classifyInput, sanitizeStem } from "./doc-to-md-options.ts";
15
+ import { type Bundle, abortBundle, commitBundle, openBundle, publishAttachments, publishSidecars, publishWords, writePageStats, publishPageImages, publishSheetCsvs, publishSheetImages, publishStaged, rewriteLinks, tempBundleRoot, validateImageLinks } from "./doc-to-md-bundle.ts";
16
+ import { type Engine, type HandleData, type InfoData, type OcrInfo, type PageStat, type Tier, formatHandle, formatInfoHandle, scanOutline, type SheetInfo, type TocEntry } from "./doc-to-md-handle.ts";
17
+ import { type DocToMdOptions, type InputType, IMAGE_EXTS, TUNABLE_DEFAULTS, UsageError, classifyInput, sanitizeStem } from "./doc-to-md-options.ts";
18
18
 
19
19
  export * from "./doc-to-md-options.ts";
20
20
  export { compactRanges, formatHandle, formatInfoHandle, formatSize, scanOutline } from "./doc-to-md-handle.ts";
@@ -466,8 +466,8 @@ const lacksDocx = (b: Backend) => b.kind === "none" || !b.docx;
466
466
  const clearStaging = (b: Pick<Bundle, "stagingDir">) => { for (const f of readdirSync(b.stagingDir)) rmSync(join(b.stagingDir, f), { recursive: true, force: true }); };
467
467
  export const EXCEL_REMEDY = "Remedy: install uv, or pip install openpyxl xlrd pillow";
468
468
 
469
- export type Mode = "html" | "image" | "info" | "pdf-primary" | "pdf-fallback" | "xlsx" | "pdf-text" | "render-pages" | "docx" | "email";
470
- export interface TierJson { pageImages?: { page: number; file: string }[]; ocr?: OcrInfo; markdown?: string; pages?: number[]; pageCount?: number; emptyPages?: number[]; failedPages?: { page: number; error: string }[]; notes?: string[]; images?: { sheetIndex: number; file: string }[]; metadata?: Record<string, string>; toc?: [number, string, number | null][]; explicitBreaks?: number; engine?: string; degraded?: boolean; fallbackReason?: string | null; sheets?: SheetInfo[]; renderPages?: number[]; sheetCount?: number; ok?: boolean; reason?: string; rendered?: { idx: number; file: string; dpi: number }[]; failed?: { idx: number; reason: string }[]; }
469
+ export type Mode = "html" | "image" | "info" | "pdf-primary" | "pdf-fallback" | "xlsx" | "pdf-text" | "render-pages" | "docx" | "email" | "ocr-pages";
470
+ export interface TierJson { words?: boolean; wordsErrors?: Record<string, string>; pageStats?: PageStat[]; status?: string; written?: number[]; noText?: number[]; ocrFailed?: number[]; ocrErrors?: Record<string, string>; budgetStopped?: number[]; pageImages?: { page: number; file: string }[]; ocr?: OcrInfo; markdown?: string; pages?: number[]; pageCount?: number; emptyPages?: number[]; failedPages?: { page: number; error: string }[]; notes?: string[]; images?: { sheetIndex: number; file: string }[]; metadata?: Record<string, string>; toc?: [number, string, number | null][]; explicitBreaks?: number; engine?: string; degraded?: boolean; fallbackReason?: string | null; sheets?: SheetInfo[]; renderPages?: number[]; sheetCount?: number; ok?: boolean; reason?: string; rendered?: { idx: number; file: string; dpi: number }[]; failed?: { idx: number; reason: string }[]; }
471
471
  export type TierResult = { ok: true; json: TierJson } | { ok: false; reason: string; detail?: string } | { ok: false; userError: string; pageCount?: number };
472
472
 
473
473
  export interface PipelineSeams {
@@ -490,7 +490,7 @@ function unpdfWorkerPath(): string {
490
490
  ], existsSync);
491
491
  }
492
492
 
493
- async function runTierReal(mode: Mode, childOptions: Record<string, unknown>, _b: Pick<Bundle, "stagingDir">, signal: AbortSignal | undefined, timeoutMs: number, backend: Backend): Promise<TierResult> {
493
+ export async function runTierReal(mode: Mode, childOptions: Record<string, unknown>, _b: Pick<Bundle, "stagingDir">, signal: AbortSignal | undefined, timeoutMs: number, backend: Backend): Promise<TierResult> {
494
494
  const cfg = { pymupdfVersion: String(childOptions.pymupdfVersion), warmTimeoutMs: 0 };
495
495
  let cmd: string, args: string[];
496
496
  if (mode === "pdf-text" || (mode === "info" && backend.kind === "none")) { cmd = process.execPath; args = [unpdfWorkerPath(), mode]; }
@@ -536,7 +536,39 @@ export function reconcileRenderMarkers(md: string, renderPages: number[], fmt: s
536
536
  return md;
537
537
  }
538
538
 
539
- export const emptyOcr = (lang: string): OcrInfo => ({ status: "off", lang, textless: [], pages: [], noText: [], ocrFailed: [], budgetStopped: [], reason: null, tesseract: null });
539
+ export const emptyOcr = (lang: string): OcrInfo => ({ status: "off", lang, textless: [], pages: [], noText: [], ocrFailed: [], budgetStopped: [], reason: null, tesseract: null, mode: "textless", sidecars: {}, wordSidecars: {}, ocrErrors: {}, killed: null, notAttempted: [], childError: null });
540
+
541
+ export interface OcrPagesOutcome { written: number[]; noText: number[]; ocrFailed: number[]; ocrErrors: Record<number, string>; budgetStopped: number[]; killed: number | null; notAttempted: number[]; childError: string | null; }
542
+ const emptyOutcome = (): OcrPagesOutcome => ({ written: [], noText: [], ocrFailed: [], ocrErrors: {}, budgetStopped: [], killed: null, notAttempted: [], childError: null });
543
+ const ocrPagesTag = (n: number) => `p${String(n).padStart(3, "0")}`;
544
+ const SIDE_MARKER_RE = /^--- end of page\.page_number=\d+ ---$/;
545
+
546
+ function sidecarHasText(dir: string): boolean {
547
+ const file = readdirSync(dir).find((f) => f.endsWith(".md"));
548
+ if (!file) return false;
549
+ return readFileSync(join(dir, file), "utf8").split("\n").slice(1).some((l) => l.trim() && !SIDE_MARKER_RE.test(l));
550
+ }
551
+
552
+ /** Rebuild the outcome from staging markers after the child died or returned garbage. */
553
+ export function recoverOcrPages(stagingDir: string, pages: number[], detail: string): OcrPagesOutcome {
554
+ const out = emptyOutcome();
555
+ const activePath = join(stagingDir, "active");
556
+ const active = existsSync(activePath) ? Number(readFileSync(activePath, "utf8").trim()) || null : null;
557
+ let sawPage = false;
558
+ for (const n of pages) {
559
+ const dir = join(stagingDir, ocrPagesTag(n));
560
+ if (existsSync(join(dir, ".done"))) { sawPage = true; (sidecarHasText(dir) ? out.written : out.noText).push(n); }
561
+ else if (existsSync(join(dir, ".failed"))) { sawPage = true; out.ocrFailed.push(n); out.ocrErrors[n] = readFileSync(join(dir, ".failed"), "utf8").trim() || "unknown error"; }
562
+ else if (n === active) out.killed = n;
563
+ else out.notAttempted.push(n);
564
+ }
565
+ if (active === null && !sawPage) out.childError = detail;
566
+ return out;
567
+ }
568
+
569
+ function outcomeFromChild(j: TierJson): OcrPagesOutcome {
570
+ return { ...emptyOutcome(), written: j.written ?? [], noText: j.noText ?? [], ocrFailed: j.ocrFailed ?? [], ocrErrors: j.ocrErrors ?? {}, budgetStopped: j.budgetStopped ?? [] };
571
+ }
540
572
  const OCR_SENTINEL_RE = /\x00OCR ([^\x00]*)\x00/g;
541
573
 
542
574
  /** Child OCR labels name staged files; the published name exists only after publishStaged. */
@@ -551,7 +583,7 @@ function handleOcr(tier: Tier, type: InputType, o: DocToMdOptions, json: TierJso
551
583
  if (tier === "unpdf") return o.ocr ? { ...emptyOcr(o.ocrLanguage), status: "unavailable", reason: "no Python backend" } : null;
552
584
  const x = json.ocr;
553
585
  if (!x || (type !== "image" && !x.textless.length && !x.pages.length && !x.ocrFailed.length)) return null;
554
- return x;
586
+ return { ...emptyOcr(o.ocrLanguage), ...x };
555
587
  }
556
588
 
557
589
  export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, seams?: Partial<PipelineSeams>): Promise<ConvertOutcome> {
@@ -561,10 +593,17 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
561
593
  if (!st || !st.isFile()) throw new Error(`Not a readable file: ${o.path}`);
562
594
  if (st.size === 0) throw new Error(`empty file: ${o.path}`);
563
595
  const type = classifyInput(inputPath);
596
+ const forced = o.ocrMode === "all";
597
+ if (forced) {
598
+ if (!o.ocr) throw new UsageError("--ocr-mode all requires --ocr");
599
+ if (o.pages === null) throw new UsageError('--ocr-mode all requires an explicit --pages selection (e.g. --pages 2,7); omitted pages and --pages "" mean all pages and are refused to keep OCR cost bounded');
600
+ if (type !== "pdf" && type !== "pptx" && type !== "doc") throw new UsageError("--ocr-mode all applies to PDF, PPTX and DOC inputs only (DOCX pages are page-break segments, not PDF pages; convert the DOCX to PDF first)");
601
+ }
564
602
  if (o.pages && (type === "html" || type === "image" || type === "email")) throw new Error(`--pages does not apply to ${type === "html" ? "HTML files" : type === "image" ? "images" : "email"}`);
565
603
  const isExcel = type === "xlsx" || type === "xlsm" || type === "xls";
566
604
  if (isExcel && o.pages) throw new Error("--pages does not apply to spreadsheets: worksheets have no stable page numbering");
567
605
  const backend = await s.backend({ pymupdfVersion: o.pymupdfVersion, warmTimeoutMs: o.warmTimeoutMs });
606
+ if (forced && backend.kind === "none") throw new Error(`--ocr-mode all cannot run: no Python backend (${backend.reason})`);
568
607
  if (isExcel && (backend.kind === "none" || !backend.xlsx)) throw new Error(`Excel conversion needs a Python backend with openpyxl, xlrd and pillow (${backend.kind === "none" ? backend.reason : `${backend.kind} lacks the Excel packages`}). ${EXCEL_REMEDY}`);
569
608
  if (type === "email" && extname(inputPath).toLowerCase() === ".msg" && (backend.kind === "none" || !backend.email)) throw new Error(`MSG conversion needs the extract-msg package. Python backend: ${backendState(backend, "found without extract-msg")}. Remedy: install uv, or pip install extract-msg markdownify into that Python`);
570
609
  if (type === "email" && extname(inputPath).toLowerCase() !== ".msg" && lacksDocx(backend)) throw new Error(`EML conversion needs the Python DOCX/HTML packages (mammoth, markdownify, python-docx). Python backend: ${backendState(backend, "found without mammoth/markdownify/python-docx")}. ${docxRemedy}`);
@@ -573,11 +612,12 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
573
612
  let office: { pdfPath: string; cleanup: () => void } | null = null;
574
613
  try {
575
614
  let pdfPath = inputPath;
576
- const base = { path: inputPath, pages: o.pages, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
615
+ const base = { path: inputPath, pages: o.pages, ...(o.words && (type === "pdf" || type === "image") ? { words: true } : {}), stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr && !forced, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
577
616
  let tier: Tier | undefined, engine: Engine | undefined, json: TierJson | undefined, degraded: string | null = null, fallbackReason: string | null = null;
578
617
  let explicitBreaks: number | null = null;
579
618
  let notes: string[] = [];
580
619
  let officeRoute: string | null = null;
620
+ let copyReason: string | null = null;
581
621
  if (type === "html") {
582
622
  const prepared = await prepareHtml(inputPath, b.stagingDir);
583
623
  if (signal?.aborted) throw new Error("aborted");
@@ -603,6 +643,7 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
603
643
  }
604
644
  if (signal?.aborted) throw new Error("aborted");
605
645
  if (json === undefined) {
646
+ copyReason = reason;
606
647
  clearStaging(b);
607
648
  const file = `original${extname(inputPath).toLowerCase()}`;
608
649
  const dir = join(b.stagingDir, "p1");
@@ -704,6 +745,35 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
704
745
  if (type === "doc" || officeRoute !== null) degraded = DEGRADED_DOCX_OFFICE;
705
746
  if (officeRoute !== null) fallbackReason = fallbackReason ? `${officeRoute}; ${fallbackReason}` : officeRoute;
706
747
  if (tier === undefined || engine === undefined || json === undefined) throw new Error("internal: no tier produced output");
748
+ const pageStats = json.pageStats ?? null;
749
+ if (pageStats) writePageStats(b, pageStats);
750
+ let wordsPath: string | null = null, wordsReason: string | null = null;
751
+ const wordsErrors: Record<number, string> = {};
752
+ const takeWordsErrors = (j: TierJson | undefined) => {
753
+ for (const [k, v] of Object.entries(j?.wordsErrors ?? {})) {
754
+ const page = Number(k);
755
+ if (Number.isFinite(page)) wordsErrors[page] = v;
756
+ }
757
+ };
758
+ if (o.words) {
759
+ if (type !== "pdf" && type !== "image") wordsReason = `none - word positions apply to PDF and image inputs only (${type})`;
760
+ else if (tier === "unpdf") wordsReason = "none - unpdf tier has no page geometry";
761
+ else if (engine === "copy") wordsReason = `none - image copied without conversion (${copyReason})`;
762
+ else if (json?.words === true) { wordsReason = publishWords(b); if (wordsReason === null) wordsPath = b.wordsPath; }
763
+ else wordsReason = `write failed - ${json?.wordsErrors?.file ?? "child reported no words document"}`;
764
+ takeWordsErrors(json);
765
+ }
766
+ let ocr = handleOcr(tier, type, o, json);
767
+ if (forced) {
768
+ const r = await s.runTier("ocr-pages", { path: pdfPath, pages: o.pages, ...(o.words && type === "pdf" ? { words: true } : {}), stem: b.stem, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs, stagingDir: b.ocrStagingDir, dpi: o.imageDpi, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion }, b, signal, o.primaryTimeoutMs, backend);
769
+ if (signal?.aborted || (!r.ok && "reason" in r && r.reason === "aborted")) throw new Error("aborted");
770
+ if (r.ok && r.json.status === "unavailable") throw new Error(`OCR unavailable: ${r.json.reason} (install Tesseract; see doc/doc-to-md.md)`);
771
+ const outcome = r.ok && r.json.status === "ran" ? outcomeFromChild(r.json)
772
+ : recoverOcrPages(b.ocrStagingDir, o.pages!, !r.ok ? ("userError" in r ? r.userError : `${r.reason}${detailSuffix(r)}`) : "malformed child output");
773
+ const { sidecars, wordSidecars } = publishSidecars(b);
774
+ ocr = { ...emptyOcr(o.ocrLanguage), ...json.ocr, status: "ran", reason: null, mode: "all", pages: outcome.written, noText: outcome.noText, ocrFailed: outcome.ocrFailed, ocrErrors: outcome.ocrErrors, budgetStopped: outcome.budgetStopped, killed: outcome.killed, notAttempted: outcome.notAttempted, childError: outcome.childError, sidecars: Object.fromEntries(sidecars), wordSidecars: Object.fromEntries(wordSidecars) };
775
+ if (o.words && r.ok) takeWordsErrors(r.json);
776
+ }
707
777
  if (!isExcel) notes = [...notes, ...(json.notes ?? [])];
708
778
  if (b.renamedFrom) notes.splice(notes[0]?.startsWith("preview truncated:") ? 1 : 0, 0, `renamed to ${b.stem} (${b.renameReason})`);
709
779
  const pageImagesReason = !o.pageImages || b.pageManifest.size || tier === "primary" || tier === "fallback" ? null
@@ -719,7 +789,7 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
719
789
  const markdown = (head.length ? `${head.join("\n")}\n\n` : "") + body;
720
790
  commitBundle(b, markdown);
721
791
  const outline = scanOutline(markdown, o.outlineMaxEntries);
722
- const details: DocToMdDetails = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr: handleOcr(tier, type, o, json) };
792
+ const details: DocToMdDetails = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr, pageStats, pageStatsPath: pageStats ? b.pageStatsPath : null, ocrDir: b.ocrManifest.size ? b.ocrDir : null, wordsPath, wordsReason, wordsErrors };
723
793
  return { output: formatHandle(details), details };
724
794
  } catch (e) { abortBundle(b); throw e; }
725
795
  finally { office?.cleanup(); }
@@ -1,4 +1,4 @@
1
- import type { InputType } from "./doc-to-md-options.ts";
1
+ import type { InputType, OcrMode } from "./doc-to-md-options.ts";
2
2
 
3
3
  export type Tier = "primary" | "fallback" | "unpdf" | "excel" | "docx" | "html" | "image" | "email";
4
4
  export type Engine = "pymupdf4llm" | "pymupdf-text" | "unpdf" | "openpyxl" | "xlrd" | "mammoth" | "python-docx" | "markdownify" | "turndown" | "copy" | "extract-msg" | "email";
@@ -18,6 +18,8 @@ export interface SheetInfo {
18
18
  csv: string | null;
19
19
  }
20
20
 
21
+ export type PageStat = { page: number; chars: number; images: number; imageCoverage: number } | { page: number; error: string };
22
+
21
23
  export interface OcrInfo {
22
24
  status: "off" | "unavailable" | "skipped" | "ran";
23
25
  lang: string;
@@ -28,6 +30,13 @@ export interface OcrInfo {
28
30
  budgetStopped: number[];
29
31
  reason: string | null;
30
32
  tesseract: boolean | null;
33
+ mode: OcrMode;
34
+ sidecars: Record<number, string>;
35
+ wordSidecars: Record<number, string>;
36
+ ocrErrors: Record<number, string>;
37
+ killed: number | null;
38
+ notAttempted: number[];
39
+ childError: string | null;
31
40
  }
32
41
 
33
42
  export interface HandleData {
@@ -35,6 +44,8 @@ export interface HandleData {
35
44
  pageCount: number | null; pages: number[] | null; explicitBreaks: number | null; imageCount: number; pageImageCount: number; pageImagesReason: string | null; bytes: number; lines: number;
36
45
  degraded: string | null; fallbackReason: string | null; failedPages: number[]; emptyPages: number[];
37
46
  notes: string[]; outline: OutlineEntry[]; outlineTotal: number; ocr: OcrInfo | null;
47
+ pageStats: PageStat[] | null; pageStatsPath: string | null; ocrDir: string | null;
48
+ wordsPath: string | null; wordsReason: string | null; wordsErrors: Record<number, string>;
38
49
  }
39
50
 
40
51
  export interface InfoData {
@@ -72,7 +83,23 @@ export function compactRanges(nums: number[], maxEntries = 20): string {
72
83
  const INSTALL_HINT = "install Tesseract (see doc/doc-to-md.md), then rerun with ocr=true";
73
84
  const BARE_REASONS = ["fallback tier", "no Python backend"];
74
85
 
86
+ const pageList = (pages: number[]) => `page${pages.length === 1 ? "" : "s"} ${pages.join(", ")}`;
87
+
88
+ function forcedOcrLine(ocr: OcrInfo): string {
89
+ const clauses: string[] = [];
90
+ if (ocr.pages.length) clauses.push(`sidecars for ${pageList(ocr.pages)}`);
91
+ if (ocr.noText.length) clauses.push(`no text on ${pageList(ocr.noText)}`);
92
+ for (const page of ocr.ocrFailed) clauses.push(`failed on page ${page} (${ocr.ocrErrors[page] ?? "unknown error"})`);
93
+ if (ocr.budgetStopped.length) clauses.push(`budget-stopped ${pageList(ocr.budgetStopped)}`);
94
+ if (ocr.killed !== null) clauses.push(`page ${ocr.killed} killed the OCR child (timeout or crash - likely a compression bomb)`);
95
+ if (ocr.childError !== null) clauses.push(`OCR child failed before processing pages: ${ocr.childError}`);
96
+ if (ocr.notAttempted.length) clauses.push(`${pageList(ocr.notAttempted)} not attempted`);
97
+ const rerun = [...ocr.budgetStopped, ...ocr.notAttempted].sort((a, b) => a - b);
98
+ return `OCR: forced (${ocr.lang}) - ${clauses.join("; ")}${rerun.length ? ` - re-run with --pages ${rerun.join(",")}` : ""}`;
99
+ }
100
+
75
101
  export function ocrLine(ocr: OcrInfo, type: InputType): string {
102
+ if (ocr.mode === "all") return forcedOcrLine(ocr);
76
103
  const image = type === "image";
77
104
  const rerun = ocr.tesseract ? "rerun with ocr=true" : INSTALL_HINT;
78
105
  switch (ocr.status) {
@@ -143,6 +170,12 @@ export function formatHandle(h: HandleData): string {
143
170
  if (h.sheetsDir) lines.push(`Sheets-Dir: ${h.sheetsDir}`);
144
171
  if (h.pagesDir && h.pageImageCount > 0) lines.push(`Pages-Dir: ${h.pagesDir} (${h.pageImageCount} pages)`);
145
172
  else if (h.pageImagesReason) lines.push(`Pages-Dir: none - ${h.pageImagesReason}`);
173
+ if (h.pageStatsPath) lines.push(`Page-Stats: ${h.pageStatsPath}`);
174
+ if (h.wordsPath) {
175
+ const bad = Object.keys(h.wordsErrors).map(Number).sort((a, b) => a - b);
176
+ lines.push(`Words: ${h.wordsPath}${bad.length ? ` (extraction failed for pages ${compactRanges(bad)}: ${h.wordsErrors[bad[0]]})` : ""}`);
177
+ } else if (h.wordsReason) lines.push(`Words: ${h.wordsReason}`);
178
+ if (h.ocrDir) lines.push(`OCR-Dir: ${h.ocrDir}`);
146
179
  lines.push(`Type: ${h.type} Engine: ${h.engine} Tier: ${h.tier}`);
147
180
  lines.push(`Page-Count: ${pageCountLabel(h)} Pages: ${h.pages ? compactRanges(h.pages) : "all"} Images: ${h.imageCount} Size: ${formatSize(h.bytes)} / ${h.lines} lines`);
148
181
  if (h.degraded) lines.push(`Degraded: ${h.degraded}`);
@@ -6,6 +6,7 @@ import { extname } from "node:path";
6
6
 
7
7
  export type InputType = "pdf" | "docx" | "doc" | "pptx" | "xlsx" | "xlsm" | "xls" | "html" | "image" | "email";
8
8
  export type ImageFormat = "png" | "jpg";
9
+ export type OcrMode = "textless" | "all";
9
10
 
10
11
  export interface Tunables {
11
12
  primaryTimeoutMs: number;
@@ -29,6 +30,8 @@ export interface DocToMdOptions extends Tunables {
29
30
  outputDir: string | null;
30
31
  overwrite: boolean;
31
32
  pageImages: boolean;
33
+ words: boolean;
34
+ ocrMode: OcrMode;
32
35
  }
33
36
 
34
37
  /** What adapters pass in: intents as raw strings/booleans, tunables optional. */
@@ -39,6 +42,8 @@ export interface PerCallInput extends Partial<Tunables> {
39
42
  outputDir?: string | null;
40
43
  overwrite?: boolean;
41
44
  pageImages?: boolean;
45
+ words?: boolean;
46
+ ocrMode?: OcrMode;
42
47
  }
43
48
 
44
49
  export type DescriptorType = "string" | "int" | "bool" | "pages" | "enum" | "version" | "lang";
@@ -68,6 +73,8 @@ export const DOC_TO_MD_OPTIONS: readonly OptionDescriptor[] = [
68
73
  { key: "outputDir", flag: "--output-dir", type: "string", default: null, settable: false, help: "Bundle root for <stem>.md + images/; default a per-call temp dir. <stem> = basename without extension with [^A-Za-z0-9._-]+ -> _ (empty -> document); a second call on the same stem writes <stem>-2.md" },
69
74
  { key: "overwrite", flag: "--overwrite", type: "bool", default: false, settable: false, help: "Replace an existing completed <stem>.md bundle" },
70
75
  { key: "pageImages", flag: "--page-images", type: "bool", default: false, settable: false, help: "Also render every selected page to pages/<stem>-pNNN.<imageFormat> at imageDpi (PDF, PPTX, .doc, DOCX via LibreOffice); off by default" },
76
+ { key: "words", flag: "--words", type: "bool", default: false, settable: false, help: "Write word positions: <stem>.words.json beside the Markdown lists every text-layer word of each selected page with its bbox (PDF points, top-left origin, display orientation; image inputs in source pixels) and the words inline OCR recognized, tagged source \"text\" or \"ocr\"; under --ocr-mode all the OCR words go to ocr/<stem>-pNNN.words.json beside each sidecar. Never triggers OCR. PDF and image inputs only." },
77
+ { key: "ocrMode", flag: "--ocr-mode", type: "enum", default: "textless", settable: false, enumValues: ["textless", "all"], help: "OCR policy: textless (default) OCRs only pages with an empty text layer, inline; all OCRs every selected page and writes the recognized text to ocr/<stem>-pNNN.md sidecars, leaving the Markdown untouched. all requires --ocr and an explicit --pages selection (PDF, PPTX, DOC)." },
71
78
  { key: "primaryTimeoutMs", flag: "--primary-timeout", type: "int", default: 60000, settable: true, env: "PI_DOC_TO_MD_CONVERT_TIMEOUT_MS", help: "pymupdf4llm tier and DOCX child (docx mode); also the unpdf tier" },
72
79
  { key: "fallbackTimeoutMs", flag: "--fallback-timeout", type: "int", default: 30000, settable: true, help: "PyMuPDF get_text tier (including DOCX LibreOffice fallback); also PDF and DOCX info and Excel rendered views" },
73
80
  { key: "sofficeTimeoutMs", flag: "--soffice-timeout", type: "int", default: 120000, settable: true, env: "PI_DOC_TO_MD_SOFFICE_TIMEOUT_MS", help: "DOCX/PPTX -> PDF via LibreOffice; also Excel rendered views" },
@@ -177,12 +184,47 @@ export function resolveOptions(perCall: PerCallInput, settings: Partial<Tunables
177
184
  out[d.key] = value;
178
185
  }
179
186
  const o = out as unknown as DocToMdOptions;
180
- if (o.info && (o.pages !== null || o.outputDir !== null || o.overwrite || o.pageImages)) {
181
- throw new UsageError("--info cannot be combined with --pages, --output-dir, --overwrite or --page-images");
187
+ if (o.info && (o.pages !== null || o.outputDir !== null || o.overwrite || o.pageImages || o.words || o.ocrMode === "all")) {
188
+ throw new UsageError("--info cannot be combined with --pages, --output-dir, --overwrite, --page-images, --words or --ocr-mode all");
182
189
  }
183
190
  return o;
184
191
  }
185
192
 
193
+ /** Two-pass OCR recipe, rendered into --help and the generated skill; `<cmd>` is the command prefix. */
194
+ export const USAGE_PATTERNS = [
195
+ "Two-pass OCR (PDF, PPTX, DOC):",
196
+ " 1. <cmd> report.pdf --output-dir out --json",
197
+ ' -> "pageStatsPath" points at out/report.pages.json; pages with few',
198
+ ' chars and high imageCoverage are scans. "savedTo" is the Markdown.',
199
+ " 2. <cmd> report.pdf --output-dir out --ocr --ocr-mode all --pages 2,7 --json",
200
+ ' -> "ocr"."sidecars" maps 2 and 7 to out/ocr/report-2-p002.md and',
201
+ " ...-p007.md (a second run in the same dir gets stem report-2);",
202
+ " the Markdown of this run holds pages 2 and 7 only and equals what",
203
+ " --pages 2,7 without --ocr would produce. Read sidecars and Markdown",
204
+ " by the returned paths, never by guessing names.",
205
+ " 3. --ocr-mode all refuses to run without --ocr and an explicit --pages.",
206
+ ' The "ocr" object (OCR: line) names failed, budget-stopped, killed and',
207
+ " not-attempted pages and the exact --pages to re-run.",
208
+ " Details: doc/doc-to-md.md (bundle contract, failure buckets).",
209
+ ].join("\n");
210
+
211
+ export const usagePatterns = (cmd: string): string => USAGE_PATTERNS.replaceAll("<cmd>", cmd);
212
+
213
+ /** Every artifact a bundle can contain; shared by --help and the generated skill. */
214
+ export const BUNDLE_LAYOUT: readonly { artifact: string; trigger: string; content: string; namedBy: string }[] = [
215
+ { artifact: "<stem>.md", trigger: "always", content: "the Markdown", namedBy: "Saved-To: / savedTo" },
216
+ { artifact: "images/", trigger: "embedded or extracted figures", content: "image files linked from the Markdown", namedBy: "Images-Dir: / imagesDir" },
217
+ { artifact: "pages/<stem>-pNNN.<fmt>", trigger: "--page-images", content: "page renders at --image-dpi", namedBy: "Pages-Dir: / pagesDir" },
218
+ { artifact: "sheets/", trigger: "Excel input", content: "one CSV per non-empty worksheet", namedBy: "Sheets-Dir: / sheetsDir" },
219
+ { artifact: "attachments/", trigger: "email input", content: "saved attachments", namedBy: "Markdown attachment list" },
220
+ { artifact: "<stem>.pages.json", trigger: "Python PDF tiers (PDF, PPTX, DOC, DOCX via LibreOffice; not unpdf)", content: "per-page chars, image count, image coverage", namedBy: "Page-Stats: / pageStatsPath" },
221
+ { artifact: "<stem>.words.json", trigger: "--words", content: "per-page word boxes, source text/ocr", namedBy: "Words: / wordsPath" },
222
+ { artifact: "ocr/<stem>-pNNN.md", trigger: "--ocr --ocr-mode all", content: "recognized text of a forced page", namedBy: "OCR-Dir: / ocr.sidecars" },
223
+ { artifact: "ocr/<stem>-pNNN.words.json", trigger: "--ocr --ocr-mode all --words", content: "word boxes of that OCR", namedBy: "ocr.wordSidecars" },
224
+ ];
225
+
226
+ const bundleLayoutText = (): string => BUNDLE_LAYOUT.map((r) => ` ${r.artifact.padEnd(28)} ${r.trigger}; ${r.content}; named by ${r.namedBy}`).join("\n");
227
+
186
228
  export function renderHelp(): string {
187
229
  const row = (d: OptionDescriptor) => ` ${(d.flag ?? "<path>").padEnd(26)} ${d.help}${d.default !== null && d.key !== "info" && d.key !== "overwrite" ? ` (default ${d.default})` : ""}`;
188
230
  return [
@@ -190,7 +232,8 @@ export function renderHelp(): string {
190
232
  "", "Per-call:", ...DOC_TO_MD_OPTIONS.filter((d) => !d.settable).map(row),
191
233
  "", "Tunables (also settable under quiver.docToMd in settings.json; per-call > settings > default):",
192
234
  ...DOC_TO_MD_OPTIONS.filter((d) => d.settable).map(row),
193
- "", "Result: a handle (Saved-To, Images-Dir, Page-Count, Outline ...). Read the Saved-To file for the Markdown.",
235
+ "", "Result: a handle (Saved-To, Images-Dir, Page-Stats, Page-Count, Outline ...). Read the Saved-To file for the Markdown.",
194
236
  "Exit codes: 0 success, 1 runtime error, 2 usage error.",
237
+ "", "Bundle layout:", bundleLayoutText(), "", usagePatterns("pi-quiver doc-to-md"),
195
238
  ].join("\n");
196
239
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-quiver",
3
- "version": "6.8.0",
3
+ "version": "6.10.0",
4
4
  "description": "Personal pack of Pi coding-agent extensions: context-safe fetch, doc_to_md PDF/DOCX/PPTX-to-Markdown conversion, session naming, a themed ASCII startup header, Opus 4.8 fast mode, and a provider-stall watchdog.",
5
5
  "author": "Jacek Juraszek",
6
6
  "license": "MIT",
@@ -205,11 +205,22 @@ def mode_image(o):
205
205
  info = new_ocr(lang)
206
206
  status = ocr_status(bool(o.get("ocr")), lang)
207
207
  apply_status(info, status)
208
- if status["status"] == "ready":
208
+ words_on = bool(o.get("words"))
209
+ word_pages, words_errors = [], {}
210
+ w_px = h_px = None
211
+ if words_on or status["status"] == "ready":
209
212
  try:
210
213
  pix = pymupdf.Pixmap(path)
211
214
  w_px, h_px = pix.width, pix.height
212
215
  del pix
216
+ except Exception as exc: # noqa: BLE001
217
+ if words_on:
218
+ words_errors["1"] = words_error(exc)
219
+ if status["status"] == "ready":
220
+ info["ocrFailed"].append(1)
221
+ status = {**status, "status": "failed"}
222
+ if status["status"] == "ready":
223
+ try:
213
224
  if min(w_px, h_px) < MIN_OCR_SIDE_PX:
214
225
  info["status"], info["reason"] = "skipped", "image too small"
215
226
  else:
@@ -218,6 +229,19 @@ def mode_image(o):
218
229
  text = pymupdf4llm.to_markdown(pdf, pages=[0], write_images=False, use_ocr=True, force_ocr=True,
219
230
  ocr_language=lang, ocr_dpi=image_ocr_dpi(w_px, r.width, r.height),
220
231
  page_separators=False).strip()
232
+ if words_on:
233
+ try:
234
+ pg = pdf[0]
235
+ entry = words_page(pg, 1, 0, page_words(pg, True))
236
+ sx, sy = w_px / r.width, h_px / r.height
237
+ for word in entry["words"]:
238
+ b = word["bbox"]
239
+ word["bbox"] = [round(b[0] * sx, 1), round(b[1] * sy, 1),
240
+ round(b[2] * sx, 1), round(b[3] * sy, 1)]
241
+ entry["width"], entry["height"] = w_px, h_px
242
+ word_pages.append(entry)
243
+ except Exception as exc: # noqa: BLE001
244
+ words_errors["1"] = words_error(exc)
221
245
  if text:
222
246
  info["pages"].append(1)
223
247
  md += "\n\n" + ocr_block(f"p1/{name}", text)
@@ -225,7 +249,17 @@ def mode_image(o):
225
249
  info["noText"].append(1)
226
250
  except Exception: # noqa: BLE001 - OCR never fails the conversion
227
251
  info["ocrFailed"].append(1)
228
- return {"markdown": md + "\n", "pageCount": 1, "emptyPages": [], "failedPages": [], "notes": [], "ocr": info}
252
+ if words_on and not o.get("ocr") and w_px is not None and "1" not in words_errors:
253
+ try:
254
+ if os.environ.get("DOC_TO_MD_WORDS_FAIL") == "1": # tests only
255
+ raise RuntimeError("words injected failure")
256
+ word_pages.append({"page": 1, "width": w_px, "height": h_px, "rotation": 0, "words": []})
257
+ except Exception as exc: # noqa: BLE001
258
+ words_errors["1"] = words_error(exc)
259
+ result = {"markdown": md + "\n", "pageCount": 1, "emptyPages": [], "failedPages": [], "notes": [], "ocr": info}
260
+ if words_on:
261
+ result.update(write_words(o["stagingDir"], "px", word_pages, words_errors))
262
+ return result
229
263
 
230
264
 
231
265
  def mode_info(o):
@@ -253,6 +287,71 @@ def primary_page_markdown(doc, n, d, o, kw, write_images):
253
287
  return rewrite_image_destinations(md, sources)
254
288
 
255
289
 
290
+ def page_stats(doc, n):
291
+ try:
292
+ page = doc[n - 1]
293
+ chars = len(page.get_text("text").strip())
294
+ infos = page.get_image_info()
295
+ area = page.rect.width * page.rect.height
296
+ covered = sum(max(0.0, (b[2] - b[0]) * (b[3] - b[1])) for b in (i["bbox"] for i in infos))
297
+ coverage = round(min(1.0, covered / area), 2) if area > 0 else 0.0
298
+ return {"page": n, "chars": chars, "images": len(infos), "imageCoverage": coverage}
299
+ except Exception as exc: # noqa: BLE001 - stats never cost a page its Markdown
300
+ return {"page": n, "error": f"{type(exc).__name__}: {exc}"[:300]}
301
+
302
+
303
+ GLYPHLESS_FONT = "GlyphLessFont"
304
+
305
+
306
+ def word_entries(page, words, ocr_lines):
307
+ import pymupdf
308
+ matrix = page.rotation_matrix
309
+ out = []
310
+ for x0, y0, x1, y1, text, bno, lno, _ in words:
311
+ r = pymupdf.Rect(x0, y0, x1, y1) * matrix
312
+ out.append({"text": text, "bbox": [round(r.x0, 1), round(r.y0, 1), round(r.x1, 1), round(r.y1, 1)],
313
+ "source": "ocr" if ocr_lines is None or (bno, lno) in ocr_lines else "text"})
314
+ return out
315
+
316
+
317
+ def page_words(page, inline_ocr_ran):
318
+ if not inline_ocr_ran:
319
+ return word_entries(page, page.get_text("words"), set())
320
+ # Shared textpage keeps block/line indexes aligned despite differing default image flags.
321
+ tp = page.get_textpage()
322
+ ocr_lines = {(bi, li) for bi, b in enumerate(page.get_text("dict", textpage=tp)["blocks"])
323
+ for li, line in enumerate(b.get("lines", []))
324
+ if line["spans"] and all(s["font"] == GLYPHLESS_FONT for s in line["spans"])}
325
+ return word_entries(page, page.get_text("words", textpage=tp), ocr_lines)
326
+
327
+
328
+ def ocr_words(page, tp):
329
+ return word_entries(page, page.get_text("words", textpage=tp), None)
330
+
331
+
332
+ def words_page(page, n, rotation, words):
333
+ if os.environ.get("DOC_TO_MD_WORDS_FAIL") == str(n): # tests only
334
+ raise RuntimeError("words injected failure")
335
+ return {"page": n, "width": round(page.rect.width, 1), "height": round(page.rect.height, 1),
336
+ "rotation": rotation, "words": words}
337
+
338
+
339
+ def words_error(exc):
340
+ return f"{type(exc).__name__}: {exc}"[:300]
341
+
342
+
343
+ def write_words(staging, unit, pages, errors):
344
+ try:
345
+ os.makedirs(staging, exist_ok=True)
346
+ with open(os.path.join(staging, "words.json"), "w", encoding="utf-8") as fh:
347
+ json.dump({"unit": unit, "pages": pages}, fh, indent=2)
348
+ fh.write("\n")
349
+ return {"words": True, "wordsErrors": errors}
350
+ except Exception as exc: # noqa: BLE001 - geometry never fails the conversion
351
+ errors["file"] = words_error(exc)
352
+ return {"words": False, "wordsErrors": errors}
353
+
354
+
256
355
  def mode_pdf_primary(o):
257
356
  import time
258
357
  start = time.monotonic()
@@ -260,14 +359,18 @@ def mode_pdf_primary(o):
260
359
  pages = check_pages(o.get("pages"), doc.page_count)
261
360
  staging, out, empty, failed, notes = o["stagingDir"], [], [], [], []
262
361
  page_images = []
362
+ stats = []
263
363
  lang, budget = o.get("ocrLanguage", "eng"), o.get("ocrBudgetMs", 60000)
264
364
  info = new_ocr(lang)
265
365
  status = ocr_status(True, lang) if o.get("ocr") else None
266
366
  ocr_ms, plain_ms = [], []
367
+ word_pages, words_errors = [], {}
267
368
  for i, n in enumerate(pages):
369
+ stats.append(page_stats(doc, n))
268
370
  d = page_dir(staging, n)
269
371
  try:
270
372
  page = doc[n - 1]
373
+ rotation = page.rotation
271
374
  textless = not page.get_text("text").strip()
272
375
  if textless:
273
376
  info["textless"].append(n)
@@ -311,6 +414,11 @@ def mode_pdf_primary(o):
311
414
  if page_pic:
312
415
  md = "\n\n".join(x for x in [md.rstrip(), f"![page {n}](pages/{page_pic})"] if x)
313
416
  out.append(md.rstrip())
417
+ if o.get("words") and not (textless and o.get("ocr") and not kw["use_ocr"]):
418
+ try:
419
+ word_pages.append(words_page(page, n, rotation, page_words(page, kw["use_ocr"])))
420
+ except Exception as exc: # noqa: BLE001
421
+ words_errors[str(n)] = words_error(exc)
314
422
  except Exception as exc: # noqa: BLE001
315
423
  shutil.rmtree(d, ignore_errors=True)
316
424
  failed.append({"page": n, "error": f"{type(exc).__name__}: {exc}"[:300]})
@@ -324,8 +432,11 @@ def mode_pdf_primary(o):
324
432
  missing = len(pages) - len(page_images) if o.get("pageImages") else 0
325
433
  if missing:
326
434
  notes.append(f"Page images: {missing} of {len(pages)} unavailable")
327
- return {"markdown": "\n\n".join(out) + "\n", "pages": pages, "pageCount": doc.page_count,
328
- "emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": info, "pageImages": page_images}
435
+ result = {"markdown": "\n\n".join(out) + "\n", "pages": pages, "pageCount": doc.page_count,
436
+ "emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": info, "pageImages": page_images, "pageStats": stats}
437
+ if o.get("words"):
438
+ result.update(write_words(staging, "pt", word_pages, words_errors))
439
+ return result
329
440
 
330
441
 
331
442
  def mode_pdf_fallback(o):
@@ -335,15 +446,20 @@ def mode_pdf_fallback(o):
335
446
  keep = {int(k): v for k, v in (o.get("keepPages") or {}).items()}
336
447
  staging, out, empty, failed = o["stagingDir"], [], [], []
337
448
  page_images = []
449
+ stats = []
338
450
  lang = o.get("ocrLanguage", "eng")
339
451
  ocr_info = new_ocr(lang)
452
+ word_pages, words_errors = [], {}
340
453
  status = {"status": "unavailable", "reason": "fallback tier", "tesseract": None} if o.get("ocr") else None
341
454
  for n in pages:
455
+ stats.append(page_stats(doc, n))
342
456
  links = [f"![](images/{f})" for f in keep.get(n, [])]
343
457
  text = ""
344
458
  page_pic = None
459
+ page_ok = False
345
460
  try:
346
461
  page = doc[n - 1]
462
+ rotation = page.rotation
347
463
  text = page.get_text("text").strip()
348
464
  if not text:
349
465
  ocr_info["textless"].append(n)
@@ -371,11 +487,17 @@ def mode_pdf_fallback(o):
371
487
  links.append(f"![](p{n}/{name})")
372
488
  mark_done(d)
373
489
  page_pic = render_page_image(page, n, o, page_images) if o.get("pageImages") else None
490
+ page_ok = True
374
491
  except Exception as exc: # noqa: BLE001
375
492
  shutil.rmtree(os.path.join(staging, f"p{n}"), ignore_errors=True)
376
493
  text = ""
377
494
  links = [f"![](images/{f})" for f in keep.get(n, [])]
378
495
  failed.append({"page": n, "error": f"{type(exc).__name__}: {exc}"[:300]})
496
+ if o.get("words") and page_ok and not (not text and o.get("ocr")):
497
+ try:
498
+ word_pages.append(words_page(page, n, rotation, page_words(page, False)))
499
+ except Exception as exc: # noqa: BLE001
500
+ words_errors[str(n)] = words_error(exc)
379
501
  if not text:
380
502
  empty.append(n)
381
503
  out.append("\n\n".join(x for x in [text, "\n".join(links), f"![page {n}](pages/{page_pic})" if page_pic else ""] if x))
@@ -388,8 +510,88 @@ def mode_pdf_fallback(o):
388
510
  missing = len(pages) - len(page_images) if o.get("pageImages") else 0
389
511
  if missing:
390
512
  notes.append(f"Page images: {missing} of {len(pages)} unavailable")
391
- return {"markdown": "\n\n".join(out) + "\n", "pages": pages, "pageCount": doc.page_count,
392
- "emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": ocr_info, "pageImages": page_images}
513
+ result = {"markdown": "\n\n".join(out) + "\n", "pages": pages, "pageCount": doc.page_count,
514
+ "emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": ocr_info, "pageImages": page_images, "pageStats": stats}
515
+ if o.get("words"):
516
+ result.update(write_words(staging, "pt", word_pages, words_errors))
517
+ return result
518
+
519
+
520
+ def mode_ocr_pages(o):
521
+ import time
522
+ start = time.monotonic()
523
+ lang, budget = o.get("ocrLanguage", "eng"), o.get("ocrBudgetMs", 60000)
524
+ status = ocr_status(True, lang)
525
+ if status["status"] != "ready":
526
+ return {"status": "unavailable", "reason": status["reason"]}
527
+ doc = open_pdf(o["path"])
528
+ pages = check_pages(o.get("pages"), doc.page_count)
529
+ staging, stem, dpi = o["stagingDir"], o["stem"], o.get("dpi", 150)
530
+ os.makedirs(staging, exist_ok=True)
531
+ active = os.path.join(staging, "active")
532
+ out = {"status": "ran", "written": [], "noText": [], "ocrFailed": [], "ocrErrors": {}, "budgetStopped": []}
533
+ if o.get("words"):
534
+ out["wordsErrors"] = {}
535
+ ocr_ms = []
536
+ for i, n in enumerate(pages):
537
+ with open(active, "w") as fh:
538
+ fh.write(str(n))
539
+ if os.environ.get("DOC_TO_MD_OCR_STALL_PAGE") == str(n): # tests only: simulate a wedged page
540
+ time.sleep(3600)
541
+ elapsed = (time.monotonic() - start) * 1000
542
+ est = max(ocr_ms) if ocr_ms else OCR_EST_INITIAL_MS
543
+ if not ocr_admit(elapsed, est, 0, 0, budget):
544
+ out["budgetStopped"].extend(pages[i:])
545
+ os.remove(active)
546
+ break
547
+ tag = f"p{n:03d}"
548
+ d = os.path.join(staging, tag)
549
+ os.makedirs(d, exist_ok=True)
550
+ sidecar = os.path.join(d, f"{stem}-{tag}.md")
551
+ t0 = time.monotonic()
552
+ try:
553
+ page = doc[n - 1]
554
+ eff = clamped_dpi(page.rect.width, page.rect.height, dpi)
555
+ if eff is None:
556
+ raise RuntimeError(f"page cannot be rendered at a usable DPI ({page.rect.width:.0f} x {page.rect.height:.0f} pt)")
557
+ tp = page.get_textpage_ocr(full=True, language=lang, dpi=eff)
558
+ text = page.get_text("text", textpage=tp).strip()
559
+ header = f"<!-- OCR of page {n} (tesseract {lang}); recognized text, not the text layer -->"
560
+ with open(sidecar, "w", encoding="utf-8") as fh:
561
+ fh.write(header + "\n\n" + (text + "\n\n" if text else "") + SEP.format(n=n).strip("\n") + "\n")
562
+ if o.get("words"):
563
+ wpath = os.path.join(d, f"{stem}-{tag}.words.json")
564
+ try:
565
+ entry = words_page(page, n, page.rotation, ocr_words(page, tp))
566
+ entry["unit"] = "pt"
567
+ with open(wpath, "w", encoding="utf-8") as fh:
568
+ json.dump(entry, fh, indent=2)
569
+ fh.write("\n")
570
+ except Exception as exc: # noqa: BLE001 - geometry never changes the OCR outcome
571
+ try:
572
+ os.remove(wpath)
573
+ except OSError:
574
+ pass
575
+ out["wordsErrors"][str(n)] = words_error(exc)
576
+ mark_done(d)
577
+ (out["written"] if text else out["noText"]).append(n)
578
+ except Exception as exc: # noqa: BLE001 - one page never stops the pass
579
+ msg = f"{type(exc).__name__}: {exc}"
580
+ with open(os.path.join(d, ".failed"), "w", encoding="utf-8") as fh:
581
+ fh.write(msg)
582
+ try:
583
+ os.remove(sidecar)
584
+ except OSError:
585
+ pass
586
+ try:
587
+ os.remove(os.path.join(d, f"{stem}-{tag}.words.json"))
588
+ except OSError:
589
+ pass
590
+ out["ocrFailed"].append(n)
591
+ out["ocrErrors"][str(n)] = msg
592
+ ocr_ms.append((time.monotonic() - t0) * 1000)
593
+ os.remove(active)
594
+ return out
393
595
 
394
596
 
395
597
  def esc(v):
@@ -1249,12 +1451,12 @@ def mode_render_pages(o):
1249
1451
  return {"ok": True, "rendered": rendered, "failed": failed}
1250
1452
 
1251
1453
 
1252
- MODES = {"info": None, "pdf-primary": mode_pdf_primary, "pdf-fallback": mode_pdf_fallback, "xlsx": mode_xlsx, "render-pages": mode_render_pages, "docx": mode_docx, "html": mode_html, "image": mode_image, "email": mode_email}
1454
+ MODES = {"info": None, "pdf-primary": mode_pdf_primary, "pdf-fallback": mode_pdf_fallback, "xlsx": mode_xlsx, "render-pages": mode_render_pages, "docx": mode_docx, "html": mode_html, "image": mode_image, "email": mode_email, "ocr-pages": mode_ocr_pages}
1253
1455
 
1254
1456
 
1255
1457
  def main():
1256
1458
  if len(sys.argv) != 2 or sys.argv[1] not in MODES:
1257
- print("usage: doc_to_md.py <info|pdf-primary|pdf-fallback|xlsx|render-pages|docx|html|image|email> (options JSON on stdin)", file=sys.stderr)
1459
+ print("usage: doc_to_md.py <info|pdf-primary|pdf-fallback|xlsx|render-pages|docx|html|image|email|ocr-pages> (options JSON on stdin)", file=sys.stderr)
1258
1460
  return 1
1259
1461
  mode = sys.argv[1]
1260
1462
  o = json.loads(sys.stdin.read() or "{}")