pi-quiver 6.8.0 → 6.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -8,6 +8,12 @@ Published to npm as `pi-quiver` (`pi install npm:pi-quiver`). Pushing a
8
8
  via OIDC trusted publishing. The release helper at
9
9
  `.agents/skills/release/scripts/release.sh` cuts the tag; CI publishes.
10
10
 
11
+ ## v6.9.0 - 2026-10-01
12
+
13
+ - `doc_to_md`: every conversion that goes through the Python PDF tiers (PDF, PPTX, `.doc`, DOCX via LibreOffice) writes `<stem>.pages.json` (per-page `chars`, `images`, `imageCoverage`) and the handle prints `Page-Stats:`; CLI `--json` gains `pageStats`, `pageStatsPath`, `ocrDir`. The unpdf tier writes no stats (#26).
14
+ - `doc_to_md`: new per-call `ocrMode` (`--ocr-mode textless|all`). `all` forces OCR on an explicit `pages` selection in a second, kill-capped Python child and writes `ocr/<stem>-pNNN.md` sidecars; the Markdown stays byte-identical to the same selected-page call without OCR. Refused without `--ocr`, without explicit pages, or on non-PDF/PPTX/DOC inputs (exit 2); missing Tesseract data is a hard error. The `OCR:` line names written, no-text, failed, budget-stopped, killed and not-attempted pages, with the exact `--pages` to re-run for budget-stopped and not-attempted pages (#26).
15
+ - `doc_to_md`: `--help` and the generated `skills/doc-to-md/SKILL.md` end with the two-pass OCR usage block (#26).
16
+
11
17
  ## v6.8.0 - 2026-10-01
12
18
 
13
19
  - `session-name`: opt-in automatic naming starts in the background after three completed model/tool rounds. Shorter runs start a non-blocking attempt at run end; initial generation is best-effort with a 30-second local deadline and no automatic retry per session activation.
package/README.md CHANGED
@@ -65,7 +65,7 @@ A 300 KB changelog page never touches your context window - you get a preview an
65
65
  | Extension | Tool | What it does |
66
66
  | --- | --- | --- |
67
67
  | `extensions/fetch.ts` | `fetch` | Retrieve URLs over HTTP(S). HTML -> Markdown (Readability extraction, Turndown conversion). Binary saved untouched to a temp file. GitHub issue/PR/repo/actions-run/actions-job URLs auto-route through `gh` (falls back to HTTP); failed runs/jobs include failed-step logs (best-effort, summary-only otherwise). Same size gate as `fetch`. Behavior lives in `lib/fetch-core.ts`; also exposed as the `pi-quiver fetch` CLI (see [Claude Code support](#claude-code-support)). |
68
- | `extensions/doc_to_md.ts` | `doc_to_md` | Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, MSG/EML, HTML, or image file to a Markdown bundle on disk (`<stem>.md` + `images/`, spreadsheet `sheets/`, optional `pageImages` in `pages/`, and email `attachments/`) and return a handle (paths, page count, outline, diagnostics) - never inline Markdown. `info` inspects PDF/Office inputs; `pages` selects 1-based PDF/Office pages (DOCX: explicit-page-break segments); HTML, image, and email inputs reject both; every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Tiers: pymupdf4llm -> PyMuPDF text (degraded) -> unpdf worker (no Python only). Excel -> sheet inventory, full CSVs, bounded previews, and optional rendered views. HTML keeps local and `data:` images in the bundle and remote images as links; image inputs keep the original image. Scanned PDF pages keep a page picture on both Python tiers; opt-in OCR needs Tesseract language data. DOCX converts directly (mammoth, python-docx fallback); LibreOffice pagination marks the degraded route. The Outline lists `L<line>` and `p<page>` per heading. Settings under `quiver.docToMd`. |
68
+ | `extensions/doc_to_md.ts` | `doc_to_md` | Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, MSG/EML, HTML, or image file to a Markdown bundle on disk (`<stem>.md` + `images/`, spreadsheet `sheets/`, optional `pageImages` in `pages/`, and email `attachments/`) and return a handle (paths, page count, outline, diagnostics) - never inline Markdown. `info` inspects PDF/Office inputs; `pages` selects 1-based PDF/Office pages (DOCX: explicit-page-break segments); HTML, image, and email inputs reject both; every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Tiers: pymupdf4llm -> PyMuPDF text (degraded) -> unpdf worker (no Python only). Excel -> sheet inventory, full CSVs, bounded previews, and optional rendered views. HTML keeps local and `data:` images in the bundle and remote images as links; image inputs keep the original image. Scanned PDF pages keep a page picture on both Python tiers; opt-in OCR needs Tesseract language data. DOCX converts directly (mammoth, python-docx fallback); LibreOffice pagination marks the degraded route. The Outline lists `L<line>` and `p<page>` per heading. Settings under `quiver.docToMd`. Every conversion that goes through the Python PDF tiers (PDF, PPTX, `.doc`, DOCX via LibreOffice) writes `<stem>.pages.json` (per-page chars, image count, image coverage; handle line `Page-Stats:`); unpdf writes no stats. `ocrMode: "all"` with `ocr: true` and an explicit `pages` forces OCR on those pages into `ocr/<stem>-pNNN.md` sidecars, leaving the Markdown byte-identical to the same selected-page call without OCR. |
69
69
  | `extensions/session-name.ts` | `/session-name` | Manual + opt-in automatic session naming, naming rules and deny list, long-session revisits, and Ghostty/Herdr tab rename. OFF by default. |
70
70
  | `extensions/sword-header.ts` | `/builtin-header` | Themed ASCII startup header replacing pi's default logo. OFF by default. |
71
71
  | `extensions/fast-mode.ts` | `/fast` | Inject Anthropic fast-mode (`speed: "fast"` + `anthropic-beta: fast-mode-2026-02-01`) into every Claude Opus 4.8 / Opus 5 request, any thinking level. `--fast` flag + `/fast [on\|off\|status]`. OFF by default. |
@@ -310,7 +310,9 @@ Each setting can also be overridden per-process via `PI_QUIVER_SLACK_ENABLED`, `
310
310
  | `ocr` | `false` | Opt-in OCR for scanned pages and images when Tesseract language data is installed. |
311
311
  | `ocrLanguage` | `eng` | Plain `+`-joined Tesseract language codes for OCR. |
312
312
 
313
- A bundle is `<outputDir>/<stem>.md` plus `<outputDir>/images/` and, for spreadsheets with data, `<outputDir>/sheets/`; without `outputDir`, the tool creates a per-call temp root. A conversion owns `<stem>.md.lock` until it atomically publishes the Markdown. An existing `<stem>.md` fails the call unless `overwrite` is set; `overwrite` replaces that Markdown and the files it owns (`images/<stem>-p<N>-<n>.*`, `images/<stem>-s<idx>[-<n>].*`, `sheets/<stem>-s<idx>-<slug>.csv`), nothing else. Temp bundles are caller-owned - the tool never deletes a bundle it produced.
313
+ `ocrMode` (`textless` | `all`) is a per-call parameter and `--ocr-mode` flag only; a `quiver.docToMd.ocrMode` key is reported as unknown and ignored, so persisted settings cannot enable forced OCR or bypass its explicit-pages guard.
314
+
315
+ A bundle is `<outputDir>/<stem>.md` plus `<outputDir>/<stem>.pages.json` on Python PDF tiers, optional OCR sidecars under `<outputDir>/ocr/`, and `<outputDir>/images/` and, for spreadsheets with data, `<outputDir>/sheets/`; without `outputDir`, the tool creates a per-call temp root. A conversion owns `<stem>.md.lock` until it atomically publishes the Markdown. An existing `<stem>.md` fails the call unless `overwrite` is set; `overwrite` replaces that Markdown and the files it owns (`images/<stem>-p<N>-<n>.*`, `images/<stem>-s<idx>[-<n>].*`, `sheets/<stem>-s<idx>-<slug>.csv`, `<stem>.pages.json`, `ocr/<stem>-pNNN.md`), nothing else. Temp bundles are caller-owned - the tool never deletes a bundle it produced.
314
316
 
315
317
  Excel needs a Python backend with openpyxl, xlrd and pillow - otherwise the call fails with `Remedy: install uv, or pip install openpyxl xlrd pillow`. The Markdown opens with a `## Sheets` table listing every sheet in workbook order (0-based `#`, `worksheet`/`chartsheet`, size, hidden, chart and image counts, rendered view, CSV link for non-empty worksheets), then one section per sheet: a `Data:` line linking the full-content CSV under `sheets/` for non-empty worksheets, chart metadata from the workbook model, embedded images, an optional rendered view, a preview of at most the first 100 rows x 50 columns, and - only when the preview is truncated - a `Columns:` profile (type, non-empty count, min/max, distinct up to 50). Sizes are the extent of non-empty cells (the `info` handle reports the raw worksheet dimensions instead, which may be larger). Rendered views (`images/<stem>-s<idx>.<fmt>`) are produced for sheets carrying charts or images when LibreOffice is on `PATH`: the workbook is exported one PDF page per sheet and rasterized under a 16 Mpx budget. Any LibreOffice or rasterization failure degrades to `Rendered view: unavailable (<reason>)` and a handle note; it never fails the conversion. Workbooks whose chartsheet drawings carry a zero-size anchor (openpyxl-authored files; Excel-authored files are unaffected) render as a degenerate page and are reported as such. `.xls` gets the inventory, CSVs and previews but no visual detection or rendering.
316
318
 
@@ -581,6 +581,9 @@ function ownedCsvPattern(stem) {
581
581
  function ownedPagePattern(stem) {
582
582
  return new RegExp(`^${escRe(stem)}-p\\d+\\.[a-z0-9]+$`);
583
583
  }
584
+ function ownedOcrPattern(stem) {
585
+ return new RegExp(`^${escRe(stem)}-p\\d+\\.md$`);
586
+ }
584
587
  var FILE_LINK_RE = /\[[^\]]*\]\(\s*((?:sheets|attachments)\/[^)\s]+)\s*\)/g;
585
588
  var IMG_LINK_RE = /!\[[^\]]*\]\(\s*(?:<([^>]*)>|([^)]*?))\s*\)|<img\b[^>]*\bsrc\s*=\s*(?:"([^"]*)"|'([^']*)'|([^\s>"']+))/gi;
586
589
  function imageTarget(m) {
@@ -627,6 +630,7 @@ function openBundle(root, requested, overwrite) {
627
630
  const sheetsDir = join2(root, "sheets");
628
631
  const pagesDir = join2(root, "pages");
629
632
  const attachmentsDir = join2(root, "attachments");
633
+ const ocrDir = join2(root, "ocr");
630
634
  try {
631
635
  if (existsSync(mdPath) && overwrite) {
632
636
  const owned = ownedPattern(stem), ownedCsv = ownedCsvPattern(stem), ownedPage = ownedPagePattern(stem);
@@ -644,6 +648,11 @@ function openBundle(root, requested, overwrite) {
644
648
  if (existsSync(attachmentsDir)) {
645
649
  for (const f of readdirSync(attachmentsDir)) if (attachments.has(f)) rmSync(join2(attachmentsDir, f), { force: true });
646
650
  }
651
+ rmSync(join2(root, `${stem}.pages.json`), { force: true });
652
+ const ownedOcr = ownedOcrPattern(stem);
653
+ if (existsSync(ocrDir)) {
654
+ for (const f of readdirSync(ocrDir)) if (ownedOcr.test(f)) rmSync(join2(ocrDir, f), { force: true });
655
+ }
647
656
  }
648
657
  const lockId = randomBytes(6).toString("hex");
649
658
  const stagingDir = join2(imagesDir, `.stage-${lockId}`);
@@ -651,7 +660,8 @@ function openBundle(root, requested, overwrite) {
651
660
  mkdirSync2(stagingDir, { recursive: true });
652
661
  const pagesStagingDir = join2(pagesDir, `.stage-${lockId}`);
653
662
  const attachmentsStagingDir = join2(attachmentsDir, `.stage-${lockId}`);
654
- return { root, stem, renamedFrom, renameReason, mdPath, lockPath, imagesDir, stagingDir, lockId, sheetsDir, sheetsStagingDir, pagesDir, pagesStagingDir, attachmentsDir, attachmentsStagingDir, manifest: /* @__PURE__ */ new Set(), csvManifest: /* @__PURE__ */ new Set(), pageManifest: /* @__PURE__ */ new Set(), attachmentManifest: /* @__PURE__ */ new Set(), sourceMap: /* @__PURE__ */ new Map() };
663
+ const ocrStagingDir = join2(ocrDir, `.stage-${lockId}`);
664
+ return { root, stem, renamedFrom, renameReason, mdPath, lockPath, imagesDir, stagingDir, lockId, sheetsDir, sheetsStagingDir, pagesDir, pagesStagingDir, attachmentsDir, attachmentsStagingDir, ocrDir, ocrStagingDir, pageStatsPath: join2(root, `${stem}.pages.json`), ocrManifest: /* @__PURE__ */ new Set(), manifest: /* @__PURE__ */ new Set(), csvManifest: /* @__PURE__ */ new Set(), pageManifest: /* @__PURE__ */ new Set(), attachmentManifest: /* @__PURE__ */ new Set(), sourceMap: /* @__PURE__ */ new Map() };
655
665
  } catch (e) {
656
666
  rmSync(lockPath, { force: true });
657
667
  throw e;
@@ -715,6 +725,28 @@ function publishPageImages(b, pageCount) {
715
725
  }
716
726
  rmSync(b.pagesStagingDir, { recursive: true, force: true });
717
727
  }
728
+ function writePageStats(b, stats) {
729
+ writeFileSync2(b.pageStatsPath, `${JSON.stringify(stats, null, 2)}
730
+ `, "utf8");
731
+ }
732
+ function publishSidecars(b) {
733
+ const out = /* @__PURE__ */ new Map();
734
+ if (!existsSync(b.ocrStagingDir)) return out;
735
+ for (const dir of readdirSync(b.ocrStagingDir).sort()) {
736
+ const m = dir.match(/^p(\d+)$/);
737
+ if (!m) continue;
738
+ const pageDir = join2(b.ocrStagingDir, dir);
739
+ if (!existsSync(join2(pageDir, ".done"))) continue;
740
+ const file = `${b.stem}-${dir}.md`;
741
+ if (!existsSync(join2(pageDir, file))) continue;
742
+ mkdirSync2(b.ocrDir, { recursive: true });
743
+ renameSync(join2(pageDir, file), join2(b.ocrDir, file));
744
+ b.ocrManifest.add(file);
745
+ out.set(Number(m[1]), join2(b.ocrDir, file));
746
+ }
747
+ rmSync(b.ocrStagingDir, { recursive: true, force: true });
748
+ return out;
749
+ }
718
750
  function publishAttachments(b) {
719
751
  if (!existsSync(b.attachmentsStagingDir)) return;
720
752
  for (const f of readdirSync(b.attachmentsStagingDir).sort()) {
@@ -772,6 +804,10 @@ function commitBundle(b, markdown) {
772
804
  fs.rmSync(b.attachmentsStagingDir, { recursive: true, force: true });
773
805
  } catch {
774
806
  }
807
+ try {
808
+ fs.rmSync(b.ocrStagingDir, { recursive: true, force: true });
809
+ } catch {
810
+ }
775
811
  try {
776
812
  fs.rmSync(b.lockPath, { force: true });
777
813
  } catch {
@@ -782,6 +818,9 @@ function abortBundle(b) {
782
818
  for (const f of b.csvManifest) rmSync(join2(b.sheetsDir, f), { force: true });
783
819
  for (const f of b.pageManifest) rmSync(join2(b.pagesDir, f), { force: true });
784
820
  for (const f of b.attachmentManifest) rmSync(join2(b.attachmentsDir, f), { force: true });
821
+ for (const f of b.ocrManifest) rmSync(join2(b.ocrDir, f), { force: true });
822
+ rmSync(b.ocrStagingDir, { recursive: true, force: true });
823
+ rmSync(b.pageStatsPath, { force: true });
785
824
  rmSync(`${b.mdPath}.tmp`, { force: true });
786
825
  rmSync(b.stagingDir, { recursive: true, force: true });
787
826
  rmSync(b.sheetsStagingDir, { recursive: true, force: true });
@@ -820,7 +859,21 @@ function compactRanges(nums, maxEntries = 20) {
820
859
  }
821
860
  var INSTALL_HINT = "install Tesseract (see doc/doc-to-md.md), then rerun with ocr=true";
822
861
  var BARE_REASONS = ["fallback tier", "no Python backend"];
862
+ var pageList = (pages) => `page${pages.length === 1 ? "" : "s"} ${pages.join(", ")}`;
863
+ function forcedOcrLine(ocr) {
864
+ const clauses = [];
865
+ if (ocr.pages.length) clauses.push(`sidecars for ${pageList(ocr.pages)}`);
866
+ if (ocr.noText.length) clauses.push(`no text on ${pageList(ocr.noText)}`);
867
+ for (const page of ocr.ocrFailed) clauses.push(`failed on page ${page} (${ocr.ocrErrors[page] ?? "unknown error"})`);
868
+ if (ocr.budgetStopped.length) clauses.push(`budget-stopped ${pageList(ocr.budgetStopped)}`);
869
+ if (ocr.killed !== null) clauses.push(`page ${ocr.killed} killed the OCR child (timeout or crash - likely a compression bomb)`);
870
+ if (ocr.childError !== null) clauses.push(`OCR child failed before processing pages: ${ocr.childError}`);
871
+ if (ocr.notAttempted.length) clauses.push(`${pageList(ocr.notAttempted)} not attempted`);
872
+ const rerun = [...ocr.budgetStopped, ...ocr.notAttempted].sort((a, b) => a - b);
873
+ return `OCR: forced (${ocr.lang}) - ${clauses.join("; ")}${rerun.length ? ` - re-run with --pages ${rerun.join(",")}` : ""}`;
874
+ }
823
875
  function ocrLine(ocr, type) {
876
+ if (ocr.mode === "all") return forcedOcrLine(ocr);
824
877
  const image = type === "image";
825
878
  const rerun = ocr.tesseract ? "rerun with ocr=true" : INSTALL_HINT;
826
879
  switch (ocr.status) {
@@ -896,6 +949,8 @@ function formatHandle(h) {
896
949
  if (h.sheetsDir) lines.push(`Sheets-Dir: ${h.sheetsDir}`);
897
950
  if (h.pagesDir && h.pageImageCount > 0) lines.push(`Pages-Dir: ${h.pagesDir} (${h.pageImageCount} pages)`);
898
951
  else if (h.pageImagesReason) lines.push(`Pages-Dir: none - ${h.pageImagesReason}`);
952
+ if (h.pageStatsPath) lines.push(`Page-Stats: ${h.pageStatsPath}`);
953
+ if (h.ocrDir) lines.push(`OCR-Dir: ${h.ocrDir}`);
899
954
  lines.push(`Type: ${h.type} Engine: ${h.engine} Tier: ${h.tier}`);
900
955
  lines.push(`Page-Count: ${pageCountLabel(h)} Pages: ${h.pages ? compactRanges(h.pages) : "all"} Images: ${h.imageCount} Size: ${formatSize2(h.bytes)} / ${h.lines} lines`);
901
956
  if (h.degraded) lines.push(`Degraded: ${h.degraded}`);
@@ -945,6 +1000,7 @@ var DOC_TO_MD_OPTIONS = [
945
1000
  { key: "outputDir", flag: "--output-dir", type: "string", default: null, settable: false, help: "Bundle root for <stem>.md + images/; default a per-call temp dir. <stem> = basename without extension with [^A-Za-z0-9._-]+ -> _ (empty -> document); a second call on the same stem writes <stem>-2.md" },
946
1001
  { key: "overwrite", flag: "--overwrite", type: "bool", default: false, settable: false, help: "Replace an existing completed <stem>.md bundle" },
947
1002
  { key: "pageImages", flag: "--page-images", type: "bool", default: false, settable: false, help: "Also render every selected page to pages/<stem>-pNNN.<imageFormat> at imageDpi (PDF, PPTX, .doc, DOCX via LibreOffice); off by default" },
1003
+ { key: "ocrMode", flag: "--ocr-mode", type: "enum", default: "textless", settable: false, enumValues: ["textless", "all"], help: "OCR policy: textless (default) OCRs only pages with an empty text layer, inline; all OCRs every selected page and writes the recognized text to ocr/<stem>-pNNN.md sidecars, leaving the Markdown untouched. all requires --ocr and an explicit --pages selection (PDF, PPTX, DOC)." },
948
1004
  { key: "primaryTimeoutMs", flag: "--primary-timeout", type: "int", default: 6e4, settable: true, env: "PI_DOC_TO_MD_CONVERT_TIMEOUT_MS", help: "pymupdf4llm tier and DOCX child (docx mode); also the unpdf tier" },
949
1005
  { key: "fallbackTimeoutMs", flag: "--fallback-timeout", type: "int", default: 3e4, settable: true, help: "PyMuPDF get_text tier (including DOCX LibreOffice fallback); also PDF and DOCX info and Excel rendered views" },
950
1006
  { key: "sofficeTimeoutMs", flag: "--soffice-timeout", type: "int", default: 12e4, settable: true, env: "PI_DOC_TO_MD_SOFFICE_TIMEOUT_MS", help: "DOCX/PPTX -> PDF via LibreOffice; also Excel rendered views" },
@@ -1051,11 +1107,28 @@ function resolveOptions(perCall, settings, env) {
1051
1107
  out[d.key] = value;
1052
1108
  }
1053
1109
  const o = out;
1054
- if (o.info && (o.pages !== null || o.outputDir !== null || o.overwrite || o.pageImages)) {
1055
- throw new UsageError("--info cannot be combined with --pages, --output-dir, --overwrite or --page-images");
1110
+ if (o.info && (o.pages !== null || o.outputDir !== null || o.overwrite || o.pageImages || o.ocrMode === "all")) {
1111
+ throw new UsageError("--info cannot be combined with --pages, --output-dir, --overwrite, --page-images or --ocr-mode all");
1056
1112
  }
1057
1113
  return o;
1058
1114
  }
1115
+ var USAGE_PATTERNS = [
1116
+ "Two-pass OCR (PDF, PPTX, DOC):",
1117
+ " 1. <cmd> report.pdf --output-dir out --json",
1118
+ ' -> "pageStatsPath" points at out/report.pages.json; pages with few',
1119
+ ' chars and high imageCoverage are scans. "savedTo" is the Markdown.',
1120
+ " 2. <cmd> report.pdf --output-dir out --ocr --ocr-mode all --pages 2,7 --json",
1121
+ ' -> "ocr"."sidecars" maps 2 and 7 to out/ocr/report-2-p002.md and',
1122
+ " ...-p007.md (a second run in the same dir gets stem report-2);",
1123
+ " the Markdown of this run holds pages 2 and 7 only and equals what",
1124
+ " --pages 2,7 without --ocr would produce. Read sidecars and Markdown",
1125
+ " by the returned paths, never by guessing names.",
1126
+ " 3. --ocr-mode all refuses to run without --ocr and an explicit --pages.",
1127
+ ' The "ocr" object (OCR: line) names failed, budget-stopped, killed and',
1128
+ " not-attempted pages and the exact --pages to re-run.",
1129
+ " Details: doc/doc-to-md.md (bundle contract, failure buckets)."
1130
+ ].join("\n");
1131
+ var usagePatterns = (cmd) => USAGE_PATTERNS.replaceAll("<cmd>", cmd);
1059
1132
  function renderHelp() {
1060
1133
  const row = (d) => ` ${(d.flag ?? "<path>").padEnd(26)} ${d.help}${d.default !== null && d.key !== "info" && d.key !== "overwrite" ? ` (default ${d.default})` : ""}`;
1061
1134
  return [
@@ -1067,8 +1140,10 @@ function renderHelp() {
1067
1140
  "Tunables (also settable under quiver.docToMd in settings.json; per-call > settings > default):",
1068
1141
  ...DOC_TO_MD_OPTIONS.filter((d) => d.settable).map(row),
1069
1142
  "",
1070
- "Result: a handle (Saved-To, Images-Dir, Page-Count, Outline ...). Read the Saved-To file for the Markdown.",
1071
- "Exit codes: 0 success, 1 runtime error, 2 usage error."
1143
+ "Result: a handle (Saved-To, Images-Dir, Page-Stats, Page-Count, Outline ...). Read the Saved-To file for the Markdown.",
1144
+ "Exit codes: 0 success, 1 runtime error, 2 usage error.",
1145
+ "",
1146
+ usagePatterns("pi-quiver doc-to-md")
1072
1147
  ].join("\n");
1073
1148
  }
1074
1149
 
@@ -1547,7 +1622,38 @@ function reconcileRenderMarkers(md, renderPages, fmt, sourceMap, reason) {
1547
1622
  if (/<!--rvs?:\d+-->/.test(md)) throw new Error("internal: unresolved render marker");
1548
1623
  return md;
1549
1624
  }
1550
- var emptyOcr = (lang) => ({ status: "off", lang, textless: [], pages: [], noText: [], ocrFailed: [], budgetStopped: [], reason: null, tesseract: null });
1625
+ var emptyOcr = (lang) => ({ status: "off", lang, textless: [], pages: [], noText: [], ocrFailed: [], budgetStopped: [], reason: null, tesseract: null, mode: "textless", sidecars: {}, ocrErrors: {}, killed: null, notAttempted: [], childError: null });
1626
+ var emptyOutcome = () => ({ written: [], noText: [], ocrFailed: [], ocrErrors: {}, budgetStopped: [], killed: null, notAttempted: [], childError: null });
1627
+ var ocrPagesTag = (n) => `p${String(n).padStart(3, "0")}`;
1628
+ var SIDE_MARKER_RE = /^--- end of page\.page_number=\d+ ---$/;
1629
+ function sidecarHasText(dir) {
1630
+ const file = readdirSync2(dir).find((f) => f.endsWith(".md"));
1631
+ if (!file) return false;
1632
+ return readFileSync2(join3(dir, file), "utf8").split("\n").slice(1).some((l) => l.trim() && !SIDE_MARKER_RE.test(l));
1633
+ }
1634
+ function recoverOcrPages(stagingDir, pages, detail) {
1635
+ const out = emptyOutcome();
1636
+ const activePath = join3(stagingDir, "active");
1637
+ const active = existsSync2(activePath) ? Number(readFileSync2(activePath, "utf8").trim()) || null : null;
1638
+ let sawPage = false;
1639
+ for (const n of pages) {
1640
+ const dir = join3(stagingDir, ocrPagesTag(n));
1641
+ if (existsSync2(join3(dir, ".done"))) {
1642
+ sawPage = true;
1643
+ (sidecarHasText(dir) ? out.written : out.noText).push(n);
1644
+ } else if (existsSync2(join3(dir, ".failed"))) {
1645
+ sawPage = true;
1646
+ out.ocrFailed.push(n);
1647
+ out.ocrErrors[n] = readFileSync2(join3(dir, ".failed"), "utf8").trim() || "unknown error";
1648
+ } else if (n === active) out.killed = n;
1649
+ else out.notAttempted.push(n);
1650
+ }
1651
+ if (active === null && !sawPage) out.childError = detail;
1652
+ return out;
1653
+ }
1654
+ function outcomeFromChild(j) {
1655
+ return { ...emptyOutcome(), written: j.written ?? [], noText: j.noText ?? [], ocrFailed: j.ocrFailed ?? [], ocrErrors: j.ocrErrors ?? {}, budgetStopped: j.budgetStopped ?? [] };
1656
+ }
1551
1657
  var OCR_SENTINEL_RE = /\x00OCR ([^\x00]*)\x00/g;
1552
1658
  function resolveOcrLabels(md, sourceMap) {
1553
1659
  return md.replace(OCR_SENTINEL_RE, (_, key) => {
@@ -1559,7 +1665,7 @@ function handleOcr(tier, type, o, json) {
1559
1665
  if (tier === "unpdf") return o.ocr ? { ...emptyOcr(o.ocrLanguage), status: "unavailable", reason: "no Python backend" } : null;
1560
1666
  const x = json.ocr;
1561
1667
  if (!x || type !== "image" && !x.textless.length && !x.pages.length && !x.ocrFailed.length) return null;
1562
- return x;
1668
+ return { ...emptyOcr(o.ocrLanguage), ...x };
1563
1669
  }
1564
1670
  async function convertDocument(o, signal, seams) {
1565
1671
  const s = { backend: (c) => getBackend(c, void 0, signal), runTier: runTierReal, office: tryConvertOffice, ...seams };
@@ -1568,10 +1674,17 @@ async function convertDocument(o, signal, seams) {
1568
1674
  if (!st || !st.isFile()) throw new Error(`Not a readable file: ${o.path}`);
1569
1675
  if (st.size === 0) throw new Error(`empty file: ${o.path}`);
1570
1676
  const type = classifyInput(inputPath);
1677
+ const forced = o.ocrMode === "all";
1678
+ if (forced) {
1679
+ if (!o.ocr) throw new UsageError("--ocr-mode all requires --ocr");
1680
+ if (o.pages === null) throw new UsageError('--ocr-mode all requires an explicit --pages selection (e.g. --pages 2,7); omitted pages and --pages "" mean all pages and are refused to keep OCR cost bounded');
1681
+ if (type !== "pdf" && type !== "pptx" && type !== "doc") throw new UsageError("--ocr-mode all applies to PDF, PPTX and DOC inputs only (DOCX pages are page-break segments, not PDF pages; convert the DOCX to PDF first)");
1682
+ }
1571
1683
  if (o.pages && (type === "html" || type === "image" || type === "email")) throw new Error(`--pages does not apply to ${type === "html" ? "HTML files" : type === "image" ? "images" : "email"}`);
1572
1684
  const isExcel = type === "xlsx" || type === "xlsm" || type === "xls";
1573
1685
  if (isExcel && o.pages) throw new Error("--pages does not apply to spreadsheets: worksheets have no stable page numbering");
1574
1686
  const backend = await s.backend({ pymupdfVersion: o.pymupdfVersion, warmTimeoutMs: o.warmTimeoutMs });
1687
+ if (forced && backend.kind === "none") throw new Error(`--ocr-mode all cannot run: no Python backend (${backend.reason})`);
1575
1688
  if (isExcel && (backend.kind === "none" || !backend.xlsx)) throw new Error(`Excel conversion needs a Python backend with openpyxl, xlrd and pillow (${backend.kind === "none" ? backend.reason : `${backend.kind} lacks the Excel packages`}). ${EXCEL_REMEDY}`);
1576
1689
  if (type === "email" && extname3(inputPath).toLowerCase() === ".msg" && (backend.kind === "none" || !backend.email)) throw new Error(`MSG conversion needs the extract-msg package. Python backend: ${backendState(backend, "found without extract-msg")}. Remedy: install uv, or pip install extract-msg markdownify into that Python`);
1577
1690
  if (type === "email" && extname3(inputPath).toLowerCase() !== ".msg" && lacksDocx(backend)) throw new Error(`EML conversion needs the Python DOCX/HTML packages (mammoth, markdownify, python-docx). Python backend: ${backendState(backend, "found without mammoth/markdownify/python-docx")}. ${docxRemedy}`);
@@ -1580,7 +1693,7 @@ async function convertDocument(o, signal, seams) {
1580
1693
  let office = null;
1581
1694
  try {
1582
1695
  let pdfPath = inputPath;
1583
- const base = { path: inputPath, pages: o.pages, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
1696
+ const base = { path: inputPath, pages: o.pages, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr && !forced, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
1584
1697
  let tier, engine, json, degraded = null, fallbackReason = null;
1585
1698
  let explicitBreaks = null;
1586
1699
  let notes = [];
@@ -1757,6 +1870,17 @@ async function convertDocument(o, signal, seams) {
1757
1870
  if (type === "doc" || officeRoute !== null) degraded = DEGRADED_DOCX_OFFICE;
1758
1871
  if (officeRoute !== null) fallbackReason = fallbackReason ? `${officeRoute}; ${fallbackReason}` : officeRoute;
1759
1872
  if (tier === void 0 || engine === void 0 || json === void 0) throw new Error("internal: no tier produced output");
1873
+ const pageStats = json.pageStats ?? null;
1874
+ if (pageStats) writePageStats(b, pageStats);
1875
+ let ocr = handleOcr(tier, type, o, json);
1876
+ if (forced) {
1877
+ const r = await s.runTier("ocr-pages", { path: pdfPath, pages: o.pages, stem: b.stem, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs, stagingDir: b.ocrStagingDir, dpi: o.imageDpi, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion }, b, signal, o.primaryTimeoutMs, backend);
1878
+ if (signal?.aborted || !r.ok && "reason" in r && r.reason === "aborted") throw new Error("aborted");
1879
+ if (r.ok && r.json.status === "unavailable") throw new Error(`OCR unavailable: ${r.json.reason} (install Tesseract; see doc/doc-to-md.md)`);
1880
+ const outcome = r.ok && r.json.status === "ran" ? outcomeFromChild(r.json) : recoverOcrPages(b.ocrStagingDir, o.pages, !r.ok ? "userError" in r ? r.userError : `${r.reason}${detailSuffix(r)}` : "malformed child output");
1881
+ const sidecars = publishSidecars(b);
1882
+ ocr = { ...emptyOcr(o.ocrLanguage), ...json.ocr, status: "ran", reason: null, mode: "all", pages: outcome.written, noText: outcome.noText, ocrFailed: outcome.ocrFailed, ocrErrors: outcome.ocrErrors, budgetStopped: outcome.budgetStopped, killed: outcome.killed, notAttempted: outcome.notAttempted, childError: outcome.childError, sidecars: Object.fromEntries(sidecars) };
1883
+ }
1760
1884
  if (!isExcel) notes = [...notes, ...json.notes ?? []];
1761
1885
  if (b.renamedFrom) notes.splice(notes[0]?.startsWith("preview truncated:") ? 1 : 0, 0, `renamed to ${b.stem} (${b.renameReason})`);
1762
1886
  const pageImagesReason = !o.pageImages || b.pageManifest.size || tier === "primary" || tier === "fallback" ? null : tier === "unpdf" ? "page images need the Python backend" : `${type} has no page geometry`;
@@ -1773,7 +1897,7 @@ async function convertDocument(o, signal, seams) {
1773
1897
  ` : "") + body;
1774
1898
  commitBundle(b, markdown);
1775
1899
  const outline = scanOutline(markdown, o.outlineMaxEntries);
1776
- const details = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr: handleOcr(tier, type, o, json) };
1900
+ const details = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr, pageStats, pageStatsPath: pageStats ? b.pageStatsPath : null, ocrDir: b.ocrManifest.size ? b.ocrDir : null };
1777
1901
  return { output: formatHandle(details), details };
1778
1902
  } catch (e) {
1779
1903
  abortBundle(b);
@@ -1819,7 +1943,7 @@ async function inspectDocument(o, signal, seams) {
1819
1943
  }
1820
1944
 
1821
1945
  // bin/pi-quiver.ts
1822
- var USAGE = 'Usage: pi-quiver fetch <url> [--method GET|HEAD|POST] [--header "K: V"]... [--body <str>] [--raw] [--timeout-ms <n>]\n pi-quiver doc-to-md [--json] [--info] [--page-images] [--pages <spec>] [--output-dir <dir>] [--overwrite] [tunable flags] <path> (--help for all flags)';
1946
+ var USAGE = 'Usage: pi-quiver fetch <url> [--method GET|HEAD|POST] [--header "K: V"]... [--body <str>] [--raw] [--timeout-ms <n>]\n pi-quiver doc-to-md [--json] [--info] [--page-images] [--pages <spec>] [--output-dir <dir>] [--overwrite] [--ocr] [--ocr-mode textless|all] [tunable flags] <path> (--help for all flags)';
1823
1947
  function parseDocToMd(rest) {
1824
1948
  if (rest.includes("--help") || rest.includes("-h")) return { ok: true, cmd: "doc-to-md-help" };
1825
1949
  const json = rest.includes("--json");
@@ -1974,6 +2098,12 @@ ${USAGE}
1974
2098
  `);
1975
2099
  return 0;
1976
2100
  } catch (err) {
2101
+ if (err instanceof UsageError) {
2102
+ process.stderr.write(`${err.message}
2103
+ ${USAGE}
2104
+ `);
2105
+ return 2;
2106
+ }
1977
2107
  process.stderr.write(`doc-to-md failed: ${err instanceof Error ? err.message : String(err)}
1978
2108
  `);
1979
2109
  return 1;
@@ -35,7 +35,7 @@ export default function docToMdExtension(pi: ExtensionAPI) {
35
35
  name: "doc_to_md",
36
36
  label: "Convert doc to Markdown bundle",
37
37
  description:
38
- "Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, email (.msg/.eml), HTML (.html/.htm) or image (.png .jpg .jpeg .tif .tiff .bmp .gif) to a Markdown bundle on disk and return a handle (Saved-To, Images-Dir, Page-Count, Outline, diagnostics) - the Markdown itself is never inlined; read the Saved-To file (offset/limit) for content. `info: true` returns page count, metadata and TOC (or the sheet inventory) without converting - use it to pick `pages`. `pages` selects inclusive 1-based pages (PDF/PPTX) or explicit-page-break segments (DOCX; rejected when the file has none); every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Images are always extracted into images/ with relative links. Primary engine pymupdf4llm, fallback PyMuPDF text (degraded, marked), pure-JS unpdf only when no Python backend exists. Excel yields a sheet inventory (worksheets + chartsheets), a full CSV per non-empty worksheet under sheets/, a bounded preview with formulas and cached values, merged/hidden disclosure, and rendered chart views when LibreOffice is available. DOCX converts directly (mammoth; python-docx text fallback) and keeps heading styles, so the handle Outline lists `L<line>` and `p<segment>` per heading; LibreOffice (soffice) is optional for DOCX (fallback route, degraded) and Excel rendered views, and required for PPTX. Use `outputDir` for a durable bundle; without it the bundle lands in a per-call temp dir. Input must be a local file path (use fetch first for URLs). HTML converts without Readability (markdownify; Turndown fallback): local and data: images are copied into images/, remote images stay links. Pages without a text layer keep their picture in images/, or pages/ with `pageImages: true`; image inputs keep their picture in images/. `pageImages: true` writes page renders to pages/<stem>-pNNN, and email attachments are saved under attachments/. OCR is off by default: `ocr: true` adds recognized text (labeled as OCR) when Tesseract language data is installed; the handle's OCR: line says what ran.",
38
+ "Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, email (.msg/.eml), HTML (.html/.htm) or image (.png .jpg .jpeg .tif .tiff .bmp .gif) to a Markdown bundle on disk and return a handle (Saved-To, Images-Dir, Page-Count, Outline, diagnostics) - the Markdown itself is never inlined; read the Saved-To file (offset/limit) for content. `info: true` returns page count, metadata and TOC (or the sheet inventory) without converting - use it to pick `pages`. `pages` selects inclusive 1-based pages (PDF/PPTX) or explicit-page-break segments (DOCX; rejected when the file has none); every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Images are always extracted into images/ with relative links. Primary engine pymupdf4llm, fallback PyMuPDF text (degraded, marked), pure-JS unpdf only when no Python backend exists. Excel yields a sheet inventory (worksheets + chartsheets), a full CSV per non-empty worksheet under sheets/, a bounded preview with formulas and cached values, merged/hidden disclosure, and rendered chart views when LibreOffice is available. DOCX converts directly (mammoth; python-docx text fallback) and keeps heading styles, so the handle Outline lists `L<line>` and `p<segment>` per heading; LibreOffice (soffice) is optional for DOCX (fallback route, degraded) and Excel rendered views, and required for PPTX. Use `outputDir` for a durable bundle; without it the bundle lands in a per-call temp dir. Input must be a local file path (use fetch first for URLs). HTML converts without Readability (markdownify; Turndown fallback): local and data: images are copied into images/, remote images stay links. Pages without a text layer keep their picture in images/, or pages/ with `pageImages: true`; image inputs keep their picture in images/. `pageImages: true` writes page renders to pages/<stem>-pNNN, and email attachments are saved under attachments/. OCR is off by default: `ocr: true` adds recognized text (labeled as OCR) when Tesseract language data is installed; the handle's OCR: line says what ran. Two-pass OCR: read Page-Stats first, then re-run with ocr: true, ocrMode: \"all\" and an explicit pages selection to get ocr/ sidecars for the pages you name.",
39
39
  promptSnippet: "Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS/MSG/EML/HTML/image to a Markdown bundle (handle returned; read Saved-To)",
40
40
  parameters: buildSchema(),
41
41
 
@@ -1,23 +1,26 @@
1
1
  /**
2
2
  * Bundle protocol: a call owns `<stem>` for its whole duration via `<stem>.md.lock`;
3
- * children stage assets under `images/`, `sheets/`, `pages/`, and `attachments/` staging dirs;
4
- * Node publishes them to stem-prefixed files in those four dirs and records every file it
3
+ * children stage assets under `images/`, `sheets/`, `pages/`, `attachments/`, and `ocr/` staging dirs;
4
+ * Node publishes them to stem-prefixed files in those asset dirs and records every file it
5
5
  * wrote in a manifest, and commits `<stem>.md` atomically (tmp + rename).
6
6
  */
7
7
  import fs, { closeSync, existsSync, mkdirSync, openSync, readdirSync, readFileSync, renameSync, rmSync, statSync, unlinkSync, writeFileSync } from "node:fs";
8
8
  import { tmpdir } from "node:os";
9
9
  import { extname, join, resolve } from "node:path";
10
10
  import { randomBytes } from "node:crypto";
11
+ import type { PageStat } from "./doc-to-md-handle.ts";
11
12
 
12
13
  export interface Bundle {
13
14
  root: string; stem: string; renamedFrom: string | null; renameReason: string | null; mdPath: string; lockPath: string; imagesDir: string; stagingDir: string; lockId: string;
14
15
  sheetsDir: string; sheetsStagingDir: string;
15
16
  pagesDir: string; pagesStagingDir: string;
16
17
  attachmentsDir: string; attachmentsStagingDir: string;
18
+ ocrDir: string; ocrStagingDir: string; pageStatsPath: string;
17
19
  manifest: Set<string>;
18
20
  csvManifest: Set<string>;
19
21
  pageManifest: Set<string>;
20
22
  attachmentManifest: Set<string>;
23
+ ocrManifest: Set<string>;
21
24
  sourceMap: Map<string, string>;
22
25
  }
23
26
 
@@ -33,6 +36,8 @@ export function ownedCsvPattern(stem: string): RegExp {
33
36
 
34
37
  export function ownedPagePattern(stem: string): RegExp { return new RegExp(`^${escRe(stem)}-p\\d+\\.[a-z0-9]+$`); }
35
38
 
39
+ export function ownedOcrPattern(stem: string): RegExp { return new RegExp(`^${escRe(stem)}-p\\d+\\.md$`); }
40
+
36
41
  const FILE_LINK_RE = /\[[^\]]*\]\(\s*((?:sheets|attachments)\/[^)\s]+)\s*\)/g;
37
42
  const IMG_LINK_RE = /!\[[^\]]*\]\(\s*(?:<([^>]*)>|([^)]*?))\s*\)|<img\b[^>]*\bsrc\s*=\s*(?:"([^"]*)"|'([^']*)'|([^\s>"']+))/gi;
38
43
 
@@ -75,6 +80,7 @@ export function openBundle(root: string, requested: string, overwrite: boolean):
75
80
  const sheetsDir = join(root, "sheets");
76
81
  const pagesDir = join(root, "pages");
77
82
  const attachmentsDir = join(root, "attachments");
83
+ const ocrDir = join(root, "ocr");
78
84
  try {
79
85
  if (existsSync(mdPath) && overwrite) {
80
86
  const owned = ownedPattern(stem), ownedCsv = ownedCsvPattern(stem), ownedPage = ownedPagePattern(stem);
@@ -87,6 +93,9 @@ export function openBundle(root: string, requested: string, overwrite: boolean):
87
93
  if (existsSync(sheetsDir)) for (const f of readdirSync(sheetsDir)) if (ownedCsv.test(f)) rmSync(join(sheetsDir, f), { force: true });
88
94
  if (existsSync(pagesDir)) for (const f of readdirSync(pagesDir)) if (ownedPage.test(f)) rmSync(join(pagesDir, f), { force: true });
89
95
  if (existsSync(attachmentsDir)) for (const f of readdirSync(attachmentsDir)) if (attachments.has(f)) rmSync(join(attachmentsDir, f), { force: true });
96
+ rmSync(join(root, `${stem}.pages.json`), { force: true });
97
+ const ownedOcr = ownedOcrPattern(stem);
98
+ if (existsSync(ocrDir)) for (const f of readdirSync(ocrDir)) if (ownedOcr.test(f)) rmSync(join(ocrDir, f), { force: true });
90
99
  }
91
100
  const lockId = randomBytes(6).toString("hex");
92
101
  const stagingDir = join(imagesDir, `.stage-${lockId}`);
@@ -94,7 +103,8 @@ export function openBundle(root: string, requested: string, overwrite: boolean):
94
103
  mkdirSync(stagingDir, { recursive: true });
95
104
  const pagesStagingDir = join(pagesDir, `.stage-${lockId}`);
96
105
  const attachmentsStagingDir = join(attachmentsDir, `.stage-${lockId}`);
97
- return { root, stem, renamedFrom, renameReason, mdPath, lockPath, imagesDir, stagingDir, lockId, sheetsDir, sheetsStagingDir, pagesDir, pagesStagingDir, attachmentsDir, attachmentsStagingDir, manifest: new Set(), csvManifest: new Set(), pageManifest: new Set(), attachmentManifest: new Set(), sourceMap: new Map() };
106
+ const ocrStagingDir = join(ocrDir, `.stage-${lockId}`);
107
+ return { root, stem, renamedFrom, renameReason, mdPath, lockPath, imagesDir, stagingDir, lockId, sheetsDir, sheetsStagingDir, pagesDir, pagesStagingDir, attachmentsDir, attachmentsStagingDir, ocrDir, ocrStagingDir, pageStatsPath: join(root, `${stem}.pages.json`), ocrManifest: new Set(), manifest: new Set(), csvManifest: new Set(), pageManifest: new Set(), attachmentManifest: new Set(), sourceMap: new Map() };
98
108
  } catch (e) { rmSync(lockPath, { force: true }); throw e; }
99
109
  }
100
110
 
@@ -160,6 +170,30 @@ export function publishPageImages(b: Bundle, pageCount: number): void {
160
170
  rmSync(b.pagesStagingDir, { recursive: true, force: true });
161
171
  }
162
172
 
173
+ export function writePageStats(b: Bundle, stats: PageStat[]): void {
174
+ writeFileSync(b.pageStatsPath, `${JSON.stringify(stats, null, 2)}\n`, "utf8");
175
+ }
176
+
177
+ /** Move every `.done`-gated `pNNN/<stem>-pNNN.md` into `ocr/`; partial page dirs and the checkpoint are dropped with the staging dir. Returns page -> absolute sidecar path. */
178
+ export function publishSidecars(b: Bundle): Map<number, string> {
179
+ const out = new Map<number, string>();
180
+ if (!existsSync(b.ocrStagingDir)) return out;
181
+ for (const dir of readdirSync(b.ocrStagingDir).sort()) {
182
+ const m = dir.match(/^p(\d+)$/);
183
+ if (!m) continue;
184
+ const pageDir = join(b.ocrStagingDir, dir);
185
+ if (!existsSync(join(pageDir, ".done"))) continue;
186
+ const file = `${b.stem}-${dir}.md`;
187
+ if (!existsSync(join(pageDir, file))) continue;
188
+ mkdirSync(b.ocrDir, { recursive: true });
189
+ renameSync(join(pageDir, file), join(b.ocrDir, file));
190
+ b.ocrManifest.add(file);
191
+ out.set(Number(m[1]), join(b.ocrDir, file));
192
+ }
193
+ rmSync(b.ocrStagingDir, { recursive: true, force: true });
194
+ return out;
195
+ }
196
+
163
197
  export function publishAttachments(b: Bundle): void {
164
198
  if (!existsSync(b.attachmentsStagingDir)) return;
165
199
  for (const f of readdirSync(b.attachmentsStagingDir).sort()) {
@@ -208,6 +242,7 @@ export function commitBundle(b: Bundle, markdown: string): void {
208
242
  try { fs.rmSync(b.sheetsStagingDir, { recursive: true, force: true }); } catch { /* best-effort */ }
209
243
  try { fs.rmSync(b.pagesStagingDir, { recursive: true, force: true }); } catch { /* best-effort */ }
210
244
  try { fs.rmSync(b.attachmentsStagingDir, { recursive: true, force: true }); } catch { /* best-effort */ }
245
+ try { fs.rmSync(b.ocrStagingDir, { recursive: true, force: true }); } catch { /* best-effort */ }
211
246
  try { fs.rmSync(b.lockPath, { force: true }); } catch { /* Markdown is published; cleanup is best-effort. */ }
212
247
  }
213
248
 
@@ -216,6 +251,9 @@ export function abortBundle(b: Bundle): void {
216
251
  for (const f of b.csvManifest) rmSync(join(b.sheetsDir, f), { force: true });
217
252
  for (const f of b.pageManifest) rmSync(join(b.pagesDir, f), { force: true });
218
253
  for (const f of b.attachmentManifest) rmSync(join(b.attachmentsDir, f), { force: true });
254
+ for (const f of b.ocrManifest) rmSync(join(b.ocrDir, f), { force: true });
255
+ rmSync(b.ocrStagingDir, { recursive: true, force: true });
256
+ rmSync(b.pageStatsPath, { force: true });
219
257
  rmSync(`${b.mdPath}.tmp`, { force: true });
220
258
  rmSync(b.stagingDir, { recursive: true, force: true });
221
259
  rmSync(b.sheetsStagingDir, { recursive: true, force: true });
@@ -12,9 +12,9 @@ import { type ChildProcess, spawn } from "node:child_process";
12
12
  import { homedir, tmpdir } from "node:os";
13
13
  import { basename, dirname, extname, isAbsolute, join, resolve } from "node:path";
14
14
  import { fileURLToPath, pathToFileURL } from "node:url";
15
- import { type Bundle, abortBundle, commitBundle, openBundle, publishAttachments, publishPageImages, publishSheetCsvs, publishSheetImages, publishStaged, rewriteLinks, tempBundleRoot, validateImageLinks } from "./doc-to-md-bundle.ts";
16
- import { type Engine, type HandleData, type InfoData, type OcrInfo, type Tier, formatHandle, formatInfoHandle, scanOutline, type SheetInfo, type TocEntry } from "./doc-to-md-handle.ts";
17
- import { type DocToMdOptions, type InputType, IMAGE_EXTS, TUNABLE_DEFAULTS, classifyInput, sanitizeStem } from "./doc-to-md-options.ts";
15
+ import { type Bundle, abortBundle, commitBundle, openBundle, publishAttachments, publishSidecars, writePageStats, publishPageImages, publishSheetCsvs, publishSheetImages, publishStaged, rewriteLinks, tempBundleRoot, validateImageLinks } from "./doc-to-md-bundle.ts";
16
+ import { type Engine, type HandleData, type InfoData, type OcrInfo, type PageStat, type Tier, formatHandle, formatInfoHandle, scanOutline, type SheetInfo, type TocEntry } from "./doc-to-md-handle.ts";
17
+ import { type DocToMdOptions, type InputType, IMAGE_EXTS, TUNABLE_DEFAULTS, UsageError, classifyInput, sanitizeStem } from "./doc-to-md-options.ts";
18
18
 
19
19
  export * from "./doc-to-md-options.ts";
20
20
  export { compactRanges, formatHandle, formatInfoHandle, formatSize, scanOutline } from "./doc-to-md-handle.ts";
@@ -466,8 +466,8 @@ const lacksDocx = (b: Backend) => b.kind === "none" || !b.docx;
466
466
  const clearStaging = (b: Pick<Bundle, "stagingDir">) => { for (const f of readdirSync(b.stagingDir)) rmSync(join(b.stagingDir, f), { recursive: true, force: true }); };
467
467
  export const EXCEL_REMEDY = "Remedy: install uv, or pip install openpyxl xlrd pillow";
468
468
 
469
- export type Mode = "html" | "image" | "info" | "pdf-primary" | "pdf-fallback" | "xlsx" | "pdf-text" | "render-pages" | "docx" | "email";
470
- export interface TierJson { pageImages?: { page: number; file: string }[]; ocr?: OcrInfo; markdown?: string; pages?: number[]; pageCount?: number; emptyPages?: number[]; failedPages?: { page: number; error: string }[]; notes?: string[]; images?: { sheetIndex: number; file: string }[]; metadata?: Record<string, string>; toc?: [number, string, number | null][]; explicitBreaks?: number; engine?: string; degraded?: boolean; fallbackReason?: string | null; sheets?: SheetInfo[]; renderPages?: number[]; sheetCount?: number; ok?: boolean; reason?: string; rendered?: { idx: number; file: string; dpi: number }[]; failed?: { idx: number; reason: string }[]; }
469
+ export type Mode = "html" | "image" | "info" | "pdf-primary" | "pdf-fallback" | "xlsx" | "pdf-text" | "render-pages" | "docx" | "email" | "ocr-pages";
470
+ export interface TierJson { pageStats?: PageStat[]; status?: string; written?: number[]; noText?: number[]; ocrFailed?: number[]; ocrErrors?: Record<string, string>; budgetStopped?: number[]; pageImages?: { page: number; file: string }[]; ocr?: OcrInfo; markdown?: string; pages?: number[]; pageCount?: number; emptyPages?: number[]; failedPages?: { page: number; error: string }[]; notes?: string[]; images?: { sheetIndex: number; file: string }[]; metadata?: Record<string, string>; toc?: [number, string, number | null][]; explicitBreaks?: number; engine?: string; degraded?: boolean; fallbackReason?: string | null; sheets?: SheetInfo[]; renderPages?: number[]; sheetCount?: number; ok?: boolean; reason?: string; rendered?: { idx: number; file: string; dpi: number }[]; failed?: { idx: number; reason: string }[]; }
471
471
  export type TierResult = { ok: true; json: TierJson } | { ok: false; reason: string; detail?: string } | { ok: false; userError: string; pageCount?: number };
472
472
 
473
473
  export interface PipelineSeams {
@@ -490,7 +490,7 @@ function unpdfWorkerPath(): string {
490
490
  ], existsSync);
491
491
  }
492
492
 
493
- async function runTierReal(mode: Mode, childOptions: Record<string, unknown>, _b: Pick<Bundle, "stagingDir">, signal: AbortSignal | undefined, timeoutMs: number, backend: Backend): Promise<TierResult> {
493
+ export async function runTierReal(mode: Mode, childOptions: Record<string, unknown>, _b: Pick<Bundle, "stagingDir">, signal: AbortSignal | undefined, timeoutMs: number, backend: Backend): Promise<TierResult> {
494
494
  const cfg = { pymupdfVersion: String(childOptions.pymupdfVersion), warmTimeoutMs: 0 };
495
495
  let cmd: string, args: string[];
496
496
  if (mode === "pdf-text" || (mode === "info" && backend.kind === "none")) { cmd = process.execPath; args = [unpdfWorkerPath(), mode]; }
@@ -536,7 +536,39 @@ export function reconcileRenderMarkers(md: string, renderPages: number[], fmt: s
536
536
  return md;
537
537
  }
538
538
 
539
- export const emptyOcr = (lang: string): OcrInfo => ({ status: "off", lang, textless: [], pages: [], noText: [], ocrFailed: [], budgetStopped: [], reason: null, tesseract: null });
539
+ export const emptyOcr = (lang: string): OcrInfo => ({ status: "off", lang, textless: [], pages: [], noText: [], ocrFailed: [], budgetStopped: [], reason: null, tesseract: null, mode: "textless", sidecars: {}, ocrErrors: {}, killed: null, notAttempted: [], childError: null });
540
+
541
+ export interface OcrPagesOutcome { written: number[]; noText: number[]; ocrFailed: number[]; ocrErrors: Record<number, string>; budgetStopped: number[]; killed: number | null; notAttempted: number[]; childError: string | null; }
542
+ const emptyOutcome = (): OcrPagesOutcome => ({ written: [], noText: [], ocrFailed: [], ocrErrors: {}, budgetStopped: [], killed: null, notAttempted: [], childError: null });
543
+ const ocrPagesTag = (n: number) => `p${String(n).padStart(3, "0")}`;
544
+ const SIDE_MARKER_RE = /^--- end of page\.page_number=\d+ ---$/;
545
+
546
+ function sidecarHasText(dir: string): boolean {
547
+ const file = readdirSync(dir).find((f) => f.endsWith(".md"));
548
+ if (!file) return false;
549
+ return readFileSync(join(dir, file), "utf8").split("\n").slice(1).some((l) => l.trim() && !SIDE_MARKER_RE.test(l));
550
+ }
551
+
552
+ /** Rebuild the outcome from staging markers after the child died or returned garbage. */
553
+ export function recoverOcrPages(stagingDir: string, pages: number[], detail: string): OcrPagesOutcome {
554
+ const out = emptyOutcome();
555
+ const activePath = join(stagingDir, "active");
556
+ const active = existsSync(activePath) ? Number(readFileSync(activePath, "utf8").trim()) || null : null;
557
+ let sawPage = false;
558
+ for (const n of pages) {
559
+ const dir = join(stagingDir, ocrPagesTag(n));
560
+ if (existsSync(join(dir, ".done"))) { sawPage = true; (sidecarHasText(dir) ? out.written : out.noText).push(n); }
561
+ else if (existsSync(join(dir, ".failed"))) { sawPage = true; out.ocrFailed.push(n); out.ocrErrors[n] = readFileSync(join(dir, ".failed"), "utf8").trim() || "unknown error"; }
562
+ else if (n === active) out.killed = n;
563
+ else out.notAttempted.push(n);
564
+ }
565
+ if (active === null && !sawPage) out.childError = detail;
566
+ return out;
567
+ }
568
+
569
+ function outcomeFromChild(j: TierJson): OcrPagesOutcome {
570
+ return { ...emptyOutcome(), written: j.written ?? [], noText: j.noText ?? [], ocrFailed: j.ocrFailed ?? [], ocrErrors: j.ocrErrors ?? {}, budgetStopped: j.budgetStopped ?? [] };
571
+ }
540
572
  const OCR_SENTINEL_RE = /\x00OCR ([^\x00]*)\x00/g;
541
573
 
542
574
  /** Child OCR labels name staged files; the published name exists only after publishStaged. */
@@ -551,7 +583,7 @@ function handleOcr(tier: Tier, type: InputType, o: DocToMdOptions, json: TierJso
551
583
  if (tier === "unpdf") return o.ocr ? { ...emptyOcr(o.ocrLanguage), status: "unavailable", reason: "no Python backend" } : null;
552
584
  const x = json.ocr;
553
585
  if (!x || (type !== "image" && !x.textless.length && !x.pages.length && !x.ocrFailed.length)) return null;
554
- return x;
586
+ return { ...emptyOcr(o.ocrLanguage), ...x };
555
587
  }
556
588
 
557
589
  export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, seams?: Partial<PipelineSeams>): Promise<ConvertOutcome> {
@@ -561,10 +593,17 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
561
593
  if (!st || !st.isFile()) throw new Error(`Not a readable file: ${o.path}`);
562
594
  if (st.size === 0) throw new Error(`empty file: ${o.path}`);
563
595
  const type = classifyInput(inputPath);
596
+ const forced = o.ocrMode === "all";
597
+ if (forced) {
598
+ if (!o.ocr) throw new UsageError("--ocr-mode all requires --ocr");
599
+ if (o.pages === null) throw new UsageError('--ocr-mode all requires an explicit --pages selection (e.g. --pages 2,7); omitted pages and --pages "" mean all pages and are refused to keep OCR cost bounded');
600
+ if (type !== "pdf" && type !== "pptx" && type !== "doc") throw new UsageError("--ocr-mode all applies to PDF, PPTX and DOC inputs only (DOCX pages are page-break segments, not PDF pages; convert the DOCX to PDF first)");
601
+ }
564
602
  if (o.pages && (type === "html" || type === "image" || type === "email")) throw new Error(`--pages does not apply to ${type === "html" ? "HTML files" : type === "image" ? "images" : "email"}`);
565
603
  const isExcel = type === "xlsx" || type === "xlsm" || type === "xls";
566
604
  if (isExcel && o.pages) throw new Error("--pages does not apply to spreadsheets: worksheets have no stable page numbering");
567
605
  const backend = await s.backend({ pymupdfVersion: o.pymupdfVersion, warmTimeoutMs: o.warmTimeoutMs });
606
+ if (forced && backend.kind === "none") throw new Error(`--ocr-mode all cannot run: no Python backend (${backend.reason})`);
568
607
  if (isExcel && (backend.kind === "none" || !backend.xlsx)) throw new Error(`Excel conversion needs a Python backend with openpyxl, xlrd and pillow (${backend.kind === "none" ? backend.reason : `${backend.kind} lacks the Excel packages`}). ${EXCEL_REMEDY}`);
569
608
  if (type === "email" && extname(inputPath).toLowerCase() === ".msg" && (backend.kind === "none" || !backend.email)) throw new Error(`MSG conversion needs the extract-msg package. Python backend: ${backendState(backend, "found without extract-msg")}. Remedy: install uv, or pip install extract-msg markdownify into that Python`);
570
609
  if (type === "email" && extname(inputPath).toLowerCase() !== ".msg" && lacksDocx(backend)) throw new Error(`EML conversion needs the Python DOCX/HTML packages (mammoth, markdownify, python-docx). Python backend: ${backendState(backend, "found without mammoth/markdownify/python-docx")}. ${docxRemedy}`);
@@ -573,7 +612,7 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
573
612
  let office: { pdfPath: string; cleanup: () => void } | null = null;
574
613
  try {
575
614
  let pdfPath = inputPath;
576
- const base = { path: inputPath, pages: o.pages, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
615
+ const base = { path: inputPath, pages: o.pages, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr && !forced, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
577
616
  let tier: Tier | undefined, engine: Engine | undefined, json: TierJson | undefined, degraded: string | null = null, fallbackReason: string | null = null;
578
617
  let explicitBreaks: number | null = null;
579
618
  let notes: string[] = [];
@@ -704,6 +743,18 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
704
743
  if (type === "doc" || officeRoute !== null) degraded = DEGRADED_DOCX_OFFICE;
705
744
  if (officeRoute !== null) fallbackReason = fallbackReason ? `${officeRoute}; ${fallbackReason}` : officeRoute;
706
745
  if (tier === undefined || engine === undefined || json === undefined) throw new Error("internal: no tier produced output");
746
+ const pageStats = json.pageStats ?? null;
747
+ if (pageStats) writePageStats(b, pageStats);
748
+ let ocr = handleOcr(tier, type, o, json);
749
+ if (forced) {
750
+ const r = await s.runTier("ocr-pages", { path: pdfPath, pages: o.pages, stem: b.stem, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs, stagingDir: b.ocrStagingDir, dpi: o.imageDpi, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion }, b, signal, o.primaryTimeoutMs, backend);
751
+ if (signal?.aborted || (!r.ok && "reason" in r && r.reason === "aborted")) throw new Error("aborted");
752
+ if (r.ok && r.json.status === "unavailable") throw new Error(`OCR unavailable: ${r.json.reason} (install Tesseract; see doc/doc-to-md.md)`);
753
+ const outcome = r.ok && r.json.status === "ran" ? outcomeFromChild(r.json)
754
+ : recoverOcrPages(b.ocrStagingDir, o.pages!, !r.ok ? ("userError" in r ? r.userError : `${r.reason}${detailSuffix(r)}`) : "malformed child output");
755
+ const sidecars = publishSidecars(b);
756
+ ocr = { ...emptyOcr(o.ocrLanguage), ...json.ocr, status: "ran", reason: null, mode: "all", pages: outcome.written, noText: outcome.noText, ocrFailed: outcome.ocrFailed, ocrErrors: outcome.ocrErrors, budgetStopped: outcome.budgetStopped, killed: outcome.killed, notAttempted: outcome.notAttempted, childError: outcome.childError, sidecars: Object.fromEntries(sidecars) };
757
+ }
707
758
  if (!isExcel) notes = [...notes, ...(json.notes ?? [])];
708
759
  if (b.renamedFrom) notes.splice(notes[0]?.startsWith("preview truncated:") ? 1 : 0, 0, `renamed to ${b.stem} (${b.renameReason})`);
709
760
  const pageImagesReason = !o.pageImages || b.pageManifest.size || tier === "primary" || tier === "fallback" ? null
@@ -719,7 +770,7 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
719
770
  const markdown = (head.length ? `${head.join("\n")}\n\n` : "") + body;
720
771
  commitBundle(b, markdown);
721
772
  const outline = scanOutline(markdown, o.outlineMaxEntries);
722
- const details: DocToMdDetails = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr: handleOcr(tier, type, o, json) };
773
+ const details: DocToMdDetails = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr, pageStats, pageStatsPath: pageStats ? b.pageStatsPath : null, ocrDir: b.ocrManifest.size ? b.ocrDir : null };
723
774
  return { output: formatHandle(details), details };
724
775
  } catch (e) { abortBundle(b); throw e; }
725
776
  finally { office?.cleanup(); }
@@ -1,4 +1,4 @@
1
- import type { InputType } from "./doc-to-md-options.ts";
1
+ import type { InputType, OcrMode } from "./doc-to-md-options.ts";
2
2
 
3
3
  export type Tier = "primary" | "fallback" | "unpdf" | "excel" | "docx" | "html" | "image" | "email";
4
4
  export type Engine = "pymupdf4llm" | "pymupdf-text" | "unpdf" | "openpyxl" | "xlrd" | "mammoth" | "python-docx" | "markdownify" | "turndown" | "copy" | "extract-msg" | "email";
@@ -18,6 +18,8 @@ export interface SheetInfo {
18
18
  csv: string | null;
19
19
  }
20
20
 
21
+ export type PageStat = { page: number; chars: number; images: number; imageCoverage: number } | { page: number; error: string };
22
+
21
23
  export interface OcrInfo {
22
24
  status: "off" | "unavailable" | "skipped" | "ran";
23
25
  lang: string;
@@ -28,6 +30,12 @@ export interface OcrInfo {
28
30
  budgetStopped: number[];
29
31
  reason: string | null;
30
32
  tesseract: boolean | null;
33
+ mode: OcrMode;
34
+ sidecars: Record<number, string>;
35
+ ocrErrors: Record<number, string>;
36
+ killed: number | null;
37
+ notAttempted: number[];
38
+ childError: string | null;
31
39
  }
32
40
 
33
41
  export interface HandleData {
@@ -35,6 +43,7 @@ export interface HandleData {
35
43
  pageCount: number | null; pages: number[] | null; explicitBreaks: number | null; imageCount: number; pageImageCount: number; pageImagesReason: string | null; bytes: number; lines: number;
36
44
  degraded: string | null; fallbackReason: string | null; failedPages: number[]; emptyPages: number[];
37
45
  notes: string[]; outline: OutlineEntry[]; outlineTotal: number; ocr: OcrInfo | null;
46
+ pageStats: PageStat[] | null; pageStatsPath: string | null; ocrDir: string | null;
38
47
  }
39
48
 
40
49
  export interface InfoData {
@@ -72,7 +81,23 @@ export function compactRanges(nums: number[], maxEntries = 20): string {
72
81
  const INSTALL_HINT = "install Tesseract (see doc/doc-to-md.md), then rerun with ocr=true";
73
82
  const BARE_REASONS = ["fallback tier", "no Python backend"];
74
83
 
84
+ const pageList = (pages: number[]) => `page${pages.length === 1 ? "" : "s"} ${pages.join(", ")}`;
85
+
86
+ function forcedOcrLine(ocr: OcrInfo): string {
87
+ const clauses: string[] = [];
88
+ if (ocr.pages.length) clauses.push(`sidecars for ${pageList(ocr.pages)}`);
89
+ if (ocr.noText.length) clauses.push(`no text on ${pageList(ocr.noText)}`);
90
+ for (const page of ocr.ocrFailed) clauses.push(`failed on page ${page} (${ocr.ocrErrors[page] ?? "unknown error"})`);
91
+ if (ocr.budgetStopped.length) clauses.push(`budget-stopped ${pageList(ocr.budgetStopped)}`);
92
+ if (ocr.killed !== null) clauses.push(`page ${ocr.killed} killed the OCR child (timeout or crash - likely a compression bomb)`);
93
+ if (ocr.childError !== null) clauses.push(`OCR child failed before processing pages: ${ocr.childError}`);
94
+ if (ocr.notAttempted.length) clauses.push(`${pageList(ocr.notAttempted)} not attempted`);
95
+ const rerun = [...ocr.budgetStopped, ...ocr.notAttempted].sort((a, b) => a - b);
96
+ return `OCR: forced (${ocr.lang}) - ${clauses.join("; ")}${rerun.length ? ` - re-run with --pages ${rerun.join(",")}` : ""}`;
97
+ }
98
+
75
99
  export function ocrLine(ocr: OcrInfo, type: InputType): string {
100
+ if (ocr.mode === "all") return forcedOcrLine(ocr);
76
101
  const image = type === "image";
77
102
  const rerun = ocr.tesseract ? "rerun with ocr=true" : INSTALL_HINT;
78
103
  switch (ocr.status) {
@@ -143,6 +168,8 @@ export function formatHandle(h: HandleData): string {
143
168
  if (h.sheetsDir) lines.push(`Sheets-Dir: ${h.sheetsDir}`);
144
169
  if (h.pagesDir && h.pageImageCount > 0) lines.push(`Pages-Dir: ${h.pagesDir} (${h.pageImageCount} pages)`);
145
170
  else if (h.pageImagesReason) lines.push(`Pages-Dir: none - ${h.pageImagesReason}`);
171
+ if (h.pageStatsPath) lines.push(`Page-Stats: ${h.pageStatsPath}`);
172
+ if (h.ocrDir) lines.push(`OCR-Dir: ${h.ocrDir}`);
146
173
  lines.push(`Type: ${h.type} Engine: ${h.engine} Tier: ${h.tier}`);
147
174
  lines.push(`Page-Count: ${pageCountLabel(h)} Pages: ${h.pages ? compactRanges(h.pages) : "all"} Images: ${h.imageCount} Size: ${formatSize(h.bytes)} / ${h.lines} lines`);
148
175
  if (h.degraded) lines.push(`Degraded: ${h.degraded}`);
@@ -6,6 +6,7 @@ import { extname } from "node:path";
6
6
 
7
7
  export type InputType = "pdf" | "docx" | "doc" | "pptx" | "xlsx" | "xlsm" | "xls" | "html" | "image" | "email";
8
8
  export type ImageFormat = "png" | "jpg";
9
+ export type OcrMode = "textless" | "all";
9
10
 
10
11
  export interface Tunables {
11
12
  primaryTimeoutMs: number;
@@ -29,6 +30,7 @@ export interface DocToMdOptions extends Tunables {
29
30
  outputDir: string | null;
30
31
  overwrite: boolean;
31
32
  pageImages: boolean;
33
+ ocrMode: OcrMode;
32
34
  }
33
35
 
34
36
  /** What adapters pass in: intents as raw strings/booleans, tunables optional. */
@@ -39,6 +41,7 @@ export interface PerCallInput extends Partial<Tunables> {
39
41
  outputDir?: string | null;
40
42
  overwrite?: boolean;
41
43
  pageImages?: boolean;
44
+ ocrMode?: OcrMode;
42
45
  }
43
46
 
44
47
  export type DescriptorType = "string" | "int" | "bool" | "pages" | "enum" | "version" | "lang";
@@ -68,6 +71,7 @@ export const DOC_TO_MD_OPTIONS: readonly OptionDescriptor[] = [
68
71
  { key: "outputDir", flag: "--output-dir", type: "string", default: null, settable: false, help: "Bundle root for <stem>.md + images/; default a per-call temp dir. <stem> = basename without extension with [^A-Za-z0-9._-]+ -> _ (empty -> document); a second call on the same stem writes <stem>-2.md" },
69
72
  { key: "overwrite", flag: "--overwrite", type: "bool", default: false, settable: false, help: "Replace an existing completed <stem>.md bundle" },
70
73
  { key: "pageImages", flag: "--page-images", type: "bool", default: false, settable: false, help: "Also render every selected page to pages/<stem>-pNNN.<imageFormat> at imageDpi (PDF, PPTX, .doc, DOCX via LibreOffice); off by default" },
74
+ { key: "ocrMode", flag: "--ocr-mode", type: "enum", default: "textless", settable: false, enumValues: ["textless", "all"], help: "OCR policy: textless (default) OCRs only pages with an empty text layer, inline; all OCRs every selected page and writes the recognized text to ocr/<stem>-pNNN.md sidecars, leaving the Markdown untouched. all requires --ocr and an explicit --pages selection (PDF, PPTX, DOC)." },
71
75
  { key: "primaryTimeoutMs", flag: "--primary-timeout", type: "int", default: 60000, settable: true, env: "PI_DOC_TO_MD_CONVERT_TIMEOUT_MS", help: "pymupdf4llm tier and DOCX child (docx mode); also the unpdf tier" },
72
76
  { key: "fallbackTimeoutMs", flag: "--fallback-timeout", type: "int", default: 30000, settable: true, help: "PyMuPDF get_text tier (including DOCX LibreOffice fallback); also PDF and DOCX info and Excel rendered views" },
73
77
  { key: "sofficeTimeoutMs", flag: "--soffice-timeout", type: "int", default: 120000, settable: true, env: "PI_DOC_TO_MD_SOFFICE_TIMEOUT_MS", help: "DOCX/PPTX -> PDF via LibreOffice; also Excel rendered views" },
@@ -177,12 +181,32 @@ export function resolveOptions(perCall: PerCallInput, settings: Partial<Tunables
177
181
  out[d.key] = value;
178
182
  }
179
183
  const o = out as unknown as DocToMdOptions;
180
- if (o.info && (o.pages !== null || o.outputDir !== null || o.overwrite || o.pageImages)) {
181
- throw new UsageError("--info cannot be combined with --pages, --output-dir, --overwrite or --page-images");
184
+ if (o.info && (o.pages !== null || o.outputDir !== null || o.overwrite || o.pageImages || o.ocrMode === "all")) {
185
+ throw new UsageError("--info cannot be combined with --pages, --output-dir, --overwrite, --page-images or --ocr-mode all");
182
186
  }
183
187
  return o;
184
188
  }
185
189
 
190
+ /** Two-pass OCR recipe, rendered into --help and the generated skill; `<cmd>` is the command prefix. */
191
+ export const USAGE_PATTERNS = [
192
+ "Two-pass OCR (PDF, PPTX, DOC):",
193
+ " 1. <cmd> report.pdf --output-dir out --json",
194
+ ' -> "pageStatsPath" points at out/report.pages.json; pages with few',
195
+ ' chars and high imageCoverage are scans. "savedTo" is the Markdown.',
196
+ " 2. <cmd> report.pdf --output-dir out --ocr --ocr-mode all --pages 2,7 --json",
197
+ ' -> "ocr"."sidecars" maps 2 and 7 to out/ocr/report-2-p002.md and',
198
+ " ...-p007.md (a second run in the same dir gets stem report-2);",
199
+ " the Markdown of this run holds pages 2 and 7 only and equals what",
200
+ " --pages 2,7 without --ocr would produce. Read sidecars and Markdown",
201
+ " by the returned paths, never by guessing names.",
202
+ " 3. --ocr-mode all refuses to run without --ocr and an explicit --pages.",
203
+ ' The "ocr" object (OCR: line) names failed, budget-stopped, killed and',
204
+ " not-attempted pages and the exact --pages to re-run.",
205
+ " Details: doc/doc-to-md.md (bundle contract, failure buckets).",
206
+ ].join("\n");
207
+
208
+ export const usagePatterns = (cmd: string): string => USAGE_PATTERNS.replaceAll("<cmd>", cmd);
209
+
186
210
  export function renderHelp(): string {
187
211
  const row = (d: OptionDescriptor) => ` ${(d.flag ?? "<path>").padEnd(26)} ${d.help}${d.default !== null && d.key !== "info" && d.key !== "overwrite" ? ` (default ${d.default})` : ""}`;
188
212
  return [
@@ -190,7 +214,8 @@ export function renderHelp(): string {
190
214
  "", "Per-call:", ...DOC_TO_MD_OPTIONS.filter((d) => !d.settable).map(row),
191
215
  "", "Tunables (also settable under quiver.docToMd in settings.json; per-call > settings > default):",
192
216
  ...DOC_TO_MD_OPTIONS.filter((d) => d.settable).map(row),
193
- "", "Result: a handle (Saved-To, Images-Dir, Page-Count, Outline ...). Read the Saved-To file for the Markdown.",
217
+ "", "Result: a handle (Saved-To, Images-Dir, Page-Stats, Page-Count, Outline ...). Read the Saved-To file for the Markdown.",
194
218
  "Exit codes: 0 success, 1 runtime error, 2 usage error.",
219
+ "", usagePatterns("pi-quiver doc-to-md"),
195
220
  ].join("\n");
196
221
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-quiver",
3
- "version": "6.8.0",
3
+ "version": "6.9.0",
4
4
  "description": "Personal pack of Pi coding-agent extensions: context-safe fetch, doc_to_md PDF/DOCX/PPTX-to-Markdown conversion, session naming, a themed ASCII startup header, Opus 4.8 fast mode, and a provider-stall watchdog.",
5
5
  "author": "Jacek Juraszek",
6
6
  "license": "MIT",
@@ -253,6 +253,19 @@ def primary_page_markdown(doc, n, d, o, kw, write_images):
253
253
  return rewrite_image_destinations(md, sources)
254
254
 
255
255
 
256
+ def page_stats(doc, n):
257
+ try:
258
+ page = doc[n - 1]
259
+ chars = len(page.get_text("text").strip())
260
+ infos = page.get_image_info()
261
+ area = page.rect.width * page.rect.height
262
+ covered = sum(max(0.0, (b[2] - b[0]) * (b[3] - b[1])) for b in (i["bbox"] for i in infos))
263
+ coverage = round(min(1.0, covered / area), 2) if area > 0 else 0.0
264
+ return {"page": n, "chars": chars, "images": len(infos), "imageCoverage": coverage}
265
+ except Exception as exc: # noqa: BLE001 - stats never cost a page its Markdown
266
+ return {"page": n, "error": f"{type(exc).__name__}: {exc}"[:300]}
267
+
268
+
256
269
  def mode_pdf_primary(o):
257
270
  import time
258
271
  start = time.monotonic()
@@ -260,11 +273,13 @@ def mode_pdf_primary(o):
260
273
  pages = check_pages(o.get("pages"), doc.page_count)
261
274
  staging, out, empty, failed, notes = o["stagingDir"], [], [], [], []
262
275
  page_images = []
276
+ stats = []
263
277
  lang, budget = o.get("ocrLanguage", "eng"), o.get("ocrBudgetMs", 60000)
264
278
  info = new_ocr(lang)
265
279
  status = ocr_status(True, lang) if o.get("ocr") else None
266
280
  ocr_ms, plain_ms = [], []
267
281
  for i, n in enumerate(pages):
282
+ stats.append(page_stats(doc, n))
268
283
  d = page_dir(staging, n)
269
284
  try:
270
285
  page = doc[n - 1]
@@ -325,7 +340,7 @@ def mode_pdf_primary(o):
325
340
  if missing:
326
341
  notes.append(f"Page images: {missing} of {len(pages)} unavailable")
327
342
  return {"markdown": "\n\n".join(out) + "\n", "pages": pages, "pageCount": doc.page_count,
328
- "emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": info, "pageImages": page_images}
343
+ "emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": info, "pageImages": page_images, "pageStats": stats}
329
344
 
330
345
 
331
346
  def mode_pdf_fallback(o):
@@ -335,10 +350,12 @@ def mode_pdf_fallback(o):
335
350
  keep = {int(k): v for k, v in (o.get("keepPages") or {}).items()}
336
351
  staging, out, empty, failed = o["stagingDir"], [], [], []
337
352
  page_images = []
353
+ stats = []
338
354
  lang = o.get("ocrLanguage", "eng")
339
355
  ocr_info = new_ocr(lang)
340
356
  status = {"status": "unavailable", "reason": "fallback tier", "tesseract": None} if o.get("ocr") else None
341
357
  for n in pages:
358
+ stats.append(page_stats(doc, n))
342
359
  links = [f"![](images/{f})" for f in keep.get(n, [])]
343
360
  text = ""
344
361
  page_pic = None
@@ -389,7 +406,64 @@ def mode_pdf_fallback(o):
389
406
  if missing:
390
407
  notes.append(f"Page images: {missing} of {len(pages)} unavailable")
391
408
  return {"markdown": "\n\n".join(out) + "\n", "pages": pages, "pageCount": doc.page_count,
392
- "emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": ocr_info, "pageImages": page_images}
409
+ "emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": ocr_info, "pageImages": page_images, "pageStats": stats}
410
+
411
+
412
+ def mode_ocr_pages(o):
413
+ import time
414
+ start = time.monotonic()
415
+ lang, budget = o.get("ocrLanguage", "eng"), o.get("ocrBudgetMs", 60000)
416
+ status = ocr_status(True, lang)
417
+ if status["status"] != "ready":
418
+ return {"status": "unavailable", "reason": status["reason"]}
419
+ doc = open_pdf(o["path"])
420
+ pages = check_pages(o.get("pages"), doc.page_count)
421
+ staging, stem, dpi = o["stagingDir"], o["stem"], o.get("dpi", 150)
422
+ os.makedirs(staging, exist_ok=True)
423
+ active = os.path.join(staging, "active")
424
+ out = {"status": "ran", "written": [], "noText": [], "ocrFailed": [], "ocrErrors": {}, "budgetStopped": []}
425
+ ocr_ms = []
426
+ for i, n in enumerate(pages):
427
+ with open(active, "w") as fh:
428
+ fh.write(str(n))
429
+ if os.environ.get("DOC_TO_MD_OCR_STALL_PAGE") == str(n): # tests only: simulate a wedged page
430
+ time.sleep(3600)
431
+ elapsed = (time.monotonic() - start) * 1000
432
+ est = max(ocr_ms) if ocr_ms else OCR_EST_INITIAL_MS
433
+ if not ocr_admit(elapsed, est, 0, 0, budget):
434
+ out["budgetStopped"].extend(pages[i:])
435
+ os.remove(active)
436
+ break
437
+ tag = f"p{n:03d}"
438
+ d = os.path.join(staging, tag)
439
+ os.makedirs(d, exist_ok=True)
440
+ sidecar = os.path.join(d, f"{stem}-{tag}.md")
441
+ t0 = time.monotonic()
442
+ try:
443
+ page = doc[n - 1]
444
+ eff = clamped_dpi(page.rect.width, page.rect.height, dpi)
445
+ if eff is None:
446
+ raise RuntimeError(f"page cannot be rendered at a usable DPI ({page.rect.width:.0f} x {page.rect.height:.0f} pt)")
447
+ tp = page.get_textpage_ocr(full=True, language=lang, dpi=eff)
448
+ text = page.get_text("text", textpage=tp).strip()
449
+ header = f"<!-- OCR of page {n} (tesseract {lang}); recognized text, not the text layer -->"
450
+ with open(sidecar, "w", encoding="utf-8") as fh:
451
+ fh.write(header + "\n\n" + (text + "\n\n" if text else "") + SEP.format(n=n).strip("\n") + "\n")
452
+ mark_done(d)
453
+ (out["written"] if text else out["noText"]).append(n)
454
+ except Exception as exc: # noqa: BLE001 - one page never stops the pass
455
+ msg = f"{type(exc).__name__}: {exc}"
456
+ with open(os.path.join(d, ".failed"), "w", encoding="utf-8") as fh:
457
+ fh.write(msg)
458
+ try:
459
+ os.remove(sidecar)
460
+ except OSError:
461
+ pass
462
+ out["ocrFailed"].append(n)
463
+ out["ocrErrors"][str(n)] = msg
464
+ ocr_ms.append((time.monotonic() - t0) * 1000)
465
+ os.remove(active)
466
+ return out
393
467
 
394
468
 
395
469
  def esc(v):
@@ -1249,12 +1323,12 @@ def mode_render_pages(o):
1249
1323
  return {"ok": True, "rendered": rendered, "failed": failed}
1250
1324
 
1251
1325
 
1252
- MODES = {"info": None, "pdf-primary": mode_pdf_primary, "pdf-fallback": mode_pdf_fallback, "xlsx": mode_xlsx, "render-pages": mode_render_pages, "docx": mode_docx, "html": mode_html, "image": mode_image, "email": mode_email}
1326
+ MODES = {"info": None, "pdf-primary": mode_pdf_primary, "pdf-fallback": mode_pdf_fallback, "xlsx": mode_xlsx, "render-pages": mode_render_pages, "docx": mode_docx, "html": mode_html, "image": mode_image, "email": mode_email, "ocr-pages": mode_ocr_pages}
1253
1327
 
1254
1328
 
1255
1329
  def main():
1256
1330
  if len(sys.argv) != 2 or sys.argv[1] not in MODES:
1257
- print("usage: doc_to_md.py <info|pdf-primary|pdf-fallback|xlsx|render-pages|docx|html|image|email> (options JSON on stdin)", file=sys.stderr)
1331
+ print("usage: doc_to_md.py <info|pdf-primary|pdf-fallback|xlsx|render-pages|docx|html|image|email|ocr-pages> (options JSON on stdin)", file=sys.stderr)
1258
1332
  return 1
1259
1333
  mode = sys.argv[1]
1260
1334
  o = json.loads(sys.stdin.read() or "{}")