pi-quiver 6.10.0 → 6.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -8,6 +8,14 @@ Published to npm as `pi-quiver` (`pi install npm:pi-quiver`). Pushing a
8
8
  via OIDC trusted publishing. The release helper at
9
9
  `.agents/skills/release/scripts/release.sh` cuts the tag; CI publishes.
10
10
 
11
+ ## v6.11.0 - 2026-10-02
12
+
13
+ - `doc_to_md`: textless PDF pages holding one eligible full-page image deliver the embedded stream without re-encoding; the handle prints `Native-Images:` and tool details / CLI JSON carry `nativeImages` (#27).
14
+ - `doc_to_md`: native probing, PDF page renders, and textless-page inline OCR run in a spawned raster worker with per-job budgets; a toxic page becomes a `Failed pages` note instead of failing the conversion when other pages succeed (#27).
15
+ - `doc_to_md`: the render ceiling rises from 16 to 50 Mpx; clamped PDF renders add Notes and report effective `pageImages[].dpi` plus `requestedDpi` (#27).
16
+ - `doc_to_md`: `hideAnnotations` / `--hide-annotations` hides PDF annotations and form widgets on renders and lets annotated scans use native delivery; default runs keep their painted render (#27).
17
+ - `doc_to_md`: `.done` markers retain native-image and clamp metadata for completed pages across primary-tier failure and fallback; corrupt metadata does not discard completed images (#27).
18
+
11
19
  ## v6.10.0 - 2026-10-02
12
20
 
13
21
  - `doc_to_md`: new per-call `words` option (`--words`; not settable) writes `<stem>.words.json` - per selected PDF/image page, every text-layer word and every word inline OCR recognized in the same run, with display-space bbox (points; source pixels for images) and `source` `text`/`ocr`. Under `--ocr-mode all` each sidecar gets `ocr/<stem>-pNNN.words.json`. Never triggers OCR; Markdown, page stats and OCR outcome are unchanged. Handle gains `Words:`, `--json` gains `wordsPath`, `wordsReason`, `wordsErrors`, `ocr.wordSidecars`; `--info --words` is a usage error (#28).
package/README.md CHANGED
@@ -65,7 +65,7 @@ A 300 KB changelog page never touches your context window - you get a preview an
65
65
  | Extension | Tool | What it does |
66
66
  | --- | --- | --- |
67
67
  | `extensions/fetch.ts` | `fetch` | Retrieve URLs over HTTP(S). HTML -> Markdown (Readability extraction, Turndown conversion). Binary saved untouched to a temp file. GitHub issue/PR/repo/actions-run/actions-job URLs auto-route through `gh` (falls back to HTTP); failed runs/jobs include failed-step logs (best-effort, summary-only otherwise). Same size gate as `fetch`. Behavior lives in `lib/fetch-core.ts`; also exposed as the `pi-quiver fetch` CLI (see [Claude Code support](#claude-code-support)). |
68
- | `extensions/doc_to_md.ts` | `doc_to_md` | Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, MSG/EML, HTML, or image file to a Markdown bundle on disk (`<stem>.md` + `images/`, spreadsheet `sheets/`, optional `pageImages` in `pages/`, and email `attachments/`) and return a handle (paths, page count, outline, diagnostics) - never inline Markdown. `info` inspects PDF/Office inputs; `pages` selects 1-based PDF/Office pages (DOCX: explicit-page-break segments); HTML, image, and email inputs reject both; every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Tiers: pymupdf4llm -> PyMuPDF text (degraded) -> unpdf worker (no Python only). Excel -> sheet inventory, full CSVs, bounded previews, and optional rendered views. HTML keeps local and `data:` images in the bundle and remote images as links; image inputs keep the original image. Scanned PDF pages keep a page picture on both Python tiers; opt-in OCR needs Tesseract language data. DOCX converts directly (mammoth, python-docx fallback); LibreOffice pagination marks the degraded route. The Outline lists `L<line>` and `p<page>` per heading. Settings under `quiver.docToMd`. Every conversion that goes through the Python PDF tiers (PDF, PPTX, `.doc`, DOCX via LibreOffice) writes `<stem>.pages.json` (per-page chars, image count, image coverage; handle line `Page-Stats:`); unpdf writes no stats. `ocrMode: "all"` with `ocr: true` and an explicit `pages` forces OCR on those pages into `ocr/<stem>-pNNN.md` sidecars, leaving the Markdown byte-identical to the same selected-page call without OCR. `words: true` (CLI `--words`, per-call only - no settings key) writes `<stem>.words.json` with every word's bbox per selected PDF/image page (`source` `text` or `ocr`), plus `ocr/<stem>-pNNN.words.json` beside each forced-OCR sidecar; it never triggers OCR. |
68
+ | `extensions/doc_to_md.ts` | `doc_to_md` | Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, MSG/EML, HTML, or image file to a Markdown bundle on disk (`<stem>.md` + `images/`, spreadsheet `sheets/`, optional `pageImages` in `pages/`, and email `attachments/`) and return a handle (paths, page count, outline, diagnostics) - never inline Markdown. `info` inspects PDF/Office inputs; `pages` selects 1-based PDF/Office pages (DOCX: explicit-page-break segments); HTML, image, and email inputs reject both; every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Tiers: pymupdf4llm -> PyMuPDF text (degraded) -> unpdf worker (no Python only). Excel -> sheet inventory, full CSVs, bounded previews, and optional rendered views. HTML keeps local and `data:` images in the bundle and remote images as links; image inputs keep the original image. Scanned PDF pages keep a page picture on both Python tiers; eligible single full-page images are delivered as their embedded stream, without re-encoding or render DPI; opt-in OCR needs Tesseract language data. DOCX converts directly (mammoth, python-docx fallback); LibreOffice pagination marks the degraded route. The Outline lists `L<line>` and `p<page>` per heading. Settings under `quiver.docToMd`. Every conversion that goes through the Python PDF tiers (PDF, PPTX, `.doc`, DOCX via LibreOffice) writes `<stem>.pages.json` (per-page chars, image count, image coverage; handle line `Page-Stats:`); unpdf writes no stats. `ocrMode: "all"` with `ocr: true` and an explicit `pages` forces OCR on those pages into `ocr/<stem>-pNNN.md` sidecars, leaving the Markdown byte-identical to the same selected-page call without OCR. `words: true` (CLI `--words`, per-call only - no settings key) writes `<stem>.words.json` with every word's bbox per selected PDF/image page (`source` `text` or `ocr`), plus `ocr/<stem>-pNNN.words.json` beside each forced-OCR sidecar; it never triggers OCR. |
69
69
  | `extensions/session-name.ts` | `/session-name` | Manual + opt-in automatic session naming, naming rules and deny list, long-session revisits, and Ghostty/Herdr tab rename. OFF by default. |
70
70
  | `extensions/sword-header.ts` | `/builtin-header` | Themed ASCII startup header replacing pi's default logo. OFF by default. |
71
71
  | `extensions/fast-mode.ts` | `/fast` | Inject Anthropic fast-mode (`speed: "fast"` + `anthropic-beta: fast-mode-2026-02-01`) into every Claude Opus 4.8 / Opus 5 request, any thinking level. `--fast` flag + `/fast [on\|off\|status]`. OFF by default. |
@@ -303,8 +303,9 @@ Each setting can also be overridden per-process via `PI_QUIVER_SLACK_ENABLED`, `
303
303
  | `excelTimeoutMs` | `60000` | Excel child and Excel info deadline. |
304
304
  | `warmTimeoutMs` | `120000` | Absolute first-call backend discovery/bootstrap deadline. |
305
305
  | `pymupdfVersion` | `1.27.2.3` | pymupdf4llm pin, minimum `1.27.0`. |
306
- | `imageDpi` | `150` | Render DPI for page images and Excel rendered views (capped by a 16 Mpx budget). |
306
+ | `imageDpi` | `150` | Render DPI for page images and Excel rendered views (capped by a 50 Mpx budget). |
307
307
  | `imageFormat` | `png` | Rendered image format: `png` or `jpg`; embedded images retain their extension. |
308
+ | `hideAnnotations` | `false` | Hide annotations on PDF textless-page and `pages/` renders, including form-field widgets (filled values disappear). Default paints them. Does not affect OCR text or embedded images; also lets an annotated scan be delivered as its embedded image. |
308
309
  | `maxOutputBytes` | `20000000` | Child stdout cap in bytes. |
309
310
  | `outlineMaxEntries` | `40` | Heading outline, TOC, or sheet inventory entries in the handle. |
310
311
  | `ocr` | `false` | Opt-in OCR for scanned pages and images when Tesseract language data is installed. |
@@ -314,7 +315,7 @@ Each setting can also be overridden per-process via `PI_QUIVER_SLACK_ENABLED`, `
314
315
 
315
316
  A bundle is `<outputDir>/<stem>.md` plus `<outputDir>/<stem>.pages.json` on Python PDF tiers, optional OCR sidecars under `<outputDir>/ocr/`, and `<outputDir>/images/` and, for spreadsheets with data, `<outputDir>/sheets/`; without `outputDir`, the tool creates a per-call temp root. A conversion owns `<stem>.md.lock` until it atomically publishes the Markdown. An existing `<stem>.md` fails the call unless `overwrite` is set; `overwrite` replaces that Markdown and the files it owns (`images/<stem>-p<N>-<n>.*`, `images/<stem>-s<idx>[-<n>].*`, `sheets/<stem>-s<idx>-<slug>.csv`, `<stem>.pages.json`, `ocr/<stem>-pNNN.md`), nothing else. Temp bundles are caller-owned - the tool never deletes a bundle it produced.
316
317
 
317
- Excel needs a Python backend with openpyxl, xlrd and pillow - otherwise the call fails with `Remedy: install uv, or pip install openpyxl xlrd pillow`. The Markdown opens with a `## Sheets` table listing every sheet in workbook order (0-based `#`, `worksheet`/`chartsheet`, size, hidden, chart and image counts, rendered view, CSV link for non-empty worksheets), then one section per sheet: a `Data:` line linking the full-content CSV under `sheets/` for non-empty worksheets, chart metadata from the workbook model, embedded images, an optional rendered view, a preview of at most the first 100 rows x 50 columns, and - only when the preview is truncated - a `Columns:` profile (type, non-empty count, min/max, distinct up to 50). Sizes are the extent of non-empty cells (the `info` handle reports the raw worksheet dimensions instead, which may be larger). Rendered views (`images/<stem>-s<idx>.<fmt>`) are produced for sheets carrying charts or images when LibreOffice is on `PATH`: the workbook is exported one PDF page per sheet and rasterized under a 16 Mpx budget. Any LibreOffice or rasterization failure degrades to `Rendered view: unavailable (<reason>)` and a handle note; it never fails the conversion. Workbooks whose chartsheet drawings carry a zero-size anchor (openpyxl-authored files; Excel-authored files are unaffected) render as a degenerate page and are reported as such. `.xls` gets the inventory, CSVs and previews but no visual detection or rendering.
318
+ Excel needs a Python backend with openpyxl, xlrd and pillow - otherwise the call fails with `Remedy: install uv, or pip install openpyxl xlrd pillow`. The Markdown opens with a `## Sheets` table listing every sheet in workbook order (0-based `#`, `worksheet`/`chartsheet`, size, hidden, chart and image counts, rendered view, CSV link for non-empty worksheets), then one section per sheet: a `Data:` line linking the full-content CSV under `sheets/` for non-empty worksheets, chart metadata from the workbook model, embedded images, an optional rendered view, a preview of at most the first 100 rows x 50 columns, and - only when the preview is truncated - a `Columns:` profile (type, non-empty count, min/max, distinct up to 50). Sizes are the extent of non-empty cells (the `info` handle reports the raw worksheet dimensions instead, which may be larger). Rendered views (`images/<stem>-s<idx>.<fmt>`) are produced for sheets carrying charts or images when LibreOffice is on `PATH`: the workbook is exported one PDF page per sheet and rasterized under a 50 Mpx budget. Any LibreOffice or rasterization failure degrades to `Rendered view: unavailable (<reason>)` and a handle note; it never fails the conversion. Workbooks whose chartsheet drawings carry a zero-size anchor (openpyxl-authored files; Excel-authored files are unaffected) render as a degenerate page and are reported as such. `.xls` gets the inventory, CSVs and previews but no visual detection or rendering.
318
319
 
319
320
  Worst-case wall time: PDF `warmTimeoutMs (first call) + primaryTimeoutMs + fallbackTimeoutMs`; PPTX adds `sofficeTimeoutMs`; DOCX on the Python path `warmTimeoutMs + primaryTimeoutMs` (success or a terminal child failure), DOCX child exit 1 then LibreOffice `warmTimeoutMs + primaryTimeoutMs + sofficeTimeoutMs + primaryTimeoutMs + fallbackTimeoutMs`, DOCX without a DOCX-capable backend `warmTimeoutMs + sofficeTimeoutMs + primaryTimeoutMs + fallbackTimeoutMs`; Excel `warmTimeoutMs + excelTimeoutMs + sofficeTimeoutMs + fallbackTimeoutMs`. Add `KILL_GRACE_MS` (2000 ms) per kill. There is no cap on image count, image bytes, cell count or workbook memory - deliberately; the per-tier timeouts, the rendered-view pixel budget and `maxOutputBytes` are the bounds.
320
321
 
@@ -680,6 +680,13 @@ function publishStaged(b) {
680
680
  continue;
681
681
  }
682
682
  const page = Number(m[1]);
683
+ const raw = readFileSync(join2(pageDir, ".done"), "utf8").trim();
684
+ let meta = {};
685
+ try {
686
+ const v = raw ? JSON.parse(raw) : {};
687
+ if (v && typeof v === "object" && !Array.isArray(v)) meta = v;
688
+ } catch {
689
+ }
683
690
  const files = readdirSync(pageDir).filter((f) => f !== ".done" && statSync(join2(pageDir, f)).isFile()).sort();
684
691
  const names = [];
685
692
  files.forEach((f, i) => {
@@ -689,8 +696,9 @@ function publishStaged(b) {
689
696
  b.sourceMap.set(`${dir}/${f}`, `images/${name}`);
690
697
  names.push(name);
691
698
  });
699
+ if (meta.native) meta.native = { ...meta.native, file: names[files.indexOf(meta.native.file)] ?? meta.native.file };
692
700
  rmSync(pageDir, { recursive: true, force: true });
693
- out.set(page, names);
701
+ out.set(page, { files: names, meta });
694
702
  }
695
703
  return out;
696
704
  }
@@ -969,6 +977,7 @@ function formatHandle(h) {
969
977
  if (h.pagesDir && h.pageImageCount > 0) lines.push(`Pages-Dir: ${h.pagesDir} (${h.pageImageCount} pages)`);
970
978
  else if (h.pageImagesReason) lines.push(`Pages-Dir: none - ${h.pageImagesReason}`);
971
979
  if (h.pageStatsPath) lines.push(`Page-Stats: ${h.pageStatsPath}`);
980
+ if (h.nativeImages.length) lines.push(`Native-Images: ${h.nativeImages.length === 1 ? "page" : "pages"} ${compactRanges(h.nativeImages.map((n) => n.page))} (embedded image streams, no render DPI)`);
972
981
  if (h.wordsPath) {
973
982
  const bad = Object.keys(h.wordsErrors).map(Number).sort((a, b) => a - b);
974
983
  lines.push(`Words: ${h.wordsPath}${bad.length ? ` (extraction failed for pages ${compactRanges(bad)}: ${h.wordsErrors[bad[0]]})` : ""}`);
@@ -1036,7 +1045,8 @@ var DOC_TO_MD_OPTIONS = [
1036
1045
  { key: "maxOutputBytes", flag: "--max-output-bytes", type: "int", default: 2e7, settable: true, help: "Child stdout cap in bytes" },
1037
1046
  { key: "outlineMaxEntries", flag: "--outline-max-entries", type: "int", default: 40, settable: true, help: "Heading outline / TOC / sheet inventory cap in the handle" },
1038
1047
  { key: "ocr", flag: "--ocr", type: "bool", default: false, settable: true, help: "Run OCR on pages without a text layer and on image inputs when Tesseract language data is installed; off by default (--no-ocr turns a settings-level true off)" },
1039
- { key: "ocrLanguage", flag: "--ocr-language", type: "lang", default: "eng", settable: true, help: "Tesseract language code(s), +-joined, e.g. deu+eng" }
1048
+ { key: "ocrLanguage", flag: "--ocr-language", type: "lang", default: "eng", settable: true, help: "Tesseract language code(s), +-joined, e.g. deu+eng" },
1049
+ { key: "hideAnnotations", flag: "--hide-annotations", type: "bool", default: false, settable: true, help: "Render PDF pages without annotations (sticky notes, highlights, stamps - and form-field widgets, so filled form values disappear); default paints them, as PyMuPDF does. Applies to pages/ renders and textless-page renders, not to OCR text or embedded images; also lets an annotated scan be delivered as its embedded image." }
1040
1050
  ];
1041
1051
  var TUNABLE_DEFAULTS = Object.fromEntries(
1042
1052
  DOC_TO_MD_OPTIONS.filter((d) => d.settable).map((d) => [d.key, d.default])
@@ -1706,6 +1716,22 @@ function handleOcr(tier, type, o, json) {
1706
1716
  if (!x || type !== "image" && !x.textless.length && !x.pages.length && !x.ocrFailed.length) return null;
1707
1717
  return { ...emptyOcr(o.ocrLanguage), ...x };
1708
1718
  }
1719
+ var nativeFromChild = (b, json, notes) => (json.nativeImages ?? []).flatMap((e) => {
1720
+ const file = b.sourceMap.get(`p${e.page}/${e.file}`);
1721
+ if (!file) {
1722
+ notes.push(`Native image p${e.page}/${e.file} not published`);
1723
+ return [];
1724
+ }
1725
+ return [{ ...e, file: join3(b.root, file) }];
1726
+ });
1727
+ function retainedNative(b, kept, notes) {
1728
+ const out = [];
1729
+ for (const [page, k] of kept) {
1730
+ if (k.meta.native) out.push({ page, ...k.meta.native, file: join3(b.imagesDir, k.meta.native.file) });
1731
+ if (k.meta.dpi !== void 0 && k.meta.requestedDpi !== void 0 && k.meta.dpi < k.meta.requestedDpi) notes.push(`Page ${page} rendered at ${k.meta.dpi} dpi (requested ${k.meta.requestedDpi}; 50 Mpx ceiling)`);
1732
+ }
1733
+ return out;
1734
+ }
1709
1735
  async function convertDocument(o, signal, seams) {
1710
1736
  const s = { backend: (c) => getBackend(c, void 0, signal), runTier: runTierReal, office: tryConvertOffice, ...seams };
1711
1737
  const inputPath = resolve2(o.path);
@@ -1732,10 +1758,11 @@ async function convertDocument(o, signal, seams) {
1732
1758
  let office = null;
1733
1759
  try {
1734
1760
  let pdfPath = inputPath;
1735
- const base = { path: inputPath, pages: o.pages, ...o.words && (type === "pdf" || type === "image") ? { words: true } : {}, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr && !forced, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
1761
+ const base = { path: inputPath, pages: o.pages, ...o.words && (type === "pdf" || type === "image") ? { words: true } : {}, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr && !forced, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs, hideAnnotations: o.hideAnnotations };
1736
1762
  let tier, engine, json, degraded = null, fallbackReason = null;
1737
1763
  let explicitBreaks = null;
1738
1764
  let notes = [];
1765
+ let nativeImages = [];
1739
1766
  let officeRoute = null;
1740
1767
  let copyReason = null;
1741
1768
  if (type === "html") {
@@ -1890,11 +1917,12 @@ async function convertDocument(o, signal, seams) {
1890
1917
  tier = "primary";
1891
1918
  engine = "pymupdf4llm";
1892
1919
  json = p.json;
1920
+ nativeImages = nativeFromChild(b, json, notes);
1893
1921
  if (json.pageImages?.length) publishPageImages(b, json.pageCount ?? 0);
1894
1922
  } else if ("userError" in p) throw new Error(p.userError);
1895
1923
  else {
1896
1924
  if (signal?.aborted) throw new Error("aborted");
1897
- const keepPages = Object.fromEntries([...kept.entries()].map(([k, v]) => [String(k), v]));
1925
+ const keepPages = Object.fromEntries([...kept.entries()].map(([k, v]) => [String(k), v.files]));
1898
1926
  rmSync2(b.pagesStagingDir, { recursive: true, force: true });
1899
1927
  const f = await s.runTier("pdf-fallback", { ...pdfBase, keepPages }, b, signal, o.fallbackTimeoutMs, backend);
1900
1928
  publishStaged(b);
@@ -1905,6 +1933,7 @@ async function convertDocument(o, signal, seams) {
1905
1933
  json = f.json;
1906
1934
  degraded = DEGRADED_TEXT;
1907
1935
  fallbackReason = `primary ${p.reason}`;
1936
+ nativeImages = [...retainedNative(b, kept, notes), ...nativeFromChild(b, json, notes)].sort((x, y) => x.page - y.page);
1908
1937
  }
1909
1938
  }
1910
1939
  }
@@ -1957,7 +1986,7 @@ async function convertDocument(o, signal, seams) {
1957
1986
  ` : "") + body;
1958
1987
  commitBundle(b, markdown);
1959
1988
  const outline = scanOutline(markdown, o.outlineMaxEntries);
1960
- const details = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr, pageStats, pageStatsPath: pageStats ? b.pageStatsPath : null, ocrDir: b.ocrManifest.size ? b.ocrDir : null, wordsPath, wordsReason, wordsErrors };
1989
+ const details = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr, pageStats, pageStatsPath: pageStats ? b.pageStatsPath : null, ocrDir: b.ocrManifest.size ? b.ocrDir : null, nativeImages, wordsPath, wordsReason, wordsErrors };
1961
1990
  return { output: formatHandle(details), details };
1962
1991
  } catch (e) {
1963
1992
  abortBundle(b);
@@ -109,9 +109,12 @@ export function openBundle(root: string, requested: string, overwrite: boolean):
109
109
  } catch (e) { rmSync(lockPath, { force: true }); throw e; }
110
110
  }
111
111
 
112
- /** Publish every `p<N>/` staging dir carrying `.done`; discard partial ones. Returns page -> published filenames. */
113
- export function publishStaged(b: Bundle): Map<number, string[]> {
114
- const out = new Map<number, string[]>();
112
+ export interface DoneMeta { native?: { file: string; width: number; height: number }; dpi?: number; requestedDpi?: number }
113
+ export interface StagedPage { files: string[]; meta: DoneMeta }
114
+
115
+ /** Publish completed pages, retaining metadata with native filenames renamed; discard partial pages. */
116
+ export function publishStaged(b: Bundle): Map<number, StagedPage> {
117
+ const out = new Map<number, StagedPage>();
115
118
  if (!existsSync(b.stagingDir)) return out;
116
119
  for (const dir of readdirSync(b.stagingDir).sort()) {
117
120
  const m = dir.match(/^p(\d+)$/);
@@ -119,6 +122,12 @@ export function publishStaged(b: Bundle): Map<number, string[]> {
119
122
  const pageDir = join(b.stagingDir, dir);
120
123
  if (!existsSync(join(pageDir, ".done"))) { rmSync(pageDir, { recursive: true, force: true }); continue; }
121
124
  const page = Number(m[1]);
125
+ const raw = readFileSync(join(pageDir, ".done"), "utf8").trim();
126
+ let meta: DoneMeta = {};
127
+ try {
128
+ const v = raw ? JSON.parse(raw) : {};
129
+ if (v && typeof v === "object" && !Array.isArray(v)) meta = v;
130
+ } catch { /* Completed images survive truncated metadata. */ }
122
131
  const files = readdirSync(pageDir).filter((f) => f !== ".done" && statSync(join(pageDir, f)).isFile()).sort();
123
132
  const names: string[] = [];
124
133
  files.forEach((f, i) => {
@@ -128,8 +137,9 @@ export function publishStaged(b: Bundle): Map<number, string[]> {
128
137
  b.sourceMap.set(`${dir}/${f}`, `images/${name}`);
129
138
  names.push(name);
130
139
  });
140
+ if (meta.native) meta.native = { ...meta.native, file: names[files.indexOf(meta.native.file)] ?? meta.native.file };
131
141
  rmSync(pageDir, { recursive: true, force: true });
132
- out.set(page, names);
142
+ out.set(page, { files: names, meta });
133
143
  }
134
144
  return out;
135
145
  }
@@ -12,8 +12,8 @@ import { type ChildProcess, spawn } from "node:child_process";
12
12
  import { homedir, tmpdir } from "node:os";
13
13
  import { basename, dirname, extname, isAbsolute, join, resolve } from "node:path";
14
14
  import { fileURLToPath, pathToFileURL } from "node:url";
15
- import { type Bundle, abortBundle, commitBundle, openBundle, publishAttachments, publishSidecars, publishWords, writePageStats, publishPageImages, publishSheetCsvs, publishSheetImages, publishStaged, rewriteLinks, tempBundleRoot, validateImageLinks } from "./doc-to-md-bundle.ts";
16
- import { type Engine, type HandleData, type InfoData, type OcrInfo, type PageStat, type Tier, formatHandle, formatInfoHandle, scanOutline, type SheetInfo, type TocEntry } from "./doc-to-md-handle.ts";
15
+ import { type StagedPage, type Bundle, abortBundle, commitBundle, openBundle, publishAttachments, publishSidecars, publishWords, writePageStats, publishPageImages, publishSheetCsvs, publishSheetImages, publishStaged, rewriteLinks, tempBundleRoot, validateImageLinks } from "./doc-to-md-bundle.ts";
16
+ import { type NativeImage, type Engine, type HandleData, type InfoData, type OcrInfo, type PageStat, type Tier, formatHandle, formatInfoHandle, scanOutline, type SheetInfo, type TocEntry } from "./doc-to-md-handle.ts";
17
17
  import { type DocToMdOptions, type InputType, IMAGE_EXTS, TUNABLE_DEFAULTS, UsageError, classifyInput, sanitizeStem } from "./doc-to-md-options.ts";
18
18
 
19
19
  export * from "./doc-to-md-options.ts";
@@ -467,7 +467,7 @@ const clearStaging = (b: Pick<Bundle, "stagingDir">) => { for (const f of readdi
467
467
  export const EXCEL_REMEDY = "Remedy: install uv, or pip install openpyxl xlrd pillow";
468
468
 
469
469
  export type Mode = "html" | "image" | "info" | "pdf-primary" | "pdf-fallback" | "xlsx" | "pdf-text" | "render-pages" | "docx" | "email" | "ocr-pages";
470
- export interface TierJson { words?: boolean; wordsErrors?: Record<string, string>; pageStats?: PageStat[]; status?: string; written?: number[]; noText?: number[]; ocrFailed?: number[]; ocrErrors?: Record<string, string>; budgetStopped?: number[]; pageImages?: { page: number; file: string }[]; ocr?: OcrInfo; markdown?: string; pages?: number[]; pageCount?: number; emptyPages?: number[]; failedPages?: { page: number; error: string }[]; notes?: string[]; images?: { sheetIndex: number; file: string }[]; metadata?: Record<string, string>; toc?: [number, string, number | null][]; explicitBreaks?: number; engine?: string; degraded?: boolean; fallbackReason?: string | null; sheets?: SheetInfo[]; renderPages?: number[]; sheetCount?: number; ok?: boolean; reason?: string; rendered?: { idx: number; file: string; dpi: number }[]; failed?: { idx: number; reason: string }[]; }
470
+ export interface TierJson { words?: boolean; wordsErrors?: Record<string, string>; pageStats?: PageStat[]; status?: string; written?: number[]; noText?: number[]; ocrFailed?: number[]; ocrErrors?: Record<string, string>; budgetStopped?: number[]; pageImages?: { page: number; file: string; dpi?: number; requestedDpi?: number }[]; nativeImages?: NativeImage[]; ocr?: OcrInfo; markdown?: string; pages?: number[]; pageCount?: number; emptyPages?: number[]; failedPages?: { page: number; error: string }[]; notes?: string[]; images?: { sheetIndex: number; file: string }[]; metadata?: Record<string, string>; toc?: [number, string, number | null][]; explicitBreaks?: number; engine?: string; degraded?: boolean; fallbackReason?: string | null; sheets?: SheetInfo[]; renderPages?: number[]; sheetCount?: number; ok?: boolean; reason?: string; rendered?: { idx: number; file: string; dpi: number }[]; failed?: { idx: number; reason: string }[]; }
471
471
  export type TierResult = { ok: true; json: TierJson } | { ok: false; reason: string; detail?: string } | { ok: false; userError: string; pageCount?: number };
472
472
 
473
473
  export interface PipelineSeams {
@@ -586,6 +586,25 @@ function handleOcr(tier: Tier, type: InputType, o: DocToMdOptions, json: TierJso
586
586
  return { ...emptyOcr(o.ocrLanguage), ...x };
587
587
  }
588
588
 
589
+ const nativeFromChild = (b: Bundle, json: TierJson, notes: string[]): NativeImage[] => (json.nativeImages ?? []).flatMap((e) => {
590
+ const file = b.sourceMap.get(`p${e.page}/${e.file}`);
591
+ if (!file) {
592
+ notes.push(`Native image p${e.page}/${e.file} not published`);
593
+ return [];
594
+ }
595
+ return [{ ...e, file: join(b.root, file) }];
596
+ });
597
+
598
+ // A dead primary has no JSON response; retained pages carry their image and clamp facts in .done.
599
+ function retainedNative(b: Bundle, kept: Map<number, StagedPage>, notes: string[]): NativeImage[] {
600
+ const out: NativeImage[] = [];
601
+ for (const [page, k] of kept) {
602
+ if (k.meta.native) out.push({ page, ...k.meta.native, file: join(b.imagesDir, k.meta.native.file) });
603
+ if (k.meta.dpi !== undefined && k.meta.requestedDpi !== undefined && k.meta.dpi < k.meta.requestedDpi) notes.push(`Page ${page} rendered at ${k.meta.dpi} dpi (requested ${k.meta.requestedDpi}; 50 Mpx ceiling)`);
604
+ }
605
+ return out;
606
+ }
607
+
589
608
  export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, seams?: Partial<PipelineSeams>): Promise<ConvertOutcome> {
590
609
  const s: PipelineSeams = { backend: (c) => getBackend(c, undefined, signal), runTier: runTierReal, office: tryConvertOffice, ...seams };
591
610
  const inputPath = resolve(o.path);
@@ -612,10 +631,11 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
612
631
  let office: { pdfPath: string; cleanup: () => void } | null = null;
613
632
  try {
614
633
  let pdfPath = inputPath;
615
- const base = { path: inputPath, pages: o.pages, ...(o.words && (type === "pdf" || type === "image") ? { words: true } : {}), stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr && !forced, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
634
+ const base = { path: inputPath, pages: o.pages, ...(o.words && (type === "pdf" || type === "image") ? { words: true } : {}), stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr && !forced, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs, hideAnnotations: o.hideAnnotations };
616
635
  let tier: Tier | undefined, engine: Engine | undefined, json: TierJson | undefined, degraded: string | null = null, fallbackReason: string | null = null;
617
636
  let explicitBreaks: number | null = null;
618
637
  let notes: string[] = [];
638
+ let nativeImages: NativeImage[] = [];
619
639
  let officeRoute: string | null = null;
620
640
  let copyReason: string | null = null;
621
641
  if (type === "html") {
@@ -728,17 +748,18 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
728
748
  } else {
729
749
  const p = await s.runTier("pdf-primary", pdfBase, b, signal, o.primaryTimeoutMs, backend);
730
750
  const kept = publishStaged(b);
731
- if (p.ok) { tier = "primary"; engine = "pymupdf4llm"; json = p.json; if (json.pageImages?.length) publishPageImages(b, json.pageCount ?? 0); }
751
+ if (p.ok) { tier = "primary"; engine = "pymupdf4llm"; json = p.json; nativeImages = nativeFromChild(b, json, notes); if (json.pageImages?.length) publishPageImages(b, json.pageCount ?? 0); }
732
752
  else if ("userError" in p) throw new Error(p.userError);
733
753
  else {
734
754
  if (signal?.aborted) throw new Error("aborted");
735
- const keepPages = Object.fromEntries([...kept.entries()].map(([k, v]) => [String(k), v]));
755
+ const keepPages = Object.fromEntries([...kept.entries()].map(([k, v]) => [String(k), v.files]));
736
756
  rmSync(b.pagesStagingDir, { recursive: true, force: true });
737
757
  const f = await s.runTier("pdf-fallback", { ...pdfBase, keepPages }, b, signal, o.fallbackTimeoutMs, backend);
738
758
  publishStaged(b);
739
759
  if (f.ok && f.json.pageImages?.length) publishPageImages(b, f.json.pageCount ?? 0);
740
760
  if (!f.ok) throw new Error("userError" in f ? f.userError : `Conversion failed: primary ${p.reason}; fallback ${f.reason}${detailSuffix(f)}`);
741
761
  tier = "fallback"; engine = "pymupdf-text"; json = f.json; degraded = DEGRADED_TEXT; fallbackReason = `primary ${p.reason}`;
762
+ nativeImages = [...retainedNative(b, kept, notes), ...nativeFromChild(b, json, notes)].sort((x, y) => x.page - y.page);
742
763
  }
743
764
  }
744
765
  }
@@ -789,7 +810,7 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
789
810
  const markdown = (head.length ? `${head.join("\n")}\n\n` : "") + body;
790
811
  commitBundle(b, markdown);
791
812
  const outline = scanOutline(markdown, o.outlineMaxEntries);
792
- const details: DocToMdDetails = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr, pageStats, pageStatsPath: pageStats ? b.pageStatsPath : null, ocrDir: b.ocrManifest.size ? b.ocrDir : null, wordsPath, wordsReason, wordsErrors };
813
+ const details: DocToMdDetails = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr, pageStats, pageStatsPath: pageStats ? b.pageStatsPath : null, ocrDir: b.ocrManifest.size ? b.ocrDir : null, nativeImages, wordsPath, wordsReason, wordsErrors };
793
814
  return { output: formatHandle(details), details };
794
815
  } catch (e) { abortBundle(b); throw e; }
795
816
  finally { office?.cleanup(); }
@@ -39,12 +39,15 @@ export interface OcrInfo {
39
39
  childError: string | null;
40
40
  }
41
41
 
42
+ export interface NativeImage { page: number; file: string; width: number; height: number }
43
+
42
44
  export interface HandleData {
43
45
  savedTo: string; imagesDir: string | null; sheetsDir: string | null; pagesDir: string | null; type: InputType; engine: Engine; tier: Tier;
44
46
  pageCount: number | null; pages: number[] | null; explicitBreaks: number | null; imageCount: number; pageImageCount: number; pageImagesReason: string | null; bytes: number; lines: number;
45
47
  degraded: string | null; fallbackReason: string | null; failedPages: number[]; emptyPages: number[];
46
48
  notes: string[]; outline: OutlineEntry[]; outlineTotal: number; ocr: OcrInfo | null;
47
49
  pageStats: PageStat[] | null; pageStatsPath: string | null; ocrDir: string | null;
50
+ nativeImages: NativeImage[];
48
51
  wordsPath: string | null; wordsReason: string | null; wordsErrors: Record<number, string>;
49
52
  }
50
53
 
@@ -171,6 +174,7 @@ export function formatHandle(h: HandleData): string {
171
174
  if (h.pagesDir && h.pageImageCount > 0) lines.push(`Pages-Dir: ${h.pagesDir} (${h.pageImageCount} pages)`);
172
175
  else if (h.pageImagesReason) lines.push(`Pages-Dir: none - ${h.pageImagesReason}`);
173
176
  if (h.pageStatsPath) lines.push(`Page-Stats: ${h.pageStatsPath}`);
177
+ if (h.nativeImages.length) lines.push(`Native-Images: ${h.nativeImages.length === 1 ? "page" : "pages"} ${compactRanges(h.nativeImages.map((n) => n.page))} (embedded image streams, no render DPI)`);
174
178
  if (h.wordsPath) {
175
179
  const bad = Object.keys(h.wordsErrors).map(Number).sort((a, b) => a - b);
176
180
  lines.push(`Words: ${h.wordsPath}${bad.length ? ` (extraction failed for pages ${compactRanges(bad)}: ${h.wordsErrors[bad[0]]})` : ""}`);
@@ -21,6 +21,7 @@ export interface Tunables {
21
21
  outlineMaxEntries: number;
22
22
  ocr: boolean;
23
23
  ocrLanguage: string;
24
+ hideAnnotations: boolean;
24
25
  }
25
26
 
26
27
  export interface DocToMdOptions extends Tunables {
@@ -87,6 +88,7 @@ export const DOC_TO_MD_OPTIONS: readonly OptionDescriptor[] = [
87
88
  { key: "outlineMaxEntries", flag: "--outline-max-entries", type: "int", default: 40, settable: true, help: "Heading outline / TOC / sheet inventory cap in the handle" },
88
89
  { key: "ocr", flag: "--ocr", type: "bool", default: false, settable: true, help: "Run OCR on pages without a text layer and on image inputs when Tesseract language data is installed; off by default (--no-ocr turns a settings-level true off)" },
89
90
  { key: "ocrLanguage", flag: "--ocr-language", type: "lang", default: "eng", settable: true, help: "Tesseract language code(s), +-joined, e.g. deu+eng" },
91
+ { key: "hideAnnotations", flag: "--hide-annotations", type: "bool", default: false, settable: true, help: "Render PDF pages without annotations (sticky notes, highlights, stamps - and form-field widgets, so filled form values disappear); default paints them, as PyMuPDF does. Applies to pages/ renders and textless-page renders, not to OCR text or embedded images; also lets an annotated scan be delivered as its embedded image." },
90
92
  ];
91
93
 
92
94
  export const TUNABLE_DEFAULTS: Tunables = Object.fromEntries(
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-quiver",
3
- "version": "6.10.0",
3
+ "version": "6.11.0",
4
4
  "description": "Personal pack of Pi coding-agent extensions: context-safe fetch, doc_to_md PDF/DOCX/PPTX-to-Markdown conversion, session naming, a themed ASCII startup header, Opus 4.8 fast mode, and a provider-stall watchdog.",
5
5
  "author": "Jacek Juraszek",
6
6
  "license": "MIT",
@@ -90,8 +90,9 @@ def page_dir(staging, n):
90
90
  return d
91
91
 
92
92
 
93
- def mark_done(d):
94
- open(os.path.join(d, ".done"), "w").close()
93
+ def mark_done(d, meta=None):
94
+ with open(os.path.join(d, ".done"), "w", encoding="utf-8") as fh:
95
+ json.dump(meta or {}, fh)
95
96
 
96
97
 
97
98
  OCR_SENTINEL = "\x00OCR {}\x00"
@@ -150,16 +151,232 @@ def clamped_dpi(w, h, dpi):
150
151
  return eff if eff >= MIN_RENDER_DPI else None
151
152
 
152
153
 
153
- def render_textless_page(page, d, o):
154
+ RASTER_BUDGET_S, OCR_JOB_BUDGET_S, OCR_DPI = 20, 30, 300
155
+ RASTER_SPAWN_BUDGET_S = 30
156
+ MIN_OCR_JOB_S = 1
157
+ NATIVE_MIN_COVERAGE, NATIVE_MAX_COVERAGE = 0.9, 1.1
158
+ NATIVE_EXTS = ("png", "jpeg", "jpg")
159
+ JOB_VERB = {"native": "render", "render": "render", "images": "render", "ocr": "OCR"}
160
+ RASTER_INLINE = False # tests only: run raster jobs in-process so monkeypatches reach them
161
+
162
+
163
+ class RasterError(Exception):
164
+ """A raster job did not complete; str(exc) is the failedPages message."""
165
+
166
+
167
+ class RasterUnavailable(Exception):
168
+ pass
169
+
170
+
171
+ def raster_budget(default):
172
+ return float(os.environ.get("DOC_TO_MD_RASTER_BUDGET_S", default))
173
+
174
+
175
+ def clamp_note(n, eff, requested):
176
+ return f"Page {n} rendered at {eff} dpi (requested {requested}; {MAX_RENDER_PX // 1_000_000} Mpx ceiling)"
177
+
178
+
179
+ def native_page_image(doc, page, target_dir, allow_annots=False):
180
+ """One full-page, unrotated, unmasked PNG/JPEG stream with no vector overlay -> written untouched as page.<ext>."""
181
+ import pymupdf
182
+ if not allow_annots and (page.first_annot is not None or page.first_widget is not None):
183
+ return None
184
+ infos = page.get_image_info(xrefs=True)
185
+ if len(infos) != 1 or infos[0].get("xref", 0) <= 0 or page.rotation != 0:
186
+ return None
187
+ drawings = page.get_drawings()
188
+ if allow_annots:
189
+ # MuPDF includes appearance streams, clipped to annotation/widget rectangles.
190
+ annot_rects = [a.rect + (-1, -1, 1, 1) for items in (page.annots(), page.widgets()) for a in items]
191
+ drawings = [drawing for drawing in drawings if not any(drawing["rect"] in rect for rect in annot_rects)]
192
+ if drawings:
193
+ return None
194
+ info = infos[0]
195
+ a, b, c, d = info["transform"][:4]
196
+ if b != 0 or c != 0 or a <= 0 or d <= 0:
197
+ return None
198
+ area, bbox = page.rect.get_area(), pymupdf.Rect(info["bbox"])
199
+ if area <= 0 or (bbox & page.rect).get_area() / area < NATIVE_MIN_COVERAGE or bbox.get_area() / area > NATIVE_MAX_COVERAGE:
200
+ return None
201
+ img = doc.extract_image(info["xref"])
202
+ ext, cs = img["ext"].lower(), img.get("cs-name")
203
+ plain = cs in ("DeviceRGB", "DeviceGray") or (cs == "DeviceCMYK" and ext in ("jpeg", "jpg"))
204
+ if img.get("smask", 0) or ext not in NATIVE_EXTS or not img["image"] or not plain:
205
+ return None
206
+ name = f"page.{ext}"
207
+ with open(os.path.join(target_dir, name), "wb") as fh:
208
+ fh.write(img["image"])
209
+ return {"file": name, "width": img["width"], "height": img["height"]}
210
+
211
+
212
+ def run_raster_job(doc, job):
213
+ import time
214
+ import pymupdf
215
+ n, kind = job["page"], job["kind"]
216
+ try:
217
+ if os.environ.get("DOC_TO_MD_RASTER_STALL_PAGE") in (str(n), f"{kind}:{n}"): # tests only: a wedged page
218
+ time.sleep(3600)
219
+ if os.environ.get("DOC_TO_MD_RASTER_CRASH_PAGE") == str(n): # tests only: partial output, stdout chatter, hard exit
220
+ if kind == "render":
221
+ with open(job["target"], "wb") as fh:
222
+ fh.write(b"\x89PNG partial")
223
+ print("raster worker crash hook", flush=True)
224
+ os._exit(1)
225
+ page = doc[n - 1]
226
+ if kind == "native":
227
+ r = native_page_image(doc, page, job["target"], job.get("allowAnnots", False))
228
+ return {"ok": True, "eligible": r is not None, **(r or {})}
229
+ if kind == "render":
230
+ kw = {"dpi": job["dpi"], "annots": job["annots"]}
231
+ if job.get("clip") is not None:
232
+ kw["clip"] = pymupdf.Rect(job["clip"])
233
+ pix = page.get_pixmap(**kw)
234
+ pix.save(job["target"])
235
+ return {"ok": True, "width": pix.width, "height": pix.height}
236
+ if kind == "images":
237
+ files = []
238
+ for i, image in enumerate(page.get_image_info(xrefs=True), 1):
239
+ xref = image.get("xref", 0)
240
+ if xref > 0:
241
+ img = doc.extract_image(xref)
242
+ name = f"img{i}.{img['ext'].lower()}"
243
+ with open(os.path.join(job["target"], name), "wb") as fh:
244
+ fh.write(img["image"])
245
+ else:
246
+ if job["dpi"] is None:
247
+ continue
248
+ name = f"img{i}.{job['format']}"
249
+ pix = page.get_pixmap(clip=pymupdf.Rect(image["bbox"]), dpi=job["dpi"], annots=job["annots"])
250
+ pix.save(os.path.join(job["target"], name))
251
+ files.append(name)
252
+ return {"ok": True, "files": files}
253
+ import pymupdf4llm
254
+ md = pymupdf4llm.to_markdown(doc, pages=[n - 1], write_images=False, page_separators=False, use_ocr=True, force_ocr=True,
255
+ ocr_language=job["lang"], ocr_dpi=job["ocrDpi"])
256
+ result = {"ok": True, "markdown": md}
257
+ if job.get("wordsSnapshot"):
258
+ # OCR mutates only the worker's document; carry that text layer back for parent-side word extraction.
259
+ try:
260
+ snapshot = pymupdf.open()
261
+ snapshot.insert_pdf(doc, from_page=n - 1, to_page=n - 1)
262
+ snapshot.save(job["wordsSnapshot"])
263
+ snapshot.close()
264
+ except Exception as exc: # noqa: BLE001 - geometry never changes the OCR outcome
265
+ result["wordsSnapshotError"] = words_error(exc)
266
+ return result
267
+ except Exception as exc: # noqa: BLE001 - the parent decides what a failed job costs
268
+ return {"ok": False, "error": f"{type(exc).__name__}: {exc}"[:300]}
269
+
270
+
271
+ def redirect_worker_stdout():
272
+ # redirect_stdout in the parent is a Python-object swap the spawned process does not inherit; pymupdf4llm prints to stdout.
273
+ try:
274
+ os.dup2(sys.stderr.fileno(), sys.stdout.fileno())
275
+ except (AttributeError, OSError, ValueError):
276
+ sys.stdout = sys.stderr
277
+
278
+
279
+ def raster_worker(conn, path):
280
+ redirect_worker_stdout()
281
+ import pymupdf
282
+ doc = pymupdf.open(path)
283
+ conn.send({"ready": True})
284
+ while True:
285
+ job = conn.recv()
286
+ if job["kind"] == "stop":
287
+ break
288
+ conn.send(run_raster_job(doc, job))
289
+ doc.close()
290
+
291
+
292
+ class RasterWorker:
293
+ """One spawned PyMuPDF process per conversion. MuPDF spins at C level, so only an OS kill ends a toxic page; a job that overruns its budget or takes the process down costs that job, and the next job respawns."""
294
+
295
+ def __init__(self, path, doc):
296
+ self.path, self.doc, self.proc, self.conn = path, doc, None, None
297
+
298
+ def run(self, job, budget_s):
299
+ if RASTER_INLINE:
300
+ return run_raster_job(self.doc, job)
301
+ if self.proc is None:
302
+ try:
303
+ import multiprocessing
304
+ ctx = multiprocessing.get_context("spawn")
305
+ self.conn, child = ctx.Pipe()
306
+ self.proc = ctx.Process(target=raster_worker, args=(child, self.path))
307
+ try:
308
+ self.proc.start()
309
+ finally:
310
+ child.close()
311
+ except Exception as exc:
312
+ if self.conn is not None:
313
+ self.conn.close()
314
+ self.proc = self.conn = None
315
+ raise RasterUnavailable(f"raster worker unavailable: {exc}") from exc
316
+ try:
317
+ ready = self.conn.poll(RASTER_SPAWN_BUDGET_S) and self.conn.recv() == {"ready": True}
318
+ except (EOFError, OSError):
319
+ ready = False
320
+ if not ready:
321
+ self.kill()
322
+ return {"ok": False, "error": "renderer failed to start"}
323
+ try:
324
+ self.conn.send(job)
325
+ if self.conn.poll(budget_s):
326
+ return self.conn.recv()
327
+ error = f"{JOB_VERB[job['kind']]} timed out after {budget_s:g}s"
328
+ except (EOFError, OSError):
329
+ error = "renderer crashed"
330
+ self.kill()
331
+ if job["kind"] == "render":
332
+ with contextlib.suppress(OSError):
333
+ os.remove(job["target"])
334
+ if job["kind"] == "ocr" and job.get("wordsSnapshot"):
335
+ with contextlib.suppress(OSError):
336
+ os.remove(job["wordsSnapshot"])
337
+ return {"ok": False, "error": error}
338
+
339
+ def kill(self):
340
+ if self.proc is not None:
341
+ if self.proc.is_alive():
342
+ self.proc.kill()
343
+ self.proc.join()
344
+ self.conn.close()
345
+ self.proc = self.conn = None
346
+
347
+ def stop(self):
348
+ if self.proc is None:
349
+ return
350
+ with contextlib.suppress(OSError):
351
+ self.conn.send({"kind": "stop"})
352
+ self.proc.join(2)
353
+ self.kill()
354
+
355
+
356
+ def textless_picture(worker, page, n, d, o, notes, render):
357
+ """Native stream when the page is one full-page image, else (when render) a worker render at the clamped DPI. Returns (file name or None, .done metadata)."""
358
+ r = worker.run({"kind": "native", "page": n, "target": d, "allowAnnots": bool(o.get("hideAnnotations"))}, raster_budget(RASTER_BUDGET_S))
359
+ if r.get("ok") and r.get("eligible"):
360
+ return r["file"], {"native": {"file": r["file"], "width": r["width"], "height": r["height"]}}
361
+ for f in os.listdir(d):
362
+ if f.startswith("page."): # a failed native job may have left a partial stream
363
+ os.remove(os.path.join(d, f))
364
+ if not render:
365
+ return None, {}
154
366
  eff = clamped_dpi(page.rect.width, page.rect.height, o["imageDpi"])
155
367
  if eff is None:
156
- return None
368
+ return None, {}
157
369
  name = f"page.{o['imageFormat']}"
158
- page.get_pixmap(dpi=eff).save(os.path.join(d, name))
159
- return name
370
+ r = worker.run({"kind": "render", "page": n, "dpi": eff, "annots": not o.get("hideAnnotations"), "target": os.path.join(d, name)}, raster_budget(RASTER_BUDGET_S))
371
+ if not r["ok"]:
372
+ raise RasterError(r["error"])
373
+ if eff < o["imageDpi"]:
374
+ notes.append(clamp_note(n, eff, o["imageDpi"]))
375
+ return name, {"dpi": eff, "requestedDpi": o["imageDpi"]}
376
+ return name, {}
160
377
 
161
378
 
162
- def render_page_image(page, n, o, page_images):
379
+ def render_page_image(worker, page, n, o, page_images, notes):
163
380
  d = o["pagesStagingDir"]
164
381
  eff = clamped_dpi(page.rect.width, page.rect.height, o["imageDpi"])
165
382
  if eff is None:
@@ -167,18 +384,24 @@ def render_page_image(page, n, o, page_images):
167
384
  os.makedirs(d, exist_ok=True)
168
385
  name = f"p{n}.{o['imageFormat']}"
169
386
  target = os.path.join(d, name)
170
- try:
171
- page.get_pixmap(dpi=eff).save(target)
172
- except Exception:
173
- try:
387
+ r = worker.run({"kind": "render", "page": n, "dpi": eff, "annots": not o.get("hideAnnotations"), "target": target}, raster_budget(RASTER_BUDGET_S))
388
+ if not r["ok"]:
389
+ notes.append(f"Page {n} render unavailable: {r['error']}")
390
+ with contextlib.suppress(OSError):
174
391
  os.remove(target)
175
- except OSError:
176
- pass
177
392
  return None
178
- page_images.append({"page": n, "file": name})
393
+ entry = {"page": n, "file": name, "dpi": eff}
394
+ if eff < o["imageDpi"]:
395
+ entry["requestedDpi"] = o["imageDpi"]
396
+ notes.append(clamp_note(n, eff, o["imageDpi"]))
397
+ page_images.append(entry)
179
398
  return name
180
399
 
181
400
 
401
+ def page_failure(n, exc):
402
+ return {"page": n, "error": str(exc) if isinstance(exc, RasterError) else f"{type(exc).__name__}: {exc}"[:300]}
403
+
404
+
182
405
  def page_ocr_kwargs(textless, lang):
183
406
  if textless:
184
407
  return {"use_ocr": True, "force_ocr": True, "ocr_language": lang}
@@ -359,72 +582,113 @@ def mode_pdf_primary(o):
359
582
  pages = check_pages(o.get("pages"), doc.page_count)
360
583
  staging, out, empty, failed, notes = o["stagingDir"], [], [], [], []
361
584
  page_images = []
585
+ native_images = []
586
+ worker = RasterWorker(o["path"], doc)
362
587
  stats = []
363
588
  lang, budget = o.get("ocrLanguage", "eng"), o.get("ocrBudgetMs", 60000)
364
589
  info = new_ocr(lang)
365
590
  status = ocr_status(True, lang) if o.get("ocr") else None
366
591
  ocr_ms, plain_ms = [], []
367
592
  word_pages, words_errors = [], {}
368
- for i, n in enumerate(pages):
369
- stats.append(page_stats(doc, n))
370
- d = page_dir(staging, n)
371
- try:
372
- page = doc[n - 1]
373
- rotation = page.rotation
374
- textless = not page.get_text("text").strip()
375
- if textless:
376
- info["textless"].append(n)
377
- if status is None:
378
- status = ocr_status(False, lang)
379
- kw = page_ocr_kwargs(textless, lang) if status and status["status"] == "ready" else {"use_ocr": False}
380
- if kw["use_ocr"]:
381
- elapsed = (time.monotonic() - start) * 1000
382
- est_ocr = max(ocr_ms) if ocr_ms else OCR_EST_INITIAL_MS
383
- est_page = sum(plain_ms) / len(plain_ms) if plain_ms else PAGE_EST_INITIAL_MS
384
- if not ocr_admit(elapsed, est_ocr, len(pages) - i - 1, est_page, budget):
385
- info["budgetStopped"].append(n)
386
- kw = {"use_ocr": False}
387
- t0 = time.monotonic()
593
+ try:
594
+ for i, n in enumerate(pages):
595
+ stats.append(page_stats(doc, n))
596
+ d = page_dir(staging, n)
388
597
  try:
389
- md = primary_page_markdown(doc, n, d, o, kw, not textless) if kw["use_ocr"] or not textless else ""
390
- except Exception: # noqa: BLE001
391
- if not kw["use_ocr"]:
392
- raise
393
- kw, t0, md = {"use_ocr": False}, time.monotonic(), ""
394
- info["ocrFailed"].append(n)
395
- (ocr_ms if kw["use_ocr"] else plain_ms).append((time.monotonic() - t0) * 1000)
396
- if textless:
397
- pic = None if o.get("pageImages") else render_textless_page(page, d, o)
398
- text = md.strip()
399
- if kw["use_ocr"] and text:
400
- info["pages"].append(n)
401
- else:
598
+ page = doc[n - 1]
599
+ rotation = page.rotation
600
+ textless = not page.get_text("text").strip()
601
+ if textless:
602
+ info["textless"].append(n)
603
+ if status is None:
604
+ status = ocr_status(False, lang)
605
+ kw = page_ocr_kwargs(textless, lang) if status and status["status"] == "ready" else {"use_ocr": False}
606
+ if kw["use_ocr"]:
607
+ elapsed = (time.monotonic() - start) * 1000
608
+ est_ocr = max(ocr_ms) if ocr_ms else OCR_EST_INITIAL_MS
609
+ est_page = sum(plain_ms) / len(plain_ms) if plain_ms else PAGE_EST_INITIAL_MS
610
+ if not ocr_admit(elapsed, est_ocr, len(pages) - i - 1, est_page, budget):
611
+ info["budgetStopped"].append(n)
612
+ kw = {"use_ocr": False}
613
+ text, meta = "", {}
614
+ ocr_words = None
615
+ if textless:
616
+ pic, meta = textless_picture(worker, page, n, d, o, notes, render=not o.get("pageImages"))
617
+ t0 = time.monotonic()
402
618
  if kw["use_ocr"]:
403
- info["noText"].append(n)
619
+ remaining = (budget - (time.monotonic() - start) * 1000 - OCR_BUDGET_RESERVE_MS) / 1000
620
+ if remaining < MIN_OCR_JOB_S:
621
+ info["budgetStopped"].append(n)
622
+ kw = {"use_ocr": False}
623
+ else:
624
+ ocr_dpi = clamped_dpi(page.rect.width, page.rect.height, OCR_DPI)
625
+ words_snapshot = os.path.join(staging, f".ocr-words-p{n}.pdf") if o.get("words") else None
626
+ try:
627
+ r = worker.run({"kind": "ocr", "page": n, "lang": lang, "ocrDpi": ocr_dpi, **({"wordsSnapshot": words_snapshot} if words_snapshot else {})}, min(raster_budget(OCR_JOB_BUDGET_S), remaining)) if ocr_dpi is not None else {"ok": False}
628
+ if words_snapshot and r["ok"]:
629
+ if r.get("wordsSnapshotError"):
630
+ words_errors[str(n)] = r["wordsSnapshotError"]
631
+ else:
632
+ try:
633
+ with open_pdf(words_snapshot) as snapshot:
634
+ ocr_words = words_page(snapshot[0], n, rotation, page_words(snapshot[0], True))
635
+ except Exception as exc: # noqa: BLE001
636
+ words_errors[str(n)] = words_error(exc)
637
+ finally:
638
+ if words_snapshot:
639
+ with contextlib.suppress(OSError):
640
+ os.remove(words_snapshot)
641
+ if r["ok"]:
642
+ text = r["markdown"].strip()
643
+ else:
644
+ kw = {"use_ocr": False}
645
+ info["ocrFailed"].append(n)
646
+ md = ""
647
+ else:
648
+ t0 = time.monotonic()
649
+ md = primary_page_markdown(doc, n, d, o, kw, True)
650
+ if not (textless and not kw["use_ocr"]):
651
+ (ocr_ms if kw["use_ocr"] else plain_ms).append((time.monotonic() - t0) * 1000)
652
+ if textless:
653
+ if kw["use_ocr"] and text:
654
+ info["pages"].append(n)
655
+ else:
656
+ if kw["use_ocr"]:
657
+ info["noText"].append(n)
658
+ empty.append(n)
659
+ elif not md.strip():
404
660
  empty.append(n)
405
- elif not md.strip():
661
+ mark_done(d, meta)
662
+ page_pic = render_page_image(worker, page, n, o, page_images, notes) if o.get("pageImages") else None
663
+ if textless:
664
+ parts = [f"![page {n}](p{n}/{pic})"] if pic else []
665
+ if kw["use_ocr"] and text:
666
+ parts.append(ocr_block(f"p{n}/{pic}" if pic else f"pages/{page_pic}" if page_pic else "-", text))
667
+ md = "\n\n".join(parts)
668
+ if page_pic:
669
+ md = "\n\n".join(x for x in [md.rstrip(), f"![page {n}](pages/{page_pic})"] if x)
670
+ out.append(md.rstrip())
671
+ if o.get("words") and not (textless and o.get("ocr") and not kw["use_ocr"]):
672
+ try:
673
+ if textless and kw["use_ocr"]:
674
+ if ocr_words is not None:
675
+ word_pages.append(ocr_words)
676
+ else:
677
+ word_pages.append(words_page(page, n, rotation, page_words(page, kw["use_ocr"])))
678
+ except Exception as exc: # noqa: BLE001
679
+ words_errors[str(n)] = words_error(exc)
680
+ if meta.get("native"):
681
+ native_images.append({"page": n, **meta["native"]})
682
+ except RasterUnavailable:
683
+ raise
684
+ except Exception as exc: # noqa: BLE001
685
+ shutil.rmtree(d, ignore_errors=True)
686
+ failed.append(page_failure(n, exc))
406
687
  empty.append(n)
407
- mark_done(d)
408
- page_pic = render_page_image(page, n, o, page_images) if o.get("pageImages") else None
409
- if textless:
410
- parts = [f"![page {n}](p{n}/{pic})"] if pic else []
411
- if kw["use_ocr"] and text:
412
- parts.append(ocr_block(f"pages/{page_pic}" if page_pic else f"p{n}/{pic}" if pic else "-", text))
413
- md = "\n\n".join(parts)
414
- if page_pic:
415
- md = "\n\n".join(x for x in [md.rstrip(), f"![page {n}](pages/{page_pic})"] if x)
416
- out.append(md.rstrip())
417
- if o.get("words") and not (textless and o.get("ocr") and not kw["use_ocr"]):
418
- try:
419
- word_pages.append(words_page(page, n, rotation, page_words(page, kw["use_ocr"])))
420
- except Exception as exc: # noqa: BLE001
421
- words_errors[str(n)] = words_error(exc)
422
- except Exception as exc: # noqa: BLE001
423
- shutil.rmtree(d, ignore_errors=True)
424
- failed.append({"page": n, "error": f"{type(exc).__name__}: {exc}"[:300]})
425
- empty.append(n)
426
- out.append("")
427
- out.append(SEP.format(n=n).strip("\n"))
688
+ out.append("")
689
+ out.append(SEP.format(n=n).strip("\n"))
690
+ finally:
691
+ worker.stop()
428
692
  if failed and len(failed) == len(pages):
429
693
  raise RuntimeError("every selected page failed: " + failed[0]["error"])
430
694
  if status is not None:
@@ -433,7 +697,7 @@ def mode_pdf_primary(o):
433
697
  if missing:
434
698
  notes.append(f"Page images: {missing} of {len(pages)} unavailable")
435
699
  result = {"markdown": "\n\n".join(out) + "\n", "pages": pages, "pageCount": doc.page_count,
436
- "emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": info, "pageImages": page_images, "pageStats": stats}
700
+ "emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": info, "pageImages": page_images, "nativeImages": native_images, "pageStats": stats}
437
701
  if o.get("words"):
438
702
  result.update(write_words(staging, "pt", word_pages, words_errors))
439
703
  return result
@@ -446,72 +710,75 @@ def mode_pdf_fallback(o):
446
710
  keep = {int(k): v for k, v in (o.get("keepPages") or {}).items()}
447
711
  staging, out, empty, failed = o["stagingDir"], [], [], []
448
712
  page_images = []
713
+ native_images = []
714
+ worker = RasterWorker(o["path"], doc)
449
715
  stats = []
450
716
  lang = o.get("ocrLanguage", "eng")
451
717
  ocr_info = new_ocr(lang)
452
718
  word_pages, words_errors = [], {}
453
719
  status = {"status": "unavailable", "reason": "fallback tier", "tesseract": None} if o.get("ocr") else None
454
- for n in pages:
455
- stats.append(page_stats(doc, n))
456
- links = [f"![](images/{f})" for f in keep.get(n, [])]
457
- text = ""
458
- page_pic = None
459
- page_ok = False
460
- try:
461
- page = doc[n - 1]
462
- rotation = page.rotation
463
- text = page.get_text("text").strip()
464
- if not text:
465
- ocr_info["textless"].append(n)
466
- if status is None:
467
- status = ocr_status(False, lang)
468
- if n not in keep:
469
- d = page_dir(staging, n)
470
- if not text and not o.get("pageImages"):
471
- pic = render_textless_page(page, d, o)
472
- if pic:
473
- links.append(f"![page {n}](p{n}/{pic})")
474
- elif text:
475
- i = 0
476
- for image in page.get_image_info(xrefs=True):
477
- i += 1
478
- xref = image.get("xref", 0)
479
- if xref > 0:
480
- img = doc.extract_image(xref)
481
- name = f"img{i}.{img['ext'].lower()}"
482
- with open(os.path.join(d, name), "wb") as fh:
483
- fh.write(img["image"])
484
- else:
485
- name = f"img{i}.{o['imageFormat']}"
486
- page.get_pixmap(clip=pymupdf.Rect(image["bbox"]), dpi=o["imageDpi"]).save(os.path.join(d, name))
487
- links.append(f"![](p{n}/{name})")
488
- mark_done(d)
489
- page_pic = render_page_image(page, n, o, page_images) if o.get("pageImages") else None
490
- page_ok = True
491
- except Exception as exc: # noqa: BLE001
492
- shutil.rmtree(os.path.join(staging, f"p{n}"), ignore_errors=True)
493
- text = ""
720
+ notes = [DEGRADED_NOTE]
721
+ try:
722
+ for n in pages:
723
+ stats.append(page_stats(doc, n))
494
724
  links = [f"![](images/{f})" for f in keep.get(n, [])]
495
- failed.append({"page": n, "error": f"{type(exc).__name__}: {exc}"[:300]})
496
- if o.get("words") and page_ok and not (not text and o.get("ocr")):
725
+ text = ""
726
+ page_pic = None
727
+ meta = {}
728
+ page_ok = False
497
729
  try:
498
- word_pages.append(words_page(page, n, rotation, page_words(page, False)))
730
+ page = doc[n - 1]
731
+ rotation = page.rotation
732
+ text = page.get_text("text").strip()
733
+ if not text:
734
+ ocr_info["textless"].append(n)
735
+ if status is None:
736
+ status = ocr_status(False, lang)
737
+ if n not in keep:
738
+ d = page_dir(staging, n)
739
+ if not text:
740
+ pic, meta = textless_picture(worker, page, n, d, o, notes, render=not o.get("pageImages"))
741
+ if pic:
742
+ links.append(f"![page {n}](p{n}/{pic})")
743
+ else:
744
+ eff = clamped_dpi(page.rect.width, page.rect.height, o["imageDpi"])
745
+ r = worker.run({"kind": "images", "page": n, "dpi": eff, "annots": not o.get("hideAnnotations"), "format": o["imageFormat"], "target": d}, raster_budget(RASTER_BUDGET_S))
746
+ if not r["ok"]:
747
+ raise RasterError(r["error"])
748
+ links.extend(f"![](p{n}/{name})" for name in r["files"])
749
+ mark_done(d, meta)
750
+ page_pic = render_page_image(worker, page, n, o, page_images, notes) if o.get("pageImages") else None
751
+ page_ok = True
752
+ except RasterUnavailable:
753
+ raise
499
754
  except Exception as exc: # noqa: BLE001
500
- words_errors[str(n)] = words_error(exc)
501
- if not text:
502
- empty.append(n)
503
- out.append("\n\n".join(x for x in [text, "\n".join(links), f"![page {n}](pages/{page_pic})" if page_pic else ""] if x))
504
- out.append(SEP.format(n=n).strip("\n"))
755
+ shutil.rmtree(os.path.join(staging, f"p{n}"), ignore_errors=True)
756
+ text = ""
757
+ meta = {}
758
+ links = [f"![](images/{f})" for f in keep.get(n, [])]
759
+ failed.append(page_failure(n, exc))
760
+ if o.get("words") and page_ok and not (not text and o.get("ocr")):
761
+ try:
762
+ word_pages.append(words_page(page, n, rotation, page_words(page, False)))
763
+ except Exception as exc: # noqa: BLE001
764
+ words_errors[str(n)] = words_error(exc)
765
+ if not text:
766
+ empty.append(n)
767
+ out.append("\n\n".join(x for x in [text, "\n".join(links), f"![page {n}](pages/{page_pic})" if page_pic else ""] if x))
768
+ if meta.get("native"):
769
+ native_images.append({"page": n, **meta["native"]})
770
+ out.append(SEP.format(n=n).strip("\n"))
771
+ finally:
772
+ worker.stop()
505
773
  if failed and len(failed) == len(pages):
506
774
  raise RuntimeError("every selected page failed: " + failed[0]["error"])
507
775
  if status is not None:
508
776
  apply_status(ocr_info, status)
509
- notes = [DEGRADED_NOTE]
510
777
  missing = len(pages) - len(page_images) if o.get("pageImages") else 0
511
778
  if missing:
512
779
  notes.append(f"Page images: {missing} of {len(pages)} unavailable")
513
780
  result = {"markdown": "\n\n".join(out) + "\n", "pages": pages, "pageCount": doc.page_count,
514
- "emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": ocr_info, "pageImages": page_images, "pageStats": stats}
781
+ "emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": ocr_info, "pageImages": page_images, "nativeImages": native_images, "pageStats": stats}
515
782
  if o.get("words"):
516
783
  result.update(write_words(staging, "pt", word_pages, words_errors))
517
784
  return result
@@ -614,7 +881,7 @@ def col_letter(i):
614
881
 
615
882
  PREVIEW_ROWS, PREVIEW_COLS = 100, 50
616
883
  PROFILE_MAJORITY, DISTINCT_CAP, SLUG_MAX = 0.6, 50, 40
617
- MAX_RENDER_PX, MIN_RENDER_DPI, MIN_PAGE_PT = 16_000_000, 36, 72
884
+ MAX_RENDER_PX, MIN_RENDER_DPI, MIN_PAGE_PT = 50_000_000, 36, 72
618
885
  INV_HEADER = "| # | name | kind | size | hidden | charts | images | rendered | data |\n|---|---|---|---|---|---|---|---|---|\n"
619
886
  XLS_NOTE = "Rendered views: unavailable (visual detection not supported for .xls)"
620
887
 
@@ -1424,7 +1691,6 @@ def mode_info_excel(o):
1424
1691
  return {"sheets": sheets}
1425
1692
 
1426
1693
  def mode_render_pages(o):
1427
- import math
1428
1694
  import pymupdf
1429
1695
  doc = pymupdf.open(o["path"])
1430
1696
  expected = o["expectedPages"]
@@ -1439,8 +1705,8 @@ def mode_render_pages(o):
1439
1705
  if w < MIN_PAGE_PT or h < MIN_PAGE_PT:
1440
1706
  failed.append({"idx": idx, "reason": f"rendered view degenerate (page {w:.0f} x {h:.0f} pt)"})
1441
1707
  continue
1442
- eff = min(dpi, math.floor(math.sqrt(MAX_RENDER_PX / (w * h / 72 ** 2))))
1443
- if eff < MIN_RENDER_DPI:
1708
+ eff = clamped_dpi(w, h, dpi)
1709
+ if eff is None:
1444
1710
  failed.append({"idx": idx, "reason": f"rendered view too large (page {w:.0f} x {h:.0f} pt)"})
1445
1711
  continue
1446
1712
  name = f"s{idx}.{fmt}"