pi-quiver 6.10.0 → 6.11.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +8 -0
- package/README.md +4 -3
- package/dist/bin/pi-quiver.js +34 -5
- package/lib/doc-to-md-bundle.ts +14 -4
- package/lib/doc-to-md-core.ts +28 -7
- package/lib/doc-to-md-handle.ts +4 -0
- package/lib/doc-to-md-options.ts +2 -0
- package/package.json +1 -1
- package/scripts/doc_to_md.py +391 -125
package/CHANGELOG.md
CHANGED
|
@@ -8,6 +8,14 @@ Published to npm as `pi-quiver` (`pi install npm:pi-quiver`). Pushing a
|
|
|
8
8
|
via OIDC trusted publishing. The release helper at
|
|
9
9
|
`.agents/skills/release/scripts/release.sh` cuts the tag; CI publishes.
|
|
10
10
|
|
|
11
|
+
## v6.11.0 - 2026-10-02
|
|
12
|
+
|
|
13
|
+
- `doc_to_md`: textless PDF pages holding one eligible full-page image deliver the embedded stream without re-encoding; the handle prints `Native-Images:` and tool details / CLI JSON carry `nativeImages` (#27).
|
|
14
|
+
- `doc_to_md`: native probing, PDF page renders, and textless-page inline OCR run in a spawned raster worker with per-job budgets; a toxic page becomes a `Failed pages` note instead of failing the conversion when other pages succeed (#27).
|
|
15
|
+
- `doc_to_md`: the render ceiling rises from 16 to 50 Mpx; clamped PDF renders add Notes and report effective `pageImages[].dpi` plus `requestedDpi` (#27).
|
|
16
|
+
- `doc_to_md`: `hideAnnotations` / `--hide-annotations` hides PDF annotations and form widgets on renders and lets annotated scans use native delivery; default runs keep their painted render (#27).
|
|
17
|
+
- `doc_to_md`: `.done` markers retain native-image and clamp metadata for completed pages across primary-tier failure and fallback; corrupt metadata does not discard completed images (#27).
|
|
18
|
+
|
|
11
19
|
## v6.10.0 - 2026-10-02
|
|
12
20
|
|
|
13
21
|
- `doc_to_md`: new per-call `words` option (`--words`; not settable) writes `<stem>.words.json` - per selected PDF/image page, every text-layer word and every word inline OCR recognized in the same run, with display-space bbox (points; source pixels for images) and `source` `text`/`ocr`. Under `--ocr-mode all` each sidecar gets `ocr/<stem>-pNNN.words.json`. Never triggers OCR; Markdown, page stats and OCR outcome are unchanged. Handle gains `Words:`, `--json` gains `wordsPath`, `wordsReason`, `wordsErrors`, `ocr.wordSidecars`; `--info --words` is a usage error (#28).
|
package/README.md
CHANGED
|
@@ -65,7 +65,7 @@ A 300 KB changelog page never touches your context window - you get a preview an
|
|
|
65
65
|
| Extension | Tool | What it does |
|
|
66
66
|
| --- | --- | --- |
|
|
67
67
|
| `extensions/fetch.ts` | `fetch` | Retrieve URLs over HTTP(S). HTML -> Markdown (Readability extraction, Turndown conversion). Binary saved untouched to a temp file. GitHub issue/PR/repo/actions-run/actions-job URLs auto-route through `gh` (falls back to HTTP); failed runs/jobs include failed-step logs (best-effort, summary-only otherwise). Same size gate as `fetch`. Behavior lives in `lib/fetch-core.ts`; also exposed as the `pi-quiver fetch` CLI (see [Claude Code support](#claude-code-support)). |
|
|
68
|
-
| `extensions/doc_to_md.ts` | `doc_to_md` | Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, MSG/EML, HTML, or image file to a Markdown bundle on disk (`<stem>.md` + `images/`, spreadsheet `sheets/`, optional `pageImages` in `pages/`, and email `attachments/`) and return a handle (paths, page count, outline, diagnostics) - never inline Markdown. `info` inspects PDF/Office inputs; `pages` selects 1-based PDF/Office pages (DOCX: explicit-page-break segments); HTML, image, and email inputs reject both; every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Tiers: pymupdf4llm -> PyMuPDF text (degraded) -> unpdf worker (no Python only). Excel -> sheet inventory, full CSVs, bounded previews, and optional rendered views. HTML keeps local and `data:` images in the bundle and remote images as links; image inputs keep the original image. Scanned PDF pages keep a page picture on both Python tiers; opt-in OCR needs Tesseract language data. DOCX converts directly (mammoth, python-docx fallback); LibreOffice pagination marks the degraded route. The Outline lists `L<line>` and `p<page>` per heading. Settings under `quiver.docToMd`. Every conversion that goes through the Python PDF tiers (PDF, PPTX, `.doc`, DOCX via LibreOffice) writes `<stem>.pages.json` (per-page chars, image count, image coverage; handle line `Page-Stats:`); unpdf writes no stats. `ocrMode: "all"` with `ocr: true` and an explicit `pages` forces OCR on those pages into `ocr/<stem>-pNNN.md` sidecars, leaving the Markdown byte-identical to the same selected-page call without OCR. `words: true` (CLI `--words`, per-call only - no settings key) writes `<stem>.words.json` with every word's bbox per selected PDF/image page (`source` `text` or `ocr`), plus `ocr/<stem>-pNNN.words.json` beside each forced-OCR sidecar; it never triggers OCR. |
|
|
68
|
+
| `extensions/doc_to_md.ts` | `doc_to_md` | Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, MSG/EML, HTML, or image file to a Markdown bundle on disk (`<stem>.md` + `images/`, spreadsheet `sheets/`, optional `pageImages` in `pages/`, and email `attachments/`) and return a handle (paths, page count, outline, diagnostics) - never inline Markdown. `info` inspects PDF/Office inputs; `pages` selects 1-based PDF/Office pages (DOCX: explicit-page-break segments); HTML, image, and email inputs reject both; every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Tiers: pymupdf4llm -> PyMuPDF text (degraded) -> unpdf worker (no Python only). Excel -> sheet inventory, full CSVs, bounded previews, and optional rendered views. HTML keeps local and `data:` images in the bundle and remote images as links; image inputs keep the original image. Scanned PDF pages keep a page picture on both Python tiers; eligible single full-page images are delivered as their embedded stream, without re-encoding or render DPI; opt-in OCR needs Tesseract language data. DOCX converts directly (mammoth, python-docx fallback); LibreOffice pagination marks the degraded route. The Outline lists `L<line>` and `p<page>` per heading. Settings under `quiver.docToMd`. Every conversion that goes through the Python PDF tiers (PDF, PPTX, `.doc`, DOCX via LibreOffice) writes `<stem>.pages.json` (per-page chars, image count, image coverage; handle line `Page-Stats:`); unpdf writes no stats. `ocrMode: "all"` with `ocr: true` and an explicit `pages` forces OCR on those pages into `ocr/<stem>-pNNN.md` sidecars, leaving the Markdown byte-identical to the same selected-page call without OCR. `words: true` (CLI `--words`, per-call only - no settings key) writes `<stem>.words.json` with every word's bbox per selected PDF/image page (`source` `text` or `ocr`), plus `ocr/<stem>-pNNN.words.json` beside each forced-OCR sidecar; it never triggers OCR. |
|
|
69
69
|
| `extensions/session-name.ts` | `/session-name` | Manual + opt-in automatic session naming, naming rules and deny list, long-session revisits, and Ghostty/Herdr tab rename. OFF by default. |
|
|
70
70
|
| `extensions/sword-header.ts` | `/builtin-header` | Themed ASCII startup header replacing pi's default logo. OFF by default. |
|
|
71
71
|
| `extensions/fast-mode.ts` | `/fast` | Inject Anthropic fast-mode (`speed: "fast"` + `anthropic-beta: fast-mode-2026-02-01`) into every Claude Opus 4.8 / Opus 5 request, any thinking level. `--fast` flag + `/fast [on\|off\|status]`. OFF by default. |
|
|
@@ -303,8 +303,9 @@ Each setting can also be overridden per-process via `PI_QUIVER_SLACK_ENABLED`, `
|
|
|
303
303
|
| `excelTimeoutMs` | `60000` | Excel child and Excel info deadline. |
|
|
304
304
|
| `warmTimeoutMs` | `120000` | Absolute first-call backend discovery/bootstrap deadline. |
|
|
305
305
|
| `pymupdfVersion` | `1.27.2.3` | pymupdf4llm pin, minimum `1.27.0`. |
|
|
306
|
-
| `imageDpi` | `150` | Render DPI for page images and Excel rendered views (capped by a
|
|
306
|
+
| `imageDpi` | `150` | Render DPI for page images and Excel rendered views (capped by a 50 Mpx budget). |
|
|
307
307
|
| `imageFormat` | `png` | Rendered image format: `png` or `jpg`; embedded images retain their extension. |
|
|
308
|
+
| `hideAnnotations` | `false` | Hide annotations on PDF textless-page and `pages/` renders, including form-field widgets (filled values disappear). Default paints them. Does not affect OCR text or embedded images; also lets an annotated scan be delivered as its embedded image. |
|
|
308
309
|
| `maxOutputBytes` | `20000000` | Child stdout cap in bytes. |
|
|
309
310
|
| `outlineMaxEntries` | `40` | Heading outline, TOC, or sheet inventory entries in the handle. |
|
|
310
311
|
| `ocr` | `false` | Opt-in OCR for scanned pages and images when Tesseract language data is installed. |
|
|
@@ -314,7 +315,7 @@ Each setting can also be overridden per-process via `PI_QUIVER_SLACK_ENABLED`, `
|
|
|
314
315
|
|
|
315
316
|
A bundle is `<outputDir>/<stem>.md` plus `<outputDir>/<stem>.pages.json` on Python PDF tiers, optional OCR sidecars under `<outputDir>/ocr/`, and `<outputDir>/images/` and, for spreadsheets with data, `<outputDir>/sheets/`; without `outputDir`, the tool creates a per-call temp root. A conversion owns `<stem>.md.lock` until it atomically publishes the Markdown. An existing `<stem>.md` fails the call unless `overwrite` is set; `overwrite` replaces that Markdown and the files it owns (`images/<stem>-p<N>-<n>.*`, `images/<stem>-s<idx>[-<n>].*`, `sheets/<stem>-s<idx>-<slug>.csv`, `<stem>.pages.json`, `ocr/<stem>-pNNN.md`), nothing else. Temp bundles are caller-owned - the tool never deletes a bundle it produced.
|
|
316
317
|
|
|
317
|
-
Excel needs a Python backend with openpyxl, xlrd and pillow - otherwise the call fails with `Remedy: install uv, or pip install openpyxl xlrd pillow`. The Markdown opens with a `## Sheets` table listing every sheet in workbook order (0-based `#`, `worksheet`/`chartsheet`, size, hidden, chart and image counts, rendered view, CSV link for non-empty worksheets), then one section per sheet: a `Data:` line linking the full-content CSV under `sheets/` for non-empty worksheets, chart metadata from the workbook model, embedded images, an optional rendered view, a preview of at most the first 100 rows x 50 columns, and - only when the preview is truncated - a `Columns:` profile (type, non-empty count, min/max, distinct up to 50). Sizes are the extent of non-empty cells (the `info` handle reports the raw worksheet dimensions instead, which may be larger). Rendered views (`images/<stem>-s<idx>.<fmt>`) are produced for sheets carrying charts or images when LibreOffice is on `PATH`: the workbook is exported one PDF page per sheet and rasterized under a
|
|
318
|
+
Excel needs a Python backend with openpyxl, xlrd and pillow - otherwise the call fails with `Remedy: install uv, or pip install openpyxl xlrd pillow`. The Markdown opens with a `## Sheets` table listing every sheet in workbook order (0-based `#`, `worksheet`/`chartsheet`, size, hidden, chart and image counts, rendered view, CSV link for non-empty worksheets), then one section per sheet: a `Data:` line linking the full-content CSV under `sheets/` for non-empty worksheets, chart metadata from the workbook model, embedded images, an optional rendered view, a preview of at most the first 100 rows x 50 columns, and - only when the preview is truncated - a `Columns:` profile (type, non-empty count, min/max, distinct up to 50). Sizes are the extent of non-empty cells (the `info` handle reports the raw worksheet dimensions instead, which may be larger). Rendered views (`images/<stem>-s<idx>.<fmt>`) are produced for sheets carrying charts or images when LibreOffice is on `PATH`: the workbook is exported one PDF page per sheet and rasterized under a 50 Mpx budget. Any LibreOffice or rasterization failure degrades to `Rendered view: unavailable (<reason>)` and a handle note; it never fails the conversion. Workbooks whose chartsheet drawings carry a zero-size anchor (openpyxl-authored files; Excel-authored files are unaffected) render as a degenerate page and are reported as such. `.xls` gets the inventory, CSVs and previews but no visual detection or rendering.
|
|
318
319
|
|
|
319
320
|
Worst-case wall time: PDF `warmTimeoutMs (first call) + primaryTimeoutMs + fallbackTimeoutMs`; PPTX adds `sofficeTimeoutMs`; DOCX on the Python path `warmTimeoutMs + primaryTimeoutMs` (success or a terminal child failure), DOCX child exit 1 then LibreOffice `warmTimeoutMs + primaryTimeoutMs + sofficeTimeoutMs + primaryTimeoutMs + fallbackTimeoutMs`, DOCX without a DOCX-capable backend `warmTimeoutMs + sofficeTimeoutMs + primaryTimeoutMs + fallbackTimeoutMs`; Excel `warmTimeoutMs + excelTimeoutMs + sofficeTimeoutMs + fallbackTimeoutMs`. Add `KILL_GRACE_MS` (2000 ms) per kill. There is no cap on image count, image bytes, cell count or workbook memory - deliberately; the per-tier timeouts, the rendered-view pixel budget and `maxOutputBytes` are the bounds.
|
|
320
321
|
|
package/dist/bin/pi-quiver.js
CHANGED
|
@@ -680,6 +680,13 @@ function publishStaged(b) {
|
|
|
680
680
|
continue;
|
|
681
681
|
}
|
|
682
682
|
const page = Number(m[1]);
|
|
683
|
+
const raw = readFileSync(join2(pageDir, ".done"), "utf8").trim();
|
|
684
|
+
let meta = {};
|
|
685
|
+
try {
|
|
686
|
+
const v = raw ? JSON.parse(raw) : {};
|
|
687
|
+
if (v && typeof v === "object" && !Array.isArray(v)) meta = v;
|
|
688
|
+
} catch {
|
|
689
|
+
}
|
|
683
690
|
const files = readdirSync(pageDir).filter((f) => f !== ".done" && statSync(join2(pageDir, f)).isFile()).sort();
|
|
684
691
|
const names = [];
|
|
685
692
|
files.forEach((f, i) => {
|
|
@@ -689,8 +696,9 @@ function publishStaged(b) {
|
|
|
689
696
|
b.sourceMap.set(`${dir}/${f}`, `images/${name}`);
|
|
690
697
|
names.push(name);
|
|
691
698
|
});
|
|
699
|
+
if (meta.native) meta.native = { ...meta.native, file: names[files.indexOf(meta.native.file)] ?? meta.native.file };
|
|
692
700
|
rmSync(pageDir, { recursive: true, force: true });
|
|
693
|
-
out.set(page, names);
|
|
701
|
+
out.set(page, { files: names, meta });
|
|
694
702
|
}
|
|
695
703
|
return out;
|
|
696
704
|
}
|
|
@@ -969,6 +977,7 @@ function formatHandle(h) {
|
|
|
969
977
|
if (h.pagesDir && h.pageImageCount > 0) lines.push(`Pages-Dir: ${h.pagesDir} (${h.pageImageCount} pages)`);
|
|
970
978
|
else if (h.pageImagesReason) lines.push(`Pages-Dir: none - ${h.pageImagesReason}`);
|
|
971
979
|
if (h.pageStatsPath) lines.push(`Page-Stats: ${h.pageStatsPath}`);
|
|
980
|
+
if (h.nativeImages.length) lines.push(`Native-Images: ${h.nativeImages.length === 1 ? "page" : "pages"} ${compactRanges(h.nativeImages.map((n) => n.page))} (embedded image streams, no render DPI)`);
|
|
972
981
|
if (h.wordsPath) {
|
|
973
982
|
const bad = Object.keys(h.wordsErrors).map(Number).sort((a, b) => a - b);
|
|
974
983
|
lines.push(`Words: ${h.wordsPath}${bad.length ? ` (extraction failed for pages ${compactRanges(bad)}: ${h.wordsErrors[bad[0]]})` : ""}`);
|
|
@@ -1036,7 +1045,8 @@ var DOC_TO_MD_OPTIONS = [
|
|
|
1036
1045
|
{ key: "maxOutputBytes", flag: "--max-output-bytes", type: "int", default: 2e7, settable: true, help: "Child stdout cap in bytes" },
|
|
1037
1046
|
{ key: "outlineMaxEntries", flag: "--outline-max-entries", type: "int", default: 40, settable: true, help: "Heading outline / TOC / sheet inventory cap in the handle" },
|
|
1038
1047
|
{ key: "ocr", flag: "--ocr", type: "bool", default: false, settable: true, help: "Run OCR on pages without a text layer and on image inputs when Tesseract language data is installed; off by default (--no-ocr turns a settings-level true off)" },
|
|
1039
|
-
{ key: "ocrLanguage", flag: "--ocr-language", type: "lang", default: "eng", settable: true, help: "Tesseract language code(s), +-joined, e.g. deu+eng" }
|
|
1048
|
+
{ key: "ocrLanguage", flag: "--ocr-language", type: "lang", default: "eng", settable: true, help: "Tesseract language code(s), +-joined, e.g. deu+eng" },
|
|
1049
|
+
{ key: "hideAnnotations", flag: "--hide-annotations", type: "bool", default: false, settable: true, help: "Render PDF pages without annotations (sticky notes, highlights, stamps - and form-field widgets, so filled form values disappear); default paints them, as PyMuPDF does. Applies to pages/ renders and textless-page renders, not to OCR text or embedded images; also lets an annotated scan be delivered as its embedded image." }
|
|
1040
1050
|
];
|
|
1041
1051
|
var TUNABLE_DEFAULTS = Object.fromEntries(
|
|
1042
1052
|
DOC_TO_MD_OPTIONS.filter((d) => d.settable).map((d) => [d.key, d.default])
|
|
@@ -1706,6 +1716,22 @@ function handleOcr(tier, type, o, json) {
|
|
|
1706
1716
|
if (!x || type !== "image" && !x.textless.length && !x.pages.length && !x.ocrFailed.length) return null;
|
|
1707
1717
|
return { ...emptyOcr(o.ocrLanguage), ...x };
|
|
1708
1718
|
}
|
|
1719
|
+
var nativeFromChild = (b, json, notes) => (json.nativeImages ?? []).flatMap((e) => {
|
|
1720
|
+
const file = b.sourceMap.get(`p${e.page}/${e.file}`);
|
|
1721
|
+
if (!file) {
|
|
1722
|
+
notes.push(`Native image p${e.page}/${e.file} not published`);
|
|
1723
|
+
return [];
|
|
1724
|
+
}
|
|
1725
|
+
return [{ ...e, file: join3(b.root, file) }];
|
|
1726
|
+
});
|
|
1727
|
+
function retainedNative(b, kept, notes) {
|
|
1728
|
+
const out = [];
|
|
1729
|
+
for (const [page, k] of kept) {
|
|
1730
|
+
if (k.meta.native) out.push({ page, ...k.meta.native, file: join3(b.imagesDir, k.meta.native.file) });
|
|
1731
|
+
if (k.meta.dpi !== void 0 && k.meta.requestedDpi !== void 0 && k.meta.dpi < k.meta.requestedDpi) notes.push(`Page ${page} rendered at ${k.meta.dpi} dpi (requested ${k.meta.requestedDpi}; 50 Mpx ceiling)`);
|
|
1732
|
+
}
|
|
1733
|
+
return out;
|
|
1734
|
+
}
|
|
1709
1735
|
async function convertDocument(o, signal, seams) {
|
|
1710
1736
|
const s = { backend: (c) => getBackend(c, void 0, signal), runTier: runTierReal, office: tryConvertOffice, ...seams };
|
|
1711
1737
|
const inputPath = resolve2(o.path);
|
|
@@ -1732,10 +1758,11 @@ async function convertDocument(o, signal, seams) {
|
|
|
1732
1758
|
let office = null;
|
|
1733
1759
|
try {
|
|
1734
1760
|
let pdfPath = inputPath;
|
|
1735
|
-
const base = { path: inputPath, pages: o.pages, ...o.words && (type === "pdf" || type === "image") ? { words: true } : {}, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr && !forced, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
|
|
1761
|
+
const base = { path: inputPath, pages: o.pages, ...o.words && (type === "pdf" || type === "image") ? { words: true } : {}, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr && !forced, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs, hideAnnotations: o.hideAnnotations };
|
|
1736
1762
|
let tier, engine, json, degraded = null, fallbackReason = null;
|
|
1737
1763
|
let explicitBreaks = null;
|
|
1738
1764
|
let notes = [];
|
|
1765
|
+
let nativeImages = [];
|
|
1739
1766
|
let officeRoute = null;
|
|
1740
1767
|
let copyReason = null;
|
|
1741
1768
|
if (type === "html") {
|
|
@@ -1890,11 +1917,12 @@ async function convertDocument(o, signal, seams) {
|
|
|
1890
1917
|
tier = "primary";
|
|
1891
1918
|
engine = "pymupdf4llm";
|
|
1892
1919
|
json = p.json;
|
|
1920
|
+
nativeImages = nativeFromChild(b, json, notes);
|
|
1893
1921
|
if (json.pageImages?.length) publishPageImages(b, json.pageCount ?? 0);
|
|
1894
1922
|
} else if ("userError" in p) throw new Error(p.userError);
|
|
1895
1923
|
else {
|
|
1896
1924
|
if (signal?.aborted) throw new Error("aborted");
|
|
1897
|
-
const keepPages = Object.fromEntries([...kept.entries()].map(([k, v]) => [String(k), v]));
|
|
1925
|
+
const keepPages = Object.fromEntries([...kept.entries()].map(([k, v]) => [String(k), v.files]));
|
|
1898
1926
|
rmSync2(b.pagesStagingDir, { recursive: true, force: true });
|
|
1899
1927
|
const f = await s.runTier("pdf-fallback", { ...pdfBase, keepPages }, b, signal, o.fallbackTimeoutMs, backend);
|
|
1900
1928
|
publishStaged(b);
|
|
@@ -1905,6 +1933,7 @@ async function convertDocument(o, signal, seams) {
|
|
|
1905
1933
|
json = f.json;
|
|
1906
1934
|
degraded = DEGRADED_TEXT;
|
|
1907
1935
|
fallbackReason = `primary ${p.reason}`;
|
|
1936
|
+
nativeImages = [...retainedNative(b, kept, notes), ...nativeFromChild(b, json, notes)].sort((x, y) => x.page - y.page);
|
|
1908
1937
|
}
|
|
1909
1938
|
}
|
|
1910
1939
|
}
|
|
@@ -1957,7 +1986,7 @@ async function convertDocument(o, signal, seams) {
|
|
|
1957
1986
|
` : "") + body;
|
|
1958
1987
|
commitBundle(b, markdown);
|
|
1959
1988
|
const outline = scanOutline(markdown, o.outlineMaxEntries);
|
|
1960
|
-
const details = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr, pageStats, pageStatsPath: pageStats ? b.pageStatsPath : null, ocrDir: b.ocrManifest.size ? b.ocrDir : null, wordsPath, wordsReason, wordsErrors };
|
|
1989
|
+
const details = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr, pageStats, pageStatsPath: pageStats ? b.pageStatsPath : null, ocrDir: b.ocrManifest.size ? b.ocrDir : null, nativeImages, wordsPath, wordsReason, wordsErrors };
|
|
1961
1990
|
return { output: formatHandle(details), details };
|
|
1962
1991
|
} catch (e) {
|
|
1963
1992
|
abortBundle(b);
|
package/lib/doc-to-md-bundle.ts
CHANGED
|
@@ -109,9 +109,12 @@ export function openBundle(root: string, requested: string, overwrite: boolean):
|
|
|
109
109
|
} catch (e) { rmSync(lockPath, { force: true }); throw e; }
|
|
110
110
|
}
|
|
111
111
|
|
|
112
|
-
|
|
113
|
-
export
|
|
114
|
-
|
|
112
|
+
export interface DoneMeta { native?: { file: string; width: number; height: number }; dpi?: number; requestedDpi?: number }
|
|
113
|
+
export interface StagedPage { files: string[]; meta: DoneMeta }
|
|
114
|
+
|
|
115
|
+
/** Publish completed pages, retaining metadata with native filenames renamed; discard partial pages. */
|
|
116
|
+
export function publishStaged(b: Bundle): Map<number, StagedPage> {
|
|
117
|
+
const out = new Map<number, StagedPage>();
|
|
115
118
|
if (!existsSync(b.stagingDir)) return out;
|
|
116
119
|
for (const dir of readdirSync(b.stagingDir).sort()) {
|
|
117
120
|
const m = dir.match(/^p(\d+)$/);
|
|
@@ -119,6 +122,12 @@ export function publishStaged(b: Bundle): Map<number, string[]> {
|
|
|
119
122
|
const pageDir = join(b.stagingDir, dir);
|
|
120
123
|
if (!existsSync(join(pageDir, ".done"))) { rmSync(pageDir, { recursive: true, force: true }); continue; }
|
|
121
124
|
const page = Number(m[1]);
|
|
125
|
+
const raw = readFileSync(join(pageDir, ".done"), "utf8").trim();
|
|
126
|
+
let meta: DoneMeta = {};
|
|
127
|
+
try {
|
|
128
|
+
const v = raw ? JSON.parse(raw) : {};
|
|
129
|
+
if (v && typeof v === "object" && !Array.isArray(v)) meta = v;
|
|
130
|
+
} catch { /* Completed images survive truncated metadata. */ }
|
|
122
131
|
const files = readdirSync(pageDir).filter((f) => f !== ".done" && statSync(join(pageDir, f)).isFile()).sort();
|
|
123
132
|
const names: string[] = [];
|
|
124
133
|
files.forEach((f, i) => {
|
|
@@ -128,8 +137,9 @@ export function publishStaged(b: Bundle): Map<number, string[]> {
|
|
|
128
137
|
b.sourceMap.set(`${dir}/${f}`, `images/${name}`);
|
|
129
138
|
names.push(name);
|
|
130
139
|
});
|
|
140
|
+
if (meta.native) meta.native = { ...meta.native, file: names[files.indexOf(meta.native.file)] ?? meta.native.file };
|
|
131
141
|
rmSync(pageDir, { recursive: true, force: true });
|
|
132
|
-
out.set(page, names);
|
|
142
|
+
out.set(page, { files: names, meta });
|
|
133
143
|
}
|
|
134
144
|
return out;
|
|
135
145
|
}
|
package/lib/doc-to-md-core.ts
CHANGED
|
@@ -12,8 +12,8 @@ import { type ChildProcess, spawn } from "node:child_process";
|
|
|
12
12
|
import { homedir, tmpdir } from "node:os";
|
|
13
13
|
import { basename, dirname, extname, isAbsolute, join, resolve } from "node:path";
|
|
14
14
|
import { fileURLToPath, pathToFileURL } from "node:url";
|
|
15
|
-
import { type Bundle, abortBundle, commitBundle, openBundle, publishAttachments, publishSidecars, publishWords, writePageStats, publishPageImages, publishSheetCsvs, publishSheetImages, publishStaged, rewriteLinks, tempBundleRoot, validateImageLinks } from "./doc-to-md-bundle.ts";
|
|
16
|
-
import { type Engine, type HandleData, type InfoData, type OcrInfo, type PageStat, type Tier, formatHandle, formatInfoHandle, scanOutline, type SheetInfo, type TocEntry } from "./doc-to-md-handle.ts";
|
|
15
|
+
import { type StagedPage, type Bundle, abortBundle, commitBundle, openBundle, publishAttachments, publishSidecars, publishWords, writePageStats, publishPageImages, publishSheetCsvs, publishSheetImages, publishStaged, rewriteLinks, tempBundleRoot, validateImageLinks } from "./doc-to-md-bundle.ts";
|
|
16
|
+
import { type NativeImage, type Engine, type HandleData, type InfoData, type OcrInfo, type PageStat, type Tier, formatHandle, formatInfoHandle, scanOutline, type SheetInfo, type TocEntry } from "./doc-to-md-handle.ts";
|
|
17
17
|
import { type DocToMdOptions, type InputType, IMAGE_EXTS, TUNABLE_DEFAULTS, UsageError, classifyInput, sanitizeStem } from "./doc-to-md-options.ts";
|
|
18
18
|
|
|
19
19
|
export * from "./doc-to-md-options.ts";
|
|
@@ -467,7 +467,7 @@ const clearStaging = (b: Pick<Bundle, "stagingDir">) => { for (const f of readdi
|
|
|
467
467
|
export const EXCEL_REMEDY = "Remedy: install uv, or pip install openpyxl xlrd pillow";
|
|
468
468
|
|
|
469
469
|
export type Mode = "html" | "image" | "info" | "pdf-primary" | "pdf-fallback" | "xlsx" | "pdf-text" | "render-pages" | "docx" | "email" | "ocr-pages";
|
|
470
|
-
export interface TierJson { words?: boolean; wordsErrors?: Record<string, string>; pageStats?: PageStat[]; status?: string; written?: number[]; noText?: number[]; ocrFailed?: number[]; ocrErrors?: Record<string, string>; budgetStopped?: number[]; pageImages?: { page: number; file: string }[]; ocr?: OcrInfo; markdown?: string; pages?: number[]; pageCount?: number; emptyPages?: number[]; failedPages?: { page: number; error: string }[]; notes?: string[]; images?: { sheetIndex: number; file: string }[]; metadata?: Record<string, string>; toc?: [number, string, number | null][]; explicitBreaks?: number; engine?: string; degraded?: boolean; fallbackReason?: string | null; sheets?: SheetInfo[]; renderPages?: number[]; sheetCount?: number; ok?: boolean; reason?: string; rendered?: { idx: number; file: string; dpi: number }[]; failed?: { idx: number; reason: string }[]; }
|
|
470
|
+
export interface TierJson { words?: boolean; wordsErrors?: Record<string, string>; pageStats?: PageStat[]; status?: string; written?: number[]; noText?: number[]; ocrFailed?: number[]; ocrErrors?: Record<string, string>; budgetStopped?: number[]; pageImages?: { page: number; file: string; dpi?: number; requestedDpi?: number }[]; nativeImages?: NativeImage[]; ocr?: OcrInfo; markdown?: string; pages?: number[]; pageCount?: number; emptyPages?: number[]; failedPages?: { page: number; error: string }[]; notes?: string[]; images?: { sheetIndex: number; file: string }[]; metadata?: Record<string, string>; toc?: [number, string, number | null][]; explicitBreaks?: number; engine?: string; degraded?: boolean; fallbackReason?: string | null; sheets?: SheetInfo[]; renderPages?: number[]; sheetCount?: number; ok?: boolean; reason?: string; rendered?: { idx: number; file: string; dpi: number }[]; failed?: { idx: number; reason: string }[]; }
|
|
471
471
|
export type TierResult = { ok: true; json: TierJson } | { ok: false; reason: string; detail?: string } | { ok: false; userError: string; pageCount?: number };
|
|
472
472
|
|
|
473
473
|
export interface PipelineSeams {
|
|
@@ -586,6 +586,25 @@ function handleOcr(tier: Tier, type: InputType, o: DocToMdOptions, json: TierJso
|
|
|
586
586
|
return { ...emptyOcr(o.ocrLanguage), ...x };
|
|
587
587
|
}
|
|
588
588
|
|
|
589
|
+
const nativeFromChild = (b: Bundle, json: TierJson, notes: string[]): NativeImage[] => (json.nativeImages ?? []).flatMap((e) => {
|
|
590
|
+
const file = b.sourceMap.get(`p${e.page}/${e.file}`);
|
|
591
|
+
if (!file) {
|
|
592
|
+
notes.push(`Native image p${e.page}/${e.file} not published`);
|
|
593
|
+
return [];
|
|
594
|
+
}
|
|
595
|
+
return [{ ...e, file: join(b.root, file) }];
|
|
596
|
+
});
|
|
597
|
+
|
|
598
|
+
// A dead primary has no JSON response; retained pages carry their image and clamp facts in .done.
|
|
599
|
+
function retainedNative(b: Bundle, kept: Map<number, StagedPage>, notes: string[]): NativeImage[] {
|
|
600
|
+
const out: NativeImage[] = [];
|
|
601
|
+
for (const [page, k] of kept) {
|
|
602
|
+
if (k.meta.native) out.push({ page, ...k.meta.native, file: join(b.imagesDir, k.meta.native.file) });
|
|
603
|
+
if (k.meta.dpi !== undefined && k.meta.requestedDpi !== undefined && k.meta.dpi < k.meta.requestedDpi) notes.push(`Page ${page} rendered at ${k.meta.dpi} dpi (requested ${k.meta.requestedDpi}; 50 Mpx ceiling)`);
|
|
604
|
+
}
|
|
605
|
+
return out;
|
|
606
|
+
}
|
|
607
|
+
|
|
589
608
|
export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, seams?: Partial<PipelineSeams>): Promise<ConvertOutcome> {
|
|
590
609
|
const s: PipelineSeams = { backend: (c) => getBackend(c, undefined, signal), runTier: runTierReal, office: tryConvertOffice, ...seams };
|
|
591
610
|
const inputPath = resolve(o.path);
|
|
@@ -612,10 +631,11 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
|
|
|
612
631
|
let office: { pdfPath: string; cleanup: () => void } | null = null;
|
|
613
632
|
try {
|
|
614
633
|
let pdfPath = inputPath;
|
|
615
|
-
const base = { path: inputPath, pages: o.pages, ...(o.words && (type === "pdf" || type === "image") ? { words: true } : {}), stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr && !forced, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
|
|
634
|
+
const base = { path: inputPath, pages: o.pages, ...(o.words && (type === "pdf" || type === "image") ? { words: true } : {}), stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr && !forced, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs, hideAnnotations: o.hideAnnotations };
|
|
616
635
|
let tier: Tier | undefined, engine: Engine | undefined, json: TierJson | undefined, degraded: string | null = null, fallbackReason: string | null = null;
|
|
617
636
|
let explicitBreaks: number | null = null;
|
|
618
637
|
let notes: string[] = [];
|
|
638
|
+
let nativeImages: NativeImage[] = [];
|
|
619
639
|
let officeRoute: string | null = null;
|
|
620
640
|
let copyReason: string | null = null;
|
|
621
641
|
if (type === "html") {
|
|
@@ -728,17 +748,18 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
|
|
|
728
748
|
} else {
|
|
729
749
|
const p = await s.runTier("pdf-primary", pdfBase, b, signal, o.primaryTimeoutMs, backend);
|
|
730
750
|
const kept = publishStaged(b);
|
|
731
|
-
if (p.ok) { tier = "primary"; engine = "pymupdf4llm"; json = p.json; if (json.pageImages?.length) publishPageImages(b, json.pageCount ?? 0); }
|
|
751
|
+
if (p.ok) { tier = "primary"; engine = "pymupdf4llm"; json = p.json; nativeImages = nativeFromChild(b, json, notes); if (json.pageImages?.length) publishPageImages(b, json.pageCount ?? 0); }
|
|
732
752
|
else if ("userError" in p) throw new Error(p.userError);
|
|
733
753
|
else {
|
|
734
754
|
if (signal?.aborted) throw new Error("aborted");
|
|
735
|
-
const keepPages = Object.fromEntries([...kept.entries()].map(([k, v]) => [String(k), v]));
|
|
755
|
+
const keepPages = Object.fromEntries([...kept.entries()].map(([k, v]) => [String(k), v.files]));
|
|
736
756
|
rmSync(b.pagesStagingDir, { recursive: true, force: true });
|
|
737
757
|
const f = await s.runTier("pdf-fallback", { ...pdfBase, keepPages }, b, signal, o.fallbackTimeoutMs, backend);
|
|
738
758
|
publishStaged(b);
|
|
739
759
|
if (f.ok && f.json.pageImages?.length) publishPageImages(b, f.json.pageCount ?? 0);
|
|
740
760
|
if (!f.ok) throw new Error("userError" in f ? f.userError : `Conversion failed: primary ${p.reason}; fallback ${f.reason}${detailSuffix(f)}`);
|
|
741
761
|
tier = "fallback"; engine = "pymupdf-text"; json = f.json; degraded = DEGRADED_TEXT; fallbackReason = `primary ${p.reason}`;
|
|
762
|
+
nativeImages = [...retainedNative(b, kept, notes), ...nativeFromChild(b, json, notes)].sort((x, y) => x.page - y.page);
|
|
742
763
|
}
|
|
743
764
|
}
|
|
744
765
|
}
|
|
@@ -789,7 +810,7 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
|
|
|
789
810
|
const markdown = (head.length ? `${head.join("\n")}\n\n` : "") + body;
|
|
790
811
|
commitBundle(b, markdown);
|
|
791
812
|
const outline = scanOutline(markdown, o.outlineMaxEntries);
|
|
792
|
-
const details: DocToMdDetails = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr, pageStats, pageStatsPath: pageStats ? b.pageStatsPath : null, ocrDir: b.ocrManifest.size ? b.ocrDir : null, wordsPath, wordsReason, wordsErrors };
|
|
813
|
+
const details: DocToMdDetails = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr, pageStats, pageStatsPath: pageStats ? b.pageStatsPath : null, ocrDir: b.ocrManifest.size ? b.ocrDir : null, nativeImages, wordsPath, wordsReason, wordsErrors };
|
|
793
814
|
return { output: formatHandle(details), details };
|
|
794
815
|
} catch (e) { abortBundle(b); throw e; }
|
|
795
816
|
finally { office?.cleanup(); }
|
package/lib/doc-to-md-handle.ts
CHANGED
|
@@ -39,12 +39,15 @@ export interface OcrInfo {
|
|
|
39
39
|
childError: string | null;
|
|
40
40
|
}
|
|
41
41
|
|
|
42
|
+
export interface NativeImage { page: number; file: string; width: number; height: number }
|
|
43
|
+
|
|
42
44
|
export interface HandleData {
|
|
43
45
|
savedTo: string; imagesDir: string | null; sheetsDir: string | null; pagesDir: string | null; type: InputType; engine: Engine; tier: Tier;
|
|
44
46
|
pageCount: number | null; pages: number[] | null; explicitBreaks: number | null; imageCount: number; pageImageCount: number; pageImagesReason: string | null; bytes: number; lines: number;
|
|
45
47
|
degraded: string | null; fallbackReason: string | null; failedPages: number[]; emptyPages: number[];
|
|
46
48
|
notes: string[]; outline: OutlineEntry[]; outlineTotal: number; ocr: OcrInfo | null;
|
|
47
49
|
pageStats: PageStat[] | null; pageStatsPath: string | null; ocrDir: string | null;
|
|
50
|
+
nativeImages: NativeImage[];
|
|
48
51
|
wordsPath: string | null; wordsReason: string | null; wordsErrors: Record<number, string>;
|
|
49
52
|
}
|
|
50
53
|
|
|
@@ -171,6 +174,7 @@ export function formatHandle(h: HandleData): string {
|
|
|
171
174
|
if (h.pagesDir && h.pageImageCount > 0) lines.push(`Pages-Dir: ${h.pagesDir} (${h.pageImageCount} pages)`);
|
|
172
175
|
else if (h.pageImagesReason) lines.push(`Pages-Dir: none - ${h.pageImagesReason}`);
|
|
173
176
|
if (h.pageStatsPath) lines.push(`Page-Stats: ${h.pageStatsPath}`);
|
|
177
|
+
if (h.nativeImages.length) lines.push(`Native-Images: ${h.nativeImages.length === 1 ? "page" : "pages"} ${compactRanges(h.nativeImages.map((n) => n.page))} (embedded image streams, no render DPI)`);
|
|
174
178
|
if (h.wordsPath) {
|
|
175
179
|
const bad = Object.keys(h.wordsErrors).map(Number).sort((a, b) => a - b);
|
|
176
180
|
lines.push(`Words: ${h.wordsPath}${bad.length ? ` (extraction failed for pages ${compactRanges(bad)}: ${h.wordsErrors[bad[0]]})` : ""}`);
|
package/lib/doc-to-md-options.ts
CHANGED
|
@@ -21,6 +21,7 @@ export interface Tunables {
|
|
|
21
21
|
outlineMaxEntries: number;
|
|
22
22
|
ocr: boolean;
|
|
23
23
|
ocrLanguage: string;
|
|
24
|
+
hideAnnotations: boolean;
|
|
24
25
|
}
|
|
25
26
|
|
|
26
27
|
export interface DocToMdOptions extends Tunables {
|
|
@@ -87,6 +88,7 @@ export const DOC_TO_MD_OPTIONS: readonly OptionDescriptor[] = [
|
|
|
87
88
|
{ key: "outlineMaxEntries", flag: "--outline-max-entries", type: "int", default: 40, settable: true, help: "Heading outline / TOC / sheet inventory cap in the handle" },
|
|
88
89
|
{ key: "ocr", flag: "--ocr", type: "bool", default: false, settable: true, help: "Run OCR on pages without a text layer and on image inputs when Tesseract language data is installed; off by default (--no-ocr turns a settings-level true off)" },
|
|
89
90
|
{ key: "ocrLanguage", flag: "--ocr-language", type: "lang", default: "eng", settable: true, help: "Tesseract language code(s), +-joined, e.g. deu+eng" },
|
|
91
|
+
{ key: "hideAnnotations", flag: "--hide-annotations", type: "bool", default: false, settable: true, help: "Render PDF pages without annotations (sticky notes, highlights, stamps - and form-field widgets, so filled form values disappear); default paints them, as PyMuPDF does. Applies to pages/ renders and textless-page renders, not to OCR text or embedded images; also lets an annotated scan be delivered as its embedded image." },
|
|
90
92
|
];
|
|
91
93
|
|
|
92
94
|
export const TUNABLE_DEFAULTS: Tunables = Object.fromEntries(
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-quiver",
|
|
3
|
-
"version": "6.
|
|
3
|
+
"version": "6.11.0",
|
|
4
4
|
"description": "Personal pack of Pi coding-agent extensions: context-safe fetch, doc_to_md PDF/DOCX/PPTX-to-Markdown conversion, session naming, a themed ASCII startup header, Opus 4.8 fast mode, and a provider-stall watchdog.",
|
|
5
5
|
"author": "Jacek Juraszek",
|
|
6
6
|
"license": "MIT",
|
package/scripts/doc_to_md.py
CHANGED
|
@@ -90,8 +90,9 @@ def page_dir(staging, n):
|
|
|
90
90
|
return d
|
|
91
91
|
|
|
92
92
|
|
|
93
|
-
def mark_done(d):
|
|
94
|
-
open(os.path.join(d, ".done"), "w")
|
|
93
|
+
def mark_done(d, meta=None):
|
|
94
|
+
with open(os.path.join(d, ".done"), "w", encoding="utf-8") as fh:
|
|
95
|
+
json.dump(meta or {}, fh)
|
|
95
96
|
|
|
96
97
|
|
|
97
98
|
OCR_SENTINEL = "\x00OCR {}\x00"
|
|
@@ -150,16 +151,232 @@ def clamped_dpi(w, h, dpi):
|
|
|
150
151
|
return eff if eff >= MIN_RENDER_DPI else None
|
|
151
152
|
|
|
152
153
|
|
|
153
|
-
|
|
154
|
+
RASTER_BUDGET_S, OCR_JOB_BUDGET_S, OCR_DPI = 20, 30, 300
|
|
155
|
+
RASTER_SPAWN_BUDGET_S = 30
|
|
156
|
+
MIN_OCR_JOB_S = 1
|
|
157
|
+
NATIVE_MIN_COVERAGE, NATIVE_MAX_COVERAGE = 0.9, 1.1
|
|
158
|
+
NATIVE_EXTS = ("png", "jpeg", "jpg")
|
|
159
|
+
JOB_VERB = {"native": "render", "render": "render", "images": "render", "ocr": "OCR"}
|
|
160
|
+
RASTER_INLINE = False # tests only: run raster jobs in-process so monkeypatches reach them
|
|
161
|
+
|
|
162
|
+
|
|
163
|
+
class RasterError(Exception):
|
|
164
|
+
"""A raster job did not complete; str(exc) is the failedPages message."""
|
|
165
|
+
|
|
166
|
+
|
|
167
|
+
class RasterUnavailable(Exception):
|
|
168
|
+
pass
|
|
169
|
+
|
|
170
|
+
|
|
171
|
+
def raster_budget(default):
|
|
172
|
+
return float(os.environ.get("DOC_TO_MD_RASTER_BUDGET_S", default))
|
|
173
|
+
|
|
174
|
+
|
|
175
|
+
def clamp_note(n, eff, requested):
|
|
176
|
+
return f"Page {n} rendered at {eff} dpi (requested {requested}; {MAX_RENDER_PX // 1_000_000} Mpx ceiling)"
|
|
177
|
+
|
|
178
|
+
|
|
179
|
+
def native_page_image(doc, page, target_dir, allow_annots=False):
|
|
180
|
+
"""One full-page, unrotated, unmasked PNG/JPEG stream with no vector overlay -> written untouched as page.<ext>."""
|
|
181
|
+
import pymupdf
|
|
182
|
+
if not allow_annots and (page.first_annot is not None or page.first_widget is not None):
|
|
183
|
+
return None
|
|
184
|
+
infos = page.get_image_info(xrefs=True)
|
|
185
|
+
if len(infos) != 1 or infos[0].get("xref", 0) <= 0 or page.rotation != 0:
|
|
186
|
+
return None
|
|
187
|
+
drawings = page.get_drawings()
|
|
188
|
+
if allow_annots:
|
|
189
|
+
# MuPDF includes appearance streams, clipped to annotation/widget rectangles.
|
|
190
|
+
annot_rects = [a.rect + (-1, -1, 1, 1) for items in (page.annots(), page.widgets()) for a in items]
|
|
191
|
+
drawings = [drawing for drawing in drawings if not any(drawing["rect"] in rect for rect in annot_rects)]
|
|
192
|
+
if drawings:
|
|
193
|
+
return None
|
|
194
|
+
info = infos[0]
|
|
195
|
+
a, b, c, d = info["transform"][:4]
|
|
196
|
+
if b != 0 or c != 0 or a <= 0 or d <= 0:
|
|
197
|
+
return None
|
|
198
|
+
area, bbox = page.rect.get_area(), pymupdf.Rect(info["bbox"])
|
|
199
|
+
if area <= 0 or (bbox & page.rect).get_area() / area < NATIVE_MIN_COVERAGE or bbox.get_area() / area > NATIVE_MAX_COVERAGE:
|
|
200
|
+
return None
|
|
201
|
+
img = doc.extract_image(info["xref"])
|
|
202
|
+
ext, cs = img["ext"].lower(), img.get("cs-name")
|
|
203
|
+
plain = cs in ("DeviceRGB", "DeviceGray") or (cs == "DeviceCMYK" and ext in ("jpeg", "jpg"))
|
|
204
|
+
if img.get("smask", 0) or ext not in NATIVE_EXTS or not img["image"] or not plain:
|
|
205
|
+
return None
|
|
206
|
+
name = f"page.{ext}"
|
|
207
|
+
with open(os.path.join(target_dir, name), "wb") as fh:
|
|
208
|
+
fh.write(img["image"])
|
|
209
|
+
return {"file": name, "width": img["width"], "height": img["height"]}
|
|
210
|
+
|
|
211
|
+
|
|
212
|
+
def run_raster_job(doc, job):
|
|
213
|
+
import time
|
|
214
|
+
import pymupdf
|
|
215
|
+
n, kind = job["page"], job["kind"]
|
|
216
|
+
try:
|
|
217
|
+
if os.environ.get("DOC_TO_MD_RASTER_STALL_PAGE") in (str(n), f"{kind}:{n}"): # tests only: a wedged page
|
|
218
|
+
time.sleep(3600)
|
|
219
|
+
if os.environ.get("DOC_TO_MD_RASTER_CRASH_PAGE") == str(n): # tests only: partial output, stdout chatter, hard exit
|
|
220
|
+
if kind == "render":
|
|
221
|
+
with open(job["target"], "wb") as fh:
|
|
222
|
+
fh.write(b"\x89PNG partial")
|
|
223
|
+
print("raster worker crash hook", flush=True)
|
|
224
|
+
os._exit(1)
|
|
225
|
+
page = doc[n - 1]
|
|
226
|
+
if kind == "native":
|
|
227
|
+
r = native_page_image(doc, page, job["target"], job.get("allowAnnots", False))
|
|
228
|
+
return {"ok": True, "eligible": r is not None, **(r or {})}
|
|
229
|
+
if kind == "render":
|
|
230
|
+
kw = {"dpi": job["dpi"], "annots": job["annots"]}
|
|
231
|
+
if job.get("clip") is not None:
|
|
232
|
+
kw["clip"] = pymupdf.Rect(job["clip"])
|
|
233
|
+
pix = page.get_pixmap(**kw)
|
|
234
|
+
pix.save(job["target"])
|
|
235
|
+
return {"ok": True, "width": pix.width, "height": pix.height}
|
|
236
|
+
if kind == "images":
|
|
237
|
+
files = []
|
|
238
|
+
for i, image in enumerate(page.get_image_info(xrefs=True), 1):
|
|
239
|
+
xref = image.get("xref", 0)
|
|
240
|
+
if xref > 0:
|
|
241
|
+
img = doc.extract_image(xref)
|
|
242
|
+
name = f"img{i}.{img['ext'].lower()}"
|
|
243
|
+
with open(os.path.join(job["target"], name), "wb") as fh:
|
|
244
|
+
fh.write(img["image"])
|
|
245
|
+
else:
|
|
246
|
+
if job["dpi"] is None:
|
|
247
|
+
continue
|
|
248
|
+
name = f"img{i}.{job['format']}"
|
|
249
|
+
pix = page.get_pixmap(clip=pymupdf.Rect(image["bbox"]), dpi=job["dpi"], annots=job["annots"])
|
|
250
|
+
pix.save(os.path.join(job["target"], name))
|
|
251
|
+
files.append(name)
|
|
252
|
+
return {"ok": True, "files": files}
|
|
253
|
+
import pymupdf4llm
|
|
254
|
+
md = pymupdf4llm.to_markdown(doc, pages=[n - 1], write_images=False, page_separators=False, use_ocr=True, force_ocr=True,
|
|
255
|
+
ocr_language=job["lang"], ocr_dpi=job["ocrDpi"])
|
|
256
|
+
result = {"ok": True, "markdown": md}
|
|
257
|
+
if job.get("wordsSnapshot"):
|
|
258
|
+
# OCR mutates only the worker's document; carry that text layer back for parent-side word extraction.
|
|
259
|
+
try:
|
|
260
|
+
snapshot = pymupdf.open()
|
|
261
|
+
snapshot.insert_pdf(doc, from_page=n - 1, to_page=n - 1)
|
|
262
|
+
snapshot.save(job["wordsSnapshot"])
|
|
263
|
+
snapshot.close()
|
|
264
|
+
except Exception as exc: # noqa: BLE001 - geometry never changes the OCR outcome
|
|
265
|
+
result["wordsSnapshotError"] = words_error(exc)
|
|
266
|
+
return result
|
|
267
|
+
except Exception as exc: # noqa: BLE001 - the parent decides what a failed job costs
|
|
268
|
+
return {"ok": False, "error": f"{type(exc).__name__}: {exc}"[:300]}
|
|
269
|
+
|
|
270
|
+
|
|
271
|
+
def redirect_worker_stdout():
|
|
272
|
+
# redirect_stdout in the parent is a Python-object swap the spawned process does not inherit; pymupdf4llm prints to stdout.
|
|
273
|
+
try:
|
|
274
|
+
os.dup2(sys.stderr.fileno(), sys.stdout.fileno())
|
|
275
|
+
except (AttributeError, OSError, ValueError):
|
|
276
|
+
sys.stdout = sys.stderr
|
|
277
|
+
|
|
278
|
+
|
|
279
|
+
def raster_worker(conn, path):
|
|
280
|
+
redirect_worker_stdout()
|
|
281
|
+
import pymupdf
|
|
282
|
+
doc = pymupdf.open(path)
|
|
283
|
+
conn.send({"ready": True})
|
|
284
|
+
while True:
|
|
285
|
+
job = conn.recv()
|
|
286
|
+
if job["kind"] == "stop":
|
|
287
|
+
break
|
|
288
|
+
conn.send(run_raster_job(doc, job))
|
|
289
|
+
doc.close()
|
|
290
|
+
|
|
291
|
+
|
|
292
|
+
class RasterWorker:
|
|
293
|
+
"""One spawned PyMuPDF process per conversion. MuPDF spins at C level, so only an OS kill ends a toxic page; a job that overruns its budget or takes the process down costs that job, and the next job respawns."""
|
|
294
|
+
|
|
295
|
+
def __init__(self, path, doc):
|
|
296
|
+
self.path, self.doc, self.proc, self.conn = path, doc, None, None
|
|
297
|
+
|
|
298
|
+
def run(self, job, budget_s):
|
|
299
|
+
if RASTER_INLINE:
|
|
300
|
+
return run_raster_job(self.doc, job)
|
|
301
|
+
if self.proc is None:
|
|
302
|
+
try:
|
|
303
|
+
import multiprocessing
|
|
304
|
+
ctx = multiprocessing.get_context("spawn")
|
|
305
|
+
self.conn, child = ctx.Pipe()
|
|
306
|
+
self.proc = ctx.Process(target=raster_worker, args=(child, self.path))
|
|
307
|
+
try:
|
|
308
|
+
self.proc.start()
|
|
309
|
+
finally:
|
|
310
|
+
child.close()
|
|
311
|
+
except Exception as exc:
|
|
312
|
+
if self.conn is not None:
|
|
313
|
+
self.conn.close()
|
|
314
|
+
self.proc = self.conn = None
|
|
315
|
+
raise RasterUnavailable(f"raster worker unavailable: {exc}") from exc
|
|
316
|
+
try:
|
|
317
|
+
ready = self.conn.poll(RASTER_SPAWN_BUDGET_S) and self.conn.recv() == {"ready": True}
|
|
318
|
+
except (EOFError, OSError):
|
|
319
|
+
ready = False
|
|
320
|
+
if not ready:
|
|
321
|
+
self.kill()
|
|
322
|
+
return {"ok": False, "error": "renderer failed to start"}
|
|
323
|
+
try:
|
|
324
|
+
self.conn.send(job)
|
|
325
|
+
if self.conn.poll(budget_s):
|
|
326
|
+
return self.conn.recv()
|
|
327
|
+
error = f"{JOB_VERB[job['kind']]} timed out after {budget_s:g}s"
|
|
328
|
+
except (EOFError, OSError):
|
|
329
|
+
error = "renderer crashed"
|
|
330
|
+
self.kill()
|
|
331
|
+
if job["kind"] == "render":
|
|
332
|
+
with contextlib.suppress(OSError):
|
|
333
|
+
os.remove(job["target"])
|
|
334
|
+
if job["kind"] == "ocr" and job.get("wordsSnapshot"):
|
|
335
|
+
with contextlib.suppress(OSError):
|
|
336
|
+
os.remove(job["wordsSnapshot"])
|
|
337
|
+
return {"ok": False, "error": error}
|
|
338
|
+
|
|
339
|
+
def kill(self):
|
|
340
|
+
if self.proc is not None:
|
|
341
|
+
if self.proc.is_alive():
|
|
342
|
+
self.proc.kill()
|
|
343
|
+
self.proc.join()
|
|
344
|
+
self.conn.close()
|
|
345
|
+
self.proc = self.conn = None
|
|
346
|
+
|
|
347
|
+
def stop(self):
|
|
348
|
+
if self.proc is None:
|
|
349
|
+
return
|
|
350
|
+
with contextlib.suppress(OSError):
|
|
351
|
+
self.conn.send({"kind": "stop"})
|
|
352
|
+
self.proc.join(2)
|
|
353
|
+
self.kill()
|
|
354
|
+
|
|
355
|
+
|
|
356
|
+
def textless_picture(worker, page, n, d, o, notes, render):
|
|
357
|
+
"""Native stream when the page is one full-page image, else (when render) a worker render at the clamped DPI. Returns (file name or None, .done metadata)."""
|
|
358
|
+
r = worker.run({"kind": "native", "page": n, "target": d, "allowAnnots": bool(o.get("hideAnnotations"))}, raster_budget(RASTER_BUDGET_S))
|
|
359
|
+
if r.get("ok") and r.get("eligible"):
|
|
360
|
+
return r["file"], {"native": {"file": r["file"], "width": r["width"], "height": r["height"]}}
|
|
361
|
+
for f in os.listdir(d):
|
|
362
|
+
if f.startswith("page."): # a failed native job may have left a partial stream
|
|
363
|
+
os.remove(os.path.join(d, f))
|
|
364
|
+
if not render:
|
|
365
|
+
return None, {}
|
|
154
366
|
eff = clamped_dpi(page.rect.width, page.rect.height, o["imageDpi"])
|
|
155
367
|
if eff is None:
|
|
156
|
-
return None
|
|
368
|
+
return None, {}
|
|
157
369
|
name = f"page.{o['imageFormat']}"
|
|
158
|
-
|
|
159
|
-
|
|
370
|
+
r = worker.run({"kind": "render", "page": n, "dpi": eff, "annots": not o.get("hideAnnotations"), "target": os.path.join(d, name)}, raster_budget(RASTER_BUDGET_S))
|
|
371
|
+
if not r["ok"]:
|
|
372
|
+
raise RasterError(r["error"])
|
|
373
|
+
if eff < o["imageDpi"]:
|
|
374
|
+
notes.append(clamp_note(n, eff, o["imageDpi"]))
|
|
375
|
+
return name, {"dpi": eff, "requestedDpi": o["imageDpi"]}
|
|
376
|
+
return name, {}
|
|
160
377
|
|
|
161
378
|
|
|
162
|
-
def render_page_image(page, n, o, page_images):
|
|
379
|
+
def render_page_image(worker, page, n, o, page_images, notes):
|
|
163
380
|
d = o["pagesStagingDir"]
|
|
164
381
|
eff = clamped_dpi(page.rect.width, page.rect.height, o["imageDpi"])
|
|
165
382
|
if eff is None:
|
|
@@ -167,18 +384,24 @@ def render_page_image(page, n, o, page_images):
|
|
|
167
384
|
os.makedirs(d, exist_ok=True)
|
|
168
385
|
name = f"p{n}.{o['imageFormat']}"
|
|
169
386
|
target = os.path.join(d, name)
|
|
170
|
-
|
|
171
|
-
|
|
172
|
-
|
|
173
|
-
|
|
387
|
+
r = worker.run({"kind": "render", "page": n, "dpi": eff, "annots": not o.get("hideAnnotations"), "target": target}, raster_budget(RASTER_BUDGET_S))
|
|
388
|
+
if not r["ok"]:
|
|
389
|
+
notes.append(f"Page {n} render unavailable: {r['error']}")
|
|
390
|
+
with contextlib.suppress(OSError):
|
|
174
391
|
os.remove(target)
|
|
175
|
-
except OSError:
|
|
176
|
-
pass
|
|
177
392
|
return None
|
|
178
|
-
|
|
393
|
+
entry = {"page": n, "file": name, "dpi": eff}
|
|
394
|
+
if eff < o["imageDpi"]:
|
|
395
|
+
entry["requestedDpi"] = o["imageDpi"]
|
|
396
|
+
notes.append(clamp_note(n, eff, o["imageDpi"]))
|
|
397
|
+
page_images.append(entry)
|
|
179
398
|
return name
|
|
180
399
|
|
|
181
400
|
|
|
401
|
+
def page_failure(n, exc):
|
|
402
|
+
return {"page": n, "error": str(exc) if isinstance(exc, RasterError) else f"{type(exc).__name__}: {exc}"[:300]}
|
|
403
|
+
|
|
404
|
+
|
|
182
405
|
def page_ocr_kwargs(textless, lang):
|
|
183
406
|
if textless:
|
|
184
407
|
return {"use_ocr": True, "force_ocr": True, "ocr_language": lang}
|
|
@@ -359,72 +582,113 @@ def mode_pdf_primary(o):
|
|
|
359
582
|
pages = check_pages(o.get("pages"), doc.page_count)
|
|
360
583
|
staging, out, empty, failed, notes = o["stagingDir"], [], [], [], []
|
|
361
584
|
page_images = []
|
|
585
|
+
native_images = []
|
|
586
|
+
worker = RasterWorker(o["path"], doc)
|
|
362
587
|
stats = []
|
|
363
588
|
lang, budget = o.get("ocrLanguage", "eng"), o.get("ocrBudgetMs", 60000)
|
|
364
589
|
info = new_ocr(lang)
|
|
365
590
|
status = ocr_status(True, lang) if o.get("ocr") else None
|
|
366
591
|
ocr_ms, plain_ms = [], []
|
|
367
592
|
word_pages, words_errors = [], {}
|
|
368
|
-
|
|
369
|
-
|
|
370
|
-
|
|
371
|
-
|
|
372
|
-
page = doc[n - 1]
|
|
373
|
-
rotation = page.rotation
|
|
374
|
-
textless = not page.get_text("text").strip()
|
|
375
|
-
if textless:
|
|
376
|
-
info["textless"].append(n)
|
|
377
|
-
if status is None:
|
|
378
|
-
status = ocr_status(False, lang)
|
|
379
|
-
kw = page_ocr_kwargs(textless, lang) if status and status["status"] == "ready" else {"use_ocr": False}
|
|
380
|
-
if kw["use_ocr"]:
|
|
381
|
-
elapsed = (time.monotonic() - start) * 1000
|
|
382
|
-
est_ocr = max(ocr_ms) if ocr_ms else OCR_EST_INITIAL_MS
|
|
383
|
-
est_page = sum(plain_ms) / len(plain_ms) if plain_ms else PAGE_EST_INITIAL_MS
|
|
384
|
-
if not ocr_admit(elapsed, est_ocr, len(pages) - i - 1, est_page, budget):
|
|
385
|
-
info["budgetStopped"].append(n)
|
|
386
|
-
kw = {"use_ocr": False}
|
|
387
|
-
t0 = time.monotonic()
|
|
593
|
+
try:
|
|
594
|
+
for i, n in enumerate(pages):
|
|
595
|
+
stats.append(page_stats(doc, n))
|
|
596
|
+
d = page_dir(staging, n)
|
|
388
597
|
try:
|
|
389
|
-
|
|
390
|
-
|
|
391
|
-
|
|
392
|
-
|
|
393
|
-
|
|
394
|
-
|
|
395
|
-
|
|
396
|
-
|
|
397
|
-
|
|
398
|
-
|
|
399
|
-
|
|
400
|
-
|
|
401
|
-
|
|
598
|
+
page = doc[n - 1]
|
|
599
|
+
rotation = page.rotation
|
|
600
|
+
textless = not page.get_text("text").strip()
|
|
601
|
+
if textless:
|
|
602
|
+
info["textless"].append(n)
|
|
603
|
+
if status is None:
|
|
604
|
+
status = ocr_status(False, lang)
|
|
605
|
+
kw = page_ocr_kwargs(textless, lang) if status and status["status"] == "ready" else {"use_ocr": False}
|
|
606
|
+
if kw["use_ocr"]:
|
|
607
|
+
elapsed = (time.monotonic() - start) * 1000
|
|
608
|
+
est_ocr = max(ocr_ms) if ocr_ms else OCR_EST_INITIAL_MS
|
|
609
|
+
est_page = sum(plain_ms) / len(plain_ms) if plain_ms else PAGE_EST_INITIAL_MS
|
|
610
|
+
if not ocr_admit(elapsed, est_ocr, len(pages) - i - 1, est_page, budget):
|
|
611
|
+
info["budgetStopped"].append(n)
|
|
612
|
+
kw = {"use_ocr": False}
|
|
613
|
+
text, meta = "", {}
|
|
614
|
+
ocr_words = None
|
|
615
|
+
if textless:
|
|
616
|
+
pic, meta = textless_picture(worker, page, n, d, o, notes, render=not o.get("pageImages"))
|
|
617
|
+
t0 = time.monotonic()
|
|
402
618
|
if kw["use_ocr"]:
|
|
403
|
-
|
|
619
|
+
remaining = (budget - (time.monotonic() - start) * 1000 - OCR_BUDGET_RESERVE_MS) / 1000
|
|
620
|
+
if remaining < MIN_OCR_JOB_S:
|
|
621
|
+
info["budgetStopped"].append(n)
|
|
622
|
+
kw = {"use_ocr": False}
|
|
623
|
+
else:
|
|
624
|
+
ocr_dpi = clamped_dpi(page.rect.width, page.rect.height, OCR_DPI)
|
|
625
|
+
words_snapshot = os.path.join(staging, f".ocr-words-p{n}.pdf") if o.get("words") else None
|
|
626
|
+
try:
|
|
627
|
+
r = worker.run({"kind": "ocr", "page": n, "lang": lang, "ocrDpi": ocr_dpi, **({"wordsSnapshot": words_snapshot} if words_snapshot else {})}, min(raster_budget(OCR_JOB_BUDGET_S), remaining)) if ocr_dpi is not None else {"ok": False}
|
|
628
|
+
if words_snapshot and r["ok"]:
|
|
629
|
+
if r.get("wordsSnapshotError"):
|
|
630
|
+
words_errors[str(n)] = r["wordsSnapshotError"]
|
|
631
|
+
else:
|
|
632
|
+
try:
|
|
633
|
+
with open_pdf(words_snapshot) as snapshot:
|
|
634
|
+
ocr_words = words_page(snapshot[0], n, rotation, page_words(snapshot[0], True))
|
|
635
|
+
except Exception as exc: # noqa: BLE001
|
|
636
|
+
words_errors[str(n)] = words_error(exc)
|
|
637
|
+
finally:
|
|
638
|
+
if words_snapshot:
|
|
639
|
+
with contextlib.suppress(OSError):
|
|
640
|
+
os.remove(words_snapshot)
|
|
641
|
+
if r["ok"]:
|
|
642
|
+
text = r["markdown"].strip()
|
|
643
|
+
else:
|
|
644
|
+
kw = {"use_ocr": False}
|
|
645
|
+
info["ocrFailed"].append(n)
|
|
646
|
+
md = ""
|
|
647
|
+
else:
|
|
648
|
+
t0 = time.monotonic()
|
|
649
|
+
md = primary_page_markdown(doc, n, d, o, kw, True)
|
|
650
|
+
if not (textless and not kw["use_ocr"]):
|
|
651
|
+
(ocr_ms if kw["use_ocr"] else plain_ms).append((time.monotonic() - t0) * 1000)
|
|
652
|
+
if textless:
|
|
653
|
+
if kw["use_ocr"] and text:
|
|
654
|
+
info["pages"].append(n)
|
|
655
|
+
else:
|
|
656
|
+
if kw["use_ocr"]:
|
|
657
|
+
info["noText"].append(n)
|
|
658
|
+
empty.append(n)
|
|
659
|
+
elif not md.strip():
|
|
404
660
|
empty.append(n)
|
|
405
|
-
|
|
661
|
+
mark_done(d, meta)
|
|
662
|
+
page_pic = render_page_image(worker, page, n, o, page_images, notes) if o.get("pageImages") else None
|
|
663
|
+
if textless:
|
|
664
|
+
parts = [f""] if pic else []
|
|
665
|
+
if kw["use_ocr"] and text:
|
|
666
|
+
parts.append(ocr_block(f"p{n}/{pic}" if pic else f"pages/{page_pic}" if page_pic else "-", text))
|
|
667
|
+
md = "\n\n".join(parts)
|
|
668
|
+
if page_pic:
|
|
669
|
+
md = "\n\n".join(x for x in [md.rstrip(), f""] if x)
|
|
670
|
+
out.append(md.rstrip())
|
|
671
|
+
if o.get("words") and not (textless and o.get("ocr") and not kw["use_ocr"]):
|
|
672
|
+
try:
|
|
673
|
+
if textless and kw["use_ocr"]:
|
|
674
|
+
if ocr_words is not None:
|
|
675
|
+
word_pages.append(ocr_words)
|
|
676
|
+
else:
|
|
677
|
+
word_pages.append(words_page(page, n, rotation, page_words(page, kw["use_ocr"])))
|
|
678
|
+
except Exception as exc: # noqa: BLE001
|
|
679
|
+
words_errors[str(n)] = words_error(exc)
|
|
680
|
+
if meta.get("native"):
|
|
681
|
+
native_images.append({"page": n, **meta["native"]})
|
|
682
|
+
except RasterUnavailable:
|
|
683
|
+
raise
|
|
684
|
+
except Exception as exc: # noqa: BLE001
|
|
685
|
+
shutil.rmtree(d, ignore_errors=True)
|
|
686
|
+
failed.append(page_failure(n, exc))
|
|
406
687
|
empty.append(n)
|
|
407
|
-
|
|
408
|
-
|
|
409
|
-
|
|
410
|
-
|
|
411
|
-
if kw["use_ocr"] and text:
|
|
412
|
-
parts.append(ocr_block(f"pages/{page_pic}" if page_pic else f"p{n}/{pic}" if pic else "-", text))
|
|
413
|
-
md = "\n\n".join(parts)
|
|
414
|
-
if page_pic:
|
|
415
|
-
md = "\n\n".join(x for x in [md.rstrip(), f""] if x)
|
|
416
|
-
out.append(md.rstrip())
|
|
417
|
-
if o.get("words") and not (textless and o.get("ocr") and not kw["use_ocr"]):
|
|
418
|
-
try:
|
|
419
|
-
word_pages.append(words_page(page, n, rotation, page_words(page, kw["use_ocr"])))
|
|
420
|
-
except Exception as exc: # noqa: BLE001
|
|
421
|
-
words_errors[str(n)] = words_error(exc)
|
|
422
|
-
except Exception as exc: # noqa: BLE001
|
|
423
|
-
shutil.rmtree(d, ignore_errors=True)
|
|
424
|
-
failed.append({"page": n, "error": f"{type(exc).__name__}: {exc}"[:300]})
|
|
425
|
-
empty.append(n)
|
|
426
|
-
out.append("")
|
|
427
|
-
out.append(SEP.format(n=n).strip("\n"))
|
|
688
|
+
out.append("")
|
|
689
|
+
out.append(SEP.format(n=n).strip("\n"))
|
|
690
|
+
finally:
|
|
691
|
+
worker.stop()
|
|
428
692
|
if failed and len(failed) == len(pages):
|
|
429
693
|
raise RuntimeError("every selected page failed: " + failed[0]["error"])
|
|
430
694
|
if status is not None:
|
|
@@ -433,7 +697,7 @@ def mode_pdf_primary(o):
|
|
|
433
697
|
if missing:
|
|
434
698
|
notes.append(f"Page images: {missing} of {len(pages)} unavailable")
|
|
435
699
|
result = {"markdown": "\n\n".join(out) + "\n", "pages": pages, "pageCount": doc.page_count,
|
|
436
|
-
"emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": info, "pageImages": page_images, "pageStats": stats}
|
|
700
|
+
"emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": info, "pageImages": page_images, "nativeImages": native_images, "pageStats": stats}
|
|
437
701
|
if o.get("words"):
|
|
438
702
|
result.update(write_words(staging, "pt", word_pages, words_errors))
|
|
439
703
|
return result
|
|
@@ -446,72 +710,75 @@ def mode_pdf_fallback(o):
|
|
|
446
710
|
keep = {int(k): v for k, v in (o.get("keepPages") or {}).items()}
|
|
447
711
|
staging, out, empty, failed = o["stagingDir"], [], [], []
|
|
448
712
|
page_images = []
|
|
713
|
+
native_images = []
|
|
714
|
+
worker = RasterWorker(o["path"], doc)
|
|
449
715
|
stats = []
|
|
450
716
|
lang = o.get("ocrLanguage", "eng")
|
|
451
717
|
ocr_info = new_ocr(lang)
|
|
452
718
|
word_pages, words_errors = [], {}
|
|
453
719
|
status = {"status": "unavailable", "reason": "fallback tier", "tesseract": None} if o.get("ocr") else None
|
|
454
|
-
|
|
455
|
-
|
|
456
|
-
|
|
457
|
-
|
|
458
|
-
page_pic = None
|
|
459
|
-
page_ok = False
|
|
460
|
-
try:
|
|
461
|
-
page = doc[n - 1]
|
|
462
|
-
rotation = page.rotation
|
|
463
|
-
text = page.get_text("text").strip()
|
|
464
|
-
if not text:
|
|
465
|
-
ocr_info["textless"].append(n)
|
|
466
|
-
if status is None:
|
|
467
|
-
status = ocr_status(False, lang)
|
|
468
|
-
if n not in keep:
|
|
469
|
-
d = page_dir(staging, n)
|
|
470
|
-
if not text and not o.get("pageImages"):
|
|
471
|
-
pic = render_textless_page(page, d, o)
|
|
472
|
-
if pic:
|
|
473
|
-
links.append(f"")
|
|
474
|
-
elif text:
|
|
475
|
-
i = 0
|
|
476
|
-
for image in page.get_image_info(xrefs=True):
|
|
477
|
-
i += 1
|
|
478
|
-
xref = image.get("xref", 0)
|
|
479
|
-
if xref > 0:
|
|
480
|
-
img = doc.extract_image(xref)
|
|
481
|
-
name = f"img{i}.{img['ext'].lower()}"
|
|
482
|
-
with open(os.path.join(d, name), "wb") as fh:
|
|
483
|
-
fh.write(img["image"])
|
|
484
|
-
else:
|
|
485
|
-
name = f"img{i}.{o['imageFormat']}"
|
|
486
|
-
page.get_pixmap(clip=pymupdf.Rect(image["bbox"]), dpi=o["imageDpi"]).save(os.path.join(d, name))
|
|
487
|
-
links.append(f"")
|
|
488
|
-
mark_done(d)
|
|
489
|
-
page_pic = render_page_image(page, n, o, page_images) if o.get("pageImages") else None
|
|
490
|
-
page_ok = True
|
|
491
|
-
except Exception as exc: # noqa: BLE001
|
|
492
|
-
shutil.rmtree(os.path.join(staging, f"p{n}"), ignore_errors=True)
|
|
493
|
-
text = ""
|
|
720
|
+
notes = [DEGRADED_NOTE]
|
|
721
|
+
try:
|
|
722
|
+
for n in pages:
|
|
723
|
+
stats.append(page_stats(doc, n))
|
|
494
724
|
links = [f"" for f in keep.get(n, [])]
|
|
495
|
-
|
|
496
|
-
|
|
725
|
+
text = ""
|
|
726
|
+
page_pic = None
|
|
727
|
+
meta = {}
|
|
728
|
+
page_ok = False
|
|
497
729
|
try:
|
|
498
|
-
|
|
730
|
+
page = doc[n - 1]
|
|
731
|
+
rotation = page.rotation
|
|
732
|
+
text = page.get_text("text").strip()
|
|
733
|
+
if not text:
|
|
734
|
+
ocr_info["textless"].append(n)
|
|
735
|
+
if status is None:
|
|
736
|
+
status = ocr_status(False, lang)
|
|
737
|
+
if n not in keep:
|
|
738
|
+
d = page_dir(staging, n)
|
|
739
|
+
if not text:
|
|
740
|
+
pic, meta = textless_picture(worker, page, n, d, o, notes, render=not o.get("pageImages"))
|
|
741
|
+
if pic:
|
|
742
|
+
links.append(f"")
|
|
743
|
+
else:
|
|
744
|
+
eff = clamped_dpi(page.rect.width, page.rect.height, o["imageDpi"])
|
|
745
|
+
r = worker.run({"kind": "images", "page": n, "dpi": eff, "annots": not o.get("hideAnnotations"), "format": o["imageFormat"], "target": d}, raster_budget(RASTER_BUDGET_S))
|
|
746
|
+
if not r["ok"]:
|
|
747
|
+
raise RasterError(r["error"])
|
|
748
|
+
links.extend(f"" for name in r["files"])
|
|
749
|
+
mark_done(d, meta)
|
|
750
|
+
page_pic = render_page_image(worker, page, n, o, page_images, notes) if o.get("pageImages") else None
|
|
751
|
+
page_ok = True
|
|
752
|
+
except RasterUnavailable:
|
|
753
|
+
raise
|
|
499
754
|
except Exception as exc: # noqa: BLE001
|
|
500
|
-
|
|
501
|
-
|
|
502
|
-
|
|
503
|
-
|
|
504
|
-
|
|
755
|
+
shutil.rmtree(os.path.join(staging, f"p{n}"), ignore_errors=True)
|
|
756
|
+
text = ""
|
|
757
|
+
meta = {}
|
|
758
|
+
links = [f"" for f in keep.get(n, [])]
|
|
759
|
+
failed.append(page_failure(n, exc))
|
|
760
|
+
if o.get("words") and page_ok and not (not text and o.get("ocr")):
|
|
761
|
+
try:
|
|
762
|
+
word_pages.append(words_page(page, n, rotation, page_words(page, False)))
|
|
763
|
+
except Exception as exc: # noqa: BLE001
|
|
764
|
+
words_errors[str(n)] = words_error(exc)
|
|
765
|
+
if not text:
|
|
766
|
+
empty.append(n)
|
|
767
|
+
out.append("\n\n".join(x for x in [text, "\n".join(links), f"" if page_pic else ""] if x))
|
|
768
|
+
if meta.get("native"):
|
|
769
|
+
native_images.append({"page": n, **meta["native"]})
|
|
770
|
+
out.append(SEP.format(n=n).strip("\n"))
|
|
771
|
+
finally:
|
|
772
|
+
worker.stop()
|
|
505
773
|
if failed and len(failed) == len(pages):
|
|
506
774
|
raise RuntimeError("every selected page failed: " + failed[0]["error"])
|
|
507
775
|
if status is not None:
|
|
508
776
|
apply_status(ocr_info, status)
|
|
509
|
-
notes = [DEGRADED_NOTE]
|
|
510
777
|
missing = len(pages) - len(page_images) if o.get("pageImages") else 0
|
|
511
778
|
if missing:
|
|
512
779
|
notes.append(f"Page images: {missing} of {len(pages)} unavailable")
|
|
513
780
|
result = {"markdown": "\n\n".join(out) + "\n", "pages": pages, "pageCount": doc.page_count,
|
|
514
|
-
"emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": ocr_info, "pageImages": page_images, "pageStats": stats}
|
|
781
|
+
"emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": ocr_info, "pageImages": page_images, "nativeImages": native_images, "pageStats": stats}
|
|
515
782
|
if o.get("words"):
|
|
516
783
|
result.update(write_words(staging, "pt", word_pages, words_errors))
|
|
517
784
|
return result
|
|
@@ -614,7 +881,7 @@ def col_letter(i):
|
|
|
614
881
|
|
|
615
882
|
PREVIEW_ROWS, PREVIEW_COLS = 100, 50
|
|
616
883
|
PROFILE_MAJORITY, DISTINCT_CAP, SLUG_MAX = 0.6, 50, 40
|
|
617
|
-
MAX_RENDER_PX, MIN_RENDER_DPI, MIN_PAGE_PT =
|
|
884
|
+
MAX_RENDER_PX, MIN_RENDER_DPI, MIN_PAGE_PT = 50_000_000, 36, 72
|
|
618
885
|
INV_HEADER = "| # | name | kind | size | hidden | charts | images | rendered | data |\n|---|---|---|---|---|---|---|---|---|\n"
|
|
619
886
|
XLS_NOTE = "Rendered views: unavailable (visual detection not supported for .xls)"
|
|
620
887
|
|
|
@@ -1424,7 +1691,6 @@ def mode_info_excel(o):
|
|
|
1424
1691
|
return {"sheets": sheets}
|
|
1425
1692
|
|
|
1426
1693
|
def mode_render_pages(o):
|
|
1427
|
-
import math
|
|
1428
1694
|
import pymupdf
|
|
1429
1695
|
doc = pymupdf.open(o["path"])
|
|
1430
1696
|
expected = o["expectedPages"]
|
|
@@ -1439,8 +1705,8 @@ def mode_render_pages(o):
|
|
|
1439
1705
|
if w < MIN_PAGE_PT or h < MIN_PAGE_PT:
|
|
1440
1706
|
failed.append({"idx": idx, "reason": f"rendered view degenerate (page {w:.0f} x {h:.0f} pt)"})
|
|
1441
1707
|
continue
|
|
1442
|
-
eff =
|
|
1443
|
-
if eff
|
|
1708
|
+
eff = clamped_dpi(w, h, dpi)
|
|
1709
|
+
if eff is None:
|
|
1444
1710
|
failed.append({"idx": idx, "reason": f"rendered view too large (page {w:.0f} x {h:.0f} pt)"})
|
|
1445
1711
|
continue
|
|
1446
1712
|
name = f"s{idx}.{fmt}"
|