pi-quiver 6.8.0 → 6.10.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +11 -0
- package/README.md +4 -2
- package/dist/bin/pi-quiver.js +200 -10
- package/extensions/doc_to_md.ts +1 -1
- package/lib/doc-to-md-bundle.ts +62 -3
- package/lib/doc-to-md-core.ts +80 -10
- package/lib/doc-to-md-handle.ts +34 -1
- package/lib/doc-to-md-options.ts +46 -3
- package/package.json +1 -1
- package/scripts/doc_to_md.py +210 -8
package/CHANGELOG.md
CHANGED
|
@@ -8,6 +8,17 @@ Published to npm as `pi-quiver` (`pi install npm:pi-quiver`). Pushing a
|
|
|
8
8
|
via OIDC trusted publishing. The release helper at
|
|
9
9
|
`.agents/skills/release/scripts/release.sh` cuts the tag; CI publishes.
|
|
10
10
|
|
|
11
|
+
## v6.10.0 - 2026-10-02
|
|
12
|
+
|
|
13
|
+
- `doc_to_md`: new per-call `words` option (`--words`; not settable) writes `<stem>.words.json` - per selected PDF/image page, every text-layer word and every word inline OCR recognized in the same run, with display-space bbox (points; source pixels for images) and `source` `text`/`ocr`. Under `--ocr-mode all` each sidecar gets `ocr/<stem>-pNNN.words.json`. Never triggers OCR; Markdown, page stats and OCR outcome are unchanged. Handle gains `Words:`, `--json` gains `wordsPath`, `wordsReason`, `wordsErrors`, `ocr.wordSidecars`; `--info --words` is a usage error (#28).
|
|
14
|
+
- `doc_to_md`: `--help` and the generated skill list every bundle artifact under `Bundle layout` (#28).
|
|
15
|
+
|
|
16
|
+
## v6.9.0 - 2026-10-01
|
|
17
|
+
|
|
18
|
+
- `doc_to_md`: every conversion that goes through the Python PDF tiers (PDF, PPTX, `.doc`, DOCX via LibreOffice) writes `<stem>.pages.json` (per-page `chars`, `images`, `imageCoverage`) and the handle prints `Page-Stats:`; CLI `--json` gains `pageStats`, `pageStatsPath`, `ocrDir`. The unpdf tier writes no stats (#26).
|
|
19
|
+
- `doc_to_md`: new per-call `ocrMode` (`--ocr-mode textless|all`). `all` forces OCR on an explicit `pages` selection in a second, kill-capped Python child and writes `ocr/<stem>-pNNN.md` sidecars; the Markdown stays byte-identical to the same selected-page call without OCR. Refused without `--ocr`, without explicit pages, or on non-PDF/PPTX/DOC inputs (exit 2); missing Tesseract data is a hard error. The `OCR:` line names written, no-text, failed, budget-stopped, killed and not-attempted pages, with the exact `--pages` to re-run for budget-stopped and not-attempted pages (#26).
|
|
20
|
+
- `doc_to_md`: `--help` and the generated `skills/doc-to-md/SKILL.md` end with the two-pass OCR usage block (#26).
|
|
21
|
+
|
|
11
22
|
## v6.8.0 - 2026-10-01
|
|
12
23
|
|
|
13
24
|
- `session-name`: opt-in automatic naming starts in the background after three completed model/tool rounds. Shorter runs start a non-blocking attempt at run end; initial generation is best-effort with a 30-second local deadline and no automatic retry per session activation.
|
package/README.md
CHANGED
|
@@ -65,7 +65,7 @@ A 300 KB changelog page never touches your context window - you get a preview an
|
|
|
65
65
|
| Extension | Tool | What it does |
|
|
66
66
|
| --- | --- | --- |
|
|
67
67
|
| `extensions/fetch.ts` | `fetch` | Retrieve URLs over HTTP(S). HTML -> Markdown (Readability extraction, Turndown conversion). Binary saved untouched to a temp file. GitHub issue/PR/repo/actions-run/actions-job URLs auto-route through `gh` (falls back to HTTP); failed runs/jobs include failed-step logs (best-effort, summary-only otherwise). Same size gate as `fetch`. Behavior lives in `lib/fetch-core.ts`; also exposed as the `pi-quiver fetch` CLI (see [Claude Code support](#claude-code-support)). |
|
|
68
|
-
| `extensions/doc_to_md.ts` | `doc_to_md` | Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, MSG/EML, HTML, or image file to a Markdown bundle on disk (`<stem>.md` + `images/`, spreadsheet `sheets/`, optional `pageImages` in `pages/`, and email `attachments/`) and return a handle (paths, page count, outline, diagnostics) - never inline Markdown. `info` inspects PDF/Office inputs; `pages` selects 1-based PDF/Office pages (DOCX: explicit-page-break segments); HTML, image, and email inputs reject both; every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Tiers: pymupdf4llm -> PyMuPDF text (degraded) -> unpdf worker (no Python only). Excel -> sheet inventory, full CSVs, bounded previews, and optional rendered views. HTML keeps local and `data:` images in the bundle and remote images as links; image inputs keep the original image. Scanned PDF pages keep a page picture on both Python tiers; opt-in OCR needs Tesseract language data. DOCX converts directly (mammoth, python-docx fallback); LibreOffice pagination marks the degraded route. The Outline lists `L<line>` and `p<page>` per heading. Settings under `quiver.docToMd`. |
|
|
68
|
+
| `extensions/doc_to_md.ts` | `doc_to_md` | Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, MSG/EML, HTML, or image file to a Markdown bundle on disk (`<stem>.md` + `images/`, spreadsheet `sheets/`, optional `pageImages` in `pages/`, and email `attachments/`) and return a handle (paths, page count, outline, diagnostics) - never inline Markdown. `info` inspects PDF/Office inputs; `pages` selects 1-based PDF/Office pages (DOCX: explicit-page-break segments); HTML, image, and email inputs reject both; every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Tiers: pymupdf4llm -> PyMuPDF text (degraded) -> unpdf worker (no Python only). Excel -> sheet inventory, full CSVs, bounded previews, and optional rendered views. HTML keeps local and `data:` images in the bundle and remote images as links; image inputs keep the original image. Scanned PDF pages keep a page picture on both Python tiers; opt-in OCR needs Tesseract language data. DOCX converts directly (mammoth, python-docx fallback); LibreOffice pagination marks the degraded route. The Outline lists `L<line>` and `p<page>` per heading. Settings under `quiver.docToMd`. Every conversion that goes through the Python PDF tiers (PDF, PPTX, `.doc`, DOCX via LibreOffice) writes `<stem>.pages.json` (per-page chars, image count, image coverage; handle line `Page-Stats:`); unpdf writes no stats. `ocrMode: "all"` with `ocr: true` and an explicit `pages` forces OCR on those pages into `ocr/<stem>-pNNN.md` sidecars, leaving the Markdown byte-identical to the same selected-page call without OCR. `words: true` (CLI `--words`, per-call only - no settings key) writes `<stem>.words.json` with every word's bbox per selected PDF/image page (`source` `text` or `ocr`), plus `ocr/<stem>-pNNN.words.json` beside each forced-OCR sidecar; it never triggers OCR. |
|
|
69
69
|
| `extensions/session-name.ts` | `/session-name` | Manual + opt-in automatic session naming, naming rules and deny list, long-session revisits, and Ghostty/Herdr tab rename. OFF by default. |
|
|
70
70
|
| `extensions/sword-header.ts` | `/builtin-header` | Themed ASCII startup header replacing pi's default logo. OFF by default. |
|
|
71
71
|
| `extensions/fast-mode.ts` | `/fast` | Inject Anthropic fast-mode (`speed: "fast"` + `anthropic-beta: fast-mode-2026-02-01`) into every Claude Opus 4.8 / Opus 5 request, any thinking level. `--fast` flag + `/fast [on\|off\|status]`. OFF by default. |
|
|
@@ -310,7 +310,9 @@ Each setting can also be overridden per-process via `PI_QUIVER_SLACK_ENABLED`, `
|
|
|
310
310
|
| `ocr` | `false` | Opt-in OCR for scanned pages and images when Tesseract language data is installed. |
|
|
311
311
|
| `ocrLanguage` | `eng` | Plain `+`-joined Tesseract language codes for OCR. |
|
|
312
312
|
|
|
313
|
-
|
|
313
|
+
`ocrMode` (`textless` | `all`) is a per-call parameter and `--ocr-mode` flag only; a `quiver.docToMd.ocrMode` key is reported as unknown and ignored, so persisted settings cannot enable forced OCR or bypass its explicit-pages guard.
|
|
314
|
+
|
|
315
|
+
A bundle is `<outputDir>/<stem>.md` plus `<outputDir>/<stem>.pages.json` on Python PDF tiers, optional OCR sidecars under `<outputDir>/ocr/`, and `<outputDir>/images/` and, for spreadsheets with data, `<outputDir>/sheets/`; without `outputDir`, the tool creates a per-call temp root. A conversion owns `<stem>.md.lock` until it atomically publishes the Markdown. An existing `<stem>.md` fails the call unless `overwrite` is set; `overwrite` replaces that Markdown and the files it owns (`images/<stem>-p<N>-<n>.*`, `images/<stem>-s<idx>[-<n>].*`, `sheets/<stem>-s<idx>-<slug>.csv`, `<stem>.pages.json`, `ocr/<stem>-pNNN.md`), nothing else. Temp bundles are caller-owned - the tool never deletes a bundle it produced.
|
|
314
316
|
|
|
315
317
|
Excel needs a Python backend with openpyxl, xlrd and pillow - otherwise the call fails with `Remedy: install uv, or pip install openpyxl xlrd pillow`. The Markdown opens with a `## Sheets` table listing every sheet in workbook order (0-based `#`, `worksheet`/`chartsheet`, size, hidden, chart and image counts, rendered view, CSV link for non-empty worksheets), then one section per sheet: a `Data:` line linking the full-content CSV under `sheets/` for non-empty worksheets, chart metadata from the workbook model, embedded images, an optional rendered view, a preview of at most the first 100 rows x 50 columns, and - only when the preview is truncated - a `Columns:` profile (type, non-empty count, min/max, distinct up to 50). Sizes are the extent of non-empty cells (the `info` handle reports the raw worksheet dimensions instead, which may be larger). Rendered views (`images/<stem>-s<idx>.<fmt>`) are produced for sheets carrying charts or images when LibreOffice is on `PATH`: the workbook is exported one PDF page per sheet and rasterized under a 16 Mpx budget. Any LibreOffice or rasterization failure degrades to `Rendered view: unavailable (<reason>)` and a handle note; it never fails the conversion. Workbooks whose chartsheet drawings carry a zero-size anchor (openpyxl-authored files; Excel-authored files are unaffected) render as a degenerate page and are reported as such. `.xls` gets the inventory, CSVs and previews but no visual detection or rendering.
|
|
316
318
|
|
package/dist/bin/pi-quiver.js
CHANGED
|
@@ -581,6 +581,9 @@ function ownedCsvPattern(stem) {
|
|
|
581
581
|
function ownedPagePattern(stem) {
|
|
582
582
|
return new RegExp(`^${escRe(stem)}-p\\d+\\.[a-z0-9]+$`);
|
|
583
583
|
}
|
|
584
|
+
function ownedOcrPattern(stem) {
|
|
585
|
+
return new RegExp(`^${escRe(stem)}-p\\d+(?:\\.words\\.json|\\.md)$`);
|
|
586
|
+
}
|
|
584
587
|
var FILE_LINK_RE = /\[[^\]]*\]\(\s*((?:sheets|attachments)\/[^)\s]+)\s*\)/g;
|
|
585
588
|
var IMG_LINK_RE = /!\[[^\]]*\]\(\s*(?:<([^>]*)>|([^)]*?))\s*\)|<img\b[^>]*\bsrc\s*=\s*(?:"([^"]*)"|'([^']*)'|([^\s>"']+))/gi;
|
|
586
589
|
function imageTarget(m) {
|
|
@@ -627,6 +630,7 @@ function openBundle(root, requested, overwrite) {
|
|
|
627
630
|
const sheetsDir = join2(root, "sheets");
|
|
628
631
|
const pagesDir = join2(root, "pages");
|
|
629
632
|
const attachmentsDir = join2(root, "attachments");
|
|
633
|
+
const ocrDir = join2(root, "ocr");
|
|
630
634
|
try {
|
|
631
635
|
if (existsSync(mdPath) && overwrite) {
|
|
632
636
|
const owned = ownedPattern(stem), ownedCsv = ownedCsvPattern(stem), ownedPage = ownedPagePattern(stem);
|
|
@@ -644,6 +648,12 @@ function openBundle(root, requested, overwrite) {
|
|
|
644
648
|
if (existsSync(attachmentsDir)) {
|
|
645
649
|
for (const f of readdirSync(attachmentsDir)) if (attachments.has(f)) rmSync(join2(attachmentsDir, f), { force: true });
|
|
646
650
|
}
|
|
651
|
+
rmSync(join2(root, `${stem}.pages.json`), { force: true });
|
|
652
|
+
rmSync(join2(root, `${stem}.words.json`), { force: true });
|
|
653
|
+
const ownedOcr = ownedOcrPattern(stem);
|
|
654
|
+
if (existsSync(ocrDir)) {
|
|
655
|
+
for (const f of readdirSync(ocrDir)) if (ownedOcr.test(f)) rmSync(join2(ocrDir, f), { force: true });
|
|
656
|
+
}
|
|
647
657
|
}
|
|
648
658
|
const lockId = randomBytes(6).toString("hex");
|
|
649
659
|
const stagingDir = join2(imagesDir, `.stage-${lockId}`);
|
|
@@ -651,7 +661,8 @@ function openBundle(root, requested, overwrite) {
|
|
|
651
661
|
mkdirSync2(stagingDir, { recursive: true });
|
|
652
662
|
const pagesStagingDir = join2(pagesDir, `.stage-${lockId}`);
|
|
653
663
|
const attachmentsStagingDir = join2(attachmentsDir, `.stage-${lockId}`);
|
|
654
|
-
|
|
664
|
+
const ocrStagingDir = join2(ocrDir, `.stage-${lockId}`);
|
|
665
|
+
return { root, stem, renamedFrom, renameReason, mdPath, lockPath, imagesDir, stagingDir, lockId, sheetsDir, sheetsStagingDir, pagesDir, pagesStagingDir, attachmentsDir, attachmentsStagingDir, ocrDir, ocrStagingDir, pageStatsPath: join2(root, `${stem}.pages.json`), wordsPath: join2(root, `${stem}.words.json`), ocrManifest: /* @__PURE__ */ new Set(), manifest: /* @__PURE__ */ new Set(), csvManifest: /* @__PURE__ */ new Set(), pageManifest: /* @__PURE__ */ new Set(), attachmentManifest: /* @__PURE__ */ new Set(), sourceMap: /* @__PURE__ */ new Map() };
|
|
655
666
|
} catch (e) {
|
|
656
667
|
rmSync(lockPath, { force: true });
|
|
657
668
|
throw e;
|
|
@@ -715,6 +726,45 @@ function publishPageImages(b, pageCount) {
|
|
|
715
726
|
}
|
|
716
727
|
rmSync(b.pagesStagingDir, { recursive: true, force: true });
|
|
717
728
|
}
|
|
729
|
+
function writePageStats(b, stats) {
|
|
730
|
+
writeFileSync2(b.pageStatsPath, `${JSON.stringify(stats, null, 2)}
|
|
731
|
+
`, "utf8");
|
|
732
|
+
}
|
|
733
|
+
function publishWords(b) {
|
|
734
|
+
const staged = join2(b.stagingDir, "words.json");
|
|
735
|
+
try {
|
|
736
|
+
if (!existsSync(staged)) throw new Error("child staged no words.json");
|
|
737
|
+
renameSync(staged, b.wordsPath);
|
|
738
|
+
return null;
|
|
739
|
+
} catch (e) {
|
|
740
|
+
rmSync(b.wordsPath, { force: true });
|
|
741
|
+
return `write failed - ${e.message}`;
|
|
742
|
+
}
|
|
743
|
+
}
|
|
744
|
+
function publishSidecars(b) {
|
|
745
|
+
const sidecars = /* @__PURE__ */ new Map(), wordSidecars = /* @__PURE__ */ new Map();
|
|
746
|
+
if (!existsSync(b.ocrStagingDir)) return { sidecars, wordSidecars };
|
|
747
|
+
for (const dir of readdirSync(b.ocrStagingDir).sort()) {
|
|
748
|
+
const m = dir.match(/^p(\d+)$/);
|
|
749
|
+
if (!m) continue;
|
|
750
|
+
const pageDir = join2(b.ocrStagingDir, dir);
|
|
751
|
+
if (!existsSync(join2(pageDir, ".done"))) continue;
|
|
752
|
+
const file = `${b.stem}-${dir}.md`;
|
|
753
|
+
if (!existsSync(join2(pageDir, file))) continue;
|
|
754
|
+
mkdirSync2(b.ocrDir, { recursive: true });
|
|
755
|
+
renameSync(join2(pageDir, file), join2(b.ocrDir, file));
|
|
756
|
+
b.ocrManifest.add(file);
|
|
757
|
+
sidecars.set(Number(m[1]), join2(b.ocrDir, file));
|
|
758
|
+
const wfile = `${b.stem}-${dir}.words.json`;
|
|
759
|
+
if (existsSync(join2(pageDir, wfile))) {
|
|
760
|
+
renameSync(join2(pageDir, wfile), join2(b.ocrDir, wfile));
|
|
761
|
+
b.ocrManifest.add(wfile);
|
|
762
|
+
wordSidecars.set(Number(m[1]), join2(b.ocrDir, wfile));
|
|
763
|
+
}
|
|
764
|
+
}
|
|
765
|
+
rmSync(b.ocrStagingDir, { recursive: true, force: true });
|
|
766
|
+
return { sidecars, wordSidecars };
|
|
767
|
+
}
|
|
718
768
|
function publishAttachments(b) {
|
|
719
769
|
if (!existsSync(b.attachmentsStagingDir)) return;
|
|
720
770
|
for (const f of readdirSync(b.attachmentsStagingDir).sort()) {
|
|
@@ -772,6 +822,10 @@ function commitBundle(b, markdown) {
|
|
|
772
822
|
fs.rmSync(b.attachmentsStagingDir, { recursive: true, force: true });
|
|
773
823
|
} catch {
|
|
774
824
|
}
|
|
825
|
+
try {
|
|
826
|
+
fs.rmSync(b.ocrStagingDir, { recursive: true, force: true });
|
|
827
|
+
} catch {
|
|
828
|
+
}
|
|
775
829
|
try {
|
|
776
830
|
fs.rmSync(b.lockPath, { force: true });
|
|
777
831
|
} catch {
|
|
@@ -782,6 +836,10 @@ function abortBundle(b) {
|
|
|
782
836
|
for (const f of b.csvManifest) rmSync(join2(b.sheetsDir, f), { force: true });
|
|
783
837
|
for (const f of b.pageManifest) rmSync(join2(b.pagesDir, f), { force: true });
|
|
784
838
|
for (const f of b.attachmentManifest) rmSync(join2(b.attachmentsDir, f), { force: true });
|
|
839
|
+
for (const f of b.ocrManifest) rmSync(join2(b.ocrDir, f), { force: true });
|
|
840
|
+
rmSync(b.ocrStagingDir, { recursive: true, force: true });
|
|
841
|
+
rmSync(b.pageStatsPath, { force: true });
|
|
842
|
+
rmSync(b.wordsPath, { force: true });
|
|
785
843
|
rmSync(`${b.mdPath}.tmp`, { force: true });
|
|
786
844
|
rmSync(b.stagingDir, { recursive: true, force: true });
|
|
787
845
|
rmSync(b.sheetsStagingDir, { recursive: true, force: true });
|
|
@@ -820,7 +878,21 @@ function compactRanges(nums, maxEntries = 20) {
|
|
|
820
878
|
}
|
|
821
879
|
var INSTALL_HINT = "install Tesseract (see doc/doc-to-md.md), then rerun with ocr=true";
|
|
822
880
|
var BARE_REASONS = ["fallback tier", "no Python backend"];
|
|
881
|
+
var pageList = (pages) => `page${pages.length === 1 ? "" : "s"} ${pages.join(", ")}`;
|
|
882
|
+
function forcedOcrLine(ocr) {
|
|
883
|
+
const clauses = [];
|
|
884
|
+
if (ocr.pages.length) clauses.push(`sidecars for ${pageList(ocr.pages)}`);
|
|
885
|
+
if (ocr.noText.length) clauses.push(`no text on ${pageList(ocr.noText)}`);
|
|
886
|
+
for (const page of ocr.ocrFailed) clauses.push(`failed on page ${page} (${ocr.ocrErrors[page] ?? "unknown error"})`);
|
|
887
|
+
if (ocr.budgetStopped.length) clauses.push(`budget-stopped ${pageList(ocr.budgetStopped)}`);
|
|
888
|
+
if (ocr.killed !== null) clauses.push(`page ${ocr.killed} killed the OCR child (timeout or crash - likely a compression bomb)`);
|
|
889
|
+
if (ocr.childError !== null) clauses.push(`OCR child failed before processing pages: ${ocr.childError}`);
|
|
890
|
+
if (ocr.notAttempted.length) clauses.push(`${pageList(ocr.notAttempted)} not attempted`);
|
|
891
|
+
const rerun = [...ocr.budgetStopped, ...ocr.notAttempted].sort((a, b) => a - b);
|
|
892
|
+
return `OCR: forced (${ocr.lang}) - ${clauses.join("; ")}${rerun.length ? ` - re-run with --pages ${rerun.join(",")}` : ""}`;
|
|
893
|
+
}
|
|
823
894
|
function ocrLine(ocr, type) {
|
|
895
|
+
if (ocr.mode === "all") return forcedOcrLine(ocr);
|
|
824
896
|
const image = type === "image";
|
|
825
897
|
const rerun = ocr.tesseract ? "rerun with ocr=true" : INSTALL_HINT;
|
|
826
898
|
switch (ocr.status) {
|
|
@@ -896,6 +968,12 @@ function formatHandle(h) {
|
|
|
896
968
|
if (h.sheetsDir) lines.push(`Sheets-Dir: ${h.sheetsDir}`);
|
|
897
969
|
if (h.pagesDir && h.pageImageCount > 0) lines.push(`Pages-Dir: ${h.pagesDir} (${h.pageImageCount} pages)`);
|
|
898
970
|
else if (h.pageImagesReason) lines.push(`Pages-Dir: none - ${h.pageImagesReason}`);
|
|
971
|
+
if (h.pageStatsPath) lines.push(`Page-Stats: ${h.pageStatsPath}`);
|
|
972
|
+
if (h.wordsPath) {
|
|
973
|
+
const bad = Object.keys(h.wordsErrors).map(Number).sort((a, b) => a - b);
|
|
974
|
+
lines.push(`Words: ${h.wordsPath}${bad.length ? ` (extraction failed for pages ${compactRanges(bad)}: ${h.wordsErrors[bad[0]]})` : ""}`);
|
|
975
|
+
} else if (h.wordsReason) lines.push(`Words: ${h.wordsReason}`);
|
|
976
|
+
if (h.ocrDir) lines.push(`OCR-Dir: ${h.ocrDir}`);
|
|
899
977
|
lines.push(`Type: ${h.type} Engine: ${h.engine} Tier: ${h.tier}`);
|
|
900
978
|
lines.push(`Page-Count: ${pageCountLabel(h)} Pages: ${h.pages ? compactRanges(h.pages) : "all"} Images: ${h.imageCount} Size: ${formatSize2(h.bytes)} / ${h.lines} lines`);
|
|
901
979
|
if (h.degraded) lines.push(`Degraded: ${h.degraded}`);
|
|
@@ -945,6 +1023,8 @@ var DOC_TO_MD_OPTIONS = [
|
|
|
945
1023
|
{ key: "outputDir", flag: "--output-dir", type: "string", default: null, settable: false, help: "Bundle root for <stem>.md + images/; default a per-call temp dir. <stem> = basename without extension with [^A-Za-z0-9._-]+ -> _ (empty -> document); a second call on the same stem writes <stem>-2.md" },
|
|
946
1024
|
{ key: "overwrite", flag: "--overwrite", type: "bool", default: false, settable: false, help: "Replace an existing completed <stem>.md bundle" },
|
|
947
1025
|
{ key: "pageImages", flag: "--page-images", type: "bool", default: false, settable: false, help: "Also render every selected page to pages/<stem>-pNNN.<imageFormat> at imageDpi (PDF, PPTX, .doc, DOCX via LibreOffice); off by default" },
|
|
1026
|
+
{ key: "words", flag: "--words", type: "bool", default: false, settable: false, help: 'Write word positions: <stem>.words.json beside the Markdown lists every text-layer word of each selected page with its bbox (PDF points, top-left origin, display orientation; image inputs in source pixels) and the words inline OCR recognized, tagged source "text" or "ocr"; under --ocr-mode all the OCR words go to ocr/<stem>-pNNN.words.json beside each sidecar. Never triggers OCR. PDF and image inputs only.' },
|
|
1027
|
+
{ key: "ocrMode", flag: "--ocr-mode", type: "enum", default: "textless", settable: false, enumValues: ["textless", "all"], help: "OCR policy: textless (default) OCRs only pages with an empty text layer, inline; all OCRs every selected page and writes the recognized text to ocr/<stem>-pNNN.md sidecars, leaving the Markdown untouched. all requires --ocr and an explicit --pages selection (PDF, PPTX, DOC)." },
|
|
948
1028
|
{ key: "primaryTimeoutMs", flag: "--primary-timeout", type: "int", default: 6e4, settable: true, env: "PI_DOC_TO_MD_CONVERT_TIMEOUT_MS", help: "pymupdf4llm tier and DOCX child (docx mode); also the unpdf tier" },
|
|
949
1029
|
{ key: "fallbackTimeoutMs", flag: "--fallback-timeout", type: "int", default: 3e4, settable: true, help: "PyMuPDF get_text tier (including DOCX LibreOffice fallback); also PDF and DOCX info and Excel rendered views" },
|
|
950
1030
|
{ key: "sofficeTimeoutMs", flag: "--soffice-timeout", type: "int", default: 12e4, settable: true, env: "PI_DOC_TO_MD_SOFFICE_TIMEOUT_MS", help: "DOCX/PPTX -> PDF via LibreOffice; also Excel rendered views" },
|
|
@@ -1051,11 +1131,40 @@ function resolveOptions(perCall, settings, env) {
|
|
|
1051
1131
|
out[d.key] = value;
|
|
1052
1132
|
}
|
|
1053
1133
|
const o = out;
|
|
1054
|
-
if (o.info && (o.pages !== null || o.outputDir !== null || o.overwrite || o.pageImages)) {
|
|
1055
|
-
throw new UsageError("--info cannot be combined with --pages, --output-dir, --overwrite or --
|
|
1134
|
+
if (o.info && (o.pages !== null || o.outputDir !== null || o.overwrite || o.pageImages || o.words || o.ocrMode === "all")) {
|
|
1135
|
+
throw new UsageError("--info cannot be combined with --pages, --output-dir, --overwrite, --page-images, --words or --ocr-mode all");
|
|
1056
1136
|
}
|
|
1057
1137
|
return o;
|
|
1058
1138
|
}
|
|
1139
|
+
var USAGE_PATTERNS = [
|
|
1140
|
+
"Two-pass OCR (PDF, PPTX, DOC):",
|
|
1141
|
+
" 1. <cmd> report.pdf --output-dir out --json",
|
|
1142
|
+
' -> "pageStatsPath" points at out/report.pages.json; pages with few',
|
|
1143
|
+
' chars and high imageCoverage are scans. "savedTo" is the Markdown.',
|
|
1144
|
+
" 2. <cmd> report.pdf --output-dir out --ocr --ocr-mode all --pages 2,7 --json",
|
|
1145
|
+
' -> "ocr"."sidecars" maps 2 and 7 to out/ocr/report-2-p002.md and',
|
|
1146
|
+
" ...-p007.md (a second run in the same dir gets stem report-2);",
|
|
1147
|
+
" the Markdown of this run holds pages 2 and 7 only and equals what",
|
|
1148
|
+
" --pages 2,7 without --ocr would produce. Read sidecars and Markdown",
|
|
1149
|
+
" by the returned paths, never by guessing names.",
|
|
1150
|
+
" 3. --ocr-mode all refuses to run without --ocr and an explicit --pages.",
|
|
1151
|
+
' The "ocr" object (OCR: line) names failed, budget-stopped, killed and',
|
|
1152
|
+
" not-attempted pages and the exact --pages to re-run.",
|
|
1153
|
+
" Details: doc/doc-to-md.md (bundle contract, failure buckets)."
|
|
1154
|
+
].join("\n");
|
|
1155
|
+
var usagePatterns = (cmd) => USAGE_PATTERNS.replaceAll("<cmd>", cmd);
|
|
1156
|
+
var BUNDLE_LAYOUT = [
|
|
1157
|
+
{ artifact: "<stem>.md", trigger: "always", content: "the Markdown", namedBy: "Saved-To: / savedTo" },
|
|
1158
|
+
{ artifact: "images/", trigger: "embedded or extracted figures", content: "image files linked from the Markdown", namedBy: "Images-Dir: / imagesDir" },
|
|
1159
|
+
{ artifact: "pages/<stem>-pNNN.<fmt>", trigger: "--page-images", content: "page renders at --image-dpi", namedBy: "Pages-Dir: / pagesDir" },
|
|
1160
|
+
{ artifact: "sheets/", trigger: "Excel input", content: "one CSV per non-empty worksheet", namedBy: "Sheets-Dir: / sheetsDir" },
|
|
1161
|
+
{ artifact: "attachments/", trigger: "email input", content: "saved attachments", namedBy: "Markdown attachment list" },
|
|
1162
|
+
{ artifact: "<stem>.pages.json", trigger: "Python PDF tiers (PDF, PPTX, DOC, DOCX via LibreOffice; not unpdf)", content: "per-page chars, image count, image coverage", namedBy: "Page-Stats: / pageStatsPath" },
|
|
1163
|
+
{ artifact: "<stem>.words.json", trigger: "--words", content: "per-page word boxes, source text/ocr", namedBy: "Words: / wordsPath" },
|
|
1164
|
+
{ artifact: "ocr/<stem>-pNNN.md", trigger: "--ocr --ocr-mode all", content: "recognized text of a forced page", namedBy: "OCR-Dir: / ocr.sidecars" },
|
|
1165
|
+
{ artifact: "ocr/<stem>-pNNN.words.json", trigger: "--ocr --ocr-mode all --words", content: "word boxes of that OCR", namedBy: "ocr.wordSidecars" }
|
|
1166
|
+
];
|
|
1167
|
+
var bundleLayoutText = () => BUNDLE_LAYOUT.map((r) => ` ${r.artifact.padEnd(28)} ${r.trigger}; ${r.content}; named by ${r.namedBy}`).join("\n");
|
|
1059
1168
|
function renderHelp() {
|
|
1060
1169
|
const row = (d) => ` ${(d.flag ?? "<path>").padEnd(26)} ${d.help}${d.default !== null && d.key !== "info" && d.key !== "overwrite" ? ` (default ${d.default})` : ""}`;
|
|
1061
1170
|
return [
|
|
@@ -1067,8 +1176,13 @@ function renderHelp() {
|
|
|
1067
1176
|
"Tunables (also settable under quiver.docToMd in settings.json; per-call > settings > default):",
|
|
1068
1177
|
...DOC_TO_MD_OPTIONS.filter((d) => d.settable).map(row),
|
|
1069
1178
|
"",
|
|
1070
|
-
"Result: a handle (Saved-To, Images-Dir, Page-Count, Outline ...). Read the Saved-To file for the Markdown.",
|
|
1071
|
-
"Exit codes: 0 success, 1 runtime error, 2 usage error."
|
|
1179
|
+
"Result: a handle (Saved-To, Images-Dir, Page-Stats, Page-Count, Outline ...). Read the Saved-To file for the Markdown.",
|
|
1180
|
+
"Exit codes: 0 success, 1 runtime error, 2 usage error.",
|
|
1181
|
+
"",
|
|
1182
|
+
"Bundle layout:",
|
|
1183
|
+
bundleLayoutText(),
|
|
1184
|
+
"",
|
|
1185
|
+
usagePatterns("pi-quiver doc-to-md")
|
|
1072
1186
|
].join("\n");
|
|
1073
1187
|
}
|
|
1074
1188
|
|
|
@@ -1547,7 +1661,38 @@ function reconcileRenderMarkers(md, renderPages, fmt, sourceMap, reason) {
|
|
|
1547
1661
|
if (/<!--rvs?:\d+-->/.test(md)) throw new Error("internal: unresolved render marker");
|
|
1548
1662
|
return md;
|
|
1549
1663
|
}
|
|
1550
|
-
var emptyOcr = (lang) => ({ status: "off", lang, textless: [], pages: [], noText: [], ocrFailed: [], budgetStopped: [], reason: null, tesseract: null });
|
|
1664
|
+
var emptyOcr = (lang) => ({ status: "off", lang, textless: [], pages: [], noText: [], ocrFailed: [], budgetStopped: [], reason: null, tesseract: null, mode: "textless", sidecars: {}, wordSidecars: {}, ocrErrors: {}, killed: null, notAttempted: [], childError: null });
|
|
1665
|
+
var emptyOutcome = () => ({ written: [], noText: [], ocrFailed: [], ocrErrors: {}, budgetStopped: [], killed: null, notAttempted: [], childError: null });
|
|
1666
|
+
var ocrPagesTag = (n) => `p${String(n).padStart(3, "0")}`;
|
|
1667
|
+
var SIDE_MARKER_RE = /^--- end of page\.page_number=\d+ ---$/;
|
|
1668
|
+
function sidecarHasText(dir) {
|
|
1669
|
+
const file = readdirSync2(dir).find((f) => f.endsWith(".md"));
|
|
1670
|
+
if (!file) return false;
|
|
1671
|
+
return readFileSync2(join3(dir, file), "utf8").split("\n").slice(1).some((l) => l.trim() && !SIDE_MARKER_RE.test(l));
|
|
1672
|
+
}
|
|
1673
|
+
function recoverOcrPages(stagingDir, pages, detail) {
|
|
1674
|
+
const out = emptyOutcome();
|
|
1675
|
+
const activePath = join3(stagingDir, "active");
|
|
1676
|
+
const active = existsSync2(activePath) ? Number(readFileSync2(activePath, "utf8").trim()) || null : null;
|
|
1677
|
+
let sawPage = false;
|
|
1678
|
+
for (const n of pages) {
|
|
1679
|
+
const dir = join3(stagingDir, ocrPagesTag(n));
|
|
1680
|
+
if (existsSync2(join3(dir, ".done"))) {
|
|
1681
|
+
sawPage = true;
|
|
1682
|
+
(sidecarHasText(dir) ? out.written : out.noText).push(n);
|
|
1683
|
+
} else if (existsSync2(join3(dir, ".failed"))) {
|
|
1684
|
+
sawPage = true;
|
|
1685
|
+
out.ocrFailed.push(n);
|
|
1686
|
+
out.ocrErrors[n] = readFileSync2(join3(dir, ".failed"), "utf8").trim() || "unknown error";
|
|
1687
|
+
} else if (n === active) out.killed = n;
|
|
1688
|
+
else out.notAttempted.push(n);
|
|
1689
|
+
}
|
|
1690
|
+
if (active === null && !sawPage) out.childError = detail;
|
|
1691
|
+
return out;
|
|
1692
|
+
}
|
|
1693
|
+
function outcomeFromChild(j) {
|
|
1694
|
+
return { ...emptyOutcome(), written: j.written ?? [], noText: j.noText ?? [], ocrFailed: j.ocrFailed ?? [], ocrErrors: j.ocrErrors ?? {}, budgetStopped: j.budgetStopped ?? [] };
|
|
1695
|
+
}
|
|
1551
1696
|
var OCR_SENTINEL_RE = /\x00OCR ([^\x00]*)\x00/g;
|
|
1552
1697
|
function resolveOcrLabels(md, sourceMap) {
|
|
1553
1698
|
return md.replace(OCR_SENTINEL_RE, (_, key) => {
|
|
@@ -1559,7 +1704,7 @@ function handleOcr(tier, type, o, json) {
|
|
|
1559
1704
|
if (tier === "unpdf") return o.ocr ? { ...emptyOcr(o.ocrLanguage), status: "unavailable", reason: "no Python backend" } : null;
|
|
1560
1705
|
const x = json.ocr;
|
|
1561
1706
|
if (!x || type !== "image" && !x.textless.length && !x.pages.length && !x.ocrFailed.length) return null;
|
|
1562
|
-
return x;
|
|
1707
|
+
return { ...emptyOcr(o.ocrLanguage), ...x };
|
|
1563
1708
|
}
|
|
1564
1709
|
async function convertDocument(o, signal, seams) {
|
|
1565
1710
|
const s = { backend: (c) => getBackend(c, void 0, signal), runTier: runTierReal, office: tryConvertOffice, ...seams };
|
|
@@ -1568,10 +1713,17 @@ async function convertDocument(o, signal, seams) {
|
|
|
1568
1713
|
if (!st || !st.isFile()) throw new Error(`Not a readable file: ${o.path}`);
|
|
1569
1714
|
if (st.size === 0) throw new Error(`empty file: ${o.path}`);
|
|
1570
1715
|
const type = classifyInput(inputPath);
|
|
1716
|
+
const forced = o.ocrMode === "all";
|
|
1717
|
+
if (forced) {
|
|
1718
|
+
if (!o.ocr) throw new UsageError("--ocr-mode all requires --ocr");
|
|
1719
|
+
if (o.pages === null) throw new UsageError('--ocr-mode all requires an explicit --pages selection (e.g. --pages 2,7); omitted pages and --pages "" mean all pages and are refused to keep OCR cost bounded');
|
|
1720
|
+
if (type !== "pdf" && type !== "pptx" && type !== "doc") throw new UsageError("--ocr-mode all applies to PDF, PPTX and DOC inputs only (DOCX pages are page-break segments, not PDF pages; convert the DOCX to PDF first)");
|
|
1721
|
+
}
|
|
1571
1722
|
if (o.pages && (type === "html" || type === "image" || type === "email")) throw new Error(`--pages does not apply to ${type === "html" ? "HTML files" : type === "image" ? "images" : "email"}`);
|
|
1572
1723
|
const isExcel = type === "xlsx" || type === "xlsm" || type === "xls";
|
|
1573
1724
|
if (isExcel && o.pages) throw new Error("--pages does not apply to spreadsheets: worksheets have no stable page numbering");
|
|
1574
1725
|
const backend = await s.backend({ pymupdfVersion: o.pymupdfVersion, warmTimeoutMs: o.warmTimeoutMs });
|
|
1726
|
+
if (forced && backend.kind === "none") throw new Error(`--ocr-mode all cannot run: no Python backend (${backend.reason})`);
|
|
1575
1727
|
if (isExcel && (backend.kind === "none" || !backend.xlsx)) throw new Error(`Excel conversion needs a Python backend with openpyxl, xlrd and pillow (${backend.kind === "none" ? backend.reason : `${backend.kind} lacks the Excel packages`}). ${EXCEL_REMEDY}`);
|
|
1576
1728
|
if (type === "email" && extname3(inputPath).toLowerCase() === ".msg" && (backend.kind === "none" || !backend.email)) throw new Error(`MSG conversion needs the extract-msg package. Python backend: ${backendState(backend, "found without extract-msg")}. Remedy: install uv, or pip install extract-msg markdownify into that Python`);
|
|
1577
1729
|
if (type === "email" && extname3(inputPath).toLowerCase() !== ".msg" && lacksDocx(backend)) throw new Error(`EML conversion needs the Python DOCX/HTML packages (mammoth, markdownify, python-docx). Python backend: ${backendState(backend, "found without mammoth/markdownify/python-docx")}. ${docxRemedy}`);
|
|
@@ -1580,11 +1732,12 @@ async function convertDocument(o, signal, seams) {
|
|
|
1580
1732
|
let office = null;
|
|
1581
1733
|
try {
|
|
1582
1734
|
let pdfPath = inputPath;
|
|
1583
|
-
const base = { path: inputPath, pages: o.pages, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
|
|
1735
|
+
const base = { path: inputPath, pages: o.pages, ...o.words && (type === "pdf" || type === "image") ? { words: true } : {}, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr && !forced, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
|
|
1584
1736
|
let tier, engine, json, degraded = null, fallbackReason = null;
|
|
1585
1737
|
let explicitBreaks = null;
|
|
1586
1738
|
let notes = [];
|
|
1587
1739
|
let officeRoute = null;
|
|
1740
|
+
let copyReason = null;
|
|
1588
1741
|
if (type === "html") {
|
|
1589
1742
|
const prepared = await prepareHtml(inputPath, b.stagingDir);
|
|
1590
1743
|
if (signal?.aborted) throw new Error("aborted");
|
|
@@ -1624,6 +1777,7 @@ async function convertDocument(o, signal, seams) {
|
|
|
1624
1777
|
}
|
|
1625
1778
|
if (signal?.aborted) throw new Error("aborted");
|
|
1626
1779
|
if (json === void 0) {
|
|
1780
|
+
copyReason = reason;
|
|
1627
1781
|
clearStaging(b);
|
|
1628
1782
|
const file = `original${extname3(inputPath).toLowerCase()}`;
|
|
1629
1783
|
const dir = join3(b.stagingDir, "p1");
|
|
@@ -1757,6 +1911,36 @@ async function convertDocument(o, signal, seams) {
|
|
|
1757
1911
|
if (type === "doc" || officeRoute !== null) degraded = DEGRADED_DOCX_OFFICE;
|
|
1758
1912
|
if (officeRoute !== null) fallbackReason = fallbackReason ? `${officeRoute}; ${fallbackReason}` : officeRoute;
|
|
1759
1913
|
if (tier === void 0 || engine === void 0 || json === void 0) throw new Error("internal: no tier produced output");
|
|
1914
|
+
const pageStats = json.pageStats ?? null;
|
|
1915
|
+
if (pageStats) writePageStats(b, pageStats);
|
|
1916
|
+
let wordsPath = null, wordsReason = null;
|
|
1917
|
+
const wordsErrors = {};
|
|
1918
|
+
const takeWordsErrors = (j) => {
|
|
1919
|
+
for (const [k, v] of Object.entries(j?.wordsErrors ?? {})) {
|
|
1920
|
+
const page = Number(k);
|
|
1921
|
+
if (Number.isFinite(page)) wordsErrors[page] = v;
|
|
1922
|
+
}
|
|
1923
|
+
};
|
|
1924
|
+
if (o.words) {
|
|
1925
|
+
if (type !== "pdf" && type !== "image") wordsReason = `none - word positions apply to PDF and image inputs only (${type})`;
|
|
1926
|
+
else if (tier === "unpdf") wordsReason = "none - unpdf tier has no page geometry";
|
|
1927
|
+
else if (engine === "copy") wordsReason = `none - image copied without conversion (${copyReason})`;
|
|
1928
|
+
else if (json?.words === true) {
|
|
1929
|
+
wordsReason = publishWords(b);
|
|
1930
|
+
if (wordsReason === null) wordsPath = b.wordsPath;
|
|
1931
|
+
} else wordsReason = `write failed - ${json?.wordsErrors?.file ?? "child reported no words document"}`;
|
|
1932
|
+
takeWordsErrors(json);
|
|
1933
|
+
}
|
|
1934
|
+
let ocr = handleOcr(tier, type, o, json);
|
|
1935
|
+
if (forced) {
|
|
1936
|
+
const r = await s.runTier("ocr-pages", { path: pdfPath, pages: o.pages, ...o.words && type === "pdf" ? { words: true } : {}, stem: b.stem, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs, stagingDir: b.ocrStagingDir, dpi: o.imageDpi, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion }, b, signal, o.primaryTimeoutMs, backend);
|
|
1937
|
+
if (signal?.aborted || !r.ok && "reason" in r && r.reason === "aborted") throw new Error("aborted");
|
|
1938
|
+
if (r.ok && r.json.status === "unavailable") throw new Error(`OCR unavailable: ${r.json.reason} (install Tesseract; see doc/doc-to-md.md)`);
|
|
1939
|
+
const outcome = r.ok && r.json.status === "ran" ? outcomeFromChild(r.json) : recoverOcrPages(b.ocrStagingDir, o.pages, !r.ok ? "userError" in r ? r.userError : `${r.reason}${detailSuffix(r)}` : "malformed child output");
|
|
1940
|
+
const { sidecars, wordSidecars } = publishSidecars(b);
|
|
1941
|
+
ocr = { ...emptyOcr(o.ocrLanguage), ...json.ocr, status: "ran", reason: null, mode: "all", pages: outcome.written, noText: outcome.noText, ocrFailed: outcome.ocrFailed, ocrErrors: outcome.ocrErrors, budgetStopped: outcome.budgetStopped, killed: outcome.killed, notAttempted: outcome.notAttempted, childError: outcome.childError, sidecars: Object.fromEntries(sidecars), wordSidecars: Object.fromEntries(wordSidecars) };
|
|
1942
|
+
if (o.words && r.ok) takeWordsErrors(r.json);
|
|
1943
|
+
}
|
|
1760
1944
|
if (!isExcel) notes = [...notes, ...json.notes ?? []];
|
|
1761
1945
|
if (b.renamedFrom) notes.splice(notes[0]?.startsWith("preview truncated:") ? 1 : 0, 0, `renamed to ${b.stem} (${b.renameReason})`);
|
|
1762
1946
|
const pageImagesReason = !o.pageImages || b.pageManifest.size || tier === "primary" || tier === "fallback" ? null : tier === "unpdf" ? "page images need the Python backend" : `${type} has no page geometry`;
|
|
@@ -1773,7 +1957,7 @@ async function convertDocument(o, signal, seams) {
|
|
|
1773
1957
|
` : "") + body;
|
|
1774
1958
|
commitBundle(b, markdown);
|
|
1775
1959
|
const outline = scanOutline(markdown, o.outlineMaxEntries);
|
|
1776
|
-
const details = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr:
|
|
1960
|
+
const details = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr, pageStats, pageStatsPath: pageStats ? b.pageStatsPath : null, ocrDir: b.ocrManifest.size ? b.ocrDir : null, wordsPath, wordsReason, wordsErrors };
|
|
1777
1961
|
return { output: formatHandle(details), details };
|
|
1778
1962
|
} catch (e) {
|
|
1779
1963
|
abortBundle(b);
|
|
@@ -1819,7 +2003,7 @@ async function inspectDocument(o, signal, seams) {
|
|
|
1819
2003
|
}
|
|
1820
2004
|
|
|
1821
2005
|
// bin/pi-quiver.ts
|
|
1822
|
-
var USAGE = 'Usage: pi-quiver fetch <url> [--method GET|HEAD|POST] [--header "K: V"]... [--body <str>] [--raw] [--timeout-ms <n>]\n pi-quiver doc-to-md [--json] [--info] [--page-images] [--pages <spec>] [--output-dir <dir>] [--overwrite] [tunable flags] <path> (--help for all flags)';
|
|
2006
|
+
var USAGE = 'Usage: pi-quiver fetch <url> [--method GET|HEAD|POST] [--header "K: V"]... [--body <str>] [--raw] [--timeout-ms <n>]\n pi-quiver doc-to-md [--json] [--info] [--page-images] [--words] [--pages <spec>] [--output-dir <dir>] [--overwrite] [--ocr] [--ocr-mode textless|all] [tunable flags] <path> (--help for all flags)';
|
|
1823
2007
|
function parseDocToMd(rest) {
|
|
1824
2008
|
if (rest.includes("--help") || rest.includes("-h")) return { ok: true, cmd: "doc-to-md-help" };
|
|
1825
2009
|
const json = rest.includes("--json");
|
|
@@ -1974,6 +2158,12 @@ ${USAGE}
|
|
|
1974
2158
|
`);
|
|
1975
2159
|
return 0;
|
|
1976
2160
|
} catch (err) {
|
|
2161
|
+
if (err instanceof UsageError) {
|
|
2162
|
+
process.stderr.write(`${err.message}
|
|
2163
|
+
${USAGE}
|
|
2164
|
+
`);
|
|
2165
|
+
return 2;
|
|
2166
|
+
}
|
|
1977
2167
|
process.stderr.write(`doc-to-md failed: ${err instanceof Error ? err.message : String(err)}
|
|
1978
2168
|
`);
|
|
1979
2169
|
return 1;
|
package/extensions/doc_to_md.ts
CHANGED
|
@@ -35,7 +35,7 @@ export default function docToMdExtension(pi: ExtensionAPI) {
|
|
|
35
35
|
name: "doc_to_md",
|
|
36
36
|
label: "Convert doc to Markdown bundle",
|
|
37
37
|
description:
|
|
38
|
-
"Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, email (.msg/.eml), HTML (.html/.htm) or image (.png .jpg .jpeg .tif .tiff .bmp .gif) to a Markdown bundle on disk and return a handle (Saved-To, Images-Dir, Page-Count, Outline, diagnostics) - the Markdown itself is never inlined; read the Saved-To file (offset/limit) for content. `info: true` returns page count, metadata and TOC (or the sheet inventory) without converting - use it to pick `pages`. `pages` selects inclusive 1-based pages (PDF/PPTX) or explicit-page-break segments (DOCX; rejected when the file has none); every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Images are always extracted into images/ with relative links. Primary engine pymupdf4llm, fallback PyMuPDF text (degraded, marked), pure-JS unpdf only when no Python backend exists. Excel yields a sheet inventory (worksheets + chartsheets), a full CSV per non-empty worksheet under sheets/, a bounded preview with formulas and cached values, merged/hidden disclosure, and rendered chart views when LibreOffice is available. DOCX converts directly (mammoth; python-docx text fallback) and keeps heading styles, so the handle Outline lists `L<line>` and `p<segment>` per heading; LibreOffice (soffice) is optional for DOCX (fallback route, degraded) and Excel rendered views, and required for PPTX. Use `outputDir` for a durable bundle; without it the bundle lands in a per-call temp dir. Input must be a local file path (use fetch first for URLs). HTML converts without Readability (markdownify; Turndown fallback): local and data: images are copied into images/, remote images stay links. Pages without a text layer keep their picture in images/, or pages/ with `pageImages: true`; image inputs keep their picture in images/. `pageImages: true` writes page renders to pages/<stem>-pNNN, and email attachments are saved under attachments/. OCR is off by default: `ocr: true` adds recognized text (labeled as OCR) when Tesseract language data is installed; the handle's OCR: line says what ran.",
|
|
38
|
+
"Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, email (.msg/.eml), HTML (.html/.htm) or image (.png .jpg .jpeg .tif .tiff .bmp .gif) to a Markdown bundle on disk and return a handle (Saved-To, Images-Dir, Page-Count, Outline, diagnostics) - the Markdown itself is never inlined; read the Saved-To file (offset/limit) for content. `info: true` returns page count, metadata and TOC (or the sheet inventory) without converting - use it to pick `pages`. `pages` selects inclusive 1-based pages (PDF/PPTX) or explicit-page-break segments (DOCX; rejected when the file has none); every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Images are always extracted into images/ with relative links. Primary engine pymupdf4llm, fallback PyMuPDF text (degraded, marked), pure-JS unpdf only when no Python backend exists. Excel yields a sheet inventory (worksheets + chartsheets), a full CSV per non-empty worksheet under sheets/, a bounded preview with formulas and cached values, merged/hidden disclosure, and rendered chart views when LibreOffice is available. DOCX converts directly (mammoth; python-docx text fallback) and keeps heading styles, so the handle Outline lists `L<line>` and `p<segment>` per heading; LibreOffice (soffice) is optional for DOCX (fallback route, degraded) and Excel rendered views, and required for PPTX. Use `outputDir` for a durable bundle; without it the bundle lands in a per-call temp dir. Input must be a local file path (use fetch first for URLs). HTML converts without Readability (markdownify; Turndown fallback): local and data: images are copied into images/, remote images stay links. Pages without a text layer keep their picture in images/, or pages/ with `pageImages: true`; image inputs keep their picture in images/. `pageImages: true` writes page renders to pages/<stem>-pNNN, and email attachments are saved under attachments/. OCR is off by default: `ocr: true` adds recognized text (labeled as OCR) when Tesseract language data is installed; the handle's OCR: line says what ran. Two-pass OCR: read Page-Stats first, then re-run with ocr: true, ocrMode: \"all\" and an explicit pages selection to get ocr/ sidecars for the pages you name.",
|
|
39
39
|
promptSnippet: "Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS/MSG/EML/HTML/image to a Markdown bundle (handle returned; read Saved-To)",
|
|
40
40
|
parameters: buildSchema(),
|
|
41
41
|
|
package/lib/doc-to-md-bundle.ts
CHANGED
|
@@ -1,23 +1,26 @@
|
|
|
1
1
|
/**
|
|
2
2
|
* Bundle protocol: a call owns `<stem>` for its whole duration via `<stem>.md.lock`;
|
|
3
|
-
* children stage assets under `images/`, `sheets/`, `pages/`, and `
|
|
4
|
-
* Node publishes them to stem-prefixed files in those
|
|
3
|
+
* children stage assets under `images/`, `sheets/`, `pages/`, `attachments/`, and `ocr/` staging dirs;
|
|
4
|
+
* Node publishes them to stem-prefixed files in those asset dirs and records every file it
|
|
5
5
|
* wrote in a manifest, and commits `<stem>.md` atomically (tmp + rename).
|
|
6
6
|
*/
|
|
7
7
|
import fs, { closeSync, existsSync, mkdirSync, openSync, readdirSync, readFileSync, renameSync, rmSync, statSync, unlinkSync, writeFileSync } from "node:fs";
|
|
8
8
|
import { tmpdir } from "node:os";
|
|
9
9
|
import { extname, join, resolve } from "node:path";
|
|
10
10
|
import { randomBytes } from "node:crypto";
|
|
11
|
+
import type { PageStat } from "./doc-to-md-handle.ts";
|
|
11
12
|
|
|
12
13
|
export interface Bundle {
|
|
13
14
|
root: string; stem: string; renamedFrom: string | null; renameReason: string | null; mdPath: string; lockPath: string; imagesDir: string; stagingDir: string; lockId: string;
|
|
14
15
|
sheetsDir: string; sheetsStagingDir: string;
|
|
15
16
|
pagesDir: string; pagesStagingDir: string;
|
|
16
17
|
attachmentsDir: string; attachmentsStagingDir: string;
|
|
18
|
+
ocrDir: string; ocrStagingDir: string; pageStatsPath: string; wordsPath: string;
|
|
17
19
|
manifest: Set<string>;
|
|
18
20
|
csvManifest: Set<string>;
|
|
19
21
|
pageManifest: Set<string>;
|
|
20
22
|
attachmentManifest: Set<string>;
|
|
23
|
+
ocrManifest: Set<string>;
|
|
21
24
|
sourceMap: Map<string, string>;
|
|
22
25
|
}
|
|
23
26
|
|
|
@@ -33,6 +36,8 @@ export function ownedCsvPattern(stem: string): RegExp {
|
|
|
33
36
|
|
|
34
37
|
export function ownedPagePattern(stem: string): RegExp { return new RegExp(`^${escRe(stem)}-p\\d+\\.[a-z0-9]+$`); }
|
|
35
38
|
|
|
39
|
+
export function ownedOcrPattern(stem: string): RegExp { return new RegExp(`^${escRe(stem)}-p\\d+(?:\\.words\\.json|\\.md)$`); }
|
|
40
|
+
|
|
36
41
|
const FILE_LINK_RE = /\[[^\]]*\]\(\s*((?:sheets|attachments)\/[^)\s]+)\s*\)/g;
|
|
37
42
|
const IMG_LINK_RE = /!\[[^\]]*\]\(\s*(?:<([^>]*)>|([^)]*?))\s*\)|<img\b[^>]*\bsrc\s*=\s*(?:"([^"]*)"|'([^']*)'|([^\s>"']+))/gi;
|
|
38
43
|
|
|
@@ -75,6 +80,7 @@ export function openBundle(root: string, requested: string, overwrite: boolean):
|
|
|
75
80
|
const sheetsDir = join(root, "sheets");
|
|
76
81
|
const pagesDir = join(root, "pages");
|
|
77
82
|
const attachmentsDir = join(root, "attachments");
|
|
83
|
+
const ocrDir = join(root, "ocr");
|
|
78
84
|
try {
|
|
79
85
|
if (existsSync(mdPath) && overwrite) {
|
|
80
86
|
const owned = ownedPattern(stem), ownedCsv = ownedCsvPattern(stem), ownedPage = ownedPagePattern(stem);
|
|
@@ -87,6 +93,10 @@ export function openBundle(root: string, requested: string, overwrite: boolean):
|
|
|
87
93
|
if (existsSync(sheetsDir)) for (const f of readdirSync(sheetsDir)) if (ownedCsv.test(f)) rmSync(join(sheetsDir, f), { force: true });
|
|
88
94
|
if (existsSync(pagesDir)) for (const f of readdirSync(pagesDir)) if (ownedPage.test(f)) rmSync(join(pagesDir, f), { force: true });
|
|
89
95
|
if (existsSync(attachmentsDir)) for (const f of readdirSync(attachmentsDir)) if (attachments.has(f)) rmSync(join(attachmentsDir, f), { force: true });
|
|
96
|
+
rmSync(join(root, `${stem}.pages.json`), { force: true });
|
|
97
|
+
rmSync(join(root, `${stem}.words.json`), { force: true });
|
|
98
|
+
const ownedOcr = ownedOcrPattern(stem);
|
|
99
|
+
if (existsSync(ocrDir)) for (const f of readdirSync(ocrDir)) if (ownedOcr.test(f)) rmSync(join(ocrDir, f), { force: true });
|
|
90
100
|
}
|
|
91
101
|
const lockId = randomBytes(6).toString("hex");
|
|
92
102
|
const stagingDir = join(imagesDir, `.stage-${lockId}`);
|
|
@@ -94,7 +104,8 @@ export function openBundle(root: string, requested: string, overwrite: boolean):
|
|
|
94
104
|
mkdirSync(stagingDir, { recursive: true });
|
|
95
105
|
const pagesStagingDir = join(pagesDir, `.stage-${lockId}`);
|
|
96
106
|
const attachmentsStagingDir = join(attachmentsDir, `.stage-${lockId}`);
|
|
97
|
-
|
|
107
|
+
const ocrStagingDir = join(ocrDir, `.stage-${lockId}`);
|
|
108
|
+
return { root, stem, renamedFrom, renameReason, mdPath, lockPath, imagesDir, stagingDir, lockId, sheetsDir, sheetsStagingDir, pagesDir, pagesStagingDir, attachmentsDir, attachmentsStagingDir, ocrDir, ocrStagingDir, pageStatsPath: join(root, `${stem}.pages.json`), wordsPath: join(root, `${stem}.words.json`), ocrManifest: new Set(), manifest: new Set(), csvManifest: new Set(), pageManifest: new Set(), attachmentManifest: new Set(), sourceMap: new Map() };
|
|
98
109
|
} catch (e) { rmSync(lockPath, { force: true }); throw e; }
|
|
99
110
|
}
|
|
100
111
|
|
|
@@ -160,6 +171,49 @@ export function publishPageImages(b: Bundle, pageCount: number): void {
|
|
|
160
171
|
rmSync(b.pagesStagingDir, { recursive: true, force: true });
|
|
161
172
|
}
|
|
162
173
|
|
|
174
|
+
export function writePageStats(b: Bundle, stats: PageStat[]): void {
|
|
175
|
+
writeFileSync(b.pageStatsPath, `${JSON.stringify(stats, null, 2)}\n`, "utf8");
|
|
176
|
+
}
|
|
177
|
+
|
|
178
|
+
/** Rename within the bundle root keeps the complete words document atomic. */
|
|
179
|
+
export function publishWords(b: Bundle): string | null {
|
|
180
|
+
const staged = join(b.stagingDir, "words.json");
|
|
181
|
+
try {
|
|
182
|
+
if (!existsSync(staged)) throw new Error("child staged no words.json");
|
|
183
|
+
renameSync(staged, b.wordsPath);
|
|
184
|
+
return null;
|
|
185
|
+
} catch (e) {
|
|
186
|
+
rmSync(b.wordsPath, { force: true });
|
|
187
|
+
return `write failed - ${(e as Error).message}`;
|
|
188
|
+
}
|
|
189
|
+
}
|
|
190
|
+
|
|
191
|
+
/** Move `.done`-gated Markdown and words sidecars into `ocr/`; drop partial dirs and the checkpoint. Returns page -> absolute paths for each kind. */
|
|
192
|
+
export function publishSidecars(b: Bundle): { sidecars: Map<number, string>; wordSidecars: Map<number, string> } {
|
|
193
|
+
const sidecars = new Map<number, string>(), wordSidecars = new Map<number, string>();
|
|
194
|
+
if (!existsSync(b.ocrStagingDir)) return { sidecars, wordSidecars };
|
|
195
|
+
for (const dir of readdirSync(b.ocrStagingDir).sort()) {
|
|
196
|
+
const m = dir.match(/^p(\d+)$/);
|
|
197
|
+
if (!m) continue;
|
|
198
|
+
const pageDir = join(b.ocrStagingDir, dir);
|
|
199
|
+
if (!existsSync(join(pageDir, ".done"))) continue;
|
|
200
|
+
const file = `${b.stem}-${dir}.md`;
|
|
201
|
+
if (!existsSync(join(pageDir, file))) continue;
|
|
202
|
+
mkdirSync(b.ocrDir, { recursive: true });
|
|
203
|
+
renameSync(join(pageDir, file), join(b.ocrDir, file));
|
|
204
|
+
b.ocrManifest.add(file);
|
|
205
|
+
sidecars.set(Number(m[1]), join(b.ocrDir, file));
|
|
206
|
+
const wfile = `${b.stem}-${dir}.words.json`;
|
|
207
|
+
if (existsSync(join(pageDir, wfile))) {
|
|
208
|
+
renameSync(join(pageDir, wfile), join(b.ocrDir, wfile));
|
|
209
|
+
b.ocrManifest.add(wfile);
|
|
210
|
+
wordSidecars.set(Number(m[1]), join(b.ocrDir, wfile));
|
|
211
|
+
}
|
|
212
|
+
}
|
|
213
|
+
rmSync(b.ocrStagingDir, { recursive: true, force: true });
|
|
214
|
+
return { sidecars, wordSidecars };
|
|
215
|
+
}
|
|
216
|
+
|
|
163
217
|
export function publishAttachments(b: Bundle): void {
|
|
164
218
|
if (!existsSync(b.attachmentsStagingDir)) return;
|
|
165
219
|
for (const f of readdirSync(b.attachmentsStagingDir).sort()) {
|
|
@@ -208,6 +262,7 @@ export function commitBundle(b: Bundle, markdown: string): void {
|
|
|
208
262
|
try { fs.rmSync(b.sheetsStagingDir, { recursive: true, force: true }); } catch { /* best-effort */ }
|
|
209
263
|
try { fs.rmSync(b.pagesStagingDir, { recursive: true, force: true }); } catch { /* best-effort */ }
|
|
210
264
|
try { fs.rmSync(b.attachmentsStagingDir, { recursive: true, force: true }); } catch { /* best-effort */ }
|
|
265
|
+
try { fs.rmSync(b.ocrStagingDir, { recursive: true, force: true }); } catch { /* best-effort */ }
|
|
211
266
|
try { fs.rmSync(b.lockPath, { force: true }); } catch { /* Markdown is published; cleanup is best-effort. */ }
|
|
212
267
|
}
|
|
213
268
|
|
|
@@ -216,6 +271,10 @@ export function abortBundle(b: Bundle): void {
|
|
|
216
271
|
for (const f of b.csvManifest) rmSync(join(b.sheetsDir, f), { force: true });
|
|
217
272
|
for (const f of b.pageManifest) rmSync(join(b.pagesDir, f), { force: true });
|
|
218
273
|
for (const f of b.attachmentManifest) rmSync(join(b.attachmentsDir, f), { force: true });
|
|
274
|
+
for (const f of b.ocrManifest) rmSync(join(b.ocrDir, f), { force: true });
|
|
275
|
+
rmSync(b.ocrStagingDir, { recursive: true, force: true });
|
|
276
|
+
rmSync(b.pageStatsPath, { force: true });
|
|
277
|
+
rmSync(b.wordsPath, { force: true });
|
|
219
278
|
rmSync(`${b.mdPath}.tmp`, { force: true });
|
|
220
279
|
rmSync(b.stagingDir, { recursive: true, force: true });
|
|
221
280
|
rmSync(b.sheetsStagingDir, { recursive: true, force: true });
|
package/lib/doc-to-md-core.ts
CHANGED
|
@@ -12,9 +12,9 @@ import { type ChildProcess, spawn } from "node:child_process";
|
|
|
12
12
|
import { homedir, tmpdir } from "node:os";
|
|
13
13
|
import { basename, dirname, extname, isAbsolute, join, resolve } from "node:path";
|
|
14
14
|
import { fileURLToPath, pathToFileURL } from "node:url";
|
|
15
|
-
import { type Bundle, abortBundle, commitBundle, openBundle, publishAttachments, publishPageImages, publishSheetCsvs, publishSheetImages, publishStaged, rewriteLinks, tempBundleRoot, validateImageLinks } from "./doc-to-md-bundle.ts";
|
|
16
|
-
import { type Engine, type HandleData, type InfoData, type OcrInfo, type Tier, formatHandle, formatInfoHandle, scanOutline, type SheetInfo, type TocEntry } from "./doc-to-md-handle.ts";
|
|
17
|
-
import { type DocToMdOptions, type InputType, IMAGE_EXTS, TUNABLE_DEFAULTS, classifyInput, sanitizeStem } from "./doc-to-md-options.ts";
|
|
15
|
+
import { type Bundle, abortBundle, commitBundle, openBundle, publishAttachments, publishSidecars, publishWords, writePageStats, publishPageImages, publishSheetCsvs, publishSheetImages, publishStaged, rewriteLinks, tempBundleRoot, validateImageLinks } from "./doc-to-md-bundle.ts";
|
|
16
|
+
import { type Engine, type HandleData, type InfoData, type OcrInfo, type PageStat, type Tier, formatHandle, formatInfoHandle, scanOutline, type SheetInfo, type TocEntry } from "./doc-to-md-handle.ts";
|
|
17
|
+
import { type DocToMdOptions, type InputType, IMAGE_EXTS, TUNABLE_DEFAULTS, UsageError, classifyInput, sanitizeStem } from "./doc-to-md-options.ts";
|
|
18
18
|
|
|
19
19
|
export * from "./doc-to-md-options.ts";
|
|
20
20
|
export { compactRanges, formatHandle, formatInfoHandle, formatSize, scanOutline } from "./doc-to-md-handle.ts";
|
|
@@ -466,8 +466,8 @@ const lacksDocx = (b: Backend) => b.kind === "none" || !b.docx;
|
|
|
466
466
|
const clearStaging = (b: Pick<Bundle, "stagingDir">) => { for (const f of readdirSync(b.stagingDir)) rmSync(join(b.stagingDir, f), { recursive: true, force: true }); };
|
|
467
467
|
export const EXCEL_REMEDY = "Remedy: install uv, or pip install openpyxl xlrd pillow";
|
|
468
468
|
|
|
469
|
-
export type Mode = "html" | "image" | "info" | "pdf-primary" | "pdf-fallback" | "xlsx" | "pdf-text" | "render-pages" | "docx" | "email";
|
|
470
|
-
export interface TierJson { pageImages?: { page: number; file: string }[]; ocr?: OcrInfo; markdown?: string; pages?: number[]; pageCount?: number; emptyPages?: number[]; failedPages?: { page: number; error: string }[]; notes?: string[]; images?: { sheetIndex: number; file: string }[]; metadata?: Record<string, string>; toc?: [number, string, number | null][]; explicitBreaks?: number; engine?: string; degraded?: boolean; fallbackReason?: string | null; sheets?: SheetInfo[]; renderPages?: number[]; sheetCount?: number; ok?: boolean; reason?: string; rendered?: { idx: number; file: string; dpi: number }[]; failed?: { idx: number; reason: string }[]; }
|
|
469
|
+
export type Mode = "html" | "image" | "info" | "pdf-primary" | "pdf-fallback" | "xlsx" | "pdf-text" | "render-pages" | "docx" | "email" | "ocr-pages";
|
|
470
|
+
export interface TierJson { words?: boolean; wordsErrors?: Record<string, string>; pageStats?: PageStat[]; status?: string; written?: number[]; noText?: number[]; ocrFailed?: number[]; ocrErrors?: Record<string, string>; budgetStopped?: number[]; pageImages?: { page: number; file: string }[]; ocr?: OcrInfo; markdown?: string; pages?: number[]; pageCount?: number; emptyPages?: number[]; failedPages?: { page: number; error: string }[]; notes?: string[]; images?: { sheetIndex: number; file: string }[]; metadata?: Record<string, string>; toc?: [number, string, number | null][]; explicitBreaks?: number; engine?: string; degraded?: boolean; fallbackReason?: string | null; sheets?: SheetInfo[]; renderPages?: number[]; sheetCount?: number; ok?: boolean; reason?: string; rendered?: { idx: number; file: string; dpi: number }[]; failed?: { idx: number; reason: string }[]; }
|
|
471
471
|
export type TierResult = { ok: true; json: TierJson } | { ok: false; reason: string; detail?: string } | { ok: false; userError: string; pageCount?: number };
|
|
472
472
|
|
|
473
473
|
export interface PipelineSeams {
|
|
@@ -490,7 +490,7 @@ function unpdfWorkerPath(): string {
|
|
|
490
490
|
], existsSync);
|
|
491
491
|
}
|
|
492
492
|
|
|
493
|
-
async function runTierReal(mode: Mode, childOptions: Record<string, unknown>, _b: Pick<Bundle, "stagingDir">, signal: AbortSignal | undefined, timeoutMs: number, backend: Backend): Promise<TierResult> {
|
|
493
|
+
export async function runTierReal(mode: Mode, childOptions: Record<string, unknown>, _b: Pick<Bundle, "stagingDir">, signal: AbortSignal | undefined, timeoutMs: number, backend: Backend): Promise<TierResult> {
|
|
494
494
|
const cfg = { pymupdfVersion: String(childOptions.pymupdfVersion), warmTimeoutMs: 0 };
|
|
495
495
|
let cmd: string, args: string[];
|
|
496
496
|
if (mode === "pdf-text" || (mode === "info" && backend.kind === "none")) { cmd = process.execPath; args = [unpdfWorkerPath(), mode]; }
|
|
@@ -536,7 +536,39 @@ export function reconcileRenderMarkers(md: string, renderPages: number[], fmt: s
|
|
|
536
536
|
return md;
|
|
537
537
|
}
|
|
538
538
|
|
|
539
|
-
export const emptyOcr = (lang: string): OcrInfo => ({ status: "off", lang, textless: [], pages: [], noText: [], ocrFailed: [], budgetStopped: [], reason: null, tesseract: null });
|
|
539
|
+
export const emptyOcr = (lang: string): OcrInfo => ({ status: "off", lang, textless: [], pages: [], noText: [], ocrFailed: [], budgetStopped: [], reason: null, tesseract: null, mode: "textless", sidecars: {}, wordSidecars: {}, ocrErrors: {}, killed: null, notAttempted: [], childError: null });
|
|
540
|
+
|
|
541
|
+
export interface OcrPagesOutcome { written: number[]; noText: number[]; ocrFailed: number[]; ocrErrors: Record<number, string>; budgetStopped: number[]; killed: number | null; notAttempted: number[]; childError: string | null; }
|
|
542
|
+
const emptyOutcome = (): OcrPagesOutcome => ({ written: [], noText: [], ocrFailed: [], ocrErrors: {}, budgetStopped: [], killed: null, notAttempted: [], childError: null });
|
|
543
|
+
const ocrPagesTag = (n: number) => `p${String(n).padStart(3, "0")}`;
|
|
544
|
+
const SIDE_MARKER_RE = /^--- end of page\.page_number=\d+ ---$/;
|
|
545
|
+
|
|
546
|
+
function sidecarHasText(dir: string): boolean {
|
|
547
|
+
const file = readdirSync(dir).find((f) => f.endsWith(".md"));
|
|
548
|
+
if (!file) return false;
|
|
549
|
+
return readFileSync(join(dir, file), "utf8").split("\n").slice(1).some((l) => l.trim() && !SIDE_MARKER_RE.test(l));
|
|
550
|
+
}
|
|
551
|
+
|
|
552
|
+
/** Rebuild the outcome from staging markers after the child died or returned garbage. */
|
|
553
|
+
export function recoverOcrPages(stagingDir: string, pages: number[], detail: string): OcrPagesOutcome {
|
|
554
|
+
const out = emptyOutcome();
|
|
555
|
+
const activePath = join(stagingDir, "active");
|
|
556
|
+
const active = existsSync(activePath) ? Number(readFileSync(activePath, "utf8").trim()) || null : null;
|
|
557
|
+
let sawPage = false;
|
|
558
|
+
for (const n of pages) {
|
|
559
|
+
const dir = join(stagingDir, ocrPagesTag(n));
|
|
560
|
+
if (existsSync(join(dir, ".done"))) { sawPage = true; (sidecarHasText(dir) ? out.written : out.noText).push(n); }
|
|
561
|
+
else if (existsSync(join(dir, ".failed"))) { sawPage = true; out.ocrFailed.push(n); out.ocrErrors[n] = readFileSync(join(dir, ".failed"), "utf8").trim() || "unknown error"; }
|
|
562
|
+
else if (n === active) out.killed = n;
|
|
563
|
+
else out.notAttempted.push(n);
|
|
564
|
+
}
|
|
565
|
+
if (active === null && !sawPage) out.childError = detail;
|
|
566
|
+
return out;
|
|
567
|
+
}
|
|
568
|
+
|
|
569
|
+
function outcomeFromChild(j: TierJson): OcrPagesOutcome {
|
|
570
|
+
return { ...emptyOutcome(), written: j.written ?? [], noText: j.noText ?? [], ocrFailed: j.ocrFailed ?? [], ocrErrors: j.ocrErrors ?? {}, budgetStopped: j.budgetStopped ?? [] };
|
|
571
|
+
}
|
|
540
572
|
const OCR_SENTINEL_RE = /\x00OCR ([^\x00]*)\x00/g;
|
|
541
573
|
|
|
542
574
|
/** Child OCR labels name staged files; the published name exists only after publishStaged. */
|
|
@@ -551,7 +583,7 @@ function handleOcr(tier: Tier, type: InputType, o: DocToMdOptions, json: TierJso
|
|
|
551
583
|
if (tier === "unpdf") return o.ocr ? { ...emptyOcr(o.ocrLanguage), status: "unavailable", reason: "no Python backend" } : null;
|
|
552
584
|
const x = json.ocr;
|
|
553
585
|
if (!x || (type !== "image" && !x.textless.length && !x.pages.length && !x.ocrFailed.length)) return null;
|
|
554
|
-
return x;
|
|
586
|
+
return { ...emptyOcr(o.ocrLanguage), ...x };
|
|
555
587
|
}
|
|
556
588
|
|
|
557
589
|
export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, seams?: Partial<PipelineSeams>): Promise<ConvertOutcome> {
|
|
@@ -561,10 +593,17 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
|
|
|
561
593
|
if (!st || !st.isFile()) throw new Error(`Not a readable file: ${o.path}`);
|
|
562
594
|
if (st.size === 0) throw new Error(`empty file: ${o.path}`);
|
|
563
595
|
const type = classifyInput(inputPath);
|
|
596
|
+
const forced = o.ocrMode === "all";
|
|
597
|
+
if (forced) {
|
|
598
|
+
if (!o.ocr) throw new UsageError("--ocr-mode all requires --ocr");
|
|
599
|
+
if (o.pages === null) throw new UsageError('--ocr-mode all requires an explicit --pages selection (e.g. --pages 2,7); omitted pages and --pages "" mean all pages and are refused to keep OCR cost bounded');
|
|
600
|
+
if (type !== "pdf" && type !== "pptx" && type !== "doc") throw new UsageError("--ocr-mode all applies to PDF, PPTX and DOC inputs only (DOCX pages are page-break segments, not PDF pages; convert the DOCX to PDF first)");
|
|
601
|
+
}
|
|
564
602
|
if (o.pages && (type === "html" || type === "image" || type === "email")) throw new Error(`--pages does not apply to ${type === "html" ? "HTML files" : type === "image" ? "images" : "email"}`);
|
|
565
603
|
const isExcel = type === "xlsx" || type === "xlsm" || type === "xls";
|
|
566
604
|
if (isExcel && o.pages) throw new Error("--pages does not apply to spreadsheets: worksheets have no stable page numbering");
|
|
567
605
|
const backend = await s.backend({ pymupdfVersion: o.pymupdfVersion, warmTimeoutMs: o.warmTimeoutMs });
|
|
606
|
+
if (forced && backend.kind === "none") throw new Error(`--ocr-mode all cannot run: no Python backend (${backend.reason})`);
|
|
568
607
|
if (isExcel && (backend.kind === "none" || !backend.xlsx)) throw new Error(`Excel conversion needs a Python backend with openpyxl, xlrd and pillow (${backend.kind === "none" ? backend.reason : `${backend.kind} lacks the Excel packages`}). ${EXCEL_REMEDY}`);
|
|
569
608
|
if (type === "email" && extname(inputPath).toLowerCase() === ".msg" && (backend.kind === "none" || !backend.email)) throw new Error(`MSG conversion needs the extract-msg package. Python backend: ${backendState(backend, "found without extract-msg")}. Remedy: install uv, or pip install extract-msg markdownify into that Python`);
|
|
570
609
|
if (type === "email" && extname(inputPath).toLowerCase() !== ".msg" && lacksDocx(backend)) throw new Error(`EML conversion needs the Python DOCX/HTML packages (mammoth, markdownify, python-docx). Python backend: ${backendState(backend, "found without mammoth/markdownify/python-docx")}. ${docxRemedy}`);
|
|
@@ -573,11 +612,12 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
|
|
|
573
612
|
let office: { pdfPath: string; cleanup: () => void } | null = null;
|
|
574
613
|
try {
|
|
575
614
|
let pdfPath = inputPath;
|
|
576
|
-
const base = { path: inputPath, pages: o.pages, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
|
|
615
|
+
const base = { path: inputPath, pages: o.pages, ...(o.words && (type === "pdf" || type === "image") ? { words: true } : {}), stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr && !forced, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
|
|
577
616
|
let tier: Tier | undefined, engine: Engine | undefined, json: TierJson | undefined, degraded: string | null = null, fallbackReason: string | null = null;
|
|
578
617
|
let explicitBreaks: number | null = null;
|
|
579
618
|
let notes: string[] = [];
|
|
580
619
|
let officeRoute: string | null = null;
|
|
620
|
+
let copyReason: string | null = null;
|
|
581
621
|
if (type === "html") {
|
|
582
622
|
const prepared = await prepareHtml(inputPath, b.stagingDir);
|
|
583
623
|
if (signal?.aborted) throw new Error("aborted");
|
|
@@ -603,6 +643,7 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
|
|
|
603
643
|
}
|
|
604
644
|
if (signal?.aborted) throw new Error("aborted");
|
|
605
645
|
if (json === undefined) {
|
|
646
|
+
copyReason = reason;
|
|
606
647
|
clearStaging(b);
|
|
607
648
|
const file = `original${extname(inputPath).toLowerCase()}`;
|
|
608
649
|
const dir = join(b.stagingDir, "p1");
|
|
@@ -704,6 +745,35 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
|
|
|
704
745
|
if (type === "doc" || officeRoute !== null) degraded = DEGRADED_DOCX_OFFICE;
|
|
705
746
|
if (officeRoute !== null) fallbackReason = fallbackReason ? `${officeRoute}; ${fallbackReason}` : officeRoute;
|
|
706
747
|
if (tier === undefined || engine === undefined || json === undefined) throw new Error("internal: no tier produced output");
|
|
748
|
+
const pageStats = json.pageStats ?? null;
|
|
749
|
+
if (pageStats) writePageStats(b, pageStats);
|
|
750
|
+
let wordsPath: string | null = null, wordsReason: string | null = null;
|
|
751
|
+
const wordsErrors: Record<number, string> = {};
|
|
752
|
+
const takeWordsErrors = (j: TierJson | undefined) => {
|
|
753
|
+
for (const [k, v] of Object.entries(j?.wordsErrors ?? {})) {
|
|
754
|
+
const page = Number(k);
|
|
755
|
+
if (Number.isFinite(page)) wordsErrors[page] = v;
|
|
756
|
+
}
|
|
757
|
+
};
|
|
758
|
+
if (o.words) {
|
|
759
|
+
if (type !== "pdf" && type !== "image") wordsReason = `none - word positions apply to PDF and image inputs only (${type})`;
|
|
760
|
+
else if (tier === "unpdf") wordsReason = "none - unpdf tier has no page geometry";
|
|
761
|
+
else if (engine === "copy") wordsReason = `none - image copied without conversion (${copyReason})`;
|
|
762
|
+
else if (json?.words === true) { wordsReason = publishWords(b); if (wordsReason === null) wordsPath = b.wordsPath; }
|
|
763
|
+
else wordsReason = `write failed - ${json?.wordsErrors?.file ?? "child reported no words document"}`;
|
|
764
|
+
takeWordsErrors(json);
|
|
765
|
+
}
|
|
766
|
+
let ocr = handleOcr(tier, type, o, json);
|
|
767
|
+
if (forced) {
|
|
768
|
+
const r = await s.runTier("ocr-pages", { path: pdfPath, pages: o.pages, ...(o.words && type === "pdf" ? { words: true } : {}), stem: b.stem, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs, stagingDir: b.ocrStagingDir, dpi: o.imageDpi, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion }, b, signal, o.primaryTimeoutMs, backend);
|
|
769
|
+
if (signal?.aborted || (!r.ok && "reason" in r && r.reason === "aborted")) throw new Error("aborted");
|
|
770
|
+
if (r.ok && r.json.status === "unavailable") throw new Error(`OCR unavailable: ${r.json.reason} (install Tesseract; see doc/doc-to-md.md)`);
|
|
771
|
+
const outcome = r.ok && r.json.status === "ran" ? outcomeFromChild(r.json)
|
|
772
|
+
: recoverOcrPages(b.ocrStagingDir, o.pages!, !r.ok ? ("userError" in r ? r.userError : `${r.reason}${detailSuffix(r)}`) : "malformed child output");
|
|
773
|
+
const { sidecars, wordSidecars } = publishSidecars(b);
|
|
774
|
+
ocr = { ...emptyOcr(o.ocrLanguage), ...json.ocr, status: "ran", reason: null, mode: "all", pages: outcome.written, noText: outcome.noText, ocrFailed: outcome.ocrFailed, ocrErrors: outcome.ocrErrors, budgetStopped: outcome.budgetStopped, killed: outcome.killed, notAttempted: outcome.notAttempted, childError: outcome.childError, sidecars: Object.fromEntries(sidecars), wordSidecars: Object.fromEntries(wordSidecars) };
|
|
775
|
+
if (o.words && r.ok) takeWordsErrors(r.json);
|
|
776
|
+
}
|
|
707
777
|
if (!isExcel) notes = [...notes, ...(json.notes ?? [])];
|
|
708
778
|
if (b.renamedFrom) notes.splice(notes[0]?.startsWith("preview truncated:") ? 1 : 0, 0, `renamed to ${b.stem} (${b.renameReason})`);
|
|
709
779
|
const pageImagesReason = !o.pageImages || b.pageManifest.size || tier === "primary" || tier === "fallback" ? null
|
|
@@ -719,7 +789,7 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
|
|
|
719
789
|
const markdown = (head.length ? `${head.join("\n")}\n\n` : "") + body;
|
|
720
790
|
commitBundle(b, markdown);
|
|
721
791
|
const outline = scanOutline(markdown, o.outlineMaxEntries);
|
|
722
|
-
const details: DocToMdDetails = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr:
|
|
792
|
+
const details: DocToMdDetails = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr, pageStats, pageStatsPath: pageStats ? b.pageStatsPath : null, ocrDir: b.ocrManifest.size ? b.ocrDir : null, wordsPath, wordsReason, wordsErrors };
|
|
723
793
|
return { output: formatHandle(details), details };
|
|
724
794
|
} catch (e) { abortBundle(b); throw e; }
|
|
725
795
|
finally { office?.cleanup(); }
|
package/lib/doc-to-md-handle.ts
CHANGED
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
import type { InputType } from "./doc-to-md-options.ts";
|
|
1
|
+
import type { InputType, OcrMode } from "./doc-to-md-options.ts";
|
|
2
2
|
|
|
3
3
|
export type Tier = "primary" | "fallback" | "unpdf" | "excel" | "docx" | "html" | "image" | "email";
|
|
4
4
|
export type Engine = "pymupdf4llm" | "pymupdf-text" | "unpdf" | "openpyxl" | "xlrd" | "mammoth" | "python-docx" | "markdownify" | "turndown" | "copy" | "extract-msg" | "email";
|
|
@@ -18,6 +18,8 @@ export interface SheetInfo {
|
|
|
18
18
|
csv: string | null;
|
|
19
19
|
}
|
|
20
20
|
|
|
21
|
+
export type PageStat = { page: number; chars: number; images: number; imageCoverage: number } | { page: number; error: string };
|
|
22
|
+
|
|
21
23
|
export interface OcrInfo {
|
|
22
24
|
status: "off" | "unavailable" | "skipped" | "ran";
|
|
23
25
|
lang: string;
|
|
@@ -28,6 +30,13 @@ export interface OcrInfo {
|
|
|
28
30
|
budgetStopped: number[];
|
|
29
31
|
reason: string | null;
|
|
30
32
|
tesseract: boolean | null;
|
|
33
|
+
mode: OcrMode;
|
|
34
|
+
sidecars: Record<number, string>;
|
|
35
|
+
wordSidecars: Record<number, string>;
|
|
36
|
+
ocrErrors: Record<number, string>;
|
|
37
|
+
killed: number | null;
|
|
38
|
+
notAttempted: number[];
|
|
39
|
+
childError: string | null;
|
|
31
40
|
}
|
|
32
41
|
|
|
33
42
|
export interface HandleData {
|
|
@@ -35,6 +44,8 @@ export interface HandleData {
|
|
|
35
44
|
pageCount: number | null; pages: number[] | null; explicitBreaks: number | null; imageCount: number; pageImageCount: number; pageImagesReason: string | null; bytes: number; lines: number;
|
|
36
45
|
degraded: string | null; fallbackReason: string | null; failedPages: number[]; emptyPages: number[];
|
|
37
46
|
notes: string[]; outline: OutlineEntry[]; outlineTotal: number; ocr: OcrInfo | null;
|
|
47
|
+
pageStats: PageStat[] | null; pageStatsPath: string | null; ocrDir: string | null;
|
|
48
|
+
wordsPath: string | null; wordsReason: string | null; wordsErrors: Record<number, string>;
|
|
38
49
|
}
|
|
39
50
|
|
|
40
51
|
export interface InfoData {
|
|
@@ -72,7 +83,23 @@ export function compactRanges(nums: number[], maxEntries = 20): string {
|
|
|
72
83
|
const INSTALL_HINT = "install Tesseract (see doc/doc-to-md.md), then rerun with ocr=true";
|
|
73
84
|
const BARE_REASONS = ["fallback tier", "no Python backend"];
|
|
74
85
|
|
|
86
|
+
const pageList = (pages: number[]) => `page${pages.length === 1 ? "" : "s"} ${pages.join(", ")}`;
|
|
87
|
+
|
|
88
|
+
function forcedOcrLine(ocr: OcrInfo): string {
|
|
89
|
+
const clauses: string[] = [];
|
|
90
|
+
if (ocr.pages.length) clauses.push(`sidecars for ${pageList(ocr.pages)}`);
|
|
91
|
+
if (ocr.noText.length) clauses.push(`no text on ${pageList(ocr.noText)}`);
|
|
92
|
+
for (const page of ocr.ocrFailed) clauses.push(`failed on page ${page} (${ocr.ocrErrors[page] ?? "unknown error"})`);
|
|
93
|
+
if (ocr.budgetStopped.length) clauses.push(`budget-stopped ${pageList(ocr.budgetStopped)}`);
|
|
94
|
+
if (ocr.killed !== null) clauses.push(`page ${ocr.killed} killed the OCR child (timeout or crash - likely a compression bomb)`);
|
|
95
|
+
if (ocr.childError !== null) clauses.push(`OCR child failed before processing pages: ${ocr.childError}`);
|
|
96
|
+
if (ocr.notAttempted.length) clauses.push(`${pageList(ocr.notAttempted)} not attempted`);
|
|
97
|
+
const rerun = [...ocr.budgetStopped, ...ocr.notAttempted].sort((a, b) => a - b);
|
|
98
|
+
return `OCR: forced (${ocr.lang}) - ${clauses.join("; ")}${rerun.length ? ` - re-run with --pages ${rerun.join(",")}` : ""}`;
|
|
99
|
+
}
|
|
100
|
+
|
|
75
101
|
export function ocrLine(ocr: OcrInfo, type: InputType): string {
|
|
102
|
+
if (ocr.mode === "all") return forcedOcrLine(ocr);
|
|
76
103
|
const image = type === "image";
|
|
77
104
|
const rerun = ocr.tesseract ? "rerun with ocr=true" : INSTALL_HINT;
|
|
78
105
|
switch (ocr.status) {
|
|
@@ -143,6 +170,12 @@ export function formatHandle(h: HandleData): string {
|
|
|
143
170
|
if (h.sheetsDir) lines.push(`Sheets-Dir: ${h.sheetsDir}`);
|
|
144
171
|
if (h.pagesDir && h.pageImageCount > 0) lines.push(`Pages-Dir: ${h.pagesDir} (${h.pageImageCount} pages)`);
|
|
145
172
|
else if (h.pageImagesReason) lines.push(`Pages-Dir: none - ${h.pageImagesReason}`);
|
|
173
|
+
if (h.pageStatsPath) lines.push(`Page-Stats: ${h.pageStatsPath}`);
|
|
174
|
+
if (h.wordsPath) {
|
|
175
|
+
const bad = Object.keys(h.wordsErrors).map(Number).sort((a, b) => a - b);
|
|
176
|
+
lines.push(`Words: ${h.wordsPath}${bad.length ? ` (extraction failed for pages ${compactRanges(bad)}: ${h.wordsErrors[bad[0]]})` : ""}`);
|
|
177
|
+
} else if (h.wordsReason) lines.push(`Words: ${h.wordsReason}`);
|
|
178
|
+
if (h.ocrDir) lines.push(`OCR-Dir: ${h.ocrDir}`);
|
|
146
179
|
lines.push(`Type: ${h.type} Engine: ${h.engine} Tier: ${h.tier}`);
|
|
147
180
|
lines.push(`Page-Count: ${pageCountLabel(h)} Pages: ${h.pages ? compactRanges(h.pages) : "all"} Images: ${h.imageCount} Size: ${formatSize(h.bytes)} / ${h.lines} lines`);
|
|
148
181
|
if (h.degraded) lines.push(`Degraded: ${h.degraded}`);
|
package/lib/doc-to-md-options.ts
CHANGED
|
@@ -6,6 +6,7 @@ import { extname } from "node:path";
|
|
|
6
6
|
|
|
7
7
|
export type InputType = "pdf" | "docx" | "doc" | "pptx" | "xlsx" | "xlsm" | "xls" | "html" | "image" | "email";
|
|
8
8
|
export type ImageFormat = "png" | "jpg";
|
|
9
|
+
export type OcrMode = "textless" | "all";
|
|
9
10
|
|
|
10
11
|
export interface Tunables {
|
|
11
12
|
primaryTimeoutMs: number;
|
|
@@ -29,6 +30,8 @@ export interface DocToMdOptions extends Tunables {
|
|
|
29
30
|
outputDir: string | null;
|
|
30
31
|
overwrite: boolean;
|
|
31
32
|
pageImages: boolean;
|
|
33
|
+
words: boolean;
|
|
34
|
+
ocrMode: OcrMode;
|
|
32
35
|
}
|
|
33
36
|
|
|
34
37
|
/** What adapters pass in: intents as raw strings/booleans, tunables optional. */
|
|
@@ -39,6 +42,8 @@ export interface PerCallInput extends Partial<Tunables> {
|
|
|
39
42
|
outputDir?: string | null;
|
|
40
43
|
overwrite?: boolean;
|
|
41
44
|
pageImages?: boolean;
|
|
45
|
+
words?: boolean;
|
|
46
|
+
ocrMode?: OcrMode;
|
|
42
47
|
}
|
|
43
48
|
|
|
44
49
|
export type DescriptorType = "string" | "int" | "bool" | "pages" | "enum" | "version" | "lang";
|
|
@@ -68,6 +73,8 @@ export const DOC_TO_MD_OPTIONS: readonly OptionDescriptor[] = [
|
|
|
68
73
|
{ key: "outputDir", flag: "--output-dir", type: "string", default: null, settable: false, help: "Bundle root for <stem>.md + images/; default a per-call temp dir. <stem> = basename without extension with [^A-Za-z0-9._-]+ -> _ (empty -> document); a second call on the same stem writes <stem>-2.md" },
|
|
69
74
|
{ key: "overwrite", flag: "--overwrite", type: "bool", default: false, settable: false, help: "Replace an existing completed <stem>.md bundle" },
|
|
70
75
|
{ key: "pageImages", flag: "--page-images", type: "bool", default: false, settable: false, help: "Also render every selected page to pages/<stem>-pNNN.<imageFormat> at imageDpi (PDF, PPTX, .doc, DOCX via LibreOffice); off by default" },
|
|
76
|
+
{ key: "words", flag: "--words", type: "bool", default: false, settable: false, help: "Write word positions: <stem>.words.json beside the Markdown lists every text-layer word of each selected page with its bbox (PDF points, top-left origin, display orientation; image inputs in source pixels) and the words inline OCR recognized, tagged source \"text\" or \"ocr\"; under --ocr-mode all the OCR words go to ocr/<stem>-pNNN.words.json beside each sidecar. Never triggers OCR. PDF and image inputs only." },
|
|
77
|
+
{ key: "ocrMode", flag: "--ocr-mode", type: "enum", default: "textless", settable: false, enumValues: ["textless", "all"], help: "OCR policy: textless (default) OCRs only pages with an empty text layer, inline; all OCRs every selected page and writes the recognized text to ocr/<stem>-pNNN.md sidecars, leaving the Markdown untouched. all requires --ocr and an explicit --pages selection (PDF, PPTX, DOC)." },
|
|
71
78
|
{ key: "primaryTimeoutMs", flag: "--primary-timeout", type: "int", default: 60000, settable: true, env: "PI_DOC_TO_MD_CONVERT_TIMEOUT_MS", help: "pymupdf4llm tier and DOCX child (docx mode); also the unpdf tier" },
|
|
72
79
|
{ key: "fallbackTimeoutMs", flag: "--fallback-timeout", type: "int", default: 30000, settable: true, help: "PyMuPDF get_text tier (including DOCX LibreOffice fallback); also PDF and DOCX info and Excel rendered views" },
|
|
73
80
|
{ key: "sofficeTimeoutMs", flag: "--soffice-timeout", type: "int", default: 120000, settable: true, env: "PI_DOC_TO_MD_SOFFICE_TIMEOUT_MS", help: "DOCX/PPTX -> PDF via LibreOffice; also Excel rendered views" },
|
|
@@ -177,12 +184,47 @@ export function resolveOptions(perCall: PerCallInput, settings: Partial<Tunables
|
|
|
177
184
|
out[d.key] = value;
|
|
178
185
|
}
|
|
179
186
|
const o = out as unknown as DocToMdOptions;
|
|
180
|
-
if (o.info && (o.pages !== null || o.outputDir !== null || o.overwrite || o.pageImages)) {
|
|
181
|
-
throw new UsageError("--info cannot be combined with --pages, --output-dir, --overwrite or --
|
|
187
|
+
if (o.info && (o.pages !== null || o.outputDir !== null || o.overwrite || o.pageImages || o.words || o.ocrMode === "all")) {
|
|
188
|
+
throw new UsageError("--info cannot be combined with --pages, --output-dir, --overwrite, --page-images, --words or --ocr-mode all");
|
|
182
189
|
}
|
|
183
190
|
return o;
|
|
184
191
|
}
|
|
185
192
|
|
|
193
|
+
/** Two-pass OCR recipe, rendered into --help and the generated skill; `<cmd>` is the command prefix. */
|
|
194
|
+
export const USAGE_PATTERNS = [
|
|
195
|
+
"Two-pass OCR (PDF, PPTX, DOC):",
|
|
196
|
+
" 1. <cmd> report.pdf --output-dir out --json",
|
|
197
|
+
' -> "pageStatsPath" points at out/report.pages.json; pages with few',
|
|
198
|
+
' chars and high imageCoverage are scans. "savedTo" is the Markdown.',
|
|
199
|
+
" 2. <cmd> report.pdf --output-dir out --ocr --ocr-mode all --pages 2,7 --json",
|
|
200
|
+
' -> "ocr"."sidecars" maps 2 and 7 to out/ocr/report-2-p002.md and',
|
|
201
|
+
" ...-p007.md (a second run in the same dir gets stem report-2);",
|
|
202
|
+
" the Markdown of this run holds pages 2 and 7 only and equals what",
|
|
203
|
+
" --pages 2,7 without --ocr would produce. Read sidecars and Markdown",
|
|
204
|
+
" by the returned paths, never by guessing names.",
|
|
205
|
+
" 3. --ocr-mode all refuses to run without --ocr and an explicit --pages.",
|
|
206
|
+
' The "ocr" object (OCR: line) names failed, budget-stopped, killed and',
|
|
207
|
+
" not-attempted pages and the exact --pages to re-run.",
|
|
208
|
+
" Details: doc/doc-to-md.md (bundle contract, failure buckets).",
|
|
209
|
+
].join("\n");
|
|
210
|
+
|
|
211
|
+
export const usagePatterns = (cmd: string): string => USAGE_PATTERNS.replaceAll("<cmd>", cmd);
|
|
212
|
+
|
|
213
|
+
/** Every artifact a bundle can contain; shared by --help and the generated skill. */
|
|
214
|
+
export const BUNDLE_LAYOUT: readonly { artifact: string; trigger: string; content: string; namedBy: string }[] = [
|
|
215
|
+
{ artifact: "<stem>.md", trigger: "always", content: "the Markdown", namedBy: "Saved-To: / savedTo" },
|
|
216
|
+
{ artifact: "images/", trigger: "embedded or extracted figures", content: "image files linked from the Markdown", namedBy: "Images-Dir: / imagesDir" },
|
|
217
|
+
{ artifact: "pages/<stem>-pNNN.<fmt>", trigger: "--page-images", content: "page renders at --image-dpi", namedBy: "Pages-Dir: / pagesDir" },
|
|
218
|
+
{ artifact: "sheets/", trigger: "Excel input", content: "one CSV per non-empty worksheet", namedBy: "Sheets-Dir: / sheetsDir" },
|
|
219
|
+
{ artifact: "attachments/", trigger: "email input", content: "saved attachments", namedBy: "Markdown attachment list" },
|
|
220
|
+
{ artifact: "<stem>.pages.json", trigger: "Python PDF tiers (PDF, PPTX, DOC, DOCX via LibreOffice; not unpdf)", content: "per-page chars, image count, image coverage", namedBy: "Page-Stats: / pageStatsPath" },
|
|
221
|
+
{ artifact: "<stem>.words.json", trigger: "--words", content: "per-page word boxes, source text/ocr", namedBy: "Words: / wordsPath" },
|
|
222
|
+
{ artifact: "ocr/<stem>-pNNN.md", trigger: "--ocr --ocr-mode all", content: "recognized text of a forced page", namedBy: "OCR-Dir: / ocr.sidecars" },
|
|
223
|
+
{ artifact: "ocr/<stem>-pNNN.words.json", trigger: "--ocr --ocr-mode all --words", content: "word boxes of that OCR", namedBy: "ocr.wordSidecars" },
|
|
224
|
+
];
|
|
225
|
+
|
|
226
|
+
const bundleLayoutText = (): string => BUNDLE_LAYOUT.map((r) => ` ${r.artifact.padEnd(28)} ${r.trigger}; ${r.content}; named by ${r.namedBy}`).join("\n");
|
|
227
|
+
|
|
186
228
|
export function renderHelp(): string {
|
|
187
229
|
const row = (d: OptionDescriptor) => ` ${(d.flag ?? "<path>").padEnd(26)} ${d.help}${d.default !== null && d.key !== "info" && d.key !== "overwrite" ? ` (default ${d.default})` : ""}`;
|
|
188
230
|
return [
|
|
@@ -190,7 +232,8 @@ export function renderHelp(): string {
|
|
|
190
232
|
"", "Per-call:", ...DOC_TO_MD_OPTIONS.filter((d) => !d.settable).map(row),
|
|
191
233
|
"", "Tunables (also settable under quiver.docToMd in settings.json; per-call > settings > default):",
|
|
192
234
|
...DOC_TO_MD_OPTIONS.filter((d) => d.settable).map(row),
|
|
193
|
-
"", "Result: a handle (Saved-To, Images-Dir, Page-Count, Outline ...). Read the Saved-To file for the Markdown.",
|
|
235
|
+
"", "Result: a handle (Saved-To, Images-Dir, Page-Stats, Page-Count, Outline ...). Read the Saved-To file for the Markdown.",
|
|
194
236
|
"Exit codes: 0 success, 1 runtime error, 2 usage error.",
|
|
237
|
+
"", "Bundle layout:", bundleLayoutText(), "", usagePatterns("pi-quiver doc-to-md"),
|
|
195
238
|
].join("\n");
|
|
196
239
|
}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-quiver",
|
|
3
|
-
"version": "6.
|
|
3
|
+
"version": "6.10.0",
|
|
4
4
|
"description": "Personal pack of Pi coding-agent extensions: context-safe fetch, doc_to_md PDF/DOCX/PPTX-to-Markdown conversion, session naming, a themed ASCII startup header, Opus 4.8 fast mode, and a provider-stall watchdog.",
|
|
5
5
|
"author": "Jacek Juraszek",
|
|
6
6
|
"license": "MIT",
|
package/scripts/doc_to_md.py
CHANGED
|
@@ -205,11 +205,22 @@ def mode_image(o):
|
|
|
205
205
|
info = new_ocr(lang)
|
|
206
206
|
status = ocr_status(bool(o.get("ocr")), lang)
|
|
207
207
|
apply_status(info, status)
|
|
208
|
-
|
|
208
|
+
words_on = bool(o.get("words"))
|
|
209
|
+
word_pages, words_errors = [], {}
|
|
210
|
+
w_px = h_px = None
|
|
211
|
+
if words_on or status["status"] == "ready":
|
|
209
212
|
try:
|
|
210
213
|
pix = pymupdf.Pixmap(path)
|
|
211
214
|
w_px, h_px = pix.width, pix.height
|
|
212
215
|
del pix
|
|
216
|
+
except Exception as exc: # noqa: BLE001
|
|
217
|
+
if words_on:
|
|
218
|
+
words_errors["1"] = words_error(exc)
|
|
219
|
+
if status["status"] == "ready":
|
|
220
|
+
info["ocrFailed"].append(1)
|
|
221
|
+
status = {**status, "status": "failed"}
|
|
222
|
+
if status["status"] == "ready":
|
|
223
|
+
try:
|
|
213
224
|
if min(w_px, h_px) < MIN_OCR_SIDE_PX:
|
|
214
225
|
info["status"], info["reason"] = "skipped", "image too small"
|
|
215
226
|
else:
|
|
@@ -218,6 +229,19 @@ def mode_image(o):
|
|
|
218
229
|
text = pymupdf4llm.to_markdown(pdf, pages=[0], write_images=False, use_ocr=True, force_ocr=True,
|
|
219
230
|
ocr_language=lang, ocr_dpi=image_ocr_dpi(w_px, r.width, r.height),
|
|
220
231
|
page_separators=False).strip()
|
|
232
|
+
if words_on:
|
|
233
|
+
try:
|
|
234
|
+
pg = pdf[0]
|
|
235
|
+
entry = words_page(pg, 1, 0, page_words(pg, True))
|
|
236
|
+
sx, sy = w_px / r.width, h_px / r.height
|
|
237
|
+
for word in entry["words"]:
|
|
238
|
+
b = word["bbox"]
|
|
239
|
+
word["bbox"] = [round(b[0] * sx, 1), round(b[1] * sy, 1),
|
|
240
|
+
round(b[2] * sx, 1), round(b[3] * sy, 1)]
|
|
241
|
+
entry["width"], entry["height"] = w_px, h_px
|
|
242
|
+
word_pages.append(entry)
|
|
243
|
+
except Exception as exc: # noqa: BLE001
|
|
244
|
+
words_errors["1"] = words_error(exc)
|
|
221
245
|
if text:
|
|
222
246
|
info["pages"].append(1)
|
|
223
247
|
md += "\n\n" + ocr_block(f"p1/{name}", text)
|
|
@@ -225,7 +249,17 @@ def mode_image(o):
|
|
|
225
249
|
info["noText"].append(1)
|
|
226
250
|
except Exception: # noqa: BLE001 - OCR never fails the conversion
|
|
227
251
|
info["ocrFailed"].append(1)
|
|
228
|
-
|
|
252
|
+
if words_on and not o.get("ocr") and w_px is not None and "1" not in words_errors:
|
|
253
|
+
try:
|
|
254
|
+
if os.environ.get("DOC_TO_MD_WORDS_FAIL") == "1": # tests only
|
|
255
|
+
raise RuntimeError("words injected failure")
|
|
256
|
+
word_pages.append({"page": 1, "width": w_px, "height": h_px, "rotation": 0, "words": []})
|
|
257
|
+
except Exception as exc: # noqa: BLE001
|
|
258
|
+
words_errors["1"] = words_error(exc)
|
|
259
|
+
result = {"markdown": md + "\n", "pageCount": 1, "emptyPages": [], "failedPages": [], "notes": [], "ocr": info}
|
|
260
|
+
if words_on:
|
|
261
|
+
result.update(write_words(o["stagingDir"], "px", word_pages, words_errors))
|
|
262
|
+
return result
|
|
229
263
|
|
|
230
264
|
|
|
231
265
|
def mode_info(o):
|
|
@@ -253,6 +287,71 @@ def primary_page_markdown(doc, n, d, o, kw, write_images):
|
|
|
253
287
|
return rewrite_image_destinations(md, sources)
|
|
254
288
|
|
|
255
289
|
|
|
290
|
+
def page_stats(doc, n):
|
|
291
|
+
try:
|
|
292
|
+
page = doc[n - 1]
|
|
293
|
+
chars = len(page.get_text("text").strip())
|
|
294
|
+
infos = page.get_image_info()
|
|
295
|
+
area = page.rect.width * page.rect.height
|
|
296
|
+
covered = sum(max(0.0, (b[2] - b[0]) * (b[3] - b[1])) for b in (i["bbox"] for i in infos))
|
|
297
|
+
coverage = round(min(1.0, covered / area), 2) if area > 0 else 0.0
|
|
298
|
+
return {"page": n, "chars": chars, "images": len(infos), "imageCoverage": coverage}
|
|
299
|
+
except Exception as exc: # noqa: BLE001 - stats never cost a page its Markdown
|
|
300
|
+
return {"page": n, "error": f"{type(exc).__name__}: {exc}"[:300]}
|
|
301
|
+
|
|
302
|
+
|
|
303
|
+
GLYPHLESS_FONT = "GlyphLessFont"
|
|
304
|
+
|
|
305
|
+
|
|
306
|
+
def word_entries(page, words, ocr_lines):
|
|
307
|
+
import pymupdf
|
|
308
|
+
matrix = page.rotation_matrix
|
|
309
|
+
out = []
|
|
310
|
+
for x0, y0, x1, y1, text, bno, lno, _ in words:
|
|
311
|
+
r = pymupdf.Rect(x0, y0, x1, y1) * matrix
|
|
312
|
+
out.append({"text": text, "bbox": [round(r.x0, 1), round(r.y0, 1), round(r.x1, 1), round(r.y1, 1)],
|
|
313
|
+
"source": "ocr" if ocr_lines is None or (bno, lno) in ocr_lines else "text"})
|
|
314
|
+
return out
|
|
315
|
+
|
|
316
|
+
|
|
317
|
+
def page_words(page, inline_ocr_ran):
|
|
318
|
+
if not inline_ocr_ran:
|
|
319
|
+
return word_entries(page, page.get_text("words"), set())
|
|
320
|
+
# Shared textpage keeps block/line indexes aligned despite differing default image flags.
|
|
321
|
+
tp = page.get_textpage()
|
|
322
|
+
ocr_lines = {(bi, li) for bi, b in enumerate(page.get_text("dict", textpage=tp)["blocks"])
|
|
323
|
+
for li, line in enumerate(b.get("lines", []))
|
|
324
|
+
if line["spans"] and all(s["font"] == GLYPHLESS_FONT for s in line["spans"])}
|
|
325
|
+
return word_entries(page, page.get_text("words", textpage=tp), ocr_lines)
|
|
326
|
+
|
|
327
|
+
|
|
328
|
+
def ocr_words(page, tp):
|
|
329
|
+
return word_entries(page, page.get_text("words", textpage=tp), None)
|
|
330
|
+
|
|
331
|
+
|
|
332
|
+
def words_page(page, n, rotation, words):
|
|
333
|
+
if os.environ.get("DOC_TO_MD_WORDS_FAIL") == str(n): # tests only
|
|
334
|
+
raise RuntimeError("words injected failure")
|
|
335
|
+
return {"page": n, "width": round(page.rect.width, 1), "height": round(page.rect.height, 1),
|
|
336
|
+
"rotation": rotation, "words": words}
|
|
337
|
+
|
|
338
|
+
|
|
339
|
+
def words_error(exc):
|
|
340
|
+
return f"{type(exc).__name__}: {exc}"[:300]
|
|
341
|
+
|
|
342
|
+
|
|
343
|
+
def write_words(staging, unit, pages, errors):
|
|
344
|
+
try:
|
|
345
|
+
os.makedirs(staging, exist_ok=True)
|
|
346
|
+
with open(os.path.join(staging, "words.json"), "w", encoding="utf-8") as fh:
|
|
347
|
+
json.dump({"unit": unit, "pages": pages}, fh, indent=2)
|
|
348
|
+
fh.write("\n")
|
|
349
|
+
return {"words": True, "wordsErrors": errors}
|
|
350
|
+
except Exception as exc: # noqa: BLE001 - geometry never fails the conversion
|
|
351
|
+
errors["file"] = words_error(exc)
|
|
352
|
+
return {"words": False, "wordsErrors": errors}
|
|
353
|
+
|
|
354
|
+
|
|
256
355
|
def mode_pdf_primary(o):
|
|
257
356
|
import time
|
|
258
357
|
start = time.monotonic()
|
|
@@ -260,14 +359,18 @@ def mode_pdf_primary(o):
|
|
|
260
359
|
pages = check_pages(o.get("pages"), doc.page_count)
|
|
261
360
|
staging, out, empty, failed, notes = o["stagingDir"], [], [], [], []
|
|
262
361
|
page_images = []
|
|
362
|
+
stats = []
|
|
263
363
|
lang, budget = o.get("ocrLanguage", "eng"), o.get("ocrBudgetMs", 60000)
|
|
264
364
|
info = new_ocr(lang)
|
|
265
365
|
status = ocr_status(True, lang) if o.get("ocr") else None
|
|
266
366
|
ocr_ms, plain_ms = [], []
|
|
367
|
+
word_pages, words_errors = [], {}
|
|
267
368
|
for i, n in enumerate(pages):
|
|
369
|
+
stats.append(page_stats(doc, n))
|
|
268
370
|
d = page_dir(staging, n)
|
|
269
371
|
try:
|
|
270
372
|
page = doc[n - 1]
|
|
373
|
+
rotation = page.rotation
|
|
271
374
|
textless = not page.get_text("text").strip()
|
|
272
375
|
if textless:
|
|
273
376
|
info["textless"].append(n)
|
|
@@ -311,6 +414,11 @@ def mode_pdf_primary(o):
|
|
|
311
414
|
if page_pic:
|
|
312
415
|
md = "\n\n".join(x for x in [md.rstrip(), f""] if x)
|
|
313
416
|
out.append(md.rstrip())
|
|
417
|
+
if o.get("words") and not (textless and o.get("ocr") and not kw["use_ocr"]):
|
|
418
|
+
try:
|
|
419
|
+
word_pages.append(words_page(page, n, rotation, page_words(page, kw["use_ocr"])))
|
|
420
|
+
except Exception as exc: # noqa: BLE001
|
|
421
|
+
words_errors[str(n)] = words_error(exc)
|
|
314
422
|
except Exception as exc: # noqa: BLE001
|
|
315
423
|
shutil.rmtree(d, ignore_errors=True)
|
|
316
424
|
failed.append({"page": n, "error": f"{type(exc).__name__}: {exc}"[:300]})
|
|
@@ -324,8 +432,11 @@ def mode_pdf_primary(o):
|
|
|
324
432
|
missing = len(pages) - len(page_images) if o.get("pageImages") else 0
|
|
325
433
|
if missing:
|
|
326
434
|
notes.append(f"Page images: {missing} of {len(pages)} unavailable")
|
|
327
|
-
|
|
328
|
-
|
|
435
|
+
result = {"markdown": "\n\n".join(out) + "\n", "pages": pages, "pageCount": doc.page_count,
|
|
436
|
+
"emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": info, "pageImages": page_images, "pageStats": stats}
|
|
437
|
+
if o.get("words"):
|
|
438
|
+
result.update(write_words(staging, "pt", word_pages, words_errors))
|
|
439
|
+
return result
|
|
329
440
|
|
|
330
441
|
|
|
331
442
|
def mode_pdf_fallback(o):
|
|
@@ -335,15 +446,20 @@ def mode_pdf_fallback(o):
|
|
|
335
446
|
keep = {int(k): v for k, v in (o.get("keepPages") or {}).items()}
|
|
336
447
|
staging, out, empty, failed = o["stagingDir"], [], [], []
|
|
337
448
|
page_images = []
|
|
449
|
+
stats = []
|
|
338
450
|
lang = o.get("ocrLanguage", "eng")
|
|
339
451
|
ocr_info = new_ocr(lang)
|
|
452
|
+
word_pages, words_errors = [], {}
|
|
340
453
|
status = {"status": "unavailable", "reason": "fallback tier", "tesseract": None} if o.get("ocr") else None
|
|
341
454
|
for n in pages:
|
|
455
|
+
stats.append(page_stats(doc, n))
|
|
342
456
|
links = [f"" for f in keep.get(n, [])]
|
|
343
457
|
text = ""
|
|
344
458
|
page_pic = None
|
|
459
|
+
page_ok = False
|
|
345
460
|
try:
|
|
346
461
|
page = doc[n - 1]
|
|
462
|
+
rotation = page.rotation
|
|
347
463
|
text = page.get_text("text").strip()
|
|
348
464
|
if not text:
|
|
349
465
|
ocr_info["textless"].append(n)
|
|
@@ -371,11 +487,17 @@ def mode_pdf_fallback(o):
|
|
|
371
487
|
links.append(f"")
|
|
372
488
|
mark_done(d)
|
|
373
489
|
page_pic = render_page_image(page, n, o, page_images) if o.get("pageImages") else None
|
|
490
|
+
page_ok = True
|
|
374
491
|
except Exception as exc: # noqa: BLE001
|
|
375
492
|
shutil.rmtree(os.path.join(staging, f"p{n}"), ignore_errors=True)
|
|
376
493
|
text = ""
|
|
377
494
|
links = [f"" for f in keep.get(n, [])]
|
|
378
495
|
failed.append({"page": n, "error": f"{type(exc).__name__}: {exc}"[:300]})
|
|
496
|
+
if o.get("words") and page_ok and not (not text and o.get("ocr")):
|
|
497
|
+
try:
|
|
498
|
+
word_pages.append(words_page(page, n, rotation, page_words(page, False)))
|
|
499
|
+
except Exception as exc: # noqa: BLE001
|
|
500
|
+
words_errors[str(n)] = words_error(exc)
|
|
379
501
|
if not text:
|
|
380
502
|
empty.append(n)
|
|
381
503
|
out.append("\n\n".join(x for x in [text, "\n".join(links), f"" if page_pic else ""] if x))
|
|
@@ -388,8 +510,88 @@ def mode_pdf_fallback(o):
|
|
|
388
510
|
missing = len(pages) - len(page_images) if o.get("pageImages") else 0
|
|
389
511
|
if missing:
|
|
390
512
|
notes.append(f"Page images: {missing} of {len(pages)} unavailable")
|
|
391
|
-
|
|
392
|
-
|
|
513
|
+
result = {"markdown": "\n\n".join(out) + "\n", "pages": pages, "pageCount": doc.page_count,
|
|
514
|
+
"emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": ocr_info, "pageImages": page_images, "pageStats": stats}
|
|
515
|
+
if o.get("words"):
|
|
516
|
+
result.update(write_words(staging, "pt", word_pages, words_errors))
|
|
517
|
+
return result
|
|
518
|
+
|
|
519
|
+
|
|
520
|
+
def mode_ocr_pages(o):
|
|
521
|
+
import time
|
|
522
|
+
start = time.monotonic()
|
|
523
|
+
lang, budget = o.get("ocrLanguage", "eng"), o.get("ocrBudgetMs", 60000)
|
|
524
|
+
status = ocr_status(True, lang)
|
|
525
|
+
if status["status"] != "ready":
|
|
526
|
+
return {"status": "unavailable", "reason": status["reason"]}
|
|
527
|
+
doc = open_pdf(o["path"])
|
|
528
|
+
pages = check_pages(o.get("pages"), doc.page_count)
|
|
529
|
+
staging, stem, dpi = o["stagingDir"], o["stem"], o.get("dpi", 150)
|
|
530
|
+
os.makedirs(staging, exist_ok=True)
|
|
531
|
+
active = os.path.join(staging, "active")
|
|
532
|
+
out = {"status": "ran", "written": [], "noText": [], "ocrFailed": [], "ocrErrors": {}, "budgetStopped": []}
|
|
533
|
+
if o.get("words"):
|
|
534
|
+
out["wordsErrors"] = {}
|
|
535
|
+
ocr_ms = []
|
|
536
|
+
for i, n in enumerate(pages):
|
|
537
|
+
with open(active, "w") as fh:
|
|
538
|
+
fh.write(str(n))
|
|
539
|
+
if os.environ.get("DOC_TO_MD_OCR_STALL_PAGE") == str(n): # tests only: simulate a wedged page
|
|
540
|
+
time.sleep(3600)
|
|
541
|
+
elapsed = (time.monotonic() - start) * 1000
|
|
542
|
+
est = max(ocr_ms) if ocr_ms else OCR_EST_INITIAL_MS
|
|
543
|
+
if not ocr_admit(elapsed, est, 0, 0, budget):
|
|
544
|
+
out["budgetStopped"].extend(pages[i:])
|
|
545
|
+
os.remove(active)
|
|
546
|
+
break
|
|
547
|
+
tag = f"p{n:03d}"
|
|
548
|
+
d = os.path.join(staging, tag)
|
|
549
|
+
os.makedirs(d, exist_ok=True)
|
|
550
|
+
sidecar = os.path.join(d, f"{stem}-{tag}.md")
|
|
551
|
+
t0 = time.monotonic()
|
|
552
|
+
try:
|
|
553
|
+
page = doc[n - 1]
|
|
554
|
+
eff = clamped_dpi(page.rect.width, page.rect.height, dpi)
|
|
555
|
+
if eff is None:
|
|
556
|
+
raise RuntimeError(f"page cannot be rendered at a usable DPI ({page.rect.width:.0f} x {page.rect.height:.0f} pt)")
|
|
557
|
+
tp = page.get_textpage_ocr(full=True, language=lang, dpi=eff)
|
|
558
|
+
text = page.get_text("text", textpage=tp).strip()
|
|
559
|
+
header = f"<!-- OCR of page {n} (tesseract {lang}); recognized text, not the text layer -->"
|
|
560
|
+
with open(sidecar, "w", encoding="utf-8") as fh:
|
|
561
|
+
fh.write(header + "\n\n" + (text + "\n\n" if text else "") + SEP.format(n=n).strip("\n") + "\n")
|
|
562
|
+
if o.get("words"):
|
|
563
|
+
wpath = os.path.join(d, f"{stem}-{tag}.words.json")
|
|
564
|
+
try:
|
|
565
|
+
entry = words_page(page, n, page.rotation, ocr_words(page, tp))
|
|
566
|
+
entry["unit"] = "pt"
|
|
567
|
+
with open(wpath, "w", encoding="utf-8") as fh:
|
|
568
|
+
json.dump(entry, fh, indent=2)
|
|
569
|
+
fh.write("\n")
|
|
570
|
+
except Exception as exc: # noqa: BLE001 - geometry never changes the OCR outcome
|
|
571
|
+
try:
|
|
572
|
+
os.remove(wpath)
|
|
573
|
+
except OSError:
|
|
574
|
+
pass
|
|
575
|
+
out["wordsErrors"][str(n)] = words_error(exc)
|
|
576
|
+
mark_done(d)
|
|
577
|
+
(out["written"] if text else out["noText"]).append(n)
|
|
578
|
+
except Exception as exc: # noqa: BLE001 - one page never stops the pass
|
|
579
|
+
msg = f"{type(exc).__name__}: {exc}"
|
|
580
|
+
with open(os.path.join(d, ".failed"), "w", encoding="utf-8") as fh:
|
|
581
|
+
fh.write(msg)
|
|
582
|
+
try:
|
|
583
|
+
os.remove(sidecar)
|
|
584
|
+
except OSError:
|
|
585
|
+
pass
|
|
586
|
+
try:
|
|
587
|
+
os.remove(os.path.join(d, f"{stem}-{tag}.words.json"))
|
|
588
|
+
except OSError:
|
|
589
|
+
pass
|
|
590
|
+
out["ocrFailed"].append(n)
|
|
591
|
+
out["ocrErrors"][str(n)] = msg
|
|
592
|
+
ocr_ms.append((time.monotonic() - t0) * 1000)
|
|
593
|
+
os.remove(active)
|
|
594
|
+
return out
|
|
393
595
|
|
|
394
596
|
|
|
395
597
|
def esc(v):
|
|
@@ -1249,12 +1451,12 @@ def mode_render_pages(o):
|
|
|
1249
1451
|
return {"ok": True, "rendered": rendered, "failed": failed}
|
|
1250
1452
|
|
|
1251
1453
|
|
|
1252
|
-
MODES = {"info": None, "pdf-primary": mode_pdf_primary, "pdf-fallback": mode_pdf_fallback, "xlsx": mode_xlsx, "render-pages": mode_render_pages, "docx": mode_docx, "html": mode_html, "image": mode_image, "email": mode_email}
|
|
1454
|
+
MODES = {"info": None, "pdf-primary": mode_pdf_primary, "pdf-fallback": mode_pdf_fallback, "xlsx": mode_xlsx, "render-pages": mode_render_pages, "docx": mode_docx, "html": mode_html, "image": mode_image, "email": mode_email, "ocr-pages": mode_ocr_pages}
|
|
1253
1455
|
|
|
1254
1456
|
|
|
1255
1457
|
def main():
|
|
1256
1458
|
if len(sys.argv) != 2 or sys.argv[1] not in MODES:
|
|
1257
|
-
print("usage: doc_to_md.py <info|pdf-primary|pdf-fallback|xlsx|render-pages|docx|html|image|email> (options JSON on stdin)", file=sys.stderr)
|
|
1459
|
+
print("usage: doc_to_md.py <info|pdf-primary|pdf-fallback|xlsx|render-pages|docx|html|image|email|ocr-pages> (options JSON on stdin)", file=sys.stderr)
|
|
1258
1460
|
return 1
|
|
1259
1461
|
mode = sys.argv[1]
|
|
1260
1462
|
o = json.loads(sys.stdin.read() or "{}")
|