pi-quiver 6.8.0 → 6.9.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +6 -0
- package/README.md +4 -2
- package/dist/bin/pi-quiver.js +140 -10
- package/extensions/doc_to_md.ts +1 -1
- package/lib/doc-to-md-bundle.ts +41 -3
- package/lib/doc-to-md-core.ts +61 -10
- package/lib/doc-to-md-handle.ts +28 -1
- package/lib/doc-to-md-options.ts +28 -3
- package/package.json +1 -1
- package/scripts/doc_to_md.py +78 -4
package/CHANGELOG.md
CHANGED
|
@@ -8,6 +8,12 @@ Published to npm as `pi-quiver` (`pi install npm:pi-quiver`). Pushing a
|
|
|
8
8
|
via OIDC trusted publishing. The release helper at
|
|
9
9
|
`.agents/skills/release/scripts/release.sh` cuts the tag; CI publishes.
|
|
10
10
|
|
|
11
|
+
## v6.9.0 - 2026-10-01
|
|
12
|
+
|
|
13
|
+
- `doc_to_md`: every conversion that goes through the Python PDF tiers (PDF, PPTX, `.doc`, DOCX via LibreOffice) writes `<stem>.pages.json` (per-page `chars`, `images`, `imageCoverage`) and the handle prints `Page-Stats:`; CLI `--json` gains `pageStats`, `pageStatsPath`, `ocrDir`. The unpdf tier writes no stats (#26).
|
|
14
|
+
- `doc_to_md`: new per-call `ocrMode` (`--ocr-mode textless|all`). `all` forces OCR on an explicit `pages` selection in a second, kill-capped Python child and writes `ocr/<stem>-pNNN.md` sidecars; the Markdown stays byte-identical to the same selected-page call without OCR. Refused without `--ocr`, without explicit pages, or on non-PDF/PPTX/DOC inputs (exit 2); missing Tesseract data is a hard error. The `OCR:` line names written, no-text, failed, budget-stopped, killed and not-attempted pages, with the exact `--pages` to re-run for budget-stopped and not-attempted pages (#26).
|
|
15
|
+
- `doc_to_md`: `--help` and the generated `skills/doc-to-md/SKILL.md` end with the two-pass OCR usage block (#26).
|
|
16
|
+
|
|
11
17
|
## v6.8.0 - 2026-10-01
|
|
12
18
|
|
|
13
19
|
- `session-name`: opt-in automatic naming starts in the background after three completed model/tool rounds. Shorter runs start a non-blocking attempt at run end; initial generation is best-effort with a 30-second local deadline and no automatic retry per session activation.
|
package/README.md
CHANGED
|
@@ -65,7 +65,7 @@ A 300 KB changelog page never touches your context window - you get a preview an
|
|
|
65
65
|
| Extension | Tool | What it does |
|
|
66
66
|
| --- | --- | --- |
|
|
67
67
|
| `extensions/fetch.ts` | `fetch` | Retrieve URLs over HTTP(S). HTML -> Markdown (Readability extraction, Turndown conversion). Binary saved untouched to a temp file. GitHub issue/PR/repo/actions-run/actions-job URLs auto-route through `gh` (falls back to HTTP); failed runs/jobs include failed-step logs (best-effort, summary-only otherwise). Same size gate as `fetch`. Behavior lives in `lib/fetch-core.ts`; also exposed as the `pi-quiver fetch` CLI (see [Claude Code support](#claude-code-support)). |
|
|
68
|
-
| `extensions/doc_to_md.ts` | `doc_to_md` | Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, MSG/EML, HTML, or image file to a Markdown bundle on disk (`<stem>.md` + `images/`, spreadsheet `sheets/`, optional `pageImages` in `pages/`, and email `attachments/`) and return a handle (paths, page count, outline, diagnostics) - never inline Markdown. `info` inspects PDF/Office inputs; `pages` selects 1-based PDF/Office pages (DOCX: explicit-page-break segments); HTML, image, and email inputs reject both; every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Tiers: pymupdf4llm -> PyMuPDF text (degraded) -> unpdf worker (no Python only). Excel -> sheet inventory, full CSVs, bounded previews, and optional rendered views. HTML keeps local and `data:` images in the bundle and remote images as links; image inputs keep the original image. Scanned PDF pages keep a page picture on both Python tiers; opt-in OCR needs Tesseract language data. DOCX converts directly (mammoth, python-docx fallback); LibreOffice pagination marks the degraded route. The Outline lists `L<line>` and `p<page>` per heading. Settings under `quiver.docToMd`. |
|
|
68
|
+
| `extensions/doc_to_md.ts` | `doc_to_md` | Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, MSG/EML, HTML, or image file to a Markdown bundle on disk (`<stem>.md` + `images/`, spreadsheet `sheets/`, optional `pageImages` in `pages/`, and email `attachments/`) and return a handle (paths, page count, outline, diagnostics) - never inline Markdown. `info` inspects PDF/Office inputs; `pages` selects 1-based PDF/Office pages (DOCX: explicit-page-break segments); HTML, image, and email inputs reject both; every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Tiers: pymupdf4llm -> PyMuPDF text (degraded) -> unpdf worker (no Python only). Excel -> sheet inventory, full CSVs, bounded previews, and optional rendered views. HTML keeps local and `data:` images in the bundle and remote images as links; image inputs keep the original image. Scanned PDF pages keep a page picture on both Python tiers; opt-in OCR needs Tesseract language data. DOCX converts directly (mammoth, python-docx fallback); LibreOffice pagination marks the degraded route. The Outline lists `L<line>` and `p<page>` per heading. Settings under `quiver.docToMd`. Every conversion that goes through the Python PDF tiers (PDF, PPTX, `.doc`, DOCX via LibreOffice) writes `<stem>.pages.json` (per-page chars, image count, image coverage; handle line `Page-Stats:`); unpdf writes no stats. `ocrMode: "all"` with `ocr: true` and an explicit `pages` forces OCR on those pages into `ocr/<stem>-pNNN.md` sidecars, leaving the Markdown byte-identical to the same selected-page call without OCR. |
|
|
69
69
|
| `extensions/session-name.ts` | `/session-name` | Manual + opt-in automatic session naming, naming rules and deny list, long-session revisits, and Ghostty/Herdr tab rename. OFF by default. |
|
|
70
70
|
| `extensions/sword-header.ts` | `/builtin-header` | Themed ASCII startup header replacing pi's default logo. OFF by default. |
|
|
71
71
|
| `extensions/fast-mode.ts` | `/fast` | Inject Anthropic fast-mode (`speed: "fast"` + `anthropic-beta: fast-mode-2026-02-01`) into every Claude Opus 4.8 / Opus 5 request, any thinking level. `--fast` flag + `/fast [on\|off\|status]`. OFF by default. |
|
|
@@ -310,7 +310,9 @@ Each setting can also be overridden per-process via `PI_QUIVER_SLACK_ENABLED`, `
|
|
|
310
310
|
| `ocr` | `false` | Opt-in OCR for scanned pages and images when Tesseract language data is installed. |
|
|
311
311
|
| `ocrLanguage` | `eng` | Plain `+`-joined Tesseract language codes for OCR. |
|
|
312
312
|
|
|
313
|
-
|
|
313
|
+
`ocrMode` (`textless` | `all`) is a per-call parameter and `--ocr-mode` flag only; a `quiver.docToMd.ocrMode` key is reported as unknown and ignored, so persisted settings cannot enable forced OCR or bypass its explicit-pages guard.
|
|
314
|
+
|
|
315
|
+
A bundle is `<outputDir>/<stem>.md` plus `<outputDir>/<stem>.pages.json` on Python PDF tiers, optional OCR sidecars under `<outputDir>/ocr/`, and `<outputDir>/images/` and, for spreadsheets with data, `<outputDir>/sheets/`; without `outputDir`, the tool creates a per-call temp root. A conversion owns `<stem>.md.lock` until it atomically publishes the Markdown. An existing `<stem>.md` fails the call unless `overwrite` is set; `overwrite` replaces that Markdown and the files it owns (`images/<stem>-p<N>-<n>.*`, `images/<stem>-s<idx>[-<n>].*`, `sheets/<stem>-s<idx>-<slug>.csv`, `<stem>.pages.json`, `ocr/<stem>-pNNN.md`), nothing else. Temp bundles are caller-owned - the tool never deletes a bundle it produced.
|
|
314
316
|
|
|
315
317
|
Excel needs a Python backend with openpyxl, xlrd and pillow - otherwise the call fails with `Remedy: install uv, or pip install openpyxl xlrd pillow`. The Markdown opens with a `## Sheets` table listing every sheet in workbook order (0-based `#`, `worksheet`/`chartsheet`, size, hidden, chart and image counts, rendered view, CSV link for non-empty worksheets), then one section per sheet: a `Data:` line linking the full-content CSV under `sheets/` for non-empty worksheets, chart metadata from the workbook model, embedded images, an optional rendered view, a preview of at most the first 100 rows x 50 columns, and - only when the preview is truncated - a `Columns:` profile (type, non-empty count, min/max, distinct up to 50). Sizes are the extent of non-empty cells (the `info` handle reports the raw worksheet dimensions instead, which may be larger). Rendered views (`images/<stem>-s<idx>.<fmt>`) are produced for sheets carrying charts or images when LibreOffice is on `PATH`: the workbook is exported one PDF page per sheet and rasterized under a 16 Mpx budget. Any LibreOffice or rasterization failure degrades to `Rendered view: unavailable (<reason>)` and a handle note; it never fails the conversion. Workbooks whose chartsheet drawings carry a zero-size anchor (openpyxl-authored files; Excel-authored files are unaffected) render as a degenerate page and are reported as such. `.xls` gets the inventory, CSVs and previews but no visual detection or rendering.
|
|
316
318
|
|
package/dist/bin/pi-quiver.js
CHANGED
|
@@ -581,6 +581,9 @@ function ownedCsvPattern(stem) {
|
|
|
581
581
|
function ownedPagePattern(stem) {
|
|
582
582
|
return new RegExp(`^${escRe(stem)}-p\\d+\\.[a-z0-9]+$`);
|
|
583
583
|
}
|
|
584
|
+
function ownedOcrPattern(stem) {
|
|
585
|
+
return new RegExp(`^${escRe(stem)}-p\\d+\\.md$`);
|
|
586
|
+
}
|
|
584
587
|
var FILE_LINK_RE = /\[[^\]]*\]\(\s*((?:sheets|attachments)\/[^)\s]+)\s*\)/g;
|
|
585
588
|
var IMG_LINK_RE = /!\[[^\]]*\]\(\s*(?:<([^>]*)>|([^)]*?))\s*\)|<img\b[^>]*\bsrc\s*=\s*(?:"([^"]*)"|'([^']*)'|([^\s>"']+))/gi;
|
|
586
589
|
function imageTarget(m) {
|
|
@@ -627,6 +630,7 @@ function openBundle(root, requested, overwrite) {
|
|
|
627
630
|
const sheetsDir = join2(root, "sheets");
|
|
628
631
|
const pagesDir = join2(root, "pages");
|
|
629
632
|
const attachmentsDir = join2(root, "attachments");
|
|
633
|
+
const ocrDir = join2(root, "ocr");
|
|
630
634
|
try {
|
|
631
635
|
if (existsSync(mdPath) && overwrite) {
|
|
632
636
|
const owned = ownedPattern(stem), ownedCsv = ownedCsvPattern(stem), ownedPage = ownedPagePattern(stem);
|
|
@@ -644,6 +648,11 @@ function openBundle(root, requested, overwrite) {
|
|
|
644
648
|
if (existsSync(attachmentsDir)) {
|
|
645
649
|
for (const f of readdirSync(attachmentsDir)) if (attachments.has(f)) rmSync(join2(attachmentsDir, f), { force: true });
|
|
646
650
|
}
|
|
651
|
+
rmSync(join2(root, `${stem}.pages.json`), { force: true });
|
|
652
|
+
const ownedOcr = ownedOcrPattern(stem);
|
|
653
|
+
if (existsSync(ocrDir)) {
|
|
654
|
+
for (const f of readdirSync(ocrDir)) if (ownedOcr.test(f)) rmSync(join2(ocrDir, f), { force: true });
|
|
655
|
+
}
|
|
647
656
|
}
|
|
648
657
|
const lockId = randomBytes(6).toString("hex");
|
|
649
658
|
const stagingDir = join2(imagesDir, `.stage-${lockId}`);
|
|
@@ -651,7 +660,8 @@ function openBundle(root, requested, overwrite) {
|
|
|
651
660
|
mkdirSync2(stagingDir, { recursive: true });
|
|
652
661
|
const pagesStagingDir = join2(pagesDir, `.stage-${lockId}`);
|
|
653
662
|
const attachmentsStagingDir = join2(attachmentsDir, `.stage-${lockId}`);
|
|
654
|
-
|
|
663
|
+
const ocrStagingDir = join2(ocrDir, `.stage-${lockId}`);
|
|
664
|
+
return { root, stem, renamedFrom, renameReason, mdPath, lockPath, imagesDir, stagingDir, lockId, sheetsDir, sheetsStagingDir, pagesDir, pagesStagingDir, attachmentsDir, attachmentsStagingDir, ocrDir, ocrStagingDir, pageStatsPath: join2(root, `${stem}.pages.json`), ocrManifest: /* @__PURE__ */ new Set(), manifest: /* @__PURE__ */ new Set(), csvManifest: /* @__PURE__ */ new Set(), pageManifest: /* @__PURE__ */ new Set(), attachmentManifest: /* @__PURE__ */ new Set(), sourceMap: /* @__PURE__ */ new Map() };
|
|
655
665
|
} catch (e) {
|
|
656
666
|
rmSync(lockPath, { force: true });
|
|
657
667
|
throw e;
|
|
@@ -715,6 +725,28 @@ function publishPageImages(b, pageCount) {
|
|
|
715
725
|
}
|
|
716
726
|
rmSync(b.pagesStagingDir, { recursive: true, force: true });
|
|
717
727
|
}
|
|
728
|
+
function writePageStats(b, stats) {
|
|
729
|
+
writeFileSync2(b.pageStatsPath, `${JSON.stringify(stats, null, 2)}
|
|
730
|
+
`, "utf8");
|
|
731
|
+
}
|
|
732
|
+
function publishSidecars(b) {
|
|
733
|
+
const out = /* @__PURE__ */ new Map();
|
|
734
|
+
if (!existsSync(b.ocrStagingDir)) return out;
|
|
735
|
+
for (const dir of readdirSync(b.ocrStagingDir).sort()) {
|
|
736
|
+
const m = dir.match(/^p(\d+)$/);
|
|
737
|
+
if (!m) continue;
|
|
738
|
+
const pageDir = join2(b.ocrStagingDir, dir);
|
|
739
|
+
if (!existsSync(join2(pageDir, ".done"))) continue;
|
|
740
|
+
const file = `${b.stem}-${dir}.md`;
|
|
741
|
+
if (!existsSync(join2(pageDir, file))) continue;
|
|
742
|
+
mkdirSync2(b.ocrDir, { recursive: true });
|
|
743
|
+
renameSync(join2(pageDir, file), join2(b.ocrDir, file));
|
|
744
|
+
b.ocrManifest.add(file);
|
|
745
|
+
out.set(Number(m[1]), join2(b.ocrDir, file));
|
|
746
|
+
}
|
|
747
|
+
rmSync(b.ocrStagingDir, { recursive: true, force: true });
|
|
748
|
+
return out;
|
|
749
|
+
}
|
|
718
750
|
function publishAttachments(b) {
|
|
719
751
|
if (!existsSync(b.attachmentsStagingDir)) return;
|
|
720
752
|
for (const f of readdirSync(b.attachmentsStagingDir).sort()) {
|
|
@@ -772,6 +804,10 @@ function commitBundle(b, markdown) {
|
|
|
772
804
|
fs.rmSync(b.attachmentsStagingDir, { recursive: true, force: true });
|
|
773
805
|
} catch {
|
|
774
806
|
}
|
|
807
|
+
try {
|
|
808
|
+
fs.rmSync(b.ocrStagingDir, { recursive: true, force: true });
|
|
809
|
+
} catch {
|
|
810
|
+
}
|
|
775
811
|
try {
|
|
776
812
|
fs.rmSync(b.lockPath, { force: true });
|
|
777
813
|
} catch {
|
|
@@ -782,6 +818,9 @@ function abortBundle(b) {
|
|
|
782
818
|
for (const f of b.csvManifest) rmSync(join2(b.sheetsDir, f), { force: true });
|
|
783
819
|
for (const f of b.pageManifest) rmSync(join2(b.pagesDir, f), { force: true });
|
|
784
820
|
for (const f of b.attachmentManifest) rmSync(join2(b.attachmentsDir, f), { force: true });
|
|
821
|
+
for (const f of b.ocrManifest) rmSync(join2(b.ocrDir, f), { force: true });
|
|
822
|
+
rmSync(b.ocrStagingDir, { recursive: true, force: true });
|
|
823
|
+
rmSync(b.pageStatsPath, { force: true });
|
|
785
824
|
rmSync(`${b.mdPath}.tmp`, { force: true });
|
|
786
825
|
rmSync(b.stagingDir, { recursive: true, force: true });
|
|
787
826
|
rmSync(b.sheetsStagingDir, { recursive: true, force: true });
|
|
@@ -820,7 +859,21 @@ function compactRanges(nums, maxEntries = 20) {
|
|
|
820
859
|
}
|
|
821
860
|
var INSTALL_HINT = "install Tesseract (see doc/doc-to-md.md), then rerun with ocr=true";
|
|
822
861
|
var BARE_REASONS = ["fallback tier", "no Python backend"];
|
|
862
|
+
var pageList = (pages) => `page${pages.length === 1 ? "" : "s"} ${pages.join(", ")}`;
|
|
863
|
+
function forcedOcrLine(ocr) {
|
|
864
|
+
const clauses = [];
|
|
865
|
+
if (ocr.pages.length) clauses.push(`sidecars for ${pageList(ocr.pages)}`);
|
|
866
|
+
if (ocr.noText.length) clauses.push(`no text on ${pageList(ocr.noText)}`);
|
|
867
|
+
for (const page of ocr.ocrFailed) clauses.push(`failed on page ${page} (${ocr.ocrErrors[page] ?? "unknown error"})`);
|
|
868
|
+
if (ocr.budgetStopped.length) clauses.push(`budget-stopped ${pageList(ocr.budgetStopped)}`);
|
|
869
|
+
if (ocr.killed !== null) clauses.push(`page ${ocr.killed} killed the OCR child (timeout or crash - likely a compression bomb)`);
|
|
870
|
+
if (ocr.childError !== null) clauses.push(`OCR child failed before processing pages: ${ocr.childError}`);
|
|
871
|
+
if (ocr.notAttempted.length) clauses.push(`${pageList(ocr.notAttempted)} not attempted`);
|
|
872
|
+
const rerun = [...ocr.budgetStopped, ...ocr.notAttempted].sort((a, b) => a - b);
|
|
873
|
+
return `OCR: forced (${ocr.lang}) - ${clauses.join("; ")}${rerun.length ? ` - re-run with --pages ${rerun.join(",")}` : ""}`;
|
|
874
|
+
}
|
|
823
875
|
function ocrLine(ocr, type) {
|
|
876
|
+
if (ocr.mode === "all") return forcedOcrLine(ocr);
|
|
824
877
|
const image = type === "image";
|
|
825
878
|
const rerun = ocr.tesseract ? "rerun with ocr=true" : INSTALL_HINT;
|
|
826
879
|
switch (ocr.status) {
|
|
@@ -896,6 +949,8 @@ function formatHandle(h) {
|
|
|
896
949
|
if (h.sheetsDir) lines.push(`Sheets-Dir: ${h.sheetsDir}`);
|
|
897
950
|
if (h.pagesDir && h.pageImageCount > 0) lines.push(`Pages-Dir: ${h.pagesDir} (${h.pageImageCount} pages)`);
|
|
898
951
|
else if (h.pageImagesReason) lines.push(`Pages-Dir: none - ${h.pageImagesReason}`);
|
|
952
|
+
if (h.pageStatsPath) lines.push(`Page-Stats: ${h.pageStatsPath}`);
|
|
953
|
+
if (h.ocrDir) lines.push(`OCR-Dir: ${h.ocrDir}`);
|
|
899
954
|
lines.push(`Type: ${h.type} Engine: ${h.engine} Tier: ${h.tier}`);
|
|
900
955
|
lines.push(`Page-Count: ${pageCountLabel(h)} Pages: ${h.pages ? compactRanges(h.pages) : "all"} Images: ${h.imageCount} Size: ${formatSize2(h.bytes)} / ${h.lines} lines`);
|
|
901
956
|
if (h.degraded) lines.push(`Degraded: ${h.degraded}`);
|
|
@@ -945,6 +1000,7 @@ var DOC_TO_MD_OPTIONS = [
|
|
|
945
1000
|
{ key: "outputDir", flag: "--output-dir", type: "string", default: null, settable: false, help: "Bundle root for <stem>.md + images/; default a per-call temp dir. <stem> = basename without extension with [^A-Za-z0-9._-]+ -> _ (empty -> document); a second call on the same stem writes <stem>-2.md" },
|
|
946
1001
|
{ key: "overwrite", flag: "--overwrite", type: "bool", default: false, settable: false, help: "Replace an existing completed <stem>.md bundle" },
|
|
947
1002
|
{ key: "pageImages", flag: "--page-images", type: "bool", default: false, settable: false, help: "Also render every selected page to pages/<stem>-pNNN.<imageFormat> at imageDpi (PDF, PPTX, .doc, DOCX via LibreOffice); off by default" },
|
|
1003
|
+
{ key: "ocrMode", flag: "--ocr-mode", type: "enum", default: "textless", settable: false, enumValues: ["textless", "all"], help: "OCR policy: textless (default) OCRs only pages with an empty text layer, inline; all OCRs every selected page and writes the recognized text to ocr/<stem>-pNNN.md sidecars, leaving the Markdown untouched. all requires --ocr and an explicit --pages selection (PDF, PPTX, DOC)." },
|
|
948
1004
|
{ key: "primaryTimeoutMs", flag: "--primary-timeout", type: "int", default: 6e4, settable: true, env: "PI_DOC_TO_MD_CONVERT_TIMEOUT_MS", help: "pymupdf4llm tier and DOCX child (docx mode); also the unpdf tier" },
|
|
949
1005
|
{ key: "fallbackTimeoutMs", flag: "--fallback-timeout", type: "int", default: 3e4, settable: true, help: "PyMuPDF get_text tier (including DOCX LibreOffice fallback); also PDF and DOCX info and Excel rendered views" },
|
|
950
1006
|
{ key: "sofficeTimeoutMs", flag: "--soffice-timeout", type: "int", default: 12e4, settable: true, env: "PI_DOC_TO_MD_SOFFICE_TIMEOUT_MS", help: "DOCX/PPTX -> PDF via LibreOffice; also Excel rendered views" },
|
|
@@ -1051,11 +1107,28 @@ function resolveOptions(perCall, settings, env) {
|
|
|
1051
1107
|
out[d.key] = value;
|
|
1052
1108
|
}
|
|
1053
1109
|
const o = out;
|
|
1054
|
-
if (o.info && (o.pages !== null || o.outputDir !== null || o.overwrite || o.pageImages)) {
|
|
1055
|
-
throw new UsageError("--info cannot be combined with --pages, --output-dir, --overwrite or --
|
|
1110
|
+
if (o.info && (o.pages !== null || o.outputDir !== null || o.overwrite || o.pageImages || o.ocrMode === "all")) {
|
|
1111
|
+
throw new UsageError("--info cannot be combined with --pages, --output-dir, --overwrite, --page-images or --ocr-mode all");
|
|
1056
1112
|
}
|
|
1057
1113
|
return o;
|
|
1058
1114
|
}
|
|
1115
|
+
var USAGE_PATTERNS = [
|
|
1116
|
+
"Two-pass OCR (PDF, PPTX, DOC):",
|
|
1117
|
+
" 1. <cmd> report.pdf --output-dir out --json",
|
|
1118
|
+
' -> "pageStatsPath" points at out/report.pages.json; pages with few',
|
|
1119
|
+
' chars and high imageCoverage are scans. "savedTo" is the Markdown.',
|
|
1120
|
+
" 2. <cmd> report.pdf --output-dir out --ocr --ocr-mode all --pages 2,7 --json",
|
|
1121
|
+
' -> "ocr"."sidecars" maps 2 and 7 to out/ocr/report-2-p002.md and',
|
|
1122
|
+
" ...-p007.md (a second run in the same dir gets stem report-2);",
|
|
1123
|
+
" the Markdown of this run holds pages 2 and 7 only and equals what",
|
|
1124
|
+
" --pages 2,7 without --ocr would produce. Read sidecars and Markdown",
|
|
1125
|
+
" by the returned paths, never by guessing names.",
|
|
1126
|
+
" 3. --ocr-mode all refuses to run without --ocr and an explicit --pages.",
|
|
1127
|
+
' The "ocr" object (OCR: line) names failed, budget-stopped, killed and',
|
|
1128
|
+
" not-attempted pages and the exact --pages to re-run.",
|
|
1129
|
+
" Details: doc/doc-to-md.md (bundle contract, failure buckets)."
|
|
1130
|
+
].join("\n");
|
|
1131
|
+
var usagePatterns = (cmd) => USAGE_PATTERNS.replaceAll("<cmd>", cmd);
|
|
1059
1132
|
function renderHelp() {
|
|
1060
1133
|
const row = (d) => ` ${(d.flag ?? "<path>").padEnd(26)} ${d.help}${d.default !== null && d.key !== "info" && d.key !== "overwrite" ? ` (default ${d.default})` : ""}`;
|
|
1061
1134
|
return [
|
|
@@ -1067,8 +1140,10 @@ function renderHelp() {
|
|
|
1067
1140
|
"Tunables (also settable under quiver.docToMd in settings.json; per-call > settings > default):",
|
|
1068
1141
|
...DOC_TO_MD_OPTIONS.filter((d) => d.settable).map(row),
|
|
1069
1142
|
"",
|
|
1070
|
-
"Result: a handle (Saved-To, Images-Dir, Page-Count, Outline ...). Read the Saved-To file for the Markdown.",
|
|
1071
|
-
"Exit codes: 0 success, 1 runtime error, 2 usage error."
|
|
1143
|
+
"Result: a handle (Saved-To, Images-Dir, Page-Stats, Page-Count, Outline ...). Read the Saved-To file for the Markdown.",
|
|
1144
|
+
"Exit codes: 0 success, 1 runtime error, 2 usage error.",
|
|
1145
|
+
"",
|
|
1146
|
+
usagePatterns("pi-quiver doc-to-md")
|
|
1072
1147
|
].join("\n");
|
|
1073
1148
|
}
|
|
1074
1149
|
|
|
@@ -1547,7 +1622,38 @@ function reconcileRenderMarkers(md, renderPages, fmt, sourceMap, reason) {
|
|
|
1547
1622
|
if (/<!--rvs?:\d+-->/.test(md)) throw new Error("internal: unresolved render marker");
|
|
1548
1623
|
return md;
|
|
1549
1624
|
}
|
|
1550
|
-
var emptyOcr = (lang) => ({ status: "off", lang, textless: [], pages: [], noText: [], ocrFailed: [], budgetStopped: [], reason: null, tesseract: null });
|
|
1625
|
+
var emptyOcr = (lang) => ({ status: "off", lang, textless: [], pages: [], noText: [], ocrFailed: [], budgetStopped: [], reason: null, tesseract: null, mode: "textless", sidecars: {}, ocrErrors: {}, killed: null, notAttempted: [], childError: null });
|
|
1626
|
+
var emptyOutcome = () => ({ written: [], noText: [], ocrFailed: [], ocrErrors: {}, budgetStopped: [], killed: null, notAttempted: [], childError: null });
|
|
1627
|
+
var ocrPagesTag = (n) => `p${String(n).padStart(3, "0")}`;
|
|
1628
|
+
var SIDE_MARKER_RE = /^--- end of page\.page_number=\d+ ---$/;
|
|
1629
|
+
function sidecarHasText(dir) {
|
|
1630
|
+
const file = readdirSync2(dir).find((f) => f.endsWith(".md"));
|
|
1631
|
+
if (!file) return false;
|
|
1632
|
+
return readFileSync2(join3(dir, file), "utf8").split("\n").slice(1).some((l) => l.trim() && !SIDE_MARKER_RE.test(l));
|
|
1633
|
+
}
|
|
1634
|
+
function recoverOcrPages(stagingDir, pages, detail) {
|
|
1635
|
+
const out = emptyOutcome();
|
|
1636
|
+
const activePath = join3(stagingDir, "active");
|
|
1637
|
+
const active = existsSync2(activePath) ? Number(readFileSync2(activePath, "utf8").trim()) || null : null;
|
|
1638
|
+
let sawPage = false;
|
|
1639
|
+
for (const n of pages) {
|
|
1640
|
+
const dir = join3(stagingDir, ocrPagesTag(n));
|
|
1641
|
+
if (existsSync2(join3(dir, ".done"))) {
|
|
1642
|
+
sawPage = true;
|
|
1643
|
+
(sidecarHasText(dir) ? out.written : out.noText).push(n);
|
|
1644
|
+
} else if (existsSync2(join3(dir, ".failed"))) {
|
|
1645
|
+
sawPage = true;
|
|
1646
|
+
out.ocrFailed.push(n);
|
|
1647
|
+
out.ocrErrors[n] = readFileSync2(join3(dir, ".failed"), "utf8").trim() || "unknown error";
|
|
1648
|
+
} else if (n === active) out.killed = n;
|
|
1649
|
+
else out.notAttempted.push(n);
|
|
1650
|
+
}
|
|
1651
|
+
if (active === null && !sawPage) out.childError = detail;
|
|
1652
|
+
return out;
|
|
1653
|
+
}
|
|
1654
|
+
function outcomeFromChild(j) {
|
|
1655
|
+
return { ...emptyOutcome(), written: j.written ?? [], noText: j.noText ?? [], ocrFailed: j.ocrFailed ?? [], ocrErrors: j.ocrErrors ?? {}, budgetStopped: j.budgetStopped ?? [] };
|
|
1656
|
+
}
|
|
1551
1657
|
var OCR_SENTINEL_RE = /\x00OCR ([^\x00]*)\x00/g;
|
|
1552
1658
|
function resolveOcrLabels(md, sourceMap) {
|
|
1553
1659
|
return md.replace(OCR_SENTINEL_RE, (_, key) => {
|
|
@@ -1559,7 +1665,7 @@ function handleOcr(tier, type, o, json) {
|
|
|
1559
1665
|
if (tier === "unpdf") return o.ocr ? { ...emptyOcr(o.ocrLanguage), status: "unavailable", reason: "no Python backend" } : null;
|
|
1560
1666
|
const x = json.ocr;
|
|
1561
1667
|
if (!x || type !== "image" && !x.textless.length && !x.pages.length && !x.ocrFailed.length) return null;
|
|
1562
|
-
return x;
|
|
1668
|
+
return { ...emptyOcr(o.ocrLanguage), ...x };
|
|
1563
1669
|
}
|
|
1564
1670
|
async function convertDocument(o, signal, seams) {
|
|
1565
1671
|
const s = { backend: (c) => getBackend(c, void 0, signal), runTier: runTierReal, office: tryConvertOffice, ...seams };
|
|
@@ -1568,10 +1674,17 @@ async function convertDocument(o, signal, seams) {
|
|
|
1568
1674
|
if (!st || !st.isFile()) throw new Error(`Not a readable file: ${o.path}`);
|
|
1569
1675
|
if (st.size === 0) throw new Error(`empty file: ${o.path}`);
|
|
1570
1676
|
const type = classifyInput(inputPath);
|
|
1677
|
+
const forced = o.ocrMode === "all";
|
|
1678
|
+
if (forced) {
|
|
1679
|
+
if (!o.ocr) throw new UsageError("--ocr-mode all requires --ocr");
|
|
1680
|
+
if (o.pages === null) throw new UsageError('--ocr-mode all requires an explicit --pages selection (e.g. --pages 2,7); omitted pages and --pages "" mean all pages and are refused to keep OCR cost bounded');
|
|
1681
|
+
if (type !== "pdf" && type !== "pptx" && type !== "doc") throw new UsageError("--ocr-mode all applies to PDF, PPTX and DOC inputs only (DOCX pages are page-break segments, not PDF pages; convert the DOCX to PDF first)");
|
|
1682
|
+
}
|
|
1571
1683
|
if (o.pages && (type === "html" || type === "image" || type === "email")) throw new Error(`--pages does not apply to ${type === "html" ? "HTML files" : type === "image" ? "images" : "email"}`);
|
|
1572
1684
|
const isExcel = type === "xlsx" || type === "xlsm" || type === "xls";
|
|
1573
1685
|
if (isExcel && o.pages) throw new Error("--pages does not apply to spreadsheets: worksheets have no stable page numbering");
|
|
1574
1686
|
const backend = await s.backend({ pymupdfVersion: o.pymupdfVersion, warmTimeoutMs: o.warmTimeoutMs });
|
|
1687
|
+
if (forced && backend.kind === "none") throw new Error(`--ocr-mode all cannot run: no Python backend (${backend.reason})`);
|
|
1575
1688
|
if (isExcel && (backend.kind === "none" || !backend.xlsx)) throw new Error(`Excel conversion needs a Python backend with openpyxl, xlrd and pillow (${backend.kind === "none" ? backend.reason : `${backend.kind} lacks the Excel packages`}). ${EXCEL_REMEDY}`);
|
|
1576
1689
|
if (type === "email" && extname3(inputPath).toLowerCase() === ".msg" && (backend.kind === "none" || !backend.email)) throw new Error(`MSG conversion needs the extract-msg package. Python backend: ${backendState(backend, "found without extract-msg")}. Remedy: install uv, or pip install extract-msg markdownify into that Python`);
|
|
1577
1690
|
if (type === "email" && extname3(inputPath).toLowerCase() !== ".msg" && lacksDocx(backend)) throw new Error(`EML conversion needs the Python DOCX/HTML packages (mammoth, markdownify, python-docx). Python backend: ${backendState(backend, "found without mammoth/markdownify/python-docx")}. ${docxRemedy}`);
|
|
@@ -1580,7 +1693,7 @@ async function convertDocument(o, signal, seams) {
|
|
|
1580
1693
|
let office = null;
|
|
1581
1694
|
try {
|
|
1582
1695
|
let pdfPath = inputPath;
|
|
1583
|
-
const base = { path: inputPath, pages: o.pages, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
|
|
1696
|
+
const base = { path: inputPath, pages: o.pages, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr && !forced, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
|
|
1584
1697
|
let tier, engine, json, degraded = null, fallbackReason = null;
|
|
1585
1698
|
let explicitBreaks = null;
|
|
1586
1699
|
let notes = [];
|
|
@@ -1757,6 +1870,17 @@ async function convertDocument(o, signal, seams) {
|
|
|
1757
1870
|
if (type === "doc" || officeRoute !== null) degraded = DEGRADED_DOCX_OFFICE;
|
|
1758
1871
|
if (officeRoute !== null) fallbackReason = fallbackReason ? `${officeRoute}; ${fallbackReason}` : officeRoute;
|
|
1759
1872
|
if (tier === void 0 || engine === void 0 || json === void 0) throw new Error("internal: no tier produced output");
|
|
1873
|
+
const pageStats = json.pageStats ?? null;
|
|
1874
|
+
if (pageStats) writePageStats(b, pageStats);
|
|
1875
|
+
let ocr = handleOcr(tier, type, o, json);
|
|
1876
|
+
if (forced) {
|
|
1877
|
+
const r = await s.runTier("ocr-pages", { path: pdfPath, pages: o.pages, stem: b.stem, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs, stagingDir: b.ocrStagingDir, dpi: o.imageDpi, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion }, b, signal, o.primaryTimeoutMs, backend);
|
|
1878
|
+
if (signal?.aborted || !r.ok && "reason" in r && r.reason === "aborted") throw new Error("aborted");
|
|
1879
|
+
if (r.ok && r.json.status === "unavailable") throw new Error(`OCR unavailable: ${r.json.reason} (install Tesseract; see doc/doc-to-md.md)`);
|
|
1880
|
+
const outcome = r.ok && r.json.status === "ran" ? outcomeFromChild(r.json) : recoverOcrPages(b.ocrStagingDir, o.pages, !r.ok ? "userError" in r ? r.userError : `${r.reason}${detailSuffix(r)}` : "malformed child output");
|
|
1881
|
+
const sidecars = publishSidecars(b);
|
|
1882
|
+
ocr = { ...emptyOcr(o.ocrLanguage), ...json.ocr, status: "ran", reason: null, mode: "all", pages: outcome.written, noText: outcome.noText, ocrFailed: outcome.ocrFailed, ocrErrors: outcome.ocrErrors, budgetStopped: outcome.budgetStopped, killed: outcome.killed, notAttempted: outcome.notAttempted, childError: outcome.childError, sidecars: Object.fromEntries(sidecars) };
|
|
1883
|
+
}
|
|
1760
1884
|
if (!isExcel) notes = [...notes, ...json.notes ?? []];
|
|
1761
1885
|
if (b.renamedFrom) notes.splice(notes[0]?.startsWith("preview truncated:") ? 1 : 0, 0, `renamed to ${b.stem} (${b.renameReason})`);
|
|
1762
1886
|
const pageImagesReason = !o.pageImages || b.pageManifest.size || tier === "primary" || tier === "fallback" ? null : tier === "unpdf" ? "page images need the Python backend" : `${type} has no page geometry`;
|
|
@@ -1773,7 +1897,7 @@ async function convertDocument(o, signal, seams) {
|
|
|
1773
1897
|
` : "") + body;
|
|
1774
1898
|
commitBundle(b, markdown);
|
|
1775
1899
|
const outline = scanOutline(markdown, o.outlineMaxEntries);
|
|
1776
|
-
const details = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr
|
|
1900
|
+
const details = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr, pageStats, pageStatsPath: pageStats ? b.pageStatsPath : null, ocrDir: b.ocrManifest.size ? b.ocrDir : null };
|
|
1777
1901
|
return { output: formatHandle(details), details };
|
|
1778
1902
|
} catch (e) {
|
|
1779
1903
|
abortBundle(b);
|
|
@@ -1819,7 +1943,7 @@ async function inspectDocument(o, signal, seams) {
|
|
|
1819
1943
|
}
|
|
1820
1944
|
|
|
1821
1945
|
// bin/pi-quiver.ts
|
|
1822
|
-
var USAGE = 'Usage: pi-quiver fetch <url> [--method GET|HEAD|POST] [--header "K: V"]... [--body <str>] [--raw] [--timeout-ms <n>]\n pi-quiver doc-to-md [--json] [--info] [--page-images] [--pages <spec>] [--output-dir <dir>] [--overwrite] [tunable flags] <path> (--help for all flags)';
|
|
1946
|
+
var USAGE = 'Usage: pi-quiver fetch <url> [--method GET|HEAD|POST] [--header "K: V"]... [--body <str>] [--raw] [--timeout-ms <n>]\n pi-quiver doc-to-md [--json] [--info] [--page-images] [--pages <spec>] [--output-dir <dir>] [--overwrite] [--ocr] [--ocr-mode textless|all] [tunable flags] <path> (--help for all flags)';
|
|
1823
1947
|
function parseDocToMd(rest) {
|
|
1824
1948
|
if (rest.includes("--help") || rest.includes("-h")) return { ok: true, cmd: "doc-to-md-help" };
|
|
1825
1949
|
const json = rest.includes("--json");
|
|
@@ -1974,6 +2098,12 @@ ${USAGE}
|
|
|
1974
2098
|
`);
|
|
1975
2099
|
return 0;
|
|
1976
2100
|
} catch (err) {
|
|
2101
|
+
if (err instanceof UsageError) {
|
|
2102
|
+
process.stderr.write(`${err.message}
|
|
2103
|
+
${USAGE}
|
|
2104
|
+
`);
|
|
2105
|
+
return 2;
|
|
2106
|
+
}
|
|
1977
2107
|
process.stderr.write(`doc-to-md failed: ${err instanceof Error ? err.message : String(err)}
|
|
1978
2108
|
`);
|
|
1979
2109
|
return 1;
|
package/extensions/doc_to_md.ts
CHANGED
|
@@ -35,7 +35,7 @@ export default function docToMdExtension(pi: ExtensionAPI) {
|
|
|
35
35
|
name: "doc_to_md",
|
|
36
36
|
label: "Convert doc to Markdown bundle",
|
|
37
37
|
description:
|
|
38
|
-
"Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, email (.msg/.eml), HTML (.html/.htm) or image (.png .jpg .jpeg .tif .tiff .bmp .gif) to a Markdown bundle on disk and return a handle (Saved-To, Images-Dir, Page-Count, Outline, diagnostics) - the Markdown itself is never inlined; read the Saved-To file (offset/limit) for content. `info: true` returns page count, metadata and TOC (or the sheet inventory) without converting - use it to pick `pages`. `pages` selects inclusive 1-based pages (PDF/PPTX) or explicit-page-break segments (DOCX; rejected when the file has none); every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Images are always extracted into images/ with relative links. Primary engine pymupdf4llm, fallback PyMuPDF text (degraded, marked), pure-JS unpdf only when no Python backend exists. Excel yields a sheet inventory (worksheets + chartsheets), a full CSV per non-empty worksheet under sheets/, a bounded preview with formulas and cached values, merged/hidden disclosure, and rendered chart views when LibreOffice is available. DOCX converts directly (mammoth; python-docx text fallback) and keeps heading styles, so the handle Outline lists `L<line>` and `p<segment>` per heading; LibreOffice (soffice) is optional for DOCX (fallback route, degraded) and Excel rendered views, and required for PPTX. Use `outputDir` for a durable bundle; without it the bundle lands in a per-call temp dir. Input must be a local file path (use fetch first for URLs). HTML converts without Readability (markdownify; Turndown fallback): local and data: images are copied into images/, remote images stay links. Pages without a text layer keep their picture in images/, or pages/ with `pageImages: true`; image inputs keep their picture in images/. `pageImages: true` writes page renders to pages/<stem>-pNNN, and email attachments are saved under attachments/. OCR is off by default: `ocr: true` adds recognized text (labeled as OCR) when Tesseract language data is installed; the handle's OCR: line says what ran.",
|
|
38
|
+
"Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS, email (.msg/.eml), HTML (.html/.htm) or image (.png .jpg .jpeg .tif .tiff .bmp .gif) to a Markdown bundle on disk and return a handle (Saved-To, Images-Dir, Page-Count, Outline, diagnostics) - the Markdown itself is never inlined; read the Saved-To file (offset/limit) for content. `info: true` returns page count, metadata and TOC (or the sheet inventory) without converting - use it to pick `pages`. `pages` selects inclusive 1-based pages (PDF/PPTX) or explicit-page-break segments (DOCX; rejected when the file has none); every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Images are always extracted into images/ with relative links. Primary engine pymupdf4llm, fallback PyMuPDF text (degraded, marked), pure-JS unpdf only when no Python backend exists. Excel yields a sheet inventory (worksheets + chartsheets), a full CSV per non-empty worksheet under sheets/, a bounded preview with formulas and cached values, merged/hidden disclosure, and rendered chart views when LibreOffice is available. DOCX converts directly (mammoth; python-docx text fallback) and keeps heading styles, so the handle Outline lists `L<line>` and `p<segment>` per heading; LibreOffice (soffice) is optional for DOCX (fallback route, degraded) and Excel rendered views, and required for PPTX. Use `outputDir` for a durable bundle; without it the bundle lands in a per-call temp dir. Input must be a local file path (use fetch first for URLs). HTML converts without Readability (markdownify; Turndown fallback): local and data: images are copied into images/, remote images stay links. Pages without a text layer keep their picture in images/, or pages/ with `pageImages: true`; image inputs keep their picture in images/. `pageImages: true` writes page renders to pages/<stem>-pNNN, and email attachments are saved under attachments/. OCR is off by default: `ocr: true` adds recognized text (labeled as OCR) when Tesseract language data is installed; the handle's OCR: line says what ran. Two-pass OCR: read Page-Stats first, then re-run with ocr: true, ocrMode: \"all\" and an explicit pages selection to get ocr/ sidecars for the pages you name.",
|
|
39
39
|
promptSnippet: "Convert a local PDF/DOCX/DOC/PPTX/XLSX/XLSM/XLS/MSG/EML/HTML/image to a Markdown bundle (handle returned; read Saved-To)",
|
|
40
40
|
parameters: buildSchema(),
|
|
41
41
|
|
package/lib/doc-to-md-bundle.ts
CHANGED
|
@@ -1,23 +1,26 @@
|
|
|
1
1
|
/**
|
|
2
2
|
* Bundle protocol: a call owns `<stem>` for its whole duration via `<stem>.md.lock`;
|
|
3
|
-
* children stage assets under `images/`, `sheets/`, `pages/`, and `
|
|
4
|
-
* Node publishes them to stem-prefixed files in those
|
|
3
|
+
* children stage assets under `images/`, `sheets/`, `pages/`, `attachments/`, and `ocr/` staging dirs;
|
|
4
|
+
* Node publishes them to stem-prefixed files in those asset dirs and records every file it
|
|
5
5
|
* wrote in a manifest, and commits `<stem>.md` atomically (tmp + rename).
|
|
6
6
|
*/
|
|
7
7
|
import fs, { closeSync, existsSync, mkdirSync, openSync, readdirSync, readFileSync, renameSync, rmSync, statSync, unlinkSync, writeFileSync } from "node:fs";
|
|
8
8
|
import { tmpdir } from "node:os";
|
|
9
9
|
import { extname, join, resolve } from "node:path";
|
|
10
10
|
import { randomBytes } from "node:crypto";
|
|
11
|
+
import type { PageStat } from "./doc-to-md-handle.ts";
|
|
11
12
|
|
|
12
13
|
export interface Bundle {
|
|
13
14
|
root: string; stem: string; renamedFrom: string | null; renameReason: string | null; mdPath: string; lockPath: string; imagesDir: string; stagingDir: string; lockId: string;
|
|
14
15
|
sheetsDir: string; sheetsStagingDir: string;
|
|
15
16
|
pagesDir: string; pagesStagingDir: string;
|
|
16
17
|
attachmentsDir: string; attachmentsStagingDir: string;
|
|
18
|
+
ocrDir: string; ocrStagingDir: string; pageStatsPath: string;
|
|
17
19
|
manifest: Set<string>;
|
|
18
20
|
csvManifest: Set<string>;
|
|
19
21
|
pageManifest: Set<string>;
|
|
20
22
|
attachmentManifest: Set<string>;
|
|
23
|
+
ocrManifest: Set<string>;
|
|
21
24
|
sourceMap: Map<string, string>;
|
|
22
25
|
}
|
|
23
26
|
|
|
@@ -33,6 +36,8 @@ export function ownedCsvPattern(stem: string): RegExp {
|
|
|
33
36
|
|
|
34
37
|
export function ownedPagePattern(stem: string): RegExp { return new RegExp(`^${escRe(stem)}-p\\d+\\.[a-z0-9]+$`); }
|
|
35
38
|
|
|
39
|
+
export function ownedOcrPattern(stem: string): RegExp { return new RegExp(`^${escRe(stem)}-p\\d+\\.md$`); }
|
|
40
|
+
|
|
36
41
|
const FILE_LINK_RE = /\[[^\]]*\]\(\s*((?:sheets|attachments)\/[^)\s]+)\s*\)/g;
|
|
37
42
|
const IMG_LINK_RE = /!\[[^\]]*\]\(\s*(?:<([^>]*)>|([^)]*?))\s*\)|<img\b[^>]*\bsrc\s*=\s*(?:"([^"]*)"|'([^']*)'|([^\s>"']+))/gi;
|
|
38
43
|
|
|
@@ -75,6 +80,7 @@ export function openBundle(root: string, requested: string, overwrite: boolean):
|
|
|
75
80
|
const sheetsDir = join(root, "sheets");
|
|
76
81
|
const pagesDir = join(root, "pages");
|
|
77
82
|
const attachmentsDir = join(root, "attachments");
|
|
83
|
+
const ocrDir = join(root, "ocr");
|
|
78
84
|
try {
|
|
79
85
|
if (existsSync(mdPath) && overwrite) {
|
|
80
86
|
const owned = ownedPattern(stem), ownedCsv = ownedCsvPattern(stem), ownedPage = ownedPagePattern(stem);
|
|
@@ -87,6 +93,9 @@ export function openBundle(root: string, requested: string, overwrite: boolean):
|
|
|
87
93
|
if (existsSync(sheetsDir)) for (const f of readdirSync(sheetsDir)) if (ownedCsv.test(f)) rmSync(join(sheetsDir, f), { force: true });
|
|
88
94
|
if (existsSync(pagesDir)) for (const f of readdirSync(pagesDir)) if (ownedPage.test(f)) rmSync(join(pagesDir, f), { force: true });
|
|
89
95
|
if (existsSync(attachmentsDir)) for (const f of readdirSync(attachmentsDir)) if (attachments.has(f)) rmSync(join(attachmentsDir, f), { force: true });
|
|
96
|
+
rmSync(join(root, `${stem}.pages.json`), { force: true });
|
|
97
|
+
const ownedOcr = ownedOcrPattern(stem);
|
|
98
|
+
if (existsSync(ocrDir)) for (const f of readdirSync(ocrDir)) if (ownedOcr.test(f)) rmSync(join(ocrDir, f), { force: true });
|
|
90
99
|
}
|
|
91
100
|
const lockId = randomBytes(6).toString("hex");
|
|
92
101
|
const stagingDir = join(imagesDir, `.stage-${lockId}`);
|
|
@@ -94,7 +103,8 @@ export function openBundle(root: string, requested: string, overwrite: boolean):
|
|
|
94
103
|
mkdirSync(stagingDir, { recursive: true });
|
|
95
104
|
const pagesStagingDir = join(pagesDir, `.stage-${lockId}`);
|
|
96
105
|
const attachmentsStagingDir = join(attachmentsDir, `.stage-${lockId}`);
|
|
97
|
-
|
|
106
|
+
const ocrStagingDir = join(ocrDir, `.stage-${lockId}`);
|
|
107
|
+
return { root, stem, renamedFrom, renameReason, mdPath, lockPath, imagesDir, stagingDir, lockId, sheetsDir, sheetsStagingDir, pagesDir, pagesStagingDir, attachmentsDir, attachmentsStagingDir, ocrDir, ocrStagingDir, pageStatsPath: join(root, `${stem}.pages.json`), ocrManifest: new Set(), manifest: new Set(), csvManifest: new Set(), pageManifest: new Set(), attachmentManifest: new Set(), sourceMap: new Map() };
|
|
98
108
|
} catch (e) { rmSync(lockPath, { force: true }); throw e; }
|
|
99
109
|
}
|
|
100
110
|
|
|
@@ -160,6 +170,30 @@ export function publishPageImages(b: Bundle, pageCount: number): void {
|
|
|
160
170
|
rmSync(b.pagesStagingDir, { recursive: true, force: true });
|
|
161
171
|
}
|
|
162
172
|
|
|
173
|
+
export function writePageStats(b: Bundle, stats: PageStat[]): void {
|
|
174
|
+
writeFileSync(b.pageStatsPath, `${JSON.stringify(stats, null, 2)}\n`, "utf8");
|
|
175
|
+
}
|
|
176
|
+
|
|
177
|
+
/** Move every `.done`-gated `pNNN/<stem>-pNNN.md` into `ocr/`; partial page dirs and the checkpoint are dropped with the staging dir. Returns page -> absolute sidecar path. */
|
|
178
|
+
export function publishSidecars(b: Bundle): Map<number, string> {
|
|
179
|
+
const out = new Map<number, string>();
|
|
180
|
+
if (!existsSync(b.ocrStagingDir)) return out;
|
|
181
|
+
for (const dir of readdirSync(b.ocrStagingDir).sort()) {
|
|
182
|
+
const m = dir.match(/^p(\d+)$/);
|
|
183
|
+
if (!m) continue;
|
|
184
|
+
const pageDir = join(b.ocrStagingDir, dir);
|
|
185
|
+
if (!existsSync(join(pageDir, ".done"))) continue;
|
|
186
|
+
const file = `${b.stem}-${dir}.md`;
|
|
187
|
+
if (!existsSync(join(pageDir, file))) continue;
|
|
188
|
+
mkdirSync(b.ocrDir, { recursive: true });
|
|
189
|
+
renameSync(join(pageDir, file), join(b.ocrDir, file));
|
|
190
|
+
b.ocrManifest.add(file);
|
|
191
|
+
out.set(Number(m[1]), join(b.ocrDir, file));
|
|
192
|
+
}
|
|
193
|
+
rmSync(b.ocrStagingDir, { recursive: true, force: true });
|
|
194
|
+
return out;
|
|
195
|
+
}
|
|
196
|
+
|
|
163
197
|
export function publishAttachments(b: Bundle): void {
|
|
164
198
|
if (!existsSync(b.attachmentsStagingDir)) return;
|
|
165
199
|
for (const f of readdirSync(b.attachmentsStagingDir).sort()) {
|
|
@@ -208,6 +242,7 @@ export function commitBundle(b: Bundle, markdown: string): void {
|
|
|
208
242
|
try { fs.rmSync(b.sheetsStagingDir, { recursive: true, force: true }); } catch { /* best-effort */ }
|
|
209
243
|
try { fs.rmSync(b.pagesStagingDir, { recursive: true, force: true }); } catch { /* best-effort */ }
|
|
210
244
|
try { fs.rmSync(b.attachmentsStagingDir, { recursive: true, force: true }); } catch { /* best-effort */ }
|
|
245
|
+
try { fs.rmSync(b.ocrStagingDir, { recursive: true, force: true }); } catch { /* best-effort */ }
|
|
211
246
|
try { fs.rmSync(b.lockPath, { force: true }); } catch { /* Markdown is published; cleanup is best-effort. */ }
|
|
212
247
|
}
|
|
213
248
|
|
|
@@ -216,6 +251,9 @@ export function abortBundle(b: Bundle): void {
|
|
|
216
251
|
for (const f of b.csvManifest) rmSync(join(b.sheetsDir, f), { force: true });
|
|
217
252
|
for (const f of b.pageManifest) rmSync(join(b.pagesDir, f), { force: true });
|
|
218
253
|
for (const f of b.attachmentManifest) rmSync(join(b.attachmentsDir, f), { force: true });
|
|
254
|
+
for (const f of b.ocrManifest) rmSync(join(b.ocrDir, f), { force: true });
|
|
255
|
+
rmSync(b.ocrStagingDir, { recursive: true, force: true });
|
|
256
|
+
rmSync(b.pageStatsPath, { force: true });
|
|
219
257
|
rmSync(`${b.mdPath}.tmp`, { force: true });
|
|
220
258
|
rmSync(b.stagingDir, { recursive: true, force: true });
|
|
221
259
|
rmSync(b.sheetsStagingDir, { recursive: true, force: true });
|
package/lib/doc-to-md-core.ts
CHANGED
|
@@ -12,9 +12,9 @@ import { type ChildProcess, spawn } from "node:child_process";
|
|
|
12
12
|
import { homedir, tmpdir } from "node:os";
|
|
13
13
|
import { basename, dirname, extname, isAbsolute, join, resolve } from "node:path";
|
|
14
14
|
import { fileURLToPath, pathToFileURL } from "node:url";
|
|
15
|
-
import { type Bundle, abortBundle, commitBundle, openBundle, publishAttachments, publishPageImages, publishSheetCsvs, publishSheetImages, publishStaged, rewriteLinks, tempBundleRoot, validateImageLinks } from "./doc-to-md-bundle.ts";
|
|
16
|
-
import { type Engine, type HandleData, type InfoData, type OcrInfo, type Tier, formatHandle, formatInfoHandle, scanOutline, type SheetInfo, type TocEntry } from "./doc-to-md-handle.ts";
|
|
17
|
-
import { type DocToMdOptions, type InputType, IMAGE_EXTS, TUNABLE_DEFAULTS, classifyInput, sanitizeStem } from "./doc-to-md-options.ts";
|
|
15
|
+
import { type Bundle, abortBundle, commitBundle, openBundle, publishAttachments, publishSidecars, writePageStats, publishPageImages, publishSheetCsvs, publishSheetImages, publishStaged, rewriteLinks, tempBundleRoot, validateImageLinks } from "./doc-to-md-bundle.ts";
|
|
16
|
+
import { type Engine, type HandleData, type InfoData, type OcrInfo, type PageStat, type Tier, formatHandle, formatInfoHandle, scanOutline, type SheetInfo, type TocEntry } from "./doc-to-md-handle.ts";
|
|
17
|
+
import { type DocToMdOptions, type InputType, IMAGE_EXTS, TUNABLE_DEFAULTS, UsageError, classifyInput, sanitizeStem } from "./doc-to-md-options.ts";
|
|
18
18
|
|
|
19
19
|
export * from "./doc-to-md-options.ts";
|
|
20
20
|
export { compactRanges, formatHandle, formatInfoHandle, formatSize, scanOutline } from "./doc-to-md-handle.ts";
|
|
@@ -466,8 +466,8 @@ const lacksDocx = (b: Backend) => b.kind === "none" || !b.docx;
|
|
|
466
466
|
const clearStaging = (b: Pick<Bundle, "stagingDir">) => { for (const f of readdirSync(b.stagingDir)) rmSync(join(b.stagingDir, f), { recursive: true, force: true }); };
|
|
467
467
|
export const EXCEL_REMEDY = "Remedy: install uv, or pip install openpyxl xlrd pillow";
|
|
468
468
|
|
|
469
|
-
export type Mode = "html" | "image" | "info" | "pdf-primary" | "pdf-fallback" | "xlsx" | "pdf-text" | "render-pages" | "docx" | "email";
|
|
470
|
-
export interface TierJson { pageImages?: { page: number; file: string }[]; ocr?: OcrInfo; markdown?: string; pages?: number[]; pageCount?: number; emptyPages?: number[]; failedPages?: { page: number; error: string }[]; notes?: string[]; images?: { sheetIndex: number; file: string }[]; metadata?: Record<string, string>; toc?: [number, string, number | null][]; explicitBreaks?: number; engine?: string; degraded?: boolean; fallbackReason?: string | null; sheets?: SheetInfo[]; renderPages?: number[]; sheetCount?: number; ok?: boolean; reason?: string; rendered?: { idx: number; file: string; dpi: number }[]; failed?: { idx: number; reason: string }[]; }
|
|
469
|
+
export type Mode = "html" | "image" | "info" | "pdf-primary" | "pdf-fallback" | "xlsx" | "pdf-text" | "render-pages" | "docx" | "email" | "ocr-pages";
|
|
470
|
+
export interface TierJson { pageStats?: PageStat[]; status?: string; written?: number[]; noText?: number[]; ocrFailed?: number[]; ocrErrors?: Record<string, string>; budgetStopped?: number[]; pageImages?: { page: number; file: string }[]; ocr?: OcrInfo; markdown?: string; pages?: number[]; pageCount?: number; emptyPages?: number[]; failedPages?: { page: number; error: string }[]; notes?: string[]; images?: { sheetIndex: number; file: string }[]; metadata?: Record<string, string>; toc?: [number, string, number | null][]; explicitBreaks?: number; engine?: string; degraded?: boolean; fallbackReason?: string | null; sheets?: SheetInfo[]; renderPages?: number[]; sheetCount?: number; ok?: boolean; reason?: string; rendered?: { idx: number; file: string; dpi: number }[]; failed?: { idx: number; reason: string }[]; }
|
|
471
471
|
export type TierResult = { ok: true; json: TierJson } | { ok: false; reason: string; detail?: string } | { ok: false; userError: string; pageCount?: number };
|
|
472
472
|
|
|
473
473
|
export interface PipelineSeams {
|
|
@@ -490,7 +490,7 @@ function unpdfWorkerPath(): string {
|
|
|
490
490
|
], existsSync);
|
|
491
491
|
}
|
|
492
492
|
|
|
493
|
-
async function runTierReal(mode: Mode, childOptions: Record<string, unknown>, _b: Pick<Bundle, "stagingDir">, signal: AbortSignal | undefined, timeoutMs: number, backend: Backend): Promise<TierResult> {
|
|
493
|
+
export async function runTierReal(mode: Mode, childOptions: Record<string, unknown>, _b: Pick<Bundle, "stagingDir">, signal: AbortSignal | undefined, timeoutMs: number, backend: Backend): Promise<TierResult> {
|
|
494
494
|
const cfg = { pymupdfVersion: String(childOptions.pymupdfVersion), warmTimeoutMs: 0 };
|
|
495
495
|
let cmd: string, args: string[];
|
|
496
496
|
if (mode === "pdf-text" || (mode === "info" && backend.kind === "none")) { cmd = process.execPath; args = [unpdfWorkerPath(), mode]; }
|
|
@@ -536,7 +536,39 @@ export function reconcileRenderMarkers(md: string, renderPages: number[], fmt: s
|
|
|
536
536
|
return md;
|
|
537
537
|
}
|
|
538
538
|
|
|
539
|
-
export const emptyOcr = (lang: string): OcrInfo => ({ status: "off", lang, textless: [], pages: [], noText: [], ocrFailed: [], budgetStopped: [], reason: null, tesseract: null });
|
|
539
|
+
export const emptyOcr = (lang: string): OcrInfo => ({ status: "off", lang, textless: [], pages: [], noText: [], ocrFailed: [], budgetStopped: [], reason: null, tesseract: null, mode: "textless", sidecars: {}, ocrErrors: {}, killed: null, notAttempted: [], childError: null });
|
|
540
|
+
|
|
541
|
+
export interface OcrPagesOutcome { written: number[]; noText: number[]; ocrFailed: number[]; ocrErrors: Record<number, string>; budgetStopped: number[]; killed: number | null; notAttempted: number[]; childError: string | null; }
|
|
542
|
+
const emptyOutcome = (): OcrPagesOutcome => ({ written: [], noText: [], ocrFailed: [], ocrErrors: {}, budgetStopped: [], killed: null, notAttempted: [], childError: null });
|
|
543
|
+
const ocrPagesTag = (n: number) => `p${String(n).padStart(3, "0")}`;
|
|
544
|
+
const SIDE_MARKER_RE = /^--- end of page\.page_number=\d+ ---$/;
|
|
545
|
+
|
|
546
|
+
function sidecarHasText(dir: string): boolean {
|
|
547
|
+
const file = readdirSync(dir).find((f) => f.endsWith(".md"));
|
|
548
|
+
if (!file) return false;
|
|
549
|
+
return readFileSync(join(dir, file), "utf8").split("\n").slice(1).some((l) => l.trim() && !SIDE_MARKER_RE.test(l));
|
|
550
|
+
}
|
|
551
|
+
|
|
552
|
+
/** Rebuild the outcome from staging markers after the child died or returned garbage. */
|
|
553
|
+
export function recoverOcrPages(stagingDir: string, pages: number[], detail: string): OcrPagesOutcome {
|
|
554
|
+
const out = emptyOutcome();
|
|
555
|
+
const activePath = join(stagingDir, "active");
|
|
556
|
+
const active = existsSync(activePath) ? Number(readFileSync(activePath, "utf8").trim()) || null : null;
|
|
557
|
+
let sawPage = false;
|
|
558
|
+
for (const n of pages) {
|
|
559
|
+
const dir = join(stagingDir, ocrPagesTag(n));
|
|
560
|
+
if (existsSync(join(dir, ".done"))) { sawPage = true; (sidecarHasText(dir) ? out.written : out.noText).push(n); }
|
|
561
|
+
else if (existsSync(join(dir, ".failed"))) { sawPage = true; out.ocrFailed.push(n); out.ocrErrors[n] = readFileSync(join(dir, ".failed"), "utf8").trim() || "unknown error"; }
|
|
562
|
+
else if (n === active) out.killed = n;
|
|
563
|
+
else out.notAttempted.push(n);
|
|
564
|
+
}
|
|
565
|
+
if (active === null && !sawPage) out.childError = detail;
|
|
566
|
+
return out;
|
|
567
|
+
}
|
|
568
|
+
|
|
569
|
+
function outcomeFromChild(j: TierJson): OcrPagesOutcome {
|
|
570
|
+
return { ...emptyOutcome(), written: j.written ?? [], noText: j.noText ?? [], ocrFailed: j.ocrFailed ?? [], ocrErrors: j.ocrErrors ?? {}, budgetStopped: j.budgetStopped ?? [] };
|
|
571
|
+
}
|
|
540
572
|
const OCR_SENTINEL_RE = /\x00OCR ([^\x00]*)\x00/g;
|
|
541
573
|
|
|
542
574
|
/** Child OCR labels name staged files; the published name exists only after publishStaged. */
|
|
@@ -551,7 +583,7 @@ function handleOcr(tier: Tier, type: InputType, o: DocToMdOptions, json: TierJso
|
|
|
551
583
|
if (tier === "unpdf") return o.ocr ? { ...emptyOcr(o.ocrLanguage), status: "unavailable", reason: "no Python backend" } : null;
|
|
552
584
|
const x = json.ocr;
|
|
553
585
|
if (!x || (type !== "image" && !x.textless.length && !x.pages.length && !x.ocrFailed.length)) return null;
|
|
554
|
-
return x;
|
|
586
|
+
return { ...emptyOcr(o.ocrLanguage), ...x };
|
|
555
587
|
}
|
|
556
588
|
|
|
557
589
|
export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, seams?: Partial<PipelineSeams>): Promise<ConvertOutcome> {
|
|
@@ -561,10 +593,17 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
|
|
|
561
593
|
if (!st || !st.isFile()) throw new Error(`Not a readable file: ${o.path}`);
|
|
562
594
|
if (st.size === 0) throw new Error(`empty file: ${o.path}`);
|
|
563
595
|
const type = classifyInput(inputPath);
|
|
596
|
+
const forced = o.ocrMode === "all";
|
|
597
|
+
if (forced) {
|
|
598
|
+
if (!o.ocr) throw new UsageError("--ocr-mode all requires --ocr");
|
|
599
|
+
if (o.pages === null) throw new UsageError('--ocr-mode all requires an explicit --pages selection (e.g. --pages 2,7); omitted pages and --pages "" mean all pages and are refused to keep OCR cost bounded');
|
|
600
|
+
if (type !== "pdf" && type !== "pptx" && type !== "doc") throw new UsageError("--ocr-mode all applies to PDF, PPTX and DOC inputs only (DOCX pages are page-break segments, not PDF pages; convert the DOCX to PDF first)");
|
|
601
|
+
}
|
|
564
602
|
if (o.pages && (type === "html" || type === "image" || type === "email")) throw new Error(`--pages does not apply to ${type === "html" ? "HTML files" : type === "image" ? "images" : "email"}`);
|
|
565
603
|
const isExcel = type === "xlsx" || type === "xlsm" || type === "xls";
|
|
566
604
|
if (isExcel && o.pages) throw new Error("--pages does not apply to spreadsheets: worksheets have no stable page numbering");
|
|
567
605
|
const backend = await s.backend({ pymupdfVersion: o.pymupdfVersion, warmTimeoutMs: o.warmTimeoutMs });
|
|
606
|
+
if (forced && backend.kind === "none") throw new Error(`--ocr-mode all cannot run: no Python backend (${backend.reason})`);
|
|
568
607
|
if (isExcel && (backend.kind === "none" || !backend.xlsx)) throw new Error(`Excel conversion needs a Python backend with openpyxl, xlrd and pillow (${backend.kind === "none" ? backend.reason : `${backend.kind} lacks the Excel packages`}). ${EXCEL_REMEDY}`);
|
|
569
608
|
if (type === "email" && extname(inputPath).toLowerCase() === ".msg" && (backend.kind === "none" || !backend.email)) throw new Error(`MSG conversion needs the extract-msg package. Python backend: ${backendState(backend, "found without extract-msg")}. Remedy: install uv, or pip install extract-msg markdownify into that Python`);
|
|
570
609
|
if (type === "email" && extname(inputPath).toLowerCase() !== ".msg" && lacksDocx(backend)) throw new Error(`EML conversion needs the Python DOCX/HTML packages (mammoth, markdownify, python-docx). Python backend: ${backendState(backend, "found without mammoth/markdownify/python-docx")}. ${docxRemedy}`);
|
|
@@ -573,7 +612,7 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
|
|
|
573
612
|
let office: { pdfPath: string; cleanup: () => void } | null = null;
|
|
574
613
|
try {
|
|
575
614
|
let pdfPath = inputPath;
|
|
576
|
-
const base = { path: inputPath, pages: o.pages, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
|
|
615
|
+
const base = { path: inputPath, pages: o.pages, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, pageImages: o.pageImages, pagesStagingDir: b.pagesStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr && !forced, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
|
|
577
616
|
let tier: Tier | undefined, engine: Engine | undefined, json: TierJson | undefined, degraded: string | null = null, fallbackReason: string | null = null;
|
|
578
617
|
let explicitBreaks: number | null = null;
|
|
579
618
|
let notes: string[] = [];
|
|
@@ -704,6 +743,18 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
|
|
|
704
743
|
if (type === "doc" || officeRoute !== null) degraded = DEGRADED_DOCX_OFFICE;
|
|
705
744
|
if (officeRoute !== null) fallbackReason = fallbackReason ? `${officeRoute}; ${fallbackReason}` : officeRoute;
|
|
706
745
|
if (tier === undefined || engine === undefined || json === undefined) throw new Error("internal: no tier produced output");
|
|
746
|
+
const pageStats = json.pageStats ?? null;
|
|
747
|
+
if (pageStats) writePageStats(b, pageStats);
|
|
748
|
+
let ocr = handleOcr(tier, type, o, json);
|
|
749
|
+
if (forced) {
|
|
750
|
+
const r = await s.runTier("ocr-pages", { path: pdfPath, pages: o.pages, stem: b.stem, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs, stagingDir: b.ocrStagingDir, dpi: o.imageDpi, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion }, b, signal, o.primaryTimeoutMs, backend);
|
|
751
|
+
if (signal?.aborted || (!r.ok && "reason" in r && r.reason === "aborted")) throw new Error("aborted");
|
|
752
|
+
if (r.ok && r.json.status === "unavailable") throw new Error(`OCR unavailable: ${r.json.reason} (install Tesseract; see doc/doc-to-md.md)`);
|
|
753
|
+
const outcome = r.ok && r.json.status === "ran" ? outcomeFromChild(r.json)
|
|
754
|
+
: recoverOcrPages(b.ocrStagingDir, o.pages!, !r.ok ? ("userError" in r ? r.userError : `${r.reason}${detailSuffix(r)}`) : "malformed child output");
|
|
755
|
+
const sidecars = publishSidecars(b);
|
|
756
|
+
ocr = { ...emptyOcr(o.ocrLanguage), ...json.ocr, status: "ran", reason: null, mode: "all", pages: outcome.written, noText: outcome.noText, ocrFailed: outcome.ocrFailed, ocrErrors: outcome.ocrErrors, budgetStopped: outcome.budgetStopped, killed: outcome.killed, notAttempted: outcome.notAttempted, childError: outcome.childError, sidecars: Object.fromEntries(sidecars) };
|
|
757
|
+
}
|
|
707
758
|
if (!isExcel) notes = [...notes, ...(json.notes ?? [])];
|
|
708
759
|
if (b.renamedFrom) notes.splice(notes[0]?.startsWith("preview truncated:") ? 1 : 0, 0, `renamed to ${b.stem} (${b.renameReason})`);
|
|
709
760
|
const pageImagesReason = !o.pageImages || b.pageManifest.size || tier === "primary" || tier === "fallback" ? null
|
|
@@ -719,7 +770,7 @@ export async function convertDocument(o: DocToMdOptions, signal?: AbortSignal, s
|
|
|
719
770
|
const markdown = (head.length ? `${head.join("\n")}\n\n` : "") + body;
|
|
720
771
|
commitBundle(b, markdown);
|
|
721
772
|
const outline = scanOutline(markdown, o.outlineMaxEntries);
|
|
722
|
-
const details: DocToMdDetails = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr
|
|
773
|
+
const details: DocToMdDetails = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, pagesDir: b.pageManifest.size ? b.pagesDir : null, pageImageCount: b.pageManifest.size, pageImagesReason, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr, pageStats, pageStatsPath: pageStats ? b.pageStatsPath : null, ocrDir: b.ocrManifest.size ? b.ocrDir : null };
|
|
723
774
|
return { output: formatHandle(details), details };
|
|
724
775
|
} catch (e) { abortBundle(b); throw e; }
|
|
725
776
|
finally { office?.cleanup(); }
|
package/lib/doc-to-md-handle.ts
CHANGED
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
import type { InputType } from "./doc-to-md-options.ts";
|
|
1
|
+
import type { InputType, OcrMode } from "./doc-to-md-options.ts";
|
|
2
2
|
|
|
3
3
|
export type Tier = "primary" | "fallback" | "unpdf" | "excel" | "docx" | "html" | "image" | "email";
|
|
4
4
|
export type Engine = "pymupdf4llm" | "pymupdf-text" | "unpdf" | "openpyxl" | "xlrd" | "mammoth" | "python-docx" | "markdownify" | "turndown" | "copy" | "extract-msg" | "email";
|
|
@@ -18,6 +18,8 @@ export interface SheetInfo {
|
|
|
18
18
|
csv: string | null;
|
|
19
19
|
}
|
|
20
20
|
|
|
21
|
+
export type PageStat = { page: number; chars: number; images: number; imageCoverage: number } | { page: number; error: string };
|
|
22
|
+
|
|
21
23
|
export interface OcrInfo {
|
|
22
24
|
status: "off" | "unavailable" | "skipped" | "ran";
|
|
23
25
|
lang: string;
|
|
@@ -28,6 +30,12 @@ export interface OcrInfo {
|
|
|
28
30
|
budgetStopped: number[];
|
|
29
31
|
reason: string | null;
|
|
30
32
|
tesseract: boolean | null;
|
|
33
|
+
mode: OcrMode;
|
|
34
|
+
sidecars: Record<number, string>;
|
|
35
|
+
ocrErrors: Record<number, string>;
|
|
36
|
+
killed: number | null;
|
|
37
|
+
notAttempted: number[];
|
|
38
|
+
childError: string | null;
|
|
31
39
|
}
|
|
32
40
|
|
|
33
41
|
export interface HandleData {
|
|
@@ -35,6 +43,7 @@ export interface HandleData {
|
|
|
35
43
|
pageCount: number | null; pages: number[] | null; explicitBreaks: number | null; imageCount: number; pageImageCount: number; pageImagesReason: string | null; bytes: number; lines: number;
|
|
36
44
|
degraded: string | null; fallbackReason: string | null; failedPages: number[]; emptyPages: number[];
|
|
37
45
|
notes: string[]; outline: OutlineEntry[]; outlineTotal: number; ocr: OcrInfo | null;
|
|
46
|
+
pageStats: PageStat[] | null; pageStatsPath: string | null; ocrDir: string | null;
|
|
38
47
|
}
|
|
39
48
|
|
|
40
49
|
export interface InfoData {
|
|
@@ -72,7 +81,23 @@ export function compactRanges(nums: number[], maxEntries = 20): string {
|
|
|
72
81
|
const INSTALL_HINT = "install Tesseract (see doc/doc-to-md.md), then rerun with ocr=true";
|
|
73
82
|
const BARE_REASONS = ["fallback tier", "no Python backend"];
|
|
74
83
|
|
|
84
|
+
const pageList = (pages: number[]) => `page${pages.length === 1 ? "" : "s"} ${pages.join(", ")}`;
|
|
85
|
+
|
|
86
|
+
function forcedOcrLine(ocr: OcrInfo): string {
|
|
87
|
+
const clauses: string[] = [];
|
|
88
|
+
if (ocr.pages.length) clauses.push(`sidecars for ${pageList(ocr.pages)}`);
|
|
89
|
+
if (ocr.noText.length) clauses.push(`no text on ${pageList(ocr.noText)}`);
|
|
90
|
+
for (const page of ocr.ocrFailed) clauses.push(`failed on page ${page} (${ocr.ocrErrors[page] ?? "unknown error"})`);
|
|
91
|
+
if (ocr.budgetStopped.length) clauses.push(`budget-stopped ${pageList(ocr.budgetStopped)}`);
|
|
92
|
+
if (ocr.killed !== null) clauses.push(`page ${ocr.killed} killed the OCR child (timeout or crash - likely a compression bomb)`);
|
|
93
|
+
if (ocr.childError !== null) clauses.push(`OCR child failed before processing pages: ${ocr.childError}`);
|
|
94
|
+
if (ocr.notAttempted.length) clauses.push(`${pageList(ocr.notAttempted)} not attempted`);
|
|
95
|
+
const rerun = [...ocr.budgetStopped, ...ocr.notAttempted].sort((a, b) => a - b);
|
|
96
|
+
return `OCR: forced (${ocr.lang}) - ${clauses.join("; ")}${rerun.length ? ` - re-run with --pages ${rerun.join(",")}` : ""}`;
|
|
97
|
+
}
|
|
98
|
+
|
|
75
99
|
export function ocrLine(ocr: OcrInfo, type: InputType): string {
|
|
100
|
+
if (ocr.mode === "all") return forcedOcrLine(ocr);
|
|
76
101
|
const image = type === "image";
|
|
77
102
|
const rerun = ocr.tesseract ? "rerun with ocr=true" : INSTALL_HINT;
|
|
78
103
|
switch (ocr.status) {
|
|
@@ -143,6 +168,8 @@ export function formatHandle(h: HandleData): string {
|
|
|
143
168
|
if (h.sheetsDir) lines.push(`Sheets-Dir: ${h.sheetsDir}`);
|
|
144
169
|
if (h.pagesDir && h.pageImageCount > 0) lines.push(`Pages-Dir: ${h.pagesDir} (${h.pageImageCount} pages)`);
|
|
145
170
|
else if (h.pageImagesReason) lines.push(`Pages-Dir: none - ${h.pageImagesReason}`);
|
|
171
|
+
if (h.pageStatsPath) lines.push(`Page-Stats: ${h.pageStatsPath}`);
|
|
172
|
+
if (h.ocrDir) lines.push(`OCR-Dir: ${h.ocrDir}`);
|
|
146
173
|
lines.push(`Type: ${h.type} Engine: ${h.engine} Tier: ${h.tier}`);
|
|
147
174
|
lines.push(`Page-Count: ${pageCountLabel(h)} Pages: ${h.pages ? compactRanges(h.pages) : "all"} Images: ${h.imageCount} Size: ${formatSize(h.bytes)} / ${h.lines} lines`);
|
|
148
175
|
if (h.degraded) lines.push(`Degraded: ${h.degraded}`);
|
package/lib/doc-to-md-options.ts
CHANGED
|
@@ -6,6 +6,7 @@ import { extname } from "node:path";
|
|
|
6
6
|
|
|
7
7
|
export type InputType = "pdf" | "docx" | "doc" | "pptx" | "xlsx" | "xlsm" | "xls" | "html" | "image" | "email";
|
|
8
8
|
export type ImageFormat = "png" | "jpg";
|
|
9
|
+
export type OcrMode = "textless" | "all";
|
|
9
10
|
|
|
10
11
|
export interface Tunables {
|
|
11
12
|
primaryTimeoutMs: number;
|
|
@@ -29,6 +30,7 @@ export interface DocToMdOptions extends Tunables {
|
|
|
29
30
|
outputDir: string | null;
|
|
30
31
|
overwrite: boolean;
|
|
31
32
|
pageImages: boolean;
|
|
33
|
+
ocrMode: OcrMode;
|
|
32
34
|
}
|
|
33
35
|
|
|
34
36
|
/** What adapters pass in: intents as raw strings/booleans, tunables optional. */
|
|
@@ -39,6 +41,7 @@ export interface PerCallInput extends Partial<Tunables> {
|
|
|
39
41
|
outputDir?: string | null;
|
|
40
42
|
overwrite?: boolean;
|
|
41
43
|
pageImages?: boolean;
|
|
44
|
+
ocrMode?: OcrMode;
|
|
42
45
|
}
|
|
43
46
|
|
|
44
47
|
export type DescriptorType = "string" | "int" | "bool" | "pages" | "enum" | "version" | "lang";
|
|
@@ -68,6 +71,7 @@ export const DOC_TO_MD_OPTIONS: readonly OptionDescriptor[] = [
|
|
|
68
71
|
{ key: "outputDir", flag: "--output-dir", type: "string", default: null, settable: false, help: "Bundle root for <stem>.md + images/; default a per-call temp dir. <stem> = basename without extension with [^A-Za-z0-9._-]+ -> _ (empty -> document); a second call on the same stem writes <stem>-2.md" },
|
|
69
72
|
{ key: "overwrite", flag: "--overwrite", type: "bool", default: false, settable: false, help: "Replace an existing completed <stem>.md bundle" },
|
|
70
73
|
{ key: "pageImages", flag: "--page-images", type: "bool", default: false, settable: false, help: "Also render every selected page to pages/<stem>-pNNN.<imageFormat> at imageDpi (PDF, PPTX, .doc, DOCX via LibreOffice); off by default" },
|
|
74
|
+
{ key: "ocrMode", flag: "--ocr-mode", type: "enum", default: "textless", settable: false, enumValues: ["textless", "all"], help: "OCR policy: textless (default) OCRs only pages with an empty text layer, inline; all OCRs every selected page and writes the recognized text to ocr/<stem>-pNNN.md sidecars, leaving the Markdown untouched. all requires --ocr and an explicit --pages selection (PDF, PPTX, DOC)." },
|
|
71
75
|
{ key: "primaryTimeoutMs", flag: "--primary-timeout", type: "int", default: 60000, settable: true, env: "PI_DOC_TO_MD_CONVERT_TIMEOUT_MS", help: "pymupdf4llm tier and DOCX child (docx mode); also the unpdf tier" },
|
|
72
76
|
{ key: "fallbackTimeoutMs", flag: "--fallback-timeout", type: "int", default: 30000, settable: true, help: "PyMuPDF get_text tier (including DOCX LibreOffice fallback); also PDF and DOCX info and Excel rendered views" },
|
|
73
77
|
{ key: "sofficeTimeoutMs", flag: "--soffice-timeout", type: "int", default: 120000, settable: true, env: "PI_DOC_TO_MD_SOFFICE_TIMEOUT_MS", help: "DOCX/PPTX -> PDF via LibreOffice; also Excel rendered views" },
|
|
@@ -177,12 +181,32 @@ export function resolveOptions(perCall: PerCallInput, settings: Partial<Tunables
|
|
|
177
181
|
out[d.key] = value;
|
|
178
182
|
}
|
|
179
183
|
const o = out as unknown as DocToMdOptions;
|
|
180
|
-
if (o.info && (o.pages !== null || o.outputDir !== null || o.overwrite || o.pageImages)) {
|
|
181
|
-
throw new UsageError("--info cannot be combined with --pages, --output-dir, --overwrite or --
|
|
184
|
+
if (o.info && (o.pages !== null || o.outputDir !== null || o.overwrite || o.pageImages || o.ocrMode === "all")) {
|
|
185
|
+
throw new UsageError("--info cannot be combined with --pages, --output-dir, --overwrite, --page-images or --ocr-mode all");
|
|
182
186
|
}
|
|
183
187
|
return o;
|
|
184
188
|
}
|
|
185
189
|
|
|
190
|
+
/** Two-pass OCR recipe, rendered into --help and the generated skill; `<cmd>` is the command prefix. */
|
|
191
|
+
export const USAGE_PATTERNS = [
|
|
192
|
+
"Two-pass OCR (PDF, PPTX, DOC):",
|
|
193
|
+
" 1. <cmd> report.pdf --output-dir out --json",
|
|
194
|
+
' -> "pageStatsPath" points at out/report.pages.json; pages with few',
|
|
195
|
+
' chars and high imageCoverage are scans. "savedTo" is the Markdown.',
|
|
196
|
+
" 2. <cmd> report.pdf --output-dir out --ocr --ocr-mode all --pages 2,7 --json",
|
|
197
|
+
' -> "ocr"."sidecars" maps 2 and 7 to out/ocr/report-2-p002.md and',
|
|
198
|
+
" ...-p007.md (a second run in the same dir gets stem report-2);",
|
|
199
|
+
" the Markdown of this run holds pages 2 and 7 only and equals what",
|
|
200
|
+
" --pages 2,7 without --ocr would produce. Read sidecars and Markdown",
|
|
201
|
+
" by the returned paths, never by guessing names.",
|
|
202
|
+
" 3. --ocr-mode all refuses to run without --ocr and an explicit --pages.",
|
|
203
|
+
' The "ocr" object (OCR: line) names failed, budget-stopped, killed and',
|
|
204
|
+
" not-attempted pages and the exact --pages to re-run.",
|
|
205
|
+
" Details: doc/doc-to-md.md (bundle contract, failure buckets).",
|
|
206
|
+
].join("\n");
|
|
207
|
+
|
|
208
|
+
export const usagePatterns = (cmd: string): string => USAGE_PATTERNS.replaceAll("<cmd>", cmd);
|
|
209
|
+
|
|
186
210
|
export function renderHelp(): string {
|
|
187
211
|
const row = (d: OptionDescriptor) => ` ${(d.flag ?? "<path>").padEnd(26)} ${d.help}${d.default !== null && d.key !== "info" && d.key !== "overwrite" ? ` (default ${d.default})` : ""}`;
|
|
188
212
|
return [
|
|
@@ -190,7 +214,8 @@ export function renderHelp(): string {
|
|
|
190
214
|
"", "Per-call:", ...DOC_TO_MD_OPTIONS.filter((d) => !d.settable).map(row),
|
|
191
215
|
"", "Tunables (also settable under quiver.docToMd in settings.json; per-call > settings > default):",
|
|
192
216
|
...DOC_TO_MD_OPTIONS.filter((d) => d.settable).map(row),
|
|
193
|
-
"", "Result: a handle (Saved-To, Images-Dir, Page-Count, Outline ...). Read the Saved-To file for the Markdown.",
|
|
217
|
+
"", "Result: a handle (Saved-To, Images-Dir, Page-Stats, Page-Count, Outline ...). Read the Saved-To file for the Markdown.",
|
|
194
218
|
"Exit codes: 0 success, 1 runtime error, 2 usage error.",
|
|
219
|
+
"", usagePatterns("pi-quiver doc-to-md"),
|
|
195
220
|
].join("\n");
|
|
196
221
|
}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-quiver",
|
|
3
|
-
"version": "6.
|
|
3
|
+
"version": "6.9.0",
|
|
4
4
|
"description": "Personal pack of Pi coding-agent extensions: context-safe fetch, doc_to_md PDF/DOCX/PPTX-to-Markdown conversion, session naming, a themed ASCII startup header, Opus 4.8 fast mode, and a provider-stall watchdog.",
|
|
5
5
|
"author": "Jacek Juraszek",
|
|
6
6
|
"license": "MIT",
|
package/scripts/doc_to_md.py
CHANGED
|
@@ -253,6 +253,19 @@ def primary_page_markdown(doc, n, d, o, kw, write_images):
|
|
|
253
253
|
return rewrite_image_destinations(md, sources)
|
|
254
254
|
|
|
255
255
|
|
|
256
|
+
def page_stats(doc, n):
|
|
257
|
+
try:
|
|
258
|
+
page = doc[n - 1]
|
|
259
|
+
chars = len(page.get_text("text").strip())
|
|
260
|
+
infos = page.get_image_info()
|
|
261
|
+
area = page.rect.width * page.rect.height
|
|
262
|
+
covered = sum(max(0.0, (b[2] - b[0]) * (b[3] - b[1])) for b in (i["bbox"] for i in infos))
|
|
263
|
+
coverage = round(min(1.0, covered / area), 2) if area > 0 else 0.0
|
|
264
|
+
return {"page": n, "chars": chars, "images": len(infos), "imageCoverage": coverage}
|
|
265
|
+
except Exception as exc: # noqa: BLE001 - stats never cost a page its Markdown
|
|
266
|
+
return {"page": n, "error": f"{type(exc).__name__}: {exc}"[:300]}
|
|
267
|
+
|
|
268
|
+
|
|
256
269
|
def mode_pdf_primary(o):
|
|
257
270
|
import time
|
|
258
271
|
start = time.monotonic()
|
|
@@ -260,11 +273,13 @@ def mode_pdf_primary(o):
|
|
|
260
273
|
pages = check_pages(o.get("pages"), doc.page_count)
|
|
261
274
|
staging, out, empty, failed, notes = o["stagingDir"], [], [], [], []
|
|
262
275
|
page_images = []
|
|
276
|
+
stats = []
|
|
263
277
|
lang, budget = o.get("ocrLanguage", "eng"), o.get("ocrBudgetMs", 60000)
|
|
264
278
|
info = new_ocr(lang)
|
|
265
279
|
status = ocr_status(True, lang) if o.get("ocr") else None
|
|
266
280
|
ocr_ms, plain_ms = [], []
|
|
267
281
|
for i, n in enumerate(pages):
|
|
282
|
+
stats.append(page_stats(doc, n))
|
|
268
283
|
d = page_dir(staging, n)
|
|
269
284
|
try:
|
|
270
285
|
page = doc[n - 1]
|
|
@@ -325,7 +340,7 @@ def mode_pdf_primary(o):
|
|
|
325
340
|
if missing:
|
|
326
341
|
notes.append(f"Page images: {missing} of {len(pages)} unavailable")
|
|
327
342
|
return {"markdown": "\n\n".join(out) + "\n", "pages": pages, "pageCount": doc.page_count,
|
|
328
|
-
"emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": info, "pageImages": page_images}
|
|
343
|
+
"emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": info, "pageImages": page_images, "pageStats": stats}
|
|
329
344
|
|
|
330
345
|
|
|
331
346
|
def mode_pdf_fallback(o):
|
|
@@ -335,10 +350,12 @@ def mode_pdf_fallback(o):
|
|
|
335
350
|
keep = {int(k): v for k, v in (o.get("keepPages") or {}).items()}
|
|
336
351
|
staging, out, empty, failed = o["stagingDir"], [], [], []
|
|
337
352
|
page_images = []
|
|
353
|
+
stats = []
|
|
338
354
|
lang = o.get("ocrLanguage", "eng")
|
|
339
355
|
ocr_info = new_ocr(lang)
|
|
340
356
|
status = {"status": "unavailable", "reason": "fallback tier", "tesseract": None} if o.get("ocr") else None
|
|
341
357
|
for n in pages:
|
|
358
|
+
stats.append(page_stats(doc, n))
|
|
342
359
|
links = [f"" for f in keep.get(n, [])]
|
|
343
360
|
text = ""
|
|
344
361
|
page_pic = None
|
|
@@ -389,7 +406,64 @@ def mode_pdf_fallback(o):
|
|
|
389
406
|
if missing:
|
|
390
407
|
notes.append(f"Page images: {missing} of {len(pages)} unavailable")
|
|
391
408
|
return {"markdown": "\n\n".join(out) + "\n", "pages": pages, "pageCount": doc.page_count,
|
|
392
|
-
"emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": ocr_info, "pageImages": page_images}
|
|
409
|
+
"emptyPages": empty, "failedPages": failed, "notes": notes, "ocr": ocr_info, "pageImages": page_images, "pageStats": stats}
|
|
410
|
+
|
|
411
|
+
|
|
412
|
+
def mode_ocr_pages(o):
|
|
413
|
+
import time
|
|
414
|
+
start = time.monotonic()
|
|
415
|
+
lang, budget = o.get("ocrLanguage", "eng"), o.get("ocrBudgetMs", 60000)
|
|
416
|
+
status = ocr_status(True, lang)
|
|
417
|
+
if status["status"] != "ready":
|
|
418
|
+
return {"status": "unavailable", "reason": status["reason"]}
|
|
419
|
+
doc = open_pdf(o["path"])
|
|
420
|
+
pages = check_pages(o.get("pages"), doc.page_count)
|
|
421
|
+
staging, stem, dpi = o["stagingDir"], o["stem"], o.get("dpi", 150)
|
|
422
|
+
os.makedirs(staging, exist_ok=True)
|
|
423
|
+
active = os.path.join(staging, "active")
|
|
424
|
+
out = {"status": "ran", "written": [], "noText": [], "ocrFailed": [], "ocrErrors": {}, "budgetStopped": []}
|
|
425
|
+
ocr_ms = []
|
|
426
|
+
for i, n in enumerate(pages):
|
|
427
|
+
with open(active, "w") as fh:
|
|
428
|
+
fh.write(str(n))
|
|
429
|
+
if os.environ.get("DOC_TO_MD_OCR_STALL_PAGE") == str(n): # tests only: simulate a wedged page
|
|
430
|
+
time.sleep(3600)
|
|
431
|
+
elapsed = (time.monotonic() - start) * 1000
|
|
432
|
+
est = max(ocr_ms) if ocr_ms else OCR_EST_INITIAL_MS
|
|
433
|
+
if not ocr_admit(elapsed, est, 0, 0, budget):
|
|
434
|
+
out["budgetStopped"].extend(pages[i:])
|
|
435
|
+
os.remove(active)
|
|
436
|
+
break
|
|
437
|
+
tag = f"p{n:03d}"
|
|
438
|
+
d = os.path.join(staging, tag)
|
|
439
|
+
os.makedirs(d, exist_ok=True)
|
|
440
|
+
sidecar = os.path.join(d, f"{stem}-{tag}.md")
|
|
441
|
+
t0 = time.monotonic()
|
|
442
|
+
try:
|
|
443
|
+
page = doc[n - 1]
|
|
444
|
+
eff = clamped_dpi(page.rect.width, page.rect.height, dpi)
|
|
445
|
+
if eff is None:
|
|
446
|
+
raise RuntimeError(f"page cannot be rendered at a usable DPI ({page.rect.width:.0f} x {page.rect.height:.0f} pt)")
|
|
447
|
+
tp = page.get_textpage_ocr(full=True, language=lang, dpi=eff)
|
|
448
|
+
text = page.get_text("text", textpage=tp).strip()
|
|
449
|
+
header = f"<!-- OCR of page {n} (tesseract {lang}); recognized text, not the text layer -->"
|
|
450
|
+
with open(sidecar, "w", encoding="utf-8") as fh:
|
|
451
|
+
fh.write(header + "\n\n" + (text + "\n\n" if text else "") + SEP.format(n=n).strip("\n") + "\n")
|
|
452
|
+
mark_done(d)
|
|
453
|
+
(out["written"] if text else out["noText"]).append(n)
|
|
454
|
+
except Exception as exc: # noqa: BLE001 - one page never stops the pass
|
|
455
|
+
msg = f"{type(exc).__name__}: {exc}"
|
|
456
|
+
with open(os.path.join(d, ".failed"), "w", encoding="utf-8") as fh:
|
|
457
|
+
fh.write(msg)
|
|
458
|
+
try:
|
|
459
|
+
os.remove(sidecar)
|
|
460
|
+
except OSError:
|
|
461
|
+
pass
|
|
462
|
+
out["ocrFailed"].append(n)
|
|
463
|
+
out["ocrErrors"][str(n)] = msg
|
|
464
|
+
ocr_ms.append((time.monotonic() - t0) * 1000)
|
|
465
|
+
os.remove(active)
|
|
466
|
+
return out
|
|
393
467
|
|
|
394
468
|
|
|
395
469
|
def esc(v):
|
|
@@ -1249,12 +1323,12 @@ def mode_render_pages(o):
|
|
|
1249
1323
|
return {"ok": True, "rendered": rendered, "failed": failed}
|
|
1250
1324
|
|
|
1251
1325
|
|
|
1252
|
-
MODES = {"info": None, "pdf-primary": mode_pdf_primary, "pdf-fallback": mode_pdf_fallback, "xlsx": mode_xlsx, "render-pages": mode_render_pages, "docx": mode_docx, "html": mode_html, "image": mode_image, "email": mode_email}
|
|
1326
|
+
MODES = {"info": None, "pdf-primary": mode_pdf_primary, "pdf-fallback": mode_pdf_fallback, "xlsx": mode_xlsx, "render-pages": mode_render_pages, "docx": mode_docx, "html": mode_html, "image": mode_image, "email": mode_email, "ocr-pages": mode_ocr_pages}
|
|
1253
1327
|
|
|
1254
1328
|
|
|
1255
1329
|
def main():
|
|
1256
1330
|
if len(sys.argv) != 2 or sys.argv[1] not in MODES:
|
|
1257
|
-
print("usage: doc_to_md.py <info|pdf-primary|pdf-fallback|xlsx|render-pages|docx|html|image|email> (options JSON on stdin)", file=sys.stderr)
|
|
1331
|
+
print("usage: doc_to_md.py <info|pdf-primary|pdf-fallback|xlsx|render-pages|docx|html|image|email|ocr-pages> (options JSON on stdin)", file=sys.stderr)
|
|
1258
1332
|
return 1
|
|
1259
1333
|
mode = sys.argv[1]
|
|
1260
1334
|
o = json.loads(sys.stdin.read() or "{}")
|