pi-quiver 6.5.0 → 6.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +8 -0
- package/README.md +5 -2
- package/dist/bin/pi-quiver.js +314 -93
- package/extensions/doc_to_md.ts +2 -2
- package/lib/doc-to-md-bundle.ts +2 -1
- package/lib/doc-to-md-core.ts +113 -11
- package/lib/doc-to-md-handle.ts +43 -3
- package/lib/doc-to-md-options.ts +12 -5
- package/lib/fetch-core.ts +5 -0
- package/package.json +1 -1
- package/scripts/doc_to_md.py +274 -48
package/CHANGELOG.md
CHANGED
|
@@ -8,6 +8,14 @@ Published to npm as `pi-quiver` (`pi install npm:pi-quiver`). Pushing a
|
|
|
8
8
|
via OIDC trusted publishing. The release helper at
|
|
9
9
|
`.agents/skills/release/scripts/release.sh` cuts the tag; CI publishes.
|
|
10
10
|
|
|
11
|
+
## v6.6.0 - 2026-09-29
|
|
12
|
+
|
|
13
|
+
- `doc_to_md` converts local `.html`/`.htm` (markdownify; Readability-free Turndown fallback) with local and `data:` images copied into the bundle and remote images kept as links.
|
|
14
|
+
- `doc_to_md` accepts `.png .jpg .jpeg .tif .tiff .bmp .gif`; the image is copied into the bundle.
|
|
15
|
+
- Pages without a text layer now keep a page picture on both Python tiers (a scanned PDF used to convert to nothing).
|
|
16
|
+
- Opt-in OCR: `ocr` (default `false`) and `ocrLanguage` (default `eng`) tunables, `--ocr`/`--no-ocr`; needs Tesseract language data (optional OS dependency). A small-image gate and a time budget keep OCR from eating the primary timeout; the handle's `OCR:` line reports what ran.
|
|
17
|
+
- DOCX tables with `|` in a cell now render as `\|` instead of breaking the table; code blocks with a `language-*` class keep their fence language.
|
|
18
|
+
|
|
11
19
|
## v6.5.0 - 2026-09-29
|
|
12
20
|
|
|
13
21
|
- `doc_to_md` converts DOCX directly in the Python child (mammoth -> markdownify, python-docx text fallback) instead of LibreOffice -> PDF: heading styles survive as `#` headings, hyperlinks and footnotes are kept, pictures land under `images/` (#24). DOCX page semantics change: only author-inserted page breaks become `--- end of page.page_number=N ---` markers, `Page-Count` is suffixed `(explicit page breaks, not printed pages)` / `(no explicit page breaks)`, `pages` selects those segments and is rejected on a break-less file, and DOCX that falls back to LibreOffice (no DOCX-capable Python, or both DOCX engines failed) is marked degraded with `(LibreOffice pagination)` and rejects `pages`. LibreOffice is now optional for DOCX and still required for PPTX. Backend probe gains a `DOCX` line and pins `mammoth==1.13.0`, `markdownify==1.2.3`, `python-docx==1.2.0`; the managed venv moves to `doc-to-md-venv-v3` and the older venvs are removed after the first successful build.
|
package/README.md
CHANGED
|
@@ -16,7 +16,7 @@ But the moment an agent does that, one `fetch` or PDF read can dump hundreds of
|
|
|
16
16
|
|
|
17
17
|
## Why pi-quiver exists
|
|
18
18
|
|
|
19
|
-
`fetch` brings real web pages and GitHub issues/PRs into context and is size-gated by construction: over 32 KB or 1000 lines spills to a temp file with a preview and a grep/read hint. `doc_to_md` converts local PDF/DOCX/PPTX/XLSX/XLS files into a Markdown bundle on disk and returns only a concise handle. Ingestion is what makes data-driven work possible; bounded tool results keep it safe.
|
|
19
|
+
`fetch` brings real web pages and GitHub issues/PRs into context and is size-gated by construction: over 32 KB or 1000 lines spills to a temp file with a preview and a grep/read hint. `doc_to_md` converts local PDF/DOCX/PPTX/XLSX/XLS, HTML, and image files into a Markdown bundle on disk and returns only a concise handle; OCR is opt-in and uses optional Tesseract language data. Ingestion is what makes data-driven work possible; bounded tool results keep it safe.
|
|
20
20
|
|
|
21
21
|
`session-name`, `sword-header`, `fast-mode`, `provider-stall-watchdog`, and `slack` are opt-in ergonomics, recovery, and integration controls: session labeling, a themed startup header, Anthropic fast mode, semantic-stall recovery, and context-safe Slack search/threads/posting with repo-policy injection, `@name` mention resolution, cached emails, and per-call unfurl control.
|
|
22
22
|
|
|
@@ -65,7 +65,7 @@ A 300 KB changelog page never touches your context window - you get a preview an
|
|
|
65
65
|
| Extension | Tool | What it does |
|
|
66
66
|
| --- | --- | --- |
|
|
67
67
|
| `extensions/fetch.ts` | `fetch` | Retrieve URLs over HTTP(S). HTML -> Markdown (Readability extraction, Turndown conversion). Binary saved untouched to a temp file. GitHub issue/PR/repo/actions-run/actions-job URLs auto-route through `gh` (falls back to HTTP); failed runs/jobs include failed-step logs (best-effort, summary-only otherwise). Same size gate as `fetch`. Behavior lives in `lib/fetch-core.ts`; also exposed as the `pi-quiver fetch` CLI (see [Claude Code support](#claude-code-support)). |
|
|
68
|
-
| `extensions/doc_to_md.ts` | `doc_to_md` | Convert a local PDF/DOCX/PPTX/XLSX/XLS to a Markdown bundle on disk (`<stem>.md` + `images/` and spreadsheet `sheets/`) and return a handle (paths, page count, outline, diagnostics) - never inline Markdown. `info`
|
|
68
|
+
| `extensions/doc_to_md.ts` | `doc_to_md` | Convert a local PDF/DOCX/PPTX/XLSX/XLS, HTML, or image file to a Markdown bundle on disk (`<stem>.md` + `images/` and spreadsheet `sheets/`) and return a handle (paths, page count, outline, diagnostics) - never inline Markdown. `info` inspects PDF/Office inputs; `pages` selects 1-based PDF/Office pages (DOCX: explicit-page-break segments); HTML and image inputs reject both; every selected PDF/PPTX page or DOCX segment when `pageCount > 1` ends with `--- end of page.page_number=N ---`. Tiers: pymupdf4llm -> PyMuPDF text (degraded) -> unpdf worker (no Python only). Excel -> sheet inventory, full CSVs, bounded previews, and optional rendered views. HTML keeps local and `data:` images in the bundle and remote images as links; image inputs keep the original image. Scanned PDF pages keep a page picture on both Python tiers; opt-in OCR needs Tesseract language data. DOCX converts directly (mammoth, python-docx fallback); LibreOffice pagination marks the degraded route. The Outline lists `L<line>` and `p<page>` per heading. Settings under `quiver.docToMd`. |
|
|
69
69
|
| `extensions/session-name.ts` | `/session-name` | Manual + opt-in automatic session naming, naming rules and deny list, long-session revisits, and Ghostty/Herdr tab rename. OFF by default. |
|
|
70
70
|
| `extensions/sword-header.ts` | `/builtin-header` | Themed ASCII startup header replacing pi's default logo. OFF by default. |
|
|
71
71
|
| `extensions/fast-mode.ts` | `/fast` | Inject Anthropic fast-mode (`speed: "fast"` + `anthropic-beta: fast-mode-2026-02-01`) into every Claude Opus 4.8 / Opus 5 request, any thinking level. `--fast` flag + `/fast [on\|off\|status]`. OFF by default. |
|
|
@@ -134,6 +134,7 @@ The npm package's bundled JS deps install automatically on `pi install`. A few *
|
|
|
134
134
|
| --- | --- | --- |
|
|
135
135
|
| `gh` (GitHub CLI, installed + `gh auth login`) | `fetch` GitHub issue/PR/repo/actions-run/actions-job routing | Falls back to an HTTP fetch of the rendered page (private repos hit a login wall). |
|
|
136
136
|
| `uv` (+ managed Python 3.14, fetched on first use) | `doc_to_md` high-fidelity PDF and Excel conversion and DOCX (preferred route), with `pymupdf4llm`, `openpyxl`, `xlrd`, `pillow`, `mammoth`, `markdownify`, and `python-docx` | Falls back to a system Python >= 3.12 or one-time managed venv; PDF degrades to `unpdf` only when no capable Python exists. Excel requires the Python backend (no JS fallback). |
|
|
137
|
+
| Tesseract language data (`tesseract` on `PATH` or `TESSDATA_PREFIX`) | `doc_to_md` OCR with `ocr: true` for scanned pages and images | OCR is skipped and the handle reports `OCR: unavailable` (or an install hint when OCR is off); pages keep their picture. |
|
|
137
138
|
| LibreOffice (`soffice` on `PATH`) | `doc_to_md` PPTX conversion, the DOCX fallback route, and Excel rendered views | PPTX errors with a remedy; DOCX converts directly via the Python backend (LibreOffice fills in when that backend lacks the DOCX packages or its DOCX child exits 1); Excel omits rendered views. |
|
|
138
139
|
|
|
139
140
|
None is a hard install-time dependency of the package; they are tools you provide in the environment where pi runs.
|
|
@@ -304,6 +305,8 @@ Each setting can also be overridden per-process via `PI_QUIVER_SLACK_ENABLED`, `
|
|
|
304
305
|
| `imageFormat` | `png` | Rendered image format: `png` or `jpg`; embedded images retain their extension. |
|
|
305
306
|
| `maxOutputBytes` | `20000000` | Child stdout cap in bytes. |
|
|
306
307
|
| `outlineMaxEntries` | `40` | Heading outline, TOC, or sheet inventory entries in the handle. |
|
|
308
|
+
| `ocr` | `false` | Opt-in OCR for scanned pages and images when Tesseract language data is installed. |
|
|
309
|
+
| `ocrLanguage` | `eng` | Plain `+`-joined Tesseract language codes for OCR. |
|
|
307
310
|
|
|
308
311
|
A bundle is `<outputDir>/<stem>.md` plus `<outputDir>/images/` and, for spreadsheets with data, `<outputDir>/sheets/`; without `outputDir`, the tool creates a per-call temp root. A conversion owns `<stem>.md.lock` until it atomically publishes the Markdown. An existing `<stem>.md` fails the call unless `overwrite` is set; `overwrite` replaces that Markdown and the files it owns (`images/<stem>-p<N>-<n>.*`, `images/<stem>-s<idx>[-<n>].*`, `sheets/<stem>-s<idx>-<slug>.csv`), nothing else. Temp bundles are caller-owned - the tool never deletes a bundle it produced.
|
|
309
312
|
|
package/dist/bin/pi-quiver.js
CHANGED
|
@@ -1,12 +1,39 @@
|
|
|
1
1
|
#!/usr/bin/env node
|
|
2
|
-
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
2
|
+
var __defProp = Object.defineProperty;
|
|
3
|
+
var __getOwnPropNames = Object.getOwnPropertyNames;
|
|
4
|
+
var __esm = (fn, res, err) => function __init() {
|
|
5
|
+
if (err) throw err[0];
|
|
6
|
+
try {
|
|
7
|
+
return fn && (res = (0, fn[__getOwnPropNames(fn)[0]])(fn = 0)), res;
|
|
8
|
+
} catch (e) {
|
|
9
|
+
throw err = [e], e;
|
|
10
|
+
}
|
|
11
|
+
};
|
|
12
|
+
var __export = (target, all) => {
|
|
13
|
+
for (var name in all)
|
|
14
|
+
__defProp(target, name, { get: all[name], enumerable: true });
|
|
15
|
+
};
|
|
8
16
|
|
|
9
17
|
// lib/fetch-core.ts
|
|
18
|
+
var fetch_core_exports = {};
|
|
19
|
+
__export(fetch_core_exports, {
|
|
20
|
+
DEFAULT_TIMEOUT_MS: () => DEFAULT_TIMEOUT_MS,
|
|
21
|
+
applyGate: () => applyGate,
|
|
22
|
+
binaryExtension: () => binaryExtension,
|
|
23
|
+
buildGhArgs: () => buildGhArgs,
|
|
24
|
+
buildGhLogArgs: () => buildGhLogArgs,
|
|
25
|
+
categorize: () => categorize,
|
|
26
|
+
classifyGitHubTarget: () => classifyGitHubTarget,
|
|
27
|
+
collectBody: () => collectBody,
|
|
28
|
+
executeGhRouting: () => executeGhRouting,
|
|
29
|
+
fetchUrl: () => fetchUrl,
|
|
30
|
+
formatSize: () => formatSize,
|
|
31
|
+
htmlToMarkdown: () => htmlToMarkdown,
|
|
32
|
+
htmlToMarkdownRaw: () => htmlToMarkdownRaw,
|
|
33
|
+
planGhRouting: () => planGhRouting,
|
|
34
|
+
prettyJson: () => prettyJson,
|
|
35
|
+
runGh: () => runGh
|
|
36
|
+
});
|
|
10
37
|
import { mkdirSync, writeFileSync, createWriteStream } from "node:fs";
|
|
11
38
|
import { rm } from "node:fs/promises";
|
|
12
39
|
import { tmpdir } from "node:os";
|
|
@@ -27,24 +54,6 @@ function formatSize(bytes) {
|
|
|
27
54
|
return `${(bytes / (1024 * 1024)).toFixed(1)}MB`;
|
|
28
55
|
}
|
|
29
56
|
}
|
|
30
|
-
var RESERVED_OWNERS = /* @__PURE__ */ new Set([
|
|
31
|
-
"orgs",
|
|
32
|
-
"users",
|
|
33
|
-
"sponsors",
|
|
34
|
-
"topics",
|
|
35
|
-
"marketplace",
|
|
36
|
-
"apps",
|
|
37
|
-
"collections",
|
|
38
|
-
"stars",
|
|
39
|
-
"settings",
|
|
40
|
-
"notifications",
|
|
41
|
-
"codespaces",
|
|
42
|
-
"features",
|
|
43
|
-
"trending",
|
|
44
|
-
"security",
|
|
45
|
-
"customer-stories"
|
|
46
|
-
]);
|
|
47
|
-
var GH_NAME = /^[A-Za-z0-9._-]+$/;
|
|
48
57
|
function classifyGitHubTarget(url) {
|
|
49
58
|
const host = url.hostname.toLowerCase();
|
|
50
59
|
if (host !== "github.com" && host !== "www.github.com") return null;
|
|
@@ -82,22 +91,6 @@ function buildGhLogArgs(target) {
|
|
|
82
91
|
if (target.kind === "job") return ["run", "view", "--job", target.jobId, "--log-failed", "--repo", target.slug];
|
|
83
92
|
return null;
|
|
84
93
|
}
|
|
85
|
-
var GH_MAX_BUFFER = 1e7;
|
|
86
|
-
var execFileAsync = promisify(execFile);
|
|
87
|
-
var runGh = async (args, timeoutMs, signal) => {
|
|
88
|
-
try {
|
|
89
|
-
const { stdout } = await execFileAsync("gh", args, {
|
|
90
|
-
timeout: timeoutMs,
|
|
91
|
-
signal,
|
|
92
|
-
maxBuffer: GH_MAX_BUFFER,
|
|
93
|
-
encoding: "utf8"
|
|
94
|
-
});
|
|
95
|
-
if (!stdout.trim()) return { ok: false };
|
|
96
|
-
return { ok: true, stdout };
|
|
97
|
-
} catch {
|
|
98
|
-
return { ok: false };
|
|
99
|
-
}
|
|
100
|
-
};
|
|
101
94
|
function planGhRouting(params, url) {
|
|
102
95
|
if (params.raw) return null;
|
|
103
96
|
if ((params.method ?? "GET") !== "GET") return null;
|
|
@@ -176,16 +169,6 @@ async function executeGhRouting(params, url, signal, runner = runGh) {
|
|
|
176
169
|
}
|
|
177
170
|
return renderGhResult(target, gh.stdout, failedLogs);
|
|
178
171
|
}
|
|
179
|
-
var PARSABLE_MAX_BYTES = 1e6;
|
|
180
|
-
var BINARY_MAX_BYTES = 5e7;
|
|
181
|
-
var SNIFF_MAX_BYTES = 64e3;
|
|
182
|
-
var DEFAULT_TIMEOUT_MS = 2e4;
|
|
183
|
-
var INLINE_MAX_BYTES = 32e3;
|
|
184
|
-
var INLINE_MAX_LINES = 1e3;
|
|
185
|
-
var PREVIEW_LINES = 60;
|
|
186
|
-
var PREVIEW_MAX_BYTES = 4e3;
|
|
187
|
-
var FIREFOX_UA = "Mozilla/5.0 (Macintosh; Intel Mac OS X 14.7; rv:135.0) Gecko/20100101 Firefox/135.0";
|
|
188
|
-
var DEFAULT_ACCEPT = "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8";
|
|
189
172
|
function parseCharset(contentType) {
|
|
190
173
|
const m = /charset\s*=\s*"?([^";\s]+)"?/i.exec(contentType);
|
|
191
174
|
return (m?.[1] ?? "utf-8").trim().toLowerCase();
|
|
@@ -205,27 +188,9 @@ function buildPreview(body) {
|
|
|
205
188
|
}
|
|
206
189
|
return preview;
|
|
207
190
|
}
|
|
208
|
-
var turndownService = new TurndownService({
|
|
209
|
-
headingStyle: "atx",
|
|
210
|
-
codeBlockStyle: "fenced",
|
|
211
|
-
bulletListMarker: "-"
|
|
212
|
-
});
|
|
213
|
-
turndownService.use(gfm);
|
|
214
191
|
function mimeType(contentType) {
|
|
215
192
|
return contentType.split(";")[0].trim().toLowerCase();
|
|
216
193
|
}
|
|
217
|
-
var TEXT_ALLOWLIST = [
|
|
218
|
-
/^text\//,
|
|
219
|
-
/^application\/(json|xml|xhtml\+xml|javascript)$/,
|
|
220
|
-
/\+json$/,
|
|
221
|
-
/\+xml$/
|
|
222
|
-
];
|
|
223
|
-
var KNOWN_BINARY = [
|
|
224
|
-
/^audio\//,
|
|
225
|
-
/^video\//,
|
|
226
|
-
/^font\//,
|
|
227
|
-
/^application\/(pdf|zip|gzip|x-tar|x-7z-compressed|x-rar-compressed|wasm)$/
|
|
228
|
-
];
|
|
229
194
|
function categorize(contentType, sniff, raw) {
|
|
230
195
|
const mime = mimeType(contentType);
|
|
231
196
|
if (/^image\//.test(mime)) return "binary";
|
|
@@ -265,6 +230,9 @@ function htmlToMarkdown(html, url) {
|
|
|
265
230
|
${md}`;
|
|
266
231
|
return md;
|
|
267
232
|
}
|
|
233
|
+
function htmlToMarkdownRaw(html) {
|
|
234
|
+
return turndownService.turndown(html).trim();
|
|
235
|
+
}
|
|
268
236
|
function prettyJson(text) {
|
|
269
237
|
try {
|
|
270
238
|
return JSON.stringify(JSON.parse(text), null, 2);
|
|
@@ -300,18 +268,6 @@ function textExtension(category, contentType) {
|
|
|
300
268
|
if (category === "json") return "json";
|
|
301
269
|
return mimeType(contentType).includes("xml") ? "xml" : "txt";
|
|
302
270
|
}
|
|
303
|
-
var BINARY_EXT = {
|
|
304
|
-
"application/pdf": "pdf",
|
|
305
|
-
"application/zip": "zip",
|
|
306
|
-
"application/vnd.openxmlformats-officedocument.wordprocessingml.document": "docx",
|
|
307
|
-
"application/vnd.openxmlformats-officedocument.presentationml.presentation": "pptx",
|
|
308
|
-
"application/gzip": "gz",
|
|
309
|
-
"image/png": "png",
|
|
310
|
-
"image/jpeg": "jpg",
|
|
311
|
-
"image/gif": "gif",
|
|
312
|
-
"image/webp": "webp",
|
|
313
|
-
"image/svg+xml": "svg"
|
|
314
|
-
};
|
|
315
271
|
function binaryExtension(contentType) {
|
|
316
272
|
const mime = mimeType(contentType);
|
|
317
273
|
if (BINARY_EXT[mime]) return BINARY_EXT[mime];
|
|
@@ -515,12 +471,99 @@ async function fetchUrl(opts) {
|
|
|
515
471
|
opts.signal?.removeEventListener("abort", onAbort);
|
|
516
472
|
}
|
|
517
473
|
}
|
|
474
|
+
var RESERVED_OWNERS, GH_NAME, GH_MAX_BUFFER, execFileAsync, runGh, PARSABLE_MAX_BYTES, BINARY_MAX_BYTES, SNIFF_MAX_BYTES, DEFAULT_TIMEOUT_MS, INLINE_MAX_BYTES, INLINE_MAX_LINES, PREVIEW_LINES, PREVIEW_MAX_BYTES, FIREFOX_UA, DEFAULT_ACCEPT, turndownService, TEXT_ALLOWLIST, KNOWN_BINARY, BINARY_EXT;
|
|
475
|
+
var init_fetch_core = __esm({
|
|
476
|
+
"lib/fetch-core.ts"() {
|
|
477
|
+
"use strict";
|
|
478
|
+
RESERVED_OWNERS = /* @__PURE__ */ new Set([
|
|
479
|
+
"orgs",
|
|
480
|
+
"users",
|
|
481
|
+
"sponsors",
|
|
482
|
+
"topics",
|
|
483
|
+
"marketplace",
|
|
484
|
+
"apps",
|
|
485
|
+
"collections",
|
|
486
|
+
"stars",
|
|
487
|
+
"settings",
|
|
488
|
+
"notifications",
|
|
489
|
+
"codespaces",
|
|
490
|
+
"features",
|
|
491
|
+
"trending",
|
|
492
|
+
"security",
|
|
493
|
+
"customer-stories"
|
|
494
|
+
]);
|
|
495
|
+
GH_NAME = /^[A-Za-z0-9._-]+$/;
|
|
496
|
+
GH_MAX_BUFFER = 1e7;
|
|
497
|
+
execFileAsync = promisify(execFile);
|
|
498
|
+
runGh = async (args, timeoutMs, signal) => {
|
|
499
|
+
try {
|
|
500
|
+
const { stdout } = await execFileAsync("gh", args, {
|
|
501
|
+
timeout: timeoutMs,
|
|
502
|
+
signal,
|
|
503
|
+
maxBuffer: GH_MAX_BUFFER,
|
|
504
|
+
encoding: "utf8"
|
|
505
|
+
});
|
|
506
|
+
if (!stdout.trim()) return { ok: false };
|
|
507
|
+
return { ok: true, stdout };
|
|
508
|
+
} catch {
|
|
509
|
+
return { ok: false };
|
|
510
|
+
}
|
|
511
|
+
};
|
|
512
|
+
PARSABLE_MAX_BYTES = 1e6;
|
|
513
|
+
BINARY_MAX_BYTES = 5e7;
|
|
514
|
+
SNIFF_MAX_BYTES = 64e3;
|
|
515
|
+
DEFAULT_TIMEOUT_MS = 2e4;
|
|
516
|
+
INLINE_MAX_BYTES = 32e3;
|
|
517
|
+
INLINE_MAX_LINES = 1e3;
|
|
518
|
+
PREVIEW_LINES = 60;
|
|
519
|
+
PREVIEW_MAX_BYTES = 4e3;
|
|
520
|
+
FIREFOX_UA = "Mozilla/5.0 (Macintosh; Intel Mac OS X 14.7; rv:135.0) Gecko/20100101 Firefox/135.0";
|
|
521
|
+
DEFAULT_ACCEPT = "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8";
|
|
522
|
+
turndownService = new TurndownService({
|
|
523
|
+
headingStyle: "atx",
|
|
524
|
+
codeBlockStyle: "fenced",
|
|
525
|
+
bulletListMarker: "-"
|
|
526
|
+
});
|
|
527
|
+
turndownService.use(gfm);
|
|
528
|
+
TEXT_ALLOWLIST = [
|
|
529
|
+
/^text\//,
|
|
530
|
+
/^application\/(json|xml|xhtml\+xml|javascript)$/,
|
|
531
|
+
/\+json$/,
|
|
532
|
+
/\+xml$/
|
|
533
|
+
];
|
|
534
|
+
KNOWN_BINARY = [
|
|
535
|
+
/^audio\//,
|
|
536
|
+
/^video\//,
|
|
537
|
+
/^font\//,
|
|
538
|
+
/^application\/(pdf|zip|gzip|x-tar|x-7z-compressed|x-rar-compressed|wasm)$/
|
|
539
|
+
];
|
|
540
|
+
BINARY_EXT = {
|
|
541
|
+
"application/pdf": "pdf",
|
|
542
|
+
"application/zip": "zip",
|
|
543
|
+
"application/vnd.openxmlformats-officedocument.wordprocessingml.document": "docx",
|
|
544
|
+
"application/vnd.openxmlformats-officedocument.presentationml.presentation": "pptx",
|
|
545
|
+
"application/gzip": "gz",
|
|
546
|
+
"image/png": "png",
|
|
547
|
+
"image/jpeg": "jpg",
|
|
548
|
+
"image/gif": "gif",
|
|
549
|
+
"image/webp": "webp",
|
|
550
|
+
"image/svg+xml": "svg"
|
|
551
|
+
};
|
|
552
|
+
}
|
|
553
|
+
});
|
|
554
|
+
|
|
555
|
+
// bin/pi-quiver.ts
|
|
556
|
+
init_fetch_core();
|
|
557
|
+
import { existsSync as existsSync3, readFileSync as readFileSync2, realpathSync } from "node:fs";
|
|
558
|
+
import { homedir as homedir2 } from "node:os";
|
|
559
|
+
import { join as join4, resolve as resolve3 } from "node:path";
|
|
560
|
+
import { fileURLToPath as fileURLToPath2 } from "node:url";
|
|
518
561
|
|
|
519
562
|
// lib/doc-to-md-core.ts
|
|
520
|
-
import { existsSync as existsSync2, mkdtempSync, readdirSync as readdirSync2, renameSync as renameSync2, rmSync as rmSync2, statSync as statSync2 } from "node:fs";
|
|
563
|
+
import { copyFileSync, existsSync as existsSync2, mkdirSync as mkdirSync3, mkdtempSync, readFileSync, readdirSync as readdirSync2, renameSync as renameSync2, rmSync as rmSync2, statSync as statSync2, writeFileSync as writeFileSync3 } from "node:fs";
|
|
521
564
|
import { spawn } from "node:child_process";
|
|
522
565
|
import { homedir, tmpdir as tmpdir3 } from "node:os";
|
|
523
|
-
import { basename, dirname, extname as extname3, join as join3, resolve as resolve2 } from "node:path";
|
|
566
|
+
import { basename, dirname, extname as extname3, isAbsolute, join as join3, resolve as resolve2 } from "node:path";
|
|
524
567
|
import { fileURLToPath, pathToFileURL } from "node:url";
|
|
525
568
|
|
|
526
569
|
// lib/doc-to-md-bundle.ts
|
|
@@ -634,9 +677,10 @@ function rewriteLinks(md, sourceMap) {
|
|
|
634
677
|
return dest ? whole.split(target).join(dest) : whole;
|
|
635
678
|
});
|
|
636
679
|
}
|
|
637
|
-
function validateImageLinks(md, manifest, csvManifest = /* @__PURE__ */ new Set()) {
|
|
680
|
+
function validateImageLinks(md, manifest, csvManifest = /* @__PURE__ */ new Set(), html = false) {
|
|
638
681
|
for (const m of md.matchAll(IMG_LINK_RE)) {
|
|
639
682
|
const target = imageTarget(m);
|
|
683
|
+
if (html && !target.replace(/^\.\//, "").startsWith("p1/")) continue;
|
|
640
684
|
if (!target.startsWith("images/") || !manifest.has(target.slice("images/".length))) throw new Error(`unexpected image reference in output: ${target}`);
|
|
641
685
|
}
|
|
642
686
|
for (const m of md.matchAll(SHEET_LINK_RE)) {
|
|
@@ -697,6 +741,33 @@ function compactRanges(nums, maxEntries = 20) {
|
|
|
697
741
|
if (parts.length <= maxEntries) return parts.join(", ");
|
|
698
742
|
return `${parts.slice(0, maxEntries).join(", ")} (+${parts.length - maxEntries} more)`;
|
|
699
743
|
}
|
|
744
|
+
var INSTALL_HINT = "install Tesseract (see doc/doc-to-md.md), then rerun with ocr=true";
|
|
745
|
+
var BARE_REASONS = ["fallback tier", "no Python backend"];
|
|
746
|
+
function ocrLine(ocr, type) {
|
|
747
|
+
const image = type === "image";
|
|
748
|
+
const rerun = ocr.tesseract ? "rerun with ocr=true" : INSTALL_HINT;
|
|
749
|
+
switch (ocr.status) {
|
|
750
|
+
case "off":
|
|
751
|
+
return image ? `OCR: off; ${rerun}` : `OCR: off - ${ocr.textless.length} page(s) without a text layer; ${rerun}`;
|
|
752
|
+
case "unavailable": {
|
|
753
|
+
const reason = ocr.reason ?? "";
|
|
754
|
+
const bare = BARE_REASONS.includes(reason) || reason.startsWith("OCR child failed: ");
|
|
755
|
+
return `OCR: unavailable - ${reason}${bare ? "" : " (install Tesseract; see doc/doc-to-md.md)"}`;
|
|
756
|
+
}
|
|
757
|
+
case "skipped":
|
|
758
|
+
return "OCR: skipped - image too small";
|
|
759
|
+
case "ran": {
|
|
760
|
+
const clauses = [`OCR: ${ocr.pages.length} ${image ? "image" : "page(s)"} (${ocr.lang})`];
|
|
761
|
+
if (ocr.noText.length) clauses.push(`${ocr.noText.length} returned no text`);
|
|
762
|
+
if (ocr.ocrFailed.length) clauses.push(`${ocr.ocrFailed.length} failed and were converted without OCR`);
|
|
763
|
+
if (ocr.budgetStopped.length) {
|
|
764
|
+
const r = compactRanges(ocr.budgetStopped, Number.POSITIVE_INFINITY).replaceAll(", ", ",");
|
|
765
|
+
clauses.push(`time budget reached for pages=${r}; rerun with pages=${r} or raise primaryTimeoutMs`);
|
|
766
|
+
}
|
|
767
|
+
return clauses.join("; ");
|
|
768
|
+
}
|
|
769
|
+
}
|
|
770
|
+
}
|
|
700
771
|
var PAGE_MARKER_RE = /^--- end of page\.page_number=(\d+) ---$/;
|
|
701
772
|
function scanOutline(md, max) {
|
|
702
773
|
const entries = [];
|
|
@@ -754,6 +825,7 @@ function formatHandle(h) {
|
|
|
754
825
|
if (h.failedPages.length) fe.push(`Failed-Pages: ${compactRanges(h.failedPages)}`);
|
|
755
826
|
if (h.emptyPages.length) fe.push(`Empty-Pages: ${compactRanges(h.emptyPages)}`);
|
|
756
827
|
if (fe.length) lines.push(fe.join(" "));
|
|
828
|
+
if (h.ocr) lines.push(ocrLine(h.ocr, h.type));
|
|
757
829
|
h.notes.slice(0, NOTE_MAX_LINES).forEach((n, i) => lines.push(`${i === 0 ? "Notes: " : " "}${trunc(n, NOTE_MAX_CHARS)}`));
|
|
758
830
|
lines.push(...outlineLines(h.outline, h.outlineTotal));
|
|
759
831
|
return lines.join("\n");
|
|
@@ -785,8 +857,10 @@ var UsageError = class extends Error {
|
|
|
785
857
|
};
|
|
786
858
|
var MIN_PYMUPDF4LLM = "1.27.0";
|
|
787
859
|
var VERSION_RE = /^\d+(\.\d+)*$/;
|
|
860
|
+
var OCR_LANGUAGE_RE = /^[a-z][a-z0-9_]*(\+[a-z][a-z0-9_]*)*$/;
|
|
861
|
+
var IMAGE_EXTS = [".png", ".jpg", ".jpeg", ".tif", ".tiff", ".bmp", ".gif"];
|
|
788
862
|
var DOC_TO_MD_OPTIONS = [
|
|
789
|
-
{ key: "path", flag: null, type: "string", default: null, settable: false, help: "Local .pdf .docx .pptx .xlsx .xls file" },
|
|
863
|
+
{ key: "path", flag: null, type: "string", default: null, settable: false, help: "Local .pdf .docx .pptx .xlsx .xls .html .htm .png .jpg .jpeg .tif .tiff .bmp .gif file" },
|
|
790
864
|
{ key: "info", flag: "--info", type: "bool", default: false, settable: false, help: "Inspect only (page count, metadata, TOC or sheet inventory); no bundle" },
|
|
791
865
|
{ key: "pages", flag: "--pages", type: "pages", default: null, settable: false, help: 'Inclusive 1-based pages, e.g. "12-15" or "3,7,10-12" (PDF/DOCX/PPTX only); default all. DOCX: selects explicit-page-break segments; rejected when the file has none' },
|
|
792
866
|
{ key: "outputDir", flag: "--output-dir", type: "string", default: null, settable: false, help: "Bundle root for <stem>.md + images/; default a per-call temp dir" },
|
|
@@ -800,7 +874,9 @@ var DOC_TO_MD_OPTIONS = [
|
|
|
800
874
|
{ key: "imageDpi", flag: "--image-dpi", type: "int", default: 150, settable: true, help: "Render DPI for page images and Excel rendered views" },
|
|
801
875
|
{ key: "imageFormat", flag: "--image-format", type: "enum", default: "png", settable: true, enumValues: ["png", "jpg"], help: "Rendered image format (embedded images keep their native extension)" },
|
|
802
876
|
{ key: "maxOutputBytes", flag: "--max-output-bytes", type: "int", default: 2e7, settable: true, help: "Child stdout cap in bytes" },
|
|
803
|
-
{ key: "outlineMaxEntries", flag: "--outline-max-entries", type: "int", default: 40, settable: true, help: "Heading outline / TOC / sheet inventory cap in the handle" }
|
|
877
|
+
{ key: "outlineMaxEntries", flag: "--outline-max-entries", type: "int", default: 40, settable: true, help: "Heading outline / TOC / sheet inventory cap in the handle" },
|
|
878
|
+
{ key: "ocr", flag: "--ocr", type: "bool", default: false, settable: true, help: "Run OCR on pages without a text layer and on image inputs when Tesseract language data is installed; off by default (--no-ocr turns a settings-level true off)" },
|
|
879
|
+
{ key: "ocrLanguage", flag: "--ocr-language", type: "lang", default: "eng", settable: true, help: "Tesseract language code(s), +-joined, e.g. deu+eng" }
|
|
804
880
|
];
|
|
805
881
|
var TUNABLE_DEFAULTS = Object.fromEntries(
|
|
806
882
|
DOC_TO_MD_OPTIONS.filter((d) => d.settable).map((d) => [d.key, d.default])
|
|
@@ -827,6 +903,8 @@ function coerceValue(d, raw, fromString = false) {
|
|
|
827
903
|
if (typeof raw !== "string" || !VERSION_RE.test(raw)) return { ok: false, reason: "must be digits and dots" };
|
|
828
904
|
if (!versionAtLeast(raw, MIN_PYMUPDF4LLM)) return { ok: false, reason: `must be >= ${MIN_PYMUPDF4LLM}` };
|
|
829
905
|
return { ok: true, value: raw };
|
|
906
|
+
case "lang":
|
|
907
|
+
return typeof raw === "string" && OCR_LANGUAGE_RE.test(raw) ? { ok: true, value: raw } : { ok: false, reason: "must be Tesseract language codes joined by + (e.g. eng, deu+eng)" };
|
|
830
908
|
case "string":
|
|
831
909
|
return typeof raw === "string" && raw.length > 0 ? { ok: true, value: raw } : { ok: false, reason: "must be a non-empty string" };
|
|
832
910
|
case "pages":
|
|
@@ -865,10 +943,10 @@ function sanitizeStem(base) {
|
|
|
865
943
|
const s = base.replace(/[^A-Za-z0-9._-]+/g, "_");
|
|
866
944
|
return s.length ? s : "document";
|
|
867
945
|
}
|
|
868
|
-
var SUPPORTED = { ".pdf": "pdf", ".docx": "docx", ".pptx": "pptx", ".xlsx": "xlsx", ".xls": "xls" };
|
|
946
|
+
var SUPPORTED = { ".pdf": "pdf", ".docx": "docx", ".pptx": "pptx", ".xlsx": "xlsx", ".xls": "xls", ".html": "html", ".htm": "html", ...Object.fromEntries(IMAGE_EXTS.map((ext) => [ext, "image"])) };
|
|
869
947
|
function classifyInput(filePath) {
|
|
870
948
|
const t = SUPPORTED[extname2(filePath).toLowerCase()];
|
|
871
|
-
if (!t) throw new Error(`Unsupported file type "${extname2(filePath) || "(none)"}"; supported: .
|
|
949
|
+
if (!t) throw new Error(`Unsupported file type "${extname2(filePath) || "(none)"}"; supported: ${Object.keys(SUPPORTED).join(", ")}`);
|
|
872
950
|
return t;
|
|
873
951
|
}
|
|
874
952
|
function resolveOptions(perCall, settings, env) {
|
|
@@ -1247,6 +1325,76 @@ function officeFailure(r) {
|
|
|
1247
1325
|
return new Error(`soffice failed (code=${r.code} timedOut=${r.timedOut}): ${r.stderr}`);
|
|
1248
1326
|
}
|
|
1249
1327
|
var DEGRADED_TEXT = "PyMuPDF text extraction - layout/tables not preserved";
|
|
1328
|
+
var DEGRADED_HTML_TURNDOWN = "Turndown HTML conversion - definition lists and headerless tables not preserved";
|
|
1329
|
+
var DATA_IMAGE_RE = /^data:image\/(png|jpeg|gif|bmp|tiff);base64,([A-Za-z0-9+/=\s]+)$/i;
|
|
1330
|
+
var DATA_EXT = { png: ".png", jpeg: ".jpg", gif: ".gif", bmp: ".bmp", tiff: ".tif" };
|
|
1331
|
+
var DATA_SIGNATURES = {
|
|
1332
|
+
png: [[137, 80, 78, 71]],
|
|
1333
|
+
jpeg: [[255, 216, 255]],
|
|
1334
|
+
gif: [[71, 73, 70, 56]],
|
|
1335
|
+
bmp: [[66, 77]],
|
|
1336
|
+
tiff: [[73, 73, 42, 0], [77, 77, 0, 42]]
|
|
1337
|
+
};
|
|
1338
|
+
async function prepareHtml(inputPath, stagingDir) {
|
|
1339
|
+
const { JSDOM: JSDOM2 } = await import("jsdom");
|
|
1340
|
+
const bytes = readFileSync(inputPath);
|
|
1341
|
+
let source = bytes;
|
|
1342
|
+
try {
|
|
1343
|
+
source = new TextDecoder("utf-8", { fatal: true }).decode(bytes);
|
|
1344
|
+
} catch {
|
|
1345
|
+
}
|
|
1346
|
+
const dom = new JSDOM2(source);
|
|
1347
|
+
const doc = dom.window.document;
|
|
1348
|
+
const title = doc.querySelector("title")?.textContent?.trim();
|
|
1349
|
+
if (title && !doc.body.querySelector("h1")) {
|
|
1350
|
+
const h1 = doc.createElement("h1");
|
|
1351
|
+
h1.textContent = title;
|
|
1352
|
+
doc.body.prepend(h1);
|
|
1353
|
+
}
|
|
1354
|
+
for (const el of [...doc.querySelectorAll("head, script, style, noscript, template")]) el.remove();
|
|
1355
|
+
const dir = join3(stagingDir, "p1");
|
|
1356
|
+
mkdirSync3(dir, { recursive: true });
|
|
1357
|
+
let k = 0, missing = 0;
|
|
1358
|
+
for (const img of [...doc.querySelectorAll("img")]) {
|
|
1359
|
+
const src = (img.getAttribute("src") ?? "").trim();
|
|
1360
|
+
const alt = img.getAttribute("alt") ?? "";
|
|
1361
|
+
if (/^(?:https?:)?\/\//i.test(src)) {
|
|
1362
|
+
const a = doc.createElement("a");
|
|
1363
|
+
const url = src.startsWith("//") ? `https:${src}` : src;
|
|
1364
|
+
a.setAttribute("href", url);
|
|
1365
|
+
a.textContent = alt || url;
|
|
1366
|
+
img.replaceWith(a);
|
|
1367
|
+
continue;
|
|
1368
|
+
}
|
|
1369
|
+
let target = null;
|
|
1370
|
+
const data = src.match(DATA_IMAGE_RE);
|
|
1371
|
+
if (data) {
|
|
1372
|
+
const buf = Buffer.from(data[2].replace(/\s+/g, ""), "base64");
|
|
1373
|
+
if (DATA_SIGNATURES[data[1].toLowerCase()].some((signature) => signature.every((byte, i) => buf[i] === byte))) {
|
|
1374
|
+
target = `p1/${++k}${DATA_EXT[data[1].toLowerCase()]}`;
|
|
1375
|
+
writeFileSync3(join3(stagingDir, target), buf);
|
|
1376
|
+
}
|
|
1377
|
+
} else if (src && !/^[a-z][a-z0-9+.-]*:/i.test(src) && !isAbsolute(src)) {
|
|
1378
|
+
let file = null;
|
|
1379
|
+
try {
|
|
1380
|
+
file = resolve2(dirname(inputPath), decodeURIComponent(src.split(/[?#]/)[0]));
|
|
1381
|
+
} catch {
|
|
1382
|
+
}
|
|
1383
|
+
const ext = file ? extname3(file).toLowerCase() : "";
|
|
1384
|
+
if (file && IMAGE_EXTS.includes(ext) && statSync2(file, { throwIfNoEntry: false })?.isFile()) {
|
|
1385
|
+
target = `p1/${++k}${ext}`;
|
|
1386
|
+
copyFileSync(file, join3(stagingDir, target));
|
|
1387
|
+
}
|
|
1388
|
+
}
|
|
1389
|
+
if (target) img.setAttribute("src", target);
|
|
1390
|
+
else {
|
|
1391
|
+
missing++;
|
|
1392
|
+
img.replaceWith(doc.createTextNode(alt));
|
|
1393
|
+
}
|
|
1394
|
+
}
|
|
1395
|
+
writeFileSync3(join3(dir, ".done"), "");
|
|
1396
|
+
return { html: dom.serialize(), missing };
|
|
1397
|
+
}
|
|
1250
1398
|
var DEGRADED_UNPDF = "unpdf text extraction - structure not preserved";
|
|
1251
1399
|
var DEGRADED_DOCX_TEXT = "python-docx text extraction - footnotes, hyperlinks, images not preserved";
|
|
1252
1400
|
var DEGRADED_DOCX_OFFICE = "LibreOffice PDF route - heading styles and explicit page breaks not preserved; page numbers are LibreOffice pagination";
|
|
@@ -1312,12 +1460,27 @@ function reconcileRenderMarkers(md, renderPages, fmt, sourceMap, reason) {
|
|
|
1312
1460
|
if (/<!--rvs?:\d+-->/.test(md)) throw new Error("internal: unresolved render marker");
|
|
1313
1461
|
return md;
|
|
1314
1462
|
}
|
|
1463
|
+
var emptyOcr = (lang) => ({ status: "off", lang, textless: [], pages: [], noText: [], ocrFailed: [], budgetStopped: [], reason: null, tesseract: null });
|
|
1464
|
+
var OCR_SENTINEL_RE = /\x00OCR ([^\x00]*)\x00/g;
|
|
1465
|
+
function resolveOcrLabels(md, sourceMap) {
|
|
1466
|
+
return md.replace(OCR_SENTINEL_RE, (_, key) => {
|
|
1467
|
+
const dest = sourceMap.get(key);
|
|
1468
|
+
return dest ? `> Text recognized in ${dest} (OCR, may contain recognition errors):` : "> Text recognized by OCR (source image missing):";
|
|
1469
|
+
});
|
|
1470
|
+
}
|
|
1471
|
+
function handleOcr(tier, type, o, json) {
|
|
1472
|
+
if (tier === "unpdf") return o.ocr ? { ...emptyOcr(o.ocrLanguage), status: "unavailable", reason: "no Python backend" } : null;
|
|
1473
|
+
const x = json.ocr;
|
|
1474
|
+
if (!x || type !== "image" && !x.textless.length && !x.pages.length && !x.ocrFailed.length) return null;
|
|
1475
|
+
return x;
|
|
1476
|
+
}
|
|
1315
1477
|
async function convertDocument(o, signal, seams) {
|
|
1316
1478
|
const s = { backend: (c) => getBackend(c, void 0, signal), runTier: runTierReal, office: tryConvertOffice, ...seams };
|
|
1317
1479
|
const inputPath = resolve2(o.path);
|
|
1318
1480
|
const st = statSync2(inputPath, { throwIfNoEntry: false });
|
|
1319
1481
|
if (!st || !st.isFile()) throw new Error(`Not a readable file: ${o.path}`);
|
|
1320
1482
|
const type = classifyInput(inputPath);
|
|
1483
|
+
if (o.pages && (type === "html" || type === "image")) throw new Error(`--pages does not apply to ${type === "html" ? "HTML files" : "images"}`);
|
|
1321
1484
|
const isExcel = type === "xlsx" || type === "xls";
|
|
1322
1485
|
if (isExcel && o.pages) throw new Error("--pages does not apply to spreadsheets: worksheets have no stable page numbering");
|
|
1323
1486
|
const backend = await s.backend({ pymupdfVersion: o.pymupdfVersion, warmTimeoutMs: o.warmTimeoutMs });
|
|
@@ -1327,11 +1490,63 @@ async function convertDocument(o, signal, seams) {
|
|
|
1327
1490
|
let office = null;
|
|
1328
1491
|
try {
|
|
1329
1492
|
let pdfPath = inputPath;
|
|
1330
|
-
const base = { path: inputPath, pages: o.pages, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion };
|
|
1493
|
+
const base = { path: inputPath, pages: o.pages, stagingDir: b.stagingDir, sheetsStagingDir: b.sheetsStagingDir, imageDpi: o.imageDpi, imageFormat: o.imageFormat, maxOutputBytes: o.maxOutputBytes, pymupdfVersion: o.pymupdfVersion, ocr: o.ocr, ocrLanguage: o.ocrLanguage, ocrBudgetMs: o.primaryTimeoutMs };
|
|
1331
1494
|
let tier, engine, json, degraded = null, fallbackReason = null;
|
|
1332
1495
|
let explicitBreaks = null;
|
|
1333
1496
|
let notes = [];
|
|
1334
1497
|
let officeRoute = null;
|
|
1498
|
+
if (type === "html") {
|
|
1499
|
+
const prepared = await prepareHtml(inputPath, b.stagingDir);
|
|
1500
|
+
if (signal?.aborted) throw new Error("aborted");
|
|
1501
|
+
if (prepared.missing) notes.push(`${prepared.missing} image(s) not found; replaced with alt text`);
|
|
1502
|
+
if (!lacksDocx(backend)) {
|
|
1503
|
+
const r = await s.runTier("html", { ...base, html: prepared.html }, b, signal, o.primaryTimeoutMs, backend);
|
|
1504
|
+
if (r.ok) {
|
|
1505
|
+
tier = "html";
|
|
1506
|
+
engine = "markdownify";
|
|
1507
|
+
json = r.json;
|
|
1508
|
+
} else if ("userError" in r) throw new Error(r.userError);
|
|
1509
|
+
else if (signal?.aborted || r.reason === "aborted") throw new Error("aborted");
|
|
1510
|
+
else fallbackReason = `html ${r.reason}${detailSuffix(r)}`;
|
|
1511
|
+
}
|
|
1512
|
+
if (json === void 0) {
|
|
1513
|
+
const { htmlToMarkdownRaw: htmlToMarkdownRaw2 } = await Promise.resolve().then(() => (init_fetch_core(), fetch_core_exports));
|
|
1514
|
+
tier = "html";
|
|
1515
|
+
engine = "turndown";
|
|
1516
|
+
degraded = DEGRADED_HTML_TURNDOWN;
|
|
1517
|
+
json = { markdown: `${htmlToMarkdownRaw2(prepared.html)}
|
|
1518
|
+
`, notes: [] };
|
|
1519
|
+
}
|
|
1520
|
+
publishStaged(b);
|
|
1521
|
+
}
|
|
1522
|
+
if (type === "image") {
|
|
1523
|
+
let reason = "no Python backend";
|
|
1524
|
+
if (backend.kind !== "none") {
|
|
1525
|
+
const r = await s.runTier("image", { ...base, stem: b.stem }, b, signal, o.primaryTimeoutMs, backend);
|
|
1526
|
+
if (r.ok) {
|
|
1527
|
+
publishStaged(b);
|
|
1528
|
+
tier = "image";
|
|
1529
|
+
engine = "pymupdf4llm";
|
|
1530
|
+
json = r.json;
|
|
1531
|
+
} else if ("userError" in r) throw new Error(r.userError);
|
|
1532
|
+
else if (signal?.aborted || r.reason === "aborted") throw new Error("aborted");
|
|
1533
|
+
else reason = `OCR child failed: ${r.reason}`;
|
|
1534
|
+
}
|
|
1535
|
+
if (signal?.aborted) throw new Error("aborted");
|
|
1536
|
+
if (json === void 0) {
|
|
1537
|
+
clearStaging(b);
|
|
1538
|
+
const file = `original${extname3(inputPath).toLowerCase()}`;
|
|
1539
|
+
const dir = join3(b.stagingDir, "p1");
|
|
1540
|
+
mkdirSync3(dir, { recursive: true });
|
|
1541
|
+
copyFileSync(inputPath, join3(dir, file));
|
|
1542
|
+
writeFileSync3(join3(dir, ".done"), "");
|
|
1543
|
+
publishStaged(b);
|
|
1544
|
+
tier = "image";
|
|
1545
|
+
engine = "copy";
|
|
1546
|
+
json = { markdown: `
|
|
1547
|
+
`, pageCount: 1, notes: [], ocr: { ...emptyOcr(o.ocrLanguage), status: "unavailable", reason } };
|
|
1548
|
+
}
|
|
1549
|
+
}
|
|
1335
1550
|
if (type === "docx" && !lacksDocx(backend)) {
|
|
1336
1551
|
const d = await s.runTier("docx", base, b, signal, o.primaryTimeoutMs, backend);
|
|
1337
1552
|
if (d.ok) {
|
|
@@ -1441,9 +1656,9 @@ async function convertDocument(o, signal, seams) {
|
|
|
1441
1656
|
fallbackReason = fallbackReason ? `${officeRoute}; ${fallbackReason}` : officeRoute;
|
|
1442
1657
|
}
|
|
1443
1658
|
if (tier === void 0 || engine === void 0 || json === void 0) throw new Error("internal: no tier produced output");
|
|
1444
|
-
if (!isExcel) notes = json.notes ?? [];
|
|
1445
|
-
const body = rewriteLinks(json.markdown ?? "", b.sourceMap);
|
|
1446
|
-
validateImageLinks(body, b.manifest, b.csvManifest);
|
|
1659
|
+
if (!isExcel) notes = [...notes, ...json.notes ?? []];
|
|
1660
|
+
const body = resolveOcrLabels(rewriteLinks(json.markdown ?? "", b.sourceMap), b.sourceMap);
|
|
1661
|
+
validateImageLinks(body, b.manifest, b.csvManifest, type === "html");
|
|
1447
1662
|
const head = [];
|
|
1448
1663
|
if (degraded) head.push(`Degraded: ${degraded}`);
|
|
1449
1664
|
if (fallbackReason) head.push(`Fallback-Reason: ${fallbackReason}`);
|
|
@@ -1455,7 +1670,7 @@ async function convertDocument(o, signal, seams) {
|
|
|
1455
1670
|
` : "") + body;
|
|
1456
1671
|
commitBundle(b, markdown);
|
|
1457
1672
|
const outline = scanOutline(markdown, o.outlineMaxEntries);
|
|
1458
|
-
const details = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total };
|
|
1673
|
+
const details = { path: inputPath, backend: backend.kind, pymupdfVersion: o.pymupdfVersion, inputType: type, file: b.mdPath, outputDir: b.root, savedTo: b.mdPath, imagesDir: b.imagesDir, sheetsDir: b.csvManifest.size ? b.sheetsDir : null, type, engine, tier, pageCount: json.pageCount ?? null, pages: o.pages, explicitBreaks, imageCount: b.manifest.size, bytes: Buffer.byteLength(markdown, "utf8"), lines: markdown.split("\n").length, degraded, fallbackReason, failedPages: (json.failedPages ?? []).map((f) => f.page), emptyPages: json.emptyPages ?? [], notes, outline: outline.entries, outlineTotal: outline.total, ocr: handleOcr(tier, type, o, json) };
|
|
1459
1674
|
return { output: formatHandle(details), details };
|
|
1460
1675
|
} catch (e) {
|
|
1461
1676
|
abortBundle(b);
|
|
@@ -1470,6 +1685,7 @@ async function inspectDocument(o, signal, seams) {
|
|
|
1470
1685
|
const st = statSync2(inputPath, { throwIfNoEntry: false });
|
|
1471
1686
|
if (!st || !st.isFile()) throw new Error(`Not a readable file: ${o.path}`);
|
|
1472
1687
|
const type = classifyInput(inputPath);
|
|
1688
|
+
if (type === "html" || type === "image") throw new Error(`info does not apply to ${type === "html" ? "HTML files" : "images"}; convert directly`);
|
|
1473
1689
|
const isExcel = type === "xlsx" || type === "xls";
|
|
1474
1690
|
const backend = await s.backend({ pymupdfVersion: o.pymupdfVersion, warmTimeoutMs: o.warmTimeoutMs });
|
|
1475
1691
|
if (isExcel && (backend.kind === "none" || !backend.xlsx)) throw new Error(`Excel inspection needs a Python backend with openpyxl, xlrd and pillow. ${EXCEL_REMEDY}`);
|
|
@@ -1511,6 +1727,11 @@ function parseDocToMd(rest) {
|
|
|
1511
1727
|
path = arg;
|
|
1512
1728
|
continue;
|
|
1513
1729
|
}
|
|
1730
|
+
const negated = arg.startsWith("--no-") ? DOC_TO_MD_OPTIONS.find((o) => o.type === "bool" && o.settable && o.flag === `--${arg.slice(5)}`) : void 0;
|
|
1731
|
+
if (negated) {
|
|
1732
|
+
perCall[negated.key] = false;
|
|
1733
|
+
continue;
|
|
1734
|
+
}
|
|
1514
1735
|
const d = DOC_TO_MD_OPTIONS.find((o) => o.flag === arg);
|
|
1515
1736
|
if (!d) return { ok: false, error: `unknown flag: ${arg}` };
|
|
1516
1737
|
if (d.type === "bool") {
|
|
@@ -1544,7 +1765,7 @@ function readCliSettings(cwd, env, warn) {
|
|
|
1544
1765
|
if (!existsSync3(file)) continue;
|
|
1545
1766
|
let raw;
|
|
1546
1767
|
try {
|
|
1547
|
-
raw = JSON.parse(
|
|
1768
|
+
raw = JSON.parse(readFileSync2(file, "utf8")).quiver?.docToMd;
|
|
1548
1769
|
} catch {
|
|
1549
1770
|
warn(`pi-quiver: ${file} is not valid JSON; ignored.`);
|
|
1550
1771
|
continue;
|