@jiaoqsh/dsh-document 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 jiaoqsh
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md ADDED
@@ -0,0 +1,165 @@
1
+ # dsh-document
2
+
3
+ A [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) plugin bundle (`@jiaoqsh/dsh-document`) that gives the model a `read_document` tool: Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF files are converted to line-numbered Markdown the model can page through with `offset`/`limit`.
4
+
5
+ Conversion runs locally: office formats through [`@firecrawl/anydoc`](https://github.com/firecrawl/anydoc) (Rust core, Node bindings) and PDFs through [`@firecrawl/pdf-inspector`](https://github.com/firecrawl/pdf-inspector) (WASM build, run in a worker thread) — no API key, no network, no external binaries; both MIT. PDFs get page selection, `<!-- Page N -->` markers, and page facts: total pages, text-based / scanned / mixed, and which pages have no extractable text. Files are read through the harness `ctx.fs` seam, so whatever filesystem provider and sandbox policy a deployment mounts applies unchanged.
6
+
7
+ ## Install
8
+
9
+ Into an existing profile (`web`, `headless`, or your own), from npm:
10
+
11
+ ```sh
12
+ dsh plugin --profile web add @jiaoqsh/dsh-document
13
+ ```
14
+
15
+ The npm package ships built code, so nothing runs at install time. Releases are published from this repository's `release.yml` through npm trusted publishing and carry provenance attestations.
16
+
17
+ ### From GitHub instead
18
+
19
+ ```sh
20
+ dsh plugin --profile web add github:jiaoqsh/dsh-document#<commit-sha>
21
+ ```
22
+
23
+ A git install fetches sources, so pnpm ≥ 10 refuses to run the package's `prepare` (a `tsdown` transpile of `src/`) until you allow it: the first `add` fails and prints the exact key to copy into the profile's `pnpm-workspace.yaml` (`$DSH_HOME/profiles/<name>/pnpm-workspace.yaml`). The key names the resolved tarball, so a bare package name does not match:
24
+
25
+ ```yaml
26
+ allowBuilds:
27
+ '@jiaoqsh/dsh-document@https://codeload.github.com/jiaoqsh/dsh-document/tar.gz/<commit-sha>': true
28
+ ```
29
+
30
+ Re-run the `add`. Allowing the build means executing this package's code on your machine at install time; pin a commit so a later push cannot change what runs.
31
+
32
+ `dsh plugin` prints "missing peer" warnings for the `@deepseek-ai/*` packages: expected. The dsh installation supplies them at runtime; profiles deliberately do not install peers.
33
+
34
+ Verify, then boot:
35
+
36
+ ```sh
37
+ dsh --profile web --dump-config # shows a "# == @jiaoqsh/dsh-document" layer
38
+ dsh --profile web
39
+ ```
40
+
41
+ Remove with `dsh plugin --profile web remove @jiaoqsh/dsh-document`.
42
+
43
+ ### From a source checkout of the harness
44
+
45
+ ```sh
46
+ pnpm dsh web --patch /absolute/path/to/dsh-document/overlay.yml
47
+ ```
48
+
49
+ where the overlay inserts the built entry by absolute path:
50
+
51
+ ```yaml
52
+ - insert:
53
+ - id: document-tools
54
+ name: '/absolute/path/to/dsh-document/lib/index.js'
55
+ ```
56
+
57
+ ## Configuration
58
+
59
+ The bundle's layer inserts one row, `document-tools`, with schema defaults. Override it by id in your profile's `cordis.patch.yml`; a patch replaces the whole `config`, so restate every key you need:
60
+
61
+ ```yaml
62
+ - id: document-tools
63
+ config:
64
+ maxInputBytes: 104857600 # 100 MiB
65
+ readLimit: 2000
66
+ maxLineLength: 2000
67
+ maxOutputBytes: 51200
68
+ pdfMaxPages: 100
69
+ pdfProfile: fidelity
70
+ ```
71
+
72
+ | Key | Default | Meaning |
73
+ |---|---|---|
74
+ | `maxInputBytes` | 52428800 (50 MiB) | Inclusive byte cap on the source file. Enforced by the filesystem provider before any bytes are buffered; larger files are refused. |
75
+ | `readLimit` | 2000 | Default and maximum number of Markdown lines returned by one call. |
76
+ | `maxLineLength` | 2000 | Maximum characters per returned line; overflow is cut with `… [line truncated]`. |
77
+ | `maxOutputBytes` | 51200 (50 KiB) | Maximum bytes of line text returned by one call; the window stops early and the footer says how to continue. |
78
+ | `pdfMaxPages` | 100 | Maximum distinct pages one `pages` selection may name. |
79
+ | `pdfProfile` | `fidelity` | PDF Markdown profile: `fidelity` keeps source structure, `compact` spends fewer tokens. |
80
+
81
+ Every numeric value must be a positive integer and `pdfProfile` one of the two names; anything else fails the plugin load with a message naming the key.
82
+
83
+ ## The tool
84
+
85
+ `read_document(file_path, offset?, limit?, pages?)`
86
+
87
+ - `file_path` — resolved by the filesystem backend; relative paths resolve against the calling session's workspace.
88
+ - `offset` — 1-based first line of the converted Markdown (default 1).
89
+ - `limit` — lines to return (default and maximum `readLimit`).
90
+ - `pages` — PDF only: 1-based pages to convert, as numbers and ranges like `"1-3,7"` (at most `pdfMaxPages`). Default: every page. Naming a page beyond the last one is an error that states the page count.
91
+
92
+ Supported extensions: `.pdf`, `.doc`, `.docm`, `.docx`, `.ppt`, `.pps`, `.pot`, `.pptx`, `.pptm`, `.ppsx`, `.ppsm`, `.xls`, `.xlsx`, `.xlsm`, `.xlsb`, `.odt`, `.ods`, `.odp`, `.rtf`, `.epub`, `.csv`. The format comes from the extension, never from content sniffing (CSV has no signature).
93
+
94
+ Canonical value (what Code Mode receives):
95
+
96
+ ```ts
97
+ { path: string, format: 'pdf' | 'docx' | ..., offset: number,
98
+ lines: { number: number, text: string }[], totalLines: number, truncatedByBytes: boolean,
99
+ pdf?: { pageCount: number, kind: 'text' | 'scanned' | 'image' | 'mixed',
100
+ pagesNeedingOcr: number[], pages?: number[], title?: string } }
101
+ ```
102
+
103
+ Model-facing text (a PDF, pages 2 and 4 of 5):
104
+
105
+ ```text
106
+ <path>/work/report.pdf</path>
107
+ <format>pdf</format>
108
+ <pdf>5 pages, text-based; showing pages 2, 4</pdf>
109
+ <content>
110
+ 1: <!-- Page 2 -->
111
+ 2:
112
+ 3: Revenue grew 12% year over year.
113
+
114
+ (Showing lines 1-3 of 8. Use offset=4 to continue.)
115
+ </content>
116
+ ```
117
+
118
+ A mixed PDF adds `<warning>Pages 3, 7-8 contain no extractable text (scanned or image content); their content is missing below and would need OCR.</warning>` before `<content>`. Non-PDF formats omit the `<pdf>` line.
119
+
120
+ Failures are tool errors in model terms: unsupported extension (pointing at `read` for plain text), `pages` on a non-PDF, a malformed `pages` value, not found, not a regular file, over `maxInputBytes`, encrypted, damaged or incomplete, engine resource limit, a page beyond the last page, cancellation, or no extractable text (a scanned or image-only PDF says so and that OCR is needed; this tool performs no OCR).
121
+
122
+ ## Model Experience
123
+
124
+ ### System prompt section `tool:read_document`
125
+
126
+ #### What the model sees
127
+
128
+ One fixed sentence, order 100 beside the shipped `tool:read` guidance:
129
+
130
+ ```markdown
131
+ Use the read_document tool — not read or shell commands — to inspect PDF, Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, and CSV files. It returns the document converted to line-numbered Markdown; use offset and limit to continue reading long documents. For a PDF, pass pages (for example "1-3,7") to read only those pages; page markers like <!-- Page 4 --> show where each page starts.
132
+ ```
133
+
134
+ #### Token effect
135
+
136
+ Fixed: the section and the tool schema add a constant number of tokens to every request; results add up to `maxOutputBytes` per call.
137
+
138
+ #### KV Cache effect
139
+
140
+ Prefix-stable: the section text and schema never change between requests, so they do not invalidate a cached prompt prefix.
141
+
142
+ ## Known Limitations and Deferred Work
143
+
144
+ - **No OCR** — pages without a text layer are reported (`pagesNeedingOcr`, the `<warning>` line) but not read; a fully scanned or image-only PDF is an error naming the cause.
145
+ - **PDF conversion runs in a fresh worker thread per call** — cancellation terminates it, and the harness event loop stays free, at the cost of ~50 ms of WASM start-up per call. Office-format conversion (anydoc) runs on the libuv thread pool and cannot be cancelled once started; `maxInputBytes` is its bound.
146
+ - **Layout-heavy PDFs may collapse paragraphs into long lines** (then cut by `maxLineLength`); `pdfProfile: compact` trades structure for tokens.
147
+ - **Engines are fixed** — the `DocumentConverter` interface and `routeConverters` in `src/converter.ts` are the seam for a hosted or OCR-capable engine; none is wired today.
148
+ - **The pdf-inspector native binary is not used** — its npm build ships no `darwin-x64` binary, so the WASM build runs everywhere for one code path.
149
+
150
+ ## Development
151
+
152
+ ```sh
153
+ pnpm install # also builds lib/ via prepare
154
+ pnpm run typecheck
155
+ pnpm test # real Cordis Context + real registry + real local fs; no API key
156
+ pnpm run build
157
+ ```
158
+
159
+ Release: bump `version` in `package.json` on `main`, then push the matching tag (`git tag v0.2.0 && git push origin v0.2.0`). `release.yml` checks the tag against the version, runs the checks, publishes to npm via trusted publishing, and creates the GitHub release with generated notes.
160
+
161
+ Fixtures under `tests/fixtures/` were generated once with macOS `textutil` and `cupsfilter` (including a five-page PDF and an image-only PDF) and are committed so the suite runs anywhere. In source mode the PDF worker is spawned as `.ts` with `--experimental-strip-types`, so it stays free of TypeScript-only runtime syntax.
162
+
163
+ ## License
164
+
165
+ MIT
@@ -0,0 +1,11 @@
1
+ # The layer this bundle contributes when a profile lists it. Config is
2
+ # omitted so the plugin's schema defaults apply; override this row by id
3
+ # in your profile's cordis.patch.yml to change any bound:
4
+ #
5
+ # - id: document-tools
6
+ # name: '@jiaoqsh/dsh-document'
7
+ # config:
8
+ # maxInputBytes: 104857600
9
+ - insert:
10
+ - id: document-tools
11
+ name: '@jiaoqsh/dsh-document'
package/lib/index.d.ts ADDED
@@ -0,0 +1,187 @@
1
+ import "./pdf-worker-CTqxWga-.js";
2
+ import Schema from "@deepseek-ai/schemastery";
3
+ import { MarkdownProfile } from "@firecrawl/pdf-inspector-wasm";
4
+ import { Context } from "@deepseek-ai/cordis";
5
+ //#region src/converter.d.ts
6
+ /**
7
+ * Document-to-Markdown conversion behind one small interface, plus the
8
+ * format router the tool uses. Engines: `@firecrawl/anydoc` for every
9
+ * office format (local Rust napi) and `@firecrawl/pdf-inspector` (WASM in a
10
+ * worker thread) for PDF, where page selection and scanned-page facts matter.
11
+ * @module @jiaoqsh/dsh-document/converter
12
+ */
13
+ /** Formats the engines convert; the value doubles as the model-visible `format` field. */
14
+ declare const DOCUMENT_FORMATS: readonly ["doc", "docx", "odt", "pdf", "ppt", "pptx", "rtf", "epub", "xlsx", "ods", "odp", "csv"];
15
+ /** One convertible document format. */
16
+ type DocumentFormat = typeof DOCUMENT_FORMATS[number];
17
+ /** File extensions (lowercase, no dot) that map to a {@link DocumentFormat}. */
18
+ declare const DOCUMENT_EXTENSIONS: readonly ["pdf", "doc", "docm", "docx", "ppt", "pps", "pot", "pptx", "pptm", "ppsx", "ppsm", "xls", "xlsx", "xlsm", "xlsb", "odt", "ods", "odp", "rtf", "epub", "csv"];
19
+ /**
20
+ * Failure classes a conversion can report. The first six mirror anydoc's
21
+ * `ConvertErrorCode`; `empty` means the engine returned no text; `scanned`
22
+ * means a PDF has no extractable text on any requested page; `aborted` means
23
+ * the caller's signal fired; `unknown` wraps an engine failure without a
24
+ * recognized code.
25
+ */
26
+ type ConversionErrorCode = 'unsupported' | 'malformed' | 'encrypted' | 'resourceLimit' | 'missingPart' | 'io' | 'empty' | 'scanned' | 'pageRange' | 'aborted' | 'unknown';
27
+ /** A conversion failure with a stable code for callers and a model-readable message. */
28
+ declare class DocumentConversionError extends Error {
29
+ readonly code: ConversionErrorCode;
30
+ constructor(message: string, code: ConversionErrorCode);
31
+ }
32
+ /** PDF classification, as the model should hear it. */
33
+ type PdfKind = 'text' | 'scanned' | 'image' | 'mixed';
34
+ /** Facts a PDF engine reports beside the Markdown. Page numbers are 1-based. */
35
+ interface PdfFacts {
36
+ /** Total pages in the document. */
37
+ pageCount: number;
38
+ /** Whether the text layer covers the document. */
39
+ kind: PdfKind;
40
+ /** Pages with no extractable text; their content is absent from the Markdown. */
41
+ pagesNeedingOcr: number[];
42
+ /** The pages that were converted, when the request selected some. */
43
+ pages?: number[];
44
+ /** Document title from metadata, when present. */
45
+ title?: string;
46
+ }
47
+ /** One conversion request. */
48
+ interface ConvertRequest {
49
+ /** The whole file. */
50
+ bytes: Uint8Array;
51
+ /** The format to parse the bytes as; never sniffed (CSV has no signature). */
52
+ format: DocumentFormat;
53
+ /** 1-based pages to convert; PDF only. Absent means every page. */
54
+ pages?: readonly number[];
55
+ /** Cancels a running conversion where the engine allows it. */
56
+ signal: AbortSignal;
57
+ }
58
+ /** One conversion result. */
59
+ interface ConvertResult {
60
+ /** Non-empty GitHub-Flavored Markdown. */
61
+ markdown: string;
62
+ /** Present for PDF engines. */
63
+ pdf?: PdfFacts;
64
+ }
65
+ /** One conversion engine. */
66
+ interface DocumentConverter {
67
+ /** Engine identifier for diagnostics. */
68
+ readonly name: string;
69
+ /** Formats this engine converts; the router picks the first engine listing a format. */
70
+ readonly formats: readonly DocumentFormat[];
71
+ /**
72
+ * Convert complete document bytes to Markdown.
73
+ * @param request - bytes, format, optional page selection, and cancellation.
74
+ * @returns non-empty Markdown plus engine facts.
75
+ * @throws DocumentConversionError for every engine failure and for cancellation.
76
+ */
77
+ convert(request: ConvertRequest): Promise<ConvertResult>;
78
+ }
79
+ /**
80
+ * Map a file path to its document format by extension, case-insensitively.
81
+ * @param path - any path; only the extension is inspected.
82
+ * @returns the format, or `undefined` for an extension no engine handles.
83
+ */
84
+ declare function formatOf(path: string): DocumentFormat | undefined;
85
+ /**
86
+ * Pick the engine for a format: the first converter that lists it.
87
+ * @param converters - engines in priority order.
88
+ * @returns a lookup that throws for a format no engine lists.
89
+ */
90
+ declare function routeConverters(converters: readonly DocumentConverter[]): (format: DocumentFormat) => DocumentConverter;
91
+ /**
92
+ * The anydoc engine: pure local conversion on the libuv thread pool. It lists
93
+ * every format, so it is the last-resort engine behind any specialized one.
94
+ * @returns the converter.
95
+ */
96
+ declare function anydocConverter(): DocumentConverter;
97
+ //#endregion
98
+ //#region src/pdf-inspector.d.ts
99
+ /** Deployment choices for the PDF engine. */
100
+ interface PdfInspectorOptions {
101
+ /** `fidelity` keeps source structure; `compact` spends fewer tokens. */
102
+ profile: MarkdownProfile;
103
+ }
104
+ /**
105
+ * Build the PDF engine.
106
+ * @param options - deployment choices.
107
+ * @returns a converter that handles only `pdf`.
108
+ */
109
+ declare function pdfInspectorConverter(options: PdfInspectorOptions): DocumentConverter;
110
+ //#endregion
111
+ //#region src/window.d.ts
112
+ /** One returned line. */
113
+ interface WindowLine {
114
+ /** 1-based line number in the converted document. */
115
+ number: number;
116
+ /** Line text after per-line truncation. */
117
+ text: string;
118
+ }
119
+ //#endregion
120
+ //#region src/tool.d.ts
121
+ /** Deployment bounds after defaulting (see `Config` in index.ts). */
122
+ interface ReadDocumentCaps {
123
+ /** Inclusive byte cap on the source file; larger files are refused before conversion. */
124
+ maxInputBytes: number;
125
+ /** Default and maximum number of lines returned by one call. */
126
+ readLimit: number;
127
+ /** Maximum characters returned for a single line. */
128
+ maxLineLength: number;
129
+ /** Maximum bytes of line text returned by one call. */
130
+ maxOutputBytes: number;
131
+ /** Maximum distinct PDF pages one `pages` selection may name. */
132
+ pdfMaxPages: number;
133
+ }
134
+ /** Canonical value of one successful call. */
135
+ interface ReadDocumentOutcome {
136
+ path: string;
137
+ format: DocumentFormat;
138
+ offset: number;
139
+ lines: WindowLine[];
140
+ totalLines: number;
141
+ truncatedByBytes: boolean;
142
+ /** PDF facts; absent for other formats. */
143
+ pdf?: PdfFacts;
144
+ }
145
+ //#endregion
146
+ //#region src/index.d.ts
147
+ declare const name = "document-tools";
148
+ declare const inject: string[];
149
+ /** Default inclusive byte cap on a source document. */
150
+ declare const DEFAULT_MAX_INPUT_BYTES: number;
151
+ /** Default and maximum lines per call. */
152
+ declare const DEFAULT_READ_LIMIT = 2000;
153
+ /** Default per-line character cap. */
154
+ declare const DEFAULT_MAX_LINE_LENGTH = 2000;
155
+ /** Default byte cap on returned line text per call. */
156
+ declare const DEFAULT_MAX_OUTPUT_BYTES: number;
157
+ /** Default cap on distinct pages one PDF `pages` selection may name. */
158
+ declare const DEFAULT_PDF_MAX_PAGES = 100;
159
+ /** PDF Markdown profiles the engine offers. */
160
+ type PdfProfile = 'fidelity' | 'compact';
161
+ /**
162
+ * Deployment-owned bounds and choices. Every field is optional on the input
163
+ * side because the schema fills defaults; `apply` receives the resolved record.
164
+ */
165
+ interface Config {
166
+ /** Inclusive byte cap on the source file; larger files are refused before conversion. */
167
+ maxInputBytes?: number;
168
+ /** Default and maximum number of Markdown lines returned by one call. */
169
+ readLimit?: number;
170
+ /** Maximum characters returned for a single line; overflow is cut with a marker. */
171
+ maxLineLength?: number;
172
+ /** Maximum bytes of line text returned by one call. */
173
+ maxOutputBytes?: number;
174
+ /** Maximum distinct pages one PDF `pages` selection may name. */
175
+ pdfMaxPages?: number;
176
+ /** PDF Markdown profile: `fidelity` keeps source structure, `compact` spends fewer tokens. */
177
+ pdfProfile?: PdfProfile;
178
+ }
179
+ declare const Config: Schema<Config>;
180
+ /**
181
+ * Register `read_document` over the anydoc and pdf-inspector engines.
182
+ * @param ctx - context carrying the tool registry, filesystem seam, and system-prompt registry.
183
+ * @param rawConfig - schema-validated bounds after defaulting.
184
+ */
185
+ declare function apply(ctx: Context, rawConfig: Config): void;
186
+ //#endregion
187
+ export { Config, type ConversionErrorCode, type ConvertRequest, type ConvertResult, DEFAULT_MAX_INPUT_BYTES, DEFAULT_MAX_LINE_LENGTH, DEFAULT_MAX_OUTPUT_BYTES, DEFAULT_PDF_MAX_PAGES, DEFAULT_READ_LIMIT, DOCUMENT_EXTENSIONS, DOCUMENT_FORMATS, DocumentConversionError, type DocumentConverter, type DocumentFormat, type PdfFacts, type PdfInspectorOptions, type PdfKind, PdfProfile, type ReadDocumentCaps, type ReadDocumentOutcome, anydocConverter, apply, formatOf, inject, name, pdfInspectorConverter, routeConverters };
package/lib/index.js ADDED
@@ -0,0 +1,613 @@
1
+ import Schema from "@deepseek-ai/schemastery";
2
+ import { formatFromPath, toMarkdownBytes } from "@firecrawl/anydoc";
3
+ import { Worker } from "node:worker_threads";
4
+ import { defineTool } from "@deepseek-ai/dsh-tools";
5
+ //#region src/converter.ts
6
+ /**
7
+ * Document-to-Markdown conversion behind one small interface, plus the
8
+ * format router the tool uses. Engines: `@firecrawl/anydoc` for every
9
+ * office format (local Rust napi) and `@firecrawl/pdf-inspector` (WASM in a
10
+ * worker thread) for PDF, where page selection and scanned-page facts matter.
11
+ * @module @jiaoqsh/dsh-document/converter
12
+ */
13
+ /** Formats the engines convert; the value doubles as the model-visible `format` field. */
14
+ const DOCUMENT_FORMATS = [
15
+ "doc",
16
+ "docx",
17
+ "odt",
18
+ "pdf",
19
+ "ppt",
20
+ "pptx",
21
+ "rtf",
22
+ "epub",
23
+ "xlsx",
24
+ "ods",
25
+ "odp",
26
+ "csv"
27
+ ];
28
+ /** File extensions (lowercase, no dot) that map to a {@link DocumentFormat}. */
29
+ const DOCUMENT_EXTENSIONS = [
30
+ "pdf",
31
+ "doc",
32
+ "docm",
33
+ "docx",
34
+ "ppt",
35
+ "pps",
36
+ "pot",
37
+ "pptx",
38
+ "pptm",
39
+ "ppsx",
40
+ "ppsm",
41
+ "xls",
42
+ "xlsx",
43
+ "xlsm",
44
+ "xlsb",
45
+ "odt",
46
+ "ods",
47
+ "odp",
48
+ "rtf",
49
+ "epub",
50
+ "csv"
51
+ ];
52
+ /** A conversion failure with a stable code for callers and a model-readable message. */
53
+ var DocumentConversionError = class extends Error {
54
+ code;
55
+ constructor(message, code) {
56
+ super(message);
57
+ this.name = "DocumentConversionError";
58
+ this.code = code;
59
+ }
60
+ };
61
+ const FORMAT_SET = new Set(DOCUMENT_FORMATS);
62
+ const ANYDOC_CODES = /* @__PURE__ */ new Set([
63
+ "unsupported",
64
+ "malformed",
65
+ "encrypted",
66
+ "resourceLimit",
67
+ "missingPart",
68
+ "io"
69
+ ]);
70
+ /**
71
+ * Map a file path to its document format by extension, case-insensitively.
72
+ * @param path - any path; only the extension is inspected.
73
+ * @returns the format, or `undefined` for an extension no engine handles.
74
+ */
75
+ function formatOf(path) {
76
+ const format = formatFromPath(path);
77
+ return format !== null && FORMAT_SET.has(format) ? format : void 0;
78
+ }
79
+ /**
80
+ * Pick the engine for a format: the first converter that lists it.
81
+ * @param converters - engines in priority order.
82
+ * @returns a lookup that throws for a format no engine lists.
83
+ */
84
+ function routeConverters(converters) {
85
+ return (format) => {
86
+ const converter = converters.find((candidate) => candidate.formats.includes(format));
87
+ if (converter === void 0) throw new Error(`document-tools: no engine converts ${format}`);
88
+ return converter;
89
+ };
90
+ }
91
+ function anydocError(error) {
92
+ const code = typeof error === "object" && error !== null && "code" in error ? String(error.code) : "";
93
+ return new DocumentConversionError(error instanceof Error ? error.message : String(error), ANYDOC_CODES.has(code) ? code : "unknown");
94
+ }
95
+ /**
96
+ * The anydoc engine: pure local conversion on the libuv thread pool. It lists
97
+ * every format, so it is the last-resort engine behind any specialized one.
98
+ * @returns the converter.
99
+ */
100
+ function anydocConverter() {
101
+ return {
102
+ name: "anydoc",
103
+ formats: DOCUMENT_FORMATS,
104
+ async convert({ bytes, format, signal }) {
105
+ signal.throwIfAborted();
106
+ let markdown;
107
+ try {
108
+ markdown = await toMarkdownBytes(bytes, format);
109
+ } catch (error) {
110
+ throw anydocError(error);
111
+ }
112
+ if (markdown.trim().length === 0) throw new DocumentConversionError("the document contains no extractable text", "empty");
113
+ return { markdown };
114
+ }
115
+ };
116
+ }
117
+ //#endregion
118
+ //#region src/pdf-inspector.ts
119
+ /**
120
+ * The pdf-inspector engine: PDF → Markdown with page selection, page markers,
121
+ * and scanned-page facts. Each conversion runs in a fresh worker thread so a
122
+ * large document never blocks the harness event loop, and cancellation
123
+ * terminates the worker.
124
+ * @module @jiaoqsh/dsh-document/pdf-inspector
125
+ */
126
+ const KIND = {
127
+ TextBased: "text",
128
+ Scanned: "scanned",
129
+ ImageBased: "image",
130
+ Mixed: "mixed"
131
+ };
132
+ const WORKER_IS_SOURCE = import.meta.url.endsWith(".ts");
133
+ const WORKER_URL = new URL(WORKER_IS_SOURCE ? "pdf-worker.ts" : "pdf-worker.js", import.meta.url);
134
+ const WORKER_EXEC_ARGV = WORKER_IS_SOURCE ? [
135
+ ...process.execArgv,
136
+ "--experimental-strip-types",
137
+ "--no-warnings"
138
+ ] : void 0;
139
+ /** Classify an engine failure message; the engine reports plain strings, not codes. */
140
+ function engineError(message) {
141
+ const text = message.replace(/^process PDF: /u, "");
142
+ const lower = text.toLowerCase();
143
+ if (lower.includes("not a pdf") || lower.includes("malformed") || lower.includes("invalid")) return new DocumentConversionError(text, "malformed");
144
+ if (lower.includes("encrypt") || lower.includes("password")) return new DocumentConversionError(text, "encrypted");
145
+ return new DocumentConversionError(text, "unknown");
146
+ }
147
+ /**
148
+ * Run one conversion in a worker and settle with its reply.
149
+ * @param request - bytes, pages, and profile for the worker.
150
+ * @param signal - terminates the worker when it fires.
151
+ * @returns the worker's facts and Markdown.
152
+ */
153
+ function runPdfWorker(request, signal) {
154
+ return new Promise((resolve, reject) => {
155
+ const worker = new Worker(WORKER_URL, {
156
+ workerData: request,
157
+ ...WORKER_EXEC_ARGV === void 0 ? {} : { execArgv: WORKER_EXEC_ARGV }
158
+ });
159
+ let settled = false;
160
+ const settle = (outcome) => {
161
+ if (settled) return;
162
+ settled = true;
163
+ signal.removeEventListener("abort", onAbort);
164
+ worker.terminate();
165
+ outcome();
166
+ };
167
+ const onAbort = () => settle(() => reject(new DocumentConversionError("the conversion was cancelled", "aborted")));
168
+ signal.addEventListener("abort", onAbort, { once: true });
169
+ worker.once("message", (reply) => settle(() => {
170
+ if (reply.ok) resolve(reply.result);
171
+ else reject(engineError(reply.message));
172
+ }));
173
+ worker.once("error", (error) => settle(() => reject(new DocumentConversionError(error.message, "unknown"))));
174
+ worker.once("exit", (code) => settle(() => reject(new DocumentConversionError(`the PDF worker exited with code ${code} before replying`, "unknown"))));
175
+ });
176
+ }
177
+ function formatPages(pages) {
178
+ const sorted = [...pages].sort((a, b) => a - b);
179
+ const parts = [];
180
+ for (let i = 0; i < sorted.length;) {
181
+ let j = i;
182
+ while (j + 1 < sorted.length && sorted[j + 1] === sorted[j] + 1) j += 1;
183
+ parts.push(j === i ? `${sorted[i]}` : `${sorted[i]}-${sorted[j]}`);
184
+ i = j + 1;
185
+ }
186
+ return parts.join(", ");
187
+ }
188
+ /**
189
+ * Build the PDF engine.
190
+ * @param options - deployment choices.
191
+ * @returns a converter that handles only `pdf`.
192
+ */
193
+ function pdfInspectorConverter(options) {
194
+ return {
195
+ name: "pdf-inspector",
196
+ formats: ["pdf"],
197
+ async convert({ bytes, pages, signal }) {
198
+ if (signal.aborted) throw new DocumentConversionError("the conversion was cancelled", "aborted");
199
+ const raw = await runPdfWorker({
200
+ bytes,
201
+ profile: options.profile,
202
+ ...pages === void 0 ? {} : { pages: [...pages] }
203
+ }, signal);
204
+ if (pages !== void 0) {
205
+ const beyond = pages.filter((page) => page > raw.pageCount);
206
+ if (beyond.length > 0) throw new DocumentConversionError(`the document has ${raw.pageCount} page${raw.pageCount === 1 ? "" : "s"}; requested page${beyond.length === 1 ? "" : "s"} ${formatPages(beyond)} ${beyond.length === 1 ? "does" : "do"} not exist`, "pageRange");
207
+ }
208
+ const kind = KIND[raw.pdfType];
209
+ if (raw.markdown.trim().length === 0) {
210
+ const plural = pages === void 0 ? raw.pageCount !== 1 : pages.length !== 1;
211
+ const scope = pages === void 0 ? `all ${raw.pageCount} page${plural ? "s" : ""}` : `the requested page${plural ? "s" : ""} (${formatPages(pages)})`;
212
+ const verb = plural ? "contain" : "contains";
213
+ if (kind === "scanned" || kind === "image") throw new DocumentConversionError(`${scope} ${verb} no extractable text: this is ${kind === "image" ? "an image-only" : "a scanned"} PDF and needs OCR`, "scanned");
214
+ if (raw.pagesNeedingOcr.length > 0) throw new DocumentConversionError(`${scope} ${verb} no extractable text (pages needing OCR: ${formatPages(raw.pagesNeedingOcr)})`, "scanned");
215
+ throw new DocumentConversionError("the document contains no extractable text", "empty");
216
+ }
217
+ return {
218
+ markdown: raw.markdown,
219
+ pdf: {
220
+ pageCount: raw.pageCount,
221
+ kind,
222
+ pagesNeedingOcr: raw.pagesNeedingOcr,
223
+ ...pages === void 0 ? {} : { pages: [...pages] },
224
+ ...raw.title === void 0 ? {} : { title: raw.title }
225
+ }
226
+ };
227
+ }
228
+ };
229
+ }
230
+ //#endregion
231
+ //#region src/pages.ts
232
+ /**
233
+ * The `pages` argument grammar: comma-separated 1-based page numbers and
234
+ * inclusive ranges, e.g. `"1-3,7,10-12"`.
235
+ * @module @jiaoqsh/dsh-document/pages
236
+ */
237
+ const TOKEN = /^(\d+)(?:-(\d+))?$/;
238
+ /**
239
+ * Parse a page selection into a sorted, de-duplicated list.
240
+ * @param spec - the model-supplied text.
241
+ * @param maxPages - the largest number of distinct pages one call may select.
242
+ * @returns ascending 1-based page numbers.
243
+ * @throws Error naming the offending token, or the count when it exceeds `maxPages`.
244
+ */
245
+ function parsePages(spec, maxPages) {
246
+ const pages = /* @__PURE__ */ new Set();
247
+ for (const rawToken of spec.split(",")) {
248
+ const token = rawToken.trim();
249
+ const match = TOKEN.exec(token);
250
+ if (match === null) throw new Error(`pages must be comma-separated page numbers or ranges like "1-3,7", got "${token}"`);
251
+ const first = Number(match[1]);
252
+ const last = match[2] === void 0 ? first : Number(match[2]);
253
+ if (first < 1 || last < first) throw new Error(`pages contains an invalid range "${token}"`);
254
+ if (last - first + 1 > maxPages || pages.size + (last - first + 1) > maxPages) throw new Error(`pages selects more than ${maxPages} pages`);
255
+ for (let page = first; page <= last; page += 1) pages.add(page);
256
+ }
257
+ if (pages.size === 0) throw new Error("pages must select at least one page");
258
+ return [...pages].sort((a, b) => a - b);
259
+ }
260
+ //#endregion
261
+ //#region src/window.ts
262
+ const TRUNCATION_MARKER = "… [line truncated]";
263
+ const encoder = new TextEncoder();
264
+ const decoder = new TextDecoder("utf-8", { fatal: false });
265
+ /**
266
+ * Split converted text into lines: a trailing newline does not open an empty
267
+ * final line, and empty text has zero lines.
268
+ * @param text - the converted document.
269
+ * @returns lines without their terminators; CRLF is normalized.
270
+ */
271
+ function splitLines(text) {
272
+ if (text.length === 0) return [];
273
+ return (text.endsWith("\n") ? text.slice(0, -1) : text).split("\n").map((line) => line.endsWith("\r") ? line.slice(0, -1) : line);
274
+ }
275
+ /** Cut a string to at most `maxBytes` UTF-8 bytes on a character boundary. */
276
+ function cutToBytes(text, maxBytes) {
277
+ const bytes = encoder.encode(text);
278
+ if (bytes.byteLength <= maxBytes) return text;
279
+ let end = maxBytes;
280
+ while (end > 0 && (bytes[end] & 192) === 128) end -= 1;
281
+ return decoder.decode(bytes.subarray(0, end));
282
+ }
283
+ /**
284
+ * Select one window of lines under every bound.
285
+ * @param text - the converted document.
286
+ * @param request - paging and bounds; callers validate that all are positive integers.
287
+ * @returns the window; a first line larger than `maxBytes` is cut rather than dropped so
288
+ * the caller always makes progress.
289
+ */
290
+ function windowLines(text, request) {
291
+ const all = splitLines(text);
292
+ const lines = [];
293
+ let bytes = 0;
294
+ let truncatedByBytes = false;
295
+ const last = Math.min(all.length, request.offset - 1 + request.limit);
296
+ for (let index = request.offset - 1; index < last; index += 1) {
297
+ let line = all[index];
298
+ if (line.length > request.maxLineLength) line = `${line.slice(0, request.maxLineLength)}${TRUNCATION_MARKER}`;
299
+ let size = encoder.encode(line).byteLength;
300
+ if (bytes + size > request.maxBytes) {
301
+ if (lines.length > 0) {
302
+ truncatedByBytes = true;
303
+ break;
304
+ }
305
+ line = cutToBytes(line, request.maxBytes);
306
+ size = encoder.encode(line).byteLength;
307
+ truncatedByBytes = true;
308
+ }
309
+ lines.push({
310
+ number: index + 1,
311
+ text: line
312
+ });
313
+ bytes += size;
314
+ }
315
+ return {
316
+ lines,
317
+ totalLines: all.length,
318
+ truncatedByBytes
319
+ };
320
+ }
321
+ //#endregion
322
+ //#region src/tool.ts
323
+ const EXTENSION_LIST = DOCUMENT_EXTENSIONS.map((ext) => `.${ext}`).join(", ");
324
+ const PDF_KINDS = [
325
+ "text",
326
+ "scanned",
327
+ "image",
328
+ "mixed"
329
+ ];
330
+ function parsePositiveInteger(value, name) {
331
+ if (!Number.isFinite(value) || !Number.isInteger(value) || value < 1) throw new Error(`${name} must be a positive integer`);
332
+ return value;
333
+ }
334
+ /**
335
+ * Validate the constraints the schema DSL does not express.
336
+ * @param args - schema-validated raw arguments.
337
+ * @param caps - the deployment line cap and PDF page cap.
338
+ * @returns validated input with `offset` defaulted to 1 and `limit` to the cap.
339
+ */
340
+ function parseReadArgs(args, caps) {
341
+ if (args.file_path.trim().length === 0) throw new Error("file_path must be a non-empty string");
342
+ const offset = args.offset === void 0 ? 1 : parsePositiveInteger(args.offset, "offset");
343
+ const limit = args.limit === void 0 ? caps.readLimit : parsePositiveInteger(args.limit, "limit");
344
+ if (limit > caps.readLimit) throw new Error(`limit must be less than or equal to ${caps.readLimit}`);
345
+ const pages = args.pages === void 0 ? void 0 : parsePages(args.pages, caps.pdfMaxPages);
346
+ return {
347
+ filePath: args.file_path,
348
+ offset,
349
+ limit,
350
+ ...pages === void 0 ? {} : { pages }
351
+ };
352
+ }
353
+ const KIND_LABEL = {
354
+ text: "text-based",
355
+ scanned: "scanned",
356
+ image: "image-only",
357
+ mixed: "mixed text and scanned pages"
358
+ };
359
+ /**
360
+ * Model-facing text for one outcome: an envelope naming the path and format,
361
+ * PDF facts when present, numbered lines, and a footer that says how to
362
+ * continue.
363
+ * @param outcome - the canonical value.
364
+ * @returns the rendered text.
365
+ */
366
+ function formatReadOutput(outcome) {
367
+ const endLine = outcome.lines.at(-1)?.number ?? Math.max(0, outcome.offset - 1);
368
+ let footer;
369
+ if (outcome.truncatedByBytes) footer = `(Output capped. Showing lines ${outcome.offset}-${endLine}. Use offset=${endLine + 1} to continue.)`;
370
+ else if (endLine < outcome.totalLines) footer = `(Showing lines ${outcome.offset}-${endLine} of ${outcome.totalLines}. Use offset=${endLine + 1} to continue.)`;
371
+ else footer = `(End of document - total ${outcome.totalLines} lines)`;
372
+ const body = outcome.lines.length > 0 ? `${outcome.lines.map((line) => `${line.number}: ${line.text}`).join("\n")}\n\n${footer}` : footer;
373
+ const head = [`<path>${outcome.path}</path>`, `<format>${outcome.format}</format>`];
374
+ if (outcome.pdf !== void 0) {
375
+ const { pageCount, kind, pages, pagesNeedingOcr, title } = outcome.pdf;
376
+ const shown = pages === void 0 ? "all pages" : `showing page${pages.length === 1 ? "" : "s"} ${formatPages(pages)}`;
377
+ head.push(`<pdf>${pageCount} page${pageCount === 1 ? "" : "s"}, ${KIND_LABEL[kind]}; ${shown}${title === void 0 ? "" : `; title: ${title}`}</pdf>`);
378
+ if (pagesNeedingOcr.length > 0) head.push(`<warning>Page${pagesNeedingOcr.length === 1 ? "" : "s"} ${formatPages(pagesNeedingOcr)} contain${pagesNeedingOcr.length === 1 ? "s" : ""} no extractable text (scanned or image content); ${pagesNeedingOcr.length === 1 ? "its" : "their"} content is missing below and would need OCR.</warning>`);
379
+ }
380
+ return `${head.join("\n")}\n<content>\n${body}\n</content>`;
381
+ }
382
+ /** Model-facing message for one conversion failure. */
383
+ function describeFailure(displayPath, error) {
384
+ switch (error.code) {
385
+ case "encrypted": return `cannot read "${displayPath}": the document is encrypted; supply a decrypted copy`;
386
+ case "unsupported": return `cannot read "${displayPath}": ${error.message}`;
387
+ case "empty": return `cannot read "${displayPath}": ${error.message}`;
388
+ case "scanned": return `cannot read "${displayPath}": ${error.message}, which this tool does not perform`;
389
+ case "pageRange": return `cannot read "${displayPath}": ${error.message}`;
390
+ case "malformed":
391
+ case "missingPart": return `cannot read "${displayPath}": the file is damaged or incomplete (${error.message})`;
392
+ case "resourceLimit": return `cannot read "${displayPath}": the document exceeds the converter's internal safety limits (${error.message})`;
393
+ case "aborted": return `cannot read "${displayPath}": ${error.message}`;
394
+ case "io":
395
+ case "unknown": return `cannot read "${displayPath}": ${error.message}`;
396
+ }
397
+ }
398
+ /**
399
+ * Register the `read_document` tool and its system-prompt guidance.
400
+ * @param ctx - the plugin context; registrations are effects scoped to it.
401
+ * @param converterFor - the format router over the mounted engines.
402
+ * @param caps - resolved deployment bounds.
403
+ */
404
+ function applyReadDocumentTool(ctx, converterFor, caps) {
405
+ ctx.systemPrompt.section({
406
+ name: "tool:read_document",
407
+ order: 100,
408
+ text: `Use the read_document tool — not read or shell commands — to inspect PDF, Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, and CSV files. It returns the document converted to line-numbered Markdown; use offset and limit to continue reading long documents. For a PDF, pass pages (for example "1-3,7") to read only those pages; page markers like <!-- Page 4 --> show where each page starts.`
409
+ });
410
+ ctx.tools.register(defineTool({
411
+ name: "read_document",
412
+ description: `Read a document file as Markdown: PDF, Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, and CSV files are converted to line-numbered Markdown text. Supported: ${EXTENSION_LIST}. Returns text only and writes no files. PDFs report page count, whether pages are scanned, and accept a pages selection; scanned or image-only pages are not OCRed. For plain text or source files use read; for images use read_image.`,
413
+ parameters: {
414
+ file_path: {
415
+ type: "string",
416
+ required: true,
417
+ description: "Path to the document, resolved by the filesystem backend; relative paths resolve against the session workspace."
418
+ },
419
+ offset: {
420
+ type: "number",
421
+ description: "1-based first line of the converted Markdown to return. Defaults to 1."
422
+ },
423
+ limit: {
424
+ type: "number",
425
+ description: `Maximum number of lines to return. Defaults to ${caps.readLimit}.`
426
+ },
427
+ pages: {
428
+ type: "string",
429
+ description: `PDF only: 1-based pages to convert, as numbers and ranges like "1-3,7" (at most ${caps.pdfMaxPages} pages). Defaults to every page.`
430
+ }
431
+ },
432
+ output: {
433
+ schema: {
434
+ type: "object",
435
+ additionalProperties: false,
436
+ properties: {
437
+ path: {
438
+ type: "string",
439
+ required: true
440
+ },
441
+ format: {
442
+ type: "string",
443
+ enum: [...DOCUMENT_FORMATS],
444
+ required: true
445
+ },
446
+ offset: {
447
+ type: "integer",
448
+ required: true
449
+ },
450
+ lines: {
451
+ type: "array",
452
+ required: true,
453
+ items: {
454
+ type: "object",
455
+ additionalProperties: false,
456
+ properties: {
457
+ number: {
458
+ type: "integer",
459
+ required: true
460
+ },
461
+ text: {
462
+ type: "string",
463
+ required: true
464
+ }
465
+ }
466
+ }
467
+ },
468
+ totalLines: {
469
+ type: "integer",
470
+ required: true
471
+ },
472
+ truncatedByBytes: {
473
+ type: "boolean",
474
+ required: true
475
+ },
476
+ pdf: {
477
+ type: "object",
478
+ additionalProperties: false,
479
+ properties: {
480
+ pageCount: {
481
+ type: "integer",
482
+ required: true
483
+ },
484
+ kind: {
485
+ type: "string",
486
+ enum: [...PDF_KINDS],
487
+ required: true
488
+ },
489
+ pagesNeedingOcr: {
490
+ type: "array",
491
+ required: true,
492
+ items: { type: "integer" }
493
+ },
494
+ pages: {
495
+ type: "array",
496
+ items: { type: "integer" }
497
+ },
498
+ title: { type: "string" }
499
+ }
500
+ }
501
+ }
502
+ },
503
+ render: (_args, value) => [{
504
+ type: "text",
505
+ text: formatReadOutput(value)
506
+ }]
507
+ },
508
+ isConcurrencySafe: () => true,
509
+ async execute(args, exec) {
510
+ const input = parseReadArgs(args, caps);
511
+ const format = formatOf(input.filePath);
512
+ if (format === void 0) throw new Error(`cannot read "${input.filePath}": unsupported extension. Supported: ${EXTENSION_LIST}. For plain text or source files use the read tool.`);
513
+ if (input.pages !== void 0 && format !== "pdf") throw new Error(`pages applies to PDF files only; "${input.filePath}" is ${format}`);
514
+ const cwd = exec.agent?.session.header.cwd;
515
+ const target = await ctx.fs.resolve(input.filePath, {
516
+ ...cwd === void 0 ? {} : { cwd },
517
+ signal: exec.signal
518
+ });
519
+ const info = await ctx.fs.stat(target, exec.signal);
520
+ if (info === void 0) throw new Error(`cannot read "${target.displayPath}": not found`);
521
+ if (info.type !== "file") throw new Error(`cannot read "${target.displayPath}": not a regular file`);
522
+ const bytes = await ctx.fs.readBytes(target, exec.signal, caps.maxInputBytes);
523
+ let converted;
524
+ try {
525
+ converted = await converterFor(format).convert({
526
+ bytes,
527
+ format,
528
+ signal: exec.signal,
529
+ ...input.pages === void 0 ? {} : { pages: input.pages }
530
+ });
531
+ } catch (error) {
532
+ if (error instanceof DocumentConversionError) throw new Error(describeFailure(target.displayPath, error));
533
+ throw error;
534
+ }
535
+ const window = windowLines(converted.markdown, {
536
+ offset: input.offset,
537
+ limit: input.limit,
538
+ maxLineLength: caps.maxLineLength,
539
+ maxBytes: caps.maxOutputBytes
540
+ });
541
+ return {
542
+ path: target.displayPath,
543
+ format,
544
+ offset: input.offset,
545
+ lines: window.lines,
546
+ totalLines: window.totalLines,
547
+ truncatedByBytes: window.truncatedByBytes,
548
+ ...converted.pdf === void 0 ? {} : { pdf: converted.pdf }
549
+ };
550
+ },
551
+ presentCall(args) {
552
+ const { offset, limit, pages } = args;
553
+ const window = pages !== void 0 && pages.length > 0 ? ` (pages ${pages})` : limit !== void 0 && limit > 0 ? ` (lines ${offset ?? 1} - ${(offset ?? 1) + limit - 1})` : offset !== void 0 ? ` (from line ${offset})` : "";
554
+ return {
555
+ card: "generic",
556
+ title: `Read ${args.file_path} as Markdown${window}`,
557
+ kind: "read",
558
+ locations: [{ path: args.file_path }]
559
+ };
560
+ }
561
+ }));
562
+ }
563
+ //#endregion
564
+ //#region src/index.ts
565
+ const name = "document-tools";
566
+ const inject = [
567
+ "tools",
568
+ "fs",
569
+ "systemPrompt"
570
+ ];
571
+ /** Default inclusive byte cap on a source document. */
572
+ const DEFAULT_MAX_INPUT_BYTES = 52428800;
573
+ /** Default and maximum lines per call. */
574
+ const DEFAULT_READ_LIMIT = 2e3;
575
+ /** Default per-line character cap. */
576
+ const DEFAULT_MAX_LINE_LENGTH = 2e3;
577
+ /** Default byte cap on returned line text per call. */
578
+ const DEFAULT_MAX_OUTPUT_BYTES = 51200;
579
+ /** Default cap on distinct pages one PDF `pages` selection may name. */
580
+ const DEFAULT_PDF_MAX_PAGES = 100;
581
+ const Config = Schema.object({
582
+ maxInputBytes: Schema.number().default(DEFAULT_MAX_INPUT_BYTES),
583
+ readLimit: Schema.number().default(DEFAULT_READ_LIMIT),
584
+ maxLineLength: Schema.number().default(DEFAULT_MAX_LINE_LENGTH),
585
+ maxOutputBytes: Schema.number().default(DEFAULT_MAX_OUTPUT_BYTES),
586
+ pdfMaxPages: Schema.number().default(100),
587
+ pdfProfile: Schema.union(["fidelity", "compact"]).default("fidelity")
588
+ });
589
+ function assertPositiveInteger(name, value) {
590
+ if (!Number.isInteger(value) || value < 1) throw new Error(`document-tools: ${name} must be a positive integer`);
591
+ }
592
+ /**
593
+ * Register `read_document` over the anydoc and pdf-inspector engines.
594
+ * @param ctx - context carrying the tool registry, filesystem seam, and system-prompt registry.
595
+ * @param rawConfig - schema-validated bounds after defaulting.
596
+ */
597
+ function apply(ctx, rawConfig) {
598
+ const config = rawConfig;
599
+ assertPositiveInteger("maxInputBytes", config.maxInputBytes);
600
+ assertPositiveInteger("readLimit", config.readLimit);
601
+ assertPositiveInteger("maxLineLength", config.maxLineLength);
602
+ assertPositiveInteger("maxOutputBytes", config.maxOutputBytes);
603
+ assertPositiveInteger("pdfMaxPages", config.pdfMaxPages);
604
+ applyReadDocumentTool(ctx, routeConverters([pdfInspectorConverter({ profile: config.pdfProfile }), anydocConverter()]), {
605
+ maxInputBytes: config.maxInputBytes,
606
+ readLimit: config.readLimit,
607
+ maxLineLength: config.maxLineLength,
608
+ maxOutputBytes: config.maxOutputBytes,
609
+ pdfMaxPages: config.pdfMaxPages
610
+ });
611
+ }
612
+ //#endregion
613
+ export { Config, DEFAULT_MAX_INPUT_BYTES, DEFAULT_MAX_LINE_LENGTH, DEFAULT_MAX_OUTPUT_BYTES, DEFAULT_PDF_MAX_PAGES, DEFAULT_READ_LIMIT, DOCUMENT_EXTENSIONS, DOCUMENT_FORMATS, DocumentConversionError, anydocConverter, apply, formatOf, inject, name, pdfInspectorConverter, routeConverters };
@@ -0,0 +1,28 @@
1
+ import { MarkdownProfile } from "@firecrawl/pdf-inspector-wasm";
2
+ //#region src/pdf-worker.d.ts
3
+ /** What the parent posts as `workerData`. */
4
+ interface PdfWorkerRequest {
5
+ bytes: Uint8Array;
6
+ /** 1-based pages to convert; absent means all. */
7
+ pages?: number[];
8
+ profile: MarkdownProfile;
9
+ }
10
+ /** The engine facts the parent needs; everything else stays in the worker. */
11
+ interface PdfWorkerResult {
12
+ markdown: string;
13
+ pdfType: 'TextBased' | 'Scanned' | 'ImageBased' | 'Mixed';
14
+ pageCount: number;
15
+ /** 1-based. */
16
+ pagesNeedingOcr: number[];
17
+ title?: string;
18
+ }
19
+ /** Worker → parent message. */
20
+ type PdfWorkerReply = {
21
+ ok: true;
22
+ result: PdfWorkerResult;
23
+ } | {
24
+ ok: false;
25
+ message: string;
26
+ };
27
+ //#endregion
28
+ export { PdfWorkerRequest as n, PdfWorkerResult as r, PdfWorkerReply as t };
@@ -0,0 +1,2 @@
1
+ import { n as PdfWorkerRequest, r as PdfWorkerResult, t as PdfWorkerReply } from "./pdf-worker-CTqxWga-.js";
2
+ export { PdfWorkerReply, PdfWorkerRequest, PdfWorkerResult };
@@ -0,0 +1,47 @@
1
+ import { createRequire } from "node:module";
2
+ import { parentPort, workerData } from "node:worker_threads";
3
+ import { readFileSync } from "node:fs";
4
+ import { initSync, processPdf } from "@firecrawl/pdf-inspector-wasm";
5
+ //#region src/pdf-worker.ts
6
+ /**
7
+ * Worker-thread entry for PDF conversion. Runs `@firecrawl/pdf-inspector`'s
8
+ * WASM build, whose `processPdf` is synchronous, off the main event loop;
9
+ * the parent terminates the worker to cancel. This module is executed by
10
+ * Node directly (never bundled into the plugin entry), so it stays free of
11
+ * TypeScript-only runtime syntax and imports nothing from the plugin.
12
+ * @module @jiaoqsh/dsh-document/pdf-worker
13
+ */
14
+ const require = createRequire(import.meta.url);
15
+ function run(request) {
16
+ initSync({ module: readFileSync(require.resolve("@firecrawl/pdf-inspector-wasm/pdf_inspector_wasm_bg.wasm")) });
17
+ const result = processPdf(new Uint8Array(request.bytes), {
18
+ profile: request.profile,
19
+ includePageMarkers: true,
20
+ includeImages: false,
21
+ ...request.pages === void 0 ? {} : { pages: request.pages }
22
+ });
23
+ return {
24
+ markdown: result.markdown ?? "",
25
+ pdfType: result.pdfType,
26
+ pageCount: result.pageCount,
27
+ pagesNeedingOcr: result.pagesNeedingOcr,
28
+ ...result.title === void 0 ? {} : { title: result.title }
29
+ };
30
+ }
31
+ if (parentPort !== null) {
32
+ let reply;
33
+ try {
34
+ reply = {
35
+ ok: true,
36
+ result: run(workerData)
37
+ };
38
+ } catch (error) {
39
+ reply = {
40
+ ok: false,
41
+ message: error instanceof Error ? error.message : String(error)
42
+ };
43
+ }
44
+ parentPort.postMessage(reply);
45
+ }
46
+ //#endregion
47
+ export {};
package/package.json ADDED
@@ -0,0 +1,76 @@
1
+ {
2
+ "name": "@jiaoqsh/dsh-document",
3
+ "version": "0.1.0",
4
+ "description": "DeepSeek Harness bundle: the read_document tool reads Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF files as Markdown for the model",
5
+ "keywords": [
6
+ "deepseek-harness",
7
+ "dsh",
8
+ "plugin",
9
+ "document",
10
+ "pdf",
11
+ "docx",
12
+ "markdown",
13
+ "anydoc"
14
+ ],
15
+ "license": "MIT",
16
+ "publishConfig": {
17
+ "access": "public"
18
+ },
19
+ "author": "jiaoqsh",
20
+ "repository": {
21
+ "type": "git",
22
+ "url": "git+https://github.com/jiaoqsh/dsh-document.git"
23
+ },
24
+ "type": "module",
25
+ "main": "lib/index.js",
26
+ "types": "lib/index.d.ts",
27
+ "exports": {
28
+ ".": {
29
+ "types": "./lib/index.d.ts",
30
+ "default": "./lib/index.js"
31
+ },
32
+ "./package.json": "./package.json"
33
+ },
34
+ "files": [
35
+ "lib",
36
+ "cordis.patch.yml"
37
+ ],
38
+ "engines": {
39
+ "node": ">=22.19"
40
+ },
41
+ "dsh": {
42
+ "bundle": {
43
+ "patch": "./cordis.patch.yml"
44
+ }
45
+ },
46
+ "scripts": {
47
+ "build": "tsdown",
48
+ "prepare": "tsdown",
49
+ "typecheck": "tsc -p tsconfig.json",
50
+ "test": "vitest run",
51
+ "check": "pnpm run typecheck && pnpm run test && pnpm run build"
52
+ },
53
+ "peerDependencies": {
54
+ "@deepseek-ai/cordis": "^4.0.1",
55
+ "@deepseek-ai/dsh-fs": ">=0.1.0-rc.5",
56
+ "@deepseek-ai/dsh-system-prompt": ">=0.1.0-rc.5",
57
+ "@deepseek-ai/dsh-tools": ">=0.1.0-rc.5"
58
+ },
59
+ "dependencies": {
60
+ "@deepseek-ai/schemastery": "^3.18.1",
61
+ "@firecrawl/anydoc": "0.1.9",
62
+ "@firecrawl/pdf-inspector-wasm": "1.14.2"
63
+ },
64
+ "devDependencies": {
65
+ "@deepseek-ai/cordis": "^4.0.1",
66
+ "@deepseek-ai/dsh-fs": "0.1.0-rc.6",
67
+ "@deepseek-ai/dsh-fs-local": "0.1.0-rc.6",
68
+ "@deepseek-ai/dsh-llm": "0.1.0-rc.6",
69
+ "@deepseek-ai/dsh-system-prompt": "0.1.0-rc.6",
70
+ "@deepseek-ai/dsh-tools": "0.1.0-rc.6",
71
+ "@types/node": "^22.20.0",
72
+ "tsdown": "^0.22.2",
73
+ "typescript": "^6.0.3",
74
+ "vitest": "^4.1.8"
75
+ }
76
+ }