@jiaoqsh/dsh-document 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +165 -0
- package/cordis.patch.yml +11 -0
- package/lib/index.d.ts +187 -0
- package/lib/index.js +613 -0
- package/lib/pdf-worker-CTqxWga-.d.ts +28 -0
- package/lib/pdf-worker.d.ts +2 -0
- package/lib/pdf-worker.js +47 -0
- package/package.json +76 -0
package/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 jiaoqsh
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
package/README.md
ADDED
|
@@ -0,0 +1,165 @@
|
|
|
1
|
+
# dsh-document
|
|
2
|
+
|
|
3
|
+
A [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) plugin bundle (`@jiaoqsh/dsh-document`) that gives the model a `read_document` tool: Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF files are converted to line-numbered Markdown the model can page through with `offset`/`limit`.
|
|
4
|
+
|
|
5
|
+
Conversion runs locally: office formats through [`@firecrawl/anydoc`](https://github.com/firecrawl/anydoc) (Rust core, Node bindings) and PDFs through [`@firecrawl/pdf-inspector`](https://github.com/firecrawl/pdf-inspector) (WASM build, run in a worker thread) — no API key, no network, no external binaries; both MIT. PDFs get page selection, `<!-- Page N -->` markers, and page facts: total pages, text-based / scanned / mixed, and which pages have no extractable text. Files are read through the harness `ctx.fs` seam, so whatever filesystem provider and sandbox policy a deployment mounts applies unchanged.
|
|
6
|
+
|
|
7
|
+
## Install
|
|
8
|
+
|
|
9
|
+
Into an existing profile (`web`, `headless`, or your own), from npm:
|
|
10
|
+
|
|
11
|
+
```sh
|
|
12
|
+
dsh plugin --profile web add @jiaoqsh/dsh-document
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
The npm package ships built code, so nothing runs at install time. Releases are published from this repository's `release.yml` through npm trusted publishing and carry provenance attestations.
|
|
16
|
+
|
|
17
|
+
### From GitHub instead
|
|
18
|
+
|
|
19
|
+
```sh
|
|
20
|
+
dsh plugin --profile web add github:jiaoqsh/dsh-document#<commit-sha>
|
|
21
|
+
```
|
|
22
|
+
|
|
23
|
+
A git install fetches sources, so pnpm ≥ 10 refuses to run the package's `prepare` (a `tsdown` transpile of `src/`) until you allow it: the first `add` fails and prints the exact key to copy into the profile's `pnpm-workspace.yaml` (`$DSH_HOME/profiles/<name>/pnpm-workspace.yaml`). The key names the resolved tarball, so a bare package name does not match:
|
|
24
|
+
|
|
25
|
+
```yaml
|
|
26
|
+
allowBuilds:
|
|
27
|
+
'@jiaoqsh/dsh-document@https://codeload.github.com/jiaoqsh/dsh-document/tar.gz/<commit-sha>': true
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
Re-run the `add`. Allowing the build means executing this package's code on your machine at install time; pin a commit so a later push cannot change what runs.
|
|
31
|
+
|
|
32
|
+
`dsh plugin` prints "missing peer" warnings for the `@deepseek-ai/*` packages: expected. The dsh installation supplies them at runtime; profiles deliberately do not install peers.
|
|
33
|
+
|
|
34
|
+
Verify, then boot:
|
|
35
|
+
|
|
36
|
+
```sh
|
|
37
|
+
dsh --profile web --dump-config # shows a "# == @jiaoqsh/dsh-document" layer
|
|
38
|
+
dsh --profile web
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
Remove with `dsh plugin --profile web remove @jiaoqsh/dsh-document`.
|
|
42
|
+
|
|
43
|
+
### From a source checkout of the harness
|
|
44
|
+
|
|
45
|
+
```sh
|
|
46
|
+
pnpm dsh web --patch /absolute/path/to/dsh-document/overlay.yml
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
where the overlay inserts the built entry by absolute path:
|
|
50
|
+
|
|
51
|
+
```yaml
|
|
52
|
+
- insert:
|
|
53
|
+
- id: document-tools
|
|
54
|
+
name: '/absolute/path/to/dsh-document/lib/index.js'
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
## Configuration
|
|
58
|
+
|
|
59
|
+
The bundle's layer inserts one row, `document-tools`, with schema defaults. Override it by id in your profile's `cordis.patch.yml`; a patch replaces the whole `config`, so restate every key you need:
|
|
60
|
+
|
|
61
|
+
```yaml
|
|
62
|
+
- id: document-tools
|
|
63
|
+
config:
|
|
64
|
+
maxInputBytes: 104857600 # 100 MiB
|
|
65
|
+
readLimit: 2000
|
|
66
|
+
maxLineLength: 2000
|
|
67
|
+
maxOutputBytes: 51200
|
|
68
|
+
pdfMaxPages: 100
|
|
69
|
+
pdfProfile: fidelity
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
| Key | Default | Meaning |
|
|
73
|
+
|---|---|---|
|
|
74
|
+
| `maxInputBytes` | 52428800 (50 MiB) | Inclusive byte cap on the source file. Enforced by the filesystem provider before any bytes are buffered; larger files are refused. |
|
|
75
|
+
| `readLimit` | 2000 | Default and maximum number of Markdown lines returned by one call. |
|
|
76
|
+
| `maxLineLength` | 2000 | Maximum characters per returned line; overflow is cut with `… [line truncated]`. |
|
|
77
|
+
| `maxOutputBytes` | 51200 (50 KiB) | Maximum bytes of line text returned by one call; the window stops early and the footer says how to continue. |
|
|
78
|
+
| `pdfMaxPages` | 100 | Maximum distinct pages one `pages` selection may name. |
|
|
79
|
+
| `pdfProfile` | `fidelity` | PDF Markdown profile: `fidelity` keeps source structure, `compact` spends fewer tokens. |
|
|
80
|
+
|
|
81
|
+
Every numeric value must be a positive integer and `pdfProfile` one of the two names; anything else fails the plugin load with a message naming the key.
|
|
82
|
+
|
|
83
|
+
## The tool
|
|
84
|
+
|
|
85
|
+
`read_document(file_path, offset?, limit?, pages?)`
|
|
86
|
+
|
|
87
|
+
- `file_path` — resolved by the filesystem backend; relative paths resolve against the calling session's workspace.
|
|
88
|
+
- `offset` — 1-based first line of the converted Markdown (default 1).
|
|
89
|
+
- `limit` — lines to return (default and maximum `readLimit`).
|
|
90
|
+
- `pages` — PDF only: 1-based pages to convert, as numbers and ranges like `"1-3,7"` (at most `pdfMaxPages`). Default: every page. Naming a page beyond the last one is an error that states the page count.
|
|
91
|
+
|
|
92
|
+
Supported extensions: `.pdf`, `.doc`, `.docm`, `.docx`, `.ppt`, `.pps`, `.pot`, `.pptx`, `.pptm`, `.ppsx`, `.ppsm`, `.xls`, `.xlsx`, `.xlsm`, `.xlsb`, `.odt`, `.ods`, `.odp`, `.rtf`, `.epub`, `.csv`. The format comes from the extension, never from content sniffing (CSV has no signature).
|
|
93
|
+
|
|
94
|
+
Canonical value (what Code Mode receives):
|
|
95
|
+
|
|
96
|
+
```ts
|
|
97
|
+
{ path: string, format: 'pdf' | 'docx' | ..., offset: number,
|
|
98
|
+
lines: { number: number, text: string }[], totalLines: number, truncatedByBytes: boolean,
|
|
99
|
+
pdf?: { pageCount: number, kind: 'text' | 'scanned' | 'image' | 'mixed',
|
|
100
|
+
pagesNeedingOcr: number[], pages?: number[], title?: string } }
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
Model-facing text (a PDF, pages 2 and 4 of 5):
|
|
104
|
+
|
|
105
|
+
```text
|
|
106
|
+
<path>/work/report.pdf</path>
|
|
107
|
+
<format>pdf</format>
|
|
108
|
+
<pdf>5 pages, text-based; showing pages 2, 4</pdf>
|
|
109
|
+
<content>
|
|
110
|
+
1: <!-- Page 2 -->
|
|
111
|
+
2:
|
|
112
|
+
3: Revenue grew 12% year over year.
|
|
113
|
+
|
|
114
|
+
(Showing lines 1-3 of 8. Use offset=4 to continue.)
|
|
115
|
+
</content>
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
A mixed PDF adds `<warning>Pages 3, 7-8 contain no extractable text (scanned or image content); their content is missing below and would need OCR.</warning>` before `<content>`. Non-PDF formats omit the `<pdf>` line.
|
|
119
|
+
|
|
120
|
+
Failures are tool errors in model terms: unsupported extension (pointing at `read` for plain text), `pages` on a non-PDF, a malformed `pages` value, not found, not a regular file, over `maxInputBytes`, encrypted, damaged or incomplete, engine resource limit, a page beyond the last page, cancellation, or no extractable text (a scanned or image-only PDF says so and that OCR is needed; this tool performs no OCR).
|
|
121
|
+
|
|
122
|
+
## Model Experience
|
|
123
|
+
|
|
124
|
+
### System prompt section `tool:read_document`
|
|
125
|
+
|
|
126
|
+
#### What the model sees
|
|
127
|
+
|
|
128
|
+
One fixed sentence, order 100 beside the shipped `tool:read` guidance:
|
|
129
|
+
|
|
130
|
+
```markdown
|
|
131
|
+
Use the read_document tool — not read or shell commands — to inspect PDF, Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, and CSV files. It returns the document converted to line-numbered Markdown; use offset and limit to continue reading long documents. For a PDF, pass pages (for example "1-3,7") to read only those pages; page markers like <!-- Page 4 --> show where each page starts.
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
#### Token effect
|
|
135
|
+
|
|
136
|
+
Fixed: the section and the tool schema add a constant number of tokens to every request; results add up to `maxOutputBytes` per call.
|
|
137
|
+
|
|
138
|
+
#### KV Cache effect
|
|
139
|
+
|
|
140
|
+
Prefix-stable: the section text and schema never change between requests, so they do not invalidate a cached prompt prefix.
|
|
141
|
+
|
|
142
|
+
## Known Limitations and Deferred Work
|
|
143
|
+
|
|
144
|
+
- **No OCR** — pages without a text layer are reported (`pagesNeedingOcr`, the `<warning>` line) but not read; a fully scanned or image-only PDF is an error naming the cause.
|
|
145
|
+
- **PDF conversion runs in a fresh worker thread per call** — cancellation terminates it, and the harness event loop stays free, at the cost of ~50 ms of WASM start-up per call. Office-format conversion (anydoc) runs on the libuv thread pool and cannot be cancelled once started; `maxInputBytes` is its bound.
|
|
146
|
+
- **Layout-heavy PDFs may collapse paragraphs into long lines** (then cut by `maxLineLength`); `pdfProfile: compact` trades structure for tokens.
|
|
147
|
+
- **Engines are fixed** — the `DocumentConverter` interface and `routeConverters` in `src/converter.ts` are the seam for a hosted or OCR-capable engine; none is wired today.
|
|
148
|
+
- **The pdf-inspector native binary is not used** — its npm build ships no `darwin-x64` binary, so the WASM build runs everywhere for one code path.
|
|
149
|
+
|
|
150
|
+
## Development
|
|
151
|
+
|
|
152
|
+
```sh
|
|
153
|
+
pnpm install # also builds lib/ via prepare
|
|
154
|
+
pnpm run typecheck
|
|
155
|
+
pnpm test # real Cordis Context + real registry + real local fs; no API key
|
|
156
|
+
pnpm run build
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
Release: bump `version` in `package.json` on `main`, then push the matching tag (`git tag v0.2.0 && git push origin v0.2.0`). `release.yml` checks the tag against the version, runs the checks, publishes to npm via trusted publishing, and creates the GitHub release with generated notes.
|
|
160
|
+
|
|
161
|
+
Fixtures under `tests/fixtures/` were generated once with macOS `textutil` and `cupsfilter` (including a five-page PDF and an image-only PDF) and are committed so the suite runs anywhere. In source mode the PDF worker is spawned as `.ts` with `--experimental-strip-types`, so it stays free of TypeScript-only runtime syntax.
|
|
162
|
+
|
|
163
|
+
## License
|
|
164
|
+
|
|
165
|
+
MIT
|
package/cordis.patch.yml
ADDED
|
@@ -0,0 +1,11 @@
|
|
|
1
|
+
# The layer this bundle contributes when a profile lists it. Config is
|
|
2
|
+
# omitted so the plugin's schema defaults apply; override this row by id
|
|
3
|
+
# in your profile's cordis.patch.yml to change any bound:
|
|
4
|
+
#
|
|
5
|
+
# - id: document-tools
|
|
6
|
+
# name: '@jiaoqsh/dsh-document'
|
|
7
|
+
# config:
|
|
8
|
+
# maxInputBytes: 104857600
|
|
9
|
+
- insert:
|
|
10
|
+
- id: document-tools
|
|
11
|
+
name: '@jiaoqsh/dsh-document'
|
package/lib/index.d.ts
ADDED
|
@@ -0,0 +1,187 @@
|
|
|
1
|
+
import "./pdf-worker-CTqxWga-.js";
|
|
2
|
+
import Schema from "@deepseek-ai/schemastery";
|
|
3
|
+
import { MarkdownProfile } from "@firecrawl/pdf-inspector-wasm";
|
|
4
|
+
import { Context } from "@deepseek-ai/cordis";
|
|
5
|
+
//#region src/converter.d.ts
|
|
6
|
+
/**
|
|
7
|
+
* Document-to-Markdown conversion behind one small interface, plus the
|
|
8
|
+
* format router the tool uses. Engines: `@firecrawl/anydoc` for every
|
|
9
|
+
* office format (local Rust napi) and `@firecrawl/pdf-inspector` (WASM in a
|
|
10
|
+
* worker thread) for PDF, where page selection and scanned-page facts matter.
|
|
11
|
+
* @module @jiaoqsh/dsh-document/converter
|
|
12
|
+
*/
|
|
13
|
+
/** Formats the engines convert; the value doubles as the model-visible `format` field. */
|
|
14
|
+
declare const DOCUMENT_FORMATS: readonly ["doc", "docx", "odt", "pdf", "ppt", "pptx", "rtf", "epub", "xlsx", "ods", "odp", "csv"];
|
|
15
|
+
/** One convertible document format. */
|
|
16
|
+
type DocumentFormat = typeof DOCUMENT_FORMATS[number];
|
|
17
|
+
/** File extensions (lowercase, no dot) that map to a {@link DocumentFormat}. */
|
|
18
|
+
declare const DOCUMENT_EXTENSIONS: readonly ["pdf", "doc", "docm", "docx", "ppt", "pps", "pot", "pptx", "pptm", "ppsx", "ppsm", "xls", "xlsx", "xlsm", "xlsb", "odt", "ods", "odp", "rtf", "epub", "csv"];
|
|
19
|
+
/**
|
|
20
|
+
* Failure classes a conversion can report. The first six mirror anydoc's
|
|
21
|
+
* `ConvertErrorCode`; `empty` means the engine returned no text; `scanned`
|
|
22
|
+
* means a PDF has no extractable text on any requested page; `aborted` means
|
|
23
|
+
* the caller's signal fired; `unknown` wraps an engine failure without a
|
|
24
|
+
* recognized code.
|
|
25
|
+
*/
|
|
26
|
+
type ConversionErrorCode = 'unsupported' | 'malformed' | 'encrypted' | 'resourceLimit' | 'missingPart' | 'io' | 'empty' | 'scanned' | 'pageRange' | 'aborted' | 'unknown';
|
|
27
|
+
/** A conversion failure with a stable code for callers and a model-readable message. */
|
|
28
|
+
declare class DocumentConversionError extends Error {
|
|
29
|
+
readonly code: ConversionErrorCode;
|
|
30
|
+
constructor(message: string, code: ConversionErrorCode);
|
|
31
|
+
}
|
|
32
|
+
/** PDF classification, as the model should hear it. */
|
|
33
|
+
type PdfKind = 'text' | 'scanned' | 'image' | 'mixed';
|
|
34
|
+
/** Facts a PDF engine reports beside the Markdown. Page numbers are 1-based. */
|
|
35
|
+
interface PdfFacts {
|
|
36
|
+
/** Total pages in the document. */
|
|
37
|
+
pageCount: number;
|
|
38
|
+
/** Whether the text layer covers the document. */
|
|
39
|
+
kind: PdfKind;
|
|
40
|
+
/** Pages with no extractable text; their content is absent from the Markdown. */
|
|
41
|
+
pagesNeedingOcr: number[];
|
|
42
|
+
/** The pages that were converted, when the request selected some. */
|
|
43
|
+
pages?: number[];
|
|
44
|
+
/** Document title from metadata, when present. */
|
|
45
|
+
title?: string;
|
|
46
|
+
}
|
|
47
|
+
/** One conversion request. */
|
|
48
|
+
interface ConvertRequest {
|
|
49
|
+
/** The whole file. */
|
|
50
|
+
bytes: Uint8Array;
|
|
51
|
+
/** The format to parse the bytes as; never sniffed (CSV has no signature). */
|
|
52
|
+
format: DocumentFormat;
|
|
53
|
+
/** 1-based pages to convert; PDF only. Absent means every page. */
|
|
54
|
+
pages?: readonly number[];
|
|
55
|
+
/** Cancels a running conversion where the engine allows it. */
|
|
56
|
+
signal: AbortSignal;
|
|
57
|
+
}
|
|
58
|
+
/** One conversion result. */
|
|
59
|
+
interface ConvertResult {
|
|
60
|
+
/** Non-empty GitHub-Flavored Markdown. */
|
|
61
|
+
markdown: string;
|
|
62
|
+
/** Present for PDF engines. */
|
|
63
|
+
pdf?: PdfFacts;
|
|
64
|
+
}
|
|
65
|
+
/** One conversion engine. */
|
|
66
|
+
interface DocumentConverter {
|
|
67
|
+
/** Engine identifier for diagnostics. */
|
|
68
|
+
readonly name: string;
|
|
69
|
+
/** Formats this engine converts; the router picks the first engine listing a format. */
|
|
70
|
+
readonly formats: readonly DocumentFormat[];
|
|
71
|
+
/**
|
|
72
|
+
* Convert complete document bytes to Markdown.
|
|
73
|
+
* @param request - bytes, format, optional page selection, and cancellation.
|
|
74
|
+
* @returns non-empty Markdown plus engine facts.
|
|
75
|
+
* @throws DocumentConversionError for every engine failure and for cancellation.
|
|
76
|
+
*/
|
|
77
|
+
convert(request: ConvertRequest): Promise<ConvertResult>;
|
|
78
|
+
}
|
|
79
|
+
/**
|
|
80
|
+
* Map a file path to its document format by extension, case-insensitively.
|
|
81
|
+
* @param path - any path; only the extension is inspected.
|
|
82
|
+
* @returns the format, or `undefined` for an extension no engine handles.
|
|
83
|
+
*/
|
|
84
|
+
declare function formatOf(path: string): DocumentFormat | undefined;
|
|
85
|
+
/**
|
|
86
|
+
* Pick the engine for a format: the first converter that lists it.
|
|
87
|
+
* @param converters - engines in priority order.
|
|
88
|
+
* @returns a lookup that throws for a format no engine lists.
|
|
89
|
+
*/
|
|
90
|
+
declare function routeConverters(converters: readonly DocumentConverter[]): (format: DocumentFormat) => DocumentConverter;
|
|
91
|
+
/**
|
|
92
|
+
* The anydoc engine: pure local conversion on the libuv thread pool. It lists
|
|
93
|
+
* every format, so it is the last-resort engine behind any specialized one.
|
|
94
|
+
* @returns the converter.
|
|
95
|
+
*/
|
|
96
|
+
declare function anydocConverter(): DocumentConverter;
|
|
97
|
+
//#endregion
|
|
98
|
+
//#region src/pdf-inspector.d.ts
|
|
99
|
+
/** Deployment choices for the PDF engine. */
|
|
100
|
+
interface PdfInspectorOptions {
|
|
101
|
+
/** `fidelity` keeps source structure; `compact` spends fewer tokens. */
|
|
102
|
+
profile: MarkdownProfile;
|
|
103
|
+
}
|
|
104
|
+
/**
|
|
105
|
+
* Build the PDF engine.
|
|
106
|
+
* @param options - deployment choices.
|
|
107
|
+
* @returns a converter that handles only `pdf`.
|
|
108
|
+
*/
|
|
109
|
+
declare function pdfInspectorConverter(options: PdfInspectorOptions): DocumentConverter;
|
|
110
|
+
//#endregion
|
|
111
|
+
//#region src/window.d.ts
|
|
112
|
+
/** One returned line. */
|
|
113
|
+
interface WindowLine {
|
|
114
|
+
/** 1-based line number in the converted document. */
|
|
115
|
+
number: number;
|
|
116
|
+
/** Line text after per-line truncation. */
|
|
117
|
+
text: string;
|
|
118
|
+
}
|
|
119
|
+
//#endregion
|
|
120
|
+
//#region src/tool.d.ts
|
|
121
|
+
/** Deployment bounds after defaulting (see `Config` in index.ts). */
|
|
122
|
+
interface ReadDocumentCaps {
|
|
123
|
+
/** Inclusive byte cap on the source file; larger files are refused before conversion. */
|
|
124
|
+
maxInputBytes: number;
|
|
125
|
+
/** Default and maximum number of lines returned by one call. */
|
|
126
|
+
readLimit: number;
|
|
127
|
+
/** Maximum characters returned for a single line. */
|
|
128
|
+
maxLineLength: number;
|
|
129
|
+
/** Maximum bytes of line text returned by one call. */
|
|
130
|
+
maxOutputBytes: number;
|
|
131
|
+
/** Maximum distinct PDF pages one `pages` selection may name. */
|
|
132
|
+
pdfMaxPages: number;
|
|
133
|
+
}
|
|
134
|
+
/** Canonical value of one successful call. */
|
|
135
|
+
interface ReadDocumentOutcome {
|
|
136
|
+
path: string;
|
|
137
|
+
format: DocumentFormat;
|
|
138
|
+
offset: number;
|
|
139
|
+
lines: WindowLine[];
|
|
140
|
+
totalLines: number;
|
|
141
|
+
truncatedByBytes: boolean;
|
|
142
|
+
/** PDF facts; absent for other formats. */
|
|
143
|
+
pdf?: PdfFacts;
|
|
144
|
+
}
|
|
145
|
+
//#endregion
|
|
146
|
+
//#region src/index.d.ts
|
|
147
|
+
declare const name = "document-tools";
|
|
148
|
+
declare const inject: string[];
|
|
149
|
+
/** Default inclusive byte cap on a source document. */
|
|
150
|
+
declare const DEFAULT_MAX_INPUT_BYTES: number;
|
|
151
|
+
/** Default and maximum lines per call. */
|
|
152
|
+
declare const DEFAULT_READ_LIMIT = 2000;
|
|
153
|
+
/** Default per-line character cap. */
|
|
154
|
+
declare const DEFAULT_MAX_LINE_LENGTH = 2000;
|
|
155
|
+
/** Default byte cap on returned line text per call. */
|
|
156
|
+
declare const DEFAULT_MAX_OUTPUT_BYTES: number;
|
|
157
|
+
/** Default cap on distinct pages one PDF `pages` selection may name. */
|
|
158
|
+
declare const DEFAULT_PDF_MAX_PAGES = 100;
|
|
159
|
+
/** PDF Markdown profiles the engine offers. */
|
|
160
|
+
type PdfProfile = 'fidelity' | 'compact';
|
|
161
|
+
/**
|
|
162
|
+
* Deployment-owned bounds and choices. Every field is optional on the input
|
|
163
|
+
* side because the schema fills defaults; `apply` receives the resolved record.
|
|
164
|
+
*/
|
|
165
|
+
interface Config {
|
|
166
|
+
/** Inclusive byte cap on the source file; larger files are refused before conversion. */
|
|
167
|
+
maxInputBytes?: number;
|
|
168
|
+
/** Default and maximum number of Markdown lines returned by one call. */
|
|
169
|
+
readLimit?: number;
|
|
170
|
+
/** Maximum characters returned for a single line; overflow is cut with a marker. */
|
|
171
|
+
maxLineLength?: number;
|
|
172
|
+
/** Maximum bytes of line text returned by one call. */
|
|
173
|
+
maxOutputBytes?: number;
|
|
174
|
+
/** Maximum distinct pages one PDF `pages` selection may name. */
|
|
175
|
+
pdfMaxPages?: number;
|
|
176
|
+
/** PDF Markdown profile: `fidelity` keeps source structure, `compact` spends fewer tokens. */
|
|
177
|
+
pdfProfile?: PdfProfile;
|
|
178
|
+
}
|
|
179
|
+
declare const Config: Schema<Config>;
|
|
180
|
+
/**
|
|
181
|
+
* Register `read_document` over the anydoc and pdf-inspector engines.
|
|
182
|
+
* @param ctx - context carrying the tool registry, filesystem seam, and system-prompt registry.
|
|
183
|
+
* @param rawConfig - schema-validated bounds after defaulting.
|
|
184
|
+
*/
|
|
185
|
+
declare function apply(ctx: Context, rawConfig: Config): void;
|
|
186
|
+
//#endregion
|
|
187
|
+
export { Config, type ConversionErrorCode, type ConvertRequest, type ConvertResult, DEFAULT_MAX_INPUT_BYTES, DEFAULT_MAX_LINE_LENGTH, DEFAULT_MAX_OUTPUT_BYTES, DEFAULT_PDF_MAX_PAGES, DEFAULT_READ_LIMIT, DOCUMENT_EXTENSIONS, DOCUMENT_FORMATS, DocumentConversionError, type DocumentConverter, type DocumentFormat, type PdfFacts, type PdfInspectorOptions, type PdfKind, PdfProfile, type ReadDocumentCaps, type ReadDocumentOutcome, anydocConverter, apply, formatOf, inject, name, pdfInspectorConverter, routeConverters };
|
package/lib/index.js
ADDED
|
@@ -0,0 +1,613 @@
|
|
|
1
|
+
import Schema from "@deepseek-ai/schemastery";
|
|
2
|
+
import { formatFromPath, toMarkdownBytes } from "@firecrawl/anydoc";
|
|
3
|
+
import { Worker } from "node:worker_threads";
|
|
4
|
+
import { defineTool } from "@deepseek-ai/dsh-tools";
|
|
5
|
+
//#region src/converter.ts
|
|
6
|
+
/**
|
|
7
|
+
* Document-to-Markdown conversion behind one small interface, plus the
|
|
8
|
+
* format router the tool uses. Engines: `@firecrawl/anydoc` for every
|
|
9
|
+
* office format (local Rust napi) and `@firecrawl/pdf-inspector` (WASM in a
|
|
10
|
+
* worker thread) for PDF, where page selection and scanned-page facts matter.
|
|
11
|
+
* @module @jiaoqsh/dsh-document/converter
|
|
12
|
+
*/
|
|
13
|
+
/** Formats the engines convert; the value doubles as the model-visible `format` field. */
|
|
14
|
+
const DOCUMENT_FORMATS = [
|
|
15
|
+
"doc",
|
|
16
|
+
"docx",
|
|
17
|
+
"odt",
|
|
18
|
+
"pdf",
|
|
19
|
+
"ppt",
|
|
20
|
+
"pptx",
|
|
21
|
+
"rtf",
|
|
22
|
+
"epub",
|
|
23
|
+
"xlsx",
|
|
24
|
+
"ods",
|
|
25
|
+
"odp",
|
|
26
|
+
"csv"
|
|
27
|
+
];
|
|
28
|
+
/** File extensions (lowercase, no dot) that map to a {@link DocumentFormat}. */
|
|
29
|
+
const DOCUMENT_EXTENSIONS = [
|
|
30
|
+
"pdf",
|
|
31
|
+
"doc",
|
|
32
|
+
"docm",
|
|
33
|
+
"docx",
|
|
34
|
+
"ppt",
|
|
35
|
+
"pps",
|
|
36
|
+
"pot",
|
|
37
|
+
"pptx",
|
|
38
|
+
"pptm",
|
|
39
|
+
"ppsx",
|
|
40
|
+
"ppsm",
|
|
41
|
+
"xls",
|
|
42
|
+
"xlsx",
|
|
43
|
+
"xlsm",
|
|
44
|
+
"xlsb",
|
|
45
|
+
"odt",
|
|
46
|
+
"ods",
|
|
47
|
+
"odp",
|
|
48
|
+
"rtf",
|
|
49
|
+
"epub",
|
|
50
|
+
"csv"
|
|
51
|
+
];
|
|
52
|
+
/** A conversion failure with a stable code for callers and a model-readable message. */
|
|
53
|
+
var DocumentConversionError = class extends Error {
|
|
54
|
+
code;
|
|
55
|
+
constructor(message, code) {
|
|
56
|
+
super(message);
|
|
57
|
+
this.name = "DocumentConversionError";
|
|
58
|
+
this.code = code;
|
|
59
|
+
}
|
|
60
|
+
};
|
|
61
|
+
const FORMAT_SET = new Set(DOCUMENT_FORMATS);
|
|
62
|
+
const ANYDOC_CODES = /* @__PURE__ */ new Set([
|
|
63
|
+
"unsupported",
|
|
64
|
+
"malformed",
|
|
65
|
+
"encrypted",
|
|
66
|
+
"resourceLimit",
|
|
67
|
+
"missingPart",
|
|
68
|
+
"io"
|
|
69
|
+
]);
|
|
70
|
+
/**
|
|
71
|
+
* Map a file path to its document format by extension, case-insensitively.
|
|
72
|
+
* @param path - any path; only the extension is inspected.
|
|
73
|
+
* @returns the format, or `undefined` for an extension no engine handles.
|
|
74
|
+
*/
|
|
75
|
+
function formatOf(path) {
|
|
76
|
+
const format = formatFromPath(path);
|
|
77
|
+
return format !== null && FORMAT_SET.has(format) ? format : void 0;
|
|
78
|
+
}
|
|
79
|
+
/**
|
|
80
|
+
* Pick the engine for a format: the first converter that lists it.
|
|
81
|
+
* @param converters - engines in priority order.
|
|
82
|
+
* @returns a lookup that throws for a format no engine lists.
|
|
83
|
+
*/
|
|
84
|
+
function routeConverters(converters) {
|
|
85
|
+
return (format) => {
|
|
86
|
+
const converter = converters.find((candidate) => candidate.formats.includes(format));
|
|
87
|
+
if (converter === void 0) throw new Error(`document-tools: no engine converts ${format}`);
|
|
88
|
+
return converter;
|
|
89
|
+
};
|
|
90
|
+
}
|
|
91
|
+
function anydocError(error) {
|
|
92
|
+
const code = typeof error === "object" && error !== null && "code" in error ? String(error.code) : "";
|
|
93
|
+
return new DocumentConversionError(error instanceof Error ? error.message : String(error), ANYDOC_CODES.has(code) ? code : "unknown");
|
|
94
|
+
}
|
|
95
|
+
/**
|
|
96
|
+
* The anydoc engine: pure local conversion on the libuv thread pool. It lists
|
|
97
|
+
* every format, so it is the last-resort engine behind any specialized one.
|
|
98
|
+
* @returns the converter.
|
|
99
|
+
*/
|
|
100
|
+
function anydocConverter() {
|
|
101
|
+
return {
|
|
102
|
+
name: "anydoc",
|
|
103
|
+
formats: DOCUMENT_FORMATS,
|
|
104
|
+
async convert({ bytes, format, signal }) {
|
|
105
|
+
signal.throwIfAborted();
|
|
106
|
+
let markdown;
|
|
107
|
+
try {
|
|
108
|
+
markdown = await toMarkdownBytes(bytes, format);
|
|
109
|
+
} catch (error) {
|
|
110
|
+
throw anydocError(error);
|
|
111
|
+
}
|
|
112
|
+
if (markdown.trim().length === 0) throw new DocumentConversionError("the document contains no extractable text", "empty");
|
|
113
|
+
return { markdown };
|
|
114
|
+
}
|
|
115
|
+
};
|
|
116
|
+
}
|
|
117
|
+
//#endregion
|
|
118
|
+
//#region src/pdf-inspector.ts
|
|
119
|
+
/**
|
|
120
|
+
* The pdf-inspector engine: PDF → Markdown with page selection, page markers,
|
|
121
|
+
* and scanned-page facts. Each conversion runs in a fresh worker thread so a
|
|
122
|
+
* large document never blocks the harness event loop, and cancellation
|
|
123
|
+
* terminates the worker.
|
|
124
|
+
* @module @jiaoqsh/dsh-document/pdf-inspector
|
|
125
|
+
*/
|
|
126
|
+
const KIND = {
|
|
127
|
+
TextBased: "text",
|
|
128
|
+
Scanned: "scanned",
|
|
129
|
+
ImageBased: "image",
|
|
130
|
+
Mixed: "mixed"
|
|
131
|
+
};
|
|
132
|
+
const WORKER_IS_SOURCE = import.meta.url.endsWith(".ts");
|
|
133
|
+
const WORKER_URL = new URL(WORKER_IS_SOURCE ? "pdf-worker.ts" : "pdf-worker.js", import.meta.url);
|
|
134
|
+
const WORKER_EXEC_ARGV = WORKER_IS_SOURCE ? [
|
|
135
|
+
...process.execArgv,
|
|
136
|
+
"--experimental-strip-types",
|
|
137
|
+
"--no-warnings"
|
|
138
|
+
] : void 0;
|
|
139
|
+
/** Classify an engine failure message; the engine reports plain strings, not codes. */
|
|
140
|
+
function engineError(message) {
|
|
141
|
+
const text = message.replace(/^process PDF: /u, "");
|
|
142
|
+
const lower = text.toLowerCase();
|
|
143
|
+
if (lower.includes("not a pdf") || lower.includes("malformed") || lower.includes("invalid")) return new DocumentConversionError(text, "malformed");
|
|
144
|
+
if (lower.includes("encrypt") || lower.includes("password")) return new DocumentConversionError(text, "encrypted");
|
|
145
|
+
return new DocumentConversionError(text, "unknown");
|
|
146
|
+
}
|
|
147
|
+
/**
|
|
148
|
+
* Run one conversion in a worker and settle with its reply.
|
|
149
|
+
* @param request - bytes, pages, and profile for the worker.
|
|
150
|
+
* @param signal - terminates the worker when it fires.
|
|
151
|
+
* @returns the worker's facts and Markdown.
|
|
152
|
+
*/
|
|
153
|
+
function runPdfWorker(request, signal) {
|
|
154
|
+
return new Promise((resolve, reject) => {
|
|
155
|
+
const worker = new Worker(WORKER_URL, {
|
|
156
|
+
workerData: request,
|
|
157
|
+
...WORKER_EXEC_ARGV === void 0 ? {} : { execArgv: WORKER_EXEC_ARGV }
|
|
158
|
+
});
|
|
159
|
+
let settled = false;
|
|
160
|
+
const settle = (outcome) => {
|
|
161
|
+
if (settled) return;
|
|
162
|
+
settled = true;
|
|
163
|
+
signal.removeEventListener("abort", onAbort);
|
|
164
|
+
worker.terminate();
|
|
165
|
+
outcome();
|
|
166
|
+
};
|
|
167
|
+
const onAbort = () => settle(() => reject(new DocumentConversionError("the conversion was cancelled", "aborted")));
|
|
168
|
+
signal.addEventListener("abort", onAbort, { once: true });
|
|
169
|
+
worker.once("message", (reply) => settle(() => {
|
|
170
|
+
if (reply.ok) resolve(reply.result);
|
|
171
|
+
else reject(engineError(reply.message));
|
|
172
|
+
}));
|
|
173
|
+
worker.once("error", (error) => settle(() => reject(new DocumentConversionError(error.message, "unknown"))));
|
|
174
|
+
worker.once("exit", (code) => settle(() => reject(new DocumentConversionError(`the PDF worker exited with code ${code} before replying`, "unknown"))));
|
|
175
|
+
});
|
|
176
|
+
}
|
|
177
|
+
function formatPages(pages) {
|
|
178
|
+
const sorted = [...pages].sort((a, b) => a - b);
|
|
179
|
+
const parts = [];
|
|
180
|
+
for (let i = 0; i < sorted.length;) {
|
|
181
|
+
let j = i;
|
|
182
|
+
while (j + 1 < sorted.length && sorted[j + 1] === sorted[j] + 1) j += 1;
|
|
183
|
+
parts.push(j === i ? `${sorted[i]}` : `${sorted[i]}-${sorted[j]}`);
|
|
184
|
+
i = j + 1;
|
|
185
|
+
}
|
|
186
|
+
return parts.join(", ");
|
|
187
|
+
}
|
|
188
|
+
/**
|
|
189
|
+
* Build the PDF engine.
|
|
190
|
+
* @param options - deployment choices.
|
|
191
|
+
* @returns a converter that handles only `pdf`.
|
|
192
|
+
*/
|
|
193
|
+
function pdfInspectorConverter(options) {
|
|
194
|
+
return {
|
|
195
|
+
name: "pdf-inspector",
|
|
196
|
+
formats: ["pdf"],
|
|
197
|
+
async convert({ bytes, pages, signal }) {
|
|
198
|
+
if (signal.aborted) throw new DocumentConversionError("the conversion was cancelled", "aborted");
|
|
199
|
+
const raw = await runPdfWorker({
|
|
200
|
+
bytes,
|
|
201
|
+
profile: options.profile,
|
|
202
|
+
...pages === void 0 ? {} : { pages: [...pages] }
|
|
203
|
+
}, signal);
|
|
204
|
+
if (pages !== void 0) {
|
|
205
|
+
const beyond = pages.filter((page) => page > raw.pageCount);
|
|
206
|
+
if (beyond.length > 0) throw new DocumentConversionError(`the document has ${raw.pageCount} page${raw.pageCount === 1 ? "" : "s"}; requested page${beyond.length === 1 ? "" : "s"} ${formatPages(beyond)} ${beyond.length === 1 ? "does" : "do"} not exist`, "pageRange");
|
|
207
|
+
}
|
|
208
|
+
const kind = KIND[raw.pdfType];
|
|
209
|
+
if (raw.markdown.trim().length === 0) {
|
|
210
|
+
const plural = pages === void 0 ? raw.pageCount !== 1 : pages.length !== 1;
|
|
211
|
+
const scope = pages === void 0 ? `all ${raw.pageCount} page${plural ? "s" : ""}` : `the requested page${plural ? "s" : ""} (${formatPages(pages)})`;
|
|
212
|
+
const verb = plural ? "contain" : "contains";
|
|
213
|
+
if (kind === "scanned" || kind === "image") throw new DocumentConversionError(`${scope} ${verb} no extractable text: this is ${kind === "image" ? "an image-only" : "a scanned"} PDF and needs OCR`, "scanned");
|
|
214
|
+
if (raw.pagesNeedingOcr.length > 0) throw new DocumentConversionError(`${scope} ${verb} no extractable text (pages needing OCR: ${formatPages(raw.pagesNeedingOcr)})`, "scanned");
|
|
215
|
+
throw new DocumentConversionError("the document contains no extractable text", "empty");
|
|
216
|
+
}
|
|
217
|
+
return {
|
|
218
|
+
markdown: raw.markdown,
|
|
219
|
+
pdf: {
|
|
220
|
+
pageCount: raw.pageCount,
|
|
221
|
+
kind,
|
|
222
|
+
pagesNeedingOcr: raw.pagesNeedingOcr,
|
|
223
|
+
...pages === void 0 ? {} : { pages: [...pages] },
|
|
224
|
+
...raw.title === void 0 ? {} : { title: raw.title }
|
|
225
|
+
}
|
|
226
|
+
};
|
|
227
|
+
}
|
|
228
|
+
};
|
|
229
|
+
}
|
|
230
|
+
//#endregion
|
|
231
|
+
//#region src/pages.ts
|
|
232
|
+
/**
|
|
233
|
+
* The `pages` argument grammar: comma-separated 1-based page numbers and
|
|
234
|
+
* inclusive ranges, e.g. `"1-3,7,10-12"`.
|
|
235
|
+
* @module @jiaoqsh/dsh-document/pages
|
|
236
|
+
*/
|
|
237
|
+
const TOKEN = /^(\d+)(?:-(\d+))?$/;
|
|
238
|
+
/**
|
|
239
|
+
* Parse a page selection into a sorted, de-duplicated list.
|
|
240
|
+
* @param spec - the model-supplied text.
|
|
241
|
+
* @param maxPages - the largest number of distinct pages one call may select.
|
|
242
|
+
* @returns ascending 1-based page numbers.
|
|
243
|
+
* @throws Error naming the offending token, or the count when it exceeds `maxPages`.
|
|
244
|
+
*/
|
|
245
|
+
function parsePages(spec, maxPages) {
|
|
246
|
+
const pages = /* @__PURE__ */ new Set();
|
|
247
|
+
for (const rawToken of spec.split(",")) {
|
|
248
|
+
const token = rawToken.trim();
|
|
249
|
+
const match = TOKEN.exec(token);
|
|
250
|
+
if (match === null) throw new Error(`pages must be comma-separated page numbers or ranges like "1-3,7", got "${token}"`);
|
|
251
|
+
const first = Number(match[1]);
|
|
252
|
+
const last = match[2] === void 0 ? first : Number(match[2]);
|
|
253
|
+
if (first < 1 || last < first) throw new Error(`pages contains an invalid range "${token}"`);
|
|
254
|
+
if (last - first + 1 > maxPages || pages.size + (last - first + 1) > maxPages) throw new Error(`pages selects more than ${maxPages} pages`);
|
|
255
|
+
for (let page = first; page <= last; page += 1) pages.add(page);
|
|
256
|
+
}
|
|
257
|
+
if (pages.size === 0) throw new Error("pages must select at least one page");
|
|
258
|
+
return [...pages].sort((a, b) => a - b);
|
|
259
|
+
}
|
|
260
|
+
//#endregion
|
|
261
|
+
//#region src/window.ts
|
|
262
|
+
const TRUNCATION_MARKER = "… [line truncated]";
|
|
263
|
+
const encoder = new TextEncoder();
|
|
264
|
+
const decoder = new TextDecoder("utf-8", { fatal: false });
|
|
265
|
+
/**
|
|
266
|
+
* Split converted text into lines: a trailing newline does not open an empty
|
|
267
|
+
* final line, and empty text has zero lines.
|
|
268
|
+
* @param text - the converted document.
|
|
269
|
+
* @returns lines without their terminators; CRLF is normalized.
|
|
270
|
+
*/
|
|
271
|
+
function splitLines(text) {
|
|
272
|
+
if (text.length === 0) return [];
|
|
273
|
+
return (text.endsWith("\n") ? text.slice(0, -1) : text).split("\n").map((line) => line.endsWith("\r") ? line.slice(0, -1) : line);
|
|
274
|
+
}
|
|
275
|
+
/** Cut a string to at most `maxBytes` UTF-8 bytes on a character boundary. */
|
|
276
|
+
function cutToBytes(text, maxBytes) {
|
|
277
|
+
const bytes = encoder.encode(text);
|
|
278
|
+
if (bytes.byteLength <= maxBytes) return text;
|
|
279
|
+
let end = maxBytes;
|
|
280
|
+
while (end > 0 && (bytes[end] & 192) === 128) end -= 1;
|
|
281
|
+
return decoder.decode(bytes.subarray(0, end));
|
|
282
|
+
}
|
|
283
|
+
/**
|
|
284
|
+
* Select one window of lines under every bound.
|
|
285
|
+
* @param text - the converted document.
|
|
286
|
+
* @param request - paging and bounds; callers validate that all are positive integers.
|
|
287
|
+
* @returns the window; a first line larger than `maxBytes` is cut rather than dropped so
|
|
288
|
+
* the caller always makes progress.
|
|
289
|
+
*/
|
|
290
|
+
function windowLines(text, request) {
|
|
291
|
+
const all = splitLines(text);
|
|
292
|
+
const lines = [];
|
|
293
|
+
let bytes = 0;
|
|
294
|
+
let truncatedByBytes = false;
|
|
295
|
+
const last = Math.min(all.length, request.offset - 1 + request.limit);
|
|
296
|
+
for (let index = request.offset - 1; index < last; index += 1) {
|
|
297
|
+
let line = all[index];
|
|
298
|
+
if (line.length > request.maxLineLength) line = `${line.slice(0, request.maxLineLength)}${TRUNCATION_MARKER}`;
|
|
299
|
+
let size = encoder.encode(line).byteLength;
|
|
300
|
+
if (bytes + size > request.maxBytes) {
|
|
301
|
+
if (lines.length > 0) {
|
|
302
|
+
truncatedByBytes = true;
|
|
303
|
+
break;
|
|
304
|
+
}
|
|
305
|
+
line = cutToBytes(line, request.maxBytes);
|
|
306
|
+
size = encoder.encode(line).byteLength;
|
|
307
|
+
truncatedByBytes = true;
|
|
308
|
+
}
|
|
309
|
+
lines.push({
|
|
310
|
+
number: index + 1,
|
|
311
|
+
text: line
|
|
312
|
+
});
|
|
313
|
+
bytes += size;
|
|
314
|
+
}
|
|
315
|
+
return {
|
|
316
|
+
lines,
|
|
317
|
+
totalLines: all.length,
|
|
318
|
+
truncatedByBytes
|
|
319
|
+
};
|
|
320
|
+
}
|
|
321
|
+
//#endregion
|
|
322
|
+
//#region src/tool.ts
|
|
323
|
+
const EXTENSION_LIST = DOCUMENT_EXTENSIONS.map((ext) => `.${ext}`).join(", ");
|
|
324
|
+
const PDF_KINDS = [
|
|
325
|
+
"text",
|
|
326
|
+
"scanned",
|
|
327
|
+
"image",
|
|
328
|
+
"mixed"
|
|
329
|
+
];
|
|
330
|
+
function parsePositiveInteger(value, name) {
|
|
331
|
+
if (!Number.isFinite(value) || !Number.isInteger(value) || value < 1) throw new Error(`${name} must be a positive integer`);
|
|
332
|
+
return value;
|
|
333
|
+
}
|
|
334
|
+
/**
|
|
335
|
+
* Validate the constraints the schema DSL does not express.
|
|
336
|
+
* @param args - schema-validated raw arguments.
|
|
337
|
+
* @param caps - the deployment line cap and PDF page cap.
|
|
338
|
+
* @returns validated input with `offset` defaulted to 1 and `limit` to the cap.
|
|
339
|
+
*/
|
|
340
|
+
function parseReadArgs(args, caps) {
|
|
341
|
+
if (args.file_path.trim().length === 0) throw new Error("file_path must be a non-empty string");
|
|
342
|
+
const offset = args.offset === void 0 ? 1 : parsePositiveInteger(args.offset, "offset");
|
|
343
|
+
const limit = args.limit === void 0 ? caps.readLimit : parsePositiveInteger(args.limit, "limit");
|
|
344
|
+
if (limit > caps.readLimit) throw new Error(`limit must be less than or equal to ${caps.readLimit}`);
|
|
345
|
+
const pages = args.pages === void 0 ? void 0 : parsePages(args.pages, caps.pdfMaxPages);
|
|
346
|
+
return {
|
|
347
|
+
filePath: args.file_path,
|
|
348
|
+
offset,
|
|
349
|
+
limit,
|
|
350
|
+
...pages === void 0 ? {} : { pages }
|
|
351
|
+
};
|
|
352
|
+
}
|
|
353
|
+
const KIND_LABEL = {
|
|
354
|
+
text: "text-based",
|
|
355
|
+
scanned: "scanned",
|
|
356
|
+
image: "image-only",
|
|
357
|
+
mixed: "mixed text and scanned pages"
|
|
358
|
+
};
|
|
359
|
+
/**
|
|
360
|
+
* Model-facing text for one outcome: an envelope naming the path and format,
|
|
361
|
+
* PDF facts when present, numbered lines, and a footer that says how to
|
|
362
|
+
* continue.
|
|
363
|
+
* @param outcome - the canonical value.
|
|
364
|
+
* @returns the rendered text.
|
|
365
|
+
*/
|
|
366
|
+
function formatReadOutput(outcome) {
|
|
367
|
+
const endLine = outcome.lines.at(-1)?.number ?? Math.max(0, outcome.offset - 1);
|
|
368
|
+
let footer;
|
|
369
|
+
if (outcome.truncatedByBytes) footer = `(Output capped. Showing lines ${outcome.offset}-${endLine}. Use offset=${endLine + 1} to continue.)`;
|
|
370
|
+
else if (endLine < outcome.totalLines) footer = `(Showing lines ${outcome.offset}-${endLine} of ${outcome.totalLines}. Use offset=${endLine + 1} to continue.)`;
|
|
371
|
+
else footer = `(End of document - total ${outcome.totalLines} lines)`;
|
|
372
|
+
const body = outcome.lines.length > 0 ? `${outcome.lines.map((line) => `${line.number}: ${line.text}`).join("\n")}\n\n${footer}` : footer;
|
|
373
|
+
const head = [`<path>${outcome.path}</path>`, `<format>${outcome.format}</format>`];
|
|
374
|
+
if (outcome.pdf !== void 0) {
|
|
375
|
+
const { pageCount, kind, pages, pagesNeedingOcr, title } = outcome.pdf;
|
|
376
|
+
const shown = pages === void 0 ? "all pages" : `showing page${pages.length === 1 ? "" : "s"} ${formatPages(pages)}`;
|
|
377
|
+
head.push(`<pdf>${pageCount} page${pageCount === 1 ? "" : "s"}, ${KIND_LABEL[kind]}; ${shown}${title === void 0 ? "" : `; title: ${title}`}</pdf>`);
|
|
378
|
+
if (pagesNeedingOcr.length > 0) head.push(`<warning>Page${pagesNeedingOcr.length === 1 ? "" : "s"} ${formatPages(pagesNeedingOcr)} contain${pagesNeedingOcr.length === 1 ? "s" : ""} no extractable text (scanned or image content); ${pagesNeedingOcr.length === 1 ? "its" : "their"} content is missing below and would need OCR.</warning>`);
|
|
379
|
+
}
|
|
380
|
+
return `${head.join("\n")}\n<content>\n${body}\n</content>`;
|
|
381
|
+
}
|
|
382
|
+
/** Model-facing message for one conversion failure. */
|
|
383
|
+
function describeFailure(displayPath, error) {
|
|
384
|
+
switch (error.code) {
|
|
385
|
+
case "encrypted": return `cannot read "${displayPath}": the document is encrypted; supply a decrypted copy`;
|
|
386
|
+
case "unsupported": return `cannot read "${displayPath}": ${error.message}`;
|
|
387
|
+
case "empty": return `cannot read "${displayPath}": ${error.message}`;
|
|
388
|
+
case "scanned": return `cannot read "${displayPath}": ${error.message}, which this tool does not perform`;
|
|
389
|
+
case "pageRange": return `cannot read "${displayPath}": ${error.message}`;
|
|
390
|
+
case "malformed":
|
|
391
|
+
case "missingPart": return `cannot read "${displayPath}": the file is damaged or incomplete (${error.message})`;
|
|
392
|
+
case "resourceLimit": return `cannot read "${displayPath}": the document exceeds the converter's internal safety limits (${error.message})`;
|
|
393
|
+
case "aborted": return `cannot read "${displayPath}": ${error.message}`;
|
|
394
|
+
case "io":
|
|
395
|
+
case "unknown": return `cannot read "${displayPath}": ${error.message}`;
|
|
396
|
+
}
|
|
397
|
+
}
|
|
398
|
+
/**
|
|
399
|
+
* Register the `read_document` tool and its system-prompt guidance.
|
|
400
|
+
* @param ctx - the plugin context; registrations are effects scoped to it.
|
|
401
|
+
* @param converterFor - the format router over the mounted engines.
|
|
402
|
+
* @param caps - resolved deployment bounds.
|
|
403
|
+
*/
|
|
404
|
+
function applyReadDocumentTool(ctx, converterFor, caps) {
|
|
405
|
+
ctx.systemPrompt.section({
|
|
406
|
+
name: "tool:read_document",
|
|
407
|
+
order: 100,
|
|
408
|
+
text: `Use the read_document tool — not read or shell commands — to inspect PDF, Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, and CSV files. It returns the document converted to line-numbered Markdown; use offset and limit to continue reading long documents. For a PDF, pass pages (for example "1-3,7") to read only those pages; page markers like <!-- Page 4 --> show where each page starts.`
|
|
409
|
+
});
|
|
410
|
+
ctx.tools.register(defineTool({
|
|
411
|
+
name: "read_document",
|
|
412
|
+
description: `Read a document file as Markdown: PDF, Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, and CSV files are converted to line-numbered Markdown text. Supported: ${EXTENSION_LIST}. Returns text only and writes no files. PDFs report page count, whether pages are scanned, and accept a pages selection; scanned or image-only pages are not OCRed. For plain text or source files use read; for images use read_image.`,
|
|
413
|
+
parameters: {
|
|
414
|
+
file_path: {
|
|
415
|
+
type: "string",
|
|
416
|
+
required: true,
|
|
417
|
+
description: "Path to the document, resolved by the filesystem backend; relative paths resolve against the session workspace."
|
|
418
|
+
},
|
|
419
|
+
offset: {
|
|
420
|
+
type: "number",
|
|
421
|
+
description: "1-based first line of the converted Markdown to return. Defaults to 1."
|
|
422
|
+
},
|
|
423
|
+
limit: {
|
|
424
|
+
type: "number",
|
|
425
|
+
description: `Maximum number of lines to return. Defaults to ${caps.readLimit}.`
|
|
426
|
+
},
|
|
427
|
+
pages: {
|
|
428
|
+
type: "string",
|
|
429
|
+
description: `PDF only: 1-based pages to convert, as numbers and ranges like "1-3,7" (at most ${caps.pdfMaxPages} pages). Defaults to every page.`
|
|
430
|
+
}
|
|
431
|
+
},
|
|
432
|
+
output: {
|
|
433
|
+
schema: {
|
|
434
|
+
type: "object",
|
|
435
|
+
additionalProperties: false,
|
|
436
|
+
properties: {
|
|
437
|
+
path: {
|
|
438
|
+
type: "string",
|
|
439
|
+
required: true
|
|
440
|
+
},
|
|
441
|
+
format: {
|
|
442
|
+
type: "string",
|
|
443
|
+
enum: [...DOCUMENT_FORMATS],
|
|
444
|
+
required: true
|
|
445
|
+
},
|
|
446
|
+
offset: {
|
|
447
|
+
type: "integer",
|
|
448
|
+
required: true
|
|
449
|
+
},
|
|
450
|
+
lines: {
|
|
451
|
+
type: "array",
|
|
452
|
+
required: true,
|
|
453
|
+
items: {
|
|
454
|
+
type: "object",
|
|
455
|
+
additionalProperties: false,
|
|
456
|
+
properties: {
|
|
457
|
+
number: {
|
|
458
|
+
type: "integer",
|
|
459
|
+
required: true
|
|
460
|
+
},
|
|
461
|
+
text: {
|
|
462
|
+
type: "string",
|
|
463
|
+
required: true
|
|
464
|
+
}
|
|
465
|
+
}
|
|
466
|
+
}
|
|
467
|
+
},
|
|
468
|
+
totalLines: {
|
|
469
|
+
type: "integer",
|
|
470
|
+
required: true
|
|
471
|
+
},
|
|
472
|
+
truncatedByBytes: {
|
|
473
|
+
type: "boolean",
|
|
474
|
+
required: true
|
|
475
|
+
},
|
|
476
|
+
pdf: {
|
|
477
|
+
type: "object",
|
|
478
|
+
additionalProperties: false,
|
|
479
|
+
properties: {
|
|
480
|
+
pageCount: {
|
|
481
|
+
type: "integer",
|
|
482
|
+
required: true
|
|
483
|
+
},
|
|
484
|
+
kind: {
|
|
485
|
+
type: "string",
|
|
486
|
+
enum: [...PDF_KINDS],
|
|
487
|
+
required: true
|
|
488
|
+
},
|
|
489
|
+
pagesNeedingOcr: {
|
|
490
|
+
type: "array",
|
|
491
|
+
required: true,
|
|
492
|
+
items: { type: "integer" }
|
|
493
|
+
},
|
|
494
|
+
pages: {
|
|
495
|
+
type: "array",
|
|
496
|
+
items: { type: "integer" }
|
|
497
|
+
},
|
|
498
|
+
title: { type: "string" }
|
|
499
|
+
}
|
|
500
|
+
}
|
|
501
|
+
}
|
|
502
|
+
},
|
|
503
|
+
render: (_args, value) => [{
|
|
504
|
+
type: "text",
|
|
505
|
+
text: formatReadOutput(value)
|
|
506
|
+
}]
|
|
507
|
+
},
|
|
508
|
+
isConcurrencySafe: () => true,
|
|
509
|
+
async execute(args, exec) {
|
|
510
|
+
const input = parseReadArgs(args, caps);
|
|
511
|
+
const format = formatOf(input.filePath);
|
|
512
|
+
if (format === void 0) throw new Error(`cannot read "${input.filePath}": unsupported extension. Supported: ${EXTENSION_LIST}. For plain text or source files use the read tool.`);
|
|
513
|
+
if (input.pages !== void 0 && format !== "pdf") throw new Error(`pages applies to PDF files only; "${input.filePath}" is ${format}`);
|
|
514
|
+
const cwd = exec.agent?.session.header.cwd;
|
|
515
|
+
const target = await ctx.fs.resolve(input.filePath, {
|
|
516
|
+
...cwd === void 0 ? {} : { cwd },
|
|
517
|
+
signal: exec.signal
|
|
518
|
+
});
|
|
519
|
+
const info = await ctx.fs.stat(target, exec.signal);
|
|
520
|
+
if (info === void 0) throw new Error(`cannot read "${target.displayPath}": not found`);
|
|
521
|
+
if (info.type !== "file") throw new Error(`cannot read "${target.displayPath}": not a regular file`);
|
|
522
|
+
const bytes = await ctx.fs.readBytes(target, exec.signal, caps.maxInputBytes);
|
|
523
|
+
let converted;
|
|
524
|
+
try {
|
|
525
|
+
converted = await converterFor(format).convert({
|
|
526
|
+
bytes,
|
|
527
|
+
format,
|
|
528
|
+
signal: exec.signal,
|
|
529
|
+
...input.pages === void 0 ? {} : { pages: input.pages }
|
|
530
|
+
});
|
|
531
|
+
} catch (error) {
|
|
532
|
+
if (error instanceof DocumentConversionError) throw new Error(describeFailure(target.displayPath, error));
|
|
533
|
+
throw error;
|
|
534
|
+
}
|
|
535
|
+
const window = windowLines(converted.markdown, {
|
|
536
|
+
offset: input.offset,
|
|
537
|
+
limit: input.limit,
|
|
538
|
+
maxLineLength: caps.maxLineLength,
|
|
539
|
+
maxBytes: caps.maxOutputBytes
|
|
540
|
+
});
|
|
541
|
+
return {
|
|
542
|
+
path: target.displayPath,
|
|
543
|
+
format,
|
|
544
|
+
offset: input.offset,
|
|
545
|
+
lines: window.lines,
|
|
546
|
+
totalLines: window.totalLines,
|
|
547
|
+
truncatedByBytes: window.truncatedByBytes,
|
|
548
|
+
...converted.pdf === void 0 ? {} : { pdf: converted.pdf }
|
|
549
|
+
};
|
|
550
|
+
},
|
|
551
|
+
presentCall(args) {
|
|
552
|
+
const { offset, limit, pages } = args;
|
|
553
|
+
const window = pages !== void 0 && pages.length > 0 ? ` (pages ${pages})` : limit !== void 0 && limit > 0 ? ` (lines ${offset ?? 1} - ${(offset ?? 1) + limit - 1})` : offset !== void 0 ? ` (from line ${offset})` : "";
|
|
554
|
+
return {
|
|
555
|
+
card: "generic",
|
|
556
|
+
title: `Read ${args.file_path} as Markdown${window}`,
|
|
557
|
+
kind: "read",
|
|
558
|
+
locations: [{ path: args.file_path }]
|
|
559
|
+
};
|
|
560
|
+
}
|
|
561
|
+
}));
|
|
562
|
+
}
|
|
563
|
+
//#endregion
|
|
564
|
+
//#region src/index.ts
|
|
565
|
+
const name = "document-tools";
|
|
566
|
+
const inject = [
|
|
567
|
+
"tools",
|
|
568
|
+
"fs",
|
|
569
|
+
"systemPrompt"
|
|
570
|
+
];
|
|
571
|
+
/** Default inclusive byte cap on a source document. */
|
|
572
|
+
const DEFAULT_MAX_INPUT_BYTES = 52428800;
|
|
573
|
+
/** Default and maximum lines per call. */
|
|
574
|
+
const DEFAULT_READ_LIMIT = 2e3;
|
|
575
|
+
/** Default per-line character cap. */
|
|
576
|
+
const DEFAULT_MAX_LINE_LENGTH = 2e3;
|
|
577
|
+
/** Default byte cap on returned line text per call. */
|
|
578
|
+
const DEFAULT_MAX_OUTPUT_BYTES = 51200;
|
|
579
|
+
/** Default cap on distinct pages one PDF `pages` selection may name. */
|
|
580
|
+
const DEFAULT_PDF_MAX_PAGES = 100;
|
|
581
|
+
const Config = Schema.object({
|
|
582
|
+
maxInputBytes: Schema.number().default(DEFAULT_MAX_INPUT_BYTES),
|
|
583
|
+
readLimit: Schema.number().default(DEFAULT_READ_LIMIT),
|
|
584
|
+
maxLineLength: Schema.number().default(DEFAULT_MAX_LINE_LENGTH),
|
|
585
|
+
maxOutputBytes: Schema.number().default(DEFAULT_MAX_OUTPUT_BYTES),
|
|
586
|
+
pdfMaxPages: Schema.number().default(100),
|
|
587
|
+
pdfProfile: Schema.union(["fidelity", "compact"]).default("fidelity")
|
|
588
|
+
});
|
|
589
|
+
function assertPositiveInteger(name, value) {
|
|
590
|
+
if (!Number.isInteger(value) || value < 1) throw new Error(`document-tools: ${name} must be a positive integer`);
|
|
591
|
+
}
|
|
592
|
+
/**
|
|
593
|
+
* Register `read_document` over the anydoc and pdf-inspector engines.
|
|
594
|
+
* @param ctx - context carrying the tool registry, filesystem seam, and system-prompt registry.
|
|
595
|
+
* @param rawConfig - schema-validated bounds after defaulting.
|
|
596
|
+
*/
|
|
597
|
+
function apply(ctx, rawConfig) {
|
|
598
|
+
const config = rawConfig;
|
|
599
|
+
assertPositiveInteger("maxInputBytes", config.maxInputBytes);
|
|
600
|
+
assertPositiveInteger("readLimit", config.readLimit);
|
|
601
|
+
assertPositiveInteger("maxLineLength", config.maxLineLength);
|
|
602
|
+
assertPositiveInteger("maxOutputBytes", config.maxOutputBytes);
|
|
603
|
+
assertPositiveInteger("pdfMaxPages", config.pdfMaxPages);
|
|
604
|
+
applyReadDocumentTool(ctx, routeConverters([pdfInspectorConverter({ profile: config.pdfProfile }), anydocConverter()]), {
|
|
605
|
+
maxInputBytes: config.maxInputBytes,
|
|
606
|
+
readLimit: config.readLimit,
|
|
607
|
+
maxLineLength: config.maxLineLength,
|
|
608
|
+
maxOutputBytes: config.maxOutputBytes,
|
|
609
|
+
pdfMaxPages: config.pdfMaxPages
|
|
610
|
+
});
|
|
611
|
+
}
|
|
612
|
+
//#endregion
|
|
613
|
+
export { Config, DEFAULT_MAX_INPUT_BYTES, DEFAULT_MAX_LINE_LENGTH, DEFAULT_MAX_OUTPUT_BYTES, DEFAULT_PDF_MAX_PAGES, DEFAULT_READ_LIMIT, DOCUMENT_EXTENSIONS, DOCUMENT_FORMATS, DocumentConversionError, anydocConverter, apply, formatOf, inject, name, pdfInspectorConverter, routeConverters };
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
import { MarkdownProfile } from "@firecrawl/pdf-inspector-wasm";
|
|
2
|
+
//#region src/pdf-worker.d.ts
|
|
3
|
+
/** What the parent posts as `workerData`. */
|
|
4
|
+
interface PdfWorkerRequest {
|
|
5
|
+
bytes: Uint8Array;
|
|
6
|
+
/** 1-based pages to convert; absent means all. */
|
|
7
|
+
pages?: number[];
|
|
8
|
+
profile: MarkdownProfile;
|
|
9
|
+
}
|
|
10
|
+
/** The engine facts the parent needs; everything else stays in the worker. */
|
|
11
|
+
interface PdfWorkerResult {
|
|
12
|
+
markdown: string;
|
|
13
|
+
pdfType: 'TextBased' | 'Scanned' | 'ImageBased' | 'Mixed';
|
|
14
|
+
pageCount: number;
|
|
15
|
+
/** 1-based. */
|
|
16
|
+
pagesNeedingOcr: number[];
|
|
17
|
+
title?: string;
|
|
18
|
+
}
|
|
19
|
+
/** Worker → parent message. */
|
|
20
|
+
type PdfWorkerReply = {
|
|
21
|
+
ok: true;
|
|
22
|
+
result: PdfWorkerResult;
|
|
23
|
+
} | {
|
|
24
|
+
ok: false;
|
|
25
|
+
message: string;
|
|
26
|
+
};
|
|
27
|
+
//#endregion
|
|
28
|
+
export { PdfWorkerRequest as n, PdfWorkerResult as r, PdfWorkerReply as t };
|
|
@@ -0,0 +1,47 @@
|
|
|
1
|
+
import { createRequire } from "node:module";
|
|
2
|
+
import { parentPort, workerData } from "node:worker_threads";
|
|
3
|
+
import { readFileSync } from "node:fs";
|
|
4
|
+
import { initSync, processPdf } from "@firecrawl/pdf-inspector-wasm";
|
|
5
|
+
//#region src/pdf-worker.ts
|
|
6
|
+
/**
|
|
7
|
+
* Worker-thread entry for PDF conversion. Runs `@firecrawl/pdf-inspector`'s
|
|
8
|
+
* WASM build, whose `processPdf` is synchronous, off the main event loop;
|
|
9
|
+
* the parent terminates the worker to cancel. This module is executed by
|
|
10
|
+
* Node directly (never bundled into the plugin entry), so it stays free of
|
|
11
|
+
* TypeScript-only runtime syntax and imports nothing from the plugin.
|
|
12
|
+
* @module @jiaoqsh/dsh-document/pdf-worker
|
|
13
|
+
*/
|
|
14
|
+
const require = createRequire(import.meta.url);
|
|
15
|
+
function run(request) {
|
|
16
|
+
initSync({ module: readFileSync(require.resolve("@firecrawl/pdf-inspector-wasm/pdf_inspector_wasm_bg.wasm")) });
|
|
17
|
+
const result = processPdf(new Uint8Array(request.bytes), {
|
|
18
|
+
profile: request.profile,
|
|
19
|
+
includePageMarkers: true,
|
|
20
|
+
includeImages: false,
|
|
21
|
+
...request.pages === void 0 ? {} : { pages: request.pages }
|
|
22
|
+
});
|
|
23
|
+
return {
|
|
24
|
+
markdown: result.markdown ?? "",
|
|
25
|
+
pdfType: result.pdfType,
|
|
26
|
+
pageCount: result.pageCount,
|
|
27
|
+
pagesNeedingOcr: result.pagesNeedingOcr,
|
|
28
|
+
...result.title === void 0 ? {} : { title: result.title }
|
|
29
|
+
};
|
|
30
|
+
}
|
|
31
|
+
if (parentPort !== null) {
|
|
32
|
+
let reply;
|
|
33
|
+
try {
|
|
34
|
+
reply = {
|
|
35
|
+
ok: true,
|
|
36
|
+
result: run(workerData)
|
|
37
|
+
};
|
|
38
|
+
} catch (error) {
|
|
39
|
+
reply = {
|
|
40
|
+
ok: false,
|
|
41
|
+
message: error instanceof Error ? error.message : String(error)
|
|
42
|
+
};
|
|
43
|
+
}
|
|
44
|
+
parentPort.postMessage(reply);
|
|
45
|
+
}
|
|
46
|
+
//#endregion
|
|
47
|
+
export {};
|
package/package.json
ADDED
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
{
|
|
2
|
+
"name": "@jiaoqsh/dsh-document",
|
|
3
|
+
"version": "0.1.0",
|
|
4
|
+
"description": "DeepSeek Harness bundle: the read_document tool reads Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF files as Markdown for the model",
|
|
5
|
+
"keywords": [
|
|
6
|
+
"deepseek-harness",
|
|
7
|
+
"dsh",
|
|
8
|
+
"plugin",
|
|
9
|
+
"document",
|
|
10
|
+
"pdf",
|
|
11
|
+
"docx",
|
|
12
|
+
"markdown",
|
|
13
|
+
"anydoc"
|
|
14
|
+
],
|
|
15
|
+
"license": "MIT",
|
|
16
|
+
"publishConfig": {
|
|
17
|
+
"access": "public"
|
|
18
|
+
},
|
|
19
|
+
"author": "jiaoqsh",
|
|
20
|
+
"repository": {
|
|
21
|
+
"type": "git",
|
|
22
|
+
"url": "git+https://github.com/jiaoqsh/dsh-document.git"
|
|
23
|
+
},
|
|
24
|
+
"type": "module",
|
|
25
|
+
"main": "lib/index.js",
|
|
26
|
+
"types": "lib/index.d.ts",
|
|
27
|
+
"exports": {
|
|
28
|
+
".": {
|
|
29
|
+
"types": "./lib/index.d.ts",
|
|
30
|
+
"default": "./lib/index.js"
|
|
31
|
+
},
|
|
32
|
+
"./package.json": "./package.json"
|
|
33
|
+
},
|
|
34
|
+
"files": [
|
|
35
|
+
"lib",
|
|
36
|
+
"cordis.patch.yml"
|
|
37
|
+
],
|
|
38
|
+
"engines": {
|
|
39
|
+
"node": ">=22.19"
|
|
40
|
+
},
|
|
41
|
+
"dsh": {
|
|
42
|
+
"bundle": {
|
|
43
|
+
"patch": "./cordis.patch.yml"
|
|
44
|
+
}
|
|
45
|
+
},
|
|
46
|
+
"scripts": {
|
|
47
|
+
"build": "tsdown",
|
|
48
|
+
"prepare": "tsdown",
|
|
49
|
+
"typecheck": "tsc -p tsconfig.json",
|
|
50
|
+
"test": "vitest run",
|
|
51
|
+
"check": "pnpm run typecheck && pnpm run test && pnpm run build"
|
|
52
|
+
},
|
|
53
|
+
"peerDependencies": {
|
|
54
|
+
"@deepseek-ai/cordis": "^4.0.1",
|
|
55
|
+
"@deepseek-ai/dsh-fs": ">=0.1.0-rc.5",
|
|
56
|
+
"@deepseek-ai/dsh-system-prompt": ">=0.1.0-rc.5",
|
|
57
|
+
"@deepseek-ai/dsh-tools": ">=0.1.0-rc.5"
|
|
58
|
+
},
|
|
59
|
+
"dependencies": {
|
|
60
|
+
"@deepseek-ai/schemastery": "^3.18.1",
|
|
61
|
+
"@firecrawl/anydoc": "0.1.9",
|
|
62
|
+
"@firecrawl/pdf-inspector-wasm": "1.14.2"
|
|
63
|
+
},
|
|
64
|
+
"devDependencies": {
|
|
65
|
+
"@deepseek-ai/cordis": "^4.0.1",
|
|
66
|
+
"@deepseek-ai/dsh-fs": "0.1.0-rc.6",
|
|
67
|
+
"@deepseek-ai/dsh-fs-local": "0.1.0-rc.6",
|
|
68
|
+
"@deepseek-ai/dsh-llm": "0.1.0-rc.6",
|
|
69
|
+
"@deepseek-ai/dsh-system-prompt": "0.1.0-rc.6",
|
|
70
|
+
"@deepseek-ai/dsh-tools": "0.1.0-rc.6",
|
|
71
|
+
"@types/node": "^22.20.0",
|
|
72
|
+
"tsdown": "^0.22.2",
|
|
73
|
+
"typescript": "^6.0.3",
|
|
74
|
+
"vitest": "^4.1.8"
|
|
75
|
+
}
|
|
76
|
+
}
|