@docx-editor.dev/docx-to-markdown 2.20.0 → 2.21.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,15 +1,12 @@
1
1
  # DOCX to Markdown
2
2
 
3
- Use `@docx-editor.dev/docx-to-markdown` to convert DOCX files to Markdown.
4
- The result includes the document body, individual pages, and separate headers, footers, comments, and tracked changes.
3
+ Use `@docx-editor.dev/docx-to-markdown` to convert DOCX files to Markdown. The result includes the document body, individual pages, and separate headers, footers, comments, and tracked changes.
5
4
 
6
- [Try the DOCX to Markdown demo](https://docx-to-markdown.docx-editor.dev/) or read the
7
- [Markdown export guide](https://www.docx-editor.dev/docs/2.x/export/markdown).
5
+ [Try the DOCX to Markdown demo](https://docx-to-markdown.docx-editor.dev/) or read the [Markdown export guide](https://www.docx-editor.dev/docs/2.x/export/markdown).
8
6
 
9
7
  ## Before you begin
10
8
 
11
- For Node.js, use version 20.16.0 or later in the 20.x release line, or version 22.3.0 or later.
12
- The converter requires WebAssembly and uses bundled fonts by default.
9
+ For Node.js, use version 20.16.0 or later in the 20.x release line, or version 22.3.0 or later. The converter requires WebAssembly and uses bundled fonts by default.
13
10
 
14
11
  ## Install the package
15
12
 
@@ -40,18 +37,11 @@ Set `images: true` to include image links and extracted bytes:
40
37
  const result = await exportMarkdown(docxBytes, { images: true });
41
38
  ```
42
39
 
43
- `result.media` contains image bytes and their page occurrences.
44
- See [image workflows](docs/images.md) to save a folder, download a ZIP, or return hosted URLs.
40
+ `result.media` contains image bytes and their page occurrences. See [image workflows](docs/images.md) to save a folder, download a ZIP, or return hosted URLs.
45
41
 
46
42
  ### Preserve displayed image sizes
47
43
 
48
- Use `images: { syntax: 'html' }` to include each image's displayed width and height in generated `<img>` tags.
49
- Dimensions use whole CSS pixels. Configure your Markdown renderer to allow sanitized HTML and retain `width` and `height`.
50
- The default `images: true` uses standard Markdown image syntax, which has no size attributes.
51
- For custom previews, use each occurrence's `displayWidthPx` and `displayHeightPx`.
52
- Asset `pixelWidth` and `pixelHeight` describe the image file's dimensions.
53
- Crop, rotation, and floating text wrapping are not reproduced.
54
- See [displayed image sizes and custom previews](docs/images.md#preserve-displayed-image-sizes).
44
+ Use `images: { syntax: 'html' }` to include each image's displayed width and height in generated `<img>` tags. Dimensions use whole CSS pixels. Configure your Markdown renderer to allow sanitized HTML and retain `width` and `height`. The default `images: true` uses standard Markdown image syntax, which has no size attributes. For custom previews, use each occurrence's `displayWidthPx` and `displayHeightPx`. Asset `pixelWidth` and `pixelHeight` describe the image file's dimensions. Crop, rotation, and floating text wrapping are not reproduced. See [displayed image sizes and custom previews](docs/images.md#preserve-displayed-image-sizes).
55
45
 
56
46
  ## Read page output
57
47
 
@@ -82,8 +72,7 @@ For Next.js, use the Node.js runtime and [server package configuration](docs/int
82
72
 
83
73
  Page breaks depend on fonts, document features, and revision mode; they can differ from Microsoft Word. Store the document version with page citations. `result.warnings` reports omitted content and font problems. Images are omitted unless enabled. See [output limits](docs/api.md#markdown-limitations) before using the output as a complete transcription.
84
74
 
85
- The package uses the Apache 2.0 license, including comment and tracked-change extraction.
86
- Bundled fonts retain their own open-source licenses.
75
+ The package uses the Apache 2.0 license, including comment and tracked-change extraction. Bundled fonts retain their own open-source licenses.
87
76
 
88
77
  ## Next steps
89
78
 
package/docs/api.md CHANGED
@@ -1,7 +1,6 @@
1
1
  # DOCX to Markdown API reference
2
2
 
3
- Use this reference to choose an export function, configure conversion, and inspect the result.
4
- For installation and a first export, see the [package quickstart](../README.md).
3
+ Use this reference to choose an export function, configure conversion, and inspect the result. For installation and a first export, see the [package quickstart](../README.md).
5
4
 
6
5
  ## Public interface
7
6
 
@@ -157,12 +156,7 @@ Tracked changes also participate in layout through `displayMode`: `all-markup` (
157
156
 
158
157
  Enable `images: true` for relative image links and `result.media` bytes. Use `{ images: { resolveUrl, maxTotalBytes } }` for custom delivery. The default extracted-byte limit is 64 MiB; image extraction is opt-in. `exportMarkdownFrom(session, options)` accepts the same image options and an abort signal.
159
158
 
160
- Set `images: { syntax: 'html' }` to emit `<img>` tags with each occurrence's displayed width and height in whole CSS pixels.
161
- The default `syntax: 'markdown'` emits standard image links without size attributes.
162
- Your renderer must support sanitized HTML and retain `width` and `height`.
163
- Occurrences expose exact `displayWidthPx`, `displayHeightPx`, and `kind` (`inline` or `anchored`).
164
- Asset `pixelWidth` and `pixelHeight` describe the image bytes, not their displayed size.
165
- Crop, rotation, and floating text wrapping are not reproduced.
159
+ Set `images: { syntax: 'html' }` to emit `<img>` tags with each occurrence's displayed width and height in whole CSS pixels. The default `syntax: 'markdown'` emits standard image links without size attributes. Your renderer must support sanitized HTML and retain `width` and `height`. Occurrences expose exact `displayWidthPx`, `displayHeightPx`, and `kind` (`inline` or `anchored`). Asset `pixelWidth` and `pixelHeight` describe the image bytes, not their displayed size. Crop, rotation, and floating text wrapping are not reproduced.
166
160
 
167
161
  `createMarkdownZip(result)` and `toMarkdownJSON(result)` are exported from the main package. `writeMarkdownBundle(result, { directory })` comes from `@docx-editor.dev/docx-to-markdown/node`. See [image APIs, errors, ownership, and runnable workflows](images.md).
168
162
 
package/docs/images.md CHANGED
@@ -1,7 +1,6 @@
1
1
  # Include images in Markdown exports
2
2
 
3
- Before you begin, [install the converter](../README.md#install-the-package).
4
- Set `images: true` to include image links and extracted bytes. For this Node.js example, save the code as `convert.mjs`, place `document.docx` beside it, and run `node convert.mjs`:
3
+ Before you begin, [install the converter](../README.md#install-the-package). Set `images: true` to include image links and extracted bytes. For this Node.js example, save the code as `convert.mjs`, place `document.docx` beside it, and run `node convert.mjs`:
5
4
 
6
5
  ```ts
7
6
  import { readFile } from 'node:fs/promises';
@@ -15,19 +14,13 @@ console.log(result.media); // Unique bytes, paths, URLs, dimensions, and occurre
15
14
 
16
15
  `images: true` and `images: {}` use relative URLs. If you omit `images`, the result contains no image links and returns `media: []`.
17
16
 
18
- Each image contains its ID, path, URL, MIME type, bytes, intrinsic dimensions, and occurrences.
19
- The ID is a hexadecimal SHA-256 digest without the `sha256:` prefix. MIME types and extensions describe the exported bytes, including converted images.
17
+ Each image contains its ID, path, URL, MIME type, bytes, intrinsic dimensions, and occurrences. The ID is a hexadecimal SHA-256 digest without the `sha256:` prefix. MIME types and extensions describe the exported bytes, including converted images.
20
18
 
21
- Each occurrence records its page, story, source position, displayed dimensions, drawing kind, and alternative text.
22
- Repeated uses of identical bytes share one asset with separate occurrences. Source offsets use UTF-16 code units.
23
- Keep the document version with stored citations: occurrence and page identifiers belong to that export snapshot.
24
- See the [image types](https://github.com/eigenpal/docx-editor/blob/main/packages/docx-to-markdown/src/media-types.ts) for all returned fields.
19
+ Each occurrence records its page, story, source position, displayed dimensions, drawing kind, and alternative text. Repeated uses of identical bytes share one asset with separate occurrences. Source offsets use UTF-16 code units. Keep the document version with stored citations: occurrence and page identifiers belong to that export snapshot. See the [image types](https://github.com/eigenpal/docx-editor/blob/main/packages/docx-to-markdown/src/media-types.ts) for all returned fields.
25
20
 
26
21
  ## Preserve displayed image sizes
27
22
 
28
- An image file's `pixelWidth` and `pixelHeight` describe its intrinsic pixels. Word can display the same file at different sizes.
29
- Use each occurrence's `displayWidthPx` and `displayHeightPx` for the document's displayed size, in CSS pixels at 96 pixels per inch.
30
- These values retain fractional pixels. `kind` distinguishes inline and anchored drawings.
23
+ An image file's `pixelWidth` and `pixelHeight` describe its intrinsic pixels. Word can display the same file at different sizes. Use each occurrence's `displayWidthPx` and `displayHeightPx` for the document's displayed size, in CSS pixels at 96 pixels per inch. These values retain fractional pixels. `kind` distinguishes inline and anchored drawings.
31
24
 
32
25
  Standard Markdown image syntax has no width or height attributes. To carry each occurrence's size into your Markdown renderer, replace the conversion call in the first example with:
33
26
 
@@ -38,14 +31,9 @@ const result = await exportMarkdown(docxBytes, {
38
31
  // <img src="media/<digest>.png" alt="Banner" width="300" height="80">
39
32
  ```
40
33
 
41
- The converter escapes HTML attributes and rounds the display dimensions to whole CSS pixels. Extents smaller than half a pixel round to zero.
42
- This option works with local folders, ZIP downloads, and `resolveUrl` for server storage.
43
- The default `syntax: 'markdown'` keeps standard image links without dimensions.
44
- Both options return full occurrence metadata.
34
+ The converter escapes HTML attributes and rounds the display dimensions to whole CSS pixels. Extents smaller than half a pixel round to zero. This option works with local folders, ZIP downloads, and `resolveUrl` for server storage. The default `syntax: 'markdown'` keeps standard image links without dimensions. Both options return full occurrence metadata.
45
35
 
46
- Configure your renderer to parse HTML, sanitize it, and retain `img` attributes `src`, `alt`, `width`, and `height`.
47
- For example, [`react-markdown`](https://github.com/remarkjs/react-markdown#appendix-a-html-in-markdown) supports `rehype-raw` followed by [`rehype-sanitize`](https://github.com/rehypejs/rehype-sanitize); the default sanitizer retains these attributes.
48
- If your renderer disables HTML or removes size attributes, use the metadata in a custom preview.
36
+ Configure your renderer to parse HTML, sanitize it, and retain `img` attributes `src`, `alt`, `width`, and `height`. For example, [`react-markdown`](https://github.com/remarkjs/react-markdown#appendix-a-html-in-markdown) supports `rehype-raw` followed by [`rehype-sanitize`](https://github.com/rehypejs/rehype-sanitize); the default sanitizer retains these attributes. If your renderer disables HTML or removes size attributes, use the metadata in a custom preview.
49
37
 
50
38
  For a custom React preview, pass the selected occurrence and its asset's trusted preview URL:
51
39
 
@@ -78,17 +66,11 @@ function ImagePreview({
78
66
  }
79
67
  ```
80
68
 
81
- The explicit aspect ratio preserves Word's displayed proportions when the image shrinks, even if they differ from its intrinsic proportions.
82
- Use the same styles in a custom Markdown image component, with `width` and `height` from the generated HTML.
83
- Resolve `src` through your known asset URLs, as shown in the browser example.
69
+ The explicit aspect ratio preserves Word's displayed proportions when the image shrinks, even if they differ from its intrinsic proportions. Use the same styles in a custom Markdown image component, with `width` and `height` from the generated HTML. Resolve `src` through your known asset URLs, as shown in the browser example.
84
70
 
85
- Do not use an asset's first occurrence to size every image with the same URL.
86
- Occurrences describe physical layout and include repeated headers; their order is not a Markdown image index.
87
- Use HTML syntax to attach the correct dimensions directly to each rendered image.
88
- For a preview built from occurrence metadata, identify the occurrence by its page, story, part, and drawing node.
71
+ Do not use an asset's first occurrence to size every image with the same URL. Occurrences describe physical layout and include repeated headers; their order is not a Markdown image index. Use HTML syntax to attach the correct dimensions directly to each rendered image. For a preview built from occurrence metadata, identify the occurrence by its page, story, part, and drawing node.
89
72
 
90
- Displayed dimensions describe the drawing's extent before crop and rotation. They do not reproduce cropping, rotation, effects, alignment, or floating text wrapping.
91
- Anchored images appear at their source paragraph positions in Markdown.
73
+ Displayed dimensions describe the drawing's extent before crop and rotation. They do not reproduce cropping, rotation, effects, alignment, or floating text wrapping. Anchored images appear at their source paragraph positions in Markdown.
92
74
 
93
75
  ## Save a local folder
94
76
 
@@ -103,9 +85,7 @@ const result = await exportMarkdown(await readFile('input.docx'), { images: true
103
85
  await writeMarkdownBundle(result, { directory: './output' });
104
86
  ```
105
87
 
106
- The helper writes `document.md`, `document.json`, and `media/`.
107
- The parent directory must exist, and the output directory must be new or empty. Existing files are not overwritten.
108
- Files use `0o600` permissions where supported. JSON includes page, review, and image metadata without image bytes.
88
+ The helper writes `document.md`, `document.json`, and `media/`. The parent directory must exist, and the output directory must be new or empty. Existing files are not overwritten. Files use `0o600` permissions where supported. JSON includes page, review, and image metadata without image bytes.
109
89
 
110
90
  ## Save a ZIP file
111
91
 
@@ -123,9 +103,7 @@ The ZIP includes `document.md`, `document.json`, and `media/`. Keep generated re
123
103
 
124
104
  ## Convert and download in the browser
125
105
 
126
- Configure your bundler to serve the package's font and WebAssembly assets.
127
- For a working configuration, see the [browser demo](https://github.com/eigenpal/docx-editor/tree/main/examples/docx-to-markdown).
128
- Image extraction runs in the browser without a server or storage service.
106
+ Configure your bundler to serve the package's font and WebAssembly assets. For a working configuration, see the [browser demo](https://github.com/eigenpal/docx-editor/tree/main/examples/docx-to-markdown). Image extraction runs in the browser without a server or storage service.
129
107
 
130
108
  ```ts
131
109
  import { exportMarkdown, createMarkdownZip } from '@docx-editor.dev/docx-to-markdown';
@@ -1,7 +1,6 @@
1
1
  # Integrate DOCX to Markdown
2
2
 
3
- Before you begin, [install the converter and check the runtime requirements](../README.md#before-you-begin).
4
- Use `result.markdown` for text and `result.pages` for page citations. See [image delivery](images.md) for browser ZIP downloads, local folders, and server URLs with JSON metadata.
3
+ Before you begin, [install the converter and check the runtime requirements](../README.md#before-you-begin). Use `result.markdown` for text and `result.pages` for page citations. See [image delivery](images.md) for browser ZIP downloads, local folders, and server URLs with JSON metadata.
5
4
 
6
5
  ## LangChain
7
6
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@docx-editor.dev/docx-to-markdown",
3
- "version": "2.20.0",
3
+ "version": "2.21.0",
4
4
  "private": false,
5
5
  "publishConfig": {
6
6
  "access": "public"
@@ -46,10 +46,10 @@
46
46
  "check:consumer": "node ../../scripts/check-markdown-media-consumer.mjs"
47
47
  },
48
48
  "peerDependencies": {
49
- "@docx-editor.dev/core": "~2.20.0"
49
+ "@docx-editor.dev/core": "~2.21.0"
50
50
  },
51
51
  "dependencies": {
52
- "@docx-editor.dev/fonts": "~2.20.0",
52
+ "@docx-editor.dev/fonts": "~2.21.0",
53
53
  "fflate": "^0.8.2"
54
54
  },
55
55
  "devDependencies": {