@onodocs/canvas 0.2.1 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/SDK.md CHANGED
@@ -1,219 +1,407 @@
1
- # OnoDocs SDK
2
-
3
- `@onodocs/sdk` processes Word documents. Its default entry performs analysis in Node or a browser; `/browser` adds layout and immutable page publication using browser font services; `/server` runs that processing in headless Chromium. `@onodocs/canvas` paints prepared pages and provides the viewer, selection and DOM attachments. The Canvas package runs independently of the parser, semantic engine and layout engine.
4
-
5
- Supported Word inputs are DOCX, DOCM, DOTX and DOTM. Macro-enabled files are rendered without executing VBA. Older binary DOC/DOT files and RTF are not supported.
6
-
7
- Install with `npm install @onodocs/sdk @onodocs/canvas`. Use matching versions of both packages. A backend can install only the SDK; a frontend receiving prepared pages can install only Canvas. Both packages include TypeScript declarations and this guide. Archives are also available from the [public releases](https://github.com/onodocs/onodocs/releases). Server rendering requires an installed Chromium-family browser. The server adapter includes the Canvas runtime it uses for image and PDF export.
8
-
9
- Without a key, the SDK runs in free, non-production evaluation mode with full features and a watermark on rendered output. Pass `{ licenseKey }` to `openDocument` for a 30-day trial or commercial entitlement; this works in all three entry points. Verification is entirely offline. Commercial keys permit covered releases indefinitely, with renewal for later releases and support. Browser verification requires HTTPS or a secure local context. See the included `LICENSING.md` and `LICENSE` for details.
10
-
11
- ## Inspect and query
12
-
13
- ```ts
14
- import { openDocument } from "@onodocs/sdk";
15
-
16
- const doc = await openDocument(bytes);
17
- const customer = doc.query.contentControls().where({ tag: "customer-name" }).one();
18
- const cells = doc.query.tables().cells().all();
19
- const matches = doc.query.findText("Invoice total").all();
20
- const headers = doc.query.stories().where({ story: "header" }).paragraphs().all();
21
- ```
22
-
23
- TypeScript completion shows valid paths and filters. Select a kind, narrow with `where`, then choose how many results you expect:
24
-
25
- | Method | Result |
26
- | --- | --- |
27
- | `one()` | One match; throws for zero or multiple matches |
28
- | `optional()` | One match or undefined; throws for multiple matches |
29
- | `first()` / `at(index)` | An explicitly chosen match or undefined |
30
- | `all()` / iteration | Every match in document traversal order |
31
- | `count()` | Number of matches |
32
- | `map(project)` | Project matches into application data |
33
-
34
- Selectors (`paragraphs`, `runs`, `tables`, `contentControls`, `bookmarks`, `stories`) search descendants. `rows` selects a table's own rows; `cells` selects a table or row's own cells. Nested tables are selected separately with `tables()`. Typed element collections also allow `table.rows[0].cells[1]`. `children()` and `descendants()` provide generic traversal. Scope an existing element with `doc.query.within(element)`. Nested or overlapping selections never duplicate the same element. A story represents an authored story; a reused header remains one semantic story with several physical occurrences.
35
-
36
- `where({ text: "Total" })` matches exactly and case sensitively. Use `where({ text: { contains: "Total" } })`, `startsWith`, or `endsWith` for other matches. Multiple fields and chained filters combine with AND. Paragraphs support `style`; content controls support `tag` and `title`; bookmarks support `name`; stories support `story` (`body`, `header`, `footer`, `footnote`, `endnote`, `textbox`). Unknown keys fail instead of silently returning unexpected content. Use `.filter(element => ...)` for application predicates.
37
-
38
- `doc.query.get(id)` resolves an element in the current snapshot. Every element has a `parent`; `doc.query.within(element).closest("cell")` finds its containing cell, including the input itself if it is a cell. Negative query indices count from the end.
39
-
40
- `findText("...")` or `findText(/pattern/i)` searches contiguous text in selected paragraphs or runs, including text split across runs. It returns a paragraph and a half-open `[start, end)` range in JavaScript UTF-16 offsets. It does not join separate paragraphs. Regular expressions search each selected contiguous interval, preserve the caller's lastIndex, and omit zero-length matches. Queries preserve the selected revision view. Semantic text includes authored content; generated fields and physical appearances can differ after layout. Rasterized picture/chart text has no text query model.
41
-
42
- Bookmarks expose their published start target through `bookmark.run`; they do not represent an enclosing bookmark text range. Symbols without a Unicode text representation use the object replacement character in query text.
43
-
44
- ## Replace text and fill templates
45
-
46
- ```ts
47
- const customer = doc.query.contentControls().where({ tag: "customer" }).one();
48
- await doc.update({ target: customer, text: "Willow Design" });
49
- ```
50
-
51
- `update` accepts one update or a batch. Targets are runs, paragraphs, cells, content controls, their IDs, or search ranges. Element targets must contain ordinary text in one paragraph. Replacement text inherits the first selected run’s formatting; other selected runs become empty. Tabs and newlines are supported. Batches reject overlapping or missing targets, generated fields, and multi-paragraph replacements without changing the document.
52
-
53
- Use search ranges for partial replacements while keeping surrounding text and formatting:
54
-
55
- ```ts
56
- await doc.update(doc.query.findText("{{customer}}").map(target => ({ target, text: "Willow Design" })));
57
- ```
58
-
59
- Batch range offsets refer to the original paragraph text. Empty ranges insert text; offsets must not split an emoji or other surrogate pair. Unchanged updates preserve existing formatting, snapshots, and attachments.
60
-
61
- Browser updates rebuild layout and refresh mounted views. Query, semantic, layout, page and geometry values obtained earlier remain snapshots; read them again after updating. Application attachments are removed during refresh and can be reattached using fresh geometry. Updates run in submission order and accept `{ signal }` for cancellation. The original `source` is unchanged.
62
-
63
- Node inspection exposes the same update method. For server rendering, inspect the input with the default entry point to select elements or ranges, then pass those updates to the server document before `renderPage` or `pdf`. Exports reflect the updated text. Saving an updated DOCX, structural editing, and field recalculation are not supported.
64
-
65
- ## Render and attach ordinary DOM
66
-
67
- Mounted views include read-only text selection and plain-text copying. Drag or Shift-click to select text, extend with Shift and the arrow/Home/End keys, select all with Ctrl/Command+A, and copy with Ctrl/Command+C or the Copy context menu. Selection spans pages and survives scrolling and zoom, including pages whose canvases have been released. Document updates clear the previous selection. Attached inputs retain their normal editing and clipboard behavior.
68
-
69
- ```ts
70
- import { openDocument } from "@onodocs/sdk/browser";
71
- import { createDocument } from "@onodocs/canvas";
72
-
73
- const doc = await openDocument(file);
74
- const canvasDocument = createDocument(doc);
75
- const view = canvasDocument.mount(container);
76
- const field = doc.query.contentControls().where({ tag: "customer-name" }).one();
77
- const input = document.createElement("input");
78
- input.name = "customerName";
79
- const attachment = view.attach(input, { anchor: doc.geometry.fragments(field)[0] });
80
- ```
81
-
82
- The example requires a tagged control that produces one visible fragment. Call `doc.geometry.fragments(field)` for content spanning lines, pages, or repeated stories, and attach to an explicitly chosen fragment. Fragments retain transforms and clipping. Table/cell anchors use their final physical rectangles. `offset` and `size` optionally adjust attachment placement in page units. Values, validation, submission, focus, and accessibility labels belong to the application. Attaching an element moves it into the view. Detachment removes it and restores its original inline style. Reattaching the same HTML element disposes its previous attachment; old handles become inert. An attachment exposes its disposed state.
83
-
84
- Pages are anchors too. Placement avoids manual rectangle arithmetic:
85
-
86
- ```js
87
- view.attach(checkbox, { anchor: doc.geometry.fragments(doc.query.findText("Client approval:").one())[0], placement: "outside-right", gap: 80, size: { width: 420, height: 420 } });
88
- view.attach(submitButton, { anchor: doc.pages[0], placement: "inside-bottom-right", inset: 480, size: { width: 2400, height: 600 } });
89
- ```
90
-
91
- Inside placements use `inside-{top|center|bottom}-{left|center|right}` with optional `inset`. Outside placements use `outside-{top|right|bottom|left}` with optional `gap` and center the element along the selected edge. The default is `overlay`. All sizes, spacing and offsets use twips (1/1440 inch). Without a size, the element fills the anchor, reduced by the inset for inside placement. Offsets apply after placement. Content anchors retain their transforms and clips; page anchors use the full page. Attachments do not reflow text, so reserve room in the document.
92
-
93
- Adjacent styled runs on one line share a fragment. Fragment bounds enclose the unclipped geometry; the supplied clips determine which portions are visible.
94
-
95
- `view.setZoom("fit-width")` tracks container width. Numeric zoom uses 96 CSS pixels per inch at `1`. `view.toClient(pageIndex, point)` and `view.fromClient({ x: event.clientX, y: event.clientY })` translate coordinates. `view.hitTest(clientPoint)` returns `{ elementId, fragment, caret? }`. Resolve `elementId` with `doc.query.get(elementId)` when the processing SDK is available. Text hits resolve to a run or paragraph; table whitespace resolves to a cell or table. `caret` contains `{ paragraphId, offset, point }`, with the same UTF-16 offset convention as search ranges. Generated field text omits a caret when no source position can be represented. Raster drawings without a semantic hit identity and blank page space return no element.
96
-
97
- `view.scrollTo(anchor, { block: "center" })` reveals a page or a physical fragment. Resolve elements and text ranges through `doc.geometry.fragments(anchor)` first, then choose the occurrence to reveal.
98
-
99
- A run's optional `link` is `{ kind: "external", target, tooltip }` or `{ kind: "bookmark", target }`. Use bookmark queries and scrollTo for internal links. The host application decides whether and how to open external destinations.
100
-
101
- All page and query indices are **zero based**. Page geometry uses **twips: 1440 units per inch**. `await canvasDocument.pages[index].render(canvas, { dpi: 144 })` prepares the page resources and paints the canvas; `await doc.pages[index].load(signal)` returns the immutable page commands and their resources. Containers should have a usable width. Multiple views can share one document.
102
-
103
- Call `attachment.dispose()`, `view.dispose()`, `canvasDocument.dispose()`, or `doc.dispose()` when done. Canvas disposal removes its views and releases its font registrations. Processing disposal also disposes subscribed Canvas documents and releases processing resources. Models remain inspectable; further painting is rejected. Pass `{ signal }` when opening. Opening copies byte inputs; a File/Blob is read once by either the browser or analysis entry point. Fetch URLs in host code with the authentication policy your application needs, then pass bytes.
104
-
105
- ## Show pages while processing
106
-
107
- Progress events expose a document source that the Canvas package can observe. Create the view once. Its subscription follows provisional page replacements and the completed layout.
108
-
109
- ```ts
110
- import { openDocument } from "@onodocs/sdk/browser";
111
- import { createDocument, type CanvasDocument } from "@onodocs/canvas";
112
-
113
- let canvasDocument: CanvasDocument | undefined;
114
- const doc = await openDocument(bytes, {
115
- signal,
116
- onProgress(progress) {
117
- canvasDocument ??= createDocument(progress.document, { container });
118
- status.textContent = progress.stage;
119
- },
120
- });
121
- ```
122
-
123
- Cancelling or failing processing disposes the source and its subscribed view. After success, dispose `doc` when the application closes the document. A Canvas document can be disposed earlier without closing the processor.
124
-
125
- ## Render on the server
126
-
127
- ```ts
128
- import { readFile, writeFile } from "node:fs/promises";
129
- import { createRenderer } from "@onodocs/sdk/server";
130
-
131
- const renderer = await createRenderer({ channel: "chrome" });
132
- try {
133
- const doc = await renderer.openDocument(await readFile("invoice.docx"));
134
- try {
135
- await writeFile("invoice.png", await doc.renderPage(0, { dpi: 144 }));
136
- await writeFile("invoice.pdf", await doc.pdf({ dpi: 150 }));
137
- } finally { await doc.dispose(); }
138
- } finally { await renderer.dispose(); }
139
- ```
140
-
141
- Install Chromium/Chrome separately; use `executablePath` for a deployment-managed binary or `channel: "chrome"`/`"msedge"` for an installed browser. The renderer reuses one browser, with a separate context per document. PNG and JPEG output use the browser engine's Canvas path. PDF contains rasterized pages at their document sizes, including mixed sizes. It does not provide searchable PDF text or tagged-PDF accessibility.
142
-
143
- Contexts make no outbound network requests. Supply fonts as data URLs through `fonts`, or install them on the rendering host. Browser consumers can also supply font URLs. Embedded document fonts use the engine's existing resource path. Use identical browser versions and fonts when reproducible output matters.
144
-
145
- Opening and output methods accept an AbortSignal. Cancelling an active operation closes that document context; open a new document to retry. Cancelling an operation still waiting in the queue leaves the active operation and document intact. Both server documents and renderers expose `disposed`. Disposing the renderer closes all remaining documents. A renderer can open multiple documents concurrently; output requests on one document are serialized. Applications control scheduling and process isolation.
146
-
147
- ## Process on the backend and render in the frontend
148
-
149
- The [complete backend viewer sample](https://github.com/onodocs/onodocs/tree/main/examples/backend-viewer) includes a Node.js server, working endpoints, a browser frontend and an authored DOCX. Install Node.js 22.18 or newer and Google Chrome, then run:
150
-
151
- ```sh
152
- git clone https://github.com/onodocs/onodocs.git
153
- cd onodocs/examples/backend-viewer
154
- npm install
155
- npm start
156
- ```
157
-
158
- Open http://127.0.0.1:5175. Use `npm start -- /path/to/document.docx` to open another server-side file. Stop the server with Ctrl+C. The [website guide](https://onodocs.com/developers/#combined) includes the complete server and frontend source.
159
-
160
- The server's `manifest({ format: "json" })` and `page(index, { format: "json" })` methods return response bodies ready to send with `Content-Type: application/json`. For a document at `/document`, serve its manifest there and individual pages at `/document/pages/:index`. The original methods without a format option still return typed objects for in-process use.
161
-
162
- ```js
163
- import { openDocument } from "@onodocs/canvas";
164
-
165
- const doc = await openDocument("/document", {
166
- container: document.querySelector("#pages"),
167
- viewOptions: { zoom: "fit-width" }
168
- });
169
- await doc.view.whenRendered();
170
- ```
171
-
172
- Canvas handles loading and decoding. It requests page content as needed. `signal` cancels requests; `request` accepts fetch options such as `headers` and `credentials` for both the manifest and page requests. Same-origin cookies work by default. Requests default to no-store, and non-successful HTTP responses reject with the status code. Dispose the Canvas document when closing or replacing the view.
173
-
174
- The application owns the endpoints and document authorization. Keep server documents open while their endpoints are in use, then dispose them. For updates, publish each completed snapshot at its own URL and reopen the frontend view. Use matching SDK and Canvas releases.
175
-
176
- Selection, copying, hit testing, zoom and attachments use the published geometry. Queries and updates stay on the backend. The frontend receives the visible text and resources needed to display the document. Supplied and embedded fonts travel with the pages; fonts resolved from the server's installed fonts must also be available on the client.
177
-
178
- ## Errors and application data
179
-
180
- Opening, updating, and exporting can reject. Handle errors at your application boundary and dispose browser/server resources in finally blocks. The analysis and browser entries export `PackageOperationError`, `SemanticOperationError`, and `QueryCardinalityError`; cancellation uses the signal's reason.
181
-
182
- Query elements are readonly in-memory models with parent references, source metadata, and bigint values. Project the fields your application needs instead of serializing the entire model:
183
-
184
- ```ts
185
- const paragraphs = doc.query.stories().where({ story: "body" }).paragraphs()
186
- .map(({ id, text, style }) => ({ id, text, style }));
187
- const json = JSON.stringify(paragraphs);
188
- ```
189
-
190
- The same projection supports AI context, search indexing, and backend messages without coupling your data format to engine internals.
191
-
192
- ## Low-level stages
193
-
194
- `doc.source` is the typed source model and `doc.semantic` contains completed document meaning. Element `source` locations refer into `doc.source.sourceFacts` and provenance. These identifiers belong to that loaded document, not arbitrary later revisions. The models describe supported DOCX content; they do not preserve every XML detail for round-trip writing.
195
-
196
- The main SDK also exports the core stage APIs: `parseDocumentSource`, `createSemanticDocument`, `parseDocumentForRendering`, font preparation, `createLayoutDocument`, and `createPageDisplayList`. Low-level `materializePageDisplayListToCanvas` is exported by `@onodocs/canvas`. Advanced consumers can stop at a stage or replay paint through their own adapter. Use canonical readonly contracts; do not mutate stage outputs. Layout remains the authority for all physical geometry.
197
-
198
- ### Extract a table as text
199
-
200
- ```ts
201
- const data = doc.query.tables()
202
- .where({ headers: ["Deliverable", "Owner", "Due"] })
203
- .one().textRows;
204
- ```
205
-
206
- `textRows` includes the first row and keeps cells grouped by their own table rows. The header filter matches the first row exactly, in order and case sensitively; it does not infer header formatting. Nested tables remain text inside their containing cells. Merged cells retain authored topology; the result is not padded into a rectangular grid. Use `data.slice(1)` to omit the header. The readonly matrix is a snapshot; query again after an update.
207
-
208
- ## Large documents
209
-
210
- Use the progress subscription shown above to display early pages while layout continues. Early pages are provisional: page counts, fields, notes and placement can change as formatting settles. Opening resolves with the complete document. Abort the loading signal to dispose the source and its subscribed views.
211
-
212
- Mounted views paint the visible neighborhood and release distant canvas pixels. Zoom and scrolling render pages on demand. Decoded images are acquired only for a page render and released afterward. Parsing and document layout still require memory proportional to document content; virtualization does not make the document model constant-memory. Auto-fit tables and document-wide fields can require later content before their output settles.
213
-
214
- Await page renders before reading or exporting the canvas:
215
-
216
- ```js
217
- await canvasDocument.pages[0].render(canvas, { dpi: 144, signal: controller.signal });
218
- const png = canvas.toDataURL("image/png");
219
- ```
1
+ # OnoDocs SDK
2
+
3
+ Try the [proposal studio](proposal.md) for a runnable business-data, Word-template, editing and DOCX/PDF export workflow.
4
+
5
+ See [document modes](document-modes.md) for viewing, editing, reviewing and form filling in one session.
6
+
7
+ For form authoring, controlled answer regions and completed DOCX plus structured answers, see [document forms](document-forms.md).
8
+
9
+ For application events, autosave, recovery, framework lifecycle and self-hosted HTTP, see [application integration](application-integration.md).
10
+
11
+ `@onodocs/sdk` processes Word documents. Its default entry performs analysis in Node or a browser; `/browser` adds layout and immutable page publication using browser font services; `/server` runs that processing in headless Chromium. `@onodocs/canvas` paints prepared pages and provides the viewer, selection and DOM attachments. The Canvas package runs independently of the parser, semantic engine and layout engine.
12
+
13
+ Supported Word inputs are DOCX, DOCM, DOTX and DOTM. Macro-enabled files are rendered without executing VBA. Older binary DOC/DOT files and RTF are not supported.
14
+
15
+ Install with `npm install @onodocs/sdk @onodocs/canvas`. Use matching versions of both packages. A backend can install only the SDK; a frontend receiving prepared pages can install only Canvas. Both packages include TypeScript declarations and this guide. Archives are also available from the [public releases](https://github.com/onodocs/onodocs/releases). Server rendering requires an installed Chromium-family browser. The server adapter includes the Canvas runtime it uses for image and PDF export.
16
+
17
+ Without a key, the SDK runs in free, non-production evaluation mode with full features and a watermark on rendered output. Pass `{ licenseKey }` to `openDocument` for a 30-day trial or commercial entitlement; this works in all three entry points. Verification is entirely offline. Commercial keys permit covered releases indefinitely, with renewal for later releases and support. Browser verification requires HTTPS or a secure local context. See the included `LICENSING.md` and `LICENSE` for details.
18
+
19
+ Optional [shared document collaboration](document-collaboration.md) adds multi-user editing and review with customer authentication, atomic persistence and recoverable local drafts. Ordinary browser sessions remain independent.
20
+
21
+ ## Inspect and query
22
+
23
+ ```ts
24
+ import { openDocument } from "@onodocs/sdk";
25
+
26
+ const doc = await openDocument(bytes);
27
+ const customer = doc.query.contentControls().where({ tag: "customer-name" }).one();
28
+ const cells = doc.query.tables().cells().all();
29
+ const matches = doc.query.findText("Invoice total").all();
30
+ const headers = doc.query.stories().where({ story: "header" }).paragraphs().all();
31
+ ```
32
+
33
+ TypeScript completion shows valid paths and filters. Select a kind, narrow with `where`, then choose how many results you expect:
34
+
35
+ | Method | Result |
36
+ | --- | --- |
37
+ | `one()` | One match; throws for zero or multiple matches |
38
+ | `optional()` | One match or undefined; throws for multiple matches |
39
+ | `first()` / `at(index)` | An explicitly chosen match or undefined |
40
+ | `all()` / iteration | Every match in document traversal order |
41
+ | `count()` | Number of matches |
42
+ | `map(project)` | Project matches into application data |
43
+
44
+ Selectors (`paragraphs`, `runs`, `tables`, `contentControls`, `bookmarks`, `stories`) search descendants. `rows` selects a table's own rows; `cells` selects a table or row's own cells. Nested tables are selected separately with `tables()`. Typed element collections also allow `table.rows[0].cells[1]`. `children()` and `descendants()` provide generic traversal. Scope an existing element with `doc.query.within(element)`. Nested or overlapping selections never duplicate the same element. A story represents an authored story; a reused header remains one semantic story with several physical occurrences.
45
+
46
+ `where({ text: "Total" })` matches exactly and case sensitively. Use `where({ text: { contains: "Total" } })`, `startsWith`, or `endsWith` for other matches. Multiple fields and chained filters combine with AND. Paragraphs support `style`; content controls support `tag` and `title`; bookmarks support `name`; stories support `story` (`body`, `header`, `footer`, `footnote`, `endnote`, `textbox`). Unknown keys fail instead of silently returning unexpected content. Use `.filter(element => ...)` for application predicates.
47
+
48
+ `doc.query.get(id)` resolves an element in the current snapshot. Every element has a `parent`; `doc.query.within(element).closest("cell")` finds its containing cell, including the input itself if it is a cell. Negative query indices count from the end.
49
+
50
+ `findText("...")` or `findText(/pattern/i)` searches contiguous text in selected paragraphs or runs, including text split across runs. It returns a paragraph and a half-open `[start, end)` range in JavaScript UTF-16 offsets. It does not join separate paragraphs. Regular expressions search each selected contiguous interval, preserve the caller's lastIndex, and omit zero-length matches. Queries preserve the selected revision view. Semantic text includes authored content; generated fields and physical appearances can differ after layout. Rasterized picture/chart text has no text query model.
51
+
52
+ Bookmarks expose their published start target through `bookmark.run`; they do not represent an enclosing bookmark text range. Symbols without a Unicode text representation use the object replacement character in query text.
53
+
54
+ ## Extract content for search and automation
55
+
56
+ `extractDocument` returns JSON-safe content and Markdown with links back to the source document. It is available from the default SDK entry and `/browser` and makes no network requests. Applications can pass its output to their own search or AI service.
57
+
58
+ ```ts
59
+ import { extractDocument } from "@onodocs/sdk";
60
+
61
+ const extraction = extractDocument(doc, { stories: ["body"] });
62
+ const json = JSON.stringify(extraction.content);
63
+ const markdown = extraction.markdown;
64
+ ```
65
+
66
+ The content tree retains stories, paragraphs, runs, tables, cells, content controls, hyperlink metadata and source references. Tables preserve authored header and merge information. Markdown uses pipe tables for simple tables and ordered cell passages for nested or merged tables. JSON preserves the full table structure. Paragraph headings use selected style IDs `Heading1` through `Heading9` by default; pass `headingStyles: { ReportTitle: 1, SectionTitle: 2 }` to replace that mapping. This does not infer headings from visual formatting. Omit `stories` to include all supported stories. Text has the same semantic and cached-field behavior as queries.
67
+
68
+ Each extracted element has a `source` string. In a browser document, resolve it into the existing geometry API to scroll to or highlight the original passage:
69
+
70
+ ```ts
71
+ const anchor = extraction.resolve(sourceFromSearchResult);
72
+ const fragments = doc.geometry.fragments(anchor);
73
+ ```
74
+
75
+ For partial passages, `extraction.reference(paragraphSource, start, end)` creates a source reference with UTF-16 text offsets. To apply a reviewed automation result, use the existing edit commands without a selection field:
76
+
77
+ ```ts
78
+ await extraction.edit(sourceFromReviewedResult, { kind: "replace", text: replacement });
79
+ const bytes = await doc.save();
80
+ ```
81
+
82
+ Editing accepts paragraph or text-range references and has the same supported-content restrictions as `doc.edit`. Formatting and authoring commands work through the same method. `extraction.selection(source)` returns a logical selection for application controls; `editor.select(selection)` followed by `editor.execute(command)` preserves the editor's undo history.
83
+
84
+ References belong to one extraction snapshot. After a committed change or disposal, `extraction.stale` is true and reference operations throw `StaleDocumentReferenceError`. Re-extract before applying another result. The check also runs inside the edit queue, so a concurrent edit cannot redirect a delayed automation result to shifted text. Failed changes and no-op text updates preserve references. Saving alone preserves them; reopening requires a new extraction. Source links are not persistent document identities, and content remains readable after its references become stale.
85
+
86
+ The [report review example](https://github.com/ionoy/onodocs/blob/main/docs/ai-report-demo.md) combines extraction, citations, reviewed edits and Word/PDF export. Its local sample reviewer can be replaced with the host's AI service.
87
+
88
+ ## Design and generate Word templates
89
+
90
+ Use `createTemplateEditor` from `@onodocs/sdk/editor` for visual field authoring and preview. `openTemplate` and `templateTag` from the default SDK and browser entries support fields, repeated content, conditions, images, formatted values, reusable sections and batch generation. See [Word templates and generation](templates.md) for the binding model, authoring commands and output workflow.
91
+
92
+ ## Replace text and fill templates
93
+
94
+ ```ts
95
+ const customer = doc.query.contentControls().where({ tag: "customer" }).one();
96
+ await doc.update({ target: customer, text: "Willow Design" });
97
+ ```
98
+
99
+ `update` accepts one update or a batch. Targets are runs, paragraphs, cells, content controls, their IDs, or search ranges. Element targets must contain ordinary text in one paragraph. Replacement text inherits the first selected run’s formatting; other selected runs become empty. Tabs and newlines are supported. Batches reject overlapping or missing targets, generated fields, and multi-paragraph replacements without changing the document.
100
+
101
+ Use search ranges for partial replacements while keeping surrounding text and formatting:
102
+
103
+ ```ts
104
+ await doc.update(doc.query.findText("{{customer}}").map(target => ({ target, text: "Willow Design" })));
105
+ ```
106
+
107
+ Batch range offsets refer to the original paragraph text. Empty ranges insert text; offsets must not split an emoji or other surrogate pair. Unchanged updates preserve existing formatting, snapshots, and attachments.
108
+
109
+ Browser updates rebuild layout and refresh mounted views. Query, semantic, layout, page and geometry values obtained earlier remain snapshots; read them again after updating. Application attachments are removed during refresh and can be reattached using fresh geometry. Updates run in submission order and accept `{ signal }` for cancellation. The original `source` is unchanged.
110
+
111
+ Node inspection exposes the same update method. For server rendering, inspect the input with the default entry point to select elements or ranges, then pass those updates to the server document before `renderPage`, `pdf`, or `save`. Exports reflect the updated text.
112
+
113
+ ## Save an edited Word document
114
+
115
+ `save()` returns DOCX bytes in Node, the browser, and server rendering. It refreshes supported fields, including after text edits or template generation. Use `save({ fields: "preserve" })` to retain cached field results; an unedited document then returns its original bytes. Text edits preserve surrounding run formatting and unedited package content. Saving and updating run in submission order. Both accept `{ signal }`; cancellation leaves the committed document available. Call `dispose()` when finished, including for Node documents, to release the preserved source package.
116
+
117
+ `await doc.fieldStatus()` reports which fields can be updated, which require layout, and which are locked, static or unsupported. Node inspection updates supported sequences, text bookmark references and document properties. Browser and server documents also calculate page numbers, page references and supported tables of contents. Unsupported fields keep their cached values. Saving reports these through `onFieldStatus`, or a console warning when no callback is provided. Signed or protected packages remain unchanged when their fields cannot be written. `save({ fields: "require-current" })` throws `DocumentFieldUpdateError` if a field cannot be refreshed. PDF export accepts the same field policy and diagnostic callback.
118
+
119
+ ```js
120
+ import { readFile, writeFile } from "node:fs/promises";
121
+ import { openDocument } from "@onodocs/sdk";
122
+
123
+ const doc = await openDocument(await readFile("template.docx"));
124
+ try {
125
+ await doc.update({
126
+ target: doc.query.contentControls().where({ tag: "customer" }).one(),
127
+ text: "Willow Design"
128
+ });
129
+ await writeFile("completed.docx", await doc.save());
130
+ } finally {
131
+ doc.dispose();
132
+ }
133
+ ```
134
+
135
+ In a browser, create a Blob from the returned bytes with type `application/vnd.openxmlformats-officedocument.wordprocessingml.document`. In a combined deployment, serve `await doc.save()` from an application-owned download endpoint with that content type and a `.docx` filename. The backend processing document owns saving; a Canvas-only frontend does not contain the original Word package.
136
+
137
+ The SDK package includes a complete command-line example. After installing it, run `node node_modules/@onodocs/sdk/examples/fill-template.mjs template.docx completed.docx customer "Willow Design"`. Set `ONODOCS_LICENSE_KEY` for trial or commercial use. The source workspace keeps the same example at `examples/fill-template.mjs`.
138
+
139
+ The initial editing profile covers ordinary text and unbound text or rich-text content controls within one paragraph, including supported tables, headers, footers and notes. Edits inside bound or content-locked controls, temporary or specialized controls, fields, revisions, alternate-content branches, and source runs with generated or unsupported content are rejected before publication. Plain-text controls that disallow multiple lines reject newlines. Filling an unbound text control clears its placeholder-display flag. Edits elsewhere preserve these constructs. `DocumentSaveError.reason` identifies source-writing restrictions; existing semantic target errors remain `RangeError`.
140
+
141
+ Signed packages, enforced document protection, macro-enabled documents and templates retain their editing restrictions. Ordinary text updates reject documents with change tracking enabled; use the review operations below to author tracked replacements and formatting. Saving refreshes supported fields under the selected field policy. Original authored formatting is preserved independently of rendering compatibility decisions.
142
+
143
+ Evaluation exports contain the full Word document without adding a Word watermark. The existing non-production evaluation terms apply; rendered evaluation output continues to include its watermark.
144
+
145
+ ## Edit formatting and paragraphs
146
+
147
+ `edit()` accepts a logical selection with paragraph IDs and UTF-16 offsets. It returns the selection in the resulting document. Use fresh queries after editing because structural operations rebuild the document and its IDs. The `source` getter then describes the new committed package.
148
+
149
+ ```js
150
+ const paragraph = doc.query.paragraphs().first();
151
+ let selection = {
152
+ start: { paragraphId: paragraph.id, offset: 0 },
153
+ end: { paragraphId: paragraph.id, offset: paragraph.text.length }
154
+ };
155
+ selection = await doc.edit({ kind: "format", selection, formatting: { bold: true, fontSize: 18 } });
156
+ selection = await doc.edit({ kind: "replace", selection, text: "First paragraph\nSecond paragraph" });
157
+ await doc.edit({ kind: "align", selection, alignment: "center" });
158
+ const bytes = await doc.save();
159
+ ```
160
+
161
+ Character formatting supports `bold`, `italic`, `underline`, `strike`, `fontSize` in points, `fontFamily`, and `color` as `#RRGGBB`. Values change direct formatting on the selected text while preserving the other authored properties. `run.formatting` exposes the common effective character properties for controls; `paragraph.alignment` exposes left, center, right or justified alignment. These convenience properties describe ordinary Latin text; they do not replace the full semantic typography model for script-specific inspection.
162
+
163
+ A replacement with a collapsed selection inserts text. Newlines split paragraphs; set `paragraphBreaks: false` for line breaks within a paragraph. Replacing a selection across adjacent paragraphs joins their surviving prefix and suffix. Empty text deletes the selection. New text inherits the starting run's formatting unless the operation includes `formatting`. Split paragraphs inherit the starting paragraph's properties. The first paragraph's properties win when joining.
164
+
165
+ Rich editing supports ordinary text paragraphs, including empty paragraphs and paragraphs in table cells and secondary stories. Text replacement within one paragraph and selected character formatting preserve surrounding fields, comments, bookmarks, drawings and unbound controls. Selected generated fields, locked or bound controls and tracked runs reject atomically. Structural operations such as splitting or joining still reject paragraphs containing these structures. Splitting or joining section boundaries and joining across cells or stories also reject. Such content remains preserved elsewhere in the document.
166
+
167
+ The [Word interoperability evidence](../spec/features/word-interoperability.md) covers reviewed agreements, completed forms and generated invoices through actual Word edit/save/reopen journeys. It does not imply compatibility with every Word feature. For paragraph styles without primary names, rendering follows LibreOffice recovery and can omit direct paragraph formatting that Word retains; saved unedited source preserves those properties.
168
+
169
+ Node and browser documents expose `edit`; server documents forward the same operation. Parsing, semantic construction and browser layout finish before publication, so cancellation or failure leaves the previous document usable. This first implementation rebuilds the document for structural edits. Browser resources stay with the document lifetime, and the sample keeps undo history as document snapshots. Large documents need further editing-performance work before this can serve as a general-purpose editor.
170
+
171
+ ## Embed the Word editor
172
+
173
+ ```js
174
+ import { createEditor } from "@onodocs/sdk/editor";
175
+
176
+ const editor = createEditor({ container: document.querySelector("#editor") });
177
+ await editor.newDocument();
178
+ const bytes = await editor.save();
179
+ editor.dispose();
180
+ ```
181
+
182
+ The component includes file opening, a new-document command, formatting, imported paragraph styles, lists, regular tables, paragraph images, links, page settings, default headers and footers, find/replace, clipboard operations, undo/redo and Word/PDF downloads. It runs in the browser. `examples/browser-editor/` is the standalone application; run `npm run example:editor` from the source workspace.
183
+
184
+ Customize the toolbar by passing an ordered list of control names, or `false` to supply your own controls. Available names are `open`, `new`, `save`, `pdf`, `undo`, `redo`, `bold`, `italic`, `underline`, `strike`, `font`, `size`, `color`, `alignment`, `style`, `list`, `table`, `tableTools`, `link`, `image`, `imageTools`, `page`, `header`, `footer`, `find` and `replace`. `fonts` supplies font-family choices; `styles` supplies `{ id, label }` choices using existing Word style IDs. `commands` adds `{ id, label, execute(editor) }` actions. `onChange` receives the editor after committed changes; `onError` reports queued operation failures. Use `document` for browser SDK options such as licensing and font sources, and `readOnly: true` to disable authoring.
185
+
186
+ The editor supports native text selection and IME input. F6 switches between the document and toolbar, Escape focuses the toolbar, and Shift+Tab leaves the editing area. On touch screens, tap to place the caret, hold to select a word, and drag the selection endpoints. Narrow toolbars scroll horizontally. Mounted pages expose text and headings to browser accessibility tools, including pages whose canvases have been evicted.
187
+
188
+ Each instance has isolated styles and history. Customize `--onodocs-color`, `--onodocs-background`, `--onodocs-toolbar-background`, `--onodocs-accent` and `--onodocs-selection` on the host container. The `toolbar`, `pages` and `status` shadow parts are available for host styling. Call `dispose()` when removing an editor.
189
+
190
+ For application controls, use `editor.execute(command)` without a selection field, `editor.select(selection)`, `editor.find(text)`, `editor.replaceAll(text, replacement)`, `editor.undo()`, `editor.redo()`, `editor.save()` and `editor.pdf()`. The current browser document and logical selection are available as `editor.document` and `editor.selection`. Direct document mutations bypass component history; use `execute` for edits that need undo.
191
+
192
+ The additional SDK `doc.edit` commands use the same selection contract:
193
+
194
+ | Kind | Arguments | Behavior |
195
+ | --- | --- | --- |
196
+ | `style` | `style: string \| null` | Apply a Word paragraph style ID or remove the direct style |
197
+ | `link` | `target: string \| null` | Link selected text to HTTP, HTTPS, mailto or `#bookmark`; null removes links |
198
+ | `list` | `list: "bullet" \| "decimal" \| null`, optional `level` 0-8 | Set paragraph numbering or remove it |
199
+ | `insertTable` | `rows`, `columns` | Insert a regular table after the selected paragraph |
200
+ | `table` | `action` | `insertRow`, `deleteRow`, `insertColumn`, `deleteColumn` or `delete` |
201
+ | `image` | `bytes`, `mediaType`, `width`, `height`, optional `description` | Insert a PNG/JPEG picture paragraph after the selected paragraph; dimensions are points |
202
+ | `page` | optional `width`, `height`, `margins: { top, right, bottom, left }` | Change the selected section using point dimensions |
203
+ | `header`, `footer` | `text` | Set default story text for the selected section |
204
+
205
+ The authoring profile supports regular unmerged tables and paragraph-image insertion. Select inline images with `doc.query.images()` and use `imageProperties` with `imageId`, optional point dimensions `width`/`height`, or `remove: true`. The component offers these commands through Image properties after clicking a picture. Floating-image manipulation and merged-table restructuring are not included. Configure review to author tracked text replacements and formatting. Clipboard copy, cut and paste preserve common text formatting, links, paragraph alignment and nested bullet/decimal lists. Tables and image alternative text paste as text, with a conversion message. External stylesheets and custom list formats are not imported. Plain-text clipboard content remains supported. History uses full document snapshots. Structural edits reimport and lay out the candidate document; ordinary typing reuses unchanged semantic data and font measurements. Run `npm run validate:editor` for the component and standalone browser journeys.
206
+
207
+ ## Render and attach ordinary DOM
208
+
209
+ Mounted views include read-only text selection and plain-text copying. Drag or Shift-click to select text, extend with Shift and the arrow/Home/End keys, select all with Ctrl/Command+A, and copy with Ctrl/Command+C or the Copy context menu. Selection spans pages and survives scrolling and zoom, including pages whose canvases have been released. Document updates clear the previous selection. Attached inputs retain their normal editing and clipboard behavior.
210
+
211
+ ```ts
212
+ import { openDocument } from "@onodocs/sdk/browser";
213
+ import { createDocument } from "@onodocs/canvas";
214
+
215
+ const doc = await openDocument(file);
216
+ const canvasDocument = createDocument(doc);
217
+ const view = canvasDocument.mount(container);
218
+ const field = doc.query.contentControls().where({ tag: "customer-name" }).one();
219
+ const input = document.createElement("input");
220
+ input.name = "customerName";
221
+ const attachment = view.attach(input, { anchor: doc.geometry.fragments(field)[0] });
222
+ ```
223
+
224
+ The example requires a tagged control that produces one visible fragment. Call `doc.geometry.fragments(field)` for content spanning lines, pages, or repeated stories, and attach to an explicitly chosen fragment. Fragments retain transforms and clipping. Table/cell anchors use their final physical rectangles. `offset` and `size` optionally adjust attachment placement in page units. Values, validation, submission, focus, and accessibility labels belong to the application. Attaching an element moves it into the view. Detachment removes it and restores its original inline style. Reattaching the same HTML element disposes its previous attachment; old handles become inert. An attachment exposes its disposed state.
225
+
226
+ Pages are anchors too. Placement avoids manual rectangle arithmetic:
227
+
228
+ ```js
229
+ view.attach(checkbox, { anchor: doc.geometry.fragments(doc.query.findText("Client approval:").one())[0], placement: "outside-right", gap: 80, size: { width: 420, height: 420 } });
230
+ view.attach(submitButton, { anchor: doc.pages[0], placement: "inside-bottom-right", inset: 480, size: { width: 2400, height: 600 } });
231
+ ```
232
+
233
+ Inside placements use `inside-{top|center|bottom}-{left|center|right}` with optional `inset`. Outside placements use `outside-{top|right|bottom|left}` with optional `gap` and center the element along the selected edge. The default is `overlay`. All sizes, spacing and offsets use twips (1/1440 inch). Without a size, the element fills the anchor, reduced by the inset for inside placement. Offsets apply after placement. Content anchors retain their transforms and clips; page anchors use the full page. Attachments do not reflow text, so reserve room in the document.
234
+
235
+ Adjacent styled runs on one line share a fragment. Fragment bounds enclose the unclipped geometry; the supplied clips determine which portions are visible.
236
+
237
+ `view.setZoom("fit-width")` tracks container width. Numeric zoom uses 96 CSS pixels per inch at `1`. `view.toClient(pageIndex, point)` and `view.fromClient({ x: event.clientX, y: event.clientY })` translate coordinates. `view.hitTest(clientPoint)` returns `{ elementId, fragment, caret? }`. Resolve `elementId` with `doc.query.get(elementId)` when the processing SDK is available. Text hits resolve to a run or paragraph; table whitespace resolves to a cell or table. `caret` contains `{ paragraphId, offset, point }`, with the same UTF-16 offset convention as search ranges. Generated field text omits a caret when no source position can be represented. Raster drawings without a semantic hit identity and blank page space return no element.
238
+
239
+ `view.scrollTo(anchor, { block: "center" })` reveals a page or a physical fragment. Resolve elements and text ranges through `doc.geometry.fragments(anchor)` first, then choose the occurrence to reveal.
240
+
241
+ A run's optional `link` is `{ kind: "external", target, tooltip }` or `{ kind: "bookmark", target }`. Use bookmark queries and scrollTo for internal links. The host application decides whether and how to open external destinations.
242
+
243
+ All page and query indices are **zero based**. Page geometry uses **twips: 1440 units per inch**. `await canvasDocument.pages[index].render(canvas, { dpi: 144 })` prepares the page resources and paints the canvas; `await doc.pages[index].load(signal)` returns the immutable page commands and their resources. Containers should have a usable width. Multiple views can share one document.
244
+
245
+ Call `attachment.dispose()`, `view.dispose()`, `canvasDocument.dispose()`, or `doc.dispose()` when done. Canvas disposal removes its views and releases its font registrations. Processing disposal also disposes subscribed Canvas documents and releases processing resources. Models remain inspectable; further painting is rejected. Pass `{ signal }` when opening. Opening copies byte inputs; a File/Blob is read once by either the browser or analysis entry point. Fetch URLs in host code with the authentication policy your application needs, then pass bytes.
246
+
247
+ ## Show pages while processing
248
+
249
+ Progress events expose a document source that the Canvas package can observe. Create the view once. Its subscription follows provisional page replacements and the completed layout.
250
+
251
+ ```ts
252
+ import { openDocument } from "@onodocs/sdk/browser";
253
+ import { createDocument, type CanvasDocument } from "@onodocs/canvas";
254
+
255
+ let canvasDocument: CanvasDocument | undefined;
256
+ const doc = await openDocument(bytes, {
257
+ signal,
258
+ onProgress(progress) {
259
+ canvasDocument ??= createDocument(progress.document, { container });
260
+ status.textContent = progress.stage;
261
+ },
262
+ });
263
+ ```
264
+
265
+ Cancelling or failing processing disposes the source and its subscribed view. After success, dispose `doc` when the application closes the document. A Canvas document can be disposed earlier without closing the processor.
266
+
267
+ ## Render on the server
268
+
269
+ ```ts
270
+ import { readFile, writeFile } from "node:fs/promises";
271
+ import { createRenderer } from "@onodocs/sdk/server";
272
+
273
+ const renderer = await createRenderer({ channel: "chrome" });
274
+ try {
275
+ const doc = await renderer.openDocument(await readFile("invoice.docx"));
276
+ try {
277
+ await writeFile("invoice.png", await doc.renderPage(0, { dpi: 144 }));
278
+ await writeFile("invoice.pdf", await doc.pdf({ dpi: 150 }));
279
+ } finally { await doc.dispose(); }
280
+ } finally { await renderer.dispose(); }
281
+ ```
282
+
283
+ Install Chromium/Chrome separately; use `executablePath` for a deployment-managed binary or `channel: "chrome"`/`"msedge"` for an installed browser. The renderer reuses one browser, with a separate context per document. PNG and JPEG output use the browser engine's Canvas path. PDF preserves document page sizes, including mixed sizes, and includes searchable/selectable Unicode text and external and internal links. Supported solid graphics remain vector content. Native text and complex effects use cropped lossless browser-rendered layers at the requested DPI. The shared browser/server exporter includes tagged headings, lists, tables with merged-cell spans, authored figure descriptions and links. PDF/UA and PDF/A conformance are not claimed.
284
+
285
+ Contexts make no outbound network requests. Supply fonts as data URLs through `fonts`, or install them on the rendering host. Browser consumers can also supply font URLs. Embedded document fonts use the engine's existing resource path. Use identical browser versions and fonts when reproducible output matters.
286
+
287
+ Opening and output methods accept an AbortSignal. Cancelling an active operation closes that document context; open a new document to retry. Cancelling an operation still waiting in the queue leaves the active operation and document intact. Both server documents and renderers expose `disposed`. Disposing the renderer closes all remaining documents. A renderer can open multiple documents concurrently; output requests on one document are serialized. Applications control scheduling and process isolation.
288
+
289
+ ## Process on the backend and render in the frontend
290
+
291
+ The [complete backend viewer sample](https://github.com/onodocs/onodocs/tree/main/examples/backend-viewer) includes a Node.js server, working endpoints, a browser frontend and an authored DOCX. Install Node.js 22.18 or newer and Google Chrome, then run:
292
+
293
+ ```sh
294
+ git clone https://github.com/onodocs/onodocs.git
295
+ cd onodocs/examples/backend-viewer
296
+ npm install
297
+ npm start
298
+ ```
299
+
300
+ Open http://127.0.0.1:5175. Use `npm start -- /path/to/document.docx` to open another server-side file. Stop the server with Ctrl+C. The [website guide](https://onodocs.com/developers/#combined) includes the complete server and frontend source.
301
+
302
+ The server's `manifest({ format: "json" })` and `page(index, { format: "json" })` methods return response bodies ready to send with `Content-Type: application/json`. For a document at `/document`, serve its manifest there and individual pages at `/document/pages/:index`. The original methods without a format option still return typed objects for in-process use.
303
+
304
+ ```js
305
+ import { openDocument } from "@onodocs/canvas";
306
+
307
+ const doc = await openDocument("/document", {
308
+ container: document.querySelector("#pages"),
309
+ viewOptions: { zoom: "fit-width" }
310
+ });
311
+ await doc.view.whenRendered();
312
+ ```
313
+
314
+ Canvas handles loading and decoding. It requests page content as needed. `signal` cancels requests; `request` accepts fetch options such as `headers` and `credentials` for both the manifest and page requests. Same-origin cookies work by default. Requests default to no-store, and non-successful HTTP responses reject with the status code. Dispose the Canvas document when closing or replacing the view.
315
+
316
+ The application owns the endpoints and document authorization. Keep server documents open while their endpoints are in use, then dispose them. For updates, publish each completed snapshot at its own URL and reopen the frontend view. Use matching SDK and Canvas releases.
317
+
318
+ Selection, copying, hit testing, zoom and attachments use the published geometry. Queries and updates stay on the backend. The frontend receives the visible text and resources needed to display the document. Supplied and embedded fonts travel with the pages; fonts resolved from the server's installed fonts must also be available on the client.
319
+
320
+ ## Errors and application data
321
+
322
+ Opening, updating, and exporting can reject. Handle errors at your application boundary and dispose browser/server resources in finally blocks. The analysis and browser entries export `PackageOperationError`, `SemanticOperationError`, and `QueryCardinalityError`; cancellation uses the signal's reason.
323
+
324
+ Query elements are readonly in-memory models with parent references, source metadata, and bigint values. Project the fields your application needs instead of serializing the entire model:
325
+
326
+ ```ts
327
+ const paragraphs = doc.query.stories().where({ story: "body" }).paragraphs()
328
+ .map(({ id, text, style }) => ({ id, text, style }));
329
+ const json = JSON.stringify(paragraphs);
330
+ ```
331
+
332
+ The same projection supports AI context, search indexing, and backend messages without coupling your data format to engine internals.
333
+
334
+ ## Low-level stages
335
+
336
+ `doc.source` is the typed source model and `doc.semantic` contains completed document meaning. Element `source` locations refer into `doc.source.sourceFacts` and provenance. These identifiers belong to that loaded document, not arbitrary later revisions. The models describe supported DOCX content; they do not preserve every XML detail for round-trip writing.
337
+
338
+ The main SDK also exports the core stage APIs: `parseDocumentSource`, `createSemanticDocument`, `parseDocumentForRendering`, font preparation, `createLayoutDocument`, and `createPageDisplayList`. Low-level `materializePageDisplayListToCanvas` is exported by `@onodocs/canvas`. Advanced consumers can stop at a stage or replay paint through their own adapter. Use canonical readonly contracts; do not mutate stage outputs. Layout remains the authority for all physical geometry.
339
+
340
+ ### Extract a table as text
341
+
342
+ ```ts
343
+ const data = doc.query.tables()
344
+ .where({ headers: ["Deliverable", "Owner", "Due"] })
345
+ .one().textRows;
346
+ ```
347
+
348
+ `textRows` includes the first row and keeps cells grouped by their own table rows. The header filter matches the first row exactly, in order and case sensitively; it does not infer header formatting. Nested tables remain text inside their containing cells. Merged cells retain authored topology; the result is not padded into a rectangular grid. Use `data.slice(1)` to omit the header. The readonly matrix is a snapshot; query again after an update.
349
+
350
+ ## Large documents
351
+
352
+ Use the progress subscription shown above to display early pages while layout continues. Early pages are provisional: page counts, fields, notes and placement can change as formatting settles. Opening resolves with the complete document. Abort the loading signal to dispose the source and its subscribed views.
353
+
354
+ Mounted views paint the visible neighborhood and release distant canvas pixels. Zoom and scrolling render pages on demand. Decoded images are acquired only for a page render and released afterward. Parsing and document layout still require memory proportional to document content; virtualization does not make the document model constant-memory. Auto-fit tables and document-wide fields can require later content before their output settles.
355
+
356
+ Await page renders before reading or exporting the canvas:
357
+
358
+ ```js
359
+ await canvasDocument.pages[0].render(canvas, { dpi: 144, signal: controller.signal });
360
+ const png = canvas.toDataURL("image/png");
361
+ ```
362
+
363
+ For documents that omit font declarations, browser and server opening options accept `defaultFont` and `defaultFontSize` (points). They default to Calibri and 11pt; authored font choices and sizes take precedence.
364
+
365
+ ## Review and saved versions
366
+
367
+ Node and browser documents expose the same review operations. Supply your application's user identity and timestamp for authored comments and changes:
368
+
369
+ ```js
370
+ const identity = { author: currentUser.name, initials: currentUser.initials, date: new Date().toISOString() };
371
+ const paragraph = doc.query.paragraphs().first();
372
+ const selection = { start: { paragraphId: paragraph.id, offset: 0 }, end: { paragraphId: paragraph.id, offset: 6 } };
373
+ await doc.review({ kind: "comment", selection, text: "Please verify this wording.", identity });
374
+ const thread = (await doc.readReview()).comments.at(-1);
375
+ await doc.review({ kind: "reply", id: thread.id, text: "Confirmed.", identity });
376
+ await doc.review({ kind: "resolveComment", id: thread.id, resolved: true });
377
+ await doc.review({ kind: "trackedReplace", selection: thread.selection, text: "Revised", identity });
378
+ await doc.review({ kind: "revision", accept: true });
379
+ const bytes = await doc.save();
380
+ ```
381
+
382
+ Open with `revisionView: "accepted"` when editing current text. `readReview()` includes hidden deleted changes, which may have no visible selection. Pass revision `ids` to accept or reject individual changes; omit them for all changes. `trackedFormat` records selected run formatting with a rejectable property snapshot. `trackChanges` enables or disables Word's tracking setting. Review mutations support `signal` and `ifUnchangedSince` like ordinary edits. Text replacement stays within a paragraph; comments can span paragraphs in one story. Protected documents retain their write restrictions.
383
+
384
+ Enable the editor panel with `review: { identity: () => identity, history }`. The optional `history` object implements `list(signal)`, `load(id, signal)` and `save(version, bytes, signal)`. Your application stores the complete Word bytes and version metadata. The SDK adds no persistence service or user account system. Restoring a version is undoable and raises the editor's change event, so application autosave can persist it. `compareDocuments(savedBytes, doc)` returns paragraph text and formatting differences without modifying the document.
385
+
386
+ ## Searchable PDF in the browser
387
+
388
+ ```js
389
+ import { openDocument } from "@onodocs/sdk/browser";
390
+
391
+ const doc = await openDocument(await file.arrayBuffer());
392
+ try {
393
+ const bytes = await doc.pdf({ dpi: 150, title: "Quarterly report", language: "en-US" });
394
+ const url = URL.createObjectURL(new Blob([bytes], { type: "application/pdf" }));
395
+ const link = document.createElement("a");
396
+ link.href = url;
397
+ link.download = "document.pdf";
398
+ link.click();
399
+ setTimeout(() => URL.revokeObjectURL(url), 1000);
400
+ } finally {
401
+ doc.dispose();
402
+ }
403
+ ```
404
+
405
+ PDF export runs locally and uses the same implementation as server `doc.pdf()`. It waits for earlier edits to finish. Pass `signal` to cancel an export without closing the document. The PDF writer loads on demand; deploy the browser entry's generated chunks alongside it. Text extraction order for bidirectional paragraphs and selection across complex clusters depend on the PDF reader. Solid graphics use vector PDF paths where supported. Native text and complex effects retain browser rendering in cropped lossless image layers, including evaluation marks. Higher DPI increases the sharpness, file size and temporary memory use of those layers; compressed output remains in memory until export finishes.
406
+
407
+ The PDF includes tagged headings, paragraphs, nested lists, tables, header cells, merged-cell row/column spans, authored figure descriptions and links, plus heading and bookmark navigation. Set `title` and `language` for reader metadata. DrawingML descriptions use `wp:docPr/@descr`, falling back to `title`; VML uses authored `alt` text. Undescribed graphics remain artifacts. Reader support varies. Firefox PDF.js exposes named images, list nesting and table spans in its accessibility tree. This does not claim PDF/UA conformance or validation with every screen reader.