@onodocs/canvas 0.2.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,229 +1,281 @@
1
- # OnoDocs SDK
2
-
3
- `@onodocs/sdk` processes Word documents. Its default entry performs analysis in Node or a browser; `/browser` adds layout and immutable page publication using browser font services; `/server` runs that processing in headless Chromium. `@onodocs/canvas` paints prepared pages and provides the viewer, selection and DOM attachments. The Canvas package runs independently of the parser, semantic engine and layout engine.
4
-
5
- Supported Word inputs are DOCX, DOCM, DOTX and DOTM. Macro-enabled files are rendered without executing VBA. Older binary DOC/DOT files and RTF are not supported.
6
-
7
- Install with `npm install @onodocs/sdk @onodocs/canvas`. Use matching versions of both packages. A backend can install only the SDK; a frontend receiving prepared pages can install only Canvas. Both packages include TypeScript declarations and this guide. Archives are also available from the [public releases](https://github.com/onodocs/onodocs/releases). Server rendering requires an installed Chromium-family browser. The server adapter includes the Canvas runtime it uses for image and PDF export.
8
-
9
- Without a key, the SDK runs in free, non-production evaluation mode with full features and a watermark on rendered output. Pass `{ licenseKey }` to `openDocument` for a 30-day trial or commercial entitlement; this works in all three entry points. Verification is entirely offline. Commercial keys permit covered releases indefinitely, with renewal for later releases and support. Browser verification requires HTTPS or a secure local context. See the included `LICENSING.md` and `LICENSE` for details.
10
-
11
- ## Inspect and query
12
-
13
- ```ts
14
- import { openDocument } from "@onodocs/sdk";
15
-
16
- const doc = await openDocument(bytes);
17
- const customer = doc.query.contentControls().where({ tag: "customer-name" }).one();
18
- const cells = doc.query.tables().cells().all();
19
- const matches = doc.query.findText("Invoice total").all();
20
- const headers = doc.query.stories().where({ story: "header" }).paragraphs().all();
21
- ```
22
-
23
- TypeScript completion shows valid paths and filters. Select a kind, narrow with `where`, then choose how many results you expect:
24
-
25
- | Method | Result |
26
- | --- | --- |
27
- | `one()` | One match; throws for zero or multiple matches |
28
- | `optional()` | One match or undefined; throws for multiple matches |
29
- | `first()` / `at(index)` | An explicitly chosen match or undefined |
30
- | `all()` / iteration | Every match in document traversal order |
31
- | `count()` | Number of matches |
32
- | `map(project)` | Project matches into application data |
33
-
34
- Selectors (`paragraphs`, `runs`, `tables`, `contentControls`, `bookmarks`, `stories`) search descendants. `rows` selects a table's own rows; `cells` selects a table or row's own cells. Nested tables are selected separately with `tables()`. Typed element collections also allow `table.rows[0].cells[1]`. `children()` and `descendants()` provide generic traversal. Scope an existing element with `doc.query.within(element)`. Nested or overlapping selections never duplicate the same element. A story represents an authored story; a reused header remains one semantic story with several physical occurrences.
35
-
36
- `where({ text: "Total" })` matches exactly and case sensitively. Use `where({ text: { contains: "Total" } })`, `startsWith`, or `endsWith` for other matches. Multiple fields and chained filters combine with AND. Paragraphs support `style`; content controls support `tag` and `title`; bookmarks support `name`; stories support `story` (`body`, `header`, `footer`, `footnote`, `endnote`, `textbox`). Unknown keys fail instead of silently returning unexpected content. Use `.filter(element => ...)` for application predicates.
37
-
38
- `doc.query.get(id)` resolves an element in the current snapshot. Every element has a `parent`; `doc.query.within(element).closest("cell")` finds its containing cell, including the input itself if it is a cell. Negative query indices count from the end.
39
-
40
- `findText("...")` or `findText(/pattern/i)` searches contiguous text in selected paragraphs or runs, including text split across runs. It returns a paragraph and a half-open `[start, end)` range in JavaScript UTF-16 offsets. It does not join separate paragraphs. Regular expressions search each selected contiguous interval, preserve the caller's lastIndex, and omit zero-length matches. Queries preserve the selected revision view. Semantic text includes authored content; generated fields and physical appearances can differ after layout. Rasterized picture/chart text has no text query model.
41
-
42
- Bookmarks expose their published start target through `bookmark.run`; they do not represent an enclosing bookmark text range. Symbols without a Unicode text representation use the object replacement character in query text.
43
-
44
- ## Replace text and fill templates
45
-
46
- ```ts
47
- const customer = doc.query.contentControls().where({ tag: "customer" }).one();
48
- await doc.update({ target: customer, text: "Willow Design" });
49
- ```
50
-
51
- `update` accepts one update or a batch. Targets are runs, paragraphs, cells, content controls, their IDs, or search ranges. Element targets must contain ordinary text in one paragraph. Replacement text inherits the first selected run’s formatting; other selected runs become empty. Tabs and newlines are supported. Batches reject overlapping or missing targets, generated fields, and multi-paragraph replacements without changing the document.
52
-
53
- Use search ranges for partial replacements while keeping surrounding text and formatting:
54
-
55
- ```ts
56
- await doc.update(doc.query.findText("{{customer}}").map(target => ({ target, text: "Willow Design" })));
57
- ```
58
-
59
- Batch range offsets refer to the original paragraph text. Empty ranges insert text; offsets must not split an emoji or other surrogate pair. Unchanged updates preserve existing formatting, snapshots, and attachments.
60
-
61
- Browser updates rebuild layout and refresh mounted views. Query, semantic, layout, page and geometry values obtained earlier remain snapshots; read them again after updating. Application attachments are removed during refresh and can be reattached using fresh geometry. Updates run in submission order and accept `{ signal }` for cancellation. The original `source` is unchanged.
62
-
63
- Node inspection exposes the same update method. For server rendering, inspect the input with the default entry point to select elements or ranges, then pass those updates to the server document before `renderPage` or `pdf`. Exports reflect the updated text. Saving an updated DOCX, structural editing, and field recalculation are not supported.
64
-
65
- ## Render and attach ordinary DOM
66
-
67
- Mounted views include read-only text selection and plain-text copying. Drag or Shift-click to select text, extend with Shift and the arrow/Home/End keys, select all with Ctrl/Command+A, and copy with Ctrl/Command+C or the Copy context menu. Selection spans pages and survives scrolling and zoom, including pages whose canvases have been released. Document updates clear the previous selection. Attached inputs retain their normal editing and clipboard behavior.
68
-
69
- ```ts
70
- import { openDocument } from "@onodocs/sdk/browser";
71
- import { createDocument } from "@onodocs/canvas";
72
-
73
- const doc = await openDocument(file);
74
- const canvasDocument = createDocument(doc);
75
- const view = canvasDocument.mount(container);
76
- const field = doc.query.contentControls().where({ tag: "customer-name" }).one();
77
- const input = document.createElement("input");
78
- input.name = "customerName";
79
- const attachment = view.attach(input, { anchor: doc.geometry.fragments(field)[0] });
80
- ```
81
-
82
- The example requires a tagged control that produces one visible fragment. Call `doc.geometry.fragments(field)` for content spanning lines, pages, or repeated stories, and attach to an explicitly chosen fragment. Fragments retain transforms and clipping. Table/cell anchors use their final physical rectangles. `offset` and `size` optionally adjust attachment placement in page units. Values, validation, submission, focus, and accessibility labels belong to the application. Attaching an element moves it into the view. Detachment removes it and restores its original inline style. Reattaching the same HTML element disposes its previous attachment; old handles become inert. An attachment exposes its disposed state.
83
-
84
- Pages are anchors too. Placement avoids manual rectangle arithmetic:
1
+ # OnoDocs SDK
2
+
3
+ `@onodocs/sdk` processes Word documents. Its default entry performs analysis in Node or a browser; `/browser` adds layout and immutable page publication using browser font services; `/server` runs that processing in headless Chromium. `@onodocs/canvas` paints prepared pages and provides the viewer, selection and DOM attachments. The Canvas package runs independently of the parser, semantic engine and layout engine.
4
+
5
+ Supported Word inputs are DOCX, DOCM, DOTX and DOTM. Macro-enabled files are rendered without executing VBA. Older binary DOC/DOT files and RTF are not supported.
6
+
7
+ Install with `npm install @onodocs/sdk @onodocs/canvas`. Use matching versions of both packages. A backend can install only the SDK; a frontend receiving prepared pages can install only Canvas. Both packages include TypeScript declarations and this guide. Archives are also available from the [public releases](https://github.com/onodocs/onodocs/releases). Server rendering requires an installed Chromium-family browser. The server adapter includes the Canvas runtime it uses for image and PDF export.
8
+
9
+ Without a key, the SDK runs in free, non-production evaluation mode with full features and a watermark on rendered output. Pass `{ licenseKey }` to `openDocument` for a 30-day trial or commercial entitlement; this works in all three entry points. Verification is entirely offline. Commercial keys permit covered releases indefinitely, with renewal for later releases and support. Browser verification requires HTTPS or a secure local context. See the included `LICENSING.md` and `LICENSE` for details.
10
+
11
+ ## Inspect and query
12
+
13
+ ```ts
14
+ import { openDocument } from "@onodocs/sdk";
15
+
16
+ const doc = await openDocument(bytes);
17
+ const customer = doc.query.contentControls().where({ tag: "customer-name" }).one();
18
+ const cells = doc.query.tables().cells().all();
19
+ const matches = doc.query.findText("Invoice total").all();
20
+ const headers = doc.query.stories().where({ story: "header" }).paragraphs().all();
21
+ ```
22
+
23
+ TypeScript completion shows valid paths and filters. Select a kind, narrow with `where`, then choose how many results you expect:
24
+
25
+ | Method | Result |
26
+ | --- | --- |
27
+ | `one()` | One match; throws for zero or multiple matches |
28
+ | `optional()` | One match or undefined; throws for multiple matches |
29
+ | `first()` / `at(index)` | An explicitly chosen match or undefined |
30
+ | `all()` / iteration | Every match in document traversal order |
31
+ | `count()` | Number of matches |
32
+ | `map(project)` | Project matches into application data |
33
+
34
+ Selectors (`paragraphs`, `runs`, `tables`, `contentControls`, `bookmarks`, `stories`) search descendants. `rows` selects a table's own rows; `cells` selects a table or row's own cells. Nested tables are selected separately with `tables()`. Typed element collections also allow `table.rows[0].cells[1]`. `children()` and `descendants()` provide generic traversal. Scope an existing element with `doc.query.within(element)`. Nested or overlapping selections never duplicate the same element. A story represents an authored story; a reused header remains one semantic story with several physical occurrences.
35
+
36
+ `where({ text: "Total" })` matches exactly and case sensitively. Use `where({ text: { contains: "Total" } })`, `startsWith`, or `endsWith` for other matches. Multiple fields and chained filters combine with AND. Paragraphs support `style`; content controls support `tag` and `title`; bookmarks support `name`; stories support `story` (`body`, `header`, `footer`, `footnote`, `endnote`, `textbox`). Unknown keys fail instead of silently returning unexpected content. Use `.filter(element => ...)` for application predicates.
37
+
38
+ `doc.query.get(id)` resolves an element in the current snapshot. Every element has a `parent`; `doc.query.within(element).closest("cell")` finds its containing cell, including the input itself if it is a cell. Negative query indices count from the end.
39
+
40
+ `findText("...")` or `findText(/pattern/i)` searches contiguous text in selected paragraphs or runs, including text split across runs. It returns a paragraph and a half-open `[start, end)` range in JavaScript UTF-16 offsets. It does not join separate paragraphs. Regular expressions search each selected contiguous interval, preserve the caller's lastIndex, and omit zero-length matches. Queries preserve the selected revision view. Semantic text includes authored content; generated fields and physical appearances can differ after layout. Rasterized picture/chart text has no text query model.
41
+
42
+ Bookmarks expose their published start target through `bookmark.run`; they do not represent an enclosing bookmark text range. Symbols without a Unicode text representation use the object replacement character in query text.
43
+
44
+ ## Replace text and fill templates
45
+
46
+ ```ts
47
+ const customer = doc.query.contentControls().where({ tag: "customer" }).one();
48
+ await doc.update({ target: customer, text: "Willow Design" });
49
+ ```
50
+
51
+ `update` accepts one update or a batch. Targets are runs, paragraphs, cells, content controls, their IDs, or search ranges. Element targets must contain ordinary text in one paragraph. Replacement text inherits the first selected run’s formatting; other selected runs become empty. Tabs and newlines are supported. Batches reject overlapping or missing targets, generated fields, and multi-paragraph replacements without changing the document.
52
+
53
+ Use search ranges for partial replacements while keeping surrounding text and formatting:
54
+
55
+ ```ts
56
+ await doc.update(doc.query.findText("{{customer}}").map(target => ({ target, text: "Willow Design" })));
57
+ ```
58
+
59
+ Batch range offsets refer to the original paragraph text. Empty ranges insert text; offsets must not split an emoji or other surrogate pair. Unchanged updates preserve existing formatting, snapshots, and attachments.
60
+
61
+ Browser updates rebuild layout and refresh mounted views. Query, semantic, layout, page and geometry values obtained earlier remain snapshots; read them again after updating. Application attachments are removed during refresh and can be reattached using fresh geometry. Updates run in submission order and accept `{ signal }` for cancellation. The original `source` is unchanged.
62
+
63
+ Node inspection exposes the same update method. For server rendering, inspect the input with the default entry point to select elements or ranges, then pass those updates to the server document before `renderPage`, `pdf`, or `save`. Exports reflect the updated text.
64
+
65
+ ## Save an edited Word document
66
+
67
+ `save()` returns DOCX bytes in Node, the browser, and server rendering. Saving without edits returns the original file bytes. Text edits preserve surrounding run formatting and unedited package content. Saving and updating run in submission order. Both accept `{ signal }`; cancellation leaves the committed document available. Call `dispose()` when finished, including for Node documents, to release the preserved source package.
68
+
69
+ ```js
70
+ import { readFile, writeFile } from "node:fs/promises";
71
+ import { openDocument } from "@onodocs/sdk";
72
+
73
+ const doc = await openDocument(await readFile("template.docx"));
74
+ try {
75
+ await doc.update({
76
+ target: doc.query.contentControls().where({ tag: "customer" }).one(),
77
+ text: "Willow Design"
78
+ });
79
+ await writeFile("completed.docx", await doc.save());
80
+ } finally {
81
+ doc.dispose();
82
+ }
83
+ ```
84
+
85
+ In a browser, create a Blob from the returned bytes with type `application/vnd.openxmlformats-officedocument.wordprocessingml.document`. In a combined deployment, serve `await doc.save()` from an application-owned download endpoint with that content type and a `.docx` filename. The backend processing document owns saving; a Canvas-only frontend does not contain the original Word package.
86
+
87
+ The SDK package includes a complete command-line example. After installing it, run `node node_modules/@onodocs/sdk/examples/fill-template.mjs template.docx completed.docx customer "Willow Design"`. Set `ONODOCS_LICENSE_KEY` for trial or commercial use. The source workspace keeps the same example at `examples/fill-template.mjs`.
88
+
89
+ The initial editing profile covers ordinary text and unbound text or rich-text content controls within one paragraph, including supported tables, headers, footers and notes. Edits inside bound or content-locked controls, temporary or specialized controls, fields, revisions, alternate-content branches, and source runs with generated or unsupported content are rejected before publication. Plain-text controls that disallow multiple lines reject newlines. Filling an unbound text control clears its placeholder-display flag. Edits elsewhere preserve these constructs. `DocumentSaveError.reason` identifies source-writing restrictions; existing semantic target errors remain `RangeError`.
90
+
91
+ Signed packages, enforced document protection, enabled change tracking, macro-enabled documents and templates can be saved unchanged, but editing them is not supported. Field recalculation and revision authoring are not included. Field caches remain as imported, and saved bookmark/comment markers retain their original run boundaries. Original authored formatting is preserved independently of rendering compatibility decisions.
92
+
93
+ Evaluation exports contain the full Word document without adding a Word watermark. The existing non-production evaluation terms apply; rendered evaluation output continues to include its watermark.
94
+
95
+ ## Edit formatting and paragraphs
96
+
97
+ `edit()` accepts a logical selection with paragraph IDs and UTF-16 offsets. It returns the selection in the resulting document. Use fresh queries after editing because structural operations rebuild the document and its IDs. The `source` getter then describes the new committed package.
85
98
 
86
99
  ```js
87
- view.attach(checkbox, { anchor: doc.geometry.fragments(doc.query.findText("Client approval:").one())[0], placement: "outside-right", gap: 80, size: { width: 420, height: 420 } });
88
- view.attach(submitButton, { anchor: doc.pages[0], placement: "inside-bottom-right", inset: 480, size: { width: 2400, height: 600 } });
89
- ```
90
-
91
- Inside placements use `inside-{top|center|bottom}-{left|center|right}` with optional `inset`. Outside placements use `outside-{top|right|bottom|left}` with optional `gap` and center the element along the selected edge. The default is `overlay`. All sizes, spacing and offsets use twips (1/1440 inch). Without a size, the element fills the anchor, reduced by the inset for inside placement. Offsets apply after placement. Content anchors retain their transforms and clips; page anchors use the full page. Attachments do not reflow text, so reserve room in the document.
92
-
93
- Adjacent styled runs on one line share a fragment. Fragment bounds enclose the unclipped geometry; the supplied clips determine which portions are visible.
94
-
95
- `view.setZoom("fit-width")` tracks container width. Numeric zoom uses 96 CSS pixels per inch at `1`. `view.toClient(pageIndex, point)` and `view.fromClient({ x: event.clientX, y: event.clientY })` translate coordinates. `view.hitTest(clientPoint)` returns `{ elementId, fragment, caret? }`. Resolve `elementId` with `doc.query.get(elementId)` when the processing SDK is available. Text hits resolve to a run or paragraph; table whitespace resolves to a cell or table. `caret` contains `{ paragraphId, offset, point }`, with the same UTF-16 offset convention as search ranges. Generated field text omits a caret when no source position can be represented. Raster drawings without a semantic hit identity and blank page space return no element.
96
-
97
- `view.scrollTo(anchor, { block: "center" })` reveals a page or a physical fragment. Resolve elements and text ranges through `doc.geometry.fragments(anchor)` first, then choose the occurrence to reveal.
98
-
99
- A run's optional `link` is `{ kind: "external", target, tooltip }` or `{ kind: "bookmark", target }`. Use bookmark queries and scrollTo for internal links. The host application decides whether and how to open external destinations.
100
-
101
- All page and query indices are **zero based**. Page geometry uses **twips: 1440 units per inch**. `await canvasDocument.pages[index].render(canvas, { dpi: 144 })` prepares the page resources and paints the canvas; `await doc.pages[index].load(signal)` returns the immutable page commands and their resources. Containers should have a usable width. Multiple views can share one document.
102
-
103
- Call `attachment.dispose()`, `view.dispose()`, `canvasDocument.dispose()`, or `doc.dispose()` when done. Canvas disposal removes its views and releases its font registrations. Processing disposal also disposes subscribed Canvas documents and releases processing resources. Models remain inspectable; further painting is rejected. Pass `{ signal }` when opening. Opening copies byte inputs; a File/Blob is read once by either the browser or analysis entry point. Fetch URLs in host code with the authentication policy your application needs, then pass bytes.
104
-
105
- ## Show pages while processing
106
-
107
- Progress events expose a document source that the Canvas package can observe. Create the view once. Its subscription follows provisional page replacements and the completed layout.
108
-
109
- ```ts
110
- import { openDocument } from "@onodocs/sdk/browser";
111
- import { createDocument, type CanvasDocument } from "@onodocs/canvas";
112
-
113
- let canvasDocument: CanvasDocument | undefined;
114
- const doc = await openDocument(bytes, {
115
- signal,
116
- onProgress(progress) {
117
- canvasDocument ??= createDocument(progress.document, { container });
118
- status.textContent = progress.stage;
119
- },
120
- });
121
- ```
122
-
123
- Cancelling or failing processing disposes the source and its subscribed view. After success, dispose `doc` when the application closes the document. A Canvas document can be disposed earlier without closing the processor.
124
-
125
- ## Render on the server
126
-
127
- ```ts
128
- import { readFile, writeFile } from "node:fs/promises";
129
- import { createRenderer } from "@onodocs/sdk/server";
130
-
131
- const renderer = await createRenderer({ channel: "chrome" });
132
- try {
133
- const doc = await renderer.openDocument(await readFile("invoice.docx"));
134
- try {
135
- await writeFile("invoice.png", await doc.renderPage(0, { dpi: 144 }));
136
- await writeFile("invoice.pdf", await doc.pdf({ dpi: 150 }));
137
- } finally { await doc.dispose(); }
138
- } finally { await renderer.dispose(); }
139
- ```
140
-
141
- Install Chromium/Chrome separately; use `executablePath` for a deployment-managed binary or `channel: "chrome"`/`"msedge"` for an installed browser. The renderer reuses one browser, with a separate context per document. PNG and JPEG output use the browser engine's Canvas path. PDF contains rasterized pages at their document sizes, including mixed sizes. It does not provide searchable PDF text or tagged-PDF accessibility.
142
-
143
- Contexts make no outbound network requests. Supply fonts as data URLs through `fonts`, or install them on the rendering host. Browser consumers can also supply font URLs. Embedded document fonts use the engine's existing resource path. Use identical browser versions and fonts when reproducible output matters.
144
-
145
- Opening and output methods accept an AbortSignal. Cancelling an active operation closes that document context; open a new document to retry. Cancelling an operation still waiting in the queue leaves the active operation and document intact. Both server documents and renderers expose `disposed`. Disposing the renderer closes all remaining documents. A renderer can open multiple documents concurrently; output requests on one document are serialized. Applications control scheduling and process isolation.
146
-
147
- ## Process on the backend and render in the frontend
148
-
149
- The server renderer publishes a manifest with page sizes and interaction data. Its `page(index)` method returns the same `PageDisplayList` used by local Canvas rendering. Expose these results through authenticated routes in your application. The package does not install an HTTP server or choose an authentication policy.
150
-
151
- ```ts
152
- import { createRenderer } from "@onodocs/sdk/server";
153
- import { encode } from "@onodocs/canvas";
154
-
155
- const renderer = await createRenderer({ channel: "chrome" });
156
- const doc = await renderer.openDocument(bytes, { fonts });
157
-
158
- const manifestResponse = encode(await doc.manifest());
159
- const firstPageResponse = encode(await doc.page(0));
160
- ```
161
-
162
- Use `encode` and `decode` for HTTP bodies because page geometry contains bigint coordinates and resources contain byte arrays. They encode the canonical contracts directly. The transport has no separate document model or schema version. Both packages should come from the same release.
163
-
164
- The frontend imports only the Canvas package. This example expects `/document/manifest` and `/document/pages/:index` routes supplied by the application:
165
-
166
- ```ts
167
- import { createDocument, decode, type DocumentManifest, type PageDisplayList } from "@onodocs/canvas";
168
-
169
- const response = await fetch("/document/manifest");
170
- if (!response.ok) throw new Error("Document is unavailable.");
171
- const manifest = decode<DocumentManifest>(await response.text());
172
- const canvasDocument = createDocument({
173
- pages: manifest.pages.map(page => ({
174
- ...page,
175
- async load(signal) {
176
- const response = await fetch(`/document/pages/${page.index}`, { signal });
177
- if (!response.ok) throw new Error("Page is unavailable.");
178
- return decode<PageDisplayList>(await response.text());
179
- },
180
- })),
181
- }, { container });
100
+ const paragraph = doc.query.paragraphs().first();
101
+ let selection = {
102
+ start: { paragraphId: paragraph.id, offset: 0 },
103
+ end: { paragraphId: paragraph.id, offset: paragraph.text.length }
104
+ };
105
+ selection = await doc.edit({ kind: "format", selection, formatting: { bold: true, fontSize: 18 } });
106
+ selection = await doc.edit({ kind: "replace", selection, text: "First paragraph\nSecond paragraph" });
107
+ await doc.edit({ kind: "align", selection, alignment: "center" });
108
+ const bytes = await doc.save();
182
109
  ```
183
110
 
184
- Selection, copying, hit testing, zoom and attachments work from the published geometry. Semantic queries and updates remain on the backend. After an update, publish a fresh manifest and page set; do not mix pages from different document snapshots. Applications can cache immutable page responses or share resource transfers to suit their transport. Dispose the Canvas document when replacing its source, and dispose server documents and the renderer when they are no longer needed.
185
-
186
- Published pages contain visible text and resources. Keeping the source DOCX on the backend does not hide the displayed content from the frontend. Supplied and embedded font bytes travel with the pages. Fonts resolved from the server's installed fonts remain explicit local font declarations and require the same fonts on the client; supply font files for deployment across different hosts. Backend image/PDF export remains available when a client cannot reproduce the required font environment.
187
-
188
- ## Errors and application data
189
-
190
- Opening, updating, and exporting can reject. Handle errors at your application boundary and dispose browser/server resources in finally blocks. The analysis and browser entries export `PackageOperationError`, `SemanticOperationError`, and `QueryCardinalityError`; cancellation uses the signal's reason.
191
-
192
- Query elements are readonly in-memory models with parent references, source metadata, and bigint values. Project the fields your application needs instead of serializing the entire model:
193
-
194
- ```ts
195
- const paragraphs = doc.query.stories().where({ story: "body" }).paragraphs()
196
- .map(({ id, text, style }) => ({ id, text, style }));
197
- const json = JSON.stringify(paragraphs);
198
- ```
199
-
200
- The same projection supports AI context, search indexing, and backend messages without coupling your data format to engine internals.
201
-
202
- ## Low-level stages
203
-
204
- `doc.source` is the typed source model and `doc.semantic` contains completed document meaning. Element `source` locations refer into `doc.source.sourceFacts` and provenance. These identifiers belong to that loaded document, not arbitrary later revisions. The models describe supported DOCX content; they do not preserve every XML detail for round-trip writing.
205
-
206
- The main SDK also exports the core stage APIs: `parseDocumentSource`, `createSemanticDocument`, `parseDocumentForRendering`, font preparation, `createLayoutDocument`, and `createPageDisplayList`. Low-level `materializePageDisplayListToCanvas` is exported by `@onodocs/canvas`. Advanced consumers can stop at a stage or replay paint through their own adapter. Use canonical readonly contracts; do not mutate stage outputs. Layout remains the authority for all physical geometry.
207
-
208
- ### Extract a table as text
209
-
210
- ```ts
211
- const data = doc.query.tables()
212
- .where({ headers: ["Deliverable", "Owner", "Due"] })
213
- .one().textRows;
214
- ```
215
-
216
- `textRows` includes the first row and keeps cells grouped by their own table rows. The header filter matches the first row exactly, in order and case sensitively; it does not infer header formatting. Nested tables remain text inside their containing cells. Merged cells retain authored topology; the result is not padded into a rectangular grid. Use `data.slice(1)` to omit the header. The readonly matrix is a snapshot; query again after an update.
217
-
218
- ## Large documents
219
-
220
- Use the progress subscription shown above to display early pages while layout continues. Early pages are provisional: page counts, fields, notes and placement can change as formatting settles. Opening resolves with the complete document. Abort the loading signal to dispose the source and its subscribed views.
221
-
222
- Mounted views paint the visible neighborhood and release distant canvas pixels. Zoom and scrolling render pages on demand. Decoded images are acquired only for a page render and released afterward. Parsing and document layout still require memory proportional to document content; virtualization does not make the document model constant-memory. Auto-fit tables and document-wide fields can require later content before their output settles.
223
-
224
- Await page renders before reading or exporting the canvas:
225
-
226
- ```js
227
- await canvasDocument.pages[0].render(canvas, { dpi: 144, signal: controller.signal });
228
- const png = canvas.toDataURL("image/png");
229
- ```
111
+ Character formatting supports `bold`, `italic`, `underline`, `strike`, `fontSize` in points, `fontFamily`, and `color` as `#RRGGBB`. Values change direct formatting on the selected text while preserving the other authored properties. `run.formatting` exposes the common effective character properties for controls; `paragraph.alignment` exposes left, center, right or justified alignment. These convenience properties describe ordinary Latin text; they do not replace the full semantic typography model for script-specific inspection.
112
+
113
+ A replacement with a collapsed selection inserts text. Newlines split paragraphs; set `paragraphBreaks: false` for line breaks within a paragraph. Replacing a selection across adjacent paragraphs joins their surviving prefix and suffix. Empty text deletes the selection. New text inherits the starting run's formatting unless the operation includes `formatting`. Split paragraphs inherit the starting paragraph's properties. The first paragraph's properties win when joining.
114
+
115
+ Rich editing supports ordinary text paragraphs, including empty paragraphs and paragraphs in table cells and secondary stories. Operations affecting paragraphs with fields, links, controls, drawings, comments, bookmarks, tracked revisions or unsupported children reject atomically. Splitting or joining section boundaries and joining across cells or stories also reject. Such content remains preserved elsewhere in the document. The narrower `update()` operation continues to support safe text replacement within unbound controls.
116
+
117
+ Node and browser documents expose `edit`; server documents forward the same operation. Parsing, semantic construction and browser layout finish before publication, so cancellation or failure leaves the previous document usable. This first implementation rebuilds the document for structural edits. Browser resources stay with the document lifetime, and the sample keeps undo history as document snapshots. Large documents need further editing-performance work before this can serve as a general-purpose editor.
118
+
119
+ ## Run the inline browser editor sample
120
+
121
+ From the source workspace, run `npm run build` and `npm run example:editor`, then open `http://127.0.0.1:5190`. The sample opens its included Word document and accepts local DOCX files. Click directly on a rendered page to type, drag to select text, and use the toolbar to change formatting. Enter splits paragraphs; Backspace at a paragraph's start joins it to the previous paragraph. Shift+Enter inserts a line break. The sample also supports keyboard selection, copy/cut, plain-text paste, undo/redo and Word download. Files remain in the browser.
122
+
123
+ The sample is in `examples/browser-editor/`. It consumes only exported SDK and Canvas APIs. Canvas `createDocument` accepts a native textarea through `viewOptions.input`, plus an `onSelectionChange` callback in the same options. The view positions this input at the published caret, handles selection using layout geometry, and exposes `selection` and `setSelection`. The application handles input events and submits `doc.edit()` operations. Text, line wrapping and pagination remain rendered by OnoDocs.
124
+
125
+ Run `npm run validate:editor` for the browser editing journey. The sample currently uses the workspace build; it requires the new editing APIs and does not work against the earlier published package.
126
+
127
+ ## Render and attach ordinary DOM
128
+
129
+ Mounted views include read-only text selection and plain-text copying. Drag or Shift-click to select text, extend with Shift and the arrow/Home/End keys, select all with Ctrl/Command+A, and copy with Ctrl/Command+C or the Copy context menu. Selection spans pages and survives scrolling and zoom, including pages whose canvases have been released. Document updates clear the previous selection. Attached inputs retain their normal editing and clipboard behavior.
130
+
131
+ ```ts
132
+ import { openDocument } from "@onodocs/sdk/browser";
133
+ import { createDocument } from "@onodocs/canvas";
134
+
135
+ const doc = await openDocument(file);
136
+ const canvasDocument = createDocument(doc);
137
+ const view = canvasDocument.mount(container);
138
+ const field = doc.query.contentControls().where({ tag: "customer-name" }).one();
139
+ const input = document.createElement("input");
140
+ input.name = "customerName";
141
+ const attachment = view.attach(input, { anchor: doc.geometry.fragments(field)[0] });
142
+ ```
143
+
144
+ The example requires a tagged control that produces one visible fragment. Call `doc.geometry.fragments(field)` for content spanning lines, pages, or repeated stories, and attach to an explicitly chosen fragment. Fragments retain transforms and clipping. Table/cell anchors use their final physical rectangles. `offset` and `size` optionally adjust attachment placement in page units. Values, validation, submission, focus, and accessibility labels belong to the application. Attaching an element moves it into the view. Detachment removes it and restores its original inline style. Reattaching the same HTML element disposes its previous attachment; old handles become inert. An attachment exposes its disposed state.
145
+
146
+ Pages are anchors too. Placement avoids manual rectangle arithmetic:
147
+
148
+ ```js
149
+ view.attach(checkbox, { anchor: doc.geometry.fragments(doc.query.findText("Client approval:").one())[0], placement: "outside-right", gap: 80, size: { width: 420, height: 420 } });
150
+ view.attach(submitButton, { anchor: doc.pages[0], placement: "inside-bottom-right", inset: 480, size: { width: 2400, height: 600 } });
151
+ ```
152
+
153
+ Inside placements use `inside-{top|center|bottom}-{left|center|right}` with optional `inset`. Outside placements use `outside-{top|right|bottom|left}` with optional `gap` and center the element along the selected edge. The default is `overlay`. All sizes, spacing and offsets use twips (1/1440 inch). Without a size, the element fills the anchor, reduced by the inset for inside placement. Offsets apply after placement. Content anchors retain their transforms and clips; page anchors use the full page. Attachments do not reflow text, so reserve room in the document.
154
+
155
+ Adjacent styled runs on one line share a fragment. Fragment bounds enclose the unclipped geometry; the supplied clips determine which portions are visible.
156
+
157
+ `view.setZoom("fit-width")` tracks container width. Numeric zoom uses 96 CSS pixels per inch at `1`. `view.toClient(pageIndex, point)` and `view.fromClient({ x: event.clientX, y: event.clientY })` translate coordinates. `view.hitTest(clientPoint)` returns `{ elementId, fragment, caret? }`. Resolve `elementId` with `doc.query.get(elementId)` when the processing SDK is available. Text hits resolve to a run or paragraph; table whitespace resolves to a cell or table. `caret` contains `{ paragraphId, offset, point }`, with the same UTF-16 offset convention as search ranges. Generated field text omits a caret when no source position can be represented. Raster drawings without a semantic hit identity and blank page space return no element.
158
+
159
+ `view.scrollTo(anchor, { block: "center" })` reveals a page or a physical fragment. Resolve elements and text ranges through `doc.geometry.fragments(anchor)` first, then choose the occurrence to reveal.
160
+
161
+ A run's optional `link` is `{ kind: "external", target, tooltip }` or `{ kind: "bookmark", target }`. Use bookmark queries and scrollTo for internal links. The host application decides whether and how to open external destinations.
162
+
163
+ All page and query indices are **zero based**. Page geometry uses **twips: 1440 units per inch**. `await canvasDocument.pages[index].render(canvas, { dpi: 144 })` prepares the page resources and paints the canvas; `await doc.pages[index].load(signal)` returns the immutable page commands and their resources. Containers should have a usable width. Multiple views can share one document.
164
+
165
+ Call `attachment.dispose()`, `view.dispose()`, `canvasDocument.dispose()`, or `doc.dispose()` when done. Canvas disposal removes its views and releases its font registrations. Processing disposal also disposes subscribed Canvas documents and releases processing resources. Models remain inspectable; further painting is rejected. Pass `{ signal }` when opening. Opening copies byte inputs; a File/Blob is read once by either the browser or analysis entry point. Fetch URLs in host code with the authentication policy your application needs, then pass bytes.
166
+
167
+ ## Show pages while processing
168
+
169
+ Progress events expose a document source that the Canvas package can observe. Create the view once. Its subscription follows provisional page replacements and the completed layout.
170
+
171
+ ```ts
172
+ import { openDocument } from "@onodocs/sdk/browser";
173
+ import { createDocument, type CanvasDocument } from "@onodocs/canvas";
174
+
175
+ let canvasDocument: CanvasDocument | undefined;
176
+ const doc = await openDocument(bytes, {
177
+ signal,
178
+ onProgress(progress) {
179
+ canvasDocument ??= createDocument(progress.document, { container });
180
+ status.textContent = progress.stage;
181
+ },
182
+ });
183
+ ```
184
+
185
+ Cancelling or failing processing disposes the source and its subscribed view. After success, dispose `doc` when the application closes the document. A Canvas document can be disposed earlier without closing the processor.
186
+
187
+ ## Render on the server
188
+
189
+ ```ts
190
+ import { readFile, writeFile } from "node:fs/promises";
191
+ import { createRenderer } from "@onodocs/sdk/server";
192
+
193
+ const renderer = await createRenderer({ channel: "chrome" });
194
+ try {
195
+ const doc = await renderer.openDocument(await readFile("invoice.docx"));
196
+ try {
197
+ await writeFile("invoice.png", await doc.renderPage(0, { dpi: 144 }));
198
+ await writeFile("invoice.pdf", await doc.pdf({ dpi: 150 }));
199
+ } finally { await doc.dispose(); }
200
+ } finally { await renderer.dispose(); }
201
+ ```
202
+
203
+ Install Chromium/Chrome separately; use `executablePath` for a deployment-managed binary or `channel: "chrome"`/`"msedge"` for an installed browser. The renderer reuses one browser, with a separate context per document. PNG and JPEG output use the browser engine's Canvas path. PDF contains rasterized pages at their document sizes, including mixed sizes. It does not provide searchable PDF text or tagged-PDF accessibility.
204
+
205
+ Contexts make no outbound network requests. Supply fonts as data URLs through `fonts`, or install them on the rendering host. Browser consumers can also supply font URLs. Embedded document fonts use the engine's existing resource path. Use identical browser versions and fonts when reproducible output matters.
206
+
207
+ Opening and output methods accept an AbortSignal. Cancelling an active operation closes that document context; open a new document to retry. Cancelling an operation still waiting in the queue leaves the active operation and document intact. Both server documents and renderers expose `disposed`. Disposing the renderer closes all remaining documents. A renderer can open multiple documents concurrently; output requests on one document are serialized. Applications control scheduling and process isolation.
208
+
209
+ ## Process on the backend and render in the frontend
210
+
211
+ The [complete backend viewer sample](https://github.com/onodocs/onodocs/tree/main/examples/backend-viewer) includes a Node.js server, working endpoints, a browser frontend and an authored DOCX. Install Node.js 22.18 or newer and Google Chrome, then run:
212
+
213
+ ```sh
214
+ git clone https://github.com/onodocs/onodocs.git
215
+ cd onodocs/examples/backend-viewer
216
+ npm install
217
+ npm start
218
+ ```
219
+
220
+ Open http://127.0.0.1:5175. Use `npm start -- /path/to/document.docx` to open another server-side file. Stop the server with Ctrl+C. The [website guide](https://onodocs.com/developers/#combined) includes the complete server and frontend source.
221
+
222
+ The server's `manifest({ format: "json" })` and `page(index, { format: "json" })` methods return response bodies ready to send with `Content-Type: application/json`. For a document at `/document`, serve its manifest there and individual pages at `/document/pages/:index`. The original methods without a format option still return typed objects for in-process use.
223
+
224
+ ```js
225
+ import { openDocument } from "@onodocs/canvas";
226
+
227
+ const doc = await openDocument("/document", {
228
+ container: document.querySelector("#pages"),
229
+ viewOptions: { zoom: "fit-width" }
230
+ });
231
+ await doc.view.whenRendered();
232
+ ```
233
+
234
+ Canvas handles loading and decoding. It requests page content as needed. `signal` cancels requests; `request` accepts fetch options such as `headers` and `credentials` for both the manifest and page requests. Same-origin cookies work by default. Requests default to no-store, and non-successful HTTP responses reject with the status code. Dispose the Canvas document when closing or replacing the view.
235
+
236
+ The application owns the endpoints and document authorization. Keep server documents open while their endpoints are in use, then dispose them. For updates, publish each completed snapshot at its own URL and reopen the frontend view. Use matching SDK and Canvas releases.
237
+
238
+ Selection, copying, hit testing, zoom and attachments use the published geometry. Queries and updates stay on the backend. The frontend receives the visible text and resources needed to display the document. Supplied and embedded fonts travel with the pages; fonts resolved from the server's installed fonts must also be available on the client.
239
+
240
+ ## Errors and application data
241
+
242
+ Opening, updating, and exporting can reject. Handle errors at your application boundary and dispose browser/server resources in finally blocks. The analysis and browser entries export `PackageOperationError`, `SemanticOperationError`, and `QueryCardinalityError`; cancellation uses the signal's reason.
243
+
244
+ Query elements are readonly in-memory models with parent references, source metadata, and bigint values. Project the fields your application needs instead of serializing the entire model:
245
+
246
+ ```ts
247
+ const paragraphs = doc.query.stories().where({ story: "body" }).paragraphs()
248
+ .map(({ id, text, style }) => ({ id, text, style }));
249
+ const json = JSON.stringify(paragraphs);
250
+ ```
251
+
252
+ The same projection supports AI context, search indexing, and backend messages without coupling your data format to engine internals.
253
+
254
+ ## Low-level stages
255
+
256
+ `doc.source` is the typed source model and `doc.semantic` contains completed document meaning. Element `source` locations refer into `doc.source.sourceFacts` and provenance. These identifiers belong to that loaded document, not arbitrary later revisions. The models describe supported DOCX content; they do not preserve every XML detail for round-trip writing.
257
+
258
+ The main SDK also exports the core stage APIs: `parseDocumentSource`, `createSemanticDocument`, `parseDocumentForRendering`, font preparation, `createLayoutDocument`, and `createPageDisplayList`. Low-level `materializePageDisplayListToCanvas` is exported by `@onodocs/canvas`. Advanced consumers can stop at a stage or replay paint through their own adapter. Use canonical readonly contracts; do not mutate stage outputs. Layout remains the authority for all physical geometry.
259
+
260
+ ### Extract a table as text
261
+
262
+ ```ts
263
+ const data = doc.query.tables()
264
+ .where({ headers: ["Deliverable", "Owner", "Due"] })
265
+ .one().textRows;
266
+ ```
267
+
268
+ `textRows` includes the first row and keeps cells grouped by their own table rows. The header filter matches the first row exactly, in order and case sensitively; it does not infer header formatting. Nested tables remain text inside their containing cells. Merged cells retain authored topology; the result is not padded into a rectangular grid. Use `data.slice(1)` to omit the header. The readonly matrix is a snapshot; query again after an update.
269
+
270
+ ## Large documents
271
+
272
+ Use the progress subscription shown above to display early pages while layout continues. Early pages are provisional: page counts, fields, notes and placement can change as formatting settles. Opening resolves with the complete document. Abort the loading signal to dispose the source and its subscribed views.
273
+
274
+ Mounted views paint the visible neighborhood and release distant canvas pixels. Zoom and scrolling render pages on demand. Decoded images are acquired only for a page render and released afterward. Parsing and document layout still require memory proportional to document content; virtualization does not make the document model constant-memory. Auto-fit tables and document-wide fields can require later content before their output settles.
275
+
276
+ Await page renders before reading or exporting the canvas:
277
+
278
+ ```js
279
+ await canvasDocument.pages[0].render(canvas, { dpi: 144, signal: controller.signal });
280
+ const png = canvas.toDataURL("image/png");
281
+ ```