document-content-model 4.3.2 → 7.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (96) hide show
  1. package/README.md +274 -95
  2. package/dist/border-weight.cjs +26 -0
  3. package/dist/border-weight.d.cts +12 -0
  4. package/dist/border-weight.d.ts +12 -0
  5. package/dist/border-weight.js +23 -0
  6. package/dist/codec.d.cts +1 -1
  7. package/dist/codec.d.ts +1 -1
  8. package/dist/{construct-ETSlnMle.d.cts → construct-GVdTKrE7.d.cts} +242 -2
  9. package/dist/{construct-ETSlnMle.d.ts → construct-GVdTKrE7.d.ts} +242 -2
  10. package/dist/construct.cjs +13 -6
  11. package/dist/construct.d.cts +1 -1
  12. package/dist/construct.d.ts +1 -1
  13. package/dist/construct.js +13 -6
  14. package/dist/content-BCOxY0Eb.d.ts +10273 -0
  15. package/dist/content-BzkhrS6P.d.cts +10273 -0
  16. package/dist/content-json-schema-defs.cjs +1056 -189
  17. package/dist/content-json-schema-defs.js +1056 -189
  18. package/dist/content.cjs +514 -36
  19. package/dist/content.d.cts +2 -2
  20. package/dist/content.d.ts +2 -2
  21. package/dist/content.js +488 -37
  22. package/dist/decompose.cjs +7 -17
  23. package/dist/decompose.d.cts +5 -5
  24. package/dist/decompose.d.ts +5 -5
  25. package/dist/decompose.js +7 -17
  26. package/dist/definitions.cjs +5 -1
  27. package/dist/definitions.d.cts +37 -1
  28. package/dist/definitions.d.ts +37 -1
  29. package/dist/definitions.js +5 -1
  30. package/dist/factor-styles.cjs +14 -8
  31. package/dist/factor-styles.d.cts +6 -6
  32. package/dist/factor-styles.d.ts +6 -6
  33. package/dist/factor-styles.js +15 -9
  34. package/dist/flatten.cjs +5 -5
  35. package/dist/flatten.d.cts +4 -4
  36. package/dist/flatten.d.ts +4 -4
  37. package/dist/flatten.js +5 -5
  38. package/dist/font-port.d.cts +2 -2
  39. package/dist/font-port.d.ts +2 -2
  40. package/dist/index.cjs +54 -14
  41. package/dist/index.d.cts +14 -12
  42. package/dist/index.d.ts +14 -12
  43. package/dist/index.js +12 -10
  44. package/dist/{math-BlG8Tjk-.d.cts → math-BScHxedi.d.cts} +97 -14
  45. package/dist/{math-BlG8Tjk-.d.ts → math-BScHxedi.d.ts} +97 -14
  46. package/dist/math-layout.d.cts +6 -6
  47. package/dist/math-layout.d.ts +6 -6
  48. package/dist/math.cjs +32 -6
  49. package/dist/math.d.cts +2 -2
  50. package/dist/math.d.ts +2 -2
  51. package/dist/math.js +29 -7
  52. package/dist/{mathml-B1oTCtzc.d.cts → mathml-Czqc1mdZ.d.cts} +3 -3
  53. package/dist/{mathml-B1oTCtzc.d.ts → mathml-Czqc1mdZ.d.ts} +3 -3
  54. package/dist/mathml.cjs +9 -2
  55. package/dist/mathml.d.cts +1 -1
  56. package/dist/mathml.d.ts +1 -1
  57. package/dist/mathml.js +9 -2
  58. package/dist/metadata.cjs +12 -1
  59. package/dist/metadata.d.cts +13 -0
  60. package/dist/metadata.d.ts +13 -0
  61. package/dist/metadata.js +12 -1
  62. package/dist/package-node-C4mh1hLY.d.ts +1989 -0
  63. package/dist/package-node-RQMDBzwd.d.cts +1989 -0
  64. package/dist/package-node.cjs +21 -21
  65. package/dist/package-node.d.cts +2 -2
  66. package/dist/package-node.d.ts +2 -2
  67. package/dist/package-node.js +14 -14
  68. package/dist/package.cjs +5 -3
  69. package/dist/package.d.cts +253 -8
  70. package/dist/package.d.ts +253 -8
  71. package/dist/package.js +5 -3
  72. package/dist/schema-io.cjs +21 -11
  73. package/dist/schema-io.d.cts +13 -9
  74. package/dist/schema-io.d.ts +13 -9
  75. package/dist/schema-io.js +21 -12
  76. package/dist/source-Bas5ol6f.d.cts +43 -0
  77. package/dist/source-Bas5ol6f.d.ts +43 -0
  78. package/dist/source.cjs +27 -0
  79. package/dist/source.d.cts +2 -0
  80. package/dist/source.d.ts +2 -0
  81. package/dist/source.js +25 -0
  82. package/dist/style-D06NwxAJ.d.cts +29 -0
  83. package/dist/style-D06NwxAJ.d.ts +29 -0
  84. package/dist/style.cjs +2 -0
  85. package/dist/style.d.cts +2 -24
  86. package/dist/style.d.ts +2 -24
  87. package/dist/style.js +2 -1
  88. package/dist/text-layout.d.cts +1 -1
  89. package/dist/text-layout.d.ts +1 -1
  90. package/package.json +28 -29
  91. package/schemas/content-document.schema.json +6623 -1694
  92. package/schemas/{document-package.schema.json → document-tree.schema.json} +1826 -143
  93. package/dist/content-BLwxI5ba.d.cts +0 -2481
  94. package/dist/content-BdzcBsxm.d.ts +0 -2481
  95. package/dist/package-node-C-ewnL0v.d.cts +0 -459
  96. package/dist/package-node-D3scv-Cl.d.ts +0 -459
package/README.md CHANGED
@@ -1,8 +1,8 @@
1
1
  # document-schema.js
2
2
 
3
- [![GitHub](https://img.shields.io/badge/GitHub-181717?logo=github&logoColor=white)](https://github.com/ExaDev/document-schema.js) [![npm](https://img.shields.io/badge/npm-CB3837?logo=npm&logoColor=white)](https://www.npmjs.com/package/document-schema.js) [![Release](https://img.shields.io/github/v/release/ExaDev/document-schema.js)](https://github.com/ExaDev/document-schema.js/releases/latest) [![CI](https://img.shields.io/github/actions/workflow/status/ExaDev/document-schema.js/ci.yml?branch=main)](https://github.com/ExaDev/document-schema.js/actions)
3
+ [![GitHub](https://img.shields.io/badge/GitHub-181717?logo=github&logoColor=white)](https://github.com/ExaDev/documents.js/tree/main/packages/document-schema.js) [![npm](https://img.shields.io/badge/npm-CB3837?logo=npm&logoColor=white)](https://www.npmjs.com/package/document-schema.js) [![npm version](https://img.shields.io/npm/v/document-schema.js)](https://www.npmjs.com/package/document-schema.js) [![CI](https://img.shields.io/github/actions/workflow/status/ExaDev/documents.js/ci.yml?branch=main)](https://github.com/ExaDev/documents.js/actions)
4
4
 
5
- > The canonical, format-agnostic content and document-package schema pivot shared by [ooxml.js](https://github.com/ExaDev/ooxml.js), [odf.js](https://github.com/ExaDev/odf.js), [documents.js](https://github.com/ExaDev/documents.js), [pdf-codec](https://github.com/ExaDev/pdf-codec), and [markdown-codec](https://github.com/ExaDev/markdown-codec).
5
+ > The canonical, format-agnostic content and document-tree schema pivot shared by [ooxml.js](../ooxml.js/README.md), [odf.js](../odf.js/README.md), [documents.js](https://github.com/ExaDev/documents.js), [pdf-codec](../pdf-codec/README.md), and [markdown-codec](../markdown-codec/README.md).
6
6
 
7
7
  Both `ooxml.js` and `documents.js` independently arrived at the same content vocabulary, producing two field-identical copies in two places. This package is the fix: one schema, imported by every format package instead of redefined by each. It also sidesteps a circular dependency (`documents.js` depends on both `ooxml.js` and `odf.js`).
8
8
 
@@ -35,47 +35,60 @@ graph TD
35
35
  odf --> cli
36
36
  pdfcodec --> cli
37
37
 
38
- click schema "https://github.com/ExaDev/document-schema.js" "document-schema.js"
39
- click ooxml "https://github.com/ExaDev/ooxml.js" "ooxml.js"
40
- click odf "https://github.com/ExaDev/odf.js" "odf.js"
41
- click pdfcodec "https://github.com/ExaDev/pdf-codec" "pdf-codec"
42
- click mdcodec "https://github.com/ExaDev/markdown-codec" "markdown-codec"
43
- click bytecodec "https://github.com/ExaDev/byte-codec" "byte-codec"
38
+ click schema "https://github.com/ExaDev/documents.js/tree/main/packages/document-schema.js" "document-schema.js"
39
+ click ooxml "https://github.com/ExaDev/documents.js/tree/main/packages/ooxml.js" "ooxml.js"
40
+ click odf "https://github.com/ExaDev/documents.js/tree/main/packages/odf.js" "odf.js"
41
+ click pdfcodec "https://github.com/ExaDev/documents.js/tree/main/packages/pdf-codec" "pdf-codec"
42
+ click mdcodec "https://github.com/ExaDev/documents.js/tree/main/packages/markdown-codec" "markdown-codec"
43
+ click bytecodec "https://github.com/ExaDev/documents.js/tree/main/packages/byte-codec" "byte-codec"
44
44
  click documents "https://github.com/ExaDev/documents.js" "documents.js"
45
- click mcp "https://github.com/ExaDev/document-mcp" "document-mcp"
46
- click cli "https://github.com/ExaDev/document-cli" "document-cli"
45
+ click mcp "https://github.com/ExaDev/documents.js/tree/main/packages/document-mcp" "document-mcp"
46
+ click cli "https://github.com/ExaDev/documents.js/tree/main/packages/document-cli" "document-cli"
47
47
 
48
48
  style schema fill:#f9a825,stroke:#333,stroke-width:3px
49
49
  ```
50
50
 
51
- `ContentDocument` (the semantic pivot) is a discriminated union of five kinds: `wordprocessing` (docx/odt sections of paragraphs/runs/tables/images), `presentation` (pptx/odp slides of shapes), `spreadsheet` (xlsx/ods sheets of cells, columns, rows, print settings), `drawing` (odg pages of shapes plus vector primitives — rect/ellipse/line/path), and `formula` (an equation carrying its own MathML node tree plus StarMath source when the producing format had one, extended with the two-layer math model: an optional verbatim-LaTeX `presentation` authoritative for rendering, an optional semantic `content: MathExpression` tree authoritative for computation, and provenance — neither layer stored derived from the other, so editing one never silently mutates the other). `ContentEmbeddedObjectSchema` lets any of the five embed another whole `ContentDocument`. Every paragraph/run/image/table/shape/vector/spreadsheet-cell leaf also carries its own canonical `headingLevel`-or-position fields directly: a `ContentParagraph`'s optional `headingLevel` (1 = the outermost heading, independent of the round-trip-only `styleId`), and every such leaf's optional `frames: LayoutFrame[]` — that node's own rendered page position(s) (`pageIndex` plus PDF user-space `xPt`/`yPt`/`widthPt`/`heightPt`), fused directly onto the content tree once a layout pass has run. `DocumentPackage` is the single hierarchical artefact — structure, layout, and content fused in one tree (see [The package tree](#the-package-tree)): the root carries `kind`, `metadata`, the optional document-level `symbolTable` and rendered `pages`, the optional package-level `styles`/`definitions`/`layers`/`attachments`/`destinations` tables (see [Definitions tables and styles](#definitions-tables-and-styles)), and `children` — one group per top-level container with the content tree grouped inside it, where a group's node may also be one of the six fidelity construct descriptors, which the flat form carries instead as a matched `constructStart`/`constructEnd` block pair (see [Fidelity constructs](#fidelity-constructs)); the schema does not keep populated `frames` fields and `pages` in sync or detect staleness, and does not check that a tree's `style` refs name table entries (both are producer responsibilities, exactly as the frames/pages pairing always was). Every one of the five kinds also accepts an optional document-level `symbolTable` — the math curation layer mapping each written symbol glyph (within a scope) to its id, quantity kind, preferred unit, and definition source, alongside the unit registry (SI dimension-exponent vectors, exact rational conversions, per-unit-system normalisation contexts) that the `qty` nodes of lowered formulas resolve against.
51
+ `ContentDocument` (the semantic pivot) is a discriminated union of five kinds: `wordprocessing` (docx/odt sections of paragraphs/runs/tables/images), `presentation` (pptx/odp slides of shapes), `spreadsheet` (xlsx/ods sheets of cells, columns, rows, print settings), `drawing` (odg pages of shapes plus vector primitives — rect/ellipse/line/path), and `formula` (an equation carrying its own MathML node tree plus StarMath source when the producing format had one, extended with the two-layer math model: an optional verbatim-LaTeX `presentation` authoritative for rendering, an optional semantic `content: MathExpression` tree authoritative for computation, and provenance — neither layer stored derived from the other, so editing one never silently mutates the other). `ContentEmbeddedObjectSchema` lets any of the five embed another whole `ContentDocument`; its one exception is the `'chart'` objectKind, which names a chart graphic frame's cached series/category model carried as a small spreadsheet document — a chart is not itself a document kind, and the objectKind/document pairing being a convention rather than a constraint is what lets the two halves each say their own truth. Every paragraph/run/image/table/shape/vector/spreadsheet-cell leaf also carries its own canonical `headingLevel`-or-position fields directly: a `ContentParagraph`'s optional `headingLevel` (1 = the outermost heading, independent of the round-trip-only `styleId`), and every such leaf's optional `frames: LayoutFrame[]` — that node's own rendered page position(s) (`pageIndex` plus PDF user-space `xPt`/`yPt`/`widthPt`/`heightPt`), fused directly onto the content tree once a layout pass has run. `DocumentTree` is the single hierarchical artefact — structure, layout, and content fused in one tree (see [The package tree](#the-package-tree)): the root carries `kind`, `metadata`, the optional document-level `symbolTable` and rendered `pages`, the optional package-level `styles`/`definitions`/`layers`/`attachments`/`destinations` tables (see [Definitions tables and styles](#definitions-tables-and-styles)) and the keyed `source` residue table (see [The residue channel](#the-residue-channel)), and `children` — one group per top-level container with the content tree grouped inside it, where a group's node may also be one of the six fidelity construct descriptors, which the flat form carries instead as a matched `constructStart`/`constructEnd` block pair (see [Fidelity constructs](#fidelity-constructs)); the schema does not keep populated `frames` fields and `pages` in sync or detect staleness, and does not check that a tree's `style` refs name table entries (both are producer responsibilities, exactly as the frames/pages pairing always was). Every one of the five kinds also accepts an optional document-level `symbolTable` — the math curation layer mapping each written symbol glyph (within a scope) to its id, quantity kind, preferred unit, and definition source, alongside the unit registry (SI dimension-exponent vectors, exact rational conversions, per-unit-system normalisation contexts) that the `qty` nodes of lowered formulas resolve against.
52
52
 
53
53
  The `LayoutDocument` family (pages of positioned `LayoutItem`s — `text`/`image`/`rect`/`line`/`ellipse`/`path`/`link` in PDF user-space coordinates) no longer lives here: 4.0.0 demoted it to a pdf-codec-private model ([pdf-codec#65](https://github.com/ExaDev/pdf-codec/issues/65)), where the only codec that ever read or wrote it owns it outright. `documentFromJson` recognises old layout-document `$schema` URIs and throws a tombstone pointing at pdf-codec rather than failing as if the value were unrelated. Dependents stay on document-schema.js 3.x via semver until their own majors, so the demotion is not a cascade-breaker.
54
54
 
55
- The package contains [Zod](https://zod.dev) schemas, their inferred types, trivial schema-attached helpers (hex-colour conversion, recursive structural type guards, the style-resolution helpers of `src/definitions.ts`, the construct-marker balance check of `src/content.ts`), one small structural interface (`ContentCodec`, see [Codecs](#codecs)), and the structural transform between the two encodings it defines (`decompose`/`flattenPackage`/`factorStyles`/`assemblePackage`, see [The package boundary](#the-package-boundary)). No format-specific behaviour and no I/O of any kind — no XML, ZIP, PDF, or binary handling; the sole dependency is `zod`.
55
+ The package contains [Zod](https://zod.dev) schemas, their inferred types, trivial schema-attached helpers (hex-colour conversion, recursive structural type guards, the style-resolution helpers of `src/definitions.ts`, the construct-marker balance check of `src/content.ts`), one small structural interface (`ContentCodec`, see [Codecs](#codecs)), and the structural transform between the two encodings it defines (`decompose`/`flattenTree`/`factorStyles`/`assembleTree`, see [The package boundary](#the-package-boundary)). No format-specific behaviour and no I/O of any kind — no XML, ZIP, PDF, or binary handling; the sole dependency is `zod`.
56
56
 
57
57
  Two format-agnostic helpers live here because they operate on the content model itself: cell-addressing utilities in `src/a1.ts` (0-based row/column indices, row-first order matching `ContentSheetCell`'s `{row, column}`) and the `FontFace` interface in `src/font-port.ts` (`{family, bold, italic}`).
58
58
 
59
59
  ## Usage
60
60
 
61
61
  ```ts
62
- import { ContentDocumentSchema, DocumentPackageSchema } from 'document-schema.js';
62
+ import { ContentDocumentSchema, DocumentTreeSchema } from "document-schema.js";
63
63
 
64
64
  // The codec-exchange form: what every format's reader produces and every writer consumes -- always flat,
65
65
  // always fully materialised (no styles table, no refs), never versioned (that lives on the serialised artefact).
66
- const content = ContentDocumentSchema.parse(someWordprocessingOrPresentationValue);
66
+ const content = ContentDocumentSchema.parse(
67
+ someWordprocessingOrPresentationValue,
68
+ );
67
69
 
68
70
  // The package tree: what a serialised dump carries. `children` holds one group per top-level container,
69
71
  // with the content grouped inside it (see "The package tree" below).
70
- const pkg = DocumentPackageSchema.parse({
71
- kind: 'wordprocessing',
72
- metadata: { title: 'Example' },
72
+ const pkg = DocumentTreeSchema.parse({
73
+ kind: "wordprocessing",
74
+ metadata: { title: "Example" },
73
75
  children: [
74
76
  {
75
- node: { kind: 'section', pageSize: { widthPt: 612, heightPt: 792 }, margins: { topPt: 72, rightPt: 72, bottomPt: 72, leftPt: 72 } },
77
+ node: {
78
+ kind: "section",
79
+ pageSize: { widthPt: 612, heightPt: 792 },
80
+ margins: { topPt: 72, rightPt: 72, bottomPt: 72, leftPt: 72 },
81
+ },
76
82
  children: [
77
- { node: { kind: 'paragraph', headingLevel: 1, runs: [{ text: 'Heading' }] }, children: [] },
78
- { kind: 'paragraph', runs: [{ text: 'Body.' }] },
83
+ {
84
+ node: {
85
+ kind: "paragraph",
86
+ headingLevel: 1,
87
+ runs: [{ text: "Heading" }],
88
+ },
89
+ children: [],
90
+ },
91
+ { kind: "paragraph", runs: [{ text: "Body." }] },
79
92
  ],
80
93
  },
81
94
  ],
@@ -83,16 +96,19 @@ const pkg = DocumentPackageSchema.parse({
83
96
 
84
97
  // Once a layout pass has fused rendered positions onto the tree's own nodes (each via its own `frames` array)
85
98
  // and reported each page's own size, `pages` is populated to match:
86
- const laidOut = DocumentPackageSchema.parse({ ...pkg, pages: [{ widthPt: 612, heightPt: 792 }] });
99
+ const laidOut = DocumentTreeSchema.parse({
100
+ ...pkg,
101
+ pages: [{ widthPt: 612, heightPt: 792 }],
102
+ });
87
103
  ```
88
104
 
89
105
  ## The package tree
90
106
 
91
- `DocumentPackage` ([#20](https://github.com/ExaDev/document-schema.js/issues/20)) is the promoted single hierarchical artefact — one tree where 3.x carried `{ formatVersion, content, pages }` with the content flat. The tree's vocabulary is defined in `src/package-node.ts` and was proven first as [document-outline.js](https://github.com/ExaDev/document-outline.js)'s phase-1 `decompose`/`flatten` implementation ([document-outline.js#2](https://github.com/ExaDev/document-outline.js/issues/2)); this package's schemas are that shape's schema-home port, matching it node for node:
107
+ `DocumentTree` ([#20](https://github.com/ExaDev/document-schema.js/issues/20)) is the promoted single hierarchical artefact — one tree where 3.x carried `{ formatVersion, content, pages }` with the content flat. The tree's vocabulary is defined in `src/package-node.ts` and was proven first as [document-outline.js](../document-outline.js/README.md)'s phase-1 `decompose`/`flatten` implementation ([document-outline.js#2](https://github.com/ExaDev/document-outline.js/issues/2)); this package's schemas are that shape's schema-home port, matching it node for node:
92
108
 
93
109
  - **Groups** are `{ node, children }` where `node` embeds either an anchor paragraph (heading groups and list-item groups carry the full `ContentParagraph` — runs, formatting, frames — never a projected text label) or a container descriptor: `{ kind: 'section', pageSize, margins }`, `{ kind: 'slide', size, notes }`, `{ kind: 'sheet', name, cells, columns, rows, printSettings }`, `{ kind: 'drawPage', size }`, each tagged with a `kind` the flat container type does not carry, or a shape group's untagged frame descriptor, or — since 4.1.0 — a **construct descriptor** (see [Fidelity constructs](#fidelity-constructs)).
94
110
  - **Bare leaves** carry their own `kind` and never `children`. Discrimination is structural on `node`+`children`, not on the presence of a `kind`. Every `ContentBlock` kind is a legal leaf except the two construct boundary markers, which are the flat form's encoding of something the tree carries as a group (see [Constructs in the flat form](#constructs-in-the-flat-form)).
95
- - **Section groups are mandatory** — one per `ContentSection` — because a section carries pre-layout page geometry (`pageSize`/`margins`) that a rendered `pages` array cannot hold.
111
+ - **Section groups are mandatory** — one per `ContentSection` — because a section carries pre-layout page geometry (`pageSize`/`margins`, plus the optional `breakType` naming how the section begins — nextPage/continuous/evenPage/oddPage, absent meaning the producer's own default) and its optional page furniture (`headers`/`footers`, per-slot block flows in WordprocessingML's own default/even/first reference vocabulary — ExaDev/documents.js#1128) that a rendered `pages` array cannot hold.
96
112
  - **Grouping never crosses container boundaries**: a shape is its own group with its inner blocks grouped inside it (never a slide's paragraphs flattened across its shapes — that is a TOC projection, not a decomposition); a sheet's grid rides on the sheet node with images and embedded documents as children; embedded documents stay intact as one leaf.
97
113
  - **Style refs ride on group wrappers only** — a group may carry `style: string` naming a `styles` table entry; `ContentDocument` nodes carry no ref field, so the flat codec-exchange form is always fully materialised.
98
114
 
@@ -109,24 +125,29 @@ The codecs do not change: they keep producing flat `ContentDocument`s (their nat
109
125
  The transform between the two encodings lives in this package, alongside the schemas that define them:
110
126
 
111
127
  ```ts
112
- import { assemblePackage, decompose, factorStyles, flattenPackage } from 'document-schema.js';
128
+ import {
129
+ assembleTree,
130
+ decompose,
131
+ factorStyles,
132
+ flattenTree,
133
+ } from "document-schema.js";
113
134
 
114
135
  // The one call a construction site makes: decompose the flat content into the tree, splice the envelope
115
136
  // onto the root, and mint a styles table over the result. `pages` is optional -- pass it once a layout
116
137
  // pass has produced each rendered page's own size.
117
- const pkg = assemblePackage(content, pages);
138
+ const pkg = assembleTree(content, pages);
118
139
 
119
140
  // The inverse, with every style ref resolved away: a fully materialised, ref-free ContentDocument.
120
- const flat = flattenPackage(pkg);
141
+ const flat = flattenTree(pkg);
121
142
 
122
143
  // The two halves on their own, for a caller composing its own boundary.
123
- const children = decompose(content); // flat -> the tree `children` a package carries
124
- const reminted = factorStyles(pkg); // re-mint an already-assembled tree (idempotent)
144
+ const children = decompose(content); // flat -> the TreeChildren a DocumentTree's children field carries
145
+ const reminted = factorStyles(pkg); // re-mint an already-assembled tree (idempotent)
125
146
  ```
126
147
 
127
148
  `decompose` throws `ConstructMarkerImbalanceError` — carrying `src/content.ts`'s own `ConstructMarkerImbalance` payload, so a caller narrows with `instanceof` and reads the offending block index rather than parsing a message — when a container's `constructStart`/`constructEnd` markers do not pair up. Promotion is defined only over a balanced stream, so an unbalanced one is refused rather than repaired into a plausible tree.
128
149
 
129
- **This is a deliberate amendment to the "schemas only" charter, not a drift from it.** The transform is not business logic and not format-specific behaviour: it is the canonical, purely structural, zero-I/O relationship between the two shapes this package already defines, and its correctness contract *is* the three laws above. It lives here because it has to: `ooxml.js`, `odf.js`, `markdown-codec`, and `pdf-codec` all depend on this package and none of them depends on `documents.js`, so a codec whose public read/write functions speak `DocumentPackage` directly can only reach the transform if the transform sits at or below the schema layer. Everything the charter actually guards against — XML, ZIP, PDF, fonts, layout, bytes, filesystem — remains firmly out.
150
+ **This is a deliberate amendment to the "schemas only" charter, not a drift from it.** The transform is not business logic and not format-specific behaviour: it is the canonical, purely structural, zero-I/O relationship between the two shapes this package already defines, and its correctness contract _is_ the three laws above. It lives here because it has to: `ooxml.js`, `odf.js`, `markdown-codec`, and `pdf-codec` all depend on this package and none of them depends on `documents.js`, so a codec whose public read/write functions speak `DocumentTree` directly can only reach the transform if the transform sits at or below the schema layer. Everything the charter actually guards against — XML, ZIP, PDF, fonts, layout, bytes, filesystem — remains firmly out.
130
151
 
131
152
  The laws are pinned in `src/bijection.test.ts` over a corpus spanning every document kind, every leaf the tree vocabulary admits, and every grouping signal `decompose` reads (headings, list levels, and construct boundaries in each block flow that admits them). `documents.js` runs the same law harness over its own real-format corpus — reader output for every format it supports, editor builds, and conversion captures carrying a layout pass's real frames and pages — which is the complement this package cannot host, since every reader in it belongs to a package that depends on this one.
132
153
 
@@ -138,22 +159,51 @@ In the tree, a construct is a group like any other: `{ node: <descriptor>, child
138
159
 
139
160
  ```ts
140
161
  // A tracked insertion inside a docx content control, and a footnote marker whose body lives in the definitions table.
141
- const pkg = DocumentPackageSchema.parse({
142
- kind: 'wordprocessing',
162
+ const pkg = DocumentTreeSchema.parse({
163
+ kind: "wordprocessing",
143
164
  metadata: {},
144
- definitions: { n1: { kind: 'footnote', blocks: [{ kind: 'paragraph', runs: [{ text: 'The note body.' }] }] } },
165
+ definitions: {
166
+ n1: {
167
+ kind: "footnote",
168
+ blocks: [{ kind: "paragraph", runs: [{ text: "The note body." }] }],
169
+ },
170
+ },
145
171
  children: [
146
172
  {
147
- node: { kind: 'section', pageSize: { widthPt: 612, heightPt: 792 }, margins: { topPt: 72, rightPt: 72, bottomPt: 72, leftPt: 72 } },
173
+ node: {
174
+ kind: "section",
175
+ pageSize: { widthPt: 612, heightPt: 792 },
176
+ margins: { topPt: 72, rightPt: 72, bottomPt: 72, leftPt: 72 },
177
+ },
148
178
  children: [
149
179
  {
150
- node: { kind: 'contentControl', controlType: 'richText', tag: 'ClientBlock', lock: 'container' },
180
+ node: {
181
+ kind: "contentControl",
182
+ controlType: "richText",
183
+ tag: "ClientBlock",
184
+ lock: "container",
185
+ },
151
186
  children: [
152
187
  {
153
- node: { kind: 'provenance', change: 'insertion', author: 'A. Reviewer', dateIso: '2026-08-18T09:00:00Z' },
154
- children: [{ kind: 'paragraph', runs: [{ text: 'Inserted sentence.' }] }],
188
+ node: {
189
+ kind: "provenance",
190
+ change: "insertion",
191
+ author: "A. Reviewer",
192
+ dateIso: "2026-08-18T09:00:00Z",
193
+ },
194
+ children: [
195
+ { kind: "paragraph", runs: [{ text: "Inserted sentence." }] },
196
+ ],
197
+ },
198
+ {
199
+ node: {
200
+ kind: "anchor",
201
+ anchorType: "footnote",
202
+ name: "1",
203
+ definition: "n1",
204
+ },
205
+ children: [],
155
206
  },
156
- { node: { kind: 'anchor', anchorType: 'footnote', name: '1', definition: 'n1' }, children: [] },
157
207
  ],
158
208
  },
159
209
  ],
@@ -164,53 +214,153 @@ const pkg = DocumentPackageSchema.parse({
164
214
 
165
215
  The six kinds (`src/construct.ts`):
166
216
 
167
- | Kind | Carries | Where it comes from |
168
- | --- | --- | --- |
169
- | `contentControl` | `controlType`, `tag`, `alias`, `lock`, `value`, `checked`, `options` | docx block and inline SDTs, docx legacy `w:ffData` form fields, ODF `office:forms` controls and TOC/index wrappers, PDF AcroForm widgets and their field tree |
170
- | `field` | `instruction`, `cachedResult` | docx `w:fldChar`/`w:instrText` and `w:fldSimple`, ODF field masters and simple fields, ODF cross-reference displays, pptx `a:fld` |
171
- | `anchor` | `anchorType`, `name`, `definition` | docx bookmarks, comment extents, footnote/endnote references; ODF `text:bookmark`, `text:reference-mark*`, `text:note`, `office:annotation`; PDF sticky notes and markup annotations; markdown footnote markers |
172
- | `link` | `target` (external URI or internal anchor name), `title` | docx `@w:anchor`, pptx slide jumps, PDF `GoTo`/`/Dest` and link annotations, markdown link/image titles |
173
- | `provenance` | `change`, `author`, `dateIso` | docx `w:ins`/`w:del` and move tracking, ODF `text:tracked-changes`/`text:changed-region` |
174
- | `division` | `name`, `columnCount`, `protected`, `source` | ODF `text:section` and `text:section-source`, tagged PDF `/Sect` and `/Div` |
217
+ | Kind | Carries | Where it comes from |
218
+ | ---------------- | -------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
219
+ | `contentControl` | `controlType`, `tag`, `alias`, `lock`, `value`, `checked`, `options` | docx block and inline SDTs, docx legacy `w:ffData` form fields, ODF `office:forms` controls and TOC/index wrappers, PDF AcroForm widgets and their field tree |
220
+ | `field` | `instruction`, `cachedResult` | docx `w:fldChar`/`w:instrText` and `w:fldSimple`, ODF field masters and simple fields, ODF cross-reference displays, pptx `a:fld` |
221
+ | `anchor` | `anchorType`, `name`, `definition` | docx bookmarks, comment extents, footnote/endnote references; ODF `text:bookmark`, `text:reference-mark*`, `text:note`, `office:annotation`; PDF sticky notes and markup annotations; markdown footnote markers |
222
+ | `link` | `target` (external URI or internal anchor name), `title` | docx `@w:anchor`, pptx slide jumps, PDF `GoTo`/`/Dest` and link annotations, markdown link/image titles |
223
+ | `provenance` | `change`, `author`, `dateIso` | docx `w:ins`/`w:del` and move tracking, ODF `text:tracked-changes`/`text:changed-region` |
224
+ | `division` | `name`, `columnCount`, `protected`, `linked`, `source` | ODF `text:section` and `text:section-source`, tagged PDF `/Sect` and `/Div` |
175
225
 
176
- Four things bound the vocabulary, and each is a decision rather than an omission:
226
+ Five things bound the vocabulary, and each is a decision rather than an omission:
177
227
 
178
- - **Extents are block-scoped.** A construct group wraps the block flow of a section, heading group, shape, or list item. It cannot wrap a sub-sequence of one paragraph's runs, because a run-level extent is not expressible without changing `ContentParagraph`, and 4.1.0 changes no content shape. Run-level constructs keep their existing homes: **an external hyperlink stays on `ContentRun.hyperlink`** — `link` groups are for block-scoped and annotated extents a flat run field cannot express, never a replacement for it — and the inline field, bookmark, and tracked-change cases wait on a run-level mechanism rather than being forced into a wrapper that would split their paragraph.
228
+ - **Block-scoped extents are groups; run-scoped extents are a paragraph field.** A construct group wraps the block flow of a section, heading group, shape, or list item — and since the run-level extent mechanism landed ([#741](https://github.com/ExaDev/documents.js/issues/741)), a construct covering a sub-sequence of one paragraph's runs is a `RunConstructExtent` on that paragraph's own optional `constructs` field: a descriptor plus a half-open `startRun`/`endRun` range, so the paragraph it sits in is never split to host a wrapper (see [Run-level construct extents](#run-level-construct-extents) below). One occurrence, one scope, one encoding: a construct bracketing whole blocks is a marker pair in the flat form and a group in the tree; a construct covering a run sub-sequence is a field on the paragraph in both. An external hyperlink stays on `ContentRun.hyperlink` regardless — `link` groups are for block-scoped and annotated extents a flat run field cannot express, never a replacement for it.
229
+ - **Crossing and boundary-straddling block extents are a ratified drop.** Two block-scoped constructs whose extents cross (the first ends inside the second, the second inside the first) have no encoding in either form, structurally: the tree states a construct as a group and no tree holds crossing subtrees, while the flat form's bracket matching re-pairs a crossing couple into a different nesting than the source meant. The id-keyed pairing that could express them (WordprocessingML's own `w:id`) is exactly what the marker contract refuses — an id has no home on the tree side and no deterministic way back through `flatten`. The same holds for an extent straddling a block-list boundary (a section break, a table cell's wall), because each block list is its own bracket scope and cross-list pairing is ids again. Within one paragraph, crossing extents _are_ encodable — run ranges are data, not brackets — so this ratification covers block scope only; a codec reading a crossing or straddling pair drops the crossing extent and keeps the properly nested one.
179
230
  - **Two group variants, one per block flow.** `SectionConstructGroup` sits in a section's or heading group's flow and admits heading children; `ShapeConstructGroup` sits in a shape's or list item's and does not — exactly the `SectionChild`/`ShapeChild` split that already existed. A construct nests in and around every other group, so a `provenance` wrapper inside a `contentControl` inside a `division` is a legal (and real) docx shape. Constructs are **not** legal as direct children of a slide, sheet, drawing page, or the package root: those hold containers and leaves, not block flow. The flat form's marker pair follows the same rule by construction: it is a `ContentBlock`, so it can only appear where block flow already runs.
180
231
  - **`division` is first-class, not degraded.** [#24](https://github.com/ExaDev/document-schema.js/issues/24) posed ODF `text:section` as a choice between a new generic kind and degrading to `contentControl` with the specifics in residue. It is first-class, on the odf inventory's own recommendation: `ContentSection` cannot host it (that is page geometry, one `pageSize`/`margins` pair, and it does not nest, while a division nests arbitrarily and usually changes no page geometry at all), and burying a structural container in the form-control vocabulary would make `contentControl` mean two unrelated things. It clears #22's no-format-specific-kinds bar on a real analogue — tagged PDF's `/Sect` and `/Div` are the same construct — not on ODF's say-so. It is spelled `division` rather than `section` because `{ kind: 'section' }` is already the page-geometry container descriptor.
181
- - **Residue is not here.** #22's channel 2 — a per-node and package-level `source: { format, xml }` facility for what has no cross-format meaning — spans the whole content model and remains open on #22. Descriptors are closed objects with no escape hatch, deliberately: a descriptor-only residue field would mint exactly the parallel shape the general facility exists to avoid.
232
+ - **Residue rides the descriptors, spelt identically to everywhere else.** #22's channel 2 has landed (see [The residue channel](#the-residue-channel)): every descriptor, `division` included, carries the same optional `source: { format, xml }` every content node carries, because a descriptor is the construct group's node payload — a node position, not a special case. A matched marker pair moves it across the flat/tree boundary inside the descriptor the open marker already embeds, so the markers themselves stay bare and the construct group keeps its strict `{ node, style, children }` shape. `division` was the one refusal until this major: its `source` already named the external-chapter link (`DivisionSource`, landed 4.1.0), one name could not mean two facts, and renaming the landed field wanted a major. [#743](https://github.com/ExaDev/documents.js/issues/743) does exactly that — the link now rides `linked`, and `source` carries division's own residue rows (ODF `text:filter-name`) like every other descriptor.
182
233
 
183
234
  ### Constructs in the flat form
184
235
 
185
236
  The tree has a wrapper node to hang an extent off; the flat `ContentDocument` does not — a section's, shape's, or table cell's content is one block list and nothing else. So the flat encoding of a construct is a **matched pair of boundary markers** bracketing the blocks it spans, added to `ContentBlock` as two new kinds (`src/content.ts`):
186
237
 
187
238
  ```ts
188
- import type { ContentBlock } from 'document-schema.js';
239
+ import type { ContentBlock } from "document-schema.js";
189
240
 
190
241
  // The flat form of the same construct region the tree example above carries as nested groups: a tracked
191
242
  // insertion and a footnote marker, both inside one content control.
192
243
  const blocks: ContentBlock[] = [
193
- { kind: 'constructStart', descriptor: { kind: 'contentControl', controlType: 'richText', tag: 'ClientBlock', lock: 'container' } },
194
- { kind: 'constructStart', descriptor: { kind: 'provenance', change: 'insertion', author: 'A. Reviewer', dateIso: '2026-08-18T09:00:00Z' } },
195
- { kind: 'paragraph', runs: [{ text: 'Inserted sentence.' }] },
196
- { kind: 'constructEnd' },
197
- { kind: 'constructStart', descriptor: { kind: 'anchor', anchorType: 'footnote', name: '1', definition: 'n1' } },
198
- { kind: 'constructEnd' },
199
- { kind: 'constructEnd' },
244
+ {
245
+ kind: "constructStart",
246
+ descriptor: {
247
+ kind: "contentControl",
248
+ controlType: "richText",
249
+ tag: "ClientBlock",
250
+ lock: "container",
251
+ },
252
+ },
253
+ {
254
+ kind: "constructStart",
255
+ descriptor: {
256
+ kind: "provenance",
257
+ change: "insertion",
258
+ author: "A. Reviewer",
259
+ dateIso: "2026-08-18T09:00:00Z",
260
+ },
261
+ },
262
+ { kind: "paragraph", runs: [{ text: "Inserted sentence." }] },
263
+ { kind: "constructEnd" },
264
+ {
265
+ kind: "constructStart",
266
+ descriptor: {
267
+ kind: "anchor",
268
+ anchorType: "footnote",
269
+ name: "1",
270
+ definition: "n1",
271
+ },
272
+ },
273
+ { kind: "constructEnd" },
274
+ { kind: "constructEnd" },
200
275
  ];
201
276
  ```
202
277
 
203
278
  This is what makes the descriptor vocabulary reachable at all from the shape a codec produces: every codec reads and writes `ContentDocument`, so a construct facility wired only onto the tree is a facility no codec can emit into. `decompose` promotes each matched pair into the construct group the tree already has, and `flatten` emits the pair back.
204
279
 
205
280
  - **Pairing is ordinary bracket matching.** A `constructEnd` closes the nearest preceding still-open `constructStart` in the **same block list**, and the blocks between them are the extent. That is the entire mechanism — there is deliberately **no id, name, or other pairing key** on either marker. An id would have to be minted by whichever producer emitted the pair and then reproduced byte-for-byte by `flatten` to satisfy law 1 above, and a construct group carries a descriptor and its children and nothing else, so the id would be a value with no home on the tree side and no deterministic way back. A bare bracket has nothing to reproduce and nothing to get wrong, and bracket matching already generalises to arbitrary nesting depth and to different construct kinds nested inside each other.
206
- - **Matching never straddles a block list.** A pair opened in a section's blocks closes in that same array; a pair opened inside a table cell closes inside that cell. That cell is also where a construct inside a table is expressible in *either* encoding, because decomposition treats a table as one leaf and never descends into it — so a cell's block list stays flat in a tree too, markers and all.
281
+ - **Matching never straddles a block list.** A pair opened in a section's blocks closes in that same array; a pair opened inside a table cell closes inside that cell. That cell is also where a construct inside a table is expressible in _either_ encoding, because decomposition treats a table as one leaf and never descends into it — so a cell's block list stays flat in a tree too, markers and all.
207
282
  - **An unbalanced list is invalid input, not a shape to repair.** A close with nothing open, or a start still open when the list ends, is malformed: `decompose` throws rather than inventing a boundary. `findConstructMarkerImbalance(blocks)` is the one shared definition of that check — it returns `{ kind: 'unmatchedEnd' | 'unclosedStart', index }` for the first fault or `undefined` when the list balances, and it exists here because a codec emitting pairs, another reading them, and `decompose` promoting them must all agree on exactly one answer, and no schema can express balance (it is a property of a list's sequence, not of any block in it). It is deliberately non-recursive: each block list is its own bracket scope, so a caller walking nested lists calls it once per list. Balance is checked here, not by the schema — `ContentDocumentSchema.parse` still accepts a section whose blocks carry an unmatched marker (pinned by a dedicated test), since a Zod-only refinement expressing balance would validate against a rule the published `content-document.schema.json` fragment cannot express and so would silently diverge from it.
208
283
  - **Balance is necessary but not sufficient — an extent must not cross a heading-group or list-group scope boundary.** Headings and list items have no delimiter of their own in the flat form; a heading paragraph's scope runs until the next paragraph whose `headingLevel` is shallower than or equal to its own, and a list item's scope runs until nesting shallows back out — exactly the nesting `decompose` infers when it builds `HeadingGroupNode`/`ListGroupNode`. A pair whose extent contains a paragraph that closes a scope open before the pair started gives `decompose` no single correct tree: nesting the construct group inside the closing scope strands that paragraph with no legal parent, and closing the scope at the `constructStart` and hoisting the construct group out silently moves everything after that paragraph out of a scope it belonged to. `findConstructMarkerImbalance` cannot see this — balance is a property of the marker pair alone, not of what sits between the two markers — so this package does not check it either; `decompose` is the sole enforcement point, since only it walks the heading/list nesting needed to detect a crossing, and it must reject a crossing extent exactly as it already rejects an unbalanced one, rather than silently picking between the two divergent trees. A producer (`ooxml.js`, `odf.js`, `markdown-codec`, `pdf-codec`) must never open a marker pair inside a heading or list scope that some other block inside the extent goes on to close.
209
- - **A marker carries nothing but its kind and (on the open half) its descriptor.** No `frames` or `sourcePath` — a boundary renders nothing, occupies no space, and has no position — and no `style` ref, since refs are tree-only and a construct group's own ref never resolves onto the construct anyway; it only extends the chain passed to its children, which already carry their own resolved properties.
210
- - **The tree refuses markers at leaf positions.** `PackageBlockLeaf` (`src/package-node.ts`) is `ContentBlock` minus the two marker kinds, and every block-flow child position uses it. A construct is a *group* in the tree; admitting the marker pair there as well would put one fact in two encodings inside one tree, and `decompose(flatten(x)) === x` could never hold for it. Table cells are not an exception to this — a cell's blocks are flat in both encodings, so nothing there crosses the boundary.
284
+ - **A marker carries nothing but its kind and (on the open half) its descriptor.** No `frames` or `sourcePath` — a boundary renders nothing, occupies no space, and has no position — and no `style` ref, since refs are tree-only and a construct group's own ref never resolves onto the construct anyway; it only extends the chain passed to its children, which already carry their own resolved properties. A construct's residue is not a marker field either: it rides _inside_ the descriptor the open marker embeds, the same `source` field that descriptor carries as a tree node.
285
+ - **The tree refuses markers at leaf positions.** `TreeBlockLeaf` (`src/package-node.ts`) is `ContentBlock` minus the two marker kinds, and every block-flow child position uses it. A construct is a _group_ in the tree; admitting the marker pair there as well would put one fact in two encodings inside one tree, and `decompose(flatten(x)) === x` could never hold for it. Table cells are not an exception to this — a cell's blocks are flat in both encodings, so nothing there crosses the boundary.
286
+
287
+ ### Run-level construct extents
288
+
289
+ A construct whose extent is a **sub-sequence of one paragraph's runs** — an inline field, a mid-paragraph bookmark, a comment reference site, a footnote reference marker — is not a marker pair and not a group; it is an entry on the paragraph it sits inside ([#741](https://github.com/ExaDev/documents.js/issues/741)):
290
+
291
+ ```ts
292
+ import type { ContentParagraph } from "document-schema.js";
293
+
294
+ // A bookmark over the middle two runs of four, and a point footnote reference at the paragraph's end.
295
+ const paragraph: ContentParagraph = {
296
+ kind: "paragraph",
297
+ runs: [
298
+ { text: "before " },
299
+ { text: "marked " },
300
+ { text: "words" },
301
+ { text: " after" },
302
+ ],
303
+ constructs: [
304
+ {
305
+ descriptor: { kind: "anchor", anchorType: "bookmark", name: "midway" },
306
+ startRun: 1,
307
+ endRun: 3,
308
+ },
309
+ {
310
+ descriptor: {
311
+ kind: "anchor",
312
+ anchorType: "footnote",
313
+ name: "1",
314
+ definition: "n1",
315
+ },
316
+ startRun: 4,
317
+ endRun: 4,
318
+ },
319
+ ],
320
+ };
321
+ ```
322
+
323
+ The shape is an **extent array, not markers spliced into the runs**, and the choice is load-bearing rather than stylistic:
324
+
325
+ - **It is additive in both runtime and type space.** An optional field on `ContentParagraph` parses unchanged for a pre-4.5.0 document and compiles unchanged for a consumer that never reads it; widening the runs union to carry markers would break every runs consumer the way the `ContentBlock` marker addition broke exhaustive switchers — and `ContentRunSchema` is reused by sheet cells, which have no use for run-level markers.
326
+ - **The bijection laws extend by the existing embed discipline alone.** A paragraph is atomic to decomposition — a bare leaf, or a heading/list group's anchor; its runs are never regrouped — so `decompose`, `flattenTree`, and minting's strip-copy carry the field on the same node object, exactly the way a table cell's own markers ride through untouched. `bijection.test.ts` pins all three laws over every placement (leaf, heading anchor, list anchor, table cell, and minted paragraphs) with no transform change.
327
+ - **Ranges are data, not brackets, so two extents may cross freely** — WordprocessingML's own bookmarks overlap by `w:id` with no imposed nesting, and the run-level mechanism encodes that reality rather than dropping it. This is also precisely why the block-level crossing ratification above cannot be lifted by reusing this mechanism at block scope: flat block indices do not survive the tree's heading/list regrouping, so a block-range array would be a fact only one encoding can read.
328
+
329
+ Well-formedness is a per-entry range check — `0 <= startRun <= endRun <= runs.length` — stated by one shared helper, `findRunConstructFault(paragraph)`, the run-level twin of `findConstructMarkerImbalance`: it returns `{ kind: 'invertedRange' | 'beyondRuns', index }` for the first faulty entry or `undefined` when every extent names real runs. As with marker balance, the schema deliberately does not enforce it (`ContentParagraphSchema.parse` accepts an inverted range, pinned by a dedicated test): the bound is the paragraph's own `runs.length`, a cross-object fact no single node's schema states, and a Zod refinement would validate against a rule the published `content-document.schema.json` fragment cannot express. A run extent can never cross a heading-group or list-group scope either — it is inside one paragraph by construction — so the no-scope-crossing rule the block markers obey cannot even be phrased here.
211
330
 
212
331
  Two harmonisations the inventories asked for are also deliberately absent. There is **no `fieldType` enum**: `instruction` is required and verbatim, so nothing is lost without one, and #24 asks for exactly one new vocabulary — `link`'s internal targets — rather than one per kind, with the four inventories' own corpus gate still standing before a field-type member set could be frozen honestly. And a **sheet-scoped named range** (xlsx defined names and tables, ODF `table:named-expressions`) is a definitions-table entry, not an `anchor`: a sheet group's children are its images and embedded documents, never a block flow, so there is no extent for an anchor to wrap — which is the odf inventory's own verdict for the identical construct.
213
332
 
333
+ ## The residue channel
334
+
335
+ [#22](https://github.com/ExaDev/document-schema.js/issues/22)'s second channel, landed by [#718](https://github.com/ExaDev/documents.js/issues/718) (`src/source.ts`): a quarantined **`source: { format, xml }`** value for what has no cross-format meaning — a WordprocessingML proofing mark, a custom XML payload, a markdown raw-HTML block's remainder, a PDF XMP packet. `format` is a closed enum naming the producing format (one member per reader the workspace has today), so a consumer can tell residue it may restore from residue it must leave alone without reading the text; `xml` is that format's own serialisation, validated as opaque text and nothing more.
336
+
337
+ It is one field with one shape at every position a producer fact can sit:
338
+
339
+ - **Per-node**, on every content node a format reader produces — the same node set that carries `sourcePath` and `frames`, plus the containers and the formula. In the tree, the containers' descriptors inherit it automatically (omit+extend, `src/package-node.ts`), and `decompose`/`flattenTree` carry it verbatim because they embed node objects rather than copying them — the bijection laws run over it unchanged.
340
+ - **On construct descriptors** (every kind, including `division` since [#743](https://github.com/ExaDev/documents.js/issues/743)): a construct with no cross-format analogue degrades to the nearest semantic kind with its specifics in residue — a docx SDT whose gallery is not the TOC degrades to `richText` with its `w:docPartObj` in the descriptor's `source`. The flat form carries it inside the `constructStart` marker's own descriptor payload, so the markers stay bare and the construct group keeps its strict `{ node, style, children }` shape. `division`'s external-chapter link rides its own `linked` field instead, freeing `source` for residue exactly like every other descriptor.
341
+ - **At the package root**, as a keyed `source` table for whole-package facts no content node owns (an unmapped frontmatter key, an XMP packet, a custom XML store), keyed by the producer's own identifier for what each entry reconstructs — a part path, a named store, `frontmatter`. Its own root field rather than a `definitions` tenant, for the same reason `styles` has one: separate key namespaces, and a consumer reaches residue without filtering kind-tagged entries. Like the other tables it is tree-only, and `factorStyles` re-carries it beside `definitions` and `pages`.
342
+
343
+ **The quarantine contract.** Residue is never semantically interpreted. Within this package that is structural: no module reads, resolves, normalises, factors, or branches on a `source` value — the styles table's strict entry objects reject the key outright, so minting can no more factor residue than it can `frames` or `sourcePath`. Outside this package it is a usage rule: a **same-format writer may re-emit its own residue verbatim** (that re-emission is the restorable tier's whole mechanism, and re-serialising opaque text is not interpreting it), and no consumer derives semantics from it — renders it, converts it into semantic nodes, or lets it change content behaviour.
344
+
345
+ ### The three fidelity tiers
346
+
347
+ "Full fidelity" is three testable tiers, not one ([#22](https://github.com/ExaDev/document-schema.js/issues/22)'s definition, restated now that both channels exist):
348
+
349
+ 1. **Semantic fidelity** — every translatable construct survives as a first-class node of the harmonised vocabulary, so cross-format conversion loses no _meaning_. This is channel 1 alone: source → schema → target never touches residue, because residue names the source format precisely so a cross-format consumer can tell it is not its concern.
350
+ 2. **Restorable fidelity** — same-format re-emission rebuilds the original from semantic nodes plus residue, verified by round-trip tests over each codec's corpus. This is the tier the residue channel exists for: a semantic pivot alone cannot promise it (a degraded gallery name has nowhere to come back from), and byte identity is not required for it.
351
+ 3. **Byte fidelity** — the lossless layer's job (`ooxml.js`'s `decodePackage`/`encodePackage`, `odf.js`'s package model): live views over the raw parts, for any format, forever. No semantic pivot achieves byte identity, and this package does not pretend otherwise — the tiers are stacked, not competing: byte fidelity subsumes restorable, which subsumes semantic.
352
+
353
+ The construct verdicts (semantic / residue / derivable — derivable meaning recomputable and dropped without loss, e.g. pivot caches, statistics parts) are recorded per format in the four codec inventories ([ooxml.js#65](https://github.com/ExaDev/ooxml.js/issues/65), [odf.js#59](https://github.com/ExaDev/odf.js/issues/59), [markdown-codec#63](https://github.com/ExaDev/markdown-codec/issues/63), [pdf-codec#66](https://github.com/ExaDev/pdf-codec/issues/66)).
354
+
355
+ ## Sheet input-validation and conditional-formatting rules
356
+
357
+ [ExaDev/documents.js#758](https://github.com/ExaDev/documents.js/issues/758) promotes `ContentSheet`'s dataValidation and conditionalFormatting rules from ooxml.js's own anchor-cell residue landing (its construct inventory's own corpus gate, which deferred freezing a semantic shape until a real producer file existed to verify it against) to real vocabulary, verified against a real LibreOffice-produced xlsx carrying both rule families:
358
+
359
+ - `dataValidations: ContentSheetDataValidation[]` — a cell range's input constraint (xlsx `dataValidation`, ODF `table:content-validation`): a closed `type` (`whole`/`decimal`/`list`/`date`/`time`/`textLength`/`custom`), the shared `SheetRuleOperator` comparison vocabulary (ECMA-376's own `ST_DataValidationOperator`, reused verbatim by conditionalFormatting's `cellIs` rule below), and the input/error message fields. `formula1`/`formula2` stay raw formula text in whatever spelling the producer used (a literal quoted list, a cell-range reference, a numeric/date literal) — Excel's own formula language has no closed grammar this package could parse without a general formula engine, the same reason `ContentSheetCell.formula` is carried verbatim rather than structurally modelled.
360
+ - `conditionalFormats: ContentSheetConditionalFormat[]` — a range's conditional display rule (xlsx `conditionalFormatting`/`cfRule`, LibreOffice's `calcext:conditional-formats` ODF extension), one discriminated-union member per closed-form ECMA-376 rule type: `cellIs` (the shared operator vocabulary), the text-predicate family (`containsText`/`notContainsText`/`beginsWith`/`endsWith`), the operand-free family (`containsBlanks`/`notContainsBlanks`/`containsErrors`/`notContainsErrors`/`uniqueValues`/`duplicateValues`), `top10`, `aboveAverage`, `timePeriod`, and the three visual-scale rules (`colorScale`/`dataBar`/`iconSet`, whose thresholds and colours are the whole rule rather than a separate trigger/style split). A rule's resulting style, when structured at all, is `ContentSheetConditionalFormatStyle` — limited to the two properties actually observed on a real producer's differential format (font colour, fill background), with everything else a dxf may carry (borders, alignment, font weight) riding its own `source` residue. **`expression` is deliberately not a member**: an arbitrary boolean formula has no closed-form structure to model without a general formula engine, so it is not promoted by this union at all and continues to land through ooxml.js's pre-existing anchor-cell residue mechanism, unchanged.
361
+
362
+ Both arrays are sheet-level, like `images`/`embeddedObjects`, rather than anchored to one cell: a rule's own `ranges: ContentSheetRange[]` names every range it applies to (confirmed against a real file: one rule's own `sqref` can and does carry several ranges, not always one), and two rules may legitimately target overlapping cells (a real file was found carrying both a `dataValidation` and a `conditionalFormatting` rule on the same cell simultaneously).
363
+
214
364
  ## Definitions tables and styles
215
365
 
216
366
  The package root carries a generic definitions-table facility ([#21](https://github.com/ExaDev/document-schema.js/issues/21)): named tables whose entries tree nodes reference by string id. **Styles were the first tenant**; link, footnote, and comment definitions ride the tenant-generic `definitions` table (entries tagged with a `kind` discriminator and an open body, `src/definitions.ts`) alongside it — which is why that generic table exists rather than the facility being shaped around styles.
@@ -223,30 +373,30 @@ The package root carries a generic definitions-table facility ([#21](https://git
223
373
 
224
374
  Each is its own root field rather than three more tenants of `definitions`, for the reason `styles` is its own field despite being the facility's first tenant: separate key namespaces, so a layer and a destination may share a name without colliding. The `kind` discriminator still earns its keep inside each, because each holds more than one tenant — a layers table carries group definitions alongside their configuration, and a destinations table carries named destinations alongside outline entries. Per-tenant entry fields stay the tenant's own, never this package's.
225
375
 
226
- A styles entry carries `{ paragraph?, run? }` sub-objects of **resolved canonical properties only**: paragraph `alignment`/`list`/`spacingBeforePt`/`spacingAfterPt`/`lineSpacing`/`indentLeftPt`/`indentFirstLinePt`, run `bold`/`italic`/`underline`/`strike`/`fontFamily`/`sizePt`/`color`. Never `frames`, never `sourcePath`, never `styleId` (per-node facts — a position is a fact about a node, not a style), never a `basedOn` graph (the table is a dictionary, not a program) — and the ban list is **enforced by schema shape** (strict objects that reject those keys outright), not merely documented.
376
+ A styles entry carries `{ paragraph?, run? }` sub-objects of **resolved canonical properties only**: paragraph `alignment`/`list`/`spacingBeforePt`/`spacingAfterPt`/`lineSpacing`/`indentLeftPt`/`indentFirstLinePt`/`pageBreakBefore`/`pageBreakAfter`, run `bold`/`italic`/`underline`/`strike`/`fontFamily`/`sizePt`/`color`. Never `frames`, never `sourcePath`, never `styleId` (per-node facts — a position is a fact about a node, not a style), never a `basedOn` graph (the table is a dictionary, not a program) — and the ban list is **enforced by schema shape** (strict objects that reject those keys outright), not merely documented.
227
377
 
228
378
  Resolution is one overlay chain — outermost ancestor group's style, each nearer group's style, the node's own direct properties; innermost wins, with the resolved run half applying one level further down as run defaults under each run's own properties. `src/definitions.ts` exports the pure helpers that implement it (`overlayStyleEntries`, `resolveStyleChain`, `applyParagraphStyleProperties`, `applyRunStyleProperties`), and `factorStyles` (`src/factor-styles.ts`) is the deterministic frequency pass that mints entries, factoring repeated property tuples into `s1`, `s2`, … refs — see [The package boundary](#the-package-boundary).
229
379
 
230
380
  Every module is also importable directly — `tsdown` builds one file per source module, and `package.json`'s `"./*"` export makes each individually resolvable:
231
381
 
232
382
  ```ts
233
- import { schemaUriFor } from 'document-schema.js/schema-io';
234
- import { ColorSchema } from 'document-schema.js/color';
383
+ import { schemaUriFor } from "document-schema.js/schema-io";
384
+ import { ColorSchema } from "document-schema.js/color";
235
385
  ```
236
386
 
237
387
  ## Codecs
238
388
 
239
- `ContentCodec` (`src/codec.ts`) is the format-agnostic *interface* a sibling package's docx/pptx/odt/odp/ods/odg/xlsx/markdown codec can implement, so a caller working across formats holds one of these instead of a format-specific function pair:
389
+ `ContentCodec` (`src/codec.ts`) is the format-agnostic _interface_ a sibling package's docx/pptx/odt/odp/ods/odg/xlsx/markdown codec can implement, so a caller working across formats holds one of these instead of a format-specific function pair:
240
390
 
241
391
  ```ts
242
- import type { ContentCodec } from 'document-schema.js';
392
+ import type { ContentCodec } from "document-schema.js";
243
393
 
244
394
  declare const docxCodec: ContentCodec; // read(bytes) -> ContentDocument; write(content) -> bytes -- write is optional
245
395
  ```
246
396
 
247
397
  `ContentCodec.write` is optional (`odf` has a reader but no builder — recovering MathML from glyphs is OCR-adjacent), and the interface is generic over its own `TOptions`. There is no `LayoutCodec` any more: it modelled the one format that produces layout cheaply on read — PDF — and the whole `LayoutDocument` family it described moved to pdf-codec in 4.0.0 (see [pdf-codec#65](https://github.com/ExaDev/pdf-codec/issues/65)).
248
398
 
249
- The interface constructs no `DocumentPackage`; composing one (decomposing a codec's flat `ContentDocument` into the tree) is the caller's job (`documents.js`'s `DOCUMENT_FORMAT_CODECS` registry is the concrete example).
399
+ The interface constructs no `DocumentTree`; composing one (decomposing a codec's flat `ContentDocument` into the tree) is the caller's job (`documents.js`'s `DOCUMENT_FORMAT_CODECS` registry is the concrete example).
250
400
 
251
401
  This package also hosts the **port contracts** a layout engine consumes: `TextMeasurer`/`StyledRun`/`WrappedLine` (`src/text-layout.ts`), `ProvidedFont`/`FontSubstitution` (`src/font-port.ts`), `MathBox`/`MathFontMetrics`/`PositionedFormula` (`src/math-layout.ts`), and `Point` (`src/geometry.ts`).
252
402
 
@@ -255,41 +405,51 @@ This package also hosts the **port contracts** a layout engine consumes: `TextMe
255
405
  Two plain [JSON Schema](https://json-schema.org) files are published — generated from the Zod definitions via [`z.toJSONSchema()`](https://zod.dev/json-schema) at build time (`scripts/generate-json-schemas.mjs`) — for non-TypeScript consumers:
256
406
 
257
407
  ```ts
258
- const documentPackageSchema = require('document-schema.js/schemas/document-package.schema.json');
408
+ const documentTreeSchema = require("document-schema.js/schemas/document-tree.schema.json");
259
409
  // or, from a bundler/toolchain that supports JSON module imports:
260
- import documentPackageSchema from 'document-schema.js/schemas/document-package.schema.json' with { type: 'json' };
410
+ import documentTreeSchema from "document-schema.js/schemas/document-tree.schema.json" with { type: "json" };
261
411
  ```
262
412
 
263
413
  or from any language/tool that can read a file out of `node_modules`:
264
414
 
265
- ```
266
- node_modules/document-schema.js/schemas/document-package.schema.json
415
+ ```text
416
+ node_modules/document-schema.js/schemas/document-tree.schema.json
267
417
  node_modules/document-schema.js/schemas/content-document.schema.json
268
418
  ```
269
419
 
270
- Each file's `$id` is a jsdelivr URL pinned to the exact npm version — immutable and live on publish, and (see [Versioning by `$schema`](#versioning-by-schema)) the version of anything stamped with it. Both files carry the same hand-authored `$defs` block (the same object emitted twice in one generator run, so the copies cannot drift), covering the recursive paragraph/table/embedded-object, MathML, and package-tree node models that Zod's converter cannot express directly; `content-json-schema-defs.ts` holds those fragments, and a regression test compares each fragment that has a real Zod counterpart against a live `z.toJSONSchema()` of that schema so a field changed without updating its fragment fails a test. The one deliberate cross-file `$ref` is the embedded-object cycle back to a whole `ContentDocument`. Fragments downstream of a `z.custom()` node (`ContentBlock` and the tree's marker-free `PackageBlockLeaf`, `ContentTable`/`Cell`/`Row`, `ContentEmbeddedObject(Block)`, the nine package-tree group wrappers, `MathMlNode`/`Element`/`Attribute`, `ContentFormula`, `MathExpression` and its recursive variants) still need hand re-verification — the construct descriptors and the two boundary markers do not, since each is a plain `z.object`/`z.strictObject` reaching no opaque node and is held to the live comparison against `src/content.ts`/`src/package-node.ts`/`src/mathml.ts`/`src/math.ts` — see below.
420
+ Each file's `$id` is a jsdelivr URL pinned to the exact npm version — immutable and live on publish, and (see [Versioning by `$schema`](#versioning-by-schema)) the version of anything stamped with it. Both files carry the same `$defs` block (the same object emitted twice in one generator run, so the copies cannot drift), covering the recursive paragraph/table/embedded-object, MathML, and package-tree node models that Zod's converter cannot express directly; `content-json-schema-defs.ts` holds those fragments — MathML's own (`MathMlAttribute`/`MathMlElement`/`MathMlNode`) computed from a live `z.toJSONSchema()` call rather than hand-authored, everything else still transcribed by hand — and a regression test compares each fragment that has a real Zod counterpart against a live `z.toJSONSchema()` of that schema so a field changed without updating its fragment fails a test. The one deliberate cross-file `$ref` is the embedded-object cycle back to a whole `ContentDocument`. Fragments downstream of a `z.custom()` node (`ContentBlock` and the tree's marker-free `TreeBlockLeaf`, `ContentTable`/`Cell`/`Row`, `ContentEmbeddedObject(Block)`, the nine package-tree group wrappers, `ContentFormula`, `MathExpression` and its recursive variants) still need hand re-verification — the construct descriptor vocabulary (`src/construct.ts`) and the two boundary markers (`src/content.ts`) do not, since each is a plain `z.object`/`z.strictObject` reaching no opaque node and is held to a live comparison against its own real schema. `MathMlNode`/`Element`/`Attribute` need no hand re-verification either, but on different grounds: since [ExaDev/documents.js#937](https://github.com/ExaDev/documents.js/issues/937) they are generated outright from a live `z.toJSONSchema()` call over `src/mathml.ts` rather than transcribed by hand, so there is no hand-authored fragment left for a comparison to catch drifting away from its own schema — run over the generated fragments, that same comparison instead proves only that the generation itself is deterministic and reproducible, and genuine coverage of `src/mathml.ts`'s own field shapes comes from a separate hard-coded expected-shape test instead — see below.
271
421
 
272
422
  ### `z.custom()` vs `z.lazy()` for recursive schemas
273
423
 
274
- `ContentBlockSchema`, `ContentEmbeddedObjectSchema`, the package tree's per-kind group schemas (`src/package-node.ts`), `MathMlNodeSchema`, and `MathExpressionSchema` are `z.custom()` type-guard predicates rather than real Zod schemas, because `z.lazy()` was believed to collapse to `unknown` for recursive children. A throwaway spike (reverted) re-tested `MathMlNodeSchema` (the simplest case) against `zod@4.4.3`.
424
+ The package tree's per-kind group schemas (`src/package-node.ts`) are still `z.custom()` type-guard predicates rather than real Zod schemas, because `z.lazy()` was believed to collapse to `unknown` for recursive children — a genuinely separate recursion axis (tree-of-groups, not a flat block list or a math expression's own grammar) that neither #937 nor #1009 below touched. `ContentBlockSchema`/`ContentEmbeddedObjectSchema` and `MathExpressionSchema` used to share that same z.custom() treatment; both have since been converted for real (see "Landed for #1009" below). A throwaway spike (reverted) first re-tested `MathMlNodeSchema` (the simplest case) against `zod@4.4.3`, and [ExaDev/documents.js#937](https://github.com/ExaDev/documents.js/issues/937) later applied the finding for real.
275
425
 
276
426
  **Finding: `z.lazy()` works now, with one constructional gotcha.** The naive rewrite —
277
427
 
278
428
  ```ts
279
429
  export const MathMlElementSchema: z.ZodType<MathMlElement> = z.object({
280
- type: z.literal('element'),
430
+ type: z.literal("element"),
281
431
  tag: z.string(),
282
432
  attributes: z.array(MathMlAttributeSchema),
283
433
  children: z.lazy(() => z.array(MathMlNodeSchema)),
284
434
  });
285
- export const MathMlNodeSchema: z.ZodType<MathMlNode> = z.discriminatedUnion('type', [
286
- MathMlTextSchema, MathMlCdataSchema, MathMlCommentSchema, MathMlDeclarationSchema, MathMlPiSchema, MathMlElementSchema,
287
- ]);
435
+ export const MathMlNodeSchema: z.ZodType<MathMlNode> = z.discriminatedUnion(
436
+ "type",
437
+ [
438
+ MathMlTextSchema,
439
+ MathMlCdataSchema,
440
+ MathMlCommentSchema,
441
+ MathMlDeclarationSchema,
442
+ MathMlPiSchema,
443
+ MathMlElementSchema,
444
+ ],
445
+ );
288
446
  ```
289
447
 
290
448
  — fails to typecheck: annotating `MathMlElementSchema` as `z.ZodType<MathMlElement>` widens it so `z.discriminatedUnion` (which needs each member's internal `propValues`) rejects it, and dropping the annotation hits TypeScript's circular-inference error. The fix: annotate **only the outer union's binding** (`MathMlNodeSchema`), leaving every member schema unannotated and fully inferred — all tests passed, and `z.toJSONSchema()` produced a real `oneOf` with `{ "$ref": "#" }` at the recursion point.
291
449
 
292
- **This is a tracked follow-up, not carried out here.** Converting `MathMlNodeSchema` for real would let the JSON-schema generator drop its hand-authored `$defs` entries; `ContentBlockSchema`/`ContentEmbeddedObjectSchema` are harder (mutual recursion across table/cell/row, plus the full `ContentDocument` cycle) and were not spiked.
450
+ **Landed for real in [ExaDev/documents.js#937](https://github.com/ExaDev/documents.js/issues/937), with a second gotcha the original spike missed.** `z.ZodType<MathMlNode>` supplies only the Output type parameter, leaving Input at its own default of `unknown` — invisible within this package (`z.infer<>` and every test here read Output alone), so the spike's own from-scratch tests passed. It broke `markdown-codec`'s typecheck instead: `z.codec()`'s `encode()` callback is typed against a schema's _input_, so `markdownCodec`/`markdownContentCodec`'s `encode()` saw `mathml: unknown[]` in the value handed to `writeMarkdown`/`writeMarkdownContent`, not `MathMlNode[]` — caught only by a full-workspace typecheck, since nothing in this package itself calls `z.codec()` over `MathMlNodeSchema`. The fix is annotating both parameters — `z.ZodType<MathMlNode, MathMlNode>` — which is correct rather than merely defensive, since this schema has no transform and Input and Output are genuinely identical. `src/mathml.ts` now defines `MathMlNodeSchema` this way, and `content-json-schema-defs.ts` computes its `$defs.MathMlAttribute`/`MathMlElement`/`MathMlNode` fragments from a small local registry (`z.toJSONSchema` with a `#/$defs/<id>` uri callback, the same cross-reference mechanism `content-json-schema-defs.test.ts`'s own live comparison already used) rather than transcribing them by hand — confirmed to reproduce the identical nested-`$ref` shape the hand-authored fragment used to carry. `ContentFormulaSchema` itself was unaffected at the time: its `mathml` field reached a real schema, but its `content` field still reached the opaque `MathExpressionSchema`, so the whole fragment stayed hand-transcribed regardless (see below for how that changed too).
451
+
452
+ **The harder half landed in [ExaDev/documents.js#1009](https://github.com/ExaDev/documents.js/issues/1009): `ContentBlockSchema`/`ContentEmbeddedObjectSchema` (mutual recursion across table/cell/row, plus the full `ContentDocument` cycle) and `MathExpressionSchema` (mutual recursion across app args, binder bounds/bodies, and matrix rows), the two families #937 explicitly deferred as harder and unspiked.** Both gotchas above applied again, confirmed the identical way: `z.discriminatedUnion` member schemas (`ContentParagraphSchema`, `ContentTableSchema`, `ContentEmbeddedObjectBlockSchema`, ... for `ContentBlockSchema`; `MathAppSchema`, `MathSumSchema`, `MathProdSchema`, `MathMatrixSchema`, ... for `MathExpressionSchema`) stay unannotated, only the outer union binding carries `z.ZodType<T, T>` (both parameters — the Input-defaults-to-`unknown` gotcha bites identically here), and `pnpm exec turbo run _typecheck --affected` was run for real this time, over the whole workspace, both before committing to the design and after landing it. It caught a genuine regression: `rtf-codec`'s `readEmbeddedObjectData` relied on `ContentEmbeddedObjectSchema`'s old `z.custom()` guard never inspecting `source` at all (independently re-validating and dropping it after a successful parse) — once `source` became a real, validated field of the schema itself, a malformed `source` started rejecting the WHOLE embedded object rather than just that one field, caught only by rtf-codec's own test suite once the schema was real. The fix strips an invalid `source` key before validation runs, restoring the original lenient contract rather than the newly-strict one. A separate, unrelated bug surfaced from the same real-corpus bijection gate (`documents.js`'s `bijection.test.ts`, run against actual reader output for the first time `ContentBlockSchema` was capable of validating anything): `ContentTableSchema.columnWidthsPt` was `z.array(z.number().positive())`, but both ooxml.js's and odf.js's own table readers deliberately default an unresolvable column's own width to `0` — a real shape neither `isContentBlock`'s own old hand-written guard (which never checked positivity at all) nor anything else had ever actually validated against a live document before. Loosened to `.nonnegative()`. `content-json-schema-defs.ts`'s own `ContentBlock`/`ContentTableCell`/`ContentTableRow`/`ContentTable`/`ContentFormula`/`MathExpression`/`MathApp`/`MathSum`/`MathProd`/`MathMatrix` fragments stay hand-transcribed (unlike `MathMlNode`'s own computed-`get`-accessor treatment) but are now held to `content-json-schema-defs.test.ts`'s live `z.toJSONSchema()` comparison, registered alongside their own already-covered sibling fragments so cross-references resolve correctly — all ten matched byte-for-byte on the first attempt. `ContentEmbeddedObjectSchema`/`ContentEmbeddedObjectBlockSchema` are the one exception: their `document` field's genuine cross-file cycle back to `ContentDocumentSchema` produces an anonymous, Zod-internal `#/$defs/__shared#/$defs/schemaN`-shaped ref once `ContentDocumentSchema` is registered alongside that test's ~90 other entries (confirmed working correctly in an isolated two-schema registry, so this is a real, narrow interaction specific to a large combined registry, not a sign either schema is wrong) — both stay registered (so `ContentBlockSchema`'s own union member `$ref` still resolves correctly) but excluded from the direct comparison, a gap left open for a future, narrower investigation into Zod's own multi-schema cycle handling rather than blocking #1009 on it. The package tree's own nine group wrappers (`src/package-node.ts`) remain untouched — a genuinely separate recursion axis (tree-of-groups, not a flat block list or a math expression's own grammar), out of #1009's own stated scope.
293
453
 
294
454
  ### Versioning by `$schema`
295
455
 
@@ -298,35 +458,44 @@ There is no `formatVersion` field anywhere in a 4.0.0 dump. A serialised value s
298
458
  - a URI from the **same major** as the installed release parses (patch and minor releases are semver-compatible with their major's schema generation);
299
459
  - an **older major's** URI throws `SchemaVersionMismatchError` naming the change — pre-4.0.0 dumps carry the retired `formatVersion` field and the flat `{ formatVersion, content, pages }` package shape, replaced by the tree form ([#20](https://github.com/ExaDev/document-schema.js/issues/20));
300
460
  - a **newer major's** URI throws the same error with the upgrade pointer;
301
- - a **layout-document** URI (any release) throws `LayoutSchemaDemotedError` pointing at pdf-codec — the demotion tombstone.
461
+ - a **layout-document** URI (any release) throws `LayoutSchemaDemotedError` pointing at pdf-codec — the demotion tombstone;
462
+ - a **document-package** URI (any release) throws `DocumentPackageRenamedError` pointing at `DocumentTree` — the rename tombstone for [#661](https://github.com/ExaDev/documents.js/issues/661).
302
463
 
303
- A bare `DocumentPackageSchema.parse(value)` does **not** version-discriminate — it structurally validates whatever it is handed against the installed schema, full stop — so a caller ingesting a dump it did not itself produce must go through `documentFromJson`, not a direct parse. `documentSchemaKindOf(value)` still answers "which kind does this URI name" version-agnostically without parsing. Content hashes and structural comparisons over serialised dumps must exclude `$schema` — it is envelope metadata that names the dumper, not content; two dumps of one document by two releases hash equal once it is excluded.
464
+ A bare `DocumentTreeSchema.parse(value)` does **not** version-discriminate — it structurally validates whatever it is handed against the installed schema, full stop — so a caller ingesting a dump it did not itself produce must go through `documentFromJson`, not a direct parse. `documentSchemaKindOf(value)` still answers "which kind does this URI name" version-agnostically without parsing. Content hashes and structural comparisons over serialised dumps must exclude `$schema` — it is envelope metadata that names the dumper, not content; two dumps of one document by two releases hash equal once it is excluded.
304
465
 
305
466
  ### Self-describing JSON
306
467
 
307
- `documentPackageWithSchema`/`contentDocumentWithSchema` each stamp a `$schema` property pointing at the `.schema.json` file for the currently installed version:
468
+ `documentTreeWithSchema`/`contentDocumentWithSchema` each stamp a `$schema` property pointing at the `.schema.json` file for the currently installed version:
308
469
 
309
470
  ```ts
310
- import { documentPackageWithSchema } from 'document-schema.js';
471
+ import { documentTreeWithSchema } from "document-schema.js";
311
472
 
312
- const tagged = documentPackageWithSchema(pkg);
313
- // { $schema: 'https://cdn.jsdelivr.net/npm/document-schema.js@4.0.0/schemas/document-package.schema.json', kind: 'wordprocessing', metadata: {...}, children: [...] }
314
- writeFileSync('package.json.doc', JSON.stringify(tagged, null, 2));
473
+ const tagged = documentTreeWithSchema(pkg);
474
+ // { $schema: 'https://cdn.jsdelivr.net/npm/document-schema.js@5.0.0/schemas/document-tree.schema.json', kind: 'wordprocessing', metadata: {...}, children: [...] }
475
+ writeFileSync("package.json.doc", JSON.stringify(tagged, null, 2));
315
476
  ```
316
477
 
317
- A caller who already knows the kind can keep using the schemas directly — `DocumentPackageSchema.parse(value)` tolerates and strips an incoming `$schema` (none are `.strict()`). `documentFromJson` is for the "don't yet know the kind or provenance" case, reading `$schema` to decide which schema to run and whether this release may run it:
478
+ A caller who already knows the kind can keep using the schemas directly — `DocumentTreeSchema.parse(value)` tolerates and strips an incoming `$schema` (none are `.strict()`). `documentFromJson` is for the "don't yet know the kind or provenance" case, reading `$schema` to decide which schema to run and whether this release may run it:
318
479
 
319
480
  ```ts
320
- import { documentFromJson, SchemaVersionMismatchError, UnrecognizedDocumentSchemaError } from 'document-schema.js';
481
+ import {
482
+ documentFromJson,
483
+ SchemaVersionMismatchError,
484
+ UnrecognizedDocumentSchemaError,
485
+ } from "document-schema.js";
321
486
 
322
487
  try {
323
- const { kind, value } = documentFromJson(JSON.parse(readFileSync('some-file.json', 'utf8')));
324
- // kind: 'DocumentPackage' | 'ContentDocument'
488
+ const { kind, value } = documentFromJson(
489
+ JSON.parse(readFileSync("some-file.json", "utf8")),
490
+ );
491
+ // kind: 'DocumentTree' | 'ContentDocument'
325
492
  } catch (error) {
326
493
  if (error instanceof UnrecognizedDocumentSchemaError) {
327
- console.error('not a document-schema.js value:', error.schema);
494
+ console.error("not a document-schema.js value:", error.schema);
328
495
  } else if (error instanceof SchemaVersionMismatchError) {
329
- console.error(`dump is @${error.dumpVersion}, installed is @${error.installedVersion}`);
496
+ console.error(
497
+ `dump is @${error.dumpVersion}, installed is @${error.installedVersion}`,
498
+ );
330
499
  }
331
500
  }
332
501
  ```
@@ -335,11 +504,11 @@ try {
335
504
 
336
505
  ## Used by
337
506
 
338
- - [ooxml.js](https://github.com/ExaDev/ooxml.js) — `readDocx`/`readPptx`/`readXlsxContent` return types are typed against this package's schemas, not a local lookalike.
339
- - [odf.js](https://github.com/ExaDev/odf.js) — ODF typed readers return the same shared types, so ODF and OOXML speak the identical pivot.
340
- - [documents.js](https://github.com/ExaDev/documents.js) — primary consumer of `ContentDocument` and `DocumentPackage`; its `DOCUMENT_FORMAT_CODECS` registry implements `ContentCodec` per format, and its conversion pipeline calls `assemblePackage`/`flattenPackage` from here at every package construction site.
341
- - [pdf-codec](https://github.com/ExaDev/pdf-codec) — owns its layout item model outright since 4.0.0; `readPdf`/`writePdf` operate on pdf-codec's own `LayoutDocument`, and this package's `ContentDocument` remains its content pivot.
342
- - [markdown-codec](https://github.com/ExaDev/markdown-codec) — `readMarkdown`/`writeMarkdown` read and write this package's `ContentDocument` directly.
507
+ - [ooxml.js](../ooxml.js/README.md) — `readDocx`/`readPptx`/`readXlsxContent` return types are typed against this package's schemas, not a local lookalike.
508
+ - [odf.js](../odf.js/README.md) — ODF typed readers return the same shared types, so ODF and OOXML speak the identical pivot.
509
+ - [documents.js](https://github.com/ExaDev/documents.js) — primary consumer of `ContentDocument` and `DocumentTree`; its `DOCUMENT_FORMAT_CODECS` registry implements `ContentCodec` per format, and its conversion pipeline calls `assembleTree`/`flattenTree` from here at every package construction site.
510
+ - [pdf-codec](../pdf-codec/README.md) — owns its layout item model outright since 4.0.0; `readPdf`/`writePdf` operate on pdf-codec's own `LayoutDocument`, and this package's `ContentDocument` remains its content pivot.
511
+ - [markdown-codec](../markdown-codec/README.md) — `readMarkdown`/`writeMarkdown` read and write this package's `ContentDocument` directly.
343
512
 
344
513
  None depend on each other for this vocabulary — each depends on `document-schema.js` directly.
345
514
 
@@ -360,9 +529,17 @@ pnpm test:smoke # turbo run _test:smoke -> rebuilds dist/ and schemas/ first,
360
529
 
361
530
  To run a single test file: `pnpm vitest run src/path/to/file.test.ts`.
362
531
 
532
+ ## Release and publishing
533
+
534
+ Release, CI, and commit-message conventions are all workspace-wide, not package-local — see the [monorepo root README](../../README.md#releases) for the mechanism (topological per-package `semantic-release` via `@exadev/semantic-release-workspace`, OIDC trusted npm publishing, automatic sibling dependency-range rewriting) and its [post-release republishing and attestation](../../README.md#releases) note on the restored GitHub Packages mirrors, npm aliases, and SBOM/provenance signing.
535
+
536
+ ## Contributing
537
+
538
+ Conventional Commits, enforced workspace-wide by commitlint through a root `commit-msg` hook. Work inside `packages/document-schema.js/`; see [CONTRIBUTING.md](../../CONTRIBUTING.md) for the shared git hooks and history conventions.
539
+
363
540
  ## npm aliases
364
541
 
365
- This package also publishes under the following alternate npm names — the identical build, same version, republished by CI alongside the primary `document-schema.js` package:
542
+ This package also published under the following alternate npm names from the pre-monorepo pipeline:
366
543
 
367
544
  - [document-content-model](https://www.npmjs.com/package/document-content-model)
368
545
  - [doc-model.js](https://www.npmjs.com/package/doc-model.js)
@@ -370,6 +547,8 @@ This package also publishes under the following alternate npm names — the iden
370
547
  - [document-schema](https://www.npmjs.com/package/document-schema)
371
548
  - [document-model.js](https://www.npmjs.com/package/document-model.js)
372
549
 
550
+ **Frozen since the monorepo migration** — see the [root README's release note](../../README.md#releases): the alias republish step was dropped along with GitHub Packages mirroring and SBOM/provenance signing, and nothing today keeps any of the five in sync with `document-schema.js`'s own releases. Tracked in [ExaDev/documents.js#730](https://github.com/ExaDev/documents.js/issues/730).
551
+
373
552
  ## License
374
553
 
375
554
  MIT