@orkestrel/scaffold 0.0.67 → 0.0.68

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (74) hide show
  1. package/dist/bin/main.js +67 -44
  2. package/dist/bin/main.js.map +1 -1
  3. package/dist/host/agents/templates/brief.md +9 -0
  4. package/dist/host/claude/agents/orkestrel.md +4 -4
  5. package/dist/host/claude/rules/names.md +15 -0
  6. package/dist/host/claude/rules/tests.md +33 -4
  7. package/dist/host/claude/rules/workspace.md +14 -2
  8. package/dist/host/dotfiles/prettierignore +3 -0
  9. package/dist/host/guides/README.md +65 -0
  10. package/dist/host/guides/abort.md +169 -0
  11. package/dist/host/guides/agent.md +1509 -0
  12. package/dist/host/guides/brief.md +1266 -0
  13. package/dist/host/guides/browser.md +2200 -0
  14. package/dist/host/guides/budget.md +196 -0
  15. package/dist/host/guides/codec.md +519 -0
  16. package/dist/host/guides/console.md +785 -0
  17. package/dist/host/guides/contract.md +1193 -0
  18. package/dist/host/guides/csv.md +541 -0
  19. package/dist/host/guides/database.md +2518 -0
  20. package/dist/host/guides/emitter.md +233 -0
  21. package/dist/host/guides/form.md +1791 -0
  22. package/dist/host/guides/html.md +717 -0
  23. package/dist/host/guides/indexeddb.md +505 -0
  24. package/dist/host/guides/interpret.md +1029 -0
  25. package/dist/host/guides/lsp.md +515 -0
  26. package/dist/host/guides/markdown.md +964 -0
  27. package/dist/host/guides/mcp.md +5554 -0
  28. package/dist/host/guides/middleware.md +927 -0
  29. package/dist/host/guides/msg.md +440 -0
  30. package/dist/host/guides/ndjson.md +120 -0
  31. package/dist/host/guides/ollama.md +380 -0
  32. package/dist/host/guides/pool.md +280 -0
  33. package/dist/host/guides/probe.md +1210 -0
  34. package/dist/host/guides/process.md +1620 -0
  35. package/dist/host/guides/program.md +1110 -0
  36. package/dist/host/guides/qualifier.md +854 -0
  37. package/dist/host/guides/queue.md +370 -0
  38. package/dist/host/guides/rater.md +330 -0
  39. package/dist/host/guides/reason.md +1122 -0
  40. package/dist/host/guides/relation.md +373 -0
  41. package/dist/host/guides/router.md +753 -0
  42. package/dist/host/guides/scaffold.md +192 -31
  43. package/dist/host/guides/sea.md +383 -0
  44. package/dist/host/guides/server.md +752 -0
  45. package/dist/host/guides/sqlite.md +330 -0
  46. package/dist/host/guides/sse.md +187 -0
  47. package/dist/host/guides/supervisor.md +4890 -0
  48. package/dist/host/guides/table.md +1556 -0
  49. package/dist/host/guides/template.md +280 -0
  50. package/dist/host/guides/terminal.md +1145 -0
  51. package/dist/host/guides/test.md +2969 -0
  52. package/dist/host/guides/timeout.md +252 -0
  53. package/dist/host/guides/tool.md +311 -0
  54. package/dist/host/guides/toolbox.md +1038 -0
  55. package/dist/host/guides/websocket.md +282 -0
  56. package/dist/host/guides/worker.md +615 -0
  57. package/dist/host/guides/workflow.md +1507 -0
  58. package/dist/host/guides/workspace.md +595 -0
  59. package/dist/host/manifest.json +1218 -10
  60. package/dist/host/tests/policy.test.ts +279 -2
  61. package/dist/host/tests/setupPolicy.ts +437 -6
  62. package/dist/src/core/index.cjs +38 -16
  63. package/dist/src/core/index.cjs.map +1 -1
  64. package/dist/src/core/index.d.cts +33 -9
  65. package/dist/src/core/index.d.ts +33 -9
  66. package/dist/src/core/index.js +37 -17
  67. package/dist/src/core/index.js.map +1 -1
  68. package/dist/src/server/index.cjs +1750 -1567
  69. package/dist/src/server/index.cjs.map +1 -1
  70. package/dist/src/server/index.d.cts +106 -24
  71. package/dist/src/server/index.d.ts +106 -24
  72. package/dist/src/server/index.js +1751 -1570
  73. package/dist/src/server/index.js.map +1 -1
  74. package/package.json +3 -3
@@ -0,0 +1,964 @@
1
+ # Markdown
2
+
3
+ > A types-first markdown layer over `@orkestrel/html`: a linear-time scanner that parses
4
+ > GitHub-Flavored Markdown into a typed AST, a stateful `Markdown` workspace that queries,
5
+ > rewrites, folds, and streams that AST, and standalone projections that carry it out to
6
+ > sanitized HTML or canonical markdown source and carry an HTML AST back in.
7
+
8
+ One parse is the whole contract: every later output is a projection of the AST it produced, never a second read of the source. `parseDocument` runs a block phase (headings / paragraphs / lists / GFM tables / fenced code / blockquotes / thematic breaks) then an inline phase (emphasis / inline code / links / images / hard breaks) over each block's text, and returns a render-agnostic `MarkdownDocument` — a discriminated union of node values keyed by `element` (the axis that varies: never `kind` / `type`). The AST itself is the primary contract — render-agnostic and exhaustively testable — with a from-unknown validation surface (`isInlineNode` / `isBlockNode` / `isMarkdownNode` / `isMarkdownDocument`) for when an AST arrives from outside `parseDocument` (a deserialized document, a value crossing a process/RPC boundary). Source: [`src/core`](../src/core). Surfaced through the `@src/core` barrel.
9
+
10
+ **Each conversion direction lives here**, because what an HTML subtree becomes in markdown — and what a markdown node becomes in HTML — is markdown-format knowledge, not HTML knowledge. `@orkestrel/html` owns the HTML AST, its total parser, its canonical serializer, and its sanitize floor; this package owns the two projections across the boundary and never asks html to know a markdown word. Outbound: `markdownToHTML` projects a `MarkdownNode` onto html's AST, `renderHTML` composes that projection with html's sanitizer and serializer into one sanitized string, and `renderMarkdown` writes canonical markdown source instead (§ [`renderMarkdown` round-trip](#rendermarkdown-round-trip)). Inbound: `htmlToMarkdown` folds an html `HTMLNode` back down to a `MarkdownDocument` (§ [`htmlToMarkdown` projection](#htmltomarkdown-projection)). None of them assumes its input came from a trusted parse, and none of them throws: malformed markdown degrades to literal text, while at the outbound depth cap value-bearing nodes degrade to text and structural nodes degrade to nothing; the inbound trip inherits html's own cap rather than exhausting the call stack (no ReDoS, no stack overflow).
11
+
12
+ ## Surface
13
+
14
+ ### Types
15
+
16
+ The full node shape and workspace contract, from [`types.ts`](../src/core/types.ts). `element` is the discriminant every node carries; block nodes carry document structure, inline nodes carry the inline content of a heading / paragraph / list item / table cell. `MarkdownInterface`'s call-signature members are documented under [`## Methods`](#methods).
17
+
18
+ A `Shape` cell holds an interface's data members as bare names in braces, `?` marking an optional member and `plus` introducing its call-signature members, and a type alias's own type literal with a union's arms escaped as `\|`.
19
+
20
+ | Type | Kind | Shape | Summary |
21
+ | --------------------------- | --------- | --------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
22
+ | `TableAlign` | type | `'left' \| 'right' \| 'center'` | Names the horizontal alignment of a GFM table column, as declared by its delimiter row (`:---` left, `---:` right, `:---:` center). A bare `---` delimiter is represented by `null` in `TableNode.align`: the positional array requires one entry per column, JSON cannot carry `undefined` in an array, and the bare delimiter is an explicit no-alignment marker rather than an omitted value. |
23
+ | `ListItemMatch` | interface | `{ ordered, start, content, indent, marker }` | Represents the parsed parts of a single list-item line — the value the block phase's list detector returns for a `-` / `*` / `+` bullet or a `1.` / `1)` ordinal line. |
24
+ | `HeadingMatch` | interface | `{ level, text, offset }` | Represents the parsed parts of a single ATX heading line — the value the block phase's heading detector returns for a `#` … `######` line. |
25
+ | `FenceMatch` | interface | `{ marker, lang }` | Represents the parsed parts of a fenced-code opening line — the value the block phase's fence detector returns for a \`\`\`\` \`\`\` \`\`\`\` or \`~~~\` opener. |
26
+ | `CodeSpanMatch` | interface | `{ value, end }` | Represents the located extent of one inline code span — the value the inline phase's code scanner returns for a matched backtick run. |
27
+ | `LinkBounds` | interface | `{ close, end }` | Represents the located syntax bounds of one `[text](href)` link — the value the inline phase's link locator returns for a balanced label followed by a destination. |
28
+ | `EmphasisBounds` | interface | `{ strong, open, close, end }` | Represents the located content and syntax bounds of one emphasis run — the value the inline phase's emphasis locator returns for a matched marker run. |
29
+ | `LinkScan` | interface | `{ node, end }` | Represents the scanned result of one `[text](href)` link — the node the inline phase's link scanner built from `LinkBounds` and where the scan resumes. |
30
+ | `EmphasisScan` | interface | `{ node, end }` | Represents the scanned result of one emphasis run — the node the inline phase's emphasis scanner built from `EmphasisBounds` and where the scan resumes. |
31
+ | `TableCollection` | interface | `{ node, next }` | Represents the result of collecting one GFM table — the node the construct scanner built and where the block phase resumes. |
32
+ | `ListCollection` | interface | `{ node, next }` | Represents the result of collecting one list — the node the construct scanner built and where the block phase resumes. |
33
+ | `MarkdownSpan` | interface | `{ start, end }` | Addresses a half-open region of the original markdown string, in UTF-16 code units — `start` inclusive, `end` exclusive. The provenance a parse records for a node and `MarkdownInterface.span` reads back. |
34
+ | `MarkdownSegment` | interface | `{ offset, start, end }` | Maps one run of a `MarkdownSource` back to the region of the original markdown string it was taken from. |
35
+ | `MarkdownSource` | interface | `{ text, segments }` | Pairs a piece of derived markdown text with the runs mapping it back to the original string — what `splitLines` returns per line, so every phase downstream of it keeps original coordinates instead of reconstructing them from node values. |
36
+ | `TextNode` | interface | `{ element, value }` | Represents a run of plain text — the leaf inline node. `value` is the decoded text with markdown escapes (`\*`, `\_`, …) already resolved to their literal characters; html's text encoder escapes `&`, `<`, `>` on the way out; `"` and `'` stay literal in character data. |
37
+ | `EmphasisNode` | interface | `{ element, strong, children }` | Represents emphasized inline content — `*italic*` / `_italic_` (`strong: false`) or `**bold**` / `__bold__` (`strong: true`). `children` are the nested inline nodes, so emphasis composes (a `**bold _and italic_**` is a strong node wrapping a text node and an emphasis node). |
38
+ | `CodeSpanNode` | interface | `{ element, value }` | Represents an inline code span — \`\` \`code\` \`\`. \`value\` is the verbatim span text; no inner markdown is parsed (code is literal), and the renderer HTML-escapes it inside a \`<code>\` element. |
39
+ | `LineBreakNode` | interface | `{ element }` | Represents a GFM hard line break — two or more trailing spaces before a newline in markdown source, a `br` element in HTML. |
40
+ | `LinkNode` | interface | `{ element, href, children }` | Represents an inline link — `[text](href)`. `children` are the inline nodes of the link text. At render, html's floor removes a refused `href` attribute and the link keeps its text; `htmlToMarkdown` instead stores a refused destination as `''`. |
41
+ | `ImageNode` | interface | `{ element, src, children }` | Represents an inline image — `![alt](src)`. `children` are the inline nodes of the alternative content and `src` is the image destination. |
42
+ | `InlineNode` | type | `TextNode \| EmphasisNode \| CodeSpanNode \| LineBreakNode \| LinkNode \| ImageNode` | Represents a node that can appear inside inline content (a heading / paragraph / cell / list item / link text). |
43
+ | `HeadingNode` | interface | `{ element, level, children }` | Represents an ATX heading — `#` … `######`. `level` is 1–6 (the number of leading `#`), `children` the inline content of the heading text. |
44
+ | `ParagraphNode` | interface | `{ element, children }` | Represents a paragraph — a run of non-blank lines that is not another block; `children` its inline content. |
45
+ | `ListItemNode` | interface | `{ element, children }` | Represents one item of a `ListNode` — `children` the block content of the item (typically one paragraph, plus any nested list). |
46
+ | `ListNode` | interface | `{ element, ordered, start, items }` | Represents a list — bulleted (`-` / `*` / `+`, `ordered: false`) or numbered (`1.` / `1)`, `ordered: true`). `start` is the first ordinal of an ordered list (usually `1`). Nesting is expressed by a `ListNode` appearing in a `ListItemNode`'s `children`. |
47
+ | `TableNode` | interface | `{ element, header, rows, align }` | Represents a GFM table — `header` the inline content of each header cell, `rows` the body rows (each a list of cells, each cell inline content), `align` the per-column alignment from the delimiter row. A short body row is padded with empty cells; an over-long one is truncated to the header's column count. |
48
+ | `CodeBlockNode` | interface | `{ element, lang?, code }` | Represents a fenced code block — \`\`\`\` \`\`\`lang \`\`\`\`. \`code\` is the verbatim block content (no inner markdown; the closing fence and the trailing newline are stripped), \`lang\` the info-string language tag (the first word after the opening fence), absent when none was given. |
49
+ | `BlockquoteNode` | interface | `{ element, children }` | Represents a blockquote — `>`-prefixed lines; `children` the block content parsed from the de-quoted lines (so quotes nest). |
50
+ | `ThematicBreakNode` | interface | `{ element }` | Represents a thematic break — a horizontal rule (`---` / `***` / `___` on its own line). |
51
+ | `BlockNode` | type | `HeadingNode \| ParagraphNode \| ListNode \| TableNode \| CodeBlockNode \| BlockquoteNode \| ThematicBreakNode` | Represents a node that can appear at the block level of a document (or inside a list item / blockquote). |
52
+ | `MarkdownDocument` | interface | `{ element, children }` | Represents the root of a parsed markdown AST — the ordered block children of the whole document. The value `MarkdownInterface.document` holds. |
53
+ | `MarkdownNode` | type | `MarkdownDocument \| BlockNode \| ListItemNode \| InlineNode` | Represents any node in a markdown AST — the `MarkdownDocument` root, a `BlockNode`, a `ListItemNode`, or an `InlineNode`. The exhaustive set every projection's `switch` covers. |
54
+ | `MarkdownCell` | interface | `{ align, inlines }` | Represents one projected table cell — the inline content and alignment of a `th` / `td`. |
55
+ | `MarkdownProjection` | interface | `{ blocks, inlines, text, cells, rows }` | Represents what one HTML node projects to on the way to markdown — the fold value `htmlToMarkdown` carries up the AST. |
56
+ | `MarkdownHandler<TNode, T>` | type | `(node: TNode, children: readonly T[]) => T` | Represents a fold handler for one AST element — receives the node and its children already folded to `T`, and produces the node's own `T`. The building block of a `MarkdownHandlerMap` catamorphism table. |
57
+ | `MarkdownHandlerMap<T>` | interface | `{ document, heading, paragraph, thematicBreak, blockquote, codeBlock, list, listItem, table, text, emphasis, codeSpan, break, link, image }` | Represents the total catamorphism table for `MarkdownInterface.fold` — one `MarkdownHandler` per AST element, keyed by its `element` discriminant. Every key is required: a fold is total over the AST, so there is no element it can skip. |
58
+ | `MarkdownRewriteHandler` | type | `(node: MarkdownNode) => MarkdownNode` | Represents a copy-on-write node rewrite applied bottom-up by `MarkdownInterface.map` — receives one node (its own children already rewritten) and returns its replacement (the same node, unchanged, or a new node). |
59
+ | `MarkdownParseResult` | type | `readonly [document: MarkdownDocument, spans: ReadonlyMap<MarkdownNode, MarkdownSpan>]` | Pairs a parsed document with the `MarkdownSpan` of each of its nodes — what `parseProvenance` returns, and what `parseDocument` projects the document out of. |
60
+ | `MarkdownDerivation<T>` | type | `readonly [value: T, derivations: ReadonlyMap<MarkdownNode, MarkdownNode \| undefined>]` | Pairs a rewritten value with the input node each rewritten node was produced from — what `rewriteDocument` returns, so provenance survives a rewrite instead of ending at it. `T` is the rewritten value: the document for a whole-document rewrite. |
61
+ | `MarkdownInterface` | interface | `{ document } plus walk, find, filter, span, map, reduce, fold, stream` | Represents a stateful, parsed markdown document: the typed `MarkdownDocument` AST plus the query, rewrite, and fold operations over it. |
62
+
63
+ ### Constants
64
+
65
+ From [`constants.ts`](../src/core/constants.ts).
66
+
67
+ A `Shape` cell holds the constant's declared type.
68
+
69
+ | Constant | Kind | Shape | Summary |
70
+ | ------------------ | ----- | -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
71
+ | `MAX_DEPTH` | const | `number` | Caps the recursion depth the parse pipeline (`parseDocument` and its `parsers.ts` helpers), the `helpers.ts` traversal / projection functions (`markdownToHTML`, `renderMarkdown`, `walkNodes`, `foldNode`, `rewriteDocument`), and the `compilers.ts` renderer (`renderHTML`) honor before degrading, at 64. It bounds blockquote nesting, inline nesting (emphasis / links), and traversal / projection recursion so pathological or hostile input cannot exhaust the call stack. `htmlToMarkdown` is the inherited exception: its fold and depth cap belong to `@orkestrel/html`. |
72
+ | `EMPTY_PROJECTION` | const | `MarkdownProjection` | Holds the frozen empty HTML-to-markdown projection from which projection factories default every absent field. |
73
+
74
+ ### Parsers
75
+
76
+ The block/inline parsing pipeline, from [`parsers.ts`](../src/core/parsers.ts) — the orchestration `parseDocument` composes out of `helpers.ts`'s pure scanning leaves. `parseBlocks` is the recursive spine; each parser is exported and independently testable. The construct scanners it composes (`collectTable` / `collectList`) are leaves and live in [`helpers.ts`](../src/core/helpers.ts) with the other scanners; each calls back into the phase entry above it, so the two files are mutually recursive by design.
77
+
78
+ | Parser | Kind | Signature | Summary |
79
+ | ----------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
80
+ | `parseBlocks` | function | `(lines: readonly MarkdownSource[], depth: number, spans?: Map<MarkdownNode, MarkdownSpan>, end?: number) => readonly BlockNode[]` | Parses a run of markdown lines into a block AST, recursing into nested blockquotes, list items, and depth-capped degrade paragraphs. |
81
+ | `parseDocument` | function | `(markdown: string) => MarkdownDocument` | Parses a markdown string into a typed `MarkdownDocument` AST through the block phase — the document half of what `parseProvenance` returns. Malformed markdown degrades to literal text, so the parse never throws. |
82
+ | `parseProvenance` | function | `(markdown: string) => MarkdownParseResult` | Parses a markdown string into a document and its original-source spans. Malformed markdown degrades to literal text, so the parse never throws. |
83
+ | `parseInline` | function | `(text: string) => readonly InlineNode[]` | Parses inline markdown text (emphasis, code spans, links, images, and hard breaks) into inline AST nodes, coalescing adjacent text runs and reading no block structure. Malformed markdown degrades to literal text, so the parse never throws. |
84
+
85
+ ### Helpers
86
+
87
+ Pure, total leaves from [`helpers.ts`](../src/core/helpers.ts) — the line and character structural predicates the block and inline phases test raw lines with, the scanning functional core `parsers.ts` composes, the AST-crossing projections (`markdownToHTML`, `renderMarkdown`, `htmlToMarkdown`) plus the projection leaves they are built from, and the traversal engines `Markdown` delegates to. Every function is unit-testable in isolation; malformed input degrades to text, never throws. A predicate over a raw `string` narrows no type, so it is a leaf here rather than a guard in `validators.ts`; `renderHTML` drives the `HTML` class, so it sits in [`compilers.ts`](../src/core/compilers.ts) instead (§ [Compilers](#compilers)). `projectHTMLLeaf`'s leaf parameter is html's own `TextNode`, written `HTMLTextNode` in the signature because this package declares a `TextNode` of its own.
88
+
89
+ | Helper | Kind | Signature | Summary |
90
+ | ------------------------ | -------- | ------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
91
+ | `splitLines` | function | `(markdown: string) => readonly MarkdownSource[]` | Splits a markdown document into offset-bearing lines while normalizing CRLF and bare CR terminators at the line boundary. A single trailing terminator does not yield a final empty line. |
92
+ | `sliceSource` | function | `(source: MarkdownSource, from: number, to: number) => MarkdownSource` | Slices derived markdown text and narrows each intersecting source segment to the same text-relative range. |
93
+ | `joinSources` | function | `(sources: readonly MarkdownSource[], separator: string) => MarkdownSource` | Joins offset-bearing markdown sources while mapping a separator to the original region between adjacent mapped sources. |
94
+ | `projectSpan` | function | `(source: MarkdownSource, from: number, to: number) => MarkdownSpan \| undefined` | Projects a derived text range through its segments to a half-open region of the original markdown string. |
95
+ | `trimSource` | function | `(source: MarkdownSource) => MarkdownSource` | Trims an offset-bearing source without losing the coordinates of its retained text. |
96
+ | `normalizeParagraphLine` | function | `(source: MarkdownSource, breaks: boolean) => MarkdownSource` | Normalizes one paragraph line while retaining the full source run consumed by a trailing-space hard break. |
97
+ | `countIndent` | function | `(line: string) => number` | Counts the leading space / tab characters on `line` (a tab counts as one) — the indent that decides whether a list item's continuation belongs to the item. |
98
+ | `isFlankingWhitespace` | function | `(character: string) => boolean` | Checks whether `character` is whitespace under the emphasis flanking rule — a space, a tab, or a newline. |
99
+ | `isEscapable` | function | `(character: string) => boolean` | Checks whether `character` is escapable by a leading backslash — the ASCII punctuation markdown gives meaning to (so `\*` becomes `*` but `\.` stays `\.`). |
100
+ | `isBlankLine` | function | `(line: string) => boolean` | Checks whether `line` is blank — empty, or containing only whitespace — the markdown definition of a blank line that block parsing uses to separate paragraphs, skip gaps, and end list continuations. |
101
+ | `isQuote` | function | `(line: string) => boolean` | Checks whether `line` is a blockquote line (`>` optionally indented up to three spaces) — its content is de-quoted by `stripQuote`. |
102
+ | `isFenceClose` | function | `(line: string, marker: string) => boolean` | Checks whether `line` closes a fence opened by `marker` — the same fence character, a run at least as long, and nothing else but surrounding whitespace. |
103
+ | `isFenceWhitespace` | function | `(character: string \| undefined) => boolean` | Checks whether `character` is a regex-`\s`-equivalent whitespace character — the character class `isFenceClose`'s scan treats as surrounding padding. |
104
+ | `isThematicBreak` | function | `(line: string) => boolean` | Checks whether `line` is a thematic break (horizontal rule) — three or more of the same marker `-`, `*`, or `_` (optionally space-separated) and nothing else (`---`, `***`, `___`, `- - -`). |
105
+ | `isTableStart` | function | `(header: string, delimiter: string \| undefined) => boolean` | Checks whether the pair (`header`, `delimiter`) opens a GFM table — `delimiter` is a row of `\|`-separated cells each matching `:?-+:?`, the GFM rule that a table requires a header row immediately followed by a delimiter row. |
106
+ | `extractHeading` | function | `(line: string) => HeadingMatch \| undefined` | Extracts an ATX heading line (`#` … `######` followed by text) into its level, trimmed text, and the text's offset inside the line. A run of more than 6 `#`s, or `#`s not followed by whitespace + text, is not a heading; an optional closing `###` run is stripped. |
107
+ | `extractFence` | function | `(line: string) => FenceMatch \| undefined` | Extracts a fenced-code opening line (\`\`\`\` \`\`\` \`\`\`\` or \`~~~\`, optionally with an info string) into its \`{ marker, lang }\`, or \`undefined\` when \`line\` is not a fence opener. \`marker\` is the exact fence run (the closer must match the same character + at least the same length); \`lang\` is the first word of the info string. |
108
+ | `extractListItem` | function | `(line: string) => ListItemMatch \| undefined` | Extracts a list-item line (`-` / `*` / `+` bullet, or `1.` / `1)` ordinal, followed by a space) into its `ListItemMatch`, or `undefined` when `line` is not a list item. `content` is the text after the marker; `marker` is the full marker-plus-space width (for measuring a continuation's indent). |
109
+ | `stripQuote` | function | `(source: MarkdownSource) => MarkdownSource` | Strips one level of blockquote marker (`>` plus one optional following space) from an offset-bearing blockquote line, so the de-quoted source re-parses as nested blocks without losing its original coordinates. |
110
+ | `splitTableRow` | function | `(row: string) => readonly string[]` | Splits one GFM table row into its cell strings — outer pipes are optional, a pipe escaped by a leading backslash inside a cell is not a separator (it becomes a literal pipe character), and the empty leading / trailing cell an outer pipe produces is dropped. Derives the string form from `splitTableSources`, which owns the escaped-pipe splitting rule. |
111
+ | `splitTableSources` | function | `(row: MarkdownSource) => readonly MarkdownSource[]` | Splits an offset-bearing GFM table row into offset-bearing cells, retaining the complete source spelling of an escaped pipe while exposing its literal value. |
112
+ | `delimiterToAlignments` | function | `(delimiter: string) => readonly (TableAlign \| null)[]` | Derives the per-column `TableAlign` list from a GFM delimiter row — `:---` left, `---:` right, `:---:` center, and `---` as the explicit no-alignment marker represented by `null`. |
113
+ | `startsBlock` | function | `(lines: readonly string[], index: number) => boolean` | Checks whether the line at `index` starts a new block kind (heading / fence / thematic break / blockquote / list / table) — the paragraph collector stops at such a line so a block following a paragraph without a blank line still parses (a trusted-input caller writing a `##` heading directly under a paragraph, with no intervening blank line). |
114
+ | `unescapeText` | function | `(text: string) => string` | Resolves backslash escapes in a raw string to their literal characters — used for a link `href` (which is not otherwise inline-parsed) and any plain text run. |
115
+ | `coalesceText` | function | `(nodes: readonly InlineNode[], spans?: Map<MarkdownNode, MarkdownSpan>) => readonly InlineNode[]` | Merges adjacent text nodes into one — the inline scanner emits a text node per unrecognized character, so coalescing keeps the AST clean and assertion-friendly. |
116
+ | `scanCode` | function | `(source, start, to) => CodeSpanMatch \| undefined` | Scans an inline code span at `start` (a \`\` \` \`\`-run … a matching \`\` \` \`\`-run of the same length, the CommonMark rule that lets a span contain backticks). Returns the span's literal text + end index, or \`undefined\` when no matching closer exists (it then degrades to literal backticks). |
117
+ | `scanLink` | function | `(source, start, to, depth = 0) => LinkScan \| undefined` | Scans a link `[text](href)` at `start` — the text runs to a balanced `]`, then `(` must immediately follow and the destination runs to the matching `)` (both respect nested delimiters + escapes) through `locateLink`, and returns the parsed node and end index. Returns `undefined` when the shape does not hold (it then degrades to a literal `[`). |
118
+ | `scanEmphasis` | function | `(source, start, to, depth = 0) => EmphasisScan \| undefined` | Scans an emphasis run at `start` (`*` / `_`, doubled for strong) — finds the nearest matching closing run of the same marker + width while skipping complete nested runs from the other marker family, and requires non-space immediately inside both delimiters (the CommonMark flanking simplification that blocks `* x *`) through `locateEmphasis`, and returns the parsed node and end index. Returns `undefined` when no valid closer exists (it then degrades to a literal marker). |
119
+ | `locateLink` | function | `(source: string, start: number, to: number) => LinkBounds \| undefined` | Locates a link `[text](href)` at `start` — the text runs to a balanced `]`, then `(` must immediately follow and the destination runs to the matching `)` (both respect nested delimiters + escapes). Returns the label close and syntax end, or `undefined` when the shape does not hold (it then degrades to a literal `[`). |
120
+ | `locateEmphasis` | function | `(source: string, start: number, to: number) => EmphasisBounds \| undefined` | Locates an emphasis run at `start` (`*` / `_`, doubled for strong) — finds the nearest matching closing run of the same marker + width while skipping complete nested runs from the other marker family, and requires non-space immediately inside both delimiters (the CommonMark flanking simplification that blocks `* x *`). Returns the content and syntax bounds, or `undefined` when no valid closer exists (it then degrades to a literal marker). |
121
+ | `scanInline` | function | `(source: string, from: number, to: number, depth = 0) => readonly InlineNode[]` | Scans the window `[from, to)` of `source` into inline nodes — the single recursive engine the inline phase runs on (emphasis, link text, and image alternative content recurse through it). Linear: each character is consumed once; a failed construct emits its opening character as text and advances by one, so there is no re-scan (no ReDoS). |
122
+ | `scanInlineSource` | function | `(source: MarkdownSource, from: number, to: number, spans: Map<MarkdownNode, MarkdownSpan>, depth = 0) => readonly InlineNode[]` | Scans an offset-bearing inline window with the same engine as `scanInline` and records each emitted node against the original markdown string. |
123
+ | `collectTable` | function | `(lines: readonly MarkdownSource[], start: number, spans?: Map<MarkdownNode, MarkdownSpan>) => TableCollection` | Collects a GFM table starting at a header row, parsing the header, the alignment row, and every contiguous body row that follows. |
124
+ | `collectList` | function | `(lines: readonly MarkdownSource[], start: number, depth: number, spans?: Map<MarkdownNode, MarkdownSpan>, end?: number) => ListCollection` | Collects a list starting at the first item, gathering sibling items at the same indent/ordering and recursing into each item's own block content. |
125
+ | `markdownToHTML` | function | `(node: MarkdownNode) => HTMLDocument` | Projects a `MarkdownNode` into an unsanitized `HTMLDocument`. |
126
+ | `renderMarkdown` | function | `(node: MarkdownNode) => string` | Renders a `MarkdownNode` to its canonical markdown source — the inverse projection of `renderHTML`. It is the serializer a `parse(renderMarkdown(doc))` round-trip is built on. Canonical forms: `*` / `**` emphasis at even emphasis nesting depths and `_` / `__` at odd depths, `-` bullets, `N.` sequential ordinals (from the list's `start`), `---` thematic breaks, fenced code blocks (backtick run widened past any 3+ backtick run inside the body), ATX headings, `>`-prefixed blockquote lines, GFM tables (1-space-padded cells, a backslash before each literal pipe, an alignment delimiter row), `[text](href)` links, `![alt](src)` images, and two-space hard breaks. A `text` node's literal content is backslash-escaped wherever it would otherwise re-parse as markup, so parsing the rendered source returns the node it was rendered from. |
127
+ | `htmlToMarkdown` | function | `(node: HTMLNode) => MarkdownDocument` | Projects an `@orkestrel/html` `HTMLNode` into a `MarkdownDocument` — the HTML→markdown direction, and the inverse of `markdownToHTML`. |
128
+ | `createProjection` | function | `(parts?: Partial<MarkdownProjection>) => MarkdownProjection` | Builds an HTML-to-markdown projection with absent fields defaulted from `EMPTY_PROJECTION` and the block/inline exclusivity invariant enforced. |
129
+ | `trimInlines` | function | `(nodes: readonly InlineNode[]) => readonly InlineNode[]` | Trims the whitespace at the two ends of an inline run — the leading whitespace of a leading text node and the trailing whitespace of a trailing one — dropping either node when nothing survives. |
130
+ | `normalizeInlines` | function | `(nodes: readonly InlineNode[], breaks: boolean) => readonly InlineNode[]` | Reduces an inline run to the shape markdown can actually write back: adjacent text coalesced, empty text dropped, and every hard break either kept as a real line ending or spent as a space. |
131
+ | `mergeProjections` | function | `(children: readonly MarkdownProjection[]) => MarkdownProjection` | Combines the projections of one node's children into the projection of that node — the single place inline runs become paragraphs, so no ancestor has to decide it twice. |
132
+ | `projectHTMLLeaf` | function | `(leaf: CommentNode \| DoctypeNode \| HTMLTextNode) => MarkdownProjection` | Projects one HTML leaf — a text node, a comment, or a doctype — to its `MarkdownProjection`. |
133
+ | `projectHTMLNode` | function | `(node: ElementNode \| HTMLDocument, children: readonly MarkdownProjection[]) => MarkdownProjection` | Projects one HTML container — the document root or an element — from its children's already-computed projections. The element mapping, and the only place that decides what an HTML tag becomes in markdown. |
134
+ | `projectionToBlocks` | function | `(projection: MarkdownProjection) => readonly BlockNode[]` | Reads a projection as block content — the view a document, a blockquote, and a list item each need. |
135
+ | `projectionToInlines` | function | `(projection: MarkdownProjection) => readonly InlineNode[]` | Reads a projection as inline content — the view a link, an emphasis, and a table cell each need. |
136
+ | `walkNodes` | function | `(node: MarkdownNode) => Generator<MarkdownNode>` | Walks a `MarkdownNode` depth-first, pre-order, root-inclusive — yields the node itself, then recurses into its children (block children, list items, image/link inline children, table header/row cells' inline nodes) in walk order. |
137
+ | `foldNode` | function | `<T>(node: MarkdownNode, handlers: MarkdownHandlerMap<T>, depth: number) => T` | Folds a `MarkdownNode` into a `T` through a total catamorphism — children are folded first (post-order), then the node's own `MarkdownHandler` is invoked with the already-folded children. |
138
+ | `rewriteDocument` | function | `(document: MarkdownDocument, rewrite: MarkdownRewriteHandler) => MarkdownDerivation<MarkdownDocument>` | Rewrites a `MarkdownDocument` bottom-up (copy-on-write) — each node's children are rewritten first (post-order), then `rewrite` is applied to the node itself; the document root is never passed to `rewrite` (the `element: 'document'` invariant always holds). A table's inline cells and a list's items are rewritten too. |
139
+ | `flattenText` | function | `(node: MarkdownNode) => string` | Concatenates the `value` / `code` content of every descendant text / code-span / code-block node under `node`, including image alternative content, in walk order — the plain-text projection of an AST (search indexing, word counts, a text-only preview). |
140
+
141
+ ### Compilers
142
+
143
+ From [`compilers.ts`](../src/core/compilers.ts) — the class-driving half of the outbound direction. `helpers.ts` owns the pure `markdownToHTML` projection and imports no implementation class; the compiler below constructs `@orkestrel/html`'s `HTML` class, so it sits above the leaves and consumes them.
144
+
145
+ | Compiler | Kind | Signature | Summary |
146
+ | ------------ | -------- | -------------------------------- | ----------------------------------------------------- |
147
+ | `renderHTML` | function | `(node: MarkdownNode) => string` | Renders a `MarkdownNode` to sanitized canonical HTML. |
148
+
149
+ ### Shapers
150
+
151
+ Declarative `ContractShape` values (from `@orkestrel/contract`) from [`shapers.ts`](../src/core/shapers.ts) — one shape compiles into a guard, coercing parser, JSON Schema, and seeded generator (the compilers live in `@orkestrel/contract`, invoked here through `createContract` in `factories.ts`). Only the non-recursive node types shape here; any type whose fields recurse into `BlockNode` / `InlineNode` / `MarkdownNode` stays guard-only (`validators.ts`, through `lazyOf`) — see [Relationship with @orkestrel/contract](#relationship-with-orkestrelcontract).
152
+
153
+ A `Shape` cell holds the constant's declared type.
154
+
155
+ | Shaper | Kind | Shape | Summary |
156
+ | -------------------- | ----- | ---------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
157
+ | `textShape` | const | `ObjectShape<{ element, value }>` | Describes the shape of a `TextNode` — a plain-text leaf inline run. |
158
+ | `codeSpanShape` | const | `ObjectShape<{ element, value }>` | Describes the shape of a `CodeSpanNode` — an inline code span (\`\` \`code\` \`\`). |
159
+ | `lineBreakShape` | const | `ObjectShape<{ element }>` | Describes the shape of a `LineBreakNode` — a GFM hard line-break leaf. |
160
+ | `codeBlockShape` | const | `ObjectShape<{ element, lang?, code }>` | Describes the shape of a `CodeBlockNode` — a fenced code block. `lang` is optional (absent when the opening fence carries no info-string). |
161
+ | `thematicBreakShape` | const | `ObjectShape<{ element }>` | Describes the shape of a `ThematicBreakNode` — a horizontal rule. Carries no fields beyond its `element` discriminant. |
162
+ | `tableAlignShape` | const | `LiteralShape<'left' \| 'right' \| 'center'>` | Describes the shape of a `TableAlign` — the per-column GFM table alignment literal. Absence is no member of it, so the shape refuses the `null` a bare `---` delimiter takes in a `TableNode`'s `align` list. |
163
+ | `listItemMatchShape` | const | `ObjectShape<{ ordered, start, content, indent, marker }>` | Describes the shape of `ListItemMatch` — the parsed parts of a single list-item line the block phase's list detector returns. Fully non-recursive (no nested node fields), so every field shapes directly. |
164
+
165
+ ### Validators
166
+
167
+ Node guards, from [`validators.ts`](../src/core/validators.ts). The `is{Element}Node` guards narrow an already-parsed `MarkdownNode` by its `element` tag; the from-unknown guards (`isInlineNode` / `isBlockNode` / `isMarkdownNode` / `isMarkdownDocument`) instead validate an arbitrary `unknown` value against the full node shape, composed from `@orkestrel/contract` combinators. Two distinct guard families: the **from-unknown boundary guards** (`isInlineNode` / `isBlockNode` / `isMarkdownNode` / `isMarkdownDocument`) take `unknown` and validate an entire untrusted value from scratch; the **narrowing guards** (`is{Element}Node`, for example `isTableNode`) take an already-typed `MarkdownNode` and narrow it to one member of the union by its `element` tag — they assume the value is already a valid node shape. The line and character structural predicates the parser tests raw strings with narrow nothing, so they are pure leaves in `helpers.ts` (§ [Helpers](#helpers)).
168
+
169
+ In a guard table a `Shape` cell holds the type the guard narrows to.
170
+
171
+ | Guard | Kind | Shape | Summary |
172
+ | --------------------- | -------- | ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
173
+ | `isHeadingNode` | function | `HeadingNode` | Determines whether a node is a heading block. |
174
+ | `isParagraphNode` | function | `ParagraphNode` | Determines whether a node is a paragraph block. |
175
+ | `isListNode` | function | `ListNode` | Determines whether a node is a list block. |
176
+ | `isTableNode` | function | `TableNode` | Determines whether a node is a GFM table block. |
177
+ | `isCodeBlockNode` | function | `CodeBlockNode` | Determines whether a node is a fenced code block. |
178
+ | `isBlockquoteNode` | function | `BlockquoteNode` | Determines whether a node is a blockquote block. |
179
+ | `isThematicBreakNode` | function | `ThematicBreakNode` | Determines whether a node is a thematic break (horizontal rule) block. |
180
+ | `isTextNode` | function | `TextNode` | Determines whether a node is a plain text run. |
181
+ | `isEmphasisNode` | function | `EmphasisNode` | Determines whether a node is an emphasis run (`*em*` / `**strong**`). |
182
+ | `isCodeSpanNode` | function | `CodeSpanNode` | Determines whether a node is an inline code span. |
183
+ | `isLineBreakNode` | function | `LineBreakNode` | Determines whether a node is a GFM hard line break. |
184
+ | `isLinkNode` | function | `LinkNode` | Determines whether a node is a link. |
185
+ | `isImageNode` | function | `ImageNode` | Determines whether a node is an image. |
186
+ | `isInlineNode` | const | `InlineNode` | Determines whether an arbitrary value is a valid `InlineNode` — a text run, emphasis, code span, hard break, link, or image, recursively validated. |
187
+ | `isBlockNode` | const | `BlockNode` | Determines whether an arbitrary value is a valid `BlockNode` — a heading, paragraph, list, table, code block, blockquote, or thematic break, recursively validated. |
188
+ | `isMarkdownNode` | const | `MarkdownNode` | Determines whether an arbitrary value is a valid `MarkdownNode` — the `MarkdownDocument` root, a `BlockNode`, a `ListItemNode`, or an `InlineNode`, recursively validated. |
189
+ | `isMarkdownDocument` | const | `MarkdownDocument` | Determines whether an arbitrary value is a valid `MarkdownDocument` — the parsed-AST root `parseDocument` returns, recursively validated. |
190
+
191
+ ### Classes
192
+
193
+ The implementing class of `MarkdownInterface`, from [`Markdown.ts`](../src/core/Markdown.ts) —
194
+ documented in full under its own heading following this table.
195
+
196
+ | Name | Kind | Summary |
197
+ | ---------- | ----- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
198
+ | `Markdown` | class | Wraps a typed `MarkdownDocument` AST as a stateful, parsed markdown document with the query (`find` / `filter` / `reduce` / iteration), rewrite (`map`), fold, and streaming operations `MarkdownInterface` declares. |
199
+
200
+ ### `Markdown`
201
+
202
+ A stateful, parsed markdown workspace, constructed from a markdown `string` (runs `parseDocument`) or an already-parsed `MarkdownDocument` (adopted as-is, not re-validated). Exposes its AST through the `readonly document` member (documented here in Surface prose, per the `ContractInterface` precedent, alongside `walk` — both are part of the documented surface even though `document` carries no row in the [`## Methods`](#methods) table below, which lists only call-signature members). `walk` is the deep traversal — a lazy, depth-first, pre-order, root-inclusive generator over every node; its sync `for (const node of markdown.walk())` surface is also consumable by `for await (const node of markdown.walk())` (JavaScript accepts a sync iterable in a `for await`), so an async pipeline needs no separate iterator. Contrast with `stream`: `walk` is deep (every node) and sync; `stream` is shallow (top-level blocks only) and backpressure-respecting. Immutable — `map` never mutates the stored AST, it returns a new `Markdown`. See [`## Methods`](#methods) for its public call-signature surface.
203
+
204
+ ### Factories
205
+
206
+ From [`factories.ts`](../src/core/factories.ts).
207
+
208
+ | Factory | Kind | Signature | Summary |
209
+ | ----------------------------- | -------- | ---------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
210
+ | `createMarkdown` | function | `(input: string \| MarkdownDocument) => MarkdownInterface` | Creates a stateful markdown handle from a markdown string or an already-parsed `MarkdownDocument` — a typed AST plus the query, rewrite, and fold operations `MarkdownInterface` exposes. |
211
+ | `createTextContract` | function | `() => ContractInterface<TextNode>` | Compiles the `textShape` into a `ContractInterface` for `TextNode` — a guard, coercing parser, JSON Schema, and seeded generator from one shape declaration. |
212
+ | `createCodeSpanContract` | function | `() => ContractInterface<CodeSpanNode>` | Compiles the `codeSpanShape` into a `ContractInterface` for `CodeSpanNode` — a guard, coercing parser, JSON Schema, and seeded generator from one shape declaration. |
213
+ | `createLineBreakContract` | function | `() => ContractInterface<LineBreakNode>` | Compiles the `lineBreakShape` into a `ContractInterface` for `LineBreakNode`. |
214
+ | `createCodeBlockContract` | function | `() => ContractInterface<CodeBlockNode>` | Compiles the `codeBlockShape` into a `ContractInterface` for `CodeBlockNode` — a guard, coercing parser, JSON Schema, and seeded generator from one shape declaration. |
215
+ | `createThematicBreakContract` | function | `() => ContractInterface<ThematicBreakNode>` | Compiles the `thematicBreakShape` into a `ContractInterface` for `ThematicBreakNode` — a guard, coercing parser, JSON Schema, and seeded generator from one shape declaration. |
216
+
217
+ ## Methods
218
+
219
+ The public methods of each behavioral interface — one table per type, keyed by its backticked name. The `readonly document` member is Surface-documented above, not listed here — this table lists exactly `MarkdownInterface`'s call-signature members.
220
+
221
+ #### `MarkdownInterface`
222
+
223
+ | Method | Returns | Summary |
224
+ | -------- | ----------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
225
+ | `walk` | `Generator<MarkdownNode>` | Returns the deep traversal — a lazy, depth-first, pre-order, root-inclusive `Generator` over every `MarkdownNode` in the document. The sync `for (const node of markdown.walk())` surface is also consumable by `for await (const node of markdown.walk())` (JavaScript accepts a sync iterable in a `for await`), so async pipelines need no separate iterator. Contrast with `stream`: `walk` is deep, every-node, and sync; `stream` is shallow (top-level blocks only) and backpressure-respecting. |
226
+ | `find` | `T \| MarkdownNode \| undefined` | Finds the first node (depth-first, pre-order) narrowed by a type guard, and returns `undefined` when no node matches; a second overload takes a plain predicate. |
227
+ | `filter` | `readonly T[] \| readonly MarkdownNode[]` | Collects every node (depth-first, pre-order) narrowed by a type guard; a second overload takes a plain predicate. |
228
+ | `span` | `MarkdownSpan \| undefined` | Reads the region of the original markdown string a node was produced from. |
229
+ | `map` | `MarkdownInterface` | Rewrites the AST bottom-up (copy-on-write) and returns a new `MarkdownInterface`. |
230
+ | `reduce` | `T` | Folds the AST depth-first, pre-order into an accumulator through a reducer callback. |
231
+ | `fold` | `T` | Runs a total catamorphism over the document using a `MarkdownHandlerMap` table. |
232
+ | `stream` | `ReadableStream<BlockNode>` | Returns a web-standard `ReadableStream` over the document's top-level block nodes (shallow, source order) — a lazy, pull-based, backpressure-respecting source. A fresh, independently-replayable stream every call; never mutates the document. |
233
+
234
+ ## The AST model
235
+
236
+ Every node is plain, readonly data with no behavior — a discriminated union keyed by `element` (never `kind` / `type`). Block nodes and inline nodes:
237
+
238
+ - **Block nodes** (`BlockNode`) carry document structure: `heading`, `paragraph`, `list` (of `listItem`s), `table`, `codeBlock`, `blockquote`, `thematicBreak`. A `MarkdownDocument` is the root — `{ element: 'document', children: readonly BlockNode[] }`.
239
+ - **Inline nodes** (`InlineNode`) carry the inline content of a heading / paragraph / list item / table cell: `text`, `emphasis` (nests further inline children — `**bold _and italic_**` is a strong node wrapping a text node and an emphasis node), `codeSpan` (verbatim, no inner markdown), `break` (a GFM hard line break), `link` (nests inline children for its text), and `image`.
240
+
241
+ Recursion in the AST is structural, not incidental: a `blockquote`'s `children` re-parse the de-quoted lines as blocks (so quotes nest), a `list`'s `items` each carry `BlockNode[]` (so a nested list is a `list` block inside a `listItem`'s children), and `emphasis` / `link` / `image` nest `InlineNode[]`. `MarkdownNode` is the exhaustive union every projection's `switch` covers: `MarkdownDocument | BlockNode | ListItemNode | InlineNode`.
242
+
243
+ ### Images and hard breaks
244
+
245
+ An `ImageNode` carries its destination in `src` and its alternative content in `children`, exactly as a `LinkNode` carries `href` and its text — an image is a link that renders its target rather than pointing at it, and giving alt text the same inline children a link's text has means `walk`, `filter`, `map`, `fold`, and `flattenText` reach it without a special case. HTML's `alt` is a flat attribute, so the two directions meet in the middle: `markdownToHTML` writes `alt` from `flattenText`, and `htmlToMarkdown` reads `alt` back into a single text child.
246
+
247
+ A syntax hazard comes with each of them. An image is written `![alt](src)`, which is a `!` immediately followed by link syntax — so a `text` node that ends in `!` directly before a link would re-parse as an image it never was, and `renderMarkdown` backslash-escapes exactly that `!` and no other. A hard break is written as two trailing spaces before a newline, which is invisible and fragile: a break at either edge of a run has no line to end, a run of breaks reads as the blank line that would end the paragraph, and adjacent whitespace is eaten by line trimming. `renderMarkdown` writes what the AST holds, so keeping a break writable is the producer's job: `normalizeInlines` is the shared leaf that drops, merges, and trims breaks into the one shape markdown can carry, and spends a break as the space it stood for wherever the target is a single line (a heading, a table cell). The inbound projection runs every inline run through it for exactly that reason.
248
+
249
+ ### Alignment and absence
250
+
251
+ `TableAlign` is `'left' | 'right' | 'center'` — three real GFM delimiter forms, and nothing else. A column that declares no alignment is an absence, not another mode, and the two places absence appears differ because their containers differ:
252
+
253
+ - `TableNode.align` is positional — one entry per column, in column order — and JSON cannot carry `undefined` inside an array without changing the array's length on a round trip. It uses `null`, which is also the honest reading of the source: GFM's bare `---` is an explicit "no alignment here" marker written by the author, not an omitted field.
254
+ - `MarkdownCell.align` is a plain optional property on one cell, so absence is `undefined` there, per the ordinary rule.
255
+
256
+ `renderHTML` emits an `align` attribute only for those literals; a `null` column emits no attribute at all. `renderMarkdown` writes `:---` / `---:` / `:---:` for them and a bare `---` for `null`, so the delimiter row round-trips exactly.
257
+
258
+ ## The parse pipeline
259
+
260
+ `parseDocument(markdown)` runs its phases in order:
261
+
262
+ 1. **Block phase** — splits the document into lines (`splitLines`, CRLF/CR normalized) and walks them, detecting fences, thematic breaks, ATX headings, blockquotes, GFM tables, and lists (`parseBlocks`, `collectTable`, `collectList`); anything left over collects into a paragraph. `startsBlock` lets a new block interrupt a paragraph without a separating blank line.
263
+ 2. **Inline phase** — each block's raw text runs through `scanInline` (backslash escapes, code spans, links, images, emphasis, and the two-trailing-space hard break) through `parseInline`, then `coalesceText` merges adjacent text runs.
264
+
265
+ **Lines carry coordinates, not only text.** `splitLines` returns a `MarkdownSource` per line rather than a bare string: `{ text, segments }`, where `text` is the line a parser reads and `segments` are the `MarkdownSegment` runs mapping that text back to the original string. Each run is `{ offset, start, end }` — `offset` addresses `text`, `start` and `end` address the original, and the run's original length derives from `end - start` rather than being stored beside them. Every strip, trim, and join the block phase performs (`sliceSource`, `trimSource`, `stripQuote`, `joinSources`, `normalizeParagraphLine`, `splitTableSources`) narrows or remaps those runs instead of discarding them, so `projectSpan` turns a derived range back into an original one by resolving the range's two boundaries against those runs. It reports `undefined` when either boundary lands in a position the runs leave uncovered: joining two abutting regions with a separator leaves that separator's derived position uncovered. It bridges an uncovered interior when both boundaries resolve. That is what lets provenance be read off the source rather than reconstructed from node values, which would be wrong the moment a value differs from its spelling (§ [Source provenance](#source-provenance)).
266
+
267
+ `new Markdown(markdown)` (or `createMarkdown(markdown)`) calls `parseProvenance` internally and stores the resulting document as its `document`, alongside the span map `span` reads back. `markdownToHTML(node)`, `renderHTML(node)`, and `renderMarkdown(node)` are **separate**, downstream, standalone projections out of an AST — never fused into parsing, so a caller can inspect, transform, or fold the AST (through `Markdown`'s `find` / `filter` / `map` / `reduce` / `fold`) before ever calling one, or never call one at all. `htmlToMarkdown(node)` is the standalone projection in, and produces the same `MarkdownDocument` shape `parseDocument` does, so everything downstream of a parse works identically on a projection.
268
+
269
+ **Total / never-throw.** `parseDocument`, `markdownToHTML`, `renderHTML`, `renderMarkdown`, and `htmlToMarkdown` are all total functions: malformed markdown degrades to literal text (an unterminated `**` stays literal, a broken table falls back to a paragraph) rather than throwing, and hostile, cyclic, or pathologically deep HTML degrades rather than throwing. Inline scanning is index-based (no backtracking regex), so it is linear-time — no ReDoS on adversarial input.
270
+
271
+ ### Depth degrade semantics
272
+
273
+ `MAX_DEPTH` (`64`) bounds several independent recursions, each degrading to a fixed, cheap fallback instead of recursing further:
274
+
275
+ - **Block recursion** (blockquote / list nesting, `parsers.ts`'s `parseBlocks`) — past the cap, the remaining lines collapse into **one literal paragraph** containing those lines joined by `\n`, instead of continuing to parse nested structure.
276
+ - **Inline recursion** (`scanInline` and `scanInlineSource`) — the engine recurses into `scanInlineSource` itself for a link's text, an image's alternative content, and an emphasis run's children, incrementing `depth` at each descent. Past the cap, the scan window is not scanned for markup at all; it emits as a **single literal text node**.
277
+ - **`markdownToHTML` / `renderHTML` recursion** — past the cap, a node is not projected structurally; it yields a text node carrying the `value` of a node that has one (a `TextNode`, `CodeSpanNode`, …), and **nothing at all** for a node with no `value` field. A table reserves the four levels its `thead` / `tbody` / `tr` / cell scaffolding costs and contributes nothing when they would not fit, and a code block reserves the two its `pre > code` costs, so generated structure is charged to the same budget as authored structure and cannot escape the cap. The internal `switch` also carries a `default` arm, so a fabricated node with an `element` outside the exhaustive set (bypassing the type system, for example through an untyped/deserialized value) contributes nothing rather than `undefined` — the projection is total even against a hostile `MarkdownNode`.
278
+ - **`renderMarkdown` recursion** — the same cap and the same value-bearing-vs-empty degrade rule, applied to canonical markdown source instead of HTML.
279
+ - **`walkNodes` / `foldNode` recursion** — descent stops at the cap; the node at the cap is still yielded/folded (with an empty children list for `foldNode`), its children are not.
280
+ - **`rewriteDocument` / `Markdown.map` recursion** — the same cap, because `map` delegates to `rewriteDocument`: at the cap, the subtree is passed through unchanged (by reference — not rebuilt, and `rewrite` is not invoked on it) instead of recursing further, so a pathologically deep adopted document cannot exhaust the call stack.
281
+ - **`htmlToMarkdown` recursion — the one inherited cap.** This is the only recursion here markdown does not own: the fold is `@orkestrel/html`'s `foldNode`, so its depth bound is html's, and html happens to cap at `64` as well. Nesting past it truncates on html's side before markdown ever sees the content, and because the projected chain can end a level or two deeper than `MAX_DEPTH`, `renderMarkdown` may then truncate the result a second time. Deeply nested HTML is therefore bounded by html's cap on the fold and then by `renderMarkdown`'s own, in sequence rather than by a single cap: the anchor law that follows holds within the depth budget, and beyond it only totality is promised.
282
+
283
+ Together these bound pathological or hostile input (deeply nested blockquotes, runaway emphasis, adversarially deep ASTs) so no parsing, projecting, or writing function can ever exhaust the call stack.
284
+
285
+ ## Source provenance
286
+
287
+ A parse records the region of the source each node was produced from, and the `Markdown` instance
288
+ that ran the parse reads those regions back through `span`:
289
+
290
+ ```ts
291
+ import { Markdown, isHeadingNode } from '@orkestrel/markdown'
292
+
293
+ const source = '# Title\n\nA **bold** word.'
294
+ const markdown = new Markdown(source)
295
+
296
+ const heading = markdown.find(isHeadingNode)
297
+ const region = heading === undefined ? undefined : markdown.span(heading)
298
+ region // { start: 0, end: 7 }
299
+ if (region !== undefined) source.slice(region.start, region.end) // '# Title'
300
+ ```
301
+
302
+ These rules fix what a region means, and every one of them is about the string the instance was
303
+ constructed from:
304
+
305
+ - **A region addresses the original constructor string**, never the line text a later phase walks.
306
+ `source.slice(region.start, region.end)` returns the original source the node was produced from,
307
+ which is not the node's value: the region carries the syntax the value drops and the characters
308
+ that normalization removed.
309
+ - **A region is half-open and counted in UTF-16 code units** — `start` inclusive, `end` exclusive.
310
+ Its length derives from `end - start`; no length member exists to drift from the two offsets.
311
+ - **A parsed node covers the source it was produced from**, markers included, and sometimes more.
312
+ The heading in the preceding fence starts at `0`, not at `2` where its text starts, and an
313
+ emphasis node covers its `**` delimiters. A text node can also cover source its value drops: in
314
+ `'a \nb'` the paragraph phase trims the trailing space, so the text node's `value` is `a\nb`
315
+ while its region is `{ start: 0, end: 4 }` — the whole `a \nb`, trimmed space included.
316
+ - **A one-source rewrite keeps its source's region through `map`.** A node the handler replaced
317
+ from one input node reports that input's region; a rebuilt ancestor reports its original's.
318
+ - **A node with no region in this handle reports `undefined`** — a foreign node, every node of an
319
+ adopted document, and a rewrite output the handler built fresh from separate source nodes.
320
+
321
+ **A region always slices the constructor string verbatim; the derived text a phase reads does not.**
322
+ `splitLines` treats `\r\n` and a lone `\r` as terminators and drops them with the line boundary, so
323
+ a region can span a terminator the derived text no longer holds — that is what the separator segment
324
+ `joinSources` records is for. Inside a line the phases rewrite derived text too:
325
+ `normalizeParagraphLine` trims trailing spaces, `splitTableSources` turns an escaped `\|` into one
326
+ `|`, and the inline scan decodes every escape into the node's `value`. Each rewrite keeps the
327
+ original coordinates of what it retained, so a region still slices the constructor string verbatim
328
+ while the value it belongs to does not match that slice. Read provenance off the source; never
329
+ reconstruct it from a node value.
330
+
331
+ ```ts
332
+ import { Markdown, isTextNode } from '@orkestrel/markdown'
333
+
334
+ const source = 'a \\* b *c*'
335
+ const markdown = new Markdown(source)
336
+
337
+ const [text] = markdown.filter(isTextNode)
338
+ text?.value // 'a * b ' — the escape decoded
339
+ const region = text === undefined ? undefined : markdown.span(text)
340
+ region // { start: 0, end: 7 } — the region slices the spelling `a \* b `, not the value
341
+ ```
342
+
343
+ **`parseProvenance` is the handle-free entry point.** It returns the document and the
344
+ operation-owned span map together, so a caller that wants coordinates without a `Markdown` instance
345
+ gets the document and the map from one parse. `parseDocument` is its document projection, and constructing a `Markdown`
346
+ from a string runs it once and copies the map into the instance.
347
+
348
+ ```ts
349
+ import { parseProvenance } from '@orkestrel/markdown'
350
+
351
+ const [document, spans] = parseProvenance('# Title\n\nA **bold** word.')
352
+ spans.get(document) // { start: 0, end: 25 } — the whole input
353
+ ```
354
+
355
+ The map is keyed by node identity, so it addresses that document's nodes and no other. Two
356
+ instances over the same text hold independent maps, and a node from one reports `undefined` in the
357
+ other.
358
+
359
+ **A rewrite carries regions forward through its derivations.** `rewriteDocument` returns a
360
+ `MarkdownDerivation`, and `map` resolves each output node against the source instance's map in a
361
+ fixed order: an output identity that already holds a region keeps it, an output derived from one
362
+ input node takes that input's region, and anything else takes none. The resolution reads the direct
363
+ input the rewrite named for that output and stops there — it never follows a second derivation edge
364
+ back into an earlier rewrite's input.
365
+
366
+ ```ts
367
+ import { Markdown, isTextNode } from '@orkestrel/markdown'
368
+
369
+ const markdown = new Markdown('# Hi\n\nText.')
370
+ const lowered = markdown.map((node) =>
371
+ node.element === 'text' ? { element: 'text', value: node.value.toLowerCase() } : node,
372
+ )
373
+
374
+ const [first] = lowered.filter(isTextNode)
375
+ first?.value // 'hi'
376
+ const region = first === undefined ? undefined : lowered.span(first)
377
+ region // { start: 2, end: 4 } — where `Hi` sits in the original source
378
+ ```
379
+
380
+ **Where provenance stops, and why.** An adopted document, the inbound projection, and a rewrite
381
+ output that holds no region of its own and was assembled from separate sources each report
382
+ `undefined` rather than an approximate region, because an approximate region is worse than none — it
383
+ claims the author wrote something they did not:
384
+
385
+ - **An adopted document.** `new Markdown(document)` parsed no string, so no coordinates exist to
386
+ report. Adoption keeps the tree by reference and starts with an empty map, including where those
387
+ nodes are shared with an instance that does have regions.
388
+ - **The inbound projection.** `htmlToMarkdown` folds an HTML node, which carries no markdown
389
+ coordinates, so nothing it emits has a region — a paragraph `mergeProjections` synthesizes from a
390
+ run of pending siblings least of all. That paragraph has no single source even in principle: it
391
+ was assembled from an HTML wrapper's children, not written as a paragraph anywhere.
392
+ - **A rewrite output assembled from separate sources, where that output holds no region of its
393
+ own.** A node the handler built fresh and returned for several input nodes is covered by no single
394
+ region of the original, so `map` resolves it to `undefined`. An identity that already carries a
395
+ region is the exception: returning an existing spanned node for several inputs keeps that node's
396
+ own region, because own-region resolution runs first.
397
+
398
+ ```ts
399
+ import { Markdown, htmlToMarkdown, isTextNode } from '@orkestrel/markdown'
400
+ import { parseDocument as parseHTML } from '@orkestrel/html'
401
+
402
+ const imported = new Markdown(htmlToMarkdown(parseHTML('<div>text<p>para</p></div>')))
403
+ imported.span(imported.document) // undefined — adopted, and projected from HTML
404
+
405
+ const markdown = new Markdown('a *b* c')
406
+ const joined = { element: 'text', value: 'joined' } as const
407
+ const merged = markdown.map((node) => (node.element === 'text' ? joined : node))
408
+ const [text] = merged.filter(isTextNode)
409
+ text === undefined ? undefined : merged.span(text) // undefined — one output, separate sources
410
+ ```
411
+
412
+ `span` builds its return value fresh on every call, so a caller who mutates the object it hands
413
+ back changes nothing the next call reports.
414
+
415
+ ### Coordinates inside a line
416
+
417
+ The block phase carries `MarkdownSource` values rather than strings (§ [The parse
418
+ pipeline](#the-parse-pipeline)), and the inline engine has an offset-bearing entry point of its own.
419
+ `scanInlineSource` runs the same scan `scanInline` does and records each node it emits into a
420
+ recorder the caller owns:
421
+
422
+ ```ts
423
+ import { scanInlineSource, splitLines } from '@orkestrel/markdown'
424
+ import type { MarkdownNode, MarkdownSpan } from '@orkestrel/markdown'
425
+
426
+ const [line] = splitLines('> a *b*')
427
+ const spans = new Map<MarkdownNode, MarkdownSpan>()
428
+ const nodes = line === undefined ? [] : scanInlineSource(line, 2, line.text.length, spans, 0)
429
+
430
+ const [text, emphasis] = nodes
431
+ text === undefined ? undefined : spans.get(text) // { start: 2, end: 4 } — `a `
432
+ emphasis === undefined ? undefined : spans.get(emphasis) // { start: 4, end: 7 } — `*b*`
433
+ ```
434
+
435
+ Every coordinate leaf is exported for the same reason the projection leaves are: a caller writing
436
+ their own block or inline phase over the same lines needs `sliceSource`, `trimSource`,
437
+ `joinSources`, `normalizeParagraphLine`, `splitTableSources`, and `projectSpan` to keep coordinates
438
+ the way the shipped phases do.
439
+
440
+ ## Sanitization policy
441
+
442
+ There is one URL floor here, and markdown does not own it. `@orkestrel/html` owns HTML escaping, scheme judgement, and the sanitize floor; this package holds no escaper, no scheme list, and no sanitizer of its own, and composes with html's instead. A second copy would be a second thing to keep correct, and the failure mode of a sanitizer that has drifted from the one it was copied from is silent.
443
+
444
+ **`renderHTML` sanitizes, unconditionally.** It takes one argument and exposes no options, so there is no call shape that emits unsanitized HTML by accident:
445
+
446
+ ```text
447
+ markdownToHTML(node) → new HTML(document).sanitize({ attributes: [...SAFE_ATTRIBUTES, 'src'] }) → renderHTML(document)
448
+ ```
449
+
450
+ Only the middle step judges anything. `markdownToHTML` is deliberately inert: it leaves text literal and destinations unsanitized so that the projection stays a pure AST-to-AST mapping and a caller who wants a different policy can supply one. `markdownToHTML` projects only markdown's own node shapes — headings, paragraphs, links, images, emphasis, code, tables, and the rest — and never an `UNSAFE_ELEMENTS` tag such as `script`, the case the fence below exercises: raw HTML written in markdown source has no node shape of its own, so the parser leaves it as literal text and it reaches the output escaped rather than as an element. html's `UNSAFE_ELEMENTS` removal therefore has nothing to remove in this pipeline; what actually judges the projection's output is html's attribute floor, whose full refusal list — the always-stripped attributes and the hard-banned schemes — is [`guides/html.md`](./html.md)'s to state. The fence below shows one member of that floor: a `javascript:` destination stripped from a link's `href` and from an image's `src` while the element and its remaining content survive. This still runs unconditionally: `renderHTML` accepts any `MarkdownNode`, including one a caller constructed by hand, rewrote through `map`, or accepted from elsewhere, so it can never assume its input came from `parseDocument` on trusted markdown, and a hand-built node whose destination or attribute is hostile is still caught at this floor.
451
+
452
+ **The one widening: `src`.** html's `SAFE_ATTRIBUTES` deliberately omits resource `src`, because a sanitized page that keeps its `alt` text and loses its download is the safer default for a general HTML sanitizer. Markdown cannot accept that default — `![alt](src)` is syntax whose entire content is a destination — so `renderHTML` widens the attribute allowlist by exactly `src`, and by nothing else. The widening is narrow by construction: `src` is a member of html's `URL_ATTRIBUTES`, so every widened value still goes through `sanitizeURL`, and the hard refusals are not part of the allowlist axis at all and cannot be widened by anyone. A refused image keeps its element and its alt text and loses only the destination.
453
+
454
+ ```ts
455
+ import { parseDocument, renderHTML } from '@orkestrel/markdown'
456
+
457
+ const source = [
458
+ '<script>alert(1)</script>',
459
+ '',
460
+ '[link](javascript:alert(1))',
461
+ '',
462
+ '![alt](javascript:alert(1))',
463
+ '',
464
+ '![alt](https://x.dev/pic.png)',
465
+ ].join('\n')
466
+
467
+ const html = renderHTML(parseDocument(source))
468
+ // '<p>&lt;script&gt;alert(1)&lt;/script&gt;</p><p><a>link</a></p><p><img alt="alt"></p><p><img src="https://x.dev/pic.png" alt="alt"></p>'
469
+ // — the script line has no element shape of its own and renders as escaped text, so no
470
+ // script tag ever reaches the output for html's UNSAFE_ELEMENTS floor to remove; the
471
+ // javascript: link keeps its text and drops its href; the javascript: image keeps its
472
+ // element and its alt and drops only its src; the https: image keeps its src.
473
+ ```
474
+
475
+ **What the composed output looks like.** These differences are worth stating plainly, because they are visible in any byte-level comparison against a hand-rolled markdown renderer:
476
+
477
+ - **`align`, not `style`.** A table cell declares alignment as `align="left"`, not `style="text-align:left"`. html strips `style` unconditionally before any allowlist check, so a style-carrying cell would arrive at the browser with its alignment silently gone; `align` survives because html narrows it to the closed `TABLE_ALIGNMENTS` set on table cells only.
478
+ - **Literal quotes in text.** html's text encoder emits `&`, `<`, and `>` and leaves `"` and `'` alone, which is correct for character data and keeps prose readable: `alert("x" & 'y')` renders as `alert("x" &amp; 'y')`. Attribute values are encoded separately and do get their quotes handled.
479
+ - **Compact bytes.** Canonical serialization writes no whitespace between blocks: `# Hi\n\nText.` renders as `<h1>Hi</h1><p>Text.</p>`, not as two newline-separated lines.
480
+ - **A refused URL loses the whole attribute.** html removes a URL attribute it refuses rather than emptying it, so a hostile link renders `<a>x</a>` and a hostile image `<img alt="x">` — the words survive, the attribute does not appear at all.
481
+
482
+ **The inbound rule is different, and that asymmetry is deliberate.** Outbound, sanitizing at the very end is right: `markdownToHTML` can stay inert because `renderHTML` is the only door to a string and it always sanitizes. Inbound there is no such door. `htmlToMarkdown` produces a `MarkdownDocument`, and the serializer that document eventually reaches — `renderMarkdown` — is not a sanitization boundary and must not become one: it writes markdown source, where a destination is content rather than an executed attribute, and where a value dropped late could not be told apart from a value the author wrote. So the projection bakes html's `sanitizeURL(value, SAFE_URL_SCHEMES)` in at projection time, on every `href` and every `src`, whether or not the AST was ever sanitized — a hand-built one never was. A refused destination empties to `''` and the link or image is kept (`[text]()`), because a bad URL is no reason to lose the words around it. Two pipelines, two last responsible moments; the rule sits at each one rather than in the same place twice.
483
+
484
+ **Wanting a stricter policy.** Because the composition is exposed rather than hidden, a caller who needs a narrower floor does not need an option on `renderHTML`: project with `markdownToHTML`, sanitize with `@orkestrel/html`'s own `HTML` class and whatever `SanitizeOptions` they want, and serialize with html's `renderHTML`. That path can narrow the element set, drop `src` again, or replace the scheme list — and it still cannot go below html's floor, which is the point. The inbound counterpart is [bringing your own element policy](#bringing-your-own-element-policy): the same composition seam placed where that direction's element mapping varies.
485
+
486
+ ## `renderMarkdown` round-trip
487
+
488
+ `renderMarkdown` is the inverse of `parseDocument`: for any `MarkdownDocument` produced by `parseDocument`, `parseDocument(renderMarkdown(doc))` deep-equals `doc`, and `renderMarkdown` is idempotent — `renderMarkdown(parseDocument(renderMarkdown(doc))) === renderMarkdown(doc)`. It writes every node to one canonical markdown form, never the source's original (possibly variant) spelling:
489
+
490
+ | Construct | Canonical form |
491
+ | ------------------ | ----------------------------------------------------------------------------------------------------------------------------------------- |
492
+ | Emphasis | One marker family per nesting parity: `*em*` / `**strong**` at even emphasis depth, `_em_` / `__strong__` at odd. See below. |
493
+ | Bulleted list item | `- ` (a single hyphen + space), regardless of source marker (`*` / `+`). |
494
+ | Ordered list item | `N. ` — sequential ordinals starting from the list's `start`, `.`-style (never `)`) |
495
+ | Thematic break | `---`, regardless of source marker (`***` / `___` / spaced variants). |
496
+ | Fenced code block | Backtick fences, widened past any 3+ backtick run already inside the body. |
497
+ | Blockquote | `> `-prefixed lines (`>` alone for an otherwise-empty line). |
498
+ | GFM table | 1-space-padded cells, `\|`-escaped literal pipes, an explicit alignment delimiter row (bare `---` for a `null` column). |
499
+ | Link | `[text](href)` — `href` with `\`, `(`, `)` backslash-escaped (mirroring the parser's unescape) so a paren in the destination round-trips. |
500
+ | Image | `![alt](src)` — `src` escaped exactly like a link `href`, alt written from the image's inline children. |
501
+ | Hard break | Exactly two spaces then a newline, and only where a line can end (see [Images and hard breaks](#images-and-hard-breaks)). |
502
+ | Block separation | Exactly one blank line between top-level blocks; a document with zero blocks renders `''`. |
503
+
504
+ Sanitization is not a markdown-writing concern and does not appear in that table. A destination that needed judging was judged earlier — by html's floor on the way out to HTML, or by `htmlToMarkdown` on the way in — so `renderMarkdown` writes what the AST holds and escapes only what re-parsing requires.
505
+
506
+ **Why emphasis alternates.** A single canonical marker would be canonical and wrong: `**b *c***` closes ambiguously, because three identical markers in a row have more than one reading. Alternating families by nesting parity removes the ambiguity structurally — a nested run never shares a delimiter with the run enclosing it, so `**b _c_**` and `*x _a **c** b_ y*` each have exactly one parse. The form is still canonical in the sense that matters: it is a function of the AST's emphasis depth alone, never of the source's original spelling, so `_a **c** b_` and `*a __c__ b*` both write as `*a __c__ b*` and re-parse to the same tree.
507
+
508
+ A `text` node's literal content is backslash-escaped wherever it would otherwise re-parse as different markup (a leading `#`, a leading list marker, a leading `---` / `~~~` run, a literal `*`/`_`/`` ` ``/`[`/`]`, a trailing `!` before a link); a heading whose inline text ends in a `#` run (with or without leading whitespace) has that run's first `#` backslash-escaped so it cannot be mistaken for an ATX closing sequence on reparse — the round-trip soundness a parser owes its inverse, so parsing the rendered source returns the document it was rendered from.
509
+
510
+ This guarantee is scoped to documents `parseDocument` produced (or an equivalent well-formed `MarkdownDocument`). A value fabricated through `map` (or constructed by hand) that stuffs block-significant content or an embedded newline into a node field `renderMarkdown` treats as literal text (a `TextNode.value`, a `LinkNode.href`, …) has no round-trip guarantee — `renderMarkdown` still never throws, but the resulting source is not guaranteed to reparse back to the same AST.
511
+
512
+ ## `htmlToMarkdown` projection
513
+
514
+ `htmlToMarkdown` is the inbound direction: an `@orkestrel/html` `HTMLNode` in, a `MarkdownDocument` out, structurally identical to one `parseDocument` produces. It lives here rather than in html for the same reason `markdownToHTML` does — deciding that a `<pre>` is a fenced code block, that a `<div>` is nothing at all, and that a `<td>` holding two paragraphs must become one line of text is markdown-format knowledge, and html has no business carrying it.
515
+
516
+ **The engine is borrowed, the projection is not.** The traversal is html's own `foldNode` catamorphism, driven by a total handler table: `projectHTMLNode` for the containers (`document`, `element`) and `projectHTMLLeaf` for the leaves (`text`, `comment`, `doctype`). That is a deliberate reuse rather than a rebuild: html's fold already owns bottom-up ordering, cycle termination, and depth capping over its own AST, and reimplementing them here would mean maintaining a second, subtly different traversal of somebody else's data structure. What this package contributes is the fold value and the element mapping.
517
+
518
+ **The fold value is `MarkdownProjection`.** A node cannot know what it will become, because markdown decides late: a `<td>`'s content is inline inside a table and a paragraph outside one, and a `<code>` body is a code span in prose and a verbatim block under a `<pre>`. Rather than guess, every node reports each view an ancestor could want, and the ancestor that knows the context takes the one it needs:
519
+
520
+ | Field | What it carries | Who consumes it |
521
+ | --------- | ------------------------------------------------------------------- | ------------------------------------------------ |
522
+ | `blocks` | block content, with surrounding inline runs already made paragraphs | a document, a `blockquote`, an `li` |
523
+ | `inlines` | inline content; empty whenever `blocks` is not | a link, an emphasis, a table cell |
524
+ | `text` | raw subtree text — whitespace uncollapsed, escapes unresolved | a `code` span, a `pre > code` body |
525
+ | `cells` | the cells this node gives an enclosing row | a `tr` |
526
+ | `rows` | the rows this node gives an enclosing table | a `table`, through the `thead` / `tbody` between |
527
+
528
+ `blocks` and `inlines` are exclusive by construction — `mergeProjections` wraps a pending inline run into a paragraph the moment any sibling contributes a block, at that exact source position — so interleaving is never lost and no ancestor has to decide the same question twice. `projectionToBlocks` and `projectionToInlines` are the two readers, and `createProjection` is the one constructor that enforces the exclusivity invariant.
529
+
530
+ **The element mapping** is `projectHTMLNode`, and it is the only place that decides what an HTML tag becomes:
531
+
532
+ | HTML | Markdown |
533
+ | ----------------------------- | -------------------------------------------------------------------------------------------------------------------------- |
534
+ | `h1`–`h6` | a heading at that level |
535
+ | `p`, `li` | their block content, with a bare inline run wrapped in a paragraph |
536
+ | `blockquote` | a blockquote over that same block view |
537
+ | `hr`, `br` | a thematic break, a hard break |
538
+ | `strong` / `b`, `em` / `i` | strong and ordinary emphasis; whitespace padding moves outside the markers, because markdown refuses `* x *` |
539
+ | `code` | a code span, from the raw subtree text with newline runs collapsed to one space |
540
+ | `pre` | a code block: verbatim through a first `code` element child (its `language-` class naming the language), else `renderText` |
541
+ | `a`, `img` | a link and an image, each destination re-sanitized; an `img`'s `alt` becomes one text child |
542
+ | `ul` / `ol` | a list, ordered from the tag and numbered from `start`; one item per `li`, so an empty `li` is still an item |
543
+ | `th` / `td`, `tr`, `table` | a cell (alignment from its `align`), a row, and a GFM table whose header is the first `th`-bearing row |
544
+ | any `UNSAFE_ELEMENTS` element | nothing at all, text included |
545
+ | anything else | unwraps to its children |
546
+
547
+ The `pre`, list, and table rows read their own node rather than only their children's projections, because HTML puts the fact in a position rather than in a value: a `pre` takes its body from its `code` child's raw text, a list takes one item per `li` child, and a `tr` accepts only its own direct cells while a `table` derives the header row from its own source structure.
548
+
549
+ **The anchor law.** HTML is richer than markdown, so bytes cannot round-trip and the input document is the wrong fixpoint to promise. The right one is the projected AST:
550
+
551
+ ```text
552
+ parseDocument(renderMarkdown(htmlToMarkdown(x))) deep-equals htmlToMarkdown(x)
553
+ ```
554
+
555
+ Whatever the projection emits, markdown can write it and re-read it as the same tree. That is why the projection normalizes rather than translates literally — whitespace collapsed, edges trimmed, a blank paragraph dropped, a hard break kept only where a line can end, an emphasis's padding moved outside its markers: a shape markdown cannot write back is a shape this projection has no business producing. The law is proved over a corpus with one entry per construct the projection can emit, and again over the whole corpus concatenated into one document.
556
+
557
+ **What it loses, honestly.** The projection is lossy by construction, and these are the losses worth knowing before you rely on it:
558
+
559
+ - **Comments and doctypes vanish.** Neither carries anything markdown can represent, so both project to nothing.
560
+ - **Unknown wrappers unwrap.** An element with no markdown meaning contributes its children and disappears, so `<section><div>text</div><p>para</p></section>` keeps two blocks and loses two tags. Wrapper soup melts; content keeps its shape.
561
+ - **`UNSAFE_ELEMENTS` subtrees contribute nothing at all — text included.** A `<script>` body is not prose that lost its tag; it is content that never existed. Dropping the subtree whole is what stops it resurfacing.
562
+ - **Block content in a table cell flattens.** Markdown has no way to put a paragraph inside a cell, so a cell's blocks become one text node of their words, joined and whitespace-collapsed: a `<td>` holding `<p>a</p><p>b</p>` becomes the cell `a b`.
563
+ - **Presentation generally.** Attributes outside the small set the mapping reads (`href`, `src`, `alt`, `class` for a code language, `align` on a cell, `start` on an `ol`) have no markdown home and are not preserved.
564
+
565
+ Depth is the one bound markdown does not set here; see the inherited cap in [Depth degrade semantics](#depth-degrade-semantics).
566
+
567
+ ### Bringing your own element policy
568
+
569
+ `projectHTMLNode` and `projectHTMLLeaf` are exported as handlers, not hidden inside `htmlToMarkdown`, precisely so that the element mapping is replaceable without forking the projection. The vocabulary a replacement needs is exported alongside them: `createProjection` builds one with the exclusivity invariant enforced, `mergeProjections` combines a node's children into it, `projectionToBlocks` and `projectionToInlines` read one back out, and `trimInlines` and `normalizeInlines` reduce an inline run to a shape markdown can actually write. So a caller with a house rule — an element markdown has no opinion about, a wrapper that belongs in a blockquote, a `<kbd>` that reads better as code — writes one handler, delegates everything else to the default, and folds with html's `foldNode` exactly as `htmlToMarkdown` does:
570
+
571
+ The outbound counterpart is the stricter-policy recipe in [Sanitization policy](#sanitization-policy): the same composition seam placed where that direction's sanitize floor varies.
572
+
573
+ ```ts
574
+ import type { ElementNode, HTMLDocument, HTMLNode } from '@orkestrel/html'
575
+ import type { MarkdownDocument, MarkdownProjection } from '@orkestrel/markdown'
576
+ import { foldNode, parseDocument as parseHTML } from '@orkestrel/html'
577
+ import {
578
+ createProjection,
579
+ mergeProjections,
580
+ projectHTMLLeaf,
581
+ projectHTMLNode as projectDefaultHTMLNode,
582
+ projectionToBlocks,
583
+ renderMarkdown,
584
+ } from '@orkestrel/markdown'
585
+
586
+ // House rule: <kbd>Esc</kbd> reads as a code span. Every other element keeps the default.
587
+ function projectHTMLNode(
588
+ node: ElementNode | HTMLDocument,
589
+ children: readonly MarkdownProjection[],
590
+ ): MarkdownProjection {
591
+ if (node.category === 'element' && node.name === 'kbd') {
592
+ const merged = mergeProjections(children)
593
+ return createProjection({
594
+ inlines: [{ element: 'codeSpan', value: merged.text }],
595
+ text: merged.text,
596
+ })
597
+ }
598
+ return projectDefaultHTMLNode(node, children)
599
+ }
600
+
601
+ function project(node: HTMLNode): MarkdownDocument {
602
+ return {
603
+ element: 'document',
604
+ children: projectionToBlocks(
605
+ foldNode<MarkdownProjection>(node, {
606
+ document: projectHTMLNode,
607
+ element: projectHTMLNode,
608
+ text: projectHTMLLeaf,
609
+ comment: projectHTMLLeaf,
610
+ doctype: projectHTMLLeaf,
611
+ }),
612
+ ),
613
+ }
614
+ }
615
+
616
+ renderMarkdown(project(parseHTML('<p>Press <kbd>Esc</kbd> twice.</p>'))) // 'Press `Esc` twice.'
617
+ ```
618
+
619
+ The custom policy inherits everything the default has: the same fold, the same depth bound, the same totality, and the same anchor law for every element it did not override.
620
+
621
+ ## Relationship with `@orkestrel/contract`
622
+
623
+ Markdown's validation surface is a thin, purpose-built layer over `@orkestrel/contract`'s guard/combinator/shape machinery:
624
+
625
+ - **From-unknown guards for untrusted ASTs.** `isInlineNode` / `isBlockNode` / `isMarkdownNode` / `isMarkdownDocument` (`validators.ts`) are `Guard<T>` values composed from `recordOf` / `arrayOf` / `unionOf` / `literalOf` / `lazyOf` — each is total (never throws, even on cyclic or adversarially deep input) because every combinator involved is throw-contained by `@orkestrel/contract`'s guard contract. These validate a value that did **not** necessarily come from `parseDocument` — a deserialized document, a value crossing a process/RPC boundary.
626
+ - **Leaf shapes + compiled contracts, in lockstep.** `shapers.ts` declares `ContractShape` values (`textShape`, `codeSpanShape`, `lineBreakShape`, `codeBlockShape`, `thematicBreakShape`, `tableAlignShape`, `listItemMatchShape`) for the AST's non-recursive node types. `factories.ts` compiles the node shapes among them through `createContract` into `ContractInterface<T>` bundles — `schema` / `is` / `parse` / `generate` derived from one declaration, so they can never drift from each other.
627
+ - **Why recursive nodes are guard-only.** A `ContractShape` tree has no lazy/self-referential node — it is a finite, developer-authored tree the compilers can walk exhaustively. Any AST type whose fields recurse into `BlockNode` / `InlineNode` / `MarkdownNode` (`EmphasisNode`, `LinkNode`, `ImageNode`, `HeadingNode`, `ParagraphNode`, `ListItemNode`, `ListNode`, `TableNode`, `BlockquoteNode`, `MarkdownDocument`) is therefore **not** shaped — it stays guard-only, expressed directly in `validators.ts` with `@orkestrel/contract`'s `lazyOf` (the sanctioned recursion entry point: the thunk defers construction so a self-referential guard never references itself before it exists).
628
+
629
+ ## Patterns
630
+
631
+ Every feature below has a compact, runnable example. Together they cover every `MarkdownInterface`
632
+ method, every standalone projection and traversal helper, and the contract-factory fixture path.
633
+
634
+ ### Construct from a string and narrow with a guard
635
+
636
+ Construct a `Markdown` from a source string and narrow a found node with a guard:
637
+
638
+ ```ts
639
+ import { Markdown, isHeadingNode } from '@orkestrel/markdown'
640
+
641
+ const markdown = new Markdown('# Title\n\nA **bold** [link](https://x.dev).')
642
+ markdown.document.children[0] // { element: 'heading', level: 1, children: [...] }
643
+
644
+ const heading = markdown.find(isHeadingNode) // HeadingNode | undefined, narrowed
645
+ if (heading !== undefined) heading.level // number — narrowed to HeadingNode
646
+ ```
647
+
648
+ ### Construct from an adopted document
649
+
650
+ Adopt an already-parsed `MarkdownDocument` after validating it with a guard:
651
+
652
+ ```ts
653
+ import { Markdown, isMarkdownDocument } from '@orkestrel/markdown'
654
+ import type { MarkdownDocument } from '@orkestrel/markdown'
655
+
656
+ function adopt(candidate: unknown): Markdown | undefined {
657
+ if (!isMarkdownDocument(candidate)) return undefined // total guard - never throws
658
+ return new Markdown(candidate) // adopted as-is, not re-validated
659
+ }
660
+
661
+ const good: MarkdownDocument = { element: 'document', children: [] }
662
+ adopt(good) // Markdown instance
663
+ adopt({ element: 'bogus' }) // undefined - rejected before Markdown ever adopts it
664
+ ```
665
+
666
+ ### Filter and flatten
667
+
668
+ Filter every link node and flatten each one down to its link text:
669
+
670
+ ```ts
671
+ import { Markdown, isLinkNode, flattenText } from '@orkestrel/markdown'
672
+
673
+ const markdown = new Markdown('See [one](https://a.dev) and [two](https://b.dev).')
674
+ const links = markdown.filter(isLinkNode) // readonly LinkNode[]
675
+ const labels = links.map((link) => flattenText(link)) // ['one', 'two']
676
+ ```
677
+
678
+ ### Chain `map` rewrites, then write back with `renderMarkdown`
679
+
680
+ Chain two `map` rewrites and write the result back out with `renderMarkdown`:
681
+
682
+ ```ts
683
+ import { Markdown, renderMarkdown } from '@orkestrel/markdown'
684
+
685
+ const markdown = new Markdown('See [one](https://a.dev) and [two](https://b.dev).')
686
+
687
+ const shouted = markdown.map((node) =>
688
+ node.element === 'text' ? { element: 'text', value: node.value.toUpperCase() } : node,
689
+ )
690
+ const linked = shouted.map((node) =>
691
+ node.element === 'link' ? { ...node, href: `${node.href}?ref=guide` } : node,
692
+ )
693
+
694
+ renderMarkdown(linked.document) // 'SEE [ONE](https://a.dev?ref=guide) AND [TWO](https://b.dev?ref=guide).'
695
+ ```
696
+
697
+ Each `map` call returns a new `MarkdownInterface` — the original `markdown` is never mutated, so a
698
+ transform pipeline is a chain of small, composable, side-effect-free rewrites ending in a projection.
699
+
700
+ ### Reduce into an accumulator
701
+
702
+ Reduce over every heading node into a plain array of heading levels:
703
+
704
+ ```ts
705
+ import { Markdown, isHeadingNode } from '@orkestrel/markdown'
706
+
707
+ const markdown = new Markdown('# One\n\n## Two\n\nBody text.')
708
+
709
+ const levels = markdown.reduce<readonly number[]>(
710
+ (accumulator, node) => (isHeadingNode(node) ? [...accumulator, node.level] : accumulator),
711
+ [],
712
+ ) // [1, 2]
713
+ ```
714
+
715
+ ### Environment-agnostic fold
716
+
717
+ Fold a document through a total `MarkdownHandlerMap` that projects it to a plain HTML string:
718
+
719
+ ```ts
720
+ import { Markdown } from '@orkestrel/markdown'
721
+ import type { MarkdownHandlerMap } from '@orkestrel/markdown'
722
+
723
+ // A fold is total: one handler per element, no default arm, so a new AST node is a
724
+ // compile error here rather than a silent omission at runtime. Reach for `renderHTML`
725
+ // for real HTML — this table is the shape of an arbitrary projection, not a renderer.
726
+ const toHTML: MarkdownHandlerMap<string> = {
727
+ document: (_, children) => children.join('\n'),
728
+ heading: (node, children) => `<h${node.level}>${children.join('')}</h${node.level}>`,
729
+ paragraph: (_, children) => `<p>${children.join('')}</p>`,
730
+ thematicBreak: () => '<hr>',
731
+ blockquote: (_, children) => `<blockquote>${children.join('\n')}</blockquote>`,
732
+ codeBlock: (node) => `<pre><code>${node.code}</code></pre>`,
733
+ list: (node, children) =>
734
+ node.ordered ? `<ol>${children.join('')}</ol>` : `<ul>${children.join('')}</ul>`,
735
+ listItem: (_, children) => `<li>${children.join('')}</li>`,
736
+ table: (_, children) => `<table>${children.join('')}</table>`,
737
+ text: (node) => node.value,
738
+ emphasis: (node, children) =>
739
+ node.strong ? `<strong>${children.join('')}</strong>` : `<em>${children.join('')}</em>`,
740
+ codeSpan: (node) => `<code>${node.value}</code>`,
741
+ break: () => '<br>',
742
+ link: (node, children) => `<a href="${node.href}">${children.join('')}</a>`,
743
+ image: (node, children) => `<img src="${node.src}" alt="${children.join('')}">`,
744
+ }
745
+
746
+ const markdown = new Markdown('# Hi')
747
+ markdown.fold(toHTML) // '<h1>Hi</h1>'
748
+ ```
749
+
750
+ ### Shallow streaming with `stream()`
751
+
752
+ `stream()` returns a web-standard `ReadableStream<BlockNode>` — a fresh, pull-based stream every
753
+ call (one block enqueued per `pull`, so a slow reader's backpressure is respected). Each of these
754
+ ways consumes it:
755
+
756
+ ```ts
757
+ import { Markdown } from '@orkestrel/markdown'
758
+
759
+ const markdown = new Markdown('# Title\n\nFirst.\n\nSecond.')
760
+
761
+ // universal — a reader loop works in every ReadableStream-supporting environment
762
+ const reader = markdown.stream().getReader()
763
+ const tops: string[] = []
764
+ for (let result = await reader.read(); !result.done; result = await reader.read()) {
765
+ tops.push(result.value.element) // shallow — top-level blocks only
766
+ }
767
+ // tops: ['heading', 'paragraph', 'paragraph']
768
+
769
+ // Node / Deno / Firefox support native async iteration of ReadableStream
770
+ const topsAsync: string[] = []
771
+ for await (const block of markdown.stream()) topsAsync.push(block.element)
772
+ ```
773
+
774
+ ### Sync deep iteration
775
+
776
+ Walk every node synchronously with the deep, depth-first `walk` generator:
777
+
778
+ ```ts
779
+ import { Markdown } from '@orkestrel/markdown'
780
+
781
+ const markdown = new Markdown('# Title\n\nA **bold** word.')
782
+
783
+ const all: string[] = []
784
+ for (const node of markdown.walk()) all.push(node.element) // deep, depth-first, pre-order
785
+ ```
786
+
787
+ ### Async iteration with `for await…of`
788
+
789
+ Consume `walk()` and `stream()` alike with `for await…of`, in a writer that only needs an async iterable:
790
+
791
+ ```ts
792
+ import { Markdown } from '@orkestrel/markdown'
793
+
794
+ const markdown = new Markdown('# Title\n\nA **bold** word.')
795
+
796
+ async function writeAll(writer: { write(chunk: string): void }): Promise<void> {
797
+ for await (const node of markdown.walk()) writer.write(node.element) // sync generator, for-await composes fine
798
+ }
799
+
800
+ // `for await…of` also works over `stream()` — `ReadableStream` is natively async-iterable in
801
+ // Node / Deno / Firefox. Environments without that support use the reader loop above instead.
802
+ async function streamAll(writer: { write(chunk: string): void }): Promise<void> {
803
+ for await (const block of markdown.stream()) writer.write(block.element)
804
+ }
805
+ ```
806
+
807
+ `walk()` is a single lazy, sync generator over every node (deep, depth-first, pre-order,
808
+ root-inclusive) — a `for await…of` over it composes naturally with any async pipeline (a stream
809
+ writer, a queue) without first collecting the whole traversal into memory or needing a separate
810
+ async iterator.
811
+
812
+ ### Standalone projections and traversal on a bare node
813
+
814
+ Run the class-free projections, traversal, and rewrite functions directly against a bare node or document:
815
+
816
+ ```ts
817
+ import { parseDocument as parseHTML } from '@orkestrel/html'
818
+ import {
819
+ Markdown,
820
+ markdownToHTML,
821
+ renderHTML,
822
+ renderMarkdown,
823
+ htmlToMarkdown,
824
+ walkNodes,
825
+ foldNode,
826
+ rewriteDocument,
827
+ parseInline,
828
+ parseDocument,
829
+ } from '@orkestrel/markdown'
830
+ import type { MarkdownHandlerMap } from '@orkestrel/markdown'
831
+
832
+ const markdown = new Markdown('# Hi\n\nText.')
833
+
834
+ renderHTML(markdown.document) // '<h1>Hi</h1><p>Text.</p>' — sanitized, canonical, compact
835
+
836
+ // The intermediate AST, for a caller who wants to apply their own HTML policy.
837
+ markdownToHTML(markdown.document) // { category: 'document', children: [...] }
838
+
839
+ // renderMarkdown round-trip: parseDocument(renderMarkdown(doc)) deep-equals doc.
840
+ const roundTripped = parseDocument(renderMarkdown(markdown.document))
841
+
842
+ // The inbound direction: HTML in, the same MarkdownDocument shape a parse produces.
843
+ const imported = htmlToMarkdown(parseHTML('<h1>Release notes</h1><p>Ship <b>fast</b>.</p>'))
844
+ renderMarkdown(imported) // '# Release notes\n\nShip **fast**.'
845
+
846
+ // The class-free path: walkNodes / foldNode / rewriteDocument all operate on a bare MarkdownNode,
847
+ // no Markdown instance required.
848
+ const heading = markdown.document.children[0]
849
+ const elements = [...walkNodes(heading)].map((node) => node.element) // ['heading', 'text']
850
+
851
+ const countHandlers: MarkdownHandlerMap<number> = {
852
+ document: (_, children) => children.reduce((a, b) => a + b, 0),
853
+ heading: (_, children) => 1 + children.reduce((a, b) => a + b, 0),
854
+ paragraph: (_, children) => 1 + children.reduce((a, b) => a + b, 0),
855
+ thematicBreak: () => 1,
856
+ blockquote: (_, children) => 1 + children.reduce((a, b) => a + b, 0),
857
+ codeBlock: () => 1,
858
+ list: (_, children) => 1 + children.reduce((a, b) => a + b, 0),
859
+ listItem: (_, children) => 1 + children.reduce((a, b) => a + b, 0),
860
+ table: (_, children) => 1 + children.reduce((a, b) => a + b, 0),
861
+ text: () => 1,
862
+ emphasis: (_, children) => 1 + children.reduce((a, b) => a + b, 0),
863
+ codeSpan: () => 1,
864
+ break: () => 1,
865
+ link: (_, children) => 1 + children.reduce((a, b) => a + b, 0),
866
+ image: (_, children) => 1 + children.reduce((a, b) => a + b, 0),
867
+ }
868
+ const nodeCount = foldNode(heading, countHandlers, 0) // 2
869
+
870
+ // rewriteDocument is copy-on-write and returns a MarkdownDerivation: the rewritten value, plus
871
+ // the input node each rewritten node was produced from. An unchanged subtree keeps its
872
+ // identity, so an identity rewrite returns the very same document object and records no
873
+ // derivation at all.
874
+ const [rewritten, derivations] = rewriteDocument(markdown.document, (node) =>
875
+ node.element === 'text' ? { element: 'text', value: node.value.toLowerCase() } : node,
876
+ )
877
+ rewritten.children.length === markdown.document.children.length // true — same shape, lowercased text
878
+ derivations.size > 0 // true — the rebuilt spine, keyed by the output nodes
879
+
880
+ const [unchanged] = rewriteDocument(markdown.document, (node) => node)
881
+ unchanged === markdown.document // true — nothing changed, so nothing was rebuilt
882
+
883
+ const fragment = parseInline('a **bold** span') // readonly InlineNode[], no block structure
884
+ ```
885
+
886
+ ### Scan one inline construct
887
+
888
+ The inline phase's construct scanners are exported leaves, so a caller can run one against a bare
889
+ string window without a document around it. Each returns its parsed node and the index its syntax
890
+ ended at, or `undefined` when the shape does not hold — the point at which the engine degrades the
891
+ opening marker to literal text.
892
+
893
+ ```ts
894
+ import { scanEmphasis, scanInline, scanLink } from '@orkestrel/markdown'
895
+
896
+ scanInline('a **b** c', 0, 9) // readonly InlineNode[] — text, emphasis, text
897
+
898
+ const link = scanLink('[docs](https://x.dev)', 0, 21)
899
+ link?.node.href // 'https://x.dev'
900
+ link?.end // 21
901
+
902
+ const emphasis = scanEmphasis('**bold**', 0, 8)
903
+ emphasis?.node.strong // true
904
+ emphasis?.end // 8
905
+
906
+ scanLink('[unclosed', 0, 9) // undefined — degrades to a literal `[`
907
+ ```
908
+
909
+ `locateLink` and `locateEmphasis` are the same scans without the node construction: they return the
910
+ syntax bounds alone, which is what the offset-bearing path needs to record a region before it builds
911
+ anything (§ [Source provenance](#source-provenance)).
912
+
913
+ `scanLink` and `scanEmphasis` are standalone leaves rather than steps of the engine. A parse runs
914
+ `scanInlineSource`, which locates each construct with `locateLink` / `locateEmphasis`, recurses into
915
+ itself for the construct's children, and passes that child run through `coalesceText` before storing
916
+ it. `scanLink` and `scanEmphasis` skip that step: `node.children` is the raw scan output, which
917
+ carries no coalescing guarantee. Apply `coalesceText` yourself when you depend on one.
918
+
919
+ ### Guide-parity extraction
920
+
921
+ Extract every `## Surface`-table identifier from this guide's own markdown text:
922
+
923
+ ```ts
924
+ import { Markdown, isTableNode, flattenText } from '@orkestrel/markdown'
925
+
926
+ // Extract every Surface-table first-column identifier from this very guide.
927
+ function extractSurfaceNames(source: string): readonly string[] {
928
+ const markdown = new Markdown(source)
929
+ const tables = markdown.filter(isTableNode) // readonly TableNode[] — narrowed, no cast needed
930
+ return tables.flatMap((table) =>
931
+ table.rows.map((row) => flattenText({ element: 'paragraph', children: row[0] ?? [] })),
932
+ )
933
+ }
934
+ ```
935
+
936
+ ### Contract-backed fixture generation
937
+
938
+ Compile a shape into a contract and generate a reproducible fixture from a seed:
939
+
940
+ ```ts
941
+ import { createTextContract } from '@orkestrel/markdown'
942
+ import { seededRandom } from '@orkestrel/contract'
943
+
944
+ const text = createTextContract()
945
+ text.schema // the compiled JSON Schema for TextNode
946
+ const fixture = text.generate(seededRandom(42)) // reproducible seed data
947
+ text.is(fixture) // true — guard / generator stay in lockstep
948
+ ```
949
+
950
+ ## Tests
951
+
952
+ - [`tests/guides.test.ts`](../tests/guides.test.ts) — the `## Surface` ↔ `src/core` bijection (value + type exports), the `MarkdownInterface` ↔ `Markdown` method bijection, and the equality gate: every `Summary` cell against its declaration's description paragraph, the titled `Construct from a string and narrow with a guard` fence against the `@example` block of that title (pinned so the titled pair cannot be retired silently), and the README pitch against this guide's tagline. It also runs the flagship fences and asserts the values their comments claim.
953
+ - [`tests/src/core/Markdown.test.ts`](../tests/src/core/Markdown.test.ts) — `walk` / `find` / `filter` / `map` / `reduce` / `fold` / `stream` behavior, construction from a string vs. an already-parsed document.
954
+ - [`tests/src/core/parsers.test.ts`](../tests/src/core/parsers.test.ts) — `parseDocument` / `parseInline` / `parseBlocks`, incl. degrade semantics at `MAX_DEPTH`.
955
+ - [`tests/src/core/validators.test.ts`](../tests/src/core/validators.test.ts) — structural predicates + per-node guards + the from-unknown AST guards (soundness on cyclic / adversarial input).
956
+ - [`tests/src/core/helpers.test.ts`](../tests/src/core/helpers.test.ts) — the pure line/block/inline scanning leaves, including the `collectTable` / `collectList` construct scanners; `markdownToHTML` (the unsanitized projection) and `renderMarkdown` (canonical forms, the emphasis-parity corpus, the parse↔render round-trip); the projection leaves (`trimInlines` / `normalizeInlines` / `mergeProjections` / `projectionToBlocks` / `projectionToInlines` / `projectHTMLLeaf` / `projectHTMLNode`) and `htmlToMarkdown` end to end — element mapping, adversarial and cyclic input, the round-trip anchor law over the projection corpus, and the grand markdown → HTML → markdown trip; plus `walkNodes` / `foldNode` / `rewriteDocument` / `flattenText`.
957
+ - [`tests/src/core/compilers.test.ts`](../tests/src/core/compilers.test.ts) — `renderHTML` as the composed pipeline: structure, escaping and sanitization, the `@orkestrel/html` URL floor, the `src` widening, composed elements, and the `MAX_DEPTH` degrade arms.
958
+ - [`tests/src/core/shapers.test.ts`](../tests/src/core/shapers.test.ts) — per-shape guard exactness, JSON Schema essentials, seeded generate round-trips, parse rebuilds, and bidirectional `Infer` ↔ interface type parity.
959
+ - [`tests/src/core/factories.test.ts`](../tests/src/core/factories.test.ts) — `createMarkdown` + the compiled node contracts (`is` / `parse` / `schema` / `generate` round-trips).
960
+
961
+ ## See also
962
+
963
+ - [`AGENTS.md`](../AGENTS.md) — the repository rules, including the documentation contract every guide here is held to.
964
+ - [`README.md`](README.md) — the guides index.