samemark 0.0.0-stage → 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +13 -0
- package/KNOWN_GAPS.md +84 -0
- package/LICENSE +21 -0
- package/README.md +676 -2
- package/core.d.ts +10 -0
- package/core.js +10 -0
- package/dist/LICENSES.txt +217 -0
- package/dist/core.cjs +7178 -0
- package/dist/stream.cjs +6967 -0
- package/examples/react-chat/StreamMarkdown.jsx +67 -0
- package/index.d.ts +10 -0
- package/index.js +14 -0
- package/lib/api.js +155 -0
- package/lib/probe.js +109 -0
- package/lib/rows.js +47 -0
- package/lib/signature.js +77 -0
- package/lib/types.d.ts +46 -0
- package/lib/verify.js +38 -0
- package/package.json +140 -4
- package/src/block.js +1807 -0
- package/src/inline.js +1597 -0
- package/src/stream.js +178 -0
- package/stream.d.ts +26 -0
- package/stream.js +7 -0
package/README.md
CHANGED
|
@@ -1,3 +1,677 @@
|
|
|
1
|
-
#
|
|
1
|
+
# samemark
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
The same mdast as remark-parse, much faster.
|
|
4
|
+
|
|
5
|
+
<!-- demo:start -->
|
|
6
|
+
**[Live demo](https://anzal1.github.io/samemark/)**: a long chat answer streamed into react-markdown (left) and into
|
|
7
|
+
`samemark/stream` (right), with per-update timings measured in your browser and a check that both DOMs are identical,
|
|
8
|
+
plus a playground that compares samemark's tree with `mdast-util-from-markdown` as you type.
|
|
9
|
+
|
|
10
|
+
<a href="https://anzal1.github.io/samemark/"><img alt="Screen capture of the live demo: the same 10,000 character chat answer streams into react-markdown on the left and samemark/stream on the right, with per-update timings, a sparkline for each side and a badge reporting that both panes hold identical HTML." src="assets/demo-stream.gif" width="880"></a>
|
|
11
|
+
|
|
12
|
+
The capture above is the live demo, recorded in headless Chromium on a GitHub Actions runner ([capture run 38066505899](https://github.com/anzal1/samemark/actions/runs/38066505899): 250 updates, 1.76 s of update time on the left, 94 ms on the right, identical HTML at all 25 checks). On your own device the numbers will differ; they are measured in your browser.
|
|
13
|
+
<!-- demo:end -->
|
|
14
|
+
|
|
15
|
+
<picture>
|
|
16
|
+
<source media="(prefers-color-scheme: dark)" srcset="assets/e2e-speedup-dark.svg">
|
|
17
|
+
<img alt="Horizontal bar chart of end to end markdown to HTML time for six inputs, remark-parse against samemark. samemark is 6.2x to 10.3x faster, for example 5.00 s down to 486 ms on the Node.js API docs." src="assets/e2e-speedup-light.svg" width="880">
|
|
18
|
+
</picture>
|
|
19
|
+
|
|
20
|
+
samemark is a markdown parser that returns the same mdast tree as `mdast-util-from-markdown` with GFM (what
|
|
21
|
+
`remark-parse` + `remark-gfm` produce), positions included, and falls back to that real parser for anything it
|
|
22
|
+
does not reproduce exactly.
|
|
23
|
+
|
|
24
|
+
[micromark](https://github.com/micromark/micromark) by Titus Wormer defines the behavior; this parser matches it and
|
|
25
|
+
falls back to it. Its test suite and its output are the specification.
|
|
26
|
+
|
|
27
|
+
**How it was made.** AI agents wrote most of the code. Every change is checked against micromark's own output: the
|
|
28
|
+
full tree must be equal, node by node, on thousands of real documents, the CommonMark spec examples, and tens of
|
|
29
|
+
millions of generated inputs. Nothing here relies on trust in the author or the model; the checks are in this
|
|
30
|
+
repository and the commands to run them are below.
|
|
31
|
+
|
|
32
|
+
## What you get
|
|
33
|
+
|
|
34
|
+
<!-- bench:headline -->
|
|
35
|
+
In a real unified pipeline (markdown to HTML with `remark-gfm`, `remark-rehype`, `rehype-stringify`) the whole run is
|
|
36
|
+
**6.2x to 10.3x faster** with samemark as the parser, across six inputs on a GitHub Actions runner (Node 24, median of
|
|
37
|
+
7 rounds, identical HTML on every file). Parsing alone is 28x to 47x faster; the rest of the pipeline is unchanged,
|
|
38
|
+
which is why the end-to-end number is smaller.
|
|
39
|
+
|
|
40
|
+
Streaming a 30,000 character chat answer into a real DOM in Chromium, `samemark/stream` with memoised blocks spends
|
|
41
|
+
**60x less main-thread time** than react-markdown re-rendering the whole message on every chunk, and 77x less with the
|
|
42
|
+
CPU throttled 6x as a phone proxy (79.4 s down to 1.03 s over 750 pushes; 103 ms down to 1.1 ms per push), with the same
|
|
43
|
+
DOM. Swapping only the parser inside react-markdown gives 5x to 6x.
|
|
44
|
+
|
|
45
|
+
Building a 239-page Astro site took 29.6 s with `remark-parse` and 22.9 s with samemark (23% less).
|
|
46
|
+
|
|
47
|
+
All numbers come from one CI run ([38036885270](https://github.com/anzal1/samemark/actions/runs/38036885270)) and are
|
|
48
|
+
reproducible with `npm run bench`. Details and method: [Benchmarks](#benchmarks).
|
|
49
|
+
<!-- /bench:headline -->
|
|
50
|
+
|
|
51
|
+
- `parse(source)` returns an mdast `Root` equal to `fromMarkdown(source, {extensions: [gfm()], mdastExtensions: [gfmFromMarkdown()]})`.
|
|
52
|
+
- `remarkParseFast` is a drop-in for `remark-parse` in a unified processor.
|
|
53
|
+
- `remark-gfm`, `remark-frontmatter` (yaml, toml or both) and `remark-math` (with or without `singleDollarTextMath`)
|
|
54
|
+
are implemented natively, alone or together. Anything else (`remark-directive`, MDX, custom frontmatter fences,
|
|
55
|
+
other micromark extensions) is detected and given to the real parser. You get micromark's answer, at micromark's
|
|
56
|
+
speed.
|
|
57
|
+
- No micromark at parse time. The fast path imports six small pure helper packages (four `micromark-util-*`
|
|
58
|
+
packages plus `ccount` and `decode-named-character-reference`) and nothing else. `mdast-util-from-markdown` is
|
|
59
|
+
only needed for the fallback, as an optional peer dependency.
|
|
60
|
+
|
|
61
|
+
## Install
|
|
62
|
+
|
|
63
|
+
```sh
|
|
64
|
+
npm install samemark
|
|
65
|
+
# for the fallback (recommended; the main entry imports it):
|
|
66
|
+
npm install mdast-util-from-markdown unified remark-gfm
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
Node 18 or newer, or any runtime with ES2020 and `TextDecoder` (browsers, Deno, Bun, edge runtimes).
|
|
70
|
+
|
|
71
|
+
## Usage
|
|
72
|
+
|
|
73
|
+
With unified, replace `remark-parse`:
|
|
74
|
+
|
|
75
|
+
```js
|
|
76
|
+
import {unified} from 'unified'
|
|
77
|
+
import remarkGfm from 'remark-gfm'
|
|
78
|
+
import remarkRehype from 'remark-rehype'
|
|
79
|
+
import rehypeStringify from 'rehype-stringify'
|
|
80
|
+
import remarkParseFast from 'samemark'
|
|
81
|
+
|
|
82
|
+
const file = await unified()
|
|
83
|
+
.use(remarkParseFast) // instead of remark-parse
|
|
84
|
+
.use(remarkGfm)
|
|
85
|
+
.use(remarkRehype)
|
|
86
|
+
.use(rehypeStringify)
|
|
87
|
+
.process('# Hello *world*')
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
Inside a framework that already installs `remark-parse` (Astro, MDX, Next.js plugins), add it to `remarkPlugins`;
|
|
91
|
+
it replaces the parser after `remark-gfm` has registered, which is the order it needs:
|
|
92
|
+
|
|
93
|
+
```js
|
|
94
|
+
// astro.config.mjs
|
|
95
|
+
import remarkParseFast from 'samemark'
|
|
96
|
+
export default {markdown: {remarkPlugins: [remarkParseFast]}}
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
With react-markdown (it builds a unified processor on every render with `remark-parse` first, so a plugin that sets the
|
|
100
|
+
parser takes over; verified by the benchmark below):
|
|
101
|
+
|
|
102
|
+
```jsx
|
|
103
|
+
import Markdown from 'react-markdown'
|
|
104
|
+
import remarkGfm from 'remark-gfm'
|
|
105
|
+
import remarkParseFast from 'samemark'
|
|
106
|
+
|
|
107
|
+
const plugins = [remarkParseFast, remarkGfm] // define once, outside the component
|
|
108
|
+
|
|
109
|
+
export const Answer = ({text}) => <Markdown remarkPlugins={plugins}>{text}</Markdown>
|
|
110
|
+
```
|
|
111
|
+
|
|
112
|
+
Chat apps re-render the whole message on every streamed token, so the parser runs again over a growing string each
|
|
113
|
+
time; see the react-markdown numbers in [Benchmarks](#benchmarks).
|
|
114
|
+
|
|
115
|
+
Without a processor:
|
|
116
|
+
|
|
117
|
+
```js
|
|
118
|
+
import {parse} from 'samemark'
|
|
119
|
+
const tree = parse('# Hello *world*') // mdast Root
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
### Entries
|
|
123
|
+
|
|
124
|
+
| Entry | Fallback | Imports |
|
|
125
|
+
|---|---|---|
|
|
126
|
+
| `samemark` | the real parser, `mdast-util-from-markdown` (optional peer) | needs `mdast-util-from-markdown` installed |
|
|
127
|
+
| `samemark/stream` | none; incremental parser for streamed text, stock GFM with optional frontmatter and math | the six helper packages; also CommonJS |
|
|
128
|
+
| `samemark/core` | none: throws on unsupported extensions, unless `remark-parse` was set up before the plugin (it then takes over) or `{fallback}` is passed | only the six helper packages; also `require('samemark/core')` from CommonJS |
|
|
129
|
+
|
|
130
|
+
The main entry also runs a one-time check per processor: three small documents are parsed by both parsers with the
|
|
131
|
+
processor's own extensions; any difference sends that processor to the fallback (reason: "self-check ... failed").
|
|
132
|
+
That guards against a micromark release that changes behavior. It is a tripwire, not a proof.
|
|
133
|
+
|
|
134
|
+
### Options
|
|
135
|
+
|
|
136
|
+
Both `parse(source, options)` and `remarkParseFast(options)` take:
|
|
137
|
+
|
|
138
|
+
| Option | Meaning |
|
|
139
|
+
|---|---|
|
|
140
|
+
| `strict` | Throw instead of falling back (also when the fast parser itself throws). Default `false`. |
|
|
141
|
+
| `onFallback(reason)` | Called with a reason string for every document parsed by the fallback. |
|
|
142
|
+
| `warn` | One `console.warn` per distinct fallback reason. |
|
|
143
|
+
| `fallback(text, ctx)` | Parser to use instead of the default fallback. `ctx` has `extensions`, `mdastExtensions`, `settings`. |
|
|
144
|
+
|
|
145
|
+
`parse` additionally accepts `extensions` and `mdastExtensions` (the arrays you would give `fromMarkdown`). Left out,
|
|
146
|
+
stock GFM is assumed. `source` may be a string, a `Uint8Array` (UTF-8) or an object with `toString()`.
|
|
147
|
+
|
|
148
|
+
Nobody should find out about a fallback in a profiler. `stats` is exported by both entries:
|
|
149
|
+
|
|
150
|
+
```js
|
|
151
|
+
import {stats, check} from 'samemark'
|
|
152
|
+
stats // {fast: 120, fallback: 0, lastFallbackReason: null}
|
|
153
|
+
check(micromarkExtensions, fromMarkdownExtensions) // null when the fast path is exact, else the reason
|
|
154
|
+
```
|
|
155
|
+
|
|
156
|
+
## Streaming for chat UIs
|
|
157
|
+
|
|
158
|
+
Chat UIs re-parse the whole message every time a token arrives. With samemark as the parser that re-parse is 3x to 7x
|
|
159
|
+
cheaper depending on the runtime, with identical output (react-markdown numbers in [Benchmarks](#benchmarks)).
|
|
160
|
+
`samemark/stream` removes the re-parse: it is incremental, and it stays exactly identical to parsing the whole text.
|
|
161
|
+
|
|
162
|
+
```js
|
|
163
|
+
import {createStreamParser} from 'samemark/stream'
|
|
164
|
+
|
|
165
|
+
const stream = createStreamParser() // options: {frontmatter, math}, same meaning as remark-frontmatter / remark-math
|
|
166
|
+
for await (const chunk of tokens) {
|
|
167
|
+
const tree = stream.push(chunk) // mdast Root, equal to parse(everything pushed so far), positions included
|
|
168
|
+
render(tree)
|
|
169
|
+
}
|
|
170
|
+
stream.end() // final tree; stream.reset() starts a new message
|
|
171
|
+
```
|
|
172
|
+
|
|
173
|
+
- After every push the tree equals what `parse()` returns for all text so far, which equals micromark's tree. Chunks
|
|
174
|
+
can split anything: a token, a fence, a table row, a link.
|
|
175
|
+
- Blocks that can no longer change keep their object identity from push to push (an unfinished tail is re-parsed, the
|
|
176
|
+
rest is reused), so `React.memo` keyed on the node skips them. The cost of a push is the new text plus the unfinished
|
|
177
|
+
tail, not the whole message.
|
|
178
|
+
- A later reference or footnote definition can change how an earlier block resolves; the affected blocks are rebuilt
|
|
179
|
+
as new objects, so identity changes exactly when the content can have changed.
|
|
180
|
+
- Supports stock GFM, optionally with frontmatter and math. For other extensions parse the whole message with
|
|
181
|
+
the main entry. Also `require('samemark/stream')` from CommonJS.
|
|
182
|
+
|
|
183
|
+
A React integration lives in [`examples/react-chat/StreamMarkdown.jsx`](examples/react-chat/StreamMarkdown.jsx): a
|
|
184
|
+
`useMarkdownStream()` hook and a `<StreamMarkdown tree>` component that renders the mdast with `mdast-util-to-hast` and
|
|
185
|
+
`hast-util-to-jsx-runtime`, one memoised block per top-level node, applying the same URL sanitising and raw HTML
|
|
186
|
+
handling as react-markdown. Its HTML equals react-markdown's for the whole message (`test/stream.test.mjs`, and the DOM
|
|
187
|
+
comparison in the Chromium benchmark below).
|
|
188
|
+
|
|
189
|
+
```jsx
|
|
190
|
+
import {StreamMarkdown, useMarkdownStream} from './StreamMarkdown.jsx'
|
|
191
|
+
|
|
192
|
+
function Answer({tokens}) { // tokens: an async iterable of strings
|
|
193
|
+
const {tree, push, reset} = useMarkdownStream()
|
|
194
|
+
useEffect(() => {
|
|
195
|
+
reset()
|
|
196
|
+
;(async () => { for await (const chunk of tokens) push(chunk) })()
|
|
197
|
+
}, [tokens])
|
|
198
|
+
return <StreamMarkdown tree={tree} />
|
|
199
|
+
}
|
|
200
|
+
```
|
|
201
|
+
|
|
202
|
+
### Streaming cost
|
|
203
|
+
|
|
204
|
+
Parser only, from `node bench/stream-bench.mjs` in CI (ubuntu-latest, 12 characters per push, a generated LLM-style
|
|
205
|
+
answer): total time to push the whole answer, a full re-parse per push against `createStreamParser`.
|
|
206
|
+
|
|
207
|
+
| Answer length | Full re-parse per push | `samemark/stream` | Speedup |
|
|
208
|
+
|---|---|---|---|
|
|
209
|
+
| 10,451 characters | 422 ms | 66 ms | 6.4x |
|
|
210
|
+
| 30,416 characters | 981 ms | 54 ms | 18.2x |
|
|
211
|
+
| 100,193 characters | 9,232 ms | 133 ms | 69.4x |
|
|
212
|
+
|
|
213
|
+
Identity of that parser, in CI: 48,196,453 pushes over 465,378 documents (corpus replays, fuzz strings, 1-character,
|
|
214
|
+
random, token-like and whole-line chunkings, with and without frontmatter and math), each tree compared with a full
|
|
215
|
+
parse; 0 failing.
|
|
216
|
+
|
|
217
|
+
<picture>
|
|
218
|
+
<source media="(prefers-color-scheme: dark)" srcset="assets/stream-growth-dark.svg">
|
|
219
|
+
<img alt="Line chart of the total time to stream an answer in 12 character pushes at 10,451, 30,416 and 100,193 characters. A full re-parse per push takes 422 ms, 981 ms and 9,232 ms; samemark/stream takes 66 ms, 54 ms and 133 ms." src="assets/stream-growth-light.svg" width="880">
|
|
220
|
+
</picture>
|
|
221
|
+
|
|
222
|
+
React render included (DOM commit, memoised blocks):
|
|
223
|
+
|
|
224
|
+
<!-- stream-bench:start -->
|
|
225
|
+
From CI run [38036885270](https://github.com/anzal1/samemark/actions/runs/38036885270) (Chromium, GitHub Actions runner); reproduce locally with `npm run bench:stream` (jsdom, smaller numbers) or run the `bench` workflow.
|
|
226
|
+
|
|
227
|
+
One chat message streamed 40 characters per push into a real DOM node (Chromium 156.0.8078.4, React 19 production, `flushSync` per push, so each push is parse + React render + DOM commit; layout and paint are not included). (a) react-markdown re-rendering the whole message per push, (b) the same with samemark as the parser, (c) the stream parser with per-block memoisation (`examples/react-chat`). "Total" is the sum of all push times on the main thread. The DOM after the final push, and after every 25th push, is compared with (a): the last column.
|
|
228
|
+
|
|
229
|
+
| CPU | Variant | Pushes | Total | Per push p50 | p95 | worst | Same DOM as (a) |
|
|
230
|
+
|---|---|---|---|---|---|---|---|
|
|
231
|
+
| no throttling | (a) react-markdown, whole message per push | 750 | 11.29 s | 15 ms | 28 ms | 32 ms | 31/31 |
|
|
232
|
+
| no throttling | (b) react-markdown + samemark | 750 | 2.24 s (**5.0x** less) | 2.9 ms | 5.0 ms | 31 ms | 31/31 |
|
|
233
|
+
| no throttling | (c) stream parser + memoised blocks | 750 | 187 ms (**60.3x** less) | 0.2 ms | 0.6 ms | 6.4 ms | 31/31 |
|
|
234
|
+
| 4x slower | (a) react-markdown, whole message per push | 750 | 50.91 s | 67 ms | 126 ms | 141 ms | 31/31 |
|
|
235
|
+
| 4x slower | (b) react-markdown + samemark | 750 | 9.85 s (**5.2x** less) | 13 ms | 24 ms | 122 ms | 31/31 |
|
|
236
|
+
| 4x slower | (c) stream parser + memoised blocks | 750 | 669 ms (**76.1x** less) | 0.8 ms | 2.2 ms | 30 ms | 31/31 |
|
|
237
|
+
| 6x slower | (a) react-markdown, whole message per push | 750 | 79.43 s | 103 ms | 198 ms | 219 ms | 31/31 |
|
|
238
|
+
| 6x slower | (b) react-markdown + samemark | 750 | 15.33 s (**5.2x** less) | 20 ms | 36 ms | 207 ms | 31/31 |
|
|
239
|
+
| 6x slower | (c) stream parser + memoised blocks | 750 | 1.03 s (**77.0x** less) | 1.1 ms | 3.2 ms | 45 ms | 31/31 |
|
|
240
|
+
|
|
241
|
+
|
|
242
|
+
Environment: Node v24.21.0, GitHub Actions ubuntu24 X64 (github-hosted), INTEL(R) XEON(R) PLATINUM 8573C x4, linux x64. Each cell is the median of 7 rounds after one warm-up round, one fresh process per cell.
|
|
243
|
+
|
|
244
|
+
(a) is what a chat UI does today: react-markdown re-parses and re-renders the whole message on every push. (b) swaps the parser only. (c) is `examples/react-chat`.
|
|
245
|
+
<!-- stream-bench:end -->
|
|
246
|
+
|
|
247
|
+
<picture>
|
|
248
|
+
<source media="(prefers-color-scheme: dark)" srcset="assets/stream-dom-dark.svg">
|
|
249
|
+
<img alt="Grouped horizontal bars of main-thread time streaming a 30,000 character answer into the DOM in Chromium, at no CPU throttling, 4x and 6x. react-markdown takes 11.29 s, 50.91 s and 79.43 s; with samemark 2.24 s, 9.85 s and 15.33 s; samemark/stream with memoised blocks 187 ms, 669 ms and 1.03 s." src="assets/stream-dom-light.svg" width="880">
|
|
250
|
+
</picture>
|
|
251
|
+
|
|
252
|
+
## Examples
|
|
253
|
+
|
|
254
|
+
Each snippet was run against this repository; the unified, frontmatter and fallback ones print what is shown.
|
|
255
|
+
|
|
256
|
+
**unified pipeline** (markdown to HTML):
|
|
257
|
+
|
|
258
|
+
```js
|
|
259
|
+
import {unified} from 'unified'
|
|
260
|
+
import remarkGfm from 'remark-gfm'
|
|
261
|
+
import remarkRehype from 'remark-rehype'
|
|
262
|
+
import rehypeStringify from 'rehype-stringify'
|
|
263
|
+
import remarkParseFast from 'samemark'
|
|
264
|
+
|
|
265
|
+
const file = await unified()
|
|
266
|
+
.use(remarkParseFast)
|
|
267
|
+
.use(remarkGfm)
|
|
268
|
+
.use(remarkRehype)
|
|
269
|
+
.use(rehypeStringify)
|
|
270
|
+
.process('# Hello *world*\n\n| a | b |\n|---|---|\n| ~~x~~ | y |')
|
|
271
|
+
console.log(String(file))
|
|
272
|
+
// <h1>Hello <em>world</em></h1>
|
|
273
|
+
// <table> ... <td><del>x</del></td> ... </table>
|
|
274
|
+
```
|
|
275
|
+
|
|
276
|
+
**react-markdown, one line changed:**
|
|
277
|
+
|
|
278
|
+
```jsx
|
|
279
|
+
import Markdown from 'react-markdown'
|
|
280
|
+
import remarkGfm from 'remark-gfm'
|
|
281
|
+
import remarkParseFast from 'samemark'
|
|
282
|
+
|
|
283
|
+
const plugins = [remarkParseFast, remarkGfm] // module level: not recreated per render
|
|
284
|
+
export const Answer = ({text}) => <Markdown remarkPlugins={plugins}>{text}</Markdown>
|
|
285
|
+
```
|
|
286
|
+
|
|
287
|
+
**Streaming chat with `samemark/stream`.** In React use [`examples/react-chat`](examples/react-chat/StreamMarkdown.jsx)
|
|
288
|
+
(the [live demo](https://anzal1.github.io/samemark/) runs it). Without a framework:
|
|
289
|
+
|
|
290
|
+
```js
|
|
291
|
+
import {createStreamParser} from 'samemark/stream'
|
|
292
|
+
|
|
293
|
+
const stream = createStreamParser()
|
|
294
|
+
let tree
|
|
295
|
+
for (const chunk of ['# Hel', 'lo\n\nSome **bo', 'ld** text\n\n- a\n- b\n']) {
|
|
296
|
+
tree = stream.push(chunk) // same tree as parse(everything so far); finished blocks keep their identity
|
|
297
|
+
}
|
|
298
|
+
console.log(tree.children.map(node => node.type)) // [ 'heading', 'paragraph', 'list' ]
|
|
299
|
+
```
|
|
300
|
+
|
|
301
|
+
**Astro:**
|
|
302
|
+
|
|
303
|
+
```js
|
|
304
|
+
// astro.config.mjs
|
|
305
|
+
import remarkParseFast from 'samemark'
|
|
306
|
+
export default {markdown: {remarkPlugins: [remarkParseFast]}}
|
|
307
|
+
```
|
|
308
|
+
|
|
309
|
+
**Frontmatter and math** stay on the fast path:
|
|
310
|
+
|
|
311
|
+
```js
|
|
312
|
+
import remarkFrontmatter from 'remark-frontmatter'
|
|
313
|
+
import remarkMath from 'remark-math'
|
|
314
|
+
import {stats} from 'samemark'
|
|
315
|
+
|
|
316
|
+
const file = await unified()
|
|
317
|
+
.use(remarkParseFast)
|
|
318
|
+
.use(remarkGfm)
|
|
319
|
+
.use(remarkFrontmatter)
|
|
320
|
+
.use(remarkMath)
|
|
321
|
+
.use(remarkRehype)
|
|
322
|
+
.use(rehypeStringify)
|
|
323
|
+
.process('---\ntitle: Notes\n---\n\nEuler: $e^{i\\pi}+1=0$\n')
|
|
324
|
+
console.log(String(file)) // <p>Euler: <code class="language-math math-inline">e^{i\pi}+1=0</code></p>
|
|
325
|
+
console.log(stats) // { fast: 1, fallback: 0, lastFallbackReason: null }
|
|
326
|
+
```
|
|
327
|
+
|
|
328
|
+
**`samemark/core` in the browser** (no fallback import; the `browser` build is about 20 KB gzipped, see
|
|
329
|
+
[Bundle size](#bundle-size)). It needs a bundler that honours package `exports`, such as esbuild or Vite:
|
|
330
|
+
|
|
331
|
+
```js
|
|
332
|
+
import {parse} from 'samemark/core'
|
|
333
|
+
|
|
334
|
+
const tree = parse('# Hello *world*') // mdast Root, GFM
|
|
335
|
+
```
|
|
336
|
+
|
|
337
|
+
**Checking fallbacks** so nobody finds one in a profiler:
|
|
338
|
+
|
|
339
|
+
```js
|
|
340
|
+
import remarkDirective from 'remark-directive'
|
|
341
|
+
import remarkParseFast, {stats} from 'samemark'
|
|
342
|
+
|
|
343
|
+
const reasons = []
|
|
344
|
+
await unified()
|
|
345
|
+
.use(remarkParseFast, {onFallback: reason => reasons.push(reason)})
|
|
346
|
+
.use(remarkGfm)
|
|
347
|
+
.use(remarkDirective) // not in the supported set: the real parser takes over
|
|
348
|
+
.use(remarkRehype)
|
|
349
|
+
.use(rehypeStringify)
|
|
350
|
+
.process(':::note\nhi\n:::')
|
|
351
|
+
console.log(reasons) // ['micromark extension not in the supported set (GFM, frontmatter, math): ...']
|
|
352
|
+
console.log(stats) // { fast: 0, fallback: 1, lastFallbackReason: '...' }
|
|
353
|
+
```
|
|
354
|
+
|
|
355
|
+
Pass `{strict: true}` to throw instead of falling back.
|
|
356
|
+
|
|
357
|
+
## Compatibility
|
|
358
|
+
|
|
359
|
+
Exact statement. The fast path is verified, and only verified, against these pinned versions
|
|
360
|
+
(`package.json` `devDependencies`, installed by `npm ci`):
|
|
361
|
+
|
|
362
|
+
| Package | Version |
|
|
363
|
+
|---|---|
|
|
364
|
+
| micromark | 4.0.3 |
|
|
365
|
+
| micromark-extension-gfm | 3.0.0 |
|
|
366
|
+
| mdast-util-gfm | 3.1.0 |
|
|
367
|
+
| mdast-util-from-markdown | 2.1.0 |
|
|
368
|
+
| remark-parse | 11.0.0 |
|
|
369
|
+
| remark-gfm | 4.0.1 |
|
|
370
|
+
| unified | 11.0.5 |
|
|
371
|
+
|
|
372
|
+
Other versions are not claimed. A newer micromark that changes output is caught by the main entry's self-check
|
|
373
|
+
(it then uses the fallback), and by `npm test` in this repository, which fails if the stock GFM construct set
|
|
374
|
+
changes (`node scripts/gen-signature.mjs --check`).
|
|
375
|
+
|
|
376
|
+
What counts as supported is decided by construct names, not by function source, so minified bundles of
|
|
377
|
+
micromark keep the fast path (tested with an esbuild `--minify` bundle that includes the frontmatter and math
|
|
378
|
+
extensions). Settings that hide in closures (`singleTilde: false`, `singleDollarTextMath: false`, which fences a
|
|
379
|
+
frontmatter matter uses, `anywhere: true`) are read by running the real tokenizers on probe strings; the result is
|
|
380
|
+
compared with what the stock configuration does. Any unknown construct, any `disable` entry, a custom or `anywhere`
|
|
381
|
+
frontmatter matter, or a missing GFM piece means fallback.
|
|
382
|
+
|
|
383
|
+
### Which setups take the fast path
|
|
384
|
+
|
|
385
|
+
Generated by `node scripts/compat-table.mjs` from the same cases `npm test` asserts (each fast row is also compared with `remark-parse` on a document that exercises it):
|
|
386
|
+
|
|
387
|
+
<!-- compat:start -->
|
|
388
|
+
| Setup | Path | Reported reason |
|
|
389
|
+
|---|---|---|
|
|
390
|
+
| remark-gfm | **fast** | |
|
|
391
|
+
| remark-gfm twice | **fast** | |
|
|
392
|
+
| remark-gfm with `tablePipeAlign` / `tableCellPadding` (serialisation options only) | **fast** | |
|
|
393
|
+
| remark-gfm + remark-breaks (tree transform only) | **fast** | |
|
|
394
|
+
| remark-gfm + remark-smartypants (tree transform only) | **fast** | |
|
|
395
|
+
| remark-gfm with `singleTilde: false` | fallback (real parser) | `strikethrough is not the stock GFM one (singleTilde: false or unknown tokenizer)` |
|
|
396
|
+
| remark-gfm + remark-frontmatter (yaml) | **fast** | |
|
|
397
|
+
| remark-gfm + remark-frontmatter `['yaml', 'toml']` | **fast** | |
|
|
398
|
+
| remark-gfm + remark-frontmatter `'toml'` | **fast** | |
|
|
399
|
+
| remark-gfm + remark-frontmatter with a custom `fence` (same type, other fence) | fallback (real parser) | `frontmatter fences differ from the stock yaml matters (custom or `anywhere` matter)` |
|
|
400
|
+
| remark-gfm + remark-frontmatter with a custom matter type | fallback (real parser) | `micromark extension not in the supported set (GFM, frontmatter, math): flow\|unnamed construct at code 60 ('<')` |
|
|
401
|
+
| remark-gfm + remark-frontmatter with `anywhere: true` | fallback (real parser) | `frontmatter fences differ from the stock yaml matters (custom or `anywhere` matter)` |
|
|
402
|
+
| remark-gfm + remark-math | **fast** | |
|
|
403
|
+
| remark-gfm + remark-math with `singleDollarTextMath: false` | **fast** | |
|
|
404
|
+
| remark-gfm + remark-frontmatter + remark-math | **fast** | |
|
|
405
|
+
| remark-frontmatter + remark-math without remark-gfm | fallback (real parser) | `GFM is not fully enabled (missing micromark construct attentionMarkers\|126)` |
|
|
406
|
+
| remark-gfm + remark-directive | fallback (real parser) | `micromark extension not in the supported set (GFM, frontmatter, math): flow\|unnamed construct at code 58 (':')` |
|
|
407
|
+
| remark-gfm + remark-frontmatter + remark-directive | fallback (real parser) | `micromark extension not in the supported set (GFM, frontmatter, math): flow\|unnamed construct at code 45 ('-')` |
|
|
408
|
+
| remark-mdx (MDX) | fallback (real parser) | `GFM is not fully enabled (missing micromark construct attentionMarkers\|126)` |
|
|
409
|
+
| remark-gfm + remark-mdx | fallback (real parser) | `micromark extension not in the supported set (GFM, frontmatter, math): disable\|{"null":["autolink","codeIndented","htmlFlow","htmlText"]}` |
|
|
410
|
+
| no remark-gfm (CommonMark only) | fallback (real parser) | `no GFM extensions (CommonMark-only parsing needs the fallback)` |
|
|
411
|
+
| unknown micromark construct | fallback (real parser) | `micromark extension not in the supported set (GFM, frontmatter, math): text\|atSign` |
|
|
412
|
+
| micromark `disable` entry | fallback (real parser) | `micromark extension not in the supported set (GFM, frontmatter, math): disable\|{"null":["codeIndented"]}` |
|
|
413
|
+
<!-- compat:end -->
|
|
414
|
+
|
|
415
|
+
Frameworks. **Astro** (plain markdown, `remark-gfm` default): fast. **Starlight** adds `remark-directive` for asides,
|
|
416
|
+
so it falls back and gains nothing. **Docusaurus** and **MDX** compile through `micromark-extension-mdxjs`: fallback.
|
|
417
|
+
`@mdx-js/mdx` with `format: 'md'` and `remark-gfm`: fast.
|
|
418
|
+
|
|
419
|
+
## How identity is verified
|
|
420
|
+
|
|
421
|
+
Correctness at a glance (every row compares the full tree, positions included, against `mdast-util-from-markdown` + GFM; details and commands in the table below):
|
|
422
|
+
|
|
423
|
+
| Check | Scale | Result | Reproduce |
|
|
424
|
+
|---|---|---|---|
|
|
425
|
+
| CommonMark and GFM spec examples | 652 CommonMark 0.31.2, 672 GFM spec and 30 GFM extension examples | all identical | `npm run check` |
|
|
426
|
+
| Pinned docs corpora and synthetic corpus | Node.js API docs and changelog, Rust book, Vite docs, React changelog, 240 generated documents | all identical | `npm run corpus && npm run check` |
|
|
427
|
+
| README corpus | 4,780 npm package READMEs | all identical at the last run | `node bench/readme-corpus.mjs && node bench/run.mjs --readme --nospeed` |
|
|
428
|
+
| Held-out files | 10,893 non-README markdown files | all identical at the last run | `node bench/heldout.mjs --roots=DIR` |
|
|
429
|
+
| Differential fuzz | tens of millions of generated inputs (the weekly full matrix is about 60M) | 0 failing; known corner shapes are listed in [KNOWN_GAPS.md](KNOWN_GAPS.md) | `node bench/fuzz.mjs --mode=regress`, [`ci.yml`](.github/workflows/ci.yml), [`full.yml`](.github/workflows/full.yml) |
|
|
430
|
+
| Stream parser, tree after every push | 48,196,453 pushes over 465,378 documents | 0 failing | `node bench/stream-diff.mjs --mode=regress`, [`full.yml`](.github/workflows/full.yml) |
|
|
431
|
+
|
|
432
|
+
"Identical" means `JSON`-equal trees including every `position`, not "looks the same" (`firstDiff` in
|
|
433
|
+
`bench/common.mjs`; `undefined` fields are treated as absent).
|
|
434
|
+
|
|
435
|
+
| Check | What | Command |
|
|
436
|
+
|---|---|---|
|
|
437
|
+
| Spec | all 652 CommonMark 0.31.2 examples, the 672 GFM spec examples and the 30 GFM extension examples, full tree compared | `npm run check` |
|
|
438
|
+
| Docs and synthetic corpora | Node.js API docs and changelog, Rust book, Vite docs, a React changelog (pinned commits, fetched in CI, `bench/corpora.json`) and 240 generated documents (`bench/synthetic.mjs`), full tree compared | `npm run corpus && npm run check` |
|
|
439
|
+
| README corpus (development) | 4,780 README.md files of npm packages (30 MB) collected on the author's machine, full tree compared (all identical at the last run). 4,648 of them are in `bench/readme-corpus-manifest.json` (package, version, sha256) and can be rebuilt from npm, slowly | `node bench/readme-corpus.mjs && node bench/run.mjs --readme --nospeed` |
|
|
440
|
+
| CommonMark spec | all 652 examples of spec 0.31.2 | `npm run check` |
|
|
441
|
+
| Held-out | 10,893 non-README markdown files from the author's machine, never used while writing the parser (all identical at the last run). The file list is machine specific, so others cannot repeat this exact set; pass your own list | `node bench/heldout.mjs --roots=DIR` |
|
|
442
|
+
| Differential fuzzing | exhaustive strings (17 characters, length up to 6), random, token-biased, mutations of corpus paragraphs, spec crossovers | `node bench/fuzz.mjs --mode=exhaustive --maxlen=5 --shard=0/64` (the full matrix, about 60M cases in 60 jobs, is `.github/workflows/full.yml`, weekly or by hand; `ci.yml` runs about 2M cases on every push) |
|
|
443
|
+
| Frontmatter and math | the corpus and spec examples with yaml / toml / custom frontmatter prepended (CRLF, unclosed fences, not at the start) and with math enabled, plus extension-aware fuzzing | `node bench/frontmatter.mjs --sample=24` |
|
|
444
|
+
| Streaming | `samemark/stream` replays documents in random chunk sizes; the tree after every push must equal a full parse, and micromark's at checkpoints (corpus, spec examples, edge inputs, definitions arriving late) | `node bench/stream-diff.mjs --mode=regress`, `node --test test/stream.test.mjs` |
|
|
445
|
+
| Regressions | every input that ever failed | `node bench/fuzz.mjs --mode=regress` |
|
|
446
|
+
| Edge inputs | CRLF, CR, BOM, NUL, lone surrogates, empty, 50k-character lines, deep nesting, unicode, tabs; 10 MB inputs in CI | `node --test test/edge.test.mjs`; `BIG=1` for 10 MB |
|
|
447
|
+
| Packaging | minified bundle keeps the fast path, edge-runtime VM, Chromium, Deno, Node 18 to 24, packed tarball in an empty project | `npm test`; `ci.yml` (Node 18 to 24, packed tarball), `full.yml` (Deno, Chromium, 10 MB inputs) |
|
|
448
|
+
|
|
449
|
+
Third-party text is never committed. The docs corpora are fetched at pinned commits (licences: MIT for Node.js, Vite and
|
|
450
|
+
React, MIT or Apache-2.0 for the Rust book, BSD-2-Clause code with CC-BY-SA spec text for the GFM spec examples), the
|
|
451
|
+
README corpus is rebuilt from npm by sha256 manifest, and the synthetic corpus is generated from a seed.
|
|
452
|
+
|
|
453
|
+
## Known differences
|
|
454
|
+
|
|
455
|
+
On the tested set (spec examples, corpora, held-out, edge inputs) there are none. Differential fuzzing did find
|
|
456
|
+
shapes that differ, all block-level corners around footnote definitions and lists; they are listed with minimal
|
|
457
|
+
inputs in [KNOWN_GAPS.md](KNOWN_GAPS.md) and none occurs in a real document so far. What is not covered by the
|
|
458
|
+
identity claim: non-GFM extensions (they fall back, by design), non-string input types, and pipelines other than the
|
|
459
|
+
ones in the compatibility table.
|
|
460
|
+
|
|
461
|
+
Behavior that is different on purpose: micromark throws a `RangeError` on roughly 10,000 levels of nesting; samemark
|
|
462
|
+
parses it. Large single inputs: micromark's time grows faster than linearly with document size (a mixed document took
|
|
463
|
+
1.6 s at 400 KB, 5.1 s at 800 KB and 19 s at 1.6 MB on a fast laptop; one 500 KB paragraph of inline markup takes 16 s;
|
|
464
|
+
a 10 MB comparison ran 42 minutes and exhausted a 6 GB heap), while samemark parsed a 10 MB mixed document in about
|
|
465
|
+
0.4 s. That is why the full-tree comparison in the edge tests stops at 2 to 3 MB and the 10 MB cases only check that
|
|
466
|
+
samemark finishes and does not throw.
|
|
467
|
+
|
|
468
|
+
## Benchmarks
|
|
469
|
+
|
|
470
|
+
**Which numbers come from which input.** The tables below were produced by the `bench` workflow (and the parser-level
|
|
471
|
+
stream benchmark by `bench/stream-bench.mjs`) while the project lived in a private repository. Rows named
|
|
472
|
+
"README corpus" ran on the 4,780-file development corpus of npm package READMEs (a 1 in 4 or 1 in 16 sample for the
|
|
473
|
+
slower tests; 4,648 of the files are rebuildable from `bench/readme-corpus-manifest.json`). Rows named after a
|
|
474
|
+
repository (`node-api-docs`, `rust-book`, `vite-docs`, `node-changelog-v22`, `react-changelog`) ran on the pinned
|
|
475
|
+
commits in `bench/corpora.json`; anyone can fetch exactly those with `npm run corpus`. The react-markdown tables use
|
|
476
|
+
the generated chat-style answer (`bench/perf/chat-text.mjs`), the Astro site uses the three docs sets, and
|
|
477
|
+
`npm run bench` on a fresh checkout measures the pinned docs sets plus the synthetic corpus (and the README corpus
|
|
478
|
+
with `--readme` once it is rebuilt). Runner CPUs differ between jobs and runs, so compare the ratios within a table, not
|
|
479
|
+
absolute times across tables.
|
|
480
|
+
|
|
481
|
+
<!-- bench:start -->
|
|
482
|
+
|
|
483
|
+
Environment: Node v24.21.0, GitHub Actions ubuntu24 X64 (github-hosted), AMD EPYC 7763 64-Core Processor x4, linux x64. Each cell is the median of 7 rounds after one warm-up round, one fresh process per cell.
|
|
484
|
+
|
|
485
|
+
### End to end: markdown to HTML with unified
|
|
486
|
+
|
|
487
|
+
`unified().use(parser).use(remarkGfm).use(remarkRehype).use(rehypeStringify)`, with `remark-parse` or with samemark as the parser. Output HTML is compared file by file.
|
|
488
|
+
|
|
489
|
+
| Input | Files | Size | remark-parse | samemark | Speedup | Parse share before | Parse share after | Identical HTML |
|
|
490
|
+
|---|---|---|---|---|---|---|---|---|
|
|
491
|
+
| node-api-docs | 70 | 4.75 MB | 5.00 s | 486 ms | **10.3x** | 92% | 33% | 70/70 |
|
|
492
|
+
| node-changelog-v22 | 1 | 0.88 MB | 1.33 s | 171 ms | **7.8x** | 89% | 21% | 1/1 |
|
|
493
|
+
| rust-book | 112 | 1.22 MB | 1.07 s | 125 ms | **8.6x** | 93% | 40% | 112/112 |
|
|
494
|
+
| vite-docs | 57 | 0.56 MB | 570 ms | 73 ms | **7.8x** | 92% | 29% | 57/57 |
|
|
495
|
+
| react-changelog | 1 | 0.30 MB | 498 ms | 71 ms | **7.0x** | 90% | 19% | 1/1 |
|
|
496
|
+
| synthetic corpus (240 generated documents) | 66 | 0.25 MB | 546 ms | 89 ms | **6.2x** | 88% | 22% | 66/66 |
|
|
497
|
+
|
|
498
|
+
Content pipeline, as a docs or blog site runs it: `remark-gfm`, `remark-frontmatter` (yaml) and `remark-math` before `remark-rehype`. `vite-docs` has real frontmatter in every page; the synthetic (or README) documents get a generated block so every file has one. Both sides are compared on the HTML.
|
|
499
|
+
|
|
500
|
+
| Input | Files | Size | remark-parse | samemark | Speedup | Identical HTML | samemark path |
|
|
501
|
+
|---|---|---|---|---|---|---|---|
|
|
502
|
+
| synthetic corpus (240 generated documents) with a frontmatter block added | 66 | 0.26 MB | 559 ms | 95 ms | **5.9x** | 66/66 | fast |
|
|
503
|
+
| vite-docs | 57 | 0.56 MB | 543 ms | 66 ms | **8.2x** | 57/57 | fast |
|
|
504
|
+
| react-changelog | 1 | 0.30 MB | 517 ms | 70 ms | **7.3x** | 1/1 | fast |
|
|
505
|
+
|
|
506
|
+
For context, other converters on the same input (different pipeline and output: HTML strings, no mdast, not drop-in replacements):
|
|
507
|
+
|
|
508
|
+
| Input | markdown-it `render` (HTML) | marked `parse` (HTML) |
|
|
509
|
+
|---|---|---|
|
|
510
|
+
| node-api-docs | 477 ms | 449 ms |
|
|
511
|
+
| node-changelog-v22 | 182 ms | 144 ms |
|
|
512
|
+
| rust-book | 122 ms | 103 ms |
|
|
513
|
+
| vite-docs | 74 ms | 66 ms |
|
|
514
|
+
| react-changelog | 71 ms | 57 ms |
|
|
515
|
+
| synthetic corpus (240 generated documents) | 89 ms | 70 ms |
|
|
516
|
+
|
|
517
|
+
|
|
518
|
+
Environment: Node v24.21.0, GitHub Actions ubuntu24 X64 (github-hosted), AMD EPYC 7763 64-Core Processor x4, linux x64. Each cell is the median of 7 rounds after one warm-up round, one fresh process per cell.
|
|
519
|
+
|
|
520
|
+
### react-markdown (desktop, Node, `renderToStaticMarkup`, React production build)
|
|
521
|
+
|
|
522
|
+
react-markdown 10.1.0 builds a new unified processor on every render (`unified().use(remarkParse).use(remarkPlugins)...`), so `remarkPlugins={[remarkParseFast, remarkGfm]}` replaces its parser on every render. The input is a synthetic chat-style answer (headings, lists, code, tables, links; `bench/perf/chat-text.mjs`). Output HTML is compared at 120 prefix lengths, partial constructs included: 120/120 identical. samemark path: 6148 fast, 0 fallback.
|
|
523
|
+
|
|
524
|
+
One update at a fixed message length (this is also the cost of rendering a finished long answer):
|
|
525
|
+
|
|
526
|
+
| Message length | react-markdown default | with samemark | Speedup |
|
|
527
|
+
|---|---|---|---|
|
|
528
|
+
| 10,000 characters | 36 ms | 5 ms | **6.5x** |
|
|
529
|
+
| 30,000 characters | 92 ms | 13 ms | **7.0x** |
|
|
530
|
+
|
|
531
|
+
Streaming simulation: the message grows by a few characters at a time and the whole message is re-rendered after each step, as chat apps do:
|
|
532
|
+
|
|
533
|
+
| Stream | Updates | default: total / mean / worst update | samemark: total / mean / worst update | Speedup |
|
|
534
|
+
|---|---|---|---|---|
|
|
535
|
+
| to 10,000 characters, 20 per step | 500 | 3.25 s / 6 ms / 16 ms | 676 ms / 1 ms / 4 ms | **4.8x** |
|
|
536
|
+
| to 30,000 characters, 20 per step | 1500 | 27.98 s / 19 ms / 47 ms | 5.31 s / 4 ms / 9 ms | **5.3x** |
|
|
537
|
+
|
|
538
|
+
|
|
539
|
+
### react-markdown in Chromium, CPU throttled as a phone proxy
|
|
540
|
+
|
|
541
|
+
Chromium 156.0.8078.4, react-markdown 10.1.0 production bundle, `renderToStaticMarkup` in the page. CPU throttling uses DevTools `Emulation.setCPUThrottlingRate` (4x and 6x slowdown of the main thread). It is a proxy for a slower device, not a measurement of one. Same chat-style input as the desktop table; streaming steps are 100 characters here to keep the throttled runs short.
|
|
542
|
+
|
|
543
|
+
| CPU | Update at 10,000 characters: default / samemark | Update at 30,000 characters: default / samemark | Stream to 10,000: mean update default / samemark | Stream to 30,000: mean update default / samemark | Identical HTML |
|
|
544
|
+
|---|---|---|---|---|---|
|
|
545
|
+
| no throttling | 9.6 ms / 2.0 ms (**4.8x**) | 29 ms / 5.4 ms (**5.4x**) | 5.1 ms / 1.0 ms (**4.9x**) | 14 ms / 2.7 ms (**5.3x**) | 60/60 |
|
|
546
|
+
| 4x slower | 45 ms / 7.9 ms (**5.7x**) | 145 ms / 22 ms (**6.6x**) | 24 ms / 4.3 ms (**5.5x**) | 65 ms / 12 ms (**5.5x**) | 60/60 |
|
|
547
|
+
| 6x slower | 70 ms / 12 ms (**6.0x**) | 200 ms / 35 ms (**5.8x**) | 35 ms / 6.5 ms (**5.3x**) | 102 ms / 18 ms (**5.8x**) | 60/60 |
|
|
548
|
+
|
|
549
|
+
### Streaming a 30,000 character answer into the DOM
|
|
550
|
+
|
|
551
|
+
One chat message streamed 40 characters per push into a real DOM node (Chromium 156.0.8078.4, React 19 production, `flushSync` per push, so each push is parse + React render + DOM commit; layout and paint are not included). (a) react-markdown re-rendering the whole message per push, (b) the same with samemark as the parser, (c) the stream parser with per-block memoisation (`examples/react-chat`). "Total" is the sum of all push times on the main thread. The DOM after the final push, and after every 25th push, is compared with (a): the last column.
|
|
552
|
+
|
|
553
|
+
| CPU | Variant | Pushes | Total | Per push p50 | p95 | worst | Same DOM as (a) |
|
|
554
|
+
|---|---|---|---|---|---|---|---|
|
|
555
|
+
| no throttling | (a) react-markdown, whole message per push | 750 | 11.29 s | 15 ms | 28 ms | 32 ms | 31/31 |
|
|
556
|
+
| no throttling | (b) react-markdown + samemark | 750 | 2.24 s (**5.0x** less) | 2.9 ms | 5.0 ms | 31 ms | 31/31 |
|
|
557
|
+
| no throttling | (c) stream parser + memoised blocks | 750 | 187 ms (**60.3x** less) | 0.2 ms | 0.6 ms | 6.4 ms | 31/31 |
|
|
558
|
+
| 4x slower | (a) react-markdown, whole message per push | 750 | 50.91 s | 67 ms | 126 ms | 141 ms | 31/31 |
|
|
559
|
+
| 4x slower | (b) react-markdown + samemark | 750 | 9.85 s (**5.2x** less) | 13 ms | 24 ms | 122 ms | 31/31 |
|
|
560
|
+
| 4x slower | (c) stream parser + memoised blocks | 750 | 669 ms (**76.1x** less) | 0.8 ms | 2.2 ms | 30 ms | 31/31 |
|
|
561
|
+
| 6x slower | (a) react-markdown, whole message per push | 750 | 79.43 s | 103 ms | 198 ms | 219 ms | 31/31 |
|
|
562
|
+
| 6x slower | (b) react-markdown + samemark | 750 | 15.33 s (**5.2x** less) | 20 ms | 36 ms | 207 ms | 31/31 |
|
|
563
|
+
| 6x slower | (c) stream parser + memoised blocks | 750 | 1.03 s (**77.0x** less) | 1.1 ms | 3.2 ms | 45 ms | 31/31 |
|
|
564
|
+
|
|
565
|
+
|
|
566
|
+
Environment: Node v24.21.0, GitHub Actions ubuntu24 X64 (github-hosted), INTEL(R) XEON(R) PLATINUM 8573C x4, linux x64. Each cell is the median of 7 rounds after one warm-up round, one fresh process per cell.
|
|
567
|
+
|
|
568
|
+
### Parse only
|
|
569
|
+
|
|
570
|
+
Same mdast on both sides (`mdast-util-from-markdown` with `micromark-extension-gfm` and `mdast-util-gfm`, the reference). "Identical" counts files whose full tree, positions included, is equal.
|
|
571
|
+
|
|
572
|
+
| Input | Files | Size | mdast-util-from-markdown + GFM | samemark | Speedup | Identical |
|
|
573
|
+
|---|---|---|---|---|---|---|
|
|
574
|
+
| node-api-docs | 70 | 4.75 MB | 4.34 s | 121 ms | **35.7x** | 70/70 |
|
|
575
|
+
| node-changelog-v22 | 1 | 0.88 MB | 1.06 s | 23 ms | **46.8x** | 1/1 |
|
|
576
|
+
| rust-book | 112 | 1.22 MB | 921 ms | 32 ms | **29.1x** | 112/112 |
|
|
577
|
+
| vite-docs | 57 | 0.56 MB | 486 ms | 17 ms | **27.9x** | 57/57 |
|
|
578
|
+
| react-changelog | 1 | 0.30 MB | 414 ms | 13 ms | **31.7x** | 1/1 |
|
|
579
|
+
| synthetic corpus (240 generated documents) | 240 | 1.08 MB | 1.54 s | 51 ms | **30.3x** | 240/240 |
|
|
580
|
+
|
|
581
|
+
Other parsers on the same input. Different output: markdown-it returns a token list and marked returns its own token tree, neither is mdast, and neither is a drop-in replacement. Shown for scale only.
|
|
582
|
+
|
|
583
|
+
| Input | markdown-it `parse` (tokens) | marked `lexer` (tokens) |
|
|
584
|
+
|---|---|---|
|
|
585
|
+
| node-api-docs | 423 ms | 298 ms |
|
|
586
|
+
| node-changelog-v22 | 154 ms | 81 ms |
|
|
587
|
+
| rust-book | 94 ms | 63 ms |
|
|
588
|
+
| vite-docs | 64 ms | 34 ms |
|
|
589
|
+
| react-changelog | 61 ms | 30 ms |
|
|
590
|
+
| synthetic corpus (240 generated documents) | 210 ms | 139 ms |
|
|
591
|
+
|
|
592
|
+
|
|
593
|
+
Environment: Node v24.21.0, GitHub Actions ubuntu24 X64 (github-hosted), AMD EPYC 7763 64-Core Processor x4, linux x64. Each cell is the median of 7 rounds after one warm-up round, one fresh process per cell.
|
|
594
|
+
|
|
595
|
+
### MDX (`@mdx-js/mdx` `compile`)
|
|
596
|
+
|
|
597
|
+
Plain `format: 'md'` with `remark-gfm`, default parser versus samemark as the parser. `format: 'mdx'` loads `micromark-extension-mdxjs`, which is not in the stock GFM set, so samemark falls back there by design; it is measured on the files that are valid MDX to show the overhead is nil. Compiled output is compared.
|
|
598
|
+
|
|
599
|
+
| Input | Format | Files compiled | default | with samemark | Speedup | Identical output | samemark path |
|
|
600
|
+
|---|---|---|---|---|---|---|---|
|
|
601
|
+
| node-api-docs | md | 70 of 70 | 6.21 s | 1.55 s | 4.0x | yes | 280 fast, 0 fallback |
|
|
602
|
+
| node-api-docs | mdx | 0 of 70 | n/a | n/a | n/a | n/a | no file in this set is valid MDX |
|
|
603
|
+
| node-changelog-v22 | md | 1 of 1 | 1.68 s | 509 ms | 3.3x | yes | 4 fast, 0 fallback |
|
|
604
|
+
| node-changelog-v22 | mdx | 0 of 1 | n/a | n/a | n/a | n/a | no file in this set is valid MDX |
|
|
605
|
+
| rust-book | md | 112 of 112 | 1.34 s | 367 ms | 3.7x | yes | 448 fast, 0 fallback |
|
|
606
|
+
| rust-book | mdx | 29 of 112 | 245 ms | 251 ms | 1.0x | yes | 0 fast, 116 fallback |
|
|
607
|
+
|
|
608
|
+
|
|
609
|
+
### Real site: Astro 6.4.8 building 239 real markdown pages
|
|
610
|
+
|
|
611
|
+
rust-book, vite-docs, node-api-docs (6.53 MB of markdown), stock Astro markdown pipeline (remark-parse, remark-gfm, smartypants, Shiki, rehype-raw). Cold build each time (`dist`, `.astro` and the Vite cache removed), 3 builds per variant alternating, median shown, Node v24.21.0, AMD EPYC 7763 64-Core Processor x4.
|
|
612
|
+
|
|
613
|
+
| Parser | Build wall clock | Time inside the markdown parser | Parser share of build | Documents parsed | Fell back |
|
|
614
|
+
|---|---|---|---|---|---|
|
|
615
|
+
| remark-parse (default) | 29.6 s | 7.07 s | 23.9% | 239 | n/a |
|
|
616
|
+
| samemark | 22.9 s | 0.50 s | 2.2% | 239 | 0 |
|
|
617
|
+
|
|
618
|
+
Build time saved: 6.7 s (22.7%). Parser time saved: 6.57 s.
|
|
619
|
+
|
|
620
|
+
Individual builds (s): default 29.2, 29.6, 29.7; samemark 22.7, 22.9, 23.3.
|
|
621
|
+
|
|
622
|
+
<!-- bench:end -->
|
|
623
|
+
|
|
624
|
+
**Reading the react-markdown numbers.** The speedup per update is smaller than the parse-only number because React rendering and the hast to JSX step are unchanged. It differs by runtime: Node 6.5x to 7.0x, unthrottled Chromium 4.8x to 5.4x, throttled Chromium 5.3x to 6.6x. I do not have a verified explanation for why the ratio grows under CPU throttling (throttling slows the main thread; garbage collection and compilation threads are not slowed the same way, which is a guess, not a measurement). The absolute numbers are the practical result: a finished 30,000 character answer renders in 35 ms instead of 200 ms at 6x throttle.
|
|
625
|
+
|
|
626
|
+
Reproduce:
|
|
627
|
+
|
|
628
|
+
```sh
|
|
629
|
+
npm ci
|
|
630
|
+
npm run bench:small # a few seconds: the synthetic corpus sample plus one changelog, laptop sized
|
|
631
|
+
npm run bench # the full run (what CI ran), fetches the pinned documents first
|
|
632
|
+
```
|
|
633
|
+
|
|
634
|
+
`npm run bench` prints the exact tables above. Method: each cell is one fresh Node process (own JIT and heap),
|
|
635
|
+
one warm-up round, then the median of the stated number of rounds over all files of the input; GC is forced between
|
|
636
|
+
rounds; both sides are compared for identical output before timings are reported (the run fails otherwise).
|
|
637
|
+
markdown-it and marked are included for scale only: they produce tokens or HTML, not mdast, and are not drop-in
|
|
638
|
+
replacements. Documents: `bench/corpora.json` pins repository, commit and sha256 of every set.
|
|
639
|
+
|
|
640
|
+
## Bundle size
|
|
641
|
+
|
|
642
|
+
Minified with esbuild, ESM, platform neutral. `node scripts/size.mjs` reproduces the table.
|
|
643
|
+
|
|
644
|
+
<!-- size:start -->
|
|
645
|
+
|
|
646
|
+
| bundle | minified | gzip | brotli |
|
|
647
|
+
|---|---|---|---|
|
|
648
|
+
| samemark/core (parser + 6 helper packages, everything the fast path needs) | 87.4 KB | 31.2 KB | 27.0 KB |
|
|
649
|
+
| samemark/core for browsers (bundler uses the `browser` condition: entities decoded by the DOM) | 52.6 KB | 19.5 KB | 17.4 KB |
|
|
650
|
+
| samemark main entry, mdast-util-from-markdown left external (peer) | 88.4 KB | 31.7 KB | 27.5 KB |
|
|
651
|
+
| samemark main entry, everything bundled (includes mdast-util-from-markdown + micromark) | 144.0 KB | 47.1 KB | 40.9 KB |
|
|
652
|
+
| of which: decode-named-character-reference (HTML entity table) | 35.0 KB | 11.6 KB | 9.6 KB |
|
|
653
|
+
| baseline: mdast-util-from-markdown + micromark-extension-gfm + mdast-util-gfm | 114.6 KB | 34.6 KB | 29.7 KB |
|
|
654
|
+
| baseline: unified + remark-parse + remark-gfm | 143.6 KB | 44.1 KB | 37.8 KB |
|
|
655
|
+
|
|
656
|
+
The `browser` row is what a bundler produces for browsers: `decode-named-character-reference` then decodes entities through the DOM instead of shipping its 35 KB table. Edge runtimes (`workerd`, `worker`, `edge-light`, `deno`) get the table.
|
|
657
|
+
|
|
658
|
+
<!-- size:end -->
|
|
659
|
+
|
|
660
|
+
## Where it fits
|
|
661
|
+
|
|
662
|
+
Sätteri, the Rust engine that is the default markdown processor in Astro 7, says in its docs that it
|
|
663
|
+
["isn't a drop-in replacement for unified"](https://satteri.bruits.org/docs/) and that existing remark and rehype
|
|
664
|
+
plugins won't run unmodified. samemark is for the projects that stay on unified: the same mdast, existing plugins run
|
|
665
|
+
unchanged, and it is plain JavaScript with no native binary.
|
|
666
|
+
|
|
667
|
+
## Credits
|
|
668
|
+
|
|
669
|
+
- [micromark](https://github.com/micromark/micromark), [mdast-util-from-markdown](https://github.com/syntax-tree/mdast-util-from-markdown)
|
|
670
|
+
and the GFM extensions by Titus Wormer and contributors define the behavior this parser reproduces, and supply the
|
|
671
|
+
helper packages it imports (`micromark-util-*`, `ccount`, `decode-named-character-reference`).
|
|
672
|
+
- The [CommonMark](https://commonmark.org) spec and its examples.
|
|
673
|
+
- [unified](https://unifiedjs.com) and [remark](https://remark.js.org).
|
|
674
|
+
|
|
675
|
+
## License
|
|
676
|
+
|
|
677
|
+
MIT. Third-party licences of the code bundled in `dist/core.cjs` are in `dist/LICENSES.txt`.
|