mdq-cli 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md ADDED
@@ -0,0 +1,187 @@
1
+ # mdq
2
+
3
+ Query and edit markdown with a selector language — jq, for markdown.
4
+
5
+ ## Install
6
+
7
+ ```bash
8
+ npx mdq-cli 'h2' README.md # no install
9
+ npm install mdq-cli # as a library
10
+ npm install -g mdq-cli # as a command
11
+ ```
12
+
13
+ Node 18 or newer. Two dependencies: `marked` and `yaml`.
14
+
15
+ ## Use
16
+
17
+ ```js
18
+ import { mdq } from 'mdq-cli';
19
+
20
+ mdq(readme).query('section("Install") code[0]').text();
21
+ mdq(plan).comment(/^test/).nodes();
22
+ mdq(doc).table().addRow({ Method: 'POST', Path: '/sessions' }).toString();
23
+ ```
24
+
25
+ There is a CLI too:
26
+
27
+ ```bash
28
+ mdq 'section("API") table' --json README.md
29
+ ```
30
+
31
+ ## The one rule
32
+
33
+ **Reads narrow, writes return the document.**
34
+
35
+ `mdq(source)` gives a `MarkdownQuery` — one class, holding the document and the set of
36
+ blocks currently selected. `query()` and the sugar methods narrow that set. Every write
37
+ returns a fresh `MarkdownQuery` over the edited document, so edits chain and end with
38
+ `toString()`:
39
+
40
+ ```js
41
+ mdq(source)
42
+ .query('section("API")').append('## Notes\n')
43
+ .query('blockquote[0]').remove()
44
+ .toString();
45
+ ```
46
+
47
+ `toString()` is always the whole document; `text()` is the markdown of the current
48
+ selection.
49
+
50
+ ## Selectors
51
+
52
+ | Selector | Matches |
53
+ | --- | --- |
54
+ | `section` | a heading and everything under it, until the next heading of the same or shallower depth |
55
+ | `section1` … `section6` | the same, restricted to one heading depth |
56
+ | `heading`, `h1` … `h6` | the heading line alone |
57
+ | `paragraph` | a paragraph |
58
+ | `table` | a GFM table |
59
+ | `list` | a bullet or ordered list |
60
+ | `item` | one item of a list |
61
+ | `code` | a fenced code block |
62
+ | `blockquote` | a `>` block |
63
+ | `hr` | a thematic break |
64
+ | `html` | an HTML block, matched on its raw text |
65
+ | `comment` | an HTML comment, matched on its **inner** body |
66
+
67
+ Text matchers go in parentheses, and `!` negates any of them:
68
+
69
+ ```
70
+ section("Install") exact
71
+ section(~"Inst") contains
72
+ section(/^inst/i) regex, with its own flags
73
+ section(!~"Draft") negated
74
+ ```
75
+
76
+ Index and slice with brackets, and compose with spaces to scope one selector inside
77
+ another:
78
+
79
+ ```
80
+ heading[0] first
81
+ heading[-1] last
82
+ blockquote[2:5] a slice
83
+ section("API") table every table inside that section
84
+ ```
85
+
86
+ A leading `.` is accepted and ignored, so `.h2` works if that is your habit from jq.
87
+
88
+ An unknown selector throws `MdqSelectorError`, which carries the `index` of the offending
89
+ character. It never silently matches nothing.
90
+
91
+ ## Matchers as values
92
+
93
+ Passing a JavaScript value avoids escaping a dynamic string into a selector:
94
+
95
+ ```js
96
+ mdq(doc).query('section2', section.name); // exact
97
+ mdq(doc).heading(/^summary/i); // regex, own flags
98
+ mdq(doc).item((text) => text.length > 80); // predicate
99
+ ```
100
+
101
+ A `string` matches exactly, a `RegExp` honors its own flags, and a function is a predicate
102
+ over the node's text. Every sugar method takes one: `section` `heading` `paragraph` `table`
103
+ `list` `item` `code` `blockquote` `comment` `html` `hr`. Each is exactly
104
+ `query(selector, matcher)`; `section` and `heading` also take `{ depth }`.
105
+
106
+ ## Reading
107
+
108
+ | Method | Returns |
109
+ | --- | --- |
110
+ | `text()` | raw markdown of every match, joined |
111
+ | `nodes()` | `{ type, depth, text }` per match |
112
+ | `rows()` | table rows as objects, keyed by header |
113
+ | `entries()` | `Key: value` lines of a block, keys lowercased |
114
+ | `count()` / `exists()` | how many matched / whether any did |
115
+ | `first()` / `last()` / `at(n)` / `slice(from, to)` | narrow the selection |
116
+ | `each()` | one single-match query per match |
117
+ | `preceding()` / `following()` | everything before the first / after the last match |
118
+
119
+ ## Writing
120
+
121
+ Every one returns a `MarkdownQuery` over the edited document.
122
+
123
+ | Method | Effect |
124
+ | --- | --- |
125
+ | `replace(md)` | replace each match |
126
+ | `replaceEach(fn)` | replace each match with `fn(selection, index)` |
127
+ | `remove()` | delete each match, and its blank line |
128
+ | `insertBefore(md)` / `insertAfter(md)` | add a sibling block |
129
+ | `prepend(md)` / `append(md)` | add a block inside a section or list |
130
+ | `addRow(obj)` | append a table row, re-aligning the columns |
131
+ | `addItem(text)` | append a list item, copying the existing marker |
132
+ | `setEntry(key, value)` | set a `Key: value` line; `null` deletes it |
133
+
134
+ `prepend` and `append` need a section or list; on any other block they raise
135
+ `MdqOperationError`. Anything that takes markdown also takes a `MarkdownQuery`. Writes
136
+ never leave zero blank lines between blocks, and never more than one — including inside
137
+ fenced code blocks, which are left exactly as they are.
138
+
139
+ ## Frontmatter
140
+
141
+ A leading `---` block is parsed as YAML, kept out of the token index, and exposed as data.
142
+ Without this, `marked` reads `url: /login` as a setext heading.
143
+
144
+ ```js
145
+ const doc = mdq(page);
146
+ doc.frontmatter(); // { url: '/login', wait: 1000, tags: ['auth'] }
147
+ doc.query('h2').count(); // 0 — the --- block is not a heading
148
+ doc.setFrontmatter('wait', 2000); // comments and formatting survive
149
+ ```
150
+
151
+ ## Prior art
152
+
153
+ The selector grammar here is **bespoke**. It is not a standard, and there is no upstream
154
+ parser for it — element names and descendant-by-space come from CSS, `[2:5]` slices from
155
+ Python, and a tolerated leading `.` from jq.
156
+
157
+ The established alternative is the **unified/remark** stack: parse to
158
+ [mdast](https://github.com/syntax-tree/mdast), then select with
159
+ [`unist-util-select`](https://github.com/syntax-tree/unist-util-select), which implements
160
+ real CSS selectors on top of `css-selector-parser`. If you want a standards-based tool,
161
+ use that.
162
+
163
+ mdq exists because three things do not fall out of that stack:
164
+
165
+ - **Sections.** mdast is flat: a heading and the blocks beneath it are siblings, so the CSS
166
+ descendant combinator cannot express "this heading and everything under it until the next
167
+ heading of the same depth". [`remark-sectionize`](https://github.com/jake-low/remark-sectionize)
168
+ adds the nesting, but its synthetic `section` nodes carry no `position`, so their source
169
+ range has to be derived from their children before anything can be edited in place.
170
+ - **Editing by byte range.** mdq records each block's offset in the original source and
171
+ splices text, so anything it does not touch stays byte-identical. mdast nodes do carry
172
+ offsets, so this is achievable there too — it is a thing to build, not a thing you get.
173
+ - **Text and pattern matching.** mdast headings have no flat text field (the text is a child
174
+ node), CSS dropped `:contains()`, and `unist-util-select` parses the attribute `i` flag
175
+ but does not apply it — so `/^summary/i` has no selector form at all.
176
+
177
+ GFM tables are not in core remark either; they need `remark-gfm`.
178
+
179
+
180
+ ## Limitations
181
+
182
+ - **Block-level comments only.** A comment inside a paragraph (`text <!-- x --> more`) is
183
+ part of that paragraph's token and is not reachable as a `comment`.
184
+ - **No row or item selectors.** `addRow` and `addItem` append; there is no `removeRow`,
185
+ because there is nothing to select.
186
+ - **YAML frontmatter only.** TOML (`+++`) and JSON blocks are skipped from the token index
187
+ but not parsed.